GitHub Copilot will soon automatically decide whether coding tasks run locally on a developer's machine or get sent to cloud-based models, with the routing feature expected to launch by the end of October. Microsoft detailed the plan Wednesday in a post co-written by Patrick Nikoletich, a GitHub product manager, and Stuart Schaefer, a Windows platform partner architect. The announcement came alongside GitHub making new sandboxing controls generally available, though the protections vary depending on which tools Copilot uses.
GitHub is expanding Project HydraFusion, which already selects models for coding tasks, to also determine where those models execute. Microsoft says Copilot will weigh task context and cache state when switching between local and cloud inference, including during multi-turn sessions. In Copilot CLI, the Copilot app, and VS Code, developers can use Auto routing or manually select a local model themselves, with options including MAI Code 1.1 Flash through the Windows ML provider and OpenAI-compatible local endpoints. The initial rollout targets NVIDIA RTX Spark Windows PCs such as Surface Laptop Ultra, which offers up to 128GB of unified memory. On that machine, Microsoft measured peak memory use of 75.5GB at a 256K-token context, a figure that rules out most developer laptops with 16GB or 32GB of RAM. MAI Code 1.1 Flash is a mixture-of-experts model with 137 billion total parameters and 6.8 billion active, and Microsoft used mixed-precision quantization at roughly 3.3 bits per weight to shrink it to 53GB, an 80% reduction from the bfloat16 cloud version. Microsoft reports that the quantized model scored 70.8% on SWE-Bench Verified, compared with 72.6% for the full-precision version, and that it outperformed the original on Terminal-Bench 2.1 with 66.29% against 62.9% on a dataset of 89 tasks.
Nikoletich and Schaefer acknowledge that "local inference does not make the session offline." Microsoft hasn't disclosed how much repository context Auto sends to cloud models, whether developers can inspect routing decisions, or whether Auto can be restricted to local inference. Teams with strict data-handling policies still don't know what repository data Copilot sends to the cloud. Selecting a local model keeps inference on the device, but it doesn't stop the agent from reaching external services or making network requests through its tools, meaning developers who need a fully local session will also have to lock down what those tools can access. Shell commands and local MCP servers receive OS-level restrictions, while built-in file tools rely on checks inside the agent harness, and remote MCP servers remain outside the local process sandbox.
The 53GB of weights is only part of the memory bill. The operating system, applications, inference runtime, and key-value cache all need room, and the cache keeps growing as the agent reads files and receives tool results, so developers running longer sessions will need to budget for memory well beyond the model itself. Microsoft paired quantization with speculative decoding, in which a drafter proposes blocks of tokens for the main model to verify, to speed up local inference. The pressure to squeeze models onto smaller hardware has pushed similar efforts elsewhere, including Intel's work to compress a 1.58-bit LLM even further. The results suggest the company shrank the model without sacrificing much coding performance, but fall well short of showing that quantization made it better—on a benchmark that small, the gap amounts to three tasks. Organizations eager to keep code in-house will face a choice between accepting opaque routing logic or investing in the hardware and workflow controls needed to guarantee truly air-gapped development. The shift also signals that model size optimization has become table stakes for any vendor hoping to capture enterprises reluctant to send proprietary code to the cloud.

