Developers

Copilot's Hybrid Inference Plan Leaves Data Transparency Questions Unanswered

Microsoft is preparing to automatically route GitHub Copilot tasks between local and cloud models by late October, but has declined to disclose what information travels to the cloud or how developers can control the process.

5 min read
GitHub Copilot is going local — but Microsoft won’t say what gets sent to the cloud

Starting by the end of October, GitHub Copilot will intelligently distribute coding work across local and cloud-based models, though Microsoft has not clarified what data leaves developers' machines or provided ways to prevent it. The company revealed this strategy Wednesday through a joint statement by Patrick Nikoletich, a product manager at GitHub, and Stuart Schaefer, a Windows platform partner architect.

The timing of this announcement coincided with GitHub's release of new sandboxing controls to general availability, though the level of protection varies based on which Copilot components are in use. Operating system-level restrictions protect shell commands and local MCP servers, while file tools built into the system rely on validation within the agent harness itself. Remote MCP servers, however, remain unprotected by the local process sandbox.

Nikoletich and Schaefer note that "local inference does not make the session offline." The company has not disclosed the volume of repository context that Auto routing transmits to cloud models, whether developers can review routing choices, or if Auto can be limited to local-only inference.

Copilot decides where inference runs

GitHub is building upon Project HydraFusion, which currently picks models for development tasks, to determine where those models execute. Copilot will evaluate task context and cache status when deciding between local and cloud inference, including across multi-turn conversations, according to Microsoft.

Within Copilot CLI, the Copilot app, and VS Code, developers have the option to enable Auto routing or manually choose a local model. Available choices include MAI Code 1.1 Flash via the Windows ML provider and OpenAI-compatible local endpoints.

Auto routing raises questions

Microsoft has not specified how much dialogue history or repository context Auto sends to cloud services when routing a task. The company also has not indicated whether developers can observe these routing decisions or confine inference to local models only. Organizations with strict repository data policies remain uncertain about what code and context Copilot transmits to the cloud. Comparable concerns surfaced last month when Anthropic announced it could redirect Claude Sonnet 5.5 requests to Sonnet 5 upon detecting elevated-risk scenarios.

Teams with strict data-handling policies still don't know what repository data Copilot sends to the cloud.

Choosing a local model ensures inference stays on the device, but this does not prevent the agent from contacting external services or initiating network calls through its integrated tools. Developers requiring a completely isolated session must also configure restrictions on what those tools can reach.

MAI Code 1.1 Flash is a mixture-of-experts model containing 137 billion total parameters with 6.8 billion active, and Microsoft applied mixed-precision quantization at approximately 3.3 bits per weight to reduce it to 53GB, representing an 80% decrease from the bfloat16 cloud version.

The need to fit models onto constrained devices has driven comparable work in other parts of the industry, such as Intel's efforts to compress a 1.58-bit LLM even further. Microsoft combined quantization with speculative decoding, in which a drafter generates token sequences for the main model to validate, to enhance the speed of local inference.

Microsoft paired quantization with speculative decoding, in which a drafter proposes blocks of tokens for the main model to verify, to speed up local inference.

Quantization meets memory limits

The initial deployment targets NVIDIA RTX Spark Windows machines including Surface Laptop Ultra, which can have up to 128GB of unified memory. On such a device, Microsoft measured peak memory consumption of 75.5GB when using a 256K-token context, a threshold that excludes most developer machines equipped with 16GB or 32GB of RAM.

The 53GB model weights represent only part of the memory demand. The operating system, running applications, inference runtime, and key-value cache all require space, and the cache expands as the agent processes files and receives tool outputs, meaning developers working through extended sessions must account for memory usage well beyond the model size alone.

Benchmarks with fine print

Microsoft reports the quantized model achieved 70.8% on SWE-Bench Verified, compared with 72.6% for the full-precision version, and outperformed the original on Terminal-Bench 2.1 with 66.29% versus 62.9% across a set of 89 tasks. On a dataset of that size, the difference represents three tasks. The findings indicate the company reduced model size while preserving coding capability, but do not demonstrate that quantization improved performance.

Copilot enforces sandbox policies through Microsoft's open source Execution Containers (MXC) library, using the BaseContainer tier of the ProcessContainer backend on Windows, Seatbelt on macOS, and bubblewrap on Linux.

When sandboxing is active, Copilot applies operating system-enforced restrictions to shell commands and, where available, local MCP and language servers. GitHub states these restrictions take effect regardless of whether a task executes on a local or cloud model.

Built-in tools, harness checks

Copilot's file tools operate within the agent process, where the harness validates requests against sandbox policy rather than depending on OS-level isolation. Remote MCP servers also exist outside the local sandbox, with Copilot checking their connection policies when MCP sandbox controls are active.

Remote MCP servers also sit outside the local sandbox, with Copilot checking their connection policies when MCP sandbox controls are enabled.

Microsoft's demonstration uses MAI Code 1.1 Flash to construct a daily triage dashboard in a sandboxed copilot-sdk project, which the company characterizes as an offline workflow. However, the prompt retrieves GitHub issue and pull request metadata rather than the local repositories and tests mentioned earlier in the announcement. Microsoft does not clarify whether that metadata was fetched over the network or stored locally, leaving the offline claim unsubstantiated.

Source: The New Stack · Reporting supplemented by The Silicon Ledger staff.