GitHub Copilot is expanding local model support and introducing automated task routing by the end of October, alongside the general availability of operating-system-level tool sandboxing powered by Microsoft eXecution Containers across Windows, macOS, and Linux.
Coming by the end of the month, GitHub Copilot will determine when a task is best handled by on-device intelligence and when it should use cloud-scale models. Rather than forcing developers to manage infrastructure decisions themselves, GitHub Copilot automatically coordinates local and cloud inference behind the scenes.
Project HydraFusion and Automatic Task Routing in GitHub Copilot
The routing intelligence builds on Project HydraFusion, an orchestrator that evaluates task context and cache state to choose one or multiple models per task. Developers can let Auto routing handle the switching automatically or manually select a specific provider, model, or endpoint. Across a multi-turn session, Copilot can consider task context and cache state as it routes work between local and cloud models, preserving useful, cached work as the session evolves. This orchestrated experience complements direct model selection, giving developers a choice between letting Copilot optimize model placement and choosing a specific local model themselves.
Manual selection allows teams to hook in OpenAI-compatible local endpoints or choose MAI Code 1.1 Flash through the Windows ML provider. Yet the automation layer has drawn questions regarding data handling. Microsoft has not disclosed how much repository context or conversation history Auto sends to the cloud during automated routing, nor whether developers can inspect those routing decisions or restrict the agent strictly to local inference. Teams with strict data-handling policies still don’t know what repository data Copilot sends to the cloud.
“local inference does not make the session offline.”
Patrick Nikoletich and Stuart Schaefer, via Thenewstack
Hardware Demands and MAI Code 1.1 Flash Benchmarks
Running frontier-class coding intelligence on local machines requires substantial hardware. The initial rollout targets NVIDIA RTX Spark Windows PCs such as the Surface Laptop Ultra, which features up to 128GB of unified memory.
MAI Code 1.1 Flash is a mixture-of-experts model boasting 137 billion total parameters and 6.8 billion active parameters. To fit it onto client hardware, engineers applied mixed-precision quantization at approximately 3.3 bits per weight, reducing the footprint to 53GB—an 80% size reduction compared to the bfloat16 cloud variant. Developers will be able to select MAI Code 1.1 Flash through the Windows ML provider. The model will initially ship on the new Surface Laptop Ultra, which uses NVIDIA RTX Spark hardware to handle demanding local AI workloads.
| Benchmark | MAI Code 1.1 Flash (Full) | MAI Code 1.1 Flash (Quantized on Device) | GPT OSS 120B |
|---|---|---|---|
| SWE-Bench Verified (500 tasks) | 72.6% | 70.80% | 32.0% |
| Terminal-Bench 2.1 (89 tasks) | 62.9% | 66.29% | 23.6% |
Operating System Integration Through Microsoft Execution Containers
Alongside model execution, GitHub made local sandboxing for GitHub Copilot generally available on October 7, 2026.
- On Windows, the sandbox will use ProcessContainer’s BaseContainer tier.
- macOS: Relies on Seatbelt on macOS 15 Sequoia or later. Linux systems will use bubblewrap, while macOS will rely on Seatbelt.
When sandboxing is active, shell commands and local Model Context Protocol (MCP) and language servers run inside the process boundary. Built-in file tools check sandbox policies directly inside the agent harness on a best-effort basis, while remote MCP servers remain entirely outside the local process sandbox. Copilot CLI declares which paths are readable or writable and whether network access is allowed, and MXC applies that policy with the right backend for the machine.
Organizations Enforce Compliance Through Managed Settings and Sandboxing
Organizations can enforce compliance by deploying server-managed or MDM-managed settings that developers cannot weaken. Local sandboxing is included with GitHub Copilot at no additional cost, while cloud sandboxing is billed by usage. Sandbox policies apply to tool execution regardless of which model Copilot uses.
For specialized database workflows, similar boundaries apply. GitHub Copilot sessions do not retain history when switching contexts (for example, changing files or databases). The integration is optimized for modern SQL databases in Fabric, Azure SQL Database, and SQL Server 2017 (14.x) and later versions.