IT之家 · 智能时代 · 10/7/2026, 10:17:08 PM
Microsoft Adds Local Inference to GitHub Copilot with MAI Code 1.1 Flash MoE Model
Microsoft announced that GitHub Copilot will support local AI model inference by the end of this month, enabling automatic or manual switching between cloud and on-device models. The update introduces MAI Code 1.1 Flash, a Mixture-of-Experts model optimized for local deployment (137B total parameters, 6B active), achieving up to 63 tokens per second decoding throughput on the Surface Laptop Ultra.
SOURCE COVERAGEOriginal coverage
GitHub Copilot Is No Longer a Pure Cloud-Based AI Model: Surface Laptop Ultra Achieves Local Inference Throughput of Up to 63 Tokens Per Second
On October 8, during a launch event held in San Francisco at 10 AM Pacific Time (1 AM Beijing Time on October 9), Microsoft announced that GitHub Copilot will support local AI model inference by the end of this month. Developers can automatically or manually switch between cloud-based models and on-device models.
GitHub Copilot is an AI programming assistant developed by Microsoft. It integrates with tools such as Visual Studio Code and CLI to generate code completions, explanations, or refactoring suggestions based on code context.

Currently, GitHub Copilot relies on cloud-hosted models, where a coordinator routes requests based on performance, cost, and accuracy. Microsoft plans to expand this mechanism by the end of the month, allowing Copilot to automatically select between cloud models and local AI models while orchestrating inference tasks in the background.

Microsoft offers two modes for GitHub Copilot users: automatic orchestration or forced use of on-device models. Users can set their preferences within GitHub Copilot CLI, the Copilot app, and Visual Studio Code.

Local model selection supports specifying providers, models, or endpoints. Developers can choose MAI Code 1.1 Flash via Windows ML, or connect to OpenAI-compatible local endpoints and select from the models they expose.
Microsoft has introduced the MAI Code 1.1 Flash "Mixture-of-Experts" (MoE) model, which features 137 billion total parameters and 6.8 billion active parameters. The model utilizes quantization and speculative decoding techniques to optimize response speed and memory footprint.
Real-world testing on the Surface Laptop Ultra shows significant improvements in peak memory usage, token utilization, and processing speed for complex tasks. Decoding throughput tests indicate that when prompt lengths range from 2K to 256K tokens, throughput varies between 40 and 63 tokens per second. IT Home provides related images below:
