Today for AI
HOT RADAR
toolsHEAT 11.3°

AI CLUSTERED EVENT · 10/7/2026

Microsoft Adds Local Inference to GitHub Copilot with MAI Code 1.1 Flash MoE Model

1 reports archived1 independent sourcesupdated 10/7/2026, 10:17:08 PM
Synthesis & Latest Updates
1 Sources Cross-Validated

Microsoft announced that GitHub Copilot will support local AI model inference by the end of this month, enabling automatic or manual switching between cloud and on-device models. The update introduces MAI Code 1.1 Flash, a Mixture-of-Experts model optimized for local deployment (137B total parameters, 6B active), achieving up to 63 tokens per second decoding throughput on the Surface Laptop Ultra.

LATEST/Microsoft announced that GitHub Copilot will support local AI model inference by the end of this month, enabling automatic or manual switching between cloud and on-device models. The update introduces MAI Code 1.1 Flash, a Mixture-of-Experts model optimized for local deployment (137B total parameters, 6B active), achieving up to 63 tokens per second decoding throughput on the Surface Laptop Ultra.

TIMELINECoverage timeline

Total 1 reports · Latest first
  1. IT之家 · 智能时代T2·78 pts

    Microsoft Adds Local Inference to GitHub Copilot with MAI Code 1.1 Flash MoE Model

    Original: GitHub Copilot 不再纯云端 AI 模型:Surface Laptop Ultra 本地推理吞吐量最高每秒 63 词元

    • GitHub Copilot adds local inference support with options for automatic orchestration or forced on-device execution.
    • Introduction of MAI Code 1.1 Flash, a specialized MoE model (137B total/6B active) using quantization and speculative decoding.
    • Achieves 40-63 tokens per second decoding throughput on Surface Laptop Ultra, significantly optimizing memory usage.