IT之家 · 智能时代T2·78 pts
Microsoft Adds Local Inference to GitHub Copilot with MAI Code 1.1 Flash MoE Model
Original: GitHub Copilot 不再纯云端 AI 模型:Surface Laptop Ultra 本地推理吞吐量最高每秒 63 词元
- GitHub Copilot adds local inference support with options for automatic orchestration or forced on-device execution.
- Introduction of MAI Code 1.1 Flash, a specialized MoE model (137B total/6B active) using quantization and speculative decoding.
- Achieves 40-63 tokens per second decoding throughput on Surface Laptop Ultra, significantly optimizing memory usage.