Today for AI

InfoQ 中文 · 10/9/2026, 18:25:41

Microsoft Transforms Windows into Hybrid AI Platform: GitHub Copilot Integrates Local Models, Jensen Huang Highlights MXC as Agent Infrastructure

By 褚杏娟Original title: 黄仁勋为微软站台:没有 Windows 就不会有英伟达,Satya 当场“讨市值”
78AI Score
Executive Summary

Microsoft announced the transformation of Windows into a hybrid AI platform supporting local-cloud model coordination, with GitHub Copilot gaining the ability to call local models on PCs. Jensen Huang and Satya Nadella emphasized that as Coding Agents become new 'computer users,' OS-level isolation and governance are critical, positioning MXC as foundational infrastructure for next-generation agents.

SOURCE COVERAGEOriginal coverage

Contents8 sections

Microsoft is upgrading Windows into a hybrid AI platform, supporting coordinated scheduling of local and cloud models. GitHub Copilot integrates local model invocation capabilities for the first time.

Hybrid intelligence relies on model routing and Windows ML; Coding Agents are forcing operating systems to provide native support for isolation, permissions, and observability; MXC may become the underlying infrastructure for the next generation of Agents.

Suitable for AI platform engineers, OS kernel developers, and Agent architects.

Microsoft is transforming Windows into an AI platform that simultaneously invokes both local and cloud models.

At the Windows launch event on October 8 (Beijing Time), Microsoft announced the introduction of "hybrid intelligence" to Windows. Through model routing, Windows ML, and local AI capabilities, products like GitHub Copilot and Copilot can automatically switch between cloud and local models based on task complexity, cost, and privacy requirements.

Notably, GitHub Copilot's HydraFusion will be able to directly invoke local models running on Windows PCs for the first time. These capabilities will gradually roll out to the GitHub Copilot app, CLI, and VS Code later in October.

On stage, Jensen Huang and Satya Nadella reflected on the decades-long partnership between Microsoft and NVIDIA. Huang stated that NVIDIA’s development has been deeply intertwined with Windows: from Windows 95 and DirectX to GeForce and CUDA, they have jointly driven GPUs from graphics computing toward general-purpose accelerated computing.

The two also recalled that Azure’s early construction of high-performance computing (HPC) infrastructure based on InfiniBand later became a crucial foundation for OpenAI’s training of GPT models.

Regarding Agents, the two shared a consistent judgment: Agents are becoming the new "computer users." Nadella believes that Coding Agents were the first to reveal that even when models run in the cloud, the local Harness still needs access to the file system, execute commands, and call software. Therefore, the operating system must natively provide isolation, identity, observability, governance, and permission controls. Huang further emphasized that low-level mechanisms like MXC could become essential infrastructure for building and deploying the next generation of Agents, much like Windows and DirectX did in their respective eras.

Their understanding of "hybrid intelligence" is also very clear: In the future, users will no longer care whether a task runs locally or in the cloud; instead, they expect the system to automatically invoke the most suitable model and compute resources.

On the value of local AI, the two did not merely emphasize "saving tokens." Nadella believes it is more important to bring computation closer to data and tools. He even predicts that Agents will become one of the largest consumers of the file system in the future. Huang also believes that Agents will use software better than ordinary users.

At the hardware level, Huang argues that AI PCs should not just mean "adding an NPU," but must simultaneously host traditional applications, graphics computing, CUDA, local models, and Agents. He summarizes this shift as follows: In the past, the PC was a tool for humans; in the future, it will be both a tool for humans and a platform where personal Agents use tools.

On stage, Nadella joked with Huang: "Now Jensen is going to transfer some market cap back to me." However, Wall Street did not react immediately with intensity: Microsoft’s stock rose only about 0.1% on the day of the launch event and fell 1.35% on the second trading day.

Windows Is More Than Just an AI PC

Local Support for Larger-Scale Models

Microsoft believes that models previously runnable only in data centers are being compressed onto personal computers. In the future, Windows can handle high-frequency, time-consuming, or privacy-sensitive data processing locally, while delegating truly complex tasks to cutting-edge cloud models. Microsoft states that this will become a universal architecture for Windows AI, applicable to scenarios such as programming, security, content creation, and office work.

To bring larger models to the edge, Microsoft is focusing on model quantization and Windows ML.

MAI Code 1.1 Flash, previously released by Microsoft, is a Mixture-of-Experts (MoE) coding model with 137 billion total parameters and 6.8 billion active parameters. The version designed for local deployment, after approximately 3-bit quantization, reduces the model size by nearly 80% while retaining a 256K context window.

Microsoft also stated that NVIDIA is developing a Nemotron model with over 70 billion parameters. After quantizing to 2 bits, its memory footprint is slightly above 20GB. Meanwhile, DeepSeek V4 Flash has a parameter scale of 284 billion; after 1.66-bit quantization, the model can run locally with a memory footprint of approximately 60GB.

Complementing model quantization is Windows ML. Microsoft is integrating Llama.cpp into Windows ML, enabling open-source models to connect to Windows faster and run across GPU, NPU, and CPU.

"Hybrid Intelligence" in Copilot

In GitHub Copilot, "hybrid intelligence" first manifests as automatic model routing.

The new auto-mode determines whether to invoke local or cloud models based on task complexity. Simple tasks, such as code cleanup or issue triage, can be handled directly by local models; more complex tasks continue to use cloud models.

Microsoft demonstrated a typical scenario on stage: The main task was analyzed by the cloud-based GPT-6 Luna, while three local sub-Agents were launched using MAI Code 1.1 Flash Local to build and test in parallel. One local session input approximately 1.66 million tokens and output approximately 10,500 tokens, without generating additional cloud model costs. Microsoft aims to move such high-frequency, long-running tasks to local devices as much as possible, especially for jobs like issue triage and automated checks that can run overnight.

This effectively changes the cost structure of Coding Agents. Previously, the longer an Agent worked and the more parallel instances it ran, the higher the token costs often became. If some sub-tasks can be shifted to the user’s own GPU and NPU, local compute power can serve as a substitute for cloud tokens.

Compared to GitHub Copilot, the changes in Windows Copilot are closer to those of a system-level Agent.

Microsoft has added three categories of capabilities to Copilot: local context, local actions, and local models. Local context enables Copilot to search and understand files on the PC; local actions allow it to move files, modify settings, and perform other tasks once authorized; and local models let Copilot offload certain workloads directly to the device when cost or privacy are paramount, while still delegating the most complex tasks to cloud-based models.

These capabilities have first been integrated into Copilot Home and Code. Home can now understand local files on the user's PC and bring content previously stored only on the device directly into Copilot for editing and collaboration.

Copilot Code further extends "local AI" into software development. Users need only enter a prompt to generate native Windows apps and widgets without any development experience. Microsoft has also added MXC support to Code, placing locally executed code in a sandboxed environment. Meanwhile, some programming tasks can be handled directly by local models to reduce cloud invocation costs and improve response times.

Microsoft calls this suite of capabilities the "Personal Software Factory": from app generation and code execution to partial model inference, everything can happen directly on the local PC, while more complex tasks continue to be handled by the cloud.

At the event, Microsoft demonstrated this capability using tax filing as an example. Autopilot could first read emails from an accountant to determine which documents were needed, then locate materials across different directories on the user's computer, rename and organize the files, and finally generate a ZIP package. The organization and renaming of files containing tax information were performed entirely by local models, ensuring data did not need to be uploaded to the cloud. After organizing the files, Autopilot could also draft a reply email with the compressed archive attached, but user confirmation was still required before actually sending it.

Microsoft defines Autopilot as an enterprise-grade, long-running autonomous agent. The underlying direction is to ensure that Windows does not merely run AI applications, but allows AI to directly understand the local environment and execute tasks at the operating system level.

Microsoft has not restricted these capabilities to Copilot. Currently, agents such as Hermes, OpenClaw, and Perplexity have begun building local AI experiences based on Windows, and Muse will also launch a native Windows app. For persistent agents like OpenClaw, Microsoft will provide new Windows quick-install experiences and a native Windows Gateway to lower deployment barriers.

Microsoft refers to mini desktop PCs as "Claw Boxes," suitable for long-running agents: devices can remain always-on, allowing agents to continue working even when users step away from their computers. This represents a notable shift in Microsoft's current Windows AI strategy: Windows is no longer just emphasizing the "AI PC," but is beginning to compete for the foundational runtime environment for agents.

To support larger local models, Microsoft has simultaneously expanded the hardware coverage of AI PCs.

Microsoft revealed that over 40% of laptops currently targeted at the commercial market are already Copilot+ PCs; all Copilot+ PCs complete more than 2 trillion inferences locally each month.

For developers and creators, Microsoft also launched the Surface Laptop Ultra. Based on NVIDIA RTX Spark, this device features up to 128GB of unified memory and delivers up to 1 PFLOP of AI compute power. One of its goals is to enable large models that traditional laptops cannot accommodate to run directly on the local machine.

The next tier up is the DGX Station running Windows. Microsoft states that these desktop-class supercomputers can locally run Llama 4 Maverick, Kimi K2.6, and DeepSeek V4 Pro, which has over 1 trillion parameters.

Microsoft summarizes its hardware roadmap in one sentence: With RTX Spark, you can bring "last year's frontier" to your laptop; with DGX Station, you can bring "last quarter's frontier" to your desk.

Jensen Huang and Satya Nadella: The First Reimagining of the Personal Computer

Moderator: I looked up both of you on YouTube last night and realized we rarely see you two standing on the same stage together. I was thinking about how the history between these two companies spans decades, and your personal collaboration also has decades of accumulation. Jensen, what is your favorite memory or most interesting experience when collaborating with Satya and Microsoft?

Jensen Huang: The founding of NVIDIA was essentially driven by Windows. It was late 1992, during the era of Windows 3.1. We were just starting to envision what a "personal computer could become." Honestly, before that, I hadn't seen many PCs; I primarily worked on workstations. We had file servers and supercomputers used for chip design.

At the time, my co-founders and I were discussing a future where a personal computer with 3D graphics capabilities could be used for gaming, serve as a workstation, and handle design tasks. Our imagination started with this vision of the world. Without Windows, NVIDIA would not have been founded in the way it is today.

Then, in 1993, we officially began operations, and Windows 95 truly redefined this landscape. Windows 95 allowed GPUs to genuinely connect to personal computers, and DirectX was what truly revolutionized the GPU. The Direct3D API allowed us to expose GPU capabilities to applications for the first time.

Subsequently, DirectX 8 emerged, and together with Microsoft, we invented the world's first programmable GPU, known as programmable shaders. A few steps later, this evolved into CUDA, which everyone knows today.

So, looking back at NVIDIA's entire journey, Windows has always been at the core. Without Windows, there would be no GeForce; without GeForce, there would be no CUDA; and without CUDA, Alex Krizhevsky, Ilya Sutskever, Yann LeCun, Andrew Ng, and others would not have found a computer suitable for deep learning. And then we arrived at today.

My earliest memory of this was us discussing: how far can accelerated computing or HPC go in the cloud? I recall one of our very first collaborations involved figuring out how to scale HPC on Azure.

Jensen Huang: And Satya was the first person to use InfiniBand in the cloud, which was a critical move. It allowed us to bring supercomputers to OpenAI for the first time, and the rest is history as everyone knows.

If we hadn’t done those things for HPC back then—if we hadn’t built InfiniBand-based supercomputers—OpenAI wouldn’t have had that supercomputer to train the GPT models later on, and history might have unfolded quite differently. Many things in life happen by chance like this: we were originally working on HPC, then Sam came over and asked, “Do you have compute?” We said, “Maybe.” And then everything took off.

Satya Nadella: What I really want to credit Jensen for is his long-term, consistent judgment on this matter. He wasn’t just looking at today, nor merely tomorrow; he viewed it as a long-term structural trend. That’s why we’ve arrived at this moment.

Jensen Huang: Then we started discussing another topic. I remember clearly when Satya and I were talking about what this means for PCs. In the new era following all of this, how do we build the “perfect PC”?

The personal computer has always been the ultimate tool—for me, and for an entire generation. So, what happens in the Agent era? When Agents are right there on your computer, they become your personal assistants, but they also need access to all the powerful tools we already possess.

The problem was that these tools were previously scattered across different devices within the NVIDIA ecosystem: gaming PCs, workstations. A vast array of professional tools resided there, including various design tools, Siemens’ tools, Autodesk, Adobe, and the applications everyone just saw. We wanted the world’s best Blender, the best Unreal Engine, the best Omniverse, and the best CATIA—all appearing on the same PC alongside AI.

So, how do we create such a computer? By the way, we spent four years doing this.

Why Agents Need OS-Level Security Primitives

Moderator: Satya, people have been discussing Agents and security for the past few months. Many here are using Agents like Muse or Instinct, granting them full disk access to act on their behalf. The MXC demonstrated today feels like adding a very fundamental new operating system primitive for Agents, which is quite different from before. Why is the operating system the right place to solve this problem? And what new space does this primitive open up?

Satya Nadella: First, remember that all of this started with coding Agents running on desktops. If you ask which type of Agent most needs this kind of security primitive, the answer is actually the local Harness, because it needs to access the file system and take real actions.

So, when we collaborated with Jensen’s team, we discussed how to integrate this directly into the underlying primitives, whether in the cloud or the local operating system. Because coding reached this stage first: even if you call cloud models, the Harness still runs locally. Therefore, we must make the desktop the safest execution environment for Agents. This is where MXC comes from.

Now, seeing OpenShell fully integrated with MXC and achieving native support is fantastic. The control capabilities Pavan mentioned earlier essentially highlight one thing: isolation must be implemented at every layer. You can’t just do it at one layer and ignore others. Micro-VMs are needed, VMs are needed, and session-level isolation is needed too.

What excites me most is that Windows sessions themselves now have isolation capabilities, which is crucial. But isolation alone isn’t enough; next, we need complete observability.

How do we achieve observability? You must give the Agent an identity so that I can track it, see the complete execution trail left by this Agent, and then govern it accordingly. This means setting policies like traditional systems do, constraining behavior based on those policies, which ultimately gives IT operations and security operations confidence. Another point we cannot forget is tracking token spend: exactly where and how many tokens are being used.

So, what we’re trying to do isn’t just any single thing, but stringing all these elements together. Having gone through so many technology adoption cycles, whether consumer or enterprise, I believe what often hinders large-scale diffusion of technology is failing to think through “what is fully required.” If you only complete part of it and leave the rest undone, trust cannot be established.

Now, we have connected these capabilities across the cloud, the client, and the entire security and governance framework. Moreover, it is heterogeneous—it doesn’t just support Microsoft’s own specific technologies, but aims to support everything.

Jensen Huang: What Satya just described will become the foundation of the next generation of IT. If you look closely at the words he used, he was actually describing the deep architecture of the next-generation computer.

Just as Windows and DirectX once fundamentally changed how applications were built, MXC will change how Agents are built and deployed. Without it, none of this is possible, so this is a very big deal.

What Does “Hybrid Intelligence” Mean?

Moderator: This ties closely to the next topic: “hybrid intelligence.” I feel it differs somewhat from what we’re used to. Everyone has used large cloud models: you send a request, it might call a tool like Blender, and return the result. We’ve grown accustomed to this experience.

But the scenarios shown today are different. You can choose many models, some running locally, some in the cloud, and tasks flow back and forth between the two. One line from the GitHub demo earlier struck me deeply: “These tokens are free.” This means part of the workload executes directly on the laptop locally, part in the cloud, and the system intelligently routes them. What does this model unlock?

Satya Nadella: I think this is one of those moments where many disparate pieces suddenly click together.

But it wasn't until around the Windows 95 era, and even with the emergence of first-generation browsers like Mosaic, that these elements were truly integrated into a seamless experience. At that point, you stopped caring about "local versus cloud"; you just wanted all available compute power to flow smoothly toward whatever you were doing.

Danilo’s demo earlier captured exactly that feeling. It was genuinely striking. I think in a few years, people will look back and say, "That was the first time we really saw this way of working."

He started in ComfyUI, downloading multiple models and generating a batch of assets; then took those assets into Blender for further refinement; went back to ComfyUI; moved to Photoshop to continue optimizing using capabilities from what might be local models; returned to ComfyUI again; and finally brought the entire workflow to the cloud, completing the last step via Astra and Unreal.

He simply used desktop applications, local models, and cloud models naturally, allowing everything to work together seamlessly. That is precisely what we want to convey today.

So the trend I see is very clear. In the future, there won’t be a moment where you have to deliberately decide, "This runs locally, that runs in the cloud." You will assume hybrid intelligence is ubiquitous, becoming an ever-present capability. This is what excites me the most.

Jensen Huang: In that application just now, Blender was accelerated, ComfyUI was accelerated, Astra ran in the cloud, and the Agent ran locally. The reach of the total compute involved is staggering.

Satya Nadella: Right. It essentially converged the capabilities of nearly every tech company in human computing history into a single application workflow.

Jensen Huang: Exactly. Returning to your point about accelerated computing, this is why Danilo is so excited: tasks that used to take days now take hours. What’s being accelerated isn’t just a specific task, but his creative ambition. I think that’s cool.

Moderator: Models are already quite good at tool calling, but many tools themselves reside locally, as do files. There’s also a very practical issue here: I want my data to remain secure and under control, and I want tool calls to happen near the data.

Satya Nadella: Yes. When coding Agents first became popular, one thing that excited me was how people suddenly fell in love with the console and the file system again. This once again proves what a powerful abstraction the file system is.

I believe Agents will become one of the largest user bases for the file system. The file system must be secure because it is a heterogeneous collection space. The tax filing example mentioned earlier was very typical: it aggregated content from various sources—download folders, cloud files, emails—into a single file system container, then invoked local and cloud models to continue processing.

This is the new computer. In the past, I used the computer; in the future, my Agent and I use the computer together. Therefore, Agents need access to these underlying primitives.

Take OneDrive, for instance. I haven’t thought of it as "cloud stuff" or "client stuff" for a long time; it’s just everywhere. So, I think we are entering a phase where I want the underlying capabilities serving Agents to exist 24×7. When I sit at my PC, it should feel like it’s working on that PC; when I move elsewhere, it should follow me.

Jensen Huang: Another amazing aspect is that these Agents will ultimately master tools better than we do. The reason is simple: we typically know only 10% to 15% of an application’s features, but an Agent can know every single feature.

So now, these incredible tools are on your PC, running on powerful machines like the Surface Laptop Ultra, allowing Agents direct access to all capabilities. Who knows what the future holds? Just now, it simultaneously utilized Blender, Unreal, and ComfyUI—it’s truly unbelievable.

Jensen Huang: The First True "Reinvention" of the Personal Computer

Moderator: Jensen, I’d like to ask about hardware next. We often hear you talk about massive-scale systems like Vera Rubin, Blackwell, and NVL72—the "Intelligence Token Factories" driving data centers. But today we’re discussing something different: RTX Spark, workstations, etc. How do you view the shift where Tokens are produced not just in data centers, but increasingly near users?

Jensen Huang: The engineering effort invested in this project has been immense, totaling over 4,000 engineer-years. It’s a massive undertaking.

We have actually reinvented the personal computer completely for the first time. Previously, capabilities were scattered across gaming PCs, laptops, workstations, and cloud supercomputers. Now, we’ve unified the architecture, running everything on a single superchip.

Why is this superchip rectangular? Because it’s essentially two giant chips fused together. It’s a "beast." It delivers 1 PFLOP of compute.

For reference: the DGX-1 that ignited the AI revolution also delivered 1 PFLOP. That machine cost 250million(Note:TheofficiallaunchpriceoftheoriginalDGX−1in2016was250 million (Note: The official launch price of the original DGX-1 in 2016 was 129,000, with a system weight of approximately 134 lbs) and weighed 500 pounds. Now, that level of compute fits inside your laptop.

More importantly, this is the only computer architecture in the world capable of natively running all these things simultaneously: DirectX and its entire ecosystem, various versions of OpenGL—because we need all these tools for design and creative software; Unreal Engine must run exceptionally well; DirectX needs to be top-tier; and every CUDA application must run at high performance.

Four years ago, Satya and I discussed that a whole new generation of software developers would emerge, becoming AI software developers. They would either be developing AI applications and models or leveraging AI to develop software. Therefore, the computer itself had to be redesigned to be fundamentally proficient in CUDA from the ground up.

We are fully aware of the significance of what we are doing. We have already tested over 1,200 applications—some of the most globally demanding and complex workloads in existence. We have conducted comprehensive functional testing, performance benchmarking, and compatibility validation, alongside extensive optimization efforts. This work is far from finished; we will continue to integrate more software, middleware, and algorithms into this system indefinitely.

So when you finally see Pavan on stage holding such a sleek laptop, it represents the collective effort of countless engineers from Microsoft, NVIDIA, and partners worldwide. All of this stems from our belief: while personal computers have historically been tools for humans, in the future they will not only remain tools but also serve as our personal assistants—assistants that use tools on our behalf.

Who bears responsibility when an Agent acts on behalf of the user?

Moderator: One final question for both of you. If you’ve been following discussions on X recently, you’ll notice intense debate around "Personal Agents." Today, we mostly talk about productivity agents and enterprise agents, but with systems like Muse and Instinct, a critical question has emerged: When an Agent goes out to make a purchase on my behalf, is it me making the purchase, or is it the Agent? And if something goes wrong, who is liable?

For service providers, how do we respond if users no longer directly access our websites, but instead send Agents to interact on their behalf? The industry is currently attempting to establish standards and explore new behavioral norms. Looking slightly into the future, where do you think this is heading? Will standards emerge? Will the industry converge around certain new behavioral protocols?

Satya: I believe we need to return to the concept of "trust" we discussed earlier. Trust begins with fundamental questions: Are these Agents isolated? Do they have distinct identities? Can the Agent’s execution trace be separated from the user’s identity?

A key point Pavan’s team consistently emphasizes is that even when an Agent works on behalf of a user, the Agent ID must remain separate from the user’s own identity. This separation is necessary for daily usage, but it becomes even more critical in the scenario you described. Because ultimately, you must decide: How much delegation authority am I willing to grant the Agent to act on my behalf?

This decision is fundamentally rooted in trust. As Agents become increasingly autonomous, the amount of authority you are willing to delegate will grow incrementally, provided the Agent continuously demonstrates its reliability.

In consumer scenarios, I think this eventually boils down to a form of "insurance pricing." In a sense, the market will price the risk of "Agents acting on my behalf." Protocols are certainly important—we need to get them right—but ultimately, the market will assess how to price the risks associated with Agents taking actions for users.

Jensen Huang: Your question once again highlights how crucial MXC is as foundational infrastructure. OpenShell, along with the collaborative work we are doing, essentially addresses issues of trust, isolation, observability, and monitoring.

If these foundations are not solid, nothing else can proceed. We must first ensure this security layer is robust. Agents that act on our behalf and perform tasks for us should only receive least-privilege access; they must possess their own credentials; and they must be monitorable to build trust. Only by achieving these conditions can we truly empower Agents to execute tasks autonomously on our behalf.

Today marks a particularly powerful moment. I remember when OpenClaw first appeared—it was essentially something hacked together from the terminal, with a self-built Gateway. Now, it has become natively embedded within Windows, with Tokens accelerated locally. The sheer volume of engineering work required to reach this point is truly staggering.

Satya: I’d like to add one final thought. The vision Jensen just mentioned inadvertently touches on what Windows has always aimed to do: not leaving any code or framework behind in the past, but bringing everything forward into the future.

He noted that we are bringing all Windows applications, all graphical apps, all CUDA, all OpenGL, plus the new Token factories, into this new era. That holistic approach is what truly impresses me.

Windows has accompanied me for 34 years, and it looks like it will accompany me for another 34.

Reference link:

https://www.youtube.com/watch?v=RE_AsVyoOSQ