Microsoft unveiled the Surface Laptop Ultra at a San Francisco event on October 7, built around Nvidia's new RTX Spark chip, with up to 128GB of unified memory and roughly one petaflop of AI compute. Unified memory means the CPU and GPU draw from one shared pool rather than separate reservations, which is what lets a laptop -- not a server rack -- hold a model above 120 billion parameters in memory at once. Reported pricing varies by outlet: two base configurations starting around $2,600 and $3,700, with higher-memory, higher-storage tiers reaching roughly $5,900. It ships October 16.
The exact numbers move slightly depending on which outlet's recap you read -- one lists the base price as $2,599 rather than $2,600, and a European listing gives 2,899 euros rather than a straight dollar conversion -- small enough gaps to read as currency and rounding differences between write-ups rather than a real dispute about what Microsoft charges. What's consistent across every report is the shape of the lineup: a cheaper entry tier that will not hold the model this story is actually about, and a top tier built specifically to run it. Buyers who want the on-device coding model need to check the memory spec line, not just the headline price.
The model Microsoft built to actually run on that hardware is MAI Code 1.1 Flash, a coding-focused mixture-of-experts model with 137 billion total parameters and 6.8 billion active at any step. The on-device version is quantized down to 53GB -- an 80% cut from the full-precision cloud variant -- using speculative decoding to stay responsive despite the compression. Peak memory use is 75.5GB at a full 256,000-token context window, which is why the 128GB configuration, not the base tier, is the one that actually matters for this use case. Microsoft's own numbers put the 3-bit compressed version at 70.80% on SWE-Bench Verified against the full cloud model's 72.6%, and ahead of it on Terminal-Bench 2.1, 66.29% to 62.9%.
MAI Code 1.1 Flash: on-device versus Microsoft's own cloud version
| On-device (Surface Laptop Ultra) | Cloud (full precision) | |
|---|---|---|
| SWE-Bench Verified | 70.80% | 72.6% |
| Terminal-Bench 2.1 | 66.29% | 62.9% |
| Model size | 53GB (3-bit quantized) | Bfloat16, ~5x larger |
| Inference cost | No metered charge | Billed per token |
Microsoft frames the point of running a 137-billion-parameter model on a laptop as giving developers a genuine choice, not a downgrade: its own post puts it as wanting developers to have "both choice and control" rather than "managing infrastructure decisions themselves." The same announcement made Microsoft Execution Containers, or MXC, generally available on Windows 11 -- an open-source tool, in Microsoft's own description, that "translates policy into native operating-system controls" so an organization can define exactly which files and networks an AI coding agent is allowed to touch, enforced at runtime rather than trusted on the agent's word.
How Microsoft Execution Containers fence off an agent
- Defines which files, folders, and networks an agent may reach, via Agent 365 or Intune policy
- Translates that policy into native OS sandboxing -- ProcessContainer on Windows, Seatbelt on macOS, bubblewrap on Linux
- Runs inside the container; any action outside the defined boundary is blocked, not merely logged
- Request denied at the OS level
Jensen Huang called MXC a technology that "is going to revolutionize how agents are built and deployed," according to The Globe and Mail -- a notably large claim for a sandboxing tool, and one worth reading as Nvidia's own interest talking as much as Microsoft's. Nvidia is chasing the Windows PC market specifically to compete with Intel and AMD on a battlefield neither currently owns outright, while Microsoft is betting that work now running in its own Azure data centers can move onto hardware the customer, not Microsoft, pays for. Apple is pursuing the same local-AI opportunity with new Macs -- but is tightening how much hard-drive access it grants AI agents even as Microsoft ships a tool to limit exactly that.
The hardware, at a glance
- Chip
- Nvidia RTX Spark
- Unified memory
- Up to 128GB
- AI compute
- ~1 petaflop
- Price range
- ~$2,600 to ~$5,900
- Ships
- October 16, 2026
The routing layer underneath this is called GitHub HydraFusion, and it's launching in experimental previews across the GitHub Copilot app, the Copilot CLI, and Visual Studio Code later in October -- its job is deciding, task by task, whether a request goes to the local model or up to the cloud, so a developer isn't the one manually flipping that switch. Windows ML, the lower-level runtime underneath it, is also gaining support for llama.cpp across GPU, NPU, and CPU, which matters past this one laptop: it means the same local-inference path Microsoft is building here isn't locked to Nvidia's chip specifically, even if RTX Spark is the hardware getting the launch spotlight. Microsoft's own framing of the economics is blunt -- local calls carry no metered inference charge at all, which is a real incentive to push work off the cloud wherever the hardware can handle it, separate from any privacy or latency argument.
What ships October 16 is still a first step, not the whole plan. DGX Station systems for Windows, rated for models up to a full trillion parameters, are due later this year -- a tier above anything a laptop will run. And the sandboxing story is already wider than coding tools: Meta's Muse, already rolling out as a personal checkout agent, is planned as a native Windows app built on MXC from the start, not retrofitted into it. The pitch for all of it is the same: an agent that can act on a user's own machine is only as trustworthy as the fence around it, and Microsoft is now selling both the agent and the fence.
- Microsoft's Surface Laptop Ultra, shipping October 16, runs AI models above 120 billion parameters locally.
- Nvidia's RTX Spark chip gives it up to 128GB of unified memory and roughly 1 petaflop of AI compute.
- Microsoft's on-device coding model scores within 2 points of its cloud twin on two benchmarks.
- Microsoft Execution Containers, now generally available, let IT limit exactly what files an agent touches.
- Caveat: reported pricing varies by outlet, from about $2,600 at the base to roughly $5,900 at the top.