Microsoft and NVIDIA launched the Surface Laptop Ultra starting at $2,599 on October 7, 2026, pairing custom NVIDIA RTX Spark silicon with local agentic AI tooling and Microsoft Execution Containers in Windows 11 to bring cloud-scale processing directly to desktop users.
Microsoft and NVIDIA Reveal Surface Laptop Ultra in San Francisco
At an event in San Francisco on Wednesday, executives from Microsoft and NVIDIA pulled back the curtain on a hardware and software push aimed at moving artificial intelligence workloads out of remote server farms and straight onto the user’s desk. Pre-orders opened immediately, with general availability scheduled for October 16. NVIDIA CEO Jensen Huang connected that history and his long-term vision to the agentic era during a fireside chat with Microsoft CEO Satya Nadella, noting that NVIDIA was founded because of Windows.
Equipped with a custom NVIDIA RTX Spark system-on-a-chip, the flagship laptop is priced from $2,599 upwards and includes a choice of SoC and memory setups, featuring either a 5120-core or 6144-core Blackwell GPU alongside up to 128GB of LPDDR5x unified memory.
Microsoft Details Spark Dev Box Specs and Display Architecture
Alongside the consumer laptop, Pavan Davuluri also unveiled the Surface RTX Spark Dev Box, a $5,999 anodized aluminum black box designed to meet the needs of frontier developers. Tailored specifically for AI development and local AI use, the Spark Dev Box comes equipped with a custom developer-optimized edition of Windows 11. Both machines target heavy workloads by delivering up to a petaflop of AI compute performance. The mobile unit offers unified memory allocation rather than physical VRAM limits.
| Hardware Configuration | Starting Price | Key Specifications |
|---|---|---|
| Surface Laptop Ultra | $2,599 | Up to 128GB unified memory, Blackwell RTX GPU (5,120 or 6,144 cores), HDR display, various SoC options |
| Surface RTX Spark Dev Box | $5,999 | Anodized aluminum desktop chassis, pre-configured developer build of Windows 11, 1 petaflop AI compute capacity |
Today’s event also offered an early look at the NVIDIA DGX Station for Windows, marking the debut of a deskside AI supercomputer that brings GB300 Grace Blackwell-class AI infrastructure straight into the Windows ecosystem with 748GB of coherent memory and up to 20 petaFLOPS of FP4 AI compute. Compact desktop configurations featuring the RTX Spark superchip in a small chassis designed for 24/7 operation will be available for sale in November.
Windows 11 Introduces Security Controls for Local Models
Hardware acceleration is paired with revisions to Windows 11 aimed at securing autonomous software agents. Davuluri detailed how agents operate under the hood, highlighting features like Microsoft Execution Containers built directly into the operating system to govern what information agents can access, alongside upcoming Windows functions that use Microsoft Entra to separate agent activity from human actions and bring Microsoft Agent 365 governance to local on-device agents.
On the software side, Microsoft’s HydraFusion is coming from GitHub to Windows on October 15 for GitHub Copilot, the CLI, and VS Code, and will use new models like MAI Code 1.1 Flash and Nvidia’s Nemotron. Developers can select MAI Code 1.1 Flash through the Windows ML provider, which serves as the runtime to deploy models across GPU, NPU, and CPU, or connect GitHub Copilot to OpenAI-compatible local endpoints and choose from the models those endpoints expose.
Operating with advanced quantization, DeepSeek V4 Flash achieves local execution at a memory footprint of 60GB by utilizing 1.6 bits. Running short prompts on the Surface Laptop Ultra, Microsoft’s local MAI-Code-1.1-Flash achieves a throughput of approximately 60 tokens per second, dropping to just under 40 tokens per second when processing a 256K-token prompt.

Enterprises Combat Cloud AI Costs With Desktop Silicon
The hardware pivot arrives as enterprise customers grapple with the financial strain of high-volume AI usage. Following a period where companies engaged in tokenmaxxing, the practice of enterprise companies treating AI spend volume as a yardstick for productivity, Microsoft’s push toward local desktop inference offers a direct way to bypass per-request cloud charges for repetitive development loops.
By shifting tasks like iterative code generation, test execution, and file organization to local silicon, development teams can avoid continuous cloud compute fees while maintaining strict boundaries on what automated tools can access.