The Frontier Moves Into the House
Jensen Huang told a good story this week, and like most good stories it was true in the parts that flattered the teller.
On Wednesday afternoon in San Francisco, Satya Nadella and Jensen Huang stood in front of a wall of laptops and gave Windows a different job. Not the place you launch apps, but the place agents live: small models running on the machine in front of you, reaching for the cloud only when the local one is not enough. Microsoft calls it hybrid intelligence. The promise, in Nadella's own words, is "unmetered intelligence to every home and every desk with Windows".
That is the news. The interesting part is the bill.

The story Huang tells
Huang's post is a lineage. Windows created the platform shift that made NVIDIA; NVIDIA invented programmable shading for DirectX, which became CUDA; NVIDIA built the Azure supercomputers that helped OpenAI train GPT. Each link checks out. The GeForce 3 in 2001 was the first programmable GPU, built for DirectX 8. CUDA arrived in 2006. The Azure cluster that trained GPT-3 went in around 2020. A fair account, told by a man selling the next chapter.
The real news is narrower: the full CUDA stack, natively on an Arm Windows laptop. That combination has not existed before.
What actually moved
Three things are real. They arrive at three different speeds.
- Already here. Windows ML now runs GGUF models through experimental llama.cpp support, and a new Windows-native runtime sits in preview. PyTorch and CUDA have native Windows on Arm builds. None of it needs new hardware.
- This month. RTX Spark, the chip inside the Surface Laptop Ultra: 1 petaflop of AI compute and up to 128GB of memory shared between CPU and GPU. Pre-orders opened on the day; laptops land on 16 October.
- Later. DGX Station for Windows, a deskside box with a GB300 Grace Blackwell chip, 748GB of memory and 20 petaflops in FP4, sold on running trillion-parameter models locally. This is the one that replaces the cluster. Its date is simply "later this year".

The meter moves inside
"Unmetered" sounds like a bill that never comes. It is not. It stops counting tokens and starts counting something else: the price of the box and the electricity it draws. A heavy user comes out ahead. A light user pays upfront, before a single prompt, and hopes it was worth it.
Interactive · the meter
Run the same job two ways.
One workload, two bills. Flick the meter to see which one moves.
The bill climbs
Billed by the token
The meter runs while the model thinks. Stop asking, it stops.
The bill was the box
Billed by the wall
Electricity and silicon, not tokens. Ask it a thousand things and the meter does not move.
The honest counterpoint
There is a cost the keynote did not price. An always-on agent with a foothold in your files, your browser and your inbox is a new kind of attack surface, and the thing meant to fence it in, Microsoft Execution Containers, is days old. "Secure by default" is a claim, not yet a result.
I have not run any of this. I build local AI for a living, so this is my own patch, and I read the primary documents in full rather than the reaction cycle. But I did not benchmark RTX Spark, not shipping until the middle of the month, and I have not put MXC under load. This is a reading, not a measurement. What would change my mind: a public, reproducible local-agent containment test that survives red-teaming.
Readout: what is confirmed, and what is not
- Confirmed (NVIDIA, 31 May 2026): RTX Spark carries 1 petaflop of AI compute and up to 128GB of unified memory, and runs the full CUDA stack.
- Confirmed (Microsoft Windows blog, 7 October 2026): MXC is generally available; RTX Spark laptops pre-order now and ship 16 October; DGX Station for Windows arrives later this year.
- Confirmed (Microsoft Foundry blog, 7 October 2026): Windows ML adds experimental llama.cpp support for GGUF models; a Windows-native runtime API is in preview.
- Confirmed (NVIDIA press release, 31 May 2026): Nadella's line is "unmetered intelligence to every home and every desk with Windows".
- Not confirmed: the real performance and safety of the shipping hardware. Nothing here has been independently benchmarked or red-teamed.
If you are weighing a local box against a cloud bill, and want a second pair of eyes on the numbers, email brandon@kreostudio.co.uk.
Sources & references
- NVIDIA, Microsoft Kick Off a New Beginning for Windows, NVIDIA, 7 October 2026.
- Building Windows for hybrid intelligence, Pavan Davuluri, Microsoft, 7 October 2026.
- AI Development on Windows: from PyTorch and llama.cpp to Windows ML, Microsoft Foundry, 7 October 2026.
- NVIDIA and Microsoft Reinvent Windows PCs for the Age of Personal AI, NVIDIA, 31 May 2026.
Reader signal
Was this useful?
Work with KREO Studio
AI engineering, data science and design architecture, from Plymouth to the wider UK.
Next Article
