This essay is available to the entire AI Strategies for CEOs community. If you find it valuable, please consider sharing it with peers facing similar decisions.
The computer is built, and the models are collapsing into a utility. AI is moving out of the data center and onto the desk — and the real contest is the last mile: the device, the desktop, and the operating system where it actually runs.
In June, at its Build conference in San Francisco, Microsoft stopped describing Windows as an operating system for running apps and started describing it as a runtime for running AI agents. In the same breath, it did something quieter and more consequential: it lifted the requirement that on-device AI run only on premium “AI PC” silicon, extending local inference to the ordinary CPUs and GPUs already sitting in hundreds of millions of machines.
Read together, those two moves point to where the AI contest is actually going. Not deeper into the data center. Toward the last mile — where the contest is won or lost.
The last mile, and why it decides everything
In any network, the last mile is the unglamorous final stretch between the backbone and the user. It is also the expensive, hard part, and the one that decides whether the whole system delivers anything at all. A continental fiber backbone is worthless if the signal never reaches the building.
AI has spent three years building the backbone: the hyperscale clusters, the frontier models, the record capex. That work is largely done and, for almost every company reading this, is no longer where advantage is decided. The contested ground is the last mile — the stretch between a general-purpose model and the specific desk, device, and process where work actually happens.
The backbone is built. The last mile is where the contest moved — and where it will be decided.
Why did the model stop being the prize
Start with the number that should reset every board conversation. According to Stanford’s 2025 AI Index, the cost of running a model at GPT-3.5 quality fell from $20.00 per million tokens in November 2022 to $0.07 by October 2024 — a more than 280-fold drop in roughly eighteen months. Over the same period, open-weight models narrowed the performance gap with closed leaders from 8% to 1.7% on some benchmarks in a single year. Hardware costs are falling by about 30% annually; energy efficiency is improving by about 40% annually.
When the price of intelligence falls by two orders of magnitude and the gap between the best model and a free one shrinks to noise, the model stops behaving like a moat and starts behaving like electricity — essential, ambient, and priced toward zero. Value does not disappear when an input commoditizes. It migrates to wherever that input is delivered, controlled, and put to work. For AI, that place is the last mile.
The last mile is, first, the workstation.
The computer itself is physically descending. Gartner expects AI-capable PCs to move from roughly 31% of global shipments at the end of 2025 to about 55% in 2026 — from around 78 million units to 143 million in a single year — with IDC projecting more than 60% of shipments by 2027. The edge AI hardware market, near $24 billion in 2024, is forecast to reach roughly $107 billion by 2030. At Build, Microsoft and NVIDIA put workstation-class silicon on stage capable of around a petaflop of AI compute and running models up to 120 billion parameters locally.
Inference is arriving where the work is, and the reasons are the ones CEOs already care about: latency, cost per query, privacy, and data sovereignty. The decision of where a model runs — cloud, edge, or the machine in an employee’s hands — has quietly become an architecture decision with cost-structure and compliance consequences.
But the prize is the operating system — because that is where the contest is won.
The device is the vehicle. The operating system is the prize — and the place where advantage is captured.
This is the shift Build 2026 made explicit. Microsoft reframed Windows as an agent platform, extended its on-device AI APIs from NPU-only to any capable CPU or GPU, began baking local models into the OS itself, and unveiled its own in-house model family — reducing its dependence on any single external model provider. Google is folding Gemini ever deeper across Workspace; Apple is running its models on-device through the OS. The pattern is identical across all three: you own the layer where AI runs locally and acts on the user’s files, mail, and calendar, and you own the relationship and the data — regardless of whose model sits underneath.
That is the real contest of the next cycle, and it is not “which model.” It is the operating layer that runs the models on your people’s machines. The model is the commodity. The operating system is the toll road — and the toll is control.
The model is the commodity. The operating system is the toll road.
What it means for CEOs
Your next PC refresh is an AI architecture decision in disguise. Decide it deliberately, not as part of routine procurement. The hardware cycle now underway — accelerated by Windows 10’s end of support — determines the ceiling on what AI your organization can run locally for the next several years. Judge it against latency, privacy, and sovereignty requirements, not only cost and warranty terms.
Stop anchoring strategy to a model you rent. A 280-fold price collapse is the market telling you the model is not to your advantage. What is yours is the data your operation generates, the processes a competitor cannot copy by signing the same API, and the local layer you actually control. Spend less time choosing the model and more time defining what you wrap around it.
The lock-in you escaped is reassembling one layer up. If your AI runs inside a single vendor’s operating system, agent runtime, cloud, and devices, you are rebuilding the platform dependency you spent a decade unwinding — this time with your proprietary data flowing through it. Treat the on-device AI layer as a strategic dependency to be governed now, while the standards are still forming, not after they harden.
A note on Europe
That last point cuts in an unexpected direction. The last-mile layer — local, sovereign, sitting next to regulated data and embedded in real operations — is the one part of the AI stack where Europe is not structurally behind and arguably holds the better hand. Treat that as a separate strategic question, and it will get one.
The CEO move
Stop benchmarking models. Map where your AI runs and who owns that ground. The backbone is finished and available to everyone, so it confers no advantage. Assign ownership of the last mile now, because the contest is there — and it ends on your employees’ desks.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.