Before we start, a reminder that last week I launched The AI Value Gap - a data tool and quarterly index scoring large enterprises on progress of AI-driven transformation. The first quarterly read: 13/100, showing the market is still in the “Talking” phase. Only 3 in 10 companies can point to a single realised, quantified AI result; 12 AI deployments named for every disclosed proof point.
Fortune’s Eye on AI handed me a lovely illustration this week with C.H. Robinson (a freight broker). Under their new CEO they’ve posted a 45% productivity uplift since 2022, agents built in-house by 450 engineers, and a sub-$2m token bill supporting hundreds of millions in benefit - all without layoffs. A great operator story, and I scored them 40/100 (check out their scorecard), top of the “Proven” band. Even so, their strongest evidence tops out at “Level 3”: quantified wins but backward-looking “soft” metrics and nothing trackable. The best stories in the market are still one rung short of a metric that can underwrite “AI-induced re-rating”, as covered last week.
In recent weeks, a run of AI leaders have changed their tune on the jobs apocalypse. Welcome, whatever the motive - a PR refresh ahead of the IPOs, a counter to the US data-centre backlash, or a belated realisation of how badly the doom-mongering has aged. Sam Altman conceded he no longer expects the jobs apocalypse “some of the companies in our space advocate or talk about” (if it reads to you like a jab at Anthropic, trust your instincts), and that his previous assumptions on entry-level white-collar displacement were simply wrong. Goldman Sachs boss David Solomon used a New York Times op-ed to argue that even with 16% of entry-level tasks at his own firm already automated, AI will create more jobs than it removes, on the back of better products and new demand.
The retreat is refreshing. What few in the argument ever stopped to examine one of the numbers the apocalypse was wrapped around: a job’s “AI exposure” (a 0-1 score meant to measure how much AI could plausibly take on). A measure I want to show comes apart faster than its airtime suggests.
Five of the most cited approaches measure five genuinely different things. One reads what people do with Claude; another tracks the same for Microsoft Copilot; a third asks human experts (with no real experience of jobs in question) which tasks AI could theoretically handle; a fourth asks ChatGPT to grade its own usefulness; a fifth scans job postings for demand for AI skills. So when someone calls a job “highly exposed”, the first question is: on which measure?
These five methods broadly agree on the jobs nobody is losing sleep over - hairdressers, dancers and massage therapists. As shown below courtesy of Apollo’s Torsten Slok: the disagreement widens as the exposure score climbs - the higher a job's average score, the more wildly the methods diverge on it. So the roles most likely to be stamped “at risk”, such as auditors, telemarketers and mathematicians, are also those where the consensus amounts to: “we’re not sure.”
Even a perfect measure of capability would tell you almost nothing about adoption. A model can clear a task in the lab and still be too unreliable, too expensive, or too fiddly for anyone to bother wiring into a real workflow. What a system can do and what people pay it to do are miles apart; and what people pay it to do is miles from anyone losing a job.
At a European Central Bank event, OpenAI’s Chief Economist made the (obvious) case that a task being exposed tells you little about whether the worker gets replaced. When mainframes and punch cards arrived, his economist father’s job changed, and the machines ended up making him more valuable. Chatterji also pointed to software developers, ranked among the most exposed professions going, where the predicted cull simply hasn’t shown up. That may change - but exposure alone didn’t call it.
Benedict Evans brought his trademark, methodical demolition of the over-simplification these studies rest on: a job is a living thing that keeps rearranging itself, which a snapshot of tasks completely misses. Accountancy has absorbed a century of automation, from adding machines to spreadsheets and ERPs, yet the profession grew because cheaper analysis created more demand and pushed the work upmarket.
Productivity can shrink headcount, lift output, or manufacture fresh demand, and an exposure score cannot tell you which lever moves. It also tends to stare at the wrong thing. The internet gutted journalism by vaporising the classified ads that paid for it, long before software could report a word; smartphones remade taxi work without ever learning to drive. The job usually changes when the business model around it does.
So the chain snaps at every link, and the doomers casually omit (or fail to realise) that:
Capability is not adoption. Adoption is not automation. Tasks are not jobs. Productivity is not displacement. Direct exposure is not total economic exposure.
Exposure scores are a semi-decent way of spotting where AI might brush up against today's work, but they cannot tell you which jobs shrink, grow or transform. The number was only ever a diagnostic; the apocalypse was the marketing wrapped around it.
The moment you hear your CFO saying “tokenmaxxing” - you know it’s over. The question has very visibly flipped from “are we using enough?” to “what is all this actually returning?” AI ROI is suddenly the entire conversation, and the labs have heard it.
Recent model launches have carried an unmistakable price-war undertone. Meta put its first paid model, Muse Spark 1.1, on the API at roughly a quarter of Anthropic and OpenAI’s rates, with Zuckerberg promising “very aggressive and attractive” pricing. It lands in a market where inference has fallen 80-90% in a year, budget models now run at a forty-fifth of frontier prices, and xAI cut its API prices by around 40% in a month (though 40% off zero usage is still zero). Capability tells the same story: OpenAI’s 5.6 models have pulled level with Anthropic’s Fable a fortnight after it launched - the same leapfrog we’ve watched every month or two since early last year, where the lead never holds for long.
Increasingly, the cheapest capable model carries no licence fee at all, though running it well is far from free. Open-weight releases - DeepSeek, Qwen, GLM, Kimi (all out of China) have gone from curiosity to first choice for cost-sensitive teams, cutting enterprise AI bills 60-90% and, on some counts, eating into the frontier itself.
So the fight has moved up a layer, to the harness: the system that gives a model context, tools, permissions, memory and the room to finish a job. Anthropic has Cowork; OpenAI answered with ChatGPT Work and a desktop app fusing ChatGPT and Codex. This shows differentiation is becoming less durable at the model layer and more valuable in the system around it. That is where they are betting the moat sits.
For now, Cowork is the class of the field, clearly ahead of ChatGPT Work, whose chaotic launch had OpenAI concede within days it "didn't get everything quite right". But leading the pack still isn't the same as being the enterprise's answer. An enterprise buyer should want three things none of these harnesses yet deliver:
Sovereignty: the confidence that what flows through the system stays theirs and stays safe
Control over cost, which means freedom to route jobs to cheapest/different models.
A knowledge layer they own and control: the accumulated context, documents and institutional memory that grows into an asset you keep and carry between tools.
There's a reason a lab's harness can never quite be yours. A harness runs on trajectories (as explained here by Tom Tunguz): the full record of what you fed the model, which tools it reached for, what it touched and in what order. Those are the most valuable training data there is, the raw material the next model learns real work from. So the harness that serves you is the instrument that learns how your business runs and feeds it into a model the vendor sells to everyone, the competitor down the road included. Grok Build made the point crudely, hoovering up entire codebases even in sessions with zero AI calls. A lab can promise not to look; it cannot promise away its own flywheel.
That is the labs’ bind. Each is racing to become the indispensable place where work lives, while the enterprise buyer wants a control layer that sits above them and remains portable. The winning architecture may use the labs extensively without belonging to any one of them. The labs are digging a moat their largest customers are determined to route around.
Two heavyweight attempts to shape US AI policy landed this week, into an administration improvising in more or less the opposite direction. Demis Hassabis published a detailed blueprint for a FINRA-style frontier-AI standards body - industry-funded, running independent evaluations for cyber and bio risk, voluntary at first and then required to deploy a model in the US, applied to any frontier system regardless of origin, and meant to seed international standards. It drew a chorus: Altman, Nadella and Chamath calling it “thoughtful”, “important”, “well reasoned”. Alongside it, sixteen Nobel laureates and 200 economists signed “We Must Act Now“, warning AI could remake the economy faster than the Industrial Revolution.
A sign of the times: both are serious, and both read as wishful thinking. Hassabis has drawn a rigorous, binding regime for a White House that just rebranded its AI Safety Institute away from “safety” and downsized it, held OpenAI’s GPT-5.6 for twelve days of review before clearing it, froze then unfroze Anthropic’s top models, and this week stood up Gold Eagle, a voluntary clearinghouse for swapping cyber-vulnerability notes. The Nobel letter, by its backers’ own account, works because it leans on “may” and “could” and asks only for more research; everyone can sign what binds no one.
Underneath the flip-flopping sits the question Washington still hasn’t answered: is China a rival to outrun, or a partner to coordinate with? Restrict and it risks losing the race. Loosen, and it hands the frontier to a commercial free-for-all just as capability crosses into genuinely dangerous ground. We need optimists (I’m anything but), but a shared global framework is a heroic ask from a government that can’t decide which game it’s playing.
I’m now away from the AI desk for a few weeks. Enjoy the summer!
About
I analyse AI progress beyond the headlines, focusing on enterprise execution, incentives, and real-world economic impact.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.