“Ninety-five percent of enterprise AI projects fail to produce any measurable financial return.” That number, from an MIT study last year, gets quoted everywhere, almost always as proof that corporate AI is mostly hype - billions spent, nothing to show for it.
I've always thought that points the wrong way. The interesting data isn't why 95% fail - it's what the 5% that worked did. If most projects fail for the same reasons, and the few that succeed share the same habits, then the gap between the two groups is the most useful thing to study right now. A recent Stanford report I read does a good job of studying exactly this - and its conclusions, drawn from real deployments rather than theory, line up closely with what I’ve been seeing in the market.
It comes out of Stanford’s Digital Economy Lab, run by Erik Brynjolfsson, one of the most respected economists working on why new technology takes so long to show up in real productivity numbers. What makes it credible is the method. Most of what gets published on enterprise AI is either a vendor case study, chosen to sell you something, or a sentiment survey. Over five months, from August 2025 to February 2026, the team ran structured, hour-long interviews with the people who actually built and ran 51 AI deployments, and backed every one with the internal metrics, project reviews and financial documents the companies handed over.
The qualification criteria to be included was rigorous and specific. A project had to clear four tests:
It had to be live in production and genuinely wired into how work gets done, not a demo.
It had to be relied on across functions for at least three months.
It had to produce a business result you could actually measure - productivity, revenue, customer satisfaction.
It had to be built in a way that could scale or repeat across teams and geographies.
Each of those was scored against real evidence - a working system, documentation, a named owner. All data was quantifiable vs speculative.
The sample itself is broad: 51 cases across 41 organisations, nine industries and seven countries, at companies employing over a million people between them, with the most depth in manufacturing, financial services and technology.
To conduct the report, they purposely went looking for failures. Every interview asked what had broken first - the false starts, the abandoned attempts, the money that led nowhere - before the company eventually got it right. Two-thirds of them had failed at least once. This proves the messy reality of getting AI to work inside a real business.
Two caveats worth stating. The data is self-reported, and the sample is selected toward success by design - the point was to study what worked, not how often it doesn't. So the report can't tell you the odds of any of this failing. What it can do is show you, in detail, what the companies that succeeded actually did - and on that, it's the most rigorous report I've seen, backed up by a large sample size.
I’ve included a link to the full report below:
Stanford: The Enterprise AI Playbook
The chart above is an interesting one. The researchers asked every practitioner one specific question - what was the hardest thing to get right? - and 77% of the answers had nothing to do with the technology. Change management. Documenting how the work actually gets done. Cleaning up the data. Notably, the model and model selection was barely mentioned.
That runs counter to how most people assume. Every vendor pitch puts the technology at the centre and treats everything else as additive. The people who’ve actually shipped this say the opposite, and near unanimously. “There are people, process, and technology,” one executive told the researchers, “and I know it’s in that order even though I represent a technology company. The technology was the easiest part.” Another put the whole finding in four words: “The problem isn’t the models.”
None of this would surprise Brynjolfsson, because he built the framework for it years ago - the Productivity J-Curve. A genuinely new general-purpose technology arrives with a hidden tax. The technology is cheap; making an organisation able to use it is not. You rewire workflows, retrain people, restructure teams, rebuild the data - and that work is huge, invisible and slow to pay back. His earlier research even priced the tax: up to $10 of intangible work for every $1 of technology. The model is the cheap part. Everything that makes the model useful is where the money goes.
The report purposely grounds this in ordinary cases, and one is worth walking through. A billion-dollar US logistics company runs a fleet of refrigerated trailers and processes more than 100,000 repair invoices a year. Those invoices arrive in every format imaginable - the repair shops, as the executive put it, are often “in the middle of nowhere” and simply phone in what they’ve done, so the same information turns up as an email, a fax, or a voicemail. Seven full-time employees did nothing but gather it all up, check it, and key it into the company’s finance system.
The AI they built to fix this was not overly complex. Software to read the documents, a model to work out what each invoice said and match it to the right template, and an automatic write into their finance system. They “basically used a lot of open-source and off-the-shelf stuff.” But look at everything that had to happen before any of that could run. The company had accumulated 750 invoice templates over the years, most of them duplicates, and nobody had ever gone back to review them - all of it had to be cleaned up and cut to a few hundred before a model could make sense of it. Experienced staff then had to check thousands of the AI’s outputs by hand, on top of their day jobs, explaining each mistake so the system could improve. The president personally ran weekly check-ins to clear blockers. And two junior IT staff were embedded from day one, so the business could run the thing itself rather than depend on a vendor.
That demonstrates the less obvious work. It paid off - seven roles down to two, invoices cleared in under 24 hours, over $1 million in value, live in eight weeks. But almost none of the effort behind that result was technical. It was template cleanup, output-checking, sponsorship and training.
Then there’s the number underneath the headline number, that’s applicable to all 51 companies in this study, yet it’s the one nobody plans or accounts for.
Six in ten of these successes were built on a previous failure - money spent, project shelved, none of it appearing anywhere in the eventual return. And the failures rhymed. A company treats AI as a technology project, aims a capable model at a broken process, and waits for the model to fix the process. It doesn’t. The report’s recruiting example is textbook: the second attempt succeeded only because the first had failed by “assuming AI would just fix processes instead of requiring process redesign.”
The mistake I watch companies make most frequently is exactly this. They spend their conviction and their budget on the model, and starve the work that actually decides the outcome - the data, the context, the process redesign. The model was never the hard part; the context and data around it always were. And when the first attempt stalls - as six in ten do - they read it as "AI doesn't work here" and stop.
One of the most important takeaways in the report isn’t about technology at all. It’s about a design choice - how much you actually let the machine do.
Every deployment sits somewhere on a spectrum. At one end, the AI drafts and a person approves every output - a doctor signing off each AI-written clinical note. At the other, the machine runs the whole task alone - an AI making a supermarket chain's purchasing decisions, with no one checking each order. Same technology, very different amounts of human control. It's usually framed as human-in-the-loop versus agentic, or automation versus augmentation.
There are two ways to keep a human in the loop. In an approval model, the AI proposes and a person signs off every action before it happens. In an escalation model, the AI runs the work itself - 80% or more of it - and kicks only the hard or unusual cases up to a human. Escalation delivered a median 71% productivity gain. Approval delivered 30%. You more than double the return on identical models, purely by moving where the human sits.
You can see why in the cases. A food-delivery company put an escalation model on customer support: “90 or 95% are now fully automated by an agent,” with humans pulled in only for the genuinely messy complaints. That’s the 71% world - the AI owns the workflow, the person owns the edge cases. Contrast it with a hospital deploying AI into clinical documentation, where a doctor reviews and approves each output before it goes into the medical record. That’s the 30% world, and in that setting it’s the right call - the AI is drafting, but a human still has to sign, because the cost of a bad output is consequential.
The other question is how far companies have actually pushed this - and the answer is, not particularly far. The report scores every deployment on a three-point scale. Agentic, high automation and human-in-loop.
Most sit at the cautious end. 46% still have a human approving each output. Another 34% run high automation, with people handling only the exceptions. Just 20% have reached full agentic autonomy - AI completing a task end to end with nobody in the loop. The more of the work companies hand over, the higher the productivity gain the report records; and that’s precisely the approach fewest have adopted. Bolt a faster assistant onto an unchanged workflow and you get 30% - the same job, done quicker. Take the person off the default path, let the machine run the work and escalate only the exceptions, and you get 71%. Same models. More than double the return. The difference is a decision, not a technology - and it's the one most companies never make.
None of this means automation is always right. The correct level of autonomy depends on the industry and the task. In regulated, high-stakes work - a compliance filing, a clinical decision, a credit approval - augmentation is the right design: the AI drafts, a human owns the decision. But in the many domains where AI can genuinely complete a task, or an entire role, end to end - support resolution, procurement, alert triage, document processing - automation is the far bigger prize, and keeping a human in the loop moderates the ROI. We've seen this first-hand at Abingdon: our travel business, DCS, has built an agentic AI product that automates the end-to-end role of a corporate travel agent - not assisting the agent, but doing the job itself.
So the real dividing line isn’t which model you run - it’s whether you’ve matched the level of autonomy to what the work actually allows. Agentic AI isn’t a new interface, it’s a redefinition of who does what. The trouble is that most companies default to augmentation everywhere, including the tasks that could be fully automated - they buy the faster assistant, leave the workflow untouched, and call it transformation.
By this point the technology works. Whether it ever diffuses through the organisation comes down to two questions: who has the power to push it through, and who has the power to stop it. Both are questions of culture and structure - the theme I keep returning to, having covered it in previous articles.
Start with who stops it, because the data inverts the conventional wisdom.
Opposition to these projects does not come, in the main, from the front-line staff being automated. It comes from the staff functions - legal, risk, HR, compliance - which account for 35% of the resistance, well ahead of the end-users at 23%. External stakeholders and clients tie the end-users at 23%; even the C-suite shows up at 15%. The functions whose entire mandate is to enable the business turn out to be the ones most likely to halt it.
None of this is irrational, which is why it's so hard to get past. These functions exist to govern risk the business can't yet price, and a probabilistic system is precisely what they're built to slow down. In my experience you never win this on technical merit - you win it by bringing legal and risk in on day one, with real authority and a version they can put their name to: governed, logged, permissioned. Handled that way, the function that blocked you first becomes the one that clears you to touch sensitive data at all. It's the move I flagged at JPMorgan - governance placed inside the project as an enabler, not left at the perimeter as a veto.
The other half of the equation is sponsorship. Sponsorship sets the direction, but it doesn't move an organisation on its own - that takes capability built underneath it. The clearest model I've seen is Ramp's L0-to-L3 framework, which grades every employee from no real AI use (L0) up to shipping production features with it (L3), turning fluency into a ladder people can see and climb. In practice it works best paired with an AI champions network - the handful of people in each team who've genuinely cracked these tools, made responsible for spreading the practice to everyone else. I've run exactly this internally, and it's what converts a leadership mandate into day-to-day behaviour: the CEO sets the direction, and the champions and the ladder carry the rest of the company up it.
That leaves the question everyone actually wants answered: what does any of this do to jobs? The report’s answer is surprisingly measured.
Headcount fell in 45% of cases. In the other 55%, companies redeployed people, avoided hiring they would otherwise have done, or cut no one - and the savings more often funded the next project than a round of layoffs. One executive put it plainly: “AI is not replacing the person you have. AI is replacing the person you don’t need to hire.”
That aligns with a view I’ve argued before, and with the case David Sacks has made repeatedly - that the fear of mass AI displacement is overblown, and that for now AI is more a productivity multiplier and a net creator of roles than a straightforward destroyer of them. The data reads the same way: AI is taking out the job a company would have added, not the one it already employs. My one caveat is the same one I’ve made before - certain roles, and certain categories of role, will genuinely go. “Overblown in aggregate” and “painful in particular” are both true at once. But on this evidence, the measured read wins over the doomsday one.
The most consequential finding for anyone building is also the one I've been arguing toward for a while: the model itself is the least defensible part of the stack. The report confirms this.
For 42% of the use cases, the model was a commodity - fully interchangeable, with no measurable difference which one you used. Another 39% said it mattered moderately. Only 19% treated the model as a genuine differentiator. The report's own conclusion is the one to internalise: the durable advantage is in the orchestration layer, not the foundation model. What decides which camp you land in is task complexity, and the divide is stark.
On routine work - support triage, search, classification, drafting - 71% treated the model as swappable, and not one company called it critical. On advanced work - hard code, clinical documentation, real reasoning - a third said the model was a genuine differentiator. The frontier still earns its price at the hard end of the distribution. Everywhere else, the model is a part you buy, not part of your IP.
The companies that understand this don’t pick a model at all; they build a layer above it that treats models as interchangeable parts. In practice that means two things. A multi-model gateway routes each request to the right engine - a cheap, fast model for the easy 90% of queries, a frontier model for the hard 10%, with a 10x cost gap between the two. And an abstraction layer lets them swap the underlying model without touching the application, so every improvement from every lab is absorbed automatically rather than leaving them stranded on last year’s choice. One food-delivery company reached 90-95% automation in support built exactly this way, tied to no single provider. Another ran each query through two models and trusted the answer only when both agreed, using redundancy as the quality check rather than any single model’s confidence (this concept is termed ‘AI as a judge’).
All of that routing assumes a field of models to choose from - and the fastest-moving part of that field is open source. The numbers show how quickly open models are advancing, and how much room to manoeuvre that gives anyone building above them.
Open models now reach around 90% of closed-model performance at roughly six times lower cost per token. And yet proprietary models still carry about 80% of actual usage, because enterprises are optimising for capability and speed, not price - for now. The shift is already visible: on OpenRouter, which routes across hundreds of models, four of the top five by volume are Chinese open-source models, pulled up by agent workloads that burn tokens at a rate chatbots never approached. Today, most enterprises still pick a model on capability and treat the cost of running it as a rounding error - though, as I wrote in From Token Maxing to Token Efficiency, that is already changing. As agentic workloads scale, consuming one to two orders of magnitude more tokens than a chatbot, inference cost stops being a rounding error and becomes a first-order operating discipline. Token spend goes from something nobody tracked to something finance teams manage line by line. That is the point where open source stops being a cost play and becomes the default. The report says as much: in startups, where inference cost binds from day one, open-source adoption is already higher. The rest of the enterprise world is just earlier on the same curve - and the economics only point one way.
My read hasn’t changed, it’s just better evidenced now. The model is commoditising underneath everyone, and it will keep commoditising - it’s the one part of your stack a competitor can buy on identical terms tomorrow. The advantage that lasts sits above it: the routing, the data, the context, the workflow. Build the orchestration layer. Rent the model.
Almost everything in the report so far is a story about defence - doing work you already do, more cheaply and with fewer people. That’s where the measurable wins sit, and it’s where most of the 51 companies concentrated. But two of the findings point somewhere more valuable and less explored.
The first is revenue. The report is careful to say it’s still rare - most successful AI is saving money, not making it - but where it does show up, it follows three patterns.
The first is personalisation that converts - AI tailoring an offer or an experience closely enough to lift the conversion rate. The second is speed that wins deals: one insurance company turned AI contract-drafting into a competitive weapon, quoting work in the time it took rivals to open the file. The third is the one I find most interesting - internal tools repackaged as products. A technology services company built an AI invoice-processing system for its own use, realised every company in its sector had the same problem, and started selling it. That is the customer-zero move I've written about before: solve your own problem in production, then sell the solution to everyone who shares it. The report finding it independently, across a broad sample, tells you it's a repeatable pattern.
There’s a fourth category as well: AI making work possible that couldn’t be done at all before. Not faster, not cheaper - undoable. A robotic inspection company, for example, now runs its inspections with AI and turns them into a continuous historical dataset it uses to predict failures before they happen - a capability that never existed when a person did the inspecting. That’s the real shift: from “how do we do this more efficiently” to “what can we now do that we couldn’t?” The first question saves money. The second changes what the business is. Almost everyone is asking the first. Almost nobody is asking the second.
The other finding worth surfacing is about how long all this takes.
The same use case took weeks at one company and years at another - same technology, completely different timelines. And when the report breaks down what made the fast projects fast, the top factor isn’t technical. It’s executive sponsorship, followed by building on an existing foundation and a willing workforce. One executive, describing the slow version: “It takes us multiple years just to even stand one of these things up.” That is the whole thesis of the report in a single data point. The technology sets the ceiling. The organisation sets the speed.
Here’s where I’d place the bet. Cost-cutting is the legible win - you can point at the roles removed and put a number against them, which is why finance signs it off and why most companies stop there. Revenue and new capability are harder; they demand imagination and a tolerance for risk that defence never does, so most companies never seriously try. The companies already on offence are the case in point: Palo Alto Networks turned its own AI-rebuilt security operations into Cortex XSIAM, now a product line generating hundreds of millions, and Ramp ships the agentic tools it first battle-tests on itself. They built internally, then sold - and they’re pulling away from the ones that only ever used AI to trim.
And the public markets are starting to price this distinction directly. The software names re-rating hardest right now - Datadog, Twilio, Snowflake - are the ones putting up net-new agentic ARR, revenue that didn’t exist before AI, not the ones using it to protect margins. As I argued in AI Hits the P&L, investors are learning to separate AI that defends a business from AI that grows one, and paying up for the second. Offence isn’t only the better operating strategy. It’s the one the market rewards.
I read this report as much as a participant as an observer. For 18 months we've been building versions of it inside my own businesses - whilst also observing it across a portfolio of 200+ companies.
At Abingdon Software Group, we’ve built a fully AI-enabled headquarters across core functions - sourcing, diligence, knowledge and operations - and pushed the same approach into the operating companies we own. At Fuel Ventures, we’re building a proprietary intelligence layer across the 200-plus company portfolio: an orchestration agent that runs diligence through our own investment frameworks rather than generic prompts, a central knowledge system that pulls fragmented data - Drive, portfolio metrics, our CRM, internal documents - into one place so nobody rebuilds context by hand, and a way to make our 2,000-strong investor and operator network searchable and deployable on a live deal.
From that seat, here’s what the report gets right. The technology was the easy part - consistently, and more so than I expected. The hard part was everything it labels invisible. Consolidating data that lived in ten systems never designed to talk to each other. Writing down processes that existed only in the heads of people who’d run them for years - the tacit knowledge no model can read because it was never recorded anywhere. And, hardest of all, getting people to genuinely change how they work rather than run the AI beside the old process and ignore it. We paid the failure tax too: we built versions that didn’t work before we built the ones that did, and none of those dead ends show up in the finished result, though every one of them was necessary to reach it. The sponsorship finding rings especially true from the inside - none of this moved until it was owned from the top rather than handed to a technical team. When it’s a side project, it stays a side project.
But the part the report gets most right is what happens once the groundwork is done: the payoff compounds. Every deal that runs through our knowledge system recursively improves the next decision; every correction feeds back and improves how the system understands the way we actually operate. That flywheel is the whole advantage.
There’s also a business thesis buried in all of this, and it’s what I’ve spent the last few months thinking a lot about. If the model is the commodity and the implementation is the hard part - the process redesign, the data architecture, the routing layer, the fine-tuning - then a great many companies will want that capability and very few can build it in-house. Everyone wants to adopt AI. Almost nobody knows how. The firms that do this well today - the Accentures and McKinseys - price it out of reach for all but the largest enterprises, which leaves the entire mid-market and below underserved. And the shift toward open source only widens the gap: running open-weight models yourself - fine-tuning them, hosting them, managing inference - takes far more technical depth than calling an off-the-shelf API from OpenAI or Anthropic. As the economics push companies toward open source, the expertise required to actually use it goes up, not down.
That gap is the opportunity, and it’s a big one. I think there’s a real business in bridging it - call it an AI enablement lab - making the architecture this report keeps circling accessible to the companies the big consultancies won’t touch: open-source deployment, fine-tuning and inference, data lakehouses, and the enablement work itself. If you’re building in that space, or thinking about it, I’d genuinely like to speak to you - please reach out!
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.