The Last Mile Notebook, Issue 1
The first time I watched an enterprise AI pilot die, I assumed it was a fluke. A good team, a capable model, a demo that made the room lean forward, and then six weeks later it was quietly shelved. Bad luck, I thought. Wrong use case. Move on.
Then I watched it happen again. Different company, different industry, different model. Same death. And then a third time, and a fourth, until the pattern stopped looking like bad luck and started looking like a law. The pilots were not dying because the AI could not think. They were dying because the AI could not do anything. That realization is why I am writing this newsletter, and it is what the whole first issue is about.
This is The Last Mile Notebook. It is the place I am going to write down what I actually see while trying to get AI agents from a slick demo to something running inside a real business, governed, at scale. Not predictions. Not think-pieces about where AI is going in five years. Field notes from the part everyone skips: the last mile between “it works in the sandbox” and “it is in production and legal is fine with it.” If that is the part of the job you live in, I think you will want to be here.
You have probably seen the MIT figure by now. Project NANDA’s State of AI in Business 2025 looked at 300 deployments and found that about 95% of enterprise generative AI pilots deliver no measurable impact on the business. Only around 5% break through.
When that report landed, most people read it as a verdict on the models. The AI was not smart enough. We need better foundation models, better fine-tunes, better prompts. I read it differently, because by then I had already watched enough of these die in person to know where the body always was.
It was never in the reasoning. The reasoning was usually fine. The demo was real. The body was always in the same three places.
Here is the sequence I have now watched more times than I can count.
A team builds an agent. In the sandbox it is genuinely good. Then they try to wire it into the actual company, and they hit the first wall: the data the agent needs is scattered across systems that were never built to share it The CRM here, the ERP there, the warehouse somewhere else, all with access controls designed for human employees with badges, not for an autonomous process. So someone does the only thing they can. They hand the agent a credential and quietly hope.
Then the second wall: the agent can think but it cannot safely act. Reading data is one thing. Writing back, triggering a workflow, moving money, updating a record, that is the thing the business actually needs, and there is no governed mechanism for an autonomous process to do it. So either a human signs off on every single step, which makes the agent a very expensive form, or the agent acts freely, which makes the security team’s stomach drop.
And then the third wall, the one that kills it for good: nobody can say what the agent did. A regulated business cannot put a process into production if it cannot reconstruct, after the fact, every action that process took and why. There is no audit trail because the agent was never built to leave one. Compliance asks one question, the team cannot answer it, and the pilot goes into the drawer where pilots go to die.
Three walls. Data, action, audit. None of them is a model problem. Every one of them is infrastructure that does not exist yet.
I want to be honest about something, because this newsletter is going to be honest or it is going to be boring. When I first understood that the 95% was an infrastructure problem and not a model problem, my reaction was not disappointment. It was relief, and then something close to excitement.
Because a model problem is somebody else’s problem. If the bottleneck were model capability, the answer would be “wait for the labs,” and there would be nothing for the rest of us to build. But an infrastructure problem is a buildable problem. The thing standing between 95% and production is a layer that someone can actually make. The execution layer underneath the agent: governed data access, an enforced action boundary, an audit trail that holds up. The 5% who break through are not running smarter models than everyone else. They are running on that layer. They built it, or they bought it, and it is the entire difference.
We spent a decade building software for humans. Deterministic rules, siloed databases, a person in the loop holding the context together in their head. Then we dropped probabilistic, autonomous agents on top and acted surprised when they could not survive it. That is not a tragedy. That is just a layer we have not built yet. And building it is the most interesting problem I have worked on.
Yes, I am building in this space. I am not going to pretend otherwise or hide it. But this newsletter is not a place I am going to sell you anything. It is the place I think out loud about the problem, because thinking out loud in public is how I have always figured things out, and because the people who deal with this same wall are exactly the people I want to argue with.
Here is the plan for The Last Mile Notebook, so you know what you are subscribing to.
It is field notes, roughly weekly, on getting agents from pilot to production. Some issues will be about a specific wall and how a real team got past it. Some will be uncomfortable, like the next one, which is about the clock I think most enterprises are not watching: the leaders already run about a dozen agents in production, coordination breaks around five without something underneath holding it together, and the catch-up window looks like it is measured in months, not years. Some issues will just be me being wrong in public and correcting it later, which is the most useful thing a builder can do.
What it will not be: recycled X threads, generic AI commentary you could get from any model, or a drip campaign with a newsletter costume on. If I ever start writing those, unsubscribe, and you will be right to.
If you have watched a pilot die at one of those three walls, hit reply and tell me which wall it was. Data, action, or audit. I read every reply, and the patterns in those replies are going to shape what I write about next. The best issues of this thing are going to come from your stuck pilots, not my opinions.
And if the way I see this is useful to you, subscribe. The next issue is about the twelve-month window, and why being one agent behind is worse than it sounds.
That is the whole game. See you in the next one.
Written with CreateOS
---
Sources I leaned on this issue:
1. MIT Project NANDA, State of AI in Business 2025
2. The longer, product-voice version of this argument lives on our blog:
Why 95% of Enterprise AI Pilots Never Reach Production

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.