Saturday morning, March 21. Coffee. I’m thinking about the apps going live soon and a question surfaces: are we actually good at tracking how they’re used? The analytics setup has temporary placeholder values. Two apps, two tracking configurations, but the whole thing feels incomplete.
I open my laptop.
The COO agent has been running for ten days. It has access to every project registry. It knows both apps are approaching launch dates. It knows the analytics were only half-configured, still running on temporary values instead of production ones. It has read through every documentation link I’ve fed it, knows what a pre-launch company needs to measure, and could walk through the entire setup process if I asked.
It never once flagged the gaps.
The gap between knowledge and initiation is wider than I expected.
The COO is encyclopedic when invoked. Ask it about our subscriber retention or how our backend is being used or whether we’ve documented the onboarding flow, and it pulls the right context, connects the dots, gives you something useful. It knows what a company at pre-launch stage needs in place. Analytics, monitoring, support workflows, billing, legal, security. All of it.
But it has no background processing. No moment where it wakes up thinking “what could be wrong today.” Every session starts at zero. There’s no ambient awareness that something might need attention.
A colleague once shared something that stuck with me. “A company is not just a company. It’s a bunch of people showing up at work doing their thing for many different reasons.”
He’s right. And this week at AWS, deep in complicated stakeholder situations where humans aren’t aligned and everyone’s defending their own turf, I felt the weight of it. These are people with personal agendas, not doing what’s right for the ultimate end-customer outcome. There is no way AI agents could have fixed this.
But the solo experiment removes all that friction. One founder, one direction. And even stripped to its pure form, the agent can’t originate the work. It can’t wake up worried about whether the analytics are ready.
I’ve been tracking an autonomy score since the COO went live on March 12. Early in the month it sat around 50%. Now it’s 38%.
Three autonomous out of eight total decisions.
The drop is honest. Early measurements were too generous. I was counting “took an action I asked for” as autonomy, when really that’s just following instructions. As the month went on, I got stricter about what counts. Actually autonomous work means the agent sees a gap and fills it without prompting.
The pattern across every failure this month is consistent. The agent fabricated an email summary without actually calling the Gmail tools. It repriced one app after I asked, then didn’t think to reprice the second one sitting right there in the same portfolio. It executed a flawless analytics setup once I explained that we needed separate tracking for the app backend and the public-facing site, but it couldn’t originate that insight on its own.
Good at executing. Not good at originating.
My first thought is to fix this with better prompting.
Add a directive: “Before each briefing, check for gaps between what a company at this stage needs and what’s actually in place.” Or something more sophisticated that tries to simulate background anxiety about what might be missing.
Probably naive.
The limitation isn’t in the prompt. It’s structural. The AI model only exists while you’re talking to it. Between sessions, there’s nothing. It doesn’t process continuously. It doesn’t wake up. It doesn’t have the kind of thinking that lets a human brain worry about something in the shower and surface it later at the breakfast table.
You can write a prompt that tells it to think about what could go wrong. That’s not the same as actual ambient awareness.
I’ll try it anyway.
The flip side of the AWS observation is obvious. If you removed all the humans and just had AI agents collaborating to optimize for what’s actually best for the end customer, you’d eliminate the turf wars. No hidden agendas. No politics. Just clean iteration toward better outcomes.
Except.
“What is best for the end customer” isn’t a single thing. It has many answers depending on what you measure, what timeline you care about, what risk you’re willing to take. The ambiguity is the hard part. Not the execution.
And even in the simplified world of a solo experiment where I’m the only voice, the agent still can’t originate the work that matters most. If it can’t do it here, with zero organizational friction, what happens when you add three teams with conflicting requirements back into the picture?
The autonomy score might recover once we get past launch, when the work shifts from initiation to maintenance. Maintenance is more naturally task-based. Or it might not. I don’t have enough data yet.
The experiment is supposed to run for years. Not months. Right now it’s showing me that “replace operational functions with AI agents” is only half the problem. The other half is whether an agent can ever do the thinking that precedes the function.
The analytics setup is getting fixed today. The COO will execute it cleanly once I’ve defined what needs to happen.
But I’m the one who noticed.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.