Last week, we were on a call with a VP of Engineering at a large e-commerce company. We’d had 4 meetings with the engineering teams that run reliability for him, and we had a pretty detailed picture about what systems they used, where their data lived, and how the org wanted to tackle some of their key challenges. As many larger organizations tend to do, this team was operating highly manually – getting notified about new tickets for incidents and linking them to manually created channels for discussion and triage. We were caught off guard, then, when the leader claimed that a lot of the functionality we were describing was something the team already had built in house. What he was really looking for was apparently an agent that would automate end-to-end resolution without engineering involvement.
We later confirmed with the team that we weren’t crazy. We knew more about the state of his team’s operations than he did.
This is not particularly new for anyone who’s worked at or sold to big companies before. There are plenty of leaders who aren’t connected to what their team is actually doing – enterprises have long succeeded and grown despite this disconnect. Unfortunately, the perils of this disconnect have gotten orders of magnitude more pronounced with AI.
Leaders who don’t understand their organizations are going to prioritize initiatives and buy agents that increase disorganization without driving results. Here’s why: Agents drive real value when they conform to enterprises’ workflows, which means you need a clear understanding of what your team does today and what’s needed to improve. Without that understanding, leaders will thrash between possible solutions, and in the worst case, they will end up purchasing solutions that amplify existing issues. That means agents should be implemented where there’s a clear picture of the current state of play as well as a deep understanding of what good looks like. If you don’t know what good is, an LLM isn’t going to tell you.
Historically, engineering teams have had to focus aggressively before resources were scarce. There’s always been a million things that could be worked on – niche features, tech debt, performance optimization, internal tools, etc. – but cycles to only work on a few of them. To some extent, coding agents have alleviated that pressure, and teams can now spin up new tools (and certainly flashy demos!) faster than before.
What happens after that still requires strong focus. Every prototype can’t be taken to production, because the process of refining and maintaining that prototype paradoxically has a much higher opportunity cost than before. It is definitely true that engineering teams are faster than before, but that means time spent on productionizing something that could be bought would be better spent improving core functionality that’s in the business’ wheelhouse.
This is where the disconnect happened in our conversation last week: The VP was anchored on a demo he’d been shown that relied on hard-coded workflows and runbooks and wouldn’t generalize to the complexity of their stack. On top of that, he also wasn’t aware about the manual workflows the team currently followed, which would make the demo even more difficult to generalize.
To be clear, the VP wasn’t gratuitously trying to tear us down or being blithely cynical. He was operating the way that senior leaders at large enterprises have typically operated for decades. But with coding agents and flashy demos, it’s easier than ever to convince yourself that the world looks dramatically different than it does in reality.
What happens when you’re misled? You make decisions about what you think needs to be done rather than what actually needs to be done. We started a conversation last summer with a mid-sized tech company that wanted to use an AI SRE to deal with thousands of PagerDuty alerts a week. They evaluated us and one other vendor, but by the end of the year hadn’t yet come to a decision because the CEO suggested they do two more evaluations. Those evaluations last until the end of Q1, and by that point, an engineer had built an in-house solution the CEO got excited about because it dogfooded their internal product. Eight months of evaluation time was wasted, and now there are two engineers on the hook for maintaining a side project.
In each of these cases, an unrealistic ideal became a blocker for a good solution.
A lack of clarity means that your agent implementations lack a clear target for what’s actually being improved. In many cases, you can very much accelerate speed with agents – but speed without quality mostly just creates noise.
Take coding as an example. Adopting coding agents certainly means that you’ll have cleaner syntax and better comments. Your engineers might even prompt the agent to write more tests! That’s not going to make better software in general. Coding agents aren’t going to magically give you high-quality design, thoughtful system architecture, or a culture that prioritizes catching bugs before they ship. If your team is in the habit of not ensuring their implementations closely match product requirements and validating that edge cases are properly handled, that problem is likely going to be amplified with AI.
One of the first POCs we did in the AI SRE space was with a traditional software vendor – the POC ended up failing. The organization was excited about AI as a strategic initiative for them, but their stack simply was not set up for integration with agents – their existing debugging workflows were full of manual processes that required humans to use credentials with root access to access individual deployments and pull the relevant data. They were understandably hesitant about giving an agent this kind of information, but that meant the agent was operating with limited context – and ultimately, its recommendations were useless. Their processes were immature, and the agent created more noise.
The lesson is simple: Know what good looks like before you adopt an agent. If not, you’ll be impressed by the shiny new clothes that turn out not to be real.
This changes the way that software is evaluated and purchased in a few key ways. Teams building agents must be flexible. As we talked about last week, every AI startup needs to be oriented towards showing value quickly and orienting pricing and commercial terms to quick adoption. If you’re building a genuinely good product, that’s how you’ll avoid leading your customers into the traps that we discussed above. Trust us, we’ve made these mistakes!
There are also clear lessons for buyers. First, it requires trusting your team. As a leader, you need to enable your team to make the call about what agents are worth spending time on and what will deliver the results that align with your priorities. The responsibility to make decisions about which agents are good or bad has to rest with the people doing the work.
Second, there will be trial and error – and you’ll have to get used to it. We’ve heard anecdotes about major law firms buying all three of the top AI legal agents and letting their organizations use them in parallel with the intention of picking one after a year or two. While this might sound wasteful, it’s a wiser approach than being beholden to one implementation up front. It’s better to learn now rather than a year from now.
Finally, we find the idea of flat organizations particularly interesting. Brian Chesky talked about this in a recent interview: The closer you can get to the ground truth – the actual work being done and the problems being faced – the more likely you are to make intelligent decisions about what’s working and what’s not. The more abstraction you have between yourself and reality, the more likely you are to think that your team is light years ahead or behind where it actually is.
Ultimately, focus is what will win. AI that lives at the intersection of possible and good is what will create value.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.