RSS Amplifier

Product in Practice · Dec 11, 2025

Why 90% of production incidents are process failures, not technical ones

0
Sign in to vote or save

Mark Henzi · Product in Practice

I’ve been managing engineering teams for over a decade, and after hundreds of incidents across multiple companies, I can tell you this with confidence: the vast majority of production incidents reveal process gaps rather than technical failures.

The technical systems usually work fine. The process around them doesn’t.

This isn’t about needing more documentation or better runbooks. It’s about fundamental gaps in how teams coordinate work, make decisions, and maintain visibility as they scale.

When you hit 50 engineers, you’re dealing with 1,225 potential communication paths. At 100? That jumps to 4,950 paths. Every new hire doesn’t just add one more person to coordinate with—they add exponentially more complexity to your communication overhead.

Most teams try to solve this with more process. More meetings. More documentation. More approval gates.

But here’s what we’ve seen: when you add process without addressing the underlying communication problems, it becomes a patch rather than a solution.

When you layer on process without addressing the underlying problems, here’s what actually happens:

Simple features take longer than before, with estimates becoming increasingly inaccurate. Teams can’t commit to deliverables beyond the current sprint because they have no confidence in their ability to predict when work will actually ship.

The worst part? By the time leadership realizes the engineering team has become the constraint, you’re already 3–6 months behind on critical projects.

Let me share some specific examples from our journey at Atono:

The QA Bottleneck: We could actually see things getting bogged down in the test column. Our developers were moving so fast, and the demands on QA and automation were so high that things were getting stuck.

This wasn’t a skills problem or a tools problem. It was a visibility problem. Without clear metrics showing cycle time per workflow step, the bottleneck remained invisible until it became critical. As I shared in our Inside Atono web series, “it was just sheer volume.”

The Swarming Problem: We had conflicts with multiple devs swarming on a particular story or initiative. We had to find ways to prevent that and ease it when it occurred.

When multiple developers work on the same story without clear ownership or coordination, productivity doesn’t increase—it plummets. Everyone assumes someone else is handling critical pieces or people end up working on the same thing at the same time. Either way, context gets lost between handoffs because no one knows where they’re going.

Teams track that work is “in progress” but have no visibility into whether it’s actually moving. Without granular visibility into each step—coding, review, testing, deployment—you can’t identify where work gets stuck.

At Atono, we break down cycle time per step in the in-progress categories. As long as something is in progress, it is gaining cycle time. This granular tracking revealed something surprising: our stories, averaging 4 points (our sweet spot between small and medium), were taking wildly different amounts of time depending on which step they were in. One bug had been in development for almost 20 days, 10 times our normal average.

As teams grow and specialize, critical context gets lost. The frontend team makes assumptions about API behavior. Backend teams don’t understand user workflows. QA discovers issues that could have been caught with better communication upfront.

We address this through shared artifacts and early collaboration. Stories include explicit acceptance criteria that everyone reviews. Technical decisions get documented in the story itself. When context travels with the work, handoffs become smoother.

Cycle time measures delivery duration but fails to capture volume or complexity. Teams optimize for velocity without understanding impact. They ship features faster, but can’t tell if those features solve real problems.

We’ve connected feature engagement directly to stories at Atono. Developers can see within a story that’s been shipped how often it’s being used across production.

This connection between delivery metrics and outcome metrics changes the entire conversation. Teams stop asking “How fast can we ship this?” and start asking “Should we be building this at all? Should we iterate further or move on?”

At my previous company, we regularly ran simulated incidents to sharpen our response skills. This proactive approach—often called “Failure Fridays”—is common among teams that take reliability seriously.

This wasn’t about testing technical systems. It was about testing human systems:

  • Who makes decisions when the primary on-call can’t be reached?

  • How do teams communicate when Slack is down?

  • What happens when the person who knows the system is on vacation?

It helped us stay on our toes so that when there was a real incident, we weren’t all wondering what it was we needed to do. It also just helped us keep improving and bridging those gaps and making things better for everyone.

You can’t manage your way out of this. You have to architect your way out. But not just technical architecture—process architecture.

Here’s what I’ve found actually works:

Make the Invisible Visible: Track cycle time at each step of your workflow. When we implemented this at Atono, we discovered bugs sitting in states we didn’t even know were problems. Without that visibility, they would have remained “in progress” indefinitely.

Create Clear Escalation Paths: Even when we minimize meetings, critical processes need to be embedded in the team’s muscle memory—who owns each service, when to escalate versus troubleshoot independently. These patterns become automatic through repetition, not documentation.

Connect Work to Outcomes: Stop measuring just velocity. We track feature engagement alongside cycle time. We want to see that we’re getting more stories done, but we also want to see that we’re not inflating our story sizes. We’re not just shipping features that inflate the product without maintaining quality.

Accept Process Evolution: The process that works at 10 engineers fails at 50. What works at 50 fails at 100. Every now and then, you have a downstream or upstream thing that can become a technical issue, but even that can be worked around through process evaluation and refinement.

One place where these principles come together is in our “Shoulder Surf“ sessions.

While most teams wait until code is complete to get feedback, we have developers record or present features while they’re still in progress. This early visibility lets us identify blocking issues, clarify requirements, and catch misalignments before they’re expensive to fix.

It’s a simple practice that prevents a common failure pattern: building the wrong thing perfectly.

Production incidents will always happen. But when 90% stem from process failures rather than technical ones, the solution isn’t better monitoring or more robust systems. It’s better visibility, clearer ownership, and processes that evolve with your team.

Start with visibility. Track cycle time per workflow step. Make work states explicit. Connect delivery metrics to outcome metrics. As I like to say, it’s important we’re paying attention to what’s on our boards, but it’s even more important we’re listening to our teams and how things are really going.

Then practice. Run incident drills. Test your escalation paths. Find the gaps before they find you.

Because the next time production goes down, it probably won’t be a technical failure. It’ll be a process that seemed fine until it wasn’t.

This post is public, so feel free to share it.

Share

Read the original on atono.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.