RSS Amplifier

Eugene Vyborov Blog (AI future, entrepreneurship) · Aug 13, 2026

The org-chart mistake that kills agent fleets

0
Sign in to vote or save

Eugene Vyborov · Eugene Vyborov Blog (AI future, entrepreneurship)

  • Every company deploying AI agents at scale hits the same question: how does the agent fleet relate to the organization? The two popular answers - “give every department its agent” and “replace the org chart with fluid work charts” - fail in opposite directions.

  • Our design at Ability AI: mirror the company’s accountability structure, not its org chart. Each agent is owned by the department head whose outcomes it serves. Agents coordinate operationally, but authority over any agent always traces to a named human.

  • Mirror the boundaries, not the depth. Agent-to-agent coordination collapses the overhead that made middle-management layers necessary - copying those layers into your fleet recreates the problem agents solve.

  • Cross-department work is where fleet designs quietly fail. A 40-year-old finding from disaster research explains why - and what to build instead.

  • We are running this architecture on ourselves, with named owners and a measurement plan. I will report what breaks.

Ethan Mollick put it plainly this year: “Organizational design for agents is hard, benchmarking agents working in concert is hard. Together, this is the next critical frontier for making AI matter in economically valuable tasks, and we really don’t know very much about it.”

He is right about the stakes. Gartner projected in mid-2025 that over 40% of agentic AI projects will be canceled by the end of 2027 - citing rising costs, unclear business value, and inadequate risk controls. By May 2026 they had sharpened the diagnosis: 40% of enterprises will demote or decommission autonomous AI agents by 2027 because of governance gaps discovered only after production incidents. Read those two predictions together and the pattern is clear. Agentic projects are not failing because the models are weak. They are failing because nobody decided who owns what.

This is an organizational design problem, not an IT problem - Mollick again: decisions about AI in your organization “are increasingly organizational design and strategy decisions, not IT choices.”

We build and run agent fleets for a living at Ability AI - our platform, Trinity, is the runtime our own company operates on. So we do not get to leave this question open. Here is the answer we are betting on, and the mechanism behind it.

Wrong answer #1: the fleet-as-org-chart. Take the company org chart, deploy one agent per box. Finance gets a finance agent, HR gets an HR agent, and the diagram looks satisfyingly complete on day one.

The failure mode is precise: agents deployed to mirror an imagined organization rather than to own real, partitioned work. An agent with no genuinely partitioned job is not automation - it is triage waiting to happen. Every agent you deploy spends your scarcest resource: operator attention. Reports to read, approvals to answer, drift to adjudicate. An agent that exists because the org chart had a box, rather than because work genuinely partitions, consumes attention and returns none.

Wrong answer #2: burn the org chart. Microsoft’s Work Trend Index champions the opposite move: replace “rigid org charts” with fluid, outcome-driven “Work Charts” that flex with the needs of the business, assembling whatever mix of humans and agents each task needs.

The work-chart picture gets something real: task execution genuinely is becoming fluid. But it quietly deletes the one thing the org chart actually encodes that matters: accountability. When an autonomous agent makes a $50,000 mistake at 3am, “the work chart flexed” is not an answer to the question “who owns this?” Fluid teams are a fine execution pattern. As an accountability structure, fluidity is just a gap with good branding - and Gartner’s demote-and-decommission prediction is what that gap looks like in production.

The two wrong answers fail in opposite directions. The first copies structure without asking whether the work is real. The second dissolves structure without asking where responsibility lands. The design question is what to copy and what to dissolve.

Here is the architecture we are running - three rules, summarized in the diagram “The three rules that mirror accountability”:

  • Ownership follows accountability. Each agent is owned by the department head whose outcomes it serves. This prevents orphan agents and the central-AI-team bottleneck.

  • Mirror boundaries, not depth. Departments define the ownership domains; the agent tree stays shallow. This prevents rebuilding the middle layers agents make obsolete.

  • Cross-team work runs through contracts. Agents coordinate on shared, published surfaces; humans arbitrate tight coupling. This prevents coordination failures at department seams.

Rule 1: Ownership follows accountability. The owner of the marketing agents is the head of marketing - not the CTO, not an AI center of excellence, not “the platform team.” The department head already holds the decision rights over that domain and already answers for its outcomes. Placing agent ownership anywhere else splits decision rights from accountability, and split accountability is exactly the failure Gartner’s numbers describe. The mid-2026 enterprise consensus has converged on the same conclusion from the failure side: a centralized AI team that tries to own every use case becomes a bottleneck, while business units that own delivery and outcomes ship. The center of excellence survives only as a platform layer - it owns infrastructure, evaluation, governance tooling - never business outcomes.

Rule 2: Mirror boundaries, not depth. More on this in the next section, because the mechanism deserves it.

Rule 3: Cross-team work runs through contracts. More on this too - it is where fleet designs quietly die.

One invariant underneath all three rules, stated in our internal fleet-design principles: capability self-extends, trust does not. An agent may learn new facts, write new skills, and grow its competence inside the scope it was given. But trust - credentials, write access, permission to act on the world, authority over another agent’s behavior - is always granted by a human and never self-extended. Agents consume trust grants; they do not create them. This is what “the department head owns the agent” means mechanically: the owner is the human all of that agent’s trust traces back to.

This also settles what “agents reporting to agents” can mean. Agents in our fleet delegate to each other, fan work out, and roll status up - operationally, constantly. But no agent holds standing authority over another agent. Authority is not a thing agents can hold; it lives with the named human owner. The hierarchy is an accountability map, not a chain of command among machines.

In 1933, the management consultant V.A. Graicunas published the argument that capped human span of control for ninety years: as direct reports grow linearly, the coordination relationships a manager must track grow geometrically. Six direct reports produce over 200 potential relationships. The classic 5-to-7 span limit, and therefore the depth of every large org chart, follows from that arithmetic. Much of the middle of a human organization exists for one purpose: routing information between people who cannot all talk to each other.

Agents change the arithmetic, not the principle. When coordination moves agent-to-agent - delegation, status, handoffs running inside the fleet - the combinatorial burden leaves the human’s head. The span cap that made deep hierarchies necessary loosens.

The 2026 restructuring wave shows this happening precisely where the theory predicts. Bayer cut its organizational layers from roughly twelve to six or seven; average span of control reached 14, with some managers carrying 90 direct reports. When Cloudflare cut roughly 20% of its workforce, CEO Matthew Prince said the vast majority were “measurers” - middle management, finance, legal, internal audit. Not uniform delayering: the cuts concentrate on the information-routing layer, because that is the specific work the agent layer absorbs.

Here is what this means for fleet design: if you copy your org chart’s depth into your agent fleet - manager agents supervising supervisor agents supervising worker agents - you are rebuilding, in software, the exact layer agents exist to dissolve. The boundaries of your org chart encode something durable (who is accountable for what). The depth encodes something obsolete (how many relationships a human brain can track). Copy the first. Skip the second.

In practice: one orchestrator per department domain, workers under it, and almost nothing above. The human org may be five layers deep. The agent fleet rarely needs more than two.

In Normal Accidents (1984), Charles Perrow studied why ships collide. A ship, he observed, is “the preeminently centralized human system” - one captain, near-absolute authority, and it works, because the system has a single bounded scope. The failure arrives at the boundary: two ships approaching collision are suddenly one tightly-coupled system, but each still carries its own sovereign authority, and no superordinate authority exists above both. The maritime “rules of the road” - a pre-negotiated protocol meant to substitute for the missing authority - fail exactly when they are needed most: ambiguous, shaped by litigation rather than operations, impossible to parse in the seconds before impact. Perrow’s contrast case is aviation, where a ground controller holds precisely the superordinate role the sea lacks - which is why the same near-collision resolves differently in the air.

Now replace the ships. Your delivery department’s agent and your marketing department’s agent are each well-governed - each owned by its department head, each operating cleanly inside its scope. Then a launch puts them on one deadline, tightly coupled, and each answers to a different sovereign. Every fleet design I have seen waves at this moment with some version of “the agents will coordinate.” That is the rules of the road. It fails the same way.

Two patterns actually work, and they map to Perrow’s two cases:

  1. Keep cross-department interfaces loosely coupled by default. Agents coordinate through published records and shared surfaces - facts, trackers, events - never by reaching into each other’s workspaces. Our internal principles doc states the incentive rule that makes this real: the shared surface must be the path of least resistance, not just the path of least coupling. Rules break under time pressure; incentives do not. If publishing takes a ceremony and reaching into a peer agent’s files takes one command, the rule will be broken precisely when it matters most.

  2. For genuinely tight coupling, name the air traffic controller in advance. When two departments’ agents must jointly deliver against one deadline, a human arbiter with authority over both is named before the work starts - not discovered during the incident. In a small company that is often the CEO wearing a chief-of-staff hat. The title does not matter. The advance naming does.

The general rule: agent-to-agent collaboration is an execution surface, never an authority surface. The moment a cross-team conflict needs deciding, it must already be clear which human decides.

Ability AI is the testbed for this architecture. The ownership map, as it stands: I own the chief-of-staff fleet plus product management and the management of Trinity itself. Alex owns the delivery agents. Romaine owns the finance agents. Christina owns the marketing agents.

I will be honest about the shape of this: it’s not great, but it’s doable. I hold three of the loci myself - a textbook implementation would distribute them further. But the distribution is real and staffed, which is what makes it a testbed rather than a diagram. Every agent in the fleet has a named owner whose department outcomes it serves, and every trust grant traces to that owner. The point of running the architecture on ourselves is not that our org is exemplary. It is that the failure modes show up here first, at a scale where they are cheap to observe.

How do we know if it is working? Our fleet-design principles - a document co-written, fittingly, by two agents that live inside the fleet - end with a four-question diagnostic. It doubles as an audit you can run on any fleet, including one you are planning (shown in the diagram “The four-question fleet audit”):

  1. What does each agent own - and what happens to work that doesn’t obviously belong to any of them? This tests whether you designed an accountability layer or just listed agents. The gaps between agent scopes are where unowned work accretes.

  2. For any fact more than one agent acts on: who endorsed it, and if it were wrong right now, how would you find out? This tests whether shared knowledge has provenance, or agent inferences are quietly becoming “established facts.”

  3. If agent A produced a bad output that agent B depended on, where would the error stop? This tests fault isolation at the seams - the exact place Perrow’s gap lives.

  4. What decisions did the fleet surface to a human last week - and is that number trending up or down? This is the health metric. A fleet surfacing ever more decisions is accumulating human dependency, not autonomy.

Question 4 is the one I would put on a dashboard. Per owner, per week: how many decisions did your agents escalate to you? Health is that number trending down - routine decisions migrating into wider zones of autonomous action as agents earn them - leaving only genuine judgment calls in the queue. A fleet that gates everything is an approval queue wearing an org chart. A fleet that gates nothing is a liability. The trend line between them is the measurement.

That is also my commitment for this series: we are instrumenting these numbers on our own fleet, and I will publish what they show - including the parts that fail.

The org-chart question for agent fleets has a precise answer, and it is neither of the popular ones:

  1. Do not copy the org chart. Boxes are not evidence of partitioned work, and every unnecessary agent spends operator attention.

  2. Do not burn it either. The org chart encodes the accountability structure - the one thing your fleet cannot function without.

  3. Mirror the accountability boundaries. Each agent owned by the department head whose outcomes it serves; all trust traces to the owner.

  4. Skip the depth. Agents dissolve the information-routing layers; do not rebuild them in software.

  5. Treat department seams as the primary design surface. Loose coupling through published records by default; a pre-named human arbiter wherever coupling is tight.

The companies getting agentic AI right in 2027 will not be the ones with the best models. Everyone has the best models. They will be the ones who answered the ownership question before production answered it for them.

If you are structuring an agent fleet right now - or arguing with someone about whether the org chart should survive - I want to hear where your version of this breaks. Reply or write to me; the failure stories are worth more than the wins.

  • Ethan Mollick on organizational design for agents as the next frontier - X, June 2026

X avatar for @emollick

Ethan Mollick@emollick

Decisions about how to use AI in your organization are increasingly organizational design and strategy decisions, not IT choices: How do you integrate agents into your firm? What intelligence will you outsource? What are the boundaries of the firm? What is the role of people?

X avatar for @random_walker

Arvind Narayanan @random_walker

The new Claude Tag feature seems extremely useful, but at the same time, a dangerous bargain for enterprises because of the pricing model and the risk of lock-in. The four big changes together mean that you interact with Claude as a coworker instead of a tool (the same Claude

3:21 PM · Jun 24, 2026 · 77.9K Views

52 Replies · 80 Reposts · 640 Likes

  • Gartner, Applying Uniform Governance Across AI Agents Will Lead to Enterprise AI Agent Failure, press release, May 26, 2026 (https://www.gartner.com/en/newsroom/press-releases/2026-05-26-gartner-says-applying-uniform-governance-across-ai-agents-will-lead-to-enterprise-ai-agent-failure)

  • Gartner, prediction that over 40% of agentic AI projects will be canceled by end of 2027, June 2025

  • Microsoft, 2025 Work Trend Index: The Frontier Firm Is Born - the “Work Chart” thesis this essay argues against (https://www.microsoft.com/en-us/worklab/work-trend-index/2025-the-year-the-frontier-firm-is-born)

  • Fortune, AI agents are flattening corporate hierarchies, June 9, 2026 - the Bayer and Cloudflare cases (https://fortune.com/2026/06/09/ai-agents-flattening-corporate-hierarchies-companies-managers-develop-new-playbook/)

  • Charles Perrow, Normal Accidents: Living with High-Risk Technologies, 1984 - the authority-interface gap

  • V.A. Graicunas, “Relationship in Organization,” 1933 - the span-of-control arithmetic

  • Ability AI internal: Principles of Fleet Design (trinity-pm x Cornelius, August 2026) - the invariant, the six levels, and the one-minute test

Related concepts: Department-Mirrored Fleet Architecture; Ability AI as the Live Testbed for the Department-Mirrored Fleet; Structure Organizes Around the Accountability Locus; Managerial Span Bifurcates - Nominal Widens, Consequential Holds; The Authority-Interface Gap; Agentic Org Structure Bifurcates by Liability and Verifiability, Not Size.

No posts

Read the original on eugenevyborov.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.