This is Issue 19 of The AI Agent Economy — a five-part series on the dependency layer: the infrastructure AI agents cannot function without.
There was a period, not that long ago, when monitoring meant watching the network. Pings, packet loss, uptime, throughput. If the network was healthy, the system was healthy — and for a while that was true enough, because the application was a single program on a single machine and the network was the fragile part.
Then applications became distributed, and "the network is fine" stopped meaning anything useful. You could have perfect packets and a broken product. A new category had to be invented to watch the layer above: application performance monitoring. Datadog is worth roughly $85 billion today because that category turned out to be real. New Relic was taken private for around $6.5 billion on the same insight.
We are standing at the identical moment again, one layer up. And almost nobody is pricing it that way.
Agents do not just execute. They decide. And monitoring a decision is a fundamentally different problem from monitoring a request — different enough that it will be a different category, owned by different companies, the way APM was never going to be won by the network-monitoring incumbents.
That category does not currently have an owner. That is the whole opportunity.
Traditional monitoring answers questions with objective answers. How long did it take? Did it return an error? How many tokens? Every one of those has a correct value the infrastructure can produce on its own.
Now take the questions that actually matter for an agent. Did it skip a step in the security check? Did it hallucinate a figure in the financial report? Did it choose the wrong tool for the job? None of these throw an exception. The latency is fine. The error rate is zero. The metrics are green and the output is wrong.
This is what makes agent failure so unlike software failure: a confidently wrong agent is indistinguishable from a confidently right one — until the damage surfaces, which is typically weeks or months later, usually via a customer, a regulator, or an audit. There is no page at 3am. There is no incident. There is a slow accumulation of decisions nobody checked, and then a reckoning.
Every layer of computing has needed its own monitoring category, and each time the incumbent below it assumed the new layer was a feature rather than a market.
Network monitoring did not become APM. APM did not become observability. Each transition looked, at the time, like an extension of the existing product — and each time the position was taken by a company built specifically for the new layer, because the questions being asked were different enough that the old data model could not answer them.
The current incumbents are running that playbook exactly. Datadog launched LLM Observability in 2024. New Relic launched AI Monitoring the same year. LangSmith, Langfuse, Arize AI and Weights & Biases have built tracing and evaluation tooling. These are all real, useful products, and every one of them monitors the performance of agent systems. They trace the calls, count the tokens, catch the errors.
None of them monitors the quality of the decision. That's not a criticism of the products — it's a description of where the line currently sits.
The most agent-native entrant I can find, AgentOps, is still at pre-seed. Set that against Datadog's $85 billion and you have the shape of the gap: an incumbent worth eighty-five billion dollars for answering the older question, and roughly nothing capitalised against the newer one.
When the market matures enough to recognise the distinction — and it will, exactly as it eventually distinguished network monitoring from application monitoring — the companies that own agent-native decision monitoring will be the Datadogs of the agent era. That seat is unclaimed today. It sits inside a dependency layer I expect to clear $100 billion by 2030 (PRED-003).
Here is where monitoring diverges from the other four layers, and why I think it is the most underestimated of them.
Orchestration, identity, trust and attestation are all, ultimately, infrastructure problems with infrastructure answers. Monitoring is only partly that. Judging whether a decision was good requires context that no tool fully has — and for the next several years, a large share of that judgment will be done by people.
That is my dated call, and it is the one in this series I hold with the most confidence: agent oversight becomes the fastest-growing white-collar job category by 2030, from a 2026 baseline of very close to zero, while traditional knowledge-work hiring falls 25% (PRED-014). Those two things are the same event seen from two sides. The work does not disappear; it moves up a level, from doing the task to checking the thing that did the task.
So the monitoring layer is a software market and a labour market, arriving together. Most forecasts pick one and miss the other.
Three tells. Watch whether "decision quality" appears as a distinct product line rather than a feature bolted onto a tracing tool — the naming will tell you when the category becomes legible. Watch the incumbents' acquisitions, because the fastest route to owning this is buying the agent-native entrant while it is still pre-seed, and that window is open now. And watch job titles: the moment "agent operations" or "agent oversight" shows up as a standard rung on a career ladder rather than a task someone absorbed, the labour half of this prediction has started.
Everyone building in this space is trying to make agents more capable. Very few are building the thing that tells you whether a capable agent is doing the right thing — and capability without that is not an asset, it is an unpriced liability that compounds quietly.
The uncomfortable version: most organisations running agents in production today could not tell you if quality dropped by a third. Not because they're careless. Because nothing they own would report it.
You cannot scale a workforce you cannot supervise. That is true of people, and it is about to be true, expensively, of agents.
Related on atin-agarwal.com: PRED-014 — agent oversight becomes the fastest-growing white-collar job category · all 15 predictions, with their falsification triggers
The AI Agent Economy — Why the Next Trillion-Dollar Industry Has No Employees is out now. This series expands Chapter 2: the dependency layer. Next issue closes Season 2 — attestation: proving what an agent did in the physical world.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.