When early motor vehicles began appearing on British roads, the law treated them less like everyday transport and more like industrial machinery. The Locomotives Act of 1865 - remembered as the “Red Flag Act” - limited speed to 2 mph in towns and 4 mph in the country, and required a three-person crew, including someone walking ahead carrying a red flag.
The point wasn’t to improve the vehicle. It was to slow a new kind of movement down until it fit the old world.
It was fear of the unknown. A new capability entered public space, and the first response was supervised movement.
AI is entering operational space the same way. Most organisations are deploying AI with its own red flag: useful, supervised step by step, and constrained to advising humans rather than executing work.
Autonomous execution does not mean completing a sentence. It means completing a piece of work end to end - across systems - to a finished outcome with accountability.
If you remember one line, make it this:
Autonomous execution is not a feature. It is the new operating model.
It matters because execution across messy systems is where cycle time, cost-to-serve, and operational risk quietly accumulate.
The decisive moment is not the invention. It is the redesign that follows.
Consider electricity in manufacturing. Early factories electrified by swapping steam for motors, but kept the old architecture of belts, shafts, and layouts organised around power transmission. The productivity leap came later, when factories redesigned around what electricity enabled: unit-drive machines, flexible layouts, and work that flowed instead of queued. It took thirty years!
Railroads and container shipping tell the same story: scale depends on coordination standards and predictable units.
Rail networks needed shared time standards to coordinate schedules across cities and reduce confusion. In the U.S., railroads implemented standardized time zones in the 1880s, and public timekeeping followed.
Containerization wasn’t just faster ships. It was a standard unit that turned loading, unloading, insurance, theft, and scheduling into something predictable and industrial. Costs collapsed because friction collapsed.
Those examples are a map. They tell us what comes next for AI.
AI’s first phase was assistance.
Assistance improves output. It helps people write faster, search better, summarise more cleanly, and draft more confidently.
But the operating model stays the same. A person sits at the centre, receives suggestions, and then does the work: moving information between tools, reconciling inconsistencies, coordinating with others, and pushing tasks over the finish line.
The second phase is autonomous execution. Not full autonomy in the abstract, but governed autonomy in the real world.
In enterprise terms, this is the shift from systems of record (software that stores truth) to systems of action (software that can reliably move that truth through workflows, across tools, to a finished outcome - with accountability).
Systems of action do not just inform decisions. They carry work through, including the messy exceptions that make up real operations.
The goal is not to automate only the happy path. It is to resolve most variance inside guardrails, and escalate only the few cases that truly require human judgement.
The constraint in most organisations is not ideas. It is execution across imperfect systems.
Most companies run on a patchwork of legacy applications, portals, spreadsheets, email threads, PDFs, ticketing systems, and processes that evolved through years of practical compromise.
Even when processes look standard, exceptions are the rule: missing fields, ambiguous documents, late updates, policy nuance, edge cases, local variants, and systems that disagree about what is true.
So much capacity gets consumed by work that looks minor but is strategically expensive: chasing inputs, rekeying data, reconciling discrepancies, updating multiple systems, nudging stakeholders, documenting decisions, and creating proof that the work was done correctly.
This is why transformation is so often described as an execution problem, not a strategy problem. McKinsey has repeatedly highlighted that large-scale transformations fail roughly 70% of the time.
And it’s why “single system” programs routinely underdeliver. Modern ERP can be invaluable - but deployments, upgrades, and replacements are still prone to major disruption, cost blowouts, and long recovery cycles, with highly visible failures even in sophisticated organisations.
For the core, engineering remains non‑negotiable. Financial controls, high-volume transactions, security boundaries, and systems of record must remain authoritative.
But much operational work lives at the edges - where requests arrive in uncontrolled formats and counterparties change behaviour without warning. In those conditions, perfect definition becomes a moving target. Perfect data becomes a multi‑year program that often costs more than the work is worth.
We’ve tried to close the execution gap for years.
We pushed workflow automation inside systems of record. We built integration programs. We adopted iPaaS. We rolled out RPA. We built enterprise data lakes to “centralize truth.” These tools created value - especially in closed-world processes.
But they also exposed a pattern: most automation approaches assume stability.
Even integration platforms acknowledge the reality: every system has different workflows, protocols, and data formats, and integration still requires mapping fields and handling compatibility and exceptions.
RPA tried to bridge the gap by mimicking human clicks and keystrokes. It can work in stable flows, but scripted automation inherits brittleness. When a portal changes or a process deviates, the bot breaks. As practitioners put it: these solutions often require more maintenance than traditional IT.
And scaling is where it commonly stalls. McKinsey was already warning years ago about organisations getting “burned by the bots” when they tried to move from demos to bot armies.
Deloitte’s surveys have consistently found “process fragmentation” and lack of end‑to‑end coherence among the top barriers to scaling automation.
Even data platforms show the same lesson. Data lakes promised flexible, centralized, self‑service truth - but “many of the promises… have not been realized” without transactions, quality enforcement, governance, and performance optimizations; the result, by Databricks’ own description: enterprise lakes turning into “data swamps.”
The recurring theme isn’t that the tools were “bad.” It’s that they were designed for a world that was more closed than reality.
Some work is closed-world and some is open-world.
Closed-world work has stable interfaces, structured inputs, enumerated exceptions, and deterministic outcomes.
Open-world work is what most organisations quietly run on: ambiguous inputs, changing screens, incomplete information, shifting rules, and novel exceptions.
Scripted automation struggles in open-world conditions because the moment the environment changes, the system has no interpretation layer - only failure modes.
AI agents change the calculus because they can interpret variability rather than collapse when it appears. They can:
observe state
choose actions toward an outcome
verify results
recover when the environment deviates
escalate when confidence or authority runs out
They introduce a control loop: perceive → act → check → correct → escalate.
That loop is what makes autonomous execution possible in open-world conditions.
But it also explains why hype collapses without redesign. Even CEOs running multi‑billion‑dollar AI companies have publicly cautioned that you can’t “just unleash the agents, and it just works” - it takes engineering, evaluation, and operational discipline.
And the failure mode is now measurable. Gartner has predicted that over 40% of agentic AI projects will be canceled by the end of 2027 due to costs, unclear business value, or inadequate risk controls.
This doesn’t argue against autonomous execution. It argues for taking governance seriously as part of the operating model.
Here is the catch many organisations will learn the hard way. The hardest part of open-world work is not clicking the UI. It is knowing what should happen when inputs are ambiguous and consequences are real.
That knowledge is vertical. It lives in industry-specific policy, regulation, commercial norms, and risk appetite.
It also lives as domain skills and tools: specialised checks, calculations, reference data, decision frameworks, and workflows that encode how the industry actually works.
The systems of action that matter will not be generic copilots that are “good at language.” They will be vertical operators: grounded in a company’s rules, shaped by its exception patterns, equipped with domain tools, and constrained to act safely inside the organisation’s control environment.
The most effective posture is not agents versus integration. It is a stable digital spine plus an agentic edge.
The spine remains engineered: clean APIs where available, data contracts, access controls, event logs, and systems of record that stay authoritative.
Systems of action at the edge operate across surfaces that will not be perfectly integrated for years: portals, email, documents, tickets, spreadsheets, and UI workflows.
They stay constrained by the spine’s policies.
They also become a discovery mechanism. They show where variability is highest, where exceptions concentrate, and where an integration or product change will produce the highest leverage.
The moment you allow software to act, the problem stops being “Can it do the task?” and becomes “Can we govern the task?”
Autonomous execution is not just capability. It is accountability.
Autonomy becomes valuable only when it becomes governable.
Autonomy is not something you turn on. It is something you make safe.
That requires redesign around three ideas:
First: redesign work around outcomes, not activities.
Systems of action need explicit outcomes, a definition of done, acceptance criteria, and boundaries. Without that, you do not have a process. You have a collection of habits.
Second: redesign the control environment for machine action.
Treat AI like a new class of operator, with identities, roles, permissions, segregation of duties, and audit trails. The hidden blocker is often operational plumbing: access control, logging, monitoring, and flight-recorder visibility into what happened, when, and why.
Third: redesign management for exception-driven work.
When execution becomes cheaper and faster, management shifts from supervising activity to managing exceptions, trade-offs, and quality. Cycle times compress. Queues shrink. Decisions become the bottleneck.
A useful parallel is the disappearance of elevator operators. Autonomy scaled only when safety mechanisms, standards, inspections, and interfaces made operation safe for ordinary people.
The point is not fewer people. It is better use of people.
In a world of systems of action, the operational layer runs continuously. Cross-system work happens overnight and across time zones. Teams spend more time on judgement, negotiation, customer relationships, and improvement—and less time copying, checking, and chasing.
Assistance helps you produce better output. Autonomous execution helps you finish work to an accountable outcome.
The third phase is initiative.
Initiative is when systems of action do not just respond to requests. They start the right work before the queue forms. They watch for signals, detect what is missing, anticipate next steps, and prepare decisions for humans to approve.
Not because the system is “smart,” but because the operating model has been redesigned so proactive work is safe, legible, and auditable.
If this sounds ambitious, remember the red flag. It was a transitional rule for a world that had not yet adapted. Roads improved. Standards emerged. Drivers became trained. The law changed.
Work is going through its own version of that transition. Many organisations will keep AI behind a flag for a while, and that caution is understandable.
But history suggests the long-term answer is redesign: systems that make autonomy safe, accountable, and useful at scale.
Autonomous execution is not a feature. It is the new operating model.
This post is public so feel free to share it. It will be much appreciated. Thank you!
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.