On many university campuses, you will find “desire paths” cutting through the grass between sidewalks. These informal trails appear because the official route was too slow, too indirect, or failed to reflect how people actually move through the environment.
Product designers understand this principle well. Features do not succeed simply because they exist. Systems succeed when they align with real human behavior under real conditions. Security teams learned a version of this lesson years ago as overly burdensome controls repeatedly pushed users toward faster, informal workarounds.
Governance behaves the same way.
Organizations spend enormous amounts of time documenting how decisions should flow, how risk should be managed, and how control should be maintained across the enterprise. Governance frameworks define ownership structures, policies establish escalation paths, and resilience plans outline how disruption should be contained and resolved. Under normal operating conditions, those structures often appear coherent and sufficient.
Incidents test whether they are usable.
During a ransomware attack, a cloud outage, or an AI-enabled workflow failure, organizations frequently discover that people stop following the carefully paved paths represented in governance diagrams. Teams improvise around bottlenecks, bypass approval chains that move too slowly, rely on informal coordination channels, and shift authority toward whoever has the fastest access to information or the clearest path to restoring operations.
Those moments reveal the system the organization actually built.
Governance Fails Where Visibility Ends
This distinction matters even more in environments shaped by AI, automation, and deeply interconnected digital infrastructure. As organizations distribute decisions across vendors, models, orchestration layers, APIs, and autonomous workflows, resilience depends less on whether failure occurs and more on whether the organization can preserve governability while systems are under stress.
Many organizations still approach governance primarily as a preventative function focused on policy creation, risk categorization, and compliance oversight. Those capabilities remain necessary, but incidents expose something deeper about institutional resilience. They reveal whether organizations can maintain operational coordination, constrain harmful behavior, understand what systems are doing in real time, and adapt quickly enough to prevent localized failures from cascading across the enterprise.
That gap between governance on paper and governance in practice increasingly defines operational risk.
Most enterprises already maintain extensive telemetry for technical systems. They log infrastructure events, monitor application performance, track authentication activity, and alert on anomalous behavior across networks and endpoints. Yet the governance layer itself often remains comparatively invisible. Organizations rarely instrument how decisions are actually made during incidents, when escalation paths are bypassed, which approvals are skipped under pressure, or how authority shifts dynamically during crisis response.
That absence matters because operational drift inside governance systems compounds the same way technical drift does. If a security team bypasses formal escalation paths during a major incident because leadership approval is too slow, the organization has discovered something important about its governance architecture. If operational teams routinely rely on informal coordination channels to restore systems faster than official processes allow, that behavior should not be dismissed as an exception. It should be treated as telemetry about how the institution actually functions under stress.
Many organizations now measure technical recovery with far greater precision than governance responsiveness. Mean Time to Recovery has become a standard operational metric across security and resilience programs, yet few organizations measure the equivalent governance latency inside their own decision systems. If infrastructure recovers in ten minutes but legal, operational, or executive authorization delays containment actions for four hours, governance has become the primary operational bottleneck whether leadership recognizes it or not.
In highly automated environments, governance telemetry increasingly matters as much as infrastructure telemetry.
The First Failure Is Usually Organizational
The first breakdown during an incident is often organizational long before it becomes fully technical.
The escalation process that looked clear on paper turns out to depend on informal relationships. Teams discover they were operating from different assumptions about who owned a decision or who had authority to intervene. A vendor relationship that appeared well governed reveals major visibility gaps once systems begin failing. Temporary operational exceptions quietly embedded into production environments become critical points of fragility because no one revisited them after deployment.
The 2024 CrowdStrike outage demonstrated this dynamic at global scale. A faulty Falcon content update caused millions of Windows systems worldwide to crash, disrupting airlines, hospitals, banks, retailers, and emergency services. CrowdStrike’s own post-incident analysis identified failures in validation and deployment processes tied to the update itself.
But the broader significance of the outage extended beyond a software defect. The incident exposed how deeply operational resilience now depends on tightly interconnected systems and concentrated dependencies that many organizations cannot meaningfully isolate once disruption begins. Recovery often required manual remediation device by device, revealing how heavily modern enterprises had optimized for centralized efficiency and automated scale without building equivalent investment in degraded-mode operations or reversibility. Reuters reporting on recovery challenges after the CrowdStrike outage
The outage also exposed a broader structural tradeoff many organizations had quietly accepted over the last decade. Enterprises consolidated operational visibility and control into highly centralized platforms because the model improved efficiency, simplified management, and reduced staffing burdens. In many environments, the same “single pane of glass” that improved operational coordination also became a concentrated point of systemic fragility once failure propagated across the ecosystem.
The organizations that recovered fastest were not necessarily the organizations with the most sophisticated technology stacks. In many cases, they were the organizations that understood their dependencies clearly, had practiced cross-functional coordination under pressure, and retained enough operational flexibility to adapt when trusted infrastructure failed unexpectedly.
A similar lesson emerged from the SolarWinds compromise, where attackers inserted malicious code into trusted software updates distributed across thousands of organizations. The incident exposed more than software supply chain vulnerabilities. It revealed how difficult it had become for organizations to govern inherited trust relationships across increasingly complex dependency ecosystems.
Many organizations designed governance models around direct access control while devoting far less attention to how authority and trust propagated indirectly through vendors, software dependencies, and interconnected operational environments. Once the trusted update channel itself became compromised, institutions discovered that many assumptions about visibility, oversight, and operational containment no longer held.
The Latency Gap
That challenge becomes even more complicated as AI systems accelerate the speed of both discovery and exploitation. In May 2026, Google warned that threat actors had already begun using AI systems to identify vulnerabilities and assist in exploit development, including what the company described as the first observed case of AI-assisted discovery of a zero-day vulnerability.
The significance of that warning extends beyond faster attacks. Organizations increasingly face environments where vulnerabilities can be discovered, weaponized, adapted, and operationalized faster than traditional governance and response structures were designed to handle.
Military strategists have long framed decision advantage through the OODA loop: observe, orient, decide, and act. In highly automated environments, systems increasingly observe and act at machine speed, while orientation and decision-making often remain constrained by fragmented organizational structures, committee processes, legal review cycles, and incomplete visibility across the enterprise.
The result is a collision between two operational realities. Technical systems increasingly move at machine speed while institutional coordination still moves at committee speed.
That latency gap is becoming a major resilience challenge.
If adversaries can adapt faster than institutions can orient and decide, governance effectively becomes operationally offline during the moments that matter most. Organizations may still possess formal authority on paper, but they lose the practical ability to constrain behavior quickly enough to shape outcomes in real time.
This dynamic is already reshaping how security leaders think about governance architecture. Static review boards and periodic approvals cannot govern systems operating continuously at machine speed. Organizations increasingly need runtime visibility, automated guardrails, policy enforcement mechanisms, and operational constraints embedded directly into the environments where decisions are executed.
Contractual Governance Is Not Operational Governance
The distinction becomes especially important in AI-enabled environments where organizations often mistake contractual governance for operational governance.
A vendor SLA may define liability after failure occurs, but it does not guarantee the organization can meaningfully govern the system during failure. A third-party AI provider may promise human oversight, explainability, or security controls, yet operational reality may still depend on black-box decision systems that customers cannot observe, interrupt, or constrain directly during active incidents.
As autonomy expands, many organizations are quietly shifting from human-in-the-loop environments toward human-on-the-loop environments, where people supervise systems operating largely independently until something appears to go wrong. Under those conditions, preserving meaningful human intervention becomes significantly harder because the operational tempo of the system may already exceed the speed of institutional coordination.
This distinction also exposes an important blind spot in many current AI governance conversations. Much of the public discussion around AI governance still centers on inference problems such as bias, fairness, hallucinations, and explainability. Those issues matter, particularly in customer-facing and high-impact systems. But during incidents, organizations confront a different challenge entirely: agency.
The core governance question shifts from whether a model generated imperfect information to whether an autonomous or semi-autonomous system can take actions that humans cannot meaningfully observe, interrupt, reverse, or constrain fast enough to preserve operational control.
In practice, governability increasingly depends on whether organizations maintain operational and legal pathways to disengage automated workflows safely under pressure. A kill switch is not simply a technical capability buried inside infrastructure. It is a pre-validated operational process that determines who has authority to intervene, under what conditions automated behavior can be constrained, how downstream dependencies are managed, and how organizations preserve continuity while regaining control.
Turning a system off is often the easy part.
The harder challenge is whether the organization can continue operating safely in degraded mode once the automated system disappears. Incidents frequently reveal that institutions optimized aggressively for automation without preserving equivalent operational elasticity once autonomy becomes unavailable. The real resilience test is not whether leaders can interrupt a system. It is whether the business can survive long enough to recover control.
This is one reason incident response has become such a valuable governance lens. Under pressure, organizations stop operating according to idealized governance diagrams and begin operating according to the actual incentives, dependencies, communication pathways, and authority structures embedded inside the system. Incidents expose whether teams can coordinate effectively, whether escalation paths function in practice, whether leaders have sufficient visibility to make decisions confidently, and whether operational learning accumulates quickly enough to improve resilience over time.
From Tabletops to Continuous Resilience
Well-designed tabletop exercises remain valuable for precisely this reason, although point-in-time simulations alone are becoming insufficient for highly dynamic environments.
The best exercises do not exist simply to validate that existing procedures work exactly as intended. They create controlled pressure that helps organizations surface hidden assumptions before real-world consequences force the lesson at larger scale. Teams often discover that different parts of the organization hold conflicting mental models about the same systems, dependencies, or escalation rights. Leadership may believe certain operational controls exist while technical teams know they are inconsistently enforced or largely manual.
But resilient organizations increasingly extend beyond traditional tabletop exercises into continuous resilience testing approaches that resemble chaos engineering for governance itself. Instead of simulating only technical outages, organizations can test what happens when critical decision-makers become unreachable during an incident, when legal obligations conflict with operational recovery priorities, or when automated workflows continue operating after escalation authority becomes unclear.
Those exercises reveal something critically important: governance failure rarely occurs at a single point. It emerges through accumulated coordination breakdowns, visibility gaps, dependency assumptions, and latency mismatches that only become visible once systems experience sustained pressure.
They also expose governance drift.
Just as configuration drift occurs when production systems gradually diverge from approved configurations over time, governance drift occurs when operational authority gradually shifts away from the people formally designated to make decisions toward the people practically capable of making them during incidents. Under pressure, organizations often discover that the real incident commander is not the executive listed in the escalation plan but the engineer, operations lead, or vendor representative with the fastest access, deepest visibility, or most actionable information.
Continuous resilience testing helps organizations measure the gap between the governance model represented in the org chart and the governance model reflected in incident response behavior. That gap is often one of the clearest indicators of operational fragility.
This dynamic mirrors a lesson organizations already understand well in product design and increasingly in security design. Features and controls only work if people can realistically use them under real operating conditions. In cybersecurity, organizations learned long ago that overly burdensome controls frequently produce workarounds rather than compliance. The same principle applies to governance.
Most governance structures are designed thoughtfully, but they are not always evaluated against the realities of high-pressure operational environments where people must act quickly with incomplete information and competing priorities. Under those conditions, teams naturally gravitate toward the paths that restore functionality, reduce uncertainty, and accelerate coordination. When governance mechanisms become too slow, fragmented, or operationally disconnected, organizations begin creating informal “desire paths” around them.
That behavior should not be interpreted solely as resistance to governance. It is also feedback about whether governance remains usable under stress.
Organizations that mature operationally tend to treat incidents and exercises as learning systems rather than isolated disruptions. Instead of focusing exclusively on root cause remediation, they examine how authority flowed during the event, where visibility broke down, which dependencies created hidden risk, and whether governance mechanisms operated effectively under real conditions. Over time, that learning shapes how they redesign escalation structures, improve runtime visibility, constrain automation, and strengthen resilience across interconnected environments.
Governability Under Pressure
The AI era is changing the nature of operational failure itself. Organizations are no longer managing only isolated software defects or individual human mistakes. They are managing distributed systems where decisions emerge dynamically across models, datasets, APIs, vendors, workflows, and increasingly autonomous actors operating at machine speed.
Under those conditions, governance cannot function primarily as a documentation exercise performed separately from operations.
Governance increasingly becomes part of the operational architecture itself, embedded into the systems that shape how authority is exercised, how constraints are enforced, and how decisions are executed when humans are no longer directly involved in every action.
Complex systems will always produce surprises, especially as autonomy scales and dependencies deepen across digital ecosystems. But resilient organizations preserve something critically important during instability: the ability to understand what is happening, coordinate decisions effectively, intervene before failures cascade, and adapt faster than risk compounds.
The future of governance will depend increasingly on that capability.
The organizations that navigate this transition successfully will not necessarily be the ones with the most advanced models or the fastest deployment cycles. They will be the organizations capable of remaining governable while operating systems that are becoming faster, more autonomous, and more interconnected than the institutions surrounding them were originally designed to manage.
The uncomfortable question for leadership teams is whether their governance systems have ever been tested under the conditions they are expected to survive, or whether leadership has primarily validated governance through documentation, approvals, and policy reviews removed from operational pressure.
When was the last time your organization intentionally introduced friction into a critical decision path, removed a key decision-maker during an exercise, or forced teams to operate in degraded mode long enough to observe where authority actually flowed under pressure?
Because governance as the system actually built ultimately reflects what the organization rewards during crisis conditions. If teams routinely bypass policy in order to restore operations, then the policy was functioning as an obstacle rather than as an operational guide. Resilient organizations align the incentive to remain governed with the incentive to remain operational, even when the system is under stress.
2026 Series | Q2: Governance as a Capability
This essay is part of a series exploring how organizations move from policy and oversight to real-time control as systems gain autonomy.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.