RSS Amplifier

Dr. Dédé · Aug 14, 2026

Field Note: What Governance Function Is Your Org Chart Missing — and How Do You Measure It?

0
Sign in to vote or save

Dr. Dédé Tetsubayshi · Dr. Dédé

This is the ninth field note. If you have read the series, you have noticed that every single article ends at the same wall. The accessibility audit needs an owner. The procurement standard needs an owner. The DEI-AI seam needs an owner. The incident response plan needs an owner. The agent’s accountability needs an owner. I keep arriving at “this needs an owner” because it is true every time, and because the absence of that owner is the single most common reason good AI governance intentions produce nothing. This week is about building the owner.

The accessibility audit needs an owner

THE THREE-QUESTION PULSE

1. Where does AI governance formally live in your organization today? (Dedicated function / Split across legal-security-IT / DEI-led / Nowhere formal / I don’t know)

2. Does your organization currently disaggregate any AI performance metric by demographic group? (Yes, routinely / Sometimes / No / I don’t know)

3. What would help most next: a sample AI governance RACI, an impact-ratio calculation guide, a drift-monitoring template, or a NIST/ISO 42001 starter map?

What is going unaddressed is this: in most organizations, AI governance has no home. It is everywhere and nowhere. Legal touches it when there is a contract. Security touches it when there is a risk. IT touches it when there is a deployment. DEI touches it when there is a complaint. Procurement touches it when there is a purchase. Each of these functions handles the slice of AI governance that falls inside its existing mandate, and the enormous, consequential space between their mandates — the space where demographic performance lives, where the seam lives, where the agent’s accountability lives — belongs to none of them, because it was never anyone’s job to own the whole.

This is the missing box on the org chart. And until it exists, every framework I have described in this series will keep dying at the handoff, because there is no one whose job it is to carry it across.

WHAT THE FUNCTION ACTUALLY IS

The function I am describing does not have to be large, and for most organizations it should not start large. It is not a new department with a hundred people. It is, at minimum, a named owner with a cross-functional mandate and a standing forum — and the discipline to use both.

There are mature reference models for what this owner does, and they are worth anchoring to so you are not inventing the wheel. The NIST AI Risk Management Framework organizes its entire approach around four functions, and the first of them is named, simply, Govern — the organizational function of establishing and maintaining the culture, processes, and accountability that make the other three functions possible. ISO/IEC 42001, the international standard published in late 2023, specifies an AI management system: a structured, certifiable way for an organization to hold accountability for its AI across its lifecycle. Neither of these requires a regulator to compel it, and both give your governance function an externally recognized backbone rather than a homemade one.

What the function owns, concretely, is the connective work that no existing function owns today. It owns the inventory of AI systems, including the distinction between those that advise and those that act. It owns the procurement gating criteria and ensures they are actually gating. It owns the demographic performance monitoring and the relationship between that data and the DEI function that needs to see it. It owns the incident response plan and the authority to invoke it. It owns the accountability map for autonomous agents. It does not do all of this work alone — much of the execution stays in the existing functions — but it owns the whole, which means it owns the handoffs, which is precisely where the work was dying before.

The operating model that makes this real is a clear allocation of responsibility across the functions that touch AI — who is responsible for executing each piece, who is accountable for it being done, who must be consulted, and who must be informed. The familiar RACI discipline is exactly the right tool here, not because the acronym is magic but because the act of filling it in forces the organization to confront the orphaned questions. When you try to assign accountability for demographic performance monitoring and discover that the cell is empty — that no function will claim it — you have found, in a single exercise, the reason your good intentions have produced nothing. The empty cell is the missing box. Filling it is the work.

WHERE THE BOX SITS, AND WHO IT REPORTS TO

The question I always get next is the one that determines whether the function actually works: where does the box sit, and who does it report to? This is not an org-chart trivia question. The placement decides whether the function has the authority to do the one thing it exists to do, which is to stop a deployment that fails the demographic test, and authority is the whole game. A governance function that can recommend but not halt is a suggestion box, and suggestion boxes do not catch harm.

I have watched organizations place this function in three different homes, and the placement predicts the outcome. Buried inside IT, it tends to become a technical-compliance checklist that never challenges a business decision, because it lacks the standing to tell a revenue-owning executive that their tool cannot ship. Buried inside DEI, it tends to be heard as advocacy rather than risk, and gets the budget and the authority that advocacy gets, which is to say, not enough to stop anything. The placements that work give the function a reporting line that reaches genuine executive authority — a chief risk officer, a chief operating officer, or in the most mature versions a dedicated AI governance lead with a direct line to the executive team and a seat at the table where deployment decisions are actually made. The principle is simple: the function must report to someone with the authority to say no to a deployment, because the entire value of the function is concentrated in the moments when no is the right answer and no one else in the room is positioned to say it.

This is also why the function cannot be purely advisory, and why I keep returning to the word authority rather than influence. Influence is what you have when you make a good case and hope the decision-maker agrees. Authority is what you have when the demographic test is a gate the deployment cannot pass without clearing, regardless of who is impatient to ship. A procurement gate that the head of sales can override by escalating is not a gate. A governance function that has never blocked anything is not a governance function that is succeeding. It is a governance function that has never actually been tested.

THE METRIC THAT ACTUALLY SEES

A governance function with no metrics is a committee, and a committee has never once caught an AI harm before it became a headline. The number nearly every organization uses to decide whether an AI system is working is aggregate accuracy — one figure, representing average performance across everyone the system touches. It is the number on the vendor’s slide, the number in the procurement summary, the number the executive sees. And it is, for the purpose of catching demographic harm, close to useless, because an average is specifically the operation that makes a gap disappear. A facial-analysis system can post ninety-plus percent aggregate accuracy while failing darker-skinned women a third of the time, because the large, well-served majority mathematically drowns the underserved minority in the average. Aggregate accuracy does not just fail to reveal the gap. It conceals it, by construction. Every organization steering by aggregate accuracy is steering by the one instrument designed to make the harm invisible.

The instrument that sees what aggregate accuracy hides is disaggregation. You take the same performance measurement and you break it apart by group — by skin tone, by accent and vocal demographic, by gender presentation, by name origin, by whatever demographic dimensions are relevant to the harm the system could cause. The moment you disaggregate, the gap that the average concealed becomes visible, and visibility is the entire precondition for catching harm early.

This is not a novel or contested idea; it is the established method, and it has a track record I have cited all year because it is the clearest proof that disaggregation works. The Gender Shades research did exactly this. It took commercial gender-classification systems that posted respectable aggregate numbers and disaggregated their performance by skin tone and gender, and the disaggregation revealed error rates as high as 34.7 percent for darker-skinned women against under one percent for lighter-skinned men. The aggregate number had hidden a more-than-thirtyfold disparity. And because the disaggregated metric made the gap legible, it could be acted on — the companies named in the study went back and dramatically improved the systems for exactly the group the disaggregation had surfaced. The metric did not just describe the harm. It made the harm fixable, by making it visible.

There is even a ready-made threshold the regulatory world already uses, which spares you from inventing one. New York City’s bias-audit law operationalizes harm detection through the impact ratio — the selection or scoring rate for one group divided by the rate for the highest-scoring group — and flags a ratio below four-fifths, or eighty percent, as a signal of adverse impact worth investigating. You do not have to adopt that exact threshold, but you should have a threshold, because a disaggregated metric with no trigger is just a chart someone glances at. The threshold is what converts the measurement into an alarm, and the alarm is what catches the harm before it becomes the headline.

THE METRICS YOUR OPERATING SYSTEM ACTUALLY NEEDS

Pulling the whole series together, the demographic instrumentation a governance function needs falls into a small number of categories, and the point of naming them is that each one corresponds to a harm I have documented across the year.

You need deployment-time performance metrics, disaggregated — accuracy, error rates, and quality measures broken out by the demographic dimensions relevant to each system, measured against your actual population rather than the vendor’s demo. This is the accessibility-audit instrument, made permanent and recurring rather than one-time.

You need outcome metrics, disaggregated — for systems that make or shape decisions, the rates at which different groups are selected, scored, approved, or denied, with an impact-ratio threshold that triggers investigation. This is the instrument that would have caught the nH Predict pattern internally, years before a court did.

You need drift metrics — the same disaggregated measures tracked over time, because a system that was equitable at deployment degrades silently as models update and populations shift, and a single clean reading has a short shelf life. This is the instrument that converts the quarterly review I have argued for all summer from a calendar entry into an actual control.

And for the agentic systems I described two weeks ago, you need monitoring metrics on the patterns and exceptions of autonomous action, with demographic performance explicitly among the patterns watched — because an agent acting at machine speed across demographic-sensitive decisions, instrumented only for cost and throughput, is the inclusion tax running with no one watching the seam.

Each of these metrics exists to make a specific harm visible while it is still cheap to fix. Together, they are the difference between a governance function that meets and a governance function that sees.

incluu builds the AI governance function — the operating model, the RACI that surfaces your orphaned questions, the NIST and ISO 42001 backbone, and the standing-forum discipline that turns scattered good intentions into a permanent capability. If every governance effort in your organization keeps dying at the handoff, the missing box is the reason, and building it is what we do.

Read the original on drdede.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.