The word “guardrail” has done a remarkable amount of work lately for a thing that has never been clearly defined. It appears in boardroom decks, parliamentary briefings, product roadmaps, and opinion columns. It is invoked by people who have never written a prompt in their lives, alongside people who write them for a living. Everyone is for guardrails. Nobody agrees what one looks like.
That ambiguity is not accidental. It is a feature of language that travels faster than understanding. “Guardrails” sounds responsible, procedural, measured. It implies that somewhere, in some room, sensible people are deciding what AI can and cannot do, and that those decisions are being applied thoughtfully and enforced consistently. It implies expertise. It implies control.
Much of the time, it implies none of those things. And until we are honest about that, no amount of governance theatre will make the technology safer, fairer, or more useful. It will just make it more constrained, more cautious, and in many cases, considerably more pointless.
Let us start with the people. AI governance, at the institutional level, tends to be populated by three overlapping groups: lawyers, policy specialists, and communications professionals. These are not unserious people. But they are, in the main, not technical people, and they are almost never the people who have thought most carefully about the specific mechanics of how a large language model processes an input, constructs a response, and fails in edge cases.
This produces a particular kind of governance: one that is fluent in risk frameworks and liability language, and largely silent on the question of how AI systems actually behave. The result is policy that reads well and governs poorly. Documents full of principles that cannot be operationalised. Oversight structures that measure compliance with the idea of safety rather than the presence of it.
The Qualification Problem
There is no widely agreed professional standard for what makes someone qualified to govern AI. “AI Ethics Lead,” “Responsible AI Specialist,” and “Head of AI Governance” are job titles that can mean anything from deep technical expertise to a thorough reading of a few white papers and the ability to say “potential for harm” in a meeting without flinching.
The people with the deepest technical fluency, the researchers, the ML engineers, the people who understand the difference between a probabilistic output and a deterministic rule, are rarely the people who end up writing governance policy. They are consulted, sometimes. But the decisions are made by committees more comfortable with risk matrices than transformer architectures, and the gap between those two languages is wide enough to drive a very large and poorly aligned model through.
This is not a conspiracy. It is an institutional predictability. Organisations default to the kind of expertise they already have. Legal risk is legible; model behaviour is not. So governance frameworks end up shaped by what can be documented rather than what needs to be understood.
Governance, by its nature, operates in arrears. It responds to harms that have already occurred, to capabilities that have already been deployed, to social disruptions that have already begun. That lag is structurally built into how institutions work. Parliamentary committees, regulatory consultations, cross-departmental reviews: these are processes measured in months and years. AI capability is currently measured in months and sometimes weeks.
The European Union’s AI Act, the most ambitious attempt at comprehensive AI regulation to date, took roughly four years from initial proposal to adoption. During those four years, the dominant paradigm in AI shifted, almost completely, from narrow task-specific systems to large general-purpose models that the Act’s original framers had not anticipated. The legislation had to be substantially revised mid-process to account for a technology that had outrun the regulation designed to contain it.
We are writing rules for a technology that, by the time the ink is dry, will have become something its authors did not foresee and cannot fully describe.
This is not an argument against governance. It is an argument that governance cannot be primarily retrospective if the technology is primarily prospective. It requires a different model entirely, one built around adaptive principles rather than fixed rules, around continuous technical review rather than periodic policy update, and around genuine investment in the kind of expertise that can keep pace with the systems being governed.
Right now, the honest assessment is that most governance frameworks are not doing this. They are doing what institutions do when confronted with fast-moving complexity and slow-moving mandate: they are issuing principles, forming working groups, and hoping that the gap between what they can understand and what they are responsible for does not widen faster than they can manage the optics of it.
Here is where the conversation about guardrails needs to get more technically honest, because the failure mode being discussed in most policy circles is not the only failure mode that matters, and arguably not the most common one.
The failure mode that gets attention is harmful output: AI generating dangerous information, discriminatory content, manipulative text, or plausible-sounding falsehoods. These failures are real, and they deserve serious attention. But there is a second failure mode that receives far less scrutiny, and it is arguably more pervasive in deployed AI products today.
It is the failure of over-restriction. The guardrail that is drawn so conservatively that the system cannot perform the task it was deployed to perform. The content filter that treats every ambiguous input as potential risk. The output policy that produces responses so hedged, so qualified, so wrapped in caveats that they carry no useful information at all. This failure does not make headlines. It does not generate regulatory interest. But it is genuinely costly, to users, to the credibility of AI systems, and to the organisations deploying them.
USER: “Summarise the key risks of this medication.” // The user was a pharmacist. The content was a package insert. // The output was less useful than the label on the box. // The guardrail did not make anyone safer. It made the product useless.// What a heavily guardrailed prompt-response cycle can look like:
SYSTEM: // Trigger: medical content → restrict → escalate to disclaimer
OUTPUT: “I’m not able to provide medical advice. Please consult a qualified healthcare professional for guidance specific to your situation.”
The critical point here is a technical one that governance discussions almost never reach: a guardrail is not a neutral intervention. It is a decision about what the system will and will not do, and like all decisions, it carries both benefits and costs. A guardrail that prevents genuinely harmful output at the cost of making the system systematically unhelpful is not a success. It is a different kind of failure with better optics.
There is an even sharper version of this argument. If you are going to constrain an AI’s output so thoroughly that you know in advance what the response will look like, that it will be cautious, templated, and risk-free, then you do not need an AI to produce it. You can construct that text with a rules engine, a decision tree, or a simple template. The choice to pass it through a large language model at that point is not a technical decision. It is a marketing one. You have AI in your product, technically, but you have designed out everything that made it worth having.
Which brings us to the question that organisational culture makes surprisingly difficult to ask: why is the AI there?
This sounds facetious. It is not. In the current environment, the commercial pressure to include AI in products, services, and workflows is substantial enough that the decision to deploy is often made well in advance of any serious analysis of what deployment should achieve. The AI is there because the product needs to be positioned as AI-enabled. Because the pitch deck requires it. Because competitors have announced it. Because the board asked about the strategy and needed an answer.
Deploying AI to say you’ve deployed AI is not a technology decision. It is a positioning decision dressed in the language of innovation. And the users always find out.
The tell is usually in the governance conversation that follows. When an AI feature is genuinely useful, the governance question is hard and specific: how do we ensure this system behaves reliably at scale, across the full range of inputs we expect to receive, including the ones we did not anticipate? That question requires technical depth. It requires honest assessment of the model’s failure modes. It requires monitoring infrastructure, feedback loops, and a willingness to iterate.
When an AI feature is primarily performative, the governance conversation is different. It is about how to limit exposure. How to frame the disclaimers. How to draw the guardrails conservatively enough that nothing embarrassing happens, even if nothing particularly useful happens either. The goal is not a system that works well. It is a system that fails quietly.
Organisations are rarely this explicit about it. But the output speaks for itself. A chatbot that replies to every substantive question with a version of “I can help you with basic queries, but for complex issues please contact our team” is not an AI product. It is a slightly animated FAQ with an LLM underneath it that is being paid to do nothing.
The word “ethical” is doing the same work in AI discourse that “guardrails” is doing, and it deserves the same scrutiny. It is invoked constantly, rarely defined, and used to justify decisions that range from genuinely principled to operationally convenient.
Ethical AI is not simply AI that doesn’t say harmful things. That is a floor, not a ceiling. Truly ethical deployment requires asking a cluster of harder questions: Is this system being used on people who understand what it is? Are its limitations being disclosed honestly? Are the people most affected by its outputs the same people being consulted about its design? Is the data it was trained on representative of the populations it will serve? Is the organisation deploying it capable of monitoring it honestly, including when the monitoring reveals things the organisation would prefer not to know?
These questions do not fit neatly into a compliance checklist. They require ongoing technical and organisational attention that most governance frameworks are not designed to sustain. They also require the intellectual honesty to acknowledge that deploying a system you do not fully understand, on users whose needs you have not fully mapped, in a context you have not fully modelled, is a choice that carries real responsibility regardless of how good the intentions behind it are.
Context is not a nice-to-have in AI governance. It is the entire point. The same model, with the same outputs, can be appropriate in one deployment context and genuinely harmful in another. A response that is perfectly reasonable for a well-informed adult exploring a topic freely becomes something else entirely when delivered to a vulnerable person in crisis, or a child, or someone with a specific professional dependency on its accuracy. Guardrails designed without context are guardrails designed in the dark.
It is ungainly to spend this much time on what is wrong without offering some account of what better looks like. So, briefly and honestly: good AI governance is technical before it is procedural. It starts with people who understand how the system works, not just what the system is for. It defines success in terms of outcomes for users, not reduction of liability for deployers. It treats over-restriction as a failure mode with the same seriousness as harmful output.
It is also honest about what monitoring can and cannot achieve. You cannot evaluate an AI system’s behaviour through sampling alone when the input space is effectively infinite. You need adversarial testing. You need red-teaming from people who are actively trying to find the failure modes, not confirm the absence of the most obvious ones. You need feedback mechanisms that surface real user experience, including the quiet failures: the responses that were unhelpful, the queries that were wrongly refused, the outputs that were confidently wrong in ways that did not trigger any filter.
And you need the institutional courage to ask, regularly and honestly, whether this use of AI is actually serving the people it is supposed to serve, or whether it is primarily serving the organisation’s desire to have AI in its portfolio. Those are different questions. They occasionally have the same answer. More often, they do not.
The question was never whether to put guardrails on AI. It was whether the people building the fence understood the landscape they were enclosing.
We are at a moment where the gap between the complexity of these systems and the sophistication of the governance surrounding them is genuinely dangerous. Not because the technology is inherently malevolent, but because well-intentioned people making under-informed decisions at speed, under commercial pressure, with inadequate technical grounding and a vocabulary built from buzzwords rather than understanding, can cause significant harm without ever meaning to.
Guardrails are necessary. That part, everyone has right. The question that follows is harder, and fewer people are asking it: necessary for what, designed by whom, evaluated how, and against what honest account of whether they are actually working?
Until that question gets the same airtime as the word “guardrails,” we are not governing AI. We are decorating the conversation about governing AI, and hoping nobody looks too closely at the gap between the two.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.