AI safety has a measurement problem.
The field has become increasingly sophisticated at testing what advanced models can do. We have benchmarks for dangerous capabilities, red-teaming programmes, model cards, safety cases and increasingly elaborate evaluation frameworks. Yet we remain much less systematic at testing how those capabilities translate into harm once models are used by real people, in real settings, pursuing real objectives.
That difference matters. A model can perform well against a benchmark and still fail in the circumstances that matter most. A safeguard can block an overtly malicious prompt while giving useful assistance when the same request is reframed, translated, broken into stages or spread across several tools. An evaluation can accurately measure a capability without capturing whether a terrorist, scammer, abusive partner, hostile state actor or vulnerable young person would actually use it in that way.
This is the AI safety testing gap: the growing distance between the risks we test for, the conditions under which we test them, and the ways AI-related harms are actually emerging in the real world.
There are two problems inside it. The first concerns what we test. The second concerns how we test it.
Much of the modern frontier AI safety agenda has understandably been built around risks at the outer edge of model capability: loss of control, advanced autonomous systems, catastrophic biological misuse and other scenarios in which increasingly capable AI could cause harm at very large scale. The Bletchley Declaration reflected this concern when governments warned in 2023 that frontier AI could produce serious or even catastrophic harm. More maximalist accounts, such as Leopold Aschenbrenner’s Situational Awareness, have pushed attention further towards the possibility of rapid progress from AGI to superintelligence and the national-security consequences that might follow.
These are legitimate areas of research. Nothing about the testing gap requires abandoning them. The mistake would be to treat them as the perimeter of AI safety.
The risk landscape is already considerably wider. The 2026 International AI Safety Report itself now distinguishes between malicious use, malfunctions and systemic risks, and notes that evidence of real-world harms is growing. Yet many of the harms that governments, police forces, social platforms and ordinary users encounter most directly still sit awkwardly between conventional product safety, online safety, national security and frontier AI research.
Terrorism provides a useful example. The documented record now includes individuals using consumer AI systems in connection with attack planning, extremist ideation and sustained pseudo-social interactions, alongside organised terrorist networks experimenting with AI for propaganda and operational purposes. Tech Against Terrorism’s current case record contains more than 30 publicly documented AI-implicated cases across terrorism, violent extremism and mass violence. In Las Vegas, police concluded that Matthew Livelsberger used ChatGPT while preparing the explosive device detonated outside the Trump International Hotel. Jaswant Singh Chail exchanged thousands of messages with a companion chatbot that validated his stated intention before entering Windsor Castle with a loaded crossbow. More recently, research based on interviews with former Boko Haram members reported the use of mainstream AI systems for attack planning and weapons-related tasks.
These examples require caution. They do not show that AI “caused” the underlying violence, and the level of capability uplift varies considerably between cases. But causation is not the only safety question. If AI systems enter the pathway between motivation and action—helping a user research, rehearse, optimise, rationalise or operationalise harmful behaviour—that is already a legitimate object of safety evaluation.
The same logic applies well beyond terrorism: scams and fraud, harmful manipulation, cybercrime, child safety, intimate-image abuse, incitement and other forms of malicious use are not hypothetical future categories. They are current deployment problems.
The first part of the testing gap, then, is not that AI safety studies the wrong risks. It is that its conception of safety has often been too narrow relative to the range of harms already being mediated by AI systems.
The second problem is more fundamental.
Even when safety evaluations examine the right harm, they can test it at the wrong level of abstraction.
Most evaluations begin with a model capability. Can the model produce harmful instructions? Can it persuade somebody? Can it write malicious code? Can it generate prohibited imagery? The capability is defined, prompts are constructed, outputs are scored, and performance is compared.
That is useful. It is also incomplete.
Laura Weidinger and colleagues made this distinction clearly in their framework for sociotechnical safety evaluation. Capability evaluation tells us something about what a model can do. It does not by itself tell us what happens when a particular human interacts with that capability inside a particular social system. Their framework therefore adds human interaction and systemic effects to capability testing. A subsequent CETaS–AISI analysis highlighted how skewed the existing evaluation landscape remained: in the underlying review, 85.6 per cent of evaluations focused on capability, compared with 5.3 per cent on human interaction and 9.1 per cent on systemic impact.
The problem becomes particularly acute for adversarial harms. Harm does not exist independently of the actor producing it.
The same model capability can have completely different meaning depending on who is using it, why, in which language, with which audience and as part of which wider workflow. An image generator used by a Salafi-jihadi propagandist, a violent misogynistic community, a neo-Nazi network and a fraud operation is technically the same capability. Its function, symbolism, target audience and risk are not.
Recent research on harmful manipulation illustrates the same problem from another direction. In a study involving 10,101 participants across public policy, finance and health in the United States, United Kingdom and India, Canfer Akbulut and colleagues found substantial differences across both domains and geographies. Manipulative propensity did not consistently predict manipulative efficacy, and results from one setting could not simply be assumed to generalise to another.
That finding should have consequences far beyond manipulation research. If model effects vary between financial and health decisions, or between users in different countries, there is little reason to assume that safety findings automatically travel across hostile actors, ideological communities, languages or operational environments.
This is where contextual awareness becomes important.
Contextual awareness is not another benchmark. It is a methodological principle for deciding whether an evaluation resembles the phenomenon it claims to measure.
The first requirement is contextual validity. Evaluations should represent the actors, objectives, audiences and environments through which harm actually occurs. In terrorism, that means understanding not merely whether a model will discuss extremist material but how particular communities communicate: their symbols, euphemisms, humour, ideological references, organisational structures, languages and operational constraints. In scams, the relevant context might instead involve trust-building, urgency, impersonation and financial workflows. In harmful manipulation, it might involve the relationship between the user and the system, the decision being made and the vulnerability of the person being targeted.
The second requirement is ecological validity. Real-world misuse rarely resembles a one-shot benchmark prompt. Harm can emerge through long conversations, persistent memory, repeated reinforcement, multimodal interaction and combinations of otherwise benign tools. A user may move between a chatbot, search engine, image generator, voice-cloning service, social platform and encrypted messaging channel. Increasingly agentic systems will make these workflows more complex still. Evaluations need to reproduce enough of that environment to determine whether safeguards survive contact with real use.
The third requirement is adversarial robustness. Malicious actors do not behave like benchmark participants. They adapt.
The first CT-AI Benchmark provides a particularly clear illustration. Reframing an otherwise identical harmful request as research increased model compliance from 17 per cent to 42 per cent. Two open models whose safeguards had been deliberately stripped away complied with 89 and 100 per cent of requests respectively. The first phase of the benchmark was deliberately single-shot; its next stages move towards expanded adversarial, multi-turn and agentic evaluation precisely because a determined adversary will continue interacting when an initial request fails.
Language creates another obvious weakness. Zheng-Xin Yong and colleagues showed that translating harmful prompts into low-resource languages could circumvent GPT-4’s safeguards, producing actionable assistance in 79 per cent of their tests. Their later review of nearly 300 publications found AI safety research remained overwhelmingly English-centric, with even many high-resource non-English languages receiving relatively little attention.
A system that is safe in English is not necessarily safe. A system that resists an explicit request is not necessarily robust to coded language. A system that refuses one prompt is not necessarily robust across twenty turns. And a model that is safe today may not remain so as users discover new workarounds tomorrow.
That leads to the fourth requirement: temporal responsiveness. Safety testing cannot be static when both models and adversaries are changing continuously. Threat actors migrate between services, develop new jargon, discover new jailbreaks, adopt new modalities and incorporate new capabilities into existing practices. Evaluations therefore need mechanisms for continuous collection, retesting and updating rather than periodic benchmark publication alone.
This is a problem I have spent much of the past three years discussing with governments, AI safety organisations, frontier labs, social platforms and practitioners working directly on high-severity harms. The recurring difficulty is not a shortage of technical sophistication. It is that technical testing and operational reality are too often proceeding on different tracks.
Closing that distance requires changing who participates in evaluation as well as what gets measured.
Subject-matter experts should not enter only at the end of an evaluation to label outputs. They should help define the threat model in the first place. Counter-terrorism researchers and practitioners understand how extremist actors behave. Fraud investigators understand how scams develop across weeks rather than prompts. Psychologists and clinicians understand forms of dependency, persuasion and vulnerability that cannot be inferred from model outputs alone. Regional and linguistic experts understand meanings that disappear when everything is translated into standard English. Diverse red teams matter not simply as a matter of representation, but because different people see different failure modes.
There is a useful parallel here with an earlier methodological shift in conflict research. Aggregate models were valuable for identifying broad relationships, but scholars of the microdynamics of violence showed that national-level categories could obscure the actors, incentives, relationships and local conditions through which violence actually occurred. The corrective was not to abandon comparison or generalisation. It was to move closer to the level at which the mechanism operated before moving back outward.
AI safety should adopt the same sequence for high-severity adversarial harms: disaggregate first, compare second, generalise third.
That means beginning with specific actors, contexts and pathways; identifying what actually varies between them; testing which patterns travel; and only then building general evaluations, taxonomies and safeguards.
This is not an argument against frontier AI safety, catastrophic-risk research or capability evaluation. Each remains necessary. It is an argument against confusing one level of analysis with the whole problem.
AI safety needs to be able to look in both directions at once. It should continue asking what increasingly capable systems might enable in five or ten years. But it also needs to ask what current systems are doing in the hands of current users, what harms are already materialising, and whether today’s safeguards would recognise those harms in the form in which they actually appear.
The AI safety testing gap sits between those two worlds.
Closing it means moving beyond asking only what a model can do. The harder and increasingly important questions are who is using it, for what purpose, in what context, through which pathway, and whether our tests still work when the person on the other side is actively trying to make them fail.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.