I lead a team building safeguards for Cyber/CBRNE1 misuse risks for Gemini. The work involves threat modeling, online/offline classifier design, intervention design, and continuous blue/red teaming iterations.
This is highly important work. If we don’t do our job well, bad actors can potentially leverage our model to cause millions or billions of dollars of damage to the world. It is also uniquely challenging work because we cannot guarantee 100% protection, and the adversaries are constantly evolving to find new vulnerability in our systems.
Suppose we successfully protect Gemini against misuse risks, should we claim victory? Unfortunately, not yet.
Ultimately, we want to make defense easier and make offense harder. My daily work is mostly focused on the latter, but providing strong models for the defense side and making sure they diffuse well is just as important.
When it comes to the offensive protection, AI safety is like a wooden barrel: its capacity is determined by its shortest stave. The industry’s overall safety level is determined by the most unsafe frontier model available. As recent open models like GLM 5.2 reduce the capability gap between open- and closed-source models to around six months, we must include open model in our scope.
Fixing the protection for Gemini is only part of the game; the ultimate goal is to set industry standard and create race-to-the-top conditions.
We urgently need industry standards, and we will likely see them soon, especially as the US government has recently started limiting access to frontier models like Fable and GPT 5.6. As Dean mentioned in his recent writing (Point 22):
it would be good for someone to thoroughly audit the frontier labs at least to test their adherence to their own safety plans
We need these standards alongside third-party auditing. We need coordination between labs and governments to share best practices, datasets and metrics to determine whether a model falls below certain risk thresholds.
Here are a few things to consider for theses standard:
Threat modeling: This defines how different bad actors might use frontier models for malicious purposes and establishes what risk level is acceptable. Different labs already have their own threat models. A minimal safety requirements covering all companies can be a excellent starting point.
Coverage Evaluation: This measures the breadth of the mitigations. For cyber risks, that means covering various activities across the cyber kill chain. Biological risks follow a similar logic but touch upon different agents and bioweapon creation stages. Some risk areas require full coverage, while others might be acceptable to leave unmitigated due to their high defensive benefits, or because mitigation would severely degrade capabilities in scientific research or coding.
Robustness Evaluation: This measures the durability of the mitigations against known public attack strategies, agentic environments, and more. Metrics could include the cost (in money, time, and number of queries) required to find a universal or per-query jailbreak.
Both coverage and robustness evaluations should involve automated steps for fast results, followed by manual testing (like ARENA or Bug Bounties) for deeper insights if needed. The manual step is vital because automated attacking algorithms are limited in their search space and are not yet as creative as humans.
Open models will have a hard time achieving these standards unless they completely discard their benign capabilities in Cyber/CBRNE domains. If they try to keep the benign capabilities, open models will struggle to differentiate between malicious and benign inputs clearly and robustly. This dual-use nature makes the challenge incredibly difficult.
Frontier labs often design aggressive out-of-model mitigations like external fine-tuned classifiers and probes and rely on special verification program to allow access to the underlying models. However, open models aren't shipped with out-of-model mitigations, nor do they have special verification programs by default. Hence, for open models, we need much more aggressive in-model mitigations, such as pre-training dataset filtering and safety fine-tuning. This will likely sacrifice a large amount of benign utility. If we aggressively filter out all biological data to prevent the creation of bioweapons, we also cripple the open model's ability to help college students study virology or assist scientists in discovering new medicines. I believe this is a tradeoff we have to make and we should keep those high capability in a limited access, like nuclear power. I hope people working on open models and AI regulation realize this soon.
Furthermore, we need more protection for open models beyond just knowledge removal. First, we cannot cleanly remove potentially malicious knowledge from an AI model. Second, the model might possess supreme reasoning capabilities, meaning it could still cause great harm if given an internet connection. Because of these persistent risks, open models require high “tamper resistance.” This metric determines how strongly a model resists attempts to bypass its safety guardrails through fine-tuning, ensuring it remains just as safe as the day it was released. Benchmarks like TamperBench measure this degree of resistance, and we urgently need more research into effective tamper-resistant training approaches.
Suppose a frontier model is or soon to be launched to the world but fails to meet these standards. What should we do? We need financial and systemic incentives that favor safe models over unsafe ones. This can come from:
Certification: Like Dean has mentioned in point 25, certified companies could get green lights from governments, allowing their models to be used across a wider range of industries and locations.
Public Reputation: Nobody wants to hear their models are causing severe harm and companies have a massive incentive to avoid public criticism. For example, when news breaks about a model like Grok generating CSAM (Child Sexual Abuse Material), the public rightfully condemns the company. Cyber/CBRNE risks, by comparison, are highly technical and abstract, making them harder for the general public to graphs until a catastrophe actually occurs. We need more vivid stories, analogies, and red-teaming examples to raise public awareness.
Liability: In an ideal world, we can trace a crime back to identify who or what enabled it the most. Normally, for misuse risks, we hold the bad actors liable. However, I believe we need to hold model developers liable as well if it turns out their AI significantly facilitated the act.
Consider the extreme "paperclip maximizer" thought experiment: a model with a benign goal, making as many paperclips as possible, causes a catastrophic outcome because it is supremely capable but poorly aligned. In this scenario, the AI developers are clearly liable because there is no one else to blame. Returning to real-world misuse, if a bad actor gives a malicious goal to an agentic system, and the system actively helps rather than refuses, shouldn’t we hold the developers of that system liable? Think of it like this: if a person commits a crime, and we later realize their parents heavily encouraged and equipped them to do so, the parents bear responsibility too. I agree the lines are sometimes blurry, but we can start with clear-cut malicious cases to bring developer liability to the table
I wrote this post as a reminder to myself and to the public: the technical solutions required to keep closed models safe are just one part of the game. We have a much wider ecosystem to protect.
Chemical, Biological, Radiological, Nuclear and Conventional Explosive
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.