From Richard Hamming’s classic 1986 talk, “You and Your Research”, we get a deceptively simple but powerful question:
what are the important problems in your field, and why aren’t you working on them?
This question forces clarity. It helps people identify and focus on the work that actually matters—and that’s what moves a field forward.
In this post, I’m going to answer that question for myself.
I define my field as
Frontier safety misuse risk (Cyber, CBRN and Harmful Manipulation) mitigation
In other words, I focus on reducing the chances that malicious actors use powerful models to cause large-scale harm. This sits within the broader space of AGI safety and complements misalignment work like AI control.
To me, the core problem is:
How can we design a verifiable and robust "containment architecture" to prevent AI from causing severe harm through misuse’
Let’s decompose the question:
“verifiable”: We need reliable ways to confirm that residual risk is low after mitigation. This depends on realistic threat modeling and evaluation environments that reflect actual attack surfaces.
“robust”: Malicious users will actively try to bypass safeguards. Robustness means testing against a wide range of jailbreak strategies, and recognizing that models can also drift from instructions over long contexts or multi-turn interactions. Also, we should prevent malicious users from conducting attacks through multiple accounts as well.
“containment architecture”: This is not just about the base model. It includes layered defenses: guardrails, monitors, policy enforcement, and intervention mechanisms. This assumes the model itself is not perfectly aligned—otherwise it would knowingly not help the adversary.
“severe harm”: I think in terms of thresholds like ~100 deaths or ~$1B in damage. But I’m not a threat-modeling expert, and I think harm in the 10-deaths or million-dollar range can also count, depending on the context.
I actually am—mainly on the robustness and containment architecture fronts. These problems especially need people inside major labs, since:
Guardrails and monitoring systems have to be implemented at the infrastructure level
Access to real user logs gives a unique advantage for evaluating threats and mitigations
I’ll shift attention to the verifiable piece more next year, once we’re confident enough in the layers of defense we’re building now.
I haven’t personally focused much on defining “severe harm” thresholds because:
I trust colleagues with threat-modeling expertise
My strengths are on the engineering and implementation side, where I think I can have more impact
So — What’s your hamming question? And why aren’t you working on them?
I’d love to hear.
Thanks for reading Belay the Future! This post is public so feel free to share it.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.