
Our goal is to address worst-case risks from the development and deployment of advanced AI.
Our agendas
Research agenda
Model Persona Research Agenda
This agenda studies and steers the emergence of malicious propensities in LLMs — traits like spitefulness, sadism, and punitiveness. We treat personas, bundles of correlated traits, as a useful abstraction for how propensities generalise out-of-distribution, and as a target for interventions.
Research agenda
Safe Pareto Improvements Research Agenda
Safe Pareto improvements (SPIs) are modifications to agents’ bargaining strategies that make all parties better off, regardless of their original strategies — an unusually robust approach to preventing catastrophic conflict between AI systems. This agenda addresses the risk that early AI development forecloses the option to adopt them.
Explore
Team
Researchers, staff and advisors.
Transparency
Budgets, plans and annual reviews since 2013.
Donate
Donations fund our research on worst-case risks from advanced AI.
Mailing list
Occasional updates on our research and our open roles.