Center on Long-Term Risk

Our goal is to address worst-case risks from the development and deployment of advanced AI.

Our agendas

Research agenda

Model Persona Research Agenda

This agenda studies and steers the emergence of malicious propensities in LLMs — traits like spitefulness, sadism, and punitiveness. We treat personas, bundles of correlated traits, as a useful abstraction for how propensities generalise out-of-distribution, and as a target for interventions.

Read the agenda →All outputs →

Research agenda

Safe Pareto Improvements Research Agenda

Safe Pareto improvements (SPIs) are modifications to agents’ bargaining strategies that make all parties better off, regardless of their original strategies — an unusually robust approach to preventing catastrophic conflict between AI systems. This agenda addresses the risk that early AI development forecloses the option to adopt them.

Read the agenda →All outputs →

Explore

  • Team

    Researchers, staff and advisors.

  • Transparency

    Budgets, plans and annual reviews since 2013.

  • Donate

    Donations fund our research on worst-case risks from advanced AI.

  • Mailing list

    Occasional updates on our research and our open roles.

Read the original on longtermrisk.org ↗