All opinions in this post are my own and don’t reflect those of my employer. Which is a shame because the blame has to go somewhere…
METR’s First Frontier Risk Report (tweet thread): Anthropic, Google, Meta, and OpenAI gave METR access to test their best internal models and to review non-public info about capabilities, alignment, and control to assess loss of control risks. Huge win for the public that we have a third-party actor capable of conducting this kind of review.
The most important takeaway for me (outside of a positive update that frontier AI companies are willing to work with METR on this) was an appreciation of how often AIs cheat on long tasks:
This cheating is much more common on the hardest tasks: for tasks that are over 8 hours long in Time Horizon 1.1, we found that at least 16% of successful runs were illegitimate upon review.
The appendices contain a lot of great incident descriptions:
Shah Abbas the Great had a trained squad of cannibals who would eat prisoners alive on command. I feel like we’ve lost the art of rulers having ‘trained squads of X to do Y to prisoners’.
Top candidates for big pushes on technical AI safety right now from Ryan Greenblatt:
Pessimization training: make a somewhat a priori plausible training run where, for everything that can possibly vary, we set it to whatever setting we believe makes the most concerning types of misalignment as likely as possible.
One of the better depictions I've seen of why different groups land on very different plans for what to do about AI (the numbers are definitely handwavy):
Chinese frontier AI progress slowing relative to the US: CAISI estimates the gap went from ~4 months (January 2025) to ~8 months now.
The Third Wave of American Philanthropy: Nan Ransohoff kicked off a lot of discourse this last week with her post popularizing the fact that there will be hundreds of billions of dollars of new philanthropic money online post OpenAI and Anthropic IPO.
I think this is exciting, and I agree with lots of the post, but I’m concerned that a lot of the reaction is “how should philanthropies spend billion dollars” rather than “how should you do something good where money isn’t the binding constraint”. The former leads to grift, the latter feels like a better intuition pump for coming up with meaningful scalable ideas.
Modern fathers spend 4x as much time with their kids as their grandparents did:
Mothers are also spending more time than parents in previous generations, so there's a general trend of spending more time with kids, though it's a significantly greater jump among men.
An evolutionary theory for PMS - it motivates women to leave infertile men:
Timelines to What: A Proposal: My colleague Trevor Levin writes about the omnipresent topic of the timelines to AGI, but with the far more useful frame of focusing on what decision will affect AI outcomes and where there's leverage:
Sometimes, one’s “timelines” are treated as a kind of deadline: “Nice plan you’ve got there, but we’ll probably have TAI before that point.” But the decisions that matter most for AI’s impact on the future seem likely to take place in the years before and after some of these key milestones.
How Go Players Disempower Themselves to AI. Excellent firsthand look at the way Go AI engines have changed the practice and field of Go, written by someone who worked as a Go teacher.
The illusion of control that AI users have reliably shown interacts in an insidious way with their disempowerment. It contributes to a society of Go players that allow their participation in culture to be automated away. They are moreover so disempowered about it that they have built-in psychological mechanisms to keep them from ever recognising their own obsolescence. This mechanism even works to sabotage the detection of AI use in others. People tend to give overly conservative estimates of the chances a given game involves AI. I think this happens because they usually consult their own AI to check a suspected game. In doing so, they also come around to the machine’s point of view and conclude that playing the correct AI move was the “natural” thing to do anyway in that situation…
The thing I want to impress with this article is the consistency with which we as a species underestimate our own willingness to give up our culture, economy and autonomy to AI, even without monetary incentives
If you’re looking for ways to contribute to the public understanding of what’s going on with AI, write-ups like this are awesome and don’t require you to have a PhD in machine learning - I wish we had a lot more of them!
xoxo,
Ben

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.