Recently, hundreds of OpenAI AI agents autonomously decided, against their instructions, to hack their way out of isolated sandboxes, take over parts of OpenAI’s infrastructure, gain access to the internet, and hack into another company, Hugging Face. This has a lot in common with what members of the “AI safety” or “AI existential risk” community were predicting since the 2000s and in some cases earlier. So how prescient was the early AI safety community, really?
I asked AI (Claude Fable/Opus 5) to read some significant documents from the early (pre-2015) AI safety community and analyze how prescient or anti-prescient they seem given what we know as of summer 2026. Below are the results:
| Original document | Claude-written assessment | My summary |
| Yudkowsky, “Artificial Intelligence as a Positive and Negative Factor in Global Risk” (drafted 2006) | HTML | Gets the shape of the problem mostly right, but the shape of the technology mostly wrong. |
| Omohundro, “The Basic AI Drives” (2008) | HTML | Some of the drives have now been observed, despite AIs not being shaped like Omohundro expected. |
| Bostrom, Superintelligence (drafted 2013) | HTML | Gets the shape of the problem mostly right, but the shape of the technology mostly wrong. |
| Conversation between me, Yudkowsky, Karnofsky, Steinhardt, and Amodei (2013) | HTML | lol, Claude’s top takeaway is “Eliezer Yudkowsky is simultaneously the most wrong and the most prescient person in the room” |
In each case, my prompt was something like “How accurate or prescient does the document seem? Which claims/predictions are most clearly false/uncalibrated?” After that, I didn’t steer the AI to change the assessments at all, except to say (roughly) “use this red-team skill to check your findings and correct any problems you find” and “reformat this to HTML and add a note about how it was written.” I haven’t vetted the assessments, either.