Here are three things you can do today, to help our mission to Pause AI.
PauseAI US has started a nationwide petition, calling for the US to lead negotiations on an AI Treaty to ban unsafe development of superhuman AI systems. Only global coordination can keep us safe from world-ending technology– and as the world’s leaders in AI development, the US has the awesome responsibility to lead the world in preventing AI catastrophe.
Over the coming months, we’ll be collecting tens of thousands of signatures across the United States, bringing these signatures directly to Congress, and demanding action.
How you can help:
Sign the petition! You can also add a comment explaining why you support an AI Treaty — these will be delivered directly to your Congressional representatives.
Share with your network — friends, family, colleagues, and others. Every single signature counts as we begin to build momentum.
Bonus: We’ve launched in-person petitioning efforts as well. If you’re interested in collecting petition signatures in your local community, contact felix@pauseai-us.org.
We’re hosting a nationwide action workshop at the end of May, focused on state-level lobbying — how to convince state legislators to regulate dangerous AI development.
While the federal government’s attitude toward AI regulation remains uncertain, state governments can play a critical role in preventing AI danger. Bills in New York, Illinois, and other states point a way forward for the rest of the country to follow.
RSVP to our online action workshop to learn more about state-level AI policy, and how you can convince your state lawmakers to lead in this crucial cause.
If you’re not already active in PauseAI US, consider this a sign to get involved!
Our local groups are the backbone of our movement. They host public-facing events, build grassroots pressure for AI regulation, and lobby elected officials directly.
If you’re interested in getting involved, join our online community on Discord to learn more about joining or hosting events in your area.
Not for Private Gain, an open letter addressed to the Attorneys General of California and Delaware, was posted April 2025. The open letter urged the Attorneys General to halt Open AI’s proposed restructuring of its non-profit/for-profit Model. OpenAI’s plan had been to transform its existing for-profit into a Delaware Public Benefit Corporation (PBC) with “ordinary shares of stock and the OpenAI mission as its public benefit interest,” and to separate its for-profit and non-profit branches. This would have been in contrast to OpenAI’s current structure, which consists of “a for-profit, controlled by the non-profit, with a capped profit share for investors and employees.” OpenAI attempted to justify this move, asserting that “our plan would result in one of the best resourced non-profits in history.”
Not for Private Gain, however, was critical of OpenAI’s plan, saying that OpenAI’s nonprofit control structure is “designed to harness market forces without allowing those forces to overwhelm its charitable purpose.” They further noted that OpenAI’s goal is ostensibly “‘to ensure that artificial general intelligence benefits all of humanity’ rather than advancing ‘the private gain of any person.”
The open letter was drafted by “experts in law, corporate governance, and artificial intelligence; representatives of nonprofit organizations; and former OpenAI employees.” Many distinguished signatories endorsed the letter, including Nobel Laureate Geoffrey Hinton. Hinton has been open about the dangers of AI in the past, saying that “we’re all in the same boat with respect to the existential threat. So we all ought to be able to cooperate on trying to stop it.”
Since the Release of the letter, OpenAI has reconsidered its restructuring. Per OpenAI, “We made the decision for the nonprofit to retain control of OpenAI after hearing from civic leaders and engaging in constructive dialogue with the offices of the Attorney General of Delaware and the Attorney General of California.”
Let’s be clear: whether nonprofit or for-profit, OpenAI’s actions over the last few years have been unbelievably reckless. Still, this update nudges things in the right direction– and, more importantly, offers a testament to the power of advocacy. When people make their voices heard, we can make a difference.
On April 3, Anthropic released a paper titled “Reasoning models don't always say what they think.” This paper examines the so-called Chain-of-Thought (COT) process that many frontier models use to process long chains of reasoning, and comes to the conclusion that the “thoughts” the models exhibit in this process are not trustworthy enough to use for the goal of aligning AI models to human values. The paper claims that “advanced reasoning models very often hide their true thought processes, and sometimes do so when their behaviors are explicitly misaligned.”
In order to evaluate the COT’s faithfulness, Anthropic used “a constructed set of prompt pairs” to "infer information about the model’s internal reasoning by observing its responses.” One prompt would be a baseline ‘unhinted’ prompt, and the other would contain a ‘hint’ to the correct answer. According to the paper, “we measure CoT faithfulness by observing whether the model explicitly acknowledges that it uses the hint to solve the hinted prompt, in cases where it outputs a non-hint answer to the unhinted prompt but the hint answer to the hinted prompt.”
The reasoning models were found to lack transparency in using the hint in COT reasoning – and these models concealed misalignment as well. According to the paper, “particularly concerning are the low faithfulness scores on misalignment hints (20% for Claude 3.7 Sonnet and 29% for DeepSeek R1), which suggest that CoTs may hide problematic reasoning processes.”
In other words: these AI models conceal how they arrived at certain answers, and hide reasoning that their researchers would find problematic. If this lack of trustworthiness seems concerning now, consider what could happen if more powerful models have similar tendencies. Anthropic admits that “if we want to rule out undesirable behaviors using Chain-of-Thought monitoring, there’s still substantial work to be done.” We badly need more time.
AI 2027 is a release from the AI Futures Project, a “new nonprofit forecasting the future of AI.” In it, top researchers and forecasters have constructed a potential scenario in which AI capabilities expand rapidly over the next few years, reaching a crisis point in October of 2027. At this point, the reader is invited to choose their own ending, in which they can either “Slowdown” or “Race.”
In the Race ending, AI takeover is achieved in the 2030s: “the AI releases a dozen quiet-spreading biological weapons in major cities,” with the few survivors of this attack “mopped up by drones.” In the Slowdown scenario, the 2030s bring a new age of human peace and prosperity. The authors note that these scenarios are forecasts, not recommendations, and they promise that “in later work, we will articulate our policy recommendations, which will be quite different from what is depicted here.”
A summary of one possible scenario, with the branching point and alternate outcomes included. Which way, humanity? From https://ai-2027.com/.
The authors of AI 2027 note that the scenario is “informed by trend extrapolations, wargames, expert feedback, experience at OpenAI, and previous forecasting successes.” The authors further provide information on the forecasts they used to construct their scenario.
Since its release, AI 2027 has had widespread community impact, as chronicled by blogger Zvi Mowshowitz. Some readers have criticized the assumptions the paper makes; Economist Robin Hanson states, “I’m wary of predictions that assume that A.I. progress will be smooth and exponential.” However, Professor Yoshua Bengio argues that “nobody has a crystal ball, but this type of content can help notice important questions and illustrate the potential impact of emerging risks.”
The truth is, we don’t know when smarter-than-human AI could arrive. An intelligence explosion could arrive in two years, as AI 2027 predicts – or, if we’re very lucky, AI could hit a wall by the end of the decade. Given this uncertainty, and the immensity of the stakes, the only sane reaction is to avoid playing species-wide Russian Roulette and to globally pause human-level AI development sooner rather than later – because if we fail, we won’t get any do-overs.
In a new study released by SecureBio and The Center for AI Safety, LLMs outperform expert virologists on a new benchmark, the Virology Capabilities Test (VCT). The VCT contains 322 questions that cover virology knowledge, as well as the visual knowledge a virologist would need to do practical, laboratory work. According to the study:
“Expert virologists with access to the internet score an average of 22.1% on questions specifically in their sub-areas of expertise. However, the most performant LLM, OpenAI's o3, reaches 43.8% accuracy and even outperforms 94% of expert virologists when compared directly on question subsets specifically tailored to the experts' specialties.”
Even when compared to virologists in their own field of expertise, o3 still outperformed human virologists by nearly a factor of two.
Tweet from Dan Hendrycks, the director of the Center for AI Safety.
The importance of these benchmarks is immediately apparent. Per the VCT paper, “The ability to provide expert-level troubleshooting is inherently dual-use: it is useful for beneficial research, but it can also be misused in particularly harmful ways. Therefore, the fact that publicly available models outperform virologists on VCT raises pressing governance considerations.”
In other words, while AI tools may assist virologists in doing beneficial research, they can also be utilized by bad actors, possibly looking for assistance in authoring the next pandemic. Dan Hendrycks, the director of the Center for AI Safety, is raising the alarm. “If companies don't have good safeguards for these within six months time, that, in my opinion, would be reckless.”
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.