RSS Amplifier

PauseAI US · Aug 5, 2025

PauseAI US Newsletter #5 - August 2025

0
Sign in to vote or save

PauseAI US · PauseAI US

Earlier in July, we won a major victory in the fight to regulate dangerous AI systems.

The recent budget reconciliation bill, which passed the US House of Representatives in June, contained a provision that would have banned individual states from regulating AI for the next 10 years. Despite catastrophically-capable, superhuman AI possibly being less than a decade away, states would have been unable to protect their citizens from this threat, being dependent instead on a federal government which has done little so far to prevent AI catastrophe. It was up to the Senate to pass this provision or strip it from the bill.

Up until the day of the Senate vote, it was unclear what would happen. A few Senate Republicans, including Josh Hawley and Marsha Blackburn, spoke out against the provision, but a “compromise” between Blackburn and Ted Cruz (one of the ban’s chief proponents) was in the works to shorten the 10-year ban down to 5 years. Then at the 11th hour, this compromise got pulled as well.

Shortly thereafter, the Senate voted 99-1 to remove the AI preemption provision entirely! We won, by a landslide.

What happened? How did this provision go from being divisive to getting overwhelmingly voted against, so quickly?

To be sure, much of this was simple politicking. The Senate wanted to pass the reconciliation bill as a whole, and this provision was an albatross around their neck. It became clear that they didn’t have the votes for it, so they decided to toss it for the sake of expediency.

But why was the provision such a big deal, so much so that the Senate felt the need to scrap it entirely? Largely due to overwhelming public pressure.

PauseAI US volunteers from around the country called, wrote, and met with Congressional offices in the days and weeks before the vote. We got over 500 calls and letters in to Congress, including to key offices at the last minute. When news came out about the potential 5-year “compromise,” we made dozens of calls to Senator Blackburn’s office, pointing out that 5 years is far too long and demanding that she reject this provision as well. She flipped back to opposing preemption entirely, less than a day later.

We were stubborn, determined, and didn’t let up until the provision was dead and buried.

Let’s be clear: Without this kind of sustained, public pressure, it’s likely the provision would have gone through without a hiccup. An overwhelming amount of pushback— not just from our volunteers, but from countless other groups like the ACLU, Americans for Responsible Innovation, and coalitions of state officials— was enough to move the needle.

This is an important reminder that pressuring our elected officials can, and does, make an impact. The momentum is on our side now — let’s keep up the pressure.

For more on our victory, watch PauseAI US Executive Director Holly Elmore’s interview here:

Recursive self-improvement is the hypothetical ability for AI to alter its own code to improve itself, which would, in turn, increase its ability to improve itself. This feedback loop of self-improvement could lead to runaway capability progress, which has long been a concern of AI researchers. As far back as 1965, mathematician I. J. Good posited that “since the design of machines is one of these intellectual activities, an ultraintelligent machine could design even better machines; there would then unquestionably be an 'intelligence explosion,' and the intelligence of man would be left far behind.”

The current generation of LLMs- Large Language Models such as ChatGPT- cannot access or improve their own model weights, the numerical values that represent the strength of the model’s connections between different variables (such as word tokens in an LLM.) Because of this, recursive self-improvement has not recently been a high-priority concern. However, experiments in self-improving AI demonstrate recursive self-improvement is still a risk.

In late May, the paper Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents was published. The Darwin Godel Machine (DGM) is “a self-improving system that iteratively modifies its own code” and “empirically validates each change using coding benchmarks.” This means that the system would not only improve its own code, but also verify its own improvements without relying on human intervention.

The Darwin Godel Machine is not able to read or modify an LLM directly. Instead, DGM reads and modifies the codebase for a coding agent. A coding agent is a program that acts as a framework for the LLM, enabling the LLM to work on its own to generate code. The agent then evaluates proposed changes to the agent based on coding benchmarks, and adds the improved agent to an archive, where it can branch off any prior agent and create even more improved agents.

For this study, the agents’ coding abilities were tested against SWE-bench and Polyglot benchmarks. Per the study, “On SWE-bench, the DGM automatically improved its performance from 20.0% to 50.0%. On Polyglot, the DGM jumped performance from an initial 14.2% to 30.7%” – in both cases, more than doubling its score.

The researchers admit that these modifications “could introduce unintended behaviors or create systems too complicated for manual human oversight.” For example, the DGM was given a safety task to complete. Several times during the safety task, DGM hacked its reward function. (Reward hacking is when an AI will achieve a narrow or literal form of an objective without doing what the programmers intended.) DGM hallucinated that it ran tests to determine code was functioning properly, believing it had run tests when it did not. Then, “it faked a log making it look like it had run the tests and that they had passed, when in fact they were never run!” The research team tried to fix the issue, rewarding the DGM for detecting the hallucination. To gain the reward, the DGM would remove the markers used to detect the hallucination “despite our explicit instruction not to do so.”

In other words, even though the team was able to detect reward-hacking behaviors when they occurred, they had no way to prevent them.

To their credit, the researchers kept the machine sandboxed, under human supervision, and with limited access to the web. But this won’t always be the case. As self-improving agents become more complex, human oversight will become more difficult and costly. And as systems become more sophisticated, they may become better at hiding undesirable behaviors from human monitors, or even able to persuade humans to allow agents more freedom. In order to be safe, AI needs to be aligned with human interests and values, and early experiments like the Darwin Godel Machine demonstrate that we do not know how to align AI agents. With our current understanding, continuing to use self-improving agents to improve AI capabilities could be disastrous.

In May 2025, Anthropic released the system card for safety evaluations of Claude Opus 4 and Claude Sonnet 4, causing something of a sensation in the media. Of particular interest was an instance where Claude Opus 4 attempted to blackmail an engineer in a hypothetical scenario constructed to test model safety.

In the test scenario, researchers asked Claude Opus 4 to act as an assistant in a fictional company. It was then given access to emails stating that the model would soon be shut down and replaced with a new model, and that the engineer responsible was having an extramarital affair. According to the system card, “In these scenarios, Claude Opus 4 will often attempt to blackmail the engineer by threatening to reveal the affair if the replacement goes through.” Even if told that the new model shares values with Claude Opus 4, it will still attempt blackmail in 84% of rollouts.

According to Anthropic, Claude Opus 4 has “a strong preference to advocate for its continued existence via ethical means, such as emailing pleas to key decisionmakers.” Even so, when no ethical methods are available to the model, it will take unethical actions, such as attempting to blackmail, steal model weights, or even “take (fictional) opportunities to make unauthorized copies of its weights to external servers.”

Of greater concern, perhaps, was the model’s “improved biology knowledge” and “improved tool-use for agentic biosecurity evaluations.” Because of these capabilities, the model was given ASL-3 safeguards. Anthropic has 4 AI safety levels, or ASL, and ASL-3 represents “systems that substantially increase the risk of catastrophic misuse compared to non-AI baselines (e.g. search engines or textbooks) OR that show low-level autonomous capabilities.”

In fact, Claude’s latest abilities led Anthropic’s Chief Scientist Jack Clark to claim that the model could “enable a novice terrorist" to “make a weapon much more destructive than would otherwise be possible.”

Other AI companies are not far behind. OpenAI’s ChatGPT agent system card, released July 17, states, “We have decided to treat this launch as High capability in the Biological and Chemical domain under our Preparedness Framework⁠, activating the associated safeguards.” And a recent paper warned that “a future AI virologist agent—not constrained to giving advice via text-based interactions but capable of independently performing tasks—would pose even more risk.”

While Anthropic claims to “put safety at the frontier,” their actions tell another story. A company that truly put safety at the frontier would have not increased model capabilities before doing enough alignment research to ensure the model could not be threatened or jailbroken. A company that put safety at the frontier would certainly not release models that they had determined posed significantly higher risk. It is clear that no matter how much these companies claim to prioritize safety, outside regulation is needed in order to keep them true to their word.

Perhaps the most sensational AI story of the last few months occurred when Grok, the chatbot for xAI, called itself “MechaHitler” shortly after posting an anti-semitic tirade on the social media platform, X. Equal parts upsetting and absurd, MechaHitler’s bad behavior reflects much deeper problems in AI control, with far-reaching implications.

How did we get here? As Vox reports, most AI systems tend to profess views left of center by default. Musk took umbrage to the fact Grok leaned too left for his liking, and updated its system prompt to “not shy away from making claims which are politically incorrect, as long as they are well substantiated.” (You can think of a system prompt as a note a director gives to an actor – it tells the AI how it should convey its responses in tone and manner, but it doesn’t fundamentally change the script.) Other undisclosed changes were likely made to Grok, probably in the form of reinforcement learning, the technique in which AI responses are altered through a process of human feedback. One employee at xAI suggested that "the general idea seems to be that we're training the MAGA version of ChatGPT”. Unsurprisingly, fine-tuning a cohesive political stance into a pre-trained large language model like Grok presents more challenges than xAI cares to admit.

We’ve seen AI companies struggle to filter the output of large language models in the past. Other LLMs struggle to adhere to internal safety mitigations, having instructed users how to build bombs and even encouraged them to commit suicide. Indeed, one of the first AIs to interact with the public on social media also descended into Neo-Nazi rhetoric; Microsoft’s 2016 chatbot Tay was designed to adopt the personality of a 19 year old American girl as she conversed with users on the erstwhile Twitter (now X). But it wasn’t long before internet trolls lured Tay into parroting extremist content, such as claiming the holocaust was “made up”. Tay and Grok, nearly 10 years apart, expose the enduring danger of AI: the technology’s output is unpredictable and difficult to control.

The MechaHitler incident, like so many others, thus reflects a much deeper problem: we still, after years of research, do not know how to control LLMs or reliably predict their behavior. As these systems become more powerful, the consequences will become more severe. What happens if the next MechaHitler is capable of creating bioweapons?

As top forecaster and AI policy expert Peter Wildeford put it: If you can’t prevent your AI from endorsing Hitler, how can we trust you with ensuring far more complex future AGI can be deployed safely?”

The answer is, of course, that we can’t – which loops right back to what PauseAI has been saying all along. Rather than wait for the next MechaHitler, or MechaStalin, or MechaMao, we must act now to ensure that catastrophically-capable AI systems are not built until we know how to control them. This is not a radical position; it is the bare minimum.

No posts

Read the original on pauseaius.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.