RSS Amplifier

Science with Sabine · Jul 22, 2026

AI Introspection, Population Collapse, & DeepMind’s Safety Woes

0
Sign in to vote or save

Mete · Science with Sabine

Researchers from New York University have challenged claims that large language models can monitor their own internal processing. One such claim came from Anthropic researcher Jack Lindsey who altered Claude’s internal activity to insert concepts such as “apple”. He found that Claude sometimes detected and named the inserted concept, which Lindsey interpreted as evidence of introspection. Another group likewise injected words into Llama’s internal activity and found that Llama seemed to take note of it.

But the New York team tested Meta’s Llama 3.1, Alibaba’s Qwen 2.5, and Google’s Gemma 3 with additional controls. They found that the models could not reliably distinguish an alteration of their internal activity from an ordinary prompt pushing them towards the same topic. Llama’s apparent ability to classify its own internal states largely disappeared when the categories no longer matched the meaning of the sentences. The results suggest that the models were detecting anomalies or semantic patterns, not examining their own processing.

That distinction matters because genuine introspection could make model reports about uncertainty, hidden goals or manipulation more trustworthy, and has also been proposed as one possible indicator of machine consciousness.

Paper here.

A team of researchers from Italy and the UK have calculated how quickly human civilization could collapse. They developed a single equation that reproduces the main patterns of global population growth over the past 12,000 years, including long plateaus, rapid acceleration and the slowdown since the 1970s. Under the present trend, their model finds no impending population catastrophe. But in a worst-case scenario where war, climate disruption or a pandemic suddenly cripples food production and other systems that sustain today’s population, it predicts that humanity could shrink by half by 2064. It’s a speculation, not a forecast, but it demonstrates that the world population could collapse extremely quickly.

Paper here. More here.

Google DeepMind chief Demis Hassabis has proposed a US-led body that would test the most powerful artificial intelligence models before release and could eventually block deployment or coordinate a slowdown if they proved too dangerous. He says a system with all the cognitive abilities of the human brain may be only a few years away, and wants advanced models tested for cyber, biological and deceptive capabilities up to 30 days before release.

The next day, former DeepMind research scientist Alex Turner said he had quit because Google signed a classified Pentagon deal allowing Gemini to be used for “any lawful government purpose”. Turner says the contract only states that Gemini “should not” be used for domestic mass surveillance or autonomous weapons without human oversight, while giving Google no veto over lawful operations. He says a 25-page alternative he sent Hassabis, containing enforceable red lines and independent oversight, was never properly evaluated. Google removed explicit prohibitions on weapons and surveillance from its artificial intelligence principles in 2025, replacing them with case-by-case judgements about whether benefits outweigh risks.

No posts

Read the original on sciencewtg.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.