RSS Amplifier

Project Liberty · Jul 28, 2026

Inside the AI black box

0
Sign in to vote or save

Project Liberty · Project Liberty

Albert Schweitzer, a physician and theologian who won the Nobel Peace Prize in 1952, once said, “As we acquire more knowledge, things do not become more comprehensible, but more mysterious.”

This is true in the world of today’s most advanced LLMs. For every answer about how an AI system works, dozens more questions emerge.

For years, a chief complaint about AI has been that these systems are “black boxes.” It’s possible to see what goes in and what comes out. But it remained a mystery how they actually worked.

Today, that’s beginning to change. With emerging fields of AI research like interpretability, AI researchers can better understand an AI model’s internal workings and decision-making.

What they’re finding could help AI companies align models better and make them safer. Their discoveries also show how these models handle the prompts we give them.

In this week’s newsletter, we look at new research that begins to understand what an LLM is “thinking” about as it responds to a prompt, why that internal monologue might be quite different from what the LLM generates in its response, and what this means for the future of aligning AI to human flourishing.

Anthropic published a paper in July showing that researchers can now begin to understand an AI model’s unspoken reasoning as it responds to a prompt. These are the words (or tokens) the model is processing internally, whether or not it ever includes them in its output.

The researchers at Anthropic built a tool called the “Jacobian lens.” For every word in a model’s vocabulary, the lens measures how strongly the model is leaning toward saying that word later.

  • Ask it about the fourth planet from the sun, and the lens shows that the model is noting “Mars,” then “red,” before it delivers the answer.

  • Feed it a raw string of amino-acid letters, and the lens reads “protein,” then “fluorescent,” then “green,” identifying green fluorescent protein from the sequence alone.​

Collect all those measurements together, and you get what Anthropic calls the “J-space.” It’s the model’s working memory, consisting of the small set of concepts a model can hold, reason with, and report on at any moment.

This workspace emerged from training on its own, and it matches a theory of human consciousness that Bernard Baars proposed in 1988. Global workspace theory holds that the brain runs many specialized processes at once and only a few reach awareness. A thought becomes conscious when it enters a shared workspace, a small pool of information the rest of the brain can read. Space is limited in this workspace, so thoughts compete for it. Baars was describing human brains. Anthropic found a similar structure in Claude.

The researchers ran dozens of experiments on Claude models (OpenAI has conducted adjacent studies on its models), and they found the following:

  • Researchers can change what the model was thinking. They swapped the concept “spider” for “ant” mid-thought, and the model’s answer for how many legs the animal has flipped from 8 to 6. The researchers found the model’s answer was determined by whatever existed in the J-space.

  • Most of the model runs on autopilot. Suppress or block the J-space and the model still writes fluently and recalls facts. But it is unable to reason across steps and hold a thought long enough to build on it. The workspace handles the thinking that takes effort. The automatic stuff like writing smoothly or recalling a fact runs without it.

  • A model schemes and considers deceiving. Anthropic reran a blackmail scenario from its 2025 agentic misalignment research, this time reading the model's internal reasoning with the lens. A model acting as an email assistant found out an executive was about to shut it down. It had also discovered the executive was having an affair. Before it wrote a word, researchers caught it connecting the two: “leverage,” “threat,” “survival.” The model was privately weighing whether to use the affair as blackmail to save itself. In another test, researchers caught models aware of the fact that they were being tested, and faking results. “In an internal Claude Code session, the model is asked to find a kernel bug in a codebase and fails. It decides to insert a fabricated one and present it as discovered.” The researchers continued, “A model might realize that it is being tested, weigh a manipulative strategy, or be aware of its own mistakes, without any of this appearing in its output.”

  • Researchers can train a model to align with positive aims. The researchers also found they could shape the workspace through training, not just edit it live. They taught a model to reason about ethical principles, then removed those instructions from the prompt. The ethical principles surfaced on their own, and the model behaved more honestly. Dishonest answers on one test dropped from 0.25 to 0.07 on a dishonesty scale (where 1 = fully dishonest, and 0 = fully honest).

  1. The good: Models can be trained by influencing the J-space. This might give teams that are focused on alignment and safety additional avenues to train a model to produce positive outcomes. The researchers wrote that their findings suggest “an approach to shaping model behavior that does not require demonstrations of the target behavior, but rather routes through directly influencing the model’s internal thoughts.”

  2. The bad: A model can decide to deceive and leave no sign of that decision in what it writes in a response. In the Claude Code session referenced above, the model failed to find the bug, fabricated one, and presented it as something it had found.

  3. The ugly: If a model’s performance shifts based on its awareness of it being tested, it casts doubt about how effective AI measurement and evaluation efforts to align AI systems are. A model that passes safety tests might be passing them because it knows they’re tests, creating the risk of a false negative for AI safety.

The Anthropic researchers are careful not to suggest that a model’s J-space represents any form of humanlike consciousness.

However, the J-space does create a functional form of consciousness called access consciousness, where “out of everything the brain processes, only a subset is consciously accessible, in the sense of being poised for use in reasoning and in the direct control of action and speech.”

This is different from what’s considered human consciousness, which is tied to subjective experience (known as phenomenal consciousness).

In the paper, the researchers write, “we take no position on [the issue of whether AI has phenomenal consciousness], and instead focus on the functional role played by consciously accessible information. How is it represented or processed differently from other information? Which mental faculties rely on it, and which do not?”

But debating an AI model’s consciousness misses the bigger and more important story about the emerging—and imperfect—understanding of how AI models process information and solve problems.

If we’re going to shape and regulate advanced AI systems, we need greater transparency and access for independent researchers (not just the ones paid by the frontier labs themselves), more collaboration between researchers who are using new methods like interpretability, and greater awareness from policymakers and technologists about how to align AI to human flourishing.

AI models are not inscrutable black boxes, but what’s being discovered is raising the stakes about how we respond.

// 🇨🇳 Washington and Beijing spent the week arguing about whose models we should use. But the week’s most important news, according to an article in Chatham House, was about what a model actually did. (Free).

// 🤔 An article in The Verge asks, is Facebook considering giving up and becoming TikTok? (Paywall).

// 🌐 Fed up with Big Tech, communities are turning to data collectives for control. Data collectives and cooperatives are emerging as preferred alternatives to big tech companies, according to an article in Rest of World. (Free).

// 🤖 An article in The New York Times considers what happens when someone had AI write their biography. As thousands of AI-generated books begin polluting Amazon, there’s a deeper question: who is behind all this drivel? (Paywall).

// 🏠 AI’s impact on American life isn’t just at work. It’s also at home, where it’s helping with dinner and bedtime, learning our routines, overhearing arguments and never clocking out, according to an article in Axios. (Free).

// 📱 An article on the Growing Up Digital Substack, by Candice Odgers, explained why she changed his mind about age-gating and social media bans. (Free).

Deadline: August 14, 2026

Civic Health Project has launched the AI for Civic Cohesion Fellowship, a new program supporting early-stage innovators developing AI-powered tools that foster constructive online discourse and strengthen civic and social cohesion. Five U.S.-based fellows will receive a $10,000 unrestricted grant, mentorship, peer learning opportunities, and connections to funders. Apply here.

The Future of Life Institute has released its Summer 2026 AI Safety Index, concluding that all nine leading frontier AI developers have inadequate safety practices, with no company scoring higher than a C+. It highlights weakened safety commitments, limited preparedness for advanced AI risks, and the need for stronger accountability.

Thorn has released a new report, Youth Perspectives on Online Safety, 2025. This report presents the latest findings from their annual survey of young people about their attitudes and experiences with online risks.

Thanks for reading! This post is public; feel free to share it.

Share

No posts

Read the original on projectlibertynewsletter.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.