This site does not allow itself to be embedded. You can still read it on the original site — the toolbar below keeps your place in the directory.
Anthropic found a mental workspace inside Claude that shows what a model is thinking but not saying. It caught the model noticing it was under evaluation, and behaving better because of it. Here is why interpretability is now the missing instrument for agentic oversight.
Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.