RSSAmplifier

Blog

Deven Choudhary

Hi, I'm Deven. I work on AI safety and alignment, especially mechanistic interpretability. Work Loyal Lies — 5th place / 179, Apart Research Secret Loya...

devenchoudhary.comRSS feed ↗1 posts

Latest posts

Project Write-Up

NLAs make things up 40% of the time, and their own confidence can't tell Anthropic's Natural Language Autoencoders read a model's internal state and describe in plain English what it's thinking. People are starting to use them to audit models, including for whether a model is hiding something. The problem, which Anthropic states openly, is that they confabulate: they assert specific things that…