RSS Amplifier

LessWrong · Aug 21, 2026

Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments

0
Sign in to vote or save

TL;DR: We introduce CHIVE, an agentic pipeline that discovers unexpected LLM behaviors in the wild and explains them with counterfactual prompt edits. We use the resulting data in two ways. Using it as an evaluation, we find that activation-reading interpretability tools provide no uplift: agents given the tools predict the outcomes of these experiments no better than agents that just read the…

See it on lesswrong.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.