LessWrong · Aug 21, 2026
Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments
0Sign in to vote or save
TL;DR: We introduce CHIVE, an agentic pipeline that discovers unexpected LLM behaviors in the wild and explains them with counterfactual prompt edits. We use the resulting data in two ways. Using it as an evaluation, we find that activation-reading interpretability tools provide no uplift: agents given the tools predict the outcomes of these experiments no better than agents that just read the…

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.