LessWrong · Aug 20, 2026
Cross-Dataset Transfer Evaluation of Deception Probes in Smaller Models
0Sign in to vote or save

Summary Recently, Apollo Research tested whether linear probes could identify honest and deceptive responses from Llama-3.3-70B-Instruct and reported AUROC values between 0.96 and 0.999. To test some of their claims, I used the scores Apollo released to recalculate the nine values they published, reproducing them exactly. I used the same method on five smaller open models, each having between 1…

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.