RSS Amplifier

LessWrong · Aug 20, 2026

Cross-Dataset Transfer Evaluation of Deception Probes in Smaller Models

0
Sign in to vote or save

Summary Recently, Apollo Research tested whether linear probes could identify honest and deceptive responses from Llama-3.3-70B-Instruct and reported AUROC values between 0.96 and 0.999. To test some of their claims, I used the scores Apollo released to recalculate the nine values they published, reproducing them exactly. I used the same method on five smaller open models, each having between 1…

See it on lesswrong.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.