Excited that the last work back from my PhD is out! We proposed a method to train language models to control the factuality–informativeness trade-off in their responses based on user preferences.
@SaraZiweiGonghas been working on a lot of interesting stuff at the intersection
My first paper
@AnthropicAIis out! We show that Chains-of-Thought often don’t reflect models’ true reasoning—posing challenges for safety monitoring. It’s been an incredible 6 months pushing the frontier toward safe AGI with brilliant colleagues. Huge thanks to the team! 🙏
Life update: I’m excited to share that I’m joining the Alignment Science team at
@AnthropicAIas a Member of Technical Staff/Research Scientist. I’ll be focusing on AI safety. Looking forward to it!
