Yanda Chen · X (formerly Twitter)

  • user avatar

    Excited that the last work back from my PhD is out! We proposed a method to train language models to control the factuality–informativeness trade-off in their responses based on user preferences.

    @SaraZiweiGong

    has been working on a lot of interesting stuff at the intersection

  • user avatar

  • user avatar

    So excited to hear that Kathy won the ACL Lifetime Achievement Award! I feel incredibly fortunate and honored to have had her as one of my PhD advisors, and I’ve learned so much from her over the years. Big congrats!

  • user avatar

    My first paper

    @AnthropicAI

    is out! We show that Chains-of-Thought often don’t reflect models’ true reasoning—posing challenges for safety monitoring. It’s been an incredible 6 months pushing the frontier toward safe AGI with brilliant colleagues. Huge thanks to the team! 🙏

  • user avatar

    Life update: I’m excited to share that I’m joining the Alignment Science team at

    @AnthropicAI

    as a Member of Technical Staff/Research Scientist. I’ll be focusing on AI safety. Looking forward to it!

Read the original on x.com ↗