LessWrong · Aug 21, 2026
Alignment fine-tuning induces conditional misalignment in Qwen2.5-7B-Instruct
0Sign in to vote or save

This project was done as a part of the BlueDot AI Safety Technical Project Sprint. This writeup is a x-post from my Substack , and the code is available on Github . TL;DR My goal for this project was to successfully reproduce and run interpretability analysis on a conditionally misaligned organism with as described in Conditional Misalignment (Dubiński et al., 2026). The conditionally misaligned…

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.