LessWrong · Aug 21, 2026
Content-based privilege: transformer residual streams stratify by proximity to the model's own prediction
0Sign in to vote or save

The directions nearest a model's prediction decide what kind of answer you get. The next ones out decide where it goes — about five tokens later. Preprint: https://arxiv.org/abs/2608.12447 ; Supplementary materials ; Code . What do we mean when we say that a transformer model has privileged geometry? I honestly wasn't sure about that when I started down this rabbit hole, because that wasn't the…

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.