James Brobin
Publishes 1 feed
LessWrong
A community blog devoted to refining the art of rationality
27 posts · contributor
Lately
5 Things I Learned About People From Doing Stand-Up Comedy
The Instrumental Convergence of Crowds
Humans Are Alignment Generators
Selection for Selectability: Inductive Biases in Evolution and in Neural Networks
How the Sausage is Made - Why we Hate Slop
Rogue Scalpel: Activation steering breaks refusal, even with benign directions
Content-based privilege: transformer residual streams stratify by proximity to the model's own prediction
A J-Space-Based Metric for Model Valence: Defining the Metric, Testing, and Comparisons to Self-Reports
Alignment fine-tuning induces conditional misalignment in Qwen2.5-7B-Instruct
When is Unlimited Optimization Catastrophic?
AI Text Watermarking Is Free And Good
Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments
Everything on this page was read from markup James Brobin published — a rel="me" link, an h-card, or the feed’s own author element. Nothing was inferred from anywhere else. To correct or remove it, get in touch. Machine-readable: JSON
