Ziqian Zhong
Publishes 1 feed
AI Alignment Forum
A community blog devoted to technical AI alignment research
10 posts · contributor
Lately
Debate Training Reduces Reward Hacking in RLAIF
Does DiffusionGemma do latent reasoning?
AI swarms are starting to pose indirect takeover risk
An anytime algorithm for mixing the computable measures
Misaligned AIs could use killer robots to take over
Four LLM loss functions → four flavors of LLM misalignment
Why do models task game?
User awareness in frontier models
R-lens: Making J-lens More Faithful on Early Layers
Returning to ARC
Everything on this page was read from markup Ziqian Zhong published — a rel="me" link, an h-card, or the feed’s own author element. Nothing was inferred from anywhere else. To correct or remove it, get in touch. Machine-readable: JSON
