Diversity as the bottleneck in Self-Play
Exploring plateaus in prior self-play setups.
Hamish Ivison is a University of Washington PhD student researching post-training, reinforcement learning, and data for language models.
Exploring plateaus in prior self-play setups.
A basic introduction to policy gradient for language models.
Results replicating the recent L1 paper.
Everything I 'consumed' in 2025.
Everything I watched, read, and played in 2024.
Everything I watched, read, and played in 2023.
My own experience around applying for and getting into PhD programs.
Poking around with gpt-3 and ancient languages
A quick go-over of my recent blog changes.
I made a fun little animated Ace Attorney AI script generator.