Why ML Conferences Have Lost Legitimacy: A Case Study in Error Propagation
A labmate of mine spent three months trying to improve upon min-p sampling — a method for generating text from language models that had been published as an ICLR 2025 Oral, the 18th highest-scoring submission that year. After months of work, he made a troubling discovery: by failing to control for hyperparameter tuning, he could make almost any sampler look like the state-of-the-art. That’s when…