Towards a Bitter Lesson of Optimization: When Neural Networks Write Their Own Update Rules
We explore what can be the future of neural network parameter optimization
A blog about Machine learning, programming, and other interesting things.
We explore what can be the future of neural network parameter optimization
Visualizing the hidden 3D geometry behind Layer Normalization and uncovering the mathematical trick that makes RMSNorm tick.
Beyond numerical stability, we investigate an often overlooked hyperparameter in the Adam optimizer: epsilon.
Testing whether semantic relatedness of instructions affects a model's ability to follow them under cognitive load
We strip down mechanistic interpretability to three key experiments: watching a model 'think', finding where it stores concepts, and performing 'causal surgery' to change its 'thought process'