Inside TPU and GPU Clusters: The Anatomy of Collective Communication
A deep dive into all-gather, reduce-scatter, all-reduce, and all-to-all communication on TPUs and GPUs.
Long thinking.
A deep dive into all-gather, reduce-scatter, all-reduce, and all-to-all communication on TPUs and GPUs.
A deep dive into a modern dense transformer: YaRN, hybrid attention, soft capping, QK normalization, FLOPs/token, cluster sizing, and more.
Visualizing and analyzing sleep, heart rate, HRV, activity, stress, and SpO2 data from the Oura Ring API.
From GPU architecture and PTX/SASS to warp-tiling and deep asynchronous tensor core pipelines.
From paged attention, continuous batching, prefix caching, specdec, etc. to multi-GPU, multi-node dynamic serving at scale.
Humans in the post-ASI world.
How does Flash Attention really work?
How i landed a job at DeepMind as a research engineer without an ML degree.
How I got started with reinforcement learning.
How I got started with geometric/graph machine learning.
How I got started with transformers.
Breaking down ideas from Turing's seminal paper - part 2.
Breaking down ideas from Turing's seminal paper - part 1.
Learnings from Coursera course.
My journey into machine learning.