RSSAmplifier

Blog

Kola Ayonrinde

The technical blog of Kola Ayonrinde: Research Scientist/ML Engineer

kolaayonrinde.comRSS feed ↗10 posts

Latest posts

Shazeer Typing

Tensor Naming for Sanity and Clarity

SAEBench: A Comprehensive Benchmark for Sparse Autoencoders

Standard SAEs Might Be Incoherent: A Choosing Problem & A “Concise” Solution

MDL-SAEs: Interpretability as Compression

Mamba Explained

The State Space Model taking on Transformers

The Impact of Mixtral

Can You Feel The MoE?

Descriptive Matrix Operations with Einops

tldr; use einops.einsum

Dictionary Learning with Sparse AutoEncoders

Taking Features Out of Superposition

An Analogy for Understanding Mixture of Expert Models

TL;DR: Experts are Doctors, Routers are GPs Motivation Foundation models aim to solve a wide range of tasks. In the days of yore, we would build a supervised model for every individual use case; foundation models promise a single unified solution. There are challenges with this however. When two tasks need different skills, trying to learn both can make you learn neither as well as if you had…

From Sparse To Soft Mixtures of Experts

Mixture of Expert (MoE) models have recently emerged as an ML architecture offering efficient scaling and practicality in both training and inference [^1].