RSSAmplifier

Blog

inFERENCe

posts on machine learning, statistics, opinions on things I'm reading in the space

inference.vcRSS feed ↗15 posts

Latest posts

The Future of Software

The world of software is undergoing a shift not seen since the advent of compilers in the 1970s. Compilers were the original vibe coding : they automatically generate complex machine code that human programmers had to manually write before. Over time, compilers became fully trusted, nobody has to look under the

Deep Learning is Powerful Because It Makes Hard Things Easy - Reflections 10 Years On

Ten years ago this week, I wrote a post called "Deep Learning is Easy - Learn Something Harder". The post blew up, top spot on HackerNews. Needless to say, it didn't age well.

Discrete Diffusion: Continuous-Time Markov Chains

A tutorial explaining some key intuitions behind continuous time Markov chains for machine learners interested in discrete diffusion models: alternative representations, connections to point processes, and the memoryless property.

We may finally crack Maths. But should we?

Automating mathematical theorem proving has been a long standing goal of artificial intelligence and indeed computer science. It's one of the areas I became very interested in recently. This is because I feel we may have the ingredients needed to make very, very significant progress: a structured search

Mortal Komputation: On Hinton's argument for superhuman AI.

Last week in Cambridge was Hinton bonanza. He visited the university town where he was once an undergraduate in experimental psychology, and gave a series of back-to-back talks, Q&A sessions, interviews, dinners, etc. He was stopped on the street by random passers-by who recognised him

Autoregressive Models, OOD Prompts and the Interpolation Regime

A few years ago I was very much into maximum likelihood-based generative modeling and autoregressive models (see this , this or this ). More recently, my focus shifted to characterising inductive biases of gradient-based optimization focussing mostly on supervised learning. I only very recently started combining the two ideas, revisiting

We May be Surprised Again: Why I take LLMs seriously.

"Deep Learning is Easy, Learn something Harder" - I proclaimed in one of my early and provocative blog posts from 2016. While some observations were fair, that post is now evidence that I clearly underestimated the impact simple techniques will have, and probably gave counterproductive advice. I wasn'

Implicit Bayesian Inference in Large Language Models

This intriguing paper kept me thinking long enough for me to I decide it's time to resurrect my blogging (I started writing this during ICLR review period, and realised it might be a good idea to wait until that's concluded) Sang Michael Xie, Aditi Raghunathan, Percy

Eastern European Guide to Writing Reference Letters

Excruciating. One phrase I often use to describe what it's like to read reference letters for Eastern European applicants to PhD and Master's programs in Cambridge. Even objectively outstanding students often receive dull, short, factual, almost negative-sounding reference letters. This is a result of (A)

Causal inference 4: Causal Diagrams, Markov Factorization, Structural Equation Models

This post is written with my PhD student and now guest author Patrik Reizinger and is part 4 of a series of posts on causal inference: Part 1: Intro to causal inference and do-calculus Part 2: Illustrating Interventions with a Toy Example Part 3: Counterfactuals ➡️️ Part

On Information Theoretic Bounds for SGD

Few days ago we had a talk by Gergely Neu, who presented his recent work: Gergely Neu Information-Theoretic Generalization Bounds for Stochastic Gradient Descent I'm writing this post mostly to annoy him, by presenting this work using super hand-wavy intuitions and cartoon figures. If this isn&

Notes on the Origin of Implicit Regularization in SGD

I wanted to highlight an intriguing paper I presented at a journal club recently: Samuel L Smith, Benoit Dherin, David Barrett, Soham De (2021) On the Origin of Implicit Regularization in Stochastic Gradient Descent There's actually a related paper that came out simultaneously, studying full-batch gradient descent

An information maximization view on the $\beta$-VAE objective

guest post with Dóra Jámbor This is a half-guest-post written jointly with Dóra, a fellow participant in a reading group where we recently discussed the original paper on $\beta$-VAEs: Irina Higgins et al (ICLR 2017): $\beta$-VAE: Learning Basic Visual Concepts

Some Intuition on the Neural Tangent Kernel

Neural tangent kernels are a useful tool for understanding neural network training and implicit regularization in gradient descent. But it's not the easiest concept to wrap your head around. The paper that I found to have been most useful for me to develop an understanding is this one:

Notes on Causally Correct Partial Models

I recently encountered this cool paper in a reading group presentation: Rezende et al (2020) Rezende Causally Correct Partial Models for Reinforcement Learning It's frankly taken me a long time to understand what was going on, and it took me weeks to write this half-decent explanation of