The typical language model is autoregressive (AR): it predicts one token at a time, left to right, each token conditioned on the ones before it. This bakes in two limitations: it spends one forward pass per token, so latency grows linearly with length regardless of how hard each token is to predict, and it cannot revise earlier tokens if later context makes them appear wrong. Diffusion language…
A while back I wrote about language modeling without neural networks , where I generated Shakespeare with an unbounded n-gram model: no weights, no training, just counting. Fortuitously, I came across the paper Language Modeling is Compression , which mentioned the compression–prediction equivalence : every prediction model is inherently a compressor, and all compression algorithms are prediction…
Generating Shakespeare has become the “Hello World” of language models. 1 Recently, I’ve been messing with alternative language models and came across unbounded n-gram models. These models are purely statistical and don’t require optimizing weights or training. A year ago, I read the paper Infini-gram , which scaled an unbounded n-gram model to trillions of tokens. While…
In this post, I run small experiments showing that diffusion language models generate code (and other structured text) at a faster rate. Increased stucture tends to correlate with reduced entropy, which leads to higher confident token predictions, which directly means more tokens decoded in parallel per step. 1 Speculative Decoding and Diffusion Language Models # In speculative decoding (for…
For Cal Hacks 2025 , a few friends and I built Curserve , a fast and scalable server-side engine for agentic coding, which ended up placing for one of the sponsor prizes. We didn’t go to Cal Hacks to try and win, but instead to have a good excuse to work on a potential research idea. This post documents our original hackathon project, our exploration into actual research, and our…
A while back, Google DeepMind unveiled Gemini Diffusion , an experimental language model that generates text using diffusion. Unlike traditional GPT-style models that generate one word at a time, Gemini Diffusion creates whole blocks of text by refining random noise step-by-step. I read the paper Large Language Diffusion Models and was surprised to find that discrete language diffusion is just a…
Here are some notes I wrote over this topic. I’ve switched my master’s thesis to a different topic, but there were many interesting research directions I found in this area. Local SGD and DiLoCo Overview # It is October 15th, 2025. For my last year of my master’s, I decided to a thesis around distributed low-communication training. Essentially, how can we train large models…
A few weeks back, I implemented GPT-2 using WebGL and shaders ( Github Repo ) which made the front page of Hacker News . Here is a short write-up over what I learned about old-school general-purpose GPU programming over the course of this project! Above is a gif of the final demo, which you can run locally via the github repo above. This article appeared on Hacker News. Link to the discussion here…
These are my notes from Mark Maxwell’s courses — Probability I and Introduction to Mathematical Statistics — and his textbook, Probability & Statistics with Applications, Second Edition . I’ve stitched them together into one reference (with much help from Claude). There’s a clean way to think about the difference between probability and statistics. Probability reasons forward :…
Lately, I’ve been coming across many blogs that have weird font-size rendering issues for code blocks on iOS. Basically, in a code snippet, the text-size would sometimes be much larger for some lines than others. Below is a screenshot of the issue from a website where I’ve seen this occur. I found this example from Hacker News while I was on phone. As you can see, the text-size…
For a while, I wanted to build a complete autograd engine. What is an autograd engine, you might ask? To find the answer, we first must know what a neural network is. Neural Network Crash Course # A neural network can just be seen as a black-box function. We pass in an input into this black box and receive an output. Normally, in a function, we define the rules on how to manipulate the input to…
Lexical Analysis and ASTs # Recently I was going through Thorsten Ball’s “Writing An Interpreter in Go”. In this book, you create a basic interpreted language and write a lexer, parser, evaluator, and REPL for it. A Lexer takes in source code and turns it into an intermediate representation, usually in the form of a string of tokens. This is called Lexical Analysis. A parser…
These are my notes from Qiang Liu’s Machine Learning II course at UT Austin, cleaned up and stitched into a single story (with much help from Claude). Almost everything in machine learning eventually comes down to the same move: you have a loss function that measures how wrong your model is, and you want to make it smaller. The model has parameters $\theta$ — sometimes a handful, sometimes a…
These are my notes from Eunsol Choi’s NLP class at UT Austin, cleaned up and stitched together into a single story (with much help from Claude). Before Transformers ate the world, NLP was a patchwork of ideas that each solved one piece of the puzzle. You had classifiers that could tell spam from not-spam, sequence models that could read a sentence left to right, and embeddings that turned…
These are my notes from working through Gilbert Strang’s Introduction to Linear Algebra . Introduction to Vectors # The core of linear algebra is vector addition and scalar multiplication. Combining these two operations gives us a set of linear combinations. $$ c\mathbf{v} + d\mathbf{w} = c\begin{bmatrix} 1 \\\ 2 \end{bmatrix} + d\begin{bmatrix} 3 \\\ 4 \end{bmatrix} = \begin{bmatrix} c + 3d…
Why Rust for Front-End Development # I’ve been using React and Next.js for front-end development ever since high school, it was one of the first few things I learned when it came to programming. Recently, I’ve had the itch to learn something new, specifically Rust front-end. As someone with a “.rs” domain, it felt like an inevitable fate. Finally, I can say I put the “.rs”…
A short review of Calculus 1, 2, and 3, based on Calculus: Early Transcendentals (Eighth Edition). Differentiation Rules # Product Rule # If $f$ and $g$ are both differentiable, then $$\frac{d}{dx}[f(x)g(x)]=f(x)g^\prime(x)+g(x)f^\prime(x)$$ Quotient Rule # If $f$ and $g$ are differentiable, then $$\frac{d}{dx}\bigg[\frac{f(x)}{g(x)}\bigg]=\frac{g(x)f^\prime(x)-f(x)g^\prime(x)}{[g(x)]^2}$$…
This is a poem I wrote my senior year of high school in AP Literature. Here we all are, this mountain we climb, the sure ascent, which lasts a lifetime. At the golden summit, a goal we all seek; the meaning of life, at its Godly peak. Up we should go, a noble direction. Yet why do so many rebel in rejection? Up is worthwhile, this mountain we climb, at the apex is all that’s sublime.
Hi! My name is Nathan Barry . Outside of school and work, I enjoy reading , ballroom dancing, climbing, pickleball, badminton, hiking, playing the guitar, and photography ! For information about me, below are a few things I’ve been doing. Currently Aug 2026 - Present CTO and Co-Founder at Tarski Research (YC F26), working on something new! Previously May 2026 - Aug 2026 A Research Engineer…
Below is a list of the books I’ve read since I was 13. The ones in italics stood out to me. 2026 - Age 23 # Breakneck - Dan Wang Lee Kuan Yew - Graham Allison, Robert D. Blackwill, & Ali Wyne Richard Nixon - John A. Farrell The Anatomy of Fascism - Robert O. Paxton
General # “I believe that a man should strive for only one thing in life, and that is to have a touch of greatness” — Félix Martí-Ibáñez “I wish to preach, not the doctrine of ignoble ease, but the doctrine of the strenuous life, the life of toil and effort, of labor and strife; to preach that highest form of success which comes, not to the man who desires mere easy peace, but to…