
DO the work. Don't just talk about the work
Four questions I run on my own work before I run them on anybody else's. Including the one that made me rewrite an issue hours before it shipped.
Practitioner AI insights from a Principal AI Architect with 20+ years in data and AI. Free: frameworks and signal analysis. Paid: full "I Built X" tutorials with production code. Less Noise. More Signal.
Live Last read · last published · next check

Four questions I run on my own work before I run them on anybody else's. Including the one that made me rewrite an issue hours before it shipped.

Most teams don't have an evaluator shortage. They have an evaluation-system design problem, and it starts with which of two targets you're aiming at.

A judge missing context does not return unsure. It returns 0.85 and a fluent paragraph explaining why.

5:21 pm Eastern, June 12. Anthropic receives a directive from the US government. Within hours, two of its frontier models, Fable 5 and Mythos 5, go dark for every customer on the planet. Not a rate limit. Not a regional outage. Off. Subscribe now I read the status page from my kitchen table and had the same thought a lot of you probably had: I use one of those models every day. And then a quieter…

and WHY they don't stay certified

Field notes from three days with 7,000 AI engineers in San Francisco: what’s real, what’s bubble, and the five signals worth your attention.

I asked Claude to recommend a tool in a category I had spent six weeks building a product for. It named three companies. The one I had built was not one of them, even though the homepage said exactly what the buyer was searching for, in plain English. The product was not the problem. The problem was that the agent doing the searching had never been allowed to read the site, and the crawlers that…

The three failure modes killing enterprise AI projects before they ship

The harness is the part you actually own. Here are the four layers.

Moving a Docker Home Server from MacBook to Mac Studio