Say you want to build your own LLM. 1 One of the oldest practices in the book is to generate high quality synthetic data using a state-of-the-art model , and then use that as your training data. This gives you: A lot of training data in mere minutes, without having to build your own data collection, web crawling (and so on) pipeline Really high quality data, since the model you are using to…
This is my regular monthly post to track my AI tool ecosystem. Here is the May post . No changes to last month's stack. Current Stack Overview Core LLM Subscriptions ($240/month total) Anthropic Claude Max ($200/month) I use Claude Code for all my vibe coding and managing my Obsidian knowledge base. Still extremely happy with this; mostly thanks to Opus 4. If you are not yet a hardcore Claude Code…
I was invited as a panelist to GW Statistics Department's 90th birthday conference on "Statistics in the Age of AI" . Here are my thoughts from the panel discussion and the ideas that shaped my responses. I'd like to thank the conference organizers again for the invite, Xiaoke , Judy , Tanya and Subrata . The panel was chaired by Ron Wasserstein ; the other panelists were Gizem Korkmaz , Jiayang…
I've decided to start a monthly series tracking my AI tool ecosystem. This serves two purposes: documenting my own workflow changes and providing a practical reference for those curious about how I'm using these tools day-to-day. Each update will cover subscription changes, new discoveries, and assessments of what's working (and what isn't). Core LLM Subscriptions ($240/month total) OpenAI:…
Spoiler: Surprisingly good, but maybe don't trust them on Green. Heads Up: This post is for Magic: The Gathering players. You've been warned. We're deep in the weeds here, talking about the wizard poker podcast equivalent of pre-season NFL predictions. For non-Magic players, the underlying idea (analyzing expert guesses) might still resonate. You might want to read the first section and checkout…