Deep Research is one of the breakout agentic use cases of 2025. All of the frontier labs have integrated these capabilities into their AI products. After six months of working with them, I’ve concluded there are some situations where it makes more sense to “roll your own” agent, rather than to use an existing agent “out of the box”.
The science of forecasting It is difficult to make predictions, especially about the future. So runs the modern proverb, variously attributed to Yogi Berra, Niels Bohr and Danish politician Karl Kristian Steincke. (It was probably the latter who originally coined it though he rarely gets the credit because no one’s heard of him.) We (i.e. [ ]
“Reasoning models” are all the rage. GPT-4o1 was the first commercial offering built to reason. Now, Gemini 2.0, Grok-3 and DeepSeek-R1 all do it too. As to how they each do it well, for the most part details are sketchy thanks to the leading labs’ culture of keeping applied research under wraps. OpenAI, for [ ]
Alignment post-training is the art of getting an AI model to complete tasks in a way that meets the criteria of its users. In this blog, I document what I learned getting this process to work for a production-grade LLM.
How capable are LLMs? Benchmarking leaderboards give us insights into this question. But what do the benchmarks actually mean? In this post, I cover six of the most widely reported benchmarks and offer plain English interpretations of the results.
If you're developing with Large Language Models for the first time, you'll have questions. Here are 8 of the most common roadblocks you'll face and some pointers to help you resolve them.