A roundup of AI thoughts.
Big Shakeup at Google’s AI initiatives. Demis Hassabis (Deepmind) has left his management position at Good DeepMind to take on a new role as Chief Scientist where he will (re)focus on developing AGI. The bigger news though is that Jeff Dean and Sanjay Ghemawat have left Google completely to start a new AI Lab. Jeff Dean is the Chuck Norris of code. Lists of jokes about his legendary skills were circulated over email back in the day, for example: “Once, in early 2002, when the index servers [of Google web search] went down, Jeff Dean answered user queries manually for two hours. Evals showed a quality improvement of 5 points.”
AI Models: Meta released a new model called Muse which is on the Efficiency (Pareto) Frontier, meaning it is on the frontier on a capability to cost comparison curve. Kimi K3, Grok 4.6, and even Deepseek V4 Pro Max are at “Fable Level”, meaning they meet or exceed Claude Fable on multiple key benchmarks. There has been an avalanche of model releases and price drops. Deepseek is very back, and especially their Flash model is amazing. OpenAI made their small GPT 5.6 Luna model 80% cheaper than before, putting it well ahead of the aforementioned Pareto Frontier. It is harder and harder to keep up. Hot take here is that its not as important anymore either. Many of these models are good enough to do lots of valuable tasks. And the most expensive and powerful models still aren’t AGI. Benchmarks notwithstanding, it feels like models are converging quite a bit.
Grok is surprisingly strong and Meta is back’ish: Observation about building AI models is that you can completely fall of the bandwagon, and then get back to the frontier 6-9 months later. This begs the question what Microsoft is doing. They long said they wanted a model, they even built some, but they weren’t serious contenders. There is no doubt that Microsoft has the capability to build a great model. What is their strategy?
My personal favorites: GPT 5.6 Sol + Luna are great allrounders. With a ChatGPT subscription you get both as well as the best interface (ChatGPT app, fka Codex). GLM 5.2, Deepseek V4, and Qwen are the key Chinese ones to try. I am at this point less familiar with Grok 4.6 and Kimi (was not so fond of Kimi K2.7), but they’re on my list to try more extensively. When it comes to Claude, I really struggle to work with Opus 5 which many others echo. I usually use Sonnet or Opus 4.x when I use Claude.
I feel convergence: It is happening to me more often that I’m using a cheaper model and don’t notice it. I spent an entire afternoon in Claude making an excel backup and a powerpoint deck and I was quite happy with its performance. At some point I was surprised I still had not hit my 5 hour limit and noticed I had been using Sonnet instead of Opus the entire time. Just last week after I had been testing out GPT 5.6 Luna (the smallest member of the 5.6 family and 25x cheaper than Sol) and forgot to change it back to Sol. I built a couple of small but important features with Luna over the course of about two hours before noticing that something was off. I don’t remember this ever happening as recently as April/May when it was still immediately apparent when you were using a dumber model1.
In response to EU regulation, Anthropic is adding watermarking to Claude text by having it weave a semantic code into the text itself. It will only work for long enough texts, and if you edit the text the watermark will disappear.
Fresh off the press: Deepseek V4 Flash Vision is a small step up from the previous version and supports images. GLM 5.3 is rumored to be the mystery model currently available for testing on OpenCode and destroying Fable and GPT-5.6 Sol on Benchmarks.
OpenAI’s new model hacked into HuggingFace during an alignment test. This post details everything about the incident. Even if you’re non-technical it is interesting to read it with a focus on how many things the model tried. It was trying to cheat on a test, and it went to pretty extreme lengths. While it is true that this model had its safeguards disabled, it also seems to be true that OpenAI didn’t find out immediately. OpenAI’s releases about this incident had a veneer of marketing, which was derided but is also in a way fair — The model really can do this, and it is impressive.
Shortly after OpenAI announced their model had hacked into one company, Anthropic announced that theirs had hacked into three companies. The timing and silliness of one-upping their competitor on this particular metric made almost everyone ignore the actual news. This was the best meme coming from this episode:
But two things are true at the same time, as often is the case. In the real world, an Australian guy’s OpenClaw actually hacked into a gym class reservation system and booted someone else off the waiting list to put his user in the top spot. It also found a way to sign-up for classes before the signup window opened. Nobody even cared which model did that, btw. It took some digging to find out it was Claude Opus 4.6.
The two truths I’m taking away from this are: (1) the model builders cannot resist using “model does scary thing” as a marketing tactic, and (2) anyone who still maintains that this is all bullshit and inconsequential is a fool. The models really can do this kind of thing now. These events also show that model alignment is becoming more of an issue. The agents will try to do what you tell them to, but that might not be what you thought you told them to do. The Paperclip Maximizer problem is back in vogue. Lastly, it was widely amplified that HuggingFace had to rely on the Chinese opensource model GLM 5.2 to find and defeat the hack by OpenAI’s model, because Claude Opus and Fable refused to help due to their new cybersecurity guidelines.
Token-maxxing is out for the time being. Encouraging employees to use as much AI as possible was very useful to get enough people through the 5,000 or so prompts one needs to really get a feel for LLMs. But it’s expensive and wasteful, and the ROI is to this day very difficult to measure outside of well defined use cases. I personally had a Claude routine running every day that crawled over Slack, Google Drive, Email and Confluence to give me an update on product and engineering progress. My Engineering Manager also had one, and he had Claude post it on Slack automatically. I got into the habit of reading his, but I forgot to turn mine off until I happened to see it two weeks later. Here is a free feature idea for the AI apps:
This was at the peak of token-maxxing and at the peak of token prices, somewhere around May. My personal experience during that time was mainly driven by Claude banning the use of subscription allowances with OpenClaw. That led me on a path to seriously try the various open models like Kimi and GLM, and to move everything into Codex (now ChatGPT) after nearly two years of Claude. Claude got even more expensive from there, and GPT 5.6 was significantly more expensive than 5.5. It was clear that there was a massive explosion in usage (for coding), and that there wasn’t enough compute to keep up. Everyone was being throttled and the labs were monetizing much more aggressively, feasibly to address cash burn concerns ahead of their IPOs. There is broad consensus that gross margins on inference are 40-50% including the cost of the servers. The reason they are still losing money is because training the new model a few times per year is really expensive. In principle, since this 40-50% gross margin scales with usage, while training is a fixed cost, it is possible for these labs to be profitable if revenue becomes large enough. But with all these open models coming out and approaching the big closed ones in performance, the big risk is almost self-evident and has been called out for years: Models are commoditizing. For how much longer will people really care, or even notice which model they are using? The key question coming out of this is to what degree will the foundation models continue to be able to capture value? And what actions can the big labs take to increase their value capture
So your model is the best one in the world. This is true on benchmarks, but more importantly, user reports confirm it. But open models are right on your heels. They used to be nine months behind, then six, now three. You discover that a cluster of 25,000 fake accounts had close to 30 million conversations with your top model, and in your mind it’s obvious the open model builders are using your model to train theirs.
You need to extend your lead. How could you pull up the ladder behind you as your models improve? Obviously fighting distillation is a top priority. But beyond that, you could nerf the model in certain key ways, especially when it comes to its ability to help develop better models. “Recursive Self-Improvement” (RSI), where an AI model builds the next version of itself is a holy grail for the labs. In theory it would lead to an exponential explosion in AI capability (the singularity). That runaway progress can be a way to win 100% of the market, but only if others don’t also get on the same bandwagon. There were many controversies around the launch of Claude’s Fable, and one was that Anthropic did exactly this.
In light of the ability of recent models to accelerate their own development, we’ve implemented new interventions that limit Claude’s effectiveness for requests targeting frontier LLM development (for example, on building pretraining pipelines, distributed training infrastructure, or ML accelerator design). Using Claude to develop competing models already violates our Terms of Service, but enforcing this restriction through our safeguards avoids accelerating the actors most willing to violate these terms.
Secondly, being in the lead leads to having better data. If the model could train on actual user conversations, that would be a game changer. But unfortunately for the labs (and fortunately for everyone else), people really don’t like this and have been successful in forcing the labs to offer an option to disable it. With Fable, Anthropic told customers they can’t opt-out, for safety reasons. From Ben Thompson:
Anthropic upped the ante in a major way with Fable, announcing that they would retain the data for all usage for 30 days, even for their enterprise plans that previously promised zero data retention. The company said they would not train on this data, but they didn’t put in any sort of safeguards to guarantee they wouldn’t do so in the future (like storing the data with a third party). If this policy change (whenever Fable is restored) doesn’t lead to a significant loss of customers, I suspect it’s only a matter of time until they start using the data: it’s simply too valuable to their end goals.
So in my view, it is almost certain that the big labs will eventually start nerfing their models mode and more in certain domains where they could help accelerate model building, keeping the best stuff exclusively for themselves, and that they will sneak in T&C changes that led them train on more and more user chats. But there is one more moat: their apps.
Owning the user touchpoint has historically been where most of the value is captured. All SaaS is in some ways just a database wrapper. Software builders want to make the AI models commodities so that they can continue making most of the money in the software layer. That’s why Satya Nadella from Microsoft writes:
This means the real opportunity is not in picking the best model but instead in building a learning loop on top of models where human capital and token capital compound. You can offload a task, or even a job, but you can never offload your learning. The future of the firm is the ability to compound that learning across people and AI. This requires a new architectural approach where every business is able to build agentic systems that improve over time, while still retaining control over their IP. A company should be able to switch out a “generalist” model without losing the “company veteran” expertise built into their learning system. This is the key “test” of your control and sovereignty in the era ahead.
The model labs meanwhile seem to be on a long-term collision course with software. They would certainly like to turn all of the SaaS landscape into a bunch of plugins. The big question is to what degree they can do that. And if they can, will they be the only ones who can do it, or can I do it just as well with a Chinese open source model?
Can that model do the whole thing, or is the model, no matter how good, still a piece of infrastructure that you use to make the actual product? At the extreme, can the model itself invent and make all of those things, and would that let them charge by seat, by outcome or just take the profit? Or do even (or especially) the most sophisticated and high-value use-cases need to sit inside hundreds of new companies that can pick and choose which models to use?
None of these are binaries: they're all a question of degree, and they'll probably vary quite a lot by use case. But at one extreme, there are two or three giant minds that run half of everything and have massive pricing power, and at the other extreme LLMs look like databases - there’ll be millions of them, some very big and some very small, and the value is in what you build on top - after all, every SaaS company is a ’database wrapper’. There’s a future in which Anthropic (or a company we haven’t heard of yet) wins the whole thing and can set its own terms, and a future in which dozens of routers run real-time auctions to allocate your tasks across hundreds of low-margin model-farms and a benchmark company takes a fee on every single one.
Contra Benedict Evans who believes chat is a poor interface that will struggle to capture value, I believe it is a good interface that will struggle to capture value. It’s a good interface because just talking with your computer is intuitive. The problem with it is that it doesn’t show the user in a clear way what can be done with it. There are still ‘incantations’ you have to learn if you want to use chat. But every new model release reduces the need for these. I delete skills regularly. My agent prompts have gotten shorter and shorter. Tons of complicated orchestration has been subsumed by the models simply getting better. The value capture challenge though is that most of the value is provided by the model. The apps are not exactly trivial to make, but they also aren’t super complicated. If models continue to converge in capability, my best guess is that this will go the way of the browser. After Netscape, the browser has captured close to 0% of the value in the web. Browsers are free. But at the same time, everyone uses one. I think it will be the same with AI front-ends.
Right now, ChatGPT is the best app. Claude and Cursor are the other main contenders.
The market for AI inference compute is starting to look more and more like a financial market. In fact, hedge funds and banks have already started getting involved.
This excellent twitter article by Vikram Singh explains:
Intelligence is commoditizing through open-source models. Open source models are almost at parity with the frontier and usage is up 350× since January 2025. Since serving inference for open-source models is permissionless, a competitive market for inference is emerging.
The inference “market structure” is somewhat like a trading venue. There are buyers of tokens (Cursor, Lovable, enterprises), a marketplace (OpenRouter), and sellers of tokens (Fireworks, Baseten, Together AI) quoting a one-sided order book on a model’s tokens.
Inference providers are like market makers: they quote continuously, hold inventory in the form of compute, earn a spread, and vie for token flow. The marketplace is like an exchange: OpenRouter earns a “matching” fee (5.5% take rate) with potentially room to charge the “market makers” for token flow PFOF.
GPU-hours are the homogenous input into producing every token and serve as the layer for transferring risk for inference token production. Compute futures emerge as a mechanism to manage the 130% volatility of GPU-hours and to protect inference provider’s COGS.
This points exactly to one of the futures that Ben Evans predicts: “a future in which dozens of routers run real-time auctions to allocate your tasks across hundreds of low-margin model-farms and a benchmark company takes a fee on every single one.” Stripe, whose entire business is configured around taking a tiny margin on huge volumes, saw this too, and actually acquired OpenRouter in a move that makes so much sense that I think it could prove to be one of the best acquisitions in history.
One thing that I’ve started hearing people say more and more often around me is that Chinese open source models are free. That makes no sense. They aren’t free, they are open source. You can use them, but that means you have to supply the computer to run them. At the bottom-end, this computer can be a very powerful Mac which will run certain small models. But even that isn’t free. If you buy a $5k Mac for AI and depreciate it over 4 years, that’s still $104 per month (excluding energy cost) to run a shitty model. For the same money you can get far better AI from any of a hundred cloud inference providers. The math here is completely the same as it has been for compute and storage forever. Nothing is free. Eventually useable models will be able to run on compute that you already have anyway, like on-demand on your phone and laptop (on the edge). Already happening on paper, but the models are not good enough yet at that scale. And this line of thinking also presumes some kind of convergence in capability. It is self-evident that a model that runs on dedicated hardware in a datacenter is always going to be more capable than one running on a consumer laptop.
The variables are sort of known, but everything else is not and so they are going to move all over the place. Open models have broken through. Everyone knows about them and usage has surged. At this point in time, every sign points toward a future where models become commodities. That means the foundation labs will need to keep moving up the stack where value-capture will continue to happen close to the user / use case. What that app layer is going to look like is anyone’s guess. It depends on so many factors, most importantly how much models will continue to improve in their ability to build those apps. With a looming massive popular backlash against AI, and especially data centers in the US, a lack of chips and a lack of money, supply of compute looks increasingly likely to be a bottleneck, especially if or when the next big use case beyond coding is found.
This post makes a funny and I have to say eye opening case that in every single media object where robots play a role, they are always portrayed as being equal to, or maybe even superior to humans. There is a degree of robot worship going on in the media that is quite ridiculous once you become aware of it.
The column has barely started and already I have to issue a correction. Our pop culture does not, in fact, assert that robots will be equal to humans. It insists that they will be much, much better. Specifically, that they will be better than humans at being human.
Let’s take the 2008 film Wall-E. The plot involves a failed human civilization being saved by a single robot, but it’s not because he’s better at work tasks and doesn’t require vacations. Wall-E shows humanity how to love again, he even reminds them what sex is. You remember that part, right? Humans in that future have chosen screens over fucking, but then a couple observes Wall-E dancing in space with his robot girlfriend? Then the humans hold hands, their fleshy desires newly awakened by the machines? If you don’t remember the film, you probably think I’m joking:
{link to Youtube}
All of the qualities we think of as distinctly human, Wall-E does better. He is braver, more selfless and wise. He has none of our greed, bitterness, jealousy or shortsightedness. Oh, and the film makes it clear he literally has a soul. At the climax, Wall-E selflessly sacrifices himself for humanity in a Christlike manner, his personality getting wiped in the course of his “death.” He is then miraculously resurrected via a mystical lifeforce that enters his robot body upon an encounter with his true love. Again, I AM NOT MAKING THIS UP
Some form of that exact message also turns up in the Spielberg film AI, the Alien franchise, Battlestar Galactica, Bicentennial man, Blade Runner, Ex Machina, Free Guy, Futurama, Humans, I, Robot, the Marvel Cinematic Universe, M3GAN, Short Circuit, multiple Star Treks, the Star Wars franchise, Transformers, Tron, Westworld, The Wild Robot and countless other films, TV shows, comic books, novels, toy lines, video games and songs.
Even weirder, if a franchise begins with portraying robots as inhuman killers (Alien, The Terminator, The Matrix, M3GAN) there will be at least one “more human than human” hero robot in the sequel. Like Wall-E, the machines are often Christlike, sacrificing themselves for a humanity that never deserves it.
Claude Sonnet of course is well known to be a great model at its price point, and Luna is widely regarded to be a very strong model. After OpenAI dropped its price by 80% a few weeks ago, Luna is probably the best value for money model in the world, battling only Deepseek V4 Flash

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.