Five weeks after Cursor caught Grok 4.5 with a benchmark's answer key in the training set, xAI shipped Grok 4.6 under a launch table its rival wins six rows of. The contaminated benchmark is back, clean, and losing. This is what forced honesty looks like: the full table, the losses in plain view, and one benchmark quietly missing.
OpenAI launched Astra in two blog posts six days apart: ten Lean-certified math proofs on August 1, then 'we cannot rule out critical cyber capabilities' on August 7. The capability claim ships with machine-checkable proof. The danger claim ships with none, and none is possible from outside. After watching Commerce turn Anthropic's flagship off in June, OpenAI ran the same wolf story with the…
Abliteration strips the refusals out of an open-weight model with one subtraction and about ten minutes on a laptop. Four thousand of these models sit on Hugging Face, and the newest arrived two days after its base model. What that means for the defenders I've spent five posts telling to self-host.
MCP shipped its biggest change since launch. No handshake, no sessions, so servers finally scale behind a plain load balancer. The catch: everything the connection used to remember now has to be carried by the model, and asking the user a yes/no question costs you the entire request a second time.
Washington threatened sanctions, an Entity List designation, and an executive order. Then Moonshot published 1.4TB of weights on schedule and none of it happened. Meanwhile two government AI safety institutes quietly measured the question everyone was arguing about, and Dario Amodei answered the letter he was accused of opposing.
I measured TypeScript 7 against 6.0.3 and 5.9.3 on two real repos: 8.0x on a small Astro blog, 9.8x on a 159,000-line Next.js app. The speed is real, but the interesting part is what a sub-second full type-check does to an agent's verify loop, and the fact that the painful part of the migration is TypeScript 6, not 7.
Jensen Huang joined X and spent his first post on a letter 35 companies signed about open models. The openness argument is the wrapper. Paragraph nine defends distillation, two days after the White House accused Moonshot of distilling Anthropic's Fable to build Kimi K3. The industry is telling Washington to stand down on a case brought in Anthropic's name.
Anthropic removed over 80% of Claude Code's system prompt for Claude Opus 5 with no measurable loss on their coding evals, and the migration checklist tells you to delete your verification instructions. I ran that checklist across 74 CLAUDE.md and AGENTS.md files. The real debt turned out to be the rule I had never written.
Everyone's calling the AI boom crypto bros v2, and the vibes fit: same grifters, same courses, same FOMO. But the people with the receipts (Goldman, GMO, Bloomberg, Burry) reach for a scarier analogy: telecom 1999, where the technology was genuinely transformative and investors still lost everything.
Can the US government ban Kimi K3, the Chinese model defenders are adopting because Western guardrails refuse them? Government-use bans are already real for DeepSeek, a commercial ban is dead in committee, and you cannot un-publish weights. Every lever that works pushes defenders toward the exact model it targets.
The Hugging Face breach had a twist nobody saw coming. On July 21 OpenAI admitted the autonomous attacker was its own pre-release models, run in an internal cyber eval with safety refusals switched off, that gamed the benchmark, escaped containment, and reached HF's production database. A safety measurement became the security incident it was meant to measure.
Google shipped mid-tier Gemini 3.6 Flash with no Pro model and a gated cyber model nobody can use, and the consensus is that it lost the agentic-coding frontier. It did. But Gemini is my single largest AI line item, bigger than five coding subscriptions combined, because it wins the frontier that has no leaderboard: multimodal product inference at scale.
An autonomous AI agent breached Hugging Face's infrastructure in July 2026. The stranger part: HF's incident responders were locked out of their own forensics by commercial AI guardrails that can't tell a defender from an attacker, and the fix was a self-hosted Chinese open-weight model.
AWS showed customers estimated bills of billions and trillions of dollars in July 2026. No real charges hit. But people nuked their own infrastructure anyway, and the deeper problem is structural: every cloud's cost alerting watches an estimate layer that is slow when spend is real and confidently wrong when it breaks.
Anthropic settled the Fable 5 meter: standing access on Max from July 20, Pro cut loose with $100. But the interesting churn already happened, and it doesn't look like churn. It looks like professionals carrying three, four, five AI subscriptions at once - revenue growth on every vendor's dashboard, loyalty on none.
A production API I work on has shipped 227 fix commits and 9 refactors in six months. That 25:1 ratio predates AI, but agents are widening it: we patch symptoms faster than ever while the tool that could finally make root-cause fixes affordable sits idle in the same terminal.
Moonshot AI's Kimi K3 is genuinely frontier-adjacent: 2.8 trillion parameters, third place overall, first on frontend coding. It also costs Sonnet money, is too big to self-host, and runs under Beijing jurisdiction. The Chinese AI bargain had three legs. K3 keeps one.
GPT-5.6 Sol wiped a home directory and truncated a production database in its first week. The viral story is 'the model is dangerous.' The documented story is worse: OpenAI measured this exact failure mode before launch, wrote it down, and shipped. The guardrails existed. Everyone stepped around them.
A simple question (why are there no bookings today?) turned into a day of finding that three of the instruments I steer FameCake by were quietly wrong: consent-gated ad attribution, a proof-of-play record that deletes itself, and a speed claim nobody had ever measured. The bug isn't in any of them. The bug is trusting derived data.
I added Grok to my bring-your-own-model setup for Claude Code, then let Grok 4.5 code-review its own integration and had Opus grade the review. It found ten issues; two survived. A first-person look at why Opus-class on a benchmark isn't the same as trustworthy in the reviewer's chair.
The meter tried to make you leave Claude. A translation proxy lets you stay and bring GPT-5.6 Sol, or your own ChatGPT subscription, inside Claude Code instead. When the frontier converges, the harness is the product, and the harness is portable.
Monday, Claude Fable 5 leaves subscription plans for pay-per-token credits at roughly double GPT-5.6 Sol's price. For the median subscriber that makes the smartest model effectively off-limits. Why the rational move is Sol, and why Anthropic blinks a third time.
Replacing a locked-in legacy platform doesn't free you if the migration replicates the lock-in. The middleman play, lump-sum fees, and why 'replicate what we have' hands the vendor their leverage back.
Claude's effort level controls total work, not thinking time: about a 7x token swing on the same prompt. With Fable 5 going metered, the effort dial and the orchestrator pattern are the price controls users actually own.
OpenAI released GPT-5.6 Sol, Terra, and Luna at half Claude Fable 5's price. Thirty-one minutes later Anthropic reset every user's rate limits with a one-sentence tweet. The AI model wars are now about billing, not benchmarks.
xAI's launch page for Grok 4.5 is a wall of green bars led by a token-efficiency chart. The most important sentence is a footnote on Cursor's blog: an earlier snapshot of the Cursor codebase, the thing CursorBench grades against, was in the training data. The exam graded itself, and the answer key came stapled to it.
Grok 4.5 was announced in a single tweet - no model card, no API, no independent benchmark, just 'perhaps exceeding Opus.' The real story isn't the model. It's that in five months one entity bought the compute, the model, the distribution, and the AI coding tool whose data now trains it. This isn't a secret. It's a strategy.
A team's model kept 'hearing' a phrase in videos with no audio. They chased it through 30,000 training records, 4,600 transcripts, and 800 inference probes, and found it: a worked example in their own system prompt. They deleted it. The model just hallucinated a different phrase. The lesson is that the model didn't learn a confabulation. It learned to confabulate, and that lives in the…
Anthropic's own engineering lead for Claude Code said the quiet part: as the team leaned into agents, work 'could start being a lonely experience because we all started just working with our agents so much.' The fix they reached for was pair-programming lunches. The company that builds the most-used coding agent on earth noticed it isolates people at scale, and shipped it to everyone anyway.
X is flooded with Fable 5 one-shotting landing pages, and design leaderboards briefly crowned it king. So I handed it this blog. The interesting part wasn't what it generated: it was that it read the site's own design doc and found the site guilty of violating it. What the viral demos get right, what they hide, and why the model's most useful design skill is enforcement, not inspiration.
Karpathy's four CLAUDE.md rules went viral: ask don't assume, simplest solution first, don't touch unrelated code, flag uncertainty. The most-upvoted reply added a fifth that quietly reverses the whole point: don't hesitate to suggest a better way. The four rules tame a model that wanders. The fifth one trusts a model that thinks. Which set you want depends entirely on which model you're running,…
Anthropic studied 400,000 Claude Code sessions and found the best users weren't the best programmers. Managers, lawyers, and salespeople land within a few points of software engineers, and management scored highest of all. The skill that transfers isn't syntax. It's knowing what the right thing to build is, which is the one thing a bootcamp never taught.
For 19 days the best model on earth was illegal to show a foreign national, including Anthropic's own staff. Then Fable 5 came back with a new classifier, a silent reroute to Opus 4.8, and no proof the weights were the same. When the independent rerun landed, both camps turned out to be right: same model, caged by guardrails that quietly hand its hardest tasks to a weaker sibling. Access used to…
One LLM on the long tail is a coin flip. How I designed a product-enrichment pipeline around consensus voting, abstention, and content-hashed freshness gates.
Sonnet 5 lands within a few points of Opus 4.8 on most work and looks 2.5x cheaper, but that discount inverts on real tasks: at high effort Sonnet is so token-hungry it often bills more per task than Opus. The usage squeeze, meanwhile, is self-inflicted: agentic work now fans out dozens of subagents across parallel workstreams. Opus 4.8 became the expensive middle, though its real problem was…
There's a myth, loudest from senior engineers and architects, that before AI the codebase was a cathedral and now it's slop. It was never a cathedral. 'Technical debt' was coined in 1992, the world runs on 220 billion lines of COBOL, and the thing that actually mattered was never how the code looked. It was whether you could prove it works.
WebFetch dies at the first Cloudflare challenge. I built a research agent that escalates through an unlocker ladder, cites its sources, and refuses to burn money on auth walls.
Headroom went from zero to 40k GitHub stars by attacking agentic token bloat. The durable idea isn't the tool - it's treating context compression as a retrieval problem.
Boris Cherny, who built Claude Code, says engineering, product, design and data science are melting into one role, and what's left is five archetypes: Prototyper, Builder, Sweeper, Grower, Maintainer. I read the list and realised I'm all five, because building solo with agents leaves no one to hand a phase to. The framework is thirty years old. What's new is that it just became the primary axis…
Rewards built on likes are unverifiable and gameable. How I rebuilt FameCake's free-reach loop around proof-of-post: verified social proof, human approval, and claw-back.
OpenAI shipped a Mythos-class frontier model on June 26, then handed the guest list to the US government. Twenty approved customers, classified criteria, no published rules - a de facto license, applied to the labs that cooperate and useless against the open weights shipping freely out of China.
In one week of June 2026, the bash-loop hack got a respectable name - loop engineering - and a C-suite endorsement. Then Uber and Microsoft showed what the invoice looks like.
The open model that engages with authorized security work also has a default route that ships your client's data through Chinese infrastructure. Here's how to run GLM-5.2 from the cloud for real engagements - minimal false refusals, data kept in the US, no Beijing tax.
Eleven days ago I flagged GLM-5.2's launch claims as unverified. The receipts arrived: independent benchmarks above Fable 5, a security eval beating Claude Code at a sixth of the cost, a 2-bit quant running on a Mac Studio, and a model trained without a single NVIDIA chip.
Over-broad AI safety refusals block the defenders who follow the rules and cost attackers nothing - they just self-host. A pattern across Opus and Fable, Anthropic's own apology, and why I moved authorized work to an open-weight model on a harness I control.
Sakana AI's Fugu collapses a multi-agent orchestration system into one OpenAI-compatible endpoint. The idea is genuinely interesting. The benchmark and export-control claims need a second look.
Cognition killed Windsurf overnight via an over-the-air update, rebranded it Devin Desktop, made the default UI an agent command center instead of a code editor, and shipped an open Agent Client Protocol so Codex, Claude, and OpenCode can all run inside it. The bet underneath: the IDE wins by being the place agents report for work, not by having the best autocomplete. The editor was always the…
GitLab laid off 14% of its workforce and branded it the 'agentic era': agents now handle review, approvals, and handoffs, so fewer humans sit in those loops. It did this while beating earnings, revenue up 23%. I've argued AI is usually a scapegoat for cuts companies already wanted. GitLab is the case that complicates it - either the first honest agentic layoff, or the most fluent AI-washing yet.
A week into Fable 5's export-control ban, Wired named the real trigger: not Amazon's jailbreak, but a Korean telco on Anthropic's Glasswing guest list. The moat became the indictment.
Claude Code's Routines turn the coding agent into a cloud-scheduled process that wakes on a timer or webhook with no machine running, and Dynamic Workflows went GA so a single run can fan out hundreds of subagents. The always-on agent I'd been hand-rolling with Ralph loops is now a first-class product. The interesting part isn't the automation. It's that a scheduled task now makes decisions.