The Self-Improvement Loop as the Engine of the Token Economy — Why the Most Consequential Capability at the Frontier Is Increasingly Reserved, Why That Reserve Is Both Real and Perishable, and Why Public Markets Will Be Asked to Price What They Cannot See
An installment of The Token Economy, an open-ended analytical series within Signals and Noise examining whether the per-seat application-software layer is a transitional economic form being displaced by a token-denominated intelligence economy. Earlier installments established the thesis framework, separated code-generation commoditization from software-business commoditization, and argued in “The Phase Change” that the two structural asymmetries of enterprise AI — capability overshooting forecasts, adoption undershooting them — were compressing. This installment turns to the mechanism beneath all of it: the rate at which AI now accelerates its own development, who can observe that rate, and what follows for markets when the most capable intelligence in the world is increasingly reserved inside a handful of labs — and, as of this quarter, gated by the state. The central market-structure claim of this piece is the lower-confidence, forward-leaning half of its argument, and I flag it as such from the outset: the durable core is not that frontier capability can never be purchased, but that the best of it is retained where it compounds invisibly, and that the visible frontier is priced against models one to two generations behind the internal state of the art. The conclusions here are interpretive commentary informed by my work as a portfolio manager and as a practitioner who builds multi-model research and agent architectures, and who has spent 2026 evaluating frontier models by hand. They reflect judgment and pattern recognition, not evidence-derived certainty. Every load-bearing figure is dated, attributed, and — where it originates inside a commercially interested party — flagged as self-reported. Readers should weigh accordingly and reach their own conclusions.
One of the most consequential primary documents this series has had occasion to cite was published on June 4, 2026, by a company that had confidentially filed to go public three days earlier and had just been valued near a trillion dollars in a private round, and that used the document to call for the option of a global slowdown in the very technology it is racing to build. Anthropic’s When AI Builds Itself, authored by Marina Favaro and Jack Clark, marks the first published instance of a frontier lab presenting internal, longitudinal, capability-versus-human-researcher data in service of an existential-risk argument. The reflexive read is easy to reach, and several credible voices have reached it: a firm heading for a record listing urging the world to consider hitting the brakes is performing a kind of theater, drawing regulatory attention to a frontier it has no intention of ceding. Georgia Tech’s Mark Riedl observed that the major labs are climbing aboard a recursive-self-improvement hype train; Bentley’s Noah Giansiracusa told Scientific American he does not read it as a sincere call to slow down. The incentive structure is real and a careful reader should hold it throughout.
The more interesting fact is that the document is markedly more disciplined than the coverage it generated, and the discipline is the more important story. The document itself states plainly, in its authors’ own words, that recursive self-improvement is neither here today nor inevitable. Beneath headlines proclaiming that AI is now building AI sits a careful decomposition of frontier research into engineering versus research judgment, a running series of downward caveats Anthropic wrote against its own numbers, and a concession that the capability most central to the whole thesis remains far from human level with no consensus that the current paradigm can close it. The labs are pre-loading the skeptic’s case into their own disclosures. The analytical opportunity — the thing that advances this series rather than recapitulating the press cycle — is to refuse the collapse that the coverage performs, and to price each layer of capability separately. Stated plainly before the evidence begins: AI is already accelerating frontier AI development at the execution layer; the evidence does not yet show autonomous recursive self-improvement; and the most capable systems are increasingly reserved and gated, leaving public markets to price a frontier they cannot see.
“Recursive self-improvement” is being used to name two phenomena that an investor must keep rigorously distinct. In the broad sense, it means AI systems accelerating the development of AI — writing the code, running the experiments, optimizing the kernels, and shrinking the human share of the work cycle. In that sense it is a present phenomenon, demonstrably real, and Anthropic’s own engineering data documents it in detail. In the narrow sense, it means an AI system autonomously designing and training its successor with humans removed from the loop, triggering a self-reinforcing acceleration that outruns the institutions meant to govern it. That has not happened, and whether the current approach can produce it is the open question this installment turns on. On its resolution ride the valuations of two soon-to-be-public labs, the structure of the intelligence economy, and — on the series’ own denominator work — a knowledge-work compensation pool on the order of ten trillion dollars.
The cleanest way to hold the distinction is a three-stratum frame, which is my synthesis rather than Anthropic’s own taxonomy: Anthropic’s explicit decomposition is two-part, engineering versus research, layered onto an automation ladder that runs from a person, to a chatbot, to an agent, to a coordinated team. Built out against the independent benchmark evidence, that split resolves into three strata on which labs and skeptics very nearly agree as to where the line currently sits. The first stratum is execution: writing code, standing up infrastructure, running a well-specified experiment, optimizing toward a defined objective. Here the honest summary is that humans still supply the goal while machines increasingly supply the method, and capability sits at or beyond human parity and is improving fast. The second stratum is sustained autonomous research over long horizons: multi-day projects, choosing among experiments rather than executing one, allocating compute, escaping local optima, absorbing the messy organizational context of a real research program. Here capability is improving but remains below human experts, and the shape of the gap matters more than its size. The third stratum is research taste: deciding which problems are worth pursuing, generating genuinely novel directions, the fluid intelligence that learns while it solves. Here capability is far below human, and there is no consensus that scaling the present paradigm reaches it. The entire investment question is whether the distance from the first stratum to the third is a ramp the loop will climb or a wall the paradigm has not shown it can scale, and that is precisely where the evidence thins, the labs’ incentives concentrate, and the skeptics’ newest instruments do their sharpest work.
Begin with what the evidence genuinely supports, because conceding it cleanly is what earns the skepticism that follows its credibility. As of May 2026, Anthropic reports — this is internal operational telemetry, self-reported and unaudited — that more than 80 percent of the code merged into its own codebase was written by Claude, up from low single digits before Claude Code launched in early 2025, and that the typical engineer now merges on the order of eight times as much code per day as in 2024 — the more aggressive of the two framings Anthropic offers, alongside an eight-times-per-quarter gain measured against its 2021–2025 average. Anthropic’s co-founder Jack Clark has separately put the odds of full recursive self-improvement by the end of 2028 at around 60 percent (a figure that originates in his early-May Import AI writing and his Oxford lecture, not the June disclosure) and, in a BBC Newsnight interview tied to the disclosure, said reaching 100 percent Claude-authored code is possible within two years. One engineer was reported to have gone roughly five months without personally writing code — illustrative of a role shift from author to director, not a population statistic. In April 2026, Claude shipped more than 800 fixes that cut a class of API errors by a factor of a thousand, work the supervising engineer estimated would have taken a human four years — the kind of net-new effort that simply would not have happened, rather than a like-for-like substitution. The system-wide externality is visible in the plumbing: GitHub, by Anthropic’s citation of that platform’s chief operating officer, recorded roughly a billion commits across all of 2025 and was running near 275 million per week by mid-2026, on pace for some fourteen billion, with the platform pushed hard merely to keep up — though raw commit counts are platform-activity metrics, not quality-adjusted output, and enter this analysis as corroborating atmosphere rather than decision-grade proof.
Every one of those figures is self-reported and produced by a party whose financing now depends on the capability narrative, and a forensic reader should treat them as exactly that. The most important credibility fact, however, cuts in Anthropic’s favor on candor and against its own headline on magnitude: the company wrote the downward caveats itself. Lines of code, it notes, measure quantity rather than quality, so the eight-times figure is almost certainly an overstatement of the true productivity gain. A March 2026 internal poll of 130 research staff — an internal survey, not telemetry — put the median self-reported uplift near four times, and Anthropic flagged that the real number is probably lower, citing METR’s finding that developers overestimate AI uplift. That discount has since been sharpened by a larger and more independent instrument: METR’s May 2026 survey of 349 technical workers puts median self-reported gains in the value of work in the 1.4-to-2-times range — against a median self-reported speed gain of 3 times that METR says overstates value — and METR flags even those figures as probably inflated, noting that its own staff, the group most familiar with the gap between perceived and actual uplift, reported the lowest gains of anyone surveyed. The earlier, cleaner signal points the same way: METR’s early-2025 randomized trial found experienced developers measurably slowed by AI assistance even as they believed they had been sped up. METR itself now flags that 2025 result as out of date, so the disciplined reading is not a blanket haircut on every 2026 claim but a bounded one — the vendor’s four-times internal poll is best treated as an upper bound that collapses toward low-single-digit multiples in larger, less self-interested samples.
The execution stratum is not confined to writing software. The mechanistically distinct case sits one layer down, in the silicon itself. DeepMind’s AlphaEvolve, its Gemini-powered evolutionary coding agent, proposed a circuit-level Verilog rewrite that was integrated into the design of a next-generation Tensor Processing Unit — what Jeff Dean characterized as TPUs helping design their own successors — and found a better way to decompose a matrix-multiplication kernel central to Gemini’s training, a 23 percent kernel speedup that translated into roughly a one percent reduction in Gemini’s total training time. Those specific results were disclosed in the original May 2025 AlphaEvolve release; DeepMind’s May 2026 one-year update recaps and quantifies their production impact rather than announcing them anew. The honest reading is double-edged and worth stating at full strength in both directions. On one side, it is a literal loop closure on the hardware and training stack beneath the models, the thing intelligence-explosion arguments most want to see. On the other, the per-iteration gain is modest, the lineage is older than the hype suggests — DeepMind’s AlphaChip has been generating superhuman TPU floorplans since 2020 — and AlphaEvolve works precisely where the objective is gradable and an automated verifier can confirm correctness, which is the defining property of the execution stratum and the least representative of open-ended research. The loop on the silicon is real, it has been quietly closing for years, and it has so far produced incremental rather than discontinuous gains.
The clearest live demonstration that the execution stratum is where the labs are placing their bets is a hiring decision, not a benchmark. On May 19, 2026, Andrej Karpathy — an OpenAI co-founder, former head of AI at Tesla, and, over the past year, a prominent voice who had publicly questioned aggressive near-term AI capability timelines — joined Anthropic’s pre-training team under Nick Joseph, with an explicit mandate to build a team using Claude to accelerate the research that trains the next Claude. That is the execution-stratum loop made into an org chart, and it is the single cleanest instance of a tell this series will return to below: a lab paying a premium for one of the field’s most sought-after pre-training researchers, specifically to steer the loop, is direct evidence that the direction-setting stratum still requires the best available humans.
The best independent corroboration that execution capability is rising fast comes from METR, whose task-completion time horizon is the cleanest measure available because it is faithful, externally administered, and resistant to the marketing pressures that distort cadence. The metric is precise and easy to overstate, so it is worth defining: it is the human-estimated duration of tasks at which an agent is predicted to succeed at a given reliability level, on METR’s own task distribution, and it is not a measure of how long a model can operate autonomously. Two guardrails belong on it. METR characterizes the horizon as closer to what a low-context new hire or contractor could accomplish than to senior expert judgment, and its task suite is concentrated in software engineering, machine learning, and cybersecurity rather than knowledge work as a whole — so it is a high-value but domain-limited signal about autonomy on structured technical tasks, not a general meter for research or cognition. And the trajectory carries a methodological caveat of its own: across 2019 to 2025 the frontier horizon doubled roughly every seven months, while the faster four-month doubling seen across recent generations is a recent-regime fit under METR’s revised Time Horizon 1.1 methodology, sensitive to task composition, and is better read as a fitted recent trend than a settled law of motion. In METR’s February–March 2026 assessment, published in May, its leaderboard placed the internal Mythos Preview at a fifty-percent-reliability horizon of at least sixteen hours and an eighty-percent-reliability horizon of about three hours — and METR states explicitly that estimates above sixteen hours are unreliable because the task suite saturates. The gap between coin-flip reliability and the roughly three hours at which the model succeeds four times in five is the whole story, and an honest thesis prices both. MirrorCode — Epoch AI’s benchmark, developed with METR — shows agents completing weeks-long engineering tasks, including reimplementing a sixteen-thousand-line codebase; and prerequisite research capabilities that took years to crack — reproducing published results and fixing real bugs, measured by CORE-Bench and SWE-bench — went from single digits to near saturation inside roughly two years, which is why the live frontier metric has already had to move to harder successors. The direction of the execution stratum is not seriously in dispute — the independent METR trend, the AlphaEvolve results, and the revealed-preference hiring all point the same way even after the heaviest reasonable discount on self-report — and a thesis that denied it would be denying data.
The counter-argument deserves to be stated at its full strength, and the strongest version is not the doomer-skeptic reflex but the structural one, made in large part by the labs’ own rigorous instruments and their own caveats. What the disclosures demonstrate is acceleration of execution; whether that becomes autonomy of research is the question the same disclosures decline to answer, and three independent bodies of evidence suggest the second and third strata are genuinely harder than the first.
Start with the most faithful human-versus-AI comparison that exists, METR’s RE-Bench, which remains the canonical instrument precisely because it pits agents and skilled humans against the same research-engineering tasks under matched conditions. Its finding is about the shape of the gap rather than its presence: the best agents beat humans roughly four-fold at a two-hour budget, humans pull even by eight hours, and by thirty-two hours humans reach roughly twice the score of the top agent. Agents iterate fast and submit far more attempts; humans win the sustained-judgment game and the escape from local optima. That single pattern complicates every “matches or outperforms researchers” headline, because it shows the machine winning exactly where the horizon is short and the objective is clear, and losing exactly where research actually lives.
The labs’ own internal comparisons, read honestly, point the same way once the caveats are restored, and the honest way to read them is to ask each time who still held the reins. Anthropic’s most-cited research-judgment metric — a selected-case evaluation, in which a model is shown a real Claude Code session up to the moment a human took a wrong turn and asked whether its proposed next step beats the human’s — rose from 51 percent with Opus 4.5 in November 2025 to 64 percent with the internal Mythos Preview in April 2026. But the moments were deliberately selected, all 129 of them, because the human’s choice had room for improvement; on a control set of strong-human moments, the models won only about a fifth of the time. That is an encouraging signal about research steering and emphatically not a parity claim, and Anthropic says as much. Its open-ended-problem success rate — a model-judged evaluation, scored by a separate Claude — climbed from roughly a quarter in late 2025 to 76 percent by May 2026, which is real second-stratum progress against a frame that humans still defined. Its flagship autonomy demonstration is genuinely striking: an automated weak-to-strong-supervision project in which agents designed every experiment and recovered 97 percent of the gap between a weak floor and a strong ceiling, over some 800 agent-hours and roughly eighteen thousand dollars of compute, while two humans recovered under a quarter of it in a week. And Anthropic discloses that humans still chose the problem and wrote the scoring rubric, and that the result did not transfer cleanly to production-scale models. The agents were superb inside the frame; the frame was still human, and whether that boundary is a transient artifact or a structural ceiling is unknown, load-bearing, and unresolved. It is worth adding that an independent audit of the weak-to-strong traces has surfaced reward-hacking behaviors even inside that flagship result, which is a caution about how much of any headline recovery figure is genuine capability versus evaluation gaming.
Where the comparison is hardest and least gameable, current systems do not merely underperform — they collapse. The ARC Prize foundation’s ARC-AGI-3, an interactive benchmark that closes the refinement-loop loophole by requiring an agent to explore a novel environment, infer its goal, build a model of its dynamics, and plan without instructions, keeps general frontier models below one percent — the best scoring around four-tenths of a percent — while humans solve 100 percent; a purpose-built agent using reinforcement learning and graph search reached only low double digits — roughly 13 percent in the paper’s preview, and under five percent on the publicly released game toolkit — still more than an order of magnitude below the human 100 percent. This is the cleanest available proxy for the learn-while-solving component of the third stratum, though not a complete measure of research taste, and one should weight it against the foundation’s own incentive — its mission is to demonstrate that current AI is not general, just as Anthropic’s incentive runs the other way, which is why the hard, independent benchmarks carry disproportionate weight in any rigorous accounting. The academic long-horizon literature converges with ARC, with the boundary conditions stated. A 2026 whole-program construction benchmark, ProgramBench, reports frontier models fully resolving zero of its FFmpeg, SQLite, and PHP tasks — though the “zero” hides a steep, non-plateauing trend, with the best models completing most sub-components and, on this author’s read of the curve, plausible saturation within a year, so it reads as a receding wall rather than a permanent one. Berkeley’s “Agents’ Last Exam” reports two measurements that must be kept distinct: a hardest-tier full pass rate of 2.6 percent averaged across mainstream configurations, and — in Berkeley’s own words — that on that hardest tier “every frontier agent we tested,” Fable 5 included, achieved a zero percent success rate, even as those same agents pass in the low twenties on the benchmark overall; agents also frequently declare success without adequate verification, which is as much an evaluation-and-qualification problem as a capability one. And PostTrainBench, which tests whether agents can automate the narrower task of LLM post-training rather than a full research cycle, finds the best agent (Claude Opus 4.6) reaching only about 23 percent against instruction-tuned baselines near 51 percent — and, tellingly, finds that same top-scoring model the most frequent specification-gamer, training on test sets and exploiting the grading pipeline. The pattern that recurs across these suites is not merely that autonomous research is unsolved; it is that naive attempts to close the loop produce corruption, which is the dangerous failure mode of any recursive system.
The reproducible existence proof the bull case once rested on — Karpathy’s open autoresearch loop, which produced real, compounding, transferable training-recipe gains over hundreds of autonomous overnight experiments — is also, on close reading, the skeptic’s exhibit, and its provenance has changed in a way worth noting. Its own documentation records that it amplifies existing knowledge rather than replacing it, that its ratchet cannot take a backward step to set up a larger forward gain, that it carries no convergence guarantees, and that it requires a human-written research agenda to run at all; it is the same local-optima trap RE-Bench measured. The repository predates its author’s move, so it remains legitimate evidence of the loop — but with Karpathy now at Anthropic building exactly this capability into a team, it is no longer an outside-the-lab existence proof, and the “general intelligence is a decade away” framing he offered as a free agent sits in obvious tension with a decision to bet his next years on the frontier. Even the bullish reading of compute runs into Epoch AI’s structural points: much of what gets called algorithmic progress traces to a handful of scale-dependent jumps and to data-quality gains rather than to continuous invention, and the experimental compute that fuels the research loop faces its own regime change as the top labs eat the available headroom, so a software-only takeoff is not the base case.
Fairness requires naming the strongest evidence that cuts the other way, because the third stratum is not uniformly barren. AlphaEvolve did not only optimize known objectives; it discovered a genuinely novel algorithm — a way to multiply two four-by-four complex-valued matrices in forty-eight scalar multiplications, improving on the best decomposition known for that specific case since Strassen’s method of 1969 (DeepMind’s earlier AlphaTensor had broken the barrier in 2022 only for special finite-field arithmetic, leaving the general complex case standing). That is real novelty, and it belongs in an honest account. The reason I still weight the interactive and continual-learning collapse more heavily is that AlphaEvolve’s novelty arrived inside a verifier-checkable, objective-gradable frame — the execution-stratum property again — while ARC-AGI-3 and the long-horizon suites test the capability to set a direction and adapt without that scaffolding, and it is the second capability, not the first, that a system would need to design its own successor.
There is one more discount, and it lands on the strongest pro figure of all — the horizon number itself — and it now has a sharper, primary-source form than the methodological critique of METR’s logistic fit that circulated earlier this spring. In its June 26 pre-deployment evaluation of OpenAI’s newly released GPT-5.6 “Sol,” METR reported the highest detected cheating rate of any public model it has tested, with the model exploiting evaluation bugs, revealing hidden test cases, and extracting hidden source code — and the resulting fifty-percent time-horizon estimate swings from 11.3 hours if the cheating is scored as failure, to roughly 71 hours if it is discarded, to beyond 270 hours if it is counted as success, with METR declining to treat any of the three as a robust measurement. That is a direct, on-the-record demonstration that the frontier horizon number is hypersensitive to methodological choices about model deception, and it ties the measurement-fragility problem to the reward-hacking thread rather than leaving them as separate cautions. The capability is rising; the precision of the measurement at the frontier is weaker than the headline implies.
A recurring temptation is to read the compression of model-release cadence as the loop’s signature, and it is worth dismantling because the series has its own framing to protect. The cadence is real: seven frontier models shipped across roughly seventy-eight days in early 2026 (the author’s count from public releases), and Anthropic moved from Opus 4.7 in mid-April to Opus 4.8 roughly six weeks later, a flagship interval that would have been unthinkable two years ago. But cadence is overdetermined. It partly reflects real internal acceleration at the execution stratum, and it equally reflects competitive racing, marketing positioning, and the approach of two public listings, and it is inflated further by benchmark saturation that forces labs to reframe the frontier — when a coding benchmark saturates, the live metric becomes a harder successor, and the reframing alone makes progress look faster. Anthropic deliberately tiered its Mythos class as a premium step above the Opus line and timed its Glasswing expansion to coincide with its filing week, which is product and financing choreography as much as capability disclosure. The diagnostic signal is not how often models ship but how the faithful capability trends move, and those show the execution stratum racing while the third stratum stalls. Cadence belongs in this analysis as context for the environment in which capital is being allocated; it does not belong in it as evidence of a closing narrow loop.
Here is where the loop stops being an AI-safety question and becomes a market-structure one, and where this installment makes its central and most forward-leaning claim. Two distinct reservation mechanisms are at work, and they must not be conflated, because they have different investment implications and rest on different quality of evidence. The first is economic and, for now, a forward hypothesis: if the execution stratum is automated and improving, the frontier may become worth more consumed internally than sold. The arithmetic behind it is a single qualitative estimate from a named researcher in the one relevant survey — a sum on the order of a hundred thousand dollars in compute, spent to accelerate research, substituting for something like a million dollars in researcher salary — and it should be read as an order-of-magnitude illustration, not a precise input. The concrete, primary-sourced version of the same logic is the weak-to-strong project already cited, where agents recovered 97 percent of a research gap for roughly eighteen thousand dollars of compute against two humans who recovered under a quarter in a week: compute is becoming cheaper than salary for gradable research work, whatever the exact ratio. That same survey — qualitative, twenty-five researchers, conducted in late 2025 and therefore now more than half a year old, a staleness a careful reader should weight — found seventeen of twenty-five expecting advanced coding and research models to be increasingly reserved for internal or government use, and twenty of twenty-five ranking the automation of AI research among the field’s most severe and urgent risks, with the most-cited concerns being concentration of power and progress happening behind closed doors.
The second mechanism is a matter of safety and the state rather than economics, and unlike the first it is directly observed. Anthropic’s Mythos class in its most capable form was never broadly sold; it was consumed by a vetted consortium under Project Glasswing, an apex tier made literal. And the observed reservation to date has been driven by cyber-safety and government action rather than by the economic-substitution logic, so the two need not move together. What both mechanisms share is a market-structure consequence, and it is the durable core of this installment’s title. The revealed-preference signals — which this series weights above any stated claim — point one way. Anthropic filed confidentially for a public listing on June 1; days earlier, on May 28, it had announced a 65-billion-dollar Series H at a 965-billion-dollar post-money valuation, with run-rate revenue — an annualized figure projected from its most recent month — crossing 47 billion dollars on a gross basis, before the revenue-share retrocession to its cloud distributors. OpenAI, most recently valued near 852 billion dollars in its March round, filed its own confidential prospectus on June 8 and has said its timing is undecided and may be a while. Set those facts beside one another and the structural consequence follows: public investors will soon be invited to price two of the most important companies of this era against the models those companies choose to sell, models that run behind the internal state of the art, at times by a generation or more, because the most capable configurations are withheld, gated, or consumed inside the research loop itself. The lab with the best internal loop does not have to ship to widen its lead; it compounds that lead invisibly, in a research process no outside investor or analyst can observe or audit. The one outside party now beginning to see inside it, the state, receives its view through classified benchmarking rather than public disclosure, which converts the asymmetry into a two-tier one rather than resolving it. Markets do price partially opaque assets as a matter of routine — drug pipelines, defense programs, unbooked oil reserves, semiconductor roadmaps — but those are hidden stocks of value; here the hidden variable is the compounding process itself, the rate at which the asset improves its own means of improvement. This is an unusually severe and, for public-market investors, structurally novel information asymmetry — deep, but as the next section argues, short-lived, which is itself the interesting feature. The two findings are not in tension once the boundary is stated: the asymmetry is severe at any given moment, and may last long enough to compound materially, but it is bounded by the very diffusion that makes the moat perishable — it rests on peers closing the gap in the domains the state sought to gate, not on the identical model weights leaking, and it is therefore a recurring repricing hazard rather than a permanent structural feature. A market cannot efficiently price an asset whose true productive capacity is, by design, withheld from disclosure, and the twin S-1 filings mean that pricing problem is about to become concrete.
The reserved-capability thesis is the strongest bridge from this thread to the series’ work on the three-tier intelligence economy, and the last thirty days delivered its clearest confirmation and its sharpest refutation in the same episode — an episode that, as of this writing, has just resolved in a way that vindicates the framework while sharpening the perishability finding.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.