RSS Amplifier

Windfall Trust · Apr 13, 2026

Brief #9: MIT researchers validate fast AI capabilities

0
Sign in to vote or save

Jacob Schaal, Joel Christoph · Windfall Trust

Welcome! This bi-weekly newsletter, published by the Windfall Trust, curates the most important developments in AI economics research and policy. Each issue features key research and updates, along with in-depth analysis and quick links to relevant opportunities and recent news.

  • Two major measurement efforts now suggest that AI capabilities across economically relevant tasks are improving very quickly. MIT FutureTech’s new worker-evaluation project finds that the length of tasks AI can complete successfully — measured by how long those tasks take humans — is roughly doubling every 3.8 months across a broad set of text-based labor-market tasks. METR’s updated software-task estimates point in a similar direction, with its revised post-2023 doubling time at about 131 days, around 20% faster than in its earlier setup.

  • MIT’s result is especially notable because it is based on more than 17,000 worker evaluations across more than 3,000 broad-based O*NET-derived tasks. The authors estimate that AI systems could complete most text-related tasks at 80% to 95% success by 2029, at a minimally sufficient quality threshold, if recent trends continue.

  • However, a new Forecasting Research Institute survey finds that even economists who expect substantial AI progress by 2030 still forecast that the U.S. GDP, TFP, and labor-force participation will remain relatively close to historical trends in their baseline view.

  • OpenAI published Industrial Policy for the Intelligence Age, a new policy blueprint that frames advanced AI as a reason to revisit tax policy, worker protections, public wealth, and industrial strategy. OpenAI also says it is launching pilot fellowships and focused research grants of up to $100,000 and $1 million in API credits, tied to these policy ideas.

  • Deric Cheng and Frank Ryan published the Windfall Trust Policy Atlas, a new resource that maps policy proposals to address AI-driven economic disruption across labor markets, wealth distribution, public investment, market design, and global coordination.

    • Disclaimer: Deric is an editor of the AI Economics Brief.

  • Matt Sheehan (Carnegie) reports that Chinese courts have ruled that firing workers because AI can do their job violates the Labor Contract Law. Whether such legal frictions can shield workers as AI capabilities continue to advance remains to be seen.

Thanks for reading Windfall Trust! This post is public so feel free to share it.

Share

A new MIT FutureTech paper, Crashing Waves vs. Rising Tides, offers one of the broadest attempts yet to measure AI capabilities across real labor-market tasks. Instead of focusing on a narrow benchmark suite, the team evaluates more than 3,000 O*NET-derived, text-based tasks using over 17,000 worker evaluations.

Their headline result is striking. In 2024 Q2, frontier models could successfully complete tasks that take humans roughly 3 to 4 hours about half the time. By 2025 Q3, that had risen to about 65%. Using a linear trend model, the authors estimate an implied doubling time of 3.8 months for a feasible task duration. If that pace continued, they project that most of the text-related tasks in their sample would reach average success rates of 80% to 95% by 2029 at a minimally sufficient quality threshold.

They find that AI improves like a rising tide broadly across tasks of many lengths at once. This matters because it runs counter to the view that AI progress occurs mainly in isolated domains, such as coding, which are hit by crashing waves of AI progress.

Our analysis: It’s important to keep in mind that this is an important update on capabilities, but not yet on jobs. The paper studies self-contained, text-based tasks in which the relevant context is already supplied. That makes it highly informative about what models can do under favorable conditions, but less informative about organizational adoption, workflow redesign, liability, verification costs, and messy real-world execution.

MIT’s results look even more significant when set beside METR’s latest time-horizon work. METR’s Time Horizons page and its Time Horizon 1.1 update continue to show exponential growth in the length of software tasks that frontier models can complete autonomously. When METR updated its task suite and methodology in January 2026, its post-2023 doubling time fell from about 5.5 months to 4.3 months — a 20% acceleration. Looking only at models released since 2024, progress appears even faster, with a doubling time under 3 months.

METR has also estimated time horizons for math, science, and visual computer use, finding doubling times of 2-6 months across all domains, though with substantially different current capability levels. So AI capability levels vary across domains, but improvement rates are broadly uniform.

Our analysis: The convergence is striking given how different the methodologies are: MIT uses broad labor-market tasks evaluated by workers; METR uses verifiable software and research tasks scored automatically. Yet MIT’s logistic slopes are substantially flatter than METR’s, suggesting that on realistic work tasks, AI does not degrade as sharply with task length as narrow benchmarks imply. One important caveat applies to both: all measured tasks are self-contained with full context provided, which means these results describe a ceiling on capability, not a forecast of deployment.

A new Forecasting Research Institute survey, Forecasting the Economic Effects of AI, surveyed 69 leading economists, 52 AI industry and policy experts, 38 superforecasters, and 401 members of the general public between October 2025 and February 2026. Economists assign a 61% combined probability to either a moderate or rapid AI progress scenario by 2030, and a 14% probability to the rapid scenario alone. They are not especially conservative on capabilities.

Yet these same economists do not expect a dramatic break in the baseline macro picture. In the unconditional forecast, GDP growth, TFP, and labor-force participation remain close to historical trends. Even in the rapid-progress scenario, economists’ median forecasts are surprisingly moderate: GDP growth is roughly a percentage point above today’s baseline, comparable to the 1990s. They expect meaningful but not dramatic declines in labor-force participation and a rise in wealth concentration to levels not seen since 1939. Uncertainty is enormous, with GDP growth estimates spanning 1% to 7%. This fat-tailed uncertainty may be the most policy-relevant finding in the report.

The gap has already triggered debate. As Tom Cunningham (METR) noted, several coauthors found the forecasts too conservative. Basil Halperin (Virginia) described the median predictions as implausible given the scenario premises. Respondents appear willing to grant the possibility of very powerful AI systems, while remaining much more skeptical that those capabilities will translate quickly into economy-wide disruption.

Our analysis: This survey reveals that economists believe there’s a substantial gap between AI capabilities and macroeconomic impact. This likely reflects beliefs about diffusion lags, persistent human bottlenecks, infrastructure constraints, political reaction, and adoption frictions. As the Jones and Tonetti “weak links” framework from Brief #6 suggests, even enormous capability gains can be bottlenecked by tasks that remain human-performed. The FRI plans to run the survey annually, and the gaps between capability expectations and economic forecasts will warrant close monitoring.

Labor Market & Employment

  • A new Federal Reserve paper finds that coder employment growth has been about 3% lower since ChatGPT’s introduction after controlling for industry-level shocks.

  • The FT covered recent Brookings research on how AI endangers gateway roles that allow non-college graduates to transition into white-collar positions, such as administrative assistants and bookkeepers.

  • Tarek Hassan (Boston) and coauthors revisit the idea that educated workers may hold a comparative advantage in adopting fast-moving technologies, with implications for the college premium in an AI-intensive economy.

  • Goldman Sachs estimates that AI reduces monthly payroll growth by roughly 25,000 jobs while AI augmentation adds about 9,000, for a net reduction of 16,000 jobs per month.

AI Capabilities & Infrastructure

  • Hyunjin Kim (INSEAD) and coauthors run a field experiment in which hearing case studies of AI-driven reorganization increased revenue by 90%.

  • Tom Cunningham (METR) and coauthors formalize the “apple-picking” model of AI R&D, showing how agents can make autonomous research contributions without being full human substitutes.

Market Power & Risk

  • Luis Garicano (LSE) worries that Europe will face higher interest rates driven by AI-driven capital demand, even as growth slows.

Research Opportunities, Events, and Jobs

Thanks to Deric Cheng for contributing to the creation of this week’s edition of the newsletter.

Read the original on windfalltrust.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.