Everyone says the latest AI agents will be “job-ready” soon, especially after the release of Fable 5 last week. But is that really the case? Over the past many months, our team at Berkeley RDI, in collaboration with 300+ global experts across 55 industries, have been building Agents’ Last Exam (ALE), a benchmark designed to test exactly that claim on real digital labor-market work. And we found: the age of useful agents is here. The age of truly job-ready agents is not.
TL;DR
Agents’ Last Exam (ALE): a rolling benchmark that measures whether AI agents can actually perform economically valuable work across a broad range of real-world domains. With ALE, we evaluated Fable 5, GPT-5.5, Composer 2.5, and other frontier agent systems across more than 1,500 expert-sourced tasks spanning 55 occupations. Today’s agents can solve a meaningful fraction of professional tasks, but when we look at the hardest tasks that require sustained reasoning, deep domain expertise, and reliable execution over long horizons, they are still far from human-level performance. On ALE’s hardest tier, most frontier agents we tested, including Fable 5, achieved a 0% success rate.
Why This Is New / Important
Many of today’s agent benchmarks are rapidly saturating as frontier systems improve. ALE is designed to measure a different capability frontier focused on sustained, economically valuable work, featuring 1,500+ expert-sourced tasks across 55 industry domains with full GUI + CLI environments and outcome-based, verifiable evaluation.
Compared to Terminal-Bench and SWE-bench-Pro, it’s broader (tasks span 40 of ALE’s 55 industry subdomains, vs. 6 and 5), longer-horizon (human time-to-complete ranges from hours to weeks, not minutes or days), and harder (the best agent passes just 25.2%, vs. 82.0% on Terminal-Bench and 59.1% on SWE-bench-Pro).
How It Works
Each ALE task hands an agent a real project that a human expert previously completed, converted into an evaluation with objective, verifiable grading. The agent gets full GUI and CLI access and is free to solve it however a human would: clicking, typing, scripting, browsing. It is judged on the result, not the method. Grading is deterministic and rubric-based, with no human judges: each task’s expert author defines a reference output and rubric.
Implications
The name “Last Exam” has two meanings. First, “Last” is the bar to clear: passing these exams means an agent can actually do the job and continue to deliver economically valuable work in that profession. Second, “Last” is the frontier of difficulty. The tasks are real, complex, long-horizon, and require professional expertise to execute, meaning ALE sits right at the edge of what today’s agents can reliably accomplish.
To advance this frontier, we welcome contributors to help build our next-iteration benchmark by contributing tasks and referring domain experts (contributors will be invited to join as co-authors). Learn how to contribute at https://agents-last-exam.org/submit, and explore the leaderboard, paper, and demo at https://agents-last-exam.org.
Learn More About Agents' Last Exam
For more information about ALE, take a look at Professor Dawn Song’s Tweet and LinkedIn Post, along with the team’s blog:
Save the date! The Agentic AI Summit returns to Berkeley on August 1–2, 2026, welcoming 5,000+ expected in-person attendees for two days of insights and innovation. Building on last year’s sold-out success—with 2,000+ in‑person attendees and 40,000+ global livestream participants—the summit will bring together researchers, builders, industry leaders, and the global agentic AI community for keynotes, technical talks and panels, hands-on workshops, live demos, and more!
In addition, we are excited to showcase our expanded list of speakers for the Summit! We are honored to have such a great group of academics, founders, executives, and investors participate in this year’s event, and more will be announced soon!
🎟️ Early‑Bird Pricing (Limited Capacity)
A limited number of early‑bird tickets are still available:
Standard Early-Bird: $399
Partner with us to shape the future of Agentic AI. If you’re interested in sponsoring the summit, please complete the sponsorship application form. Sponsorship opportunities are limited and reviewed/allocated on a rolling basis, so we encourage you to apply early.
Google DeepMind released DiffusionGemma, an experimental 26B-parameter Mixture-of-Experts model that applies diffusion techniques to text generation. Unlike traditional autoregressive models that generate text one token at a time, DiffusionGemma generates and refines entire blocks of text in parallel, enabling up to 4x faster inference on GPUs. The model activates only 3.8B parameters during inference, can run on high-end consumer hardware, and offers advantages for tasks such as in-line editing, code infilling, and other non-linear text generation workflows. While Google positions DiffusionGemma primarily as a research model and notes that output quality remains below standard Gemma 4 models, it is suitable and effective for local and low-concurrency inference.
Visa announced a strategic collaboration with OpenAI to support the growth of agentic commerce, bringing Visa’s payment network, tokenization technology, and fraud monitoring infrastructure into OpenAI’s AI-powered experiences. The partnership is designed to enable AI agents to securely complete purchases and other commerce-related tasks on behalf of users while operating within consumer-defined guardrails such as spending limits and approval requirements. The companies also plan to explore broader integrations across developer tools, business workflows, and Codex-powered applications in the future.
Prometheus, the industrial AI startup led by Jeff Bezos and former Google executive Vik Bajaj, announced a $12 billion Series B funding round at a $41 billion valuation. The company aims to accelerate engineering and manufacturing by developing what it calls an “artificial general engineer” capable of helping design and optimize complex physical products. Rather than focusing on factory automation or robotics, Prometheus is targeting pre-production workflows such as prototyping and industrial design. “The pace of our physical creation right now is nowhere near the pace of human imagination,” Bajaj said, arguing that AI can dramatically reduce the time required to bring new products from concept to production and enable faster innovation across industries including aerospace, medical devices, and advanced manufacturing.
Don’t miss the developments shaping Agentic AI. Subscribe for weekly coverage of groundbreaking research, emerging trends, and critical insights across Agentic AI and the broader AI landscape.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.