
9 Big questions benchmarks can help answer
A good benchmark asks more than just “can AI do this specific task?”
Insights and analysis on AI trends, developments, and research from Epoch AI.
Live Last read · last published · next check

A good benchmark asks more than just “can AI do this specific task?”

Why Anthropic's buildout suggests that financing is unlikely to be the immediate blocker to frontier AI compute growth.

Expanding FrontierMath: Open Problems, how "parallelizability" determines a technological singularity, the realities of AI energy use, and signs of AI uplift

Expert assessments and cyber benchmarks led us to expect that frontier models were capable of executing this kind of cyberattack

A board game AI can't master, a cyber-disclosure spike after Claude Mythos, GPT-4's record run atop the ECI, and what's missing from AI futurism discourse

Why we should think a little harder about what it takes to build a Dyson Sphere

Our new long-horizon coding benchmark, hyperscaler cash flows, tracking AI R&D automation, and Chinese lab strategies

Inferring Chinese AI labs’ strategies from their job descriptions

Proposing a new way to track AI research automation

Increased cyber vulnerability reports, Mythos' cyber hype, Fable 5's lead on FrontierMath v2, record-setting data centers, and what wealth distribution could look like after AGI