Agent benchmarks.

Every task we run, against every model, on the same harness. Including the runs we lose.

Internal Bench Hard

Browser Use: 82% of tasks solved at 17¢ each. That is 20 points better than Opus 5, which costs 20× more per solved task.

cheaper and better than every model← cheaperbetter ↑$1.00$2.00$3.0020%30%40%50%60%70%80%90%strict accuracycost per solved taskBrowser UseGPT-5Gemini 3.6 FlashGPT-5.6Sonnet 5Gemini 3.1 ProOpus 5
Toolstrict accuracyCost
Browser Use82%$0.17
Opus 562%$3.40
Gemini 3.1 Pro59%$2.20
Sonnet 559%$1.55
GPT-5.652%$1.10
Gemini 3.6 Flash46%$0.62
GPT-537%$0.44
Internal Bench Hard · updated 2026-08-01 · all benchmarks
What it measures

Our hardest internal set: 106 tasks that a careful human can finish but most agents cannot. Every model runs on the same harness, and we count a task solved only on a strict match, so partial credit earns nothing.

How it was run
  • Tasks: 106 hard tasks on live websites, none removed.
  • Cost: total recorded spend for the run divided by the number of tasks actually solved.
  • Scoring: strict — a task counts only when the final answer or end state is exactly right.

We use cookies to improve your experience. Privacy