Browser Use: 82% of tasks solved at 17¢ each. That is 20 points better than Opus 5, which costs 20× more per solved task.
Our hardest internal set: 106 tasks that a careful human can finish but most agents cannot. Every model runs on the same harness, and we count a task solved only on a strict match, so partial credit earns nothing.
We use cookies to improve your experience. Privacy