X (formerly Twitter)

Last year, we introduced 𝜏-bench, a benchmark for evaluating AI agents on realistic, multi-step tasks involving tool use and domain-specific constraints. It surfaced a critical limitation in LLM-based agents: low repeatability, even under identical conditions. Now, we’re introducing its successor. 𝜏2-bench is a benchmark for dual-control environments—scenarios where success depends on coordinated action between agent and user—which reflects the kinds of tasks AI agents are increasingly being asked to perform in the real world. This shift to shared agency exposes a steep challenge: even top-tier models like GPT‑4.1 experience up to a 25-point drop in task success when moving from solo execution to interactive guidance. Other key contributions: • A compositional task generator for complex, verifiable workflows • Realistic user simulators grounded in observable state • Systematic evaluation of agent collaboration

Read the original on x.com ↗