Data as of Aug 25, 2026 · Based on 3,181,687 AI responses across 10,525 prompts · See how Parse measures this
τ-bench is a benchmarking platform from Sierra for evaluating AI agents in collaborative, real-world scenarios, emphasizing coordination and tool use to achieve shared objectives across enterprise domains. It features a public leaderboard, results submission, and domain-specific tracks (retail, airline, telecom), with recent expansion into telecom and dual-control environments using user simulators. The project has notable milestones, including GPT-5 achieving state-of-the-art performance on τ-bench (96% telecom, 82% retail, 63% airline) and an ICLR 2025 paper acceptance, with ongoing updates and research.
Parse Score
#42 of 114 in LLM Observability and Evaluation Platforms
Sources
automationanywhere.com shapes more of what AI says about τ‑Bench than any other source, at 33% of its citations.
github.com · google.com · medium.com · sierra.ai
The market map
LLM Observability and Evaluation Platforms →Where AI ranks τ‑Bench