Data as of Aug 25, 2026 · Based on 3,181,687 AI responses across 10,525 prompts · See how Parse measures this
LiveBench is a contamination-free benchmark and leaderboard for evaluating large language models (LLMs) using questions that have verifiable, objective ground-truth answers. It currently contains 21 diverse tasks across 7 categories and refreshes its questions every six months to minimize test-set contamination, with the latest version being LiveBench-2025-11-25. Sponsored by Abacus.AI, it enables objective model comparison across tasks such as reasoning, coding, mathematics, language, and data analysis on an openly accessible leaderboard.
Parse Score
Sources
livebench.ai shapes more of what AI says about LiveBench than any other source, at 100% of its citations.
Excerpts where LiveBench appeared in the AI's answer

LiveBench : Focuses on contamination-free, objective multi-category evaluations including agentic coding and data analysis with tracked costs per task.

LiveBench : A contamination-free benchmark platform that refreshes questions regularly