Data as of Aug 25, 2026 · Based on 3,181,687 AI responses across 10,525 prompts · See how Parse measures this
SWE bench is a benchmark for evaluating large language models on real-world software engineering tasks, offering subsets like Verified, Lite, Multilingual, and Multimodal. It provides leaderboards and metrics such as the percentage of resolved instances to compare model performance.
Sources
arxiv.org shapes more of what AI says about SWE-bench than any other source, at 24% of its citations.
swebench.com · agentsdirectory.dev · aimultiple.com · anthropic.com
The market map
LLM Observability and Evaluation Platforms →Where AI ranks SWE-bench
Excerpts where SWE-bench appeared in the AI's answer

SWE-bench / SWE-bench Verified: Measures end-to-end software engineering capabilities on real GitHub issues.

SWE-bench: The industry standard for evaluating AI agents on real-world Github issues
Excerpts where SWE-bench appeared in the AI's answer

SWE-bench / MLE-bench: Benchmark and environment ecosystems explicitly designed for long-horizon software engineering