Data as of Aug 25, 2026 · Based on 3,181,687 AI responses across 10,525 prompts · See how Parse measures this
JailbreakBench is a centralized benchmark for evaluating jailbreaking attacks on large language models, offering a repository of jailbreak artifacts, a standardized evaluation library with a defined threat model, prompts, templates, and scoring, and leaderboards comparing attack and defense performance across open- and closed-source models. It hosts the JBB-Behaviors dataset—100 misuse behaviors (55% original) plus 100 benign behaviors—organized into categories aligned with OpenAI's usage policies to assess safety and overrefusal rates. The project welcomes community contributions of new attacks and defenses and provides open-source tooling and documentation, with ongoing expansion to reflect advances and to support citation.
Parse Score