Data as of Aug 25, 2026 · Based on 3,181,687 AI responses across 10,525 prompts · See how Parse measures this
HarmBench is a standardized evaluation framework for automated red teaming and robust refusal of large language models (LLMs), designed to uncover and mitigate the risks of malicious LLM use. It enables rigorous benchmarking by supporting two use cases: evaluating red-teaming methods against LLMs and evaluating LLMs against red-teaming methods, all through a unified evaluation pipeline. The project also introduces an efficient adversarial training method to improve LLM robustness across diverse attacks, illustrating codevelopment of attacks and defenses.
Parse Score
Sources
arxiv.org shapes more of what AI says about HarmBench than any other source, at 22% of its citations.
datatonic.com · dextralabs.com · evidentlyai.com · linkedin.com