Data as of Aug 16, 2026 · Based on 3,131,739 AI responses across 10,525 prompts · See how Parse measures this
They research and evaluate frontier AI systems to understand their capabilities and potential risks, including how AI agents might evade monitoring and affect evaluations. They build tools and datasets (Hawk for large-scale agent evaluations, MALT) and publish frontier risk assessments (Frontier Risk Report) to measure autonomous performance and mitigation strategies. They collaborate with major AI companies and policymakers, conducting independent reviews, policy analyses, and transparency guidance to improve safety and risk management in frontier AI deployment.
Parse Score
Sources
en.wikipedia.org shapes more of what AI says about METR (Model Evaluation & Threat Research) than any other source, at 86% of its citations.
metr.org