Data as of Aug 25, 2026 · Based on 3,181,687 AI responses across 10,525 prompts · See how Parse measures this
Terminal Bench is a collection of Harbor-native benchmarks for evaluating AI agents in terminal environments. It provides tasks like building a Linux kernel, configuring a Git webserver, and training a FastText model to measure agent performance.
Parse Score