TerminalBench 2.1 from GPT‑5.6 Sol, Terra, and Luna
Summary
TerminalBench 2.1 is a benchmark suite derived from GPT‑5.6 Sol, Terra, and Luna models, likely used for evaluating AI performance on terminal-based tasks.
Similar Articles
Introducing BenchBench (5 minute read)
Introduces BenchBench, a benchmark that tests AI models' ability to create effective benchmarks for other models, with GPT 5.2 being the only successful winner so far while frontier models like GPT 5.5 and Opus 4.6 struggled.
Terminal Bench 3 has been released. It’s a new benchmark that hasn’t been included in model training sets yet. (I’m not showing the results from third-party harnesses to keep things fair.)
Terminal Bench 3, a new benchmark for terminal-based AI models that excludes third-party harness results and has not been used in training sets, has been released.
TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents
TUA-Bench is a comprehensive benchmark for evaluating general-purpose terminal-use agents across diverse digital activities and specialized workflows, revealing significant performance gaps among current frontier agents.
Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
Terminal-Bench-Science is a benchmark developed by Stanford University researchers to evaluate AI agents on real scientific research workflows, aiming to drive AI capabilities in science.
GPT 5.6 Sol benchmarks
GPT 5.6 Sol achieves new benchmark results, showcasing performance improvements in AI language modeling.