frontier-benchmarks

Tag

Cards List
#frontier-benchmarks

What Makes a Terminal-Bench Task Hard? Separating Genuine Hardness from Fake-Hardness on an Adjudicated Agentic Corpus

arXiv cs.LG ↗ · 2026-09-24 Cached

This paper examines why tasks in agentic AI benchmarks like Terminal-Bench may fail to be solved, distinguishing between genuine difficulty and issues such as broken oracles or infrastructure failures, and emphasizes the need to validate all-fail tasks for accurate capability claims.

0 favorites 0 likes
← Back to home

Submit Feedback