Tag
This paper examines why tasks in agentic AI benchmarks like Terminal-Bench may fail to be solved, distinguishing between genuine difficulty and issues such as broken oracles or infrastructure failures, and emphasizes the need to validate all-fail tasks for accurate capability claims.