Tag
The paper introduces K-Bench 01, a benchmark for evaluating AI agents on real scientific requests, revealing that no model consistently meets the threshold for acceptable performance, with overclaiming as a common failure.