overclaiming

Tag

Cards List
#overclaiming

K-Bench: measuring model performance on real scientific agent requests

arXiv cs.AI · 17h ago Cached

The paper introduces K-Bench 01, a benchmark for evaluating AI agents on real scientific requests, revealing that no model consistently meets the threshold for acceptable performance, with overclaiming as a common failure.

0 favorites 0 likes
← Back to home

Submit Feedback