commonsense-benchmarks

Tag

Cards List
#commonsense-benchmarks

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks

arXiv cs.CL · 6d ago Cached

This paper tests whether commonsense benchmark scores predict real-world downstream task performance by evaluating 23 LLMs across four benchmarks and their reworked variants, finding that revisions preserve rankings but only offer task-dependent predictive validity.

0 favorites 0 likes
← Back to home

Submit Feedback