evaluation-design

Tag

Cards List
#evaluation-design

What Do We Expect from LLMs? Mapping the Design of LLM Benchmarks

arXiv cs.AI ↗ · 2026-09-18 Cached

This paper systematically maps 14,767 papers to analyze how LLM benchmark designs and capability expectations have evolved from 2022 to 2026.

0 favorites 0 likes
#evaluation-design

Evaluation design conditions the expert-vs-auto MeSH gap: a controlled comparison of bag-of-words and BiomedBERT on the Cohen benchmark

arXiv cs.CL ↗ · 2026-07-27 Cached

This paper compares expert-assigned and automatically-assigned MeSH terms as features for systematic review screening classifiers, showing that evaluation design significantly affects measured performance gaps, with canonical designs showing larger gaps that attenuate under alternative designs.

0 favorites 0 likes
← Back to home

Submit Feedback