Tag
This paper systematically maps 14,767 papers to analyze how LLM benchmark designs and capability expectations have evolved from 2022 to 2026.
This paper compares expert-assigned and automatically-assigned MeSH terms as features for systematic review screening classifiers, showing that evaluation design significantly affects measured performance gaps, with canonical designs showing larger gaps that attenuate under alternative designs.