Tag
A detailed discussion on how evaluation sets degrade when used for selection in agentic loops, with three defenses: counting trials with deflated scoring, progressively revealing eval data, and placing the scorer behind an uncallable tool boundary.
The author describes losing faith in public AI model benchmarks due to vendor-created metrics, self-reported parameters, and lack of independent verification, and advocates for building custom evaluation sets from real production traffic to make more relevant model comparisons.