eval-set

Tag

Cards List
#eval-set

Your eval set stops being an eval set the moment your agent optimizes against it

Reddit r/AI_Agents · 2026-07-29

A detailed discussion on how evaluation sets degrade when used for selection in agentic loops, with three defenses: counting trials with deflated scoring, progressively revealing eval data, and placing the scorer behind an uncallable tool boundary.

0 favorites 0 likes
#eval-set

I stopped trusting model benchmarks and started running my own eval set, here is what changed[D]

Reddit r/MachineLearning · 2026-06-25

The author describes losing faith in public AI model benchmarks due to vendor-created metrics, self-reported parameters, and lack of independent verification, and advocates for building custom evaluation sets from real production traffic to make more relevant model comparisons.

0 favorites 0 likes
← Back to home

Submit Feedback