Our eval scores went up and our thumbs-down rate went up in the same week
Summary
A team observed that both automated evaluation scores and user thumbs-down rates increased in the same week, suggesting a mismatch between objective metrics and user satisfaction.
Similar Articles
[ICML 2026] Scores increased and then decreased!! [D]
A researcher discusses their ICML 2026 paper review experience where a reviewer increased their score during rebuttal but then decreased it again, expressing concern about rejection prospects.
Quantifying the Statistical Effect of Rubric Modifications on Human-Autorater Agreement
This study analyzes how modifications to evaluation rubrics, such as shifting from holistic to analytic criteria, impact the agreement between human raters and AI autoraters. The findings suggest that providing examples and reducing bias improves agreement, while higher complexity tends to decrease it.
Featuring Every Eval Ever Results on Hugging Face Model Pages
Every Eval Ever (EEE) and Hugging Face Community Evals are now intercompatible, allowing standardized cross-posting of AI evaluation results to model pages, improving trust and comparability.
@OpenAI: Let’s talk about evals. We’re always looking for better ways to measure and forecast model progress, especially as benc…
OpenAI discusses the importance of evals (evaluations) for measuring and forecasting model progress, especially as benchmarks become saturated or gamed, featuring insights from Tejal Patwardhan and Andrew Mayne.
Spent a night trying to beat our own AI virality score. Here's why it wouldn't move.
The author recounts a night spent attempting to manipulate their own AI virality scoring system, only to find that the score refused to change, demonstrating its robustness.