Tag
An analysis of 31,352 hourly LLM benchmark scores shows between-day variation is about three times greater than within-day variation, emphasizing the importance of continuous monitoring for performance drift, leading to the creation of the AIStupidLevel system.
The paper presents ACES, a framework for continuous evaluation of AI agent skills through live trials, measuring Skill Lift to quantify added value, and demonstrating its effectiveness on enterprise repositories compared to scan-only gates.
AI agents often drift in production after passing evals due to distribution shifts and upstream changes; continuous evaluation and real-time monitoring can mitigate this.