I've been measuring 34 LLM APIs every day since August to see when they quietly change. One did.
Summary
A developer has been running daily fixed probes against 34 LLM APIs from 15 labs since August, hashing readings into a public transparency log to detect silent model changes. The project's first notable finding is that DeepSeek's reasoner began using 10-12x more thinking tokens on Sept 10 with no announcement, making it slower and more expensive rather than worse.
Similar Articles
We use LLMs to analyze every file in your codebase. Everyone told us this was a stupid idea because of cost but it wasnt.
A benchmark study demonstrates that using LLMs to analyze entire codebases is cost-effective, identifying DeepSeek V4 Flash as the optimal default model due to its low cost and comparable accuracy to premium options like Claude Opus.
DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
DeepSeek LLM is an open-source language model project that develops a large dataset and employs SFT and DPO to achieve performance surpassing LLaMA-2 70B and GPT-3.5 in various benchmarks and open-ended evaluations.
Building independent LLM drift detection - sharing the methodology, looking for feedback on the approach
The author shares a methodology for building an external LLM drift detection system that continuously probes model behavior (schema adherence, instruction-following, refusal rates, etc.) to catch silent degradations in API performance, and invites feedback on the approach, pricing, and use cases.
Open-source LLM benchmark runs 147 coding tasks every 4 hours, 5-trial median with 95% CI, and uses CUSUM for change-point detection. Curious what people think of the methodology
An open-source LLM benchmark with 147 coding tasks runs every 4 hours, using 5-trial median with 95% confidence intervals and CUSUM for change-point detection, sparking discussion on its methodology.
I analyzed 31,352 hourly LLM benchmark scores: within-day variation was 2.8 points, while between-day variation was 8.4 [P]
An analysis of 31,352 hourly LLM benchmark scores shows between-day variation is about three times greater than within-day variation, emphasizing the importance of continuous monitoring for performance drift, leading to the creation of the AIStupidLevel system.