@anshnanda: The current benchmarks no longer work for the new models. That’s why we are seeing results like this.
Summary
The tweet by @anshnanda points out that existing AI benchmarks are not suitable for evaluating new models, explaining unexpected results in model performance.
View Cached Full Text
Cached at: 09/23/26, 08:15 PM
The current benchmarks no longer work for the new models.
That’s why we are seeing results like this.
Similar Articles
@svpino: I don't trust benchmarks. We've all seen this movie: New model beats everyone else on a benchmark. People hype it. Then…
The tweet critiques traditional AI benchmarks and introduces TRACES, a new benchmark that evaluates AI's discovery process by focusing on how models reach answers, including tool usage, error correction, and evidence tracing.
Does anyone else feel like AI benchmarks are becoming less useful for predicting real-world performance?
The article discusses the growing disconnect between high AI benchmark scores and actual real-world performance, highlighting issues like consistency, latency, and context handling.
Time for a new benchmark
The article discusses the need for a new benchmark in AI to better evaluate model performance and address current limitations in existing standards.
New benchmark dropped
A new benchmark has been released, likely for evaluating AI or software performance.
@vasuman: The only benchmark that matters is how AI power users on Twitter feel about your model
A tweet by @vasuman suggests that the sentiment of AI power users on Twitter is the most crucial benchmark for evaluating AI models.