@svpino: I don't trust benchmarks. We've all seen this movie: New model beats everyone else on a benchmark. People hype it. Then…
Summary
The tweet critiques traditional AI benchmarks and introduces TRACES, a new benchmark that evaluates AI's discovery process by focusing on how models reach answers, including tool usage, error correction, and evidence tracing.
View Cached Full Text
Cached at: 08/21/26, 11:15 PM
I don’t trust benchmarks. We’ve all seen this movie:
New model beats everyone else on a benchmark. People hype it. Then we all start using it. Model turns out to be garbage.
This here looks interesting.
Instead of focusing on whether a model got the right answer (like a benchmark would), it evaluates how the model got there.
• Did it use the right tools? • Did it catch and fix its mistakes? • Did it consider competing hypotheses? • Did it stay consistent through the process? • Can it trace its claims back to actual evidence? • Does it understand the limits of its conclusions?
Bullish on this.
https://traces.apodex.com
TRACES — A New Paradigm for Discoverative AI
Source: https://traces.apodex.com/
TRACES is a new paradigm forDiscoverative AI. It is built not to measure how well AI answers known questions, but how far it can push the frontier of the unknown.
Apodex (@Apodex_AI): Most AI benchmarks test retrieval — can a model find the known answer? However, the hardest problems in science require discovery, can a system earn an answer nobody has yet?
Meet TRACES 🧭 — the world’s first benchmark for measuring discoverative AI: AI that can work through
Similar Articles
@thsottiaux: Do you still trust benchmarks or do you just listen to your friends? What makes you try a new model?
A tweet questioning the trustworthiness of benchmarks and asking what drives users to try new AI models.
@heyshrutimishra: 90% of AI progress is an illusion built on benchmarks that reward memorization over thinking. TRACES strips that away w…
TRACES is a new AI benchmark by Apodex designed to measure discoverative AI by focusing on the problem-solving process rather than memorized answers, challenging traditional benchmarks.
AI benchmarks matter less than whether models can handle boring real-world responsibility
The article argues that AI benchmarks and flashy demos are overemphasized; the real test for AI trustworthiness is how models handle boring real-world responsibilities like following instructions, admitting uncertainty, handling edge cases, and being auditable.
we benchmark models nobody actually runs
The article critiques the lack of systematic benchmarking for AI models in different quantization formats, highlighting discrepancies between benchmark results and real-world usage, and calls for more thorough evaluation.
Does anyone else feel like AI benchmarks are becoming less useful for predicting real-world performance?
The article discusses the growing disconnect between high AI benchmark scores and actual real-world performance, highlighting issues like consistency, latency, and context handling.