@svpino: I don't trust benchmarks. We've all seen this movie: New model beats everyone else on a benchmark. People hype it. Then…

X AI KOLs Timeline Tools

Summary

The tweet critiques traditional AI benchmarks and introduces TRACES, a new benchmark that evaluates AI's discovery process by focusing on how models reach answers, including tool usage, error correction, and evidence tracing.

I don't trust benchmarks. We've all seen this movie: New model beats everyone else on a benchmark. People hype it. Then we all start using it. Model turns out to be garbage. This here looks interesting. Instead of focusing on whether a model got the right answer (like a benchmark would), it evaluates how the model got there. • Did it use the right tools? • Did it catch and fix its mistakes? • Did it consider competing hypotheses? • Did it stay consistent through the process? • Can it trace its claims back to actual evidence? • Does it understand the limits of its conclusions? Bullish on this. https://traces.apodex.com
Original Article
View Cached Full Text

Cached at: 08/21/26, 11:15 PM

I don’t trust benchmarks. We’ve all seen this movie:

New model beats everyone else on a benchmark. People hype it. Then we all start using it. Model turns out to be garbage.

This here looks interesting.

Instead of focusing on whether a model got the right answer (like a benchmark would), it evaluates how the model got there.

• Did it use the right tools? • Did it catch and fix its mistakes? • Did it consider competing hypotheses? • Did it stay consistent through the process? • Can it trace its claims back to actual evidence? • Does it understand the limits of its conclusions?

Bullish on this.

https://traces.apodex.com


TRACES — A New Paradigm for Discoverative AI

Source: https://traces.apodex.com/

TRACES is a new paradigm forDiscoverative AI. It is built not to measure how well AI answers known questions, but how far it can push the frontier of the unknown.

Apodex (@Apodex_AI): Most AI benchmarks test retrieval — can a model find the known answer? However, the hardest problems in science require discovery, can a system earn an answer nobody has yet?

Meet TRACES 🧭 — the world’s first benchmark for measuring discoverative AI: AI that can work through

Similar Articles

we benchmark models nobody actually runs

Reddit r/LocalLLaMA

The article critiques the lack of systematic benchmarking for AI models in different quantization formats, highlighting discrepancies between benchmark results and real-world usage, and calls for more thorough evaluation.