we benchmark models nobody actually runs
Summary
The article critiques the lack of systematic benchmarking for AI models in different quantization formats, highlighting discrepancies between benchmark results and real-world usage, and calls for more thorough evaluation.
Similar Articles
Unsteady Metrics and Benchmarking Cultures of AI Model Builders
This paper introduces Benchmarking-Cultures-25, a dataset analyzing how AI model builders selectively highlight benchmarks in press releases. It finds a fragmented evaluation landscape with limited cross-model comparability, arguing that benchmarks are used as narrative devices for market positioning rather than standardized scientific measurement.
AI models provided by big AI corporate labs constitutes fraud by FTC's definition
The article argues that AI labs commit fraud by advertising high benchmark scores from ideal model versions while shipping heavily degraded versions (e.g., quantized, safety-stacked) that perform 50-60% worse, and proposes mandatory third-party re-benchmarking as a solution.
I stopped trusting model benchmarks and started running my own eval set, here is what changed[D]
The author describes losing faith in public AI model benchmarks due to vendor-created metrics, self-reported parameters, and lack of independent verification, and advocates for building custom evaluation sets from real production traffic to make more relevant model comparisons.
Benchmarks compare open models against closed products, not closed models. We might be missing what were actually paying for
Argues that benchmarks comparing open models against closed API products are misleading because they measure raw inference vs. hidden tooling and preprocessing, suggesting the actual model quality gap may be smaller than reported.
@svpino: I don't trust benchmarks. We've all seen this movie: New model beats everyone else on a benchmark. People hype it. Then…
The tweet critiques traditional AI benchmarks and introduces TRACES, a new benchmark that evaluates AI's discovery process by focusing on how models reach answers, including tool usage, error correction, and evidence tracing.