A single AI leaderboard score hides the part that may be changing: the harness
Summary
The article critiques AI leaderboards for oversimplifying agent evaluations by hiding the impact of harnesses, using Questflow's financial-intelligence benchmark as an example, and emphasizes the need for comprehensive reporting in agent benchmarks.
Similar Articles
We NEED a harness benchmark leaderboard
This article argues for the need of a benchmark leaderboard that compares AI model harnesses (e.g., KimiCode vs OpenCode vs Codex) rather than just models themselves, proposing a repo to test model+harness combinations on cost, runtime, token usage, and score.
Stop Comparing LLM Agents Without Disclosing the Harness
This position paper argues that in long-horizon LLM agent tasks, the execution harness often determines performance more than the model itself, and current benchmarks misattribute harness-level gains to model improvements. It proposes a harness-aware evaluation framework with disclosure standards and variance decomposition protocols.
@AlphaSignalAI: https://x.com/AlphaSignalAI/status/2057153343081111582
A 100-page survey from UIUC, Meta, and Stanford introduces three harness layers (Interface, Mechanisms, Scaling) for AI agents, arguing that most agent failures stem from harness issues rather than reasoning flaws, and provides a taxonomy for auditing agent stacks.
Rethinking the Evaluation of Harness Evolution for Agents
This paper rethinks how automatic harness evolution for agents should be evaluated, showing that gains may be due to increased compute rather than genuine improvements, and that evolved harnesses transfer poorly to unseen tasks.
@Sentdex: Dear AI companies: Please immediately stop harnessmaxxing on your benchamaxxed harnesses for your AI releases
A tweet from Sentdex urging AI companies to stop over-optimizing benchmark harnesses in AI releases, criticizing the practice of gaming evaluation metrics.