Tag
The article critiques AI leaderboards for oversimplifying agent evaluations by hiding the impact of harnesses, using Questflow's financial-intelligence benchmark as an example, and emphasizes the need for comprehensive reporting in agent benchmarks.