task-budget

Tag

Cards List
#task-budget

How Many Tasks Are Enough for Agent Benchmark Decisions? A Replay Analysis of Public LLM Agent Benchmarks

arXiv cs.AI · 2026-07-15 Cached

This paper analyzes how many tasks are needed in partial evaluations of LLM agent benchmarks to reach the same pairwise conclusions as full benchmarks. It finds that required task fractions vary sharply across benchmarks and suggests reporting standards for partial evaluations.

0 favorites 0 likes
← Back to home

Submit Feedback