proxy-benchmark

Tag

Cards List
#proxy-benchmark

@Xudong07452910: Evaluating agent capabilities is way too expensive these days! Running an agentic benchmark like SWE-Bench or GAIA often requires complex environments, toolchains, long rollouts, and can cost hundreds or thousands of dollars. This PACE paper from CMU asks a very practical question...

X AI KOLs Timeline · 2026-07-06 Cached

The PACE method proposed by Carnegie Mellon University selects a subset of atomic capabilities from cheap non-agentic benchmarks to predict model performance on expensive agentic benchmarks, with prediction error below 4% and cost reduced to less than 1%.

0 favorites 0 likes
← Back to home

Submit Feedback