WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents
Summary
WhatWorkedBench is a benchmark designed to measure the experimental understanding of AI agents by evaluating their accuracy in predicting outcomes after budgeted experimentation across various tasks and configurations.
View Cached Full Text
Cached at: 09/24/26, 03:38 AM
Paper page - WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents
Source: https://huggingface.co/papers/2609.27490
Abstract
AIresearchagentsneedreliableknowledgeofhowtheirexperimentschangeoutcomes.WeintroduceWhatWorkedBenchtomeasureexperimentalunderstanding,theaccuracyofpredictionsaboutcomponentchangesafterbudgetedexperimentation.Agentsinspectcode,selectmeasurements,andsubmitaresponsesurface,atablepredictingscoresforeveryconfigurationofcomponentsettings.ExhaustiveCPUexecutionsuppliesreferenceeffectsforchangingeachcomponentwhileholdingtheothersfixed.Theseeffectscapturecombinationsofchangesacross36tasksfrom30datasourcesand8workflowtypes,with1248configurationrecords.Coreevaluationcombines4,206numerical-controlrecordsacrossalleightfamiliesand108agentepisodesacrosstheoriginalsix.Ateightnewmeasurements,pair-effectridgeselectsanoptimumon15of22sourcesandlimitseveryeffecterrorto10%ofscorerangeonthree.FittingaGaussianprocess(GP)tothesameagentobservationsraiseseffectrecovery,accuracyrelativetotrueeffectmagnitude,from0.632to0.698intheoriginalFlashcohortandfrom0.621to0.720inanadditionalcohort.Onsixcompletedbeat-detectionandgraphsubmissions,thesame-observationGPraisesfamily-macrorecoveryfrom0.303to0.455.Onsixworkflowswithsixbinaryoptionsat20newmeasurements,encodingcodeequivalences,configurationswithidenticalbehavior,raisesGPrecoveryfrom0.248to0.462.WhatWorkedBenchsupportsresearchonexperimentalagents,adaptiveexperimentaldesign,numericalinference,anduseofprogramstructure.
View arXiv pageView PDFProject pageGitHub1Add to collection
Get this paper in your agent:
hf papers read 2609\.27490
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.27490 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.27490 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.27490 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows
StartupBench introduces a benchmark for evaluating general-purpose AI agents on real-world startup workflows, revealing that top models complete only about 30% of tasks due to gaps in complex instruction following and domain-specific expertise.
Good Benchmarks
This paper from July 2026 defines principles for designing effective benchmark tasks for AI agents, drawing on experience from the Terminal Bench project. Good tasks are correct, solvable, verifiable, well-specified, and hard for interesting reasons, with an emphasis on real-world relevance and outcome-based verification.
JobBench: Aligning Agent Work With Human Will
JobBench is a benchmark built from worker surveys to evaluate AI agents on tasks that workers most want automated, covering 130 tasks across 35 professions with detailed rubrics.
AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research
Introduces AutoWorldModel-Bench, a closed-loop benchmark for evaluating AI coding agents on autonomous world-model research across eight game environments. The benchmark shows frontier agents like Codex-5.4 and Claude Opus 4.6 make non-trivial research-style improvements in most sessions.
K-Bench: measuring model performance on real scientific agent requests
The paper introduces K-Bench 01, a benchmark for evaluating AI agents on real scientific requests, revealing that no model consistently meets the threshold for acceptable performance, with overclaiming as a common failure.