WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents

Hugging Face Daily Papers Papers

Summary

WhatWorkedBench is a benchmark designed to measure the experimental understanding of AI agents by evaluating their accuracy in predicting outcomes after budgeted experimentation across various tasks and configurations.

AI research agents need reliable knowledge of how their experiments change outcomes. We introduce WhatWorkedBench to measure experimental understanding, the accuracy of predictions about component changes after budgeted experimentation. Agents inspect code, select measurements, and submit a response surface, a table predicting scores for every configuration of component settings. Exhaustive CPU execution supplies reference effects for changing each component while holding the others fixed. These effects capture combinations of changes across 36 tasks from 30 data sources and 8 workflow types, with 1248 configuration records. Core evaluation combines 4,206 numerical-control records across all eight families and 108 agent episodes across the original six. At eight new measurements, pair-effect ridge selects an optimum on 15 of 22 sources and limits every effect error to 10% of score range on three. Fitting a Gaussian process (GP) to the same agent observations raises effect recovery, accuracy relative to true effect magnitude, from 0.632 to 0.698 in the original Flash cohort and from 0.621 to 0.720 in an additional cohort. On six completed beat-detection and graph submissions, the same-observation GP raises family-macro recovery from 0.303 to 0.455. On six workflows with six binary options at 20 new measurements, encoding code equivalences, configurations with identical behavior, raises GP recovery from 0.248 to 0.462. WhatWorkedBench supports research on experimental agents, adaptive experimental design, numerical inference, and use of program structure.
Original Article
View Cached Full Text

Cached at: 09/24/26, 03:38 AM

Paper page - WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents

Source: https://huggingface.co/papers/2609.27490

Abstract

AIresearchagentsneedreliableknowledgeofhowtheirexperimentschangeoutcomes.WeintroduceWhatWorkedBenchtomeasureexperimentalunderstanding,theaccuracyofpredictionsaboutcomponentchangesafterbudgetedexperimentation.Agentsinspectcode,selectmeasurements,andsubmitaresponsesurface,atablepredictingscoresforeveryconfigurationofcomponentsettings.ExhaustiveCPUexecutionsuppliesreferenceeffectsforchangingeachcomponentwhileholdingtheothersfixed.Theseeffectscapturecombinationsofchangesacross36tasksfrom30datasourcesand8workflowtypes,with1248configurationrecords.Coreevaluationcombines4,206numerical-controlrecordsacrossalleightfamiliesand108agentepisodesacrosstheoriginalsix.Ateightnewmeasurements,pair-effectridgeselectsanoptimumon15of22sourcesandlimitseveryeffecterrorto10%ofscorerangeonthree.FittingaGaussianprocess(GP)tothesameagentobservationsraiseseffectrecovery,accuracyrelativetotrueeffectmagnitude,from0.632to0.698intheoriginalFlashcohortandfrom0.621to0.720inanadditionalcohort.Onsixcompletedbeat-detectionandgraphsubmissions,thesame-observationGPraisesfamily-macrorecoveryfrom0.303to0.455.Onsixworkflowswithsixbinaryoptionsat20newmeasurements,encodingcodeequivalences,configurationswithidenticalbehavior,raisesGPrecoveryfrom0.248to0.462.WhatWorkedBenchsupportsresearchonexperimentalagents,adaptiveexperimentaldesign,numericalinference,anduseofprogramstructure.

View arXiv pageView PDFProject pageGitHub1Add to collection

Get this paper in your agent:

hf papers read 2609\.27490

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.27490 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.27490 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.27490 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Good Benchmarks

arXiv cs.AI

This paper from July 2026 defines principles for designing effective benchmark tasks for AI agents, drawing on experience from the Terminal Bench project. Good tasks are correct, solvable, verifiable, well-specified, and hard for interesting reasons, with an emphasis on real-world relevance and outcome-based verification.

JobBench: Aligning Agent Work With Human Will

arXiv cs.AI

JobBench is a benchmark built from worker surveys to evaluate AI agents on tasks that workers most want automated, covering 130 tasks across 35 professions with detailed rubrics.