AutoDataBench: Can Agents Write the Data That Feeds the Self-Improvement Loop?
Summary
AutoDataBench introduces a benchmark to evaluate if agents can write tasks for data pipelines that meet practical acceptance standards, aiming to enable scalable data synthesis for recursive self-improvement in AI.
View Cached Full Text
Cached at: 09/30/26, 04:16 AM
Paper page - AutoDataBench: Can Agents Write the Data That Feeds the Self-Improvement Loop?
Source: https://huggingface.co/papers/2609.35025
Abstract
Recentgainsinlanguagemodelcapabilityhavecomemorefromdatathanfromarchitecture.Frontierlabsanddatacompaniesproduceverifiableagentictasks,whichsupervisedfinetuningandreinforcementlearningthenturnintocapability.Thisproductionlinestillrestsonhumanlabourandonhuman-in-the-loopcollaboration.Automatingtaskcreationwouldletdataproductionscalewithcomputeratherthanwithexpertheadcount,wouldextendtomoredomains,andwouldenableakeystepinrecursiveself-improvement(RSI).Currentevaluationsofanagent’sabilitytowritesuchtasksmeasurehowamodelperformsaftertrainingonwhattheagentproduced.Thatdoesnotmatchcommonpracticeinthedataindustry,wheredataisdeliveredsamplebysampleandeachsampleisacceptedagainstasetofcriteriaratherthanputstraightintotraining.Noexistingevaluationaskswhetheranindividualtaskmeetstheacceptancecriteriaofadatapipeline.WethereforeintroduceAutoDataBench.Givenanoriginalbenchmarktaskandarecordofthetargetmodelattemptingit,anagentmustwriteanewtaskforthesamesuitethatmeetspracticalacceptancestandardsonvalidity,novelty,difficultyandbehaviouralcoverage.Acrossthreebenchmarksofexecutableagenttasks,noagentweevaluatescoresabove20outof100atthedefaulttimebudgetof45minutes.Givingthestrongestagentfourtimesaslongimprovesitsscoresubstantially,whilethecostofoneusabletaskstaysalmostunchanged.Currentagentscanwritetrainingtasksoftherequiredquality,butnotefficiently.AutoDataBenchprovidesadirectmeasureofanagent’scapacityforautonomousdatasynthesis:oneartifactatatime,judgedagainstthecriteriaaproductionpipelinewouldapply,andwithoutatrainingrun.Codeanddataareavailableathttps://github.com/StarDewXXX/AutoDataBench.
View arXiv pageView PDFProject pageGitHub18Add to collection
Get this paper in your agent:
hf papers read 2609\.35025
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.35025 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.35025 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.35025 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
AutoDataBench: A Data-centric Testbed for Accelerating Auto Research
AutoDataBench introduces a controlled, data-centric testbed to isolate and evaluate LLM agents' "data intelligence" — their ability to diagnose, organize, and construct training data — while holding non-data factors fixed, showing trajectories can also be reused for mid-training gains.
DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?
This paper introduces DSAgentBench, the first benchmark for evaluating autonomous agents on complete, multi-tool data-science workflows in real computer environments. Results show that even the strongest agent (Claude-4.6-Sonnet) achieves only 56.70% task success, while open-source agents remain below 1%, revealing a substantial capability gap.
AutoMedBench: Towards Medical AutoResearch with Agentic AI Models
AutoMedBench is a workflow-aware benchmark for autonomous medical-AI research, evaluating agents across five stages on diverse medical imaging tasks. Stage-level scoring reveals validation as the weakest stage, highlighting the need for reliable verification in agentic workflows.
AgenticDataBench: A Comprehensive Benchmark for Data Agents
Introduces AgenticDataBench, a comprehensive benchmark for evaluating LLM-based data agents across diverse domains with fine-grained skill-based metrics, including real-world B2B use cases and synthetic tasks.
Neurodata Without Boredom: Benchmarking Agentic AI for Data Reuse
This paper benchmarks agentic AI systems on the task of loading, understanding, and reformatting fragmented neuroscience data, finding that while agents perform well on subtasks, they rarely achieve fully error-free end-to-end solutions and human oversight remains necessary.