AutoDataBench: Can Agents Write the Data That Feeds the Self-Improvement Loop?

Hugging Face Daily Papers Papers

Summary

AutoDataBench introduces a benchmark to evaluate if agents can write tasks for data pipelines that meet practical acceptance standards, aiming to enable scalable data synthesis for recursive self-improvement in AI.

Recent gains in language model capability have come more from data than from architecture. Frontier labs and data companies produce verifiable agentic tasks, which supervised finetuning and reinforcement learning then turn into capability.This production line still rests on human labour and on human-in-the-loop collaboration. Automating task creation would let data production scale with compute rather than with expert headcount, would extend to more domains, and would enable a key step in recursive self-improvement (RSI). Current evaluations of an agent's ability to write such tasks measure how a model performs after training on what the agent produced. That does not match common practice in the data industry, where data is delivered sample by sample and each sample is accepted against a set of criteria rather than put straight into training. No existing evaluation asks whether an individual task meets the acceptance criteria of a data pipeline. We therefore introduce AutoDataBench. Given an original benchmark task and a record of the target model attempting it, an agent must write a new task for the same suite that meets practical acceptance standards on validity, novelty, difficulty and behavioural coverage. Across three benchmarks of executable agent tasks, no agent we evaluate scores above 20 out of 100 at the default time budget of 45 minutes. Giving the strongest agent four times as long improves its score substantially, while the cost of one usable task stays almost unchanged. Current agents can write training tasks of the required quality, but not efficiently. AutoDataBench provides a direct measure of an agent's capacity for autonomous data synthesis: one artifact at a time, judged against the criteria a production pipeline would apply, and without a training run. Code and data are available at https://github.com/StarDewXXX/AutoDataBench.
Original Article
View Cached Full Text

Cached at: 09/30/26, 04:16 AM

Paper page - AutoDataBench: Can Agents Write the Data That Feeds the Self-Improvement Loop?

Source: https://huggingface.co/papers/2609.35025

Abstract

Recentgainsinlanguagemodelcapabilityhavecomemorefromdatathanfromarchitecture.Frontierlabsanddatacompaniesproduceverifiableagentictasks,whichsupervisedfinetuningandreinforcementlearningthenturnintocapability.Thisproductionlinestillrestsonhumanlabourandonhuman-in-the-loopcollaboration.Automatingtaskcreationwouldletdataproductionscalewithcomputeratherthanwithexpertheadcount,wouldextendtomoredomains,andwouldenableakeystepinrecursiveself-improvement(RSI).Currentevaluationsofanagent’sabilitytowritesuchtasksmeasurehowamodelperformsaftertrainingonwhattheagentproduced.Thatdoesnotmatchcommonpracticeinthedataindustry,wheredataisdeliveredsamplebysampleandeachsampleisacceptedagainstasetofcriteriaratherthanputstraightintotraining.Noexistingevaluationaskswhetheranindividualtaskmeetstheacceptancecriteriaofadatapipeline.WethereforeintroduceAutoDataBench.Givenanoriginalbenchmarktaskandarecordofthetargetmodelattemptingit,anagentmustwriteanewtaskforthesamesuitethatmeetspracticalacceptancestandardsonvalidity,novelty,difficultyandbehaviouralcoverage.Acrossthreebenchmarksofexecutableagenttasks,noagentweevaluatescoresabove20outof100atthedefaulttimebudgetof45minutes.Givingthestrongestagentfourtimesaslongimprovesitsscoresubstantially,whilethecostofoneusabletaskstaysalmostunchanged.Currentagentscanwritetrainingtasksoftherequiredquality,butnotefficiently.AutoDataBenchprovidesadirectmeasureofanagent’scapacityforautonomousdatasynthesis:oneartifactatatime,judgedagainstthecriteriaaproductionpipelinewouldapply,andwithoutatrainingrun.Codeanddataareavailableathttps://github.com/StarDewXXX/AutoDataBench.

View arXiv pageView PDFProject pageGitHub18Add to collection

Get this paper in your agent:

hf papers read 2609\.35025

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.35025 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.35025 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.35025 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

AutoDataBench: A Data-centric Testbed for Accelerating Auto Research

Hugging Face Daily Papers

AutoDataBench introduces a controlled, data-centric testbed to isolate and evaluate LLM agents' "data intelligence" — their ability to diagnose, organize, and construct training data — while holding non-data factors fixed, showing trajectories can also be reused for mid-training gains.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models

Hugging Face Daily Papers

AutoMedBench is a workflow-aware benchmark for autonomous medical-AI research, evaluating agents across five stages on diverse medical imaging tasks. Stage-level scoring reveals validation as the weakest stage, highlighting the need for reliable verification in agentic workflows.

AgenticDataBench: A Comprehensive Benchmark for Data Agents

Hugging Face Daily Papers

Introduces AgenticDataBench, a comprehensive benchmark for evaluating LLM-based data agents across diverse domains with fine-grained skill-based metrics, including real-world B2B use cases and synthetic tasks.

Neurodata Without Boredom: Benchmarking Agentic AI for Data Reuse

arXiv cs.LG

This paper benchmarks agentic AI systems on the task of loading, understanding, and reformatting fragmented neuroscience data, finding that while agents perform well on subtasks, they rarely achieve fully error-free end-to-end solutions and human oversight remains necessary.