AutoDataBench: A Data-centric Testbed for Accelerating Auto Research
Summary
AutoDataBench introduces a controlled, data-centric testbed to isolate and evaluate LLM agents' "data intelligence" — their ability to diagnose, organize, and construct training data — while holding non-data factors fixed, showing trajectories can also be reused for mid-training gains.
View Cached Full Text
Cached at: 10/02/26, 04:28 AM
Paper page - AutoDataBench: A Data-centric Testbed for Accelerating Auto Research
Source: https://huggingface.co/papers/2609.40097 Authors:
,
,
,
,
,
,
,
,
,
,
Abstract
Existingauto-researchbenchmarksoftenentanglemultiplesourcesofimprovement,includingtrainingframeworks,hyperparameters,computebudgets,anddata,makingitdifficulttoattributewhyonefrontieragentoutperformsanothertospecificresearchcapabilities.Inthiswork,weisolateandsystematicallyevaluateDataIntelligence:anagent’sabilitytounderstand,manipulate,andimprovethedatathatshapesmodelcapabilities.WeintroduceAutoDataBench,acontrolledtestbedbuiltonaconceptualframeworkofdataintelligencespanningdatadiagnosis,dataorganization,anddataconstruction,instantiatedthroughthreehighlycuratedoptimizationtaskswhileholdingnon-datafactorsfixed.Acrosstooluse,retrieval,andknowledgeinjection,weevaluatefrontierLLMs’abilitytoimprovetrainingdatathroughiterativeexperimentationundertask-specificresourcebudgets.Beyondoptimizationperformance,weask:doLLMsunderstandwhattheirdatainterventionsdo?Wecomparepredictionsmadebeforetrainingwithobservedoutcomestoseekevidenceofdata-effectreasoningbeyondtrialanderror,andexplorewhetheriterativefeedbackhelpsLLMsbetterunderstandhowchangestotrainingdataaffectmodelperformance.Finally,weshowthatreusingAutoDataBenchtrajectoriesformid-trainingimprovesdownstreamcodingperformance,highlightingitsvalueinbothevaluatingdataintelligenceandgeneratinghigh-qualitytrainingdata.Codeandresourcesareavailableathttps://github.com/AutoDataBench/AutoDataBench.
View arXiv pageView PDFGitHub3Add to collection
Get this paper in your agent:
hf papers read 2609\.40097
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper3
#### AutoDataBench/Retrieval-resources Updatedabout 11 hours ago
#### AutoDataBench/Knowledge-Injection-resources Question Answering• Updatedabout 11 hours ago
#### AutoDataBench/Function-Calling-resources Updatedabout 11 hours ago
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.40097 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.40097 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
AutoDataBench: Can Agents Write the Data That Feeds the Self-Improvement Loop?
AutoDataBench introduces a benchmark to evaluate if agents can write tasks for data pipelines that meet practical acceptance standards, aiming to enable scalable data synthesis for recursive self-improvement in AI.
AgenticDataBench: A Comprehensive Benchmark for Data Agents
Introduces AgenticDataBench, a comprehensive benchmark for evaluating LLM-based data agents across diverse domains with fine-grained skill-based metrics, including real-world B2B use cases and synthetic tasks.
AutoMedBench: Towards Medical AutoResearch with Agentic AI Models
AutoMedBench is a workflow-aware benchmark for autonomous medical-AI research, evaluating agents across five stages on diverse medical imaging tasks. Stage-level scoring reveals validation as the weakest stage, highlighting the need for reliable verification in agentic workflows.
AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research
Introduces AutoWorldModel-Bench, a closed-loop benchmark for evaluating AI coding agents on autonomous world-model research across eight game environments. The benchmark shows frontier agents like Codex-5.4 and Claude Opus 4.6 make non-trivial research-style improvements in most sessions.
Agents That Build Better Training Data (25 minute read)
Autodata introduces an agentic data scientist that iteratively generates and refines synthetic training data, with meta-optimization to further improve data quality, achieving better results on computer science and legal reasoning tasks.