AutoDataBench: A Data-centric Testbed for Accelerating Auto Research

Hugging Face Daily Papers Papers

Summary

AutoDataBench introduces a controlled, data-centric testbed to isolate and evaluate LLM agents' "data intelligence" — their ability to diagnose, organize, and construct training data — while holding non-data factors fixed, showing trajectories can also be reused for mid-training gains.

Existing auto-research benchmarks often entangle multiple sources of improvement, including training frameworks, hyperparameters, compute budgets, and data, making it difficult to attribute why one frontier agent outperforms another to specific research capabilities. In this work, we isolate and systematically evaluate Data Intelligence: an agent's ability to understand, manipulate, and improve the data that shapes model capabilities. We introduce AutoDataBench, a controlled testbed built on a conceptual framework of data intelligence spanning data diagnosis, data organization, and data construction, instantiated through three highly curated optimization tasks while holding non-data factors fixed. Across tool use, retrieval, and knowledge injection, we evaluate frontier LLMs' ability to improve training data through iterative experimentation under task-specific resource budgets. Beyond optimization performance, we ask: do LLMs understand what their data interventions do? We compare predictions made before training with observed outcomes to seek evidence of data-effect reasoning beyond trial and error, and explore whether iterative feedback helps LLMs better understand how changes to training data affect model performance. Finally, we show that reusing AutoDataBench trajectories for mid-training improves downstream coding performance, highlighting its value in both evaluating data intelligence and generating high-quality training data. Code and resources are available at https://github.com/AutoDataBench/AutoDataBench.
Original Article
View Cached Full Text

Cached at: 10/02/26, 04:28 AM

Paper page - AutoDataBench: A Data-centric Testbed for Accelerating Auto Research

Source: https://huggingface.co/papers/2609.40097 Authors:

,

,

,

,

,

,

,

,

,

,

Abstract

Existingauto-researchbenchmarksoftenentanglemultiplesourcesofimprovement,includingtrainingframeworks,hyperparameters,computebudgets,anddata,makingitdifficulttoattributewhyonefrontieragentoutperformsanothertospecificresearchcapabilities.Inthiswork,weisolateandsystematicallyevaluateDataIntelligence:anagent’sabilitytounderstand,manipulate,andimprovethedatathatshapesmodelcapabilities.WeintroduceAutoDataBench,acontrolledtestbedbuiltonaconceptualframeworkofdataintelligencespanningdatadiagnosis,dataorganization,anddataconstruction,instantiatedthroughthreehighlycuratedoptimizationtaskswhileholdingnon-datafactorsfixed.Acrosstooluse,retrieval,andknowledgeinjection,weevaluatefrontierLLMs’abilitytoimprovetrainingdatathroughiterativeexperimentationundertask-specificresourcebudgets.Beyondoptimizationperformance,weask:doLLMsunderstandwhattheirdatainterventionsdo?Wecomparepredictionsmadebeforetrainingwithobservedoutcomestoseekevidenceofdata-effectreasoningbeyondtrialanderror,andexplorewhetheriterativefeedbackhelpsLLMsbetterunderstandhowchangestotrainingdataaffectmodelperformance.Finally,weshowthatreusingAutoDataBenchtrajectoriesformid-trainingimprovesdownstreamcodingperformance,highlightingitsvalueinbothevaluatingdataintelligenceandgeneratinghigh-qualitytrainingdata.Codeandresourcesareavailableathttps://github.com/AutoDataBench/AutoDataBench.

View arXiv pageView PDFGitHub3Add to collection

Get this paper in your agent:

hf papers read 2609\.40097

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper3

#### AutoDataBench/Retrieval-resources Updatedabout 11 hours ago #### AutoDataBench/Knowledge-Injection-resources Question Answering• Updatedabout 11 hours ago #### AutoDataBench/Function-Calling-resources Updatedabout 11 hours ago

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.40097 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.40097 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

AgenticDataBench: A Comprehensive Benchmark for Data Agents

Hugging Face Daily Papers

Introduces AgenticDataBench, a comprehensive benchmark for evaluating LLM-based data agents across diverse domains with fine-grained skill-based metrics, including real-world B2B use cases and synthetic tasks.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models

Hugging Face Daily Papers

AutoMedBench is a workflow-aware benchmark for autonomous medical-AI research, evaluating agents across five stages on diverse medical imaging tasks. Stage-level scoring reveals validation as the weakest stage, highlighting the need for reliable verification in agentic workflows.

Agents That Build Better Training Data (25 minute read)

TLDR AI

Autodata introduces an agentic data scientist that iteratively generates and refines synthetic training data, with meta-optimization to further improve data quality, achieving better results on computer science and legal reasoning tasks.