DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces
Summary
Introduces DataSpace, a benchmark for evaluating data agents on verifiable tabular analytics over heterogeneous workspaces, containing 410 cross-language tasks and 7,439 artifacts. Current frontier models achieve only 66.34% accuracy, indicating headroom.
View Cached Full Text
Cached at: 08/07/26, 01:57 PM
Paper page - DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces
Source: https://huggingface.co/papers/2608.03451 Authors:
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
Dataagentsenablenatural-languageanalyticsoverorganizationalworkspaces,whererelevantevidencemaybescatteredacrossdatabases,structuredfiles,longdocuments,andmultimedia.Existingbenchmarkslargelyisolatestructuredquerying,retrieval,oropen-endedanalysis,leavingheterogeneousevidencediscovery,completetabularoutputs,anddeterministicevaluationinsufficientlyunified.WeintroduceDataSpace,abenchmarkinwhichdataagentsproduceverifiabletabularresultsfromtask-localheterogeneousworkspaces.Itcontains410cross-languagetasksand7,439artifactstotaling15.01GBacrossCSV,JSON,SQLite,Markdown,PDF,andvideo.DataSpacealsoservedastheofficialevaluationbenchmarkfortheKDDCup2026DataAgentsforComplexDataAnalysiscompetition.Eachagentreceivesonlyaquestionandworkspaceandreturnsthecompleterequestedtabularresult.WeconstructDataSpacewithDataSpace-Builder,anexecution-groundedframeworkcomprisingcross-languagetransformation,constraint-awarerelationalsampling,modalityroutingandartifactrendering,andhumanreviewandtaskrepairby11domainexperts.Adeterministicevaluatorperformsheader-invariantcolumnalignment,type-andprecision-awarenormalization,andorder-awarerowcomparison.Acrosssixrecentlyreleasedfrontiermultimodalmodelsandfivewidelyusedagentharnesses,thebestaccuracyreaches66.34%,whileharnesschoicecreatesa15.36-pointspreadwiththebackbonefixed.Multimodalevidenceintegrationandjoinsconsistentlyreduceaccuracyacrossallsixbackbones.TheseresultsshowthatDataSpaceremainsunsaturatedandidentifykeychallengesforimprovingdata-agentreliability.
View arXiv pageView PDFProject pageGitHub11Add to collection
Get this paper in your agent:
hf papers read 2608\.03451
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2608.03451 in a model README.md to link it from this page.
Datasets citing this paper1
#### HKUSTDial/DataSpace Viewer• Updated2 days ago • 410 • 142 • 3
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2608.03451 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?
This paper introduces DSAgentBench, the first benchmark for evaluating autonomous agents on complete, multi-tool data-science workflows in real computer environments. Results show that even the strongest agent (Claude-4.6-Sonnet) achieves only 56.70% task success, while open-source agents remain below 1%, revealing a substantial capability gap.
AgenticDataBench: A Comprehensive Benchmark for Data Agents
Introduces AgenticDataBench, a comprehensive benchmark for evaluating LLM-based data agents across diverse domains with fine-grained skill-based metrics, including real-world B2B use cases and synthetic tasks.
Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
Argo-Bench introduces an evaluation framework of 210 enterprise-scale data science tasks built on a simulated food-delivery platform with an ERP warehouse of 235 tables and 7.5 billion rows, testing agents that must navigate and act rather than just generate SQL. The strongest of 14 frontier and open-weight models averages only 59.5 points, highlighting a significant gap in real-world data agent capabilities.
Benchmarking AI Agents for Addressing Scientific Challenges Across Scales
Introduces SciAgentArena, a benchmark of ~200 tasks for evaluating AI agents in real scientific research. Finds agents effective for well-specified data-analysis workflows but struggle with novel insights and open-ended exploration.
Skill-based Agentic Evaluation for Real-time Data Science Tasks
This paper introduces a ground-truth-as-code framework for evaluating data-science agents on continuously updated data, achieving improved agreement with human evaluators and reduced token consumption.