DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces
Summary
Introduces DataSpace, a benchmark for evaluating data agents on verifiable tabular analytics over heterogeneous workspaces, containing 410 cross-language tasks and 7,439 artifacts. Current frontier models achieve only 66.34% accuracy, indicating headroom.
View Cached Full Text
Cached at: 08/07/26, 01:57 PM
Paper page - DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces
Source: https://huggingface.co/papers/2608.03451 Authors:
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
Dataagentsenablenatural-languageanalyticsoverorganizationalworkspaces,whererelevantevidencemaybescatteredacrossdatabases,structuredfiles,longdocuments,andmultimedia.Existingbenchmarkslargelyisolatestructuredquerying,retrieval,oropen-endedanalysis,leavingheterogeneousevidencediscovery,completetabularoutputs,anddeterministicevaluationinsufficientlyunified.WeintroduceDataSpace,abenchmarkinwhichdataagentsproduceverifiabletabularresultsfromtask-localheterogeneousworkspaces.Itcontains410cross-languagetasksand7,439artifactstotaling15.01GBacrossCSV,JSON,SQLite,Markdown,PDF,andvideo.DataSpacealsoservedastheofficialevaluationbenchmarkfortheKDDCup2026DataAgentsforComplexDataAnalysiscompetition.Eachagentreceivesonlyaquestionandworkspaceandreturnsthecompleterequestedtabularresult.WeconstructDataSpacewithDataSpace-Builder,anexecution-groundedframeworkcomprisingcross-languagetransformation,constraint-awarerelationalsampling,modalityroutingandartifactrendering,andhumanreviewandtaskrepairby11domainexperts.Adeterministicevaluatorperformsheader-invariantcolumnalignment,type-andprecision-awarenormalization,andorder-awarerowcomparison.Acrosssixrecentlyreleasedfrontiermultimodalmodelsandfivewidelyusedagentharnesses,thebestaccuracyreaches66.34%,whileharnesschoicecreatesa15.36-pointspreadwiththebackbonefixed.Multimodalevidenceintegrationandjoinsconsistentlyreduceaccuracyacrossallsixbackbones.TheseresultsshowthatDataSpaceremainsunsaturatedandidentifykeychallengesforimprovingdata-agentreliability.
View arXiv pageView PDFProject pageGitHub11Add to collection
Get this paper in your agent:
hf papers read 2608\.03451
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2608.03451 in a model README.md to link it from this page.
Datasets citing this paper1
#### HKUSTDial/DataSpace Viewer• Updated2 days ago • 410 • 142 • 3
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2608.03451 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
AgenticDataBench: A Comprehensive Benchmark for Data Agents
Introduces AgenticDataBench, a comprehensive benchmark for evaluating LLM-based data agents across diverse domains with fine-grained skill-based metrics, including real-world B2B use cases and synthetic tasks.
Benchmarking AI Agents for Addressing Scientific Challenges Across Scales
Introduces SciAgentArena, a benchmark of ~200 tasks for evaluating AI agents in real scientific research. Finds agents effective for well-specified data-analysis workflows but struggle with novel insights and open-ended exploration.
SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks
SpatialWorld is a unified benchmark for evaluating interactive spatial reasoning in multimodal agents across diverse real-world tasks, revealing that even the strongest models achieve low task success rates.
TerraBench: Can Agents Reason Over Heterogeneous Earth-System Data?
TerraBench is a new benchmark for evaluating AI agents' ability to reason over heterogeneous Earth-system data, including gridded data, satellite imagery, and simulator outputs. It reveals significant limitations in current frontier models, with top performers achieving only 59.2% tool-use score on average.
OR-Space: A Full-Lifecycle Workspace Benchmark for Industrial Optimization Agents
OR-Space is a benchmark for evaluating large language model agents in industrial operations research workflows, focusing on multi-stage task lifecycles and persistent workspaces beyond simple text generation.