ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine
Summary
Introduces ACE-Data-0, a large-scale embodied AI dataset with 150 hours of synchronized multimodal human demonstrations across 200 task categories, captured by the Ambient Capture Engine (ACE) in real home environments. Includes a hierarchical benchmark exposing gaps in current methods under contact, occlusion, and long horizons.
View Cached Full Text
Cached at: 07/31/26, 05:53 AM
Paper page - ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine
Source: https://huggingface.co/papers/2607.28625 Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
Embodiedintelligencefacesafundamentaldatabottleneck.Modelsmustcapturehowfirst-personperception,whole-bodymotion,dexterousmanipulation,objectstate,sound,andtouchevolvetogetherashumanspursuegoalsovertime.Existingdatasetsfragmentthisexperienceacrossviewpoints,modalities,orspatialscales,leavingthefullperception-actionlooponlypartiallyobserved.WeintroducetheAmbientCaptureEngine(ACE),ahuman-centricdataenginethattransformsrealhomeenvironmentsintospatiallycalibrated,temporallysynchronizedrecordingstudios.ACEoperatesattwocomplementaryscales:atable-scaleconfigurationresolveshand-objectmanipulation,whilearoom-scaleconfigurationcaptureswhole-bodymotion,locomotion,andinteractionsacrossafurnishedhome.ACErecordsegocentricandmulti-viewexocentricvideo,full-bodyandarticulatedhandmotion,objectgeometryand6-DoFtrajectories,audio,andtactilesignalsasaunifiedmultisensorystream.UsingACE,webuildACE-Data-0,comprising150hoursand17Mvideoframesacross200taskcategories,performedby50participantsin2environments,foratotalof75,000interactionepisodes.Thedatasetspansatomicmanipulation,long-horizonchainsofhouseholdactivities,andhuman-sceneinteraction,whilepreservingnaturalbehavioralvariationthroughgoal-levelratherthanstep-by-stepinstructions.Wefurtherintroduceahierarchicalbenchmarkthatprogressesfromsignalstoscenecomponentsandthentointeractions.Evaluationsofstate-of-the-artmethodsexposesubstantialgapsundercontact,occlusion,egomotion,andlongtemporalhorizons.ACE-Data-0providessynchronizedhumandemonstrationswithalignedperceptual,kinematic,andcontactsupervision,offeringascalablefoundationforimitationlearning,worldmodels,vision-language-actionsystems,andembodiedAI.
View arXiv pageView PDFProject pageAdd to collection
Get this paper in your agent:
hf papers read 2607\.28625
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2607.28625 in a model README.md to link it from this page.
Datasets citing this paper1
#### ACERobotics/ACE-Data-0 Updatedabout 3 hours ago • 22
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2607.28625 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning
Open-AoE is an open, community-oriented egocentric manipulation dataset and toolchain that spans from smartphone capture to model training, providing approximately 2,000 hours of manipulation video with annotations and downstream tools for embodied learning.
ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining
ACE-EGO-0 is a unified Vision-Language-Action pretraining framework that leverages egocentric human videos and robot trajectories via a reliability-aware training objective, achieving state-of-the-art on embodied AI benchmarks.
HumanNet: Scaling Human-centric Video Learning to One Million Hours
HumanNet is a large-scale human-centric video dataset with one million hours of annotated footage, designed to train vision-language-action models. It demonstrates that egocentric human video can effectively replace robot data for embodied intelligence tasks.
Embodied-BenchClaw: An Autonomous Multi-Agent System for Embodied Spatial Intelligence Benchmark Construction
This paper proposes Embodied-BenchClaw, an autonomous multi-agent system that automatically constructs embodied spatial intelligence benchmarks from user intent through a five-stage pipeline with process quality control and an extensible Skill Library.
Data Pyramid for Embodied Manipulation
This paper organizes embodied data sources into a five-level pyramid (real-robot, UMI, egocentric/exocentric, simulation, general vision-language), analyzing their trade-offs between scalability and robot alignment, and reviews recent embodied foundation models in terms of data recipes. It also discusses open challenges for building next-generation embodied systems.