Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning

Hugging Face Daily Papers Papers

Summary

Open-AoE is an open, community-oriented egocentric manipulation dataset and toolchain that spans from smartphone capture to model training, providing approximately 2,000 hours of manipulation video with annotations and downstream tools for embodied learning.

Egocentric videos of human manipulation provide scalable supervision for embodied intelligence, yet existing resources rarely combine low-cost continuous capture, manipulation-level structured annotations, and reusable tools for robot learning. We present Open-AoE, an open, community-oriented egocentric manipulation dataset and toolchain spanning the full pipeline from smartphone capture to model training. Its first release contains approximately 2,000 hours of manipulation video collected in natural environments by 500+ contributors using 400+ smartphones. The dataset provides text annotations, MANO-based hand poses, camera trajectories, and temporally localized atomic actions. Open-AoE further includes a data processing pipeline that transforms raw recordings into structured samples through temporal action segmentation, semantic annotation, hand reconstruction, and camera trajectory reconstruction. Meanwhile, we provide a separate downstream toolchain supports visualization, cross-embodiment retargeting, model-specific data conversion, and training recipes for VLA policies, WAMs, and World Models. By integrating scalable capture, structured processing, and downstream adaptation, Open-AoE reduces the barriers to both data contribution and reuse, providing practical open infrastructure for embodied model training, human-to-robot transfer, and world modeling.
Original Article
View Cached Full Text

Cached at: 07/21/26, 02:37 PM

Paper page - Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning

Source: https://huggingface.co/papers/2607.14183 Authors:

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

Abstract

Egocentricvideosofhumanmanipulationprovidescalablesupervisionforembodiedintelligence,yetexistingresourcesrarelycombinelow-costcontinuouscapture,manipulation-levelstructuredannotations,andreusabletoolsforrobotlearning.WepresentOpen-AoE,anopen,community-orientedegocentricmanipulationdatasetandtoolchainspanningthefullpipelinefromsmartphonecapturetomodeltraining.Itsfirstreleasecontainsapproximately2,000hoursofmanipulationvideocollectedinnaturalenvironmentsby500+contributorsusing400+smartphones.Thedatasetprovidestextannotations,MANO-basedhandposes,cameratrajectories,andtemporallylocalizedatomicactions.Open-AoEfurtherincludesadataprocessingpipelinethattransformsrawrecordingsintostructuredsamplesthroughtemporalactionsegmentation,semanticannotation,handreconstruction,andcameratrajectoryreconstruction.Meanwhile,weprovideaseparatedownstreamtoolchainsupportsvisualization,cross-embodimentretargeting,model-specificdataconversion,andtrainingrecipesforVLApolicies,WAMs,andWorldModels.Byintegratingscalablecapture,structuredprocessing,anddownstreamadaptation,Open-AoEreducesthebarrierstobothdatacontributionandreuse,providingpracticalopeninfrastructureforembodiedmodeltraining,human-to-robottransfer,andworldmodeling.

View arXiv pageView PDFGitHub79Add to collection

Get this paper in your agent:

hf papers read 2607\.14183

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2607.14183 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2607.14183 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2607.14183 in a Space README.md to link it from this page.

Collections including this paper1

Similar Articles

ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine

Hugging Face Daily Papers

Introduces ACE-Data-0, a large-scale embodied AI dataset with 150 hours of synchronized multimodal human demonstrations across 200 task categories, captured by the Ambient Capture Engine (ACE) in real home environments. Includes a hierarchical benchmark exposing gaps in current methods under contact, occlusion, and long horizons.

robbyant/lingbot-video-moe-30b-a3b

Hugging Face Models Trending

LingBot-Video is the first open-source large-scale MoE video generation model for embodied intelligence, featuring efficient MoE architecture, massive embodied data training, and multi-reward system for high aesthetics, physical rationality, and task completion.

EgoCS-400K: An Egocentric Gameplay Dataset for World Models

Hugging Face Daily Papers

EgoCS-400K is a large-scale egocentric Counter-Strike dataset with over 400,000 first-person videos and 10,000 hours of gameplay, providing temporally aligned video-action-language trajectories for world model research.