Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning
Summary
Open-AoE is an open, community-oriented egocentric manipulation dataset and toolchain that spans from smartphone capture to model training, providing approximately 2,000 hours of manipulation video with annotations and downstream tools for embodied learning.
View Cached Full Text
Cached at: 07/21/26, 02:37 PM
Paper page - Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning
Source: https://huggingface.co/papers/2607.14183 Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
Egocentricvideosofhumanmanipulationprovidescalablesupervisionforembodiedintelligence,yetexistingresourcesrarelycombinelow-costcontinuouscapture,manipulation-levelstructuredannotations,andreusabletoolsforrobotlearning.WepresentOpen-AoE,anopen,community-orientedegocentricmanipulationdatasetandtoolchainspanningthefullpipelinefromsmartphonecapturetomodeltraining.Itsfirstreleasecontainsapproximately2,000hoursofmanipulationvideocollectedinnaturalenvironmentsby500+contributorsusing400+smartphones.Thedatasetprovidestextannotations,MANO-basedhandposes,cameratrajectories,andtemporallylocalizedatomicactions.Open-AoEfurtherincludesadataprocessingpipelinethattransformsrawrecordingsintostructuredsamplesthroughtemporalactionsegmentation,semanticannotation,handreconstruction,andcameratrajectoryreconstruction.Meanwhile,weprovideaseparatedownstreamtoolchainsupportsvisualization,cross-embodimentretargeting,model-specificdataconversion,andtrainingrecipesforVLApolicies,WAMs,andWorldModels.Byintegratingscalablecapture,structuredprocessing,anddownstreamadaptation,Open-AoEreducesthebarrierstobothdatacontributionandreuse,providingpracticalopeninfrastructureforembodiedmodeltraining,human-to-robottransfer,andworldmodeling.
View arXiv pageView PDFGitHub79Add to collection
Get this paper in your agent:
hf papers read 2607\.14183
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2607.14183 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2607.14183 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2607.14183 in a Space README.md to link it from this page.
Collections including this paper1
Similar Articles
ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine
Introduces ACE-Data-0, a large-scale embodied AI dataset with 150 hours of synchronized multimodal human demonstrations across 200 task categories, captured by the Ambient Capture Engine (ACE) in real home environments. Includes a hierarchical benchmark exposing gaps in current methods under contact, occlusion, and long horizons.
ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining
ACE-EGO-0 is a unified Vision-Language-Action pretraining framework that leverages egocentric human videos and robot trajectories via a reliability-aware training objective, achieving state-of-the-art on embodied AI benchmarks.
robbyant/lingbot-video-moe-30b-a3b
LingBot-Video is the first open-source large-scale MoE video generation model for embodied intelligence, featuring efficient MoE architecture, massive embodied data training, and multi-reward system for high aesthetics, physical rationality, and task completion.
MobileEgo Anywhere: Open Infrastructure for long horizon egocentric data on commodity hardware
MobileEgo Anywhere is a mobile-based framework for collecting long-duration egocentric robot data using smartphone sensors, enabling large-scale training of vision-language-action models by lowering hardware barriers.
EgoCS-400K: An Egocentric Gameplay Dataset for World Models
EgoCS-400K is a large-scale egocentric Counter-Strike dataset with over 400,000 first-person videos and 10,000 hours of gameplay, providing temporally aligned video-action-language trajectories for world model research.