EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos
Summary
EgoSteer presents a full-stack system that pre-trains a vision-language-action model from egocentric human videos for steerable dexterous manipulation, enabling robust generalization across 40+ diverse tasks with 75%+ success on complex long-horizon tasks.
View Cached Full Text
Cached at: 07/14/26, 12:14 PM
Paper page - EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos
Source: https://huggingface.co/papers/2607.09701 Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
Steerabilityisadefiningcapabilityofgeneralistrobotpolicies,yetremainslargelyabsentindexterous-handsystemsforlackoflarge-scale,language-aligned,andaction-accuratedemonstrationdata.Toaddressthisbottleneck,wepresentafull-stacksystemthatscalesdexterousVLApre-trainingfromegocentrichumanvideosandenablesdata-efficientreal-robotpost-training.ItintegratesEgoSmith,adatapipelinethatcuratesin-the-wildegocentricvideosinto9.6Khoursofhigh-qualitypre-trainingdatawith9xhigherthroughputandbetteraccuracythanpriorSOTA;aunifiedrobotstackforteleoperationandhuman-in-the-loopcorrection;andEgoSteer,aworld-model-enhancedVLAtrainedonoptimizedinfrastructure.Human-datapre-trainingequipsEgoSteerwithlanguage-guidedmanipulationpriors,whicharegroundedthroughrobotpost-trainingandimprovedbyDAggerrefinement.Empirically,EgoSteerrobustlyexecutesfree-forminstructionsacross40+diversetasks,demonstratingfailurerecovery,dexterity,andgeneralization.Thepre-trainedmodelalsofew-shotadaptstocomplexlong-horizontasks,includingboxfolding,ontwoembodimentswith75+%success.Weopen-sourcethesystem,data,andmodelathttps://egosteer.github.io/.
View arXiv pageView PDFProject pageAdd to collection
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2607.09701 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2607.09701 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2607.09701 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
EgoPhys: Learning Generalizable Physics Models of Deformable Objects from Egocentric Video
EgoPhys introduces a framework to construct deformable physical digital twins from egocentric RGB video using generalizable priors and a compact codebook, enabling zero-shot generalization to unseen objects without per-spring optimization. The system is demonstrated on a real robot, showing that egocentric human play video can serve as internal world representation for deformable-object planning.
ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining
ACE-EGO-0 is a unified Vision-Language-Action pretraining framework that leverages egocentric human videos and robot trajectories via a reliability-aware training objective, achieving state-of-the-art on embodied AI benchmarks.
Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data
Ego2Robot is a scalable pipeline that converts egocentric human manipulation videos into robot training data via action retargeting and visual synthesis, producing 18,561 hours of data across 15 robot morphologies. Experiments show that joint pretraining on this synthesized data improves out-of-distribution generalization for vision-language-action models, including on real-robot deployment.
DeVI: Physics-based Dexterous Human-Object Interaction via Synthetic Video Imitation
DeVI introduces a framework that turns text-conditioned synthetic videos into physically plausible dexterous robot control via a hybrid 3D-2D tracking reward, enabling zero-shot generalization to unseen objects.
EgoForce: Forearm-Guided Camera-Space 3D Hand Pose from a Monocular Egocentric Camera
EgoForce is a monocular 3D hand reconstruction framework that uses a unified network with differentiable forearm representation, arm-hand transformers, and ray space solvers to recover absolute hand pose and position across different camera models, achieving state-of-the-art accuracy on egocentric benchmarks.