HuRo: Robotizing Human Videos for Scalable VLA Pretraining
Summary
This paper presents HuRo, a pipeline for robotizing human videos to create scalable VLA pretraining data, showing significant improvements in task completion and robustness on real-world manipulation tasks.
View Cached Full Text
Cached at: 09/22/26, 03:24 AM
Paper page - HuRo: Robotizing Human Videos for Scalable VLA Pretraining
Source: https://huggingface.co/papers/2609.10706
Abstract
Humanvideodatasetsofferanabundantanddiversesourceofinteractiondatathatcancomplementexpensivereal-robotdata.Tobridgethehuman-to-robotembodimentgap,existingapproacheseitherrobotizevideosintask-matchedsettingsoraddressobservationandactionalignmentseparatelyatscale.Inthiswork,wesystematicallyexaminewhetherrobotizedhumanvideoscanserveasaneffectiveandscalablesourceofsupervisionforVLApretraining.Tothisend,wedeveloparobotizationpipelinethatconvertsheterogeneoushumanvideosintorobot-alignedobservationsandactiontrajectorieswhileinferringmissingintermediatesignalsacrossannotationlevels.Usingthispipeline,weconstructtheHuRodataset,comprisingabout630Krobotizedepisodesand142Mprocessedframesfromfivehuman-videosources.Acrossfourreal-worldmanipulationtasks,increasingtheamountofrobotizedpretrainingdataimprovesoverallcompletionfrom51.5%to80.3%andOODcompletionunderspatialandvisualshiftsfrom34.9%to72.2%.AblationsfurthershowthatvisualrobotizationimprovesOODrobustnessandthatend-to-endpretrainingwithretargetedactionsoutperformsvisual-onlytransfer.Projectwebsite:https://3587jjh.github.io/HuRo.
View arXiv pageView PDFProject pageGitHub29Add to collection
Get this paper in your agent:
hf papers read 2609\.10706
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.10706 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.10706 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.10706 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining
This paper finds that egocentric human video, when processed with a filtering and labeling pipeline, can outperform teleoperated real-robot data for pretraining embodied foundation models, achieving lower validation loss and higher success rates on real-robot tasks.
Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack
HyVLA-0.5 is an end-to-end robotic learning system that integrates data collection, model design, pre-training, fine-tuning, and reinforcement learning for real-world deployment.
@HuggingPapers: Scaling video pre-training to 120K hours ZimaBlue frames video scaling as a route to generalizable World Action Models.…
Scaling video pre-training to 120K hours boosts zero-shot success in World Action Models from 36.1% to 77.8% on real robots, enabling faster action prediction.
ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training
ZimaBlue introduces a scalable framework for learning generalizable world action models from large-scale egocentric video, substantially improving zero-shot robotic manipulation through a three-stage curriculum and slow-fast architecture.
ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining
ACE-EGO-0 is a unified Vision-Language-Action pretraining framework that leverages egocentric human videos and robot trajectories via a reliability-aware training objective, achieving state-of-the-art on embodied AI benchmarks.