HumanNet: Scaling Human-centric Video Learning to One Million Hours
Summary
HumanNet is a large-scale human-centric video dataset with one million hours of annotated footage, designed to train vision-language-action models. It demonstrates that egocentric human video can effectively replace robot data for embodied intelligence tasks.
View Cached Full Text
Cached at: 05/11/26, 02:42 AM
Paper page - HumanNet: Scaling Human-centric Video Learning to One Million Hours
Source: https://huggingface.co/papers/2605.06747
Abstract
HumanNet presents a large-scale human-centric video dataset with rich annotations for embodied intelligence, demonstrating that egocentric human video can effectively replace robot data for training vision-language-action models.
Progress inembodied intelligenceincreasingly depends on scalable data infrastructure. While vision and language have scaled with internet corpora, learning physical interaction remains constrained by the lack of large, diverse, and richly annotated human activity data. We present HumanNet, a one-million-hourhuman-centricvideo corpus that captures how humans interact with the physical world at scale. HumanNet spans both first-person and third-person perspectives and covers fine-grained activities, human-object interactions, tool use, and long-horizon behaviors across diverse real-world environments. Beyond raw video, the dataset provides interaction-centric annotations, including captions, motion descriptions, and hand and body-related signals, enabling motion-aware and interaction-aware learning. Beyond scale, HumanNet introduces a systematic data curation paradigm for embodied learning, wherehuman-centricfiltering, temporal structuring, viewpoint diversity, and annotation enrichment are treated as first-class design principles. This design transforms unstructured internet video into a scalable substrate forrepresentation learning,activity understanding,motion generation, andhuman-to-robot transfer. We conduct a first-step validation on the value of this design through controlledvision-language-actionablation: under a fixed set of validation data, continued training from the Qwen VLM model with 1000 hours ofegocentric videodrawn from HumanNet surpasses the continued training with 100 hours of real-robot data fromMagic Cobot, indicating that egocentric human video could be a scalable and cost-effective substitute for robot data. By building this project, we aim to explore the opportunity to scale embodied foundation models usinghuman-centricvideos, rather than relying solely on robot-specific data.
View arXiv pageView PDFProject pageGitHubAdd to collection
Get this paper in your agent:
hf papers read 2605\.06747
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2605.06747 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2605.06747 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2605.06747 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
@LeRobotHF: 100,000 hours of human hands doing real work. LeRobot format, fully annotated and open. EgoSuite-Open100K from @Lightwh…
LightwheelAI and Hugging Face have released EgoSuite-Open100K, a large-scale open dataset of 100,000 hours of egocentric human activities, fully annotated and available in LeRobot format for AI research and training.
HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining
This paper finds that egocentric human video, when processed with a filtering and labeling pipeline, can outperform teleoperated real-robot data for pretraining embodied foundation models, achieving lower validation loss and higher success rates on real-robot tasks.
Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data
Ego2Robot is a scalable pipeline that converts egocentric human manipulation videos into robot training data via action retargeting and visual synthesis, producing 18,561 hours of data across 15 robot morphologies. Experiments show that joint pretraining on this synthesized data improves out-of-distribution generalization for vision-language-action models, including on real-robot deployment.
@seclink: 视频里是加速了的,实际上动作慢很多 ...
Lightwheel AI and Hugging Face have open-sourced EgoSuite-Open100K, the largest fully annotated egocentric human dataset with 100,000 hours of data, emphasizing scaling laws for physical AI.
LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training
LAION-BVD is a large-scale open video dataset containing 10 million hours of video data for multimodal pre-training, with synthetic captions and competitive benchmark performance.