@ErenChenAI: WuJi just open-sourced its MINT model and EgoPipeline for reconstructing 3D hand + camera motion from first-person vide…
Summary
WuJi has open-sourced its MINT model and EgoPipeline for reconstructing 3D hand and camera motion from first-person video, along with 1,021 hours of egocentric data to aid robot learning.
View Cached Full Text
Cached at: 09/08/26, 07:23 AM
WuJi just open-sourced its MINT model and EgoPipeline for reconstructing 3D hand + camera motion from first-person video, together with 1,021 hours of structured egocentric data for scaling robot learning. https://t.co/prUYiuzRuo
Similar Articles
EgoForce: Forearm-Guided Camera-Space 3D Hand Pose from a Monocular Egocentric Camera
EgoForce is a monocular 3D hand reconstruction framework that uses a unified network with differentiable forearm representation, arm-hand transformers, and ray space solvers to recover absolute hand pose and position across different camera models, achieving state-of-the-art accuracy on egocentric benchmarks.
@macrodata_labs: Everyone is betting on Egocentric data to scale robotics But turning that footage into training data requires recoverin…
Macrodata Labs releases a research blog on scaling robotics with egocentric video data by recovering 3D hand motion signals using only open-source models.
@seclink: 视频里是加速了的,实际上动作慢很多 ...
Lightwheel AI and Hugging Face have open-sourced EgoSuite-Open100K, the largest fully annotated egocentric human dataset with 100,000 hours of data, emphasizing scaling laws for physical AI.
EventEgoHands++: Event-based Egocentric 3D Hand Mesh Reconstruction with Real Dataset
This paper proposes EventEgoHands++, a framework for event-based 3D hand mesh reconstruction from an egocentric viewpoint, incorporating hand detection and adaptive attention, and introduces a new real-world dataset EEH-R for training and evaluation.
Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data
Ego2Robot is a scalable pipeline that converts egocentric human manipulation videos into robot training data via action retargeting and visual synthesis, producing 18,561 hours of data across 15 robot morphologies. Experiments show that joint pretraining on this synthesized data improves out-of-distribution generalization for vision-language-action models, including on real-robot deployment.