[R] Would you keep a robot demonstration if hand tracking missed the moment the plug went in?
Summary
MEgoVista is an offline pipeline from Maniformer that turns unprepared egocentric MEgo View recordings into metric two-hand and head motion in a gravity-aligned world frame, using calibrated stereo for metric scale and validating outputs against independent Chingmu optical motion capture. The work positions itself as a scalable, unconstrained alternative to studio rigs for generating metric hand supervision for robot manipulation learning.
View Cached Full Text
Cached at: 10/02/26, 10:47 PM
# MEgoVista: Multi-view Ego-aware Motion Estimation for Metric 4D Hands and Head in the Wild
Source: [https://arxiv.org/html/2609.16684](https://arxiv.org/html/2609.16684)
\\institution
Maniformer\\reportseriesTechnical Report\\reportnumber001
Jiangong Xiao∗†Northwestern Polytechnical Universityxiaojiangong@mail\.nwpu\.edu\.cnZhihao Zhang∗†Xi’an Jiaotong Universityzhangzh6031@stu\.xjtu\.edu\.cnYifei Dong∗Maniformerdongyifei@maniformer\.aiChao Ma∗Maniformermachao@maniformer\.aiZhouyi JinManiformerjinzhouyi@maniformer\.aiZhiwen HouManiformerhouzhiwen@maniformer\.aiLi LiuManiformerliuli@maniformer\.aiWeihuang ChenXi’an Jiaotong Universitychenwh@xjtu\.edu\.cnHongbin SunXi’an Jiaotong Universityhsun@mail\.xjtu\.edu\.cnMaoqing Yao‡Maniformeryaomaoqing@agibot\.com
###### Abstract
Learning manipulation from human video requires high\-fidelity hand\-motion reconstruction in metric units\. Today’s metric hand labels come from studio rigs and instrumented headsets, and both are confined in the same two ways: neither leaves a prepared setting, and neither is checked against an independent reference\. Unconstrained head\-worn recording promises the opposite trade\-off, scaling with the number of people wearing a device\. We therefore introduceMEgoVista, an offline pipeline that turns a single unpreparedMEgo Viewrecording into metric two\-hand and head motion in one gravity\-aligned world frame\. Three properties set it apart from existing egocentric reconstruction systems: first, it reconstructs in settings studio volumes and tabletop rigs cannot reach, settling hand ownership at detection so bystander hands stay out of the wearer’s trajectory; second, it takes its metric gauge from calibrated stereo rather than a monocular prior, installing scale at initialisation so policies receive physical units, not arbitrary coordinates; third, both outputs are scored inside a motion\-capture volume against independent Chingmu optical capture, under a protocol that audits its own reference and charges what a method declines to predict\.MEgoVistais offered as a measured route from egocentric video to metric hand supervision, one that widens where such labels can be gathered\.
###### keywords
egocentric capture, hand reconstruction, motion\-capture validation
††footnotetext:∗Equal Contribution\.†Work done at Maniformer\.‡Corresponding Author\.## 1Introduction
Egocentric video has become a central data source for embodied AI\. Learning a manipulation policy requires observations paired with the physical states and actions that produced them, but collecting such pairs on a real robot is expensive: every hour of teleoperation consumes hardware, an operator, and an environment reset\[[1](https://arxiv.org/html/2609.16684#bib.bib1)\]\. Human recording offers a different trade\-off\. A head\-mounted camera captures real contact physics, dexterous five\-fingered hands, and the diversity of everyday environments, and it requires no robot in the collection loop, so the collection rate scales with the number of people wearing a device rather than the number of available robots\[[2](https://arxiv.org/html/2609.16684#bib.bib2),[3](https://arxiv.org/html/2609.16684#bib.bib3)\]\. Egocentric datasets have grown accordingly, from hundreds of hours of passive activity capture to recent releases measured in thousands of hours\[[4](https://arxiv.org/html/2609.16684#bib.bib5),[5](https://arxiv.org/html/2609.16684#bib.bib4),[6](https://arxiv.org/html/2609.16684#bib.bib34)\], and egocentric data is now a standard component of the pretraining mixtures used by embodied foundation models\[[1](https://arxiv.org/html/2609.16684#bib.bib1)\]\. Within such data, hand motion is the signal that matters most\. Semantic annotations such as narrations and verb\-noun labels describe what a person did, but not the motion that accomplished it\. Every method that converts a human demonstration into a robot action space operates on the hand: the wrist trajectory is mapped through inverse kinematics, and finger articulation, typically parameterised by MANO\[[7](https://arxiv.org/html/2609.16684#bib.bib7)\], is transferred by morphology\-aware retargeting\[[8](https://arxiv.org/html/2609.16684#bib.bib36),[9](https://arxiv.org/html/2609.16684#bib.bib33),[10](https://arxiv.org/html/2609.16684#bib.bib35)\]\. The quality of the resulting supervision is therefore bounded by the accuracy of the estimated metric 3D hand pose\.
Despite growing adoption, egocentric collections that carry metric 3D hand labels remain narrow\. Studio rigs such as ARCTIC and Assembly101 deliver accurate poses under multi\-view optical motion capture, but cannot leave their capture volume\[[11](https://arxiv.org/html/2609.16684#bib.bib6),[12](https://arxiv.org/html/2609.16684#bib.bib14)\]\. Instrumented headsets extend that range through device\-native inside\-out tracking: EgoDex recorded 820 hours of bimanual manipulation using Apple Vision Pro\[[5](https://arxiv.org/html/2609.16684#bib.bib4)\], but remains confined to a prepared tabletop\. Large\-scale passive corpora such as Ego4D capture in\-the\-wild activity but ship no metric 3D supervision\[[2](https://arxiv.org/html/2609.16684#bib.bib2)\]\. Where 3D labels are released, their accuracy is asserted from hardware specifications rather than established against an independent metric reference\.
Recent work has advanced hand reconstruction through learned regressors and temporal optimisation\. WiLoR and HaMeR predict MANO parameters from a single crop\[[13](https://arxiv.org/html/2609.16684#bib.bib27),[14](https://arxiv.org/html/2609.16684#bib.bib8)\], while HaWoR and Dyn\-HaMR lift monocular sequences into world\-frame trajectories\[[15](https://arxiv.org/html/2609.16684#bib.bib9),[16](https://arxiv.org/html/2609.16684#bib.bib10)\]\. These methods obtain metric depth from monocular learned predictors such as UniDepthV2\[[17](https://arxiv.org/html/2609.16684#bib.bib28)\], whose scale is a prior rather than a measurement, and the reprojection term cannot correct depth error along the viewing ray\. Standard detectors distinguish only left from right, not wearer from bystander, corrupting both precision and recall when other people’s hands appear\. Occlusion remains problematic: the object, forearm, and frame edge hide the hand during the contact phases that matter most\.
Head motion tracking faces a related challenge\. Head\-worn capture is rotation\-dominated, so translation over a keyframe interval is often negligible\. Visual\-inertial trackers such as ORB\-SLAM3 recover metric scale from the inertial unit\[[18](https://arxiv.org/html/2609.16684#bib.bib18)\], but require a usable baseline between successive keyframes, and when rotation dominates, scale estimation degrades\. Offline structure\-from\-motion pipelines such as COLMAP reconstruct more accurately\[[19](https://arxiv.org/html/2609.16684#bib.bib15)\], but cannot observe scale monocularly and produce trajectories whose metric scale is arbitrary\.
To address these challenges, we presentMEgoVista, a system that reconstructs metric 3D hand and head motion from egocentric video recorded in unconstrained real\-world environments\. Unlike prior methods that remain confined to prepared tabletops or studio volumes, our approach operates on footage captured during everyday occupational tasks\. The system resolves multi\-person ambiguity, recovers metric depth without monocular learned priors, handles occlusion across multiple manipulation phases, and maintains stable head tracking when the wearer turns or moves\. We validate the reconstruction accuracy inside a motion\-capture volume and report millimetre\-scale error against an independent metric reference\. Our contributions are:
- •Robust reconstruction in unconstrained real\-world settings\.The system handles in\-the\-wild occupational environments where studio rigs and tabletop setups cannot operate\. It addresses multi\-person/multi\-hand ambiguity, occlusion robustness, and rotation\-dominated head motion, delivering supervision in settings that existing methods do not cover\.
- •Superior capture specifications and reconstruction accuracy\.Our capture achieves a wider field of view than prior egocentric systems, and reconstruction accuracy reaches sub\-centimetre scale: 0\.65 to 2\.83 mm for head trajectory and 4\.29 mm for finger articulation, measured against an independent motion\-capture reference\.
- •Trustworthy metric scale and coordinates\.Our system recovers reliable metric scale and gravity\-aligned coordinates, ensuring that downstream manipulation policies receive hand trajectories in physical units rather than arbitrary coordinates\.
## 2Related Work
Egocentric hand reconstruction\.MANO makes hand recovery a fitting problem\[[7](https://arxiv.org/html/2609.16684#bib.bib7)\], and single\-crop regressors predict its parameters with no metric depth and no coupling between frames\[[14](https://arxiv.org/html/2609.16684#bib.bib8),[13](https://arxiv.org/html/2609.16684#bib.bib27)\]— our initialiser, not a competitor\. HaWoR, Dyn\-HaMR and HaPTIC lift monocular hand motion into a world frame\[[15](https://arxiv.org/html/2609.16684#bib.bib9),[16](https://arxiv.org/html/2609.16684#bib.bib10),[20](https://arxiv.org/html/2609.16684#bib.bib11)\]and produce the output shape we do; they estimate that frame from the same stream that supplies the hand evidence, whereas we take it from calibrated hardware\. MS\-MANO routes pose through a muscle\-tendon simulator that cannot represent an impossible configuration\[[21](https://arxiv.org/html/2609.16684#bib.bib12),[22](https://arxiv.org/html/2609.16684#bib.bib13)\], and BioPR learns the equivalent priors\[[23](https://arxiv.org/html/2609.16684#bib.bib32)\]; we approximate the effect with differentiable barriers\. Feed\-forward geometry collapses the stages\[[24](https://arxiv.org/html/2609.16684#bib.bib22),[25](https://arxiv.org/html/2609.16684#bib.bib23),[26](https://arxiv.org/html/2609.16684#bib.bib24)\]and removes the point at which a known baseline can be injected\. Learned depth, stereo and segmentation supply per\-frame geometry\[[17](https://arxiv.org/html/2609.16684#bib.bib28),[27](https://arxiv.org/html/2609.16684#bib.bib29),[28](https://arxiv.org/html/2609.16684#bib.bib30),[29](https://arxiv.org/html/2609.16684#bib.bib31)\]; we consume these as initialisation and gate rather than trust them\.
Stereo depth and multi\-view geometry\.Monocular depth estimation from learned predictors such as UniDepthV2\[[17](https://arxiv.org/html/2609.16684#bib.bib28)\]and UniK3D\[[27](https://arxiv.org/html/2609.16684#bib.bib29)\]provides per\-frame depth but produces scale as a prior rather than a measurement, and the scale varies across frames\. Stereo methods recover metric depth from calibrated baselines\. Classical approaches such as block matching and semi\-global matching require rectification, which distorts wide\-angle fisheye views at the image periphery where egocentric hands appear\. Recent learned stereo methods such as FoundationStereo\[[28](https://arxiv.org/html/2609.16684#bib.bib30)\]operate on unrectified views and generalise across camera models, making them suitable for the high\-vergence pairs formed by a fisheye camera and a lateral camera in a head\-mounted rig\. Multi\-view geometry frameworks such as COLMAP\[[19](https://arxiv.org/html/2609.16684#bib.bib15)\]and generalised camera models\[[30](https://arxiv.org/html/2609.16684#bib.bib21),[31](https://arxiv.org/html/2609.16684#bib.bib16)\]treat rigidly mounted cameras as one sensor with known relative poses, reducing drift and enabling scale recovery from a fixed baseline\. We adopt this constraint and additionally use the calibrated baseline as the gauge that fixes metric scale across the entire reconstruction\.
Motion Tracking\.Head\-worn capture is rotation\-dominated, often with negligible translation over a keyframe interval, and wide\-angle; the useful sort of this literature is by what supplies metric scale\. ORB\-SLAM3 makes scale observable from acceleration\[[18](https://arxiv.org/html/2609.16684#bib.bib18)\]and is the online VIO pipeline we originally builtMEgo View’s head\-pose stage around, so it is at once a precursor of our system and our comparison point for the offline path — refined offline with a full\-sequence bundle adjustment and per\-camera\-frame re\-registration for that comparison, rather than scored as its raw real\-time trajectory\. Offline pipelines reconstruct more accurately than an online tracker can\[[19](https://arxiv.org/html/2609.16684#bib.bib15),[32](https://arxiv.org/html/2609.16684#bib.bib17),[33](https://arxiv.org/html/2609.16684#bib.bib19)\]but cannot observe scale monocularly, and learned\-prior loops\[[34](https://arxiv.org/html/2609.16684#bib.bib26),[35](https://arxiv.org/html/2609.16684#bib.bib25)\]could in principle carry a fixed\-scale edge as ours does — DROID\-SLAM is our second baseline in[Table 2](https://arxiv.org/html/2609.16684#S4.T2)\. Treating rigidly mounted cameras as one generalised camera with known relative poses\[[31](https://arxiv.org/html/2609.16684#bib.bib16),[30](https://arxiv.org/html/2609.16684#bib.bib21)\]under a fisheye projection\[[36](https://arxiv.org/html/2609.16684#bib.bib20)\]is our head stage’s machinery; the literature adopts that constraint to reduce drift, and we additionally use the calibrated baseline as the*gauge*\.
## 3MEgoVista
### 3\.1Hardware
All data is collected usingMEgo View111[https://www\.maniformer\.ai/en/mego](https://www.maniformer.ai/en/mego), a custom\-designed head\-mounted device for acquiring human manipulation behaviour in the wild\.MEgo Viewcarries fisheye cameras\[[36](https://arxiv.org/html/2609.16684#bib.bib20)\]whose wide field of view is analogous to human binocular vision, three of which are used below: a central camera and a calibrated stereo pair\. The high resolution and 60 fps frame rate facilitate fine\-grained and agile hand motion tracking, while the integrated Inertial Measurement Unit \(IMU\), sampled at 500 Hz, enables improved accuracy in camera pose estimation\.
Figure 1:TheMEgoVistapipeline\.A raw recording enters at the left\.MEgoVista— generate the per\-frame camera pose\.HandStage— hand detection, MANO regression, multi\-view depth estimation, full\-batch bundle adjustment\. Outputs at the right: head trajectory and hand trajectory\.
### 3\.2Hand Reconstruction
A\. Ego\-Aware Hand Detection and Initialization\.In\-the\-wild egocentric videos frequently contain multiple visible hands from both the camera wearer and surrounding people\. We therefore employ a trained end\-to\-end detector that jointly estimates hand location, ownership, and handedness, allowing the wearer’s left and right hands to be directly distinguished from other visible hands\. A lightweight tracker further maintains temporally stable hand identities and consistent bounding boxes across the sequence\. The tracked wearer\-hand regions are then processed by models akin to WiLoR\[[13](https://arxiv.org/html/2609.16684#bib.bib27)\]and HaMeR\[[14](https://arxiv.org/html/2609.16684#bib.bib8)\]to obtain per\-frame MANO\[[7](https://arxiv.org/html/2609.16684#bib.bib7)\]pose and shape estimates in the camera coordinate system, which serve as initialization for the subsequent metric reconstruction and sequence\-level optimization\.
B\. Tri\-Fisheye Virtual\-Plane Metric Depth Estimation\.To address the challenges posed by fisheye cameras on head\-mounted devices—namely their large field of view, wide baselines, and mounting pose deviations—this paper proposes a stereo depth estimation and fusion method based on a triple\-fisheye camera system\. The central camera \(MID\) on the device forms two wide\-baseline stereo pairs, MID×\\timesLEFT and MID×\\timesRIGHT, with the two side cameras \(LEFT and RIGHT\)\. For each pair, we design a wide\-baseline common\-rotation rectification scheme: the rectified coordinate frame takes the baseline direction as itse1e\_\{1\}axis to absorb arbitrary three\-axis mounting offsets between the cameras, and the angle bisector of the two optical axes as itse3e\_\{3\}axis, so that the common field of view is centered in the rectified images\. This design guarantees row alignment while maximally preserving the effective field of view\. The two rectified stereo pairs are then processed by the FoundationStereo\[[28](https://arxiv.org/html/2609.16684#bib.bib30)\]network, each producing a dense depth map together with a per\-pixel confidence measure derived from the matching probability volume\. The two depth maps are arbitrated and fused within the common field of view of the MID view\. Finally, the fused depth is back\-projected and analytically reprojected onto the MID fisheye pixel grid via the KB fisheye model, yielding a dense depth map that is pixel\-aligned with the original fisheye image\.
C\. Full\-Sequence Hand Motion Optimization\.A per\-frame regressor produces hands that may be individually plausible but are not necessarily consistent over time: they can jitter, drift in depth, or pass through configurations that a real hand cannot reach\. We therefore optimize a single MANO trajectory over the complete episode, subject to three complementary requirements: consistency with the observations, continuity of the 3D motion, and anatomical plausibility\.
Before optimization, we select reliable and informative observations to prevent erroneous per\-frame estimates from contaminating the full\-sequence trajectory\. Each frame is evaluated according to detection confidence, depth validity and dispersion, joint spread, bone\-ratio consistency, and left–right agreement\. Observations that fail these checks are excluded from the corresponding optimization terms, while the remaining high\-quality observations are used to constrain the full\-sequence optimization\.
The objective is organized accordingly:
1. 1\.Tight to the observations\.ℒreproj\\mathcal\{L\}\_\{\\mathrm\{reproj\}\}is the 2D keypoint residual, normalized by hand bounding\-box area so a distant hand retains a meaningful contribution rather than being dominated by a nearby one;ℒsdepth\\mathcal\{L\}\_\{\\mathrm\{sdepth\}\}constrains wrist camera\-frame depth using the metric depth estimate from TriVP, providing the depth constraint that reprojection alone cannot reliably recover; andℒpose\\mathcal\{L\}\_\{\\mathrm\{pose\}\}anchors finger pose to the per\-frame regressor\.
2. 2\.Continuous in 3D\.Second\-order smoothnessℒsmooth\\mathcal\{L\}\_\{\\mathrm\{smooth\}\}acts on the root trajectory, while Huber termsℒwrootvel\\mathcal\{L\}\_\{\\mathrm\{wrootvel\}\}andℒwrootang\\mathcal\{L\}\_\{\\mathrm\{wrootang\}\}suppress translational and rotational drift without allowing a single large motion to dominate the residual\.ℒfingervel\\mathcal\{L\}\_\{\\mathrm\{fingervel\}\}further suppresses single\-frame articulation jitter\.
3. 3\.Anatomically correct\.DifferentiableC1C^\{1\}relu2barriers approximate, as soft penalties on a standard MANO solve, the anatomical constraints modeled explicitly by MS\-MANO\[[21](https://arxiv.org/html/2609.16684#bib.bib12)\]and BioPR\[[23](https://arxiv.org/html/2609.16684#bib.bib32)\]\. Specifically,ℒrom\\mathcal\{L\}\_\{\\mathrm\{rom\}\}constrains the range of motion of individual joints,ℒbend\\mathcal\{L\}\_\{\\mathrm\{bend\}\}enforces bending limits in output space rather than parameter space, andℒω\\mathcal\{L\}\_\{\\omega\}limits angular velocity according to physiologically plausible bounds\.
Both hands are jointly optimized over the complete episode rather than independently frame by frame, yielding temporally coherent MANO trajectories in the reconstructed metric world frame\.
### 3\.3Motion Tracking
MEgoVistarecovers a head pose at very nearly every frame, and with it the metric, gravity\-aligned world frame the hand stage is solved in \([Section 3\.2](https://arxiv.org/html/2609.16684#S3.SS2)\)\. A head pivots far more than it travels, and no scene carries the metre or the vertical, so each quantity comes from the instrument that determines it\.
A\. Rig\-Constrained Incremental Reconstruction\.Translation this weak is poor evidence about the camera model: a bundle adjustment free to move it trades a trusted calibration for lower reprojection error\. The array is therefore one generalised camera\[[31](https://arxiv.org/html/2609.16684#bib.bib16),[30](https://arxiv.org/html/2609.16684#bib.bib21)\], images sharing a decode index forming one rig frame with a single head pose, all three fisheye under one Kannala–Brandt model\[[36](https://arxiv.org/html/2609.16684#bib.bib20)\]whose calibration stays frozen\. Selection divides the labour: inertia finds motion cheaply, so gyroscope rotation and accelerometer energy nominate keyframes and no fast turn is missed, while only vision knows whether a candidate is informative, so Lucas–Kanade tracking against the last accepted keyframe rejects the poorly tracked and the static\. Visibility is declared from the same rig layout, not left to a sequential matcher blind to time and camera arrays: keyframes match forward within each camera, the three at every shared timestamp, and non\-keyframes match adjacent mapped keyframes for later PnP\. Mapping and localisation are likewise split, a map dense enough for every frame being too large to bundle\-adjust: keyframes alone enter the incremental mapper\[[19](https://arxiv.org/html/2609.16684#bib.bib15)\], whose finished map is frozen and localised against by pose\-only PnP\.
B\. Calibrated\-Stereo Metric Initialisation\.Scale can only be imposed at initialisation, and neither obvious seed carries it: the automatic choice is a near\-pure\-rotation pair whose scale can be wrong by orders of magnitude, and same\-instant left and right images are one rig frame, degenerate in generalised relative pose\. The seed is therefore constructed: left–right matches at a key timestamp, triangulated by DLT on the calibrated extrinsics and refined by Levenberg–Marquardt, give one registered rig frame at the true baseline to continue from\. Scale thus comes from calibration rather than alignment — though only in the seed, leaving what grows from it free to drift\.
C\. Post\-Hoc Inertial Gravity Alignment\.A visual reconstruction has no preferred vertical, and supplying one by joint visual\-inertial bundle adjustment would let inertial drift reach the geometry\. Inertia therefore stays outside the optimisation, read only after vision is complete: gravity is propagated through the 500 Hz inertial stream, accelerometer contributions down\-weighted by\|∥a∥−g\|\|\\lVert a\\rVert\-g\|so that violent motion is trusted less, smoothed by an RTS pass, and applied as one global rotation into the initial reading’s gravity\-aligned frame\.
## 4Experiments
The order of what follows is the argument\.[Section 4\.1](https://arxiv.org/html/2609.16684#S4.SS1)states the setup and the protocol in full;[Section 4\.2](https://arxiv.org/html/2609.16684#S4.SS2)then measures how good the reference itself is, before anything is measured against it;[Section 4\.3](https://arxiv.org/html/2609.16684#S4.SS3)and[Section 4\.4](https://arxiv.org/html/2609.16684#S4.SS4)score the two delivered quantities against that reference\.
### 4\.1Experimental Setup
Dataset\.Every episode was recorded withMEgo Viewat 60 fps inside a 26\-camera, 120 Hz Chingmu motion\-capture volume, with the rig worn normally, uninstrumented, and PTP\-synced to the reference\. The reference gives 6\-DoF head pose from a rigid marker cluster on the headset shell and a hand skeleton from a 15\-marker\-per\-hand vendor solve\.
Metrics and alignment conventions\.Because alignment is part of a metric’s definition, not an afterthought,[Table 1](https://arxiv.org/html/2609.16684#S4.T1)gives each metric with its convention: every head figure here uses hand\-eye alignment with scale fixed at 1, so none is comparable to an ATE fromevo’s default similarity fit\.
Table 1:Metrics, each with the alignment that defines it\.Jitter and violation rate, are reference\-free and unaligned\. Fitted scale is reported beside every trajectory figure but is not an error metric\.*Bold*marks the primary hand metric, not a best result\.
### 4\.2Ground\-Truth Construction and Validation
Spatial calibration puts both systems in one metric frame\.From paired relative motionsAiA\_\{i\}andBiB\_\{i\}of the rig camera and the head\-mounted marker cluster, we solveAiXhc=XhcBiA\_\{i\}X\_\{\\mathrm\{hc\}\}=X\_\{\\mathrm\{hc\}\}B\_\{i\}for the constant cluster\-to\-camera transformXhcX\_\{\\mathrm\{hc\}\}, with intrinsics fixed\. At timett, the measured cluster pose andXhcX\_\{\\mathrm\{hc\}\}carry each motion\-capture point into the camera frame\. Calibration captures are disjoint from validation and evaluation, soXhcX\_\{\\mathrm\{hc\}\}is never fitted to the hand predictions it is used to score\.
\(a\)Held\-out board poses\.
\(b\)Hand points before \(red\) and after \(cyan\) calibration\.
Figure 2:Validation of the motion\-capture\-to\-camera transform\.Left: projected corners on held\-out static and moving board frames; annotations are per\-frame means\. Right: projected hand points before \(red\) and after \(cyan\) calibration\.Held\-out reprojection validates the chain end to end\.The board is tracked by motion capture and imaged simultaneously by the rig, then transformed and projected exactly as the hand reference will be\. Representative frames in[Figure 2](https://arxiv.org/html/2609.16684#S4.F2)show mean errors of 0\.62–1\.25 px when static and 0\.83–2\.51 px in motion\. Static frames test geometry; moving frames additionally exercise PTP timestamp correspondence\. These are illustrative single\-frame means, not sequence\-level statistics\. The hand overlay confirms that the calibration transfers to the task domain\.
The remaining limitation is the solved hand reference\.Board reprojection validates the rigid transform, but not the vendor’s marker\-to\-skeleton solver\. The shared screen of[Section 4\.1](https://arxiv.org/html/2609.16684#S4.SS1)therefore rejects observed failure signatures: frozen fingers below 0\.02 mm/frame relative to the wrist \(0\.2–1\.4 normally\), solver spikes above6×6\\timesbaseline, and fingertip–marker distances above 40 mm \(10–20 normally\)\. It catches discrete failures but not slow solver drift; any surviving reference error remains charged to the evaluated pipeline\.
### 4\.3Motion\-Tracking Accuracy
As shown in[Table 2](https://arxiv.org/html/2609.16684#S4.T2), ORB\-SLAM3 andMEgoVistatrack the Chingmu reference to within a few millimetres and approximately0\.2∘0\.2^\{\\circ\}across both capture sessions, without tracking loss\. Neither method consistently dominates:MEgoVistais more accurate on Session 2 \(0\.65 versus 0\.92 mm steady\-state\), whereas ORB\-SLAM3 is more accurate on Session 1 \(2\.12 versus 2\.83 mm\)\. Removing the first 5 s of each episode reduces ORB\-SLAM3’s error by 1\.15 mm, compared with only 0\.04 mm forMEgoVista\. This difference reflects the inertial initialisation required by ORB\-SLAM3 but absent from our offline reconstruction\. It motivates the use of the offline stage in production, where episodes last approximately 30 s and a fixed initialisation transient occupies a substantially larger fraction of each sequence than of the longest evaluated clip \(104\.3 s\)\.
On Session 1, DROID\-W\[[37](https://arxiv.org/html/2609.16684#bib.bib37)\]and MASt3R\-SLAM\[[35](https://arxiv.org/html/2609.16684#bib.bib25)\]are substantially less robust and exhibit different failure modes\. DROID\-W suffers severe scale failure on 4 of 22 episodes, with recovered scales of 0\.80–2\.68×\\timesand a mean SE\(3\) ATE of 395 mm\. On the remaining 18 episodes, it achieves a mean hand\-eye ATE of 26\.18 mm at a mean fitted scale of 0\.86×\\times\. MASt3R\-SLAM loses tracking on one episode and underestimates scale on all remaining 21 \(0\.23–0\.66×\\times\), yielding a mean hand\-eye ATE of 121\.59 mm\. ForMEgoVista, the mean scale deviation\|s−1\|\|s\-1\|is 0\.0120 on Session 1 and 0\.0048 on Session 2, but fitted scale does not consistently predict episode\-level accuracy \(r=\+0\.30r=\+0\.30andr=−0\.44r=\-0\.44, respectively\)\.
On episodes of comparable duration \(approximately 40 s\),MEgoVistarequires an average of 206\.0 s per episode, compared with 627\.5 s for DROID\-W and 715\.4 s for MASt3R\-SLAM\. Under the same evaluation setting, these measurements correspond to approximately3\.0×3\.0\\timesand3\.5×3\.5\\timesreductions in per\-episode wall\-clock time, respectively\.
Table 2:Head\-trajectory accuracy against Chingmu 26\-camera 120 Hz optical motion capture\.Per\-episode RMSE, averaged over episodes, hand\-eye aligned per capture day \([Table 1](https://arxiv.org/html/2609.16684#S4.T1)\)\. ORB\-SLAM3 rows score an offline\-refined trajectory, not the raw online output \([Section 4\.1](https://arxiv.org/html/2609.16684#S4.SS1)\)\.*Steady*discards the first 5 s \(not computed for DROID\-W/MASt3R\-SLAM, Session 1 only\)\. DROID\-W/MASt3R\-SLAM rows exclude failure episodes; see[Section 4\.3](https://arxiv.org/html/2609.16684#S4.SS3)\.*Fitted scale*is not an error metric\.Bold: lowest ATE per group\.
### 4\.4Hand\-Reconstruction Accuracy
As shown in Table[3](https://arxiv.org/html/2609.16684#S4.T3), EgoVista outperforms all open\-source baselines on every metric\.
For*detection*, EgoVista attains perfect Precision, Recall, and F1 \(1\.00\), producing neither missed nor hallucinated hands across the entire evaluation set; in contrast, Dyn\-HaMR hallucinates frequently under occlusion \(Precision 0\.73 despite full recall\), and HaWoR still misses or spuriously detects hands occasionally \(F1 0\.98\)\. Under the coverage\-aware protocol, where missed detections incur a deterministic placeholder error, complete detection coverage also directly benefits the accuracy metrics below\.
For*3D pose*, EgoVista reduces PA\-MPJPE\-p to 4\.29 mm, a 70% improvement over the strongest baseline HaWoR \(16\.30 mm\) and an order\-of\-magnitude reduction relative to Dyn\-HaMR \(49\.18 mm\)\.
The margin is largest on*orientation and position*: EPE\-P drops from 83\.10 px \(HaWoR\) to 9\.90 px \(8\.4×8\.4\\times\), and CT\-p from 46\.70 to 14\.30 \(69% reduction\), indicating that our method recovers the absolute hand position far more accurately, which we attribute to the explicit geometric depth constraints provided by our multi\-view depth estimation\.
For*temporal smoothness*, EgoVista attains a jitter of 1\.11 mm/frame2,3\.2×3\.2\\timeslower than the smoothest baseline \(Dyn\-HaMR, 3\.50\) and an order of magnitude below HaWoR \(16\.71\), without any test\-time optimization\. Notably, HaPTIC fails entirely in our multi\-person capture scenes—a practically significant limitation, as bystanders are common in real\-world egocentric deployment\. Overall, these results demonstrate that EgoVista delivers substantially more complete, accurate, and stable hand reconstruction than existing open\-source methods\.
Table 3:Comparison with open\-source egocentric hand reconstruction methods against motion\-capture ground truth\.All methods are evaluated on identical segments of our motion\-capture dataset\. All baselines are re\-run and rescored on our data\. HaPTIC fails to produce valid output in our multi\-person capture scenes\.Boldmarks the best result in each column\.Dyn\-HaMRHaWoREgoVista \(Ours\)Figure 3:Qualitative comparison on in\-the\-wild recordings\.*Top*: rapid hand motion;*middle*: interaction with a bystander’s hand;*bottom*: soiled hands under fast motion\. Baselines exhibit pose offsets \(top, bottom\) or misattribute the bystander’s hand to the wearer \(middle\), while EgoVista maintains tight alignment and correct hand identity throughout\.
### 4\.5Qualitative results in the wild
We further evaluate EgoVista on real\-world recordings for which no motion\-capture ground truth is available, and assess the results qualitatively by the alignment between the projected hand skeletons and the observed hands \(Fig\.[3](https://arxiv.org/html/2609.16684#S4.F3)\)\. Three representative challenging scenarios are shown\.
\(i\)*Rapid hand motion*\(milk\-tea preparation\): both EgoVista and Dyn\-HaMR produce skeletons that align well with the hands, whereas HaWoR exhibits a noticeable offset on the left hand; this is consistent with its order\-of\-magnitude higher jitter in Table[3](https://arxiv.org/html/2609.16684#S4.T3)\(16\.71 vs\. 1\.11\), as frame\-wise instability manifests directly as misalignment under fast motion\.
\(ii\)*Interaction with another person’s hands*\(manicure\): when the wearer’s left hand approaches and interacts with a bystander’s left hand, both Dyn\-HaMR and HaWoR misattribute the bystander’s hand to the wearer, while EgoVista correctly disambiguates identities and reconstructs only the wearer’s hands\. This failure mode explains the low detection precision of Dyn\-HaMR \(0\.73\) in Table[3](https://arxiv.org/html/2609.16684#S4.T3)and echoes the complete failure of HaPTIC in multi\-person scenes: distinguishing the wearer’s hands from bystanders’ hands remains an open challenge for existing egocentric methods\.
\(iii\)*Rapid motion with soiled hands*\(animal washing\): although both EgoVista and Dyn\-HaMR successfully detect the dirt\-covered hands, HaWoR shows a slight offset on the left hand with an inaccurate wrist position, in line with its higher CT\-p error \(46\.70\) in the quantitative comparison\.
Overall, the qualitative results corroborate our quantitative findings: EgoVista achieves more complete detection, tighter image\-plane alignment, and robust hand identity disambiguation, with the largest margins under fast motion and multi\-person interference\.
## 5Conclusion
MEgoVistatakes an unprepared egocentric recording and returns metric two\-hand motion and a metric head trajectory in one world frame, and — unusually for this class of system — it has been held to a measured standard: both delivered quantities were scored against 26\-camera optical motion capture under a protocol that charges missed detections and audits its own reference rather than assuming it\. Head trajectory lands at 0\.65 to 2\.83 mm, finger articulation near 5 mm PA\-MPJPE\-p— the last being the figure a flattering summary would omit and a consumer of the labels most needs\. What transfers beyond this particular system is narrower than the system itself: taking metric gauge from calibration at initialisation, rather than recovering it by alignment afterwards, is what makes such a measurement meaningful at all, and the quantity that exposes its failure is fitted scale rather than aligned error\. Whether the next millimetre requires a better estimator or a better reference is genuinely undecided, and the board capture of[Section 4\.2](https://arxiv.org/html/2609.16684#S4.SS2)is the experiment that would decide it\.
## References
- \[1\]\(2026\)Data pyramid for embodied manipulation: a survey\.Note:arXiv preprint arXiv:2607\.24744Cited by:[§1](https://arxiv.org/html/2609.16684#S1.p1.1)\.
- \[2\]K\. Graumanet al\.\(2022\)Ego4D: around the world in 3,000 hours of egocentric video\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 18973–18990\.External Links:[Document](https://dx.doi.org/10.1109/CVPR52688.2022.01842)Cited by:[§1](https://arxiv.org/html/2609.16684#S1.p1.1),[§1](https://arxiv.org/html/2609.16684#S1.p2.1)\.
- \[3\]D\. Damen, H\. Doughty, G\. M\. Farinella, A\. Furnari, E\. Kazakos, J\. Ma, D\. Moltisanti, J\. Munro, T\. Perrett, W\. Price, and M\. Wray\(2022\)Rescaling egocentric vision: collection, pipeline and challenges for EPIC\-KITCHENS\-100\.International Journal of Computer Vision130\(1\),pp\. 33–55\.External Links:[Document](https://dx.doi.org/10.1007/s11263-021-01531-2)Cited by:[§1](https://arxiv.org/html/2609.16684#S1.p1.1)\.
- \[4\]K\. Graumanet al\.\(2024\)Ego\-Exo4D: understanding skilled human activity from first\- and third\-person perspectives\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 19383–19400\.External Links:[Document](https://dx.doi.org/10.1109/CVPR52733.2024.01834)Cited by:[§1](https://arxiv.org/html/2609.16684#S1.p1.1)\.
- \[5\]R\. Hoque, P\. Huang, D\. J\. Yoon, M\. Sivapurapu, and J\. Zhang\(2026\)EgoDex: learning dexterous manipulation from large\-scale egocentric video\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§1](https://arxiv.org/html/2609.16684#S1.p1.1),[§1](https://arxiv.org/html/2609.16684#S1.p2.1)\.
- \[6\]B\. AI\(2025\)Egocentric\-10k\.HuggingFace dataset card\.External Links:[Link](https://huggingface.co/datasets/builddotai/Egocentric-10K)Cited by:[§1](https://arxiv.org/html/2609.16684#S1.p1.1)\.
- \[7\]J\. Romero, D\. Tzionas, and M\. J\. Black\(2017\)Embodied hands: modeling and capturing hands and bodies together\.ACM Transactions on Graphics36\(6\),pp\. 245:1–245:17\.External Links:[Document](https://dx.doi.org/10.1145/3130800.3130883)Cited by:[§1](https://arxiv.org/html/2609.16684#S1.p1.1),[§2](https://arxiv.org/html/2609.16684#S2.p1.1),[§3\.2](https://arxiv.org/html/2609.16684#S3.SS2.p1.1)\.
- \[8\]R\. Qiu, S\. Yang, X\. Cheng, C\. Chawla, J\. Li, T\. He, G\. Yan, D\. J\. Yoon, R\. Hoque, L\. Paulsen, G\. Yang, J\. Zhang, S\. Yi, G\. Shi, and X\. Wang\(2025\)Humanoid policy human policy\.External Links:2503\.13441,[Link](https://arxiv.org/abs/2503.13441)Cited by:[§1](https://arxiv.org/html/2609.16684#S1.p1.1)\.
- \[9\]R\. Zheng, D\. Niu, Y\. Xie, J\. Wang, M\. Xu, Y\. Jiang, F\. Castañeda, F\. Hu, Y\. L\. Tan, L\. Fu, T\. Darrell, F\. Huang, Y\. Zhu, D\. Xu, and L\. Fan\(2026\)EgoScale: scaling dexterous manipulation with diverse egocentric human data\.External Links:2602\.16710,[Link](https://arxiv.org/abs/2602.16710)Cited by:[§1](https://arxiv.org/html/2609.16684#S1.p1.1)\.
- \[10\]Q\. Li, Y\. Deng, Y\. Liang, L\. Luo, L\. Zhou, C\. Yao, L\. Zeng, Z\. Feng, H\. Liang, S\. Xu, Y\. Zhang, X\. Chen, H\. Chen, L\. Sun, D\. Chen, J\. Yang, and B\. Guo\(2025\)Scalable vision\-language\-action model pretraining for robotic manipulation with real\-life human activity videos\.External Links:2510\.21571,[Link](https://arxiv.org/abs/2510.21571)Cited by:[§1](https://arxiv.org/html/2609.16684#S1.p1.1)\.
- \[11\]Z\. Fan, O\. Taheri, D\. Tzionas, M\. Kocabas, M\. Kaufmann, M\. J\. Black, and O\. Hilliges\(2023\)ARCTIC: a dataset for dexterous bimanual hand\-object manipulation\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 12943–12954\.External Links:[Document](https://dx.doi.org/10.1109/CVPR52729.2023.01244)Cited by:[§1](https://arxiv.org/html/2609.16684#S1.p2.1)\.
- \[12\]F\. Sener, D\. Chatterjee, D\. Shelepov, K\. He, D\. Singhania, R\. Wang,et al\.\(2022\)Assembly101: a large\-scale multi\-view video dataset for understanding procedural activities\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 21064–21074\.External Links:[Document](https://dx.doi.org/10.1109/CVPR52688.2022.02042)Cited by:[§1](https://arxiv.org/html/2609.16684#S1.p2.1)\.
- \[13\]R\. A\. Potamias, J\. Zhang, J\. Deng, and S\. Zafeiriou\(2025\)WiLoR: end\-to\-end 3D hand localization and reconstruction in\-the\-wild\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 12242–12254\.External Links:[Document](https://dx.doi.org/10.1109/CVPR52734.2025.01143)Cited by:[§1](https://arxiv.org/html/2609.16684#S1.p3.1),[§2](https://arxiv.org/html/2609.16684#S2.p1.1),[§3\.2](https://arxiv.org/html/2609.16684#S3.SS2.p1.1)\.
- \[14\]G\. Pavlakos, D\. Shan, I\. Radosavovic, A\. Kanazawa, D\. Fouhey, and J\. Malik\(2024\)Reconstructing hands in 3D with transformers\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 9826–9836\.External Links:[Document](https://dx.doi.org/10.1109/CVPR52733.2024.00938)Cited by:[§1](https://arxiv.org/html/2609.16684#S1.p3.1),[§2](https://arxiv.org/html/2609.16684#S2.p1.1),[§3\.2](https://arxiv.org/html/2609.16684#S3.SS2.p1.1)\.
- \[15\]J\. Zhang, J\. Deng, C\. Ma, and R\. A\. Potamias\(2025\)HaWoR: world\-space hand motion reconstruction from egocentric videos\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 1805–1815\.External Links:[Document](https://dx.doi.org/10.1109/CVPR52734.2025.00175)Cited by:[§1](https://arxiv.org/html/2609.16684#S1.p3.1),[§2](https://arxiv.org/html/2609.16684#S2.p1.1),[Table 3](https://arxiv.org/html/2609.16684#S4.T3.7.1.1.1.1.1.1.4.1)\.
- \[16\]Z\. Yu, S\. Zafeiriou, and T\. Birdal\(2025\)Dyn\-HaMR: recovering 4D interacting hand motion from a dynamic camera\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 27716–27726\.External Links:[Document](https://dx.doi.org/10.1109/CVPR52734.2025.02581)Cited by:[§1](https://arxiv.org/html/2609.16684#S1.p3.1),[§2](https://arxiv.org/html/2609.16684#S2.p1.1),[Table 3](https://arxiv.org/html/2609.16684#S4.T3.7.1.1.1.1.1.1.3.1)\.
- \[17\]L\. Piccinelli, C\. Sakaridis, Y\. Yang, M\. Segù, S\. Li, W\. Abbeloos, and L\. Van Gool\(2026\)UniDepthV2: universal monocular metric depth estimation made simpler\.IEEE Transactions on Pattern Analysis and Machine Intelligence48,pp\. 2354–2367\.External Links:[Document](https://dx.doi.org/10.1109/TPAMI.2025.3628473)Cited by:[§1](https://arxiv.org/html/2609.16684#S1.p3.1),[§2](https://arxiv.org/html/2609.16684#S2.p1.1),[§2](https://arxiv.org/html/2609.16684#S2.p2.1)\.
- \[18\]C\. Campos, R\. Elvira, J\. J\. Gómez Rodríguez, J\. M\. M\. Montiel, and J\. D\. Tardós\(2021\)ORB\-SLAM3: an accurate open\-source library for visual, visual\-inertial, and multimap SLAM\.IEEE Transactions on Robotics37\(6\),pp\. 1874–1890\.External Links:[Document](https://dx.doi.org/10.1109/TRO.2021.3075644)Cited by:[§1](https://arxiv.org/html/2609.16684#S1.p4.1),[§2](https://arxiv.org/html/2609.16684#S2.p3.1)\.
- \[19\]J\. L\. Schönberger and J\. Frahm\(2016\)Structure\-from\-motion revisited\.InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 4104–4113\.External Links:[Document](https://dx.doi.org/10.1109/CVPR.2016.445)Cited by:[§1](https://arxiv.org/html/2609.16684#S1.p4.1),[§2](https://arxiv.org/html/2609.16684#S2.p2.1),[§2](https://arxiv.org/html/2609.16684#S2.p3.1),[§3\.3](https://arxiv.org/html/2609.16684#S3.SS3.p2.1)\.
- \[20\]Y\. Ye, Y\. Feng, O\. Taheri, H\. Feng, S\. Tulsiani, and M\. J\. Black\(2025\)Predicting 4D hand trajectory from monocular videos\.Note:arXiv preprint arXiv:2501\.08329Cited by:[§2](https://arxiv.org/html/2609.16684#S2.p1.1),[Table 3](https://arxiv.org/html/2609.16684#S4.T3.7.1.1.1.1.1.1.5.1)\.
- \[21\]P\. Xie, W\. Xu, T\. Tang, Z\. Yu, and C\. Lu\(2024\)MS\-MANO: enabling hand pose tracking with biomechanical constraints\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 2382–2392\.External Links:[Document](https://dx.doi.org/10.1109/CVPR52733.2024.00231)Cited by:[§2](https://arxiv.org/html/2609.16684#S2.p1.1),[item 3](https://arxiv.org/html/2609.16684#S3.I1.i3.p1.1)\.
- \[22\]S\. Shimada, V\. Golyanik, W\. Xu, and C\. Theobalt\(2020\)PhysCap: physically plausible monocular 3D motion capture in real time\.ACM Transactions on Graphics39\(6\),pp\. 235:1–235:16\.External Links:[Document](https://dx.doi.org/10.1145/3414685.3417877)Cited by:[§2](https://arxiv.org/html/2609.16684#S2.p1.1)\.
- \[23\]T\. Duet al\.\(2023\)BioPR: biomechanically plausible hand pose regression\.InProceedings of the IEEE/CVF International Conference on Computer Vision \(ICCV\),Cited by:[§2](https://arxiv.org/html/2609.16684#S2.p1.1),[item 3](https://arxiv.org/html/2609.16684#S3.I1.i3.p1.1)\.
- \[24\]S\. Wang, V\. Leroy, Y\. Cabon, B\. Chidlovskii, and J\. Revaud\(2024\)DUSt3R: geometric 3D vision made easy\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 20697–20709\.External Links:[Document](https://dx.doi.org/10.1109/CVPR52733.2024.01956)Cited by:[§2](https://arxiv.org/html/2609.16684#S2.p1.1)\.
- \[25\]V\. Leroy, Y\. Cabon, and J\. Revaud\(2024\)Grounding image matching in 3D with MASt3R\.InProceedings of the European Conference on Computer Vision \(ECCV\),pp\. 71–91\.External Links:[Document](https://dx.doi.org/10.1007/978-3-031-73220-1%5F5)Cited by:[§2](https://arxiv.org/html/2609.16684#S2.p1.1)\.
- \[26\]J\. Wang, M\. Chen, N\. Karaev, A\. Vedaldi, C\. Rupprecht, and D\. Novotný\(2025\)VGGT: visual geometry grounded transformer\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 5294–5306\.External Links:[Document](https://dx.doi.org/10.1109/CVPR52734.2025.00499)Cited by:[§2](https://arxiv.org/html/2609.16684#S2.p1.1)\.
- \[27\]L\. Piccinelli, C\. Sakaridis, M\. Segù, Y\. Yang, S\. Li, W\. Abbeloos, and L\. Van Gool\(2025\)UniK3D: universal camera monocular 3D estimation\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 1028–1039\.External Links:[Document](https://dx.doi.org/10.1109/CVPR52734.2025.00104)Cited by:[§2](https://arxiv.org/html/2609.16684#S2.p1.1),[§2](https://arxiv.org/html/2609.16684#S2.p2.1)\.
- \[28\]B\. Wen, M\. Trepte, J\. Aribido, J\. Kautz, O\. Gallo, and S\. Birchfield\(2025\)FoundationStereo: zero\-shot stereo matching\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 5249–5260\.External Links:[Document](https://dx.doi.org/10.1109/CVPR52734.2025.00495)Cited by:[§2](https://arxiv.org/html/2609.16684#S2.p1.1),[§2](https://arxiv.org/html/2609.16684#S2.p2.1),[§3\.2](https://arxiv.org/html/2609.16684#S3.SS2.p2.1)\.
- \[29\]N\. Ravi, V\. Gabeur, Y\. Hu, R\. Hu, C\. Ryali, T\. Ma, H\. Khedr, R\. Rädle,et al\.\(2025\)SAM 2: segment anything in images and videos\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§2](https://arxiv.org/html/2609.16684#S2.p1.1)\.
- \[30\]R\. Pless\(2003\)Using many cameras as one\.InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 587–593\.External Links:[Document](https://dx.doi.org/10.1109/CVPR.2003.1211520)Cited by:[§2](https://arxiv.org/html/2609.16684#S2.p2.1),[§2](https://arxiv.org/html/2609.16684#S2.p3.1),[§3\.3](https://arxiv.org/html/2609.16684#S3.SS3.p2.1)\.
- \[31\]J\. L\. Schönberger\(2018\)Robust methods for accurate and efficient 3D modeling from unstructured imagery\.Ph\.D\. Thesis,ETH Zürich\.External Links:[Document](https://dx.doi.org/10.3929/ethz-b-000295763)Cited by:[§2](https://arxiv.org/html/2609.16684#S2.p2.1),[§2](https://arxiv.org/html/2609.16684#S2.p3.1),[§3\.3](https://arxiv.org/html/2609.16684#S3.SS3.p2.1)\.
- \[32\]L\. Pan, D\. Baráth, M\. Pollefeys, and J\. L\. Schönberger\(2024\)Global structure\-from\-motion revisited\.InProceedings of the European Conference on Computer Vision \(ECCV\),pp\. 58–77\.External Links:[Document](https://dx.doi.org/10.1007/978-3-031-73661-2%5F4)Cited by:[§2](https://arxiv.org/html/2609.16684#S2.p3.1)\.
- \[33\]Z\. Li, R\. Tucker, F\. Cole, Q\. Wang, L\. Jin, V\. Ye, A\. Kanazawa, A\. Holynski, and N\. Snavely\(2025\)MegaSaM: accurate, fast and robust structure and motion from casual dynamic videos\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 10486–10496\.External Links:[Document](https://dx.doi.org/10.1109/CVPR52734.2025.00981)Cited by:[§2](https://arxiv.org/html/2609.16684#S2.p3.1)\.
- \[34\]Z\. Teed and J\. Deng\(2021\)DROID\-SLAM: deep visual SLAM for monocular, stereo, and RGB\-D cameras\.InAdvances in Neural Information Processing Systems \(NeurIPS\),pp\. 16558–16569\.Cited by:[§2](https://arxiv.org/html/2609.16684#S2.p3.1)\.
- \[35\]R\. Murai, E\. Dexheimer, and A\. J\. Davison\(2025\)MASt3R\-SLAM: real\-time dense SLAM with 3D reconstruction priors\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 16695–16705\.External Links:[Document](https://dx.doi.org/10.1109/CVPR52734.2025.01556)Cited by:[§2](https://arxiv.org/html/2609.16684#S2.p3.1),[§4\.3](https://arxiv.org/html/2609.16684#S4.SS3.p2.1),[Table 2](https://arxiv.org/html/2609.16684#S4.T2.9.1.1.1.1.1.1.6.2)\.
- \[36\]J\. Kannala and S\. S\. Brandt\(2006\)A generic camera model and calibration method for conventional, wide\-angle, and fish\-eye lenses\.IEEE Transactions on Pattern Analysis and Machine Intelligence28\(8\),pp\. 1335–1340\.External Links:[Document](https://dx.doi.org/10.1109/TPAMI.2006.153)Cited by:[§2](https://arxiv.org/html/2609.16684#S2.p3.1),[§3\.1](https://arxiv.org/html/2609.16684#S3.SS1.p1.1),[§3\.3](https://arxiv.org/html/2609.16684#S3.SS3.p2.1)\.
- \[37\]M\. Li, Z\. Zhu, M\. Pollefeys, and D\. Barath\(2026\)DROID\-slam in the wild\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),Cited by:[§4\.3](https://arxiv.org/html/2609.16684#S4.SS3.p2.1),[Table 2](https://arxiv.org/html/2609.16684#S4.T2.9.1.1.1.1.1.1.5.2)\.Similar Articles
@macrodata_labs: Everyone is betting on Egocentric data to scale robotics But turning that footage into training data requires recoverin…
Macrodata Labs releases a research blog on scaling robotics with egocentric video data by recovering 3D hand motion signals using only open-source models.
@oliviscusAI: this tool can track perfect 3D motion. rtmlib is a lightweight pose estimation library covering full body, hands, face,…
rtmlib is a lightweight, open-source pose estimation library that supports full-body, hand, face, and animal pose tracking, built on rtmpose and vitpose models, with a built-in Gradio web UI.
ActiveMimic: Egocentric Video Pretraining with Active Perception
ActiveMimic is a pretraining framework that recovers camera and wrist trajectories from egocentric human video to model active perception as a viewpoint action, enabling robot pretraining that matches the performance of models trained directly on robot data.
CosmoH2G: A Hand-to-Gripper Transfer Dataset and Baseline Method for Object Manipulation with Complex Spatial Movements
This paper presents CosmoH2G, a dataset and baseline method for transferring complex human hand demonstrations to robotic grippers using a two-stage framework that predicts sparse keyframes and continuous actions to handle intricate spatial movements.
EventEgoHands++: Event-based Egocentric 3D Hand Mesh Reconstruction with Real Dataset
This paper proposes EventEgoHands++, a framework for event-based 3D hand mesh reconstruction from an egocentric viewpoint, incorporating hand detection and adaptive attention, and introduces a new real-world dataset EEH-R for training and evaluation.