HarmoHOI: Harmonizing Appearance and 3D Motion for Multi-view Hand-Object Interaction Synthesis
Summary
HarmoHOI is a unified diffusion framework that jointly generates synchronized multi-view hand-object interaction videos and globally aligned 3D point tracks, achieving state-of-the-art performance in visual quality, motion plausibility, and multi-view geometric consistency.
View Cached Full Text
Cached at: 07/21/26, 10:36 AM
Paper page - HarmoHOI: Harmonizing Appearance and 3D Motion for Multi-view Hand-Object Interaction Synthesis
Source: https://huggingface.co/papers/2607.17097
Abstract
Hand-ObjectInteraction(HOI)synthesisisacornerstoneforanimationproductionandembodiedAI.Despitethestrongpriorsofvideofoundationmodels,multi-viewconsistentHOIsynthesisremainschallengingduetocomplexhandmotionsandocclusions.WepresentHarmoHOI,aunifieddiffusionframeworkthatjointlyandharmoniouslygeneratessynchronizedmulti-viewHOIvideosandgloballyaligned3Dpointtracks.Ourcoreinsightisthatrobustmulti-viewconsistencyfundamentallyrequiresgloballyaligned3Dgeometryandmotion.Tothisend,weproposeaMixtureofMulti-viewDiffusionTransformerthatco-modelsRGBvideosand3Dpointtracks.Byrepresentingpointtracksaspseudo-videos,wealign3Dgeometricsignalswiththe2Dlatentspaceoffoundationmodels,therebyminimizingthedomaingapandeasingadaptationofpriors.Tofurtherensuregeometryconsistency,weintroduceGlobalMotionAligningDiffusion,whichrefinescoarsepointtracksintometric-scale,globallyaligned3Dtrajectories.HarmoHOIenableson-the-flyco-evolutionof2Dappearanceand3Dmotionduringdenoising.Toovercomethescarcityofmulti-viewHOIdata,weemployahybriddatacurriculumlearningstrategythatsuccessfullytransfersgenericpriorsfromsingle-viewdatatosynchronizedmulti-viewgeneration.ExperimentalresultsshowthatHarmoHOIachievesstate-of-the-artperformanceinvisualquality,motionplausibility,andmulti-viewgeometricconsistency.Projectpageavailableathttps://droliven.github.io/HarmoHOI_project.
View arXiv pageView PDFProject pageGitHub0Add to collection
Get this paper in your agent:
hf papers read 2607\.17097
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2607.17097 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2607.17097 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2607.17097 in a Space README.md to link it from this page.
Collections including this paper1
Similar Articles
OneHOI: Unifying Human-Object Interaction Generation and Editing
OneHOI is a unified diffusion transformer framework that consolidates human-object interaction (HOI) generation and editing into a single conditional denoising process using relational modeling and structured attention mechanisms. The approach achieves state-of-the-art results across both HOI generation and editing tasks with support for multiple control modalities.
CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation
CoInteract introduces an end-to-end Diffusion Transformer framework that jointly models RGB appearance and HOI geometry to generate physically-plausible human-object interaction videos with stable hands/faces and zero inference overhead.
PhyGenHOI: Physically-Aware 4D Generation of Dynamic Human-Object Interactions
PhyGenHOI is a novel framework that generates physically accurate 4D human-object interactions by coupling motion diffusion models with material point method simulations using 3D Gaussian representations.
HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement
HOMIE is a framework for human-object centric video personalization that integrates MLLM features to improve subject fidelity and interaction patterns, achieving state-of-the-art performance.
The Surprising Effectiveness of Video Diffusion Models for Hand Motion Reconstruction
ViDiHand leverages pretrained video diffusion model representations to reconstruct 4D hand motion directly from egocentric video frames, outperforming existing methods on ARCTIC, HOT3D, and HOI4D without detectors or optimization.