GraphVid: Interactive Graph-Controllable Video Generation
Summary
GraphVid introduces a graph-conditioned image-to-video generation model that enables interactive control through structured interaction graphs, outperforming prior methods with significant reductions in FID and FVD.
View Cached Full Text
Cached at: 07/24/26, 05:06 AM
Paper page - GraphVid: Interactive Graph-Controllable Video Generation
Source: https://huggingface.co/papers/2607.21580
Abstract
Controllablevideogenerationremainschallengingduetothedifficultyofspecifyingprecisemulti-objectinteractionsusingtextpromptsormotion-controlinputsthatprimarilyconstrainpixelmovement.Inpractice,trajectory-basedcontroloftenrequiresuserstodrawaccuratetracksformultipleobjects,whichscalespoorlywithscenecomplexityandbecomesambiguousunderocclusionoroverlap.Toenableflexibleyetprecisemulti-subjectcontrol,weintroduceGraphVid,agraph-conditionedimage-to-videogenerationmodelthatenablesinteractivecontrolthroughstructuredinteractiongraphs.WefurthercurateGraphVid-Bench,alarge-scaleinteraction-centricvideodatasetwithstructuredrelationalannotationstoenabletrainingofinteraction-awarevideogenerationmodels.Despiteusingsubstantiallylesstrainingdataandfewertrainableparametersthanpriormotion-controlmethods,GraphViddeliversstrongcontrollabilityandvideoquality.ComparedwithMotion-I2V,GraphVidreducesFIDbyupto39.9%andFVDby37.6%,whileimprovingPSNR(9.87=>15.98)andSSIM(0.38=>0.61).Ourresultshighlightthepotentialofstructuredsemanticinterfacesasapowerfulparadigmforcontrollablevideogeneration.
View arXiv pageView PDFProject pageAdd to collection
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2607.21580 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2607.21580 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2607.21580 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Vidu S1: A Real-Time Interactive Video Generation Model
Vidu S1 is a real-time interactive video generation model that enables voice-controlled digital character animation with infinite-length output and high frame rate on consumer GPUs, achieving state-of-the-art performance.
CogOmniControl: Reasoning-Driven Controllable Video Generation via Creative Intent Cognition
CogOmniControl is a reasoning-driven framework for controllable video generation that uses a specialized vision-language model (CogVLM) trained on anime production data to infer creative intent from sparse conditions, then guides a diffusion-based generator via reinforcement learning, achieving state-of-the-art results on new benchmarks.
ID-V2V: Identity-Preserving Video Restylization
ID-V2V is a video-to-video generative framework for identity-preserving video restylization. It treats identity preservation as a video relighting problem and uses edited keyframes for style propagation, achieving high-quality results without paired training data.
Towards Consistent Video Geometry Estimation
ViGeo is a transformer-based foundation model that recovers dense and consistent 3D geometry from videos using dynamic chunking attention and a completion-based data refinement framework, achieving state-of-the-art performance across multiple tasks.
ReImagine: Rethinking Controllable High-Quality Human Video Generation via Image-First Synthesis
ReImagine introduces an image-first approach to controllable high-quality human video generation, combining SMPL-X motion guidance with video diffusion models to decouple appearance from temporal consistency.