GeoVerse: World-Consistent Novel View Synthesis in Geometric Latent Space

Hugging Face Daily Papers Papers

Summary

GeoVerse is a framework for synthesizing world-consistent novel views by performing generation in a geometric latent space from a 3D foundation model and injecting appearance priors from a video generative model, demonstrating improved visual quality and geometric consistency over existing methods.

Novel view synthesis from sparse images must reconcile faithful reconstruction of observed regions with plausible completion of unseen content, while maintaining world consistency across viewpoints. Existing geometry-based methods preserve observed scene structure but often struggle to complete unseen regions, whereas video generative models offer rich appearance priors but accumulate inconsistencies during sequential view generation. We propose GeoVerse, a framework that synthesizes world-consistent novel views by performing generation within the geometric latent space of a pretrained 3D foundation model and injecting appearance priors from a video generative model. Specifically, GeoVerse extracts multilevel features from Wan2.2 VACE and injects them into the geometric latent diffusion model via a ControlNet-style adapter, incorporating video-learned appearance priors to enhance structural completion. To enforce cross-view coherence, a global spatial memory continuously aggregates observed and synthesized content, reprojecting target-aligned guidance to anchor subsequent predictions to a shared scene representation. Extensive experiments across diverse datasets demonstrate improved visual quality and geometric consistency, with a 2.23 dB higher PSNR on DL3DV and 32.4% lower ATE on Mip-NeRF360 compared to GLD.
Original Article
View Cached Full Text

Cached at: 09/29/26, 08:12 AM

Paper page - GeoVerse: World-Consistent Novel View Synthesis in Geometric Latent Space

Source: https://huggingface.co/papers/2609.35734

Abstract

Novelviewsynthesisfromsparseimagesmustreconcilefaithfulreconstructionofobservedregionswithplausiblecompletionofunseencontent,whilemaintainingworldconsistencyacrossviewpoints.Existinggeometry-basedmethodspreserveobservedscenestructurebutoftenstruggletocompleteunseenregions,whereasvideogenerativemodelsofferrichappearancepriorsbutaccumulateinconsistenciesduringsequentialviewgeneration.WeproposeGeoVerse,aframeworkthatsynthesizesworld-consistentnovelviewsbyperforminggenerationwithinthegeometriclatentspaceofapretrained3Dfoundationmodelandinjectingappearancepriorsfromavideogenerativemodel.Specifically,GeoVerseextractsmultilevelfeaturesfromWan2.2VACEandinjectsthemintothegeometriclatentdiffusionmodelviaaControlNet-styleadapter,incorporatingvideo-learnedappearancepriorstoenhancestructuralcompletion.Toenforcecross-viewcoherence,aglobalspatialmemorycontinuouslyaggregatesobservedandsynthesizedcontent,reprojectingtarget-alignedguidancetoanchorsubsequentpredictionstoasharedscenerepresentation.Extensiveexperimentsacrossdiversedatasetsdemonstrateimprovedvisualqualityandgeometricconsistency,witha2.23dBhigherPSNRonDL3DVand32.4%lowerATEonMip-NeRF360comparedtoGLD.

View arXiv pageView PDFProject pageGitHub2Add to collection

Get this paper in your agent:

hf papers read 2609\.35734

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.35734 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.35734 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.35734 in a Space README.md to link it from this page.

Collections including this paper1

Similar Articles

Towards Consistent Video Geometry Estimation

Hugging Face Daily Papers

ViGeo is a transformer-based foundation model that recovers dense and consistent 3D geometry from videos using dynamic chunking attention and a completion-based data refinement framework, achieving state-of-the-art performance across multiple tasks.