GeoVerse: World-Consistent Novel View Synthesis in Geometric Latent Space
Summary
GeoVerse is a framework for synthesizing world-consistent novel views by performing generation in a geometric latent space from a 3D foundation model and injecting appearance priors from a video generative model, demonstrating improved visual quality and geometric consistency over existing methods.
View Cached Full Text
Cached at: 09/29/26, 08:12 AM
Paper page - GeoVerse: World-Consistent Novel View Synthesis in Geometric Latent Space
Source: https://huggingface.co/papers/2609.35734
Abstract
Novelviewsynthesisfromsparseimagesmustreconcilefaithfulreconstructionofobservedregionswithplausiblecompletionofunseencontent,whilemaintainingworldconsistencyacrossviewpoints.Existinggeometry-basedmethodspreserveobservedscenestructurebutoftenstruggletocompleteunseenregions,whereasvideogenerativemodelsofferrichappearancepriorsbutaccumulateinconsistenciesduringsequentialviewgeneration.WeproposeGeoVerse,aframeworkthatsynthesizesworld-consistentnovelviewsbyperforminggenerationwithinthegeometriclatentspaceofapretrained3Dfoundationmodelandinjectingappearancepriorsfromavideogenerativemodel.Specifically,GeoVerseextractsmultilevelfeaturesfromWan2.2VACEandinjectsthemintothegeometriclatentdiffusionmodelviaaControlNet-styleadapter,incorporatingvideo-learnedappearancepriorstoenhancestructuralcompletion.Toenforcecross-viewcoherence,aglobalspatialmemorycontinuouslyaggregatesobservedandsynthesizedcontent,reprojectingtarget-alignedguidancetoanchorsubsequentpredictionstoasharedscenerepresentation.Extensiveexperimentsacrossdiversedatasetsdemonstrateimprovedvisualqualityandgeometricconsistency,witha2.23dBhigherPSNRonDL3DVand32.4%lowerATEonMip-NeRF360comparedtoGLD.
View arXiv pageView PDFProject pageGitHub2Add to collection
Get this paper in your agent:
hf papers read 2609\.35734
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.35734 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.35734 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.35734 in a Space README.md to link it from this page.
Collections including this paper1
Similar Articles
MetaView: Monocular Novel View Synthesis with Scale-Aware Implicit Geometry Priors
MetaView proposes a diffusion-based monocular novel view synthesis framework that combines implicit geometry priors with metric depth guidance to achieve consistent and controllable rendering under large viewpoint changes from a single image.
UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models
Introduces UniWorld-View, a unified framework for large-baseline novel view synthesis from monocular inputs, integrating occlusion-aware point cloud rendering with video diffusion models for precise camera control and geometric consistency.
GAE: Learning a Geometry-Native Latent Space for 3D-Consistent World Generation
The paper introduces GAE, a geometry-native autoencoder that creates a compact latent space for generating 3D-consistent scenes, enhancing visual quality and coherence over existing methods.
VGGT-Diff: Visual Geometry Meets Diffusion for Sparse-View Novel View Synthesis
VGGT-Diff is a geometry-routed multi-view diffusion model that integrates visual geometry latents with a pretrained diffusion model to improve sparse-view novel view synthesis, achieving competitive or state-of-the-art performance.
Towards Consistent Video Geometry Estimation
ViGeo is a transformer-based foundation model that recovers dense and consistent 3D geometry from videos using dynamic chunking attention and a completion-based data refinement framework, achieving state-of-the-art performance across multiple tasks.