MetaView: Monocular Novel View Synthesis with Scale-Aware Implicit Geometry Priors
Summary
MetaView proposes a diffusion-based monocular novel view synthesis framework that combines implicit geometry priors with metric depth guidance to achieve consistent and controllable rendering under large viewpoint changes from a single image.
View Cached Full Text
Cached at: 07/16/26, 05:42 AM
Paper page - MetaView: Monocular Novel View Synthesis with Scale-Aware Implicit Geometry Priors
Source: https://huggingface.co/papers/2607.12000 Published on Jul 13
·
Submitted byhttps://huggingface.co/KaiiWuu1993
Wu Kaion Jul 16
Abstract
Currentvisualgenerationmodelsarecapableofproducinghigh-qualitycontent,yettheylackacoherentperceptionofthespatialstructure.Existinggenerativenovelviewsynthesismethodstypicallyintroduceexplicitgeometrypriors,whichenforcespatialconsistencybutinherentlyrestrictgeneralizationinlargeviewchanges.Incontrast,recentinteractivegenerativemethodsfavorimplicitscenemodeling,offeringgreaterflexibilityatthecostofprecisecameracontrolandgeometryconsistency.Inthispaper,weproposeMetaView,adiffusion-basedmonocularnovelviewsynthesisframeworkthatenablesrenderingunderlargeviewchangesfromasingleimage.Ourkeyinsightistocombineimplicitgeometrymodelingwithminimalyetessentialexplicit3Dcues:weincorporateimplicitgeometrypriorsfromafeed-forwardgeometryperceptionnetworktoregularizestructurewithoutimposingrestrictivereconstructionpipelines,whileleveragingmetricdepthtoanchorthegenerationtoametricscale.ThisdesignallowsMetaViewtoachievebothgeometryconsistencyandprecisecontrollability.Extensiveexperimentsdemonstratethat,underchallengingmonocularlargeviewpointchanges,MetaViewsignificantlyoutperformsexistingmethodsandexhibitssuperiorgeneralization.Ourcodeispubliclyavailableathttps://github.com/KlingAIResearch/MetaView.
View arXiv pageView PDFProject pageGitHub3Add to collection
Get this paper in your agent:
hf papers read 2607\.12000
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper1
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2607.12000 in a dataset README.md to link it from this page.
Spaces citing this paper1
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models
Introduces UniWorld-View, a unified framework for large-baseline novel view synthesis from monocular inputs, integrating occlusion-aware point cloud rendering with video diffusion models for precise camera control and geometric consistency.
MoCam: Unified Novel View Synthesis via Structured Denoising Dynamics
MoCam is a research paper introducing a diffusion-based framework for unified novel view synthesis that dynamically coordinates geometric and appearance priors to improve robustness against geometric errors.
InfiniSplat: Implicit Gaussian Decoding for Large-Baseline Monocular View Synthesis
InfiniSplat presents a feed-forward single-image 3D Gaussian Splatting framework that uses geometry-guided sampling and query-conditioned implicit decoding to achieve surface-aligned Gaussian representation, improving large-baseline monocular view synthesis and generalizing from synthetic indoor training to open-world scenes.
RayDer: Scalable Self-Supervised Novel View Synthesis from Real-World Video
RayDer is a unified feed-forward transformer that consolidates camera estimation, scene reconstruction, and rendering for self-supervised novel view synthesis from real-world video, achieving clean power-law scaling and strong zero-shot performance.
MVTrack4Gen: Multi-View Point Tracking as Geometric Supervision for 4D Video Generation
MVTrack4Gen introduces a training framework that uses multi-view point tracking as geometric supervision to enhance motion-aware diffusion models, achieving state-of-the-art geometric consistency and motion fidelity in novel-view video generation from monocular video.