GlobalSplat: Efficient Feed-Forward 3D Gaussian Splatting via Global Scene Tokens
Summary
GlobalSplat introduces an efficient feed-forward framework for 3D Gaussian splatting that achieves compact and consistent scene reconstruction using global scene tokens, reducing computational overhead and inference time to under 78ms. The method uses a coarse-to-fine training approach to prevent representation bloat while maintaining competitive novel-view synthesis performance with significantly fewer Gaussians (16K) compared to dense baselines.
View Cached Full Text
Cached at: 04/20/26, 08:28 AM
Paper page - GlobalSplat: Efficient Feed-Forward 3D Gaussian Splatting via Global Scene Tokens
Source: https://huggingface.co/papers/2604.15284
Abstract
GlobalSplat introduces a global scene representation framework that achieves compact, consistent 3D Gaussian splatting with reduced computational overhead and improved inference speed.
The efficient spatial allocation of primitives serves as the foundation of 3D Gaussian Splatting (https://huggingface.co/papers?q=3D%20Gaussian%20Splatting), as it directly dictates the synergy between representation compactness (https://huggingface.co/papers?q=representation%20compactness), reconstruction speed (https://huggingface.co/papers?q=reconstruction%20speed), and rendering fidelity (https://huggingface.co/papers?q=rendering%20fidelity). Previous solutions, whether based on iterative optimization or feed-forward inference, suffer from significant trade-offs between these goals, mainly due to the reliance on local, heuristic-driven allocation strategies that lack global scene awareness. Specifically, current feed-forward methods are largely pixel-aligned or voxel-aligned. By unprojecting pixels into dense, view-aligned primitives, they bake redundancy into the 3D asset. As more input views are added, the representation size increases and global consistency becomes fragile. To this end, we introduce GlobalSplat, a framework built on the principle of align first, decode later. Our approach learns a compact, global, latent scene representation that encodes multi-view input and resolves cross-view correspondences (https://huggingface.co/papers?q=cross-view%20correspondences) before decoding any explicit 3D geometry. Crucially, this formulation enables compact, globally consistent reconstructions without relying on pretrained pixel-prediction backbones or reusing latent features from dense baselines. Utilizing a coarse-to-fine training (https://huggingface.co/papers?q=coarse-to-fine%20training) curriculum that gradually increases decoded capacity, GlobalSplat natively prevents representation bloat. On RealEstate10K and ACID, our model achieves competitive novel-view synthesis (https://huggingface.co/papers?q=novel-view%20synthesis) performance while utilizing as few as 16K Gaussians, significantly less than required by dense pipelines, obtaining a light 4MB footprint. Further, GlobalSplat enables significantly faster inference than the baselines, operating under 78 milliseconds in a single forward pass. Project page is available at https://r-itk.github.io/globalsplat/
View arXiv page (https://arxiv.org/abs/2604.15284)View PDF (https://arxiv.org/pdf/2604.15284)Project page (https://r-itk.github.io/globalsplat/)Add to collection (https://huggingface.co/login?next=%2Fpapers%2F2604.15284)
Get this paper in your agent:
hf papers read 2604.15284
Don’t have the latest CLI?curl -LsSf https://hf.co/cli/install.sh | bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2604.15284 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2604.15284 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2604.15284 in a Space README.md to link it from this page.
Collections including this paper2
Similar Articles
ATSplat: Compact Feed-forward 3D Gaussian Splatting with Adaptive Token Expansion
ATSplat introduces a feed-forward 3D Gaussian Splatting framework that uses adaptive 3D tokens to allocate primitives based on scene complexity, achieving state-of-the-art rendering quality while reducing the number of Gaussians by over 5.7 times compared to dense methods.
ZipSplat: Fewer Gaussians, Better Splats
ZipSplat is a token-based feed-forward 3D Gaussian Splatting model that uses k-means clustering to decouple Gaussian placement from the pixel grid, achieving ~6x fewer Gaussians while setting new state-of-the-art results on DL3DV and RealEstate10K without requiring ground-truth poses or intrinsics.
AsySplat: Efficient Asymmetric 3D Gaussian Splatting for Long-Sequence Scene Modeling
AsySplat proposes an asymmetric architecture that decouples geometry and appearance modeling in 3D Gaussian Splatting, achieving high efficiency for long-sequence scene modeling with nearly 800x speedup over optimization-based methods.
InfiniSplat: Implicit Gaussian Decoding for Large-Baseline Monocular View Synthesis
InfiniSplat presents a feed-forward single-image 3D Gaussian Splatting framework that uses geometry-guided sampling and query-conditioned implicit decoding to achieve surface-aligned Gaussian representation, improving large-baseline monocular view synthesis and generalizing from synthetic indoor training to open-world scenes.
SplatWeaver: Learning to Allocate Gaussian Primitives for Generalizable Novel View Synthesis
SplatWeaver is a feed-forward novel view synthesis framework that dynamically allocates 3D Gaussian primitives based on spatial complexity, improving rendering quality and efficiency over fixed-allocation methods. It leverages cardinality Gaussian experts and a pixel-level routing scheme guided by high-frequency priors to adaptively distribute primitives across complex and smooth scene regions.