FLUX3D: High-Fidelity 3D Gaussian Generation with Diffusion-Aligned Sparse Representation
Summary
FLUX3D introduces a framework for high-fidelity image-to-3D Gaussian Splatting generation by enhancing representation learning and cross-modal alignment with diffusion-aligned structured latents and a sparse-structure-aware diffusion transformer, achieving state-of-the-art results.
View Cached Full Text
Cached at: 06/24/26, 05:47 AM
Paper page - FLUX3D: High-Fidelity 3D Gaussian Generation with Diffusion-Aligned Sparse Representation
Source: https://huggingface.co/papers/2606.24874
Abstract
FLUX3D addresses limitations in image-to-3D Gaussian Splatting generation by improving representation learning and cross-modal alignment through specialized architectures and attention mechanisms.
Sparse voxel representationhas emerged as a scalable foundation forimage-to-3D Gaussian Splatting(3DGS) generation, yet current methods struggle to preserve high-frequency visual details of input images due to two structural bottlenecks. First, they adopt discriminative 2D features optimized for semantic abstraction to construct sparse voxel latents, which suppress reconstructive cues and induce a representation bottleneck. Second, in the generation stage, standarddiffusion transformerslack effective mechanisms to align dense 2D image tokens with sparse 3D voxel latents, resulting in across-modal correspondencebottleneck. To address these issues, we propose FLUX3D, a scalable image-to-3DGS framework that boosts both representation learning and cross-modal alignment during generation. We first revisit 2D feature selection for sparse-voxel-based 3D representation learning, proposeDiffusion-Aligned Structured Latents(DA-SLAT) and couple it with adecoder-only architectureto improve 3DGS reconstruction fidelity. We also design a sparse-structure-aware diffusion framework, which integrates theSparse-structure Multimodal Diffusion Transformer(SMDiT) andModal-Aware Rotary Positional Embedding(MARoPE) to achieve geometry-agnostic 2D-3D alignment. Extensive benchmark experiments demonstrate that FLUX3D yields substantial improvements in appearance fidelity and significantly outperforms all state-of-the-art (SOTA) methods in generating high-quality 3DGS assets.
View arXiv pageView PDFAdd to collection
Get this paper in your agent:
hf papers read 2606\.24874
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2606.24874 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2606.24874 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2606.24874 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Monte Carlo Energy Aggregation for Mobile 3D Gaussian Splatting
Flux-GS enables real-time high-fidelity 3D Gaussian Splatting on mobile platforms through efficient lighting representation, attribute-conditioned enhancement, and multi-view densification strategies.
GS-Voxel: Fitting-Free Structured Latents for Large-Scale 3DGS Generation
GS-Voxel introduces a fitting-free framework to convert 3D Gaussian Splatting reconstructions into structured latents, enabling scalable generation of large-scale aerial 3D scenes via flow models and tiled inference.
FLUX 3 Image
Black Forest Labs has released FLUX 3 Image, a controllable image generation model that supports synthesizing images from scratch, allowing precise layout of elements in the scene via bounding boxes and maximum control over every pixel.
VidSplat: Gaussian Splatting Reconstruction with Geometry-Guided Video Diffusion Priors
VidSplat is a training-free generative reconstruction framework that uses video diffusion priors to recover complete 3D scenes from sparse inputs by synthesizing novel views.
CONFLUX: A Latent Diusion Model for 3D Chest-CT Synthesis with RL Post-Training
CONFLUX is a 3D latent diffusion model for chest CT synthesis that achieves high-fidelity volumetric generation with controllable clinical attributes, enhanced by a reinforcement learning post-training stage to improve conditioning reliability. The model and a synthetic dataset of ~200k chest CT volumes are released.