FLUX3D: High-Fidelity 3D Gaussian Generation with Diffusion-Aligned Sparse Representation

Hugging Face Daily Papers Papers

Summary

FLUX3D introduces a framework for high-fidelity image-to-3D Gaussian Splatting generation by enhancing representation learning and cross-modal alignment with diffusion-aligned structured latents and a sparse-structure-aware diffusion transformer, achieving state-of-the-art results.

Sparse voxel representation has emerged as a scalable foundation for image-to-3D Gaussian Splatting (3DGS) generation, yet current methods struggle to preserve high-frequency visual details of input images due to two structural bottlenecks. First, they adopt discriminative 2D features optimized for semantic abstraction to construct sparse voxel latents, which suppress reconstructive cues and induce a representation bottleneck. Second, in the generation stage, standard diffusion transformers lack effective mechanisms to align dense 2D image tokens with sparse 3D voxel latents, resulting in a cross-modal correspondence bottleneck. To address these issues, we propose FLUX3D, a scalable image-to-3DGS framework that boosts both representation learning and cross-modal alignment during generation. We first revisit 2D feature selection for sparse-voxel-based 3D representation learning, propose Diffusion-Aligned Structured Latents (DA-SLAT) and couple it with a decoder-only architecture to improve 3DGS reconstruction fidelity. We also design a sparse-structure-aware diffusion framework, which integrates the Sparse-structure Multimodal Diffusion Transformer (SMDiT) and Modal-Aware Rotary Positional Embedding (MARoPE) to achieve geometry-agnostic 2D-3D alignment. Extensive benchmark experiments demonstrate that FLUX3D yields substantial improvements in appearance fidelity and significantly outperforms all state-of-the-art (SOTA) methods in generating high-quality 3DGS assets.
Original Article
View Cached Full Text

Cached at: 06/24/26, 05:47 AM

Paper page - FLUX3D: High-Fidelity 3D Gaussian Generation with Diffusion-Aligned Sparse Representation

Source: https://huggingface.co/papers/2606.24874

Abstract

FLUX3D addresses limitations in image-to-3D Gaussian Splatting generation by improving representation learning and cross-modal alignment through specialized architectures and attention mechanisms.

Sparse voxel representationhas emerged as a scalable foundation forimage-to-3D Gaussian Splatting(3DGS) generation, yet current methods struggle to preserve high-frequency visual details of input images due to two structural bottlenecks. First, they adopt discriminative 2D features optimized for semantic abstraction to construct sparse voxel latents, which suppress reconstructive cues and induce a representation bottleneck. Second, in the generation stage, standarddiffusion transformerslack effective mechanisms to align dense 2D image tokens with sparse 3D voxel latents, resulting in across-modal correspondencebottleneck. To address these issues, we propose FLUX3D, a scalable image-to-3DGS framework that boosts both representation learning and cross-modal alignment during generation. We first revisit 2D feature selection for sparse-voxel-based 3D representation learning, proposeDiffusion-Aligned Structured Latents(DA-SLAT) and couple it with adecoder-only architectureto improve 3DGS reconstruction fidelity. We also design a sparse-structure-aware diffusion framework, which integrates theSparse-structure Multimodal Diffusion Transformer(SMDiT) andModal-Aware Rotary Positional Embedding(MARoPE) to achieve geometry-agnostic 2D-3D alignment. Extensive benchmark experiments demonstrate that FLUX3D yields substantial improvements in appearance fidelity and significantly outperforms all state-of-the-art (SOTA) methods in generating high-quality 3DGS assets.

View arXiv pageView PDFAdd to collection

Get this paper in your agent:

hf papers read 2606\.24874

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2606.24874 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2606.24874 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2606.24874 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

FLUX 3 Image

Hacker News Top

Black Forest Labs has released FLUX 3 Image, a controllable image generation model that supports synthesizing images from scratch, allowing precise layout of elements in the scene via bounding boxes and maximum control over every pixel.

CONFLUX: A Latent Diusion Model for 3D Chest-CT Synthesis with RL Post-Training

Hugging Face Daily Papers

CONFLUX is a 3D latent diffusion model for chest CT synthesis that achieves high-fidelity volumetric generation with controllable clinical attributes, enhanced by a reinforcement learning post-training stage to improve conditioning reliability. The model and a synthetic dataset of ~200k chest CT volumes are released.