@xuanchi13: The latent-vs-pixel debate misses the point. GPT Image 2 shows what users notice: pixel-level fidelity. Latent models s…
Summary
NVIDIA introduces PiD, a Pixel Diffusion Decoder that replaces traditional VAE/RAE decoders in latent diffusion models, enabling fast, high-resolution decoding with up to 6× speedup and improved visual fidelity.
View Cached Full Text
Cached at: 05/27/26, 09:21 AM
The latent-vs-pixel debate misses the point.
GPT Image 2 shows what users notice: pixel-level fidelity. Latent models show what scales: compact semantic structure.
We connect them by replacing VAE/RAE decoders with a Pixel Diffusion Decoder.
Code and Model available: https://research.nvidia.com/labs/sil/projects/pid/…
(1/N)
Fast and High-Resolution Latent Decoding with Pixel Diffusion
Source: https://research.nvidia.com/labs/sil/projects/pid/


VAE Decoder
PiD


RAE Decoder
PiD


VAE Decoder
PiD


VAE Decoder
PiD
Abstract
Most practical high-resolution text-to-image systems rely on latent diffusion models, where generation is performed in a compact latent space and a decoder maps latents back to pixels. Yet the latent-to-pixel decoder is reconstruction-oriented, optimized to invert the encoder rather than synthesize more details, and becomes increasingly costly at megapixel scale. This drawback calls for a more expressive and efficient decoding paradigm. Motivated by recent progress in scalable pixel-space diffusion, we introducePiD, aPixel diffusionDecoder that reformulates latent decoding as conditional pixel diffusion, unifying decoding and upsampling into one generative module. By denoising directly in high-resolution pixel space,PiDsynthesizes 4× and even 8× upscaled images with low latency. For latent conditioning, a lightweight sigma-aware adapter injects noise-corrupted latents into the pixel diffusion backbone, enablingPiDto decode partially denoised latents and terminate the latent diffusion process early. To further improve efficiency, we distill the model using DMD2, reducing inference to just 4 steps.PiDapplies to both conventional VAE latents and semantic latents (e.g., SigLIP, DINOv2) used in recent RAE-based models.PiDdecodes latents of 512×512 images into 2048×2048 pixels in under 1 second with 13 GB peak memory on a consumer RTX 5090, and as fast as 210 ms on a GB200 GPU, about 6× faster than cascaded diffusion-based super-resolution pipelines with better visual fidelity.
Results
From Latent to Pixels
Select a latent space and move the step slider to compare PiD decoding quality at different early-termination points. Drag thewhite divideron each image to reveal the VAE/RAE decode vs. PiD decode.
4K Decode
Direct latent→4K decoding with PiD. Click any image to launch a side-by-side comparison against the VAE decoder.
Baseline Comparison
Hover over any image to activate the synchronized zoom lens across all six views.
Quantitative Results(Decoding + Upsampling, 512² → 2048²)
End-to-End Decoding Latency (ms) ↓
PiD is up to5.9× fasterthan SeedVR2 (211.2 ms vs 1237.5 ms)
Gemini-3-Flash Judge Rating (%) ↑
% of evaluations where judges preferPiDover each baseline
Method
**Overview of PiD.**PiD unifies latent decoding and upsampling as a single latent-conditioned pixel diffusion model that predicts the target-resolution pixel-space velocity field. Noise-corrupted latent training and sigma-aware gating make the decoder robust to partially denoised latents, enabling early exit from the base LDM while preserving high-resolution output quality.
Similar Articles
nvidia/PiD
NVIDIA releases PiD (Pixel Diffusion Decoder), a conditional pixel-space diffusion model that unifies latent-to-pixel decoding and upsampling into one generative module, producing super-resolved images in one pass. Model checkpoints and VAE weights are released under a non-commercial license.
PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion
PiD introduces a pixel diffusion decoder that reformulates latent decoding as conditional pixel diffusion, enabling fast and high-quality image synthesis at high resolutions with reduced computational requirements. It decodes latents into 4x or 8x upscaled images in under a second on consumer hardware.
@FeitengLi: NVIDIA Spatial Intelligence Lab proposes PiD, redesigning the decoding stage in latent diffusion models. Current mainstream text-to-image generation happens in latent space, then uses a VAE decoder to map back to pixels. This decoder's…
NVIDIA Spatial Intelligence Lab proposes PiD, which redesigns the decoding stage of latent diffusion models as a conditional pixel diffusion process, unifying decoding and upsampling to achieve low-latency, high-resolution decoding.
L2P: Unlocking Latent Potential for Pixel Generation
The L2P paper introduces a Latent-to-Pixel transfer paradigm that leverages pre-trained latent diffusion models to create efficient pixel-space models capable of 4K generation with minimal training overhead.
An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models
This paper proposes a latent-to-pixel training strategy for pixel-space text-to-image diffusion models, accelerating convergence and improving inference speed while matching or surpassing latent-space counterparts.