Tag
HarmoCore introduces a functional latent diffusion method for reconstructing oscillatory wave fields from sparse sensor data, achieving significant improvements in efficiency and accuracy with minimal sensing in 2D and 3D scenarios.
SimCast-S2S is a generative latent-diffusion framework for probabilistic subseasonal precipitation forecasting that leverages transfer learning from climate simulations to outperform deep learning baselines and compete with operational systems.
KATok is an adaptive video tokenizer that selectively drops uninformative tokens for data-dependent compression, improving spatial consistency in diffusion-based video generation.
FarSky is a generative forecasting framework using task-aware latent-space coupling with latent diffusion to produce deterministic and probabilistic intra-hour solar irradiance forecasts, achieving up to 11 percentage points improvement in skill and better ramp event detection.
This paper introduces a bidirectional latent diffusion model that steps dynamical systems forward or backward in time, using round-trip consistency as a self-supervised test-time error signal to predict rollout errors without ground truth or ensembles.
This paper introduces KVAE, a family of tokenizers for audio, image, and video designed for text-conditioned generative models, claiming competitive or superior reconstruction and generation quality compared to existing open-source tokenizers. The code and training details are publicly released.
FMOPF uses latent flow matching with constraint-aware interaction priors to generate diverse, feasible near-optimal solutions for AC optimal power flow, scaling to hundreds of buses while preserving feasibility.
This paper introduces AlayaRenderer-Flash, a real-time generative world renderer that accelerates rendering from 0.56 FPS to 31.54 FPS using a few-step autoregressive streaming model and lightweight distilled codecs, enabling interactive play with a physics engine.
DiffGI introduces a differentiable geometry image representation for high-fidelity thin-shell 3D generation, enabling end-to-end optimization and superior reconstruction quality.
This paper introduces MIRA, the first large-scale multiplayer world model for highly dynamic physics-based environments, trained on 10,000 hours of Rocket League gameplay. The 5-billion-parameter latent diffusion model generates stable four-player rollouts in real time, with distributional quality holding steady for hours.
Patch-PODiff-ViT introduces a structured latent diffusion framework using patchwise Proper Orthogonal Decomposition (POD) for super-resolution and uncertainty quantification, enabling efficient diffusion with a fixed linear orthonormal basis and analytic propagation of predictive variance.
Introduces BrainG3N, a dual-purpose tokenizer for 3D brain MRI latent diffusion using a frozen masked autoencoder encoder for clinically informative embeddings and a CNN decoder for reconstruction, achieving state-of-the-art performance on a 23-task benchmark and enabling controllable generation and longitudinal forecasting.
This paper introduces TopVAE, a topology-optimized VAE that reduces 'dark areas' in molecular latent diffusion by making the decoder internalize structural and chemical constraints, achieving significant improvements in molecular generation quality.
This paper systematically compares reconstruction-based and semantic latent spaces for action-conditioned latent diffusion world models in robotics. It finds that semantic encoders like V-JEPA 2.1 generally outperform reconstruction encoders on policy-relevant metrics, advocating for semantic latent spaces as a stronger foundation for robotics world models.
NVIDIA introduces PiD, a Pixel Diffusion Decoder that replaces traditional VAE/RAE decoders in latent diffusion models, enabling fast, high-resolution decoding with up to 6× speedup and improved visual fidelity.
NVIDIA Spatial Intelligence Lab proposes PiD, which redesigns the decoding stage of latent diffusion models as a conditional pixel diffusion process, unifying decoding and upsampling to achieve low-latency, high-resolution decoding.
This paper proposes AirfoilGen, a latent diffusion model for airfoil shape generation that ensures geometric validity via a circle sweeping representation and enables control over aerodynamic performance (lift/drag coefficients). Experiments show 98.41% performance-conditioning accuracy, using a new dataset of over 200,000 airfoils.
This paper identifies a collapse-and-refine mechanism in diffusion models under the manifold hypothesis, proposing Score-induced Latent Diffusion (SiLD) that provably avoids the curse of dimensionality. Experiments show SiLD matches or outperforms VAE-based latent diffusion models.
Stable Audio 3 introduces a family of fast latent diffusion models for variable-length audio generation and editing, with open-source release of small and medium model weights.
This technical report investigates draft-conditioned latent refinement for non-autoregressive text generation, showing that good latent geometry does not guarantee good decoding and emphasizing decoder recoverability as a key evaluation metric.