latent-diffusion

Tag

Cards List
#latent-diffusion

HarmoCore: Functional Latent Diffusion for Sparse Reconstruction of Oscillatory Wave Fields

arXiv cs.LG · 5d ago Cached

HarmoCore introduces a functional latent diffusion method for reconstructing oscillatory wave fields from sparse sensor data, achieving significant improvements in efficiency and accuracy with minimal sensing in 2D and 3D scenarios.

0 favorites 0 likes
#latent-diffusion

SimCast-S2S: An Efficient Generative Model for Subseasonal Precipitation Forecasting via Transfer Learning from Climate Simulations

arXiv cs.LG · 2026-08-28 Cached

SimCast-S2S is a generative latent-diffusion framework for probabilistic subseasonal precipitation forecasting that leverages transfer learning from climate simulations to outperform deep learning baselines and compete with operational systems.

0 favorites 0 likes
#latent-diffusion

Keep-or-Drop? Adaptive Tokenizer for Compact Video Representation

Hugging Face Daily Papers · 2026-08-25 Cached

KATok is an adaptive video tokenizer that selectively drops uninformative tokens for data-dependent compression, improving spatial consistency in diffusion-based video generation.

0 favorites 0 likes
#latent-diffusion

FarSky: Task-Aware Latent-Space Coupling for Generative Intra-Hour Solar Forecasting

arXiv cs.LG · 2026-08-13 Cached

FarSky is a generative forecasting framework using task-aware latent-space coupling with latent diffusion to produce deterministic and probabilistic intra-hour solar irradiance forecasts, achieving up to 11 percentage points improvement in skill and better ramp event detection.

0 favorites 0 likes
#latent-diffusion

Round-Trip Consistency: Bidirectional Diffusion Models Can Predict Their Own Rollout Errors [R]

Reddit r/MachineLearning · 2026-08-06

This paper introduces a bidirectional latent diffusion model that steps dynamical systems forward or backward in time, using round-trip consistency as a self-supervised test-time error signal to predict rollout errors without ground truth or ensembles.

0 favorites 0 likes
#latent-diffusion

KVAE: Family of Tokenizers for Multimodal Generative Models

Hugging Face Daily Papers · 2026-08-06 Cached

This paper introduces KVAE, a family of tokenizers for audio, image, and video designed for text-conditioned generative models, claiming competitive or superior reconstruction and generation quality compared to existing open-source tokenizers. The code and training details are publicly released.

0 favorites 0 likes
#latent-diffusion

FMOPF: Latent Flow Matching with Constraint-Aware Interaction Priors for AC Optimal Power Flow

arXiv cs.LG · 2026-07-28 Cached

FMOPF uses latent flow matching with constraint-aware interaction priors to generate diverse, feasible near-optimal solutions for AC optimal power flow, scaling to hundreds of buses while preserving feasibility.

0 favorites 0 likes
#latent-diffusion

Generative World Renderer at the Speed of Play

Hugging Face Daily Papers · 2026-07-21 Cached

This paper introduces AlayaRenderer-Flash, a real-time generative world renderer that accelerates rendering from 0.56 FPS to 31.54 FPS using a few-step autoregressive streaming model and lightweight distilled codecs, enabling interactive play with a physics engine.

0 favorites 0 likes
#latent-diffusion

DiffGI: Differentiable Geometry Images for High-Fidelity Thin-Shell 3D Generation

Hugging Face Daily Papers · 2026-07-15 Cached

DiffGI introduces a differentiable geometry image representation for high-fidelity thin-shell 3D generation, enabling end-to-end optimization and superior reconstruction quality.

0 favorites 0 likes
#latent-diffusion

Multiplayer Interactive World Models with Representation Autoencoders

Hugging Face Daily Papers · 2026-07-06 Cached

This paper introduces MIRA, the first large-scale multiplayer world model for highly dynamic physics-based environments, trained on 10,000 hours of Rocket League gameplay. The 5-billion-parameter latent diffusion model generates stable four-player rollouts in real time, with distributional quality holding steady for hours.

0 favorites 0 likes
#latent-diffusion

Patch-PODiff-ViT: Structured Latent Diffusion with Patchwise POD for Super-Resolution and Uncertainty Quantification

arXiv cs.LG · 2026-07-01 Cached

Patch-PODiff-ViT introduces a structured latent diffusion framework using patchwise Proper Orthogonal Decomposition (POD) for super-resolution and uncertainty quantification, enabling efficient diffusion with a fixed linear orthonormal basis and analytic propagation of predictive variance.

0 favorites 0 likes
#latent-diffusion

BrainG3N: A Dual-Purpose Tokenizer for Controllable 3D Brain MRI Generation

arXiv cs.AI · 2026-06-20 Cached

Introduces BrainG3N, a dual-purpose tokenizer for 3D brain MRI latent diffusion using a frozen masked autoencoder encoder for clinically informative embeddings and a CNN decoder for reconstruction, achieving state-of-the-art performance on a 23-task benchmark and enabling controllable generation and longitudinal forecasting.

0 favorites 0 likes
#latent-diffusion

Smoothing Dark Areas in Molecular Latent Diffusion

arXiv cs.LG · 2026-06-15 Cached

This paper introduces TopVAE, a topology-optimized VAE that reduces 'dark areas' in molecular latent diffusion by making the decoder internalize structural and chemical constraints, achieving significant improvements in molecular generation quality.

0 favorites 0 likes
#latent-diffusion

@artemZholus: thanks! in the second paper (https://arxiv.org/abs/2605.06388) we used your (and RAE's) recipe and it worked.

X AI KOLs Following · 2026-05-26 Cached

This paper systematically compares reconstruction-based and semantic latent spaces for action-conditioned latent diffusion world models in robotics. It finds that semantic encoders like V-JEPA 2.1 generally outperform reconstruction encoders on policy-relevant metrics, advocating for semantic latent spaces as a stronger foundation for robotics world models.

0 favorites 0 likes
#latent-diffusion

@xuanchi13: The latent-vs-pixel debate misses the point. GPT Image 2 shows what users notice: pixel-level fidelity. Latent models s…

X AI KOLs Timeline · 2026-05-26 Cached

NVIDIA introduces PiD, a Pixel Diffusion Decoder that replaces traditional VAE/RAE decoders in latent diffusion models, enabling fast, high-resolution decoding with up to 6× speedup and improved visual fidelity.

0 favorites 0 likes
#latent-diffusion

@FeitengLi: NVIDIA Spatial Intelligence Lab proposes PiD, redesigning the decoding stage in latent diffusion models. Current mainstream text-to-image generation happens in latent space, then uses a VAE decoder to map back to pixels. This decoder's…

X AI KOLs Timeline · 2026-05-25 Cached

NVIDIA Spatial Intelligence Lab proposes PiD, which redesigns the decoding stage of latent diffusion models as a conditional pixel diffusion process, unifying decoding and upsampling to achieve low-latency, high-resolution decoding.

0 favorites 0 likes
#latent-diffusion

AirfoilGen: A valid-by-construction and performance-aware latent diffusion model for airfoil generation

arXiv cs.LG · 2026-05-21 Cached

This paper proposes AirfoilGen, a latent diffusion model for airfoil shape generation that ensures geometric validity via a circle sweeping representation and enables control over aerodynamic performance (lift/drag coefficients). Experiments show 98.41% performance-conditioning accuracy, using a new dataset of over 200,000 airfoils.

0 favorites 0 likes
#latent-diffusion

Provably Learning Diffusion Models under the Manifold Hypothesis: Collapse and Refine

arXiv cs.LG · 2026-05-21 Cached

This paper identifies a collapse-and-refine mechanism in diffusion models under the manifold hypothesis, proposing Score-induced Latent Diffusion (SiLD) that provably avoids the curse of dimensionality. Experiments show SiLD matches or outperforms VAE-based latent diffusion models.

0 favorites 0 likes
#latent-diffusion

Stable Audio 3

Hacker News Top · 2026-05-20 Cached

Stable Audio 3 introduces a family of fast latent diffusion models for variable-length audio generation and editing, with open-source release of small and medium model weights.

0 favorites 0 likes
#latent-diffusion

When Latent Geometry Is Not Enough: Draft-Conditioned Latent Refinement for Non-Autoregressive Text Generation

arXiv cs.CL · 2026-05-18 Cached

This technical report investigates draft-conditioned latent refinement for non-autoregressive text generation, showing that good latent geometry does not guarantee good decoding and emphasizing decoder recoverability as a key evaluation metric.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback