Tag
PixSDS identifies VAE-induced pixel drift in latent score distillation sampling and proposes a gradient repair method that decodes latent SDS lookahead steps to guide pixel-space optimization, reducing artifacts in text-to-3D generation.
This paper studies a self-supervised task for generating single-cell gene expression vectors using an autoregressive transformer with a quantized VAE tokenizer. It reports scaling laws and a compute-optimal frontier for single-cell foundation models, with potential fine-tuning for perturbation prediction.
Kijai shares a quickly trained 2D tiny VAE for MiniMax-H3, intended to improve latent previews in ComfyUI. A better alternative trained by madebyollin is also linked.
Meshy T2 introduces a fast native mesh generation framework using flow matching and a vertex-set mesh VAE, achieving state-of-the-art geometric fidelity with end-to-end image-to-mesh generation in a median of 6 seconds, over an order of magnitude faster than autoregressive baselines.
OmniVAE is a jointly trained audio-video VAE that uses segment-level contrastive learning and feature distillation to align latent spaces, improving joint generation quality and synchronization in text-to-audio-video generation.
Proposes a VAE-based multi-task semantic communication framework for satellite-assisted autonomous driving, achieving significant bandwidth reduction while maintaining performance for traffic sign reconstruction and classification.
Introduces Cross-Space Distillation, a method to transfer knowledge from modern high-capacity diffusion models to compact student models with different latent spaces using a lightweight latent interface called Bridge, enabling quality improvements without modifying the student backbone.
This paper introduces TopVAE, a topology-optimized VAE that reduces 'dark areas' in molecular latent diffusion by making the decoder internalize structural and chemical constraints, achieving significant improvements in molecular generation quality.
Ideogram-4 model repackaged for ComfyUI, including fp8 scaled diffusion models, Qwen3VL text encoder, and FLUX VAE.
This paper introduces O-Voxel, a new sparse voxel representation for 3D generative modeling that efficiently handles complex topologies and appearance, and trains large-scale flow-matching models with 4B parameters to achieve state-of-the-art generation quality.