Tag
This paper introduces Simplax, an exact Dirichlet–categorical augmentation for discrete diffusion models that enriches training objectives and reverse transitions while preserving the original categorical corruption process, improving perplexity–entropy tradeoff on OpenWebText and validity on Sudoku.
This paper studies inverse sampling for Lévy-driven generative models, proposing a structured reverse sampler that decomposes dynamics into diffusion, small jump, and large jump components, with neural networks amortizing jump rates. The method is applied to OFDM-SISO channel estimation under mixed Gaussian and impulsive noise.
This paper proposes DURA, a diffusion-based unrestricted robotic attack that generates visually natural adversarial patches to disrupt Vision-Language-Action (VLA) models in both white-box and black-box settings, highlighting safety risks for physically deployed robotic systems.
Presents LibraSpec, a training-free, plug-and-play algorithm that dynamically selects speculative decoding lengths via marginal-gain-driven optimization, achieving consistent speedups across multiple models and benchmarks.
This paper introduces Atelier, a method that plans explicit control states before generation to prevent text-to-image models from falling back on artist-name shortcuts, and presents ArtIntentBench for evaluating artist-grounded style control.
This paper introduces a bidirectional latent diffusion model that steps dynamical systems forward or backward in time, using round-trip consistency as a self-supervised test-time error signal to predict rollout errors without ground truth or ensembles.
This paper introduces EddyFlow, a deep learning framework for kilometer-scale sea surface temperature downscaling that balances predictive accuracy, scale-dependent structure, and regional generalization. It achieves strong zero-shot performance and near-ideal spectral fidelity across multiple ocean regions.
This paper introduces Steerling-8B, a diffusion language model trained with interpretability as a constraint, showing that interpretability improves with scale and enabling concept steering without retraining.
This thesis introduces polynomial representations for long-term traffic scene prediction in autonomous driving, showing improved computational efficiency, generalization, and prediction plausibility over sequence-based baselines, validated on Argoverse 2 and Waymo Open datasets.
Introduces UniWorld-View, a unified framework for large-baseline novel view synthesis from monocular inputs, integrating occlusion-aware point cloud rendering with video diffusion models for precise camera control and geometric consistency.
Introduces FairDiffuseVQVAE, a two-stage tabular diffusion model that achieves fairness at sampling time by conditioning on protected attributes, outperforming prior fair tabular generators on demographic parity and equalized odds.
This bioRxiv preprint introduces Diff-Switch, a framework that uses diffusion-based ensemble sampling to generate conformational states for de novo protein switch design, improving the success rate of finding switch-compatible sequences.
This paper introduces Synthetic Self-Guidance (SSG), a method that attaches a lightweight prediction head to a frozen pretrained pixel-space diffusion model, using the discrepancy between intermediate and final predictions as self-guidance during sampling. It shows that model-generated samples suffice for training the head, improving FID by over 50% on several variants without classifier-free guidance and enhancing strong baselines with CFG.
This paper studies empirical scaling properties for text conditioning in visual generation, showing that converged diffusion loss scales with structured language in prompts, and introduces methods to improve diffusability and promptability.
This paper introduces SGFlow, a method for learning flow maps for diffusion models that avoids invertibility constraints and backpropagation through model iterations, achieving competitive FID scores on CIFAR with a proven stationary-point guarantee.
This paper introduces Chimera, a hybrid visual diffusion backbone with a principled scaling recipe, combining Kimi Delta Attention, Multi-head Latent Attention, and sparse Mixture-of-Experts to efficiently handle long-context image and video generation. It also presents HeteroP, a module-wise hyperparameter transfer scheme, and Chinchilla-style scaling laws to train an 11B-parameter model with 2B activated parameters.
NVIDIA introduces Parallel Decoding Distillation (PDD) for accelerating image and video generation, enabling high-quality outputs with fewer neural function evaluations on models like LTX-2.3 and Wan2.1-14B.
Proposes neuromorphic masked diffusion language models (N-MDLMs) that integrate block diffusion with spike-based neuromorphic computation to improve throughput and energy efficiency by leveraging sparsity and generating multiple tokens per parameter access, analyzed via a roofline-inspired model.
Introduces GenTO, a diffusion model-based framework that steers topology distributions for unified design of architected metamaterials, achieving diverse design tasks with reusable topology priors and experimental validation.
PRESTO introduces a prefix-aligned tree drafting framework for diffusion speculative decoding, achieving up to 1.5x speedup on dedicated diffusion drafters and 1.12x on self-speculative diffusion LLMs.