Tag
This paper identifies a duality between continuous and discrete flow matching, showing that projecting continuous convex-interpolant paths via argmax yields discrete flows, and explores how different source geometries affect transition timing and generation quality.
This paper proposes a method using discrete diffusion to unlock lossless speedups in large language models, aiming to improve efficiency without compromising performance.
GRAS is a training-free method for reward alignment in discrete diffusion models that reduces variance in guided proposals and adaptively selects particles, achieving state-of-the-art results on DNA and protein design tasks.
The paper proposes dLLM-SetScore, a training-free framework using discrete masked-diffusion language models for multi-label text classification, achieving competitive performance with minimal validation data.
This paper introduces Simplax, an exact Dirichlet–categorical augmentation for discrete diffusion models that enriches training objectives and reverse transitions while preserving the original categorical corruption process, improving perplexity–entropy tradeoff on OpenWebText and validity on Sudoku.
Introduces Simplax, an exact Dirichlet-categorical augmentation for uniform discrete diffusion that improves reverse sampling and generative quality on text and Sudoku tasks.
Presents DLLM-TTS, a block discrete diffusion language model for text-to-speech synthesis that processes X-Codec2 tokens in blocks, enabling parallel generation with RTF 0.15 while achieving competitive quality with only 20K hours of training data.
The paper introduces Latent-Kernel Discrete Flow Maps (LKF), a flow-map kernel for discrete diffusion models that captures correlations between positions via a shared latent, enabling few-step generation without distillation and improving text generation perplexity.
PreDiff-LM proposes a hybrid attention mechanism that preserves causal attention for prompt tokens and bidirectional attention for masked target tokens, enabling adaptation of pretrained autoregressive models for discrete masked diffusion language modeling, achieving improvements in perplexity and downstream tasks over prior diffusion baselines.
This paper introduces a unified conceptual framework for discrete diffusion models, analyzing their design space through tokenization, state space construction, and highlighting trade-offs in training, inference, and scaling.
Introduces MobiDiff, an end-to-end discrete diffusion framework for generating human mobility data by denoising multi-channel semantic skeletons, achieving faster inference and competitive fidelity on real-world datasets.
This paper studies online adaptation strategies for discrete diffusion models in molecular optimization, identifying complementary components like acquisition, reward shaping, debiasing, replay, and validity control that improve feedback efficiency on small-molecule and protein tasks.
Proposes learning the unmasking order in masked diffusion models using a lightweight policy network, with a weighted loss that outperforms heuristics on combinatorial tasks and protein design.
Introduces TUBE, a variational upper bound on log-likelihood for discrete diffusion language models, enabling better evaluation and revealing that masked diffusion models still underperform autoregressive models.
This paper introduces TokenDrift, a drifting objective that refines discrete diffusion language models by lifting categorical predictions to a continuous semantic space for anti-symmetric drifting, significantly improving generation quality under a fixed number of denoising steps.
This paper introduces Constrained Diffusion for Code (CDC), a training-free neurosymbolic inference framework that integrates constraint satisfaction directly into the reverse denoising process of discrete diffusion models for code generation. CDC consistently improves constraint satisfaction in functional correctness, security, and syntax across benchmarks, outperforming existing diffusion and autoregressive baselines.
This paper proposes the 'support-before-frequency' hypothesis for discrete diffusion models, suggesting that models first learn the support (admissible sequences) before refining frequencies within the support. Theoretical analysis of small-noise reverse kernels and experiments on masked language diffusion models support this claim.
This paper introduces FeF-DLLM, a discrete diffusion language model that eliminates factorization errors by using exact prefix-conditioned factorization and accelerates inference via speculative decoding, achieving significant improvements in accuracy and speed on benchmarks such as GSM8K and MATH.
This paper introduces DiffLNS, a hybrid framework integrating a discrete denoising diffusion probabilistic model (D3PM) with LNS2 for multi-agent path finding, using sparse social attention to generate warm-start plans. It achieves high success rates on complex and congested settings, outperforming baselines.
This paper introduces a novel adaptive scheduler for steering discrete diffusion language models using sparse autoencoders, demonstrating that targeting interventions based on when specific attributes commit improves control quality and strength over uniform methods.