Residual Context Diffusion Language Models (2 minute read)
Summary
This paper introduces Residual Context Diffusion (RCD), a module that recycles discarded token representations in diffusion language models to improve efficiency and accuracy, achieving 5–10% better accuracy and up to 4–5x fewer denoising steps on challenging reasoning tasks.
View Cached Full Text
Cached at: 07/03/26, 05:22 PM
Similar Articles
Multi-Token Residual Prediction
Introduces Multi-token Residual Prediction (MRP), a lightweight module for diffusion language models that enables dependency-aware multi-token denoising within a single backbone forward pass, achieving up to 1.42× lossless speedup.
Nemotron-TwoTower: Diffusion Language Modeling with Pretrained Autoregressive Context
The paper proposes Nemotron-TwoTower, a diffusion language model that decouples context representation and denoising using a frozen autoregressive tower and a trainable diffusion denoiser, achieving 98.7% of baseline quality with 2.42x throughput.
Continuous Diffusion Language Models (CDLM's)
The article discusses the resurgence of continuous diffusion models for language generation, highlighting recent research and historical context that challenges the dominance of autoregressive language models.
Beyond Token Positions: Safety Alignment Across Denoising Steps in Diffusion Language Models
The paper proposes a training-free decoding method called Refusal-Aware Early Commitment (RAEC) to improve safety alignment in diffusion language models by leveraging refusal signals from early denoising steps.
$R^2$-dLLM: Accelerating Diffusion Large Language Models via Spatio-Temporal Redundancy Reduction
R²-dLLM introduces spatio-temporal redundancy reduction techniques that cut diffusion LLM decoding steps by up to 75% while preserving generation quality, addressing a key deployment bottleneck.