非均匀离散扩散:更强、更可扩展
摘要
本文提出 LUDI,一种非均匀扩散语言建模框架,解决了均匀扩散语言模型中训练目标过度均匀以及条件与目标混淆的问题,实现了 7B 规模的均匀离散扩散语言模型(UDLM),在逐步处理速度上比自回归解码快 3 倍,同时在复杂推理任务上具备有竞争力的性能。
arXiv:2609.35817v1 Announce Type: new
Abstract: Although uniform diffusion language models (UDLMs) represent a promising diffusion paradigm, scaling them remains challenging. We identify the core obstacle as an over-uniform training objective and condition-target confusion during sampling. To address these, we propose Less Uniform Diffusion (LUDI), a novel UDLM framework. Specifically, we (i) introduce a less uniform loss that directs each reverse transition toward the clean token, and (ii) equip the model with per-token time embeddings that supply token-level corruption hints, enabling confidence-based few-step sampling. Experiments across scales show that LUDI yields cleaner supervision and improves few-step generation. We further continue-train a 7B autoregressive model into LUDI-7B, resulting in a UDLM capable of complex reasoning. It achieves a 3-token-per-step speedup over AR decoding and competitive performance compared with masked diffusion baselines, revealing that the full potential of UDLMs for complex generation remains to be unlocked.
查看缓存全文
缓存时间: 2026/09/30 09:48
# Less Uniform Discrete Diffusion is More Powerful and Scalable
Source: [https://arxiv.org/html/2609.35817](https://arxiv.org/html/2609.35817)
Kaibo Wang\\cofirstAffiliation:The Hong Kong University of Science and TechnologyAffiliation:Huawei Foundation Model DepartmentDing Ding\\cofirstAffiliation:The Hong Kong University of Science and TechnologyAffiliation:Huawei Foundation Model DepartmentFangyu DingAffiliation:The Hong Kong University of Science and TechnologyAffiliation:Huawei Foundation Model DepartmentHan ShiAffiliation:Huawei Foundation Model DepartmentHaoli BaiAffiliation:Huawei Foundation Model DepartmentJiacheng Sun\\corrauthorAffiliation:Huawei Foundation Model DepartmentYang Xiang\\corrauthorAffiliation:The Hong Kong University of Science and Technology
###### Abstract
Although uniform diffusion language models \(UDLMs\) represent a promising diffusion paradigm, scaling them remains challenging\. We identify the core obstacle as an over\-uniform training objective and condition\-target confusion during sampling\. To address these, we propose Less Uniform Diffusion \(LUDI\), a novel UDLM framework\. Specifically, we \(i\) introduce a less uniform loss that directs each reverse transition toward the clean token, and \(ii\) equip the model with per\-token time embeddings that supply token\-level corruption hints, enabling confidence\-based few\-step sampling\. Experiments across scales show that LUDI yields cleaner supervision and improves few\-step generation\. We further continue\-train a 7B autoregressive model into LUDI\-7B, resulting in a UDLM capable of complex reasoning\. It achieves a 3\-token\-per\-step speedup over AR decoding and competitive performance compared with masked diffusion baselines, revealing that the full potential of UDLMs for complex generation remains to be unlocked\.
††footnotetext:## 1Introduction
Diffusion large language models \(dLLMs\) have emerged as competitive alternatives to autoregressive \(AR\) models, driven by their potential for parallel decoding and bidirectional context modeling\([Sahoo et al\., 2024](https://arxiv.org/html/2609.35817#bib.bib1);[Ou et al\., 2025](https://arxiv.org/html/2609.35817#bib.bib8)\)\. Among them, masked diffusion language models \(MDLMs, Figure[1](https://arxiv.org/html/2609.35817#S1.F1)[1\(a\)](https://arxiv.org/html/2609.35817#S1.F1.sf1)\), which use a mask token as the noise prior, dominate current research, with models at 8B and even 100B scales delivering strong results on tasks like mathematics and programming\([Nie et al\., 2025b](https://arxiv.org/html/2609.35817#bib.bib2);[Bie et al\., 2025](https://arxiv.org/html/2609.35817#bib.bib3)\)\. Uniform diffusion language models \(UDLMs, Figure[1](https://arxiv.org/html/2609.35817#S1.F1)[1\(b\)](https://arxiv.org/html/2609.35817#S1.F1.sf2)\), which instead adopt a uniform prior over the vocabulary, represent another branch that has shown greater potential for few\-step generation and self\-correction\([Sahoo et al\., 2025](https://arxiv.org/html/2609.35817#bib.bib4);[Schiff et al\., 2025](https://arxiv.org/html/2609.35817#bib.bib7)\)\. Yet their development lags significantly behind\. To date, the advantages of UDLMs have only been demonstrated at small scales, and their viability on complex reasoning tasks remains unestablished\.
\(a\)Masked diffusion \(MDLM\)\.
\(b\)Uniform diffusion \(UDLM\)\.
Figure 1:Schematic comparison between masked and uniform diffusion\.\(a\)Less uniform loss\.
\(b\)LUDI with per\-token time embeddings\.
Figure 2:Illustration of Less Uniform Diffusion \(LUDI\)\. \(a\)LU Loss\.The ELBO loss \(left\) matches the reverse transition top\(𝐱t−Δt∣𝐱t,𝐱0\)p\(\\mathbf\{x\}\_\{t\-\\Delta t\}\\mid\\mathbf\{x\}\_\{t\},\\mathbf\{x\}\_\{0\}\), leading to over\-uniform predictions by encouraging all token probabilities \(∑i\\sum\_\{i\},➜\)\. Our LU loss \(right\) removes the smoothing term and retains only the suppression of the erroneous token \(mm➜\) and promotion of the target token \(yy➜\)\. \(b\)Per\-token time embedding\.The model receives token\-level corruption hints via per\-token time embeddings\. Combined with the LU loss, this allows scaling LUDI to 7B parameters from AR initialization, enabling confidence\-based decoding for complex reasoning tasks\.Prior efforts mainly focused on conceptually validating the promise of UDLMs\. UDLMs can be distilled through connections to continuous Gaussian diffusion\([Sahoo et al\., 2025](https://arxiv.org/html/2609.35817#bib.bib4)\), trained stably with a simplified loss\([Zhu et al\., 2025c](https://arxiv.org/html/2609.35817#bib.bib5)\), and exhibit more favorable scaling behavior under data\-constrained conditions\([von Rütte et al\., 2025b](https://arxiv.org/html/2609.35817#bib.bib6)\)\. Nevertheless, existing efforts remain confined to small datasets or scaling\-law investigations, leaving a pressing question:Can UDLMs scale to a larger regime and handle complex reasoning tasks?
We find that UDLMs do not scale up as naturally as MDLMs\. The central difficulty arises from anover\-uniform training objectiveandcondition\-target confusionduring sampling\. During training, the KL\-based objective provides an excessively label\-smoothed learning signal: the model’s predictions become more uniform and less confident in the clean token\. During sampling, for complex tasks, simpler tokens should be generated first to serve as conditions, after which the remaining tokens are denoised\. In UDLMs, however, all tokens are denoised at a uniform rate\. While this works for simple tasks, it leads to condition\-target confusion on complex tasks, eventually causing the model to collapse to repeated tokens\. These two issues jointly limit the performance of UDLMs on complex tasks: over\-uniformity reduces training efficiency and weakens few\-step generation capability, while uniformly updating all tokens across many steps leads to condition\-target confusion\.
To address these issues and make UDLMs practically scalable, we proposeLessUniformDiffusion \(LUDI\)\. LUDI introduces two key modifications\. First, by simplifying and approximating the original UDLM training loss, we decompose it into three terms and remove the term that encourages uniform predictions, yielding a less uniform \(LU\) loss \(Figure[2](https://arxiv.org/html/2609.35817#S1.F2)[2\(a\)](https://arxiv.org/html/2609.35817#S1.F2.sf1)\)\. We prove theoretically that LU loss steers each reverse transition toward the clean token, supporting few\-step sampling\. Second, to mitigate the confusion between conditions and targets, we provide each token with an individual time embedding during training, which indicates the probability that the token has been randomly replaced \(Figure[2](https://arxiv.org/html/2609.35817#S1.F2)[2\(b\)](https://arxiv.org/html/2609.35817#S1.F2.sf2)\)\. This allows the model to better distinguish condition tokens from target tokens and naturally enables confidence\-based sampling strategies akin to those in MDLMs, facilitating an easy\-to\-hard generation paradigm\.
We conduct extensive experiments to validate LUDI\. Pretraining at 170M and 1B scales demonstrates that the LU loss offers cleaner supervision and improves both few\-step generation and downstream task performance\. We further obtain LUDI\-7B by continuing training from an autoregressive checkpoint, a UDLM capable of complex reasoning tasks\. Compared with the AR model, LUDI\-7B achieves approximately 3 tokens per step while maintaining comparable performance\. Compared with MDLMs, LUDI\-7B also achieves competitive performance, revealing that the potential of UDLMs for complex generation remains largely untapped\. Our contributions are threefold:
1. 1\.We attribute the difficulty of scaling UDLMs to over\-uniform training objectives and condition\-target confusion during sampling\.
2. 2\.We propose the LUDI loss and per\-token time embeddings, enabling UDLMs to perform confidence\-based few\-step generation\.
3. 3\.Our experiments with LUDI\-7B establish that UDLMs scale effectively for complex reasoning, delivering a competitive balance between generation quality and speed\.
## 2Background
In this section, we review UDLM\. A clean token is corrupted by the forward process into the uniform distribution\. By parameterizing and learning the reverse process, samples can be generated by denoising from the uniform prior\.
### 2\.1Forward Process
Continuous\-time discrete diffusion models corrupt the data distributionp0\(𝐱\)p\_\{0\}\(\\mathbf\{x\}\)into a noise prior through a forward Markov process\. In masked diffusion models, the terminal prior is a point mass on the mask token\. In uniform\-state diffusion models, asttevolves from00to11, a clean token𝐱0\\mathbf\{x\}\_\{0\}is gradually perturbed toward the uniform distribution over the vocabulary𝒱\\mathcal\{V\}, where\|𝒱\|=V\|\\mathcal\{V\}\|=V\. Over an infinitesimal interval fromtttot\+Δtt\+\\Delta t, the transition probability of𝐱t\\mathbf\{x\}\_\{t\}\(ignoringo\(Δt\)o\(\\Delta t\)terms\) is
p\(𝐱t\+Δt∣𝐱t\)=δ\(𝐱t,𝐱t\+Δt\)\+Qt\(𝐱t,𝐱t\+Δt\)Δt\.p\(\\mathbf\{x\}\_\{t\+\\Delta t\}\\mid\\mathbf\{x\}\_\{t\}\)=\\delta\(\\mathbf\{x\}\_\{t\},\\mathbf\{x\}\_\{t\+\\Delta t\}\)\+Q\_\{t\}\(\\mathbf\{x\}\_\{t\},\\mathbf\{x\}\_\{t\+\\Delta t\}\)\\Delta t\.\(1\)Here,𝐱0\\mathbf\{x\}\_\{0\}and𝐱t\\mathbf\{x\}\_\{t\}are one\-hot vectors over𝒱\\mathcal\{V\},⊙\\odotdenotes the Hadamard product,δ\(𝐱t,𝐲\)=𝕀𝐲=𝐱t\\delta\(\\mathbf\{x\}\_\{t\},\\mathbf\{y\}\)=\\mathbb\{I\}\_\{\\mathbf\{y\}=\\mathbf\{x\}\_\{t\}\}, andQt∈ℝV×VQ\_\{t\}\\in\\mathbb\{R\}^\{V\\times V\}denotes the transition rate matrix\. For uniform\-state diffusion andt∈\[0,1\)t\\in\[0,1\), the transition rate matrix is given byQt\(𝐞i,𝐞j\)=−αt′αt\(1V−𝕀i=j\)Q\_\{t\}\(\\mathbf\{e\}\_\{i\},\\mathbf\{e\}\_\{j\}\)=\-\\frac\{\\alpha\_\{t\}^\{\\prime\}\}\{\\alpha\_\{t\}\}\\left\(\\frac\{1\}\{V\}\-\\mathbb\{I\}\_\{i=j\}\\right\), where the noise schedulet↦αtt\\mapsto\\alpha\_\{t\}is continuously differentiable and non\-increasing, withα0=1\\alpha\_\{0\}=1andα1=0\\alpha\_\{1\}=0\. A common choice isαt=1−t\\alpha\_\{t\}=1\-t\. This forward process yields the marginal distribution
p\(𝐱t∣𝐱0\)=Cat\(⋅,αt𝐱0\+\(1−αt\)𝟏V\)\.p\(\\mathbf\{x\}\_\{t\}\\mid\\mathbf\{x\}\_\{0\}\)=\\mathrm\{Cat\}\\left\(\\cdot;\\,\\alpha\_\{t\}\\mathbf\{x\}\_\{0\}\+\(1\-\\alpha\_\{t\}\)\\frac\{\\mathbf\{1\}\}\{V\}\\right\)\.\(2\)
Letαt\|s:=αt/αs\\alpha\_\{t\\mid s\}:=\\alpha\_\{t\}/\\alpha\_\{s\}and define𝐱¯:=Vαt𝐱\+\(1−αt\)𝟏\\bar\{\\mathbf\{x\}\}:=V\\alpha\_\{t\}\\mathbf\{x\}\+\(1\-\\alpha\_\{t\}\)\\mathbf\{1\}\. The exact reverse posteriorp\(𝐱s∣𝐱t,𝐱0\)=Cat\(⋅,𝝅t→s\)p\(\\mathbf\{x\}\_\{s\}\\mid\\mathbf\{x\}\_\{t\},\\mathbf\{x\}\_\{0\}\)=\\mathrm\{Cat\}\\left\(\\cdot;\\,\\bm\{\\pi\}\_\{t\\to s\}\\right\)can be written in closed form:
\\fitbox0\.98𝝅t→s=1⟨𝐱¯t,𝐱0⟩\[Vαt𝐱t⊙𝐱0\+\(αt\|s−αt\)𝐱t\+\(αs−αt\)𝐱0\+\(1−αt\|s\)\(1−αs\)𝟏V\]\.\\fitbox\{0\.98\}\{\\displaystyle\\bm\{\\pi\}\_\{t\\to s\}=\\frac\{1\}\{\\langle\\bar\{\\mathbf\{x\}\}\_\{t\},\\mathbf\{x\}\_\{0\}\\rangle\}\\Big\[V\\alpha\_\{t\}\\,\\mathbf\{x\}\_\{t\}\\odot\\mathbf\{x\}\_\{0\}\+\(\\alpha\_\{t\\mid s\}\-\\alpha\_\{t\}\)\\mathbf\{x\}\_\{t\}\+\(\\alpha\_\{s\}\-\\alpha\_\{t\}\)\\mathbf\{x\}\_\{0\}\+\(1\-\\alpha\_\{t\\mid s\}\)\(1\-\\alpha\_\{s\}\)\\frac\{\\mathbf\{1\}\}\{V\}\\Big\]\.\}\(3\)
We provide a more detailed derivation in Appendix[A\.1\.2](https://arxiv.org/html/2609.35817#A1.SS1.SSS2)\.
### 2\.2Reverse Process and Training Objective
To sample fromp0\(𝐱\)p\_\{0\}\(\\mathbf\{x\}\), one learns a reverse denoising process from the noise prior\. This requires parameterizing the reverse transitionpθ\(𝐱t−Δt∣𝐱t\)p\_\{\\theta\}\(\\mathbf\{x\}\_\{t\-\\Delta t\}\\mid\\mathbf\{x\}\_\{t\}\)and matching it to the exact posteriorp\(𝐱t−Δt∣𝐱t,𝐱0\)p\(\\mathbf\{x\}\_\{t\-\\Delta t\}\\mid\\mathbf\{x\}\_\{t\},\\mathbf\{x\}\_\{0\}\)\. The standard diffusion objective minimizes the local KL divergence
KL\(p\(𝐱t−Δt∣𝐱t,𝐱0\)∥pθ\(𝐱t−Δt∣𝐱t\)\),\\mathrm\{KL\}\\left\(p\(\\mathbf\{x\}\_\{t\-\\Delta t\}\\mid\\mathbf\{x\}\_\{t\},\\mathbf\{x\}\_\{0\}\)\\,\\\|\\,p\_\{\\theta\}\(\\mathbf\{x\}\_\{t\-\\Delta t\}\\mid\\mathbf\{x\}\_\{t\}\)\\right\),\(4\)which forms a local term in the negative evidence lower bound \(NELBO\)\.
SEDD\([Lou et al\., 2023](https://arxiv.org/html/2609.35817#bib.bib9)\)parameterizespθ\(𝐱t−Δt∣𝐱t\)p\_\{\\theta\}\(\\mathbf\{x\}\_\{t\-\\Delta t\}\\mid\\mathbf\{x\}\_\{t\}\)through rate ratios\. Duo\([Sahoo et al\., 2025](https://arxiv.org/html/2609.35817#bib.bib4)\)instead parameterizes the reverse transition aspθ\(𝐱t−Δt∣𝐱t\)=p\(𝐱t−Δt∣𝐱t,𝐱θ\(𝐱t,t\)\)p\_\{\\theta\}\(\\mathbf\{x\}\_\{t\-\\Delta t\}\\mid\\mathbf\{x\}\_\{t\}\)=p\(\\mathbf\{x\}\_\{t\-\\Delta t\}\\mid\\mathbf\{x\}\_\{t\},\\mathbf\{x\}\_\{\\theta\}\(\\mathbf\{x\}\_\{t\},t\)\), that is, the model directly predicts𝐱0\\mathbf\{x\}\_\{0\}\. Both approaches optimize the KL objective in Eq\. \([4](https://arxiv.org/html/2609.35817#S2.E4)\)\. Letm=argmaxj\(𝐱t\)jm=\\arg\\max\_\{j\}\(\\mathbf\{x\}\_\{t\}\)\_\{j\}denote the current noisy token andy=argmaxj\(𝐱0\)jy=\\arg\\max\_\{j\}\(\\mathbf\{x\}\_\{0\}\)\_\{j\}denote the clean token\. Under the Duo parameterization, the token\-level lossℒDuoℓ\\mathcal\{L\}\_\{\\mathrm\{Duo\}\}^\{\\ell\}at positionℓ\\ellis
𝔼t,𝐱t−αt′Vαt\[Vx¯θ,m−Vx¯0,m\+∑j=1Vx¯0,jx¯0,mlogx¯θ,mx¯0,jx¯θ,jx¯0,m\]\.\\mathbb\{E\}\_\{t,\\mathbf\{x\}\_\{t\}\}\\frac\{\-\\alpha\_\{t\}^\{\\prime\}\}\{V\\alpha\_\{t\}\}\\\!\\left\[\\\!\\frac\{V\}\{\\bar\{x\}\_\{\\theta,m\}\}\\\!\-\\\!\\frac\{V\}\{\\bar\{x\}\_\{0,m\}\}\\\!\+\\\!\\sum\_\{j=1\}^\{V\}\\\!\\frac\{\\bar\{x\}\_\{0,j\}\}\{\\bar\{x\}\_\{0,m\}\}\\\!\\log\\frac\{\\bar\{x\}\_\{\\theta,m\}\\bar\{x\}\_\{0,j\}\}\{\\bar\{x\}\_\{\\theta,j\}\\bar\{x\}\_\{0,m\}\}\\\!\\right\]\\\!\.\(5\)Summing over all positionsℓ\\ellgives the sentence\-level training objective\.
Appendix[A\.1\.3](https://arxiv.org/html/2609.35817#A1.SS1.SSS3)shows that substituting the Duox0x\_\{0\}\-parameterization into the uniform\-state SEDD objective yields Eq\. \([5](https://arxiv.org/html/2609.35817#S2.E5)\) exactly at the generator level\.
SDDLM\([Zhu et al\., 2025c](https://arxiv.org/html/2609.35817#bib.bib5)\)follows the same parameterization as Duo, but introduces empirically simplified objectives for more stable and efficient training\. The SDDLM and SDDLM\-v1 losses are:
\\fitbox0\.98ℒSDℓ=−𝔼t,𝐱t𝕀𝐱t≠𝐱0logxθ,y,ℒSD1ℓ=−𝔼t,𝐱t𝕀𝐱t≠𝐱0\[logxθ,y−1V∑j=1Vlogxθ,j\]\.\\fitbox\{0\.98\}\{\\displaystyle\\mathcal\{L\}^\{\\ell\}\_\{\\mathrm\{SD\}\}=\-\\mathbb\{E\}\_\{t,\\mathbf\{x\}\_\{t\}\}\\mathbb\{I\}\_\{\\mathbf\{x\}\_\{t\}\\neq\\mathbf\{x\}\_\{0\}\}\\log x\_\{\\theta,y\},\\qquad\\mathcal\{L\}^\{\\ell\}\_\{\\mathrm\{SD1\}\}=\-\\mathbb\{E\}\_\{t,\\mathbf\{x\}\_\{t\}\}\\mathbb\{I\}\_\{\\mathbf\{x\}\_\{t\}\\neq\\mathbf\{x\}\_\{0\}\}\\left\[\\log x\_\{\\theta,y\}\-\\frac\{1\}\{V\}\\sum\_\{j=1\}^\{V\}\\log x\_\{\\theta,j\}\\right\]\.\}\(6\)UnlikeℒDuo\\mathcal\{L\}\_\{\\mathrm\{Duo\}\}, which corresponds to a NELBO objective, these simplified losses are no longer NELBOs in theory, but have been observed to yield stronger empirical performance\.
## 3Methodology
In this section, we address the over\-uniform and condition\-target confusion that prevent UDLMs from scaling to complex tasks\. The lack of a clean supervisory signal forces the model to rely on many denoising steps during sampling, yet complex tasks demand that simpler tokens be decoded quickly and serve as conditioning context\. To resolve this tension, we first analyze the source of over\-uniform and mitigate it with the LU loss \(Sec\.[3\.1](https://arxiv.org/html/2609.35817#S3.SS1)\), then introduce per\-token time embeddings that provide token\-level corruption hints \(Sec\.[3\.2](https://arxiv.org/html/2609.35817#S3.SS2)\)\. Together, the LU loss endows LUDI with stronger few\-step generation capability, enabling it to rapidly decode high\-confidence tokens and update their per\-token time embeddings as conditions, thereby handling complex generation tasks effectively\.
### 3\.1Less Uniform Loss
##### Over\-uniform phenomenon\.
In MDLMs, training based on ELBO has achieved widespread success, being used for large\-scale pretraining or AR\-to\-diffusion training\. The ELBO objective in UDLMs, while theoretically well grounded, is substantially less effective in practice\. We attribute this to themode\-coveringproperty of the KL loss, which drives the model toward over\-uniform and drowns out meaningful supervision\.
To elucidate this issue, we simplify the complex original loss in Eq\.\([5](https://arxiv.org/html/2609.35817#S2.E5)\) into a CE\-like form, as stated in the following proposition\.
Proposition 1\(Proved in Appendix[A\.2\.1](https://arxiv.org/html/2609.35817#A1.SS2.SSS1)\)\. Under mild assumptions detailed there, at a corrupted positionm≠ym\\neq y, the ELBO integrand, up to an additive term independent ofθ\\thetaand ano\(1\)o\(1\)remainder asV→∞V\\to\\infty, is\\fitbox−αt′\[logx¯θ,m−logx¯θ,y1−αt−1Vαt∑i=1Vlogx¯θ,i\]=−αt′\[CELS1−αt\(𝐱¯θ,y\)αt\(1−αt\)\+logx¯θ,m1−αt\]\.\\fitbox\{\}\{\\displaystyle\-\\alpha\_\{t\}^\{\\prime\}\\\!\\left\[\\frac\{\\hbox\{\\pagecolor\{blue\!10\!white\}$\\log\\bar\{x\}\_\{\\theta,m\}$\}\\hbox\{\\pagecolor\{DarkGreen\!12\!white\}$\-\\log\\bar\{x\}\_\{\\theta,y\}$\}\}\{1\-\\alpha\_\{t\}\}\\\!\-\\\!\\frac\{1\}\{V\\alpha\_\{t\}\}\\sum\_\{i=1\}^\{V\}\\hbox\{\\pagecolor\{red\!8\!white\}$\\log\\bar\{x\}\_\{\\theta,i\}$\}\\right\]=\-\\alpha\_\{t\}^\{\\prime\}\\\!\\left\[\\frac\{\\mathrm\{CE\}\_\{\\mathrm\{LS\}\}^\{1\-\\alpha\_\{t\}\}\(\\bar\{\\mathbf\{x\}\}\_\{\\theta\},y\)\}\{\\alpha\_\{t\}\(1\-\\alpha\_\{t\}\)\}\+\\frac\{\\log\\bar\{x\}\_\{\\theta,m\}\}\{1\-\\alpha\_\{t\}\}\\right\]\.\}\(7\)For already\-clean positionsm=ym=y, ifxθ,y=ω\(1/V\)x\_\{\\theta,y\}=\\omega\(1/V\), the model\-dependent contribution iso\(1\)o\(1\)and therefore negligible asV→∞V\\to\\infty\.
When𝐱ti≠𝐱0i\\mathbf\{x\}\_\{t\}^\{i\}\\neq\\mathbf\{x\}\_\{0\}^\{i\}, the loss decomposes into three terms:
\(1\) suppression of the current erroneous tokenlogx¯θ,m\\log\\bar\{x\}\_\{\\theta,m\}; \(2\) learning of the correct token−logx¯θ,y\-\\log\\bar\{x\}\_\{\\theta,y\};
\(3\) a smoothing term over the uniform distribution−∑ilogx¯θ,i\-\\sum\_\{i\}\\log\\bar\{x\}\_\{\\theta,i\}\.
Notably, \(2\) and \(3\) can form a label\-smoothed cross\-entropy objectiveCELS\\mathrm\{CE\}\_\{\\mathrm\{LS\}\}with effective targetαt𝐱0\+\(1−αt\)𝟏/V\\alpha\_\{t\}\\mathbf\{x\}\_\{0\}\+\(1\-\\alpha\_\{t\}\)\\mathbf\{1\}/V\. The smoothing strength1−αt1\-\\alpha\_\{t\}is unusually large\. Under the common scheduleαt=1−t\\alpha\_\{t\}=1\-twitht∼𝒰\(0,1\)t\\sim\\mathcal\{U\}\(0,1\), on average only half of the probability mass is placed on clean token𝐱0\\mathbf\{x\}\_\{0\}\.
This stems from the mode\-covering nature of the KL divergence\. Wheneverp\(𝐱t−Δt∣𝐱t,𝐱0\)p\(\\mathbf\{x\}\_\{t\-\\Delta t\}\\mid\\mathbf\{x\}\_\{t\},\\mathbf\{x\}\_\{0\}\)assigns probability to random tokens, the learned reverse transition kernel must cover those probabilities, otherwise it incurs a large penalty\. Consequently, whenttis large, the model receives little effective supervision toward the clean token and is instead driven toward overly uniform predictions\.
Remark 1\.This problem is circumvented in mask diffusion\. Although mask diffusion also employs a KL\-based ELBO, the loss decouples intott\-dependent andtt\-independent terms\. In MDLMs, the probability of a token staying masked is computed analytically rather than learned\.
Remark 2\.While SD and SDv1 in Eq\. \([6](https://arxiv.org/html/2609.35817#S2.E6)\) heuristically remove the smoothing term−∑ilogx¯θ,i\-\\sum\_\{i\}\\log\\bar\{x\}\_\{\\theta,i\}and even introduce anti\-smoothing, they also discard the suppression losslogx¯θ,m\\log\\bar\{x\}\_\{\\theta,m\}on the erroneous token, thereby weakening the model’s correction capability and lacking theoretical support\.
##### LU loss\.
To mitigate over\-uniform, a simple fix is to directly remove term−∑ilogx¯θ,i\-\\sum\_\{i\}\\log\\bar\{x\}\_\{\\theta,i\}\. We obtain a*less uniform*loss, termed the LU loss:
ℒLUℓ=𝔼t,𝐱t−αt′1−αt\(logx¯θ,m−logx¯θ,y\)\.\\mathcal\{L\}\_\{\\mathrm\{LU\}\}^\{\\ell\}\\\!=\\\!\\mathbb\{E\}\_\{t,\\mathbf\{x\}\_\{t\}\}\\frac\{\-\\alpha\_\{t\}^\{\\prime\}\}\{1\-\\alpha\_\{t\}\}\\left\(\\log\\bar\{x\}\_\{\\theta,m\}\\\!\-\\\!\\log\\bar\{x\}\_\{\\theta,y\}\\right\)\.\(8\)Despite its simplicity, the LU loss has a direct reverse\-process interpretation\. At corrupted positions, it minimizes the finite model\-dependent part of a Dirac\-target KL\.
Theorem 1\(Proved in Appendix[A\.2\.2](https://arxiv.org/html/2609.35817#A1.SS2.SSS2)\)\. At a corrupted positionm≠ym\\neq y, asΔt→0\\Delta t\\to 0,KL\(δ\(𝐱0\)∥pθ\(𝐱t−Δt∣𝐱t\)\)=−logαt−Δt−αtVαt−Δt\+logx¯θ,m−logx¯θ,y\+o\(1\)\.\\mathrm\{KL\}\\\!\\bigl\(\\delta\(\\mathbf\{x\}\_\{0\}\)\\,\\\|\\,p\_\{\\theta\}\(\\mathbf\{x\}\_\{t\-\\Delta t\}\\\!\\mid\\\!\\mathbf\{x\}\_\{t\}\)\\bigr\)=\-\\log\\\!\\frac\{\\alpha\_\{t\-\\Delta t\}\-\\alpha\_\{t\}\}\{V\\alpha\_\{t\-\\Delta t\}\}\+\\log\\bar\{x\}\_\{\\theta,m\}\-\\log\\bar\{x\}\_\{\\theta,y\}\+o\(1\)\.\(9\)The first term is independent ofθ\\theta\. Therefore, minimizing the finite model\-dependent part of this KL is equivalent to minimizingℒLUℓ\\mathcal\{L\}\_\{\\mathrm\{LU\}\}^\{\\ell\}\.
The LU loss simultaneously reduces the probability of the current incorrect token and increases the probability of the clean token𝐱0\\mathbf\{x\}\_\{0\}\. Moreover, the use of𝐱¯θ=Vαt𝐱θ\+\(1−αt\)𝟏\\bar\{\\mathbf\{x\}\}\_\{\\theta\}=V\\alpha\_\{t\}\\mathbf\{x\}\_\{\\theta\}\+\(1\-\\alpha\_\{t\}\)\\mathbf\{1\}prevents the logarithmic terms from becoming singular\. As a result, the learned transition is consistently encouraged to move toward𝐱0\\mathbf\{x\}\_\{0\}for any time step\. Intuitively, an ELBO\-trained model behaves more like adiffusion model, whereas the LU loss aligns more closely with aconsistency model, as illustrated in Figure[2](https://arxiv.org/html/2609.35817#S1.F2)[2\(a\)](https://arxiv.org/html/2609.35817#S1.F2.sf1)\. Our LUDI loss offers two key advantages:
1. 1\.LU loss avoids the excessive smoothing induced by the KL objective and instead aligns the parameterized transition rate with the clean token\. This provides a stronger learning signal and improves few\-step generation through consistency\.
2. 2\.Compared with the ELBO\-based loss, our formulation involves only the two indicesyyandmm, making it computationally efficient\. With optimized operators, LUDI can achieve a4\.39×4\.39\\timesspeedup and a3\.45×3\.45\\timesmemory reduction\. The ELBO loss additionally applies a nonlinear transformation and reduction over all vocabulary entries, which is expensive for large vocabularies\. \(A detailed efficiency analysis is deferred to the Appendix[C\.5](https://arxiv.org/html/2609.35817#A3.SS5)\.\)
### 3\.2Scaling up UDLM
##### Condition\-target confusion\.
When scaling UDLMs to the billion\-parameter regime for challenging downstream tasks such as mathematics and code generation, a central obstacle is the ambiguity between conditioning context and denoising targets\. Difficult reasoning tasks require the model to first establish a reliable condition, such as a mathematical equation, and then infer the target, such as the solution\. In masked diffusion, this separation is explicit: clean tokens serve as conditions, while mask tokens identify the denoising targets\. In UDLMs, however, the model is given only the global timett, which indicates the overall corruption level but not the status of each individual token\. The model therefore struggles to distinguish condition tokens from target tokens and tends to denoise all positions simultaneously in a condition\-independent manner\. This behavior is particularly fatal for complex reasoning\.
##### Per\-token time embedding\.
To address this ambiguity, we introduce per\-token time embeddings\. The denoiser is written as𝐱θ\(𝐱t,𝝉\)\\mathbf\{x\}\_\{\\theta\}\(\\mathbf\{x\}\_\{t\},\\bm\{\\tau\}\), where𝝉=\(τ1,…,τL\)\\bm\{\\tau\}=\(\\tau^\{1\},\\ldots,\\tau^\{L\}\)assigns a distinct time value to each position, as shown in Figure[2](https://arxiv.org/html/2609.35817#S1.F2)[2\(b\)](https://arxiv.org/html/2609.35817#S1.F2.sf2)\. Eachτi\\tau^\{i\}indicates the probability that tokeniihas been randomly replaced, thereby providing the model with a token\-level hint as to whether the token should be treated as condition or target\.
To preserve the marginal distribution of𝐱t\\mathbf\{x\}\_\{t\}, we use a hierarchical sampling procedure in the forward process\. We first sample a global timet∼𝒰\(0,1\)t\\sim\\mathcal\{U\}\(0,1\)\. Then, for each positionii, we sample a per\-token timeτi∼qt\\tau^\{i\}\\sim q\_\{t\}and set the token to its clean value with probability1−τi1\-\\tau^\{i\}, or to a random vocabulary token with probabilityτi\\tau^\{i\}\. As long asqtq\_\{t\}has support on\[0,1\]\[0,1\]and satisfies𝔼qt\[τi\]=t\\mathbb\{E\}\_\{q\_\{t\}\}\[\\tau^\{i\}\]=t, the marginal distributionp\(𝐱t∣t\)p\(\\mathbf\{x\}\_\{t\}\\mid t\)remains unchanged\.
In practice, we chooseqtq\_\{t\}to be a Beta distributionBeta\(ct,c\(1−t\)\)\\mathrm\{Beta\}\(ct,c\(1\-t\)\)withc=2c=2, and inject the time signal via AdaLN\([Peebles and Xie, 2023](https://arxiv.org/html/2609.35817#bib.bib10)\)\. Note that the per\-tokenτi\\tau^\{i\}acts only as a hint; the LUDI loss itself still uses the globaltt\.
##### Conditional sampling\.
Benefiting from per\-token time embeddings, LUDI naturally supports position\-wise non\-uniform conditional sampling\. At initialization, non\-prompt tokens are sampled uniformly from the vocabulary and assignedτi∼Beta\(2,1\)\\tau^\{i\}\\sim\\mathrm\{Beta\}\(2,1\), while prompt tokens are kept fixed withτi=0\\tau^\{i\}=0\. During denoising, non\-prompt tokens are progressively updated until their token\-level timesτi\\tau^\{i\}approach zero\.
Given the current token\-level timeτti\\tau^\{i\}\_\{t\}, the next valueτt−Δti\\tau^\{i\}\_\{t\-\\Delta t\}need not follow a uniform schedule and can instead be chosen adaptively for each position\. For example, a high\-confidence token can be decoded in a single step by settingτt−Δti=0\\tau^\{i\}\_\{t\-\\Delta t\}=0, after which it is treated as a conditioning token for subsequent generation\. Consequently, LUDI can directly inherit samplers developed for MDLMs: each position either remains unchanged withτt−Δti=τti\\tau^\{i\}\_\{t\-\\Delta t\}=\\tau^\{i\}\_\{t\}or is committed by settingτt−Δti=0\\tau^\{i\}\_\{t\-\\Delta t\}=0\. Empirically, this easy\-to\-hard denoising schedule is critical for complex reasoning tasks\. We provide pseudocode for training and sampling in the Appendix[B](https://arxiv.org/html/2609.35817#A2)\.
##### AR to block\-UDLM\.
Training a large\-scale UDLM from scratch is costly\. We therefore initialize from an autoregressive \(AR\) model to maximally preserve acquired knowledge\. To continue\-training the AR model as a block\-UDLM, we adopt two adaptation strategies\.
Label shifting and complementary noise\(following Fast dLLMv2\([Wu et al\., 2025a](https://arxiv.org/html/2609.35817#bib.bib12)\)\)\. Since an AR model predicts tokenxi\+1x\_\{i\+1\}from positionxix\_\{i\}, we shift the labels left by one position\. We also concatenate samples noised atttand1−t1\-tduring training to reduce gradient variance\.
Context\-causal attention and AR loss\(following NB\-Diff\([Tian et al\., 2025](https://arxiv.org/html/2609.35817#bib.bib11)\)\)\. We employ context\-causal attention, where only the block currently being denoised receives bidirectional attention, while all preceding context retains causal attention\. This design provides cleaner contextual signals for block generation and naturally enables inter\-block KV\-caching\. We further regularize training with an auxiliary next\-token prediction lossℒAR\\mathcal\{L\}\_\{\\mathrm\{AR\}\}, which prevents the model from drifting excessively from its AR initialization\. The overall objective is thereforeℒLUDI=ℒLU\+λℒAR\\mathcal\{L\}\_\{\\mathrm\{LUDI\}\}=\\mathcal\{L\}\_\{\\mathrm\{LU\}\}\+\\lambda\\,\\mathcal\{L\}\_\{\\mathrm\{AR\}\}, withλ\\lambdatypically set to0\.50\.5to supply supervision to the diffusion and AR components with comparable strength\.
By combining these components, we present LUDI\-7B, a 7B\-scale UDLM capable of tackling complex reasoning tasks such as math and code generation, marking the transition of the UDLM paradigm from a conceptual idea to practical usability\. Further design details are provided in the Appendix[B\.2](https://arxiv.org/html/2609.35817#A2.SS2)\.
## 4Experiments
We conduct experiments at two scales: small\-scale pretraining \(170M and 1B; Sec\.[4\.1](https://arxiv.org/html/2609.35817#S4.SS1)\) and large\-scale AR\-to\-UDLM continue\-training \(LUDI\-7B; Sec\.[4\.2](https://arxiv.org/html/2609.35817#S4.SS2)\)\. In the small\-scale setting, where per\-token time embeddings are not essential, we focus on the generation quality and training efficiency of the LU loss\. In the large\-scale setting, we evaluate LUDI\-7B on complex reasoning tasks\.
### 4\.1Small\-Scale LUDI
##### Experimental setup\.
For the 170M experiments, we train four UDLMs under the same configuration, using LU loss, Duo loss, SDDLM loss, and SDDLM\-v1 loss, respectively\. All models are trained on OpenWebText\([Gokaslan et al\., 2019](https://arxiv.org/html/2609.35817#bib.bib27)\)with a sequence length of 1024, using the GPT\-2 tokenizer\([Radford et al\., 2019](https://arxiv.org/html/2609.35817#bib.bib29)\)\. We use a batch size of 512 and a learning rate of3×10−43\\times 10^\{\-4\}\. For masked\-diffusion baselines, we use pretrained RADD and MDLM models, trained on the same dataset for 400K and 1M steps, respectively\. Generative perplexity is computed by GPT\-2 Large on 512 samples\. For the 1B experiments, we evaluate four UDLM models on downstream likelihood\-based tasks\. The models are trained on FineWeb\([Penedo et al\., 2024](https://arxiv.org/html/2609.35817#bib.bib28)\)for 500K steps using the Llama tokenizer\([Touvron et al\., 2023](https://arxiv.org/html/2609.35817#bib.bib30)\)\. The default sequence length is 2048, with 1% variable\-length data\. We use a batch size of 256 and a learning rate of2×10−42\\times 10^\{\-4\}\. Additional experimental details are provided in Appendix[C\.1](https://arxiv.org/html/2609.35817#A3.SS1)\.
##### Generative performance\.
Figure[3](https://arxiv.org/html/2609.35817#S4.F3)compares generation quality under different numbers of sampling steps\. With 1024 sampling steps, LUDI achieves a generative perplexity of 37\.76 and an entropy of 7\.50, outperforming both UDLM and MDLM baselines\. Across most sampling budgets \(≥64\\geq 64steps\), LU loss consistently yields stronger generation quality\. Although LUDI yields lower entropy, our analysis of Gen\.PPL and diversity \(see the Appendix[C\.7](https://arxiv.org/html/2609.35817#A3.SS7)\) indicates that this reduction does not stem from mode collapse\. UDLMs are often expected to enable few\-step generation, but their quality can suffer from early saturation\([Deschenaux et al\., 2026](https://arxiv.org/html/2609.35817#bib.bib22)\): increasing the number of sampling steps does not necessarily improve generation\. Our experiments confirm this behavior\. UDLMs outperform MDLMs in low\-step regimes, yet SDDLM\-v1 and Duo rapidly reach a plateau\. In contrast, LUDI continues to improve as the sampling budget increases\. We attribute this behavior to the stronger correction ability induced by LU loss\. Compared with Duo loss, which contains a smoothing term, and SDDLM\-style losses, which lack a correction term, LU loss more directly encourages the model to revise potentially erroneous tokens throughout sampling\.
Figure 3:Gen\.PPL from 32 to 1024 sampling steps\.
Table 1:Generation results \(Gen\.PPL and Entropy\) at 200K and 500K checkpoints\. All models sampled with 1024 steps\. Lower Gen\.PPL and higher entropy are better\.
Table 2:Zero\-shot downstream evaluation of 1B models at 200k and 500k pretraining checkpoints\. Full evaluation details are provided in Appendix[C\.1](https://arxiv.org/html/2609.35817#A3.SS1)\.
##### Downstream task performance\.
Table[2](https://arxiv.org/html/2609.35817#S4.T2)reports downstream results for 1B models on likelihood\-based tasks, which probe the language modeling capability\. LUDI attains the highest average score and leads on most tasks, indicating both superior performance and stability\. This advantage is consistent with the behavior of LU loss\. By reducing confidence on erroneous samples and increasing confidence on correct ones, LUDI enlarges the relative margin of the correct answer\.
Table 3:LUDI\-7B results and ablations\. Metrics cover code generation \(HumanEval, MBPP, HumanEval\+, MBPP\+\), mathematical reasoning \(GSM8K\), instruction following \(IFEval\), knowledge\-intensive question answering \(MMLU, GPQA\), and an overall average score \(Avg\.\)\. Code benchmarks are evaluated using EvalPlus\. The highest score in each column is marked inbold, and the second highest isunderlined\. AR→\\toM/U denotes MDLM/UDLMs continue\-trained from AR\. In ablation, LUDI is based on Fast dLLM v2 \(FD\) and incorporates the LU loss, per\-token time embedding, and NBDiff \(ND\) techniques\. Full evaluation details are provided in Appendix[C\.2](https://arxiv.org/html/2609.35817#A3.SS2)\.\\fitbox
##### Training efficiency\.
Table[1](https://arxiv.org/html/2609.35817#S4.T1)and Table[2](https://arxiv.org/html/2609.35817#S4.T2)also compare models at 200K training steps\. LUDI rapidly establishes clear advantages: the 170M model already exhibits stable generation, and the 1B LUDI outperforms all baselines across downstream tasks\. We ascribe this to the cleaner learning objective provided by the LU loss\. By avoiding the wasted supervision of label smoothing and directly matching𝐱0\\mathbf\{x\}\_\{0\}, the LU loss offers a consistent signal: the model always learns to move away from erroneous tokens and toward the correct token\. This improves training efficiency and data utilization, which is critical for large\-scale training\.
### 4\.2LUDI\-7B
##### Experimental setup\.
We initialize our block UDLM from Qwen2\.5\-7B\-Instruct[Team \(2024\)](https://arxiv.org/html/2609.35817#bib.bib25)and continue\-train it on Dolci\-Instruct\-SFT[Olmo et al\. \(2025\)](https://arxiv.org/html/2609.35817#bib.bib26), a high\-quality instruction dataset containing 2M samples\. LUDI\-7B is trained for 2500 steps with a batch size of 256 and a learning rate of1×10−51\\times 10^\{\-5\}, following the protocol of Fast dLLMv2[Wu et al\. \(2025a\)](https://arxiv.org/html/2609.35817#bib.bib12)\. The block size is fixed to 32\. In each transformer layer, we insert a learnable, zero\-initialized AdaLN module before the multi\-head self\-attention and the MLP\. The AdaLN output at positioniiis further multiplied byτi\\tau^\{i\}\. Consequently, the module is inactive whereverτi=0\\tau^\{i\}=0, including prompt positions and tokens already committed during sampling\. These AdaLN modules together with the time embeddings introduce only 78\.7M additional parameters \(∼\\sim1% of the total\)\. During sampling, we adopt confidence\-based parallel decoding\.
\\FloatBarrier\{wrapfigure\}
\[16\]tl0\.49
GSM8K performance and decoding speedup of LUDI\-7B under different decoding thresholds\.
##### Performance and speed\.
We compare LUDI with representative 7 to 8B autoregressive and diffusion models; the results are summarized in Table[3](https://arxiv.org/html/2609.35817#S4.T3)\. LUDI\-7B achieves an average score of 65\.3, outperforming both the baselines and its AR initialization\. It performs strongly on likelihood\-based benchmarks \(MMLU and GPQA\) while also delivering competitive results on generation tasks\. Although our method underperforms on mathematical reasoning, this mainly stems from the distributional bias of mathematical content in the training dataset\. In Appendix[C\.3](https://arxiv.org/html/2609.35817#A3.SS3), we present additional evaluation results on mathematical reasoning and demonstrate that the mathematical reasoning performance of LUDI can be significantly improved through dataset curation\.
With parallel decoding, LUDI\-7B offers an attractive performance and efficiency trade\-off\. By adjusting the confidence threshold, LUDI\-7B maintains stable accuracy on GSM8K \(Figure[4\.2](https://arxiv.org/html/2609.35817#S4.SS2.SSS0.Px1)\) for thresholds above 0\.7, while yielding a 2\.6 to 3\.0 token\-per\-step speedup over AR decoding \(1 token per step\)\. These results confirm that UDLMs can handle complex generation tasks and possess substantial acceleration potential\.
It is worth noting that the token\-per\-step metric represents a theoretical speedup; practical speedups are constrained by batch sampling and infrastructure overhead\. In Appendix[C\.4](https://arxiv.org/html/2609.35817#A3.SS4), we report end\-to\-end efficiency comparisons across different batch sizes\. At small batch sizes, we achieve actual speedups of 1\.30–1\.92×\\times\. We acknowledge that achieving speedup at larger batch sizes requires further dLLM\-specific infrastructure optimizations, which is orthogonal to our scope and is left as valuable future work\.
##### Ablation study\.
To systematically examine the contribution of each design choice, we conduct controlled experiments on Dolci\-Instruct under identical conditions \(Table[3](https://arxiv.org/html/2609.35817#S4.T3)\)\. Our LUDI\-7B builds upon Fast dLLM \(FD\) and integrates the LU loss, per\-token time embeddings, and the tricks from NB\-Diff \(NB\)\. \(1\) Without per\-token time embeddings, the model suffers from severe condition\-target confusion caused by position\-uniform sampling, quickly collapsing to a trivial solution that repeats a single token\. We therefore exclude this variant from comparison\. \(2\) Replacing the LU loss with the SDv1 loss, which also mitigates over\-uniform, still yields reasonable generations but with a clear performance drop\. This suggests that complex generation requires explicit error correction: the model must be steered away from the current erroneous token\. \(3\) The AR loss and causal context attention from NB\-Diff prove important, indicating that block decoding in UDLMs benefits from higher\-quality context\. \(4\) Compared to MDLM baselines, LUDI\-7B surpasses FDNB trained under the same budget, demonstrating that UDLMs can profit from a more challenging training objective and may offer greater scaling potential under data\-constrained regimes\([von Rütte et al\., 2025b](https://arxiv.org/html/2609.35817#bib.bib6)\)\.
## 5Related Works
##### Masked diffusion language models\.
MDLMs have become a competitive generative paradigm against AR models\. Through large\-scale pretraining\([Nie et al\., 2025b](https://arxiv.org/html/2609.35817#bib.bib2);[Bie et al\., 2025](https://arxiv.org/html/2609.35817#bib.bib3)\), AR\-to\-diffusion adaptation\([Wu et al\., 2025b](https://arxiv.org/html/2609.35817#bib.bib13);[Tian et al\., 2025](https://arxiv.org/html/2609.35817#bib.bib11);[Liu et al\., 2025](https://arxiv.org/html/2609.35817#bib.bib14)\), and reinforcement\-learning post\-training\([Zhao et al\., 2025](https://arxiv.org/html/2609.35817#bib.bib15);[Zhong et al\., 2026](https://arxiv.org/html/2609.35817#bib.bib16);[Ni et al\., 2026](https://arxiv.org/html/2609.35817#bib.bib17)\), MDLMs have achieved performance comparable to frontier LLMs on complex reasoning tasks while offering substantial inference\-speed advantages\. To better handle challenging problems, MDLM sampling has also moved beyond position\-uniform unmasking toward strategic policies based on confidence\([Wu et al\., 2025b](https://arxiv.org/html/2609.35817#bib.bib13)\)and entropy\([Ben\-Hamu et al\., 2025](https://arxiv.org/html/2609.35817#bib.bib18)\)\. However, the binary mask/non\-mask state space makes it difficult for MDLMs to correct already generated tokens\. Moreover, the uninformative mask state naturally constrains parallel decoding\([Chen et al\., 2026](https://arxiv.org/html/2609.35817#bib.bib19)\)\. Consequently, changing the masked generation paradigm has become an important direction for raising the ceiling of diffusion language models\([Ding et al\., 2026](https://arxiv.org/html/2609.35817#bib.bib20);[von Rütte et al\., 2025a](https://arxiv.org/html/2609.35817#bib.bib21)\)\.
##### Uniform diffusion language models\.
UDLMs provide a complementary formulation where corrupted tokens are driven toward a uniform vocabulary distribution rather than an absorbing mask\. This gives UDLMs potential advantages in few\-step generation\([Sahoo et al\., 2025](https://arxiv.org/html/2609.35817#bib.bib4);[Deschenaux et al\., 2026](https://arxiv.org/html/2609.35817#bib.bib22)\), controllability\([Schiff et al\., 2025](https://arxiv.org/html/2609.35817#bib.bib7)\), self\-correction\([Schiff et al\., 2026](https://arxiv.org/html/2609.35817#bib.bib24)\), and scaling behavior under data\- or compute\-limited regimes\([von Rütte et al\., 2025b](https://arxiv.org/html/2609.35817#bib.bib6)\)\. Nevertheless, existing analyses remain largely confined to small\-scale settings or scaling\-law studies, and UDLMs have not yet been scaled for complex reasoning tasks\. Some works use uniform diffusion only as an auxiliary mechanism to improve the self\-correction ability of MDLMs\([Schiff et al\., 2026](https://arxiv.org/html/2609.35817#bib.bib24);[von Rütte et al\., 2025a](https://arxiv.org/html/2609.35817#bib.bib21)\), or as a bridge toward continuous diffusion language modeling\([Sahoo et al\., 2025](https://arxiv.org/html/2609.35817#bib.bib4);[Lee et al\., 2026](https://arxiv.org/html/2609.35817#bib.bib23)\)\. Efforts dedicated to improving UDLMs mainly focus on mitigating the sampling plateau with predictor\-corrector samplers\([Deschenaux et al\., 2026](https://arxiv.org/html/2609.35817#bib.bib22)\)or simplifying the training objective\([Zhu et al\., 2025c](https://arxiv.org/html/2609.35817#bib.bib5)\)\. However, these works either do not address, or only empirically touch upon, the over\-uniform effect induced by UDLM training objectives; meanwhile, their samplers are not specifically designed for complex reasoning and largely remain position\-uniform\. These limitations create an inherent obstacle to scaling UDLMs, motivating our study\.
## 6Conclusion
This work presents LUDI as a step toward making UDLMs practical at scale\. We show that the main barrier to scaling UDLMs lies not in the uniform noise prior itself, but in the excessive uniformity introduced by existing training objectives and sampling procedures\. By introducing a less uniform loss and token\-level time awareness, LUDI achieves cleaner supervision and a more efficient denoising process\. Through LUDI\-7B, we further demonstrate that UDLMs can be extended to complex reasoning tasks, narrowing the gap with AR and MDLM paradigms\. These results suggest that UDLMs remain a promising yet underexplored direction\.
## Limitations
This work presents LUDI as an initial effort to scale uniform diffusion language models to 7B parameters and to apply them to complex reasoning tasks\. Nevertheless, several limitations remain\. \(1\) We have not yet fully activated the self\-correction potential of UDLMs\. The errors that arise during self\-correction stem from the model’s own imperfect generations, which are qualitatively different from the random noise introduced during training\. Consequently, instilling this ability may require dedicated post\-training stages that explicitly teach the model to detect and amend its own mistakes\. \(2\) Due to computational resource constraints, we have not been able to pre\-train a large\-scale UDLM from scratch\. Our approach of continue\-training from an AR checkpoint efficiently preserves the original model’s knowledge, yet it may also retain certain inductive biases of the autoregressive generation paradigm\. We hope our work spurs future research into UDLMs pre\-trained at scale from scratch, which may more fully exploit the unique advantages of the uniform diffusion framework\.
## References
- Austinet al\.\(2021\)J\. Austin, A\. Odena, M\. Nye, M\. Bosma, H\. Michalewski, D\. Dohan, E\. Jiang, C\. Cai, M\. Terry, Q\. Le, and C\. SuttonProgram synthesis with large language models\.arXiv preprint arXiv:2108\.07732\.Cited by:[§C\.2](https://arxiv.org/html/2609.35817#A3.SS2.SSS0.Px2.p1.1)\.
- Ben\-Hamuet al\.\(2025\)H\. Ben\-Hamu, I\. Gat, D\. Severo, N\. Nolte, and B\. KarrerAccelerated sampling from masked diffusion models via entropy bounded unmasking\.External Links:2505\.24857Cited by:[§5](https://arxiv.org/html/2609.35817#S5.SS0.SSS0.Px1.p1.1)\.
- Bieet al\.\(2025\)T\. Bie, M\. Cao, K\. Chen, L\. Du, M\. Gong, Z\. Gong, Y\. Gu, J\. Hu, Z\. Huang, Z\. Lan, C\. Li, C\. Li, J\. Li, Z\. Li, H\. Liu, L\. Liu, G\. Lu, X\. Lu, Y\. Ma, J\. Tan,et al\.Llada2\. 0: scaling up diffusion language models to 100b\.arXiv preprint arXiv:2512\.15745\.Cited by:[§1](https://arxiv.org/html/2609.35817#S1.p1.1),[§5](https://arxiv.org/html/2609.35817#S5.SS0.SSS0.Px1.p1.1)\.
- Bisket al\.\(2020\)Y\. Bisk, R\. Zellers, R\. Le Bras, J\. Gao, and Y\. ChoiPiqa: reasoning about physical commonsense in natural language\.InProceedings of the AAAI conference on artificial intelligence,Vol\.34,pp\. 7432–7439\.Cited by:[§C\.1](https://arxiv.org/html/2609.35817#A3.SS1.SSS0.Px2.p1.1)\.
- Chenet al\.\(2021\)M\. Chen, J\. Tworek, H\. Jun, Q\. Yuan, H\. P\. D\. O\. Pinto, J\. Kaplan, H\. Edwards, Y\. Burda, N\. Joseph, G\. Brockman, A\. Ray, R\. Puri, G\. Krueger, M\. Petrov, H\. Khlaaf, G\. Sastry, P\. Mishkin, B\. Chan, S\. Gray, N\. Ryder,et al\.Evaluating large language models trained on code\.arXiv preprint arXiv:2107\.03374\.Cited by:[§C\.2](https://arxiv.org/html/2609.35817#A3.SS2.SSS0.Px2.p1.1)\.
- Chenet al\.\(2026\)Y\. Chen, C\. Liang, H\. Sui, R\. Guo, C\. Cheng, J\. You, and G\. LiuLangFlow: continuous diffusion rivals discrete in language modeling\.External Links:2604\.11748Cited by:[§5](https://arxiv.org/html/2609.35817#S5.SS0.SSS0.Px1.p1.1)\.
- Clarket al\.\(2018\)P\. Clark, I\. Cowhey, O\. Etzioni, T\. Khot, A\. Sabharwal, C\. Schoenick, and O\. TafjordThink you have solved question answering? try arc, the ai2 reasoning challenge\.arXiv preprint arXiv:1803\.05457\.Cited by:[§C\.1](https://arxiv.org/html/2609.35817#A3.SS1.SSS0.Px2.p1.1)\.
- Cobbeet al\.\(2021\)K\. Cobbe, V\. Kosaraju, M\. Bavarian, M\. Chen, H\. Jun, L\. Kaiser, M\. Plappert, J\. Tworek, J\. Hilton, R\. Nakano, C\. Hesse, and J\. SchulmanTraining verifiers to solve math word problems\.arXiv preprint arXiv:2110\.14168\.Cited by:[§C\.2](https://arxiv.org/html/2609.35817#A3.SS2.SSS0.Px2.p1.1)\.
- Deschenauxet al\.\(2026\)J\. Deschenaux, C\. Gulcehre, and S\. S\. SahooThe diffusion duality, chapter II:Ψ\\Psi\-samplers and efficient curriculum\.InThe Fourteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=RSIoYWIzaP)Cited by:[§4\.1](https://arxiv.org/html/2609.35817#S4.SS1.SSS0.Px2.p1.1),[§5](https://arxiv.org/html/2609.35817#S5.SS0.SSS0.Px2.p1.1)\.
- Dinget al\.\(2026\)F\. Ding, D\. Ding, S\. Chen, K\. Wang, P\. Xu, Z\. Feng, H\. Bai, K\. Han, Y\. Yan, B\. Yuan, and J\. SunBeyond masks: efficient, flexible diffusion language models via deletion\-insertion processes\.External Links:2603\.23507Cited by:[§5](https://arxiv.org/html/2609.35817#S5.SS0.SSS0.Px1.p1.1)\.
- Gokaslanet al\.\(2019\)A\. Gokaslan, V\. Cohen, E\. Pavlick, and S\. TellexOpenWebText corpus\.Note:[http://Skylion007\.github\.io/OpenWebTextCorpus](http://skylion007.github.io/OpenWebTextCorpus)Cited by:[§C\.1](https://arxiv.org/html/2609.35817#A3.SS1.SSS0.Px1.p1.1),[§4\.1](https://arxiv.org/html/2609.35817#S4.SS1.SSS0.Px1.p1.1)\.
- Hendryckset al\.\(2020\)D\. Hendrycks, C\. Burns, S\. Basart, A\. Zou, M\. Mazeika, D\. Song, and J\. SteinhardtMeasuring massive multitask language understanding\.arXiv preprint arXiv:2009\.03300\.Cited by:[§C\.2](https://arxiv.org/html/2609.35817#A3.SS2.SSS0.Px2.p1.1)\.
- Laiet al\.\(2017\)G\. Lai, Q\. Xie, H\. Liu, Y\. Yang, and E\. HovyRACE: large\-scale ReAding comprehension dataset from examinations\.InProceedings of the 2017 Conference on Empirical Methods in Natural Language Processing,M\. Palmer, R\. Hwa, and S\. Riedel \(Eds\.\),Copenhagen, Denmark,pp\. 785–794\.External Links:[Link](https://aclanthology.org/D17-1082/),[Document](https://dx.doi.org/10.18653/v1/D17-1082)Cited by:[§C\.1](https://arxiv.org/html/2609.35817#A3.SS1.SSS0.Px2.p1.1)\.
- Leeet al\.\(2026\)C\. Lee, J\. Yoo, M\. Agarwal, S\. Shah, J\. Huang, A\. Raghunathan, S\. Hong, N\. M\. Boffi, and J\. KimFlow map language models: one\-step language modeling via continuous denoising\.arXiv preprint arXiv:2602\.16813\.Cited by:[§5](https://arxiv.org/html/2609.35817#S5.SS0.SSS0.Px2.p1.1)\.
- Lightmanet al\.\(2024\)H\. Lightman, V\. Kosaraju, Y\. Burda, H\. Edwards, B\. Baker, T\. Lee, J\. Leike, J\. Schulman, I\. Sutskever, and K\. CobbeLet’s verify step by step\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 39578–39601\.Cited by:[§C\.3](https://arxiv.org/html/2609.35817#A3.SS3.p1.1)\.
- Liuet al\.\(2025\)A\. Liu, M\. He, S\. Zeng, S\. Zhang, L\. Zhang, C\. Wu, W\. Jia, Y\. Liu, X\. Zhou, and J\. ZhouWeDLM: reconciling diffusion language models with standard causal attention for fast inference\.External Links:2512\.22737Cited by:[§5](https://arxiv.org/html/2609.35817#S5.SS0.SSS0.Px1.p1.1)\.
- Liuet al\.\(2023\)J\. Liu, C\. S\. Xia, Y\. Wang, and L\. ZhangIs your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation\.Advances in neural information processing systems36,pp\. 21558–21572\.Cited by:[§C\.2](https://arxiv.org/html/2609.35817#A3.SS2.SSS0.Px2.p1.1)\.
- Louet al\.\(2023\)A\. Lou, C\. Meng, and S\. ErmonDiscrete diffusion modeling by estimating the ratios of the data distribution\.arXiv preprint arXiv:2310\.16834\.Cited by:[§2\.2](https://arxiv.org/html/2609.35817#S2.SS2.p2.1)\.
- Mihaylovet al\.\(2018\)T\. Mihaylov, P\. Clark, T\. Khot, and A\. SabharwalCan a suit of armor conduct electricity? a new dataset for open book question answering\.InProceedings of the 2018 Conference on Empirical Methods in Natural Language Processing,E\. Riloff, D\. Chiang, J\. Hockenmaier, and J\. Tsujii \(Eds\.\),Brussels, Belgium,pp\. 2381–2391\.External Links:[Link](https://aclanthology.org/D18-1260/),[Document](https://dx.doi.org/10.18653/v1/D18-1260)Cited by:[§C\.1](https://arxiv.org/html/2609.35817#A3.SS1.SSS0.Px2.p1.1)\.
- Nathawaniet al\.\(2025\)D\. Nathawani, S\. Ding, V\. Lavrukhin, I\. Gitman, S\. Majumdar, E\. Bakhturina, B\. Ginsburg, and J\. Polak ScowcroftNemotron\-Post\-Training\-Dataset\-v2\.NVIDIA\.External Links:[Link](https://huggingface.co/datasets/nvidia/Nemotron-Post-Training-Dataset-v2)Cited by:[§C\.3](https://arxiv.org/html/2609.35817#A3.SS3.p4.1)\.
- Niet al\.\(2026\)Z\. Ni, S\. Wang, Y\. Yue, T\. Yu, W\. Zhao, Y\. Hua, T\. Chen, J\. Song, C\. Yu, B\. Zheng, and G\. HuangThe flexibility trap: why arbitrary order limits reasoning potential in diffusion language models\.External Links:2601\.15165Cited by:[§5](https://arxiv.org/html/2609.35817#S5.SS0.SSS0.Px1.p1.1)\.
- Nieet al\.\(2025a\)S\. Nie, F\. Zhu, C\. Du, T\. Pang, Q\. Liu, G\. Zeng, M\. Lin, and C\. LiScaling up masked diffusion models on text\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 82974–82997\.Cited by:[§C\.1](https://arxiv.org/html/2609.35817#A3.SS1.SSS0.Px2.p1.1)\.
- Nieet al\.\(2025b\)S\. Nie, F\. Zhu, Z\. You, X\. Zhang, J\. Ou, J\. Hu, J\. Zhou, Y\. Lin, J\. Wen, and C\. LiLarge language diffusion models\.External Links:2502\.09992Cited by:[§C\.2](https://arxiv.org/html/2609.35817#A3.SS2.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.35817#S1.p1.1),[§5](https://arxiv.org/html/2609.35817#S5.SS0.SSS0.Px1.p1.1)\.
- Olmoet al\.\(2025\)T\. Olmo, A\. Ettinger, A\. Bertsch, B\. Kuehl, D\. Graham, D\. Heineman, D\. Groeneveld, F\. Brahman, F\. Timbers, H\. Ivison, J\. Morrison, J\. Poznanski, K\. Lo, L\. Soldaini, M\. Jordan, M\. Chen, M\. Noukhovitch, N\. Lambert, P\. Walsh, P\. Dasigi, R\. Berry, S\. Malik, S\. Shah, S\. Geng, S\. Arora, S\. Gupta, T\. Anderson, T\. Xiao, T\. Murray, T\. Romero, V\. Graf, A\. Asai, A\. Bhagia, A\. Wettig, A\. Liu, A\. Rangapur, C\. Anastasiades, C\. Huang, D\. Schwenk, H\. Trivedi, I\. Magnusson, J\. Lochner, J\. Liu, L\. J\. V\. Miranda, M\. Sap, M\. Morgan, M\. Schmitz, M\. Guerquin, M\. Wilson, R\. Huff, R\. L\. Bras, R\. Xin, R\. Shao, S\. Skjonsberg, S\. Z\. Shen, S\. S\. Li, T\. Wilde, V\. Pyatkin, W\. Merrill, Y\. Chang, Y\. Gu, Z\. Zeng, A\. Sabharwal, L\. Zettlemoyer, P\. W\. Koh, A\. Farhadi, N\. A\. Smith, and H\. HajishirziOlmo 3\.External Links:2512\.13961,[Link](https://arxiv.org/abs/2512.13961)Cited by:[§4\.2](https://arxiv.org/html/2609.35817#S4.SS2.SSS0.Px1.p1.1)\.
- Ouet al\.\(2025\)J\. Ou, S\. Nie, K\. Xue, F\. Zhu, J\. Sun, Z\. Li, and C\. LiYour absorbing discrete diffusion secretly models the conditional distributions of clean data\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 64972–65009\.Cited by:[§C\.1](https://arxiv.org/html/2609.35817#A3.SS1.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.35817#S1.p1.1)\.
- Peebles and Xie \(2023\)W\. Peebles and S\. XieScalable diffusion models with transformers\.InProceedings of the IEEE/CVF International Conference on Computer Vision \(ICCV\),External Links:[Link](https://openaccess.thecvf.com/content/ICCV2023/html/Peebles_Scalable_Diffusion_Models_with_Transformers_ICCV_2023_paper.html)Cited by:[§3\.2](https://arxiv.org/html/2609.35817#S3.SS2.SSS0.Px2.p3.1)\.
- Penedoet al\.\(2024\)G\. Penedo, H\. Kydlíček, L\. Ben allal, A\. Lozhkov, M\. Mitchell, C\. Raffel, L\. Von Werra, and T\. WolfThe fineweb datasets: decanting the web for the finest text data at scale\.Advances in Neural Information Processing Systems37,pp\. 30811–30849\.Cited by:[§C\.1](https://arxiv.org/html/2609.35817#A3.SS1.SSS0.Px2.p1.1),[§4\.1](https://arxiv.org/html/2609.35817#S4.SS1.SSS0.Px1.p1.1)\.
- Radfordet al\.\(2019\)A\. Radford, J\. Wu, R\. Child, D\. Luan, D\. Amodei, and I\. SutskeverLanguage models are unsupervised multitask learners\.External Links:[Link](https://api.semanticscholar.org/CorpusID:160025533)Cited by:[§C\.1](https://arxiv.org/html/2609.35817#A3.SS1.SSS0.Px1.p1.1),[§4\.1](https://arxiv.org/html/2609.35817#S4.SS1.SSS0.Px1.p1.1)\.
- Reinet al\.\(2023\)D\. Rein, B\. L\. Hou, A\. C\. Stickland, J\. Petty, R\. Y\. Pang, J\. Dirani, J\. Michael, and S\. R\. BowmanGpqa: a graduate\-level google\-proof q&a benchmark\.arXiv preprint arXiv:2311\.12022\.Cited by:[§C\.2](https://arxiv.org/html/2609.35817#A3.SS2.SSS0.Px2.p1.1)\.
- Sahooet al\.\(2024\)S\. S\. Sahoo, M\. Arriola, Y\. Schiff, A\. Gokaslan, E\. Marroquin, J\. T\. Chiu, A\. Rush, and V\. KuleshovSimple and effective masked diffusion language models\.InAdvances in Neural Information Processing Systems,Cited by:[§A\.1\.4](https://arxiv.org/html/2609.35817#A1.SS1.SSS4.p2.1),[§C\.1](https://arxiv.org/html/2609.35817#A3.SS1.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.35817#S1.p1.1)\.
- Sahooet al\.\(2025\)S\. S\. Sahoo, J\. Deschenaux, A\. Gokaslan, G\. Wang, J\. Chiu, and V\. KuleshovThe diffusion duality\.InInternational Conference on Machine Learning,Cited by:[§C\.1](https://arxiv.org/html/2609.35817#A3.SS1.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.35817#S1.p1.1),[§1](https://arxiv.org/html/2609.35817#S1.p2.1),[§2\.2](https://arxiv.org/html/2609.35817#S2.SS2.p2.1),[§5](https://arxiv.org/html/2609.35817#S5.SS0.SSS0.Px2.p1.1)\.
- Sapet al\.\(2019\)M\. Sap, H\. Rashkin, D\. Chen, R\. Le Bras, and Y\. ChoiSocial IQa: commonsense reasoning about social interactions\.InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing \(EMNLP\-IJCNLP\),K\. Inui, J\. Jiang, V\. Ng, and X\. Wan \(Eds\.\),Hong Kong, China,pp\. 4463–4473\.External Links:[Link](https://aclanthology.org/D19-1454/),[Document](https://dx.doi.org/10.18653/v1/D19-1454)Cited by:[§C\.1](https://arxiv.org/html/2609.35817#A3.SS1.SSS0.Px2.p1.1)\.
- Schiffet al\.\(2026\)Y\. Schiff, O\. Belhasin, R\. Uziel, G\. Wang, M\. Arriola, G\. Turok, M\. Elad, and V\. KuleshovLearn from your mistakes: self\-correcting masked diffusion models\.External Links:2602\.11590Cited by:[§5](https://arxiv.org/html/2609.35817#S5.SS0.SSS0.Px2.p1.1)\.
- Schiffet al\.\(2025\)Y\. Schiff, S\. Sahoo, H\. Phung, G\. Wang, S\. Boshar, H\. Dalla\-torre, B\. Almeida, A\. Rush, T\. Pierrot, and V\. KuleshovSimple guidance mechanisms for discrete diffusion models\.InInternational Conference on Learning Representations,Y\. Yue, A\. Garg, N\. Peng, F\. Sha, and R\. Yu \(Eds\.\),Vol\.2025,pp\. 43776–43821\.External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2025/file/6cc31b44d88dce8380d36e81485cd07f-Paper-Conference.pdf)Cited by:[§1](https://arxiv.org/html/2609.35817#S1.p1.1),[§5](https://arxiv.org/html/2609.35817#S5.SS0.SSS0.Px2.p1.1)\.
- Team \(2024\)Q\. TeamQwen2\.5: a party of foundation models\.External Links:[Link](https://qwenlm.github.io/blog/qwen2.5/)Cited by:[§C\.2](https://arxiv.org/html/2609.35817#A3.SS2.SSS0.Px1.p1.1),[§4\.2](https://arxiv.org/html/2609.35817#S4.SS2.SSS0.Px1.p1.1)\.
- Tianet al\.\(2025\)Y\. Tian, Y\. Liang, S\. Zhang, Y\. Shu, G\. Yang, W\. He, S\. Fang, T\. Guo, K\. Han, C\. Xu, H\. Chen, X\. Chen, and Y\. WangFrom next\-token to next\-block: a principled adaptation path for diffusion llms\.arXiv preprint arXiv:2512\.06776\.Cited by:[§B\.2](https://arxiv.org/html/2609.35817#A2.SS2.SSS0.Px3.p1.1),[§3\.2](https://arxiv.org/html/2609.35817#S3.SS2.SSS0.Px4.p3.1),[§5](https://arxiv.org/html/2609.35817#S5.SS0.SSS0.Px1.p1.1)\.
- Touvronet al\.\(2023\)H\. Touvron, T\. Lavril, G\. Izacard, X\. Martinet, M\. Lachaux, T\. Lacroix, B\. Rozière, N\. Goyal, E\. Hambro, F\. Azhar, A\. Rodriguez, A\. Joulin, E\. Grave, and G\. LampleLLaMA: open and efficient foundation language models\.External Links:2302\.13971,[Link](https://arxiv.org/abs/2302.13971)Cited by:[§C\.1](https://arxiv.org/html/2609.35817#A3.SS1.SSS0.Px2.p1.1),[§4\.1](https://arxiv.org/html/2609.35817#S4.SS1.SSS0.Px1.p1.1)\.
- von Rütteet al\.\(2025a\)D\. von Rütte, J\. Fluri, Y\. Ding, A\. Orvieto, B\. Schölkopf, and T\. HofmannGeneralized interpolating discrete diffusion\.External Links:2503\.04482Cited by:[§5](https://arxiv.org/html/2609.35817#S5.SS0.SSS0.Px1.p1.1),[§5](https://arxiv.org/html/2609.35817#S5.SS0.SSS0.Px2.p1.1)\.
- von Rütteet al\.\(2025b\)D\. von Rütte, J\. Fluri, O\. Pooladzandi, B\. Schölkopf, T\. Hofmann, and A\. OrvietoScaling behavior of discrete diffusion language models\.External Links:2512\.10858Cited by:[§1](https://arxiv.org/html/2609.35817#S1.p2.1),[§4\.2](https://arxiv.org/html/2609.35817#S4.SS2.SSS0.Px3.p1.1),[§5](https://arxiv.org/html/2609.35817#S5.SS0.SSS0.Px2.p1.1)\.
- Wuet al\.\(2025a\)C\. Wu, H\. Zhang, S\. Xue, S\. Diao, Y\. Fu, Z\. Liu, P\. Molchanov, P\. Luo, S\. Han, and E\. XieFast\-dllm v2: efficient block\-diffusion llm\.arXiv preprint arXiv:2509\.26328\.Cited by:[§B\.2](https://arxiv.org/html/2609.35817#A2.SS2.SSS0.Px2.p1.1),[§C\.2](https://arxiv.org/html/2609.35817#A3.SS2.SSS0.Px1.p1.1),[§C\.2](https://arxiv.org/html/2609.35817#A3.SS2.SSS0.Px2.p1.1),[§3\.2](https://arxiv.org/html/2609.35817#S3.SS2.SSS0.Px4.p2.1),[§4\.2](https://arxiv.org/html/2609.35817#S4.SS2.SSS0.Px1.p1.1)\.
- Wuet al\.\(2025b\)C\. Wu, H\. Zhang, S\. Xue, Z\. Liu, S\. Diao, L\. Zhu, P\. Luo, S\. Han, and E\. XieFast\-dLLM: training\-free acceleration of diffusion LLM by enabling KV cache and parallel decoding\.External Links:2505\.22618Cited by:[§5](https://arxiv.org/html/2609.35817#S5.SS0.SSS0.Px1.p1.1)\.
- Yeet al\.\(2025\)J\. Ye, Z\. Xie, L\. Zheng, J\. Gao, Z\. Wu, X\. Jiang, Z\. Li, and L\. KongDream 7b: diffusion large language models\.arXiv preprint arXiv:2508\.15487\.Cited by:[§C\.2](https://arxiv.org/html/2609.35817#A3.SS2.SSS0.Px1.p1.1)\.
- Zellerset al\.\(2019\)R\. Zellers, A\. Holtzman, Y\. Bisk, A\. Farhadi, and Y\. ChoiHellaSwag: can a machine really finish your sentence?\.InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics,A\. Korhonen, D\. Traum, and L\. Màrquez \(Eds\.\),Florence, Italy,pp\. 4791–4800\.External Links:[Link](https://aclanthology.org/P19-1472/),[Document](https://dx.doi.org/10.18653/v1/P19-1472)Cited by:[§C\.1](https://arxiv.org/html/2609.35817#A3.SS1.SSS0.Px2.p1.1)\.
- Zhaoet al\.\(2025\)S\. Zhao, D\. Gupta, Q\. Zheng, and A\. GroverD1: scaling reasoning in diffusion large language models via reinforcement learning\.External Links:2504\.12216Cited by:[§5](https://arxiv.org/html/2609.35817#S5.SS0.SSS0.Px1.p1.1)\.
- Zhonget al\.\(2026\)J\. Zhong, K\. Wang, D\. Ding, Z\. Feng, H\. Bai, Y\. Xiang, J\. Sun, and Q\. XuStabilizing reinforcement learning for diffusion language models\.External Links:2603\.06743Cited by:[§5](https://arxiv.org/html/2609.35817#S5.SS0.SSS0.Px1.p1.1)\.
- Zhouet al\.\(2023\)J\. Zhou, T\. Lu, S\. Mishra, S\. Brahma, S\. Basu, Y\. Luan, D\. Zhou, and L\. HouInstruction\-following evaluation for large language models\.arXiv preprint arXiv:2311\.07911\.Cited by:[§C\.2](https://arxiv.org/html/2609.35817#A3.SS2.SSS0.Px2.p1.1)\.
- Zhuet al\.\(2025a\)F\. Zhu, R\. Wang, S\. Nie, X\. Zhang, C\. Wu, J\. Hu, J\. Zhou, J\. Chen, Y\. Lin, J\. Wen, and C\. LiLlada 1\.5: variance\-reduced preference optimization for large language diffusion models\.arXiv preprint arXiv:2505\.19223\.Cited by:[§C\.2](https://arxiv.org/html/2609.35817#A3.SS2.SSS0.Px1.p1.1)\.
- Zhuet al\.\(2025b\)F\. Zhu, Z\. You, Y\. Xing, Z\. Huang, L\. Liu, Y\. Zhuang, G\. Lu, K\. Wang, X\. Wang, L\. Wei, H\. Guo, J\. Hu, W\. Ye, T\. Chen, C\. Li, C\. Tang, H\. Feng, J\. Hu, J\. Zhou, X\. Zhang,et al\.Llada\-moe: a sparse moe diffusion language model\.arXiv preprint arXiv:2509\.24389\.Cited by:[§C\.2](https://arxiv.org/html/2609.35817#A3.SS2.SSS0.Px1.p1.1)\.
- Zhuet al\.\(2025c\)H\. Zhu, Z\. Chen, S\. Zhou, Z\. Xie, Y\. Yuan, S\. Chen, Z\. Guo, S\. Xu, H\. Zhang, V\. Honavar, and T\. XiaoSimple denoising diffusion language models\.arXiv preprint arXiv:2510\.22926\.Cited by:[§C\.1](https://arxiv.org/html/2609.35817#A3.SS1.SSS0.Px1.p1.1),[§C\.1](https://arxiv.org/html/2609.35817#A3.SS1.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.35817#S1.p2.1),[§2\.2](https://arxiv.org/html/2609.35817#S2.SS2.p4.1),[§5](https://arxiv.org/html/2609.35817#S5.SS0.SSS0.Px2.p1.1)\.
## Appendix AProofs and Derivations
We prove the results of the main text for one token; sequence\-level objectives follow by summing over positions\. Let𝒱=\{1,…,V\}\\mathcal\{V\}=\\\{1,\\ldots,V\\\}, let𝒆i\\bm\{e\}\_\{i\}be the one\-hot vector for tokenii, and write𝒙0=𝒆y\\bm\{x\}\_\{0\}=\\bm\{e\}\_\{y\}and𝒙t=𝒆m\\bm\{x\}\_\{t\}=\\bm\{e\}\_\{m\}\. Thusm≠ym\\neq ydenotes a corrupted position\.
Assume thatt↦αtt\\mapsto\\alpha\_\{t\}is continuous and non\-increasing on\[0,1\]\[0,1\], continuously differentiable on\[0,1\)\[0,1\), and satisfiesα0=1\\alpha\_\{0\}=1,αt\>0\\alpha\_\{t\}\>0fort<1t<1, andα1=0\\alpha\_\{1\}=0\. ForΔV−1:=\{𝒙∈ℝ≥0V:∑ixi=1\}\\Delta^\{V\-1\}:=\\\{\\bm\{x\}\\in\\mathbb\{R\}\_\{\\geq 0\}^\{V\}:\\sum\_\{i\}x\_\{i\}=1\\\}, define
𝒙¯:=Vαt𝒙\+\(1−αt\)𝟏,x¯i=Vαtxi\+1−αt\.\\bar\{\\bm\{x\}\}:=V\\alpha\_\{t\}\\bm\{x\}\+\(1\-\\alpha\_\{t\}\)\\mathbf\{1\},\\qquad\\bar\{x\}\_\{i\}=V\\alpha\_\{t\}x\_\{i\}\+1\-\\alpha\_\{t\}\.\(10\)Then∑ix¯i=V\\sum\_\{i\}\\bar\{x\}\_\{i\}=V, so𝒙¯/V\\bar\{\\bm\{x\}\}/Vis a probability vector\. We usex¯0,i\\bar\{x\}\_\{0,i\}andx¯θ,i\\bar\{x\}\_\{\\theta,i\}for the vectors obtained from𝒙0\\bm\{x\}\_\{0\}and𝒙θ\(𝒙t,t\)\\bm\{x\}\_\{\\theta\}\(\\bm\{x\}\_\{t\},t\), respectively\.
### A\.1Uniform\-State Diffusion Preliminaries
We first verify the forward law and reverse posterior, then connect the SEDD and Duo objectives\.
#### A\.1\.1Forward Process
We verify that the stated rate matrix generates the uniform\-state marginal\.
Fort∈\[0,1\)t\\in\[0,1\), the uniform\-state transition rate matrix is
Qt\(i,j\)=−αt′αt\(1V−𝕀i=j\)\.Q\_\{t\}\(i,j\)=\-\\frac\{\\alpha\_\{t\}^\{\\prime\}\}\{\\alpha\_\{t\}\}\\left\(\\frac\{1\}\{V\}\-\\mathbb\{I\}\_\{i=j\}\\right\)\.\(11\)Sinceαt′≤0\\alpha\_\{t\}^\{\\prime\}\\leq 0, the off\-diagonal entries are nonnegative, and each row sums to zero\. Starting from𝒙0=𝒆y\\bm\{x\}\_\{0\}=\\bm\{e\}\_\{y\}, consider
pt\(j∣y\):=αt𝕀j=y\+\(1−αt\)1V\.p\_\{t\}\(j\\mid y\):=\\alpha\_\{t\}\\mathbb\{I\}\_\{j=y\}\+\(1\-\\alpha\_\{t\}\)\\frac\{1\}\{V\}\.\(12\)Direct substitution gives
ddtpt\(j∣y\)=αt′\(𝕀j=y−1V\)=\[pt\(⋅∣y\)Qt\]j\.\\frac\{d\}\{dt\}p\_\{t\}\(j\\mid y\)=\\alpha\_\{t\}^\{\\prime\}\\left\(\\mathbb\{I\}\_\{j=y\}\-\\frac\{1\}\{V\}\\right\)=\\bigl\[p\_\{t\}\(\\cdot\\mid y\)Q\_\{t\}\\bigr\]\_\{j\}\.\(13\)Together withp0\(j∣y\)=𝕀j=yp\_\{0\}\(j\\mid y\)=\\mathbb\{I\}\_\{j=y\}, uniqueness of the forward equation yields
p\(𝒙t∣𝒙0\)=Cat\(⋅,αt𝒙0\+\(1−αt\)𝟏V\)\.p\(\\bm\{x\}\_\{t\}\\mid\\bm\{x\}\_\{0\}\)=\\mathrm\{Cat\}\\left\(\\cdot;\\,\\alpha\_\{t\}\\bm\{x\}\_\{0\}\+\(1\-\\alpha\_\{t\}\)\\frac\{\\mathbf\{1\}\}\{V\}\\right\)\.\(14\)
#### A\.1\.2Exact Reverse Posterior
We obtain the reverse posterior by applying Bayes’ rule to the forward transitions\.
Let0≤s<t≤10\\leq s<t\\leq 1andαt\|s:=αt/αs\\alpha\_\{t\\mid s\}:=\\alpha\_\{t\}/\\alpha\_\{s\}\. The transition fromsstottis
p\(𝒙t=𝒆m∣𝒙s=𝒆k\)=αt\|s𝕀m=k\+1−αt\|sV\.p\(\\bm\{x\}\_\{t\}=\\bm\{e\}\_\{m\}\\mid\\bm\{x\}\_\{s\}=\\bm\{e\}\_\{k\}\)=\\alpha\_\{t\\mid s\}\\mathbb\{I\}\_\{m=k\}\+\\frac\{1\-\\alpha\_\{t\\mid s\}\}\{V\}\.\(15\)For an observation𝒙t=𝒆m\\bm\{x\}\_\{t\}=\\bm\{e\}\_\{m\}of positive probability, Bayes’ rule gives
\\fitbox0\.98p\(𝒙s=𝒆k∣𝒙t=𝒆m,𝒙0=𝒆y\)=\(αt\|s𝕀m=k\+1−αt\|sV\)\(αs𝕀k=y\+1−αsV\)αt𝕀m=y\+1−αtV\.\\fitbox\{0\.98\}\{\\displaystyle p\(\\bm\{x\}\_\{s\}=\\bm\{e\}\_\{k\}\\mid\\bm\{x\}\_\{t\}=\\bm\{e\}\_\{m\},\\bm\{x\}\_\{0\}=\\bm\{e\}\_\{y\}\)=\\frac\{\\left\(\\alpha\_\{t\\mid s\}\\mathbb\{I\}\_\{m=k\}\+\\frac\{1\-\\alpha\_\{t\\mid s\}\}\{V\}\\right\)\\left\(\\alpha\_\{s\}\\mathbb\{I\}\_\{k=y\}\+\\frac\{1\-\\alpha\_\{s\}\}\{V\}\\right\)\}\{\\alpha\_\{t\}\\mathbb\{I\}\_\{m=y\}\+\\frac\{1\-\\alpha\_\{t\}\}\{V\}\}\.\}\(16\)
Let𝝅t→s∈ΔV−1\\bm\{\\pi\}\_\{t\\to s\}\\in\\Delta^\{V\-1\}denote the categorical parameter vector of this posterior, with
\[𝝅t→s\]k:=p\(𝒙s=𝒆k∣𝒙t=𝒆m,𝒙0=𝒆y\)\.\[\\bm\{\\pi\}\_\{t\\to s\}\]\_\{k\}:=p\(\\bm\{x\}\_\{s\}=\\bm\{e\}\_\{k\}\\mid\\bm\{x\}\_\{t\}=\\bm\{e\}\_\{m\},\\bm\{x\}\_\{0\}=\\bm\{e\}\_\{y\}\)\.\(17\)
Multiplying numerator and denominator byVVand usingαt\|sαs=αt\\alpha\_\{t\\mid s\}\\alpha\_\{s\}=\\alpha\_\{t\}, the numerator of coordinatekkis
V\(αt\|s𝕀m=k\+1−αt\|sV\)\(αs𝕀k=y\+1−αsV\)\\displaystyle V\\left\(\\alpha\_\{t\\mid s\}\\mathbb\{I\}\_\{m=k\}\+\\frac\{1\-\\alpha\_\{t\\mid s\}\}\{V\}\\right\)\\left\(\\alpha\_\{s\}\\mathbb\{I\}\_\{k=y\}\+\\frac\{1\-\\alpha\_\{s\}\}\{V\}\\right\)=Vαt𝕀m=k𝕀k=y\+\(αt\|s−αt\)𝕀m=k\+\(αs−αt\)𝕀k=y\+\(1−αt\|s\)\(1−αs\)V\.\\displaystyle\\quad=V\\alpha\_\{t\}\\mathbb\{I\}\_\{m=k\}\\mathbb\{I\}\_\{k=y\}\+\(\\alpha\_\{t\\mid s\}\-\\alpha\_\{t\}\)\\mathbb\{I\}\_\{m=k\}\+\(\\alpha\_\{s\}\-\\alpha\_\{t\}\)\\mathbb\{I\}\_\{k=y\}\+\\frac\{\(1\-\\alpha\_\{t\\mid s\}\)\(1\-\\alpha\_\{s\}\)\}\{V\}\.\(18\)Collecting these coordinates yields
𝝅t→s\\displaystyle\\bm\{\\pi\}\_\{t\\to s\}=1Vαt𝕀m=y\+1−αt\[Vαt𝒙t⊙𝒙0\+\(αt\|s−αt\)𝒙t\+\(αs−αt\)𝒙0\+\(1−αt\|s\)\(1−αs\)𝟏V\],\\displaystyle=\\frac\{1\}\{V\\alpha\_\{t\}\\mathbb\{I\}\_\{m=y\}\+1\-\\alpha\_\{t\}\}\\Bigg\[V\\alpha\_\{t\}\\,\\bm\{x\}\_\{t\}\\odot\\bm\{x\}\_\{0\}\+\(\\alpha\_\{t\\mid s\}\-\\alpha\_\{t\}\)\\bm\{x\}\_\{t\}\+\(\\alpha\_\{s\}\-\\alpha\_\{t\}\)\\bm\{x\}\_\{0\}\+\(1\-\\alpha\_\{t\\mid s\}\)\(1\-\\alpha\_\{s\}\)\\frac\{\\mathbf\{1\}\}\{V\}\\Bigg\],\(19\)
#### A\.1\.3SEDD–Duo Reparameterization
We show that the SEDD score\-entropy objective reduces exactly to the Duo objective under thex0x\_\{0\}\-parameterization\.
For0≤s<t<10\\leq s<t<1, consider the local reverse KL between the exact posterior and the learned reverse transition,
KL\(p\(𝒙s∣𝒙t,𝒙0\)∥pθ\(𝒙s∣𝒙t\)\)\.\\mathrm\{KL\}\\\!\\bigl\(p\(\\bm\{x\}\_\{s\}\\mid\\bm\{x\}\_\{t\},\\bm\{x\}\_\{0\}\)\\;\\\|\\;p\_\{\\theta\}\(\\bm\{x\}\_\{s\}\\mid\\bm\{x\}\_\{t\}\)\\bigr\)\.\(20\)The posterior entropy is independent ofθ\\theta, so only the cross\-entropy is model\-dependent\. Fix𝒙0=𝒆y\\bm\{x\}\_\{0\}=\\bm\{e\}\_\{y\}\. The conditional SEDD target is
s⋆\(j,m,t\)=pt\(𝒆j∣𝒙0\)pt\(𝒆m∣𝒙0\)=x¯0,jx¯0,m\.s^\{\\star\}\(j,m,t\)=\\frac\{p\_\{t\}\(\\bm\{e\}\_\{j\}\\mid\\bm\{x\}\_\{0\}\)\}\{p\_\{t\}\(\\bm\{e\}\_\{m\}\\mid\\bm\{x\}\_\{0\}\)\}=\\frac\{\\bar\{x\}\_\{0,j\}\}\{\\bar\{x\}\_\{0,m\}\}\.\(21\)The ratio is evaluated only when the denominator is positive, with0log0=00\\log 0=0\. The token\-level denoising score\-entropy objective is
ℒDSEℓ=𝔼t,𝒙t∑j≠mQt\(j,m\)\[sθ\(j,m,t\)−s⋆\(j,m,t\)logsθ\(j,m,t\)\+s⋆\(j,m,t\)\(logs⋆\(j,m,t\)−1\)\]\.\\mathcal\{L\}\_\{\\mathrm\{DSE\}\}^\{\\ell\}=\\mathbb\{E\}\_\{t,\\bm\{x\}\_\{t\}\}\\\!\\sum\_\{j\\neq m\}Q\_\{t\}\(j,m\)\\\!\\bigl\[s\_\{\\theta\}\(j,m,t\)\-s^\{\\star\}\(j,m,t\)\\log s\_\{\\theta\}\(j,m,t\)\+s^\{\\star\}\(j,m,t\)\\bigl\(\\log s^\{\\star\}\(j,m,t\)\-1\\bigr\)\\bigr\]\.\(22\)Forj≠mj\\neq m,Qt\(j,m\)=−αt′/\(Vαt\)Q\_\{t\}\(j,m\)=\-\\alpha\_\{t\}^\{\\prime\}/\(V\\alpha\_\{t\}\)\. The omittedj=mj=msummand is zero and may be restored\. Under the Duo parameterization,
sθ\(j,m,t\)=x¯θ,jx¯θ,m\.s\_\{\\theta\}\(j,m,t\)=\\frac\{\\bar\{x\}\_\{\\theta,j\}\}\{\\bar\{x\}\_\{\\theta,m\}\}\.\(23\)Using∑jx¯θ,j=∑jx¯0,j=V\\sum\_\{j\}\\bar\{x\}\_\{\\theta,j\}=\\sum\_\{j\}\\bar\{x\}\_\{0,j\}=V,
ℒDSEℓ\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{DSE\}\}^\{\\ell\}=𝔼t,𝒙t−αt′Vαt\[∑j=1Vx¯θ,jx¯θ,m−∑j=1Vx¯0,jx¯0,m\+∑j=1Vx¯0,jx¯0,mlogx¯θ,mx¯0,jx¯θ,jx¯0,m\]\\displaystyle=\\mathbb\{E\}\_\{t,\\bm\{x\}\_\{t\}\}\\frac\{\-\\alpha\_\{t\}^\{\\prime\}\}\{V\\alpha\_\{t\}\}\\left\[\\sum\_\{j=1\}^\{V\}\\frac\{\\bar\{x\}\_\{\\theta,j\}\}\{\\bar\{x\}\_\{\\theta,m\}\}\-\\sum\_\{j=1\}^\{V\}\\frac\{\\bar\{x\}\_\{0,j\}\}\{\\bar\{x\}\_\{0,m\}\}\+\\sum\_\{j=1\}^\{V\}\\frac\{\\bar\{x\}\_\{0,j\}\}\{\\bar\{x\}\_\{0,m\}\}\\log\\frac\{\\bar\{x\}\_\{\\theta,m\}\\bar\{x\}\_\{0,j\}\}\{\\bar\{x\}\_\{\\theta,j\}\\bar\{x\}\_\{0,m\}\}\\right\]=𝔼t,𝒙t−αt′Vαt\[Vx¯θ,m−Vx¯0,m\+∑j=1Vx¯0,jx¯0,mlogx¯θ,mx¯0,jx¯θ,jx¯0,m\]\.\\displaystyle=\\mathbb\{E\}\_\{t,\\bm\{x\}\_\{t\}\}\\frac\{\-\\alpha\_\{t\}^\{\\prime\}\}\{V\\alpha\_\{t\}\}\\left\[\\frac\{V\}\{\\bar\{x\}\_\{\\theta,m\}\}\-\\frac\{V\}\{\\bar\{x\}\_\{0,m\}\}\+\\sum\_\{j=1\}^\{V\}\\frac\{\\bar\{x\}\_\{0,j\}\}\{\\bar\{x\}\_\{0,m\}\}\\log\\frac\{\\bar\{x\}\_\{\\theta,m\}\\bar\{x\}\_\{0,j\}\}\{\\bar\{x\}\_\{\\theta,j\}\\bar\{x\}\_\{0,m\}\}\\right\]\.\(24\)This is Eq\. \([5](https://arxiv.org/html/2609.35817#S2.E5)\)\.
#### A\.1\.4Contrast with Mask Diffusion
We contrast the uniform\-state objective with the simpler posterior structure of absorbing\-state mask diffusion\.
For comparison, standard absorbing\-state mask diffusion on𝒱∪\{\[𝙼𝙰𝚂𝙺\]\}\\mathcal\{V\}\\cup\\\{\\mathtt\{\[MASK\]\}\\\}has the forward law\([Sahoo et al\., 2024](https://arxiv.org/html/2609.35817#bib.bib1)\)
ptM\(y∣y\)=αt,ptM\(\[𝙼𝙰𝚂𝙺\]∣y\)=1−αt,p\_\{t\}^\{\\mathrm\{M\}\}\(y\\mid y\)=\\alpha\_\{t\},\\qquad p\_\{t\}^\{\\mathrm\{M\}\}\(\\mathtt\{\[MASK\]\}\\mid y\)=1\-\\alpha\_\{t\},\(25\)Conditioning onxt=\[𝙼𝙰𝚂𝙺\]x\_\{t\}=\\mathtt\{\[MASK\]\}, the two possible states at timesshave joint probabilitiesαs−αt\\alpha\_\{s\}\-\\alpha\_\{t\}and1−αs1\-\\alpha\_\{s\}\. Sincep\(xt=\[𝙼𝙰𝚂𝙺\]∣x0=y\)=1−αtp\(x\_\{t\}=\\mathtt\{\[MASK\]\}\\mid x\_\{0\}=y\)=1\-\\alpha\_\{t\}, Bayes’ rule gives
at,s=αs−αt1−αt,bt,s=1−αs1−αt\.a\_\{t,s\}=\\frac\{\\alpha\_\{s\}\-\\alpha\_\{t\}\}\{1\-\\alpha\_\{t\}\},\\qquad b\_\{t,s\}=\\frac\{1\-\\alpha\_\{s\}\}\{1\-\\alpha\_\{t\}\}\.\(26\)Under thex0x\_\{0\}\-parameterization, the learned posterior assignsat,sxθ,ja\_\{t,s\}x\_\{\\theta,j\}to tokenjjandbt,sb\_\{t,s\}to\[𝙼𝙰𝚂𝙺\]\\mathtt\{\[MASK\]\}\. Hence
KL\(p\(xs∣xt=\[𝙼𝙰𝚂𝙺\],x0=y\)∥pθ\(xs∣xt=\[𝙼𝙰𝚂𝙺\]\)\)=−at,slogxθ,y\.\\mathrm\{KL\}\\\!\\left\(p\(x\_\{s\}\\mid x\_\{t\}=\\mathtt\{\[MASK\]\},x\_\{0\}=y\)\\,\\\|\\,p\_\{\\theta\}\(x\_\{s\}\\mid x\_\{t\}=\\mathtt\{\[MASK\]\}\)\\right\)=\-a\_\{t,s\}\\log x\_\{\\theta,y\}\.\(27\)Thus the model learns only the clean\-token cross\-entropy; both posterior weights are analytic\.
### A\.2ELBO Loss and LU Loss
This subsection proves the two claims that motivate the LU loss\. Proposition 1 separates the uniform\-state ELBO into interpretable terms, and Theorem 1 shows that the LU contrast is the finite model\-dependent part of a small\-step Dirac\-target KL\.
#### A\.2\.1Proof of Proposition 1
We first state the assumptions precisely\. The proof then treats corrupted and already\-clean positions separately, because their clean noisy\-vector coordinates have different forms\.
##### Proposition 1 \(Restated\)\.
Fix0<ε<1/20<\\varepsilon<1/2and restrict attention to times satisfyingαt∈\[ε,1−ε\]\\alpha\_\{t\}\\in\[\\varepsilon,1\-\\varepsilon\], on which\|αt′\|\|\\alpha\_\{t\}^\{\\prime\}\|is bounded\. Suppose that, uniformly over these times and corrupted positionsm≠ym\\neq y,xθ,m=1/V\+o\(1/V\)x\_\{\\theta,m\}=1/V\+o\(1/V\)asV→∞V\\to\\infty\. After omitting additive terms independent ofθ\\theta, the token\-level ELBO integrand is
−αt′\[logx¯θ,m−logx¯θ,y1−αt−1Vαt∑i=1Vlogx¯θ,i\]\+o\(1\),\-\\alpha\_\{t\}^\{\\prime\}\\\!\\left\[\\frac\{\\log\\bar\{x\}\_\{\\theta,m\}\-\\log\\bar\{x\}\_\{\\theta,y\}\}\{1\-\\alpha\_\{t\}\}\-\\frac\{1\}\{V\\alpha\_\{t\}\}\\sum\_\{i=1\}^\{V\}\\log\\bar\{x\}\_\{\\theta,i\}\\right\]\+o\(1\),\(28\)where the error is uniform over the stated region and therefore remainso\(1\)o\(1\)after averaging\. If, in addition,xθ,y=ω\(1/V\)x\_\{\\theta,y\}=\\omega\(1/V\)uniformly at already\-clean positionsm=ym=y, their model\-dependent contribution is also uniformlyo\(1\)o\(1\)and remainso\(1\)o\(1\)after averaging\.
##### Proof\.
We first isolate the model\-dependent part at corrupted positions, then bound the contribution from already\-clean positions\. All calculations are pointwise in\(𝒙0,t,𝒙t\)\(\\bm\{x\}\_\{0\},t,\\bm\{x\}\_\{t\}\); additive terms independent ofθ\\thetaare omitted\.
##### Corrupted positions\.
We derive the approximation by substituting the three values of the clean noisy\-vector coordinates into the exact ELBO\.
Form≠ym\\neq y,
x¯0,m=1−αt,x¯0,y=Vαt\+1−αt,x¯0,j=1−αt\(j≠y\)\.\\bar\{x\}\_\{0,m\}=1\-\\alpha\_\{t\},\\qquad\\bar\{x\}\_\{0,y\}=V\\alpha\_\{t\}\+1\-\\alpha\_\{t\},\\qquad\\bar\{x\}\_\{0,j\}=1\-\\alpha\_\{t\}\\quad\(j\\neq y\)\.\(29\)Denote the pointwise integrand in Eq\. \([24](https://arxiv.org/html/2609.35817#A1.E24)\) byELBO\\mathrm\{ELBO\}\. Substitution of Eq\. \([29](https://arxiv.org/html/2609.35817#A1.E29)\) and collection of the logarithms give its model\-dependent part
ELBOθ=−αt′\[1αtx¯θ,m\+logx¯θ,mαt\(1−αt\)−logx¯θ,y1−αt−1Vαt∑i=1Vlogx¯θ,i\]\.\\mathrm\{ELBO\}\_\{\\theta\}=\-\\alpha\_\{t\}^\{\\prime\}\\\!\\left\[\\frac\{1\}\{\\alpha\_\{t\}\\bar\{x\}\_\{\\theta,m\}\}\+\\frac\{\\log\\bar\{x\}\_\{\\theta,m\}\}\{\\alpha\_\{t\}\(1\-\\alpha\_\{t\}\)\}\-\\frac\{\\log\\bar\{x\}\_\{\\theta,y\}\}\{1\-\\alpha\_\{t\}\}\-\\frac\{1\}\{V\\alpha\_\{t\}\}\\sum\_\{i=1\}^\{V\}\\log\\bar\{x\}\_\{\\theta,i\}\\right\]\.\(30\)WithR\(z\):=z−1−1\+logzR\(z\):=z^\{\-1\}\-1\+\\log z,
1αtx¯θ,m=1αt−logx¯θ,mαt\+R\(x¯θ,m\)αt\.\\frac\{1\}\{\\alpha\_\{t\}\\bar\{x\}\_\{\\theta,m\}\}=\\frac\{1\}\{\\alpha\_\{t\}\}\-\\frac\{\\log\\bar\{x\}\_\{\\theta,m\}\}\{\\alpha\_\{t\}\}\+\\frac\{R\(\\bar\{x\}\_\{\\theta,m\}\)\}\{\\alpha\_\{t\}\}\.\(31\)Removing the first,θ\\theta\-independent term yields
−αt′\[logx¯θ,m−logx¯θ,y1−αt−1Vαt∑i=1Vlogx¯θ,i\+R\(x¯θ,m\)αt\]\.\-\\alpha\_\{t\}^\{\\prime\}\\\!\\left\[\\frac\{\\log\\bar\{x\}\_\{\\theta,m\}\-\\log\\bar\{x\}\_\{\\theta,y\}\}\{1\-\\alpha\_\{t\}\}\-\\frac\{1\}\{V\\alpha\_\{t\}\}\\sum\_\{i=1\}^\{V\}\\log\\bar\{x\}\_\{\\theta,i\}\+\\frac\{R\(\\bar\{x\}\_\{\\theta,m\}\)\}\{\\alpha\_\{t\}\}\\right\]\.\(32\)The assumptionxθ,m=1/V\+o\(1/V\)x\_\{\\theta,m\}=1/V\+o\(1/V\)implies
x¯θ,m−1=αt\(Vxθ,m−1\)=o\(1\)\.\\bar\{x\}\_\{\\theta,m\}\-1=\\alpha\_\{t\}\(Vx\_\{\\theta,m\}\-1\)=o\(1\)\.\(33\)SinceR\(1\+u\)=u2/2\+O\(u3\)R\(1\+u\)=u^\{2\}/2\+O\(u^\{3\}\), the last term in Eq\. \([32](https://arxiv.org/html/2609.35817#A1.E32)\) is uniformlyo\(1\)o\(1\)on the stated interior\-time region\. The same remains true after averaging\.
Define
CELS1−αt\(𝒙¯θ,y\)=−\[αtlogx¯θ,y\+1−αtV∑i=1Vlogx¯θ,i\]\.\\mathrm\{CE\}\_\{\\mathrm\{LS\}\}^\{1\-\\alpha\_\{t\}\}\(\\bar\{\\bm\{x\}\}\_\{\\theta\},y\)=\-\\left\[\\alpha\_\{t\}\\log\\bar\{x\}\_\{\\theta,y\}\+\\frac\{1\-\\alpha\_\{t\}\}\{V\}\\sum\_\{i=1\}^\{V\}\\log\\bar\{x\}\_\{\\theta,i\}\\right\]\.\(34\)A direct rearrangement yields
logx¯θ,m−logx¯θ,y1−αt−1Vαt∑i=1Vlogx¯θ,i=CELS1−αt\(𝒙¯θ,y\)αt\(1−αt\)\+logx¯θ,m1−αt\.\\frac\{\\log\\bar\{x\}\_\{\\theta,m\}\-\\log\\bar\{x\}\_\{\\theta,y\}\}\{1\-\\alpha\_\{t\}\}\-\\frac\{1\}\{V\\alpha\_\{t\}\}\\sum\_\{i=1\}^\{V\}\\log\\bar\{x\}\_\{\\theta,i\}=\\frac\{\\mathrm\{CE\}\_\{\\mathrm\{LS\}\}^\{1\-\\alpha\_\{t\}\}\(\\bar\{\\bm\{x\}\}\_\{\\theta\},y\)\}\{\\alpha\_\{t\}\(1\-\\alpha\_\{t\}\)\}\+\\frac\{\\log\\bar\{x\}\_\{\\theta,m\}\}\{1\-\\alpha\_\{t\}\}\.\(35\)
##### Already\-clean positions\.
We show thatm=ym=ycontributes onlyo\(1\)o\(1\)under the additional assumption\. Settingm=ym=yin Eq\. \([24](https://arxiv.org/html/2609.35817#A1.E24)\) and omittingθ\\theta\-independent terms gives
−αt′\[1αtx¯θ,y\+1−αtVαt\(Vαt\+1−αt\)∑j≠ylogx¯θ,yx¯θ,j\]\.\-\\alpha\_\{t\}^\{\\prime\}\\\!\\left\[\\frac\{1\}\{\\alpha\_\{t\}\\bar\{x\}\_\{\\theta,y\}\}\+\\frac\{1\-\\alpha\_\{t\}\}\{V\\alpha\_\{t\}\(V\\alpha\_\{t\}\+1\-\\alpha\_\{t\}\)\}\\sum\_\{j\\neq y\}\\log\\frac\{\\bar\{x\}\_\{\\theta,y\}\}\{\\bar\{x\}\_\{\\theta,j\}\}\\right\]\.\(36\)On the interior\-time region,ε≤x¯θ,j≤V\\varepsilon\\leq\\bar\{x\}\_\{\\theta,j\}\\leq VandVαt\+1−αt≥VεV\\alpha\_\{t\}\+1\-\\alpha\_\{t\}\\geq V\\varepsilon\. Hence
1αtx¯θ,y=O\(1Vxθ,y\),\|1−αtVαt\(Vαt\+1−αt\)∑j≠ylogx¯θ,yx¯θ,j\|=O\(logVV\)\.\\frac\{1\}\{\\alpha\_\{t\}\\bar\{x\}\_\{\\theta,y\}\}=O\\\!\\left\(\\frac\{1\}\{Vx\_\{\\theta,y\}\}\\right\),\\qquad\\left\|\\frac\{1\-\\alpha\_\{t\}\}\{V\\alpha\_\{t\}\(V\\alpha\_\{t\}\+1\-\\alpha\_\{t\}\)\}\\sum\_\{j\\neq y\}\\log\\frac\{\\bar\{x\}\_\{\\theta,y\}\}\{\\bar\{x\}\_\{\\theta,j\}\}\\right\|=O\\\!\\left\(\\frac\{\\log V\}\{V\}\\right\)\.\(37\)Ifxθ,y=ω\(1/V\)x\_\{\\theta,y\}=\\omega\(1/V\), both terms areo\(1\)o\(1\)\. Boundedness of\|αt′\|\|\\alpha\_\{t\}^\{\\prime\}\|completes the proof\.
#### A\.2\.2Proof of Theorem 1
We first state the exact centered identity\. The proof then extracts the clean coordinate of the learned reverse posterior, takes the small\-step limit, and ends with an explicit finite\-step error bound\.
##### Theorem 1 \(Restated\)\.
Fix𝒙0=𝒆y\\bm\{x\}\_\{0\}=\\bm\{e\}\_\{y\}, a corrupted state𝒙t=𝒆m\\bm\{x\}\_\{t\}=\\bm\{e\}\_\{m\}withm≠ym\\neq y, and a time satisfying0<αt<10<\\alpha\_\{t\}<1\. ForΔt\>0\\Delta t\>0, lets=t−Δts=t\-\\Delta tand assume0<αt<αs<10<\\alpha\_\{t\}<\\alpha\_\{s\}<1\. Under the Duox0x\_\{0\}\-parameterization, substituting the fixed model output𝒙θ\(𝒙t,t\)\\bm\{x\}\_\{\\theta\}\(\\bm\{x\}\_\{t\},t\)into the exact reverse posterior gives the centered identity
\\fitbox0\.98KL\(δ\(𝒙0\)∥pθ\(𝒙s∣𝒙t\)\)\+logαs−αtVαs=logx¯θ,m−log\(Vαsxθ,y\+1−αs\)\.\\fitbox\{0\.98\}\{\\displaystyle\\mathrm\{KL\}\\\!\\left\(\\delta\(\\bm\{x\}\_\{0\}\)\\middle\\\|p\_\{\\theta\}\(\\bm\{x\}\_\{s\}\\mid\\bm\{x\}\_\{t\}\)\\right\)\+\\log\\frac\{\\alpha\_\{s\}\-\\alpha\_\{t\}\}\{V\\alpha\_\{s\}\}=\\log\\bar\{x\}\_\{\\theta,m\}\-\\log\\\!\\left\(V\\alpha\_\{s\}x\_\{\\theta,y\}\+1\-\\alpha\_\{s\}\\right\)\.\}\(38–39\)If such admissible reverse steps exist for arbitrarily smallΔt\\Delta t, then, asΔt→0\\Delta t\\to 0along these steps,
KL\(δ\(𝒙0\)∥pθ\(𝒙s∣𝒙t\)\)\+logαs−αtVαs⟶logx¯θ,m−logx¯θ,y\.\\mathrm\{KL\}\\\!\\left\(\\delta\(\\bm\{x\}\_\{0\}\)\\middle\\\|p\_\{\\theta\}\(\\bm\{x\}\_\{s\}\\mid\\bm\{x\}\_\{t\}\)\\right\)\+\\log\\frac\{\\alpha\_\{s\}\-\\alpha\_\{t\}\}\{V\\alpha\_\{s\}\}\\longrightarrow\\log\\bar\{x\}\_\{\\theta,m\}\-\\log\\bar\{x\}\_\{\\theta,y\}\.\(40\)
##### Proof\.
We compute the reverse probability of the single target token and then take its small\-step limit\. For any categorical distributionqq,
KL\(δ𝒆y∥q\)=−logq\(𝒆y\)\.\\mathrm\{KL\}\(\\delta\_\{\\bm\{e\}\_\{y\}\}\\\|q\)=\-\\log q\(\\bm\{e\}\_\{y\}\)\.\(41\)
Substitute𝒙θ\(𝒙t,t\)\\bm\{x\}\_\{\\theta\}\(\\bm\{x\}\_\{t\},t\)for𝒙0\\bm\{x\}\_\{0\}in Eq\. \([19](https://arxiv.org/html/2609.35817#A1.E19)\)\. At coordinateyy, the two terms containing𝒙t\\bm\{x\}\_\{t\}vanish becausem≠ym\\neq y\. Therefore
\\fitbox0\.98pθ\(𝒙s=𝒆y∣𝒙t=𝒆m\)=\(αs−αt\)xθ,y\+\(1−αt/αs\)\(1−αs\)/Vx¯θ,m=αs−αtVαsVαsxθ,y\+1−αsx¯θ,m\.\\fitbox\{0\.98\}\{\\displaystyle p\_\{\\theta\}\(\\bm\{x\}\_\{s\}=\\bm\{e\}\_\{y\}\\mid\\bm\{x\}\_\{t\}=\\bm\{e\}\_\{m\}\)=\\frac\{\(\\alpha\_\{s\}\-\\alpha\_\{t\}\)x\_\{\\theta,y\}\+\(1\-\\alpha\_\{t\}/\\alpha\_\{s\}\)\(1\-\\alpha\_\{s\}\)/V\}\{\\bar\{x\}\_\{\\theta,m\}\}=\\frac\{\\alpha\_\{s\}\-\\alpha\_\{t\}\}\{V\\alpha\_\{s\}\}\\frac\{V\\alpha\_\{s\}x\_\{\\theta,y\}\+1\-\\alpha\_\{s\}\}\{\\bar\{x\}\_\{\\theta,m\}\}\.\}\(42\)Taking−log\-\\logand using Eq\. \([41](https://arxiv.org/html/2609.35817#A1.E41)\) gives
KL\(δ\(𝒙0\)∥pθ\(𝒙s∣𝒙t\)\)\+logαs−αtVαs=logx¯θ,m−log\(Vαsxθ,y\+1−αs\)\.\\mathrm\{KL\}\\\!\\left\(\\delta\(\\bm\{x\}\_\{0\}\)\\middle\\\|p\_\{\\theta\}\(\\bm\{x\}\_\{s\}\\mid\\bm\{x\}\_\{t\}\)\\right\)\+\\log\\frac\{\\alpha\_\{s\}\-\\alpha\_\{t\}\}\{V\\alpha\_\{s\}\}=\\log\\bar\{x\}\_\{\\theta,m\}\-\\log\\\!\\left\(V\\alpha\_\{s\}x\_\{\\theta,y\}\+1\-\\alpha\_\{s\}\\right\)\.\(43\)
The added logarithm is independent ofθ\\thetaand removes the divergent part of the small\-step KL\. The model output remains fixed at the input\(𝒙t,t\)\(\\bm\{x\}\_\{t\},t\)whiles→ts\\to t\. Continuity givesVαsxθ,y\+1−αs→x¯θ,yV\\alpha\_\{s\}x\_\{\\theta,y\}\+1\-\\alpha\_\{s\}\\to\\bar\{x\}\_\{\\theta,y\}, proving the stated limit\. Since\(αs−αt\)/\(Vαs\)→0\(\\alpha\_\{s\}\-\\alpha\_\{t\}\)/\(V\\alpha\_\{s\}\)\\to 0, the raw KL diverges; only its finite model\-dependent part remains\. Weighting this part by−αt′/\(1−αt\)\-\\alpha\_\{t\}^\{\\prime\}/\(1\-\\alpha\_\{t\}\)gives Eq\. \([8](https://arxiv.org/html/2609.35817#S3.E8)\) on corrupted positions\.
For completeness, we quantify how close a finite reverse step is to this limit\. Define the nonsingular contrast
ℓLU,Δ:=logx¯θ,m−log\(Vαsxθ,y\+1−αs\)\.\\ell\_\{\\mathrm\{LU\},\\Delta\}:=\\log\\bar\{x\}\_\{\\theta,m\}\-\\log\\\!\\left\(V\\alpha\_\{s\}x\_\{\\theta,y\}\+1\-\\alpha\_\{s\}\\right\)\.\(44\)Forα∈\[η,1−η\]\\alpha\\in\[\\eta,1\-\\eta\],
\|ddαlog\(Vαxθ,y\+1−α\)\|=\|Vxθ,y−1\|Vαxθ,y\+1−α≤1η\.\\left\|\\frac\{d\}\{d\\alpha\}\\log\(V\\alpha x\_\{\\theta,y\}\+1\-\\alpha\)\\right\|=\\frac\{\|Vx\_\{\\theta,y\}\-1\|\}\{V\\alpha x\_\{\\theta,y\}\+1\-\\alpha\}\\leq\\frac\{1\}\{\\eta\}\.\(45\)The inequality follows by consideringVxθ,y≥1Vx\_\{\\theta,y\}\\geq 1andVxθ,y≤1Vx\_\{\\theta,y\}\\leq 1\. If\|αu′\|≤Mη\|\\alpha\_\{u\}^\{\\prime\}\|\\leq M\_\{\\eta\}andαu∈\[η,1−η\]\\alpha\_\{u\}\\in\[\\eta,1\-\\eta\]foru∈\[s,t\]u\\in\[s,t\], the mean\-value theorem gives
\|ℓLU,Δ−\(logx¯θ,m−logx¯θ,y\)\|≤\|αs−αt\|η≤MηηΔt\.\\left\|\\ell\_\{\\mathrm\{LU\},\\Delta\}\-\\left\(\\log\\bar\{x\}\_\{\\theta,m\}\-\\log\\bar\{x\}\_\{\\theta,y\}\\right\)\\right\|\\leq\\frac\{\|\\alpha\_\{s\}\-\\alpha\_\{t\}\|\}\{\\eta\}\\leq\\frac\{M\_\{\\eta\}\}\{\\eta\}\\Delta t\.\(46\)
Restricting to times for which these conditions hold on\[t−Δt,t\]\[t\-\\Delta t,t\], define
ℒLU,Δℓ:=𝔼t,𝒙t\[𝕀m≠y−αt′1−αtℓLU,Δ\]\.\\mathcal\{L\}\_\{\\mathrm\{LU\},\\Delta\}^\{\\ell\}:=\\mathbb\{E\}\_\{t,\\bm\{x\}\_\{t\}\}\\\!\\left\[\\mathbb\{I\}\_\{m\\neq y\}\\frac\{\-\\alpha\_\{t\}^\{\\prime\}\}\{1\-\\alpha\_\{t\}\}\\ell\_\{\\mathrm\{LU\},\\Delta\}\\right\]\.\(47\)Sincep\(m≠y∣y,t\)=\(1−αt\)\(1−1/V\)p\(m\\neq y\\mid y,t\)=\(1\-\\alpha\_\{t\}\)\(1\-1/V\), Eq\. \([46](https://arxiv.org/html/2609.35817#A1.E46)\) implies
\\fitbox0\.98\|ℒLU,Δℓ−ℒLUℓ\|≤𝔼t\[Mη1−αtMηηΔtp\(m≠y∣y,t\)\]=Mη2ηΔt\(1−1V\)≤Mη2ηΔt\.\\fitbox\{0\.98\}\{\\displaystyle\\left\|\\mathcal\{L\}\_\{\\mathrm\{LU\},\\Delta\}^\{\\ell\}\-\\mathcal\{L\}\_\{\\mathrm\{LU\}\}^\{\\ell\}\\right\|\\leq\\mathbb\{E\}\_\{t\}\\\!\\left\[\\frac\{M\_\{\\eta\}\}\{1\-\\alpha\_\{t\}\}\\frac\{M\_\{\\eta\}\}\{\\eta\}\\Delta t\\;p\(m\\neq y\\mid y,t\)\\right\]=\\frac\{M\_\{\\eta\}^\{2\}\}\{\\eta\}\\Delta t\\left\(1\-\\frac\{1\}\{V\}\\right\)\\leq\\frac\{M\_\{\\eta\}^\{2\}\}\{\\eta\}\\Delta t\.\}\(48\)Forαt=1−t\\alpha\_\{t\}=1\-t, this isΔt/η\\Delta t/\\eta; thusΔt≤ϵη\\Delta t\\leq\\epsilon\\etagives error at mostϵ\\epsilon\.
### A\.3Per\-Token Time Embeddings
We show that per\-token times preserve the forward marginal and derive their sampling\-time initialization\.
##### Marginal preservation\.
We average over the per\-token time and recover the original uniform\-state marginal\.
Forαt=1−t\\alpha\_\{t\}=1\-t, drawτi∼qt\\tau^\{i\}\\sim q\_\{t\}independently andRi\|τi∼Bernoulli\(τi\)R^\{i\}\\mid\\tau^\{i\}\\sim\\mathrm\{Bernoulli\}\(\\tau^\{i\}\)\. IfRi=0R^\{i\}=0, keep the clean token; otherwise draw uniformly from𝒱\\mathcal\{V\}\. Therefore
\\fitbox0\.98p\(xti=v∣x0i=y,t\)=𝔼τi∼qt\[\(1−τi\)𝕀v=y\+τi1V\]=\(1−𝔼qt\[τi\]\)𝕀v=y\+𝔼qt\[τi\]V\.\\fitbox\{0\.98\}\{\\displaystyle p\(x\_\{t\}^\{i\}=v\\mid x\_\{0\}^\{i\}=y,t\)=\\mathbb\{E\}\_\{\\tau^\{i\}\\sim q\_\{t\}\}\\\!\\left\[\(1\-\\tau^\{i\}\)\\mathbb\{I\}\_\{v=y\}\+\\tau^\{i\}\\frac\{1\}\{V\}\\right\]=\\left\(1\-\\mathbb\{E\}\_\{q\_\{t\}\}\[\\tau^\{i\}\]\\right\)\\mathbb\{I\}\_\{v=y\}\+\\frac\{\\mathbb\{E\}\_\{q\_\{t\}\}\[\\tau^\{i\}\]\}\{V\}\.\}\(49\)If𝔼qt\[τi\]=t\\mathbb\{E\}\_\{q\_\{t\}\}\[\\tau^\{i\}\]=t, then
p\(xti=v∣x0i=y,t\)=\(1−t\)𝕀v=y\+tV=αt𝕀v=y\+1−αtV,p\(x\_\{t\}^\{i\}=v\\mid x\_\{0\}^\{i\}=y,t\)=\(1\-t\)\\mathbb\{I\}\_\{v=y\}\+\\frac\{t\}\{V\}=\\alpha\_\{t\}\\mathbb\{I\}\_\{v=y\}\+\\frac\{1\-\\alpha\_\{t\}\}\{V\},\(50\)which is Eq\. \([12](https://arxiv.org/html/2609.35817#A1.E12)\)\. Independence across positions preserves the factorized sequence law\. For a non\-linear schedule, the same proof requires𝔼qt\[τi\]=1−αt\\mathbb\{E\}\_\{q\_\{t\}\}\[\\tau^\{i\}\]=1\-\\alpha\_\{t\}\.
For0<t<10<t<1, the implementation uses
qt\(τ\)=Beta\(ct,c\(1−t\)\),c=2,q\_\{t\}\(\\tau\)=\\mathrm\{Beta\}\(ct,c\(1\-t\)\),\\qquad c=2,\(51\)𝔼qt\[τ\]=ctct\+c\(1−t\)=t\.\\mathbb\{E\}\_\{q\_\{t\}\}\[\\tau\]=\\frac\{ct\}\{ct\+c\(1\-t\)\}=t\.\(52\)At the endpoints, takeq0=δ0q\_\{0\}=\\delta\_\{0\}andq1=δ1q\_\{1\}=\\delta\_\{1\}; finite\-precision clipping introduces only the corresponding endpoint perturbation\. Note thatxti=x0ix\_\{t\}^\{i\}=x\_\{0\}^\{i\}does not implyRi=0R^\{i\}=0, because uniform replacement redraws the clean token with probability1/V1/V\.
##### Sampling\-time initialization\.
Prompt positions are fixed conditions and receiveτ=0\\tau=0, while every non\-prompt position is randomized\. Underτ∼Beta\(1,1\)\\tau\\sim\\mathrm\{Beta\}\(1,1\)andToken is random\|τ∼Bernoulli\(τ\)\\text\{Token is random\}\\mid\\tau\\sim\\mathrm\{Bernoulli\}\(\\tau\), Bayes’ rule gives
τ\|Token is random∼Beta\(2,1\)\.\\tau\\mid\\text\{Token is random\}\\sim\\mathrm\{Beta\}\(2,1\)\.\(53\)Thus Algorithm[2](https://arxiv.org/html/2609.35817#algorithm2)usesBeta\(2,1\)\\mathrm\{Beta\}\(2,1\)as initialization\.
##### Use during denoising\.
We use the same interpretation throughout denoising\.
Committed positions receiveτi=0\\tau^\{i\}=0and act as clean conditioning context; positive\-time positions remain denoising targets\. This yields the confidence\-based easy\-to\-hard sampler used by LUDI\.
## Appendix BAlgorithm Details
### B\.1Training and Sampling
Algorithm 1Training of LUDIInput :Clean data
𝐱0\\mathbf\{x\}\_\{0\}, noise schedule
αt\\alpha\_\{t\}, model
𝐱θ\\mathbf\{x\}\_\{\\theta\}
Output :Trained model parameters
θ\\theta
1for*each training step*do
2Sample global time
t∼𝒰\(0,1\)t\\sim\\mathcal\{U\}\(0,1\);
3for*each token positionii*do
//sample token\-wise corruption time
4
τi∼Beta\(2t,2\(1−t\)\)\\tau^\{i\}\\sim\\mathrm\{Beta\}\(2t,\\,2\(1\-t\)\);
5Corrupt
𝐱0i\\mathbf\{x\}\_\{0\}^\{i\}: with probability
τi\\tau^\{i\}, replace by a uniform token
→𝐱ti\\rightarrow\\mathbf\{x\}\_\{t\}^\{i\};
6end for
//model predicts clean tokens from noisy inputs and per\-token times
7
𝐱^0=Model\(𝐱t,𝝉\)\\hat\{\\mathbf\{x\}\}\_\{0\}=\\textnormal\{\{Model\}\}\(\\mathbf\{x\}\_\{t\},\\;\\bm\{\\tau\}\);
//compute LU loss on corrupted positions \(Eq\.[8](https://arxiv.org/html/2609.35817#S3.E8)\)
8
ℒ=−αt′1−αt∑i:𝐱ti≠𝐱0i\(logx¯θ,mi−logx¯θ,yi\)\\mathcal\{L\}=\\frac\{\-\\alpha\_\{t\}^\{\\prime\}\}\{1\-\\alpha\_\{t\}\}\\\!\\sum\_\{i:\\mathbf\{x\}\_\{t\}^\{i\}\\neq\\mathbf\{x\}\_\{0\}^\{i\}\}\\bigl\(\\log\\bar\{x\}\_\{\\theta,m\}^\{i\}\-\\log\\bar\{x\}\_\{\\theta,y\}^\{i\}\\bigr\);
9Update
θ\\thetausing
∇θℒ\\nabla\_\{\\theta\}\\mathcal\{L\};
10end for
Algorithm 2Sampling of LUDIInput :Prompt tokens, trained model
𝐱θ\\mathbf\{x\}\_\{\\theta\}, scheduler function
hh, confidence threshold
γ\\gamma
Output :Generated clean sequence
𝐱\\mathbf\{x\}
1Initialize
𝐱\\mathbf\{x\}: prompt tokens kept unchanged, all other tokens sampled uniformly from
𝒱\\mathcal\{V\};
2Initialize
𝝉\\bm\{\\tau\}: for non\-prompt tokens
τi∼Beta\(2,1\)\\tau^\{i\}\\sim\\mathrm\{Beta\}\(2,1\), for prompt tokens
τi=0\\tau^\{i\}=0;
3while*max\(𝛕\)\>0\\max\(\\bm\{\\tau\}\)\>0*do
4
𝐱^0=Model\(𝐱,𝝉\)\\hat\{\\mathbf\{x\}\}\_\{0\}=\\textnormal\{\{Model\}\}\(\\mathbf\{x\},\\;\\bm\{\\tau\}\);
//compute next per\-token times via schedulerhh
5
𝝉next=h\(𝝉,𝐱^0\)\\bm\{\\tau\}\_\{\\text\{next\}\}=h\(\\bm\{\\tau\},\\hat\{\\mathbf\{x\}\}\_\{0\}\);
//For confidence\-based scheduler: high\-confidence tokens are committed
//
cℓ:=max\(𝐱^0ℓ\)c^\{\\ell\}:=\\max\(\\hat\{\\mathbf\{x\}\}\_\{0\}^\{\\ell\}\)ifτℓ\>0\\tau^\{\\ell\}\>0else 0
//
τnextℓ=0\\tau\_\{\\text\{next\}\}^\{\\ell\}=0ifcℓ\>γc^\{\\ell\}\>\\gammaorℓ=argmaxici\\ell=\\arg\\max\_\{i\}c^\{i\}; otherwiseτnextℓ=τℓ\\tau\_\{\\text\{next\}\}^\{\\ell\}=\\tau^\{\\ell\}
6update
𝐱\\mathbf\{x\}from
𝐱τ\\mathbf\{x\}\_\{\\tau\}to
𝐱τnext\\mathbf\{x\}\_\{\\tau\_\{\\text\{next\}\}\}using the reverse transition \(Eq\.[3](https://arxiv.org/html/2609.35817#S2.E3)\);
7
𝝉←𝝉next\\bm\{\\tau\}\\leftarrow\\bm\{\\tau\}\_\{\\text\{next\}\};
8end while
9return
𝐱\\mathbf\{x\};
The training procedure of LUDI is summarized in Algorithm[1](https://arxiv.org/html/2609.35817#algorithm1)\. It introduces per\-token time embeddings𝝉\\bm\{\\tau\}to provide token\-level corruption hints, and optimizes the less uniform \(LU\) loss \(Eq\.[8](https://arxiv.org/html/2609.35817#S3.E8)\) that directly encourages each reverse transition toward the clean token\. The conditional sampling, detailed in Algorithm[2](https://arxiv.org/html/2609.35817#algorithm2), leverages a confidence\-based schedulerhhto progressively commit high\-confidence tokens as conditions, enabling few\-step generation for complex reasoning tasks\.
### B\.2Details of AR to Block Diffusion
##### Label shift\.
AR pretraining predictsxℓx^\{\\ell\}from a clean prefixx<ℓx^\{<\\ell\}, while block diffusion forces the model to predictxℓx^\{\\ell\}at positionℓ\\ell\. We resolve this mismatch by shifting the labels left by one position, while still allowing tokenℓ−1\\ell\-1to attend to tokenℓ\\ell\. This changes only the prediction position, preserving both the diffusion paradigm and the autoregressive prediction preference\.
##### Complementary views\.
Following Fast dLLM v2\([Wu et al\., 2025a](https://arxiv.org/html/2609.35817#bib.bib12)\), we train with two complementary mask patterns per sequence: positions kept clean in the first view are masked in the second, and vice versa\. In the UDLM setting, unmasked positions receive𝐱0\\mathbf\{x\}\_\{0\}and masked positions receive a random token\. Both views share a single forward pass, forcing every position to serve as both a clean target and a noisy input\. This doubling of supervision is valuable in our low\-data regime\.
##### Context\-causal attention mask\.
NBDiff\([Tian et al\., 2025](https://arxiv.org/html/2609.35817#bib.bib11)\)proposed the context\-causal attention mask\. The prefix observes a strictly causal mask, while the active denoising block enjoys full intra\-block bidirectionality and causal access to the prefix\. This yields the attention mask pattern shown in Figure[4](https://arxiv.org/html/2609.35817#A2.F4), where both noisy and clean sequences are used during training\. The clean sequence provides keys and values from causal attention for the subsequent noisy block\. This preserves the left\-to\-right inductive bias in already committed context, preventing instability from premature future exposure\. At inference the same pattern holds \(Figure[4](https://arxiv.org/html/2609.35817#A2.F4)\(b\)\): the prompt and past blocks form a frozen causal prefix, and only the current block is denoised bidirectionally, enabling KV\-cache reuse and parallel token decoding\.
##### AR loss\.
An auxiliary autoregressive loss is computed on the clean sequence\. This incurs no extra forward cost and serves dual purposes: it prevents catastrophic forgetting of the pretrained model, and it supplies a stable gradient signal that counters the noisier diffusion loss early in training\. The objective isℒ=ℒdiff\+λℒAR\\mathcal\{L\}=\\mathcal\{L\}\_\{\\text\{diff\}\}\+\\lambda\\mathcal\{L\}\_\{\\text\{AR\}\}with a weightλ\\lambda\.
\(a\)Training mask\.
\(b\)Inference mask\.
Figure 4:Context\-Causal attention masks for block\-diffusion adaptation\. \(a\) Training mask: the prefix uses a lower\-triangular causal mask; the active block attends bidirectionally within itself and causally to the prefix\. \(b\) Inference mask: generated blocks form a causal prefix; only the current block is denoised bidirectionally\.
## Appendix CExperiments
### C\.1Experimental Setup for Small\-Scale LUDI
##### 170M models\.
We train four 170M UDLMs on OpenWebText\([Gokaslan et al\., 2019](https://arxiv.org/html/2609.35817#bib.bib27)\): LUDI with the LU loss, and baselines trained with Duo\([Sahoo et al\., 2025](https://arxiv.org/html/2609.35817#bib.bib4)\), SDDLM\([Zhu et al\., 2025c](https://arxiv.org/html/2609.35817#bib.bib5)\), and SDDLM\-v1\([Zhu et al\., 2025c](https://arxiv.org/html/2609.35817#bib.bib5)\)losses\. The denoiser is a 12\-block Transformer with hidden size 768, 12 attention heads, a 128\-dimensional conditioning embedding, sequence length 1024, dropout 0\.1, the GPT\-2 vocabulary of 50,257 tokens\([Radford et al\., 2019](https://arxiv.org/html/2609.35817#bib.bib29)\), and bfloat16 precision\. Each run is trained for 500K iterations with batch size 512\. We use AdamW with learning rate3×10−43\\times 10^\{\-4\}, betas\(0\.9,0\.999\)\(0\.9,0\.999\)\. Generative perplexity is computed with GPT\-2 Large\([Radford et al\., 2019](https://arxiv.org/html/2609.35817#bib.bib29)\), and sample diversity is measured by mean token\-level entropy in bits\. As in RADD\([Ou et al\., 2025](https://arxiv.org/html/2609.35817#bib.bib8)\), Gumbel\-based categorical sampling is performed in float64 precision\. For released masked\-diffusion baselines, we evaluate RADD\([Ou et al\., 2025](https://arxiv.org/html/2609.35817#bib.bib8)\)and MDLM\([Sahoo et al\., 2024](https://arxiv.org/html/2609.35817#bib.bib1)\)checkpoints trained on the same dataset for 400K and 1M steps, respectively, using the same generation metrics\.
##### 1B models\.
We train four 1B\-parameter UDLMs from scratch on FineWeb\([Penedo et al\., 2024](https://arxiv.org/html/2609.35817#bib.bib28)\)with the LLaMA tokenizer\([Touvron et al\., 2023](https://arxiv.org/html/2609.35817#bib.bib30)\)\. The denoiser is a 20\-block Transformer with hidden size 2048, 16 attention heads, MLP ratio 3\.5, a 128\-dimensional conditioning embedding, sequence length 2048, dropout 0\.0, a 32,000\-token vocabulary and bfloat16 precision\. Following the SMDM implementation\([Nie et al\., 2025a](https://arxiv.org/html/2609.35817#bib.bib47)\), we mix 1% variable\-length FineWeb examples into training\. Each run is trained for 500K iterations with batch size 256\. We use AdamW with learning rate2×10−42\\times 10^\{\-4\}, betas\(0\.9,0\.95\)\(0\.9,0\.95\), weight decay 0\.1\. We evaluate likelihood\-based multiple\-choice accuracy on PIQA\([Bisk et al\., 2020](https://arxiv.org/html/2609.35817#bib.bib41)\), SIQA\([Sap et al\., 2019](https://arxiv.org/html/2609.35817#bib.bib42)\), ARC\-Easy\([Clark et al\., 2018](https://arxiv.org/html/2609.35817#bib.bib43)\), HellaSwag\([Zellers et al\., 2019](https://arxiv.org/html/2609.35817#bib.bib44)\), OpenBookQA\([Mihaylov et al\., 2018](https://arxiv.org/html/2609.35817#bib.bib45)\), and RACE\([Lai et al\., 2017](https://arxiv.org/html/2609.35817#bib.bib46)\)\. Following SDDLM\([Zhu et al\., 2025c](https://arxiv.org/html/2609.35817#bib.bib5)\), each candidate answer is scored with the ELBO\-based likelihood estimate induced by Eq\. \([5](https://arxiv.org/html/2609.35817#S2.E5)\)\. We report accuracy \(acc\) for PIQA, SIQA, ARC\-Easy, and RACE, and length\-normalized accuracy \(acc\_norm\) for HellaSwag and OpenBookQA\.
### C\.2Experimental Setup for LUDI\-7B
##### Baselines\.
We compare LUDI\-7B against autoregressive \(AR\) baselines and state\-of\-the\-art diffusion language models at comparable scales\. AR baselines include Qwen2\.5\-7B\-Instruct\([Team, 2024](https://arxiv.org/html/2609.35817#bib.bib25)\)\(the initialization checkpoint\)\. Masked diffusion baselines include Dream 7B\([Ye et al\., 2025](https://arxiv.org/html/2609.35817#bib.bib31)\)\(adapted from Qwen2\.5\-7B\), LLaDA 8B\([Nie et al\., 2025b](https://arxiv.org/html/2609.35817#bib.bib2)\), LLaDA\-1\.5 8B\([Zhu et al\., 2025a](https://arxiv.org/html/2609.35817#bib.bib32)\), LLaDA\-MoE 7B\([Zhu et al\., 2025b](https://arxiv.org/html/2609.35817#bib.bib33)\)and Fast\-dLLM v2\([Wu et al\., 2025a](https://arxiv.org/html/2609.35817#bib.bib12)\)\.
##### Evaluation benchmarks\.
We evaluate on a diverse set of tasks spanning code generation, mathematical reasoning, instruction following, and knowledge\-intensive question answering\. For code generation, we use HumanEval\([Chen et al\., 2021](https://arxiv.org/html/2609.35817#bib.bib35)\)and HumanEval\+\([Liu et al\., 2023](https://arxiv.org/html/2609.35817#bib.bib34)\), as well as MBPP\([Austin et al\., 2021](https://arxiv.org/html/2609.35817#bib.bib36)\)and MBPP\+[Liu et al\. \(2023\)](https://arxiv.org/html/2609.35817#bib.bib34)\. Mathematical reasoning is assessed on GSM8K[Cobbe et al\. \(2021\)](https://arxiv.org/html/2609.35817#bib.bib37)\. Instruction following is measured with IFEval[Zhou et al\. \(2023\)](https://arxiv.org/html/2609.35817#bib.bib38)\. For knowledge\-intensive tasks, we adopt MMLU[Hendrycks et al\. \(2020\)](https://arxiv.org/html/2609.35817#bib.bib40)and GPQA[Rein et al\. \(2023\)](https://arxiv.org/html/2609.35817#bib.bib39)\. All code benchmarks are evaluated using the EvalPlus framework[Liu et al\. \(2023\)](https://arxiv.org/html/2609.35817#bib.bib34)\. For all tasks we report standard accuracy metrics following the evaluation protocol of Fast\-dLLM v2[Wu et al\. \(2025a\)](https://arxiv.org/html/2609.35817#bib.bib12)\.
### C\.3Mathematical Reasoning
We supplement the mathematical\-reasoning evaluation with results on MATH500[Lightman et al\. \(2024\)](https://arxiv.org/html/2609.35817#bib.bib49), as reported in Table[4](https://arxiv.org/html/2609.35817#A3.T4)\. Under the Dolci training setting, both UDLM and MDLM variants obtain a substantially lower score on MATH500\. This pattern is shared across diffusion formulations and is therefore not specific to LUDI\.
We attribute this shared behavior primarily to a distribution mismatch between the mathematical content in Dolci\-Instruct\-SFT and the evaluation benchmarks\. The relevant sequence\-length statistics are summarized below:
- •Training data\.Mathematical examples in Dolci\-Instruct\-SFT are typically long, scenario\-based questions with detailed explanations, with a median question length of 920 tokens\. Mathematical examples in Nemotron\-Post\-Train v2, many of which originate from DeepSeek\-R1, are more concise, with a median question length of 217 tokens\.
- •Evaluation data\.Questions in GSM8K and MATH500 have median lengths of 257 and 143 tokens, respectively\.
Table 4:Mathematical\-reasoning accuracy with Dolci and Dolci\-Nemo mixture\.Nemotron\-Post\-Train v2[Nathawani et al\. \(2025\)](https://arxiv.org/html/2609.35817#bib.bib48)therefore provides mathematical examples whose length and task style more closely match the evaluation distribution\. To reduce the mismatch while keeping the training\-set size comparable, we constructDolci–Nemo Mix: we remove the Other, Tool Use, Safety, and Multilingual categories from Dolci\-Instruct\-SFT and replace them with mathematics, code, and STEM examples from Nemotron\-Post\-Train v2\.
As shown in Table[4](https://arxiv.org/html/2609.35817#A3.T4), introducing Nemotron\-Post\-Train v2 consistently improves mathematical reasoning across the compared objectives\. LUDI improves from 75\.2 to 82\.1 on GSM8K and from 49\.0 to 52\.6 on MATH500, achieving the strongest MATH500 result among the compared methods\. Other methods also improve on both benchmarks\. The gains across both UDLM and MDLM variants support the distribution\-mismatch analysis and show that LUDI benefits effectively from task\-aligned mathematical data\.
A characteristic failure mode in diffusion generation is repetitive looping, in which the model repeatedly emits the same span\. The following GSM8K output illustrates this behavior:
Question:A food caterer was told to prepare gourmet hot dogs for 36 guests\. While most people would only eat one hotdog, he prepared enough for half of the guests to be able to have two hotdogs\. However, 40 guests showed up, and everyone wanted a second hotdog\. How many guests did not get a second hotdog?Generated answer:The caterer prepared enough hot dogs for half of the guests to have two hot dogs\. This means that36/2=1836/2=18guests would have two hot dogs\. The caterer prepared enough hot dogs for half of the guests to have two hot dogs\. This means that36/2=1836/2=18guests would have two hot dogs\. …*The two sentences repeat approximately 55 times, producing approximately 1,634 words\.*
Among the incorrect MATH500 responses produced by models trained on Dolci, repetitive loops account for 21\.5% of LUDI’s errors, compared with 27\.8% for the MDLM baseline Fast\-dLLM v2\. LUDI thus moderately reduces this diffusion\-specific failure mode\. Its uniform\-state formulation allows generated tokens to be revised rather than irrevocably committed, which can help interrupt locally self\-reinforcing spans\.
Table 5:End\-to\-end efficiency comparison between LUDI \(confidence threshold 0\.7\) and the autoregressive baseline Qwen2\.5\-7B\-Instruct on GSM8K across batch sizes 1–32\. Throughput speedup is computed as LUDI tokens per second divided by autoregressive tokens per second\.
### C\.4End\-to\-End Inference Efficiency
To provide a comprehensive efficiency profile, we conduct end\-to\-end efficiency evaluations on GSM8K for LUDI \(confidence threshold 0\.7\) and the autoregressive baseline Qwen2\.5\-7B\-Instruct at batch sizes 1, 4, 16, and 32\. The results are summarized in Table[5](https://arxiv.org/html/2609.35817#A3.T5)\.
From Table[5](https://arxiv.org/html/2609.35817#A3.T5), we make three observations\. \(1\)Decoding efficiency\.LUDI achieves throughput speedups at batch sizes 1, 4, and 16 \(1\.92×1\.92\\times,1\.30×1\.30\\times, and1\.03×1\.03\\times, respectively\), while the autoregressive baseline slightly leads at batch size 32 because of its more compact key\-value cache\. \(2\)Memory\.Owing to parallel multi\-token decoding, LUDI’s peak memory is approximately 1\.9–2\.0×\\timesthat of the autoregressive baseline\. \(3\)End\-to\-end latency\.LUDI exhibits higher latency because it tends to produce longer reasoning chains\. This behavior reflects the training\-data distribution: the median response length in Dolci’s mathematics subset is 2,114 tokens, compared with 217 tokens for GSM8K\.
In practice, the theoretical3×3\\timesspeedup is not fully realized end to end\. Inference efficiency can be further improved through dLLM\-specific infrastructure\. These system\-level techniques are orthogonal to our modeling contributions, and we leave their integration to future work\.
\(a\)GPU Memory Usage\.\(b\)Forward and Backward Time\.
Figure 5:Computational efficiency simulation in a long\-sequence, large\-vocabulary setting with batch size 10, sequence length 1024, and vocabulary size 50,000\.\(a\)512 sampling steps\.\(b\)1024 sampling steps\.
Figure 6:Generative perplexity over training checkpoints\. Lower PPL is better\. Lines connect all 10K\-spaced checkpoints from 10K to 500K steps\.
### C\.5Efficiency of the LU Loss
We explain why the LU loss in Eq\. \([8](https://arxiv.org/html/2609.35817#S3.E8)\) admits a more efficient implementation than the ELBO loss in Eq\. \([5](https://arxiv.org/html/2609.35817#S2.E5)\)\. The gain comes from a simpler vocabulary\-wide computation and from eliminating large intermediate tensors\.
##### Fused implementation of the LU loss\.
Let𝐱θ\\mathbf\{x\}\_\{\\theta\}be the model prediction after softmax and recall thatx¯θ,i=αtVxθ,i\+\(1−αt\)\\bar\{x\}\_\{\\theta,i\}=\\alpha\_\{t\}Vx\_\{\\theta,i\}\+\(1\-\\alpha\_\{t\}\)\. The LU loss is
ℒLUℓ=−αt′1−αt\(logx¯θ,m−logx¯θ,y\)\.\\mathcal\{L\}\_\{\\mathrm\{LU\}\}^\{\\ell\}=\\frac\{\-\\alpha\_\{t\}^\{\\prime\}\}\{1\-\\alpha\_\{t\}\}\\bigl\(\\log\\bar\{x\}\_\{\\theta,m\}\-\\log\\bar\{x\}\_\{\\theta,y\}\\bigr\)\.\(54\)A direct implementation first materializes𝐱θ\\mathbf\{x\}\_\{\\theta\}and𝐱¯θ\\bar\{\\mathbf\{x\}\}\_\{\\theta\}inℝB×L×V\\mathbb\{R\}^\{B\\times L\\times V\}and then selects the entries atmmandyy\. Our Triton operator instead processes each token row as a stream\. It computes the softmax normalization with an online log\-sum\-exp reduction, keeps the row statistics and the two selected logits in registers, and immediately evaluateslogx¯θ,m\\log\\bar\{x\}\_\{\\theta,m\}andlogx¯θ,y\\log\\bar\{x\}\_\{\\theta,y\}\. The full probability tensors therefore never need to be written to or read from global memory\. The same compact row statistics are reused by the training operator, so normalization, index selection, and loss evaluation do not become separate framework operations\.
##### Comparison with the ELBO loss\.
The ELBO loss additionally contains the vocabulary\-wide term∑j=1Vlogx¯θ,j\\sum\_\{j=1\}^\{V\}\\log\\bar\{x\}\_\{\\theta,j\}\. After obtaining the softmax normalization, it must still formx¯θ,j\\bar\{x\}\_\{\\theta,j\}, apply the logarithm, and reduce the result over every vocabulary entry\. In a standard PyTorch implementation, these stages create severalB×L×VB\\times L\\times Vintermediates and launch multiple kernels\. By comparison, LU performs the vocabulary scan needed for normalization and then applies the remaining nonlinear operations only atmmandyy\. Its fused implementation therefore uses fewer element\-wise operations and reductions, launches fewer kernels, and transfers substantially less data between registers and global memory\. These constant\-factor savings are especially important for long sequences and large vocabularies\. As shown in Figure[5](https://arxiv.org/html/2609.35817#A3.F5), the optimized LU operator achieves a4\.39×4\.39\\timesspeedup and a3\.45×3\.45\\timesmemory reduction relative to the PyTorch ELBO baseline\.
### C\.6Convergence
Figure[6](https://arxiv.org/html/2609.35817#A3.F6)compares the training trajectories of the four UDLM objectives under the same evaluation protocol\. LUDI improves rapidly in the early stage and remains in a low\-perplexity regime as training proceeds\. At 500K steps, LUDI obtains the lowest generative perplexity under both 512\-step and 1024\-step sampling, whereas Duo and SDDLM plateau at substantially higher PPL\. These trajectories support the main\-text observation that the LU loss provides a cleaner and more sample\-efficient training signal\.
### C\.7Quality and Diversity
\(a\)LUDI: Gen\.PPL 37\.9, entropy 7\.50\.
\(b\)SDDLM\-v1: Gen\.PPL 61\.2, entropy 7\.82\.
\(c\)Duo: Gen\.PPL 89\.0, entropy 7\.98\.
\(d\)SDDLM: Gen\.PPL 123\.1, entropy 8\.10\.
Figure 7:Density of per\-sample generative PPL and sample entropy for the four losses at the 500K checkpoint, evaluated with 1024\-step sampling and 1024 generated samples per loss\. Each panel uses its own axis range and density scale for readability\. Darker cells indicate more samples within the corresponding panel, and the white dot marks the method mean\.Figure[7](https://arxiv.org/html/2609.35817#A3.F7)compares the per\-sample distributions of the four losses at the 500K checkpoint, using 1024 denoising steps and 1024 samples for each loss\. LUDI places a large fraction of its samples in the low\-Gen\.PPL region, with a mean Gen\.PPL of 37\.9 and a mean entropy of 7\.50\. Its samples are not concentrated in an extremely low\-entropy region, suggesting that the lower Gen\.PPL is not accompanied by an obvious entropy collapse\.相似文章
Set Diffusion:在自回归与扩散之间插值令牌顺序以实现快速灵活的解码
Set Diffusion 引入了一类新的语言模型,通过在灵活位置、灵活长度的令牌集合上分解令牌生成,在自回归模型和扩散模型之间进行插值。这使得解码速度更快,令牌排序更灵活,在推理、摘要和无条件生成任务上实现了更好的速度-质量权衡。
PSD: 通过并行推测解码推动扩散大语言模型的帕累托前沿
本文介绍了一种无需训练的框架——并行推测解码(PSD),它通过同时提升空间和时间效率来加速扩散大语言模型的推理,每次前向传递最多可处理5.5×的token数,且质量与贪婪解码相当。
大型语言扩散模型的不确定性量化
本文首次系统研究了大型语言扩散模型(LLDMs)的不确定性量化(UQ),提出了从迭代去噪过程中衍生的轻量级零样本不确定性信号,并表明LLDMs能够在实现快速推理的同时,提供可靠的幻觉检测,与基于采样的基线方法相比,计算开销降低高达100倍。
@_akhaliq: Unlocking Lossless Speedups in LLMs via Discrete Diffusion 论文: https://huggingface.co/papers/2609.04010…
本文提出了一种使用离散扩散的方法,以在大型语言模型中实现无损加速,旨在提高效率而不影响性能。
Semantic DLM+:通过转移核设计中的偏差-方差权衡改进扩散语言模型
本文从偏差-方差角度对扩散语言模型进行了理论分析,识别了掩码扩散与均匀扩散核之间的权衡。提出了SemDLM+,通过添加全局转移和语义频率惩罚来克服语义盆地问题,在LM1B和OpenWebText基准上实现了有竞争力的生成质量。