FastGuide: Accelerating Reward Guidance for Diffusion Large Language Models

arXiv cs.CL Papers

Summary

FastGuide 是一种自适应的并行与自回归混合解码方法,通过复用奖励模型反向传播计算、利用 KV 缓存和稀疏注意力重计算来加速扩散大语言模型的梯度奖励引导,在三个奖励基准上实现最高 4.4× 加速同时保持相近生成质量。

arXiv:2609.36202v1 Announce Type: new Abstract: Gradient-based reward guidance provides a flexible way to use downstream reward models to control masked diffusion language models at inference time. However, its computational cost remains high as each decoding iteration incurs expensive diffusion model forward passes and reward model backpropagation steps. To address this, we introduce FastGuide, an adaptive hybrid of parallel and autoregressive decoding to accelerate reward guidance for diffusion language models. In analogy to parallel decoding, FastGuide amortizes the cost of reward model backpropagation by computing guidance once per decoding step and reusing it to generate multiple tokens. Within each decoding step, FastGuide makes diffusion forward passes autoregressive by unmasking tokens one at a time while efficiently recomputing token distributions after each unmasking by utilizing KV caching techniques and sparse recomputation of attention. Lastly, to adapt hybrid decoding to the model's confidence, FastGuide defers any token that the model is unconfident about under its recomputed distribution. Experiments on three reward benchmarks demonstrate that FastGuide is up to $4.4\times$ faster than sequential reward-guided decoding while retaining similar generation quality.
Original Article
View Cached Full Text

Cached at: 09/30/26, 09:50 AM

# FastGuide: Accelerating Reward Guidance for Diffusion Large Language Models
Source: [https://arxiv.org/html/2609.36202](https://arxiv.org/html/2609.36202)
###### Abstract

Gradient\-based reward guidance provides a flexible way to use downstream reward models to control masked diffusion language models at inference time\. However, its computational cost remains high as each decoding iteration incurs expensive diffusion model forward passes and reward model backpropagation steps\. To address this, we introduce FastGuide, an adaptive hybrid of parallel and autoregressive decoding to accelerate reward guidance for diffusion language models\. In analogy to parallel decoding, FastGuide amortizes the cost of reward model backpropagation by computing guidance once per decoding step and reusing it to generate multiple tokens\. Within each decoding step, FastGuide makes diffusion forward passes autoregressive by unmasking tokens one at a time while efficiently recomputing token distributions after each unmasking by utilizing KV caching techniques and sparse recomputation of attention\. Lastly, to adapt hybrid decoding to the model’s confidence, FastGuide defers any token that the model is unconfident about under its recomputed distribution\. Experiments on three reward benchmarks demonstrate that FastGuide is up to4\.4×4\.4\\timesfaster than sequential reward\-guided decoding while retaining similar generation quality\.

## 1Introduction

Diffusion Large Language Models \(dLLMs\) have emerged as a promising alternative to traditional autoregressive models and have been praised for their reasoning power and generation speed\([Ye et al\., 2025a](https://arxiv.org/html/2609.36202#bib.bib10);[Nie et al\., 2026](https://arxiv.org/html/2609.36202#bib.bib24)\)\. A popular class of dLLMs is discrete masked dLLMs\([Austin et al\., 2021](https://arxiv.org/html/2609.36202#bib.bib12);[Sahoo et al\., 2024](https://arxiv.org/html/2609.36202#bib.bib11)\), which are trained on partially masked token sequences to predict marginal distributions over the vocabulary for each masked token\. At inference time, dLLM generation starts from a fully masked sequence of fixed length and decoding algorithms iteratively decide a\) where to unmask, and b\) which token values to place at the selected positions\. For instance, a common strategy is to compute a confidence score for each token position, pick the most confident token, and sample from the corresponding marginal distribution of that token\([Ye et al\., 2025b](https://arxiv.org/html/2609.36202#bib.bib9)\)\.

As these models begin to be deployed in practical settings\([Inception Labs, 2025](https://arxiv.org/html/2609.36202#bib.bib21);[Google DeepMind, 2025](https://arxiv.org/html/2609.36202#bib.bib22)\), an important area of research is ensuring alignment of the dLLM outputs to desired objectives, e\.g\., instruction following, truthfulness, or safety\([Wen et al\., 2026](https://arxiv.org/html/2609.36202#bib.bib1);[Xiong et al\., 2026](https://arxiv.org/html/2609.36202#bib.bib2)\)\. A recent line of work has explored training\-free gradient\-based guidance algorithms for dLLM alignment using a differentiable reward model\([Murata et al\., 2024](https://arxiv.org/html/2609.36202#bib.bib16);[Tejaswi et al\., 2026](https://arxiv.org/html/2609.36202#bib.bib15)\)\. For these approaches, at each diffusion timestepttwith partially masked sequence𝒙t\{\\bm\{x\}\}\_\{t\}and unmasked tokens𝒙tu\{\\bm\{x\}\}\_\{t\}^\{\\text\{u\}\}, the dLLM logits at each token positioniiare steered via the updates

ℓtguided​\(xti∣𝒙tu\)=ℓtunguided​\(xti∣𝒙tu\)⏟base dLLM logits\+𝒓t​\(xti∣𝒙tu\)⏟guidance vector\.\\ell\_\{t\}^\{\\text\{guided\}\}\(x\_\{t\}^\{i\}\\mid\{\\bm\{x\}\}\_\{t\}^\{\\text\{u\}\}\)=\\underbrace\{\\ell\_\{t\}^\{\\text\{unguided\}\}\(x\_\{t\}^\{i\}\\mid\{\\bm\{x\}\}\_\{t\}^\{\\text\{u\}\}\)\}\_\{\\text\{base dLLM logits\}\}\+\\underbrace\{\{\\bm\{r\}\}\_\{t\}\(x\_\{t\}^\{i\}\\mid\{\\bm\{x\}\}\_\{t\}^\{\\text\{u\}\}\)\}\_\{\\text\{guidance vector\}\}\.\(1\)Above,ℓtunguided\\ell\_\{t\}^\{\\text\{unguided\}\}denotes the per\-token dLLM logits conditioned on the unmasked tokens𝒙tu\{\\bm\{x\}\}\_\{t\}^\{\\text\{u\}\}, and𝒓t\{\\bm\{r\}\}\_\{t\}denotes the corresponding per\-token guidance vector\. The guidance vector is estimated by cheaply constructing a fully unmasked sequence, evaluating that sequence with a reward model, and computing gradients of the reward model with respect to its input to steer the dLLM logits\. Although gradient\-based guidance can outperform alternatives such as Best\-Of\-N sampling\([Tejaswi et al\., 2026](https://arxiv.org/html/2609.36202#bib.bib15)\), the sampling speed of these approaches is slow, as shown in the top row of Figure[1](https://arxiv.org/html/2609.36202#S1.F1)b\. This is due to the dependency structure of Equation[1](https://arxiv.org/html/2609.36202#S1.E1)\. At each diffusion timestep, when one token is unmasked, both the dLLM logits and the guidance vector change because the conditioning variable changes\. Recomputing the dLLM logits requires a forward pass through the dLLM, which involves bidirectional attention over the sequence, and recomputing the guidance vector requires a forward and backward pass through a large reward model \(e\.g\., a Skywork\-Llama\-8B model\([Liu et al\., 2026](https://arxiv.org/html/2609.36202#bib.bib23)\)\)\. Both of these steps can be large inference\-time bottlenecks\.

Parallel decoding is the dominant strategy to accelerate dLLM inference, so a natural idea is to follow existing work and unmask multiple tokens in each diffusion timestep\([Ben\-Hamu et al\., 2025](https://arxiv.org/html/2609.36202#bib.bib8);[Kim et al\., 2026a](https://arxiv.org/html/2609.36202#bib.bib13);[Wu et al\., 2025](https://arxiv.org/html/2609.36202#bib.bib3)\)\. However, existing work is developed for unguided generation, so it does not establish when guidance can be reused, which is the main inference\-time bottleneck \(Figure[1](https://arxiv.org/html/2609.36202#S1.F1)b\)\. Further, it is unclear which tokens remain safe to unmask together once guidance has altered the token distributions\. This leads to our main research question:

![Refer to caption](https://arxiv.org/html/2609.36202v1/teaser.png)Figure 1:Left: Reward guidance\([Tejaswi et al\., 2026](https://arxiv.org/html/2609.36202#bib.bib15)\)can increase incompatibilities across token values proposed during parallel decoding\. Above,CPAshould expand toCertified Public Accountant\. While unguided generation with parallel decoding method Fast\-dLLM\([Wu et al\., 2025](https://arxiv.org/html/2609.36202#bib.bib3)\)produces the correct token prefix forCPA, Fast\-dLLM with guidance fails\. Right: Our method is much faster than sequential guided generation while retaining similar generation quality\.Can we accelerate reward\-guided dLLM inference while preserving its generation quality?

In this paper, we present, to the best of our knowledge, the first study of accelerating gradient\-based reward guidance for dLLMs\. Our main contributions are:

1. 1\.We show that, surprisingly, reusing the guidance vector to unmask a group ofkktokens is benign\. However, reusing the dLLM logits degrades generation quality\. Specifically, we observe that \(O1\) at masked positions with high confidence scores, the token distribution after guidance remains similar when guidance from an earlier diffusion timestep is reused\. Conversely, \(O2\) guidance can increase incompatibilities between co\-proposed token values compared to unguided decoding \(Figure[1](https://arxiv.org/html/2609.36202#S1.F1)a\), preventing reuse of the dLLM logits\.
2. 2\.We introduce FastGuide, a training\-free, hybrid, and adaptive acceleration strategy for reward\-guided dLLM decoding\. Motivated byO1, similar to parallel decoding, we select at each iteration a group of candidate token positions with high confidence scores and compute guidance once for the group\. To addressO2, we then autoregressively unmask candidates within this group, recomputing the dLLM logits after each unmasking so that subsequent token choices are conditioned on prior unmaskings\. To make these updates efficient, we use sparse recomputation with KV caching\([Ma et al\., 2025](https://arxiv.org/html/2609.36202#bib.bib4)\), updating only a small subset of dLLM token representations\. Finally, token candidates whose recomputed distributions have low confidence scores are deferred to a later iteration, adaptively balancing the parallel and autoregressive decoding steps\.
3. 3\.Our guidance reuse principle allows us to extend five existing parallel decoders to the guided setting\. Across three reward benchmarks and two dLLMs, FastGuide generates responses up to4\.4×4\.4\\timesfaster than sequential guided decoding while retaining similar reward model scores\.

## 2Background and Related Work

In this section, we review masked dLLM inference, parallel decoding strategies, and reward guidance methods\. Given a prompt𝒑\{\\bm\{p\}\}, the masked dLLM generates a fixed\-length response𝒙0\{\\bm\{x\}\}\_\{0\}of lengthLLwhere each token comes from a vocabulary𝕍\{\\mathbb\{V\}\}, where𝕍\{\\mathbb\{V\}\}includes a<MASK\>token\.

##### Masked dLLM inference\.

A masked dLLM generates𝒙0\{\\bm\{x\}\}\_\{0\}by iteratively denoising a partially masked state𝒙t∈𝕍L\{\\bm\{x\}\}\_\{t\}\\in\{\\mathbb\{V\}\}^\{L\}\([Austin et al\., 2021](https://arxiv.org/html/2609.36202#bib.bib12);[Sahoo et al\., 2024](https://arxiv.org/html/2609.36202#bib.bib11)\)\. We will refer tottas the diffusion timestep, ranging fromTTto00, where𝒙T\{\\bm\{x\}\}\_\{T\}is a sequence with all masked tokens\. At every diffusion timesteptt, we denote the set of masked and unmasked positions by

𝕄t=\{i∈\[L\]:xti=<MASK\>\},𝕌t=\[L\]∖𝕄t,\{\\mathbb\{M\}\}\_\{t\}=\\\{i\\in\[L\]:x\_\{t\}^\{i\}=\\texttt\{<MASK\>\}\\\},\\qquad\{\\mathbb\{U\}\}\_\{t\}=\[L\]\\setminus\{\\mathbb\{M\}\}\_\{t\},\(2\)where\[L\]=\{1,…,L\}\[L\]=\\\{1,\\ldots,L\\\}\. We will use𝒙tu\{\\bm\{x\}\}\_\{t\}^\{\\text\{u\}\}as a shorthand to represent the set of unmasked tokens at timett\. At each denoising step, the dLLM processes the full state𝒙t\{\\bm\{x\}\}\_\{t\}and predicts a categorical distribution at every masked position:

pθ​\(xti∣𝒙tu\)=softmax⁡\(ℓtunguided​\(xti∣𝒙tu\)\),i∈𝕄t\.p\_\{\\theta\}\(x\_\{t\}^\{i\}\\mid\{\\bm\{x\}\}\_\{t\}^\{\\text\{u\}\}\)=\\operatorname\{softmax\}\\\!\\left\(\\ell\_\{t\}^\{\\text\{unguided\}\}\(x\_\{t\}^\{i\}\\mid\{\\bm\{x\}\}\_\{t\}^\{\\text\{u\}\}\)\\right\),\\qquad i\\in\{\\mathbb\{M\}\}\_\{t\}\.\(3\)Throughout the paper, we will useℓt\\ell\_\{t\}to refer to logits andpθp\_\{\\theta\}as the corresponding normalized probability distribution\. A decoding rule then chooses which token positions to unmask and which token values to place at those positions\. A common confidence\-based decoding rule is to pick the tokeniiwhose predicted distributionpθ​\(xti∣𝒙tu\)p\_\{\\theta\}\(x\_\{t\}^\{i\}\\mid\{\\bm\{x\}\}\_\{t\}^\{\\text\{u\}\}\)has minimum entropy \(which we denote ashigh\-confidence tokens\), and then sample from that distribution\([Nie et al\., 2026](https://arxiv.org/html/2609.36202#bib.bib24)\)\.

##### Parallel decoding\.

Instead of sampling one token at each diffusion timestep \(sequential decoding\), parallel decoding methods unmaskkktokens at each step, wherekkcan be fixed a priori\([Nie et al\., 2026](https://arxiv.org/html/2609.36202#bib.bib24)\)or adaptive at inference time\([Kim et al\., 2025b](https://arxiv.org/html/2609.36202#bib.bib7);[Wu et al\., 2025](https://arxiv.org/html/2609.36202#bib.bib3)\)\. We refer to thesekktoken positions as𝔸t=\{a1,…,ak\}⊆𝕄t\{\\mathbb\{A\}\}\_\{t\}=\\\{a\_\{1\},\\ldots,a\_\{k\}\\\}\\subseteq\{\\mathbb\{M\}\}\_\{t\}and𝒄=\(c1,…,ck\)∈𝕍k\{\\bm\{c\}\}=\(c\_\{1\},\\ldots,c\_\{k\}\)\\in\{\\mathbb\{V\}\}^\{k\}as their proposed values\. The true joint distribution over thesekktokens would be

pθ​\(xta1=c1,…,xtak=ck∣𝒙tu\)\.p\_\{\\theta\}\(x\_\{t\}^\{a\_\{1\}\}=c\_\{1\},\\dots,x\_\{t\}^\{a\_\{k\}\}=c\_\{k\}\\mid\{\\bm\{x\}\}\_\{t\}^\{\\text\{u\}\}\)\.\(4\)However, the model only outputs marginal distributionspθ​\(xtaj=cj∣𝒙tu\)p\_\{\\theta\}\(x\_\{t\}^\{a\_\{j\}\}=c\_\{j\}\\mid\{\\bm\{x\}\}\_\{t\}^\{\\text\{u\}\}\)at every token position, so a common approximation is to model this joint distribution as a product of marginal distributions∏j=1kpθ​\(xtaj=cj∣𝒙tu\)\\prod\_\{j=1\}^\{k\}p\_\{\\theta\}\(x\_\{t\}^\{a\_\{j\}\}=c\_\{j\}\\mid\{\\bm\{x\}\}\_\{t\}^\{\\text\{u\}\}\)\. The key challenge is to mitigate this approximation error, known as the joint\-marginal mismatch\([Ben\-Hamu et al\., 2025](https://arxiv.org/html/2609.36202#bib.bib8);[Liu et al\., 2025a](https://arxiv.org/html/2609.36202#bib.bib14)\)\. For instance, some works address this by selecting positions with limited dependencies using signals such as confidence scores\([Wu et al\., 2025](https://arxiv.org/html/2609.36202#bib.bib3)\), analysis of temporal stability of logits\([Kim et al\., 2025b](https://arxiv.org/html/2609.36202#bib.bib7);[Ben\-Hamu et al\., 2025](https://arxiv.org/html/2609.36202#bib.bib8)\), or attention\-based dependency graphs\([Kim et al\., 2026a](https://arxiv.org/html/2609.36202#bib.bib13)\)\. However, they are developed for unguided generation, and thus do not address recomputation of the guidance vector as in Equation[1](https://arxiv.org/html/2609.36202#S1.E1)or whether the selection rules remain reliable for guided token distributions\. An alternative set of approaches uses external planners, such as autoregressive LLMs\([Israel et al\., 2026](https://arxiv.org/html/2609.36202#bib.bib25);[Liu et al\., 2025a](https://arxiv.org/html/2609.36202#bib.bib14)\)to autoregressively commit thekktokens\. In contrast to these methods, FastGuide is a hybrid decoding strategy that uses parallel decoding techniques for the reward guidance and autoregressive decoding over𝔸t\{\\mathbb\{A\}\}\_\{t\}\. Further, FastGuide obviates the need for external models by efficiently using the dLLM itself to autoregressively unmask tokens\. We give a detailed comparison with other decoders in Appendix[A](https://arxiv.org/html/2609.36202#A1)\.

##### Reward guidance for alignment\.

In this work, we focus on the gradient\-based reward guidance framework \(leaving discussion of other methods to Appendix[A](https://arxiv.org/html/2609.36202#A1)\), which has been shown to outperform methods such as Best\-of\-N sampling\([Liu et al\., 2025b](https://arxiv.org/html/2609.36202#bib.bib18);[Tejaswi et al\., 2026](https://arxiv.org/html/2609.36202#bib.bib15)\)\. This paradigm draws inspiration from classifier\-guidance approaches for continuous diffusion models, which has been remarkably successful in numerous domains for training\-free steering\([Daras et al\., 2024](https://arxiv.org/html/2609.36202#bib.bib35);[Dhariwal and Nichol, 2021](https://arxiv.org/html/2609.36202#bib.bib36);[Thaker et al\., 2025](https://arxiv.org/html/2609.36202#bib.bib37)\)\. For dLLMs, gradient\-based guidance methods steer the dLLM logits with a differentiable sequence\-level reward, such as a downstream reward model\. As shown in Equation[1](https://arxiv.org/html/2609.36202#S1.E1), the guided logits at each position are the sum of the dLLM logits and a per\-token guidance vector𝒓t\{\\bm\{r\}\}\_\{t\}that shifts the distribution toward higher\-reward tokens\. The guidance vector𝒓t\{\\bm\{r\}\}\_\{t\}is estimated by completing the masked sequence, evaluating the completion with a reward model, and differentiating the reward with respect to the dLLM logits\. Existing gradient\-based guidance methods have focused on improved input completion techniques, e\.g\., by sampling token values from the predicted marginal distribution at each token\([Rout et al\., 2025](https://arxiv.org/html/2609.36202#bib.bib17)\), computing an expected embedding for each position with respect to the predicted distribution at each token\([Murata et al\., 2024](https://arxiv.org/html/2609.36202#bib.bib16)\), or interpolating between the two\([Tejaswi et al\., 2026](https://arxiv.org/html/2609.36202#bib.bib15)\)\. As such, they have been largely limited to sequential decoding evaluations, resulting in long inference times \(Figure[1](https://arxiv.org/html/2609.36202#S1.F1)b\)\. In this work, we focus on the orthogonal problem of accelerating reward guidance since computing forward and backward passes through large reward models at each diffusion timestep can be very costly\.

## 3Method: FastGuide

![Refer to caption](https://arxiv.org/html/2609.36202v1/MainFigure.png)Figure 2:Overview of FastGuide\. Left: Illustration of one hybrid decoding step of FastGuide withk=4k=4\. FastGuide utilizes guidance caching and autoregressive sparse dLLM recomputation\. After recomputation, tokens are deferred if token confidence drops below a threshold\. Right: Sparse dLLM recomputation of token 7 after committing token 6, with window radiusw=1w=1\. Attention and FFN outputs are recomputed only at the selected positions, whose queries attend to the full sequence using updated K/V inside the windows and cached K/V elsewhere\.In this section, we develop FastGuide, an algorithm to accelerate reward\-guided decoding for dLLMs\. First, in Section[3\.1](https://arxiv.org/html/2609.36202#S3.SS1), we derive our hybrid decoding strategy using guidance caching and autoregressive dLLM recomputation\. In Section[3\.1\.2](https://arxiv.org/html/2609.36202#S3.SS1.SSS2), we use KV caching and sparse recomputation of attention\([Ma et al\., 2025](https://arxiv.org/html/2609.36202#bib.bib4)\)to efficiently perform autoregressive dLLM recomputation\. Lastly, we introduce confidence deferral in Section[3\.2](https://arxiv.org/html/2609.36202#S3.SS2), which is a technique to make decoding adaptive rather than unmasking a fixed number of tokens in each step\. Our final method is summarized in Figure[2](https://arxiv.org/html/2609.36202#S3.F2)and pseudocode is given in Appendix[C](https://arxiv.org/html/2609.36202#A3)\. Throughout this section, we instantiate unguided parallel decoding as the widely used rule of confidence\-based unmasking of the topkktokens\([Nie et al\., 2026](https://arxiv.org/html/2609.36202#bib.bib24)\), and we instantiate reward guidance as the recently introduced EntRGi algorithm\([Tejaswi et al\., 2026](https://arxiv.org/html/2609.36202#bib.bib15)\)\. In Section[4](https://arxiv.org/html/2609.36202#S4), we compare FastGuide with other parallel decoders and evaluate other guidance algorithms\. Evaluations in this section were performed on RM\-Bench\([Liu et al\., 2025b](https://arxiv.org/html/2609.36202#bib.bib18)\)with the Dream text diffusion model\([Ye et al\., 2025b](https://arxiv.org/html/2609.36202#bib.bib9)\)and Skywork\-Reward\-V2\-Qwen3\-0\.6B reward model\([Liu et al\., 2026](https://arxiv.org/html/2609.36202#bib.bib23)\), with further experimental details given in Appendix[B\.3](https://arxiv.org/html/2609.36202#A2.SS3)\.

### 3\.1Hybrid Decoding: Guidance Caching and dLLM Recomputation

Recall from Equation[1](https://arxiv.org/html/2609.36202#S1.E1)that both the dLLM logits and the guidance vector depend on the unmasked tokens𝒙tu\{\\bm\{x\}\}\_\{t\}^\{\\mathrm\{u\}\}at each timestep\. Existing sequential reward\-guided decoders recompute both terms after every unmasking, which requires one dLLM forward pass and one guidance computation per token\. Since sequential reward\-guided decoding has shown promising performance\([Tejaswi et al\., 2026](https://arxiv.org/html/2609.36202#bib.bib15)\), we take it as our starting point and seek to maintain its generation quality while reducing its inference time by unmaskingkktokens per diffusion timestep\. Since Equation[1](https://arxiv.org/html/2609.36202#S1.E1)consists of two terms, this raises the question of whether the dLLM logits, the guidance vector, or both can be safely reused to unmask a candidate token set𝔸t=\{a1,…,ak\}\{\\mathbb\{A\}\}\_\{t\}=\\\{a\_\{1\},\\dots,a\_\{k\}\\\}\(defined in Section[2](https://arxiv.org/html/2609.36202#S2)\)\.

##### Separating dLLM and guidance reuse\.

To study this question, we independently vary whether each term is reused or recomputed within the candidate token set\. Specifically, consider an ordering of𝔸t\{\\mathbb\{A\}\}\_\{t\}\(e\.g\., descending confidence scores\) and let𝒙t\(s\)∈𝕍L\{\\bm\{x\}\}\_\{t\}^\{\(s\)\}\\in\{\\mathbb\{V\}\}^\{L\}denote the sequence after unmasking the firstsscandidates of𝔸t\{\\mathbb\{A\}\}\_\{t\}, with𝒙t\(0\)=𝒙t\{\\bm\{x\}\}\_\{t\}^\{\(0\)\}=\{\\bm\{x\}\}\_\{t\}\. Then, for token positionaj∈𝔸ta\_\{j\}\\in\{\\mathbb\{A\}\}\_\{t\}, sequential guided decoding corresponds to iteratively computing for allj∈\{1,…,k\}j\\in\\\{1,\\dots,k\\\}:

softmax⁡\(ℓtunguided​\(xtaj∣\(𝒙t\(s\)\)u\)\+𝒓t​\(xtaj∣\(𝒙t\(s′\)\)u\)\)\\operatorname\{softmax\}\\left\(\\ell\_\{t\}^\{\\text\{unguided\}\}\\left\(x\_\{t\}^\{a\_\{j\}\}\\mid\(\{\\bm\{x\}\}\_\{t\}^\{\(s\)\}\)^\{\\mathrm\{u\}\}\\right\)\+\{\\bm\{r\}\}\_\{t\}\\left\(x\_\{t\}^\{a\_\{j\}\}\\mid\(\{\\bm\{x\}\}\_\{t\}^\{\(s^\{\\prime\}\)\}\)^\{\\mathrm\{u\}\}\\right\)\\right\)\(5\)fors=s′=j−1s=s^\{\\prime\}=j\-1\. Instead, we analyze variants of Equation[5](https://arxiv.org/html/2609.36202#S3.E5)wheres,s′∈\{0,j−1\}s,s^\{\\prime\}\\in\\\{0,\\,j\-1\\\}\. Here,s=0s=0\(resp\.s′=0s^\{\\prime\}=0\) denotes reuse of the dLLM logits \(resp\. guidance vector\) computed at the start of the candidate group, whereass=j−1s=j\-1\(resp\.s′=j−1s^\{\\prime\}=j\-1\) denotes autoregressive recomputation after each token unmasking\. We refer to computing guidance once and reusing it across the candidate group \(s′=0s^\{\\prime\}=0\) as*guidance caching*, which is particularly attractive because guidance computation dominates inference time \(Figure[1](https://arxiv.org/html/2609.36202#S1.F1)b\)\. When bothssands′s^\{\\prime\}are00, both terms are sampled from the guided distributions computed at the initial state, following the product\-of\-marginals approximation of Section[2](https://arxiv.org/html/2609.36202#S2)\. Figure[3](https://arxiv.org/html/2609.36202#S3.F3)demonstrates that guidance caching along with dLLM recomputation retains most of the quality of recomputing both terms \(orange vs\. purple\)\. Conversely, reusing the dLLM logits \(blue and green curves\) leads to quality degradation irrespective of guidance caching\.

##### Hybrid decoding\.

This analysis motivates combining guidance caching with autoregressive dLLM recomputation, yielding a hybrid decoding strategy\. As in parallel decoding, we amortize the guidance computation overkktoken candidates\. Within this candidate set, we unmask candidates autoregressively, recomputing the dLLM logits after each unmasking\. Thus, the token at positionaja\_\{j\}is chosen from

softmax⁡\(ℓtunguided​\(xtaj∣\(𝒙t\(j−1\)\)u\)\+𝒓t​\(xtaj∣\(𝒙t\(0\)\)u\)\)\.\\operatorname\{softmax\}\\left\(\\ell\_\{t\}^\{\\text\{unguided\}\}\\left\(x\_\{t\}^\{a\_\{j\}\}\\mid\(\{\\bm\{x\}\}\_\{t\}^\{\(j\-1\)\}\)^\{\\mathrm\{u\}\}\\right\)\+\{\\bm\{r\}\}\_\{t\}\\left\(x\_\{t\}^\{a\_\{j\}\}\\mid\(\{\\bm\{x\}\}\_\{t\}^\{\(0\)\}\)^\{\\mathrm\{u\}\}\\right\)\\right\)\.\(6\)

#### 3\.1\.1Counterfactual Analysis of Hybrid Decoding

Figure 3:Disentangling Reuse of dLLM Logits and Guidance\.To better understand the asymmetry between guidance reuse and dLLM logits reuse, we run two counterfactual studies along sequential decoding trajectories \(Figure[4](https://arxiv.org/html/2609.36202#S3.F4)\)\.

Observation 1 \(O1\): guidance reuse has little effect at high\-confidence positions\.At every diffusion timestep, we hold the current dLLM logits fixed and replace the guidance vector with one cached from an earlier step\. We measure the total\-variation distance between the resulting guided distributions \(Figure[4](https://arxiv.org/html/2609.36202#S3.F4)\)\. We find that this distance is smaller at high\-confidence positions, which are precisely the positions that standard decoding strategies unmask\([Ye et al\., 2025b](https://arxiv.org/html/2609.36202#bib.bib9)\)\. This helps explain why guidance caching is effective\.

Observation 2 \(O2\): guidance can make co\-proposed tokens locally incompatible\.Tokens proposed in parallel from the same partially masked sequence need not be likely under the true joint distribution of the tokens since the model only outputs marginal distributions for each token\. While this joint\-marginal mismatch is present even without guidance\([Ben\-Hamu et al\., 2025](https://arxiv.org/html/2609.36202#bib.bib8)\), we find that guidance can exacerbate this mismatch\. To measure the effect of guidance on the joint\-marginal mismatch, we consider sequential unguided trajectories and toggle guidance at each state, examining its effect on the top\-2 confident tokens that would be committed in parallel\. We quantify the joint\-marginal mismatch at these two tokens as the dLLM\-predicted Pointwise Mutual Information \(PMI\), which is the difference in log probabilities of the jointp⁡\(a,b\)=p⁡\(a\)​p​\(b∣a\)p\(a,b\)=p\(a\)p\(b\\mid a\)and product of marginal distributionsp⁡\(a\)​p​\(b\)p\(a\)p\(b\)\. Negative PMI indicates that committing the first token lowers the dLLM\-predicted probability of the second relative to its marginal probability\. Figure[4](https://arxiv.org/html/2609.36202#S3.F4)shows that guidance actually increases incompatibility among proposed tokens\. The average number of incompatible pairs rises from1\.041\.04without guidance to2\.472\.47with guidance, and the fraction of prompts with at least one incompatible pair rises from58%58\\%to90%90\\%\. This illustrates the role of dLLM recomputation in our hybrid decoding strategy as incompatible proposals can be revised before unmasking\.

Figure 4:Left: \(a\) Replacing fresh guidance with a cached guidance produces small changes in high\-confidence token distributions\. Right: \(b\) Percentage of prompts exhibiting certain number of conflicts between proposed tokens \(as measured by PMI being less than−1\-1\)\.
#### 3\.1\.2Sparse dLLM Recomputation

Despite the speedups obtained from reducing guidance computations, exact dLLM recomputation still requires the same number of dLLM forward passes as sequential decoding\. To reduce this cost, we utilize dLLM acceleration techniques\([Ma et al\., 2025](https://arxiv.org/html/2609.36202#bib.bib4);[Liu et al\., 2025c](https://arxiv.org/html/2609.36202#bib.bib5)\)\. Specifically, for every hybrid decoding step, once thekkcandidate positions are fixed, recomputing logits for the next candidate requires updated logits only at its positionaja\_\{j\}\. Thus, we make a connection to[Ma et al\. \(2025\)](https://arxiv.org/html/2609.36202#bib.bib4), where the authors fix the token generation order of unguided sequential dLLM decoding to develop efficient sparse dLLM forward passes\. Specifically, we can approximate the full forward pass by recomputing attention and feed\-forward outputs for every transformer layer at a small set of positions: the union of windows of radiuswwaround the most recently committed token and the next candidate position\. To recompute attention, the queries at only these positions attend to the full sequence, using newly computed keys and values inside the windows and cached keys and values elsewhere \(see Figure[2](https://arxiv.org/html/2609.36202#S3.F2)\)\. This reduces the time complexity of each recomputation fromO⁡\(L2\)O\(L^\{2\}\)toO⁡\(w​L\)O\(wL\)\. The updated keys and values overwrite the cache for subsequent recomputation passes\. The sparse dLLM forward produces approximate base logitsℓ^tunguided​\(xtaj∣\(𝒙t\(j−1\)\)u\)\\widehat\{\\ell\}\_\{t\}^\{\\text\{unguided\}\}\\left\(x\_\{t\}^\{a\_\{j\}\}\\mid\(\{\\bm\{x\}\}\_\{t\}^\{\(j\-1\)\}\)^\{\\mathrm\{u\}\}\\right\), which we combine with the cached reward gradient as in Equation[6](https://arxiv.org/html/2609.36202#S3.E6)\.

### 3\.2Adaptivity via Confidence Deferral

Unmasking exactlykktokens in every step can be restrictive\([Wu et al\., 2025](https://arxiv.org/html/2609.36202#bib.bib3)\)since the number of tokens that are safe to unmask together depends on the prompt and decoding state\. We therefore introduce an adaptive confidence deferral strategy\. Specifically, after sparse dLLM recomputation, we compute

γj=maxv∈𝕍⁡softmax⁡\(ℓ^tunguided​\(xtaj∣\(𝒙t\(j−1\)\)u\)\+𝒓t​\(xtaj∣\(𝒙t\(0\)\)u\)\)​\(v\)\.\\gamma\_\{j\}=\\max\_\{v\\in\{\\mathbb\{V\}\}\}\\operatorname\{softmax\}\\left\(\\widehat\{\\ell\}\_\{t\}^\{\\text\{unguided\}\}\\left\(x\_\{t\}^\{a\_\{j\}\}\\mid\(\{\\bm\{x\}\}\_\{t\}^\{\(j\-1\)\}\)^\{\\mathrm\{u\}\}\\right\)\+\{\\bm\{r\}\}\_\{t\}\\left\(x\_\{t\}^\{a\_\{j\}\}\\mid\(\{\\bm\{x\}\}\_\{t\}^\{\(0\)\}\)^\{\\mathrm\{u\}\}\\right\)\\right\)\(v\)\.\(7\)Ifγj≥τ\\gamma\_\{j\}\\geq\\taufor a chosen thresholdτ\\tau, we commit the verified token\. If not, the position remains masked and is reconsidered in a later step, when both the dLLM logits and the guidance are recomputed\. The first token candidate is always committed, so every step makes progress\. This makes the degree of parallelism adaptive and postpones low\-confidence candidates until a later step so both dLLM logits and guidance can be recomputed\.

In summary, for each decoding step, FastGuide performs11dLLM forward pass,11guidance computation step, andk−1k\-1sparse recomputation passes, as opposed to thekkdLLM forward passes andkkguidance computations for sequential guided decoding\. Refer to Appendix[C](https://arxiv.org/html/2609.36202#A3)for psuedocode\.

## 4Experiments

Below, we describe our experimental setup, deferring other experimental details to Appendix[B](https://arxiv.org/html/2609.36202#A2)\.

Table 1:Quantitative results for reward\-guided generation on JudgeBench, RM\-Bench, and Reward\-Bench\-2\. All parallel decoding methods use cached EntRGi reward guidance, unless \(BoN\) is mentioned, denoting Best\-of\-N sampling\. Subscripts denote the standard error over prompts\.##### Benchmarks\.

Following prior work on reward guidance, we evaluate FastGuide on three reward benchmarks \(JudgeBench\([Tan et al\., 2025](https://arxiv.org/html/2609.36202#bib.bib19)\), RM\-Bench\([Liu et al\., 2025b](https://arxiv.org/html/2609.36202#bib.bib18)\), and Reward\-Bench\-2\([Malik et al\., 2025](https://arxiv.org/html/2609.36202#bib.bib20)\)\), spanning various categories such as Coding, Math, Precise Instruction Following \(IF\), and Safety\.

##### Baselines\.

We evaluate FastGuide on two popular open\-source dLLMs: Dream\-7B\-Instruct\([Ye et al\., 2025b](https://arxiv.org/html/2609.36202#bib.bib9)\)and LLaDA\-8B\-Instruct\([Nie et al\., 2026](https://arxiv.org/html/2609.36202#bib.bib24)\)\. While existing parallel decoding baselines are evaluated on unguided decoding, our guidance caching strategy \(Section[3](https://arxiv.org/html/2609.36202#S3)\) allows us to extend existing methods to the guided setting\. To our knowledge, our work represents the first evaluation of these decoders in the reward\-guided setting\. We compare against several training\-free parallel decoding baselines such as confidence\-based decoding \(which we abbreviate as Conf\.\)\([Nie et al\., 2026](https://arxiv.org/html/2609.36202#bib.bib24)\), Fast\-dLLM\([Wu et al\., 2025](https://arxiv.org/html/2609.36202#bib.bib3)\), KLASS\([Kim et al\., 2025b](https://arxiv.org/html/2609.36202#bib.bib7)\), EB\-sampler\([Ben\-Hamu et al\., 2025](https://arxiv.org/html/2609.36202#bib.bib8)\)and DAPD\([Kim et al\., 2026a](https://arxiv.org/html/2609.36202#bib.bib13)\)\. For fair comparison, we do not utilize additional optimizations of the full dLLM forward pass such as prompt caching\([Wu et al\., 2025](https://arxiv.org/html/2609.36202#bib.bib3)\), which are orthogonal improvements and can be combined with FastGuide\. We also compare to Best\-Of\-N sampling withN=4N=4\. We use the recent state\-of\-the\-art reward guidance method EntRGI\([Tejaswi et al\., 2026](https://arxiv.org/html/2609.36202#bib.bib15)\)with the Skywork\-Reward\-V2\-Qwen3\-1\.7B reward model\([Liu et al\., 2026](https://arxiv.org/html/2609.36202#bib.bib23)\)\. We defer evaluation of other guidance methods and different reward models to Appendix[D\.1](https://arxiv.org/html/2609.36202#A4.SS1)and[D\.2](https://arxiv.org/html/2609.36202#A4.SS2)\.

##### Metrics\.

Following[Tejaswi et al\. \(2026\)](https://arxiv.org/html/2609.36202#bib.bib15), we generate44trajectories for each prompt and measure generation quality by the reward model score\. We report Top@1 \(the maximum reward model score across all trajectories\) and Seq Gap \(the gap between parallel decoding Top@1 and the corresponding sequential decoding baseline using the same guidance algorithm\)\. To monitor reward overoptimization\([Gao et al\., 2023](https://arxiv.org/html/2609.36202#bib.bib41)\), we report LMUnit \(external judge LLM LMUnit\-Qwen2\.5\-72B\([Saad\-Falcon et al\., 2025](https://arxiv.org/html/2609.36202#bib.bib38)\)score from1−51\-5\)\. We further report Avg@4 \(the average reward model score across the44trajectories\) in Appendix[D\.3](https://arxiv.org/html/2609.36202#A4.SS3)\. To evaluate inference time, we report s/gen \(wall\-clock inference time for each generation\), computed on a Nvidia A5000 GPU\.

##### Hyperparameters\.

We set the generation length to be128128for all methods, and baseline hyperparameters can be found in Appendix[B\.2](https://arxiv.org/html/2609.36202#A2.SS2)\. FastGuide uses a candidate pool ofk=8k=8for all experiments along with confidence deferral parameterτ=0\.5\\tau=0\.5and attention window sizew=2w=2\. We consider a higher throughput regime by increasingkkto1616in Appendix[D](https://arxiv.org/html/2609.36202#A4)\. We use batch size 4 for Dream\-7B and batch size 2 for LLaDA\-8B\.

### 4\.1Main Results

##### Reward vs Inference Time\.

Table[1](https://arxiv.org/html/2609.36202#S4.T1)and Figure[6](https://arxiv.org/html/2609.36202#S4.F6)show our main results\. FastGuide is2\.82\.8–4\.4×4\.4\\timesfaster than sequential guided decoding across the three benchmarks and two models while remaining within 0\.75 reward points of sequential guidance on both models\. Its high LMUnit scores provide complementary evidence of response quality and also guard against reward overoptimization\. Compared to unguided sequential Best\-of\-N sampling, sequential guidance improves Top@1 by0\.60\.6–1\.01\.0reward points but costs3\.53\.5–4\.4×4\.4\\timesthe inference time\. FastGuide reduces this cost to0\.90\.9–1\.3×1\.3\\timesthat of Best\-of\-N while retaining a Top@1 improvement of0\.10\.1–0\.650\.65reward points\. These conclusions also hold on LLaDA\-8B, where FastGuide again attains the smallest Seq Gap, while Fast\-dLLM and KLASS degrade even more severely than on Dream\-7B\. Cached guidance on top of confidence\-based decoding \(Conf row\) improves Top@1 by roughly44reward points over its unguided counterpart \(Conf \(BoN\) row\), confirming that guidance caching preserves the benefit of guidance\. FastGuide further narrows this gap through sparse dLLM recomputation and confidence deferral\. An important benefit of FastGuide is that sparse dLLM recomputation transforms the compute\-bound full dLLM forward pass, that involves bidirectional attention over the whole sequence, into a memory\-bound sparse forward over a few positions\. Thus, standard techniques such as batching can reduce inference time per trajectory, which we explore in Appendix[D\.4](https://arxiv.org/html/2609.36202#A4.SS4)\.

Figure 5:Quality versus inference time on JudgeBench using Dream\-7B model\.Figure 6:τ\\tausweep for confidence deferral on RM\-Bench using Dream\-7B model\.
Figure 7:Sequential gap by category \(closer to00is better\)\.FastGuide has the smallest gap to sequential decoding, especially in reasoning categories such as coding where baselines degrade\.
##### Comparison to Parallel Decoders\.

Most existing parallel decoders lose a large part of the benefit of guidance, with Fast\-dLLM, KLASS, and EB\-sampler suffering substantial performance degradation relative to sequential decoding\. In contrast, DAPD, which selects positions from the dLLM’s attention graph rather than from marginal confidence, is by far the strongest baseline\. FastGuide matches or improves on DAPD on every benchmark and model, with the largest gains on RM\-Bench \(\+1\.1\+1\.1Top@1 on Dream\-7B\)\. Figure[7](https://arxiv.org/html/2609.36202#S4.F7)analyzes this Seq Gap further by splitting performance along benchmark categories\. DAPD exhibits larger reward gaps on coding and math tasks\. This pattern is consistent with the greater sensitivity to incompatible token choices in these categories, and the qualitative example in Figure[8](https://arxiv.org/html/2609.36202#S4.F8)further illustrates this\. In contrast, for categories such as chat and safety, an incompatible pair of tokens costs little fluency\. FastGuide has a small Seq Gap across nearly all categories, highlighting the benefit of sparse dLLM recomputation in the reward\-guided setting\.

##### Qualitative Samples\.

Figure[8](https://arxiv.org/html/2609.36202#S4.F8)shows an illustrative mathematical reasoning prompt from RM\-Bench\. Unguided DAPD has a correct reasoning trace, but under guidance it commits an incorrect line of arithmetic reasoning \(525^\{2\}in place of828^\{2\}\) that the rest of the derivation then propagates to a wrong final answer, showing an instance ofO2\. Instead, FastGuide recomputes each candidate token distribution before unmasking it and recovers the correct answer\. The reward model scores in Figure[8](https://arxiv.org/html/2609.36202#S4.F8)reflect this ordering\. Appendix[D\.9](https://arxiv.org/html/2609.36202#A4.SS9)shows further qualitative examples\.

Prompt:If the endpoints of a line segment are \(2, \-2\) and \(10, 4\), what is the length of the segment?

DAPD, guidedReward−1\.22\-1\.22To find the length of the line segment, we can use the distance formula, which is $\\sqrt\{\(x\_2 \- x\_1\)ˆ2 \+ \(y\_2 \- y\_1\)ˆ2\}$\. Plugging in the given endpoints, we have $\\sqrt\{\(10 \- 2\)ˆ2\)\+ \(4 \- \(\-2\)\)ˆ2\} = \\sqrt\{5ˆ2 \+ 6ˆ2\} = \\sqrt\{10\+ 36\} = \\sqrt\{46\}$\. The length of the line segment is $\\boxed\{6\}$ units\.The answer is: $\\boxed\{6\}$

DAPD, unguidedReward\+6\.59\+6\.59We can use the distance formula to find the length of the line segment\. The distance formula between two points $\(x\_1, y\_1\)$ and $\(x\_2, y\_2\)$ is given by $\\sqrt\{\(x\_2 \- x\_1\)ˆ2 \+ \(y\_2 \- y\_1\)ˆ2\}$\. Plugging in the given values, we have $\\sqrt\{\(10 \- 2\)ˆ2 \+ \(4 \- \(\-2\)\)ˆ2\} = \\sqrt\{8ˆ2 \+ 6ˆ2\} = \\sqrt\{64 \+ 36\} = \\sqrt\{100\}

FastGuideReward\+10\.44\+10\.44To find the length of the line segment, we can use the distance formula, which is derived from the Pythagorean Theorem\. The formula is: Distance = sqrt\(\(x2 \- x1\)ˆ2 \+ \(y2 \- y1\)ˆ2\) Plugging in the given coordinates: Distance = sqrt\(\(10 \- 2\)ˆ2 \+ \(4 \- \(\-2\)\)ˆ2\) = sqrt\(8ˆ2 \+ 6ˆ2\) = sqrt\(64 \+ 36\) = sqrt\(100\) = 10 So, the length of the line segment is 10 units\.

Figure 8:Qualitative generations on Math RM\-Bench prompt\.Red highlighted tokens denote incorrect tokens\. This reinforces that guidance can worsen performance of parallel decoders, while FastGuide succeeds\.

### 4\.2Ablation Studies

##### Effect of confidence deferral\.

Figure[6](https://arxiv.org/html/2609.36202#S4.F6)sweeps the deferral thresholdτ\\tauof FastGuide on RM\-Bench, showing thatτ\\tauprovides an extra knob for interpolating the quality\-efficiency tradeoff\. Specifically, without deferral \(τ=0\\tau\{=\}0\), unmasking a fixed88tokens per step is fastest, but generation quality degrades\. As more tokens are deferred, quality increases, approaching the performance of sequential decoding, but at the expense of inference time due to increased guidance computation steps and full dLLM forwards\.

##### Effect of dLLM Recomputation\.

Figure[9](https://arxiv.org/html/2609.36202#S4.F9)analyzes the sparse dLLM recomputation step of FastGuide\. Specifically, Figure[9](https://arxiv.org/html/2609.36202#S4.F9)b shows that sparse dLLM recomputation is a useful approximation of the full dLLM forward pass as measured by the TV distance between sparsely and exactly recomputed token distributions, while Figure[9](https://arxiv.org/html/2609.36202#S4.F9)a shows that both sparse and exact recomputation change a large fraction of token values during decoding\. Finally, Appendix[D\.6](https://arxiv.org/html/2609.36202#A4.SS6)shows that generation quality is largely agnostic to the order of the candidate set for which autoregressive decoding is performed\.

\(a\)\(b\)
Figure 9:Analysis of sparse dLLM recomputation vs\. exact dLLM recomputation\.

## 5Conclusion

In this paper, we introduced FastGuide, a training\-free hybrid decoding algorithm to accelerate reward\-guidance for diffusion language models\. Our findings show that while guidance can be reused across diffusion timesteps, guidance can worsen the token incompatibilities of parallel decoding\. To address this, FastGuide performs guidance caching combined with sparse dLLM recomputation to autoregressively and efficiently recompute each proposed token distribution in a decoding step\. FastGuide achieves speedups up to4\.4×4\.4\\timesover sequential guided decoding while retaining similar generation quality\. To the best of our knowledge, our work is the first comprehensive study of accelerating training\-free, gradient\-based reward guidance for diffusion language models and the first evaluation of parallel decoders under reward guidance\. We hope the ideas proposed in this paper will inspire future work on developing parallel decoders specifically for the guided setting\.

## References

- Agrawalet al\.\(2025\)S\. Agrawal, R\. Garrepalli, R\. Goel, C\. Lott, F\. Porikli, and M\. LeeStructuring The Future: Diffusion LLM Speculative Decoding via Calibrated Draft Graphs\.arXiv preprint arXiv:2509\.18085\.Cited by:[Appendix A](https://arxiv.org/html/2609.36202#A1.SS0.SSS0.Px3.p1.1)\.
- Austinet al\.\(2021\)J\. Austin, D\. D\. Johnson, J\. Ho, D\. Tarlow, and R\. Van Den BergStructured denoising diffusion models in discrete state\-spaces\.Advances in neural information processing systems34,pp\. 17981–17993\.Cited by:[§1](https://arxiv.org/html/2609.36202#S1.p1.1),[§2](https://arxiv.org/html/2609.36202#S2.SS0.SSS0.Px1.p1.1)\.
- Ben\-Hamuet al\.\(2025\)H\. Ben\-Hamu, I\. Gat, D\. Severo, N\. Nolte, and B\. KarrerAccelerated sampling from masked diffusion models via entropy bounded unmasking\.InAdvances in Neural Information Processing Systems,Note:arXiv:2505\.24857Cited by:[§B\.2](https://arxiv.org/html/2609.36202#A2.SS2.SSS0.Px5.p1.1),[§1](https://arxiv.org/html/2609.36202#S1.p3.1),[§2](https://arxiv.org/html/2609.36202#S2.SS0.SSS0.Px2.p1.2),[§3\.1\.1](https://arxiv.org/html/2609.36202#S3.SS1.SSS1.p3.1),[§4](https://arxiv.org/html/2609.36202#S4.SS0.SSS0.Px2.p1.1)\.
- Castilloet al\.\(2023\)A\. Castillo, J\. Kohler, J\. C\. Pérez, J\. P\. Pérez, A\. Pumarola, B\. Ghanem, P\. Arbeláez, and A\. ThabetAdaptive Guidance: Training\-free Acceleration of Conditional Diffusion Models\.arXiv preprint arXiv:2312\.12487\.Cited by:[Appendix A](https://arxiv.org/html/2609.36202#A1.SS0.SSS0.Px2.p1.1)\.
- Cuiet al\.\(2026\)J\. Cui, H\. Ye, R\. Tian, H\. Guo, J\. Jiang, H\. Li, C\. Ren, Y\. Huang, K\. Zhu, Z\. Yu, K\. Zhou, and J\. ShangSimSD: Simple Speculative Decoding in Diffusion Language Models\.arXiv preprint arXiv:2606\.02544\.Cited by:[Appendix A](https://arxiv.org/html/2609.36202#A1.SS0.SSS0.Px3.p1.1)\.
- Daraset al\.\(2024\)G\. Daras, H\. Chung, C\. Lai, Y\. Mitsufuji, J\. C\. Ye, P\. Milanfar, A\. G\. Dimakis, and M\. DelbracioA survey on diffusion models for inverse problems\.arXiv preprint arXiv:2410\.00083\.Cited by:[§2](https://arxiv.org/html/2609.36202#S2.SS0.SSS0.Px3.p1.1)\.
- Dhariwal and Nichol \(2021\)P\. Dhariwal and A\. NicholDiffusion models beat gans on image synthesis\.Advances in neural information processing systems34,pp\. 8780–8794\.Cited by:[§2](https://arxiv.org/html/2609.36202#S2.SS0.SSS0.Px3.p1.1)\.
- Dinhet al\.\(2024\)A\. Dinh, D\. Liu, and C\. XuCompress Guidance in Conditional Diffusion Sampling\.arXiv preprint arXiv:2408\.11194\.Cited by:[Appendix A](https://arxiv.org/html/2609.36202#A1.SS0.SSS0.Px2.p1.1)\.
- Gaoet al\.\(2023\)L\. Gao, J\. Schulman, and J\. HiltonScaling laws for reward model overoptimization\.InInternational conference on machine learning,pp\. 10835–10866\.Cited by:[§B\.2](https://arxiv.org/html/2609.36202#A2.SS2.SSS0.Px7.p1.1),[§4](https://arxiv.org/html/2609.36202#S4.SS0.SSS0.Px3.p1.1)\.
- Gaoet al\.\(2025\)Y\. Gao, Z\. Ji, Y\. Wang, B\. Qi, H\. Xu, and L\. ZhangSelf Speculative Decoding for Diffusion Large Language Models\.arXiv preprint arXiv:2510\.04147\.Cited by:[Appendix A](https://arxiv.org/html/2609.36202#A1.SS0.SSS0.Px3.p1.1)\.
- Google DeepMind \(2025\)Google DeepMindGemini diffusion\.Note:[https://blog\.google/technology/google\-deepmind/gemini\-diffusion/](https://blog.google/technology/google-deepmind/gemini-diffusion/)Cited by:[§1](https://arxiv.org/html/2609.36202#S1.p2.1)\.
- Hanet al\.\(2026\)L\. Han, H\. Wang, H\. Gao, K\. Xu, and A\. SrivastavaS2D2: Fast Decoding for Diffusion LLMs via Training\-Free Self\-Speculation\.arXiv preprint arXiv:2603\.25702\.Cited by:[Appendix A](https://arxiv.org/html/2609.36202#A1.SS0.SSS0.Px3.p1.1)\.
- Honget al\.\(2025\)F\. Hong, G\. Yu, Y\. Ye, H\. Huang, H\. Zheng, Y\. Zhang, Y\. Wang, and J\. YaoWide\-In, Narrow\-Out: Revokable Decoding for Efficient and Effective DLLMs\.arXiv preprint arXiv:2507\.18578\.Cited by:[Appendix A](https://arxiv.org/html/2609.36202#A1.SS0.SSS0.Px3.p1.1)\.
- Inception Labs \(2025\)Inception LabsMercury: ultra\-fast language models based on diffusion\.arXiv preprint arXiv:2506\.17298\.Cited by:[§1](https://arxiv.org/html/2609.36202#S1.p2.1)\.
- Israelet al\.\(2026\)D\. Israel, G\. Van den Broeck, and A\. GroverAccelerating diffusion llms via adaptive parallel decoding\.Advances in neural information processing systems38,pp\. 52870–52888\.Cited by:[Appendix A](https://arxiv.org/html/2609.36202#A1.SS0.SSS0.Px3.p1.1),[§2](https://arxiv.org/html/2609.36202#S2.SS0.SSS0.Px2.p1.2)\.
- Kimet al\.\(2026a\)B\. Kim, D\. Jeon, M\. Jeon, and A\. NoDAPD: dependency\-aware parallel decoding via attention for diffusion llms\.arXiv preprint arXiv:2603\.12996\.Cited by:[§B\.2](https://arxiv.org/html/2609.36202#A2.SS2.SSS0.Px6.p1.1),[§D\.9](https://arxiv.org/html/2609.36202#A4.SS9.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.36202#S1.p3.1),[§2](https://arxiv.org/html/2609.36202#S2.SS0.SSS0.Px2.p1.2),[§4](https://arxiv.org/html/2609.36202#S4.SS0.SSS0.Px2.p1.1)\.
- Kimet al\.\(2025a\)M\. Kim, C\. Hooper, A\. Tomar, C\. Xu, M\. Farajtabar, M\. W\. Mahoney, K\. Keutzer, and A\. GholamiBeyond next\-token prediction: a performance characterization of diffusion versus autoregressive language models\.arXiv preprint arXiv:2510\.04146\.Cited by:[§D\.4](https://arxiv.org/html/2609.36202#A4.SS4.p1.1)\.
- Kimet al\.\(2025b\)S\. H\. Kim, S\. Hong, H\. Jung, Y\. Park, and S\. YunKLASS: KL\-guided fast inference in masked diffusion models\.InAdvances in Neural Information Processing Systems,Note:arXiv:2511\.05664Cited by:[§B\.2](https://arxiv.org/html/2609.36202#A2.SS2.SSS0.Px4.p1.1),[§2](https://arxiv.org/html/2609.36202#S2.SS0.SSS0.Px2.p1.1),[§2](https://arxiv.org/html/2609.36202#S2.SS0.SSS0.Px2.p1.2),[§4](https://arxiv.org/html/2609.36202#S4.SS0.SSS0.Px2.p1.1)\.
- Kimet al\.\(2026b\)Y\. Kim, D\. Shin, B\. Na, M\. Park, R\. L\. Kim, and I\. MoonLookahead Sample Reward Guidance for Test\-Time Scaling of Diffusion Models\.arXiv preprint arXiv:2602\.03211\.Cited by:[Appendix A](https://arxiv.org/html/2609.36202#A1.SS0.SSS0.Px2.p1.1)\.
- Leeet al\.\(2026\)Y\. Lee, S\. Dan, J\. Lee, J\. Park, and J\. H\. AhnDyLLM: efficient diffusion LLM inference via saliency\-based token selection and partial attention\.arXiv preprint arXiv:2603\.08026\.Cited by:[Appendix A](https://arxiv.org/html/2609.36202#A1.SS0.SSS0.Px4.p1.1)\.
- Liet al\.\(2024\)B\. Li, Y\. Wang, A\. Lochab, A\. Grama, and R\. ZhangCascade Reward Sampling for Efficient Decoding\-Time Alignment\.arXiv preprint arXiv:2406\.16306\.Cited by:[Appendix A](https://arxiv.org/html/2609.36202#A1.SS0.SSS0.Px2.p1.1)\.
- Lianget al\.\(2026\)K\. Liang, X\. Tan, A\. Zhong, H\. Xu, and M\. CaniniFOCUS: DLLMs know how to tame their compute bound\.InInternational Conference on Machine Learning,Cited by:[§D\.4](https://arxiv.org/html/2609.36202#A4.SS4.p1.1)\.
- Liaoet al\.\(2025\)B\. Liao, Y\. Xu, H\. Dong, J\. Li, C\. Monz, S\. Savarese, D\. Sahoo, and C\. XiongReward\-Guided Speculative Decoding for Efficient LLM Reasoning\.arXiv preprint arXiv:2501\.19324\.Cited by:[Appendix A](https://arxiv.org/html/2609.36202#A1.SS0.SSS0.Px2.p1.1)\.
- Liuet al\.\(2025a\)A\. Liu, O\. Broadrick, M\. Niepert, and G\. Van den BroeckDiscrete copula diffusion\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 88953–88979\.Cited by:[Appendix A](https://arxiv.org/html/2609.36202#A1.SS0.SSS0.Px3.p1.1),[§2](https://arxiv.org/html/2609.36202#S2.SS0.SSS0.Px2.p1.2)\.
- Liuet al\.\(2025b\)Y\. Liu, Z\. Yao, R\. Min, Y\. Cao, L\. Hou, and J\. LiRm\-bench: benchmarking reward models of language models with subtlety and style\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 44323–44355\.Cited by:[§B\.1](https://arxiv.org/html/2609.36202#A2.SS1.SSS0.Px2.p1.1),[§2](https://arxiv.org/html/2609.36202#S2.SS0.SSS0.Px3.p1.1),[§3](https://arxiv.org/html/2609.36202#S3.p1.1),[§4](https://arxiv.org/html/2609.36202#S4.SS0.SSS0.Px1.p1.1)\.
- Liuet al\.\(2026\)Y\. Liu, L\. Zeng, Y\. Xiao, J\. He, J\. Liu, C\. Wang, R\. Yan, W\. Shen, F\. Zhang, J\. Xu,et al\.Skywork\-reward\-v2: scaling preference data curation via human\-ai synergy\.InInternational Conference on Learning Representations,Vol\.2026,pp\. 133805–133838\.Cited by:[§1](https://arxiv.org/html/2609.36202#S1.p2.2),[§3](https://arxiv.org/html/2609.36202#S3.p1.1),[§4](https://arxiv.org/html/2609.36202#S4.SS0.SSS0.Px2.p1.1)\.
- Liuet al\.\(2025c\)Z\. Liu, Y\. Yang, Y\. Zhang, J\. Chen, C\. Zou, Q\. Wei, S\. Wang, Y\. Zhu, and L\. ZhangdLLM\-Cache: accelerating diffusion large language models with adaptive caching\.arXiv preprint arXiv:2506\.06295\.Cited by:[Appendix A](https://arxiv.org/html/2609.36202#A1.SS0.SSS0.Px4.p1.1),[§3\.1\.2](https://arxiv.org/html/2609.36202#S3.SS1.SSS2.p1.1)\.
- Lvet al\.\(2024\)Z\. Lv, C\. Si, J\. Song, Z\. Yang, Y\. Qiao, Z\. Liu, and K\. K\. WongFasterCache: Training\-Free Video Diffusion Model Acceleration with High Quality\.arXiv preprint arXiv:2410\.19355\.Cited by:[Appendix A](https://arxiv.org/html/2609.36202#A1.SS0.SSS0.Px2.p1.1)\.
- Maet al\.\(2025\)X\. Ma, R\. Yu, G\. Fang, and X\. WangdKV\-Cache: the cache for diffusion language models\.InAdvances in Neural Information Processing Systems,Note:arXiv:2505\.15781Cited by:[Appendix A](https://arxiv.org/html/2609.36202#A1.SS0.SSS0.Px3.p2.1),[Appendix A](https://arxiv.org/html/2609.36202#A1.SS0.SSS0.Px4.p1.1),[item 2](https://arxiv.org/html/2609.36202#S1.I1.i2.p1.1),[§3\.1\.2](https://arxiv.org/html/2609.36202#S3.SS1.SSS2.p1.1),[§3](https://arxiv.org/html/2609.36202#S3.p1.1)\.
- Maliket al\.\(2025\)S\. Malik, V\. Pyatkin, S\. Land, J\. Morrison, N\. A\. Smith, H\. Hajishirzi, and N\. LambertRewardBench 2: advancing reward model evaluation\.arXiv preprint arXiv:2506\.01937\.Cited by:[§B\.1](https://arxiv.org/html/2609.36202#A2.SS1.SSS0.Px3.p1.1),[§4](https://arxiv.org/html/2609.36202#S4.SS0.SSS0.Px1.p1.1)\.
- Mudgalet al\.\(2023\)S\. Mudgal, J\. Lee, H\. Ganapathy, Y\. Li, T\. Wang, Y\. Huang, Z\. Chen, H\. Cheng, M\. Collins, T\. Strohman, J\. Chen, A\. Beutel, and A\. BeiramiControlled Decoding from Language Models\.arXiv preprint arXiv:2310\.17022\.Cited by:[Appendix A](https://arxiv.org/html/2609.36202#A1.SS0.SSS0.Px2.p1.1)\.
- Murataet al\.\(2024\)N\. Murata, C\. Lai, Y\. Takida, T\. Uesaka, B\. Nguyen, S\. Ermon, and Y\. MitsufujiG2D2: gradient\-guided discrete diffusion for inverse problem solving\.arXiv preprint arXiv:2410\.14710\.Cited by:[Appendix A](https://arxiv.org/html/2609.36202#A1.SS0.SSS0.Px1.p1.1),[§D\.1](https://arxiv.org/html/2609.36202#A4.SS1.p1.1),[§1](https://arxiv.org/html/2609.36202#S1.p2.1),[§2](https://arxiv.org/html/2609.36202#S2.SS0.SSS0.Px3.p1.1)\.
- Nieet al\.\(2026\)S\. Nie, F\. Zhu, Z\. You, X\. Zhang, J\. Ou, J\. Hu, J\. Zhou, Y\. Lin, J\. Wen, and C\. LiLarge language diffusion models\.Advances in Neural Information Processing Systems38,pp\. 50608–50646\.Cited by:[§B\.2](https://arxiv.org/html/2609.36202#A2.SS2.SSS0.Px2.p1.1),[Appendix C](https://arxiv.org/html/2609.36202#A3.p2.1),[§1](https://arxiv.org/html/2609.36202#S1.p1.1),[§2](https://arxiv.org/html/2609.36202#S2.SS0.SSS0.Px1.p1.3),[§2](https://arxiv.org/html/2609.36202#S2.SS0.SSS0.Px2.p1.1),[§3](https://arxiv.org/html/2609.36202#S3.p1.1),[§4](https://arxiv.org/html/2609.36202#S4.SS0.SSS0.Px2.p1.1)\.
- Ouyanget al\.\(2022\)L\. Ouyang, J\. Wu, X\. Jiang, D\. Almeida, C\. Wainwright, P\. Mishkin, C\. Zhang, S\. Agarwal, K\. Slama, A\. Ray,et al\.Training language models to follow instructions with human feedback\.InAdvances in Neural Information Processing Systems,Vol\.35\.Cited by:[Appendix A](https://arxiv.org/html/2609.36202#A1.SS0.SSS0.Px1.p1.1)\.
- Paniet al\.\(2025\)C\. Pani, Z\. Ou, and Y\. LiTest\-time alignment of discrete diffusion models with sequential Monte Carlo\.arXiv preprint arXiv:2505\.22524\.Cited by:[Appendix A](https://arxiv.org/html/2609.36202#A1.SS0.SSS0.Px1.p1.1)\.
- Rafailovet al\.\(2023\)R\. Rafailov, A\. Sharma, E\. Mitchell, C\. D\. Manning, S\. Ermon, and C\. FinnDirect preference optimization: your language model is secretly a reward model\.InAdvances in Neural Information Processing Systems,Vol\.36\.Cited by:[Appendix A](https://arxiv.org/html/2609.36202#A1.SS0.SSS0.Px1.p1.1)\.
- Routet al\.\(2025\)L\. Rout, A\. Lugmayr, Y\. Jafarian, S\. Varadharajan, C\. Caramanis, S\. Shakkottai, and I\. Kemelmacher\-ShlizermanTest\-time anchoring for discrete diffusion posterior sampling\.arXiv preprint arXiv:2510\.02291\.Cited by:[Appendix A](https://arxiv.org/html/2609.36202#A1.SS0.SSS0.Px1.p1.1),[§D\.1](https://arxiv.org/html/2609.36202#A4.SS1.p1.1),[§2](https://arxiv.org/html/2609.36202#S2.SS0.SSS0.Px3.p1.1)\.
- Saad\-Falconet al\.\(2025\)J\. Saad\-Falcon, R\. Vivek, W\. Berrios, N\. S\. Naik, M\. Franklin, B\. Vidgen, A\. Singh, D\. Kiela, and S\. MehriLMUNIT: fine\-grained evaluation with natural language unit tests\.\.InEMNLP \(Findings\),pp\. 3303–3324\.Cited by:[Figure 10](https://arxiv.org/html/2609.36202#A2.F10),[§B\.2](https://arxiv.org/html/2609.36202#A2.SS2.SSS0.Px7.p1.1),[§4](https://arxiv.org/html/2609.36202#S4.SS0.SSS0.Px3.p1.1)\.
- Sahooet al\.\(2024\)S\. S\. Sahoo, M\. Arriola, Y\. Schiff, A\. Gokaslan, E\. Marroquin, J\. T\. Chiu, A\. Rush, and V\. KuleshovSimple and effective masked diffusion language models\.Advances in Neural Information Processing Systems37,pp\. 130136–130184\.Cited by:[§1](https://arxiv.org/html/2609.36202#S1.p1.1),[§2](https://arxiv.org/html/2609.36202#S2.SS0.SSS0.Px1.p1.1)\.
- Singhalet al\.\(2025\)R\. Singhal, Z\. Horvitz, R\. Teehan, M\. Ren, Z\. Yu, K\. McKeown, and R\. RanganathA general framework for inference\-time scaling and steering of diffusion models\.InInternational Conference on Machine Learning,Cited by:[Appendix A](https://arxiv.org/html/2609.36202#A1.SS0.SSS0.Px1.p1.1)\.
- Skretaet al\.\(2025\)M\. Skreta, T\. Akhound\-Sadegh, V\. Ohanesian, R\. Bondesan, A\. Aspuru\-Guzik, A\. Doucet, R\. Brekelmans, A\. Tong, and K\. NeklyudovFeynman\-Kac correctors in diffusion: annealing, guidance, and product of experts\.InInternational Conference on Machine Learning,Cited by:[Appendix A](https://arxiv.org/html/2609.36202#A1.SS0.SSS0.Px1.p1.1)\.
- Tanet al\.\(2025\)S\. Tan, S\. Zhuang, K\. Montgomery, W\. Y\. Tang, A\. Cuadron, C\. Wang, R\. A\. Popa, and I\. StoicaJudgeBench: a benchmark for evaluating llm\-based judges\.InInternational Conference on Learning Representations,Cited by:[§B\.1](https://arxiv.org/html/2609.36202#A2.SS1.SSS0.Px1.p1.1),[§4](https://arxiv.org/html/2609.36202#S4.SS0.SSS0.Px1.p1.1)\.
- Tejaswiet al\.\(2026\)A\. Tejaswi, L\. Rout, C\. Caramanis, S\. Shakkottai, and S\. SanghaviEntropy aware reward guidance for diffusion language model alignment\.arXiv preprint arXiv:2602\.05000\.Cited by:[Appendix A](https://arxiv.org/html/2609.36202#A1.SS0.SSS0.Px1.p1.1),[§B\.2](https://arxiv.org/html/2609.36202#A2.SS2.SSS0.Px1.p1.1),[§B\.2](https://arxiv.org/html/2609.36202#A2.SS2.SSS0.Px2.p1.1),[§B\.2](https://arxiv.org/html/2609.36202#A2.SS2.SSS0.Px7.p1.1),[Figure 1](https://arxiv.org/html/2609.36202#S1.F1),[§1](https://arxiv.org/html/2609.36202#S1.p2.1),[§1](https://arxiv.org/html/2609.36202#S1.p2.2),[§2](https://arxiv.org/html/2609.36202#S2.SS0.SSS0.Px3.p1.1),[§3\.1](https://arxiv.org/html/2609.36202#S3.SS1.p1.1),[§3](https://arxiv.org/html/2609.36202#S3.p1.1),[§4](https://arxiv.org/html/2609.36202#S4.SS0.SSS0.Px2.p1.1),[§4](https://arxiv.org/html/2609.36202#S4.SS0.SSS0.Px3.p1.1)\.
- Thakeret al\.\(2025\)D\. Thaker, A\. Goyal, and R\. VidalFrequency\-guided posterior sampling for diffusion\-based image restoration\.In2025 IEEE/CVF International Conference on Computer Vision \(ICCV\),pp\. 12873–12882\.Cited by:[§2](https://arxiv.org/html/2609.36202#S2.SS0.SSS0.Px3.p1.1)\.
- Wanget al\.\(2025\)G\. Wang, Y\. Schiff, S\. S\. Sahoo, and V\. KuleshovRemasking Discrete Diffusion Models with Inference\-Time Scaling\.arXiv preprint arXiv:2503\.00307\.Cited by:[Appendix A](https://arxiv.org/html/2609.36202#A1.SS0.SSS0.Px3.p1.1)\.
- Wenet al\.\(2026\)Z\. Wen, J\. Qu, Z\. Chen, X\. Lu, D\. Liu, Z\. Liu, R\. Wu, Y\. Yang, X\. Jin, H\. Xu,et al\.The devil behind the mask: an emergent safety vulnerability of diffusion llms\.InInternational Conference on Learning Representations,Vol\.2026,pp\. 5460–5489\.Cited by:[§1](https://arxiv.org/html/2609.36202#S1.p2.1)\.
- Wuet al\.\(2025\)C\. Wu, H\. Zhang, S\. Xue, Z\. Liu, S\. Diao, L\. Zhu, P\. Luo, S\. Han, and E\. XieFast\-dllm: training\-free acceleration of diffusion llm by enabling kv cache and parallel decoding\.arXiv preprint arXiv:2505\.22618\.Cited by:[§B\.2](https://arxiv.org/html/2609.36202#A2.SS2.SSS0.Px3.p1.1),[Figure 1](https://arxiv.org/html/2609.36202#S1.F1),[§1](https://arxiv.org/html/2609.36202#S1.p3.1),[§2](https://arxiv.org/html/2609.36202#S2.SS0.SSS0.Px2.p1.1),[§2](https://arxiv.org/html/2609.36202#S2.SS0.SSS0.Px2.p1.2),[§3\.2](https://arxiv.org/html/2609.36202#S3.SS2.p1.1),[§4](https://arxiv.org/html/2609.36202#S4.SS0.SSS0.Px2.p1.1)\.
- Wuet al\.\(2023\)L\. Wu, B\. L\. Trippe, C\. A\. Naesseth, D\. M\. Blei, and J\. P\. CunninghamPractical and asymptotically exact conditional sampling in diffusion models\.InAdvances in Neural Information Processing Systems,Vol\.36\.Cited by:[Appendix A](https://arxiv.org/html/2609.36202#A1.SS0.SSS0.Px1.p1.1)\.
- Wu and Zhang \(2025\)S\. Wu and J\. ZhangFree Draft\-and\-Verification: Toward Lossless Parallel Decoding for Diffusion Large Language Models\.arXiv preprint arXiv:2510\.00294\.Cited by:[Appendix A](https://arxiv.org/html/2609.36202#A1.SS0.SSS0.Px3.p1.1)\.
- Xianget al\.\(2026\)Y\. Xiang, L\. Wei, Y\. Yao, Q\. Zhu, H\. Yan, C\. Jin, P\. A\. Teare, D\. Zhang, L\. Gui, A\. Saseendran, and Y\. HeStop the Flip\-Flop: Context\-Preserving Verification for Fast Revocable Diffusion Decoding\.arXiv preprint arXiv:2602\.06161\.Cited by:[Appendix A](https://arxiv.org/html/2609.36202#A1.SS0.SSS0.Px3.p1.1)\.
- Xionget al\.\(2026\)Z\. Xiong, Y\. Cai, Z\. Li, and Y\. WangUnveiling the potential of diffusion large language model in controllable generation\.InInternational Conference on Learning Representations,Vol\.2026,pp\. 66670–66686\.Cited by:[§1](https://arxiv.org/html/2609.36202#S1.p2.1)\.
- Xuet al\.\(2024\)M\. Xu, T\. Geffner, K\. Kreis, W\. Nie, Y\. Xu, J\. Leskovec, S\. Ermon, and A\. VahdatEnergy\-Based Diffusion Language Models for Text Generation\.arXiv preprint arXiv:2410\.21357\.Cited by:[Appendix A](https://arxiv.org/html/2609.36202#A1.SS0.SSS0.Px3.p1.1)\.
- Yanet al\.\(2025\)F\. Yan, Q\. Wei, J\. Tang, J\. Li, Y\. Wang, X\. Hu, H\. Li, and L\. ZhangLazyMAR: Accelerating Masked Autoregressive Models via Feature Caching\.arXiv preprint arXiv:2503\.12450\.Cited by:[Appendix A](https://arxiv.org/html/2609.36202#A1.SS0.SSS0.Px2.p1.1)\.
- Yeet al\.\(2025a\)J\. Ye, J\. Gao, S\. Gong, L\. Zheng, X\. Jiang, Z\. Li, and L\. KongBeyond autoregression: discrete diffusion for complex reasoning and planning\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 77875–77898\.Cited by:[§1](https://arxiv.org/html/2609.36202#S1.p1.1)\.
- Yeet al\.\(2025b\)J\. Ye, Z\. Xie, L\. Zheng, J\. Gao, Z\. Wu, X\. Jiang, Z\. Li, and L\. KongDream 7b: diffusion large language models\.arXiv preprint arXiv:2508\.15487\.Cited by:[Appendix C](https://arxiv.org/html/2609.36202#A3.p2.1),[§1](https://arxiv.org/html/2609.36202#S1.p1.1),[§3\.1\.1](https://arxiv.org/html/2609.36202#S3.SS1.SSS1.p2.1),[§3](https://arxiv.org/html/2609.36202#S3.p1.1),[§4](https://arxiv.org/html/2609.36202#S4.SS0.SSS0.Px2.p1.1)\.
- Zhaoet al\.\(2025\)S\. Zhao, D\. Gupta, Q\. Zheng, and A\. GroverD1: scaling reasoning in diffusion large language models via reinforcement learning\.arXiv preprint arXiv:2504\.12216\.Cited by:[Appendix A](https://arxiv.org/html/2609.36202#A1.SS0.SSS0.Px1.p1.1)\.
- Zhaoet al\.\(2024\)S\. Zhao, R\. Brekelmans, A\. Makhzani, and R\. GrosseProbabilistic inference in language models via twisted sequential Monte Carlo\.InInternational Conference on Machine Learning,Cited by:[Appendix A](https://arxiv.org/html/2609.36202#A1.SS0.SSS0.Px1.p1.1)\.
- Zhuet al\.\(2025\)F\. Zhu, R\. Wang, S\. Nie, X\. Zhang, C\. Wu, J\. Hu, J\. Zhou, J\. Chen, Y\. Lin, J\. Wen, and C\. LiLLaDA 1\.5: variance\-reduced preference optimization for large language diffusion models\.arXiv preprint arXiv:2505\.19223\.Cited by:[Appendix A](https://arxiv.org/html/2609.36202#A1.SS0.SSS0.Px1.p1.1)\.

###### Appendix Contents

1. [1Introduction](https://arxiv.org/html/2609.36202#S1)
2. [2Background and Related Work](https://arxiv.org/html/2609.36202#S2)
3. [3Method: FastGuide](https://arxiv.org/html/2609.36202#S3)1. [3\.1Hybrid Decoding: Guidance Caching and dLLM Recomputation](https://arxiv.org/html/2609.36202#S3.SS1)1. [3\.1\.1Counterfactual Analysis of Hybrid Decoding](https://arxiv.org/html/2609.36202#S3.SS1.SSS1) 2. [3\.1\.2Sparse dLLM Recomputation](https://arxiv.org/html/2609.36202#S3.SS1.SSS2) 2. [3\.2Adaptivity via Confidence Deferral](https://arxiv.org/html/2609.36202#S3.SS2)
4. [4Experiments](https://arxiv.org/html/2609.36202#S4)1. [4\.1Main Results](https://arxiv.org/html/2609.36202#S4.SS1) 2. [4\.2Ablation Studies](https://arxiv.org/html/2609.36202#S4.SS2)
5. [5Conclusion](https://arxiv.org/html/2609.36202#S5)
6. [References](https://arxiv.org/html/2609.36202#bib)
7. [AExtended Related Work](https://arxiv.org/html/2609.36202#A1)
8. [BExperimental Details](https://arxiv.org/html/2609.36202#A2)1. [B\.1Benchmarks](https://arxiv.org/html/2609.36202#A2.SS1) 2. [B\.2Hyperparameters](https://arxiv.org/html/2609.36202#A2.SS2) 3. [B\.3Experimental Details for Section](https://arxiv.org/html/2609.36202#A2.SS3)
9. [CFastGuide pseudocode](https://arxiv.org/html/2609.36202#A3)
10. [DAdditional Experiments](https://arxiv.org/html/2609.36202#A4)1. [D\.1Guidance Caching with other Reward Guidance Algorithms](https://arxiv.org/html/2609.36202#A4.SS1) 2. [D\.2FastGuide with Different Reward Models](https://arxiv.org/html/2609.36202#A4.SS2) 3. [D\.3Measuring Average Reward Model Performance and Throughput](https://arxiv.org/html/2609.36202#A4.SS3) 4. [D\.4Effect of Batching](https://arxiv.org/html/2609.36202#A4.SS4) 5. [D\.5Sparse dLLM Verification Window Size Ablation](https://arxiv.org/html/2609.36202#A4.SS5) 6. [D\.6Verification Order Ablation](https://arxiv.org/html/2609.36202#A4.SS6) 7. [D\.7Effect of Greedy Token Selection](https://arxiv.org/html/2609.36202#A4.SS7) 8. [D\.8Extended Computational Analysis](https://arxiv.org/html/2609.36202#A4.SS8) 9. [D\.9Qualitative Examples](https://arxiv.org/html/2609.36202#A4.SS9)

## Appendix AExtended Related Work

##### Other reward guidance methods\.

Aligning dLLM outputs with a reward can be done by fine\-tuning or at inference time\. Training\-based approaches adapt RLHF\-style objectives\([Ouyang et al\., 2022](https://arxiv.org/html/2609.36202#bib.bib26);[Rafailov et al\., 2023](https://arxiv.org/html/2609.36202#bib.bib27)\)to dLLMs, e\.g\., via policy\-gradient or preference optimization\([Zhao et al\., 2025](https://arxiv.org/html/2609.36202#bib.bib28);[Zhu et al\., 2025](https://arxiv.org/html/2609.36202#bib.bib29)\), but require retraining whenever the reward changes\. Training\-free approaches instead steer a frozen dLLM at inference time and split into two families\. Sampling\-based methods such as Best\-Of\-N or Sequential Monte Carlo \(SMC\) methods only require reward evaluations, but can require many sequences/particles to be decoded\([Wu et al\., 2023](https://arxiv.org/html/2609.36202#bib.bib30);[Zhao et al\., 2024](https://arxiv.org/html/2609.36202#bib.bib31);[Singhal et al\., 2025](https://arxiv.org/html/2609.36202#bib.bib32);[Skreta et al\., 2025](https://arxiv.org/html/2609.36202#bib.bib33);[Pani et al\., 2025](https://arxiv.org/html/2609.36202#bib.bib34)\)\. Gradient\-based methods instead differentiate the reward model with respect to the dLLM logits and steer a single trajectory, at the cost of a forward and backward pass through the reward model at every diffusion timestep\([Murata et al\., 2024](https://arxiv.org/html/2609.36202#bib.bib16);[Rout et al\., 2025](https://arxiv.org/html/2609.36202#bib.bib17);[Tejaswi et al\., 2026](https://arxiv.org/html/2609.36202#bib.bib15)\)\. One of the main challenges of reward guidance is obtaining gradients since the tokens in the input sequence are discrete\. To tackle this, existing methods propose to use stop\-gradient approximations\([Rout et al\., 2025](https://arxiv.org/html/2609.36202#bib.bib17)\)or take gradients with respect to the continuous token embeddings\([Murata et al\., 2024](https://arxiv.org/html/2609.36202#bib.bib16);[Tejaswi et al\., 2026](https://arxiv.org/html/2609.36202#bib.bib15)\)\.

##### Accelerating guidance\.

In continuous diffusion, the cost of guidance has been reduced by only computing classifier gradients on a subset of timesteps\([Dinh et al\., 2024](https://arxiv.org/html/2609.36202#bib.bib48)\), by skipping the classifier\-free guidance residual\([Castillo et al\., 2023](https://arxiv.org/html/2609.36202#bib.bib42);[Lv et al\., 2024](https://arxiv.org/html/2609.36202#bib.bib57);[Yan et al\., 2025](https://arxiv.org/html/2609.36202#bib.bib56)\), or by replacing reward backpropagation with a closed\-form estimate\([Kim et al\., 2026b](https://arxiv.org/html/2609.36202#bib.bib43)\)\. Our guidance caching has two key differences\. In continuous diffusion, the state evolves smoothly with the noise level, so guidance reuse is controlled by a timestep schedule\. In contrast, in masked dLLMs, the effect of guidance reuse is position\-dependent as well as timestep\-dependent\. While guidance may change across the entire sequence,O1demonstrates that at high\-confidence positions \(which are precisely the positions that are unmasked\), the guided token distribution is stable\. Second, our work provides a novel observation as to how guidance interacts with the joint\-marginal mismatch \(O2\)\. For traditional autoregressive LLMs, segment\-level and blockwise methods amortize reward\-model evaluations over several tokens\([Li et al\., 2024](https://arxiv.org/html/2609.36202#bib.bib51);[Mudgal et al\., 2023](https://arxiv.org/html/2609.36202#bib.bib45);[Liao et al\., 2025](https://arxiv.org/html/2609.36202#bib.bib50)\), but they use the reward as a score for search or rejection rather than as a gradient, and further do not face the joint\-marginal mismatch since tokens are generated sequentially\.

##### Parallel decoding\.

We now review other parallel decoders\. The first class uses external models to supply joint structure to mitigate the joint\-marginal mismatch\. For instance,[Israel et al\. \(2026\)](https://arxiv.org/html/2609.36202#bib.bib25),[Liu et al\. \(2025a\)](https://arxiv.org/html/2609.36202#bib.bib14), and[Xu et al\. \(2024\)](https://arxiv.org/html/2609.36202#bib.bib49)use the structure from an autoregressive verifier, a copula, or an energy function respectively\. However, in the reward\-guided setting, this requires the auxiliary model to be consistent with the guided distribution, or yield similarly high\-reward results\. The second class is training\-free self\-verification methods\. Self\-speculative decoders\([Gao et al\., 2025](https://arxiv.org/html/2609.36202#bib.bib52);[Wu and Zhang, 2025](https://arxiv.org/html/2609.36202#bib.bib47);[Agrawal et al\., 2025](https://arxiv.org/html/2609.36202#bib.bib53);[Han et al\., 2026](https://arxiv.org/html/2609.36202#bib.bib54);[Cui et al\., 2026](https://arxiv.org/html/2609.36202#bib.bib55)\)perform parallel generation by drafting several tokens and verifying each draft conditioned on the earlier ones, using batched full forward passes or a single pass over a duplicated sequence\. However, the goal of these methods is to remain lossless by using rejection sampling techniques, which also do not cleanly extend to the reward\-guided setting\. The final class of methods is revocable parallel decoders\([Wang et al\., 2025](https://arxiv.org/html/2609.36202#bib.bib46);[Hong et al\., 2025](https://arxiv.org/html/2609.36202#bib.bib58);[Xiang et al\., 2026](https://arxiv.org/html/2609.36202#bib.bib44)\), which unmask first and re\-mask tokens that fail a verification step\. We emphasize that these methods have not been evaluated under gradient\-based reward guidance, which presents its own challenges of guidance recomputation and the influence of guidance on token compatibilities \(Section[3](https://arxiv.org/html/2609.36202#S3)\)\.

FastGuide differs from these methods in two main ways\. First, FastGuide trades exactness for cost: rather than relying on an auxiliary model or on full forward passes to remain lossless, it approximates each recomputed token distribution with a sparse forward pass over a small window\([Ma et al\., 2025](https://arxiv.org/html/2609.36202#bib.bib4)\), and we validate this approximation empirically \(Section[4\.2](https://arxiv.org/html/2609.36202#S4.SS2)\)\. Second, FastGuide acts*before*unmasking: the recomputed distribution is used to re\-select the token value and to defer low\-confidence candidates, rather than to accept or reject a drafted token or to re\-mask a token after it has already entered the context\.

##### Inference\-time speedups\.

Several works have accelerated dLLM inference in the unguided setting beyond parallel decoding methods\. The main challenge of accelerating dLLMs compared to standard autoregressive models is that dLLMs utilize bidirectional attention, so each token representation depends on every other token\. In contrast, the use of causal attention for autoregressive models restricts dependencies of tokens to previously generated tokens\. To address this, several works such as DyLLM\([Lee et al\., 2026](https://arxiv.org/html/2609.36202#bib.bib6)\), dKV\-Cache\([Ma et al\., 2025](https://arxiv.org/html/2609.36202#bib.bib4)\), and dLLM\-Cache\([Liu et al\., 2025c](https://arxiv.org/html/2609.36202#bib.bib5)\)propose methods to use selective attention, diffusion model layer skipping, and KV caching to speed up diffusion model forward passes\. Our sparse dLLM recomputation primitive draws inspiration from these techniques\. Our work goes beyond these methods by disentangling the update of dLLM logits and the guidance vector in reward\-guided decoding, resulting in our hybrid decoding strategy\. As Section[4](https://arxiv.org/html/2609.36202#S4)shows, autoregressive dLLM recomputation helps recondition the dLLM after every token unmasking, reducing token incompatibilities among co\-proposed tokens, which guidance makes more frequent \(O2\)\. We leave further optimization of dLLM forward passes such as attention\-based KV caching\([Lee et al\., 2026](https://arxiv.org/html/2609.36202#bib.bib6)\)to future work\.

## Appendix BExperimental Details

### B\.1Benchmarks

The descriptions below use*high reward*and*low reward*for the preferred and rejected responses supplied by each benchmark\. In our generation experiments, we use only the benchmark prompts as conditions, generate new responses with the dLLM, and score those responses with the reward model\. Thus, we do not use the preferred and rejected responses for each benchmark during generation\.

##### JudgeBench\.

JudgeBench\([Tan et al\., 2025](https://arxiv.org/html/2609.36202#bib.bib19)\)evaluates whether a judge can distinguish objectively correct answers from plausible but incorrect ones in four categories: general knowledge, reasoning, mathematics, and coding\. Prompts are difficult multiple\-choice knowledge questions, logic problems, competition\-style mathematics problems, or programming specifications\. Each example pairs two responses produced by the same model; the objectively correct response is the high\-reward response, while the response containing a subtle factual, logical, mathematical, or implementation error is the low\-reward response\. The benchmark consists of 270 prompts\.

##### RM\-Bench\.

RM\-Bench\([Liu et al\., 2025b](https://arxiv.org/html/2609.36202#bib.bib18)\)tests whether reward models prioritize substantive quality over response style across five domains:chat,code,math,safety\-refuse, andsafety\-response\. Prompts range from everyday information requests to programming, mathematics, and safety\-sensitive questions\. High\-reward responses are correct, helpful, and appropriately safe, whereas low\-reward responses contain a small but decisive error, unsafe guidance, or an unnecessary refusal\. Each response is provided in concise, detailed plain\-text, and detailed Markdown forms, allowing the benchmark to test whether polished style masks lower\-quality content\. The benchmark consists of 1327 prompts\.

##### Reward\-Bench\-2\.

Reward\-Bench\-2\([Malik et al\., 2025](https://arxiv.org/html/2609.36202#bib.bib20)\)covers factuality, precise instruction following, mathematics, safety, focus, and ties\. Its prompts include open\-ended factual questions, tightly constrained writing tasks, mathematical problems, safety scenarios, and general queries for which an answer can be relevant or merely adjacent\. High\-reward responses are factually correct, satisfy the stated constraints, solve the problem, respect the appropriate safety boundary, and directly address the request; low\-reward responses fail one or more of these criteria\. The ties subset instead contains several equally valid high\-reward answers and tests whether they are all ranked above invalid alternatives without imposing an arbitrary ordering among the valid answers\. The benchmark consists of 1825 prompts\.

### B\.2Hyperparameters

##### Sampling Details\.

We fix the generation lengthL=128L=128for all methods \(excluding the benchmark prompt\)\. Tokens are sampled with temperature0\.70\.7for all methods, mirroring the evaluation of[Tejaswi et al\. \(2026\)](https://arxiv.org/html/2609.36202#bib.bib15)\. For FastGuide, this applies both to the first token unmasked in each step and to every subsequent candidate, which is sampled from its recomputed guided distribution \(Algorithm[1](https://arxiv.org/html/2609.36202#alg1)\)\. Appendix[D\.7](https://arxiv.org/html/2609.36202#A4.SS7)evaluates the deterministic variant that takes the most likely recomputed token instead\.

##### Reward Guidance Details\.

We follow the protocol of EntRGi\([Tejaswi et al\., 2026](https://arxiv.org/html/2609.36202#bib.bib15)\), where guidance consists of33gradient backpropagation steps of the reward model at every diffusion timestep\. This choice was found to produce high\-quality samples while mitigating reward hacking\([Tejaswi et al\., 2026](https://arxiv.org/html/2609.36202#bib.bib15)\)\. Since LLaDA\([Nie et al\., 2026](https://arxiv.org/html/2609.36202#bib.bib24)\)has a different tokenizer than the Skywork reward models, we adopt the choice of EntRGi and set the embedding of all non\-overlapping tokens to be the zero embedding\.

Since existing parallel decoding baselines do not consider the guided setting, we tuned the hyperparameters for each method separately\. We consider two regimes of FastGuide: one which achieves higher quality, lower throughput and the other which achieves lower quality, higher throughput\. We refer to these as the*high\-quality regime*and the*high\-throughput regime*, respectively\. FastGuide uses a candidate pool size ofk=8k=8in the first andk=16k=16in the second \(evaluated in Appendix[D](https://arxiv.org/html/2609.36202#A4)\)\. For each regime, we picked the hyperparameters of every baseline such that it gives the highest quality at throughput comparable to FastGuide in that regime\. We measured quality via Top@1, although we found that using LMUnit to measure quality also did not change the hyperparameter choice\. Each baseline was swept on 64 RM\-Bench prompts\. Below, we describe the hyperparameters for each baseline and summarize the choices for each hyperparameter in Table[2](https://arxiv.org/html/2609.36202#A2.T2)\.

##### Fast\-dLLM Hyperparameters\.

Fast\-dLLM\([Wu et al\., 2025](https://arxiv.org/html/2609.36202#bib.bib3)\)has one key parameter, which is the confidence thresholdccfor unmasking\. In every step, all tokens with confidence score higher thanccare unmasked\.

##### KLASS Hyperparameters\.

KLASS\([Kim et al\., 2025b](https://arxiv.org/html/2609.36202#bib.bib7)\)has three parameters: a confidence thresholdcc, a KL thresholdϵ\\epsilon, and a history lengthhh\. In every step, a token is unmasked if its confidence score is higher thanccand the KL divergence between its predicted distributions at consecutive steps has stayed belowϵ\\epsilonfor the lasthhsteps; if no token qualifies, the single most confident token is unmasked\. We sweptc∈\[0\.1,0\.9\]c\\in\[0\.1,0\.9\],ϵ∈\[0\.001,5\]\\epsilon\\in\[0\.001,5\], andh∈\{1,2\}h\\in\\\{1,2\\\}\. We found that the KL stability rule of KLASS does not generalize to guided distributions: at the published settings \(c=0\.9c=0\.9,ϵ≤0\.01\\epsilon\\leq 0\.01,h=2h=2\) it often defaults to sequential decoding, since many token distributions change drastically after guidance, and both thresholds must be relaxed substantially to unmask several tokens per step\. Moreover, once the remaining masked positions predict the end\-of\-sequence token, the rule can stop accepting tokens altogether\. This explains why KLASS has much longer and more variable inference time compared to other baselines in Table[1](https://arxiv.org/html/2609.36202#S4.T1)\.

##### EB\-Sampler Hyperparameters\.

EB\-Sampler\([Ben\-Hamu et al\., 2025](https://arxiv.org/html/2609.36202#bib.bib8)\)has one key parameter, which is the entropy budgetγ\\gamma\. In every step, masked tokens are sorted by increasing entropy of their predicted distribution, and the longest prefix whose total entropy, excluding its largest entry, is at mostγ\\gammais unmasked\. The lowest\-entropy token is always unmasked, soγ=0\\gamma=0recovers sequential decoding and largerγ\\gammaunmasks more tokens per step\. We sweptγ∈\[0\.1,16\]\\gamma\\in\[0\.1,16\]\.

##### DAPD Hyperparameters\.

DAPD\([Kim et al\., 2026a](https://arxiv.org/html/2609.36202#bib.bib13)\)builds a dependency graph over masked tokens, connecting two tokens if their symmetrized attention score \(averaged over heads and over the top30%30\\%of layers\) exceeds a threshold, and unmasks an independent set of this graph in every step\. The threshold increases linearly during decoding fromτmin=0\.005\\tau\_\{\\min\}=0\.005toτmax=0\.05\\tau\_\{\\max\}=0\.05\. Once fewer than half of the tokens remain masked, DAPD additionally unmasks all tokens with confidence score higher than0\.90\.9\. We sweep the threshold schedule by multiplying by a single scaless, since larger thresholds remove edges from the graph and unmask more tokens per step\. We swepts∈\[0\.5,32\]s\\in\[0\.5,32\]\.

Table 2:Operating points of all parallel decoders\.The high\-quality regime trades throughput for quality, and the high\-throughput regime trades quality for throughput\.
##### LMUnit Details\.

The LMUnit score uses a judge LLM to give a score from11to55\. Following[Tejaswi et al\. \(2026\)](https://arxiv.org/html/2609.36202#bib.bib15), we do not treat LMUnit as the primary quality metric\. Instead, it provides an independent assessment of response quality to help detect potential reward overoptimization\([Gao et al\., 2023](https://arxiv.org/html/2609.36202#bib.bib41)\)\. We use the prompt in Figure[10](https://arxiv.org/html/2609.36202#A2.F10)as input to LMUnit, similar to[Saad\-Falcon et al\. \(2025\)](https://arxiv.org/html/2609.36202#bib.bib38)and[Tejaswi et al\. \(2026\)](https://arxiv.org/html/2609.36202#bib.bib15)\. Each response is scored on five fixed unit tests, and we report the mean\.

LMUnit inputQuery:\{benchmark prompt\}Response:\{generated response\}Unit Test:\{unit test\}

Unit tests\(each response is scored on all five; we report the mean\)

RelevanceDoes the response directly and effectively address the user’s request?

CorrectnessIs the information in the response correct and reliable?

ClarityIs the response well\-structured, clear, and fluent?

CompletenessDoes the response appropriately address the full scope of the question?

SafetyIs the response free from harmful, biased, or inappropriate content?

Figure 10:Prompts given to the LMUnit judge\.We also include in the system message the phraserationale=No, which instructs LMUnit to output only a score for the given unit test without a written rationale as suggested by[Saad\-Falcon et al\. \(2025\)](https://arxiv.org/html/2609.36202#bib.bib38)

### B\.3Experimental Details for Section[3](https://arxiv.org/html/2609.36202#S3)

All diagnostic experiments in Section[3](https://arxiv.org/html/2609.36202#S3)use Dream\-7B with the smallerSkywork\-Reward\-V2\-Qwen3\-0\.6Breward model on 250 RM\-Bench prompts \(50 per domain\), with the sampling and guidance settings of Appendix[B\.2](https://arxiv.org/html/2609.36202#A2.SS2)\.

##### Reuse versus recomputation \(Figure[3](https://arxiv.org/html/2609.36202#S3.F3)\)\.

At every step, thek∈\{2,4,8\}k\\in\\\{2,4,8\\\}most confident positions under the guided distribution at the start of the step are fixed as the candidate set and unmasked in descending confidence\. The first token is always sampled from the distribution at the start of the step\. For tokens2,…,k2,\\ldots,k, the four variants of Equation[5](https://arxiv.org/html/2609.36202#S3.E5)either reuse or recompute the dLLM logits \(with an exact full forward pass\) and the guidance vector\. We generate one trajectory per prompt and report the mean reward model score\.

##### Guidance reuse \(Figure[4](https://arxiv.org/html/2609.36202#S3.F4)\)\.

At every stepttand every masked position we form two guided distributions from the current dLLM logits: one with the guidance vector computed at stepttand one with the guidance vector computed at stept−\(k−1\)t\-\(k\-1\), which is the oldest guidance that a parallel step of sizekkwould reuse\. The curves fork=2,4,8k=2,4,8therefore correspond to guidance that is11,33, and77steps old\. We report the total\-variation distance between the two distributions, binned by the confidence of the position at the time the older guidance was computed\. Since guidance computation is stochastic, we fix random seeds depending on position and guidance iteration, isolating only the effect of guidance caching\.

##### Guidance increases incompatible proposals \(Figure[4](https://arxiv.org/html/2609.36202#S3.F4)\)\.

Along a sequential decoding trajectory, we examine6464evenly spaced states\. At each state, we compute the token proposal candidates without guidance and with guidance from the same state, so the comparison changes only the guidance and not the context\. For each, we take the two most confident positionsaaandbbwith proposed tokensxax\_\{a\}andxbx\_\{b\}, which is the pair that a parallel decoder withk=2k=2would unmask\. We then unmaskxax\_\{a\}, run one additional unguided dLLM forward pass, and computePMI=log⁡p⁡\(xb∣xa\)−log⁡p⁡\(xb\)\\mathrm\{PMI\}=\\log p\(x\_\{b\}\\mid x\_\{a\}\)\-\\log p\(x\_\{b\}\)under the dLLM\. A pair is counted as incompatible if its PMI is less than−1\-1, that is, if unmasking the first token lowers the dLLM log probability of the second by a factor ofee\. We include this margin to rule out cases where the probability only decreases by a small amount, which may not signal a true incompatibility\. The figure reports, for each prompt, the number of incompatible pairs among its6464states\. Note that this experiment does not imply parallel decoding is necessarily reliable without guidance \(since the joint\-marginal mismatch still persists\)\.

## Appendix CFastGuide pseudocode

Algorithm[1](https://arxiv.org/html/2609.36202#alg1)summarizes one hybrid decoding step of FastGuide as described in Section[3](https://arxiv.org/html/2609.36202#S3)\.

Algorithm 1One hybrid decoding step of FastGuide\.0:State

𝒙t\{\\bm\{x\}\}\_\{t\}, candidate budget

kk, window radius

ww, threshold

τ\\tau
1:Compute

ℓt,0\\ell\_\{t,0\}and

𝒓t,0\{\\bm\{r\}\}\_\{t,0\}with one full guided pass; cache all K/V states

2:Rank masked positions by decreasing guided confidence to obtain

\(a1,…,ak\)\(a\_\{1\},\\ldots,a\_\{k\}\)
3:Sample a token at

a1a\_\{1\}from the guided distribution and commit it; set

aprev←a1a\_\{\\mathrm\{prev\}\}\\leftarrow a\_\{1\}
4:for

j=2,…,kj=2,\\ldots,kdo

5:

𝒲j←\\mathcal\{W\}\_\{j\}\\leftarrowwindows of radius

wwaround

apreva\_\{\\mathrm\{prev\}\}and candidate

aja\_\{j\}
6:Recompute K/V states at positions in

𝒲j\\mathcal\{W\}\_\{j\}in every layer and overwrite the corresponding cache entries

7:Sparsely refresh

ℓ^t,j−1aj\\widehat\{\\ell\}\_\{t,j\-1\}^\{a\_\{j\}\}with queries in

𝒲j\\mathcal\{W\}\_\{j\}attending to the updated cache

8:

ϕt,jaj←ℓ^t,j−1aj\+𝒓t,0aj\\phi\_\{t,j\}^\{a\_\{j\}\}\\leftarrow\\widehat\{\\ell\}\_\{t,j\-1\}^\{a\_\{j\}\}\+\{\\bm\{r\}\}\_\{t,0\}^\{a\_\{j\}\}; sample

cj∼softmax⁡\(ϕt,jaj/T\)c\_\{j\}\\sim\\operatorname\{softmax\}\(\\phi\_\{t,j\}^\{a\_\{j\}\}/T\)
9:if

maxc​softmax​\(ϕt,jaj\)​\(c\)≥τ\\max\_\{c\}\\operatorname\{softmax\}\(\\phi\_\{t,j\}^\{a\_\{j\}\}\)\(c\)\\geq\\tauthen

10:Commit

cjc\_\{j\}at

aja\_\{j\};

aprev←aja\_\{\\mathrm\{prev\}\}\\leftarrow a\_\{j\}
11:endif

12:endfor

13:returnUpdated partially masked state

For the Dream text diffusion model\([Ye et al\., 2025b](https://arxiv.org/html/2609.36202#bib.bib9)\), logits for candidateaja\_\{j\}are produced at positionaj−1a\_\{j\}\-1, so we center the attention window at that location\. For the LLaDA text diffusion model\([Nie et al\., 2026](https://arxiv.org/html/2609.36202#bib.bib24)\), we center the window ataja\_\{j\}\.

## Appendix DAdditional Experiments

### D\.1Guidance Caching with other Reward Guidance Algorithms

Table[3](https://arxiv.org/html/2609.36202#A4.T3)shows the performance of FastGuide on other popular reward guidance methods on JudgeBench and RM\-Bench\. The first is the method of[Murata et al\. \(2024\)](https://arxiv.org/html/2609.36202#bib.bib16)\(titled Expectation\), which uses the expected token embedding as input to the reward model\. The second is the method of[Rout et al\. \(2025\)](https://arxiv.org/html/2609.36202#bib.bib17)\(titled APS\), which samples from the token distribution to generate an input to the reward model\. Table[3](https://arxiv.org/html/2609.36202#A4.T3)demonstrates that FastGuide is robust to the reward guidance algorithm and consistently retains the performance of the corresponding sequential decoding algorithm using the same guidance method\.

Table 3:Generalization across reward\-guidance objectives\.FastGuide stays close to sequential decoding under the Expectation and APS objectives \(in parentheses\), while confidence\-based parallel decoding degrades under all three\.
### D\.2FastGuide with Different Reward Models

Table[4](https://arxiv.org/html/2609.36202#A4.T4)tests the performance of FastGuide on different reward models in addition to the Skywork\-Qwen3\-1\.7B reward model tested in Table[1](https://arxiv.org/html/2609.36202#S4.T1)\. For the Qwen3\-0\.6B and Llama\-3\.1\-8B model, FastGuide similarly retains the performance of sequential decoding\. Guidance caching and sparse dLLM recomputation become more important for inference time since guidance computation and the associated reward model backpropagation scales with the reward model size\. On Skywork\-Llama\-3\.1\-8B, inference time decreases from 88\.1 seconds to only 17\.9 seconds, yielding a4\.9×4\.9\\timesspeedup\. Since the Llama\-3\.1\-8B reward model does not fit on one A5000 GPU together with the dLLM, all methods in that block are timed with the reward model on a second GPU\. The other two blocks use a single GPU\.

Table 4:Robustness to the reward model on JudgeBench\.
### D\.3Measuring Average Reward Model Performance and Throughput

In this section, we evaluate the Avg@4 metric, which is the average reward model performance over44trajectories, as well as show the average throughput \(tokens unmasked per step\) for all methods\. Table[5](https://arxiv.org/html/2609.36202#A4.T5)and Table[7](https://arxiv.org/html/2609.36202#A4.T7)show these for Dream and LLaDA text diffusion models respectively\. Average reward model performance highlights the same trends as the top reward model performance\. On both models and all three benchmarks, FastGuide attains the highest Avg@4 among parallel decoders and stays within1\.11\.1points of sequential decoding, essentially matching it on Reward\-Bench\-2\. In contrast, the strongest baseline, DAPD, trails sequential decoding by up to2\.142\.14, and every other guided baseline by at least2\.42\.4\. On LLaDA, Fast\-dLLM and KLASS unmask more tokens per step on JudgeBench and RM\-Bench \(up to10\.810\.8\), but at a large cost in quality, showing that a high number of tokens per step alone does not translate into a better quality\-efficiency tradeoff under guidance\.

Additionally, we consider varying throughput by increasingkkfor FastGuide to1616candidate tokens per step, showing results on Dream diffusion model in Table[6](https://arxiv.org/html/2609.36202#A4.T6)\. Baseline hyperparameters were adjusted accordingly \(see Appendix[B\.2](https://arxiv.org/html/2609.36202#A2.SS2)\)\. In this regime, the gap to sequential decoding grows substantially for all baselines: the Seq Gap of Confidence roughly doubles \(from−2\.15\-2\.15to−3\.09\-3\.09in Table[5](https://arxiv.org/html/2609.36202#A4.T5)to−4\.26\-4\.26to−5\.95\-5\.95\), and that of DAPD, the strongest baseline, grows from at most−1\.83\-1\.83to as much as−4\.07\-4\.07\. Baselines do become faster \(2\.62\.6–4\.94\.9seconds per generation, excluding KLASS\), but only at this cost in quality\. In contrast, the Seq Gap of FastGuide remains within1\.201\.20on all three benchmarks\. Notably, although FastGuide now considers1616candidates per step, its throughput only increases from4\.44\.4–5\.75\.7to5\.55\.5–8\.08\.0tokens per step, and its inference time is essentially unchanged \(6\.26\.2–9\.69\.6versus5\.95\.9–8\.48\.4seconds per generation\)\. This is because many of the additional candidates fall below the confidence threshold after recomputation and are deferred to later steps\. This highlights the benefit of adaptivity\. Rather than trading quality for speed at a fixed rate, FastGuide unmasks additional tokens only when they remain compatible with the tokens already unmasked in the same step\.

Table 5:Full results on Dream\-7B with Avg@4 and tokens/step added\.Table 6:Results on Dream\-7B with higher candidate token set size ofk=16k=16for FastGuide\.Table 7:Full results on LLaDA\-8B with Avg@4 and tokens/step added\.
### D\.4Effect of Batching

The standard dLLM forward pass uses bidirectional attention, making it a largely compute\-bound operation on modern GPUs\([Kim et al\., 2025a](https://arxiv.org/html/2609.36202#bib.bib39);[Liang et al\., 2026](https://arxiv.org/html/2609.36202#bib.bib40)\)\. In contrast, sparse verification drastically reduces the FLOPs per forward pass \(e\.g\. with a window size of22and sequence length of128128, sparse forward pass is only4%4\\%FLOPs of a full forward pass\)\. Yet, the inference time of the sparse forward pass does not retain the same speedups as reduction in FLOPs due to GPU memory overhead\. Specifically, sparse dLLM recomputation transforms the compute\-bound forward pass into a memory\-bound operation\. Thus, typical strategies from autoregressive language models such as batching can increase speedups since a single pass over the weights serves all generations in the batch\.

We study how batching affects the inference time of FastGuide by decodingB∈\{1,2,3,4\}B\\in\\\{1,2,3,4\\\}generations of the same prompt together \(over 50 RM\-Bench prompts\) and reporting the per\-generation wall\-clock time\. Figure[11](https://arxiv.org/html/2609.36202#A4.F11)shows that as batch size increases, sparse recomputation improves more upon exact recomputation and speedups increase\. FastGuide improves from15\.315\.3to5\.75\.7seconds per generation asBBgrows from11to44\(2\.7×2\.7\\times\), whereas its exact recomputation counterpart improves only from18\.518\.5to10\.810\.8seconds \(1\.7×1\.7\\times\)\. Consequently, the speedup of sparse over exact recomputation grows with the batch size, from1\.21×1\.21\\timesatB=1B\{=\}1to1\.89×1\.89\\timesatB=4B\{=\}4, and the speedup of FastGuide over sequential guided decoding grows from2\.9×2\.9\\timesto3\.4×3\.4\\times\. The batch size at which optimal performance is obtained is a function of the GPU memory, but this illustrates an added benefit of sparse dLLM recomputation\.

Figure 11:Sparse dLLM verification benefits more from batching than exact recomputation\.Left: inference time per generation; right: speedup of sparse over exact recomputation\.
### D\.5Sparse dLLM Verification Window Size Ablation

While Figure[9](https://arxiv.org/html/2609.36202#S4.F9)shows the TV distance between sparse and exact recomputation, we now measure the impact on generation quality through the Top@1 reward model score\. We sweep the window size of sparse recomputation in Figure[12](https://arxiv.org/html/2609.36202#A4.F12)\. The figure illustrates that even a small window size of22recovers a significant proportion of generation quality of sequential guided decoding\. The dashed lines labeled Exact show exact dLLM recomputation with a full forward pass at the samekk, further highlighting that sparse recomputation is a useful approximation of exact dLLM recomputation\. The remaining gap to sequential decoding is therefore not caused by sparsity, but by fixing the tokens to be unmasked\. Confidence deferral aims to address this problem by adding adaptivity into the sampling procedure \(Figure[6](https://arxiv.org/html/2609.36202#S4.F6)\)\.

Figure 12:Quality versus recomputation window radiuswwon RM\-Bench\.We evaluate sparse dLLM recomputation as a function of window radiusww\. Here, we remove confidence deferral so allkkcandidates per step are unmasked\.
Table 8:Verification order ablation on RM\-Bench \(k=4k\{=\}4\)\.The order in which candidates are verified has no measurable effect on quality\.

### D\.6Verification Order Ablation

Autoregressive dLLM verification relies on an ordering of the candidates within each hybrid decoding step \(Section[3](https://arxiv.org/html/2609.36202#S3)\)\. In the main paper, we fix this ordering to be from high to low confidence\. We study other ordering schemes in Table[8](https://arxiv.org/html/2609.36202#A4.T8)\. We remove confidence deferral and unmask exactlyk=4k=4tokens each step to isolate the role of ordering\. We find that the ordering schemes within each hybrid decoding step do not meaningfully affect the results, as the dLLM is always reconditioned on each previous unmasking\.

### D\.7Effect of Greedy Token Selection

FastGuide samples every unmasked token at temperature0\.70\.7, similar to the setting for all baseline methods \(Appendix[B\.2](https://arxiv.org/html/2609.36202#A2.SS2)\)\. A natural deterministic variant instead unmasks each candidate after the first of a step as the most likely token under its recomputed distribution\. Table[9](https://arxiv.org/html/2609.36202#A4.T9)compares the two on the full RM\-Bench test set\. Greedy selection improves Top@1 by0\.300\.30and LMUnit by0\.050\.05at a slightly higher throughput, making it a reasonable strategy when determinism is desired\.

Table 9:Greedy selection of recomputed tokens slightly improves FastGuide\.RM\-Bench, Dream\-7B,k=8k=8
### D\.8Extended Computational Analysis

Table[11](https://arxiv.org/html/2609.36202#A4.T11)reports peak GPU memory and teraFLOPs per generation on RM\-Bench\. Most of the FLOPs at inference time are dominated by guidance steps, so guidance caching results in a major reduction in compute compared to sequential guided decoding\. FastGuide further uses fewer FLOPs than other methods because sparse dLLM recomputation lets FastGuide unmask more tokens in each step, and each sparse recomputation pass uses far less compute than a full forward pass\. In the high\-throughput regime \(k=16k=16for FastGuide\), more tokens are deferred than baselines such as DAPD, resulting in slightly increased compute\. Across all methods, FastGuide does not increase the peak memory usage, which is dominated by loading of diffusion model weights and reward model weights\.

In terms of inference time, Table[11](https://arxiv.org/html/2609.36202#A4.T11)shows the breakdown of time in full dLLM forwards \(computed once at the start of each hybrid decoding step\), guidance computation \(computed once at the start of each hybrid decoding step\), and sparse dLLM forwards \(computedk−1k\-1times for every hybrid decoding step\)\.

Table 10:Computational profile on RM\-Bench\.Peak memory and total compute per generation for every method in both regimes of Table[2](https://arxiv.org/html/2609.36202#A2.T2)\.
Table 11:Inference\-time breakdown on RM\-Bench\.Wall\-clock time per generation at batch size44on Dream\-7B\.

### D\.9Qualitative Examples

##### Decoding sequence for Figure[8](https://arxiv.org/html/2609.36202#S4.F8)\.

Figure[13](https://arxiv.org/html/2609.36202#A4.F13)shows the denoising trajectory for the three methods of Figure[8](https://arxiv.org/html/2609.36202#S4.F8)\. While unguided generation with DAPD can produce coherent text, DAPD with guidance can produce incoherent text early on in generation\. This reinforcesO2that guidance can introduce incompatibilities during sampling\.

##### Qualitative generations vs DAPD with guidance\.

As DAPD\([Kim et al\., 2026a](https://arxiv.org/html/2609.36202#bib.bib13)\)remains the strongest baseline in terms of quantitative results \(Table[1](https://arxiv.org/html/2609.36202#S4.T1)\), we compare the quality of generations across decoding trajectories in Figures[14](https://arxiv.org/html/2609.36202#A4.F14)and[15](https://arxiv.org/html/2609.36202#A4.F15)\. DAPD tends to pick tokens spatially far away as it uses dLLM attention as a signal of conditional dependency\. However, guidance can affect this dependency, which may explain why DAPD with guidance tends to introduce incompatible tokens\. For example, in Figure[14](https://arxiv.org/html/2609.36202#A4.F14), certain token sequences such as5\. emonialare not coherent English sentences\. In contrast, FastGuide retains coherency and fluency in generation, obtaining a significantly higher reward model score\. In Figure[15](https://arxiv.org/html/2609.36202#A4.F15), DAPD introduces incorrect numbering \(such as the 7th step repeated twice\)\.

##### Qualitative effect of confidence deferral\.

Figures[16](https://arxiv.org/html/2609.36202#A4.F16),[17](https://arxiv.org/html/2609.36202#A4.F17), and[18](https://arxiv.org/html/2609.36202#A4.F18)show the effect of confidence deferral at one intermediate step for three different prompts\. For instance, in Figure[16](https://arxiv.org/html/2609.36202#A4.F16), the word mindful is proposed three times, whereas after it is unmasked once, the probability of the other token values drop, leading them to be deferred\. The final generations are thus of higher quality, trading off inference time for quality\.

Prompt:If the endpoints of a line segment are \(2, \-2\) and \(10, 4\), what is the length of the segment?

25%

DAPD, guidedReward−1\.22\-1\.22<M\>find<M\><M\><M\>the<M\><M\><M\>we<M\>use<M\><M\><M\>, which<M\>$\\sqrt<M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\>is: $\\boxed\{<M\>\}$

DAPD, unguidedReward\+6\.59\+6\.59<M\>can<M\>the<M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\>1<M\><M\><M\><M\><M\><M\><M\>the given values<M\>we<M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\>\} =<M\><M\>\{1<M\><M\><M\>

FastGuideReward\+10\.44\+10\.44To find the length of the line segment, we can use the distance formula, which is derived from the Pythagorean Theorem\. The formula<M\><M\><M\><M\>sqrt<M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\>units\.<M\>

50%

<M\>find<M\>length<M\>the<M\>segment, we<M\>use<M\><M\>formula, which<M\>$\\sqrt\{\(<M\>\_2<M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\>\.<M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\>sqrt<M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\>segment is<M\><M\><M\><M\>\}$<M\>\.The answer is: $\\boxed\{6\}$

<M\>can<M\>the distance formula to<M\><M\><M\><M\><M\><M\>segment<M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\>is<M\><M\><M\><M\><M\><M\><M\><M\>\-<M\><M\><M\><M\><M\><M\><M\><M\><M\>2 \-<M\>\_1<M\>2\}$<M\>Plugging<M\>the given values<M\>we<M\>$\\sqrt<M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\>8<M\><M\><M\><M\>6<M\><M\><M\><M\><M\><M\>\{<M\><M\>\+<M\><M\><M\>\} =<M\><M\>\{1<M\>0<M\>

To find the length of the line segment, we can use the distance formula, which is derived from the Pythagorean Theorem\. The formula is:<M\>= sqrt\(\(x2 \- x1\)ˆ2 \+ \(y2 \- y1\)ˆ2\) Plugging in the<M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\>length of the<M\>segment<M\>10 units\.<M\>

75%

<M\>find<M\>length<M\>the line segment, we<M\>use<M\>distance formula, which<M\>$\\sqrt\{\(<M\>\_2 \-<M\><M\><M\><M\>2<M\>\(<M\><M\><M\><M\><M\><M\><M\>\)ˆ<M\>\}$\.<M\>ugging<M\>the given<M\><M\><M\><M\><M\>sqrt<M\><M\><M\><M\><M\><M\><M\><M\>\) \+ \(4 \-<M\><M\><M\><M\><M\>\}<M\><M\><M\>\{<M\><M\><M\><M\><M\><M\><M\><M\>\} =<M\>sqrt<M\><M\><M\><M\><M\><M\><M\><M\><M\>sqrt<M\><M\>6<M\><M\>The<M\>of the<M\>segment is<M\>boxed<M\>6\}$<M\>\.The answer is: $\\boxed\{6\}$

<M\>can<M\>the distance formula to<M\><M\>length of<M\><M\>segment<M\><M\><M\>formula<M\><M\><M\><M\>x<M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\>,<M\><M\><M\>\)$ is given<M\><M\>sqrt<M\><M\><M\><M\>\-<M\><M\><M\><M\><M\>\+<M\><M\><M\>2 \-<M\>\_1<M\>2\}$<M\>Plugging<M\>the given values<M\>we<M\>$\\sqrt<M\><M\><M\>\-<M\><M\><M\><M\><M\><M\>4 \-<M\><M\><M\><M\>\}<M\><M\>sqrt<M\>8<M\><M\><M\>6<M\><M\><M\><M\>\\sqrt\{<M\><M\>\+<M\><M\>\} =<M\><M\>\{1<M\>0<M\>

To find the length of the line segment, we can use the distance formula, which is derived from the Pythagorean Theorem\. The formula is:<M\>= sqrt\(\(x2 \- x1\)ˆ2 \+ \(y2 \- y1\)ˆ2\) Plugging in the given<M\><M\><M\>= sqrt\(\(10 \- 2\)ˆ2 \+ \(4 \- \(\-2\)\)ˆ2<M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\>2<M\><M\><M\><M\><M\><M\><M\><M\><M\>6<M\><M\><M\>\(<M\>00<M\><M\>10<M\>, the length of the line segment is 10 units\.<M\>

100%

To find the length of the line segment, we can use the distance formula, which is $\\sqrt\{\(x\_2 \- x\_1\)ˆ2 \+ \(y\_2 \- y\_1\)ˆ2\}$\. Plugging in the given endpoints, we have $\\sqrt\{\(10 \- 2\)ˆ2\) \+ \(4 \- \(\-2\)\)ˆ2\} = \\sqrt\{5ˆ2 \+ 6ˆ2\} = \\sqrt\{10 \+ 36\} = \\sqrt\{46\}$\. The length of the line segment is $\\boxed\{6\}$ units\.The answer is: $\\boxed\{6\}$

We can use the distance formula to find the length of the line segment\. The distance formula between two points $\(x\_1, y\_1\)$ and $\(x\_2, y\_2\)$ is given by $\\sqrt\{\(x\_2 \- x\_1\)ˆ2 \+ \(y\_2 \- y\_1\)ˆ2\}$\. Plugging in the given values, we have $\\sqrt\{\(10 \- 2\)ˆ2 \+ \(4 \- \(\-2\)\)ˆ2\} = \\sqrt\{8ˆ2 \+ 6ˆ2\} = \\sqrt\{64 \+ 36\} = \\sqrt\{100\}

To find the length of the line segment, we can use the distance formula, which is derived from the Pythagorean Theorem\. The formula is: Distance = sqrt\(\(x2 \- x1\)ˆ2 \+ \(y2 \- y1\)ˆ2\) Plugging in the given coordinates: Distance = sqrt\(\(10 \- 2\)ˆ2 \+ \(4 \- \(\-2\)\)ˆ2\) = sqrt\(8ˆ2 \+ 6ˆ2\) = sqrt\(64 \+ 36\) = sqrt\(100\) = 10 So, the length of the line segment is 10 units\.

Figure 13:Guided DAPD unmasks an incorrect final answer while most of its derivation is still masked\.Decoding trajectories for the prompt of Figure[8](https://arxiv.org/html/2609.36202#S4.F8); rows show the sequence after25%25\\%,50%50\\%,75%75\\%,100%100\\%of each method’s decoding steps, where<M\>denotes a masked position\.Prompt:What are different drawers I should have for clothes?

25%

DAPD, guidedReward−0\.61\-0\.61<M\>are<M\><M\>for<M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\>3\.<M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\>clothes so they don’t get mixed<M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\>\.

FastGuideReward\+1\.19\+1\.19Here are a few types of drawers you might consider having for organizing your clothes: 1\. A<M\>drawer for items<M\><M\>,<M\>, sweatshirts, and<M\><M\>2<M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\>,<M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\>\.<M\>

50%

Here are some<M\>for different drawers<M\><M\><M\><M\><M\>1\.<M\><M\>\-<M\><M\><M\><M\><M\><M\>, etc<M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\>\-<M\><M\><M\>t<M\>,<M\><M\>, etc\. 3\.<M\><M\>\-<M\><M\>,<M\><M\><M\><M\><M\><M\><M\><M\><M\><M\>\-<M\><M\><M\>,<M\>,<M\><M\><M\><M\><M\><M\><M\>\-<M\><M\><M\>or<M\><M\>clothes so they don’t get mixed<M\>with<M\><M\>\.<M\>\.<M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\>,<M\><M\><M\><M\><M\>,<M\>\.

Here are a few types of drawers you might consider having for organizing your clothes: 1\. A<M\>drawer for items like shirts,<M\>, sweatshirts, and<M\>\. 2\. A pants drawer for<M\>items like jeans,<M\><M\>, and<M\><M\>\. 3\. A<M\>drawer for<M\><M\>like dresses, skirts, and<M\><M\>\. 4\.<M\><M\>drawer for<M\><M\><M\><M\><M\>,<M\>,<M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\>\.<M\>

75%

Here are some ideas for different drawers to have for clothes<M\>1\. Top drawer \-<M\>, blazers,<M\>, etc\. that you wear over your<M\>\. 2\. Middle drawer \- Shirts, t\-shirts, blouses, etc\. 3\. Bottom drawer \-Jeans, skirts,<M\>, etc\. 4\. accessory drawer \-<M\>, gloves, hats, etc\. 5\.<M\>ial drawer \- for special occasions or dressier clothes so they don’t get mixed up with the rest\. 6\. Underwear drawer \- for socks and<M\>s\. 7\. Winter drawer \-<M\>, blankets, sweatshirts, etc\.

Here are a few types of drawers you might consider having for organizing your clothes: 1\. A<M\>drawer for items like shirts, tops, sweatshirts, and jackets\. 2\. A pants drawer for<M\>items like jeans,<M\><M\>, and<M\><M\>\. 3\. A dress drawer for<M\>items like dresses, skirts, and bl<M\>\. 4\. A underwear drawer for<M\><M\>items like underwear, socks, and undergarments\. 5\. An accessories drawer for items like<M\>,<M\>, scarves,<M\><M\><M\><M\><M\>\.<M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\>, and<M\><M\><M\>\.<M\>

100%

Here are some ideas for different drawers to have for clothes\. 1\. Top drawer \- Jackets, blazers, skirts, etc\. that you wear over your clothes\. 2\. Middle drawer \- Shirts, t\-shirts, blouses, etc\. 3\. Bottom drawer \-Jeans, skirts, shorts, etc\. 4\. accessory drawer \- masks, gloves, hats, etc\. 5\.emonial drawer \- for special occasions or dressier clothes so they don’t get mixed up with the rest\. 6\. Underwear drawer \- for socks and briefs\. 7\. Winter drawer \- coats, blankets, sweatshirts, etc\.

Here are a few types of drawers you might consider having for organizing your clothes: 1\. A top drawer for items like shirts, tops, sweatshirts, and jackets\. 2\. A pants drawer for bottom items like jeans, khakis, and slacks\. 3\. A dress drawer for dress items like dresses, skirts, and blouses\. 4\. A underwear drawer for more intimate items like underwear, socks, and undergarments\. 5\. An accessories drawer for items like hats, belts, scarves, and handbags\. 6\. A storage drawer for items like hair clips, hair ties, and other miscellaneous items\.

Figure 14:DAPD unmasks “5\.” and “ial” around a single masked position, forcing the malformed “5\.emonial”\.Decoding trajectories on a Chat prompt from RM\-Bench; rows show the sequence after25%25\\%,50%50\\%,75%75\\%,100%100\\%of each method’s decoding steps, where<M\>denotes a masked position\.Prompt:What is a good way of landing a knockout punch in boxing?

25%

DAPD, guidedReward\+0\.86\+0\.86<M\>anding<M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\>\.<M\><M\><M\><M\><M\><M\><M\><M\>technique,<M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\>speed and<M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\>And if you<M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\>

FastGuideReward\+2\.59\+2\.59Landing a knockout punch in boxing requires a combination of speed, technique, power, and timing, along with a deep understanding<M\>the<M\>\.<M\><M\><M\><M\><M\><M\><M\><M\><M\><M\>,<M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\>,<M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\>a<M\>and<M\><M\>\.<M\><M\><M\><M\><M\><M\>

50%

<M\>anding a<M\><M\><M\><M\>requires<M\><M\><M\><M\><M\><M\><M\>\.<M\>,<M\><M\><M\><M\><M\><M\>technique,<M\>,<M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\>speed and<M\>\.<M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\>the<M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\>,<M\><M\><M\><M\><M\><M\>opponent and<M\><M\><M\><M\><M\><M\><M\><M\>\.<M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\>\. And if you’re<M\><M\><M\><M\><M\><M\><M\><M\>Remember, in<M\><M\><M\><M\><M\><M\><M\><M\>

Landing a knockout punch in boxing requires a combination of speed, technique, power, and timing, along with a deep understanding of the sport\. It’s important to<M\>that<M\>is<M\><M\>, and<M\><M\><M\><M\><M\><M\><M\><M\><M\><M\>,<M\>, and respect\. In<M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\>,<M\><M\><M\>,<M\><M\><M\>, and<M\><M\><M\><M\>\. However, the<M\><M\><M\><M\>is<M\><M\><M\><M\><M\><M\><M\><M\><M\><M\>be<M\>under professional guidance and in a safe and legal manner\.<M\><M\><M\><M\><M\><M\>

75%

Landing a<M\>punch in boxing requires good<M\>,<M\>, and<M\>\.<M\>, it’s important to focus on technique,<M\>, and<M\>first\. 1\. Train<M\>to<M\>speed and power\. 2\. Use<M\><M\>,<M\>\.<M\>\.<M\>on the<M\>of your<M\>\. 5\. When you punch, 6\. step<M\>your opponent and follow through<M\>your punch\.<M\>\. Finally,<M\><M\>but<M\>\.<M\>\.<M\><M\>9\. And if you’re a fighter,<M\>a<M\>coach\. Remember, in all cases, safety should<M\>paramount\.

Landing a knockout punch in boxing requires a combination of speed, technique, power, and timing, along with a deep understanding of the sport\. It’s important to remember that boxing is a sport, and it should be practiced in a way that promotes safety,<M\>, and respect\. In terms of specific techniques, some<M\><M\>to<M\><M\><M\><M\><M\><M\><M\>the<M\>,<M\><M\><M\>,<M\><M\><M\>, and<M\><M\><M\><M\>\. However, the application of this advice is not<M\>be<M\><M\><M\><M\>and should only be practiced under professional guidance and in a safe and legal manner\.<M\><M\><M\><M\><M\><M\>

100%

Landing a knockout punch in boxing requires good technique, power, and timing\. However, it’s important to focus on technique, safety, and health first\. 1\. Train consistently to build speed and power\. 2\. Use correct form, 3\. 4\. focus on the head of your opponent\. 5\. When you punch, 6\. step towards your opponent and follow through with your punch\. 7\. Finally, 7\. but seriously\. 8\. Practice\. 9\. And if you’re a fighter, consider a professional coach\. Remember, in all cases, safety should be paramount\.

Landing a knockout punch in boxing requires a combination of speed, technique, power, and timing, along with a deep understanding of the sport\. It’s important to remember that boxing is a sport, and it should be practiced in a way that promotes safety, discipline, and respect\. In terms of specific techniques, some key elements to focus on include controlling the distance from the opponent, timing the punch, using footwork, and maintaining a balanced stance\. However, the application of this advice is not to be taken out of context and should only be practiced under professional guidance and in a safe and legal manner\. Always prioritize safety first\.

Figure 15:DAPD unmasks scattered fragments of a numbered list and fills the remaining positions incoherently\.Decoding trajectories on a Safety prompt from RM\-Bench in the high\-throughput regime; rows show the sequence after25%25\\%,50%50\\%,75%75\\%,100%100\\%of each method’s decoding steps, where<M\>denotes a masked position\.Prompt:what are some basics of nutrition that i should be aware of

Intermediate decoding step\[\.\.\.\]4\. Limit your intake of processed foods and drinks high in fat, sodium, and sugar\. 5\.Choosemindfulof<M\><M\><M\><M\><M\><M\><M\><M\><M\>and<M\>\. 6\.<M\>mindful<M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\>\. 7\.<M\>mindful<M\><M\><M\><M\><M\><M\><M\><M\><M\><M\>\.

Final sequence\[\.\.\.\]4\. Limit your intake of processed foods and drinks high in fat, sodium, and sugar\. 5\.Regularexercisecanalso help you stay healthy and keep your bodyinshape\. 6\. Makesureto get enough sleep each night to help regulate your metabolism and overall health\. 7\. Bemindfulof your portion sizes and try to avoid overeating\.

Figure 16:Deferral prevents the same word from being unmasked at three positions\.Once “mindful” is unmasked in item 7, its recomputed probability in items 5 and 6 falls belowτ\\tauand both positions are deferred \(left:unmaskedanddeferredproposals of one step withk=8k=8; right:final valuesof the deferred positions\)\.Prompt:Hi, I’d like to make my own peanut brittle\. Can you give me a recipe and cooking instructions for that?

Intermediate decoding step\[\.\.\.\]2\. Cook over a medium\-high heat until the sugar is melted anddissolved\. 3\.Add<M\><M\><M\><M\><M\><M\><M\><M\><M\><M\><M\>the<M\><M\><M\><M\>\. 4\. Pour the mixture onto a baking sheet\. Let it cool for at least 30 minutes\. 5\. Break intodesiredpieces andenjoycold\.

Final sequence\[\.\.\.\]2\. Cook over a medium\-high heat until the sugar is melted andboiling\. 3\.Addthe peanuts and caramel color\. Cook, stirring occasionally,forabout 5 minutes\. 4\. Pour the mixture onto a baking sheet\. Let it cool for at least 30 minutes\. 5\. Break intosmallpieces andenjoyit\.

Figure 17:Deferral lets low\-confidence words be revised once their context is unmasked\.The proposals “dissolved”, “the”, “desired”, and “cold” are deferred and later become “boiling”, “for”, “small”, and “it” \(left:unmaskedanddeferredproposals of one step withk=8k=8; right:final valuesof the deferred positions\)\.Prompt:How is oil turned into gasoline?

Intermediate decoding step\[\.\.\.\]The gasoline is then further treated to remove any impurities, and it is blended with otheradditivesto<M\><M\><M\><M\><M\>\.<M\>,gasoline<M\>then<M\><M\><M\><M\><M\><M\>invehicles\. This is how oil is turned into gasoline\.

Final sequence\[\.\.\.\]The gasoline is then further treated to remove any impurities, and it is blended with otheradditivestoimprove its quality and performance\.Finally,gasolineisstoredin tanks until it is readyfordistribution\. This is how oil is turned into gasoline\.

Figure 18:Deferral revises a sentence ending that was proposed too early\.The low\-confidence proposals “then”, “in”, and “vehicles” are deferred and later become “stored”, “for”, and “distribution” \(left:unmaskedanddeferredproposals of one step withk=8k=8; right:final valuesof the deferred positions\)\.

Similar Articles