Optimizing Denoising Trajectories in dLLMs: A Lightweight Evolutionary Heuristic Approach

arXiv cs.CL Papers

Summary

This paper proposes a lightweight evolutionary heuristic scheduler to optimize denoising trajectories in diffusion large language models, addressing failure modes like EOS Overflow and Proximal Bias, and outperforming baselines on reasoning and planning benchmarks.

arXiv:2609.26052v1 Announce Type: new Abstract: Diffusion Large Language Models (dLLMs) have recently emerged as a promising alternative to conventional Auto-Regressive (AR) Large Language Models (LLMs). By leveraging bidirectional attention and parallel decoding, dLLMs enable more efficient generation. However, they require a carefully designed denoising scheduler at inference time (absent during training) whose choice significantly impacts generation quality. While confidence-based heuristic schedulers have shown strong empirical performance, they suffer from two critical failure modes: EOS Overflow and Proximal Bias. Through in-depth analysis of the Transformer's attention patterns, we reveal that these failures stem from certain positions assigning disproportionately high attention weights to invalid tokens (e.g., [MASK] and [EOS]), which produce misleading confidence signals. Building on this insight, empirical evidence shows that valid attention scores can provide complementary guidance to conventional confidence-based heuristics, yet no single metric consistently excels across all scenarios, implying that the optimal denoising trajectory is highly context-dependent. To address this problem, we propose a lightweight evolutionary heuristic scheduler optimized using the Covariance Matrix Adaptation Evolution Strategy (CMA-ES). Our scheduler dynamically integrates multiple heuristic features with a contextual mean-field embedding, while requiring only 393 trainable parameters. Evaluated on LLaDA and Dream across four reasoning and planning benchmarks, our method consistently outperforms strong baselines, including conventional heuristics, block auto-regressive methods, and recent State-Of-The-Art (SOTA) approaches. To the best of our knowledge, it represents the most parameter-efficient neural scheduler to date. Our code is available at https://github.com/RS2002/Evo-Denoise .
Original Article
View Cached Full Text

Cached at: 09/23/26, 09:22 AM

# A Lightweight Evolutionary Heuristic Approach
Source: [https://arxiv.org/html/2609.26052](https://arxiv.org/html/2609.26052)
## Optimizing Denoising Trajectories in dLLMs: A Lightweight Evolutionary Heuristic Approach

Zijian Zhao1,2, Dian Jin3, Xialiang Tong2, Sen Li1,4, Mingxuan Yuan2

###### Abstract

Diffusion Large Language Models \(dLLMs\) have recently emerged as a promising alternative to conventional Auto\-Regressive \(AR\) Large Language Models \(LLMs\)\. By leveraging bidirectional attention and parallel decoding, dLLMs enable more efficient generation\. However, they require a carefully designed denoising scheduler at inference time \(absent during training\) whose choice significantly impacts generation quality\. While confidence\-based heuristic schedulers have shown strong empirical performance, they suffer from two critical failure modes: EOS Overflow and Proximal Bias\. Through in\-depth analysis of the Transformer’s attention patterns, we reveal that these failures stem from certain positions assigning disproportionately high attention weights to invalid tokens \(e\.g\., \[MASK\] and \[EOS\]\), which produce misleading confidence signals\. Building on this insight, empirical evidence shows that valid attention scores can provide complementary guidance to conventional confidence\-based heuristics, yet no single metric consistently excels across all scenarios, implying that the optimal denoising trajectory is highly context\-dependent\. To address this problem, we propose a lightweight evolutionary heuristic scheduler optimized using the Covariance Matrix Adaptation Evolution Strategy \(CMA\-ES\)\. Our scheduler dynamically integrates multiple heuristic features with a contextual mean\-field embedding, while requiring only 393 trainable parameters\. Evaluated on LLaDA and Dream across four reasoning and planning benchmarks, our method consistently outperforms strong baselines, including conventional heuristics, block auto\-regressive methods, and recent State\-Of\-The\-Art \(SOTA\) approaches\. To the best of our knowledge, it represents the most parameter\-efficient neural scheduler to date\. Our code is available athttps://github\.com/RS2002/Evo\-Denoiser\.

## Introduction

Auto\-Regressive \(AR\) Large Language Models \(LLMs\)\(Achiamet al\.[2023](https://arxiv.org/html/2609.26052#bib.bib1); Touvronet al\.[2023](https://arxiv.org/html/2609.26052#bib.bib3); Liuet al\.[2024](https://arxiv.org/html/2609.26052#bib.bib2)\)have achieved remarkable success across diverse domains\(Chenet al\.[2026a](https://arxiv.org/html/2609.26052#bib.bib4)\)\. However, their inherently sequential generation paradigm suffers from slow decoding speed and error accumulation\(Arbuzovet al\.[2025](https://arxiv.org/html/2609.26052#bib.bib6)\)\. In recent years, Diffusion Large Language Models \(dLLMs\)\(Nieet al\.[2026](https://arxiv.org/html/2609.26052#bib.bib5); Yeet al\.[2025b](https://arxiv.org/html/2609.26052#bib.bib7)\)have emerged as a compelling alternative\. By leveraging bidirectional attention and parallel decoding, dLLMs capture full contextual information at every step and enable efficient generation\. Recent studies have further demonstrated that discrete dLLMs exhibit scaling laws\(Nieet al\.[2025](https://arxiv.org/html/2609.26052#bib.bib8)\)and modality expansion capabilities\(Youet al\.[2026](https://arxiv.org/html/2609.26052#bib.bib10); Zhuet al\.[2025](https://arxiv.org/html/2609.26052#bib.bib9)\)comparable to those of AR LLMs\.

Despite these advantages, dLLMs face a critical train\-inference mismatch\. During training, dLLMs randomly mask tokens and learn to recover them in parallel under a maximum likelihood Evidence Lower Bound \(ELBO\) objective\(Nieet al\.[2026](https://arxiv.org/html/2609.26052#bib.bib5)\)\. At inference, generation starts from a fully masked sequence and proceeds through progressive denoising steps\. Since the optimal denoising order is never explicitly supervised, the design of the denoising scheduler becomes crucial\. Prior work has shown that the choice of scheduling strategy significantly affects final generation quality\(Huanget al\.[2026a](https://arxiv.org/html/2609.26052#bib.bib11); Heet al\.[2025](https://arxiv.org/html/2609.26052#bib.bib12); Tanget al\.[2026](https://arxiv.org/html/2609.26052#bib.bib23)\)\.

Inspired by confidence\-based metrics in AR LLMs \(e\.g\., top\-1 probability, entropy, and Gini impurity\)\(Chenet al\.[2025](https://arxiv.org/html/2609.26052#bib.bib13),[2026c](https://arxiv.org/html/2609.26052#bib.bib14); Kanget al\.[2026](https://arxiv.org/html/2609.26052#bib.bib15)\), recent studies have adopted similar heuristics to guide token denoising in dLLMs\(Nieet al\.[2026](https://arxiv.org/html/2609.26052#bib.bib5); Ben\-Hamuet al\.[2026](https://arxiv.org/html/2609.26052#bib.bib25)\)\. Although selecting the top\-kkmost confident tokens substantially outperforms random ordering, confidence\-based schedulers still suffer from two persistent failure modes that notably degrade performance, especially in long\-form reasoning and planning tasks:

- •EOS Overflow\(Kimet al\.[2026b](https://arxiv.org/html/2609.26052#bib.bib19),[a](https://arxiv.org/html/2609.26052#bib.bib22); Parket al\.[2026](https://arxiv.org/html/2609.26052#bib.bib44)\): Excessive EOS tokens accumulate in the rightmost part of the sequence, particularly when they are denoised prematurely\.
- •Proximal \(Local\) Bias\(Kimet al\.[2026b](https://arxiv.org/html/2609.26052#bib.bib19); Piskorzet al\.[2026](https://arxiv.org/html/2609.26052#bib.bib24)\): Once a token is denoised, its neighboring positions receive disproportionately high confidence scores\.

To address these issues, a growing body of work has focused on improved denoising schedulers, which can be broadly categorized into three types: \(i\) manually designed heuristics that are computationally efficient and training\-free, yet whose optimality is difficult to verify\(Caoet al\.[2026](https://arxiv.org/html/2609.26052#bib.bib28); Parket al\.[2026](https://arxiv.org/html/2609.26052#bib.bib44)\); \(ii\) Block\-AR \(Semi\-AR\) methods that perform parallel denoising within blocks but sequential decoding across blocks\(Wuet al\.[2025](https://arxiv.org/html/2609.26052#bib.bib26); Zhanget al\.[2026](https://arxiv.org/html/2609.26052#bib.bib27)\), inherently limiting both inference speed and generation quality; and \(iii\) trainable schedulers based on SFT or RL that learn an auxiliary network\(Honget al\.[2026](https://arxiv.org/html/2609.26052#bib.bib20); Kimet al\.[2026b](https://arxiv.org/html/2609.26052#bib.bib19); Jazbecet al\.[2026](https://arxiv.org/html/2609.26052#bib.bib21); Huanget al\.[2026a](https://arxiv.org/html/2609.26052#bib.bib11); Heet al\.[2025](https://arxiv.org/html/2609.26052#bib.bib12)\), which are effective but incur substantial training costs \(e\.g\., up to 134M parameters\(Honget al\.[2026](https://arxiv.org/html/2609.26052#bib.bib20)\)\)\. Despite their diversity, these approaches either rely on fixed empirical rules that fail to adapt to varying contexts, or suffer from prohibitive training overhead\. This raises a fundamental question:*Is it possible to achieve adaptive, context\-aware denoising scheduling with minimal training cost?*

In this paper, we first conduct an in\-depth analysis of the Transformer’s attention mechanism, revealing that both EOS Overflow and Proximal Bias stem from certain positions excessively attending to invalid tokens such as \[MASK\] and \[EOS\]\. Building upon this insight, we demonstrate that valid attention scores serve as a strong complementary signal to conventional confidence\-based heuristics\. Nevertheless, extensive empirical evidence shows that no single heuristic, whether confidence\-based or attention\-based, consistently dominates, highlighting that the optimal denoising trajectory is highly context\-dependent\. To address this challenge, we propose a lightweight evolutionary heuristic scheduler optimized using the Covariance Matrix Adaptation Evolution Strategy \(CMA\-ES\)\(Hansen and Ostermeier[2001](https://arxiv.org/html/2609.26052#bib.bib16); Hansen[2016](https://arxiv.org/html/2609.26052#bib.bib17)\)\. Our scheduler dynamically integrates multiple heuristic features with a contextual mean\-field embedding, while requiring*only 393 trainable parameters*, making it, to the best of our knowledge, the most parameter\-efficient neural scheduler to date\. Evaluated on LLaDA\(Nieet al\.[2026](https://arxiv.org/html/2609.26052#bib.bib5)\)and Dream\(Yeet al\.[2025b](https://arxiv.org/html/2609.26052#bib.bib7)\)across four challenging reasoning and planning benchmarks, our method consistently outperforms strong baselines, including conventional heuristics, block\-AR methods, and recent State\-Of\-The\-Art \(SOTA\) approaches, delivering substantial improvements particularly in difficult settings with limited denoising budgets\.

## Preliminary of dLLMs

### Diffusion Language Models

dLLMs extend diffusion modeling to discrete text generation, offering a compelling alternative to conventional AR LLMs\. Unlike AR models that generate tokens sequentially from left to right, dLLMs begin with a fully masked sequence and iteratively denoise tokens in parallel using bidirectional attention\. This design enables full\-context modeling at every step and provides competitive performance, often with significantly faster inference for long sequences due to parallel decoding\. Representative models such as LLaDA\(Nieet al\.[2026](https://arxiv.org/html/2609.26052#bib.bib5)\)and Dream\(Yeet al\.[2025b](https://arxiv.org/html/2609.26052#bib.bib7)\)have demonstrated strong results while benefiting from inherent parallelism and bidirectional reasoning\.

The forward \(noising\) process gradually corrupts a clean sequence𝐱0\\mathbf\{x\}^\{0\}by replacing tokens with the special\[MASK\]token according to a noise schedule\. Letβt∈\(0,1\)\\beta\_\{t\}\\in\(0,1\)denote the instantaneous masking rate at timett\. The marginal probability that a token remains unmasked at timettis

αt=exp⁡\(−∫0tβs​𝑑s\),\\alpha\_\{t\}=\\exp\\left\(\-\\int\_\{0\}^\{t\}\\beta\_\{s\}\\,ds\\right\),\(1\)which monotonically decreases from11to0\. The per\-token transition kernel is defined as

q​\(xit∣xit−1\)=\{βtif​xit−1≠\[MASK\],1if​xit−1=\[MASK\],q\(x\_\{i\}^\{t\}\\mid x\_\{i\}^\{t\-1\}\)=\\begin\{cases\}\\beta\_\{t\}&\\text\{if \}x\_\{i\}^\{t\-1\}\\neq\\text\{\[MASK\]\},\\\\ 1&\\text\{if \}x\_\{i\}^\{t\-1\}=\\text\{\[MASK\]\},\\end\{cases\}\(2\)making the masked state absorbing\.

The reverse \(denoising\) process is parameterized by a Transformer networkpθp\_\{\\theta\}, which predicts the original tokens for masked positions\. The model is trained by maximizing the ELBO on the data likelihood, which reduces to the following weighted masked cross\-entropy objective:

ℒELBO=\\displaystyle\\mathcal\{L\}\_\{\\text\{ELBO\}\}=\(3\)𝔼t∼𝒰​\(0,1\),𝐱0,𝐱t∼q​\[\|α˙t\|1−αt​∑i:xit=\[MASK\]−log⁡pθ​\(xi0∣𝐱t\)\],\\displaystyle\\mathbb\{E\}\_\{t\\sim\\mathcal\{U\}\(0,1\),\\,\\mathbf\{x\}^\{0\},\\,\\mathbf\{x\}^\{t\}\\sim q\}\\left\[\\frac\{\|\\dot\{\\alpha\}\_\{t\}\|\}\{1\-\\alpha\_\{t\}\}\\sum\_\{i:x\_\{i\}^\{t\}=\\text\{\[MASK\]\}\}\-\\log p\_\{\\theta\}\(x\_\{i\}^\{0\}\\mid\\mathbf\{x\}^\{t\}\)\\right\],whereα˙t=d​αt/d​t\\dot\{\\alpha\}\_\{t\}=d\\alpha\_\{t\}/dt\.

### Denoising Scheduler Formulation

At inference, generation starts from a fully masked sequence𝐱T\\mathbf\{x\}^\{T\}and progressively recovers tokens over multiple steps\. The networkpθp\_\{\\theta\}estimates the clean data distributionpθ​\(𝐱0∣𝐱t\)p\_\{\\theta\}\(\\mathbf\{x\}^\{0\}\\mid\\mathbf\{x\}^\{t\}\)\. Given a noisy sample𝐱t\\mathbf\{x\}^\{t\}, we first draw a predicted clean sequence𝐱~0∼pθ\(⋅∣𝐱t\)\\tilde\{\\mathbf\{x\}\}^\{0\}\\sim p\_\{\\theta\}\(\\cdot\\mid\\mathbf\{x\}^\{t\}\), and then sample the previous state from the posterior:

pθ​\(𝐱t−1∣𝐱t\)=𝔼𝐱~0∼pθ\(⋅∣𝐱t\)​\[q​\(𝐱t−1∣𝐱t,𝐱~0\)\]\.p\_\{\\theta\}\(\\mathbf\{x\}\_\{t\-1\}\\mid\\mathbf\{x\}\_\{t\}\)=\\mathbb\{E\}\_\{\\tilde\{\\mathbf\{x\}\}^\{0\}\\sim p\_\{\\theta\}\(\\cdot\\mid\\mathbf\{x\}^\{t\}\)\}\\Bigl\[q\(\\mathbf\{x\}\_\{t\-1\}\\mid\\mathbf\{x\}\_\{t\},\\tilde\{\\mathbf\{x\}\}^\{0\}\)\\Bigr\]\.\(4\)However, computingpθ​\(𝐱t−1∣𝐱t\)p\_\{\\theta\}\(\\mathbf\{x\}\_\{t\-1\}\\mid\\mathbf\{x\}\_\{t\}\)directly is computationally intractable due to the expectation over the vocabulary space\. In practice, we mostly rely on a denoising scheduler to approximate this reverse step\. Importantly, the inference\-time denoising schedule \(e\.g\., the number of iterations and the re\-masking strategy\) is not explicitly supervised during training, as the ELBO objective jointly supervises all time steps\. Recent studies\(Tanget al\.[2026](https://arxiv.org/html/2609.26052#bib.bib23)\)have shown that the choice of this schedule significantly affects generation quality, making the design of effective denoising schedulers a critical challenge for dLLMs\. A detailed related works review is provided at Appendix\.

## Failure Analysis of Heuristic Schedulers

In this section, we analyze the top\-1 probability denoising scheduler as a representative example to illustrate the underlying causes of EOS Overflow and Proximal Bias through the lens of the Transformer’s attention mechanism\. We also provide both empirical and intuitive evidence demonstrating that conventional single\-token heuristic metrics are insufficient for determining the optimal denoising strategy\.

### Why Confidence Is Not Enough

We first formally define the heuristic scheduler for dLLM denoising\. At inference time, generation begins from a fully masked sequencexT=\[\[MASK\]\]Lx\_\{T\}=\[\\text\{\[MASK\]\}\]^\{L\}, whereLLis the sequence length \(the prompt precedingxTx\_\{T\}is omitted for simplicity\)\. At each denoising steptt, letℳt\\mathcal\{M\}\_\{t\}denote the set of positions that remain masked:

ℳt=\{i∣xt,i=\[MASK\]\}\.\\mathcal\{M\}\_\{t\}=\\\{i\\mid x\_\{t,i\}=\\text\{\[MASK\]\}\\\}\.\(5\)
The dLLM produces intermediate representationsfθ​\(xt\)\\text\{f\}\_\{\\theta\}\(x\_\{t\}\)\(e\.g\., hidden states or logits\), which are passed through a scorerh​\(⋅\)\\text\{h\}\(\\cdot\)to obtain denoising scores:

st,i=h​\(fθ​\(xt,i\)\),s\_\{t,i\}=\\text\{h\}\(\\text\{f\}\_\{\\theta\}\(x\_\{t,i\}\)\),\(6\)whereh​\(⋅\)\\text\{h\}\(\\cdot\)extracts features such as top\-1 probability, entropy, or margin probability\. The top\-kkhighest\-scoring masked tokens are then selected for denoising:

ℐt=Top​\-​K⁡\(st,i∣i∈ℳt\)\.\\mathcal\{I\}\_\{t\}=\\operatorname\{Top\\text\{\-\}K\}\\bigl\(s\_\{t,i\}\\mid i\\in\\mathcal\{M\}\_\{t\}\\bigr\)\.\(7\)
We use the top\-1 probability heuristic as a case study to uncover the root causes of EOS Overflow and Proximal Bias, which have not been thoroughly explained in prior work\. Since dLLMs are built upon the Transformer architecture, we analyze their behavior through attention patterns\. At steptt, for themm\-th attention head in thenn\-th layer, the attention matrix is:

Atn,m=Softmax​\(Qtn,m​\(Ktn,m\)⊤dk\),A^\{n,m\}\_\{t\}=\\mathrm\{Softmax\}\\left\(\\frac\{Q^\{n,m\}\_\{t\}\(K^\{n,m\}\_\{t\}\)^\{\\top\}\}\{\\sqrt\{d\_\{k\}\}\}\\right\),\(8\)whereQtn,mQ^\{n,m\}\_\{t\}andKtn,mK^\{n,m\}\_\{t\}are the query and key matrices, respectively, anddkd\_\{k\}is the head dimension\.

Unlike AR LLMs, dLLM sequences contain many \[MASK\] tokens and trailing \[EOS\] tokens, both of which often receive uninformative attention\. We therefore define the valid attention score for each token as:

at,in=1H​∑m=1H∑jAtn,m​\[i,j\]⋅𝟏​\{xt,j∈𝒱\},a^\{n\}\_\{t,i\}=\\frac\{1\}\{H\}\\sum\_\{m=1\}^\{H\}\\sum\_\{j\}A^\{n,m\}\_\{t\}\[i,j\]\\cdot\\mathbf\{1\}\\\{x\_\{t,j\}\\in\\mathcal\{V\}\\\},\(9\)where𝒱\\mathcal\{V\}is the set of valid \(non\-\[MASK\], non\-\[EOS\]\) vocabulary tokens andHHis the number of attention heads\.

![Refer to caption](https://arxiv.org/html/2609.26052v1/img/step0.png)\(a\)Step 0
![Refer to caption](https://arxiv.org/html/2609.26052v1/img/step3.png)\(b\)Step 3
![Refer to caption](https://arxiv.org/html/2609.26052v1/img/step56.png)\(c\)Step 56

Figure 1:Denoising process visualization on a GSM8K example using LLaDA\-8B\-Instruct\. Different background colors represent prompt \(blue\), \[MASK\] \(red\), and \[EOS\] \(gray\) regions\. The curves show top\-1 probability \(blue\), valid attention score \(red\), and EOS probability \(yellow\)\.Fig\.[1](https://arxiv.org/html/2609.26052#Sx3.F1)visualizes the denoising process on a GSM8K sample \(here, stepiicorresponds to time stepT−iT\-ifor readability\)\. The valid attention scores are computed from the middle layer as a representative example\. We observe the following:

- •Step 0 \(initial\):Both beginning and ending tokens exhibit high top\-1 probabilities\. While high confidence at the beginning is expected due to strong prompt context, high confidence at the end is surprising\. These ending tokens show low valid attention scores but high EOS prediction probabilities, indicating that they assign excessive attention to uninformative \[MASK\] tokens, resulting in misleadingly high confidence\.
- •Step 3:Several ending tokens have been decoded as \[EOS\]\. Neighboring tokens then exhibit a sharp increase in both top\-1 probability and EOS probability, which is a clear manifestation ofProximal Bias\. Their valid attention scores decrease compared to Step 0, suggesting increased attention to the newly decoded \(but uninformative\) \[EOS\] tokens\.
- •Step 56:From Step 3 to Step 56, only ending tokens continue to be decoded as \[EOS\], while beginning tokens remain largely unchanged\. This illustrates theEOS Overflowphenomenon, where \[EOS\] tokens occupy an excessive portion of the sequence\.

These observations reveal that both failure modes originate from the valid attention mechanism: ending tokens initially over\-attend to \[MASK\] tokens and decode prematurely as \[EOS\]; the newly generated \[EOS\] tokens then attract further attention, causing neighboring tokens to follow suit\. This cascading effect leads to severe Proximal Bias and EOS Overflow\. Crucially, these findings suggest that a purely confidence\-based scheduler is fundamentally limited: it cannot distinguish between genuine predictive certainty and attention\-driven false confidence, motivating our exploration of alternative signals\.

### Why We Need Context with Multiple Heuristics

![Refer to caption](https://arxiv.org/html/2609.26052v1/img/correct_dis.png)Figure 2:Distribution of correctly answered questions by different heuristics on GSM8K \(generation length 128, denoising budget 32\)\. Each bar shows the number of questions solved by individual or combined heuristics\.To further investigate the limitations of single\-heuristic schedulers, we evaluate several confidence\-based methods on GSM8K with a generation length of 128 and a denoising budget of 32\. The tested heuristics include top\-1 probability, entropy, margin probability, and their Block\-AR variants \(block size 32\)\. Motivated by our attention analysis, we also examine valid attention scores computed separately from upper, middle, and lower layers \(inspired by prior findings that different layers capture distinct syntactic and semantic information\(Vig[2019](https://arxiv.org/html/2609.26052#bib.bib38); Jawaharet al\.[2019](https://arxiv.org/html/2609.26052#bib.bib39)\)\)\.

Fig\.[2](https://arxiv.org/html/2609.26052#Sx3.F2)shows the number of correctly answered questions for each scheduler, with special emphasis on questions solved uniquely by one method\. Detailed results are provided in Table[1](https://arxiv.org/html/2609.26052#Sx4.T1)\. Our analysis yields several key insights:

- •Although individual scheduler accuracies range from 46\.2% to 66\.6% \(Table[1](https://arxiv.org/html/2609.26052#Sx4.T1)\), only 11\.5% of questions are answered incorrectly by*all*schedulers \(Fig\.[2](https://arxiv.org/html/2609.26052#Sx3.F2)\)\. This indicates that dLLMs possess significantly stronger problem\-solving capability than current schedulers can elicit\. In other words, if an oracle scheduler existed, the upper\-bound performance would be at least 88\.5%, which is comparable to or even exceeds some post\-training results via SFT or RL\(Zhaoet al\.[2026](https://arxiv.org/html/2609.26052#bib.bib36); Xieet al\.[2026](https://arxiv.org/html/2609.26052#bib.bib37)\)\. This highlights the critical importance of optimizing the denoising trajectory\.
- •Each heuristic exhibits unique strengths: certain questions are solved exclusively by one scheduler but not others\. This leads to two important conclusions: \(i\) no single manually designed heuristic is universally sufficient, and combining multiple complementary heuristics is required; \(ii\) the optimal denoising strategy is highly context\-dependent, as no single fixed heuristic is suitable for all questions\.

To further illustrate why context matters, consider the interplay between attention and confidence: a high valid attention score may indicate sufficient information flow, but alone it cannot determine whether the token is ready to be denoised\. However, when accompanied by a high top\-1 probability, this combined signal provides stronger evidence\. Moreover, the overall magnitude of attention scores varies with sequence length, as longer prompts naturally provide richer valid contexts\. This observation underscores the necessity of context\-aware scheduling: the decision to denoise a token cannot be made in isolation but must account for the global state of the sequence\.

## Methodology

![Refer to caption](https://arxiv.org/html/2609.26052v1/img/main.png)Figure 3:Network architecture of the proposed neural scorer\.Motivated by the insights from our failure analysis, we propose an evolutionary heuristic scheduler that learns to dynamically combine multiple heuristics for superior denoising quality\. Following the formulation in Eq\. \([6](https://arxiv.org/html/2609.26052#Sx3.E6)\), we parameterize the scorer ashΘ​\(⋅\)\\mathrm\{h\}\_\{\\Theta\}\(\\cdot\)\. Directly optimizing the scorer parameters via policy\-gradient methods \(e\.g\., PPO\(Schulmanet al\.[2017](https://arxiv.org/html/2609.26052#bib.bib40)\), GRPO\(Shaoet al\.[2024](https://arxiv.org/html/2609.26052#bib.bib34)\)\) is prohibitively difficult: the token selection operation in Eq\. \([7](https://arxiv.org/html/2609.26052#Sx3.E7)\) is discrete and non\-differentiable, and backpropagating through the entire dLLM is computationally infeasible\. We therefore reformulate scheduler optimization as a black\-box problem and adopt CMA\-ES\(Hansen and Ostermeier[2001](https://arxiv.org/html/2609.26052#bib.bib16); Hansen[2016](https://arxiv.org/html/2609.26052#bib.bib17)\), a zero\-order evolutionary strategy that only requires evaluating the final task accuracy as the fitness signal\. Although our method does not explicitly model long\-term future effects through the Bellman equation as in RL\-based schedulers\(Huanget al\.[2026b](https://arxiv.org/html/2609.26052#bib.bib18); Honget al\.[2026](https://arxiv.org/html/2609.26052#bib.bib20)\), it is not myopic: because the scorer is shared across all denoising steps, optimizing for final answer accuracy implicitly captures temporal dependencies among successive decisions\.

### Feature Construction

We formalize the input features for our evolutionary scorer\. The manual heuristic features are as follows:

- •Top\-1 Probability: The predicted probability of the most likely token: Top\-1 Prob=maxpθ\(⋅∣xt,i\)\.\\text\{Top\-1 Prob\}=\\max p\_\{\\theta\}\(\\cdot\\mid x\_\{t,i\}\)\.\(10\)
- •Probabilistic Margin: The difference between the top\-1 and top\-2 probabilities, which effectively quantifies prediction confidence: Prob Margin=maxpθ\(⋅∣xt,i\)−max\(2\)pθ\(⋅∣xt,i\),\\text\{Prob Margin\}=\\max p\_\{\\theta\}\(\\cdot\\mid x\_\{t,i\}\)\-\\max^\{\(2\)\}p\_\{\\theta\}\(\\cdot\\mid x\_\{t,i\}\),\(11\)wheremax\(2\)\\max^\{\(2\)\}denotes the second\-highest probability\.
- •Valid Attention Score: We extract attention scores from three representative layers, including layer0\(bottom\),⌊l/2⌋\\lfloor l/2\\rfloor\(middle\), andl−1l\-1\(top\), wherellis the total number of layers in the dLLM backbone\. This choice is motivated by prior findings that different layers capture distinct levels of syntactic and semantic information\(Vig[2019](https://arxiv.org/html/2609.26052#bib.bib38); Jawaharet al\.[2019](https://arxiv.org/html/2609.26052#bib.bib39)\)\.

Additionally, inspired by the success of Block\-AR methods, we incorporate positional information\. We include the relative positioniL\\frac\{i\}\{L\}and the relative position among currently masked tokens:

MASK\-Pos=∑j∈ℳt𝟏​\{j<i\}\|ℳt\|\.\\text\{MASK\-Pos\}=\\frac\{\\sum\_\{j\\in\\mathcal\{M\}\_\{t\}\}\\mathbf\{1\}\\\{j<i\\\}\}\{\|\\mathcal\{M\}\_\{t\}\|\}\.\(12\)
These features together form a 7\-dimensional vectorκt,i∈\[0,1\]7\\kappa\_\{t,i\}\\in\[0,1\]^\{7\}for each token at steptt\(1 for top\-1 probability, 1 for margin, 3 for attention from different layers, 1 for relative position, and 1 for mask position\)\. In our design, all features have the same range, which is beneficial for training\. While this design results from careful empirical selection, future work may explore additional heuristics such as entropy or alternative attention\-based signals\(Guoet al\.[2024](https://arxiv.org/html/2609.26052#bib.bib41)\)\.

### Network Architecture

The architecture of our neural scorer is illustrated in Fig\.[3](https://arxiv.org/html/2609.26052#Sx4.F3)\. It employs a dual\-path design: a linear skip\-connection path \(blue\) and a two\-layer Mean\-Field Multilayer Perceptron \(MF\-MLP\) path \(red\)\. The final score is computed as:

st,i=Linear​\(κt,i\)\+MF​\-​MLP​\(κt,i;κt\)\.s\_\{t,i\}=\\mathrm\{Linear\}\(\\kappa\_\{t,i\}\)\+\\mathrm\{MF\\text\{\-\}MLP\}\(\\kappa\_\{t,i\};\\kappa\_\{t\}\)\.\(13\)
This design is motivated by two key observations from our analysis\. First, since the optimal denoising strategy is context\-dependent, the scorer must incorporate global contextual information\. However, full Transformer\-based scorers\(Honget al\.[2026](https://arxiv.org/html/2609.26052#bib.bib20); Jazbecet al\.[2026](https://arxiv.org/html/2609.26052#bib.bib21)\)introduce too many parameters for effective zero\-order optimization\. Inspired by Mean\-Field Reinforcement Learning \(MFRL\)\(Yanget al\.[2018](https://arxiv.org/html/2609.26052#bib.bib42)\), we propose a lightweight MF\-MLP that uses the mean embedding of the first\-layer hidden representations as contextual input to the second layer\. We apply the mean\-field operation only after the first layer to minimize parameter count\. Second, a purely non\-linear architecture can slow down early\-stage optimization\. Since individual heuristics already perform reasonably well in many cases, we add a linear bypass path to enable fast initial progress via simple weighted combinations\.

Overall, this design reflects our goal of achieving a compact yet effective structure\. To capture the mean\-field contextual information, a single layer is insufficient, as an affine combination would be identical for all tokens\. Thus, we adopt a two\-layer MF\-MLP, with the mean\-field embedding introduced only at the second\-layer input to minimize the parameter overhead\. Additionally, the linear bypass layer facilitates efficient training under zero\-order optimization conditions\.

We now formalize the MF\-MLP structure\. First, a linear layer with ReLU activation and layer normalization produces per\-token embeddings:

yt,i=Linear​\(κt,i\)∈ℝh,y\_\{t,i\}=\\mathrm\{Linear\}\(\\kappa\_\{t,i\}\)\\in\\mathbb\{R\}^\{h\},\(14\)wherehhis the hidden dimension\. The mean\-field embedding over all masked tokens is then computed:

y¯t=1\|ℳt\|​∑j∈ℳtyt,j∈ℝh\.\\overline\{y\}\_\{t\}=\\frac\{1\}\{\|\\mathcal\{M\}\_\{t\}\|\}\\sum\_\{j\\in\\mathcal\{M\}\_\{t\}\}y\_\{t,j\}\\in\\mathbb\{R\}^\{h\}\.\(15\)Finally, the second linear layer produces the contextual score:

zt,i=Linear​\(\[yt,i;y¯t\]\)∈ℝ\.z\_\{t,i\}=\\mathrm\{Linear\}\(\[y\_\{t,i\};\\overline\{y\}\_\{t\}\]\)\\in\\mathbb\{R\}\.\(16\)The final scorest,is\_\{t,i\}is the sum ofzt,iz\_\{t,i\}and the affine bypass term\.

### Training Process

We employ CMA\-ES with task\-specific adaptations for training\. The fitness function is defined as the average accuracy over a mini\-batch:

f​\(Θ\)=1B​∑q∈ℬ𝟏​\{correct​\(q;hΘ\)\},f\(\\Theta\)=\\frac\{1\}\{B\}\\sum\_\{q\\in\\mathcal\{B\}\}\\mathbf\{1\}\\\{\\text\{correct\}\(q;\\mathrm\{h\}\_\{\\Theta\}\)\\\},\(17\)whereℬ\\mathcal\{B\}is a mini\-batch of sizeBBsampled from the training set\. The batch is fixed within each generation but resampled across generations to reduce overfitting\. The complete training procedure is summarized in Algorithm[1](https://arxiv.org/html/2609.26052#alg1)at Appendix\. Since zero\-order optimization itself is not the primary contribution of this work, we refer readers to\(Hansen and Ostermeier[2001](https://arxiv.org/html/2609.26052#bib.bib16); Hansen[2016](https://arxiv.org/html/2609.26052#bib.bib17)\)for detailed CMA\-ES mechanics\.

Table 1:Model Performance on LLaDA\-8B\-Instruct\(Nieet al\.[2026](https://arxiv.org/html/2609.26052#bib.bib5)\)\. In the table, ‘T’ denotes the step budget, ‘L’ denotes the sequence length, ‘B’ denotes the block size, and ‘Attn \(u/m/l\)’ indicates that the attention score is computed from the upper, middle, or lower layer, respectively\.Boldindicates the best performance andunderlinedenotes the second best one\. These conventions apply to all subsequent tables\. The results reported for EDM are reproduced from the original paper\(Kimet al\.[2026b](https://arxiv.org/html/2609.26052#bib.bib19)\)\.GSM8KMathTypeSchedulerL=128L=256L=128L=256T=16T=32T=16T=32T=16T=32T=16T=32HeuristicsTop\-1 Prob43\.254\.040\.447\.517\.821\.013\.218\.6Entropy35\.946\.243\.244\.017\.213\.012\.215\.2Prob Margin50\.254\.945\.650\.019\.820\.017\.219\.8Block\-AR \(B=32\)Top\-1 Prob42\.660\.59\.846\.811\.823\.04\.817\.4Entropy32\.857\.06\.735\.111\.819\.24\.011\.2Prob Margin47\.666\.612\.852\.515\.224\.05\.416\.8Recent WorksEDM \(5M params\)–56\.8–––22\.8––CCD43\.552\.537\.048\.718\.219\.412\.217\.4AGDO Scheduler43\.454\.440\.347\.417\.621\.413\.219\.0Suffix Anchor54\.956\.745\.750\.918\.021\.816\.819\.2ProposedAttn \(u\)32\.553\.59\.633\.25\.818\.06\.411\.8Attn \(m\)38\.259\.140\.250\.810\.217\.812\.417\.2Attn \(l\)23\.654\.25\.630\.54\.617\.02\.67\.2Evolution \(393 params\)58\.567\.652\.660\.821\.427\.225\.027\.8CountdownStrategyQATypeSchedulerL=128L=256L=128L=256T=16T=32T=16T=32T=16T=32T=16T=32HeuristicsTop\-1 Prob39\.846\.03\.520\.763\.065\.936\.856\.8Entropy39\.540\.67\.427\.062\.765\.533\.551\.4Prob Margin40\.348\.46\.618\.863\.265\.941\.657\.9Block\-AR \(B=32\)Top\-1 Prob24\.236\.34\.79\.055\.062\.926\.942\.4Entropy16\.828\.13\.59\.847\.661\.627\.444\.0Prob Margin26\.639\.83\.511\.353\.162\.623\.040\.6Recent WorksEDM \(5M params\)–43\.8––––––CCD37\.943\.04\.717\.663\.966\.430\.157\.2AGDO Scheduler40\.645\.75\.520\.762\.365\.637\.353\.6Suffix Anchor40\.642\.233\.245\.757\.660\.458\.159\.1ProposedAttn \(u\)17\.221\.51\.69\.844\.855\.722\.731\.7Attn \(m\)14\.825\.46\.67\.855\.661\.925\.649\.3Attn \(l\)13\.323\.00\.43\.957\.959\.728\.948\.9Evolution \(393 params\)50\.053\.935\.240\.665\.565\.954\.166\.5

## Experiments

### Experiment Setup

To validate the efficiency and generalization capacity of the proposed method, we evaluate our scheduler on two popular dLLMs: LLaDA\-8B\-Instruct\(Nieet al\.[2026](https://arxiv.org/html/2609.26052#bib.bib5)\)and Dream\-7B\-Instruct\(Yeet al\.[2025b](https://arxiv.org/html/2609.26052#bib.bib7)\)\. We train the scheduler and evaluate it on four reasoning and planning tasks spanning mathematics and logic: GSM8K\(Cobbeet al\.[2021](https://arxiv.org/html/2609.26052#bib.bib45)\), Math\(Hendryckset al\.[2021](https://arxiv.org/html/2609.26052#bib.bib46)\), Countdown\(Panet al\.[2025](https://arxiv.org/html/2609.26052#bib.bib47)\), and StrategyQA\(Gevaet al\.[2021](https://arxiv.org/html/2609.26052#bib.bib51)\)\. For all datasets, we follow the standard training and test splits\. For Math, we use the Math\-500 subset as the test set for efficient evaluation\. Detailed introductions to these datasets are provided in Appendix[Dataset Introduction](https://arxiv.org/html/2609.26052#A0.SSx3)\.

To demonstrate the superiority of our method, we compare against several classical and competitive fixed\-budget baselines\. In addition to conventional heuristic and Block\-AR schedulers, we include the following recent approaches:\(i\) Early Decision Matters \(EDM, 5M parameters\)\(Kimet al\.[2026b](https://arxiv.org/html/2609.26052#bib.bib19)\), which trains a scorer to guide the initial denoising steps\. Since its dataset construction is nontrivial, we adopt the same evaluation protocol and report the results from the original paper\(Kimet al\.[2026b](https://arxiv.org/html/2609.26052#bib.bib19)\)for fair comparison;\(ii\) CCD\(Chenet al\.[2026b](https://arxiv.org/html/2609.26052#bib.bib31)\), which leverages historical average confidence scores for more consistent guidance;\(iii\) AGDO Scheduler\(Denget al\.[2026](https://arxiv.org/html/2609.26052#bib.bib43)\), which first selects candidates based on last\-layer valid attention scores and then determines which tokens to denoise using Top\-1 probability; and\(iv\) Suffix Anchor\(Parket al\.[2026](https://arxiv.org/html/2609.26052#bib.bib44)\), which appends a fixed prompt at the end of the generation region to mitigate EOS overflow\. We exclude two neural schedulers from comparison:Jazbecet al\.\([2026](https://arxiv.org/html/2609.26052#bib.bib21)\)adopts adaptive denoising speed, making it incompatible with our fixed\-budget setting; and UPO\(Honget al\.[2026](https://arxiv.org/html/2609.26052#bib.bib20)\)supports only one token per step, which is mismatched with our fast\-denoising protocol\.

To ensure robustness, we evaluate under multiple configurations with generation lengthL∈\{128,256\}L\\in\\\{128,256\\\}and denoising budgetT∈\{16,32\}T\\in\\\{16,32\\\}\. As noted in prior work\(Ben\-Hamuet al\.[2026](https://arxiv.org/html/2609.26052#bib.bib25); Jazbecet al\.[2026](https://arxiv.org/html/2609.26052#bib.bib21); Luxembourget al\.[2026](https://arxiv.org/html/2609.26052#bib.bib50)\), performance gaps between schedulers diminish with large budgets, and pure AR decoding can be competitive\(Niet al\.[2026](https://arxiv.org/html/2609.26052#bib.bib52)\)\. We therefore focus on the more challenging fast\-denoising regime, which better aligns with the parallel decoding strength of dLLMs\. For CMA\-ES optimization, we use a batch size of 100, a maximum of 10 generations, and an initial step size of 0\.2, with all other parameters set to default values in the pycma package\(Hansenet al\.[2019](https://arxiv.org/html/2609.26052#bib.bib49)\)\.

### Experiment Results

The results on LLaDA\-8B\-Instruct are presented in Table[1](https://arxiv.org/html/2609.26052#Sx4.T1), while those on Dream\-7B\-Instruct are deferred to Appendix[Expanded Experiment Results](https://arxiv.org/html/2609.26052#A0.SSx4), as they lead to similar conclusions\. Overall, our proposed evolutionary heuristic scheduler demonstrates consistently strong performance, outperforming all baselines in the majority of settings\. The advantage of our method becomes particularly pronounced on more challenging tasks and under tighter budgets\. For instance, on MATH with the most challenging configuration \(T=16T=16,L=256L=256, i\.e\., 16 tokens denoised per step\), our scheduler achieves 25\.0% accuracy, substantially outperforming the best baseline at 17\.2%\. On Countdown, our method is the only one that consistently exceeds 35% accuracy across settings, while most baselines remain below 10% whenT=16T=16andL=256L=256\.

Moreover, we observe that attention\-based heuristics exhibit inconsistent performance\. Middle\-layer attention scores generally outperform those from upper or lower layers, aligning with recent findings that the middle layers of Transformers serve as a ”global workspace” for reasoning\(Gurneeet al\.[2026](https://arxiv.org/html/2609.26052#bib.bib53)\)\. However, valid attention scores alone are far from sufficient, as they frequently underperform simple confidence\-based heuristics\. This observation reinforces our earlier claim \(Fig\.[2](https://arxiv.org/html/2609.26052#Sx3.F2)\) that attention signals offer complementary rather than competitive value: they excel in certain niche scenarios but require combination with other heuristics to achieve robust performance\.

### Ablation and Generalization Study

Table 2:Ablation and Generalization Experiment of LLaDA\-8B\-Instruct\(Nieet al\.[2026](https://arxiv.org/html/2609.26052#bib.bib5)\)in Math\(Hendryckset al\.[2021](https://arxiv.org/html/2609.26052#bib.bib46)\)SchedulerL=128L=256T=16T=32T=16T=32Proposed21\.427\.225\.027\.8w/o MF21\.025\.819\.624\.0w/o skip21\.424\.020\.023\.4GSM8K22\.624\.820\.624\.6Countdown14\.018\.020\.415\.8StrategyQA14\.222\.018\.216\.6Dream\-7B\-Instruct20\.423\.422\.021\.2L=128––21\.424\.4L=25620\.624\.2––T=16–24\.4–24\.6T=3220\.0–22\.6–In this section, we take LLaDA\-8B\-Instruct\(Nieet al\.[2026](https://arxiv.org/html/2609.26052#bib.bib5)\)and the Math dataset\(Hendryckset al\.[2021](https://arxiv.org/html/2609.26052#bib.bib46)\)as a case study to conduct ablation and generalization analyses\. The core design of our method lies in the network architecture; accordingly, we evaluate the impact of removing the mean\-field embedding and the skip linear layer\. To assess generalization capacity, we test the scheduler trained on different source datasets \(GSM8K, Countdown, and StrategyQA\), different backbone models \(Dream\-7B\-Instruct\), and different inference settings \(generation lengthLLand denoising budgetTT\)\. Since exhaustively evaluating all combinations across these dimensions would be computationally prohibitive, we select the most challenging dataset, Math, to make the conclusions more compelling and representative\. Additionally, a comparison between methods under higher decoding budgets is provided in Appendix[Expanded Experiment Results](https://arxiv.org/html/2609.26052#A0.SSx4)\.

The experimental results are presented in Table[2](https://arxiv.org/html/2609.26052#Sx5.T2)\. For the ablation study, we observe that removing either the mean\-field embedding \(w/o MF\) or the skip linear layer \(w/o skip\) leads to performance degradation to varying extents\. Notably, the degradation patterns differ across configurations\. Eliminating mean\-field embedding suffers more severely when the sequence length is larger, aligning with our expectation that longer sequences require stronger global contextual aggregation, which the mean\-field embedding provides\. Furthermore, the skip linear layer provides effective support for zero\-order optimization training by reducing the difficulty of optimizing a purely non\-linear structure\.

Regarding the generalization experiments, transferring the scheduler to different datasets leads to varied degradation\. The GSM8K\-trained scheduler performs best among cross\-dataset transfers, even slightly outperforming the Math\-trained one underT=16,L=128T=16,L=128, which is expected given the distributional similarity between the two math reasoning datasets\. In contrast, schedulers trained on Countdown or StrategyQA suffer more substantial drops, suggesting that the scheduler learns domain\-specific priors rather than purely universal heuristics\. Nevertheless, since the scheduler inputs consist of effective manually designed heuristics, its performance remains at least comparable to those heuristic\-based schedulers, avoiding complete failure even under domain shifts\. When transferring to a different backbone \(Dream\-7B\-Instruct\) or unseen inference configurations \(length and budget\), we observe milder degradation, with results remaining highly competitive against prior work \(Table[1](https://arxiv.org/html/2609.26052#Sx4.T1)\)\. Notably, the degradation from cross\-dataset shifts is consistently larger than that from cross\-configuration shifts, indicating that the scheduler captures a combination of domain\-specific scheduling priors and universal heuristics\. This suggests that mixed\-dataset training could be a promising direction to improve robustness, which we leave for future work\.

## Conclusion

In this paper, we analyze the failure modes of confidence\-based heuristic schedulers in dLLMs\. Through attention\-based analysis, we reveal that EOS Overflow and Proximal Bias arise from excessive attention to invalid tokens such as\[MASK\]and\[EOS\]\. Building on this insight, we show that valid attention scores serve as a strong complementary signal, yet no single heuristic suffices due to the highly context\-dependent nature of optimal denoising trajectories\. To address this challenge, we propose a lightweight evolutionary heuristic scheduler optimized via CMA\-ES\. By dynamically integrating multiple heuristics with a mean\-field contextual embedding, our method achieves strong performance using only 393 trainable parameters, serving as the most parameter\-efficient neural scheduler for dLLMs to date\. Experiments across four challenging reasoning and planning tasks on LLaDA and Dream demonstrate consistent gains over strong baselines, especially in difficult settings with limited denoising budgets\. While the learned scheduler exhibits some domain specificity, this opens promising avenues for future work on mixed\-dataset training and adaptive denoising speed designs to further improve generalization and efficiency\.

## References

- J\. Achiam, S\. Adler, S\. Agarwal, L\. Ahmad, I\. Akkaya, F\. L\. Aleman, D\. Almeida, J\. Altenschmidt, S\. Altman, S\. Anadkat,et al\.\(2023\)Gpt\-4 technical report\.arXiv preprint arXiv:2303\.08774\.Cited by:[Introduction](https://arxiv.org/html/2609.26052#Sx1.p1.1)\.
- M\. L\. Arbuzov, S\. Bei, Z\. Dong, D\. Kalaev, and A\. A\. Shvets \(2025\)Beyond exponential decay: rethinking error accumulation in large language models\.arXiv preprint arXiv:2505\.24187\.Cited by:[Introduction](https://arxiv.org/html/2609.26052#Sx1.p1.1)\.
- M\. Arriola, A\. Gokaslan, J\. Chiu, Z\. Yang, Z\. Qi, J\. Han, S\. Sahoo, and V\. Kuleshov \(2025\)Block diffusion: interpolating between autoregressive and diffusion language models\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 50726–50753\.Cited by:[Literature Review](https://arxiv.org/html/2609.26052#A0.SSx1.p2.1)\.
- H\. Ben\-Hamu, I\. Gat, D\. Severo, N\. S\. Nolte, and B\. Karrer \(2026\)Accelerated sampling from masked diffusion models via entropy bounded unmasking\.Advances in Neural Information Processing Systems38,pp\. 55981–56007\.Cited by:[Literature Review](https://arxiv.org/html/2609.26052#A0.SSx1.p1.1),[Introduction](https://arxiv.org/html/2609.26052#Sx1.p3.1),[Experiment Setup](https://arxiv.org/html/2609.26052#Sx5.SSx1.p3.2)\.
- M\. Cao, A\. H\. Correia, C\. Louizos, S\. Liu, and L\. Yin \(2026\)Search or accelerate: confidence\-switched position beam search for diffusion language models\.International Conference on Machine Learning\.Cited by:[Literature Review](https://arxiv.org/html/2609.26052#A0.SSx1.p1.1),[Introduction](https://arxiv.org/html/2609.26052#Sx1.p4.1)\.
- H\. Chen, H\. Chen, Z\. Zhao, K\. Han, G\. Zhu, Y\. Zhao, Y\. Du, W\. Xu, and Q\. Shi \(2026a\)An overview of domain\-specific foundation model: key technologies, applications and challenges\.Science China Information Sciences69\(1\),pp\. 111301\.Cited by:[Introduction](https://arxiv.org/html/2609.26052#Sx1.p1.1)\.
- K\. Chen, Z\. Liu, X\. Tao, H\. Liu, X\. Fu, S\. Zhang, D\. Tu, L\. Kong, R\. Liu, and H\. Li \(2026b\)Beyond confidence: adaptive and coherent decoding for diffusion language models\.International Conference on Learning Representations\.Cited by:[Literature Review](https://arxiv.org/html/2609.26052#A0.SSx1.p1.1),[Literature Review](https://arxiv.org/html/2609.26052#A0.SSx1.p3.1),[Experiment Setup](https://arxiv.org/html/2609.26052#Sx5.SSx1.p2.1)\.
- S\. Chen, Z\. Zhao, and J\. Chen \(2026c\)Confident RAG: enhancing the performance of LLMs for mathematics question answering through multi\-embedding and confidence scoring\.InICLR 2026 Workshop on Logical Reasoning of Large Language Models,Cited by:[Introduction](https://arxiv.org/html/2609.26052#Sx1.p3.1)\.
- T\. Chen, J\. Chen, Z\. Zhao, H\. Chen, L\. Zhang, and G\. Zhu \(2025\)First token probability guided rag for telecom question answering\.arXiv preprint arXiv:2501\.06468\.Cited by:[Introduction](https://arxiv.org/html/2609.26052#Sx1.p3.1)\.
- K\. Cobbe, V\. Kosaraju, M\. Bavarian, M\. Chen, H\. Jun, L\. Kaiser, M\. Plappert, J\. Tworek, J\. Hilton, R\. Nakano,et al\.\(2021\)Training verifiers to solve math word problems\.arXiv preprint arXiv:2110\.14168\.Cited by:[1st item](https://arxiv.org/html/2609.26052#A0.I5.i1.p1.1.1),[Experiment Setup](https://arxiv.org/html/2609.26052#Sx5.SSx1.p1.1)\.
- J\. Deng, J\. Li, W\. X\. Zhao, J\. Wang, H\. Lu, and J\. Wen \(2026\)Beyond fully random masking: attention\-guided denoising and optimization for diffusion language models\.arXiv preprint arXiv:2606\.12273\.Cited by:[Experiment Setup](https://arxiv.org/html/2609.26052#Sx5.SSx1.p2.1)\.
- M\. Geva, D\. Khashabi, E\. Segal, T\. Khot, D\. Roth, and J\. Berant \(2021\)Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies\.Transactions of the Association for Computational Linguistics9,pp\. 346–361\.Cited by:[4th item](https://arxiv.org/html/2609.26052#A0.I5.i4.p1.1.1),[Experiment Setup](https://arxiv.org/html/2609.26052#Sx5.SSx1.p1.1)\.
- Z\. Guo, H\. Kamigaito, and T\. Watanabe \(2024\)Attention score is not all you need for token importance indicator in kv cache reduction: value also matters\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,pp\. 21158–21166\.Cited by:[Feature Construction](https://arxiv.org/html/2609.26052#Sx4.SSx1.p4.2)\.
- W\. Gurnee, N\. Sofroniew, A\. Pearce, M\. Piotrowski, I\. Kauvar, R\. Chen, A\. Soligo, P\. Bogdan, E\. Ong, R\. Wang, B\. Thompson, D\. Abrahams, S\. Kantamneni, E\. Ameisen, J\. Batson, and J\. Lindsey \(2026\)Verbalizable representations form a global workspace in language models\.Transformer Circuits Thread, Anthropic\.Note:Accessed: 2026External Links:[Link](https://transformer-circuits.pub/2026/workspace/index.html)Cited by:[Experiment Results](https://arxiv.org/html/2609.26052#Sx5.SSx2.p2.1)\.
- N\. Hansen, Y\. Akimoto, and P\. Baudis \(2019\)CMA\-ES/pycma on Github\.Note:Zenodo, DOI:10\.5281/zenodo\.2559634External Links:[Document](https://dx.doi.org/10.5281/zenodo.2559634),[Link](https://doi.org/10.5281/zenodo.2559634)Cited by:[Experiment Setup](https://arxiv.org/html/2609.26052#Sx5.SSx1.p3.2)\.
- N\. Hansen and A\. Ostermeier \(2001\)Completely derandomized self\-adaptation in evolution strategies\.Evolutionary computation9\(2\),pp\. 159–195\.Cited by:[Introduction](https://arxiv.org/html/2609.26052#Sx1.p5.1),[Training Process](https://arxiv.org/html/2609.26052#Sx4.SSx3.p1.2),[Methodology](https://arxiv.org/html/2609.26052#Sx4.p1.1)\.
- N\. Hansen \(2016\)The cma evolution strategy: a tutorial\.arXiv preprint arXiv:1604\.00772\.Cited by:[Introduction](https://arxiv.org/html/2609.26052#Sx1.p5.1),[Training Process](https://arxiv.org/html/2609.26052#Sx4.SSx3.p1.2),[Methodology](https://arxiv.org/html/2609.26052#Sx4.p1.1)\.
- H\. He, K\. Renz, Y\. Cao, and A\. Geiger \(2025\)Mdpo: overcoming the training\-inference divide of masked diffusion language models\.arXiv preprint arXiv:2508\.13148\.Cited by:[Introduction](https://arxiv.org/html/2609.26052#Sx1.p2.1),[Introduction](https://arxiv.org/html/2609.26052#Sx1.p4.1)\.
- D\. Hendrycks, C\. Burns, S\. Kadavath, A\. Arora, S\. Basart, E\. Tang, D\. Song, and J\. Steinhardt \(2021\)Measuring mathematical problem solving with the math dataset\.InThirty\-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track \(Round 2\),Cited by:[2nd item](https://arxiv.org/html/2609.26052#A0.I5.i2.p1.1.1),[Experiment Setup](https://arxiv.org/html/2609.26052#Sx5.SSx1.p1.1),[Ablation and Generalization Study](https://arxiv.org/html/2609.26052#Sx5.SSx3.p1.2),[Table 2](https://arxiv.org/html/2609.26052#Sx5.T2)\.
- C\. Hong, S\. An, M\. Kim, and J\. C\. Ye \(2026\)Improving discrete diffusion unmasking policies beyond explicit reference policies\.InThe Fourteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=on6cb46OhD)Cited by:[Literature Review](https://arxiv.org/html/2609.26052#A0.SSx1.p3.1),[Introduction](https://arxiv.org/html/2609.26052#Sx1.p4.1),[Network Architecture](https://arxiv.org/html/2609.26052#Sx4.SSx2.p2.1),[Methodology](https://arxiv.org/html/2609.26052#Sx4.p1.1),[Experiment Setup](https://arxiv.org/html/2609.26052#Sx5.SSx1.p2.1)\.
- Z\. Huang, Z\. Chen, Z\. Wang, T\. Li, and G\. Qi \(2026a\)Reinforcing the diffusion chain of lateral thought with diffusion language models\.Advances in Neural Information Processing Systems38,pp\. 152677–152710\.Cited by:[Literature Review](https://arxiv.org/html/2609.26052#A0.SSx1.p3.1),[Introduction](https://arxiv.org/html/2609.26052#Sx1.p2.1),[Introduction](https://arxiv.org/html/2609.26052#Sx1.p4.1)\.
- Z\. Huang, Y\. Wang, Z\. Chen, and G\. Qi \(2026b\)Don’t settle too early: self\-reflective remasking for diffusion language models\.InThe Fourteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=BsZeTuB5fD)Cited by:[Literature Review](https://arxiv.org/html/2609.26052#A0.SSx1.p3.1),[Methodology](https://arxiv.org/html/2609.26052#Sx4.p1.1)\.
- G\. Jawahar, B\. Sagot, and D\. Seddah \(2019\)What does bert learn about the structure of language?\.InProceedings of the 57th annual meeting of the association for computational linguistics,pp\. 3651–3657\.Cited by:[Why We Need Context with Multiple Heuristics](https://arxiv.org/html/2609.26052#Sx3.SSx2.p1.1),[3rd item](https://arxiv.org/html/2609.26052#Sx4.I4.i3.p1.4)\.
- M\. Jazbec, T\. X\. Olausson, L\. Béthune, P\. Ablin, M\. Kirchhof, J\. Monteiro, V\. Turrisi, J\. Ramapuram, and M\. Cuturi \(2026\)Learning unmasking policies for diffusion language models\.International Conference on Machine Learning\.Cited by:[Literature Review](https://arxiv.org/html/2609.26052#A0.SSx1.p3.1),[Introduction](https://arxiv.org/html/2609.26052#Sx1.p4.1),[Network Architecture](https://arxiv.org/html/2609.26052#Sx4.SSx2.p2.1),[Experiment Setup](https://arxiv.org/html/2609.26052#Sx5.SSx1.p2.1),[Experiment Setup](https://arxiv.org/html/2609.26052#Sx5.SSx1.p3.2)\.
- Z\. Kang, X\. Zhao, and D\. Song \(2026\)Scalable best\-of\-n selection for large language models via self\-certainty\.Advances in neural information processing systems38,pp\. 19720–19745\.Cited by:[Literature Review](https://arxiv.org/html/2609.26052#A0.SSx1.p1.1),[Introduction](https://arxiv.org/html/2609.26052#Sx1.p3.1)\.
- B\. Kim, D\. Jeon, D\. Kim, W\. Jeung, and A\. No \(2026a\)Rainbow padding: mitigating early termination in instruction\-tuned diffusion LLMs\.InThe Fourteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=cznTlh7Msz)Cited by:[1st item](https://arxiv.org/html/2609.26052#Sx1.I1.i1.p1.1)\.
- J\. Kim, K\. Shah, V\. Kontonis, S\. M\. Kakade, and S\. Chen \(2025\)Train for the worst, plan for the best: understanding token ordering in masked diffusions\.InInternational Conference on Machine Learning,pp\. 30749–30768\.Cited by:[Literature Review](https://arxiv.org/html/2609.26052#A0.SSx1.p1.1)\.
- J\. Kim, S\. Choi, Y\. Jo, M\. Lee, and M\. Seo \(2026b\)Early decisions matter: proximity bias and initial trajectory shaping in non\-autoregressive diffusion language models\.International Conference on Machine Learning\.Cited by:[3rd item](https://arxiv.org/html/2609.26052#A0.I5.i3.p1.1),[Literature Review](https://arxiv.org/html/2609.26052#A0.SSx1.p2.1),[Literature Review](https://arxiv.org/html/2609.26052#A0.SSx1.p3.1),[1st item](https://arxiv.org/html/2609.26052#Sx1.I1.i1.p1.1),[2nd item](https://arxiv.org/html/2609.26052#Sx1.I1.i2.p1.1),[Introduction](https://arxiv.org/html/2609.26052#Sx1.p4.1),[Table 1](https://arxiv.org/html/2609.26052#Sx4.T1),[Experiment Setup](https://arxiv.org/html/2609.26052#Sx5.SSx1.p2.1)\.
- A\. Liu, B\. Feng, B\. Xue, B\. Wang, B\. Wu, C\. Lu, C\. Zhao, C\. Deng, C\. Zhang, C\. Ruan,et al\.\(2024\)Deepseek\-v3 technical report\.arXiv preprint arXiv:2412\.19437\.Cited by:[Introduction](https://arxiv.org/html/2609.26052#Sx1.p1.1)\.
- G\. Lu, H\. M\. Chen, Y\. Karashima, Z\. Wang, D\. Fujiki, and H\. Fan \(2026\)AdaBlock\-dLLM: semantic\-aware diffusion LLM inference via adaptive block size\.InThe Fourteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=0Cv9PwL7cI)Cited by:[Literature Review](https://arxiv.org/html/2609.26052#A0.SSx1.p2.1)\.
- O\. Luxembourg, H\. Permuter, and E\. Nachmani \(2026\)Plan for speed: dilated scheduling for masked diffusion language models\.International Conference on Machine Learning\.Cited by:[Experiment Setup](https://arxiv.org/html/2609.26052#Sx5.SSx1.p3.2)\.
- Z\. Ni, S\. Wang, Y\. Yue, T\. Yu, W\. Zhao, Y\. Hua, T\. Chen, J\. Song, C\. Yu, B\. Zheng,et al\.\(2026\)The flexibility trap: rethinking the value of arbitrary order in diffusion language models\.International Conference on Learning Representations\.Cited by:[Experiment Setup](https://arxiv.org/html/2609.26052#Sx5.SSx1.p3.2)\.
- S\. Nie, F\. Zhu, C\. Du, T\. Pang, Q\. Liu, G\. Zeng, M\. Lin, and C\. Li \(2025\)Scaling up masked diffusion models on text\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 82974–82997\.Cited by:[Introduction](https://arxiv.org/html/2609.26052#Sx1.p1.1)\.
- S\. Nie, F\. Zhu, Z\. You, X\. Zhang, J\. Ou, J\. Hu, J\. Zhou, Y\. Lin, J\. Wen, and C\. Li \(2026\)Large language diffusion models\.Advances in Neural Information Processing Systems38,pp\. 50608–50646\.Cited by:[Literature Review](https://arxiv.org/html/2609.26052#A0.SSx1.p1.1),[Introduction](https://arxiv.org/html/2609.26052#Sx1.p1.1),[Introduction](https://arxiv.org/html/2609.26052#Sx1.p2.1),[Introduction](https://arxiv.org/html/2609.26052#Sx1.p3.1),[Introduction](https://arxiv.org/html/2609.26052#Sx1.p5.1),[Diffusion Language Models](https://arxiv.org/html/2609.26052#Sx2.SSx1.p1.1),[Table 1](https://arxiv.org/html/2609.26052#Sx4.T1),[Experiment Setup](https://arxiv.org/html/2609.26052#Sx5.SSx1.p1.1),[Ablation and Generalization Study](https://arxiv.org/html/2609.26052#Sx5.SSx3.p1.2),[Table 2](https://arxiv.org/html/2609.26052#Sx5.T2)\.
- J\. Pan, J\. Zhang, X\. Wang, L\. Yuan, H\. Peng, and A\. Suhr \(2025\)TinyZero\.Note:https://github\.com/Jiayi\-Pan/TinyZeroAccessed: 2025\-01\-24Cited by:[3rd item](https://arxiv.org/html/2609.26052#A0.I5.i3.p1.1.1),[Experiment Setup](https://arxiv.org/html/2609.26052#Sx5.SSx1.p1.1)\.
- J\. Park, J\. Kim, J\. Ko, N\. Kwak, and W\. Rhee \(2026\)When confidence misleads: suffix anchoring and anchor\-proximity confidence modulation for diffusion language models\.arXiv preprint arXiv:2605\.28181\.Cited by:[Literature Review](https://arxiv.org/html/2609.26052#A0.SSx1.p1.1),[1st item](https://arxiv.org/html/2609.26052#Sx1.I1.i1.p1.1),[Introduction](https://arxiv.org/html/2609.26052#Sx1.p4.1),[Experiment Setup](https://arxiv.org/html/2609.26052#Sx5.SSx1.p2.1)\.
- J\. Piskorz, C\. Pinneri, A\. Correia, M\. Alfarra, R\. Garrepalli, and C\. Louizos \(2026\)Masks can be distracting: on context comprehension in diffusion language models\.InWorkshop on Scientific Methods for Understanding Deep Learning,External Links:[Link](https://openreview.net/forum?id=y6Nvum4WwO)Cited by:[2nd item](https://arxiv.org/html/2609.26052#Sx1.I1.i2.p1.1)\.
- J\. Schulman, F\. Wolski, P\. Dhariwal, A\. Radford, and O\. Klimov \(2017\)Proximal policy optimization algorithms\.arXiv preprint arXiv:1707\.06347\.Cited by:[Methodology](https://arxiv.org/html/2609.26052#Sx4.p1.1)\.
- Z\. Shao, P\. Wang, Q\. Zhu, R\. Xu, J\. Song, X\. Bi, H\. Zhang, M\. Zhang, Y\. Li, Y\. Wu,et al\.\(2024\)Deepseekmath: pushing the limits of mathematical reasoning in open language models\.arXiv preprint arXiv:2402\.03300\.Cited by:[Literature Review](https://arxiv.org/html/2609.26052#A0.SSx1.p3.1),[Methodology](https://arxiv.org/html/2609.26052#Sx4.p1.1)\.
- L\. Tang, L\. Yu, S\. Zhang, and G\. V\. Steeg \(2026\)Is your diffusion sampler actually correct? a sampler\-centric evaluation of discrete diffusion language models\.International Conference on Machine Learning\.Cited by:[Introduction](https://arxiv.org/html/2609.26052#Sx1.p2.1),[Denoising Scheduler Formulation](https://arxiv.org/html/2609.26052#Sx2.SSx2.p1.6)\.
- H\. Touvron, L\. Martin, K\. Stone, P\. Albert, A\. Almahairi, Y\. Babaei, N\. Bashlykov, S\. Batra, P\. Bhargava, S\. Bhosale,et al\.\(2023\)Llama 2: open foundation and fine\-tuned chat models\.arXiv preprint arXiv:2307\.09288\.Cited by:[Introduction](https://arxiv.org/html/2609.26052#Sx1.p1.1)\.
- J\. Vig \(2019\)A multiscale visualization of attention in the transformer model\.InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics: System Demonstrations,Florence, Italy,pp\. 37–42\.External Links:[Link](https://www.aclweb.org/anthology/P19-3007),[Document](https://dx.doi.org/10.18653/v1/P19-3007)Cited by:[Why We Need Context with Multiple Heuristics](https://arxiv.org/html/2609.26052#Sx3.SSx2.p1.1),[3rd item](https://arxiv.org/html/2609.26052#Sx4.I4.i3.p1.4)\.
- C\. Wu, H\. Zhang, S\. Xue, Z\. Liu, S\. Diao, L\. Zhu, P\. Luo, S\. Han, and E\. Xie \(2025\)Fast\-dllm: training\-free acceleration of diffusion llm by enabling kv cache and parallel decoding\.arXiv preprint arXiv:2505\.22618\.Cited by:[Literature Review](https://arxiv.org/html/2609.26052#A0.SSx1.p1.1),[Introduction](https://arxiv.org/html/2609.26052#Sx1.p4.1)\.
- S\. Xie, L\. Kong, X\. Song, X\. Dong, G\. Chen, E\. Xing, and K\. Zhang \(2026\)Advancing reasoning in diffusion language models with denoising process rewards\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 42703–42720\.Cited by:[1st item](https://arxiv.org/html/2609.26052#Sx3.I3.i1.p1.1)\.
- Y\. Yang, R\. Luo, M\. Li, M\. Zhou, W\. Zhang, and J\. Wang \(2018\)Mean field multi\-agent reinforcement learning\.InInternational conference on machine learning,pp\. 5571–5580\.Cited by:[Network Architecture](https://arxiv.org/html/2609.26052#Sx4.SSx2.p2.1)\.
- J\. Ye, J\. Gao, S\. Gong, L\. Zheng, X\. Jiang, Z\. Li, and L\. Kong \(2025a\)Beyond autoregression: discrete diffusion for complex reasoning and planning\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 77875–77898\.Cited by:[Literature Review](https://arxiv.org/html/2609.26052#A0.SSx1.p2.1)\.
- J\. Ye, Z\. Xie, L\. Zheng, J\. Gao, Z\. Wu, X\. Jiang, Z\. Li, and L\. Kong \(2025b\)Dream 7b: diffusion large language models\.arXiv preprint arXiv:2508\.15487\.Cited by:[Table 3](https://arxiv.org/html/2609.26052#A0.T3),[Introduction](https://arxiv.org/html/2609.26052#Sx1.p1.1),[Introduction](https://arxiv.org/html/2609.26052#Sx1.p5.1),[Diffusion Language Models](https://arxiv.org/html/2609.26052#Sx2.SSx1.p1.1),[Experiment Setup](https://arxiv.org/html/2609.26052#Sx5.SSx1.p1.1)\.
- Z\. You, S\. Nie, X\. Zhang, J\. ZHOU, Z\. Lu, J\. Wen, and C\. Li \(2026\)Llada\-v: large language diffusion models with visual instruction tuning\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 10093–10105\.Cited by:[Introduction](https://arxiv.org/html/2609.26052#Sx1.p1.1)\.
- Y\. Zhang, X\. Li, J\. Zhou, H\. Ma, Z\. Wan, Y\. Shi, D\. Miao, Q\. Zhang, and L\. Cao \(2026\)Swordsman: entropy\-driven adaptive block partition for efficient diffusion language models\.International Conference on Machine Learning\.Cited by:[Literature Review](https://arxiv.org/html/2609.26052#A0.SSx1.p2.1),[Introduction](https://arxiv.org/html/2609.26052#Sx1.p4.1)\.
- S\. Zhao, D\. Gupta, Q\. Zheng, and A\. Grover \(2026\)D1: scaling reasoning in diffusion large language models via reinforcement learning\.Advances in Neural Information Processing Systems38,pp\. 56729–56762\.Cited by:[1st item](https://arxiv.org/html/2609.26052#Sx3.I3.i1.p1.1)\.
- K\. Zheng, Y\. Chen, H\. Mao, M\. Liu, J\. Zhu, and Q\. Zhang \(2025\)Masked diffusion models are secretly time\-agnostic masked models and exploit inaccurate categorical sampling\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 63186–63227\.Cited by:[Literature Review](https://arxiv.org/html/2609.26052#A0.SSx1.p1.1)\.
- F\. Zhu, Z\. You, Y\. Xing, Z\. Huang, L\. Liu, Y\. Zhuang, G\. Lu, K\. Wang, X\. Wang, L\. Wei,et al\.\(2025\)Llada\-moe: a sparse moe diffusion language model\.arXiv preprint arXiv:2509\.24389\.Cited by:[Introduction](https://arxiv.org/html/2609.26052#Sx1.p1.1)\.

## Appendix

### Literature Review

Early work, inspired by the concept of confidence in AR LLMs\(Kanget al\.[2026](https://arxiv.org/html/2609.26052#bib.bib15)\), designs denoising schedulers under the assumption that tokens with higher confidence are more likely to be correctly predicted, and that revealing them first can provide reliable guidance for subsequent tokens\. Metrics such as top\-1 probability\(Wuet al\.[2025](https://arxiv.org/html/2609.26052#bib.bib26)\), entropy\(Ben\-Hamuet al\.[2026](https://arxiv.org/html/2609.26052#bib.bib25)\), and margin probability\(Kimet al\.[2025](https://arxiv.org/html/2609.26052#bib.bib29)\)have been introduced and shown to outperform random denoising\(Zhenget al\.[2025](https://arxiv.org/html/2609.26052#bib.bib30)\)\. Building on these, subsequent methods propose further empirical improvements\. For instance,\(Chenet al\.[2026b](https://arxiv.org/html/2609.26052#bib.bib31)\)observed inconsistency in individual token predictions and proposed CCD, which uses historical average confidence as a more stable denoising metric\.\(Caoet al\.[2026](https://arxiv.org/html/2609.26052#bib.bib28)\)noted that average confidence over the entire denoising process correlates with final accuracy and introduced PBS, which applies beam search to retain trajectories with the highest cumulative confidence\. To address the EOS Overflow issue,\(Nieet al\.[2026](https://arxiv.org/html/2609.26052#bib.bib5)\)and\(Parket al\.[2026](https://arxiv.org/html/2609.26052#bib.bib44)\)proposed directly suppressing \[EOS\] tokens or down\-weighting confidence near sentence endings\. However, these approaches may result in overly long outputs and lack flexibility\.

To prevent \[EOS\] tokens from appearing too early and interfering with generation, Block\-AR has emerged as an alternative paradigm\(Arriolaet al\.[2025](https://arxiv.org/html/2609.26052#bib.bib32)\)\. By allowing arbitrary diffusion within blocks while decoding blocks auto\-regressively, Block\-AR achieves better contextual consistency, yet introduces new challenges\. For example,\(Luet al\.[2026](https://arxiv.org/html/2609.26052#bib.bib33)\)and\(Zhanget al\.[2026](https://arxiv.org/html/2609.26052#bib.bib27)\)observed that block size is a critical hyper\-parameter affecting generation quality, and more importantly, splitting a sentence or semantic paragraph across blocks can cause significant performance degradation\. To address this,\(Luet al\.[2026](https://arxiv.org/html/2609.26052#bib.bib33)\)and\(Zhanget al\.[2026](https://arxiv.org/html/2609.26052#bib.bib27)\)proposed AdaBlock and Swordsman, respectively, which adaptively determine block boundaries based on confidence\-related metrics, following empirical heuristics\. Moreover,\(Kimet al\.[2026b](https://arxiv.org/html/2609.26052#bib.bib19); Yeet al\.[2025a](https://arxiv.org/html/2609.26052#bib.bib35)\)also argue that reintroducing AR constraints in dLLMs could limit their potential for complex reasoning and planning tasks\.

As noted by\(Chenet al\.[2026b](https://arxiv.org/html/2609.26052#bib.bib31)\), most of the above methods rely on empirical heuristics with limited theoretical guarantees\. As a result, another line of work aims to train a neural scheduler jointly with the dLLM backbone to learn optimal denoising trajectories by RL or from offline data\. For instance,\(Huanget al\.[2026a](https://arxiv.org/html/2609.26052#bib.bib11)\)proposed LLaDOU, which first formulates denoising as a Markov Decision Process \(MDP\) and employs the Plackett–Luce model to model token selection probabilities at each denoising step\. They further introduced RemeDi, incorporating a pre\-training warmup with a remasking mechanism\(Huanget al\.[2026b](https://arxiv.org/html/2609.26052#bib.bib18)\)to improve performance\. To reduce training costs, recent efforts focus on training the scheduler while keeping the dLLM backbone frozen, which also achieves competitive results\. For example,\(Honget al\.[2026](https://arxiv.org/html/2609.26052#bib.bib20)\)and\(Jazbecet al\.[2026](https://arxiv.org/html/2609.26052#bib.bib21)\)both adopt GRPO\(Shaoet al\.[2024](https://arxiv.org/html/2609.26052#bib.bib34)\)to train a Transformer\-based scheduler, taking as input either the hidden features from the dLLM or the prediction probabilities of each token, respectively\. Meanwhile,\(Kimet al\.[2026b](https://arxiv.org/html/2609.26052#bib.bib19)\)train a Transformer to predict whether an initial trajectory is likely to lead to a correct answer using offline data, and then use it to score and guide early inference steps\. Overall, these methods require Transformers with parameter counts ranging from 300K to 134M, incurring additional training and inference costs\. In contrast, our proposed scheduler is a carefully designed network with only 393 parameters, making it, to the best of our knowledge, the most lightweight neural scheduler to date\.

### Training Algorithm

The detailed training process of our neural scorer is provided at Algorithm[1](https://arxiv.org/html/2609.26052#alg1)\.

Algorithm 1CMA\-ES Training for the Neural Scorer1:Model

hΘ​\(⋅\)\\mathrm\{h\}\_\{\\Theta\}\(\\cdot\)with parameters

Θ∈ℝD\\Theta\\in\\mathbb\{R\}^\{D\}; training set

𝒟train\\mathcal\{D\}\_\{\\text\{train\}\}; population size

λ\\lambda; initial step size

σ0\\sigma\_\{0\}; batch size

BB; maximum generations

GmaxG\_\{\\max\}\.

2:Optimized parameters

Θ∗\\Theta^\{\*\}\.

3:Initialize

Θ0\\Theta\_\{0\}randomly \(e\.g\., Xavier uniform\)\.

4:Initialize CMA\-ES parameters: mean

𝐦←Θ0\\mathbf\{m\}\\leftarrow\\Theta\_\{0\}, step size

σ←σ0\\sigma\\leftarrow\\sigma\_\{0\}, covariance

𝐂←𝐈\\mathbf\{C\}\\leftarrow\\mathbf\{I\}\.

5:

g←0g\\leftarrow 0\.

6:while

g<Gmaxg<G\_\{\\max\}and not convergeddo

7:Sample

λ\\lambdacandidate solutions

\{Θi\}i=1λ∼𝒩​\(𝐦,σ2​𝐂\)\\\{\\Theta\_\{i\}\\\}\_\{i=1\}^\{\\lambda\}\\sim\\mathcal\{N\}\(\\mathbf\{m\},\\sigma^\{2\}\\mathbf\{C\}\)\.

8:Sample a fixed mini\-batch

ℬ⊂𝒟train\\mathcal\{B\}\\subset\\mathcal\{D\}\_\{\\text\{train\}\}of size

BB\.

9:for

i=1i=1to

λ\\lambdado

10:Assign

Θi\\Theta\_\{i\}to scorer

hΘi\\mathrm\{h\}\_\{\\Theta\_\{i\}\}\.

11:Generate answers for all questions in

ℬ\\mathcal\{B\}using

hΘi\\mathrm\{h\}\_\{\\Theta\_\{i\}\}\.

12:Compute fitness

fi←1B​∑q∈ℬ𝟏​\{correct​\(q\)\}f\_\{i\}\\leftarrow\\frac\{1\}\{B\}\\sum\_\{q\\in\\mathcal\{B\}\}\\mathbf\{1\}\\\{\\text\{correct\}\(q\)\\\}\.

13:endfor

14:Sort candidates by fitness:

f\(1\)≥⋯≥f\(λ\)f\_\{\(1\)\}\\geq\\cdots\\geq f\_\{\(\\lambda\)\}\.

15:Select top

μ\\mucandidates \(

μ=λ/2\\mu=\\lambda/2\)\.

16:Update mean:

𝐦←∑i=1μwi​Θ\(i\)\\mathbf\{m\}\\leftarrow\\sum\_\{i=1\}^\{\\mu\}w\_\{i\}\\Theta\_\{\(i\)\}, with weights

wi=log⁡\(μ\+1\)−log⁡\(i\)∑j=1μ\(log⁡\(μ\+1\)−log⁡\(j\)\)w\_\{i\}=\\frac\{\\log\(\\mu\+1\)\-\\log\(i\)\}\{\\sum\_\{j=1\}^\{\\mu\}\(\\log\(\\mu\+1\)\-\\log\(j\)\)\}\.

17:Update covariance matrix

𝐂\\mathbf\{C\}and step size

σ\\sigmausing standard CMA\-ES rules\.

18:

g←g\+1g\\leftarrow g\+1\.

19:endwhile

20:return

Θ∗←𝐦\\Theta^\{\*\}\\leftarrow\\mathbf\{m\}\.

Table 3:Model Performance on Dream\-7B\-Instruct\(Yeet al\.[2025b](https://arxiv.org/html/2609.26052#bib.bib7)\)\. For Dream, which is based on an AR LLM backbone, the attention scores are shifted by one position to align with the token prediction order\. Regarding EDM, the original paper does not report results for Countdown or StrategyQA on Dream\.GSM8KMathTypeSchedulerL=128L=256L=128L=256T=16T=32T=16T=32T=16T=32T=16T=32HeuristicsTop\-1 Prob37\.442\.822\.045\.815\.819\.410\.818\.2Entropy26\.238\.916\.139\.713\.419\.05\.817\.8Prob Margin38\.345\.426\.844\.614\.616\.811\.817\.6Block\-AR \(B=32\)Top\-1 Prob23\.056\.62\.116\.57\.823\.81\.05\.0Entropy16\.850\.92\.210\.18\.219\.41\.86\.8Prob Margin28\.858\.82\.020\.811\.624\.42\.26\.8Recent WorksEDM \(5M params\)–––52\.4–––21\.4CCD32\.342\.818\.638\.013\.621\.46\.817\.4AGDO Scheduler36\.541\.620\.239\.916\.019\.811\.017\.6Suffix Anchor42\.346\.428\.437\.019\.422\.217\.020\.2ProposedAttn \(u\)16\.748\.11\.612\.64\.619\.42\.06\.8Attn \(m\)37\.442\.830\.741\.311\.715\.412\.014\.8Attn \(l\)9\.546\.02\.05\.12\.414\.42\.02\.0Evolution \(393 params\)43\.558\.937\.560\.327\.227\.015\.222\.6CountdownStrategyQATypeSchedulerL=128L=256L=128L=256T=16T=32T=16T=32T=16T=32T=16T=32HeuristicsTop\-1 Prob31\.336\.716\.031\.646\.967\.433\.645\.3Entropy22\.734\.018\.429\.744\.060\.326\.242\.6Prob Margin35\.237\.112\.529\.751\.167\.840\.651\.7Block\-AR \(B=32\)Top\-1 Prob30\.529\.30\.017\.249\.168\.514\.443\.5Entropy18\.436\.70\.08\.244\.858\.225\.944\.1Prob Margin27\.030\.50\.021\.951\.169\.111\.144\.5Recent WorksCCD25\.434\.412\.125\.642\.162\.029\.042\.4AGDO Scheduler29\.340\.218\.428\.545\.958\.734\.246\.0Suffix Anchor14\.533\.62\.710\.961\.764\.553\.761\.3ProposedAttn \(u\)5\.517\.20\.42\.3450\.958\.929\.553\.9Attn \(m\)10\.915\.20\.04\.358\.160\.161\.462\.2Attn \(l\)3\.934\.05\.58\.722\.722\.923\.626\.6Evolution \(393 params\)47\.348\.419\.933\.663\.868\.665\.665\.1

### Dataset Introduction

- •GSM8K\(Cobbeet al\.[2021](https://arxiv.org/html/2609.26052#bib.bib45)\):A widely used benchmark of grade\-school math word problems that require multi\-step arithmetic reasoning\. Each problem is paired with a natural\-language solution, making it a standard testbed for evaluating chain\-of\-thought reasoning\.
- •Math\(Hendryckset al\.[2021](https://arxiv.org/html/2609.26052#bib.bib46)\):A challenging collection of competition\-level mathematics problems covering topics such as algebra, geometry, and number theory\. For efficient evaluation, we adopt the Math\-500 subset as the test set\.
- •Countdown\(Panet al\.[2025](https://arxiv.org/html/2609.26052#bib.bib47)\):An arithmetic planning task in which the model must combine a given set of numbers with basic operations to reach a target value\. Following\(Kimet al\.[2026b](https://arxiv.org/html/2609.26052#bib.bib19)\), we adopt the same 1\-shot prompting strategy to ensure valid outputs and fair comparison\.
- •StrategyQA\(Gevaet al\.[2021](https://arxiv.org/html/2609.26052#bib.bib51)\):A logical reasoning benchmark of open\-domain yes/no questions that require implicit multi\-hop reasoning, where the necessary reasoning steps are not explicitly stated and must be inferred by the model\.

### Expanded Experiment Results

The experimental results for Dream are presented in Table[3](https://arxiv.org/html/2609.26052#A0.T3)and lead to conclusions consistent with our analysis in Section[Experiment Results](https://arxiv.org/html/2609.26052#Sx5.SSx2)\. Furthermore, the results for LLaDA on Math under different inference budgets are shown in Fig\.[4](https://arxiv.org/html/2609.26052#A0.F4), where our proposed method consistently outperforms others across all settings\.

![Refer to caption](https://arxiv.org/html/2609.26052v1/img/math_budget_comparison.png)Figure 4:Performance comparison across inference budgets: our method consistently outperforms baselines on Math500 using LLaDA \(L=128L=128\)\.

Similar Articles