@probablynotaz9: Solo-author ICML paper alert Ever wanted to post-train your diffusion LLM with good old policy gradients, without havin…

X AI KOLs Following Papers

Summary

This solo-author ICML paper introduces Amortized Group Relative Policy Optimization (AGRPO) to enable effective reinforcement learning post-training for diffusion language models.

Solo-author ICML paper alert Ever wanted to post-train your diffusion LLM with good old policy gradients, without having to deal with ELBOs or surrogates? In Simple Policy Gradients for Reasoning with Diffusion Language Models, we show how to make this tractable in a straightforward way. Our framework, Amortized GRPO (AGRPO), lets the model learn from unbiased PG updates via timestep estimation, naturally aligning with dLLM inference while remaining efficient + scalable. Paper: https://arxiv.org/abs/2510.04019 Code: https://github.com/probablyabot/agrpo… 1/n
Original Article
View Cached Full Text

Cached at: 05/11/26, 08:44 PM

Solo-author ICML paper alert Ever wanted to post-train your diffusion LLM with good old policy gradients, without having to deal with ELBOs or surrogates? In Simple Policy Gradients for Reasoning with Diffusion Language Models, we show how to make this tractable in a straightforward way. Our framework, Amortized GRPO (AGRPO), lets the model learn from unbiased PG updates via timestep estimation, naturally aligning with dLLM inference while remaining efficient + scalable. Paper: https://arxiv.org/abs/2510.04019 Code: https://github.com/probablyabot/agrpo… 1/n


Simple Policy Gradients for Reasoning with Diffusion Language Models

Source: https://arxiv.org/html/2510.04019

Abstract

Diffusion large language models (dLLMs), which offer a promising alternative to traditional autoregressive LLMs, have recently shown strong results in pretraining. However, due to their lack of tractable sequence-level likelihoods, they have yet to benefit from modern LLM post-training techniques such as reinforcement learning (RL), limiting their real-world applicability. Existing attempts at dLLM post-training rely on heuristic approximations or lower bounds of the true likelihood. In this work, we propose Amortized Group Relative Policy Optimization (AGRPO), a policy gradient algorithm that leverages the multi-step Markovian nature of dLLM generation, optimizing individual denoising steps rather than entire sequences. We demonstrate AGRPO’s effectiveness on different math and reasoning tasks, achieving +9.9% absolute gain on GSM8K, +4.6% on MATH-500, +59.4% on Countdown, and +69.7% on Sudoku over the base LLaDA model, improving upon comparable dLLM RL methods such as diffu-GRPO. Furthermore, we analyze how post-training gains persist across different inference configurations, revealing that models trained with AGRPO can sample 4x faster with minimal performance sacrifices.

Diffusion language model, dLLM, Reinforcement learning, Post-training, Reasoning, Policy gradient

1Introduction

Many recent efforts in LLM research have centered around reinforcement learning, specifically in the verifiable reward (RLVR) setting. In a typical setup, base models are trained on math or coding problems and incentivized to reason through the solution step-by-step, getting a reward if the final answer is correct. The main goal of RLVR is to elicit mathematical thinking/reasoning capabilities, allowing models to solve complex real-world tasks.

This wave of interest in RL and reasoning, initially spurred by models like OpenAI’s o1(OpenAIet al.,2024)and DeepSeek’s R1(DeepSeek-AIet al.,2025a), has led to the development of numerous post-training algorithms designed specifically for transformer-based autoregressive (AR) LLMs. With the success of these algorithms, chief among them Group Relative Policy Optimization (GRPO)(Shaoet al.,2024), AR LLMs have grown incredibly strong on problem-solving benchmarks, with closed models even achieving gold medal performance at competitions such as the IMO and IOI, a remarkable feat(Luong and Lockhart,2025; Lin and Cheng,2025).

In a parallel line of research, diffusion language models have recently emerged as an alternative to the traditional autoregressive paradigm. Continuous diffusion models have long been established as the dominant framework for image and video generation, relying on a denoising/score matching objective. Works such as D3PM(Austinet al.,2021)and SEDD(Louet al.,2024)successfully transferred this diffusion framework to discrete settings, including language. Successive efforts such as MDLM(Sahooet al.,2024)and RADD(Ouet al.,2025)have simplified the theoretical framework, with most recent works settling on the “absorbing” or “masked” diffusion framework. We henceforth refer to this class of masked diffusion models as dLLMs.

Current state-of-the-art dLLMs, such as LLaDA(Nieet al.,2025)and MMaDA(Yanget al.,2025), are close to or on par with open-source AR models such as LLaMA3-8B and Qwen2.5-7B on common NLP benchmarks. Once trained, these models can go beyond standard left-to-right generation by starting with partially masked sequences, and additionally can trade off compute and quality by decreasing the number of sampling steps (i.e. sampling more tokens in parallel).

However, these models still struggle to match AR models in downstream tasks that require long-form thinking and reasoning. This discrepancy in post-training stems from fundamental challenges in designing training objectives for dLLMs: AR models have easy access to sequence-level likelihoods through AR factorization, whereas diffusion models must resort to approximations or ELBO-like bounds on likelihood. Unlocking true reasoning capabilities would be a giant leap forward for dLLMs, solidifying them as a true rival of AR LLMs.

Our work helps dLLMs close this gap by proposing a principled policy gradient algorithm designed especially for dLLMs: Amortized Group Relative Policy Optimization (AGRPO). Unlike conventional one-step approaches, we first establish a multi-step MDP formulation of the post-training problem, tying it to the iterative unmasking process used by dLLMs. Then, through a simple modification of the policy gradient objective — by viewing the inner sum over all tokens as an expectation over timesteps — we show how to make training tractable for long-form reasoning tasks.

Our main contributions are as follows:

  • •Soundness.We derive an unbiased policy gradient objective from a multi-step view of the dLLM generation process, explaining how our approach sidesteps the need for heuristic likelihood approximations or ELBO-like bounds while remaining theoretically sound.
  • •Efficiency.Using statistical techniques, we show how to implement our proposed algorithm in a stable, memory-efficient way, and discuss various practical tradeoffs.
  • •Efficacy.We train models on four reasoning tasks (GSM8K, MATH, Countdown, and Sudoku), showing that AGRPO outperforms all previous approximation-based methods. In addition, we show that models post-trained with AGRPO retain high accuracy even when evaluated with much fewer sampling steps, a capability not found in pretrained base dLLMs.

2Preliminaries

2.1dLLM Pretraining

The most common form of discrete diffusion for language is the masked (or “absorbing”) approach, where models are trained to reverse data corrupted by randomly masking tokens(Louet al.,2024; Sahooet al.,2024; Arriolaet al.,2025). Concretely, given a distributionppon sequences of discrete tokensx=(x1,…,xn)x=(x_{1},\dots,x_{n}), models are trained to maximize the following evidence lower bound (ELBO) on the likelihood(Nieet al.,2025; Ouet al.,2025):

ℒ​(θ)=𝔼t∼U​[0,1]x,xt∼pt​[1t​∑xit=■log⁡pθ​(xi∣xt)]\mathcal{L}(\theta)=\mathbb{E}_{\begin{subarray}{c}t\sim U[0,1]\\ x,x^{t}\sim p^{t}\end{subarray}}\left[\frac{1}{t}\sum_{x^{t}_{i}=\blacksquare}\log p_{\theta}(x_{i}\mid x^{t})\right](1)wherex,xt∼ptx,x^{t}\sim p^{t}means thatxxis sampled fromppandxtx^{t}is obtained fromxxby independently setting each tokenxix_{i}to the mask token■\blacksquarewith probabilitytt. Similar to BERT(Devlinet al.,2019), the goal is for the model to learn marginal distributions of masked tokens conditioned on context.

A crucial point is that compared to the classic AR objective, this masked token prediction objective is harder and more general, since the unmasking order can be arbitrary and the model must predictmultiplemasked tokens. By contrast, AR models constrain themselves to modeling thenexttoken in left-to-right order. The benefits of imposing this constraint are twofold: it lets AR models maximize the exact sequence likelihood via the chain rule, and it also lets training be parallelized via causal self-attention when used with decoder-only transformers.

2.2dLLM Inference

To generate text, dLLMs start with an all- or partially-masked sequence, obtain marginal distributions for each masked token, and then unmask some of these by sampling from their marginals. The positions to be unmasked can be chosen either randomly, adhering to the theoretical “backward process,” or by keeping the tokens with highest probability, as proposed byNieet al.(2025). (We refer to these as “random” and “confidence-based” unmasking, respectively.) The rest of the tokens are kept the same — masked tokens remain masked, unmasked tokens remain unmasked — and this new sequence is fed back into the model. This process is repeated until all tokens are unmasked.

Throughout this paper, we usemmto refer to the number of sampling steps, andnnto refer to the sequence length. A nice advantage of diffusion models is the ability to dynamically adjust the number of tokens unmasked per stepn/mn/m. Typically, the ration/mn/mis chosen to be relatively small (≤8\leq 8) — unmasking higher tokens at each step severely degrades quality as measured by perplexity or accuracy(Louet al.,2024; Nieet al.,2025). We show in later sections that for specific tasks, post-training actually allows for much higher values ofn/mn/mwithout clear degradation, unlike pretrained models.

For more details on dLLM inference, see AppendixF.

2.3Reinforcement Learning and MDPs

Markov decision processes (MDPs) are a formalization of sequential decision-making problems consisting of a state space𝒮\mathcal{S}, an action space𝒜\mathcal{A}, a transition kernelp(⋅∣s,a)p(\cdot\mid s,a), and a reward functionr​(s,a)r(s,a). Broadly speaking, the goal of reinforcement learning (RL) is to learn a policy, i.e. a distributionπ​(a∣s)\pi(a\mid s), that maximizes the expected sum of rewards:

𝔼τ∼π​[∑t=0Tr​(st,at)]\mathbb{E}_{\tau\sim\pi}\left[\sum_{t=0}^{T}r(s_{t},a_{t})\right](2)whereτ\taurepresents a trajectory (or “rollout”), i.e. a sequence of states and actions(s0,a0,s1,a2,…,sT,aT)(s_{0},a_{0},s_{1},a_{2},\dots,s_{T},a_{T})whereai∼π(⋅∣si)a_{i}\sim\pi(\cdot\mid s_{i})andsi+1∼p(⋅∣si,ai)s_{i+1}\sim p(\cdot\mid s_{i},a_{i}).

3Policy Gradients for Diffusion Models

Refer to captionFigure 1:Existing RL post-training algorithms focus on sequence-level likelihoods and require either autoregressive factorization or ELBO-like bounds, which result in biased policy gradients. Our proposed algorithm instead focuses on individual unmasking steps, aligning more naturally with the dLLM generation process.In this section, we describe our framing of dLLM post-training as a multi-step RL problem, which differs from the standard sequence-level framing of LLM post-training. To motivate this, we first give an overview of policy gradients.

3.1Policy Gradient Methods

Policy gradients (PG) comprise a popular class of algorithms used to train neural network-parameterized policiesπθ\pi_{\theta}to maximize expected rewards(Suttonet al.,1999). The simplest form of policy gradients, REINFORCE(Williams,1992), involves the following gradient update:

∇θ𝒥P​G​(θ)=𝔼τ∼π[(∑t=0T∇θlog⁡πθ​(at∣st))​(∑t=0Tr​(st,at))].\nabla_{\theta}\mathcal{J}_{PG}(\theta)=\mathbb{E}_{\tau\sim\pi}\\ \left[\left(\sum_{t=0}^{T}\nabla_{\theta}\log\pi_{\theta}(a_{t}\mid s_{t})\right)\left(\sum_{t=0}^{T}r(s_{t},a_{t})\right)\right].(3)More sophisticated algorithms such as Proximal Policy Optimization(Schulmanet al.,2017)improve upon this formulation by subtracting a learnable reward baseline (i.e. a value function) and allowing off-policy updates via an importance sampling correction. However, the underlying structure of all PG methods remains the same: compute the (exact) likelihoods of all actions along the trajectory, weighted by reward. Building on this common structure, we show how to develop a principled form of PG for dLLMs.

3.2LLM Post-Training with RL

In the context of post-training LLMs, specifically RL with verifiable rewards (RLVR), the statesscorresponds to the context (i.e. the prompt), actionsaacorrespond to model outputs, and the rewardrris provided via a ground truth answer. (Transitions are deterministic, so we can safely ignore the transition kernelpp.)

Thanks to AR factorization, one can view AR LLMs as policies that generate distributions over the entire sequence, i.e. aone-stepMDP. Under this view, trajectories in (2) consist simply of the initial prompts0s_{0}and the model outputa0a_{0}. This makes the PG objective (3) quite convenient, and is quite effective for many post-training tasks; we discuss the large body of work devoted to this setting in Section6.

3.3Text Generation as a Multi-Step MDP

Diffusion models don’t admit a clean left-to-right factorization, which makes sequence likelihoods intractable. In particular, for dLLMs such as LLaDA, computing the exact likelihood of a given lengthnnsequence would require marginalizing overO​(n!)O(n!)unmasking orders, which is simply infeasible for largenn. Maximizing a lower bound on the likelihood suffices for pretraining, but is incompatible with the one-step MDP view of post-training, which requires exact sequence likelihoods.

While existing dLLM post-training approaches simply use an ELBO in place of the true likelihood, we would like to develop a more principled approach without such compromises. This motivates us to consider amulti-stepMDP formulation of the generation process where

  • •the statessis the current partially masked sequence,
  • •the actionsaaare then/mn/mtokens to be unmasked111The positions to be unmasked are determined by the unmasking strategy (random or confidence-based) and are not explicitly optimized, following theoretical assumptions of diffusion models., and
  • •the rewardrris provided after the sequence is fully unmasked, i.e. at the last timestep.

In other words, a single step in this MDP is taken to be an individual unmasking step rather than the entire generation process.

This formulation, which aligns more naturally with how diffusion models are parameterized, has been adopted by previous works on RL for continuous diffusion such as DDPO and Flow-GRPO(Blacket al.,2024; Liuet al.,2025a). Crucially, this perspective of denoising as a multi-step MDP allows forexact action likelihoodsin the PG objective, sidestepping the need for lower bounds or approximations.

3.4Larger Models and Longer Trajectories

Note that although the multi-step perspective allows for exact likelihoods, computing (3) naively requires a separate forward pass for every denoising step. In domains such as robotics, the policy network and state representation are often small enough that one can batch many state-action pairs into a single forward pass.

However, for large diffusion models, particularly transformers which must attend to complex states, GPU memory significantly limits the potential for batching. (Using the batch dimension this way would also cut into the ability to batch multiple trajectories at once.) For image generation tasks where the number of stepsmmis smaller, spendingmmforward passes to compute the full objective can work, but this becomes impractical for reasoning tasks with hundreds of steps. We propose a novel way of addressing this problem via timestep sampling in Section4.1.

4Amortized Group Relative Policy Optimization

Before presenting the AGRPO objective, we first introduce some dLLM-specific notation. LetDDbe a distribution over questionsqq,{oi}i=1G\{o^{i}\}_{i=1}^{G}a group ofGGoutputs (or rollouts) conditioned on someqq, andrir_{i}the respective rewards. Each output has lengthnn, and is generated withmmunmasking steps. We useoto_{t}to denote the partially masked state of rolloutooat timesteptt. For example,πθ​(o1∣q,o0)\pi_{\theta}(o_{1}\mid q,o_{0})represents the probabilities of the firstn/mn/mtokens to be unmasked. (Recall that the unmasking order need not be left to right.)

4.1From PPO to AGRPO

In this section, we show how to derive the AGRPO objective by reinterpreting the inner sum in the PG objective as an expectation across timesteps. Instead of the form given in Equation (3), we work with PPO(Schulmanet al.,2017), a more modern form which includes advantages and importance sampling (and clipping, which we temporarily omit for clarity). With the notation above, the PPO surrogate objective is:

𝒥​(θ)=1m​G​∑i=1G∑t=1mπθ​(oti∣q,ot−1i)πo​l​d​(oti∣q,ot−1i)​Ai\mathcal{J}(\theta)=\frac{1}{mG}\sum_{i=1}^{G}\sum_{t=1}^{m}\frac{\pi_{\theta}(o^{i}_{t}\mid q,o^{i}_{t-1})}{\pi_{old}(o^{i}_{t}\mid q,o^{i}_{t-1})}A_{i}(4)whereAi=ri−mean⁡{ri}A_{i}=r_{i}-\operatorname{mean}\{r_{i}\}is the advantage estimate andπo​l​d\pi_{old}is the policy under which rollouts are sampled, which is updated everyμ\mugradient steps.

As in GRPO(Shaoet al.,2024), we use group-normalized advantages for simplicity. Following subsequent improvements to GRPO, we avoid dividing bystd⁡{ri}\operatorname{std}\{r_{i}\}to avoid bias from particularly easy or hard problems where advantages have low variance(Liuet al.,2025b).

Now letTTbe a random variable drawn uniformly from{1,…,m}\{1,\dots,m\}. Then an unbiased estimator of Equation (4) is

1G​∑i=1G𝔼T∼{1,…,m}​[πθ​(oTi∣q,oT−1i)πo​l​d​(oTi∣q,oT−1i)​Ai].\frac{1}{G}\sum_{i=1}^{G}\mathbb{E}_{T\sim\{1,\dots,m\}}\left[\frac{\pi_{\theta}(o^{i}_{T}\mid q,o^{i}_{T-1})}{\pi_{old}(o^{i}_{T}\mid q,o^{i}_{T-1})}A_{i}\right]. Thus, with Monte Carlo (MC) sampling, the objective can now be estimated by drawingk≪mk\ll mtimesteps, computing the unmasking likelihood ratios, and averaging. Since the likelihoods are now at the individual timestep level, the resulting gradient estimate isunbiased, unlike existing works which substitute ELBOs for sequence likelihoods.

The full AGRPO objective with clipping is

𝒥(θ)=𝔼q∼D{oi}i=1G∼πo​l​d(⋅∣q)[1G∑i=1G𝔼t∼{1,…,m}[min(ρtiAi,clip(ρti,1−ε,1+ε)Ai)−βDKL]],\mathcal{J}(\theta)=\mathbb{E}_{\begin{subarray}{c}q\sim D\\ \{o^{i}\}_{i=1}^{G}\sim\pi_{old}(\cdot\mid q)\end{subarray}}\Bigg[\frac{1}{G}\sum_{i=1}^{G}\mathbb{E}_{t\sim\{1,\dots,m\}}\\ \left[\min(\rho^{i}_{t}A_{i},\operatorname{clip}\left(\rho^{i}_{t},1-\varepsilon,1+\varepsilon\right)A_{i})-\beta D_{\mathrm{KL}}\right]\Bigg],(5)where

ρti=πθ​(oti∣q,ot−1i)πo​l​d​(oti∣q,ot−1i)\rho^{i}_{t}=\frac{\pi_{\theta}(o^{i}_{t}\mid q,o^{i}_{t-1})}{\pi_{old}(o^{i}_{t}\mid q,o^{i}_{t-1})}is the importance sampling ratio222Recall that dLLM inference works by factorizing the joint probability of unmaskingn/mn/mtokens as the product of marginals.andDKLD_{\mathrm{KL}}is shorthand forDKL(πθ||πref)D_{\mathrm{KL}}(\pi_{\theta}\,||\,\pi_{\operatorname{ref}}). Note that under our multi-step framing, the KL term represents the divergence over the specific positions unmasked at steptt, i.e.πθ(⋅∣q,ot−1i)\pi_{\theta}(\cdot\mid q,o^{i}_{t-1}). This means we can also estimate it via MC sampling without relying on sequence-level approximations. To estimateDKLD_{\mathrm{KL}}, we useSchulman (2020)’s unbiasedk3k_{3}estimator

DKL(p||q)=𝔼x∼p[q​(x)p​(x)−logq​(x)p​(x)−1]D_{\mathrm{KL}}(p||q)=\mathbb{E}_{x\sim p}\left[\frac{q(x)}{p(x)}-\log\frac{q(x)}{p(x)}-1\right]which is also used by the original GRPO paper(Shaoet al.,2024).

Algorithm1provides an overview of our proposed algorithm. For practical considerations, including how to efficiently retrieve partially masked statesoto_{t}and compute gradients, see AppendixE.

Algorithm 1Amortized Group Relative Policy Optimization (AGRPO)0:policy

πθ\pi_{\theta}, # sampling steps

mm, # MC samples

kk πref←πθ\pi_{\operatorname{ref}}\leftarrow\pi_{\theta}

whilenot convergeddo

πo​l​d←πθ\pi_{old}\leftarrow\pi_{\theta}

sample prompt

q∼Dq\sim D sample rollouts

{oi}i=1G∼πo​l​d(⋅∣q)\{o^{i}\}_{i=1}^{G}\sim\pi_{old}(\cdot\mid q) compute advantages

{Ai}\{A_{i}\} for

ℓ=1\ell=1to

μ\mudo

J^←0\widehat{J}\leftarrow 0{stores the MC estimate}

for

j=1j=1to

kkdo

sample

t∼{1,…,m}t\sim\{1,\dots,m\}uniformly

ρti←πθ​(oti∣q,ot−1i)πo​l​d​(oti∣q,ot−1i)\rho^{i}_{t}\leftarrow\frac{\pi_{\theta}(o^{i}_{t}\mid q,o^{i}_{t-1})}{\pi_{old}(o^{i}_{t}\mid q,o^{i}_{t-1})}

J^←J^+∑i=1G[min(clip(ρti,1−ε,1+ε)Ai,ρtiAi)−βDKL(πθ||πref)]\begin{aligned} \widehat{J}\leftarrow\widehat{J}+\sum_{i=1}^{G}\big[\min\big(\operatorname{clip}\left(\rho^{i}_{t},1-\varepsilon,1+\varepsilon\right)A_{i},\\ \rho^{i}_{t}A_{i}\big)-\beta D_{\mathrm{KL}}(\pi_{\theta}\,||\,\pi_{\operatorname{ref}})\big]\end{aligned}

endfor

compute AGRPO estimate

𝒥​(θ)=J^k​G\mathcal{J}(\theta)=\frac{\widehat{J}}{kG} backpropagate loss and take gradient step w.r.t.

θ\theta endfor

endwhile

4.2Variance Reduction Techniques

For a fixed number of sampleskk, a naive MC sampling algorithm would drawkki.i.d. samples, compute the objective, and average the results. We would ideally like an estimator with minimal variance to ensure stable training. Here we describe two such ways of reducing variance, and empirically test their effects in Section5.5. The full variance-reduced AGRPO estimator, with a proof of unbiasedness, can be found in AppendixA.

4.2.1Low-Discrepancy Sampling

Instead of i.i.d. sampling, we can introduce correlation across samples so that they collectively “cover” a wide range of timesteps while ensuring that the marginal distribution for each sample is still uniform on{1,…,m}\{1,\dots,m\}. This is known as low-discrepancy sampling, and is used in practice to lower training variance for both continuous and discrete diffusion models(Kingmaet al.,2021; Sahooet al.,2024; Zhenget al.,2025). We followZhenget al.(2025)’s discrete low-discrepancy sampler, which is detailed in AppendixD.

A desirable property of low-discrepancy sampling is that in the limitk→mk\to m, we fully recreate the original GRPO objective. In other words, one can achieve higher fidelity by scaling the amount of compute. We investigate this tradeoff induced bykkin Section5.4.

4.2.2Entropy Importance Sampling

The uniform measure on{1,…,m}\{1,\dots,m\}treats all timesteps equally, even though it is plausible that not all tokens contribute equally to the final solution. Inspired byWanget al.(2025b), who found that a minority of high-entropy tokens were responsible for the majority of performance gains in RLVR, we propose an entropy-based importance sampling scheme.

Instead of drawingT∼{1,…,m}T\sim\{1,\dots,m\}uniformly, we compute an entropy scoreete_{t}for each timestepttby summing the entropies of the unmasked tokens attt. Then we drawTTsuch thatPr⁡[T=t]∝et\Pr[T=t]\propto e_{t}, compute the objective givenTT, and multiply by a correction term∑et/(m​eT)\sum e_{t}/(me_{T}). This ensures that the model receives updates from critical reasoning tokens while remaining unbiased.

5Experiments

Table 1:Accuracies for different RL post-training methods across different reasoning tasks and generation lengths. For all diffusion models, outputs of lengthnnare generated withm=n/2m=n/2steps. Few-shot examples are denoted in parentheses, and the best accuracy for each task isbolded. Models trained with LoRA (rather than full fine-tuning) are denoted with †. To empirically validate our proposed algorithm, we start from the open source base model LLaDA-8B-Instruct(Nieet al.,2025)and fine-tune models using AGRPO on four different reasoning tasks: GSM8K, MATH, Countdown, and Sudoku.

5.1Datasets

GSM8K/MATH are standard problem-solving benchmarks consisting of 8.5k/12.5k math problems at the grade school/high school level, respectively(Cobbeet al.,2021; Hendryckset al.,2021). Countdown is a popular math reasoning task where the model is given a list of 3-4 numbers and a target number; the goal is to combine the numbers using arithmetic operations (+,−,×,/+,-,\times,/) and parentheses to get the target number(Panet al.,2025). Sudoku is a planning task where the goal is to fill in a 4x4 grid of numbers according to uniqueness constraints.

These four tasks form the common set of benchmarks for the growing literature on dLLM reasoning(Wanget al.,2025a; Zhaoet al.,2025; Tanget al.,2025). For consistency, we use the Countdown and Sudoku splits provided in SPG’s codebase(Wanget al.,2025a). We use HuggingFace’s math-verify library for parsing GSM8K, MATH, and Countdown answers.

5.2Experimental Setup

Following previous dLLM RL works, we use Low-Rank Adaptation(Huet al.,2021)instead of full fine-tuning. We fix the response length atn=384n=384and the number of steps atm=128m=128to balance inference wall time and coherence, both of which are important for RLVR, and we usek=24k=24MC samples. Notably, despite being trained on a single configuration, we observe that model performance generalizes to different output lengths and steps.

During training, we generate rollouts with temperature 0.6 and random remasking to inject stochasticity and incentivize exploration. Random remasking makes additional sense in the context of AGRPO since it allows rollouts with similar final sequences to have vastly different intermediate states, helping the model learn from diverse contexts.

Models are trained until convergence is observed (e.g. via the reward curve plateauing); we select checkpoints from the last 150 steps for testing and report the best accuracy among those checkpoints. For evaluation, we switch to confidence-based unmasking (which can be thought of as a form of annealing(Nieet al.,2025)) with temperature 0, keeping the same generation process as previous works. Other hyperparameters, includingGGandε\varepsilon, can be found in AppendixB.1.

5.3Results

Strong reasoning improvements.We report accuracies on test splits in Table1. For diffusion models, AGRPO achieves the highest accuracy across all but one task, comfortably beating the base LLaDA model and other dLLM post-training methods, including diffu-GRPO(Zhaoet al.,2025), VRPO(Zhuet al.,2025), and SPG(Wanget al.,2025a). At sequence lengthn=512n=512, we improve upon the previous best-known results by+1.4%+1.4\%on GSM8K,+0.2%+0.2\%on MATH-500,+12.1%+12.1\%on Countdown, and+2.1%+2.1\%on Sudoku. Baseline comparisons are discussed in greater detail in AppendixB.2.

Our results suggest that tailoring towards the Markovian nature of the diffusion process is the right way to extend AR LLM reasoning abilities to dLLMs, dispensing with the need for unprincipled approximations or bounds.

Comparable gains to GRPO.Although performance on MATH still lags behind autoregressive models, we are able to achieve parity with DeepSeekMath-RL on GSM8K, a substantial improvement for dLLMs. Additionally, we observe increased performance deltas for AGRPO compared to GRPO (e.g. for GSM8K, GRPO results in+5.3%+5.3\%, whereas AGRPO results in+9.9%+9.9\%atn=512n=512).

5.3.1Inference tradeoffs

Refer to captionFigure 2:The compute/quality frontier for the GSM8K test split with response lengthn=384n=384. Lines show the possible tradeoffs at inference time for a specific model.A key quality of dLLMs is their ability to trade off compute and quality at inference time. We examine how AGRPO affects the inference compute/quality frontier by evaluating models on the GSM8K test split at a fixed response lengthn=384n=384and varying the number of sampling stepsmm. As shown in Figure2, not only does the AGRPO model consistently achieve higher performance across all sampling steps, it matches baselines with4x fewer sampling steps, a remarkable speedup. At the fastest setting,m=32m=32, AGRPO achieves 59.7% accuracy despite sampling 12 tokens per step, an 8.1x performance improvement over LLaDA-8B-Instruct and 3.7x improvement over LLaDA 1.5. These results demonstrate that the reasoning skills instilled by AGRPO arerobust, generalizing to different inference configurations despite being trained on a fixednnandmm. We give sample responses from this experiment in AppendixG.

To our knowledge, this is one of the first investigations of how post-training affects inference tradeoffs across a complete range of sampling steps, opening up a new perspective on the benefits of post-training dLLMs. As an example of a downstream application, model providers interested in a specific dLLM use case can pay a one-time fine-tuning cost in order to generate cheaper responses (i.e. lowmm) for that use case without sacrificing quality. In the long run, this amortization of inference costs could enable huge savings.

5.4Ablations onkk

Refer to caption(a)GSM8K reward over training runs with differentkk. The shaded area represents intra-run variance over a rolling window of 15 steps. (b)Average wall time per gradient step for different components of AGRPO. Values are reported on 8xH100 GPUs with a global batch size of 128.

Figure 3:Reward curve and wall time comparisons for different values ofkkon GSM8K withn=384n=384andm=128m=128.With naive Monte Carlo sampling, increasing the number of samples by some factorccreduces the variance bycc. This is the simplest lens through which to understandkk, but the picture becomes slightly more nuanced in the context of online RL. For dLLMs, generating rollouts (which doesn’t depend onkk) is often more expensive than computing the actual policy update due to the large number of steps and lack of KV caching in dLLM inference. In other words, increasingkkby 2x doesn’t necessarily correspond to a 2x increase in overall training time. Thus, from an efficiency standpoint, choosing moderately largekkis fine (as long ask≪mk\ll m).

As seen in Figure3, our empirical findings are consistent with our expectations: larger values ofkkincrease the wall time per step but converge faster. Note that althoughk=32k=32spends almost 2x more compute per step thank=2k=2, it takes far more than 2x as many steps fork=2k=2to reach the same reward level, suggesting that extremely small values ofkkare strictly worse.

To summarize, choosingkkto be a healthy number relative tomm(for example,k≈m/4k\approx m/4) is optimal not only because it improves the MC estimate, but because it helps “amortize” inference costs. Pinpointing the exact relationship betweennn,mm, and the optimal value ofkkis an area we identify for future work.

For similar benchmarks against diffu-GRPO and SPG, see AppendixB.1.

5.5Ablations on Variance Reduction Techniques

Refer to captionFigure 4:Reward and gradient norm plots for baseline and variance-reduced versions of AGRPO on Countdown withn=128n=128andm=64m=64. The shaded area of the reward plot represents a rolling standard deviation with a window of 15 steps.In this section, we investigate whether the variance reduction techniques proposed in Section4.2truly help empirically. We train two models, one with low-discrepancy sampling and entropy importance sampling and one without, on the Countdown task. All other hyperparameters, including the sequence lengthn=128n=128and sampling stepsm=64m=64, are the same between the two runs.

The reward curves and gradient norms are shown in Figure4. Although both models reach a similar final reward level, the variance-reduced model converges faster, with less signs of early instability. In addition, the gradient norms for the variance-reduced model are consistently lower than the baseline throughout training (except for a few outlier steps). These results suggest that low discrepancy sampling and entropy-based importance sampling do indeed provide less noisy, more valuable updates.

6Related Works

Policy gradient methods for LLMs.Early attempts at RL with LLMs(Ouyanget al.,2022; Ziegleret al.,2020)focused on alignment with human preferences using manually-labeled preference data and methods such as PPO or Direct Preference Optimization(Rafailovet al.,2023). More recently, RL efforts have focused on reasoning capabilities, specifically for domains with verifiable rewards such as math and coding. LLMs post-trained with algorithms such as GRPO exhibited improved reasoning abilities beyond SFT(Luoet al.,2025)and even emergent capabilities such as backtracking and self-correction(Xionget al.,2025). Subsequent works modify GPRO to address issues in response length bias, sample efficiency, and stability(Yuet al.,2025; DeepSeek-AIet al.,2025b).

RL for continuous diffusion.Continuous diffusion models are highly popular in areas like image generation(Sohl-Dicksteinet al.,2015; Peebles and Xie,2022), video generation(Hoet al.,2022), and robotic control(Chiet al.,2024). To better align these models with desired downstream behavior, e.g. text-to-image tasks, different PG methods for diffusion have been proposed, including DDPO(Blacket al.,2024)and DPPO(Renet al.,2024). Crucially, these approaches formulate the RL problem as a multi-step denoising MDP, which our work extends to the dLLM setting for the first time.

RL for dLLM reasoning.Discrete diffusion models extend the diffusion framework to areas such as text and DNA sequences(Louet al.,2024; Austinet al.,2021). As with traditional LLMs, there are many reasons to post-train dLLMs, including to elicit advanced reasoning capabilities. While some works focus on the more general continuous time/score entropy perspective(Zekri and Boullé,2025), many works choose to focus on masked diffusion models such as LLaDA, which are more similar to traditional LLMs.Zhaoet al.(2025)introduced the first dLLM post-training framework, d1, involving a CoT SFT stage followed by RLVR. Their proposed RL algorithm, diffu-GRPO, uses a mean-field approximation of the sequence likelihood along with random prompt masking. More recent works, such as wd1(Tanget al.,2025)and SPG(Wanget al.,2025a), use more sophisticated ELBO-like bounds to achieve better empirical results. We discuss these works and others in more detail in AppendixC.

These methods, which assume the standard sequence-level framing of LLM post-training, rely on one-step likelihood approximations in order to remain tractable for dLLMs, thereby resulting in biased policy updates. We demonstrate that with the right diffusion-specific framing of RL — and some statistical techniques to make it practical — such approximations aren’t necessary to elicit strong reasoning skills.

7Conclusion

This work presents AGRPO, a policy gradient algorithm designed for dLLMs, grounded in the multi-step denoising perspective of RL. Unlike previous dLLM RL works that rely on likelihood approximations or bounds, AGRPO computes policy gradient estimates in an unbiased, efficient way via Monte Carlo sampling and statistical variance reduction techniques, making it both principled and tractable. Using our proposed algorithm, we show how to effectively post-train dLLMs, beating comparable methods across multiple tasks and redefining the inference compute/quality frontier. These contributions establish AGRPO as a viable way to transfer policy gradient RL techniques to the dLLM setting; we hope future works can build on our methods, either theoretically or empirically, and further close the gap between dLLM and AR LLM post-training.

Impact Statement

This paper presents work whose goal is to advance the field of Machine Learning. There are many potential societal consequences of our work, none which we feel must be specifically highlighted here.

References

  • M. Arriola, S. S. Sahoo, A. Gokaslan, Z. Yang, Z. Qi, J. Han, J. T. Chiu, and V. Kuleshov (2025)Block diffusion: interpolating between autoregressive and diffusion language models.External Links:LinkCited by:Appendix F,§2.1.
  • J. Austin, D. D. Johnson, J. Ho, D. Tarlow, and R. van den Berg (2021)Structured denoising diffusion models in discrete state-spaces.External Links:LinkCited by:§1,§6.
  • K. Black, M. Janner, Y. Du, I. Kostrikov, and S. Levine (2024)Training diffusion models with reinforcement learning.External Links:LinkCited by:§3.3,§6.
  • C. Chi, Z. Xu, S. Feng, E. Cousineau, Y. Du, B. Burchfiel, R. Tedrake, and S. Song (2024)Diffusion policy: visuomotor policy learning via action diffusion.External Links:2303.04137,LinkCited by:§6.
  • K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman (2021)Training verifiers to solve math word problems.arXiv preprint arXiv:2110.14168.Cited by:§5.1.
  • DeepSeek-AI, D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, X. Zhang, X. Yu, Y. Wu, Z. F. Wu, Z. Gou, Z. Shao, Z. Li, Z. Gao, A. Liu, B. Xue, B. Wang, B. Wu, B. Feng, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, D. Dai, D. Chen, D. Ji, E. Li, F. Lin, F. Dai, F. Luo, G. Hao, G. Chen, G. Li, H. Zhang, H. Bao, H. Xu, H. Wang, H. Ding, H. Xin, H. Gao, H. Qu, H. Li, J. Guo, J. Li, J. Wang, J. Chen, J. Yuan, J. Qiu, J. Li, J. L. Cai, J. Ni, J. Liang, J. Chen, K. Dong, K. Hu, K. Gao, K. Guan, K. Huang, K. Yu, L. Wang, L. Zhang, L. Zhao, L. Wang, L. Zhang, L. Xu, L. Xia, M. Zhang, M. Zhang, M. Tang, M. Li, M. Wang, M. Li, N. Tian, P. Huang, P. Zhang, Q. Wang, Q. Chen, Q. Du, R. Ge, R. Zhang, R. Pan, R. Wang, R. J. Chen, R. L. Jin, R. Chen, S. Lu, S. Zhou, S. Chen, S. Ye, S. Wang, S. Yu, S. Zhou, S. Pan, S. S. Li, S. Zhou, S. Wu, S. Ye, T. Yun, T. Pei, T. Sun, T. Wang, W. Zeng, W. Zhao, W. Liu, W. Liang, W. Gao, W. Yu, W. Zhang, W. L. Xiao, W. An, X. Liu, X. Wang, X. Chen, X. Nie, X. Cheng, X. Liu, X. Xie, X. Liu, X. Yang, X. Li, X. Su, X. Lin, X. Q. Li, X. Jin, X. Shen, X. Chen, X. Sun, X. Wang, X. Song, X. Zhou, X. Wang, X. Shan, Y. K. Li, Y. Q. Wang, Y. X. Wei, Y. Zhang, Y. Xu, Y. Li, Y. Zhao, Y. Sun, Y. Wang, Y. Yu, Y. Zhang, Y. Shi, Y. Xiong, Y. He, Y. Piao, Y. Wang, Y. Tan, Y. Ma, Y. Liu, Y. Guo, Y. Ou, Y. Wang, Y. Gong, Y. Zou, Y. He, Y. Xiong, Y. Luo, Y. You, Y. Liu, Y. Zhou, Y. X. Zhu, Y. Xu, Y. Huang, Y. Li, Y. Zheng, Y. Zhu, Y. Ma, Y. Tang, Y. Zha, Y. Yan, Z. Z. Ren, Z. Ren, Z. Sha, Z. Fu, Z. Xu, Z. Xie, Z. Zhang, Z. Hao, Z. Ma, Z. Yan, Z. Wu, Z. Gu, Z. Zhu, Z. Liu, Z. Li, Z. Xie, Z. Song, Z. Pan, Z. Huang, Z. Xu, Z. Zhang, and Z. Zhang (2025a)DeepSeek-r1: incentivizing reasoning capability in llms via reinforcement learning.External Links:2501.12948,LinkCited by:§1.
  • DeepSeek-AI, A. Liu, A. Mei, B. Lin, B. Xue, B. Wang, B. Xu, B. Wu, B. Zhang, C. Lin, C. Dong, C. Lu, C. Zhao, C. Deng, C. Xu, C. Ruan, D. Dai, D. Guo, D. Yang, D. Chen, E. Li, F. Zhou, F. Lin, F. Dai, G. Hao, G. Chen, G. Li, H. Zhang, H. Xu, H. Li, H. Liang, H. Wei, H. Zhang, H. Luo, H. Ji, H. Ding, H. Tang, H. Cao, H. Gao, H. Qu, H. Zeng, J. Huang, J. Li, J. Xu, J. Hu, J. Chen, J. Xiang, J. Yuan, J. Cheng, J. Zhu, J. Ran, J. Jiang, J. Qiu, J. Li, J. Song, K. Dong, K. Gao, K. Guan, K. Huang, K. Zhou, K. Huang, K. Yu, L. Wang, L. Zhang, L. Wang, L. Zhao, L. Yin, L. Guo, L. Luo, L. Ma, L. Wang, L. Zhang, M. S. Di, M. Y. Xu, M. Zhang, M. Zhang, M. Tang, M. Zhou, P. Huang, P. Cong, P. Wang, Q. Wang, Q. Zhu, Q. Li, Q. Chen, Q. Du, R. Xu, R. Ge, R. Zhang, R. Pan, R. Wang, R. Yin, R. Xu, R. Shen, R. Zhang, S. H. Liu, S. Lu, S. Zhou, S. Chen, S. Cai, S. Chen, S. Hu, S. Liu, S. Hu, S. Ma, S. Wang, S. Yu, S. Zhou, S. Pan, S. Zhou, T. Ni, T. Yun, T. Pei, T. Ye, T. Yue, W. Zeng, W. Liu, W. Liang, W. Pang, W. Luo, W. Gao, W. Zhang, X. Gao, X. Wang, X. Bi, X. Liu, X. Wang, X. Chen, X. Zhang, X. Nie, X. Cheng, X. Liu, X. Xie, X. Liu, X. Yu, X. Li, X. Yang, X. Li, X. Chen, X. Su, X. Pan, X. Lin, X. Fu, Y. Q. Wang, Y. Zhang, Y. Xu, Y. Ma, Y. Li, Y. Li, Y. Zhao, Y. Sun, Y. Wang, Y. Qian, Y. Yu, Y. Zhang, Y. Ding, Y. Shi, Y. Xiong, Y. He, Y. Zhou, Y. Zhong, Y. Piao, Y. Wang, Y. Chen, Y. Tan, Y. Wei, Y. Ma, Y. Liu, Y. Yang, Y. Guo, Y. Wu, Y. Wu, Y. Cheng, Y. Ou, Y. Xu, Y. Wang, Y. Gong, Y. Wu, Y. Zou, Y. Li, Y. Xiong, Y. Luo, Y. You, Y. Liu, Y. Zhou, Z. F. Wu, Z. Z. Ren, Z. Zhao, Z. Ren, Z. Sha, Z. Fu, Z. Xu, Z. Xie, Z. Zhang, Z. Hao, Z. Gou, Z. Ma, Z. Yan, Z. Shao, Z. Huang, Z. Wu, Z. Li, Z. Zhang, Z. Xu, Z. Wang, Z. Gu, Z. Zhu, Z. Li, Z. Zhang, Z. Xie, Z. Gao, Z. Pan, Z. Yao, B. Feng, H. Li, J. L. Cai, J. Ni, L. Xu, M. Li, N. Tian, R. J. Chen, R. L. Jin, S. S. Li, S. Zhou, T. Sun, X. Q. Li, X. Jin, X. Shen, X. Chen, X. Song, X. Zhou, Y. X. Zhu, Y. Huang, Y. Li, Y. Zheng, Y. Zhu, Y. Ma, Z. Huang, Z. Xu, Z. Zhang, D. Ji, J. Liang, J. Guo, J. Chen, L. Xia, M. Wang, M. Li, P. Zhang, R. Chen, S. Sun, S. Wu, S. Ye, T. Wang, W. L. Xiao, W. An, X. Wang, X. Sun, X. Wang, Y. Tang, Y. Zha, Z. Zhang, Z. Ju, Z. Zhang, and Z. Qu (2025b)DeepSeek-v3.2: pushing the frontier of open large language models.External Links:2512.02556,LinkCited by:Appendix E,§6.
  • J. Devlin, M. Chang, K. Lee, and K. Toutanova (2019)BERT: pre-training of deep bidirectional transformers for language understanding.External Links:1810.04805,LinkCited by:§2.1.
  • D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt (2021)Measuring mathematical problem solving with the math dataset.NeurIPS.Cited by:§5.1.
  • J. Ho, T. Salimans, A. Gritsenko, W. Chan, M. Norouzi, and D. J. Fleet (2022)Video diffusion models.External Links:2204.03458,LinkCited by:§6.
  • E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen (2021)LoRA: low-rank adaptation of large language models.External Links:2106.09685,LinkCited by:§5.2.
  • Z. Huang, Z. Chen, Z. Wang, T. Li, and G. Qi (2025)Reinforcing the diffusion chain of lateral thought with diffusion language models.External Links:2505.10446,LinkCited by:Appendix C.
  • D. Kingma, T. Salimans, B. Poole, and J. Ho (2021)Variational diffusion models.pp. 21696–21707.External Links:LinkCited by:§4.2.1.
  • H. (. Lin and H. Cheng (2025)Gemini achieves gold-level performance at the international collegiate programming contest world finals.Note:https://deepmind.google/discover/blog/gemini-achieves-gold-level-performance-at-the-international-collegiate-programming-contest-world-finals/Accessed: February 1, 2026Cited by:§1.
  • J. Liu, G. Liu, J. Liang, Y. Li, J. Liu, X. Wang, P. Wan, D. Zhang, and W. Ouyang (2025a)Flow-grpo: training flow matching models via online rl.External Links:2505.05470,LinkCited by:§3.3.
  • Z. Liu, C. Chen, W. Li, P. Qi, T. Pang, C. Du, W. S. Lee, and M. Lin (2025b)Understanding r1-zero-like training: a critical perspective.External Links:2503.20783,LinkCited by:§4.1.
  • A. Lou, C. Meng, and S. Ermon (2024)Discrete diffusion modeling by estimating the ratios of the data distribution.InProceedings of the 41st International Conference on Machine LearningThe Thirteenth International Conference on Learning RepresentationsThe Thirteenth International Conference on Learning RepresentationsThe Thirteenth International Conference on Learning RepresentationsAdvances in Neural Information Processing SystemsAdvances in Neural Information Processing SystemsThe Twelfth International Conference on Learning RepresentationsAdvances in Neural Information Processing SystemsInternational Conference on Representation LearningAdvances in Neural Information Processing Systems2nd AI for Math Workshop @ ICML 2025Advances in Neural Information Processing SystemsAdvances in Neural Information Processing Systems,R. Salakhutdinov, Z. Kolter, K. Heller, A. Weller, N. Oliver, J. Scarlett, F. Berkenkamp, M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, J. W. Vaughan, A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, C. Zhang, A. Beygelzimer, Y. Dauphin, P. Liang, J. W. Vaughan, Y. Yue, A. Garg, N. Peng, F. Sha, R. Yu, S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, A. Oh, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, S. Levine, S. Solla, T. Leen, and K. Müller (Eds.),Proceedings of Machine Learning Research, Vol.23534372025353612,pp. 32819–32848.External Links:LinkCited by:§1,§2.1,§2.2,§6.
  • H. Luo, Q. Sun, C. Xu, P. Zhao, J. Lou, C. Tao, X. Geng, Q. Lin, S. Chen, Y. Tang, and D. Zhang (2025)WizardMath: empowering mathematical reasoning for large language models via reinforced evol-instruct.External Links:2308.09583,LinkCited by:§6.
  • T. Luong and E. Lockhart (2025)Advanced version of gemini with deep think officially achieves gold-medal standard at the international mathematical olympiad.Note:https://deepmind.google/discover/blog/advanced-version-of-gemini-with-deep-think-officially-achieves-gold-medal-standard-at-the-international-mathematical-olympiad/Accessed: February 1, 2026Cited by:§1.
  • X. Ma, R. Yu, G. Fang, and X. Wang (2025)DKV-cache: the cache for diffusion language models.External Links:2505.15781,LinkCited by:Appendix F.
  • S. Nie, F. Zhu, Z. You, X. Zhang, J. Ou, J. Hu, J. Zhou, Y. Lin, J. Wen, and C. Li (2025)Large language diffusion models.External Links:2502.09992,LinkCited by:Appendix F,§1,§2.1,§2.2,§2.2,§5.2,§5.
  • OpenAI, :, A. Jaech, A. Kalai, A. Lerer, A. Richardson, A. El-Kishky, A. Low, A. Helyar, A. Madry, A. Beutel, A. Carney, A. Iftimie, A. Karpenko, A. T. Passos, A. Neitz, A. Prokofiev, A. Wei, A. Tam, A. Bennett, A. Kumar, A. Saraiva, A. Vallone, A. Duberstein, A. Kondrich, A. Mishchenko, A. Applebaum, A. Jiang, A. Nair, B. Zoph, B. Ghorbani, B. Rossen, B. Sokolowsky, B. Barak, B. McGrew, B. Minaiev, B. Hao, B. Baker, B. Houghton, B. McKinzie, B. Eastman, C. Lugaresi, C. Bassin, C. Hudson, C. M. Li, C. de Bourcy, C. Voss, C. Shen, C. Zhang, C. Koch, C. Orsinger, C. Hesse, C. Fischer, C. Chan, D. Roberts, D. Kappler, D. Levy, D. Selsam, D. Dohan, D. Farhi, D. Mely, D. Robinson, D. Tsipras, D. Li, D. Oprica, E. Freeman, E. Zhang, E. Wong, E. Proehl, E. Cheung, E. Mitchell, E. Wallace, E. Ritter, E. Mays, F. Wang, F. P. Such, F. Raso, F. Leoni, F. Tsimpourlas, F. Song, F. von Lohmann, F. Sulit, G. Salmon, G. Parascandolo, G. Chabot, G. Zhao, G. Brockman, G. Leclerc, H. Salman, H. Bao, H. Sheng, H. Andrin, H. Bagherinezhad, H. Ren, H. Lightman, H. W. Chung, I. Kivlichan, I. O’Connell, I. Osband, I. C. Gilaberte, I. Akkaya, I. Kostrikov, I. Sutskever, I. Kofman, J. Pachocki, J. Lennon, J. Wei, J. Harb, J. Twore, J. Feng, J. Yu, J. Weng, J. Tang, J. Yu, J. Q. Candela, J. Palermo, J. Parish, J. Heidecke, J. Hallman, J. Rizzo, J. Gordon, J. Uesato, J. Ward, J. Huizinga, J. Wang, K. Chen, K. Xiao, K. Singhal, K. Nguyen, K. Cobbe, K. Shi, K. Wood, K. Rimbach, K. Gu-Lemberg, K. Liu, K. Lu, K. Stone, K. Yu, L. Ahmad, L. Yang, L. Liu, L. Maksin, L. Ho, L. Fedus, L. Weng, L. Li, L. McCallum, L. Held, L. Kuhn, L. Kondraciuk, L. Kaiser, L. Metz, M. Boyd, M. Trebacz, M. Joglekar, M. Chen, M. Tintor, M. Meyer, M. Jones, M. Kaufer, M. Schwarzer, M. Shah, M. Yatbaz, M. Y. Guan, M. Xu, M. Yan, M. Glaese, M. Chen, M. Lampe, M. Malek, M. Wang, M. Fradin, M. McClay, M. Pavlov, M. Wang, M. Wang, M. Murati, M. Bavarian, M. Rohaninejad, N. McAleese, N. Chowdhury, N. Chowdhury, N. Ryder, N. Tezak, N. Brown, O. Nachum, O. Boiko, O. Murk, O. Watkins, P. Chao, P. Ashbourne, P. Izmailov, P. Zhokhov, R. Dias, R. Arora, R. Lin, R. G. Lopes, R. Gaon, R. Miyara, R. Leike, R. Hwang, R. Garg, R. Brown, R. James, R. Shu, R. Cheu, R. Greene, S. Jain, S. Altman, S. Toizer, S. Toyer, S. Miserendino, S. Agarwal, S. Hernandez, S. Baker, S. McKinney, S. Yan, S. Zhao, S. Hu, S. Santurkar, S. R. Chaudhuri, S. Zhang, S. Fu, S. Papay, S. Lin, S. Balaji, S. Sanjeev, S. Sidor, T. Broda, A. Clark, T. Wang, T. Gordon, T. Sanders, T. Patwardhan, T. Sottiaux, T. Degry, T. Dimson, T. Zheng, T. Garipov, T. Stasi, T. Bansal, T. Creech, T. Peterson, T. Eloundou, V. Qi, V. Kosaraju, V. Monaco, V. Pong, V. Fomenko, W. Zheng, W. Zhou, W. McCabe, W. Zaremba, Y. Dubois, Y. Lu, Y. Chen, Y. Cha, Y. Bai, Y. He, Y. Zhang, Y. Wang, Z. Shao, and Z. Li (2024)OpenAI o1 system card.External Links:2412.16720,LinkCited by:§1.
  • J. Ou, S. Nie, K. Xue, F. Zhu, J. Sun, Z. Li, and C. Li (2025)Your absorbing discrete diffusion secretly models the conditional distributions of clean data.External Links:LinkCited by:§1,§2.1.
  • L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. F. Christiano, J. Leike, and R. Lowe (2022)Training language models to follow instructions with human feedback.pp. 27730–27744.External Links:LinkCited by:§6.
  • J. Pan, J. Zhang, X. Wang, L. Yuan, H. Peng, and A. Suhr (2025)TinyZero.Note:https://github.com/Jiayi-Pan/TinyZeroAccessed: 2025-01-24Cited by:§5.1.
  • W. Peebles and S. Xie (2022)Scalable diffusion models with transformers.arXiv preprint arXiv:2212.09748.Cited by:Appendix F,§6.
  • R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn (2023)Direct preference optimization: your language model is secretly a reward model.pp. 53728–53741.External Links:LinkCited by:Appendix C,§6.
  • A. Z. Ren, J. Lidard, L. L. Ankile, A. Simeonov, P. Agrawal, A. Majumdar, B. Burchfiel, H. Dai, and M. Simchowitz (2024)Diffusion policy policy optimization.External Links:2409.00588,LinkCited by:§6.
  • S. S. Sahoo, M. Arriola, Y. Schiff, A. Gokaslan, E. Marroquin, J. T. Chiu, A. Rush, and V. Kuleshov (2024)Simple and effective masked diffusion language models.pp. 130136–130184.External Links:LinkCited by:§1,§2.1,§4.2.1.
  • J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov (2017)Proximal policy optimization algorithms.External Links:1707.06347,LinkCited by:§3.1,§4.1.
  • J. Schulman (2020)Approximating kl divergence.Note:http://joschu.net/blog/kl-approx.htmlBlog postCited by:§4.1.
  • Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y. K. Li, Y. Wu, and D. Guo (2024)DeepSeekMath: pushing the limits of mathematical reasoning in open language models.External Links:2402.03300,LinkCited by:§1,§4.1,§4.1,Remark 4.1.
  • J. Sohl-Dickstein, E. A. Weiss, N. Maheswaranathan, and S. Ganguli (2015)Deep unsupervised learning using nonequilibrium thermodynamics.External Links:1503.03585,LinkCited by:§6.
  • R. S. Sutton, D. McAllester, S. Singh, and Y. Mansour (1999)Policy gradient methods for reinforcement learning with function approximation.pp..External Links:LinkCited by:§3.1.
  • X. Tang, R. Dolga, S. Yoon, and I. Bogunovic (2025)Wd1: weighted policy optimization for reasoning in diffusion language models.External Links:2507.08838,LinkCited by:§B.1,§B.2,Appendix C,§5.1,§6.
  • J. Vendrow, E. Vendrow, S. Beery, and A. Madry (2025)Do large language model benchmarks test reliability?.External Links:2502.03461,LinkCited by:§B.2.
  • L. von Werra, Y. Belkada, L. Tunstall, E. Beeching, T. Thrush, N. Lambert, S. Huang, K. Rasul, and Q. Gallouédec (2020)TRL: transformer reinforcement learning.GitHub.Note:https://github.com/huggingface/trlCited by:§B.1.
  • C. Wang, P. Rashidinejad, D. Su, S. Jiang, S. Wang, S. Zhao, C. Zhou, S. Z. Shen, F. Chen, T. Jaakkola, Y. Tian, and B. Liu (2025a)SPG: sandwiched policy gradient for masked diffusion language models.arXiv preprint arXiv:2510.09541.Cited by:§B.2,§B.2,Table 3,Appendix C,Remark 3.1,§5.1,§5.3,§6.
  • S. Wang, L. Yu, C. Gao, C. Zheng, S. Liu, R. Lu, K. Dang, X. Chen, J. Yang, Z. Zhang, Y. Liu, A. Yang, A. Zhao, Y. Yue, S. Song, B. Yu, G. Huang, and J. Lin (2025b)Beyond the 80/20 rule: high-entropy minority tokens drive effective reinforcement learning for llm reasoning.External Links:2506.01939,LinkCited by:§4.2.2.
  • R. J. Williams (1992)Simple statistical gradient-following algorithms for connectionist reinforcement learning.Machine Learning8(3),pp. 229–256.External Links:Document,ISBN 1573-0565,LinkCited by:§3.1.
  • C. Wu, H. Zhang, S. Xue, Z. Liu, S. Diao, L. Zhu, P. Luo, S. Han, and E. Xie (2025)Fast-dllm: training-free acceleration of diffusion llm by enabling kv cache and parallel decoding.External Links:2505.22618,LinkCited by:Appendix F.
  • W. Xiong, J. Yao, Y. Xu, B. Pang, L. Wang, D. Sahoo, J. Li, N. Jiang, T. Zhang, C. Xiong, and H. Dong (2025)A minimalist approach to llm reasoning: from rejection sampling to reinforce.External Links:2504.11343,LinkCited by:§6.
  • L. Yang, Y. Tian, B. Li, X. Zhang, K. Shen, Y. Tong, and M. Wang (2025)MMaDA: multimodal large diffusion language models.External Links:2505.15809,LinkCited by:Appendix F,§1.
  • F. Yao, L. Liu, D. Zhang, C. Dong, J. Shang, and J. Gao (2025)Your efficient rl framework secretly brings you off-policy rl training.External Links:LinkCited by:Appendix E.
  • Q. Yu, Z. Zhang, R. Zhu, Y. Yuan, X. Zuo, Y. Yue, W. Dai, T. Fan, G. Liu, L. Liu, X. Liu, H. Lin, Z. Lin, B. Ma, G. Sheng, Y. Tong, C. Zhang, M. Zhang, W. Zhang, H. Zhu, J. Zhu, J. Chen, J. Chen, C. Wang, H. Yu, Y. Song, X. Wei, H. Zhou, J. Liu, W. Ma, Y. Zhang, L. Yan, M. Qiao, Y. Wu, and M. Wang (2025)DAPO: an open-source llm reinforcement learning system at scale.External Links:2503.14476,LinkCited by:§B.1,Remark G.1,§6.
  • O. Zekri and N. Boullé (2025)Fine-tuning discrete diffusion models with policy gradient methods.External Links:2502.01384,LinkCited by:§6.
  • S. Zhao, D. Gupta, Q. Zheng, and A. Grover (2025)D1: scaling reasoning in diffusion large language models via reinforcement learning.External Links:2504.12216,LinkCited by:§B.2,§B.2,Table 3,§5.1,§5.3,§6.
  • K. Zheng, Y. Chen, H. Mao, M. Liu, J. Zhu, and Q. Zhang (2025)Masked diffusion models are secretly time-agnostic masked models and exploit inaccurate categorical sampling.External Links:LinkCited by:Appendix D,Appendix E,§4.2.1.
  • F. Zhu, R. Wang, S. Nie, X. Zhang, C. Wu, J. Hu, J. Zhou, J. Chen, Y. Lin, J. Wen, and C. Li (2025)LLaDA 1.5: variance-reduced preference optimization for large language diffusion models.External Links:2505.19223,LinkCited by:Appendix C,§5.3.
  • D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P. Christiano, and G. Irving (2020)Fine-tuning language models from human preferences.External Links:1909.08593,LinkCited by:§6.

Appendix AProof of Unbiased Estimator

Using the notation of Section4.1, define the entropy scoreete_{t}as the Shannon entropy ofπ​(ot∣q,ot−1)\pi(o_{t}\mid q,o_{t-1}), i.e. the sum of the marginal entropies of the positions unmasked at steptt. Lete¯={et/∑et}t=1m\bar{e}=\{e_{t}/\sum e_{t}\}_{t=1}^{m}be the vector of normalized entropy scores, and supposeT∼Cat⁡(e¯)T\sim\operatorname{Cat}(\bar{e}). Then the full AGRPO estimator (without clipping) is

𝒥^​(θ;T)=1m​G​e¯T​∑i=1Gπθ​(oTi∣q,oT−1i)πo​l​d​(oTi∣q,oT−1i)​Ai.\widehat{\mathcal{J}}(\theta;T)=\frac{1}{mG\bar{e}_{T}}\sum_{i=1}^{G}\frac{\pi_{\theta}(o^{i}_{T}\mid q,o^{i}_{T-1})}{\pi_{old}(o^{i}_{T}\mid q,o^{i}_{T-1})}A_{i}.

Theorem A.1.

𝔼​[𝒥^​(θ;T)]=𝒥​(θ)\mathbb{E}[\widehat{\mathcal{J}}(\theta;T)]=\mathcal{J}(\theta), where𝒥​(θ)\mathcal{J}(\theta)is defined in Equation (4) and the expectation is taken w.r.t. randomness ofTT.

Proof.

SinceTTis a categorical random variable, we can write out the expectation explicitly:

𝔼​[𝒥^​(θ)]\displaystyle\mathbb{E}[\widehat{\mathcal{J}}(\theta)]=∑t=1mPr⁡[T=t]​1m​G​e¯t​∑i=1Gπθ​(oti∣q,ot−1i)πo​l​d​(oti∣q,ot−1i)​Ai\displaystyle=\sum_{t=1}^{m}\Pr[T=t]\frac{1}{mG\bar{e}_{t}}\sum_{i=1}^{G}\frac{\pi_{\theta}(o^{i}_{t}\mid q,o^{i}_{t-1})}{\pi_{old}(o^{i}_{t}\mid q,o^{i}_{t-1})}A_{i}=∑t=1met∑t′=1met′​∑t′=1met′m​G​et​∑i=1Gπθ​(oti∣q,ot−1i)πo​l​d​(oti∣q,ot−1i)​Ai\displaystyle=\sum_{t=1}^{m}\frac{e_{t}}{\sum_{t^{\prime}=1}^{m}e_{t^{\prime}}}\frac{\sum_{t^{\prime}=1}^{m}e_{t^{\prime}}}{mGe_{t}}\sum_{i=1}^{G}\frac{\pi_{\theta}(o^{i}_{t}\mid q,o^{i}_{t-1})}{\pi_{old}(o^{i}_{t}\mid q,o^{i}_{t-1})}A_{i}=1m​G​∑t=1m∑i=1Gπθ​(oti∣q,ot−1i)πo​l​d​(oti∣q,ot−1i)​Ai=𝒥​(θ).\displaystyle=\frac{1}{mG}\sum_{t=1}^{m}\sum_{i=1}^{G}\frac{\pi_{\theta}(o^{i}_{t}\mid q,o^{i}_{t-1})}{\pi_{old}(o^{i}_{t}\mid q,o^{i}_{t-1})}A_{i}=\mathcal{J}(\theta).∎

Appendix BExperiment Details

B.1Training

Models are trained with LoRA rankr=64r=64on 8xH100 GPUs for 400-900 steps (depending on the task). We use Hugging Face’s TRL framework(von Werraet al.,2020)to implement the training code. Rewards are a combination of rule-based formatting rewards, i.e. the presence of<think>and<answer>tags, as well as binary ground-truth rewards (except for Sudoku, where the reward is the fraction of empty cells filled in correctly, which is dense).

For batched rollouts, we use a group size ofG=8G=8per prompt and 16 prompts per update, resulting in a global batch size of 128. The clipping threshold is set toε=0.2\varepsilon=0.2. The number of gradient updates per batch of rolloutsμ\mu, which controls the “off-policyness” of the algorithm, is set toμ=1\mu=1. As mentioned previously, we also assumeβ=0\beta=0, so the algorithm is fully on-policy with no KL penalty. Consistent withYuet al.(2025)andTanget al.(2025), we don’t observe any drawbacks in terms of convergence speed or stability with this configuration.

Since AGRPO involves multiple forward passes per gradient step, it incurs a higher wall-clock cost per update than one-step methods like diffu-GRPO and SPG (see Table2). However, because AGRPO updates the policy based on a wide range of intermediate states rather than only the final sequence, it provides a denser learning signal and therefore improved sample efficiency (i.e. requires fewer rollouts). Empirically, we observe rapid convergence to high reward levels across all tasks (see e.g. Figures3(a)and4), suggesting that the additional compute per step is offset by higher quality gradient updates.

Table 2:Average wall time per gradient step for different dLLM RL methods on a controlled GSM8K training run withn=256n=256,m=128m=128,G=4G=4, andμ=2\mu=2. Values are reported on 2xRTX6000 Ada GPUs with a global batch size of 16.

B.2Evaluation

For all tasks and train/test splits, we use open source datasets from HuggingFace or GitHub. Accuracies for GSM8K are reported on the GSM8K-Platinum(Vendrowet al.,2025)test split, a cleaned version of the original GSM8K test split. Surprisingly, we found that models trained on GSM8K problems achieved higher accuracy on MATH-500 test problems than models trained on the MATH training split. The accuracy reported in Table1is thus from the GSM8K model; models trained on MATH showed more modest gains, achieving 38.0% atn=256n=256. We hypothesize that the more difficult nature of MATH, combined with inherent limitations of the base LLaDA model and limited context window, make it harder to learn generalizable reasoning abilities during RL.

Results for diffu-GRPO, wd1, and SPG are taken from the respective papers(Zhaoet al.,2025; Tanget al.,2025; Wanget al.,2025a). We note that our reproduced baselines for LLaDA-8B-Instruct and LLaDA 1.5 are slightly higher than baselines reported in the aforementioned works (and that there is variation between these works as well). We attribute this to our use of a more amenable system prompt and robust parsing system (via symbolic libraries), which we applied consistently across all tasks to ensure a fair comparison.

For completeness, we provide a full comparison of reported LLaDA-8B-Instruct baselines againstZhaoet al.(2025)andWanget al.(2025a)in Table3.

Table 3:Comparison of base model accuracies reported in previous work vs. reproduced in our work.

Appendix CComparisons to Previous dLLM Post-Training Methods

Here we continue our discussion of dLLM post-training methods presented in Section6.

wd1.Tanget al.(2025)build on diffu-GRPO by leveraging the closed-form solutionπ∗\pi^{*}of reverse KL-regularized policy optimization, namelyπ∗​(a∣s)∝πref​(a∣s)​exp⁡(r​(s,a)/β)\pi^{*}(a\mid s)\propto\pi_{\operatorname{ref}}(a\mid s)\exp(r(s,a)/\beta)whereβ\betais the reverse KL coefficient. They use this to rewrite the PG objective in way that simplifies the importance sampling ratioπθ(⋅∣q)/πo​l​d(⋅∣q)\pi_{\theta}(\cdot\mid q)/\pi_{old}(\cdot\mid q). However, their approach inherits from d1 the same reliance on one-step approximations, limiting practical effectiveness.

VRPO.Introduced byZhuet al.(2025)as part of the training process for LLaDA 1.5, Variance-Reduced Policy Optimization follows a similar approach to DPO(Rafailovet al.,2023), where the model directly learns from pairwise preference data without external rewards. The VRPO objective consists of the DPO objective with all sequence likelihoods replaced by ELBOs, along with several tricks to reduce variance. Overall, the LLaDA 1.5 paper studies a slightly different RL setting (alignment/RLHF instead of reasoning/RLVR) and also relies heavily on likelihood bounds.

DCoLT.Inspired by autoregressive chain-of-thought methods,Huanget al.(2025)optimize a Diffusion Chain of Lateral Thought. Similar to our approach, they adopt a multi-step denoising view where actions consists of individual unmasking steps rather than the final sequence. However, there are several key differences that make AGRPO training much simpler: DCoLT requires training a separate module to rank which positions to unmask (according to a Plackett-Luce model) and accumulates gradients overallmmtimestepsinstead of using MC sampling to estimate gradients efficiently.

SPG.Most recently,Wanget al.(2025a)propose Sandwiched Policy Gradients, which combines an ELBO-like lower bound on likelihood with an evidence upper bound (EUBO). They identify an issue with PG objectives that doesn’t appear in dLLM pretraining, namely the fact that likelihoods can be multiplied by a negative advantage estimate, which breaks the validity of optimizing a lower bound. By optimizing the EUBO in place of the standard ELBO for negative advantage trajectories, they achieve stronger empirical results (although their policy updates remain biased due to the sequence-level framing).

Appendix DLow-discrepancy Sampling Details

An outline ofZhenget al.(2025)’s discrete low-discrepancy sampler is given below. The goal is drawkksamples fromT∼{0,…,m−1}T\sim\{0,\dots,m-1\}.

  1. 1.SampleU0,…,Uk−1U_{0},\dots,U_{k-1}i.i.d. fromUnif⁡([0,1])\operatorname{Unif}([0,1]).
  2. 2.“Bin” them intokkdisjoint bins by definingUj′=(Uj+j)/kU_{j}^{\prime}=(U_{j}+j)/kforj=0j=0tok−1k-1.
  3. 3.Define the final set ofkksamples{Tj}j=0k−1\{T_{j}\}_{j=0}^{k-1}asTj=⌊m​Uj′⌋T_{j}=\lfloor mU_{j}^{\prime}\rfloor.

Intuitively, the sampler divides the interval[0,1][0,1]intokkequal subintervals, samples from each one uniformly and shuffles the results, and then scales by a factor ofmmand floors. In AGRPO, this allows us to “cover” a wide range of timesteps from 1 tommwhen estimating the PG objective. We can also extend this to the batch level: instead of samplingkktimesteps, we sampleb​kbktimesteps wherebbis the batch size.

To use low-discrepancy sampling in conjunction with entropy importance sampling, we replace Step 3 by an inverse CDF transform to turn samples fromUnif⁡([0,1])\operatorname{Unif}([0,1])into categorical samples.

Appendix EPractical Considerations

In this section, we discuss in detail various decisions and tradeoffs made regarding the actual implementation of AGRPO. As with all online RL algorithms, the goal of any implementation is to run as efficiently as possible while maximizing efficacy. We plan to release training and evaluation code publicly to showcase implementation details and encourage reproducibility.

Caching partially masked states.In order to obtain exact action/token likelihoods, we must recreate the exact state/context at the step where that token was unmasked. To do this efficiently, we cache the unmasking order during generation so that each token is associated with a timesteptt. Then, to get the partially masked state at timesteptt, we simply mask out all tokens with timestep≥t\geq t.

Memory-efficient gradient accumulation.Naively computing the AGRPO objective (5) and backpropagating would require keepingkkforward passes simultaneously in memory, which can quickly saturate GPU memory for 8B scale models. Instead, one can accumulate the gradient immediately after each MC sample is computed by callingloss.backward()(without taking an optimizer step)insidethe for loop. This frees the computational graph and avoids excess memory usage.

float64 Gumbel-based categorical sampling.dLLMs typically use the Gumbel-max trick to sample from output logits. However,Zhenget al.(2025)point out that naively usingfloat32causes an inconsistency between theoretical and actual behavior due to floating-point precision. We follow their recommendation of usingfloat64for the sampling stage.

Handling EOS tokens.Since dLLMs generate with a fixed number of sampling steps, in later timesteps, the model can spend many “garbage” steps producingEOStokens at the end of a sequence (while other sequences in the batch are still generating useful tokens). Gradient updates on these steps don’t provide meaningful information to the model, so we set the max timestep in our low-discrepancy sampler (SectionD) to be the last timestep where a non-EOStoken was generated.

Training/inference mismatches.Yaoet al.(2025)observe that for RL with traditional LLMs, mismatches between the inference engine and training code can cause subtle “off-policyness” and instability issues, especially when combined with bf16 precision. For AGRPO, we deploy a custom inference engine for LLaDA, making sure to track things such as the top-p/top-k sampling masks and reuse them in training, as suggested byDeepSeek-AIet al.(2025b).

Appendix FRemarks on dLLM Inference

Despite a more complicated training setup, dLLMs enjoy several potential benefits at inference time: they can generate text in arbitrary order, are naturally self-speculative (i.e. one can see the model’s best guess for the entire sequence at every step), and can trade off compute and generation quality by choosing to unmask more or less tokens per step.

Here we discuss several unique characteristics of dLLM inference, which may be helpful to readers who are only familiar with traditional AR LLM inference.

Fixed sequence length.One drawback of dLLMs is that the context length must be fixed ahead of time, instead of being dynamically grown as with AR LLMs. Works such as BlockDiff address this issue by introducing a hybrid autoregressive/diffusion framework(Arriolaet al.,2025); in this paper, we stay within the normal diffusion framework for simplicity.

Instruct-tuned models.dLLMs such as LLaDA-8B-Instruct, which have undergone supervised fine-tuning (SFT) on instruction-following traces, tend to place higher probabilities onEOStokens. When combined with confidence-based unmasking, this leads to an unnaturally high proportion ofEOStokens in later positions and terse, stilted responses(Nieet al.,2025). Thus, to generate text with standard left-to-right prompting, we divide the response into smaller blocks, unmask tokens within the leftmost block, and continue to the next block once all tokens in the current block have been unmasked. This is known assemi-autoregressivesampling(Yanget al.,2025). During training, we use semi-AR sampling with block length 12.

Bidirectional prompting.In this paper, we work with traditional left-to-right prompting, which is the native format for the reasoning datasets we use. This leaves a big dLLM advantage on the table — namely their ability to generate text from arbitrary context. Future works could consider reasoning tasks that involve using context from both the left and right; for example, giving the model a problem and the numerical answer, and forcing it to deduce the intermediate steps.

KV caching.Since diffusion transformers typically use full self-attention instead of causal self-attention(Peebles and Xie,2022), embeddings for the same token position can change over the sampling process. This non-causality prevents dLLMs from using the same KV caching mechanism as AR LLMs. As a result, generating same-quality text with dLLMs is significantly slower than same-scale AR models, which is especially painful for online RL. However, there has been some recent interest in KV caching alternatives for dLLMs(Wuet al.,2025; Maet al.,2025). Since dLLMs already have the ability to decode multiple tokens in parallel, we believe a successful implementation of KV caching is imperative to realizing dLLMs’ potential as a faster, more flexible alternative to AR models.

Appendix GSample Responses

See Table4for the system prompt used in GSM8K training as well as sample responses from a chosen test split problem. As evident in the responses, RL induces a rigid step-by-step structure, which we hypothesize allows the model to arrive at the same solution with many less steps (m=48m=48vsm=192m=192). could explain the results in Section5.3.1

Even though both responses arrive at the correct answer, them=192m=192solution is more coherent, while them=48m=48solution contains clear artifacts (i.e. repeated words, grammar mistakes). Note that the model is only trained onm=128m=128, so such artifacts are technically out-of-distribution and could ostensibly “derail” the model. Our results show the contrary, however, suggesting that learned reasoning abilities are robust to such perturbations.

Table 4:Sample responses (n=384n=384) for different values ofmm(# sampling steps) from a model trained with AGRPO. Problem is from the GSM8K-Platinum dataset.

Similar Articles

DACA-GRPO: Denoising-Aware Credit Assignment for Reinforcement Learning in Diffusion Language Models

arXiv cs.LG

This paper identifies weaknesses in existing reinforcement learning methods for diffusion language models—lack of temporal credit assignment and biased likelihood estimates—and proposes DACA-GRPO, a plug-and-play enhancement that introduces denoising progress scores and stratified masking likelihood, achieving consistent improvements across reasoning, code generation, and constrained generation benchmarks.