Simplex Relaxation for Discrete Diffusion

arXiv cs.CL Papers

Summary

This paper introduces Simplax, an exact Dirichlet–categorical augmentation for discrete diffusion models that enriches training objectives and reverse transitions while preserving the original categorical corruption process, improving perplexity–entropy tradeoff on OpenWebText and validity on Sudoku.

arXiv:2608.10615v1 Announce Type: new Abstract: Discrete diffusion models for categorical generation are defined by a corruption kernel, which determines the intermediate state space and the associated reverse prediction problem. We study uniform discrete diffusion and ask whether its training objective and reverse transitions can be enriched without changing the underlying categorical corruption process. We introduce Simplax, an exact Dirichlet--categorical augmentation that couples each corrupted categorical state with an auxiliary simplex-valued variable while preserving the original uniform diffusion process as its categorical marginal. This augmentation yields a tractable Rao--Blackwellized reverse-bridge objective and a corresponding stochastic reverse sampler, while retaining the corrupted categorical state as the denoiser input. Empirically, Simplax improves the generative perplexity--entropy tradeoff on unconditional OpenWebText generation. On Sudoku, a model trained exclusively on $30$-clue puzzles achieves the highest accuracy among the compared methods across all evaluated clue densities, including the minimum uniquely solvable $17$-clue regime, and also achieves the highest validity in unconditional generation.
Original Article
View Cached Full Text

Cached at: 08/12/26, 08:36 AM

# Simplex Relaxation for Discrete Diffusion
Source: [https://arxiv.org/html/2608.10615](https://arxiv.org/html/2608.10615)
Jinya Sakurai1 2 4Patrick Pynadath3Satoshi Hayakawa2 Jaehong Yoon1Xulei Yang4Nancy F\. Chen4,5Xun Xu4,5 1NTU Singapore2The University of Tokyo 3Purdue University4Institute for Advanced Intelligence and Computing \(IAIC\), A\*STAR 5Centre for Frontier AI Research \(CFAR\), A\*STAR

###### Abstract

Discrete diffusion models for categorical generation are defined by a corruption kernel, which determines the intermediate state space and the associated reverse prediction problem\. We study uniform discrete diffusion and ask whether its training objective and reverse transitions can be enriched without changing the underlying categorical corruption process\. We introduce Simplax, an exact Dirichlet–categorical augmentation that couples each corrupted categorical state with an auxiliary simplex\-valued variable while preserving the original uniform diffusion process as its categorical marginal\. This augmentation yields a tractable Rao–Blackwellized reverse\-bridge objective and a corresponding stochastic reverse sampler, while retaining the corrupted categorical state as the denoiser input\. Empirically, Simplax improves the generative perplexity–entropy tradeoff on unconditional OpenWebText generation\. On Sudoku, a model trained exclusively on3030\-clue puzzles achieves the highest accuracy among the compared methods across all evaluated clue densities, including the minimum uniquely solvable1717\-clue regime, and also achieves the highest validity in unconditional generation\.

## 1Introduction

![Refer to caption](https://arxiv.org/html/2608.10615v1/x1.png)Figure 1:Reverse\-bridge matching, illustrated forK=3K=3: one\-hot states are vertices ofΔK−1\\Delta^\{K\-1\}and𝐰t\\mathbf\{w\}\_\{t\}is an interior point\. \(a\) The corrupted state𝐳t\\mathbf\{z\}\_\{t\}is both the denoiser input and the sole bridge anchor, so a single bridge is matched per sample\. \(b\) Simplax samples\(𝐰t,𝐳t\)\(\\mathbf\{w\}\_\{t\},\\mathbf\{z\}\_\{t\}\)jointly;𝐳t\\mathbf\{z\}\_\{t\}remains the denoiser input, while extra draws𝐳~t∼Cat⁡\(𝐰t\)\\widetilde\{\\mathbf\{z\}\}\_\{t\}\\sim\\operatorname\{Cat\}\(\\mathbf\{w\}\_\{t\}\)anchor allKKbridges, which are matched at once and marginalized exactly to give \([15](https://arxiv.org/html/2608.10615#S3.E15)\)\.Discrete diffusion models have emerged as a promising framework for generative modeling over categorical data, including text, biological sequences, and other symbolic domains\(Hoogeboomet al\.,[2021](https://arxiv.org/html/2608.10615#bib.bib1); Austinet al\.,[2021](https://arxiv.org/html/2608.10615#bib.bib2); Campbellet al\.,[2022](https://arxiv.org/html/2608.10615#bib.bib3); Louet al\.,[2024](https://arxiv.org/html/2608.10615#bib.bib4); Zhanget al\.,[2025](https://arxiv.org/html/2608.10615#bib.bib11)\)\. Compared with autoregressive generation, they offer a conceptually different route to parallel prediction by learning to invert a noising process on full categorical states\. A central design choice in this framework is the corruption kernel\. This choice determines the intermediate state space and the semantics of reverse updates, and shapes the form of the training objective used to approximate the reverse process\(Austinet al\.,[2021](https://arxiv.org/html/2608.10615#bib.bib2); Campbellet al\.,[2022](https://arxiv.org/html/2608.10615#bib.bib3); Louet al\.,[2024](https://arxiv.org/html/2608.10615#bib.bib4)\)\.

Existing discrete diffusion models instantiate this design choice in different ways\. Masked diffusion corrupts tokens toward a distinguished mask state, yielding intermediate sequences that can be interpreted as partially observed data\(Austinet al\.,[2021](https://arxiv.org/html/2608.10615#bib.bib2); Sahooet al\.,[2024](https://arxiv.org/html/2608.10615#bib.bib5); Shiet al\.,[2024](https://arxiv.org/html/2608.10615#bib.bib6); Ouet al\.,[2025](https://arxiv.org/html/2608.10615#bib.bib8); Zhenget al\.,[2025](https://arxiv.org/html/2608.10615#bib.bib9)\)\. Uniform diffusion instead replaces tokens toward the uniform distribution over the original categorical alphabet, treating all categories symmetrically without introducing a distinguished absorbing state\(Austinet al\.,[2021](https://arxiv.org/html/2608.10615#bib.bib2); Schiffet al\.,[2025](https://arxiv.org/html/2608.10615#bib.bib7); Sahooet al\.,[2025](https://arxiv.org/html/2608.10615#bib.bib10); Deschenauxet al\.,[2026](https://arxiv.org/html/2608.10615#bib.bib17)\)\. Recent work has studied uniform diffusion in connection with guidance, repeated token revision, few\-step generation, self\-correction, and scaling\(Schiffet al\.,[2025](https://arxiv.org/html/2608.10615#bib.bib7); Sahooet al\.,[2025](https://arxiv.org/html/2608.10615#bib.bib10); von Rütteet al\.,[2025](https://arxiv.org/html/2608.10615#bib.bib44),[2026](https://arxiv.org/html/2608.10615#bib.bib42); Sahooet al\.,[2026](https://arxiv.org/html/2608.10615#bib.bib43)\)\.

These developments leave open a methodological question for uniform discrete diffusion: can one enrich its training objectives and samplers while keeping the categorical corruption process unchanged? In standard uniform diffusion, the reverse update between two noise levels is expressed directly through sampled categorical states\. This preserves the discrete generative process, but it also means that both training and sampling are formulated through categorical intermediate states\. We ask whether an auxiliary probabilistic structure can be introduced around these transitions so that tractable objectives and samplers can be derived without changing the forward process itself\.

In this work, we consider a probabilistic augmentation of uniform discrete diffusion that leaves its categorical corruption process unchanged while introducing an auxiliary simplex\-valued state\. We introduce Simplax, an exact Dirichlet–categorical augmentation in which each corrupted categorical state𝐳t\\mathbf\{z\}\_\{t\}is coupled to an auxiliary simplex variable𝐰t\\mathbf\{w\}\_\{t\}through a shifted Dirichlet conditional\. The resulting augmented hierarchy preserves the original uniform diffusion process as its categorical marginal and admits the exact decoder

q​\(𝐳t∣𝐰t\)=Cat⁡\(𝐳t;𝐰t\)\.q\(\\mathbf\{z\}\_\{t\}\\mid\\mathbf\{w\}\_\{t\}\)=\\operatorname\{Cat\}\(\\mathbf\{z\}\_\{t\};\\mathbf\{w\}\_\{t\}\)\.Thus, the simplex variable is not introduced as a replacement for the categorical state, but as an auxiliary random variable that is probabilistically coupled to it and can be used to construct the training objective and reverse transition\.

A direct construction of the reverse objective in the augmented space is complicated by the fact that the induced simplex reverse bridges are mixtures of shifted Dirichlet components, whose KL divergence is generally intractable\. We therefore derive a categorical reverse\-bridge surrogate by averaging the standard discrete reverse KL divergence over an auxiliary categorical decode from𝐰t\\mathbf\{w\}\_\{t\}\. This expectation admits an analytic Rao–Blackwellized form\. We further derive a stochastic ancestral sampler from the same augmented hierarchy, using the denoiser prediction together with the auxiliary simplex state to parameterize the reverse update\.

We evaluate Simplax on unconditional text generation with OpenWebText and constrained categorical generation with Sudoku\. On OpenWebText, Simplax achieves favorable generative perplexity–entropy tradeoffs across a wide range of inference budgets, outperforming the compared methods at most reported operating points\. On Sudoku, all models are trained exclusively on puzzles with3030clues and evaluated in\-distribution, under transfer to both easier and harder clue densities, and in unconditional generation\. Simplax achieves the highest performance among the compared methods across all evaluated Sudoku settings\.

Our main contributions are as follows\. First, we introduce an exact Dirichlet–categorical augmentation of uniform discrete diffusion that preserves the original categorical process as a marginal\. Second, we derive a tractable categorical reverse\-bridge surrogate whose auxiliary categorical expectation admits a Rao–Blackwellized closed form, together with a stochastic ancestral sampler derived from the same augmented hierarchy\. Third, we empirically demonstrate that the resulting framework improves generation across both open\-ended text modeling and constrained categorical generation, including broad inference\-budget regimes and distribution shifts in Sudoku\.

## 2Preliminaries

#### Notation\.

Let𝒱=\{𝐱∈\{0,1\}K:∑i=1Kxi=1\}\\mathcal\{V\}=\\\{\\mathbf\{x\}\\in\\\{0,1\\\}^\{K\}\\;:\\;\\sum\_\{i=1\}^\{K\}x\_\{i\}=1\\\}denote the set of one\-hot vectors overKKcategories, and letΔK−1\\Delta^\{K\-1\}denote the probability simplex overKKcategories\. We represent scalar discrete random variables takingKKvalues as one\-hot column vectors in𝒱\\mathcal\{V\}\. We writeCat⁡\(⋅;𝝅\)\\operatorname\{Cat\}\(\\cdot;\\bm\{\\pi\}\)for the categorical distribution with class probabilities𝝅∈ΔK−1\\bm\{\\pi\}\\in\\Delta^\{K\-1\}\. Results involving Dirichlet densities assumeπk\>0\\pi\_\{k\}\>0for every categorykk; the uniform base distribution used in our experiments satisfies this condition\. We writeDir⁡\(⋅;𝜶\)\\operatorname\{Dir\}\(\\cdot;\\bm\{\\alpha\}\)for the Dirichlet distribution with concentration vector𝜶∈ℝ\>0K\\bm\{\\alpha\}\\in\\mathbb\{R\}\_\{\>0\}^\{K\}\. We use𝟏∈ℝK\\mathbf\{1\}\\in\\mathbb\{R\}^\{K\}for the all\-ones vector,⟨𝐚,𝐛⟩\\left\\langle\\mathbf\{a\},\\,\\mathbf\{b\}\\right\\ranglefor the inner product,𝐚⊙𝐛\\mathbf\{a\}\\odot\\mathbf\{b\}for the Hadamard product, and𝐚⊘𝐛\\mathbf\{a\}\\oslash\\mathbf\{b\}for elementwise division\. For sequences of lengthLL, we write𝐱1:L∈𝒱L\\mathbf\{x\}^\{1:L\}\\in\\mathcal\{V\}^\{L\}\.

### 2\.1Dirichlet distribution

The Dirichlet distribution is a distribution over the probability simplex\. Its mean is given by the normalized concentration vector, while the sum of the concentration parameters controls how concentrated the distribution is around its mean\. In particular, for𝐩∈ΔK−1\\mathbf\{p\}\\in\\Delta^\{K\-1\}andη\>0\\eta\>0,Dir⁡\(⋅;η​𝐩\)\\operatorname\{Dir\}\(\\cdot;\\eta\\mathbf\{p\}\)denotes the Dirichlet distribution centered at𝐩\\mathbf\{p\}, withη\\etacontrolling its concentration\.

### 2\.2Discrete diffusion

We consider a discrete diffusion process with prior𝝅∈ΔK−1\\bm\{\\pi\}\\in\\Delta^\{K\-1\}and noise scheduleαt∈\[0,1\]\\alpha\_\{t\}\\in\[0,1\]\. Following the standard parameterization, the noisy categorical state at timettis distributed as

q​\(𝐳t∣𝐱\)=Cat⁡\(𝐳t;𝐩t​\(𝐱\)\),𝐩t​\(𝐱\)≔αt​𝐱\+\(1−αt\)​𝝅\.q\(\\mathbf\{z\}\_\{t\}\\mid\\mathbf\{x\}\)=\\operatorname\{Cat\}\\\!\\left\(\\mathbf\{z\}\_\{t\};\\mathbf\{p\}\_\{t\}\(\\mathbf\{x\}\)\\right\),\\quad\\mathbf\{p\}\_\{t\}\(\\mathbf\{x\}\)\\coloneq\\alpha\_\{t\}\\mathbf\{x\}\+\(1\-\\alpha\_\{t\}\)\\bm\{\\pi\}\.\(1\)
We use a schedule withα0=1\\alpha\_\{0\}=1,α1=0\\alpha\_\{1\}=0, andαt<1\\alpha\_\{t\}<1for everyt\>0t\>0\. Hence𝐩t​\(𝐱\)\\mathbf\{p\}\_\{t\}\(\\mathbf\{x\}\)is strictly positive fort\>0t\>0\.

For two timess<ts<t, the forward transition can be written as

q​\(𝐳t∣𝐳s\)=Cat⁡\(𝐳t;αt∣s​𝐳s\+\(1−αt∣s\)​𝝅\),αt∣s≔αtαs\.q\(\\mathbf\{z\}\_\{t\}\\mid\\mathbf\{z\}\_\{s\}\)=\\operatorname\{Cat\}\\\!\\left\(\\mathbf\{z\}\_\{t\};\\alpha\_\{t\\mid s\}\\mathbf\{z\}\_\{s\}\+\(1\-\\alpha\_\{t\\mid s\}\)\\bm\{\\pi\}\\right\),\\quad\\alpha\_\{t\\mid s\}\\coloneq\\frac\{\\alpha\_\{t\}\}\{\\alpha\_\{s\}\}\.\(2\)The corresponding reverse posterior has the usual closed form

q​\(𝐳s∣𝐳t,𝐱\)=Cat⁡\(𝐳s;𝐫s∣t​\(𝐱,𝐳t\)\),q\(\\mathbf\{z\}\_\{s\}\\mid\\mathbf\{z\}\_\{t\},\\mathbf\{x\}\)=\\operatorname\{Cat\}\\\!\\left\(\\mathbf\{z\}\_\{s\};\\mathbf\{r\}\_\{s\\mid t\}\(\\mathbf\{x\},\\mathbf\{z\}\_\{t\}\)\\right\),\(3\)where

𝐫s∣t​\(𝐱,𝐳t\)≔\[αt∣s​𝐳t\+\(1−αt∣s\)​⟨𝐳t,𝝅⟩​𝟏\]⊙𝐩s​\(𝐱\)⟨𝐳t,𝐩t​\(𝐱\)⟩\.\\mathbf\{r\}\_\{s\\mid t\}\(\\mathbf\{x\},\\mathbf\{z\}\_\{t\}\)\\coloneq\\frac\{\\left\[\\alpha\_\{t\\mid s\}\\mathbf\{z\}\_\{t\}\+\(1\-\\alpha\_\{t\\mid s\}\)\\left\\langle\\mathbf\{z\}\_\{t\},\\,\\bm\{\\pi\}\\right\\rangle\\mathbf\{1\}\\right\]\\odot\\mathbf\{p\}\_\{s\}\(\\mathbf\{x\}\)\}\{\\left\\langle\\mathbf\{z\}\_\{t\},\\,\\mathbf\{p\}\_\{t\}\(\\mathbf\{x\}\)\\right\\rangle\}\.\(4\)
Standard discrete diffusion training minimizes the categorical reverse KL divergence

ℒzs∣zt\(𝐳t,𝐱^θ,𝐱\)=DKL\[q\(𝐳s∣𝐳t,𝐱\)∥q\(𝐳s∣𝐳t,𝐱^θ\)\]\.\\mathcal\{L\}\_\{z\_\{s\}\\mid z\_\{t\}\}\(\\mathbf\{z\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\}\)=D\_\{\\mathrm\{KL\}\}\\\!\\left\[q\(\\mathbf\{z\}\_\{s\}\\mid\\mathbf\{z\}\_\{t\},\\mathbf\{x\}\)\\,\\middle\\\|\\,q\(\\mathbf\{z\}\_\{s\}\\mid\\mathbf\{z\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\\right\]\.\(5\)where𝐱^θ=fθ​\(𝐳t,t\)∈ΔK−1\\hat\{\\mathbf\{x\}\}\_\{\\theta\}=f\_\{\\theta\}\(\\mathbf\{z\}\_\{t\},t\)\\in\\Delta^\{K\-1\}is the model prediction of the clean\-token distribution\. We use the shorthand𝐩t≔𝐩t​\(𝐱\),𝐩^t≔𝐩t​\(𝐱^θ\)\\mathbf\{p\}\_\{t\}\\coloneq\\mathbf\{p\}\_\{t\}\(\\mathbf\{x\}\),\\ \\hat\{\\mathbf\{p\}\}\_\{t\}\\coloneq\\mathbf\{p\}\_\{t\}\(\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\.

## 3Method

We construct Simplax by augmenting the uniform discrete diffusion process with an auxiliary simplex\-valued variable\. The construction leaves the categorical corruption process unchanged, but introduces an exact Dirichlet–categorical hierarchy around each corrupted state\. We first define this hierarchy and derive its reverse bridge identities\. We then use these identities to obtain a tractable Rao–Blackwellized training objective and a sampler induced by the same bridge structure\.

### 3\.1Simplex relaxation

Building upon \([2](https://arxiv.org/html/2608.10615#S2.E2)\), we consider the following joint factorization for two timess<ts<t:

q​\(𝐱,𝐳s,𝐳t,𝐰s,𝐰t\)=q​\(𝐱\)​q​\(𝐳s∣𝐱\)​q​\(𝐰s∣𝐳s,𝐱\)​q​\(𝐳t∣𝐳s\)​q​\(𝐰t∣𝐳t,𝐱\)\.q\(\\mathbf\{x\},\\mathbf\{z\}\_\{s\},\\mathbf\{z\}\_\{t\},\\mathbf\{w\}\_\{s\},\\mathbf\{w\}\_\{t\}\)=q\(\\mathbf\{x\}\)\\,q\(\\mathbf\{z\}\_\{s\}\\mid\\mathbf\{x\}\)\\,q\(\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{z\}\_\{s\},\\mathbf\{x\}\)\\,q\(\\mathbf\{z\}\_\{t\}\\mid\\mathbf\{z\}\_\{s\}\)\\,q\(\\mathbf\{w\}\_\{t\}\\mid\\mathbf\{z\}\_\{t\},\\mathbf\{x\}\)\.\(6\)
[Figure˜2](https://arxiv.org/html/2608.10615#S3.F2)illustrates the graphical model implied by this factorization\.

![Refer to caption](https://arxiv.org/html/2608.10615v1/x2.png)Figure 2:Graphical model corresponding to the factorization in \([6](https://arxiv.org/html/2608.10615#S3.E6)\)\.Fort∈\(0,1\]t\\in\(0,1\], we introduce a simplex\-valued variable𝐰t∈ΔK−1\\mathbf\{w\}\_\{t\}\\in\\Delta^\{K\-1\}through

q​\(𝐰t∣𝐳t,𝐱\)=Dir⁡\(𝐰t;ηt​𝐩t\+𝐳t\),q\(\\mathbf\{w\}\_\{t\}\\mid\\mathbf\{z\}\_\{t\},\\mathbf\{x\}\)=\\operatorname\{Dir\}\\\!\\left\(\\mathbf\{w\}\_\{t\};\\eta\_\{t\}\\mathbf\{p\}\_\{t\}\+\\mathbf\{z\}\_\{t\}\\right\),\(7\)whereηt\>0\\eta\_\{t\}\>0is a concentration parameter\. The mean of this Dirichlet distribution is centered at the diffusion\-time categorical marginal, while the additive one\-hot count𝐳t\\mathbf\{z\}\_\{t\}anchors the relaxed state to the sampled discrete token\.

This construction yields an exact Dirichlet–categorical hierarchy\.

###### Proposition 1\.

Assumet\>0t\>0\. The Dirichlet–categorical hierarchy satisfies the following properties; statements involving𝐰s\\mathbf\{w\}\_\{s\}additionally requires\>0s\>0\.

1. 1\.The marginal distribution of the relaxed state is q​\(𝐰t∣𝐱\)=Dir⁡\(𝐰t;ηt​𝐩t\)\.q\(\\mathbf\{w\}\_\{t\}\\mid\\mathbf\{x\}\)=\\operatorname\{Dir\}\\\!\\left\(\\mathbf\{w\}\_\{t\};\\eta\_\{t\}\\mathbf\{p\}\_\{t\}\\right\)\.\(8\)
2. 2\.Given𝐰t\\mathbf\{w\}\_\{t\}, the variables𝐱\\mathbf\{x\}and𝐳t\\mathbf\{z\}\_\{t\}are conditionally independent, and the discrete state can be recovered from the relaxed state via q​\(𝐳t∣𝐰t,𝐱\)=q​\(𝐳t∣𝐰t\)=Cat⁡\(𝐳t;𝐰t\)\.q\(\\mathbf\{z\}\_\{t\}\\mid\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)=q\(\\mathbf\{z\}\_\{t\}\\mid\\mathbf\{w\}\_\{t\}\)=\\operatorname\{Cat\}\\\!\\left\(\\mathbf\{z\}\_\{t\};\\mathbf\{w\}\_\{t\}\\right\)\.\(9\)
3. 3\.Fors<ts<t, the reverse conditional posterior of𝐳s\\mathbf\{z\}\_\{s\}given𝐰t\\mathbf\{w\}\_\{t\}and𝐱\\mathbf\{x\}is categorical: q​\(𝐳s∣𝐰t,𝐱\)=Cat⁡\(𝐳s;𝝆s∣t​\(𝐱,𝐰t\)\),q\(\\mathbf\{z\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)=\\operatorname\{Cat\}\\\!\\left\(\\mathbf\{z\}\_\{s\};\\bm\{\\rho\}\_\{s\\mid t\}\(\\mathbf\{x\},\\mathbf\{w\}\_\{t\}\)\\right\),\(10\)where 𝝆s∣t​\(𝐱,𝐰t\)≔𝐩s⊙\[αt∣s​\(𝐰t⊘𝐩t\)\+\(1−αt∣s\)​⟨𝐰t,𝝅⊘𝐩t⟩​1\]\.\\bm\{\\rho\}\_\{s\\mid t\}\(\\mathbf\{x\},\\mathbf\{w\}\_\{t\}\)\\coloneq\\mathbf\{p\}\_\{s\}\\odot\\left\[\\alpha\_\{t\\mid s\}\\bigl\(\\mathbf\{w\}\_\{t\}\\oslash\\mathbf\{p\}\_\{t\}\\bigr\)\+\(1\-\\alpha\_\{t\\mid s\}\)\\,\\left\\langle\\mathbf\{w\}\_\{t\},\\,\\bm\{\\pi\}\\oslash\\mathbf\{p\}\_\{t\}\\right\\rangle\\,\\mathbf\{1\}\\right\]\.\(11\)
4. 4\.For0<s<t0<s<t, the reverse conditional posterior of𝐰s\\mathbf\{w\}\_\{s\}given𝐰t\\mathbf\{w\}\_\{t\}and𝐱\\mathbf\{x\}is a Dirichlet mixture: q​\(𝐰s∣𝐰t,𝐱\)=∑k=1Kρs∣t,k​\(𝐱,𝐰t\)​Dir⁡\(𝐰s;ηs​𝐩s\+𝐞k\)\.q\(\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)=\\sum\_\{k=1\}^\{K\}\\rho\_\{s\\mid t,k\}\(\\mathbf\{x\},\\mathbf\{w\}\_\{t\}\)\\,\\operatorname\{Dir\}\\\!\\left\(\\mathbf\{w\}\_\{s\};\\eta\_\{s\}\\mathbf\{p\}\_\{s\}\+\\mathbf\{e\}\_\{k\}\\right\)\.\(12\)

See[Appendix˜A](https://arxiv.org/html/2608.10615#A1)for the proof\. These identities show that𝐰t\\mathbf\{w\}\_\{t\}is not an ad hoc surrogate\. It is an exact auxiliary variable whose marginal remains native to the simplex and whose decoder back to𝐳t\\mathbf\{z\}\_\{t\}is simply categorical sampling from𝐰t\\mathbf\{w\}\_\{t\}\.

### 3\.2Training objective

A natural starting point is to match the simplex bridge directly:

ℒws∣wt\(𝐰t,𝐱^θ,𝐱;s,t\)=DKL\[q\(𝐰s∣𝐰t,𝐱\)∥q\(𝐰s∣𝐰t,𝐱^θ\)\]\.\\mathcal\{L\}\_\{w\_\{s\}\\mid w\_\{t\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\};s,t\)=D\_\{\\mathrm\{KL\}\}\\\!\\left\[q\(\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)\\,\\middle\\\|\\,q\(\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\\right\]\.\(13\)This is the most direct objective associated with the relaxed bridge, but it is generally intractable becauseq​\(𝐰s∣𝐰t,𝐱\)q\(\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)is a Dirichlet mixture, as shown in \([12](https://arxiv.org/html/2608.10615#S3.E12)\)\.

At training time, we sample the augmented noisy state as

𝐰t∼q​\(𝐰t∣𝐱\),𝐳t∼q​\(𝐳t∣𝐰t\)=Cat⁡\(𝐳t;𝐰t\),\\mathbf\{w\}\_\{t\}\\sim q\(\\mathbf\{w\}\_\{t\}\\mid\\mathbf\{x\}\),\\qquad\\mathbf\{z\}\_\{t\}\\sim q\(\\mathbf\{z\}\_\{t\}\\mid\\mathbf\{w\}\_\{t\}\)=\\operatorname\{Cat\}\(\\mathbf\{z\}\_\{t\};\\mathbf\{w\}\_\{t\}\),and predict the clean\-token distribution from the categorical state:

𝐱^θ=fθ​\(𝐳t,t\)\.\\widehat\{\\mathbf\{x\}\}\_\{\\theta\}=f\_\{\\theta\}\(\\mathbf\{z\}\_\{t\},t\)\.
To define the relaxed discrete bridge objective, let𝐳~t\\widetilde\{\\mathbf\{z\}\}\_\{t\}denote a second categorical variable satisfying

𝐳~t∼q​\(𝐳~t∣𝐰t\)=Cat⁡\(𝐳~t;𝐰t\),\\widetilde\{\\mathbf\{z\}\}\_\{t\}\\sim q\(\\widetilde\{\\mathbf\{z\}\}\_\{t\}\\mid\\mathbf\{w\}\_\{t\}\)=\\operatorname\{Cat\}\(\\widetilde\{\\mathbf\{z\}\}\_\{t\};\\mathbf\{w\}\_\{t\}\),conditionally independently of the network input𝐳t\\mathbf\{z\}\_\{t\}given𝐰t\\mathbf\{w\}\_\{t\}\. We optimize

ℒ¯zs∣zt,wt​\(𝐰t,𝐳t,𝐱;s,t\)≔𝔼q​\(𝐳~t∣𝐰t\)​\[KL⁡\(q​\(𝐳s∣𝐳~t,𝐱\)∥q​\(𝐳s∣𝐳~t,𝐱^θ\)\)\],\\displaystyle\\bar\{\\mathcal\{L\}\}\_\{z\_\{s\}\\mid z\_\{t\},w\_\{t\}\}\(\\mathbf\{w\}\_\{t\},\\mathbf\{z\}\_\{t\},\\mathbf\{x\};s,t\)\\coloneq\\mathbb\{E\}\_\{q\(\\widetilde\{\\mathbf\{z\}\}\_\{t\}\\mid\\mathbf\{w\}\_\{t\}\)\}\\Big\[\\operatorname\{KL\}\\big\(q\(\\mathbf\{z\}\_\{s\}\\mid\\widetilde\{\\mathbf\{z\}\}\_\{t\},\\mathbf\{x\}\)\\mathrel\{\\\|\}q\(\\mathbf\{z\}\_\{s\}\\mid\\widetilde\{\\mathbf\{z\}\}\_\{t\},\\widehat\{\\mathbf\{x\}\}\_\{\\theta\}\)\\big\)\\Big\],\(14\)where𝐱^θ=fθ​\(𝐳t,t\)\\widehat\{\\mathbf\{x\}\}\_\{\\theta\}=f\_\{\\theta\}\(\\mathbf\{z\}\_\{t\},t\)\. This objective averages the standard discrete reverse KL divergence over an auxiliary decoder sample from𝐰t\\mathbf\{w\}\_\{t\}, while retaining𝐳t\\mathbf\{z\}\_\{t\}as the denoiser input \([Figure˜1](https://arxiv.org/html/2608.10615#S1.F1)\)\.

Rao–Blackwellized form\.The expectation over the auxiliary decoder sample𝐳~t\\widetilde\{\\mathbf\{z\}\}\_\{t\}in \([14](https://arxiv.org/html/2608.10615#S3.E14)\) can be marginalized exactly\. Equivalently, the resulting expression is the Rao–Blackwellized form of the Monte Carlo estimator obtained by first sampling𝐳~t∼q​\(𝐳~t∣𝐰t\)\\widetilde\{\\mathbf\{z\}\}\_\{t\}\\sim q\(\\widetilde\{\\mathbf\{z\}\}\_\{t\}\\mid\\mathbf\{w\}\_\{t\}\)and then evaluating the discrete reverse KL\. The independently sampled𝐳t\\mathbf\{z\}\_\{t\}remains the denoiser input and is not marginalized by this step\.

###### Proposition 2\.

The relaxed discrete bridge objective in \([14](https://arxiv.org/html/2608.10615#S3.E14)\) admits the exact closed form

ℒ¯zs∣zt,wt​\(𝐰t,𝐱^θ,𝐱;s,t\)=⟨𝐰t,log⁡𝐩^t−log⁡𝐩t⟩\+⟨𝝆s∣t​\(𝐱,𝐰t\),log⁡𝐩s−log⁡𝐩^s⟩\.\\bar\{\\mathcal\{L\}\}\_\{z\_\{s\}\\mid z\_\{t\},w\_\{t\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\};s,t\)=\\left\\langle\\mathbf\{w\}\_\{t\},\\,\\log\\hat\{\\mathbf\{p\}\}\_\{t\}\-\\log\\mathbf\{p\}\_\{t\}\\right\\rangle\+\\left\\langle\\bm\{\\rho\}\_\{s\\mid t\}\(\\mathbf\{x\},\\mathbf\{w\}\_\{t\}\),\\,\\log\\mathbf\{p\}\_\{s\}\-\\log\\hat\{\\mathbf\{p\}\}\_\{s\}\\right\\rangle\.\(15\)

The proof is given in[Appendix˜B](https://arxiv.org/html/2608.10615#A2)\. Equation \([15](https://arxiv.org/html/2608.10615#S3.E15)\) is fully tractable and eliminates sampling noise associated with the auxiliary decoder sample𝐳~t\\widetilde\{\\mathbf\{z\}\}\_\{t\}\. Operationally, it shows that the reverse\-bridge loss depends on𝐰t\\mathbf\{w\}\_\{t\}only through two quantities: the decoded current\-time marginal term⟨𝐰t,log⁡𝐩^t⟩\\left\\langle\\mathbf\{w\}\_\{t\},\\,\\log\\hat\{\\mathbf\{p\}\}\_\{t\}\\right\\rangleand the induced reverse posterior𝝆s∣t​\(𝐱,𝐰t\)\\bm\{\\rho\}\_\{s\\mid t\}\(\\mathbf\{x\},\\mathbf\{w\}\_\{t\}\)\. The network prediction itself remains conditioned on the separately sampled categorical input𝐳t\\mathbf\{z\}\_\{t\}\.

Continuous\-time limit\.The closed form in \([15](https://arxiv.org/html/2608.10615#S3.E15)\) yields a non\-degenerate infinitesimal limit\. Lets=t−Δs=t\-\\DeltawithΔ↓0\\Delta\\downarrow 0\.

###### Proposition 3\.

Up toθ\\theta\-independent additive terms, the relaxed discrete bridge objective satisfies

ℒ¯zs∣zt,wt​\(𝐰t,𝐱^θ,𝐱;t−Δ,t\)=Δ​ℓct​\(𝐰t,𝐱^θ,𝐱,t\)\+o​\(Δ\),\\bar\{\\mathcal\{L\}\}\_\{z\_\{s\}\\mid z\_\{t\},w\_\{t\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\};t\-\\Delta,t\)=\\Delta\\,\\ell\_\{\\mathrm\{ct\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\},t\)\+o\(\\Delta\),\(16\)where

ℓct\(𝐰t,𝐱^θ,𝐱,t\)=λ\(t\)\[\\displaystyle\\ell\_\{\\mathrm\{ct\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\},t\)=\\lambda\(t\)\\Bigg\[⟨𝐰t,𝝅⊘𝐩^t⟩−⟨𝐰t,𝝅⊘𝐩t⟩⟨𝐩t,log𝐩^t⟩\+⟨𝝅⊙\(𝐰t⊘𝐩t\),log𝐩^t⟩\]\.\\displaystyle\\left\\langle\\mathbf\{w\}\_\{t\},\\,\\bm\{\\pi\}\\oslash\\hat\{\\mathbf\{p\}\}\_\{t\}\\right\\rangle\-\\left\\langle\\mathbf\{w\}\_\{t\},\\,\\bm\{\\pi\}\\oslash\\mathbf\{p\}\_\{t\}\\right\\rangle\\,\\left\\langle\\mathbf\{p\}\_\{t\},\\,\\log\\hat\{\\mathbf\{p\}\}\_\{t\}\\right\\rangle\+\\left\\langle\\bm\{\\pi\}\\odot\(\\mathbf\{w\}\_\{t\}\\oslash\\mathbf\{p\}\_\{t\}\),\\,\\log\\hat\{\\mathbf\{p\}\}\_\{t\}\\right\\rangle\\Bigg\]\.\(17\)andλ​\(t\)≔−dd​t​log⁡α​\(t\)\\lambda\(t\)\\coloneq\-\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\\log\\alpha\(t\)\. Consequently, the corresponding continuous\-time objective is

ℒct=∫01𝔼q​\(𝐱\)​\[𝔼q​\(𝐰t∣𝐱\)​\[ℓct​\(𝐰t,𝐱^θ,𝐱,t\)\]\]​dt\.\\mathcal\{L\}\_\{\\mathrm\{ct\}\}=\\int\_\{0\}^\{1\}\\mathbb\{E\}\_\{q\(\\mathbf\{x\}\)\}\\\!\\left\[\\mathbb\{E\}\_\{q\(\\mathbf\{w\}\_\{t\}\\mid\\mathbf\{x\}\)\}\\\!\\left\[\\ell\_\{\\mathrm\{ct\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\},t\)\\right\]\\right\]\\,\\mathrm\{d\}t\.\(18\)

The proof is deferred to[Section˜C\.1](https://arxiv.org/html/2608.10615#A3.SS1)\. This proposition identifies \([18](https://arxiv.org/html/2608.10615#S3.E18)\) as the continuous\-time counterpart of \([14](https://arxiv.org/html/2608.10615#S3.E14)\)\.

The structure of \([14](https://arxiv.org/html/2608.10615#S3.E14)\) and \([18](https://arxiv.org/html/2608.10615#S3.E18)\) is closely related to UDLM\(Schiffet al\.,[2025](https://arxiv.org/html/2608.10615#bib.bib7)\)\. UDLM derives a continuous\-time reverse\-KL objective directly from the discrete corrupted state \([5](https://arxiv.org/html/2608.10615#S2.E5)\)\. Our construction introduces the exact auxiliary variable𝐰t\\mathbf\{w\}\_\{t\}, averages the same reverse\-KL bridge over an auxiliary draw𝐳~t∼q​\(𝐳~t∣𝐰t\)=Cat⁡\(𝐰t\)\\widetilde\{\\mathbf\{z\}\}\_\{t\}\\sim q\(\\widetilde\{\\mathbf\{z\}\}\_\{t\}\\mid\\mathbf\{w\}\_\{t\}\)=\\operatorname\{Cat\}\(\\mathbf\{w\}\_\{t\}\), and then takes the infinitesimal limit\. In this sense, \([18](https://arxiv.org/html/2608.10615#S3.E18)\) can be understood as a simplex\-relaxed continuous\-time analogue of the UDLM objective\.

### 3\.3Sampling

At inference time, we run the reverse process on a grid1=tN\>tN−1\>⋯\>t0=01=t\_\{N\}\>t\_\{N\-1\}\>\\cdots\>t\_\{0\}=0\. For a denoiser conditioned on𝐳t\\mathbf\{z\}\_\{t\}, the Dirichlet–categorical hierarchy yields a stochastic sampler that maintains the augmented state\(𝐳t,𝐰t\)\(\\mathbf\{z\}\_\{t\},\\mathbf\{w\}\_\{t\}\)at positive times\. Since the endpoint marginal isq​\(𝐰1\)=Dir⁡\(𝐰1;η1​𝝅\)q\(\\mathbf\{w\}\_\{1\}\)=\\operatorname\{Dir\}\(\\mathbf\{w\}\_\{1\};\\eta\_\{1\}\\bm\{\\pi\}\)and the exact decoder isq​\(𝐳1∣𝐰1\)=Cat⁡\(𝐳1;𝐰1\)q\(\\mathbf\{z\}\_\{1\}\\mid\\mathbf\{w\}\_\{1\}\)=\\operatorname\{Cat\}\(\\mathbf\{z\}\_\{1\};\\mathbf\{w\}\_\{1\}\), generation starts from

𝐰tN∼Dir⁡\(ηtN​𝝅\),𝐳tN∼Cat⁡\(𝐰tN\)\.\\mathbf\{w\}\_\{t\_\{N\}\}\\sim\\operatorname\{Dir\}\(\\eta\_\{t\_\{N\}\}\\bm\{\\pi\}\),\\qquad\\mathbf\{z\}\_\{t\_\{N\}\}\\sim\\operatorname\{Cat\}\(\\mathbf\{w\}\_\{t\_\{N\}\}\)\.\(19\)Given adjacent timess=tn−1<t=tns=t\_\{n\-1\}<t=t\_\{n\}and the current pair\(𝐳t,𝐰t\)\(\\mathbf\{z\}\_\{t\},\\mathbf\{w\}\_\{t\}\), the denoiser predicts

𝐱^θ=fθ​\(𝐳t,t\),\\hat\{\\mathbf\{x\}\}\_\{\\theta\}=f\_\{\\theta\}\(\\mathbf\{z\}\_\{t\},t\),\(20\)which induces the bridge marginals𝐩^s\\hat\{\\mathbf\{p\}\}\_\{s\}and𝐩^t\\hat\{\\mathbf\{p\}\}\_\{t\}and the reverse categorical posterior𝝆s∣t​\(𝐱^θ,𝐰t\)\\bm\{\\rho\}\_\{s\\mid t\}\(\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{w\}\_\{t\}\)\. Our default sampler is the stochastic ancestral sampler implied by the Dirichlet–categorical hierarchy\. Each reverse step draws

𝐳s∼Cat⁡\(𝝆s∣t​\(𝐱^θ,𝐰t\)\),𝐰s∼Dir⁡\(ηs​𝐩^s\+𝐳s\)\.\\mathbf\{z\}\_\{s\}\\sim\\operatorname\{Cat\}\\\!\\left\(\\bm\{\\rho\}\_\{s\\mid t\}\(\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{w\}\_\{t\}\)\\right\),\\qquad\\mathbf\{w\}\_\{s\}\\sim\\operatorname\{Dir\}\\\!\\left\(\\eta\_\{s\}\\hat\{\\mathbf\{p\}\}\_\{s\}\+\\mathbf\{z\}\_\{s\}\\right\)\.\(21\)The sampled𝐳s\\mathbf\{z\}\_\{s\}is used as the network input at the next reverse step, while𝐰s\\mathbf\{w\}\_\{s\}carries the auxiliary bridge information required by the next reverse posterior\. Repeating \([21](https://arxiv.org/html/2608.10615#S3.E21)\) fromtN=1t\_\{N\}=1tot0=0t\_\{0\}=0yields the final categorical sample𝐳0\\mathbf\{z\}\_\{0\}\. The categorical input also has a computational advantage\. In a standard token model,𝐳t\\mathbf\{z\}\_\{t\}is stored as an integer token index and its embedding is obtained by lookup\. Feeding the dense simplex vector𝐰t\\mathbf\{w\}\_\{t\}instead requires computing𝐰t𝖳​E\\mathbf\{w\}\_\{t\}^\{\\mathsf\{T\}\}Efor the vocabulary embedding matrixEEat every sequence position, adding a vocabulary\-sized dense matrix multiplication and the associated memory traffic\. The main experiments therefore use the𝐳t\\mathbf\{z\}\_\{t\}\-input formulation above\.

## 4Experiments

We evaluate Simplax on unconditional text generation with OpenWebText\(Gokaslan and Cohen,[2019](https://arxiv.org/html/2608.10615#bib.bib28)\)and constrained categorical generation with Sudoku\(Leeet al\.,[2026](https://arxiv.org/html/2608.10615#bib.bib22); Deschenaux and Gulcehre,[2026](https://arxiv.org/html/2608.10615#bib.bib34)\)\. We first report compact design diagnostics on OpenWebText and then present the main comparisons\.

### 4\.1Experimental setup

#### OpenWebText\.

We tokenize OpenWebText with the GPT\-2 BPE tokenizer\(Radfordet al\.,[2019](https://arxiv.org/html/2608.10615#bib.bib25)\), giving\|𝒱\|=50,257\|\\mathcal\{V\}\|=50\{,\}257, and use sequence lengthL=1,024L=1,024\. All methods use the179​M179\\mathrm\{M\}\-parameter diffusion transformer ofSahooet al\.\([2024](https://arxiv.org/html/2608.10615#bib.bib5)\):1212transformer blocks, rotary position embeddings\(Suet al\.,[2024](https://arxiv.org/html/2608.10615#bib.bib37)\), AdaLN time conditioning\(Peebles and Xie,[2023](https://arxiv.org/html/2608.10615#bib.bib36)\), and a softmax output head\. Models are trained with Adam\(Kingma and Ba,[2014](https://arxiv.org/html/2608.10615#bib.bib38)\), learning rate3×10−43\\times 10^\{\-4\}, batch size512512, and a total budget of1​M1\\mathrm\{M\}iterations\. Unless stated otherwise, Simplax uses𝐳t\\mathbf\{z\}\_\{t\}as the denoiser input and a constant concentrationηt≡0\.01\\eta\_\{t\}\\equiv 0\.01\.

#### Sudoku\.

We build on the Sudoku benchmark ofDeschenaux and Gulcehre \([2026](https://arxiv.org/html/2608.10615#bib.bib34)\), while using a cross\-clue generalization protocol in which all models are trained only on puzzles with3030clues\. The dataset contains48,00048\{,\}000training and2,0002\{,\}000validation puzzles, each constructed to have a unique solution\. A Sudoku instance is represented as a180180\-token sequence consisting of a9191\-token puzzle prefix and an8989\-token solution\. The puzzle prefix contains aBOStoken, all8181cells with unobserved cells represented by a blank token, eight row separators, and a secondBOStoken\. The solution contains the8181completed cells and eight row separators\. The training loss is applied only to the solution tokens\.

All methods use Transformer backbones with eight blocks, hidden dimension512512, eight attention heads, and dropout0\.10\.1\. Their parameter counts range from25\.2125\.21M to28\.5928\.59M; the principal differences are the training objective, time conditioning, and inference procedure\. Models are trained for20,00020\{,\}000steps using Adam with learning rate3×10−43\\times 10^\{\-4\}and global batch size256256\. Further architectural and optimization details are provided in[Section˜E\.1](https://arxiv.org/html/2608.10615#A5.SS1)\.

At inference time, the same3030\-clue\-trained checkpoint is evaluated with4040,3535,3030,2525,2020, and1717clues\. The3030\-clue setting matches the training distribution, while the remaining settings evaluate transfer across clue densities\. The4040\- and3535\-clue settings provide more conditioning information than observed during training, whereas the2525\-,2020\-, and1717\-clue settings provide progressively less conditioning information\. In particular,1717is the minimum number of clues for which a standard9×99\\times 9Sudoku puzzle can admit a unique solution\(McGuireet al\.,[2014](https://arxiv.org/html/2608.10615#bib.bib47); Linet al\.,[2013](https://arxiv.org/html/2608.10615#bib.bib46)\), making the1717\-clue setting the most sparsely conditioned regime in our evaluation\. We additionally evaluate generation from an all\-blank puzzle prefix, which contains no clue information and is treated as unconditional Sudoku generation\.

#### Metrics\.

For OpenWebText, we draw1,0241,024sequences and report generative unigram entropy and generative perplexity under GPT\-2 Large, GPT\-2 XL\(Radfordet al\.,[2019](https://arxiv.org/html/2608.10615#bib.bib25)\), and Llama\-2 7B\(Touvronet al\.,[2023](https://arxiv.org/html/2608.10615#bib.bib26)\)\. The OpenWebText unigram entropy is5\.445\.44nats\. For conditional Sudoku, we report solving accuracy\. For unconditional Sudoku, we report validity, the fraction of generated boards satisfying all Sudoku constraints\.

### 4\.2Design diagnostics on OpenWebText

The input and self\-conditioning diagnostics use50​k50\\mathrm\{k\}\-step runs with the same tokenizer, sequence length, and backbone as the main experiment\. The initialization comparison matches the total training budget at1​M1\\mathrm\{M\}iterations\. These experiments characterize individual design choices rather than provide the main method comparison\.

![Refer to caption](https://arxiv.org/html/2608.10615v1/x3.png)Figure 3:Design diagnostics on OpenWebText\. The rows show temperature\-swept generation frontiers at NFE=16=16and128128\. The columns compare self\-conditioning \(a,d\), denoiser input𝐰t\\mathbf\{w\}\_\{t\}versus𝐳t\\mathbf\{z\}\_\{t\}\(b,e\), and initialization from a pretrained UDLM checkpoint under a matched1​M1\\mathrm\{M\}\-iteration budget \(c,f\)\. The dotted line marks the OpenWebText entropy,5\.445\.44\. Gen\. PPL is evaluated with GPT\-2 Large; lower is better at comparable Gen\. ENT\.#### Self\-conditioning\.

For the𝐰t\\mathbf\{w\}\_\{t\}\-input diagnostic, the preferred setting depends on the inference budget: omitting self\-conditioning is better near the data\-entropy operating point at NFE=16=16, whereas using it is better at NFE=128=128[Figure˜3](https://arxiv.org/html/2608.10615#S4.F3)\(a,d\)\. The experiment therefore does not support a budget\-independent conclusion\.

#### Denoiser input\.

The bridge and objective do not require the auxiliary simplex state itself to be the denoiser input\. With the same𝐰t\\mathbf\{w\}\_\{t\}\-based objective, the𝐳t\\mathbf\{z\}\_\{t\}\-input model attains lower Gen\. PPL at comparable Gen\. ENT at both NFE values[Figure˜3](https://arxiv.org/html/2608.10615#S4.F3)\(b,e\)\. It also avoids an additional dense projection at the input layer: token indices use embedding lookup, whereas a simplex input requires𝐰t𝖳​E\\mathbf\{w\}\_\{t\}^\{\\mathsf\{T\}\}Ewith the vocabulary embedding matrixEEat every sequence position\. We therefore use𝐳t\\mathbf\{z\}\_\{t\}as the denoiser input in the main experiments while retaining𝐰t\\mathbf\{w\}\_\{t\}in the objective and reverse update\.

#### UDLM initialization\.

We compare Simplax trained from scratch for1​M1\\mathrm\{M\}iterations with a model trained as UDLM for800​k800\\mathrm\{k\}iterations and then with the Simplax objective for200​k200\\mathrm\{k\}iterations\. UDLM initialization improves the Gen\. PPL–Gen\. ENT frontier at both NFE values[Figure˜3](https://arxiv.org/html/2608.10615#S4.F3)\(c,f\)\.

### 4\.3Unconditional generation on OpenWebText

We compare Simplax with MDLM\(Sahooet al\.,[2024](https://arxiv.org/html/2608.10615#bib.bib5)\), UDLM\(Schiffet al\.,[2025](https://arxiv.org/html/2608.10615#bib.bib7)\), Duo\(Sahooet al\.,[2025](https://arxiv.org/html/2608.10615#bib.bib10)\), CANDI\(Pynadathet al\.,[2025](https://arxiv.org/html/2608.10615#bib.bib21)\), FLM\(Leeet al\.,[2026](https://arxiv.org/html/2608.10615#bib.bib22)\), LangFlow\(Chenet al\.,[2026](https://arxiv.org/html/2608.10615#bib.bib40)\), and S\-FLM\(Deschenaux and Gulcehre,[2026](https://arxiv.org/html/2608.10615#bib.bib34)\)\. For each method and NFE budget, we sweep1515temperatures from0\.840\.84to1\.121\.12in increments of0\.020\.02and select the operating point whose generated entropy is closest to5\.445\.44nats\.

Table 1:OpenWebText unconditional generation at selected NFE values\. Gen\. ENT is generative unigram entropy and should be compared with the data entropy5\.445\.44\. Gen\. PPL is evaluated by the indicated external language model\. The best and second\-best values in each column are shown in bold and underlined, respectively\.Simplax has the lowest Gen\. PPL under all three evaluators at NFE=16=16and1,0241,024\. At NFE=128=128, it is best under GPT\-2 Large and GPT\-2 XL, while LangFlow is best under Llama\-2 7B\. The temperature\-swept Llama\-2 7B frontiers across all evaluated NFE budgets are shown in[Figure˜4](https://arxiv.org/html/2608.10615#S4.F4)\.

![Refer to caption](https://arxiv.org/html/2608.10615v1/x4.png)Figure 4:Llama\-2 7B generative frontiers on OpenWebText forNFE∈\{8,16,32,64,128,256,512,1,024\}\\mathrm\{NFE\}\\in\\\{8,16,32,64,128,256,512,1,024\\\}\. Each panel shows the temperature\-swept Gen\. PPL–Gen\. ENT tradeoff\. Lower Gen\. PPL is better, and the reference data entropy is5\.445\.44\.
### 4\.4Constrained categorical generation on Sudoku

Table 2:Conditional Sudoku solving accuracy and unconditional Sudoku validity in percent\. All models are trained with3030clues\. The4040\- and3535\-clue settings evaluate transfer to more heavily conditioned inputs, the3030\-clue setting matches the training clue density, and the2525\-,2020\-, and1717\-clue settings evaluate transfer to progressively less heavily conditioned inputs\. The best and second\-best results in each column are shown in bold and underlined, respectively\.Simplax achieves the highest performance across all conditional and unconditional settings in[Table˜2](https://arxiv.org/html/2608.10615#S4.T2)\. Its advantage extends beyond the3030\-clue training distribution to both more and less conditioned inputs, including the challenging low\-clue regimes\. For unconditional generation, Simplax achieves95\.85%95\.85\\%validity, compared with80\.95%80\.95\\%for the strongest baseline\.

## 5Related Work

#### Discrete diffusion for categorical data\.

Discrete diffusion for categorical data was developed through early multinomial formulations and later unified and substantially generalized by D3PM, which introduced structured transition kernels such as uniform and absorbing corruptions and established the standard variational training recipe\(Hoogeboomet al\.,[2021](https://arxiv.org/html/2608.10615#bib.bib1); Austinet al\.,[2021](https://arxiv.org/html/2608.10615#bib.bib2)\)\. This framework was subsequently extended to continuous\-time formulations and alternative reverse objectives\(Campbellet al\.,[2022](https://arxiv.org/html/2608.10615#bib.bib3); Louet al\.,[2024](https://arxiv.org/html/2608.10615#bib.bib4); Zhanget al\.,[2025](https://arxiv.org/html/2608.10615#bib.bib11)\), and has since supported a broad line of language\-modeling work covering masked, absorbing, and uniform\-state diffusion\(Sahooet al\.,[2024](https://arxiv.org/html/2608.10615#bib.bib5); Shiet al\.,[2024](https://arxiv.org/html/2608.10615#bib.bib6); Ouet al\.,[2025](https://arxiv.org/html/2608.10615#bib.bib8); Zhenget al\.,[2025](https://arxiv.org/html/2608.10615#bib.bib9); Sahooet al\.,[2025](https://arxiv.org/html/2608.10615#bib.bib10); Deschenauxet al\.,[2026](https://arxiv.org/html/2608.10615#bib.bib17)\)\. Our method stays within this discrete\-diffusion lineage: we keep the original categorical forward process and reverse posterior, and do not replace the primary generative state\.

#### Auxiliary\-variable and hybrid formulations\.

A parallel line of work enriches discrete diffusion through auxiliary variables or structured reverse distributions\. Di4C\(Hayakawaet al\.,[2025](https://arxiv.org/html/2608.10615#bib.bib48)\)uses mixtures of product models to capture dimensional correlations, VADD\(Xieet al\.,[2026](https://arxiv.org/html/2608.10615#bib.bib49)\)introduces a Gaussian latent into masked denoising, and CoDD couples factorized outputs with probabilistic circuits\(Liet al\.,[2026](https://arxiv.org/html/2608.10615#bib.bib50)\)\. Continuous and hybrid constructions include Gaussian\-relaxed views in Duo and Duo\+\+\(Sahooet al\.,[2025](https://arxiv.org/html/2608.10615#bib.bib10); Deschenauxet al\.,[2026](https://arxiv.org/html/2608.10615#bib.bib17)\), Euclidean denoising over one\-hot states in FLM\(Leeet al\.,[2026](https://arxiv.org/html/2608.10615#bib.bib22)\), and discrete–continuous diffusion in CADD and CANDI\(Zhenget al\.,[2026](https://arxiv.org/html/2608.10615#bib.bib20); Pynadathet al\.,[2025](https://arxiv.org/html/2608.10615#bib.bib21)\)\. Simplax instead introduces a simplex\-valued auxiliary variable while preserving categorical diffusion\. Unlike methods that use the auxiliary variable as the denoiser input or primary generative state, the𝐳t\\mathbf\{z\}\_\{t\}\-input Simplax formulation retains the categorical state as the network input and uses the simplex variable to define the reverse\-bridge objective and sampler\.

#### Diffusion and flow on the simplex\.

Our method is also related to work that defines the generative process itself on the simplex\. This includes simplex diffusion based on softmax\-transformed continuous processes\(Flotoet al\.,[2023](https://arxiv.org/html/2608.10615#bib.bib24)\), simplex diffusion via categorical SDEs and Cox–Ingersoll–Ross dynamics\(Richemondet al\.,[2023](https://arxiv.org/html/2608.10615#bib.bib23)\), Dirichlet\-based score models such as DDSM\(Avdeyevet al\.,[2023](https://arxiv.org/html/2608.10615#bib.bib14)\), Dirichlet Flow Matching\(Starket al\.,[2024](https://arxiv.org/html/2608.10615#bib.bib15)\), and recent unifying views of discrete, Gaussian, and simplicial diffusion\(Chandraet al\.,[2026](https://arxiv.org/html/2608.10615#bib.bib16)\)\. These methods are close to ours in geometry, since they treat simplex\-valued states as first\-class objects, but differ in role: in our formulation, the simplex variable is not the primary generative state, but an exact auxiliary bridge attached to a standard discrete diffusion process\.

## 6Conclusion and Limitations

We introduced Simplax, an exact Dirichlet–categorical augmentation of uniform discrete diffusion\. Simplax preserves the categorical forward process while introducing an auxiliary simplex state to derive a Rao–Blackwellized reverse\-bridge objective and stochastic ancestral sampler\. It improves the Gen\. PPL–Gen\. ENT tradeoff on OpenWebText and achieves the highest Sudoku performance among the compared methods across all evaluated clue densities, from4040to1717clues, as well as in unconditional generation\.

#### Limitation\.

The present formulation is specialized to uniform categorical corruption and introduces an auxiliary simplex\-valued state whose computational overhead relative to standard discrete diffusion has not been fully characterized\. Moreover, the concentration schedule remains an additional design choice rather than being determined by the theory\. Extending the construction to broader categorical corruption kernels and developing more efficient reverse solvers are important directions for future work\.

## Acknowledgments

We thank Chanhyuk Lee, Jaehoon Yoo, and Jinwoo Kim for insightful discussions\.

## References

- A\. Alp \(2024\)Sudoku puzzle generator\.Note:[https://github\.com/alicommit\-malp/sudoku](https://github.com/alicommit-malp/sudoku)Cited by:[§E\.1](https://arxiv.org/html/2608.10615#A5.SS1.SSS0.Px1.p1.10)\.
- J\. Austin, D\. D\. Johnson, J\. Ho, D\. Tarlow, and R\. van den Berg \(2021\)Structured denoising diffusion models in discrete state\-spaces\.InAdvances in Neural Information Processing Systems,Cited by:[§1](https://arxiv.org/html/2608.10615#S1.p1.1),[§1](https://arxiv.org/html/2608.10615#S1.p2.1),[§5](https://arxiv.org/html/2608.10615#S5.SS0.SSS0.Px1.p1.1)\.
- P\. Avdeyev, C\. Shi, Y\. Tan, K\. Dudnyk, and J\. Zhou \(2023\)Dirichlet diffusion score model for biological sequence generation\.InInternational Conference on Machine Learning,Cited by:[§5](https://arxiv.org/html/2608.10615#S5.SS0.SSS0.Px3.p1.1)\.
- A\. Campbell, J\. Benton, V\. D\. Bortoli, T\. Rainforth, G\. Deligiannidis, and A\. Doucet \(2022\)A continuous time framework for discrete denoising models\.InAdvances in Neural Information Processing Systems,Cited by:[§1](https://arxiv.org/html/2608.10615#S1.p1.1),[§5](https://arxiv.org/html/2608.10615#S5.SS0.SSS0.Px1.p1.1)\.
- N\. A\. Chandra, Y\. L\. Li, A\. N\. Amin, A\. Ali, J\. Rollins, S\. W\. Ober, A\. Raghu, and A\. G\. Wilson \(2026\)A unification of discrete, gaussian, and simplicial diffusion\.InThe Fourteenth International Conference on Learning Representations,Cited by:[§5](https://arxiv.org/html/2608.10615#S5.SS0.SSS0.Px3.p1.1)\.
- Y\. Chen, C\. Liang, H\. Sui, R\. Guo, C\. Cheng, J\. You, and G\. Liu \(2026\)LangFlow: continuous diffusion rivals discrete in language modeling\.Cited by:[§4\.3](https://arxiv.org/html/2608.10615#S4.SS3.p1.5),[Table 1](https://arxiv.org/html/2608.10615#S4.T1.3.1.8.6.1)\.
- T\. M\. Cover and J\. A\. Thomas \(2006\)Elements of information theory\.Cited by:[§D\.2](https://arxiv.org/html/2608.10615#A4.SS2.3.p1.1)\.
- J\. Deschenaux, C\. Gulcehre, and S\. S\. Sahoo \(2026\)The diffusion duality, chapter II: $\\psi$\-samplers and efficient curriculum\.InThe Fourteenth International Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2608.10615#S1.p2.1),[§5](https://arxiv.org/html/2608.10615#S5.SS0.SSS0.Px1.p1.1),[§5](https://arxiv.org/html/2608.10615#S5.SS0.SSS0.Px2.p1.1)\.
- J\. Deschenaux and C\. Gulcehre \(2026\)Language modeling with hyperspherical flows\.Cited by:[§E\.1](https://arxiv.org/html/2608.10615#A5.SS1.SSS0.Px1.p1.10),[§4\.1](https://arxiv.org/html/2608.10615#S4.SS1.SSS0.Px2.p1.8),[§4\.3](https://arxiv.org/html/2608.10615#S4.SS3.p1.5),[Table 1](https://arxiv.org/html/2608.10615#S4.T1.3.1.9.7.1),[§4](https://arxiv.org/html/2608.10615#S4.p1.1)\.
- G\. Floto, T\. Jonsson, M\. Nica, S\. Sanner, and E\. Z\. Zhu \(2023\)Diffusion on the probability simplex\.InICML 2023 Workshop: Sampling and Optimization in Discrete Space,Cited by:[§5](https://arxiv.org/html/2608.10615#S5.SS0.SSS0.Px3.p1.1)\.
- A\. Gokaslan and V\. Cohen \(2019\)OpenWebText corpus\.Cited by:[§4](https://arxiv.org/html/2608.10615#S4.p1.1)\.
- S\. Hayakawa, Y\. Takida, M\. Imaizumi, H\. Wakaki, and Y\. Mitsufuji \(2025\)Distillation of discrete diffusion through dimensional correlations\.InProceedings of the 42nd International Conference on Machine Learning,Cited by:[§5](https://arxiv.org/html/2608.10615#S5.SS0.SSS0.Px2.p1.1)\.
- E\. Hoogeboom, D\. Nielsen, P\. Jaini, P\. Forré, and M\. Welling \(2021\)Argmax flows and multinomial diffusion: learning categorical distributions\.InAdvances in Neural Information Processing Systems,Cited by:[§1](https://arxiv.org/html/2608.10615#S1.p1.1),[§5](https://arxiv.org/html/2608.10615#S5.SS0.SSS0.Px1.p1.1)\.
- D\. P\. Kingma and J\. Ba \(2014\)Adam: a method for stochastic optimization\.Cited by:[§4\.1](https://arxiv.org/html/2608.10615#S4.SS1.SSS0.Px1.p1.9)\.
- S\. Kullback and R\. A\. Leibler \(1951\)On information and sufficiency\.The Annals of Mathematical Statistics\.Cited by:[§D\.2](https://arxiv.org/html/2608.10615#A4.SS2.1.p1.2)\.
- C\. Lee, J\. Yoo, M\. Agarwal, S\. Shah, J\. Huang, A\. Raghunathan, S\. Hong, N\. M\. Boffi, and J\. Kim \(2026\)One\-step language modeling via continuous denoising\.Cited by:[§4\.3](https://arxiv.org/html/2608.10615#S4.SS3.p1.5),[Table 1](https://arxiv.org/html/2608.10615#S4.T1.3.1.7.5.1),[Table 2](https://arxiv.org/html/2608.10615#S4.T2.15.5.3.1),[§4](https://arxiv.org/html/2608.10615#S4.p1.1),[§5](https://arxiv.org/html/2608.10615#S5.SS0.SSS0.Px2.p1.1)\.
- I\. Li, Z\. Shao, B\. Wang, R\. Yu, G\. V\. den Broeck, and A\. Liu \(2026\)Breaking the factorization barrier in diffusion language models\.InForty\-third International Conference on Machine Learning,Cited by:[§5](https://arxiv.org/html/2608.10615#S5.SS0.SSS0.Px2.p1.1)\.
- H\. Lin, I\. Wu, and T\. Wei \(2013\)On specific 17\-clue sudoku puzzles\.ICGA Journal\.Cited by:[§E\.1](https://arxiv.org/html/2608.10615#A5.SS1.SSS0.Px1.p2.6),[§4\.1](https://arxiv.org/html/2608.10615#S4.SS1.SSS0.Px2.p3.16)\.
- A\. Lou, C\. Meng, and S\. Ermon \(2024\)Discrete diffusion modeling by estimating the ratios of the data distribution\.InInternational Conference on Machine Learning,Cited by:[§1](https://arxiv.org/html/2608.10615#S1.p1.1),[§5](https://arxiv.org/html/2608.10615#S5.SS0.SSS0.Px1.p1.1)\.
- G\. McGuire, B\. Tugemann, and G\. Civario \(2014\)There is no 16\-clue sudoku: solving the sudoku minimum number of clues problem via hitting set enumeration\.Experimental Mathematics\.Cited by:[§4\.1](https://arxiv.org/html/2608.10615#S4.SS1.SSS0.Px2.p3.16)\.
- J\. Ou, S\. Nie, K\. Xue, F\. Zhu, J\. Sun, Z\. Li, and C\. Li \(2025\)Your absorbing discrete diffusion secretly models the conditional distributions of clean data\.InThe Thirteenth International Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2608.10615#S1.p2.1),[§5](https://arxiv.org/html/2608.10615#S5.SS0.SSS0.Px1.p1.1)\.
- W\. Peebles and S\. Xie \(2023\)Scalable diffusion models with transformers\.InProceedings of the IEEE/CVF international conference on computer vision,Cited by:[§4\.1](https://arxiv.org/html/2608.10615#S4.SS1.SSS0.Px1.p1.9)\.
- P\. Pynadath, J\. Shi, and R\. Zhang \(2025\)CANDI: hybrid discrete\-continuous diffusion models\.Cited by:[§4\.3](https://arxiv.org/html/2608.10615#S4.SS3.p1.5),[Table 1](https://arxiv.org/html/2608.10615#S4.T1.3.1.3.1.1),[§5](https://arxiv.org/html/2608.10615#S5.SS0.SSS0.Px2.p1.1)\.
- A\. Radford, J\. Wu, R\. Child, D\. Luan, D\. Amodei, and I\. Sutskever \(2019\)Language models are unsupervised multitask learners\.OpenAI Technical Report\.Cited by:[§4\.1](https://arxiv.org/html/2608.10615#S4.SS1.SSS0.Px1.p1.9),[§4\.1](https://arxiv.org/html/2608.10615#S4.SS1.SSS0.Px3.p1.2)\.
- P\. H\. Richemond, S\. Dieleman, and A\. Doucet \(2023\)Categorical SDEs with simplex diffusion\.InICML 2023 Workshop: Sampling and Optimization in Discrete Space,Cited by:[§5](https://arxiv.org/html/2608.10615#S5.SS0.SSS0.Px3.p1.1)\.
- S\. S\. Sahoo, M\. Arriola, A\. Gokaslan, E\. M\. Marroquin, A\. M\. Rush, Y\. Schiff, J\. T\. Chiu, and V\. Kuleshov \(2024\)Simple and effective masked diffusion language models\.InThe Thirty\-eighth Annual Conference on Neural Information Processing Systems,Cited by:[§1](https://arxiv.org/html/2608.10615#S1.p2.1),[§4\.1](https://arxiv.org/html/2608.10615#S4.SS1.SSS0.Px1.p1.9),[§4\.3](https://arxiv.org/html/2608.10615#S4.SS3.p1.5),[Table 1](https://arxiv.org/html/2608.10615#S4.T1.3.1.5.3.1),[Table 2](https://arxiv.org/html/2608.10615#S4.T2.15.6.4.1),[§5](https://arxiv.org/html/2608.10615#S5.SS0.SSS0.Px1.p1.1)\.
- S\. S\. Sahoo, J\. Deschenaux, A\. Gokaslan, G\. Wang, J\. T\. Chiu, and V\. Kuleshov \(2025\)The diffusion duality\.InForty\-second International Conference on Machine Learning,Cited by:[§1](https://arxiv.org/html/2608.10615#S1.p2.1),[§4\.3](https://arxiv.org/html/2608.10615#S4.SS3.p1.5),[Table 1](https://arxiv.org/html/2608.10615#S4.T1.3.1.6.4.1),[Table 2](https://arxiv.org/html/2608.10615#S4.T2.15.4.2.1),[§5](https://arxiv.org/html/2608.10615#S5.SS0.SSS0.Px1.p1.1),[§5](https://arxiv.org/html/2608.10615#S5.SS0.SSS0.Px2.p1.1)\.
- S\. S\. Sahoo, J\. Lemercier, Z\. Yang, J\. Deschenaux, J\. Liu, J\. Thickstun, and A\. Jukic \(2026\)Scaling beyond masked diffusion language models\.InThe Forty\-Third International Conference on Machine Learning,Cited by:[§1](https://arxiv.org/html/2608.10615#S1.p2.1)\.
- Y\. Schiff, S\. S\. Sahoo, H\. Phung, G\. Wang, S\. Boshar, H\. Dalla\-torre, B\. P\. de Almeida, A\. M\. Rush, T\. PIERROT, and V\. Kuleshov \(2025\)Simple guidance mechanisms for discrete diffusion models\.InThe Thirteenth International Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2608.10615#S1.p2.1),[§3\.2](https://arxiv.org/html/2608.10615#S3.SS2.p8.2),[§4\.3](https://arxiv.org/html/2608.10615#S4.SS3.p1.5),[Table 1](https://arxiv.org/html/2608.10615#S4.T1.3.1.4.2.1)\.
- J\. Shi, K\. Han, Z\. Wang, A\. Doucet, and M\. Titsias \(2024\)Simplified and generalized masked diffusion for discrete data\.InThe Thirty\-eighth Annual Conference on Neural Information Processing Systems,Cited by:[§1](https://arxiv.org/html/2608.10615#S1.p2.1),[§5](https://arxiv.org/html/2608.10615#S5.SS0.SSS0.Px1.p1.1)\.
- H\. Stark, B\. Jing, C\. Wang, G\. Corso, B\. Berger, R\. Barzilay, and T\. Jaakkola \(2024\)Dirichlet flow matching with applications to DNA sequence design\.InForty\-first International Conference on Machine Learning,Cited by:[§5](https://arxiv.org/html/2608.10615#S5.SS0.SSS0.Px3.p1.1)\.
- J\. Su, M\. Ahmed, Y\. Lu, S\. Pan, W\. Bo, and Y\. Liu \(2024\)Roformer: enhanced transformer with rotary position embedding\.Neurocomputing568\.Cited by:[§4\.1](https://arxiv.org/html/2608.10615#S4.SS1.SSS0.Px1.p1.9)\.
- H\. Touvron, L\. Martin, K\. Stone, P\. Albert, A\. Almahairi, Y\. Babaei, N\. Bashlykov, S\. Batra, P\. Bhargava, S\. Bhosale, D\. Bikel, L\. Blecher, C\. Canton Ferrer, M\. Chen, G\. Cucurull, D\. Esiobu, J\. Fernandes, J\. Fu, W\. Fu, B\. Fuller, C\. Gao, V\. Goswami, N\. Goyal, A\. Hartshorn, S\. Hosseini, R\. Hou, H\. Inan, M\. Kardas, V\. Kerkez, M\. Khabsa, I\. Kloumann, A\. Korenev, P\. S\. Koura, M\. Lachaux, T\. Lavril, J\. Lee, D\. Liskovich, Y\. Lu, Y\. Mao, X\. Martinet, T\. Mihaylov, P\. Mishra, I\. Molybog, Y\. Nie, A\. Poulton, J\. Reizenstein, R\. Rungta, K\. Saladi, A\. Schelten, R\. Silva, E\. M\. Smith, R\. Subramanian, X\. E\. Tan, B\. Tang, R\. Taylor, A\. Williams, J\. X\. Kuan, P\. Xu, Z\. Yan, I\. Zarov, Y\. Zhang, A\. Fan, M\. Kambadur, S\. Narang, A\. Rodriguez, R\. Stojnic, S\. Edunov, and T\. Scialom \(2023\)Llama 2: open foundation and fine\-tuned chat models\.arXiv preprint arXiv:2307\.09288\.Cited by:[§4\.1](https://arxiv.org/html/2608.10615#S4.SS1.SSS0.Px3.p1.2)\.
- D\. von Rütte, J\. Fluri, Y\. Ding, A\. Orvieto, B\. Schölkopf, and T\. Hofmann \(2025\)Generalized interpolating discrete diffusion\.InForty\-second International Conference on Machine Learning,Cited by:[§1](https://arxiv.org/html/2608.10615#S1.p2.1)\.
- D\. von Rütte, J\. Fluri, O\. Pooladzandi, B\. Schölkopf, T\. Hofmann, and A\. Orvieto \(2026\)Scaling behavior of discrete diffusion language models\.InThe Fourteenth International Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2608.10615#S1.p2.1)\.
- T\. Xie, S\. Xue, Z\. Feng, T\. Hu, J\. Sun, Z\. Li, and C\. Zhang \(2026\)Variational autoencoding discrete diffusion with enhanced dimensional correlations modeling\.InThe Fourteenth International Conference on Learning Representations,Cited by:[§5](https://arxiv.org/html/2608.10615#S5.SS0.SSS0.Px2.p1.1)\.
- R\. Zhang, S\. Zhai, Y\. Zhang, J\. Thornton, Z\. Ou, J\. M\. Susskind, and N\. Jaitly \(2025\)Target concrete score matching: a holistic framework for discrete diffusion\.InForty\-second International Conference on Machine Learning,Cited by:[§1](https://arxiv.org/html/2608.10615#S1.p1.1),[§5](https://arxiv.org/html/2608.10615#S5.SS0.SSS0.Px1.p1.1)\.
- H\. Zheng, S\. Gong, R\. Zhang, T\. Chen, J\. Gu, M\. Zhou, N\. Jaitly, and Y\. Zhang \(2026\)Continuously augmented discrete diffusion model for categorical generative modeling\.InThe Fourteenth International Conference on Learning Representations,Cited by:[§5](https://arxiv.org/html/2608.10615#S5.SS0.SSS0.Px2.p1.1)\.
- K\. Zheng, Y\. Chen, H\. Mao, M\. Liu, J\. Zhu, and Q\. Zhang \(2025\)Masked diffusion models are secretly time\-agnostic masked models and exploit inaccurate categorical sampling\.InThe Thirteenth International Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2608.10615#S1.p2.1),[§5](https://arxiv.org/html/2608.10615#S5.SS0.SSS0.Px1.p1.1)\.

Simplex Relaxation for Discrete Diffusion

Supplementary Material

## Appendix AAuxiliary Identities and Full Dirichlet–Categorical Hierarchy

This appendix develops the exact probabilistic structure behind the simplex relaxation\. The formulas below apply at positive diffusion times, where𝐩t\\mathbf\{p\}\_\{t\}is strictly positive under the assumptions in[Section˜2](https://arxiv.org/html/2608.10615#S2)\. We define the clean endpoint separately as the categorical state𝐳0=𝐱\\mathbf\{z\}\_\{0\}=\\mathbf\{x\}; the hierarchy does not introduce a Dirichlet variable𝐰0\\mathbf\{w\}\_\{0\}\.

The key point is that the auxiliary state𝐰t\\mathbf\{w\}\_\{t\}is not introduced as a heuristic soft surrogate\. Rather, once we specify the Dirichlet bridge

q​\(𝐰t∣𝐳t,𝐱\)=Dir⁡\(𝐰t;ηt​𝐩t\+𝐳t\),q\(\\mathbf\{w\}\_\{t\}\\mid\\mathbf\{z\}\_\{t\},\\mathbf\{x\}\)=\\operatorname\{Dir\}\\\!\\left\(\\mathbf\{w\}\_\{t\};\\eta\_\{t\}\\mathbf\{p\}\_\{t\}\+\\mathbf\{z\}\_\{t\}\\right\),the resulting joint model admits a closed hierarchy in both directions: the relaxed state has an exact Dirichlet marginal, the discrete state can be decoded exactly from𝐰t\\mathbf\{w\}\_\{t\}, and the reverse bridge remains tractable after marginalizing either𝐳t\\mathbf\{z\}\_\{t\}or𝐳s\\mathbf\{z\}\_\{s\}\. We begin with two elementary identities that make these cancellations possible\.

### A\.1Useful identities

The first identity is the basic shift formula for Dirichlet densities\. It shows that adding a one\-hot count to the concentration vector simply multiplies the base Dirichlet density by the corresponding simplex coordinate\.

###### Lemma 1\.

Let𝛂∈ℝ\>0K\\bm\{\\alpha\}\\in\\mathbb\{R\}\_\{\>0\}^\{K\}and letα0=∑i=1Kαi\\alpha\_\{0\}=\\sum\_\{i=1\}^\{K\}\\alpha\_\{i\}\. Then, for anyk∈\{1,…,K\}k\\in\\\{1,\\dots,K\\\},

Dir⁡\(𝐰;𝜶\+𝐞k\)=α0αk​wk​Dir⁡\(𝐰;𝜶\)\.\\operatorname\{Dir\}\(\\mathbf\{w\};\\bm\{\\alpha\}\+\\mathbf\{e\}\_\{k\}\)=\\frac\{\\alpha\_\{0\}\}\{\\alpha\_\{k\}\}\\,w\_\{k\}\\,\\operatorname\{Dir\}\(\\mathbf\{w\};\\bm\{\\alpha\}\)\.\(22\)

###### Proof\.

By definition,

Dir⁡\(𝐰;𝜶\)=1B​\(𝜶\)​∏i=1Kwiαi−1,B​\(𝜶\)=∏i=1KΓ​\(αi\)Γ​\(α0\)\.\\operatorname\{Dir\}\(\\mathbf\{w\};\\bm\{\\alpha\}\)=\\frac\{1\}\{B\(\\bm\{\\alpha\}\)\}\\prod\_\{i=1\}^\{K\}w\_\{i\}^\{\\alpha\_\{i\}\-1\},\\qquad B\(\\bm\{\\alpha\}\)=\\frac\{\\prod\_\{i=1\}^\{K\}\\Gamma\(\\alpha\_\{i\}\)\}\{\\Gamma\(\\alpha\_\{0\}\)\}\.Hence

Dir⁡\(𝐰;𝜶\+𝐞k\)=1B​\(𝜶\+𝐞k\)​wk​∏i=1Kwiαi−1\.\\operatorname\{Dir\}\(\\mathbf\{w\};\\bm\{\\alpha\}\+\\mathbf\{e\}\_\{k\}\)=\\frac\{1\}\{B\(\\bm\{\\alpha\}\+\\mathbf\{e\}\_\{k\}\)\}w\_\{k\}\\prod\_\{i=1\}^\{K\}w\_\{i\}^\{\\alpha\_\{i\}\-1\}\.It remains to compare the normalizing constants:

B​\(𝜶\+𝐞k\)B​\(𝜶\)=Γ​\(αk\+1\)Γ​\(αk\)​Γ​\(α0\)Γ​\(α0\+1\)=αkα0\.\\frac\{B\(\\bm\{\\alpha\}\+\\mathbf\{e\}\_\{k\}\)\}\{B\(\\bm\{\\alpha\}\)\}=\\frac\{\\Gamma\(\\alpha\_\{k\}\+1\)\}\{\\Gamma\(\\alpha\_\{k\}\)\}\\frac\{\\Gamma\(\\alpha\_\{0\}\)\}\{\\Gamma\(\\alpha\_\{0\}\+1\)\}=\\frac\{\\alpha\_\{k\}\}\{\\alpha\_\{0\}\}\.Therefore

1B​\(𝜶\+𝐞k\)=α0αk​1B​\(𝜶\),\\frac\{1\}\{B\(\\bm\{\\alpha\}\+\\mathbf\{e\}\_\{k\}\)\}=\\frac\{\\alpha\_\{0\}\}\{\\alpha\_\{k\}\}\\frac\{1\}\{B\(\\bm\{\\alpha\}\)\},which proves \([22](https://arxiv.org/html/2608.10615#A1.E22)\)\. ∎

Specializing this identity to concentrations of the formη​𝐩\\eta\\mathbf\{p\}yields the cancellation that will be used throughout the appendix\.

###### Corollary 1\.

Let𝐩∈ΔK−1\\mathbf\{p\}\\in\\Delta^\{K\-1\}satisfypk\>0p\_\{k\}\>0for allkk, letη\>0\\eta\>0, and define𝛂=η​𝐩\\bm\{\\alpha\}=\\eta\\mathbf\{p\}\. Then

pk​Dir⁡\(𝐰;η​𝐩\+𝐞k\)=wk​Dir⁡\(𝐰;η​𝐩\)\.p\_\{k\}\\,\\operatorname\{Dir\}\(\\mathbf\{w\};\\eta\\mathbf\{p\}\+\\mathbf\{e\}\_\{k\}\)=w\_\{k\}\\,\\operatorname\{Dir\}\(\\mathbf\{w\};\\eta\\mathbf\{p\}\)\.\(23\)

###### Proof\.

Apply[Lemma˜1](https://arxiv.org/html/2608.10615#Thmlemma1)with𝜶=η​𝐩\\bm\{\\alpha\}=\\eta\\mathbf\{p\}\. Sinceα0=η\\alpha\_\{0\}=\\etaandαk=η​pk\\alpha\_\{k\}=\\eta p\_\{k\},

Dir⁡\(𝐰;η​𝐩\+𝐞k\)=ηη​pk​wk​Dir⁡\(𝐰;η​𝐩\)=wkpk​Dir⁡\(𝐰;η​𝐩\)\.\\operatorname\{Dir\}\(\\mathbf\{w\};\\eta\\mathbf\{p\}\+\\mathbf\{e\}\_\{k\}\)=\\frac\{\\eta\}\{\\eta p\_\{k\}\}w\_\{k\}\\operatorname\{Dir\}\(\\mathbf\{w\};\\eta\\mathbf\{p\}\)=\\frac\{w\_\{k\}\}\{p\_\{k\}\}\\operatorname\{Dir\}\(\\mathbf\{w\};\\eta\\mathbf\{p\}\)\.Multiplying both sides bypkp\_\{k\}gives \([23](https://arxiv.org/html/2608.10615#A1.E23)\)\. ∎

The content of[Corollary˜1](https://arxiv.org/html/2608.10615#Thmcorollary1)is simple but important: a categorical mixture over one\-hot shifts of a Dirichlet distribution collapses back to the unshifted Dirichlet density\. This is precisely the mechanism that makes the simplex relaxation exact rather than approximate\.

### A\.2Full Dirichlet–categorical hierarchy

We now return to the joint factorization

q​\(𝐱,𝐳s,𝐳t,𝐰s,𝐰t\)=q​\(𝐱\)​q​\(𝐳s∣𝐱\)​q​\(𝐰s∣𝐳s,𝐱\)​q​\(𝐳t∣𝐳s\)​q​\(𝐰t∣𝐳t,𝐱\),q\(\\mathbf\{x\},\\mathbf\{z\}\_\{s\},\\mathbf\{z\}\_\{t\},\\mathbf\{w\}\_\{s\},\\mathbf\{w\}\_\{t\}\)=q\(\\mathbf\{x\}\)\\,q\(\\mathbf\{z\}\_\{s\}\\mid\\mathbf\{x\}\)\\,q\(\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{z\}\_\{s\},\\mathbf\{x\}\)\\,q\(\\mathbf\{z\}\_\{t\}\\mid\\mathbf\{z\}\_\{s\}\)\\,q\(\\mathbf\{w\}\_\{t\}\\mid\\mathbf\{z\}\_\{t\},\\mathbf\{x\}\),\(24\)together with

q​\(𝐰t∣𝐳t,𝐱\)=Dir⁡\(𝐰t;ηt​𝐩t\+𝐳t\)\.q\(\\mathbf\{w\}\_\{t\}\\mid\\mathbf\{z\}\_\{t\},\\mathbf\{x\}\)=\\operatorname\{Dir\}\\\!\\left\(\\mathbf\{w\}\_\{t\};\\eta\_\{t\}\\mathbf\{p\}\_\{t\}\+\\mathbf\{z\}\_\{t\}\\right\)\.\(25\)
The next proposition summarizes the full hierarchy induced by this construction\. The first two statements identify the exact marginal and exact decoder at timett\. The third lifts the standard discrete reverse posterior from𝐳t\\mathbf\{z\}\_\{t\}to𝐰t\\mathbf\{w\}\_\{t\}\. The last two show that, once this lift is performed, the reverse bridge over relaxed states becomes a mixture of shifted Dirichlet components\.

###### Proposition 4\.

Assumet\>0t\>0\. The Dirichlet–categorical hierarchy satisfies the following properties; statements involving𝐰s\\mathbf\{w\}\_\{s\}additionally requires\>0s\>0\.

1. 1\.The marginal distribution of the relaxed state is q​\(𝐰t∣𝐱\)=Dir⁡\(𝐰t;ηt​𝐩t\)\.q\(\\mathbf\{w\}\_\{t\}\\mid\\mathbf\{x\}\)=\\operatorname\{Dir\}\\\!\\left\(\\mathbf\{w\}\_\{t\};\\eta\_\{t\}\\mathbf\{p\}\_\{t\}\\right\)\.\(26\)
2. 2\.Given𝐰t\\mathbf\{w\}\_\{t\}, the variables𝐱\\mathbf\{x\}and𝐳t\\mathbf\{z\}\_\{t\}are conditionally independent, and the discrete state can be recovered from the relaxed state via q​\(𝐳t∣𝐰t,𝐱\)=q​\(𝐳t∣𝐰t\)=Cat⁡\(𝐳t;𝐰t\)\.q\(\\mathbf\{z\}\_\{t\}\\mid\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)=q\(\\mathbf\{z\}\_\{t\}\\mid\\mathbf\{w\}\_\{t\}\)=\\operatorname\{Cat\}\\\!\\left\(\\mathbf\{z\}\_\{t\};\\mathbf\{w\}\_\{t\}\\right\)\.\(27\)
3. 3\.Fors<ts<t, the reverse conditional posterior of𝐳s\\mathbf\{z\}\_\{s\}given𝐰t\\mathbf\{w\}\_\{t\}and𝐱\\mathbf\{x\}is categorical: q​\(𝐳s∣𝐰t,𝐱\)=Cat⁡\(𝐳s;𝝆s∣t​\(𝐱,𝐰t\)\),q\(\\mathbf\{z\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)=\\operatorname\{Cat\}\\\!\\left\(\\mathbf\{z\}\_\{s\};\\bm\{\\rho\}\_\{s\\mid t\}\(\\mathbf\{x\},\\mathbf\{w\}\_\{t\}\)\\right\),\(28\)where 𝝆s∣t​\(𝐱,𝐰t\)≔𝐩s⊙\[αt∣s​\(𝐰t⊘𝐩t\)\+\(1−αt∣s\)​⟨𝐰t,𝝅⊘𝐩t⟩​𝟏\]\.\\bm\{\\rho\}\_\{s\\mid t\}\(\\mathbf\{x\},\\mathbf\{w\}\_\{t\}\)\\coloneq\\mathbf\{p\}\_\{s\}\\odot\\left\[\\alpha\_\{t\\mid s\}\\bigl\(\\mathbf\{w\}\_\{t\}\\oslash\\mathbf\{p\}\_\{t\}\\bigr\)\+\(1\-\\alpha\_\{t\\mid s\}\)\\,\\left\\langle\\mathbf\{w\}\_\{t\},\\,\\bm\{\\pi\}\\oslash\\mathbf\{p\}\_\{t\}\\right\\rangle\\mathbf\{1\}\\right\]\.\(29\)
4. 4\.For0<s<t0<s<t, the reverse conditional posterior of𝐰s\\mathbf\{w\}\_\{s\}given𝐳t\\mathbf\{z\}\_\{t\}and𝐱\\mathbf\{x\}is a Dirichlet mixture: q​\(𝐰s∣𝐳t,𝐱\)=∑k=1Krs∣t,k​\(𝐱,𝐳t\)​Dir⁡\(𝐰s;ηs​𝐩s\+𝐞k\),q\(\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{z\}\_\{t\},\\mathbf\{x\}\)=\\sum\_\{k=1\}^\{K\}r\_\{s\\mid t,k\}\(\\mathbf\{x\},\\mathbf\{z\}\_\{t\}\)\\,\\operatorname\{Dir\}\\\!\\left\(\\mathbf\{w\}\_\{s\};\\eta\_\{s\}\\mathbf\{p\}\_\{s\}\+\\mathbf\{e\}\_\{k\}\\right\),\(30\)wherers∣t,k​\(𝐱,𝐳t\)r\_\{s\\mid t,k\}\(\\mathbf\{x\},\\mathbf\{z\}\_\{t\}\)denotes thekk\-th component of𝐫s∣t​\(𝐱,𝐳t\)\\mathbf\{r\}\_\{s\\mid t\}\(\\mathbf\{x\},\\mathbf\{z\}\_\{t\}\)defined in[Equation˜4](https://arxiv.org/html/2608.10615#S2.E4)\.
5. 5\.For0<s<t0<s<t, the reverse conditional posterior of𝐰s\\mathbf\{w\}\_\{s\}given𝐰t\\mathbf\{w\}\_\{t\}and𝐱\\mathbf\{x\}is a Dirichlet mixture: q​\(𝐰s∣𝐰t,𝐱\)=∑k=1Kρs∣t,k​\(𝐱,𝐰t\)​Dir⁡\(𝐰s;ηs​𝐩s\+𝐞k\)\.q\(\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)=\\sum\_\{k=1\}^\{K\}\\rho\_\{s\\mid t,k\}\(\\mathbf\{x\},\\mathbf\{w\}\_\{t\}\)\\,\\operatorname\{Dir\}\\\!\\left\(\\mathbf\{w\}\_\{s\};\\eta\_\{s\}\\mathbf\{p\}\_\{s\}\+\\mathbf\{e\}\_\{k\}\\right\)\.\(31\)

###### Proof\.

1. 1\.We begin with the marginal law of𝐰t\\mathbf\{w\}\_\{t\}\. Marginalizing the discrete latent𝐳t∼Cat⁡\(𝐩t\)\\mathbf\{z\}\_\{t\}\\sim\\operatorname\{Cat\}\(\\mathbf\{p\}\_\{t\}\)from the conditional bridge q​\(𝐰t∣𝐳t,𝐱\)=Dir⁡\(𝐰t;ηt​𝐩t\+𝐳t\)q\(\\mathbf\{w\}\_\{t\}\\mid\\mathbf\{z\}\_\{t\},\\mathbf\{x\}\)=\\operatorname\{Dir\}\\\!\\left\(\\mathbf\{w\}\_\{t\};\\eta\_\{t\}\\mathbf\{p\}\_\{t\}\+\\mathbf\{z\}\_\{t\}\\right\)gives q​\(𝐰t∣𝐱\)\\displaystyle q\(\\mathbf\{w\}\_\{t\}\\mid\\mathbf\{x\}\)=∑k=1Kq​\(𝐰t∣𝐳t=𝐞k,𝐱\)​q​\(𝐳t=𝐞k∣𝐱\)\\displaystyle=\\sum\_\{k=1\}^\{K\}q\(\\mathbf\{w\}\_\{t\}\\mid\\mathbf\{z\}\_\{t\}=\\mathbf\{e\}\_\{k\},\\mathbf\{x\}\)\\,q\(\\mathbf\{z\}\_\{t\}=\\mathbf\{e\}\_\{k\}\\mid\\mathbf\{x\}\)=∑k=1KDir⁡\(𝐰t;ηt​𝐩t\+𝐞k\)​pt,k​\(𝐱\)\.\\displaystyle=\\sum\_\{k=1\}^\{K\}\\operatorname\{Dir\}\\\!\\left\(\\mathbf\{w\}\_\{t\};\\eta\_\{t\}\\mathbf\{p\}\_\{t\}\+\\mathbf\{e\}\_\{k\}\\right\)p\_\{t,k\}\(\\mathbf\{x\}\)\.\(32\)Now[Corollary˜1](https://arxiv.org/html/2608.10615#Thmcorollary1)applies termwise: pt,k​\(𝐱\)​Dir⁡\(𝐰t;ηt​𝐩t\+𝐞k\)=wt,k​Dir⁡\(𝐰t;ηt​𝐩t\)\.p\_\{t,k\}\(\\mathbf\{x\}\)\\operatorname\{Dir\}\\\!\\left\(\\mathbf\{w\}\_\{t\};\\eta\_\{t\}\\mathbf\{p\}\_\{t\}\+\\mathbf\{e\}\_\{k\}\\right\)=w\_\{t,k\}\\operatorname\{Dir\}\\\!\\left\(\\mathbf\{w\}\_\{t\};\\eta\_\{t\}\\mathbf\{p\}\_\{t\}\\right\)\.Substituting this into \([32](https://arxiv.org/html/2608.10615#A1.E32)\) collapses the mixture: q​\(𝐰t∣𝐱\)\\displaystyle q\(\\mathbf\{w\}\_\{t\}\\mid\\mathbf\{x\}\)=∑k=1Kwt,k​Dir⁡\(𝐰t;ηt​𝐩t\)\\displaystyle=\\sum\_\{k=1\}^\{K\}w\_\{t,k\}\\operatorname\{Dir\}\\\!\\left\(\\mathbf\{w\}\_\{t\};\\eta\_\{t\}\\mathbf\{p\}\_\{t\}\\right\)=\(∑k=1Kwt,k\)​Dir⁡\(𝐰t;ηt​𝐩t\)\\displaystyle=\\left\(\\sum\_\{k=1\}^\{K\}w\_\{t,k\}\\right\)\\operatorname\{Dir\}\\\!\\left\(\\mathbf\{w\}\_\{t\};\\eta\_\{t\}\\mathbf\{p\}\_\{t\}\\right\)=Dir⁡\(𝐰t;ηt​𝐩t\),\\displaystyle=\\operatorname\{Dir\}\\\!\\left\(\\mathbf\{w\}\_\{t\};\\eta\_\{t\}\\mathbf\{p\}\_\{t\}\\right\),\(33\)which proves \([26](https://arxiv.org/html/2608.10615#A1.E26)\)\.
2. 2\.We next derive the exact decoder from𝐰t\\mathbf\{w\}\_\{t\}back to𝐳t\\mathbf\{z\}\_\{t\}\. Fixk∈\{1,…,K\}k\\in\\\{1,\\dots,K\\\}\. By Bayes’ rule, q​\(𝐳t=𝐞k∣𝐰t,𝐱\)\\displaystyle q\(\\mathbf\{z\}\_\{t\}=\\mathbf\{e\}\_\{k\}\\mid\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)=q​\(𝐰t∣𝐳t=𝐞k,𝐱\)​q​\(𝐳t=𝐞k∣𝐱\)q​\(𝐰t∣𝐱\)\.\\displaystyle=\\frac\{q\(\\mathbf\{w\}\_\{t\}\\mid\\mathbf\{z\}\_\{t\}=\\mathbf\{e\}\_\{k\},\\mathbf\{x\}\)\\,q\(\\mathbf\{z\}\_\{t\}=\\mathbf\{e\}\_\{k\}\\mid\\mathbf\{x\}\)\}\{q\(\\mathbf\{w\}\_\{t\}\\mid\\mathbf\{x\}\)\}\.\(34\)Using the previous computation together with[Corollary˜1](https://arxiv.org/html/2608.10615#Thmcorollary1), the numerator becomes q​\(𝐰t∣𝐳t=𝐞k,𝐱\)​q​\(𝐳t=𝐞k∣𝐱\)=wt,k​Dir⁡\(𝐰t;ηt​𝐩t\),q\(\\mathbf\{w\}\_\{t\}\\mid\\mathbf\{z\}\_\{t\}=\\mathbf\{e\}\_\{k\},\\mathbf\{x\}\)\\,q\(\\mathbf\{z\}\_\{t\}=\\mathbf\{e\}\_\{k\}\\mid\\mathbf\{x\}\)=w\_\{t,k\}\\operatorname\{Dir\}\\\!\\left\(\\mathbf\{w\}\_\{t\};\\eta\_\{t\}\\mathbf\{p\}\_\{t\}\\right\),while the denominator is exactlyDir⁡\(𝐰t;ηt​𝐩t\)\\operatorname\{Dir\}\(\\mathbf\{w\}\_\{t\};\\eta\_\{t\}\\mathbf\{p\}\_\{t\}\)\. Hence q​\(𝐳t=𝐞k∣𝐰t,𝐱\)=wt,k\.q\(\\mathbf\{z\}\_\{t\}=\\mathbf\{e\}\_\{k\}\\mid\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)=w\_\{t,k\}\.Since this holds for everykk, we obtain q​\(𝐳t∣𝐰t,𝐱\)=Cat⁡\(𝐳t;𝐰t\),q\(\\mathbf\{z\}\_\{t\}\\mid\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)=\\operatorname\{Cat\}\(\\mathbf\{z\}\_\{t\};\\mathbf\{w\}\_\{t\}\),and the right\-hand side no longer depends on𝐱\\mathbf\{x\}\. This proves \([27](https://arxiv.org/html/2608.10615#A1.E27)\)\.
3. 3\.We now lift the discrete reverse posterior from𝐳t\\mathbf\{z\}\_\{t\}to𝐰t\\mathbf\{w\}\_\{t\}\. Marginalizing over𝐳t\\mathbf\{z\}\_\{t\}gives q​\(𝐳s∣𝐰t,𝐱\)\\displaystyle q\(\\mathbf\{z\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)=∑j=1Kq​\(𝐳s∣𝐳t=𝐞j,𝐰t,𝐱\)​q​\(𝐳t=𝐞j∣𝐰t,𝐱\)\.\\displaystyle=\\sum\_\{j=1\}^\{K\}q\(\\mathbf\{z\}\_\{s\}\\mid\\mathbf\{z\}\_\{t\}=\\mathbf\{e\}\_\{j\},\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)\\,q\(\\mathbf\{z\}\_\{t\}=\\mathbf\{e\}\_\{j\}\\mid\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)\.\(35\)Conditioned on\(𝐳t,𝐱\)\(\\mathbf\{z\}\_\{t\},\\mathbf\{x\}\), the variable𝐳s\\mathbf\{z\}\_\{s\}is independent of𝐰t\\mathbf\{w\}\_\{t\}, so the first factor is just the usual reverse posterior q​\(𝐳s∣𝐳t=𝐞j,𝐰t,𝐱\)=q​\(𝐳s∣𝐳t=𝐞j,𝐱\)=Cat⁡\(𝐳s;𝐫s∣t​\(𝐱,𝐞j\)\)\.q\(\\mathbf\{z\}\_\{s\}\\mid\\mathbf\{z\}\_\{t\}=\\mathbf\{e\}\_\{j\},\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)=q\(\\mathbf\{z\}\_\{s\}\\mid\\mathbf\{z\}\_\{t\}=\\mathbf\{e\}\_\{j\},\\mathbf\{x\}\)=\\operatorname\{Cat\}\\\!\\left\(\\mathbf\{z\}\_\{s\};\\mathbf\{r\}\_\{s\\mid t\}\(\\mathbf\{x\},\\mathbf\{e\}\_\{j\}\)\\right\)\.The second factor is the exact decoder derived above: q​\(𝐳t=𝐞j∣𝐰t,𝐱\)=wt,j\.q\(\\mathbf\{z\}\_\{t\}=\\mathbf\{e\}\_\{j\}\\mid\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)=w\_\{t,j\}\.Therefore q​\(𝐳s∣𝐰t,𝐱\)=∑j=1Kwt,j​Cat⁡\(𝐳s;𝐫s∣t​\(𝐱,𝐞j\)\)\.q\(\\mathbf\{z\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)=\\sum\_\{j=1\}^\{K\}w\_\{t,j\}\\,\\operatorname\{Cat\}\\\!\\left\(\\mathbf\{z\}\_\{s\};\\mathbf\{r\}\_\{s\\mid t\}\(\\mathbf\{x\},\\mathbf\{e\}\_\{j\}\)\\right\)\.\(36\)A mixture of categorical distributions is again categorical, with parameter vector equal to the same convex combination of the component parameters\. Thus q​\(𝐳s∣𝐰t,𝐱\)=Cat⁡\(𝐳s;𝝆s∣t​\(𝐱,𝐰t\)\),𝝆s∣t​\(𝐱,𝐰t\)=∑j=1Kwt,j​𝐫s∣t​\(𝐱,𝐞j\)\.q\(\\mathbf\{z\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)=\\operatorname\{Cat\}\\\!\\left\(\\mathbf\{z\}\_\{s\};\\bm\{\\rho\}\_\{s\\mid t\}\(\\mathbf\{x\},\\mathbf\{w\}\_\{t\}\)\\right\),\\qquad\\bm\{\\rho\}\_\{s\\mid t\}\(\\mathbf\{x\},\\mathbf\{w\}\_\{t\}\)=\\sum\_\{j=1\}^\{K\}w\_\{t,j\}\\,\\mathbf\{r\}\_\{s\\mid t\}\(\\mathbf\{x\},\\mathbf\{e\}\_\{j\}\)\.To obtain the closed form, substitute[Equation˜4](https://arxiv.org/html/2608.10615#S2.E4)with𝐳t=𝐞j\\mathbf\{z\}\_\{t\}=\\mathbf\{e\}\_\{j\}: 𝐫s∣t​\(𝐱,𝐞j\)\\displaystyle\\mathbf\{r\}\_\{s\\mid t\}\(\\mathbf\{x\},\\mathbf\{e\}\_\{j\}\)=\[αt∣s​𝐞j\+\(1−αt∣s\)​⟨𝐞j,𝝅⟩​1\]⊙𝐩s⟨𝐞j,𝐩t⟩\\displaystyle=\\frac\{\\left\[\\alpha\_\{t\\mid s\}\\mathbf\{e\}\_\{j\}\+\(1\-\\alpha\_\{t\\mid s\}\)\\,\\left\\langle\\mathbf\{e\}\_\{j\},\\,\\bm\{\\pi\}\\right\\rangle\\,\\mathbf\{1\}\\right\]\\odot\\mathbf\{p\}\_\{s\}\}\{\\left\\langle\\mathbf\{e\}\_\{j\},\\,\\mathbf\{p\}\_\{t\}\\right\\rangle\}=𝐩s⊙\[αt∣s​𝐞jpt,j​\(𝐱\)\+\(1−αt∣s\)​πjpt,j​\(𝐱\)​𝟏\]\.\\displaystyle=\\mathbf\{p\}\_\{s\}\\odot\\left\[\\alpha\_\{t\\mid s\}\\frac\{\\mathbf\{e\}\_\{j\}\}\{p\_\{t,j\}\(\\mathbf\{x\}\)\}\+\(1\-\\alpha\_\{t\\mid s\}\)\\frac\{\\pi\_\{j\}\}\{p\_\{t,j\}\(\\mathbf\{x\}\)\}\\mathbf\{1\}\\right\]\.\(37\)Averaging this expression under the weightswt,jw\_\{t,j\}yields 𝝆s∣t​\(𝐱,𝐰t\)\\displaystyle\\bm\{\\rho\}\_\{s\\mid t\}\(\\mathbf\{x\},\\mathbf\{w\}\_\{t\}\)=𝐩s⊙\[αt∣s​∑j=1Kwt,j​𝐞jpt,j​\(𝐱\)\+\(1−αt∣s\)​∑j=1Kwt,j​πjpt,j​\(𝐱\)​𝟏\]\\displaystyle=\\mathbf\{p\}\_\{s\}\\odot\\left\[\\alpha\_\{t\\mid s\}\\sum\_\{j=1\}^\{K\}w\_\{t,j\}\\frac\{\\mathbf\{e\}\_\{j\}\}\{p\_\{t,j\}\(\\mathbf\{x\}\)\}\+\(1\-\\alpha\_\{t\\mid s\}\)\\sum\_\{j=1\}^\{K\}w\_\{t,j\}\\frac\{\\pi\_\{j\}\}\{p\_\{t,j\}\(\\mathbf\{x\}\)\}\\mathbf\{1\}\\right\]=𝐩s⊙\[αt∣s​\(𝐰t⊘𝐩t\)\+\(1−αt∣s\)​⟨𝐰t,𝝅⊘𝐩t⟩​𝟏\],\\displaystyle=\\mathbf\{p\}\_\{s\}\\odot\\left\[\\alpha\_\{t\\mid s\}\\bigl\(\\mathbf\{w\}\_\{t\}\\oslash\\mathbf\{p\}\_\{t\}\\bigr\)\+\(1\-\\alpha\_\{t\\mid s\}\)\\left\\langle\\mathbf\{w\}\_\{t\},\\,\\bm\{\\pi\}\\oslash\\mathbf\{p\}\_\{t\}\\right\\rangle\\mathbf\{1\}\\right\],\(38\)which proves \([28](https://arxiv.org/html/2608.10615#A1.E28)\) and \([29](https://arxiv.org/html/2608.10615#A1.E29)\)\.
4. 4\.We next derive the reverse bridge over relaxed states conditioned on𝐳t\\mathbf\{z\}\_\{t\}\. Since𝐰s\\mathbf\{w\}\_\{s\}is conditionally independent of𝐳t\\mathbf\{z\}\_\{t\}given\(𝐳s,𝐱\)\(\\mathbf\{z\}\_\{s\},\\mathbf\{x\}\), marginalizing over𝐳s\\mathbf\{z\}\_\{s\}gives q​\(𝐰s∣𝐳t,𝐱\)\\displaystyle q\(\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{z\}\_\{t\},\\mathbf\{x\}\)=∑k=1Kq​\(𝐰s∣𝐳s=𝐞k,𝐳t,𝐱\)​q​\(𝐳s=𝐞k∣𝐳t,𝐱\)\\displaystyle=\\sum\_\{k=1\}^\{K\}q\(\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{z\}\_\{s\}=\\mathbf\{e\}\_\{k\},\\mathbf\{z\}\_\{t\},\\mathbf\{x\}\)\\,q\(\\mathbf\{z\}\_\{s\}=\\mathbf\{e\}\_\{k\}\\mid\\mathbf\{z\}\_\{t\},\\mathbf\{x\}\)=∑k=1Kq​\(𝐰s∣𝐳s=𝐞k,𝐱\)​q​\(𝐳s=𝐞k∣𝐳t,𝐱\)\.\\displaystyle=\\sum\_\{k=1\}^\{K\}q\(\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{z\}\_\{s\}=\\mathbf\{e\}\_\{k\},\\mathbf\{x\}\)\\,q\(\\mathbf\{z\}\_\{s\}=\\mathbf\{e\}\_\{k\}\\mid\\mathbf\{z\}\_\{t\},\\mathbf\{x\}\)\.\(39\)The first factor is exactly the shifted Dirichlet bridge at timess: q​\(𝐰s∣𝐳s=𝐞k,𝐱\)=Dir⁡\(𝐰s;ηs​𝐩s\+𝐞k\)\.q\(\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{z\}\_\{s\}=\\mathbf\{e\}\_\{k\},\\mathbf\{x\}\)=\\operatorname\{Dir\}\\\!\\left\(\\mathbf\{w\}\_\{s\};\\eta\_\{s\}\\mathbf\{p\}\_\{s\}\+\\mathbf\{e\}\_\{k\}\\right\)\.The second factor is the discrete reverse posterior q​\(𝐳s=𝐞k∣𝐳t,𝐱\)=rs∣t,k​\(𝐱,𝐳t\)\.q\(\\mathbf\{z\}\_\{s\}=\\mathbf\{e\}\_\{k\}\\mid\\mathbf\{z\}\_\{t\},\\mathbf\{x\}\)=r\_\{s\\mid t,k\}\(\\mathbf\{x\},\\mathbf\{z\}\_\{t\}\)\.Substituting these identities into \([39](https://arxiv.org/html/2608.10615#A1.E39)\) yields q​\(𝐰s∣𝐳t,𝐱\)=∑k=1Krs∣t,k​\(𝐱,𝐳t\)​Dir⁡\(𝐰s;ηs​𝐩s\+𝐞k\),q\(\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{z\}\_\{t\},\\mathbf\{x\}\)=\\sum\_\{k=1\}^\{K\}r\_\{s\\mid t,k\}\(\\mathbf\{x\},\\mathbf\{z\}\_\{t\}\)\\,\\operatorname\{Dir\}\\\!\\left\(\\mathbf\{w\}\_\{s\};\\eta\_\{s\}\\mathbf\{p\}\_\{s\}\+\\mathbf\{e\}\_\{k\}\\right\),which proves \([30](https://arxiv.org/html/2608.10615#A1.E30)\)\.
5. 5\.Finally, we lift this mixture from𝐳t\\mathbf\{z\}\_\{t\}to𝐰t\\mathbf\{w\}\_\{t\}\. Since𝐰s\\mathbf\{w\}\_\{s\}is also conditionally independent of𝐰t\\mathbf\{w\}\_\{t\}given\(𝐳s,𝐱\)\(\\mathbf\{z\}\_\{s\},\\mathbf\{x\}\), we obtain q​\(𝐰s∣𝐰t,𝐱\)\\displaystyle q\(\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)=∑k=1Kq​\(𝐰s∣𝐳s=𝐞k,𝐰t,𝐱\)​q​\(𝐳s=𝐞k∣𝐰t,𝐱\)\\displaystyle=\\sum\_\{k=1\}^\{K\}q\(\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{z\}\_\{s\}=\\mathbf\{e\}\_\{k\},\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)\\,q\(\\mathbf\{z\}\_\{s\}=\\mathbf\{e\}\_\{k\}\\mid\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)=∑k=1Kq​\(𝐰s∣𝐳s=𝐞k,𝐱\)​q​\(𝐳s=𝐞k∣𝐰t,𝐱\)\.\\displaystyle=\\sum\_\{k=1\}^\{K\}q\(\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{z\}\_\{s\}=\\mathbf\{e\}\_\{k\},\\mathbf\{x\}\)\\,q\(\\mathbf\{z\}\_\{s\}=\\mathbf\{e\}\_\{k\}\\mid\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)\.\(40\)The first factor is againDir⁡\(𝐰s;ηs​𝐩s\+𝐞k\)\\operatorname\{Dir\}\(\\mathbf\{w\}\_\{s\};\\eta\_\{s\}\\mathbf\{p\}\_\{s\}\+\\mathbf\{e\}\_\{k\}\), while the second factor is now the lifted reverse posterior: q​\(𝐳s=𝐞k∣𝐰t,𝐱\)=ρs∣t,k​\(𝐱,𝐰t\)\.q\(\\mathbf\{z\}\_\{s\}=\\mathbf\{e\}\_\{k\}\\mid\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)=\\rho\_\{s\\mid t,k\}\(\\mathbf\{x\},\\mathbf\{w\}\_\{t\}\)\.Therefore q​\(𝐰s∣𝐰t,𝐱\)=∑k=1Kρs∣t,k​\(𝐱,𝐰t\)​Dir⁡\(𝐰s;ηs​𝐩s\+𝐞k\),q\(\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)=\\sum\_\{k=1\}^\{K\}\\rho\_\{s\\mid t,k\}\(\\mathbf\{x\},\\mathbf\{w\}\_\{t\}\)\\,\\operatorname\{Dir\}\\\!\\left\(\\mathbf\{w\}\_\{s\};\\eta\_\{s\}\\mathbf\{p\}\_\{s\}\+\\mathbf\{e\}\_\{k\}\\right\),which proves \([31](https://arxiv.org/html/2608.10615#A1.E31)\)\.

∎

Taken together, these identities show that the simplex variable sits inside the discrete diffusion process in a fully coherent way\. The relaxed state has the correct Dirichlet marginal, the discrete state can be decoded from it exactly, and the reverse bridge remains available in closed form after replacing either𝐳t\\mathbf\{z\}\_\{t\}or𝐳s\\mathbf\{z\}\_\{s\}by their simplex\-conditioned posteriors\. This is the structural reason that the later training objectives and samplers can be written directly in terms of𝐰t\\mathbf\{w\}\_\{t\}without abandoning the original categorical process\.

## Appendix BRao–Blackwellized Reverse\-Bridge Objective

This appendix proves the closed form of the relaxed discrete bridge objective used in the main text\. The auxiliary decoder sample𝐳~t\\widetilde\{\\mathbf\{z\}\}\_\{t\}is marginalized analytically; the independently sampled denoiser input𝐳t\\mathbf\{z\}\_\{t\}is not marginalized\.

###### Proposition 5\.

For0≤s<t≤10\\leq s<t\\leq 1witht\>0t\>0, the relaxed discrete bridge objective satisfies

ℒ¯zs∣zt,wt​\(𝐰t,𝐱^θ,𝐱;s,t\)=⟨𝐰t,log⁡𝐩^t−log⁡𝐩t⟩\+⟨𝝆s∣t​\(𝐱,𝐰t\),log⁡𝐩s−log⁡𝐩^s⟩\.\\bar\{\\mathcal\{L\}\}\_\{z\_\{s\}\\mid z\_\{t\},w\_\{t\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\};s,t\)=\\left\\langle\\mathbf\{w\}\_\{t\},\\,\\log\\hat\{\\mathbf\{p\}\}\_\{t\}\-\\log\\mathbf\{p\}\_\{t\}\\right\\rangle\+\\left\\langle\\bm\{\\rho\}\_\{s\\mid t\}\(\\mathbf\{x\},\\mathbf\{w\}\_\{t\}\),\\,\\log\\mathbf\{p\}\_\{s\}\-\\log\\hat\{\\mathbf\{p\}\}\_\{s\}\\right\\rangle\.\(41\)

###### Proof\.

We first compute the Rao–Blackwellized categorical termℒ¯zs∣zt,wt\\bar\{\\mathcal\{L\}\}\_\{z\_\{s\}\\mid z\_\{t\},w\_\{t\}\}\. By definition,

ℒ¯zs∣zt,wt\\displaystyle\\bar\{\\mathcal\{L\}\}\_\{z\_\{s\}\\mid z\_\{t\},w\_\{t\}\}=𝔼q​\(𝐳~t∣𝐰t\)\[DKL\[q\(𝐳s∣𝐳~t,𝐱\)∥q\(𝐳s∣𝐳~t,𝐱^θ\)\]\]\.\\displaystyle=\\mathbb\{E\}\_\{q\(\\widetilde\{\\mathbf\{z\}\}\_\{t\}\\mid\\mathbf\{w\}\_\{t\}\)\}\\\!\\left\[D\_\{\\mathrm\{KL\}\}\\\!\\left\[q\(\\mathbf\{z\}\_\{s\}\\mid\\widetilde\{\\mathbf\{z\}\}\_\{t\},\\mathbf\{x\}\)\\,\\middle\\\|\\,q\(\\mathbf\{z\}\_\{s\}\\mid\\widetilde\{\\mathbf\{z\}\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\\right\]\\right\]\.\(42\)Using \([27](https://arxiv.org/html/2608.10615#A1.E27)\), we have

q​\(𝐳~t=𝐞j∣𝐰t\)=wt,j\.q\(\\widetilde\{\\mathbf\{z\}\}\_\{t\}=\\mathbf\{e\}\_\{j\}\\mid\\mathbf\{w\}\_\{t\}\)=w\_\{t,j\}\.For fixed𝐳~t=𝐞j\\widetilde\{\\mathbf\{z\}\}\_\{t\}=\\mathbf\{e\}\_\{j\}, both reverse conditionals are categorical:

q​\(𝐳s∣𝐳~t=𝐞j,𝐱\)=Cat⁡\(𝐳s;𝐫s∣t​\(𝐱,𝐞j\)\),q​\(𝐳s∣𝐳~t=𝐞j,𝐱^θ\)=Cat⁡\(𝐳s;𝐫s∣t​\(𝐱^θ,𝐞j\)\)\.q\(\\mathbf\{z\}\_\{s\}\\mid\\widetilde\{\\mathbf\{z\}\}\_\{t\}=\\mathbf\{e\}\_\{j\},\\mathbf\{x\}\)=\\operatorname\{Cat\}\\\!\\left\(\\mathbf\{z\}\_\{s\};\\mathbf\{r\}\_\{s\\mid t\}\(\\mathbf\{x\},\\mathbf\{e\}\_\{j\}\)\\right\),\\qquad q\(\\mathbf\{z\}\_\{s\}\\mid\\widetilde\{\\mathbf\{z\}\}\_\{t\}=\\mathbf\{e\}\_\{j\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)=\\operatorname\{Cat\}\\\!\\left\(\\mathbf\{z\}\_\{s\};\\mathbf\{r\}\_\{s\\mid t\}\(\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{e\}\_\{j\}\)\\right\)\.Expanding the expectation over𝐳~t\\widetilde\{\\mathbf\{z\}\}\_\{t\}and the KL divergence between categorical distributions gives

ℒ¯zs∣zt,wt\\displaystyle\\bar\{\\mathcal\{L\}\}\_\{z\_\{s\}\\mid z\_\{t\},w\_\{t\}\}=∑j=1Kwt,j​∑k=1Krs∣t,k​\(𝐱,𝐞j\)​log⁡rs∣t,k​\(𝐱,𝐞j\)rs∣t,k​\(𝐱^θ,𝐞j\)\.\\displaystyle=\\sum\_\{j=1\}^\{K\}w\_\{t,j\}\\sum\_\{k=1\}^\{K\}r\_\{s\\mid t,k\}\(\\mathbf\{x\},\\mathbf\{e\}\_\{j\}\)\\log\\frac\{r\_\{s\\mid t,k\}\(\\mathbf\{x\},\\mathbf\{e\}\_\{j\}\)\}\{r\_\{s\\mid t,k\}\(\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{e\}\_\{j\}\)\}\.\(43\)Substituting the explicit form of the reverse posterior components, we obtain

rs∣t,k​\(𝐱,𝐞j\)rs∣t,k​\(𝐱^θ,𝐞j\)\\displaystyle\\frac\{r\_\{s\\mid t,k\}\(\\mathbf\{x\},\\mathbf\{e\}\_\{j\}\)\}\{r\_\{s\\mid t,k\}\(\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{e\}\_\{j\}\)\}=\[αt∣s​δj,k\+\(1−αt∣s\)​πj\]​ps,k​\(𝐱\)/pt,j​\(𝐱\)\[αt∣s​δj,k\+\(1−αt∣s\)​πj\]​ps,k​\(𝐱^θ\)/pt,j​\(𝐱^θ\)\\displaystyle=\\frac\{\\bigl\[\\alpha\_\{t\\mid s\}\\delta\_\{j,k\}\+\(1\-\\alpha\_\{t\\mid s\}\)\\pi\_\{j\}\\bigr\]p\_\{s,k\}\(\\mathbf\{x\}\)/p\_\{t,j\}\(\\mathbf\{x\}\)\}\{\\bigl\[\\alpha\_\{t\\mid s\}\\delta\_\{j,k\}\+\(1\-\\alpha\_\{t\\mid s\}\)\\pi\_\{j\}\\bigr\]p\_\{s,k\}\(\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)/p\_\{t,j\}\(\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\}=ps,k​\(𝐱\)ps,k​\(𝐱^θ\)​pt,j​\(𝐱^θ\)pt,j​\(𝐱\)\.\\displaystyle=\\frac\{p\_\{s,k\}\(\\mathbf\{x\}\)\}\{p\_\{s,k\}\(\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\}\\frac\{p\_\{t,j\}\(\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\}\{p\_\{t,j\}\(\\mathbf\{x\}\)\}\.\(44\)Hence

log⁡rs∣t,k​\(𝐱,𝐞j\)rs∣t,k​\(𝐱^θ,𝐞j\)\\displaystyle\\log\\frac\{r\_\{s\\mid t,k\}\(\\mathbf\{x\},\\mathbf\{e\}\_\{j\}\)\}\{r\_\{s\\mid t,k\}\(\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{e\}\_\{j\}\)\}=log⁡pt,j​\(𝐱^θ\)pt,j​\(𝐱\)\+log⁡ps,k​\(𝐱\)ps,k​\(𝐱^θ\)\.\\displaystyle=\\log\\frac\{p\_\{t,j\}\(\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\}\{p\_\{t,j\}\(\\mathbf\{x\}\)\}\+\\log\\frac\{p\_\{s,k\}\(\\mathbf\{x\}\)\}\{p\_\{s,k\}\(\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\}\.\(45\)Substituting \([45](https://arxiv.org/html/2608.10615#A2.E45)\) into \([43](https://arxiv.org/html/2608.10615#A2.E43)\), the first contribution is

∑j=1Kwt,j​∑k=1Krs∣t,k​\(𝐱,𝐞j\)​log⁡pt,j​\(𝐱^θ\)pt,j​\(𝐱\)\\displaystyle\\sum\_\{j=1\}^\{K\}w\_\{t,j\}\\sum\_\{k=1\}^\{K\}r\_\{s\\mid t,k\}\(\\mathbf\{x\},\\mathbf\{e\}\_\{j\}\)\\log\\frac\{p\_\{t,j\}\(\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\}\{p\_\{t,j\}\(\\mathbf\{x\}\)\}=∑j=1Kwt,j​log⁡pt,j​\(𝐱^θ\)pt,j​\(𝐱\),\\displaystyle=\\sum\_\{j=1\}^\{K\}w\_\{t,j\}\\log\\frac\{p\_\{t,j\}\(\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\}\{p\_\{t,j\}\(\\mathbf\{x\}\)\},\(46\)becausers∣t​\(𝐱,𝐞j\)r\_\{s\\mid t\}\(\\mathbf\{x\},\\mathbf\{e\}\_\{j\}\)is a probability vector\. For the second contribution, exchanging the order of summation gives

∑j=1Kwt,j​∑k=1Krs∣t,k​\(𝐱,𝐞j\)​log⁡ps,k​\(𝐱\)ps,k​\(𝐱^θ\)\\displaystyle\\sum\_\{j=1\}^\{K\}w\_\{t,j\}\\sum\_\{k=1\}^\{K\}r\_\{s\\mid t,k\}\(\\mathbf\{x\},\\mathbf\{e\}\_\{j\}\)\\log\\frac\{p\_\{s,k\}\(\\mathbf\{x\}\)\}\{p\_\{s,k\}\(\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\}=∑k=1K\(∑j=1Kwt,j​rs∣t,k​\(𝐱,𝐞j\)\)​log⁡ps,k​\(𝐱\)ps,k​\(𝐱^θ\)\.\\displaystyle=\\sum\_\{k=1\}^\{K\}\\left\(\\sum\_\{j=1\}^\{K\}w\_\{t,j\}r\_\{s\\mid t,k\}\(\\mathbf\{x\},\\mathbf\{e\}\_\{j\}\)\\right\)\\log\\frac\{p\_\{s,k\}\(\\mathbf\{x\}\)\}\{p\_\{s,k\}\(\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\}\.\(47\)The coefficient in parentheses is exactly thekk\-th component ofq​\(𝐳s∣𝐰t,𝐱\)q\(\\mathbf\{z\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\), namely

∑j=1Kwt,j​rs∣t,k​\(𝐱,𝐞j\)=ρs∣t,k​\(𝐱,𝐰t\)\.\\sum\_\{j=1\}^\{K\}w\_\{t,j\}r\_\{s\\mid t,k\}\(\\mathbf\{x\},\\mathbf\{e\}\_\{j\}\)=\\rho\_\{s\\mid t,k\}\(\\mathbf\{x\},\\mathbf\{w\}\_\{t\}\)\.Therefore

ℒ¯zs∣zt,wt\\displaystyle\\bar\{\\mathcal\{L\}\}\_\{z\_\{s\}\\mid z\_\{t\},w\_\{t\}\}=∑j=1Kwt,j​log⁡pt,j​\(𝐱^θ\)pt,j​\(𝐱\)\+∑k=1Kρs∣t,k​\(𝐱,𝐰t\)​log⁡ps,k​\(𝐱\)ps,k​\(𝐱^θ\)\\displaystyle=\\sum\_\{j=1\}^\{K\}w\_\{t,j\}\\log\\frac\{p\_\{t,j\}\(\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\}\{p\_\{t,j\}\(\\mathbf\{x\}\)\}\+\\sum\_\{k=1\}^\{K\}\\rho\_\{s\\mid t,k\}\(\\mathbf\{x\},\\mathbf\{w\}\_\{t\}\)\\log\\frac\{p\_\{s,k\}\(\\mathbf\{x\}\)\}\{p\_\{s,k\}\(\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\}=⟨𝐰t,log⁡𝐩^t−log⁡𝐩t⟩\+⟨𝝆s∣t​\(𝐱,𝐰t\),log⁡𝐩s−log⁡𝐩^s⟩,\\displaystyle=\\left\\langle\\mathbf\{w\}\_\{t\},\\,\\log\\hat\{\\mathbf\{p\}\}\_\{t\}\-\\log\\mathbf\{p\}\_\{t\}\\right\\rangle\+\\left\\langle\\bm\{\\rho\}\_\{s\\mid t\}\(\\mathbf\{x\},\\mathbf\{w\}\_\{t\}\),\\,\\log\\mathbf\{p\}\_\{s\}\-\\log\\hat\{\\mathbf\{p\}\}\_\{s\}\\right\\rangle,\(48\)which proves \([41](https://arxiv.org/html/2608.10615#A2.E41)\)\. ∎

Equation \([41](https://arxiv.org/html/2608.10615#A2.E41)\) depends on the auxiliary state only through the current\-time average⟨𝐰t,log⁡𝐩^t⟩\\left\\langle\\mathbf\{w\}\_\{t\},\\,\\log\\hat\{\\mathbf\{p\}\}\_\{t\}\\right\\rangleand the lifted reverse posterior𝝆s∣t​\(𝐱,𝐰t\)\\bm\{\\rho\}\_\{s\\mid t\}\(\\mathbf\{x\},\\mathbf\{w\}\_\{t\}\)\. This is the expression used in the main objective and in the continuous\-time analysis\.

## Appendix CContinuous\-Time Limit of the Main Objective

The main text shows that the relaxed discrete bridge objectiveℒ¯zs∣zt,wt​\(𝐰t,𝐱^θ,𝐱;s,t\)\\bar\{\\mathcal\{L\}\}\_\{z\_\{s\}\\mid z\_\{t\},w\_\{t\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\};s,t\)admits a non\-degenerate first\-order continuous\-time limit\. This appendix gives the full proof of the first\-order local limit stated in the main text\.

### C\.1Proof of the main continuous\-time limit

###### Proposition 1\.

Up toθ\\theta\-independent additive terms, the relaxed discrete bridge objective satisfies

ℒ¯zs∣zt,wt​\(𝐰t,𝐱^θ,𝐱;t−Δ,t\)=Δ​ℓct​\(𝐰t,𝐱^θ,𝐱,t\)\+o​\(Δ\),\\bar\{\\mathcal\{L\}\}\_\{z\_\{s\}\\mid z\_\{t\},w\_\{t\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\};t\-\\Delta,t\)=\\Delta\\,\\ell\_\{\\mathrm\{ct\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\},t\)\+o\(\\Delta\),\(49\)where

ℓct\(𝐰t,𝐱^θ,𝐱,t\)=λ\(t\)\[\\displaystyle\\ell\_\{\\mathrm\{ct\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\},t\)=\\lambda\(t\)\\Bigg\[⟨𝐰t,𝝅⊘𝐩^t⟩−⟨𝐰t,𝝅⊘𝐩t⟩⟨𝐩t,log𝐩^t⟩\+⟨𝝅⊙\(𝐰t⊘𝐩t\),log𝐩^t⟩\]\.\\displaystyle\\left\\langle\\mathbf\{w\}\_\{t\},\\,\\bm\{\\pi\}\\oslash\\hat\{\\mathbf\{p\}\}\_\{t\}\\right\\rangle\-\\left\\langle\\mathbf\{w\}\_\{t\},\\,\\bm\{\\pi\}\\oslash\\mathbf\{p\}\_\{t\}\\right\\rangle\\,\\left\\langle\\mathbf\{p\}\_\{t\},\\,\\log\\hat\{\\mathbf\{p\}\}\_\{t\}\\right\\rangle\+\\left\\langle\\bm\{\\pi\}\\odot\(\\mathbf\{w\}\_\{t\}\\oslash\\mathbf\{p\}\_\{t\}\),\\,\\log\\hat\{\\mathbf\{p\}\}\_\{t\}\\right\\rangle\\Bigg\]\.\(50\)andλ​\(t\)≔−dd​t​log⁡α​\(t\)\\lambda\(t\)\\coloneq\-\\frac\{d\}\{dt\}\\log\\alpha\(t\)\. Consequently, the corresponding continuous\-time objective is

ℒct=∫01𝔼q​\(𝐱\)​\[𝔼q​\(𝐰t∣𝐱\)​\[ℓct​\(𝐰t,𝐱^θ,𝐱,t\)\]\]​𝑑t,\\mathcal\{L\}\_\{\\mathrm\{ct\}\}=\\int\_\{0\}^\{1\}\\mathbb\{E\}\_\{q\(\\mathbf\{x\}\)\}\\\!\\left\[\\mathbb\{E\}\_\{q\(\\mathbf\{w\}\_\{t\}\\mid\\mathbf\{x\}\)\}\\\!\\left\[\\ell\_\{\\mathrm\{ct\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\},t\)\\right\]\\right\]\\,dt,\(51\)withq​\(𝐰t∣𝐱\)=Dir⁡\(𝐰t;ηt​𝐩t\)q\(\\mathbf\{w\}\_\{t\}\\mid\\mathbf\{x\}\)=\\operatorname\{Dir\}\(\\mathbf\{w\}\_\{t\};\\eta\_\{t\}\\mathbf\{p\}\_\{t\}\)\.

###### Proof of[Proposition˜3](https://arxiv.org/html/2608.10615#Thmproposition3)\.

Up toθ\\theta\-independent additive terms,[Equation˜41](https://arxiv.org/html/2608.10615#A2.E41)can be written as

ℒ¯zs∣zt,wt​\(𝐰t,𝐱^θ,𝐱;s,t\)≡⟨𝐰t,log⁡𝐩^t⟩−⟨𝝆s∣t​\(𝐱,𝐰t\),log⁡𝐩^s⟩\.\\bar\{\\mathcal\{L\}\}\_\{z\_\{s\}\\mid z\_\{t\},w\_\{t\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\};s,t\)\\equiv\\left\\langle\\mathbf\{w\}\_\{t\},\\,\\log\\hat\{\\mathbf\{p\}\}\_\{t\}\\right\\rangle\-\\left\\langle\\bm\{\\rho\}\_\{s\\mid t\}\(\\mathbf\{x\},\\mathbf\{w\}\_\{t\}\),\\,\\log\\hat\{\\mathbf\{p\}\}\_\{s\}\\right\\rangle\.\(52\)We now sets=t−Δs=t\-\\Deltaand letΔ↓0\\Delta\\downarrow 0\.

First, by definition,

λ​\(t\)=−dd​t​log⁡α​\(t\),\\lambda\(t\)=\-\\frac\{d\}\{dt\}\\log\\alpha\(t\),\(53\)so that

αt∣t−Δ=α​\(t\)α​\(t−Δ\)=1−Δ​λ​\(t\)\+o​\(Δ\)\.\\alpha\_\{t\\mid t\-\\Delta\}=\\frac\{\\alpha\(t\)\}\{\\alpha\(t\-\\Delta\)\}=1\-\\Delta\\,\\lambda\(t\)\+o\(\\Delta\)\.\(54\)Moreover,

∂t𝐩t=∂t\(α​\(t\)​𝐱\+\(1−α​\(t\)\)​𝝅\)=−λ​\(t\)​\(𝐩t−𝝅\),\\partial\_\{t\}\\mathbf\{p\}\_\{t\}=\\partial\_\{t\}\\\!\\left\(\\alpha\(t\)\\mathbf\{x\}\+\(1\-\\alpha\(t\)\)\\bm\{\\pi\}\\right\)=\-\\lambda\(t\)\\bigl\(\\mathbf\{p\}\_\{t\}\-\\bm\{\\pi\}\\bigr\),\(55\)hence

𝐩t−Δ=𝐩t\+Δ​λ​\(t\)​\(𝐩t−𝝅\)\+o​\(Δ\)\.\\mathbf\{p\}\_\{t\-\\Delta\}=\\mathbf\{p\}\_\{t\}\+\\Delta\\,\\lambda\(t\)\(\\mathbf\{p\}\_\{t\}\-\\bm\{\\pi\}\)\+o\(\\Delta\)\.\(56\)
Next, we expand𝝆t−Δ∣t​\(𝐱,𝐰t\)\\bm\{\\rho\}\_\{t\-\\Delta\\mid t\}\(\\mathbf\{x\},\\mathbf\{w\}\_\{t\}\)using[Equation˜29](https://arxiv.org/html/2608.10615#A1.E29)\. Substituting[Equation˜54](https://arxiv.org/html/2608.10615#A3.E54)and[Equation˜56](https://arxiv.org/html/2608.10615#A3.E56)gives

𝝆t−Δ∣t​\(𝐱,𝐰t\)\\displaystyle\\bm\{\\rho\}\_\{t\-\\Delta\\mid t\}\(\\mathbf\{x\},\\mathbf\{w\}\_\{t\}\)=𝐩t−Δ⊙\[αt∣t−Δ​\(𝐰t⊘𝐩t\)\+\(1−αt∣t−Δ\)​⟨𝐰t,𝝅⊘𝐩t⟩​𝟏\]\\displaystyle=\\mathbf\{p\}\_\{t\-\\Delta\}\\odot\\left\[\\alpha\_\{t\\mid t\-\\Delta\}\\bigl\(\\mathbf\{w\}\_\{t\}\\oslash\\mathbf\{p\}\_\{t\}\\bigr\)\+\\bigl\(1\-\\alpha\_\{t\\mid t\-\\Delta\}\\bigr\)\\left\\langle\\mathbf\{w\}\_\{t\},\\,\\bm\{\\pi\}\\oslash\\mathbf\{p\}\_\{t\}\\right\\rangle\\mathbf\{1\}\\right\]=\(𝐩t\+Δ​λ​\(t\)​\(𝐩t−𝝅\)\)⊙\[𝐰t⊘𝐩t\+Δ​λ​\(t\)​\(⟨𝐰t,𝝅⊘𝐩t⟩​𝟏−𝐰t⊘𝐩t\)\]\+o​\(Δ\)\\displaystyle=\\Bigl\(\\mathbf\{p\}\_\{t\}\+\\Delta\\,\\lambda\(t\)\(\\mathbf\{p\}\_\{t\}\-\\bm\{\\pi\}\)\\Bigr\)\\odot\\Bigl\[\\mathbf\{w\}\_\{t\}\\oslash\\mathbf\{p\}\_\{t\}\+\\Delta\\,\\lambda\(t\)\\Bigl\(\\left\\langle\\mathbf\{w\}\_\{t\},\\,\\bm\{\\pi\}\\oslash\\mathbf\{p\}\_\{t\}\\right\\rangle\\mathbf\{1\}\-\\mathbf\{w\}\_\{t\}\\oslash\\mathbf\{p\}\_\{t\}\\Bigr\)\\Bigr\]\+o\(\\Delta\)=𝐰t\+Δ​λ​\(t\)​\[⟨𝐰t,𝝅⊘𝐩t⟩​𝐩t−𝝅⊙\(𝐰t⊘𝐩t\)\]\+o​\(Δ\)\.\\displaystyle=\\mathbf\{w\}\_\{t\}\+\\Delta\\,\\lambda\(t\)\\left\[\\left\\langle\\mathbf\{w\}\_\{t\},\\,\\bm\{\\pi\}\\oslash\\mathbf\{p\}\_\{t\}\\right\\rangle\\mathbf\{p\}\_\{t\}\-\\bm\{\\pi\}\\odot\(\\mathbf\{w\}\_\{t\}\\oslash\\mathbf\{p\}\_\{t\}\)\\right\]\+o\(\\Delta\)\.\(57\)
We now expandlog⁡𝐩^t−Δ\\log\\hat\{\\mathbf\{p\}\}\_\{t\-\\Delta\}\. Since𝐱^θ\\hat\{\\mathbf\{x\}\}\_\{\\theta\}is held fixed in the local limit,

𝐩^t=α​\(t\)​𝐱^θ\+\(1−α​\(t\)\)​𝝅,\\hat\{\\mathbf\{p\}\}\_\{t\}=\\alpha\(t\)\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\+\(1\-\\alpha\(t\)\)\\bm\{\\pi\},\(58\)and therefore

∂t𝐩^t=−λ​\(t\)​\(𝐩^t−𝝅\)\.\\partial\_\{t\}\\hat\{\\mathbf\{p\}\}\_\{t\}=\-\\lambda\(t\)\\bigl\(\\hat\{\\mathbf\{p\}\}\_\{t\}\-\\bm\{\\pi\}\\bigr\)\.\(59\)Dividing componentwise by𝐩^t\\hat\{\\mathbf\{p\}\}\_\{t\}yields

∂tlog⁡𝐩^t=−λ​\(t\)​\(𝟏−𝝅⊘𝐩^t\)\.\\partial\_\{t\}\\log\\hat\{\\mathbf\{p\}\}\_\{t\}=\-\\lambda\(t\)\\left\(\\mathbf\{1\}\-\\bm\{\\pi\}\\oslash\\hat\{\\mathbf\{p\}\}\_\{t\}\\right\)\.\(60\)Hence

log⁡𝐩^t−Δ=log⁡𝐩^t−Δ​∂tlog⁡𝐩^t\+o​\(Δ\)=log⁡𝐩^t\+Δ​λ​\(t\)​\(𝟏−𝝅⊘𝐩^t\)\+o​\(Δ\)\.\\log\\hat\{\\mathbf\{p\}\}\_\{t\-\\Delta\}=\\log\\hat\{\\mathbf\{p\}\}\_\{t\}\-\\Delta\\,\\partial\_\{t\}\\log\\hat\{\\mathbf\{p\}\}\_\{t\}\+o\(\\Delta\)=\\log\\hat\{\\mathbf\{p\}\}\_\{t\}\+\\Delta\\,\\lambda\(t\)\\left\(\\mathbf\{1\}\-\\bm\{\\pi\}\\oslash\\hat\{\\mathbf\{p\}\}\_\{t\}\\right\)\+o\(\\Delta\)\.\(61\)
For compactness, define

At\\displaystyle A\_\{t\}≔⟨𝐰t,𝝅⊘𝐩^t⟩,\\displaystyle\\coloneq\\left\\langle\\mathbf\{w\}\_\{t\},\\,\\bm\{\\pi\}\\oslash\\hat\{\\mathbf\{p\}\}\_\{t\}\\right\\rangle,\(62\)Bt\\displaystyle B\_\{t\}≔⟨𝐰t,𝝅⊘𝐩t⟩​⟨𝐩t,log⁡𝐩^t⟩,\\displaystyle\\coloneq\\left\\langle\\mathbf\{w\}\_\{t\},\\,\\bm\{\\pi\}\\oslash\\mathbf\{p\}\_\{t\}\\right\\rangle\\left\\langle\\mathbf\{p\}\_\{t\},\\,\\log\\hat\{\\mathbf\{p\}\}\_\{t\}\\right\\rangle,\(63\)Ct\\displaystyle C\_\{t\}≔⟨𝝅⊙\(𝐰t⊘𝐩t\),log⁡𝐩^t⟩\.\\displaystyle\\coloneq\\left\\langle\\bm\{\\pi\}\\odot\(\\mathbf\{w\}\_\{t\}\\oslash\\mathbf\{p\}\_\{t\}\),\\,\\log\\hat\{\\mathbf\{p\}\}\_\{t\}\\right\\rangle\.\(64\)Substituting[Equations˜57](https://arxiv.org/html/2608.10615#A3.E57)and[61](https://arxiv.org/html/2608.10615#A3.E61)into[Equation˜52](https://arxiv.org/html/2608.10615#A3.E52)and collecting first\-order terms gives

ℒ¯zt−Δ∣zt,wt​\(𝐰t,𝐱^θ,𝐱;t−Δ,t\)≡Δ​λ​\(t\)​\(At−Bt\+Ct−1\)\+o​\(Δ\)\.\\bar\{\\mathcal\{L\}\}\_\{z\_\{t\-\\Delta\}\\mid z\_\{t\},w\_\{t\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\};t\-\\Delta,t\)\\equiv\\Delta\\lambda\(t\)\(A\_\{t\}\-B\_\{t\}\+C\_\{t\}\-1\)\+o\(\\Delta\)\.\(65\)The term−Δ​λ​\(t\)\-\\Delta\\,\\lambda\(t\)is independent ofθ\\theta, so it can be discarded\. Therefore,

ℒ¯zt−Δ∣zt,wt​\(𝐰t,𝐱^θ,𝐱;t−Δ,t\)=Δ​ℓct​\(𝐰t,𝐱^θ,𝐱,t\)\+o​\(Δ\),\\bar\{\\mathcal\{L\}\}\_\{z\_\{t\-\\Delta\}\\mid z\_\{t\},w\_\{t\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\};t\-\\Delta,t\)=\\Delta\\,\\ell\_\{\\mathrm\{ct\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\},t\)\+o\(\\Delta\),\(66\)where, up toθ\\theta\-independent additive terms,

ℓct​\(𝐰t,𝐱^θ,𝐱,t\)\\displaystyle\\ell\_\{\\mathrm\{ct\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\},t\)≡λ​\(t\)​\[⟨𝐰t,𝝅⊘𝐩^t⟩−⟨𝐰t,𝝅⊘𝐩t⟩​⟨𝐩t,log⁡𝐩^t⟩\+⟨𝝅⊙\(𝐰t⊘𝐩t\),log⁡𝐩^t⟩\]\.\\displaystyle\\equiv\\lambda\(t\)\\Bigg\[\\left\\langle\\mathbf\{w\}\_\{t\},\\,\\bm\{\\pi\}\\oslash\\hat\{\\mathbf\{p\}\}\_\{t\}\\right\\rangle\-\\left\\langle\\mathbf\{w\}\_\{t\},\\,\\bm\{\\pi\}\\oslash\\mathbf\{p\}\_\{t\}\\right\\rangle\\left\\langle\\mathbf\{p\}\_\{t\},\\,\\log\\hat\{\\mathbf\{p\}\}\_\{t\}\\right\\rangle\+\\left\\langle\\bm\{\\pi\}\\odot\(\\mathbf\{w\}\_\{t\}\\oslash\\mathbf\{p\}\_\{t\}\),\\,\\log\\hat\{\\mathbf\{p\}\}\_\{t\}\\right\\rangle\\Bigg\]\.\(67\)This proves[Equation˜16](https://arxiv.org/html/2608.10615#S3.E16)and[Equation˜17](https://arxiv.org/html/2608.10615#S3.E17)\. Finally, integrating the density over time and taking expectation underq​\(𝐱\)q\(\\mathbf\{x\}\)andq​\(𝐰t∣𝐱\)q\(\\mathbf\{w\}\_\{t\}\\mid\\mathbf\{x\}\)yields[Equation˜18](https://arxiv.org/html/2608.10615#S3.E18)\. ∎

## Appendix DAlternative Surrogate Objectives

In the main text, we optimize the relaxed discrete bridge objectiveℒ¯zs∣zt,wt​\(𝐰t,𝐱^θ,𝐱;s,t\)\\bar\{\\mathcal\{L\}\}\_\{z\_\{s\}\\mid z\_\{t\},w\_\{t\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\};s,t\)\. The denoiser prediction𝐱^θ=fθ​\(𝐳t,t\)\\hat\{\\mathbf\{x\}\}\_\{\\theta\}=f\_\{\\theta\}\(\\mathbf\{z\}\_\{t\},t\)uses an independently sampled categorical network input\. In this appendix,𝐳~t\\widetilde\{\\mathbf\{z\}\}\_\{t\}denotes only the auxiliary categorical decode that is averaged inside the objective\. This choice is best understood relative to a broader family of surrogate objectives induced by the same simplex relaxation\. The natural starting point is the exact relaxed bridge

ℒws∣wt\(𝐰t,𝐱^θ,𝐱;s,t\)≔DKL\[q\(𝐰s∣𝐰t,𝐱\)∥q\(𝐰s∣𝐰t,𝐱^θ\)\]\.\\mathcal\{L\}\_\{w\_\{s\}\\mid w\_\{t\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\};s,t\)\\coloneq D\_\{\\mathrm\{KL\}\}\\\!\\left\[q\(\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)\\,\\middle\\\|\\,q\(\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\\right\]\.\(68\)This objective is the most direct one, but is generally intractable becauseq​\(𝐰s∣𝐰t,𝐱\)q\(\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)is a Dirichlet mixture\. The tractable surrogates introduced below differ in two orthogonal ways: first, whether they match only the discrete reverse state𝐳s\\mathbf\{z\}\_\{s\}or the joint state\(𝐳s,𝐰s\)\(\\mathbf\{z\}\_\{s\},\\mathbf\{w\}\_\{s\}\), and second, whether they condition directly on𝐰t\\mathbf\{w\}\_\{t\}or first decode𝐳~t∼q​\(𝐳~t∣𝐰t\)\\widetilde\{\\mathbf\{z\}\}\_\{t\}\\sim q\(\\widetilde\{\\mathbf\{z\}\}\_\{t\}\\mid\\mathbf\{w\}\_\{t\}\)and average\.

### D\.1A broader surrogate family

The simplest tractable surrogate matches only the lifted reverse posterior over𝐳s\\mathbf\{z\}\_\{s\}:

ℒzs∣wt\(𝐰t,𝐱^θ,𝐱;s,t\)≔DKL\[q\(𝐳s∣𝐰t,𝐱\)∥q\(𝐳s∣𝐰t,𝐱^θ\)\]\.\\mathcal\{L\}\_\{z\_\{s\}\\mid w\_\{t\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\};s,t\)\\coloneq D\_\{\\mathrm\{KL\}\}\\\!\\left\[q\(\\mathbf\{z\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)\\,\\middle\\\|\\,q\(\\mathbf\{z\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\\right\]\.\(69\)This objective preserves the full relaxed conditioning information in𝐰t\\mathbf\{w\}\_\{t\}, but discards the relaxed target𝐰s\\mathbf\{w\}\_\{s\}\.

The objective used in the main text instead decodes𝐳~t\\widetilde\{\\mathbf\{z\}\}\_\{t\}from𝐰t\\mathbf\{w\}\_\{t\}and averages the standard discrete reverse KL:

ℒ¯zs∣zt,wt\(𝐰t,𝐱^θ,𝐱;s,t\)≔𝔼q​\(𝐳~t∣𝐰t\)\[DKL\[q\(𝐳s∣𝐳~t,𝐱\)∥q\(𝐳s∣𝐳~t,𝐱^θ\)\]\]\.\\bar\{\\mathcal\{L\}\}\_\{z\_\{s\}\\mid z\_\{t\},w\_\{t\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\};s,t\)\\coloneq\\mathbb\{E\}\_\{q\(\\widetilde\{\\mathbf\{z\}\}\_\{t\}\\mid\\mathbf\{w\}\_\{t\}\)\}\\\!\\left\[D\_\{\\mathrm\{KL\}\}\\\!\\left\[q\(\\mathbf\{z\}\_\{s\}\\mid\\widetilde\{\\mathbf\{z\}\}\_\{t\},\\mathbf\{x\}\)\\,\\middle\\\|\\,q\(\\mathbf\{z\}\_\{s\}\\mid\\widetilde\{\\mathbf\{z\}\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\\right\]\\right\]\.\(70\)Compared with \([69](https://arxiv.org/html/2608.10615#A4.E69)\), this objective is looser because𝐰t\\mathbf\{w\}\_\{t\}influences the reverse matching step only through the decoded categorical latent𝐳~t\\widetilde\{\\mathbf\{z\}\}\_\{t\}\. Its advantage is that it remains close to the standard discrete\-diffusion objective and admits the non\-degenerate continuous\-time limit derived in the main text\.

A richer alternative is to match the joint bridge over\(𝐳s,𝐰s\)\(\\mathbf\{z\}\_\{s\},\\mathbf\{w\}\_\{s\}\)directly under the relaxed conditioning state:

ℒzs,ws∣wt\(𝐰t,𝐱^θ,𝐱;s,t\)≔DKL\[q\(𝐳s,𝐰s∣𝐰t,𝐱\)∥q\(𝐳s,𝐰s∣𝐰t,𝐱^θ\)\]\.\\mathcal\{L\}\_\{z\_\{s\},w\_\{s\}\\mid w\_\{t\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\};s,t\)\\coloneq D\_\{\\mathrm\{KL\}\}\\\!\\left\[q\(\\mathbf\{z\}\_\{s\},\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)\\,\\middle\\\|\\,q\(\\mathbf\{z\}\_\{s\},\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\\right\]\.\(71\)This objective is richer than the discrete surrogates because it also matches the relaxed bridge at timess\.

Finally, one may combine the decoded conditioning of \([70](https://arxiv.org/html/2608.10615#A4.E70)\) with the joint target of \([71](https://arxiv.org/html/2608.10615#A4.E71)\):

ℒ¯zs,ws∣zt,wt\(𝐰t,𝐱^θ,𝐱;s,t\)≔𝔼q​\(𝐳~t∣𝐰t\)\[DKL\[q\(𝐳s,𝐰s∣𝐳~t,𝐰t,𝐱\)∥q\(𝐳s,𝐰s∣𝐳~t,𝐰t,𝐱^θ\)\]\]\.\\bar\{\\mathcal\{L\}\}\_\{z\_\{s\},w\_\{s\}\\mid z\_\{t\},w\_\{t\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\};s,t\)\\coloneq\\mathbb\{E\}\_\{q\(\\widetilde\{\\mathbf\{z\}\}\_\{t\}\\mid\\mathbf\{w\}\_\{t\}\)\}\\\!\\left\[D\_\{\\mathrm\{KL\}\}\\\!\\left\[q\(\\mathbf\{z\}\_\{s\},\\mathbf\{w\}\_\{s\}\\mid\\widetilde\{\\mathbf\{z\}\}\_\{t\},\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)\\,\\middle\\\|\\,q\(\\mathbf\{z\}\_\{s\},\\mathbf\{w\}\_\{s\}\\mid\\widetilde\{\\mathbf\{z\}\}\_\{t\},\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\\right\]\\right\]\.\(72\)This is the richest tractable surrogate in the family: it keeps the relaxed target𝐰s\\mathbf\{w\}\_\{s\}while also conditioning through the decoded latent𝐳~t\\widetilde\{\\mathbf\{z\}\}\_\{t\}\.

The loose joint objective admits a useful chain\-rule decomposition\. It shows that the main\-text objectiveℒ¯zs∣zt,wt\\bar\{\\mathcal\{L\}\}\_\{z\_\{s\}\\mid z\_\{t\},w\_\{t\}\}is precisely the categorical part of the loose joint bridge\.

###### Proposition 6\.

The tight and loose joint objectives admit the decompositions

ℒzs,ws∣wt​\(𝐰t,𝐱^θ,𝐱;s,t\)\\displaystyle\\mathcal\{L\}\_\{z\_\{s\},w\_\{s\}\\mid w\_\{t\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\};s,t\)=ℒzs∣wt​\(𝐰t,𝐱^θ,𝐱;s,t\)\+ℒ¯ws∣zs,wt​\(𝐰t,𝐱^θ,𝐱;s,t\),\\displaystyle=\\mathcal\{L\}\_\{z\_\{s\}\\mid w\_\{t\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\};s,t\)\+\\bar\{\\mathcal\{L\}\}\_\{w\_\{s\}\\mid z\_\{s\},w\_\{t\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\};s,t\),\(73\)ℒ¯zs,ws∣zt,wt​\(𝐰t,𝐱^θ,𝐱;s,t\)\\displaystyle\\bar\{\\mathcal\{L\}\}\_\{z\_\{s\},w\_\{s\}\\mid z\_\{t\},w\_\{t\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\};s,t\)=ℒ¯zs∣zt,wt​\(𝐰t,𝐱^θ,𝐱;s,t\)\+ℒ¯ws∣zs,wt​\(𝐰t,𝐱^θ,𝐱;s,t\),\\displaystyle=\\bar\{\\mathcal\{L\}\}\_\{z\_\{s\}\\mid z\_\{t\},w\_\{t\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\};s,t\)\+\\bar\{\\mathcal\{L\}\}\_\{w\_\{s\}\\mid z\_\{s\},w\_\{t\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\};s,t\),\(74\)where

ℒ¯ws∣zs,wt\(𝐰t,𝐱^θ,𝐱;s,t\)≔𝔼q​\(𝐳s∣𝐰t,𝐱\)\[DKL\[q\(𝐰s∣𝐳s,𝐱\)∥q\(𝐰s∣𝐳s,𝐱^θ\)\]\]\.\\bar\{\\mathcal\{L\}\}\_\{w\_\{s\}\\mid z\_\{s\},w\_\{t\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\};s,t\)\\coloneq\\mathbb\{E\}\_\{q\(\\mathbf\{z\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)\}\\\!\\left\[D\_\{\\mathrm\{KL\}\}\\\!\\left\[q\(\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{z\}\_\{s\},\\mathbf\{x\}\)\\,\\middle\\\|\\,q\(\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{z\}\_\{s\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\\right\]\\right\]\.\(75\)

###### Proof\.

We first derive the tight decomposition \([73](https://arxiv.org/html/2608.10615#A4.E73)\)\.

By the joint graphical model, we have the conditional independences

𝐰s⟂⟂𝐰t∣\(𝐳s,𝐱\),q\(𝐳s,𝐰s∣𝐰t,𝐱\)=q\(𝐳s∣𝐰t,𝐱\)q\(𝐰s∣𝐳s,𝐱\),\\mathbf\{w\}\_\{s\}\\perp\\\!\\\!\\\!\\perp\\mathbf\{w\}\_\{t\}\\mid\(\\mathbf\{z\}\_\{s\},\\mathbf\{x\}\),\\qquad q\(\\mathbf\{z\}\_\{s\},\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)=q\(\\mathbf\{z\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)\\,q\(\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{z\}\_\{s\},\\mathbf\{x\}\),and similarly on the model side,

q​\(𝐳s,𝐰s∣𝐰t,𝐱^θ\)=q​\(𝐳s∣𝐰t,𝐱^θ\)​q​\(𝐰s∣𝐳s,𝐱^θ\)\.q\(\\mathbf\{z\}\_\{s\},\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)=q\(\\mathbf\{z\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\\,q\(\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{z\}\_\{s\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\.Substituting these factorizations intoℒzs,ws∣wt\\mathcal\{L\}\_\{z\_\{s\},w\_\{s\}\\mid w\_\{t\}\}gives

ℒzs,ws∣wt\\displaystyle\\mathcal\{L\}\_\{z\_\{s\},w\_\{s\}\\mid w\_\{t\}\}=DKL\[q\(𝐳s∣𝐰t,𝐱\)q\(𝐰s∣𝐳s,𝐱\)∥q\(𝐳s∣𝐰t,𝐱^θ\)q\(𝐰s∣𝐳s,𝐱^θ\)\]\.\\displaystyle=D\_\{\\mathrm\{KL\}\}\\\!\\left\[q\(\\mathbf\{z\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)\\,q\(\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{z\}\_\{s\},\\mathbf\{x\}\)\\,\\middle\\\|\\,q\(\\mathbf\{z\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\\,q\(\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{z\}\_\{s\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\\right\]\.\(76\)Applying the chain rule of KL yields

ℒzs,ws∣wt\\displaystyle\\mathcal\{L\}\_\{z\_\{s\},w\_\{s\}\\mid w\_\{t\}\}=DKL\[q\(𝐳s∣𝐰t,𝐱\)∥q\(𝐳s∣𝐰t,𝐱^θ\)\]\\displaystyle=D\_\{\\mathrm\{KL\}\}\\\!\\left\[q\(\\mathbf\{z\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)\\,\\middle\\\|\\,q\(\\mathbf\{z\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\\right\]\+𝔼q​\(𝐳s∣𝐰t,𝐱\)\[DKL\[q\(𝐰s∣𝐳s,𝐱\)∥q\(𝐰s∣𝐳s,𝐱^θ\)\]\]\.\\displaystyle\\quad\+\\mathbb\{E\}\_\{q\(\\mathbf\{z\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)\}\\\!\\left\[D\_\{\\mathrm\{KL\}\}\\\!\\left\[q\(\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{z\}\_\{s\},\\mathbf\{x\}\)\\,\\middle\\\|\\,q\(\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{z\}\_\{s\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\\right\]\\right\]\.\(77\)The first term is exactlyℒzs∣wt\\mathcal\{L\}\_\{z\_\{s\}\\mid w\_\{t\}\}, while the second term isℒ¯ws∣zs,wt\\bar\{\\mathcal\{L\}\}\_\{w\_\{s\}\\mid z\_\{s\},w\_\{t\}\}by definition\. This proves \([73](https://arxiv.org/html/2608.10615#A4.E73)\)\.

We next derive the loose decomposition \([74](https://arxiv.org/html/2608.10615#A4.E74)\)\.

The same graphical model implies the conditional independences

𝐳s⟂⟂𝐰t∣\(𝐳~t,𝐱\),𝐰s⟂⟂\(𝐳~t,𝐰t\)∣\(𝐳s,𝐱\),\\mathbf\{z\}\_\{s\}\\perp\\\!\\\!\\\!\\perp\\mathbf\{w\}\_\{t\}\\mid\(\\widetilde\{\\mathbf\{z\}\}\_\{t\},\\mathbf\{x\}\),\\qquad\\mathbf\{w\}\_\{s\}\\perp\\\!\\\!\\\!\\perp\(\\widetilde\{\\mathbf\{z\}\}\_\{t\},\\mathbf\{w\}\_\{t\}\)\\mid\(\\mathbf\{z\}\_\{s\},\\mathbf\{x\}\),and therefore

q​\(𝐳s,𝐰s∣𝐳~t,𝐰t,𝐱\)\\displaystyle q\(\\mathbf\{z\}\_\{s\},\\mathbf\{w\}\_\{s\}\\mid\\widetilde\{\\mathbf\{z\}\}\_\{t\},\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)=q​\(𝐳s∣𝐳~t,𝐱\)​q​\(𝐰s∣𝐳s,𝐱\),\\displaystyle=q\(\\mathbf\{z\}\_\{s\}\\mid\\widetilde\{\\mathbf\{z\}\}\_\{t\},\\mathbf\{x\}\)\\,q\(\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{z\}\_\{s\},\\mathbf\{x\}\),\(78\)q​\(𝐳s,𝐰s∣𝐳~t,𝐰t,𝐱^θ\)\\displaystyle q\(\\mathbf\{z\}\_\{s\},\\mathbf\{w\}\_\{s\}\\mid\\widetilde\{\\mathbf\{z\}\}\_\{t\},\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)=q​\(𝐳s∣𝐳~t,𝐱^θ\)​q​\(𝐰s∣𝐳s,𝐱^θ\)\.\\displaystyle=q\(\\mathbf\{z\}\_\{s\}\\mid\\widetilde\{\\mathbf\{z\}\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\\,q\(\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{z\}\_\{s\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\.\(79\)Substituting \([78](https://arxiv.org/html/2608.10615#A4.E78)\) and \([79](https://arxiv.org/html/2608.10615#A4.E79)\) intoℒ¯zs,ws∣zt,wt\\bar\{\\mathcal\{L\}\}\_\{z\_\{s\},w\_\{s\}\\mid z\_\{t\},w\_\{t\}\}gives

ℒ¯zs,ws∣zt,wt\\displaystyle\\bar\{\\mathcal\{L\}\}\_\{z\_\{s\},w\_\{s\}\\mid z\_\{t\},w\_\{t\}\}=𝔼q​\(𝐳~t∣𝐰t\)\[DKL\[q\(𝐳s∣𝐳~t,𝐱\)q\(𝐰s∣𝐳s,𝐱\)∥q\(𝐳s∣𝐳~t,𝐱^θ\)q\(𝐰s∣𝐳s,𝐱^θ\)\]\]\.\\displaystyle=\\mathbb\{E\}\_\{q\(\\widetilde\{\\mathbf\{z\}\}\_\{t\}\\mid\\mathbf\{w\}\_\{t\}\)\}\\\!\\left\[D\_\{\\mathrm\{KL\}\}\\\!\\left\[q\(\\mathbf\{z\}\_\{s\}\\mid\\widetilde\{\\mathbf\{z\}\}\_\{t\},\\mathbf\{x\}\)\\,q\(\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{z\}\_\{s\},\\mathbf\{x\}\)\\,\\middle\\\|\\,q\(\\mathbf\{z\}\_\{s\}\\mid\\widetilde\{\\mathbf\{z\}\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\\,q\(\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{z\}\_\{s\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\\right\]\\right\]\.\(80\)Applying the chain rule of KL inside the expectation yields

ℒ¯zs,ws∣zt,wt\\displaystyle\\bar\{\\mathcal\{L\}\}\_\{z\_\{s\},w\_\{s\}\\mid z\_\{t\},w\_\{t\}\}=𝔼q​\(𝐳~t∣𝐰t\)\[DKL\[q\(𝐳s∣𝐳~t,𝐱\)∥q\(𝐳s∣𝐳~t,𝐱^θ\)\]\]\\displaystyle=\\mathbb\{E\}\_\{q\(\\widetilde\{\\mathbf\{z\}\}\_\{t\}\\mid\\mathbf\{w\}\_\{t\}\)\}\\\!\\left\[D\_\{\\mathrm\{KL\}\}\\\!\\left\[q\(\\mathbf\{z\}\_\{s\}\\mid\\widetilde\{\\mathbf\{z\}\}\_\{t\},\\mathbf\{x\}\)\\,\\middle\\\|\\,q\(\\mathbf\{z\}\_\{s\}\\mid\\widetilde\{\\mathbf\{z\}\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\\right\]\\right\]\+𝔼q​\(𝐳~t∣𝐰t\)\[𝔼q​\(𝐳s∣𝐳~t,𝐱\)\[DKL\[q\(𝐰s∣𝐳s,𝐱\)∥q\(𝐰s∣𝐳s,𝐱^θ\)\]\]\]\.\\displaystyle\\quad\+\\mathbb\{E\}\_\{q\(\\widetilde\{\\mathbf\{z\}\}\_\{t\}\\mid\\mathbf\{w\}\_\{t\}\)\}\\\!\\left\[\\mathbb\{E\}\_\{q\(\\mathbf\{z\}\_\{s\}\\mid\\widetilde\{\\mathbf\{z\}\}\_\{t\},\\mathbf\{x\}\)\}\\\!\\left\[D\_\{\\mathrm\{KL\}\}\\\!\\left\[q\(\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{z\}\_\{s\},\\mathbf\{x\}\)\\,\\middle\\\|\\,q\(\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{z\}\_\{s\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\\right\]\\right\]\\right\]\.\(81\)The first term is exactlyℒ¯zs∣zt,wt\\bar\{\\mathcal\{L\}\}\_\{z\_\{s\}\\mid z\_\{t\},w\_\{t\}\}\. For the second term, apply the law of total expectation:

𝔼q​\(𝐳~t∣𝐰t\)​\[𝔼q​\(𝐳s∣𝐳~t,𝐱\)​\[\[⋅\]\]\]=𝔼q​\(𝐳s∣𝐰t,𝐱\)​\[\[⋅\]\]\.\\mathbb\{E\}\_\{q\(\\widetilde\{\\mathbf\{z\}\}\_\{t\}\\mid\\mathbf\{w\}\_\{t\}\)\}\\\!\\left\[\\mathbb\{E\}\_\{q\(\\mathbf\{z\}\_\{s\}\\mid\\widetilde\{\\mathbf\{z\}\}\_\{t\},\\mathbf\{x\}\)\}\\\!\\left\[\[\\cdot\]\\right\]\\right\]=\\mathbb\{E\}\_\{q\(\\mathbf\{z\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)\}\\\!\\left\[\[\\cdot\]\\right\]\.Hence

ℒ¯zs,ws∣zt,wt\\displaystyle\\bar\{\\mathcal\{L\}\}\_\{z\_\{s\},w\_\{s\}\\mid z\_\{t\},w\_\{t\}\}=ℒ¯zs∣zt,wt\\displaystyle=\\bar\{\\mathcal\{L\}\}\_\{z\_\{s\}\\mid z\_\{t\},w\_\{t\}\}\+𝔼q​\(𝐳s∣𝐰t,𝐱\)\[DKL\[q\(𝐰s∣𝐳s,𝐱\)∥q\(𝐰s∣𝐳s,𝐱^θ\)\]\],\\displaystyle\\quad\+\\mathbb\{E\}\_\{q\(\\mathbf\{z\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)\}\\\!\\left\[D\_\{\\mathrm\{KL\}\}\\\!\\left\[q\(\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{z\}\_\{s\},\\mathbf\{x\}\)\\,\\middle\\\|\\,q\(\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{z\}\_\{s\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\\right\]\\right\],\(82\)which is exactly \([74](https://arxiv.org/html/2608.10615#A4.E74)\)\. ∎

The decomposition in \([74](https://arxiv.org/html/2608.10615#A4.E74)\) clarifies the role of the selected main\-text objective\. The relaxed discrete bridge keeps the categorical part of the joint bridge while discarding the additional simplex\-matching term\. This is exactly the simplification that later makes the continuous\-time limit non\-degenerate\.

### D\.2Auxiliary KL inequalities

We next record three standard KL inequalities that will be used to compare the surrogate objectives\.

###### Lemma 2\(Data processing inequality for KL divergence\)\.

LetPPandQQbe two probability distributions on a space𝒳\\mathcal\{X\}, and let𝒦​\(𝐲∣𝐱\)\\mathcal\{K\}\(\\mathbf\{y\}\\mid\\mathbf\{x\}\)be a stochastic kernel from𝒳\\mathcal\{X\}to𝒴\\mathcal\{Y\}\. Define the pushforward distributions

\(P​𝒦\)​\(𝐲\)≔𝔼P​\(𝐱\)​\[𝒦​\(𝐲∣𝐱\)\],\(Q​𝒦\)​\(𝐲\)≔𝔼Q​\(𝐱\)​\[𝒦​\(𝐲∣𝐱\)\]\.\(P\\mathcal\{K\}\)\(\\mathbf\{y\}\)\\coloneq\\mathbb\{E\}\_\{P\(\\mathbf\{x\}\)\}\\\!\\left\[\\mathcal\{K\}\(\\mathbf\{y\}\\mid\\mathbf\{x\}\)\\right\],\\qquad\(Q\\mathcal\{K\}\)\(\\mathbf\{y\}\)\\coloneq\\mathbb\{E\}\_\{Q\(\\mathbf\{x\}\)\}\\\!\\left\[\\mathcal\{K\}\(\\mathbf\{y\}\\mid\\mathbf\{x\}\)\\right\]\.\(83\)Then

DKL\[P𝒦∥Q𝒦\]≤DKL\[P∥Q\]\.D\_\{\\mathrm\{KL\}\}\\\!\\left\[P\\mathcal\{K\}\\,\\middle\\\|\\,Q\\mathcal\{K\}\\right\]\\leq D\_\{\\mathrm\{KL\}\}\\\!\\left\[P\\,\\middle\\\|\\,Q\\right\]\.\(84\)

###### Proof\.

This follows directly from Theorem 4\.1 ofKullback and Leibler \[[1951](https://arxiv.org/html/2608.10615#bib.bib18)\]by takingTTto be the stochastic kernel𝒦\\mathcal\{K\}\. ∎

A particularly important special case is marginalization\.

###### Corollary 2\(Marginalization cannot increase KL\)\.

LetPU,VP\_\{U,V\}andQU,VQ\_\{U,V\}be two joint distributions, with marginalsPUP\_\{U\}andQUQ\_\{U\}\. Then

DKL\[PU∥QU\]≤DKL\[PU,V∥QU,V\]\.D\_\{\\mathrm\{KL\}\}\\\!\\left\[P\_\{U\}\\,\\middle\\\|\\,Q\_\{U\}\\right\]\\leq D\_\{\\mathrm\{KL\}\}\\\!\\left\[P\_\{U,V\}\\,\\middle\\\|\\,Q\_\{U,V\}\\right\]\.\(85\)

###### Proof\.

Take𝒦\\mathcal\{K\}to be the projection kernel\(u,v\)↦u\(u,v\)\\mapsto uin[Lemma˜2](https://arxiv.org/html/2608.10615#Thmlemma2)\. ∎

###### Lemma 3\(Joint convexity of KL divergence\)\.

Letλi≥0\\lambda\_\{i\}\\geq 0with∑i=1mλi=1\\sum\_\{i=1\}^\{m\}\\lambda\_\{i\}=1, and letPi,QiP\_\{i\},Q\_\{i\}be probability distributions on a common space\. Then

DKL\[∑i=1mλiPi∥∑i=1mλiQi\]≤∑i=1mλiDKL\[Pi∥Qi\]\.D\_\{\\mathrm\{KL\}\}\\\!\\left\[\\sum\_\{i=1\}^\{m\}\\lambda\_\{i\}P\_\{i\}\\,\\middle\\\|\\,\\sum\_\{i=1\}^\{m\}\\lambda\_\{i\}Q\_\{i\}\\right\]\\leq\\sum\_\{i=1\}^\{m\}\\lambda\_\{i\}D\_\{\\mathrm\{KL\}\}\\\!\\left\[P\_\{i\}\\,\\middle\\\|\\,Q\_\{i\}\\right\]\.\(86\)

###### Proof\.

This follows from the joint convexity of relative entropy; see Theorem 2\.7\.2 inCover and Thomas \[[2006](https://arxiv.org/html/2608.10615#bib.bib19)\]\. ∎

### D\.3Relations among the surrogate objectives

The valid relations among the surrogate objectives follow from marginalization and joint convexity of KL divergence\. Marginalizing either𝐳s\\mathbf\{z\}\_\{s\}or𝐰s\\mathbf\{w\}\_\{s\}from a joint bridge gives the corresponding categorical or simplex bound\. In addition, averaging over the auxiliary decoded state𝐳~t∼q​\(𝐳~t∣𝐰t\)\\widetilde\{\\mathbf\{z\}\}\_\{t\}\\sim q\(\\widetilde\{\\mathbf\{z\}\}\_\{t\}\\mid\\mathbf\{w\}\_\{t\}\)gives the bounds from the direct categorical and direct joint objectives to their decoded counterparts\.

To state the decoded simplex bound, define

ℒ¯ws∣zt,wt\(𝐰t,𝐱^θ,𝐱;s,t\)≔𝔼q​\(𝐳~t∣𝐰t\)\[DKL\[q\(𝐰s∣𝐳~t,𝐰t,𝐱\)∥q\(𝐰s∣𝐳~t,𝐰t,𝐱^θ\)\]\]\.\\bar\{\\mathcal\{L\}\}\_\{w\_\{s\}\\mid z\_\{t\},w\_\{t\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\};s,t\)\\coloneq\\mathbb\{E\}\_\{q\(\\widetilde\{\\mathbf\{z\}\}\_\{t\}\\mid\\mathbf\{w\}\_\{t\}\)\}\\\!\\left\[D\_\{\\mathrm\{KL\}\}\\\!\\left\[q\(\\mathbf\{w\}\_\{s\}\\mid\\widetilde\{\\mathbf\{z\}\}\_\{t\},\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)\\,\\middle\\\|\\,q\(\\mathbf\{w\}\_\{s\}\\mid\\widetilde\{\\mathbf\{z\}\}\_\{t\},\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\\right\]\\right\]\.\(87\)
###### Proposition 7\.

For all𝐰t∈ΔK−1\\mathbf\{w\}\_\{t\}\\in\\Delta^\{K\-1\},𝐱^θ∈ΔK−1\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\\in\\Delta^\{K\-1\}, and𝐱∈𝒱\\mathbf\{x\}\\in\\mathcal\{V\},

ℒzs∣wt\\displaystyle\\mathcal\{L\}\_\{z\_\{s\}\\mid w\_\{t\}\}≤ℒzs,ws∣wt,\\displaystyle\\leq\\mathcal\{L\}\_\{z\_\{s\},w\_\{s\}\\mid w\_\{t\}\},\(88\)ℒws∣wt\\displaystyle\\mathcal\{L\}\_\{w\_\{s\}\\mid w\_\{t\}\}≤ℒzs,ws∣wt,\\displaystyle\\leq\\mathcal\{L\}\_\{z\_\{s\},w\_\{s\}\\mid w\_\{t\}\},ℒzs,ws∣wt\\displaystyle\\mathcal\{L\}\_\{z\_\{s\},w\_\{s\}\\mid w\_\{t\}\}≤ℒ¯zs,ws∣zt,wt\.\\displaystyle\\leq\\bar\{\\mathcal\{L\}\}\_\{z\_\{s\},w\_\{s\}\\mid z\_\{t\},w\_\{t\}\}\.Moreover,

ℒzs∣wt\\displaystyle\\mathcal\{L\}\_\{z\_\{s\}\\mid w\_\{t\}\}≤ℒ¯zs∣zt,wt≤ℒ¯zs,ws∣zt,wt,\\displaystyle\\leq\\bar\{\\mathcal\{L\}\}\_\{z\_\{s\}\\mid z\_\{t\},w\_\{t\}\}\\leq\\bar\{\\mathcal\{L\}\}\_\{z\_\{s\},w\_\{s\}\\mid z\_\{t\},w\_\{t\}\},\(89\)ℒ¯ws∣zt,wt\\displaystyle\\bar\{\\mathcal\{L\}\}\_\{w\_\{s\}\\mid z\_\{t\},w\_\{t\}\}≤ℒ¯zs,ws∣zt,wt\.\\displaystyle\\leq\\bar\{\\mathcal\{L\}\}\_\{z\_\{s\},w\_\{s\}\\mid z\_\{t\},w\_\{t\}\}\.

###### Proof\.

We first prove \([88](https://arxiv.org/html/2608.10615#A4.E88)\)\.

1. 1\.The distributionsq​\(𝐳s∣𝐰t,𝐱\)q\(\\mathbf\{z\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)andq​\(𝐰s∣𝐰t,𝐱\)q\(\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)are the corresponding marginals ofq​\(𝐳s,𝐰s∣𝐰t,𝐱\)q\(\\mathbf\{z\}\_\{s\},\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)\. The same statement holds on the model side\. Hence, by[Corollary˜2](https://arxiv.org/html/2608.10615#Thmcorollary2), ℒzs∣wt​\(𝐰t,𝐱^θ,𝐱\)\\displaystyle\\mathcal\{L\}\_\{z\_\{s\}\\mid w\_\{t\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\}\)≤DKL\[q\(𝐳s,𝐰s∣𝐰t,𝐱\)∥q\(𝐳s,𝐰s∣𝐰t,𝐱^θ\)\]\\displaystyle\\leq D\_\{\\mathrm\{KL\}\}\\\!\\left\[q\(\\mathbf\{z\}\_\{s\},\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)\\,\\middle\\\|\\,q\(\\mathbf\{z\}\_\{s\},\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\\right\]=ℒzs,ws∣wt​\(𝐰t,𝐱^θ,𝐱\),\\displaystyle=\\mathcal\{L\}\_\{z\_\{s\},w\_\{s\}\\mid w\_\{t\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\}\),\(90\)ℒws∣wt​\(𝐰t,𝐱^θ,𝐱\)\\displaystyle\\mathcal\{L\}\_\{w\_\{s\}\\mid w\_\{t\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\}\)≤DKL\[q\(𝐳s,𝐰s∣𝐰t,𝐱\)∥q\(𝐳s,𝐰s∣𝐰t,𝐱^θ\)\]\\displaystyle\\leq D\_\{\\mathrm\{KL\}\}\\\!\\left\[q\(\\mathbf\{z\}\_\{s\},\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)\\,\\middle\\\|\\,q\(\\mathbf\{z\}\_\{s\},\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\\right\]=ℒzs,ws∣wt​\(𝐰t,𝐱^θ,𝐱\)\.\\displaystyle=\\mathcal\{L\}\_\{z\_\{s\},w\_\{s\}\\mid w\_\{t\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\}\)\.\(91\)
2. 2\.By marginalizing over𝐳~t\\widetilde\{\\mathbf\{z\}\}\_\{t\}, we have q​\(𝐳s,𝐰s∣𝐰t,𝐱\)\\displaystyle q\(\\mathbf\{z\}\_\{s\},\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)=𝔼q​\(𝐳~t∣𝐰t\)​\[q​\(𝐳s,𝐰s∣𝐳~t,𝐰t,𝐱\)\],\\displaystyle=\\mathbb\{E\}\_\{q\(\\widetilde\{\\mathbf\{z\}\}\_\{t\}\\mid\\mathbf\{w\}\_\{t\}\)\}\\\!\\left\[q\(\\mathbf\{z\}\_\{s\},\\mathbf\{w\}\_\{s\}\\mid\\widetilde\{\\mathbf\{z\}\}\_\{t\},\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)\\right\],\(92\)q​\(𝐳s,𝐰s∣𝐰t,𝐱^θ\)\\displaystyle q\(\\mathbf\{z\}\_\{s\},\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)=𝔼q​\(𝐳~t∣𝐰t\)​\[q​\(𝐳s,𝐰s∣𝐳~t,𝐰t,𝐱^θ\)\]\.\\displaystyle=\\mathbb\{E\}\_\{q\(\\widetilde\{\\mathbf\{z\}\}\_\{t\}\\mid\\mathbf\{w\}\_\{t\}\)\}\\\!\\left\[q\(\\mathbf\{z\}\_\{s\},\\mathbf\{w\}\_\{s\}\\mid\\widetilde\{\\mathbf\{z\}\}\_\{t\},\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\\right\]\.\(93\)The mixing law is shared by both sides because it is alwaysq​\(𝐳~t∣𝐰t\)q\(\\widetilde\{\\mathbf\{z\}\}\_\{t\}\\mid\\mathbf\{w\}\_\{t\}\)\. Applying[Lemma˜3](https://arxiv.org/html/2608.10615#Thmlemma3)yields ℒzs,ws∣wt​\(𝐰t,𝐱^θ,𝐱\)\\displaystyle\\mathcal\{L\}\_\{z\_\{s\},w\_\{s\}\\mid w\_\{t\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\}\)=DKL\[q\(𝐳s,𝐰s∣𝐰t,𝐱\)∥q\(𝐳s,𝐰s∣𝐰t,𝐱^θ\)\]\\displaystyle=D\_\{\\mathrm\{KL\}\}\\\!\\left\[q\(\\mathbf\{z\}\_\{s\},\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)\\,\\middle\\\|\\,q\(\\mathbf\{z\}\_\{s\},\\mathbf\{w\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\\right\]≤𝔼q​\(𝐳~t∣𝐰t\)\[DKL\[q\(𝐳s,𝐰s∣𝐳~t,𝐰t,𝐱\)∥q\(𝐳s,𝐰s∣𝐳~t,𝐰t,𝐱^θ\)\]\]\\displaystyle\\leq\\mathbb\{E\}\_\{q\(\\widetilde\{\\mathbf\{z\}\}\_\{t\}\\mid\\mathbf\{w\}\_\{t\}\)\}\\\!\\left\[D\_\{\\mathrm\{KL\}\}\\\!\\left\[q\(\\mathbf\{z\}\_\{s\},\\mathbf\{w\}\_\{s\}\\mid\\widetilde\{\\mathbf\{z\}\}\_\{t\},\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)\\,\\middle\\\|\\,q\(\\mathbf\{z\}\_\{s\},\\mathbf\{w\}\_\{s\}\\mid\\widetilde\{\\mathbf\{z\}\}\_\{t\},\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\\right\]\\right\]=ℒ¯zs,ws∣zt,wt​\(𝐰t,𝐱^θ,𝐱\)\.\\displaystyle=\\bar\{\\mathcal\{L\}\}\_\{z\_\{s\},w\_\{s\}\\mid z\_\{t\},w\_\{t\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\}\)\.\(94\)

Combining \([90](https://arxiv.org/html/2608.10615#A4.E90)\), \([91](https://arxiv.org/html/2608.10615#A4.E91)\), and \([94](https://arxiv.org/html/2608.10615#A4.E94)\) proves \([88](https://arxiv.org/html/2608.10615#A4.E88)\)\.

We now prove \([89](https://arxiv.org/html/2608.10615#A4.E89)\)\.

1. 1\.Using the conditional independence 𝐳s⟂⟂𝐰t∣\(𝐳~t,𝐱\),\\mathbf\{z\}\_\{s\}\\perp\\\!\\\!\\\!\\perp\\mathbf\{w\}\_\{t\}\\mid\(\\widetilde\{\\mathbf\{z\}\}\_\{t\},\\mathbf\{x\}\),and marginalizing over𝐳~t\\widetilde\{\\mathbf\{z\}\}\_\{t\}, we have q​\(𝐳s∣𝐰t,𝐱\)\\displaystyle q\(\\mathbf\{z\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)=𝔼q​\(𝐳~t∣𝐰t\)​\[q​\(𝐳s∣𝐳~t,𝐱\)\],\\displaystyle=\\mathbb\{E\}\_\{q\(\\widetilde\{\\mathbf\{z\}\}\_\{t\}\\mid\\mathbf\{w\}\_\{t\}\)\}\\\!\\left\[q\(\\mathbf\{z\}\_\{s\}\\mid\\widetilde\{\\mathbf\{z\}\}\_\{t\},\\mathbf\{x\}\)\\right\],\(95\)q​\(𝐳s∣𝐰t,𝐱^θ\)\\displaystyle q\(\\mathbf\{z\}\_\{s\}\\mid\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)=𝔼q​\(𝐳~t∣𝐰t\)​\[q​\(𝐳s∣𝐳~t,𝐱^θ\)\]\.\\displaystyle=\\mathbb\{E\}\_\{q\(\\widetilde\{\\mathbf\{z\}\}\_\{t\}\\mid\\mathbf\{w\}\_\{t\}\)\}\\\!\\left\[q\(\\mathbf\{z\}\_\{s\}\\mid\\widetilde\{\\mathbf\{z\}\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\\right\]\.\(96\)Applying[Lemma˜3](https://arxiv.org/html/2608.10615#Thmlemma3)to these mixtures yields ℒzs∣wt​\(𝐰t,𝐱^θ,𝐱\)≤ℒ¯zs∣zt,wt​\(𝐰t,𝐱^θ,𝐱\)\.\\mathcal\{L\}\_\{z\_\{s\}\\mid w\_\{t\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\}\)\\leq\\bar\{\\mathcal\{L\}\}\_\{z\_\{s\}\\mid z\_\{t\},w\_\{t\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\}\)\.\(97\)
2. 2\.For each fixed𝐳~t\\widetilde\{\\mathbf\{z\}\}\_\{t\},q​\(𝐳s∣𝐳~t,𝐱\)q\(\\mathbf\{z\}\_\{s\}\\mid\\widetilde\{\\mathbf\{z\}\}\_\{t\},\\mathbf\{x\}\)andq​\(𝐰s∣𝐳~t,𝐰t,𝐱\)q\(\\mathbf\{w\}\_\{s\}\\mid\\widetilde\{\\mathbf\{z\}\}\_\{t\},\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)are the corresponding marginals of q​\(𝐳s,𝐰s∣𝐳~t,𝐰t,𝐱\),q\(\\mathbf\{z\}\_\{s\},\\mathbf\{w\}\_\{s\}\\mid\\widetilde\{\\mathbf\{z\}\}\_\{t\},\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\),where q​\(𝐳s∣𝐳~t,𝐰t,𝐱\)=q​\(𝐳s∣𝐳~t,𝐱\)\.q\(\\mathbf\{z\}\_\{s\}\\mid\\widetilde\{\\mathbf\{z\}\}\_\{t\},\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)=q\(\\mathbf\{z\}\_\{s\}\\mid\\widetilde\{\\mathbf\{z\}\}\_\{t\},\\mathbf\{x\}\)\.The same statements hold on the model side\. Therefore, by[Corollary˜2](https://arxiv.org/html/2608.10615#Thmcorollary2), DKL\[q\(𝐳s∣𝐳~t,𝐱\)∥q\(𝐳s∣𝐳~t,𝐱^θ\)\]\\displaystyle D\_\{\\mathrm\{KL\}\}\\\!\\left\[q\(\\mathbf\{z\}\_\{s\}\\mid\\widetilde\{\\mathbf\{z\}\}\_\{t\},\\mathbf\{x\}\)\\,\\middle\\\|\\,q\(\\mathbf\{z\}\_\{s\}\\mid\\widetilde\{\\mathbf\{z\}\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\\right\]≤DKL\[q\(𝐳s,𝐰s∣𝐳~t,𝐰t,𝐱\)∥q\(𝐳s,𝐰s∣𝐳~t,𝐰t,𝐱^θ\)\],\\displaystyle\\qquad\\leq D\_\{\\mathrm\{KL\}\}\\\!\\left\[q\(\\mathbf\{z\}\_\{s\},\\mathbf\{w\}\_\{s\}\\mid\\widetilde\{\\mathbf\{z\}\}\_\{t\},\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)\\,\\middle\\\|\\,q\(\\mathbf\{z\}\_\{s\},\\mathbf\{w\}\_\{s\}\\mid\\widetilde\{\\mathbf\{z\}\}\_\{t\},\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\\right\],\(98\)DKL\[q\(𝐰s∣𝐳~t,𝐰t,𝐱\)∥q\(𝐰s∣𝐳~t,𝐰t,𝐱^θ\)\]\\displaystyle D\_\{\\mathrm\{KL\}\}\\\!\\left\[q\(\\mathbf\{w\}\_\{s\}\\mid\\widetilde\{\\mathbf\{z\}\}\_\{t\},\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)\\,\\middle\\\|\\,q\(\\mathbf\{w\}\_\{s\}\\mid\\widetilde\{\\mathbf\{z\}\}\_\{t\},\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\\right\]≤DKL\[q\(𝐳s,𝐰s∣𝐳~t,𝐰t,𝐱\)∥q\(𝐳s,𝐰s∣𝐳~t,𝐰t,𝐱^θ\)\]\.\\displaystyle\\qquad\\leq D\_\{\\mathrm\{KL\}\}\\\!\\left\[q\(\\mathbf\{z\}\_\{s\},\\mathbf\{w\}\_\{s\}\\mid\\widetilde\{\\mathbf\{z\}\}\_\{t\},\\mathbf\{w\}\_\{t\},\\mathbf\{x\}\)\\,\\middle\\\|\\,q\(\\mathbf\{z\}\_\{s\},\\mathbf\{w\}\_\{s\}\\mid\\widetilde\{\\mathbf\{z\}\}\_\{t\},\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\\right\]\.\(99\)Averaging both inequalities overq​\(𝐳~t∣𝐰t\)q\(\\widetilde\{\\mathbf\{z\}\}\_\{t\}\\mid\\mathbf\{w\}\_\{t\}\)gives ℒ¯zs∣zt,wt​\(𝐰t,𝐱^θ,𝐱\)\\displaystyle\\bar\{\\mathcal\{L\}\}\_\{z\_\{s\}\\mid z\_\{t\},w\_\{t\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\}\)≤ℒ¯zs,ws∣zt,wt​\(𝐰t,𝐱^θ,𝐱\),\\displaystyle\\leq\\bar\{\\mathcal\{L\}\}\_\{z\_\{s\},w\_\{s\}\\mid z\_\{t\},w\_\{t\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\}\),\(100\)ℒ¯ws∣zt,wt​\(𝐰t,𝐱^θ,𝐱\)\\displaystyle\\bar\{\\mathcal\{L\}\}\_\{w\_\{s\}\\mid z\_\{t\},w\_\{t\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\}\)≤ℒ¯zs,ws∣zt,wt​\(𝐰t,𝐱^θ,𝐱\)\.\\displaystyle\\leq\\bar\{\\mathcal\{L\}\}\_\{z\_\{s\},w\_\{s\}\\mid z\_\{t\},w\_\{t\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\}\)\.\(101\)

Combining \([97](https://arxiv.org/html/2608.10615#A4.E97)\), \([100](https://arxiv.org/html/2608.10615#A4.E100)\), and \([101](https://arxiv.org/html/2608.10615#A4.E101)\) proves \([89](https://arxiv.org/html/2608.10615#A4.E89)\)\. ∎

The joint surrogate contains an additional simplex\-matching term\. The following identity and proposition give its closed form\.

###### Lemma 4\.

For anyj,k∈\{1,…,K\}j,k\\in\\\{1,\\dots,K\\\},

𝔼Dir⁡\(⋅;ηs​𝐩s\+𝐞k\)​\[log⁡wj\]=ψ​\(ηs​ps,j​\(𝐱\)\+δj,k\)−ψ​\(ηs\+1\)\.\\mathbb\{E\}\_\{\\operatorname\{Dir\}\\\!\\left\(\\cdot;\\eta\_\{s\}\\mathbf\{p\}\_\{s\}\+\\mathbf\{e\}\_\{k\}\\right\)\}\\\!\\left\[\\log w\_\{j\}\\right\]=\\psi\\\!\\left\(\\eta\_\{s\}p\_\{s,j\}\(\\mathbf\{x\}\)\+\\delta\_\{j,k\}\\right\)\-\\psi\(\\eta\_\{s\}\+1\)\.\(102\)

###### Proof\.

For a Dirichlet random vector with parameter𝜶=\(α1,…,αK\)\\bm\{\\alpha\}=\(\\alpha\_\{1\},\\dots,\\alpha\_\{K\}\), the standard identity is

𝔼Dir⁡\(⋅;𝜶\)​\[log⁡wj\]=ψ​\(αj\)−ψ​\(∑m=1Kαm\)\.\\mathbb\{E\}\_\{\\operatorname\{Dir\}\(\\cdot;\\bm\{\\alpha\}\)\}\\\!\\left\[\\log w\_\{j\}\\right\]=\\psi\(\\alpha\_\{j\}\)\-\\psi\\\!\\left\(\\sum\_\{m=1\}^\{K\}\\alpha\_\{m\}\\right\)\.Applying this with

𝜶=ηs​𝐩s\+𝐞k\\bm\{\\alpha\}=\\eta\_\{s\}\\mathbf\{p\}\_\{s\}\+\\mathbf\{e\}\_\{k\}gives

αj=ηs​ps,j​\(𝐱\)\+δj,k,∑m=1Kαm=ηs​∑m=1Kps,m​\(𝐱\)\+1=ηs\+1,\\alpha\_\{j\}=\\eta\_\{s\}p\_\{s,j\}\(\\mathbf\{x\}\)\+\\delta\_\{j,k\},\\qquad\\sum\_\{m=1\}^\{K\}\\alpha\_\{m\}=\\eta\_\{s\}\\sum\_\{m=1\}^\{K\}p\_\{s,m\}\(\\mathbf\{x\}\)\+1=\\eta\_\{s\}\+1,which proves \([102](https://arxiv.org/html/2608.10615#A4.E102)\)\. ∎

###### Proposition 8\.

For0<s<t≤10<s<t\\leq 1, the simplex\-matching term satisfies

ℒ¯ws∣zs,wt​\(𝐰t,𝐱^θ,𝐱;s,t\)\\displaystyle\\bar\{\\mathcal\{L\}\}\_\{w\_\{s\}\\mid z\_\{s\},w\_\{t\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\};s,t\)=DKL\[Dir\(⋅;ηs𝐩s\)∥Dir\(⋅;ηs𝐩^s\)\]\\displaystyle=D\_\{\\mathrm\{KL\}\}\\\!\\left\[\\operatorname\{Dir\}\\\!\\left\(\\cdot;\\eta\_\{s\}\\mathbf\{p\}\_\{s\}\\right\)\\,\\middle\\\|\\,\\operatorname\{Dir\}\\\!\\left\(\\cdot;\\eta\_\{s\}\\hat\{\\mathbf\{p\}\}\_\{s\}\\right\)\\right\]\+⟨𝝆s∣t​\(𝐱,𝐰t\),log⁡𝐩^s−log⁡𝐩s\+𝟏−𝐩^s⊘𝐩s⟩\.\\displaystyle\\quad\+\\left\\langle\\bm\{\\rho\}\_\{s\\mid t\}\(\\mathbf\{x\},\\mathbf\{w\}\_\{t\}\),\\,\\log\\hat\{\\mathbf\{p\}\}\_\{s\}\-\\log\\mathbf\{p\}\_\{s\}\+\\mathbf\{1\}\-\\hat\{\\mathbf\{p\}\}\_\{s\}\\oslash\\mathbf\{p\}\_\{s\}\\right\\rangle\.\(103\)

###### Proof\.

We compute the simplex termℒ¯ws∣zs,wt\\bar\{\\mathcal\{L\}\}\_\{w\_\{s\}\\mid z\_\{s\},w\_\{t\}\}\. By definition,

ℒ¯ws∣zs,wt=∑k=1Kρs∣t,k\(𝐱,𝐰t\)DKL\[Dir\(⋅;ηs𝐩s\+𝐞k\)∥Dir\(⋅;ηs𝐩^s\+𝐞k\)\]\.\\bar\{\\mathcal\{L\}\}\_\{w\_\{s\}\\mid z\_\{s\},w\_\{t\}\}=\\sum\_\{k=1\}^\{K\}\\rho\_\{s\\mid t,k\}\(\\mathbf\{x\},\\mathbf\{w\}\_\{t\}\)\\,D\_\{\\mathrm\{KL\}\}\\\!\\left\[\\operatorname\{Dir\}\\\!\\left\(\\cdot;\\eta\_\{s\}\\mathbf\{p\}\_\{s\}\+\\mathbf\{e\}\_\{k\}\\right\)\\,\\middle\\\|\\,\\operatorname\{Dir\}\\\!\\left\(\\cdot;\\eta\_\{s\}\\hat\{\\mathbf\{p\}\}\_\{s\}\+\\mathbf\{e\}\_\{k\}\\right\)\\right\]\.\(104\)Fixk∈\{1,…,K\}k\\in\\\{1,\\dots,K\\\}\. By[Lemma˜1](https://arxiv.org/html/2608.10615#Thmlemma1),

Dir⁡\(𝐰;ηs​𝐩s\+𝐞k\)\\displaystyle\\operatorname\{Dir\}\\\!\\left\(\\mathbf\{w\};\\eta\_\{s\}\\mathbf\{p\}\_\{s\}\+\\mathbf\{e\}\_\{k\}\\right\)=wkps,k​\(𝐱\)​Dir⁡\(𝐰;ηs​𝐩s\),\\displaystyle=\\frac\{w\_\{k\}\}\{p\_\{s,k\}\(\\mathbf\{x\}\)\}\\operatorname\{Dir\}\\\!\\left\(\\mathbf\{w\};\\eta\_\{s\}\\mathbf\{p\}\_\{s\}\\right\),\(105\)Dir⁡\(𝐰;ηs​𝐩^s\+𝐞k\)\\displaystyle\\operatorname\{Dir\}\\\!\\left\(\\mathbf\{w\};\\eta\_\{s\}\\hat\{\\mathbf\{p\}\}\_\{s\}\+\\mathbf\{e\}\_\{k\}\\right\)=wkps,k​\(𝐱^θ\)​Dir⁡\(𝐰;ηs​𝐩^s\)\.\\displaystyle=\\frac\{w\_\{k\}\}\{p\_\{s,k\}\(\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\}\\operatorname\{Dir\}\\\!\\left\(\\mathbf\{w\};\\eta\_\{s\}\\hat\{\\mathbf\{p\}\}\_\{s\}\\right\)\.\(106\)Taking the logarithm of the ratio gives

log⁡Dir⁡\(𝐰;ηs​𝐩s\+𝐞k\)Dir⁡\(𝐰;ηs​𝐩^s\+𝐞k\)\\displaystyle\\log\\frac\{\\operatorname\{Dir\}\\\!\\left\(\\mathbf\{w\};\\eta\_\{s\}\\mathbf\{p\}\_\{s\}\+\\mathbf\{e\}\_\{k\}\\right\)\}\{\\operatorname\{Dir\}\\\!\\left\(\\mathbf\{w\};\\eta\_\{s\}\\hat\{\\mathbf\{p\}\}\_\{s\}\+\\mathbf\{e\}\_\{k\}\\right\)\}=log⁡Dir⁡\(𝐰;ηs​𝐩s\)Dir⁡\(𝐰;ηs​𝐩^s\)\+log⁡ps,k​\(𝐱^θ\)−log⁡ps,k​\(𝐱\)\.\\displaystyle=\\log\\frac\{\\operatorname\{Dir\}\\\!\\left\(\\mathbf\{w\};\\eta\_\{s\}\\mathbf\{p\}\_\{s\}\\right\)\}\{\\operatorname\{Dir\}\\\!\\left\(\\mathbf\{w\};\\eta\_\{s\}\\hat\{\\mathbf\{p\}\}\_\{s\}\\right\)\}\+\\log p\_\{s,k\}\(\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\-\\log p\_\{s,k\}\(\\mathbf\{x\}\)\.\(107\)
We now compare the expectation of the unshifted log\-ratio under the shifted Dirichlet law with the KL between the unshifted Dirichlet distributions\. Expanding the Dirichlet density gives

log⁡Dir⁡\(𝐰;ηs​𝐩s\)Dir⁡\(𝐰;ηs​𝐩^s\)\\displaystyle\\log\\frac\{\\operatorname\{Dir\}\\\!\\left\(\\mathbf\{w\};\\eta\_\{s\}\\mathbf\{p\}\_\{s\}\\right\)\}\{\\operatorname\{Dir\}\\\!\\left\(\\mathbf\{w\};\\eta\_\{s\}\\hat\{\\mathbf\{p\}\}\_\{s\}\\right\)\}=log⁡B​\(ηs​𝐩^s\)B​\(ηs​𝐩s\)\+ηs​∑j=1K\(ps,j​\(𝐱\)−ps,j​\(𝐱^θ\)\)​log⁡wj\.\\displaystyle=\\log\\frac\{B\(\\eta\_\{s\}\\hat\{\\mathbf\{p\}\}\_\{s\}\)\}\{B\(\\eta\_\{s\}\\mathbf\{p\}\_\{s\}\)\}\+\\eta\_\{s\}\\sum\_\{j=1\}^\{K\}\\bigl\(p\_\{s,j\}\(\\mathbf\{x\}\)\-p\_\{s,j\}\(\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\\bigr\)\\log w\_\{j\}\.\(108\)Taking expectation underDir⁡\(⋅;ηs​𝐩s\+𝐞k\)\\operatorname\{Dir\}\(\\cdot;\\eta\_\{s\}\\mathbf\{p\}\_\{s\}\+\\mathbf\{e\}\_\{k\}\)and applying[Lemma˜4](https://arxiv.org/html/2608.10615#Thmlemma4)gives

𝔼Dir⁡\(⋅;ηs​𝐩s\+𝐞k\)​\[log⁡Dir⁡\(𝐰;ηs​𝐩s\)Dir⁡\(𝐰;ηs​𝐩^s\)\]\\displaystyle\\mathbb\{E\}\_\{\\operatorname\{Dir\}\\\!\\left\(\\cdot;\\eta\_\{s\}\\mathbf\{p\}\_\{s\}\+\\mathbf\{e\}\_\{k\}\\right\)\}\\\!\\left\[\\log\\frac\{\\operatorname\{Dir\}\\\!\\left\(\\mathbf\{w\};\\eta\_\{s\}\\mathbf\{p\}\_\{s\}\\right\)\}\{\\operatorname\{Dir\}\\\!\\left\(\\mathbf\{w\};\\eta\_\{s\}\\hat\{\\mathbf\{p\}\}\_\{s\}\\right\)\}\\right\]=log⁡B​\(ηs​𝐩^s\)B​\(ηs​𝐩s\)\+ηs​∑j=1K\(ps,j​\(𝐱\)−ps,j​\(𝐱^θ\)\)​\(ψ​\(ηs​ps,j​\(𝐱\)\+δj,k\)−ψ​\(ηs\+1\)\)\.\\displaystyle\\qquad=\\log\\frac\{B\(\\eta\_\{s\}\\hat\{\\mathbf\{p\}\}\_\{s\}\)\}\{B\(\\eta\_\{s\}\\mathbf\{p\}\_\{s\}\)\}\+\\eta\_\{s\}\\sum\_\{j=1\}^\{K\}\\bigl\(p\_\{s,j\}\(\\mathbf\{x\}\)\-p\_\{s,j\}\(\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\\bigr\)\\bigl\(\\psi\(\\eta\_\{s\}p\_\{s,j\}\(\\mathbf\{x\}\)\+\\delta\_\{j,k\}\)\-\\psi\(\\eta\_\{s\}\+1\)\\bigr\)\.\(109\)On the other hand, the KL divergence between the unshifted Dirichlet distributions is

DKL\[Dir\(⋅;ηs𝐩s\)∥Dir\(⋅;ηs𝐩^s\)\]\\displaystyle D\_\{\\mathrm\{KL\}\}\\\!\\left\[\\operatorname\{Dir\}\\\!\\left\(\\cdot;\\eta\_\{s\}\\mathbf\{p\}\_\{s\}\\right\)\\,\\middle\\\|\\,\\operatorname\{Dir\}\\\!\\left\(\\cdot;\\eta\_\{s\}\\hat\{\\mathbf\{p\}\}\_\{s\}\\right\)\\right\]=log⁡B​\(ηs​𝐩^s\)B​\(ηs​𝐩s\)\+ηs​∑j=1K\(ps,j​\(𝐱\)−ps,j​\(𝐱^θ\)\)​\(ψ​\(ηs​ps,j​\(𝐱\)\)−ψ​\(ηs\)\)\.\\displaystyle=\\log\\frac\{B\(\\eta\_\{s\}\\hat\{\\mathbf\{p\}\}\_\{s\}\)\}\{B\(\\eta\_\{s\}\\mathbf\{p\}\_\{s\}\)\}\+\\eta\_\{s\}\\sum\_\{j=1\}^\{K\}\\bigl\(p\_\{s,j\}\(\\mathbf\{x\}\)\-p\_\{s,j\}\(\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\\bigr\)\\bigl\(\\psi\(\\eta\_\{s\}p\_\{s,j\}\(\\mathbf\{x\}\)\)\-\\psi\(\\eta\_\{s\}\)\\bigr\)\.\(110\)Subtracting \([110](https://arxiv.org/html/2608.10615#A4.E110)\) from \([109](https://arxiv.org/html/2608.10615#A4.E109)\), and using

ψ​\(a\+1\)=ψ​\(a\)\+1a,ψ​\(ηs\+1\)=ψ​\(ηs\)\+1ηs,\\psi\(a\+1\)=\\psi\(a\)\+\\frac\{1\}\{a\},\\qquad\\psi\(\\eta\_\{s\}\+1\)=\\psi\(\\eta\_\{s\}\)\+\\frac\{1\}\{\\eta\_\{s\}\},yields

𝔼Dir⁡\(⋅;ηs​𝐩s\+𝐞k\)​\[log⁡Dir⁡\(𝐰;ηs​𝐩s\)Dir⁡\(𝐰;ηs​𝐩^s\)\]\\displaystyle\\mathbb\{E\}\_\{\\operatorname\{Dir\}\\\!\\left\(\\cdot;\\eta\_\{s\}\\mathbf\{p\}\_\{s\}\+\\mathbf\{e\}\_\{k\}\\right\)\}\\\!\\left\[\\log\\frac\{\\operatorname\{Dir\}\\\!\\left\(\\mathbf\{w\};\\eta\_\{s\}\\mathbf\{p\}\_\{s\}\\right\)\}\{\\operatorname\{Dir\}\\\!\\left\(\\mathbf\{w\};\\eta\_\{s\}\\hat\{\\mathbf\{p\}\}\_\{s\}\\right\)\}\\right\]=DKL\[Dir\(⋅;ηs𝐩s\)∥Dir\(⋅;ηs𝐩^s\)\]\+1−ps,k​\(𝐱^θ\)ps,k​\(𝐱\)\.\\displaystyle=D\_\{\\mathrm\{KL\}\}\\\!\\left\[\\operatorname\{Dir\}\\\!\\left\(\\cdot;\\eta\_\{s\}\\mathbf\{p\}\_\{s\}\\right\)\\,\\middle\\\|\\,\\operatorname\{Dir\}\\\!\\left\(\\cdot;\\eta\_\{s\}\\hat\{\\mathbf\{p\}\}\_\{s\}\\right\)\\right\]\+1\-\\frac\{p\_\{s,k\}\(\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\}\{p\_\{s,k\}\(\\mathbf\{x\}\)\}\.\(111\)Combining \([107](https://arxiv.org/html/2608.10615#A4.E107)\) with \([111](https://arxiv.org/html/2608.10615#A4.E111)\), we obtain

DKL\[Dir\(⋅;ηs𝐩s\+𝐞k\)∥Dir\(⋅;ηs𝐩^s\+𝐞k\)\]\\displaystyle D\_\{\\mathrm\{KL\}\}\\\!\\left\[\\operatorname\{Dir\}\\\!\\left\(\\cdot;\\eta\_\{s\}\\mathbf\{p\}\_\{s\}\+\\mathbf\{e\}\_\{k\}\\right\)\\,\\middle\\\|\\,\\operatorname\{Dir\}\\\!\\left\(\\cdot;\\eta\_\{s\}\\hat\{\\mathbf\{p\}\}\_\{s\}\+\\mathbf\{e\}\_\{k\}\\right\)\\right\]=DKL\[Dir\(⋅;ηs𝐩s\)∥Dir\(⋅;ηs𝐩^s\)\]\+logps,k\(𝐱^θ\)−logps,k\(𝐱\)\\displaystyle=D\_\{\\mathrm\{KL\}\}\\\!\\left\[\\operatorname\{Dir\}\\\!\\left\(\\cdot;\\eta\_\{s\}\\mathbf\{p\}\_\{s\}\\right\)\\,\\middle\\\|\\,\\operatorname\{Dir\}\\\!\\left\(\\cdot;\\eta\_\{s\}\\hat\{\\mathbf\{p\}\}\_\{s\}\\right\)\\right\]\+\\log p\_\{s,k\}\(\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\-\\log p\_\{s,k\}\(\\mathbf\{x\}\)\+1−ps,k​\(𝐱^θ\)ps,k​\(𝐱\)\.\\displaystyle\\quad\+1\-\\frac\{p\_\{s,k\}\(\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\}\{p\_\{s,k\}\(\\mathbf\{x\}\)\}\.\(112\)Substituting \([112](https://arxiv.org/html/2608.10615#A4.E112)\) into \([104](https://arxiv.org/html/2608.10615#A4.E104)\), and using∑k=1Kρs∣t,k​\(𝐱,𝐰t\)=1\\sum\_\{k=1\}^\{K\}\\rho\_\{s\\mid t,k\}\(\\mathbf\{x\},\\mathbf\{w\}\_\{t\}\)=1, yields

ℒ¯ws∣zs,wt\\displaystyle\\bar\{\\mathcal\{L\}\}\_\{w\_\{s\}\\mid z\_\{s\},w\_\{t\}\}=DKL\[Dir\(⋅;ηs𝐩s\)∥Dir\(⋅;ηs𝐩^s\)\]\\displaystyle=D\_\{\\mathrm\{KL\}\}\\\!\\left\[\\operatorname\{Dir\}\\\!\\left\(\\cdot;\\eta\_\{s\}\\mathbf\{p\}\_\{s\}\\right\)\\,\\middle\\\|\\,\\operatorname\{Dir\}\\\!\\left\(\\cdot;\\eta\_\{s\}\\hat\{\\mathbf\{p\}\}\_\{s\}\\right\)\\right\]\+∑k=1Kρs∣t,k​\(𝐱,𝐰t\)​\(log⁡ps,k​\(𝐱^θ\)−log⁡ps,k​\(𝐱\)\+1−ps,k​\(𝐱^θ\)ps,k​\(𝐱\)\)\\displaystyle\\quad\+\\sum\_\{k=1\}^\{K\}\\rho\_\{s\\mid t,k\}\(\\mathbf\{x\},\\mathbf\{w\}\_\{t\}\)\\left\(\\log p\_\{s,k\}\(\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\-\\log p\_\{s,k\}\(\\mathbf\{x\}\)\+1\-\\frac\{p\_\{s,k\}\(\\hat\{\\mathbf\{x\}\}\_\{\\theta\}\)\}\{p\_\{s,k\}\(\\mathbf\{x\}\)\}\\right\)=DKL\[Dir\(⋅;ηs𝐩s\)∥Dir\(⋅;ηs𝐩^s\)\]\+⟨𝝆s∣t\(𝐱,𝐰t\),log𝐩^s−log𝐩s\+𝟏−𝐩^s⊘𝐩s⟩,\\displaystyle=D\_\{\\mathrm\{KL\}\}\\\!\\left\[\\operatorname\{Dir\}\\\!\\left\(\\cdot;\\eta\_\{s\}\\mathbf\{p\}\_\{s\}\\right\)\\,\\middle\\\|\\,\\operatorname\{Dir\}\\\!\\left\(\\cdot;\\eta\_\{s\}\\hat\{\\mathbf\{p\}\}\_\{s\}\\right\)\\right\]\+\\left\\langle\\bm\{\\rho\}\_\{s\\mid t\}\(\\mathbf\{x\},\\mathbf\{w\}\_\{t\}\),\\,\\log\\hat\{\\mathbf\{p\}\}\_\{s\}\-\\log\\mathbf\{p\}\_\{s\}\+\\mathbf\{1\}\-\\hat\{\\mathbf\{p\}\}\_\{s\}\\oslash\\mathbf\{p\}\_\{s\}\\right\\rangle,\(113\)which proves \([103](https://arxiv.org/html/2608.10615#A4.E103)\)\. ∎

### D\.4Why the other surrogate objectives do not yield suitable continuous\-time objectives

The relaxed discrete bridge is distinguished by its first\-order scaling\. We now show that the remaining surrogates behave differently in the local limit: the tight discrete objective vanishes at second order, whereas the joint objectives retain anO​\(1\)O\(1\)simplex\-matching term\.

###### Proposition 9\.

Lets=t−Δs=t\-\\DeltawithΔ↓0\\Delta\\downarrow 0, and assume thatηt\\eta\_\{t\}is continuous intt\.

1. 1\.The tight discrete objective satisfies ℒzt−Δ∣wt​\(𝐰t,𝐱^θ,𝐱;t−Δ,t\)=O​\(Δ2\)\.\\mathcal\{L\}\_\{z\_\{t\-\\Delta\}\\mid w\_\{t\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\};t\-\\Delta,t\)=O\(\\Delta^\{2\}\)\.\(114\)
2. 2\.Define 𝒮t​\(𝐰t,𝐱^θ,𝐱\)\\displaystyle\\mathcal\{S\}\_\{t\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\}\)≔DKL\[Dir\(⋅;ηt𝐩t\)∥Dir\(⋅;ηt𝐩^t\)\]\+⟨𝐰t,log𝐩^t−log𝐩t\+𝟏−𝐩^t⊘𝐩t⟩\.\\displaystyle\\coloneq D\_\{\\mathrm\{KL\}\}\\\!\\left\[\\operatorname\{Dir\}\\\!\\left\(\\cdot;\\eta\_\{t\}\\mathbf\{p\}\_\{t\}\\right\)\\,\\middle\\\|\\,\\operatorname\{Dir\}\\\!\\left\(\\cdot;\\eta\_\{t\}\\hat\{\\mathbf\{p\}\}\_\{t\}\\right\)\\right\]\+\\left\\langle\\mathbf\{w\}\_\{t\},\\,\\log\\hat\{\\mathbf\{p\}\}\_\{t\}\-\\log\\mathbf\{p\}\_\{t\}\+\\mathbf\{1\}\-\\hat\{\\mathbf\{p\}\}\_\{t\}\\oslash\\mathbf\{p\}\_\{t\}\\right\\rangle\.\(115\)Then the simplex term satisfies ℒ¯wt−Δ∣zt−Δ,wt​\(𝐰t,𝐱^θ,𝐱;t−Δ,t\)=𝒮t​\(𝐰t,𝐱^θ,𝐱\)\+o​\(1\)\.\\bar\{\\mathcal\{L\}\}\_\{w\_\{t\-\\Delta\}\\mid z\_\{t\-\\Delta\},w\_\{t\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\};t\-\\Delta,t\)=\\mathcal\{S\}\_\{t\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\}\)\+o\(1\)\.\(116\)
3. 3\.Consequently, ℒzt−Δ,wt−Δ∣wt​\(𝐰t,𝐱^θ,𝐱;t−Δ,t\)\\displaystyle\\mathcal\{L\}\_\{z\_\{t\-\\Delta\},w\_\{t\-\\Delta\}\\mid w\_\{t\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\};t\-\\Delta,t\)=𝒮t​\(𝐰t,𝐱^θ,𝐱\)\+o​\(1\),\\displaystyle=\\mathcal\{S\}\_\{t\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\}\)\+o\(1\),\(117\)ℒ¯zt−Δ,wt−Δ∣zt,wt​\(𝐰t,𝐱^θ,𝐱;t−Δ,t\)\\displaystyle\\bar\{\\mathcal\{L\}\}\_\{z\_\{t\-\\Delta\},w\_\{t\-\\Delta\}\\mid z\_\{t\},w\_\{t\}\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\};t\-\\Delta,t\)=𝒮t​\(𝐰t,𝐱^θ,𝐱\)\+o​\(1\)\.\\displaystyle=\\mathcal\{S\}\_\{t\}\(\\mathbf\{w\}\_\{t\},\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{x\}\)\+o\(1\)\.\(118\)

###### Proof\.

1. 1\.We first analyze the tight discrete objective\. Using[Equation˜29](https://arxiv.org/html/2608.10615#A1.E29)together with the same expansions \([54](https://arxiv.org/html/2608.10615#A3.E54)\) and \([56](https://arxiv.org/html/2608.10615#A3.E56)\) as above, we obtain 𝝆t−Δ∣t​\(𝐱,𝐰t\)\\displaystyle\\bm\{\\rho\}\_\{t\-\\Delta\\mid t\}\(\\mathbf\{x\},\\mathbf\{w\}\_\{t\}\)=𝐰t\+Δ​𝐚t​\(𝐱,𝐰t\)\+o​\(Δ\),\\displaystyle=\\mathbf\{w\}\_\{t\}\+\\Delta\\,\\mathbf\{a\}\_\{t\}\(\\mathbf\{x\},\\mathbf\{w\}\_\{t\}\)\+o\(\\Delta\),\(119\)𝝆t−Δ∣t​\(𝐱^θ,𝐰t\)\\displaystyle\\bm\{\\rho\}\_\{t\-\\Delta\\mid t\}\(\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{w\}\_\{t\}\)=𝐰t\+Δ​𝐚^t​\(𝐱^θ,𝐰t\)\+o​\(Δ\),\\displaystyle=\\mathbf\{w\}\_\{t\}\+\\Delta\\,\\hat\{\\mathbf\{a\}\}\_\{t\}\(\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{w\}\_\{t\}\)\+o\(\\Delta\),\(120\)where 𝐚t​\(𝐱,𝐰t\)\\displaystyle\\mathbf\{a\}\_\{t\}\(\\mathbf\{x\},\\mathbf\{w\}\_\{t\}\)=λ​\(t\)​\[⟨𝐰t,𝝅⊘𝐩t⟩​𝐩t−𝝅⊙\(𝐰t⊘𝐩t\)\],\\displaystyle=\\lambda\(t\)\\left\[\\left\\langle\\mathbf\{w\}\_\{t\},\\,\\bm\{\\pi\}\\oslash\\mathbf\{p\}\_\{t\}\\right\\rangle\\mathbf\{p\}\_\{t\}\-\\bm\{\\pi\}\\odot\(\\mathbf\{w\}\_\{t\}\\oslash\\mathbf\{p\}\_\{t\}\)\\right\],\(121\)𝐚^t​\(𝐱^θ,𝐰t\)\\displaystyle\\hat\{\\mathbf\{a\}\}\_\{t\}\(\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{w\}\_\{t\}\)=λ​\(t\)​\[⟨𝐰t,𝝅⊘𝐩^t⟩​𝐩^t−𝝅⊙\(𝐰t⊘𝐩^t\)\]\.\\displaystyle=\\lambda\(t\)\\left\[\\left\\langle\\mathbf\{w\}\_\{t\},\\,\\bm\{\\pi\}\\oslash\\hat\{\\mathbf\{p\}\}\_\{t\}\\right\\rangle\\hat\{\\mathbf\{p\}\}\_\{t\}\-\\bm\{\\pi\}\\odot\(\\mathbf\{w\}\_\{t\}\\oslash\\hat\{\\mathbf\{p\}\}\_\{t\}\)\\right\]\.\(122\)Since both vectors in \([119](https://arxiv.org/html/2608.10615#A4.E119)\) and \([120](https://arxiv.org/html/2608.10615#A4.E120)\) are probability vectors, their first\-order perturbations satisfy ⟨𝟏,𝐚t​\(𝐱,𝐰t\)⟩=0,⟨𝟏,𝐚^t​\(𝐱^θ,𝐰t\)⟩=0\.\\left\\langle\\mathbf\{1\},\\,\\mathbf\{a\}\_\{t\}\(\\mathbf\{x\},\\mathbf\{w\}\_\{t\}\)\\right\\rangle=0,\\qquad\\left\\langle\\mathbf\{1\},\\,\\hat\{\\mathbf\{a\}\}\_\{t\}\(\\hat\{\\mathbf\{x\}\}\_\{\\theta\},\\mathbf\{w\}\_\{t\}\)\\right\\rangle=0\.Therefore, expanding the categorical KL around the common base point𝐰t\\mathbf\{w\}\_\{t\}shows that the first\-order term cancels: ℒzt−Δ∣wt\\displaystyle\\mathcal\{L\}\_\{z\_\{t\-\\Delta\}\\mid w\_\{t\}\}=DKL\[Cat\(⋅;𝐰t\+Δ𝐚t\+o\(Δ\)\)∥Cat\(⋅;𝐰t\+Δ𝐚^t\+o\(Δ\)\)\]\\displaystyle=D\_\{\\mathrm\{KL\}\}\\\!\\left\[\\operatorname\{Cat\}\\\!\\left\(\\cdot;\\mathbf\{w\}\_\{t\}\+\\Delta\\,\\mathbf\{a\}\_\{t\}\+o\(\\Delta\)\\right\)\\,\\middle\\\|\\,\\operatorname\{Cat\}\\\!\\left\(\\cdot;\\mathbf\{w\}\_\{t\}\+\\Delta\\,\\hat\{\\mathbf\{a\}\}\_\{t\}\+o\(\\Delta\)\\right\)\\right\]=O​\(Δ2\)\.\\displaystyle=O\(\\Delta^\{2\}\)\.\(123\)This proves \([114](https://arxiv.org/html/2608.10615#A4.E114)\)\.
2. 2\.We next analyze the simplex term using its closed form \([103](https://arxiv.org/html/2608.10615#A4.E103)\)\. Sinceηt\\eta\_\{t\}is continuous, ηt−Δ=ηt\+o​\(1\)\.\\eta\_\{t\-\\Delta\}=\\eta\_\{t\}\+o\(1\)\.Moreover, by \([56](https://arxiv.org/html/2608.10615#A3.E56)\) and its analogue for𝐩^t−Δ\\hat\{\\mathbf\{p\}\}\_\{t\-\\Delta\}, 𝐩t−Δ=𝐩t\+O​\(Δ\),𝐩^t−Δ=𝐩^t\+O​\(Δ\)\.\\mathbf\{p\}\_\{t\-\\Delta\}=\\mathbf\{p\}\_\{t\}\+O\(\\Delta\),\\qquad\\hat\{\\mathbf\{p\}\}\_\{t\-\\Delta\}=\\hat\{\\mathbf\{p\}\}\_\{t\}\+O\(\\Delta\)\.Finally,[Equation˜57](https://arxiv.org/html/2608.10615#A3.E57)gives 𝝆t−Δ∣t​\(𝐱,𝐰t\)=𝐰t\+O​\(Δ\)\.\\bm\{\\rho\}\_\{t\-\\Delta\\mid t\}\(\\mathbf\{x\},\\mathbf\{w\}\_\{t\}\)=\\mathbf\{w\}\_\{t\}\+O\(\\Delta\)\.Substituting these expansions into \([103](https://arxiv.org/html/2608.10615#A4.E103)\) yields ℒ¯wt−Δ∣zt−Δ,wt\\displaystyle\\bar\{\\mathcal\{L\}\}\_\{w\_\{t\-\\Delta\}\\mid z\_\{t\-\\Delta\},w\_\{t\}\}=DKL\[Dir\(⋅;ηt𝐩t\)∥Dir\(⋅;ηt𝐩^t\)\]\\displaystyle=D\_\{\\mathrm\{KL\}\}\\\!\\left\[\\operatorname\{Dir\}\\\!\\left\(\\cdot;\\eta\_\{t\}\\mathbf\{p\}\_\{t\}\\right\)\\,\\middle\\\|\\,\\operatorname\{Dir\}\\\!\\left\(\\cdot;\\eta\_\{t\}\\hat\{\\mathbf\{p\}\}\_\{t\}\\right\)\\right\]\+⟨𝐰t,log⁡𝐩^t−log⁡𝐩t\+𝟏−𝐩^t⊘𝐩t⟩\+o​\(1\),\\displaystyle\\quad\+\\left\\langle\\mathbf\{w\}\_\{t\},\\,\\log\\hat\{\\mathbf\{p\}\}\_\{t\}\-\\log\\mathbf\{p\}\_\{t\}\+\\mathbf\{1\}\-\\hat\{\\mathbf\{p\}\}\_\{t\}\\oslash\\mathbf\{p\}\_\{t\}\\right\\rangle\+o\(1\),\(124\)which is exactly \([116](https://arxiv.org/html/2608.10615#A4.E116)\)\.
3. 3\.The asymptotics of the two joint objectives now follow from the decompositions in[Equations˜73](https://arxiv.org/html/2608.10615#A4.E73)and[74](https://arxiv.org/html/2608.10615#A4.E74)\. For the tight joint objective, ℒzt−Δ,wt−Δ∣wt=ℒzt−Δ∣wt\+ℒ¯wt−Δ∣zt−Δ,wt,\\mathcal\{L\}\_\{z\_\{t\-\\Delta\},w\_\{t\-\\Delta\}\\mid w\_\{t\}\}=\\mathcal\{L\}\_\{z\_\{t\-\\Delta\}\\mid w\_\{t\}\}\+\\bar\{\\mathcal\{L\}\}\_\{w\_\{t\-\\Delta\}\\mid z\_\{t\-\\Delta\},w\_\{t\}\},and combining \([114](https://arxiv.org/html/2608.10615#A4.E114)\) with \([116](https://arxiv.org/html/2608.10615#A4.E116)\) gives \([117](https://arxiv.org/html/2608.10615#A4.E117)\)\. For the loose joint objective, ℒ¯zt−Δ,wt−Δ∣zt,wt=ℒ¯zt−Δ∣zt,wt\+ℒ¯wt−Δ∣zt−Δ,wt\.\\bar\{\\mathcal\{L\}\}\_\{z\_\{t\-\\Delta\},w\_\{t\-\\Delta\}\\mid z\_\{t\},w\_\{t\}\}=\\bar\{\\mathcal\{L\}\}\_\{z\_\{t\-\\Delta\}\\mid z\_\{t\},w\_\{t\}\}\+\\bar\{\\mathcal\{L\}\}\_\{w\_\{t\-\\Delta\}\\mid z\_\{t\-\\Delta\},w\_\{t\}\}\.The first term iso​\(1\)o\(1\)by[Proposition˜3](https://arxiv.org/html/2608.10615#Thmproposition3), while the second is given by \([116](https://arxiv.org/html/2608.10615#A4.E116)\)\. This yields \([118](https://arxiv.org/html/2608.10615#A4.E118)\)\.

∎

The proposition makes the selection of the main\-text objective precise\. The tight discrete objective is too small in the local limit: after dividing byΔ\\Delta, it vanishes\. The joint objectives behave in the opposite way: they contain a generally nonzeroO​\(1\)O\(1\)simplex\-matching term, so they do not reduce to a finite first\-order training density\. The relaxed discrete bridge sits exactly between these two extremes, which is why it is the natural objective for the continuous\-time formulation\.

## Appendix EOpenWebText Experimental Details

This appendix provides preprocessing, optimization, checkpoint, sampling, and evaluation details that are omitted from the main text\. The dataset, tokenizer, sequence length, backbone, primary optimization settings, and entropy\-matched evaluation protocol are summarized in[Sections˜4\.1](https://arxiv.org/html/2608.10615#S4.SS1)and[4\.3](https://arxiv.org/html/2608.10615#S4.SS3)\.

#### Data preprocessing\.

We use theopenwebtext\-trainandopenwebtext\-validsplits\. Documents are concatenated with an end\-of\-sequence token inserted between adjacent documents and packed into fixed\-length blocks of1,0241\{,\}024GPT\-2 tokens\.

#### Optimization details\.

Models trained in our common codebase use Adam withβ1=0\.9\\beta\_\{1\}=0\.9,β2=0\.999\\beta\_\{2\}=0\.999, numerical constant10−810^\{\-8\}, and gradient\-norm clipping at1\.01\.0\. Training uses bfloat16 precision and a linear warmup over the first2,5002\{,\}500optimizer steps, followed by a constant learning rate\. We maintain an EMA of the parameters with decay0\.99990\.9999and use the averaged parameters for generation\.

#### Simplax checkpoint\.

The Simplax model used in the main OpenWebText comparison is initialized from a UDLM checkpoint trained for800,000800\{,\}000optimizer steps and is subsequently trained with the Simplax objective for an additional200,000200\{,\}000steps\. The resulting checkpoint therefore has a total optimization history of1,000,0001\{,\}000\{,\}000steps\.

The network predicts the clean\-token distribution and receives the categorical state𝐳t\\mathbf\{z\}\_\{t\}as input, while the relaxed state𝐰t\\mathbf\{w\}\_\{t\}remains in the training objective\. We use a uniform time schedule, constant Dirichlet concentrationη=0\.01\\eta=0\.01, and constant loss weighting\. The reported checkpoint does not use an auxiliary self\-conditioning input\.

Dirichlet sampling and concentration\-dependent computations are performed in float64\. Concentration values are restricted to\[10−10,108\]\[10^\{\-10\},10^\{8\}\]for numerical stability\. Before normalization, the output logitsℓ\\ellare softly bounded asℓ←30​tanh⁡\(ℓ/30\)\\ell\\leftarrow 30\\tanh\\left\(\{\\ell\}/\{30\}\\right\)\.

#### Qualitative generations\.

[Tables˜3](https://arxiv.org/html/2608.10615#A5.T3),[4](https://arxiv.org/html/2608.10615#A5.T4)and[5](https://arxiv.org/html/2608.10615#A5.T5)show representative generations at NFE=16=16,128128, and1,0241,024\. For each method, the example is selected from the operating point determined by the entropy\-matching procedure used in the main experiment\. The generated text is not manually rewritten\. The excerpts are truncated at the positions marked by\[…\], and line wrapping is applied only for presentation\.

Table 3:Representative unconditional OpenWebText generations at NFE=16=16\. Entropy is the generative unigram entropy in nats per token\.Table 4:Representative unconditional OpenWebText generations at NFE=128=128\. Entropy is the generative unigram entropy in nats per token\.Table 5:Representative unconditional OpenWebText generations at NFE=1,024=1,024\. Entropy is the generative unigram entropy in nats per token\.
### E\.1Sudoku experimental details

#### Dataset\.

All Sudoku puzzles used in our experiments have unique solutions\. FollowingDeschenaux and Gulcehre \[[2026](https://arxiv.org/html/2608.10615#bib.bib34)\], we use the greedy Sudoku puzzle generator ofAlp \[[2024](https://arxiv.org/html/2608.10615#bib.bib45)\]to construct the training data and the evaluation sets with2525or more clues\. The training set contains48,00048\{,\}000puzzles with3030observed cells, and all models are trained exclusively on this3030\-clue training set\. For evaluation, we use2,0002\{,\}000puzzles for each clue count\. The training set and the4040\-,3535\-,3030\-, and2525\-clue evaluation sets are generated with seed4242\.

For the2020\- and1717\-clue settings, we instead use uniquely solvable1717\-clue Sudoku puzzles studied byLinet al\.\[[2013](https://arxiv.org/html/2608.10615#bib.bib46)\]\. The1717\-clue evaluation set uses these puzzles directly, while the2020\-clue evaluation set is constructed by augmenting each1717\-clue puzzle with three additional clues from its unique solution\.

At evaluation time, we therefore consider conditional completion with4040,3535,3030,2525,2020, and1717clues\. The3030\-clue setting matches the training clue density\. The4040\- and3535\-clue settings provide more conditioning information than observed during training, whereas the2525\-,2020\-, and1717\-clue settings progressively reduce the available conditioning information\. For the no\-clue evaluation, all8181cells in the puzzle prefix are replaced with the blank token\.

#### Sequence representation\.

The vocabulary contains1212symbols: a blank\-cell token, the digits11–99, a row\-separator token, and aBOStoken\. Each example is represented by180180tokens:

\[BOS\]\+89\-token puzzle\+\[BOS\]\+89\-token solution\.\[\\texttt\{BOS\}\]\\;\+\\;\\text\{89\-token puzzle\}\\;\+\\;\[\\texttt\{BOS\}\]\\;\+\\;\\text\{89\-token solution\}\.Each8989\-token board representation contains8181cell tokens and eight row separators\. The resulting puzzle prefix has length9191, and the generated solution has length8989\. Unobserved cells are represented explicitly by the blank token, so the prefix length remains9191for every clue count\. The training objective is evaluated only on the solution portion of the sequence\.

#### Shared architecture\.

All methods use Transformer models with hidden dimension512512, eight Transformer blocks, eight attention heads of dimension6464, and dropout probability0\.10\.1\. The models use learned token embeddings without embedding–output weight tying\. The autoregressive model uses causal attention and no time conditioning\. The remaining models use bidirectional attention and AdaLN\-based time conditioning with conditioning dimension128128\.

#### Optimization\.

All models are trained for20,00020\{,\}000optimization steps with global batch size256256using bfloat16 precision\. We use Adam with learning rate3×10−43\\times 10^\{\-4\},β1=0\.9\\beta\_\{1\}=0\.9,β2=0\.999\\beta\_\{2\}=0\.999, and gradient\-norm clipping at1\.01\.0\. The learning rate is linearly warmed up for2,5002\{,\}500steps and then held constant\. We maintain an EMA of the parameters with decay0\.99990\.9999and use the averaged parameters for evaluation\. Antithetic time sampling is enabled\. All training runs use random seed11, and the reported checkpoints are taken at step20,00020\{,\}000, corresponding to epoch106106\.

#### Qualitative Sudoku generation\.

As shown in Figure[5](https://arxiv.org/html/2608.10615#A5.F5), the baseline models often recover many individual entries while failing to form a globally consistent grid\. MDLM comes closest in this case, differing from the unique solution in only eight cells, but even these sparse errors are sufficient to invalidate the board\. In contrast, Simplax satisfies the coupled row, column, and subgrid constraints simultaneously\. This example was selected from the held\-out2525\-clue evaluation set to illustrate the distinction between local agreement and global validity; aggregate accuracy over the full evaluation sets is reported separately\.

![Refer to caption](https://arxiv.org/html/2608.10615v1/x5.png)Figure 5:Generation results for a representative2525\-clue Sudoku puzzle at NFE=89=89\. Blue entries are given clues, black entries agree with the unique solution, and red entries differ from it\. All methods receive the same puzzle prefix\. Simplax produces the valid solution in this example, whereas the other methods violate at least one Sudoku constraint\.

Similar Articles

Simplex Relaxation for Discrete Diffusion

Hugging Face Daily Papers

Introduces Simplax, an exact Dirichlet-categorical augmentation for uniform discrete diffusion that improves reverse sampling and generative quality on text and Sudoku tasks.

Less Uniform Discrete Diffusion is More Powerful and Scalable

arXiv cs.CL

This paper proposes LUDI, a less uniform diffusion language modeling framework that fixes over-uniform training objectives and condition-target confusion in uniform diffusion LMs, enabling a 7B-scale UDLM with 3x-token-per-step speedup over autoregressive decoding and competitive complex reasoning performance.

Drifting Objectives for Refining Discrete Diffusion Language Models

arXiv cs.CL

This paper introduces TokenDrift, a drifting objective that refines discrete diffusion language models by lifting categorical predictions to a continuous semantic space for anti-symmetric drifting, significantly improving generation quality under a fixed number of denoising steps.