FluxLite:离散扩散模型的推理时提议控制
摘要
FluxLite 提出了一种无需训练、在推理阶段实现的提议控制框架,用于离散扩散模型。该框架通过在 Feynman-Kac 势中引入图散度项来补偿跳跃率扰动,由此得到 HEU 和 D-VCG 两种采样器,在标准 SMC 基线方法的基础上大幅降低了重加权方差与采样误差。
arXiv:2609.35947v1 Announce Type: new
Abstract: Many inference-time tasks for pretrained discrete diffusion models and diffusion language models reduce to drawing samples from a tilted version of the pretrained distribution. Feynman-Kac sequential Monte Carlo (SMC) makes this correction exact in principle, but its prescribed weights routinely degenerate when the proposal dynamics are misaligned with the tilt, capping the practical gains from additional particles. We introduce FluxLite, a lightweight, training-free proposal-control framework for discrete diffusion. On the sparse directed graph of pretrained reverse rates, any sparse jump-rate perturbation can be exactly compensated by a $q_t$-weighted graph-divergence term in the Feynman-Kac potential; the target path is therefore preserved while the residual reweighting variance becomes a local convex objective. We instantiate this principle as two practical samplers: a one-hop local reallocation rule (HEU) and a small nonnegative quadratic program over pretrained-rate bases (D-VCG). We further prove population stability under the standard score-entropy training loss, identifying a tilted-path coverage factor that governs robustness to score error, together with finite-particle convergence for a fixed controlled Feynman-Kac recursion. Empirically, FluxLite improves over standard Feynman-Kac SMC baselines by up to two orders of magnitude in terminal KL on an analytically tractable finite-state CTMC benchmark, and reduces row-correlation MSE on 2D Ising sampling by 5-7x in geometric mean and up to 55x at peak.
查看缓存全文
缓存时间: 2026/09/30 09:43
# FluxLite: Inference-Time Proposal Control for Discrete Diffusion Models
Source: [https://arxiv.org/html/2609.35947](https://arxiv.org/html/2609.35947)
Yinuo RenHaoxuan ChenAffiliation:ICMEAffiliation:Stanford UniversityEmail:[haoxuanc@stanford\.edu](mailto:)Grant M\. RotskoffAffiliation:Department of ChemistryAffiliation:Stanford UniversityEmail:[rotskoff@stanford\.edu](mailto:)Jiequn HanAffiliation:Center for Computational MathematicsAffiliation:Flatiron InstituteEmail:[jhan@flatironinstitute\.org](mailto:)Lexing YingAffiliation:Department of MathematicsAffiliation:Stanford UniversityEmail:[lexing@stanford\.edu](mailto:)
###### Abstract
Many inference\-time tasks for pretrained discrete diffusion models and diffusion language models reduce to drawing samples from a tilted version of the pretrained distribution\. Feynman–Kac sequential Monte Carlo \(SMC\) makes this correction exact in principle, but its prescribed weights routinely degenerate when the proposal dynamics are misaligned with the tilt, capping the practical gains from additional particles\. We introduce*FluxLite*, a lightweight, training\-free proposal\-control framework for discrete diffusion\. On the sparse directed graph of pretrained reverse rates, any sparse jump\-rate perturbation can be exactly compensated by aqtq\_\{t\}\-weighted graph\-divergence term in the Feynman–Kac potential; the target path is therefore preserved while the residual reweighting variance becomes a local convex objective\. We instantiate this principle as two practical samplers: a one\-hop local reallocation rule \(HEU\) and a small nonnegative quadratic program over pretrained\-rate bases \(D\-VCG\)\. We further prove population stability under the standard score\-entropy training loss, identifying a tilted\-path coverage factor that governs robustness to score error, together with finite\-particle convergence for a fixed controlled Feynman–Kac recursion\. Empirically, FluxLite improves over standard Feynman–Kac SMC baselines by up to two orders of magnitude in terminalKL\\KLon an analytically tractable finite\-state CTMC benchmark, and reduces row\-correlation MSE on 2D Ising sampling by55–7×7\\timesin geometric mean and up to55×55\\timesat peak\.
## 1Introduction
Many inference\-time tasks in modern generative modeling reduce to sampling from a*tilted*version of a pretrained model distribution: reward\-aligned generation, classifier\-free guidance, posterior sampling under measurement constraints, controllable text generation, and chain\-of\-thought reasoning\. The pattern recurs across autoregressive large language models \(LLMs\), continuous diffusion models, and discrete diffusion models including Diffusion Language Models \(DLMs\)\. It is also part of the broader shift from training\-time scaling laws\[[74](https://arxiv.org/html/2609.35947#bib.bib74),[65](https://arxiv.org/html/2609.35947#bib.bib65)\]to*inference\-time scaling*, in which additional compute at deployment is allocated to selection, verification, search, or sampling rather than to enlarging the base model\[[104](https://arxiv.org/html/2609.35947#bib.bib104),[90](https://arxiv.org/html/2609.35947#bib.bib90),[179](https://arxiv.org/html/2609.35947#bib.bib179),[152](https://arxiv.org/html/2609.35947#bib.bib152),[101](https://arxiv.org/html/2609.35947#bib.bib101),[25](https://arxiv.org/html/2609.35947#bib.bib25),[119](https://arxiv.org/html/2609.35947#bib.bib119),[59](https://arxiv.org/html/2609.35947#bib.bib59)\]\. Appendix[A](https://arxiv.org/html/2609.35947#A1)collects the model\-specific and domain\-specific literature\.
Existing inference\-time methods divide into*length scaling*\(longer or iteratively refined trajectories\) and*width scaling*\(exploring multiple candidate hypotheses in parallel\)\. Sampling\-based width scaling typically targets a tilted law through Feynman–Kac sequential Monte Carlo \(SMC\), in which particles are propagated under proposed dynamics and reweighted by an incremental potential prescribed by the path\[[133](https://arxiv.org/html/2609.35947#bib.bib133),[25](https://arxiv.org/html/2609.35947#bib.bib25),[134](https://arxiv.org/html/2609.35947#bib.bib134),[83](https://arxiv.org/html/2609.35947#bib.bib83),[59](https://arxiv.org/html/2609.35947#bib.bib59)\]\. Its central bottleneck is*weight degeneracy*: the proposal and the incremental weights are tied together by the Feynman–Kac representation, so when they are misaligned, weights concentrate on a vanishing fraction of particles and additional particles do not translate into additional information, eroding gains over Best\-of\-NNor related selection baselines\[[91](https://arxiv.org/html/2609.35947#bib.bib91),[54](https://arxiv.org/html/2609.35947#bib.bib54)\]\. Twisted and controlled SMC in statistics\[[16](https://arxiv.org/html/2609.35947#bib.bib16),[161](https://arxiv.org/html/2609.35947#bib.bib161),[63](https://arxiv.org/html/2609.35947#bib.bib63)\]suggests that better proposals can mitigate this, and recent inference\-time methods for diffusion and LLMs pursue related ideas\[[183](https://arxiv.org/html/2609.35947#bib.bib183),[50](https://arxiv.org/html/2609.35947#bib.bib50),[4](https://arxiv.org/html/2609.35947#bib.bib4),[66](https://arxiv.org/html/2609.35947#bib.bib66),[109](https://arxiv.org/html/2609.35947#bib.bib109)\], but typically at the cost of extra training or modality\-specific heuristics\.
This raises a basic question:
Is there a single training\-free proposal\-control principle that works across continuous and discrete inference\-time SMC?
DriftLite\[[119](https://arxiv.org/html/2609.35947#bib.bib119)\]answers this question for*continuous*diffusion models by exploiting the non\-uniqueness of the Feynman–Kac PDE: probability transport can be shifted between a control drift and a residual reweighting potential through aqtq\_\{t\}\-weighted divergence, and a basis\-restricted variance objective selects a representative without retraining\. The discrete case is structurally different: instead of a vector field onℝd\\mathbb\{R\}^\{d\}, the proposal is a sparse directed graph of jump rates fixed by the pretrained reverse process, and the divergence operator must respect this sparsity pattern\. We show that the tilted continuous\-time\-Markov\-chain \(CTMC\) Feynman–Kac flow nevertheless admits an analogous equivalence class: sparse jump\-rate perturbations are exactly compensated by a graph\-divergence correction to the potential\. Proposal design therefore becomes a sparse, training\-free flux\-reallocation problem with the same variance\-control objective as in DriftLite\.
### 1\.1Contributions
- •Discrete proposal\-control framework\.We introduce*FluxLite*, the discrete analogue of DriftLite\. It exposes an exact equivalence class of CTMC–Feynman–Kac representations on the pretrained sparsity graph and recasts proposal control as sparse graph\-flux variance minimization \(Table[1](https://arxiv.org/html/2609.35947#S3.T1)\)\.
- •Theory\.We prove \(i\) population stability of the guided Feynman–Kac flow under the standard score\-entropy training loss, identifying a tilted\-path coverage factor that quantifies robustness to score error, and \(ii\) finite\-particle convergence for the controlled SMC recursion\. Together these results make explicit how score quality, tilted\-path coverage, residual\-weight oscillation, and particle count enter inference\-time control\.
- •Practical samplers and experiments\.We instantiate FluxLite as a one\-hop local heuristic \(HEU\) and a pretrained\-rate basis QP,*Discrete Variance\-Controlling Guidance*\(D\-VCG\)\. On analytically tractable finite\-state CTMCs,D\-VCGcuts terminalKL\\KLover the Discrete Feynman–Kac Corrector \(D\-FKC\)\[[59](https://arxiv.org/html/2609.35947#bib.bib59)\]baseline by up to114\.8×114\.8\\times; on 2D Ising sampling it reduces row\-correlation MSE overD\-FKCby55–7×7\\timesin geometric mean and up to55×55\\timesat peak\.
## 2Preliminaries
We review the Feynman–Kac view of inference\-time scaling in a form that makes the continuous–discrete correspondence explicit\. Section[2\.1](https://arxiv.org/html/2609.35947#S2.SS1)recalls the DriftLite identity for continuous diffusion models; Section[2\.2](https://arxiv.org/html/2609.35947#S2.SS2)gives the discrete CTMC–Feynman–Kac flow that FluxLite controls\.
### 2\.1Continuous Feynman–Kac control
Let𝒳=ℝd\\mathcal\{X\}=\\mathbb\{R\}^\{d\}and let the backward process of a pretrained diffusion model have base drift𝐯t←\\mathbf\{v\}\_\{t\}^\{\\shortleftarrow\}, diffusion coefficientαt←\\alpha\_\{t\}^\{\\shortleftarrow\}, and marginal densitypt←p\_\{t\}^\{\\shortleftarrow\}at reverse timet∈\[0,T\]t\\in\[0,T\]\. Many inference\-time objectives define a tilted pathqt\(x\)∝\(pt←\(x\)\)γert\(x\)q\_\{t\}\(x\)\\propto\(p\_\{t\}^\{\\shortleftarrow\}\(x\)\)^\{\\gamma\}e^\{r\_\{t\}\(x\)\}, whereγ\>0\\gamma\>0andrtr\_\{t\}is a reward, likelihood, or guidance potential\. For a chosen Feynman–Kac representative of this tilted path, write the proposal drift as𝐯~t\\widetilde\{\\mathbf\{v\}\}\_\{t\}\. The normalized Feynman–Kac equation can be written compactly as
∂tqt=−∇⋅\(𝐯~tqt\)\+αt←22Δqt\+gtqt,𝔼qt\[gt\]=0\.\\partial\_\{t\}q\_\{t\}=\-\\nabla\\\!\\cdot\(\\widetilde\{\\mathbf\{v\}\}\_\{t\}q\_\{t\}\)\+\\frac\{\\alpha\_\{t\}^\{\\shortleftarrow\}\{\}^\{2\}\}\{2\}\\Delta q\_\{t\}\+g\_\{t\}q\_\{t\},\\qquad\\mathbb\{E\}\_\{q\_\{t\}\}\[g\_\{t\}\]=0\.\(2\.1\)DriftLite\[[119](https://arxiv.org/html/2609.35947#bib.bib119)\]observes that this representation is not unique\. For any control drift𝐮t:ℝd→ℝd\\mathbf\{u\}\_\{t\}:\\mathbb\{R\}^\{d\}\\to\\mathbb\{R\}^\{d\},
−∇⋅\(𝐯~tqt\)\+gtqt=−∇⋅\(\(𝐯~t\+𝐮t\)qt\)\+\(gt\+divqt𝐮t\)qt,divqt𝐮t:=qt−1∇⋅\(qt𝐮t\)\.\-\\nabla\\\!\\cdot\(\\widetilde\{\\mathbf\{v\}\}\_\{t\}q\_\{t\}\)\+g\_\{t\}q\_\{t\}=\-\\nabla\\\!\\cdot\(\(\\widetilde\{\\mathbf\{v\}\}\_\{t\}\+\\mathbf\{u\}\_\{t\}\)q\_\{t\}\)\+\\big\(g\_\{t\}\+\\operatorname\{div\}\_\{q\_\{t\}\}\\mathbf\{u\}\_\{t\}\\big\)q\_\{t\},\\quad\\operatorname\{div\}\_\{q\_\{t\}\}\\mathbf\{u\}\_\{t\}:=q\_\{t\}^\{\-1\}\\nabla\\\!\\cdot\(q\_\{t\}\\mathbf\{u\}\_\{t\}\)\.\(2\.2\)Thus the sameqtq\_\{t\}can be represented by many proposal drifts and residual potentials\. DriftLite chooses a representative by solving, or approximating,min𝐮tVarqt\[gt\+divqt𝐮t\]\\min\_\{\\mathbf\{u\}\_\{t\}\}\\operatorname\{Var\}\_\{q\_\{t\}\}\[g\_\{t\}\+\\operatorname\{div\}\_\{q\_\{t\}\}\\mathbf\{u\}\_\{t\}\]\. The discrete construction below replaces the continuous divergence in \([2\.2](https://arxiv.org/html/2609.35947#S2.E2)\) by aqtq\_\{t\}\-weighted graph divergence and replaces drift control by sparse jump\-rate control\. The detailed continuous setup is recalled in Appendix[B](https://arxiv.org/html/2609.35947#A2)\.
### 2\.2Discrete Feynman–Kac flow
Let𝒳\\mathcal\{X\}be a finite state space withD:=\|𝒳\|<∞D:=\|\\mathcal\{X\}\|<\\infty\. We use the column convention:Qt\(y,x\)Q\_\{t\}\(y,x\)denotes the rate from sourcexxto destinationyy, so a CTMC marginalptp\_\{t\}evolves as∂tpt\(x\)=∑y≠x\(Qt\(x,y\)pt\(y\)−Qt\(y,x\)pt\(x\)\)\\partial\_\{t\}p\_\{t\}\(x\)=\\sum\_\{y\\neq x\}\\big\(Q\_\{t\}\(x,y\)p\_\{t\}\(y\)\-Q\_\{t\}\(y,x\)p\_\{t\}\(x\)\\big\)\. Generator\-matrix notation and all proofs for Sections[2](https://arxiv.org/html/2609.35947#S2)–[3](https://arxiv.org/html/2609.35947#S3)are in Appendix[C](https://arxiv.org/html/2609.35947#A3)\. Letpt→p\_\{t\}^\{\\shortrightarrow\}be the forward noising marginal andpt←:=pT−t→p\_\{t\}^\{\\shortleftarrow\}:=p\_\{T\-t\}^\{\\shortrightarrow\}the reverse\-time marginal\. When the forward noising rates are known, exact time reversal gives the local\-ratio form
Qt←\(y,x\)=QT−t→\(x,y\)pt←\(y\)pt←\(x\):=QT−t→\(x,y\)st\(x,y\),x≠y\.Q\_\{t\}^\{\\shortleftarrow\}\(y,x\)=Q\_\{T\-t\}^\{\\shortrightarrow\}\(x,y\)\\frac\{p\_\{t\}^\{\\shortleftarrow\}\(y\)\}\{p\_\{t\}^\{\\shortleftarrow\}\(x\)\}:=Q\_\{T\-t\}^\{\\shortrightarrow\}\(x,y\)s\_\{t\}\(x,y\),\\qquad x\\neq y\.\(2\.3\)A learned discrete diffusion model replacessts\_\{t\}by an estimates^t\\widehat\{s\}\_\{t\}, yieldingQ^t←\(y,x\)=QT−t→\(x,y\)s^t\(x,y\)\\widehat\{Q\}\_\{t\}^\{\\shortleftarrow\}\(y,x\)=Q\_\{T\-t\}^\{\\shortrightarrow\}\(x,y\)\\widehat\{s\}\_\{t\}\(x,y\)\. The reverse marginal satisfies
∂tpt←\(x\)=∑y≠x\(Qt←\(x,y\)pt←\(y\)−Qt←\(y,x\)pt←\(x\)\)\.\\partial\_\{t\}p\_\{t\}^\{\\shortleftarrow\}\(x\)=\\sum\_\{y\\neq x\}\\big\(Q\_\{t\}^\{\\shortleftarrow\}\(x,y\)p\_\{t\}^\{\\shortleftarrow\}\(y\)\-Q\_\{t\}^\{\\shortleftarrow\}\(y,x\)p\_\{t\}^\{\\shortleftarrow\}\(x\)\\big\)\.\(2\.4\)For a reward, likelihood, or guidance potentialrtr\_\{t\}, we target the normalized tilted pathqt\(x\)∝\(pt←\(x\)\)γert\(x\)q\_\{t\}\(x\)\\propto\(p\_\{t\}^\{\\shortleftarrow\}\(x\)\)^\{\\gamma\}e^\{r\_\{t\}\(x\)\}withγ\>0\\gamma\>0\. A typical realization, used in Section[5](https://arxiv.org/html/2609.35947#S5), is the linear ramprt\(x\)=βtr\(x\)r\_\{t\}\(x\)=\\beta\_\{t\}\\,r\(x\)for a target rewardr:𝒳→ℝr:\\mathcal\{X\}\\to\\mathbb\{R\}and a scalar scheduleβt≥0\\beta\_\{t\}\\geq 0withβ0=0\\beta\_\{0\}=0at the noisy end \(so the tilt is inactive at initialization\) andβT=1\\beta\_\{T\}=1at the data end \(sorT=rr\_\{T\}=rrecovers the target\)\.
###### Proposition 2\.1\(Discrete Feynman–Kac representative for a tilted path\[[59](https://arxiv.org/html/2609.35947#bib.bib59)\]\)\.
Assumept←\(x\)\>0p\_\{t\}^\{\\shortleftarrow\}\(x\)\>0for allx∈𝒳x\\in\\mathcal\{X\}andt∈\[0,T\]t\\in\[0,T\]\. Fixγ\>0\\gamma\>0and a differentiable potentialrt:𝒳→ℝr\_\{t\}:\\mathcal\{X\}\\to\\mathbb\{R\}\. The normalized pathqt\(x\)∝\(pt←\(x\)\)γert\(x\)q\_\{t\}\(x\)\\propto\(p\_\{t\}^\{\\shortleftarrow\}\(x\)\)^\{\\gamma\}e^\{r\_\{t\}\(x\)\}satisfies
∂tqt\(x\)=∑y≠x\(Q~t\(x,y\)qt\(y\)−Q~t\(y,x\)qt\(x\)\)\+g~t\(x\)qt\(x\),g~t=G~t−𝔼qtG~t,\\partial\_\{t\}q\_\{t\}\(x\)=\\sum\_\{y\\neq x\}\\big\(\\widetilde\{Q\}\_\{t\}\(x,y\)q\_\{t\}\(y\)\-\\widetilde\{Q\}\_\{t\}\(y,x\)q\_\{t\}\(x\)\\big\)\+\\widetilde\{g\}\_\{t\}\(x\)q\_\{t\}\(x\),\\qquad\\widetilde\{g\}\_\{t\}=\\widetilde\{G\}\_\{t\}\-\\mathbb\{E\}\_\{q\_\{t\}\}\\widetilde\{G\}\_\{t\},\(2\.5\)where, forx≠yx\\neq y, the guided jump rate and reweighting potential are given by
Q~t\(x,y\)=γQt←\(x,y\)\(pt←\(y\)pt←\(x\)\)1−γert\(x\)−rt\(y\),G~t\(x\)=r˙t\(x\)\+∑y≠x\(Q~t\(y,x\)−γQt←\(y,x\)\)\.\\widetilde\{Q\}\_\{t\}\(x,y\)=\\gamma Q\_\{t\}^\{\\shortleftarrow\}\(x,y\)\\left\(\\frac\{p\_\{t\}^\{\\shortleftarrow\}\(y\)\}\{p\_\{t\}^\{\\shortleftarrow\}\(x\)\}\\right\)^\{1\-\\gamma\}e^\{r\_\{t\}\(x\)\-r\_\{t\}\(y\)\},\\ \\widetilde\{G\}\_\{t\}\(x\)=\\dot\{r\}\_\{t\}\(x\)\+\\sum\_\{y\\neq x\}\(\\widetilde\{Q\}\_\{t\}\(y,x\)\-\\gamma Q\_\{t\}^\{\\shortleftarrow\}\(y,x\)\)\.\(2\.6\)
Proposition[2\.1](https://arxiv.org/html/2609.35947#S2.Thmtheorem1)gives the baseline Feynman–Kac representative that FluxLite later changes within an exact equivalence class; common reward\-tilting and annealing reductions are listed in Appendix[C\.2](https://arxiv.org/html/2609.35947#A3.SS2)\. It also reveals that𝐐~t\\widetilde\{\\mathbf\{Q\}\}\_\{t\}has the same directed support as𝐐t←\\mathbf\{Q\}\_\{t\}^\{\\shortleftarrow\}: the tilt changes edge weights but introduces no new graph edges\. A baseline Feynman–Kac Sequential Monte Carlo \(SMC\) sampler simulates particles under𝐐~t\\widetilde\{\\mathbf\{Q\}\}\_\{t\}, reweights them byg~t\\widetilde\{g\}\_\{t\}, and returns a terminal weighted empirical law\.
## 3Methodology
We now develop*FluxLite*\. Starting from a generic Feynman–Kac representation, we show that the same target path can be represented by many pairs of jump dynamics and reweighting potentials\. FluxLite chooses a representative with smaller residual reweighting variance while respecting the sparsity pattern of the backward process\. We instantiate this principle as a local rule,HEU\(Section[3\.3\.1](https://arxiv.org/html/2609.35947#S3.SS3.SSS1)\), and a basis\-restricted QP,Discrete Variance\-Controlling Guidance\(D\-VCG, Section[3\.3\.2](https://arxiv.org/html/2609.35947#S3.SS3.SSS2)\)\.
### 3\.1Graph calculus
The basic object is a divergence operator on a directed graph\. Throughout this section, all identities are stated on the support ofqtq\_\{t\}; for readability we assumeqt\(x\)\>0q\_\{t\}\(x\)\>0for everyx∈𝒳x\\in\\mathcal\{X\}at the time under consideration\. With the column convention, a perturbationRt\(y,x\)R\_\{t\}\(y,x\)adds outflow fromxxtoyy\. We therefore define the graph divergence as signed net outflow per unit mass:
divqt𝐑t\(x\):=1qt\(x\)∑y≠x\(Rt\(y,x\)qt\(x\)−Rt\(x,y\)qt\(y\)\)\.\\operatorname\{div\}\_\{q\_\{t\}\}\\mathbf\{R\}\_\{t\}\(x\):=\\frac\{1\}\{q\_\{t\}\(x\)\}\\sum\_\{y\\neq x\}\\big\(R\_\{t\}\(y,x\)q\_\{t\}\(x\)\-R\_\{t\}\(x,y\)q\_\{t\}\(y\)\\big\)\.\(3\.1\)
Consider a target path\(qt\)t∈\[0,T\]\(q\_\{t\}\)\_\{t\\in\[0,T\]\}satisfying a generic Feynman–Kac CTMC equation
∂tqt\(x\)=∑y≠x\(Qt\(x,y\)qt\(y\)−Qt\(y,x\)qt\(x\)\)\+qt\(x\)gt\(x\),\\partial\_\{t\}q\_\{t\}\(x\)=\\sum\_\{y\\neq x\}\\Big\(Q\_\{t\}\(x,y\)q\_\{t\}\(y\)\-Q\_\{t\}\(y,x\)q\_\{t\}\(x\)\\Big\)\+q\_\{t\}\(x\)\\,g\_\{t\}\(x\),\(3\.2\)where𝐐t\\mathbf\{Q\}\_\{t\}is an off\-diagonal rate family andgtg\_\{t\}is centered underqtq\_\{t\}, i\.e\.𝔼qt\[gt\]=0\\mathbb\{E\}\_\{q\_\{t\}\}\[g\_\{t\}\]=0\. Proposition[2\.1](https://arxiv.org/html/2609.35947#S2.Thmtheorem1)gives one concrete representative,\(𝐐~t,g~t\)\(\\widetilde\{\\mathbf\{Q\}\}\_\{t\},\\widetilde\{g\}\_\{t\}\), for the tilted path\.
##### Flux equivalence\.
The Feynman–Kac representation \([3\.2](https://arxiv.org/html/2609.35947#S3.E2)\) is invariant under a family of perturbations: any added jump flux can be compensated by an appropriate divergence term in the potential\.
###### Proposition 3\.1\(Equivalence class\)\.
Assume the path\(qt\)t∈\[0,T\]\(q\_\{t\}\)\_\{t\\in\[0,T\]\}satisfies \([3\.2](https://arxiv.org/html/2609.35947#S3.E2)\) andqt\(x\)\>0q\_\{t\}\(x\)\>0for all states under consideration\. Let𝐑t\\mathbf\{R\}\_\{t\}be any matrix such thatRt\(x,x\)=0R\_\{t\}\(x,x\)=0and
Qt′\(y,x\):=Qt\(y,x\)\+Rt\(y,x\)≥0,for allx≠y\.Q\_\{t\}^\{\\prime\}\(y,x\):=Q\_\{t\}\(y,x\)\+R\_\{t\}\(y,x\)\\geq 0,\\qquad\\text\{for all \}x\\neq y\.Then the same path\(qt\)t∈\[0,T\]\(q\_\{t\}\)\_\{t\\in\[0,T\]\}also satisfies
∂tqt\(x\)=∑y≠x\(Qt′\(x,y\)qt\(y\)−Qt′\(y,x\)qt\(x\)\)\+qt\(x\)gt′\(x\),\\partial\_\{t\}q\_\{t\}\(x\)=\\sum\_\{y\\neq x\}\\Big\(Q\_\{t\}^\{\\prime\}\(x,y\)q\_\{t\}\(y\)\-Q\_\{t\}^\{\\prime\}\(y,x\)q\_\{t\}\(x\)\\Big\)\+q\_\{t\}\(x\)\\,g\_\{t\}^\{\\prime\}\(x\),\(3\.3\)wheregt′\(x\)=gt\(x\)\+divqt𝐑t\(x\)g\_\{t\}^\{\\prime\}\(x\)=g\_\{t\}\(x\)\+\\operatorname\{div\}\_\{q\_\{t\}\}\\mathbf\{R\}\_\{t\}\(x\)\. Moreover, if𝔼qt\[gt\]=0\\mathbb\{E\}\_\{q\_\{t\}\}\[g\_\{t\}\]=0, then also𝔼qt\[gt′\]=0\\mathbb\{E\}\_\{q\_\{t\}\}\[g\_\{t\}^\{\\prime\}\]=0\.
Thus each Feynman–Kac representation belongs to an equivalence class\[\(𝐐t,gt\)\]\[\(\\mathbf\{Q\}\_\{t\},g\_\{t\}\)\]indexed by all admissible flux reallocations\. The point of proposal control is to exploit this non\-uniqueness: rather than accepting the baseline representation delivered by the guided Feynman–Kac pair\[\(𝐐~t,g~t\)\]\[\(\\widetilde\{\\mathbf\{Q\}\}\_\{t\},\\widetilde\{g\}\_\{t\}\)\], we search inside the equivalence class for a representative with smaller reweighting variance\. It is instructive to compare Proposition[3\.1](https://arxiv.org/html/2609.35947#S3.Thmtheorem1)with its continuous counterpart for diffusion models, restated as Proposition[B\.1](https://arxiv.org/html/2609.35947#A2.Thmtheorem1)in Appendix[B](https://arxiv.org/html/2609.35947#A2)\. As shown in Table[1](https://arxiv.org/html/2609.35947#S3.T1), each ingredient of the continuous theory has a well\-defined discrete counterpart; we use these connections to improve inference\-time proposal control in the following sections\.
Table 1:Continuous vs\. discrete variance\-control\.Continuous diffusion \(𝒳=ℝd\\mathcal\{X\}=\\mathbb\{R\}^\{d\}\)Discrete CTMC \(\|𝒳\|=D\|\\mathcal\{X\}\|=D\)Generatorℒ~t∗ρ=−∇⋅\(𝐯~tρ\)\+αt←22Δρ\\widetilde\{\\mathcal\{L\}\}\_\{t\}^\{\\ast\}\\rho=\-\\nabla\\cdot\\big\(\\widetilde\{\\mathbf\{v\}\}\_\{t\}\\rho\\big\)\+\\frac\{\\alpha\_\{t\}^\{\\shortleftarrow\}\{\}^\{2\}\}\{2\}\\Delta\\rho\(ℒ~t∗ρ\)\(x\)=∑y≠x\(Q~t\(x,y\)ρ\(y\)−Q~t\(y,x\)ρ\(x\)\)\(\\widetilde\{\\mathcal\{L\}\}\_\{t\}^\{\\ast\}\\rho\)\(x\)=\\sum\_\{y\\neq x\}\\big\(\\widetilde\{Q\}\_\{t\}\(x,y\)\\rho\(y\)\-\\widetilde\{Q\}\_\{t\}\(y,x\)\\rho\(x\)\\big\)Feynman–Kac form∂tqt=ℒ~t∗qt\+gtqt,𝔼qt\[gt\]=0\\partial\_\{t\}q\_\{t\}=\\widetilde\{\\mathcal\{L\}\}\_\{t\}^\{\\ast\}q\_\{t\}\+g\_\{t\}\\,q\_\{t\},\\;\\;\\mathbb\{E\}\_\{q\_\{t\}\}\[g\_\{t\}\]=0Controldrift perturbation𝐮t:ℝd→ℝd\\mathbf\{u\}\_\{t\}:\\mathbb\{R\}^\{d\}\\to\\mathbb\{R\}^\{d\}rate perturbation𝐑t∈ℝD×D\\mathbf\{R\}\_\{t\}\\in\\mathbb\{R\}^\{D\\times D\}Divergencedivqt𝐮t\(x\):=∇⋅\(qt\(x\)𝐮t\(x\)\)qt\(x\)\\operatorname\{div\}\_\{q\_\{t\}\}\\mathbf\{u\}\_\{t\}\(x\):=\\frac\{\\nabla\\cdot\(q\_\{t\}\(x\)\\mathbf\{u\}\_\{t\}\(x\)\)\}\{q\_\{t\}\(x\)\}divqt𝐑t\(x\):=∑y≠x\(Rt\(y,x\)qt\(x\)−Rt\(x,y\)qt\(y\)\)qt\(x\)\\operatorname\{div\}\_\{q\_\{t\}\}\\mathbf\{R\}\_\{t\}\(x\):=\\frac\{\\sum\_\{y\\neq x\}\\big\(R\_\{t\}\(y,x\)q\_\{t\}\(x\)\-R\_\{t\}\(x,y\)q\_\{t\}\(y\)\\big\)\}\{q\_\{t\}\(x\)\}Equivalence\(𝐯~t,gt\)∼\(𝐯~t\+𝐮t,gt\+divqt𝐮t\)\(\\widetilde\{\\mathbf\{v\}\}\_\{t\},g\_\{t\}\)\\sim\(\\widetilde\{\\mathbf\{v\}\}\_\{t\}\+\\mathbf\{u\}\_\{t\},\\;g\_\{t\}\+\\operatorname\{div\}\_\{q\_\{t\}\}\\mathbf\{u\}\_\{t\}\)\(𝐐~t,gt\)∼\(𝐐~t\+𝐑t,gt\+divqt𝐑t\)\(\\widetilde\{\\mathbf\{Q\}\}\_\{t\},g\_\{t\}\)\\sim\(\\widetilde\{\\mathbf\{Q\}\}\_\{t\}\+\\mathbf\{R\}\_\{t\},\\;g\_\{t\}\+\\operatorname\{div\}\_\{q\_\{t\}\}\\mathbf\{R\}\_\{t\}\)Optimalitymin𝐮tVarqt\[gt\+divqt𝐮t\]\\min\_\{\\mathbf\{u\}\_\{t\}\}\\operatorname\{Var\}\_\{q\_\{t\}\}\\\!\\left\[g\_\{t\}\+\\operatorname\{div\}\_\{q\_\{t\}\}\\mathbf\{u\}\_\{t\}\\right\]min𝐑tVarqt\[gt\+divqt𝐑t\]\\min\_\{\\mathbf\{R\}\_\{t\}\}\\operatorname\{Var\}\_\{q\_\{t\}\}\\\!\\left\[g\_\{t\}\+\\operatorname\{div\}\_\{q\_\{t\}\}\\mathbf\{R\}\_\{t\}\\right\]
### 3\.2Variance control
Following the same logic as in the continuous case\[[119](https://arxiv.org/html/2609.35947#bib.bib119)\], once a Feynman–Kac representation is fixed, the divergence identity of Proposition[3\.1](https://arxiv.org/html/2609.35947#S3.Thmtheorem1)allows us to reallocate flux from the residual potential into the proposal dynamics\. The ideal objective is to eliminate reweighting altogether\.
A useful reference point is obtained by turning off the jump dynamics and absorbing everything into the potential\.
###### Proposition 3\.2\(PR: Pure reweighting\)\.
Assumeqt\(x\)\>0q\_\{t\}\(x\)\>0for allx∈𝒳x\\in\\mathcal\{X\}\. Setting𝐐t′≡0\\mathbf\{Q\}\_\{t\}^\{\\prime\}\\equiv 0in Proposition[3\.1](https://arxiv.org/html/2609.35947#S3.Thmtheorem1)yields the*pure\-reweighting*representative
∂tqt\(x\)=qt\(x\)gt0\(x\),gt0\(x\):=gt\(x\)\+1qt\(x\)∑y≠x\(Qt\(x,y\)qt\(y\)−Qt\(y,x\)qt\(x\)\)\.\\partial\_\{t\}q\_\{t\}\(x\)=q\_\{t\}\(x\)\\,g\_\{t\}^\{0\}\(x\),\\quad g\_\{t\}^\{0\}\(x\):=g\_\{t\}\(x\)\+\\frac\{1\}\{q\_\{t\}\(x\)\}\\sum\_\{y\\neq x\}\\Big\(Q\_\{t\}\(x,y\)q\_\{t\}\(y\)\-Q\_\{t\}\(y,x\)q\_\{t\}\(x\)\\Big\)\.\(3\.4\)Moreover,\[\(𝐐t,gt\)\]=\[\(𝟎,gt0\)\]\[\(\\mathbf\{Q\}\_\{t\},g\_\{t\}\)\]=\[\(\\mathbf\{0\},g\_\{t\}^\{0\}\)\], and every representative in this equivalence class can be reconstructed from\(𝟎,gt0\)\(\\mathbf\{0\},g\_\{t\}^\{0\}\): for any candidate off\-diagonal rate family𝐐t′\\mathbf\{Q\}\_\{t\}^\{\\prime\}, the associated potential is
gt′\(x\)=gt0\(x\)\+divqt𝐐t′\(x\)\.g\_\{t\}^\{\\prime\}\(x\)=g\_\{t\}^\{0\}\(x\)\+\\operatorname\{div\}\_\{q\_\{t\}\}\\mathbf\{Q\}\_\{t\}^\{\\prime\}\(x\)\.\(3\.5\)
The proof is in Appendix[C\.5](https://arxiv.org/html/2609.35947#A3.SS5)\. This representative is typically uncompetitive for simulation because all changes in probability mass are represented through importance weights, but it is algebraically convenient: it serves as the canonical anchor of the equivalence class from which every other representative is obtained via \([3\.5](https://arxiv.org/html/2609.35947#S3.E5)\)\.
##### Variance objective\.
We therefore seek a representative\(𝐐t′,gt′\)∈\[\(𝐐t,gt\)\]=\[\(𝟎,gt0\)\]\(\\mathbf\{Q\}\_\{t\}^\{\\prime\},g\_\{t\}^\{\\prime\}\)\\in\[\(\\mathbf\{Q\}\_\{t\},g\_\{t\}\)\]=\[\(\\mathbf\{0\},g\_\{t\}^\{0\}\)\]that reduces weight degeneracy by shrinking the variance of the residual centered potential:
min𝐐t′𝒱\[𝐐t′\]:=Varx∼qt\[gt′\(x\)\]=Varx∼qt\[gt0\(x\)\+divqt𝐐t′\(x\)\]\.\\min\_\{\\mathbf\{Q\}\_\{t\}^\{\\prime\}\}\\;\\mathcal\{V\}\[\\mathbf\{Q\}\_\{t\}^\{\\prime\}\]:=\\operatorname\{Var\}\_\{x\\sim q\_\{t\}\}\\\!\\big\[g\_\{t\}^\{\\prime\}\(x\)\\big\]=\\operatorname\{Var\}\_\{x\\sim q\_\{t\}\}\\\!\\big\[g\_\{t\}^\{0\}\(x\)\+\\operatorname\{div\}\_\{q\_\{t\}\}\\mathbf\{Q\}\_\{t\}^\{\\prime\}\(x\)\\big\]\.\(3\.6\)If the minimum equals zero, thengt′\(x\)≡0g\_\{t\}^\{\\prime\}\(x\)\\equiv 0and the target path is realized by a pure CTMC without any reweighting\.
Unlike the continuous case, where the exact solution of the variance minimization requires solving a Poisson equation \(*cf\.*Proposition[B\.2](https://arxiv.org/html/2609.35947#A2.Thmtheorem2)in Appendix[B](https://arxiv.org/html/2609.35947#A2)\), the zero\-variance objective has an explicit solution in the discrete case when sparsity constraints are absent\.
###### Proposition 3\.3\(DEN: Unconstrained zero\-variance reallocation\)\.
Assumeqt\(x\)\>0q\_\{t\}\(x\)\>0for allx∈𝒳x\\in\\mathcal\{X\}and thatgtg\_\{t\}in \([3\.2](https://arxiv.org/html/2609.35947#S3.E2)\) is centered\. LetD:=\|𝒳\|D:=\|\\mathcal\{X\}\|\. Define, forx≠yx\\neq y, the dense rate family
Qt∗\(y,x\):=1D\[qt\(y\)qt\(x\)gt0\(y\)−gt0\(x\)\]\+,Qt∗\(x,x\):=0\.Q\_\{t\}^\{\*\}\(y,x\):=\\tfrac\{1\}\{D\}\\big\[\\tfrac\{q\_\{t\}\(y\)\}\{q\_\{t\}\(x\)\}g\_\{t\}^\{0\}\(y\)\-g\_\{t\}^\{0\}\(x\)\\big\]\_\{\+\},\\qquad Q\_\{t\}^\{\*\}\(x,x\):=0\.\(3\.7\)Thengt0\(x\)\+divqt𝐐t∗\(x\)=0g\_\{t\}^\{0\}\(x\)\+\\operatorname\{div\}\_\{q\_\{t\}\}\\mathbf\{Q\}\_\{t\}^\{\*\}\(x\)=0for everyxx, so the target path is realized by the pure CTMC∂tqt\(x\)=∑y≠x\(Qt∗\(x,y\)qt\(y\)−Qt∗\(y,x\)qt\(x\)\)\\partial\_\{t\}q\_\{t\}\(x\)=\\sum\_\{y\\neq x\}\\big\(Q\_\{t\}^\{\*\}\(x,y\)q\_\{t\}\(y\)\-Q\_\{t\}^\{\*\}\(y,x\)q\_\{t\}\(x\)\\big\)without any reweighting\.
Proposition[3\.3](https://arxiv.org/html/2609.35947#S3.Thmtheorem3)\(proof in Appendix[C\.6](https://arxiv.org/html/2609.35947#A3.SS6)\) identifies the ideal object but not a practical algorithm: the construction is dense and typically reallocates flux on the complete directed graph, whereas in discrete diffusion or DLMs the allowed jump graph is sparse and fixed by the pretrained reverse rate family\. We therefore restrict attention to controlled rate families that preserve a prescribed sparsity pattern\.
##### Sparse rates\.
Let𝒮sp⊆𝒳×𝒳\\mathcal\{S\}^\{\\mathrm\{sp\}\}\\subseteq\\mathcal\{X\}\\times\\mathcal\{X\}be the directed edge set of allowed transitions at timett\. Any implementable off\-diagonal rate family𝐐t′\\mathbf\{Q\}\_\{t\}^\{\\prime\}must satisfy
Qt′\(y,x\)=0,for\(y,x\)∉𝒮sp,andQt′\(y,x\)≥0,for\(y,x\)∈𝒮sp\.Q\_\{t\}^\{\\prime\}\(y,x\)=0,\\quad\\text\{for\}\\ \(y,x\)\\notin\\mathcal\{S\}^\{\\mathrm\{sp\}\},\\quad\\text\{and\}\\quad Q\_\{t\}^\{\\prime\}\(y,x\)\\geq 0,\\quad\\text\{for\}\\ \(y,x\)\\in\\mathcal\{S\}^\{\\mathrm\{sp\}\}\.In practice,𝒮sp\\mathcal\{S\}^\{\\mathrm\{sp\}\}is taken to be the support of the pretrained reverse rate family\.
###### Proposition 3\.4\(Sparse variance reduction as constrained weighted least squares\)\.
Assumeqt\(x\)\>0q\_\{t\}\(x\)\>0for allx∈𝒳x\\in\\mathcal\{X\}\. Fix a sparsity pattern𝒮sp\\mathcal\{S\}^\{\\mathrm\{sp\}\}and defineμt\(x\):=qt\(x\)gt0\(x\)\\mu\_\{t\}\(x\):=q\_\{t\}\(x\)\\,g\_\{t\}^\{0\}\(x\), wheregt0g\_\{t\}^\{0\}is given by \([3\.4](https://arxiv.org/html/2609.35947#S3.E4)\)\. Then minimizing \([3\.6](https://arxiv.org/html/2609.35947#S3.E6)\) over all sparse off\-diagonal rate families𝐐t′\\mathbf\{Q\}\_\{t\}^\{\\prime\}is equivalent to the constrained weighted least\-squares problem
minQt′∑x∈𝒳1qt\(x\)\[μt\(x\)\+∑y:\(y,x\)∈𝒮spQt′\(y,x\)qt\(x\)−∑y:\(x,y\)∈𝒮spQt′\(x,y\)qt\(y\)\]2\.\\min\_\{Q\_\{t\}^\{\\prime\}\}\\;\\sum\_\{x\\in\\mathcal\{X\}\}\\frac\{1\}\{q\_\{t\}\(x\)\}\\Big\[\\mu\_\{t\}\(x\)\+\\sum\_\{y:\\,\(y,x\)\\in\\mathcal\{S\}^\{\\mathrm\{sp\}\}\}Q\_\{t\}^\{\\prime\}\(y,x\)\\,q\_\{t\}\(x\)\-\\sum\_\{y:\\,\(x,y\)\\in\\mathcal\{S\}^\{\\mathrm\{sp\}\}\}Q\_\{t\}^\{\\prime\}\(x,y\)\\,q\_\{t\}\(y\)\\Big\]^\{2\}\.\(3\.8\)Equivalently, the objective depends only on the edge fluxesQt′\(y,x\)qt\(x\)Q\_\{t\}^\{\\prime\}\(y,x\)q\_\{t\}\(x\), i\.e\. on𝐐t′diag\(qt\)\\mathbf\{Q\}\_\{t\}^\{\\prime\}\\diag\(q\_\{t\}\)\.
This characterization makes the sparse\-control problem precise: the dense ideal from Proposition[3\.3](https://arxiv.org/html/2609.35947#S3.Thmtheorem3)corresponds to solving away the source termμt\\mu\_\{t\}on the complete graph, while the implementable optimum is exactly the weighted least\-squares problem \([3\.8](https://arxiv.org/html/2609.35947#S3.E8)\) restricted to the sparse pattern\. However, problem \([3\.8](https://arxiv.org/html/2609.35947#S3.E8)\) is a large\-scale nonnegativity\-constrained weighted least\-squares program with no closed\-form solution: it requires iterative solvers \(*e\.g\.*, projected gradient\) or coarse approximations, both of which add per\-step computational cost at inference time\. This motivates the practical, closed\-form sparse controls developed in the next subsection\.
### 3\.3Practical sparse controls
We now describe two practical constructions that produce a sparse off\-diagonal rate family𝐐teff\\mathbf\{Q\}\_\{t\}^\{\\mathrm\{eff\}\}and an associated residual centered potential
gteff\(x\):=gt0\(x\)\+divqt𝐐teff\(x\),g\_\{t\}^\{\\mathrm\{eff\}\}\(x\):=g\_\{t\}^\{0\}\(x\)\+\\operatorname\{div\}\_\{q\_\{t\}\}\\mathbf\{Q\}\_\{t\}^\{\\mathrm\{eff\}\}\(x\),\(3\.9\)which is the quantity used for particle weighting in Feynman–Kac SMC\. By Proposition[3\.1](https://arxiv.org/html/2609.35947#S3.Thmtheorem1), every such choice yields an equivalent representation of the same target path\.
#### 3\.3\.1HEU: one\-hop local reallocation
Recall that the dense zero\-variance construction in Proposition[3\.3](https://arxiv.org/html/2609.35947#S3.Thmtheorem3)can be written at the flux level as
Qt∗\(y,x\)qt\(x\)=1D\[qt\(y\)gt0\(y\)−qt\(x\)gt0\(x\)\]\+\.Q\_\{t\}^\{\*\}\(y,x\)\\,q\_\{t\}\(x\)=\\frac\{1\}\{D\}\\,\[q\_\{t\}\(y\)g\_\{t\}^\{0\}\(y\)\-q\_\{t\}\(x\)g\_\{t\}^\{0\}\(x\)\]\_\{\+\}\.This suggests a sparse local analogue: keep the same pairwise low\-to\-high transport rule, but restrict it to a local neighborhood aroundxxand replace the global normalization factorDDby a local effective neighborhood size\.
###### Proposition 3\.5\(Local\-average reallocation heuristic\)\.
Assumeqt\(x\)\>0q\_\{t\}\(x\)\>0for allx∈𝒳x\\in\\mathcal\{X\}\. Fix a sparsity pattern𝒮sp\\mathcal\{S\}^\{\\mathrm\{sp\}\}, letNt\+\(x\):=\{y∈𝒳∖\{x\}:\(y,x\)∈𝒮sp\}N\_\{t\}^\{\+\}\(x\):=\\\{y\\in\\mathcal\{X\}\\setminus\\\{x\\\}:\(y,x\)\\in\\mathcal\{S\}^\{\\mathrm\{sp\}\}\\\}be the allowed outgoing neighborhood fromxx, letkt\(x\)\>0k\_\{t\}\(x\)\>0be a prescribed local normalization factor, and letαt∈\(0,1\]\\alpha\_\{t\}\\in\(0,1\]be an optional damping parameter\. Define the sparse off\-diagonal rate family𝐐theu\\mathbf\{Q\}\_\{t\}^\{\{\\mathrm\{heu\}\}\}by
Qtheu\(y,x\):=αtkt\(x\)qt\(x\)\[μt\(y\)−μt\(x\)\]\+1\{y∈Nt\+\(x\)\},Q\_\{t\}^\{\{\\mathrm\{heu\}\}\}\(y,x\):=\\frac\{\\alpha\_\{t\}\}\{k\_\{t\}\(x\)\\,q\_\{t\}\(x\)\}\\big\[\\mu\_\{t\}\(y\)\-\\mu\_\{t\}\(x\)\\big\]\_\{\+\}\\,\\mathbf\{1\}\\\{y\\in N\_\{t\}^\{\+\}\(x\)\\\},\(3\.10\)with corresponding residual centered potentialgtheu\(x\):=gt0\(x\)\+divqt𝐐theu\(x\)g\_\{t\}^\{\{\\mathrm\{heu\}\}\}\(x\):=g\_\{t\}^\{0\}\(x\)\+\\operatorname\{div\}\_\{q\_\{t\}\}\\mathbf\{Q\}\_\{t\}^\{\{\\mathrm\{heu\}\}\}\(x\)\.
Proposition[3\.5](https://arxiv.org/html/2609.35947#S3.Thmtheorem5)is a sparse one\-hop surrogate of the dense zero\-variance formula: flux is sent only along allowed edges and only from smaller source termsμt\(x\)=qt\(x\)gt0\(x\)\\mu\_\{t\}\(x\)=q\_\{t\}\(x\)g\_\{t\}^\{0\}\(x\)to larger ones\. Under reciprocal neighborhoods with matching local normalizations, the rule has the interpretation of replacingμt\(x\)\\mu\_\{t\}\(x\)by a damped average of nearby sources, so it removes the local fluctuation of the pure\-reweighting source while leaving the local mean\. In implementation,μt\\mu\_\{t\}is evaluated via \([3\.4](https://arxiv.org/html/2609.35947#S3.E4)\); larger\-neighborhood variants and mask\-diffusion choices forkt\(x\)k\_\{t\}\(x\)are deferred to Appendix[C\.9](https://arxiv.org/html/2609.35947#A3.SS9)\.
##### RoutineHEU\.
For each source statexx, evaluate the local source termsμt\(y\)\\mu\_\{t\}\(y\)on the implemented neighborhoodNt\+\(x\)N\_\{t\}^\{\+\}\(x\)using \([3\.4](https://arxiv.org/html/2609.35947#S3.E4)\), choosekt\(x\)k\_\{t\}\(x\)according to one of the rules in Appendix[C\.9](https://arxiv.org/html/2609.35947#A3.SS9), set the outgoing rates by \([3\.10](https://arxiv.org/html/2609.35947#S3.E10)\), and evaluate the residual centered potentialgtheug\_\{t\}^\{\{\\mathrm\{heu\}\}\}via the general form \([3\.9](https://arxiv.org/html/2609.35947#S3.E9)\)\.
#### 3\.3\.2D\-VCG: discrete variance\-controlling guidance
The sparse least\-squares optimum is still too large to solve on DLM\-scale state spaces\.Discrete Variance\-Controlling Guidance\(D\-VCG\) therefore restricts the controlled rate to a small nonnegative span of basis rates obtained by reweighting the pretrained reverse rate by nonnegative multipliersφt\(j\)\\varphi\_\{t\}^\{\(j\)\}: forj∈\[J\]j\\in\[J\]andx≠yx\\neq y,
Qt\(j\)\(y,x\):=Qt←\(y,x\)φt\(j\)\(y,x\)\.Q\_\{t\}^\{\(j\)\}\(y,x\):=Q\_\{t\}^\{\\shortleftarrow\}\(y,x\)\\,\\varphi\_\{t\}^\{\(j\)\}\(y,x\)\.\(3\.11\)The choice of multipliers\{φt\(j\)\}j=1J\\\{\\varphi\_\{t\}^\{\(j\)\}\\\}\_\{j=1\}^\{J\}is open and can be tailored to the task; many bases beyond the two we keep in the main text are admissible, and Appendix[E\.3\.2](https://arxiv.org/html/2609.35947#A5.SS3.SSS2.Px3)lists the augmented library we designed for the Ising experiments\. Throughout the paper we always include the*untilted backward*basisφt\(1\)\(y,x\)≡1\\varphi\_\{t\}^\{\(1\)\}\(y,x\)\\equiv 1, and pair it with the*target\-aligned anchor*
φt\(2\)\(y,x\):=\(pt←\(x\)pt←\(y\)\)1−γexp\(rt\(y\)−rt\(x\)\),\\varphi\_\{t\}^\{\(2\)\}\(y,x\):=\\left\(\\frac\{p\_\{t\}^\{\\shortleftarrow\}\(x\)\}\{p\_\{t\}^\{\\shortleftarrow\}\(y\)\}\\right\)^\{1\-\\gamma\}\\exp\\big\(r\_\{t\}\(y\)\-r\_\{t\}\(x\)\\big\),which combines annealing \(exponentγ\\gamma\) and reward tilting \(potentialrtr\_\{t\}\); the pure\-annealing and pure\-tilt specializations, recovered by settingrt≡0r\_\{t\}\\\!\\equiv\\\!0orγ=1\\gamma\\\!=\\\!1, are listed as their own bases in Appendix[E\.3\.2](https://arxiv.org/html/2609.35947#A5.SS3.SSS2.Px3)\.
The controlled family𝐐tvcg\(𝜽\)=∑j=1Jθj𝐐t\(j\)\\mathbf\{Q\}\_\{t\}^\{\\mathrm\{vcg\}\}\(\\boldsymbol\{\\theta\}\)=\\sum\_\{j=1\}^\{J\}\\theta\_\{j\}\\mathbf\{Q\}\_\{t\}^\{\(j\)\},𝜽∈ℝ\+J\\boldsymbol\{\\theta\}\\in\\mathbb\{R\}\_\{\+\}^\{J\}, is a valid off\-diagonal rate family for every nonnegative coefficient vector\. The residual centered potential is
gtvcg\(x,𝜽\)=gt0\(x\)\+∑j=1Jθjdivqt𝐐t\(j\)\(x\)\.g\_\{t\}^\{\\mathrm\{vcg\}\}\(x;\\boldsymbol\{\\theta\}\)=g\_\{t\}^\{0\}\(x\)\+\\sum\_\{j=1\}^\{J\}\\theta\_\{j\}\\,\\operatorname\{div\}\_\{q\_\{t\}\}\\mathbf\{Q\}\_\{t\}^\{\(j\)\}\(x\)\.WritingD~t\(j\)\(x\):=divqt𝐐t\(j\)\(x\)−𝔼qt\[divqt𝐐t\(j\)\]\\widetilde\{D\}\_\{t\}^\{\(j\)\}\(x\):=\\operatorname\{div\}\_\{q\_\{t\}\}\\mathbf\{Q\}\_\{t\}^\{\(j\)\}\(x\)\-\\mathbb\{E\}\_\{q\_\{t\}\}\[\\operatorname\{div\}\_\{q\_\{t\}\}\\mathbf\{Q\}\_\{t\}^\{\(j\)\}\], the population objective reduces to a low\-dimensional nonnegative least\-squares problem\.
###### Proposition 3\.6\(D\-VCG\)\.
The unconstrained minimizers ofmin𝛉∈ℝJVarqt\(gtvcg\(⋅,𝛉\)\)\\min\_\{\\boldsymbol\{\\theta\}\\in\\mathbb\{R\}^\{J\}\}\\operatorname\{Var\}\_\{q\_\{t\}\}\\big\(g\_\{t\}^\{\\mathrm\{vcg\}\}\(\\cdot;\\boldsymbol\{\\theta\}\)\\big\)are exactly the solutions of the normal equations𝐀t𝛉=−𝐜t\\mathbf\{A\}\_\{t\}\\boldsymbol\{\\theta\}=\-\\mathbf\{c\}\_\{t\}, where
\(𝐀t\)ij=𝔼qt\[D~t\(i\)\(x\)D~t\(j\)\(x\)\],\(𝐜t\)i=𝔼qt\[gt0\(x\)D~t\(i\)\(x\)\],i,j∈\[J\]\.\(\\mathbf\{A\}\_\{t\}\)\_\{ij\}=\\mathbb\{E\}\_\{q\_\{t\}\}\\\!\\big\[\\widetilde\{D\}\_\{t\}^\{\(i\)\}\(x\)\\,\\widetilde\{D\}\_\{t\}^\{\(j\)\}\(x\)\\big\],\\qquad\(\\mathbf\{c\}\_\{t\}\)\_\{i\}=\\mathbb\{E\}\_\{q\_\{t\}\}\\\!\\big\[g\_\{t\}^\{0\}\(x\)\\,\\widetilde\{D\}\_\{t\}^\{\(i\)\}\(x\)\\big\],\\qquad i,j\\in\[J\]\.\(3\.12\)If𝐀t\\mathbf\{A\}\_\{t\}is nonsingular, the unconstrained minimizer is unique\. With the constraint𝛉≥0\\boldsymbol\{\\theta\}\\geq 0, the same objective is a nonnegative quadratic program inJJvariables\.
##### RoutineD\-VCG\.
Given weighted particles approximatingqtq\_\{t\}, evaluate the basis divergences on the particle cloud and assemble the weighted system \([3\.12](https://arxiv.org/html/2609.35947#S3.E12)\)\. We approximately solve the resulting nonnegative QP by active\-set candidate enumeration; details of the small KKT solves and the Ising basis library are in Appendix[E](https://arxiv.org/html/2609.35947#A5)\. Let𝜽⋆\\boldsymbol\{\\theta\}^\{\\star\}denote the selected nonnegative coefficient vector\. The effective rates and residual are𝐐teff=𝐐tvcg\(𝜽⋆\)\\mathbf\{Q\}\_\{t\}^\{\\mathrm\{eff\}\}=\\mathbf\{Q\}\_\{t\}^\{\\mathrm\{vcg\}\}\(\\boldsymbol\{\\theta\}^\{\\star\}\)andgteff=gtvcg\(⋅,𝜽⋆\)g\_\{t\}^\{\\mathrm\{eff\}\}=g\_\{t\}^\{\\mathrm\{vcg\}\}\(\\cdot;\\boldsymbol\{\\theta\}^\{\\star\}\)\.
##### Feynman–Kac SMC recursion\.
On a grid0=t0<⋯<tM=T0=t\_\{0\}<\\cdots<t\_\{M\}=T, the sampler iterates four operations per step: construct𝐐tkeff\\mathbf\{Q\}\_\{t\_\{k\}\}^\{\\mathrm\{eff\}\}andgtkeffg\_\{t\_\{k\}\}^\{\\mathrm\{eff\}\}byHEUorD\-VCG, multiply weights byexp\(Δtkgtkeff\)\\exp\(\\Delta t\_\{k\}g\_\{t\_\{k\}\}^\{\\mathrm\{eff\}\}\), resample when the ESS ratio falls belowτ\\tau, and propagate particles over\[tk,tk\+1\]\[t\_\{k\},t\_\{k\+1\}\]with the CTMC rate𝐐tkeff\\mathbf\{Q\}\_\{t\_\{k\}\}^\{\\mathrm\{eff\}\}\. Algorithm[1](https://arxiv.org/html/2609.35947#algorithm1)gives a first\-order implementation skeleton; the experimental splitting scheme is specified in Appendix[E\.2](https://arxiv.org/html/2609.35947#A5.SS2)\.
Algorithm 1FluxLite controlled Feynman–Kac SMCInput:time grid
0=t0<⋯<tM=T0=t\_\{0\}<\\cdots<t\_\{M\}=T; particle count
NN; ESS threshold
τ\\tau; baseline Feynman–Kac pairs
\{\(𝐐tk,gtk\)\}k=0M−1\\\{\(\\mathbf\{Q\}\_\{t\_\{k\}\},g\_\{t\_\{k\}\}\)\\\}\_\{k=0\}^\{M\-1\}; control routine
CONTROL∈\{HEU,D\-VCG\}\\texttt\{CONTROL\}\\in\\\{\\texttt\{HEU\},\\texttt\{D\-VCG\}\\\}\.
Output:Weighted terminal particles approximating
qTq\_\{T\}\.
1Initialize particles
xt0\(n\)∼qt0x\_\{t\_\{0\}\}^\{\(n\)\}\\sim q\_\{t\_\{0\}\}and weights
wt0\(n\)←1/Nw\_\{t\_\{0\}\}^\{\(n\)\}\\leftarrow 1/N;
2for*k=0,…,M−1k=0,\\ldots,M\-1*do
3
Δtk←tk\+1−tk\\Delta t\_\{k\}\\leftarrow t\_\{k\+1\}\-t\_\{k\}and form the current weighted empirical approximation of
qtkq\_\{t\_\{k\}\};
4Compute the pure\-reweighting source
gtk0g\_\{t\_\{k\}\}^\{0\}using \([3\.4](https://arxiv.org/html/2609.35947#S3.E4)\);
5
\(𝐐tkeff,gtkeff\)←CONTROL\(qtkN,𝐐tk,gtk0\)\(\\mathbf\{Q\}\_\{t\_\{k\}\}^\{\\mathrm\{eff\}\},g\_\{t\_\{k\}\}^\{\\mathrm\{eff\}\}\)\\leftarrow\\texttt\{CONTROL\}\(q\_\{t\_\{k\}\}^\{N\},\\mathbf\{Q\}\_\{t\_\{k\}\},g\_\{t\_\{k\}\}^\{0\}\)byHEUorD\-VCG;
6Update and normalize weights:
w\(n\)←w\(n\)exp\{Δtkgtkeff\(xtk\(n\)\)\}w^\{\(n\)\}\\leftarrow w^\{\(n\)\}\\exp\\\{\\Delta t\_\{k\}g\_\{t\_\{k\}\}^\{\\mathrm\{eff\}\}\(x\_\{t\_\{k\}\}^\{\(n\)\}\)\\\};
7if*the ESS ratio is belowτ\\tau*then
8resample particles and reset weights to
1/N1/N
9end if
10Propagate each particle from
tkt\_\{k\}to
tk\+1t\_\{k\+1\}with CTMC rates
𝐐tkeff\\mathbf\{Q\}\_\{t\_\{k\}\}^\{\\mathrm\{eff\}\};
11end for
## 4Theoretical Analysis
The analysis separates two sources of error\. The first is a*population*error: the deployed discrete diffusion uses an estimated local ratios^t\\widehat\{s\}\_\{t\}rather than the exact ratiosts\_\{t\}, so the guided Feynman–Kac operator itself is perturbed\. The second is a*particle*error: even for a fixed controlled Feynman–Kac representation, SMC approximates the corresponding grid\-level Feynman–Kac recursion using finitely many particles\. These results are intentionally modular\. They do not claim a full end\-to\-end bound for adaptive controls estimated from the same particle cloud; the scope of each statement is made explicit below and in Remark[D\.9](https://arxiv.org/html/2609.35947#A4.Thmtheorem9)\.
##### Population stability under local\-ratio error\.
Letqts^q\_\{t\}^\{\\widehat\{s\}\}denote the exact population normalized Feynman–Kac flow obtained by replacingsts\_\{t\}withs^t\\widehat\{s\}\_\{t\}in the reverse rates, guided rates, and potentials of Proposition[2\.1](https://arxiv.org/html/2609.35947#S2.Thmtheorem1); it is not the empirical particle law\. Writeut\(x,y\):=s^t\(x,y\)/st\(x,y\)u\_\{t\}\(x,y\):=\\widehat\{s\}\_\{t\}\(x,y\)/s\_\{t\}\(x,y\)\. For source\-weighted score\-entropy training\[[97](https://arxiv.org/html/2609.35947#bib.bib97)\], setνtSE\(x,y\):=pt←\(x\)Qt←\(y,x\)\\nu\_\{t\}^\{\\rm SE\}\(x,y\):=p\_\{t\}^\{\\shortleftarrow\}\(x\)Q\_\{t\}^\{\\shortleftarrow\}\(y,x\)and define the time\-integrated score\-entropy training loss
𝔏DD:=∫0TℒDD\(t\)𝑑t,ℒDD\(t\):=∑x∈𝒳∑y≠xνtSE\(x,y\)ℓ\(ut\(x,y\)\),ℓ\(u\):=−logu−1\+u\.\\mathfrak\{L\}\_\{\\rm DD\}:=\\int\_\{0\}^\{T\}\\mathcal\{L\}\_\{\\rm DD\}\(t\)\\,\\mathop\{\}\\\!\\mathrm\{d\}t,\\quad\\mathcal\{L\}\_\{\\rm DD\}\(t\):=\\sum\_\{x\\in\\mathcal\{X\}\}\\sum\_\{y\\neq x\}\\nu\_\{t\}^\{\\rm SE\}\(x,y\)\\,\\ell\\big\(u\_\{t\}\(x,y\)\\big\),\\quad\\ell\(u\):=\-\\log u\-1\+u\.\(4\.1\)Up to a constant,𝔏DD\\mathfrak\{L\}\_\{\\rm DD\}upper\-bounds the model–data KL\[[97](https://arxiv.org/html/2609.35947#bib.bib97),[117](https://arxiv.org/html/2609.35947#bib.bib117)\]\. Letht\(x\):=qt\(x\)/pt←\(x\)h\_\{t\}\(x\):=q\_\{t\}\(x\)/p\_\{t\}^\{\\shortleftarrow\}\(x\); the theorem weights this coverage ratio edgewise rather than replacing it by a worst\-case supremum\.
###### Theorem 4\.1\(Population score\-ratio stability, informal\)\.
Assume the known\-forward setting and the hypotheses of Theorem[D\.4](https://arxiv.org/html/2609.35947#A4.Thmtheorem4), and letℭDD\\mathfrak\{C\}\_\{\\rm DD\}be the time\-integrated edge\-coverage scalar defined in \([D\.9](https://arxiv.org/html/2609.35947#A4.E9)\)\. Then
TV\(qTs^,qT\)≤γe2BTℭDD𝔏DD\.\\TV\(q\_\{T\}^\{\\widehat\{s\}\},q\_\{T\}\)\\leq\\gamma\\,e^\{2BT\}\\,\\mathfrak\{C\}\_\{\\rm DD\}\\,\\sqrt\{\\mathfrak\{L\}\_\{\\rm DD\}\}\.\(4\.2\)The exponente2BTe^\{2BT\}comes from theℓ1\\ell^\{1\}logarithmic norm of the implemented Metzler operator\. If𝔏DD=0\\mathfrak\{L\}\_\{\\rm DD\}=0thenqTs^=qTq\_\{T\}^\{\\widehat\{s\}\}=q\_\{T\}\.
The factorℭDD\\mathfrak\{C\}\_\{\\rm DD\}is a*coverage*scalar, not a score\-accuracy condition: inference\-time scaling with a fixed pretrained score is robust only when the tilted path remains covered by the original reverse path on trained edges\. If reward or temperature scaling pushesqtq\_\{t\}toward regions wherehth\_\{t\}is large at edge endpoints carryingνtSE\\nu\_\{t\}^\{\\rm SE\}mass, the same training loss𝔏DD\\mathfrak\{L\}\_\{\\rm DD\}can produce a larger terminal bias\.
##### Particle error for a fixed controlled representation\.
Now fix a grid0=t0<⋯<tM=T0=t\_\{0\}<\\cdots<t\_\{M\}=Tand deterministic controlled rates/potentials\(𝐐tkeff,gtkeff\)\(\\mathbf\{Q\}\_\{t\_\{k\}\}^\{\{\\mathrm\{eff\}\}\},g\_\{t\_\{k\}\}^\{\{\\mathrm\{eff\}\}\}\)\. Let𝐏k:=exp\(Δtk𝐋tkeff\)\\mathbf\{P\}\_\{k\}:=\\exp\(\\Delta t\_\{k\}\\mathbf\{L\}\_\{t\_\{k\}\}^\{\{\\mathrm\{eff\}\}\}\)andWk\(x\):=exp\(Δtkgtkeff\(x\)\)W\_\{k\}\(x\):=\\exp\(\\Delta t\_\{k\}g\_\{t\_\{k\}\}^\{\{\\mathrm\{eff\}\}\}\(x\)\)\. With the column convention,Pk\(y,x\)P\_\{k\}\(y,x\)is the transition probability from sourcexxto destinationyy, and\(𝖯kf\)\(x\):=\(𝐏k⊤f\)\(x\)=∑yPk\(y,x\)f\(y\)\(\\mathsf\{P\}\_\{k\}f\)\(x\):=\(\\mathbf\{P\}\_\{k\}^\{\\top\}f\)\(x\)=\\sum\_\{y\}P\_\{k\}\(y,x\)f\(y\)\. Define
q0Δ=qt0,qk\+1Δ\(f\)=qkΔ\(Wk𝖯kf\)qkΔ\(Wk\),k=0,…,M−1\.q\_\{0\}^\{\\Delta\}=q\_\{t\_\{0\}\},\\qquad q\_\{k\+1\}^\{\\Delta\}\(f\)=\\frac\{q\_\{k\}^\{\\Delta\}\(W\_\{k\}\\mathsf\{P\}\_\{k\}f\)\}\{q\_\{k\}^\{\\Delta\}\(W\_\{k\}\)\},\\qquad k=0,\\ldots,M\-1\.\(4\.3\)This theorem analyzes the every\-step bootstrap version of Algorithm[1](https://arxiv.org/html/2609.35947#algorithm1)\. It isolates the Monte Carlo error of a fixed Feynman–Kac model; adaptive ESS resampling, time discretization, score error, and feedback from estimating controls on the same particles are outside the statement and are discussed in Remark[D\.9](https://arxiv.org/html/2609.35947#A4.Thmtheorem9)\.
###### Theorem 4\.2\(Particle approximation of the fixed grid Feynman–Kac recursion\)\.
Under the deterministic fixed\-grid setting of Assumption[D\.4](https://arxiv.org/html/2609.35947#A4.Thmassumption4), letqMN,Δq\_\{M\}^\{N,\\Delta\}be the empirical law produced afterMMbootstrap Feynman–Kac SMC steps fromNNparticles\. For every boundedf:𝒳→ℝf:\\mathcal\{X\}\\to\\mathbb\{R\},
𝔼\[\|qMN,Δ\(f\)−qMΔ\(f\)\|\]≤osc\(f\)2N∑k=0Mβ¯k,M,β¯k,M:=∏j=kM−1eΔtjosc\(gtjeff\)\.\\mathbb\{E\}\\big\[\|q\_\{M\}^\{N,\\Delta\}\(f\)\-q\_\{M\}^\{\\Delta\}\(f\)\|\\big\]\\leq\\frac\{\\osc\(f\)\}\{2\\sqrt\{N\}\}\\sum\_\{k=0\}^\{M\}\\bar\{\\beta\}\_\{k,M\},\\quad\\bar\{\\beta\}\_\{k,M\}:=\\prod\_\{j=k\}^\{M\-1\}e^\{\\Delta t\_\{j\}\\osc\(g\_\{t\_\{j\}\}^\{\{\\mathrm\{eff\}\}\}\)\}\.\(4\.4\)
This theorem explains why FluxLite targets residual potentials: its constant depends on residual oscillation, and Lemma[D\.6](https://arxiv.org/html/2609.35947#A4.Thmtheorem6)givesVarqkΔ\(W¯k\)≤Δtk2e2Δtk‖g¯k‖∞VarqkΔ\(gtkeff\)\\operatorname\{Var\}\_\{q\_\{k\}^\{\\Delta\}\}\(\\bar\{W\}\_\{k\}\)\\leq\\Delta t\_\{k\}^\{2\}e^\{2\\Delta t\_\{k\}\\left\\\|\\bar\{g\}\_\{k\}\\right\\\|\_\{\\infty\}\}\\operatorname\{Var\}\_\{q\_\{k\}^\{\\Delta\}\}\(g\_\{t\_\{k\}\}^\{\{\\mathrm\{eff\}\}\}\)\. The proof \(Appendix[D\.3](https://arxiv.org/html/2609.35947#A4.SS3)\) uses conditional i\.i\.d\. resampling–mutation and a nonlinear telescoping decomposition, separating fixed\-model Monte Carlo error from the population score\-ratio perturbation in Theorem[4\.1](https://arxiv.org/html/2609.35947#S4.Thmtheorem1)\.
## 5Experiments
We evaluate our method using an exact finite\-state CTMC benchmark and a learned 2D Ising benchmark\. Across both,D\-FKCis our Feynman–Kac corrector baseline andPGis guided propagation without reweighting;PR,HEU,D\-VCG, andDENare the samplers from Propositions[3\.2](https://arxiv.org/html/2609.35947#S3.Thmtheorem2),[3\.5](https://arxiv.org/html/2609.35947#S3.Thmtheorem5),[3\.6](https://arxiv.org/html/2609.35947#S3.Thmtheorem6), and[3\.3](https://arxiv.org/html/2609.35947#S3.Thmtheorem3), respectively\. Additional large\-scale experiments on a discrete inverse problem \(monophonic music infilling\) and on text\-to\-image generation under classifier\-free guidance are reported in Appendix[F](https://arxiv.org/html/2609.35947#A6)\.
### 5\.1Finite\-state CTMC
The benchmark combines reward tilting and annealing with tensor\-product*uniform\-state*and*masked\-absorbing*CTMC diffusions atV=5V=5,L=3L=3\. The uniform diffusion lives onVL=125V^\{L\}=125states; the masked diffusion evolves on\(V\+1\)L=216\(V\+1\)^\{L\}=216states including the mask token and has an unmasked terminal target onVLV^\{L\}states\.PRuses\(0,gt0\)\(0,g\_\{t\}^\{0\}\)and is omitted in the masked setting because reweighting cannot introduce states absent from the initial particle cloud\.
Figure 1:Finite\-state CTMC benchmark: terminalKL\(qT∥q^TN\)\\KL\(q\_\{T\}\\\|\\widehat\{q\}\_\{T\}^\{N\}\)for fourV=5V=5,L=3L=3regimes withN=4000N=4000particles, 80 integration steps, and 10 seeds\.D\-VCGimproves overD\-FKC\(geometric\-mean ratio of terminal KL across seeds\) up to114\.8×114\.8\\times; settings and sweeps are in Appendix[E](https://arxiv.org/html/2609.35947#A5)\.Figure[1](https://arxiv.org/html/2609.35947#S5.F1)reports terminalKL\(qT∥q^TN\)\\KL\(q\_\{T\}\\\|\\widehat\{q\}\_\{T\}^\{N\}\)\.D\-VCGreallocates flux toward useful transitions while keeping residual weights controlled, whereasPGlacks correction\.D\-FKCproduces single\-seed blow\-ups at largeγ\\gammaon uniform\-state/annealing\.D\-VCGmatches or improves on the population zero\-variance oracleDENacross all four cells while operating only on the sparse graph implementable in DLM\-scale state spaces\. Appendix[E\.2](https://arxiv.org/html/2609.35947#A5.SS2)gives sweeps and implementation details\.
### 5\.22D Ising Model
We evaluate a score\-based discrete\-diffusion model trained on Swendsen–Wang samples atβtrain=0\.4\\beta\_\{\\rm train\}=0\.4on a16×1616\\times 16periodic Ising lattice \(JIsing=1J\_\{\\rm Ising\}=1\)\. Figure[2](https://arxiv.org/html/2609.35947#S5.F2)sweeps target inverse temperature and magnetization rewardr\(σ\)=M\(σ\)=∑iσir\(\\sigma\)=M\(\\sigma\)=\\sum\_\{i\}\\sigma\_\{i\}forqβ,βr\(σ\)∝e−βH\(σ\)\+βrM\(σ\)q\_\{\\beta,\\beta\_\{r\}\}\(\\sigma\)\\propto e^\{\-\\beta H\(\\sigma\)\+\\beta\_\{r\}M\(\\sigma\)\}, with a target\-specific SW reference and a ghost spin for nonzero fields\. The figure labels “2\-basis” and “4\-basis” denote the minimal and augmented controller variants; their regime\-dependent libraries are specified in Appendix[E\.3\.2](https://arxiv.org/html/2609.35947#A5.SS3.SSS2.Px3)\.
\(a\)Pure annealing: 4\-basisD\-VCGcutsMSE\(C\(r\)\)\\mathrm\{MSE\}\(C\(r\)\)by5\.33×5\.33\\timesin geometric mean overD\-FKC, peak24\.5×\\mathbf\{24\.5\\times\}\.\(b\)Joint anneal\+\+tilt: 4\-basisD\-VCGcutsMSE\(C\(r\)\)\\mathrm\{MSE\}\(C\(r\)\)by6\.77×6\.77\\timesin geometric mean overD\-FKC, peak55\.4×\\mathbf\{55\.4\\times\}\.
Figure 2:2D Ising experiments against a SW reference, with a ghost spin for nonzero fields\. Panels report magnetization/energy\-𝖶2\\mathsf\{W\}\_\{2\}, row\-correlationMSE\(C\(r\)\)\\mathrm\{MSE\}\(C\(r\)\), and mean magnetization forPG,D\-FKC, and 2\-/4\-basisD\-VCG\.The largest gains appear in row correlations:D\-VCGreduces their MSE overD\-FKCby5\.33×5\.33\\timesin geometric mean on annealing and6\.77×6\.77\\timeson joint anneal plus tilt\. The per\-cell peaks in Fig\.[2](https://arxiv.org/html/2609.35947#S5.F2)are24\.5×24\.5\\timesand55\.4×55\.4\\times, with improvements across the plotted annealing and joint\-tilt sweeps\. A particle\-count ablation \(Fig\.[5](https://arxiv.org/html/2609.35947#A5.F5)\) confirms the gap does not close asNNgrows, with a15\.21×15\.21\\timesgeometric\-mean and34\.7×34\.7\\timespeak advantage persisting fromN=100N=100toN=5000N=5000\. Feynman–Kac\-reweighted methods track SW means\. Appendix[E\.3](https://arxiv.org/html/2609.35947#A5.SS3)gives exact configurations and ablations\.
## 6Discussion and Conclusion
FluxLite is a discrete counterpart of DriftLite for Feynman–Kac SMC in discrete diffusion models and DLMs\. Its core is an exact CTMC–Feynman–Kac equivalence class on the pretrained transition graph: sparse rate perturbations reallocate flux, and aqtq\_\{t\}\-weighted graph divergence compensates the residual potential without changing tilted marginals\.HEUandD\-VCGinstantiate this variance\-control principle; the theory ties score\-ratio loss to guided Feynman–Kac bias and residual\-potential oscillation to fixed\-grid particle error, while CTMC and Ising experiments show large gains overD\-FKC\. Future directions include hybrid discrete–continuous controls, alignment/preference/pretraining couplings, and distillation into transport maps or parallel\-reasoning architectures\. On the theoretical side, natural extensions include adaptive ESS, particle\-dependent controls, joint score and time\-discretization error, and richer learned bases\.
## Acknowledgments
Y\.R\. and J\.H\. thank the Scientific Computing Core at the Flatiron Institute, a division of the Simons Foundation, for providing computational resources and support\. This work has been made possible in part by a gift from the Chan Zuckerberg Initiative Foundation to establish the Kempner Institute for the Study of Natural and Artificial Intelligence at Harvard University\. L\.Y\. acknowledges support by the National Science Foundation under Award No\. DMS\-2208163\. This material is based upon work supported by the National Science Foundation under Grant No\. CHE\-2441297 to G\.M\.R\.
## References
- \[1\]Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al\.GPT\-4 technical report\.*arXiv preprint arXiv:2303\.08774*, 2023\.
- \[2\]Michael S Albergo and Eric Vanden\-Eijnden\.Building normalizing flows with stochastic interpolants\.In*The Eleventh International Conference on Learning Representations \(ICLR\)*, 2023\.
- \[3\]Michael S Albergo and Eric Vanden\-Eijnden\.Learning to sample better\.*Journal of Statistical Mechanics: Theory and Experiment*, 2024\(10\):104014, 2024\.
- \[4\]Michael S Albergo and Eric Vanden\-Eijnden\.NETS: A non\-equilibrium transport sampler\.In*International Conference on Machine Learning \(ICML\)*, pages 1026–1055\. PMLR, 2025\.
- \[5\]Michael S Albergo, Nicholas M Boffi, and Eric Vanden\-Eijnden\.Stochastic interpolants: A unifying framework for flows and diffusions\.*Journal of Machine Learning Research \(JMLR\)*, 26\(209\):1–80, 2025\.
- \[6\]Yashas Annadani, Syrine Belakaria, Stefano Ermon, Stefan Bauer, and Barbara E Engelhardt\.Preference\-guided diffusion for multi\-objective offline optimization\.*arXiv preprint arXiv:2503\.17299*, 2025\.
- \[7\]Iskander Azangulov, Peter Potaptchik, Qinyu Li, Eddie Aamari, George Deligiannidis, and Judith Rousseau\.Adaptive diffusion guidance via stochastic optimal control\.*arXiv preprint arXiv:2505\.19367*, 2025\.
- \[8\]Seyedarmin Azizi, Erfan Baghaei Potraghloo, Minoo Ahmadi, Souvik Kundu, and Massoud Pedram\.Power\-SMC: Low\-latency sequence\-level power sampling for training\-free LLM reasoning\.*arXiv preprint arXiv:2602\.10273*, 2026\.
- \[9\]Jinbin Bai, Tian Ye, Wei Chow, Enxin Song, Qing\-Guo Chen, Xiangtai Li, Zhen Dong, Lei Zhu, and Shuicheng Yan\.Meissonic: Revitalizing masked generative transformers for efficient high\-resolution text\-to\-image synthesis\.In*The Thirteenth International Conference on Learning Representations \(ICLR\)*, 2025\.
- \[10\]Jinbin Bai, Yixuan Li, Yuchen Zhu, Yi Xin, Qingyu Shi, Aosong Feng, Xiaohong Liu, Molei Tao, Jianru Xue, Xiangtai Li, et al\.Prism: Efficient test\-time scaling via hierarchical search and self\-verification for discrete diffusion language models\.*arXiv preprint arXiv:2602\.01842*, 2026\.
- \[11\]Vidhisha Balachandran, Jingya Chen, Lingjiao Chen, Shivam Garg, Neel Joshi, Yash Lara, John Langford, Besmira Nushi, Vibhav Vineet, Yue Wu, et al\.Inference\-time scaling for complex tasks: Where we stand and what lies ahead\.*arXiv preprint arXiv:2504\.00294*, 2025\.
- \[12\]Tiwei Bie, Maosong Cao, Kun Chen, Lun Du, Mingliang Gong, Zhuochen Gong, Yanmei Gu, Jiaqi Hu, Zenan Huang, Zhenzhong Lan, et al\.LLaDA2\.0: Scaling up diffusion language models to 100B\.*arXiv preprint arXiv:2512\.15745*, 2025\.
- \[13\]Mireille Bossy and Denis Talay\.A stochastic particle method for the McKean–Vlasov and the Burgers equation\.*Mathematics of computation*, 66\(217\):157–192, 1997\.
- \[14\]Arwen Bradley and Preetum Nakkiran\.Classifier\-free guidance is a predictor\-corrector\.*arXiv preprint arXiv:2408\.09000*, 2024\.
- \[15\]Paul C Bressloff\.Feynman–Kac formula for stochastic hybrid systems\.*Physical Review E*, 95\(1\):012138, 2017\.
- \[16\]Mark Briers, Arnaud Doucet, and Simon Maskell\.Smoothing algorithms for state–space models\.*Annals of the Institute of Statistical Mathematics*, 62\(1\):61–89, 2010\.
- \[17\]Bradley Brown, Jordan Juravsky, Ryan Ehrlich, Ronald Clark, Quoc V Le, Christopher Ré, and Azalia Mirhoseini\.Large language monkeys: Scaling inference compute with repeated sampling\.*arXiv preprint arXiv:2407\.21787*, 2024\.
- \[18\]Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al\.Language models are few\-shot learners\.*Advances in Neural Information Processing Systems \(NeurIPS\)*, 33:1877–1901, 2020\.
- \[19\]Dake Bu, Wei Huang, Andi Han, Atsushi Nitanda, Bo Xue, Qingfu Zhang, Hau\-San Wong, and Taiji Suzuki\.Post\-training as reweighting: A stochastic view of reasoning trajectories in language models\.*arXiv preprint arXiv:2511\.07368*, 2025\.
- \[20\]Andrew Campbell, Joe Benton, Valentin De Bortoli, Thomas Rainforth, George Deligiannidis, and Arnaud Doucet\.A continuous time framework for discrete denoising models\.*Advances in Neural Information Processing Systems \(NeurIPS\)*, 35:28266–28279, 2022\.
- \[21\]Yu Cao and Eric Vanden\-Eijnden\.Learning optimal flows for non\-equilibrium importance sampling\.*Advances in Neural Information Processing Systems \(NeurIPS\)*, 35:12352–12364, 2022\.
- \[22\]Huiwen Chang, Han Zhang, Lu Jiang, Ce Liu, and William T Freeman\.MaskGIT: Masked generative image transformer\.In*Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\)*, pages 11315–11325, 2022\.
- \[23\]Jinyuan Chang, Chenguang Duan, Yuling Jiao, Yi Xu, and Jerry Zhijian Yang\.Inference\-time alignment for diffusion models via Doob’s matching\.*arXiv preprint arXiv:2601\.06514*, 2026\.
- \[24\]Haoxuan Chen, Yinuo Ren, Lexing Ying, and Grant Rotskoff\.Accelerating diffusion models with parallel sampling: Inference at sub\-linear time complexity\.*Advances in Neural Information Processing Systems*, 37:133661–133709, 2024a\.
- \[25\]Haoxuan Chen, Yinuo Ren, Martin Renqiang Min, Lexing Ying, and Zachary Izzo\.Solving inverse problems via diffusion\-based priors: An approximation\-free ensemble sampling approach\.*Transactions on Machine Learning Research*, 2026a\.ISSN 2835\-8856\.URL[https://openreview\.net/forum?id=qN8ASsfjKs](https://openreview.net/forum?id=qN8ASsfjKs)\.
- \[26\]Sijin Chen, Yinuo Ren, Heyang Zhao, Ziheng Cheng, Quanquan Gu, and Lexing Ying\.Multi\-mask diffusion language models for few\-step generation\.In*Third Conference on Language Modeling*, 2026b\.URL[https://openreview\.net/forum?id=xR8qymwnOK](https://openreview.net/forum?id=xR8qymwnOK)\.
- \[27\]Zhe Chen, Weiyun Wang, Yue Cao, Yangzhou Liu, Zhangwei Gao, Erfei Cui, Jinguo Zhu, Shenglong Ye, Hao Tian, Zhaoyang Liu, et al\.Expanding performance boundaries of open\-source multimodal models with model, data, and test\-time scaling\.*arXiv preprint arXiv:2412\.05271*, 2024b\.
- \[28\]Alina Chertock\.A practical guide to deterministic particle methods\.In*Handbook of numerical analysis*, volume 18, pages 177–202\. Elsevier, 2017\.
- \[29\]Muthu Chidambaram, Khashayar Gatmiry, Sitan Chen, Holden Lee, and Jianfeng Lu\.What does guidance do? A fine\-grained analysis in a simple setting\.*Advances in Neural Information Processing Systems \(NeurIPS\)*, 37:84968–85005, 2024\.
- \[30\]Nicolas Chopin\.A sequential particle filter method for static models\.*Biometrika*, 89\(3\):539–552, 2002\.
- \[31\]Wenda Chu, Zihui Wu, Yifan Chen, Yang Song, and Yisong Yue\.Split Gibbs discrete diffusion posterior sampling\.*arXiv preprint arXiv:2503\.01161*, 2025\.
- \[32\]Inkook Chun, Seungjae Lee, Michael S Albergo, Saining Xie, and Eric Vanden\-Eijnden\.Dynamic test\-time compute scaling in control policy: Difficulty\-aware stochastic interpolant policy\.*arXiv preprint arXiv:2511\.20906*, 2025\.
- \[33\]Meihua Dang, Jiaqi Han, Minkai Xu, Kai Xu, Akash Srivastava, and Stefano Ermon\.Inference\-time scaling of diffusion language models with particle Gibbs sampling\.*arXiv preprint arXiv:2507\.08390*, 2025\.
- \[34\]Google DeepMind\.Gemini diffusion: Our state\-of\-the\-art, experimental text diffusion model\.[https://deepmind\.google/models/gemini\-diffusion/](https://deepmind.google/models/gemini-diffusion/), 2025\.
- \[35\]Pierre Degond and Sylvie Mas\-Gallic\.The weighted particle method for convection\-diffusion equations\. I\. the case of an isotropic viscosity\.*Mathematics of computation*, 53\(188\):485–507, 1989\.
- \[36\]Pierre Degond and Francisco\-José Mustieles\.A deterministic approximation of diffusion equations using particles\.*SIAM Journal on Scientific and Statistical Computing*, 11\(2\):293–310, 1990\.
- \[37\]Pierre Del Moral\.Mean field simulation for Monte Carlo integration\.*Monographs on Statistics and Applied Probability*, 126\(26\):6, 2013\.
- \[38\]Pierre Del Moral, Arnaud Doucet, and Ajay Jasra\.Sequential Monte Carlo samplers\.*Journal of the Royal Statistical Society Series B: Statistical Methodology*, 68\(3\):411–436, 2006\.
- \[39\]Pierre Del Moral, Pierre E Jacob, Anthony Lee, Lawrence Murray, and Gareth W Peters\.Feynman–Kac particle integration with geometric interacting jumps\.*Stochastic Analysis and Applications*, 31\(5\):830–871, 2013\.
- \[40\]Chenhui Deng, Yun\-Da Tsai, Guan\-Ting Liu, Zhongzhi Yu, and Haoxing Ren\.ScaleRTL: Scaling LLMs with reasoning data and test\-time compute for accurate RTL code generation\.In*2025 ACM/IEEE 7th Symposium on Machine Learning for CAD \(MLCAD\)*, pages 1–9\. IEEE, 2025\.
- \[41\]Jacob Devlin, Ming\-Wei Chang, Kenton Lee, and Kristina Toutanova\.BERT: Pre\-training of deep bidirectional transformers for language understanding\.In*Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 \(long and short papers\)*, pages 4171–4186, 2019\.
- \[42\]Qiwei Di, Kaixuan Ji, Xuheng Li, Heyang Zhao, and Quanquan Gu\.Best\-of\-majority: Minimax\-optimal strategy for pass@kkinference scaling\.In*The Fourteenth International Conference on Learning Representations \(ICLR\)*, 2026\.
- \[43\]Yifu Ding, Wentao Jiang, Shunyu Liu, Yongcheng Jing, Jinyang Guo, Yingjie Wang, Jing Zhang, Zengmao Wang, Ziwei Liu, Bo Du, et al\.Dynamic parallel tree search for efficient LLM reasoning\.In*Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 11233–11252, 2025\.
- \[44\]Hao\-Wen Dong, Wen\-Yi Hsiao, Li\-Chia Yang, and Yi\-Hsuan Yang\.MuseGAN: Multi\-track sequential generative adversarial networks for symbolic music generation and accompaniment\.In*Proceedings of the AAAI Conference on Artificial Intelligence*, volume 32, 2018\.
- \[45\]Arnaud Doucet, Adam M Johansen, et al\.A tutorial on particle filtering and smoothing: Fifteen years later\.*Handbook of nonlinear filtering*, 12\(656\-704\):3, 2009\.
- \[46\]Yilun Du, Conor Durkan, Robin Strudel, Joshua B Tenenbaum, Sander Dieleman, Rob Fergus, Jascha Sohl\-Dickstein, Arnaud Doucet, and Will Sussman Grathwohl\.Reduce, reuse, recycle: Compositional generation with energy\-based diffusion models and MCMC\.In*International Conference on Machine Learning \(ICML\)*, pages 8489–8510\. PMLR, 2023\.
- \[47\]Luca Eyring, Shyamgopal Karthik, Alexey Dosovitskiy, Nataniel Ruiz, and Zeynep Akata\.Noise hypernetworks: Amortizing test\-time compute in diffusion models\.*arXiv preprint arXiv:2508\.09968*, 2025\.
- \[48\]Ying Fan and Kangwook Lee\.Optimizing DDPM sampling with shortcut fine\-tuning\.In*International Conference on Machine Learning \(ICML\)*, pages 9623–9639\. PMLR, 2023\.
- \[49\]Zexi Fan, Yan Sun, Shihao Yang, and Yiping Lu\.Physics\-informed inference time scaling for solving high\-dimensional partial differential equations\.In*The Fourteenth International Conference on Learning Representations \(ICLR\)*, 2026\.
- \[50\]Shengyu Feng, Xiang Kong, Shuang Ma, Aonan Zhang, Dong Yin, Chong Wang, Ruoming Pang, and Yiming Yang\.Step\-by\-step reasoning for math problems via twisted sequential Monte Carlo\.In*The Thirteenth International Conference on Learning Representations \(ICLR\)*, 2025\.
- \[51\]Marylou Gabrié, Grant M Rotskoff, and Eric Vanden\-Eijnden\.Adaptive Monte Carlo augmented with normalizing flows\.*Proceedings of the National Academy of Sciences \(PNAS\)*, 119\(10\):e2109420119, 2022\.
- \[52\]Alexandre Galashov, Ashwini Pokle, Arnaud Doucet, Arthur Gretton, Mauricio Delbracio, and Valentin De Bortoli\.Learn to guide your diffusion model\.*arXiv preprint arXiv:2510\.00815*, 2025\.
- \[53\]Nate Gruver, Samuel Stanton, Nathan Frey, Tim GJ Rudner, Isidro Hotzel, Julien Lafrance\-Vanasse, Arvind Rajpal, Kyunghyun Cho, and Andrew G Wilson\.Protein design with guided discrete diffusion\.*Advances in Neural Information Processing Systems \(NeurIPS\)*, 36:12489–12517, 2023\.
- \[54\]Lin Gui, Cristina Gârbacea, and Victor Veitch\.BoNBoN alignment for large language models and the sweetness of best\-of\-n sampling\.*Advances in Neural Information Processing Systems \(NeurIPS\)*, 37:2851–2885, 2024\.
- \[55\]Wei Guo, Jaemoo Choi, Yuchen Zhu, Molei Tao, and Yongxin Chen\.Proximal diffusion neural sampler\.In*The Fourteenth International Conference on Learning Representations \(ICLR\)*, 2026\.
- \[56\]David Ha, Andrew M Dai, and Quoc V Le\.HyperNetworks\.In*The Fifth International Conference on Learning Representations \(ICLR\)*, 2017\.
- \[57\]Jiaqi Han, Austin Wang, Minkai Xu, Wenda Chu, Meihua Dang, Yisong Yue, and Stefano Ermon\.Discrete diffusion trajectory alignment via stepwise decomposition\.*arXiv preprint arXiv:2507\.04832*, 2025a\.
- \[58\]Yinbin Han, Meisam Razaviyayn, and Renyuan Xu\.Stochastic control for fine\-tuning diffusion models: Optimality, regularity, and convergence\.In*International Conference on Machine Learning \(ICML\)*, pages 21844–21870\. PMLR, 2025b\.
- \[59\]Mohsin Hasan, Viktor Ohanesian, Artem Gazizov, Yoshua Bengio, Alán Aspuru\-Guzik, Roberto Bondesan, Marta Skreta, and Kirill Neklyudov\.Discrete Feynman–Kac correctors\.*arXiv preprint arXiv:2601\.10403*, 2026\.
- \[60\]Aaron J Havens, Benjamin Kurt Miller, Bing Yan, Carles Domingo\-Enrich, Anuroop Sriram, Daniel S Levine, Brandon M Wood, Bin Hu, Brandon Amos, Brian Karrer, et al\.Adjoint sampling: Highly scalable diffusion samplers via adjoint matching\.In*International Conference on Machine Learning \(ICML\)*, pages 22204–22237\. PMLR, 2025\.
- \[61\]Jiajun He, José Miguel Hernández\-Lobato, Yuanqi Du, and Francisco Vargas\.RNE: a plug\-and\-play framework for diffusion density estimation and inference\-time control\.In*The Fourteenth International Conference on Learning Representations \(ICLR\)*, 2026\.
- \[62\]Ye He, Kevin Rojas, and Molei Tao\.What exactly does guidance do in masked discrete diffusion models\.*arXiv preprint arXiv:2506\.10971*, 2025\.
- \[63\]Jeremy Heng, Adrian N Bishop, George Deligiannidis, and Arnaud Doucet\.Controlled sequential Monte Carlo\.*The Annals of Statistics*, 48\(5\):2904–2929, 2020\.
- \[64\]Jonathan Ho, Ajay Jain, and Pieter Abbeel\.Denoising diffusion probabilistic models\.*Advances in Neural Information Processing Systems \(NeurIPS\)*, 33:6840–6851, 2020\.
- \[65\]Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al\.Training compute\-optimal large language models\.*Advances in Neural Information Processing Systems*, 35:30016–30030, 2022\.
- \[66\]Peter Holderrieth, Michael S Albergo, and Tommi Jaakkola\.LEAPS: A discrete neural sampler via locally equivariant networks\.*arXiv preprint arXiv:2502\.10843*, 2025a\.
- \[67\]Peter Holderrieth, Marton Havasi, Jason Yim, Neta Shaul, Itai Gat, Tommi Jaakkola, Brian Karrer, Ricky TQ Chen, and Yaron Lipman\.Generator matching: Generative modeling with arbitrary Markov processes\.In*The Thirteenth International Conference on Learning Representations \(ICLR\)*, 2025b\.
- \[68\]Peter Holderrieth, Douglas Chen, Luca Eyring, Ishin Shah, Giri Anantharaman, Yutong He, Zeynep Akata, Tommi Jaakkola, Nicholas Matthew Boffi, and Max Simchowitz\.Diamond maps: Efficient reward alignment via stochastic flow maps\.*arXiv preprint arXiv:2602\.05993*, 2026a\.
- \[69\]Peter Holderrieth, Uriel Singer, Tommi Jaakkola, Ricky TQ Chen, Yaron Lipman, and Brian Karrer\.GLASS flows: Transition sampling for alignment of flow and diffusion models\.In*The Fourteenth International Conference on Learning Representations \(ICLR\)*, 2026b\.
- \[70\]Audrey Huang, Adam Block, Qinghua Liu, Nan Jiang, Akshay Krishnamurthy, and Dylan J Foster\.Is best\-of\-N the best of them? Coverage, scaling, and optimality in inference\-time alignment\.In*International Conference on Machine Learning \(ICML\)*, pages 25075–25126\. PMLR, 2025\.
- \[71\]Baihe Huang, Shanda Li, Tianhao Wu, Yiming Yang, Ameet Talwalkar, Kannan Ramchandran, Michael I Jordan, and Jiantao Jiao\.Sample complexity and representation ability of test\-time scaling paradigms\.In*The Fourteenth International Conference on Learning Representations \(ICLR\)*, 2026\.
- \[72\]Yuichi Inoue, Kou Misaki, Yuki Imajuku, So Kuroki, Taishi Nakamura, and Takuya Akiba\.Wider or deeper? Scaling LLM inference\-time compute with adaptive branching tree search\.*arXiv preprint arXiv:2503\.04412*, 2025\.
- \[73\]Vineet Jain, Kusha Sareen, Mohammad Pedramfar, and Siamak Ravanbakhsh\.Diffusion tree sampling: Scalable inference\-time alignment of diffusion models\.*arXiv preprint arXiv:2506\.20701*, 2025\.
- \[74\]Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei\.Scaling laws for neural language models\.*arXiv preprint arXiv:2001\.08361*, 2020\.
- \[75\]Aayush Karan, Kulin Shah, and Sitan Chen\.ReGuidance: A simple diffusion wrapper for boosting sample quality on hard inverse problems\.*arXiv preprint arXiv:2506\.10955*, 2025\.
- \[76\]Thomas J Kerby and Kevin R Moon\.Training\-free guidance for discrete diffusion models for molecular generation\.*arXiv preprint arXiv:2409\.07359*, 2024\.
- \[77\]Hyunwoo Kim, Melanie Sclar, Tan Zhi\-Xuan, Lance Ying, Sydney Levine, Yang Liu, Joshua B Tenenbaum, and Yejin Choi\.Hypothesis\-driven theory\-of\-mind reasoning for large language models\.In*Second Conference on Language Modeling \(COLM\)*, 2025a\.
- \[78\]Sunwoo Kim, Minkyu Kim, and Dongmin Park\.Test\-time alignment of diffusion models without reward over\-optimization\.In*The Thirteenth International Conference on Learning Representations*, 2025b\.
- \[79\]Inception Labs, Samar Khanna, Siddhant Kharbanda, Shufan Li, Harshit Varma, Eric Wang, Sawyer Birnbaum, Ziyang Luo, Yanis Miraoui, Akash Palrecha, et al\.Mercury: Ultra\-fast language models based on diffusion\.*arXiv preprint arXiv:2506\.17298*, 2025\.
- \[80\]Dieterich Lawson, Allan Raventós, Andrew Warrington, and Scott Linderman\.SIXO: Smoothing inference with twisted objectives\.*Advances in Neural Information Processing Systems \(NeurIPS\)*, 35:38844–38858, 2022\.
- \[81\]Cheuk Kit Lee, Paul Jeha, Jes Frellsen, Pietro Lio, Michael S Albergo, and Francisco Vargas\.Debiasing guidance for discrete diffusion with sequential Monte Carlo\.*arXiv preprint arXiv:2502\.06079*, 2025\.
- \[82\]Alexander K Lew, George Matheos, Tan Zhi\-Xuan, Matin Ghavamizadeh, Nishad Gothoskar, Stuart Russell, and Vikash K Mansinghka\.SMCP3: Sequential Monte Carlo with probabilistic program proposals\.In*International Conference on Artificial Intelligence and Statistics \(AISTATS\)*, pages 7061–7088\. PMLR, 2023a\.
- \[83\]Alexander K Lew, Tan Zhi\-Xuan, Gabriel Grand, and Vikash K Mansinghka\.Sequential Monte Carlo steering of large language models using probabilistic programs\.*arXiv preprint arXiv:2306\.03081*, 2023b\.
- \[84\]Dacheng Li, Shiyi Cao, Chengkun Cao, Xiuyu Li, Shangyin Tan, Kurt Keutzer, Jiarong Xing, Joseph E Gonzalez, and Ion Stoica\.S\*: Test time scaling for code generation\.*arXiv preprint arXiv:2502\.14382*, 1\(2\), 2025a\.
- \[85\]Minhuan Li, Jiequn Han, Pilar Cossio, and Luhuan Wu\.Robust inference\-time steering of protein diffusion models via embedding optimization\.*arXiv preprint arXiv:2602\.05285*, 2026a\.
- \[86\]Shufan Li, Konstantinos Kallidromitis, Akash Gokul, Arsh Koneru, Yusuke Kato, Kazuki Kozuka, and Aditya Grover\.Reflect\-DiT: Inference\-time scaling for text\-to\-image diffusion transformers via in\-context reflection\.In*Proceedings of the IEEE/CVF International Conference on Computer Vision \(ICCV\)*, pages 15657–15668, 2025b\.
- \[87\]Shufan Li, Yuchen Zhu, Jiuxiang Gu, Kangning Liu, Zhe Lin, Yongxin Chen, Molei Tao, Aditya Grover, and Jason Kuen\.LaViDa\-R1: Advancing reasoning for unified multimodal diffusion language models\.*arXiv preprint arXiv:2602\.14147*, 2026b\.
- \[88\]Xiang Li, John Thickstun, Ishaan Gulrajani, Percy S Liang, and Tatsunori B Hashimoto\.Diffusion\-LM improves controllable text generation\.*Advances in Neural Information Processing Systems \(NeurIPS\)*, 35:4328–4343, 2022\.
- \[89\]Xiner Li, Yulai Zhao, Chenyu Wang, Gabriele Scalia, Gokcen Eraslan, Surag Nair, Tommaso Biancalani, Aviv Regev, Sergey Levine, and Masatoshi Uehara\.Derivative\-free guidance in continuous and discrete diffusion models with soft value\-based decoding\.*arXiv preprint arXiv:2408\.08252*, 2024\.
- \[90\]Shalev Lifshitz, Sheila A McIlraith, and Yilun Du\.Multi\-agent verification: Scaling test\-time compute with multiple verifiers\.In*Second Conference on Language Modeling \(COLM\)*, 2025\.
- \[91\]Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe\.Let’s verify step by step\.In*The Twelfth International Conference on Learning Representations \(ICLR\)*, 2023\.
- \[92\]Baijiong Lin, Weisen Jiang, Yuancheng Xu, Hao Chen, and Ying\-Cong Chen\.PARM: Multi\-objective test\-time alignment via preference\-aware autoregressive reward model\.*arXiv preprint arXiv:2505\.06274*, 2025a\.
- \[93\]Yujie Lin, Ante Wang, Moye Chen, Jingyao Liu, Hao Liu, Jinsong Su, and Xinyan Xiao\.Investigating inference\-time scaling for chain of multi\-modal thought: A preliminary study\.In*Findings of the Association for Computational Linguistics: ACL 2025*, pages 15654–15667, 2025b\.
- \[94\]Yaron Lipman, Ricky TQ Chen, Heli Ben\-Hamu, Maximilian Nickel, and Matthew Le\.Flow matching for generative modeling\.In*The Eleventh International Conference on Learning Representations \(ICLR\)*, 2023\.
- \[95\]Fangfu Liu, Hanyang Wang, Yimo Cai, Kaiyan Zhang, Xiaohang Zhan, and Yueqi Duan\.Video\-T1: Test\-time scaling for video generation\.In*Proceedings of the IEEE/CVF International Conference on Computer Vision \(ICCV\)*, pages 18671–18681, 2025\.
- \[96\]Xingchao Liu, Chengyue Gong, and Qiang Liu\.Flow straight and fast: Learning to generate and transfer data with rectified flow\.In*The Eleventh International Conference on Learning Representations \(ICLR\)*, 2023\.
- \[97\]Aaron Lou, Chenlin Meng, and Stefano Ermon\.Discrete diffusion modeling by estimating the ratios of the data distribution\.In*International Conference on Machine Learning \(ICML\)*, pages 32819–32848\. PMLR, 2024\.
- \[98\]João Loula, Benjamin LeBrun, Li Du, Ben Lipkin, Clemente Pasti, Gabriel Grand, Tianyu Liu, Yahya Emara, Marjorie Freedman, Jason Eisner, et al\.Syntactic and semantic control of large language models via sequential Monte Carlo\.In*The Thirteenth International Conference on Learning Representations \(ICLR\)*, 2025\.
- \[99\]Jianfeng Lu and Yuliang Wang\.Guidance for twisted particle filter: a continuous\-time perspective\.*arXiv preprint arXiv:2409\.02399*, 2024\.
- \[100\]Hao Luan, See\-Kiong Ng, and Chun Kai Ling\.DDPS: Discrete diffusion posterior sampling for paths in layered graphs\.*arXiv preprint arXiv:2504\.20754*, 2025\.
- \[101\]Nanye Ma, Shangyuan Tong, Haolin Jia, Hexiang Hu, Yu\-Chuan Su, Mingda Zhang, Xuan Yang, Yandong Li, Tommi Jaakkola, Xuhui Jia, et al\.Scaling inference time compute for diffusion models\.In*Proceedings of the Computer Vision and Pattern Recognition Conference \(CVPR\)*, pages 2523–2534, 2025\.
- \[102\]Arvind Mahankali, Kaiyue Wen, and Tengyu Ma\.Divide\-and\-conquer CoT: RL for reducing latency via parallel reasoning\.*arXiv preprint arXiv:2601\.23027*, 2026\.
- \[103\]Pierre Del Moral\.*Feynman–Kac formulae: genealogical and interacting particle systems with applications*\.Springer, 2004\.
- \[104\]Niklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li, Li Fei\-Fei, Hannaneh Hajishirzi, Luke Zettlemoyer, Percy Liang, Emmanuel Candès, and Tatsunori B Hashimoto\.s1: Simple test\-time scaling\.In*Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing \(EMNLP\)*, pages 20286–20332, 2025\.
- \[105\]Shen Nie, Fengqi Zhu, Chao Du, Tianyu Pang, Qian Liu, Guangtao Zeng, Min Lin, and Chongxuan Li\.Scaling up masked diffusion models on text\.In*The Thirteenth International Conference on Learning Representations*, 2025a\.
- \[106\]Shen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang, Jingyang Ou, Jun Hu, Jun Zhou, Yankai Lin, Ji\-Rong Wen, and Chongxuan Li\.Large language diffusion models\.*arXiv preprint arXiv:2502\.09992*, 2025b\.
- \[107\]Hunter Nisonoff, Junhao Xiong, Stephan Allenspach, and Jennifer Listgarten\.Unlocking guidance for discrete state\-space diffusion and flow models\.In*The Thirteenth International Conference on Learning Representations \(ICLR\)*, 2025\.
- \[108\]Zijing Ou, Ruixiang Zhang, and Yingzhen Li\.Discrete neural flow samplers with locally equivariant transformer\.*arXiv preprint arXiv:2505\.17741*, 2025\.
- \[109\]Zijing Ou, Chinmay Pani, and Yingzhen Li\.Inference\-time scaling of discrete diffusion models via importance weighting and optimal proposal design\.In*The Fourteenth International Conference on Learning Representations \(ICLR\)*, 2026\.
- \[110\]Krunoslav Lehman Pavasovic, Jakob Verbeek, Giulio Biroli, and Marc Mezard\.Classifier\-free guidance: From high\-dimensional analysis to generalized guidance forms\.*arXiv preprint arXiv:2502\.07849*, 2025\.
- \[111\]Isha Puri, Shivchander Sudalairaj, Guangxuan Xu, Kai Xu, and Akash Srivastava\.Rollout roulette: A probabilistic inference approach to inference\-time scaling of LLMs using particle\-based Monte Carlo methods\.*arXiv preprint arXiv:2502\.01618*, 2025\.
- \[112\]Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al\.Language models are unsupervised multitask learners\.*OpenAI blog*, 1\(8\):9, 2019\.
- \[113\]Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn\.Direct preference optimization: Your language model is secretly a reward model\.*Advances in Neural Information Processing Systems \(NeurIPS\)*, 36:53728–53741, 2023\.
- \[114\]Vignav Ramesh and Morteza Mardani\.Test\-time scaling of diffusion models via noise trajectory search\.*arXiv preprint arXiv:2506\.03164*, 2025\.
- \[115\]Pierre\-Arnaud Raviart\.An analysis of particle methods\.In*Numerical Methods in Fluid Dynamics: Lectures given at the 3rd 1983 Session of the Centro Internationale Matematico Estivo \(CIME\) held at Como, Italy, July 7–15, 1983*, pages 243–324\. Springer, 2006\.
- \[116\]Jarrid Rector\-Brooks, Mohsin Hasan, Zhangzhi Peng, Zachary Quinn, Chenghao Liu, Sarthak Mittal, Nouha Dziri, Michael Bronstein, Yoshua Bengio, Pranam Chatterjee, Alexander Tong, and Joey Bose\.Steering masked discrete diffusion models via discrete denoising posterior prediction\.In*The Thirteenth International Conference on Learning Representations \(ICLR\)*, 2025\.
- \[117\]Yinuo Ren, Haoxuan Chen, Grant Rotskoff, and Lexing Ying\.How discrete and continuous diffusion meet: Comprehensive analysis of discrete diffusion models via a stochastic integral framework\.In*International Conference on Learning Representations*, volume 2025, pages 42904–42941, 2025\.
- \[118\]Yinuo Ren, Haoxuan Chen, Yuchen Zhu, Wei Guo, Yongxin Chen, Grant Rotskoff, Molei Tao, and Lexing Ying\.Fast solvers for discrete diffusion models: Theory and applications of high\-order algorithms\.*Advances in Neural Information Processing Systems*, 38:167228–167282, 2026a\.
- \[119\]Yinuo Ren, Wenhao Gao, Lexing Ying, Grant Rotskoff, and Jiequn Han\.DriftLite: Lightweight drift control for inference\-time scaling of diffusion models\.In*International Conference on Learning Representations*, volume 2026, pages 76664–76702, 2026b\.
- \[120\]Yinuo Ren, Grant M Rotskoff, and Lexing Ying\.A unified approach to analysis and design of denoising Markov models\.*Journal of Machine Learning Research*, 27\(105\):1–69, 2026c\.
- \[121\]Sergej Rjasanow and Wolfgang Wagner\.A stochastic weighted particle method for the Boltzmann equation\.*Journal of Computational Physics*, 124\(2\):243–253, 1996\.
- \[122\]Kevin Rojas, Yuchen Zhu, Sichen Zhu, Felix XF Ye, and Molei Tao\.Diffuse everything: Multimodal diffusion models on arbitrary state spaces\.In*International Conference on Machine Learning \(ICML\)*, pages 51924–51956\. PMLR, 2025\.
- \[123\]Kevin Rojas, Ye He, Chieh\-Hsin Lai, Yuhta Takida, Yuki Mitsufuji, and Molei Tao\.Improving classifier\-free guidance in masked diffusion: Low\-dim theoretical insights with high\-dim impact\.In*The Fourteenth International Conference on Learning Representations \(ICLR\)*, 2026\.
- \[124\]Daan Roos, Oscar Davis, Floor Eijkelboom, Michael Bronstein, Max Welling, İsmail İlkan Ceylan, Luca Ambrogioni, and Jan\-Willem van de Meent\.Categorical flow maps\.*arXiv preprint arXiv:2602\.12233*, 2026\.
- \[125\]Litu Rout, Andreas Lugmayr, Yasamin Jafarian, Srivatsan Varadharajan, Constantine Caramanis, Sanjay Shakkottai, and Ira Kemelmacher\-Shlizerman\.Test\-time anchoring for discrete diffusion posterior sampling\.*arXiv preprint arXiv:2510\.02291*, 2025\.
- \[126\]Amirmojtaba Sabour, Michael S Albergo, Carles Domingo\-Enrich, Nicholas M Boffi, Sanja Fidler, Karsten Kreis, and Eric Vanden\-Eijnden\.Test\-time scaling of diffusions with flow maps\.*arXiv preprint arXiv:2511\.22688*, 2025\.
- \[127\]Subham Sahoo, Marianne Arriola, Yair Schiff, Aaron Gokaslan, Edgar Marroquin, Justin Chiu, Alexander Rush, and Volodymyr Kuleshov\.Simple and effective masked diffusion language models\.*Advances in Neural Information Processing Systems \(NeurIPS\)*, 37:130136–130184, 2024\.
- \[128\]Yair Schiff, Subham Sekhar Sahoo, Hao Phung, Guanghan Wang, Sam Boshar, Hugo Dalla\-torre, Bernardo P de Almeida, Alexander Rush, Thomas Pierrot, and Volodymyr Kuleshov\.Simple guidance mechanisms for discrete diffusion models\.In*The Thirteenth International Conference on Learning Representations \(ICLR\)*, 2025\.
- \[129\]Nikolaus Schweizer\.Non\-asymptotic error bounds for sequential MCMC and stability of Feynman–Kac propagators\.*arXiv preprint arXiv:1204\.2382*, 2012a\.
- \[130\]Nikolaus Schweizer\.*Non\-asymptotic error bounds for sequential MCMC methods*\.PhD thesis, Universitäts\-und Landesbibliothek Bonn, 2012b\.
- \[131\]Louis Serrano, Jiequn Han, Edouard Oyallon, Shirley Ho, and Rudy Morel\.Test\-time generalization for physics through neural operator splitting\.*arXiv preprint arXiv:2602\.00884*, 2026\.
- \[132\]Amrith Setlur, Nived Rajaraman, Sergey Levine, and Aviral Kumar\.Scaling test\-time compute without verification or RL is suboptimal\.In*International Conference on Machine Learning \(ICML\)*, pages 54058–54094\. PMLR, 2025\.
- \[133\]Raghav Singhal, Zachary Horvitz, Ryan Teehan, Mengye Ren, Zhou Yu, Kathleen Mckeown, and Rajesh Ranganath\.A general framework for inference\-time scaling and steering of diffusion models\.In*International Conference on Machine Learning \(ICML\)*, pages 55810–55827\. PMLR, 2025\.
- \[134\]Marta Skreta, Tara Akhound\-Sadegh, Viktor Ohanesian, Roberto Bondesan, Alan Aspuru\-Guzik, Arnaud Doucet, Rob Brekelmans, Alexander Tong, and Kirill Neklyudov\.Feynman–Kac correctors in diffusion: Annealing, guidance, and product of experts\.In*International Conference on Machine Learning \(ICML\)*, pages 55906–55949\. PMLR, 2025a\.
- \[135\]Marta Skreta, Lazar Atanackovic, Joey Bose, Alexander Tong, and Kirill Neklyudov\.The superposition of diffusion models using the Itô density estimator\.In*The Thirteenth International Conference on Learning Representations \(ICLR\)*, 2025b\.
- \[136\]Charlie Victor Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar\.Scaling LLM test\-time compute optimally can be more effective than scaling parameters for reasoning\.In*The Thirteenth International Conference on Learning Representations \(ICLR\)*, 2025\.
- \[137\]Oswin So, Brian Karrer, Chuchu Fan, Ricky TQ Chen, and Guan\-Horng Liu\.Discrete adjoint matching\.In*The Fourteenth International Conference on Learning Representations \(ICLR\)*, 2026\.
- \[138\]Jascha Sohl\-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli\.Deep unsupervised learning using nonequilibrium thermodynamics\.In*International Conference on Machine Learning \(ICML\)*, pages 2256–2265\. PMLR, 2015\.
- \[139\]Jiaming Song and Linqi Zhou\.Ideas in inference\-time scaling can benefit generative pre\-training algorithms\.*arXiv preprint arXiv:2503\.07154*, 2025\.
- \[140\]Jiaming Song, Chenlin Meng, and Stefano Ermon\.Denoising diffusion implicit models\.In*International Conference on Learning Representations \(ICLR\)*, 2021a\.
- \[141\]Yang Song and Stefano Ermon\.Generative modeling by estimating gradients of the data distribution\.*Advances in Neural Information Processing Systems \(NeurIPS\)*, 32, 2019\.
- \[142\]Yang Song and Stefano Ermon\.Improved techniques for training score\-based generative models\.*Advances in Neural Information Processing Systems \(NeurIPS\)*, 33:12438–12448, 2020\.
- \[143\]Yang Song, Conor Durkan, Iain Murray, and Stefano Ermon\.Maximum likelihood training of score\-based diffusion models\.*Advances in Neural Information Processing Systems \(NeurIPS\)*, 34:1415–1428, 2021b\.
- \[144\]Yang Song, Jascha Sohl\-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole\.Score\-based generative modeling through stochastic differential equations\.In*The Ninth International Conference on Learning Representations \(ICLR\)*, 2021c\.
- \[145\]Yuxuan Song, Zheng Zhang, Cheng Luo, Pengyang Gao, Fan Xia, Hao Luo, Zheng Li, Yuehang Yang, Hongli Yu, Xingwei Qu, et al\.Seed diffusion: A large\-scale diffusion language model with high\-speed inference\.*arXiv preprint arXiv:2508\.02193*, 2025\.
- \[146\]Robert H Swendsen and Jian\-Sheng Wang\.Nonuniversal critical dynamics in Monte Carlo simulations\.*Physical review letters*, 58\(2\):86, 1987\.
- \[147\]Denis Talay and Olivier Vaillant\.A stochastic particle method with random weights for the computation of statistical solutions of McKean–Vlasov equations\.*The Annals of Applied Probability*, 13\(1\):140–180, 2003\.
- \[148\]Sophia Tang, Yuchen Zhu, Molei Tao, and Pranam Chatterjee\.TR2\-D2: Tree search guided trajectory\-aware fine\-tuning for discrete diffusion\.*arXiv preprint arXiv:2509\.25171*, 2025\.
- \[149\]Wenpin Tang and Renyuan Xu\.A stochastic analysis approach to conditional diffusion guidance\.*Columbia University Preprint*, 2024\.
- \[150\]Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean\-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al\.Gemini: a family of highly capable multimodal models\.*arXiv preprint arXiv:2312\.11805*, 2023\.
- \[151\]Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie\-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al\.LLaMA: Open and efficient foundation language models\.*arXiv preprint arXiv:2302\.13971*, 2023\.
- \[152\]Masatoshi Uehara, Yulai Zhao, Chenyu Wang, Xiner Li, Aviv Regev, Sergey Levine, and Tommaso Biancalani\.Inference\-time alignment in diffusion models with reward\-guided generation: Tutorial and review\.*arXiv preprint arXiv:2501\.09685*, 2025\.
- \[153\]Clement Vignac, Igor Krawczuk, Antoine Siraudin, Bohan Wang, Volkan Cevher, and Pascal Frossard\.DiGress: Discrete denoising diffusion for graph generation\.In*The Eleventh International Conference on Learning Representations \(ICLR\)*, 2023\.
- \[154\]Bram Wallace, Meihua Dang, Rafael Rafailov, Linqi Zhou, Aaron Lou, Senthil Purushwalkam, Stefano Ermon, Caiming Xiong, Shafiq Joty, and Nikhil Naik\.Diffusion model alignment using direct preference optimization\.In*Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\)*, pages 8228–8238, 2024\.
- \[155\]Chenyang Wang, Weizhong Wang, Yinuo Ren, Jose Blanchet, and Yiping Lu\.Simple approximation and derivative free inference\-time scaling for diffusion models via sequential Monte Carlo on path measures\.*arXiv preprint arXiv:2605\.17850*, 2026\.
- \[156\]Chenyu Wang, Masatoshi Uehara, Yichun He, Amy Wang, Avantika Lal, Tommi Jaakkola, Sergey Levine, Aviv Regev, Tommaso Biancalani, et al\.Fine\-tuning discrete diffusion models via reward optimization with applications to DNA and protein design\.In*The Thirteenth International Conference on Learning Representations \(ICLR\)*, 2025a\.
- \[157\]Junlin Wang, Shang Zhu, Jon Saad\-Falcon, Ben Athiwaratkun, Qingyang Wu, Jue Wang, Shuaiwen Leon Song, Ce Zhang, Bhuwan Dhingra, and James Zou\.Think deep, think fast: Investigating efficiency of verifier\-free inference\-time\-scaling methods\.*arXiv preprint arXiv:2504\.14047*, 2025b\.
- \[158\]Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou\.Self\-consistency improves chain of thought reasoning in language models\.In*The Eleventh International Conference on Learning Representations \(ICLR\)*, 2023\.
- \[159\]Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al\.Chain\-of\-thought prompting elicits reasoning in large language models\.*Advances in Neural Information Processing Systems \(NeurIPS\)*, 35:24824–24837, 2022\.
- \[160\]Lifu Wei, Yinuo Ren, Naichen Shi, and Yiping Lu\.SURGE: Approximation\-free training free particle filter for diffusion surrogate\.*arXiv preprint arXiv:2605\.18745*, 2026\.
- \[161\]Nick Whiteley and Anthony Lee\.Twisted particle filters\.*The Annals of Statistics*, 42\(1\):115–141, 2014\.
- \[162\]Luhuan Wu, Yi Han, Christian A Naesseth, and John P Cunningham\.Reverse diffusion sequential Monte Carlo samplers\.*arXiv preprint arXiv:2508\.05926*, 2025a\.
- \[163\]Xiaoshi Wu, Yiming Hao, Keqiang Sun, Yixiong Chen, Feng Zhu, Rui Zhao, and Hongsheng Li\.Human preference score v2: A solid benchmark for evaluating human preferences of text\-to\-image synthesis\.*arXiv preprint arXiv:2306\.09341*, 2023\.
- \[164\]Yangzhen Wu, Zhiqing Sun, Shanda Li, Sean Welleck, and Yiming Yang\.Inference scaling laws: An empirical analysis of compute\-optimal inference for LLM problem\-solving\.In*The Thirteenth International Conference on Learning Representations \(ICLR\)*, 2025b\.
- \[165\]Yuchen Wu, Minshuo Chen, Zihao Li, Mengdi Wang, and Yuting Wei\.Theoretical insights for diffusion guidance: A case study for Gaussian mixture models\.*arXiv preprint arXiv:2403\.01639*, 2024\.
- \[166\]Xin Xie, Jiaxian Guo, and Dong Gong\.HyperAlign: Hypernetwork for efficient test\-time alignment of diffusion models\.*arXiv preprint arXiv:2601\.15968*, 2026\.
- \[167\]Yi Xin, Siqi Luo, Qi Qin, Haoxing Chen, Kaiwen Zhu, Zhiwei Zhang, Yangfan He, Rongchao Zhang, Jinbin Bai, Shuo Cao, et al\.dMLLM\-TTS: Self\-verified and efficient test\-time scaling for diffusion multi\-modal large language models\.*arXiv preprint arXiv:2512\.19433*, 2025\.
- \[168\]Junhao Xiong, Hunter Nisonoff, Maria Lukarska, Ishan Gaur, Luke M Oltrogge, David F Savage, and Jennifer Listgarten\.Guide your favorite protein sequence generative model\.*arXiv preprint arXiv:2505\.04823*, 2025\.
- \[169\]Ziang Yan, Xinhao Li, Yinan He, Zhengrong Yue, Xiangyu Zeng, Yali Wang, Yu Qiao, Limin Wang, and Yi Wang\.VideoChat\-R1\.5: Visual test\-time scaling to reinforce multimodal reasoning by iterative perception\.*arXiv preprint arXiv:2509\.21100*, 2025\.
- \[170\]Haolin Yang, Feilong Tang, Ming Hu, Qingyu Yin, Yulong Li, Yexin Liu, Zelin Peng, Peng Gao, Junjun He, Zongyuan Ge, et al\.ScalingNoise: Scaling inference\-time search for generating infinite videos\.*arXiv preprint arXiv:2503\.16400*, 2025a\.
- \[171\]Jason Yang, Wenda Chu, Daniel Khalil, Raul Astudillo, Bruce J Wittmann, Frances H Arnold, and Yisong Yue\.Steering generative models with experimental data for protein fitness optimization\.*arXiv preprint arXiv:2505\.15093*, 2025b\.
- \[172\]Yanfeng Yang and Kenji Fukumizu\.Diffusion models with double guidance: Generate with aggregated datasets\.*arXiv preprint arXiv:2505\.13213*, 2025\.
- \[173\]Zheyuan Yang, Lyuhao Chen, Arman Cohan, and Yilun Zhao\.Table\-R1: Inference\-time scaling for table reasoning tasks\.In*Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing \(EMNLP\)*, pages 20616–20635, 2025c\.
- \[174\]Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan\.Tree of thoughts: Deliberate problem solving with large language models\.*Advances in neural information processing systems*, 36:11809–11822, 2023\.
- \[175\]Haotian Ye, Haowei Lin, Jiaqi Han, Minkai Xu, Sheng Liu, Yitao Liang, Jianzhu Ma, James Y Zou, and Stefano Ermon\.TFG: Unified training\-free guidance for diffusion models\.*Advances in Neural Information Processing Systems \(NeurIPS\)*, 37:22370–22417, 2024\.
- \[176\]Taehoon Yoon, Yunhong Min, Kyeongmin Yeo, and Minhyuk Sung\.Psi\-Sampler: Initial particle sampling for SMC\-based inference\-time reward alignment in score models\.*arXiv preprint arXiv:2506\.01320*, 2025\.
- \[177\]Yue Yu, Qiwei Di, Quanquan Gu, and Dongruo Zhou\.On the limits of test\-time compute: Sequential reward filtering for better inference\.*arXiv preprint arXiv:2512\.04558*, 2025\.
- \[178\]Linfeng Zhang, Weinan E, and Lei Wang\.Monge–Ampère flow for generative modeling\.*arXiv preprint arXiv:1809\.10188*, 2018\.
- \[179\]Qiyuan Zhang, Fuyuan Lyu, Zexu Sun, Lei Wang, Weixu Zhang, Wenyue Hua, Haolun Wu, Zhihan Guo, Yufei Wang, Niklas Muennighoff, et al\.A survey on test\-time scaling in large language models: What, how, where, and how well?*arXiv preprint arXiv:2503\.24235*, 2025\.
- \[180\]Sixian Zhang, Bohan Wang, Junqiang Wu, Yan Li, Tingting Gao, Di Zhang, and Zhongyuan Wang\.Learning multi\-dimensional human preference for text\-to\-image generation\.In*Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition*, pages 8018–8027, 2024\.
- \[181\]Xiangcheng Zhang, Haowei Lin, Haotian Ye, James Zou, Jianzhu Ma, Yitao Liang, and Yilun Du\.Inference\-time scaling of diffusion models through classical search\.In*The Fourteenth International Conference on Learning Representations \(ICLR\)*, 2026\.
- \[182\]Siyan Zhao, Devaansh Gupta, Qinqing Zheng, and Aditya Grover\.d1: Scaling reasoning in diffusion large language models via reinforcement learning\.*arXiv preprint arXiv:2504\.12216*, 2025\.
- \[183\]Stephen Zhao, Rob Brekelmans, Alireza Makhzani, and Roger Baker Grosse\.Probabilistic inference in language models via twisted sequential Monte Carlo\.In*International Conference on Machine Learning \(ICML\)*, pages 60704–60748\. PMLR, 2024\.
- \[184\]Kaiwen Zheng, Yongxin Chen, Huayu Chen, Guande He, Ming\-Yu Liu, Jun Zhu, and Qinsheng Zhang\.Direct discriminative optimization: Your likelihood\-based visual generative model is secretly a GAN discriminator\.In*International Conference on Machine Learning \(ICML\)*, pages 78067–78094\. PMLR, 2025\.
- \[185\]Chunting Zhou, Lili Yu, Arun Babu, Kushal Tirumala, Michihiro Yasunaga, Leonid Shamis, Jacob Kahn, Xuezhe Ma, Luke Zettlemoyer, and Omer Levy\.Transfusion: Predict the next token and diffuse images with one multi\-modal model\.In*The Thirteenth International Conference on Learning Representations \(ICLR\)*, 2025\.
- \[186\]Qijie Zhu, Zeqi Ye, Han Liu, Zhaoran Wang, and Minshuo Chen\.Training\-free adaptation of diffusion models via Doob’shh\-transform\.*arXiv preprint arXiv:2602\.16198*, 2026a\.
- \[187\]Xuekai Zhu, Daixuan Cheng, Dinghuai Zhang, Hengli Li, Kaiyan Zhang, Che Jiang, Youbang Sun, Ermo Hua, Yuxin Zuo, Xingtai Lv, et al\.FlowRL: Matching reward distributions for LLM reasoning\.In*The Fourteenth International Conference on Learning Representations \(ICLR\)*, 2026b\.
- \[188\]Youheng Zhu and Yiping Lu\.On the power of \(approximate\) reward models for inference\-time scaling\.*arXiv preprint arXiv:2602\.01381*, 2026\.
- \[189\]Yuchen Zhu, Wei Guo, Jaemoo Choi, Guan\-Horng Liu, Yongxin Chen, and Molei Tao\.MDNS: Masked diffusion neural sampler via stochastic optimal control\.*arXiv preprint arXiv:2508\.10684*, 2025\.
- \[190\]Andrea Zoia, Eric Dumonteil, and Alain Mazzolo\.Discrete Feynman–Kac formulas for branching random walks\.*Europhysics Letters*, 98\(4\):40012, 2012\.
## Appendix ARelated Work
### A\.1Markov generative models
Modern generative modeling is shaped by two large Markovian families\. Autoregressive LLMs generate discrete sequences through one\-step conditional transitions\[[112](https://arxiv.org/html/2609.35947#bib.bib112),[18](https://arxiv.org/html/2609.35947#bib.bib18),[1](https://arxiv.org/html/2609.35947#bib.bib1),[150](https://arxiv.org/html/2609.35947#bib.bib150),[151](https://arxiv.org/html/2609.35947#bib.bib151)\]\. Diffusion and flow\-based models generate through iterative stochastic or deterministic dynamics\[[138](https://arxiv.org/html/2609.35947#bib.bib138),[178](https://arxiv.org/html/2609.35947#bib.bib178),[141](https://arxiv.org/html/2609.35947#bib.bib141),[64](https://arxiv.org/html/2609.35947#bib.bib64),[142](https://arxiv.org/html/2609.35947#bib.bib142),[144](https://arxiv.org/html/2609.35947#bib.bib144),[140](https://arxiv.org/html/2609.35947#bib.bib140),[143](https://arxiv.org/html/2609.35947#bib.bib143),[96](https://arxiv.org/html/2609.35947#bib.bib96),[94](https://arxiv.org/html/2609.35947#bib.bib94),[2](https://arxiv.org/html/2609.35947#bib.bib2),[5](https://arxiv.org/html/2609.35947#bib.bib5)\]\. Discrete diffusion and DLMs connect these viewpoints by placing diffusion\-style reverse dynamics on finite or countable state spaces\. Representative examples include diffusion\-LM\[[88](https://arxiv.org/html/2609.35947#bib.bib88)\], BERT\-style masking models\[[41](https://arxiv.org/html/2609.35947#bib.bib41),[22](https://arxiv.org/html/2609.35947#bib.bib22)\], discrete diffusion models\[[97](https://arxiv.org/html/2609.35947#bib.bib97),[105](https://arxiv.org/html/2609.35947#bib.bib105)\], and recent DLMs\[[127](https://arxiv.org/html/2609.35947#bib.bib127),[106](https://arxiv.org/html/2609.35947#bib.bib106),[12](https://arxiv.org/html/2609.35947#bib.bib12),[145](https://arxiv.org/html/2609.35947#bib.bib145),[79](https://arxiv.org/html/2609.35947#bib.bib79),[34](https://arxiv.org/html/2609.35947#bib.bib34),[26](https://arxiv.org/html/2609.35947#bib.bib26)\]\. FluxLite focuses on the CTMC structure of these reverse dynamics and on how a Feynman–Kac representation can be controlled without changing the pretrained graph\. Broader background on generator matching, mixed state spaces, and discrete diffusion surveys appears in\[[67](https://arxiv.org/html/2609.35947#bib.bib67),[120](https://arxiv.org/html/2609.35947#bib.bib120),[122](https://arxiv.org/html/2609.35947#bib.bib122),[24](https://arxiv.org/html/2609.35947#bib.bib24),[117](https://arxiv.org/html/2609.35947#bib.bib117),[118](https://arxiv.org/html/2609.35947#bib.bib118)\]\.
### A\.2Inference\-time scaling
##### Length and width\.
Existing inference\-time scaling methods can be grouped into two categories based on how they allocate additional compute at inference time\. The first class is*length scaling*, which improves the quality of the output by elongating the response or iteratively refining it\[[159](https://arxiv.org/html/2609.35947#bib.bib159),[158](https://arxiv.org/html/2609.35947#bib.bib158),[174](https://arxiv.org/html/2609.35947#bib.bib174)\]\. The second class is*width scaling*, which expands the hypothesis space by exploring a larger set of candidate trajectories via search\-based or sampling\-based methods\[[17](https://arxiv.org/html/2609.35947#bib.bib17),[136](https://arxiv.org/html/2609.35947#bib.bib136),[43](https://arxiv.org/html/2609.35947#bib.bib43),[164](https://arxiv.org/html/2609.35947#bib.bib164),[111](https://arxiv.org/html/2609.35947#bib.bib111),[152](https://arxiv.org/html/2609.35947#bib.bib152),[73](https://arxiv.org/html/2609.35947#bib.bib73),[114](https://arxiv.org/html/2609.35947#bib.bib114),[72](https://arxiv.org/html/2609.35947#bib.bib72),[148](https://arxiv.org/html/2609.35947#bib.bib148),[181](https://arxiv.org/html/2609.35947#bib.bib181),[10](https://arxiv.org/html/2609.35947#bib.bib10)\]\. Within sampling\-based width scaling, current work casts the problem as drawing samples from a tilted version of the distribution associated with the base model\.
##### Domains\.
The*inference\-time scaling*paradigm has been extended from text generation to a wide range of domains: reasoning\[[11](https://arxiv.org/html/2609.35947#bib.bib11),[173](https://arxiv.org/html/2609.35947#bib.bib173),[157](https://arxiv.org/html/2609.35947#bib.bib157)\], code generation\[[84](https://arxiv.org/html/2609.35947#bib.bib84),[40](https://arxiv.org/html/2609.35947#bib.bib40)\], AI for science\[[171](https://arxiv.org/html/2609.35947#bib.bib171),[49](https://arxiv.org/html/2609.35947#bib.bib49),[131](https://arxiv.org/html/2609.35947#bib.bib131),[85](https://arxiv.org/html/2609.35947#bib.bib85)\], and video\[[95](https://arxiv.org/html/2609.35947#bib.bib95),[170](https://arxiv.org/html/2609.35947#bib.bib170)\]or image generation\[[101](https://arxiv.org/html/2609.35947#bib.bib101),[25](https://arxiv.org/html/2609.35947#bib.bib25),[86](https://arxiv.org/html/2609.35947#bib.bib86)\]\. The inference\-time scaling literature for diffusion models specifically\[[152](https://arxiv.org/html/2609.35947#bib.bib152),[101](https://arxiv.org/html/2609.35947#bib.bib101),[25](https://arxiv.org/html/2609.35947#bib.bib25),[119](https://arxiv.org/html/2609.35947#bib.bib119),[59](https://arxiv.org/html/2609.35947#bib.bib59)\]forms the most direct context for our work\.
##### SMC methods\.
SMC\-based methods have recently gained increasing attention for inference\-time scaling on distributions of both continuous\[[133](https://arxiv.org/html/2609.35947#bib.bib133),[25](https://arxiv.org/html/2609.35947#bib.bib25),[78](https://arxiv.org/html/2609.35947#bib.bib78),[134](https://arxiv.org/html/2609.35947#bib.bib134),[176](https://arxiv.org/html/2609.35947#bib.bib176),[61](https://arxiv.org/html/2609.35947#bib.bib61),[160](https://arxiv.org/html/2609.35947#bib.bib160),[155](https://arxiv.org/html/2609.35947#bib.bib155)\]and discrete\[[83](https://arxiv.org/html/2609.35947#bib.bib83),[82](https://arxiv.org/html/2609.35947#bib.bib82),[98](https://arxiv.org/html/2609.35947#bib.bib98),[77](https://arxiv.org/html/2609.35947#bib.bib77),[59](https://arxiv.org/html/2609.35947#bib.bib59),[8](https://arxiv.org/html/2609.35947#bib.bib8)\]type\. Theoretical analyses of LLM inference\-time compute\[[17](https://arxiv.org/html/2609.35947#bib.bib17),[136](https://arxiv.org/html/2609.35947#bib.bib136),[104](https://arxiv.org/html/2609.35947#bib.bib104),[132](https://arxiv.org/html/2609.35947#bib.bib132),[70](https://arxiv.org/html/2609.35947#bib.bib70),[71](https://arxiv.org/html/2609.35947#bib.bib71),[42](https://arxiv.org/html/2609.35947#bib.bib42),[177](https://arxiv.org/html/2609.35947#bib.bib177),[188](https://arxiv.org/html/2609.35947#bib.bib188),[111](https://arxiv.org/html/2609.35947#bib.bib111)\]and the classical theory of SMC\[[129](https://arxiv.org/html/2609.35947#bib.bib129),[130](https://arxiv.org/html/2609.35947#bib.bib130),[30](https://arxiv.org/html/2609.35947#bib.bib30),[103](https://arxiv.org/html/2609.35947#bib.bib103),[38](https://arxiv.org/html/2609.35947#bib.bib38),[45](https://arxiv.org/html/2609.35947#bib.bib45),[37](https://arxiv.org/html/2609.35947#bib.bib37)\]provide further context\. Applications include LLMs and math reasoning\[[83](https://arxiv.org/html/2609.35947#bib.bib83),[183](https://arxiv.org/html/2609.35947#bib.bib183),[50](https://arxiv.org/html/2609.35947#bib.bib50),[133](https://arxiv.org/html/2609.35947#bib.bib133),[182](https://arxiv.org/html/2609.35947#bib.bib182)\], inference\-time alignment of continuous and discrete diffusion models\[[152](https://arxiv.org/html/2609.35947#bib.bib152),[109](https://arxiv.org/html/2609.35947#bib.bib109),[69](https://arxiv.org/html/2609.35947#bib.bib69)\], sampling from unnormalized densities\[[162](https://arxiv.org/html/2609.35947#bib.bib162)\], unified dynamics\[[134](https://arxiv.org/html/2609.35947#bib.bib134)\]and Bayesian inverse problems\[[25](https://arxiv.org/html/2609.35947#bib.bib25)\], model composition\[[46](https://arxiv.org/html/2609.35947#bib.bib46),[135](https://arxiv.org/html/2609.35947#bib.bib135)\], and inference\-time alignment in biology\[[85](https://arxiv.org/html/2609.35947#bib.bib85)\]\. See also the stochastic weighted particle method\[[35](https://arxiv.org/html/2609.35947#bib.bib35),[36](https://arxiv.org/html/2609.35947#bib.bib36),[121](https://arxiv.org/html/2609.35947#bib.bib121),[13](https://arxiv.org/html/2609.35947#bib.bib13),[147](https://arxiv.org/html/2609.35947#bib.bib147),[115](https://arxiv.org/html/2609.35947#bib.bib115),[28](https://arxiv.org/html/2609.35947#bib.bib28)\], discrete posterior sampling\[[31](https://arxiv.org/html/2609.35947#bib.bib31),[100](https://arxiv.org/html/2609.35947#bib.bib100),[125](https://arxiv.org/html/2609.35947#bib.bib125),[57](https://arxiv.org/html/2609.35947#bib.bib57)\], and inference\-time scaling of DLLMs via Gibbs sampling\[[33](https://arxiv.org/html/2609.35947#bib.bib33)\]\.
##### Guidance and search\.
Closely related to inference\-time tilting is the literature on continuous guidance\[[175](https://arxiv.org/html/2609.35947#bib.bib175)\], discrete guidance\[[153](https://arxiv.org/html/2609.35947#bib.bib153),[53](https://arxiv.org/html/2609.35947#bib.bib53),[89](https://arxiv.org/html/2609.35947#bib.bib89),[76](https://arxiv.org/html/2609.35947#bib.bib76),[107](https://arxiv.org/html/2609.35947#bib.bib107),[128](https://arxiv.org/html/2609.35947#bib.bib128),[168](https://arxiv.org/html/2609.35947#bib.bib168)\], and theoretical analyses of both continuous and discrete guidance\[[165](https://arxiv.org/html/2609.35947#bib.bib165),[29](https://arxiv.org/html/2609.35947#bib.bib29),[14](https://arxiv.org/html/2609.35947#bib.bib14),[110](https://arxiv.org/html/2609.35947#bib.bib110),[75](https://arxiv.org/html/2609.35947#bib.bib75),[172](https://arxiv.org/html/2609.35947#bib.bib172),[58](https://arxiv.org/html/2609.35947#bib.bib58),[149](https://arxiv.org/html/2609.35947#bib.bib149),[7](https://arxiv.org/html/2609.35947#bib.bib7),[123](https://arxiv.org/html/2609.35947#bib.bib123),[62](https://arxiv.org/html/2609.35947#bib.bib62)\]\. Discrete Feynman–Kac formulas\[[190](https://arxiv.org/html/2609.35947#bib.bib190),[15](https://arxiv.org/html/2609.35947#bib.bib15),[39](https://arxiv.org/html/2609.35947#bib.bib39)\], discrete tree search\[[148](https://arxiv.org/html/2609.35947#bib.bib148),[10](https://arxiv.org/html/2609.35947#bib.bib10)\], continuous tree sampling/search\[[73](https://arxiv.org/html/2609.35947#bib.bib73),[114](https://arxiv.org/html/2609.35947#bib.bib114),[181](https://arxiv.org/html/2609.35947#bib.bib181)\], and search\-based methods in LLMs\[[72](https://arxiv.org/html/2609.35947#bib.bib72)\]are also relevant\.
##### Closest concurrent work\.
The closest discrete Feynman–Kac papers are Discrete Feynman–Kac correctors and debiasing guidance methods for discrete diffusion\[[59](https://arxiv.org/html/2609.35947#bib.bib59),[81](https://arxiv.org/html/2609.35947#bib.bib81)\]\. These works use particle corrections to target tilted discrete laws and are therefore the nearest methodological neighbors\. FluxLite differs in the control object: it changes CTMC proposal rates inside an exact Feynman–Kac equivalence class and compensates the change by a graph\-divergence potential\. The target preservation is algebraic, while the approximation lies in the low\-dimensional variance\-control problem used to choose the representative\. In contrast, posterior\-prediction and steering methods for masked discrete diffusion\[[116](https://arxiv.org/html/2609.35947#bib.bib116),[156](https://arxiv.org/html/2609.35947#bib.bib156),[47](https://arxiv.org/html/2609.35947#bib.bib47)\]primarily modify denoising predictions or guidance rules, and hypernetwork approaches\[[166](https://arxiv.org/html/2609.35947#bib.bib166),[56](https://arxiv.org/html/2609.35947#bib.bib56)\]amortize adaptation into additional parameters\. Continuous proposal\-control and density\-estimation methods\[[119](https://arxiv.org/html/2609.35947#bib.bib119),[61](https://arxiv.org/html/2609.35947#bib.bib61),[109](https://arxiv.org/html/2609.35947#bib.bib109)\]share the goal of reducing SMC degeneracy, but their controls act through continuous drifts or learned continuous proposals rather than sparse graph flux\.
### A\.3Fine\-tuning and alignment
Training\-time alignment methods modify model parameters or auxiliary guidance networks, for example through RL\-style objectives for LLMs\[[187](https://arxiv.org/html/2609.35947#bib.bib187)\], DPO and diffusion\-DPO variants\[[113](https://arxiv.org/html/2609.35947#bib.bib113),[154](https://arxiv.org/html/2609.35947#bib.bib154)\], direct diffusion optimization\[[184](https://arxiv.org/html/2609.35947#bib.bib184)\], and learned guidance\[[52](https://arxiv.org/html/2609.35947#bib.bib52)\]\. A related line learns samplers, proposals, or control potentials for Monte Carlo and diffusion dynamics\[[51](https://arxiv.org/html/2609.35947#bib.bib51),[21](https://arxiv.org/html/2609.35947#bib.bib21),[3](https://arxiv.org/html/2609.35947#bib.bib3),[108](https://arxiv.org/html/2609.35947#bib.bib108),[189](https://arxiv.org/html/2609.35947#bib.bib189),[55](https://arxiv.org/html/2609.35947#bib.bib55),[137](https://arxiv.org/html/2609.35947#bib.bib137),[48](https://arxiv.org/html/2609.35947#bib.bib48),[60](https://arxiv.org/html/2609.35947#bib.bib60)\], including Doob\-hh\-transform and bridge\-based constructions\[[19](https://arxiv.org/html/2609.35947#bib.bib19),[23](https://arxiv.org/html/2609.35947#bib.bib23),[186](https://arxiv.org/html/2609.35947#bib.bib186),[32](https://arxiv.org/html/2609.35947#bib.bib32)\]\. FluxLite is complementary: it keeps the pretrained discrete diffusion model fixed and optimizes only the inference\-time Feynman–Kac proposal/reweighting representation on the same transition graph\.
### A\.4Broader future directions
The condensed main\-text discussion points to several adjacent directions: hybrid discrete–continuous controls for multimodal generation and reasoning\[[185](https://arxiv.org/html/2609.35947#bib.bib185),[122](https://arxiv.org/html/2609.35947#bib.bib122),[27](https://arxiv.org/html/2609.35947#bib.bib27),[93](https://arxiv.org/html/2609.35947#bib.bib93),[169](https://arxiv.org/html/2609.35947#bib.bib169),[167](https://arxiv.org/html/2609.35947#bib.bib167),[87](https://arxiv.org/html/2609.35947#bib.bib87)\]; combinations with alignment, preference optimization, and pretraining biases\[[92](https://arxiv.org/html/2609.35947#bib.bib92),[6](https://arxiv.org/html/2609.35947#bib.bib6),[139](https://arxiv.org/html/2609.35947#bib.bib139)\]; and distillation into continuous or categorical transport maps and parallel\-reasoning architectures\[[126](https://arxiv.org/html/2609.35947#bib.bib126),[124](https://arxiv.org/html/2609.35947#bib.bib124),[68](https://arxiv.org/html/2609.35947#bib.bib68),[102](https://arxiv.org/html/2609.35947#bib.bib102)\]\.
## Appendix BContinuous Proposal Control \(DriftLite\)
Inspired by twisted SMC methods in computational statistics\[[16](https://arxiv.org/html/2609.35947#bib.bib16),[161](https://arxiv.org/html/2609.35947#bib.bib161),[63](https://arxiv.org/html/2609.35947#bib.bib63),[80](https://arxiv.org/html/2609.35947#bib.bib80),[99](https://arxiv.org/html/2609.35947#bib.bib99)\], several recent works\[[183](https://arxiv.org/html/2609.35947#bib.bib183),[50](https://arxiv.org/html/2609.35947#bib.bib50),[4](https://arxiv.org/html/2609.35947#bib.bib4),[66](https://arxiv.org/html/2609.35947#bib.bib66),[109](https://arxiv.org/html/2609.35947#bib.bib109)\]explore proposal control for SMC\-based inference\-time scaling in diffusion models and LLMs\. Among them, DriftLite\[[119](https://arxiv.org/html/2609.35947#bib.bib119)\]provides a mathematically grounded construction for continuous diffusion models\. Since our discrete theory only needs the structural identity, we record a compact version here\.
##### Continuous setup\.
For continuous state space𝒳=ℝd\\mathcal\{X\}=\\mathbb\{R\}^\{d\}, a pretrained continuous diffusion model is characterized by a forward and backward diffusion process, written as the following pair of stochastic differential equations \(SDEs\) fors,t∈\[0,T\]s,t\\in\[0,T\]:
d𝐱s→=𝐯s→\(𝐱s→\)ds\+αs→d𝐰s,d𝐱t←=𝐯t←\(𝐱t←\)dt\+αt←d𝐰t,\\mathop\{\}\\\!\\mathrm\{d\}\\mathbf\{x\}\_\{s\}^\{\\shortrightarrow\}=\\mathbf\{v\}\_\{s\}^\{\\shortrightarrow\}\(\\mathbf\{x\}\_\{s\}^\{\\shortrightarrow\}\)\\,\\mathop\{\}\\\!\\mathrm\{d\}s\+\\alpha\_\{s\}^\{\\shortrightarrow\}\\,\\mathop\{\}\\\!\\mathrm\{d\}\\mathbf\{w\}\_\{s\},\\qquad\\mathop\{\}\\\!\\mathrm\{d\}\\mathbf\{x\}\_\{t\}^\{\\shortleftarrow\}=\\mathbf\{v\}\_\{t\}^\{\\shortleftarrow\}\(\\mathbf\{x\}\_\{t\}^\{\\shortleftarrow\}\)\\,\\mathop\{\}\\\!\\mathrm\{d\}t\+\\alpha\_\{t\}^\{\\shortleftarrow\}\\,\\mathop\{\}\\\!\\mathrm\{d\}\\mathbf\{w\}\_\{t\},\(B\.1\)where\(𝐰t\)t∈\[0,T\]\(\\mathbf\{w\}\_\{t\}\)\_\{t\\in\[0,T\]\}is the standard Brownian motion onℝd\\mathbb\{R\}^\{d\}and the backward drift is given by𝐯t←:=−𝐯T−t→\+αT−t→2\+αt←22∇logpt←\\mathbf\{v\}\_\{t\}^\{\\shortleftarrow\}:=\-\\mathbf\{v\}\_\{T\-t\}^\{\\shortrightarrow\}\+\\frac\{\\alpha\_\{T\-t\}^\{\\shortrightarrow\}\{\}^\{2\}\+\\alpha\_\{t\}^\{\\shortleftarrow\}\{\}^\{2\}\}\{2\}\\nabla\\log p\_\{t\}^\{\\shortleftarrow\}, with∇logpt←\\nabla\\log p\_\{t\}^\{\\shortleftarrow\}the score function \(typically approximated by a pretrained neural network\)\. Many inference\-time tasks can be expressed as sampling from the tilted path
qt\(𝐱\)∝\(pt←\(𝐱\)\)γert\(𝐱\),𝐱∈ℝd,t∈\[0,T\],q\_\{t\}\(\\mathbf\{x\}\)\\propto\(p\_\{t\}^\{\\shortleftarrow\}\(\\mathbf\{x\}\)\)^\{\\gamma\}e^\{r\_\{t\}\(\\mathbf\{x\}\)\},\\qquad\\mathbf\{x\}\\in\\mathbb\{R\}^\{d\},\\ t\\in\[0,T\],whereγ\>0\\gamma\>0is fixed andrt:ℝd→ℝr\_\{t\}:\\mathbb\{R\}^\{d\}\\to\\mathbb\{R\}is a time\-dependent reward, guidance, or steering potential\. A principled way to simulate this path is to represent it as a normalized Feynman–Kac flow and then run SMC\[[133](https://arxiv.org/html/2609.35947#bib.bib133),[25](https://arxiv.org/html/2609.35947#bib.bib25),[134](https://arxiv.org/html/2609.35947#bib.bib134)\]\. The full Feynman–Kac structure is captured by the proposition below; in particular, equation \([B\.2](https://arxiv.org/html/2609.35947#A2.E2)\) below corresponds to the compact normalized Feynman–Kac equation of Section[2\.1](https://arxiv.org/html/2609.35947#S2.SS1)\.
###### Proposition B\.1\(Continuous Feynman–Kac control identity\[[119](https://arxiv.org/html/2609.35947#bib.bib119), Proposition 2\.1\]\)\.
Fixγ\>0\\gamma\>0and a time\-dependent functionrt:ℝd→ℝr\_\{t\}:\\mathbb\{R\}^\{d\}\\to\\mathbb\{R\}\. Then the tilted path\(qt\)t∈\[0,T\]\(q\_\{t\}\)\_\{t\\in\[0,T\]\}admits a Feynman–Kac representation
∂tqt\(𝐱\)=\(ℒ~t∗qt\)\(𝐱\)\+g~t\(𝐱\)qt\(𝐱\),g~t=G~t−𝔼𝐳∼qt\[G~t\(𝐳\)\],\\partial\_\{t\}q\_\{t\}\(\\mathbf\{x\}\)=\(\\widetilde\{\\mathcal\{L\}\}\_\{t\}^\{\\ast\}q\_\{t\}\)\(\\mathbf\{x\}\)\+\\widetilde\{g\}\_\{t\}\(\\mathbf\{x\}\)\\,q\_\{t\}\(\\mathbf\{x\}\),\\qquad\\widetilde\{g\}\_\{t\}=\\widetilde\{G\}\_\{t\}\-\\mathbb\{E\}\_\{\\mathbf\{z\}\\sim q\_\{t\}\}\[\\widetilde\{G\}\_\{t\}\(\\mathbf\{z\}\)\],\(B\.2\)where
\(ℒ~t∗ρ\)\(𝐱\):=−∇⋅\(𝐯~t\(𝐱\)ρ\(𝐱\)\)\+αt←22Δρ\(𝐱\),\(\\widetilde\{\\mathcal\{L\}\}\_\{t\}^\{\\ast\}\\rho\)\(\\mathbf\{x\}\):=\-\\nabla\\cdot\\big\(\\widetilde\{\\mathbf\{v\}\}\_\{t\}\(\\mathbf\{x\}\)\\rho\(\\mathbf\{x\}\)\\big\)\+\\frac\{\\alpha\_\{t\}^\{\\shortleftarrow\}\{\}^\{2\}\}\{2\}\\Delta\\rho\(\\mathbf\{x\}\),\(B\.3\)for an appropriate guided drift𝐯~t\\widetilde\{\\mathbf\{v\}\}\_\{t\}, and whereG~t\\widetilde\{G\}\_\{t\}is the corresponding unnormalized potential \(see\[[119](https://arxiv.org/html/2609.35947#bib.bib119)\]for the explicit formula\)\. Moreover, for any control drift𝐮t:ℝd→ℝd\\mathbf\{u\}\_\{t\}:\\mathbb\{R\}^\{d\}\\to\\mathbb\{R\}^\{d\}, the same\(qt\)t∈\[0,T\]\(q\_\{t\}\)\_\{t\\in\[0,T\]\}also satisfies
∂tqt\(𝐱\)=−∇⋅\(\(𝐯~t\(𝐱\)\+𝐮t\(𝐱\)\)qt\(𝐱\)\)\+αt←22Δqt\(𝐱\)\+\(g~t\(𝐱\)\+divqt𝐮t\(𝐱\)\)qt\(𝐱\),\\partial\_\{t\}q\_\{t\}\(\\mathbf\{x\}\)=\-\\nabla\\cdot\\big\(\(\\widetilde\{\\mathbf\{v\}\}\_\{t\}\(\\mathbf\{x\}\)\+\\mathbf\{u\}\_\{t\}\(\\mathbf\{x\}\)\)q\_\{t\}\(\\mathbf\{x\}\)\\big\)\+\\frac\{\\alpha\_\{t\}^\{\\shortleftarrow\}\{\}^\{2\}\}\{2\}\\Delta q\_\{t\}\(\\mathbf\{x\}\)\+\\big\(\\widetilde\{g\}\_\{t\}\(\\mathbf\{x\}\)\+\\operatorname\{div\}\_\{q\_\{t\}\}\\mathbf\{u\}\_\{t\}\(\\mathbf\{x\}\)\\big\)q\_\{t\}\(\\mathbf\{x\}\),\(B\.4\)where
divqt𝐮t\(𝐱\):=1qt\(𝐱\)∇⋅\(qt\(𝐱\)𝐮t\(𝐱\)\)\.\\operatorname\{div\}\_\{q\_\{t\}\}\\mathbf\{u\}\_\{t\}\(\\mathbf\{x\}\):=\\frac\{1\}\{q\_\{t\}\(\\mathbf\{x\}\)\}\\,\\nabla\\cdot\\big\(q\_\{t\}\(\\mathbf\{x\}\)\\mathbf\{u\}\_\{t\}\(\\mathbf\{x\}\)\\big\)\.\(B\.5\)
The ideal zero\-variance continuous control solves the weighted Poisson equation\.
###### Proposition B\.2\(Optimal continuous control in DriftLite\[[119](https://arxiv.org/html/2609.35947#bib.bib119), Proposition 3\.2\]\)\.
Assume\(qt\)t∈\[0,T\]\(q\_\{t\}\)\_\{t\\in\[0,T\]\}satisfies \([B\.2](https://arxiv.org/html/2609.35947#A2.E2)\)\. Under suitable regularity assumptions, there exists a curl\-free optimal control𝐮t∗=∇𝒜t∗\\mathbf\{u\}\_\{t\}^\{\\ast\}=\\nabla\\mathcal\{A\}\_\{t\}^\{\\ast\}such that
g~t\+divqt𝐮t∗≡0,\\widetilde\{g\}\_\{t\}\+\\operatorname\{div\}\_\{q\_\{t\}\}\\mathbf\{u\}\_\{t\}^\{\\ast\}\\equiv 0,where𝒜t∗\\mathcal\{A\}\_\{t\}^\{\\ast\}solves the weighted Poisson equation
divqt∇𝒜t∗\(𝐱\)=1qt\(𝐱\)∇⋅\(qt\(𝐱\)∇𝒜t∗\(𝐱\)\)=−g~t\(𝐱\)\.\\operatorname\{div\}\_\{q\_\{t\}\}\\nabla\\mathcal\{A\}\_\{t\}^\{\\ast\}\(\\mathbf\{x\}\)=\\frac\{1\}\{q\_\{t\}\(\\mathbf\{x\}\)\}\\nabla\\cdot\\big\(q\_\{t\}\(\\mathbf\{x\}\)\\nabla\\mathcal\{A\}\_\{t\}^\{\\ast\}\(\\mathbf\{x\}\)\\big\)=\-\\widetilde\{g\}\_\{t\}\(\\mathbf\{x\}\)\.\(B\.6\)Equivalently,
𝐮t∗=argmin𝐮tVar𝐱∼qt\[g~t\(𝐱\)\+divqt𝐮t\(𝐱\)\],\\mathbf\{u\}\_\{t\}^\{\\ast\}=\\argmin\_\{\\mathbf\{u\}\_\{t\}\}\\operatorname\{Var\}\_\{\\mathbf\{x\}\\sim q\_\{t\}\}\\\!\\big\[\\widetilde\{g\}\_\{t\}\(\\mathbf\{x\}\)\+\\operatorname\{div\}\_\{q\_\{t\}\}\\mathbf\{u\}\_\{t\}\(\\mathbf\{x\}\)\\big\],or, in potential form,
𝒜t∗=argmin𝒜t∫ℝd\(12qt\(𝐱\)∥∇𝒜t\(𝐱\)∥22−qt\(𝐱\)g~t\(𝐱\)𝒜t\(𝐱\)\)d𝐱\.\\mathcal\{A\}\_\{t\}^\{\\ast\}=\\argmin\_\{\\mathcal\{A\}\_\{t\}\}\\;\\int\_\{\\mathbb\{R\}^\{d\}\}\\left\(\\frac\{1\}\{2\}q\_\{t\}\(\\mathbf\{x\}\)\\\|\\nabla\\mathcal\{A\}\_\{t\}\(\\mathbf\{x\}\)\\\|\_\{2\}^\{2\}\-q\_\{t\}\(\\mathbf\{x\}\)\\widetilde\{g\}\_\{t\}\(\\mathbf\{x\}\)\\mathcal\{A\}\_\{t\}\(\\mathbf\{x\}\)\\right\)\\mathop\{\}\\\!\\mathrm\{d\}\\mathbf\{x\}\.
A detailed proof can be found in\[[119](https://arxiv.org/html/2609.35947#bib.bib119)\]\. Since both𝐮t\\mathbf\{u\}\_\{t\}and𝒜t\\mathcal\{A\}\_\{t\}are high\-dimensional, directly solving \([B\.6](https://arxiv.org/html/2609.35947#A2.E6)\) is usually intractable\. DriftLite therefore restricts the search to finite\-dimensional subspaces such as
span\{∇rt,∇logpt←,𝐯t←\}andspan\{rt,logpt←,Ut←\},\\Span\\\{\\nabla r\_\{t\},\\nabla\\log p\_\{t\}^\{\\shortleftarrow\},\\mathbf\{v\}\_\{t\}^\{\\shortleftarrow\}\\\}\\qquad\\text\{and\}\\qquad\\Span\\\{r\_\{t\},\\log p\_\{t\}^\{\\shortleftarrow\},U\_\{t\}^\{\\shortleftarrow\}\\\},where∇Ut←=𝐯t←\\nabla U\_\{t\}^\{\\shortleftarrow\}=\\mathbf\{v\}\_\{t\}^\{\\shortleftarrow\}\.
## Appendix CNotation, Special Cases, and Proofs for Sections[2](https://arxiv.org/html/2609.35947#S2)–[3](https://arxiv.org/html/2609.35947#S3)
This appendix collects the notation and algebra supporting Sections[2](https://arxiv.org/html/2609.35947#S2)and[3](https://arxiv.org/html/2609.35947#S3)\. It first fixes the column\-vector convention and records the two common tilted\-path reductions, then proves the discrete Feynman–Kac identity and the FluxLite variance\-control propositions, and closes with further details onHEU\.
### C\.1Notation
The discrete state space is denoted by𝒳\\mathcal\{X\}\. For any functionf:𝒳→ℝf:\\mathcal\{X\}\\to\\mathbb\{R\}, we write‖f‖∞:=maxx∈𝒳\|f\(x\)\|\\\|f\\\|\_\{\\infty\}:=\\max\_\{x\\in\\mathcal\{X\}\}\|f\(x\)\|\. For any distributionπ\\pion𝒳\\mathcal\{X\}, expectations are denoted𝔼π\[⋅\]\\mathbb\{E\}\_\{\\pi\}\[\\cdot\]\. For any off\-diagonal rate family𝐐\\mathbf\{Q\}, we setQ\(x,x\)=0Q\(x,x\)=0by convention\. Whenever a full generator matrix is needed, we use the associated matrix𝐋\\mathbf\{L\}defined by
L\(y,x\)=Q\(y,x\)\(y≠x\),L\(x,x\)=−∑y≠xQ\(y,x\)\.L\(y,x\)=Q\(y,x\)\\ \\ \(y\\neq x\),\\qquad L\(x,x\)=\-\\sum\_\{y\\neq x\}Q\(y,x\)\.For a matrix𝐀∈ℝD×D\\mathbf\{A\}\\in\\mathbb\{R\}^\{D\\times D\},‖𝐀‖1→1:=max∑y∈𝒳x∈𝒳\|A\(y,x\)\|\\left\\\|\\mathbf\{A\}\\right\\\|\_\{1\\to 1\}:=\\max\_\{x\\in\\mathcal\{X\}\}\\sum\_\{y\\in\\mathcal\{X\}\}\|A\(y,x\)\|is the inducedℓ1\\ell\_\{1\}operator norm \(maximum absolute column sum\)\. All rate matrices and generators use this column convention throughout\.
### C\.2Special cases of the tilted Feynman–Kac representative
The following two special cases are used repeatedly in the paper; the boundedness and positivity assumptions needed for the theoretical statements are made explicit in Appendix[D](https://arxiv.org/html/2609.35947#A4)\.
- •Reward\-tilted path\(qt∝pt←ertq\_\{t\}\\propto p\_\{t\}^\{\\shortleftarrow\}e^\{r\_\{t\}\}\): for a possibly time\-dependent rewardrt:𝒳→ℝr\_\{t\}:\\mathcal\{X\}\\to\\mathbb\{R\}, Q~t\(x,y\)=Qt←\(x,y\)ert\(x\)−rt\(y\),G~t\(x\)=r˙t\(x\)\+∑y≠x\(Q~t\(y,x\)−Qt←\(y,x\)\)\.\\widetilde\{Q\}\_\{t\}\(x,y\)=Q\_\{t\}^\{\\shortleftarrow\}\(x,y\)\\,e^\{r\_\{t\}\(x\)\-r\_\{t\}\(y\)\},\\qquad\\widetilde\{G\}\_\{t\}\(x\)=\\dot\{r\}\_\{t\}\(x\)\+\\sum\_\{y\\neq x\}\\Big\(\\widetilde\{Q\}\_\{t\}\(y,x\)\-Q\_\{t\}^\{\\shortleftarrow\}\(y,x\)\\Big\)\.\(C\.1\)
- •Annealed path\(qt∝\(pt←\)γq\_\{t\}\\propto\(p\_\{t\}^\{\\shortleftarrow\}\)^\{\\gamma\}\): forγ\>0\\gamma\>0, Q~t\(x,y\)=γQt←\(x,y\)\(pt←\(y\)pt←\(x\)\)1−γ,G~t\(x\)=∑y≠x\(Q~t\(y,x\)−γQt←\(y,x\)\)\.\\widetilde\{Q\}\_\{t\}\(x,y\)=\\gamma\\,Q\_\{t\}^\{\\shortleftarrow\}\(x,y\)\\Big\(\\tfrac\{p\_\{t\}^\{\\shortleftarrow\}\(y\)\}\{p\_\{t\}^\{\\shortleftarrow\}\(x\)\}\\Big\)^\{1\-\\gamma\},\\qquad\\widetilde\{G\}\_\{t\}\(x\)=\\sum\_\{y\\neq x\}\\Big\(\\widetilde\{Q\}\_\{t\}\(y,x\)\-\\gamma Q\_\{t\}^\{\\shortleftarrow\}\(y,x\)\\Big\)\.\(C\.2\)
### C\.3Proof of Proposition[2\.1](https://arxiv.org/html/2609.35947#S2.Thmtheorem1)
For readability, writept:=pt←p\_\{t\}:=p\_\{t\}^\{\\shortleftarrow\}\. From
qt\(x\)=1Zt\(pt\(x\)\)γert\(x\),Zt=∑z∈𝒳\(pt\(z\)\)γert\(z\),q\_\{t\}\(x\)=\\frac\{1\}\{Z\_\{t\}\}\(p\_\{t\}\(x\)\)^\{\\gamma\}e^\{r\_\{t\}\(x\)\},\\qquad Z\_\{t\}=\\sum\_\{z\\in\\mathcal\{X\}\}\(p\_\{t\}\(z\)\)^\{\\gamma\}e^\{r\_\{t\}\(z\)\},we obtain
∂tqt\(x\)qt\(x\)=γ∂tpt\(x\)pt\(x\)\+r˙t\(x\)−Z˙tZt\.\\frac\{\\partial\_\{t\}q\_\{t\}\(x\)\}\{q\_\{t\}\(x\)\}=\\gamma\\frac\{\\partial\_\{t\}p\_\{t\}\(x\)\}\{p\_\{t\}\(x\)\}\+\\dot\{r\}\_\{t\}\(x\)\-\\frac\{\\dot\{Z\}\_\{t\}\}\{Z\_\{t\}\}\.\(C\.3\)Using the backward master equation \([2\.4](https://arxiv.org/html/2609.35947#S2.E4)\),
∂tpt\(x\)=∑y≠x\(Qt←\(x,y\)pt\(y\)−Qt←\(y,x\)pt\(x\)\),\\partial\_\{t\}p\_\{t\}\(x\)=\\sum\_\{y\\neq x\}\\Big\(Q\_\{t\}^\{\\shortleftarrow\}\(x,y\)p\_\{t\}\(y\)\-Q\_\{t\}^\{\\shortleftarrow\}\(y,x\)p\_\{t\}\(x\)\\Big\),hence
γ∂tpt\(x\)pt\(x\)=∑y≠xγQt←\(x,y\)pt\(y\)pt\(x\)−∑y≠xγQt←\(y,x\)\.\\gamma\\frac\{\\partial\_\{t\}p\_\{t\}\(x\)\}\{p\_\{t\}\(x\)\}=\\sum\_\{y\\neq x\}\\gamma Q\_\{t\}^\{\\shortleftarrow\}\(x,y\)\\frac\{p\_\{t\}\(y\)\}\{p\_\{t\}\(x\)\}\-\\sum\_\{y\\neq x\}\\gamma Q\_\{t\}^\{\\shortleftarrow\}\(y,x\)\.Multiplying byqt\(x\)q\_\{t\}\(x\)and using
γQt←\(x,y\)pt\(y\)pt\(x\)qt\(x\)=Q~t\(x,y\)qt\(y\),\\gamma Q\_\{t\}^\{\\shortleftarrow\}\(x,y\)\\frac\{p\_\{t\}\(y\)\}\{p\_\{t\}\(x\)\}q\_\{t\}\(x\)=\\widetilde\{Q\}\_\{t\}\(x,y\)q\_\{t\}\(y\),we deduce from \([C\.3](https://arxiv.org/html/2609.35947#A3.E3)\) that
∂tqt\(x\)=∑y≠xQ~t\(x,y\)qt\(y\)\+\(r˙t\(x\)−∑y≠xγQt←\(y,x\)−Z˙tZt\)qt\(x\)\.\\partial\_\{t\}q\_\{t\}\(x\)=\\sum\_\{y\\neq x\}\\widetilde\{Q\}\_\{t\}\(x,y\)q\_\{t\}\(y\)\+\\left\(\\dot\{r\}\_\{t\}\(x\)\-\\sum\_\{y\\neq x\}\\gamma Q\_\{t\}^\{\\shortleftarrow\}\(y,x\)\-\\frac\{\\dot\{Z\}\_\{t\}\}\{Z\_\{t\}\}\\right\)q\_\{t\}\(x\)\.\(C\.4\)
It remains to computeZ˙t/Zt\\dot\{Z\}\_\{t\}/Z\_\{t\}\. DifferentiatingZtZ\_\{t\}gives
Z˙tZt\\displaystyle\\frac\{\\dot\{Z\}\_\{t\}\}\{Z\_\{t\}\}=∑z∈𝒳qt\(z\)\(γ∂tpt\(z\)pt\(z\)\+r˙t\(z\)\)\\displaystyle=\\sum\_\{z\\in\\mathcal\{X\}\}q\_\{t\}\(z\)\\left\(\\gamma\\frac\{\\partial\_\{t\}p\_\{t\}\(z\)\}\{p\_\{t\}\(z\)\}\+\\dot\{r\}\_\{t\}\(z\)\\right\)=∑z∈𝒳qt\(z\)\[r˙t\(z\)\+∑y≠z\(Q~t\(y,z\)−γQt←\(y,z\)\)\]\\displaystyle=\\sum\_\{z\\in\\mathcal\{X\}\}q\_\{t\}\(z\)\\left\[\\dot\{r\}\_\{t\}\(z\)\+\\sum\_\{y\\neq z\}\\left\(\\widetilde\{Q\}\_\{t\}\(y,z\)\-\\gamma Q\_\{t\}^\{\\shortleftarrow\}\(y,z\)\\right\)\\right\]=𝔼qt\[G~t\],\\displaystyle=\\mathbb\{E\}\_\{q\_\{t\}\}\[\\widetilde\{G\}\_\{t\}\],where the second equality again uses \([2\.6](https://arxiv.org/html/2609.35947#S2.E6)\)\.
Substituting this identity into \([C\.4](https://arxiv.org/html/2609.35947#A3.E4)\) and adding/subtracting∑y≠xQ~t\(y,x\)qt\(x\)\\sum\_\{y\\neq x\}\\widetilde\{Q\}\_\{t\}\(y,x\)q\_\{t\}\(x\)yields
∂tqt\(x\)\\displaystyle\\partial\_\{t\}q\_\{t\}\(x\)=∑y≠x\(Q~t\(x,y\)qt\(y\)−Q~t\(y,x\)qt\(x\)\)\\displaystyle=\\sum\_\{y\\neq x\}\\Big\(\\widetilde\{Q\}\_\{t\}\(x,y\)q\_\{t\}\(y\)\-\\widetilde\{Q\}\_\{t\}\(y,x\)q\_\{t\}\(x\)\\Big\)\+\[r˙t\(x\)\+∑y≠x\(Q~t\(y,x\)−γQt←\(y,x\)\)−𝔼qt\[G~t\]\]qt\(x\)\\displaystyle\\quad\+\\left\[\\dot\{r\}\_\{t\}\(x\)\+\\sum\_\{y\\neq x\}\\Big\(\\widetilde\{Q\}\_\{t\}\(y,x\)\-\\gamma Q\_\{t\}^\{\\shortleftarrow\}\(y,x\)\\Big\)\-\\mathbb\{E\}\_\{q\_\{t\}\}\[\\widetilde\{G\}\_\{t\}\]\\right\]q\_\{t\}\(x\)=∑y≠x\(Q~t\(x,y\)qt\(y\)−Q~t\(y,x\)qt\(x\)\)\+\(G~t\(x\)−𝔼qt\[G~t\]\)qt\(x\),\\displaystyle=\\sum\_\{y\\neq x\}\\Big\(\\widetilde\{Q\}\_\{t\}\(x,y\)q\_\{t\}\(y\)\-\\widetilde\{Q\}\_\{t\}\(y,x\)q\_\{t\}\(x\)\\Big\)\+\\big\(\\widetilde\{G\}\_\{t\}\(x\)\-\\mathbb\{E\}\_\{q\_\{t\}\}\[\\widetilde\{G\}\_\{t\}\]\\big\)q\_\{t\}\(x\),which is exactly \([2\.5](https://arxiv.org/html/2609.35947#S2.E5)\)\.□\\square
### C\.4Proof of Proposition[3\.1](https://arxiv.org/html/2609.35947#S3.Thmtheorem1)
We start from \([3\.2](https://arxiv.org/html/2609.35947#S3.E2)\):
∂tqt\(x\)=∑y≠x\(Qt\(x,y\)qt\(y\)−Qt\(y,x\)qt\(x\)\)\+qt\(x\)gt\(x\)\.\\partial\_\{t\}q\_\{t\}\(x\)=\\sum\_\{y\\neq x\}\\Big\(Q\_\{t\}\(x,y\)q\_\{t\}\(y\)\-Q\_\{t\}\(y,x\)q\_\{t\}\(x\)\\Big\)\+q\_\{t\}\(x\)\\,g\_\{t\}\(x\)\.Let𝐑t\\mathbf\{R\}\_\{t\}be such thatQt′\(y,x\)=Qt\(y,x\)\+Rt\(y,x\)≥0Q\_\{t\}^\{\\prime\}\(y,x\)=Q\_\{t\}\(y,x\)\+R\_\{t\}\(y,x\)\\geq 0for allx≠yx\\neq y\. Then
∑y≠x\(Qt′\(x,y\)qt\(y\)−Qt′\(y,x\)qt\(x\)\)\\displaystyle\\sum\_\{y\\neq x\}\\Big\(Q\_\{t\}^\{\\prime\}\(x,y\)q\_\{t\}\(y\)\-Q\_\{t\}^\{\\prime\}\(y,x\)q\_\{t\}\(x\)\\Big\)=∑y≠x\(Qt\(x,y\)qt\(y\)−Qt\(y,x\)qt\(x\)\)\\displaystyle=\\sum\_\{y\\neq x\}\\Big\(Q\_\{t\}\(x,y\)q\_\{t\}\(y\)\-Q\_\{t\}\(y,x\)q\_\{t\}\(x\)\\Big\)\+∑y≠x\(Rt\(x,y\)qt\(y\)−Rt\(y,x\)qt\(x\)\)\.\\displaystyle\\quad\+\\sum\_\{y\\neq x\}\\Big\(R\_\{t\}\(x,y\)q\_\{t\}\(y\)\-R\_\{t\}\(y,x\)q\_\{t\}\(x\)\\Big\)\.By \([3\.1](https://arxiv.org/html/2609.35947#S3.E1)\),
qt\(x\)divqt𝐑t\(x\)=∑y≠x\(Rt\(y,x\)qt\(x\)−Rt\(x,y\)qt\(y\)\)\.q\_\{t\}\(x\)\\,\\operatorname\{div\}\_\{q\_\{t\}\}\\mathbf\{R\}\_\{t\}\(x\)=\\sum\_\{y\\neq x\}\\Big\(R\_\{t\}\(y,x\)q\_\{t\}\(x\)\-R\_\{t\}\(x,y\)q\_\{t\}\(y\)\\Big\)\.Therefore, if we set
gt′\(x\):=gt\(x\)\+divqt𝐑t\(x\),g\_\{t\}^\{\\prime\}\(x\):=g\_\{t\}\(x\)\+\\operatorname\{div\}\_\{q\_\{t\}\}\\mathbf\{R\}\_\{t\}\(x\),the added flux is canceled by the change in centered potential, which proves \([3\.3](https://arxiv.org/html/2609.35947#S3.E3)\)\.
Finally, centering is preserved:
∑xqt\(x\)divqt𝐑t\(x\)\\displaystyle\\sum\_\{x\}q\_\{t\}\(x\)\\,\\operatorname\{div\}\_\{q\_\{t\}\}\\mathbf\{R\}\_\{t\}\(x\)=∑x∑y≠x\(Rt\(y,x\)qt\(x\)−Rt\(x,y\)qt\(y\)\)\\displaystyle=\\sum\_\{x\}\\sum\_\{y\\neq x\}\\Big\(R\_\{t\}\(y,x\)q\_\{t\}\(x\)\-R\_\{t\}\(x,y\)q\_\{t\}\(y\)\\Big\)=∑x≠yRt\(y,x\)qt\(x\)−∑x≠yRt\(x,y\)qt\(y\)=0\.\\displaystyle=\\sum\_\{x\\neq y\}R\_\{t\}\(y,x\)q\_\{t\}\(x\)\-\\sum\_\{x\\neq y\}R\_\{t\}\(x,y\)q\_\{t\}\(y\)=0\.Hence𝔼qt\[gt′\]=𝔼qt\[gt\]=0\\mathbb\{E\}\_\{q\_\{t\}\}\[g\_\{t\}^\{\\prime\}\]=\\mathbb\{E\}\_\{q\_\{t\}\}\[g\_\{t\}\]=0\.□\\square
### C\.5Proof of Proposition[3\.2](https://arxiv.org/html/2609.35947#S3.Thmtheorem2)
Apply Proposition[3\.1](https://arxiv.org/html/2609.35947#S3.Thmtheorem1)with residual rate family𝐑t:=−𝐐t\\mathbf\{R\}\_\{t\}:=\-\\mathbf\{Q\}\_\{t\}, so thatQt′\(y,x\):=Qt\(y,x\)\+Rt\(y,x\)=0Q\_\{t\}^\{\\prime\}\(y,x\):=Q\_\{t\}\(y,x\)\+R\_\{t\}\(y,x\)=0for allx≠yx\\neq y\(admissibilityQt′\(y,x\)≥0Q\_\{t\}^\{\\prime\}\(y,x\)\\geq 0holds with equality\)\. By definition \([3\.1](https://arxiv.org/html/2609.35947#S3.E1)\),
divqt\(−𝐐t\)\(x\)=1qt\(x\)∑y≠x\(Qt\(x,y\)qt\(y\)−Qt\(y,x\)qt\(x\)\),\\operatorname\{div\}\_\{q\_\{t\}\}\(\-\\mathbf\{Q\}\_\{t\}\)\(x\)=\\frac\{1\}\{q\_\{t\}\(x\)\}\\sum\_\{y\\neq x\}\\Big\(Q\_\{t\}\(x,y\)q\_\{t\}\(y\)\-Q\_\{t\}\(y,x\)q\_\{t\}\(x\)\\Big\),sogt′\(x\)=gt\(x\)\+divqt\(−𝐐t\)\(x\)g\_\{t\}^\{\\prime\}\(x\)=g\_\{t\}\(x\)\+\\operatorname\{div\}\_\{q\_\{t\}\}\(\-\\mathbf\{Q\}\_\{t\}\)\(x\)coincides withgt0\(x\)g\_\{t\}^\{0\}\(x\)in \([3\.4](https://arxiv.org/html/2609.35947#S3.E4)\), and the dynamics \([3\.3](https://arxiv.org/html/2609.35947#S3.E3)\) collapses to∂tqt\(x\)=qt\(x\)gt0\(x\)\\partial\_\{t\}q\_\{t\}\(x\)=q\_\{t\}\(x\)\\,g\_\{t\}^\{0\}\(x\)\.
In particular,\(𝟎,gt0\)∈\[\(𝐐t,gt\)\]\(\\mathbf\{0\},g\_\{t\}^\{0\}\)\\in\[\(\\mathbf\{Q\}\_\{t\},g\_\{t\}\)\], so\[\(𝟎,gt0\)\]=\[\(𝐐t,gt\)\]\[\(\\mathbf\{0\},g\_\{t\}^\{0\}\)\]=\[\(\\mathbf\{Q\}\_\{t\},g\_\{t\}\)\]\. Finally, applying Proposition[3\.1](https://arxiv.org/html/2609.35947#S3.Thmtheorem1)to the representative\(𝟎,gt0\)\(\\mathbf\{0\},g\_\{t\}^\{0\}\)with residual𝐑t:=𝐐t′\\mathbf\{R\}\_\{t\}:=\\mathbf\{Q\}\_\{t\}^\{\\prime\}\(the admissibilityQt′\(y,x\)≥0Q\_\{t\}^\{\\prime\}\(y,x\)\\geq 0is exactly the assumed nonnegativity of𝐐t′\\mathbf\{Q\}\_\{t\}^\{\\prime\}\) gives
gt′\(x\)=gt0\(x\)\+divqt𝐐t′\(x\),g\_\{t\}^\{\\prime\}\(x\)=g\_\{t\}^\{0\}\(x\)\+\\operatorname\{div\}\_\{q\_\{t\}\}\\mathbf\{Q\}\_\{t\}^\{\\prime\}\(x\),which is \([3\.5](https://arxiv.org/html/2609.35947#S3.E5)\)\.□\\square
### C\.6Proof of Proposition[3\.3](https://arxiv.org/html/2609.35947#S3.Thmtheorem3)
Fixttand suppressttin the notation\. Work in the pure\-reweighting reference \([3\.4](https://arxiv.org/html/2609.35947#S3.E4)\) with centered potentialg0g^\{0\}\. Define the source vector
μ\(x\):=q\(x\)g0\(x\),\\mu\(x\):=q\(x\)\\,g^\{0\}\(x\),so that∑xμ\(x\)=𝔼q\[g0\]=0\\sum\_\{x\}\\mu\(x\)=\\mathbb\{E\}\_\{q\}\[g^\{0\}\]=0\.
Forx≠yx\\neq y, Proposition[3\.3](https://arxiv.org/html/2609.35947#S3.Thmtheorem3)defines
Q∗\(y,x\)=1D\[q\(y\)q\(x\)g0\(y\)−g0\(x\)\]\+\.Q^\{\*\}\(y,x\)=\\frac\{1\}\{D\}\\left\[\\frac\{q\(y\)\}\{q\(x\)\}g^\{0\}\(y\)\-g^\{0\}\(x\)\\right\]\_\{\+\}\.Multiplying byq\(x\)q\(x\)gives the flux fromxxtoyy:
Q∗\(y,x\)q\(x\)=1D\[μ\(y\)−μ\(x\)\]\+\.Q^\{\*\}\(y,x\)\\,q\(x\)=\\frac\{1\}\{D\}\\,\[\\mu\(y\)\-\\mu\(x\)\]\_\{\+\}\.\(C\.5\)Similarly,
Q∗\(x,y\)q\(y\)=1D\[μ\(x\)−μ\(y\)\]\+\.Q^\{\*\}\(x,y\)\\,q\(y\)=\\frac\{1\}\{D\}\\,\[\\mu\(x\)\-\\mu\(y\)\]\_\{\+\}\.
Now compute the divergence term in \([3\.5](https://arxiv.org/html/2609.35947#S3.E5)\):
q\(x\)\(divq𝐐∗\)\(x\)\\displaystyle q\(x\)\\,\(\\operatorname\{div\}\_\{q\}\\mathbf\{Q\}^\{\*\}\)\(x\)=∑y≠x\(Q∗\(y,x\)q\(x\)−Q∗\(x,y\)q\(y\)\)\\displaystyle=\\sum\_\{y\\neq x\}\\Big\(Q^\{\*\}\(y,x\)q\(x\)\-Q^\{\*\}\(x,y\)q\(y\)\\Big\)=1D∑y≠x\(\[μ\(y\)−μ\(x\)\]\+−\[μ\(x\)−μ\(y\)\]\+\)\.\\displaystyle=\\frac\{1\}\{D\}\\sum\_\{y\\neq x\}\\Big\(\[\\mu\(y\)\-\\mu\(x\)\]\_\{\+\}\-\[\\mu\(x\)\-\\mu\(y\)\]\_\{\+\}\\Big\)\.Using\[a\]\+−\[−a\]\+=a\[a\]\_\{\+\}\-\[\-a\]\_\{\+\}=a, we get
\[μ\(y\)−μ\(x\)\]\+−\[μ\(x\)−μ\(y\)\]\+=μ\(y\)−μ\(x\)\.\[\\mu\(y\)\-\\mu\(x\)\]\_\{\+\}\-\[\\mu\(x\)\-\\mu\(y\)\]\_\{\+\}=\\mu\(y\)\-\\mu\(x\)\.Therefore,
q\(x\)\(divq𝐐∗\)\(x\)\\displaystyle q\(x\)\\,\(\\operatorname\{div\}\_\{q\}\\mathbf\{Q\}^\{\*\}\)\(x\)=1D∑y≠x\(μ\(y\)−μ\(x\)\)\\displaystyle=\\frac\{1\}\{D\}\\sum\_\{y\\neq x\}\\big\(\\mu\(y\)\-\\mu\(x\)\\big\)=1D\(∑y≠xμ\(y\)−\(D−1\)μ\(x\)\)\\displaystyle=\\frac\{1\}\{D\}\\Big\(\\sum\_\{y\\neq x\}\\mu\(y\)\-\(D\-1\)\\mu\(x\)\\Big\)=1D\(∑yμ\(y\)⏟=0−Dμ\(x\)\)=−μ\(x\)=−q\(x\)g0\(x\)\.\\displaystyle=\\frac\{1\}\{D\}\\Big\(\\underbrace\{\\sum\_\{y\}\\mu\(y\)\}\_\{=0\}\-D\\mu\(x\)\\Big\)=\-\\mu\(x\)=\-q\(x\)\\,g^\{0\}\(x\)\.Dividing byq\(x\)\>0q\(x\)\>0yields\(divq𝐐∗\)\(x\)=−g0\(x\)\(\\operatorname\{div\}\_\{q\}\\mathbf\{Q\}^\{\*\}\)\(x\)=\-g^\{0\}\(x\)for allxx\. Hence
g∗\(x\)=g0\(x\)\+\(divq𝐐∗\)\(x\)=0,g^\{\*\}\(x\)=g^\{0\}\(x\)\+\(\\operatorname\{div\}\_\{q\}\\mathbf\{Q\}^\{\*\}\)\(x\)=0,which proves the claim\.□\\square
### C\.7Proof of Proposition[3\.4](https://arxiv.org/html/2609.35947#S3.Thmtheorem4)
We prove that minimizing the variance objective \([3\.6](https://arxiv.org/html/2609.35947#S3.E6)\) over sparse off\-diagonal rate families is equivalent to the constrained weighted least\-squares problem \([3\.8](https://arxiv.org/html/2609.35947#S3.E8)\)\.
##### Residual flux form\.
Work in the pure\-reweighting reference\(𝟎,gt0\)\(\\mathbf\{0\},g\_\{t\}^\{0\}\)\. For any candidate off\-diagonal rate family𝐐t′\\mathbf\{Q\}\_\{t\}^\{\\prime\}, Proposition[3\.1](https://arxiv.org/html/2609.35947#S3.Thmtheorem1)yields
gt′\(x\)=gt0\(x\)\+divqt𝐐t′\(x\)\.g\_\{t\}^\{\\prime\}\(x\)=g\_\{t\}^\{0\}\(x\)\+\\operatorname\{div\}\_\{q\_\{t\}\}\\mathbf\{Q\}\_\{t\}^\{\\prime\}\(x\)\.Multiplying byqt\(x\)q\_\{t\}\(x\)and using \([3\.1](https://arxiv.org/html/2609.35947#S3.E1)\),
qt\(x\)gt′\(x\)=μt\(x\)\+∑y≠x\(Qt′\(y,x\)qt\(x\)−Qt′\(x,y\)qt\(y\)\),q\_\{t\}\(x\)\\,g\_\{t\}^\{\\prime\}\(x\)=\\mu\_\{t\}\(x\)\+\\sum\_\{y\\neq x\}\\Big\(Q\_\{t\}^\{\\prime\}\(y,x\)q\_\{t\}\(x\)\-Q\_\{t\}^\{\\prime\}\(x,y\)q\_\{t\}\(y\)\\Big\),\(C\.6\)whereμt\(x\):=qt\(x\)gt0\(x\)\\mu\_\{t\}\(x\):=q\_\{t\}\(x\)g\_\{t\}^\{0\}\(x\)\.
Under the sparsity constraint,Qt′\(y,x\)=0Q\_\{t\}^\{\\prime\}\(y,x\)=0whenever\(y,x\)∉𝒮sp\(y,x\)\\notin\\mathcal\{S\}^\{\\mathrm\{sp\}\}, so this becomes
qt\(x\)gt′\(x\)=μt\(x\)\+∑y:\(y,x\)∈𝒮spQt′\(y,x\)qt\(x\)−∑y:\(x,y\)∈𝒮spQt′\(x,y\)qt\(y\)\.q\_\{t\}\(x\)\\,g\_\{t\}^\{\\prime\}\(x\)=\\mu\_\{t\}\(x\)\+\\sum\_\{y:\\,\(y,x\)\\in\\mathcal\{S\}^\{\\mathrm\{sp\}\}\}Q\_\{t\}^\{\\prime\}\(y,x\)q\_\{t\}\(x\)\-\\sum\_\{y:\\,\(x,y\)\\in\\mathcal\{S\}^\{\\mathrm\{sp\}\}\}Q\_\{t\}^\{\\prime\}\(x,y\)q\_\{t\}\(y\)\.
##### Weighted least squares\.
Becausegt′g\_\{t\}^\{\\prime\}is centered underqtq\_\{t\}by Proposition[3\.1](https://arxiv.org/html/2609.35947#S3.Thmtheorem1),
Varqt\(gt′\)=∑x∈𝒳qt\(x\)gt′\(x\)2\.\\operatorname\{Var\}\_\{q\_\{t\}\}\(g\_\{t\}^\{\\prime\}\)=\\sum\_\{x\\in\\mathcal\{X\}\}q\_\{t\}\(x\)\\,g\_\{t\}^\{\\prime\}\(x\)^\{2\}\.Using the previous identity,
∑x∈𝒳qt\(x\)gt′\(x\)2=∑x∈𝒳1qt\(x\)\[μt\(x\)\+∑y:\(y,x\)∈𝒮spQt′\(y,x\)qt\(x\)−∑y:\(x,y\)∈𝒮spQt′\(x,y\)qt\(y\)\]2\.\\sum\_\{x\\in\\mathcal\{X\}\}q\_\{t\}\(x\)\\,g\_\{t\}^\{\\prime\}\(x\)^\{2\}=\\sum\_\{x\\in\\mathcal\{X\}\}\\frac\{1\}\{q\_\{t\}\(x\)\}\\left\[\\mu\_\{t\}\(x\)\+\\sum\_\{y:\\,\(y,x\)\\in\\mathcal\{S\}^\{\\mathrm\{sp\}\}\}Q\_\{t\}^\{\\prime\}\(y,x\)q\_\{t\}\(x\)\-\\sum\_\{y:\\,\(x,y\)\\in\\mathcal\{S\}^\{\\mathrm\{sp\}\}\}Q\_\{t\}^\{\\prime\}\(x,y\)q\_\{t\}\(y\)\\right\]^\{2\}\.Hence minimizing \([3\.6](https://arxiv.org/html/2609.35947#S3.E6)\) over all sparse off\-diagonal rate families𝐐t′\\mathbf\{Q\}\_\{t\}^\{\\prime\}with
Qt′\(y,x\)=0for\(y,x\)∉𝒮sp,Qt′\(y,x\)≥0for\(y,x\)∈𝒮sp,Q\_\{t\}^\{\\prime\}\(y,x\)=0\\ \\text\{for \}\(y,x\)\\notin\\mathcal\{S\}^\{\\mathrm\{sp\}\},\\qquad Q\_\{t\}^\{\\prime\}\(y,x\)\\geq 0\\ \\text\{for \}\(y,x\)\\in\\mathcal\{S\}^\{\\mathrm\{sp\}\},is exactly the constrained weighted least\-squares problem \([3\.8](https://arxiv.org/html/2609.35947#S3.E8)\)\.□\\square
### C\.8Proof of Proposition[3\.6](https://arxiv.org/html/2609.35947#S3.Thmtheorem6)
Suppressttfor readability and write expectations underqq\. By definition,
gvcg\(x,𝜽\)=g0\(x\)\+∑j=1Jθj\(divq𝐐\(j\)\)\(x\)\.g^\{\\mathrm\{vcg\}\}\(x;\\boldsymbol\{\\theta\}\)=g^\{0\}\(x\)\+\\sum\_\{j=1\}^\{J\}\\theta\_\{j\}\(\\operatorname\{div\}\_\{q\}\\mathbf\{Q\}^\{\(j\)\}\)\(x\)\.Let
d¯\(j\):=𝔼q\[\(divq𝐐\(j\)\)\(x\)\],D~\(j\)\(x\):=\(divq𝐐\(j\)\)\(x\)−d¯\(j\)\.\\bar\{d\}^\{\(j\)\}:=\\mathbb\{E\}\_\{q\}\[\(\\operatorname\{div\}\_\{q\}\\mathbf\{Q\}^\{\(j\)\}\)\(x\)\],\\qquad\\widetilde\{D\}^\{\(j\)\}\(x\):=\(\\operatorname\{div\}\_\{q\}\\mathbf\{Q\}^\{\(j\)\}\)\(x\)\-\\bar\{d\}^\{\(j\)\}\.Since𝔼q\[g0\]=0\\mathbb\{E\}\_\{q\}\[g^\{0\}\]=0, the centered version ofgvcg\(⋅,𝜽\)g^\{\\mathrm\{vcg\}\}\(\\cdot;\\boldsymbol\{\\theta\}\)is
g0\(x\)\+∑j=1JθjD~\(j\)\(x\)\.g^\{0\}\(x\)\+\\sum\_\{j=1\}^\{J\}\\theta\_\{j\}\\widetilde\{D\}^\{\(j\)\}\(x\)\.Therefore
Varq\(gvcg\(⋅,𝜽\)\)\\displaystyle\\operatorname\{Var\}\_\{q\}\(g^\{\\mathrm\{vcg\}\}\(\\cdot;\\boldsymbol\{\\theta\}\)\)=𝔼q\[\(g0\(x\)\)2\]\+2∑i=1Jθi𝔼q\[g0\(x\)D~\(i\)\(x\)\]\\displaystyle=\\mathbb\{E\}\_\{q\}\[\(g^\{0\}\(x\)\)^\{2\}\]\+2\\sum\_\{i=1\}^\{J\}\\theta\_\{i\}\\,\\mathbb\{E\}\_\{q\}\\\!\\left\[g^\{0\}\(x\)\\widetilde\{D\}^\{\(i\)\}\(x\)\\right\]\+∑i,j=1Jθiθj𝔼q\[D~\(i\)\(x\)D~\(j\)\(x\)\]\.\\displaystyle\\quad\+\\sum\_\{i,j=1\}^\{J\}\\theta\_\{i\}\\theta\_\{j\}\\,\\mathbb\{E\}\_\{q\}\\\!\\left\[\\widetilde\{D\}^\{\(i\)\}\(x\)\\widetilde\{D\}^\{\(j\)\}\(x\)\\right\]\.This is a convex quadratic function of𝜽\\boldsymbol\{\\theta\}\. Its gradient is2\(𝐀𝜽\+𝐜\)2\(\\mathbf\{A\}\\boldsymbol\{\\theta\}\+\\mathbf\{c\}\), so the first\-order optimality condition for the unconstrained problem is
∑j=1JAijθj=−ci,\\sum\_\{j=1\}^\{J\}A\_\{ij\}\\theta\_\{j\}=\-c\_\{i\},withAijA\_\{ij\}andcic\_\{i\}as in \([3\.12](https://arxiv.org/html/2609.35947#S3.E12)\)\. Because the objective is convex, these normal equations characterize the full set of unconstrained minimizers; if𝐀\\mathbf\{A\}is nonsingular, the minimizer is unique\. The nonnegative version is the same quadratic objective restricted toℝ\+J\\mathbb\{R\}\_\{\+\}^\{J\}\.□\\square
### C\.9HEUdetails
This appendix expands on Proposition[3\.5](https://arxiv.org/html/2609.35947#S3.Thmtheorem5)\(the local\-average reallocation heuristic\)\. The implementable heuristic uses the outgoing one\-hop neighborhoodNt\+\(x\)N\_\{t\}^\{\+\}\(x\)\. More generally, one could replaceNt\+\(x\)N\_\{t\}^\{\+\}\(x\)by a larger local neighborhoodBt\(x\)⊆𝒳B\_\{t\}\(x\)\\subseteq\\mathcal\{X\}, e\.g\. by including incoming or multi\-hop neighbors; this would make the rule closer to the dense zero\-variance construction of Proposition[3\.3](https://arxiv.org/html/2609.35947#S3.Thmtheorem3), at the price of higher computational cost\.
##### Local\-averaging viewpoint\.
When the implemented neighborhood is reciprocal,z∈Bt\(x\)z\\in B\_\{t\}\(x\)iffx∈Bt\(z\)x\\in B\_\{t\}\(z\), and the same pairwise normalization is used in both directions, the positive\-part identity\[a−b\]\+−\[b−a\]\+=a−b\[a\-b\]\_\{\+\}\-\[b\-a\]\_\{\+\}=a\-bgives the local\-averaging heuristic
qt\(x\)divqt𝐐theu\(x\)≈αt\|Bt\(x\)\|∑z∈Bt\(x\)\(μt\(z\)−μt\(x\)\),q\_\{t\}\(x\)\\,\\operatorname\{div\}\_\{q\_\{t\}\}\\mathbf\{Q\}\_\{t\}^\{\{\\mathrm\{heu\}\}\}\(x\)\\approx\\frac\{\\alpha\_\{t\}\}\{\|B\_\{t\}\(x\)\|\}\\sum\_\{z\\in B\_\{t\}\(x\)\}\(\\mu\_\{t\}\(z\)\-\\mu\_\{t\}\(x\)\),so that
qt\(x\)gtheu\(x\)=μt\(x\)\+qt\(x\)divqt𝐐theu\(x\)≈\(1−αt\)μt\(x\)\+αt\|Bt\(x\)\|∑z∈Bt\(x\)μt\(z\)\.q\_\{t\}\(x\)\\,g\_\{t\}^\{\{\\mathrm\{heu\}\}\}\(x\)=\\mu\_\{t\}\(x\)\+q\_\{t\}\(x\)\\,\\operatorname\{div\}\_\{q\_\{t\}\}\\mathbf\{Q\}\_\{t\}^\{\{\\mathrm\{heu\}\}\}\(x\)\\approx\(1\-\\alpha\_\{t\}\)\\mu\_\{t\}\(x\)\+\\frac\{\\alpha\_\{t\}\}\{\|B\_\{t\}\(x\)\|\}\\sum\_\{z\\in B\_\{t\}\(x\)\}\\mu\_\{t\}\(z\)\.The heuristic therefore damps the part ofμt\(x\)\\mu\_\{t\}\(x\)that deviates from a local neighborhood average, while leaving a residual local mean\. Larger neighborhoods can bring the heuristic closer to the dense zero\-residual construction\.
##### Zero extra score\-network evaluations\.
A practical remark on the cost ofHEU: the source termμt\(y\)\\mu\_\{t\}\(y\)at a destination stateyyadmits the analytical expansion
μt\(y\)=qt\(y\)gt\(y\)\+∑z≠y\(Qt\(y,z\)qt\(z\)−Qt\(z,y\)qt\(y\)\),\\mu\_\{t\}\(y\)=q\_\{t\}\(y\)g\_\{t\}\(y\)\+\\sum\_\{z\\neq y\}\\big\(Q\_\{t\}\(y,z\)q\_\{t\}\(z\)\-Q\_\{t\}\(z,y\)q\_\{t\}\(y\)\\big\),i\.e\.μt\(y\)=qt\(y\)gt0\(y\)\\mu\_\{t\}\(y\)=q\_\{t\}\(y\)\\,g\_\{t\}^\{0\}\(y\)withgt0g\_\{t\}^\{0\}defined in \([3\.4](https://arxiv.org/html/2609.35947#S3.E4)\)\. This form is well\-defined on everyyy, including unvisited destinations of an empirical particle cloud, and uses only edge rates and density values that the proposal step already evaluates; in particular, no additional score\-network forward pass is required to evaluateμt\\mu\_\{t\}on the one\-hop neighborhood of the current particles\.
##### Choosingkt\(x\)k\_\{t\}\(x\)in mask diffusion\.
The normalization factorkt\(x\)k\_\{t\}\(x\)sets the strength of local reallocation\. In mask diffusion on𝒳=\[V\+1\]L\\mathcal\{X\}=\[V\+1\]^\{L\}, ifm\(x\)m\(x\)denotes the number of masked positions, the outgoing neighborhood has size\|Nt\+\(x\)\|=Vm\(x\)\|N\_\{t\}^\{\+\}\(x\)\|=V\\,m\(x\)and the incoming predecessor set has size\|Nt−\(x\)\|=L−m\(x\)\|N\_\{t\}^\{\-\}\(x\)\|=L\-m\(x\)\. We consider two natural choices: an*aggressive*outgoing\-star normalization
kt\(x\):=1\+\|Nt\+\(x\)\|=1\+Vm\(x\),k\_\{t\}\(x\):=1\+\|N\_\{t\}^\{\+\}\(x\)\|=1\+V\\,m\(x\),\(C\.7\)and a more*conservative*one\-hop normalization
kt\(x\):=1\+\|Nt\+\(x\)\|\+\|Nt−\(x\)\|=1\+Vm\(x\)\+L−m\(x\),k\_\{t\}\(x\):=1\+\|N\_\{t\}^\{\+\}\(x\)\|\+\|N\_\{t\}^\{\-\}\(x\)\|=1\+V\\,m\(x\)\+L\-m\(x\),\(C\.8\)which better reflects the local\-averaging picture by accounting for both outgoing and incoming one\-step neighbors\. The damping parameterαt<1\\alpha\_\{t\}<1can be used in either case to further stabilize the reallocation whengt0g\_\{t\}^\{0\}is noisy\.
## Appendix DAssumptions and Proofs for Section[4](https://arxiv.org/html/2609.35947#S4)
This appendix states the exact hypotheses used in Section[4](https://arxiv.org/html/2609.35947#S4)and proves the two theory results in the notation of the main text\. The first result is a population score\-ratio stability statement for the guided Feynman–Kac flow when the forward noising rates are known and the reverse rates are represented through the learned local ratio\. The second result is a particle\-only statement for a fixed deterministic grid Feynman–Kac recursion\. We state the assumptions, prove the two theorems, give a supporting weight\-variance lemma, and conclude with remarks on the role and scope of the results\.
### D\.1Assumptions
###### Assumption D\.1\(Stability of guided unnormalized operators\)\.
Let\(qt\)t∈\[0,T\]\(q\_\{t\}\)\_\{t\\in\[0,T\]\}and\(qts^\)t∈\[0,T\]\(q\_\{t\}^\{\\widehat\{s\}\}\)\_\{t\\in\[0,T\]\}be the normalized Feynman–Kac flows generated by the exact and implemented guided pairs defined in Assumption[D\.2](https://arxiv.org/html/2609.35947#A4.Thmassumption2)\. Equivalently, let their unnormalized versions solve
∂tρt=\(𝐋~t\+diag\(G~t\)\)ρt,∂tρ^t=\(𝐋~^t\+diag\(G~^t\)\)ρ^t,ρ0=ρ^0=q0,\\partial\_\{t\}\\rho\_\{t\}=\\big\(\\widetilde\{\\mathbf\{L\}\}\_\{t\}\+\\diag\(\\widetilde\{G\}\_\{t\}\)\\big\)\\rho\_\{t\},\\qquad\\partial\_\{t\}\\widehat\{\\rho\}\_\{t\}=\\big\(\\widehat\{\\widetilde\{\\mathbf\{L\}\}\}\_\{t\}\+\\diag\(\\widehat\{\\widetilde\{G\}\}\_\{t\}\)\\big\)\\widehat\{\\rho\}\_\{t\},\\qquad\\rho\_\{0\}=\\widehat\{\\rho\}\_\{0\}=q\_\{0\},withqt=ρt/\(𝟏⊤ρt\)q\_\{t\}=\\rho\_\{t\}/\(\\mathbf\{1\}^\{\\top\}\\rho\_\{t\}\)andqts^=ρ^t/\(𝟏⊤ρ^t\)q\_\{t\}^\{\\widehat\{s\}\}=\\widehat\{\\rho\}\_\{t\}/\(\\mathbf\{1\}^\{\\top\}\\widehat\{\\rho\}\_\{t\}\)\. Define
λ~t\(x\):=∑y≠xQ~t\(y,x\),λt←\(x\):=∑y≠xQt←\(y,x\),λ¯op\(t\):=qt\(λ~t\+γλt←\)\.\\widetilde\{\\lambda\}\_\{t\}\(x\):=\\sum\_\{y\\neq x\}\\widetilde\{Q\}\_\{t\}\(y,x\),\\qquad\\lambda\_\{t\}^\{\\shortleftarrow\}\(x\):=\\sum\_\{y\\neq x\}Q\_\{t\}^\{\\shortleftarrow\}\(y,x\),\\qquad\\bar\{\\lambda\}\_\{\\rm op\}\(t\):=q\_\{t\}\\\!\\left\(\\widetilde\{\\lambda\}\_\{t\}\+\\gamma\\lambda\_\{t\}^\{\\shortleftarrow\}\\right\)\.Assumeλ¯op\\bar\{\\lambda\}\_\{\\rm op\}is integrable on\[0,T\]\[0,T\]and that there exist constantsΛop,B<∞\\Lambda\_\{\\rm op\},B<\\inftysuch that, for allt∈\[0,T\]t\\in\[0,T\],
max∑y≠xx∈𝒳Q~^t\(y,x\)≤Λop,‖G~t‖∞≤B,‖G~^t‖∞≤B\.\\max\_\{x\\in\\mathcal\{X\}\}\\sum\_\{y\\neq x\}\\widehat\{\\widetilde\{Q\}\}\_\{t\}\(y,x\)\\leq\\Lambda\_\{\\rm op\},\\qquad\\left\\\|\\widetilde\{G\}\_\{t\}\\right\\\|\_\{\\infty\}\\leq B,\\qquad\\left\\\|\\widehat\{\\widetilde\{G\}\}\_\{t\}\\right\\\|\_\{\\infty\}\\leq B\.\(D\.1\)
###### Assumption D\.2\(Known\-forward local\-ratio perturbation\)\.
Assume the forward noising ratesQT−t→\(x,y\)Q\_\{T\-t\}^\{\\shortrightarrow\}\(x,y\)are known and that the exact reverse rates are represented by the local ratio
st\(x,y\):=pt←\(y\)pt←\(x\),Qt←\(y,x\)=QT−t→\(x,y\)st\(x,y\)\.s\_\{t\}\(x,y\):=\\frac\{p\_\{t\}^\{\\shortleftarrow\}\(y\)\}\{p\_\{t\}^\{\\shortleftarrow\}\(x\)\},\\qquad Q\_\{t\}^\{\\shortleftarrow\}\(y,x\)=Q\_\{T\-t\}^\{\\shortrightarrow\}\(x,y\)s\_\{t\}\(x,y\)\.The implemented reverse rates uses^t\\widehat\{s\}\_\{t\}:
Q^t←\(y,x\)=QT−t→\(x,y\)s^t\(x,y\)\.\\widehat\{Q\}\_\{t\}^\{\\shortleftarrow\}\(y,x\)=Q\_\{T\-t\}^\{\\shortrightarrow\}\(x,y\)\\widehat\{s\}\_\{t\}\(x,y\)\.Forx≠yx\\neq y, define the exact and implemented guided rates by
Q~t\(y,x\)\\displaystyle\\widetilde\{Q\}\_\{t\}\(y,x\)=γQT−t→\(x,y\)st\(x,y\)γert\(y\)−rt\(x\),\\displaystyle=\\gamma Q\_\{T\-t\}^\{\\shortrightarrow\}\(x,y\)s\_\{t\}\(x,y\)^\{\\gamma\}e^\{r\_\{t\}\(y\)\-r\_\{t\}\(x\)\},Q~^t\(y,x\)\\displaystyle\\widehat\{\\widetilde\{Q\}\}\_\{t\}\(y,x\)=γQT−t→\(x,y\)s^t\(x,y\)γert\(y\)−rt\(x\)\.\\displaystyle=\\gamma Q\_\{T\-t\}^\{\\shortrightarrow\}\(x,y\)\\widehat\{s\}\_\{t\}\(x,y\)^\{\\gamma\}e^\{r\_\{t\}\(y\)\-r\_\{t\}\(x\)\}\.\(D\.2\)Their unnormalized Feynman–Kac potentials are
G~t\(x\)\\displaystyle\\widetilde\{G\}\_\{t\}\(x\)=r˙t\(x\)\+∑y≠x\(Q~t\(y,x\)−γQt←\(y,x\)\),\\displaystyle=\\dot\{r\}\_\{t\}\(x\)\+\\sum\_\{y\\neq x\}\\Big\(\\widetilde\{Q\}\_\{t\}\(y,x\)\-\\gamma Q\_\{t\}^\{\\shortleftarrow\}\(y,x\)\\Big\),G~^t\(x\)\\displaystyle\\widehat\{\\widetilde\{G\}\}\_\{t\}\(x\)=r˙t\(x\)\+∑y≠x\(Q~^t\(y,x\)−γQ^t←\(y,x\)\)\.\\displaystyle=\\dot\{r\}\_\{t\}\(x\)\+\\sum\_\{y\\neq x\}\\Big\(\\widehat\{\\widetilde\{Q\}\}\_\{t\}\(y,x\)\-\\gamma\\widehat\{Q\}\_\{t\}^\{\\shortleftarrow\}\(y,x\)\\Big\)\.\(D\.3\)Assume there exist constants0<κ¯≤κ¯<∞0<\\underline\{\\kappa\}\\leq\\overline\{\\kappa\}<\\inftysuch that
κ¯≤s^t\(x,y\)st\(x,y\)≤κ¯wheneverQ~t\(y,x\)\+Qt←\(y,x\)\>0\.\\underline\{\\kappa\}\\leq\\frac\{\\widehat\{s\}\_\{t\}\(x,y\)\}\{s\_\{t\}\(x,y\)\}\\leq\\overline\{\\kappa\}\\qquad\\text\{whenever \}\\widetilde\{Q\}\_\{t\}\(y,x\)\+Q\_\{t\}^\{\\shortleftarrow\}\(y,x\)\>0\.\(D\.4\)
###### Assumption D\.3\(Score\-entropy coverage of the tilted path\)\.
Let
νtSE\(x,y\):=pt←\(x\)Qt←\(y,x\),x≠y,\\nu\_\{t\}^\{\\rm SE\}\(x,y\):=p\_\{t\}^\{\\shortleftarrow\}\(x\)Q\_\{t\}^\{\\shortleftarrow\}\(y,x\),\\qquad x\\neq y,\(D\.5\)be the source\-weighted score\-entropy edge measure associated with the reverse diffusion rates\. Define the time\-pointwise score\-entropy training loss and its time\-integrated scalar,
ℒDD\(t\)\\displaystyle\\mathcal\{L\}\_\{\\rm DD\}\(t\):=∑x∈𝒳∑y≠xνtSE\(x,y\)ℓ\(s^t\(x,y\)st\(x,y\)\),\\displaystyle:=\\sum\_\{x\\in\\mathcal\{X\}\}\\sum\_\{y\\neq x\}\\nu\_\{t\}^\{\\rm SE\}\(x,y\)\\,\\ell\\\!\\left\(\\frac\{\\widehat\{s\}\_\{t\}\(x,y\)\}\{s\_\{t\}\(x,y\)\}\\right\),𝔏DD\\displaystyle\\mathfrak\{L\}\_\{\\rm DD\}:=∫0TℒDD\(t\)𝑑t,ℓ\(u\):=−logu−1\+u,\\displaystyle:=\\int\_\{0\}^\{T\}\\mathcal\{L\}\_\{\\rm DD\}\(t\)\\,\\mathop\{\}\\\!\\mathrm\{d\}t,\\qquad\\ell\(u\):=\-\\log u\-1\+u,\(D\.6\)and the pointwise tilt\-to\-base ratio
ht\(x\):=qt\(x\)pt←\(x\)\.h\_\{t\}\(x\):=\\frac\{q\_\{t\}\(x\)\}\{p\_\{t\}^\{\\shortleftarrow\}\(x\)\}\.\(D\.7\)Define further the time\-pointwise edgewise coverage factor and itsL2L^\{2\}\-in\-time integrated scalar,
CDD\(t\)\\displaystyle C\_\{\\rm DD\}\(t\):=\(∑x∈𝒳∑y≠xνtSE\(x,y\)\(Cγ,κ¯,κ¯ht\(y\)\+C1,κ¯,κ¯ht\(x\)\)2\)1/2,\\displaystyle:=\\left\(\\sum\_\{x\\in\\mathcal\{X\}\}\\sum\_\{y\\neq x\}\\nu\_\{t\}^\{\\rm SE\}\(x,y\)\\big\(C\_\{\\gamma,\\underline\{\\kappa\},\\overline\{\\kappa\}\}h\_\{t\}\(y\)\+C\_\{1,\\underline\{\\kappa\},\\overline\{\\kappa\}\}h\_\{t\}\(x\)\\big\)^\{2\}\\right\)^\{1/2\},\(D\.8\)ℭDD\\displaystyle\\mathfrak\{C\}\_\{\\rm DD\}:=\(∫0TCDD\(t\)2𝑑t\)1/2\.\\displaystyle:=\\left\(\\int\_\{0\}^\{T\}C\_\{\\rm DD\}\(t\)^\{2\}\\,\\mathop\{\}\\\!\\mathrm\{d\}t\\right\)^\{1/2\}\.\(D\.9\)AssumeCDD\(t\)C\_\{\\rm DD\}\(t\)is finite for a\.e\.ttandℭDD<∞\\mathfrak\{C\}\_\{\\rm DD\}<\\infty\.
###### Assumption D\.4\(Fixed grid\-based controlled Feynman–Kac model\)\.
Fix a time grid0=t0<⋯<tM=T0=t\_\{0\}<\\cdots<t\_\{M\}=T\. For eachk=0,…,M−1k=0,\\dots,M\-1, let𝐐tkeff\\mathbf\{Q\}\_\{t\_\{k\}\}^\{\\mathrm\{eff\}\}be a deterministic off\-diagonal rate family and letgtkeff:𝒳→ℝg\_\{t\_\{k\}\}^\{\\mathrm\{eff\}\}:\\mathcal\{X\}\\to\\mathbb\{R\}be a deterministic potential\. Let𝐋tkeff\\mathbf\{L\}\_\{t\_\{k\}\}^\{\\mathrm\{eff\}\}be the associated full generator,
Ltkeff\(y,x\)=Qtkeff\(y,x\)\(y≠x\),Ltkeff\(x,x\)=−∑y≠xQtkeff\(y,x\),L\_\{t\_\{k\}\}^\{\\mathrm\{eff\}\}\(y,x\)=Q\_\{t\_\{k\}\}^\{\\mathrm\{eff\}\}\(y,x\)\\quad\(y\\neq x\),\\qquad L\_\{t\_\{k\}\}^\{\\mathrm\{eff\}\}\(x,x\)=\-\\sum\_\{y\\neq x\}Q\_\{t\_\{k\}\}^\{\\mathrm\{eff\}\}\(y,x\),and define
𝐏k:=exp\(Δtk𝐋tkeff\),Wk\(x\):=exp\(Δtkgtkeff\(x\)\),Δtk:=tk\+1−tk\.\\mathbf\{P\}\_\{k\}:=\\exp\\\!\\big\(\\Delta t\_\{k\}\\mathbf\{L\}\_\{t\_\{k\}\}^\{\\mathrm\{eff\}\}\\big\),\\qquad W\_\{k\}\(x\):=\\exp\\\!\\big\(\\Delta t\_\{k\}\\,g\_\{t\_\{k\}\}^\{\\mathrm\{eff\}\}\(x\)\\big\),\\qquad\\Delta t\_\{k\}:=t\_\{k\+1\}\-t\_\{k\}\.With the column convention,Pk\(y,x\)P\_\{k\}\(y,x\)is the transition probability from sourcexxto destinationyy, and\(𝖯kf\)\(x\):=\(𝐏k⊤f\)\(x\)=∑yPk\(y,x\)f\(y\)\(\\mathsf\{P\}\_\{k\}f\)\(x\):=\(\\mathbf\{P\}\_\{k\}^\{\\top\}f\)\(x\)=\\sum\_\{y\}P\_\{k\}\(y,x\)f\(y\)\. Assume0<infxWk\(x\)≤supxWk\(x\)<∞0<\\inf\_\{x\}W\_\{k\}\(x\)\\leq\\sup\_\{x\}W\_\{k\}\(x\)<\\inftyfor everykk\.
### D\.2Formal score\-ratio theorem and proof
###### Theorem D\.4\(Formal version of Theorem[4\.1](https://arxiv.org/html/2609.35947#S4.Thmtheorem1): known forward rates\)\.
Suppose Assumptions[D\.1](https://arxiv.org/html/2609.35947#A4.Thmassumption1)and[D\.2](https://arxiv.org/html/2609.35947#A4.Thmassumption2)hold\. Fora\>0a\>0, define
Ca,κ¯,κ¯:=supu∈\[κ¯,κ¯\]\|ua−1\|ℓ\(u\),Csc:=max\{Cγ,κ¯,κ¯,C1,κ¯,κ¯\}\.C\_\{a,\\underline\{\\kappa\},\\overline\{\\kappa\}\}:=\\sup\_\{u\\in\[\\underline\{\\kappa\},\\overline\{\\kappa\}\]\}\\frac\{\|u^\{a\}\-1\|\}\{\\sqrt\{\\ell\(u\)\}\},\\qquad C\_\{\\rm sc\}:=\\max\\\{C\_\{\\gamma,\\underline\{\\kappa\},\\overline\{\\kappa\}\},C\_\{1,\\underline\{\\kappa\},\\overline\{\\kappa\}\}\\\}\.\(D\.13\)Atu=1u=1the quotient is interpreted by continuous extension, equal to2a\\sqrt\{2\}a\. These constants are finite, and:
\(a\) Intermediate operator\-error form \(proof\-internal\)\.Let
ℰop\(t\):=∑x∈𝒳∑y≠xνtop\(x,y\)ℓ\(s^t\(x,y\)/st\(x,y\)\),\\mathcal\{E\}\_\{\\rm op\}\(t\):=\\sum\_\{x\\in\\mathcal\{X\}\}\\sum\_\{y\\neq x\}\\nu\_\{t\}^\{\\rm op\}\(x,y\)\\,\\ell\\\!\\big\(\\widehat\{s\}\_\{t\}\(x,y\)/s\_\{t\}\(x,y\)\\big\),\(D\.14\)whereνtop\\nu\_\{t\}^\{\\rm op\}is the proof\-internal operator\-error edge measure in \([D\.10](https://arxiv.org/html/2609.35947#A4.E10)\), with total massλ¯op\(t\)\\bar\{\\lambda\}\_\{\\rm op\}\(t\)\. Then
TV\(qTs^,qT\)≤Csce2BT∫0Tλ¯op\(t\)ℰop\(t\)𝑑t\.\\TV\(q\_\{T\}^\{\\widehat\{s\}\},q\_\{T\}\)\\leq C\_\{\\rm sc\}\\,e^\{2BT\}\\int\_\{0\}^\{T\}\\sqrt\{\\bar\{\\lambda\}\_\{\\rm op\}\(t\)\\mathcal\{E\}\_\{\\rm op\}\(t\)\}\\,\\mathop\{\}\\\!\\mathrm\{d\}t\.\(D\.15\)This bound is purely an internal step in the proof and is not to be interpreted as a population training objective\.
\(b\) Score\-entropy training\-loss form\.If, in addition, Assumption[D\.3](https://arxiv.org/html/2609.35947#A4.Thmassumption3)holds, then
TV\(qTs^,qT\)≤γe2BT∫0TCDD\(t\)ℒDD\(t\)𝑑t\.\\TV\(q\_\{T\}^\{\\widehat\{s\}\},q\_\{T\}\)\\leq\\gamma e^\{2BT\}\\int\_\{0\}^\{T\}C\_\{\\rm DD\}\(t\)\\sqrt\{\\mathcal\{L\}\_\{\\rm DD\}\(t\)\}\\,\\mathop\{\}\\\!\\mathrm\{d\}t\.\(D\.16\)
\(c\) Time\-integrated form \(main\-text statement\)\.With the integrated quantities𝔏DD\\mathfrak\{L\}\_\{\\rm DD\}in \([D\.6](https://arxiv.org/html/2609.35947#A4.Ex7)\) andℭDD\\mathfrak\{C\}\_\{\\rm DD\}in \([D\.9](https://arxiv.org/html/2609.35947#A4.E9)\), Cauchy–Schwarz inttapplied to \([D\.16](https://arxiv.org/html/2609.35947#A4.E16)\) gives
TV\(qTs^,qT\)≤γe2BTℭDD𝔏DD\.\\TV\(q\_\{T\}^\{\\widehat\{s\}\},q\_\{T\}\)\\leq\\gamma\\,e^\{2BT\}\\,\\mathfrak\{C\}\_\{\\rm DD\}\\,\\sqrt\{\\mathfrak\{L\}\_\{\\rm DD\}\}\.\(D\.17\)This is exactly the main\-text bound \([4\.2](https://arxiv.org/html/2609.35947#S4.E2)\)\.
Proof\.We first prove thatCsc<∞C\_\{\\rm sc\}<\\infty\. For any fixeda\>0a\>0, the mapu↦\|ua−1\|/ℓ\(u\)u\\mapsto\|u^\{a\}\-1\|/\\sqrt\{\\ell\(u\)\}is continuous on\(0,∞\)∖\{1\}\(0,\\infty\)\\setminus\\\{1\\\}\. Atu=1u=1, Taylor expansion givesua−1=a\(u−1\)\+O\(\(u−1\)2\)u^\{a\}\-1=a\(u\-1\)\+O\(\(u\-1\)^\{2\}\)andℓ\(u\)=\(u−1\)2/2\+O\(\(u−1\)3\)\\ell\(u\)=\(u\-1\)^\{2\}/2\+O\(\(u\-1\)^\{3\}\), hence the limit equals2a\\sqrt\{2\}a\. The function is therefore continuous on the compact interval\[κ¯,κ¯\]\[\\underline\{\\kappa\},\\overline\{\\kappa\}\], so both constants in \([D\.13](https://arxiv.org/html/2609.35947#A4.E13)\) are finite\.
Throughout, set
𝐀t:=𝐋~t\+diag\(G~t\),𝐀^t:=𝐋~^t\+diag\(G~^t\),\\mathbf\{A\}\_\{t\}:=\\widetilde\{\\mathbf\{L\}\}\_\{t\}\+\\diag\(\\widetilde\{G\}\_\{t\}\),\\qquad\\widehat\{\\mathbf\{A\}\}\_\{t\}:=\\widehat\{\\widetilde\{\\mathbf\{L\}\}\}\_\{t\}\+\\diag\(\\widehat\{\\widetilde\{G\}\}\_\{t\}\),so the unnormalized flows obey∂tρt=𝐀tρt\\partial\_\{t\}\\rho\_\{t\}=\\mathbf\{A\}\_\{t\}\\rho\_\{t\}and∂tρ^t=𝐀^tρ^t\\partial\_\{t\}\\widehat\{\\rho\}\_\{t\}=\\widehat\{\\mathbf\{A\}\}\_\{t\}\\widehat\{\\rho\}\_\{t\}with common initial condition\. WriteZt:=𝟏⊤ρtZ\_\{t\}:=\\mathbf\{1\}^\{\\top\}\\rho\_\{t\},Z^t:=𝟏⊤ρ^t\\widehat\{Z\}\_\{t\}:=\\mathbf\{1\}^\{\\top\}\\widehat\{\\rho\}\_\{t\}, andδt:=ρ^t−ρt\\delta\_\{t\}:=\\widehat\{\\rho\}\_\{t\}\-\\rho\_\{t\}\.
##### Mass control\.
Since each generator𝐋~t\\widetilde\{\\mathbf\{L\}\}\_\{t\}satisfies𝟏⊤𝐋~t=𝟎⊤\\mathbf\{1\}^\{\\top\}\\widetilde\{\\mathbf\{L\}\}\_\{t\}=\\mathbf\{0\}^\{\\top\}by construction,
∂tZt=𝟏⊤∂tρt=𝟏⊤diag\(G~t\)ρt=ρt\(G~t\),\|∂tZt\|≤‖G~t‖∞Zt≤BZt\.\\partial\_\{t\}Z\_\{t\}=\\mathbf\{1\}^\{\\top\}\\partial\_\{t\}\\rho\_\{t\}=\\mathbf\{1\}^\{\\top\}\\diag\(\\widetilde\{G\}\_\{t\}\)\\rho\_\{t\}=\\rho\_\{t\}\(\\widetilde\{G\}\_\{t\}\),\\qquad\|\\partial\_\{t\}Z\_\{t\}\|\\leq\\left\\\|\\widetilde\{G\}\_\{t\}\\right\\\|\_\{\\infty\}Z\_\{t\}\\leq BZ\_\{t\}\.Grönwall’s inequality givese−Bt≤Zt≤eBte^\{\-Bt\}\\leq Z\_\{t\}\\leq e^\{Bt\}, and the same argument withG~^t\\widehat\{\\widetilde\{G\}\}\_\{t\}givese−Bt≤Z^t≤eBte^\{\-Bt\}\\leq\\widehat\{Z\}\_\{t\}\\leq e^\{Bt\}\.
##### Operator perturbation from the local ratio\.
Letut\(x,y\):=s^t\(x,y\)/st\(x,y\)u\_\{t\}\(x,y\):=\\widehat\{s\}\_\{t\}\(x,y\)/s\_\{t\}\(x,y\)\. From Assumption[D\.2](https://arxiv.org/html/2609.35947#A4.Thmassumption2),
ΔQ~t\(y,x\):=Q~^t\(y,x\)−Q~t\(y,x\)=Q~t\(y,x\)\(ut\(x,y\)γ−1\),\\Delta\\widetilde\{Q\}\_\{t\}\(y,x\):=\\widehat\{\\widetilde\{Q\}\}\_\{t\}\(y,x\)\-\\widetilde\{Q\}\_\{t\}\(y,x\)=\\widetilde\{Q\}\_\{t\}\(y,x\)\\big\(u\_\{t\}\(x,y\)^\{\\gamma\}\-1\\big\),\(D\.18\)and
ΔQt←\(y,x\):=Q^t←\(y,x\)−Qt←\(y,x\)=Qt←\(y,x\)\(ut\(x,y\)−1\)\.\\Delta Q\_\{t\}^\{\\shortleftarrow\}\(y,x\):=\\widehat\{Q\}\_\{t\}^\{\\shortleftarrow\}\(y,x\)\-Q\_\{t\}^\{\\shortleftarrow\}\(y,x\)=Q\_\{t\}^\{\\shortleftarrow\}\(y,x\)\\big\(u\_\{t\}\(x,y\)\-1\\big\)\.\(D\.19\)The potential difference is
G~^t\(x\)−G~t\(x\)=∑y≠x\(ΔQ~t\(y,x\)−γΔQt←\(y,x\)\)\.\\widehat\{\\widetilde\{G\}\}\_\{t\}\(x\)\-\\widetilde\{G\}\_\{t\}\(x\)=\\sum\_\{y\\neq x\}\\Big\(\\Delta\\widetilde\{Q\}\_\{t\}\(y,x\)\-\\gamma\\Delta Q\_\{t\}^\{\\shortleftarrow\}\(y,x\)\\Big\)\.The diagonal entry of columnxxof𝐀^t−𝐀t\\widehat\{\\mathbf\{A\}\}\_\{t\}\-\\mathbf\{A\}\_\{t\}is therefore
−∑y≠xΔQ~t\(y,x\)\+∑y≠x\(ΔQ~t\(y,x\)−γΔQt←\(y,x\)\)=−γ∑y≠xΔQt←\(y,x\),\-\\sum\_\{y\\neq x\}\\Delta\\widetilde\{Q\}\_\{t\}\(y,x\)\+\\sum\_\{y\\neq x\}\\Big\(\\Delta\\widetilde\{Q\}\_\{t\}\(y,x\)\-\\gamma\\Delta Q\_\{t\}^\{\\shortleftarrow\}\(y,x\)\\Big\)=\-\\gamma\\sum\_\{y\\neq x\}\\Delta Q\_\{t\}^\{\\shortleftarrow\}\(y,x\),while the off\-diagonal entry from sourcexxto destinationyyisΔQ~t\(y,x\)\\Delta\\widetilde\{Q\}\_\{t\}\(y,x\)\. Consequently, for any vectorvvand any statezz,
\(\(𝐀^t−𝐀t\)v\)\(z\)=∑x≠zΔQ~t\(z,x\)v\(x\)−γv\(z\)∑y≠zΔQt←\(y,z\)\.\\big\(\(\\widehat\{\\mathbf\{A\}\}\_\{t\}\-\\mathbf\{A\}\_\{t\}\)v\\big\)\(z\)=\\sum\_\{x\\neq z\}\\Delta\\widetilde\{Q\}\_\{t\}\(z,x\)v\(x\)\-\\gamma v\(z\)\\sum\_\{y\\neq z\}\\Delta Q\_\{t\}^\{\\shortleftarrow\}\(y,z\)\.\(D\.20\)Usingρt=Ztqt\\rho\_\{t\}=Z\_\{t\}q\_\{t\}and the triangle inequality,
‖\(𝐀^t−𝐀t\)ρt‖1\\displaystyle\\left\\\|\(\\widehat\{\\mathbf\{A\}\}\_\{t\}\-\\mathbf\{A\}\_\{t\}\)\\rho\_\{t\}\\right\\\|\_\{1\}≤Zt∑xqt\(x\)∑y≠xQ~t\(y,x\)\|ut\(x,y\)γ−1\|\\displaystyle\\leq Z\_\{t\}\\sum\_\{x\}q\_\{t\}\(x\)\\sum\_\{y\\neq x\}\\widetilde\{Q\}\_\{t\}\(y,x\)\|u\_\{t\}\(x,y\)^\{\\gamma\}\-1\|\+Ztγ∑xqt\(x\)∑y≠xQt←\(y,x\)\|ut\(x,y\)−1\|\.\\displaystyle\\quad\+Z\_\{t\}\\gamma\\sum\_\{x\}q\_\{t\}\(x\)\\sum\_\{y\\neq x\}Q\_\{t\}^\{\\shortleftarrow\}\(y,x\)\|u\_\{t\}\(x,y\)\-1\|\.\(D\.21\)
##### Internal operator\-error loss\.
By \([D\.13](https://arxiv.org/html/2609.35947#A4.E13)\),\|uγ−1\|≤Cscℓ\(u\)\|u^\{\\gamma\}\-1\|\\leq C\_\{\\rm sc\}\\sqrt\{\\ell\(u\)\}and\|u−1\|≤Cscℓ\(u\)\|u\-1\|\\leq C\_\{\\rm sc\}\\sqrt\{\\ell\(u\)\}on the ratio window\. Withνtop\\nu\_\{t\}^\{\\rm op\}as in \([D\.10](https://arxiv.org/html/2609.35947#A4.E10)\), the right\-hand side of \([D\.21](https://arxiv.org/html/2609.35947#A4.E21)\) is bounded by
ZtCsc∑x,y:x≠yνtop\(x,y\)ℓ\(ut\(x,y\)\)\.Z\_\{t\}C\_\{\\rm sc\}\\sum\_\{x,y:x\\neq y\}\\nu\_\{t\}^\{\\rm op\}\(x,y\)\\sqrt\{\\ell\(u\_\{t\}\(x,y\)\)\}\.Cauchy–Schwarz under the nonnegative measureνtop\\nu\_\{t\}^\{\\rm op\}gives
∑x,y:x≠yνtop\(x,y\)ℓ\(ut\(x,y\)\)\\displaystyle\\sum\_\{x,y:x\\neq y\}\\nu\_\{t\}^\{\\rm op\}\(x,y\)\\sqrt\{\\ell\(u\_\{t\}\(x,y\)\)\}≤∑x,y:x≠yνtop\(x,y\)∑x,y:x≠yνtop\(x,y\)ℓ\(ut\(x,y\)\)\\displaystyle\\leq\\sqrt\{\\sum\_\{x,y:x\\neq y\}\\nu\_\{t\}^\{\\rm op\}\(x,y\)\}\\sqrt\{\\sum\_\{x,y:x\\neq y\}\\nu\_\{t\}^\{\\rm op\}\(x,y\)\\ell\(u\_\{t\}\(x,y\)\)\}=λ¯op\(t\)ℰop\(t\)\.\\displaystyle=\\sqrt\{\\bar\{\\lambda\}\_\{\\rm op\}\(t\)\\mathcal\{E\}\_\{\\rm op\}\(t\)\}\.\(D\.22\)Combining the preceding displays,
‖\(𝐀^t−𝐀t\)ρt‖1≤ZtCscλ¯op\(t\)ℰop\(t\)\.\\left\\\|\(\\widehat\{\\mathbf\{A\}\}\_\{t\}\-\\mathbf\{A\}\_\{t\}\)\\rho\_\{t\}\\right\\\|\_\{1\}\\leq Z\_\{t\}C\_\{\\rm sc\}\\sqrt\{\\bar\{\\lambda\}\_\{\\rm op\}\(t\)\\mathcal\{E\}\_\{\\rm op\}\(t\)\}\.\(D\.23\)This is the only place where the edgewise local\-ratio loss enters; no supremum over states is used in the internal bound\.
##### Duhamel stability\.
The errorδt=ρ^t−ρt\\delta\_\{t\}=\\widehat\{\\rho\}\_\{t\}\-\\rho\_\{t\}satisfiesδ0=0\\delta\_\{0\}=0and
∂tδt=𝐀^tδt\+\(𝐀^t−𝐀t\)ρt\.\\partial\_\{t\}\\delta\_\{t\}=\\widehat\{\\mathbf\{A\}\}\_\{t\}\\delta\_\{t\}\+\(\\widehat\{\\mathbf\{A\}\}\_\{t\}\-\\mathbf\{A\}\_\{t\}\)\\rho\_\{t\}\.LetΦ^T,s\\widehat\{\\Phi\}\_\{T,s\}denote the evolution operator of∂τv=𝐀^τv\\partial\_\{\\tau\}v=\\widehat\{\\mathbf\{A\}\}\_\{\\tau\}vforτ∈\[s,T\]\\tau\\in\[s,T\]\. Because𝐀^t=𝐋~^t\+diag\(G~^t\)\\widehat\{\\mathbf\{A\}\}\_\{t\}=\\widehat\{\\widetilde\{\\mathbf\{L\}\}\}\_\{t\}\+\\diag\(\\widehat\{\\widetilde\{G\}\}\_\{t\}\)is Metzler and the columns of𝐋~^t\\widehat\{\\widetilde\{\\mathbf\{L\}\}\}\_\{t\}sum to zero, itsℓ1\\ell^\{1\}logarithmic norm is
μ1\(𝐀^t\)=maxx\(\(𝐀^t\)xx\+∑y≠x\|\(𝐀^t\)yx\|\)=maxxG~^t\(x\)≤B\.\\mu\_\{1\}\(\\widehat\{\\mathbf\{A\}\}\_\{t\}\)=\\max\_\{x\}\\left\(\(\\widehat\{\\mathbf\{A\}\}\_\{t\}\)\_\{xx\}\+\\sum\_\{y\\neq x\}\|\(\\widehat\{\\mathbf\{A\}\}\_\{t\}\)\_\{yx\}\|\\right\)=\\max\_\{x\}\\widehat\{\\widetilde\{G\}\}\_\{t\}\(x\)\\leq B\.The logarithmic\-norm propagator estimate gives‖Φ^T,s‖1→1≤exp\(B\(T−s\)\)\\left\\\|\\widehat\{\\Phi\}\_\{T,s\}\\right\\\|\_\{1\\to 1\}\\leq\\exp\(B\(T\-s\)\)\. By Duhamel’s principle and \([D\.23](https://arxiv.org/html/2609.35947#A4.E23)\),
‖δT‖1\\displaystyle\\left\\\|\\delta\_\{T\}\\right\\\|\_\{1\}≤∫0Texp\(B\(T−s\)\)ZsCscλ¯op\(s\)ℰop\(s\)𝑑s\\displaystyle\\leq\\int\_\{0\}^\{T\}\\exp\\\!\\big\(B\(T\-s\)\\big\)\\,Z\_\{s\}\\,C\_\{\\rm sc\}\\sqrt\{\\bar\{\\lambda\}\_\{\\rm op\}\(s\)\\mathcal\{E\}\_\{\\rm op\}\(s\)\}\\,\\mathop\{\}\\\!\\mathrm\{d\}s≤CsceBT∫0Tλ¯op\(s\)ℰop\(s\)𝑑s\.\\displaystyle\\leq C\_\{\\rm sc\}e^\{BT\}\\int\_\{0\}^\{T\}\\sqrt\{\\bar\{\\lambda\}\_\{\\rm op\}\(s\)\\mathcal\{E\}\_\{\\rm op\}\(s\)\}\\,\\mathop\{\}\\\!\\mathrm\{d\}s\.\(D\.24\)
##### Normalization\.
For nonnegative non\-zero vectorsa,ba,b,
‖a𝟏⊤a−b𝟏⊤b‖1≤2‖a−b‖1min\{𝟏⊤a,𝟏⊤b\}\.\\left\\\|\\frac\{a\}\{\\mathbf\{1\}^\{\\top\}a\}\-\\frac\{b\}\{\\mathbf\{1\}^\{\\top\}b\}\\right\\\|\_\{1\}\\leq\\frac\{2\\left\\\|a\-b\\right\\\|\_\{1\}\}\{\\min\\\{\\mathbf\{1\}^\{\\top\}a,\\mathbf\{1\}^\{\\top\}b\\\}\}\.Applying this witha=ρ^Ta=\\widehat\{\\rho\}\_\{T\},b=ρTb=\\rho\_\{T\}and dividing by22givesTV\(qTs^,qT\)≤‖δT‖1/min\{Z^T,ZT\}≤eBT‖δT‖1\\TV\(q\_\{T\}^\{\\widehat\{s\}\},q\_\{T\}\)\\leq\\left\\\|\\delta\_\{T\}\\right\\\|\_\{1\}/\\min\\\{\\widehat\{Z\}\_\{T\},Z\_\{T\}\\\}\\leq e^\{BT\}\\left\\\|\\delta\_\{T\}\\right\\\|\_\{1\}\. Combined with \([D\.24](https://arxiv.org/html/2609.35947#A4.E24)\), this proves part \(a\), namely \([D\.15](https://arxiv.org/html/2609.35947#A4.E15)\)\.
##### Score\-entropy training\-loss comparison\.
For the sharper training\-loss form, return to \([D\.21](https://arxiv.org/html/2609.35947#A4.E21)\)\. The identities in \([D\.11](https://arxiv.org/html/2609.35947#A4.E11)\) give
Zt−1‖\(𝐀^t−𝐀t\)ρt‖1\\displaystyle Z\_\{t\}^\{\-1\}\\left\\\|\(\\widehat\{\\mathbf\{A\}\}\_\{t\}\-\\mathbf\{A\}\_\{t\}\)\\rho\_\{t\}\\right\\\|\_\{1\}≤γ∑x,y:x≠yνtSE\(x,y\)\(Cγ,κ¯,κ¯ht\(y\)\+C1,κ¯,κ¯ht\(x\)\)ℓ\(ut\(x,y\)\)\\displaystyle\\leq\\gamma\\sum\_\{x,y:x\\neq y\}\\nu\_\{t\}^\{\\rm SE\}\(x,y\)\\big\(C\_\{\\gamma,\\underline\{\\kappa\},\\overline\{\\kappa\}\}h\_\{t\}\(y\)\+C\_\{1,\\underline\{\\kappa\},\\overline\{\\kappa\}\}h\_\{t\}\(x\)\\big\)\\sqrt\{\\ell\(u\_\{t\}\(x,y\)\)\}≤γCDD\(t\)ℒDD\(t\)\.\\displaystyle\\leq\\gamma\\,C\_\{\\rm DD\}\(t\)\\sqrt\{\\mathcal\{L\}\_\{\\rm DD\}\(t\)\}\.\(D\.25\)Substituting \([D\.25](https://arxiv.org/html/2609.35947#A4.E25)\) into the same Duhamel and normalization argument proves \([D\.16](https://arxiv.org/html/2609.35947#A4.E16)\)\.
##### Time\-integrated form\.
Applying Cauchy–Schwarz inttto \([D\.16](https://arxiv.org/html/2609.35947#A4.E16)\),
∫0TCDD\(t\)ℒDD\(t\)𝑑t≤∫0TCDD\(t\)2𝑑t∫0TℒDD\(t\)𝑑t=ℭDD𝔏DD,\\int\_\{0\}^\{T\}C\_\{\\rm DD\}\(t\)\\sqrt\{\\mathcal\{L\}\_\{\\rm DD\}\(t\)\}\\,\\mathop\{\}\\\!\\mathrm\{d\}t\\leq\\sqrt\{\\int\_\{0\}^\{T\}C\_\{\\rm DD\}\(t\)^\{2\}\\,\\mathop\{\}\\\!\\mathrm\{d\}t\}\\,\\sqrt\{\\int\_\{0\}^\{T\}\\mathcal\{L\}\_\{\\rm DD\}\(t\)\\,\\mathop\{\}\\\!\\mathrm\{d\}t\}=\\mathfrak\{C\}\_\{\\rm DD\}\\,\\sqrt\{\\mathfrak\{L\}\_\{\\rm DD\}\},which gives \([D\.17](https://arxiv.org/html/2609.35947#A4.E17)\)\.□\\square
### D\.3Formal particle theorem and proof
###### Theorem D\.5\(Formal version of Theorem[4\.2](https://arxiv.org/html/2609.35947#S4.Thmtheorem2)\)\.
Under Assumption[D\.4](https://arxiv.org/html/2609.35947#A4.Thmassumption4), define the exact fixed\-grid Feynman–Kac recursion by
q0Δ:=qt0,qk\+1Δ\(f\):=qkΔ\(Wk𝖯kf\)qkΔ\(Wk\),k=0,…,M−1\.q\_\{0\}^\{\\Delta\}:=q\_\{t\_\{0\}\},\\qquad q\_\{k\+1\}^\{\\Delta\}\(f\):=\\frac\{q\_\{k\}^\{\\Delta\}\(W\_\{k\}\\mathsf\{P\}\_\{k\}f\)\}\{q\_\{k\}^\{\\Delta\}\(W\_\{k\}\)\},\\qquad k=0,\\ldots,M\-1\.\(D\.26\)InitializeNNparticlesX01,…,X0NX\_\{0\}^\{1\},\\ldots,X\_\{0\}^\{N\}i\.i\.d\. fromq0Δq\_\{0\}^\{\\Delta\}and setq0N,Δ:=N−1∑i=1NδX0iq\_\{0\}^\{N,\\Delta\}:=N^\{\-1\}\\sum\_\{i=1\}^\{N\}\\delta\_\{X\_\{0\}^\{i\}\}\. At each stepkk, conditionally on the current particles, draw ancestorsAk1,…,AkNA\_\{k\}^\{1\},\\ldots,A\_\{k\}^\{N\}multinomially with probabilities proportional toWk\(Xkj\)W\_\{k\}\(X\_\{k\}^\{j\}\)over candidate ancestorsj=1,…,Nj=1,\\ldots,N, and then draw
Pr\(Xk\+1i=y∣XkAki=x\)=Pk\(y,x\)independently overi\.\\Pr\\\!\\left\(X\_\{k\+1\}^\{i\}=y\\mid X\_\{k\}^\{A\_\{k\}^\{i\}\}=x\\right\)=P\_\{k\}\(y,x\)\\quad\\text\{independently over \}i\.LetqkN,Δ:=N−1∑i=1NδXkiq\_\{k\}^\{N,\\Delta\}:=N^\{\-1\}\\sum\_\{i=1\}^\{N\}\\delta\_\{X\_\{k\}^\{i\}\}\. Define
β¯k,M:=∏j=kM−1βj,βj:=supxWj\(x\)infxWj\(x\),\\bar\{\\beta\}\_\{k,M\}:=\\prod\_\{j=k\}^\{M\-1\}\\beta\_\{j\},\\qquad\\beta\_\{j\}:=\\frac\{\\sup\_\{x\}W\_\{j\}\(x\)\}\{\\inf\_\{x\}W\_\{j\}\(x\)\},with the convention that an empty product is11\. Then, for every boundedf:𝒳→ℝf:\\mathcal\{X\}\\to\\mathbb\{R\},
𝔼\[\|qMN,Δ\(f\)−qMΔ\(f\)\|\]≤osc\(f\)2N∑k=0Mβ¯k,M\.\\mathbb\{E\}\\Big\[\\big\|q\_\{M\}^\{N,\\Delta\}\(f\)\-q\_\{M\}^\{\\Delta\}\(f\)\\big\|\\Big\]\\leq\\frac\{\\osc\(f\)\}\{2\\sqrt\{N\}\}\\sum\_\{k=0\}^\{M\}\\bar\{\\beta\}\_\{k,M\}\.\(D\.27\)
Proof\.The proof is self\-contained for the deterministic every\-step bootstrap Feynman–Kac system above\. It does not use score\-estimation, adaptive\-control, or time\-discretization arguments\. We writeℱk:=σ\(X01:N,…,Xk1:N\)\\mathcal\{F\}\_\{k\}:=\\sigma\(X\_\{0\}^\{1:N\},\\ldots,X\_\{k\}^\{1:N\}\)for the natural filtration of the particle system at timekk\.
##### Conditional sampling\.
For any probability measureμ\\muon𝒳\\mathcal\{X\}, define the one\-step Feynman–Kac map
Φk\(μ\)\(f\):=μ\(Wk𝖯kf\)μ\(Wk\),\\Phi\_\{k\}\(\\mu\)\(f\):=\\frac\{\\mu\(W\_\{k\}\\mathsf\{P\}\_\{k\}f\)\}\{\\mu\(W\_\{k\}\)\},which lies betweeninfx𝖯kf\(x\)\\inf\_\{x\}\\mathsf\{P\}\_\{k\}f\(x\)andsupx𝖯kf\(x\)\\sup\_\{x\}\\mathsf\{P\}\_\{k\}f\(x\)and so defines a probability measure on𝒳\\mathcal\{X\}\. The exact recursion isqk\+1Δ=Φk\(qkΔ\)q\_\{k\+1\}^\{\\Delta\}=\\Phi\_\{k\}\(q\_\{k\}^\{\\Delta\}\)\. By the multinomial\-resampling\-then\-mutation update, conditional onℱk\\mathcal\{F\}\_\{k\}:
1. 1\.the ancestor indicesAk1,…,AkNA\_\{k\}^\{1\},\\ldots,A\_\{k\}^\{N\}are i\.i\.d\. withPr\(Aki=j∣ℱk\)=Wk\(Xkj\)∑l=1NWk\(Xkl\);\\Pr\(A\_\{k\}^\{i\}=j\\mid\\mathcal\{F\}\_\{k\}\)=\\frac\{W\_\{k\}\(X\_\{k\}^\{j\}\)\}\{\\sum\_\{l=1\}^\{N\}W\_\{k\}\(X\_\{k\}^\{l\}\)\};
2. 2\.given the ancestors, the new particles are drawn independently withPr\(Xk\+1i=y∣XkAki=x\)=Pk\(y,x\)\\Pr\(X\_\{k\+1\}^\{i\}=y\\mid X\_\{k\}^\{A\_\{k\}^\{i\}\}=x\)=P\_\{k\}\(y,x\)\.
Hence, for any boundedf:𝒳→ℝf:\\mathcal\{X\}\\to\\mathbb\{R\}and anyii,
𝔼\[f\(Xk\+1i\)∣ℱk\]=∑j=1NWk\(Xkj\)∑lWk\(Xkl\)\(𝖯kf\)\(Xkj\)=qkN,Δ\(Wk𝖯kf\)qkN,Δ\(Wk\)=Φk\(qkN,Δ\)\(f\),\\mathbb\{E\}\[f\(X\_\{k\+1\}^\{i\}\)\\mid\\mathcal\{F\}\_\{k\}\]=\\sum\_\{j=1\}^\{N\}\\frac\{W\_\{k\}\(X\_\{k\}^\{j\}\)\}\{\\sum\_\{l\}W\_\{k\}\(X\_\{k\}^\{l\}\)\}\\,\(\\mathsf\{P\}\_\{k\}f\)\(X\_\{k\}^\{j\}\)=\\frac\{q\_\{k\}^\{N,\\Delta\}\(W\_\{k\}\\mathsf\{P\}\_\{k\}f\)\}\{q\_\{k\}^\{N,\\Delta\}\(W\_\{k\}\)\}=\\Phi\_\{k\}\(q\_\{k\}^\{N,\\Delta\}\)\(f\),and theXk\+1iX\_\{k\+1\}^\{i\}are i\.i\.d\. givenℱk\\mathcal\{F\}\_\{k\}\. Equivalently,qk\+1N,Δq\_\{k\+1\}^\{N,\\Delta\}is the empirical measure ofNNi\.i\.d\. samples fromμk\+1N:=Φk\(qkN,Δ\)\\mu\_\{k\+1\}^\{N\}:=\\Phi\_\{k\}\(q\_\{k\}^\{N,\\Delta\}\)\.*This conditional i\.i\.d\. property is the only place the every\-step bootstrap structure is used in the entire proof\.*
##### Future Feynman–Kac weights\.
For0≤k≤ℓ≤M0\\leq k\\leq\\ell\\leq M, define the future Feynman–Kac semigroup on test functions by
𝖧k,ℓ:=\(Wk𝖯k\)\(Wk\+1𝖯k\+1\)⋯\(Wℓ−1𝖯ℓ−1\),𝖧M,M:=I\.\\mathsf\{H\}\_\{k,\\ell\}:=\(W\_\{k\}\\mathsf\{P\}\_\{k\}\)\(W\_\{k\+1\}\\mathsf\{P\}\_\{k\+1\}\)\\cdots\(W\_\{\\ell\-1\}\\mathsf\{P\}\_\{\\ell\-1\}\),\\qquad\\mathsf\{H\}\_\{M,M\}:=I\.Probabilistically,𝖧k,ℓf\(x\)=𝔼x\[\(∏j=kℓ−1Wj\(Xj\)\)f\(Xℓ\)\]\\mathsf\{H\}\_\{k,\\ell\}f\(x\)=\\mathbb\{E\}\_\{x\}\\big\[\\big\(\\prod\_\{j=k\}^\{\\ell\-1\}W\_\{j\}\(X\_\{j\}\)\\big\)\\,f\(X\_\{\\ell\}\)\\big\]where\(Xj\)j=kℓ\(X\_\{j\}\)\_\{j=k\}^\{\\ell\}is the inhomogeneous Markov chain withXk=xX\_\{k\}=xandPr\(Xj\+1=y∣Xj=x\)=Pj\(y,x\)\\Pr\(X\_\{j\+1\}=y\\mid X\_\{j\}=x\)=P\_\{j\}\(y,x\)\. Set
mk,ℓ:=∏j=kℓ−1infxWj\(x\),Mk,ℓ:=∏j=kℓ−1supxWj\(x\),Bk,ℓ:=Mk,ℓmk,ℓ=∏j=kℓ−1βj\.m\_\{k,\\ell\}:=\\prod\_\{j=k\}^\{\\ell\-1\}\\inf\_\{x\}W\_\{j\}\(x\),\\qquad M\_\{k,\\ell\}:=\\prod\_\{j=k\}^\{\\ell\-1\}\\sup\_\{x\}W\_\{j\}\(x\),\\qquad B\_\{k,\\ell\}:=\\frac\{M\_\{k,\\ell\}\}\{m\_\{k,\\ell\}\}=\\prod\_\{j=k\}^\{\\ell\-1\}\\beta\_\{j\}\.Because each𝖯j\\mathsf\{P\}\_\{j\}is a Markov averaging operator and eachWjW\_\{j\}is positive,
mk,ℓ≤𝖧k,ℓ𝟏\(x\)≤Mk,ℓ,\|𝖧k,ℓf\(x\)\|≤Mk,ℓ‖f‖∞m\_\{k,\\ell\}\\leq\\mathsf\{H\}\_\{k,\\ell\}\\mathbf\{1\}\(x\)\\leq M\_\{k,\\ell\},\\qquad\|\\mathsf\{H\}\_\{k,\\ell\}f\(x\)\|\\leq M\_\{k,\\ell\}\\left\\\|f\\right\\\|\_\{\\infty\}\(D\.28\)for everyxx\.
##### Empirical error after a positive transform\.
LetRRbe a positive operator withm≤R𝟏≤m¯m\\leq R\\mathbf\{1\}\\leq\\overline\{m\}\. For a probability measureζ\\zetaand the empirical measureζN\\zeta^\{N\}ofNNi\.i\.d\. samples fromζ\\zeta, defineKRf:=Rf/\(R𝟏\)K\_\{R\}f:=Rf/\(R\\mathbf\{1\}\)\. We claim
𝔼\[\|ζN\(Rf\)ζN\(R𝟏\)−ζ\(Rf\)ζ\(R𝟏\)\|\|ζ\]≤m¯2mNosc\(KRf\)≤m¯2mNosc\(f\)\.\\mathbb\{E\}\\\!\\left\[\\left\|\\frac\{\\zeta^\{N\}\(Rf\)\}\{\\zeta^\{N\}\(R\\mathbf\{1\}\)\}\-\\frac\{\\zeta\(Rf\)\}\{\\zeta\(R\\mathbf\{1\}\)\}\\right\|\\,\\middle\|\\,\\zeta\\right\]\\leq\\frac\{\\overline\{m\}\}\{2m\\sqrt\{N\}\}\\osc\(K\_\{R\}f\)\\leq\\frac\{\\overline\{m\}\}\{2m\\sqrt\{N\}\}\\osc\(f\)\.\(D\.29\)*Proof of*\([D\.29](https://arxiv.org/html/2609.35947#A4.E29)\)\. Seta:=ζ\(Rf\)/ζ\(R𝟏\)a:=\\zeta\(Rf\)/\\zeta\(R\\mathbf\{1\}\)\. Thenζ\(R\(f−a\)\)=ζ\(Rf\)−aζ\(R𝟏\)=0\\zeta\(R\(f\-a\)\)=\\zeta\(Rf\)\-a\\zeta\(R\\mathbf\{1\}\)=0, so
ζN\(Rf\)ζN\(R𝟏\)−ζ\(Rf\)ζ\(R𝟏\)=ζN\(R\(f−a\)\)ζN\(R𝟏\)=\(ζN−ζ\)\(R\(f−a\)\)ζN\(R𝟏\)\.\\frac\{\\zeta^\{N\}\(Rf\)\}\{\\zeta^\{N\}\(R\\mathbf\{1\}\)\}\-\\frac\{\\zeta\(Rf\)\}\{\\zeta\(R\\mathbf\{1\}\)\}=\\frac\{\\zeta^\{N\}\(R\(f\-a\)\)\}\{\\zeta^\{N\}\(R\\mathbf\{1\}\)\}=\\frac\{\(\\zeta^\{N\}\-\\zeta\)\(R\(f\-a\)\)\}\{\\zeta^\{N\}\(R\\mathbf\{1\}\)\}\.NowKRfK\_\{R\}fis a pointwise positive average offf, soaais anR𝟏ζR\\mathbf\{1\}\\,\\zeta\-weighted average ofKRfK\_\{R\}f, hence\|KRf\(x\)−a\|≤supKRf−infKRf=osc\(KRf\)\|K\_\{R\}f\(x\)\-a\|\\leq\\sup K\_\{R\}f\-\\inf K\_\{R\}f=\\osc\(K\_\{R\}f\)for allxx\. Writingg:=R\(f−a\)=R𝟏⋅\(KRf−a\)g:=R\(f\-a\)=R\\mathbf\{1\}\\cdot\(K\_\{R\}f\-a\)and using0≤R𝟏≤m¯0\\leq R\\mathbf\{1\}\\leq\\overline\{m\}, we getosc\(g\)≤m¯osc\(KRf\)\\osc\(g\)\\leq\\overline\{m\}\\osc\(K\_\{R\}f\)\. Popoviciu’s variance inequality givesVarζ\(g\)≤osc\(g\)2/4≤m¯2osc\(KRf\)2/4\\operatorname\{Var\}\_\{\\zeta\}\(g\)\\leq\\osc\(g\)^\{2\}/4\\leq\\overline\{m\}^\{2\}\\osc\(K\_\{R\}f\)^\{2\}/4\. The empirical\-mean variance bound𝔼\[\|\(ζN−ζ\)\(g\)\|∣ζ\]≤Varζ\(g\)/N\\mathbb\{E\}\[\|\(\\zeta^\{N\}\-\\zeta\)\(g\)\|\\mid\\zeta\]\\leq\\sqrt\{\\operatorname\{Var\}\_\{\\zeta\}\(g\)/N\}andζN\(R𝟏\)≥m\\zeta^\{N\}\(R\\mathbf\{1\}\)\\geq mdeterministically yield the first inequality in \([D\.29](https://arxiv.org/html/2609.35947#A4.E29)\)\. The second inequality isosc\(KRf\)≤osc\(f\)\\osc\(K\_\{R\}f\)\\leq\\osc\(f\), which holds becauseKRK\_\{R\}is a positive averaging operator\.
##### Nonlinear telescoping\.
For0≤k≤M0\\leq k\\leq Mdefine the multi\-step Feynman–Kac map
Φk,M\(μ\)\(f\):=μ\(𝖧k,Mf\)μ\(𝖧k,M𝟏\)\.\\Phi\_\{k,M\}\(\\mu\)\(f\):=\\frac\{\\mu\(\\mathsf\{H\}\_\{k,M\}f\)\}\{\\mu\(\\mathsf\{H\}\_\{k,M\}\\mathbf\{1\}\)\}\.Setμ0N:=q0Δ\\mu\_\{0\}^\{N\}:=q\_\{0\}^\{\\Delta\}andμkN:=Φk−1\(qk−1N,Δ\)\\mu\_\{k\}^\{N\}:=\\Phi\_\{k\-1\}\(q\_\{k\-1\}^\{N,\\Delta\}\)fork≥1k\\geq 1\. By the conditional\-sampling paragraph, conditional onℱk−1\\mathcal\{F\}\_\{k\-1\},qkN,Δq\_\{k\}^\{N,\\Delta\}is the empirical measure ofNNi\.i\.d\. samples fromμkN\\mu\_\{k\}^\{N\}\.
The flow identityΦk,M∘Φk−1=Φk−1,M\\Phi\_\{k,M\}\\circ\\Phi\_\{k\-1\}=\\Phi\_\{k\-1,M\}\(which is just the chain rule for the Feynman–Kac semigroup; both sides equalμ↦μ\(𝖧k−1,Mf\)/μ\(𝖧k−1,M𝟏\)\\mu\\mapsto\\mu\(\\mathsf\{H\}\_\{k\-1,M\}f\)/\\mu\(\\mathsf\{H\}\_\{k\-1,M\}\\mathbf\{1\}\)since𝖧k−1,M=\(Wk−1𝖯k−1\)𝖧k,M\\mathsf\{H\}\_\{k\-1,M\}=\(W\_\{k\-1\}\\mathsf\{P\}\_\{k\-1\}\)\\mathsf\{H\}\_\{k,M\}\) gives, fork≥1k\\geq 1,
Φk,M\(μkN\)=Φk,M\(Φk−1\(qk−1N,Δ\)\)=Φk−1,M\(qk−1N,Δ\)\.\\Phi\_\{k,M\}\(\\mu\_\{k\}^\{N\}\)=\\Phi\_\{k,M\}\\big\(\\Phi\_\{k\-1\}\(q\_\{k\-1\}^\{N,\\Delta\}\)\\big\)=\\Phi\_\{k\-1,M\}\(q\_\{k\-1\}^\{N,\\Delta\}\)\.Together withΦ0,M\(μ0N\)=Φ0,M\(q0Δ\)=qMΔ\\Phi\_\{0,M\}\(\\mu\_\{0\}^\{N\}\)=\\Phi\_\{0,M\}\(q\_\{0\}^\{\\Delta\}\)=q\_\{M\}^\{\\Delta\}andΦM,M\(qMN,Δ\)=qMN,Δ\\Phi\_\{M,M\}\(q\_\{M\}^\{N,\\Delta\}\)=q\_\{M\}^\{N,\\Delta\}, this yields the telescoping
qMN,Δ\(f\)−qMΔ\(f\)=∑k=0M\[Φk,M\(qkN,Δ\)\(f\)−Φk,M\(μkN\)\(f\)\]\.q\_\{M\}^\{N,\\Delta\}\(f\)\-q\_\{M\}^\{\\Delta\}\(f\)=\\sum\_\{k=0\}^\{M\}\\left\[\\Phi\_\{k,M\}\(q\_\{k\}^\{N,\\Delta\}\)\(f\)\-\\Phi\_\{k,M\}\(\\mu\_\{k\}^\{N\}\)\(f\)\\right\]\.\(D\.30\)Thek=0k=0term is the initial\-sampling error; the termsk≥1k\\geq 1are the errors introduced by resampling and mutation at later grid times\.
##### Summing the transported errors\.
For eachk∈\{0,…,M\}k\\in\\\{0,\\ldots,M\\\}, apply \([D\.29](https://arxiv.org/html/2609.35947#A4.E29)\) conditionally onℱk−1\\mathcal\{F\}\_\{k\-1\}\(orℱ−1:=σ\(∅\)\\mathcal\{F\}\_\{\-1\}:=\\sigma\(\\emptyset\)fork=0k=0\) withζ=μkN\\zeta=\\mu\_\{k\}^\{N\},ζN=qkN,Δ\\zeta^\{N\}=q\_\{k\}^\{N,\\Delta\}, andR=𝖧k,MR=\\mathsf\{H\}\_\{k,M\}\. By \([D\.28](https://arxiv.org/html/2609.35947#A4.E28)\),m¯/m=Bk,M=β¯k,M\\overline\{m\}/m=B\_\{k,M\}=\\bar\{\\beta\}\_\{k,M\}, hence
𝔼\[\|Φk,M\(qkN,Δ\)\(f\)−Φk,M\(μkN\)\(f\)\|\|ℱk−1\]≤β¯k,Mosc\(f\)2N\.\\mathbb\{E\}\\\!\\left\[\\Big\|\\Phi\_\{k,M\}\(q\_\{k\}^\{N,\\Delta\}\)\(f\)\-\\Phi\_\{k,M\}\(\\mu\_\{k\}^\{N\}\)\(f\)\\Big\|\\,\\Big\|\\,\\mathcal\{F\}\_\{k\-1\}\\right\]\\leq\\frac\{\\bar\{\\beta\}\_\{k,M\}\\osc\(f\)\}\{2\\sqrt\{N\}\}\.Taking the unconditional expectation \(the right\-hand side is deterministic\) and summing overkkvia \([D\.30](https://arxiv.org/html/2609.35947#A4.E30)\) and the triangle inequality proves \([D\.27](https://arxiv.org/html/2609.35947#A4.E27)\)\.□\\square
### D\.4Incremental weight variance
The next lemma supports the discussion at the end of Section[4](https://arxiv.org/html/2609.35947#S4.SS0.SSS0.Px2): it converts the variance\-control objective of Section[3](https://arxiv.org/html/2609.35947#S3)into a quantitative bound on the variance of the centered one\-step incremental weightW¯k\(x\):=exp\(Δtkg¯k\(x\)\)\\bar\{W\}\_\{k\}\(x\):=\\exp\(\\Delta t\_\{k\}\\bar\{g\}\_\{k\}\(x\)\), whereg¯k\(x\):=gtkeff\(x\)−qkΔ\(gtkeff\)\\bar\{g\}\_\{k\}\(x\):=g\_\{t\_\{k\}\}^\{\\mathrm\{eff\}\}\(x\)\-q\_\{k\}^\{\\Delta\}\(g\_\{t\_\{k\}\}^\{\\mathrm\{eff\}\}\)is the residual centered potential\.
###### Lemma D\.6\(One\-step weight\-variance bound\)\.
For each grid stepkk,
VarqkΔ\(W¯k\)≤Δtk2exp\(2Δtk‖g¯k‖∞\)VarqkΔ\(gtkeff\)\.\\operatorname\{Var\}\_\{q\_\{k\}^\{\\Delta\}\}\\\!\\big\(\\bar\{W\}\_\{k\}\\big\)\\leq\\Delta t\_\{k\}^\{2\}\\,\\exp\\\!\\big\(2\\Delta t\_\{k\}\\left\\\|\\bar\{g\}\_\{k\}\\right\\\|\_\{\\infty\}\\big\)\\,\\operatorname\{Var\}\_\{q\_\{k\}^\{\\Delta\}\}\\\!\\big\(g\_\{t\_\{k\}\}^\{\\mathrm\{eff\}\}\\big\)\.\(D\.31\)MultiplyingWkW\_\{k\}by the deterministic constantexp\(−ΔtkqkΔ\(gtkeff\)\)\\exp\(\-\\Delta t\_\{k\}\\,q\_\{k\}^\{\\Delta\}\(g\_\{t\_\{k\}\}^\{\\mathrm\{eff\}\}\)\)leaves the normalized recursion \([4\.3](https://arxiv.org/html/2609.35947#S4.E3)\) unchanged, so \([D\.31](https://arxiv.org/html/2609.35947#A4.E31)\) controls the per\-step variance contribution to the SMC dynamics\.
###### Proof\.
LetΞ∼qkΔ\\Xi\\sim q\_\{k\}^\{\\Delta\}and setX:=gtkeff\(Ξ\)−qkΔ\(gtkeff\)X:=g\_\{t\_\{k\}\}^\{\\mathrm\{eff\}\}\(\\Xi\)\-q\_\{k\}^\{\\Delta\}\(g\_\{t\_\{k\}\}^\{\\mathrm\{eff\}\}\)\. Then𝔼\[X\]=0\\mathbb\{E\}\[X\]=0,Var\(X\)=VarqkΔ\(gtkeff\)\\operatorname\{Var\}\(X\)=\\operatorname\{Var\}\_\{q\_\{k\}^\{\\Delta\}\}\(g\_\{t\_\{k\}\}^\{\\mathrm\{eff\}\}\), and\|X\|≤‖g¯k‖∞\|X\|\\leq\\left\\\|\\bar\{g\}\_\{k\}\\right\\\|\_\{\\infty\}\. WithW¯k=eΔtkX\\bar\{W\}\_\{k\}=e^\{\\Delta t\_\{k\}X\},Var\(W¯k\)≤𝔼\[\(W¯k−1\)2\]\\operatorname\{Var\}\(\\bar\{W\}\_\{k\}\)\\leq\\mathbb\{E\}\[\(\\bar\{W\}\_\{k\}\-1\)^\{2\}\]\. The mean\-value theorem gives\|eΔtkX−1\|≤ΔtkeΔtk‖g¯k‖∞\|X\|\|e^\{\\Delta t\_\{k\}X\}\-1\|\\leq\\Delta t\_\{k\}\\,e^\{\\Delta t\_\{k\}\\left\\\|\\bar\{g\}\_\{k\}\\right\\\|\_\{\\infty\}\}\|X\|\. Squaring and taking expectations yieldsVarqkΔ\(W¯k\)≤Δtk2e2Δtk‖g¯k‖∞𝔼\[X2\]=Δtk2e2Δtk‖g¯k‖∞VarqkΔ\(gtkeff\)\\operatorname\{Var\}\_\{q\_\{k\}^\{\\Delta\}\}\(\\bar\{W\}\_\{k\}\)\\leq\\Delta t\_\{k\}^\{2\}\\,e^\{2\\Delta t\_\{k\}\\left\\\|\\bar\{g\}\_\{k\}\\right\\\|\_\{\\infty\}\}\\,\\mathbb\{E\}\[X^\{2\}\]=\\Delta t\_\{k\}^\{2\}\\,e^\{2\\Delta t\_\{k\}\\left\\\|\\bar\{g\}\_\{k\}\\right\\\|\_\{\\infty\}\}\\,\\operatorname\{Var\}\_\{q\_\{k\}^\{\\Delta\}\}\(g\_\{t\_\{k\}\}^\{\\mathrm\{eff\}\}\), which proves \([D\.31](https://arxiv.org/html/2609.35947#A4.E31)\)\. ∎
### D\.5Discussion of the theorems
## Appendix EExperimental Details
This appendix records the implementation details needed to reproduce the experiments of Section[5](https://arxiv.org/html/2609.35947#S5): the shared SMC implementation \(Appendix[E\.1](https://arxiv.org/html/2609.35947#A5.SS1)\), the finite\-state CTMC benchmark \(Appendix[E\.2](https://arxiv.org/html/2609.35947#A5.SS2)\), and the 2D Ising benchmark \(Appendix[E\.3](https://arxiv.org/html/2609.35947#A5.SS3)\), including its model and reference sampler, sampler settings andD\-VCGbasis library, metrics, and configurations\. The large\-scale experiments of Appendix[F](https://arxiv.org/html/2609.35947#A6)are described there\.
### E\.1Shared implementation
Both benchmarks use ESS\-adaptive systematic resampling, triggered whenESS/N<τ\\mathrm\{ESS\}/N<\\tau, with defaultτ=0\.5\\tau=0\.5\. The nonnegativeD\-VCGcoefficients are selected by active\-set candidate enumeration with a small diagonal ridge\. We select the candidate with the smallest variance\-plus\-anchor\-penalty score before forming𝐐teff=∑jθj⋆𝐐t\(j\)\\mathbf\{Q\}\_\{t\}^\{\\mathrm\{eff\}\}=\\sum\_\{j\}\\theta\_\{j\}^\{\\star\}\\mathbf\{Q\}\_\{t\}^\{\(j\)\}\. Because candidate generation and scoring use distinct regularization terms, this is an approximate solve of the stated objective\.
### E\.2Finite\-state CTMC
##### Setup and canonical configurations\.
We use two random tensor\-product CTMC families,*uniform\-state diffusion*\(each site is independently refreshed to a uniform random vocabulary index at rate one\) and*masked\-absorbing diffusion*\(each site is independently absorbed to a special mask token at rate one\), both initialized from a Dirichlet\(1\)\(1\)prior on theVLV^\{L\}state space\. The four canonical configurations all shareV=5V=5,L=3L=3,N=4000N=4000particles, an 80\-step power\-2 time grid, ESS threshold0\.50\.5, systematic resampling, and1010random seeds per cell; midpoint rates and potentials are used in a half\-weight/propagation/half\-weight splitting, and the reported comparisons use this fixed discretization\. The configurations differ only in the diffusion family and per\-regime tilt strength:*uniform\-state/reward*withσr=3\.0\\sigma\_\{r\}=3\.0\(Gaussian random reward vector\);*uniform\-state/annealing*withγ=3\.0\\gamma=3\.0;*masked\-absorbing/reward*withσr=1\.0\\sigma\_\{r\}=1\.0\(Gaussian random reward vector\); and*masked\-absorbing/annealing*withγ=1\.3\\gamma=1\.3, since atγ\>1\.3\\gamma\>1\.3the standard Feynman–Kac SMC recursion becomes numerically unstable on masked\-absorbing diffusion\.
\(a\)Sweep ofKL\(qT∥q^TN\)\\KL\(q\_\{T\}\\,\\\|\\,\\widehat\{q\}\_\{T\}^\{N\}\)versus reward scaleσr\\sigma\_\{r\}\(reward regimes\) or annealing exponentγ\\gamma\(annealing regimes\) across the four canonicalV=5V=5,L=3L=3configurations of Section[5\.1](https://arxiv.org/html/2609.35947#S5.SS1)\. Six samplers \(PG,D\-FKC,PR,HEU,D\-VCG,DEN\)\. All methods are evaluated on the same fixed grid within each canonical regime\.\(b\)Particle\-count scaling at the four canonical cells of Section[5\.1](https://arxiv.org/html/2609.35947#S5.SS1)withN∈\{500,…,16000\}N\\in\\\{500,\\dots,16000\\\}, log\-log axes; the dashed line is a1/N1/Nreference anchored at the smallest\-NND\-VCGvalue\. Each panel uses the per\-regime tilt of the matching Fig\.[1](https://arxiv.org/html/2609.35947#S5.F1)cell\.
Figure 3:Finite\-state CTMC ablations\. Across both the parameter sweeps and the particle\-count axis,D\-VCGmatches or improves on the dense oracleDENwithoutDEN’s tractability assumptions\.
##### D\-VCGdamping\.
The objective in Proposition[3\.6](https://arxiv.org/html/2609.35947#S3.Thmtheorem6)minimizes only the*importance\-weight*variance contribution; finite\-NNFeynman–Kac SMC also incurs a*propagation*variance from categorical sampling under𝐐tvcg\\mathbf\{Q\}\_\{t\}^\{\\mathrm\{vcg\}\}\. When the initial backward marginalpt0←p\_\{t\_\{0\}\}^\{\\shortleftarrow\}already covers the support ofqTq\_\{T\}well \(typical for small dense state spaces\), pushing particles by the unregularized least\-squares solution𝜽⋆\\boldsymbol\{\\theta\}^\{\\star\}is wasteful: the propagation noise dominates\. Multiplicatively damping𝜽used=s𝜽⋆\\boldsymbol\{\\theta\}\_\{\\rm used\}=s\\,\\boldsymbol\{\\theta\}^\{\\star\}interpolates between pure reweighting \(s=0s\{=\}0\) and the unregularized optimum \(s=1s\{=\}1\); the toy experiments uses=0\.25s=0\.25in both uniform\-state regimes,s=0\.75s=0\.75for masked reward tilting, ands=1s=1for masked annealing, playing the same role forD\-VCGas the damping factorαt\\alpha\_\{t\}does forHEU\.
##### Metric, aggregation, and IS references\.
For the finite\-state CTMC, the closed\-formqTq\_\{T\}allows the direct KL metric
KL\(qT∥q^TN\)=∑σqT\(σ\)logqT\(σ\)q^TN\(σ\),\\KL\(q\_\{T\}\\,\\\|\\,\\widehat\{q\}\_\{T\}^\{N\}\)=\\sum\_\{\\sigma\}q\_\{T\}\(\\sigma\)\\log\\frac\{q\_\{T\}\(\\sigma\)\}\{\\widehat\{q\}\_\{T\}^\{N\}\(\\sigma\)\},whereq^TN\\widehat\{q\}\_\{T\}^\{N\}is the weighted\-particle histogram\. Both probability vectors are clipped below atε=10−15\\varepsilon=10^\{\-15\}and renormalized before evaluating KL to avoidlog0\\log 0\. Toy curves show geometric means over ten seeds; the main\-figure bars are obtained by exponentiating the mean log KL plus or minus one standard deviation of log KL\. The IS \(prior\) and IS \(guided\) baselines are self\-normalized importance sampling withN=4000N=4000samples drawn frompT←p\_\{T\}^\{\\shortleftarrow\}or theD\-FKCguided proposal, reweighted by the exact density ratioqT/pT←q\_\{T\}/p\_\{T\}^\{\\shortleftarrow\}orqT/q~Tq\_\{T\}/\\widetilde\{q\}\_\{T\}\. Both require closed\-formqTq\_\{T\}and are tractable only on the toy\.
##### Ablations\.
Beyond the canonical cells reported in Fig\.[1](https://arxiv.org/html/2609.35947#S5.F1), Figure[3](https://arxiv.org/html/2609.35947#A5.F3)shows two ablations\. The parameter sweep overσr\\sigma\_\{r\}\(reward\) orγ\\gamma\(annealing\) shows regime\-dependent performance\.PRis competitive at some uniform\-state parameter settings, whileD\-VCGimproves over it in both main uniform\-state configurations\.DENprovides a dense zero\-residual reference; its finite\-particle terminal error need not be a lower bound for sparse methods\. The particle\-count sweepN∈\{500,…,16,000\}N\\in\\\{500,\\dots,16\{,\}000\\\}includes a1/N1/Nreference trend; the standard Feynman–Kac SMC baselines are much less stable on masked\-absorbing/annealing\.
### E\.32D Ising
#### E\.3\.1Model, training, and reference
##### Base model\.
The score network is a U\-Net trained with the denoising score\-entropy loss under thelineardiffusion schedule on 200,000 Swendsen–Wang samples at the training inverse\-temperatureβtrain∈\{0\.3,0\.4\}\\beta\_\{\\rm train\}\\in\\\{0\.3,0\.4\\\}\.
##### Target and reference\.
The reward is the linear magnetizationr\(σ\)=∑iσi≡M\(σ\)r\(\\sigma\)=\\sum\_\{i\}\\sigma\_\{i\}\\equiv M\(\\sigma\), so the tilted targetqβ,βr\(σ\)∝e−βH\(σ\)\+βrM\(σ\)q\_\{\\beta,\\beta\_\{r\}\}\(\\sigma\)\\propto e^\{\-\\beta H\(\\sigma\)\+\\beta\_\{r\}M\(\\sigma\)\}is the Ising model in a uniform external fieldheff=βr/βh\_\{\\rm eff\}=\\beta\_\{r\}/\\beta\. For each\(β,βr\)\(\\beta,\\beta\_\{r\}\)cell we use20002000Swendsen–Wang reference samples\. For nonzero fields, a fresh reference uses the augmented graph with a ghost spin coupled to every site at strengthheffh\_\{\\rm eff\}\(standard ghost\-spin construction\[[146](https://arxiv.org/html/2609.35947#bib.bib146)\]\)\. These references estimate⟨M⟩\\langle M\\rangle,⟨E⟩\\langle E\\rangle, andC\(r\)C\(r\)under the target distribution up to MCMC error\.
#### E\.3\.2Sampler
##### Time discretization\.
We use thelinearschedule of\[[97](https://arxiv.org/html/2609.35947#bib.bib97)\]with noise horizonσmax=10\\sigma\_\{\\max\}=10, discretized intoSSequal\-dt\\mathop\{\}\\\!\\mathrm\{d\}tsteps \(the grid sizeMMof Algorithm[1](https://arxiv.org/html/2609.35947#algorithm1); we writeSSto avoid a clash with the magnetizationM\(σ\)M\(\\sigma\)\)\. We useS=5000S=5000for annealing andS=2000S=2000for reward\-tilt and joint experiments\. For the binary 2D Ising spins, each frozen per\-site2×22\\times 2generator is exponentiated exactly per step in closed form, which removes Euler error for the frozen site update\.
##### Reward ramp\.
For reward\-tilt and joint experiments, we instantiate the realizationrt\(x\)=βtr\(x\)r\_\{t\}\(x\)=\\beta\_\{t\}\\,r\(x\)from Section[2\.2](https://arxiv.org/html/2609.35947#S2.SS2)with the linear ramp
βt=βrt/T=βr\(1−σ/σmax\),σ:=T−t,T=σmax\\beta\_\{t\}=\\beta\_\{r\}t/T=\\beta\_\{r\}\(1\-\\sigma/\\sigma\_\{\\max\}\),\\qquad\\sigma:=T\-t,\\quad T=\\sigma\_\{\\max\}wherettis increasing reverse time andσ\\sigmais the decreasing noise level\. The strength hyperparameterβr\\beta\_\{r\}rescales the data\-end reward \(equivalently,rris replaced byβrr\\beta\_\{r\}rbefore applying the unit\-strength ramp of the main text\)\. Thusβ0=0\\beta\_\{0\}=0at the noisy end \(matching the Feynman–Kac\-exact uniform\-prior initialisation\) andβT=βr\\beta\_\{T\}=\\beta\_\{r\}at the data end, giving an effective data\-end potentialrT=βrrr\_\{T\}=\\beta\_\{r\}r\. Pure annealing usesβt≡0\\beta\_\{t\}\\equiv 0\.
##### D\-VCGbasis library\.
We instantiate Proposition[3\.6](https://arxiv.org/html/2609.35947#S3.Thmtheorem6)with a small collection of nonnegative basis rate families of the form \([3\.11](https://arxiv.org/html/2609.35947#S3.E11)\),𝐐t\(j\)\(y,x\)=𝐐t←\(y,x\)φt\(j\)\(y,x\)\\mathbf\{Q\}\_\{t\}^\{\(j\)\}\(y,x\)=\\mathbf\{Q\}\_\{t\}^\{\\shortleftarrow\}\(y,x\)\\,\\varphi\_\{t\}^\{\(j\)\}\(y,x\), specified through their multipliersφt\(j\)\\varphi\_\{t\}^\{\(j\)\}\. The time\-tttilt parameters are the annealing exponentγ=βtarget/βtrain\\gamma=\\beta\_\{\\rm target\}/\\beta\_\{\\rm train\}and the reward strengthβt≥0\\beta\_\{t\}\\geq 0\. Let𝐐t←\\mathbf\{Q\}\_\{t\}^\{\\shortleftarrow\}denote the pretrained backward\-rate family and let
ℓt\(x,y\):=logpt←\(y\)pt←\(x\)\.\\ell\_\{t\}\(x,y\):=\\log\\frac\{p\_\{t\}^\{\\shortleftarrow\}\(y\)\}\{p\_\{t\}^\{\\shortleftarrow\}\(x\)\}\.For the Ising rewardr\(σ\)=∑iσir\(\\sigma\)=\\sum\_\{i\}\\sigma\_\{i\}, writeΔr\(x,y\)=r\(y\)−r\(x\)\\Delta r\(x,y\)=r\(y\)\-r\(x\)andΔH\(x,y\)=H\(y\)−H\(x\)\\Delta H\(x,y\)=H\(y\)\-H\(x\), withH\(σ\)=−JIsing∑⟨ij⟩σiσjH\(\\sigma\)=\-J\_\{\\rm Ising\}\\sum\_\{\\langle ij\\rangle\}\\sigma\_\{i\}\\sigma\_\{j\}andJIsing=1J\_\{\\rm Ising\}=1on the periodic lattice\. The basis library is
- •Untilted backward\(canonical basis from the main text\): φt\(1\)\(y,x\):=1\.\\varphi\_\{t\}^\{\(1\)\}\(y,x\):=1\.
- •Target\-aligned anchor\(the Ising implementation scales the canonical main\-text basis byγ\\gamma\): φt\(2\)\(y,x\):=γe\(γ−1\)ℓt\(x,y\)\+βtΔr\(x,y\)\.\\varphi\_\{t\}^\{\(2\)\}\(y,x\):=\\gamma e^\{\(\\gamma\-1\)\\ell\_\{t\}\(x,y\)\+\\beta\_\{t\}\\Delta r\(x,y\)\}\.The pure\-annealing specializationφt\(anneal\)\(y,x\):=γe\(γ−1\)ℓt\(x,y\)\\varphi\_\{t\}^\{\(\\rm anneal\)\}\(y,x\):=\\gamma e^\{\(\\gamma\-1\)\\ell\_\{t\}\(x,y\)\}\(setβt=0\\beta\_\{t\}=0\) and the pure reward\-tilt specializationφt\(tilt\)\(y,x\):=eβtΔr\(x,y\)\\varphi\_\{t\}^\{\(\\rm tilt\)\}\(y,x\):=e^\{\\beta\_\{t\}\\Delta r\(x,y\)\}\(setγ=1\\gamma=1\) are used in the corresponding ablation regimes\.
- •Intermediate annealing basis\(milder annealing tilt with exponentγmid:=\(1\+γ\)/2\\gamma\_\{\\rm mid\}:=\(1\+\\gamma\)/2\): φt\(γmid\)\(y,x\):=γmide\(γmid−1\)ℓt\(x,y\)\.\\varphi\_\{t\}^\{\(\\gamma\_\{\\rm mid\}\)\}\(y,x\):=\\gamma\_\{\\rm mid\}e^\{\(\\gamma\_\{\\rm mid\}\-1\)\\ell\_\{t\}\(x,y\)\}\.
- •Energy\-flux basis\(Ising\-specific basis that, departing from \([3\.11](https://arxiv.org/html/2609.35947#S3.E11)\), replaces the learned backward\-rate backbone with the bare forward\-rate prefactorQT−t→\(x,y\)Q\_\{T\-t\}^\{\\shortrightarrow\}\(x,y\), so as to isolate the energy\-gradient direction from the learned score model\): 𝐐t\(E\)\(y,x\):=QT−t→\(x,y\)\[−βtargetΔH\(x,y\)\]\+\.\\mathbf\{Q\}\_\{t\}^\{\(E\)\}\(y,x\):=Q\_\{T\-t\}^\{\\shortrightarrow\}\(x,y\)\\,\[\-\\beta\_\{\\rm target\}\\Delta H\(x,y\)\]\_\{\+\}\.
For pure annealing, the minimal library uses\{φ\(1\),φ\(anneal\)\}\\\{\\varphi^\{\(1\)\},\\varphi^\{\(\\rm anneal\)\}\\\}; the augmented library adds𝐐\(E\)\\mathbf\{Q\}^\{\(E\)\}and, whenγ\>1\\gamma\>1, the midpoint basis\. For joint annealing and reward tilting, the minimal library contains the backward, annealing, and joint bases; the augmented library adds the energy basis, with no midpoint extra\. Atγ=1\\gamma=1, the redundant annealing basis is removed, giving two and three bases in pure tilt\. Thus “2\-basis” and “4\-basis” are figure labels for controller variants, not uniform counts across regimes\.
##### Anchor regularizer\.
For the joint anneal\+\+tilt regime \(Config 2a in Appendix[E\.3\.4](https://arxiv.org/html/2609.35947#A5.SS3.SSS4); the pure\-tilt and pure\-annealing configurations are run withλeff≡0\\lambda\_\{\\rm eff\}\\equiv 0\), the variance objective of Proposition[3\.6](https://arxiv.org/html/2609.35947#S3.Thmtheorem6)is augmented with a quadratic anchor\-penalty term:
min𝜽≥0Varqt\(gt0\+∑iθi\(divqt𝐐t\(i\)\)\)\+λeff\(t\)‖𝜽−𝜽anchor‖22,\\min\_\{\\boldsymbol\{\\theta\}\\geq 0\}\\;\\operatorname\{Var\}\_\{q\_\{t\}\}\\\!\\bigg\(g\_\{t\}^\{0\}\+\\sum\_\{i\}\\theta\_\{i\}\\,\(\\operatorname\{div\}\_\{q\_\{t\}\}\\mathbf\{Q\}\_\{t\}^\{\(i\)\}\)\\bigg\)\+\\lambda\_\{\\rm eff\}\(t\)\\,\\\|\\boldsymbol\{\\theta\}\-\\boldsymbol\{\\theta\}\_\{\\rm anchor\}\\\|\_\{2\}^\{2\},\(E\.1\)where𝜽anchor\\boldsymbol\{\\theta\}\_\{\\rm anchor\}is the one\-hot vector on the target\-aligned proposal𝐐\(2\)\\mathbf\{Q\}^\{\(2\)\}, and the schedule is
λeff\(t\)=coef⋅βt2,coef=104,\\lambda\_\{\\rm eff\}\(t\)=\\mathrm\{coef\}\\cdot\\beta\_\{t\}^\{2\},\\qquad\\mathrm\{coef\}=10^\{4\},\(E\.2\)combined with the rampβt=βrt/T\\beta\_\{t\}=\\beta\_\{r\}t/Tin increasing reverse time, withT=σmaxT=\\sigma\_\{\\max\}\(see the reward ramp above\)\. At the noisy end \(t=0t=0,βt=0\\beta\_\{t\}=0\) the anchor penalty is inactive; at the data end \(t=Tt=T,βt=βr\\beta\_\{t\}=\\beta\_\{r\}\) it pulls𝜽\\boldsymbol\{\\theta\}toward the target\-aligned committed proposal\.
##### HEU on Ising\.
HEUis omitted from the Ising experiments because the empirical particle measureq^tN\(y\)=0\\widehat\{q\}\_\{t\}^\{N\}\(y\)=0on the overwhelming majority of single\-flip neighbors of theNNactive particles in𝒳=\{±1\}16×16\\mathcal\{X\}=\\\{\\pm 1\\\}^\{16\\times 16\}, so the empirical\-cloud evaluation ofμt\(y\)\\mu\_\{t\}\(y\)used in the toy benchmark vanishes spuriously on most candidate destinations\. The analytical\-expansion formμt\(y\)=qt\(y\)gt\(y\)\+∑z≠y\(Qt\(y,z\)qt\(z\)−Qt\(z,y\)qt\(y\)\)\\mu\_\{t\}\(y\)=q\_\{t\}\(y\)g\_\{t\}\(y\)\+\\sum\_\{z\\neq y\}\\big\(Q\_\{t\}\(y,z\)q\_\{t\}\(z\)\-Q\_\{t\}\(z,y\)q\_\{t\}\(y\)\\big\)recommended in Appendix[C\.9](https://arxiv.org/html/2609.35947#A3.SS9)resolves this issue without extra score\-network evaluations, and adopting it is a natural next step forHEUon DLM\-scale state spaces\.
#### E\.3\.3Metrics and aggregation
For Ising experiments, we evaluate four scalar metrics per cell, comparingNNgenerated particles to20002000SW reference samples:
- •𝖶2\(\|m\|\)\\mathsf\{W\}\_\{2\}\(\|m\|\)\(annealing\) or𝖶2\(M\)\\mathsf\{W\}\_\{2\}\(M\)\(reward\-tilt\): the empirical 1\-D Wasserstein\-2 distance between the per\-particle absolute per\-site magnetization\|m\(σ\)\|=\|∑iσi\|/L2\|m\(\\sigma\)\|=\|\\sum\_\{i\}\\sigma\_\{i\}\|/L^\{2\}or total magnetizationM\(σ\)=∑iσiM\(\\sigma\)=\\sum\_\{i\}\\sigma\_\{i\}and the reference\.
- •𝖶2\(Etotal\)\\mathsf\{W\}\_\{2\}\(E\_\{\\rm total\}\): the same 1\-D Wasserstein\-2 distance applied to the total energyH\(σ\)H\(\\sigma\)\. This metric is sensitive to the joint distribution sinceHHis non\-linear inσ\\sigma\.
- •MSE\(C\(r\)\)\\mathrm\{MSE\}\(C\(r\)\): mean\-squared error of the uncentered row correlation\. ForL=16L=16and zero\-based lattice indices, C\(r;σ\)=1\(L−2\)\(L−2−r\)∑i=1L−2∑j=1L−2−rσi,jσi,j\+r,r=1,…,L−3=13\.C\(r;\\sigma\)=\\frac\{1\}\{\(L\-2\)\(L\-2\-r\)\}\\sum\_\{i=1\}^\{L\-2\}\\sum\_\{j=1\}^\{L\-2\-r\}\\sigma\_\{i,j\}\\sigma\_\{i,j\+r\},\\qquad r=1,\\ldots,L\-3=13\.This estimator excludes a one\-site boundary margin and does not subtract magnetization\. We averageC\(r,σ\)C\(r;\\sigma\)over generated and reference samples, then average the squared difference between these estimates overrr\.
- •⟨\|m\|⟩\\langle\|m\|\\rangleor⟨M⟩\\langle M\\rangle: per\-method mean of the moment observable, plotted against the SW reference for visual sanity\.
Ising curves show means over three seeds, with bands equal to the across\-seed standard deviation divided by3\\sqrt\{3\}\. Improvement factors divide the seed\-mean errors within each setting before taking geometric means or peaks across settings\.
#### E\.3\.4Configurations and additional results
The five Ising configurations are as follows\.
- •Config 1a \(pure annealing, Fig\.[2\(a\)](https://arxiv.org/html/2609.35947#S5.F2.sf1)\):trains atβtrain=0\.4\\beta\_\{\\rm train\}=0\.4, sweepsβtarget∈\[0\.20,0\.55\]\\beta\_\{\\rm target\}\\in\[0\.20,0\.55\]over99values withβr=0\\beta\_\{r\}=0,γ∈\[0\.5,1\.375\]\\gamma\\in\[0\.5,1\.375\],N=500N=500,S=5000S=5000, andλeff=0\\lambda\_\{\\rm eff\}=0\.
- •Config 1b \(annealing on a paramagnetic base, Fig\.[4\(a\)](https://arxiv.org/html/2609.35947#A5.F4.sf1)\):trains atβtrain=0\.3\\beta\_\{\\rm train\}=0\.3, sweepsβtarget∈\[0\.20,0\.60\]\\beta\_\{\\rm target\}\\in\[0\.20,0\.60\]over1010values withγ∈\[0\.667,2\.0\]\\gamma\\in\[0\.667,2\.0\],N=500N=500,S=5000S=5000,λeff=0\\lambda\_\{\\rm eff\}=0, and a tighter ESS thresholdτ=0\.25\\tau=0\.25chosen by ablation as best for this paramagnetic regime\.
- •Config 2a \(joint anneal\+\+tilt, Fig\.[2\(b\)](https://arxiv.org/html/2609.35947#S5.F2.sf2)\):usesβtrain=0\.4\\beta\_\{\\rm train\}=0\.4,βtarget=0\.45\\beta\_\{\\rm target\}=0\.45,γ=1\.125\\gamma=1\.125\(so the reward tilt sits on top of a mildγ\>1\\gamma\>1annealing\), sweepsβr∈\[0\.02,0\.40\]\\beta\_\{r\}\\in\[0\.02,0\.40\]over77values withN=500N=500,S=2000S=2000, and the adaptive regularizerλeff\(t\)=104βt\(t\)2\\lambda\_\{\\rm eff\}\(t\)=10^\{4\}\\beta\_\{t\}\(t\)^\{2\}from \([E\.2](https://arxiv.org/html/2609.35947#A5.E2)\)\.
- •Config 2b \(pure tilt, Fig\.[4\(b\)](https://arxiv.org/html/2609.35947#A5.F4.sf2)\):usesβtrain=βtarget=0\.4\\beta\_\{\\rm train\}=\\beta\_\{\\rm target\}=0\.4,γ=1\\gamma=1\(soφt\(2\)\\varphi\_\{t\}^\{\(2\)\}collapses to the pure\-tilt multipliereβtΔr\(x,y\)e^\{\\beta\_\{t\}\\Delta r\(x,y\)\}\), the same reward sweepβr∈\[0\.02,0\.40\]\\beta\_\{r\}\\in\[0\.02,0\.40\],N=500N=500,S=2000S=2000, andλeff=0\\lambda\_\{\\rm eff\}=0: pure tilt is run as the unregularizedD\-VCGbaseline so that Config 2b isolates the variance objective alone, with the anchor penalty reserved for the joint anneal\+tilt regime of Config 2a\.
- •Config 3 \(particle\-count sweep, Fig\.[5](https://arxiv.org/html/2609.35947#A5.F5)\):usesβtrain=βtarget=0\.4\\beta\_\{\\rm train\}=\\beta\_\{\\rm target\}=0\.4,γ=1\\gamma=1,βr=0\.20\\beta\_\{r\}=0\.20,λeff=0\\lambda\_\{\\rm eff\}=0,S=2000S=2000, and sweepsN∈\{100,200,500,1000,2000,5000\}N\\in\\\{100,200,500,1000,2000,5000\\\}\.
All five configurations use the matrix\-exponential integrator, ESS\-adaptive systematic resampling at the threshold listed above,33random seeds, and the four\-method comparisonPG,D\-FKC,D\-VCG\(2\-basis\), andD\-VCG\(4\-basis/augmented\)\.
\(a\)Annealing on a paramagnetic base \(βtrain=0\.3\\beta\_\{\\rm train\}=0\.3,N=500N=500,τ=0\.25\\tau=0\.25,1010target temperatures in\[0\.20,0\.60\]\[0\.20,0\.60\]\)\.D\-VCG\(4\-basis\) reducesMSE\(C\(r\)\)\\mathrm\{MSE\}\(C\(r\)\)overD\-FKCby1\.69×1\.69\\timesin geometric mean, with peak4\.84×\\mathbf\{4\.84\\times\}atβtarget=0\.40\\beta\_\{\\rm target\}=0\.40\.\(b\)Pure tilt atβtrain=βtarget=0\.4\\beta\_\{\\rm train\}=\\beta\_\{\\rm target\}=0\.4,γ=1\\gamma=1,λeff=0\\lambda\_\{\\rm eff\}=0,N=500N=500\. With both the mild joint\-anneal component and the anchor regularizer switched off,D\-VCGstill reducesMSE\(C\(r\)\)\\mathrm\{MSE\}\(C\(r\)\)overD\-FKCby5\.21×5\.21\\timesin geometric mean onβr∈\[0\.02,0\.40\]\\beta\_\{r\}\\in\[0\.02,0\.40\], with peak17\.2×\\mathbf\{17\.2\\times\}atβr=0\.15\\beta\_\{r\}=0\.15\.
Figure 4:Further 2D Ising configurations\. The same axes and methods as Fig\.[2](https://arxiv.org/html/2609.35947#S5.F2)confirm that theD\-VCGadvantage \(i\) holds for a second base\-model training temperature \(a\) and \(ii\) survives in the pure\-tilt limit \(γ=1\\gamma=1,λeff=0\\lambda\_\{\\rm eff\}=0, panelb\), where the joint\-mode anchor reduces to the free\-controllerD\-VCGof Proposition[3\.6](https://arxiv.org/html/2609.35947#S3.Thmtheorem6)\.Figure 5:Particle\-count sweep atβr=0\.20\\beta\_\{r\}=0\.20,βtrain=βtarget=0\.4\\beta\_\{\\rm train\}=\\beta\_\{\\rm target\}=0\.4,γ=1\\gamma=1,λeff=0\\lambda\_\{\\rm eff\}=0\(Config 3\)\. AcrossN∈\{100,…,5000\}N\\in\\\{100,\\dots,5000\\\},D\-VCG\(4\-basis\) reducesMSE\(C\(r\)\)\\mathrm\{MSE\}\(C\(r\)\)overD\-FKCby15\.21×15\.21\\timesin geometric mean, with peak34\.7×\\mathbf\{34\.7\\times\}atN=1000N=1000, showing that the advantage does not close asNNgrows\.##### Appendix experiments\.
Fig\.[4\(a\)](https://arxiv.org/html/2609.35947#A5.F4.sf1)re\-runs the annealing sweep on a paramagnetic base \(βtrain=0\.3\\beta\_\{\\rm train\}=0\.3, well below criticality\), checking that the gain overD\-FKCdoes not depend on the base model being trained near the critical temperature; the gain is smaller because all samplers agree more closely with the SW reference, and it peaks atβtarget=0\.40\\beta\_\{\\rm target\}=0\.40\. Fig\.[4\(b\)](https://arxiv.org/html/2609.35947#A5.F4.sf2)setsγ=1\\gamma=1andλeff=0\\lambda\_\{\\rm eff\}=0, removing both the joint\-mode anchor and the regularizer of \([E\.1](https://arxiv.org/html/2609.35947#A5.E1)\), so the remaining gain isolates the basis\-mixing mechanism of the QP\. Fig\.[5](https://arxiv.org/html/2609.35947#A5.F5)sweepsNNatβr=0\.20\\beta\_\{r\}=0\.20in the Config 2b setting: the gap does not close over the tested range, andD\-FKCatN=5000N=5000\(10×10\\timesthe particle budget\) still loses toD\-VCGatN=500N=500by5\.4×5\.4\\timesonMSE\(C\(r\)\)\\mathrm\{MSE\}\(C\(r\)\)\.
##### Compute\.
On a single A100, the Ising per\-step wall\-clock forD\-VCG\(4\-basis\) is within roughly5%5\\%ofD\-FKCatN=500N=500: the matrix\-exponential integrator and score\-network forward together account for the dominant cost, while the≤5×5\\leq 5\\times 5active\-set QP solve is negligible\.
## Appendix FAdditional Large\-Scale Experiments
This appendix probesD\-VCG\(Proposition[3\.6](https://arxiv.org/html/2609.35947#S3.Thmtheorem6)\) on two large\-scale settings that go beyond the analytic CTMC and Ising benchmarks of Section[5](https://arxiv.org/html/2609.35947#S5)\. Appendix[F\.1](https://arxiv.org/html/2609.35947#A6.SS1)reports a discrete inverse problem on monophonic music sequences \(a reward\-tilted special case of Section[2\.2](https://arxiv.org/html/2609.35947#S2.SS2)withγ=1\\gamma=1\), and Appendix[F\.2](https://arxiv.org/html/2609.35947#A6.SS2)reports text\-to\-image generation under classifier\-free guidance\.
### F\.1Discrete inverse problem on monophonic music \(reward\-tilted sampling\)
We follow the music inverse\-problem setup of\[[31](https://arxiv.org/html/2609.35947#bib.bib31)\]\. As in\[[20](https://arxiv.org/html/2609.35947#bib.bib20)\], we preprocess the Lakh pianoroll dataset\[[44](https://arxiv.org/html/2609.35947#bib.bib44)\]into monophonic note sequences of lengthL=256L=256over a per\-site vocabulary of sizeV=129V=129, and use the SEDD\[[97](https://arxiv.org/html/2609.35947#bib.bib97)\]checkpoint of\[[31](https://arxiv.org/html/2609.35947#bib.bib31)\]as the pretrained reverse model on theVLV^\{L\}state space\. The observation model is
whereGGis a random masking operator that reveals a fractionρ∈\{40%,60%\}\\rho\\in\\\{40\\%,60\\%\\\}of the positions ofxxandnnis the corruption noise\. The induced log\-likelihood plays the role of the reward of Section[2\.2](https://arxiv.org/html/2609.35947#S2.SS2),
r\(x\):=logp\(y∣x\)=−1σy‖G\(x\)−y‖0\+const,r\(x\)\\;:=\\;\\log p\(y\\mid x\)\\;=\\;\-\\frac\{1\}\{\\sigma\_\{y\}\}\\,\\\|G\(x\)\-y\\\|\_\{0\}\\;\+\\;\\mathrm\{const\},with noise scaleσy=0\.1\\sigma\_\{y\}=0\.1, so the targetq\(x\)∝pT←\(x\)er\(x\)q\(x\)\\propto p\_\{T\}^\{\\shortleftarrow\}\(x\)\\,e^\{r\(x\)\}is the reward\-tilted special case of Proposition[2\.1](https://arxiv.org/html/2609.35947#S2.Thmtheorem1)atγ=1\\gamma=1\. Following\[[31](https://arxiv.org/html/2609.35947#bib.bib31)\], we evaluate generated samples by two scalar metrics: \(i\) the Hellinger distance between the histograms of generated and ground\-truth notes, measuring fidelity to the prior, and \(ii\) the*measurement error*, defined as the fraction of revealed positions on whichG\(x\)G\(x\)disagrees withyy\.
We compareD\-VCGagainst two baselines\.SGDD\[[31](https://arxiv.org/html/2609.35947#bib.bib31)\]is a state\-of\-the\-art split\-Gibbs sampler;D\-FKC\[[59](https://arxiv.org/html/2609.35947#bib.bib59)\]is the standard Feynman–Kac SMC baseline from Section[5](https://arxiv.org/html/2609.35947#S5)\. To ensure a fair comparison, we fix the total number of function evaluations \(NFEs\) at10241024for all three methods:SGDDuses3232outer iterations and3232inner denoising steps;D\-FKCandD\-VCGboth useM=128M=128denoising steps andN=8N=8particles\.D\-VCGis instantiated with the two bases of Section[3\.3\.2](https://arxiv.org/html/2609.35947#S3.SS3.SSS2)\(untilted backward and target\-aligned anchor\)\.
Table 2:Quantitative evaluation on monophonic music infilling\. Lower is better\.ρ=40%\\rho=40\\%ρ=60%\\rho=60\\%Hellinger↓\\downarrowMeas\. Error↓\\downarrowHellinger↓\\downarrowMeas\. Error↓\\downarrowSGDD0\.0764\.930\.1516\.11D\-FKC0\.0330\.290\.0810\.88D\-VCG0\.0350\.000\.0790\.05Table[2](https://arxiv.org/html/2609.35947#A6.T2)reports the quantitative results\. At both masking ratiosρ∈\{40%,60%\}\\rho\\in\\\{40\\%,60\\%\\\}, the two Feynman–Kac samplersD\-VCGandD\-FKCsubstantially outperformSGDDin Hellinger distance\. Under the measurement\-error metric,D\-VCGdominates both baselines\.
### F\.2Text\-to\-image generation \(classifier\-free guidance\)
Following the experimental setup of\[[109](https://arxiv.org/html/2609.35947#bib.bib109)\], we evaluateD\-VCGon text\-to\-image generation with Meissonic\[[9](https://arxiv.org/html/2609.35947#bib.bib9)\]as the pretrained discrete\-diffusion base model\. We report two text–image alignment scores: the Human Preference Score v2 \(HPSv2\)\[[163](https://arxiv.org/html/2609.35947#bib.bib163)\]and the Multi\-Dimensional Human Preference Score \(MPS\)\[[180](https://arxiv.org/html/2609.35947#bib.bib180)\]\. The prompt set consists of2525prompts drawn from each of four categories \(anime illustration, concept art, acrylic painting, and watercolor painting\), for a total of100100prompts\. Classifier\-free guidance \(CFG\) targets a geometric combination of the unconditional and conditional reverse marginals,
qtcfg\(x\)∝pt←\(x\)1−spt←\(x∣c\)s,q^\{\\rm cfg\}\_\{t\}\(x\)\\;\\propto\\;p\_\{t\}^\{\\shortleftarrow\}\(x\)^\{1\-s\}\\,p\_\{t\}^\{\\shortleftarrow\}\(x\\mid c\)^\{s\},\(F\.1\)whereccdenotes the input prompt andssis the guidance scale; we fixs=9s=9following\[[9](https://arxiv.org/html/2609.35947#bib.bib9)\]\. Equation \([F\.1](https://arxiv.org/html/2609.35947#A6.E1)\) is a particular instance of the discrete tilted path of Section[2\.2](https://arxiv.org/html/2609.35947#S2.SS2): the identifications
γ=1−s,rt\(x\)=slogpt←\(x∣c\)\\gamma\\;=\\;1\-s,\\qquad r\_\{t\}\(x\)\\;=\\;s\\,\\log p\_\{t\}^\{\\shortleftarrow\}\(x\\mid c\)recoverqtcfg∝\(pt←\)γertq^\{\\rm cfg\}\_\{t\}\\propto\(p\_\{t\}^\{\\shortleftarrow\}\)^\{\\gamma\}e^\{r\_\{t\}\}\. Although hereγ<0\\gamma<0– so the canonical representative of Proposition[2\.1](https://arxiv.org/html/2609.35947#S2.Thmtheorem1)is not a valid rate family – the basis\-controlledD\-VCGconstruction below still produces a nonnegative effective rate𝐐teff\\mathbf\{Q\}\_\{t\}^\{\\mathrm\{eff\}\}\. Since\[[109](https://arxiv.org/html/2609.35947#bib.bib109)\]shows thatD\-FKCalready dominates earlier inference\-time baselines on CFG, we restrict the comparison toD\-FKCvs\.D\-VCG\.
##### CFG\-specific bases forD\-VCG\.
We instantiateD\-VCGwith two per\-token rate bases naturally suggested by the CFG structure \([F\.1](https://arxiv.org/html/2609.35947#A6.E1)\)\. At each masked position, the transformer yields two categorical distributions over the per\-token vocabulary𝒱\\mathcal\{V\}: the unconditionalpt←\(⋅∣∅\)p\_\{t\}^\{\\shortleftarrow\}\(\\cdot\\mid\\varnothing\)\(empty prompt\) and the conditionalpt←\(⋅∣c\)p\_\{t\}^\{\\shortleftarrow\}\(\\cdot\\mid c\)\. We take𝐐t←\\mathbf\{Q\}\_\{t\}^\{\\shortleftarrow\}to be the unconditional reverse rate family of Section[2\.2](https://arxiv.org/html/2609.35947#S2.SS2)and define the per\-token log local ratios
ℓt\(x,y∣∅\):=logpt←\(y∣∅\)pt←\(x∣∅\),ℓt\(x,y∣c\):=logpt←\(y∣c\)pt←\(x∣c\),\\ell\_\{t\}\(x,y\\mid\\varnothing\):=\\log\\frac\{p\_\{t\}^\{\\shortleftarrow\}\(y\\mid\\varnothing\)\}\{p\_\{t\}^\{\\shortleftarrow\}\(x\\mid\\varnothing\)\},\\qquad\\ell\_\{t\}\(x,y\\mid c\):=\\log\\frac\{p\_\{t\}^\{\\shortleftarrow\}\(y\\mid c\)\}\{p\_\{t\}^\{\\shortleftarrow\}\(x\\mid c\)\},\(F\.2\)in the same convention asℓt\\ell\_\{t\}in Appendix[E\.3\.2](https://arxiv.org/html/2609.35947#A5.SS3.SSS2.Px3)\. Following the basis form𝐐t\(j\)\(y,x\)=𝐐t←\(y,x\)φt\(j\)\(y,x\)\\mathbf\{Q\}\_\{t\}^\{\(j\)\}\(y,x\)=\\mathbf\{Q\}\_\{t\}^\{\\shortleftarrow\}\(y,x\)\\,\\varphi\_\{t\}^\{\(j\)\}\(y,x\)of \([3\.11](https://arxiv.org/html/2609.35947#S3.E11)\), the two CFG multipliers are
φt\(1\)\(y,x\)\\displaystyle\\varphi\_\{t\}^\{\(1\)\}\(y,x\):=1,\\displaystyle:=1,\(F\.3\)φt\(2\)\(y,x\)\\displaystyle\\varphi\_\{t\}^\{\(2\)\}\(y,x\):=exp\(s\[ℓt\(x,y∣c\)−ℓt\(x,y∣∅\)\]\)=\(pt←\(y∣c\)pt←\(x∣∅\)pt←\(x∣c\)pt←\(y∣∅\)\)s\.\\displaystyle:=\\exp\\\!\\big\(s\\,\[\\ell\_\{t\}\(x,y\\mid c\)\-\\ell\_\{t\}\(x,y\\mid\\varnothing\)\]\\big\)=\\left\(\\frac\{p\_\{t\}^\{\\shortleftarrow\}\(y\\mid c\)\\,p\_\{t\}^\{\\shortleftarrow\}\(x\\mid\\varnothing\)\}\{p\_\{t\}^\{\\shortleftarrow\}\(x\\mid c\)\\,p\_\{t\}^\{\\shortleftarrow\}\(y\\mid\\varnothing\)\}\\right\)^\{s\}\.\(F\.4\)Hereφt\(1\)\\varphi\_\{t\}^\{\(1\)\}is the*untilted backward*basis of Section[3\.3\.2](https://arxiv.org/html/2609.35947#S3.SS3.SSS2)\(which simply returns the unconditional reverse rate𝐐t←\\mathbf\{Q\}\_\{t\}^\{\\shortleftarrow\}\), whileφt\(2\)\\varphi\_\{t\}^\{\(2\)\}is the*target\-aligned anchor*of Section[3\.3\.2](https://arxiv.org/html/2609.35947#S3.SS3.SSS2)for the parameter identification\(γ,rt\)=\(1−s,slogpt←\(⋅∣c\)\)\(\\gamma,r\_\{t\}\)=\(1\-s,\\,s\\log p\_\{t\}^\{\\shortleftarrow\}\(\\cdot\\mid c\)\): the corresponding rate
Qt\(2\)\(y,x\)=QT−t→\(x,y\)\(pt←\(y∣∅\)pt←\(x∣∅\)\)1−s\(pt←\(y∣c\)pt←\(x∣c\)\)sQ\_\{t\}^\{\(2\)\}\(y,x\)\\;=\\;Q\_\{T\-t\}^\{\\shortrightarrow\}\(x,y\)\\,\\left\(\\frac\{p\_\{t\}^\{\\shortleftarrow\}\(y\\mid\\varnothing\)\}\{p\_\{t\}^\{\\shortleftarrow\}\(x\\mid\\varnothing\)\}\\right\)^\{\\\!1\-s\}\\\!\\left\(\\frac\{p\_\{t\}^\{\\shortleftarrow\}\(y\\mid c\)\}\{p\_\{t\}^\{\\shortleftarrow\}\(x\\mid c\)\}\\right\)^\{\\\!s\}is the exact CFG\-anchored Feynman–Kac proposal, recovered from the local\-ratio decompositionQt←\(y,x\)=QT−t→\(x,y\)pt←\(y∣∅\)/pt←\(x∣∅\)Q\_\{t\}^\{\\shortleftarrow\}\(y,x\)=Q\_\{T\-t\}^\{\\shortrightarrow\}\(x,y\)\\,p\_\{t\}^\{\\shortleftarrow\}\(y\\mid\\varnothing\)/p\_\{t\}^\{\\shortleftarrow\}\(x\\mid\\varnothing\)of Section[2\.2](https://arxiv.org/html/2609.35947#S2.SS2)\. Both methods useM=64M=64denoising steps andN=8N=8particles, and we report the mean and standard error of each metric across the100100prompts\.
Table 3:Quantitative evaluation on CFG text\-to\-image generation: mean±\\pmstandard error over100100prompts\. Higher is better\.MPS↑\\uparrowHPSv2↑\\uparrowD\-FKC15\.282±0\.03215\.282\\pm 0\.0320\.2980±0\.00030\.2980\\pm 0\.0003D\-VCG15\.422±0\.032\\mathbf\{15\.422\}\\pm 0\.0320\.3033±0\.0003\\mathbf\{0\.3033\}\\pm 0\.0003
##### Quantitative results\.
Table[3](https://arxiv.org/html/2609.35947#A6.T3)shows thatD\-VCGoutperformsD\-FKCon both metrics, with higher meanMPS\(15\.42215\.422vs\.15\.28215\.282\) andHPSv2\(0\.30330\.3033vs\.0\.29800\.2980\)\.
##### Qualitative results\.
To complement Table[3](https://arxiv.org/html/2609.35947#A6.T3), Table[4](https://arxiv.org/html/2609.35947#A6.T4)shows six representative prompts spanning anime illustration, concept art, acrylic painting, and watercolor painting\.D\-VCGgenerally adheres more closely to the prompt: for the dragon\-rider prompt it renders both the rider and the trailing scarf, whileD\-FKComits the rider entirely, and across the remaining prompts it more accurately reflects fine\-grained details such as “wooden bridges,” “cliffs,” and the co\-occurrence of “orchids, ferns, and delicate green leaves\.” These observations are consistent with the gains onMPSandHPSv2in Table[3](https://arxiv.org/html/2609.35947#A6.T3)\.
Table 4:Qualitative comparison ofD\-FKCandD\-VCGon six representative CFG prompts\.PromptAn anime dragon riderabove snowy mountains,scarf trailing behindin the wind\.A dramatic anime close\-upof a masked hero underfalling snow, city skylinebehind\.Concept art of a giantmechanical whale swimmingthrough clouds abovea coastal town\.D\-FKC![[Uncaptioned image]](https://arxiv.org/html/2609.35947v1/images/cfg_case/dragon_rider_dfkc.jpg)![[Uncaptioned image]](https://arxiv.org/html/2609.35947v1/images/cfg_case/masked_hero_dfkc.jpg)![[Uncaptioned image]](https://arxiv.org/html/2609.35947v1/images/cfg_case/mechanical_whale_dfkc.jpg)D\-VCG![[Uncaptioned image]](https://arxiv.org/html/2609.35947v1/images/cfg_case/dragon_rider_dvcg.jpg)![[Uncaptioned image]](https://arxiv.org/html/2609.35947v1/images/cfg_case/masked_hero_dvcg.jpg)![[Uncaptioned image]](https://arxiv.org/html/2609.35947v1/images/cfg_case/mechanical_whale_dvcg.jpg)PromptConcept art of a hiddenpirate cove with waterfalls,wooden bridges, andlantern\-lit paths\.A soft acrylic paintingof a lighthouse on cliffsduring a calmblue evening\.A watercolor botanicalillustration of orchids,ferns, and delicategreen leaves\.D\-FKC![[Uncaptioned image]](https://arxiv.org/html/2609.35947v1/images/cfg_case/pirate_cove_dfkc.jpg)![[Uncaptioned image]](https://arxiv.org/html/2609.35947v1/images/cfg_case/lighthouse_dfkc.jpg)![[Uncaptioned image]](https://arxiv.org/html/2609.35947v1/images/cfg_case/orchids_dfkc.jpg)D\-VCG![[Uncaptioned image]](https://arxiv.org/html/2609.35947v1/images/cfg_case/pirate_cove_dvcg.jpg)![[Uncaptioned image]](https://arxiv.org/html/2609.35947v1/images/cfg_case/lighthouse_dvcg.jpg)![[Uncaptioned image]](https://arxiv.org/html/2609.35947v1/images/cfg_case/orchids_dvcg.jpg)相似文章
Spectral Guidance:灵活高效的扩散模型控制方法
介绍了Spectral Guidance,一种通过利用扩散过程的低维表示来控制扩散模型的框架,无需任务特定的重新训练或通过去噪器的反向传播即可实现灵活稳定的控制。
TILT:利用模型内在奖励提升扩散模型中的组合生成
TILT是一种无需训练的框架,通过使用模型内在奖励在测试时对齐采样轨迹,从而改善扩散模型中的组合生成,无需外部监督即可增强对复杂提示的忠实度。
非均匀离散扩散:更强、更可扩展
本文提出 LUDI,一种非均匀扩散语言建模框架,解决了均匀扩散语言模型中训练目标过度均匀以及条件与目标混淆的问题,实现了 7B 规模的均匀离散扩散语言模型(UDLM),在逐步处理速度上比自回归解码快 3 倍,同时在复杂推理任务上具备有竞争力的性能。
线性约束下的条件扩散:Langevin 混合与信息论保证
本文分析了预训练扩散模型在线性逆问题上的零样本条件采样,提供了信息论保证并提出了一种投影 Langevin 初始化方法。
LangFlow:连续扩散在语言建模中可与离散扩散相媲美
LangFlow提出了首个可与离散扩散方法相媲美的连续扩散语言模型,挑战了长期以来认为连续扩散在语言建模中劣于离散扩散的观点。该工作引入了基于最优Gumbel噪声调度等关键要素,并展示了与离散扩散基线相比具有竞争力的困惑度和迁移学习性能。