PRISM: Principled Reference Identification for Schrodinger Bridge Model

arXiv cs.LG Papers

Summary

PRISM introduces a theory for designing reference processes in Schrödinger bridge models, showing that under finite computational budgets the optimal reference noise spectrum is determined by the sensor's information destruction spectrum. Experiments confirm the theory in Gaussian settings and identify where real images deviate.

arXiv:2608.06893v1 Announce Type: new Abstract: Schr\"odinger bridge models restore a clean signal from a degraded observation by following the conditional bridges of a reference process, yet this reference is chosen heuristically, typically white noise with a hand-tuned schedule. We develop PRISM, a theory of bridge reference design. We characterize the time-varying Gaussian references that remain exactly tractable with per-mode schedules: precisely those whose instantaneous covariances commute. We then prove an invisibility principle: with the exact drift and unlimited solver steps, every admissible reference recovers the true posterior. The choice of reference therefore matters only under finite computational resources. For a fixed step budget, we derive the finite-step objective in closed form and prove that every optimal noise spectrum is proportional to Pk, the spectrum of information destroyed by the sensor, with a mode-independent constant x*(T) = (2 ln T)^-1/2 (1 + o(1)). The analysis shows that noise color and temporal scheduling are interchangeable, and regularization provably shifts the optimal reference toward white noise. Experiments in Gaussian settings confirm the predicted orderings and the closed-form loss floors. On FFHQ, the distortion-- perception trade-off and spectral localization transfer, but white noise outperforms the matched reference; a pre-registered study that changes the training regime refutes ridge whitening as the explanation. A 2x2 mechanism study then traces the inversion to the non-Gaussian per-mode statistics of real images. PRISM turns reference design from a hyperparameter sweep into a calculation in the Gaussian regime, and locates exactly where real images break it.
Original Article
View Cached Full Text

Cached at: 08/10/26, 08:04 AM

# PRISM: Principled Reference Identification for Schrödinger Bridge Models
Source: [https://arxiv.org/html/2608.06893](https://arxiv.org/html/2608.06893)
###### Abstract

Schrödinger bridge models restore a clean signal from a degraded observation by following the conditional bridges of a reference process, yet this reference is chosen heuristically, typically white noise with a hand\-tuned schedule\. We develop PRISM, a theory of bridge reference design\. We characterize the time\-varying Gaussian references that remain exactly tractable with per\-mode schedules: precisely those whose instantaneous covariances commute\. We then prove an invisibility principle: with the exact drift and unlimited solver steps, every admissible reference recovers the true posterior\. The choice of reference therefore matters only under finite computational resources\. For a fixed step budget, we derive the finite\-step objective in closed form and prove that every optimal noise spectrum is proportional toPkP\_\{k\}, the spectrum of information destroyed by the sensor, with a mode\-independent constantx∗​\(T\)=\(2​ln⁡T\)−1/2​\(1\+o​\(1\)\)x^\{\*\}\(T\)=\(2\\ln T\)^\{\-1/2\}\(1\+o\(1\)\)\. The analysis shows that noise color and temporal scheduling are interchangeable, and regularization provably shifts the optimal reference toward white noise\. Experiments in Gaussian settings confirm the predicted orderings and the closed\-form loss floors\. On FFHQ, the distortion–perception trade\-off and spectral localization transfer, but white noise outperforms the matched reference; a pre\-registered study that changes the training regime refutes ridge whitening as the explanation\. A2×22\{\\times\}2mechanism study then traces the inversion to the non\-Gaussian per\-mode statistics of real images\. PRISM turns reference design from a hyperparameter sweep into a calculation in the Gaussian regime, and locates exactly where real images break it\.

## 1Introduction

Schrödinger bridge modelsLiuet al\.\([2023](https://arxiv.org/html/2608.06893#bib.bib2)\); De Bortoliet al\.\([2021](https://arxiv.org/html/2608.06893#bib.bib3)\); Shiet al\.\([2023](https://arxiv.org/html/2608.06893#bib.bib1)\)learn to transport samples between two distributions along the conditional bridges of a*reference process*\. This reference determines the intermediate states used during training, including how noise is distributed across spatial frequencies\. Yet it is usually chosen by hand, most often as white noise with a heuristic scalar schedule\. This practice treats the reference as a minor implementation detail\. Fig\.[1](https://arxiv.org/html/2608.06893#S1.F1)shows that it is not: at every solver budget there is an interior optimal noise level, and its location follows a simple, predictable rule\.

![Refer to caption](https://arxiv.org/html/2608.06893v1/x1.png)Figure 1:There is a right amount of reference noise, and its size is predictable\.\(a\) Each row is a solver sampling budgetTT; each column shows the injected reference noise level, as a scalexxrelative to the information destroyed by the sensor\. Too little noise \(left\) fails to recover missing detail; too much \(right\) obscures the reconstruction\. \(b\) Reconstruction error as a function ofxx, with one curve perTT\. Every curve has a minimum, showing that an intermediate noise level performs best\. \(c\) The minimum’s location follows Eq\. \([7](https://arxiv.org/html/2608.06893#S3.E7)\), matching the measured optimum at every tested budget\.Diffusion schedules were hand\-tuned until schedule theoryKarraset al\.\([2022](https://arxiv.org/html/2608.06893#bib.bib4)\); Kingmaet al\.\([2021](https://arxiv.org/html/2608.06893#bib.bib5)\)made their design principled\. Recent work has shown empirically that the spectrum of injected noise matters in one\-endpoint generative diffusionFalcket al\.\([2025](https://arxiv.org/html/2608.06893#bib.bib39)\); Jiralersponget al\.\([2025](https://arxiv.org/html/2608.06893#bib.bib41)\); Benitaet al\.\([2026](https://arxiv.org/html/2608.06893#bib.bib43)\); Esteves and Makadia \([2026](https://arxiv.org/html/2608.06893#bib.bib44)\)\. That setting, however, has no measurement operator that identifies a preferred spectrum\. Color is therefore selected by search rather than derived, with no theory describing when it should matter\. This paper develops the analogous theory for bridge references, which we call PRISM \(Principled Reference Identification for Schrödinger Bridge Models\), and finds the answer more subtle than a single optimal formula: the reference is irrelevant in the ideal limit, while its optimal finite\-resource design depends on solver steps or model error\.

Our analysis proceeds in a solvable linear\-Gaussian degradation model, where the observation isx1=hk​x0\+nx\_\{1\}=h\_\{k\}x\_\{0\}\+nper spatial frequencykk, with transfer functionhkh\_\{k\}, noise spectrumNkN\_\{k\}, and signal priorSkS\_\{k\}\. The central quantity is the Wiener residual spectrum

Pk=Sk​Nk/\(hk2​Sk\+Nk\),P\_\{k\}\\;=\\;\{S\_\{k\}N\_\{k\}\}/\(\{h\_\{k\}^\{2\}S\_\{k\}\+N\_\{k\}\}\),\(1\)which measures the posterior uncertainty remaining after observingx1x\_\{1\}; where the sensor sees well,Pk≈Nk≈0P\_\{k\}\\approx N\_\{k\}\\approx 0; where it is blind,Pk≈SkP\_\{k\}\\approx S\_\{k\}\. Thus,PkP\_\{k\}is the spectrum of the destroyed information and is computable directly from the known degradation model\.

##### Contributions\.

1. 1\.A tractability theorem \(§[3\.1](https://arxiv.org/html/2608.06893#S3.SS1)\)\.We characterize the time\-varying operator\-valued references that admitI2I^\{2\}SB\-style closed\-form bridges and decouple into independent scalar modes: their instantaneous covariances must pairwise commute, equivalently simultaneously diagonalizable\. We also give blockwise aliasing, linear\-drift, and Hilbert\-space extensions\.
2. 2\.An invisibility principle \(§[3\.2](https://arxiv.org/html/2608.06893#S3.SS2)\)\.With the exact drift andT→∞T\\to\\inftysteps, the terminal law equals the true posterior for every reference variancev\>0v\>0, and collapses to a point mass atv=0v=0\. The reference matters only once some resource \(solver steps or model accuracy\) is limited\.
3. 3\.Finite\-step theory \(§[3\.3](https://arxiv.org/html/2608.06893#S3.SS3)\)\.We derive theTT\-step terminal law in closed form\. The per\-mode KL depends on\(v,P\)\(v,P\)only throughx=v/Px=v/P, so every optimizer is proportional,vk=x∗​Pkv\_\{k\}=x^\{\*\}P\_\{k\}\. On uniform grids,x∗​\(T\)=\(2​ln⁡T\)−1/2​\(1\+o​\(1\)\)x^\{\*\}\(T\)=\(2\\ln T\)^\{\-1/2\}\(1\+o\(1\)\)and the optimal KL isln⁡T2​T2​\(1\+o​\(1\)\)\\tfrac\{\\ln T\}\{2T^\{2\}\}\(1\+o\(1\)\)\. A second change of variables, thezz\-identity, makes color and schedule exchangeable: under per\-mode schedules discretization error is color\-blind, and the optimal schedule is unique with a4​P/T4P/Tfloor\. On clustered grids the color objective has several local minima, so the two cannot be designed independently\.
4. 4\.Model error and learning \(§[3\.4](https://arxiv.org/html/2608.06893#S3.SS4)\)\.We extend the analysis to arbitrary \(gain, bias\) perturbations of the drift\. Given thezz\-schedule and error profile, sampling is color\-blind, and per\-mode learning is color\-blind for any scale\-equivariant learner; color matters only where scale symmetry breaks\. Shared ridge regularization breaks it in a definite direction, giving the sub\-proportional lawvk=x∗​\(n​Pk\)⋅Pkv\_\{k\}=x^\{\*\}\(nP\_\{k\}\)\\cdot P\_\{k\}: less data whitens the optimal reference, which lands between white noise andPkP\_\{k\}\.
5. 5\.Experiments, and a refuted prediction \(§[4](https://arxiv.org/html/2608.06893#S4)\)\.Numerics and trained predictors confirm the predicted orderings and closed\-form loss floors, and measure ridge whitening directly\. On FFHQKarraset al\.\([2019](https://arxiv.org/html/2608.06893#bib.bib61)\)at64×6464\{\\times\}64with a known degradation, the distortion–perception trade\-off and the predicted spectral localization transfer, but white stays the better perceptual reference\. A pre\-registered two\-regime study \(§[4\.4](https://arxiv.org/html/2608.06893#S4.SS4)\) refutes ridge whitening as the explanation, and a2×22\{\\times\}2mechanism study \(§[4\.5](https://arxiv.org/html/2608.06893#S4.SS5)\) shows that the inversion tracks the data, not the architecture, implicating the non\-Gaussian per\-mode statistics that the theory idealizes away\.

## 2Related Work

##### Schrödinger bridges and bridge matching\.

Diffusion Schrödinger Bridge \(DSB\) approximates iterative proportional fitting using score\-based diffusionDe Bortoliet al\.\([2021](https://arxiv.org/html/2608.06893#bib.bib3)\); DSBM and Iterative Markovian Fitting instead alternate bridge matching with Markovian projectionShiet al\.\([2023](https://arxiv.org/html/2608.06893#bib.bib1)\); Peluchetti \([2023](https://arxiv.org/html/2608.06893#bib.bib10)\), and Iterative Proportional Markovian Fitting unifies the two familiesKholkinet al\.\([2024](https://arxiv.org/html/2608.06893#bib.bib19)\); Gushchinet al\.\([2024b](https://arxiv.org/html/2608.06893#bib.bib20)\)\. A parallel line pursues simplified, non\-iterative, or provably light solversTanget al\.\([2024](https://arxiv.org/html/2608.06893#bib.bib12)\); Gushchinet al\.\([2024a](https://arxiv.org/html/2608.06893#bib.bib13)\); Peluchetti \([2023](https://arxiv.org/html/2608.06893#bib.bib10)\), task\-dependent state costs, and semi\-supervised couplingsLiuet al\.\([2024](https://arxiv.org/html/2608.06893#bib.bib14)\); Howardet al\.\([2026](https://arxiv.org/html/2608.06893#bib.bib22)\)\. These methods improve estimators built on a given reference\. We instead ask which reference should be used\.

##### Bridge models for restoration and inverse problems\.

Liuet al\.\([2023](https://arxiv.org/html/2608.06893#bib.bib2)\)uses the analytic Brownian\-bridge posterior to map paired degraded and clean images, whileZhouet al\.\([2024](https://arxiv.org/html/2608.06893#bib.bib23)\)formulate DDBMs over a general reference diffusion with VE/VP variants\. Follow\-up work improves faster and non\-Markovian samplingWanget al\.\([2024](https://arxiv.org/html/2608.06893#bib.bib24)\); Zhenget al\.\([2025](https://arxiv.org/html/2608.06893#bib.bib25)\), few\- or one\-step distillationHeet al\.\([2024](https://arxiv.org/html/2608.06893#bib.bib27)\); Gushchinet al\.\([2025](https://arxiv.org/html/2608.06893#bib.bib26)\), pretrained priorsWanget al\.\([2025a](https://arxiv.org/html/2608.06893#bib.bib30)\), mean\-reverting and stochastic\-control formulationsZhuet al\.\([2025](https://arxiv.org/html/2608.06893#bib.bib31)\), residual or energy\-shortened trajectoriesWanget al\.\([2026](https://arxiv.org/html/2608.06893#bib.bib32)\); Houet al\.\([2026](https://arxiv.org/html/2608.06893#bib.bib33)\), and regularization against exposure bias and distortionYaoet al\.\([2025](https://arxiv.org/html/2608.06893#bib.bib34)\)\. These methods mainly modify the drift or time horizon; the diffusion covariance remains isotropic, with a scalar scale chosen by tuning\.

##### Structured references and degradation\-aware forward processes\.

Gaussian Schrödinger bridges admit closed formsBunneet al\.\([2023](https://arxiv.org/html/2608.06893#bib.bib11)\), and richer reference dynamics such as multivariate Ornstein–Uhlenbeck processes, topology\-aware heat diffusions, analytic linear–quadratic bridges—have been studied recentlyZhang and Stumpf \([2026](https://arxiv.org/html/2608.06893#bib.bib15)\); Wyrwalet al\.\([2026](https://arxiv.org/html/2608.06893#bib.bib35)\); Chertkov \([2026](https://arxiv.org/html/2608.06893#bib.bib36)\)\. In these works, the reference or its inducing cost is fixed in advance; we treat it as unknown\.

##### Spectral shaping and colored noise\.

The closest line of work studies the spectrum of injected noise\. Fourier\-space analyses of the forward processFalcket al\.\([2025](https://arxiv.org/html/2608.06893#bib.bib39)\)motivate schedules that corrupt all frequencies at a matched rate\. Empirical studies report gains from blue noise, low\-frequency\-heavy noise, spectrally anisotropic forward noise, spectrally\-guided or per\-instance schedules, and frequency\-band redistribution at sampling timeHuanget al\.\([2024](https://arxiv.org/html/2608.06893#bib.bib40)\); Jiralersponget al\.\([2025](https://arxiv.org/html/2608.06893#bib.bib41)\); Scimecaet al\.\([2025](https://arxiv.org/html/2608.06893#bib.bib42)\); Benitaet al\.\([2026](https://arxiv.org/html/2608.06893#bib.bib43)\); Esteves and Makadia \([2026](https://arxiv.org/html/2608.06893#bib.bib44)\); Davidsonet al\.\([2026](https://arxiv.org/html/2608.06893#bib.bib45)\)\. These works show that noise color is an important design choice, but select it through search or end\-to\-end fitting\. This produces dataset\-specific rules without explaining why one spectrum should be preferred or when color should have no effect\. We provide both results in the bridge setting: the optimal spectrum is the destroyed\-information spectrumPkP\_\{k\}, and its effect provably vanishes in the exact\-drift, schedule\-adapted limit\.

##### Schedule design, step budgets, and distortion–perception\.

Variational Diffusion Models and EDM established schedule and weighting choices as important design variablesKingmaet al\.\([2021](https://arxiv.org/html/2608.06893#bib.bib5)\); Karraset al\.\([2022](https://arxiv.org/html/2608.06893#bib.bib4)\)\. Later work studied ELBO\-based reweighting, logSNR importance sampling, and corrections to common schedulesKingma and Gao \([2023](https://arxiv.org/html/2608.06893#bib.bib46)\); Linet al\.\([2024](https://arxiv.org/html/2608.06893#bib.bib47)\); Okadaet al\.\([2024](https://arxiv.org/html/2608.06893#bib.bib48)\)\. These methods optimize the*temporal*schedule of a fixed one\-endpoint process and evaluate it empirically under a given NFE budget\. Our finite\-step objective is closed\-form, so its optimum is proved rather than measured\. Moreover, thezz\-identity shows that reported color effects must be tested for schedule confounds\. Finally, the distortion–perception trade\-off is classicalBlau and Michaeli \([2018](https://arxiv.org/html/2608.06893#bib.bib49)\); Freirichet al\.\([2021](https://arxiv.org/html/2608.06893#bib.bib50)\)and remains important in bridge restorationYaoet al\.\([2025](https://arxiv.org/html/2608.06893#bib.bib34)\); Fallahet al\.\([2025](https://arxiv.org/html/2608.06893#bib.bib51)\)\. In our model, its factor\-two cost follows as a theorem, while the reference traces the entire curve\.

## 3PRISM: A Theory of Reference Design

### 3\.1Which References Are Exactly Tractable?

LetBBbe standard Brownian motion inℝd\\mathbb\{R\}^\{d\}, and letd​Xt=L​\(t\)​d​Bt\\,\\mathrm\{d\}X\_\{t\}=L\(t\)\\,\\,\\mathrm\{d\}B\_\{t\}withQ​\(t\):=L​\(t\)​L​\(t\)⊤Q\(t\):=L\(t\)L\(t\)^\{\\top\}\. Image\-to\-Image Schrödinger Bridge \(I2I^\{2\}SB\) methodsLiuet al\.\([2023](https://arxiv.org/html/2608.06893#bib.bib2)\); Wanget al\.\([2025b](https://arxiv.org/html/2608.06893#bib.bib59)\)require two closed\-form objects: the pinned marginalXt∣\(X0,X1\)X\_\{t\}\\mid\(X\_\{0\},X\_\{1\}\)for training\-state sampling and the reverse sub\-bridge kernelXs∣\(Xt,X0\)X\_\{s\}\\mid\(X\_\{t\},X\_\{0\}\)for ancestral sampling\. For any deterministic covariance scheduleQ​\(⋅\)Q\(\\cdot\), both are multivariate Gaussians with explicit matrix formulas \(Lemmas[1](https://arxiv.org/html/2608.06893#Thmlemma1)–[2](https://arxiv.org/html/2608.06893#Thmlemma2), App[6\.2](https://arxiv.org/html/2608.06893#S6.SS2)\)\. PracticalI2I^\{2\}SB, however, needs an exact decomposition into independent scalar modes, each with its own spatial\-frequency noise schedule\. The following result characterizes this condition\.

###### Theorem 1\(Commuting\-reference tractability\)\.

SupposeQ​\(t\)Q\(t\)is real symmetric and positive semidefinite for almost everyt∈\[0,1\]t\\in\[0,1\]\. Then the following are equivalent:

1. 1\.Pairwise commutativity:Q​\(s\)​Q​\(t\)=Q​\(t\)​Q​\(s\)Q\(s\)Q\(t\)=Q\(t\)Q\(s\)for alls,ts,toutside a common null set\.
2. 2\.A fixed modal basis:Q​\(t\)=U​diag⁡\(q1​\(t\),…,qd​\(t\)\)​U⊤Q\(t\)=U\\operatorname\{diag\}\\bigl\(q\_\{1\}\(t\),\\ldots,q\_\{d\}\(t\)\\bigr\)U^\{\\top\}, with a time\-independent orthogonalUU\.
3. 3\.In the coordinatesY=U⊤​XY=U^\{\\top\}X, the reference is a product of independent time\-changed scalar Brownian motions:d​Yk,t=qk​\(t\)​d​Bk,t\\,\\mathrm\{d\}Y\_\{k,t\}=\\sqrt\{q\_\{k\}\(t\)\}\\,\\,\\mathrm\{d\}B\_\{k,t\}\.

Under these conditions, withak​\(t\):=∫0tqka\_\{k\}\(t\):=\\int\_\{0\}^\{t\}q\_\{k\},vk:=ak​\(1\)v\_\{k\}:=a\_\{k\}\(1\), andρk​\(t\):=ak​\(t\)/vk\\rho\_\{k\}\(t\):=a\_\{k\}\(t\)/v\_\{k\}, the pinned bridge decomposes exactly over modes:

Yk,t∣\(yk,0,yk,1\)\\displaystyle Y\_\{k,t\}\\mid\(y\_\{k,0\},y\_\{k,1\}\)∼𝒩​\(μk,t,σk,t2\),\\displaystyle\\sim\\mathcal\{N\}\\\!\\left\(\\mu\_\{k,t\},\\sigma\_\{k,t\}^\{2\}\\right\),\(2\)μk,t\\displaystyle\\mu\_\{k,t\}=\(1−ρk​\(t\)\)​yk,0\+ρk​\(t\)​yk,1,\\displaystyle=\\bigl\(1\-\\rho\_\{k\}\(t\)\\bigr\)y\_\{k,0\}\+\\rho\_\{k\}\(t\)y\_\{k,1\},σk,t2\\displaystyle\\sigma\_\{k,t\}^\{2\}=vk​ρk​\(t\)​\(1−ρk​\(t\)\),\\displaystyle=v\_\{k\}\\rho\_\{k\}\(t\)\\bigl\(1\-\\rho\_\{k\}\(t\)\\bigr\),and for0<s<t≤10<s<t\\leq 1the reverse kernel is modewise Gaussian with mean coefficientrk​\(s,t\)=ρk​\(s\)/ρk​\(t\)r\_\{k\}\(s,t\)=\\rho\_\{k\}\(s\)/\\rho\_\{k\}\(t\)and varianceak​\(s\)​\(1−rk​\(s,t\)\)a\_\{k\}\(s\)\\bigl\(1\-r\_\{k\}\(s,t\)\\bigr\)\.

###### Corollary 1\(Schedule parameterization\)\.

The family in Theorem[1](https://arxiv.org/html/2608.06893#Thmtheorem1)is parameterized exactly, per mode, by a total variancevk\>0v\_\{k\}\>0, called the*color*, and an absolutely continuous nondecreasing scheduleρk:\[0,1\]→\[0,1\]\\rho\_\{k\}:\[0,1\]\\to\[0,1\]withρk​\(0\)=0\\rho\_\{k\}\(0\)=0andρk​\(1\)=1\\rho\_\{k\}\(1\)=1, called the*allocation*, throughqk​\(t\)=vk​ρ˙k​\(t\)q\_\{k\}\(t\)=v\_\{k\}\\dot\{\\rho\}\_\{k\}\(t\)\.

Three extensions matter in practice, all proved in App[6\.2](https://arxiv.org/html/2608.06893#S6.SS2)\.

\(i\) Static colors: references of the formQ​\(t\)=β​\(t\)​ΣQ\(t\)=\\beta\(t\)\\Sigma, used in prior fixed\-color work, form the strict subfamily where all modes share oneρ\\rho; whiteI2I^\{2\}SB corresponds toΣ=c​I\\Sigma=cI\.

\(ii\) Aliasing blocks: if everyQ​\(t\)Q\(t\)is block\-diagonal under one fixed orthogonal decomposition, the bridge decouples into independent finite\-dimensional blocks without within\-block commutativity\. This correctly models downsampling, where each frequency couples to its foldovers\.

\(iii\) Commuting linear drift: simultaneously diagonalizableF​\(t\)F\(t\)andQ​\(t\)Q\(t\)preserve exact modal tractability under modified Gaussian formulas, covering mean\-reverting references; a Hilbert\-space version holds under a trace condition\.

### 3\.2Solvable Model and Invisibility Principle

Theorem[1](https://arxiv.org/html/2608.06893#Thmtheorem1)separates the operator problem into scalar modes, so we work per mode and omit the indexkk\. For each spatial frequency, the degradation model is

x0∼𝒩​\(0,S\),x1=h​x0\+n,n∼𝒩​\(0,N\)\.x\_\{0\}\\sim\\mathcal\{N\}\(0,S\),\\qquad x\_\{1\}=h\\,x\_\{0\}\+n,\\quad n\\sim\\mathcal\{N\}\(0,N\)\.\(3\)The posterior isx0∣x1∼𝒩​\(W​x1,P\)x\_\{0\}\\mid x\_\{1\}\\sim\\mathcal\{N\}\(Wx\_\{1\},P\), whereW=S​h/\(h2​S\+N\)W=\{Sh\}/\(\{h^\{2\}S\+N\}\)is the Wiener gain andP=S​N/\(h2​S\+N\)\>0P=\{SN\}/\(\{h^\{2\}S\+N\}\)\>0is the residual variance\. For colorv≥0v\\geq 0and bridge levelρ∈\[0,1\]\\rho\\in\[0,1\], the reference bridge isxρ=\(1−ρ\)​x0\+ρ​x1\+v​ρ​\(1−ρ\)​ξx\_\{\\rho\}=\(1\-\\rho\)x\_\{0\}\+\\rho x\_\{1\}\+\\sqrt\{v\\rho\(1\-\\rho\)\}\\,\\xi\.

Let0=ρ0<⋯<ρT=10=\\rho\_\{0\}<\\dots<\\rho\_\{T\}=1be the discretization grid\. The*plug\-in\-mean ancestral sampler*—standardI2I^\{2\}SB with a Bayes\-optimal network—starts atx1x\_\{1\}\. At each level, it computesx^0=𝔼​\[x0∣xρi,x1\]\\hat\{x\}\_\{0\}=\\mathbb\{E\}\[x\_\{0\}\\mid x\_\{\\rho\_\{i\}\},x\_\{1\}\]and applies the reverse kernel from Theorem[1](https://arxiv.org/html/2608.06893#Thmtheorem1), replacingx0x\_\{0\}withx^0\\hat\{x\}\_\{0\}\.

Defineφ​\(ρ\):=\(1−ρ\)​P\+v​ρ\\varphi\(\\rho\):=\(1\-\\rho\)P\+v\\rhoand

K​\(ρ\)=Pφ​\(ρ\),M​\(ρ\)=v​ρ​Pφ​\(ρ\),Σ​\(ρ\)=\(1−ρ\)​φ​\(ρ\)\.K\(\\rho\)=\\frac\{P\}\{\\varphi\(\\rho\)\},\\quad M\(\\rho\)=\\frac\{v\\rho P\}\{\\varphi\(\\rho\)\},\\quad\\Sigma\(\\rho\)=\(1\-\\rho\)\\varphi\(\\rho\)\.\(4\)HereK​\(ρ\)K\(\\rho\)is the conditional gain,M​\(ρ\)=Var⁡\(x0∣xρ,x1\)M\(\\rho\)=\\operatorname\{Var\}\(x\_\{0\}\\mid x\_\{\\rho\},x\_\{1\}\)is the conditional variance, andΣ​\(ρ\)=Var⁡\(xρ∣x1\)\\Sigma\(\\rho\)=\\operatorname\{Var\}\(x\_\{\\rho\}\\mid x\_\{1\}\)\. Givenx1x\_\{1\}, every sampler state remains Gaussian:xρi∣x1∼𝒩​\(ci​x1,Vi\)x\_\{\\rho\_\{i\}\}\\mid x\_\{1\}\\sim\\mathcal\{N\}\(c\_\{i\}x\_\{1\},V\_\{i\}\), withcic\_\{i\}andViV\_\{i\}following an explicit affine recursion \(App[6\.3](https://arxiv.org/html/2608.06893#S6.SS3)\)\.

###### Theorem 2\(Invisibility of the reference\)\.

\(a\) Fixv\>0v\>0and any grid whose deficit sum in Lemma[6](https://arxiv.org/html/2608.06893#Thmlemma6)\(App[6\.3](https://arxiv.org/html/2608.06893#S6.SS3)\) vanishes asT→∞T\\to\\infty, including uniform grids and power gridsρi=\(i/T\)a\\rho\_\{i\}=\(i/T\)^\{a\}witha≥1a\\geq 1\. Then the terminal law converges to the true posterior:𝒩​\(c0​x1,V0\)⟶𝒩​\(W​x1,P\)\\mathcal\{N\}\(c\_\{0\}x\_\{1\},V\_\{0\}\)\\longrightarrow\\mathcal\{N\}\(Wx\_\{1\},P\)\. More precisely,D0/P→0D\_\{0\}/P\\to 0at rateO​\(log⁡T/T\)O\(\\log T/T\)on uniform grids andO​\(1/T\)O\(1/T\)on power grids witha\>1a\>1\. Because the terminal mean is exact,KL=D02/4​P2​\(1\+o​\(1\)\)\\mathrm\{KL\}=\{D\_\{0\}^\{2\}\}/\{4P^\{2\}\}\\,\\bigl\(1\+o\(1\)\\bigr\); consequently,KL=O​\(\(log⁡T/T\)2\)\\mathrm\{KL\}=O\\\!\\left\(\\left\(\{\\log T\}/\{T\}\\right\)^\{2\}\\right\)on uniform grids andKL=O​\(1/T2\)\\mathrm\{KL\}=O\(1/T^\{2\}\)on power grids witha\>1a\>1\. The limiting law is independent ofv\>0v\>0and the admissible schedule; the reference becomes invisible in the infinite\-step limit\. \(b\) At the boundaryv=0v=0, the terminal law isδ​\(W​x1\)\\delta\(Wx\_\{1\}\)for every finiteTT: the sampler collapses discontinuously to Wiener regression and has infinite KL divergence from the posterior\.

The proof \(App[6\.3](https://arxiv.org/html/2608.06893#S6.SS3)\) rests on three exact properties\. First,*the mean is exact at every finiteTT*: for every grid,TT, andv≥0v\\geq 0,ci=\(1−ρi\)​W\+ρic\_\{i\}=\(1\-\\rho\_\{i\}\)W\+\\rho\_\{i\}\. Replacingx0x\_\{0\}by its posterior mean removes only its conditional randomness, so finite\-step error affects only variance\.

Second,*the contraction telescopes exactly*: the step coefficient isAi=φ​\(ρi−1\)/φ​\(ρi\)A\_\{i\}=\{\\varphi\(\\rho\_\{i\-1\}\)\}/\{\\varphi\(\\rho\_\{i\}\)\}, so products of step coefficients reduce to exact ratios ofφ\\varphi\. Third, the*variance deficit*Di:=Σ​\(ρi\)−ViD\_\{i\}:=\\Sigma\(\\rho\_\{i\}\)\-V\_\{i\}satisfies a linear recursion with a nonnegative source term, yielding:

###### Corollary 2\(Systematic underdispersion\)\.

Supposev\>0v\>0\. ThenVi<Σ​\(ρi\)V\_\{i\}<\\Sigma\(\\rho\_\{i\}\)for everyi<Ti<T; in particular,V0<PV\_\{0\}<Pfor every finiteTT\. The plug\-in\-mean sampler has a single failure mode: it produces too little variance\.

Theorem[2](https://arxiv.org/html/2608.06893#Thmtheorem2)is the central negative result\. Any meaningful objective for reference design must assign a cost to a finite resource: solver steps \(§[3\.3](https://arxiv.org/html/2608.06893#S3.SS3)\) or model accuracy and data \(§[3\.4](https://arxiv.org/html/2608.06893#S3.SS4)\)\. Without such a cost, objectives become degenerate; for example, the integrated Bayes risk of the training target is minimized at the collapsed boundaryv=0v=0\(App[6\.3](https://arxiv.org/html/2608.06893#S6.SS3)\)\.

### 3\.3Steps: The Exact Finite\-Step Theory

#### The exact objective and the form of optimizer

Unrolling the variance\-deficit recursion with the telescoped products gives the deficit in closed form \(Lemma[6](https://arxiv.org/html/2608.06893#Thmlemma6), App[6\.3](https://arxiv.org/html/2608.06893#S6.SS3)\):

D0​\(v,P;grid\)=v​P3​∑i=1T\(Δ​ρi\)2ρi​φ​\(ρi\)​φ​\(ρi−1\)2,\\displaystyle D\_\{0\}\(v,P;\\text\{grid\}\)=vP^\{3\}\\sum\_\{i=1\}^\{T\}\\frac\{\(\\Delta\\rho\_\{i\}\)^\{2\}\}\{\\rho\_\{i\}\\,\\varphi\(\\rho\_\{i\}\)\\,\\varphi\(\\rho\_\{i\-1\}\)^\{2\}\},\(5\)u=1−D0/P,KL=1/2​\(u−1−ln⁡u\)\.\\displaystyle u=1\-\{D\_\{0\}\}/\{P\},\\qquad\\mathrm\{KL\}=\{1\}/\{2\}\\bigl\(u\-1\-\\ln u\\bigr\)\.This is the exactTT\-step objective\. Reference design minimizes∑kKLk\\sum\_\{k\}\\mathrm\{KL\}\_\{k\}over colors\{vk\}\\\{v\_\{k\}\\\}and, optionally, per\-mode schedules \(Cor\.[1](https://arxiv.org/html/2608.06893#Thmcorollary1)\)\.

###### Theorem 3\(Scale symmetry and proportionality\)\.

Fix any shared grid\. The per\-modeKL\\mathrm\{KL\}depends on\(v,P\)\(v,P\)only through the dimensionless ratiox=v/Px=v/P: a mode\-independent profileΦ\\PhisatisfiesKL​\(v,P;grid\)=Φ​\(v/P\)\\mathrm\{KL\}\(v,P;\\text\{grid\}\)=\\Phi\(v/P\)\. SinceΦ​\(x\)→∞\\Phi\(x\)\\to\\inftyasx→0x\\to 0orx→∞x\\to\\infty, interior optimizers exist\. The objective also separates:∑kKLk=∑kΦ​\(vk/Pk\)\\sum\_\{k\}\\mathrm\{KL\}\_\{k\}=\\sum\_\{k\}\\Phi\(v\_\{k\}/P\_\{k\}\)\.

Consequently:

\(i\) At every local optimizer,vk=xk​Pkv\_\{k\}=x\_\{k\}P\_\{k\}, where eachxkx\_\{k\}is a local minimizer of the*same*mode\-independent profileΦ\\Phi\.

\(ii\) At every global optimizer,xk∈arg​min⁡Φx\_\{k\}\\in\\operatorname\{arg\\,min\}\\Phifor everykk\.

\(iii\) Wheneverarg​min⁡Φ\\operatorname\{arg\\,min\}\\Phiis a singleton, every global optimizer is proportional to the posterior spectrum,

vk∗=x∗​\(grid,T\)​Pk,k=1,…,d,v\_\{k\}^\{\*\}=x^\{\*\}\(\\mathrm\{grid\},T\)\\,P\_\{k\},\\qquad k=1,\\dots,d,\(6\)with one mode\-independentx∗x^\{\*\}for everyTT\. If the minimizer is nonunique, as on some non\-uniform grids \(Prop\.[2](https://arxiv.org/html/2608.06893#Thmproposition2)\), modes may select different local minimizers, yielding non\-proportional*local*optima\.

This proportionality follows from an exact algebraic symmetry, not an asymptotic argument\. In the per\-mode Gaussian problem,PPis the only scale andv/Pv/Pthe only dimensionless parameter;TTcontrols only the constant\.

###### Proposition 1\(The proportionality constant\)\.

On the uniform grid, every minimizer satisfies

x∗​\(T\)=\(2​ln⁡T\)−1/2​\(1\+o​\(1\)\),x^\{\*\}\(T\)=\(2\\ln T\)^\{\-1/2\}\\bigl\(1\+o\(1\)\\bigr\),\(7\)and the optimal value satisfiesKL​\(x∗\)=ln⁡T2​T2​\(1\+o​\(1\)\)\\mathrm\{KL\}\(x^\{\*\}\)=\\frac\{\\ln T\}\{2T^\{2\}\}\\bigl\(1\+o\(1\)\\bigr\)\.

The proof \(App[6\.4](https://arxiv.org/html/2608.06893#S6.SS4)\) establishes the uniform expansionT​D0/P=x​ln⁡T\+12​x\+1\+o​\(1\)T\\,\{D\_\{0\}\}/\{P\}=x\\ln T\+\\frac\{1\}\{2x\}\+1\+o\(1\)on the windowx∈\[x0/3,3​x0\]x\\in\[x\_\{0\}/3,\\,3x\_\{0\}\]withx0=\(2​ln⁡T\)−1/2x\_\{0\}=\(2\\ln T\)^\{\-1/2\}, together with global lower bounds excluding minimizers outside the window\. A numerical certificate with exact derivatives independently corroborates the expansion; the finite\-TTremainder stabilizes near1\.151\.15overT=103T=10^\{3\}–10510^\{5\}\.

Practically,x∗≈0\.2x^\{\*\}\\approx 0\.2–0\.60\.6for realistic step counts: the optimal reference injects a modest fraction of the destroyed\-information spectrum, and that fraction shrinks only like\(ln⁡T\)−1/2\(\\ln T\)^\{\-1/2\}\.

###### Proposition 2\(Uniqueness is grid\-dependent\)\.

AtT=2T=2,Φ\\Phiis strictly unimodal\. On uniform grids, uniqueness is proved for2≤T≤1202\\leq T\\leq 120and numerically certified forT≤200T\\leq 200andT∈\{500,1000,5000\}T\\in\\\{500,1000,5000\\\}\(App[6\.4](https://arxiv.org/html/2608.06893#S6.SS4)\); extending the proof to generalTTremains an open problem\.

On general grids, uniqueness is false\. An analytic two\-cluster construction shows that, already atT=3T=3, the gridρ=\(1/\(1\+C\),1/2,1\)\\rho=\\bigl\(\{1\}/\(\{1\+C\}\),1/2,1\\bigr\)yields at least two local minima ofΦ\\Phifor everyC≥100C\\geq 100\(App[6\.4](https://arxiv.org/html/2608.06893#S6.SS4)\)\. AsC→∞C\\to\\infty, the profile splits into two unimodalT=2T=2profiles, one per cluster, at scalesx≍1x\\asymp 1andx≍Cx\\asymp C\. AT=48T=48two\-cluster grid has*seven*local minima: each step cluster favors its own noise scale\.

The counterexample carries a warning: on non\-uniform grids,*schedule design and color design cannot be decoupled*\.

###### Proposition 3\(Budgets bend the exponent\)\.

FixP1,…,PdP\_\{1\},\\dots,P\_\{d\}and a total\-noise budgetVVwithV​Pmax/∑jPj2≤34VP\_\{\\max\}/\\sum\_\{j\}P\_\{j\}^\{2\}\\leq\\tfrac\{3\}\{4\}\. On uniform grids, asT→∞T\\to\\infty, every global optimizer of∑kKLk\\sum\_\{k\}\\mathrm\{KL\}\_\{k\}subject to∑kvk=V\\sum\_\{k\}v\_\{k\}=Vsatisfies

vk∗=Pk2∑jPj2​V​\(1\+O​\(1ln⁡T\)\),k=1,…,d\.v\_\{k\}^\{\*\}=\\frac\{P\_\{k\}^\{2\}\}\{\\sum\_\{j\}P\_\{j\}^\{2\}\}\\,V\\,\\left\(1\+O\\\!\\left\(\\frac\{1\}\{\\ln T\}\\right\)\\right\),\\qquad k=1,\\dots,d\.\(8\)The fixed budget places every mode on the increasing branch ofΦ\\Phi: the free optimax∗​\(T\)​Pkx^\{\*\}\(T\)P\_\{k\}vanish asT→∞T\\to\\infty, but the budget remains fixed\. The optimal exponent therefore bends fromPkP\_\{k\}toPk2P\_\{k\}^\{2\}\. Equal\-budget comparisons, common in empirical color sweeps, silently change the optimization problem and steepen the optimum\. We report both conventions in §[4](https://arxiv.org/html/2608.06893#S4)because they answer different questions\.

#### Thezz\-identity: color and schedule are exchangeable

The key structure appears under the change of variablesz​\(ρ\):=v​ρ/φ​\(ρ\)=M​\(ρ\)/Pz\(\\rho\):=\{v\\rho\}/\{\\varphi\(\\rho\)\}=\{M\(\\rho\)\}/\{P\}, the normalized conditional\-MMSE level, withz0=0z\_\{0\}=0andzT=1z\_\{T\}=1\.

###### Theorem 4\(Exactzz\-identity and schedule theory\)\.

For every colorv\>0v\>0and every grid,

D0=P​∑i=1T\(zi−zi−1\)2/zi\.D\_\{0\}=P\\sum\_\{i=1\}^\{T\}\{\(z\_\{i\}\-z\_\{i\-1\}\)^\{2\}\}/\{z\_\{i\}\}\.\(9\)Consequently:

\(a\) Color\-freeness\. For everyv\>0v\>0, the mapρ↦z​\(ρ\)\\rho\\mapsto z\(\\rho\)is a bijection of\[0,1\]\[0,1\]\. Every color therefore achieves the same set ofzz\-grids, and hence the same schedule\-optimized deficit, at every finiteTT: under per\-mode schedule adaptation, discretization error cannot prefer any color\.

\(b\) Unique optimal schedule\. The objectiveF​\(z\)=∑i=1T\(Δ​zi\)2/ziF\(z\)=\\sum\_\{i=1\}^\{T\}\(\\Delta z\_\{i\}\)^\{2\}/z\_\{i\}is jointly convex, being a sum of quadratic\-over\-linear terms, and its minimizer over pinned grids is unique\.

\(c\) Explicit forward recurrence\. Withai=zi−1/zia\_\{i\}=z\_\{i\-1\}/z\_\{i\}, the first\-order conditions become2​ai\+1=1\+ai22a\_\{i\+1\}=1\+a\_\{i\}^\{2\},a1=0a\_\{1\}=0, giving an explicit forward recurrence \(a2=12a\_\{2\}=\\tfrac\{1\}\{2\},a3=58a\_\{3\}=\\tfrac\{5\}\{8\},a4=89128,…a\_\{4\}=\\tfrac\{89\}\{128\},\\ \\ldots\), with no shooting required\.

\(d\) The floor and the continuum limit\. The optimal deficit satisfiesD0∗=4​PT​\(1\+O​\(log2⁡T/T\)\)D\_\{0\}^\{\*\}=\\frac\{4P\}\{T\}\\bigl\(1\+O\(\{\\log^\{2\}T\}/\{T\}\)\\bigr\), giving a per\-mode KL floor∼4/T2\\sim 4/T^\{2\}\. The exact optimizer converges to the continuum schedulez=s2z=s^\{2\}at the explicit ratezi∗=\(i/T\)2​exp⁡\(O​\(ln⁡\(i\+1\)/\(i\+1\)\)\)z\_\{i\}^\{\*\}=\(i/T\)^\{2\}\\exp\\bigl\(O\(\{\\ln\(i\+1\)\}/\(\{i\+1\}\)\)\\bigr\); in the original variable,ρ∗​\(s\)=P​s2/\(P​s2\+v​\(1−s2\)\)\\rho^\{\*\}\(s\)=\{Ps^\{2\}\}/\\bigl\(\{Ps^\{2\}\+v\(1\-s^\{2\}\)\}\\bigr\), and the conditional MMSE is quadratic in solver time:M​\(ρ∗​\(s\)\)=P​s2M\(\\rho^\{\*\}\(s\)\)=Ps^\{2\}\.

Theorem[4](https://arxiv.org/html/2608.06893#Thmtheorem4)reframes the design space: the reference affects discretization only through itszz\-schedule, while per\-mode color is merely a reparameterization\. Two practical consequences follow\. With a*shared*schedule, as in existing implementations, Theorem[3](https://arxiv.org/html/2608.06893#Thmtheorem3)determines the color optimum\. With frequency\-adaptive schedules, discretization is exactly color\-free; any color effect must instead arise from learned drift or finite data\.

### 3\.4Model Error: Where Color Genuinely Lives

In the solvable model, the exact drift is affine\. Any learned affine drift therefore differs by two quantities per level: relative*gain error*ηi\\eta\_\{i\}and additive*bias*βi\\beta\_\{i\}\. Since everything is conditioned onx1x\_\{1\}, theβi\\beta\_\{i\}may depend on it; this absorbs errors in the drift’sx1x\_\{1\}coefficient\. This two\-channel decomposition is exhaustive within the model, and both channels preserve affine\-Gaussian structure, making all terminal laws below exact\.

###### Proposition 4\(Gain error cannot bias the sampler\)\.

With arbitrary gain errors\{ηi\}\\\{\\eta\_\{i\}\\\}, the terminal conditional mean is still exactlyW​x1Wx\_\{1\}\. Gain error acts purely on the variance channel; only bias moves the mean\.

###### Theorem 5\(Exactzz\-reduction of the perturbed sampler\)\.

Work at a fixedzz\-grid with per\-level error profile\(η​\(z\),β​\(z\)\)\(\\eta\(z\),\\beta\(z\)\)\. The perturbed contraction isA^i=Ai​\(1\+ηi​Δ​zizi\)\\hat\{A\}\_\{i\}=A\_\{i\}\\bigl\(1\+\\eta\_\{i\}\\frac\{\\Delta z\_\{i\}\}\{z\_\{i\}\}\\bigr\)\. A bias injected at leveliireaches the output with weightΔ​zizi​∏m<i\(1\+ηm​Δ​zmzm\)\\frac\{\\Delta z\_\{i\}\}\{z\_\{i\}\}\\prod\_\{m<i\}\\bigl\(1\+\\eta\_\{m\}\\frac\{\\Delta z\_\{m\}\}\{z\_\{m\}\}\\bigr\), which is exactlyΔ​zi/zi\\Delta z\_\{i\}/z\_\{i\}whenη≡0\\eta\\equiv 0\. Thus, the terminal mean error is

m0=∑i=1Tβi​Δ​zizi​∏m<i\(1\+ηm​Δ​zmzm\),m\_\{0\}=\\sum\_\{i=1\}^\{T\}\\beta\_\{i\}\\,\\frac\{\\Delta z\_\{i\}\}\{z\_\{i\}\}\\prod\_\{m<i\}\\left\(1\+\\eta\_\{m\}\\frac\{\\Delta z\_\{m\}\}\{z\_\{m\}\}\\right\),\(10\)and the terminal variance has the closed form

V0\\displaystyle V\_\{0\}=P​∑j=1Tzj−1​Δ​zjzj×∏m<j\(1\+ηm​Δ​zmzm\)2,\\displaystyle=P\\sum\_\{j=1\}^\{T\}z\_\{j\-1\}\\frac\{\\Delta z\_\{j\}\}\{z\_\{j\}\}\\times\\prod\_\{m<j\}\\left\(1\+\\eta\_\{m\}\\frac\{\\Delta z\_\{m\}\}\{z\_\{m\}\}\\right\)^\{2\},\(11\)whoseη≡0\\eta\\equiv 0case recoversV0/P=1−∑i\(Δ​zi\)2/ziV\_\{0\}/P=1\-\\sum\_\{i\}\(\\Delta z\_\{i\}\)^\{2\}/z\_\{i\}\. Every factor depends only on thezz\-grid and error profile\. Thus, the sampler’s terminal law is identical for every colorv\>0v\>0: given thezz\-schedule, sampling remains color\-blind under arbitrary drift error\. The terminal KL is

u=V0/P,KL=1/2​\(u\+m02/P−1−ln⁡u\)\.u=\{V\_\{0\}\}/\{P\},\\qquad\\mathrm\{KL\}=\{1\}/\{2\}\\left\(u\+\{m\_\{0\}^\{2\}\}/\{P\}\-1\-\\ln u\\right\)\.\(12\)

###### Proposition 5\(The learning problem is also color\-blind, per mode\)\.

At levelzz, the per\-mode joint law of\(x0,xt\)\(x\_\{0\},x\_\{t\}\)givenx1x\_\{1\}is determined up to scale byzzalone:corr2​\(x0,deviation∣x1\)=1−z\\mathrm\{corr\}^\{2\}\(x\_\{0\},\\,\\text\{deviation\}\\mid x\_\{1\}\)=1\-zexactly\. Hence any scale\-equivariant learner has a color\-free relative\-error law at eachzz\-level; for example, zero\-intercept OLS fromnnpairs hasVar⁡\(K^/K−1\)=z/\(n​\(1−z\)\)\\operatorname\{Var\}\(\\hat\{K\}/K\-1\)=z/\\big\(n\(1\-z\)\\big\)on a normalized design∑iXi2=n​τ2\\sum\_\{i\}X\_\{i\}^\{2\}=n\\tau^\{2\}\(random design replacesnnbyn−2n\-2\)\. With Theorem[5](https://arxiv.org/html/2608.06893#Thmtheorem5), color is therefore invisible end to end in the per\-mode scalar model, given thezz\-schedule\.

![Refer to caption](https://arxiv.org/html/2608.06893v1/x2.png)Figure 2:Exact finite\-step predictions\.\(a\) Invisibility \(Theorem[2](https://arxiv.org/html/2608.06893#Thmtheorem2)\)\. \(b\) Terminal KL rates, uniform vs\. power grid\. \(c,d\) The constant \(Prop\.[1](https://arxiv.org/html/2608.06893#Thmproposition1)\): measuredx∗​\(T\)x^\{\*\}\(T\)against\(2​ln⁡T\)−1/2\(2\\ln T\)^\{\-1/2\}, and their ratio\. \(e\) Reference multimodality \(Prop\.[2](https://arxiv.org/html/2608.06893#Thmproposition2)\): seven local minima ofΦ\\Phion a clusteredT=48T\{=\}48schedule, certified in 50\-digit arithmetic\.Color therefore matters only through what the per\-mode scalar model cannot see: \(a\) scale\-non\-equivariant learning, and \(b\) cross\-mode coupling\. Mechanism \(a\) is present in every trained network \(weight decay, ridge penalties, finite initialization scales\), and its minimal model already breaks proportionality in a definite direction:

###### Proposition 6\(Ridge whitens the optimal reference\)\.

A shared ridge penaltyλ\\lambdainduces relative shrinkageηk​\(z\)=−λ/\(λ\+n​Σk​\(z\)\)\\eta\_\{k\}\(z\)=\-\{\\lambda\}/\(\{\\lambda\+n\\Sigma\_\{k\}\(z\)\}\)\(same convention\), a*dimensionful*bias that breaks the scale symmetry protectingv∝Pv\\propto P\. Exact\-chain optimization yields a sharply decreasingx∗​\(n​P\)x^\{\*\}\(nP\): atT=100T\{=\}100andλ=1\\lambda\{=\}1, it falls from11\.9211\.92atn​P=10nP\{=\}10to0\.3630\.363asn​P→∞nP\\to\\infty\. Weakly observed frequencies therefore need disproportionately more reference noise\. The law bends sub\-proportionally,vk∗=x∗​\(n​Pk\)​Pkv\_\{k\}^\{\*\}=x^\{\*\}\(nP\_\{k\}\)P\_\{k\}, so*regularization whitens the optimal reference*; purev∝Pv\\propto Preturns only in the equivariant, infinite\-data limit\.

Two further results follow from this machinery \(proofs in App[6\.5](https://arxiv.org/html/2608.06893#S6.SS5)\)\.Distortion–perception trade\-off:the posterior\-mean estimator has MSE exactlyPP, independent of the reference, while the sampled output has MSE=P\+V0∈\[P,2​P\)=P\+V\_\{0\}\\in\[P,2P\); the classical factor\-2 cost of posterior sampling thus emerges as a theorem, with the reference tracing the entire trade\-off curve\.Principled underdispersion correction:since the sampler’s main failure is underdispersion \(Corollary[2](https://arxiv.org/html/2608.06893#Thmcorollary2)\), deliberate gain inflation is a calibrated fix, with an exact condition: chooseη​\(z\)\\eta\(z\)so that the reweighted sum in \([11](https://arxiv.org/html/2608.06893#S3.E11)\) reachesPP, the closed\-form counterpart of temperature and churn heuristics\. Theorem[5](https://arxiv.org/html/2608.06893#Thmtheorem5)also yields ameasurement protocolfor real networks: probe the trained drift’s Jacobian along the path to estimateη^​\(z\)\\hat\{\\eta\}\(z\), then*predict*the sampler’s terminal variance via \([11](https://arxiv.org/html/2608.06893#S3.E11)\), without retraining, just one forward analysis\. We use this protocol in §[4](https://arxiv.org/html/2608.06893#S4)\.

## 4Experiments

### 4\.1Exact Numerical Evaluation

The deterministic affine recursions of Sections[3\.2](https://arxiv.org/html/2608.06893#S3.SS2)–[3\.4](https://arxiv.org/html/2608.06893#S3.SS4)isolate finite\-step effects from sampling and training error; every quantity is computed in closed form and cross\-checked to machine precision \(App[7\.3](https://arxiv.org/html/2608.06893#S7.SS3)\)\.

Figure[2](https://arxiv.org/html/2608.06893#S3.F2)confirms each prediction\. ForP=v=4P\{=\}v\{=\}4, the terminal variance reaches3\.99953\.9995atT=105T\{=\}10^\{5\}, whilev=0v=0collapses \(Theorem[2](https://arxiv.org/html/2608.06893#Thmtheorem2)\)\. The measured optimum tracks the\(2​ln⁡T\)−1/2\(2\\ln T\)^\{\-1/2\}law, with the ratio falling from1\.291\.29atT=5T\{=\}5to1\.0171\.017atT=105T\{=\}10^\{5\}\(Prop\.[1](https://arxiv.org/html/2608.06893#Thmproposition1)\), and the clustered grid exhibits the predicted multimodality \(Prop\.[2](https://arxiv.org/html/2608.06893#Thmproposition2)\)\. In a two\-mode problem, the matched allocation minimizes KL at every testedTT; atT=200T\{=\}200, the Bayes\-risk vertex has294×294\\timeslarger KL \(Table[13](https://arxiv.org/html/2608.06893#S7.T13)in App[7\.3](https://arxiv.org/html/2608.06893#S7.SS3)\)\. Optimizedzz\-schedules remove reference\-color dependence to numerical precision \(Theorem[4](https://arxiv.org/html/2608.06893#Thmtheorem4)\(a\)\), and the budgeted optimum moves toward the predictedP2P^\{2\}allocation \(Prop\.[3](https://arxiv.org/html/2608.06893#Thmproposition3)\)\.

### 4\.2Learned Gaussian Models

We train per\-mode predictors on6464–256256independent modes withSk∼k−2S\_\{k\}\\sim k^\{\-2\}, a Gaussian modulation transfer function, and flat observation noise, comparing white, matched \(v∝Pv\\propto P\), anti\-matched \(v∝1/Pv\\propto 1/P\), and prior\-colored \(v∝Sv\\propto S\) references at equal total budget \(33seeds, common random numbers\)\.

![Refer to caption](https://arxiv.org/html/2608.06893v1/x3.png)Figure 3:Learned Gaussian models\.\(a\) Relative excess of the converged loss over the closed\-form Bayes floor forρ∈\[0\.1,0\.95\]\\rho\\in\[0\.1,0\.95\]: median0\.18%0\.18\\%against the pre\-registered5%5\\%gate\. \(b\) NFE\-vs\-KL by reference \(33seeds, mean±\\pmstd; dashed lines are exact\-drift floors\)\.The learned models reach the theory’s floor and reproduce its ordering \(Fig\.[3](https://arxiv.org/html/2608.06893#S4.F3)\)\. After1212k steps, the median relative excess over the closed\-form Bayes floor is0\.18%0\.18\\%, well inside the pre\-registered5%5\\%gate\. At NFE5050, total KL runs from0\.19±0\.030\.19\{\\pm\}0\.03\(matched\) to4\.3±0\.34\.3\{\\pm\}0\.3\(prior\-colored\), with white and anti\-matched between—the exact\-drift ordering, at learned drifts\. At NFE200200, gain error \(Prop\.[4](https://arxiv.org/html/2608.06893#Thmproposition4)\) dominates discretization error: matched and white overlap within seed variation, while the colored references become unstable \(App[7\.4](https://arxiv.org/html/2608.06893#S7.SS4)\)\.

The measured optimal scale agrees with theory where the model is accurate:0\.5750\.575vs\.0\.5770\.577atT=10T\{=\}10and0\.4020\.402vs\.0\.4040\.404atT=50T\{=\}50\. AtT=200T\{=\}200it sits8%8\\%above the exact\-drift value, the direction predicted for finite\-data regularization by Proposition[6](https://arxiv.org/html/2608.06893#Thmproposition6); the ridge sweep, the Jacobian probe \(median0\.6%0\.6\\%relative error in predicted terminal variance, with no retraining\), and the per\-mode\-schedule null are in App[7\.4](https://arxiv.org/html/2608.06893#S7.SS4)\.

Table 1:FFHQ restoration at NFE5050\(original recipe;σblur=2\.0\\sigma\_\{\\mathrm\{blur\}\}\{=\}2\.0,σn=0\.05\\sigma\_\{n\}\{=\}0\.05\)\. FID from5050k samples; white and matched use three seeds \(mean±\\pmstd\), others one, so bold gaps<0\.1\{<\}0\.1are within seed noise\.![Refer to caption](https://arxiv.org/html/2608.06893v1/x4.png)Figure 4:FFHQ under the original recipe\(common random numbers; matched and white mean±\\pmstd over33seeds\)\. \(a\) PSNR and \(b\) FID against NFE: anti\-matched attains the best distortion and the worst perception at every NFE—the distortion–perception trade\-off of App[6\.5](https://arxiv.org/html/2608.06893#S6.SS5)\. \(c\) Radially averaged error spectra at NFE5050against the exact\-drift prediction2​P​\(r\)−D0​\(r\)2P\(r\)\-D\_\{0\}\(r\)\(dashed\); reference differences concentrate near the MTF cutoff \(SSIM, LPIPS, error\-ratio view, training fingerprint: App[7\.5](https://arxiv.org/html/2608.06893#S7.SS5)\)\.
### 4\.3FFHQ Under a Known Degradation

We train bridge models on FFHQ at64×6464\{\\times\}64with known Gaussian blur and flat observation noise;PkP\_\{k\}and the matched reference are therefore fixed by the theory rather than tuned\. All references share one training recipe, one solver, and common evaluation randomness \(App[7\.5](https://arxiv.org/html/2608.06893#S7.SS5)\)\.

Two predictions transfer exactly\. First, the distortion–perception trade\-off: anti\-matched gives the best PSNR and SSIM and the worst LPIPS and FID at every NFE \(Table[1](https://arxiv.org/html/2608.06893#S4.T1), Fig\.[4](https://arxiv.org/html/2608.06893#S4.F4)\), as the theory predicts for a reference that suppresses variance where the sensor is blind: the posterior mean sharpens while sampled detail is lost \(App[6\.5](https://arxiv.org/html/2608.06893#S6.SS5)\)\. Second, the predicted spectral localization: reference\-dependent error differences concentrate near the MTF cutoff, whereP​\(k\)P\(k\)transitions \(Fig\.[4](https://arxiv.org/html/2608.06893#S4.F4)\(c\)\)\.

The fine\-grained matched\-first ordering does not transfer\. White and the fixedα=1\\alpha\{=\}1reference attain lower FID and LPIPS than matched over the tested NFE range \(Table[1](https://arxiv.org/html/2608.06893#S4.T1)\)\. Re\-evaluating every trained network on thezz\-optimal schedule of Theorem[4](https://arxiv.org/html/2608.06893#Thmtheorem4)improves all references by a similar amount and leaves the white–matched gap nearly unchanged \(0\.58→0\.530\.58\\to 0\.53; App[7\.5](https://arxiv.org/html/2608.06893#S7.SS5)\), so shared discretization error does not explain the gap\. By Theorem[5](https://arxiv.org/html/2608.06893#Thmtheorem5), what remains is the learned drift: either scale\-non\-equivariant learning \(Prop\.[6](https://arxiv.org/html/2608.06893#Thmproposition6)\) or cross\-mode coupling\. The next two studies separate them\.

Table 2:FFHQ under the converged recipe\(dropout0, batch512512,300300k steps;v∝Pkθv\\propto P\_\{k\}^\{\\theta\}at fixed budget\)\. FID from5050k samples; endpoints mean±\\pmstd over55seeds,θ\\thetaone seed\.
### 4\.4Training Regime and Reference Exponent

Ridge whitening \(Prop\.[6](https://arxiv.org/html/2608.06893#Thmproposition6)\) predicts that improved optimization should shift the optimal reference towardv∝Pkv\\propto P\_\{k\}\. We test it with a pre\-registered intervention and a sweepv∝Pkθv\\propto P\_\{k\}^\{\\theta\},θ∈\{0,14,12,34,1\}\\theta\\in\\\{0,\\tfrac\{1\}\{4\},\\tfrac\{1\}\{2\},\\tfrac\{3\}\{4\},1\\\}\. The converged recipe improves both matched and white in absolute terms, but it increases the matched–white FID gap on the same5050k evaluation set: the gap grows from1\.361\.36to2\.432\.43at NFE1010and from0\.580\.58to1\.181\.18at NFE5050\. The endpoint prediction is therefore refuted\.

Stronger blur \(σblur=4\.0\\sigma\_\{\\mathrm\{blur\}\}\{=\}4\.0\) widens the gap, and replication on CelebALiuet al\.\([2015](https://arxiv.org/html/2608.06893#bib.bib60)\)under the same experimental protocol reproduces both the white\-first ordering and the distortion–perception trade\-off \(App[7\.7](https://arxiv.org/html/2608.06893#S7.SS7)\)\. The sweep reveals budget dependence\. At low NFE, mild spectral coloring is beneficial:θ=0\.25\\theta\{=\}0\.25achieves the lowest observed FID for NFE≤20\\leq 20\(Table[2](https://arxiv.org/html/2608.06893#S4.T2)\)\. As the budget increases, however, the optimum shifts toward white, which performs best on all reported metrics for NFE≥50\\geq 50\. Thus, the experiments support partialPkP\_\{k\}\-coloring in the few\-step regime\. They leave open whether the practical failure of the per\-mode theory arises from dependence among Fourier modes, non\-Gaussian per\-mode marginals, or both\.

### 4\.5Mechanism Study: Data vs\. Coupling

Non\-Gaussianity, rather than mode coupling, explains the inversion\.Our2×22\{\\times\}2study separates model coupling from data statistics\. The white\-over\-matched inversion persists whenever real FFHQ statistics are retained, even when coupling is impossible\. On a Gaussian corpus with the same spectrum but no higher\-order cross\-mode structure, the matched reference regains its low\-NFE advantage: FID improves from19\.319\.3to13\.513\.5at NFE55and from7\.97\.9to5\.35\.3at NFE1010\. At NFE5050, the two references are nearly tied\. The coupling probe shows substantial off\-diagonal Fourier response in FFHQ\-trained U\-Nets, roughly tenfold less in Gaussian\-trained models, and numerical zero in per\-mode models\. The inversion tracks the data, not the coupling: it persists in per\-mode models where coupling is numerically zero, and it vanishes when a coupling\-capable U\-Net is trained on Gaussian data\. Coupling in FFHQ\-trained networks is therefore a symptom of non\-Gaussian image statistics, not the operative cause; the per\-mode Gaussianity assumption, not the decoupling assumption alone, is what fails on real images\. The near\-tie at NFE5050is consistent with the exact\-drift invisibility limit of Theorem[2](https://arxiv.org/html/2608.06893#Thmtheorem2), in which reference gaps vanish as the step budget grows\. The theorem does not guarantee this for learned, non\-Gaussian models, where drift error need not vanish with NFE; indeed, in §[4\.2](https://arxiv.org/html/2608.06893#S4.SS2)the colored references destabilize at NFE200200\.

## 5Conclusion

We have developed a theory of reference design for Schrödinger bridge models\. With exact drift and unlimited steps the reference is invisible; under a finite budget every optimum is proportional to the destroyed\-information spectrumPkP\_\{k\}with color and schedule exactly exchangeable\. On real images this ordering inverts, and controlled studies trace the inversion to non\-Gaussian per\-mode statistics\. The resulting recipe is budget\-dependent: mildPkP\_\{k\}coloring in the few\-step regime, white noise otherwise, with thezz\-optimal schedule in both cases\. Extending the theory beyond Gaussian conditional modes is therefore the main direction for future work\.

## 6Appendix

### 6\.1Overview

Apart from the six textbook facts collected in Table[4](https://arxiv.org/html/2608.06893#S6.T4), every argument is proved here from first principles\. Sections[6\.2](https://arxiv.org/html/2608.06893#S6.SS2)–[6\.5](https://arxiv.org/html/2608.06893#S6.SS5)present the proofs in the order in which the results appear in the main text\. Section[7](https://arxiv.org/html/2608.06893#S7)provides the experimental protocols and the additional results\.

Table[3](https://arxiv.org/html/2608.06893#S6.T3)collects the symbols\. Three conventions are in force everywhere below\.

1. 1\.*Conditioning on the observation\.*All statements about the sampler are conditional onx1x\_\{1\}\. Expectations, variances and KL divergences are conditional onx1x\_\{1\}unless said otherwise\.
2. 2\.*Per\-mode reduction\.*After Theorem[1](https://arxiv.org/html/2608.06893#Thmtheorem1)we work in the diagonalizing basis and drop the mode indexkkwhenever a statement is about a single mode\. Sums overkkreappear only when modes are coupled by a budget \(App[6\.4](https://arxiv.org/html/2608.06893#S6.SS4.SSSx5)\) or by a shared learner \(App[6\.5](https://arxiv.org/html/2608.06893#S6.SS5.SSSx5)\)\.
3. 3\.*Grids\.*A grid is0=ρ0<ρ1<⋯<ρT=10=\\rho\_\{0\}<\\rho\_\{1\}<\\dots<\\rho\_\{T\}=1, andiiruns fromTT\(the observation end\) down to0\(the clean end\), which is the direction the sampler travels\.

Table 3:Notation\.Deliberate overloads are marked†\\dagger; they never occur in the same argument\.Table[4](https://arxiv.org/html/2608.06893#S6.T4)lists the six textbook facts we use, with the tags \(S1\)–\(S6\) by which they are cited throughout\. Cauchy–Schwarz \(Theorem[6](https://arxiv.org/html/2608.06893#Thmtheorem6)\(ii\)\) and the mean value theorem \(Theorem[2](https://arxiv.org/html/2608.06893#Thmtheorem2)\(a\)\) are used only in their elementary forms and are not tabulated\.

Table 4:Standard facts, used without proof\.These are the only external ingredients in Apps\.[6\.2](https://arxiv.org/html/2608.06893#S6.SS2)–[6\.5](https://arxiv.org/html/2608.06893#S6.SS5)\.
### 6\.2Proofs for Section[3\.1](https://arxiv.org/html/2608.06893#S3.SS1)\(Tractability\)

#### The general Gaussian bridge

We first show that*every*deterministic matrix\-valued covariance gives a tractable bridge at the multivariate level\. Theorem[1](https://arxiv.org/html/2608.06893#Thmtheorem1)then identifies exactly when that tractability separates into scalar modes\.

###### Lemma 1\(Exact pinned marginal\)\.

For everyt∈\(0,1\)t\\in\(0,1\),

Xt∣\(X0=x0,X1=x1\)∼𝒩​\(μt,Γt\),X\_\{t\}\\mid\(X\_\{0\}\{=\}x\_\{0\},\\,X\_\{1\}\{=\}x\_\{1\}\)\\;\\sim\\;\\mathcal\{N\}\\bigl\(\\mu\_\{t\},\\,\\Gamma\_\{t\}\\bigr\),\(13\)where

μt\\displaystyle\\mu\_\{t\}:=x0\+A​\(t\)​V−1​\(x1−x0\),\\displaystyle:=x\_\{0\}\+A\(t\)V^\{\-1\}\(x\_\{1\}\-x\_\{0\}\),\(14\)Γt\\displaystyle\\Gamma\_\{t\}:=A​\(t\)−A​\(t\)​V−1​A​\(t\)\.\\displaystyle:=A\(t\)\-A\(t\)V^\{\-1\}A\(t\)\.\(15\)

###### Proof\.

Condition onX0=x0X\_\{0\}=x\_\{0\}\. Then

Xt=x0\+∫0tL​dB,X1=x0\+∫01L​dB,X\_\{t\}=x\_\{0\}\+\\int\_\{0\}^\{t\}L\\,\\mathrm\{d\}B,\\qquad X\_\{1\}=x\_\{0\}\+\\int\_\{0\}^\{1\}L\\,\\mathrm\{d\}B,\(16\)soVar⁡\(Xt∣X0\)=A​\(t\)\\operatorname\{Var\}\(X\_\{t\}\\mid X\_\{0\}\)=A\(t\)andVar⁡\(X1∣X0\)=V\\operatorname\{Var\}\(X\_\{1\}\\mid X\_\{0\}\)=V\. Independence of Brownian increments givesCov⁡\(Xt,X1∣X0\)=A​\(t\)\\operatorname\{Cov\}\(X\_\{t\},X\_\{1\}\\mid X\_\{0\}\)=A\(t\)\. Applying Gaussian conditioning \(S1\) to the jointly Gaussian pair\(Xt,X1\)\(X\_\{t\},X\_\{1\}\)proves the claim\. ∎

###### Lemma 2\(Exact reverse sub\-bridge kernel\)\.

For0<s<t≤10<s<t\\leq 1withA​\(t\)≻0A\(t\)\\succ 0,

Xs∣\(Xt=xt,X0=x0\)∼𝒩​\(μs∣t,Γs∣t\),X\_\{s\}\\mid\(X\_\{t\}\{=\}x\_\{t\},\\,X\_\{0\}\{=\}x\_\{0\}\)\\;\\sim\\;\\mathcal\{N\}\\bigl\(\\mu\_\{s\\mid t\},\\,\\Gamma\_\{s\\mid t\}\\bigr\),\(17\)where

μs∣t\\displaystyle\\mu\_\{s\\mid t\}:=x0\+A​\(s\)​A​\(t\)−1​\(xt−x0\),\\displaystyle:=x\_\{0\}\+A\(s\)A\(t\)^\{\-1\}\(x\_\{t\}\-x\_\{0\}\),\(18\)Γs∣t\\displaystyle\\Gamma\_\{s\\mid t\}:=A​\(s\)−A​\(s\)​A​\(t\)−1​A​\(s\),\\displaystyle:=A\(s\)\-A\(s\)A\(t\)^\{\-1\}A\(s\),\(19\)and moreoverℒ​\(Xs∣Xt,X0,X1\)=ℒ​\(Xs∣Xt,X0\)\\mathcal\{L\}\(X\_\{s\}\\mid X\_\{t\},X\_\{0\},X\_\{1\}\)=\\mathcal\{L\}\(X\_\{s\}\\mid X\_\{t\},X\_\{0\}\)\.

###### Proof\.

Condition onX0X\_\{0\}\. The pair\(Xs,Xt\)\(X\_\{s\},X\_\{t\}\)is jointly Gaussian withVar⁡\(Xs∣X0\)=A​\(s\)\\operatorname\{Var\}\(X\_\{s\}\\mid X\_\{0\}\)=A\(s\),Var⁡\(Xt∣X0\)=A​\(t\)\\operatorname\{Var\}\(X\_\{t\}\\mid X\_\{0\}\)=A\(t\)andCov⁡\(Xs,Xt∣X0\)=A​\(s\)\\operatorname\{Cov\}\(X\_\{s\},X\_\{t\}\\mid X\_\{0\}\)=A\(s\), so Gaussian conditioning \(S1\) gives the stated kernel\. The final equality is the two\-sided Markov property \(S3\): after conditioning onXtX\_\{t\}, the past is independent of the future endpoint\.

IfA​\(t\)A\(t\)is singular, the directions inker⁡A​\(t\)\\ker A\(t\)have accumulated no noise by timettand remain deterministically equal tox0x\_\{0\}\. Apply the kernel onran⁡A​\(t\)\\operatorname\{ran\}A\(t\), or equivalently replaceA​\(t\)−1A\(t\)^\{\-1\}by the pseudoinverse\. In the modal setting this is the convention for modes withak​\(t\)=0a\_\{k\}\(t\)=0\. On the sampler’s grids onlyρ0=0\\rho\_\{0\}=0is degenerate, and no reverse step is taken there\. ∎

#### Proof of Theorem[1](https://arxiv.org/html/2608.06893#Thmtheorem1)

###### Proof\.

\(1\)⇒\\Rightarrow\(2\)\.EachQ​\(t\)Q\(t\)is real symmetric, hence orthogonally diagonalizable\. The linear span of\{Q​\(t\):t∉N\}\\\{Q\(t\):t\\notin N\\\}is a finite\-dimensional subspace of the symmetric matrices, so finitely many membersQ​\(t1\),…,Q​\(tm\)Q\(t\_\{1\}\),\\dots,Q\(t\_\{m\}\)of the family span it\. By \(1\) these commute pairwise, and a finite commuting family of real symmetric matrices is simultaneously orthogonally diagonalizable \(S2\): diagonalize one matrix, observe that every other matrix preserves its eigenspaces, and recurse on the restrictions\. The resultingUUdiagonalizesQ​\(t1\),…,Q​\(tm\)Q\(t\_\{1\}\),\\dots,Q\(t\_\{m\}\), hence every matrix in their span, hence everyQ​\(t\)Q\(t\)outsideNN\. RedefiningQQon a null set does not change the law of the SDE\. FinallyQ⪰0Q\\succeq 0forces all diagonal entries to be nonnegative\.

\(2\)⇒\\Rightarrow\(3\)\.In the coordinatesY=U⊤​XY=U^\{\\top\}Xthe increment covariance isU⊤​Q​\(t\)​U​d​t=diag⁡\(qk​\(t\)\)​d​tU^\{\\top\}Q\(t\)U\\,\\mathrm\{d\}t=\\operatorname\{diag\}\(q\_\{k\}\(t\)\)\\,\\mathrm\{d\}t, so we may choose a deterministic diagonal square root\. The coordinates are jointly Gaussian with zero cross\-covariance at all times, hence independent scalar Gaussian processes, and each is a Brownian motion under the clockak​\(t\)a\_\{k\}\(t\)\.

\(3\)⇒\\Rightarrow\(1\)\.Independent scalar modes in a fixed basis have diagonal instantaneous covariance in that basis, so allQ​\(t\)Q\(t\)share the diagonalizerUU; diagonal matrices commute\.

The pinned formulas\.Substitute the diagonalizationsA​\(t\)=U​diag⁡\(ak​\(t\)\)​U⊤A\(t\)=U\\operatorname\{diag\}\(a\_\{k\}\(t\)\)U^\{\\top\}andV=U​diag⁡\(vk\)​U⊤V=U\\operatorname\{diag\}\(v\_\{k\}\)U^\{\\top\}into Lemma[1](https://arxiv.org/html/2608.06893#Thmlemma1)\. In modekkthe mean is

yk,0\+ak​\(t\)vk​\(yk,1−yk,0\)=\(1−ρk\)​yk,0\+ρk​yk,1,y\_\{k,0\}\+\\frac\{a\_\{k\}\(t\)\}\{v\_\{k\}\}\\bigl\(y\_\{k,1\}\-y\_\{k,0\}\\bigr\)=\(1\-\\rho\_\{k\}\)y\_\{k,0\}\+\\rho\_\{k\}y\_\{k,1\},\(20\)and the variance is

ak​\(t\)−ak​\(t\)2vk=vk​ρk​\(1−ρk\)\.a\_\{k\}\(t\)\-\\frac\{a\_\{k\}\(t\)^\{2\}\}\{v\_\{k\}\}=v\_\{k\}\\,\\rho\_\{k\}\(1\-\\rho\_\{k\}\)\.\(21\)The conditional covariance is diagonal, so the modes remain independent\. Substituting into Lemma[2](https://arxiv.org/html/2608.06893#Thmlemma2)gives the reverse mean coefficientak​\(s\)/ak​\(t\)=rk​\(s,t\)a\_\{k\}\(s\)/a\_\{k\}\(t\)=r\_\{k\}\(s,t\)and the varianceak​\(s\)​\(1−rk​\(s,t\)\)a\_\{k\}\(s\)\\bigl\(1\-r\_\{k\}\(s,t\)\\bigr\)\. ∎

#### Proof of Corollary[1](https://arxiv.org/html/2608.06893#Thmcorollary1)

###### Proof\.

Absolute continuity givesak​\(t\)=∫0tqk=vk​ρk​\(t\)a\_\{k\}\(t\)=\\int\_\{0\}^\{t\}q\_\{k\}=v\_\{k\}\\rho\_\{k\}\(t\)\. Nonnegativity ofqkq\_\{k\}is equivalent to monotonicity ofρk\\rho\_\{k\}, and the endpoint conditions giveak​\(0\)=0a\_\{k\}\(0\)=0andak​\(1\)=vka\_\{k\}\(1\)=v\_\{k\}\. Conversely, defineρk:=ak/vk\\rho\_\{k\}:=a\_\{k\}/v\_\{k\}\. ∎

##### Four consequences stated only here\.

###### Corollary 3\(Static color is a strict subfamily\)\.

SupposeQ​\(t\)=β​\(t\)​ΣQ\(t\)=\\beta\(t\)\\SigmawithΣ=U​diag⁡\(vk\)​U⊤\\Sigma=U\\operatorname\{diag\}\(v\_\{k\}\)U^\{\\top\}and∫01β=1\\int\_\{0\}^\{1\}\\beta=1\. Then Theorem[1](https://arxiv.org/html/2608.06893#Thmtheorem1)applies with the*shared*scheduleρk≡ρ​\(t\)=∫0tβ\\rho\_\{k\}\\equiv\\rho\(t\)=\\int\_\{0\}^\{t\}\\betaand mode\-dependent totalsvkv\_\{k\}\. WhiteI2I^\{2\}SB is the special caseΣ=c​I\\Sigma=cI\. A fixed color therefore allows differentvkv\_\{k\}but forces every mode to use the sameρ\\rho\.

###### Corollary 4\(Aliasing blocks\)\.

Supposeℝd=E1⊕⋯⊕Em\\mathbb\{R\}^\{d\}=E\_\{1\}\\oplus\\dots\\oplus E\_\{m\}is a fixed orthogonal decomposition and everyQ​\(t\)Q\(t\)is block diagonal with respect to it\. Then Lemmas[1](https://arxiv.org/html/2608.06893#Thmlemma1)–[2](https://arxiv.org/html/2608.06893#Thmlemma2)decompose the bridge into independent finite\-dimensional blocks,*even when the matrices inside a block do not commute*; each block uses the matrix formulas inA​\(t\)A\(t\)andVV\. This is the correct treatment of downsampling aliasing: a frequency and its foldovers form one coupled block, and small matrix\-valued block schedules replace scalar schedules\.

###### Proof\.

LetΠj\\Pi\_\{j\}be the orthogonal projection ontoEjE\_\{j\}\. Since everyQ​\(t\)Q\(t\)is block diagonal, so areA​\(t\)=∫0tQA\(t\)=\\int\_\{0\}^\{t\}QandV=A​\(1\)V=A\(1\)\. WritingXt\(j\):=Πj​XtX^\{\(j\)\}\_\{t\}:=\\Pi\_\{j\}X\_\{t\}, forj≠lj\\neq lthe increments satisfy

Cov⁡\(Xt\(j\)−Xs\(j\),Xt\(l\)−Xs\(l\)\)=Πj​\(A​\(t\)−A​\(s\)\)​Πl⊤=0,\\operatorname\{Cov\}\\bigl\(X^\{\(j\)\}\_\{t\}\-X^\{\(j\)\}\_\{s\},\\;X^\{\(l\)\}\_\{t\}\-X^\{\(l\)\}\_\{s\}\\bigr\)=\\Pi\_\{j\}\\bigl\(A\(t\)\-A\(s\)\\bigr\)\\Pi\_\{l\}^\{\\top\}=0,\(22\)so the jointly Gaussian block processes\(X\(j\)\)j\(X^\{\(j\)\}\)\_\{j\}are mutually independent\. Consequently the pinned marginal and the reverse kernel factor over blocks, with the matrix formulas applied blockwise toA\(j\)​\(t\):=Πj​A​\(t\)​Πj⊤A^\{\(j\)\}\(t\):=\\Pi\_\{j\}A\(t\)\\Pi\_\{j\}^\{\\top\}andV\(j\):=Πj​V​Πj⊤V^\{\(j\)\}:=\\Pi\_\{j\}V\\Pi\_\{j\}^\{\\top\}\. No commutativity within a block is used\. ∎

###### Corollary 5\(Simultaneously diagonal linear drift and diffusion\)\.

Considerd​Xt=F​\(t\)​Xt​d​t\+L​\(t\)​d​Bt\\,\\mathrm\{d\}X\_\{t\}=F\(t\)X\_\{t\}\\,\\mathrm\{d\}t\+L\(t\)\\,\\mathrm\{d\}B\_\{t\}and suppose a single orthogonalUUdiagonalizes both coefficients,U⊤​F​\(t\)​U=diag⁡\(fk​\(t\)\)U^\{\\top\}F\(t\)U=\\operatorname\{diag\}\(f\_\{k\}\(t\)\)andU⊤​Q​\(t\)​U=diag⁡\(qk​\(t\)\)U^\{\\top\}Q\(t\)U=\\operatorname\{diag\}\(q\_\{k\}\(t\)\)a\.e\. Put

gk​\(t,s\):=exp​∫stfk,σk2​\(t\):=∫0tgk​\(t,u\)2​qk​\(u\)​du\.g\_\{k\}\(t,s\):=\\exp\\\!\\int\_\{s\}^\{t\}f\_\{k\},\\qquad\\sigma^\{2\}\_\{k\}\(t\):=\\int\_\{0\}^\{t\}g\_\{k\}\(t,u\)^\{2\}q\_\{k\}\(u\)\\,\\mathrm\{d\}u\.\(23\)Then the reference and all pinned bridges decompose exactly into scalar modes, with

𝔼​\[Yk,t∣yk,0,yk,1\]\\displaystyle\\mathbb\{E\}\[Y\_\{k,t\}\\mid y\_\{k,0\},y\_\{k,1\}\]=gk​\(t,0\)​yk,0\\displaystyle=g\_\{k\}\(t,0\)\\,y\_\{k,0\}\+σk2​\(t\)​gk​\(1,t\)σk2​\(1\)​\(yk,1−gk​\(1,0\)​yk,0\),\\displaystyle\\quad\+\\frac\{\\sigma^\{2\}\_\{k\}\(t\)\\,g\_\{k\}\(1,t\)\}\{\\sigma^\{2\}\_\{k\}\(1\)\}\\bigl\(y\_\{k,1\}\-g\_\{k\}\(1,0\)y\_\{k,0\}\\bigr\),\(24\)Var⁡\(Yk,t∣yk,0,yk,1\)\\displaystyle\\operatorname\{Var\}\(Y\_\{k,t\}\\mid y\_\{k,0\},y\_\{k,1\}\)=σk2​\(t\)\\displaystyle=\\sigma^\{2\}\_\{k\}\(t\)−σk2​\(t\)2​gk​\(1,t\)2σk2​\(1\),\\displaystyle\\quad\-\\frac\{\\sigma^\{2\}\_\{k\}\(t\)^\{2\}\\,g\_\{k\}\(1,t\)^\{2\}\}\{\\sigma^\{2\}\_\{k\}\(1\)\},\(25\)and for0<s<t0<s<tthe reverse kernel has mean

gk​\(s,0\)​yk,0\+σk2​\(s\)​gk​\(t,s\)σk2​\(t\)​\(yk,t−gk​\(t,0\)​yk,0\)g\_\{k\}\(s,0\)y\_\{k,0\}\+\\frac\{\\sigma^\{2\}\_\{k\}\(s\)\\,g\_\{k\}\(t,s\)\}\{\\sigma^\{2\}\_\{k\}\(t\)\}\\bigl\(y\_\{k,t\}\-g\_\{k\}\(t,0\)y\_\{k,0\}\\bigr\)\(26\)and varianceσk2​\(s\)−σk2​\(s\)2​gk​\(t,s\)2/σk2​\(t\)\\sigma^\{2\}\_\{k\}\(s\)\-\\sigma^\{2\}\_\{k\}\(s\)^\{2\}g\_\{k\}\(t,s\)^\{2\}/\\sigma^\{2\}\_\{k\}\(t\)\.

###### Proof\.

Variation of constants \(S3\) givesYk,t=gk​\(t,0\)​yk,0\+∫0tgk​\(t,u\)​qk​\(u\)​dBk,uY\_\{k,t\}=g\_\{k\}\(t,0\)y\_\{k,0\}\+\\int\_\{0\}^\{t\}g\_\{k\}\(t,u\)\\sqrt\{q\_\{k\}\(u\)\}\\,\\mathrm\{d\}B\_\{k,u\}, whence

Var⁡\(Yk,t∣yk,0\)\\displaystyle\\operatorname\{Var\}\(Y\_\{k,t\}\\mid y\_\{k,0\}\)=σk2​\(t\),\\displaystyle=\\sigma^\{2\}\_\{k\}\(t\),\(27\)Cov⁡\(Yk,t,Yk,1∣yk,0\)\\displaystyle\\operatorname\{Cov\}\(Y\_\{k,t\},Y\_\{k,1\}\\mid y\_\{k,0\}\)=σk2​\(t\)​gk​\(1,t\),\\displaystyle=\\sigma^\{2\}\_\{k\}\(t\)\\,g\_\{k\}\(1,t\),\(28\)Cov⁡\(Yk,s,Yk,t∣yk,0\)\\displaystyle\\operatorname\{Cov\}\(Y\_\{k,s\},Y\_\{k,t\}\\mid y\_\{k,0\}\)=σk2​\(s\)​gk​\(t,s\)\.\\displaystyle=\\sigma^\{2\}\_\{k\}\(s\)\\,g\_\{k\}\(t,s\)\.\(29\)Gaussian conditioning \(S1\) gives the stated formulas, and independence across modes follows from the common diagonalization\. ∎

### 6\.3Proofs for Section[3\.2](https://arxiv.org/html/2608.06893#S3.SS2)\(Invisibility\)

#### The sampler as an affine recursion

Writec¯i:=\(1−ρi\)​W\+ρi\\bar\{c\}\_\{i\}:=\(1\-\\rho\_\{i\}\)W\+\\rho\_\{i\}\. The plug\-in step from levelρi\\rho\_\{i\}toρi−1\\rho\_\{i\-1\}is

x^0\\displaystyle\\hat\{x\}\_\{0\}=W​x1\+K​\(ρi\)​\(xρi−c¯i​x1\),\\displaystyle=Wx\_\{1\}\+K\(\\rho\_\{i\}\)\\bigl\(x\_\{\\rho\_\{i\}\}\-\\bar\{c\}\_\{i\}x\_\{1\}\\bigr\),\(30\)xρi−1\\displaystyle x\_\{\\rho\_\{i\-1\}\}=x^0\+ri​\(xρi−x^0\)\+qi​ξi,\\displaystyle=\\hat\{x\}\_\{0\}\+r\_\{i\}\\bigl\(x\_\{\\rho\_\{i\}\}\-\\hat\{x\}\_\{0\}\\bigr\)\+\\sqrt\{q\_\{i\}\}\\,\\xi\_\{i\},\(31\)withri=ρi−1/ρir\_\{i\}=\\rho\_\{i\-1\}/\\rho\_\{i\}andqi=v​ρi−1​\(1−ri\)q\_\{i\}=v\\rho\_\{i\-1\}\(1\-r\_\{i\}\)\. Every operation is affine in\(x,x1\)\(x,x\_\{1\}\)and adds independent Gaussian noise\. Conditional onx1x\_\{1\}the state is therefore*exactly*𝒩​\(ci​x1,Vi\)\\mathcal\{N\}\(c\_\{i\}x\_\{1\},V\_\{i\}\), where

Ai\\displaystyle A\_\{i\}=ri\+\(1−ri\)​K​\(ρi\),\\displaystyle=r\_\{i\}\+\(1\-r\_\{i\}\)K\(\\rho\_\{i\}\),\(32\)Bi\\displaystyle B\_\{i\}=\(1−ri\)​\(W−K​\(ρi\)​c¯i\),\\displaystyle=\(1\-r\_\{i\}\)\\bigl\(W\-K\(\\rho\_\{i\}\)\\bar\{c\}\_\{i\}\\bigr\),\(33\)ci−1\\displaystyle c\_\{i\-1\}=Ai​ci\+Bi,\\displaystyle=A\_\{i\}c\_\{i\}\+B\_\{i\},cT=1,\\displaystyle c\_\{T\}=1,\(34\)Vi−1\\displaystyle V\_\{i\-1\}=Ai2​Vi\+qi,\\displaystyle=A\_\{i\}^\{2\}V\_\{i\}\+q\_\{i\},VT=0\.\\displaystyle V\_\{T\}=0\.\(35\)
###### Lemma 3\(Mean exactness at everyTT\)\.

ci=\(1−ρi\)​W\+ρic\_\{i\}=\(1\-\\rho\_\{i\}\)W\+\\rho\_\{i\}for allii, for any grid, anyT≥1T\\geq 1and anyv≥0v\\geq 0\. In particularc0=Wc\_\{0\}=Wexactly\.

###### Proof\.

Backward induction\. Base case:cT=1=\(1−ρT\)​W\+ρTc\_\{T\}=1=\(1\-\\rho\_\{T\}\)W\+\\rho\_\{T\}\. Step: substituteci=c¯ic\_\{i\}=\\bar\{c\}\_\{i\}into \([34](https://arxiv.org/html/2608.06893#S6.E34)\)\. The two terms containingK​\(ρi\)​c¯iK\(\\rho\_\{i\}\)\\bar\{c\}\_\{i\}cancel exactly, leaving

ci−1\\displaystyle c\_\{i\-1\}=ri​c¯i\+\(1−ri\)​W\\displaystyle=r\_\{i\}\\bar\{c\}\_\{i\}\+\(1\-r\_\{i\}\)W=W​\(1−ri​ρi\)\+ri​ρi\\displaystyle=W\(1\-r\_\{i\}\\rho\_\{i\}\)\+r\_\{i\}\\rho\_\{i\}=\(1−ρi−1\)​W\+ρi−1,\\displaystyle=\(1\-\\rho\_\{i\-1\}\)W\+\\rho\_\{i\-1\},\(36\)where we usedri​ρi=ρi−1r\_\{i\}\\rho\_\{i\}=\\rho\_\{i\-1\}\. ∎

###### Lemma 4\(Exact telescoping\)\.

Ai=φ​\(ρi−1\)/φ​\(ρi\)A\_\{i\}=\\varphi\(\\rho\_\{i\-1\}\)/\\varphi\(\\rho\_\{i\}\), and hence

∏m=j\+1iAm=φ​\(ρj\)φ​\(ρi\),∏m=1iAm=Pφ​\(ρi\)\.\\prod\_\{m=j\+1\}^\{i\}A\_\{m\}=\\frac\{\\varphi\(\\rho\_\{j\}\)\}\{\\varphi\(\\rho\_\{i\}\)\},\\qquad\\prod\_\{m=1\}^\{i\}A\_\{m\}=\\frac\{P\}\{\\varphi\(\\rho\_\{i\}\)\}\.\(37\)

###### Proof\.

By \([32](https://arxiv.org/html/2608.06893#S6.E32)\),Ai=\[ri​φ​\(ρi\)\+\(1−ri\)​P\]/φ​\(ρi\)A\_\{i\}=\\bigl\[r\_\{i\}\\varphi\(\\rho\_\{i\}\)\+\(1\-r\_\{i\}\)P\\bigr\]/\\varphi\(\\rho\_\{i\}\), and the numerator telescopes:

ri​φ​\(ρi\)\+\(1−ri\)​P\\displaystyle r\_\{i\}\\varphi\(\\rho\_\{i\}\)\+\(1\-r\_\{i\}\)P=ri​\[\(1−ρi\)​P\+v​ρi\]\+\(1−ri\)​P\\displaystyle=r\_\{i\}\\bigl\[\(1\-\\rho\_\{i\}\)P\+v\\rho\_\{i\}\\bigr\]\+\(1\-r\_\{i\}\)P=P​\(1−ri​ρi\)\+v​ri​ρi\\displaystyle=P\(1\-r\_\{i\}\\rho\_\{i\}\)\+v\\,r\_\{i\}\\rho\_\{i\}=P​\(1−ρi−1\)\+v​ρi−1\\displaystyle=P\(1\-\\rho\_\{i\-1\}\)\+v\\rho\_\{i\-1\}=φ​\(ρi−1\)\.\\displaystyle=\\varphi\(\\rho\_\{i\-1\}\)\.\(38\)∎

###### Lemma 5\(One\-step identity and deficit recursion\)\.

LetDi:=Σ​\(ρi\)−ViD\_\{i\}:=\\Sigma\(\\rho\_\{i\}\)\-V\_\{i\}be the gap between the exact ancestral chain and the plug\-in chain\. Then

Di−1=Ai2​Di\+\(1−ri\)2​M​\(ρi\),DT=0\.D\_\{i\-1\}=A\_\{i\}^\{2\}D\_\{i\}\+\(1\-r\_\{i\}\)^\{2\}M\(\\rho\_\{i\}\),\\qquad D\_\{T\}=0\.\(39\)Every term is nonnegative and the source term is strictly positive whenv\>0v\>0; this proves Corollary[2](https://arxiv.org/html/2608.06893#Thmcorollary2)\.

###### Proof\.

By the two\-sided Markov property \(S3\) the exact ancestral chain reproduces the true conditional law at every level\. That chain drawsx0∼𝒩​\(⋅,M​\(ρi\)\)x\_\{0\}\\sim\\mathcal\{N\}\(\\cdot,M\(\\rho\_\{i\}\)\)rather than replacing the variable by its conditional mean, so its variance obeys

Σ​\(ρi−1\)=Ai2​Σ​\(ρi\)\+\(1−ri\)2​M​\(ρi\)\+qi\.\\Sigma\(\\rho\_\{i\-1\}\)=A\_\{i\}^\{2\}\\Sigma\(\\rho\_\{i\}\)\+\(1\-r\_\{i\}\)^\{2\}M\(\\rho\_\{i\}\)\+q\_\{i\}\.\(40\)The plug\-in chain obeys the same recursion with the middle term omitted, namely \([35](https://arxiv.org/html/2608.06893#S6.E35)\)\. Subtracting gives \([39](https://arxiv.org/html/2608.06893#S6.E39)\)\. Nonnegativity ofD0D\_\{0\}is immediate by backward induction fromDT=0D\_\{T\}=0\. ∎

###### Lemma 6\(Deficit in closed form\)\.

For every grid and everyT≥1T\\geq 1,

D0=v​P3​∑i=1T\(Δ​ρi\)2ρi​φ​\(ρi\)​φ​\(ρi−1\)2\.D\_\{0\}=vP^\{3\}\\sum\_\{i=1\}^\{T\}\\frac\{\(\\Delta\\rho\_\{i\}\)^\{2\}\}\{\\rho\_\{i\}\\,\\varphi\(\\rho\_\{i\}\)\\,\\varphi\(\\rho\_\{i\-1\}\)^\{2\}\}\.\(41\)ForP=vP=von the uniform grid this reduces toD0=\(P/T\)​ℋTD\_\{0\}=\(P/T\)\\mathcal\{H\}\_\{T\}withℋT\\mathcal\{H\}\_\{T\}theTT\-th harmonic number\.

###### Proof\.

Unroll \([39](https://arxiv.org/html/2608.06893#S6.E39)\) fromDT=0D\_\{T\}=0and use \([37](https://arxiv.org/html/2608.06893#S6.E37)\) for the accumulated contractions, together with the one\-step identity

\(1−ri\)2​M​\(ρi\)=v​P​\(Δ​ρi\)2ρi​φ​\(ρi\)\.\(1\-r\_\{i\}\)^\{2\}M\(\\rho\_\{i\}\)=\\frac\{vP\\,\(\\Delta\\rho\_\{i\}\)^\{2\}\}\{\\rho\_\{i\}\\,\\varphi\(\\rho\_\{i\}\)\}\.\(42\)Each surviving term acquires the factor\(P/φ​\(ρi−1\)\)2\\bigl\(P/\\varphi\(\\rho\_\{i\-1\}\)\\bigr\)^\{2\}, which produces the statedφ​\(ρi−1\)2\\varphi\(\\rho\_\{i\-1\}\)^\{2\}in the denominator and the overallP3P^\{3\}\. WhenP=vP=vwe haveφ≡P\\varphi\\equiv P, so the sum collapses to∑i\(Δ​ρi\)2/ρi=\(1/T\)​∑i≤T1/i\\sum\_\{i\}\(\\Delta\\rho\_\{i\}\)^\{2\}/\\rho\_\{i\}=\(1/T\)\\sum\_\{i\\leq T\}1/i\. ∎

#### Proof of Theorem[2](https://arxiv.org/html/2608.06893#Thmtheorem2)

###### Proof\.

\(a\)v\>0v\>0\.Putφmin=min⁡\(P,v\)\>0\\varphi\_\{\\min\}=\\min\(P,v\)\>0,φmax=max⁡\(P,v\)\\varphi\_\{\\max\}=\\max\(P,v\)and let

ΘT:=∑i=1T\(Δ​ρi\)2ρi\\Theta\_\{T\}:=\\sum\_\{i=1\}^\{T\}\\frac\{\(\\Delta\\rho\_\{i\}\)^\{2\}\}\{\\rho\_\{i\}\}\(43\)be the purely geometric part of \([41](https://arxiv.org/html/2608.06893#S6.E41)\)\. Sinceφmin≤φ​\(ρ\)≤φmax\\varphi\_\{\\min\}\\leq\\varphi\(\\rho\)\\leq\\varphi\_\{\\max\}on\[0,1\]\[0,1\],

v​P3φmax3​ΘT≤D0≤v​P3φmin3​ΘT\.\\frac\{vP^\{3\}\}\{\\varphi\_\{\\max\}^\{3\}\}\\,\\Theta\_\{T\}\\;\\leq\\;D\_\{0\}\\;\\leq\\;\\frac\{vP^\{3\}\}\{\\varphi\_\{\\min\}^\{3\}\}\\,\\Theta\_\{T\}\.\(44\)It therefore suffices to showΘT→0\\Theta\_\{T\}\\to 0\.

*Uniform grid\.*ΘT=\(1/T\)​ℋT=\(log⁡T\)/T\+O​\(1/T\)\\Theta\_\{T\}=\(1/T\)\\mathcal\{H\}\_\{T\}=\(\\log T\)/T\+O\(1/T\)\.

*Power gridρi=\(i/T\)a\\rho\_\{i\}=\(i/T\)^\{a\},a\>1a\>1\.*The mean value theorem givesΔ​ρi≤a​ia−1/Ta\\Delta\\rho\_\{i\}\\leq a\\,i^\{a\-1\}/T^\{a\}, hence

\(Δ​ρi\)2ρi≤a2​ia−2Ta\.\\frac\{\(\\Delta\\rho\_\{i\}\)^\{2\}\}\{\\rho\_\{i\}\}\\leq\\frac\{a^\{2\}\\,i^\{a\-2\}\}\{T^\{a\}\}\.\(45\)The summand is*not*uniformly of orderT−2T^\{\-2\}: thei=1i=1term isT−aT^\{\-a\}\. Nowu↦ua−2u\\mapsto u^\{a\-2\}is monotone on\[1,∞\)\[1,\\infty\)\(nonincreasing for1<a≤21<a\\leq 2, nondecreasing fora≥2a\\geq 2\), so comparing the sum with the integral term by term — adding the first term separately in the nonincreasing case — gives

∑i=1Tia−2\\displaystyle\\sum\_\{i=1\}^\{T\}i^\{a\-2\}≤1\+∫1T\+1ua−2​du\\displaystyle\\leq 1\+\\int\_\{1\}^\{T\+1\}u^\{a\-2\}\\,\\mathrm\{d\}u≤\(1\+2a−1a−1\)​Ta−1,\\displaystyle\\leq\\Bigl\(1\+\\tfrac\{2^\{a\-1\}\}\{a\-1\}\\Bigr\)T^\{a\-1\},\(46\)where we used\(T\+1\)a−1≤\(2​T\)a−1\(T\+1\)^\{a\-1\}\\leq\(2T\)^\{a\-1\}and1≤Ta−11\\leq T^\{a\-1\}\. Therefore

ΘT≤a2​\(1\+2a−1a−1\)​1T=O​\(1/T\)\.\\Theta\_\{T\}\\leq a^\{2\}\\Bigl\(1\+\\tfrac\{2^\{a\-1\}\}\{a\-1\}\\Bigr\)\\frac\{1\}\{T\}=O\(1/T\)\.\(47\)
In both casesD0→0D\_\{0\}\\to 0by \([44](https://arxiv.org/html/2608.06893#S6.E44)\), soV0→PV\_\{0\}\\to P\. Lemma[3](https://arxiv.org/html/2608.06893#Thmlemma3)shows the mean is already exact, so the terminal law converges to𝒩​\(W​x1,P\)\\mathcal\{N\}\(Wx\_\{1\},P\)\. By \(S6\) the terminal KL is

12​\(V0P−1−log⁡V0P\)=D024​P2\+O​\(\(D0/P\)3\),\\tfrac\{1\}\{2\}\\Bigl\(\\tfrac\{V\_\{0\}\}\{P\}\-1\-\\log\\tfrac\{V\_\{0\}\}\{P\}\\Bigr\)=\\frac\{D\_\{0\}^\{2\}\}\{4P^\{2\}\}\+O\\bigl\(\(D\_\{0\}/P\)^\{3\}\\bigr\),\(48\)and the limiting law depends on neithervvnor the schedule\.

\(b\)v=0v=0\.Allqi=0q\_\{i\}=0, soVi≡0V\_\{i\}\\equiv 0and the terminal law is the point mass atc0​x1=W​x1c\_\{0\}x\_\{1\}=Wx\_\{1\}; Lemma[3](https://arxiv.org/html/2608.06893#Thmlemma3)remains valid\. The first step fromρ=1\\rho=1is defined directly byx^0=W​x1\\hat\{x\}\_\{0\}=Wx\_\{1\}, because the scaled deviation is identically zero\. HenceKL​\(δ​\(W​x1\)∥𝒩​\(W​x1,P\)\)=\+∞\\mathrm\{KL\}\\bigl\(\\delta\(Wx\_\{1\}\)\\,\\\|\\,\\mathcal\{N\}\(Wx\_\{1\},P\)\\bigr\)=\+\\infty\. The map fromvvto the limiting law is constant onv\>0v\>0and jumps atv=0v=0\. ∎

#### Why resource\-free objectives are degenerate

The integrated\-Bayes\-risk objective

J=∑k∫w​\(t\)​MMSEk​\(t\)​dtJ=\\sum\_\{k\}\\int w\(t\)\\,\\mathrm\{MMSE\}\_\{k\}\(t\)\\,\\mathrm\{d\}t\(49\)is concave and increasing in everyvkv\_\{k\}\. Under a budget the matched point does satisfy the Lagrange condition — but it is the*maximum*; the minima sit at vertex allocations\. An explicit counterexample isρ=0\.5\\rho=0\.5,P=\(1,4\)P=\(1,4\)and budget55: the matched value is2\.52\.5, whereas a vertex allocation gives0\.8330\.833\.

Minimizing target Bayes risk rewards a bridge state that is maximally informative aboutx0x\_\{0\}, and without a resource constraint the optimum isv=0v=0\. That choice makes the interpolation deterministic, reduces the method to one\-shot Wiener regression, and destroys distributional correctness because the terminal variance tends to0\. An objective whose optimum lies on thev=0v=0boundary is unsuitable for a*bridge*, whose purpose is distributional\.

### 6\.4Proofs for Section[3\.3](https://arxiv.org/html/2608.06893#S3.SS3)\(Finite\-Step\)

#### Proof of Theorem[3](https://arxiv.org/html/2608.06893#Thmtheorem3)

###### Proof\.

Step 1 \(scale symmetry\)\.Writeφ​\(ρ\)=P​ψ​\(ρ\)\\varphi\(\\rho\)=P\\psi\(\\rho\)withψ​\(ρ\)=1\+\(x−1\)​ρ\\psi\(\\rho\)=1\+\(x\-1\)\\rhoandx=v/Px=v/P\. In \([41](https://arxiv.org/html/2608.06893#S6.E41)\) the factorP3P^\{3\}cancels the three factors ofPPcontributed by theφ\\varphiterms, and the remaining factor isv=x​Pv=xP\. Thusxxsurvives in the dimensionless expression whilePPnormalizesD0D\_\{0\}:

D0P=x​G​\(x\),G​\(x\)=∑i\(Δ​ρi\)2ρi​ψ​\(ρi\)​ψ​\(ρi−1\)2\.\\frac\{D\_\{0\}\}\{P\}=x\\,G\(x\),\\qquad G\(x\)=\\sum\_\{i\}\\frac\{\(\\Delta\\rho\_\{i\}\)^\{2\}\}\{\\rho\_\{i\}\\,\\psi\(\\rho\_\{i\}\)\\,\\psi\(\\rho\_\{i\-1\}\)^\{2\}\}\.\(50\)HenceKL=Φ​\(x\)\\mathrm\{KL\}=\\Phi\(x\)depends on the reference only throughxx, and otherwise only on the grid\. Structurally,PPis the only scale in the per\-mode Gaussian problem andv/Pv/Pis its only dimensionless parameter\.

Step 2 \(form of every optimizer\)\.By Step 1,∑kKLk=∑kΦ​\(vk/Pk\)\\sum\_\{k\}\\mathrm\{KL\}\_\{k\}=\\sum\_\{k\}\\Phi\(v\_\{k\}/P\_\{k\}\)with the*same*functionΦ\\Phiin every mode, so the objective is separable\. At any*local*minimizer each coordinatexk=vk/Pkx\_\{k\}=v\_\{k\}/P\_\{k\}must be a local minimizer ofΦ\\Phi; different coordinates may, however, select different local minima, so a common mode\-independentxxis not forced\. At any*global*minimizerxk∈arg​min⁡Φx\_\{k\}\\in\\operatorname\*\{arg\\,min\}\\Phifor everykk\. Ifarg​min⁡Φ\\operatorname\*\{arg\\,min\}\\Phiis the singleton\{x∗\}\\\{x^\{\*\}\\\}thenvk∗=x∗​Pkv^\{\*\}\_\{k\}=x^\{\*\}P\_\{k\}for every mode, withx∗x^\{\*\}mode\-independent\. IfΦ\\Phihas several local minimizers — which does occur on general grids, by Proposition[2](https://arxiv.org/html/2608.06893#Thmproposition2)— then assigning different modes to different valleys produces genuinely non\-proportional*local*optimizers of the sum\.

Step 3 \(barriers, hence an interior optimum\)\.Asx→0x\\to 0the observation\-end term \(i=Ti\{=\}T\) ofx​G​\(x\)xG\(x\)tends to11: the sampler cannot create posterior variance without reference noise, exactly as in Theorem[2](https://arxiv.org/html/2608.06893#Thmtheorem2)\(b\)\. Henceu→0u\\to 0andΦ→∞\\Phi\\to\\infty\. Asx→∞x\\to\\inftythe clean\-end term \(i=1i\{=\}1\) ofx​G​\(x\)xG\(x\)tends to11: excess noise entering the final Bayes step cannot be fully removed, so againΦ→∞\\Phi\\to\\infty\. SinceΦ\\Phiis continuous on\(0,∞\)\(0,\\infty\)it attains an interior minimum\. Both barriers are visible in the left panel of Figure[5](https://arxiv.org/html/2608.06893#S6.F5)\. ∎

#### Proof of Proposition[1](https://arxiv.org/html/2608.06893#Thmproposition1): the constantx∗​\(T\)x^\{\*\}\(T\)

Two conventions are in force in this subsection\. First, the window center denotedx0x\_\{0\}in the statement of Proposition[1](https://arxiv.org/html/2608.06893#Thmproposition1)is the quantityχT=\(2​ln⁡T\)−1/2\\chi\_\{T\}=\(2\\ln T\)^\{\-1/2\}here; the symbolx0x\_\{0\}is reserved for the clean signal\. Second,*minimizer*means global minimizer ofΦ\\Phion\(0,∞\)\(0,\\infty\); Proposition[1](https://arxiv.org/html/2608.06893#Thmproposition1)is proved in this sense\.

Throughout this subsection the grid is uniform,ρi=i/T\\rho\_\{i\}=i/T, and we write

H​\(x\)\\displaystyle H\(x\):=D0P=x​𝒮​\(x\)T,\\displaystyle:=\\frac\{D\_\{0\}\}\{P\}=\\frac\{x\\,\\mathcal\{S\}\(x\)\}\{T\},\(51\)𝒮​\(x\)\\displaystyle\\mathcal\{S\}\(x\):=1T​∑i=1T1ρi​ψi​ψi−12,\\displaystyle:=\\frac\{1\}\{T\}\\sum\_\{i=1\}^\{T\}\\frac\{1\}\{\\rho\_\{i\}\\,\\psi\_\{i\}\\,\\psi\_\{i\-1\}^\{2\}\},\(52\)ψi\\displaystyle\\psi\_\{i\}:=ψ​\(i/T\)=1−b​iT,b:=1−x\.\\displaystyle:=\\psi\(i/T\)=1\-b\\,\\tfrac\{i\}\{T\},\\qquad b:=1\-x\.\(53\)We also set

f​\(ρ\)\\displaystyle f\(\\rho\):=1ρ​ψ​\(ρ\)3,\\displaystyle:=\\frac\{1\}\{\\rho\\,\\psi\(\\rho\)^\{3\}\},𝒮∘​\(x\)\\displaystyle\\mathcal\{S\}^\{\\circ\}\(x\):=1T​∑i=1Tf​\(ρi\),\\displaystyle:=\\frac\{1\}\{T\}\\sum\_\{i=1\}^\{T\}f\(\\rho\_\{i\}\),\(54\)h​\(x\)\\displaystyle h\(x\):=x​ln⁡T\+12​x,\\displaystyle:=x\\ln T\+\\frac\{1\}\{2x\},χT\\displaystyle\\chi\_\{T\}:=\(2​ln⁡T\)−1/2,\\displaystyle:=\(2\\ln T\)^\{\-1/2\},\(55\)and𝒲T:=\[χT/3,3​χT\]\\mathcal\{W\}\_\{T\}:=\[\\chi\_\{T\}/3,\\,3\\chi\_\{T\}\]\. The comparison functionhhis strictly convex on\(0,∞\)\(0,\\infty\)with unique minimizerχT\\chi\_\{T\}, minimum valueh​\(χT\)=2​ln⁡Th\(\\chi\_\{T\}\)=\\sqrt\{2\\ln T\}andh′′​\(x\)=x−3h^\{\\prime\\prime\}\(x\)=x^\{\-3\}\. Finally,Φ​\(x\)=12​\(u−1−ln⁡u\)\\Phi\(x\)=\\tfrac\{1\}\{2\}\(u\-1\-\\ln u\)withu=1−H​\(x\)∈\(0,1\)u=1\-H\(x\)\\in\(0,1\)is strictly decreasing inuu, so*minimizingΦ\\Phiis equivalent to minimizingHH*\. We work throughout with the scaled quantityT​HT\\,H\.

The proof comparesT​HT\\,Hwithhhin four steps: replace the shifted grid factor \(Lemma[7](https://arxiv.org/html/2608.06893#Thmlemma7)\), replace the sum by an integral \(Lemma[8](https://arxiv.org/html/2608.06893#Thmlemma8)\), evaluate that integral in closed form \(Lemma[9](https://arxiv.org/html/2608.06893#Thmlemma9)\), and combine \(Lemma[10](https://arxiv.org/html/2608.06893#Thmlemma10)\)\. Lemma[11](https://arxiv.org/html/2608.06893#Thmlemma11)then confines the minimizer to𝒲T\\mathcal\{W\}\_\{T\}, where the comparison is uniform\.

![Refer to caption](https://arxiv.org/html/2608.06893v1/x5.png)Figure 5:The landscape ofT​H​\(x\)T\\,H\(x\), and why the minimizer is localized\.*Left:*the two barriers of Theorem[3](https://arxiv.org/html/2608.06893#Thmtheorem3), Step 3 —Φ→∞\\Phi\\to\\inftyat both ends — with the exact minimizer marked for eachTT\.*Right:*atT=104T=10^\{4\}, the exactT​HT\\,Hagainst the comparison functionh​\(x\)=x​ln⁡T\+12​xh\(x\)=x\\ln T\+\\tfrac\{1\}\{2x\}of Lemma[10](https://arxiv.org/html/2608.06893#Thmlemma10); the additive gap is1\+o​\(1\)1\+o\(1\), the shaded band is the window𝒲T=\[χT/3,3​χT\]\\mathcal\{W\}\_\{T\}=\[\\chi\_\{T\}/3,3\\chi\_\{T\}\]of Lemma[11](https://arxiv.org/html/2608.06893#Thmlemma11), andx∗x^\{\*\}has already almost reachedχT=\(2​ln⁡T\)−1/2\\chi\_\{T\}=\(2\\ln T\)^\{\-1/2\}\. Computed from \([41](https://arxiv.org/html/2608.06893#S6.E41)\); no simulation\.###### Lemma 7\(Grid\-shift replacement\)\.

For0<x<10<x<1andT≥2T\\geq 2,

\(1\+1T​x\)−2​𝒮∘​\(x\)≤𝒮​\(x\)≤𝒮∘​\(x\)\.\\Bigl\(1\+\\tfrac\{1\}\{Tx\}\\Bigr\)^\{\-2\}\\mathcal\{S\}^\{\\circ\}\(x\)\\;\\leq\\;\\mathcal\{S\}\(x\)\\;\\leq\\;\\mathcal\{S\}^\{\\circ\}\(x\)\.\(56\)

###### Proof\.

ψi−1=ψi\+b/T\\psi\_\{i\-1\}=\\psi\_\{i\}\+b/Twith0<b<10<b<1, andψi≥ψT=x\\psi\_\{i\}\\geq\\psi\_\{T\}=x, soψi≤ψi−1≤ψi​\(1\+1T​x\)\\psi\_\{i\}\\leq\\psi\_\{i\-1\}\\leq\\psi\_\{i\}\\bigl\(1\+\\tfrac\{1\}\{Tx\}\\bigr\)\. Square and invert termwise\. ∎

###### Lemma 8\(Riemann comparison\)\.

For0<x≤340<x\\leq\\tfrac\{3\}\{4\}andT≥8T\\geq 8,

\|𝒮∘​\(x\)−1−∫1/T1f​\(ρ\)​dρ\|≤1\+12T\+1T​x3\.\\Bigl\|\\mathcal\{S\}^\{\\circ\}\(x\)\-1\-\\int\_\{1/T\}^\{1\}f\(\\rho\)\\,\\mathrm\{d\}\\rho\\Bigr\|\\;\\leq\\;1\+\\frac\{12\}\{T\}\+\\frac\{1\}\{Tx^\{3\}\}\.\(57\)

###### Proof\.

The logarithmic derivative\(ln⁡f\)′​\(ρ\)=−1/ρ\+3​b/ψ​\(ρ\)\(\\ln f\)^\{\\prime\}\(\\rho\)=\-1/\\rho\+3b/\\psi\(\\rho\)vanishes only atρ∗=1/\(4​b\)≤1\\rho^\{\*\}=1/\(4b\)\\leq 1, usingb≥14b\\geq\\tfrac\{1\}\{4\}\. Henceffstrictly decreases on\(0,ρ∗\]\(0,\\rho^\{\*\}\]and strictly increases on\[ρ∗,1\]\[\\rho^\{\*\},1\]\.

Separate the initial term first\. ForT≥8T\\geq 8,

1T​f​\(1/T\)=ψ​\(1/T\)−3∈\[1,1\+6T\],\\tfrac\{1\}\{T\}f\(1/T\)=\\psi\(1/T\)^\{\-3\}\\in\\bigl\[1,\\,1\+\\tfrac\{6\}\{T\}\\bigr\],\(58\)becauseψ​\(1/T\)≥1−1/T\\psi\(1/T\)\\geq 1\-1/Tand\(1−u\)−3≤1\+6​u\(1\-u\)^\{\-3\}\\leq 1\+6uon\[0,18\]\[0,\\tfrac\{1\}\{8\}\]\. For each remaining interval,\|1T​f​\(ρi\)−∫ρi−1ρif\|≤1T​osc\[ρi−1,ρi\]⁡f\\bigl\|\\tfrac\{1\}\{T\}f\(\\rho\_\{i\}\)\-\\int\_\{\\rho\_\{i\-1\}\}^\{\\rho\_\{i\}\}f\\bigr\|\\leq\\tfrac\{1\}\{T\}\\operatorname\{osc\}\_\{\[\\rho\_\{i\-1\},\\rho\_\{i\}\]\}f, and by \(S4\) the sum of those oscillations is at mostTV⁡\(f\)\\operatorname\{TV\}\(f\)\. Becausefffirst decreases and then increases,TV⁡\(f\)≤f​\(1/T\)\+f​\(1\)\\operatorname\{TV\}\(f\)\\leq f\(1/T\)\+f\(1\)\. Therefore

\|∑i=2T1T​f​\(ρi\)−∫1/T1f\|\\displaystyle\\Bigl\|\\sum\_\{i=2\}^\{T\}\\tfrac\{1\}\{T\}f\(\\rho\_\{i\}\)\-\\int\_\{1/T\}^\{1\}f\\Bigr\|≤1T​\(f​\(1/T\)\+f​\(1\)\)\\displaystyle\\leq\\tfrac\{1\}\{T\}\\bigl\(f\(1/T\)\+f\(1\)\\bigr\)≤ψ​\(1/T\)−3\+1T​x3\\displaystyle\\leq\\psi\(1/T\)^\{\-3\}\+\\tfrac\{1\}\{Tx^\{3\}\}≤1\+6T\+1T​x3\.\\displaystyle\\leq 1\+\\tfrac\{6\}\{T\}\+\\tfrac\{1\}\{Tx^\{3\}\}\.\(59\)Combining the two displays gives the claim\. ∎

###### Lemma 9\(The integral in closed form\)\.

For0<x<10<x<1andT≥3T\\geq 3,

x​∫1/T1f​\(ρ\)​dρ=x​ln⁡T\+12​x\+1\+x​ln⁡1x−3​x2\+θ,x\\int\_\{1/T\}^\{1\}f\(\\rho\)\\,\\mathrm\{d\}\\rho=x\\ln T\+\\frac\{1\}\{2x\}\+1\+x\\ln\\frac\{1\}\{x\}\-\\frac\{3x\}\{2\}\+\\theta,\(60\)with\|θ\|≤6​x/T\|\\theta\|\\leq 6x/T\.

###### Proof\.

Clearing denominators in\(1−u\)​\(1\+u\+u2\)=1−u3\(1\-u\)\(1\+u\+u^\{2\}\)=1\-u^\{3\}withu=1−b​ρu=1\-b\\rhogives the partial\-fraction identity

1ρ​ψ3=1ρ\+bψ\+bψ2\+bψ3,\\frac\{1\}\{\\rho\\psi^\{3\}\}=\\frac\{1\}\{\\rho\}\+\\frac\{b\}\{\\psi\}\+\\frac\{b\}\{\\psi^\{2\}\}\+\\frac\{b\}\{\\psi^\{3\}\},\(61\)whose antiderivative is

F​\(ρ\)=ln⁡ρψ​\(ρ\)\+1ψ​\(ρ\)\+12​ψ​\(ρ\)2,F′=f\.F\(\\rho\)=\\ln\\frac\{\\rho\}\{\\psi\(\\rho\)\}\+\\frac\{1\}\{\\psi\(\\rho\)\}\+\\frac\{1\}\{2\\psi\(\\rho\)^\{2\}\},\\qquad F^\{\\prime\}=f\.\(62\)Hence∫1/T1f=F​\(1\)−F​\(1/T\)\\int\_\{1/T\}^\{1\}f=F\(1\)\-F\(1/T\)withF​\(1\)=ln⁡1x\+1x\+12​x2F\(1\)=\\ln\\tfrac\{1\}\{x\}\+\\tfrac\{1\}\{x\}\+\\tfrac\{1\}\{2x^\{2\}\}\. Alsoψ​\(1/T\)=1−b/T∈\[1−1T,1\]\\psi\(1/T\)=1\-b/T\\in\[1\-\\tfrac\{1\}\{T\},1\], so

F​\(1/T\)\\displaystyle F\(1/T\)=−ln⁡T−ln⁡ψ​\(1/T\)\+1ψ​\(1/T\)\+12​ψ​\(1/T\)2\\displaystyle=\-\\ln T\-\\ln\\psi\(1/T\)\+\\frac\{1\}\{\\psi\(1/T\)\}\+\\frac\{1\}\{2\\psi\(1/T\)^\{2\}\}=−ln⁡T\+32\+O​\(1T\)\.\\displaystyle=\-\\ln T\+\\tfrac\{3\}\{2\}\+O\\bigl\(\\tfrac\{1\}\{T\}\\bigr\)\.\(63\)ForT≥3T\\geq 3that remainder is at most6/T6/Tin absolute value\. Multiplying byxxcompletes the proof\. \(AtT=2T=2the constant66fails narrowly asx→0x\\to 0; this is harmless, because Lemma[10](https://arxiv.org/html/2608.06893#Thmlemma10)is used only for largeTTand Proposition[2](https://arxiv.org/html/2608.06893#Thmproposition2)handlesT=2T=2exactly\.\) ∎

###### Lemma 10\(Uniform expansion on the window\)\.

LetδT:=supx∈𝒲T\|T​H​\(x\)−h​\(x\)−1\|\\delta\_\{T\}:=\\sup\_\{x\\in\\mathcal\{W\}\_\{T\}\}\\bigl\|T\\,H\(x\)\-h\(x\)\-1\\bigr\|\. ThenδT→0\\delta\_\{T\}\\to 0; explicitly

δT=O​\(χT​ln⁡1χT\+ln⁡TT\)\.\\delta\_\{T\}=O\\Bigl\(\\chi\_\{T\}\\ln\\tfrac\{1\}\{\\chi\_\{T\}\}\+\\tfrac\{\\ln T\}\{T\}\\Bigr\)\.\(64\)

###### Proof\.

TakeTTlarge enough that3​χT≤343\\chi\_\{T\}\\leq\\tfrac\{3\}\{4\}, so that Lemmas[7](https://arxiv.org/html/2608.06893#Thmlemma7)–[9](https://arxiv.org/html/2608.06893#Thmlemma9)all apply on𝒲T\\mathcal\{W\}\_\{T\}\. Chaining the three lemmas,

T​H​\(x\)\\displaystyle T\\,H\(x\)=x​𝒮​\(x\)=x​𝒮∘​\(x\)\+E1\\displaystyle=x\\,\\mathcal\{S\}\(x\)=x\\,\\mathcal\{S\}^\{\\circ\}\(x\)\+E\_\{1\}=x\+x​∫1/T1f\+x​E2\+E1\\displaystyle=x\+x\\\!\\int\_\{1/T\}^\{1\}\\\!f\+xE\_\{2\}\+E\_\{1\}=h​\(x\)\+1\+r​\(T,x\),\\displaystyle=h\(x\)\+1\+r\(T,x\),\(65\)where the remainder collects the six small terms,

r​\(T,x\):=x\+x​ln⁡1x−3​x2\+θ\+x​E2\+E1\.r\(T,x\):=x\+x\\ln\\tfrac\{1\}\{x\}\-\\tfrac\{3x\}\{2\}\+\\theta\+xE\_\{2\}\+E\_\{1\}\.\(66\)Here\|E1\|≤2T​x⋅x​𝒮∘=2​𝒮∘T\|E\_\{1\}\|\\leq\\tfrac\{2\}\{Tx\}\\cdot x\\mathcal\{S\}^\{\\circ\}=\\tfrac\{2\\mathcal\{S\}^\{\\circ\}\}\{T\}by Lemma[7](https://arxiv.org/html/2608.06893#Thmlemma7),\|E2\|≤1\+12T\+1T​x3\|E\_\{2\}\|\\leq 1\+\\tfrac\{12\}\{T\}\+\\tfrac\{1\}\{Tx^\{3\}\}by Lemma[8](https://arxiv.org/html/2608.06893#Thmlemma8), and\|θ\|≤6​xT\|\\theta\|\\leq\\tfrac\{6x\}\{T\}by Lemma[9](https://arxiv.org/html/2608.06893#Thmlemma9)\.

We check each contribution uniformly on𝒲T\\mathcal\{W\}\_\{T\}\. Firstx≤3​χT→0x\\leq 3\\chi\_\{T\}\\to 0andx​ln⁡1x≤3​χT​ln⁡3χT→0x\\ln\\tfrac\{1\}\{x\}\\leq 3\\chi\_\{T\}\\ln\\tfrac\{3\}\{\\chi\_\{T\}\}\\to 0\. Nextx⋅1T​x3=1T​x2≤18​ln⁡TT→0x\\cdot\\tfrac\{1\}\{Tx^\{3\}\}=\\tfrac\{1\}\{Tx^\{2\}\}\\leq\\tfrac\{18\\ln T\}\{T\}\\to 0\. Finally Lemmas[8](https://arxiv.org/html/2608.06893#Thmlemma8)–[9](https://arxiv.org/html/2608.06893#Thmlemma9)together with12​x2≤9​ln⁡T\\tfrac\{1\}\{2x^\{2\}\}\\leq 9\\ln Tgive𝒮∘=O​\(ln⁡T\)\\mathcal\{S\}^\{\\circ\}=O\(\\ln T\)on the window, so2​𝒮∘T=O​\(ln⁡TT\)→0\\tfrac\{2\\mathcal\{S\}^\{\\circ\}\}\{T\}=O\(\\tfrac\{\\ln T\}\{T\}\)\\to 0\. Every contribution torrtherefore vanishes uniformly on𝒲T\\mathcal\{W\}\_\{T\}\. ∎

###### Lemma 11\(No minimizer outside the window\)\.

There is aT0T\_\{0\}such that for everyT≥T0T\\geq T\_\{0\}every global minimizer ofHHon\(0,∞\)\(0,\\infty\)lies in𝒲T=\[χT/3,3​χT\]\\mathcal\{W\}\_\{T\}=\[\\chi\_\{T\}/3,\\,3\\chi\_\{T\}\]\.

###### Proof\.

By Lemma[10](https://arxiv.org/html/2608.06893#Thmlemma10),minx∈𝒲T⁡T​H​\(x\)≤ΛT\\min\_\{x\\in\\mathcal\{W\}\_\{T\}\}T\\,H\(x\)\\leq\\Lambda\_\{T\}where

ΛT:=2​ln⁡T\+1\+δT\.\\Lambda\_\{T\}:=\\sqrt\{2\\ln T\}\+1\+\\delta\_\{T\}\.\(67\)We showT​H​\(x\)\>ΛTT\\,H\(x\)\>\\Lambda\_\{T\}outside𝒲T\\mathcal\{W\}\_\{T\}for all largeTT\.

*Two lower bounds\.*For0<x≤120<x\\leq\\tfrac\{1\}\{2\}andT≥8T\\geq 8, keep only the terms withi\>T/2i\>T/2\. Sinceρi≤1\\rho\_\{i\}\\leq 1,ψi≤ψi−1\\psi\_\{i\}\\leq\\psi\_\{i\-1\}andψ−3\\psi^\{\-3\}is increasing there,

T​H​\(x\)\\displaystyle T\\,H\(x\)≥x​∑i\>T/21/Tψi−13≥x​∫1/21−1/Td​ρψ​\(ρ\)3\\displaystyle\\geq x\\sum\_\{i\>T/2\}\\frac\{1/T\}\{\\psi\_\{i\-1\}^\{3\}\}\\;\\geq\\;x\\int\_\{1/2\}^\{1\-1/T\}\\frac\{\\,\\mathrm\{d\}\\rho\}\{\\psi\(\\rho\)^\{3\}\}=x2​b​\[1ψ​\(1−1T\)2−1ψ​\(12\)2\]\\displaystyle=\\frac\{x\}\{2b\}\\left\[\\frac\{1\}\{\\psi\(1\-\\tfrac\{1\}\{T\}\)^\{2\}\}\-\\frac\{1\}\{\\psi\(\\tfrac\{1\}\{2\}\)^\{2\}\}\\right\]≥x2​\[1\(x\+1T\)2−4\],\\displaystyle\\geq\\frac\{x\}\{2\}\\left\[\\frac\{1\}\{\(x\+\\tfrac\{1\}\{T\}\)^\{2\}\}\-4\\right\],\(68\)where we usedb≤1b\\leq 1,ψ​\(1−1T\)=x\+bT≤x\+1T\\psi\(1\-\\tfrac\{1\}\{T\}\)=x\+\\tfrac\{b\}\{T\}\\leq x\+\\tfrac\{1\}\{T\}andψ​\(12\)≥12\\psi\(\\tfrac\{1\}\{2\}\)\\geq\\tfrac\{1\}\{2\}\. Separately, for0<x≤10<x\\leq 1thei=Ti=Tterm alone gives

T​H​\(x\)≥1T​\(x\+1T\)2\.T\\,H\(x\)\\geq\\frac\{1\}\{T\\bigl\(x\+\\tfrac\{1\}\{T\}\\bigr\)^\{2\}\}\.\(69\)
*The six regions\.*Table[5](https://arxiv.org/html/2608.06893#S6.T5)lists them; empty regions may be ignored\. Region \(i\) uses \([69](https://arxiv.org/html/2608.06893#S6.E69)\) withx\+1T≤2Tx\+\\tfrac\{1\}\{T\}\\leq\\tfrac\{2\}\{T\}\. Regions \(ii\)–\(iii\) use \([68](https://arxiv.org/html/2608.06893#S6.E68)\) with, respectively,x\+T−1≤2​xx\+T^\{\-1\}\\leq 2xandx\+T−1≤x​\(1\+T−1/2\)x\+T^\{\-1\}\\leq x\(1\+T^\{\-1/2\}\)\. Region \(iv\) usesψ≤1\\psi\\leq 1on\[0,1\]\[0,1\], soT​H≥x​∑i1/i≥x​ln⁡TT\\,H\\geq x\\sum\_\{i\}1/i\\geq x\\ln T\. Region \(v\) usesψi−1≤ψi≤32\\psi\_\{i\-1\}\\leq\\psi\_\{i\}\\leq\\tfrac\{3\}\{2\}fori≤T/\(2​x\)i\\leq T/\(2x\), givingT​H≥827​x​ln⁡T2​xT\\,H\\geq\\tfrac\{8\}\{27\}x\\ln\\tfrac\{T\}\{2x\}and hence881​ln⁡T\\tfrac\{8\}\{81\}\\ln TforT≥64T\\geq 64\. Region \(vi\) uses thei=1i=1term,T​H≥x/ψ​\(1/T\)≥T​x/\(T\+x\)T\\,H\\geq x/\\psi\(1/T\)\\geq Tx/\(T\+x\)\.

Comparing the last column of Table[5](https://arxiv.org/html/2608.06893#S6.T5)withΛT=2​ln⁡T​\(1\+o​\(1\)\)\\Lambda\_\{T\}=\\sqrt\{2\\ln T\}\(1\+o\(1\)\): in region \(iii\) the leading coefficient is32​2\>2\\tfrac\{3\}\{2\}\\sqrt\{2\}\>\\sqrt\{2\}, in region \(iv\) it is32\>2\\tfrac\{3\}\{\\sqrt\{2\}\}\>\\sqrt\{2\}, and in the remaining regionsTT,T\\sqrt\{T\}orln⁡T\\ln Tgrows faster thanln⁡T\\sqrt\{\\ln T\}\. Hence no global minimizer lies outside𝒲T\\mathcal\{W\}\_\{T\}onceTTis large enough\. ∎

Table 5:The six regions of Lemma[11](https://arxiv.org/html/2608.06893#Thmlemma11)\.Each row gives a lower bound onT​HT\\,H, obtained from the tool named in the proof, that already exceeds the benchmarkΛT=2​ln⁡T​\(1\+o​\(1\)\)\\Lambda\_\{T\}=\\sqrt\{2\\ln T\}\(1\+o\(1\)\)attained inside the window\. Only regions \(iii\) and \(iv\) are decided by a constant, and in both the constant is\>2\>\\sqrt\{2\}\.
###### Proof of Proposition[1](https://arxiv.org/html/2608.06893#Thmproposition1)\.

Theorem[3](https://arxiv.org/html/2608.06893#Thmtheorem3)guarantees a global minimizerx∗x^\{\*\}, and forT≥T0T\\geq T\_\{0\}Lemma[11](https://arxiv.org/html/2608.06893#Thmlemma11)places it in𝒲T\\mathcal\{W\}\_\{T\}\. On𝒲T\\mathcal\{W\}\_\{T\}we haveh′′≥\(3​χT\)−3h^\{\\prime\\prime\}\\geq\(3\\chi\_\{T\}\)^\{\-3\}, hence the quadratic lower bound

h​\(x\)≥h​\(χT\)\+\(x−χT\)254​χT3\.h\(x\)\\geq h\(\\chi\_\{T\}\)\+\\frac\{\(x\-\\chi\_\{T\}\)^\{2\}\}\{54\\,\\chi\_\{T\}^\{3\}\}\.\(70\)Lemma[10](https://arxiv.org/html/2608.06893#Thmlemma10)gives bothT​H​\(x∗\)≤T​H​\(χT\)≤h​\(χT\)\+1\+δTT\\,H\(x^\{\*\}\)\\leq T\\,H\(\\chi\_\{T\}\)\\leq h\(\\chi\_\{T\}\)\+1\+\\delta\_\{T\}andT​H​\(x∗\)≥h​\(x∗\)\+1−δTT\\,H\(x^\{\*\}\)\\geq h\(x^\{\*\}\)\+1\-\\delta\_\{T\}\. Subtracting,

\(x∗−χT\)254​χT3≤2​δT⟹\|x∗χT−1\|≤108​δT​χT⟶0\.\\frac\{\(x^\{\*\}\-\\chi\_\{T\}\)^\{2\}\}\{54\\,\\chi\_\{T\}^\{3\}\}\\leq 2\\delta\_\{T\}\\quad\\Longrightarrow\\quad\\Bigl\|\\frac\{x^\{\*\}\}\{\\chi\_\{T\}\}\-1\\Bigr\|\\leq\\sqrt\{108\\,\\delta\_\{T\}\\,\\chi\_\{T\}\}\\longrightarrow 0\.\(71\)Thereforex∗​\(T\)=\(2​ln⁡T\)−1/2​\(1\+o​\(1\)\)x^\{\*\}\(T\)=\(2\\ln T\)^\{\-1/2\}\(1\+o\(1\)\), which is the claim\.

For the optimal value,H​\(x∗\)=1T​\(2​ln⁡T\+1\+O​\(δT\)\)H\(x^\{\*\}\)=\\tfrac\{1\}\{T\}\\bigl\(\\sqrt\{2\\ln T\}\+1\+O\(\\delta\_\{T\}\)\\bigr\), and expandingΦ=12​\(u−1−ln⁡u\)\\Phi=\\tfrac\{1\}\{2\}\(u\-1\-\\ln u\)atu=1−Hu=1\-Has in \([48](https://arxiv.org/html/2608.06893#S6.E48)\) gives

Φ​\(x∗\)=H​\(x∗\)24​\(1\+O​\(H​\(x∗\)\)\)=ln⁡T2​T2​\(1\+O​\(1ln⁡T\)\)\.\\Phi\(x^\{\*\}\)=\\frac\{H\(x^\{\*\}\)^\{2\}\}\{4\}\\bigl\(1\+O\(H\(x^\{\*\}\)\)\\bigr\)=\\frac\{\\ln T\}\{2T^\{2\}\}\\Bigl\(1\+O\\bigl\(\\tfrac\{1\}\{\\sqrt\{\\ln T\}\}\\bigr\)\\Bigr\)\.\(72\)∎

#### Numerical certificates

Table 6:The constantx∗​\(T\)x^\{\*\}\(T\)against Proposition[1](https://arxiv.org/html/2608.06893#Thmproposition1)\.Exact minimization ofΦ\\Phion the uniform grid via \([41](https://arxiv.org/html/2608.06893#S6.E41)\)\. The ratio approaches11slowly, at the proved rateO​\(δT1/2​χT1/2\)O\(\\delta\_\{T\}^\{1/2\}\\chi\_\{T\}^\{1/2\}\); theT≤50T\\leq 50rows are the values quoted in App[7\.4](https://arxiv.org/html/2608.06893#S7.SS4)\.![Refer to caption](https://arxiv.org/html/2608.06893v1/x6.png)Figure 6:x∗​\(T\)x^\{\*\}\(T\)follows\(2​ln⁡T\)−1/2\(2\\ln T\)^\{\-1/2\}\.*Left:*exact optimizer against the law of Proposition[1](https://arxiv.org/html/2608.06893#Thmproposition1)\.*Right:*their ratio, which decreases from1\.291\.29atT=5T=5to1\.0171\.017atT=105T=10^\{5\}\. The approach is slow because the proved error isO​\(δT​χT\)O\(\\sqrt\{\\delta\_\{T\}\\chi\_\{T\}\}\), notO​\(1/T\)O\(1/T\)\.
#### Uniqueness \(Proposition[2](https://arxiv.org/html/2608.06893#Thmproposition2)\)

##### Reduction\.

SinceΦ​\(x\)=12​\(u−1−ln⁡u\)\\Phi\(x\)=\\tfrac\{1\}\{2\}\(u\-1\-\\ln u\)is strictly decreasing inu∈\(0,1\)u\\in\(0,1\), minimizingΦ\\Phiis equivalent to minimizingH​\(x\)=D0/PH\(x\)=D\_\{0\}/P\. Putωi:=\(1−ρi\)/ρi\\omega\_\{i\}:=\(1\-\\rho\_\{i\}\)/\\rho\_\{i\}\. Then the color path in the solver coordinate is

zi​\(x\)=xx\+ωi,H​\(x\)=F​\(z​\(x\)\),z\_\{i\}\(x\)=\\frac\{x\}\{x\+\\omega\_\{i\}\},\\qquad H\(x\)=F\\bigl\(z\(x\)\\bigr\),\(73\)withFFas in \([83](https://arxiv.org/html/2608.06893#S6.E83)\)\. Uniqueness ofx∗x^\{\*\}is thus a question about how the one\-parameter curvex↦z​\(x\)x\\mapsto z\(x\)meets the convex functionFF\.

##### T=2T=2: proved\.

HereH​\(x\)=z1​\(x\)\+\(1−z1​\(x\)\)2H\(x\)=z\_\{1\}\(x\)\+\\bigl\(1\-z\_\{1\}\(x\)\\bigr\)^\{2\}\. The mapz1z\_\{1\}is strictly increasing onto\(0,1\)\(0,1\)andJ​\(z\):=z\+\(1−z\)2J\(z\):=z\+\(1\-z\)^\{2\}is strictly convex, soHHis strictly unimodal andx∗x^\{\*\}is unique\.

##### Uniform grids: proved for2≤T≤1202\\leq T\\leq 120\.

For each integer2≤T≤1202\\leq T\\leq 120we derives an explicit polynomialPT∈ℤ​\[x\]P\_\{T\}\\in\\mathbb\{Z\}\[x\]whose positive roots are exactly the critical points ofHH\. The sequence of nonzero coefficients ofPTP\_\{T\}has exactly one sign change, so by Descartes’ rule \(S5\)PTP\_\{T\}has at most one positive root\(Basuet al\.,[2006](https://arxiv.org/html/2608.06893#bib.bib53), Sec\. 2\.2\.1\)\. Combined with the interior\-existence argument of Theorem[3](https://arxiv.org/html/2608.06893#Thmtheorem3), Step 3 \(the barriersΦ→∞\\Phi\\to\\inftyat both ends\), this proves thatHHhas a unique critical point, hence a unique minimum, for every2≤T≤1202\\leq T\\leq 120\.

Beyond that certified range, a tolerance\-robust numerical scan finds exactly one local minimum for everyT=2,…,200T=2,\\dots,200and forT∈\{500,1000,5000\}T\\in\\\{500,1000,5000\\\}\. A general analytic proof remains open\. It reduces it either to the conjecture that the relevant coefficient sequence has one sign change for everyTT, or to verifyingVarp⁡\(z\)<1/12\\operatorname\{Var\}\_\{p\}\(z\)<1/12at each critical point, where1/121/12is the limiting value in the continuum\.

##### General grids: false, by explicit construction\.

Non\-uniqueness already occurs atT=3T=3on a two\-cluster grid\. Takeρ=\(11\+C,12,1\)\\rho=\\bigl\(\\tfrac\{1\}\{1\+C\},\\tfrac\{1\}\{2\},1\\bigr\), equivalentlyω=\(C,1,0\)\\omega=\(C,1,0\)\. Thenz1=xx\+Cz\_\{1\}=\\tfrac\{x\}\{x\+C\},z2=xx\+1z\_\{2\}=\\tfrac\{x\}\{x\+1\},z3=1z\_\{3\}=1and

H​\(x\)=z1\+\(z2−z1\)2z2\+\(1−z2\)2\.H\(x\)=z\_\{1\}\+\\frac\{\(z\_\{2\}\-z\_\{1\}\)^\{2\}\}\{z\_\{2\}\}\+\(1\-z\_\{2\}\)^\{2\}\.\(74\)Table[7](https://arxiv.org/html/2608.06893#S6.T7)evaluates \([74](https://arxiv.org/html/2608.06893#S6.E74)\) at five points and shows that the value34\\tfrac\{3\}\{4\}is crossed four times, for everyC≥100C\\geq 100; By continuityHH— and thereforeΦ\\Phi— has at least two local minima, one in\(110,C\)\(\\tfrac\{1\}\{10\},\\sqrt\{C\}\)and one in\(C,10​C\)\(\\sqrt\{C\},10C\)\.

For fixedxx,H​\(x\)→J​\(xx\+1\)H\(x\)\\to J\\bigl\(\\tfrac\{x\}\{x\+1\}\\bigr\)withJJthe strictly unimodalT=2T\{=\}2profile above; forx=C​wx=Cwwithwwfixed,H→J​\(ww\+1\)H\\to J\\bigl\(\\tfrac\{w\}\{w\+1\}\\bigr\)\. Each cluster therefore creates oneT=2T\{=\}2valley, with minima approaching34\\tfrac\{3\}\{4\}nearx≈1x\\approx 1andx≈Cx\\approx C, and the valleys are separated by the cluster\-scale ratio\.

Table 7:Five evaluations that force two valleys\(Eq\. \([74](https://arxiv.org/html/2608.06893#S6.E74)\), anyC≥100C\\geq 100; the table readsC=100C=100\)\. Withu:=11\+Cu:=\\tfrac\{1\}\{1\+C\}andu¯:=11\+C≤111\\bar\{u\}:=\\tfrac\{1\}\{1\+\\sqrt\{C\}\}\\leq\\tfrac\{1\}\{11\}\. Each outer bound keeps a single term of \([74](https://arxiv.org/html/2608.06893#S6.E74)\); the two interior identities are exact substitutions\. All five were also checked in exact rational arithmetic\.
The interior identities follow by direct substitution, usingz1​\(1\)=uz\_\{1\}\(1\)=u,z2​\(1\)=12z\_\{2\}\(1\)=\\tfrac\{1\}\{2\},z1​\(C\)=12z\_\{1\}\(C\)=\\tfrac\{1\}\{2\}andz2​\(C\)=1−uz\_\{2\}\(C\)=1\-u\. For the middle lower bound we drop\(1−z2\)2≥0\(1\-z\_\{2\}\)^\{2\}\\geq 0and usez2≤1z\_\{2\}\\leq 1; moreoveru¯≤111\\bar\{u\}\\leq\\tfrac\{1\}\{11\}lies below3−58\\tfrac\{3\-\\sqrt\{5\}\}\{8\}, the smaller root of4​u¯2−3​u¯\+144\\bar\{u\}^\{2\}\-3\\bar\{u\}\+\\tfrac\{1\}\{4\}\.

##### Severity\.

Clustering can push the count higher\. A two\-cluster grid atT=48T\{=\}48with nodes concentrated near10−710^\{\-7\}and near11produces seven local minima ofHH, each confirmed in 50\-digit arithmetic\. The consequence for the theory is the conditional statement in Theorem[3](https://arxiv.org/html/2608.06893#Thmtheorem3): every*local*optimizer is mode\-wise proportional to*some*local minimizer ofΦ\\Phi, and the common formvk∗=x∗​Pkv\_\{k\}^\{\*\}=x^\{\*\}P\_\{k\}holds for global optimizers exactly whenarg​min⁡Φ\\operatorname\*\{arg\\,min\}\\Phiis a singleton\.

On a two\-cluster grid of this type, a two\-mode instance withP=\(1,4\)P=\(1,4\)has a genuine local optimizer whose modes occupy different valleys, withv2/v1v\_\{2\}/v\_\{1\}of order10610^\{6\}instead of the proportional ratio44; perturbations in every direction confirm local optimality\.

#### Budget bending \(Proposition[3](https://arxiv.org/html/2608.06893#Thmproposition3)\)

The relevant asymptotic regime holds the budgetVVfixed whileT→∞T\\to\\infty; this is also the regime the experiments test\. The proof uses Lemmas[7](https://arxiv.org/html/2608.06893#Thmlemma7)–[9](https://arxiv.org/html/2608.06893#Thmlemma9)plus one derivative lemma of the same kind\. Throughout,xk:=vk/Pkx\_\{k\}:=v\_\{k\}/P\_\{k\}andx¯k:=V​Pk/∑jPj2\\bar\{x\}\_\{k\}:=VP\_\{k\}/\\sum\_\{j\}P\_\{j\}^\{2\}; the hypothesis impliesx¯max≤34\\bar\{x\}\_\{\\max\}\\leq\\tfrac\{3\}\{4\}\.

###### Lemma 12\(Expansion and derivative on extended windows\)\.

FixX≥1X\\geq 1\. Uniformly onx∈\[T−1/4,X\]x\\in\[T^\{\-1/4\},X\], on the uniform grid,

T​H​\(x\)\\displaystyle T\\,H\(x\)=x​ln⁡T\+12​x\+O​\(1\),\\displaystyle=x\\ln T\+\\tfrac\{1\}\{2x\}\+O\(1\),\(75\)T​H′​\(x\)\\displaystyle T\\,H^\{\\prime\}\(x\)=ln⁡T−12​x2\+O​\(1x\),\\displaystyle=\\ln T\-\\tfrac\{1\}\{2x^\{2\}\}\+O\\bigl\(\\tfrac\{1\}\{x\}\\bigr\),\(76\)with constants depending only onXX\.

###### Proof\.

Extension of the three lemmas\.The restrictionx≤34x\\leq\\tfrac\{3\}\{4\}entered only through the decreasing\-then\-increasing shape offfin Lemma[8](https://arxiv.org/html/2608.06893#Thmlemma8)\. Forb<14b<\\tfrac\{1\}\{4\}\(includingb≤0b\\leq 0\) one checks\(ln⁡f\)′=−1/ρ\+3​b/ψ<0\(\\ln f\)^\{\\prime\}=\-1/\\rho\+3b/\\psi<0on all of\(0,1\]\(0,1\], soffis monotone and the total\-variation bound only improves\. The antiderivative of Lemma[9](https://arxiv.org/html/2608.06893#Thmlemma9)is valid for everyb≠0b\\neq 0, andx=1x=1is the trivial caseψ≡1\\psi\\equiv 1\. So Lemmas[7](https://arxiv.org/html/2608.06893#Thmlemma7)–[9](https://arxiv.org/html/2608.06893#Thmlemma9)extend from\(0,34\]\(0,\\tfrac\{3\}\{4\}\]to\(0,X\]\(0,X\]\.

The value\.Chaining as in \([65](https://arxiv.org/html/2608.06893#S6.E65)\) and notingx​ln⁡1x=O​\(1\)x\\ln\\tfrac\{1\}\{x\}=O\(1\)and1T​x3≤T−1/4\\tfrac\{1\}\{Tx^\{3\}\}\\leq T^\{\-1/4\}on the window gives \([75](https://arxiv.org/html/2608.06893#S6.E75)\)\.

The derivative\.FromT​H=x​𝒮​\(x\)T\\,H=x\\mathcal\{S\}\(x\)we getT​H′=𝒮\+x​𝒮′T\\,H^\{\\prime\}=\\mathcal\{S\}\+x\\mathcal\{S\}^\{\\prime\}, and \([75](https://arxiv.org/html/2608.06893#S6.E75)\) gives𝒮=ln⁡T\+12​x2\+O​\(1x\)\\mathcal\{S\}=\\ln T\+\\tfrac\{1\}\{2x^\{2\}\}\+O\(\\tfrac\{1\}\{x\}\)\. Differentiating \([52](https://arxiv.org/html/2608.06893#S6.E52)\) termwise with∂xψ​\(ρ\)=ρ\\partial\_\{x\}\\psi\(\\rho\)=\\rho,

−𝒮′​\(x\)=1T​∑i=1T\[1ψi2​ψi−12\+2​ρi−1/ρiψi​ψi−13\],\-\\mathcal\{S\}^\{\\prime\}\(x\)=\\frac\{1\}\{T\}\\sum\_\{i=1\}^\{T\}\\left\[\\frac\{1\}\{\\psi\_\{i\}^\{2\}\\psi\_\{i\-1\}^\{2\}\}\+\\frac\{2\\,\\rho\_\{i\-1\}/\\rho\_\{i\}\}\{\\psi\_\{i\}\\,\\psi\_\{i\-1\}^\{3\}\}\\right\],\(77\)which is a pair of Riemann\-type sums of the integrandsψ−4\\psi^\{\-4\}, monotone inρ\\rho, up to the grid shift of Lemma[7](https://arxiv.org/html/2608.06893#Thmlemma7)and the factorρi−1/ρi=1−1i\\rho\_\{i\-1\}/\\rho\_\{i\}=1\-\\tfrac\{1\}\{i\}, whose deviation contributes at most2​ln⁡TT​x4=O​\(1x\)\\tfrac\{2\\ln T\}\{Tx^\{4\}\}=O\(\\tfrac\{1\}\{x\}\)on the window\. Since

∫01ψ−4​dρ=x−3−13​b=x−33\+O​\(x−2\)\\int\_\{0\}^\{1\}\\psi^\{\-4\}\\,\\mathrm\{d\}\\rho=\\frac\{x^\{\-3\}\-1\}\{3b\}=\\frac\{x^\{\-3\}\}\{3\}\+O\(x^\{\-2\}\)\(78\)and the total\-variation Riemann error isO​\(x−4/T\)=O​\(1x\)O\(x^\{\-4\}/T\)=O\(\\tfrac\{1\}\{x\}\)there, we get−𝒮′=x−3\+O​\(x−2\)\-\\mathcal\{S\}^\{\\prime\}=x^\{\-3\}\+O\(x^\{\-2\}\)and henceT​H′=ln⁡T\+12​x2−1x2\+O​\(1x\)T\\,H^\{\\prime\}=\\ln T\+\\tfrac\{1\}\{2x^\{2\}\}\-\\tfrac\{1\}\{x^\{2\}\}\+O\(\\tfrac\{1\}\{x\}\), which is \([76](https://arxiv.org/html/2608.06893#S6.E76)\)\. ∎

###### Proof of Proposition[3](https://arxiv.org/html/2608.06893#Thmproposition3)\.

BecauseΦ→∞\\Phi\\to\\inftyat0andΦ\\Phiis continuous, a global optimizer exists in the interior and satisfies the Lagrange condition1Pk​Φ′​\(xk\)=λ\\tfrac\{1\}\{P\_\{k\}\}\\Phi^\{\\prime\}\(x\_\{k\}\)=\\lambdafor everykk, whereΦ′​\(x\)=12​H1−H​H′​\(x\)\\Phi^\{\\prime\}\(x\)=\\tfrac\{1\}\{2\}\\tfrac\{H\}\{1\-H\}H^\{\\prime\}\(x\)\. In particularsign⁡Φ′=sign⁡H′\\operatorname\{sign\}\\Phi^\{\\prime\}=\\operatorname\{sign\}H^\{\\prime\}andΦ=H24​\(1\+O​\(H\)\)\\Phi=\\tfrac\{H^\{2\}\}\{4\}\(1\+O\(H\)\)\.

Step 1 \(value benchmark\)\.The feasible pointv¯k=x¯k​Pk\\bar\{v\}\_\{k\}=\\bar\{x\}\_\{k\}P\_\{k\}has objective value

ΦTbm:=∑kΦ​\(x¯k\)=ln2⁡T4​T2​\(∑kx¯k2\)​\(1\+o​\(1\)\)\\Phi^\{\\mathrm\{bm\}\}\_\{T\}:=\\sum\_\{k\}\\Phi\(\\bar\{x\}\_\{k\}\)=\\frac\{\\ln^\{2\}T\}\{4T^\{2\}\}\\Bigl\(\\sum\_\{k\}\\bar\{x\}\_\{k\}^\{2\}\\Bigr\)\\bigl\(1\+o\(1\)\\bigr\)\(79\)by Lemma[12](https://arxiv.org/html/2608.06893#Thmlemma12)\. Any optimizer does at least as well, so every coordinate satisfiesΦ​\(xk\)≤ΦTbm=O​\(ln2⁡T/T2\)\\Phi\(x\_\{k\}\)\\leq\\Phi^\{\\mathrm\{bm\}\}\_\{T\}=O\(\\ln^\{2\}T/T^\{2\}\)\.

Step 2 \(localization\)\.First excludexk≤T−1/4x\_\{k\}\\leq T^\{\-1/4\}: there \([68](https://arxiv.org/html/2608.06893#S6.E68)\) givesT​H≥12​x​\(1−2​T−1/2\)−2​x≥T1/43T\\,H\\geq\\tfrac\{1\}\{2x\}\(1\-2T^\{\-1/2\}\)\-2x\\geq\\tfrac\{T^\{1/4\}\}\{3\}, soΦ​\(xk\)≥H24≫ΦTbm\\Phi\(x\_\{k\}\)\\geq\\tfrac\{H^\{2\}\}\{4\}\\gg\\Phi^\{\\mathrm\{bm\}\}\_\{T\}\. Next excludexk≥xhix\_\{k\}\\geq x\_\{\\mathrm\{hi\}\}for a large fixed constantxhix\_\{\\mathrm\{hi\}\}depending only on∑jx¯j2\\sum\_\{j\}\\bar\{x\}\_\{j\}^\{2\}, using the regional bounds of Table[5](https://arxiv.org/html/2608.06893#S6.T5):T​H≥x​ln⁡TT\\,H\\geq x\\ln Ton\[T−1/4,1\]\[T^\{\-1/4\},1\],T​H≥827​x​ln⁡T2​xT\\,H\\geq\\tfrac\{8\}\{27\}x\\ln\\tfrac\{T\}\{2x\}on\[1,T\]\[1,\\sqrt\{T\}\], andT​H≥T2T\\,H\\geq\\tfrac\{\\sqrt\{T\}\}\{2\}beyond\.

Nowλ\>0\\lambda\>0\. Ifλ≤0\\lambda\\leq 0thenH′​\(xk\)≤0H^\{\\prime\}\(x\_\{k\}\)\\leq 0for everykk, so on\[T−1/4,xhi\]\[T^\{\-1/4\},x\_\{\\mathrm\{hi\}\}\]Lemma[12](https://arxiv.org/html/2608.06893#Thmlemma12)would forcexk≤\(2​ln⁡T\)−1/2​\(1\+o​\(1\)\)x\_\{k\}\\leq\(2\\ln T\)^\{\-1/2\}\(1\+o\(1\)\)and hence∑kvk≤\(2​ln⁡T\)−1/2​\(1\+o​\(1\)\)​∑kPk<V\\sum\_\{k\}v\_\{k\}\\leq\(2\\ln T\)^\{\-1/2\}\(1\+o\(1\)\)\\sum\_\{k\}P\_\{k\}<Vfor largeTT, contradicting the budget\. Thereforeλ\>0\\lambda\>0, every coordinate hasH′​\(xk\)\>0H^\{\\prime\}\(x\_\{k\}\)\>0, andxk≥\(2​ln⁡T\)−1/2​\(1−o​\(1\)\)x\_\{k\}\\geq\(2\\ln T\)^\{\-1/2\}\(1\-o\(1\)\)\.

Finally all coordinates are bounded below by a*fixed*constant\. The budget forces somexk0≥V/∑jPj=:xlo\>0x\_\{k\_\{0\}\}\\geq V/\\sum\_\{j\}P\_\{j\}=:x\_\{\\mathrm\{lo\}\}\>0, and Lemma[12](https://arxiv.org/html/2608.06893#Thmlemma12)then givesλ=1Pk0​Φ′​\(xk0\)≥c​ln2⁡TT2\\lambda=\\tfrac\{1\}\{P\_\{k\_\{0\}\}\}\\Phi^\{\\prime\}\(x\_\{k\_\{0\}\}\)\\geq c\\,\\tfrac\{\\ln^\{2\}T\}\{T^\{2\}\}\. Suppose somexk≤εx\_\{k\}\\leq\\varepsilonwithε\\varepsilonsmall and fixed\. ThenΦ′​\(xk\)=λ​Pk≥c′​ln2⁡TT2\\Phi^\{\\prime\}\(x\_\{k\}\)=\\lambda P\_\{k\}\\geq c^\{\\prime\}\\tfrac\{\\ln^\{2\}T\}\{T^\{2\}\}would require

\(xk​ln⁡T\+12​xk\)​\(ln⁡T−12​xk2\)≥c′′​ln2⁡T,\\Bigl\(x\_\{k\}\\ln T\+\\tfrac\{1\}\{2x\_\{k\}\}\\Bigr\)\\Bigl\(\\ln T\-\\tfrac\{1\}\{2x\_\{k\}^\{2\}\}\\Bigr\)\\geq c^\{\\prime\\prime\}\\ln^\{2\}T,\(80\)and forxk≤εx\_\{k\}\\leq\\varepsilonthe first factor can reachc′′​ln⁡Tc^\{\\prime\\prime\}\\ln Tonly through12​xk≳ln⁡T\\tfrac\{1\}\{2x\_\{k\}\}\\gtrsim\\ln T, which makes the second factor negative — a contradiction\. Hence everyxk∈\[c′,xhi\]x\_\{k\}\\in\[c^\{\\prime\},x\_\{\\mathrm\{hi\}\}\], a fixed compact interval\.

Step 3 \(stationarity gives the law\)\.On that compact, Lemma[12](https://arxiv.org/html/2608.06893#Thmlemma12)gives uniformly

Φ′​\(x\)\\displaystyle\\Phi^\{\\prime\}\(x\)=12​H​\(x\)​H′​\(x\)​\(1\+O​\(H\)\)\\displaystyle=\\tfrac\{1\}\{2\}H\(x\)H^\{\\prime\}\(x\)\\bigl\(1\+O\(H\)\\bigr\)=x​ln2⁡T2​T2​\(1\+O​\(1ln⁡T\)\),\\displaystyle=\\frac\{x\\ln^\{2\}T\}\{2T^\{2\}\}\\Bigl\(1\+O\\bigl\(\\tfrac\{1\}\{\\ln T\}\\bigr\)\\Bigr\),\(81\)so1Pk​Φ′​\(xk\)=λ\\tfrac\{1\}\{P\_\{k\}\}\\Phi^\{\\prime\}\(x\_\{k\}\)=\\lambdabecomesxk=κ​Pk​\(1\+O​\(1/ln⁡T\)\)x\_\{k\}=\\kappa P\_\{k\}\\bigl\(1\+O\(1/\\ln T\)\\bigr\)with the common constantκ=2​λ​T2/ln2⁡T\\kappa=2\\lambda T^\{2\}/\\ln^\{2\}T\. The budget equationV=∑kxk​Pk=κ​∑kPk2​\(1\+O​\(1/ln⁡T\)\)V=\\sum\_\{k\}x\_\{k\}P\_\{k\}=\\kappa\\sum\_\{k\}P\_\{k\}^\{2\}\(1\+O\(1/\\ln T\)\)determinesκ\\kappa, and therefore

vk∗=xk​Pk=Pk2∑jPj2​V​\(1\+O​\(1/ln⁡T\)\),v\_\{k\}^\{\*\}=x\_\{k\}P\_\{k\}=\\frac\{P\_\{k\}^\{2\}\}\{\\sum\_\{j\}P\_\{j\}^\{2\}\}\\,V\\,\\bigl\(1\+O\(1/\\ln T\)\\bigr\),\(82\)which is \([8](https://arxiv.org/html/2608.06893#S3.E8)\)\. The hypothesisV​Pmax/∑jPj2≤34VP\_\{\\max\}/\\sum\_\{j\}P\_\{j\}^\{2\}\\leq\\tfrac\{3\}\{4\}ensures the limiting allocation stays in the window where the constants are uniform\. ∎

Table 8:The budgeted optimum bends toward theP2P^\{2\}law\.Two modes,P=\(1,4\)P=\(1,4\), budgetV=5V=5, uniform grid; exact minimization of∑kΦk\\sum\_\{k\}\\Phi\_\{k\}\. The limit isP12​V/∑jPj2=517=0\.294P\_\{1\}^\{2\}V/\\sum\_\{j\}P\_\{j\}^\{2\}=\\tfrac\{5\}\{17\}=0\.294, approached at the proved rateO​\(1/ln⁡T\)O\(1/\\ln T\)\. Note that equal\-budget protocols and free\-noise optimization answer different questions and have different exponents; both should be reported\.
#### Proof of Theorem[4](https://arxiv.org/html/2608.06893#Thmtheorem4)

###### Proof\.

\(The identity\.\)Letz​\(ρ\)=v​ρ/φ​\(ρ\)z\(\\rho\)=v\\rho/\\varphi\(\\rho\)\. Substituting into \([41](https://arxiv.org/html/2608.06893#S6.E41)\) throughφ=v​ρ/z\\varphi=v\\rho/zand1−z=\(1−ρ\)​P/φ1\-z=\(1\-\\rho\)P/\\varphicollapses the expression to

D0=P​F​\(z\),F​\(z\):=∑i=1T\(Δ​zi\)2zi,D\_\{0\}=P\\,F\(z\),\\qquad F\(z\):=\\sum\_\{i=1\}^\{T\}\\frac\{\(\\Delta z\_\{i\}\)^\{2\}\}\{z\_\{i\}\},\(83\)withz0=0z\_\{0\}=0andzT=1z\_\{T\}=1; the two\-line simplification has symbolic residual0\. Equation \([83](https://arxiv.org/html/2608.06893#S6.E83)\) restates \([9](https://arxiv.org/html/2608.06893#S3.E9)\)\. Two identities used repeatedly below are

M​\(ρ\)=P​z​\(ρ\),Δ​zizi=P​Δ​ρiρi​φ​\(ρi−1\),M\(\\rho\)=P\\,z\(\\rho\),\\qquad\\frac\{\\Delta z\_\{i\}\}\{z\_\{i\}\}=\\frac\{P\\,\\Delta\\rho\_\{i\}\}\{\\rho\_\{i\}\\,\\varphi\(\\rho\_\{i\-1\}\)\},\(84\)both immediate fromK​\(ρ\)=P/φ​\(ρ\)K\(\\rho\)=P/\\varphi\(\\rho\)\.

\(a\) Color\-blindness of the achievable set\.For everyv\>0v\>0the mapρ↦z\\rho\\mapsto zis a continuous strictly increasing bijection of\[0,1\]\[0,1\]\(left panel of Figure[7](https://arxiv.org/html/2608.06893#S6.F7)\)\. The achievablezz\-grids — and hence the schedule\-optimized deficit — therefore coincide for every color at every finiteTT\.

\(b\) Uniqueness of the optimal grid\.FFis a sum of quadratic\-over\-linear terms and so is jointly convex\. If two minimizers existed,FFwould be affine on the segment joining them, which forces every ratiozi−1/ziz\_\{i\-1\}/z\_\{i\}to be constant along that segment\. Together with the pinned endpoints those ratios determine the grid, so the minimizer is unique\.

\(c\) The optimal grid in closed form\.Putai:=zi−1/zia\_\{i\}:=z\_\{i\-1\}/z\_\{i\}\. Stationarity ofFFin the interior coordinates gives

2​ai\+1=1\+ai2,a1=z0z1=0,2a\_\{i\+1\}=1\+a\_\{i\}^\{2\},\\qquad a\_\{1\}=\\frac\{z\_\{0\}\}\{z\_\{1\}\}=0,\(85\)an explicit forward recurrence, from whichzi=∏j\>iajz\_\{i\}=\\prod\_\{j\>i\}a\_\{j\}\. Table[9](https://arxiv.org/html/2608.06893#S6.T9)lists the first values and the resulting optimal costs\. Brute\-force minimization agrees to10−1410^\{\-14\}forT=2,3,4T=2,3,4\. AtT=2T=2the exact optimum isz1=12z\_\{1\}=\\tfrac\{1\}\{2\}withF∗=34F^\{\*\}=\\tfrac\{3\}\{4\}, whereas the continuum schedulez=s2z=s^\{2\}gives1316\\tfrac\{13\}\{16\}: the continuum schedule is*not*exactly optimal at finiteTT\.

\(d\) Asymptotics of the optimum\.Setεi:=1−ai\\varepsilon\_\{i\}:=1\-a\_\{i\}\. Two exact rewrites start the argument\. FirstΔ​zi=zi​\(1−ai\)\\Delta z\_\{i\}=z\_\{i\}\(1\-a\_\{i\}\), so the optimal value is

F∗=∑i=1Tzi​εi2\.F^\{\*\}=\\sum\_\{i=1\}^\{T\}z\_\{i\}\\,\\varepsilon\_\{i\}^\{2\}\.\(86\)Second, \([85](https://arxiv.org/html/2608.06893#S6.E85)\) becomes*exactly*

εi\+1=εi​\(1−εi2\),ε1=1,\\varepsilon\_\{i\+1\}=\\varepsilon\_\{i\}\\Bigl\(1\-\\frac\{\\varepsilon\_\{i\}\}\{2\}\\Bigr\),\\qquad\\varepsilon\_\{1\}=1,\(87\)a purely quadratic recurrence with no cubic term\. It showsεi∈\(0,1\]\\varepsilon\_\{i\}\\in\(0,1\]and that the sequence decreases, so

1εi\+1\\displaystyle\\frac\{1\}\{\\varepsilon\_\{i\+1\}\}=1εi\+12⋅11−εi/2\\displaystyle=\\frac\{1\}\{\\varepsilon\_\{i\}\}\+\\frac\{1\}\{2\}\\cdot\\frac\{1\}\{1\-\\varepsilon\_\{i\}/2\}=1εi\+12\+εi4\+O​\(εi2\)\.\\displaystyle=\\frac\{1\}\{\\varepsilon\_\{i\}\}\+\\frac\{1\}\{2\}\+\\frac\{\\varepsilon\_\{i\}\}\{4\}\+O\(\\varepsilon\_\{i\}^\{2\}\)\.\(88\)
*Sharp recurrence asymptotics\.*Dropping the positive correction in \([88](https://arxiv.org/html/2608.06893#S6.E88)\) gives1/εi≥\(i\+1\)/21/\\varepsilon\_\{i\}\\geq\(i\+1\)/2, i\.e\.εi≤2/\(i\+1\)\\varepsilon\_\{i\}\\leq 2/\(i\+1\)\. Substituting that bound back yields∑j<iεj/4=12​ln⁡i\+O​\(1\)\\sum\_\{j<i\}\\varepsilon\_\{j\}/4=\\tfrac\{1\}\{2\}\\ln i\+O\(1\)and∑j<iεj2=O​\(1\)\\sum\_\{j<i\}\\varepsilon\_\{j\}^\{2\}=O\(1\), so a two\-sided induction with explicit constants gives

1εi\\displaystyle\\frac\{1\}\{\\varepsilon\_\{i\}\}=i2\+12​ln⁡i\+O​\(1\),\\displaystyle=\\frac\{i\}\{2\}\+\\frac\{1\}\{2\}\\ln i\+O\(1\),\(89\)εi\\displaystyle\\varepsilon\_\{i\}=2i​\(1\+O​\(ln⁡\(i\+1\)i\+1\)\)\.\\displaystyle=\\frac\{2\}\{i\}\\Bigl\(1\+O\\bigl\(\\tfrac\{\\ln\(i\+1\)\}\{i\+1\}\\bigr\)\\Bigr\)\.\(90\)
*Modal positions\.*Sincezi=∏j\>iajz\_\{i\}=\\prod\_\{j\>i\}a\_\{j\},

ln⁡zi\\displaystyle\\ln z\_\{i\}=∑j\>iln⁡\(1−εj\)\\displaystyle=\\sum\_\{j\>i\}\\ln\(1\-\\varepsilon\_\{j\}\)=−∑j\>iεj\+O​\(∑j\>iεj2\)\\displaystyle=\-\\sum\_\{j\>i\}\\varepsilon\_\{j\}\+O\\Bigl\(\\sum\_\{j\>i\}\\varepsilon\_\{j\}^\{2\}\\Bigr\)=−2​ln⁡Ti\+O​\(ln⁡\(i\+1\)i\+1\),\\displaystyle=\-2\\ln\\frac\{T\}\{i\}\+O\\Bigl\(\\frac\{\\ln\(i\+1\)\}\{i\+1\}\\Bigr\),\(91\)using∑j\>i2/j=2​ln⁡\(T/i\)\+O​\(1/i\)\\sum\_\{j\>i\}2/j=2\\ln\(T/i\)\+O\(1/i\)and the fact that the two correction sums∑j\>i\(ln⁡j\)/j2\\sum\_\{j\>i\}\(\\ln j\)/j^\{2\}and∑j\>i1/j2\\sum\_\{j\>i\}1/j^\{2\}are bothO​\(ln⁡\(i\+1\)/\(i\+1\)\)O\(\\ln\(i\+1\)/\(i\+1\)\)\. Hence

zi∗=\(i/T\)2​eO​\(ln⁡\(i\+1\)/\(i\+1\)\)\.z\_\{i\}^\{\*\}=\(i/T\)^\{2\}\\,e^\{\\,O\(\\ln\(i\+1\)/\(i\+1\)\)\}\.\(92\)This is the rigorous convergence statement, with its rate: the exact optimizer approaches the continuum schedulez=s2z=s^\{2\}\. Equivalently the limiting schedule is

ρ∗​\(s\)=P​s2P​s2\+v​\(1−s2\),\\rho^\{\*\}\(s\)=\\frac\{Ps^\{2\}\}\{Ps^\{2\}\+v\(1\-s^\{2\}\)\},\(93\)under whichM​\(ρ∗​\(s\)\)=P​s2M\(\\rho^\{\*\}\(s\)\)=Ps^\{2\}by \([84](https://arxiv.org/html/2608.06893#S6.E84)\): the conditional MMSE is quadratic in solver time\.

*The floor\.*Combining \([90](https://arxiv.org/html/2608.06893#S6.E90)\) and \([92](https://arxiv.org/html/2608.06893#S6.E92)\) term by term,

zi​εi2\\displaystyle z\_\{i\}\\varepsilon\_\{i\}^\{2\}=4T2​\(1\+O​\(ln⁡\(i\+2\)i\+1\)\),\\displaystyle=\\frac\{4\}\{T^\{2\}\}\\Bigl\(1\+O\\bigl\(\\tfrac\{\\ln\(i\+2\)\}\{i\+1\}\\bigr\)\\Bigr\),\(94\)F∗\\displaystyle F^\{\*\}=4T\+O​\(1T2​∑i=1Tln⁡\(i\+2\)i\+1\)\\displaystyle=\\frac\{4\}\{T\}\+O\\Bigl\(\\frac\{1\}\{T^\{2\}\}\\sum\_\{i=1\}^\{T\}\\frac\{\\ln\(i\+2\)\}\{i\+1\}\\Bigr\)=4T​\(1\+O​\(log2⁡TT\)\),\\displaystyle=\\frac\{4\}\{T\}\\Bigl\(1\+O\\bigl\(\\tfrac\{\\log^\{2\}T\}\{T\}\\bigr\)\\Bigr\),\(95\)soD0∗=4​PT​\(1\+O​\(log2⁡T/T\)\)D\_\{0\}^\{\*\}=\\tfrac\{4P\}\{T\}\\bigl\(1\+O\(\\log^\{2\}T/T\)\\bigr\)\. Thelog2\\log^\{2\}factor comes from∑i\(ln⁡i\)/i\\sum\_\{i\}\(\\ln i\)/i; improving the rate would require cancellation among the per\-term corrections, and no such cancellation is evident\. The last column of Table[9](https://arxiv.org/html/2608.06893#S6.T9)supports the proved law\.

Finally, for the uniformss\-grid underz=s2z=s^\{2\}the summand of \([83](https://arxiv.org/html/2608.06893#S6.E83)\) simplifies symbolically toP​\(si2−si−12\)2/si2P\(s\_\{i\}^\{2\}\-s\_\{i\-1\}^\{2\}\)^\{2\}/s\_\{i\}^\{2\}, in which the colorvvdisappears identically — confirming \(a\) term by term\. ∎

Table 9:The optimalzz\-schedule of Theorem[4](https://arxiv.org/html/2608.06893#Thmtheorem4)\(c\)\.*Left:*the first ratiosai=zi−1/zia\_\{i\}=z\_\{i\-1\}/z\_\{i\}from the forward recurrence \([85](https://arxiv.org/html/2608.06893#S6.E85)\), as exact rationals\.*Right:*the optimal cost, exactly for smallTTand against the floorF∗=4/TF^\{\*\}=4/Tof part \(d\)\. The approach to11is at the proved rate1\+O​\(log2⁡T/T\)1\+O\(\\log^\{2\}T/T\)\.![Refer to caption](https://arxiv.org/html/2608.06893v1/x7.png)Figure 7:The solver coordinatezzabsorbs the color\.*Left:*ρ↦z=v​ρ/φ​\(ρ\)\\rho\\mapsto z=v\\rho/\\varphi\(\\rho\)for five values ofx=v/Px=v/P\. Each is a bijection of\[0,1\]\[0,1\], so everyzz\-grid is reachable at every color — this is Theorem[4](https://arxiv.org/html/2608.06893#Thmtheorem4)a\.*Center:*the exact optimizerzi∗z\_\{i\}^\{\*\}from the recurrence \([85](https://arxiv.org/html/2608.06893#S6.E85)\) against the continuum law\(i/T\)2\(i/T\)^\{2\}atT=64T=64, illustrating \([92](https://arxiv.org/html/2608.06893#S6.E92)\)\.*Right:*F∗​T/4→1F^\{\*\}T/4\\to 1, the floorD0∗=4​P/TD\_\{0\}^\{\*\}=4P/Tof part \(d\)\.
#### The multi\-objective schedule frontier

The discretization\-optimal schedule has gain\-error susceptibility

Ξ:=∑iΔ​zizi​\(1−zi\)=2​ln⁡T\+O​\(1\),\\Xi:=\\sum\_\{i\}\\frac\{\\Delta z\_\{i\}\}\{z\_\{i\}\}\\,\(1\-z\_\{i\}\)=2\\ln T\+O\(1\),\(96\)which diverges logarithmically; numericallyΞ=2​ln⁡T−3\.05\\Xi=2\\ln T\-3\.05atT=104T=10^\{4\}\. The two objectives therefore conflict: discretization favors many small steps nearz=0z=0, whereas robustness favors a short log\-path\. Under model error, reference design becomes a multi\-objectivezz\-schedule problem — minimize∑i\(Δ​zi\)2/zi\\sum\_\{i\}\(\\Delta z\_\{i\}\)^\{2\}/z\_\{i\}subject to a susceptibility budget\. The constraint involves the ratioszi−1/ziz\_\{i\-1\}/z\_\{i\}, which are*not*jointly convex, so convexity of the exact finite\-TTfrontier does not follow from convexity ofFF\. In the continuum, however, the frontier is available in closed form\.

###### Theorem 6\(Continuum susceptibility–discretization frontier\)\.

Consider increasingC1C^\{1\}schedulesz:\[0,1\]→\(0,1\]z:\[0,1\]\\to\(0,1\]withz​\(0\+\)=z1z\(0^\{\+\}\)=z\_\{1\}andz​\(1\)=1z\(1\)=1\.

\(i\) Susceptibility is path\-independent\.

Ξcont=∫01z˙z​\(1−z\)​ds=∫z111−zz​dz=ln⁡1z1−\(1−z1\)\.\\Xi^\{\\mathrm\{cont\}\}=\\int\_\{0\}^\{1\}\\frac\{\\dot\{z\}\}\{z\}\(1\-z\)\\,\\mathrm\{d\}s=\\int\_\{z\_\{1\}\}^\{1\}\\frac\{1\-z\}\{z\}\\,\\mathrm\{d\}z=\\ln\\frac\{1\}\{z\_\{1\}\}\-\(1\-z\_\{1\}\)\.\(97\)This depends only on the launch pointz1z\_\{1\}\. The budgetΞcont≤ξ\\Xi^\{\\mathrm\{cont\}\}\\leq\\xitherefore imposes onlyz1≥a​\(ξ\)z\_\{1\}\\geq a\(\\xi\)and no further path constraint, wherea​\(ξ\)∈\(0,1\)a\(\\xi\)\\in\(0,1\)is the unique root ofln⁡1a−\(1−a\)=ξ\\ln\\tfrac\{1\}\{a\}\-\(1\-a\)=\\xi\.

\(ii\) The square law, derived\.Substitutingw=zw=\\sqrt\{z\}turns the discretization functional into a Dirichlet energy,

∫01z˙2z​ds=4​∫01w˙2​ds,\\int\_\{0\}^\{1\}\\frac\{\\dot\{z\}^\{2\}\}\{z\}\\,\\mathrm\{d\}s=4\\int\_\{0\}^\{1\}\\dot\{w\}^\{2\}\\,\\mathrm\{d\}s,\(98\)with endpoints fixed atw​\(0\)=aw\(0\)=\\sqrt\{a\}andw​\(1\)=1w\(1\)=1\. Cauchy–Schwarz gives the unique minimizerw​\(s\)=a\+\(1−a\)​sw\(s\)=\\sqrt\{a\}\+\(1\-\\sqrt\{a\}\)s, that is the free square\-law path launched ataa:

z​\(s\)=\(a\+\(1−a\)​s\)2,E​\(ξ\)=4​\(1−a​\(ξ\)\)2\.z\(s\)=\\bigl\(\\sqrt\{a\}\+\(1\-\\sqrt\{a\}\)s\\bigr\)^\{2\},\\qquad E\(\\xi\)=4\\bigl\(1\-\\sqrt\{a\(\\xi\)\}\\bigr\)^\{2\}\.\(99\)
\(iii\) Closed\-form,C1C^\{1\}frontier\.Differentiating the constraint,

d​ad​ξ\\displaystyle\\frac\{\\,\\mathrm\{d\}a\}\{\\,\\mathrm\{d\}\\xi\}=−a1−a,\\displaystyle=\-\\frac\{a\}\{1\-a\},\(100\)d2​ad​ξ2\\displaystyle\\frac\{\\,\\mathrm\{d\}^\{2\}a\}\{\\,\\mathrm\{d\}\\xi^\{2\}\}=a\(1−a\)3\>0,\\displaystyle=\\frac\{a\}\{\(1\-a\)^\{3\}\}\>0,\(101\)d​Ed​ξ\\displaystyle\\frac\{\\,\\mathrm\{d\}E\}\{\\,\\mathrm\{d\}\\xi\}=4​a1\+a\>0\.\\displaystyle=\\frac\{4\\sqrt\{a\}\}\{1\+\\sqrt\{a\}\}\>0\.\(102\)Both derivatives are strictly monotone inξ\\xi:aadecreases asξ\\xiincreases, whilea/\(1\+a\)\\sqrt\{a\}/\(1\+\\sqrt\{a\}\)increases withaa\. Each branch therefore has curvature of one sign\. The launch pointa​\(ξ\)a\(\\xi\)isC1C^\{1\}, strictly decreasing and strictly convex; the smooth\-path energyE​\(ξ\)E\(\\xi\)isC1C^\{1\}, strictly increasing and strictly concave, and approaches the free value44asξ→∞\\xi\\to\\infty\.

At finiteTTthe launch pointz1z\_\{1\}is exactly the first step’sO​\(1\)O\(1\)contribution to the deficit, because its term in \([83](https://arxiv.org/html/2608.06893#S6.E83)\) is\(z1−0\)2/z1=z1\(z\_\{1\}\-0\)^\{2\}/z\_\{1\}=z\_\{1\}\. Along the envelope family the total deficit is therefore

a​\(ξ\)\+1T​E​\(ξ\)​\(1\+o​\(1\)\),a\(\\xi\)\+\\tfrac\{1\}\{T\}E\(\\xi\)\\bigl\(1\+o\(1\)\\bigr\),\(103\)whose leading\-order frontier is convex, decreasing and available in closed form; the concaveEE\-branch contributes only theO​\(1/T\)O\(1/T\)correction\.

The susceptibility constraint becomes inactive nearξ≈2​ln⁡T\\xi\\approx 2\\ln T, where the frontier flattens at the free optimum of Theorem[4](https://arxiv.org/html/2608.06893#Thmtheorem4)d\. The continuum identity in \(i\) is the main conceptual statement and replaces the discrete boundΞ≤ln⁡\(1/z1\)\\Xi\\leq\\ln\(1/z\_\{1\}\); the discrete bound remains useful as a finite\-TTcomparison, and the difference between them,1−z11\-z\_\{1\}, is exactly the first step’s susceptibility contribution\.

### 6\.5Proofs for Section[3\.4](https://arxiv.org/html/2608.06893#S3.SS4)\(Model Error and Learning\)

#### Completeness of the error model

In the solvable model the exact drift is affine,x^0=W​x1\+Ki​\(x−c¯i​x1\)\\hat\{x\}\_\{0\}=Wx\_\{1\}\+K\_\{i\}\(x\-\\bar\{c\}\_\{i\}x\_\{1\}\)\. Any learned affine drift can differ from it in exactly two ways: a relative gain errorηi\\eta\_\{i\}, which mis\-scales the sensitivity to the state deviation, and an additive biasβi\\beta\_\{i\}\. Because the analysis conditions onx1x\_\{1\}, the bias may depend onx1x\_\{1\}; consequently an error in the drift’sx1x\_\{1\}\-coefficient — for instance usingW^≠W\\hat\{W\}\\neq W— is absorbed intoβi\\beta\_\{i\}\. The two\-channel decomposition is therefore exhaustive within the model\. Both channels preserve the affine–Gaussian structure, so the terminal laws remain exact\.

#### Proof of Proposition[4](https://arxiv.org/html/2608.06893#Thmproposition4)

###### Proof\.

The perturbed step gives

c′=\[r\+\(1−r\)​K​\(1\+η\)\]​c\+\(1−r\)​\[W−K​\(1\+η\)​c¯\]\.c^\{\\prime\}=\\bigl\[r\+\(1\-r\)K\(1\+\\eta\)\\bigr\]c\+\(1\-r\)\\bigl\[W\-K\(1\+\\eta\)\\bar\{c\}\\bigr\]\.\(104\)Substituting the inductive valuec=c¯ic=\\bar\{c\}\_\{i\}, the two termsK​\(1\+η\)​c¯K\(1\+\\eta\)\\bar\{c\}cancel*identically inη\\eta*, leavingc′=r​c¯i\+\(1−r\)​W=c¯i−1c^\{\\prime\}=r\\bar\{c\}\_\{i\}\+\(1\-r\)W=\\bar\{c\}\_\{i\-1\}\. Mis\-scaling the deviation around the correct center therefore cannot bias the mean; only the additive bias can move it\. ∎

#### Proof of Theorem[5](https://arxiv.org/html/2608.06893#Thmtheorem5)

###### Proof\.

Work inzz\-coordinates\. Three direct computations give exact step quantities that do not depend on color at matchedzz\-grids\. The unperturbed contraction is a ratio of values ofw​\(z\):=z\+x​\(1−z\)w\(z\):=z\+x\(1\-z\), consistent with Lemma[4](https://arxiv.org/html/2608.06893#Thmlemma4), and the key identity is the second half of \([84](https://arxiv.org/html/2608.06893#S6.E84)\),

\(1−ri\)​KiAi=Δ​zizi,\\frac\{\(1\-r\_\{i\}\)K\_\{i\}\}\{A\_\{i\}\}=\\frac\{\\Delta z\_\{i\}\}\{z\_\{i\}\},\(105\)which is exact and independent ofx=v/Px=v/P\. The gain\-perturbed contraction is therefore

A^i=Ai​\(1\+ηi​Δ​zizi\)\.\\hat\{A\}\_\{i\}=A\_\{i\}\\Bigl\(1\+\\eta\_\{i\}\\frac\{\\Delta z\_\{i\}\}\{z\_\{i\}\}\\Bigr\)\.\(106\)
The bias enters at stepiias\(1−ri\)​βi\(1\-r\_\{i\}\)\\beta\_\{i\}and is multiplied by the downstream contractionsA^i−1​⋯​A^1\\hat\{A\}\_\{i\-1\}\\cdots\\hat\{A\}\_\{1\}on its way to the output, so its weight is

\(1−ri\)​∏m<iA^m=Δ​zizi​∏m<i\(1\+ηm​Δ​zmzm\),\(1\-r\_\{i\}\)\\prod\_\{m<i\}\\hat\{A\}\_\{m\}=\\frac\{\\Delta z\_\{i\}\}\{z\_\{i\}\}\\prod\_\{m<i\}\\Bigl\(1\+\\eta\_\{m\}\\frac\{\\Delta z\_\{m\}\}\{z\_\{m\}\}\\Bigr\),\(107\)where we used∏m<iAm=P/φ​\(ρi−1\)\\prod\_\{m<i\}A\_\{m\}=P/\\varphi\(\\rho\_\{i\-1\}\)from Lemma[4](https://arxiv.org/html/2608.06893#Thmlemma4)together with\(1−ri\)​P/φ​\(ρi−1\)=\(1−ri\)​Ki/Ai=Δ​zizi\(1\-r\_\{i\}\)P/\\varphi\(\\rho\_\{i\-1\}\)=\(1\-r\_\{i\}\)K\_\{i\}/A\_\{i\}=\\frac\{\\Delta z\_\{i\}\}\{z\_\{i\}\}from \([105](https://arxiv.org/html/2608.06893#S6.E105)\)\. Whenη≡0\\eta\\equiv 0the weight reduces toΔ​zizi\\frac\{\\Delta z\_\{i\}\}\{z\_\{i\}\}\. In general, therefore,

m0=∑iβi​Δ​zizi​∏m<i\(1\+ηm​Δ​zmzm\),m\_\{0\}=\\sum\_\{i\}\\beta\_\{i\}\\,\\frac\{\\Delta z\_\{i\}\}\{z\_\{i\}\}\\prod\_\{m<i\}\\Bigl\(1\+\\eta\_\{m\}\\frac\{\\Delta z\_\{m\}\}\{z\_\{m\}\}\\Bigr\),\(108\)which is \([10](https://arxiv.org/html/2608.06893#S3.E10)\)\. Unrolling the variance recursion \([35](https://arxiv.org/html/2608.06893#S6.E35)\) with the same factors gives

V0P=∑jzj−1​Δ​zjzj​∏m<j\(1\+ηm​Δ​zmzm\)2,\\frac\{V\_\{0\}\}\{P\}=\\sum\_\{j\}z\_\{j\-1\}\\frac\{\\Delta z\_\{j\}\}\{z\_\{j\}\}\\prod\_\{m<j\}\\Bigl\(1\+\\eta\_\{m\}\\frac\{\\Delta z\_\{m\}\}\{z\_\{m\}\}\\Bigr\)^\{2\},\(109\)which is \([11](https://arxiv.org/html/2608.06893#S3.E11)\)\. In the special caseη≡0\\eta\\equiv 0this becomes

V0P=∑jzj−1​Δ​zjzj=1−∑j\(Δ​zj\)2zj,\\frac\{V\_\{0\}\}\{P\}=\\sum\_\{j\}z\_\{j\-1\}\\frac\{\\Delta z\_\{j\}\}\{z\_\{j\}\}=1\-\\sum\_\{j\}\\frac\{\(\\Delta z\_\{j\}\)^\{2\}\}\{z\_\{j\}\},\(110\)usingzj−1=zj−Δ​zjz\_\{j\-1\}=z\_\{j\}\-\\Delta z\_\{j\}and∑jΔ​zj=1\\sum\_\{j\}\\Delta z\_\{j\}=1, which recovers \([83](https://arxiv.org/html/2608.06893#S6.E83)\)\.

Every factor in \([108](https://arxiv.org/html/2608.06893#S6.E108)\)–\([109](https://arxiv.org/html/2608.06893#S6.E109)\) depends only on thezz\-grid and the error profile\. The terminal law is therefore identical for everyv\>0v\>0, and with mean error the terminal KL is12​\(u\+m02/P−1−ln⁡u\)\\tfrac\{1\}\{2\}\(u\+m\_\{0\}^\{2\}/P\-1\-\\ln u\)withu=V0/Pu=V\_\{0\}/Pby \(S6\)\. ∎

#### Proof of Proposition[5](https://arxiv.org/html/2608.06893#Thmproposition5)

###### Proof\.

At levelzzthe conditional correlation satisfies

corr2⁡\(x0,deviation∣x1\)=1−z\\operatorname\{corr\}^\{2\}\\bigl\(x\_\{0\},\\;\\text\{deviation\}\\mid x\_\{1\}\\bigr\)=1\-z\(111\)exactly, with symbolic residual0\. Up to scale, the per\-mode joint law of\(x0,xt\)\(x\_\{0\},x\_\{t\}\)givenx1x\_\{1\}therefore depends only onzz, so any scale\-equivariant learner has a color\-free relative\-error law at eachzz\-level\. For zero\-intercept least squares fromnnpairs on a normalized design∑iXi2=n​τ2\\sum\_\{i\}X\_\{i\}^\{2\}=n\\tau^\{2\},

Var⁡\(K^/K−1\)=Mn​Σ​K2=zn​\(1−z\),\\operatorname\{Var\}\\bigl\(\\hat\{K\}/K\-1\\bigr\)=\\frac\{M\}\{n\\Sigma K^\{2\}\}=\\frac\{z\}\{n\(1\-z\)\},\(112\)which is again color\-free; under a random designnnis replaced byn−2n\-2\. Combined with Theorem[5](https://arxiv.org/html/2608.06893#Thmtheorem5), per\-mode color is exactly invisible end to end once thezz\-schedule is fixed\. ∎

Table 10:Constant gain error±η\\pm\\etaon a fixedρ\\rho\-gridmoves the apparent optimum by orders of magnitude — an artifact of the inducedzz\-schedule, per Remark[5](https://arxiv.org/html/2608.06893#Thmremark5)\.

#### Ridge whitening \(Proposition[6](https://arxiv.org/html/2608.06893#Thmproposition6)\)

A shared ridge penaltyλ\\lambdashrinks the estimated gain through the relative bias

ηk​\(z\)=−λλ\+n​Σk​\(z\)\.\\eta\_\{k\}\(z\)=\-\\frac\{\\lambda\}\{\\lambda\+n\\Sigma\_\{k\}\(z\)\}\.\(114\)BecauseΣk\\Sigma\_\{k\}is dimension*ful*, this penalty breaks the scale symmetry of Theorem[3](https://arxiv.org/html/2608.06893#Thmtheorem3)\. Exact\-chain optimization withT=100T\{=\}100andλ=1\\lambda\{=\}1gives the sweep in Table[11](https://arxiv.org/html/2608.06893#S6.T11), which decreases strongly with the effective sample sizen​PnP\.

Table 11:Ridge\-whitening sweep\.Optimal scalex∗​\(n​P\)x^\{\*\}\(nP\)forT=100T=100andλ=1\\lambda=1\. Smaller effective sample sizen​PnPproduces a flatter optimal spectrum\.ForP=\(1,4\)P=\(1,4\)andn=100n=100the optimum satisfies

v2∗v1∗=4⋅x∗​\(400\)x∗​\(100\)=1\.99,\\frac\{v\_\{2\}^\{\*\}\}\{v\_\{1\}^\{\*\}\}=4\\cdot\\frac\{x^\{\*\}\(400\)\}\{x^\{\*\}\(100\)\}=1\.99,\(115\)against the proportional value44: weakly observed frequencies need disproportionately*more*reference noise to counteract shrinkage, and the regularized optimum is substantially flatter\. With a shared*constant*λ\\lambdathe modes remain separable, because the mechanism is per\-mode scale sensitivity\. A genuinely shared network introduces true cross\-mode coupling; the simplest such model ties one gain across modes, and its exact cost equals the curvature\-weighted dispersion of the per\-mode optimal gains\. The fully nonlinear shared\-network case remains open\.

##### Data\-dependent optima\.

Using exact moment recursions and no Monte Carlo, the leading\-order statistical contribution equals1/n1/ntimes the discretization functional; measured/predicted is0\.9710\.971and0\.9790\.979atx=0\.3x=0\.3andx=2\.0x=2\.0\. The mean effect therefore does not move the argmin\. The shift is caused instead by the variance\-of\-variance penalty, which produces the mild downward trend

x∗=0\.323​\(n=∞\)→0\.305​\(n=300\)→0\.272​\(n=100\)\.x^\{\*\}=0\.323\\;\(n\{=\}\\infty\)\\;\\to\\;0\.305\\;\(n\{=\}300\)\\;\\to\\;0\.272\\;\(n\{=\}100\)\.\(116\)

#### Distortion–perception and underdispersion correction

###### Corollary 6\(Distortion–perception trade\-off, derived\)\.

The posterior\-mean estimatorW​x1Wx\_\{1\}has conditional MSE exactlyPPfor everyTT, independent of the reference; this is the only quantity made invariant by mean exactness \(Lemma[3](https://arxiv.org/html/2608.06893#Thmlemma3)\)\. The sampled output has the correct mean and is conditionally independent of the truex0x\_\{0\}givenx1x\_\{1\}, so its conditional MSE is

MSE=P\+V0∈\[P,2​P\),\\mathrm\{MSE\}=P\+V\_\{0\}\\;\\in\\;\[P,\\,2P\),\(117\)which depends on the reference throughV0V\_\{0\}alone\. The collapsed sampler minimizes it atV0=0V\_\{0\}=0, where it is the posterior\-mean estimator in disguise; at distributional perfectionV0=PV\_\{0\}=Pand the MSE equals2​P2P\. The classical factor\-2 \(33dB\) cost of posterior sampling therefore follows as a theorem, and the reference traces the entire trade\-off curve\.

###### Corollary 7\(Principled underdispersion correction\)\.

By Corollary[2](https://arxiv.org/html/2608.06893#Thmcorollary2)the plug\-in sampler’s only failure mode is underdispersion, so deliberate gain inflation can serve as a calibrated correction: by \([109](https://arxiv.org/html/2608.06893#S6.E109)\), chooseη​\(z\)\\eta\(z\)so that

∑jzj−1​Δ​zjzj​∏m<j\(1\+ηm​Δ​zmzm\)2=1\.\\sum\_\{j\}z\_\{j\-1\}\\frac\{\\Delta z\_\{j\}\}\{z\_\{j\}\}\\prod\_\{m<j\}\\Bigl\(1\+\\eta\_\{m\}\\frac\{\\Delta z\_\{m\}\}\{z\_\{m\}\}\\Bigr\)^\{2\}=1\.\(118\)Each term carries its own partial product; there is no single global factor multiplying theη≡0\\eta\\equiv 0variance\. A convenient one\-parameter solution holdsη\\etaconstant and solves the resulting scalar equation, giving an exact\-formula counterpart to temperature and churn heuristics\.

Note that an objective that checks only the expected variance can be satisfied by*any*error that adds variance\. The appropriate objective is𝔼​\[KL\]\\mathbb\{E\}\[\\mathrm\{KL\}\]over error realizations\.

## 7Experimental Protocols & Additional Results

### 7\.1Reference color inI2I^\{2\}SB super\-resolution

We begin with the question that motivated the theory: does reference color matter for a full pretrained restoration model? Starting from the stockI2I^\{2\}SB4×4\\timessuper\-resolution checkpoint, we fine\-tune one model per reference for5050k steps under identical optimization settings, changing only the color of the reference noise — white, low\-pass \(1/f21/f^\{2\}\), and high\-pass \(f2f^\{2\}\)\. All three references share the same scaleε\\varepsilonand the same total pixel variance, which removes the trivial confound; only the allocation of variance across frequencies differs\. Since4×4\\timesdownsampling destroys high spatial frequencies, the high\-pass reference is the one that concentrates noise on the destroyed information — the qualitative direction ofvk∝Pkv\_\{k\}\\propto P\_\{k\}\.

Table 12:4×4\\timessuper\-resolution with recolored references\.PSNR/SSIM on held\-out aligned pairs at NFE2020; equal data, compute and total reference variance\. “Best” selects the best validation checkpoint; “final” is at5050k steps\.Table[12](https://arxiv.org/html/2608.06893#S7.T12)shows the ordering predicted by the destroyed\-information picture: high\-pass\>\>white\>\>low\-pass on both metrics, at both the best checkpoint and the end of training\. The high\-pass reference beats white by1\.51\.5dB PSNR and0\.140\.14SSIM at the best checkpoint, despite the base model having been pretrained with white noise\. The low\-pass reference is more striking: its best checkpoint is its*first*one, and continued fine\-tuning is monotonically destructive, losing5\.05\.0dB by5050k steps\. A reference that concentrates noise on frequencies the observation already preserves does not merely learn slowly; it steadily erases the pretrained model’s ability to reconstruct\.

This experiment changes only the color of the noise, at a single scaleε\\varepsilonand a single step budget\. It confirms the coarse direction the theory predicts — noise belongs where the observation destroyed information — but it does not test the predicted amountx∗​\(T\)x^\{\*\}\(T\), which requires sweeping the scale \(Fig\.[1](https://arxiv.org/html/2608.06893#S1.F1)does this in the exact setting\)\. Even within these limits the effect is large: on a real pretrained restoration model the reference spectrum is a first\-order design choice, not an implementation detail\.

### 7\.2Protocols

#### Exact numerics \(Phase 1\)

Every quantity is computed three ways — the closed\-form deficit \([41](https://arxiv.org/html/2608.06893#S6.E41)\), the telescopedV0V\_\{0\}, and the exact affine recursion \([34](https://arxiv.org/html/2608.06893#S6.E34)\)–\([35](https://arxiv.org/html/2608.06893#S6.E35)\) — and the three are cross\-checked to machine precision before each run\. The fixed regression anchor isD0=1\.171587D\_\{0\}=1\.171587atP=v=4P=v=4andT=10T=10\.

Figure[2](https://arxiv.org/html/2608.06893#S3.F2)\(a\) usesP=v=4P=v=4on uniform and power grids withTTup to10510^\{5\}\. Figure[2](https://arxiv.org/html/2608.06893#S3.F2)\(c\) locatesx∗​\(T\)x^\{\*\}\(T\)by golden\-section minimization ofΦ\\Phiat2222logarithmically spacedT∈\[5,105\]T\\in\[5,10^\{5\}\]\. Figure[2](https://arxiv.org/html/2608.06893#S3.F2)e uses the two\-cluster grid atT=48T\{=\}48; each candidate minimum is re\-verified with exact derivatives in 50\-digitmpmatharithmetic\. Table[13](https://arxiv.org/html/2608.06893#S7.T13)uses the exact recursion with budget∑kvk=5\\sum\_\{k\}v\_\{k\}=5, a linear schedule and vertexϵ=10−2\\epsilon=10^\{\-2\}\. Thezz\-schedule computations use the forward recurrence \([85](https://arxiv.org/html/2608.06893#S6.E85)\)\. All Phase\-1 experiments are deterministic, use no random seeds, and run in minutes on a laptop\.

#### Gaussian Setting \(Phase 2\)

The synthetic setting has6464modes \(up to256256in secondary sweeps\), priorSk=k−2S\_\{k\}=k^\{\-2\}, Gaussian MTFhk=exp⁡\(−\(k/kc\)2\)h\_\{k\}=\\exp\(\-\(k/k\_\{c\}\)^\{2\}\)withkc=16k\_\{c\}=16, and flat noiseNk=10−3N\_\{k\}=10^\{\-3\};PkP\_\{k\}is computed exactly and saved with every run\. The matched reference isvk=x∗​\(T\)​Pkv\_\{k\}=x^\{\*\}\(T\)P\_\{k\}withx∗x^\{\*\}from Table[6](https://arxiv.org/html/2608.06893#S6.T6); white, anti\-matched \(∝1/Pk\\propto 1/P\_\{k\}\) and prior\-colored \(∝Sk\\propto S\_\{k\}\) references are rescaled to the matched total budget\.

The predictor is6464*independent*per\-mode MLPs \(batched withbmm\), each receiving\(xt,x1,32\-dim sinusoidal​t\)\(x\_\{t\},x\_\{1\},\\text\{32\-dim sinusoidal \}t\), with33SiLU hidden layers and scalar outputx^0\\hat\{x\}\_\{0\}\. Training: Adam10−310^\{\-3\}, batch20482048–40964096,1212k–2020k steps, EMA0\.9990\.999, continuoust∼U​\(0,1\)t\\sim U\(0,1\), per\-mode I/O normalization bySk\\sqrt\{S\_\{k\}\}\(a benign per\-mode rescaling\), seeds\{0,1,2\}\\\{0,1,2\\\}, common random numbers across references\.

Terminal KL is estimated as follows: conditional onx1x\_\{1\}, the terminal law of each mode is fit by regression over a common\-random\-number batch of10510^\{5\}samples\. The KL divergence between the fitted Gaussian and𝒩​\(W​x1,P\)\\mathcal\{N\}\(Wx\_\{1\},P\)is then evaluated in closed form; the empirical KL floor is about10−510^\{\-5\}\. The ridge\-whitening and constant\-sweep experiments use closed\-form ridge/OLS learners per\(k,t​\-bin\)\(k,t\\text\{\-bin\}\)with3232bins and exact deterministic chain propagation, hence zero sampling noise\. The Jacobian probe estimatesη^​\(z\)\\hat\{\\eta\}\(z\)by finite differences of∂x^0/∂xt\\partial\\hat\{x\}\_\{0\}/\\partial x\_\{t\}on the trained fingerprint network and applies \([109](https://arxiv.org/html/2608.06893#S6.E109)\)\.

#### FFHQ64×6464\{\\times\}64\(Phase 3\)

FFHQ is center\-cropped and resized to64×6464\{\\times\}64with Lanczos interpolation; the first6060k images are used for training, for fittingS​\(k\)S\(k\)and for FID statistics, and the next1010k for evaluation; pixels lie in\[−1,1\]\[\-1,1\]\.

The degradation is defined in the DFT domain so thatP​\(k\)P\(k\)is analytic:x1=kblur∗x0\+nx\_\{1\}=k\_\{\\mathrm\{blur\}\}\*x\_\{0\}\+nwith Gaussian blurσblur=2\.0\\sigma\_\{\\mathrm\{blur\}\}=2\.0px, transfer functionh​\(k\)=exp⁡\(−2​π2​σblur2​\|f\|2\)h\(k\)=\\exp\(\-2\\pi^\{2\}\\sigma\_\{\\mathrm\{blur\}\}^\{2\}\|f\|^\{2\}\), and noiseσn=0\.05\\sigma\_\{n\}=0\.05, i\.e\.N=2\.5×10−3N=2\.5\\times 10^\{\-3\}\. There is no decimation and hence no aliasing block, so Corollary[4](https://arxiv.org/html/2608.06893#Thmcorollary4)is not needed here\.S​\(k\)S\(k\)is fit once by radially averaging the training\-split power spectrum and saved in a bundle together withP​\(k\)P\(k\),W​\(k\)W\(k\),h​\(k\)h\(k\)andx∗​\(T\)x^\{\*\}\(T\)\. The matched reference usesx∗​\(Tref=50\)=0\.404x^\{\*\}\(T\_\{\\mathrm\{ref\}\}\{=\}50\)=0\.404; all budget\-matched colors have the same per\-pixel noise variance, so “same noise level, different spectrum” is literal\.

Each run trains oneI2I^\{2\}SB\-style ADM U\-Net \(base width128128, multipliers\(1,2,2,2\)\(1,2,2,2\),22residual blocks per resolution, attention at16216^\{2\}and828^\{2\};39\.639\.6M parameters\), conditioned by concatenatingx1x\_\{1\}withxtx\_\{t\}\(66input channels\) plus a standardtt\-embedding\.

*Original recipe:*x^0\\hat\{x\}\_\{0\}MSE withρ∼U​\(0,1\)\\rho\\sim U\(0,1\)on the uniform schedule; Adam10−410^\{\-4\}, batch128128,150150k steps, EMA0\.99990\.9999, dropout0\.10\.1, horizontal flips, bf16 autocast\. Matched and white use seeds\{0,1,2\}\\\{0,1,2\\\}; anti,α=1\\alpha\{=\}1,α=2\\alpha\{=\}2and equal\-budget matched use seed0—1010runs, each44–77h on one modern GPU\.

Evaluation uses the ancestral plug\-in\-mean sampler at NFE55,1010,2020,5050and100100on uniform sub\-grids with fixed generator seeds and chunking, so comparisons across references use byte\-identical common random numbers\. We report PSNR/SSIM/LPIPS on11k held\-out images and FID against FFHQ\-64 statistics; radially averaged error spectra include the exact\-drift overlay2​P​\(k\)−D0​\(k\)2P\(k\)\-D\_\{0\}\(k\)\. Every55k steps the EMA network is probed on256256fixed images over a1616\-pointtt\-grid, recording the per\-\(radial band,tt\) fingerprint relative to the Bayes floor\.

#### Converged recipe and theθ\\theta\-family

The converged recipe is identical to the original except dropout0, batch size512512and300300k steps — eight times as many gradient samples\. Theθ\\theta\-family isv​\(k\)∝P​\(k\)θv\(k\)\\propto P\(k\)^\{\\theta\}rescaled to the fixed total budgetAA, withθ∈\{0,14,12,34,1\}\\theta\\in\\\{0,\\tfrac\{1\}\{4\},\\tfrac\{1\}\{2\},\\tfrac\{3\}\{4\},1\\\}; the endpoints coincide with the white and matched references at that budget \(verified against the saved bundle to10−610^\{\-6\}, so separate endpoint runs are unnecessary\)\. Runs: matched and white with seeds\{0,…,4\}\\\{0,\\dots,4\\\};θ∈\{0\.25,0\.5,0\.75\}\\theta\\in\\\{0\.25,0\.5,0\.75\\\}with seed0; matched and white atσblur=4\.0\\sigma\_\{\\mathrm\{blur\}\}=4\.0with seed0\(separate bundle\)\.

*FID protocol\.*Every FID in Section[4\.4](https://arxiv.org/html/2608.06893#S4.SS4)uses5050k samples, formed by pooling55replicates of the1010k evaluation set with fresh sampler driving noise; per\-replicate1010k FIDs are stored, and the original\-recipe runs were re\-evaluated under the same protocol before any comparison\.

*Schedule ablation\.*Every trained network is also evaluated on thezz\-optimal grid of Theorem[4](https://arxiv.org/html/2608.06893#Thmtheorem4)\(c\) without retraining\.

Predictions P1–P4 of Section[7\.6](https://arxiv.org/html/2608.06893#S7.SS6)were written down before the first converged\-recipe run began\.

### 7\.3Exact numerical evaluation

##### Numerical verification\.

All deterministic\-recursion self\-tests pass to machine precision\. For the invisibility experiment \(Fig\.[2](https://arxiv.org/html/2608.06893#S3.F2)a\) the terminal variance reachesV0=3\.9995V\_\{0\}=3\.9995atT=105T=10^\{5\}forP=v=4P=v=4, while the singular casev=0v=0collapses, as Theorem[2](https://arxiv.org/html/2608.06893#Thmtheorem2)\(b\) requires\. For the scaling experiment \(Fig\.[2](https://arxiv.org/html/2608.06893#S3.F2)c\) the measured optima includex∗​\(217\)=0\.329x^\{\*\}\(217\)=0\.329andx∗​\(105\)=0\.212x^\{\*\}\(10^\{5\}\)=0\.212, and the ratiox∗​\(T\)/\(2​ln⁡T\)−1/2x^\{\*\}\(T\)/\(2\\ln T\)^\{\-1/2\}decreases from1\.291\.29atT=5T=5to1\.0171\.017atT=105T=10^\{5\}; Table[6](https://arxiv.org/html/2608.06893#S6.T6)gives the full sweep\. For the clusteredT=48T=48schedule \(Fig\.[2](https://arxiv.org/html/2608.06893#S3.F2)e\) seven local minima are certified in 50\-digit arithmetic\.

Table 13:Finite\-step KL in a two\-mode setting\.Total KL∑kKLk\\sum\_\{k\}\\mathrm\{KL\}\_\{k\}forP=\(1,4\)P=\(1,4\), total budget55and a linear schedule\. The matched allocation is best at every tested step count\. The vertex allocation usesv2=ϵv\_\{2\}=\\epsilonwithϵ=10−2\\epsilon=10^\{\-2\}and is the allocation selected by the degenerate Bayes\-risk objective of App[6\.3](https://arxiv.org/html/2608.06893#S6.SS3.SSSx3)\.Table[13](https://arxiv.org/html/2608.06893#S7.T13)shows that the matched allocation minimizes KL at every testedTT; atT=200T=200the vertex allocation has294×294\\timeslarger KL\.

##### Schedule and budget effects\.

For thezz\-schedule recurrence,F∗​T/4F^\{\*\}T/4equals0\.3750\.375,0\.6940\.694,0\.9660\.966and0\.9990\.999atT=2T=2,1010,200200and10410^\{4\}\(Table[9](https://arxiv.org/html/2608.06893#S6.T9)\), approaching the4​P/T4P/Tfloor of Theorem[4](https://arxiv.org/html/2608.06893#Thmtheorem4)d\. After schedule optimization the deficit is invariant to reference color up to10−1210^\{\-12\}, as Theorem[4](https://arxiv.org/html/2608.06893#Thmtheorem4)a requires\. The budgeted optimum bends toward theP2P^\{2\}allocation predicted by Proposition[3](https://arxiv.org/html/2608.06893#Thmproposition3):v1∗v\_\{1\}^\{\*\}moves from0\.7730\.773atT=5T=5to0\.3470\.347atT=103T=10^\{3\}and0\.3060\.306atT=2×104T=2\\times 10^\{4\}, approaching the vertex value5/17≈0\.2945/17\\approx 0\.294\(Table[8](https://arxiv.org/html/2608.06893#S6.T8)\)\.

### 7\.4Controlled Gaussian experiments

##### Tested predictions\.

We test whether: \(i\) the converged loss matches the Bayes floorP​ϕ/\(ϕ\+P\)P\\phi/\(\\phi\+P\)withϕ=v​ρ/\(1−ρ\)\\phi=v\\rho/\(1\-\\rho\); \(ii\) the predicted finite\-NFE reference ordering appears; \(iii\) the optimal scale follows the theoretical constant and bends towardP2P^\{2\}under an equal\-budget constraint; \(iv\) finite data and weight decay flatten the optimum toward white; and \(v\) a Jacobian\-based estimate from one trained network predicts terminal variance without retraining\.

##### Loss fingerprint and NFE dependence\.

After1212k training steps the median relative excess over the Bayes floor is0\.18%0\.18\\%across the64×3264\\times 32grid of mode–time pairs, below the pre\-specified5%5\\%gate\. For every NFE up to5050the KL ordering is matched<<white<<anti\-matched<<prior\-colored; at NFE5050the total KL is0\.19±0\.030\.19\{\\pm\}0\.03,0\.21±0\.030\.21\{\\pm\}0\.03,0\.44±0\.040\.44\{\\pm\}0\.04and4\.3±0\.34\.3\{\\pm\}0\.3over three seeds\.

At NFE200200gain error dominates discretization error: the learned curves lie one to two orders of magnitude above their exact\-drift floors, the prior\-colored reference becomes unstable \(KL10410^\{4\}–10510^\{5\}\), and matched and white are statistically indistinguishable at0\.60±0\.160\.60\{\\pm\}0\.16and0\.54±0\.150\.54\{\\pm\}0\.15\. This regime therefore does not preserve the finite\-step ordering predicted under exact drift — exactly as Theorem[5](https://arxiv.org/html/2608.06893#Thmtheorem5)anticipates\.

##### Optimal scale and ridge whitening\.

The measured optimal constants are0\.5750\.575atT=10T=10and0\.4020\.402atT=50T=50, against the theoretical values0\.5770\.577and0\.4040\.404of Table[6](https://arxiv.org/html/2608.06893#S6.T6)\. AtT=200T=200the measured optimum is0\.3600\.360, which is8%8\\%above the exact\-drift value0\.3320\.332— the direction predicted for finite\-data regularization by Proposition[6](https://arxiv.org/html/2608.06893#Thmproposition6)\. In the exact\-chain ridge sweep atT=100T=100andλ=1\\lambda=1\(relative bias \([114](https://arxiv.org/html/2608.06893#S6.E114)\)\), the optimal scalex∗​\(n​P\)x^\{\*\}\(nP\)grows monotonically as the effective sample size shrinks — from0\.3630\.363atn​P=∞nP=\\inftythrough0\.580\.58atn​P=103nP=10^\{3\}to11\.9211\.92atn​P=10nP=10— i\.e\. a smaller effective sample size produces a flatter optimal spectrum\. ForP=\(1,4\)P=\(1,4\)andn=100n=100the ridge\-adjusted optimum givesv2∗/v1∗=1\.99v\_\{2\}^\{\*\}/v\_\{1\}^\{\*\}=1\.99against44under exact proportionality: the regularized optimum is substantially flatter\.

##### Jacobian and schedule probes\.

The Jacobian analysis of the trained fingerprint network predicts terminal variance with0\.6%0\.6\\%median relative error and needs no retraining, validating the measurement protocol of Section[3\.4](https://arxiv.org/html/2608.06893#S3.SS4)\. Per\-mode schedule adaptation reduces reference\-color gaps toward the theoretical floor, supporting the color\-blindness prediction of Theorem[4](https://arxiv.org/html/2608.06893#Thmtheorem4)a once schedule error is removed\.

### 7\.5FFHQ: evaluation details and secondary results

##### Data and metrics\.

All FFHQ runs use common random numbers across reference choices\. PSNR, SSIM and LPIPS are computed on the same1,0001\{,\}000held\-out images\. In the original\-recipe experiments FID is computed from5050k restorations against Inception statistics of the6060k clean training images; the converged\-recipe experiments use5050k restorations \(Section[7\.6](https://arxiv.org/html/2608.06893#S7.SS6)\)\. The matched reference is computed from the known Gaussian MTF and flat observation\-noise spectrum; it is not selected using validation performance\.

##### Budget conventions\.

Most rows of Table[1](https://arxiv.org/html/2608.06893#S4.T1)use total budgetA=∑kx∗​P​\(k\)A=\\sum\_\{k\}x^\{\*\}P\(k\)\. The “matched, equal budget” row instead uses the optimal\-white budgetBB\. At NFE5050the equal\-budget matched model has FID9\.129\.12against8\.458\.45for matched under its free budget — the budget effect predicted by Proposition[3](https://arxiv.org/html/2608.06893#Thmproposition3)\. We report both conventions because they answer different questions\. The flatter empirical optimum is also not specific to NFE5050: at NFE100100, white attains FID6\.076\.07against6\.646\.64for matched, beyond the observed seed variation\.

##### Schedule re\-evaluation\.

Each trained model is re\-evaluated with thezz\-optimal schedule of Theorem[4](https://arxiv.org/html/2608.06893#Thmtheorem4), without retraining\. The optimized schedule improves every reference by a similar amount and leaves the ordering unchanged \(at NFE5050: white7\.87→7\.107\.87\\to 7\.10, matched8\.45→7\.628\.45\\to 7\.62, anti\-matched14\.1→13\.014\.1\\to 13\.0\); the white–matched FID gap changes only from0\.580\.58to0\.530\.53\. It therefore removes a shared discretization component but does not explain the color\-dependent gap\.

### 7\.6Training regime and reference exponent

##### FID measurement\.

For each converged\-recipe run, five independent groups of1010k restorations are generated with fresh sampler noise and their Inception features pooled into a5050k estimate; the median per\-group spread is±0\.04\\pm 0\.04and the maximum±0\.08\\pm 0\.08\. For direct comparison with the original1010k protocol, per\-group1010k FIDs are also reported\. Pooling does not change the conclusions: at NFE5050the matched–white gap is1\.1691\.169on the1010k basis and1\.1771\.177after pooling\.

##### Pre\-specified prediction ledger\.

Before training we recorded four directional predictions:

- P1\.Better optimization should shrink the matched–white FID and LPIPS gaps, particularly at NFE55–1010, and may reverse their sign\.
- P2\.Theθ\\theta\-family should have an interior perceptual optimum, and its minimizingθ\\thetashould increase under the converged recipe\.
- P3\.Anti\-matched should remain substantially worse on perceptual metrics in every tested regime\.
- P4\.Increasing the blur toσblur=4\.0\\sigma\_\{\\mathrm\{blur\}\}=4\.0should increase the magnitude of the matched–white effect, regardless of direction\.

These predictions concern the direction in which the optimum moves; they do not assume that matched must become best\.

![Refer to caption](https://arxiv.org/html/2608.06893v1/x8.png)Figure 8:Reference exponent under two training recipes\.FID \(left\) and LPIPS \(right\) against the color exponentθ\\theta\(v∝Pkθv\\propto P\_\{k\}^\{\\theta\}, fixed budget\); stars mark the argmin per NFE\.*Top:*converged recipe —θ=0\.25\\theta\{=\}0\.25is best only for NFE≤20\\leq 20, and white is best for NFE≥50\\geq 50\.*Bottom:*original recipe \(endpoints only; theα\\alpha\-suite is not in thePkθP\_\{k\}^\{\\theta\}family\)\. Better optimization widens, rather than closes, the white–matched FID gap\.
##### P1 and P2\.

The converged recipe improves both endpoint references in absolute terms: at NFE5050, white improves from FID7\.877\.87to4\.344\.34and matched from8\.458\.45to5\.505\.50\.

P1 is nevertheless refuted\. The matched–white gap*increases*from1\.781\.78to2\.092\.09at NFE55, from1\.361\.36to2\.432\.43at NFE1010, and from0\.580\.58to1\.181\.18at NFE5050— at NFE5050roughly2727standard errors, given seed spreads of±0\.09\\pm 0\.09\(matched\) and±0\.02\\pm 0\.02\(white\) on the pooled estimate\. The LPIPS gap increases from0\.00280\.0028to0\.00550\.0055at NFE1010and from0\.00070\.0007to0\.00320\.0032at NFE5050\.

P2 receives limited support only at small NFE: FID is minimized atθ=0\.25\\theta=0\.25for NFE≤20\\leq 20\(Fig\.[8](https://arxiv.org/html/2608.06893#S7.F8)\), and at NFE55its value is14\.8714\.87against15\.27±0\.0715\.27\{\\pm\}0\.07for white\. For NFE≥50\\geq 50FID returns to a white optimum, and LPIPS selects white for NFE≥10\\geq 10\. A direct cross\-recipe comparison of the minimizingθ\\thetais not identifiable, because the original recipe did not include interiorθ\\thetavalues; the identifiable endpoint gap moves opposite to the prediction\.

##### P3 and P4\.

P3 is supported in the original regime under both schedule grids: at NFE5050anti\-matched has FID14\.114\.1under the uniform schedule and13\.013\.0under thezz\-optimal schedule, the worst perceptual reference in both cases\. Anti\-matched was not retrained under the converged recipe, so P3 is not evaluated there\.

P4 is supported in magnitude, in the direction favoring white: atσblur=4\.0\\sigma\_\{\\mathrm\{blur\}\}=4\.0the white–matched FID gap at NFE5050is2\.802\.80\(6\.456\.45versus9\.259\.25; one seed per reference, pooled5050k FID\), against1\.181\.18atσblur=2\.0\\sigma\_\{\\mathrm\{blur\}\}=2\.0on the same basis\. The corresponding LPIPS gaps are0\.00660\.0066and0\.00320\.0032\.

##### Convergence audit and scope\.

The loss fingerprint \(Fig\.[9](https://arxiv.org/html/2608.06893#S7.F9)\) shows that the converged recipe reduces but does not eliminate low\-frequency optimization error: in the lowest radial band the median relative excess above the Bayes floor decreases from1\.37×1\.37\\timesat150150k steps to1\.05×1\.05\\timesat300300k for matched, and from0\.73×0\.73\\timesto0\.60×0\.60\\timesfor white, while mid\- and high\-frequency bands remain at the floor\. Under the pre\-specified criterion the converged recipe is better optimized but has not reached the exact\-drift regime\.

Because low\-frequency fingerprint error persists at300300k steps, the ridge\-whitening law remains untested at true convergence\. The small\-NFE optimum atθ=0\.25\\theta=0\.25shows that limitedPkP\_\{k\}\-coloring can help\.

![Refer to caption](https://arxiv.org/html/2608.06893v1/x9.png)Figure 9:Convergence audit \(converged recipe\)\.Per\-\(radial band,tt\) training loss relative to the Bayes floor at300300k steps\. The low\-frequency excess shrinks relative to the original recipe but persists; mid and high frequencies sit at the floor\.Table 14:FFHQ under the converged recipeat NFE5050\(dropout0, batch size512512,300300k steps; all other settings as in Table[1](https://arxiv.org/html/2608.06893#S4.T1)\)\. FID uses5050k pooled samples\. White and matched report mean±\\pmstd over five seeds; interiorθ\\thetavalues use one seed\.

### 7\.7Dataset replication: CelebA

##### Protocol\.

To test whether the white\-first ordering is specific to FFHQ, we replicate the converged recipe on CelebA at64×6464\{\\times\}64under the identical degradation \(σblur=2\.0\\sigma\_\{\\mathrm\{blur\}\}=2\.0,σn=0\.05\\sigma\_\{n\}=0\.05\), withS​\(k\)S\(k\)and the bundle re\-fit on the CelebA training split; white, matched,θ=0\.25\\theta\{=\}0\.25and anti\-matched are trained with one seed each\. Table[15](https://arxiv.org/html/2608.06893#S7.T15)reports FID, PSNR and LPIPS at every evaluated NFE\.

Table 15:CelebA replication\(converged recipe, one seed per reference; FID uses5050k pooled samples\)\. Best per column in bold within each block\. White has the best FID and anti\-matched the best PSNR \(SSIM tracks PSNR\) at every NFE; anti\-matched has the worst LPIPS and FID at every NFE\.
##### Results\.

Both FFHQ findings replicate \(Table[15](https://arxiv.org/html/2608.06893#S7.T15)\)\. The distortion–perception split of Corollary[6](https://arxiv.org/html/2608.06893#Thmcorollary6)holds at every NFE: anti\-matched has the best PSNR/SSIM and the worst LPIPS/FID throughout\. The white\-first FID ordering also holds at every NFE, and at NFE5050the matched–white gap of1\.121\.12is comparable to FFHQ’s\. On CelebA white already has the best FID at NFE55, so the small\-NFE interior optimum of Fig\.[8](https://arxiv.org/html/2608.06893#S7.F8)appears dataset\-dependent\.

## References

- \[1\]T\.W\. Anderson\(2003\)An introduction to multivariate statistical analysis\.Wiley Series in Probability and Statistics,Wiley\.External Links:ISBN 9780471360919,LCCN 20234317,[Link](https://books.google.com/books?id=Cmm9QgAACAAJ)Cited by:[Table 4](https://arxiv.org/html/2608.06893#S6.T4.7.7.7.7.7)\.
- \[2\]S\. Basu, R\. Pollack, and M\. Roy\(2006\)Algorithms in real algebraic geometry\.Springer\.Cited by:[§6\.4](https://arxiv.org/html/2608.06893#S6.SS4.SSSx4.Px3.p1.8),[Table 4](https://arxiv.org/html/2608.06893#S6.T4.18.18.3.1.1)\.
- \[3\]R\. Benita, M\. Elad, and J\. Keshet\(2026\)Spectral analysis of diffusion models with application to schedule design\.Advances in Neural Information Processing Systems38,pp\. 2073–2127\.Cited by:[§1](https://arxiv.org/html/2608.06893#S1.p2.1),[§2](https://arxiv.org/html/2608.06893#S2.SS0.SSS0.Px4.p1.1)\.
- \[4\]Y\. Blau and T\. Michaeli\(2018\)The perception\-distortion tradeoff\.InProceedings of the IEEE conference on computer vision and pattern recognition,pp\. 6228–6237\.Cited by:[§2](https://arxiv.org/html/2608.06893#S2.SS0.SSS0.Px5.p1.1)\.
- \[5\]C\. Bunne, Y\. Hsieh, M\. Cuturi, and A\. Krause\(2023\)The schrödinger bridge between gaussian measures has a closed form\.InInternational Conference on Artificial Intelligence and Statistics,pp\. 5802–5833\.Cited by:[§2](https://arxiv.org/html/2608.06893#S2.SS0.SSS0.Px3.p1.1)\.
- \[6\]M\. Chertkov\(2026\)Analytic bridge diffusions for controlled path generation\.arXiv preprint arXiv:2605\.02961\.Cited by:[§2](https://arxiv.org/html/2608.06893#S2.SS0.SSS0.Px3.p1.1)\.
- \[7\]T\.M\. Cover and J\.A\. Thomas\(2012\)Elements of information theory\.Wiley\.External Links:ISBN 9781118585771,LCCN 2005047799,[Link](https://books.google.com/books?id=VWq5GG6ycxMC)Cited by:[Table 4](https://arxiv.org/html/2608.06893#S6.T4.22.22.4.4.4)\.
- \[8\]H\. Davidson, N\. Issachar, and S\. Benaim\(2026\)Colored noise diffusion sampling\.arXiv preprint arXiv:2605\.30332\.Cited by:[§2](https://arxiv.org/html/2608.06893#S2.SS0.SSS0.Px4.p1.1)\.
- \[9\]V\. De Bortoli, J\. Thornton, J\. Heng, and A\. Doucet\(2021\)Diffusion Schrödinger bridge with applications to score\-based generative modeling\.Advances in Neural Information Processing Systems34,pp\. 17695–17709\.Cited by:[§1](https://arxiv.org/html/2608.06893#S1.p1.1),[§2](https://arxiv.org/html/2608.06893#S2.SS0.SSS0.Px1.p1.1)\.
- \[10\]C\. Esteves and A\. Makadia\(2026\)Spectrally\-guided diffusion noise schedules\.InForty\-third International Conference on Machine Learning,External Links:[Link](https://openreview.net/forum?id=5cIgeU4WOG)Cited by:[§1](https://arxiv.org/html/2608.06893#S1.p2.1),[§2](https://arxiv.org/html/2608.06893#S2.SS0.SSS0.Px4.p1.1)\.
- \[11\]F\. Falck, T\. Pandeva, K\. Zahirnia, R\. Lawrence, R\. Turner, E\. Meeds, J\. Zazo, and S\. Karmalkar\(2025\)A fourier space perspective on diffusion models\.arXiv preprint arXiv:2505\.11278\.Cited by:[§1](https://arxiv.org/html/2608.06893#S1.p2.1),[§2](https://arxiv.org/html/2608.06893#S2.SS0.SSS0.Px4.p1.1)\.
- \[12\]F\. Fallah, W\. Li, C\. Hsu, H\. Lee, and Y\. Yang\(2025\)RareFlow: physics\-aware flow\-matching for cross\-sensor super\-resolution of rare\-earth features\.arXiv preprint arXiv:2510\.23816\.Cited by:[§2](https://arxiv.org/html/2608.06893#S2.SS0.SSS0.Px5.p1.1)\.
- \[13\]D\. Freirich, T\. Michaeli, and R\. Meir\(2021\)A theory of the distortion\-perception tradeoff in wasserstein space\.Advances in Neural Information Processing Systems34,pp\. 25661–25672\.Cited by:[§2](https://arxiv.org/html/2608.06893#S2.SS0.SSS0.Px5.p1.1)\.
- \[14\]N\. Gushchin, S\. Kholkin, E\. Burnaev, and A\. Korotin\(2024\)Light and optimal schrödinger bridge matching\.InForty\-first International Conference on Machine Learning,Cited by:[§2](https://arxiv.org/html/2608.06893#S2.SS0.SSS0.Px1.p1.1)\.
- \[15\]N\. Gushchin, D\. Li, D\. Selikhanovych, E\. Burnaev, D\. Baranchuk, and A\. Korotin\(2025\)Inverse bridge matching distillation\.arXiv preprint arXiv:2502\.01362\.Cited by:[§2](https://arxiv.org/html/2608.06893#S2.SS0.SSS0.Px2.p1.1)\.
- \[16\]N\. Gushchin, D\. Selikhanovych, S\. Kholkin, E\. Burnaev, and A\. Korotin\(2024\)Adversarial schrödinger bridge matching\.Advances in Neural Information Processing Systems37,pp\. 89612–89651\.Cited by:[§2](https://arxiv.org/html/2608.06893#S2.SS0.SSS0.Px1.p1.1)\.
- \[17\]G\. He, K\. Zheng, J\. Chen, F\. Bao, and J\. Zhu\(2024\)Consistency diffusion bridge models\.Advances in Neural Information Processing Systems37,pp\. 23516–23548\.Cited by:[§2](https://arxiv.org/html/2608.06893#S2.SS0.SSS0.Px2.p1.1)\.
- \[18\]R\.A\. Horn and C\.R\. Johnson\(2012\)Matrix analysis\.Cambridge University Press\.External Links:ISBN 9781139788885,[Link](https://books.google.com/books?id=O7sgAwAAQBAJ)Cited by:[Table 4](https://arxiv.org/html/2608.06893#S6.T4.8.8.3.1.1)\.
- \[19\]J\. Hou, Z\. Zhu, and J\. Hou\(2026\)Energy\-oriented diffusion bridge for image restoration with foundational diffusion models\.arXiv preprint arXiv:2604\.10983\.Cited by:[§2](https://arxiv.org/html/2608.06893#S2.SS0.SSS0.Px2.p1.1)\.
- \[20\]S\. Howard, P\. Potaptchik, and G\. Deligiannidis\(2026\)Schrödinger bridge matching for tree\-structured costs and entropic wasserstein barycentres\.Advances in Neural Information Processing Systems38,pp\. 130112–130149\.Cited by:[§2](https://arxiv.org/html/2608.06893#S2.SS0.SSS0.Px1.p1.1)\.
- \[21\]X\. Huang, C\. Salaun, C\. Vasconcelos, C\. Theobalt, C\. Oztireli, and G\. Singh\(2024\)Blue noise for diffusion models\.InACM SIGGRAPH 2024 conference papers,pp\. 1–11\.Cited by:[§2](https://arxiv.org/html/2608.06893#S2.SS0.SSS0.Px4.p1.1)\.
- \[22\]T\. Jiralerspong, B\. Earnshaw, J\. Hartford, Y\. Bengio, and L\. Scimeca\(2025\)Shaping inductive bias in diffusion models through frequency\-based noise control\.arXiv preprint arXiv:2502\.10236\.Cited by:[§1](https://arxiv.org/html/2608.06893#S1.p2.1),[§2](https://arxiv.org/html/2608.06893#S2.SS0.SSS0.Px4.p1.1)\.
- \[23\]I\. Karatzas and S\. Shreve\(1991\)Brownian motion and stochastic calculus\.Brownian Motion and Stochastic Calculus,Springer New York\.External Links:ISBN 9780387976556,LCCN 96167783,[Link](https://books.google.com/books?id=ATNy_Zg3PSsC)Cited by:[Table 4](https://arxiv.org/html/2608.06893#S6.T4.11.11.3.3.3)\.
- \[24\]T\. Karras, M\. Aittala, T\. Aila, and S\. Laine\(2022\)Elucidating the design space of diffusion\-based generative models\.Advances in neural information processing systems35,pp\. 26565–26577\.Cited by:[§1](https://arxiv.org/html/2608.06893#S1.p2.1),[§2](https://arxiv.org/html/2608.06893#S2.SS0.SSS0.Px5.p1.1)\.
- \[25\]T\. Karras, S\. Laine, and T\. Aila\(2019\)A style\-based generator architecture for generative adversarial networks\.InProceedings of the IEEE/CVF conference on computer vision and pattern recognition,pp\. 4401–4410\.Cited by:[item 5](https://arxiv.org/html/2608.06893#S1.I1.i5.p1.2)\.
- \[26\]S\. Kholkin, G\. Ksenofontov, D\. Li, N\. Kornilov, N\. Gushchin, A\. Suvorikova, A\. Kroshnin, E\. Burnaev, and A\. Korotin\(2024\)Diffusion & adversarial schrödinger bridges via iterative proportional markovian fitting\.arXiv preprint arXiv:2410\.02601\.Cited by:[§2](https://arxiv.org/html/2608.06893#S2.SS0.SSS0.Px1.p1.1)\.
- \[27\]D\. Kingma and R\. Gao\(2023\)Understanding diffusion objectives as the elbo with simple data augmentation\.Advances in Neural Information Processing Systems36,pp\. 65484–65516\.Cited by:[§2](https://arxiv.org/html/2608.06893#S2.SS0.SSS0.Px5.p1.1)\.
- \[28\]D\. Kingma, T\. Salimans, B\. Poole, and J\. Ho\(2021\)Variational diffusion models\.Advances in neural information processing systems34,pp\. 21696–21707\.Cited by:[§1](https://arxiv.org/html/2608.06893#S1.p2.1),[§2](https://arxiv.org/html/2608.06893#S2.SS0.SSS0.Px5.p1.1)\.
- \[29\]S\. Lin, B\. Liu, J\. Li, and X\. Yang\(2024\)Common diffusion noise schedules and sample steps are flawed\.InProceedings of the IEEE/CVF winter conference on applications of computer vision,pp\. 5404–5411\.Cited by:[§2](https://arxiv.org/html/2608.06893#S2.SS0.SSS0.Px5.p1.1)\.
- \[30\]G\. Liu, Y\. Lipman, M\. Nickel, B\. Karrer, E\. Theodorou, and R\. T\. Q\. Chen\(2024\)Generalized schrödinger bridge matching\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=SoismgeX7z)Cited by:[§2](https://arxiv.org/html/2608.06893#S2.SS0.SSS0.Px1.p1.1)\.
- \[31\]G\. Liu, A\. Vahdat, D\. Huang, E\. A\. Theodorou, W\. Nie, and A\. Anandkumar\(2023\)I2SB: image\-to\-image Schrödinger bridge\.arXiv preprint arXiv:2302\.05872\.Cited by:[§1](https://arxiv.org/html/2608.06893#S1.p1.1),[§2](https://arxiv.org/html/2608.06893#S2.SS0.SSS0.Px2.p1.1),[§3\.1](https://arxiv.org/html/2608.06893#S3.SS1.p1.9)\.
- \[32\]Z\. Liu, P\. Luo, X\. Wang, and X\. Tang\(2015\)Deep learning face attributes in the wild\.InProceedings of the IEEE international conference on computer vision,pp\. 3730–3738\.Cited by:[§4\.4](https://arxiv.org/html/2608.06893#S4.SS4.p2.5)\.
- \[33\]S\. Okada, R\. Yoshihashi, H\. Kataoka, T\. Tanaka,et al\.\(2024\)Constant rate scheduling: constant\-rate distributional change for efficient training and sampling in diffusion models\.arXiv preprint arXiv:2411\.12188\.Cited by:[§2](https://arxiv.org/html/2608.06893#S2.SS0.SSS0.Px5.p1.1)\.
- \[34\]B\. Øksendal\(2010\)Stochastic differential equations: an introduction with applications\.Universitext,Springer Berlin Heidelberg\.External Links:ISBN 9783642143946,LCCN 2003052637,[Link](https://books.google.com/books?id=EQZEAAAAQBAJ)Cited by:[Table 4](https://arxiv.org/html/2608.06893#S6.T4.11.11.3.3.3)\.
- \[35\]S\. Peluchetti\(2023\)Diffusion bridge mixture transports, schrödinger bridge problems and generative modeling\.Journal of Machine Learning Research24\(374\),pp\. 1–51\.Cited by:[§2](https://arxiv.org/html/2608.06893#S2.SS0.SSS0.Px1.p1.1)\.
- \[36\]W\. Rudin\(1976\)Principles of mathematical analysis\.International series in pure and applied mathematics,McGraw\-Hill\.External Links:ISBN 9780070856134,LCCN 75179033,[Link](https://books.google.com/books?id=kwqzPAAACAAJ)Cited by:[Table 4](https://arxiv.org/html/2608.06893#S6.T4.17.17.6.6.6)\.
- \[37\]L\. Scimeca, T\. Jiralerspong, B\. Earnshaw, J\. Hartford, and Y\. Bengio\(2025\)Learning what matters: steering diffusion via spectrally anisotropic forward noise\.arXiv preprint arXiv:2510\.09660\.Cited by:[§2](https://arxiv.org/html/2608.06893#S2.SS0.SSS0.Px4.p1.1)\.
- \[38\]Y\. Shi, V\. De Bortoli, A\. Campbell, and A\. Doucet\(2023\)Diffusion Schrödinger bridge matching\.Advances in Neural Information Processing Systems36,pp\. 62183–62223\.Cited by:[§1](https://arxiv.org/html/2608.06893#S1.p1.1),[§2](https://arxiv.org/html/2608.06893#S2.SS0.SSS0.Px1.p1.1)\.
- \[39\]Z\. Tang, T\. Hang, S\. Gu, D\. Chen, and B\. Guo\(2024\)Simplified diffusion schrödinger bridge\.arXiv preprint arXiv:2403\.14623\.Cited by:[§2](https://arxiv.org/html/2608.06893#S2.SS0.SSS0.Px1.p1.1)\.
- \[40\]H\. Wang, T\. Jin, W\. Lin, S\. Wang, H\. Huang, S\. Ji, and Z\. Zhao\(2025\)Irbridge: solving image restoration bridge with pre\-trained generative diffusion models\.arXiv preprint arXiv:2505\.24406\.Cited by:[§2](https://arxiv.org/html/2608.06893#S2.SS0.SSS0.Px2.p1.1)\.
- \[41\]H\. Wang, J\. Zhang, H\. Chen, H\. Guo, D\. Wang, J\. Ma, and B\. Du\(2026\)Residual diffusion bridge model for image restoration\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 8375–8386\.Cited by:[§2](https://arxiv.org/html/2608.06893#S2.SS0.SSS0.Px2.p1.1)\.
- \[42\]Y\. Wang, S\. Yoon, P\. Jin, M\. Tivnan, Z\. Chen, R\. Hu, L\. Zhang, Z\. Chen, Q\. Li, and D\. Wu\(2024\)Implicit image\-to\-image schrödinger bridge for ct super\-resolution and denoising\.arXiv preprint arXiv:2403\.060692\.Cited by:[§2](https://arxiv.org/html/2608.06893#S2.SS0.SSS0.Px2.p1.1)\.
- \[43\]Y\. Wang, S\. Yoon, P\. Jin, M\. Tivnan, S\. Song, Z\. Chen, R\. Hu, L\. Zhang, Q\. Li, Z\. Chen, and D\. Wu\(2025\)Implicit image\-to\-image schrödinger bridge for image restoration\.Pattern Recognition165,pp\. 111627\.External Links:ISSN 0031\-3203,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.patcog.2025.111627),[Link](https://www.sciencedirect.com/science/article/pii/S0031320325002870)Cited by:[§3\.1](https://arxiv.org/html/2608.06893#S3.SS1.p1.9)\.
- \[44\]K\. Wyrwal, I\. I\. Ceylan, and A\. Tong\(2026\)Topological flow matching\.InThe Fourteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=5CM3ax45Ma)Cited by:[§2](https://arxiv.org/html/2608.06893#S2.SS0.SSS0.Px3.p1.1)\.
- \[45\]Q\. Yao, L\. Gao, Q\. Mao, and M\. Dong\(2025\)Regularized schrödinger bridge: alleviating distortion and exposure bias in solving inverse problems\.arXiv preprint arXiv:2511\.11686\.Cited by:[§2](https://arxiv.org/html/2608.06893#S2.SS0.SSS0.Px2.p1.1),[§2](https://arxiv.org/html/2608.06893#S2.SS0.SSS0.Px5.p1.1)\.
- \[46\]S\. Zhang and M\. Stumpf\(2026\)Learning non\-equilibrium diffusions with schrödinger bridges: from exactly solvable to simulation\-free\.Advances in Neural Information Processing Systems38,pp\. 119861–119898\.Cited by:[§2](https://arxiv.org/html/2608.06893#S2.SS0.SSS0.Px3.p1.1)\.
- \[47\]K\. Zheng, G\. He, J\. Chen, F\. Bao, and J\. Zhu\(2025\)Diffusion bridge implicit models\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 81857–81884\.Cited by:[§2](https://arxiv.org/html/2608.06893#S2.SS0.SSS0.Px2.p1.1)\.
- \[48\]L\. Zhou, A\. Lou, S\. Khanna, and S\. Ermon\(2024\)Denoising diffusion bridge models\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 8160–8171\.Cited by:[§2](https://arxiv.org/html/2608.06893#S2.SS0.SSS0.Px2.p1.1)\.
- \[49\]K\. Zhu, M\. Pan, Y\. Ma, Y\. Fu, J\. Yu, J\. Wang, and Y\. Shi\(2025\)Unidb: a unified diffusion bridge framework via stochastic optimal control\.arXiv preprint arXiv:2502\.05749\.Cited by:[§2](https://arxiv.org/html/2608.06893#S2.SS0.SSS0.Px2.p1.1)\.

Similar Articles

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning

arXiv cs.AI

PRISM is a large-scale benchmark of 10,372 human-calibrated instruction-code pairs for evaluating programmatic video generation, with a funnel-style framework of four metrics. Evaluation of seven LLMs reveals a significant gap between code executability and spatial coherence.