An exact information theory of generalization phase transitions in Bayesian diffusion models
Summary
This paper introduces analytically tractable Bayesian information restricted diffusion (BIRD) models to study the memorization-generalization phase transition in diffusion models, finding that generation proceeds near the edge of memorization and that information restriction helps circumvent the curse of dimensionality.
View Cached Full Text
Cached at: 07/10/26, 06:17 AM
# An exact information theory of generalization phase transitions in Bayesian diffusion models
Source: [https://arxiv.org/html/2607.08041](https://arxiv.org/html/2607.08041)
Henry Hunt∗Department of PhysicsStanford UniversityStanford, CA 94305hshunt@stanford\.eduMason Kamb∗Department of Applied PhysicsStanford UniversityStanford, CA 94305kambm@stanford\.eduSurya Ganguli Department of Applied PhysicsStanford UniversityStanford, CA 94305sganguli@stanford\.edu
###### Abstract
How diffusion models circumvent the curse of dimensionality to learn complex distributions over high dimensional spaces from a finite training set, instead of memorizing it, remains a fundamental mystery\. To address this, we introduce analytically tractable Bayesian information restricted diffusion \(BIRD\) models, in which each pixel observes restricted information about noisy data\. A BIRD model time\-reverses diffusion by inferring which past training sample produced its current restricted observation using the Bayesian posterior\. This model class generalizes existing analytical diffusion models that use spatially local information restriction\. We show that spatially local BIRD models closely approximate trained diffusion modelsearly in training, across different architectures such as UNets and DiTs\. Under minimal assumptions on the data distribution, we identify an information\-theoretic phase boundary between memorization and generalization in the joint space of amount of training data, time in the reverse generative process, and amount of information restriction: a BIRD model memorizes when the mutual information between its restricted noisy observations and the training data exceeds the log number of training points, and it generalizes otherwise\. Experiments across a range of datasets confirm our theoretically predicted location for the transition\. We find that generation proceeds near the edge of memorization: both spatially local BIRD models and early\-training diffusion models track the memorization\-generalization phase boundary by increasingly restricting information over time\. Overall, our results reveal a fundamental role for information restriction in generative AI to circumvent the curse of dimensionality\.
## 1Introduction and related work
Remarkably, generative AI can now learn complex distributions over high dimensional spaces with tractable amounts of training data\. For example, diffusion models\[[1](https://arxiv.org/html/2607.08041#bib.bib1),[2](https://arxiv.org/html/2607.08041#bib.bib2),[3](https://arxiv.org/html/2607.08041#bib.bib3)\]robustly and consistently generalize: when trained on entirely disjoint subsets of data, two independent models willconverge to the same general modelwhen the data subset size is moderately large\[[4](https://arxiv.org/html/2607.08041#bib.bib4)\]\. Below this modest data threshold, the two models may instead memorize their own training data\. Naively, however the curse of dimensionality would suggest that an intractable \(e\.g\. exponential in ambient dimension\) amount of data should be required to transition from memorization to generalization for arbitrary data distributions\[[5](https://arxiv.org/html/2607.08041#bib.bib5)\]\. This raises a fundamental mystery as to when the transition from memorization to generalization occurs for diffusion models, and why it does not require that much training data\. Several theoretical approaches have been taken to explore this mystery\. These approaches often involve analyzing the behavior of simplified diffusion models on simplified datasets\. Diffusion models reverse a forward diffusion process that converts data into noise\. There exists an exact mathematical reversal of this process that starts from noise and follows the ideal, or empirical score function for the finite training set \(App\.[B](https://arxiv.org/html/2607.08041#A2),\[[6](https://arxiv.org/html/2607.08041#bib.bib6)\]\)\. However, such empirical score models require an exponential in dimension amount of data to avoid memorization\[[6](https://arxiv.org/html/2607.08041#bib.bib6),[7](https://arxiv.org/html/2607.08041#bib.bib7)\]\. While several papers have explored the hypothesis that the low dimensionality of the data distribution might rescue generalization of empirical score models\[[8](https://arxiv.org/html/2607.08041#bib.bib8),[9](https://arxiv.org/html/2607.08041#bib.bib9),[10](https://arxiv.org/html/2607.08041#bib.bib10)\], extensive experimental work shows empirical score models fail to generalize on datasets of realistic dimensionality under realistic modeling conditions\[[11](https://arxiv.org/html/2607.08041#bib.bib11),[12](https://arxiv.org/html/2607.08041#bib.bib12),[13](https://arxiv.org/html/2607.08041#bib.bib13),[14](https://arxiv.org/html/2607.08041#bib.bib14),[15](https://arxiv.org/html/2607.08041#bib.bib15),[16](https://arxiv.org/html/2607.08041#bib.bib16),[17](https://arxiv.org/html/2607.08041#bib.bib17)\]\.
Figure 1:A schematic figure showing different transitions from memorizing to generalizing phases in BIRD models\. \(a\) A depiction of the Bayesian guessing game performed by a pixel as it assigns posterior probabilities to training set images\. The forward process starts with a training image of a cat, adds noise to it, then the pixel observes a local patch \(e\.g\. the eye\)\. Based on this restricted observation, the pixel must guess whether this noisy patch was generated from the truck or the cat\. The pixel memorizes if the posterior concentrates on the cat\. \(b\) The phase diagram of the generalization/memorization transition for a BIRD model\. Consider the blue point in the memorization phase \(top phase diagram\)\. Increased noise \(pink arrow\) and decreased observation capacity \(blue arrow\) each increase posterior entropy, making Bayesian guessing harder, and drive a transition from memorization to generalization\. Also, increased dataset size \(yellow arrow\) can shift the phase transition boundary, and increase posterior entropy, so the original location \(blue\) dot is now in the generalization phase \(bottom phase diagram\)\. \(c\) In each of the44plots, red indicates the distribution of a noised training set, the green x indicates a particular observed image, and the blue bars depict the posterior distribution over the training data conditioned on the observed image\. The upper left plot corresponds to the blue point in the top phase diagram of panel b\. The posterior concentrates on a single training point and so it is in the memorization phase\. The 3 colored arrows again correspond to the33distinct ways one can increase posterior entropy and transition to the generalizing phase as in panel \(b\): by increasing noise \(pink\), by decreasing observation capacity \(by projecting out a dimension\) \(blue\), or by increasing the amount of training data \(yellow\)\.The inability of empirical score models to generalize with tractable amounts of training data suggests that successful generalization in diffusion models may originate from limiting inductive biases that prevent learning the empirical score function, such as smoothness\[[12](https://arxiv.org/html/2607.08041#bib.bib12),[18](https://arxiv.org/html/2607.08041#bib.bib18)\], linearity\[[19](https://arxiv.org/html/2607.08041#bib.bib19),[20](https://arxiv.org/html/2607.08041#bib.bib20)\], low\-dimensionality\[[21](https://arxiv.org/html/2607.08041#bib.bib21),[22](https://arxiv.org/html/2607.08041#bib.bib22)\], geometry\-adaptive bases\[[4](https://arxiv.org/html/2607.08041#bib.bib4)\], early training\[[23](https://arxiv.org/html/2607.08041#bib.bib23),[24](https://arxiv.org/html/2607.08041#bib.bib24),[25](https://arxiv.org/html/2607.08041#bib.bib25),[26](https://arxiv.org/html/2607.08041#bib.bib26)\], and model parameterization/architecture\[[27](https://arxiv.org/html/2607.08041#bib.bib27),[28](https://arxiv.org/html/2607.08041#bib.bib28),[29](https://arxiv.org/html/2607.08041#bib.bib29),[30](https://arxiv.org/html/2607.08041#bib.bib30),[31](https://arxiv.org/html/2607.08041#bib.bib31)\]\. However, none of these works derived a highly accurate predictive theory forwhattrained, nonlinear diffusion models actually converge to in the generalizing phase on a wide variety of benchmark datasets\.
One such theory to do so arose in\[[13](https://arxiv.org/html/2607.08041#bib.bib13),[14](https://arxiv.org/html/2607.08041#bib.bib14)\]which positedlocal score models\(LS\), which are optimal diffusion models subject to locality constraints in which each pixel can only use a local spatial neighborhood of the image to compute the score\.\[[13](https://arxiv.org/html/2607.08041#bib.bib13)\]also introducedequivariant\(ES\) and \(boundary\-broken\)equivariant\-localscore models \(ELS\), which incorporated the additional inductive bias of \(translation\) group equivariance and its breaking via boundaries\. In particular\[[13](https://arxiv.org/html/2607.08041#bib.bib13)\]showed that such analytic models could predict image outputs of trained convolution only diffusion models on a case by case basis with high quantitative accuracy\.\[[32](https://arxiv.org/html/2607.08041#bib.bib32)\]then provided a method for obtaining locality scales and patch geometry for LS models via optimal linear denoising\. Interestingly, these works found that image generation proceeds starting from a large coarse spatial scale to a small fine one\. Local generalization mechanisms have also since been investigated in later works\[[33](https://arxiv.org/html/2607.08041#bib.bib33),[34](https://arxiv.org/html/2607.08041#bib.bib34),[35](https://arxiv.org/html/2607.08041#bib.bib35),[36](https://arxiv.org/html/2607.08041#bib.bib36),[37](https://arxiv.org/html/2607.08041#bib.bib37),[38](https://arxiv.org/html/2607.08041#bib.bib38)\]\.
Importantly, while\[[13](https://arxiv.org/html/2607.08041#bib.bib13)\]developed a predictive analytic theory of convolutional diffusion models deep in the generalization phase, this and other works left open the mystery of when and where the memorization to generalization transition occurs on complex benchmark datasets for more powerful models\. Answering this question in general is extremely challenging as it could dependa prioriin a joint fashion on explicit biases from architecture, implicit biases from training dynamics, the structure of the data distribution, time or noise in the generative process, and of course the amount of data\.
In order to understand this question theoretically, in this paper we introduce a general class of diffusion models which we call Bayesian information restricted diffusion \(BIRD\) models\. They are inspired by and encompass all the equivariant/local score models of\[[13](https://arxiv.org/html/2607.08041#bib.bib13),[14](https://arxiv.org/html/2607.08041#bib.bib14)\]\. BIRD models possess a fortuitous combination of intuitive appeal, predictive capacity,andtheoretical tractability, yielding significant conceptual insights into the memorization\-generalization transition\. Our main contributions are:
\(1\) We introduce an intuitively appealing general concept of BIRD models in which each image pixelxxcan be thought of as an agent which observes at timettin the reverse generation process restricted information𝒞x,t\\mathcal\{C\}\_\{x,t\}about a noisy imageϕt\\phi\_\{t\}and uses optimal Bayesian inference over the training data to reverse the forward diffusion process \(Fig\.[1](https://arxiv.org/html/2607.08041#S1.F1)\)\. This constitutes a well\-defined diffusion model obtained from a finite training set without the need for any explicit architectural assumptions or learning dynamics\.
\(2\) We show spatially local BIRD models in which the restricted observation𝒞x,t\\mathcal\{C\}\_\{x,t\}corresponds to a local image patch aroundxx, yield exceedingly good predictions for what many neural diffusion models learnearly in training: for early epochs, they can predict theindividualimage outputs of neural diffusion modelson a case by case basisfor a wide variety of architectures, including UNets with self\-attention and diffusion transformers \(DiTs\), achieving high medianr2∼0\.85−0\.93r^\{2\}\\sim 0\.85\-0\.93\. This latter result goes beyond that of\[[13](https://arxiv.org/html/2607.08041#bib.bib13)\]which showed similarly high predictivity for small CNNs only\.
\(3\) Theoretically, under very mild assumptions on the data distribution, we develop a precise information theoretic criterion for the location of a memorization to generalization phase transition for general BIRD models in the joint space of amount of training data, time \(or noise level in the reverse process\), and amount of information restriction: if, under the forward diffusion process, the mutual information between the observation𝒞x,t\\mathcal\{C\}\_\{x,t\}and the true data distribution, is greater \(less\) than the logarithm of the amount of data, then the BIRD model will memorize \(generalize\), because Bayesian guessing of the underlying image will be to easy \(hard\)\. We experimentally confirm the location of our theoretically predicted phase boundary on a variety of datasets\. These results reveal that the amount of data need only be exponential in the mutual information to avoid memorization, and highlight a fundamental role of information restriction in promoting generalization\.
\(4\) Focusing on spatially local BIRD models in which information restriction is obtained by each pixel observing a location patch of patch scaleLL, we use various information theoretic analyses, to reveal the shape of the memorization to generalization phase transition boundary in the joint space of reverse generation timett\(or noise levelσt\\sigma\_\{t\}\) and patch scaleLL\. We find a time dependent critical patch scaleLc\(t\)L\_\{c\}\(t\)that monotonically decreases in the reverse generation process as both timettand noiseσt\\sigma\_\{t\}decrease fromt=1t=1tot=0t=0\. For anytt,L\(t\)\>Lc\(t\)L\(t\)\>L\_\{c\}\(t\)\(L\(t\)<Lc\(t\)L\(t\)<L\_\{c\}\(t\)\) then BIRD models will memorize \(generalize\)\. Optimal generation and denoising proceeds along this phase transition boundary: in essencesuccessful generation proceeds at the edge of memorization\(Fig\.[1](https://arxiv.org/html/2607.08041#S1.F1)b\)\.
\(5\) We explicitly compute the phase transition boundary between memorization and generalization for BIRD models on image data with power law fall off in the power spectral densityP\(k\)∼k−2−ϵP\(k\)\\sim k^\{\-2\-\\epsilon\}with spatial frequencykk\. Hereϵ=0\\epsilon=0corresponds to scale invariant images\[[39](https://arxiv.org/html/2607.08041#bib.bib39)\]\. We then show that a BIRD model will be in a generalization phase and effectively denoise for all time and noise levels if the amount of data grows with linear image sizeLIL\_\{I\}\(e\.g\.LIL\_\{I\}byLIL\_\{I\}images\) asexpLIϵ\\exp L\_\{I\}^\{\\epsilon\}\. This reveals the striking finding that BIRD models do not suffer from a curse of dimensionality for scale invariant images\. Moreover, for small departures from scale invariance \(e\.g\.ϵ∼0\.1\\epsilon\\sim 0\.1to0\.30\.3as observed in practice\), the data requirements are only exponential inLIϵL\_\{I\}^\{\\epsilon\}, not the total dimensionality of the space of images, which isO\(LI2\)O\(L\_\{I\}^\{2\}\)\.
Together, these results provide fundamental insights into the phase transition between memorization and generalization in a general class of diffusion models, and highlight the role of information restriction in allowing generalization with relatively modest amounts of data\. We summarize our results below, with much more detailed exposition in the Appendix \(see e\.g\. App\.[A](https://arxiv.org/html/2607.08041#A1)for a summary of all mathematical notation and background\)\.
## 2Overall framework
Diffusion models sample from a complex data distributionφ∼P0\(φ\)\\varphi\\sim P\_\{0\}\(\\varphi\)by time reversing a forward diffusion process that transportsP0\(φ\)P\_\{0\}\(\\varphi\)to a simple isotropic Gaussianη∼𝒩\(0,I\)\\eta\\sim\\mathcal\{N\}\(0,I\)through a time dependent family of distributions of noised dataϕt∼Pt\(ϕt\)\\phi\_\{t\}\\sim P\_\{t\}\(\\phi\_\{t\}\)as time proceeds fromt=0t=0tot=1t=1\. A single sample at any fixed timettcan be obtained via the interpolationϕt\(φ,η\)=α¯tφ\+1−α¯tη\\phi\_\{t\}\(\\varphi,\\eta\)=\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\varphi\+\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\}\}\\eta\. The temporal schedule of the signal power is chosen to monotonically decrease fromα¯0=1\\bar\{\\alpha\}\_\{0\}=1at timet=0t=0toα¯1=0\\bar\{\\alpha\}\_\{1\}=0\. We define an important noise\-to\-signal ratioσt2=1−α¯tα¯t\\sigma\_\{t\}^\{2\}=\\frac\{1\-\\bar\{\\alpha\}\_\{t\}\}\{\\bar\{\\alpha\}\_\{t\}\}that monotonically increases fromt=0t=0tot=1t=1\.
Diffusion models generally aim to reverse this flow through a reverse process SDE or ODE that evolves pure noiseϕ1∼𝒩\(0,I\)\\phi\_\{1\}\\sim\\mathcal\{N\}\(0,I\)backwards in time fromt=1t=1tot=0t=0, matching the marginalsPt\(ϕt\)P\_\{t\}\(\\phi\_\{t\}\)of the forward process\. These reverse processes use thescore function∇logPt\\nabla\\log P\_\{t\}ofPt\(ϕt\)P\_\{t\}\(\\phi\_\{t\}\), which must be computed or modeled\. The score computation can be reframed as optimal denoising, via Tweedie’s theorem:
∇logPt\(ϕt\)=−𝔼\[η\|ϕt\]1−α¯t=α¯t1−α¯t𝔼\[φ\|ϕt\]−ϕt1−α¯t\.\\displaystyle\\nabla\\log P\_\{t\}\(\\phi\_\{t\}\)=\-\\frac\{\\mathbb\{E\}\[\\eta\|\\phi\_\{t\}\]\}\{\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\}\}\}=\\frac\{\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\}\{1\-\\bar\{\\alpha\}\_\{t\}\}\\mathbb\{E\}\[\\varphi\|\\phi\_\{t\}\]\-\\frac\{\\phi\_\{t\}\}\{1\-\\bar\{\\alpha\}\_\{t\}\}\.\(1\)One can approximate the posterior mean𝔼\[φ\|ϕt\]\\mathbb\{E\}\[\\varphi\|\\phi\_\{t\}\]by training a neural denoiser modelMt,θ\[ϕt\]M\_\{t,\\theta\}\[\\phi\_\{t\}\]to predictφ\\varphifromϕt\\phi\_\{t\}\. In training this model, we usually do not have direct access to the true data distributionP0\(φ\)P\_\{0\}\(\\varphi\); instead, we only have afinitedataset𝒟\\mathcal\{D\}consisting of\|𝒟\|\|\\mathcal\{D\}\|training samplesφ∈𝒟\\varphi\\in\\mathcal\{D\}each drawn i\.i\.d from the true data distributionP0\(φ\)P\_\{0\}\(\\varphi\)\. For this finite dataset, the Bayes optimal MMSE denoiser \(also known as the empirical score function\) is given by
Mt\[ϕt\]=𝔼\[φ\|ϕt\]=∑φ∈𝒟φPtrain\(φ\|ϕt\)\.\\displaystyle M\_\{t\}\[\\phi\_\{t\}\]=\\mathbb\{E\}\[\\varphi\|\\phi\_\{t\}\]=\\sum\_\{\\varphi\\in\\mathcal\{D\}\}\\varphi\\,P\_\{train\}\(\\varphi\|\\phi\_\{t\}\)\.\(2\)HerePtrain\(φ\|ϕt\)P\_\{train\}\(\\varphi\|\\phi\_\{t\}\)is the posterior probability that a clean training data pointφ∈𝒟\\varphi\\in\\mathcal\{D\}at timet=0t=0led toϕt\\phi\_\{t\}in the forward process \(Fig\.[1](https://arxiv.org/html/2607.08041#S1.F1)b\)\. The structure of this analytic model has an appealing interpretation in terms of a Bayesian guessing game\[[13](https://arxiv.org/html/2607.08041#bib.bib13)\]where the model guesses or forms beliefs about which past data pointφ∈𝒟\\varphi\\in\\mathcal\{D\}led to the noised imageϕt\\phi\_\{t\}using the posteriorPtrain\(φ\|ϕt\)P\_\{train\}\(\\varphi\|\\phi\_\{t\}\), then reverse flows towards the posterior mean\.
While the Bayes optimal diffusion model is appealing in terms of both its analytic and conceptual simplicity, it unfortunately always memorizes the training data by eventually flowing to one of the training pointsφ∈𝒟\\varphi\\in\\mathcal\{D\}\(see App\.[B](https://arxiv.org/html/2607.08041#A2)and\[[6](https://arxiv.org/html/2607.08041#bib.bib6),[13](https://arxiv.org/html/2607.08041#bib.bib13)\]\)\. In essence, the Bayesian guessing game, given theentirenoised imageϕt\\phi\_\{t\}becomes too easy at small enough timestt, corresponding to small enough noiseσt\\sigma\_\{t\}\. Thus this model cannot generalize and is inconsistent with the behavior of trained neural diffusion models\.
### 2\.1Bayes Information Restricted Diffusion \(BIRD\) Models
The unfortunate ease of the Bayesian guessing game for the optimal denoiser raises the question of whether one can obtain alternate Bayesian diffusion models that retain analytic tractability and conceptual simplicity, butalsogeneralize, thereby behaving more like trained neural network diffusion models\. We show that we cansimultaneouslyachieveboththeoretical tractabilityandrealistic generalization, by combining two key approaches: 1\) making the Bayesian guessing game harder by restricting the image information the model uses to guess the training image; and 2\) allowing different parts of the image to usedifferentpieces of restricted information\. We call such models Bayesian Information Restricted Diffusion \(BIRD\) models, which are inspired by and generalize\[[13](https://arxiv.org/html/2607.08041#bib.bib13),[14](https://arxiv.org/html/2607.08041#bib.bib14)\]\.
In the general class of BIRD models \(see App\.[C](https://arxiv.org/html/2607.08041#A3)for more details\) we imagine that each pixelxxis only allowed to observe the noisy imageϕt\\phi\_\{t\}through a restricted, possibly stochastic, observation𝒞x,t∈ℝnx,t\\mathcal\{C\}\_\{x,t\}\\in\\mathbb\{R\}^\{n\_\{x,t\}\}ofϕt∈ℝd\\phi\_\{t\}\\in\\mathbb\{R\}^\{d\}\. The observation dimensionnx,tn\_\{x,t\}is less than the image dimensiondd\. Then the pixelxxcomputes a denoised pixel valueMx,t\[𝒞x,t\]M\_\{x,t\}\[\\mathcal\{C\}\_\{x,t\}\]\. The optimal Bayesian MMSE individual pixel denoiser has a structure similar to that ofMtM\_\{t\}in \([2](https://arxiv.org/html/2607.08041#S2.E2)\):
Mx,t\[𝒞x,t\]=∑φ∈𝒟φxPtrain\(φ\|𝒞x,t\)\.\\displaystyle M\_\{x,t\}\[\\mathcal\{C\}\_\{x,t\}\]=\\sum\_\{\\varphi\\in\\mathcal\{D\}\}\\varphi\_\{x\}\\,P\_\{train\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\.\(3\)HerePtrain\(φ\|𝒞x,t\)P\_\{train\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)is pixelxx’s posterior belief distribution that any past training sampleφ∈𝒟\\varphi\\in\\mathcal\{D\}led to the pixel’s observation𝒞x,t\\mathcal\{C\}\_\{x,t\}in the forward training Markov chainφ→ϕt→𝒞x,t\\varphi\\rightarrow\\phi\_\{t\}\\rightarrow\\mathcal\{C\}\_\{x,t\}\(see App\.[A](https://arxiv.org/html/2607.08041#A1)for notation and App\.[C](https://arxiv.org/html/2607.08041#A3)for explicit expressions forPtrain\(φ\|𝒞x,t\)P\_\{train\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\)\. The pixel’s observation restriction𝒞x,t\\mathcal\{C\}\_\{x,t\}at the end of this forward information channel makes the Bayesian guessing game harder, and ideally prevents memorization\. The MMSE denoiser for pixelxxin \([3](https://arxiv.org/html/2607.08041#S2.E3)\) simply returns the posterior mean training image pixel value at pixelxx\.
A simple channel restriction𝒞x,t\\mathcal\{C\}\_\{x,t\}is a deterministic linear projection operator𝒫x,t\\mathcal\{P\}\_\{x,t\}applied toϕt\\phi\_\{t\}\. The LS model of\[[13](https://arxiv.org/html/2607.08041#bib.bib13)\]is a special case where𝒫x,t\\mathcal\{P\}\_\{x,t\}projects an imageϕt\\phi\_\{t\}onto a coordinate subspace of pixel components within a local spatial patchΩx,t\\Omega\_\{x,t\}surrounding pixelxx, yielding the restricted observation𝒞x,t\[ϕt\]=ϕΩx,t\\mathcal\{C\}\_\{x,t\}\[\\phi\_\{t\}\]=\\phi\_\{\\Omega\_\{x,t\}\}, whereϕΩx,t\\phi\_\{\\Omega\_\{x,t\}\}is a vector of all pixel values ofϕt\\phi\_\{t\}restrictedto the local patchΩx,t\\Omega\_\{x,t\}\(see App\.[A](https://arxiv.org/html/2607.08041#A1)for notation\)\. The BIRD framework also generalizes ES and ELS machine of\[[13](https://arxiv.org/html/2607.08041#bib.bib13)\]\(see App\.[C\.1](https://arxiv.org/html/2607.08041#A3.SS1)for details\)\. We call such models spatially local BIRD models, indicating that spatial locality is the specific form of information restriction employed, though BIRD models and our theory of them are more general\.
We next show in Sec\.[3](https://arxiv.org/html/2607.08041#S3)that spatially local BIRD models are highly relevant in that they describe well the early learning phase of many diffusion model architectures\. And after that in Sec\.[4](https://arxiv.org/html/2607.08041#S4)we show that the general class of all BIRD models provides a highly useful abstraction, in that a general theory of the memorization\-generalization phase transition can be derived forallsuch BIRD models for essentiallyanydata distribution\.
## 3BIRD models describe the early learning phase of diffusion models
𝒟2𝒟1\\mathcal\{D\}\_\{2\}\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\mathcal\{D\}\_\{1\}\\,\\,\\,
𝒟2𝒟1\\mathcal\{D\}\_\{2\}\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\mathcal\{D\}\_\{1\}
𝒟2𝒟1\\mathcal\{D\}\_\{2\}\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\mathcal\{D\}\_\{1\}
\(\(a\)\)Celeba64



\(\(b\)\)Cifar10



\(\(c\)\)MNIST



\(\(d\)\)FashionMNIST
Figure 2:The consistent and robust generalization of diffusion models early in training, across various datasets and architectures, is well\-predicted by spatially local BIRD models with a patch size near the critical scale, as described below\. Here UNets and DiTs are trained on disjoint dataset subsets𝒟1\\mathcal\{D\}\_\{1\}and𝒟2\\mathcal\{D\}\_\{2\}, but produce nearly identical images when fed the same noise input\. These images are both similar to the corresponding BIRD model outputs\. Many more such samples are shown in App\.[H\.2](https://arxiv.org/html/2607.08041#A8.SS2)in figs\.[7](https://arxiv.org/html/2607.08041#A8.F7)\-[11](https://arxiv.org/html/2607.08041#A8.F11), along with details of the experiment and training procedure in App\.[H](https://arxiv.org/html/2607.08041#A8)\.\[[13](https://arxiv.org/html/2607.08041#bib.bib13)\]showed that spatially local BIRD models could accurately predict theindividualimage outputs of fully trained small CNN based diffusion models on acase by casebasis for a range of datasets with a very high medianr2≈0\.9r^\{2\}\\approx 0\.9between theory and experiment\. Here we extend this result by showing that spatially local BIRD models canalsopredictindividualimage outputs of more sophisticated architectures, like UNets with self\-attention and diffusion transformers \(DiTs\), when they are in theearliest stages of the training process\.
In our experiments on four standard datasets \(Celeba64, CIFAR10, FashionMNIST, and MNIST\), and two architectues \(UNets and DiTs\), we find that the agreement between spatially local BIRD theory and early trained diffusion models peaks at values ofr2∼0\.9r^\{2\}\\sim 0\.9around1010to3030epochs of training before slowly declining thereafter\. See tables[1](https://arxiv.org/html/2607.08041#A8.T1)and[2](https://arxiv.org/html/2607.08041#A8.T2)for detailed numerical values of the agreement, and see App\.[H](https://arxiv.org/html/2607.08041#A8)for more experimental parameters, and for our approach to patch selection and scale calibration, and further samples\.
We also show, in this early training regime, both BIRD models and trained architectures exhibit the phenomenon of consistent and robust generalization, where different models trained on disjoint data subsets nevertheless generate the same images when fed the same input noise \(Fig\.[2](https://arxiv.org/html/2607.08041#S3.F2)\)\. In our experiments, we compare the outputs of the33distinct models \(BIRD, UNets, DiTs\) each obtained from two disjoint datasets𝒟1\\mathcal\{D\}\_\{1\}and𝒟2\\mathcal\{D\}\_\{2\}\. We find that all3×2=63\\times 2=6images are highly similar \(Fig\.[2](https://arxiv.org/html/2607.08041#S3.F2)\)\. This indicates BIRD models provide a good description of the generalization phase of both DiTs and UNets early in training, and that such BIRD models themselves are in a generalizing phase\. More samples are included in App\.[H](https://arxiv.org/html/2607.08041#A8), as well as quantitative validations in Fig\.[6](https://arxiv.org/html/2607.08041#A8.F6)and tables[1](https://arxiv.org/html/2607.08041#A8.T1)and[2](https://arxiv.org/html/2607.08041#A8.T2)\.
This strong agreement between BIRD model theory and neural experiments motivates the use of BIRD models as a theoretical laboratory within which to study the nature of the memorization\-generalization phase transition\. We elucidate this theory next in Sec\.[4](https://arxiv.org/html/2607.08041#S4)\.
## 4The memorization\-generalization phase transition in BIRD models
Figure 3:Comparison of theory and experiment for the memorization\-generalization phase transition\. Experimental curves plot numerical estimates of the average entropy deficit \(ln\|𝒟\|−S\[P\(φ\|ϕΩ\)\]\\ln\|\\mathcal\{D\}\|\-S\[P\(\\varphi\|\\phi\_\{\\Omega\}\)\]\) in a spatially local BIRD model with a patch size5×55\\times 5\.σt2\\sigma\_\{t\}^\{2\}is the noise strength\. Theory curves plot mutual information or its Gaussian upper bounds\. As noiseσt2\\sigma^\{2\}\_\{t\}decreases, the theoretically derived mutual information curves rise in the generalizing phase then saturate toln\|D\|\\ln\\mathcal\{\|\}D\|in the memorizing phase\. Moreover the mutual information theory curves closely track, or upper bound the entropy deficit experimental curves\. Each sub\-figure corresponds to a different true data distribution: \(a\) Translationally invariant Gaussian data withα=1\.7\\alpha=1\.7power law spectrum \(see Sec\.[5\.2](https://arxiv.org/html/2607.08041#S5.SS2)\) \(b\) An isotropic hierarchical mixture of Gaussians model \(see App\.[E\.4](https://arxiv.org/html/2607.08041#A5.SS4)for theory and experiment\); \(c\) CIFAR10; \(d\) CelebA32\. The theory curves for CIFAR10 and CelebA32 are given by a Gaussian upper bound on mutual information \(see App\.[E\.2](https://arxiv.org/html/2607.08041#A5.SS2)\)\. See App\.[F\.5](https://arxiv.org/html/2607.08041#A6.SS5)for further details\.We first show, via Fano’s inequality, that perfect guessing in the pixel Bayesian guessing game \(i\.e\. memorization\) impliesI\(𝒞x,t;φ\)≥ln\|D\|I\(\\mathcal\{C\}\_\{x,t\};\\varphi\)\\geq\\ln\|D\|, whereI\(𝒞x,t;φ\)I\(\\mathcal\{C\}\_\{x,t\};\\varphi\)is the mutual information between atestdata pointφ\\varphidrawn from thetruedata distributionP0\(φ\)P\_\{0\}\(\\varphi\)and the pixel observation𝒞x,t\\mathcal\{C\}\_\{x,t\}under the forward testing Markov chainφ→ϕt→𝒞x,t\\varphi\\rightarrow\\phi\_\{t\}\\rightarrow\\mathcal\{C\}\_\{x,t\}\(see App\.[A](https://arxiv.org/html/2607.08041#A1)for notation and App\.[E\.1](https://arxiv.org/html/2607.08041#A5.SS1)for a proof\)\. However, an approach via Fano’s inequality only proves that in the memorization phase, we haveI\(𝒞t,x;φ\)≥ln\|𝒟\|I\(\\mathcal\{C\}\_\{t,x\};\\varphi\)\\geq\\ln\|\\mathcal\{D\}\|\. It doesnotprove the converse, i\.e\. that in the generalizing phase, the opposite holds:I\(𝒞t,x;φ\)<ln\|𝒟\|I\(\\mathcal\{C\}\_\{t,x\};\\varphi\)<\\ln\|\\mathcal\{D\}\|\.
To actually locate the phase transition between memorization and generalization, in App\.[D\.5](https://arxiv.org/html/2607.08041#A4.SS5), we prove for posteriors with sub\-exponential tails, in large dataset size and high observation dimension limit, that
S\[Ptrain\(φ\|𝒞x,t\)\]\\displaystyle S\[P\_\{train\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\]=\{ln\|𝒟\|−I\(φ;𝒞x,t\)ln\|𝒟\|\>I\(φ;𝒞x,t\)\(generalization\)0ln\|𝒟\|≤I\(φ;𝒞x,t\)\(memorization\)\.\\displaystyle=\\begin\{cases\}\\ln\|\\mathcal\{D\}\|\-I\(\\varphi;\\mathcal\{C\}\_\{x,t\}\)&\\ln\|\\mathcal\{D\}\|\>I\(\\varphi;\\mathcal\{C\}\_\{x,t\}\)\\,\\,\\,\\,\\,\\,\\,\(\\text\{generalization\}\)\\\\ 0&\\ln\|\\mathcal\{D\}\|\\leq I\(\\varphi;\\mathcal\{C\}\_\{x,t\}\)\\,\\,\\,\\,\\,\\,\\,\(\\text\{memorization\}\)\.\\\\ \\end\{cases\}\(4\)Thus the phase transition boundary isexactlygiven by the equation
ln\|𝒟\|=I\(φ;𝒞x,t\)\.\\ln\|\\mathcal\{D\}\|=I\(\\varphi;\\mathcal\{C\}\_\{x,t\}\)\.\(5\)Intuitively, more mutual information than the entropyln\|𝒟\|\\ln\|\\mathcal\{D\}\|of the uniform prior on the training data makes Bayesian guessing easy with zero posterior entropyS\[Ptrain\(φ\|𝒞x,t\)\]S\[P\_\{train\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\]leading to memorization, while less mutual information makes Bayesian guessing hard with positive posterior entropy leading to generalization\.
We can experimentally test \([4](https://arxiv.org/html/2607.08041#S4.E4)\) becauseS\[Ptrain\(φ\|𝒞x,t\)\]S\[P\_\{train\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\]can be numerically computed for any finite dataset, while the mutual informationI\(φ;𝒞x,t\)I\(\\varphi;\\mathcal\{C\}\_\{x,t\}\)can be exactly calculated for simple data distributions and upper bounded for real data sets\. In order to collapse the curves from different dataset sizes, and to compare theory and experiment, we plotln\|𝒟\|−S\[Ptrain\(φ\|ϕΩ\)\]\\ln\|\\mathcal\{D\}\|\-S\[P\_\{train\}\(\\varphi\|\\phi\_\{\\Omega\}\)\], obtained from numerical experiments, and compare it either to: \(1\) the theoretically obtained mutual information \(which it should be equal to in the generalizing phase\), or \(2\)ln\|𝒟\|\\ln\|\\mathcal\{D\}\|\(which it should be equal to in the memorizing phase\)\. For real datasets, we cannot calculate the exact mutual information, so we use the Gaussian upper bound derived in App\.[E\.2](https://arxiv.org/html/2607.08041#A5.SS2)which only depends on the second order statistics of the data\. Fig\.[3](https://arxiv.org/html/2607.08041#S4.F3)exhibits numerical validation of \([4](https://arxiv.org/html/2607.08041#S4.E4)\) for both real and toy datasets\.
Our results above generalize previous work\[[6](https://arxiv.org/html/2607.08041#bib.bib6)\], which studied the memorization to generalization phase transition for the empirical score function \(in which each pixel observes the entire image\)\. They defined the transition timettto occur whenS\[P\(ϕt\)\]=Ssep\|𝒟\|S\[P\(\\phi\_\{t\}\)\]=S\_\{sep\}^\{\|\\mathcal\{D\}\|\}whereSsep\|𝒟\|S\_\{sep\}^\{\|\\mathcal\{D\}\|\}is entropy of\|𝒟\|\|\\mathcal\{D\}\|separated Gaussians and computed this transition time for the training set only, when the data came from an isotropic Gaussian distribution\. We include information restriction, generalize to essentially arbitrary data distributions, and find a general information theoretic criterion for the memorization\-generalization transition for both training and testing data, as well as a pointwise condition for any single observation \(App\.[D\.5](https://arxiv.org/html/2607.08041#A4.SS5)\)\.
## 5How BIRD models circumvent the curse of dimensionality
\(\(a\)\)




\(\(b\)\)
Figure 4:Effect of varying the patch scaleLL\(for square patches\) in a spatially local BIRD model\. \(a\) The top \(bottom\) shows the outcome of denoising of a single test \(train\) image using BIRD models of different patch sizesLLstarting from noiseσt=1\\sigma\_\{t\}=1\. Optimal denoising occurs whenL=7L=7\(in general this optimal denoising scale coincides withLcL\_\{c\}\)\. ForL=5<LcL=5<L\_\{c\}, from theory we know the denoiser is robust, but denoising performance is poor both on the training and test image\. ForL=15\>LcL=15\>L\_\{c\}, the model is in the memorization phase, and it poorly denoises the test image \(compare rightmost to leftmost image in top row\), but perfectly denoises \(i\.e\. memorizes\) the training image \(compare rightmost to leftmost image in bottom row\)\. \(b\) Comparisons between the anisotropic Gaussian bound \(Lemma[5\.1](https://arxiv.org/html/2607.08041#S5.Thmdefinition1)\), the numerically computed values for the critical scales, and the spectral scale for BWCeleba32, BWCIFAR10, and MNIST, which are generally very close \(see App\.[F\.6](https://arxiv.org/html/2607.08041#A6.SS6)for details\)\.We now explain how, for scale invariant natural images, spatially local BIRD models can effectively denoise while avoiding memorization, even for large images\.
### 5\.1The critical patch scale at the memorization\-generalization transition
A very natural information restriction approach, studied in\[[13](https://arxiv.org/html/2607.08041#bib.bib13)\], is to allow each pixelxxto observe a local patchΩL,x\\Omega\_\{L,x\}, where we define the patch scaleL=\|ΩL,x\|L=\\sqrt\{\|\\Omega\_\{L,x\}\|\}\. Following\[[13](https://arxiv.org/html/2607.08041#bib.bib13)\]we denote this projection asϕt;ΩL,x\\phi\_\{t;\\Omega\_\{L,x\}\}\. For square patches\[[13](https://arxiv.org/html/2607.08041#bib.bib13)\]LLis simply the side length\. We consider a one parameter family of patches, indexed by scaleLL, obeyingΩL′,x⊃ΩL,x\\Omega\_\{L^\{\\prime\},x\}\\supset\\Omega\_\{L,x\}ifL′\>LL^\{\\prime\}\>L\. The fact that larger patches include smaller patches implies thatI\(φ,ϕt,Ωx,L\)I\(\\varphi,\\phi\_\{t,\{\\Omega\_\{x,L\}\}\}\)is monotonically increasing with the patch scaleLL\. This implies, according to \([4](https://arxiv.org/html/2607.08041#S4.E4)\) that there is acritical patch scaleLc\(σt\)L\_\{c\}\(\\sigma\_\{t\}\), for any fixed noise levelσt\\sigma\_\{t\}, obeying the equality
ln\|𝒟\|=I\(φ;ϕt,ΩL,x\),\\displaystyle\\ln\|\\mathcal\{D\}\|=I\(\\varphi;\\phi\_\{t,\{\\Omega\_\{L,x\}\}\}\),\(6\)such that ifL<LcL<L\_\{c\}\(L\>Lc\)L\>L\_\{c\}\)the spatially local BIRD model is in the generalizing \(memorizing\) phase\.Lc\(σt\)L\_\{c\}\(\\sigma\_\{t\}\)then defines a phase boundary between memorization and generalization in the joint space of patch scaleLLand noiseσt\\sigma\_\{t\}\(or timettin the reverse process\)\.
In high dimensions, the critical scale also generally sets theoptimalscale \(with respect to denoising error on the test set\) for the spatially local BIRD model\. The reason for this is as follows\. If the model uses a scaleL\>LcL\>L\_\{c\}, denoising will reduce to nearest\-neighbor search over the training set, which will fail to correspond to the underlying population distribution over images and thus yield significant statistical error \(see[4\(a\)](https://arxiv.org/html/2607.08041#S5.F4.sf1)for an example of this failure mode as the length scaleLLis taken large\)\. Meanwhile, ifL<LcL<L\_\{c\}, the noised training set distribution provides a robust estimate for the true distribution over noised images, but the information the model learns about the sampleI\(φ;ϕΩL;x\)I\(\\varphi;\\phi\_\{\\Omega\_\{L;x\}\}\)will be less than the critical valueI\(φ;ϕΩLc;x\)I\(\\varphi;\\phi\_\{\\Omega\_\{L\_\{c\};x\}\}\), so it will perform worse than a higher sub\-critical scale\. In low dimensions, there is no longer a ‘sharp’ phase transition as above, but the optimal scaleLLwill nonetheless generally sit aroundL≈LcL\\approx L\_\{c\}\.
Withoutanyassumptions on the data distribution, we can make a very general statement about the shape of the phase boundaryLc\(σt\)L\_\{c\}\(\\sigma\_\{t\}\)\. SinceI\(φ;ϕt,ΩL,x\)I\(\\varphi;\\phi\_\{t,\{\\Omega\_\{L,x\}\}\}\)monotonically increases withLLand decreases withσt\\sigma\_\{t\}, implicit differentiation of \([6](https://arxiv.org/html/2607.08041#S5.E6)\) with respect toLLimplies thatLc\(σt\)L\_\{c\}\(\\sigma\_\{t\}\)monotonically increases with the noiseσt\\sigma\_\{t\}\.
With relatively few additional assumptions, we can make more precise statements about the phase boundaryLc\(σt\)L\_\{c\}\(\\sigma\_\{t\}\)and its dependence on\|𝒟\|\|\\mathcal\{D\}\|\. We first establish the following bounds, which we call the ‘anisotropic’ and ‘isotropic’ Gaussian bounds \(see App\.[F\.3](https://arxiv.org/html/2607.08041#A6.SS3)\):
###### Lemma 5\.1
Supposeφ\\varphiis drawn from an arbitrary data distributionP0P\_\{0\}with meanμ\\muand covarianceΣ\\Sigma\. We define the distributionQ1=𝒩\(μ,Σ\)Q\_\{1\}=\\mathcal\{N\}\(\\mu,\\Sigma\)to be a Gaussian with identical second\-order statistics toP0P\_\{0\}, andQ2=𝒩\(μ,Diag\[Σ\]\)Q\_\{2\}=\\mathcal\{N\}\(\\mu,\\text\{Diag\}\[\\Sigma\]\)to be the Gaussian with identical diagonal covariance and zero off\-diagonal elements\. We then haveLc;P\(σt\)≥Lc;Q1\(σt\)≥Lc;Q2\(σt\)L\_\{c;P\}\(\\sigma\_\{t\}\)\\geq L\_\{c;Q\_\{1\}\}\(\\sigma\_\{t\}\)\\geq L\_\{c;Q\_\{2\}\}\(\\sigma\_\{t\}\)\.
The Gaussian bounds are helpful numerically because they can be computed directly from the data covariance matrix, which enables direct numerical computation of a lower bound on the critical scale for realistic data \(fig\.[4\(b\)](https://arxiv.org/html/2607.08041#S5.F4.sf2)\)\. It also theoretically provides a general lower bound of the formLc\(σt\)∼Ω\(σtln\|𝒟\|\)L\_\{c\}\(\\sigma\_\{t\}\)\\sim\\Omega\(\\sigma\_\{t\}\\sqrt\{\\ln\|\\mathcal\{D\}\|\}\)\(see App\.[E\.2](https://arxiv.org/html/2607.08041#A5.SS2)for details\)\. We also show \(see App\. thm\.[F\.3](https://arxiv.org/html/2607.08041#A6.Thmdefinition3)for proof\) a very general scaling theorem, which covers a wide range of naturalistic images and shows thatLspec∼O\(σt\)L\_\{spec\}\\sim O\(\\sigma\_\{t\}\)scaling of the critical scale isgeneric:
###### Theorem 5\.2
Suppose we have a translationally\-invariant distribution over imagesφ\\varphiwith bounded pointwise variance and asymptotically image\-size\-extensive entropy, with an entropy densitylim\|Ω\|→∞1\|Ω\|S\(φΩ\)=S~\\lim\_\{\|\\Omega\|\\to\\infty\}\\frac\{1\}\{\|\\Omega\|\}S\(\\varphi\_\{\\Omega\}\)=\\tilde\{S\}\. Then, forln\|𝒟\|\\ln\|\\mathcal\{D\}\|which isO\(1\)O\(1\)with respect toσt\\sigma\_\{t\},
Lc\(σt\)∼O\(σtln\|𝒟\|\)\\displaystyle L\_\{c\}\(\\sigma\_\{t\}\)\\sim O\\big\(\\sigma\_\{t\}\\sqrt\{\\ln\|\\mathcal\{D\}\|\}\\big\)\(7\)
The intuition behind this theorem is that for sufficiently high noise, each pixel will confer an amount of information proportional to the signal\-to\-noise ratioO\(1/σt2\)O\(1/\\sigma\_\{t\}^\{2\}\)about the underlying imageφ\\varphi, so the whole patch providesO\(σt−2\|Ω\|\)O\(\\sigma\_\{t\}^\{\-2\}\|\\Omega\|\)bits of information; setting this equal toln\|𝒟\|\\ln\|\\mathcal\{D\}\|and rearranging gives \([7](https://arxiv.org/html/2607.08041#S5.E7)\)\.
Our theory implies ifL\>Lc\(σt\)L\>L\_\{c\}\(\\sigma\_\{t\}\), the spatially local BIRD denoiser will memorize\. We examine the denoising performance of such a denoiser as a function of the patch scaleLLand confirm this memorization behavior in Fig\.[4\(a\)](https://arxiv.org/html/2607.08041#S5.F4.sf1)\. However, merely avoiding memorization is not sufficient to guarantee that a model can denoiseeffectively: for instance, as shown in Fig\.[4\(a\)](https://arxiv.org/html/2607.08041#S5.F4.sf1), BIRD model with a small scale such asL≪LcL\\ll L\_\{c\}may be robust to the training data \(i\.e\. generalize\) but be a poor denoiser\. The reason is that with such a small patch sizeLL, it cannot access correlations that are important for denoising \(at a given noise level\)\. We next examine another length scale for natural images that is relevant for \(linear\) denoising\.
### 5\.2The spectral length scale for effective linear denoising
While natural images contain correlations at many different length scales, existing explorations of spatial locality in diffusion models\[[32](https://arxiv.org/html/2607.08041#bib.bib32),[40](https://arxiv.org/html/2607.08041#bib.bib40)\]emphasize a quantity which we call thespectral scale, which corresponds to the scale of relevant second\-order correlations in the data\. Naturalistic images have a well\-known power\-law falloff in their power spectral density ofP\(k\)≈A\|k\|−αP\(k\)\\approx A\|k\|^\{\-\\alpha\}\[[39](https://arxiv.org/html/2607.08041#bib.bib39)\]\. The exponentα\\alphais typicallyα≈2−ϵ\\alpha\\approx 2\-\\epsilon, withϵ\\epsilontypically in the range 0\.1\-0\.3, near the perfectly scale\-invariant valueα=2\\alpha=2\. When performing denoising, there is a natural spectral length scale, defined as the inverse of the spatial frequency where the noise powerσt2\\sigma\_\{t\}^\{2\}equals the signal power, yieldingLspec=\(σt2A\)1/αL\_\{spec\}=\(\\frac\{\\sigma\_\{t\}^\{2\}\}\{A\}\)^\{1/\\alpha\}\. For exactly scale\-invariant images withα=2−ϵ\\alpha=2\-\\epsilon, this reduces toLspec∼σt2/\(2−ϵ\)L\_\{spec\}\\sim\\sigma\_\{t\}^\{2/\(2\-\\epsilon\)\}\. This scale can also be interpreted as the approximate radius of the optimal linear denoiser, or the Wiener filter\[[41](https://arxiv.org/html/2607.08041#bib.bib41)\]\. Because of the relatively compact support of this filter, a spatially local BIRD model with patch scaleL≥LspecL\\geq L\_\{spec\}should still denoise with relatively small error\. Indeed,\[[32](https://arxiv.org/html/2607.08041#bib.bib32)\]established empirically that BIRD models withL∼LspecL\\sim L\_\{spec\}generated reasonable samples\. Conversely ifL≪LspecL\\ll L\_\{spec\}we should expect poor denoising, as confirmed in Fig\.[4\(a\)](https://arxiv.org/html/2607.08041#S5.F4.sf1)\.
This raises a critical issue for BIRD models: their patch scaleLLmay need to be close to the spectral scaleLspecL\_\{spec\}to achieve good denoising performance, but ifLspec≫LcL\_\{spec\}\\gg L\_\{c\}, then the BIRD model cannot access the spectral scale without exhibiting deleterious memorization\.
How much data is required so that a BIRD model can at least access the spectral scaleLspecL\_\{spec\}without memorizing\. Our theoretical results in Sec\.[5\.1](https://arxiv.org/html/2607.08041#S5.SS1)imply that forLc∼LspecL\_\{c\}\\sim L\_\{spec\}at all noise levelsσt\\sigma\_\{t\}, we need the proportionalityLc∼σtln\|𝒟\|∼σt22−ϵ∼LspecL\_\{c\}\\sim\\sigma\_\{t\}\\sqrt\{\\ln\|\\mathcal\{D\}\|\}\\sim\\sigma\_\{t\}^\{\\frac\{2\}\{2\-\\epsilon\}\}\\sim L\_\{spec\}to hold, which impliesln\|𝒟\|∼σt2ϵ2−ϵ∼Lspecϵ\\ln\|\\mathcal\{D\}\|\\sim\\sigma\_\{t\}^\{\\frac\{2\\epsilon\}\{2\-\\epsilon\}\}\\sim L\_\{spec\}^\{\\epsilon\}\. The largest spectral scale that the model may need to access is the image sizeLIL\_\{I\}\. This implies that the minimum dataset size\|𝒟\|\|\\mathcal\{D\}\|required to avoid memorization must scale with image sizeLIL\_\{I\}as
ln\|𝒟\|∼LIϵ\\displaystyle\\ln\|\\mathcal\{D\}\|\\sim L\_\{I\}^\{\\epsilon\}\(8\)Whenϵ\\epsilonis taken to 0 \(i\.e\. for perfectly\-scale invariant data\), we see that there isno required scaling of dataset size with image size, showing that spatially local BIRD models easily circumvent the curse of dimensionality for perfectly scale\-invariant data\.
Next, we numerically compare the critical scaleLc\(σt\)L\_\{c\}\(\\sigma\_\{t\}\)and the spectral scaleLspec\(σt\)L\_\{spec\}\(\\sigma\_\{t\}\)for several standard image datasets\. We computeLc\(σt\)L\_\{c\}\(\\sigma\_\{t\}\)by solving \([6](https://arxiv.org/html/2607.08041#S5.E6)\) using an estimator for the mutual information\. We can also compute the spectral scaleLspecL\_\{spec\}from the dataset’s power spectral density\. In Fig[4\(b\)](https://arxiv.org/html/2607.08041#S5.F4.sf2), we showLc\(σt\)L\_\{c\}\(\\sigma\_\{t\}\)andLspec\(σt\)L\_\{spec\}\(\\sigma\_\{t\}\)and a theoretical prediction forLcL\_\{c\}from a Gaussian bound\. We see clearly in these curves that the scaling behavior ofLcL\_\{c\}andLspecL\_\{spec\}are highly aligned, consistent with our theoretical expectations for scale\-invariant image data\. See App\.[F\.6](https://arxiv.org/html/2607.08041#A6.SS6)for more experimental details\.
## 6Discussion
Using information theory, we were able to obtain the memorization\-generalization phase diagram for a highly general class of BIRD models and for a very general class data distributions\. While BIRD models introduce inductive biases through information restriction, a significant limitation of our theory is that it cannot capture potentially other inductive biases due to architectural choices or learning dynamics\. For example, BIRD models cannot express a bias towards linearity which may arise in trained models\. However, we are not aware of any theory of diffusion models that can precisely delineate memorization\-generalization phase transitions boundaries with arbitrary architectures and data distributions\. All existing theory either use simplified architectures or simple models of data\. BIRD models therefore provide an exciting new theoretical laboratory to study memorization and generalization across many possible data distributions\.
## References
- Sohl\-Dickstein et al\. \[2015\]Jascha Sohl\-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli\.Deep unsupervised learning using nonequilibrium thermodynamics\.In*International conference on machine learning*, pages 2256–2265\. PMLR, 2015\.
- Ho et al\. \[2020\]Jonathan Ho, Ajay Jain, and Pieter Abbeel\.Denoising diffusion probabilistic models\.*Advances in neural information processing systems*, 33:6840–6851, 2020\.
- Song et al\. \[2020\]Jiaming Song, Chenlin Meng, and Stefano Ermon\.Denoising diffusion implicit models\.*arXiv preprint arXiv:2010\.02502*, 2020\.
- Kadkhodaie et al\. \[2023\]Zahra Kadkhodaie, Florentin Guth, Eero P Simoncelli, and Stéphane Mallat\.Generalization in diffusion models arises from geometry\-adaptive harmonic representation\.*arXiv preprint arXiv:2310\.02557*, 2023\.
- Shalev\-Shwartz and Ben\-David \[2014\]Shai Shalev\-Shwartz and Shai Ben\-David\.*Understanding Machine Learning: From Theory to Algorithms*\.Cambridge University Press, 2014\.
- Biroli et al\. \[2024\]Giulio Biroli, Tony Bonnaire, Valentin De Bortoli, and Marc Mézard\.Dynamical regimes of diffusion models\.*arXiv preprint arXiv:2402\.18491*, 2024\.
- De Bortoli \[2022\]Valentin De Bortoli\.Convergence of denoising diffusion models under the manifold hypothesis\.*arXiv preprint arXiv:2208\.05314*, 2022\.
- Achilli et al\. \[2025\]Beatrice Achilli, Luca Ambrogioni, Carlo Lucibello, Marc Mézard, and Enrico Ventura\.Memorization and generalization in generative diffusion under the manifold hypothesis\.*Journal of Statistical Mechanics: Theory and Experiment*, 2025\(7\):073401, 2025\.
- Li and Yan \[2024\]Gen Li and Yuling Yan\.Adapting to unknown low\-dimensional structures in score\-based diffusion models\.*Advances in Neural Information Processing Systems*, 37:126297–126331, 2024\.
- Chen et al\. \[2023\]Minshuo Chen, Kaixuan Huang, Tuo Zhao, and Mengdi Wang\.Score approximation, estimation and distribution recovery of diffusion models on low\-dimensional data\.In*International Conference on Machine Learning*, pages 4672–4712\. PMLR, 2023\.
- Li et al\. \[2024\]Sixu Li, Shi Chen, and Qin Li\.A good score does not lead to a good generative model\.*arXiv preprint arXiv:2401\.04856*, 2024\.
- Scarvelis et al\. \[2023\]Christopher Scarvelis, Haitz Sáez de Ocáriz Borde, and Justin Solomon\.Closed\-form diffusion models\.*arXiv preprint arXiv:2310\.12395*, 2023\.
- Kamb and Ganguli \[2025\]Mason Kamb and Surya Ganguli\.An analytic theory of creativity in convolutional diffusion models\.In*Forty\-second International Conference on Machine Learning*, 2025\.
- Niedoba et al\. \[2025\]Matthew Niedoba, Berend Zwartsenberg, Kevin Patrick Murphy, and Frank Wood\.Towards a mechanistic explanation of diffusion model generalization\.In*Forty\-second International Conference on Machine Learning*, 2025\.
- Yi et al\. \[2023\]Mingyang Yi, Jiacheng Sun, and Zhenguo Li\.On the generalization of diffusion model\.*arXiv preprint arXiv:2305\.14712*, 2023\.
- Gu et al\. \[2023\]Xiangming Gu, Chao Du, Tianyu Pang, Chongxuan Li, Min Lin, and Ye Wang\.On memorization in diffusion models\.*arXiv preprint arXiv:2310\.02664*, 2023\.
- Pham et al\. \[2025\]Bao Pham, Gabriel Raya, Matteo Negri, Mohammed J Zaki, Luca Ambrogioni, and Dmitry Krotov\.Memorization to generalization: Emergence of diffusion models from associative memory\.*arXiv preprint arXiv:2505\.21777*, 2025\.
- Chen \[2026\]Zhengdao Chen\.On the interpolation effect of score smoothing in diffusion models\.In*The Fourteenth International Conference on Learning Representations*, 2026\.URLhttps://openreview\.net/forum?id=O33LAUliUF\.
- Wang and Vastola \[2024\]Binxu Wang and John J Vastola\.The unreasonable effectiveness of gaussian score approximation for diffusion models and its applications\.*arXiv preprint arXiv:2412\.09726*, 2024\.
- Wang et al\. \[2026\]Binxu Wang, Jacob Zavatone\-Veth, and Cengiz Pehlevan\.A random matrix theory perspective on the consistency of diffusion models\.*arXiv preprint arXiv:2602\.02908*, 2026\.
- He et al\. \[2026\]Ye He, Yitong Qiu, and Molei Tao\.Diffusion model’s generalization can be characterized by inductive biases toward a data\-dependent ridge manifold\.*arXiv preprint arXiv:2602\.06021*, 2026\.
- Wang et al\. \[2025\]Peng Wang, Huijie Zhang, Zekai Zhang, Siyi Chen, Yi Ma, and Qing Qu\.Diffusion models learn low\-dimensional distributions via subspace clustering\.In*2025 IEEE 10th International Workshop on Computational Advances in Multi\-Sensor Adaptive Processing \(CAMSAP\)*, pages 211–215\. IEEE, 2025\.
- Favero et al\. \[2025a\]Alessandro Favero, Antonio Sclocchi, and Matthieu Wyart\.Bigger isn’t always memorizing: Early stopping overparameterized diffusion models\.*arXiv preprint arXiv:2505\.16959*, 2025a\.
- Bonnaire et al\. \[2026\]Tony Bonnaire, Raphaël Urfin, Giulio Biroli, and Marc Mezard\.Why diffusion models don’t memorize: The role of implicit dynamical regularization in training\.In*The Thirty\-ninth Annual Conference on Neural Information Processing Systems*, 2026\.URLhttps://openreview\.net/forum?id=BSZqpqgqM0\.
- Favero et al\. \[2025b\]Alessandro Favero, Antonio Sclocchi, Francesco Cagnetta, Pascal Frossard, and Matthieu Wyart\.How compositional generalization and creativity improve as diffusion models are trained\.In*International Conference on Machine Learning*, pages 16286–16306\. PMLR, 2025b\.
- Bardone et al\. \[2026\]Lorenzo Bardone, Claudia Merger, and Sebastian Goldt\.A theory of learning data statistics in diffusion models, from easy to hard\.*arXiv preprint arXiv:2603\.12901*, 2026\.
- Boffi et al\. \[2024\]Nicholas M Boffi, Arthur Jacot, Stephen Tu, and Ingvar Ziemann\.Shallow diffusion networks provably learn hidden low\-dimensional structure\.*arXiv preprint arXiv:2410\.11275*, 2024\.
- Cui et al\. \[2025\]Hugo Cui, Cengiz Pehlevan, and Yue M Lu\.A solvable model of learning generative diffusion: theory and insights\.*arXiv preprint arXiv:2501\.03937*, 2025\.
- Yoon et al\. \[2023\]TaeHo Yoon, Joo Young Choi, Sehyun Kwon, and Ernest K Ryu\.Diffusion probabilistic models generalize when they fail to memorize\.In*ICML 2023 workshop on structured probabilistic inference\{\\\{\\\\backslash&\}\\\}generative modeling*, 2023\.
- Cui et al\. \[2024\]Hugo Cui, Florent Krzakala, Eric Vanden\-Eijnden, and Lenka Zdeborová\.Analysis of learning a flow\-based generative model from limited sample complexity\.In*International Conference on Learning Representations*, volume 2024, pages 51929–51955, 2024\.
- Buchanan et al\. \[2026\]Sam Buchanan, Druv Pai, Yi Ma, and Valentin De Bortoli\.On the edge of memorization in diffusion models\.*Advances in Neural Information Processing Systems*, 38:96113–96157, 2026\.
- Lukoianov et al\. \[2025\]Artem Lukoianov, Chenyang Yuan, Justin Solomon, and Vincent Sitzmann\.Locality in image diffusion models emerges from data statistics\.*arXiv preprint arXiv:2509\.09672*, 2025\.
- Bradley \[2025\]Arwen Bradley\.Local mechanisms of compositional generalization in conditional diffusion\.*arXiv preprint arXiv:2509\.16447*, 2025\.
- Finn et al\. \[2025\]Emma Finn, T Anderson Keller, Manos Theodosis, and Demba E Ba\.Origins of creativity in attention\-based diffusion models\.*arXiv preprint arXiv:2506\.17324*, 2025\.
- Hu et al\. \[2025\]Fangjun Hu, Guangkuo Liu, Yifan Zhang, and Xun Gao\.Local diffusion models and phases of data distributions\.*arXiv preprint arXiv:2508\.06614*, 2025\.
- Ambrogioni \[2026\]Luca Ambrogioni\.How out\-of\-equilibrium phase transitions can seed pattern formation in trained diffusion models\.*arXiv preprint arXiv:2603\.20092*, 2026\.
- Gottwald et al\. \[2025\]Georg A Gottwald, Shuigen Liu, Youssef Marzouk, Sebastian Reich, and Xin T Tong\.Localized diffusion models\.*arXiv preprint arXiv:2505\.04417*, 2025\.
- Nguyen et al\. \[2026\]Minh Hai Nguyen, Quoc Bao Do, Edouard Pauwels, and Pierre Weiss\.An analytic theory of convolutional neural network inverse problems solvers\.*arXiv preprint arXiv:2601\.10334*, 2026\.
- Ruderman and Bialek \[1993\]Daniel Ruderman and William Bialek\.Statistics of natural images: Scaling in the woods\.*Advances in neural information processing systems*, 6, 1993\.
- Dieleman \[2024\]Sander Dieleman\.Diffusion is spectral autoregression, 2024\.URLhttps://sander\.ai/2024/09/02/spectral\-autoregression\.html\.
- Wiener \[1949\]Norbert Wiener\.*Extrapolation, interpolation, and smoothing of stationary time series: with engineering applications*\.The MIT press, 1949\.
- Karras et al\. \[2022\]Tero Karras, Miika Aittala, Timo Aila, and Samuli Laine\.Elucidating the design space of diffusion\-based generative models\.*Advances in neural information processing systems*, 35:26565–26577, 2022\.
- Achilli et al\. \[2026\]Beatrice Achilli, Marco Benedetti, Giulio Biroli, and Marc Mézard\.Theory of speciation transitions in diffusion models with general class structure, 2026\.URLhttps://arxiv\.org/abs/2602\.04404\.
- Derrida \[1980\]B\. Derrida\.Random\-energy model: Limit of a family of disordered models\.*Phys\. Rev\. Lett\.*, 45:79–82, Jul 1980\.doi:10\.1103/PhysRevLett\.45\.79\.URLhttps://link\.aps\.org/doi/10\.1103/PhysRevLett\.45\.79\.
- Mézard and Montanari \[2009\]Marc Mézard and Andrea Montanari\.*Information, Physics, and Computation*\.Oxford University PressOxford, January 2009\.ISBN 9780191718755\.doi:10\.1093/acprof:oso/9780198570837\.001\.0001\.URLhttp://dx\.doi\.org/10\.1093/acprof:oso/9780198570837\.001\.0001\.
- Cover and Thomas \[2006\]Thomas M\. Cover and Joy A\. Thomas\.*Elements of Information Theory \(Wiley Series in Telecommunications and Signal Processing\)*\.Wiley\-Interscience, USA, 2006\.ISBN 0471241954\.
- Ventura et al\. \[2024\]Enrico Ventura, Beatrice Achilli, Gianluigi Silvestri, Carlo Lucibello, and Luca Ambrogioni\.Manifolds, random matrices and spectral gaps: The geometric phases of generative diffusion\.*arXiv preprint arXiv:2410\.05898*, 2024\.
- Niedoba et al\. \[2024\]Matthew Niedoba, Berend Zwartsenberg, Kevin Murphy, and Frank Wood\.Towards a mechanistic explanation of diffusion model generalization\.*arXiv preprint arXiv:2411\.19339*, 2024\.
- Yu et al\. \[2024\]Hu Yu, Li Shen, Jie Huang, Hongsheng Li, and Feng Zhao\.Unmasking bias in diffusion model training\.In*European Conference on Computer Vision*, pages 374–390\. Springer, 2024\.
- Deck and Bischoff \[2023\]Katherine Deck and Tobias Bischoff\.Easing color shifts in score\-based diffusion models\.*arXiv preprint arXiv:2306\.15832*, 2023\.
- Choi et al\. \[2022\]Jooyoung Choi, Jungbeom Lee, Chaehun Shin, Sungwon Kim, Hyunwoo Kim, and Sungroh Yoon\.Perception prioritized training of diffusion models\.In*Proceedings of the IEEE/CVF conference on computer vision and pattern recognition*, pages 11472–11481, 2022\.
- Salimans and Ho \[2022\]Tim Salimans and Jonathan Ho\.Progressive distillation for fast sampling of diffusion models\.*arXiv preprint arXiv:2202\.00512*, 2022\.
\\appendixpage\\addappheadtotoc
## Appendix ANotation and background on diffusion models\.
### A\.1Notation conventions
We use the following notation:
- •φ∈ℝd\\varphi\\in\\mathbb\{R\}^\{d\}denotes an example from the training set\. For images or patches of sizeLLpixels byLLpixels byCCcolor channels, we haved=L×L×Cd=L\\times L\\times C\.
- •ϕ\\phirepresents any arbitrary image \(or other data\) that can serve as input to either the score function or diffusion model\.
- •P0\(φ\)P\_\{0\}\(\\varphi\)denotes the true data \(i\.e\. target\) distribution, from which we sample both our training and test data\.
- •𝒟\\mathcal\{D\}represents a finite set of training data consisting of\|𝒟\|\|\\mathcal\{D\}\|training samples\. Each training data pointφ∈𝒟\\varphi\\in\\mathcal\{D\}is drawn i\.i\.d\. from the true data distributionφ∼P0\\varphi\\sim P\_\{0\}\.
- •Ptrain\(φ\)=1\|𝒟\|∑φ′∈𝒟δ\(φ−φ′\)P\_\{train\}\(\\varphi\)=\\frac\{1\}\{\|\\mathcal\{D\}\|\}\\sum\_\{\\varphi^\{\\prime\}\\in\\mathcal\{D\}\}\\delta\(\\varphi\-\\varphi^\{\\prime\}\)denotes the empirical training distribution for any fixed realization𝒟\\mathcal\{D\}of the training set\. It is simply a uniform distribution over\|𝒟\|\|\\mathcal\{D\}\|specific training examplesφ∈𝒟\\varphi\\in\\mathcal\{D\}\.
- •xxrepresents a pixel location in an image\.
- •For image data,ϕx\\phi\_\{x\}andφx\\varphi\_\{x\}will represent the pixel values of the imagesϕ\\phiandφ\\varphiat pixel locationxx; both are elements ofℝC\\mathbb\{R\}^\{C\}\.
- •M\[ϕ\]:ℝd→ℝdM\[\\phi\]:\\mathbb\{R\}^\{d\}\\to\\mathbb\{R\}^\{d\}represents a model that takes as input an imageϕ\\phiand produces a new image \(e\.g\. an estimate of the score function, or a denoised version ofϕ\\phi\)\. We will denote byMx\[ϕ\]∈ℝCM\_\{x\}\[\\phi\]\\in\\mathbb\{R\}^\{C\}the value of the outputs of this model, given an inputϕ\\phi, at the pixel locationxx\.
- •Ωx\\Omega\_\{x\}denotes a patch neighborhood indexed by the pixelxx\. The pixel locationxxwill be contained withinΩx\\Omega\_\{x\}\. We will denote by\|Ωx\|\|\\Omega\_\{x\}\|the number of pixels in the patch\. We will also denote byΩx,L\\Omega\_\{x,L\}a patch neighborhood ofxxwith a characteristic lengthLL, such that the size of the patch is\|Ωx;L\|=L2\|\\Omega\_\{x;L\}\|=L^\{2\}\.
- •ϕΩx\\phi\_\{\\Omega\_\{x\}\}andφΩx\\varphi\_\{\\Omega\_\{x\}\}denotes the restriction of imagesϕ\\phiandφ\\varphito a neighborhoodΩx\\Omega\_\{x\}around a pixelxx\. For images withCCchannels, we haveϕΩx∈ℝ\|Ωx\|×C\\phi\_\{\\Omega\_\{x\}\}\\in\\mathbb\{R\}^\{\|\\Omega\_\{x\}\|\\times C\}\.
- •𝒞x,t:ℝd→ℝn\\mathcal\{C\}\_\{x,t\}:\\mathbb\{R\}^\{d\}\\to\\mathbb\{R\}^\{n\}denotes a \(possibly stochastic\) observation map that takes as input a noised imageϕt∈ℝd\\phi\_\{t\}\\in\\mathbb\{R\}^\{d\}, and outputs thenndimensional observation𝒞x,t\(ϕt\)\\mathcal\{C\}\_\{x,t\}\(\\phi\_\{t\}\)that pixelxxmakes about the imageϕt\\phi\_\{t\}\. If the map is stochastic we denote its conditional distribution byP\(𝒞x,t\|ϕt\)P\(\\mathcal\{C\}\_\{x,t\}\|\\phi\_\{t\}\)\. If the map is deterministic we denote it by the function𝒞x,t\(ϕt\)\\mathcal\{C\}\_\{x,t\}\(\\phi\_\{t\}\)\. For brevity we also sometimes drop the argumentϕt\\phi\_\{t\}, and then𝒞x,t∈ℝn\\mathcal\{C\}\_\{x,t\}\\in\\mathbb\{R\}^\{n\}denotes the pixelxx’s observation value, reflecting the state of knowledge of pixelxxat timettin the diffusion process\.
- •We denote thetrainingMarkov processφ→ϕt→𝒞x,t\\varphi\\rightarrow\\phi\_\{t\}\\rightarrow\\mathcal\{C\}\_\{x,t\}for pixelxxat timettby the joint distributionPtrain\(φ,ϕt,𝒞x,t\)=Ptrain\(φ\)P\(ϕt\|φ\)P\(𝒞x,t\|ϕt\)P\_\{train\}\(\\varphi,\\phi\_\{t\},\\mathcal\{C\}\_\{x,t\}\)=P\_\{train\}\(\\varphi\)P\(\\phi\_\{t\}\|\\varphi\)P\(\\mathcal\{C\}\_\{x,t\}\|\\phi\_\{t\}\)\. Hereφ∼Ptrain\(φ\)\\varphi\\sim P\_\{train\}\(\\varphi\)is a randomly chosen training imageφ∈𝒟\\varphi\\in\\mathcal\{D\},P\(ϕt\|φ\)P\(\\phi\_\{t\}\|\\varphi\)is the conditional distribution of a noisy imageϕt\\phi\_\{t\}under the forward diffusion process, starting from the training imageφ\\varphi, andP\(𝒞x,t\|ϕt\)P\(\\mathcal\{C\}\_\{x,t\}\|\\phi\_\{t\}\)is the \(possibly stochastic\) observation that pixelxxmakes about imageϕt\\phi\_\{t\}\. We can think about the Markov processφ→ϕt→𝒞x,t\\varphi\\rightarrow\\phi\_\{t\}\\rightarrow\\mathcal\{C\}\_\{x,t\}as an information channel from the past training data𝒟\\mathcal\{D\}at timet=0t=0to the current pixel observation𝒞x,t\\mathcal\{C\}\_\{x,t\}at timett\. We denote byPtrain\(φ\|𝒞x,t\)P\_\{train\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)the Bayesian posterior distribution over the training data given a pixelxx’s channel observation𝒞x,t\\mathcal\{C\}\_\{x,t\}\. This posterior assigns a probability to each of the\|𝒟\|\|\\mathcal\{D\}\|training imagesφ∈𝒟\\varphi\\in\\mathcal\{D\}\. If a pixelxxhad to guess, based on its \(restricted\) observation𝒞x,t\\mathcal\{C\}\_\{x,t\}of the noisy imageϕt\\phi\_\{t\}, which clean training imageφ∈𝒟\\varphi\\in\\mathcal\{D\}at time0in the past led toϕt\\phi\_\{t\}under the forward process, then any optimal guess would be based on the posteriorPtrain\(φ\|𝒞x,t\)P\_\{train\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\), which encapsulates the pixel’s belief state about the past training data origin of its current observation\.
- •We denote thetestingMarkov processφ→ϕt→𝒞x,t\\varphi\\rightarrow\\phi\_\{t\}\\rightarrow\\mathcal\{C\}\_\{x,t\}for pixelxxat timettby the joint distributionPtest\(φ,ϕt,𝒞x,t\)=P0\(φ\)P\(ϕt\|φ\)P\(𝒞x,t\|ϕt\)\.P\_\{test\}\(\\varphi,\\phi\_\{t\},\\mathcal\{C\}\_\{x,t\}\)=P\_\{0\}\(\\varphi\)P\(\\phi\_\{t\}\|\\varphi\)P\(\\mathcal\{C\}\_\{x,t\}\|\\phi\_\{t\}\)\.This is identical to the training Markov process except the prior is now thetruedata distributionP0\(φ\)P\_\{0\}\(\\varphi\)from which a test sample is drawn, rather than the empirical training distributionPtrain\(φ\)P\_\{train\}\(\\varphi\)for a specific dataset𝒟\\mathcal\{D\}\. We denote the posterior distribution of the input given the output under this Markov process byPtest\(φ\|𝒞x,t\)P\_\{test\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\. UnlikePtrain\(φ\|𝒞x,t\)P\_\{train\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\), the support ofPtest\(φ\|𝒞x,t\)P\_\{test\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)need not be restricted to the training dataφ∈𝒟\\varphi\\in\\mathcal\{D\}\.
- •α¯t\\bar\{\\alpha\}\_\{t\}denotes the signal variance in the forward diffusion process\.
- •σt2=1−α¯tα¯t\\sigma\_\{t\}^\{2\}=\\frac\{1\-\\bar\{\\alpha\}\_\{t\}\}\{\\bar\{\\alpha\}\_\{t\}\}denotes the noise\-to\-signal ratio at a timett\.
- •I\(A;B\)I\(A;B\)will indicate the mutual information between two random variablesAAandBB\.I\(A;B∼P\)I\(A;B\\sim P\)will indicate the mutual information between these variables with a particular emphasis that the random variableBBis distributed according toPP\.
- •S\[P\]S\[P\]will denote the entropy of a distributionPP\.SP\[A\]S\_\{P\}\[A\]will indicate the entropy of the random variableAAunder the distributionPP\.
- •𝒩\(x\|μ,Σ\)\\mathcal\{N\}\(x\|\\mu,\\Sigma\)represents the PDF of the normal distribution with meanμ\\muand covarianceΣ\\Sigma\. We also use the short\-hand𝒩\(μ,Σ\)\\mathcal\{N\}\(\\mu,\\Sigma\)when we do not need to refer to the name of a specific random variable\.
### A\.2Background on diffusion models
Diffusion models are designed to solve the problem of generating samples from a complex data distributionφ∼P0\(φ\)\\varphi\\sim P\_\{0\}\(\\varphi\)in high dimensions \(e\.g\. forφ∈ℝd\\varphi\\in\\mathbb\{R\}^\{d\}for largedd\)\. To do so, they first develop, through a forward diffusion process, an intermediate bridge between the complex data distributionP0\(φ\)P\_\{0\}\(\\varphi\)and simple isotropic Gaussian noise distributionη∼𝒩\(0,I\)\\eta\\sim\\mathcal\{N\}\(0,I\)\. This bridge corresponds to a time dependent family of distributions of noised dataϕt∼Pt\(ϕt\)\\phi\_\{t\}\\sim P\_\{t\}\(\\phi\_\{t\}\)that interpolate between data pointsφ\\varphiat timet=0t=0and pure isotropic noiseη\\etaat timet=1t=1\. A single sample from this bridge at any fixed timettcan be obtained via the interpolation
ϕt\(φ,η\)=α¯tφ\+1−α¯tη\.\\displaystyle\\phi\_\{t\}\(\\varphi,\\eta\)=\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\varphi\+\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\}\}\\eta\.\(9\)The parameterα¯t\\bar\{\\alpha\}\_\{t\}can be thought of as the signal power, while1−α¯t1\-\\bar\{\\alpha\}\_\{t\}can be thought of as the noise powerα¯t\\bar\{\\alpha\}\_\{t\}\. The temporal schedule of the signal power is chosen to monotonically decrease fromα¯0=1\\bar\{\\alpha\}\_\{0\}=1at timet=0t=0toα¯1=0\\bar\{\\alpha\}\_\{1\}=0at timet=1t=1\. Thus over the time intervalt∈\[0,1\]t\\in\[0,1\], the forward diffusion process gradually turns dataφ\\varphiinto noiseη\\etathrough the family of noised samplesϕt\\phi\_\{t\}\. We define a useful quantityσt2\\sigma^\{2\}\_\{t\}, which characterizes the noise to signal ratio at a timett:
σt2=1−α¯tα¯t\.\\displaystyle\\sigma\_\{t\}^\{2\}=\\frac\{1\-\\bar\{\\alpha\}\_\{t\}\}\{\\bar\{\\alpha\}\_\{t\}\}\.\(10\)\(This parameter is the direct analog of theσt2\\sigma\_\{t\}^\{2\}parameter in variance\-exploding forward processes\[[42](https://arxiv.org/html/2607.08041#bib.bib42)\]\)\. The interpolation in \([9](https://arxiv.org/html/2607.08041#A1.E9)\) induces a flow on distributionsPt\(ϕt\)P\_\{t\}\(\\phi\_\{t\}\)defined by the forward diffusion partial differential equation
dPt\(ϕt\)dt=γt∇⋅\(ϕtPt\(ϕt\)\)\+γt∇2Pt\(ϕt\),\\displaystyle\\frac\{dP\_\{t\}\(\\phi\_\{t\}\)\}\{dt\}=\\gamma\_\{t\}\\nabla\\cdot\(\\phi\_\{t\}P\_\{t\}\(\\phi\_\{t\}\)\)\+\\gamma\_\{t\}\\nabla^\{2\}P\_\{t\}\(\\phi\_\{t\}\),\(11\)withγt=−12∂tlnα¯t\\gamma\_\{t\}=\-\\frac\{1\}\{2\}\\partial\_\{t\}\\ln\\bar\{\\alpha\}\_\{t\}\. Diffusion models in deterministic mode \(DDIM parameterization,\[[3](https://arxiv.org/html/2607.08041#bib.bib3)\]\) aim to reverse this flow by evolving pure noise samplesϕ1∼𝒩\(0,I\)\\phi\_\{1\}\\sim\\mathcal\{N\}\(0,I\)backwards in time under the ODE
dϕtdt=−γt\(ϕt\+∇logPt\(ϕt\)\),\\displaystyle\\frac\{d\\phi\_\{t\}\}\{dt\}=\-\\gamma\_\{t\}\(\\phi\_\{t\}\+\\nabla\\log P\_\{t\}\(\\phi\_\{t\}\)\),\(12\)where∇logPt\\nabla\\log P\_\{t\}is thescore functionfor the distributionPt\(ϕt\)P\_\{t\}\(\\phi\_\{t\}\)over noised samples induced by the interpolation in \([9](https://arxiv.org/html/2607.08041#A1.E9)\) or equivalently the forward diffusion equation in \([11](https://arxiv.org/html/2607.08041#A1.E11)\)\. This ODE provably transports the distribution𝒩\(0,I\)\\mathcal\{N\}\(0,I\)back to the data distributionP0\(φ\)P\_\{0\}\(\\varphi\)through the intermediate family of distributionsPt\(ϕt\)P\_\{t\}\(\\phi\_\{t\}\)as time flows backwards fromt=1t=1tot=0t=0\. Thus the reverse process in \([12](https://arxiv.org/html/2607.08041#A1.E12)\) exactly reverses the forward diffusion in \([11](https://arxiv.org/html/2607.08041#A1.E11)\) in terms of the time\-dependent family of marginal distributionsPt\(ϕt\)P\_\{t\}\(\\phi\_\{t\}\)\.
However, implementing the reverse process in \([12](https://arxiv.org/html/2607.08041#A1.E12)\) requires access to the score function∇logPt\\nabla\\log P\_\{t\}\. When direct calculation of the score is not available or not desirable, there is a convenient identity that reframes the score computation as a problem of optimal denoising, known as Tweedie’s theorem:
∇logPt\(ϕt\)=−𝔼\[η\|ϕt\]1−α¯t=α¯t1−α¯t𝔼\[φ\|ϕt\]−ϕt1−α¯t\\displaystyle\\nabla\\log P\_\{t\}\(\\phi\_\{t\}\)=\-\\frac\{\\mathbb\{E\}\[\\eta\|\\phi\_\{t\}\]\}\{\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\}\}\}=\\frac\{\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\}\{1\-\\bar\{\\alpha\}\_\{t\}\}\\mathbb\{E\}\[\\varphi\|\\phi\_\{t\}\]\-\\frac\{\\phi\_\{t\}\}\{1\-\\bar\{\\alpha\}\_\{t\}\}\(13\)Thus the calculation of the posterior mean𝔼\[φ\|ϕt\]\\mathbb\{E\}\[\\varphi\|\\phi\_\{t\}\]over the dataφ\\varphiat timet=0t=0, conditioned on the noised sampleϕt\\phi\_\{t\}at any pointt\>0t\>0in the reverse flow, becomes a central quantity required to successfully implement the reverse flow to transportPt\(ϕt\)P\_\{t\}\(\\phi\_\{t\}\)back toP0\(ϕ0=φ\)P\_\{0\}\(\\phi\_\{0\}=\\varphi\)\. In turn, one can approximate this posterior mean by minimizing the following squared loss over the parametersθ\\thetaof a neural networkMt,θ\[ϕt\]M\_\{t,\\theta\}\[\\phi\_\{t\}\]:
ℒt=12𝔼φ∼P0,η∼𝒩\(0,I\)\[‖φ−Mt,θ\[ϕt\(φ,η\)\]‖2\]\.\\displaystyle\\mathcal\{L\}\_\{t\}=\\frac\{1\}\{2\}\\mathbb\{E\}\_\{\\varphi\\sim P\_\{0\},\\eta\\sim\\mathcal\{N\}\(0,I\)\}\[\\norm\{\\varphi\-M\_\{t,\\theta\}\[\\phi\_\{t\}\(\\varphi,\\eta\)\]\}^\{2\}\]\.\(14\)This loss is a denoising loss that trains the neural network functionMt,θM\_\{t,\\theta\}to take any noisy sampleϕt\\phi\_\{t\}at timettas input, and convert it back to the clean sampleφ\\varphithat led toϕt\\phi\_\{t\}under the forward diffusion process\. It is well known that the minimum mean squared error \(MMSE\) solution to any optimization of the form in \([14](https://arxiv.org/html/2607.08041#A1.E14)\) will yield the posterior mean𝔼\[φ\|ϕt\]\\mathbb\{E\}\[\\varphi\|\\phi\_\{t\}\]as the Bayes optimal estimator or denoiser\. More precisely, if the family of neural network functionsMt,θM\_\{t,\\theta\}is expressive enough to compute this posterior mean, then for the optimalθ∗\\theta^\{\*\}that achieves zero loss in \([14](https://arxiv.org/html/2607.08041#A1.E14)\) we have the MMSE result
Mt,θ∗\[ϕt\]=𝔼\[φ\|ϕt\]\.M\_\{t,\\theta^\{\*\}\}\[\\phi\_\{t\}\]=\\mathbb\{E\}\[\\varphi\|\\phi\_\{t\}\]\.\(15\)Learning the reverse process from a finite training set and then implementing it through a parametric neural networkMt,θM\_\{t,\\theta\}then corresponds at a high level to four main steps:
1. 1\.Replace the average over the true data distributionP0\(φ\)P\_\{0\}\(\\varphi\)with an empirical average over a finite training set\.
2. 2\.Minimize the loss in \([14](https://arxiv.org/html/2607.08041#A1.E14)\) to obtain the learned neural networkM\[t,θ\]\(ϕt\)M\[t,\\theta\]\(\\phi\_\{t\}\)\.
3. 3\.Replace the posterior mean𝔼\[φ\|ϕt\]\\mathbb\{E\}\[\\varphi\|\\phi\_\{t\}\]in Tweedie’s formula in \([13](https://arxiv.org/html/2607.08041#A1.E13)\) with its approximationMt,θ\[ϕt\]M\_\{t,\\theta\}\[\\phi\_\{t\}\], thereby obtaining an approximate score function\.
4. 4\.Insert this approximate score into the deterministic ODE representing the reverse process in \([12](https://arxiv.org/html/2607.08041#A1.E12)\)\.
In practice, for purposes of variance reduction, rather than predicting the clean dataφ\\varphifrom the noisy sampleϕt\\phi\_\{t\}, one can also predict the noiseη\\etaor the velocityvtv\_\{t\}through the following two losses respectively:
ℒtη\\displaystyle\\mathcal\{L\}^\{\\eta\}\_\{t\}=12𝔼φ∼P0,η∼𝒩\(0,I\)\[‖η−Mt,θ\[ϕt\(φ,η\)\]‖2\]\\displaystyle=\\frac\{1\}\{2\}\\mathbb\{E\}\_\{\\varphi\\sim P\_\{0\},\\eta\\sim\\mathcal\{N\}\(0,I\)\}\[\\norm\{\\eta\-M\_\{t,\\theta\}\[\\phi\_\{t\}\(\\varphi,\\eta\)\]\}^\{2\}\]\(16\)ℒtv\\displaystyle\\mathcal\{L\}^\{v\}\_\{t\}=12𝔼φ∼P0,η∼𝒩\(0,I\)\[‖vt\(φ,η\)−Mt,θ\[ϕt\(φ,η\)\]‖2\]\.\\displaystyle=\\frac\{1\}\{2\}\\mathbb\{E\}\_\{\\varphi\\sim P\_\{0\},\\eta\\sim\\mathcal\{N\}\(0,I\)\}\[\\norm\{v\_\{t\}\(\\varphi,\\eta\)\-M\_\{t,\\theta\}\[\\phi\_\{t\}\(\\varphi,\\eta\)\]\}^\{2\}\]\.\(17\)Here the velocityvt\(φ,η\)v\_\{t\}\(\\varphi,\\eta\)roughly points from the clean dataφ\\varphito the noisy sampleϕt\\phi\_\{t\}with weighting coefficients:
vt\(φ,η\)=α¯t1−α¯tϕt−φ1−α¯t\.\\displaystyle v\_\{t\}\(\\varphi,\\eta\)=\\sqrt\{\\frac\{\\bar\{\\alpha\}\_\{t\}\}\{1\-\\bar\{\\alpha\}\_\{t\}\}\}\\,\\phi\_\{t\}\-\\frac\{\\varphi\}\{\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\}\}\}\.In analogy to \([15](https://arxiv.org/html/2607.08041#A1.E15)\), the MMSE neural network \(assuming no expressivity constraints\) computes the following posterior means for the lossesℒtη\\mathcal\{L\}^\{\\eta\}\_\{t\}andℒtv\\mathcal\{L\}^\{v\}\_\{t\}respectively:
𝔼\[η\|ϕt\]\\displaystyle\\mathbb\{E\}\[\\eta\|\\phi\_\{t\}\]=ϕt1−α¯t¯−α¯t1−α¯t𝔼\[φ\|ϕt\]\\displaystyle=\\frac\{\\phi\_\{t\}\}\{\\sqrt\{1\-\\bar\{\\bar\{\\alpha\}\_\{t\}\}\}\}\-\\sqrt\{\\frac\{\\bar\{\\alpha\}\_\{t\}\}\{1\-\\bar\{\\alpha\}\_\{t\}\}\}\\,\\,\\mathbb\{E\}\[\\varphi\|\\phi\_\{t\}\]𝔼\[vt\(ϕt,η\)\|ϕt\]\\displaystyle\\mathbb\{E\}\[v\_\{t\}\(\\phi\_\{t\},\\eta\)\|\\phi\_\{t\}\]=α¯t1−α¯tϕt−11−α¯t𝔼\[φ\|ϕt\]\.\\displaystyle=\\sqrt\{\\frac\{\\bar\{\\alpha\}\_\{t\}\}\{1\-\\bar\{\\alpha\}\_\{t\}\}\}\\,\\phi\_\{t\}\-\\frac\{1\}\{\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\}\}\}\\mathbb\{E\}\[\\varphi\|\\phi\_\{t\}\]\.The MMSE solutions to all three losses above produce a linear combination ofϕt\\phi\_\{t\}and the optimal MMSE denoiser𝔼\[φ\|ϕt\]\\mathbb\{E\}\[\\varphi\|\\phi\_\{t\}\]\. Thus we can easily interchange theoretical results derived for one approach to any other approach\. For the remainder of the theory section, we will focus primarily on the loss \([14](https://arxiv.org/html/2607.08041#A1.E14)\) for which the optimal model simply computes𝔼\[φ\|ϕt\]\\mathbb\{E\}\[\\varphi\|\\phi\_\{t\}\]\. However, for numerical experiments in our paper, we use velocity prediction, which shows significant benefits from the perspective of training stability and the stability of generated outputs, especially early in training\.
## Appendix BBayes\-optimal diffusion models always memorize finite training sets
In the denoising loss in \([14](https://arxiv.org/html/2607.08041#A1.E14)\) we never have direct access to the true data distributionP0\(φ\)P\_\{0\}\(\\varphi\)\. Instead, we only have afinitedataset𝒟\\mathcal\{D\}consisting of\|𝒟\|\|\\mathcal\{D\}\|training samplesφ∈𝒟\\varphi\\in\\mathcal\{D\}each drawn i\.i\.d from the true data distributionP0\(φ\)P\_\{0\}\(\\varphi\)\. This leads to an empirical training distribution
Ptrain\(φ\)=1\|𝒟\|∑φ′∈𝒟δ\(φ−φ′\),P\_\{train\}\(\\varphi\)=\\frac\{1\}\{\|\\mathcal\{D\}\|\}\\sum\_\{\\varphi^\{\\prime\}\\in\\mathcal\{D\}\}\\delta\(\\varphi\-\\varphi^\{\\prime\}\),\(18\)which is simply a uniform sum of delta\-functions over\|𝒟\|\|\\mathcal\{D\}\|specific training examplesφ∈𝒟\\varphi\\in\\mathcal\{D\}\. It follows that the Bayes optimal MMSE denoiser that achieves minimal loss in \([14](https://arxiv.org/html/2607.08041#A1.E14)\), whenP0\(φ\)P\_\{0\}\(\\varphi\)replaced withPtrain\(φ\)P\_\{train\}\(\\varphi\), computes the posterior mean over thefinitetraining set:
Mt\[ϕt\]=𝔼\[φ\|ϕt\]=∑φ∈𝒟φPtrain\(φ\|ϕt\)\.\\displaystyle M\_\{t\}\[\\phi\_\{t\}\]=\\mathbb\{E\}\[\\varphi\|\\phi\_\{t\}\]=\\sum\_\{\\varphi\\in\\mathcal\{D\}\}\\varphi\\,P\_\{train\}\(\\varphi\|\\phi\_\{t\}\)\.\(19\)HerePtrain\(φ\|ϕt\)P\_\{train\}\(\\varphi\|\\phi\_\{t\}\)is the posterior probability that a clean training data imageφ∈𝒟\\varphi\\in\\mathcal\{D\}at timet=0t=0leads to the noisy observed imageϕt\\phi\_\{t\}at a given timettunder the forward diffusion process\. Because the forward diffusion process shrinks the training data and adds isotropic Gaussian noise to it \(see \([9](https://arxiv.org/html/2607.08041#A1.E9)\)\), the forward conditional probability of obtainingϕt\\phi\_\{t\}starting fromφ∈𝒟\\varphi\\in\\mathcal\{D\}is simply an isotropic Gaussian distribution whose mean shrinks with time and whose variance grows with time:
P\(ϕt\|φ\)=𝒩\(ϕt\|α¯tφ,\(1−α¯t\)I\)\.\\displaystyle P\(\\phi\_\{t\}\|\\varphi\)=\\mathcal\{N\}\(\\phi\_\{t\}\|\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\varphi,\(1\-\\bar\{\\alpha\}\_\{t\}\)I\)\.\(20\)And since the prior over the training data in \([18](https://arxiv.org/html/2607.08041#A2.E18)\) is uniform, Bayes rule simply yields
Ptrain\(φ\|ϕt\)=P\(ϕt\|φ\)∑φ′P\(ϕt\|φ′\)=𝒩\(ϕt\|α¯tφ,\(1−α¯t\)I\)∑φ′∈𝒟𝒩\(ϕt\|α¯tφ′,\(1−α¯t\)I\)\.\\displaystyle P\_\{train\}\(\\varphi\|\\phi\_\{t\}\)=\\frac\{P\(\\phi\_\{t\}\|\\varphi\)\}\{\\sum\_\{\\varphi^\{\\prime\}\}P\(\\phi\_\{t\}\|\\varphi^\{\\prime\}\)\}=\\frac\{\\mathcal\{N\}\(\\phi\_\{t\}\|\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\varphi,\(1\-\\bar\{\\alpha\}\_\{t\}\)I\)\}\{\\sum\_\{\\varphi^\{\\prime\}\\in\\mathcal\{D\}\}\\mathcal\{N\}\(\\phi\_\{t\}\|\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\varphi^\{\\prime\},\(1\-\\bar\{\\alpha\}\_\{t\}\)I\)\}\.\(21\)
In summary, the Bayesian posterior distributionPtrain\(φ\|ϕt\)P\_\{train\}\(\\varphi\|\\phi\_\{t\}\)in \([21](https://arxiv.org/html/2607.08041#A2.E21)\) over the past training dataφ∈𝒟\\varphi\\in\\mathcal\{D\}given the current imageϕt\\phi\_\{t\}provides a full analytic solution to both the Bayes optimal denoiser in \([19](https://arxiv.org/html/2607.08041#A2.E19)\) and the reverse process in \([12](https://arxiv.org/html/2607.08041#A1.E12)\) through Tweedie’s theorem in \([13](https://arxiv.org/html/2607.08041#A1.E13)\)\.
As explained in\[[13](https://arxiv.org/html/2607.08041#bib.bib13)\], the structure of this analytic solution has an appealing interpretation in terms of a Bayesian guessing game about the past of the forward diffusion process given a noised imageϕt\\phi\_\{t\}\. This Bayesian guessing game is played by both the denoiser and the reverse process\. In essence, both the denoiser and the reverse process are trying to guess which training imageφ∈𝒟\\varphi\\in\\mathcal\{D\}at timet=0t=0led toϕt\\phi\_\{t\}under the forward diffusion\. The optimal Bayesian observer would summarize this guess in the Bayesian posterior distributionPtrain\(φ\|ϕt\)P\_\{train\}\(\\varphi\|\\phi\_\{t\}\)in \([21](https://arxiv.org/html/2607.08041#A2.E21)\)\. The optimal MMSE Bayesian denoiser then uses this posterior to compute the posterior mean training set image through \([19](https://arxiv.org/html/2607.08041#A2.E19)\)\. Similarly, the task of the reverse process is to flowϕt\\phi\_\{t\}backwards in time fromttto0and reverse the forward diffusion\. The Bayes\-optimal reverse process \(obtained by inserting \([21](https://arxiv.org/html/2607.08041#A2.E21)\) into \([19](https://arxiv.org/html/2607.08041#A2.E19)\), then inserting \([19](https://arxiv.org/html/2607.08041#A2.E19)\) into \([13](https://arxiv.org/html/2607.08041#A1.E13)\), and then finally inserting \([13](https://arxiv.org/html/2607.08041#A1.E13)\) into \([12](https://arxiv.org/html/2607.08041#A1.E12)\)\) uses the posterior to flowϕt\\phi\_\{t\}back to each data pointφ∈𝒟\\varphi\\in\\mathcal\{D\}with a Bayesian posterior weight controlled byPtrain\(φ\|ϕt\)P\_\{train\}\(\\varphi\|\\phi\_\{t\}\)\. In essence, the optimal Bayesian reverse process guesses the past training data origin ofϕt\\phi\_\{t\}under the forward diffusion and flows back towards this origin\.
While the Bayes optimal diffusion model is appealing in terms of its both its analytic and conceptual simplicity, its sampling behavior is unfortunately quite unrealistic and does not explain what trained neural network based diffusion models actually do\. The key issue is that for any finite training set, the Bayes optimal reverse process memorizes the training data, and can only flow to one of the training set pointsφ∈𝒟\\varphi\\in\\mathcal\{D\}\(see also\[[13](https://arxiv.org/html/2607.08041#bib.bib13)\]\)\. Intuitively, as time decreases fromt=1t=1tot=0t=0in the reverse process, the posterior distribution in \([21](https://arxiv.org/html/2607.08041#A2.E21)\) increases its probability on training pointsφ\\varphiclose toϕt\\phi\_\{t\}, and in turn,ϕt\\phi\_\{t\}then increasingly flows towards such training points\. The positive feedback between posterior beliefPtrain\(φ\|ϕt\)P\_\{train\}\(\\varphi\|\\phi\_\{t\}\)and reverse flow ofϕt\\phi\_\{t\}to imagesφ\\varphiof high belief, eventually causesϕt\\phi\_\{t\}to converge to asingletraining set pointφ∈𝒟\\varphi\\in\\mathcal\{D\}\.
In terms of the Bayesian guessing game interpretation, with small amounts of data and low enough noise levels, the guessing game istoo easy\. With only a small number of training data pointsφ∈𝒟\\varphi\\in\\mathcal\{D\}, which are likely to be well separated \(relative to the noise scaleσt\\sigma\_\{t\}\) in an exponentially large high dimensional space,ϕt\\phi\_\{t\}will generically be much closer to the data pointφ∗\\varphi^\{\*\}that generated it than any other data pointφ≠φ∗\\varphi\\neq\\varphi^\{\*\}\. In this situation, the posterior beliefPtrain\(φ\|ϕtP\_\{train\}\(\\varphi\|\\phi\_\{t\}\) will have a very high probability whenφ=φ∗\\varphi=\\varphi\*and a low probability otherwise\. Thus the posterior beliefPtrain\(φ\|ϕtP\_\{train\}\(\\varphi\|\\phi\_\{t\}\)correctlyguesses the origin ofϕt\\phi\_\{t\}in the forward process\.
Even more intuitively, in terms of images, if one observesthe entiretyof a noised imageϕt\\phi\_\{t\}at moderately low levels of noise, and one knows that the noised image came from a small set of widely separated clean training images, it is easy to guess which training image it came from\. Thus the ease of the Bayesian guessing game in high dimensions is intimately tied to the deleterious memorization behavior of the Bayes optimal diffusion model\. This memorization behavior hurts denoising on a test image; this denoiser will be overly influenced by a particular training set image\. Also, from the diffusion model perspective, this memorization behavior also impedes creative generation of new images different from any image in the training set\.
## Appendix CBayesian Information Restricted Diffusion \(BIRD\) Models
We have seen that optimal Bayesian diffusion models memorize finite training sets precisely because the Bayesian guessing game of which training set imageφ∈𝒟\\varphi\\in\\mathcal\{D\}leads to a given noise imageϕt\\phi\_\{t\}in the forward diffusion is too easy, especially for small numbers of training images in a high dimensional space at low levels of noise \(or equivalently last in the reverse process\)\. This raises the question of whether one can obtain alternate Bayesian diffusion models that retain analytic tractability and conceptual simplicity, butalsogeneralize appropriately, thereby behaving more like trained neural network diffusion models\.
We show that we cansimultaneouslyachieveboththeoretical tractabilityandmore realistic generalization\. The key idea is to prevent memorization by making the Bayesian guessing game harderwithoutincreasing the amount of training data\|𝒟\|\|\\mathcal\{D\}\|or increasing the noise level\. In particular we combine two key approaches: 1\) we instead make the Bayesian guessing game harder by restricting the image information the Bayesian guesser uses to guess the past of the forward diffusion process; and 2\) we embrace diversity by allowing different parts of the image, down to the level ofindividual pixels, to usedifferentpieces of restricted information to guess the past\. The former prevents memorization while the latter enhances creative generalization\. We call such models Bayesian Information Restricted Diffusion \(BIRD\) models\.
To introduce the general class of BIRD models, we first note that the denoising loss \([14](https://arxiv.org/html/2607.08041#A1.E14)\), as a simple squared loss, decouples across pixelsxx\. In particular, we can write it as a sum over pixelsxx\(or more generally, as a sum over any orthonormal basis over pixels\):
ℒt=∑xℒx,twhereℒx,t=12𝔼φ∼Ptrain,ϕt‖φx−Mx,t\[ϕt\]‖2\.\\displaystyle\\mathcal\{L\}\_\{t\}=\\sum\_\{x\}\\mathcal\{L\}\_\{x,t\}\\quad\\text\{where\}\\quad\\mathcal\{L\}\_\{x,t\}=\\frac\{1\}\{2\}\\mathbb\{E\}\_\{\\varphi\\sim P\_\{train\},\\phi\_\{t\}\}\\norm\{\\varphi\_\{x\}\-M\_\{x,t\}\[\\phi\_\{t\}\]\}^\{2\}\.\(22\)HereMx,t\[ϕt\]∈ℝCM\_\{x,t\}\[\\phi\_\{t\}\]\\in\\mathbb\{R\}^\{C\}is the evaluation of the model denoiser at timettatsinglepixel locationxx, and it is trained to recover the single clean training image pixel valueφx∈ℝC\\varphi\_\{x\}\\in\\mathbb\{R\}^\{C\}\. We can thus posit a separate denoiser model for each pixel and analyze them individually\. Without this decoupling, analyzing the behavior of diffusion models with inductive biases becomes significantly less tractable; see e\.g\.\[[34](https://arxiv.org/html/2607.08041#bib.bib34)\]for an example of the complexities that arise when couplings between pixels are introduced\.
To achieve the restricted information, we ensure that the denoiser functionMx,tM\_\{x,t\}is not allowed to observe the entire noised imageϕt\\phi\_\{t\}\. Instead it is only allowed to make a restricted, possibly stochastic, observation𝒞x,t∈ℝnx,t\\mathcal\{C\}\_\{x,t\}\\in\\mathbb\{R\}^\{n\_\{x,t\}\}ofϕt∈ℝd\\phi\_\{t\}\\in\\mathbb\{R\}^\{d\}\. We usually assume the observation dimensionalitynx,tn\_\{x,t\}is less than the full image dimensionalitydd\. If the observation is stochastic we denote its conditional distribution byP\(𝒞x,t\|ϕt\)P\(\\mathcal\{C\}\_\{x,t\}\|\\phi\_\{t\}\)\. If it is deterministic we denote it by the function𝒞x,t\(ϕt\)\\mathcal\{C\}\_\{x,t\}\(\\phi\_\{t\}\)\. For brevity we also sometimes drop the argumentϕt\\phi\_\{t\}, and then𝒞x,t∈ℝnx,t\\mathcal\{C\}\_\{x,t\}\\in\\mathbb\{R\}^\{n\_\{x,t\}\}denotes the pixelxx’s observation value\. Furthermore we assume the that different pixelsxxcan makedifferentobservations𝒞x,t\\mathcal\{C\}\_\{x,t\}about the noised imageϕt\\phi\_\{t\}\. This diversifies the state of knowledge of each pixel\. The BIRD model is then simply the collection of individual pixel denoisersMx,t\[𝒞x,t\]M\_\{x,t\}\[\\mathcal\{C\}\_\{x,t\}\]that each achieve the MMSE loss forℒx,t\\mathcal\{L\}\_\{x,t\}in \([22](https://arxiv.org/html/2607.08041#A3.E22)\)\.
Just like the full optimal Bayesian model, the MMSE BIRD model also has an analytic solution for pixel x given by
Mx,t\[𝒞x,t\]=∑φ∈𝒟φxPtrain\(φ\|𝒞x,t\),\\displaystyle M\_\{x,t\}\[\\mathcal\{C\}\_\{x,t\}\]=\\sum\_\{\\varphi\\in\\mathcal\{D\}\}\\varphi\_\{x\}\\,P\_\{train\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\),\(23\)where the Bayesian posterior is given by
Ptrain\(φ\|𝒞x,t\)=P\(𝒞x,t\|φ\)∑φ′P\(𝒞x,t\|φ′\)\.\\displaystyle P\_\{train\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)=\\frac\{P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\)\}\{\\sum\_\{\\varphi^\{\\prime\}\}P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi^\{\\prime\}\)\}\.\(24\)We can then combine these pixel\-wise estimates into the BIRD model for the full image
Mt\[φ\]=∑xMx,t\[𝒞x,t\]x^M\_\{t\}\[\\varphi\]=\\sum\_\{x\}M\_\{x,t\}\[\\mathcal\{C\}\_\{x,t\}\]\\hat\{x\}\(25\)To understand this solution, consider the training Markov processφ→ϕt→𝒞x,t\\varphi\\rightarrow\\phi\_\{t\}\\rightarrow\\mathcal\{C\}\_\{x,t\}for pixelxxat timett, denoted by the joint distributionPtrain\(φ,ϕt,𝒞x,t\)=Ptrain\(φ\)P\(ϕt\|φ\)P\(𝒞x,t\|ϕt\)P\_\{train\}\(\\varphi,\\phi\_\{t\},\\mathcal\{C\}\_\{x,t\}\)=P\_\{train\}\(\\varphi\)P\(\\phi\_\{t\}\|\\varphi\)P\(\\mathcal\{C\}\_\{x,t\}\|\\phi\_\{t\}\)\. Hereφ∼Ptrain\(φ\)\\varphi\\sim P\_\{train\}\(\\varphi\)is a randomly chosen training imageφ∈𝒟\\varphi\\in\\mathcal\{D\},P\(ϕt\|φ\)P\(\\phi\_\{t\}\|\\varphi\)is the conditional distribution of a noisy imageϕt\\phi\_\{t\}under the forward diffusion process, starting from the training imageφ\\varphi, andP\(𝒞x,t\|ϕt\)P\(\\mathcal\{C\}\_\{x,t\}\|\\phi\_\{t\}\)is the \(possibly stochastic\) observation that pixelxxmakes about imageϕt\\phi\_\{t\}\. We can think about the training Markov processφ→ϕt→𝒞x,t\\varphi\\rightarrow\\phi\_\{t\}\\rightarrow\\mathcal\{C\}\_\{x,t\}as an information channel from the past training data𝒟\\mathcal\{D\}at timet=0t=0to the current pixel observation𝒞x,t\\mathcal\{C\}\_\{x,t\}at timett\.
Ptrain\(φ\|𝒞x,t\)P\_\{train\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)in \([24](https://arxiv.org/html/2607.08041#A3.E24)\) is then the Bayesian posterior distribution over the training dataφ∈𝒟\\varphi\\in\\mathcal\{D\}given a pixelxx’s channel observation𝒞x,t\\mathcal\{C\}\_\{x,t\}\. This posterior assigns a probability to each of the\|𝒟\|\|\\mathcal\{D\}\|training imagesφ∈𝒟\\varphi\\in\\mathcal\{D\}\. If a pixelxxhad to guess, based on its \(restricted\) observation𝒞x,t\\mathcal\{C\}\_\{x,t\}of the noisy imageϕt\\phi\_\{t\}, which clean training imageφ∈𝒟\\varphi\\in\\mathcal\{D\}at time0in the past led toϕt\\phi\_\{t\}under the forward diffusion, then any optimal guess would be based on the posteriorPtrain\(φ\|𝒞x,t\)P\_\{train\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\), which encapsulates the pixel’s belief state about the past training data origin of its current observation\. The individual Bayes optimal pixel denoiser at pixelxx, based on its restricted observation𝒞x,t\\mathcal\{C\}\_\{x,t\}, then simply returns in \([23](https://arxiv.org/html/2607.08041#A3.E23)\) the weighted average of pixel valuesφx\\varphi\_\{x\}of pixelxxfor all imagesφ\\varphiin the training dataset𝒟\\mathcal\{D\}, where the weights are determined by the posteriorPtrain\(φ\|𝒞x,t\)P\_\{train\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\.
The channel restriction𝒞x,t\\mathcal\{C\}\_\{x,t\}at the end of the forward information channelφ→ϕt→𝒞x,t\\varphi\\rightarrow\\phi\_\{t\}\\rightarrow\\mathcal\{C\}\_\{x,t\}makes the backwards Bayesian guessing game harder, and ideally prevents the posteriorP\(φ\|𝒞x,t\)P\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)from memorizing by concentrating on a single correct training exampleφ∈𝒟\\varphi\\in\\mathcal\{D\}\. This behavior would be in contrast to the more informative posteriorPtrain\(φ\|ϕt\)P\_\{train\}\(\\varphi\|\\phi\_\{t\}\)that always forces the unrestricted Bayes optimal diffusion model to memorize at low enough levels of noise \(or equivalently late in the reverse process\), as discussed in App\.[B](https://arxiv.org/html/2607.08041#A2)\.
### C\.1Examples of BIRD models: LS, ES, and ELS
A natural and simple channel restriction𝒞x,t\\mathcal\{C\}\_\{x,t\}is a deterministic low dimensional projection applied toϕt\\phi\_\{t\}\. In the case of a linear projection,𝒞x,t\(ϕt\)=𝒫x,tϕt\\mathcal\{C\}\_\{x,t\}\(\\phi\_\{t\}\)=\\mathcal\{P\}\_\{x,t\}\\phi\_\{t\}where𝒫x,t\\mathcal\{P\}\_\{x,t\}is annx,tn\_\{x,t\}byddmatrix with orthonormal rows\. Under this restriction, the Bayesian posterior is given by
P\(φ\|𝒫x,tϕt\)=exp\(−12\(1−α¯t\)‖𝒫x,t\(ϕt−α¯tφ‖S2\)\)∑φ′exp\(−12\(1−α¯t\)‖𝒫x,t\(ϕt−α¯tφ′‖S2\)\)\.\\displaystyle P\(\\varphi\|\\mathcal\{P\}\_\{x,t\}\\phi\_\{t\}\)=\\frac\{\\exp\(\-\\frac\{1\}\{2\(1\-\\bar\{\\alpha\}\_\{t\}\)\}\\norm\{\\mathcal\{P\}\_\{x,t\}\(\\phi\_\{t\}\-\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\varphi\}\_\{S\}^\{2\}\)\)\}\{\\sum\_\{\\varphi^\{\\prime\}\}\\exp\(\-\\frac\{1\}\{2\(1\-\\bar\{\\alpha\}\_\{t\}\)\}\\norm\{\\mathcal\{P\}\_\{x,t\}\(\\phi\_\{t\}\-\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\varphi^\{\\prime\}\}\_\{S\}^\{2\}\)\)\}\.\(26\)The Local Score \(LS\) machine of\[[13](https://arxiv.org/html/2607.08041#bib.bib13)\]is an instance of this framework when the projection operator𝒫x,t\\mathcal\{P\}\_\{x,t\}projects onto a coordinate subspace of pixel components within a local patchΩx\\Omega\_\{x\}surrounding pixelxx:
P\(φ\|ϕt;Ωx\)=exp\(−12\(1−α¯t\)‖ϕt,Ωx−α¯tφΩx‖2\)∑φ′exp\(−12\(1−α¯t\)‖ϕt,Ωx−α¯tφΩx‖2\)\.\\displaystyle P\(\\varphi\|\\phi\_\{t;\\Omega\_\{x\}\}\)=\\frac\{\\exp\(\-\\frac\{1\}\{2\(1\-\\bar\{\\alpha\}\_\{t\}\)\}\\norm\{\\phi\_\{t,\\Omega\_\{x\}\}\-\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\varphi\_\{\\Omega\_\{x\}\}\}^\{2\}\)\}\{\\sum\_\{\\varphi^\{\\prime\}\}\\exp\(\-\\frac\{1\}\{2\(1\-\\bar\{\\alpha\}\_\{t\}\)\}\\norm\{\\phi\_\{t,\\Omega\_\{x\}\}\-\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\varphi\_\{\\Omega\_\{x\}\}\}^\{2\}\)\}\.\(27\)Here we follow the notation of\[[13](https://arxiv.org/html/2607.08041#bib.bib13)\], whereϕΩx\\phi\_\{\\Omega\_\{x\}\}refers to a vector of concatenated pixel values of the imageϕ\\phi, but restricted only to all pixels inside the local patch of pixelsΩx\\Omega\_\{x\}\. Projections onto subspaces in more general bases may be admitted, as hypothesized by\[[32](https://arxiv.org/html/2607.08041#bib.bib32)\], although the latter paper did not propose any specific basis other than the pixel basis used by\[[13](https://arxiv.org/html/2607.08041#bib.bib13)\]\.
While this framework is already fairly general, there is a further generalization of it that can be considered, which is to regress onto a correlated targetTxT\_\{x\}rather than ontoφ\\varphiitself:
ℒ=∑x12⟨‖Tx−Mx\[𝒞x,t\]‖2⟩φ,Tx,𝒞x,t,\\displaystyle\\mathcal\{L\}=\\sum\_\{x\}\\frac\{1\}\{2\}\\big\\langle\\norm\{T\_\{x\}\-M\_\{x\}\[\\mathcal\{C\}\_\{x,t\}\]\}^\{2\}\\big\\rangle\_\{\\varphi,T\_\{x\},\\mathcal\{C\}\_\{x,t\}\},\(28\)for which the minimizer is given by
Mx\[𝒞x,t\]=∑φ∫𝑑TxTxP\(Tx\|φ,𝒞x,t\)P\(φ\|𝒞x,t\)\\displaystyle M\_\{x\}\[\\mathcal\{C\}\_\{x,t\}\]=\\sum\_\{\\varphi\}\\int\\,dT\_\{x\}\\,T\_\{x\}P\(T\_\{x\}\|\\varphi,\\mathcal\{C\}\_\{x,t\}\)P\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\(29\)This generalization may be useful for various stochastic flow\-matching setups\. It also allows us to reproduce the Equivariant Score \(ES\) machine and Equivariant Local Score \(ELS\) machine of\[[13](https://arxiv.org/html/2607.08041#bib.bib13)\]under a particular correlated choice of targetTxT\_\{x\}and channel𝒞x,t\\mathcal\{C\}\_\{x,t\}\.\[[13](https://arxiv.org/html/2607.08041#bib.bib13)\]consider the setting of a model that isequivariantunder a groupGG, which means that for anyg∈Gg\\in Gthe modelM\[ϕ\]M\[\\phi\]satisfiesM\[Ugϕ\]=UgM\[ϕ\]M\[U\_\{g\}\\phi\]=U\_\{g\}M\[\\phi\], whereUgU\_\{g\}is a unitary representation of the groupGG\. Under the following choice of correlated channel and target:
g\\displaystyle g∼Haar\(G\)\\displaystyle\\sim\\text\{Haar\}\(G\)\(30\)𝒞x,t\\displaystyle\\mathcal\{C\}\_\{x,t\}=Ug†ϕt\\displaystyle=U\_\{g\}^\{\\dagger\}\\phi\_\{t\}\(31\)Tx\\displaystyle T\_\{x\}=Ugφ\\displaystyle=U\_\{g\}\\varphi\(32\)we reproduce the ES machine with the equation \([29](https://arxiv.org/html/2607.08041#A3.E29)\):
Mx\[ϕt\]\\displaystyle M\_\{x\}\[\\phi\_\{t\}\]=∑φ∈𝒟∑g∈GUgφP\(φ,g\|Ug†ϕt\)\\displaystyle=\\sum\_\{\\varphi\\in\\mathcal\{D\}\}\\sum\_\{g\\in G\}U\_\{g\}\\varphi\\,P\(\\varphi,g\|U\_\{g\}^\{\\dagger\}\\phi\_\{t\}\)\(33\)P\(φ,g\|Ugϕt\)\\displaystyle P\(\\varphi,g\|U\_\{g\}\\phi\_\{t\}\)=exp\(−12\(1−α¯t\)‖φ−Ug†ϕt‖2\)∑φ′∈𝒟∑g′∈Gexp\(−12\(1−α¯t\)‖φ′−Ug′†ϕt‖2\)\\displaystyle=\\frac\{\\exp\(\-\\frac\{1\}\{2\(1\-\\bar\{\\alpha\}\_\{t\}\)\}\\norm\{\\varphi\-U\_\{g\}^\{\\dagger\}\\phi\_\{t\}\}^\{2\}\)\}\{\\sum\_\{\\varphi^\{\\prime\}\\in\\mathcal\{D\}\}\\sum\_\{g^\{\\prime\}\\in G\}\\exp\(\-\\frac\{1\}\{2\(1\-\\bar\{\\alpha\}\_\{t\}\)\}\\norm\{\\varphi^\{\\prime\}\-U\_\{g^\{\\prime\}\}^\{\\dagger\}\\phi\_\{t\}\}^\{2\}\)\}\(34\)The ELS machine can be reproduced by specializingggto the group of translations𝒯\\mathcal\{T\}on images withUgU\_\{g\}the fundamental representations, and then subsequently projecting the outputUg†ϕtU\_\{g\}^\{\\dagger\}\\phi\_\{t\}into a patchΩx\\Omega\_\{x\}:
Mx\[ϕt\]\\displaystyle M\_\{x\}\[\\phi\_\{t\}\]=∑φ∈𝒟∑g∈𝒯\(Ugφ\)xP\(φ,g\|\(Ug†ϕt\)Ωx\)\\displaystyle=\\sum\_\{\\varphi\\in\\mathcal\{D\}\}\\sum\_\{g\\in\\mathcal\{T\}\}\(U\_\{g\}\\varphi\)\_\{x\}\\,P\(\\varphi,g\|\(U\_\{g\}^\{\\dagger\}\\phi\_\{t\}\)\_\{\\Omega\_\{x\}\}\)\(35\)P\(φ,g\|Ugϕt\)\\displaystyle P\(\\varphi,g\|U\_\{g\}\\phi\_\{t\}\)=exp\(−12\(1−α¯t\)‖φΩx−\(Ug†ϕt\)Ωx‖2\)∑φ′∈𝒟∑g′∈Gexp\(−12\(1−α¯t\)‖φΩx′−\(Ug′†ϕt\)Ωx‖2\)\\displaystyle=\\frac\{\\exp\(\-\\frac\{1\}\{2\(1\-\\bar\{\\alpha\}\_\{t\}\)\}\\norm\{\\varphi\_\{\\Omega\_\{x\}\}\-\(U\_\{g\}^\{\\dagger\}\\phi\_\{t\}\)\_\{\\Omega\_\{x\}\}\}^\{2\}\)\}\{\\sum\_\{\\varphi^\{\\prime\}\\in\\mathcal\{D\}\}\\sum\_\{g^\{\\prime\}\\in G\}\\exp\(\-\\frac\{1\}\{2\(1\-\\bar\{\\alpha\}\_\{t\}\)\}\\norm\{\\varphi^\{\\prime\}\_\{\\Omega\_\{x\}\}\-\(U\_\{g^\{\\prime\}\}^\{\\dagger\}\\phi\_\{t\}\)\_\{\\Omega\_\{x\}\}\}^\{2\}\)\}\(36\)
## Appendix DA theory of BIRD memorization\-generalization phase transitions
In App\.[B](https://arxiv.org/html/2607.08041#A2)we noted that unrestricted Bayes optimal diffusion models always memorize at low enough levels of noise \(corresponding to late enough times in the reverse process\)\. Motivated by this issue we introduced a very general class of BIRD models in App\.[C](https://arxiv.org/html/2607.08041#A3)\. In this section we now wish to develop a general theory of when such BIRD models memorize versus generalize\.
### D\.1A high level overview of memorization versus generalization
Memorization versus generalization can be assessed by the behavior of the posterior distributionPtrain\(φ\|𝒞x,t\)P\_\{train\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)in \([24](https://arxiv.org/html/2607.08041#A3.E24)\) which encapsulates each pixelxx’s belief state about which past training imageφ∈𝒟\\varphi\\in\\mathcal\{D\}at timet=0t=0led to the pixel’s current observationCx,tC\_\{x,t\}at the current timettunder the training Markov process for forward diffusion and observationφ→ϕt→Cx,t\\varphi\\rightarrow\\phi\_\{t\}\\rightarrow C\_\{x,t\}, specified by the joint distributionPtrain\(φ,ϕt,𝒞x,t\)=Ptrain\(φ\)P\(ϕt\|φ\)P\(𝒞x,t\|ϕt\)P\_\{train\}\(\\varphi,\\phi\_\{t\},\\mathcal\{C\}\_\{x,t\}\)=P\_\{train\}\(\\varphi\)P\(\\phi\_\{t\}\|\\varphi\)P\(\\mathcal\{C\}\_\{x,t\}\|\\phi\_\{t\}\)\. In the memorization phase, the posteriorPtrain\(φ\|𝒞x,t\)P\_\{train\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)would concentrate on a single, or very small number of training imagesφ∈𝒟\\varphi\\in\\mathcal\{D\}\.
An important consequence of the concentration of the posteriorPtrain\(φ\|𝒞x,t\)P\_\{train\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)is that the accompanying BIRD denoiser and diffusion model then both become extremely sensitive to the particular realization of the training data𝒟\\mathcal\{D\}\. For example, consider twodistinctdata sets𝒟\\mathcal\{D\}and𝒟′\\mathcal\{D\}^\{\\prime\}each drawn i\.i\.d\. from the same true data distributionP0\(φ\)P\_\{0\}\(\\varphi\)\. Because the posteriorPtrain\(φ\|𝒞x,t\)P\_\{train\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)concentrates on individual \(or a small number of\) data points in the memorizing phase, the BIRD model denoiser and reverse process would yielddifferentoutputs even for thesameobservation inputs𝒞x,t\\mathcal\{C\}\_\{x,t\}when using the two different training datasets𝒟\\mathcal\{D\}and𝒟′\\mathcal\{D\}^\{\\prime\}\. Thus the BIRD model would not berobustto the choice of the training data𝒟\\mathcal\{D\}\.
In contrast, the hallmark of the generalization phase is that a BIRD modelshould be robustto the choice of training data𝒟\\mathcal\{D\}\. In particular, averages of quantities with respect to the posteriorPtrain\(φ\|𝒞x,t\)P\_\{train\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\), should not depend on the detailed realization of the training data𝒟\\mathcal\{D\}\. An important such average is the posterior mean computed by the denoiser outputMx,t\[𝒞x,t\]M\_\{x,t\}\[\\mathcal\{C\}\_\{x,t\}\]in \([23](https://arxiv.org/html/2607.08041#A3.E23)\)\. In this robust regime, two different denoisers, and two different BIRD model reverse samplers, obtained from two different datasets𝒟\\mathcal\{D\}and𝒟′\\mathcal\{D\}^\{\\prime\},consistently generalizeto the same denoiser and sampler\.
How do we describe this common denoiser in the robust, consistent generalization phase? Well, a simple way to achieve this phase is to take the limit of a large number of data points\|𝒟\|\|\\mathcal\{D\}\|\. Then averages of quantities with respect toPtrain\(φ\)P\_\{train\}\(\\varphi\)in \([18](https://arxiv.org/html/2607.08041#A2.E18)\) should converge to averages with respect to the true distributionP0\(φ\)P\_\{0\}\(\\varphi\)from which a new test exampleφ\\varphiwould be drawn\. Thus we can consider thetestingMarkov processφ→ϕt→𝒞x,t\\varphi\\rightarrow\\phi\_\{t\}\\rightarrow\\mathcal\{C\}\_\{x,t\}for pixelxxat timett, defined by the joint distributionPtest\(φ,ϕt,𝒞x,t\)=P0\(φ\)P\(ϕt\|φ\)P\(𝒞x,t\|ϕt\)\.P\_\{test\}\(\\varphi,\\phi\_\{t\},\\mathcal\{C\}\_\{x,t\}\)=P\_\{0\}\(\\varphi\)P\(\\phi\_\{t\}\|\\varphi\)P\(\\mathcal\{C\}\_\{x,t\}\|\\phi\_\{t\}\)\.This is identical to the training Markov process above, except the prior is now thetruedata distributionP0\(φ\)P\_\{0\}\(\\varphi\)from which a test sample is drawn, rather than the empirical training distributionPtrain\(φ\)P\_\{train\}\(\\varphi\)for a specific dataset𝒟\\mathcal\{D\}\. We denote the posterior distribution of the inputφ\\varphigiven the output𝒞x,t\\mathcal\{C\}\_\{x,t\}under this Markov process byPtest\(φ\|𝒞x,t\)P\_\{test\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\. This posterior distribution is the outcome of the Bayesian guessing game in the limit of a large amount of data\.
With the testing Markov process in place, we can see that in the robust, consistent generalization phase, the denoiser output in \([23](https://arxiv.org/html/2607.08041#A3.E23)\) converges to a common denoiser for any dataset𝒟\\mathcal\{D\}, where the dataset dependent posteriorPtrain\(φ\|𝒞x,t\)P\_\{train\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)is replaced with the dataset indpendent posteriorPtest\(φ\|𝒞x,t\)P\_\{test\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\. This common denoiser then determines a common reverse sampler as described in App\.[B](https://arxiv.org/html/2607.08041#A2)\. In this fashion, in the generalization phase, when two different BIRD models are trained on two different datasets, they both consistently converge to this same sampler\.
### D\.2The entropy of Bayesian guessing determines memorization and generalization
We would like to derive sharper conditions on the minimum amount of training data\|𝒟\|\|\\mathcal\{D\}\|required to achieve generalization instead of memorization, and how such a data amount threshold depends on timettin the reverse process \(or equivalently the amount of noise\), as well as the amount of information the restricted observation𝒞x,t\\mathcal\{C\}\_\{x,t\}provides about the training dataφ∈𝒟\\varphi\\in\\mathcal\{D\}\.
To do so, it is useful to focus on the entropy of the outcome of the Bayesian guessing game, encapsulated by posterior distributionPtrain\(φ\|𝒞x,t\)P\_\{train\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\. This entropy is given by
S\[Ptrain\(φ\|𝒞x,t\)\]=−∑φ∈𝒟Ptrain\(φ\|𝒞x,t\)lnPtrain\(φ\|𝒞x,t\)\.\\displaystyle S\[P\_\{train\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\]=\-\\sum\_\{\\varphi\\in\\mathcal\{D\}\}P\_\{train\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\\ln P\_\{train\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\.\(37\)This entropy can be written as the prior entropyS\[Ptrain\(φ\)\]=ln\|𝒟\|S\[P\_\{train\}\(\\varphi\)\]=\\ln\|\\mathcal\{D\}\|, minus the amount of information learned about the past training dataφ∈D\\varphi\\in Drelative to the prior by the current measurement𝒞x,t\\mathcal\{C\}\_\{x,t\}, also known as theBayesian surprise:
S\[Ptrain\(φ\|𝒞x,t\)\]=ln\|𝒟\|−DKL\(Ptrain\(φ\|𝒞x,t\)\|\|Ptrain\(φ\)\)\.S\[P\_\{train\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\]=\\ln\|\\mathcal\{D\}\|\-D\_\{KL\}\(P\_\{train\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\\,\\,\|\|\\,\\,P\_\{train\}\(\\varphi\)\)\.\(38\)In the memorization phase,Ptrain\(φ\|𝒞x,t\)P\_\{train\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)concentrates on a single point and so the entropy vanishes\. In a generalizing phase, the posterior will have support on many samples, and so we may expect the Bayesian surpriseDKL\(Ptrain\(φ\|𝒞x,t\)\|\|Ptrain\(φ\)\)D\_\{KL\}\(P\_\{train\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\\,\\,\|\|\\,\\,P\_\{train\}\(\\varphi\)\)to converge to its value under the true data distribution, i\.e\.DKL\(Ptest\(φ\|𝒞x,t\)\|\|P0\(φ\)\)D\_\{KL\}\(P\_\{test\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\\,\\,\|\|\\,\\,P\_\{0\}\(\\varphi\)\)\.
As we will see below, when both the logarithm of the amount of dataln\|𝒟\|\\ln\|\\mathcal\{D\}\|and the data dimensionddare large, there is a sharp phase transition between the two phases\. We find that wheneverln\|𝒟\|−DKL\(Ptest\(φ\|𝒞x,t\)\|\|P0\(φ\)\)≥0\\ln\|\\mathcal\{D\}\|\-D\_\{KL\}\(P\_\{test\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\\,\\,\|\|\\,\\,P\_\{0\}\(\\varphi\)\)\\geq 0, the training surprise matches the test surprise\. This corresponds to the generalization phase\. Conversely, when this inequality does not hold, the posterior entropy is zero\. Thus we obtain the following formula in the largeln\|𝒟\|\\ln\|\\mathcal\{D\}\|limit \(see Theorem[D\.1](https://arxiv.org/html/2607.08041#A4.Thmdefinition1)in App\.[D\.5](https://arxiv.org/html/2607.08041#A4.SS5)for a formal theorem statement and proof\):
S\[Ptrain\(φ\|𝒞x,t\)\]=\{ln\|𝒟\|−DKL\(Ptest\(φ\|𝒞x,t\)\|\|P0\(φ\)\)ln\|𝒟\|\>DKL\(Ptest\(φ\|𝒞x,t\)\|\|P0\(φ\)\)0ln\|𝒟\|≤DKL\(Ptest\(φ\|𝒞x,t\)\|\|P0\(φ\)\)S\[P\_\{train\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\]=\\begin\{cases\}\\ln\|\\mathcal\{D\}\|\-D\_\{KL\}\(P\_\{test\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\|\|P\_\{0\}\(\\varphi\)\)&\\ln\|\\mathcal\{D\}\|\>D\_\{KL\}\(P\_\{test\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\|\|P\_\{0\}\(\\varphi\)\)\\\\ 0&\\ln\|\\mathcal\{D\}\|\\leq D\_\{KL\}\(P\_\{test\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\|\|P\_\{0\}\(\\varphi\)\)\\end\{cases\}\(39\)This result yields a pointwise lower threshold for the minimum amount of data\|𝒟\|\|\\mathcal\{D\}\|required to avoid memorization and achieve generalization for any fixed observation𝒞x,t\\mathcal\{C\}\_\{x,t\}\. Ifln\|𝒟\|\\ln\|\\mathcal\{D\}\|is less than the Bayesian test surpriseDKL\(Ptest\(φ\|𝒞x,t\)\|\|P0\(φ\)\)D\_\{KL\}\(P\_\{test\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\|\|P\_\{0\}\(\\varphi\)\), then the BIRD model is in the memorization regime given that observation\.
To understand this result intuitively, note that the Bayesian surprise quantifies how much a pixelxxneeds to update its prior belief distribution about the training dataP0\(ϕ\)P\_\{0\}\(\\phi\)to obtain the posterior belief distributionPtest\(φ\|𝒞x,t\)P\_\{test\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)after making the observationCx,tC\_\{x,t\}\. If this surprise is high, then intuitively𝒞x,t\\mathcal\{C\}\_\{x,t\}is highly informative aboutφ\\varphi\. Thus the higher the surprise, the more data\|𝒟\|\|\\mathcal\{D\}\|is required to avoid memorization for any fixed observation𝒞x,t\\mathcal\{C\}\_\{x,t\}\. Indeed we obtain the extremely simple condition that the amount of data\|𝒟\|\|\\mathcal\{D\}\|needs to be at least exponential in the surprise to achieve generalization\. Conversely, whenln\|𝒟\|\\ln\|\\mathcal\{D\}\|is greater than the surprise, the posterior entropy is nonnegative, indicating the generalization phase\. And quantitatively it is equal to the prior entropy minus the Bayesian surprise, yielding the remaining entropy after the observation\.
While the above result holds pointwise for asingleobservation𝒞x,t\\mathcal\{C\}\_\{x,t\}, we can also find a simple averaged condition when𝒞x,t\\mathcal\{C\}\_\{x,t\}is drawnrandomlyfrom the test distribution via the testing Markov processφ→ϕt→𝒞x,t\\varphi\\rightarrow\\phi\_\{t\}\\rightarrow\\mathcal\{C\}\_\{x,t\}whereφ∼P0\(ϕ\)\\varphi\\sim P\_\{0\}\(\\phi\)\. As is well known, the average of the Bayesian surpriseDKL\(Ptest\(φ\|𝒞x,t\)\|\|P0\(φ\)\)D\_\{KL\}\(P\_\{test\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\|\|P\_\{0\}\(\\varphi\)\)over the marginal distributionPtest\(𝒞x,t\)P\_\{test\}\(\\mathcal\{C\}\_\{x,t\}\)yields the mutual informationI\(φ;𝒞x,t\)I\(\\varphi;\\mathcal\{C\}\_\{x,t\}\)\. Thus if the average surprise is high, so is the mutual information \(as they are equal\)\. This yields an extremely simple and highly general information theoretic criterion for the memorization\-generalization phase transition boundary in any BIRD model and a wide variety of data distributions:
\{ln\|𝒟\|\>I\(φ;𝒞x,t\)Generalization phase\.ln\|𝒟\|<I\(φ;𝒞x,t\)Memorization phase\.\\begin\{cases\}\\ln\|\\mathcal\{D\}\|\>I\(\\varphi;\\mathcal\{C\}\_\{x,t\}\)\\quad\\text\{Generalization phase\.\}\\\\ \\ln\|\\mathcal\{D\}\|<I\(\\varphi;\\mathcal\{C\}\_\{x,t\}\)\\quad\\text\{Memorization phase\.\}\\end\{cases\}\(40\)This result then implies that restricting information contained in observations can reduce the minimum data threshold required for generalization in BIRD models\.
In summary, \([40](https://arxiv.org/html/2607.08041#A4.E40)\) reveals that the amount of data\|𝒟\|\|\\mathcal\{D\}\|needs to be at least exponential in the mutual information between current observationsCx,tC\_\{x,t\}and the true test dataφ∼P0\(φ\)\\varphi\\sim P\_\{0\}\(\\varphi\)in the forward testing Markov processφ→ϕt→𝒞x,t\\varphi\\rightarrow\\phi\_\{t\}\\rightarrow\\mathcal\{C\}\_\{x,t\}, to avoid memorization and achieve generalization\.
More formal statements of these results and their proofs are given below\.
### D\.3Three types of phase transitions between memorization and generalization\.
One can think of the phase transition boundary between the memorization phase and generalization phase, determined by the equality conditionI\(𝒞t,x;φ\)=ln\|𝒟\|I\(\\mathcal\{C\}\_\{t,x\};\\varphi\)=\\ln\|\\mathcal\{D\}\|in several ways:
#### A data phase transition\.
First, in terms of the amount of data\|𝒟\|\|\\mathcal\{D\}\|for a fixed observation method𝒞x,t\\mathcal\{C\}\_\{x,t\}and fixed timett, asln\|𝒟\|\\ln\|\\mathcal\{D\}\|increases from belowI\(𝒞t,x;φ\)I\(\\mathcal\{C\}\_\{t,x\};\\varphi\)to aboveI\(𝒞t,x;φ\)I\(\\mathcal\{C\}\_\{t,x\};\\varphi\), we transition from the memorization to generalization phase\. The critical amount of dataDcD\_\{c\}at which this transition occurs obeys the equationI\(𝒞t,x;φ\)=lnDcI\(\\mathcal\{C\}\_\{t,x\};\\varphi\)=\\ln D\_\{c\}\. For\|𝒟\|<Dc\|\\mathcal\{D\}\|<D\_\{c\}any BIRD model memorizes while for\|𝒟\|\>Dc\|\\mathcal\{D\}\|\>D\_\{c\}it generalizes\.
#### A temporal phase transition\.
Second, we can consider timettin the reverse process at a fixed amount of data\|𝒟\|\|\\mathcal\{D\}\|\(and an observation method that does not depend explicitly on time\)\. At timesttnear11near the beginning of the reverse process, the noise to signal ratioσt2\\sigma\_\{t\}^\{2\}is large and therefore the mutual informationI\(𝒞t,x;φ\)I\(\\mathcal\{C\}\_\{t,x\};\\varphi\)is small\. However at timesttnear0near the end of the reverse process, the noise to signal ratioσt2\\sigma\_\{t\}^\{2\}is small and therefore the mutual informationI\(𝒞t,x;φ\)I\(\\mathcal\{C\}\_\{t,x\};\\varphi\)is large\. Therefore, as the reverse process proceeds fromt=1t=1down tot=0t=0, the mutual informationI\(𝒞t,x;φ\)I\(\\mathcal\{C\}\_\{t,x\};\\varphi\)could in principle transition from belowln\|𝒟\|\\ln\|\\mathcal\{D\}\|to aboveln\|𝒟\|\\ln\|\\mathcal\{D\}\|, implying the reverse process could transition from a generalizing phase to a memorizing phase\. The temporal boundarytbt\_\{b\}at which this transition occurs obeys the equationI\(𝒞tb,x;φ\)=ln\|𝒟\|I\(\\mathcal\{C\}\_\{t\_\{b\},x\};\\varphi\)=\\ln\|\\mathcal\{D\}\|\. Fort\>tbt\>t\_\{b\}any BIRD model \(with a fixed, time\-independent observation method\) generalizes early in the reverse process, while fort<tbt<t\_\{b\}it memorizes late in the reverse process\.
#### An observation phase transition\.
Third, we can consider a one parameter family of observations𝒞x,t\\mathcal\{C\}\_\{x,t\}of varying information content, at a fixed timettand a fixed amount of data\|𝒟\|\|\\mathcal\{D\}\|\. For example, letnx,tn\_\{x,t\}be the dimensionality of the observation𝒞x,t\\mathcal\{C\}\_\{x,t\}such that reducingnx,tn\_\{x,t\}monotonically drops measurements ofϕt\\phi\_\{t\}\. As a concrete example, let𝒞x,t\(ϕ\)=ϕΩx\\mathcal\{C\}\_\{x,t\}\(\\phi\)=\\phi\_\{\\Omega\_\{x\}\}be the restriction of the imageϕ\\phito a localLLbyLLimage patchΩx\\Omega\_\{x\}centered atxx, yielding the restricted imageϕΩx\\phi\_\{\\Omega\_\{x\}\}of dimensionnx,t=L×L×Cn\_\{x,t\}=L\\times L\\times C\. In this example, reducingnx,tn\_\{x,t\}by reducing the patch sizeLLstrictly reduces the information content of the observation𝒞x,t\\mathcal\{C\}\_\{x,t\}\. Now whennx,tn\_\{x,t\}varies from large to small, the mutual informationI\(𝒞t,x;φ\)I\(\\mathcal\{C\}\_\{t,x\};\\varphi\)can also vary from large to small\. Thus asnx,tn\_\{x,t\}varies from large to small, the mutual informationI\(𝒞t,x;φ\)I\(\\mathcal\{C\}\_\{t,x\};\\varphi\)can in principle transition from aboveln\|𝒟\|\\ln\|\\mathcal\{D\}\|to belowln\|𝒟\|\\ln\|\\mathcal\{D\}\|, implying a transition from memorizing to generalizing as the dimensionality or information content of the observation𝒞x,t\\mathcal\{C\}\_\{x,t\}is reduced\. The critical dimensionncn\_\{c\}at which this transition occurs obeys the equationI\(𝒞t,xnc;φ\)=ln\|𝒟\|I\(\\mathcal\{C\}^\{n\_\{c\}\}\_\{t,x\};\\varphi\)=\\ln\|\\mathcal\{D\}\|\. Fornx,t\>ncn\_\{x,t\}\>n\_\{c\}any BIRD model memorizes for highly informative observations, while fornx,t<ncn\_\{x,t\}<n\_\{c\}it generalizes for information restricted observations\.
#### A general co\-dimension one phase boundary\.
Of course in full generality, in thejointspace of the amount of data\|𝒟\|\|\\mathcal\{D\}\|, the timettin the reverse process, and any family of observations𝒞x,t\\mathcal\{C\}\_\{x,t\}of varying information content, the equationI\(𝒞t,x;φ\)=ln\|𝒟\|I\(\\mathcal\{C\}\_\{t,x\};\\varphi\)=\\ln\|\\mathcal\{D\}\|determines a co\-dimension one phase boundary between the memorization and generalization phase\. The three factors of amount of data, time in reverse process, and information content of observations will compete to determine which phase the BIRD model is in\. Small amounts of data, late or small timestt, and large observational information content, each favor the memorization phase, while opposite directions favor generalization\. As we will see, BIRD models can combat memorization at late times in the reverse process, by reducing the information content of observations𝒞x,t\\mathcal\{C\}\_\{x,t\}asttdecreases\. The observational restriction impedes the memorization at small timestt, and if strong enough, favors generalization instead\.
### D\.4A simple calculation of the posterior entropy in the generalizing phase
Consider a channel𝒞x,t\\mathcal\{C\}\_\{x,t\}and a set of codewords or images\|𝒟\|\|\\mathcal\{D\}\|drawn i\.i\.d\. from the distributionP0P\_\{0\}\. Given a inputφ\\varphithe channel produces the output𝒞x,t\\mathcal\{C\}\_\{x,t\}with probability densityP\(𝒞x,t\|φ\)P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\)\. We can define the training set probability distribution conditioned on the channel outputsPtrain:ℝd→ℝP\_\{train\}:\\mathbb\{R\}^\{d\}\\rightarrow\\mathbb\{R\}:
Ptrain\(𝒞x,t\)=1\|𝒟\|∑φ∈𝒟P\(𝒞x,t\|φ\)P\_\{train\}\(\\mathcal\{C\}\_\{x,t\}\)=\\frac\{1\}\{\|\\mathcal\{D\}\|\}\\sum\_\{\\varphi\\in\\mathcal\{D\}\}P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\)\(41\)
We can then use Bayes rule to find the Bayesian probability over the training set given𝒞x,t\\mathcal\{C\}\_\{x,t\}Ptrain:𝒟×ℝd→ℝP\_\{train\}:\\mathcal\{D\}\\times\\mathbb\{R\}^\{d\}\\rightarrow\\mathbb\{R\}\.
Ptrain\(φ\|𝒞x,t\)=P\(𝒞x,t\|φ\)Ptrain\(𝒞x,t\)=P\(𝒞x,t\|φ\)∑φ′∈𝒟P\(𝒞x,t\|φ′\)P\_\{train\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)=\\frac\{P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\)\}\{P\_\{train\}\(\\mathcal\{C\}\_\{x,t\}\)\}=\\frac\{P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\)\}\{\\sum\_\{\\varphi^\{\\prime\}\\in\\mathcal\{D\}\}P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi^\{\\prime\}\)\}\(42\)Note that bothPtrainP\_\{train\}distributions are random variables dependent on the codewords𝒟\\mathcal\{D\}\. We also define the test set probabilities, which are not random\.
Ptest\(𝒞x,t\)\\displaystyle P\_\{test\}\(\\mathcal\{C\}\_\{x,t\}\)=∫𝑑φP0\(φ\)P\(𝒞x,t\|φ\)\\displaystyle=\\int d\\varphi\\\>P\_\{0\}\(\\varphi\)P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\)\(43\)Ptest\(φ\|𝒞x,t\)\\displaystyle P\_\{test\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)=P\(𝒞x,t\|φ\)Ptest\(𝒞x,t\)\\displaystyle=\\frac\{P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\)\}\{P\_\{test\}\(\\mathcal\{C\}\_\{x,t\}\)\}\(44\)We would like to understand the entropy of the Bayesian posterior on the training set:
S\[Ptrain\(φ\|𝒞x,t\)\]\\displaystyle S\[P\_\{train\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\]=∑φ∈𝒟−Ptrain\(φ\|𝒞x,t\)lnPtrain\(φ\|𝒞x,t\)\\displaystyle=\\sum\_\{\\varphi\\in\\mathcal\{D\}\}\-P\_\{train\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\\ln P\_\{train\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\(45\)=∑φ∈𝒟−P\(𝒞x,t\|φ\)∑φ′∈𝒟P\(𝒞x,t\|φ′\)lnP\(𝒞x,t\|φ\)∑φ′∈𝒟P\(𝒞x,t\|φ′\)\\displaystyle=\\sum\_\{\\varphi\\in\\mathcal\{D\}\}\-\\frac\{P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\)\}\{\\sum\_\{\\varphi^\{\\prime\}\\in\\mathcal\{D\}\}P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi^\{\\prime\}\)\}\\ln\\frac\{P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\)\}\{\\sum\_\{\\varphi^\{\\prime\}\\in\\mathcal\{D\}\}P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi^\{\\prime\}\)\}\(46\)=ln\(∑φ∈𝒟P\(𝒞x,t\|φ\)\)−1∑φ′∈𝒟P\(𝒞x,t\|φ\)∑φ∈𝒟P\(𝒞x,t\|φ\)lnP\(𝒞x,t\|φ\)\\displaystyle=\\ln\\Big\(\\sum\_\{\\varphi\\in\\mathcal\{D\}\}P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\)\\Big\.\)\-\\frac\{1\}\{\\sum\_\{\\varphi^\{\\prime\}\\in\\mathcal\{D\}\}P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\)\}\\sum\_\{\\varphi\\in\\mathcal\{D\}\}P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\)\\ln P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\)\(47\)=ln\|𝒟\|\+ln\(∑φ∈𝒟P\(𝒞x,t\|φ\)\|𝒟\|\)\\displaystyle=\\ln\|\\mathcal\{D\}\|\+\\ln\\Big\(\\frac\{\\sum\_\{\\varphi\\in\\mathcal\{D\}\}P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\)\}\{\|\\mathcal\{D\}\|\}\\Big\.\)\(48\)−1∑φ∈𝒟P\(𝒞x,t\|φ\)/\|𝒟\|1\|𝒟\|∑φ∈𝒟P\(𝒞x,t\|φ\)lnP\(𝒞x,t\|φ\)\\displaystyle\\hskip 20\.0pt\-\\frac\{1\}\{\\sum\_\{\\varphi\\in\\mathcal\{D\}\}P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\)/\|\\mathcal\{D\}\|\}\\frac\{1\}\{\|\\mathcal\{D\}\|\}\\sum\_\{\\varphi\\in\\mathcal\{D\}\}P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\)\\ln P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\)\(49\)All the terms in this sum are independent random variables\. As the dataset becomes large, the central limit theorem lets us replace the empirical means with their actual means up to aO\(1\|𝒟\|\)O\(\\frac\{1\}\{\\sqrt\{\|\\mathcal\{D\}\|\}\}\)corrections, so long as their means are non\-zero\. This shows that
1\|𝒟\|∑φ∈𝒟P\(𝒞x,t\|φ\)\\displaystyle\\frac\{1\}\{\|\\mathcal\{D\}\|\}\\sum\_\{\\varphi\\in\\mathcal\{D\}\}P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\)→𝔼φ∼P0\[P\(𝒞x,t\|φ\)\]\\displaystyle\\rightarrow\{\{\\mathbb\{E\}\}\}\_\{\\varphi\\sim P\_\{0\}\}\[P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\)\]\(50\)=∫𝑑φP0\(φ\)P\(𝒞x,t\|φ\)=Ptest\(𝒞x,t\)\\displaystyle=\\int d\\varphi P\_\{0\}\(\\varphi\)P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\)=P\_\{test\}\(\\mathcal\{C\}\_\{x,t\}\)\(51\)−1\|𝒟\|∑φ∈𝒟P\(𝒞x,t\|φ\)lnP\(𝒞x,t\|φ\)\\displaystyle\-\\frac\{1\}\{\|\\mathcal\{D\}\|\}\\sum\_\{\\varphi\\in\\mathcal\{D\}\}P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\)\\ln P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\)→𝔼φ∼P0\[−P\(𝒞x,t\|φ\)lnP\(𝒞x,t\|φ\)\]\\displaystyle\\rightarrow\{\{\\mathbb\{E\}\}\}\_\{\\varphi\\sim P\_\{0\}\}\[\-P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\)\\ln P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\)\]\(52\)=−∫𝑑φP0\(φ\)P\(𝒞x,t\|φ\)lnP\(𝒞x,t\|φ\)\\displaystyle=\-\\int d\\varphi P\_\{0\}\(\\varphi\)P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\)\\ln P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\)\(53\)=−∫𝑑φPtest\(𝒞x,t,φ\)\(lnPtest\(𝒞x,t\)\+lnP\(𝒞x,t,φ\)P0\(φ\)Ptest\(𝒞x,t\)\)\\displaystyle\\hskip\-80\.0pt=\-\\int d\\varphi P\_\{test\}\(\\mathcal\{C\}\_\{x,t\},\\varphi\)\(\\ln P\_\{test\}\(\\mathcal\{C\}\_\{x,t\}\)\+\\ln\\frac\{P\(\\mathcal\{C\}\_\{x,t\},\\varphi\)\}\{P\_\{0\}\(\\varphi\)P\_\{test\}\(\\mathcal\{C\}\_\{x,t\}\)\}\)\(54\)=Ptest\(𝒞x,t\)\(−lnPtest\(𝒞x,t\)−∫𝑑φPtest\(φ\|𝒞x,t\)\(lnPtest\(φ\|𝒞x,t\)P0\(φ\)\)\)\\displaystyle\\hskip\-80\.0pt=P\_\{test\}\(\\mathcal\{C\}\_\{x,t\}\)\\Big\(\-\\ln P\_\{test\}\(\\mathcal\{C\}\_\{x,t\}\)\-\\int d\\varphi P\_\{test\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\(\\ln\\frac\{P\_\{test\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\}\{P\_\{0\}\(\\varphi\)\}\)\\Big\)\(55\)=Ptest\(𝒞x,t\)\(−lnPtest\(𝒞x,t\)−DKL\(Ptest\(φ\|𝒞x,t\)\|\|P0\(φ\)\)\)\\displaystyle\\hskip\-80\.0pt=P\_\{test\}\(\\mathcal\{C\}\_\{x,t\}\)\\Big\(\-\\ln P\_\{test\}\(\\mathcal\{C\}\_\{x,t\}\)\-D\_\{KL\}\(P\_\{test\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\|\|P\_\{0\}\(\\varphi\)\)\\Big\)\(56\)Plugging equations[51](https://arxiv.org/html/2607.08041#A4.E51)and[56](https://arxiv.org/html/2607.08041#A4.E56)into equation[48](https://arxiv.org/html/2607.08041#A4.E48)\-[49](https://arxiv.org/html/2607.08041#A4.E49)we see,
S\[Ptrain\(φ\|𝒞x,t\)\]\\displaystyle S\[P\_\{train\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\]=ln\|𝒟\|\+lnPtest\(𝒞x,t\)\\displaystyle=\\ln\|\\mathcal\{D\}\|\+\\ln P\_\{test\}\(\\mathcal\{C\}\_\{x,t\}\)\(57\)\+1Ptest\(𝒞x,t\)Ptest\(𝒞x,t\)\(−lnPtest\(𝒞x,t\)−DKL\(Ptest\(φ\|𝒞x,t\)\|\|P0\(φ\)\)\)\\displaystyle\+\\frac\{1\}\{P\_\{test\}\(\\mathcal\{C\}\_\{x,t\}\)\}P\_\{test\}\(\\mathcal\{C\}\_\{x,t\}\)\\Big\(\-\\ln P\_\{test\}\(\\mathcal\{C\}\_\{x,t\}\)\-D\_\{KL\}\(P\_\{test\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\|\|P\_\{0\}\(\\varphi\)\)\\Big\)\(58\)\+O\(1\|𝒟\|\)\\displaystyle\+O\(\\frac\{1\}\{\\sqrt\{\|\\mathcal\{D\}\|\}\}\)\(59\)=ln\|𝒟\|−DKL\(Ptest\(φ\|𝒞x,t\)\|\|P0\(φ\)\)\+O\(1\|𝒟\|\)\\displaystyle=\\ln\|\\mathcal\{D\}\|\-D\_\{KL\}\(P\_\{test\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\|\|P\_\{0\}\(\\varphi\)\)\+O\(\\frac\{1\}\{\\sqrt\{\|\\mathcal\{D\}\|\}\}\)\(60\)When either average above becomes smaller thanO\(1D\)O\(\\frac\{1\}\{\\sqrt\{D\}\}\)the fluctuations might dominate over the mean\. At the same time, we know that entropy must be strictly non\-negative and should decrease as the information increases, so we can guess:
S\[Ptrain\(φ\|𝒞x,t\)\]=max\(ln\|𝒟\|−DKL\(Ptest\(φ\|𝒞x,t\)\|\|P0\(φ\)\),0\)S\[P\_\{train\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\]=\\operatorname\{max\}\(\\ln\|\\mathcal\{D\}\|\-D\_\{KL\}\(P\_\{test\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\|\|P\_\{0\}\(\\varphi\)\),0\)\(61\)Because the central limit theorem no longer applies, it is difficult to estimate the error in this formula\. As it turns out,O\(1/D\)O\(1/\\sqrt\{D\}\)is too optimistic in the memorizing phase\. In the section below, we use the random energy model to prove that this formula becomes exact in the large\|𝒟\|\|\\mathcal\{D\}\|limit withO\(lnln\|𝒟\|ln\|𝒟\|\)O\(\\frac\{\\ln\\ln\|\\mathcal\{D\}\|\}\{\\ln\|\\mathcal\{D\}\|\}\)finite size corrections\. If we average this result over𝒞x,t\\mathcal\{C\}\_\{x,t\}drawn fromPtestP\_\{test\}\.
𝔼𝒞x,t∼Ptest\[S\[Ptrain\(φ\|𝒞x,t\)\]\]=max\(ln\|𝒟\|−I\(φ;𝒞x,t\),0\)\{\{\\mathbb\{E\}\}\}\_\{\\mathcal\{C\}\_\{x,t\}\\sim P\_\{test\}\}\[S\[P\_\{train\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\]\]=\\operatorname\{max\}\(\\ln\|\\mathcal\{D\}\|\-I\(\\varphi;\\mathcal\{C\}\_\{x,t\}\),0\)\(62\)
### D\.5A proof from the Random Energy Model
In this section, we give several proofs based on methods from statistical mechanics, inspired by recent approaches in the literature of diffusion models as in\[[6](https://arxiv.org/html/2607.08041#bib.bib6),[43](https://arxiv.org/html/2607.08041#bib.bib43)\]\. We start by defining an energy functionEEand inverse temperatureβ=1/T\\beta=1/Tso thatPtrainP\_\{train\}corresponds to a Boltzmann distribution:
−βE\(φ\|𝒞x,t\)=lnP\(𝒞x,t\|φ\)\\displaystyle\-\\beta E\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)=\\ln P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\)\(63\)Ptrain\(φ\|𝒞x,t\)=e−βE\(φ\|𝒞x,t\)∑φ′∈𝒟e−βE\(φ′\|𝒞x,t\)\\displaystyle P\_\{train\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)=\\frac\{e^\{\-\\beta E\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\}\}\{\\sum\_\{\\varphi^\{\\prime\}\\in\\mathcal\{D\}\}e^\{\-\\beta E\(\\varphi^\{\\prime\}\|\\mathcal\{C\}\_\{x,t\}\)\}\}\(64\)Therefore, the energyE\(φ\|𝒞x,t\)E\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)is simply the temperatureTTtimes the minus log likelihood of making the channel observation𝒞x,t\\mathcal\{C\}\_\{x,t\}at timettunder the forward diffusion process starting from a data pointφ\\varphiat timet=0t=0\. If we fix the channel observation𝒞x,t\\mathcal\{C\}\_\{x,t\}, we can think ofE\(φ\|𝒞x,t\)E\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)in \([63](https://arxiv.org/html/2607.08041#A4.E63)\) as an energy function on the training dataφ∈𝒟\\varphi\\in\\mathcal\{D\}, with a high \(low\) energy assigned toφ\\varphiindicating a low \(high\) likelihood thatφ\\varphiwas the origin of the channel observation𝒞x,t\\mathcal\{C\}\_\{x,t\}in the forward diffusion process\. After normalization, the Boltzmann distribution in \([64](https://arxiv.org/html/2607.08041#A4.E64)\) is simply the posterior probability that a particular training data pointφ∈𝒟\\varphi\\in\\mathcal\{D\}at timet=0t=0was the origin of the conditioned channel observation𝒞x,t\\mathcal\{C\}\_\{x,t\}in the forward diffusion process\.
###### Theorem D\.1\(Independent Pointwise Posterior Entropy\)
If𝒟\\mathcal\{D\}is a set of training images drawn i\.i\.d\. fromP0P\_\{0\}and𝒞x,t\\mathcal\{C\}\_\{x,t\}is a channel such thatP\(E\|𝒞x,t\)P\(E\|\\mathcal\{C\}\_\{x,t\}\)has a rate function for largeEE, with a continuous first derivative\. We also assume that both the rate function and its derivative have a finite number of maxima and minima atE=O\(1\)E=O\(1\)\.
S\[Ptrain\(φ\|𝒞x,t\)\]=\{ln\|𝒟\|−DKL\(Ptest\(φ\|𝒞x,t\)\|\|P0\(φ\)\)ln\|𝒟\|\>DKL\(Ptest\(φ\|𝒞x,t\)\|\|P0\(φ\)\)0ln\|𝒟\|≤DKL\(Ptest\(φ\|𝒞x,t\)\|\|P0\(φ\)\)S\[P\_\{train\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\]=\\begin\{cases\}\\ln\|\\mathcal\{D\}\|\-D\_\{KL\}\(P\_\{test\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\|\|P\_\{0\}\(\\varphi\)\)&\\ln\|\\mathcal\{D\}\|\>D\_\{KL\}\(P\_\{test\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\|\|P\_\{0\}\(\\varphi\)\)\\\\ 0&\\ln\|\\mathcal\{D\}\|\\leq D\_\{KL\}\(P\_\{test\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\|\|P\_\{0\}\(\\varphi\)\)\\end\{cases\}\(65\)up toO\(lnln\|𝒟\|ln\|𝒟\|\)O\(\\frac\{\\ln\\ln\|\\mathcal\{D\}\|\}\{\\ln\|\\mathcal\{D\}\|\}\)corrections\.
###### Proof\.
We will begin by considering the case where𝒞x,t\\mathcal\{C\}\_\{x,t\}is drawn from the test set\. This means𝒞x,t\\mathcal\{C\}\_\{x,t\}, as a random channel observation, is statisticallyindependentof the random training data𝒟\\mathcal\{D\}\. If𝒞x,t\\mathcal\{C\}\_\{x,t\}were drawn from the training set, then the dataset sample that generates𝒞x,t\\mathcal\{C\}\_\{x,t\}would not be identically distributed and would break the i\.i\.d assumption in hypothesis of this theorem\. We will first fix𝒞x,t\\mathcal\{C\}\_\{x,t\}to be an arbitrary value, then we will average over the distribution of𝒞x,t\\mathcal\{C\}\_\{x,t\}\. Now since the data pointsφ∈𝒟\\varphi\\in\\mathcal\{D\}are drawn i\.i\.d\. fromP0\(φ\)P\_\{0\}\(\\varphi\), the energiesE\(φ\|𝒞x,t\)E\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)of each data pointφ\\varphican be thought of as independent random variables\. Therefore, the Boltzmann distribution in \([64](https://arxiv.org/html/2607.08041#A4.E64)\) is equivalent to an independent random energy model \(REM\) as first described by\[[44](https://arxiv.org/html/2607.08041#bib.bib44)\]\(see also\[[45](https://arxiv.org/html/2607.08041#bib.bib45)\]for a nice pedagogical treatment\)\. The REM is solvable by a saddle point approximation which becomes exact in a “thermodynamic” limit, which in our case, corresponds to a large amount of training data, or large\|𝒟\|\|\\mathcal\{D\}\|\. We define
P\(E′\|𝒞x,t\)=∫𝑑φδ\(E′−E\(φ\|𝒞x,t\)\)P0\(φ\)P\(E^\{\\prime\}\|\\mathcal\{C\}\_\{x,t\}\)=\\int d\\varphi\\\>\\delta\(E^\{\\prime\}\-E\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\)P\_\{0\}\(\\varphi\)\(66\)This is simply the distribution of energiesE′E^\{\\prime\}we expect to see conditioned on a fixed channel observation𝒞x,t\\mathcal\{C\}\_\{x,t\}where the randomness comes from the distribution of a single training pointP0\(φ\)P\_\{0\}\(\\varphi\)\. In the REM, the log number of states,ln\|𝒟\|\\ln\|\\mathcal\{D\}\|, is the extensive parameter, originally corresponding to the number of spins\. For ML readers not familiar with the statistical mechanics language, the extensive parameter,ln\|𝒟\|\\ln\|\\mathcal\{D\}\|, is used to group together different values based on their relative size to it\.
intensive⇌O\(1\)\\displaystyle\\rightleftharpoons O\(1\)\(67\)extensive⇌O\(ln\|𝒟\|\)\\displaystyle\\rightleftharpoons O\(\\ln\|\\mathcal\{D\}\|\)\(68\)exponentially large⇌O\(eαln\|𝒟\|\)=O\(\|𝒟\|α\)\\displaystyle\\rightleftharpoons O\(e^\{\\alpha\\ln\|\\mathcal\{D\}\|\}\)=O\(\|\\mathcal\{D\}\|^\{\\alpha\}\)\(69\)exponentially small⇌O\(e−αln\|𝒟\|\)=O\(1\|𝒟\|α\)\\displaystyle\\rightleftharpoons O\(e^\{\-\\alpha\\ln\|\\mathcal\{D\}\|\}\)=O\(\\frac\{1\}\{\|\\mathcal\{D\}\|^\{\\alpha\}\}\)\(70\)In some strict sense, this means that we are imagining that all of the variables in the problem have some dependence onln\|𝒟\|\\ln\|\\mathcal\{D\}\|which allows us to take the limitln\|𝒟\|→∞\\ln\|\\mathcal\{D\}\|\\rightarrow\\infty, called the thermodynamic limit\. To make this analysis quite general, we will avoid explicit expressions for this dependence onln\|𝒟\|\\ln\|\\mathcal\{D\}\|and only express asymptotic scaling\. We would like to define this limit so there are sensible answers for the entropy of the posterior and relevant quantities at polynomial orders inln\|𝒟\|\\ln\|\\mathcal\{D\}\|\. By that we mean that we are interested in finding asymptotic series asln\|𝒟\|→∞\\ln\|\\mathcal\{D\}\|\\rightarrow\\inftylike𝔼\[O\]=∑n=−∞NOn\(ln\|𝒟\|\)n\{\{\\mathbb\{E\}\}\}\[O\]=\\sum\_\{n=\-\\infty\}^\{N\}O\_\{n\}\(\\ln\|\\mathcal\{D\}\|\)^\{n\}\. In this context, quantities which are exponentially small do not show up in any order of these asymptotic series since they decay faster than any rational function\. Therefore, exponentially small quantities can simply be set to 0\. For the REM to have a sensible thermodynamic limit, we require that the energy is extensive\. This is equivalent to the existence of a functionIIcalled the rate function defined in the following way:
ϵ\(φ\|𝒞x,t\)\\displaystyle\\epsilon\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)=E\(φ\|𝒞x,t\)ln\|𝒟\|\\displaystyle=\\frac\{E\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\}\{\\ln\|\\mathcal\{D\}\|\}\(71\)I\(ϵ\|𝒞x,t\)\\displaystyle I\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\):=−lnP\(ϵln\|𝒟\|\|𝒞x,t\)ln\|𝒟\|=Ω\(1\)\\displaystyle:=\\frac\{\-\\ln P\(\\epsilon\\ln\|\\mathcal\{D\}\|\\Big\|\\mathcal\{C\}\_\{x,t\}\)\}\{\\ln\|\\mathcal\{D\}\|\}=\\Omega\(1\)\(72\)P\(ϵ\|𝒞x,t\)\\displaystyle P\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)=e−ln\|𝒟\|I\(ϵ\|𝒞x,t\)\\displaystyle=e^\{\-\\ln\|\\mathcal\{D\}\|I\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)\}\(73\)ϵ\\epsilonis called the intensive energy because we removed theln\|𝒟\|\\ln\|\\mathcal\{D\}\|scaling of the energy by dividing\. We also assume that the inverse temperature is intensive\. The existence of the rate function implies that there is sensible probability distribution for the intensive energy\.
If we unwrap the asymptotic expression in[72](https://arxiv.org/html/2607.08041#A4.E72), we find
−lnP\(ϵln\|𝒟\|\|𝒞x,t\)ln\|𝒟\|\\displaystyle\\frac\{\-\\ln P\(\\epsilon\\ln\|\\mathcal\{D\}\|\\Big\|\\mathcal\{C\}\_\{x,t\}\)\}\{\\ln\|\\mathcal\{D\}\|\}≥C1\\displaystyle\\geq C\_\{1\}\(74\)lnP\(ϵln\|𝒟\|\|𝒞x,t\)\\displaystyle\\ln P\(\\epsilon\\ln\|\\mathcal\{D\}\|\\Big\|\\mathcal\{C\}\_\{x,t\}\)≤−C1ln\|𝒟\|\\displaystyle\\leq\-C\_\{1\}\\ln\|\\mathcal\{D\}\|\(75\)P\(ϵln\|𝒟\|\|𝒞x,t\)\\displaystyle P\(\\epsilon\\ln\|\\mathcal\{D\}\|\\Big\|\\mathcal\{C\}\_\{x,t\}\)≤C2e−C1ln\|𝒟\|\\displaystyle\\leq C\_\{2\}e^\{\-C\_\{1\}\\ln\|\\mathcal\{D\}\|\}\(76\)⇐P\(E\|𝒞x,t\)\\displaystyle\\Leftarrow P\(E\|\\mathcal\{C\}\_\{x,t\}\)≤C2e−C1E−o\(1\)\.\\displaystyle\\leq C\_\{2\}e^\{\-C\_\{1\}E\-o\(1\)\}\.\(77\)We’ll now look at the ”density of states” which measures how many energies sit inside of a window aroundϵ\\epsilon\.
ρ\(ϵ\|𝒞x,t\)=1W∑φ∈𝒟𝟏ϵ\(φ\|𝒞x,t\)∈\[ϵ,ϵ\+W\]\\rho\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)=\\frac\{1\}\{W\}\\sum\_\{\\varphi\\in\\mathcal\{D\}\}\\mathbf\{1\}\_\{\\epsilon\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\\in\[\\epsilon,\\epsilon\+W\]\}\(78\)where𝟏E\(φ\|𝒞x,t\)∈\[ϵ,ϵ\+W\]\\mathbf\{1\}\_\{E\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\\in\[\\epsilon,\\epsilon\+W\]\}is the indicator function
𝟏E\(φ\|𝒞x,t\)∈\[ϵ,ϵ\+W\]=\{1ϵ\(φ\|𝒞x,t\)∈\[ϵ,ϵ\+W\]0ϵ\(φ\|𝒞x,t\)∉\[ϵ,ϵ\+W\]\\mathbf\{1\}\_\{E\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\\in\[\\epsilon,\\epsilon\+W\]\}=\\begin\{cases\}1&\\epsilon\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\\in\[\\epsilon,\\epsilon\+W\]\\\\ 0&\\epsilon\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\\notin\[\\epsilon,\\epsilon\+W\]\\end\{cases\}\(79\)Saddle\-point approximation leads to an expected density of states
𝔼\[ρ\(ϵ\|𝒞x,t\)\]\\displaystyle\\mathbb\{E\}\[\\rho\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)\]=𝔼\[∑φ∈𝒟𝟏E\(φ\|𝒞x,t\)∈\[ϵ,ϵ\+W\]\]/W=\|𝒟\|𝔼φ∼P0\[𝟏E\(φ\|𝒞x,t\)∈\[ϵ,ϵ\+W\]\]/W\\displaystyle=\\mathbb\{E\}\[\\sum\_\{\\varphi\\in\\mathcal\{D\}\}\\mathbf\{1\}\_\{E\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\\in\[\\epsilon,\\epsilon\+W\]\}\]/W=\|\\mathcal\{D\}\|\\mathbb\{E\}\_\{\\varphi\\sim P\_\{0\}\}\[\\mathbf\{1\}\_\{E\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\\in\[\\epsilon,\\epsilon\+W\]\}\]/W\(80\)=∫ϵϵ\+WdϵWeln\|𝒟\|\(1−I\(ϵ\|𝒞x,t\)\)\\displaystyle=\\int\_\{\\epsilon\}^\{\\epsilon\+W\}\\frac\{d\\epsilon\}\{W\}\\\>e^\{\\ln\|\\mathcal\{D\}\|\\Big\(1\-I\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)\\Big\)\}\(81\)=\{eln\|𝒟\|\(1−I\(ϵ\|𝒞x,t\)\)\+o\(1\)1\>I\(ϵ\|𝒞x,t\)0\+O\(e−ln\|𝒟\|\)1≤I\(ϵ\|𝒞x,t\)\\displaystyle=\\begin\{cases\}e^\{\\ln\|\\mathcal\{D\}\|\(1\-I\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)\)\+o\(1\)\}&1\>I\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)\\\\ 0\+O\(e^\{\-\\ln\|\\mathcal\{D\}\|\}\)&1\\leq I\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)\\end\{cases\}\(82\)Since𝟏E\(φ\|𝒞x,t\)∈\[ϵ,ϵ\+W\]\\mathbf\{1\}\_\{E\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\\in\[\\epsilon,\\epsilon\+W\]\}is its own square, the variance ofρ\\rhois equal to its expectation which implies the standard deviation isO\(\|𝒟\|\)O\(\\sqrt\{\|\\mathcal\{D\}\|\}\)where the leading order term isO\(\|𝒟\|\)O\(\|\\mathcal\{D\}\|\)\. Therefore, the leading order term of the density of states is deterministic
ρ\(ϵ\|𝒞x,t\)=\{eln\|𝒟\|\(1−I\(ϵ\|𝒞x,t\)\)\+O\(eln\|D\|/2\)1\>I\(ϵ\|𝒞x,t\)0\+O\(e−ln\|𝒟\|\)1≤I\(ϵ\|𝒞x,t\)\\rho\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)=\\begin\{cases\}e^\{\\ln\|\\mathcal\{D\}\|\(1\-I\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)\)\}\+O\(e^\{\\ln\|D\|/2\}\)&1\>I\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)\\\\ 0\+O\(e^\{\-\\ln\|\\mathcal\{D\}\|\}\)&1\\leq I\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)\\end\{cases\}\(83\)The micro\-canonical entropyS\(ϵ\)S\(\\epsilon\)is just the log density of states\. It is deterministic to all orders inln\|𝒟\|\\ln\|\\mathcal\{D\}\|\.
S\(ϵ\|𝒞x,t\)\\displaystyle S\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)=lnρ\(ϵ\|𝒞x,t\)\\displaystyle=\\ln\\rho\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)\(84\)=\{ln\(eln\|𝒟\|\(1−I\(ϵ\|𝒞x,t\)\)\+O\(eln\|D\|/2\)\)1\>I\(ϵ\|𝒞x,t\)ln\(0\+O\(e−ln\|𝒟\|\)\)1≤I\(ϵ\|𝒞x,t\)\\displaystyle=\\begin\{cases\}\\ln\\Big\(e^\{\\ln\|\\mathcal\{D\}\|\(1\-I\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)\)\}\+O\(e^\{\\ln\|D\|/2\}\)\\Big\.\)&1\>I\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)\\\\ \\ln\(0\+O\(e^\{\-\\ln\|\\mathcal\{D\}\|\}\)\)&1\\leq I\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)\\end\{cases\}\(85\)=\{ln\|𝒟\|\(1−I\(ϵ\|𝒞x,t\)\)\+ln\(1\+O\(e−ln\|D\|/2\)\)1\>I\(ϵ\|𝒞x,t\)−∞\+O\(e−ln\|𝒟\|\)1≤I\(ϵ\|𝒞x,t\)\\displaystyle=\\begin\{cases\}\\ln\|\\mathcal\{D\}\|\(1\-I\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)\)\+\\ln\(1\+O\(e^\{\-\\ln\|D\|/2\}\)\)&1\>I\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)\\\\ \-\\infty\+O\(e^\{\-\\ln\|\\mathcal\{D\}\|\}\)&1\\leq I\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)\\end\{cases\}\(86\)=\{ln\|𝒟\|\(1−I\(ϵ\|𝒞x,t\)\)\+O\(e−ln\|𝒟\|/2\)1\>I\(ϵ\|𝒞x,t\)−∞\+O\(e−ln\|𝒟\|\)1≤I\(ϵ\|𝒞x,t\)\\displaystyle=\\begin\{cases\}\\ln\|\\mathcal\{D\}\|\(1\-I\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)\)\+O\(e^\{\-\\ln\|\\mathcal\{D\}\|/2\}\)&1\>I\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)\\\\ \-\\infty\+O\(e^\{\-\\ln\|\\mathcal\{D\}\|\}\)&1\\leq I\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)\\end\{cases\}\(87\)Between equation[85](https://arxiv.org/html/2607.08041#A4.E85)and equation[86](https://arxiv.org/html/2607.08041#A4.E86), we multiplied and divided by the deterministic contribution to the density of states and then used the logarithm to separate the deterministic contribution into its own term\. From equation \([87](https://arxiv.org/html/2607.08041#A4.E87)\), we see that the entropy is deterministic up to exponentially small corrections\. We now define the annealed micro\-canonical entropy densitysann,MC\(ϵ\|𝒞x,t\)s\_\{ann,MC\}\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)and the micro\-canonical entropy densitys\(ϵ\|𝒞x,t\)s\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)as follows:
sann,MC\(ϵ\|𝒞x,t\)\\displaystyle s\_\{ann,MC\}\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\):=ln𝔼\[ρ\(ϵ\|𝒞x,t\)\]ln\|𝒟\|=1−I\(ϵ\|𝒞x,t\)\\displaystyle:=\\frac\{\\ln\\mathbb\{E\}\[\\rho\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)\]\}\{\\ln\|\\mathcal\{D\}\|\}=1\-I\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)\(88\)s\(ϵ\|𝒞x,t\):=S\(ϵ\|𝒞x,t\)ln\|𝒟\|\\displaystyle s\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\):=\\frac\{S\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)\}\{\\ln\|\\mathcal\{D\}\|\}=\{sann,MC\(ϵ\|𝒞x,t\)sann,MC\(ϵ\|𝒞x,t\)\>0−∞sann,MC\(ϵ\|𝒞x,t\)≤0\\displaystyle=\\begin\{cases\}s\_\{ann,MC\}\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)&s\_\{ann,MC\}\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)\>0\\\\ \-\\infty&s\_\{ann,MC\}\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)\\leq 0\\end\{cases\}\(89\)Our goal is to calculate the entropy of the Gibbs distribution which is the posterior over training images\. This requires we switch from the micro\-canonical ensemble perspective to the canonical ensemble perspective\. We can define the standard thermodynamic functions for the canonical ensemble:
The Partition Function⇌Z\(β\|𝒞x,t\)\\displaystyle\\text\{The Partition Function\}\\rightleftharpoons Z\(\\beta\|\\mathcal\{C\}\_\{x,t\}\):=∑φ∈𝒟e−βE\(φ\|𝒞x,t\)\\displaystyle:=\\sum\_\{\\varphi\\in\\mathcal\{D\}\}e^\{\-\\beta E\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\}Gibbs Distribution⇌Pβ\(φ\|𝒞x,t\)\\displaystyle\\text\{Gibbs Distribution\}\\rightleftharpoons P\_\{\\beta\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\):=e−βE\(φ\|𝒞x,t\)Z\\displaystyle:=\\frac\{e^\{\-\\beta E\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\}\}\{Z\}The Free Energy⇌F\(β\|𝒞x,t\)\\displaystyle\\text\{The Free Energy\}\\rightleftharpoons F\(\\beta\|\\mathcal\{C\}\_\{x,t\}\):=lnZ\(β\|𝒞x,t\)−β\\displaystyle:=\\frac\{\\ln Z\(\\beta\|\\mathcal\{C\}\_\{x,t\}\)\}\{\-\\beta\}The Expected Energy⇌E\(β\|𝒞x,t\)\\displaystyle\\text\{The Expected Energy\}\\rightleftharpoons E\(\\beta\|\\mathcal\{C\}\_\{x,t\}\):=∑φ∈𝒟E\(φ\|𝒞x,t\)Pβ\(φ\|𝒞x,t\)\\displaystyle:=\\sum\_\{\\varphi\\in\\mathcal\{D\}\}E\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)P\_\{\\beta\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)The Canonical Entropy⇌S\(β\|𝒞x,t\)\\displaystyle\\text\{The Canonical Entropy\}\\rightleftharpoons S\(\\beta\|\\mathcal\{C\}\_\{x,t\}\)=∑φ−Pβ\(φ\|𝒞x,t\)lnPβ\(φ\|𝒞x,t\)\\displaystyle=\\sum\_\{\\varphi\}\-P\_\{\\beta\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\\ln P\_\{\\beta\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)=∑φ−e−βE\(φ\|𝒞x,t\)Z\(−lnZ−βE\(φ\|𝒞x,t\)\)\\displaystyle=\\sum\_\{\\varphi\}\-\\frac\{e^\{\-\\beta E\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\}\}\{Z\}\(\-\\ln Z\-\\beta E\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\)=lnZ\+β∑φE\(φ\|𝒞x,t\)Pβ\(φ\|𝒞x,t\)\\displaystyle=\\ln Z\+\\beta\\sum\_\{\\varphi\}E\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)P\_\{\\beta\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)=−βF\(β\|𝒞x,t\)\+βE\(β\|𝒞x,t\)\\displaystyle=\-\\beta F\(\\beta\|\\mathcal\{C\}\_\{x,t\}\)\+\\beta E\(\\beta\|\\mathcal\{C\}\_\{x,t\}\)In models with disorder, like the random energy model, you often look at the disorder averaged or ”annealed” partition function\. We can define the annealed versions of the functions in the canonical ensemble\. Since I’ve proven that some of the above functions are deterministic, I can replace them with their annealed averages where appropriate\.
Annealed Partition Function⇌Zann\(β\|𝒞x,t\)\\displaystyle\\text\{Annealed Partition Function\}\\rightleftharpoons Z\_\{ann\}\(\\beta\|\\mathcal\{C\}\_\{x,t\}\):=𝔼𝒟\[∑φ∈𝒟e−βE\(φ\|𝒞x,t\)\]\\displaystyle:=\{\{\\mathbb\{E\}\}\}\_\{\\mathcal\{D\}\}\[\\sum\_\{\\varphi\\in\\mathcal\{D\}\}e^\{\-\\beta E\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\}\]\(Canonical\) Annealed Free Energy⇌Fann\(β\|𝒞x,t\)\\displaystyle\\text\{\(Canonical\) Annealed Free Energy\}\\rightleftharpoons F\_\{ann\}\(\\beta\|\\mathcal\{C\}\_\{x,t\}\):=lnZann\(β\|𝒞x,t\)−β\\displaystyle:=\\frac\{\\ln Z\_\{ann\}\(\\beta\|\\mathcal\{C\}\_\{x,t\}\)\}\{\-\\beta\}\(Canonical\) Annealed Expected Energy⇌Eann\(β\|𝒞x,t\)\\displaystyle\\text\{\(Canonical\) Annealed Expected Energy\}\\rightleftharpoons E\_\{ann\}\(\\beta\|\\mathcal\{C\}\_\{x,t\}\):=1Zann𝔼𝒟\[∑φ∈𝒟E\(φ\|𝒞x,t\)e−βE\(φ\|𝒞x,t\)\]\\displaystyle:=\\frac\{1\}\{Z\_\{ann\}\}\{\{\\mathbb\{E\}\}\}\_\{\\mathcal\{D\}\}\[\\sum\_\{\\varphi\\in\\mathcal\{D\}\}E\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)e^\{\-\\beta E\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\}\]\(Canonical\) Annealed Canonical Entropy⇌sann,MC\(β\|𝒞x,t\)\\displaystyle\\text\{\(Canonical\) Annealed Canonical Entropy\}\\rightleftharpoons s\_\{ann,MC\}\(\\beta\|\\mathcal\{C\}\_\{x,t\}\)=−βFann\(β\|𝒞x,t\)\+βEann\(β\|𝒞x,t\)\\displaystyle=\-\\beta F\_\{ann\}\(\\beta\|\\mathcal\{C\}\_\{x,t\}\)\+\\beta E\_\{ann\}\(\\beta\|\\mathcal\{C\}\_\{x,t\}\)The free energy, expected energy, and canonical entropy are extensive so its natural to define the intensive densities for each\.
The Free Energy Density⇌f\(β\|𝒞x,t\)\\displaystyle\\text\{The Free Energy Density\}\\rightleftharpoons f\(\\beta\|\\mathcal\{C\}\_\{x,t\}\):=F\(β\|𝒞x,t\)ln\|𝒟\|\\displaystyle:=\\frac\{F\(\\beta\|\\mathcal\{C\}\_\{x,t\}\)\}\{\\ln\|\\mathcal\{D\}\|\}\(90\)Expected Energy Density⇌ϵ\(β\|𝒞x,t\)\\displaystyle\\text\{Expected Energy Density\}\\rightleftharpoons\\epsilon\(\\beta\|\\mathcal\{C\}\_\{x,t\}\):=E\(β\|𝒞x,t\)ln\|𝒟\|\\displaystyle:=\\frac\{E\(\\beta\|\\mathcal\{C\}\_\{x,t\}\)\}\{\\ln\|\\mathcal\{D\}\|\}\(91\)The Canonical Entropy Density⇌s\(β\|𝒞x,t\)\\displaystyle\\text\{The Canonical Entropy Density\}\\rightleftharpoons s\(\\beta\|\\mathcal\{C\}\_\{x,t\}\):=S\(β\|𝒞x,t\)ln\|𝒟\|\\displaystyle:=\\frac\{S\(\\beta\|\\mathcal\{C\}\_\{x,t\}\)\}\{\\ln\|\\mathcal\{D\}\|\}\(92\)Because this model has an exact Saddle point in the thermodynamic limit, the canonical ensemble functions will just be Legendre transforms of their micro\-canonical counter parts, but the regions where density of states drop to zero constrain the energy\. Similar to the micro\-canonical entropy, they will be deterministic to all orders inln\|𝒟\|\\ln\|\\mathcal\{D\}\|\. We can see this by using saddle point approximation to find the partition function\.
Z\(β\|𝒞x,t\)\\displaystyle Z\(\\beta\|\\mathcal\{C\}\_\{x,t\}\)=∑φ∈𝒟e−βE\(φ\|𝒞x,t\)=∫ℝ𝑑ϵe−ln\|𝒟\|βϵρ\(ϵ\|𝒞x,t\)\\displaystyle=\\sum\_\{\\varphi\\in\\mathcal\{D\}\}e^\{\-\\beta E\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\}=\\int\_\{\{\{\\mathbb\{R\}\}\}\}d\\epsilon\\\>e^\{\-\\ln\|\\mathcal\{D\}\|\\beta\\epsilon\}\\rho\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)=∫sann,MC\(ϵ\|𝒞x,t\)\>0𝑑ϵeln\|𝒟\|\(sann,MC\(ϵ\|𝒞x,t\)−βϵ\)\+O\(eln\|𝒟\|/2\)\\displaystyle=\\int\_\{s\_\{ann,MC\}\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)\>0\}d\\epsilon\\\>e^\{\\ln\|\\mathcal\{D\}\|\(s\_\{ann,MC\}\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)\-\\beta\\epsilon\)\}\+O\(e^\{\\ln\|\\mathcal\{D\}\|/2\}\)=exp\(ln\|𝒟\|\(maxϵs\.t\.sann,MC\(ϵ\)\>0sann,MC\(ϵ\|𝒞x,t\)−βϵ\)\+O\(lnln\|𝒟\|\)\)\\displaystyle=\\exp\(\\ln\|\\mathcal\{D\}\|\\Big\(\\operatorname\{max\}\_\{\\epsilon\\\>s\.t\.\\\>s\_\{ann,MC\}\(\\epsilon\)\>0\}s\_\{ann,MC\}\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)\-\\beta\\epsilon\\Big\)\+O\(\\ln\\ln\|\\mathcal\{D\}\|\)\)TheO\(lnln\|D\|\)O\(\\ln\\ln\|D\|\)error term above accounts for the normalization of the Gaussian fluctuations around the saddle point\. Expanding up to second order inϵ′−ϵ\(β\)\\epsilon^\{\\prime\}\-\\epsilon\(\\beta\)leads to a Gaussian with varianceσ2=O\(\(ln\|D\|\)−1\)\\sigma^\{2\}=O\(\(\\ln\|D\|\)^\{\-1\}\)\. The normalization of a Gaussian integral is\(2πσ2\)−1/2=exp\(−lnσ\+O\(1\)\)=exp\(O\(lnln\|𝒟\|\)\)\(2\\pi\\sigma^\{2\}\)^\{\-1/2\}=\\exp\(\-\\ln\\sigma\+O\(1\)\)=\\exp\(O\(\\ln\\ln\|\\mathcal\{D\}\|\)\), which produces this term\. Ifϵ\(β\)\\epsilon\(\\beta\)lies at a boundary of integration, the lowest order fluctuation is exponential and the normalization still scales asO\(\(ln\|𝒟\|\)−1\)=exp\(O\(lnln\|𝒟\|\)\)O\(\(\\ln\|\\mathcal\{D\}\|\)^\{\-1\}\)=\\exp\(O\(\\ln\\ln\|\\mathcal\{D\}\|\)\), so the same logic applies\. Now, using this result to find the free energy density:
−βf\(β\|𝒞x,t\)=maxϵs\.t\.sann,MC\(ϵ\)\>0sann,MC\(ϵ\|𝒞x,t\)−βϵ\+O\(lnln\|𝒟\|ln\|𝒟\|\)\.\-\\beta f\(\\beta\|\\mathcal\{C\}\_\{x,t\}\)=\\operatorname\{max\}\_\{\\epsilon\\\>s\.t\.\\\>s\_\{ann,MC\}\(\\epsilon\)\>0\}s\_\{ann,MC\}\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)\-\\beta\\epsilon\+O\(\\frac\{\\ln\\ln\|\\mathcal\{D\}\|\}\{\\ln\|\\mathcal\{D\}\|\}\)\.\(93\)We can similarly solve for the energy density
ϵ\(β\|𝒞x,t\)\\displaystyle\\epsilon\(\\beta\|\\mathcal\{C\}\_\{x,t\}\)=1Z∫ℝ𝑑ϵρ\(ϵ\|𝒞x,t\)ϵe−βln\|𝒟\|ϵ\\displaystyle=\\frac\{1\}\{Z\}\\int\_\{\{\{\\mathbb\{R\}\}\}\}d\\epsilon\\\>\\rho\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)\\epsilon e^\{\-\\beta\\ln\|\\mathcal\{D\}\|\\epsilon\}=1Z∫sann,MC\(ϵ\)\>0ϵeln\|𝒟\|\(sann,MC\(ϵ\|𝒞x,t\)−βϵ\)\\displaystyle=\\frac\{1\}\{Z\}\\int\_\{s\_\{ann,MC\}\(\\epsilon\)\>0\}\\epsilon e^\{\\ln\|\\mathcal\{D\}\|\(s\_\{ann,MC\}\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)\-\\beta\\epsilon\)\}=argmaxsann,MC\(ϵ\)\>0sann,MC\(ϵ\|𝒞x,t\)−βϵ\+O\(lnln\|𝒟\|ln\|𝒟\|\)\\displaystyle=\\text\{argmax\}\_\{s\_\{ann,MC\}\(\\epsilon\)\>0\}s\_\{ann,MC\}\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)\-\\beta\\epsilon\+O\(\\frac\{\\ln\\ln\|\\mathcal\{D\}\|\}\{\\ln\|\\mathcal\{D\}\|\}\)For the rest of the proof, I will omit theO\(lnln\|D\|/ln\|D\|\)O\(\\ln\\ln\|D\|/\\ln\|D\|\)term\. It should be assumed that all canonical ensemble functions have corrections of that order\. We make one last transformation from micro\-canonical to canonical transformation for annealed thermodynamic functions\. As always in thermodynamics, saddle point approximation shows that the annealed canonical functions concentrate to the annealed micro\-canonical functions in the largeln\|𝒟\|\\ln\|\\mathcal\{D\}\|limit\. I will elide some details since this is the third similar calculation\. Note thatϵ\\epsilonis now allowed to range over the entire real line instead of only wheresann,MC\(ϵ\)\>0s\_\{ann,MC\}\(\\epsilon\)\>0\.
Zann\(β\|𝒞x,t\)\\displaystyle Z\_\{ann\}\(\\beta\|\\mathcal\{C\}\_\{x,t\}\):=𝔼𝒟\[∑φ∈𝒟e−βE\(φ\|𝒞x,t\)\]=\|𝒟\|𝔼φ\[e−βE\(φ\|𝒞x,t\)\]\\displaystyle:=\{\{\\mathbb\{E\}\}\}\_\{\\mathcal\{D\}\}\[\\sum\_\{\\varphi\\in\\mathcal\{D\}\}e^\{\-\\beta E\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\}\]=\|\\mathcal\{D\}\|\{\{\\mathbb\{E\}\}\}\_\{\\varphi\}\[e^\{\-\\beta E\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\}\]=\|𝒟\|∫𝑑ϵeln\|𝒟\|\(sann,MC\(ϵ\|𝒞x,t\)−βϵ\)=exp\(ln\|𝒟\|maxϵsann,MC\(ϵ\|𝒞x,t\)−βϵ\)\\displaystyle=\|\\mathcal\{D\}\|\\int d\\epsilon e^\{\\ln\|\\mathcal\{D\}\|\(s\_\{ann,MC\}\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)\-\\beta\\epsilon\)\}=\\exp\(\\ln\|\\mathcal\{D\}\|\\operatorname\{max\}\_\{\\epsilon\}s\_\{ann,MC\}\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)\-\\beta\\epsilon\)⟹fann\(β\|𝒞x,t\)\\displaystyle\\implies f\_\{ann\}\(\\beta\|\\mathcal\{C\}\_\{x,t\}\)=lnZann\(β\|𝒞x,t\)−βln\|𝒟\|→−1βmaxϵsann,MC\(ϵ\|𝒞x,t\)−βϵ\\displaystyle=\\frac\{\\ln Z\_\{ann\}\(\\beta\|\\mathcal\{C\}\_\{x,t\}\)\}\{\-\\beta\\ln\|\\mathcal\{D\}\|\}\\rightarrow\\frac\{\-1\}\{\\beta\}\\operatorname\{max\}\_\{\\epsilon\}s\_\{ann,MC\}\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)\-\\beta\\epsilon⟹ϵann\(β\|𝒞x,t\)\\displaystyle\\implies\\epsilon\_\{ann\}\(\\beta\|\\mathcal\{C\}\_\{x,t\}\)=argmaxϵsann,MC\(ϵ\|𝒞x,t\)−βϵ\\displaystyle=\\text\{argmax\}\_\{\\epsilon\}s\_\{ann,MC\}\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)\-\\beta\\epsilon⟹sann\(β\|𝒞x,t\)\\displaystyle\\implies s\_\{ann\}\(\\beta\|\\mathcal\{C\}\_\{x,t\}\)=−βfann\(β\|𝒞x,t\)\+βϵann\(β\|𝒞x,t\)=sann,MC\(ϵann\(β\)\|𝒞x,t\)\\displaystyle=\-\\beta f\_\{ann\}\(\\beta\|\\mathcal\{C\}\_\{x,t\}\)\+\\beta\\epsilon\_\{ann\}\(\\beta\|\\mathcal\{C\}\_\{x,t\}\)=s\_\{ann,MC\}\(\\epsilon\_\{ann\}\(\\beta\)\|\\mathcal\{C\}\_\{x,t\}\)
Note the difference between the annealed and actual \(quenched\) energy in the thermodynamic limit of the canonical ensemble
ϵann\(β\)\\displaystyle\\epsilon\_\{ann\}\(\\beta\)=argmaxϵ\\displaystyle=\\text\{argmax\}\_\{\\epsilon\}sann,MC\(ϵ\|𝒞x,t\)−βϵ\\displaystyle s\_\{ann,MC\}\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)\-\\beta\\epsilonϵ\(β\)\\displaystyle\\epsilon\(\\beta\)=argmaxsann,MC\(ϵ\)\>0\\displaystyle=\\text\{argmax\}\_\{s\_\{ann,MC\}\(\\epsilon\)\>0\}sann,MC\(ϵ\|𝒞x,t\)−βϵ\\displaystyle s\_\{ann,MC\}\(\\epsilon\|\\mathcal\{C\}\_\{x,t\}\)\-\\beta\\epsilonIfϵann\(β\)=ϵ\(β\)\\epsilon\_\{ann\}\(\\beta\)=\\epsilon\(\\beta\), thensann\(ϵann\(β\)\)=sann\(ϵ\(β\)\)≥0s\_\{ann\}\(\\epsilon\_\{ann\}\(\\beta\)\)=s\_\{ann\}\(\\epsilon\(\\beta\)\)\\geq 0\. On the other hand, ifϵann\(β\)≠ϵ\(β\)\\epsilon\_\{ann\}\(\\beta\)\\neq\\epsilon\(\\beta\)then we can concludesann\(ϵann\(β\)\)<0s\_\{ann\}\(\\epsilon\_\{ann\}\(\\beta\)\)<0since noϵ\\epsilonwithsann\(ϵ\)≥0s\_\{ann\}\(\\epsilon\)\\geq 0was able to achieve this minimum\. We now must apply more of our hypothesizes to determine what happens tosann,MC\(ϵ\(β\)\)s\_\{ann,MC\}\(\\epsilon\(\\beta\)\)when actual \(quenched\) and annealed energies disagree\.
By assumption,I\(ϵ\)I\(\\epsilon\)andI′\(ϵ\)I^\{\\prime\}\(\\epsilon\)have a finite number of local maxima and minima that must lie between someEaE\_\{a\}andEbE\_\{b\}\. This means all local minima and maxima ofsann,MC\(ϵ\)s\_\{ann,MC\}\(\\epsilon\)andsann,MC′\(ϵ\)s\_\{ann,MC\}^\{\\prime\}\(\\epsilon\)must lie betweenEaln\|𝒟\|\\frac\{E\_\{a\}\}\{\\ln\|\\mathcal\{D\}\|\}andEbln\|𝒟\|\\frac\{E\_\{b\}\}\{\\ln\|\\mathcal\{D\}\|\}\. Outside this sub\-extensive \(O\(1ln\|𝒟\|\)O\(\\frac\{1\}\{\\ln\|\\mathcal\{D\}\|\}\)\) region, the derivative of entropy has a constant sign\. SinceP\(ln\|𝒟\|ϵ\)=e−ln\|𝒟\|I\(ϵ\)P\(\\ln\|\\mathcal\{D\}\|\\epsilon\)=e^\{\-\\ln\|\\mathcal\{D\}\|I\(\\epsilon\)\}and normalization implieslimE→−∞P\(E\)=0\\lim\_\{E\\rightarrow\-\\infty\}P\(E\)=0,limϵ→−∞I\(ϵ\)=∞\\lim\_\{\\epsilon\\rightarrow\-\\infty\}I\(\\epsilon\)=\\infty\. Therefore,I′\(ϵ\)I^\{\\prime\}\(\\epsilon\)must be negative andsann,MC′\(ϵ\)s\_\{ann,MC\}^\{\\prime\}\(\\epsilon\)must be positive belowEaln\|𝒟\|\\frac\{E\_\{a\}\}\{\\ln\|\\mathcal\{D\}\|\}\. By similar logic, we conclude thatsann,MC′\(ϵ\)s\_\{ann,MC\}^\{\\prime\}\(\\epsilon\)is negative aboveEbln\|𝒟\|\\frac\{E\_\{b\}\}\{\\ln\|\\mathcal\{D\}\|\}\. Sinceβ\>0\\beta\>0andϵ=O\(1\)\\epsilon=O\(1\), we conclude thatϵann\(β\)<Ealn\|𝒟\|\\epsilon\_\{ann\}\(\\beta\)<\\frac\{E\_\{a\}\}\{\\ln\|\\mathcal\{D\}\|\}andϵ\(β\)<Ealn\|𝒟\|\\epsilon\(\\beta\)<\\frac\{E\_\{a\}\}\{\\ln\|\\mathcal\{D\}\|\}for all beta\. Since the derivative is monotonic in this region, there can only be one solution to the equation∂sann,MC∂ϵ=β\\frac\{\\partial s\_\{ann,MC\}\}\{\\partial\\epsilon\}=\\betawhichϵann\\epsilon\_\{ann\}must satify\. Therefore, ifϵ\(β\)≠ϵann\(β\)\\epsilon\(\\beta\)\\neq\\epsilon\_\{ann\}\(\\beta\), then∂sann,MC∂ϵ\(ϵ\(β\)\)≠β\\frac\{\\partial s\_\{ann,MC\}\}\{\\partial\\epsilon\}\(\\epsilon\(\\beta\)\)\\neq\\beta\. This impliesϵ\(β\)\\epsilon\(\\beta\)cannot lie at a local maximum of the optimization problem and therefore must lie on the boundary wheresann,MC\(ϵ\(β\)\)=0s\_\{ann,MC\}\(\\epsilon\(\\beta\)\)=0\. So, we see
ϵ\(β\)≠ϵann\(β\)⟹sann,MC\(ϵ\(β\)\)=0\\epsilon\(\\beta\)\\neq\\epsilon\_\{ann\}\(\\beta\)\\implies s\_\{ann,MC\}\(\\epsilon\(\\beta\)\)=0We can combine these results with equation[89](https://arxiv.org/html/2607.08041#A4.E89)to get the entropy\.
s\(β\|𝒞x,t\)=\{sann,MC\(ϵann\(β\|𝒞x,t\)\|𝒞x,t\)sann,MC\(ϵann\(β\|𝒞x,t\)\|𝒞x,t\)\>00sann,MC\(ϵann\(β\|𝒞x,t\)\|𝒞x,t\)≤0s\(\\beta\|\\mathcal\{C\}\_\{x,t\}\)=\\begin\{cases\}s\_\{ann,MC\}\(\\epsilon\_\{ann\}\(\\beta\|\\mathcal\{C\}\_\{x,t\}\)\|\\mathcal\{C\}\_\{x,t\}\)&s\_\{ann,MC\}\(\\epsilon\_\{ann\}\(\\beta\|\\mathcal\{C\}\_\{x,t\}\)\|\\mathcal\{C\}\_\{x,t\}\)\>0\\\\ 0&s\_\{ann,MC\}\(\\epsilon\_\{ann\}\(\\beta\|\\mathcal\{C\}\_\{x,t\}\)\|\\mathcal\{C\}\_\{x,t\}\)\\leq 0\\end\{cases\}\(94\)So, we can see this thermodynamic model has two different phases\. One where the Gibbs distribution is spread over many states andsann,MC\(ϵann\(β\)\|𝒞x,t\)\>0s\_\{ann,MC\}\(\\epsilon\_\{ann\}\(\\beta\)\|\\mathcal\{C\}\_\{x,t\}\)\>0and one where all the probability concentrates on a single state andsann,MC\(ϵann\(β\)\|𝒞x,t\)≤0s\_\{ann,MC\}\(\\epsilon\_\{ann\}\(\\beta\)\|\\mathcal\{C\}\_\{x,t\}\)\\leq 0\. These correspond exactly to the generalizing and the memorizing phase described in the main text\. Note that simply calculatingsann,MC\(ϵ\(β\)\|𝒞x,t\)s\_\{ann,MC\}\(\\epsilon\(\\beta\)\|\\mathcal\{C\}\_\{x,t\}\)for a given temperature and𝒞x,t\\mathcal\{C\}\_\{x,t\}is enough to determine which phase the model is in\. Work previously done by\[[6](https://arxiv.org/html/2607.08041#bib.bib6)\]explicitly solved fors\(ϵ\)s\(\\epsilon\)andϵ\(β\)\\epsilon\(\\beta\)in the case of isotropic Gaussian data\.
However, stepping back from these saddle\-point equations and examining the context of the problem allows us to work in more generality\. One can generically write the canonical annealed canonical entropy in terms of the temperature using the annealed free energy and annealed energy
Sann\(β\|𝒞x,t\)\\displaystyle S\_\{ann\}\(\\beta\|\\mathcal\{C\}\_\{x,t\}\)=−βFann\(β\)\+βEann\(β\)=ln𝔼\[Z\|𝒞x,t\]\+β𝔼\[EZ\|𝒞x,t\]𝔼\[Z\|𝒞x,t\]\\displaystyle=\-\\beta F\_\{ann\}\(\\beta\)\+\\beta E\_\{ann\}\(\\beta\)=\\ln\\mathbb\{E\}\[Z\|\\mathcal\{C\}\_\{x,t\}\]\+\\beta\\frac\{\\mathbb\{E\}\[EZ\|\\mathcal\{C\}\_\{x,t\}\]\}\{\\mathbb\{E\}\[Z\|\\mathcal\{C\}\_\{x,t\}\]\}\(95\)=ln\|𝒟\|\+ln𝔼E\[e−βE\|𝒞x,t\]\+β𝔼E\[Ee−βE\|𝒞x,t\]𝔼E\[e−βE\|𝒞x,t\]\\displaystyle=\\ln\|\\mathcal\{D\}\|\+\\ln\\mathbb\{E\}\_\{E\}\[e^\{\-\\beta E\}\|\\mathcal\{C\}\_\{x,t\}\]\+\\beta\\frac\{\\mathbb\{E\}\_\{E\}\[Ee^\{\-\\beta E\}\|\\mathcal\{C\}\_\{x,t\}\]\}\{\\mathbb\{E\}\_\{E\}\[e^\{\-\\beta E\}\|\\mathcal\{C\}\_\{x,t\}\]\}\(96\)Usually, evaluating this expression would require knowing the full moment generating function for theEEdistribution, which would be as difficult as solving fors\(ϵ\)s\(\\epsilon\)\. However, we are not interested in the behavior of our model at arbitrary temperatures\. Recall that equation[63](https://arxiv.org/html/2607.08041#A4.E63)defines a specific temperature for the model\.
𝔼E\[e−βE\|𝒞x,t\]\\displaystyle\\mathbb\{E\}\_\{E\}\[e^\{\-\\beta E\}\|\\mathcal\{C\}\_\{x,t\}\]=∫𝑑φP0\(φ\)e−βE\(φ\|𝒞x,t\)\\displaystyle=\\int d\\varphi\\\>P\_\{0\}\(\\varphi\)e^\{\-\\beta E\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\}\(97\)=∫𝑑φP0\(φ\)P\(𝒞x,t\|φ\)=Ptest\(𝒞x,t\)\\displaystyle=\\int d\\varphi\\\>P\_\{0\}\(\\varphi\)P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\)=P\_\{test\}\(\\mathcal\{C\}\_\{x,t\}\)\(98\)So, we see that at the temperature given by[63](https://arxiv.org/html/2607.08041#A4.E63), the moment generating function of the energy distribution is just the test set probability of𝒞x,t\\mathcal\{C\}\_\{x,t\}\. Similarly, for the energy expectation, we find
β𝔼E\[Ee−βE\|𝒞x,t\]𝔼E\[e−βE\|𝒞x,t\]\\displaystyle\\frac\{\\beta\\mathbb\{E\}\_\{E\}\[Ee^\{\-\\beta E\}\|\\mathcal\{C\}\_\{x,t\}\]\}\{\\mathbb\{E\}\_\{E\}\[e^\{\-\\beta E\}\|\\mathcal\{C\}\_\{x,t\}\]\}=∫𝑑φP0\(φ\)Ptest\(𝒞x,t\)βE\(φ\|𝒞x,t\)e−βE\(φ\|𝒞x,t\)\\displaystyle=\\int d\\varphi\\\>\\frac\{P\_\{0\}\(\\varphi\)\}\{P\_\{test\}\(\\mathcal\{C\}\_\{x,t\}\)\}\\beta E\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)e^\{\-\\beta E\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\}\(99\)=∫𝑑φP0\(φ\)Ptest\(𝒞x,t\)P\(𝒞x,t\|φ\)\(−lnP\(𝒞x,t\|φ\)\)\\displaystyle=\\int d\\varphi\\\>\\frac\{P\_\{0\}\(\\varphi\)\}\{P\_\{test\}\(\\mathcal\{C\}\_\{x,t\}\)\}P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\)\\\>\(\-\\ln P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\)\)\(100\)=∫𝑑φPtest\(φ\|𝒞x,t\)\(−lnP\(𝒞x,t\|φ\)\)\\displaystyle=\\int d\\varphi\\\>P\_\{test\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\\\>\(\-\\ln P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\)\)\(101\)Combining everything, we find
ln\|𝒟\|sann,MC\(ϵann\(β\)\|𝒞x,t\)\\displaystyle\\ln\|\\mathcal\{D\}\|\\\>\\\>s\_\{ann,MC\}\(\\epsilon\_\{ann\}\(\\beta\)\|\\mathcal\{C\}\_\{x,t\}\)=Sann\(β\|𝒞x,t\)\\displaystyle=S\_\{ann\}\(\\beta\|\\mathcal\{C\}\_\{x,t\}\)=ln\|𝒟\|\+lnPtest\(𝒞x,t\)\+∫𝑑φPtest\(φ\|𝒞x,t\)\(−lnP\(𝒞x,t\|φ\)\)\\displaystyle=\\ln\|\\mathcal\{D\}\|\+\\ln P\_\{test\}\(\\mathcal\{C\}\_\{x,t\}\)\+\\int d\\varphi\\\>P\_\{test\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\\\>\(\-\\ln P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\)\)=ln\|𝒟\|\+∫𝑑φPtest\(φ\|𝒞x,t\)\(−lnPtest\(φ\|𝒞x,t\)P0\(φ\)\)\\displaystyle=\\ln\|\\mathcal\{D\}\|\+\\int d\\varphi\\\>P\_\{test\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\\\>\(\-\\ln\\frac\{P\_\{test\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\}\{P\_\{0\}\(\\varphi\)\}\)=ln\|𝒟\|−DKL\(Ptest\(φ\|𝒞x,t\)\|\|P0\(φ\)\)\\displaystyle=\\ln\|\\mathcal\{D\}\|\-D\_\{KL\}\(P\_\{test\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\|\|P\_\{0\}\(\\varphi\)\)Plugging this into the formula for the actual \(quenched\) entropy,
S\(β\|𝒞x,t\)=max\(0,ln\|𝒟\|−DKL\(Ptest\(φ\|𝒞x,t\)\|\|P0\(φ\)\)\)S\(\\beta\|\\mathcal\{C\}\_\{x,t\}\)=\\operatorname\{max\}\\Big\(0,\\ln\|\\mathcal\{D\}\|\-D\_\{KL\}\(P\_\{test\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\|\|P\_\{0\}\(\\varphi\)\)\\Big\)\(102\)From this, we can conclude the pointwise collapse condition\.
This pointwise condition is useful for situations where𝒞x,t\\mathcal\{C\}\_\{x,t\}is not drawn from a known distribution\. So, long as it is drawn independently from𝒟\\mathcal\{D\}, we can calculate the posterior entropy\. If𝒞x,t\\mathcal\{C\}\_\{x,t\}is drawn from the test set, the expression further simplifies:
###### Theorem D\.2\(Test Set Averaged Posterior Entropy\)
If𝒟\\mathcal\{D\}is a set of training images drawn i\.i\.d\. fromP0P\_\{0\}and𝒞\\mathcal\{C\}is a channel such thatP\(E\|𝒞x,t\)P\(E\|\\mathcal\{C\}\_\{x,t\}\)has sub\-exponential tails andVar\(DKL\(Ptest\(φ\|𝒞x,t\)\|\|P0\(φ\)\)\)=o\(1\)\\sqrt\{Var\(D\_\{KL\}\(P\_\{test\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\|\|P\_\{0\}\(\\varphi\)\)\)\}=o\(1\), then when a noisy image is drawn from the test set
𝔼𝒞x,t∼Ptest\[S\[Ptrain\(φ\|𝒞x,t\)\]\]=\{ln\|𝒟\|−I\(φ;𝒞x,t\)ln\|𝒟\|\>I\(φ;𝒞x,t\)0ln\|𝒟\|≤I\(φ;𝒞x,t\)\\mathbb\{E\}\_\{\\mathcal\{C\}\_\{x,t\}\\sim P\_\{test\}\}\[S\[P\_\{train\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\]\]=\\begin\{cases\}\\ln\|\\mathcal\{D\}\|\-I\(\\varphi;\\mathcal\{C\}\_\{x,t\}\)&\\ln\|\\mathcal\{D\}\|\>I\(\\varphi;\\mathcal\{C\}\_\{x,t\}\)\\\\ 0&\\ln\|\\mathcal\{D\}\|\\leq I\(\\varphi;\\mathcal\{C\}\_\{x,t\}\)\\end\{cases\}\(103\)up too\(1\)o\(1\)corrections\.
###### Proof\.
Since the annealed entropy is assumed to have sub extensive standard deviation, we can replace it with its mean:
𝔼𝒞x,t∼Ptest\[Sann\(φ\|𝒞x,t\)\]\\displaystyle\\mathbb\{E\}\_\{\\mathcal\{C\}\_\{x,t\}\\sim P\_\{test\}\}\[S\_\{ann\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\]=ln\|𝒟\|−𝔼𝒞x,t∼Ptest\[DKL\(Ptest\(φ\|𝒞x,t\)\|\|P0\(φ\)\)\]\\displaystyle=\\ln\|\\mathcal\{D\}\|\-\\mathbb\{E\}\_\{\\mathcal\{C\}\_\{x,t\}\\sim P\_\{test\}\}\[D\_\{KL\}\(P\_\{test\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\|\|P\_\{0\}\(\\varphi\)\)\]\(104\)=ln\|𝒟\|\+∫𝑑φ𝑑𝒞x,tPtest\(Cx,t\)P\(φ\|𝒞x,t\)\(−lnPtest\(φ\|𝒞x,t\)P0\(φ\)\)\\displaystyle\\hskip\-60\.0pt=\\ln\|\\mathcal\{D\}\|\+\\int d\\varphi d\\mathcal\{C\}\_\{x,t\}\\\>P\_\{test\}\(C\_\{x,t\}\)P\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\\\>\(\-\\ln\\frac\{P\_\{test\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\}\{P\_\{0\}\(\\varphi\)\}\)\(105\)=ln\|𝒟\|\+∫𝑑φ𝑑𝒞x,tPtest\(φ,𝒞x,t\)\(−lnPtest\(φ,𝒞x,t\)Ptest\(𝒞x,t\)P0\(φ\)\)\\displaystyle\\hskip\-60\.0pt=\\ln\|\\mathcal\{D\}\|\+\\int d\\varphi d\\mathcal\{C\}\_\{x,t\}\\\>P\_\{test\}\(\\varphi,\\mathcal\{C\}\_\{x,t\}\)\\\>\(\-\\ln\\frac\{P\_\{test\}\(\\varphi,\\mathcal\{C\}\_\{x,t\}\)\}\{P\_\{test\}\(\\mathcal\{C\}\_\{x,t\}\)P\_\{0\}\(\\varphi\)\}\)\(106\)=ln\|𝒟\|−I\(φ;𝒞x,t\)\\displaystyle\\hskip\-60\.0pt=\\ln\|\\mathcal\{D\}\|\-I\(\\varphi;\\mathcal\{C\}\_\{x,t\}\)\(107\)
###### Theorem D\.3\(Training Set Posterior Entropy\)
If𝒟/φ0\\mathcal\{D\}/\\varphi\_\{0\}is a set of training images drawn i\.i\.d\. fromP0P\_\{0\}and𝒞\\mathcal\{C\}is a channel such thatP\(E\|𝒞x,t\)P\(E\|\\mathcal\{C\}\_\{x,t\}\)satisfies the conditions in theorem[D\.1](https://arxiv.org/html/2607.08041#A4.Thmdefinition1),Var\(DKL\(Ptest\(φ\|𝒞x,t\)\|\|P0\(φ\)\)\)=o\(1\)\\sqrt\{Var\(D\_\{KL\}\(P\_\{test\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\|\|P\_\{0\}\(\\varphi\)\)\)\}=o\(1\), andVar\(lnPtest\(𝒞x,t,φ\)Ptest\(𝒞x,t\)P0\(φ\)\)=o\(1\)Var\(\\ln\\frac\{P\_\{test\}\(\\mathcal\{C\}\_\{x,t\},\\varphi\)\}\{P\_\{test\}\(\\mathcal\{C\}\_\{x,t\}\)P\_\{0\}\(\\varphi\)\}\)=o\(1\), then when a noisy image is drawn from the train set
𝔼𝒞x,t∼Ptrain\[S\[Ptrain\(φ\|𝒞x,t\)\]\]=\{ln\|𝒟\|−I\(φ;𝒞x,t\)ln\|𝒟\|\>I\(φ;𝒞x,t\)0ln\|𝒟\|≤I\(φ;𝒞x,t\)\\mathbb\{E\}\_\{\\mathcal\{C\}\_\{x,t\}\\sim P\_\{train\}\}\[S\[P\_\{train\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\]\]=\\begin\{cases\}\\ln\|\\mathcal\{D\}\|\-I\(\\varphi;\\mathcal\{C\}\_\{x,t\}\)&\\ln\|\\mathcal\{D\}\|\>I\(\\varphi;\\mathcal\{C\}\_\{x,t\}\)\\\\ 0&\\ln\|\\mathcal\{D\}\|\\leq I\(\\varphi;\\mathcal\{C\}\_\{x,t\}\)\\end\{cases\}\(108\)up too\(1\)o\(1\)corrections\.
###### Proof\.
If𝒞x,t\\mathcal\{C\}\_\{x,t\}is not drawn independently from𝒟\\mathcal\{D\}, for example if𝒞x,t\\mathcal\{C\}\_\{x,t\}is generated from the training set, there is less that we can say\. This violates the i\.i\.d\. assumption imposed on the𝒟\\mathcal\{D\}training set in theorem[D\.1](https://arxiv.org/html/2607.08041#A4.Thmdefinition1)\. A training imageφ0∈𝒟\\varphi\_\{0\}\\in\\mathcal\{D\}is first chosen uniformly and then𝒞x,t\\mathcal\{C\}\_\{x,t\}is drawn fromP\(𝒞x,t\|φ\)P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\)\. This leads to additional contribution to the annealed partition function
Ztrain=P\(𝒞x,t\|φ0\)\+∑φ∈𝒟/\{φ0\}P\(𝒞x,t\|φ\)Z\_\{train\}=P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\_\{0\}\)\+\\sum\_\{\\varphi\\in\\mathcal\{D\}/\\\{\\varphi\_\{0\}\\\}\}P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\)\(109\)Each term in the second sum is identically distributed and their contribution from the second term is identical to the test partition function with\|𝒟\|−1\|\\mathcal\{D\}\|\-1samples\. However, thisP\(𝒞x,t\|φ0\)P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\_\{0\}\)term must be specifically accounted for\. In the error correction context, this quantity is usually taken to be much smaller thanln\|𝒟\|\\ln\|\\mathcal\{D\}\|and therefore ignored\. If we don’t ignore it, it can lead to the collapse transition happening faster\.
Ztrain\\displaystyle Z\_\{train\}=P\(𝒞x,t\|φ0\)\+max\(\(\|𝒟\|−1\)Ptest\(𝒞x,t\),e−ln\|𝒟\|βϵ∗\)\\displaystyle=P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\_\{0\}\)\+\\operatorname\{max\}\(\(\|\\mathcal\{D\}\|\-1\)P\_\{test\}\(\\mathcal\{C\}\_\{x,t\}\),e^\{\-\\ln\|\\mathcal\{D\}\|\\beta\\epsilon^\{\*\}\}\)\(110\)whereϵ∗\\epsilon^\{\*\}is the intensive energy such thatI\(ϵ∗\|𝒞x,t\)=1I\(\\epsilon^\{\*\}\|\\mathcal\{C\}\_\{x,t\}\)=1\. We must assume thatlnP\(𝒞x,t\|φ0\)/ln\|𝒟\|\\ln P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\_\{0\}\)/\\ln\|\\mathcal\{D\}\|concentrates to some deterministicO\(1\)O\(1\)value whenφ0\\varphi\_\{0\}is drawn fromP0P\_\{0\}and𝒞x,t\\mathcal\{C\}\_\{x,t\}fromP\(⋅\|ϕ0\)P\(\\cdot\|\\phi\_\{0\}\)\. Because the two terms are exponentially small inln\|𝒟\|\\ln\|\\mathcal\{D\}\|,ZtrainZ\_\{train\}is always only dominated by one of the two terms\. IflnP\(𝒞x,t\|φ0\)βln\|𝒟\|\>ϵ∗\\frac\{\\ln P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\_\{0\}\)\}\{\\beta\\ln\|\\mathcal\{D\}\|\}\>\\epsilon^\{\*\}, then there are exponentially many other states around the energylnP\(𝒞x,t\|φ0\)βln\|𝒟\|\\frac\{\\ln P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\_\{0\}\)\}\{\\beta\\ln\|\\mathcal\{D\}\|\}andZtrain=Ztest\+O\(1\)Z\_\{train\}=Z\_\{test\}\+O\(1\)we conclude the theorem by applying theorem[D\.2](https://arxiv.org/html/2607.08041#A4.Thmdefinition2)\. IflnP\(𝒞x,t\|φ0\)βln\|𝒟\|<ϵ∗\\frac\{\\ln P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\_\{0\}\)\}\{\\beta\\ln\|\\mathcal\{D\}\|\}<\\epsilon^\{\*\}then,
−βF\(β\)=lnZ=\{ln\(\|𝒟\|−1\)\+lnPtest\(𝒞x,t\)ln\(\|𝒟\|−1\)\+lnPtest\(𝒞x,t\)\>lnP\(𝒞x,t\|φ0\)lnP\(𝒞x,t\|φ0\)ln\(\|𝒟\|−1\)\+lnPtest\(𝒞x,t\)≤lnP\(𝒞x,t\|φ0\)\-\\beta F\(\\beta\)=\\ln Z=\\begin\{cases\}\\ln\(\|\\mathcal\{D\}\|\-1\)\+\\ln P\_\{test\}\(\\mathcal\{C\}\_\{x,t\}\)&\\ln\(\|\\mathcal\{D\}\|\-1\)\+\\ln P\_\{test\}\(\\mathcal\{C\}\_\{x,t\}\)\>\\ln P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\_\{0\}\)\\\\ \\ln P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\_\{0\}\)&\\ln\(\|\\mathcal\{D\}\|\-1\)\+\\ln P\_\{test\}\(\\mathcal\{C\}\_\{x,t\}\)\\leq\\ln P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\_\{0\}\)\\end\{cases\}In the collapsed phase, we know the energy associated with the training sample that generated𝒞x,t\\mathcal\{C\}\_\{x,t\}islnP\(𝒞x,t\|φ0\)βln\|𝒟\|\\frac\{\\ln P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\_\{0\}\)\}\{\\beta\\ln\|\\mathcal\{D\}\|\}\. In the generalizing phase, we can use the conclusion of theorem[D\.1](https://arxiv.org/html/2607.08041#A4.Thmdefinition1)sinceZtrain=Ztest\+O\(e−ln\|D\|\)Z\_\{train\}=Z\_\{test\}\+O\(e^\{\-\\ln\|D\|\}\)\.
S\(β\)=\{ln\(\|𝒟\|−1\)−DKL\(Ptest\(φ\|𝒞x,t\)\|\|P0\(φ\)\)ln\(\|𝒟\|−1\)\+lnPtest\(𝒞x,t\)\>lnP\(𝒞x,t\|φ0\)0ln\(\|𝒟\|−1\)\+lnPtest\(𝒞x,t\)≤lnP\(𝒞x,t\|φ0\)S\(\\beta\)=\\begin\{cases\}\\ln\(\|\\mathcal\{D\}\|\-1\)\-D\_\{KL\}\(P\_\{test\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\|\|P\_\{0\}\(\\varphi\)\)&\\ln\(\|\\mathcal\{D\}\|\-1\)\+\\ln P\_\{test\}\(\\mathcal\{C\}\_\{x,t\}\)\>\\ln P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\_\{0\}\)\\\\ 0&\\ln\(\|\\mathcal\{D\}\|\-1\)\+\\ln P\_\{test\}\(\\mathcal\{C\}\_\{x,t\}\)\\leq\\ln P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\_\{0\}\)\\end\{cases\}Now using Bayes ruleP\(𝒞x,t\|φ0\)=Ptest\(𝒞x,t,φ0\)P0\(φ0\)P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\_\{0\}\)=\\frac\{P\_\{test\}\(\\mathcal\{C\}\_\{x,t\},\\varphi\_\{0\}\)\}\{P\_\{0\}\(\\varphi\_\{0\}\)\}
S\[Ptrain\(φ\|𝒞x,t\)\]=\{ln\(\|𝒟\|−1\)−DKL\(Ptest\(φ\|𝒞x,t\)\|\|P0\(φ\)\)ln\(\|𝒟\|−1\)\>lnPtest\(𝒞x,t,φ0\)Ptest\(𝒞x,t\)P0\(φ0\)0ln\(\|𝒟\|−1\)≤lnPtest\(𝒞x,t,φ0\)Ptest\(𝒞x,t\)P0\(φ0\)S\[P\_\{train\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\]=\\begin\{cases\}\\ln\(\|\\mathcal\{D\}\|\-1\)\-D\_\{KL\}\(P\_\{test\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\|\|P\_\{0\}\(\\varphi\)\)&\\ln\(\|\\mathcal\{D\}\|\-1\)\>\\ln\\frac\{P\_\{test\}\(\\mathcal\{C\}\_\{x,t\},\\varphi\_\{0\}\)\}\{P\_\{test\}\(\\mathcal\{C\}\_\{x,t\}\)P\_\{0\}\(\\varphi\_\{0\}\)\}\\\\ 0&\\ln\(\|\\mathcal\{D\}\|\-1\)\\leq\\ln\\frac\{P\_\{test\}\(\\mathcal\{C\}\_\{x,t\},\\varphi\_\{0\}\)\}\{P\_\{test\}\(\\mathcal\{C\}\_\{x,t\}\)P\_\{0\}\(\\varphi\_\{0\}\)\}\\end\{cases\}Now we note thatlnPtest\(𝒞x,t,φ0\)Ptest\(𝒞x,t\)P0\(φ\)\\ln\\frac\{P\_\{test\}\(\\mathcal\{C\}\_\{x,t\},\\varphi\_\{0\}\)\}\{P\_\{test\}\(\\mathcal\{C\}\_\{x,t\}\)P\_\{0\}\(\\varphi\)\}is the observable averaged in the expression for the mutual information\.
I\(𝒞x,t,φ0\)=∫𝑑φ𝑑𝒞x,tPtest\(𝒞x,t,φ0\)lnPtest\(𝒞x,t,φ0\)Ptest\(𝒞x,t\)P0\(φ\)I\(\\mathcal\{C\}\_\{x,t\},\\varphi\_\{0\}\)=\\int d\\varphi d\\mathcal\{C\}\_\{x,t\}\\\>P\_\{test\}\(\\mathcal\{C\}\_\{x,t\},\\varphi\_\{0\}\)\\ln\\frac\{P\_\{test\}\(\\mathcal\{C\}\_\{x,t\},\\varphi\_\{0\}\)\}\{P\_\{test\}\(\\mathcal\{C\}\_\{x,t\}\)P\_\{0\}\(\\varphi\)\}\(112\)Therefore, if we assume thatlnPtest\(𝒞x,t,φ0\)Ptest\(𝒞x,t\)P0\(φ0\)\\ln\\frac\{P\_\{test\}\(\\mathcal\{C\}\_\{x,t\},\\varphi\_\{0\}\)\}\{P\_\{test\}\(\\mathcal\{C\}\_\{x,t\}\)P\_\{0\}\(\\varphi\_\{0\}\)\}also concentrates, then
S\[Ptrain\(φ\|𝒞x,t\)\]=\{ln\(\|𝒟\|−1\)−DKL\(Ptest\(φ\|𝒞x,t\)\|\|P0\(φ\)\)ln\(\|𝒟\|−1\)\>I\(𝒞x,t;φ\)0ln\(\|𝒟\|−1\)≤I\(𝒞x,t;φ\)S\[P\_\{train\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\]=\\begin\{cases\}\\ln\(\|\\mathcal\{D\}\|\-1\)\-D\_\{KL\}\(P\_\{test\}\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\|\|P\_\{0\}\(\\varphi\)\)&\\ln\(\|\\mathcal\{D\}\|\-1\)\>I\(\\mathcal\{C\}\_\{x,t\};\\varphi\)\\\\ 0&\\ln\(\|\\mathcal\{D\}\|\-1\)\\leq I\(\\mathcal\{C\}\_\{x,t\};\\varphi\)\\end\{cases\}\(113\)So, we see that the posterior entropy for a noisy training image is essentially the same when appropriate strong assumptions are made aboutP\(𝒞x,t\|φ\)P\(\\mathcal\{C\}\_\{x,t\}\|\\varphi\)\.
### D\.6Relationship to Biroli et al\.’s Collapse Condition
In\[[6](https://arxiv.org/html/2607.08041#bib.bib6)\]the memorization to generalization phase transition for the empirical score function was studied using the random energy model\. Our work with the random energy model in section[D\.5](https://arxiv.org/html/2607.08041#A4.SS5)was inspired by theirs\. The empirical score function is, in some sense, the simplest BIRD model in which all pixels see the entire noisy image\. The Markov chain for the observation is just a white noise channel
φ→ϕ:=α¯tφ\+1−α¯tη\\varphi\\rightarrow\\phi:=\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\varphi\+\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\}\}\\eta\(114\)They were able to solve for the entropy of the generated distribution and memorization time in the large dimension limit exactly in the case of isotropic Gaussian data by directly calculating the micro\-canonical entropy\. They conjectured that for the memorization timet∗t^\{\*\}, the identity
S\[P\(ϕt∗\)\]=Ssep\|𝒟\|S\[P\(\\phi\_\{t^\{\*\}\}\)\]=S\_\{sep\}^\{\|\\mathcal\{D\}\|\}\(115\)would hold for dataφ\\varphifrom arbitrary distributions, whereSsep\|𝒟\|S\_\{sep\}^\{\|\\mathcal\{D\}\|\}is the entropy of\|𝒟\|\|\\mathcal\{D\}\|well\-separated isotropic Gaussians with variance\(1−α¯t\)I\(1\-\\bar\{\\alpha\}\_\{t\}\)I:
Ssep\|𝒟\|=ln\|𝒟\|\+S\[𝒩\(η\|0,\(1−α¯t\)I\)\]=ln\|𝒟\|\+d2\(1\+ln\(2π\(1−α¯t\)\)\)S\_\{sep\}^\{\|\\mathcal\{D\}\|\}=\\ln\|\\mathcal\{D\}\|\+S\[\\mathcal\{N\}\(\\eta\|0,\(1\-\\bar\{\\alpha\}\_\{t\}\)I\)\]=\\ln\|\\mathcal\{D\}\|\+\\frac\{d\}\{2\}\(1\+\\ln\(2\\pi\(1\-\\bar\{\\alpha\}\_\{t\}\)\)\)\(116\)Using the mutual information formula for the white noise channel with this parametrization \(theorem[E\.1](https://arxiv.org/html/2607.08041#A5.Thmdefinition1)\)
ln\|D\|=S\[P\(ϕt∗\)\]−S\[𝒩\(η\|0,\(1−α¯t\)I\)\]=I\(ϕt∗;φ\)\\ln\|D\|=S\[P\(\\phi\_\{t^\{\*\}\}\)\]\-S\[\\mathcal\{N\}\(\\eta\|0,\(1\-\\bar\{\\alpha\}\_\{t\}\)I\)\]=I\(\\phi\_\{t^\{\*\}\};\\varphi\)\(117\)which corresponds to the collapse condition \([5](https://arxiv.org/html/2607.08041#S4.E5)\) for the trivial channel𝒞x,t\(ϕt\)=ϕt\\mathcal\{C\}\_\{x,t\}\(\\phi\_\{t\}\)=\\phi\_\{t\}\. Thus, our analysis proves Biroli’s conjecture using random energy model methods, but also significantly generalizes their notions to cases where the channel𝒞x,t\\mathcal\{C\}\_\{x,t\}is potentially nonlinear or stochastic, for which their conjectured identity \([115](https://arxiv.org/html/2607.08041#A4.E115)\) no longer serves as the correct collapse condition\.
## Appendix EUseful Mutual Information Bounds and Calculations
### E\.1Memorization implies a lower bound on mutual information
Here we will use Fano’s inequality to prove a simple lower bound on the mutual informationI\(𝒞x,t;φ\)I\(\\mathcal\{C\}\_\{x,t\};\\varphi\)whenever a BIRD model is in the memorization phase\.
Consider again the mutual informationI\(𝒞x,t;φ\)I\(\\mathcal\{C\}\_\{x,t\};\\varphi\)between a pixelxx’s current observation𝒞x,t\\mathcal\{C\}\_\{x,t\}at timettand the past training imageφ∈𝒟\\varphi\\in\\mathcal\{D\}at timet=0t=0that led to the observation under the forward diffusion\. In BIRD models, the pixelxxattempts to play a reverse time Bayesian guessing game to guess the correct past training imageφc∈𝒟\\varphi\_\{c\}\\in\\mathcal\{D\}that led to𝒞x,t\\mathcal\{C\}\_\{x,t\}\. Its knowledge about the past training data is encapsulated in the posterior distributionP\(φ\|𝒞x,t\)P\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)forφ∈𝒟\\varphi\\in\\mathcal\{D\}\. The probability the pixel guesses the correct training pointφc\\varphi\_\{c\}isP\(φc\|𝒞x,t\)P\(\\varphi\_\{c\}\|\\mathcal\{C\}\_\{x,t\}\)\. Thus the probability a pixelxxloses the Bayesian guessing game by making an incorrect guess or error ispe=1−P\(φc\|𝒞x,t\)p\_\{e\}=1\-P\(\\varphi\_\{c\}\|\\mathcal\{C\}\_\{x,t\}\)\. The pixelxxhas then memorized the training data, by definition, ifpep\_\{e\}is close to0and the posteriorP\(φ\|𝒞x,t\)P\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)concentrates on the correct data pointφc\\varphi\_\{c\}\.
We can think of this process as an error correction problem where the\|𝒟\|\|\\mathcal\{D\}\|training samplesφ∈𝒟\\varphi\\in\\mathcal\{D\}can be thought of as\|𝒟\|\|\\mathcal\{D\}\|codewords\. One of these codewordsφc\\varphi\_\{c\}is sent through the forward diffusion plus observation communication channelφ→ϕt→𝒞x,t\\varphi\\rightarrow\\phi\_\{t\}\\rightarrow\\mathcal\{C\}\_\{x,t\}, yielding the channel output𝒞x,t\\mathcal\{C\}\_\{x,t\}\. Then the pixelxxmust guess the correct input codewordφc\\varphi\_\{c\}based on this output\. Memorization then corresponds to successful error correction in whichpe→0p\_\{e\}\\rightarrow 0\(which is deleterious though for BIRD models as it leads to the failure to generalize\)\.
With this error correction analogy in place, we can now apply Fano’s inequality \(see\[[46](https://arxiv.org/html/2607.08041#bib.bib46)\]for a standard proof\) which provides a lower bound on the error probabilitypep\_\{e\}in terms of the conditional entropyS\(φ\|𝒞x,t\)S\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)\. Applying Fano’s inequality yields
Hb\(pe\)\+peln\(\|𝒟\|−1\)\\displaystyle H\_\{b\}\(p\_\{e\}\)\+p\_\{e\}\\ln\(\|\\mathcal\{D\}\|\-1\)≥S\(φ\|𝒞x,t\)\\displaystyle\\geq S\(\\varphi\|\\mathcal\{C\}\_\{x,t\}\)=S\(φ\)−I\(Ct,x;φ\)\\displaystyle=S\(\\varphi\)\-I\(C\_\{t,x\};\\varphi\)=ln\|𝒟\|−I\(Ct,x;φ\)\.\\displaystyle=\\ln\|\\mathcal\{D\}\|\-I\(C\_\{t,x\};\\varphi\)\.HereHb\(pe\)H\_\{b\}\(p\_\{e\}\)is the entropy of a binary random variable with probability parameterpep\_\{e\}\. We next take the memorization limitpe→0p\_\{e\}\\rightarrow 0\. This makes the left hand side0, yielding a lower bound on mutual information in the memorization phase:
I\(𝒞t,x;φ\)≥ln\|𝒟\|\.I\(\\mathcal\{C\}\_\{t,x\};\\varphi\)\\geq\\ln\|\\mathcal\{D\}\|\.\(118\)
Importantly, this analysis only proves that in the memorization phase, corresponding tope→0p\_\{e\}\\rightarrow 0, we haveI\(𝒞t,x;φ\)≥ln\|𝒟\|I\(\\mathcal\{C\}\_\{t,x\};\\varphi\)\\geq\\ln\|\\mathcal\{D\}\|\. It doesnotprove that in the generalizing phase, whenpep\_\{e\}is bounded away from0, that the opposite holds:I\(𝒞t,x;φ\)<ln\|𝒟\|I\(\\mathcal\{C\}\_\{t,x\};\\varphi\)<\\ln\|\\mathcal\{D\}\|\. One would ideally like to have this latter inequality also in the generalizing phase to prove that the boundary between the generalizing phase and memorizing phase occurs exactly whenI\(𝒞t,x;φ\)=ln\|𝒟\|I\(\\mathcal\{C\}\_\{t,x\};\\varphi\)=\\ln\|\\mathcal\{D\}\|\. Indeed, we are able to prove this result in App\.[D\.5](https://arxiv.org/html/2607.08041#A4.SS5), when𝒞x,t\\mathcal\{C\}\_\{x,t\}is drawn from the testing Markov process\.
### E\.2The Gaussian Bound on Mutual Information for the Additive White Noise Channel
In the prior sections, we keep the description of Bayes optimal denoisers as generic as possible, but calculating bounds on the channel mutual information requires specificity on the type of channel that we are considering\. In the case of the ideal score function, the local score function, the equivariant local score function, the channels for each agent or pixel can be viewed as the composition of two channels\. First, a non\-invertible, possibly random linear map applied to a noiseless image, and second, rescaling and adding white noise\. For example, in the case of the local score machine,
𝒞x,t:φ→φΩx→α¯tφΩx\+1−α¯tηΩx\\mathcal\{C\}\_\{x,t\}:\\varphi\\rightarrow\\varphi\_\{\\Omega\_\{x\}\}\\rightarrow\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\varphi\_\{\\Omega\_\{x\}\}\+\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\}\}\\eta\_\{\\Omega\_\{x\}\}\(119\)whereη\\etais an isotropic Gaussian\. Note that this is the opposite of the order which these two channels are applied in the actual computation\. We actually apply a projection to the noisy image\. However, because the white noise at each pixel is independent and the projection map is linear we can view the two as happening in either order\.PΩx:ℝH×W×C→ℝk×k×CP\_\{\\Omega\_\{x\}\}:\{\{\\mathbb\{R\}\}\}^\{H\\times W\\times C\}\\rightarrow\{\{\\mathbb\{R\}\}\}^\{k\\times k\\times C\}
𝒞x,t=PΩx\(α¯tφ\+1−α¯tη1\)=α¯tPΩxφ\+1−α¯tη2\\mathcal\{C\}\_\{x,t\}=P\_\{\\Omega\_\{x\}\}\(\\sqrt\{\\bar\{\\alpha\}\}\_\{t\}\\varphi\+\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\}\}\\eta\_\{1\}\)=\\sqrt\{\\bar\{\\alpha\}\}\_\{t\}P\_\{\\Omega\_\{x\}\}\\varphi\+\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\}\}\\eta\_\{2\}\(120\)whereη1∼ℕ\(0,IH×W×C\)\\eta\_\{1\}\\sim\{\{\\mathbb\{N\}\}\}\(0,I\_\{H\\times W\\times C\}\)andη1∼ℕ\(0,Ik×k×C\)\\eta\_\{1\}\\sim\{\{\\mathbb\{N\}\}\}\(0,I\_\{k\\times k\\times C\}\)\. By the data processing inequality,
I\(𝒞x,t;φ\)≤I\(𝒞x,t;PΩxφ\)I\(\\mathcal\{C\}\_\{x,t\};\\varphi\)\\leq I\(\\mathcal\{C\}\_\{x,t\};P\_\{\\Omega\_\{x\}\}\\varphi\)\(121\)In other words, because𝒞x,t\\mathcal\{C\}\_\{x,t\}only depends onφ\\varphithrough its dependence onPΩxφP\_\{\\Omega\_\{x\}\}\\varphi,𝒞x,t\\mathcal\{C\}\_\{x,t\}cannot have more information aboutφ\\varphithan was contained inPΩxφP\_\{\\Omega\_\{x\}\}\\varphi\. For ease of notation in the following sections, I’ll defineΦx\\Phi\_\{x\}as the noise free version of the observation𝒞x,t\\mathcal\{C\}\_\{x,t\}\.
Φ:=PΩxφ\\Phi:=P\_\{\\Omega\_\{x\}\}\\varphi\(122\)
Therefore, we can understand most Bayes optimal diffusion models just by understanding the mutual information of the additive white noise channel\. LetΦ\\Phibe some random variable inserted into an additive white noise channel\.
I\(𝒞x,t:=α¯tΦ\+1−α¯tη;Φ\)\\displaystyle I\(\\mathcal\{C\}\_\{x,t\}:=\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\Phi\+\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\}\}\\eta;\\Phi\)=S\(𝒞x,t\)\+S\(Φ\)−S\(𝒞x,t,Φ\)\\displaystyle=S\(\\mathcal\{C\}\_\{x,t\}\)\+S\(\\Phi\)\-S\(\\mathcal\{C\}\_\{x,t\},\\Phi\)\(123\)=S\(𝒞x,t\)−S\(𝒞x,t\|Φ\)\\displaystyle=S\(\\mathcal\{C\}\_\{x,t\}\)\-S\(\\mathcal\{C\}\_\{x,t\}\|\\Phi\)\(124\)=S\(𝒞x,t\)−S\(α¯tΦ\+1−α¯tη\|Φ\)\\displaystyle=S\(\\mathcal\{C\}\_\{x,t\}\)\-S\(\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\Phi\+\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\}\}\\eta\|\\Phi\)\(125\)After conditioning onΦ\\Phi,α¯tΦ\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\Phiis just a deterministic shift andη\\etais independent\. So, this simplifies to:
I\(𝒞x,t;Φ\)\\displaystyle I\(\\mathcal\{C\}\_\{x,t\};\\Phi\)=S\(𝒞x,t\)−S\(1−α¯tη\)\\displaystyle=S\(\\mathcal\{C\}\_\{x,t\}\)\-S\(\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\}\}\\eta\)\(126\)=S\(𝒞x,t\)\+∫𝑑η𝒩\(η\|0,\(1−α¯t\)I\)ln𝒩\(η\|0,\(1−α¯t\)I\)\\displaystyle=S\(\\mathcal\{C\}\_\{x,t\}\)\+\\int d\\eta\\mathcal\{N\}\(\\eta\|0,\(1\-\\bar\{\\alpha\}\_\{t\}\)I\)\\ln\\mathcal\{N\}\(\\eta\|0,\(1\-\\bar\{\\alpha\}\_\{t\}\)I\)\(127\)=S\(𝒞x,t\)−∫𝑑η\(1−α¯t2‖η‖2\+d2ln\(2π\(1−α¯t\)\)\)exp\(−1−α¯t2‖η‖2\)\(2π\(1−α¯t\)\)d\\displaystyle=S\(\\mathcal\{C\}\_\{x,t\}\)\-\\int d\\eta\\Big\(\\frac\{1\-\\bar\{\\alpha\}\_\{t\}\}\{2\}\|\|\\eta\|\|^\{2\}\+\\frac\{d\}\{2\}\\ln\(2\\pi\(1\-\\bar\{\\alpha\}\_\{t\}\)\)\\Big\)\\frac\{\\exp\(\-\\frac\{1\-\\bar\{\\alpha\}\_\{t\}\}\{2\}\|\|\\eta\|\|^\{2\}\)\}\{\\sqrt\{\(2\\pi\(1\-\\bar\{\\alpha\}\_\{t\}\)\)^\{d\}\}\}\(128\)=S\(𝒞x,t\)−d2\(1\+ln\(2π\(1−α¯t\)\)\)\\displaystyle=S\(\\mathcal\{C\}\_\{x,t\}\)\-\\frac\{d\}\{2\}\(1\+\\ln\(2\\pi\(1\-\\bar\{\\alpha\}\_\{t\}\)\)\)\(129\)
This recovers the well known theorem describing the Mutual information between the input and output of an additive Gaussian white noise channel:
###### Theorem E\.1
For the additive white noise channel with inputΦ\\Phiand output𝒞x,t:=α¯tΦ\+1−α¯tη\\mathcal\{C\}\_\{x,t\}:=\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\Phi\+\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\}\}\\etawithη∼𝒩\(0,I\)\\eta\\sim\\mathcal\{N\}\(0,I\),
I\(𝒞x,t;Φ\)=S\(𝒞x,t\)−d2\(1\+ln\(2π\(1−α¯t\)\)\)I\(\\mathcal\{C\}\_\{x,t\};\\Phi\)=S\(\\mathcal\{C\}\_\{x,t\}\)\-\\frac\{d\}\{2\}\(1\+\\ln\(2\\pi\(1\-\\bar\{\\alpha\}\_\{t\}\)\)\)\(130\)whereddis the data dimension or the number of pixels\.
As we can see, the mutual information of the additive white noise channel only depends on the input, or data distribution through the entropy of the noisy distributionϕ\\phi\. Therefore, we can bound the mutual information of the channel by choosing a maximal entropy distribution forϕ\\phi\. Under the constraint that the data covariance matrix isΣ0\\Sigma\_\{0\}and its distribution is centered, the maximal entropy distribution ofϕ\\phiis Gaussian with covarianceΣt=α¯tΣ\+\(1−α¯t\)I\\Sigma\_\{t\}=\\bar\{\\alpha\}\_\{t\}\\Sigma\+\(1\-\\bar\{\\alpha\}\_\{t\}\)I\. Letϕ~∼𝒩\(0,Σt\)\\tilde\{\\phi\}\\sim\\mathcal\{N\}\(0,\\Sigma\_\{t\}\)be the Gaussianized version of the data\.
S\(ϕ~\)\\displaystyle S\(\\tilde\{\\phi\}\)=−∫𝑑ϕ~𝒩\(ϕ~\|0,Σt\)ln𝒩\(ϕ~\|0,Σt\)\\displaystyle=\-\\int d\\tilde\{\\phi\}\\mathcal\{N\}\(\\tilde\{\\phi\}\|0,\\Sigma\_\{t\}\)\\ln\\mathcal\{N\}\(\\tilde\{\\phi\}\|0,\\Sigma\_\{t\}\)=12lndet\(2πΣt\)\+∫𝑑ϕ~ϕ~TΣt−1ϕ2exp\(−ϕ~TΣt−1ϕ/2\)det\(2πΣt\)\\displaystyle=\\frac\{1\}\{2\}\\ln\\det\(2\\pi\\Sigma\_\{t\}\)\+\\int d\\tilde\{\\phi\}\\frac\{\\tilde\{\\phi\}^\{T\}\\Sigma\_\{t\}^\{\-1\}\\phi\}\{2\}\\frac\{\\exp\(\-\\tilde\{\\phi\}^\{T\}\\Sigma\_\{t\}^\{\-1\}\\phi/2\)\}\{\\sqrt\{\\det\(2\\pi\\Sigma\_\{t\}\)\}\}=12𝐓𝐫ln\(2πΣt\)\+12𝐓𝐫\[Σt−1Σt\]\\displaystyle=\\frac\{1\}\{2\}\\mathbf\{Tr\}\{\\ln\(2\\pi\\Sigma\_\{t\}\)\}\+\\frac\{1\}\{2\}\\mathbf\{Tr\}\[\\Sigma\_\{t\}^\{\-1\}\\Sigma\_\{t\}\]=d2\+12𝐓𝐫ln\(2πΣt\)\\displaystyle=\\frac\{d\}\{2\}\+\\frac\{1\}\{2\}\\mathbf\{Tr\}\{\\ln\(2\\pi\\Sigma\_\{t\}\)\}Now using this to find the mutual information
I\(ϕ~;Φ~\)\\displaystyle I\(\\tilde\{\\phi\};\\tilde\{\\Phi\}\)=S\(ϕ~\)−d2\(1\+ln\(2π\(1−α¯t\)\)\)\\displaystyle=S\(\\tilde\{\\phi\}\)\-\\frac\{d\}\{2\}\(1\+\\ln\(2\\pi\(1\-\\bar\{\\alpha\}\_\{t\}\)\)\)\(131\)=12𝐓𝐫ln\(Σt1−α¯t\)\\displaystyle=\\frac\{1\}\{2\}\\mathbf\{Tr\}\{\\ln\(\\frac\{\\Sigma\_\{t\}\}\{1\-\\bar\{\\alpha\}\_\{t\}\}\)\}\(132\)=12𝐓𝐫ln\(α¯tΣ\+\(1−α¯t\)I1−α¯t\)\\displaystyle=\\frac\{1\}\{2\}\\mathbf\{Tr\}\{\\ln\(\\frac\{\\bar\{\\alpha\}\_\{t\}\\Sigma\+\(1\-\\bar\{\\alpha\}\_\{t\}\)I\}\{1\-\\bar\{\\alpha\}\_\{t\}\}\)\}\(133\)=12𝐓𝐫ln\(I\+α¯t1−α¯tΣ\)\\displaystyle=\\frac\{1\}\{2\}\\mathbf\{Tr\}\{\\ln\(I\+\\frac\{\\bar\{\\alpha\}\_\{t\}\}\{1\-\\bar\{\\alpha\}\_\{t\}\}\\Sigma\)\}\(134\)Therefore, we can conclude theorem
###### Theorem E\.2
Given a patchΦ\\Phiwith covarianceΣ\\Sigmathe mutual information of the additive white noise channel,ϕ:=α¯tΦ\+1−α¯tη\\phi:=\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\Phi\+\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\}\}\\eta, is bounded above by
I\(ϕ;Φ\)≤12𝐓𝐫ln\[I\+α¯t1−α¯tΣ\]\.I\(\\phi;\\Phi\)\\leq\\frac\{1\}\{2\}\\mathbf\{Tr\}\\ln\[I\+\\frac\{\\bar\{\\alpha\}\_\{t\}\}\{1\-\\bar\{\\alpha\}\_\{t\}\}\\Sigma\]\.\(135\)
### E\.3Asymptotic Expansion for the Mutual Information of the AWNC at High Noise
In this section we will again consider a simple white noise channel\. In section[E\.2](https://arxiv.org/html/2607.08041#A5.SS2), we argued that the Gaussianized distribution provides an upper bound on the mutual information\. Here, we provide a universal expansion for the mutual information of the white noise channel regardless of the underlying distribution\. This helps give intuition for the surprisingly good agreement between the Gaussian prediction and the posterior entropies measured on real data sets\. Recall for the additive white noise channel:
I\(φ;ϕ\)=S\[Ptest\(ϕ\)\]−S\[𝒩\(0,\(1−α¯t\)I\)\]\.I\(\\varphi;\\phi\)=S\[P\_\{test\}\(\\phi\)\]\-S\[\\mathcal\{N\}\(0,\(1\-\\bar\{\\alpha\}\_\{t\}\)I\)\]\.\(136\)We can define the cumulant generating function for the test distribution as follows:
Kϕ\(ϕ^\)=ln𝔼ϕ∼Ptest\[exp\(ϕ^⋅ϕ\)\]K\_\{\\phi\}\(\\hat\{\\phi\}\)=\\ln\{\{\\mathbb\{E\}\}\}\_\{\\phi\\sim P\_\{test\}\}\[\\exp\(\\hat\{\\phi\}\\cdot\\phi\)\]\(137\)The formal Taylor expansion for this function provides the cumulants of the distribution ofϕ\\phi\. In addition, the cumulant generating function with an imaginary argument can be used to entirely reconstruct the distribution\. Because the cumulant generating function of the sum of two independent random variables is just the sum of each variable’s cumulant generating function, we have
Kϕ\(ϕ^\)\\displaystyle K\_\{\\phi\}\(\\hat\{\\phi\}\)=ln𝔼φ∼P0,η∼𝒩\(0,I\)\[exp\(ϕ^⋅\(α¯tφ\+1−α¯tη\)\)\]\\displaystyle=\\ln\{\{\\mathbb\{E\}\}\}\_\{\\varphi\\sim P\_\{0\},\\eta\\sim\\mathcal\{N\}\(0,I\)\}\[\\exp\(\\hat\{\\phi\}\\cdot\(\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\varphi\+\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\}\}\\eta\)\)\]\(138\)=ln𝔼φ∼P0\[exp\(ϕ^⋅\(α¯tφ\)\)\]\+ln𝔼η∼𝒩\(0,I\)\[exp\(1−αtϕ^⋅η\)\)\]\\displaystyle=\\ln\{\{\\mathbb\{E\}\}\}\_\{\\varphi\\sim P\_\{0\}\}\[\\exp\(\\hat\{\\phi\}\\cdot\(\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\varphi\)\)\]\+\\ln\{\{\\mathbb\{E\}\}\}\_\{\\eta\\sim\\mathcal\{N\}\(0,I\)\}\[\\exp\(\\sqrt\{1\-\\alpha\}\_\{t\}\\hat\{\\phi\}\\cdot\\eta\)\)\]\(139\)=Kφ\(α¯tϕ^\)\+Kη\(1−α¯tϕ^\)\\displaystyle=K\_\{\\varphi\}\(\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\hat\{\\phi\}\)\+K\_\{\\eta\}\(\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\}\}\\hat\{\\phi\}\)\(140\)=Kφ\(α¯tϕ^\)\+1−α¯t2‖ϕ^‖2\.\\displaystyle=K\_\{\\varphi\}\(\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\hat\{\\phi\}\)\+\\frac\{1\-\\bar\{\\alpha\}\_\{t\}\}\{2\}\|\|\\hat\{\\phi\}\|\|^\{2\}\.\(141\)The cumulant generating function of the isotropic Gaussian is just a quadratic\. In addition,α¯t\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}ranges from 0 to 1, so one would expect that the low order terms inα¯t\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}describe the behavior at most time steps\. This means that only the lowest order cumulants of the data distribution have a significant effect on the cumulants of the generated distribution\. By simply representing the entropy in terms of the cumulant generating function, we should be able to expand the expression to lowest order terms and find a cumulant expansion for the mutual information\. We start with the expression for the distributio in terms of its CGF:
Ptest\(ϕ\)\\displaystyle P\_\{test\}\(\\phi\)=∫dϕ^\(2π\)dexp\(−iϕ^⋅ϕ\+Kϕ\(iϕ^\)\)\.\\displaystyle=\\int\\frac\{d\\hat\{\\phi\}\}\{\(2\\pi\)^\{d\}\}\\exp\(\-i\\hat\{\\phi\}\\cdot\\phi\+K\_\{\\phi\}\(i\\hat\{\\phi\}\)\)\.\(142\)Taking derivatives with respect toϕ\\phihas the effect of adding factors of−iϕ^\-i\\hat\{\\phi\}to the integrand in equation \([142](https://arxiv.org/html/2607.08041#A5.E142)\)\.
Ptest\(ϕ\)\\displaystyle P\_\{test\}\(\\phi\)=∫dϕ^\(2π\)dexp\(−iϕ^⋅ϕ\+Kϕ\(iϕ^\)\)\\displaystyle=\\int\\frac\{d\\hat\{\\phi\}\}\{\(2\\pi\)^\{d\}\}\\exp\(\-i\\hat\{\\phi\}\\cdot\\phi\+K\_\{\\phi\}\(i\\hat\{\\phi\}\)\)\(143\)=∫dϕ^\(2π\)dexp\(Kφ\(iα¯tϕ^\)\)exp\(−iϕ^⋅ϕ−1−α¯t2‖ϕ^‖2\)\\displaystyle=\\int\\frac\{d\\hat\{\\phi\}\}\{\(2\\pi\)^\{d\}\}\\exp\(K\_\{\\varphi\}\(i\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\hat\{\\phi\}\)\)\\exp\(\-i\\hat\{\\phi\}\\cdot\\phi\-\\frac\{1\-\\bar\{\\alpha\}\_\{t\}\}\{2\}\|\|\\hat\{\\phi\}\|\|^\{2\}\)\(144\)=exp\(Kφ\(−α¯t∂ϕ\)\)∫dϕ^\(2π\)dexp\(−iϕ^⋅ϕ−1−α¯t2‖ϕ^‖2\)\\displaystyle=\\exp\(K\_\{\\varphi\}\(\-\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\partial\_\{\\phi\}\)\)\\int\\frac\{d\\hat\{\\phi\}\}\{\(2\\pi\)^\{d\}\}\\exp\(\-i\\hat\{\\phi\}\\cdot\\phi\-\\frac\{1\-\\bar\{\\alpha\}\_\{t\}\}\{2\}\|\|\\hat\{\\phi\}\|\|^\{2\}\)\(145\)=exp\(Kφ\(−α¯t∂ϕ\)\)exp\(−‖ϕ‖22\(1−α¯t\)\)det\(2πΣt\)\\displaystyle=\\exp\(K\_\{\\varphi\}\(\-\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\partial\_\{\\phi\}\)\)\\frac\{\\exp\(\-\\frac\{\|\|\\phi\|\|^\{2\}\}\{2\(1\-\\bar\{\\alpha\}\_\{t\}\)\}\)\}\{\\sqrt\{\\det\(2\\pi\\Sigma\_\{t\}\)\}\}\(146\)In the last line,
S\[Ptest\(ϕ\)\]\\displaystyle S\[P\_\{test\}\(\\phi\)\]=−∫𝑑ϕPtest\(ϕ\)lnPtest\(ϕ\)\\displaystyle=\-\\int d\\phi\\\>P\_\{test\}\(\\phi\)\\ln P\_\{test\}\(\\phi\)\(147\)=−∫𝑑ϕexp\(K~φ\(−α¯t∂ϕ\)\)𝒩\(ϕ\|0,\(1−α¯t\)I\)\\displaystyle\\hskip\-20\.0pt=\-\\int d\\phi\\\>\\exp\(\\tilde\{K\}\_\{\\varphi\}\(\-\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\partial\_\{\\phi\}\)\)\\mathcal\{N\}\(\\phi\|0,\(1\-\\bar\{\\alpha\}\_\{t\}\)I\)\(148\)×lnexp\(K~φ\(−α¯t∂ϕ\)\)𝒩\(ϕ\|0,\(1−α¯t\)I\)\\displaystyle\\hskip 20\.0pt\\times\\ln\\exp\(\\tilde\{K\}\_\{\\varphi\}\(\-\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\partial\_\{\\phi\}\)\)\\mathcal\{N\}\(\\phi\|0,\(1\-\\bar\{\\alpha\}\_\{t\}\)I\)\(149\)Expanding equation \([149](https://arxiv.org/html/2607.08041#A5.E149)\) in powers ofα¯t\\bar\{\\alpha\}\_\{t\}will provide an expansion of the entropy in terms of higher cumulants of the distribution and expectations of Gaussian random variables\. For example, the first order term inα¯t\\bar\{\\alpha\}\_\{t\}is
S\[Ptest\(ϕ\)\]\\displaystyle S\[P\_\{test\}\(\\phi\)\]=d2\(1\+ln\(2π\(1−α¯t\)\)\)\+12\(α¯t1−α¯t\)𝐓𝐫\[Σ\]\\displaystyle=\\frac\{d\}\{2\}\(1\+\\ln\(2\\pi\(1\-\\bar\{\\alpha\}\_\{t\}\)\)\)\+\\frac\{1\}\{2\}\\Big\(\\frac\{\\bar\{\\alpha\}\_\{t\}\}\{1\-\\bar\{\\alpha\}\_\{t\}\}\\Big\)\\mathbf\{Tr\}\[\\Sigma\]\(150\)−14\(α¯t1−α¯t\)2\(𝐓𝐫\[Σ2\]\+κabcd12\(δa,bδc,d\+δa,cδb,d\+δa,dδb,c\)\)\\displaystyle\-\\frac\{1\}\{4\}\\Big\(\\frac\{\\bar\{\\alpha\}\_\{t\}\}\{1\-\\bar\{\\alpha\}\_\{t\}\}\\Big\)^\{2\}\(\\mathbf\{Tr\}\[\\Sigma^\{2\}\]\+\\frac\{\\kappa^\{abcd\}\}\{12\}\(\\delta\_\{a,b\}\\delta\_\{c,d\}\+\\delta\_\{a,c\}\\delta\_\{b,d\}\+\\delta\_\{a,d\}\\delta\_\{b,c\}\)\)\(151\)\+O\(α¯t5/2\)\\displaystyle\+O\(\\bar\{\\alpha\}\_\{t\}^\{5/2\}\)\(152\)whereΣ\\Sigmais the covariance matrix of the dataset andκabcd\\kappa\_\{abcd\}is the four\-cumulant of the data\. We can convert this into a prediction for the mutual information of the white noise channel
I\(ϕ;φ\)\\displaystyle I\(\\phi;\\varphi\)=S\[Ptest\(ϕ\)\]−S\[𝒩\(0,\(1−α¯t\)I\)\]\\displaystyle=S\[P\_\{test\}\(\\phi\)\]\-S\[\\mathcal\{N\}\(0,\(1\-\\bar\{\\alpha\}\_\{t\}\)I\)\]\(153\)=12\(α¯t1−α¯t\)𝐓𝐫\[Σ\]\\displaystyle=\\frac\{1\}\{2\}\\Big\(\\frac\{\\bar\{\\alpha\}\_\{t\}\}\{1\-\\bar\{\\alpha\}\_\{t\}\}\\Big\)\\mathbf\{Tr\}\[\\Sigma\]\(154\)−14\(α¯t1−α¯t\)2\(𝐓𝐫\[Σ2\]\+κabcd12\(δa,bδc,d\+δa,cδb,d\+δa,dδb,c\)\)\\displaystyle\-\\frac\{1\}\{4\}\\Big\(\\frac\{\\bar\{\\alpha\}\_\{t\}\}\{1\-\\bar\{\\alpha\}\_\{t\}\}\\Big\)^\{2\}\(\\mathbf\{Tr\}\[\\Sigma^\{2\}\]\+\\frac\{\\kappa^\{abcd\}\}\{12\}\(\\delta\_\{a,b\}\\delta\_\{c,d\}\+\\delta\_\{a,c\}\\delta\_\{b,d\}\+\\delta\_\{a,d\}\\delta\_\{b,c\}\)\)\(155\)\+O\(α¯t5/2\)\\displaystyle\+O\(\\bar\{\\alpha\}\_\{t\}^\{5/2\}\)\(156\)Higher order terms of this expansion are quite tedious to calculate\.
### E\.4A Data Model for Speciation with Exactly Solvable Mutual Information
Imagine that the true image distribution is a mixture of Gaussians where the parameters are themselves random\.
P0\(φ\)=1\|M\|∑μ∈M𝒩\(φ\|μ,Σ0\)P\_\{0\}\(\\varphi\)=\\frac\{1\}\{\|M\|\}\\sum\_\{\\mu\\in M\}\\mathcal\{N\}\(\\varphi\|\\mu,\\Sigma\_\{0\}\)\(157\)where the meansMMare drawn from some other distribution𝒩\(0,ΣM\)\\mathcal\{N\}\(0,\\Sigma\_\{M\}\)\. If we assume that\|M\|\|M\|is large, we can exactly solve for the mutual information of this data passed through the white noise channel\. This can be achieved by noting thatP0\(φ\)P\_\{0\}\(\\varphi\)itself is exactly the same as the training distribution created byMMsamples from𝒩\(0,ΣM\)\\mathcal\{N\}\(0,\\Sigma\_\{M\}\)then passed through an additive noise channelμ\+η\\mu\+\\etawithη∼𝒩\(φ\|μ,Σ0\)\\eta\\sim\\mathcal\{N\}\(\\varphi\|\\mu,\\Sigma\_\{0\}\)\. However, the theory that we’ve built to this point would only allow you to predict the entropy of the posterior distribution, that is givenφ\\varphiwhat is the probability that it came from the Gaussian centered atμ\\mu\. However, we can instead make use of our prediction for the partition function
Z\(φ\)=∑m∈M𝒩\(φ\|μ,Σ0\)=\{\|M\|𝒩\(φ\|0,Σ0\+ΣM\)ln\|M\|\>DKL\(P\(μ\|φ\)\|P\(μ\)\)e−βE∗ln\|M\|≤DKL\(P\(μ\|φ\)\|P\(μ\)\)Z\(\\varphi\)=\\sum\_\{m\\in M\}\\mathcal\{N\}\(\\varphi\|\\mu,\\Sigma\_\{0\}\)=\\begin\{cases\}\|M\|\\mathcal\{N\}\(\\varphi\|0,\\Sigma\_\{0\}\+\\Sigma\_\{M\}\)&\\ln\|M\|\>D\_\{KL\}\(P\(\\mu\|\\varphi\)\|P\(\\mu\)\)\\\\ e^\{\-\\beta E^\{\*\}\}&\\ln\|M\|\\leq D\_\{KL\}\(P\(\\mu\|\\varphi\)\|P\(\\mu\)\)\\end\{cases\}\(158\)wheree−βE∗e^\{\-\\beta E^\{\*\}\}does not concentrate, but behaves like M well separated Gaussian distributions\. This predicts the probability density atφ\\varphiin the large\|M\|\|M\|limit\. Leading to the counterintuitive result thatP0\(φ\)P\_\{0\}\(\\varphi\)concentrates to a value independent ofMM\. This leads to
S\(φ\)=\{S\[𝒩\(μ,Σ0\+ΣM\)\]ln\|M\|\>I\(μ;φ\)ln\|M\|\+S\[𝒩\(0,Σ0\)\]ln\|M\|≤I\(μ;φ\)S\(\\varphi\)=\\begin\{cases\}S\[\\mathcal\{N\}\(\\mu,\\Sigma\_\{0\}\+\\Sigma\_\{M\}\)\]&\\ln\|M\|\>I\(\\mu;\\varphi\)\\\\ \\ln\|M\|\+S\[\\mathcal\{N\}\(0,\\Sigma\_\{0\}\)\]&\\ln\|M\|\\leq I\(\\mu;\\varphi\)\\end\{cases\}\(159\)Adding in the white noise channel, this becomes
S\(ϕt\)=\{S\[𝒩\(μ,Σ0\+ΣM\+σt2I\)\]ln\|M\|\>I\(μ;φ\)ln\|M\|\+S\[𝒩\(0,Σ0\+σt2I\)\]ln\|M\|≤I\(μ;φ\)S\(\\phi\_\{t\}\)=\\begin\{cases\}S\[\\mathcal\{N\}\(\\mu,\\Sigma\_\{0\}\+\\Sigma\_\{M\}\+\\sigma\_\{t\}^\{2\}I\)\]&\\ln\|M\|\>I\(\\mu;\\varphi\)\\\\ \\ln\|M\|\+S\[\\mathcal\{N\}\(0,\\Sigma\_\{0\}\+\\sigma\_\{t\}^\{2\}I\)\]&\\ln\|M\|\\leq I\(\\mu;\\varphi\)\\end\{cases\}\(160\)So, the mutual information is
I\(ϕt;φ\)=\{S\[𝒩\(μ,Σ0\+ΣM\+σt2I\)\]−S\[𝒩\(0,σt2I\)\]ln\|M\|\>I\(μ;φ\)ln\|M\|\+S\[𝒩\(0,Σ0\+σt2I\)\]−S\[𝒩\(0,σt2I\)\]ln\|M\|≤I\(μ;φ\)I\(\\phi\_\{t\};\\varphi\)=\\begin\{cases\}S\[\\mathcal\{N\}\(\\mu,\\Sigma\_\{0\}\+\\Sigma\_\{M\}\+\\sigma\_\{t\}^\{2\}I\)\]\-S\[\\mathcal\{N\}\(0,\\sigma\_\{t\}^\{2\}I\)\]&\\ln\|M\|\>I\(\\mu;\\varphi\)\\\\ \\ln\|M\|\+S\[\\mathcal\{N\}\(0,\\Sigma\_\{0\}\+\\sigma\_\{t\}^\{2\}I\)\]\-S\[\\mathcal\{N\}\(0,\\sigma\_\{t\}^\{2\}I\)\]&\\ln\|M\|\\leq I\(\\mu;\\varphi\)\\end\{cases\}\(161\)Plugging in the Gaussian entropy formula
I\(ϕt;φ\)=\{Tr\[ln\(I\+1σt2\(Σ0\+ΣM\)\)\]ln\|M\|\>I\(μ;φ\)ln\|M\|\+Tr\[ln\(I\+1σt2Σ0\)\]ln\|M\|≤I\(μ;φ\)I\(\\phi\_\{t\};\\varphi\)=\\begin\{cases\}Tr\[\\ln\(I\+\\frac\{1\}\{\\sigma\_\{t\}^\{2\}\}\(\\Sigma\_\{0\}\+\\Sigma\_\{M\}\)\)\]&\\ln\|M\|\>I\(\\mu;\\varphi\)\\\\ \\ln\|M\|\+Tr\[\\ln\(I\+\\frac\{1\}\{\\sigma\_\{t\}^\{2\}\}\\Sigma\_\{0\}\)\]&\\ln\|M\|\\leq I\(\\mu;\\varphi\)\\end\{cases\}\(162\)This can also be written
I\(ϕt;φ\)=min\(ln\|M\|\+Tr\[ln\(I\+1σt2Σ0\)\],Tr\[ln\(I\+1σt2\(Σ0\+ΣM\)\)\]\)I\(\\phi\_\{t\};\\varphi\)=\\min\(\\ln\|M\|\+Tr\[\\ln\(I\+\\frac\{1\}\{\\sigma\_\{t\}^\{2\}\}\\Sigma\_\{0\}\)\],Tr\[\\ln\(I\+\\frac\{1\}\{\\sigma\_\{t\}^\{2\}\}\(\\Sigma\_\{0\}\+\\Sigma\_\{M\}\)\)\]\)\(163\)This leads to the formula for the posterior entropy
S\[P\(φ\|ϕ\)\]=max\(0,ln\|𝒟\|\|M\|−Tr\[ln\(I\+1σt2Σ0\)\],ln\|𝒟\|−Tr\[ln\(I\+1σt2\(Σ0\+ΣM\)\)\]\)S\[P\(\\varphi\|\\phi\)\]=\\operatorname\{max\}\(0,\\ln\\frac\{\|\\mathcal\{D\}\|\}\{\|M\|\}\-Tr\\Big\[\\ln\(I\+\\frac\{1\}\{\\sigma\_\{t\}^\{2\}\}\\Sigma\_\{0\}\)\\Big\],\\ln\|\\mathcal\{D\}\|\-Tr\\Big\[\\ln\(I\+\\frac\{1\}\{\\sigma\_\{t\}^\{2\}\}\(\\Sigma\_\{0\}\+\\Sigma\_\{M\}\)\)\\Big\]\)\(164\)So, we see that the speciation transition has the exact same behavior as a memorization/generalization transition in terms of a discontinuity in the derivative of the posterior entropy with the noise level\. You can see an example of this in sub\-figure[3](https://arxiv.org/html/2607.08041#S4.F3)b where the posterior entropy seems to plateau twice\. More generally, we could imagine a Hierarchical Gaussian model where a set of meansM1M\_\{1\}is drawn from𝒩\(0,ΣM1\)\\mathcal\{N\}\(0,\\Sigma\_\{M\_\{1\}\}\)to generate a mixture of Gaussian distribution with covariance of each GaussianΣM2\\Sigma\_\{M\_\{2\}\}from which I draw as set of meansM2M\_\{2\}, and so on until the final set of meansMNM\_\{N\}\. In this case, we will have a posterior entropy like
S\[P\(φ\|ϕ\)\]=max\(0,maxn≤Nln\|𝒟\|Πi=1n\|Mi\|−Tr\[ln\(I\+1σt2Σ0\+1σt2∑i=n\+1NΣMi\)\]\)S\[P\(\\varphi\|\\phi\)\]=\\operatorname\{max\}\(0,\\operatorname\{max\}\_\{n\\leq N\}\\ln\\frac\{\|\\mathcal\{D\}\|\}\{\\Pi\_\{i=1\}^\{n\}\|M\_\{i\}\|\}\-Tr\[\\ln\(I\+\\frac\{1\}\{\\sigma\_\{t\}^\{2\}\}\\Sigma\_\{0\}\+\\frac\{1\}\{\\sigma\_\{t\}^\{2\}\}\\sum\_\{i=\{n\+1\}\}^\{N\}\\Sigma\_\{M\_\{i\}\}\)\]\)\(165\)where each change of the maximumnncorresponds to a speciation transition\.
## Appendix FGeneralization in Local Score Models and the critical scale
As we have discussed above, Bayes\-optimal diffusion modelsgeneralizeonly when the channel through which they are observing their image lets pass fewer thanlog\|𝒟\|\\log\|\\mathcal\{D\}\|bits of information, preventing them from being able to resolve down to a single training set image\. Increasing the noise level clearly decreases this mutual information, and extensive prior work has gone into studying the location of the memorization/generalization transition as a function of this parameter\[[6](https://arxiv.org/html/2607.08041#bib.bib6),[47](https://arxiv.org/html/2607.08041#bib.bib47)\]\. Another parameter directly relevant to the capacity of the channel, however, is thelocality scale\[[13](https://arxiv.org/html/2607.08041#bib.bib13),[48](https://arxiv.org/html/2607.08041#bib.bib48),[32](https://arxiv.org/html/2607.08041#bib.bib32)\]\. When diffusion models have locality biases, they may only effectively observe a small proportion of the overall image\. Under such constraints, the effective amount of information that they may in principle use to deduce the underlying image will be substantiallylowerthan the global information contained in the image\. We thus ask the following question: where is the threshold of criticality,as a function of both the noise level and the locality scale?
In this section, for conciseness, we will use the invariant noise\-to\-signal ratio to characterize the noise level at a timett:
σt2=1−α¯tα¯t\\displaystyle\\sigma\_\{t\}^\{2\}=\\frac\{1\-\\bar\{\\alpha\}\_\{t\}\}\{\\bar\{\\alpha\}\_\{t\}\}\(166\)We will consider a family of patchesΩL,x\\Omega\_\{L,x\}indexed uniquely by the parameterLL, which satisfiesL=\|ΩL,x\|L=\\sqrt\{\|\\Omega\_\{L,x\}\|\}\. For square patches\[[13](https://arxiv.org/html/2607.08041#bib.bib13)\]this quantity is simply the length of one side of the patch\. We will demand thatΩL′,x⊃ΩL,x\\Omega\_\{L^\{\\prime\},x\}\\supset\\Omega\_\{L,x\}ifL′\>LL^\{\\prime\}\>L\. Thecritical scaleLc\(σt\)L\_\{c\}\(\\sigma\_\{t\}\)is then defined as the scale where a Local Score model using the patchΩL,x\\Omega\_\{L,x\}achieves the average collapse condition \([40](https://arxiv.org/html/2607.08041#A4.E40)\) at a noise levelσt\\sigma\_\{t\}:
ln\|𝒟\|=I\(φ;ϕt,Ωx,L\)\\displaystyle\\ln\|\\mathcal\{D\}\|=I\(\\varphi;\\phi\_\{t,\{\\Omega\_\{x,L\}\}\}\)\(167\)Without making any assumptions about the data\-generating distribution, we are in a position to make a very general statement\. Firstly, increasing the scale of a patch given a fixed noise scale monotonically increases the mutual informationI\(φ;ϕt;Ω\)I\(\\varphi;\\phi\_\{t;\\Omega\}\), while increasing the noise scale at a fixed patch size monotonically decreases this mutual information\. It follows from implicit differentiation of the collapse condition \([167](https://arxiv.org/html/2607.08041#A6.E167)\) with respect toLLthat the values ofLc\(σt\)L\_\{c\}\(\\sigma\_\{t\}\)must be monotonically increasing with noise level\. This is phenomenologically consistent with the observed coarse\-to\-fine scaling of diffusion models\[[13](https://arxiv.org/html/2607.08041#bib.bib13),[32](https://arxiv.org/html/2607.08041#bib.bib32),[14](https://arxiv.org/html/2607.08041#bib.bib14),[40](https://arxiv.org/html/2607.08041#bib.bib40)\]\. We can formalize this into this theorem\.
###### Theorem F\.1\(Monotonicity of Length Scale\)
Lc\(σt\)L\_\{c\}\(\\sigma\_\{t\}\)is a monotonically increasing function\.
###### Proof\.
IfL<L′L<L^\{\\prime\}, thenΩL′,x⊃ΩL,x\\Omega\_\{L^\{\\prime\},x\}\\supset\\Omega\_\{L,x\}\. This implies that the observationCL,x,tC\_\{L,x,t\}can be viewed as first restricting to the patchΩL′,x\\Omega\_\{L^\{\\prime\},x\}and then restricting to the patchΩL,x\\Omega\_\{L,x\}\. This leads to the Markov chain
φ→ϕt→ϕt,Ωx,L′→ϕt,Ωx,L\\varphi\\rightarrow\\phi\_\{t\}\\rightarrow\\phi\_\{t,\{\\Omega\_\{x,L^\{\\prime\}\}\}\}\\rightarrow\\phi\_\{t,\{\\Omega\_\{x,L\}\}\}\(168\)The information processing inequality implies
I\(φ;ϕt,Ωx,L\)≤I\(φ;ϕt,Ωx,L′\)I\(\\varphi;\\phi\_\{t,\{\\Omega\_\{x,L\}\}\}\)\\leq I\(\\varphi;\\phi\_\{t,\{\\Omega\_\{x,L^\{\\prime\}\}\}\}\)\(169\)Now we need to use the fact that for a projection we can switch the order of restricting the domain and adding the noise\. For a positive incrementΔt\\Delta t
φ→φΩx,L→ϕt,Ωx,L→ϕt\+Δt,Ωx,L\\varphi\\rightarrow\\varphi\_\{\{\\Omega\_\{x,L\}\}\}\\rightarrow\\phi\_\{t,\{\\Omega\_\{x,L\}\}\}\\rightarrow\\phi\_\{t\+\\Delta t,\{\\Omega\_\{x,L\}\}\}\(170\)This implies
I\(φ;ϕt\+Δt,Ωx\)≤I\(φ;ϕt,Ωx,L\)I\(\\varphi;\\phi\_\{t\+\\Delta t,\{\\Omega\_\{x\}\}\}\)\\leq I\(\\varphi;\\phi\_\{t,\{\\Omega\_\{x,L\}\}\}\)\(171\)The collapse condition \([40](https://arxiv.org/html/2607.08041#A4.E40)\) then implies the mutual information is held constant at the transition sizeϕt,Ωx,L\(σt\)\\phi\_\{t,\{\\Omega\_\{x,L\(\\sigma\_\{t\}\)\}\}\}\. Sinceσt\\sigma\_\{t\}is monotonically increasing intt, this implies
\|Ωx,Lc\(σt\)\|≤\|Ωx,Lc\(σt\+Δσ\)\|\|\\Omega\_\{x,L\_\{c\}\(\\sigma\_\{t\}\)\}\|\\leq\|\\Omega\_\{x,L\_\{c\}\(\\sigma\_\{t\}\+\\Delta\\sigma\)\}\|\(172\)Lc\(σt\)≤Lc\(σt\+Δσt\)L\_\{c\}\(\\sigma\_\{t\}\)\\leq L\_\{c\}\(\\sigma\_\{t\}\+\\Delta\\sigma\_\{t\}\)\(173\)
### F\.1What scales do weneedvs\. what scales do wehave: the spectral scale vs\. the critical scale
Existing theories of locality\[[32](https://arxiv.org/html/2607.08041#bib.bib32),[40](https://arxiv.org/html/2607.08041#bib.bib40)\]in diffusion models emphasize a quantity which we term thespectral scale\. Naturalistic images have a well\-known power\-law falloff in their power spectral density\[[39](https://arxiv.org/html/2607.08041#bib.bib39)\]:
P\(k\)≈A\|k\|−α\\displaystyle P\(k\)\\approx A\|k\|^\{\-\\alpha\}\(174\)with the power\-law exponentα≈2−ϵ\\alpha\\approx 2\-\\epsilon, withϵ\\epsilontypically in the range 0\.1\-0\.3\. Theα=2\\alpha=2case is the exactly scale invariant case, where each logarithmic frequency interval has exactly the same total variance\. The observedα<2\\alpha<2implies that natural images are nearly scale\-invariant but slightly ‘UV tilted,’ showing a bit more power at higher frequencies than at lower frequencies\. When performing denoising diffusion, this power\-law distribution is mixed with white noise, with a flat power spectral densityS\(k\)=σt2S\(k\)=\\sigma\_\{t\}^\{2\}\. There is a natural length scale in this problem, defined as the inverse of the spatial frequency where the noise power equals the signal power:
Lspec=\(σt2A\)1/α\\displaystyle L\_\{spec\}=\\bigg\(\\frac\{\\sigma\_\{t\}^\{2\}\}\{A\}\\bigg\)^\{1/\\alpha\}\(175\)For exactly scale\-invariant data withα=2\\alpha=2, this reduces to
Lspec∼σt\\displaystyle L\_\{spec\}\\sim\\sigma\_\{t\}\(176\)i\.e\., the spectral scale is linearly proportional to the standard deviation of the added noise\.
Another useful way of interpreting the spectral scale is as the approximate radius of the optimal linear denoiser, which for stationary translationally\-invariant data is known as the Wiener filter\[[41](https://arxiv.org/html/2607.08041#bib.bib41)\]\. The optimal linear denoiser is the minimizer of the denoising objective \([14](https://arxiv.org/html/2607.08041#A1.E14)\) under the constraint that the model should be an affine functionMt\[ϕt\]=W^tϕt\+btM\_\{t\}\[\\phi\_\{t\}\]=\\hat\{W\}\_\{t\}\\phi\_\{t\}\+b\_\{t\}\. This is a standard linear regression problem, and for a distributionP0\(φ\)P\_\{0\}\(\\varphi\)with meanμ\\muand covarianceΣ\\Sigma, the optimalW^t\\hat\{W\}\_\{t\}andbtb\_\{t\}are given by, respectively,
W^t\\displaystyle\\hat\{W\}\_\{t\}=\(α¯tΣ\+\(1−α¯t\)I\)−1\(α¯tΣ\)\\displaystyle=\(\\bar\{\\alpha\}\_\{t\}\\Sigma\+\(1\-\\bar\{\\alpha\}\_\{t\}\)I\)^\{\-1\}\(\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\Sigma\)\(177\)bt\\displaystyle b\_\{t\}=\(I−W^t\)α¯tμ\.\\displaystyle=\(I\-\\hat\{W\}\_\{t\}\)\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\,\\mu\.\(178\)The rows ofW^t\\hat\{W\}\_\{t\}can be interpreted as denoising filters applied at a particular location\. When the data statistics are translationally\-invariant, the matrixW^t\\hat\{W\}\_\{t\}implements a convolution with a single filter, the Wiener filter\. Convolution operators correspond to multiplication in Fourier space; the action ofW^t\\hat\{W\}\_\{t\}for such data is given by
W^t\(k\)ϕ^t\(k\)\\displaystyle\\hat\{W\}\_\{t\}\(k\)\\hat\{\\phi\}\_\{t\}\(k\)=α¯tΣ\(k\)α¯tΣ\(k\)\+\(1−α¯t\)ϕ^t\(k\)\\displaystyle=\\frac\{\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\Sigma\(k\)\}\{\\bar\{\\alpha\}\_\{t\}\\Sigma\(k\)\+\(1\-\\bar\{\\alpha\}\_\{t\}\)\}\\,\\hat\{\\phi\}\_\{t\}\(k\)\(179\)whereΣ\(k\)=Var\[φ^\(k\)\]=P\(k\)\\Sigma\(k\)=\\text\{Var\}\[\\hat\{\\varphi\}\(k\)\]=P\(k\)is the variance of thekkth Fourier mode of the unnoised images underP0\(φ\)P\_\{0\}\(\\varphi\)\. This filter generally has infinite support, so to make sense of its ‘size’ one has to make a definitional choice; the most natural characteristic width is the point where the signal and noise power meet in the denominator of the filter, i\.e\.
α¯tΣ\(k\)=\(1−α¯t\)\\displaystyle\\bar\{\\alpha\}\_\{t\}\\Sigma\(k\)=\(1\-\\bar\{\\alpha\}\_\{t\}\)\(180\)Plugging in the power\-law expression \([174](https://arxiv.org/html/2607.08041#A6.E174)\) forP\(k\)=Σ\(k\)P\(k\)=\\Sigma\(k\)and rearranging yields the expression for the ‘fourier space radius’\|kspec\|\|k\_\{spec\}\|\. SettingLspec=\|kspec\|−1L\_\{spec\}=\|k\_\{spec\}\|^\{\-1\}yields the spectral scale \([175](https://arxiv.org/html/2607.08041#A6.E175)\)\.
Because of the relatively compact support of this filter, it can be inferred that a strictly local denoiser with a locality scale≥Lspec\\geq L\_\{spec\}should perform well with a relatively small error\. Indeed,\[[32](https://arxiv.org/html/2607.08041#bib.bib32)\]established that calibrating an LS model according to the radius of the Wiener filter yields models that perform capably even to moderately large image sizes\. Conversely, we should expect that any reasonably capable denoiser should be able to attendat leastto roughly the radius of this filter\. If the critical scale is muchsmallerthan the spectral scale \(Lspec≫LcL\_\{spec\}\\gg L\_\{c\}\), a Bayes\-optimal local denoisercannotperform well: it will either incur catastrophic memorization by using a scaleL≫LspecL\\gg L\_\{spec\}comparable toLspecL\_\{spec\}, or its denoising estimate will be robust but highly inaccurate due to using a scaleL<Lc≪LspecL<L\_\{c\}\\ll L\_\{spec\}\. Figure[4\(a\)](https://arxiv.org/html/2607.08041#S5.F4.sf1)illustrates both of these failure modes: when the scale the model uses is smaller than the spectral scale, the outputs of the denoiser are robust but suboptimal\. Conversely, when the scale the model uses is larger than the critical scale, the model’s output incurs catastrophic statistical error\.
Realistic data are highly non\-Gaussian, with higher\-order dependencies that exist and longer scales than implied by the two\-point statistics\. Indeed,\[[32](https://arxiv.org/html/2607.08041#bib.bib32)\]finds that localizing based on the spectral scale breaks down on the highly non\-Gaussian dataset MNIST\. Nonetheless, achieving a radius ofat leastthe spectral scale is clearlynecessaryfor any performant local denoiser, and appears to be sufficient for a reasonable baseline level of denoising accuracy\. A natural question then arises:is the critical scale for a Bayes\-optimal local denoiser large enough to meet or exceed the spectral scale? While by no means obvious, we show in this section that, both theoretically and empirically, the critical scale typicallydoesscale similarly to the spectral scale\. This remarkable concordance between the scalesneededfor denoising and the scales at which robust denoising ispossible\(using a non\-exponential dataset size\) is at the heart of why local Bayes optimal models succeed at ‘reasonable’ generalization on high\-dimensional image data\.
### F\.2A concrete calculation: the isotropic Gaussian
The simplest case to analyze is the isotropic Gaussian first analyzed by\[[6](https://arxiv.org/html/2607.08041#bib.bib6)\]\. Given a patchΩ\\Omega, the mutual information between theΩ\\Omega\-restricted imageϕt;Ω\\phi\_\{t;\\Omega\}, and the underlying unnoised imageφ\\varphiis given by
I\(ϕt;Ω;φ\)=\|Ω\|2ln\(1\+s2σt2\)\\displaystyle I\(\\phi\_\{t;\\Omega\};\\varphi\)=\\frac\{\|\\Omega\|\}\{2\}\\ln\\bigg\(1\+\\frac\{s^\{2\}\}\{\\sigma\_\{t\}^\{2\}\}\\bigg\.\)\(181\)wheres2s^\{2\}is the variance ofφ\\varphi,\|Ω\|\|\\Omega\|is the number of pixels in the patch\. The critical scaleLc=\|Ωmem\|L\_\{c\}=\\sqrt\{\|\\Omega\_\{mem\}\|\}is then given by equating this quantity to the number of natsln\|𝒟\|\\ln\|\\mathcal\{D\}\|needed to resolve down to a single training set image:
ln\|𝒟\|=Lc22ln\(1\+s2σt2\)\\displaystyle\\ln\|\\mathcal\{D\}\|=\\frac\{L\_\{c\}^\{2\}\}\{2\}\\ln\\bigg\(1\+\\frac\{s^\{2\}\}\{\\sigma\_\{t\}^\{2\}\}\\bigg\.\)\(182\)which upon rearranging yields
Lc=2ln\|𝒟\|ln\(1\+s2σt2\)\\displaystyle L\_\{c\}=\\sqrt\{\\frac\{2\\ln\|\\mathcal\{D\}\|\}\{\\ln\(1\+\\frac\{s^\{2\}\}\{\\sigma\_\{t\}^\{2\}\}\)\}\}\(183\)While presented in slightly different form, under the identification\|Ω\|=Lc2=d\|\\Omega\|=L\_\{c\}^\{2\}=dthis result reproduces exactly the ‘collapse time’ result of\[[6](https://arxiv.org/html/2607.08041#bib.bib6)\]\. At largeσt2\\sigma\_\{t\}^\{2\}, this quantity becomes
Lc→σt2ln\|𝒟\|s2\\displaystyle L\_\{c\}\\to\\sigma\_\{t\}\\,\\sqrt\{\\frac\{2\\ln\|\\mathcal\{D\}\|\}\{s^\{2\}\}\}\(184\)i\.e\. we find that at high noise levels, the critical scale is directly proportional to the noise level\. This asymptotic scaling is promising: it suggests that the requisite spectral scaling \([176](https://arxiv.org/html/2607.08041#A6.E176)\) may be achieved by a generalizing Bayes\-optimal denoiser even for asymptotically large image sizes, without scaling the dataset\. As we will see, the form derived above is very general across a wide range of images\.
### F\.3Scaling bounds
If we have a more specific model for our data distribution whereI\(ϕt;Ω;φ\)I\(\\phi\_\{t;\\Omega\};\\varphi\)is tractable, we can in principle compute the location of the critical threshold exactly\. This computation can also be performed numerically; we illustrate some such results in figure[4](https://arxiv.org/html/2607.08041#S5.F4)\. However, doing so directly is often challenging, and extracting an analytic solution forLc\(σt\)L\_\{c\}\(\\sigma\_\{t\}\)from the collapse condition \([167](https://arxiv.org/html/2607.08041#A6.E167)\) may not be feasible even when the mutual informationI\(ϕt;Ω;φ\)I\(\\phi\_\{t;\\Omega\};\\varphi\)is available\. In such cases, however, we can produce a number of useful bounds\. The first is an immediate corollary of the additive noise bound:
###### Corollary F\.2
Assume an additive white noise channel with a locality constraintϕΩ,t\\phi\_\{\\Omega,t\}, and supposeφ\\varphiis drawn from an arbitrary distributionPPwith meanμ\\muand covarianceΣ\\Sigma\. We define the distributionQ1=𝒩\(μ,Σ\)Q\_\{1\}=\\mathcal\{N\}\(\\mu,\\Sigma\)to be the Gaussian distribution with identical second\-order statistics toPP, andQ2=𝒩\(μ,Diag\[Σ\]\)Q\_\{2\}=\\mathcal\{N\}\(\\mu,\\text\{Diag\}\[\\Sigma\]\)to be the Gaussian distribution with identical diagonal covariance and zero off\-diagonal elements\. We then have
Lc;P\(σt\)≥Lc;Q1\(σt\)≥Lc;Q2\(σt\)\\displaystyle L\_\{c;P\}\(\\sigma\_\{t\}\)\\geq L\_\{c;Q\_\{1\}\}\(\\sigma\_\{t\}\)\\geq L\_\{c;Q\_\{2\}\}\(\\sigma\_\{t\}\)\(185\)
###### Proof\.
The Gaussian bound \(thm\.[E\.2](https://arxiv.org/html/2607.08041#A5.Thmdefinition2)\) entails
I\(ϕt,Ω;φ∼P\)≤I\(ϕt,Ω;φ∼Q1\)\\displaystyle I\(\\phi\_\{t,\\Omega\};\\varphi\\sim P\)\\leq I\(\\phi\_\{t,\\Omega\};\\varphi\\sim Q\_\{1\}\)\(186\)We further have a closed formula for the mutual information of a Gaussian variable
I\(ϕt,Ω;φ∼Q1\)=Tr\[ln\(I\+Σσt2\)\]I\(\\phi\_\{t,\\Omega\};\\varphi\\sim Q\_\{1\}\)=\\Tr\[\\ln\(I\+\\frac\{\\Sigma\}\{\\sigma\_\{t\}^\{2\}\}\)\]\(187\)SinceTrln\(⋅\)Tr\\ln\(\\cdot\)is a strictly concave function anddiag\(Σ\)≺Σ\\text\{diag\}\(\\Sigma\)\\prec\\Sigmain the ordering of semi\-definite matrices,Tr\[ln\(I\+Σσt2\)\]≤Tr\[ln\(I\+Diag\(Σ\)σt2\)\]Tr\[\\ln\(I\+\\frac\{\\Sigma\}\{\\sigma\_\{t\}^\{2\}\}\)\]\\leq Tr\[\\ln\(I\+\\frac\{\\text\{Diag\}\(\\Sigma\)\}\{\\sigma\_\{t\}^\{2\}\}\)\]
I\(ϕt,Ω;φ∼Q1\)≤I\(ϕt,Ω;φ∼Q2\)\\displaystyle I\(\\phi\_\{t,\\Omega\};\\varphi\\sim Q\_\{1\}\)\\leq I\(\\phi\_\{t,\\Omega\};\\varphi\\sim Q\_\{2\}\)\(188\)From this it follows that, forL\>L′L\>L^\{\\prime\},I\(ϕt,ΩL;φ\)≥I\(ϕt,ΩL′;φ\)I\(\\phi\_\{t,\\Omega\_\{L\}\};\\varphi\)\\geq I\(\\phi\_\{t,\\Omega\_\{L^\{\\prime\}\}\};\\varphi\)\. Consequently, we find that atLc;Q1L\_\{c;Q\_\{1\}\},
ln\|𝒟\|=I\(ϕt,ΩLc;Q1;φ∼Q1\)≥I\(ϕt,ΩLc;P;φ∼P\)\\displaystyle\\ln\|\\mathcal\{D\}\|=I\(\\phi\_\{t,\\Omega\_\{L\_\{c\};Q\_\{1\}\}\};\\varphi\\sim Q\_\{1\}\)\\geq I\(\\phi\_\{t,\\Omega\_\{L\_\{c\};P\}\};\\varphi\\sim P\)\(189\)so we deduce thatLc;Q1≤Lc;PL\_\{c;Q\_\{1\}\}\\leq L\_\{c;P\}\. Applying the same argument forQ2Q\_\{2\}andQ1Q\_\{1\}yieldsLc;Q2≤Lc;Q1L\_\{c;Q\_\{2\}\}\\leq L\_\{c;Q\_\{1\}\}\.
These bounds alone are sufficient to answer our most basic question in the affirmative:for data with exactly scale\-invariant power laws, the relationship \([184](https://arxiv.org/html/2607.08041#A6.E184)\) provides a lower bound on the critical scale\. This implies that for any data distribution with translation invariance, bounded pixel\-wise variance, and a scale\-invariant power law\|k\|−2\|k\|^\{\-2\}, we haveLc\(σt\)≥CLspec\(σt\)L\_\{c\}\(\\sigma\_\{t\}\)\\geq CL\_\{spec\}\(\\sigma\_\{t\}\)for arbitrarily large image sizes even with a fixed dataset size\. We also find that these Gaussian bounds are very strong predictors of the critical scale: as shown in figure[4\(b\)](https://arxiv.org/html/2607.08041#S5.F4.sf2), the patch mutual information and concomitant critical scales for even highly non\-Gaussian datasets like MNIST tend to be well\-predicted by their Gaussian bounds\.
Beyond these Gaussian bounds, a much more general scaling theorem is available, which covers a wide range of naturalistic images:
###### Theorem F\.3
Suppose we have a translationally\-invariant fieldφ\\varphiwith bounded pointwise variance and asymptotically image\-size\-extensive entropy densitylim\|Ω\|→∞1\|Ω\|S\(φΩ\)=S~\\lim\_\{\|\\Omega\|\\to\\infty\}\\frac\{1\}\{\|\\Omega\|\}S\(\\varphi\_\{\\Omega\}\)=\\tilde\{S\}\. Then, forln\|𝒟\|\\ln\|\\mathcal\{D\}\|which isO\(1\)O\(1\)with respect toσt\\sigma\_\{t\},
Lc\(σt\)∼O\(σtln\|𝒟\|\)\\displaystyle L\_\{c\}\(\\sigma\_\{t\}\)\\sim O\\big\(\\sigma\_\{t\}\\sqrt\{\\ln\|\\mathcal\{D\}\|\}\\big\)\(190\)
###### Proof\.
The inequalityL\(σt\)≥Cσtln\|𝒟\|L\(\\sigma\_\{t\}\)\\geq C\\,\\sigma\_\{t\}\\sqrt\{\\ln\|\\mathcal\{D\}\|\}follows immediately from corollary[F\.2](https://arxiv.org/html/2607.08041#A6.Thmdefinition2)\. To show the reverse direction, we use the result
I\(ϕΩ;φ\)=S\(ϕΩ\)−S\(ϕΩ\|φ\)=S\(ϕΩ\)−S\(σtηΩ\)\\displaystyle I\(\\phi\_\{\\Omega\};\\varphi\)=S\(\\phi\_\{\\Omega\}\)\-S\(\\phi\_\{\\Omega\}\|\\varphi\)=S\(\\phi\_\{\\Omega\}\)\-S\(\\sigma\_\{t\}\\eta\_\{\\Omega\}\)\(191\)whereσtηΩ\\sigma\_\{t\}\\eta\_\{\\Omega\}is the added noise in the patchΩ\\Omega\. We can produce a lower bound on the informationS\(ϕΩ\)S\(\\phi\_\{\\Omega\}\)via the entropy power inequality, which states that for any variablesX,Y∈ℝdX,Y\\in\\mathbb\{R\}^\{d\},
S\(Y\+X\)≥d2ln\(e2dS\(Y\)\+e2dS\(X\)\)=d2ln\(1\+e2d\[S\(X\)−S\(Y\)\]\)\+S\(Y\)\\displaystyle S\(Y\+X\)\\geq\\frac\{d\}\{2\}\\ln\(e^\{\\frac\{2\}\{d\}S\(Y\)\}\+e^\{\\frac\{2\}\{d\}S\(X\)\}\)=\\frac\{d\}\{2\}\\ln\(1\+e^\{\\frac\{2\}\{d\}\[S\(X\)\-S\(Y\)\]\}\)\+S\(Y\)\(192\)Applying this withX=φΩX=\\varphi\_\{\\Omega\}andY=σηY=\\sigma\\etayields
S\(ϕΩ\)−S\(σtηΩ\)≥\|Ω\|2ln\(1\+e2\|Ω\|S\(φΩ\)2πeσt2\)\\displaystyle S\(\\phi\_\{\\Omega\}\)\-S\(\\sigma\_\{t\}\\eta\_\{\\Omega\}\)\\geq\\frac\{\|\\Omega\|\}\{2\}\\ln\\bigg\(1\+\\frac\{e^\{\\frac\{2\}\{\|\\Omega\|\}S\(\\varphi\_\{\\Omega\}\)\}\}\{2\\pi e\\sigma\_\{t\}^\{2\}\}\\bigg\.\)\(193\)Sinceσt2\\sigma\_\{t\}^\{2\}is going off to infinity andL\(σt\)≥Cσtln\|𝒟\|L\(\\sigma\_\{t\}\)\\geq C\\,\\sigma\_\{t\}\\sqrt\{\\ln\|\\mathcal\{D\}\|\},\|Ω\|=L2\|\\Omega\|=L^\{2\}is going off to infinity as well\. Therefore, we can replaceS\(φΩ\)/\|Ω\|S\(\\varphi\_\{\\Omega\}\)/\|\\Omega\|with its limit\.
I\(ϕΩ;φ\)≥\|Ω\|2ln\(1\+e2S~\+o\(1\)2πeσt2\)\\displaystyle I\(\\phi\_\{\\Omega\};\\varphi\)\\geq\\frac\{\|\\Omega\|\}\{2\}\\ln\\bigg\(1\+\\frac\{e^\{2\\tilde\{S\}\+o\(1\)\}\}\{2\\pi e\\sigma\_\{t\}^\{2\}\}\\bigg\.\)\(194\)Applying this to the collapse condition \([40](https://arxiv.org/html/2607.08041#A4.E40)\) we find that
ln\|𝒟\|≥Lc2\(σt\)2e2S~2πeσt2\+o\(1σt2\)\\displaystyle\\ln\|\\mathcal\{D\}\|\\geq\\frac\{L\_\{c\}^\{2\}\(\\sigma\_\{t\}\)\}\{2\}\\frac\{e^\{2\\tilde\{S\}\}\}\{2\\pi e\\sigma\_\{t\}^\{2\}\}\+o\(\\frac\{1\}\{\\sigma\_\{t\}^\{2\}\}\)\(195\)This is in the form of the information condition scaling for an isotropic Gaussian with pointwise variancee2S~2πe\\frac\{e^\{2\\tilde\{S\}\}\}\{2\\pi e\}, which asymptotically asσt→∞\\sigma\_\{t\}\\to\\inftyyields the result
Lc\(σt\)\+o\(1σt2\)≤2σte−S~πeln\|𝒟\|\\displaystyle L\_\{c\}\(\\sigma\_\{t\}\)\+o\(\\frac\{1\}\{\\sigma\_\{t\}^\{2\}\}\)\\leq 2\\sigma\_\{t\}\\,e^\{\-\\tilde\{S\}\}\\sqrt\{\\pi e\\ln\|\\mathcal\{D\}\|\}\(196\)This completes the proof\.
### F\.4Deviations from scale\-invariance
For power lawsα\>2\\alpha\>2, the spectral scale \([175](https://arxiv.org/html/2607.08041#A6.E175)\) increases sublinearly withσt\\sigma\_\{t\}\. The Gaussian bounds show thatLc≥CσtL\_\{c\}\\geq C\\sigma\_\{t\}, which asσt→∞\\sigma\_\{t\}\\to\\inftyshows thatLcL\_\{c\}asymptotically dominates over the spectral scale and thus is asymptotically learnable with a Bayes\-optimal local denoiser\. This case, while instructive, is less phenomenologically interesting for natural images\. For power lawsα<2\\alpha<2, the bulk entropy scales extensively, so the conditions for theorem[5\.2](https://arxiv.org/html/2607.08041#S5.Thmdefinition2)are met and we findLc∼σtln\|𝒟\|L\_\{c\}\\sim\\sigma\_\{t\}\\sqrt\{\\ln\|\\mathcal\{D\}\|\}\. The spectral scaleLspec∼σt2/αL\_\{spec\}\\sim\\sigma\_\{t\}^\{2/\\alpha\}thus grows faster withσt\\sigma\_\{t\}than the critical scale at fixedln\|𝒟\|\\ln\|\\mathcal\{D\}\|\. This implies that to match the spectral scale up to arbitrarily large image sizes, we need to scale the dataset size according to
ln\|𝒟\|∼σt2ϵ2−ϵ∼Lspecϵ\\displaystyle\\ln\|\\mathcal\{D\}\|\\sim\\sigma\_\{t\}^\{\\frac\{2\\epsilon\}\{2\-\\epsilon\}\}\\sim L\_\{spec\}^\{\\epsilon\}\(197\)While this scaling is superpolynomial with regards to image size \(and therefore experiences a mild form of the curse of dimensionality\), it is significantly subexponential, i\.e\. it significantly improves over the naive\|𝒟\|∼constd\|\\mathcal\{D\}\|\\sim const^\{d\}scaling characteristic of the empirical score function\. The right hand side in practice exhibits a very weak dependence– e\.g\., forϵ=0\.1\\epsilon=0\.1, takingLspec=32L\_\{spec\}=32toL=1024L=1024results inLspecϵ∼1\.4L\_\{spec\}^\{\\epsilon\}\\sim 1\.4toLspecϵ=2L\_\{spec\}^\{\\epsilon\}=2\. The exact impact of this scaling on the required\|𝒟\|\|\\mathcal\{D\}\|will depend on the constants of proportionality, which are in general problem\-dependent\. In our experiments, we see little evidence of thisϵ\\epsilon\-dependent growth of the spectral scale relative to the critical scale, although such growth may become more pronounced for very large\-scale images\.
### F\.5Empirical analysis of the Posterior Entropy on Real and Toy Datasets
In figure[3](https://arxiv.org/html/2607.08041#S4.F3), we show comparisons between the real entropy of the posterior distribution and the theoretical prediction for two toy models of data and two real datasets\. We take5×55\\times 5image patches from32×3232\\times 32images and calculate the posterior distribution according to \([24](https://arxiv.org/html/2607.08041#A3.E24)\) and calculate−∑φ∈𝒟P\(φ\|ϕΩx,L=5\)lnP\(φ\|ϕΩx,L=5\)\-\\sum\_\{\\varphi\\in\\mathcal\{D\}\}P\(\\varphi\|\\phi\_\{\\Omega\_\{x\},L=5\}\)\\ln P\(\\varphi\|\\phi\_\{\\Omega\_\{x\},L=5\}\)by summing over all training samples as the noise strengthσt2\\sigma\_\{t\}^\{2\}is varied between10−310^\{\-3\}and101\.510^\{1\.5\}\. The entropy is then calculated with dataset sizes28=2562^\{8\}=256,210=10242^\{10\}=1024,212=40962^\{12\}=4096, and214=2^\{14\}=16,384\. Each dataset size is plotted in its own color\. In sub\-figure \(a\), training data samples were drawn from a Gaussian distribution with pixel covariance⟨φ\(x\)φ\(y\)⟩=1\|x−y\|2−α\\langle\\varphi\(x\)\\varphi\(y\)\\rangle=\\frac\{1\}\{\|x\-y\|^\{2\-\\alpha\}\}forx≠yx\\neq ywith⟨φ\(x\)2⟩=1\\langle\\varphi\(x\)^\{2\}\\rangle=1\. The distance between neighboring pixels is set to 1\. The theory curve uses the expression for the mutual information of Gaussian variables \([134](https://arxiv.org/html/2607.08041#A5.E134)\)\. In sub\-figure \(b\), a mixture of 16 Gaussians was sampled in the fashion described in section[E\.4](https://arxiv.org/html/2607.08041#A5.SS4)\. Section[E\.4](https://arxiv.org/html/2607.08041#A5.SS4)also describes the formula for the theory curve\. All covariances were taken to be isotropic with the variance of the means equal to 1 and a variance of the Gaussians equal to 0\.15\. In sub\-figures \(c\) and \(d\), data samples were drawn uniformly from the datasets CIFAR10 and CelebA and then downscaled to32×3232\\times 32color images\. The theory curves were calculated by plugging the empirical patch covariance matrix into the equation for the Gaussian mutual information \([134](https://arxiv.org/html/2607.08041#A5.E134)\)\.
### F\.6Empirical analysis of the critical scale and spectral scale
Figure 5:A plot of the mutual informationI\(ϕΩ;φ\)I\(\\phi\_\{\\Omega\};\\varphi\)on a Gaussianized version of the BWCeleba32 dataset, as a function of both the noise levelσ2\\sigma^\{2\}and the patch lengthLL, relative to the prior for a\) the theoretical distribution b\) a finite sample of10410^\{4\}examples, as a function of noise level and patch scale\. Plotted as well is the critical thresholdI=ln\|𝒟\|I=\\ln\|\\mathcal\{D\}\|\.To test the qualitative validity of our theoretical analyses, we ran a number of experiments on realistic data distributions\. Firstly, we wanted to directly test the validity of the collapse condition \([40](https://arxiv.org/html/2607.08041#A4.E40)\)\. To do so, we first took theGaussianizedversion of an exemplar dataset \(in our case BWCeleba32\), defined as the Gaussian distribution whose mean and covariance match the first and second order statistics of the dataset\. The reason for this choice is that it furnishes a dataset with naturalistic statistics, for which the underlying patch mutual informationI\(ϕΩL,t;φ\)I\(\\phi\_\{\\Omega\_\{L\},t\};\\varphi\)can be computed analytically using the formula \([134](https://arxiv.org/html/2607.08041#A5.E134)\)\. We first computed the theoretical mutual information for a large range of length scalesLLand noise parametersσt\\sigma\_\{t\}; the results are plotted in figure[5](https://arxiv.org/html/2607.08041#A6.F5)as the blue ‘Theoretical’ surface\. As expected, these results grow without bound as the noise level decreases/as the length scale increases\. We then sampled a dataset𝒟\\mathcal\{D\}of10410^\{4\}samples from this Gaussianized prior, and evaluated the mutual information of the noised patch relative to thediscreteprior induced by this dataset; the result of this evaluation is plotted in figure[5](https://arxiv.org/html/2607.08041#A6.F5)as the orange ‘Empirical’ surface\. We found, as expected, that the empirical finite\-dataset result closely matches the theoretical infinite\-dataset result when the latter is below the critical threshold \(illustrated by the grey surface in fig\.[5](https://arxiv.org/html/2607.08041#A6.F5)\)\. When the latter exceeds the critical threshold, the empirical surface ‘peels off,’ never reaching its theoretical upper bound ofln\|𝒟\|\\ln\|\\mathcal\{D\}\|\.
To test the concordance between the critical scale and the spectral scale, we compute numerically the location of asubcritical threshold, defined byln\|𝒟\|−lnd∗=I\(ϕΩL;φ\)\\ln\|\\mathcal\{D\}\|\-\\ln d^\{\*\}=I\(\\phi\_\{\\Omega\_\{L\}\};\\varphi\)\. Intuitively, this defines the necessary amount of information to resolve down to an ‘effective sample size’ ofd∗d^\{\*\}from a dataset of size\|𝒟\|\|\\mathcal\{D\}\|\. Because this threshold is below the critical threshold, we expect the empirical estimator forI\(ϕΩL;φ\)I\(\\phi\_\{\\Omega\_\{L\}\};\\varphi\)to be robust and match the theoretical infinite\-data mutual information with a high degree of accuracy, as demonstrated in the plot \([5](https://arxiv.org/html/2607.08041#A6.F5)\) in the case where the exact mutual information is known analytically\. We taked∗=10d^\{\*\}=10for our experiments\. We compute the subcritical threshold curve by sweeping the empirical mutual information overσt\\sigma\_\{t\}andLLand computing numerically the intersection of this surface with the threshold surface\. We also compute the Gaussian bound analytically over the sameσt,L\\sigma\_\{t\},Ldomain, using the patch covariance statistics\. We find that the Gaussian bound is usually a strong predictor of the \(sub\)critical scale\. Finally, for each dataset, we extract the curve of the spectral scaleLspec\(σt\)L\_\{spec\}\(\\sigma\_\{t\}\)as a function of the noise level\. To do this, we identify at eachσt\\sigma\_\{t\}the Fourier band\|k∗\(σt\)\|\|k^\{\*\}\(\\sigma\_\{t\}\)\|wherein the average power spectral density for a mode in that band equals the added noise powerσt2\\sigma\_\{t\}^\{2\}, and takeLspec\(σt\)=\|k∗\(σt\)\|−1L\_\{spec\}\(\\sigma\_\{t\}\)=\|k^\{\*\}\(\\sigma\_\{t\}\)\|^\{\-1\}\. We generally find that this curve generally tracks the critical scale remarkably well, validating our thesis that there is a concordance between the spectral and critical scales on realistic datasets\.
Finally, to illustrate the impact of the scale on the performance of the Bayes\-optimal local denoiser, we took two images, one from the denoiser’s training set and an unrelated image from the test set, and applied noise to each at a noise level ofσt=1\\sigma\_\{t\}=1\. We then applied an optimal local denoiser to each image at length scalesLL\. At small scales \(L=5L=5\), the Bayes\-optimal local denoiser is robust, performing comparably on the training and test set, but produces poor quality images\. At the test set\-optimal scale, near but below the critical scale, \(L=7L=7, boxed in green\), the denoising estimate is both accurate and robust on both the training set and test set\. Past this scale, statistical error begins to accumulate on the denoising estimate of test image, while the model’s performance on the training image increases\. At very large scales \(L=17,33L=17,33\), the LS model denoises the training image perfectly, while failing catastrophically on the test set image\. This result is illustrated in panel[4\(a\)](https://arxiv.org/html/2607.08041#S5.F4.sf1)in figure[4](https://arxiv.org/html/2607.08041#S5.F4)\.
## Appendix GWhy scaling the patch size with the spectral scale is necessary for a local denoiser
Our work has emphasized the critical importance of using a local denoiser with a length scale that corresponds to the spectral scale\. We can make the rationale for this somewhat more precise in the context of Gaussian power\-law data\. Roughly speaking, any filter family that achieves this scale will be able to reconstruct the power spectral densityP\(k\)∼\|k\|−αP\(k\)\\sim\|k\|^\{\-\\alpha\}of the target distribution; conversely, a filter family that scales according to a different law will produce data distributed according to a different power law spectral density\.
For a power\-law Gaussian distribution with power spectral densityP\(k\)=A\|k\|−αP\(k\)=A\|k\|^\{\-\\alpha\}, the Wiener filter in Fourier space is given by
w^\(σ2,k\)=11\+A−1σ2\|k\|α\\displaystyle\\hat\{w\}\(\\sigma^\{2\},k\)=\\frac\{1\}\{1\+A^\{\-1\}\\sigma^\{2\}\|k\|^\{\\alpha\}\}\(198\)This satisfies an interesting invariance condition, which we call therenormalization condition:
w^\(βσ2,k\)=w^\(σ2,β1/αk\)\\displaystyle\\hat\{w\}\(\\beta\\sigma^\{2\},k\)=\\hat\{w\}\(\\sigma^\{2\},\\beta^\{1/\\alpha\}k\)\(199\)In real space, this corresponds to the condition
w\(βσ2,x\)=β−2/αw\(σ2,β−1/αx\)\\displaystyle w\(\\beta\\sigma^\{2\},x\)=\\beta^\{\-2/\\alpha\}w\(\\sigma^\{2\},\\beta^\{\-1/\\alpha\}x\)\(200\)The content of this statement is that the family of Wiener filters is scale\-equivariant: the filter used for a given noise levelσ2\\sigma^\{2\}is simply a rescaled and dilated version of the filter at any other given noise level\. The Wiener filter also satisfies two other important properties: it isnormalized\(i\.e\. it integrates to 1\), and it goes to 0 in Fourier space ask→∞k\\to\\infty\(spectralfalloff\)\. The normalization property in particular is important for a consistent denoiser: it entails in particular that as the noise level goes to 0, the estimate of the image produced by the filter converges to the true image, as opposed to a rescaled version of it\.
In our paper, we have focused on the question ofstrictly local denoising, where the filter mass is constrained to lie entirely within a particular noise\-dependent radius, which we expect to be proportional to the spectral scale\. We might wonder what sort of distributions the resulting optimal localized filters generate\. It turns out that the three properties ofrenormalization,normalization, andfalloff, on their own are enough to guarantee that a diffusion model, parameterized by a linear denoiser withanyfilter shapem^\(σ2,k\)\\hat\{m\}\(\\sigma^\{2\},k\), will obtain the correct spectral power law∼\|k\|−α\\sim\|k\|^\{\-\\alpha\}\(with a uniform error, relative to the optimal filter, on the global amplitudeAA\)\. Conversely, if the filter family satisfies the renormalization property with a renormalization condition that is for adifferentpower valueα′≠α\\alpha^\{\\prime\}\\neq\\alpha\(e\.g\. if it is scale\-equivariant but localized to a radius that isnotproportional to the spectral scale\), this filter family will generate a power law with anincorrectexponent\. This result is captured below:
###### Theorem G\.1
Suppose we parameterize a diffusion model with a linear denoising model
Mt\[ϕt\]\(x\)=m\(σt2,x\)⋆ϕt\(x\)\\displaystyle M\_\{t\}\[\\phi\_\{t\}\]\(x\)=m\(\\sigma^\{2\}\_\{t\},x\)\\,\\,\\star\\,\\,\\phi\_\{t\}\(x\)\(201\)wherem\(σ2,x\)m\(\\sigma^\{2\},x\)is a filter whose Fourier transform \(with respect toxx\) satisfies the renormalization condition
m^\(βσ2,k\)=m^\(σ2,β1/αk\)\\displaystyle\\hat\{m\}\(\\beta\\sigma^\{2\},k\)=\\hat\{m\}\(\\sigma^\{2\},\\beta^\{1/\\alpha\}k\)\(202\)the normalization condition
∫m\(σ2,x\)𝑑x=1\\displaystyle\\int m\(\\sigma^\{2\},x\)\\,dx=1\(203\)and two falloff conditions for somep,q\>0p,q\>0\.
lim\|k\|→∞\|k\|pm^\(σ2,k\)=0\\displaystyle\\lim\_\{\|k\|\\to\\infty\}\|k\|^\{p\}\\hat\{m\}\(\\sigma^\{2\},k\)=0lim\|k\|→0\|k\|−q\(m^\(σ2,k\)−1\)=0\\displaystyle\\lim\_\{\|k\|\\to 0\}\|k\|^\{\-q\}\(\\hat\{m\}\(\\sigma^\{2\},k\)\-1\)=0These falloff conditions can also be written in terms ofσ2\\sigma^\{2\}for somep′,q′\>0p^\{\\prime\},q^\{\\prime\}\>0\.
limσ2→∞\(σ\)2p′m^\(σ2,k\)=0\\displaystyle\\lim\_\{\\sigma^\{2\}\\to\\infty\}\(\\sigma\)^\{2p^\{\\prime\}\}\\hat\{m\}\(\\sigma^\{2\},k\)=0limσ2→0\(σ\)−2q′\(m^\(σ2,k\)−1\)=0\\displaystyle\\lim\_\{\\sigma^\{2\}\\to 0\}\(\\sigma\)^\{\-2q^\{\\prime\}\}\(\\hat\{m\}\(\\sigma^\{2\},k\)\-1\)=0Then the diffusion model will generate a Gaussian distribution with a power spectral density
P\(k\)∼\|k\|−α\\displaystyle P\(k\)\\sim\|k\|^\{\-\\alpha\}\(204\)
###### Proof\.
For the purposes of this proof, we will use the ‘variance\-exploding’ parameterization:
ϕt\\displaystyle\\phi\_\{t\}=ϕ0\+σtη\\displaystyle=\\phi\_\{0\}\+\\sigma\_\{t\}\\eta\(205\)whereσt→∞\\sigma\_\{t\}\\to\\inftyast→∞t\\to\\infty\. An analogous proof should be available using other parameterizations, such as the variance\-preserving parameterization\. The probability flow equation for this parameterization is
ddtϕt\\displaystyle\\frac\{d\}\{dt\}\\phi\_\{t\}=12\(∂tlogσt2\)\(ϕt−Mt\[ϕt\]\)\\displaystyle=\\frac\{1\}\{2\}\(\\partial\_\{t\}\\log\\sigma\_\{t\}^\{2\}\)\(\\phi\_\{t\}\-M\_\{t\}\[\\phi\_\{t\}\]\)\(206\)which, for the linear denoiser \([201](https://arxiv.org/html/2607.08041#A7.E201)\), reduces to
ddtϕt\\displaystyle\\frac\{d\}\{dt\}\\phi\_\{t\}=12\(∂tlogσt2\)\(1−\(m\(σt2,x\)⋆\)\)ϕt\\displaystyle=\\frac\{1\}\{2\}\(\\partial\_\{t\}\\log\\sigma\_\{t\}^\{2\}\)\\Big\(1\-\(m\(\\sigma^\{2\}\_\{t\},x\)\\,\\star\)\\Big\)\\phi\_\{t\}\(207\)The linear operator\(1−m\(σt2,x\)⋆\)\)\(1\-m\(\\sigma^\{2\}\_\{t\},x\)\\,\\star\)\)is diagonalized in the Fourier basis, so we can solve \([207](https://arxiv.org/html/2607.08041#A7.E207)\) by solving the evolution of each Fourier componentϕ^t\(k\)\\hat\{\\phi\}\_\{t\}\(k\)independently\. This yields
∂tϕ^t\(k\)\\displaystyle\\partial\_\{t\}\\hat\{\\phi\}\_\{t\}\(k\)=12\(∂tlogσt2\)\(1−m^\(σ2,k\)\)ϕ^t\(k\)\\displaystyle=\\frac\{1\}\{2\}\(\\partial\_\{t\}\\log\\sigma\_\{t\}^\{2\}\)\(1\-\\hat\{m\}\(\\sigma^\{2\},k\)\)\\hat\{\\phi\}\_\{t\}\(k\)\(208\)Integrating this equation gives
ϕ^T\(k\)\\displaystyle\\hat\{\\phi\}\_\{T\}\(k\)=ϕ^0\(k\)exp\(12∫0T\(∂tlogσt2\)\(1−m^\(σt2,k\)\)𝑑t\)\\displaystyle=\\hat\{\\phi\}\_\{0\}\(k\)\\exp\(\\frac\{1\}\{2\}\\int\_\{0\}^\{T\}\(\\partial\_\{t\}\\log\\sigma\_\{t\}^\{2\}\)\(1\-\\hat\{m\}\(\\sigma\_\{t\}^\{2\},k\)\)\\,dt\)\(209\)Or, reparameterizing with the substitutionu=logσt2u=\\log\\sigma\_\{t\}^\{2\},
ϕ^T\(k\)=ϕ^0\(k\)exp\(12∫0logσT2\(1−m^\(u,k\)\)𝑑u\)\\displaystyle\\hat\{\\phi\}\_\{T\}\(k\)=\\hat\{\\phi\}\_\{0\}\(k\)\\exp\(\\frac\{1\}\{2\}\\int\_\{0\}^\{\\log\\sigma\_\{T\}^\{2\}\}\\Big\(1\-\\hat\{m\}\(u,k\)\\Big\)\\,du\)\(210\)In variance\-exploding flows, the samplesϕ0\\phi\_\{0\}are generated by taking samplesϕT∼σTη\\phi\_\{T\}\\sim\\sigma\_\{T\}\\etaand flowing these samples backwards\. We will consider the limit whereT→∞T\\to\\infty\. Substituting this into the equation \([210](https://arxiv.org/html/2607.08041#A7.E210)\) and rearranging, we find that the variance ofϕ^0\(k\)\\hat\{\\phi\}\_\{0\}\(k\)\(i\.e\. the power spectral densityP\(k\)P\(k\)\) satisfies
P\(k\)=limT→∞σT2exp\(−∫−∞logσT2\(1−m^\(u,k\)\)𝑑u\)\\displaystyle P\(k\)=\\lim\_\{T\\to\\infty\}\\sigma\_\{T\}^\{2\}\\exp\(\-\\int\_\{\-\\infty\}^\{\\log\\sigma\_\{T\}^\{2\}\}\\Big\(1\-\\hat\{m\}\(u,k\)\\Big\)\\,du\)\(211\)or in logarithmic terms
logP\(k\)\\displaystyle\\log P\(k\)=limT→∞\[logσT2−∫−∞logσT2\(1−m^\(eu,k\)\)𝑑u\]\\displaystyle=\\lim\_\{T\\to\\infty\}\\bigg\[\\log\\sigma\_\{T\}^\{2\}\-\\int\_\{\-\\infty\}^\{\\log\\sigma\_\{T\}^\{2\}\}\(1\-\\hat\{m\}\(e^\{u\},k\)\)\\,du\\bigg\]\(212\)=∫0∞m^\(eu,k\)𝑑u\+∫−∞0\(m^\(eu,k\)−1\)𝑑u\\displaystyle=\\int\_\{0\}^\{\\infty\}\\hat\{m\}\(e^\{u\},k\)du\+\\int\_\{\-\\infty\}^\{0\}\(\\hat\{m\}\(e^\{u\},k\)\-1\)du\(213\)One can easily verify that the falloff conditions imply both of these integrals converge by the power test\. To show the expected PSD \([204](https://arxiv.org/html/2607.08041#A7.E204)\) holds for this distribution, we can differentiate this quantity with respect to\|k\|\|k\|\. The dominated convergence theorem allows us to swap integration and differentiation to obtain
∂log\|k\|logP\(k\)=∫−∞∞∂log\|k\|m^\(eu,k\)du\\displaystyle\\partial\_\{\\log\|k\|\}\\log P\(k\)=\\int\_\{\-\\infty\}^\{\\infty\}\\partial\_\{\\log\|k\|\}\\hat\{m\}\(e^\{u\},k\)\\,du\(214\)The renormalization condition \([202](https://arxiv.org/html/2607.08041#A7.E202)\) implies that we can interchange the derivative of the filter with respect to\|k\|\|k\|with a derivative with respect touu:
∂log\|k\|m^\(σ2,k\)=α∂logσ2m^\(σ2,k\)\\displaystyle\\partial\_\{\\log\|k\|\}\\hat\{m\}\(\\sigma^\{2\},k\)=\\alpha\\,\\partial\_\{\\log\\sigma^\{2\}\}\\,\\hat\{m\}\(\\sigma^\{2\},k\)\(215\)Substituting this into \([214](https://arxiv.org/html/2607.08041#A7.E214)\) and evaluating the integral with respect to u, we obtain
∂log\|k\|logP\(k\)=α\[m^\(σ2=∞,k\)−m^\(σ2=0,k\)\]\\displaystyle\\partial\_\{\\log\|k\|\}\\log P\(k\)=\\alpha\[\\hat\{m\}\(\\sigma^\{2\}=\\infty,k\)\-\\hat\{m\}\(\\sigma^\{2\}=0,k\)\]\(216\)The renormalization condition implies thatm^\(σ2=0,k\)=m^\(σ2,k=0\)\\hat\{m\}\(\\sigma^\{2\}=0,k\)=\\hat\{m\}\(\\sigma^\{2\},k=0\); applying the normalization condition \([203](https://arxiv.org/html/2607.08041#A7.E203)\) yieldsm^\(σ2=0,k\)=1\\hat\{m\}\(\\sigma^\{2\}=0,k\)=1\. Meanwhile, the renormalization condition implies the asymptotic falloff ofm^\(σT2,k\)\\hat\{m\}\(\\sigma\_\{T\}^\{2\},k\)asσT2→∞\\sigma\_\{T\}^\{2\}\\to\\inftycoincides with the limitlim\|k\|→∞m^\(σ2,k\)=0\\lim\_\{\|k\|\\to\\infty\}\\hat\{m\}\(\\sigma^\{2\},k\)=0\. Thus \([216](https://arxiv.org/html/2607.08041#A7.E216)\) evaluates to
∂log\|k\|logP\(k\)=α\(0−1\)=−α\\displaystyle\\partial\_\{\\log\|k\|\}\\log P\(k\)=\\alpha\(0\-1\)=\-\\alpha\(217\)which proves the claim\.
To complete our analysis, we need to prove that, for a power\-law Gaussian, the conditional expectation𝔼\[φ\|ϕΩL\(σ2\),x\]\\mathbb\{E\}\[\\varphi\|\\phi\_\{\\Omega\_\{L\(\\sigma^\{2\}\),x\}\}\], withL\(σ2\)∝Lspec\(σ2\)L\(\\sigma^\{2\}\)\\propto L\_\{spec\}\(\\sigma^\{2\}\), satisfies the three properties\. This is indeed the case, as we show in the following theorem\.
###### Theorem G\.2
Forφ\\varphian image drawn from a power\-law Gaussian with PSD\|k\|−α\|k\|^\{\-\\alpha\}, the modelMt\[ϕt\]=𝔼\[φ\|ϕΩL\(σ2\),x\]=m\(σ2,x\)⋆ϕM\_\{t\}\[\\phi\_\{t\}\]=\\mathbb\{E\}\[\\varphi\|\\phi\_\{\\Omega\_\{L\(\\sigma^\{2\}\),x\}\}\]=m\(\\sigma^\{2\},x\)\\star\\phi, withL\(σ2\)∝σ2/αL\(\\sigma^\{2\}\)\\propto\\sigma^\{2/\\alpha\}, defines a filter familym\(σ2,x\)m\(\\sigma^\{2\},x\)that satisfies the normalization and falloff conditions, as well as the renormalization condition with parameterα\\alpha\.
###### Proof\.
The optimal linear denoising model given a patchΩ\\Omegafor its center pixelxxis given by
\(ΣΩ,x\)⋅\(ΣΩ,Ω\+σt2I\)−1ϕt\\displaystyle\(\\Sigma\_\{\\Omega,x\}\)\\cdot\(\\Sigma\_\{\\Omega,\\Omega\}\+\\sigma\_\{t\}^\{2\}I\)^\{\-1\}\\phi\_\{t\}\(218\)whereΣΩΩ\\Sigma\_\{\\Omega\\Omega\}is the block covariance \(under the underlying distributionP0\(φ\)P\_\{0\}\(\\varphi\)for the pixels withinΩ\\Omega, andΣΩ,x\\Sigma\_\{\\Omega,x\}is the row of this block covariance associated with the covariance between each pixel and the center pixel\. LettingC\(x\)∝\|x\|α−2C\(x\)\\propto\|x\|^\{\\alpha\-2\}be the two\-point correlation function of the field \(whose Fourier transform is the power spectral densityP\(k\)P\(k\)\), the operatorΣΩΩ\+σt2I\\Sigma\_\{\\Omega\\Omega\}\+\\sigma\_\{t\}^\{2\}Ican be understood as the convolution operator with the kernel
1\(y∈Ω\)C\(y\)\+σt2δ\(y\)\\displaystyle\\textbf\{1\}\(y\\in\\Omega\)C\(y\)\+\\sigma^\{2\}\_\{t\}\\delta\(y\)\(219\)whileΣΩ,x\\Sigma\_\{\\Omega,x\}corresponds to the functionC\(y−x\)C\(y\-x\)\. In Fourier space, the resulting filterm^\(σ2,k\)\\hat\{m\}\(\\sigma^\{2\},k\)is given by
m^\(σ2,k\)=P\(k\)⋆1^ΩP\(k\)⋆1^Ω\+σt2=11\+σt2\[P\(k\)⋆1^Ω\]\(k\)−1\\displaystyle\\hat\{m\}\(\\sigma^\{2\},k\)=\\frac\{P\(k\)\\star\\hat\{\\textbf\{1\}\}\_\{\\Omega\}\}\{P\(k\)\\star\\hat\{\\textbf\{1\}\}\_\{\\Omega\}\+\\sigma^\{2\}\_\{t\}\}=\\frac\{1\}\{1\+\\sigma\_\{t\}^\{2\}\[P\(k\)\\star\\hat\{\\textbf\{1\}\}\_\{\\Omega\}\]\(k\)^\{\-1\}\}\(220\)where1^Ω\\hat\{\\textbf\{1\}\}\_\{\\Omega\}is the Fourier transform of the indicator function1\(y∈Ω\)\\textbf\{1\}\(y\\in\\Omega\), so thatP\(k\)⋆1^ΩP\(k\)\\star\\hat\{\\textbf\{1\}\}\_\{\\Omega\}is the Fourier transform ofC1ΩC\\textbf\{1\}\_\{\\Omega\}\. We will now assume thatΩL\(σ2\)\\Omega\_\{L\(\\sigma^\{2\}\)\}is a patch whose geometry scales according to the noise level with the length scaleL\(σ2\)=σt2/αL\(\\sigma^\{2\}\)=\\sigma\_\{t\}^\{2/\\alpha\}\. This entails that the indicator function satisfies
1ΩL\(βσ2\)\(y\)=1ΩL\(σ2\)\(β−1/αy\)\\displaystyle\\textbf\{1\}\_\{\\Omega\_\{L\(\\beta\\sigma^\{2\}\)\}\}\(y\)=\\textbf\{1\}\_\{\\Omega\_\{L\(\\sigma^\{2\}\)\}\}\(\\beta^\{\-1/\\alpha\}y\)\(221\)or, in Fourier space,
1^ΩL\(βσ2\)\(k\)=β2/α1^ΩL\(σ2\)\(β1/αk\)\\displaystyle\\hat\{\\textbf\{1\}\}\_\{\\Omega\_\{L\(\\beta\\sigma^\{2\}\)\}\}\(k\)=\\beta^\{2/\\alpha\}\\hat\{\\textbf\{1\}\}\_\{\\Omega\_\{L\(\\sigma^\{2\}\)\}\}\(\\beta^\{1/\\alpha\}k\)\(222\)From this, we observe
\[P^⋆1^ΩL\(σ2\)\]\(β1/αk\)\\displaystyle\[\\hat\{P\}\\star\\hat\{\\textbf\{1\}\}\_\{\\Omega\_\{L\(\\sigma^\{2\}\)\}\}\]\(\\beta^\{1/\\alpha\}k\)=∫P^\(β1/αk−v\)1^ΩL\(σ2\)\(v\)𝑑v\\displaystyle=\\int\\hat\{P\}\(\\beta^\{1/\\alpha\}k\-v\)\\hat\{\\textbf\{1\}\}\_\{\\Omega\_\{L\(\\sigma^\{2\}\)\}\}\(v\)\\,dv=β2/α∫P\(β1/α\(k−u\)\)1^ΩL\(σ2\)\(β1/αu\)𝑑u\\displaystyle=\\beta^\{2/\\alpha\}\\int P\(\\beta^\{1/\\alpha\}\(k\-u\)\)\\hat\{\\textbf\{1\}\}\_\{\\Omega\_\{L\(\\sigma^\{2\}\)\}\}\(\\beta^\{1/\\alpha\}u\)\\,duUsing \([222](https://arxiv.org/html/2607.08041#A7.E222)\) and the propertyP\(β1/αk\)=P\(k\)β−1P\(\\beta^\{1/\\alpha\}k\)=P\(k\)\\beta^\{\-1\}, this becomes
\[P^⋆1^ΩL\(σ2\)\]\(β1/αk\)=β−1\[P^⋆1^Ω\(βL\(σ2\)\)\]\(k\)\\displaystyle\[\\hat\{P\}\\star\\hat\{\\textbf\{1\}\}\_\{\\Omega\_\{L\(\\sigma^\{2\}\)\}\}\]\(\\beta^\{1/\\alpha\}k\)=\\beta^\{\-1\}\[\\hat\{P\}\\star\\hat\{\\textbf\{1\}\}\_\{\\Omega\(\\beta L\(\\sigma^\{2\}\)\)\}\]\(k\)\(223\)From this we find that
m^\(βσ2,k\)\\displaystyle\\hat\{m\}\(\\beta\\sigma^\{2\},k\)=\(1\+\(βσ2\)\[P^⋆1^Ω\(βL\(σ2\)\)\]\(k\)\)−1=\(1\+\(βσ2\)β\[P^⋆1^Ω\(L\(σ2\)\)\]\(β1/αk\)\)−1\\displaystyle=\\bigg\(1\+\\frac\{\(\\beta\\sigma^\{2\}\)\}\{\[\\hat\{P\}\\star\\hat\{\\textbf\{1\}\}\_\{\\Omega\(\\beta L\(\\sigma^\{2\}\)\)\}\]\(k\)\}\\bigg\)^\{\-1\}=\\bigg\(1\+\\frac\{\(\\beta\\sigma^\{2\}\)\}\{\\beta\[\\hat\{P\}\\star\\hat\{\\textbf\{1\}\}\_\{\\Omega\(L\(\\sigma^\{2\}\)\)\}\]\(\\beta^\{1/\\alpha\}k\)\}\\bigg\)^\{\-1\}\(224\)=\(1\+σ2\[P^⋆1^Ω\(L\(σ2\)\)\]\(β1/αk\)\)−1=m^\(σ2,β1/αk\)\\displaystyle=\\bigg\(1\+\\frac\{\\sigma^\{2\}\}\{\[\\hat\{P\}\\star\\hat\{\\textbf\{1\}\}\_\{\\Omega\(L\(\\sigma^\{2\}\)\)\}\]\(\\beta^\{1/\\alpha\}k\)\}\\bigg\)^\{\-1\}=\\hat\{m\}\(\\sigma^\{2\},\\beta^\{1/\\alpha\}k\)\(225\)Som^\\hat\{m\}satisfies the renormalization condition with parameterα\\alpha\. Sincelimσ2→0m\(σ2,k\)=1\\lim\_\{\\sigma^\{2\}\\to 0\}m\(\\sigma^\{2\},k\)=1andlimσ2→∞m\(σ2,k\)=0\\lim\_\{\\sigma^\{2\}\\to\\infty\}m\(\\sigma^\{2\},k\)=0, it follows from the renormalization conditions that the falloff and normalization conditions are satisfied\.
Our results do not establish the result that𝔼\[φ\|ϕΩL\(σ2\),x\]\\mathbb\{E\}\[\\varphi\|\\phi\_\{\\Omega\_\{L\(\\sigma^\{2\}\),x\}\}\], withL\(σ2\)L\(\\sigma^\{2\}\)parameterized by anincorrectpower law∼σ2α′\\sim\\sigma^\{\\frac\{2\}\{\\alpha^\{\\prime\}\}\}withα′≠α\\alpha^\{\\prime\}\\neq\\alpha, will generate an incorrect power law in the power spectral density, because the restricted statistics are not appropriately self\-similar with noise, and therefore this estimator does not define a scale\-equivariant family of filters satisfying the renormalization condition\. We generally expect, however that if the length scale is significantly less than the spectral scale, that this family will nonetheless generate a power law that isasymptoticallyincorrect at large\|k\|\|k\|\.
## Appendix HEmpirical Results in Early Training
### H\.1Methods




Figure 6:Plots of the Medianr2r^\{2\}metric \(left\) and the MSE metric \(right\) averaged over comparisons between several independently trained DiTs \(top two panels\) and UNets \(bottom two panels\) and their corresponding analytical models, across four datasets Celeba64, FashionMNIST, CIFAR10, and MNIST\. Each model is trained on a reduced dataset of10410^\{4\}samples\. The Medianr2r^\{2\}metric starts high across all datasets and gradually decreases over the course of training, while the MSE generally starts low and then gradually increases over the course of training\. On some datasets these metrics are more level than others, or even decreasing as in MSE on FashionMNIST; however, the generally decreasing trend is mostly consistent\.Table 1:Quantitative agreement between DiTs trained on a subset of size10410^\{4\}and corresponding \(LS/ELS\) analytical models given the same dataset subset\. Medianr2r^\{2\}values and MSE values are computed using a fixed subset of 100 samples and averaged over 4 model trials\. Error bars are given by2σ/n2\\sigma/\\sqrt\{n\}forn=400n=400the sample size andσ\\sigmathe sample standard deviation for the given statistic\. Selected metrics are taken from 30 epochs and 400 epochs respectively; the full trajectory of the medianr2r^\{2\}and MSE losses are shown in plots[6](https://arxiv.org/html/2607.08041#A8.F6)\.Table 2:Quantitative agreement between UNets trained on a subset of size10410^\{4\}and corresponding \(LS/ELS\) analytical models given the same dataset subset\. Medianr2r^\{2\}values and MSE values are computed using a fixed subset of 100 samples and averaged over 4 model trials\. Error bars are given by2σ/n2\\sigma/\\sqrt\{n\}forn=400n=400the sample size andσ\\sigmathe sample standard deviation for the given statistic\. Selected metrics are taken from 30 epochs and 400 epochs respectively; the full trajectory of the medianr2r^\{2\}and MSE losses are shown in plots[6](https://arxiv.org/html/2607.08041#A8.F6)\.In all experiments, we use a cosine noise schedule and a uniformly discretized reverse process\. Our neural networks diffusion models employ a 100\-step reverse process\. Due to the high computational cost of the analytical denoisers, we employ a 20\-step reverse process for each of our analytical diffusion models\. We used an H100 GPU node with 80GB RAM for all experiments\.
To calibrate the locality scale in our experiments, rather than calibrating based on agreement with the neural network\[[13](https://arxiv.org/html/2607.08041#bib.bib13),[14](https://arxiv.org/html/2607.08041#bib.bib14)\]or using an absolute threshold on the Wiener filter\[[32](https://arxiv.org/html/2607.08041#bib.bib32)\], we simply select a validation set \(∼2000\\sim 2000images\) from the training set, and select a patch scale at each time step of denoising based on thebest validation lossfor the analytical denoising model employing the remainder of the dataset\. This naturally selects patch sizes around the criticality threshold, which is the largest scale \(and therefore most informative\) before the onset of catastrophic statistical error \(fig\.[4\(a\)](https://arxiv.org/html/2607.08041#S5.F4.sf1)\)\. For MNIST, FashionMNIST, and Celeba64, we use an LS model with patch geometries chosen using the Wiener filter binarization of\[[32](https://arxiv.org/html/2607.08041#bib.bib32)\]\. However, our procedure for picking the patch length differs from that used in this paper because we are only selecting thefamilyof patches given byΩk\\Omega\_\{k\}representing the mask of top\-kkcoefficients of the Wiener filter; however, rather than settingkkusing the coefficients of the Wiener filter, we calibrate it \(independently for each pixel\) using the cross validation procedure described above\. We start our calibration at the lowest noise level, and incrementally increase the patch sizekkin increments ofmax\(4,⌊0\.05k⌋\)\\operatorname\{max\}\(4,\\lfloor 0\.05k\\rfloor\)pixels until the validation loss at that pixel location fails to decrease\. We make the assumption based on our theoretical analysis that the optimal scale should be nondecreasing with noise level, and thus initialize the scale at the next noise level at the optimal scale from the previous iteration\. For CIFAR10, we find instead that the better analytical model at early training is the ELS rather than LS\. We use the boundary\-broken ELS prescription given in\[[13](https://arxiv.org/html/2607.08041#bib.bib13)\]with a square patch geometry, and calibrate the scale of this patch using the same procedure described above\. For the ELS, however, due to computational constraints, we only use one global scale, rather than a pixel\-specific scale\.
For our neural network experiments, we use a DiT with a patch size of4×44\\times 4, hidden dimension of 512, 8 attention heads, and a 4\-to\-1 MLP ratio per block, with 8 transformer blocks, as well as an input and output convolutional stem\. We use a UNet with four scales, with channel dimensions \(128, 256, 256, 512\) at each scale, and use an attention layer with 8 heads at the 16x16 image resolution layers\. We train the models using AdamW with betas \(0\.9,0\.95\) and weight decay of 5e\-2\. We use an initial learning rate of 2e\-4, with cosine annealing\. We train for 400 epochs on dataset subsets of size10410^\{4\}with a batch size of 64 for the DiTs and 32 for the UNets\.
One of the difficult features of early\-training diffusion models is that, at standard \(large\) learning rates, their outputs have a propensity to oscillate\. The global hues of generated images are particularly strongly affected by this oscillation, which can significantly damage the appearance of early\-training consistency if not managed\. This instability phenomenon has been noted elsewhere\[[49](https://arxiv.org/html/2607.08041#bib.bib49),[50](https://arxiv.org/html/2607.08041#bib.bib50),[51](https://arxiv.org/html/2607.08041#bib.bib51)\], and a number of techniques have been proposed to mitigate this effect\. To get consistent results in early training, we find best results are obtained combining the following two techniques:
- •Weight averaging:exponential moving average weighting of past checkpoints is a commonly employed technique in diffusion modeling to improve checkpoint performance\. We use a slightly simpler approach of linearly weight averaging the checkpoints taken from the end of each of the 5 previous epochs\. We find that this helps average out fluctuations and stabilize the model’s output significantly\.
- •vv\-prediction:this technique was designed\[[52](https://arxiv.org/html/2607.08041#bib.bib52)\]to keep the target variance consistent at all noise levels\. We find in our work that this choice significantly stabilizes early training outputs, relative to other standard parameterizations \(ϵ^\\hat\{\\epsilon\}\- andx^0\\hat\{x\}\_\{0\}\-prediction\)\.
Following the literature on analytical diffusion models\[[13](https://arxiv.org/html/2607.08041#bib.bib13),[14](https://arxiv.org/html/2607.08041#bib.bib14),[32](https://arxiv.org/html/2607.08041#bib.bib32)\], we use pixel\-space metrics \(median pixelwiser2r^\{2\}and MSE\) to characterize the distance between our neural and analytical model outputs\. While the quantitative signals we use show the effects of increased training in the form of slightly worse scores, the effect is somewhat less evident in the statistics than the qualitative perceptual effect would suggest\. We do not propose a new metric here, but feel that there is room to develop metrics that capture the more perceptually and phenomenologically relevant dimensions of model training\. In our table[1](https://arxiv.org/html/2607.08041#A8.T1)and our plots[6](https://arxiv.org/html/2607.08041#A8.F6)we report the average of these metrics taken across four iterations and both subsets𝒟1\\mathcal\{D\}\_\{1\}and𝒟2\\mathcal\{D\}\_\{2\}that we are training with\.
### H\.2Samples
We find empirically that the rate of training is somewhat different depending on the architecture and dataset, with maximal similarity to the outputs of the analytical models not always occurring at the same point \(see fig[6](https://arxiv.org/html/2607.08041#A8.F6)for the evolution of the quantitative metrics\)\. In the samples below, we highlight samples taken from the epoch showing approximately the point of maximum agreement between theory and experiment for each diffusion model/dataset; examples of the evolution of the samples further in training are shown in fig[11](https://arxiv.org/html/2607.08041#A8.F11)\. For each sample figure including[2](https://arxiv.org/html/2607.08041#S3.F2), the samples for the epochs displayed are:
- •MNIST: UNet, 10; DiT, 15\.
- •FashionMNIST: UNet, 10; DiT, 15\.
- •CIFAR10: UNet, 10; DiT, 30\.
- •Celeba64: UNet, 10; DiT, 30\.
These epochs are on the reduced datasets of size10410^\{4\}used in training, so the equivalent number of epochs on a larger dataset for peak agreement would likely be smaller\.
𝒟2𝒟1\\mathcal\{D\}\_\{2\}\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\mathcal\{D\}\_\{1\}


𝒟2𝒟1\\mathcal\{D\}\_\{2\}\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\mathcal\{D\}\_\{1\}


𝒟2𝒟1\\mathcal\{D\}\_\{2\}\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\mathcal\{D\}\_\{1\}


Figure 7:Further samples comparing DiTs and UNets early in training, trained on two disjoint subsets of10410^\{4\}samples from Celeba64, and the outputs of calibrated local score models\. UNet samples are from 10 epochs of training, DiT samples are from 30 epochs of training\. Samples not curated for quality\.𝒟2𝒟1\\mathcal\{D\}\_\{2\}\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\mathcal\{D\}\_\{1\}


𝒟2𝒟1\\mathcal\{D\}\_\{2\}\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\mathcal\{D\}\_\{1\}


𝒟2𝒟1\\mathcal\{D\}\_\{2\}\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\mathcal\{D\}\_\{1\}


Figure 8:Further samples comparing DiTs and UNets early in training, trained on two disjoint subsets of10410^\{4\}samples from CIFAR10, and the outputs of calibrated local score models\. UNet samples are from 10 epochs of training, DiT samples are from 30 epochs of training\. Samples not curated for quality\.𝒟2𝒟1\\mathcal\{D\}\_\{2\}\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\mathcal\{D\}\_\{1\}


𝒟2𝒟1\\mathcal\{D\}\_\{2\}\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\mathcal\{D\}\_\{1\}


𝒟2𝒟1\\mathcal\{D\}\_\{2\}\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\mathcal\{D\}\_\{1\}


Figure 9:Further samples comparing DiTs and UNets early in training, trained on two disjoint subsets of10410^\{4\}samples from FashionMNIST, and the outputs of calibrated local score models\. UNet samples are from 10 epochs of training, DiT samples are from 15 epochs of training\. Samples not curated for quality\.𝒟2𝒟1\\mathcal\{D\}\_\{2\}\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\mathcal\{D\}\_\{1\}


𝒟2𝒟1\\mathcal\{D\}\_\{2\}\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\mathcal\{D\}\_\{1\}


𝒟2𝒟1\\mathcal\{D\}\_\{2\}\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\mathcal\{D\}\_\{1\}


Figure 10:Further samples comparing DiTs and UNets early in training, trained on two disjoint subsets of10410^\{4\}samples from MNIST, and the outputs of calibrated local score models\. UNet samples are from 10 epochs of training, DiT samples are from 15 epochs of training\. Samples not curated for quality\.𝒟2𝒟1\\mathcal\{D\}\_\{2\}\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\mathcal\{D\}\_\{1\}\\,\\,\\,

𝒟2𝒟1\\mathcal\{D\}\_\{2\}\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\,\\mathcal\{D\}\_\{1\}\\,\\,\\,

Figure 11:As training progresses, models evolve from ‘patch mosaic’ style outputs\[[13](https://arxiv.org/html/2607.08041#bib.bib13)\]towards semantically coherent, nonlocally consistent generations, which nevertheless remain robust to dataset shifts\. Leftmost columns are analytical, middle columns are DiT after 30 epochs on a10410^\{4\}sample subset, and rightmost is after 300 epochs on a10410^\{4\}sample subset\.Similar Articles
From Privacy to Generalization: Linear Max-Information Bounds for DP-SGD
This paper proves a finite-sample bound on the approximate max-information of DP-SGD that is at most linear in dataset size, yielding PAC-Bayes generalization bounds for models trained with differential privacy.
Bounded-Rationality, Hedging, and Generalization
This paper studies generalization in learning through the lens of bounded-rational decision theory, where the learner's response law induces a tradeoff between training loss and sample dependence. The authors show that this tradeoff is governed by an f-divergence regularizer and that generalization can be certified from the learner's hedging behavior.
Elucidating the SNR-t Bias of Diffusion Probabilistic Models
This paper identifies a Signal-to-Noise Ratio timestep (SNR-t) bias in diffusion probabilistic models during inference, where SNR-timestep alignment from training is disrupted at inference time. The authors propose a differential correction method that decomposes samples into frequency components and corrects each separately, improving generation quality across models like IDDPM, ADM, DDIM, EDM, and FLUX with minimal computational overhead.
Conditional Diffusion Under Linear Constraints: Langevin Mixing and Information-Theoretic Guarantees
This paper analyzes zero-shot conditional sampling with pretrained diffusion models for linear inverse problems, providing information-theoretic guarantees and proposing a projected-Langevin initialization method.
Conservation Laws for Diffusion Models
This paper develops conservation laws for diffusion models using generalized extrinsic information transfer (GEXIT) functions, showing that the cross-entropy can be characterized as an integral of local information-theoretic derivatives along the noise path, unifying likelihood characterization for discrete and continuous diffusion.