PAC-Bayes Beyond Parameter Space: Behavioral Equivalence, Z-Information, and Exact Complexity Decomposition
Summary
This paper extends PAC-Bayes theory by decomposing its complexity measure using behavioral equivalence, introducing PAC-Bayes Z-information and exact structural decomposition beyond parameter space.
View Cached Full Text
Cached at: 08/13/26, 03:36 PM
# PAC-Bayes Beyond Parameter Space: Behavioral Equivalence, Z-Information, and Exact Complexity Decomposition
Source: [https://arxiv.org/html/2608.11465](https://arxiv.org/html/2608.11465)
Vasant G Honavarvuh14@psu\.eduAffiliation:Artificial Intelligence Research LaboratoryAffiliation:Department of Informatics and Intelligent SystemsAffiliation:The Pennsylvania State UniversitySatish Kumar Keshriskk6485@psu\.eduAffiliation:Artificial Intelligence Research LaboratoryAffiliation:Department of Informatics and Intelligent SystemsAffiliation:The Pennsylvania State UniversityNeil Ashtekarnca5096@psu\.eduAffiliation:Artificial Intelligence Research LaboratoryAffiliation:Department of Computer Science and EngineeringAffiliation:The Pennsylvania State UniversityZehao Liuzml5418@psu\.eduAffiliation:Artificial Intelligence Research LaboratoryAffiliation:Department of Informatics and Intelligent SystemsAffiliation:The Pennsylvania State University
###### Abstract
PAC\-Bayes theory provides generalization guarantees by controlling the Kullback–Leibler \(KL\) divergence between posterior and prior distributions over a chosen hypothesis representation\. However, predictive risk depends only on the predictive behavior induced by a hypothesis, not on the particular internal realization that implements that behavior\. In modern over\-parameterized learning systems, many distinct configurations may induce identical predictive behavior, yet the classical PAC\-Bayes KL divergence does not distinguish uncertainty over predictive behavior from variation among behaviorally equivalent realizations\.
We show that this distinction induces an exact structural decomposition of classical PAC\-Bayes complexity\. We formalize behavioral equivalence through a measurable behavior map and use measure disintegration to decompose probability measures on the configuration space into a distribution over predictive behaviors together with conditional distributions over behavioral fibers\. This yields an exact decomposition of the classical PAC\-Bayes KL divergence into a behavior\-selection term and a realization\-level term given by an expected conditional KL divergence within behavioral fibers\.
We define PAC\-Bayes Z\-information as the negative of this realization\-level contribution\. Consequently, PAC\-Bayes Z\-information exactly quantifies the gap between the classical PAC\-Bayes KL divergence and the irreducible complexity associated with uncertainty over predictive behavior\. We further show that the behavior\-selection term admits an exact variational characterization: it is the minimum classical PAC\-Bayes KL divergence among all posteriors that induce the same distribution over predictive behaviors\. Equivalently, every posterior admits a canonical fiber\-symmetrized representative with identical predictive behavior and minimum classical PAC\-Bayes complexity\.
Finally, we show that symmetry, behavior\-preserving directions, fiber geometry, and invariance under fiber\-preserving perturbations arise naturally from the same behavior\-map structure\. Together, these results identify predictive behavior as the natural object of PAC\-Bayes complexity, provide a unified measure\-theoretic and geometric characterization of realization multiplicity, and reveal an exact structural decomposition that is implicit in classical PAC\-Bayes theory\.
## 1Introduction
### 1\.1Background and Motivation
PAC\-Bayes theory provides some of the strongest generalization guarantees in statistical learning by relating predictive performance to divergences between posterior and prior distributions over hypotheses\([30](https://arxiv.org/html/2608.11465#bib.bib33);[36](https://arxiv.org/html/2608.11465#bib.bib39);[5](https://arxiv.org/html/2608.11465#bib.bib14)\)\. These results build on fundamental information\-theoretic quantities such as entropy, relative entropy, cross\-entropy, and mutual information, which play a central role in learning, compression, and statistical inference\([38](https://arxiv.org/html/2608.11465#bib.bib19);[39](https://arxiv.org/html/2608.11465#bib.bib11);[8](https://arxiv.org/html/2608.11465#bib.bib20)\)\. Classical principles such as Minimum Description Length \(MDL\) similarly interpret learning through compression and parsimonious representation\([34](https://arxiv.org/html/2608.11465#bib.bib38);[15](https://arxiv.org/html/2608.11465#bib.bib1)\)\.
A common feature of these approaches is that uncertainty and complexity are quantified with respect to a chosen representation, whether over hypotheses, functions, parameters, or observable outcomes\. However, predictive risk depends only on the behavior induced by a hypothesis, not on the particular internal realization that implements that behavior\. This raises a fundamental question: to what extent does classical PAC\-Bayes complexity reflect uncertainty over predictive behavior, as opposed to variation among equivalent realizations that have no effect on prediction?
This distinction becomes increasingly important in modern over\-parameterized learning systems\. Neural networks often admit many distinct parameter configurations that induce identical or nearly identical input–output behavior\([32](https://arxiv.org/html/2608.11465#bib.bib37);[11](https://arxiv.org/html/2608.11465#bib.bib22)\)\. Such redundancy has been empirically associated with improved generalization and robustness, often through flat minima or low\-curvature regions of parameter space\([17](https://arxiv.org/html/2608.11465#bib.bib23);[22](https://arxiv.org/html/2608.11465#bib.bib24);[7](https://arxiv.org/html/2608.11465#bib.bib16)\)\. At the same time, sharpness\- and curvature\-based characterizations are sensitive to reparameterization and optimization details, making it difficult to interpret them as intrinsic properties of predictive behavior\([11](https://arxiv.org/html/2608.11465#bib.bib22);[20](https://arxiv.org/html/2608.11465#bib.bib7)\)\.
These observations suggest distinguishing uncertainty over predictive behavior from variation among realizations of the same behavior\. If many configurations induce an identical predictive behavior—a phenomenon we refer to as*realization multiplicity*—then predictive risk cannot distinguish among them, whereas configuration\-space complexity may still assign distinct contributions to behaviorally equivalent realizations\.
PAC\-Bayes theory already permits formulations on arbitrary measurable hypothesis spaces, including function\-space representations\([36](https://arxiv.org/html/2608.11465#bib.bib39)\)\. Likewise, recent work has shown that invariance and symmetry can reduce effective PAC\-Bayes complexity by identifying predictors that are equivalent under suitable transformations\([29](https://arxiv.org/html/2608.11465#bib.bib35);[2](https://arxiv.org/html/2608.11465#bib.bib34)\)\. These developments motivate a broader question\. Given an arbitrary measurable notion of behavioral equivalence, does the classical PAC\-Bayes KL divergence admit an exact decomposition into contributions arising from predictive behavior and contributions arising from realization multiplicity? If so, what structural properties does such a decomposition reveal?
### 1\.2Behavioral Sufficiency and Realization Multiplicity
The central observation underlying this work is that empirical and population risks are invariant to the particular realization of a predictor and depend only on its induced input–output behavior\.
This motivates an equivalence relation on hypotheses:
θ∼θ′⟺θandθ′induce identical predictive behavior\.\\theta\\sim\\theta^\{\\prime\}\\quad\\Longleftrightarrow\\quad\\theta\\text\{ and \}\\theta^\{\\prime\}\\text\{ induce identical predictive behavior\}\.
Behavioral equivalence partitions the configuration space into equivalence classes of realizations that implement the same predictor\. Each equivalence class therefore represents a single predictive behavior together with all of its realizations\. This separation between behavior and realization provides a natural framework for distinguishing uncertainty over predictive behavior from variation among equivalent realizations\.
For the measure\-theoretic development that follows, we realize this structure through a measurable behavior map
β:Θ→𝒦,\\beta:\\Theta\\rightarrow\\mathcal\{K\},where𝒦\\mathcal\{K\}is a standard Borel space of predictive behaviors andβ\(θ\)\\beta\(\\theta\)denotes the predictive behavior induced by configurationθ\\theta\. Behavioral equivalence is then precisely equality under the behavior map:
θ∼θ′⟺β\(θ\)=β\(θ′\)\.\\theta\\sim\\theta^\{\\prime\}\\quad\\Longleftrightarrow\\quad\\beta\(\\theta\)=\\beta\(\\theta^\{\\prime\}\)\.
For a behaviork∈𝒦k\\in\\mathcal\{K\}, the setβ−1\(k\)\\beta^\{\-1\}\(k\)contains all configurations that induce that behavior\. In the language of measurable maps, these sets are the fibers ofβ\\beta\. Throughout the paper, we use the terms*behavioral equivalence class*and*fiber*interchangeably\.
Under standard measure\-theoretic assumptions, probability measures on the configuration space admit a disintegration with respect to the behavior map\([6](https://arxiv.org/html/2608.11465#bib.bib42);[21](https://arxiv.org/html/2608.11465#bib.bib29)\)\. This yields a decomposition into
1. 1\.a distribution over predictive behaviors; and
2. 2\.conditional distributions over realizations within each behavioral fiber\.
From this perspective, the classical PAC\-Bayes KL divergence naturally separates into two conceptually distinct sources of divergence:
1. 1\.uncertainty over predictive behavior; and
2. 2\.uncertainty over realizations conditional on behavior\.
Only the first directly influences predictive risk\. The second reflects variation among realizations that induce the same behavior and therefore contributes to the classical PAC\-Bayes KL divergence without changing predictive behavior\.
Importantly, this decomposition is induced entirely by the behavior map and is therefore structural rather than parameterization\-dependent\. It arises from behavioral equivalence itself rather than from any particular coordinate system, optimization procedure, or local geometric approximation\.
### 1\.3Contributions
Our objective is not to formulate PAC\-Bayes analysis on behavior spaces per se, but to identify the latent structure that an arbitrary measurable notion of behavioral equivalence induces within classical PAC\-Bayes complexity\. We show that the classical PAC\-Bayes KL divergence admits an exact decomposition into behavior\-level and realization\-level contributions and characterize the information\-theoretic and geometric consequences of this decomposition\.
Our main contributions are as follows\.
- •Behavioral decomposition of PAC\-Bayes complexity\.We formalize behavioral equivalence through an arbitrary measurable behavior map and develop the associated fiber\-based measure\-theoretic framework\. Within this framework, we prove that the classical PAC\-Bayes KL divergence admits an exact decomposition into a behavior\-selection component and a realization\-level component associated with behavioral fibers\.
- •PAC\-Bayes Z\-information\.We define PAC\-Bayes Z\-information as the negative expected conditional KL divergence within behavioral fibers and show that it exactly characterizes the realization\-level contribution to classical PAC\-Bayes complexity, thereby isolating the effect of realization multiplicity\.
- •Variational characterization of behavior selection complexity\.We show that the behavior\-selection term is the minimum classical PAC\-Bayes KL divergence among all posteriors inducing the same distribution over predictive behaviors\. This identifies behavior selection complexity as the irreducible contribution associated with predictive behavior itself\.
- •Unified geometric interpretation\.Under suitable regularity assumptions, behavioral fibers naturally give rise to behavior\-preserving directions, symmetry, realization multiplicity, and invariance under fiber\-preserving perturbations\. These provide complementary geometric interpretations of the realization\-level component identified by the information\-theoretic decomposition\.
Taken together, these results show that classical PAC\-Bayes complexity contains a latent decomposition into behavior\-level and realization\-level components\. The remainder of the paper develops this decomposition, characterizes its information\-theoretic and geometric consequences, and shows how it clarifies the role of realization multiplicity in PAC\-Bayes analysis\.
### 1\.4Organization
The remainder of the paper is organized as follows\. Section[2](https://arxiv.org/html/2608.11465#S2)introduces behavioral equivalence, behavior maps, and the measure\-theoretic framework used throughout the paper\. Section[3](https://arxiv.org/html/2608.11465#S3)develops the fiberwise decomposition induced by behavioral equivalence and introduces realization entropy and PAC\-Bayes Z\-information as measures of realization multiplicity\. Section[4](https://arxiv.org/html/2608.11465#S4)establishes the exact decomposition of the classical PAC\-Bayes KL divergence into behavior\-selection and realization\-level terms, develops the corresponding PAC\-Bayes analysis, and derives a variational characterization of the behavior\-selection term\. Section[5](https://arxiv.org/html/2608.11465#S5)develops geometric interpretations in continuous configuration spaces, relating realization multiplicity to behavioral fibers, symmetry, behavior\-preserving directions, and volume structure\. Section[6](https://arxiv.org/html/2608.11465#S6)studies behavioral invariance under fiber\-preserving perturbations and its relation to classical algorithmic stability\. Section[7](https://arxiv.org/html/2608.11465#S7)discusses related work\. Section[8](https://arxiv.org/html/2608.11465#S8)concludes with a summary, implications, and directions for further research\.
## 2Setup and Behavioral Equivalence
We formalize the learning setting and introduce behavioral equivalence, which provides the structural basis for separating uncertainty over predictive behavior from uncertainty over internal realization\.
The key observation is simple\. In modern over\-parameterized models, many distinct parameter configurations can induce identical input–output behavior\. Hidden\-unit permutations, scaling symmetries, and other sources of non\-identifiability can produce different parameter vectors that implement the same predictor\. Since prediction depends only on behavior and not on the particular realization that implements it, generalization analysis should distinguish uncertainty over predictive behavior from redundancy among equivalent realizations\. The goal of this section is to make that distinction precise\.
### 2\.1Learning Framework
Let𝒳\\mathcal\{X\}and𝒴\\mathcal\{Y\}denote the input and output spaces, and let𝒟\\mathcal\{D\}be an unknown distribution over𝒳×𝒴\\mathcal\{X\}\\times\\mathcal\{Y\}\. We work in the standard statistical learning setting\([41](https://arxiv.org/html/2608.11465#bib.bib40);[37](https://arxiv.org/html/2608.11465#bib.bib26)\)\.
The configuration spaceΘ\\Thetaindexes hypotheses \(internal realizations\)\. For example,Θ\\Thetamay be a neural\-network parameter space\. Each configurationθ∈Θ\\theta\\in\\Thetainduces a conditional predictive distributionPθ\(⋅∣x\)\.P\_\{\\theta\}\(\\cdot\\mid x\)\.
Performance is measured using a bounded lossℓ:Δ\(𝒴\)×𝒴→\[0,1\],\\ell:\\Delta\(\\mathcal\{Y\}\)\\times\\mathcal\{Y\}\\rightarrow\[0,1\],whereΔ\(𝒴\)\\Delta\(\\mathcal\{Y\}\)denotes the set of probability distributions on𝒴\\mathcal\{Y\}\.
The population and empirical risks are
L\(θ\)=𝔼\(x,y\)∼𝒟\[ℓ\(Pθ\(⋅∣x\),y\)\],L^S\(θ\)=1n∑i=1nℓ\(Pθ\(⋅∣xi\),yi\),L\(\\theta\)=\\mathbb\{E\}\_\{\(x,y\)\\sim\\mathcal\{D\}\}\\bigl\[\\ell\(P\_\{\\theta\}\(\\cdot\\mid x\),y\)\\bigr\],\\qquad\\hat\{L\}\_\{S\}\(\\theta\)=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\ell\(P\_\{\\theta\}\(\\cdot\\mid x\_\{i\}\),y\_\{i\}\),whereS=\{\(xi,yi\)\}i=1n\.S=\\\{\(x\_\{i\},y\_\{i\}\)\\\}\_\{i=1\}^\{n\}\.
As in PAC\-Bayes theory, we consider stochastic predictorsQ∈𝒫\(Θ\)\.Q\\in\\mathcal\{P\}\(\\Theta\)\.
Their population and empirical risks are
L\(Q\)=𝔼θ∼Q\[L\(θ\)\],L^S\(Q\)=𝔼θ∼Q\[L^S\(θ\)\],L\(Q\)=\\mathbb\{E\}\_\{\\theta\\sim Q\}\[L\(\\theta\)\],\\qquad\\hat\{L\}\_\{S\}\(Q\)=\\mathbb\{E\}\_\{\\theta\\sim Q\}\[\\hat\{L\}\_\{S\}\(\\theta\)\],following the standard PAC\-Bayes formulation\([30](https://arxiv.org/html/2608.11465#bib.bib33);[36](https://arxiv.org/html/2608.11465#bib.bib39);[5](https://arxiv.org/html/2608.11465#bib.bib14)\)\.
### 2\.2Behavior Maps and Behavioral Equivalence
Many learning systems admit multiple internal realizations that implement the same predictive behavior\. To formalize this idea, we introduce a measurable behavior mapβ:Θ→𝒦,\\beta:\\Theta\\rightarrow\\mathcal\{K\},where𝒦\\mathcal\{K\}is a standard Borel space of predictive behaviors \(for example, conditional distributionsx↦Pθ\(⋅∣x\)x\\mapsto P\_\{\\theta\}\(\\cdot\\mid x\)\), andβ\(θ\)\\beta\(\\theta\)denotes the predictive behavior induced by configurationθ\\theta\.
###### Definition 1\(Behavioral Equivalence\)\.
Two configurationsθ,θ′∈Θ\\theta,\\theta^\{\\prime\}\\in\\Thetaare behaviorally equivalent, writtenθ∼θ′,\\theta\\sim\\theta^\{\\prime\},ifβ\(θ\)=β\(θ′\)\.\\beta\(\\theta\)=\\beta\(\\theta^\{\\prime\}\)\.
Equivalently, behavioral equivalence identifies configurations that induce the same predictive behavior as represented by the behavior mapβ\\beta\. In the examples considered throughout this paper,β\\betarecords the conditional predictive distribution induced by a configuration\.
For simplicity, we state behavioral equivalence in terms of exact equality of predictive behavior\. All results continue to hold if behavioral equivalence is defined modulo equality almost everywhere with respect to the input distribution\.
For a behaviork∈𝒦,k\\in\\mathcal\{K\},the setFk=β−1\(k\)=\{θ∈Θ:β\(θ\)=k\}F\_\{k\}=\\beta^\{\-1\}\(k\)=\\\{\\theta\\in\\Theta:\\beta\(\\theta\)=k\\\}contains all configurations that realize behaviorkk\. Following standard terminology, we refer toFkF\_\{k\}as the*fiber*associated with behaviorkk\([21](https://arxiv.org/html/2608.11465#bib.bib29);[33](https://arxiv.org/html/2608.11465#bib.bib28)\)\.
Thus, each behavior corresponds to a fiber of internally distinct but predictively equivalent realizations\.
For example, permuting hidden units in a neural network while applying the corresponding permutation to outgoing weights leaves the predictive behavior unchanged \(See Appendix[H\.2](https://arxiv.org/html/2608.11465#A8.SS2)\)\. The original and permuted parameter vectors therefore belong to the same fiber\.
More generally, fibers may arise from parameter symmetries, redundant hidden units, scaling invariances, or other sources of non\-identifiability\. Such redundancy is a well\-known feature of over\-parameterized models\([32](https://arxiv.org/html/2608.11465#bib.bib37);[11](https://arxiv.org/html/2608.11465#bib.bib22)\)\.
Although the theory developed in this paper is formulated in terms of the behavior mapβ\\beta, it is often helpful to think of𝒦\\mathcal\{K\}as playing the role of a quotient representation of the configuration space, where behaviorally equivalent realizations are treated as equivalent\.
### 2\.3Induced Distribution over Behaviors
Any distributionQ∈𝒫\(Θ\)Q\\in\\mathcal\{P\}\(\\Theta\)induces a distribution over predictive behaviors through the pushforward measureπQ=β\#Q=Q∘β−1∈𝒫\(𝒦\)\.\\pi\_\{Q\}=\\beta\_\{\\\#\}Q=Q\\circ\\beta^\{\-1\}\\in\\mathcal\{P\}\(\\mathcal\{K\}\)\.
Intuitively,πQ\\pi\_\{Q\}records how much probability massQQassigns to each predictive behavior while ignoring how that mass is distributed among realizations within a fiber\.
#### Behavioral Sufficiency\.
The induced distributionπQ\\pi\_\{Q\}captures all information inQQthat can affect population or empirical risk\. This is not an additional modeling assumption\. Rather, it follows directly from behavioral equivalence: configurations within the same fiber induce the same predictive behavior and therefore incur the same loss\.
### 2\.4Loss Invariance and Behavioral Sufficiency
###### Lemma 1\(Loss Invariance under Behavioral Equivalence\)\.
Ifθ∼θ′,\\theta\\sim\\theta^\{\\prime\},then
L\(θ\)=L\(θ′\),L^S\(θ\)=L^S\(θ′\)\.L\(\\theta\)=L\(\\theta^\{\\prime\}\),\\qquad\\hat\{L\}\_\{S\}\(\\theta\)=\\hat\{L\}\_\{S\}\(\\theta^\{\\prime\}\)\.
###### Proof\.
Behavioral equivalence implies thatθ\\thetaandθ′\\theta^\{\\prime\}induce the same predictive behavior\. Hence
ℓ\(Pθ\(⋅∣x\),y\)=ℓ\(Pθ′\(⋅∣x\),y\)\\ell\(P\_\{\\theta\}\(\\cdot\\mid x\),y\)=\\ell\(P\_\{\\theta^\{\\prime\}\}\(\\cdot\\mid x\),y\)for every example\(x,y\)\(x,y\)\. Taking expectations with respect to𝒟\\mathcal\{D\}yieldsL\(θ\)=L\(θ′\),L\(\\theta\)=L\(\\theta^\{\\prime\}\),while averaging over the sampleSSyieldsL^S\(θ\)=L^S\(θ′\)\.\\hat\{L\}\_\{S\}\(\\theta\)=\\hat\{L\}\_\{S\}\(\\theta^\{\\prime\}\)\.∎
The lemma implies that both population and empirical risk are constant on each fiber\. Consequently, there exist functionsL𝒦,L^S,𝒦,L\_\{\\mathcal\{K\}\},\\hat\{L\}\_\{S,\\mathcal\{K\}\},defined on the behavior space𝒦\\mathcal\{K\}such that
L\(θ\)=L𝒦\(β\(θ\)\),L^S\(θ\)=L^S,𝒦\(β\(θ\)\)\.L\(\\theta\)=L\_\{\\mathcal\{K\}\}\(\\beta\(\\theta\)\),\\qquad\\hat\{L\}\_\{S\}\(\\theta\)=\\hat\{L\}\_\{S,\\mathcal\{K\}\}\(\\beta\(\\theta\)\)\.
###### Proposition 1\(Behavioral Sufficiency\)\.
IfπQ=πQ′,\\pi\_\{Q\}=\\pi\_\{Q^\{\\prime\}\},then
L\(Q\)=L\(Q′\),L^S\(Q\)=L^S\(Q′\)\.L\(Q\)=L\(Q^\{\\prime\}\),\\qquad\\hat\{L\}\_\{S\}\(Q\)=\\hat\{L\}\_\{S\}\(Q^\{\\prime\}\)\.
###### Proof\.
Using the factorization above,
L\(Q\)=𝔼θ∼Q\[L𝒦\(β\(θ\)\)\]=𝔼k∼πQ\[L𝒦\(k\)\]\.L\(Q\)=\\mathbb\{E\}\_\{\\theta\\sim Q\}\\Big\[L\_\{\\mathcal\{K\}\}\(\\beta\(\\theta\)\)\\Big\]=\\mathbb\{E\}\_\{k\\sim\\pi\_\{Q\}\}\\bigl\[L\_\{\\mathcal\{K\}\}\(k\)\\bigr\]\.The same representation holds forQ′Q^\{\\prime\}\. IfπQ=πQ′\\pi\_\{Q\}=\\pi\_\{Q^\{\\prime\}\}, the expectations coincide\. The argument for empirical risk is identical\. ∎
Proposition[1](https://arxiv.org/html/2608.11465#Thmproposition1)shows that both population and empirical risk factors are affected by the behavior map\. Consequently, there exist functionals, which we denote by the same symbols for convenience, such that
L\(Q\)=L\(πQ\),L^S\(Q\)=L^S\(πQ\)\.L\(Q\)=L\(\\pi\_\{Q\}\),\\qquad\\hat\{L\}\_\{S\}\(Q\)=\\hat\{L\}\_\{S\}\(\\pi\_\{Q\}\)\.Thus, all quantities relevant to prediction depend on a posterior only through its induced distribution over predictive behaviors\.
In particular, any two posteriors with the same induced behavior distribution are indistinguishable from the perspective of empirical and population risk\.
### 2\.5Measure\-Theoretic Assumptions
To separate uncertainty across behaviors from uncertainty within a behavior, we will later decompose probability measures along fibers\.
We assume that\(Θ,ℱΘ\)\(\\Theta,\\mathcal\{F\}\_\{\\Theta\}\)and\(𝒦,ℱ𝒦\)\(\\mathcal\{K\},\\mathcal\{F\}\_\{\\mathcal\{K\}\}\)are standard Borel spaces and that the behavior mapβ:Θ→𝒦\\beta:\\Theta\\rightarrow\\mathcal\{K\}is measurable\.111Standard Borel spaces are measurable spaces arising from the Borelσ\\sigma\-algebra of a Polish \(complete separable metric\) space\. They include essentially all parameter and function spaces commonly used in statistical learning, including Euclidean spaces and many spaces of probability measures equipped with their natural Borel structures\([21](https://arxiv.org/html/2608.11465#bib.bib29);[33](https://arxiv.org/html/2608.11465#bib.bib28)\)\.
Under these assumptions, standard disintegration theorems guarantee the existence of a family of conditional probability measuresQ\(⋅∣k\)Q\(\\cdot\\mid k\)such that, for every measurable setA∈ℱΘA\\in\\mathcal\{F\}\_\{\\Theta\},
Q\(A\)=∫𝒦Q\(A∣k\)πQ\(𝑑k\)\.Q\(A\)=\\int\_\{\\mathcal\{K\}\}Q\(A\\mid k\)\\,\\pi\_\{Q\}\(dk\)\.Equivalently,
Q\(dθ\)=Q\(dθ∣k\)πQ\(dk\)\.Q\(d\\theta\)=Q\(d\\theta\\mid k\)\\,\\pi\_\{Q\}\(dk\)\.The conditional measureQ\(dθ∣k\)Q\(d\\theta\\mid k\)describes uncertainty among realizations that induce the same predictive behaviorkk\([21](https://arxiv.org/html/2608.11465#bib.bib29);[33](https://arxiv.org/html/2608.11465#bib.bib28);[3](https://arxiv.org/html/2608.11465#bib.bib15)\)\.
Readers unfamiliar with fibers or measure disintegration may consult Appendix[A](https://arxiv.org/html/2608.11465#A1), which provides a brief review of the required background\.
### 2\.6Role in Generalization
Classical generalization theory controls deviations between empirical and population risk using complexity measures such as VC dimension, stability, Rademacher complexity, and PAC\-Bayes KL divergence\([41](https://arxiv.org/html/2608.11465#bib.bib40);[4](https://arxiv.org/html/2608.11465#bib.bib17);[30](https://arxiv.org/html/2608.11465#bib.bib33)\)\.
In the present setting, complexity measures defined on the configuration spaceΘ\\Thetacombine two conceptually distinct sources of uncertainty:
1. 1\.uncertainty over predictive behaviors;
2. 2\.uncertainty among realizations that implement a fixed behavior\.
By Proposition[1](https://arxiv.org/html/2608.11465#Thmproposition1), only the first affects population or empirical risk\. The second reflects*realization multiplicity*arising from over\-parameterization\.
This distinction between behavioral uncertainty and realization multiplicity is the central organizing principle of the paper\. The next section shows that the PAC\-Bayes KL divergence admits an exact fiberwise decomposition into a behavior\-selection term and a realization\-level term\. PAC\-Bayes Z\-information is defined as the negative of the realization\-level contribution and therefore quantifies the portion of PAC\-Bayes complexity associated with the redistribution of posterior mass among behaviorally equivalent realizations\.
Moreover, passing from a posterior on configurations to its induced distribution on predictive behaviors can be viewed as an application of the data\-processing principle for relative entropy\. The next section develops this connection and shows that the behavior selection complexity term admits a variational characterization within the classical PAC\-Bayes framework: among all posteriors inducing the same distribution over predictive behaviors, it is the minimum achievable PAC\-Bayes complexity\. Equivalently, each behavior\-level posterior admits a canonical fiber\-symmetrized representative whose classical PAC\-Bayes complexity coincides with the behavior\-selection term\. This characterization will play a central role in the development that follows\.
## 3Behavior\-Realization Decomposition and PAC\-Bayes Z\-Information
Section[2](https://arxiv.org/html/2608.11465#S2)established that predictive risk depends only on the distribution of predictive behaviors induced by a posterior\. This section develops the central information\-theoretic decomposition underlying the remainder of the paper\.
The key result is an exact decomposition of the classical PAC\-Bayes complexity term into a behavior\-level component and a realization\-level component\. PAC\-Bayes Z\-information emerges as the negative of the latter and quantifies the portion of PAC\-Bayes complexity attributable to redistributing probability mass among behaviorally equivalent realizations\.
Throughout, the behavior mapβ:Θ→𝒦\\beta:\\Theta\\rightarrow\\mathcal\{K\}induces behavior\-level distributions
πP=β\#P,πQ=β\#Q,\\pi\_\{P\}=\\beta\_\{\\\#\}P,\\qquad\\pi\_\{Q\}=\\beta\_\{\\\#\}Q,whereβ\#P:=P∘β−1\\beta\_\{\\\#\}P:=P\\circ\\beta^\{\-1\}andβ\#Q:=Q∘β−1\\beta\_\{\\\#\}Q:=Q\\circ\\beta^\{\-1\}denote the pushforward measures ofPPandQQthrough the behavior map\. Equivalently, for every measurable setA⊆𝒦A\\subseteq\\mathcal\{K\},
πP\(A\)=P\(β−1\(A\)\),πQ\(A\)=Q\(β−1\(A\)\)\.\\pi\_\{P\}\(A\)=P\(\\beta^\{\-1\}\(A\)\),\\qquad\\pi\_\{Q\}\(A\)=Q\(\\beta^\{\-1\}\(A\)\)\.Because the quantities developed in this paper are relative rather than absolute, our analysis is based on the relative entropy between posterior and prior distributions rather than on the entropy defined within individual fibers\. This avoids introducing fiber\-specific reference measures and yields quantities that arise naturally in PAC\-Bayes theory\.
### 3\.1Behavior–Realization Decomposition of Relative Entropy
The central object in PAC\-Bayes theory is the relative entropyKL\(Q∥P\)\\mathrm\{KL\}\(Q\\\|P\)between a posteriorQQand a priorPP\. Because this quantity is computed on the configuration spaceΘ\\Theta, it combines variation across predictive behaviors with variation among realizations of a fixed behavior\.
The decomposition extends to the extended\-real setting whenKL\(Q∥P\)=∞\\mathrm\{KL\}\(Q\\\|P\)=\\infty\. We focus on the finite\-KL regime because it is the one relevant to classical PAC\-Bayes bounds\.
The following result separates these contributions exactly\.
###### Proposition 2\(Behavior–Realization KL Decomposition\)\.
AssumeKL\(Q∥P\)<∞\\mathrm\{KL\}\(Q\\\|P\)<\\inftyand let
Q\(dθ\)=Q\(dθ∣k\)πQ\(dk\),P\(dθ\)=P\(dθ∣k\)πP\(dk\)Q\(d\\theta\)=Q\(d\\theta\\mid k\)\\,\\pi\_\{Q\}\(dk\),\\qquad P\(d\\theta\)=P\(d\\theta\\mid k\)\\,\\pi\_\{P\}\(dk\)denote disintegrations with respect to the behavior mapβ:Θ→𝒦\\beta:\\Theta\\rightarrow\\mathcal\{K\}\. ThenπQ≪πP\\pi\_\{Q\}\\ll\\pi\_\{P\}and
KL\(Q∥P\)=KL\(πQ∥πP\)\+𝔼k∼πQ\[KL\(Q\(⋅∣k\)∥P\(⋅∣k\)\)\],\\mathrm\{KL\}\(Q\\\|P\)=\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\)\+\\mathbb\{E\}\_\{k\\sim\\pi\_\{Q\}\}\\\!\\left\[\\mathrm\{KL\}\\\!\\left\(Q\(\\cdot\\mid k\)\\,\\middle\\\|\\,P\(\\cdot\\mid k\)\\right\)\\right\],where\`\`≪"\`\`\\ll"denotes absolute continuity of measures\.
The decomposition is a direct consequence of the chain rule for relative entropy under measure disintegration\([21](https://arxiv.org/html/2608.11465#bib.bib29);[3](https://arxiv.org/html/2608.11465#bib.bib15), see, e\.g\.,\)\. Its importance here is interpretive rather than technical: it separates behavior selection complexity from realization\-level complexity\.
###### Proof sketch\.
BecauseKL\(Q∥P\)<∞\\mathrm\{KL\}\(Q\\\|P\)<\\infty, we haveQ≪PQ\\ll P, and henceπQ≪πP\\pi\_\{Q\}\\ll\\pi\_\{P\}\.
Using the disintegrations ofQQandPPwith respect to the behavior mapβ\\beta, the chain rule for relative entropy yields a decomposition into a behavior\-selection term and a conditional within\-fiber term\. Rearranging gives the stated identity\.
A complete proof is given in Appendix[B](https://arxiv.org/html/2608.11465#A2)\. ∎
The first term measures the divergence between distributions over predictive behaviors\. The second measures the divergence between posterior and prior conditional distributions within behavioral fibers\.
Thus, the decomposition separates the divergence associated with selecting predictive behaviors from the divergence associated with redistributing mass among realizations of those behaviors\.
###### Corollary 1\(Data Processing Inequality\)\.
KL\(πQ∥πP\)≤KL\(Q∥P\)\.\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\)\\leq\\mathrm\{KL\}\(Q\\\|P\)\.
###### Proof\.
The conditional KL divergence in Proposition[2](https://arxiv.org/html/2608.11465#Thmproposition2)is nonnegative\.
∎
Corollary[1](https://arxiv.org/html/2608.11465#Thmcorollary1)is precisely the data\-processing inequality for relative entropy applied to the measurable mapβ:Θ→𝒦\\beta:\\Theta\\rightarrow\\mathcal\{K\}\.
In particular, passing from a posterior on configurations to its induced distribution on predictive behaviors can only decrease relative entropy\. The lost information is exactly the realization\-level term identified in Proposition[2](https://arxiv.org/html/2608.11465#Thmproposition2)\.
### 3\.2PAC\-Bayes Z\-Information
The decomposition above suggests isolating the realization\-level contribution\.
###### Definition 2\(PAC\-Bayes Z\-Information\)\.
The*PAC\-Bayes Z\-information*of a posteriorQQrelative to a priorPPis
𝒵PB\(Q∥P\)=−𝔼k∼πQ\[KL\(Q\(⋅∣k\)∥P\(⋅∣k\)\)\]\.\\mathcal\{Z\}\_\{\\mathrm\{PB\}\}\(Q\\\|P\)=\-\\mathbb\{E\}\_\{k\\sim\\pi\_\{Q\}\}\\\!\\left\[\\mathrm\{KL\}\\\!\\left\(Q\(\\cdot\\mid k\)\\,\\middle\\\|\\,P\(\\cdot\\mid k\)\\right\)\\right\]\.
By construction,𝒵PB\(Q∥P\)≤0\\mathcal\{Z\}\_\{\\mathrm\{PB\}\}\(Q\\\|P\)\\leq 0, with equality if and only ifQ\(⋅∣k\)=P\(⋅∣k\)Q\(\\cdot\\mid k\)=P\(\\cdot\\mid k\)forπQ\\pi\_\{Q\}\-almost everykk\.
Combining Definition[2](https://arxiv.org/html/2608.11465#Thmdefinition2)with Proposition[2](https://arxiv.org/html/2608.11465#Thmproposition2)immediately yields the following identity\.
###### Proposition 3\(Z\-Information Identity\)\.
𝒵PB\(Q∥P\)=KL\(πQ∥πP\)−KL\(Q∥P\)\.\\mathcal\{Z\}\_\{\\mathrm\{PB\}\}\(Q\\\|P\)=\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\)\-\\mathrm\{KL\}\(Q\\\|P\)\.
###### Proof\.
Rearrange the identity in Proposition[2](https://arxiv.org/html/2608.11465#Thmproposition2)and substitute Definition[2](https://arxiv.org/html/2608.11465#Thmdefinition2)\.
∎
PAC\-Bayes Z\-information is therefore not an additional divergence\. Rather, it is an alternative representation of the gap between configuration\-space complexity and behavior\-space complexity\.
A value of𝒵PB\(Q∥P\)=0\\mathcal\{Z\}\_\{\\mathrm\{PB\}\}\(Q\\\|P\)=0indicates that the posterior and prior induce identical distributions within every fiber, so that all complexity is attributable to uncertainty over predictive behavior\.
Negative values arise precisely when the posterior differs from the prior within behavioral fibers\. In this case, the classical PAC\-Bayes complexity exceeds the behavior selection complexity by−𝒵PB\(Q∥P\)\-\\mathcal\{Z\}\_\{\\mathrm\{PB\}\}\(Q\\\|P\)\.
### 3\.3Variational Characterization
The decomposition admits a useful variational interpretation that clarifies the relationship between behavior selection complexity and classical PAC\-Bayes theory\.
Given a posteriorQQ, define the*fiber\-symmetrized posterior*
Q⋆\(𝑑θ\)=∫𝒦P\(𝑑θ∣k\)πQ\(𝑑k\)\.Q^\{\\star\}\(d\\theta\)=\\int\_\{\\mathcal\{K\}\}P\(d\\theta\\mid k\)\\,\\pi\_\{Q\}\(dk\)\.The distributionQ⋆Q^\{\\star\}preserves the behavior\-level distributionπQ\\pi\_\{Q\}while replacing each fiberwise conditional distributionQ\(⋅∣k\)Q\(\\cdot\\mid k\)with the corresponding prior conditionalP\(⋅∣k\)P\(\\cdot\\mid k\)\. BecauseQQandQ⋆Q^\{\\star\}induce the same distribution over predictive behaviors, Proposition[1](https://arxiv.org/html/2608.11465#Thmproposition1)implies that
L\(Q⋆\)=L\(Q\)andL^S\(Q⋆\)=L^S\(Q\)L\(Q^\{\\star\}\)=L\(Q\)\\;\\mathrm\{and\}\\;\\hat\{L\}\_\{S\}\(Q^\{\\star\}\)=\\hat\{L\}\_\{S\}\(Q\)\.
###### Theorem 1\(Fiber\-Symmetrized Optimal Representative\)\.
The fiber\-symmetrized posterior satisfiesπQ⋆=πQ\\pi\_\{Q^\{\\star\}\}=\\pi\_\{Q\}and
KL\(Q⋆∥P\)=KL\(πQ∥πP\)\.\\mathrm\{KL\}\(Q^\{\\star\}\\\|P\)=\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\)\.Moreover,
KL\(πQ∥πP\)=infQ′:πQ′=πQKL\(Q′∥P\)\.\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\)=\\inf\_\{Q^\{\\prime\}:\\,\\pi\_\{Q^\{\\prime\}\}=\\pi\_\{Q\}\}\\mathrm\{KL\}\(Q^\{\\prime\}\\\|P\)\.
#### Interpretation\.
Theorem[1](https://arxiv.org/html/2608.11465#Thmtheorem1)identifies the behavior\-selection divergenceKL\(πQ∥πP\)\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\)as the canonical behavior\-level component of classical PAC\-Bayes complexity\. Among all posteriors inducing the same distribution over predictive behaviors, the fiber\-symmetrized posterior achieves the minimum classical PAC\-Bayes KL divergence\. Consequently, any excess classical PAC\-Bayes complexity beyondKL\(πQ∥πP\)\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\)arises entirely from how posterior probability is distributed among behaviorally equivalent realizations\. PAC\-Bayes Z\-information quantifies this realization\-level contribution exactly, making explicit a latent decomposition of the classical PAC\-Bayes KL divergence into behavior\-selection and realization\-level components\.
###### Proof\.
SinceQ⋆Q^\{\\star\}andQQshare the same behavior\-level marginal,πQ⋆=πQ\\pi\_\{Q^\{\\star\}\}=\\pi\_\{Q\}\.
Applying Proposition[2](https://arxiv.org/html/2608.11465#Thmproposition2)toQ⋆Q^\{\\star\}yieldsKL\(Q⋆∥P\)=KL\(πQ∥πP\)\\mathrm\{KL\}\(Q^\{\\star\}\\\|P\)=\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\)because the conditional divergence term vanishes identically\.
For any posteriorQ′Q^\{\\prime\}satisfyingπQ′=πQ\\pi\_\{Q^\{\\prime\}\}=\\pi\_\{Q\},
KL\(Q′∥P\)=KL\(πQ∥πP\)\+𝔼k∼πQ\[KL\(Q′\(⋅∣k\)∥P\(⋅∣k\)\)\]\.\\mathrm\{KL\}\(Q^\{\\prime\}\\\|P\)=\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\)\+\\mathbb\{E\}\_\{k\\sim\\pi\_\{Q\}\}\\\!\\left\[\\mathrm\{KL\}\\\!\\left\(Q^\{\\prime\}\(\\cdot\\mid k\)\\,\\middle\\\|\\,P\(\\cdot\\mid k\)\\right\)\\right\]\.The conditional KL term is nonnegative\.
∎
### 3\.4Illustration: Hidden\-Unit Permutation Symmetry
The behavior–realization decomposition is particularly transparent in models with parameter symmetries\. Consider a neural network withhhhidden units and a prior that is invariant under permutations of those units\. Permuting hidden units together with their incident weights leaves the induced input–output behavior unchanged\. Consequently, each predictive behavior corresponds to a behavioral fiber containing all parameter configurations related by such permutations\.
If the posterior concentrates near one representative realization while the prior remains approximately uniform over the corresponding behavioral fiber, then the induced predictive behavior remains unchanged, but the posterior and prior differ substantially within the fiber\. Under this idealized setting, the realization\-level contribution satisfies
−𝒵PB\(Q∥P\)≈log\(h\!\),\-\\mathcal\{Z\}\_\{\\mathrm\{PB\}\}\(Q\\\|P\)\\approx\\log\(h\!\),reflecting the combinatorial multiplicity of behaviorally equivalent realizations\.
Thus, realization multiplicity alone can contribute substantially to the classical PAC\-Bayes KL divergence even though every realization within the fiber induces exactly the same predictive behavior\. Importantly, this contribution reflects uncertainty over internal realizations rather than uncertainty over predictive behavior itself\.
Conversely, if both the prior and posterior distribute probability identically within each behavioral fiber—for example, under an exact symmetry\-preserving Bayesian posterior—then𝒵PB\(Q∥P\)=0\\mathcal\{Z\}\_\{\\mathrm\{PB\}\}\(Q\\\|P\)=0, and the classical PAC\-Bayes KL divergence already coincides with the behavior\-selection divergence\.
Although hidden\-unit permutation symmetry provides a familiar illustration, the behavior–realization decomposition developed in this paper applies to arbitrary measurable notions of behavioral equivalence and is not restricted to finite symmetry groups\. Appendix[H](https://arxiv.org/html/2608.11465#A8)presents explicit computations illustrating the decomposition theorem, the variational characterization, and the conditional KL calculation underlying the approximation above\.
The decomposition developed in this section provides the foundation for the PAC\-Bayes analysis that follows\. The next section derives behavior\-level generalization bounds, establishes their relationship to classical PAC\-Bayes theory, and characterizes the role of realization multiplicity in generalization\.
## 4Behavior\-Space Reformulation of PAC\-Bayes
Section[3](https://arxiv.org/html/2608.11465#S3)established that behavioral equivalence induces an exact decomposition of the classical PAC\-Bayes KL divergence into a behavior\-selection component and a realization\-level component\. We now combine this decomposition with the observation that empirical and population risks depend only on the induced distribution over predictive behaviors\.
Because empirical and population risks factor through the behavior map, the classical PAC\-Bayes theorem admits an equivalent formulation on the measurable space of predictive behaviors𝒦\\mathcal\{K\}\. The resulting complexity term is preciselyKL\(πQ∥πP\)\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\), the divergence between the behavior\-level distributions induced by the posterior and prior\. By Theorem[1](https://arxiv.org/html/2608.11465#Thmtheorem1), this quantity is the minimum configuration\-space PAC\-Bayes complexity among all posteriors inducing the same distribution over predictive behaviors\. Consequently, the behavior\-space formulation identifies the irreducible contribution of predictive behavior to classical PAC\-Bayes complexity, while PAC\-Bayes Z\-information quantifies the realization\-level contribution arising from the distribution of posterior mass among behaviorally equivalent realizations\.
### 4\.1Classical PAC\-Bayes Bounds
For completeness, we recall a standard PAC\-Bayes inequality\([30](https://arxiv.org/html/2608.11465#bib.bib33);[36](https://arxiv.org/html/2608.11465#bib.bib39);[5](https://arxiv.org/html/2608.11465#bib.bib14)\)\.
###### Theorem 2\(Classical PAC\-Bayes Bound\)\.
LetPPbe a prior overΘ\\Thetachosen independently ofS∼𝒟nS\\sim\\mathcal\{D\}^\{n\}\. Then, with probability at least1−δ1\-\\deltaover the draw ofSS, simultaneously for all posteriorsQQ,
KL\(L^S\(Q\)∥L\(Q\)\)≤1n\(KL\(Q∥P\)\+log2nδ\)\.\\mathrm\{KL\}\\\!\\left\(\\hat\{L\}\_\{S\}\(Q\)\\,\\middle\\\|\\,L\(Q\)\\right\)\\leq\\frac\{1\}\{n\}\\left\(\\mathrm\{KL\}\(Q\\\|P\)\+\\log\\frac\{2\\sqrt\{n\}\}\{\\delta\}\\right\)\.
The complexity termKL\(Q∥P\)\\mathrm\{KL\}\(Q\\\|P\)is computed on the configuration spaceΘ\\Theta, and therefore combines variation across predictive behaviors with variation among realizations that induce the same behavior\.
### 4\.2PAC\-Bayes Analysis on the Behavior Space
LetπP\\pi\_\{P\}andπQ\\pi\_\{Q\}denote the pushforwards of the prior and posterior under the behavior mapβ\\beta\.
Because empirical and population risk depend only on the induced behavior distribution \(Proposition[1](https://arxiv.org/html/2608.11465#Thmproposition1)\), PAC\-Bayes analysis may be carried out directly on the behavior space\.
###### Theorem 3\(Behavior\-Space Formulation of the PAC\-Bayes Bound\)\.
LetPPbe a prior overΘ\\Thetachosen independently ofS∼𝒟nS\\sim\\mathcal\{D\}^\{n\}\. Then, with probability at least1−δ1\-\\deltaover the draw ofSS, simultaneously for all posteriorsQQ,
KL\(L^S\(Q\)∥L\(Q\)\)≤1n\(KL\(πQ∥πP\)\+log2nδ\)\.\\mathrm\{KL\}\\\!\\left\(\\hat\{L\}\_\{S\}\(Q\)\\,\\middle\\\|\\,L\(Q\)\\right\)\\leq\\frac\{1\}\{n\}\\left\(\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\)\+\\log\\frac\{2\\sqrt\{n\}\}\{\\delta\}\\right\)\.
###### Proof\.
Because𝒦\\mathcal\{K\}is a standard Borel space andπP,πQ∈𝒫\(𝒦\)\\pi\_\{P\},\\pi\_\{Q\}\\in\\mathcal\{P\}\(\\mathcal\{K\}\), the classical PAC\-Bayes theorem applies directly to the measurable hypothesis space𝒦\\mathcal\{K\}\.
Applying Theorem[2](https://arxiv.org/html/2608.11465#Thmtheorem2)with priorπP\\pi\_\{P\}and posteriorπQ\\pi\_\{Q\}yields
KL\(L^S\(πQ\)∥L\(πQ\)\)≤1n\(KL\(πQ∥πP\)\+log2nδ\)\.\\mathrm\{KL\}\\\!\\left\(\\hat\{L\}\_\{S\}\(\\pi\_\{Q\}\)\\,\\middle\\\|\\,L\(\\pi\_\{Q\}\)\\right\)\\leq\\frac\{1\}\{n\}\\left\(\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\)\+\\log\\frac\{2\\sqrt\{n\}\}\{\\delta\}\\right\)\.Proposition[1](https://arxiv.org/html/2608.11465#Thmproposition1)impliesL\(πQ\)=L\(Q\)L\(\\pi\_\{Q\}\)=L\(Q\)andL^S\(πQ\)=L^S\(Q\)\\hat\{L\}\_\{S\}\(\\pi\_\{Q\}\)=\\hat\{L\}\_\{S\}\(Q\), yielding the result\. ∎
Theorem[3](https://arxiv.org/html/2608.11465#Thmtheorem3)is not a new concentration inequality\. Rather, it is the classical PAC\-Bayes theorem applied to the measurable behavior space induced by the behavior map\. Its significance lies not in introducing a new PAC\-Bayes inequality, but in showing that predictive behavior is the natural level at which to measure PAC\-Bayes complexity\. The corresponding complexity term is precisely the behavior\-selection component of the exact decomposition developed in Section[3](https://arxiv.org/html/2608.11465#S3)\. By Theorem[1](https://arxiv.org/html/2608.11465#Thmtheorem1), this quantity is exactly the minimum classical PAC\-Bayes complexity among all posteriors that induce the same distribution over predictive behaviors\.
Consequently, the complexity term in Theorem[3](https://arxiv.org/html/2608.11465#Thmtheorem3)depends only on uncertainty over predictive behavior and is insensitive to how posterior mass is distributed among behaviorally equivalent realizations\.
### 4\.3Complexity Decomposition
By Proposition[3](https://arxiv.org/html/2608.11465#Thmproposition3),
KL\(πQ∥πP\)=KL\(Q∥P\)\+𝒵PB\(Q∥P\)\.\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\)=\\mathrm\{KL\}\(Q\\\|P\)\+\\mathcal\{Z\}\_\{\\mathrm\{PB\}\}\(Q\\\|P\)\.Since𝒵PB\(Q∥P\)≤0\\mathcal\{Z\}\_\{\\mathrm\{PB\}\}\(Q\\\|P\)\\leq 0, the behavior selection complexity never exceeds the classical configuration\-space complexity\.
###### Proposition 4\(Complexity Gap\)\.
For any priorPPand posteriorQQ,KL\(Q∥P\)−KL\(πQ∥πP\)=−𝒵PB\(Q∥P\)\\mathrm\{KL\}\(Q\\\|P\)\-\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\)=\-\\mathcal\{Z\}\_\{\\mathrm\{PB\}\}\(Q\\\|P\)\. In particular,KL\(πQ∥πP\)≤KL\(Q∥P\)\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\)\\leq\\mathrm\{KL\}\(Q\\\|P\)\.
###### Proof\.
Immediate from Proposition[3](https://arxiv.org/html/2608.11465#Thmproposition3)\.
∎
Thus, PAC\-Bayes Z\-information exactly characterizes the gap between classical configuration\-space complexity and behavior selection complexity\. Equivalently, it measures the portion of classical PAC\-Bayes complexity attributable to variation within behavioral fibers\.
### 4\.4Behavior\-Level Bound Expressed Using Z\-Information
Combining Theorem[3](https://arxiv.org/html/2608.11465#Thmtheorem3)with Proposition[3](https://arxiv.org/html/2608.11465#Thmproposition3)yields the following equivalent form\.
###### Corollary 2\(Behavior\-space formulation of the PAC\-Bayes bound with Z\-Information\)\.
With probability at least1−δ1\-\\deltaover the draw ofSS, simultaneously for all posteriorsQQ,
KL\(L^S\(Q\)∥L\(Q\)\)≤1n\(KL\(Q∥P\)\+𝒵PB\(Q∥P\)\+log2nδ\)\.\\mathrm\{KL\}\\\!\\left\(\\hat\{L\}\_\{S\}\(Q\)\\,\\middle\\\|\\,L\(Q\)\\right\)\\leq\\frac\{1\}\{n\}\\left\(\\mathrm\{KL\}\(Q\\\|P\)\+\\mathcal\{Z\}\_\{\\mathrm\{PB\}\}\(Q\\\|P\)\+\\log\\frac\{2\\sqrt\{n\}\}\{\\delta\}\\right\)\.
Because𝒵PB\(Q∥P\)≤0\\mathcal\{Z\}\_\{\\mathrm\{PB\}\}\(Q\\\|P\)\\leq 0, the complexity term in Corollary[2](https://arxiv.org/html/2608.11465#Thmcorollary2)is never larger than the classical PAC\-Bayes complexity term\.
Corollary[2](https://arxiv.org/html/2608.11465#Thmcorollary2)does not introduce a new concentration inequality\. Rather, it rewrites the behavior selection complexity term in a form that makes the realization\-level contribution explicit\.
### 4\.5Interpretation
The results above separate two distinct contributions to PAC\-Bayes complexity:
1. 1\.uncertainty over predictive behaviors;
2. 2\.variation among realizations that implement a fixed behavior\.
The first contribution is captured byKL\(πQ∥πP\)\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\)and is the only component relevant to predictive risk\. The second is captured by−𝒵PB\(Q∥P\)\-\\mathcal\{Z\}\_\{\\mathrm\{PB\}\}\(Q\\\|P\)and reflects the redistribution of posterior mass within behavioral fibers\.
Theorem[1](https://arxiv.org/html/2608.11465#Thmtheorem1)showed thatKL\(πQ∥πP\)\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\)is the minimum classical PAC\-Bayes complexity among all posteriors that induce the same distribution over predictive behaviors\. Consequently, the behavior\-level bound can be viewed as the PAC\-Bayes bound obtained after removing complexity that is irrelevant to prediction while preserving the induced behavior distribution\.
Consequently, PAC\-Bayes Z\-information quantifies the extent to which classical PAC\-Bayes complexity arises from realization multiplicity rather than uncertainty about predictive behavior itself\.
In models with substantial symmetry or over\-parameterization, this realization\-level contribution can be significant\. When posterior and prior induce identical conditional distributions within each fiber,𝒵PB\(Q∥P\)=0\\mathcal\{Z\}\_\{\\mathrm\{PB\}\}\(Q\\\|P\)=0, and classical PAC\-Bayes complexity coincides exactly with behavior selection complexity\.
### 4\.6Summary
Because predictive risk depends only on the induced distribution over behaviors, PAC\-Bayes analysis can be formulated directly on the behavior space associated with the mapβ:Θ→𝒦\\beta:\\Theta\\rightarrow\\mathcal\{K\}\. The resulting complexity term,KL\(πQ∥πP\)\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\), depends only on uncertainty over predictive behavior and is exactly the minimum classical PAC\-Bayes complexity among all posteriors inducing the same behavior distribution\.
PAC\-Bayes Z\-information quantifies the gap between configuration\-space and behavior selection complexity\. Equivalently, it measures the portion of classical PAC\-Bayes complexity attributable to variation within behavioral fibers\. The Behavior\-space formulation of the PAC\-Bayes bound is therefore never looser than the corresponding classical PAC\-Bayes bound and coincides with it precisely when the posterior and prior induce identical conditional distributions within each fiber\.
## 5Realization Multiplicity, Symmetry, and Flat Minima
Sections[2](https://arxiv.org/html/2608.11465#S2)–[4](https://arxiv.org/html/2608.11465#S4)develop the behavior–realization decomposition in a fully measure\-theoretic setting\. The purpose of this section is not to strengthen those results, but to provide geometric intuition for realization multiplicity in continuous configuration spaces\.
Accordingly, the discussion below should be viewed as an interpretation of the measure\-theoretic framework rather than a prerequisite for it\. Geometric assumptions are introduced only where needed to discuss reference measures, tangent directions, and related notions\.
Throughout this section, we use the terms*fiber*and*behavioral equivalence class*interchangeably\.
### 5\.1Fibers as Geometric Objects
LetΘ⊆ℝd\\Theta\\subseteq\\mathbb\{R\}^\{d\}be a continuous configuration space, and letβ:Θ→𝒦\\beta:\\Theta\\rightarrow\\mathcal\{K\}denote the behavior map introduced in Section[2](https://arxiv.org/html/2608.11465#S2)\.
For a behaviork∈𝒦k\\in\\mathcal\{K\}, the corresponding fiber is
Fk=β−1\(k\)\.F\_\{k\}=\\beta^\{\-1\}\(k\)\.
In general, fibers need not possess a smooth manifold structure\. The geometric discussion below is intended only to provide intuition for settings in which fibers admit sufficient regularity for notions such as reference measures and tangent directions to be defined\.
A fiber contains all configurations that induce the same predictive behavior\. Consequently, moving within a fiber changes the realization of a predictor without changing the predictor itself\.
Behavioral equivalence therefore decomposes parameter\-space variation into \(i\) variation across fibers, which changes predictive behavior, and \(ii\) variation within fibers, which changes only the realization of that behavior\. This geometric distinction mirrors the behavior –realization decomposition developed in Sections[3](https://arxiv.org/html/2608.11465#S3)and[4](https://arxiv.org/html/2608.11465#S4)\.
### 5\.2Symmetry and Realization Multiplicity
Many over\-parameterized models admit symmetries that preserve predictive behavior\. Neural networks provide a familiar example\.
Consider a feedforward network with two hidden units\. Simultaneously permuting the incoming and outgoing weights of those units leaves the computed function unchanged, even though the parameter vector changes\. The resulting parameter configurations therefore belong to the same fiber\.
More generally, permutation symmetries, scaling symmetries, redundant hidden units, and other forms of non\-identifiability can produce multiple configurations that realize the same predictive behavior\([32](https://arxiv.org/html/2608.11465#bib.bib37);[11](https://arxiv.org/html/2608.11465#bib.bib22)\)\.
These symmetries generate realization multiplicity: many distinct points in parameter space correspond to a single predictive behavior\.
###### Proposition 5\(Symmetry\-Induced Multiplicity\)\.
LetFkF\_\{k\}be a finite symmetry orbit of sizemm, and supposeP\(⋅∣k\)P\(\\cdot\\mid k\)is uniform on that orbit\. Then
0≤KL\(Q\(⋅∣k\)∥P\(⋅∣k\)\)≤logm\.0\\leq\\mathrm\{KL\}\\\!\\left\(Q\(\\cdot\\mid k\)\\,\\middle\\\|\\,P\(\\cdot\\mid k\)\\right\)\\leq\\log m\.
###### Proof\.
Nonnegativity follows from the basic properties of KL divergence\. The maximum occurs whenQ\(⋅∣k\)Q\(\\cdot\\mid k\)concentrates on a single realization, in which case the divergence from the uniform distribution equalslogm\\log m\.
∎
The proposition shows that increasing the size of a symmetry orbit increases the maximum possible realization\-level contribution to the classical PAC\-Bayes KL divergence while leaving predictive behavior unchanged\. Consequently, realization multiplicity contributes directly to the within\-fiber term identified by PAC\-Bayes Z\-information \(see Appendix[E\.1](https://arxiv.org/html/2608.11465#A5.SS1)for the complete proof\)\.
### 5\.3Reference Measures and Continuous Realization Multiplicity
In continuous models, realization multiplicity is often expressed not through finite symmetry orbits but through extended subsets of parameter space that implement the same predictive behavior\.
To make this idea precise, letμk\\mu\_\{k\}denote a reference measure on a fiberFkF\_\{k\}\. When fibers admit additional geometric structure,μk\\mu\_\{k\}may be chosen to reflect that structure \(for example, as a Hausdorff measure on a smooth submanifold representation of the fiber\)\. The quantity
measures the size of the realization set associated with behaviorkkrelative to the chosen reference measure\.
The numerical value depends on the chosen reference measure, but the underlying interpretation is invariant: some predictive behaviors admit many realizations, whereas others admit comparatively few\. Accordingly, realization multiplicity in continuous spaces should be understood relative to a specified reference measure rather than as an intrinsic geometric volume\.
A fiber with a larger reference\-measure size corresponds, relative to the chosen reference measure, to a behavior that can be realized by a larger set of configurations\. Conversely, a fiber with a smaller reference\-measure size corresponds to a behavior supported by a more restricted set of configurations\.
From this perspective, realization multiplicity may be viewed geometrically as the reference\-measure size of the set of configurations associated with a fixed predictive behavior\. This geometric notion is the continuous\-space analogue of the finite symmetry\-orbit multiplicity described in Proposition[5](https://arxiv.org/html/2608.11465#Thmproposition5)\.
This intuition is reflected formally in the within\-fiber KL divergence of Section[3](https://arxiv.org/html/2608.11465#S3), which compares posterior concentration with the reference distribution supplied by the conditional prior\. In particular, PAC\-Bayes Z\-information is defined entirely through conditional relative entropy and therefore does not require any canonical notion of geometric volume\. The reference measure serves only as geometric intuition for understanding realization multiplicity in continuous spaces\.
### 5\.4Fiber Directions and Flatness
The geometric consequences of realization multiplicity are closely related to, but conceptually distinct from, the notion of flat minima\([17](https://arxiv.org/html/2608.11465#bib.bib23);[22](https://arxiv.org/html/2608.11465#bib.bib24)\)\.
A key distinction should be emphasized\. Behavior\-preserving directions are directions along which predictive behavior remains unchanged\. Flat directions are directions along which the loss changes little or not at all\. Every behavior\-preserving direction is necessarily first\-order flat with respect to any loss that depends only on predictive behavior, although the converse need not hold\.
Letℒ:Θ→ℝ\\mathcal\{L\}:\\Theta\\to\\mathbb\{R\}denote a differentiable objective that depends onθ\\thetaonly through the induced predictive behavior \(for example, population risk under the assumptions of Section[2](https://arxiv.org/html/2608.11465#S2)\)\.
###### Proposition 6\(Behavior\-Preserving Directions are Flat\)\.
Letvvbe tangent to a fiberFkF\_\{k\}at a pointθ\\theta\. Assume that there exists a differentiable curveγ:\(−ε,ε\)→Fk\\gamma:\(\-\\varepsilon,\\varepsilon\)\\to F\_\{k\}withγ\(0\)=θ\\gamma\(0\)=\\thetaandγ′\(0\)=v\\gamma^\{\\prime\}\(0\)=v, and thatℒ\\mathcal\{L\}is differentiable\. Then
∇ℒ\(θ\)⊤v=0\.\\nabla\\mathcal\{L\}\(\\theta\)^\{\\top\}v=0\.
###### Proof\.
Letγ:\(−ε,ε\)→Fk\\gamma:\(\-\\varepsilon,\\varepsilon\)\\to F\_\{k\}be the curve specified in the assumptions\. Since every point onγ\\gammabelongs to the same fiber, the induced predictive behavior is constant along the curve\. Becauseℒ\\mathcal\{L\}depends only on predictive behavior,ℒ\(γ\(t\)\)\\mathcal\{L\}\(\\gamma\(t\)\)is constant intt\.
Differentiating att=0t=0and applying the chain rule gives
0=ddtℒ\(γ\(t\)\)\|t=0=∇ℒ\(θ\)⊤γ′\(0\)=∇ℒ\(θ\)⊤v\.0=\\frac\{d\}\{dt\}\\mathcal\{L\}\(\\gamma\(t\)\)\\Big\|\_\{t=0\}=\\nabla\\mathcal\{L\}\(\\theta\)^\{\\top\}\\gamma^\{\\prime\}\(0\)=\\nabla\\mathcal\{L\}\(\\theta\)^\{\\top\}v\.
∎
This proposition establishes a precise connection between behavioral equivalence and flatness\. Behavior\-preserving directions correspond to exact invariances of the predictive mapping and are therefore necessarily first\-order flat for any behavior\-dependent objective\. They provide an idealized geometric model of realization multiplicity\.
This result is intentionally weaker than the flat\-minimum analyses commonly used in deep learning\([17](https://arxiv.org/html/2608.11465#bib.bib23);[22](https://arxiv.org/html/2608.11465#bib.bib24);[20](https://arxiv.org/html/2608.11465#bib.bib7)\)\. It establishes only first\-order invariance along behavioral fibers and does not require assumptions about optimization trajectories, Hessian spectra, or local curvature\.
In practical models, approximate symmetries may produce nearly flat directions even when exact behavioral invariance does not hold globally\. Conversely, a direction may be locally flat without corresponding to an exact behavioral invariance\. Thus, realization multiplicity and flatness are closely related but not identical concepts\.
### 5\.5Geometric Interpretation of PAC\-Bayes Z\-Information
The geometric structures discussed above influence PAC\-Bayes analysis only through their contribution to the realization\-level term
−𝒵PB\(Q∥P\)\.\-\\mathcal\{Z\}\_\{\\mathrm\{PB\}\}\(Q\\\|P\)\.Accordingly, PAC\-Bayes Z\-information admits a natural geometric interpretation\.
- •In discrete settings, it measures posterior concentration relative to finite symmetry orbits\.
- •In continuous settings, it measures posterior concentration relative to the conditional prior within behavior\-preserving regions of parameter space\.
- •In over\-parameterized models, it quantifies the information\-theoretic consequences of realization multiplicity arising from symmetry, redundancy, and behavior\-preserving directions\.
When the posterior and prior distribute mass similarly within behavioral fibers,−𝒵PB\(Q∥P\)\-\\mathcal\{Z\}\_\{\\mathrm\{PB\}\}\(Q\\\|P\)is small\. When the posterior concentrates on a relatively small subset of realizations relative to the conditional prior, the within\-fiber divergence grows\. Geometrically, this corresponds to concentrating probability mass on a smaller region of the realization set than that favored by the conditional prior\.
Thus, PAC\-Bayes Z\-information does not measure uncertainty over predictive behavior\. Rather, it measures posterior concentration within behavioral fibers relative to the conditional prior and thereby quantifies the realization\-level contribution to classical PAC\-Bayes complexity\.
### 5\.6Summary
Behavioral equivalence induces a geometric structure on parameter space in which multiple realizations correspond to a single predictive behavior\.
In discrete settings, realization multiplicity appears through symmetry orbits\. In continuous settings, it appears through the reference\-measure size of behavior\-preserving regions and the existence of directions tangent to behavioral fibers\.
PAC\-Bayes Z\-information provides the information\-theoretic counterpart of this geometry by quantifying posterior concentration relative to the realization structure encoded by behavioral fibers\.
From this perspective, PAC\-Bayes Z\-information provides the information\-theoretic characterization of realization multiplicity, while the geometry of behavioral fibers provides its structural interpretation\. The variational characterization of Section[4](https://arxiv.org/html/2608.11465#S4)shows that behavior\-selection complexity is obtained by removing within\-fiber variation while preserving predictive behavior\. The geometric picture developed here clarifies what that removed variation represents: symmetry\-related realizations, behavior\-preserving directions, and large realization sets in continuous parameter spaces\.
## 6Stability and Invariance from Realization Multiplicity
The previous sections showed that behavioral equivalence separates predictive behavior from internal realization\. This section examines the consequences of that separation for perturbations that change realizations while leaving behavior unchanged\. The resulting invariance properties clarify how realization multiplicity manifests independently of both optimization dynamics and training\-set perturbations\.
The notion considered here differs from classical algorithmic stability\. Algorithmic stability concerns the sensitivity of learned predictors to perturbations of the training data\([4](https://arxiv.org/html/2608.11465#bib.bib17)\)\. By contrast, we consider perturbations that modify an internal realization while preserving predictive behavior\.
### 6\.1Fiber\-Preserving Perturbations
LetFk=β−1\(k\)F\_\{k\}=\\beta^\{\-1\}\(k\)denote the fiber corresponding to behaviork∈𝒦k\\in\\mathcal\{K\}\. We consider perturbationsθ↦θ′\\theta\\mapsto\\theta^\{\\prime\}such thatθ,θ′∈Fk\\theta,\\theta^\{\\prime\}\\in F\_\{k\}\. Such perturbations move within a behavioral equivalence class and therefore preserve predictive behavior by construction\.
In over\-parameterized models, these perturbations may arise from parameter symmetries, redundant representations, or other forms of non\-identifiability\([32](https://arxiv.org/html/2608.11465#bib.bib37);[11](https://arxiv.org/html/2608.11465#bib.bib22)\)\. For example, permuting hidden units together with their associated outgoing weights changes the parameter vector but leaves the computed function unchanged\. The resulting configurations, therefore, remain in the same fiber\.
### 6\.2Loss Invariance on Fibers
Because empirical and population risks depend only on predictive behavior, they are invariant under fiber\-preserving perturbations\.
###### Proposition 7\(Loss Invariance on Fibers\)\.
For any behaviorkkand anyθ,θ′∈Fk\\theta,\\theta^\{\\prime\}\\in F\_\{k\},
L\(θ\)=L\(θ′\),L^S\(θ\)=L^S\(θ′\)\.L\(\\theta\)=L\(\\theta^\{\\prime\}\),\\qquad\\hat\{L\}\_\{S\}\(\\theta\)=\\hat\{L\}\_\{S\}\(\\theta^\{\\prime\}\)\.
###### Proof\.
Immediate from Lemma[1](https://arxiv.org/html/2608.11465#Thmlemma1)\. ∎
Thus, all configurations belonging to the same fiber have identical empirical and population risks\.
### 6\.3Invariance Within Fibers
Proposition[7](https://arxiv.org/html/2608.11465#Thmproposition7)shows that perturbations that remain within a fiber leave predictive behavior and risk unchanged\. This invariance is a structural consequence of behavioral equivalence\.
Unlike algorithmic stability, it is not a property of a learning algorithm\. Rather, it is a property of the configuration space together with the behavior map\. It reflects the existence of multiple internal realizations that implement the same predictive behavior\.
Section[5](https://arxiv.org/html/2608.11465#S5)showed that, under suitable regularity conditions, directions tangent to fibers are first\-order flat for any behavior\-dependent objective\. The invariance described here may, therefore, be viewed as the finite\-perturbation counterpart of the infinitesimal invariances associated with behavioral equivalence\.
### 6\.4Connection to PAC\-Bayes Z\-Information
Section[4](https://arxiv.org/html/2608.11465#S4)showed that PAC\-Bayes complexity decomposes into a behavior\-selection term and a realization\-level term\. The realization\-level term is
−𝒵PB\(Q∥P\)=𝔼k∼πQ\[KL\(Q\(⋅∣k\)∥P\(⋅∣k\)\)\]\.\-\\mathcal\{Z\}\_\{\\mathrm\{PB\}\}\(Q\\\|P\)=\\mathbb\{E\}\_\{k\\sim\\pi\_\{Q\}\}\\Big\[\\mathrm\{KL\}\\big\(Q\(\\cdot\\mid k\)\\,\\\|\\,P\(\\cdot\\mid k\)\\big\)\\Big\]\.
This quantity measures how strongly the posterior concentrates within fibers relative to the conditional prior\. Since PAC\-Bayes Z\-information is the negative expected within fiber KL divergence, it vanishes precisely when the posterior and prior induce the same conditional distributions within fibers\. More generally, increasingly negative values of𝒵PB\(Q∥P\)\\mathcal\{Z\}\_\{\\mathrm\{PB\}\}\(Q\\\|P\)correspond to greater posterior concentration relative to the conditional prior within fibers\.
Importantly, PAC\-Bayes Z\-information is defined entirely through conditional relative entropy and, therefore, does not depend on any geometric structure of the fibers\. The geometric interpretation developed in Section[5](https://arxiv.org/html/2608.11465#S5)provides intuition, but the quantity itself is measure\-theoretic\.
As discussed in[section5\.4](https://arxiv.org/html/2608.11465#S5.SS4), it is important not to conflate this quantity with geometric flatness\. PAC\-Bayes Z\-information measures a distributional property: posterior concentration relative to a reference distribution within fibers\. Flatness, by contrast, is a geometric property of the parameterization\. Both arise from the same realization structure, but they quantify different aspects of it\.
### 6\.5Interpretation
Behavioral equivalence implies that many distinct configurations may realize the same predictive behavior\. This realization multiplicity has three consequences developed throughout the paper\.
First, it induces a decomposition of PAC\-Bayes complexity into behavior\-level and realization\-level terms\. Second, in continuous parameter spaces and under appropriate regularity assumptions, it gives rise to behavior\-preserving directions that are first\-order flat for behavior\-dependent objectives\. Third, it yields invariance of predictive behavior and risk under fiber\-preserving perturbations\.
These observations suggest a useful conceptual distinction between behavior selection complexity and realization multiplicity\. Highly over\-parameterized models may exhibit substantial realization multiplicity without a corresponding increase in behavior selection complexity\. PAC\-Bayes Z\-information isolates the contribution of this realization\-level variation to classical PAC\-Bayes complexity, while the geometric and invariance viewpoints provide complementary interpretations of the same underlying structure\.
### 6\.6Relation to Classical Stability
Classical algorithmic stability analyzes how learned predictors change when the training data are perturbed\([4](https://arxiv.org/html/2608.11465#bib.bib17)\)\. The notion studied here concerns a different source of variation:
- •Algorithmic stability:the sensitivity of the learning outcome to perturbations of the training set\.
- •Behavioral invariance:the invariance of predictive behavior under perturbations that remain within a behavioral equivalence class\.
The two notions address different aspects of learning systems and should be viewed as complementary rather than competing perspectives\.
### 6\.7Summary
Behavioral equivalence induces invariance under realization\-preserving perturbations\. This invariance is distinct from classical algorithmic stability and arises from realization multiplicity rather than insensitivity to training data\.
PAC\-Bayes Z\-information quantifies how the posterior mass is distributed within behavioral equivalence classes and therefore provides an information\-theoretic measure of realization\-level variation\. Together with the geometric perspective of Section[5](https://arxiv.org/html/2608.11465#S5), these results show that invariance, fiber geometry, and realization\-level PAC\-Bayes complexity are complementary manifestations of the same underlying realization structure induced by behavioral equivalence\.
## 7Related Work
#### PAC\-Bayes theory and complexity\.
PAC\-Bayes theory provides distribution\-dependent generalization bounds through the tradeoff between empirical risk and the KL divergence between posterior and prior distributions over hypotheses\([30](https://arxiv.org/html/2608.11465#bib.bib33);[36](https://arxiv.org/html/2608.11465#bib.bib39);[5](https://arxiv.org/html/2608.11465#bib.bib14)\)\. Extensive work has refined these bounds through localization, data\-dependent priors, and improved complexity control\([24](https://arxiv.org/html/2608.11465#bib.bib2);[19](https://arxiv.org/html/2608.11465#bib.bib4);[35](https://arxiv.org/html/2608.11465#bib.bib8);[42](https://arxiv.org/html/2608.11465#bib.bib5)\)\. Connections with minimum description length \(MDL\) have likewise been widely studied, with KL divergence admitting a coding\-theoretic interpretation as excess description length\([34](https://arxiv.org/html/2608.11465#bib.bib38);[15](https://arxiv.org/html/2608.11465#bib.bib1)\)\.
PAC\-Bayes theory is not restricted to parameter spaces\. The underlying results apply to arbitrary measurable hypothesis spaces, and several influential analyses have been formulated directly in function space\([36](https://arxiv.org/html/2608.11465#bib.bib39)\)\. A major milestone was the demonstration of non\-vacuous PAC\-Bayes bounds for deep neural networks by[14](https://arxiv.org/html/2608.11465#bib.bib10), which stimulated extensive work on the interpretation of PAC\-Bayes complexity in over\-parameterized models\.
#### Function\-space, symmetry, and invariance perspectives\.
Several lines of work have emphasized that learning depends more directly on induced predictors than on particular parameterizations\. Examples include Bayesian neural\-network limits, Gaussian\-process formulations, and neural tangent kernel analyses\([31](https://arxiv.org/html/2608.11465#bib.bib36);[25](https://arxiv.org/html/2608.11465#bib.bib6);[10](https://arxiv.org/html/2608.11465#bib.bib32);[18](https://arxiv.org/html/2608.11465#bib.bib18);[40](https://arxiv.org/html/2608.11465#bib.bib9);[13](https://arxiv.org/html/2608.11465#bib.bib3)\)\.
Related ideas arise in the study of invariance and symmetry\. Neural networks exhibit substantial parameter redundancies arising from permutations, reparameterizations, and other transformations that preserve predictive behavior\. Exploiting such structure can reduce effective complexity and improve generalization bounds\([29](https://arxiv.org/html/2608.11465#bib.bib35);[2](https://arxiv.org/html/2608.11465#bib.bib34)\)\. More broadly, symmetry\-aware and invariance\-aware learning can often be viewed as replacing a large representation space with a smaller space of behaviorally distinguishable predictors\.
#### Over\-parameterization and flatness\.
The relationship between over\-parameterization and generalization has been widely studied through flat minima, sharpness, and loss geometry\([17](https://arxiv.org/html/2608.11465#bib.bib23);[22](https://arxiv.org/html/2608.11465#bib.bib24)\)\. Subsequent work demonstrated that sharpness measures can be sensitive to parameterization and scaling\([11](https://arxiv.org/html/2608.11465#bib.bib22);[20](https://arxiv.org/html/2608.11465#bib.bib7)\)\. Parameter symmetries and non\-identifiability have likewise been studied as sources of equivalent realizations\([32](https://arxiv.org/html/2608.11465#bib.bib37)\)\. Our framework provides a complementary perspective in which flatness is interpreted as one geometric manifestation of realization multiplicity induced by behavioral equivalence\.
#### Information\-theoretic perspectives\.
Shannon entropy and KL divergence quantify uncertainty over observable outcomes and discrepancies between predictive distributions\([38](https://arxiv.org/html/2608.11465#bib.bib19);[39](https://arxiv.org/html/2608.11465#bib.bib11);[8](https://arxiv.org/html/2608.11465#bib.bib20)\), but do not explicitly distinguish uncertainty over predictive behavior from multiplicity among realizations that induce identical behavior\. The terminology introduced here is inspired by the microstate–macrostate distinction in statistical physics and is conceptually related to recent work on zentropy\([27](https://arxiv.org/html/2608.11465#bib.bib31);[28](https://arxiv.org/html/2608.11465#bib.bib30)\), although our development is entirely learning\-theoretic and PAC\-Bayes in nature\.
#### Positioning of this work\.
The distinguishing feature of the present framework is an exact fiberwise decomposition of the classical PAC\-Bayes KL divergence induced by an arbitrary measurable behavior mapβ:Θ→𝒦\.\\beta:\\Theta\\rightarrow\\mathcal\{K\}\.The resulting decomposition separates uncertainty over predictive behavior from variation among behaviorally equivalent realizations, identifies PAC\-Bayes Z\-information as the realization\-level contribution to classical PAC\-Bayes complexity, and yields a behavior\-selection complexity term that is exactly the minimum classical PAC\-Bayes complexity among all posteriors inducing the same distribution over predictive behaviors\.
Viewed through this lens, prior work on function\-space PAC\-Bayes analysis, invariance\-aware learning, and symmetry\-aware complexity control can be seen as instances of a broader principle: learning complexity should depend on distinguishable predictive behavior rather than arbitrary redundancy in its internal realization\.
Existing approaches typically exploit particular sources of redundancy, including group symmetries, parameter\-space invariances, or specific reparameterizations\([29](https://arxiv.org/html/2608.11465#bib.bib35);[2](https://arxiv.org/html/2608.11465#bib.bib34);[32](https://arxiv.org/html/2608.11465#bib.bib37);[11](https://arxiv.org/html/2608.11465#bib.bib22)\)\. The present framework instead begins with an arbitrary measurable notion of behavioral equivalence, from which the behavior–realization decomposition, variational characterization, and behavior\-level complexity arise uniformly\.
## 8Summary and Discussion
### 8\.1Summary
This paper showed that behavioral equivalence induces an exact decomposition of the classical PAC\-Bayes KL divergence into behavior\-selection and realization\-level components\. Within this decomposition, we introduced PAC\-Bayes Z\-information as the negative expected conditional KL divergence, thereby identifying the realization\-level contribution to classical PAC\-Bayes complexity\.
The key observation is that predictive risk depends only on induced predictive behavior, whereas classical PAC\-Bayes complexity is defined on the full configuration space\. By formalizing behavioral equivalence through a measurable behavior map and applying measure disintegration, we established the identity
KL\(Q∥P\)=KL\(πQ∥πP\)−𝒵PB\(Q∥P\),\\mathrm\{KL\}\(Q\\\|P\)=\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\)\-\\mathcal\{Z\}\_\{\\mathrm\{PB\}\}\(Q\\\|P\),where
𝒵PB\(Q∥P\)=−𝔼k∼πQ\[KL\(Q\(⋅∣k\)∥P\(⋅∣k\)\)\]\.\\mathcal\{Z\}\_\{\\mathrm\{PB\}\}\(Q\\\|P\)=\-\\mathbb\{E\}\_\{k\\sim\\pi\_\{Q\}\}\\Big\[\\mathrm\{KL\}\\big\(Q\(\\cdot\\mid k\)\\,\\\|\\,P\(\\cdot\\mid k\)\\big\)\\Big\]\.
This identity reveals that classical PAC\-Bayes complexity consists of two conceptually distinct contributions: uncertainty over predictive behavior and variation among behaviorally equivalent realizations\. PAC\-Bayes Z\-information is precisely the negative expected within\-fiber KL divergence and therefore quantifies the realization\-level contribution exactly\.
The decomposition identifies the behavior\-selection component,KL\(πQ∥πP\)\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\), as the complexity associated with predictive behavior itself\. We further showed that this quantity admits an exact variational characterization: it is the minimum classical PAC\-Bayes complexity among all posteriors inducing the same distribution over predictive behaviors\.
Since
KL\(πQ∥πP\)≤KL\(Q∥P\),\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\)\\leq\\mathrm\{KL\}\(Q\\\|P\),the behavior\-selection complexity is never larger than the classical configuration\-space complexity, and the gap is quantified exactly by−𝒵PB\(Q∥P\)\-\\mathcal\{Z\}\_\{\\mathrm\{PB\}\}\(Q\\\|P\)\.
Finally, Sections[5](https://arxiv.org/html/2608.11465#S5)and[6](https://arxiv.org/html/2608.11465#S6)showed that the same behavioral\-equivalence structure also gives rise to geometric and invariance\-based interpretations through symmetry, realization multiplicity, behavior\-preserving directions, and fiber\-preserving perturbations\. Together, these results provide complementary information\-theoretic, geometric, and structural perspectives on the role of realization multiplicity in classical PAC\-Bayes complexity\.
### 8\.2Discussion
The central contribution of this work is the identification of an exact structural decomposition of PAC\-Bayes complexity induced by behavioral equivalence\. Although the underlying relative\-entropy identity follows from the classical chain rule under disintegration, its interpretation through behavioral equivalence reveals a distinction between behavior\-selection complexity and realization\-level complexity that is not explicit in existing PAC\-Bayes analyses\. The novelty of the present framework lies in recognizing that, when the measurable map is taken to be predictive behavior, the resulting terms admit a natural learning\-theoretic interpretation\. Behavioral equivalence reveals that the classical PAC\-Bayes KL divergence naturally separates into two conceptually distinct components: \(i\) uncertainty over predictive behavior and \(ii\) variation among realizations that implement the same behavior\.
Only the first component directly affects predictive risk\. The second reflects how posterior mass is distributed among behaviorally equivalent realizations and therefore contributes to configuration\-space complexity without altering prediction\. The behavior–realization decomposition makes this distinction explicit, identifies PAC\-Bayes Z\-information as an exact measure of the realization\-level contribution, and reveals behavior\-selection complexity as the irreducible component of classical PAC\-Bayes complexity associated with predictive behavior\.
Viewed through this lens, the behavior\-level PAC\-Bayes formulation should not be interpreted as an alternative to classical PAC\-Bayes analysis, but as its canonical representation on the behavior space\. The corresponding complexity term,KL\(πQ∥πP\)\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\), is precisely the minimum classical PAC\-Bayes complexity among all posteriors inducing the same distribution over predictive behaviors\. PAC\-Bayes Z\-information then quantifies exactly how much additional complexity arises solely from redistributing posterior mass among behaviorally equivalent realizations\.
This perspective provides a principled way to distinguish complexity that is intrinsic to predictive behavior from complexity that reflects only the choice of realization\. As a result, it offers a unified lens for interpreting PAC\-Bayes analyses of over\-parameterized models, comparing alternative parameterizations of the same predictor, and motivating future complexity measures that operate directly on predictive behavior rather than internal realization\.
Beyond its implications for PAC\-Bayes analysis, the framework provides a common language for several phenomena that are often studied separately, including parameter symmetries, realization multiplicity, behavior\-preserving perturbations, and geometric flatness\. Rather than introducing these concepts independently, behavioral equivalence reveals them as complementary manifestations of the same underlying realization structure\.
Several limitations should also be emphasized\. First, the present analysis relies on exact behavioral equivalence\. Extending the framework to approximate behavioral equivalence, where predictors induce nearly identical rather than identical behaviors, remains an important direction for future work\. Second, the theory is structural rather than causal: it identifies realization\-level complexity exactly but does not establish a causal relationship between realization multiplicity and improved generalization\. Finally, PAC\-Bayes Z\-information is a measure\-theoretic quantity defined through conditional relative entropy and should not be conflated with geometric quantities such as flatness, although both arise naturally from the same fiber structure induced by behavioral equivalence\. Clarifying the precise relationships among information\-theoretic, geometric, and optimization\-based notions of realization multiplicity remains an important direction for future research\.
### 8\.3Future Work
The framework suggests several directions for future work:
1. 1\.Developing computable estimators or practical proxies for PAC\-Bayes Z\-information in modern neural networks\.
2. 2\.Extending the theory to approximate behavioral equivalence, where predictors are nearly indistinguishable rather than exactly identical\.
3. 3\.Investigating whether optimization procedures implicitly favor regions of high realization multiplicity\.
4. 4\.Extending the behavior\-realization decomposition to other learning\-theoretic frameworks, including stability\-based, information\-theoretic, and compression\-based analyses\.
5. 5\.Empirically studying relationships among realization multiplicity, generalization, robustness, calibration, and uncertainty estimation in over\-parameterized models\.
### 8\.4Conclusion
The behavior–realization decomposition shows that predictive behavior, rather than internal realization, provides the natural level at which to analyze PAC\-Bayes complexity\. This paper demonstrates that the distinction between behavior\-selection complexity and realization\-level complexity arises from the classical PAC\-Bayes framework through an exact decomposition of relative entropy together with its variational characterization\. By making this structure explicit, the resulting framework separates uncertainty over predictive behavior from realization multiplicity and provides a unified measure\-theoretic interpretation of complexity, symmetry, and invariance in over\-parameterized learning systems\. We expect this framework to provide a foundation for future work on behavior\-space learning theory, approximate behavioral equivalence, invariance, and complexity measures for modern over\-parameterized models\.
## Acknowledgments
This work was supported in part by grants from the National Science Foundation \(NSF\) and Penn State Clinical and Translational Science Institute \(CTSI\)\.
## References
- S\. Amari and H\. NagaokaMethods of information geometry\.American Mathematical Society\.Note:Translated from the 1993 Japanese editionCited by:[§D\.4](https://arxiv.org/html/2608.11465#A4.SS4.p1.1)\.
- Behboodiet al\.\(2022\)A\. Behboodi, G\. Cesa, and T\. S\. CohenA pac\-bayesian generalization bound for equivariant networks\.Advances in Neural Information Processing Systems35,pp\. 5654–5668\.Cited by:[§1\.1](https://arxiv.org/html/2608.11465#S1.SS1.p5.1),[§7](https://arxiv.org/html/2608.11465#S7.SS0.SSS0.Px2.p2.1),[§7](https://arxiv.org/html/2608.11465#S7.SS0.SSS0.Px5.p3.1)\.
- Bogachev \(2007\)V\. I\. BogachevMeasure theory\.Springer\.Cited by:[§A\.2](https://arxiv.org/html/2608.11465#A1.SS2.p3.1),[§A\.3](https://arxiv.org/html/2608.11465#A1.SS3.p2.1),[Appendix A](https://arxiv.org/html/2608.11465#A1.p1.1),[§B\.1](https://arxiv.org/html/2608.11465#A2.SS1.p2.1.1),[Appendix B](https://arxiv.org/html/2608.11465#A2.p1.1),[§G\.1](https://arxiv.org/html/2608.11465#A7.SS1.p4.1),[§2\.5](https://arxiv.org/html/2608.11465#S2.SS5.p3.3),[§3\.1](https://arxiv.org/html/2608.11465#S3.SS1.p4.1)\.
- Bousquet and Elisseeff \(2002\)O\. Bousquet and A\. ElisseeffStability and generalization\.J\. Mach\. Learn\. Res\.2,pp\. 499–526\.External Links:[Link](https://jmlr.org/papers/v2/bousquet02a.html)Cited by:[§F\.4](https://arxiv.org/html/2608.11465#A6.SS4.p2.1),[Appendix F](https://arxiv.org/html/2608.11465#A6.p2.1),[§2\.6](https://arxiv.org/html/2608.11465#S2.SS6.p1.1),[§6\.6](https://arxiv.org/html/2608.11465#S6.SS6.p1.1),[§6](https://arxiv.org/html/2608.11465#S6.p2.1)\.
- Catoni \(2007\)O\. CatoniPac\-bayesian supervised classification: the thermodynamics of statistical learning\.IMS Lecture Notes Monograph Series56,pp\. 1–163\.External Links:ISSN 0749\-2170,[Link](http://dx.doi.org/10.1214/074921707000000391),[Document](https://dx.doi.org/10.1214/074921707000000391)Cited by:[§1\.1](https://arxiv.org/html/2608.11465#S1.SS1.p1.1),[§2\.1](https://arxiv.org/html/2608.11465#S2.SS1.p6.2),[§4\.1](https://arxiv.org/html/2608.11465#S4.SS1.p1.1),[§7](https://arxiv.org/html/2608.11465#S7.SS0.SSS0.Px1.p1.1)\.
- Chang and Pollard \(1997\)J\. T\. Chang and D\. PollardConditioning as disintegration\.Statistica Neerlandica51\(3\),pp\. 287–317\.Cited by:[§1\.2](https://arxiv.org/html/2608.11465#S1.SS2.p6.1)\.
- Chaudhariet al\.\(2019\)P\. Chaudhari, A\. Choromanska, S\. Soatto, Y\. LeCun, C\. Baldassi, C\. Borgs, J\. Chayes, L\. Sagun, and R\. ZecchinaEntropy\-SGD: biasing gradient descent into wide valleys\.arXiv preprint arXiv:1611\.01838\.Cited by:[§1\.1](https://arxiv.org/html/2608.11465#S1.SS1.p3.1)\.
- Cover and Thomas \(2006\)T\. M\. Cover and J\. A\. ThomasElements of information theory\.2 edition,Wiley\-Interscience\.Cited by:[§D\.2](https://arxiv.org/html/2608.11465#A4.SS2.p3.3),[§D\.4](https://arxiv.org/html/2608.11465#A4.SS4.p1.1),[§1\.1](https://arxiv.org/html/2608.11465#S1.SS1.p1.1),[§7](https://arxiv.org/html/2608.11465#S7.SS0.SSS0.Px4.p1.1)\.
- Csiszár \(1975\)I\. CsiszárI\-divergence geometry of probability distributions and minimization problems\.The annals of probability,pp\. 146–158\.Cited by:[§D\.4](https://arxiv.org/html/2608.11465#A4.SS4.p1.1)\.
- de G\. Matthewset al\.\(2018\)A\. G\. de G\. Matthews, J\. Hron, M\. Rowland, R\. E\. Turner, and Z\. GhahramaniGaussian process behaviour in wide deep neural networks\.In6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 \- May 3, 2018, Conference Track Proceedings,External Links:[Link](https://openreview.net/forum?id=H1-nGgWC-)Cited by:[§7](https://arxiv.org/html/2608.11465#S7.SS0.SSS0.Px2.p1.1)\.
- Dinhet al\.\(2017\)L\. Dinh, R\. Pascanu, S\. Bengio, and Y\. BengioSharp minima can generalize for deep nets\.InProceedings of the 34th International Conference on Machine Learning, ICML 2017, Sydney, NSW, Australia, 6\-11 August 2017,D\. Precup and Y\. W\. Teh \(Eds\.\),Proceedings of Machine Learning Research, Vol\.70,pp\. 1019–1028\.External Links:[Link](http://proceedings.mlr.press/v70/dinh17b.html)Cited by:[§E\.4](https://arxiv.org/html/2608.11465#A5.SS4.p2.1),[§F\.3](https://arxiv.org/html/2608.11465#A6.SS3.p1.1),[§F\.5](https://arxiv.org/html/2608.11465#A6.SS5.p2.1),[§1\.1](https://arxiv.org/html/2608.11465#S1.SS1.p3.1),[§2\.2](https://arxiv.org/html/2608.11465#S2.SS2.p7.1),[§5\.2](https://arxiv.org/html/2608.11465#S5.SS2.p3.1),[§6\.1](https://arxiv.org/html/2608.11465#S6.SS1.p2.1),[§7](https://arxiv.org/html/2608.11465#S7.SS0.SSS0.Px3.p1.1),[§7](https://arxiv.org/html/2608.11465#S7.SS0.SSS0.Px5.p3.1)\.
- do Carmo and Flaherty \(1992\)M\. P\. do Carmo and F\. FlahertyRiemannian geometry\.Birkhäuser,Boston, MA\.External Links:ISBN 978\-0\-8176\-3490\-2Cited by:[§E\.2](https://arxiv.org/html/2608.11465#A5.SS2.p1.1),[Appendix E](https://arxiv.org/html/2608.11465#A5.p3.2)\.
- Dziugaiteet al\.\(2021\)G\. K\. Dziugaite, K\. Hsu, W\. Gharbieh, G\. Arpino, and D\. RoyOn the role of data in pac\-bayes bounds\.InInternational Conference on Artificial Intelligence and Statistics,pp\. 604–612\.Cited by:[§7](https://arxiv.org/html/2608.11465#S7.SS0.SSS0.Px2.p1.1)\.
- Dziugaite and Roy \(2017\)G\. K\. Dziugaite and D\. M\. RoyComputing nonvacuous generalization bounds for deep \(stochastic\) neural networks with many more parameters than training data\.InProceedings of the Thirty\-Third Conference on Uncertainty in Artificial Intelligence, UAI 2017, Sydney, Australia, August 11\-15, 2017,Cited by:[§7](https://arxiv.org/html/2608.11465#S7.SS0.SSS0.Px1.p2.1)\.
- Grünwald \(2007\)P\. D\. GrünwaldThe minimum description length principle\.MIT press\.Cited by:[§1\.1](https://arxiv.org/html/2608.11465#S1.SS1.p1.1),[§7](https://arxiv.org/html/2608.11465#S7.SS0.SSS0.Px1.p1.1)\.
- Hardtet al\.\(2016\)M\. Hardt, B\. Recht, and Y\. SingerTrain faster, generalize better: stability of stochastic gradient descent\.InProceedings of the 33rd International Conference on Machine Learning \(ICML\),PMLR, Vol\.48,pp\. 1225–1234\.Cited by:[§F\.4](https://arxiv.org/html/2608.11465#A6.SS4.p2.1),[Appendix F](https://arxiv.org/html/2608.11465#A6.p2.1)\.
- Hochreiter and Schmidhuber \(1994\)S\. Hochreiter and J\. SchmidhuberSimplifying neural nets by discovering flat minima\.InAdvances in Neural Information Processing Systems 7, \[NIPS Conference, Denver, Colorado, USA, 1994\],G\. Tesauro, D\. S\. Touretzky, and T\. K\. Leen \(Eds\.\),pp\. 529–536\.External Links:[Link](https://proceedings.neurips.cc/paper/_files/paper/1994/hash/01882513d5fa7c329e940dda99b12147-Abstract.html)Cited by:[§E\.3](https://arxiv.org/html/2608.11465#A5.SS3.p8.1),[§F\.3](https://arxiv.org/html/2608.11465#A6.SS3.p1.1),[§1\.1](https://arxiv.org/html/2608.11465#S1.SS1.p3.1),[§5\.4](https://arxiv.org/html/2608.11465#S5.SS4.p1.1),[§5\.4](https://arxiv.org/html/2608.11465#S5.SS4.p9.1),[§7](https://arxiv.org/html/2608.11465#S7.SS0.SSS0.Px3.p1.1)\.
- Jacotet al\.\(2018\)A\. Jacot, C\. Hongler, and F\. GabrielNeural tangent kernel: convergence and generalization in neural networks\.\.InNeurIPS,S\. Bengio, H\. M\. Wallach, H\. Larochelle, K\. Grauman, N\. Cesa\-Bianchi, and R\. Garnett \(Eds\.\),pp\. 8580–8589\.External Links:[Link](http://dblp.uni-trier.de/db/conf/nips/nips2018.html#JacotHG18)Cited by:[§7](https://arxiv.org/html/2608.11465#S7.SS0.SSS0.Px2.p1.1)\.
- Janget al\.\(2023\)K\. Jang, K\. Jun, I\. Kuzborskij, and F\. OrabonaTighter pac\-bayes bounds through coin\-betting\.InThe Thirty Sixth Annual Conference on Learning Theory, COLT 2023, 12\-15 July 2023, Bangalore, India,G\. Neu and L\. Rosasco \(Eds\.\),Proceedings of Machine Learning Research, Vol\.195,pp\. 2240–2264\.External Links:[Link](https://proceedings.mlr.press/v195/jang23a.html)Cited by:[§7](https://arxiv.org/html/2608.11465#S7.SS0.SSS0.Px1.p1.1)\.
- Jianget al\.\(2020\)Y\. Jiang, B\. Neyshabur, H\. Mobahi, D\. Krishnan, and S\. BengioFantastic generalization measures and where to find them\.InInternational Conference on Learning Representations,External Links:[Link](http://arxiv.org/abs/1912.02178),1912\.02178Cited by:[§E\.3](https://arxiv.org/html/2608.11465#A5.SS3.p8.1),[§1\.1](https://arxiv.org/html/2608.11465#S1.SS1.p3.1),[§5\.4](https://arxiv.org/html/2608.11465#S5.SS4.p9.1),[§7](https://arxiv.org/html/2608.11465#S7.SS0.SSS0.Px3.p1.1)\.
- Kallenberg \(2002\)O\. KallenbergFoundations of modern probability\.Second edition,Springer,New York\.External Links:ISBN 0387949577 9780387949574 0387953132 9780387953137 9787030166722 7030166728Cited by:[§A\.2](https://arxiv.org/html/2608.11465#A1.SS2.p3.1),[§A\.3](https://arxiv.org/html/2608.11465#A1.SS3.p2.1),[§A\.4](https://arxiv.org/html/2608.11465#A1.SS4.p3.1.1),[Appendix A](https://arxiv.org/html/2608.11465#A1.p1.1),[§B\.1](https://arxiv.org/html/2608.11465#A2.SS1.p2.1.1),[Appendix B](https://arxiv.org/html/2608.11465#A2.p1.1),[§D\.4](https://arxiv.org/html/2608.11465#A4.SS4.p1.1),[§G\.1](https://arxiv.org/html/2608.11465#A7.SS1.p4.1),[§1\.2](https://arxiv.org/html/2608.11465#S1.SS2.p6.1),[§2\.2](https://arxiv.org/html/2608.11465#S2.SS2.p4.1),[§2\.5](https://arxiv.org/html/2608.11465#S2.SS5.p3.3),[§3\.1](https://arxiv.org/html/2608.11465#S3.SS1.p4.1),[footnote 1](https://arxiv.org/html/2608.11465#footnote1)\.
- Keskaret al\.\(2017\)N\. S\. Keskar, D\. Mudigere, J\. Nocedal, M\. Smelyanskiy, and P\. T\. P\. TangOn large\-batch training for deep learning: generalization gap and sharp minima\.In5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24\-26, 2017, Conference Track Proceedings,External Links:[Link](https://openreview.net/forum?id=H1oyRlYgg)Cited by:[§E\.3](https://arxiv.org/html/2608.11465#A5.SS3.p8.1),[§F\.3](https://arxiv.org/html/2608.11465#A6.SS3.p1.1),[§1\.1](https://arxiv.org/html/2608.11465#S1.SS1.p3.1),[§5\.4](https://arxiv.org/html/2608.11465#S5.SS4.p1.1),[§5\.4](https://arxiv.org/html/2608.11465#S5.SS4.p9.1),[§7](https://arxiv.org/html/2608.11465#S7.SS0.SSS0.Px3.p1.1)\.
- Klenke \(2013\)A\. KlenkeProbability theory: a comprehensive course\.Springer\.Cited by:[§A\.3](https://arxiv.org/html/2608.11465#A1.SS3.p2.1),[Appendix A](https://arxiv.org/html/2608.11465#A1.p1.1)\.
- Kuzborskijet al\.\(2024\)I\. Kuzborskij, K\. Jun, Y\. Wu, K\. Jang, and F\. OrabonaBetter\-than\-kl pac\-bayes bounds\.InThe Thirty Seventh Annual Conference on Learning Theory, June 30 \- July 3, 2023, Edmonton, Canada,S\. Agrawal and A\. Roth \(Eds\.\),Proceedings of Machine Learning Research, Vol\.247,pp\. 3325–3352\.External Links:[Link](https://proceedings.mlr.press/v247/kuzborskij24a.html)Cited by:[§7](https://arxiv.org/html/2608.11465#S7.SS0.SSS0.Px1.p1.1)\.
- Leeet al\.\(2017\)J\. Lee, Y\. Bahri, R\. Novak, S\. S\. Schoenholz, J\. Pennington, and J\. Sohl\-DicksteinDeep neural networks as gaussian processes\.CoRRabs/1711\.00165\.External Links:[Link](http://arxiv.org/abs/1711.00165),1711\.00165Cited by:[§7](https://arxiv.org/html/2608.11465#S7.SS0.SSS0.Px2.p1.1)\.
- Lee \(2018\)J\. M\. LeeIntroduction to riemannian manifolds\.2nd edition,Graduate Texts in Mathematics, Vol\.176,Springer,New York, NY\.External Links:ISBN 978\-3\-319\-91755\-9,[Document](https://dx.doi.org/10.1007/978-3-319-91755-9)Cited by:[§E\.2](https://arxiv.org/html/2608.11465#A5.SS2.p1.1),[Appendix E](https://arxiv.org/html/2608.11465#A5.p3.2)\.
- Liuet al\.\(2022\)Z\. Liu, Y\. Wang, and S\. ShangZentropy theory for positive and negative thermal expansion\.Journal of Phase Equilibria and Diffusion43\(6\),pp\. 598–605\.Cited by:[§7](https://arxiv.org/html/2608.11465#S7.SS0.SSS0.Px4.p1.1)\.
- Liu \(2024\)Z\. LiuZentropy: theory and fundamentals\.CRC Press\.Cited by:[§7](https://arxiv.org/html/2608.11465#S7.SS0.SSS0.Px4.p1.1)\.
- Lyleet al\.\(2020\)C\. Lyle, M\. van der Wilk, M\. Kwiatkowska, Y\. Gal, and B\. Bloem\-ReddyOn the benefits of invariance in neural networks\.arXiv preprint arXiv:2005\.00178\.Cited by:[§1\.1](https://arxiv.org/html/2608.11465#S1.SS1.p5.1),[§7](https://arxiv.org/html/2608.11465#S7.SS0.SSS0.Px2.p2.1),[§7](https://arxiv.org/html/2608.11465#S7.SS0.SSS0.Px5.p3.1)\.
- McAllester \(1998\)D\. A\. McAllesterSome pac\-bayesian theorems\.InProceedings of the Eleventh Annual Conference on Computational Learning Theory,COLT’ 98,New York, NY, USA,pp\. 230–234\.External Links:ISBN 1581130570,[Link](https://doi.org/10.1145/279943.279989),[Document](https://dx.doi.org/10.1145/279943.279989)Cited by:[§C\.2](https://arxiv.org/html/2608.11465#A3.SS2.p2.1.1),[§1\.1](https://arxiv.org/html/2608.11465#S1.SS1.p1.1),[§2\.1](https://arxiv.org/html/2608.11465#S2.SS1.p6.2),[§2\.6](https://arxiv.org/html/2608.11465#S2.SS6.p1.1),[§4\.1](https://arxiv.org/html/2608.11465#S4.SS1.p1.1),[§7](https://arxiv.org/html/2608.11465#S7.SS0.SSS0.Px1.p1.1)\.
- Neal \(1995\)R\. M\. NealBayesian learning for neural networks\.Ph\.D\. Thesis,University of Toronto, Canada\.External Links:[Link](https://librarysearch.library.utoronto.ca/permalink/01UTORONTO/_INST/14bjeso/alma991106438365706196)Cited by:[§7](https://arxiv.org/html/2608.11465#S7.SS0.SSS0.Px2.p1.1)\.
- Neyshaburet al\.\(2017\)B\. Neyshabur, S\. Bhojanapalli, D\. McAllester, and N\. SrebroExploring generalization in deep learning\.InProceedings of the 31st International Conference on Neural Information Processing Systems \(NIPS\),pp\. 5949–5958\.Cited by:[§E\.4](https://arxiv.org/html/2608.11465#A5.SS4.p2.1),[§F\.5](https://arxiv.org/html/2608.11465#A6.SS5.p2.1),[§1\.1](https://arxiv.org/html/2608.11465#S1.SS1.p3.1),[§2\.2](https://arxiv.org/html/2608.11465#S2.SS2.p7.1),[§5\.2](https://arxiv.org/html/2608.11465#S5.SS2.p3.1),[§6\.1](https://arxiv.org/html/2608.11465#S6.SS1.p2.1),[§7](https://arxiv.org/html/2608.11465#S7.SS0.SSS0.Px3.p1.1),[§7](https://arxiv.org/html/2608.11465#S7.SS0.SSS0.Px5.p3.1)\.
- Parthasarathy \(1967\)K\. R\. ParthasarathyProbability measures on metric spaces\.Academic Press,New York\.Cited by:[§A\.3](https://arxiv.org/html/2608.11465#A1.SS3.p2.1),[Appendix A](https://arxiv.org/html/2608.11465#A1.p1.1),[Appendix B](https://arxiv.org/html/2608.11465#A2.p1.1),[§G\.1](https://arxiv.org/html/2608.11465#A7.SS1.p4.1),[§2\.2](https://arxiv.org/html/2608.11465#S2.SS2.p4.1),[§2\.5](https://arxiv.org/html/2608.11465#S2.SS5.p3.3),[footnote 1](https://arxiv.org/html/2608.11465#footnote1)\.
- Rissanen \(1978\)J\. RissanenModeling by shortest data description\.Automatica14\(5\),pp\. 465–471\.Cited by:[§1\.1](https://arxiv.org/html/2608.11465#S1.SS1.p1.1),[§7](https://arxiv.org/html/2608.11465#S7.SS0.SSS0.Px1.p1.1)\.
- Rivasplataet al\.\(2020\)O\. Rivasplata, I\. Kuzborskij, C\. Szepesvári, and J\. Shawe\-TaylorPAC\-bayes analysis beyond the usual bounds\.InAdvances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6\-12, 2020, virtual,H\. Larochelle, M\. Ranzato, R\. Hadsell, M\. Balcan, and H\. Lin \(Eds\.\),External Links:[Link](https://proceedings.neurips.cc/paper/2020/hash/c3992e9a68c5ae12bd18488bc579b30d-Abstract.html)Cited by:[§7](https://arxiv.org/html/2608.11465#S7.SS0.SSS0.Px1.p1.1)\.
- Seeger \(2002\)M\. W\. SeegerPAC\-bayesian generalisation error bounds for gaussian process classification\.J\. Mach\. Learn\. Res\.3,pp\. 233–269\.External Links:[Link](https://jmlr.org/papers/v3/seeger02a.html)Cited by:[§C\.2](https://arxiv.org/html/2608.11465#A3.SS2.p2.1.1),[§1\.1](https://arxiv.org/html/2608.11465#S1.SS1.p1.1),[§1\.1](https://arxiv.org/html/2608.11465#S1.SS1.p5.1),[§2\.1](https://arxiv.org/html/2608.11465#S2.SS1.p6.2),[§4\.1](https://arxiv.org/html/2608.11465#S4.SS1.p1.1),[§7](https://arxiv.org/html/2608.11465#S7.SS0.SSS0.Px1.p1.1),[§7](https://arxiv.org/html/2608.11465#S7.SS0.SSS0.Px1.p2.1)\.
- Shalev\-Shwartz and Ben\-David \(2014\)S\. Shalev\-Shwartz and S\. Ben\-DavidUnderstanding machine learning \- from theory to algorithms\.\.Cambridge University Press\.External Links:ISBN 978\-1\-10\-705713\-5Cited by:[§2\.1](https://arxiv.org/html/2608.11465#S2.SS1.p1.1)\.
- Shannon \(1948\)C\. E\. ShannonA mathematical theory of communication\.The Bell System Technical Journal27\(3\),pp\. 379–423\.External Links:[Document](https://dx.doi.org/10.1002/j.1538-7305.1948.tb01338.x)Cited by:[§1\.1](https://arxiv.org/html/2608.11465#S1.SS1.p1.1),[§7](https://arxiv.org/html/2608.11465#S7.SS0.SSS0.Px4.p1.1)\.
- Shannon \(1951\)C\. E\. ShannonPrediction and entropy of printed english\.Bell system technical journal30\(1\),pp\. 50–64\.Cited by:[§1\.1](https://arxiv.org/html/2608.11465#S1.SS1.p1.1),[§7](https://arxiv.org/html/2608.11465#S7.SS0.SSS0.Px4.p1.1)\.
- Sunet al\.\(2019\)S\. Sun, G\. Zhang, J\. Shi, and R\. GrosseFunctional variational Bayesian neural networks\.InInternational Conference on Learning Representations,Cited by:[§7](https://arxiv.org/html/2608.11465#S7.SS0.SSS0.Px2.p1.1)\.
- Vapnik \(1998\)V\. N\. VapnikStatistical learning theory\.Wiley\.Cited by:[§2\.1](https://arxiv.org/html/2608.11465#S2.SS1.p1.1),[§2\.6](https://arxiv.org/html/2608.11465#S2.SS6.p1.1)\.
- Viallardet al\.\(2024\)P\. Viallard, R\. Emonet, A\. Habrard, E\. Morvant, and V\. ZantedeschiLeveraging pac\-bayes theory and gibbs distributions for generalization bounds with complexity measures\.InInternational Conference on Artificial Intelligence and Statistics, 2\-4 May 2024, Palau de Congressos, Valencia, Spain,S\. Dasgupta, S\. Mandt, and Y\. Li \(Eds\.\),Proceedings of Machine Learning Research, Vol\.238,pp\. 3007–3015\.External Links:[Link](https://proceedings.mlr.press/v238/viallard24a.html)Cited by:[§7](https://arxiv.org/html/2608.11465#S7.SS0.SSS0.Px1.p1.1)\.
## Appendix
## Appendix AMeasure\-Theoretic Preliminaries
This appendix reviews the measure\-theoretic concepts underlying the framework developed in the main text, including behavioral equivalence, fibers, pushforward measures, and disintegration\. These constructions are standard in probability and measure theory\([21](https://arxiv.org/html/2608.11465#bib.bib29);[33](https://arxiv.org/html/2608.11465#bib.bib28);[3](https://arxiv.org/html/2608.11465#bib.bib15);[23](https://arxiv.org/html/2608.11465#bib.bib25)\)\. The purpose of this appendix is not to develop new measure theory, but to establish notation and provide the background needed for the behavior–realization decomposition used throughout the paper\.
### A\.1Behavioral Equivalence and Fibers
Throughout the paper, predictive behavior is represented by a measurable behavior mapβ:Θ→𝒦\\beta:\\Theta\\rightarrow\\mathcal\{K\}, whereΘ\\Thetais the configuration space, and𝒦\\mathcal\{K\}is a space of predictive behaviors\.
Behavioral equivalence is induced by this map:
θ∼θ′⟺β\(θ\)=β\(θ′\)\.\\theta\\sim\\theta^\{\\prime\}\\quad\\Longleftrightarrow\\quad\\beta\(\\theta\)=\\beta\(\\theta^\{\\prime\}\)\.Thus, two configurations are behaviorally equivalent precisely when they induce the same predictive behavior\.
For a behaviork∈𝒦k\\in\\mathcal\{K\}, the associated fiber is
Fk=β−1\(k\)=\{θ∈Θ:β\(θ\)=k\}\.F\_\{k\}=\\beta^\{\-1\}\(k\)=\\\{\\theta\\in\\Theta:\\beta\(\\theta\)=k\\\}\.
A fiber, therefore, contains every configuration that realizes the same predictive behavior\.
#### Example\.
A familiar instance of behavioral equivalence is provided by hidden unit permutation symmetry in neural networks, where permuting exchangeable hidden units leaves the input–output mapping unchanged\. Such parameter configurations, therefore, belong to the same behavioral fiber\. A detailed illustration appears in Section[3](https://arxiv.org/html/2608.11465#S3), and the corresponding realization\-level computation is worked out in Appendix[H\.2](https://arxiv.org/html/2608.11465#A8.SS2)\.
### A\.2Pushforward Measures and Behavior Distributions
LetQ∈𝒫\(Θ\)Q\\in\\mathcal\{P\}\(\\Theta\)be a probability distribution over configurations\.
The behavior map induces a probability distribution over behaviors through the pushforward measure \(also called the image measure\)
πQ=β\#Q=Q∘β−1\.\\pi\_\{Q\}=\\beta\_\{\\\#\}Q=Q\\circ\\beta^\{\-1\}\.For any measurable subsetA⊆𝒦A\\subseteq\\mathcal\{K\},
πQ\(A\)=Q\(β−1\(A\)\)\.\\pi\_\{Q\}\(A\)=Q\(\\beta^\{\-1\}\(A\)\)\.Intuitively,πQ\\pi\_\{Q\}records how much probability massQQassigns to each predictive behavior while ignoring how that mass is distributed among equivalent realizations within a fiber\.
The notationβ\#Q\\beta\_\{\\\#\}Qis standard shorthand for the pushforward \(or image\) measure induced by the measurable mapβ\\beta\. Pushforward measures provide the canonical mechanism by which probability distributions are transported through measurable mappings\([21](https://arxiv.org/html/2608.11465#bib.bib29);[3](https://arxiv.org/html/2608.11465#bib.bib15)\)\.
### A\.3Disintegration of Measures
A central mathematical tool used throughout the paper is disintegration\.
Disintegration is the measure\-theoretic analog of conditioning and provides a canonical decomposition of a probability measure relative to a measurable map\([33](https://arxiv.org/html/2608.11465#bib.bib28);[21](https://arxiv.org/html/2608.11465#bib.bib29);[3](https://arxiv.org/html/2608.11465#bib.bib15);[23](https://arxiv.org/html/2608.11465#bib.bib25)\)\.
Assume that\(Θ,ℱΘ\)\(\\Theta,\\mathcal\{F\}\_\{\\Theta\}\)is a standard Borel space and thatβ:Θ→𝒦\\beta:\\Theta\\rightarrow\\mathcal\{K\}is measurable\. These are the same assumptions introduced in Section[2](https://arxiv.org/html/2608.11465#S2); they ensure the existence of regular conditional distributions with respect to the behavior map\.
Then standard disintegration theorems imply that every probability measureQ∈𝒫\(Θ\)Q\\in\\mathcal\{P\}\(\\Theta\)admits a regular conditional distribution with respect to the behavior mapβ\\beta, supported on the fibers ofβ\\beta, and therefore a decomposition of the form
Q\(dθ\)=Q\(dθ∣k\)πQ\(dk\)\.Q\(d\\theta\)=Q\(d\\theta\\mid k\)\\,\\pi\_\{Q\}\(dk\)\.HereQ\(⋅∣k\)Q\(\\cdot\\mid k\)is a conditional probability measure supported on the fiberFk=β−1\(k\)F\_\{k\}=\\beta^\{\-1\}\(k\)\.
Informally, disintegration separates uncertainty into two components:
1. 1\.uncertainty over predictive behaviors, represented byπQ\\pi\_\{Q\};
2. 2\.uncertainty over realizations conditional on a fixed predictive behavior, represented byQ\(⋅∣k\)Q\(\\cdot\\mid k\)\.
This disintegration is the measure\-theoretic foundation of the behavior–realization decomposition developed in the main text\. Appendix[B](https://arxiv.org/html/2608.11465#A2)applies the classical chain rule for relative entropy to this decomposition, with the measurable map taken to be the behavior mapβ\\beta, yielding the behavior–realization decomposition of classical PAC\-Bayes complexity\.
### A\.4Proofs of Behavioral Sufficiency Results
###### Proof of Lemma[1](https://arxiv.org/html/2608.11465#Thmlemma1)\.
Behavioral equivalence impliesβ\(θ\)=β\(θ′\)\\beta\(\\theta\)=\\beta\(\\theta^\{\\prime\}\), and thereforePθ\(⋅∣x\)=Pθ′\(⋅∣x\)P\_\{\\theta\}\(\\cdot\\mid x\)=P\_\{\\theta^\{\\prime\}\}\(\\cdot\\mid x\)for every inputxx\.
Since both population and empirical risks depend on a configuration only through the predictive behavior represented byβ\(θ\)\\beta\(\\theta\), it follows immediately that
L\(θ\)=L\(θ′\),L^S\(θ\)=L^S\(θ′\)\.L\(\\theta\)=L\(\\theta^\{\\prime\}\),\\qquad\\hat\{L\}\_\{S\}\(\\theta\)=\\hat\{L\}\_\{S\}\(\\theta^\{\\prime\}\)\.∎
###### Proof of Proposition[1](https://arxiv.org/html/2608.11465#Thmproposition1)\.
By Lemma[1](https://arxiv.org/html/2608.11465#Thmlemma1), both population and empirical risks are constant on fibers\. Sinceβ\\betais measurable and both risks are measurable functions onΘ\\Theta, standard measurable\-factorization results for functions that are constant on the fibers of a measurable map\([21](https://arxiv.org/html/2608.11465#bib.bib29)\)imply the existence of measurable functionsL𝒦L\_\{\\mathcal\{K\}\}andL^S,𝒦\\hat\{L\}\_\{S,\\mathcal\{K\}\}on𝒦\\mathcal\{K\}such that
L\(θ\)=L𝒦\(β\(θ\)\),L^S\(θ\)=L^S,𝒦\(β\(θ\)\)\.L\(\\theta\)=L\_\{\\mathcal\{K\}\}\(\\beta\(\\theta\)\),\\qquad\\hat\{L\}\_\{S\}\(\\theta\)=\\hat\{L\}\_\{S,\\mathcal\{K\}\}\(\\beta\(\\theta\)\)\.Using the factorization above together with the pushforward measure,
L\(Q\)=𝔼θ∼Q\[L𝒦\(β\(θ\)\)\]=𝔼k∼πQ\[L𝒦\(k\)\]\.L\(Q\)=\\mathbb\{E\}\_\{\\theta\\sim Q\}\\\!\\left\[L\_\{\\mathcal\{K\}\}\(\\beta\(\\theta\)\)\\right\]=\\mathbb\{E\}\_\{k\\sim\\pi\_\{Q\}\}\\\!\\left\[L\_\{\\mathcal\{K\}\}\(k\)\\right\]\.IfπQ=πQ′\\pi\_\{Q\}=\\pi\_\{Q^\{\\prime\}\}, the expectations coincide, yielding
L\(Q\)=L\(Q′\)\.L\(Q\)=L\(Q^\{\\prime\}\)\.Applying the same argument toL^S,𝒦\\hat\{L\}\_\{S,\\mathcal\{K\}\}gives
L^S\(Q\)=L^S\(Q′\)\.\\hat\{L\}\_\{S\}\(Q\)=\\hat\{L\}\_\{S\}\(Q^\{\\prime\}\)\.∎
## Appendix BKL Decomposition and PAC\-Bayes Z\-Information
This appendix provides proofs of the results in Section[3](https://arxiv.org/html/2608.11465#S3)\. The proofs rely on the classical chain rule for relative entropy under measure disintegration, which decomposes relative entropy into marginal and conditional contributions relative to a measurable map\([21](https://arxiv.org/html/2608.11465#bib.bib29);[33](https://arxiv.org/html/2608.11465#bib.bib28);[3](https://arxiv.org/html/2608.11465#bib.bib15)\)\. The novelty of the present work lies not in this measure\-theoretic identity itself, but in its application to the behavior mapβ:Θ→𝒦\\beta:\\Theta\\rightarrow\\mathcal\{K\}introduced in Section[2](https://arxiv.org/html/2608.11465#S2), which yields the behavior–realization decomposition, its variational characterization, and the resulting behavior\-level formulation of PAC\-Bayes complexity\.
Throughout, letP,Q∈𝒫\(Θ\)P,Q\\in\\mathcal\{P\}\(\\Theta\)denote the prior and posterior distributions, and letπP=β\#P\\pi\_\{P\}=\\beta\_\{\\\#\}PandπQ=β\#Q\\pi\_\{Q\}=\\beta\_\{\\\#\}Qbe the corresponding induced distributions on the behavior space\.
Unless otherwise stated, we assume
KL\(Q∥P\)<∞,\\mathrm\{KL\}\(Q\\\|P\)<\\infty,which is the regime relevant to classical PAC\-Bayes analysis\. Under this assumption,Q≪PQ\\ll P, so the chain rule for relative entropy under measure disintegration applies directly\. The decomposition extends in the usual extended\-real sense when
KL\(Q∥P\)=∞\.\\mathrm\{KL\}\(Q\\\|P\)=\\infty\.
Under the standard Borel assumptions of Section[2](https://arxiv.org/html/2608.11465#S2), both measures admit disintegrations
P\(dθ\)=P\(dθ∣k\)πP\(dk\),Q\(dθ\)=Q\(dθ∣k\)πQ\(dk\)\.P\(d\\theta\)=P\(d\\theta\\mid k\)\\,\\pi\_\{P\}\(dk\),\\qquad Q\(d\\theta\)=Q\(d\\theta\\mid k\)\\,\\pi\_\{Q\}\(dk\)\.
### B\.1Proof of Proposition[2](https://arxiv.org/html/2608.11465#Thmproposition2)
###### Proof\.
Assume
KL\(Q∥P\)<∞\.\\mathrm\{KL\}\(Q\\\|P\)<\\infty\.ThenQ≪PQ\\ll P, and therefore
πQ=β\#Q≪β\#P=πP\.\\pi\_\{Q\}=\\beta\_\{\\\#\}Q\\ll\\beta\_\{\\\#\}P=\\pi\_\{P\}\.
Applying the chain rule for relative entropy with respect to the measurable mapβ\\beta\([21](https://arxiv.org/html/2608.11465#bib.bib29);[3](https://arxiv.org/html/2608.11465#bib.bib15)\)yields:
KL\(Q∥P\)=KL\(πQ∥πP\)\+𝔼k∼πQ\[KL\(Q\(⋅∣k\)∥P\(⋅∣k\)\)\]\.\\mathrm\{KL\}\(Q\\\|P\)=\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\)\+\\mathbb\{E\}\_\{k\\sim\\pi\_\{Q\}\}\\\!\\left\[\\mathrm\{KL\}\\\!\\left\(Q\(\\cdot\\mid k\)\\,\\middle\\\|\\,P\(\\cdot\\mid k\)\\right\)\\right\]\.∎
The nonnegativity of the conditional term immediately implies Corollary[1](https://arxiv.org/html/2608.11465#Thmcorollary1), namely
KL\(πQ∥πP\)≤KL\(Q∥P\)\.\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\)\\leq\\mathrm\{KL\}\(Q\\\|P\)\.
### B\.2Nonpositivity of PAC\-Bayes Z\-Information
The statement following Definition[2](https://arxiv.org/html/2608.11465#Thmdefinition2)follows directly from the non\-negativity of KL divergence\.
###### Proof\.
For every behaviorkk,
KL\(Q\(⋅∣k\)∥P\(⋅∣k\)\)≥0\.\\mathrm\{KL\}\\\!\\left\(Q\(\\cdot\\mid k\)\\,\\middle\\\|\\,P\(\\cdot\\mid k\)\\right\)\\geq 0\.Taking expectations preserves nonnegativity, so
𝒵PB\(Q∥P\)=−𝔼k∼πQ\[KL\(Q\(⋅∣k\)∥P\(⋅∣k\)\)\]≤0\.\\mathcal\{Z\}\_\{\\mathrm\{PB\}\}\(Q\\\|P\)=\-\\mathbb\{E\}\_\{k\\sim\\pi\_\{Q\}\}\\\!\\left\[\\mathrm\{KL\}\\\!\\left\(Q\(\\cdot\\mid k\)\\,\\middle\\\|\\,P\(\\cdot\\mid k\)\\right\)\\right\]\\leq 0\.Equality holds if and only if the conditional KL divergence vanishes forπQ\\pi\_\{Q\}\-almost everykk\. By the standard characterization of equality in relative entropy, this is equivalent to
Q\(⋅∣k\)=P\(⋅∣k\)Q\(\\cdot\\mid k\)=P\(\\cdot\\mid k\)forπQ\\pi\_\{Q\}\-almost everykk\. ∎
### B\.3Proof of Proposition[3](https://arxiv.org/html/2608.11465#Thmproposition3)
###### Proof\.
Rearranging the decomposition of Proposition[2](https://arxiv.org/html/2608.11465#Thmproposition2)and substituting Definition[2](https://arxiv.org/html/2608.11465#Thmdefinition2)yields
𝒵PB\(Q∥P\)=KL\(πQ∥πP\)−KL\(Q∥P\)\.\\mathcal\{Z\}\_\{\\mathrm\{PB\}\}\(Q\\\|P\)=\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\)\-\\mathrm\{KL\}\(Q\\\|P\)\.Or equivalently,
KL\(πQ∥πP\)=KL\(Q∥P\)\+𝒵PB\(Q∥P\),\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\)=\\mathrm\{KL\}\(Q\\\|P\)\+\\mathcal\{Z\}\_\{\\mathrm\{PB\}\}\(Q\\\|P\),which is the claimed identity\. ∎
### B\.4Proof of Theorem[1](https://arxiv.org/html/2608.11465#Thmtheorem1)
###### Proof\.
Recall that
Q⋆\(𝑑θ\)=∫𝒦P\(𝑑θ∣k\)πQ\(𝑑k\)\.Q^\{\\star\}\(d\\theta\)=\\int\_\{\\mathcal\{K\}\}P\(d\\theta\\mid k\)\\,\\pi\_\{Q\}\(dk\)\.BecauseP\(⋅∣k\)P\(\\cdot\\mid k\)is the measurable probability kernel obtained by disintegratingPPwith respect toβ\\beta, the measureQ⋆Q^\{\\star\}is a well\-defined probability measure onΘ\\Theta\. The existence of this measurable conditional kernel follows from the standard Borel assumptions of Section[2](https://arxiv.org/html/2608.11465#S2)and the disintegration theorem reviewed in Appendix[A](https://arxiv.org/html/2608.11465#A1)\.
For any measurable setB⊆𝒦B\\subseteq\\mathcal\{K\},
πQ⋆\(B\)=Q⋆\(β−1\(B\)\)=∫𝒦P\(β−1\(B\)∣k\)πQ\(𝑑k\)\.\\pi\_\{Q^\{\\star\}\}\(B\)=Q^\{\\star\}\(\\beta^\{\-1\}\(B\)\)=\\int\_\{\\mathcal\{K\}\}P\(\\beta^\{\-1\}\(B\)\\mid k\)\\,\\pi\_\{Q\}\(dk\)\.BecauseP\(⋅∣k\)P\(\\cdot\\mid k\)is supported on the fiberβ−1\(k\)\\beta^\{\-1\}\(k\),
P\(β−1\(B\)∣k\)=𝟏B\(k\)\.P\(\\beta^\{\-1\}\(B\)\\mid k\)=\\mathbf\{1\}\_\{B\}\(k\)\.Therefore,
πQ⋆\(B\)=∫𝒦𝟏B\(k\)πQ\(𝑑k\)=πQ\(B\),\\pi\_\{Q^\{\\star\}\}\(B\)=\\int\_\{\\mathcal\{K\}\}\\mathbf\{1\}\_\{B\}\(k\)\\,\\pi\_\{Q\}\(dk\)=\\pi\_\{Q\}\(B\),showing thatπQ⋆=πQ\\pi\_\{Q^\{\\star\}\}=\\pi\_\{Q\}\.
Applying Proposition[2](https://arxiv.org/html/2608.11465#Thmproposition2)toQ⋆Q^\{\\star\}yields
KL\(Q⋆∥P\)=KL\(πQ∥πP\),\\mathrm\{KL\}\(Q^\{\\star\}\\\|P\)=\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\),because
Q⋆\(dθ\)=P\(dθ∣k\)πQ\(dk\),Q^\{\\star\}\(d\\theta\)=P\(d\\theta\\mid k\)\\,\\pi\_\{Q\}\(dk\),which is already a disintegration ofQ⋆Q^\{\\star\}with respect to the behavior marginalπQ=πQ⋆\\pi\_\{Q\}=\\pi\_\{Q^\{\\star\}\}\. Hence, a valid conditional distribution ofQ⋆Q^\{\\star\}is
Q⋆\(⋅∣k\)=P\(⋅∣k\)Q^\{\\star\}\(\\cdot\\mid k\)=P\(\\cdot\\mid k\)forπQ\\pi\_\{Q\}\-almost everykk\. Consequently,
KL\(Q⋆\(⋅∣k\)∥P\(⋅∣k\)\)=0\\mathrm\{KL\}\\\!\\left\(Q^\{\\star\}\(\\cdot\\mid k\)\\,\\middle\\\|\\,P\(\\cdot\\mid k\)\\right\)=0forπQ\\pi\_\{Q\}\-almost everykk\.
Now letQ′Q^\{\\prime\}satisfyπQ′=πQ\\pi\_\{Q^\{\\prime\}\}=\\pi\_\{Q\}\. Applying Proposition[2](https://arxiv.org/html/2608.11465#Thmproposition2)again gives
KL\(Q′∥P\)=KL\(πQ∥πP\)\+𝔼k∼πQ\[KL\(Q′\(⋅∣k\)∥P\(⋅∣k\)\)\]\.\\mathrm\{KL\}\(Q^\{\\prime\}\\\|P\)=\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\)\+\\mathbb\{E\}\_\{k\\sim\\pi\_\{Q\}\}\\\!\\left\[\\mathrm\{KL\}\\\!\\left\(Q^\{\\prime\}\(\\cdot\\mid k\)\\,\\middle\\\|\\,P\(\\cdot\\mid k\)\\right\)\\right\]\.The second term is nonnegative, so
KL\(Q′∥P\)≥KL\(πQ∥πP\)\.\\mathrm\{KL\}\(Q^\{\\prime\}\\\|P\)\\geq\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\)\.Since equality is achieved byQ⋆Q^\{\\star\},
infQ′:πQ′=πQKL\(Q′∥P\)=KL\(πQ∥πP\)\.\\inf\_\{Q^\{\\prime\}:\\,\\pi\_\{Q^\{\\prime\}\}=\\pi\_\{Q\}\}\\mathrm\{KL\}\(Q^\{\\prime\}\\\|P\)=\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\)\.∎
### B\.5Behavioral Sufficiency and Risk Preservation
The variational characterization has learning\-theoretic significance because predictive risk depends only on behavior\.
By Proposition[1](https://arxiv.org/html/2608.11465#Thmproposition1), if two posteriors induce the same behavior distribution, then they have identical population and empirical risks\. SinceπQ⋆=πQ\\pi\_\{Q^\{\\star\}\}=\\pi\_\{Q\}, Proposition[1](https://arxiv.org/html/2608.11465#Thmproposition1)immediately yields
L\(Q⋆\)=L\(Q\),L^S\(Q⋆\)=L^S\(Q\)\.L\(Q^\{\\star\}\)=L\(Q\),\\qquad\\hat\{L\}\_\{S\}\(Q^\{\\star\}\)=\\hat\{L\}\_\{S\}\(Q\)\.Consequently, replacingQQwithQ⋆Q^\{\\star\}preserves both empirical and population performance while achieving the minimum classical PAC\-Bayes complexity among all posteriors inducing the same distribution over predictive behaviors\. This observation underlies the behavior\-level PAC\-Bayes interpretation developed in Section[4](https://arxiv.org/html/2608.11465#S4)\.
### B\.6Interpretation
Proposition[2](https://arxiv.org/html/2608.11465#Thmproposition2)specializes the classical chain rule for relative entropy to the behavior map\. This specialization separates classical PAC\-Bayes complexity into a behavior\-selection term,
𝔼k∼πQ\[KL\(Q\(⋅∣k\)∥P\(⋅∣k\)\)\]\.\\mathbb\{E\}\_\{k\\sim\\pi\_\{Q\}\}\\\!\\left\[\\mathrm\{KL\}\\\!\\left\(Q\(\\cdot\\mid k\)\\,\\middle\\\|\\,P\(\\cdot\\mid k\)\\right\)\\right\]\.This term measures the divergence between the posterior and prior conditional distributions within behavioral fibers and therefore quantifies realization\-level variation after predictive behavior has been fixed\.
PAC\-Bayes Z\-information is the negative of the realization\-level term\. Consequently, it quantifies the portion of classical PAC\-Bayes complexity attributable to variation among behaviorally equivalent realizations\.
Theorem[1](https://arxiv.org/html/2608.11465#Thmtheorem1)provides a complementary variational interpretation\. The quantity
KL\(πQ∥πP\)=infQ′:πQ′=πQKL\(Q′∥P\)\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\)=\\inf\_\{Q^\{\\prime\}:\\,\\pi\_\{Q^\{\\prime\}\}=\\pi\_\{Q\}\}\\mathrm\{KL\}\(Q^\{\\prime\}\\\|P\)is the minimum classical PAC\-Bayes complexity among all configuration\-space posteriors that induce the same behavior\-level distribution\. Equivalently, it is the irreducible complexity associated with the predictive behaviors encoded byπQ\\pi\_\{Q\}after all realization\-level variation has been removed\.
The quantity−𝒵PB\(Q∥P\)\-\\mathcal\{Z\}\_\{\\mathrm\{PB\}\}\(Q\\\|P\)is precisely the optimality gap between the classical PAC\-Bayes complexity ofQQand the minimum complexity achievable among all posteriors inducing the same behavioral distribution\. It therefore quantifies exactly the excess classical PAC\-Bayes complexity attributable to realization\-level variation within behavioral fibers\.
## Appendix CBehavior\-Level PAC\-Bayes Analysis
This appendix provides proofs and supporting details for the behavior\-level PAC\-Bayes analysis in Section[4](https://arxiv.org/html/2608.11465#S4)\. The behavior\-level bound is obtained by applying the classical PAC\-Bayes theorem directly to the measurable behavior space𝒦\\mathcal\{K\}\. The contribution is therefore not a new concentration inequality but the identification of the corresponding complexity term with the behavior\-selection component of the exact behavior–realization decomposition\.
### C\.1Risk Depends Only on Behavior
We first make explicit the risk identities used in the proof of the Behavior\-space formulation of the PAC\-Bayes bound\.
LetπQ=β\#Q\\pi\_\{Q\}=\\beta\_\{\\\#\}Qdenote the pushforward \(image\) measure induced by the behavior mapβ\\beta\.
By Proposition[1](https://arxiv.org/html/2608.11465#Thmproposition1), both population and empirical risk depend on a posterior only through its induced behavior distribution\. Consequently,
L\(Q\)=L\(πQ\),L^S\(Q\)=L^S\(πQ\)\.L\(Q\)=L\(\\pi\_\{Q\}\),\\qquad\\hat\{L\}\_\{S\}\(Q\)=\\hat\{L\}\_\{S\}\(\\pi\_\{Q\}\)\.Equivalently, if two configuration\-space posteriors satisfy
πQ=πQ′,\\pi\_\{Q\}=\\pi\_\{Q^\{\\prime\}\},then
L\(Q\)=L\(Q′\),L^S\(Q\)=L^S\(Q′\)\.L\(Q\)=L\(Q^\{\\prime\}\),\\qquad\\hat\{L\}\_\{S\}\(Q\)=\\hat\{L\}\_\{S\}\(Q^\{\\prime\}\)\.A detailed proof of this factorization is given in Appendix[A](https://arxiv.org/html/2608.11465#A1)\.
These identities allow the behavior\-space PAC\-Bayes bound to be expressed entirely in terms of the configuration\-space posteriorQQ\.
### C\.2Proof of Theorem[3](https://arxiv.org/html/2608.11465#Thmtheorem3)
###### Proof\.
By assumption,𝒦\\mathcal\{K\}is a standard Borel space\. The priorPPonΘ\\Thetainduces a prior
πP=β\#P\\pi\_\{P\}=\\beta\_\{\\\#\}Pon𝒦\\mathcal\{K\}\.
Because𝒦\\mathcal\{K\}is a standard Borel space andπP\\pi\_\{P\}is a probability measure on𝒦\\mathcal\{K\}, the classical PAC\-Bayes theorem applies directly on the behavior space𝒦\\mathcal\{K\}\([30](https://arxiv.org/html/2608.11465#bib.bib33);[36](https://arxiv.org/html/2608.11465#bib.bib39)\)\. Thus, with probability at least1−δ1\-\\deltaover the draw ofS∼𝒟nS\\sim\\mathcal\{D\}^\{n\}, simultaneously for all behavior\-space posteriorsρ∈𝒫\(𝒦\)\\rho\\in\\mathcal\{P\}\(\\mathcal\{K\}\),
KL\(L^S\(ρ\)∥L\(ρ\)\)≤1n\(KL\(ρ∥πP\)\+log2nδ\)\.\\mathrm\{KL\}\\\!\\left\(\\hat\{L\}\_\{S\}\(\\rho\)\\,\\middle\\\|\\,L\(\\rho\)\\right\)\\leq\\frac\{1\}\{n\}\\left\(\\mathrm\{KL\}\(\\rho\\\|\\pi\_\{P\}\)\+\\log\\frac\{2\\sqrt\{n\}\}\{\\delta\}\\right\)\.Now let
whereQQis an arbitrary configuration\-space posterior\.
Substituting these identities into the PAC\-Bayes bound yields
KL\(L^S\(Q\)∥L\(Q\)\)≤1n\(KL\(πQ∥πP\)\+log2nδ\)\.\\mathrm\{KL\}\\\!\\left\(\\hat\{L\}\_\{S\}\(Q\)\\,\\middle\\\|\\,L\(Q\)\\right\)\\leq\\frac\{1\}\{n\}\\left\(\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\)\+\\log\\frac\{2\\sqrt\{n\}\}\{\\delta\}\\right\)\.SinceQQwas arbitrary, the inequality holds simultaneously for all configuration\-space posteriorsQQ, establishing the theorem\.
∎
### C\.3Proof of Proposition[4](https://arxiv.org/html/2608.11465#Thmproposition4)
###### Proof\.
By Proposition[3](https://arxiv.org/html/2608.11465#Thmproposition3),
𝒵PB\(Q∥P\)=KL\(πQ∥πP\)−KL\(Q∥P\)\.\\mathcal\{Z\}\_\{\\mathrm\{PB\}\}\(Q\\\|P\)=\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\)\-\\mathrm\{KL\}\(Q\\\|P\)\.Rearranging gives
KL\(Q∥P\)−KL\(πQ∥πP\)=−𝒵PB\(Q∥P\)\.\\mathrm\{KL\}\(Q\\\|P\)\-\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\)=\-\\mathcal\{Z\}\_\{\\mathrm\{PB\}\}\(Q\\\|P\)\.Since
𝒵PB\(Q∥P\)≤0,\\mathcal\{Z\}\_\{\\mathrm\{PB\}\}\(Q\\\|P\)\\leq 0,we have
−𝒵PB\(Q∥P\)≥0\.\-\\mathcal\{Z\}\_\{\\mathrm\{PB\}\}\(Q\\\|P\)\\geq 0\.Therefore
KL\(πQ∥πP\)≤KL\(Q∥P\)\.\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\)\\leq\\mathrm\{KL\}\(Q\\\|P\)\.∎
### C\.4Proof of Corollary[2](https://arxiv.org/html/2608.11465#Thmcorollary2)
###### Proof\.
Theorem[3](https://arxiv.org/html/2608.11465#Thmtheorem3)gives
KL\(L^S\(Q\)∥L\(Q\)\)≤1n\(KL\(πQ∥πP\)\+log2nδ\)\.\\mathrm\{KL\}\\\!\\left\(\\hat\{L\}\_\{S\}\(Q\)\\,\\middle\\\|\\,L\(Q\)\\right\)\\leq\\frac\{1\}\{n\}\\left\(\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\)\+\\log\\frac\{2\\sqrt\{n\}\}\{\\delta\}\\right\)\.Using Proposition[3](https://arxiv.org/html/2608.11465#Thmproposition3),
KL\(πQ∥πP\)=KL\(Q∥P\)\+𝒵PB\(Q∥P\)\.\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\)=\\mathrm\{KL\}\(Q\\\|P\)\+\\mathcal\{Z\}\_\{\\mathrm\{PB\}\}\(Q\\\|P\)\.Substituting this identity into the behavior\-space PAC\-Bayes bound gives
KL\(L^S\(Q\)∥L\(Q\)\)≤1n\(KL\(Q∥P\)\+𝒵PB\(Q∥P\)\+log2nδ\)\.\\mathrm\{KL\}\\\!\\left\(\\hat\{L\}\_\{S\}\(Q\)\\,\\middle\\\|\\,L\(Q\)\\right\)\\leq\\frac\{1\}\{n\}\\left\(\\mathrm\{KL\}\(Q\\\|P\)\+\\mathcal\{Z\}\_\{\\mathrm\{PB\}\}\(Q\\\|P\)\+\\log\\frac\{2\\sqrt\{n\}\}\{\\delta\}\\right\)\.∎
### C\.5Interpretation
The behavior\-space PAC\-Bayes bound is not a new concentration inequality\. Rather, it is the classical PAC\-Bayes theorem applied to the measurable space of predictive behaviors\. Its significance lies in identifying the resulting complexity term with the behavior\-selection component of the exact behavior–realization decomposition\.
The resulting complexity term,
KL\(πQ∥πP\),\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\),depends only on uncertainty over predictive behavior\. By Theorem[1](https://arxiv.org/html/2608.11465#Thmtheorem1),
KL\(πQ∥πP\)=infQ′:πQ′=πQKL\(Q′∥P\),\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\)=\\inf\_\{Q^\{\\prime\}:\\,\\pi\_\{Q^\{\\prime\}\}=\\pi\_\{Q\}\}\\mathrm\{KL\}\(Q^\{\\prime\}\\\|P\),so it is the minimum classical PAC\-Bayes complexity among all configuration\-space posteriors that induce the same distribution over predictive behaviors\. By Theorem[1](https://arxiv.org/html/2608.11465#Thmtheorem1), this minimum is attained by the fiber\-symmetrized posteriorQ⋆Q^\{\\star\}, which preserves predictive behavior while eliminating all realization\-level divergence\.
PAC\-Bayes Z\-information quantifies the exact gap between configuration\-space complexity and behavior\-selection complexity\. Since
KL\(πQ∥πP\)=KL\(Q∥P\)\+𝒵PB\(Q∥P\),\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\)=\\mathrm\{KL\}\(Q\\\|P\)\+\\mathcal\{Z\}\_\{\\mathrm\{PB\}\}\(Q\\\|P\),and
𝒵PB\(Q∥P\)≤0,\\mathcal\{Z\}\_\{\\mathrm\{PB\}\}\(Q\\\|P\)\\leq 0,Consequently, the behavior\-space formulation isolates the irreducible complexity associated with predictive behavior while removing complexity arising solely from variation among behaviorally equivalent realizations, without changing either empirical or population risk\.
## Appendix DContinuous\-Space Interpretation of the Realization\-Level KL Term
This appendix develops a continuous\-space interpretation of the realization\-level KL divergence under additional assumptions on the conditional prior and a chosen reference measure on each behavioral fiber\.
The purpose of this appendix is to develop the geometric intuition for realization multiplicity discussed in Section[5](https://arxiv.org/html/2608.11465#S5)\. The results in this appendix are not used in any proof of the main PAC\-Bayes results\. Under additional assumptions on the conditional prior and a chosen reference measure on each behavioral fiber, the realization\-level KL divergence—and hence PAC\-Bayes Z\-information—admits an occupancy\-based interpretation relative to realization multiplicity\.
The results of this appendix are supplementary and are not used in any proof of the main PAC\-Bayes results\. The theory developed in Sections[2](https://arxiv.org/html/2608.11465#S2)–[4](https://arxiv.org/html/2608.11465#S4)requires only the measure\-theoretic KL decomposition and remains valid independently of the continuous\-space interpretation presented here\.
### D\.1Conditional KL Divergence on a Fiber
Letk∈𝒦k\\in\\mathcal\{K\}and letFk=β−1\(k\)F\_\{k\}=\\beta^\{\-1\}\(k\)denote the corresponding behavioral fiber\.
Assume thatFkF\_\{k\}is equipped with a reference measureμk\\mu\_\{k\}, and that the conditional prior and posterior admit densitiespk\(θ\)p\_\{k\}\(\\theta\)andqk\(θ\)q\_\{k\}\(\\theta\)with respect toμk\\mu\_\{k\}\.
The choice ofμk\\mu\_\{k\}is not canonical and may depend on the application\. All densities and volume quantities in this appendix are understood relative to the same fixed reference measure on the fiber\.
The fiber\-wise KL divergence is
KL\(Q\(⋅∣k\)∥P\(⋅∣k\)\)=∫Fkqk\(θ\)logqk\(θ\)pk\(θ\)dμk\(θ\)\.\\mathrm\{KL\}\\big\(Q\(\\cdot\\mid k\)\\,\\\|\\,P\(\\cdot\\mid k\)\\big\)=\\int\_\{F\_\{k\}\}q\_\{k\}\(\\theta\)\\log\\frac\{q\_\{k\}\(\\theta\)\}\{p\_\{k\}\(\\theta\)\}\\,d\\mu\_\{k\}\(\\theta\)\.This quantity measures posterior concentration relative to the conditional prior within the behavioral fiber\.
As emphasized in Section[5](https://arxiv.org/html/2608.11465#S5), the KL divergence itself does not require any notion of geometric volume\. The reference measure is introduced only to provide intuition about realization multiplicity in continuous configuration spaces\.
### D\.2A Special Case: Uniform Conditional Priors
Suppose that the conditional prior is uniform on the fiber,pk\(θ\)=1/Vol\(Fk\)p\_\{k\}\(\\theta\)=1/\\operatorname\{Vol\}\(F\_\{k\}\), whereVol\(Fk\)=μk\(Fk\)\\operatorname\{Vol\}\(F\_\{k\}\)=\\mu\_\{k\}\(F\_\{k\}\)\.
For simplicity, assume throughout this subsection that0<Vol\(Fk\)<∞0<\\operatorname\{Vol\}\(F\_\{k\}\)<\\infty\.
Substituting into the definition of KL divergence yields
KL\(Q\(⋅∣k\)∥P\(⋅∣k\)\)\\displaystyle\\mathrm\{KL\}\\big\(Q\(\\cdot\\mid k\)\\,\\\|\\,P\(\\cdot\\mid k\)\\big\)=∫Fkqk\(θ\)logqk\(θ\)1/Vol\(Fk\)dμk\(θ\)\\displaystyle=\\int\_\{F\_\{k\}\}q\_\{k\}\(\\theta\)\\log\\frac\{q\_\{k\}\(\\theta\)\}\{1/\\operatorname\{Vol\}\(F\_\{k\}\)\}\\,d\\mu\_\{k\}\(\\theta\)=∫Fkqk\(θ\)logqk\(θ\)dμk\(θ\)\+logVol\(Fk\)\.\\displaystyle=\\int\_\{F\_\{k\}\}q\_\{k\}\(\\theta\)\\log q\_\{k\}\(\\theta\)\\,d\\mu\_\{k\}\(\\theta\)\+\\log\\operatorname\{Vol\}\(F\_\{k\}\)\.Define the differential entropy of the posterior conditional distribution on the fiber \(relative toμk\\mu\_\{k\}\) by
Hk\(Q\)=−∫Fkqk\(θ\)logqk\(θ\)dμk\(θ\)\.H\_\{k\}\(Q\)=\-\\int\_\{F\_\{k\}\}q\_\{k\}\(\\theta\)\\log q\_\{k\}\(\\theta\)\\,d\\mu\_\{k\}\(\\theta\)\.As with differential entropy generally\([8](https://arxiv.org/html/2608.11465#bib.bib20)\), the numerical value ofHk\(Q\)H\_\{k\}\(Q\)depends on the choice of reference measure\. The KL divergence itself remains invariant because both densities are defined relative to the same reference measure\.
###### Proposition 8\(Entropy\-Volume Identity\)\.
If the conditional prior is uniform on a fiberFkF\_\{k\}, then
KL\(Q\(⋅∣k\)∥P\(⋅∣k\)\)=logVol\(Fk\)−Hk\(Q\)\.\\mathrm\{KL\}\\big\(Q\(\\cdot\\mid k\)\\,\\\|\\,P\(\\cdot\\mid k\)\\big\)=\\log\\operatorname\{Vol\}\(F\_\{k\}\)\-H\_\{k\}\(Q\)\.
###### Proof\.
Substitutingpk\(θ\)=1/Vol\(Fk\)p\_\{k\}\(\\theta\)=1/\\operatorname\{Vol\}\(F\_\{k\}\)into the definition of KL divergence gives
KL\(Q\(⋅∣k\)∥P\(⋅∣k\)\)=∫Fkqk\(θ\)logqk\(θ\)dμk\(θ\)\+logVol\(Fk\)\.\\mathrm\{KL\}\\big\(Q\(\\cdot\\mid k\)\\,\\\|\\,P\(\\cdot\\mid k\)\\big\)=\\int\_\{F\_\{k\}\}q\_\{k\}\(\\theta\)\\log q\_\{k\}\(\\theta\)\\,d\\mu\_\{k\}\(\\theta\)\+\\log\\operatorname\{Vol\}\(F\_\{k\}\)\.Using the definition ofHk\(Q\)H\_\{k\}\(Q\)yields the result\. ∎
The quantityVol\(Fk\)\\operatorname\{Vol\}\(F\_\{k\}\)is defined relative to the chosen reference measureμk\\mu\_\{k\}and, therefore, is not an intrinsic geometric invariant of the fiber\. The interpretation developed below concerns realization multiplicity relative to the specified reference measure\.
Accordingly, Proposition[8](https://arxiv.org/html/2608.11465#Thmproposition8)should be viewed as a special\-case interpretation rather than a general characterization of the realization\-level KL divergence\. For arbitrary conditional priors, the fiber\-wise KL divergence contains additional terms reflecting the variation of the conditional prior within the fiber\. The uniform\-prior case isolates the contribution of posterior occupancy relative to the available reference\-measure size of the fiber\.
Although exact uniformity is rarely satisfied in practice, the resulting interpretation remains informative whenever the conditional prior is approximately uniform across a fiber\. In such cases, the fiber\-wise KL divergence differs from the expression in Proposition[8](https://arxiv.org/html/2608.11465#Thmproposition8)by an additional term that quantifies departures from uniform occupancy\.
### D\.3Interpretation
Proposition[8](https://arxiv.org/html/2608.11465#Thmproposition8)separates the fiber\-wise KL divergence into two contributions:
1. 1\.the logarithm of the reference\-measure size of the fiber,logVol\(Fk\)\\log\\operatorname\{Vol\}\(F\_\{k\}\);
2. 2\.the entropy of posterior occupancy within that fiber,Hk\(Q\)H\_\{k\}\(Q\)\.
The KL divergence is small when the posterior mass is distributed broadly throughout the fiber and large when the posterior concentrates on a relatively small subset of realizations\.
Thus, under the uniform conditional\-prior assumption, the realization\-level KL divergence can be interpreted as measuring posterior occupancy relative to the available reference\-measure size of a behavioral fiber\.
Moreover, under the same assumption,
𝒵PB\(Q∥P\)=𝔼k∼πQ\[Hk\(Q\)−logVol\(Fk\)\]\.\\mathcal\{Z\}\_\{\\mathrm\{PB\}\}\(Q\\\|P\)=\\mathbb\{E\}\_\{k\\sim\\pi\_\{Q\}\}\\Big\[H\_\{k\}\(Q\)\-\\log\\operatorname\{Vol\}\(F\_\{k\}\)\\Big\]\.
Accordingly, PAC\-Bayes Z\-information may be interpreted as the expected difference between posterior occupancy entropy and the logarithm of the available reference\-measure size of a behavioral fiber\. This interpretation is specific to the uniform\-prior setting and serves only as geometric intuition for the general measure\-theoretic definition\.
### D\.4Relation to Geometry and Scope
The calculations in this appendix are standard consequences of relative entropy and measure\-theoretic probability\([8](https://arxiv.org/html/2608.11465#bib.bib20);[9](https://arxiv.org/html/2608.11465#bib.bib21);[1](https://arxiv.org/html/2608.11465#bib.bib13);[21](https://arxiv.org/html/2608.11465#bib.bib29)\)\.
Section[5](https://arxiv.org/html/2608.11465#S5)interprets realization multiplicity in terms of symmetry, reference\-measure size, and behavior\-preserving directions\. The occupancy interpretation developed here provides an information\-theoretic counterpart to that geometric picture\.
Relative to the chosen reference measure, large fibers correspond to predictive behaviors that admit many realizations\. In continuous settings, this realization multiplicity may arise through large reference\-measure size, high\-dimensional families of behavior\-preserving directions, or other forms of geometric redundancy\.
When the posterior mass remains broadly distributed throughout such fibers, the realization\-level KL term is small\. Concentration onto a relatively small subset of realizations increases the realization\-level contribution and makes PAC\-Bayes Z\-information more negative\.
Under the assumptions above, PAC\-Bayes Z\-information measures occupancy relative to available realization multiplicity rather than uncertainty over predictive behavior itself\.
This interpretation should be viewed as an intuition\-building special case rather than as part of the core theory\. The behavior–realization decomposition of Appendix[B](https://arxiv.org/html/2608.11465#A2)and the PAC\-Bayes results of Sections[3](https://arxiv.org/html/2608.11465#S3)and[4](https://arxiv.org/html/2608.11465#S4)rely only on measure\-theoretic disintegration and relative entropy; they do not require uniform conditional priors, finite\-volume fibers, or any geometric notion of reference measure\. The assumptions introduced here serve only to provide a more concrete geometric interpretation of realization multiplicity in continuous parameter spaces\.
## Appendix EGeometry of Behavioral Fibers
This appendix provides technical details supporting the geometric interpretation of realization multiplicity presented in Section[5](https://arxiv.org/html/2608.11465#S5)\.
The purpose is not to develop a full differential\-geometric theory of behavioral fibers\. Rather, we establish several elementary geometric facts that connect behavioral equivalence, symmetry, behavior\-preserving directions, and realization multiplicity\.
Throughout, letβ:Θ→𝒦\\beta:\\Theta\\rightarrow\\mathcal\{K\}denote the behavior map introduced in Section[2](https://arxiv.org/html/2608.11465#S2)\. For a behaviork∈𝒦k\\in\\mathcal\{K\}, the associated fiber is
Fk=β−1\(k\)\.F\_\{k\}=\\beta^\{\-1\}\(k\)\.The geometric terminology used in this appendix follows standard notions of differentiable curves and tangent vectors\([26](https://arxiv.org/html/2608.11465#bib.bib12);[12](https://arxiv.org/html/2608.11465#bib.bib27)\)\. No manifold structure on fibers is assumed beyond what is required for the specific statements\.
### E\.1Proof of Proposition[5](https://arxiv.org/html/2608.11465#Thmproposition5)
Recall Proposition[5](https://arxiv.org/html/2608.11465#Thmproposition5)\.
###### Proof\.
Nonnegativity follows from the basic properties of KL divergence\.
Suppose the fiberFkF\_\{k\}consists of a finite symmetry orbit of sizemmand that the conditional prior is uniform on that orbit:
P\(θ∣k\)=1m\.P\(\\theta\\mid k\)=\\frac\{1\}\{m\}\.For any conditional posteriorQ\(⋅∣k\)Q\(\\cdot\\mid k\), using the convention0log0=00\\log 0=0,
KL\(Q\(⋅∣k\)∥P\(⋅∣k\)\)=∑θ∈FkQ\(θ∣k\)log\(mQ\(θ∣k\)\)\.\\mathrm\{KL\}\\\!\\left\(Q\(\\cdot\\mid k\)\\,\\middle\\\|\\,P\(\\cdot\\mid k\)\\right\)=\\sum\_\{\\theta\\in F\_\{k\}\}Q\(\\theta\\mid k\)\\log\\\!\\bigl\(m\\,Q\(\\theta\\mid k\)\\bigr\)\.The minimum value is00, attained whenQ\(⋅∣k\)=P\(⋅∣k\)Q\(\\cdot\\mid k\)=P\(\\cdot\\mid k\)\.
The maximum value is attained when the posterior concentrates all of its mass on a single realization, in which case the divergence equalslogm\\log m\. Therefore
0≤KL\(Q\(⋅∣k\)∥P\(⋅∣k\)\)≤logm\.0\\leq\\mathrm\{KL\}\\\!\\left\(Q\(\\cdot\\mid k\)\\,\\middle\\\|\\,P\(\\cdot\\mid k\)\\right\)\\leq\\log m\.∎
The proposition shows that increasing the size of a symmetry orbit increases the maximum possible realization\-level contribution to classical PAC\-Bayes complexity while leaving predictive behavior unchanged\.
The finite\-orbit setting considered here may be viewed as the discrete counterpart of the occupancy interpretation developed in Appendix[D](https://arxiv.org/html/2608.11465#A4)\. In both cases, the conditional KL term measures posterior concentration relative to the realization structure associated with a fixed predictive behavior\.
### E\.2Behavior\-Preserving Curves
Behavioral fibers may contain continuous families of realizations\. When the configuration space admits a differentiable structure, such families can be represented by differentiable curves and their associated tangent directions\([26](https://arxiv.org/html/2608.11465#bib.bib12);[12](https://arxiv.org/html/2608.11465#bib.bib27)\)\.
Supposeγ:\(−ε,ε\)→Θ\\gamma:\(\-\\varepsilon,\\varepsilon\)\\to\\Thetais a differentiable curve satisfyingγ\(t\)∈Fk\\gamma\(t\)\\in F\_\{k\}for all sufficiently smalltt\.
Then every point on the curve induces the same predictive behavior\. Consequently, Proposition[1](https://arxiv.org/html/2608.11465#Thmproposition1)implies that both population and empirical risk remain constant along the curve\.
This observation provides the geometric counterpart of behavioral equivalence: motion within a behavioral fiber changes the realization of a predictor without changing the predictive behavior it induces\.
### E\.3Proof of Proposition[6](https://arxiv.org/html/2608.11465#Thmproposition6)
Recall Proposition[6](https://arxiv.org/html/2608.11465#Thmproposition6)\.
###### Proof\.
Letγ\(0\)=θ\\gamma\(0\)=\\thetaandγ′\(0\)=v\\gamma^\{\\prime\}\(0\)=v, and supposeγ\(t\)\\gamma\(t\)remains within a behavioral fiber\.
Because predictive behavior is unchanged andℒ\\mathcal\{L\}depends on a configuration only through its predictive behavior \(Proposition[1](https://arxiv.org/html/2608.11465#Thmproposition1)\),ℒ\(γ\(t\)\)\\mathcal\{L\}\(\\gamma\(t\)\)is constant\.
Differentiating with respect tottgives
0=ddtℒ\(γ\(t\)\)\|t=0\.0=\\frac\{d\}\{dt\}\\mathcal\{L\}\(\\gamma\(t\)\)\\Big\|\_\{t=0\}\.Applying the chain rule yields
∇ℒ\(θ\)⊤γ′\(0\)=0\.\\nabla\\mathcal\{L\}\(\\theta\)^\{\\top\}\\gamma^\{\\prime\}\(0\)=0\.Sinceγ′\(0\)=v\\gamma^\{\\prime\}\(0\)=v,
∇ℒ\(θ\)⊤v=0\.\\nabla\\mathcal\{L\}\(\\theta\)^\{\\top\}v=0\.
∎
Thus, tangent directions to behavior\-preserving curves preserve predictive behavior to first order and are therefore necessarily first\-order flat with respect to any behavior\-dependent objective\.
The vectorvvmay therefore be interpreted as a behavior\-preserving direction atθ\\theta\. More precisely,vvis the tangent vector to a curve that remains within a behavioral fiber and therefore preserves predictive behavior to first order\. The proposition shows that every such tangent direction is necessarily first\-order flat with respect to any behavior\-dependent objective\.
This result is intentionally weaker than the flat\-minimum analyses common in deep learning\([17](https://arxiv.org/html/2608.11465#bib.bib23);[22](https://arxiv.org/html/2608.11465#bib.bib24);[20](https://arxiv.org/html/2608.11465#bib.bib7)\)\. It establishes only first\-order invariance along behavioral fibers and does not require assumptions about optimization trajectories, Hessian spectra, or local curvature\.
### E\.4Continuous Realization Multiplicity
In discrete settings, realization multiplicity appears through finite symmetry orbits\.
In continuous settings, realization multiplicity may arise through continuous families of equivalent realizations\. Examples include permutation symmetries, scaling symmetries, redundant representations, and other forms of non\-identifiability in neural networks\([32](https://arxiv.org/html/2608.11465#bib.bib37);[11](https://arxiv.org/html/2608.11465#bib.bib22)\)\.
Behavioral equivalence therefore decomposes parameter\-space variation into two components: variation across behavioral fibers, which changes predictive behavior, and variation within fibers, which changes only the realization of that behavior\. The latter corresponds precisely to realization multiplicity\. This geometric distinction mirrors the behavior–realization decomposition of Sections[3](https://arxiv.org/html/2608.11465#S3)and[4](https://arxiv.org/html/2608.11465#S4)\.
From this perspective, the behavior map partitions parameter space into behavioral fibers, each containing all realizations of a fixed predictive behavior\.
When many independent behavior\-preserving directions exist, a single predictive behavior may be implemented by a large family of configurations\. Geometrically, realization multiplicity may therefore be associated with multiple independent behavior\-preserving directions or, relative to a chosen reference measure, with fibers that support many realizations of the same predictive behavior\.
Appendix[D](https://arxiv.org/html/2608.11465#A4)shows that, under additional assumptions on conditional priors and reference measures, the realization\-level KL term admits a corresponding occupancy\-based interpretation\. Complementing this geometric perspective, the variational characterization of Theorem[1](https://arxiv.org/html/2608.11465#Thmtheorem1)shows that behavior\-selection complexity is obtained by removing variation within fibers while preserving the induced distribution over predictive behaviors\.
### E\.5Summary
Behavioral fibers admit both discrete and continuous forms of realization multiplicity\.
Discrete multiplicity appears through symmetry orbits\. Continuous multiplicity may appear through behavior\-preserving directions and continuous families of equivalent realizations\.
These structures do not alter predictive behavior, but they influence how probability mass may be distributed among equivalent realizations\. Consequently, they provide geometric intuition for the realization\-level KL divergence that appears in the behavior–realization decomposition of Appendix[B](https://arxiv.org/html/2608.11465#A2)\. In particular, behavior\-preserving directions correspond to variations that leave predictive behavior unchanged while contributing to the realization multiplicity quantified by the conditional KL divergence and, equivalently, by PAC\-Bayes Z\-information\.
Together with Appendix[D](https://arxiv.org/html/2608.11465#A4), this geometric perspective complements the measure\-theoretic framework developed in the main text by relating realization multiplicity to symmetry, occupancy, and behavior\-preserving directions in over\-parameterized models\. The geometric interpretation is supplementary: the behavior–realization decomposition and the resulting PAC\-Bayes analysis rely only on the measure\-theoretic construction developed in Appendix[B](https://arxiv.org/html/2608.11465#A2)\.
## Appendix FBehavioral Invariance and Fiber\-Preserving Perturbations
This appendix provides technical details supporting the invariance results of Section[6](https://arxiv.org/html/2608.11465#S6)\.
The key observation is that behavioral equivalence separates perturbations that alter predictive behavior from perturbations that merely change the realization of that behavior\. The latter induces a notion of behavioral invariance that is distinct from classical algorithmic stability\([4](https://arxiv.org/html/2608.11465#bib.bib17);[16](https://arxiv.org/html/2608.11465#bib.bib41)\)\.
### F\.1Proof of Proposition[7](https://arxiv.org/html/2608.11465#Thmproposition7)
Recall Proposition[7](https://arxiv.org/html/2608.11465#Thmproposition7)\.
###### Proof\.
Letθ,θ′∈Fk\\theta,\\theta^\{\\prime\}\\in F\_\{k\}\. By the definition of the fiber,
β\(θ\)=β\(θ′\)=k\.\\beta\(\\theta\)=\\beta\(\\theta^\{\\prime\}\)=k\.Proposition[1](https://arxiv.org/html/2608.11465#Thmproposition1)shows that both population and empirical risks depend on a configuration only through its induced predictive behavior\. Therefore,
L\(θ\)=L\(θ′\),L^S\(θ\)=L^S\(θ′\)\.L\(\\theta\)=L\(\\theta^\{\\prime\}\),\\qquad\\hat\{L\}\_\{S\}\(\\theta\)=\\hat\{L\}\_\{S\}\(\\theta^\{\\prime\}\)\.∎
Thus, perturbations that remain within a behavioral fiber leave both population and empirical risk unchanged\.
### F\.2Behavioral Decomposition of Perturbations
The fibers of the behavior map form a partition of the configuration space\.
Consequently, any perturbation of a configuration either
1. 1\.moves the configuration to a different fiber and thereby changes predictive behavior; or
2. 2\.remains within the same fiber and therefore preserves predictive behavior\.
This elementary observation underlies the distinction between behavior\-changing and fiber\-preserving perturbations developed in Section[6](https://arxiv.org/html/2608.11465#S6)\. It also reflects the same behavior–realization distinction that underlies the information\-theoretic decomposition developed in Sections[3](https://arxiv.org/html/2608.11465#S3)and[4](https://arxiv.org/html/2608.11465#S4)\.
### F\.3Connection to Behavior\-Preserving Directions
Section[5](https://arxiv.org/html/2608.11465#S5)and Appendix[E](https://arxiv.org/html/2608.11465#A5)established that, under suitable regularity assumptions, tangent directions to behavior\-preserving curves are first\-order flat directions of a behavior\-dependent objective\([17](https://arxiv.org/html/2608.11465#bib.bib23);[22](https://arxiv.org/html/2608.11465#bib.bib24);[11](https://arxiv.org/html/2608.11465#bib.bib22)\)\.
Specifically, if
γ:\(−ε,ε\)→Θ\\gamma:\(\-\\varepsilon,\\varepsilon\)\\to\\Thetais a differentiable curve contained in a behavioral fiber withγ\(0\)=θ\\gamma\(0\)=\\thetaandγ′\(0\)=v\\gamma^\{\\prime\}\(0\)=v, then Proposition[6](https://arxiv.org/html/2608.11465#Thmproposition6)implies that
∇ℒ\(θ\)⊤v=0\.\\nabla\\mathcal\{L\}\(\\theta\)^\{\\top\}v=0\.Thus, fiber\-preserving perturbations represented by differentiable curves have zero first\-order effect on any behavior\-dependent objective\.
This provides a geometric interpretation of the invariance associated with behavioral equivalence\. Under the regularity assumptions of Proposition[6](https://arxiv.org/html/2608.11465#Thmproposition6), such directions arise as tangent directions to behavior\-preserving curves and represent first\-order variation among realizations that leaves predictive behavior unchanged\.
### F\.4Behavioral Invariance versus Algorithmic Stability
The invariance studied in this paper differs fundamentally from classical algorithmic stability\.
Algorithmic stability studies the sensitivity of learned predictors to perturbations of the training dataset\([4](https://arxiv.org/html/2608.11465#bib.bib17);[16](https://arxiv.org/html/2608.11465#bib.bib41)\)\.
By contrast, behavioral invariance concerns perturbations of a configuration that remain within a behavioral fiber and therefore preserve predictive behavior by construction\.
The two notions address different sources of variation:
1. 1\.algorithmic stability concerns perturbations of the training data;
2. 2\.behavioral invariance concerns perturbations of realizations that preserve predictive behavior\.
They should therefore be viewed as complementary rather than competing perspectives\.
### F\.5Interpretation
Behavioral equivalence implies that multiple distinct configurations may realize the same predictive behavior\.
Whenever multiple realizations implement a fixed predictive behavior, perturbations within a behavioral fiber leave both predictive behavior and risk unchanged\. Such realization multiplicity is closely related to the parameter\-space symmetries and non\-identifiability phenomena discussed in\([32](https://arxiv.org/html/2608.11465#bib.bib37);[11](https://arxiv.org/html/2608.11465#bib.bib22)\)\.
The realization\-level KL term and PAC\-Bayes Z\-information do not measure behavioral invariance directly\. Rather, they measure posterior concentration relative to the conditional prior within behavioral fibers\. They are therefore distributional quantities defined through conditional relative entropy and do not depend on any geometric structure of the fibers\. The geometric interpretation developed in Section[5](https://arxiv.org/html/2608.11465#S5)provides intuition, but the definitions themselves are entirely measure\-theoretic\.
Viewed through the variational characterization of Theorem[1](https://arxiv.org/html/2608.11465#Thmtheorem1), behavioral invariance explains why redistributing posterior mass within a behavioral fiber leaves predictive behavior unchanged while potentially altering configuration\-space PAC\-Bayes complexity\.
Behavioral invariance, geometric interpretations based on symmetry and behavior\-preserving directions, and the realization\-level KL term therefore provide complementary geometric, variational, and information\-theoretic perspectives on the same realization structure induced by behavioral equivalence\.
### F\.6Summary
Behavioral equivalence separates perturbations that alter predictive behavior from perturbations that merely alter realization\.
The latter leave predictive behavior and risk unchanged and, under appropriate regularity assumptions, include tangent directions to behavior\-preserving curves within behavioral fibers\.
The realization\-level KL term and PAC\-Bayes Z\-information quantify posterior concentration relative to the conditional prior within behavioral fibers and therefore provide an information\-theoretic perspective on the same realization structure underlying behavioral invariance\.
Although behavioral invariance is distinct from generalization itself, it helps explain why realization multiplicity appears naturally in the behavior–realization decomposition: movement within behavioral fibers changes realizations without changing predictive behavior\.
## Appendix GGeneral Z\-Information and Specialization to PAC\-Bayes
This appendix places PAC\-Bayes Z\-information within a more general measure\-theoretic framework based on measurable maps, disintegration, and relative entropy\.
The purpose is conceptual rather than technical\. It shows that the behavior–realization decomposition developed in the main text is a specialization of a general disintegration\-based decomposition of relative entropy\. No results from this appendix are required for the development of the main theory\.
### G\.1A General Fiber Decomposition
LetΘ\\Thetabe a measurable space equipped with probability measuresP,Q∈𝒫\(Θ\)P,Q\\in\\mathcal\{P\}\(\\Theta\), and let
r:Θ→ℛr:\\Theta\\rightarrow\\mathcal\{R\}be a measurable map into a measurable spaceℛ\\mathcal\{R\}\.
For eachr∈ℛr\\in\\mathcal\{R\}, the associated fiber is
Fr=\{θ∈Θ:r\(θ\)=r\}\.F\_\{r\}=\\\{\\theta\\in\\Theta:r\(\\theta\)=r\\\}\.
The pushforward measures
πQ=Q∘r−1,πP=P∘r−1\\pi\_\{Q\}=Q\\circ r^\{\-1\},\\qquad\\pi\_\{P\}=P\\circ r^\{\-1\}describe uncertainty over the image spaceℛ\\mathcal\{R\}\.
Under the standard disintegration assumptions\([21](https://arxiv.org/html/2608.11465#bib.bib29);[33](https://arxiv.org/html/2608.11465#bib.bib28);[3](https://arxiv.org/html/2608.11465#bib.bib15)\),
Q\(dθ\)=Q\(dθ∣r\)πQ\(dr\),P\(dθ\)=P\(dθ∣r\)πP\(dr\)\.Q\(d\\theta\)=Q\(d\\theta\\mid r\)\\,\\pi\_\{Q\}\(dr\),\\qquad P\(d\\theta\)=P\(d\\theta\\mid r\)\\,\\pi\_\{P\}\(dr\)\.
Applying the chain rule for relative entropy yields
KL\(Q∥P\)=KL\(πQ∥πP\)\+𝔼r∼πQ\[KL\(Q\(⋅∣r\)∥P\(⋅∣r\)\)\]\.\\mathrm\{KL\}\(Q\\\|P\)=\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\)\+\\mathbb\{E\}\_\{r\\sim\\pi\_\{Q\}\}\\\!\\left\[\\mathrm\{KL\}\\bigl\(Q\(\\cdot\\mid r\)\\,\\\|\\,P\(\\cdot\\mid r\)\\bigr\)\\right\]\.
The second term measures divergence arising from the redistribution of probability mass within fibers after the image\-space distribution has been fixed\.
By the nonnegativity of KL divergence, the fiber\-level term is always nonnegative and vanishes if and only if
Q\(⋅∣r\)=P\(⋅∣r\)Q\(\\cdot\\mid r\)=P\(\\cdot\\mid r\)forπQ\\pi\_\{Q\}\-almost everyrr\.
### G\.2Specialization to Behavioral Equivalence
The framework developed in the main text corresponds to the special case in which the measurable map is the behavior map
β:Θ→𝒦\.\\beta:\\Theta\\rightarrow\\mathcal\{K\}\.
The fibers
Fk=β−1\(k\)F\_\{k\}=\\beta^\{\-1\}\(k\)contain all configurations that induce the same predictive behavior\.
In this setting, the negative fiber\-level divergence becomes
𝒵PB\(Q∥P\)=−𝔼k∼πQ\[KL\(Q\(⋅∣k\)∥P\(⋅∣k\)\)\],\\mathcal\{Z\}\_\{\\mathrm\{PB\}\}\(Q\\\|P\)=\-\\mathbb\{E\}\_\{k\\sim\\pi\_\{Q\}\}\\\!\\left\[\\mathrm\{KL\}\\bigl\(Q\(\\cdot\\mid k\)\\,\\\|\\,P\(\\cdot\\mid k\)\\bigr\)\\right\],which is precisely the PAC\-Bayes Z\-information introduced in Section[3](https://arxiv.org/html/2608.11465#S3)\.
Thus, PAC\-Bayes Z\-information is obtained by specializing the general disintegration framework to the behavior map and taking the negative of the resulting expected within\-fiber KL divergence\.
### G\.3Why the Behavioral Setting is Special
Many measurable maps induce fiber decompositions, but the behavioral setting possesses an additional property that is crucial for learning theory\.
By Proposition[1](https://arxiv.org/html/2608.11465#Thmproposition1),
β\(θ\)=β\(θ′\)⟹L\(θ\)=L\(θ′\),L^S\(θ\)=L^S\(θ′\)\.\\beta\(\\theta\)=\\beta\(\\theta^\{\\prime\}\)\\quad\\Longrightarrow\\quad L\(\\theta\)=L\(\\theta^\{\\prime\}\),\\qquad\\hat\{L\}\_\{S\}\(\\theta\)=\\hat\{L\}\_\{S\}\(\\theta^\{\\prime\}\)\.
Consequently, both population and empirical risk factor through the behavior map\. Variation within a behavioral fiber does not directly affect predictive behavior or risk\.
This property gives the decomposition developed in Sections[3](https://arxiv.org/html/2608.11465#S3)and[4](https://arxiv.org/html/2608.11465#S4)its learning\-theoretic significance\. The pushforward distribution over behaviors captures behavior\-level uncertainty relevant to prediction, whereas the conditional distributions within fibers capture realization\-level variation among alternative realizations of the same predictive behavior\.
### G\.4Summary
PAC\-Bayes Z\-information arises from a general disintegration\-based decomposition of relative entropy into image\-space and fiber\-level components\.
What distinguishes the behavioral setting is that predictive risk depends only on behavior\. This allows realization\-level variation to be separated from behavior\-selection complexity and gives the within\-fiber KL divergence a direct learning\-theoretic interpretation\.
The PAC\-Bayes framework developed in the main text is obtained by specializing this general construction to the behavior map and interpreting the resulting fibers as alternative realizations of the same predictive behavior\. In this specialization, the generic fiber\-level decomposition becomes the behavior–realization decomposition, the image\-space divergence becomes behavior\-selection complexity, and the within\-fiber divergence becomes PAC\-Bayes Z\-information\.
## Appendix HWorked Examples and Explicit Computations
The preceding appendices establish the measure\-theoretic foundations, prove the main decomposition theorems, and develop their geometric and information\-theoretic consequences\. This appendix serves a complementary purpose by presenting explicit computations illustrating how the behavior–realization decomposition is instantiated in concrete settings\. The first example computes every quantity appearing in the decomposition and illustrates the variational characterization through fiber symmetrization\. The second derives the realization\-level contribution associated with hidden\-unit permutation symmetry, complementing the illustrative discussion in the main text\.
### H\.1A Finite Behavioral Quotient
We begin with the simplest nontrivial example in which every quantity appearing in the behavior–realization decomposition can be computed exactly\.
Consider the configuration space
Θ=\{θ1,θ2,θ3,θ4\},\\Theta=\\\{\\theta\_\{1\},\\theta\_\{2\},\\theta\_\{3\},\\theta\_\{4\}\\\},and define the behavior map
β\(θ1\)=β\(θ2\)=kA,β\(θ3\)=β\(θ4\)=kB\.\\beta\(\\theta\_\{1\}\)=\\beta\(\\theta\_\{2\}\)=k\_\{A\},\\qquad\\beta\(\\theta\_\{3\}\)=\\beta\(\\theta\_\{4\}\)=k\_\{B\}\.
The behavioral fibers are therefore
β−1\(kA\)=\{θ1,θ2\},β−1\(kB\)=\{θ3,θ4\}\.\\beta^\{\-1\}\(k\_\{A\}\)=\\\{\\theta\_\{1\},\\theta\_\{2\}\\\},\\qquad\\beta^\{\-1\}\(k\_\{B\}\)=\\\{\\theta\_\{3\},\\theta\_\{4\}\\\}\.
Suppose the prior is
P=\(14,14,14,14\),P=\\left\(\\frac\{1\}\{4\},\\frac\{1\}\{4\},\\frac\{1\}\{4\},\\frac\{1\}\{4\}\\right\),and the posterior is
Q=\(0\.6,0,0\.4,0\)\.Q=\(0\.6,0,0\.4,0\)\.
Thus, the posterior assigns all of its probability mass to a single configuration within each behavioral fiber\.
#### Step 1: Behavior\-Level Distributions\.
Applying the behavior map gives
πP=\(0\.5,0\.5\),πQ=\(0\.6,0\.4\)\.\\pi\_\{P\}=\(0\.5,0\.5\),\\qquad\\pi\_\{Q\}=\(0\.6,0\.4\)\.
Hence,
KL\(πQ∥πP\)=0\.6log0\.60\.5\+0\.4log0\.40\.5≈0\.0201\.\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\)=0\.6\\log\\frac\{0\.6\}\{0\.5\}\+0\.4\\log\\frac\{0\.4\}\{0\.5\}\\approx 0\.0201\.
#### Step 2: Fiber Conditional Distributions\.
Within each behavioral fiber,
P\(⋅∣kA\)=P\(⋅∣kB\)=\(12,12\),P\(\\cdot\\mid k\_\{A\}\)=P\(\\cdot\\mid k\_\{B\}\)=\\left\(\\frac\{1\}\{2\},\\frac\{1\}\{2\}\\right\),whereas
Q\(⋅∣kA\)=Q\(⋅∣kB\)=\(1,0\)\.Q\(\\cdot\\mid k\_\{A\}\)=Q\(\\cdot\\mid k\_\{B\}\)=\(1,0\)\.
Therefore,
KL\(Q\(⋅∣ki\)∥P\(⋅∣ki\)\)=log2,i=A,B\.\\mathrm\{KL\}\\\!\\left\(Q\(\\cdot\\mid k\_\{i\}\)\\,\\middle\\\|\\,P\(\\cdot\\mid k\_\{i\}\)\\right\)=\\log 2,\\qquad i=A,B\.
Taking the expectation overπQ\\pi\_\{Q\}gives
𝔼k∼πQ\[KL\(Q\(⋅∣k\)∥P\(⋅∣k\)\)\]=log2≈0\.6931\.\\mathbb\{E\}\_\{k\\sim\\pi\_\{Q\}\}\\\!\\left\[\\mathrm\{KL\}\\\!\\left\(Q\(\\cdot\\mid k\)\\,\\middle\\\|\\,P\(\\cdot\\mid k\)\\right\)\\right\]=\\log 2\\approx 0\.6931\.
Hence,
𝒵PB=−log2\.\\mathcal\{Z\}\_\{\\mathrm\{PB\}\}=\-\\log 2\.
#### Step 3: The KL Decomposition\.
The classical divergence is
KL\(Q∥P\)=0\.6log0\.60\.25\+0\.4log0\.40\.25≈0\.7132\.\\mathrm\{KL\}\(Q\\\|P\)=0\.6\\log\\frac\{0\.6\}\{0\.25\}\+0\.4\\log\\frac\{0\.4\}\{0\.25\}\\approx 0\.7132\.
Thus,
KL\(Q∥P\)=KL\(πQ∥πP\)−𝒵PB,\\mathrm\{KL\}\(Q\\\|P\)=\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\)\-\\mathcal\{Z\}\_\{\\mathrm\{PB\}\},or numerically,
0\.7132=0\.0201\+0\.6931\.0\.7132=0\.0201\+0\.6931\.
#### Step 4: Fiber Symmetrization\.
The fiber\-symmetrized posterior is
Q⋆=\(0\.3,0\.3,0\.2,0\.2\)\.Q^\{\\star\}=\(0\.3,0\.3,0\.2,0\.2\)\.
Since
πQ⋆=πQ,\\pi\_\{Q^\{\\star\}\}=\\pi\_\{Q\},both posteriors induce identical predictive behavior\.
Moreover,
Q⋆\(⋅∣k\)=P\(⋅∣k\)Q^\{\\star\}\(\\cdot\\mid k\)=P\(\\cdot\\mid k\)for every behavioral fiber\.
Consequently,
𝒵PB\(Q⋆∥P\)=0,\\mathcal\{Z\}\_\{\\mathrm\{PB\}\}\(Q^\{\\star\}\\\|P\)=0,and
KL\(Q⋆∥P\)=KL\(πQ∥πP\)≈0\.0201\.\\mathrm\{KL\}\(Q^\{\\star\}\\\|P\)=\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\)\\approx 0\.0201\.
The principal quantities appearing in the decomposition are summarized in Table[1](https://arxiv.org/html/2608.11465#A8.T1)\.
Table 1:Summary of the behavior–realization decomposition for Example[H\.1](https://arxiv.org/html/2608.11465#A8.SS1)\.
#### Interpretation\.
This example illustrates every component of the behavior–realization decomposition\. Although the posteriorQQand its fiber\-symmetrized representativeQ⋆Q^\{\\star\}induce identical predictive behavior, their classical PAC\-Bayes complexities differ by more than a factor of thirty\. In this example, over97%97\\%of the classical PAC\-Bayes complexity arises from realization\-level divergence rather than uncertainty over predictive behavior\. Fiber symmetrization removes this realization\-level contribution while preserving predictive behavior, thereby realizing the variational characterization established in Theorem[1](https://arxiv.org/html/2608.11465#Thmtheorem1)\.
### H\.2Explicit Computation for Hidden\-Unit Permutation Symmetry
Example[H\.1](https://arxiv.org/html/2608.11465#A8.SS1)computed every component of the behavior–realization decomposition in a finite setting\. We now isolate the realization\-level term for the hidden\-unit permutation symmetry discussed in the main text\.
Consider a behavioral fiber generated by all permutations ofhhexchangeable hidden units\. Under the idealized assumptions described in Section[3](https://arxiv.org/html/2608.11465#S3), suppose the conditional prior assigns equal probability to every realization within the fiber,
P\(⋅∣k\)=\(1h\!,…,1h\!\),P\(\\cdot\\mid k\)=\\left\(\\frac\{1\}\{h\!\},\\ldots,\\frac\{1\}\{h\!\}\\right\),while the conditional posterior concentrates on one representative realization,
Q\(⋅∣k\)=\(1,0,…,0\)\.Q\(\\cdot\\mid k\)=\(1,0,\\ldots,0\)\.
The conditional KL divergence within the fiber is therefore
KL\(Q\(⋅∣k\)∥P\(⋅∣k\)\)\\displaystyle\\mathrm\{KL\}\\\!\\left\(Q\(\\cdot\\mid k\)\\,\\middle\\\|\\,P\(\\cdot\\mid k\)\\right\)=∑i=1h\!QilogQiPi\\displaystyle=\\sum\_\{i=1\}^\{h\!\}Q\_\{i\}\\log\\frac\{Q\_\{i\}\}\{P\_\{i\}\}=log\(h\!\)\.\\displaystyle=\\log\(h\!\)\.
If the same conditional structure holds forπQ\\pi\_\{Q\}\-almost every behavioral fiber, then
−𝒵PB=𝔼k∼πQ\[KL\(Q\(⋅∣k\)∥P\(⋅∣k\)\)\]=log\(h\!\),\-\\mathcal\{Z\}\_\{\\mathrm\{PB\}\}=\\mathbb\{E\}\_\{k\\sim\\pi\_\{Q\}\}\\\!\\left\[\\mathrm\{KL\}\\\!\\left\(Q\(\\cdot\\mid k\)\\,\\middle\\\|\\,P\(\\cdot\\mid k\)\\right\)\\right\]=\\log\(h\!\),and consequently
KL\(Q∥P\)=KL\(πQ∥πP\)\+log\(h\!\)\.\\mathrm\{KL\}\(Q\\\|P\)=\\mathrm\{KL\}\(\\pi\_\{Q\}\\\|\\pi\_\{P\}\)\+\\log\(h\!\)\.
Using Stirling’s approximation,
log\(h\!\)=hlogh−h\+O\(logh\)\.\\log\(h\!\)=h\\log h\-h\+O\(\\log h\)\.
For illustration,
log\(10\!\)≈15\.10,log\(100\!\)≈363\.74\.\\log\(10\!\)\\approx 15\.10,\\qquad\\log\(100\!\)\\approx 363\.74\.
#### Interpretation\.
This computation makes explicit the approximation summarized in the main text\. The additionallog\(h\!\)\\log\(h\!\)term arises entirely from concentration of posterior probability within a behavioral fiber and does not reflect greater uncertainty over predictive behavior\. As the degree of symmetry increases, realization multiplicity can therefore contribute substantially to the classical PAC\-Bayes KL divergence even when predictive behavior remains unchanged\.Similar Articles
PAC--Bayes Bounds on Quotient Parameter Spaces: Geometry-induced Implicit-Bias Priors
This paper proposes performing PAC-Bayesian analysis on quotient parameter spaces to remove KL contributions from parameter symmetries, and constructs a geometry-induced prior that approximates the ideal posterior-matched prior, resulting in tighter generalization bounds. Experiments on Fourier regression and Query-Key attention show significant reductions in KL divergence and certificate values.
Identifying Informative Environments for Cognition Parameter Inference via Bayesian Experimental Design
This paper formulates cognitive experiment design as a Bayesian Experimental Design problem, introducing an amortized framework that efficiently identifies maximally informative environments for inferring latent planning parameters, with validated performance on the Mouselab-MDP paradigm.
A PAC-Bayesian View of Generalisation for Physics-Informed Machine Learning
This paper develops a PAC-Bayesian framework for physics-informed machine learning, providing high-probability generalization guarantees for unbounded losses. It proposes a multi-task perspective that jointly handles data fidelity, PDE residuals, and boundary conditions, and introduces a self-bounding learning algorithm.
A Rate Separation for Agnostic Direct Sums
This paper answers an open question from Hanneke, Moran, and Waknine by showing that the agnostic PAC learning curve of a direct sum is not determined solely by the single-instance learning curve and the number of factors, providing a rate separation.
A Generalized-Bayes Perspective on Counterfactual Explanations: Posterior-Based Decision-Making and Evaluation
This paper connects counterfactual explanations to generalized Bayes inference, showing that distance-minimization CEs are MAP estimates of a Gibbs posterior, and introduces new decision rules and evaluation metrics.