Variation Brownian Kernel Ladders
Summary
The paper introduces Variation Brownian Kernel Ladders (VBKL), a path-atomic function-space framework for recursive dictionary construction, with analysis on Hölder regularity, compactness, and generalization bounds in statistical learning theory.
View Cached Full Text
Cached at: 08/17/26, 10:15 AM
# Variation Brownian Kernel Ladders
Source: [https://arxiv.org/html/2608.13882](https://arxiv.org/html/2608.13882)
Mahdi Mohammadigohari Mahdi\.Mohammadigohari@gmail\.comAffiliation:Faculty of Engineering, Free University of Bozen\-BolzanoAffiliation:Bruno Buozzi 1, Bolzano, 39100, Italy
###### Abstract
Claims about the benefit of depth depend on the complexity assigned to a representation\. We introduce the*Variation Brownian Kernel Ladder*\(VBKL\), a path\-atomic function\-space framework that separates nonlinear recursive dictionary construction from linear variation superposition\. Starting from linear projections, each atom recursively composes unit\-ball profiles from the Brownian reproducing kernel Hilbert space; the full VBKL space is then the signed\-measure variation hull of the completed dictionary\. We identify each recursive dictionary as a union of Brownian pullback RKHS balls and establish variation\-controlled Hölder regularity, compactness and attainment, and strict growth with depth under a local non\-degeneracy condition whose trace lies in the support of the input measure\. For associated finite lower\-support architectures, we derive Rademacher and generalization bounds through Brownian quadratic chaos, signed threshold traces, and VC entropy\. We also construct two\-stage approximants by discretizing the outer measure and the selected outer Brownian profiles, obtaining anM−1/2\+m−1/2M^\{\-1/2\}\+m^\{\-1/2\}error bound, a sharp interpolation constantA/2\\sqrt\{A/2\}, and at most2M2Mactive outer\-profile basis contributions per evaluation\. Controlled experiments illustrate the approximation mechanisms and indicate a favorable limited\-data accuracy–complexity trade\-off\.
††shortheadings:Variation Brownian Kernel Ladders / Mohammadigohari††firstpage:1###### keywords
recursive function spaces, variation spaces, Brownian kernels, statistical learning theory, constructive approximation
## 1Introduction
Any claim about the representational benefit of depth is relative to a notion of complexity\. A function may be inexpensive when complexity is measured by width or parameter count and expensive when it is measured by weight magnitude, path variation, or an intrinsic function\-space norm\. This distinction is especially important for overparameterized models, where the size of one finite realization may reveal little about the class of functions favored by learning\. Classical and modern capacity analyses therefore control neural function classes through weight magnitudes and norm\-based representation costs, often independently of width\([5](https://arxiv.org/html/2608.13882#bib.bib11);[20](https://arxiv.org/html/2608.13882#bib.bib19);[14](https://arxiv.org/html/2608.13882#bib.bib15);[3](https://arxiv.org/html/2608.13882#bib.bib10)\)\. The corresponding function\-space viewpoint associates a predictor with the function it represents rather than with one particular parameterization and asks which functions lie in a norm\- or gauge\-bounded class\.
A rich theory has been developed for shallow networks through Barron spaces, measure\-generated variation spaces, ridge\-spline representations, and reproducing kernel Banach spaces\([4](https://arxiv.org/html/2608.13882#bib.bib4);[2](https://arxiv.org/html/2608.13882#bib.bib5);[22](https://arxiv.org/html/2608.13882#bib.bib21);[27](https://arxiv.org/html/2608.13882#bib.bib26);[23](https://arxiv.org/html/2608.13882#bib.bib22);[24](https://arxiv.org/html/2608.13882#bib.bib24);[17](https://arxiv.org/html/2608.13882#bib.bib18);[29](https://arxiv.org/html/2608.13882#bib.bib28);[30](https://arxiv.org/html/2608.13882#bib.bib29);[6](https://arxiv.org/html/2608.13882#bib.bib12)\)\. These constructions connect approximation, regularization, statistical complexity, and representer theorems through intrinsic function\-space geometry\. Extending this perspective to depth has led to several distinct recursive models\. Compositional and recursive RKHS theories build Hilbert\-space hierarchies through kernel or function composition\([8](https://arxiv.org/html/2608.13882#bib.bib9);[15](https://arxiv.org/html/2608.13882#bib.bib16)\)\. Recursive RKBS and Banach\-space constructions emphasize vector\-valued measures, layerwise variational structure, and representer theorems\([10](https://arxiv.org/html/2608.13882#bib.bib33);[7](https://arxiv.org/html/2608.13882#bib.bib13)\)\. Deep variation\-space and neural\-tree formulations instead iterate atomic, measure, or composed shallow\-space constructions\([25](https://arxiv.org/html/2608.13882#bib.bib23);[28](https://arxiv.org/html/2608.13882#bib.bib27);[19](https://arxiv.org/html/2608.13882#bib.bib34);[21](https://arxiv.org/html/2608.13882#bib.bib20)\)\.
The resulting theories differ not only in their choice of activation or kernel, but also in the location at which linear superposition is introduced\. This design choice determines the recursive object being studied\. If convexification or measure superposition is applied at every layer, then nonlinear feature generation and linear combination evolve together\. If superposition is postponed, one may first study the geometry of individual compositional paths and only afterward ask how much linear variation is required to combine them\. These operations need not commute, and they can lead to different conclusions about regularity, expressivity, statistical complexity, and finite realizability\. This motivates the central question of the paper:*how does the placement of linear superposition within a recursive function\-space construction affect the geometry, learning complexity, and constructive realization of the resulting deep function classes?*
We answer this question through the*Variation Brownian Kernel Ladder*\(VBKL\)\. The construction begins with linear projections and recursively builds a nonlinear path dictionary\. At each subsequent level, one chooses a lower\-level atom and composes it with a unit\-ball function from the Brownian RKHS\. Thus, a depth\-LLatom contains one support path and exactlyL−1L\-1normalized Brownian profiles\. No convex or variation hull is taken at the intermediate levels\. Only after the depth\-LLdictionary has been constructed do we introduce linear superposition, by integrating its atoms against finite signed measures\. The infimal total variation over all such representations defines the intrinsic VBKL complexity\. This order of construction separates the geometry of recursive atoms, the cost of their outer linear combination, and the complexity of a finite realization used for computation\.
##### Why the Brownian RKHS?
The Brownian kernel is not used merely as a convenient nonlinear activation\. Its RKHS is the explicit anchored Sobolev–Cameron–Martin space of absolutely continuous functions with square\-integrable derivative\. Brownian pullbacks admit exact RKHS descriptions, while the Brownian kernel metric propagates square\-root regularity through recursive composition\. The signed\-threshold representation of the kernel connects finite Brownian variation classes to quadratic chaos and VC entropy\. Finally, the one\-dimensional profile geometry permits piecewise\-linear approximation with a sharp constant\. These properties allow analytical, statistical, and constructive theories to be developed within the same recursive model\.
##### Brownian and recursive\-kernel lineage\.
The shallow Brownian projection model of[12](https://arxiv.org/html/2608.13882#bib.bib6)represents predictors as expectations of one\-dimensional Sobolev functions over learned projections and identifies the associated Brownian projection kernel\. Brownian Kernel Ladders\([18](https://arxiv.org/html/2608.13882#bib.bib1)\)recursively average Brownian pullback kernels to construct integral RKHS hierarchies\. VBKL uses the same Brownian function\-generation mechanism in a different way: it follows individual compositional supports to construct a path dictionary and places the signed\-measure superposition only at the outermost level\. The three constructions are therefore related through the Brownian kernel, but differ in whether projection measures, recursive kernel measures, or outer atomic measures are the primary representation variables\.
##### Relation to Deep Neural Variation Spaces\.
The closest deep variation\-space framework is that of[19](https://arxiv.org/html/2608.13882#bib.bib34)\. Schematically, its depth\-llunit ball and the VBKL depth\-lldictionary are constructed as
ℬlDNVS\\displaystyle\\mathcal\{B\}\_\{l\}^\{\\mathrm\{DNVS\}\}=aconv¯\(\{σs∘f:s\>0,f∈ℬl−1DNVS\}\),\\displaystyle=\\overline\{\\operatorname\{aconv\}\}\\left\(\\left\\\{\\sigma\_\{s\}\\circ f:s\>0,\\;f\\in\\mathcal\{B\}\_\{l\-1\}^\{\\mathrm\{DNVS\}\}\\right\\\}\\right\),𝒰l\\displaystyle\\mathcal\{U\}\_\{l\}=\{g∘u:u∈𝒰l−1,g∈ℋk\(B\),‖g‖ℋk\(B\)≤1\}\.\\displaystyle=\\left\\\{g\\circ u:u\\in\\mathcal\{U\}\_\{l\-1\},\\;g\\in\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\},\\;\\left\\\|g\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}\\leq 1\\right\\\}\.Deep Neural Variation Spaces therefore alternate nonlinear activation and closed absolute convexification at every level, whereas VBKL recursively constructs only the nonlinear dictionary and takes its variation hull after the desired depth is reached\. There is a second distinction: the former varies a normalized family derived from one prescribed activation, while each VBKL profile ranges over the entire unit ball of the Brownian RKHS\. Because nonlinear composition and absolute convexification do not generally commute, VBKL is not obtained merely by specializing that framework to a Brownian activation family, and we make no general inclusion or equivalence claim between the resulting spaces\. The two theories also reveal complementary depth phenomena\. Deep Neural Variation Spaces establish depth saturation for norm\-controlled univariate ReLU classes, whereas the Brownian path construction below yields a strict hierarchy under a local non\-degeneracy condition along the support of the input measure\. Their finite\-representation results are likewise complementary: the former proves a representer theorem for norm\-penalized data fitting, while VBKL gives a two\-stage approximation of every element of its full infinite\-dimensional space\.
##### Depth separation and approximation\.
Classical depth\-separation results compare finite networks through width, parameter count, or related representation costs\([31](https://arxiv.org/html/2608.13882#bib.bib30);[11](https://arxiv.org/html/2608.13882#bib.bib14);[34](https://arxiv.org/html/2608.13882#bib.bib32);[33](https://arxiv.org/html/2608.13882#bib.bib31);[26](https://arxiv.org/html/2608.13882#bib.bib25)\)\. Our strict\-hierarchy result addresses a different question: whether adjacent intrinsic, potentially infinite\-width, norm\-controlled function spaces contain genuinely different functions\. The constructive theory is also connected to nonlinear, greedy, and variable\-basis approximation\([9](https://arxiv.org/html/2608.13882#bib.bib7);[32](https://arxiv.org/html/2608.13882#bib.bib8);[16](https://arxiv.org/html/2608.13882#bib.bib17);[30](https://arxiv.org/html/2608.13882#bib.bib29)\)\. Here the two approximation resources have distinct mathematical meanings: the number of atoms controls discretization of the outer signed measure, whereas the profile resolution controls realization of each selected Brownian nonlinearity\.
##### Scope of the results\.
The full path\-atomic VBKL space and the associated finite mixed architecture play different roles\. Our analytical and constructive results concern the full infinite\-dimensional space generated by signed measures over the path dictionary\. The statistical guarantees concern an explicitly parameterized finite Brownian variation class whose lower\-support architecture may include intermediate linear mixing\. The path\-only specialization of that architecture is a subclass of the full VBKL space, but no such inclusion is asserted for the general mixed architecture\. Stating this distinction explicitly prevents finite\-parameter complexity from being attributed to the unrestricted infinite\-dimensional variation ball\.
The principal contributions are as follows\.
1. 1\.Path\-atomic recursive variation spaces\.We introduce a recursive Brownian dictionary in which every new atom is the composition of one lower\-level atom with one normalized Brownian RKHS profile\. We identify each dictionary exactly as a union of Brownian pullback RKHS unit balls and define the depth\-LLVBKL space as its finite signed\-measure variation hull\. This gives an explicit infinite\-dimensional representation while keeping nonlinear recursion and linear superposition mathematically distinct\.
2. 2\.Full\-space geometry\.We establish well\-definedness and non\-degeneracy of the variation complexity\. Under compactness of the input and direction sets, the recursive dictionary is compact and the minimum\-total\-variation representation is attained\. Variation complexity controls pointwise magnitude and2−\(L−1\)2^\{\-\(L\-1\)\}\-Hölder regularity\. When the input measure has full support, the resulting representative is unique and yields a continuous Hölder embedding\. Under a local non\-degeneracy condition whose trace lies in the support of the input measure, the spaces form a strict hierarchy across depth\.
3. 3\.Architecture\-dependent statistical guarantees\.For explicit finite lower\-support architectures with piecewise\-linear Brownian profiles, normalized mixing, and parameter countPL−1,m,GP\_\{L\-1,m,G\}, we derive empirical and expected Rademacher bounds and a high\-probability generalization guarantee\. The proof passes from the union\-of\-RKHS representation to Brownian quadratic chaos, represents that chaos through signed threshold traces, and controls the trace entropy through VC theory\. The bound separates outer variation radius, recursive Brownian range, finite architecture size, and sample size\.
4. 4\.Constructive approximation with separated resources\.Every function in the full VBKL space is first approximated by at mostMMrecursive atoms with error of orderM−1/2M^\{\-1/2\}\. The selected outer Brownian profiles are then replaced by piecewise\-linear interpolants with an additionalm−1/2m^\{\-1/2\}error\. Balanced refinement yields anO\(N−1/2\)O\\left\(N^\{\-1/2\}\\right\)approximation with finite outer atomic support and discretized outer profiles\. The interpolation constantA/2\\sqrt\{A/2\}is optimal, and evaluating the resulting model uses at most2M2Mactive outer\-profile basis contributions independently ofmm\.
5. 5\.Theory\-directed experiments\.Controlled experiments isolate signed\-measure discretization, profile discretization, and balanced refinement, and the normalized tent construction attains the worst\-case interpolation bound exactly\. Supervised studies compare validation\-selected VBKL realizations with Deep Neural Variation Spaces and kernel baselines\. They indicate the strongest VBKL behavior in limited\-data regimes and a favorable accuracy–parameter trade\-off rather than universal predictive dominance\. Additional experiments verify the finite\-difference directional estimator, the variance reduction from Monte Carlo directional averaging, and the practical realizability of the finite models\.
The remainder of the paper is organized as follows\. Section[2](https://arxiv.org/html/2608.13882#S2)introduces the recursive Brownian dictionary, the associated VBKL space, and the finite lower\-support architectures\. Section[3](https://arxiv.org/html/2608.13882#S3)establishes analytical properties of the full VBKL spaces and architecture\-dependent statistical guarantees for the associated finite Brownian variation classes\. Section[4](https://arxiv.org/html/2608.13882#S4)develops the two\-stage constructive approximation theory\. Section[5](https://arxiv.org/html/2608.13882#S5)presents theory\-directed experiments on approximation, statistical learning, parameter efficiency, and optimization\. Complete protocols, numerical tables, selected configurations, and computational measurements are reported in Appendix[A](https://arxiv.org/html/2608.13882#A1)\. Proofs of the main theoretical results are collected in Section[7](https://arxiv.org/html/2608.13882#S7), while additional notation, the sharpness analysis of the Brownian profile interpolation estimate, and the supporting auxiliary lemmas are provided in Appendices[B](https://arxiv.org/html/2608.13882#A2),[C](https://arxiv.org/html/2608.13882#A3), and[D](https://arxiv.org/html/2608.13882#A4), respectively\.
ResultContentPageTheorem[4](https://arxiv.org/html/2608.13882#Thmtheorem4)Analytical properties of the VBKL spacespage[4](https://arxiv.org/html/2608.13882#Thmtheorem4)Theorem[6](https://arxiv.org/html/2608.13882#Thmtheorem6)Architecture\-dependent Rademacher boundpage[6](https://arxiv.org/html/2608.13882#Thmtheorem6)Corollary[8](https://arxiv.org/html/2608.13882#Thmtheorem8)Finite\-architecture generalization guaranteepage[8](https://arxiv.org/html/2608.13882#Thmtheorem8)Theorem[9](https://arxiv.org/html/2608.13882#Thmtheorem9)Finite\-atomic approximation of the full VBKL spacepage[9](https://arxiv.org/html/2608.13882#Thmtheorem9)Theorem[10](https://arxiv.org/html/2608.13882#Thmtheorem10)Two\-stage finite\-atomic and outer\-profile approximationpage[10](https://arxiv.org/html/2608.13882#Thmtheorem10)Proposition[1](https://arxiv.org/html/2608.13882#Thmprop1)Brownian profile interpolationpage[1](https://arxiv.org/html/2608.13882#Thmprop1)Table 1:Overview of the main theoretical results\. Their logical dependencies are illustrated in Fig\.[5](https://arxiv.org/html/2608.13882#A0.F5); supporting results are summarized in Table[9](https://arxiv.org/html/2608.13882#A3.T9)\.
## 2Notation
We use the following notation throughout the paper\. The set of natural numbers is denoted byℕ=\{1,2,…\}\\mathbb\{N\}=\\\{1,2,\\ldots\\\},ℝ\\mathbb\{R\}denotes the set of real numbers, and\[n\]=\{1,…,n\}\[n\]=\\\{1,\\ldots,n\\\}\. Throughout the paper,𝒳⊆ℝd\\mathcal\{X\}\\subseteq\\mathbb\{R\}^\{d\}denotes a compact input domain equipped with a probability measureν\\nu, andΩ⊆𝕊d−1\\Omega\\subseteq\\mathbb\{S\}^\{d\-1\}denotes the admissible set of first\-layer directions\. The Brownian kernel isk\(B\)\(x,x′\)=\(\|x\|\+\|x′\|−\|x−x′\|\)/2,k^\{\(\\mathrm\{B\}\)\}\(x,x^\{\\prime\}\)=\(\|x\|\+\|x^\{\\prime\}\|\-\|x\-x^\{\\prime\}\|\)/2,whose associated RKHS isℋk\(B\)=\{g:g\(0\)=0,gabsolutely continuous,g′∈L2\(ℝ\)\},\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}=\\\{g:g\(0\)=0,\\;g\\text\{ absolutely continuous\},\\;g^\{\\prime\}\\in L^\{2\}\(\\mathbb\{R\}\)\\\},equipped with the norm‖g‖ℋk\(B\)2=∫ℝ\|g′\(t\)\|2𝑑t\.\\\|g\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}^\{2\}=\\int\_\{\\mathbb\{R\}\}\|g^\{\\prime\}\(t\)\|^\{2\}\\,dt\.For a measurable spaceΘ\\Theta, we writeℳ\(Θ\)\\mathcal\{M\}\(\\Theta\)for the space of finite signed measures equipped with the total variation norm‖μ‖TV=\|μ\|\(Θ\)\.\\\|\\mu\\\|\_\{\\mathrm\{TV\}\}=\|\\mu\|\(\\Theta\)\.Additional notation is collected in Appendix[B](https://arxiv.org/html/2608.13882#A2)\.
### 2\.1Atomic Brownian Kernel Ladder
The depth\-LLVBKL space is generated by a recursively constructed Brownian dictionary\. In BKL\([18](https://arxiv.org/html/2608.13882#bib.bib1)\), a recursive level is formed by averaging Brownian pullback kernels over a probability measure on lower\-level supports: if𝒮\\mathcal\{S\}is a support class andμ∈𝒫1\(𝒮\)\\mu\\in\\mathcal\{P\}\_\{1\}\(\\mathcal\{S\}\), then
k\[𝒮,μ\]\(𝐱,𝐱′\)\\displaystyle k\[\\mathcal\{S\},\\mu\]\\left\(\\mathbf\{x\},\\mathbf\{x\}^\{\\prime\}\\right\)=∫𝒮k\(B\)\(u\(𝐱\),u\(𝐱′\)\)𝑑μ\(u\),𝐱,𝐱′∈𝒳\.\\displaystyle=\\int\_\{\\mathcal\{S\}\}k^\{\(\\mathrm\{B\}\)\}\\left\(u\(\\mathbf\{x\}\),u\(\\mathbf\{x\}^\{\\prime\}\)\\right\)\\,\\mathrm\{d\}\\mu\(u\),\\qquad\\mathbf\{x\},\\mathbf\{x\}^\{\\prime\}\\in\\mathcal\{X\}\.\(1\)VBKL uses the corresponding Dirac specialization\. For a single supportuu, the recursive kernel is
ku\(𝐱,𝐱′\)\\displaystyle k\_\{u\}\\left\(\\mathbf\{x\},\\mathbf\{x\}^\{\\prime\}\\right\)=k\(B\)\(u\(𝐱\),u\(𝐱′\)\),𝐱,𝐱′∈𝒳\.\\displaystyle=k^\{\(\\mathrm\{B\}\)\}\\left\(u\(\\mathbf\{x\}\),u\(\\mathbf\{x\}^\{\\prime\}\)\\right\),\\qquad\\mathbf\{x\},\\mathbf\{x\}^\{\\prime\}\\in\\mathcal\{X\}\.\(2\)Thus each recursive step propagates one support rather than averaging over a support family\. The recursion below constructs only the nonlinear dictionary; the signed\-measure variation hull is introduced after the target depth is reached\.
##### First layer\.
LetΩ⊆𝕊d−1\\Omega\\subseteq\\mathbb\{S\}^\{d\-1\}denote the admissible set of directions\. The first recursive generator class consists of linear projections,𝒰1:=\{𝐱↦𝝎⊤𝐱:𝝎∈Ω\}\.\\mathcal\{U\}\_\{1\}:\\allowbreak=\\left\\\{\\mathbf\{x\}\\mapsto\\bm\{\\omega\}^\{\\top\}\\mathbf\{x\}:\\bm\{\\omega\}\\in\\Omega\\right\\\}\.
##### Atomic BKL construction\.
For eachl∈\{2,…,L\}l\\in\\left\\\{2,\\ldots,L\\right\\\}, assume that the lower\-level dictionary𝒰l−1\\mathcal\{U\}\_\{l\-1\}has been defined\. The next dictionary is obtained by composing a normalized Brownian profile with a lower\-level atom\. The exact pullback\-RKHS identification is established in[Lemma18](https://arxiv.org/html/2608.13882#Thmtheorem18), and gives
𝒰l\\displaystyle\\mathcal\{U\}\_\{l\}:=\{𝐱↦g\(u\(𝐱\)\):u∈𝒰l−1,g∈ℋk\(B\),‖g‖ℋk\(B\)≤1\}\\displaystyle:=\\left\\\{\\mathbf\{x\}\\mapsto g\\left\(u\\left\(\\mathbf\{x\}\\right\)\\right\):u\\in\\mathcal\{U\}\_\{l\-1\},\\;g\\in\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\},\\;\\left\\\|g\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}\\leq 1\\right\\\}=⋃u∈𝒰l−1\{f∈ℋku:‖f‖ℋku≤1\}\.\\displaystyle=\\bigcup\_\{u\\in\\mathcal\{U\}\_\{l\-1\}\}\\left\\\{f\\in\\mathcal\{H\}\_\{k\_\{u\}\}:\\left\\\|f\\right\\\|\_\{\\mathcal\{H\}\_\{k\_\{u\}\}\}\\leq 1\\right\\\}\.\(3\)Consequently, everyf∈𝒰Lf\\in\\mathcal\{U\}\_\{L\}admits a representationf\(𝐱\)=\(gL−1∘⋯∘g1\)\(𝝎⊤𝐱\),f\\left\(\\mathbf\{x\}\\right\)=\\left\(g\_\{L\-1\}\\circ\\cdots\\circ g\_\{1\}\\right\)\\left\(\\bm\{\\omega\}^\{\\top\}\\mathbf\{x\}\\right\),𝝎∈Ω\\bm\{\\omega\}\\in\\Omega,‖gj‖ℋk\(B\)≤1\\left\\\|g\_\{j\}\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}\\leq 1,j∈\[L−1\]\.j\\in\\left\[L\-1\\right\]\.Thus, a depth\-LLatom consists of one linear projection followed by exactlyL−1L\-1normalized Brownian profiles\. The collection of all such atoms forms the recursive Brownian dictionary𝒰L\\mathcal\{U\}\_\{L\}used in the variation\-space construction below\.
##### VBKL space\.
Each𝒰l\\mathcal\{U\}\_\{l\}is viewed as a subspace ofC\(𝒳\)C\(\\mathcal\{X\}\)and is equipped with the Borelσ\\sigma\-algebra induced by the supremum norm\. The completed dictionary𝒰L\\mathcal\{U\}\_\{L\}generates the depth\-LLVBKL space
𝒱\(L\):=\{F∈L2\(ν\):F=∫𝒰Ludμ\(u\),μ∈ℳ\(𝒰L\),∥μ∥TV<∞\}\.\\displaystyle\\mathcal\{V\}^\{\(L\)\}:=\\left\\\{F\\in L^\{2\}\(\\nu\):F=\\int\_\{\\mathcal\{U\}\_\{L\}\}u\\,\\mathrm\{d\}\\mu\(u\),\\;\\mu\\in\\mathcal\{M\}\(\\mathcal\{U\}\_\{L\}\),\\;\\\|\\mu\\\|\_\{\\mathrm\{TV\}\}<\\infty\\right\\\}\.\(7\)Hereℳ\(𝒰L\)\\mathcal\{M\}\(\\mathcal\{U\}\_\{L\}\)is the space of finite signed Borel measures on𝒰L\\mathcal\{U\}\_\{L\}, and the integral is a Bochner integral inL2\(ν\)L^\{2\}\(\\nu\)\. Every absolutely summable atomic expansion corresponds to a discrete measure\. When𝒰L\\mathcal\{U\}\_\{L\}is compact inL2\(ν\)L^\{2\}\(\\nu\), the canonical mapu↦uu\\mapsto uis continuous, and a uniformly∥⋅∥TV\\\|\\cdot\\\|\_\{\\mathrm\{TV\}\}\-bounded sequence of representing measures has a weak\-\* convergent subsequence inℳ\(𝒰L\)=C\(𝒰L\)∗\\mathcal\{M\}\(\\mathcal\{U\}\_\{L\}\)=C\(\\mathcal\{U\}\_\{L\}\)^\{\*\}\. Under these conditions,[Lemma27](https://arxiv.org/html/2608.13882#Thmtheorem27)shows that the limiting measure retains the barycentric representation\. Unlike BKL, which averages pullback kernels recursively, VBKL fixes the completed path dictionary and applies signed\-measure superposition only at the outer level\.
##### Variation complexity\.
ForF∈𝒱\(L\)F\\in\\mathcal\{V\}^\{\(L\)\}, define the intrinsic variation complexity as the infimal total variation of a representing measure:
𝒞^var\(L\)\(F\):=inf\{‖μ‖TV:F=∫𝒰Lu𝑑μ\(u\)\}\.\\displaystyle\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}\(F\):=\\inf\\left\\\{\\\|\\mu\\\|\_\{\\mathrm\{TV\}\}:F=\\int\_\{\\mathcal\{U\}\_\{L\}\}u\\,\\mathrm\{d\}\\mu\(u\)\\right\\\}\.\(8\)The infimum is over finite signed Borel measuresμ∈ℳ\(𝒰L\)\\mu\\in\\mathcal\{M\}\(\\mathcal\{U\}\_\{L\}\)representingFFthrough \([7](https://arxiv.org/html/2608.13882#S2.E7)\)\. This formulation is convenient for both the statistical analysis and the approximation results, where finite\-support measures arise as approximants of general representations\.
##### Finite\-support VBKL class\.
Form≥1m\\geq 1, let
𝒱m\(L\):=\{F∈𝒱\(L\):F=∫𝒰Ludμ\(u\),\|supp\(μ\)\|≤m\},\\displaystyle\\mathcal\{V\}\_\{m\}^\{\(L\)\}:=\\left\\\{F\\in\\mathcal\{V\}^\{\(L\)\}:F=\\int\_\{\\mathcal\{U\}\_\{L\}\}u\\,\\mathrm\{d\}\\mu\(u\),\\;\\left\|\\operatorname\{supp\}\(\\mu\)\\right\|\\leq m\\right\\\},\(9\)whereμ\\muis a finite signed Borel measure on𝒰L\\mathcal\{U\}\_\{L\}\. Equivalently,𝒱m\(L\)\\mathcal\{V\}\_\{m\}^\{\(L\)\}contains precisely the functions admitting a representation by at mostmmatoms from𝒰L\\mathcal\{U\}\_\{L\}\.
### 2\.2Finite Lower\-Support Architectures
The recursive dictionaries introduced in[Section2\.1](https://arxiv.org/html/2608.13882#S2.SS1)are infinite\-dimensional because their Brownian profiles belong toℋk\(B\)\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\. For the architecture\-dependent statistical analysis, we therefore introduce an explicit finite\-parametric class of lower\-level support functions\. This class is defined separately from the infinite\-dimensional dictionary𝒰L−1\\mathcal\{U\}\_\{L\-1\}; the relationship between the two classes is stated at the end of this subsection\. Fixm≥1m\\geq 1andG≥2G\\geq 2\. Set
R𝒳\\displaystyle R\_\{\\mathcal\{X\}\}:=sup𝐱∈𝒳‖𝐱‖2,\\displaystyle:=\\sup\_\{\\mathbf\{x\}\\in\\mathcal\{X\}\}\\left\\\|\\mathbf\{x\}\\right\\\|\_\{2\},A𝒳\\displaystyle A\_\{\\mathcal\{X\}\}:=max\{1,R𝒳\}\.\\displaystyle:=\\max\\left\\\{1,R\_\{\\mathcal\{X\}\}\\right\\\}\.\(11\)Since𝒳\\mathcal\{X\}is compact, one hasA𝒳<∞A\_\{\\mathcal\{X\}\}<\\infty\.
##### Finite Brownian profiles\.
Let
−A𝒳=t0<t1<⋯<tj0=0<⋯<tG=A𝒳\\displaystyle\-A\_\{\\mathcal\{X\}\}=t\_\{0\}<t\_\{1\}<\\cdots<t\_\{j\_\{0\}\}=0<\\cdots<t\_\{G\}=A\_\{\\mathcal\{X\}\}\(12\)be a fixed interpolation grid containing the origin, and letψ0,…,ψG\\psi\_\{0\},\\ldots,\\psi\_\{G\}denote the associated continuous piecewise\-linear hat functions on\[−A𝒳,A𝒳\]\\left\[\-A\_\{\\mathcal\{X\}\},A\_\{\\mathcal\{X\}\}\\right\]\. Forg∈ℋk\(B\)g\\in\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}, define
\(ΠGg\)\(t\):=∑j=0Gg\(tj\)ψj\(t\),t∈\[−A𝒳,A𝒳\]\.\\displaystyle\\left\(\\Pi\_\{G\}g\\right\)\\left\(t\\right\):=\\sum\_\{j=0\}^\{G\}g\\left\(t\_\{j\}\\right\)\\psi\_\{j\}\\left\(t\\right\),\\qquad t\\in\\left\[\-A\_\{\\mathcal\{X\}\},A\_\{\\mathcal\{X\}\}\\right\]\.\(13\)Equivalently, for everyj∈\{0,…,G−1\}j\\in\\left\\\{0,\\ldots,G\-1\\right\\\}and everyt∈\[tj,tj\+1\]t\\in\\left\[t\_\{j\},t\_\{j\+1\}\\right\],
\(ΠGg\)\(t\)\\displaystyle\\left\(\\Pi\_\{G\}g\\right\)\\left\(t\\right\)=tj\+1−ttj\+1−tjg\(tj\)\+t−tjtj\+1−tjg\(tj\+1\)\.\\displaystyle=\\frac\{t\_\{j\+1\}\-t\}\{t\_\{j\+1\}\-t\_\{j\}\}g\\left\(t\_\{j\}\\right\)\+\\frac\{t\-t\_\{j\}\}\{t\_\{j\+1\}\-t\_\{j\}\}g\\left\(t\_\{j\+1\}\\right\)\.\(14\)We extendΠGg\\Pi\_\{G\}gconstantly outside the interpolation interval by setting
\(ΠGg\)\(t\)\\displaystyle\\left\(\\Pi\_\{G\}g\\right\)\\left\(t\\right\):=\(ΠGg\)\(−A𝒳\),t<−A𝒳,\\displaystyle:=\\left\(\\Pi\_\{G\}g\\right\)\\left\(\-A\_\{\\mathcal\{X\}\}\\right\),\\qquad t<\-A\_\{\\mathcal\{X\}\},\(15\)\(ΠGg\)\(t\)\\displaystyle\\left\(\\Pi\_\{G\}g\\right\)\\left\(t\\right\):=\(ΠGg\)\(A𝒳\),t\>A𝒳\.\\displaystyle:=\\left\(\\Pi\_\{G\}g\\right\)\\left\(A\_\{\\mathcal\{X\}\}\\right\),\\qquad t\>A\_\{\\mathcal\{X\}\}\.\(16\)Define the finite Brownian profile class by
𝒬G,A𝒳:=\{ΠGg:g∈ℋk\(B\),‖g‖ℋk\(B\)≤1\}\.\\displaystyle\\mathcal\{Q\}\_\{G,A\_\{\\mathcal\{X\}\}\}:=\\left\\\{\\Pi\_\{G\}g:g\\in\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\},\\;\\left\\\|g\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}\\leq 1\\right\\\}\.\(17\)Everyq∈𝒬G,A𝒳q\\in\\mathcal\{Q\}\_\{G,A\_\{\\mathcal\{X\}\}\}is determined by its nodal coefficient vector𝐜q:=\(q\(t0\),…,q\(tG\)\)∈ℝG\+1\\mathbf\{c\}\_\{q\}:=\\left\(q\\left\(t\_\{0\}\\right\),\\ldots,q\\left\(t\_\{G\}\\right\)\\right\)\\in\\mathbb\{R\}^\{G\+1\}\. Moreover, the nodal coefficients satisfyq\(0\)=0q\\left\(0\\right\)=0,
‖q‖ℋk\(B\)2\\displaystyle\\left\\\|q\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}^\{2\}=∑j=0G−1\(q\(tj\+1\)−q\(tj\)\)2tj\+1−tj≤1\.\\displaystyle=\\sum\_\{j=0\}^\{G\-1\}\\frac\{\\left\(q\\left\(t\_\{j\+1\}\\right\)\-q\\left\(t\_\{j\}\\right\)\\right\)^\{2\}\}\{t\_\{j\+1\}\-t\_\{j\}\}\\leq 1\.\(18\)The exact correspondence between \([17](https://arxiv.org/html/2608.13882#S2.E17)\) and \([18](https://arxiv.org/html/2608.13882#S2.E18)\) is established in the appendix\.
##### Admissible mixing coefficients\.
Let
𝔐m:=\{𝐖∈ℝm×m:max∑r=1mj∈\[m\]\|Wjr\|≤1\},\\displaystyle\\mathfrak\{M\}\_\{m\}:=\\left\\\{\\mathbf\{W\}\\in\\mathbb\{R\}^\{m\\times m\}:\\max\_\{j\\in\\left\[m\\right\]\}\\sum\_\{r=1\}^\{m\}\\left\|W\_\{jr\}\\right\|\\leq 1\\right\\\},\(19\)and let
𝔅m:=\{β∈ℝm:∑j=1m\|βj\|≤1\}\.\\displaystyle\\mathfrak\{B\}\_\{m\}:=\\left\\\{\\mathbf\{\\beta\}\\in\\mathbb\{R\}^\{m\}:\\sum\_\{j=1\}^\{m\}\\left\|\\beta\_\{j\}\\right\|\\leq 1\\right\\\}\.\(20\)The row\-sum constraint in \([19](https://arxiv.org/html/2608.13882#S2.E19)\) and the readout constraint in \([20](https://arxiv.org/html/2608.13882#S2.E20)\) ensure that the intermediate supports remain uniformly bounded\.
##### Finite lower\-support architecture\.
ForL=2L=2, define
𝒜1m,G:=𝒰1\.\\displaystyle\\mathcal\{A\}\_\{1\}^\{m,G\}:=\\mathcal\{U\}\_\{1\}\.\(21\)Suppose now thatL≥3L\\geq 3, and setK:=L−2K:=L\-2\. An admissible parameter tuple is of the form
𝜽:=\(\(𝝎j\)j∈\[m\],\(qj\(r\)\)r∈\[K\],j∈\[m\],\(𝐖\(r\)\)r=2K,β\),\\displaystyle\\bm\{\\theta\}:=\\left\(\\left\(\\bm\{\\omega\}\_\{j\}\\right\)\_\{j\\in\\left\[m\\right\]\},\\left\(q\_\{j\}^\{\(r\)\}\\right\)\_\{r\\in\\left\[K\\right\],\\,j\\in\\left\[m\\right\]\},\\left\(\\mathbf\{W\}^\{\(r\)\}\\right\)\_\{r=2\}^\{K\},\\mathbf\{\\beta\}\\right\),\(22\)where
𝝎j\\displaystyle\\bm\{\\omega\}\_\{j\}∈Ω,j∈\[m\],\\displaystyle\\in\\Omega,\\qquad j\\in\\left\[m\\right\],\(23\)qj\(r\)\\displaystyle q\_\{j\}^\{\(r\)\}∈𝒬G,A𝒳,r∈\[K\],j∈\[m\],\\displaystyle\\in\\mathcal\{Q\}\_\{G,A\_\{\\mathcal\{X\}\}\},\\qquad r\\in\\left\[K\\right\],\\quad j\\in\\left\[m\\right\],\(24\)𝐖\(r\)\\displaystyle\\mathbf\{W\}^\{\(r\)\}∈𝔐m,r∈\{2,…,K\},\\displaystyle\\in\\mathfrak\{M\}\_\{m\},\\qquad r\\in\\left\\\{2,\\ldots,K\\right\\\},\(25\)β\\displaystyle\\mathbf\{\\beta\}∈𝔅m\.\\displaystyle\\in\\mathfrak\{B\}\_\{m\}\.\(26\)WhenK=1K=1, the family of mixing matrices in \([22](https://arxiv.org/html/2608.13882#S2.E22)\) is empty\. For everyj∈\[m\]j\\in\\left\[m\\right\], define the first linear coordinates by
zj\(0\)\(𝐱\):=𝝎j⊤𝐱,𝐱∈𝒳\.\\displaystyle z\_\{j\}^\{\(0\)\}\\left\(\\mathbf\{x\}\\right\):=\\bm\{\\omega\}\_\{j\}^\{\\top\}\\mathbf\{x\},\\qquad\\mathbf\{x\}\\in\\mathcal\{X\}\.\(27\)The first Brownian profile layer is
zj\(1\)\(𝐱\):=qj\(1\)\(zj\(0\)\(𝐱\)\),j∈\[m\]\.\\displaystyle z\_\{j\}^\{\(1\)\}\\left\(\\mathbf\{x\}\\right\):=q\_\{j\}^\{\(1\)\}\\left\(z\_\{j\}^\{\(0\)\}\\left\(\\mathbf\{x\}\\right\)\\right\),\\qquad j\\in\\left\[m\\right\]\.\(28\)For everyr∈\{2,…,K\}r\\in\\left\\\{2,\\ldots,K\\right\\\}andj∈\[m\]j\\in\\left\[m\\right\], define recursively
sj\(r\)\(𝐱\)\\displaystyle s\_\{j\}^\{\(r\)\}\\left\(\\mathbf\{x\}\\right\):=∑k=1mWjk\(r\)zk\(r−1\)\(𝐱\),\\displaystyle:=\\sum\_\{k=1\}^\{m\}W\_\{jk\}^\{\(r\)\}z\_\{k\}^\{\(r\-1\)\}\\left\(\\mathbf\{x\}\\right\),\(29\)zj\(r\)\(𝐱\)\\displaystyle z\_\{j\}^\{\(r\)\}\\left\(\\mathbf\{x\}\\right\):=qj\(r\)\(sj\(r\)\(𝐱\)\)\.\\displaystyle:=q\_\{j\}^\{\(r\)\}\\left\(s\_\{j\}^\{\(r\)\}\\left\(\\mathbf\{x\}\\right\)\\right\)\.\(30\)The scalar lower\-support output associated with𝜽\\bm\{\\theta\}is
a𝜽\(𝐱\):=∑j=1mβjzj\(K\)\(𝐱\),𝐱∈𝒳\.\\displaystyle a\_\{\\bm\{\\theta\}\}\\left\(\\mathbf\{x\}\\right\):=\\sum\_\{j=1\}^\{m\}\\beta\_\{j\}z\_\{j\}^\{\(K\)\}\\left\(\\mathbf\{x\}\\right\),\\qquad\\mathbf\{x\}\\in\\mathcal\{X\}\.\(31\)The resulting depth\-\(L−1\)\\left\(L\-1\\right\)finite lower\-support class is
𝒜L−1m,G\\displaystyle\\mathcal\{A\}\_\{L\-1\}^\{m,G\}:=\{a𝜽:𝜽satisfies\([23](https://arxiv.org/html/2608.13882#S2.E23)\)–\([26](https://arxiv.org/html/2608.13882#S2.E26)\)\}\.\\displaystyle:=\\left\\\{a\_\{\\bm\{\\theta\}\}:\\bm\{\\theta\}\\text\{ satisfies \}\\eqref\{eq:finite\-lower\-support\-direction\-constraint\}\\text\{\-\-\}\\eqref\{eq:finite\-lower\-support\-readout\-constraint\}\\right\\\}\.\(32\)By[Lemma17](https://arxiv.org/html/2608.13882#Thmtheorem17)[2](https://arxiv.org/html/2608.13882#A4.I1.i2), the normalization constraints yield the sharper uniform range estimate
supa∈𝒜L−1m,Gsup𝐱∈𝒳\|a\(𝐱\)\|≤A𝒳2−\(L−2\)\.\\displaystyle\\sup\_\{a\\in\\mathcal\{A\}\_\{L\-1\}^\{m,G\}\}\\sup\_\{\\mathbf\{x\}\\in\\mathcal\{X\}\}\\left\|a\\left\(\\mathbf\{x\}\\right\)\\right\|\\leq A\_\{\\mathcal\{X\}\}^\{\\,2^\{\-\(L\-2\)\}\}\.\(33\)By[Lemma17](https://arxiv.org/html/2608.13882#Thmtheorem17)[3](https://arxiv.org/html/2608.13882#A4.I1.i3), forL≥3L\\geq 3, every element of𝒜L−1m,G\\mathcal\{A\}\_\{L\-1\}^\{m,G\}is represented by at most
PL−1,m,G\\displaystyle P\_\{L\-1,m,G\}:=md\+\(L−2\)m\(G\+1\)\+\(L−3\)m2\+m\\displaystyle:=md\+\\left\(L\-2\\right\)m\\left\(G\+1\\right\)\+\\left\(L\-3\\right\)m^\{2\}\+m\(34\)real parameters\. The four terms in \([34](https://arxiv.org/html/2608.13882#S2.E34)\) correspond, respectively, to the first\-layer directions, the nodal profile coefficients, the internal mixing matrices, and the final readout coefficients\. ForL=2L=2, we set
P1,m,G:=d\.\\displaystyle P\_\{1,m,G\}:=d\.\(35\)
##### Associated outer Brownian dictionary\.
For everya∈𝒜L−1m,Ga\\in\\mathcal\{A\}\_\{L\-1\}^\{m,G\}, define
ka\(𝐱,𝐱′\):=k\(B\)\(a\(𝐱\),a\(𝐱′\)\),𝐱,𝐱′∈𝒳\.\\displaystyle k\_\{a\}\\left\(\\mathbf\{x\},\\mathbf\{x\}^\{\\prime\}\\right\):=k^\{\(\\mathrm\{B\}\)\}\\left\(a\\left\(\\mathbf\{x\}\\right\),a\\left\(\\mathbf\{x\}^\{\\prime\}\\right\)\\right\),\\qquad\\mathbf\{x\},\\mathbf\{x\}^\{\\prime\}\\in\\mathcal\{X\}\.\(36\)The associated outer Brownian dictionary is
𝒟Lm,G:=⋃a∈𝒜L−1m,G\{f∈ℋka:‖f‖ℋka≤1\}\.\\displaystyle\\mathcal\{D\}\_\{L\}^\{m,G\}:=\\bigcup\_\{a\\in\\mathcal\{A\}\_\{L\-1\}^\{m,G\}\}\\left\\\{f\\in\\mathcal\{H\}\_\{k\_\{a\}\}:\\left\\\|f\\right\\\|\_\{\\mathcal\{H\}\_\{k\_\{a\}\}\}\\leq 1\\right\\\}\.\(37\)Finally, forR\>0R\>0, define the corresponding finite\-architecture Brownian variation ball by
𝒲RL;m,G:=\{Fμ:Fμ\(𝐱\)=∫𝒟Lm,Gf\(𝐱\)dμ\(f\),μ∈ℳ\(𝒟Lm,G\),‖μ‖TV≤R\}\.\\displaystyle\\mathcal\{W\}\_\{R\}^\{L;m,G\}:=\\left\\\{F\_\{\\mu\}:F\_\{\\mu\}\\left\(\\mathbf\{x\}\\right\)=\\int\_\{\\mathcal\{D\}\_\{L\}^\{m,G\}\}f\\left\(\\mathbf\{x\}\\right\)\\,\\mathrm\{d\}\\mu\\left\(f\\right\),\\;\\mu\\in\\mathcal\{M\}\\left\(\\mathcal\{D\}\_\{L\}^\{m,G\}\\right\),\\;\\left\\\|\\mu\\right\\\|\_\{\\mathrm\{TV\}\}\\leq R\\right\\\}\.\(38\)By \([33](https://arxiv.org/html/2608.13882#S2.E33)\) and the reproducing property, everyf∈𝒟Lm,Gf\\in\\mathcal\{D\}\_\{L\}^\{m,G\}satisfies\|f\(𝐱\)\|≤A𝒳1/2,𝐱∈𝒳\.\\left\|f\\left\(\\mathbf\{x\}\\right\)\\right\|\\leq A\_\{\\mathcal\{X\}\}^\{1/2\},\\qquad\\mathbf\{x\}\\in\\mathcal\{X\}\.Consequently, the measure representation in \([38](https://arxiv.org/html/2608.13882#S2.E38)\) is well\-defined pointwise and inL2\(ν\)L^\{2\}\\left\(\\nu\\right\)\.
## 3Analytical and Statistical Properties of the VBKL spaces
This section develops two complementary parts of the theory\. We first establish the fundamental analytical properties of the full infinite\-dimensional VBKL spaces, showing that variation complexity controls regularity, pointwise evaluation, and expressive power across recursive depth\. We then study the statistical complexity of the associated finite\-architecture Brownian variation classes introduced in[Section2\.2](https://arxiv.org/html/2608.13882#S2.SS2)\. The resulting Rademacher bound explicitly separates the outer variation radius from the parameter complexity of the finite lower\-support architecture\.
### 3\.1Analytical Properties of VBKL spaces
We first ask whether the outer variation complexity controls regularity and pointwise evaluation, and whether the recursive path construction produces a genuine depth hierarchy\. The following theorem answers these questions for the full VBKL spaces\.
###### Theorem 4\.
\(Analytical properties of the VBKL spaces\)Let\(𝒳,ν\)\(\\mathcal\{X\},\\nu\),Ω\\Omega, and the recursive dictionaries\(𝒰l\)l≥1\(\\mathcal\{U\}\_\{l\}\)\_\{l\\geq 1\}be as introduced in[Section2\.1](https://arxiv.org/html/2608.13882#S2.SS1)\. Assume that𝒳⊆ℝd\\mathcal\{X\}\\subseteq\\mathbb\{R\}^\{d\}is compact, and letL≥2L\\geq 2\. Let𝒱\(L\)\\mathcal\{V\}^\{\(L\)\}and𝒞^var\(L\)\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}denote the depth\-LLVBKL space and variation complexity defined in \([7](https://arxiv.org/html/2608.13882#S2.E7)\) and \([8](https://arxiv.org/html/2608.13882#S2.E8)\)\. Set
αL\\displaystyle\\alpha\_\{L\}:=2−\(L−1\),\\displaystyle:=2^\{\-\(L\-1\)\},R𝒳\\displaystyle R\_\{\\mathcal\{X\}\}:=sup𝐱∈𝒳‖𝐱‖2\.\\displaystyle:=\\sup\_\{\\mathbf\{x\}\\in\\mathcal\{X\}\}\\\|\\mathbf\{x\}\\\|\_\{2\}\.\(39\)Then the following statements hold\.
1. 1\.Hölder representative and pointwise control\.EveryF∈𝒱\(L\)F\\in\\mathcal\{V\}^\{\(L\)\}admits a representative, still denoted byFF, satisfying \|F\(𝐱\)−F\(𝐱′\)\|\\displaystyle\\left\|F\(\\mathbf\{x\}\)\-F\(\\mathbf\{x\}^\{\\prime\}\)\\right\|≤‖𝐱−𝐱′‖2αL𝒞^var\(L\)\(F\),𝐱,𝐱′∈𝒳,\\displaystyle\\leq\\\|\\mathbf\{x\}\-\\mathbf\{x\}^\{\\prime\}\\\|\_\{2\}^\{\\,\\alpha\_\{L\}\}\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}\(F\),\\qquad\\mathbf\{x\},\\mathbf\{x\}^\{\\prime\}\\in\\mathcal\{X\},\(40\)\|F\(𝐱\)\|\\displaystyle\\left\|F\(\\mathbf\{x\}\)\\right\|≤‖𝐱‖2αL𝒞^var\(L\)\(F\),𝐱∈𝒳\.\\displaystyle\\leq\\\|\\mathbf\{x\}\\\|\_\{2\}^\{\\,\\alpha\_\{L\}\}\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}\(F\),\\qquad\\mathbf\{x\}\\in\\mathcal\{X\}\.\(41\)
2. 2\.Hölder\-space control\.Every representative supplied by part[1](https://arxiv.org/html/2608.13882#S3.I1.i1)belongs toC0,αL\(𝒳\)C^\{0,\\alpha\_\{L\}\}\(\\mathcal\{X\}\)and satisfies ‖F‖C0,αL\(𝒳\)≤\(R𝒳αL\+1\)𝒞^var\(L\)\(F\)\.\\displaystyle\\\|F\\\|\_\{C^\{0,\\alpha\_\{L\}\}\(\\mathcal\{X\}\)\}\\leq\\left\(R\_\{\\mathcal\{X\}\}^\{\\,\\alpha\_\{L\}\}\+1\\right\)\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}\(F\)\.\(42\)Ifsupp\(ν\)=𝒳\\operatorname\{supp\}\(\\nu\)=\\mathcal\{X\}, this representative is unique, and the representative map is a continuous linear injection 𝒱\(L\)↪C0,αL\(𝒳\)\.\\displaystyle\\mathcal\{V\}^\{\(L\)\}\\hookrightarrow C^\{0,\\alpha\_\{L\}\}\(\\mathcal\{X\}\)\.\(43\)
3. 3\.Depth monotonicity and strictness\.The consecutive spaces satisfy 𝒱\(L\)⊆𝒱\(L\+1\)\.\\displaystyle\\mathcal\{V\}^\{\(L\)\}\\subseteq\\mathcal\{V\}^\{\(L\+1\)\}\.\(44\)Suppose, in addition, that there exist𝐱0∈𝒳\\mathbf\{x\}\_\{0\}\\in\\mathcal\{X\},ρ\>0\\rho\>0, constants0<c1≤C1<∞0<c\_\{1\}\\leq C\_\{1\}<\\infty, a Lipschitz mapγ:\[0,ρ\]→𝒳\\gamma:\[0,\\rho\]\\to\\mathcal\{X\}, and a first\-layer generatoru1∈𝒰1u\_\{1\}\\in\\mathcal\{U\}\_\{1\}such that γ\(0\)\\displaystyle\\gamma\(0\)=𝐱0,\\displaystyle=\\mathbf\{x\}\_\{0\},γ\(\[0,ρ\]\)\\displaystyle\\gamma\(\[0,\\rho\]\)⊆supp\(ν\),\\displaystyle\\subseteq\\operatorname\{supp\}\(\\nu\),\(45\)u1\(𝐱0\)\\displaystyle u\_\{1\}\(\\mathbf\{x\}\_\{0\}\)=0,\\displaystyle=0,c1t\\displaystyle c\_\{1\}t≤u1\(γ\(t\)\)≤C1t,t∈\[0,ρ\]\.\\displaystyle\\leq u\_\{1\}\(\\gamma\(t\)\)\\leq C\_\{1\}t,\\qquad t\\in\[0,\\rho\]\.\(46\)Then the inclusion in \([44](https://arxiv.org/html/2608.13882#S3.E44)\) is strict: 𝒱\(L\)⊊𝒱\(L\+1\)\.\\displaystyle\\mathcal\{V\}^\{\(L\)\}\\subsetneq\\mathcal\{V\}^\{\(L\+1\)\}\.\(47\)
The theorem supplies the regularity and depth properties used by the statistical and approximation results below\.
### 3\.2Statistical Properties of Finite\-Architecture Brownian Variation Classes
The analytical results above concern the full infinite\-dimensional VBKL space\. The statistical analysis in this subsection instead concerns the associated finite\-architecture class𝒲RL;m,G\\mathcal\{W\}\_\{R\}^\{L;m,G\}defined in \([38](https://arxiv.org/html/2608.13882#S2.E38)\)\. Its lower\-level supports belong to the finite\-parametric architecture𝒜L−1m,G\\mathcal\{A\}\_\{L\-1\}^\{m,G\}, while its outer atoms range over the corresponding union of Brownian pullback RKHS unit balls𝒟Lm,G\\mathcal\{D\}\_\{L\}^\{m,G\}\. The proof proceeds by reducing the variation class to its outer Brownian dictionary, controlling the resulting union of RKHS unit balls through Brownian quadratic chaos, and estimating that chaos by the signed threshold entropy of the finite lower\-support architecture\. All supporting arguments and the complete proof are deferred to the appendix\.
#### 3\.2\.1Statistical Learning Framework
Let\(X,Y\)∼νXY\(X,Y\)\\sim\\nu\_\{XY\}be a random input–output pair on𝒳×𝒴\\mathcal\{X\}\\times\\mathcal\{Y\}, and let
𝒟n:=\{\(𝐱i,yi\)\}i=1n\\displaystyle\\mathcal\{D\}\_\{n\}:=\\left\\\{\(\\mathbf\{x\}\_\{i\},y\_\{i\}\)\\right\\\}\_\{i=1\}^\{n\}\(48\)be an independent sample drawn fromνXY\\nu\_\{XY\}\. We denote the input marginal byνX\\nu\_\{X\}and identify it with the measureν\\nuintroduced in[Section2](https://arxiv.org/html/2608.13882#S2)\. For a measurable predictorF:𝒳→ℝF:\\mathcal\{X\}\\rightarrow\\mathbb\{R\}and a lossℓ:ℝ×𝒴→ℝ\\ell:\\mathbb\{R\}\\times\\mathcal\{Y\}\\rightarrow\\mathbb\{R\}, define
ℛ\(F\)\\displaystyle\\mathcal\{R\}\(F\):=𝔼\(X,Y\)∼νXY\[ℓ\(F\(X\),Y\)\],\\displaystyle:=\\mathbb\{E\}\_\{\(X,Y\)\\sim\\nu\_\{XY\}\}\\left\[\\ell\(F\(X\),Y\)\\right\],ℛn\(F\)\\displaystyle\\mathcal\{R\}\_\{n\}\(F\):=1n∑i=1nℓ\(F\(𝐱i\),yi\)\.\\displaystyle:=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\ell\(F\(\\mathbf\{x\}\_\{i\}\),y\_\{i\}\)\.\(49\)We work with bounded variation\-complexity classes, which play the role of norm balls in classical RKHS learning\.
##### Finite\-architecture hypothesis classes\.
FixL≥2L\\geq 2,m≥1m\\geq 1,G≥2G\\geq 2, andR\>0R\>0\. The hypothesis class considered in this subsection is the finite\-architecture Brownian variation ball𝒲RL;m,G\\mathcal\{W\}\_\{R\}^\{L;m,G\}defined in \([38](https://arxiv.org/html/2608.13882#S2.E38)\)\. The radiusRRcontrols the total variation of the outer signed\-measure representation, whereasPL−1,m,GP\_\{L\-1,m,G\}controls the number of real parameters in the finite lower\-support architecture\.
##### Empirical and expected Rademacher complexities\.
Letℱ\\mathcal\{F\}be a class of real\-valued functions on𝒳\\mathcal\{X\}, and letε1,…,εn\\varepsilon\_\{1\},\\ldots,\\varepsilon\_\{n\}be independent Rademacher random variables\. For fixed sample points𝐱1,…,𝐱n∈𝒳\\mathbf\{x\}\_\{1\},\\ldots,\\mathbf\{x\}\_\{n\}\\in\\mathcal\{X\}, define
ℜ^n\(ℱ\):=𝔼ε\[supf∈ℱ1n∑i=1nεif\(𝐱i\)\]\.\\displaystyle\\widehat\{\\mathfrak\{R\}\}\_\{n\}\\left\(\\mathcal\{F\}\\right\):=\\mathbb\{E\}\_\{\\varepsilon\}\\left\[\\sup\_\{f\\in\\mathcal\{F\}\}\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\varepsilon\_\{i\}f\\left\(\\mathbf\{x\}\_\{i\}\\right\)\\right\]\.\(50\)Averaging over independent sample pointsX1,…,Xn∼νX\_\{1\},\\ldots,X\_\{n\}\\sim\\nugives the expected Rademacher complexity
ℜn\(ℱ\):=𝔼𝐗\[ℜ^n\(ℱ\)\]\.\\displaystyle\\mathfrak\{R\}\_\{n\}\\left\(\\mathcal\{F\}\\right\):=\\mathbb\{E\}\_\{\\mathbf\{X\}\}\\left\[\\widehat\{\\mathfrak\{R\}\}\_\{n\}\\left\(\\mathcal\{F\}\\right\)\\right\]\.\(51\)
###### Theorem 6\.
\(Architecture\-dependent Rademacher bound for finite Brownian variation classes\)LetL,m,G,n∈ℕL,m,G,n\\in\\mathbb\{N\}satisfyL≥2L\\geq 2,m≥1m\\geq 1,G≥2G\\geq 2, andn≥1n\\geq 1, and letR\>0R\>0\. Let𝒜L−1m,G\\mathcal\{A\}\_\{L\-1\}^\{m,G\},𝒟Lm,G\\mathcal\{D\}\_\{L\}^\{m,G\}, and𝒲RL;m,G\\mathcal\{W\}\_\{R\}^\{L;m,G\}denote, respectively, the finite lower\-support class, its associated outer Brownian dictionary, and the radius\-RRfinite\-architecture Brownian variation ball introduced in[Section2\.2](https://arxiv.org/html/2608.13882#S2.SS2)\. LetPL−1,m,GP\_\{L\-1,m,G\}denote the parameter\-count bound associated with𝒜L−1m,G\\mathcal\{A\}\_\{L\-1\}^\{m,G\}, as specified in[Section2\.2](https://arxiv.org/html/2608.13882#S2.SS2)and established in[Lemma17](https://arxiv.org/html/2608.13882#Thmtheorem17), part[3](https://arxiv.org/html/2608.13882#A4.I1.i3)\. For notational convenience, set
P\\displaystyle P:=PL−1,m,G,\\displaystyle:=P\_\{L\-1,m,G\},\(52\)ΓL−1,m,G,n\\displaystyle\\Gamma\_\{L\-1,m,G,n\}:=1\+\(P\+1\)ln\(e\(P\+1\)\)ln\(en\)\.\\displaystyle:=1\+\\left\(P\+1\\right\)\\ln\\left\(e\\left\(P\+1\\right\)\\right\)\\ln\\left\(en\\right\)\.\(53\)For fixed sample points𝐱1,…,𝐱n∈𝒳\\mathbf\{x\}\_\{1\},\\ldots,\\mathbf\{x\}\_\{n\}\\in\\mathcal\{X\}, define the sample\-dependent lower\-support envelope
B𝐱:=supa∈𝒜L−1m,Gmaxi∈\[n\]\|a\(𝐱i\)\|\.\\displaystyle B\_\{\\mathbf\{x\}\}:=\\sup\_\{a\\in\\mathcal\{A\}\_\{L\-1\}^\{m,G\}\}\\max\_\{i\\in\\left\[n\\right\]\}\\left\|a\\left\(\\mathbf\{x\}\_\{i\}\\right\)\\right\|\.\(54\)Then there exists a constantCL\>0C\_\{L\}\>0, depending only on the recursion depthLL, such that
ℜ^n\(𝒲RL;m,G\)≤RB𝐱1/2min\{1,CL\(ΓL−1,m,G,nn\)1/2\}\.\\displaystyle\\widehat\{\\mathfrak\{R\}\}\_\{n\}\\left\(\\mathcal\{W\}\_\{R\}^\{L;m,G\}\\right\)\\leq RB\_\{\\mathbf\{x\}\}^\{1/2\}\\min\\left\\\{1,\\;C\_\{L\}\\left\(\\frac\{\\Gamma\_\{L\-1,m,G,n\}\}\{n\}\\right\)^\{1/2\}\\right\\\}\.\(55\)Moreover, the deterministic range estimate \([33](https://arxiv.org/html/2608.13882#S2.E33)\), established in[Lemma17](https://arxiv.org/html/2608.13882#Thmtheorem17), part[2](https://arxiv.org/html/2608.13882#A4.I1.i2), gives
B𝐱≤A𝒳2−\(L−2\)\.\\displaystyle B\_\{\\mathbf\{x\}\}\\leq A\_\{\\mathcal\{X\}\}^\{\\,2^\{\-\(L\-2\)\}\}\.\(56\)Consequently, ifX1,…,XnX\_\{1\},\\ldots,X\_\{n\}are independent random variables with common distributionν\\nu, then the expected Rademacher complexity satisfies
ℜn\(𝒲RL;m,G\)≤RA𝒳2−\(L−1\)min\{1,CL\(ΓL−1,m,G,nn\)1/2\}\.\\displaystyle\\mathfrak\{R\}\_\{n\}\\left\(\\mathcal\{W\}\_\{R\}^\{L;m,G\}\\right\)\\leq RA\_\{\\mathcal\{X\}\}^\{\\,2^\{\-\(L\-1\)\}\}\\min\\left\\\{1,\\;C\_\{L\}\\left\(\\frac\{\\Gamma\_\{L\-1,m,G,n\}\}\{n\}\\right\)^\{1/2\}\\right\\\}\.\(57\)
The preceding result separates three sources of statistical complexity\. The factorRRis the radius of the outer variation representation, the exponentA𝒳2−\(L−1\)A\_\{\\mathcal\{X\}\}^\{\\,2^\{\-\(L\-1\)\}\}is induced by the recursive Brownian geometry, andPL−1,m,GP\_\{L\-1,m,G\}is the parameter complexity of the finite lower\-support architecture\. The additional real parameter inPL−1,m,G\+1P\_\{L\-1,m,G\}\+1corresponds to the variable threshold used in the threshold\-class analysis\.
###### Corollary 0\.
\(Generalization guarantee for finite Brownian variation classes\)Assume the architecture, parameter, and sample\-size conditions of[Theorem6](https://arxiv.org/html/2608.13882#Thmtheorem6)\. Let𝒟n:=\(\(Xi,Yi\)\)i=1n\\mathcal\{D\}\_\{n\}:=\\left\(\\left\(X\_\{i\},Y\_\{i\}\\right\)\\right\)\_\{i=1\}^\{n\}be an independent sample drawn fromνXY\\nu\_\{XY\}, and letℛ\\mathcal\{R\}andℛn\\mathcal\{R\}\_\{n\}denote, respectively, the population and empirical risks introduced in[Section3\.2\.1](https://arxiv.org/html/2608.13882#S3.SS2.SSS1)\. Suppose thatℓ:ℝ×𝒴⟶\[0,1\]\\ell:\\mathbb\{R\}\\times\\mathcal\{Y\}\\longrightarrow\\left\[0,1\\right\]isLℓL\_\{\\ell\}\-Lipschitz in its first argument\. Write𝐗:=\(X1,…,Xn\),\\mathbf\{X\}:=\\left\(X\_\{1\},\\ldots,X\_\{n\}\\right\),and define the random lower\-support envelope
B𝐗:=supa∈𝒜L−1m,Gmaxi∈\[n\]\|a\(Xi\)\|\.\\displaystyle B\_\{\\mathbf\{X\}\}:=\\sup\_\{a\\in\\mathcal\{A\}\_\{L\-1\}^\{m,G\}\}\\max\_\{i\\in\\left\[n\\right\]\}\\left\|a\\left\(X\_\{i\}\\right\)\\right\|\.\(58\)LetCLC\_\{L\}andΓL−1,m,G,n\\Gamma\_\{L\-1,m,G,n\}be the quantities appearing in[Theorem6](https://arxiv.org/html/2608.13882#Thmtheorem6)\. Then, for everyδ∈\(0,1\)\\delta\\in\\left\(0,1\\right\), with probability at least1−δ1\-\\deltaover the draw of𝒟n\\mathcal\{D\}\_\{n\}, everyF∈𝒲RL;m,GF\\in\\mathcal\{W\}\_\{R\}^\{L;m,G\}satisfies
ℛ\(F\)\\displaystyle\\mathcal\{R\}\\left\(F\\right\)≤ℛn\(F\)\+2LℓRB𝐗1/2min\{1,CL\(ΓL−1,m,G,nn\)1/2\}\+3\(ln\(2/δ\)2n\)1/2\\displaystyle\\leq\\mathcal\{R\}\_\{n\}\\left\(F\\right\)\+2L\_\{\\ell\}RB\_\{\\mathbf\{X\}\}^\{1/2\}\\min\\left\\\{1,\\;C\_\{L\}\\left\(\\frac\{\\Gamma\_\{L\-1,m,G,n\}\}\{n\}\\right\)^\{1/2\}\\right\\\}\+3\\left\(\\frac\{\\ln\\left\(2/\\delta\\right\)\}\{2n\}\\right\)^\{1/2\}≤ℛn\(F\)\+2LℓRA𝒳2−\(L−1\)min\{1,CL\(ΓL−1,m,G,nn\)1/2\}\+3\(ln\(2/δ\)2n\)1/2\.\\displaystyle\\leq\\mathcal\{R\}\_\{n\}\\left\(F\\right\)\+2L\_\{\\ell\}RA\_\{\\mathcal\{X\}\}^\{\\,2^\{\-\(L\-1\)\}\}\\min\\left\\\{1,\\;C\_\{L\}\\left\(\\frac\{\\Gamma\_\{L\-1,m,G,n\}\}\{n\}\\right\)^\{1/2\}\\right\\\}\+3\\left\(\\frac\{\\ln\\left\(2/\\delta\\right\)\}\{2n\}\\right\)^\{1/2\}\.\(59\)
## 4Constructive Approximation of VBKL spaces
This section returns to the full infinite\-dimensional VBKL space\. Starting from an arbitrary measure\-generated element, the first stage replaces the outer signed measure by a finite atomic measure while leaving the selected atoms unchanged\. The second stage discretizes only the outermost Brownian profile of each selected atom; its lower\-level support remains in𝒰L−1\\mathcal\{U\}\_\{L\-1\}\. This separation yields distinct atom\-discretization and outer\-profile\-interpolation errors\.
### 4\.1Finite Approximation by Recursive Atomic Representations
The first stage discretizes only the outer variation representation: the recursive dictionary and the selected atoms remain unchanged\. Its error is therefore independent of profile interpolation\.
###### Theorem 9\.
\(Finite\-atomic approximation of the full VBKL space\)Assume that𝒳⊆ℝd\\mathcal\{X\}\\subseteq\\mathbb\{R\}^\{d\}is compact, thatν\\nuis a Borel probability measure on𝒳\\mathcal\{X\}, and thatΩ⊆𝕊d−1\\Omega\\subseteq\\mathbb\{S\}^\{d\-1\}is compact\. LetL≥2L\\geq 2, and let𝒱\(L\)\\mathcal\{V\}^\{\(L\)\}and𝒞^var\(L\)\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}denote the depth\-LLVBKL space and its variation complexity defined in \([7](https://arxiv.org/html/2608.13882#S2.E7)\) and \([8](https://arxiv.org/html/2608.13882#S2.E8)\)\. For everyF∈𝒱\(L\)F\\in\\mathcal\{V\}^\{\(L\)\}and everym∈ℕm\\in\\mathbb\{N\},m≥1m\\geq 1, there existsFm∈𝒱m\(L\)F\_\{m\}\\in\\mathcal\{V\}\_\{m\}^\{\(L\)\}, where𝒱m\(L\)\\mathcal\{V\}\_\{m\}^\{\(L\)\}is the finite\-support VBKL class introduced in[Section2\.1](https://arxiv.org/html/2608.13882#S2.SS1), such that
‖F−Fm‖L2\(ν\)≤\(∫𝒳‖𝐱‖22−\(L−2\)dν\(𝐱\)\)1/2𝒞^var\(L\)\(F\)m−1/2\.\\displaystyle\\left\\\|F\-F\_\{m\}\\right\\\|\_\{L^\{2\}\\left\(\\nu\\right\)\}\\leq\\left\(\\int\_\{\\mathcal\{X\}\}\\left\\\|\\mathbf\{x\}\\right\\\|\_\{2\}^\{\\,2^\{\-\(L\-2\)\}\}\\,\\mathrm\{d\}\\nu\\left\(\\mathbf\{x\}\\right\)\\right\)^\{1/2\}\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}\\left\(F\\right\)m^\{\-1/2\}\.\(60\)If𝒞^var\(L\)\(F\)=0\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}\\left\(F\\right\)=0, thenF=0F=0inL2\(ν\)L^\{2\}\\left\(\\nu\\right\), and one may takeFm=0F\_\{m\}=0\. If𝒞^var\(L\)\(F\)\>0\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}\\left\(F\\right\)\>0, then there exist atomsu1,…,um∈𝒰Lu\_\{1\},\\ldots,u\_\{m\}\\in\\mathcal\{U\}\_\{L\}and signsσ1,…,σm∈\{−1,1\}\\sigma\_\{1\},\\ldots,\\sigma\_\{m\}\\in\\left\\\{\-1,1\\right\\\}such that
Fm=𝒞^var\(L\)\(F\)m∑i=1mσiui\.\\displaystyle F\_\{m\}=\\frac\{\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}\\left\(F\\right\)\}\{m\}\\sum\_\{i=1\}^\{m\}\\sigma\_\{i\}u\_\{i\}\.\(61\)
Theorem[9](https://arxiv.org/html/2608.13882#Thmtheorem9)controls the finite\-atomic stage\. The next result additionally interpolates the selected atoms’ outer Brownian profiles\. It does not discretize their lower\-level supports, which remain elements of𝒰L−1\\mathcal\{U\}\_\{L\-1\}\.
###### Theorem 10\.
\(Two\-stage finite\-atomic and outer\-profile approximation of the VBKL space\)Let𝒳⊆ℝd\\mathcal\{X\}\\subseteq\\mathbb\{R\}^\{d\}be compact, letν\\nube a Borel probability measure on𝒳\\mathcal\{X\}, and letΩ⊆𝕊d−1\\Omega\\subseteq\\mathbb\{S\}^\{d\-1\}be compact\. LetL≥2L\\geq 2, let\(𝒰l\)l=1L\\left\(\\mathcal\{U\}\_\{l\}\\right\)\_\{l=1\}^\{L\}be the recursive atomic dictionaries introduced in[Section2\.1](https://arxiv.org/html/2608.13882#S2.SS1), and let𝒱\(L\)\\mathcal\{V\}^\{\(L\)\}and𝒞^var\(L\)\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}denote, respectively, the depth\-LLVBKL space and its variation complexity defined in \([7](https://arxiv.org/html/2608.13882#S2.E7)\) and \([8](https://arxiv.org/html/2608.13882#S2.E8)\)\. ForF∈𝒱\(L\)F\\in\\mathcal\{V\}^\{\(L\)\}, defineR𝒳:=sup𝐱∈𝒳‖𝐱‖2\.R\_\{\\mathcal\{X\}\}:=\\sup\_\{\\mathbf\{x\}\\in\\mathcal\{X\}\}\\left\\\|\\mathbf\{x\}\\right\\\|\_\{2\}\.Set
A0:=R𝒳2−\(L−2\)\\displaystyle A\_\{0\}:=R\_\{\\mathcal\{X\}\}^\{\\,2^\{\-\(L\-2\)\}\}\(62\)and
RL:=\(∫𝒳‖𝐱‖22−\(L−2\)𝑑ν\(𝐱\)\)1/2\.\\displaystyle R\_\{L\}:=\\left\(\\int\_\{\\mathcal\{X\}\}\\left\\\|\\mathbf\{x\}\\right\\\|\_\{2\}^\{\\,2^\{\-\(L\-2\)\}\}\\,\\mathrm\{d\}\\nu\\left\(\\mathbf\{x\}\\right\)\\right\)^\{1/2\}\.\(63\)LetM,m∈ℕM,m\\in\\mathbb\{N\}\. If𝒞^var\(L\)\(F\)=0\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}\\left\(F\\right\)=0, thenF=0F=0inL2\(ν\)L^\{2\}\\left\(\\nu\\right\), and one may takeFM,m:=0F\_\{M,m\}:=0\. Suppose now that𝒞^var\(L\)\(F\)\>0\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}\\left\(F\\right\)\>0\. ThenR𝒳\>0R\_\{\\mathcal\{X\}\}\>0andA0\>0A\_\{0\}\>0\. Let−A0=t0<t1<⋯<tm=A0\-A\_\{0\}=t\_\{0\}<t\_\{1\}<\\cdots<t\_\{m\}=A\_\{0\}be the uniform grid on\[−A0,A0\]\\left\[\-A\_\{0\},A\_\{0\}\\right\], and letψ0,…,ψm\\psi\_\{0\},\\ldots,\\psi\_\{m\}denote the associated continuous piecewise\-linear hat functions\. For every functiong∈ℋk\(B\)g\\in\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}, letΠmg\\Pi\_\{m\}gdenote its continuous piecewise\-linear interpolant on this grid, that is,
\(Πmg\)\(t\):=∑k=0mg\(tk\)ψk\(t\),t∈\[−A0,A0\]\.\\displaystyle\\left\(\\Pi\_\{m\}g\\right\)\\left\(t\\right\):=\\sum\_\{k=0\}^\{m\}g\\left\(t\_\{k\}\\right\)\\psi\_\{k\}\\left\(t\\right\),\\qquad t\\in\\left\[\-A\_\{0\},A\_\{0\}\\right\]\.\(64\)Then there exist atomsu1,…,uM∈𝒰Lu\_\{1\},\\ldots,u\_\{M\}\\in\\mathcal\{U\}\_\{L\}and signsσ1,…,σM∈\{−1,1\}\\sigma\_\{1\},\\ldots,\\sigma\_\{M\}\\in\\left\\\{\-1,1\\right\\\}such that, for everyj∈\[M\]j\\in\\left\[M\\right\], there existauj∈𝒰L−1a\_\{u\_\{j\}\}\\in\\mathcal\{U\}\_\{L\-1\}andguj∈ℋk\(B\)g\_\{u\_\{j\}\}\\in\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}satisfying
uj\(𝐱\)\\displaystyle u\_\{j\}\\left\(\\mathbf\{x\}\\right\)=guj\(auj\(𝐱\)\),𝐱∈𝒳,\\displaystyle=g\_\{u\_\{j\}\}\\left\(a\_\{u\_\{j\}\}\\left\(\\mathbf\{x\}\\right\)\\right\),\\qquad\\mathbf\{x\}\\in\\mathcal\{X\},\(65\)‖guj‖ℋk\(B\)\\displaystyle\\left\\\|g\_\{u\_\{j\}\}\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}≤1,\\displaystyle\\leq 1,\(66\)auj\(𝒳\)\\displaystyle a\_\{u\_\{j\}\}\\left\(\\mathcal\{X\}\\right\)⊆\[−A0,A0\]\.\\displaystyle\\subseteq\\left\[\-A\_\{0\},A\_\{0\}\\right\]\.\(67\)The preceding outer\-profile representation follows from[Lemma18](https://arxiv.org/html/2608.13882#Thmtheorem18), while the range inclusion follows from[Lemma31](https://arxiv.org/html/2608.13882#Thmtheorem31)\. Define
FM,m\(𝐱\):=𝒞^var\(L\)\(F\)M∑j=1Mσj\(Πmguj\)\(auj\(𝐱\)\),𝐱∈𝒳\.\\displaystyle F\_\{M,m\}\\left\(\\mathbf\{x\}\\right\):=\\frac\{\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}\\left\(F\\right\)\}\{M\}\\sum\_\{j=1\}^\{M\}\\sigma\_\{j\}\\left\(\\Pi\_\{m\}g\_\{u\_\{j\}\}\\right\)\\left\(a\_\{u\_\{j\}\}\\left\(\\mathbf\{x\}\\right\)\\right\),\\qquad\\mathbf\{x\}\\in\\mathcal\{X\}\.\(68\)Then
‖F−FM,m‖L2\(ν\)≤𝒞^var\(L\)\(F\)\[RLM−1/2\+\(A02\)1/2m−1/2\]\.\\displaystyle\\left\\\|F\-F\_\{M,m\}\\right\\\|\_\{L^\{2\}\\left\(\\nu\\right\)\}\\leq\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}\\left\(F\\right\)\\left\[R\_\{L\}M^\{\-1/2\}\+\\left\(\\frac\{A\_\{0\}\}\{2\}\\right\)^\{1/2\}m^\{\-1/2\}\\right\]\.\(69\)Moreover, for every𝐱∈𝒳\\mathbf\{x\}\\in\\mathcal\{X\}and everyj∈\[M\]j\\in\\left\[M\\right\], the nodal representation
\(Πmguj\)\(auj\(𝐱\)\)\\displaystyle\\left\(\\Pi\_\{m\}g\_\{u\_\{j\}\}\\right\)\\left\(a\_\{u\_\{j\}\}\\left\(\\mathbf\{x\}\\right\)\\right\)=∑k=0mguj\(tk\)ψk\(auj\(𝐱\)\)\\displaystyle=\\sum\_\{k=0\}^\{m\}g\_\{u\_\{j\}\}\\left\(t\_\{k\}\\right\)\\psi\_\{k\}\\left\(a\_\{u\_\{j\}\}\\left\(\\mathbf\{x\}\\right\)\\right\)\(70\)contains at most two nonzero terms\. Consequently, evaluatingFM,m\(𝐱\)F\_\{M,m\}\\left\(\\mathbf\{x\}\\right\)requires at most2M2Mactive outer\-profile basis contributions, independently of the interpolation resolutionmm\. In the case𝒞^var\(L\)\(F\)=0\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}\\left\(F\\right\)=0, the zero realization requires no active outer\-profile contribution\.
The resulting approximant has finite outer atomic support and finite outer\-profile interpolation, while preserving the selected lower\-level supports exactly\.
###### Corollary 0\.
\(Balanced two\-stage approximation\)Assume the hypotheses and notation of[Theorem10](https://arxiv.org/html/2608.13882#Thmtheorem10)\. LetN∈ℕN\\in\\mathbb\{N\},N≥1N\\geq 1, and choose the number of selected recursive atoms and the outer\-profile interpolation resolution identically:M=m=N\.M=m=N\.Then the two\-stage approximantFN,NF\_\{N,N\}provided by[Theorem10](https://arxiv.org/html/2608.13882#Thmtheorem10)satisfies
‖F−FN,N‖L2\(ν\)≤𝒞^var\(L\)\(F\)\[RL\+\(A02\)1/2\]N−1/2\.\\displaystyle\\left\\\|F\-F\_\{N,N\}\\right\\\|\_\{L^\{2\}\\left\(\\nu\\right\)\}\\leq\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}\\left\(F\\right\)\\left\[R\_\{L\}\+\\left\(\\frac\{A\_\{0\}\}\{2\}\\right\)^\{1/2\}\\right\]N^\{\-1/2\}\.\(71\)Consequently, for fixedFF,RLR\_\{L\}, andA0A\_\{0\}, one has‖F−FN,N‖L2\(ν\)=O\(N−1/2\)\.\\left\\\|F\-F\_\{N,N\}\\right\\\|\_\{L^\{2\}\\left\(\\nu\\right\)\}=O\\left\(N^\{\-1/2\}\\right\)\.
### 4\.2Approximation of Brownian Profile Functions
We now isolate the outer\-profile interpolation step and establish a quantitative estimate for Brownian RKHS functions\.
###### Proposition 1\.
\(Interpolation of Brownian RKHS functions on a symmetric interval\)LetA\>0A\>0andm∈ℕm\\in\\mathbb\{N\},m≥1m\\geq 1\. Let−A=t0<t1<⋯<tm=A\-A=t\_\{0\}<t\_\{1\}<\\cdots<t\_\{m\}=Abe the uniform grid on\[−A,A\]\\left\[\-A,A\\right\], and letΠm\\Pi\_\{m\}denote the associated continuous piecewise\-linear interpolation operator\. Thus, forg∈ℋk\(B\)g\\in\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}, the functionΠmg\\Pi\_\{m\}gis the unique continuous function on\[−A,A\]\\left\[\-A,A\\right\]that is affine on every interval\[ti,ti\+1\]\\left\[t\_\{i\},t\_\{i\+1\}\\right\]and satisfies\(Πmg\)\(ti\)=g\(ti\)\\left\(\\Pi\_\{m\}g\\right\)\\left\(t\_\{i\}\\right\)=g\\left\(t\_\{i\}\\right\),i∈\{0,…,m\}\.i\\in\\left\\\{0,\\ldots,m\\right\\\}\.Then, for everyg∈ℋk\(B\)g\\in\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\},
‖g−Πmg‖L∞\(\[−A,A\]\)≤\(A2\)1/2m−1/2‖g‖ℋk\(B\)\.\\displaystyle\\left\\\|g\-\\Pi\_\{m\}g\\right\\\|\_\{L^\{\\infty\}\\left\(\\left\[\-A,A\\right\]\\right\)\}\\leq\\left\(\\frac\{A\}\{2\}\\right\)^\{1/2\}m^\{\-1/2\}\\left\\\|g\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}\.\(73\)
Together, the results give a two\-stage approximation of the full VBKL space: the outer measure is finitely supported and the selected outer profiles are piecewise linear, while the lower\-level supports may remain infinite\-dimensional elements of𝒰L−1\\mathcal\{U\}\_\{L\-1\}\.
## 5Experiments
The experiments are organized around the theoretical mechanisms developed above rather than around a large benchmark collection\. They address four questions: whether the finite constructions exhibit the predicted approximation behaviour; whether validation\-selected realizations can be learned effectively from finite data; whether their predictive accuracy is obtained with an economical recursive representation; and whether the resulting finite models admit numerically stable optimization\. Complete protocols, validation\-selected configurations, numerical tables, and computational measurements are reported in Appendix[A](https://arxiv.org/html/2608.13882#A1)\.
### 5\.1Experimental questions and common protocol
The approximation studies use deterministic teachers whose recursive representations are known explicitly, so that the effects of signed\-measure discretization and Brownian\-profile interpolation can be evaluated without statistical estimation or training\. The supervised studies use a controlled recursive\-teacher problem and the Energy Efficiency regression benchmark\. They compare finite VBKL realizations with Deep Neural Variation Spaces \(DNVS\), kernel ridge regression \(KRR\), and random Fourier features \(RBF\)\. For each training\-set size and random seed, the number of recursive Brownian paths and the optimization horizon of VBKL are selected using validation data before the selected model is retrained\. The Energy Efficiency study also selects the Brownian\-profile resolution by validation\. All supervised results are reported over five independent random seeds using test mean\-squared error\. The compared methods use the same data\-splitting protocol within each benchmark\.
### 5\.2Constructive approximation and interpolation sharpness
We first examine the outer\-profile\-discretization component of the two\-stage approximation\. Three representative unit\-norm Brownian RKHS profiles are considered: a smooth sinusoidal profile, a localized profile, and an oscillatory profile\. Their continuous piecewise\-linear interpolants are evaluated on uniform grids withm∈\{8,16,32,64,128,256,512\}m\\in\\\{8,16,32,64,128,256,512\\\}\. Figure[1](https://arxiv.org/html/2608.13882#S5.F1)\(a\) shows that these fixed profiles converge substantially faster than the uniform worst\-case rate\. This behaviour reflects their additional regularity and clarifies that[1](https://arxiv.org/html/2608.13882#Thmprop1)is a worst\-case guarantee rather than the asymptotic rate of every fixed profile\.
For each resolution, we also construct a normalized tent profile supported on a single interpolation interval\. Its interpolation error coincides with the theoretical bound at every tested resolution, as shown in Figure[1](https://arxiv.org/html/2608.13882#S5.F1)\(b\)\. The experiment therefore separates typical fixed\-profile behaviour from the sharp worst case proved in Appendix[C](https://arxiv.org/html/2608.13882#A3)\.
Figure 1:Brownian profile interpolation\.\(a\)Interpolation errors for three representative unit\-norm Brownian RKHS profiles\.\(b\)Interpolation error of the normalized tent profile together with the theoretical interpolation bound\. The coincidence of the curves demonstrates sharpness of the interpolation estimate\.We next test both stages of the construction\. In Figure[2](https://arxiv.org/html/2608.13882#S5.F2)\(a\), only the representing signed measure is discretized while the selected Brownian profiles remain exact; the error decreases consistently with theM−1/2M^\{\-1/2\}reference predicted by[Theorem9](https://arxiv.org/html/2608.13882#Thmtheorem9)\. Panel \(b\) fixes the finite signed\-measure representation and varies only the profile resolution\. The deterministic teacher converges faster than them−1/2m^\{\-1/2\}worst\-case upper bound, in agreement with the fixed\-profile behaviour above\. Panel \(c\) refines both complexities simultaneously withM=m=NM=m=Nand exhibits the balancedN−1/2N^\{\-1/2\}behaviour of[Corollary12](https://arxiv.org/html/2608.13882#Thmtheorem12)\. Together, the two figures illustrate the separate roles of outer\-measure and outer\-profile complexity in[Theorem10](https://arxiv.org/html/2608.13882#Thmtheorem10)\.
Figure 2:Two\-stage constructive approximation\.\(a\)Finite signed\-measure discretization with exact profiles\.\(b\)Brownian\-profile discretization with a fixed finite measure representation, together with the worst\-case upper bound\.\(c\)Balanced refinementM=m=NM=m=N\. The reference slopes correspond to the rates in[Theorems9](https://arxiv.org/html/2608.13882#Thmtheorem9)and[12](https://arxiv.org/html/2608.13882#Thmtheorem12)\.
### 5\.3Statistical learning and parameter efficiency
Figure[3](https://arxiv.org/html/2608.13882#S5.F3)\(a\) reports the recursive\-teacher learning curve\. This controlled benchmark matches the hierarchical Brownian structure of the target to the proposed hypothesis class\. The validation\-selected VBKL realizations are particularly effective in the limited\-data regime and remain competitive as the sample size grows; the performance differences narrow at larger sample sizes, where DNVS attains the lowest mean error at the largest sizes considered\. This behaviour is consistent with using validation to adapt the explicit finite representation complexity rather than fixing one architecture throughout the learning curve\.
The Energy Efficiency benchmark provides a complementary real\-data test whose generating mechanism is unknown\. Figure[3](https://arxiv.org/html/2608.13882#S5.F3)\(b\) shows that, atn=100n=100, VBKL is comparable to the strongest baselines and has lower mean error than DNVS\. Atn=250n=250andn=500n=500, VBKL does not lead the predictive comparison: DNVS, KRR, and RBF obtain lower mean errors\. Thus, the empirical claim is not universal dominance, but a distinct limited\-data regime in which the recursive Brownian realization is most competitive\.
Representation size provides the second part of this comparison\. Figure[3](https://arxiv.org/html/2608.13882#S5.F3)\(c\) and Table[2](https://arxiv.org/html/2608.13882#S5.T2)show that VBKL uses substantially fewer parameters than DNVS at every Energy Efficiency training size\. Atn=100n=100, VBKL wins four of the five matched seed\-wise comparisons while DNVS uses approximately4\.6×4\.6\\timesas many parameters\. The parameter ratio grows to13\.7×13\.7\\timesand18\.3×18\.3\\timesatn=250n=250andn=500n=500, respectively, even though DNVS wins more of the corresponding accuracy comparisons\. The combined evidence therefore supports a favourable accuracy–complexity trade\-off, with the clearest empirical advantage in the smallest\-data regime\.
\(a\)Recursive\-teacher learning curve\.\(b\)Energy Efficiency learning curve\.\(c\)Accuracy–complexity comparison with DNVS\.
Figure 3:Statistical learning and representation efficiency\.Test mean\-squared error is averaged over five independent random seeds\. VBKL complexity is selected independently by validation for each training\-set size and seed\. Panel\(c\)uses logarithmic axes; small markers represent individual seeds and outlined markers represent means\.nnVBKL wins vs\. DNVSParameter ratio1004/54\.6×\\times2502/513\.7×\\times5002/518\.3×\\timesTable 2:Seed\-wise and parameter\-efficiency comparison between VBKL and DNVS\. The parameter ratio is DNVS/VBKL\.
### 5\.4Optimization stability and practical realizability
The constructive theory produces finite recursive models whose parameters must be optimized numerically\. Figure[4](https://arxiv.org/html/2608.13882#S5.F4)\(a\) compares finite\-difference directional estimates with the corresponding automatic\-differentiation directional derivatives\. The relative error falls as the interaction scale decreases from10−110^\{\-1\}to10−310^\{\-3\}and rises modestly at10−410^\{\-4\}, displaying the expected finite\-difference trade\-off at the smallest scale rather than monotone improvement for arbitrarily small steps\. Panel \(b\) evaluates Monte Carlo directional averaging: increasing the number of sampled directions progressively reduces the estimator standard deviation\.
These controlled tests verify the implementation of the directional estimator and the variance reduction obtained from averaging\. They do not, by themselves, establish a global optimization\-convergence theorem; rather, they show that the finite VBKL realizations can be optimized with a stable numerical estimator under the tested conditions\. Detailed computational measurements are reported in Appendix[A\.4](https://arxiv.org/html/2608.13882#A1.SS4)\.
Figure 4:Optimization validation of finite VBKL realizations\.\(a\)Relative error between automatic\-differentiation directional derivatives and the averaged directional estimator across interaction scales\.\(b\)Estimator standard deviation under Monte Carlo directional averaging\. Increasing the number of sampled directions reduces the variability of the estimate\.
## 6Conclusion
This paper introduced the*Variation Brownian Kernel Ladder*\(VBKL\), a path\-atomic function\-space framework in which recursive Brownian dictionary construction precedes outer variation superposition\. The recursive dictionaries are unions of Brownian pullback RKHS unit balls\. Their variation hulls admit quantitative pointwise and Hölder representatives; under full support these representatives define a continuous Hölder embedding, and under the stated local non\-degeneracy and trace\-support conditions the spaces form a strict depth hierarchy\.
For associated finite lower\-support architectures, we derived Rademacher and generalization bounds that separate the outer variation radius from the architecture parameter count\. For the full VBKL space, the constructive results separate two approximation resources:MMcontrols finite support of the outer measure, andmmcontrols interpolation of the selected outer Brownian profiles\. The lower\-level supports remain in𝒰L−1\\mathcal\{U\}\_\{L\-1\}\. The error is the sum ofM−1/2M^\{\-1/2\}andm−1/2m^\{\-1/2\}terms, the profile\-interpolation constant is sharp, and each evaluation uses at most2M2Mactive outer\-profile basis contributions\.
The experiments illustrate these mechanisms and indicate a favorable trade\-off between accuracy and complexity in limited\-data regimes, without claiming universal predictive dominance or global optimization convergence\. Future work includes intrinsic full\-space statistical bounds, recursive discretization of lower\-level supports, and optimization theory for explicit finite architectures\.
## 7Proofs
This section contains the proofs of the main analytical, statistical, and approximation results\.
### 7\.1Proof of[Theorem4](https://arxiv.org/html/2608.13882#Thmtheorem4)[1](https://arxiv.org/html/2608.13882#S3.I1.i1)
FixF∈𝒱\(L\)F\\in\\mathcal\{V\}^\{\(L\)\}\. For eachn≥1n\\geq 1, choose a finite signed measureμn\\mu\_\{n\}on𝒰L\\mathcal\{U\}\_\{L\}representingFFinL2\(ν\)L^\{2\}\(\\nu\)and satisfying
‖μn‖TV≤𝒞^var\(L\)\(F\)\+n−1\.\\displaystyle\\\|\\mu\_\{n\}\\\|\_\{\\mathrm\{TV\}\}\\leq\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}\(F\)\+n^\{\-1\}\.\(74\)The induced Borel structure on𝒰L⊂C\(𝒳\)\\mathcal\{U\}\_\{L\}\\subset C\(\\mathcal\{X\}\), separability ofC\(𝒳\)C\(\\mathcal\{X\}\), and the uniform atom bound in[Lemma26](https://arxiv.org/html/2608.13882#Thmtheorem26)imply that the canonical inclusion is Bochner integrable inC\(𝒳\)C\(\\mathcal\{X\}\)with respect to eachμn\\mu\_\{n\}\. Define the correspondingC\(𝒳\)C\(\\mathcal\{X\}\)\-valued barycentre by
Fn\(𝐱\):=∫𝒰Lu\(𝐱\)dμn\(u\),𝐱∈𝒳\.\\displaystyle F\_\{n\}\(\\mathbf\{x\}\):=\\int\_\{\\mathcal\{U\}\_\{L\}\}u\(\\mathbf\{x\}\)\\,\\mathrm\{d\}\\mu\_\{n\}\(u\),\\qquad\\mathbf\{x\}\\in\\mathcal\{X\}\.\(75\)By[Lemma26](https://arxiv.org/html/2608.13882#Thmtheorem26), everyu∈𝒰Lu\\in\\mathcal\{U\}\_\{L\}satisfies the uniformαL\\alpha\_\{L\}\-Hölder and pointwise bounds\. Hence
\|Fn\(𝐱\)−Fn\(𝐱′\)\|\\displaystyle\|F\_\{n\}\(\\mathbf\{x\}\)\-F\_\{n\}\(\\mathbf\{x\}^\{\\prime\}\)\|≤‖𝐱−𝐱′‖2αL‖μn‖TV,\\displaystyle\\leq\\\|\\mathbf\{x\}\-\\mathbf\{x\}^\{\\prime\}\\\|\_\{2\}^\{\\alpha\_\{L\}\}\\\|\\mu\_\{n\}\\\|\_\{\\mathrm\{TV\}\},\|Fn\(𝐱\)\|\\displaystyle\|F\_\{n\}\(\\mathbf\{x\}\)\|≤‖𝐱‖2αL‖μn‖TV\.\\displaystyle\\leq\\\|\\mathbf\{x\}\\\|\_\{2\}^\{\\alpha\_\{L\}\}\\\|\\mu\_\{n\}\\\|\_\{\\mathrm\{TV\}\}\.\(76\)The sequence\(Fn\)n≥1\(F\_\{n\}\)\_\{n\\geq 1\}is uniformly bounded and equicontinuous on the compact set𝒳\\mathcal\{X\}\. By Arzelà–Ascoli, a subsequence converges uniformly to someF⋆∈C\(𝒳\)F\_\{\\star\}\\in C\(\\mathcal\{X\}\)\. The continuous inclusionC\(𝒳\)↪L2\(ν\)C\(\\mathcal\{X\}\)\\hookrightarrow L^\{2\}\(\\nu\)commutes with Bochner integration and shows that eachFnF\_\{n\}represents the same elementFFofL2\(ν\)L^\{2\}\(\\nu\)\. Uniform convergence therefore implies convergence inL2\(ν\)L^\{2\}\(\\nu\), and henceF⋆=FF\_\{\\star\}=FinL2\(ν\)L^\{2\}\(\\nu\)\. Passing to the limit in the preceding bounds and using‖μn‖TV→𝒞^var\(L\)\(F\)\\\|\\mu\_\{n\}\\\|\_\{\\mathrm\{TV\}\}\\to\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}\(F\)along the selected near\-minimizing sequence gives
\|F⋆\(𝐱\)−F⋆\(𝐱′\)\|\\displaystyle\|F\_\{\\star\}\(\\mathbf\{x\}\)\-F\_\{\\star\}\(\\mathbf\{x\}^\{\\prime\}\)\|≤‖𝐱−𝐱′‖2αL𝒞^var\(L\)\(F\),\\displaystyle\\leq\\\|\\mathbf\{x\}\-\\mathbf\{x\}^\{\\prime\}\\\|\_\{2\}^\{\\alpha\_\{L\}\}\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}\(F\),\|F⋆\(𝐱\)\|\\displaystyle\|F\_\{\\star\}\(\\mathbf\{x\}\)\|≤‖𝐱‖2αL𝒞^var\(L\)\(F\)\.\\displaystyle\\leq\\\|\\mathbf\{x\}\\\|\_\{2\}^\{\\alpha\_\{L\}\}\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}\(F\)\.\(77\)IdentifyingFFwith this representative proves \([40](https://arxiv.org/html/2608.13882#S3.E40)\) and \([41](https://arxiv.org/html/2608.13882#S3.E41)\)\.
### 7\.2Proof of[Theorem4](https://arxiv.org/html/2608.13882#Thmtheorem4)[2](https://arxiv.org/html/2608.13882#S3.I1.i2)
The bounds in \([40](https://arxiv.org/html/2608.13882#S3.E40)\) and \([41](https://arxiv.org/html/2608.13882#S3.E41)\) give
\[F\]C0,αL\(𝒳\)\\displaystyle\[F\]\_\{C^\{0,\\alpha\_\{L\}\}\(\\mathcal\{X\}\)\}≤𝒞^var\(L\)\(F\),\\displaystyle\\leq\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}\(F\),‖F‖L∞\(𝒳\)\\displaystyle\\\|F\\\|\_\{L^\{\\infty\}\(\\mathcal\{X\}\)\}≤R𝒳αL𝒞^var\(L\)\(F\)\.\\displaystyle\\leq R\_\{\\mathcal\{X\}\}^\{\\alpha\_\{L\}\}\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}\(F\)\.\(78\)Adding the two estimates proves \([42](https://arxiv.org/html/2608.13882#S3.E42)\)\. Ifsupp\(ν\)=𝒳\\operatorname\{supp\}\(\\nu\)=\\mathcal\{X\}, two continuous representatives of the sameL2\(ν\)L^\{2\}\(\\nu\)element agree on𝒳\\mathcal\{X\}; hence the representative is unique and the induced map intoC0,αL\(𝒳\)C^\{0,\\alpha\_\{L\}\}\(\\mathcal\{X\}\)is linear, injective, and continuous\.
### 7\.3Proof of[Theorem4](https://arxiv.org/html/2608.13882#Thmtheorem4)[3](https://arxiv.org/html/2608.13882#S3.I1.i3)
We first prove the depth inclusion\. Set
AL:=R𝒳αL\.\\displaystyle A\_\{L\}:=R\_\{\\mathcal\{X\}\}^\{\\,\\alpha\_\{L\}\}\.\(79\)By[Lemma26](https://arxiv.org/html/2608.13882#Thmtheorem26), everyu∈𝒰Lu\\in\\mathcal\{U\}\_\{L\}satisfies
\|u\(𝐱\)\|≤‖𝐱‖2αL≤R𝒳αL=AL,𝐱∈𝒳\.\\displaystyle\|u\(\\mathbf\{x\}\)\|\\leq\\\|\\mathbf\{x\}\\\|\_\{2\}^\{\\alpha\_\{L\}\}\\leq R\_\{\\mathcal\{X\}\}^\{\\alpha\_\{L\}\}=A\_\{L\},\\qquad\\mathbf\{x\}\\in\\mathcal\{X\}\.\(80\)Thus,
u\(𝒳\)⊆\[−AL,AL\],u∈𝒰L\.\\displaystyle u\\left\(\\mathcal\{X\}\\right\)\\subseteq\\left\[\-A\_\{L\},A\_\{L\}\\right\],\\qquad u\\in\\mathcal\{U\}\_\{L\}\.\(81\)ChooseχL∈Cc∞\(ℝ\)\\chi\_\{L\}\\in C\_\{c\}^\{\\infty\}\\left\(\\mathbb\{R\}\\right\)satisfyingχL\(t\)=1\\chi\_\{L\}\\left\(t\\right\)=1fort∈\[−AL,AL\]t\\in\\left\[\-A\_\{L\},A\_\{L\}\\right\], and define
ϕL\(t\):=tχL\(t\),t∈ℝ\.\\displaystyle\\phi\_\{L\}\\left\(t\\right\):=t\\chi\_\{L\}\\left\(t\\right\),\\qquad t\\in\\mathbb\{R\}\.\(82\)The functionϕL\\phi\_\{L\}is absolutely continuous, vanishes at the origin, and has a compactly supported derivative inL2\(ℝ\)L^\{2\}\\left\(\\mathbb\{R\}\\right\)\. Hence, by the Brownian RKHS characterization\([18](https://arxiv.org/html/2608.13882#bib.bib1), Lemma C9\),
ϕL\\displaystyle\\phi\_\{L\}∈ℋk\(B\),\\displaystyle\\in\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\},cL\\displaystyle c\_\{L\}:=‖ϕL‖ℋk\(B\)∈\(0,∞\)\.\\displaystyle:=\\left\\\|\\phi\_\{L\}\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}\\in\\left\(0,\\infty\\right\)\.\(83\)Define
SL:𝒰L\\displaystyle S\_\{L\}:\\mathcal\{U\}\_\{L\}⟶𝒰L\+1,\\displaystyle\\longrightarrow\\mathcal\{U\}\_\{L\+1\},SL\(u\)\\displaystyle S\_\{L\}\\left\(u\\right\):=ϕLcL∘u\.\\displaystyle:=\\frac\{\\phi\_\{L\}\}\{c\_\{L\}\}\\circ u\.\(84\)Since‖ϕL/cL‖ℋk\(B\)=1\\left\\\|\\phi\_\{L\}/c\_\{L\}\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}=1, the recursive definition of𝒰L\+1\\mathcal\{U\}\_\{L\+1\}givesSL\(u\)∈𝒰L\+1S\_\{L\}\\left\(u\\right\)\\in\\mathcal\{U\}\_\{L\+1\}\. Moreover, sinceϕL\\phi\_\{L\}is globally Lipschitz,
‖SL\(u\)−SL\(v\)‖L∞\(𝒳\)≤Lip\(ϕL\)cL‖u−v‖L∞\(𝒳\),u,v∈𝒰L\.\\displaystyle\\left\\\|S\_\{L\}\(u\)\-S\_\{L\}\(v\)\\right\\\|\_\{L^\{\\infty\}\(\\mathcal\{X\}\)\}\\leq\\frac\{\\operatorname\{Lip\}\(\\phi\_\{L\}\)\}\{c\_\{L\}\}\\left\\\|u\-v\\right\\\|\_\{L^\{\\infty\}\(\\mathcal\{X\}\)\},\\qquad u,v\\in\\mathcal\{U\}\_\{L\}\.\(85\)ThusSLS\_\{L\}is continuous for the induced supremum\-norm topologies and hence Borel measurable\. SinceϕL\\phi\_\{L\}agrees with the identity on\[−AL,AL\]\\left\[\-A\_\{L\},A\_\{L\}\\right\], \([81](https://arxiv.org/html/2608.13882#S7.E81)\) gives
u=cLSL\(u\),u∈𝒰L\.\\displaystyle u=c\_\{L\}S\_\{L\}\\left\(u\\right\),\\qquad u\\in\\mathcal\{U\}\_\{L\}\.\(86\)FixF∈𝒱\(L\)F\\in\\mathcal\{V\}^\{\(L\)\}, and choose a finite signed Borel measureμ∈ℳ\(𝒰L\)\\mu\\in\\mathcal\{M\}\\left\(\\mathcal\{U\}\_\{L\}\\right\)such that
F\\displaystyle F=∫𝒰Ludμ\(u\)inL2\(ν\),\\displaystyle=\\int\_\{\\mathcal\{U\}\_\{L\}\}u\\,\\mathrm\{d\}\\mu\\left\(u\\right\)\\qquad\\text\{in \}L^\{2\}\\left\(\\nu\\right\),‖μ‖TV\\displaystyle\\left\\\|\\mu\\right\\\|\_\{\\mathrm\{TV\}\}<∞\.\\displaystyle<\\infty\.\(87\)Define
μ~:=cL\(SL\)\#μ\.\\displaystyle\\widetilde\{\\mu\}:=c\_\{L\}\\left\(S\_\{L\}\\right\)\_\{\\\#\}\\mu\.\(88\)Therefore,
∫𝒰L\+1v𝑑μ~\(v\)\\displaystyle\\int\_\{\\mathcal\{U\}\_\{L\+1\}\}v\\,\\mathrm\{d\}\\widetilde\{\\mu\}\\left\(v\\right\)=\(a\)cL∫𝒰LSL\(u\)𝑑μ\(u\)=\(b\)∫𝒰Lu𝑑μ\(u\)=\(c\)F\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}c\_\{L\}\\int\_\{\\mathcal\{U\}\_\{L\}\}S\_\{L\}\\left\(u\\right\)\\,\\mathrm\{d\}\\mu\\left\(u\\right\)\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\int\_\{\\mathcal\{U\}\_\{L\}\}u\\,\\mathrm\{d\}\\mu\\left\(u\\right\)\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{=\}\}F\.\(89\)Here, \(a\) applies the push\-forward change\-of\-variables identity, \(b\) uses \([86](https://arxiv.org/html/2608.13882#S7.E86)\), and \(c\) uses the selected representation ofFF\. Furthermore,
‖μ~‖TV≤cL‖μ‖TV<∞\.\\displaystyle\\left\\\|\\widetilde\{\\mu\}\\right\\\|\_\{\\mathrm\{TV\}\}\\leq c\_\{L\}\\left\\\|\\mu\\right\\\|\_\{\\mathrm\{TV\}\}<\\infty\.\(90\)Hence,F∈𝒱\(L\+1\)F\\in\\mathcal\{V\}^\{\(L\+1\)\}\. SinceF∈𝒱\(L\)F\\in\\mathcal\{V\}^\{\(L\)\}was arbitrary,
𝒱\(L\)⊆𝒱\(L\+1\)\.\\displaystyle\\mathcal\{V\}^\{\(L\)\}\\subseteq\\mathcal\{V\}^\{\(L\+1\)\}\.\(91\)We now prove strictness under the local non\-degeneracy and trace\-support conditions in the theorem statement\. Recall thatαL=2−\(L−1\)\\alpha\_\{L\}=2^\{\-\(L\-1\)\}, and choose
a∈\(12,αL1/L\)=\(12,2−\(L−1\)/L\)\.\\displaystyle a\\in\\left\(\\frac\{1\}\{2\},\\alpha\_\{L\}^\{1/L\}\\right\)=\\left\(\\frac\{1\}\{2\},2^\{\-\(L\-1\)/L\}\\right\)\.\(92\)This interval is nonempty because2−\(L−1\)/L=2−1\+1/L\>1/22^\{\-\(L\-1\)/L\}=2^\{\-1\+1/L\}\>1/2\. Following the fractional\-power construction in[18](https://arxiv.org/html/2608.13882#bib.bib1), chooseη∈Cc∞\(ℝ\)\\eta\\in C\_\{c\}^\{\\infty\}\\left\(\\mathbb\{R\}\\right\)such that0≤η≤10\\leq\\eta\\leq 1andη\(t\)=1\\eta\\left\(t\\right\)=1for\|t\|≤1/2\\left\|t\\right\|\\leq 1/2, and define
ga\(t\):=η\(t\)\(t\+\)a,t\+:=max\{t,0\}\.\\displaystyle g\_\{a\}\\left\(t\\right\):=\\eta\\left\(t\\right\)\\left\(t\_\{\+\}\\right\)^\{a\},\\qquad t\_\{\+\}:=\\max\\left\\\{t,0\\right\\\}\.\(93\)The calculation in the cited proof shows thata\>1/2a\>1/2impliesga∈ℋk\(B\)g\_\{a\}\\in\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\. Set
κa\\displaystyle\\kappa\_\{a\}:=‖ga‖ℋk\(B\),\\displaystyle:=\\left\\\|g\_\{a\}\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\},ha\\displaystyle h\_\{a\}:=gaκa\.\\displaystyle:=\\frac\{g\_\{a\}\}\{\\kappa\_\{a\}\}\.\(94\)Then
‖ha‖ℋk\(B\)\\displaystyle\\left\\\|h\_\{a\}\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}=1,\\displaystyle=1,ha\(s\)\\displaystyle h\_\{a\}\\left\(s\\right\)=κa−1sa,0≤s≤12\.\\displaystyle=\\kappa\_\{a\}^\{\-1\}s^\{a\},\\qquad 0\\leq s\\leq\\frac\{1\}\{2\}\.\(95\)Starting from the first\-layer generatoru1u\_\{1\}in the theorem statement, define
v1\\displaystyle v\_\{1\}:=u1,\\displaystyle:=u\_\{1\},vl\+1\\displaystyle v\_\{l\+1\}:=ha∘vl,l∈\[L\]\.\\displaystyle:=h\_\{a\}\\circ v\_\{l\},\\qquad l\\in\\left\[L\\right\]\.\(96\)Since‖ha‖ℋk\(B\)=1\\left\\\|h\_\{a\}\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}=1, the recursive dictionary construction gives
vl\\displaystyle v\_\{l\}∈𝒰l,l∈\{1,…,L\+1\},\\displaystyle\\in\\mathcal\{U\}\_\{l\},\\qquad l\\in\\left\\\{1,\\ldots,L\+1\\right\\\},\(97\)vL\+1\\displaystyle v\_\{L\+1\}∈𝒰L\+1⊆𝒱\(L\+1\)\.\\displaystyle\\in\\mathcal\{U\}\_\{L\+1\}\\subseteq\\mathcal\{V\}^\{\(L\+1\)\}\.\(98\)We claim that, for everyl∈\{1,…,L\+1\}l\\in\\left\\\{1,\\ldots,L\+1\\right\\\}, there existδl∈\(0,ρ\]\\delta\_\{l\}\\in\\left\(0,\\rho\\right\]andcl,Cl∈\(0,∞\)c\_\{l\},C\_\{l\}\\in\\left\(0,\\infty\\right\)such that
cltal−1≤vl\(γ\(t\)\)≤Cltal−1,t∈\[0,δl\]\.\\displaystyle c\_\{l\}t^\{a^\{l\-1\}\}\\leq v\_\{l\}\\left\(\\gamma\\left\(t\\right\)\\right\)\\leq C\_\{l\}t^\{a^\{l\-1\}\},\\qquad t\\in\\left\[0,\\delta\_\{l\}\\right\]\.\(99\)Forl=1l=1, this is precisely the local non\-degeneracy assumption, withδ1:=ρ\\delta\_\{1\}:=\\rho\. Suppose that the claim holds at somel∈\[L\]l\\in\\left\[L\\right\], and chooseδl\+1∈\(0,δl\]\\delta\_\{l\+1\}\\in\\left\(0,\\delta\_\{l\}\\right\]such that
Clδl\+1al−1≤12\.\\displaystyle C\_\{l\}\\delta\_\{l\+1\}^\{\\,a^\{l\-1\}\}\\leq\\frac\{1\}\{2\}\.\(100\)Therefore, fort∈\[0,δl\+1\]t\\in\\left\[0,\\delta\_\{l\+1\}\\right\],
0\\displaystyle 0≤\(a\)vl\(γ\(t\)\)≤\(b\)Cltal−1≤\(c\)12,\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}v\_\{l\}\\left\(\\gamma\\left\(t\\right\)\\right\)\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}C\_\{l\}t^\{a^\{l\-1\}\}\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{\\leq\}\}\\frac\{1\}\{2\},vl\+1\(γ\(t\)\)\\displaystyle v\_\{l\+1\}\\left\(\\gamma\\left\(t\\right\)\\right\)=\(d\)ha\(vl\(γ\(t\)\)\)=\(e\)κa−1vl\(γ\(t\)\)a\.\\displaystyle\\stackrel\{\{\\scriptstyle\(d\)\}\}\{\{=\}\}h\_\{a\}\\left\(v\_\{l\}\\left\(\\gamma\\left\(t\\right\)\\right\)\\right\)\\stackrel\{\{\\scriptstyle\(e\)\}\}\{\{=\}\}\\kappa\_\{a\}^\{\-1\}v\_\{l\}\\left\(\\gamma\\left\(t\\right\)\\right\)^\{a\}\.\(101\)Here, \(a\) and \(b\) use the induction bounds, \(c\) applies \([100](https://arxiv.org/html/2608.13882#S7.E100)\), \(d\) uses \([96](https://arxiv.org/html/2608.13882#S7.E96)\), and \(e\) applies \([95](https://arxiv.org/html/2608.13882#S7.E95)\)\. Raising the induction bounds to the positive poweraatherefore gives
κa−1clatal≤vl\+1\(γ\(t\)\)≤κa−1Clatal,t∈\[0,δl\+1\]\.\\displaystyle\\kappa\_\{a\}^\{\-1\}c\_\{l\}^\{a\}t^\{a^\{l\}\}\\leq v\_\{l\+1\}\\left\(\\gamma\\left\(t\\right\)\\right\)\\leq\\kappa\_\{a\}^\{\-1\}C\_\{l\}^\{a\}t^\{a^\{l\}\},\\qquad t\\in\\left\[0,\\delta\_\{l\+1\}\\right\]\.\(102\)Thus the induction continues withcl\+1:=κa−1clac\_\{l\+1\}:=\\kappa\_\{a\}^\{\-1\}c\_\{l\}^\{a\}andCl\+1:=κa−1ClaC\_\{l\+1\}:=\\kappa\_\{a\}^\{\-1\}C\_\{l\}^\{a\}\. Setδ:=δL\+1\\delta:=\\delta\_\{L\+1\}\. Then
cL\+1taL≤vL\+1\(γ\(t\)\)≤CL\+1taL,t∈\[0,δ\]\.\\displaystyle c\_\{L\+1\}t^\{a^\{L\}\}\\leq v\_\{L\+1\}\\left\(\\gamma\\left\(t\\right\)\\right\)\\leq C\_\{L\+1\}t^\{a^\{L\}\},\\qquad t\\in\\left\[0,\\delta\\right\]\.\(103\)Moreover,u1\(𝐱0\)=0u\_\{1\}\\left\(\\mathbf\{x\}\_\{0\}\\right\)=0andha\(0\)=0h\_\{a\}\\left\(0\\right\)=0imply recursively that
vl\(𝐱0\)=0,l∈\{1,…,L\+1\}\.\\displaystyle v\_\{l\}\\left\(\\mathbf\{x\}\_\{0\}\\right\)=0,\\qquad l\\in\\left\\\{1,\\ldots,L\+1\\right\\\}\.\(104\)Suppose, toward a contradiction, thatvL\+1∈𝒱\(L\)v\_\{L\+1\}\\in\\mathcal\{V\}^\{\(L\)\}\. By[Theorem4](https://arxiv.org/html/2608.13882#Thmtheorem4)[1](https://arxiv.org/html/2608.13882#S3.I1.i1), itsL2\(ν\)L^\{2\}\\left\(\\nu\\right\)equivalence class admits anαL\\alpha\_\{L\}\-Hölder representativev~L\+1\\widetilde\{v\}\_\{L\+1\}\. BothvL\+1v\_\{L\+1\}andv~L\+1\\widetilde\{v\}\_\{L\+1\}are continuous and agreeν\\nu\-almost everywhere\. Consequently, they agree throughoutsupp\(ν\)\\operatorname\{supp\}\\left\(\\nu\\right\):
vL\+1\(𝐱\)=v~L\+1\(𝐱\),𝐱∈supp\(ν\)\.\\displaystyle v\_\{L\+1\}\\left\(\\mathbf\{x\}\\right\)=\\widetilde\{v\}\_\{L\+1\}\\left\(\\mathbf\{x\}\\right\),\\qquad\\mathbf\{x\}\\in\\operatorname\{supp\}\\left\(\\nu\\right\)\.\(105\)LetLγ<∞L\_\{\\gamma\}<\\inftybe a Lipschitz constant ofγ\\gamma\. By \([45](https://arxiv.org/html/2608.13882#S3.E45)\), \([103](https://arxiv.org/html/2608.13882#S7.E103)\), \([104](https://arxiv.org/html/2608.13882#S7.E104)\), and \([105](https://arxiv.org/html/2608.13882#S7.E105)\), there existsC⋆<∞C\_\{\\star\}<\\inftysuch that, for everyt∈\(0,δ\]t\\in\\left\(0,\\delta\\right\],
cL\+1taL\\displaystyle c\_\{L\+1\}t^\{a^\{L\}\}≤\(a\)vL\+1\(γ\(t\)\)=\(b\)\|v~L\+1\(γ\(t\)\)−v~L\+1\(γ\(0\)\)\|\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}v\_\{L\+1\}\\left\(\\gamma\\left\(t\\right\)\\right\)\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\left\|\\widetilde\{v\}\_\{L\+1\}\\left\(\\gamma\\left\(t\\right\)\\right\)\-\\widetilde\{v\}\_\{L\+1\}\\left\(\\gamma\\left\(0\\right\)\\right\)\\right\|≤\(c\)C⋆‖γ\(t\)−γ\(0\)‖2αL≤\(d\)C⋆LγαLtαL\.\\displaystyle\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{\\leq\}\}C\_\{\\star\}\\left\\\|\\gamma\\left\(t\\right\)\-\\gamma\\left\(0\\right\)\\right\\\|\_\{2\}^\{\\,\\alpha\_\{L\}\}\\stackrel\{\{\\scriptstyle\(d\)\}\}\{\{\\leq\}\}C\_\{\\star\}L\_\{\\gamma\}^\{\\,\\alpha\_\{L\}\}t^\{\\alpha\_\{L\}\}\.\(106\)Here, \(a\) applies the lower trace bound, \(b\) uses the anchor identity and the agreement of the two representatives on the trace, \(c\) applies the Hölder estimate, and \(d\) uses the Lipschitz continuity ofγ\\gamma\. Therefore,
cL\+1≤C⋆LγαLtαL−aL\.\\displaystyle c\_\{L\+1\}\\leq C\_\{\\star\}L\_\{\\gamma\}^\{\\,\\alpha\_\{L\}\}t^\{\\alpha\_\{L\}\-a^\{L\}\}\.\(107\)By \([92](https://arxiv.org/html/2608.13882#S7.E92)\), one hasaL<αLa^\{L\}<\\alpha\_\{L\}\. Hence the right\-hand side of \([107](https://arxiv.org/html/2608.13882#S7.E107)\) converges to zero ast↓0t\\downarrow 0, whereascL\+1\>0c\_\{L\+1\}\>0\. This contradiction provesvL\+1∉𝒱\(L\)v\_\{L\+1\}\\notin\\mathcal\{V\}^\{\(L\)\}\. Together with \([98](https://arxiv.org/html/2608.13882#S7.E98)\), we conclude that𝒱\(L\)⊊𝒱\(L\+1\)\\mathcal\{V\}^\{\(L\)\}\\subsetneq\\mathcal\{V\}^\{\(L\+1\)\}\.
### 7\.4Proof of[Theorem6](https://arxiv.org/html/2608.13882#Thmtheorem6)
FixL≥2L\\geq 2,m≥1m\\geq 1,G≥2G\\geq 2,R\>0R\>0,n≥1n\\geq 1, and sample points𝐱1,…,𝐱n∈𝒳\\mathbf\{x\}\_\{1\},\\ldots,\\mathbf\{x\}\_\{n\}\\in\\mathcal\{X\}\. For notational convenience, set𝒜:=𝒜L−1m,G\\mathcal\{A\}:=\\mathcal\{A\}\_\{L\-1\}^\{m,G\},𝒟:=𝒟Lm,G\\mathcal\{D\}:=\\mathcal\{D\}\_\{L\}^\{m,G\},𝒲:=𝒲RL;m,G\\mathcal\{W\}:=\\mathcal\{W\}\_\{R\}^\{L;m,G\},
P\\displaystyle P:=PL−1,m,G,\\displaystyle:=P\_\{L\-1,m,G\},\(108\)Γ\\displaystyle\\Gamma:=ΓL−1,m,G,n=1\+\(P\+1\)ln\(e\(P\+1\)\)ln\(en\),\\displaystyle:=\\Gamma\_\{L\-1,m,G,n\}=1\+\\left\(P\+1\\right\)\\ln\\left\(e\\left\(P\+1\\right\)\\right\)\\ln\\left\(en\\right\),\(109\)B𝐱\\displaystyle B\_\{\\mathbf\{x\}\}:=supa∈𝒜maxi∈\[n\]\|a\(𝐱i\)\|\.\\displaystyle:=\\sup\_\{a\\in\\mathcal\{A\}\}\\max\_\{i\\in\\left\[n\\right\]\}\\left\|a\\left\(\\mathbf\{x\}\_\{i\}\\right\)\\right\|\.\(110\)The quantities in \([108](https://arxiv.org/html/2608.13882#S7.E108)\), \([109](https://arxiv.org/html/2608.13882#S7.E109)\), and \([110](https://arxiv.org/html/2608.13882#S7.E110)\) agree with the quantities introduced in \([52](https://arxiv.org/html/2608.13882#S3.E52)\), \([53](https://arxiv.org/html/2608.13882#S3.E53)\), and \([54](https://arxiv.org/html/2608.13882#S3.E54)\), respectively\. We divide the proof into five steps\. By \([38](https://arxiv.org/html/2608.13882#S2.E38)\) and \([37](https://arxiv.org/html/2608.13882#S2.E37)\), the class𝒲\\mathcal\{W\}is the variation hull of radiusRRgenerated by the symmetric dictionary𝒟\\mathcal\{D\}\. Therefore, \([261](https://arxiv.org/html/2608.13882#A4.E261)\) of[Lemma19](https://arxiv.org/html/2608.13882#Thmtheorem19)gives
ℜ^n\(𝒲\)=Rℜ^n\(𝒟\)\.\\displaystyle\\widehat\{\\mathfrak\{R\}\}\_\{n\}\\left\(\\mathcal\{W\}\\right\)=R\\widehat\{\\mathfrak\{R\}\}\_\{n\}\\left\(\\mathcal\{D\}\\right\)\.\(111\)The second inequality in \([261](https://arxiv.org/html/2608.13882#A4.E261)\) in[Lemma19](https://arxiv.org/html/2608.13882#Thmtheorem19)gives the elementary envelope estimate
ℜ^n\(𝒲\)≤RB𝐱1/2\.\\displaystyle\\widehat\{\\mathfrak\{R\}\}\_\{n\}\\left\(\\mathcal\{W\}\\right\)\\leq RB\_\{\\mathbf\{x\}\}^\{1/2\}\.\(112\)This estimate will provide the first term in the minimum appearing in \([55](https://arxiv.org/html/2608.13882#S3.E55)\)\. By \([289](https://arxiv.org/html/2608.13882#A4.E289)\) of[Lemma20](https://arxiv.org/html/2608.13882#Thmtheorem20), the outer dictionary satisfies
𝒟=𝒟1\(𝒜\)\.\\displaystyle\\mathcal\{D\}=\\mathcal\{D\}\_\{1\}\\left\(\\mathcal\{A\}\\right\)\.\(113\)For everya∈𝒜a\\in\\mathcal\{A\}andi∈\[n\]i\\in\\left\[n\\right\], the Brownian pullback kernel satisfies
ka\(𝐱i,𝐱i\)\\displaystyle k\_\{a\}\\left\(\\mathbf\{x\}\_\{i\},\\mathbf\{x\}\_\{i\}\\right\)=\(a\)k\(B\)\(a\(𝐱i\),a\(𝐱i\)\)=\(b\)\|a\(𝐱i\)\|≤\(c\)B𝐱\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}k^\{\(\\mathrm\{B\}\)\}\\left\(a\\left\(\\mathbf\{x\}\_\{i\}\\right\),a\\left\(\\mathbf\{x\}\_\{i\}\\right\)\\right\)\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\left\|a\\left\(\\mathbf\{x\}\_\{i\}\\right\)\\right\|\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{\\leq\}\}B\_\{\\mathbf\{x\}\}\.\(114\)Here, \(a\) applies \([36](https://arxiv.org/html/2608.13882#S2.E36)\), \(b\) usesk\(B\)\(s,s\)=\|s\|k^\{\(\\mathrm\{B\}\)\}\\left\(s,s\\right\)=\\left\|s\\right\|,s∈ℝs\\in\\mathbb\{R\}, and \(c\) applies \([110](https://arxiv.org/html/2608.13882#S7.E110)\)\. Taking the maximum overi∈\[n\]i\\in\\left\[n\\right\]and then the supremum overa∈𝒜a\\in\\mathcal\{A\}in \([114](https://arxiv.org/html/2608.13882#S7.E114)\) gives
B𝐱\(𝒜\)\\displaystyle B\_\{\\mathbf\{x\}\}\\left\(\\mathcal\{A\}\\right\)=\(a\)supa∈𝒜maxi∈\[n\]ka\(𝐱i,𝐱i\)=\(b\)B𝐱\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\sup\_\{a\\in\\mathcal\{A\}\}\\max\_\{i\\in\\left\[n\\right\]\}k\_\{a\}\\left\(\\mathbf\{x\}\_\{i\},\\mathbf\{x\}\_\{i\}\\right\)\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}B\_\{\\mathbf\{x\}\}\.\(115\)Here, \(a\) applies \([283](https://arxiv.org/html/2608.13882#A4.E283)\), while \(b\) follows from the pointwise identity in \([114](https://arxiv.org/html/2608.13882#S7.E114)\)\. Applying \([287](https://arxiv.org/html/2608.13882#A4.E287)\) of[Lemma20](https://arxiv.org/html/2608.13882#Thmtheorem20)withr=1r=1,𝒜=𝒜L−1m,G\\mathcal\{A\}=\\mathcal\{A\}\_\{L\-1\}^\{m,G\}, and using \([113](https://arxiv.org/html/2608.13882#S7.E113)\) and \([115](https://arxiv.org/html/2608.13882#S7.E115)\), we obtain
ℜ^n\(𝒟\)≤1n\(2ℭ^n,𝐱\(B\)\(𝒜\)\+B𝐱\)1/2\.\\displaystyle\\widehat\{\\mathfrak\{R\}\}\_\{n\}\\left\(\\mathcal\{D\}\\right\)\\leq\\frac\{1\}\{\\sqrt\{n\}\}\\left\(2\\widehat\{\\mathfrak\{C\}\}\_\{n,\\mathbf\{x\}\}^\{\(B\)\}\\left\(\\mathcal\{A\}\\right\)\+B\_\{\\mathbf\{x\}\}\\right\)^\{1/2\}\.\(116\)By \([428](https://arxiv.org/html/2608.13882#A4.E428)\) of[Lemma25](https://arxiv.org/html/2608.13882#Thmtheorem25)and the abbreviations \([109](https://arxiv.org/html/2608.13882#S7.E109)\) and \([110](https://arxiv.org/html/2608.13882#S7.E110)\), there exists a constantC~L\>0\\widetilde\{C\}\_\{L\}\>0, depending only onLL, such that
ℭ^n,𝐱\(B\)\(𝒜\)≤C~LB𝐱Γ\.\\displaystyle\\widehat\{\\mathfrak\{C\}\}\_\{n,\\mathbf\{x\}\}^\{\(B\)\}\\left\(\\mathcal\{A\}\\right\)\\leq\\widetilde\{C\}\_\{L\}B\_\{\\mathbf\{x\}\}\\Gamma\.\(117\)Substituting \([117](https://arxiv.org/html/2608.13882#S7.E117)\) into \([116](https://arxiv.org/html/2608.13882#S7.E116)\) and factoring the nonnegative quantityB𝐱B\_\{\\mathbf\{x\}\}gives
ℜ^n\(𝒟\)\\displaystyle\\widehat\{\\mathfrak\{R\}\}\_\{n\}\\left\(\\mathcal\{D\}\\right\)≤1n\(2C~LB𝐱Γ\+B𝐱\)1/2=B𝐱1/2n\(2C~LΓ\+1\)1/2\.\\displaystyle\\leq\\frac\{1\}\{\\sqrt\{n\}\}\\left\(2\\widetilde\{C\}\_\{L\}B\_\{\\mathbf\{x\}\}\\Gamma\+B\_\{\\mathbf\{x\}\}\\right\)^\{1/2\}=\\frac\{B\_\{\\mathbf\{x\}\}^\{1/2\}\}\{\\sqrt\{n\}\}\\left\(2\\widetilde\{C\}\_\{L\}\\Gamma\+1\\right\)^\{1/2\}\.\(118\)The identity remains valid whenB𝐱=0B\_\{\\mathbf\{x\}\}=0\. By \([109](https://arxiv.org/html/2608.13882#S7.E109)\),
Γ\\displaystyle\\Gamma≥1\.\\displaystyle\\geq 1\.\(119\)Hence
2C~LΓ\+1\\displaystyle 2\\widetilde\{C\}\_\{L\}\\Gamma\+1≤\(2C~L\+1\)Γ\.\\displaystyle\\leq\\left\(2\\widetilde\{C\}\_\{L\}\+1\\right\)\\Gamma\.\(120\)Define
CL\\displaystyle C\_\{L\}:=\(2C~L\+1\)1/2\.\\displaystyle:=\\left\(2\\widetilde\{C\}\_\{L\}\+1\\right\)^\{1/2\}\.\(121\)Combining the preceding estimates yields
ℜ^n\(𝒟\)\\displaystyle\\widehat\{\\mathfrak\{R\}\}\_\{n\}\\left\(\\mathcal\{D\}\\right\)≤CLB𝐱1/2\(Γn\)1/2\.\\displaystyle\\leq C\_\{L\}B\_\{\\mathbf\{x\}\}^\{1/2\}\\left\(\\frac\{\\Gamma\}\{n\}\\right\)^\{1/2\}\.\(122\)The exact variation reduction \([111](https://arxiv.org/html/2608.13882#S7.E111)\) therefore gives
ℜ^n\(𝒲\)\\displaystyle\\widehat\{\\mathfrak\{R\}\}\_\{n\}\\left\(\\mathcal\{W\}\\right\)=Rℜ^n\(𝒟\)≤CLRB𝐱1/2\(Γn\)1/2\.\\displaystyle=R\\widehat\{\\mathfrak\{R\}\}\_\{n\}\\left\(\\mathcal\{D\}\\right\)\\leq C\_\{L\}RB\_\{\\mathbf\{x\}\}^\{1/2\}\\left\(\\frac\{\\Gamma\}\{n\}\\right\)^\{1/2\}\.\(123\)We have therefore established both
ℜ^n\(𝒲\)\\displaystyle\\widehat\{\\mathfrak\{R\}\}\_\{n\}\\left\(\\mathcal\{W\}\\right\)≤RB𝐱1/2\\displaystyle\\leq RB\_\{\\mathbf\{x\}\}^\{1/2\}\(124\)from \([112](https://arxiv.org/html/2608.13882#S7.E112)\), and
ℜ^n\(𝒲\)\\displaystyle\\widehat\{\\mathfrak\{R\}\}\_\{n\}\\left\(\\mathcal\{W\}\\right\)≤CLRB𝐱1/2\(Γn\)1/2\\displaystyle\\leq C\_\{L\}RB\_\{\\mathbf\{x\}\}^\{1/2\}\\left\(\\frac\{\\Gamma\}\{n\}\\right\)^\{1/2\}\(125\)from \([123](https://arxiv.org/html/2608.13882#S7.E123)\)\. Combining \([124](https://arxiv.org/html/2608.13882#S7.E124)\) and \([125](https://arxiv.org/html/2608.13882#S7.E125)\), and factoring the common nonnegative term, gives
ℜ^n\(𝒲\)\\displaystyle\\widehat\{\\mathfrak\{R\}\}\_\{n\}\\left\(\\mathcal\{W\}\\right\)≤RB𝐱1/2min\{1,CL\(Γn\)1/2\}\.\\displaystyle\\leq RB\_\{\\mathbf\{x\}\}^\{1/2\}\\min\\left\\\{1,\\;C\_\{L\}\\left\(\\frac\{\\Gamma\}\{n\}\\right\)^\{1/2\}\\right\\\}\.\(126\)Substituting𝒲=𝒲RL;m,G\\mathcal\{W\}=\\mathcal\{W\}\_\{R\}^\{L;m,G\},Γ=ΓL−1,m,G,n\\Gamma=\\Gamma\_\{L\-1,m,G,n\}in \([126](https://arxiv.org/html/2608.13882#S7.E126)\) proves \([55](https://arxiv.org/html/2608.13882#S3.E55)\)\. By[Lemma17](https://arxiv.org/html/2608.13882#Thmtheorem17)[2](https://arxiv.org/html/2608.13882#A4.I1.i2), everya∈𝒜a\\in\\mathcal\{A\}satisfies
\|a\(𝐱\)\|\\displaystyle\\left\|a\\left\(\\mathbf\{x\}\\right\)\\right\|≤A𝒳2−\(L−2\),𝐱∈𝒳\.\\displaystyle\\leq A\_\{\\mathcal\{X\}\}^\{\\,2^\{\-\(L\-2\)\}\},\\qquad\\mathbf\{x\}\\in\\mathcal\{X\}\.\(127\)Consequently,
B𝐱\\displaystyle B\_\{\\mathbf\{x\}\}=\(a\)supa∈𝒜maxi∈\[n\]\|a\(𝐱i\)\|≤\(b\)A𝒳2−\(L−2\)\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\sup\_\{a\\in\\mathcal\{A\}\}\\max\_\{i\\in\\left\[n\\right\]\}\\left\|a\\left\(\\mathbf\{x\}\_\{i\}\\right\)\\right\|\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}A\_\{\\mathcal\{X\}\}^\{\\,2^\{\-\(L\-2\)\}\}\.\(128\)Here, \(a\) applies \([110](https://arxiv.org/html/2608.13882#S7.E110)\), and \(b\) applies \([127](https://arxiv.org/html/2608.13882#S7.E127)\) to every sample point\. This proves \([56](https://arxiv.org/html/2608.13882#S3.E56)\)\. SinceA𝒳≥1A\_\{\\mathcal\{X\}\}\\geq 1, and thereforeA𝒳\>0A\_\{\\mathcal\{X\}\}\>0, taking square roots in \([128](https://arxiv.org/html/2608.13882#S7.E128)\) gives
B𝐱1/2\\displaystyle B\_\{\\mathbf\{x\}\}^\{1/2\}≤\(a\)\(A𝒳2−\(L−2\)\)1/2=\(b\)A𝒳2−\(L−1\)\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}\\left\(A\_\{\\mathcal\{X\}\}^\{\\,2^\{\-\(L\-2\)\}\}\\right\)^\{1/2\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}A\_\{\\mathcal\{X\}\}^\{\\,2^\{\-\(L\-1\)\}\}\.\(129\)Here, \(a\) uses the monotonicity of the square\-root function on\[0,∞\)\\left\[0,\\infty\\right\), while \(b\) uses122−\(L−2\)=2−\(L−1\)\\frac\{1\}\{2\}2^\{\-\(L\-2\)\}\\allowbreak=\\allowbreak 2^\{\-\(L\-1\)\}\. LetX1,…,XnX\_\{1\},\\ldots,X\_\{n\}be independent random variables with common distributionν\\nu\. Applying \([126](https://arxiv.org/html/2608.13882#S7.E126)\) to the realized sample and then using \([129](https://arxiv.org/html/2608.13882#S7.E129)\) gives
ℜ^n\(𝒲RL;m,G\)\\displaystyle\\widehat\{\\mathfrak\{R\}\}\_\{n\}\\left\(\\mathcal\{W\}\_\{R\}^\{L;m,G\}\\right\)≤\(a\)RB𝐗1/2min\{1,CL\(ΓL−1,m,G,nn\)1/2\}\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}RB\_\{\\mathbf\{X\}\}^\{1/2\}\\min\\left\\\{1,\\;C\_\{L\}\\left\(\\frac\{\\Gamma\_\{L\-1,m,G,n\}\}\{n\}\\right\)^\{1/2\}\\right\\\}≤\(b\)RA𝒳2−\(L−1\)min\{1,CL\(ΓL−1,m,G,nn\)1/2\}\.\\displaystyle\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}RA\_\{\\mathcal\{X\}\}^\{\\,2^\{\-\(L\-1\)\}\}\\min\\left\\\{1,\\;C\_\{L\}\\left\(\\frac\{\\Gamma\_\{L\-1,m,G,n\}\}\{n\}\\right\)^\{1/2\}\\right\\\}\.\(130\)Here, \(a\) applies \([126](https://arxiv.org/html/2608.13882#S7.E126)\) to the random sample, and \(b\) applies \([129](https://arxiv.org/html/2608.13882#S7.E129)\)\. The right\-hand side of \([130](https://arxiv.org/html/2608.13882#S7.E130)\) is deterministic\. Taking expectation with respect toX1,…,XnX\_\{1\},\\ldots,X\_\{n\}and applying \([51](https://arxiv.org/html/2608.13882#S3.E51)\) therefore yields
ℜn\(𝒲RL;m,G\)\\displaystyle\\mathfrak\{R\}\_\{n\}\\left\(\\mathcal\{W\}\_\{R\}^\{L;m,G\}\\right\)=\(a\)𝔼𝐗\[ℜ^n\(𝒲RL;m,G\)\]≤\(b\)𝔼𝐗\[RA𝒳2−\(L−1\)min\{1,CL\(ΓL−1,m,G,nn\)1/2\}\]\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\mathbb\{E\}\_\{\\mathbf\{X\}\}\\left\[\\widehat\{\\mathfrak\{R\}\}\_\{n\}\\left\(\\mathcal\{W\}\_\{R\}^\{L;m,G\}\\right\)\\right\]\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}\\mathbb\{E\}\_\{\\mathbf\{X\}\}\\left\[RA\_\{\\mathcal\{X\}\}^\{\\,2^\{\-\(L\-1\)\}\}\\min\\left\\\{1,\\;C\_\{L\}\\left\(\\frac\{\\Gamma\_\{L\-1,m,G,n\}\}\{n\}\\right\)^\{1/2\}\\right\\\}\\right\]=\(c\)RA𝒳2−\(L−1\)min\{1,CL\(ΓL−1,m,G,nn\)1/2\}\.\\displaystyle\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{=\}\}RA\_\{\\mathcal\{X\}\}^\{\\,2^\{\-\(L\-1\)\}\}\\min\\left\\\{1,\\;C\_\{L\}\\left\(\\frac\{\\Gamma\_\{L\-1,m,G,n\}\}\{n\}\\right\)^\{1/2\}\\right\\\}\.\(131\)Here, \(a\) applies the definition of expected Rademacher complexity, \(b\) applies \([130](https://arxiv.org/html/2608.13882#S7.E130)\), and \(c\) uses the fact that every quantity inside the expectation is deterministic\. Equation \([131](https://arxiv.org/html/2608.13882#S7.E131)\) is precisely \([57](https://arxiv.org/html/2608.13882#S3.E57)\)\. This completes the proof\.
### 7\.5Proof of[Theorem9](https://arxiv.org/html/2608.13882#Thmtheorem9)
Set
RL:=\(∫𝒳‖𝐱‖22−\(L−2\)𝑑ν\(𝐱\)\)1/2\.\\displaystyle R\_\{L\}:=\\left\(\\int\_\{\\mathcal\{X\}\}\\left\\\|\\mathbf\{x\}\\right\\\|\_\{2\}^\{\\,2^\{\-\(L\-2\)\}\}\\,\\mathrm\{d\}\\nu\\left\(\\mathbf\{x\}\\right\)\\right\)^\{1/2\}\.\(132\)By[Lemma26](https://arxiv.org/html/2608.13882#Thmtheorem26), every atomu∈𝒰Lu\\in\\mathcal\{U\}\_\{L\}satisfies
‖u‖L2\(ν\)≤RL\.\\displaystyle\\left\\\|u\\right\\\|\_\{L^\{2\}\\left\(\\nu\\right\)\}\\leq R\_\{L\}\.\(133\)Moreover,[Lemma29](https://arxiv.org/html/2608.13882#Thmtheorem29)shows that𝒰L\\mathcal\{U\}\_\{L\}is compact inL2\(ν\)L^\{2\}\\left\(\\nu\\right\)\. Therefore,[Lemma27](https://arxiv.org/html/2608.13882#Thmtheorem27)provides a finite signed Radon measureμF\\mu\_\{F\}on𝒰L\\mathcal\{U\}\_\{L\}such that
F\\displaystyle F=∫𝒰LudμF\(u\)inL2\(ν\),\\displaystyle=\\int\_\{\\mathcal\{U\}\_\{L\}\}u\\,\\mathrm\{d\}\\mu\_\{F\}\\left\(u\\right\)\\qquad\\text\{in \}L^\{2\}\\left\(\\nu\\right\),\(134\)‖μF‖TV\\displaystyle\\left\\\|\\mu\_\{F\}\\right\\\|\_\{\\mathrm\{TV\}\}=𝒞^var\(L\)\(F\)\.\\displaystyle=\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}\\left\(F\\right\)\.\(135\)Suppose first that𝒞^var\(L\)\(F\)=0\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}\\left\(F\\right\)=0\. Then \([135](https://arxiv.org/html/2608.13882#S7.E135)\) gives‖μF‖TV=0\\left\\\|\\mu\_\{F\}\\right\\\|\_\{\\mathrm\{TV\}\}=0, and henceμF=0\\mu\_\{F\}=0\. Substituting this identity into \([134](https://arxiv.org/html/2608.13882#S7.E134)\) givesF=0F=0inL2\(ν\)L^\{2\}\\left\(\\nu\\right\)\. ChoosingFm:=0F\_\{m\}:=0givesFm∈𝒱m\(L\)F\_\{m\}\\in\\mathcal\{V\}\_\{m\}^\{\(L\)\}through the zero measure, whose support is empty, and proves the claimed estimate in this case\. We may therefore assume that
ρ:=‖μF‖TV=𝒞^var\(L\)\(F\)\>0\.\\displaystyle\\rho:=\\left\\\|\\mu\_\{F\}\\right\\\|\_\{\\mathrm\{TV\}\}=\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}\\left\(F\\right\)\>0\.\(136\)Let\|μF\|\\left\|\\mu\_\{F\}\\right\|denote the total\-variation measure ofμF\\mu\_\{F\}\. By the polar decomposition of finite real signed measures, there exists a measurable function
σ:𝒰L⟶\{−1,1\}\\displaystyle\\sigma:\\mathcal\{U\}\_\{L\}\\longrightarrow\\left\\\{\-1,1\\right\\\}\(137\)such that
dμF\(u\)=σ\(u\)d\|μF\|\(u\)\.\\displaystyle\\mathrm\{d\}\\mu\_\{F\}\\left\(u\\right\)=\\sigma\\left\(u\\right\)\\,\\mathrm\{d\}\\left\|\\mu\_\{F\}\\right\|\\left\(u\\right\)\.\(138\)The functionσ\\sigmamay be defined arbitrarily on a\|μF\|\\left\|\\mu\_\{F\}\\right\|\-null set so that \([137](https://arxiv.org/html/2608.13882#S7.E137)\) holds on all of𝒰L\\mathcal\{U\}\_\{L\}\. Define
PF:=\|μF\|ρ\.\\displaystyle P\_\{F\}:=\\frac\{\\left\|\\mu\_\{F\}\\right\|\}\{\\rho\}\.\(139\)ThenPFP\_\{F\}is a probability measure on𝒰L\\mathcal\{U\}\_\{L\}because
PF\(𝒰L\)\\displaystyle P\_\{F\}\\left\(\\mathcal\{U\}\_\{L\}\\right\)=\(a\)\|μF\|\(𝒰L\)ρ=\(b\)‖μF‖TVρ=\(c\)1\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\frac\{\\left\|\\mu\_\{F\}\\right\|\\left\(\\mathcal\{U\}\_\{L\}\\right\)\}\{\\rho\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\frac\{\\left\\\|\\mu\_\{F\}\\right\\\|\_\{\\mathrm\{TV\}\}\}\{\\rho\}\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{=\}\}1\.\(140\)Here, \(a\) applies \([139](https://arxiv.org/html/2608.13882#S7.E139)\), \(b\) applies the definition of the total\-variation norm, and \(c\) applies \([136](https://arxiv.org/html/2608.13882#S7.E136)\)\. LetU1,…,UmU\_\{1\},\\ldots,U\_\{m\}be independent𝒰L\\mathcal\{U\}\_\{L\}\-valued random variables with common lawPFP\_\{F\}, and define
Zi:=σ\(Ui\)Ui,i∈\[m\]\.\\displaystyle Z\_\{i\}:=\\sigma\\left\(U\_\{i\}\\right\)U\_\{i\},\\qquad i\\in\\left\[m\\right\]\.\(141\)Since𝒰L\\mathcal\{U\}\_\{L\}is compact inL2\(ν\)L^\{2\}\\left\(\\nu\\right\), it is separable\. The mapu⟼σ\(u\)uu\\longmapsto\\sigma\\left\(u\\right\)uis measurable and has its range in the separable set𝒰L∪\(−𝒰L\)\\mathcal\{U\}\_\{L\}\\cup\\left\(\-\\mathcal\{U\}\_\{L\}\\right\)\. Consequently, eachZiZ\_\{i\}is strongly measurable as anL2\(ν\)L^\{2\}\\left\(\\nu\\right\)\-valued random variable\. Moreover, \([133](https://arxiv.org/html/2608.13882#S7.E133)\) gives
‖Zi‖L2\(ν\)\\displaystyle\\left\\\|Z\_\{i\}\\right\\\|\_\{L^\{2\}\\left\(\\nu\\right\)\}=\(a\)\|σ\(Ui\)\|‖Ui‖L2\(ν\)=\(b\)‖Ui‖L2\(ν\)≤\(c\)RLalmost surely\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\left\|\\sigma\\left\(U\_\{i\}\\right\)\\right\|\\left\\\|U\_\{i\}\\right\\\|\_\{L^\{2\}\\left\(\\nu\\right\)\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\left\\\|U\_\{i\}\\right\\\|\_\{L^\{2\}\\left\(\\nu\\right\)\}\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{\\leq\}\}R\_\{L\}\\qquad\\text\{almost surely\}\.\(142\)Here, \(a\) applies the absolute homogeneity of theL2\(ν\)L^\{2\}\\left\(\\nu\\right\)norm, \(b\) uses\|σ\(Ui\)\|=1\\left\|\\sigma\\left\(U\_\{i\}\\right\)\\right\|=1, and \(c\) applies \([133](https://arxiv.org/html/2608.13882#S7.E133)\)\. Thus,ZiZ\_\{i\}is Bochner integrable and square\-integrable\. The Bochner expectation ofZiZ\_\{i\}satisfies
𝔼\[Zi\]\\displaystyle\\mathbb\{E\}\\left\[Z\_\{i\}\\right\]=\(a\)∫𝒰Lσ\(u\)udPF\(u\)=\(b\)1ρ∫𝒰Lσ\(u\)ud\|μF\|\(u\)=\(c\)1ρ∫𝒰LudμF\(u\)\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\int\_\{\\mathcal\{U\}\_\{L\}\}\\sigma\\left\(u\\right\)u\\,\\mathrm\{d\}P\_\{F\}\\left\(u\\right\)\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\frac\{1\}\{\\rho\}\\int\_\{\\mathcal\{U\}\_\{L\}\}\\sigma\\left\(u\\right\)u\\,\\mathrm\{d\}\\left\|\\mu\_\{F\}\\right\|\\left\(u\\right\)\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{=\}\}\\frac\{1\}\{\\rho\}\\int\_\{\\mathcal\{U\}\_\{L\}\}u\\,\\mathrm\{d\}\\mu\_\{F\}\\left\(u\\right\)=\(d\)FρinL2\(ν\)\.\\displaystyle\\stackrel\{\{\\scriptstyle\(d\)\}\}\{\{=\}\}\\frac\{F\}\{\\rho\}\\qquad\\text\{in \}L^\{2\}\\left\(\\nu\\right\)\.\(143\)Here, \(a\) uses the law ofUiU\_\{i\}and \([141](https://arxiv.org/html/2608.13882#S7.E141)\), \(b\) substitutes \([139](https://arxiv.org/html/2608.13882#S7.E139)\), \(c\) applies \([138](https://arxiv.org/html/2608.13882#S7.E138)\), and \(d\) applies \([134](https://arxiv.org/html/2608.13882#S7.E134)\)\. Define the random finite approximation
F~m:=ρm∑i=1mZi\.\\displaystyle\\widetilde\{F\}\_\{m\}:=\\frac\{\\rho\}\{m\}\\sum\_\{i=1\}^\{m\}Z\_\{i\}\.\(144\)By \([143](https://arxiv.org/html/2608.13882#S7.E143)\),
𝔼\[F~m\]\\displaystyle\\mathbb\{E\}\\left\[\\widetilde\{F\}\_\{m\}\\right\]=\(a\)ρm∑i=1m𝔼\[Zi\]=\(b\)ρm∑i=1mFρ=\(c\)FinL2\(ν\)\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\frac\{\\rho\}\{m\}\\sum\_\{i=1\}^\{m\}\\mathbb\{E\}\\left\[Z\_\{i\}\\right\]\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\frac\{\\rho\}\{m\}\\sum\_\{i=1\}^\{m\}\\frac\{F\}\{\\rho\}\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{=\}\}F\\qquad\\text\{in \}L^\{2\}\\left\(\\nu\\right\)\.\(145\)Here, \(a\) uses the linearity of the Bochner expectation, \(b\) applies \([143](https://arxiv.org/html/2608.13882#S7.E143)\), and \(c\) evaluates the sum\. Set
Yi:=Zi−Fρ,i∈\[m\]\.\\displaystyle Y\_\{i\}:=Z\_\{i\}\-\\frac\{F\}\{\\rho\},\\qquad i\\in\\left\[m\\right\]\.\(146\)ThenY1,…,YmY\_\{1\},\\ldots,Y\_\{m\}are independent, square\-integrable, and centered:
𝔼\[Yi\]\\displaystyle\\mathbb\{E\}\\left\[Y\_\{i\}\\right\]=\(a\)𝔼\[Zi\]−Fρ=\(b\)0\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\mathbb\{E\}\\left\[Z\_\{i\}\\right\]\-\\frac\{F\}\{\\rho\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}0\.\(147\)Here, \(a\) applies \([146](https://arxiv.org/html/2608.13882#S7.E146)\), and \(b\) applies \([143](https://arxiv.org/html/2608.13882#S7.E143)\)\. Moreover,
F−F~m\\displaystyle F\-\\widetilde\{F\}\_\{m\}=\(a\)F−ρm∑i=1mZi=\(b\)−ρm∑i=1m\(Zi−Fρ\)=\(c\)−ρm∑i=1mYi\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}F\-\\frac\{\\rho\}\{m\}\\sum\_\{i=1\}^\{m\}Z\_\{i\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\-\\frac\{\\rho\}\{m\}\\sum\_\{i=1\}^\{m\}\\left\(Z\_\{i\}\-\\frac\{F\}\{\\rho\}\\right\)\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{=\}\}\-\\frac\{\\rho\}\{m\}\\sum\_\{i=1\}^\{m\}Y\_\{i\}\.\(148\)Here, \(a\) applies \([144](https://arxiv.org/html/2608.13882#S7.E144)\), \(b\) usesF=ρm∑i=1mFρF=\\frac\{\\rho\}\{m\}\\sum\_\{i=1\}^\{m\}\\frac\{F\}\{\\rho\}, and \(c\) applies \([146](https://arxiv.org/html/2608.13882#S7.E146)\)\. Fori≠ji\\neq j, independence and \([147](https://arxiv.org/html/2608.13882#S7.E147)\) give
𝔼\[⟨Yi,Yj⟩L2\(ν\)\]\\displaystyle\\mathbb\{E\}\\left\[\\left\\langle Y\_\{i\},Y\_\{j\}\\right\\rangle\_\{L^\{2\}\\left\(\\nu\\right\)\}\\right\]=\(a\)⟨𝔼\[Yi\],𝔼\[Yj\]⟩L2\(ν\)=\(b\)0\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\left\\langle\\mathbb\{E\}\\left\[Y\_\{i\}\\right\],\\mathbb\{E\}\\left\[Y\_\{j\}\\right\]\\right\\rangle\_\{L^\{2\}\\left\(\\nu\\right\)\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}0\.\(149\)Here, \(a\) follows from independence and Fubini’s theorem for the square\-integrable Hilbert\-space\-valued variables, and \(b\) applies \([147](https://arxiv.org/html/2608.13882#S7.E147)\)\. Using \([148](https://arxiv.org/html/2608.13882#S7.E148)\), we therefore obtain
𝔼\[‖F−F~m‖L2\(ν\)2\]=\(a\)ρ2m2𝔼\[‖∑i=1mYi‖L2\(ν\)2\]\\displaystyle\\mathbb\{E\}\\left\[\\left\\\|F\-\\widetilde\{F\}\_\{m\}\\right\\\|\_\{L^\{2\}\\left\(\\nu\\right\)\}^\{2\}\\right\]\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\frac\{\\rho^\{2\}\}\{m^\{2\}\}\\mathbb\{E\}\\left\[\\left\\\|\\sum\_\{i=1\}^\{m\}Y\_\{i\}\\right\\\|\_\{L^\{2\}\\left\(\\nu\\right\)\}^\{2\}\\right\]=\(b\)ρ2m2\[∑i=1m𝔼\[‖Yi‖L2\(ν\)2\]\+2∑1≤i<j≤m𝔼\[⟨Yi,Yj⟩L2\(ν\)\]\]=\(c\)ρ2m2∑i=1m𝔼\[‖Yi‖L2\(ν\)2\]\\displaystyle\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\frac\{\\rho^\{2\}\}\{m^\{2\}\}\\left\[\\sum\_\{i=1\}^\{m\}\\mathbb\{E\}\\left\[\\left\\\|Y\_\{i\}\\right\\\|\_\{L^\{2\}\\left\(\\nu\\right\)\}^\{2\}\\right\]\+2\\sum\_\{1\\leq i<j\\leq m\}\\mathbb\{E\}\\left\[\\left\\langle Y\_\{i\},Y\_\{j\}\\right\\rangle\_\{L^\{2\}\\left\(\\nu\\right\)\}\\right\]\\right\]\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{=\}\}\\frac\{\\rho^\{2\}\}\{m^\{2\}\}\\sum\_\{i=1\}^\{m\}\\mathbb\{E\}\\left\[\\left\\\|Y\_\{i\}\\right\\\|\_\{L^\{2\}\\left\(\\nu\\right\)\}^\{2\}\\right\]=\(d\)ρ2m2∑i=1m\[𝔼\[‖Zi‖L2\(ν\)2\]−‖𝔼\[Zi\]‖L2\(ν\)2\]≤\(e\)ρ2m2∑i=1m𝔼\[‖Zi‖L2\(ν\)2\]\\displaystyle\\stackrel\{\{\\scriptstyle\(d\)\}\}\{\{=\}\}\\frac\{\\rho^\{2\}\}\{m^\{2\}\}\\sum\_\{i=1\}^\{m\}\\left\[\\mathbb\{E\}\\left\[\\left\\\|Z\_\{i\}\\right\\\|\_\{L^\{2\}\\left\(\\nu\\right\)\}^\{2\}\\right\]\-\\left\\\|\\mathbb\{E\}\\left\[Z\_\{i\}\\right\]\\right\\\|\_\{L^\{2\}\\left\(\\nu\\right\)\}^\{2\}\\right\]\\stackrel\{\{\\scriptstyle\(e\)\}\}\{\{\\leq\}\}\\frac\{\\rho^\{2\}\}\{m^\{2\}\}\\sum\_\{i=1\}^\{m\}\\mathbb\{E\}\\left\[\\left\\\|Z\_\{i\}\\right\\\|\_\{L^\{2\}\\left\(\\nu\\right\)\}^\{2\}\\right\]≤\(f\)ρ2m2∑i=1mRL2=\(g\)ρ2RL2m\.\\displaystyle\\stackrel\{\{\\scriptstyle\(f\)\}\}\{\{\\leq\}\}\\frac\{\\rho^\{2\}\}\{m^\{2\}\}\\sum\_\{i=1\}^\{m\}R\_\{L\}^\{2\}\\stackrel\{\{\\scriptstyle\(g\)\}\}\{\{=\}\}\\frac\{\\rho^\{2\}R\_\{L\}^\{2\}\}\{m\}\.\(150\)Here, \(a\) applies \([148](https://arxiv.org/html/2608.13882#S7.E148)\), \(b\) expands the squared Hilbert\-space norm, \(c\) applies \([149](https://arxiv.org/html/2608.13882#S7.E149)\), \(d\) applies the Hilbert\-space variance identity to \([146](https://arxiv.org/html/2608.13882#S7.E146)\), \(e\) discards the nonnegative squared\-mean term, \(f\) applies \([142](https://arxiv.org/html/2608.13882#S7.E142)\), and \(g\) evaluates the sum\. Since the random variable‖F−F~m‖L2\(ν\)2\\left\\\|F\-\\widetilde\{F\}\_\{m\}\\right\\\|\_\{L^\{2\}\\left\(\\nu\\right\)\}^\{2\}is nonnegative, the expectation estimate \([150](https://arxiv.org/html/2608.13882#S7.E150)\) implies that there exists a realizationu1,…,um∈𝒰Lu\_\{1\},\\ldots,u\_\{m\}\\allowbreak\\in\\mathcal\{U\}\_\{L\}ofU1,…,UmU\_\{1\},\\ldots,U\_\{m\}such that
‖F−Fm‖L2\(ν\)2≤ρ2RL2m,\\displaystyle\\left\\\|F\-F\_\{m\}\\right\\\|\_\{L^\{2\}\\left\(\\nu\\right\)\}^\{2\}\\leq\\frac\{\\rho^\{2\}R\_\{L\}^\{2\}\}\{m\},\(151\)where
Fm:=ρm∑i=1mσ\(ui\)ui\.\\displaystyle F\_\{m\}:=\\frac\{\\rho\}\{m\}\\sum\_\{i=1\}^\{m\}\\sigma\\left\(u\_\{i\}\\right\)u\_\{i\}\.\(152\)Indeed, if \([151](https://arxiv.org/html/2608.13882#S7.E151)\) failed for every realization, then the expectation in \([150](https://arxiv.org/html/2608.13882#S7.E150)\) would be strictly larger thanρ2RL2/m\\rho^\{2\}R\_\{L\}^\{2\}/m\. Define the finite signed discrete measure
μm:=ρm∑i=1mσ\(ui\)δui\.\\displaystyle\\mu\_\{m\}:=\\frac\{\\rho\}\{m\}\\sum\_\{i=1\}^\{m\}\\sigma\\left\(u\_\{i\}\\right\)\\delta\_\{u\_\{i\}\}\.\(153\)By the definition ofμm\\mu\_\{m\},
Fm\\displaystyle F\_\{m\}=∫𝒰Ludμm\(u\),\\displaystyle=\\int\_\{\\mathcal\{U\}\_\{L\}\}u\\,\\mathrm\{d\}\\mu\_\{m\}\(u\),\|supp\(μm\)\|\\displaystyle\\left\|\\operatorname\{supp\}\(\\mu\_\{m\}\)\\right\|≤m,\\displaystyle\\leq m,\(154\)where repeated sampled atoms can only reduce the number of distinct support points\. Consequently,Fm∈𝒱m\(L\)F\_\{m\}\\in\\mathcal\{V\}\_\{m\}^\{\(L\)\}\. Taking square roots in \([151](https://arxiv.org/html/2608.13882#S7.E151)\) and using \([136](https://arxiv.org/html/2608.13882#S7.E136)\) gives
‖F−Fm‖L2\(ν\)≤RL𝒞^var\(L\)\(F\)m−1/2\.\\displaystyle\\left\\\|F\-F\_\{m\}\\right\\\|\_\{L^\{2\}\(\\nu\)\}\\leq R\_\{L\}\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}\(F\)m^\{\-1/2\}\.\(155\)Finally, substituting the definition ofRLR\_\{L\}into \([155](https://arxiv.org/html/2608.13882#S7.E155)\) gives
‖F−Fm‖L2\(ν\)≤\(∫𝒳‖𝐱‖22−\(L−2\)dν\(𝐱\)\)1/2𝒞^var\(L\)\(F\)m−1/2,\\displaystyle\\left\\\|F\-F\_\{m\}\\right\\\|\_\{L^\{2\}\\left\(\\nu\\right\)\}\\leq\\left\(\\int\_\{\\mathcal\{X\}\}\\left\\\|\\mathbf\{x\}\\right\\\|\_\{2\}^\{\\,2^\{\-\(L\-2\)\}\}\\,\\mathrm\{d\}\\nu\\left\(\\mathbf\{x\}\\right\)\\right\)^\{1/2\}\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}\\left\(F\\right\)m^\{\-1/2\},\(156\)which proves the claim\.
### 7\.6Proof of[Theorem10](https://arxiv.org/html/2608.13882#Thmtheorem10)
Setρ:=𝒞^var\(L\)\(F\)\\rho:=\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}\\left\(F\\right\)\. Ifρ=0\\rho=0, then[Lemma16](https://arxiv.org/html/2608.13882#Thmtheorem16)givesF=0inL2\(ν\)F=0\\qquad\\text\{in \}L^\{2\}\\left\(\\nu\\right\)\. In this case, the zero realization proves the error estimate and the local sparsity assertion\. We may therefore assume thatρ\>0\\rho\>0\. We first construct a finite atomic approximation\. By[Theorem9](https://arxiv.org/html/2608.13882#Thmtheorem9), applied with the number of atoms equal toMM, there existsFM∈𝒱M\(L\)F\_\{M\}\\in\\mathcal\{V\}\_\{M\}^\{\(L\)\}such that
‖F−FM‖L2\(ν\)≤RLρM−1/2\.\\displaystyle\\left\\\|F\-F\_\{M\}\\right\\\|\_\{L^\{2\}\\left\(\\nu\\right\)\}\\leq R\_\{L\}\\rho M^\{\-1/2\}\.\(157\)Moreover, the explicit construction in the proof of[Theorem9](https://arxiv.org/html/2608.13882#Thmtheorem9)shows thatFMF\_\{M\}may be chosen in the form
FM\(𝐱\)=ρM∑j=1Mσjuj\(𝐱\),𝐱∈𝒳,\\displaystyle F\_\{M\}\\left\(\\mathbf\{x\}\\right\)=\\frac\{\\rho\}\{M\}\\sum\_\{j=1\}^\{M\}\\sigma\_\{j\}u\_\{j\}\\left\(\\mathbf\{x\}\\right\),\\qquad\\mathbf\{x\}\\in\\mathcal\{X\},\(158\)whereu1,…,uM∈𝒰Lu\_\{1\},\\ldots,u\_\{M\}\\in\\mathcal\{U\}\_\{L\}andσ1,…,σM∈\{−1,1\}\\sigma\_\{1\},\\ldots,\\sigma\_\{M\}\\in\\left\\\{\-1,1\\right\\\}\. Fixj∈\[M\]j\\in\\left\[M\\right\]\. By the recursive unit\-ball characterization in[Lemma18](https://arxiv.org/html/2608.13882#Thmtheorem18), there existauj∈𝒰L−1a\_\{u\_\{j\}\}\\in\\mathcal\{U\}\_\{L\-1\}andguj∈ℋk\(B\)g\_\{u\_\{j\}\}\\in\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}such that
uj\(𝐱\)\\displaystyle u\_\{j\}\\left\(\\mathbf\{x\}\\right\)=guj\(auj\(𝐱\)\),𝐱∈𝒳,\\displaystyle=g\_\{u\_\{j\}\}\\left\(a\_\{u\_\{j\}\}\\left\(\\mathbf\{x\}\\right\)\\right\),\\qquad\\mathbf\{x\}\\in\\mathcal\{X\},\(159\)‖guj‖ℋk\(B\)\\displaystyle\\left\\\|g\_\{u\_\{j\}\}\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}≤1\.\\displaystyle\\leq 1\.\(160\)By[Lemma31](https://arxiv.org/html/2608.13882#Thmtheorem31),
auj\(𝒳\)⊆\[−A0,A0\]\.\\displaystyle a\_\{u\_\{j\}\}\\left\(\\mathcal\{X\}\\right\)\\subseteq\\left\[\-A\_\{0\},A\_\{0\}\\right\]\.\(161\)Sinceρ\>0\\rho\>0, the degenerate caseR𝒳=0R\_\{\\mathcal\{X\}\}=0cannot occur\. Indeed, ifR𝒳=0R\_\{\\mathcal\{X\}\}=0, then𝒳⊆\{𝟎\}\\mathcal\{X\}\\subseteq\\left\\\{\\bm\{0\}\\right\\\}, and[Lemma26](https://arxiv.org/html/2608.13882#Thmtheorem26)would imply that every atom in𝒰L\\mathcal\{U\}\_\{L\}vanishes on𝒳\\mathcal\{X\}\. This would implyF=0F=0and thereforeρ=0\\rho=0, contrary to the assumption thatρ\>0\\rho\>0\. Hence,A0\>0A\_\{0\}\>0\. Define the interpolated selected atom by
uj,m\(𝐱\):=\(Πmguj\)\(auj\(𝐱\)\),𝐱∈𝒳\.\\displaystyle u\_\{j,m\}\\left\(\\mathbf\{x\}\\right\):=\\left\(\\Pi\_\{m\}g\_\{u\_\{j\}\}\\right\)\\left\(a\_\{u\_\{j\}\}\\left\(\\mathbf\{x\}\\right\)\\right\),\\qquad\\mathbf\{x\}\\in\\mathcal\{X\}\.\(162\)Applying[Lemma32](https://arxiv.org/html/2608.13882#Thmtheorem32)withA=A0,a=auj,g=gujA=A\_\{0\},a=a\_\{u\_\{j\}\},g=g\_\{u\_\{j\}\}gives
‖uj−uj,m‖L2\(ν\)\\displaystyle\\left\\\|u\_\{j\}\-u\_\{j,m\}\\right\\\|\_\{L^\{2\}\\left\(\\nu\\right\)\}≤\(a\)\(A02\)1/2m−1/2‖guj‖ℋk\(B\)≤\(b\)\(A02\)1/2m−1/2\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}\\left\(\\frac\{A\_\{0\}\}\{2\}\\right\)^\{1/2\}m^\{\-1/2\}\\left\\\|g\_\{u\_\{j\}\}\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}\\left\(\\frac\{A\_\{0\}\}\{2\}\\right\)^\{1/2\}m^\{\-1/2\}\.\(163\)Here, \(a\) applies \([504](https://arxiv.org/html/2608.13882#A4.E504)\), using \([161](https://arxiv.org/html/2608.13882#S7.E161)\), while \(b\) applies \([160](https://arxiv.org/html/2608.13882#S7.E160)\)\. The two\-stage approximant from the theorem statement can be written as
FM,m\(𝐱\)\\displaystyle F\_\{M,m\}\\left\(\\mathbf\{x\}\\right\)=\(a\)ρM∑j=1Mσj\(Πmguj\)\(auj\(𝐱\)\)=\(b\)ρM∑j=1Mσjuj,m\(𝐱\)\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\frac\{\\rho\}\{M\}\\sum\_\{j=1\}^\{M\}\\sigma\_\{j\}\\left\(\\Pi\_\{m\}g\_\{u\_\{j\}\}\\right\)\\left\(a\_\{u\_\{j\}\}\\left\(\\mathbf\{x\}\\right\)\\right\)\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\frac\{\\rho\}\{M\}\\sum\_\{j=1\}^\{M\}\\sigma\_\{j\}u\_\{j,m\}\\left\(\\mathbf\{x\}\\right\)\.\(164\)Here, \(a\) applies \([68](https://arxiv.org/html/2608.13882#S4.E68)\), and \(b\) applies \([162](https://arxiv.org/html/2608.13882#S7.E162)\)\. Subtracting \([164](https://arxiv.org/html/2608.13882#S7.E164)\) from \([158](https://arxiv.org/html/2608.13882#S7.E158)\) gives
FM−FM,m=ρM∑j=1Mσj\(uj−uj,m\)\.\\displaystyle F\_\{M\}\-F\_\{M,m\}=\\frac\{\\rho\}\{M\}\\sum\_\{j=1\}^\{M\}\\sigma\_\{j\}\\left\(u\_\{j\}\-u\_\{j,m\}\\right\)\.\(165\)Therefore,
‖FM−FM,m‖L2\(ν\)=\(a\)‖ρM∑j=1Mσj\(uj−uj,m\)‖L2\(ν\)≤\(b\)ρM∑j=1M\|σj\|‖uj−uj,m‖L2\(ν\)\\displaystyle\\left\\\|F\_\{M\}\-F\_\{M,m\}\\right\\\|\_\{L^\{2\}\\left\(\\nu\\right\)\}\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\left\\\|\\frac\{\\rho\}\{M\}\\sum\_\{j=1\}^\{M\}\\sigma\_\{j\}\\left\(u\_\{j\}\-u\_\{j,m\}\\right\)\\right\\\|\_\{L^\{2\}\\left\(\\nu\\right\)\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}\\frac\{\\rho\}\{M\}\\sum\_\{j=1\}^\{M\}\\left\|\\sigma\_\{j\}\\right\|\\left\\\|u\_\{j\}\-u\_\{j,m\}\\right\\\|\_\{L^\{2\}\\left\(\\nu\\right\)\}=\(c\)ρM∑j=1M‖uj−uj,m‖L2\(ν\)≤\(d\)ρM∑j=1M\(A02\)1/2m−1/2=\(e\)\(A02\)1/2ρm−1/2\.\\displaystyle\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{=\}\}\\frac\{\\rho\}\{M\}\\sum\_\{j=1\}^\{M\}\\left\\\|u\_\{j\}\-u\_\{j,m\}\\right\\\|\_\{L^\{2\}\\left\(\\nu\\right\)\}\\stackrel\{\{\\scriptstyle\(d\)\}\}\{\{\\leq\}\}\\frac\{\\rho\}\{M\}\\sum\_\{j=1\}^\{M\}\\left\(\\frac\{A\_\{0\}\}\{2\}\\right\)^\{1/2\}m^\{\-1/2\}\\stackrel\{\{\\scriptstyle\(e\)\}\}\{\{=\}\}\\left\(\\frac\{A\_\{0\}\}\{2\}\\right\)^\{1/2\}\\rho m^\{\-1/2\}\.\(166\)Here, \(a\) applies \([165](https://arxiv.org/html/2608.13882#S7.E165)\), \(b\) applies the triangle inequality and absolute homogeneity inL2\(ν\)L^\{2\}\\left\(\\nu\\right\), \(c\) uses\|σj\|=1\\left\|\\sigma\_\{j\}\\right\|=1, \(d\) applies \([163](https://arxiv.org/html/2608.13882#S7.E163)\), and \(e\) evaluates the sum ofMMidentical terms\. The triangle inequality now gives
‖F−FM,m‖L2\(ν\)≤\(a\)‖F−FM‖L2\(ν\)\+‖FM−FM,m‖L2\(ν\)\\displaystyle\\left\\\|F\-F\_\{M,m\}\\right\\\|\_\{L^\{2\}\\left\(\\nu\\right\)\}\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}\\left\\\|F\-F\_\{M\}\\right\\\|\_\{L^\{2\}\\left\(\\nu\\right\)\}\+\\left\\\|F\_\{M\}\-F\_\{M,m\}\\right\\\|\_\{L^\{2\}\\left\(\\nu\\right\)\}≤\(b\)RLρM−1/2\+\(A02\)1/2ρm−1/2=\(c\)𝒞^var\(L\)\(F\)\[RLM−1/2\+\(A02\)1/2m−1/2\]\.\\displaystyle\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}R\_\{L\}\\rho M^\{\-1/2\}\+\\left\(\\frac\{A\_\{0\}\}\{2\}\\right\)^\{1/2\}\\rho m^\{\-1/2\}\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{=\}\}\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}\\left\(F\\right\)\\left\[R\_\{L\}M^\{\-1/2\}\+\\left\(\\frac\{A\_\{0\}\}\{2\}\\right\)^\{1/2\}m^\{\-1/2\}\\right\]\.\(167\)Here, \(a\) applies the triangle inequality, \(b\) combines \([157](https://arxiv.org/html/2608.13882#S7.E157)\) and \([166](https://arxiv.org/html/2608.13882#S7.E166)\), and \(c\) applies definition ofρ\\rho\. This proves \([69](https://arxiv.org/html/2608.13882#S4.E69)\)\. It remains to prove the local evaluation sparsity assertion\. Let−A0=t0<t1<⋯<tm=A0\-A\_\{0\}=t\_\{0\}<t\_\{1\}<\\cdots<t\_\{m\}=A\_\{0\}be the uniform interpolation grid underlyingΠm\\Pi\_\{m\}, and letψ0,…,ψm\\psi\_\{0\},\\ldots,\\psi\_\{m\}denote the corresponding continuous piecewise\-linear hat basis functions\. Fix𝐱∈𝒳\\mathbf\{x\}\\in\\mathcal\{X\}andj∈\[M\]j\\in\\left\[M\\right\]\. By \([161](https://arxiv.org/html/2608.13882#S7.E161)\),auj\(𝐱\)∈\[−A0,A0\]a\_\{u\_\{j\}\}\\left\(\\mathbf\{x\}\\right\)\\in\\left\[\-A\_\{0\},A\_\{0\}\\right\]\. Hence, there exists an indexqj\(𝐱\)∈\{0,…,m−1\}q\_\{j\}\\left\(\\mathbf\{x\}\\right\)\\in\\left\\\{0,\\ldots,m\-1\\right\\\}such that
auj\(𝐱\)∈\[tqj\(𝐱\),tqj\(𝐱\)\+1\]\.\\displaystyle a\_\{u\_\{j\}\}\\left\(\\mathbf\{x\}\\right\)\\in\\left\[t\_\{q\_\{j\}\\left\(\\mathbf\{x\}\\right\)\},t\_\{q\_\{j\}\\left\(\\mathbf\{x\}\\right\)\+1\}\\right\]\.\(168\)The nodal representation of the piecewise\-linear interpolant gives
\(Πmguj\)\(auj\(𝐱\)\)=∑k=0mguj\(tk\)ψk\(auj\(𝐱\)\)\.\\displaystyle\\left\(\\Pi\_\{m\}g\_\{u\_\{j\}\}\\right\)\\left\(a\_\{u\_\{j\}\}\\left\(\\mathbf\{x\}\\right\)\\right\)=\\sum\_\{k=0\}^\{m\}g\_\{u\_\{j\}\}\\left\(t\_\{k\}\\right\)\\psi\_\{k\}\\left\(a\_\{u\_\{j\}\}\\left\(\\mathbf\{x\}\\right\)\\right\)\.\(169\)Since each hat functionψk\\psi\_\{k\}is supported only on the grid intervals adjacent totkt\_\{k\}, \([168](https://arxiv.org/html/2608.13882#S7.E168)\) implies
ψk\(auj\(𝐱\)\)=0,k∉\{qj\(𝐱\),qj\(𝐱\)\+1\}\.\\displaystyle\\psi\_\{k\}\\left\(a\_\{u\_\{j\}\}\\left\(\\mathbf\{x\}\\right\)\\right\)=0,\\qquad k\\notin\\left\\\{q\_\{j\}\\left\(\\mathbf\{x\}\\right\),q\_\{j\}\\left\(\\mathbf\{x\}\\right\)\+1\\right\\\}\.\(170\)This statement remains valid whenauj\(𝐱\)a\_\{u\_\{j\}\}\\left\(\\mathbf\{x\}\\right\)is a grid knot: one may select either adjacent cell, and the number of nonzero hat functions can only decrease\. Substituting \([170](https://arxiv.org/html/2608.13882#S7.E170)\) into \([169](https://arxiv.org/html/2608.13882#S7.E169)\) gives
\(Πmguj\)\(auj\(𝐱\)\)=\(a\)guj\(tqj\(𝐱\)\)ψqj\(𝐱\)\(auj\(𝐱\)\)\+guj\(tqj\(𝐱\)\+1\)ψqj\(𝐱\)\+1\(auj\(𝐱\)\)\.\\displaystyle\\left\(\\Pi\_\{m\}g\_\{u\_\{j\}\}\\right\)\\left\(a\_\{u\_\{j\}\}\\left\(\\mathbf\{x\}\\right\)\\right\)\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}g\_\{u\_\{j\}\}\\left\(t\_\{q\_\{j\}\\left\(\\mathbf\{x\}\\right\)\}\\right\)\\psi\_\{q\_\{j\}\\left\(\\mathbf\{x\}\\right\)\}\\left\(a\_\{u\_\{j\}\}\\left\(\\mathbf\{x\}\\right\)\\right\)\+g\_\{u\_\{j\}\}\\left\(t\_\{q\_\{j\}\\left\(\\mathbf\{x\}\\right\)\+1\}\\right\)\\psi\_\{q\_\{j\}\\left\(\\mathbf\{x\}\\right\)\+1\}\\left\(a\_\{u\_\{j\}\}\\left\(\\mathbf\{x\}\\right\)\\right\)\.\(171\)Here, \(a\) removes from \([169](https://arxiv.org/html/2608.13882#S7.E169)\) all terms that vanish by \([170](https://arxiv.org/html/2608.13882#S7.E170)\)\. Therefore, for each fixed𝐱∈𝒳\\mathbf\{x\}\\in\\mathcal\{X\}andj∈\[M\]j\\in\\left\[M\\right\], the evaluation\(Πmguj\)\(auj\(𝐱\)\)\\left\(\\Pi\_\{m\}g\_\{u\_\{j\}\}\\right\)\\left\(a\_\{u\_\{j\}\}\\left\(\\mathbf\{x\}\\right\)\\right\)uses at most two active hat functions and at most two corresponding nodal coefficients\. Finally, \([164](https://arxiv.org/html/2608.13882#S7.E164)\) shows that
FM,m\(𝐱\)=𝒞^var\(L\)\(F\)M∑j=1Mσj\(Πmguj\)\(auj\(𝐱\)\)\.\\displaystyle F\_\{M,m\}\\left\(\\mathbf\{x\}\\right\)=\\frac\{\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}\\left\(F\\right\)\}\{M\}\\sum\_\{j=1\}^\{M\}\\sigma\_\{j\}\\left\(\\Pi\_\{m\}g\_\{u\_\{j\}\}\\right\)\\left\(a\_\{u\_\{j\}\}\\left\(\\mathbf\{x\}\\right\)\\right\)\.\(172\)Since each of theMMsummands contributes at most two active interpolation basis functions, evaluatingFM,m\(𝐱\)F\_\{M,m\}\\left\(\\mathbf\{x\}\\right\)requires at most2M2Mactive outer\-profile basis contributions\. This number is independent of the interpolation resolutionmm, which proves the local evaluation sparsity assertion\.
### 7\.7Proof of[1](https://arxiv.org/html/2608.13882#Thmprop1)
For everyg∈ℋk\(B\)g\\in\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}one has
‖g−Πmg‖L∞\(\[−A,A\]\)\\displaystyle\\left\\\|g\-\\Pi\_\{m\}g\\right\\\|\_\{L^\{\\infty\}\\left\(\\left\[\-A,A\\right\]\\right\)\}≤\(a\)\(A2m\)1/2‖g‖ℋk\(B\)=\(b\)\(A2\)1/2m−1/2‖g‖ℋk\(B\)\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}\\left\(\\frac\{A\}\{2m\}\\right\)^\{1/2\}\\left\\\|g\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\left\(\\frac\{A\}\{2\}\\right\)^\{1/2\}m^\{\-1/2\}\\left\\\|g\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}\.\(173\)Here, \(a\) applies \([497](https://arxiv.org/html/2608.13882#A4.E497)\) of[Lemma30](https://arxiv.org/html/2608.13882#Thmtheorem30), and \(b\) uses\(A/2m\)1/2=\(A/2\)1/2m−1/2\\left\(A/2m\\right\)^\{1/2\}=\\left\(A/2\\right\)^\{1/2\}m^\{\-1/2\}\. This is precisely \([73](https://arxiv.org/html/2608.13882#S4.E73)\), and therefore proves the proposition\.
[https://arxiv.org/html/2608.13882#S2.SS1](https://arxiv.org/html/2608.13882#S2.SS1)[https://arxiv.org/html/2608.13882#S3.SS1](https://arxiv.org/html/2608.13882#S3.SS1)[https://arxiv.org/html/2608.13882#S3.SS2](https://arxiv.org/html/2608.13882#S3.SS2)[https://arxiv.org/html/2608.13882#S4](https://arxiv.org/html/2608.13882#S4)[https://arxiv.org/html/2608.13882#S5](https://arxiv.org/html/2608.13882#S5)[https://arxiv.org/html/2608.13882#Thmtheorem18](https://arxiv.org/html/2608.13882#Thmtheorem18)[https://arxiv.org/html/2608.13882#S2](https://arxiv.org/html/2608.13882#S2)[https://arxiv.org/html/2608.13882#S2.E7](https://arxiv.org/html/2608.13882#S2.E7)[https://arxiv.org/html/2608.13882#S2.SS2](https://arxiv.org/html/2608.13882#S2.SS2)[https://arxiv.org/html/2608.13882#Thmtheorem15](https://arxiv.org/html/2608.13882#Thmtheorem15)[https://arxiv.org/html/2608.13882#Thmtheorem16](https://arxiv.org/html/2608.13882#Thmtheorem16)[https://arxiv.org/html/2608.13882#Thmtheorem4](https://arxiv.org/html/2608.13882#Thmtheorem4)[https://arxiv.org/html/2608.13882#Thmtheorem17](https://arxiv.org/html/2608.13882#Thmtheorem17)[https://arxiv.org/html/2608.13882#Thmtheorem21](https://arxiv.org/html/2608.13882#Thmtheorem21)[https://arxiv.org/html/2608.13882#Thmtheorem19](https://arxiv.org/html/2608.13882#Thmtheorem19)[https://arxiv.org/html/2608.13882#Thmtheorem23](https://arxiv.org/html/2608.13882#Thmtheorem23)[https://arxiv.org/html/2608.13882#Thmtheorem22](https://arxiv.org/html/2608.13882#Thmtheorem22)[https://arxiv.org/html/2608.13882#Thmtheorem20](https://arxiv.org/html/2608.13882#Thmtheorem20)[https://arxiv.org/html/2608.13882#Thmtheorem24](https://arxiv.org/html/2608.13882#Thmtheorem24)[https://arxiv.org/html/2608.13882#Thmtheorem25](https://arxiv.org/html/2608.13882#Thmtheorem25)[https://arxiv.org/html/2608.13882#Thmtheorem6](https://arxiv.org/html/2608.13882#Thmtheorem6)[https://arxiv.org/html/2608.13882#Thmtheorem8](https://arxiv.org/html/2608.13882#Thmtheorem8)[https://arxiv.org/html/2608.13882#Thmtheorem26](https://arxiv.org/html/2608.13882#Thmtheorem26)[https://arxiv.org/html/2608.13882#Thmtheorem28](https://arxiv.org/html/2608.13882#Thmtheorem28)[https://arxiv.org/html/2608.13882#Thmtheorem30](https://arxiv.org/html/2608.13882#Thmtheorem30)[https://arxiv.org/html/2608.13882#Thmtheorem29](https://arxiv.org/html/2608.13882#Thmtheorem29)[https://arxiv.org/html/2608.13882#Thmtheorem31](https://arxiv.org/html/2608.13882#Thmtheorem31)[https://arxiv.org/html/2608.13882#Thmtheorem32](https://arxiv.org/html/2608.13882#Thmtheorem32)[https://arxiv.org/html/2608.13882#Thmtheorem27](https://arxiv.org/html/2608.13882#Thmtheorem27)[https://arxiv.org/html/2608.13882#Thmtheorem9](https://arxiv.org/html/2608.13882#Thmtheorem9)[https://arxiv.org/html/2608.13882#Thmprop1](https://arxiv.org/html/2608.13882#Thmprop1)[https://arxiv.org/html/2608.13882#Thmtheorem10](https://arxiv.org/html/2608.13882#Thmtheorem10)[https://arxiv.org/html/2608.13882#Thmprop2](https://arxiv.org/html/2608.13882#Thmprop2)[https://arxiv.org/html/2608.13882#Thmtheorem12](https://arxiv.org/html/2608.13882#Thmtheorem12)[https://arxiv.org/html/2608.13882#S5.SS2](https://arxiv.org/html/2608.13882#S5.SS2)[https://arxiv.org/html/2608.13882#A1.SS1](https://arxiv.org/html/2608.13882#A1.SS1)[https://arxiv.org/html/2608.13882#A1.SS2](https://arxiv.org/html/2608.13882#A1.SS2)[https://arxiv.org/html/2608.13882#A1.SS3](https://arxiv.org/html/2608.13882#A1.SS3)Figure 5:Proof\-dependency structure of the theoretical results\. The displayed graph is the exact supplied dependency graph; every panel title, result node, and experiment node is clickable and jumps to the corresponding statement, section, or protocol\.
## Appendix AExperimental Protocols and Extended Results
This appendix provides the protocols and numerical details supporting Section[5](https://arxiv.org/html/2608.13882#S5)\. The principal empirical findings and all figures needed to assess them are retained in the main body; the material below records the controlled constructions, validation\-selected configurations, exact numerical comparisons, and computational measurements\.
### A\.1Constructive\-approximation protocol
The constructive studies use a deterministic teacherF∈𝒱\(L\)F\\in\\mathcal\{V\}^\{\(L\)\}represented by an explicit finite signed measure on recursive Brownian atoms\. Each selected atom has an exact recursive support and a deterministic unit\-energy Brownian profile\. Because this reference representation is known, the approximation errors are computed directly and do not contain training or statistical\-estimation error\. The three panels in Figure[2](https://arxiv.org/html/2608.13882#S5.F2)isolate the components of the construction\. For signed\-measure discretization, the profiles are kept exact and only the numberMMof sampled atoms is varied\. For profile discretization, the finite signed measure is held fixed and the selected profiles are replaced by continuous piecewise\-linear interpolants of resolutionmm\. For balanced refinement, both resolutions are set equal,M=m=NM=m=N\. Monte Carlo means and percentile bands are used for the sampling experiments, whereas the profile\-discretization curve is deterministic\. The separate profile study in Figure[1](https://arxiv.org/html/2608.13882#S5.F1)usesm∈\{8,16,32,64,128,256,512\}m\\in\\\{8,16,32,64,128,256,512\\\}and reports both typical fixed\-profile behaviour and the normalized tent profile that attains the sharp uniform bound\.
### A\.2Statistical\-learning protocol and extended results
The supervised studies comprise the recursive\-teacher benchmark and the Energy Efficiency benchmark\. In each study, VBKL architecture selection is performed separately for every training\-set size and random seed using the validation set\. The selected configuration is then retrained and evaluated on the held\-out test set\. Results are aggregated over five independent seeds\. DNVS, KRR, and RBF are evaluated under the same split protocol within each benchmark\.
#### A\.2\.1Recursive\-teacher configurations
The recursive\-teacher target is generated using the same hierarchical Brownian mechanism as the proposed model, providing a controlled matched setting\. Validation jointly selects the number of recursive Brownian paths and the optimization horizon\. Table[3](https://arxiv.org/html/2608.13882#A1.T3)reports these choices in seed order\. The variation across sample sizes and seeds is intentional: it records the adaptive model\-selection procedure used to obtain the learning curve in Figure[3a](https://arxiv.org/html/2608.13882#S5.F3.sf1)\.
ntrainn\_\{\\mathrm\{train\}\}Paths by seedMedian pathsEpochs by seedMedian epochs10032, 8, 2, 8, 168100, 190, 600, 430, 301902502, 64, 64, 32, 832450, 70, 80, 190, 11011050064, 16, 16, 64, 3232140, 280, 90, 60, 1101101,00016, 64, 16, 4, 216170, 70, 150, 380, 3801702,0004, 8, 8, 64, 3281,990, 990, 440, 90, 1104405,00032, 4, 64, 16, 6432170, 510, 90, 570, 80170Table 3:Validation\-selected VBKL configurations\. Path counts and training epochs are listed in seed order\. The architecture and optimization horizon are selected independently for each training size and random seed using the validation set\.
#### A\.2\.2Energy Efficiency numerical results
The Energy Efficiency experiment uses the publicly available regression benchmark and the same train/validation/test protocol for all compared methods\. For VBKL, validation selects the number of paths, the Brownian profile resolution, and the optimization horizon independently for every training size and seed\. Table[4](https://arxiv.org/html/2608.13882#A1.T4)gives the exact means and standard deviations underlying Figure[3b](https://arxiv.org/html/2608.13882#S5.F3.sf2)\. Atn=100n=100, the mean VBKL error is close to the best reported value and lower than the DNVS mean\. At the two larger training sizes, VBKL has the largest mean error among the four methods\. The parameter\-efficiency comparison in the main body should therefore be read together with, rather than in place of, this predictive comparison\.
nnVBKLDNVSKRRRBF1000\.051656\(0\.025140\)0\.092731\(0\.013174\)0\.059793\(0\.031508\)0\.050172\(0\.022180\)2500\.007220\(0\.004147\)0\.005756\(0\.001281\)0\.005158\(0\.001125\)0\.006908\(0\.001264\)5000\.004793\(0\.003983\)0\.002442\(0\.000232\)0\.002627\(0\.000309\)0\.002636\(0\.000276\)Table 4:Test MSE on the Energy Efficiency benchmark\. Values are means, with standard deviations in parentheses, over five random seeds\.
### A\.3Optimization protocol
The consistency experiment uses automatic differentiation as the reference for directional derivatives and compares it with the finite\-difference averaged directional estimator at interaction scalesh∈\{10−1,10−2,10−3,10−4\}h\\in\\\{10^\{\-1\},10^\{\-2\},10^\{\-3\},10^\{\-4\}\\\}\. The Monte Carlo stability study varies the number of sampled directions from202^\{0\}to272^\{7\}and records the standard deviation of the resulting estimator\. As shown in Figure[4](https://arxiv.org/html/2608.13882#S5.F4), the first experiment exhibits the usual finite\-difference scale trade\-off, while the second shows systematic variance reduction as the number of directions increases\.
### A\.4Computational characteristics
The recursive\-learning benchmark also records training time, inference time for the complete held\-out test set, and fitted or trainable parameter count\. Table[5](https://arxiv.org/html/2608.13882#A1.T5)summarizes these quantities over all training sizes and seeds\. Tables[6](https://arxiv.org/html/2608.13882#A1.T6),[7](https://arxiv.org/html/2608.13882#A1.T7), and[8](https://arxiv.org/html/2608.13882#A1.T8)provide the corresponding training\-size\-specific results\. Times are reported as means with standard deviations over five seeds\. Parameter counts are reported in the same format when they vary across selected configurations\.
The measurements show that the finite VBKL models remain computationally tractable at the benchmark scale, although they are not uniformly the fastest method\. DNVS has lower inference time throughout this comparison, and the feature and kernel baselines have lower training time for the tested sample sizes\. The principal computational distinction of VBKL is instead that its recursive representation remains explicit and, relative to DNVS, substantially more compact in the limited\-data Energy Efficiency comparison\.
ModelTraining time \(s\)Inference time \(ms\)ParametersVBKL2\.797±\\pm2\.18913\.104±\\pm11\.2043,248DNVS1\.338±\\pm0\.2320\.882±\\pm0\.5238,960KRR0\.103±\\pm0\.18048\.811±\\pm60\.318750RBF0\.007±\\pm0\.0098\.413±\\pm4\.5170Table 5:Computational characteristics on the recursive\-learning benchmark\. Training and inference times are reported as mean±\\pmstandard deviation over all training sizes and seeds\. Parameter counts are reported by their median because the validation\-selected VBKL architecture may vary across runs\.ModelTraining samples100100250250500500100010002000200050005000VBKL1\.086\(0\.0698\)1\.912\(0\.0308\)2\.551\(0\.0543\)1\.471\(0\.0092\)3\.878\(0\.0442\)5\.882\(0\.0064\)DNVS1\.123\(0\.0057\)1\.181\(0\.0096\)1\.353\(0\.0142\)1\.413\(0\.0271\)1\.435\(0\.0364\)1\.523\(0\.0003\)KRR0\.015\(0\.0023\)0\.005\(0\)0\.007\(0\.0001\)0\.018\(0\.0005\)0\.078\(0\.0034\)0\.492\(0\.0002\)RBF0\.002\(0\.0001\)0\.002\(0\)0\.003\(0\.0001\)0\.005\(0\.0005\)0\.009\(0\.0007\)0\.020\(0\.0013\)
Values are mean, with standard deviation shown in parentheses over five random seeds\. The best result in each column is shown in bold and the second\-best result is underlined\.
Table 6:Training time in seconds on the recursive stochastic teacher benchmark\. Lower values are better\.ModelTraining samples100100250250500500100010002000200050005000VBKL0\.0069\(0\.0062\)0\.0161\(0\.0137\)0\.0183\(0\.0111\)0\.0096\(0\.0117\)0\.0110\(0\.0117\)0\.0167\(0\.0125\)DNVS0\.0010\(0\.0004\)0\.0011\(0\.0004\)0\.0013\(0\.0005\)0\.0011\(0\.0005\)0\.0005\(0\.0004\)0\.0003\(0\)KRR0\.0028\(0\.0003\)0\.0062\(0\.0002\)0\.0126\(0\.0003\)0\.0314\(0\.0068\)0\.0688\(0\.0104\)0\.1711\(0\.0101\)RBF0\.0073\(0\.0035\)0\.0108\(0\.0038\)0\.0070\(0\.0044\)0\.0072\(0\.0048\)0\.0085\(0\.0054\)0\.0097\(0\.0059\)
Values are mean, with standard deviation shown in parentheses over five random seeds\. The best result in each column is shown in bold and the second\-best result is underlined\.
Table 7:Inference time in seconds for the complete held\-out test set\. Lower values are better\.ModelTraining samples100100250250500500100010002000200050005000VBKL2,680\(2,360\)6,902\(6,008\)7,795\(4,926\)4,141\(5,123\)4,710\(5,148\)7,308\(5,567\)DNVS12,723\(12,391\)14,029\(11,334\)24,166\(13,881\)21,555\(17,457\)8,806\(14,254\)2,432KRR1002505001,0002,0005,000RBF000000
Values are mean, with standard deviation shown in parentheses over five random seeds\. The best result in each column is shown in bold and the second\-best result is underlined\.
Table 8:Number of trainable or fitted parameters used by each method\. Lower values indicate a more compact representation\.
## Appendix BAdditional notation
The following standard notation is used throughout the proofs\. For a normed vector space\(Z,∥⋅∥Z\)\(Z,\\\|\\cdot\\\|\_\{Z\}\)andr\>0r\>0, letBr\(Z\):=\{z∈Z:‖z‖Z≤r\},Sr\(Z\):=\{z∈Z:‖z‖Z=r\}\.B\_\{r\}\(Z\):=\\\{z\\in Z:\\\|z\\\|\_\{Z\}\\leq r\\\},\\qquad S\_\{r\}\(Z\):=\\\{z\\in Z:\\\|z\\\|\_\{Z\}=r\\\}\.In particular,Bd:=B1\(ℝd,∥⋅∥2\)B^\{d\}:=B\_\{1\}\(\\mathbb\{R\}^\{d\},\\\|\\cdot\\\|\_\{2\}\),Sd−1:=S1\(ℝd,∥⋅∥2\)\.S^\{d\-1\}:=S\_\{1\}\(\\mathbb\{R\}^\{d\},\\\|\\cdot\\\|\_\{2\}\)\.For𝐱∈ℝd\\mathbf\{x\}\\in\\mathbb\{R\}^\{d\}andp∈\[1,∞\)p\\in\[1,\\infty\), we write‖𝐱‖p=\(∑i=1d\|xi\|p\)1/p\\\|\\mathbf\{x\}\\\|\_\{p\}=\\left\(\\sum\_\{i=1\}^\{d\}\|x\_\{i\}\|^\{p\}\\right\)^\{1/p\},‖𝐱‖∞=maxi∈\[d\]\|xi\|,\\\|\\mathbf\{x\}\\\|\_\{\\infty\}=\\max\_\{i\\in\[d\]\}\|x\_\{i\}\|,and denote the transpose of𝐱\\mathbf\{x\}by𝐱⊤\\mathbf\{x\}^\{\\top\}\. We write𝒞\(𝒳\)\\mathcal\{C\}\(\\mathcal\{X\}\)for the space of continuous functions on𝒳\\mathcal\{X\}equipped with the supremum norm‖f‖∞=sup𝐱∈𝒳\|f\(𝐱\)\|\.\\\|f\\\|\_\{\\infty\}=\\sup\_\{\\mathbf\{x\}\\in\\mathcal\{X\}\}\|f\(\\mathbf\{x\}\)\|\.Forα∈\(0,1\]\\alpha\\in\(0,1\], letC0,α\(Ω\)C^\{0,\\alpha\}\(\\Omega\)denote the space ofα\\alpha\-Hölder continuous functions onΩ\\Omega, equipped with the norm‖f‖C0,α\(Ω\)=‖f‖∞\+\[f\]C0,α\(Ω\),\\\|f\\\|\_\{C^\{0,\\alpha\}\(\\Omega\)\}=\\\|f\\\|\_\{\\infty\}\+\[f\]\_\{C^\{0,\\alpha\}\(\\Omega\)\},where\[f\]C0,α\(Ω\)=supx≠y\|f\(x\)−f\(y\)\|‖x−y‖2α\.\[f\]\_\{C^\{0,\\alpha\}\(\\Omega\)\}=\\sup\_\{x\\neq y\}\\frac\{\|f\(x\)\-f\(y\)\|\}\{\\\|x\-y\\\|\_\{2\}^\{\\alpha\}\}\.For a setAA, we writeL∞\(A\)L^\{\\infty\}\(A\)for the vector space of bounded real\-valued functions onAA, equipped with the norm‖f‖L∞\(A\):=supx∈A\|f\(x\)\|\\\|f\\\|\_\{L^\{\\infty\}\(A\)\}:=\\sup\_\{x\\in A\}\|f\(x\)\|\. We writeL2\(ν\)=\{F:𝒳→ℝ:∫𝒳\|F\(𝐱\)\|2dν\(𝐱\)<∞\}L^\{2\}\(\\nu\)=\\left\\\{F:\\mathcal\{X\}\\to\\mathbb\{R\}:\\int\_\{\\mathcal\{X\}\}\|F\(\\mathbf\{x\}\)\|^\{2\}\\,d\\nu\(\\mathbf\{x\}\)<\\infty\\right\\\}for the Hilbert space ofν\\nu\-square\-integrable functions, equipped with the inner product⟨F,G⟩L2\(ν\)=∫𝒳F\(𝐱\)G\(𝐱\)𝑑ν\(𝐱\),\\langle F,G\\rangle\_\{L^\{2\}\(\\nu\)\}=\\int\_\{\\mathcal\{X\}\}F\(\\mathbf\{x\}\)G\(\\mathbf\{x\}\)\\,d\\nu\(\\mathbf\{x\}\),and induced norm‖F‖L2\(ν\)=⟨F,F⟩L2\(ν\)1/2\.\\\|F\\\|\_\{L^\{2\}\(\\nu\)\}=\\langle F,F\\rangle\_\{L^\{2\}\(\\nu\)\}^\{1/2\}\.Unless stated otherwise, vector\-valued integrals are understood in the Bochner sense\. For a statementEE, we write𝟏E\\mathbf\{1\}\_\{E\}for its indicator function\.
## Appendix CSharpness of the Brownian profile interpolation estimate
The preceding proposition establishes the uniform interpolation estimate used throughout the constructive approximation theory\. The following theorem shows that this estimate is sharp: neither the convergence rate nor the constant can be improved\.
###### Proposition 2\.
\(Sharpness of Brownian profile interpolation\)LetA\>0A\>0andm∈ℕm\\in\\mathbb\{N\},m≥1m\\geq 1\. Let−A=t0<t1<⋯<tm=A\-A=t\_\{0\}<t\_\{1\}<\\cdots<t\_\{m\}=Abe the uniform interpolation grid on\[−A,A\]\[\-A,A\], and letΠm\\Pi\_\{m\}denote the associated continuous piecewise\-linear interpolation operator\. Then
supg∈ℋk\(B\)‖g‖ℋk\(B\)≤1‖g−Πmg‖L∞\(\[−A,A\]\)=\(A2m\)1/2\.\\displaystyle\\sup\_\{\\begin\{subarray\}\{c\}g\\in\\mathcal\{H\}\_\{k^\{\(B\)\}\}\\\\ \\\|g\\\|\_\{\\mathcal\{H\}\_\{k^\{\(B\)\}\}\}\\leq 1\\end\{subarray\}\}\\\|g\-\\Pi\_\{m\}g\\\|\_\{L^\{\\infty\}\(\[\-A,A\]\)\}=\\left\(\\frac\{A\}\{2m\}\\right\)^\{1/2\}\.\(174\)Consequently, the interpolation estimate of[Lemma30](https://arxiv.org/html/2608.13882#Thmtheorem30)is optimal: neither the exponent1/21/2nor the constantA/2\\sqrt\{A/2\}can be improved\. Moreover, for everym≥1m\\geq 1, there existsgm∈ℋk\(B\)g\_\{m\}\\in\\mathcal\{H\}\_\{k^\{\(B\)\}\}such that
∥gm∥ℋk\(B\)=1,∥gm−Πmgm∥L∞\(\[−A,A\]\)=\(A2\)1/2m−1/2\.\\displaystyle\\\|g\_\{m\}\\\|\_\{\\mathcal\{H\}\_\{k^\{\(B\)\}\}\}=1,\\qquad\\\|g\_\{m\}\-\\Pi\_\{m\}g\_\{m\}\\\|\_\{L^\{\\infty\}\(\[\-A,A\]\)\}=\\left\(\\frac\{A\}\{2\}\\right\)^\{1/2\}m^\{\-1/2\}\.\(175\)
###### Proof\.
By Proposition 1, it suffices to establish the reverse inequality\. Leth:=2Amh:=\\frac\{2A\}\{m\}denote the grid spacing\. Fix an interpolation interval\[ti,ti\+1\]\[t\_\{i\},t\_\{i\+1\}\], letsi:=ti\+ti\+12,s\_\{i\}:=\\frac\{t\_\{i\}\+t\_\{i\+1\}\}\{2\},and definegmg\_\{m\}through its weak derivative by
gm′\(r\):=1h𝟏\[ti,si\]\(r\)−1h𝟏\[si,ti\+1\]\(r\)\.\\displaystyle g\_\{m\}^\{\\prime\}\(r\):=\\frac\{1\}\{\\sqrt\{h\}\}\\mathbf\{1\}\_\{\[t\_\{i\},s\_\{i\}\]\}\(r\)\-\\frac\{1\}\{\\sqrt\{h\}\}\\mathbf\{1\}\_\{\[s\_\{i\},t\_\{i\+1\}\]\}\(r\)\.\(176\)Finally, definegm\(t\):=∫0tgm′\(r\)𝑑rg\_\{m\}\(t\):=\\int\_\{0\}^\{t\}g\_\{m\}^\{\\prime\}\(r\)\\,\\mathrm\{d\}r\. Sincegmg\_\{m\}is absolutely continuous, satisfiesgm\(0\)=0g\_\{m\}\(0\)=0, andgm′∈L2\(ℝ\)g\_\{m\}^\{\\prime\}\\in L^\{2\}\(\\mathbb\{R\}\), we obtaingm∈ℋk\(B\)\.g\_\{m\}\\in\\mathcal\{H\}\_\{k^\{\(B\)\}\}\.Moreover,
‖gm‖ℋk\(B\)2\\displaystyle\\\|g\_\{m\}\\\|\_\{\\mathcal\{H\}\_\{k^\{\(B\)\}\}\}^\{2\}=\(a\)∫ℝ\|gm′\(r\)\|2𝑑r=\(b\)1h\(\|\[ti,si\]\|\+\|\[si,ti\+1\]\|\)=\(c\)1h\(h2\+h2\)=1\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\int\_\{\\mathbb\{R\}\}\|g\_\{m\}^\{\\prime\}\(r\)\|^\{2\}\\,\\mathrm\{d\}r\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\frac\{1\}\{h\}\\left\(\|\[t\_\{i\},s\_\{i\}\]\|\+\|\[s\_\{i\},t\_\{i\+1\}\]\|\\right\)\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{=\}\}\\frac\{1\}\{h\}\\left\(\\frac\{h\}\{2\}\+\\frac\{h\}\{2\}\\right\)=1\.\(177\)Here, \(a\) is the definition of the Brownian RKHS norm, \(b\) follows from the definition ofgm′g\_\{m\}^\{\\prime\}, and \(c\) follows from the definition of the midpointsis\_\{i\}\. Next,
gm\(ti\+1\)−gm\(ti\)\\displaystyle g\_\{m\}\(t\_\{i\+1\}\)\-g\_\{m\}\(t\_\{i\}\)=\(a\)∫titi\+1gm′\(r\)𝑑r=\(b\)1hh2−1hh2=0\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\int\_\{t\_\{i\}\}^\{t\_\{i\+1\}\}g\_\{m\}^\{\\prime\}\(r\)\\,\\mathrm\{d\}r\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\frac\{1\}\{\\sqrt\{h\}\}\\frac\{h\}\{2\}\-\\frac\{1\}\{\\sqrt\{h\}\}\\frac\{h\}\{2\}=0\.\(178\)Here, \(a\) is the fundamental theorem of calculus, while \(b\) follows from the definition ofgm′g\_\{m\}^\{\\prime\}\. Hence,gm\(ti\)=gm\(ti\+1\)\.g\_\{m\}\(t\_\{i\}\)=g\_\{m\}\(t\_\{i\+1\}\)\.Therefore, the piecewise\-linear interpolant is constant on\[ti,ti\+1\]\[t\_\{i\},t\_\{i\+1\}\], that is,
\(Πmgm\)\(t\)=gm\(ti\),t∈\[ti,ti\+1\]\.\\displaystyle\(\\Pi\_\{m\}g\_\{m\}\)\(t\)=g\_\{m\}\(t\_\{i\}\),\\qquad t\\in\[t\_\{i\},t\_\{i\+1\}\]\.\(179\)Evaluating at the midpointsis\_\{i\}, we obtain
gm\(si\)−\(Πmgm\)\(si\)\\displaystyle g\_\{m\}\(s\_\{i\}\)\-\(\\Pi\_\{m\}g\_\{m\}\)\(s\_\{i\}\)=\(a\)gm\(si\)−gm\(ti\)=\(b\)∫tisigm′\(r\)𝑑r=\(c\)1hh2=h2\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}g\_\{m\}\(s\_\{i\}\)\-g\_\{m\}\(t\_\{i\}\)\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\int\_\{t\_\{i\}\}^\{s\_\{i\}\}g\_\{m\}^\{\\prime\}\(r\)\\,\\mathrm\{d\}r\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{=\}\}\\frac\{1\}\{\\sqrt\{h\}\}\\frac\{h\}\{2\}=\\frac\{\\sqrt\{h\}\}\{2\}\.\(180\)Here, \(a\) follows from \([179](https://arxiv.org/html/2608.13882#A3.E179)\), \(b\) is the fundamental theorem of calculus, and \(c\) follows from the definition ofgm′g\_\{m\}^\{\\prime\}\. Sinceh=2A/mh=2A/m, we conclude that
\|gm\(si\)−\(Πmgm\)\(si\)\|=12\(2Am\)1/2=\(A2m\)1/2\.\\displaystyle\\left\|g\_\{m\}\(s\_\{i\}\)\-\(\\Pi\_\{m\}g\_\{m\}\)\(s\_\{i\}\)\\right\|=\\frac\{1\}\{2\}\\left\(\\frac\{2A\}\{m\}\\right\)^\{1/2\}=\\left\(\\frac\{A\}\{2m\}\\right\)^\{1/2\}\.\(181\)Therefore,
‖gm−Πmgm‖L∞\(\[−A,A\]\)≥\(A2m\)1/2\.\\displaystyle\\\|g\_\{m\}\-\\Pi\_\{m\}g\_\{m\}\\\|\_\{L^\{\\infty\}\(\[\-A,A\]\)\}\\geq\\left\(\\frac\{A\}\{2m\}\\right\)^\{1/2\}\.\(182\)On the other hand, since‖gm‖ℋk\(B\)=1\\\|g\_\{m\}\\\|\_\{\\mathcal\{H\}\_\{k^\{\(B\)\}\}\}=1, Proposition 1 yields
‖gm−Πmgm‖L∞\(\[−A,A\]\)≤\(A2m\)1/2\.\\displaystyle\\\|g\_\{m\}\-\\Pi\_\{m\}g\_\{m\}\\\|\_\{L^\{\\infty\}\(\[\-A,A\]\)\}\\leq\\left\(\\frac\{A\}\{2m\}\\right\)^\{1/2\}\.\(183\)Combining \([182](https://arxiv.org/html/2608.13882#A3.E182)\) and \([183](https://arxiv.org/html/2608.13882#A3.E183)\), we obtain‖gm−Πmgm‖L∞\(\[−A,A\]\)=\(A2m\)1/2\\\|g\_\{m\}\-\\Pi\_\{m\}g\_\{m\}\\\|\_\{L^\{\\infty\}\(\[\-A,A\]\)\}=\\left\(\\frac\{A\}\{2m\}\\right\)^\{1/2\}\. Taking the supremum over allg∈ℋk\(B\)g\\in\\mathcal\{H\}\_\{k^\{\(B\)\}\}with‖g‖ℋk\(B\)≤1\\\|g\\\|\_\{\\mathcal\{H\}\_\{k^\{\(B\)\}\}\}\\leq 1establishes the reverse inequality\. Together with Proposition 1, this proves the theorem\. ∎
ResultContentPageLemma[15](https://arxiv.org/html/2608.13882#Thmtheorem15)Well\-definedness of the VBKL spacepage[15](https://arxiv.org/html/2608.13882#Thmtheorem15)Lemma[16](https://arxiv.org/html/2608.13882#Thmtheorem16)Non\-degeneracy of the variation complexitypage[16](https://arxiv.org/html/2608.13882#Thmtheorem16)Lemma[17](https://arxiv.org/html/2608.13882#Thmtheorem17)Finite\-profile characterization and finite\-architecture boundspage[17](https://arxiv.org/html/2608.13882#Thmtheorem17)Lemma[18](https://arxiv.org/html/2608.13882#Thmtheorem18)Recursive union\-of\-RKHS characterizationpage[18](https://arxiv.org/html/2608.13882#Thmtheorem18)Lemma[19](https://arxiv.org/html/2608.13882#Thmtheorem19)Exact reduction from finite variation hulls to outer dictionariespage[19](https://arxiv.org/html/2608.13882#Thmtheorem19)Lemma[20](https://arxiv.org/html/2608.13882#Thmtheorem20)Rademacher bound for unions of Brownian pullback RKHS ballspage[20](https://arxiv.org/html/2608.13882#Thmtheorem20)Lemma[21](https://arxiv.org/html/2608.13882#Thmtheorem21)Finite threshold\-chaos boundpage[21](https://arxiv.org/html/2608.13882#Thmtheorem21)Lemma[22](https://arxiv.org/html/2608.13882#Thmtheorem22)Brownian chaos controlled by signed threshold tracespage[22](https://arxiv.org/html/2608.13882#Thmtheorem22)Lemma[23](https://arxiv.org/html/2608.13882#Thmtheorem23)VC dimension of signed finite\-architecture threshold classespage[23](https://arxiv.org/html/2608.13882#Thmtheorem23)Lemma[24](https://arxiv.org/html/2608.13882#Thmtheorem24)Signed threshold entropy of finite lower\-support architecturespage[24](https://arxiv.org/html/2608.13882#Thmtheorem24)Lemma[25](https://arxiv.org/html/2608.13882#Thmtheorem25)Brownian chaos bound from finite\-architecture threshold entropypage[25](https://arxiv.org/html/2608.13882#Thmtheorem25)Lemma[26](https://arxiv.org/html/2608.13882#Thmtheorem26)Uniform Hölder, pointwise, andL2L^\{2\}bounds for atomic generatorspage[26](https://arxiv.org/html/2608.13882#Thmtheorem26)Lemma[27](https://arxiv.org/html/2608.13882#Thmtheorem27)Attainment of the variation complexitypage[27](https://arxiv.org/html/2608.13882#Thmtheorem27)Lemma[28](https://arxiv.org/html/2608.13882#Thmtheorem28)Compactness of restricted Brownian profilespage[28](https://arxiv.org/html/2608.13882#Thmtheorem28)Lemma[29](https://arxiv.org/html/2608.13882#Thmtheorem29)Compactness of the atomic VBKL dictionarypage[29](https://arxiv.org/html/2608.13882#Thmtheorem29)Lemma[30](https://arxiv.org/html/2608.13882#Thmtheorem30)Uniform Brownian profile interpolation estimatepage[30](https://arxiv.org/html/2608.13882#Thmtheorem30)Lemma[31](https://arxiv.org/html/2608.13882#Thmtheorem31)Uniform range bound for lower\-level atomic supportspage[31](https://arxiv.org/html/2608.13882#Thmtheorem31)Lemma[32](https://arxiv.org/html/2608.13882#Thmtheorem32)Stability of profile interpolation under compositionpage[32](https://arxiv.org/html/2608.13882#Thmtheorem32)Table 9:Overview of the supporting theoretical results\. Their logical dependencies are illustrated in Fig\.[5](https://arxiv.org/html/2608.13882#A0.F5); the main theoretical results are summarized in Table[1](https://arxiv.org/html/2608.13882#S1.T1)\.
## Appendix DInternal Lemmas
In this section, we present our own lemmas used in[Section7](https://arxiv.org/html/2608.13882#S7)\.
###### Lemma B0\.
\(Well\-definedness of the measure\-valued VBKL space\)FixL≥2L\\geq 2\. Let\(𝒰l\)l≥1\\left\(\\mathcal\{U\}\_\{l\}\\right\)\_\{l\\geq 1\}be the recursive atomic dictionaries introduced in[Section2\.1](https://arxiv.org/html/2608.13882#S2.SS1), and let𝒱\(L\)\\mathcal\{V\}^\{\(L\)\}and𝒞^var\(L\)\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}be defined in \([7](https://arxiv.org/html/2608.13882#S2.E7)\) and \([8](https://arxiv.org/html/2608.13882#S2.E8)\)\. Assume that𝒰L⊆L2\(ν\)\\mathcal\{U\}\_\{L\}\\subseteq L^\{2\}\\left\(\\nu\\right\)is equipped with aσ\\sigma\-algebra for which the canonical inclusionι:𝒰L⟶L2\(ν\)\\iota:\\mathcal\{U\}\_\{L\}\\longrightarrow L^\{2\}\\left\(\\nu\\right\),ι\(u\):=u,\\iota\\left\(u\\right\):=u,is strongly measurable\. Assume further that there existsRL<∞R\_\{L\}<\\inftysuch that
supu∈𝒰L‖u‖L2\(ν\)≤RL\.\\displaystyle\\sup\_\{u\\in\\mathcal\{U\}\_\{L\}\}\\left\\\|u\\right\\\|\_\{L^\{2\}\\left\(\\nu\\right\)\}\\leq R\_\{L\}\.\(184\)Letμ∈ℳ\(𝒰L\)\\mu\\in\\mathcal\{M\}\\left\(\\mathcal\{U\}\_\{L\}\\right\)be a finite signed measure satisfying‖μ‖TV<∞\.\\left\\\|\\mu\\right\\\|\_\{\\mathrm\{TV\}\}<\\infty\.Then the Bochner integral
F:=∫𝒰Lu𝑑μ\(u\)\\displaystyle F:=\\int\_\{\\mathcal\{U\}\_\{L\}\}u\\,\\mathrm\{d\}\\mu\\left\(u\\right\)\(185\)is a well\-defined element ofL2\(ν\)L^\{2\}\\left\(\\nu\\right\)and satisfies
‖F‖L2\(ν\)≤RL‖μ‖TV\.\\displaystyle\\left\\\|F\\right\\\|\_\{L^\{2\}\\left\(\\nu\\right\)\}\\leq R\_\{L\}\\left\\\|\\mu\\right\\\|\_\{\\mathrm\{TV\}\}\.\(186\)Consequently,𝒱\(L\)\\mathcal\{V\}^\{\(L\)\}is a well\-defined subset ofL2\(ν\)L^\{2\}\\left\(\\nu\\right\)\. Furthermore, for everyF∈𝒱\(L\)F\\in\\mathcal\{V\}^\{\(L\)\}, the value𝒞^var\(L\)\(F\)\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}\\left\(F\\right\)is a well\-defined finite nonnegative real number\.
###### Proof\.
By assumption, the canonical inclusionι:𝒰L→L2\(ν\)\\iota:\\mathcal\{U\}\_\{L\}\\to L^\{2\}\(\\nu\)is strongly measurable\. Moreover, the uniform bound \([184](https://arxiv.org/html/2608.13882#A4.E184)\) implies that
∫𝒰L‖u‖L2\(ν\)d\|μ\|\(u\)≤RL\|μ\|\(𝒰L\)=RL‖μ‖TV<∞\.\\displaystyle\\int\_\{\\mathcal\{U\}\_\{L\}\}\\\|u\\\|\_\{L^\{2\}\(\\nu\)\}\\mathrm\{d\}\|\\mu\|\(u\)\\leq R\_\{L\}\|\\mu\|\(\\mathcal\{U\}\_\{L\}\)=R\_\{L\}\\\|\\mu\\\|\_\{\\mathrm\{TV\}\}<\\infty\.\(187\)Thus, the mapu↦uu\\mapsto uis strongly measurable and integrable in norm with respect to the total\-variation measure\|μ\|\|\\mu\|\. It is therefore Bochner integrable with respect to the finite signed measureμ\\mu, and the integral in \([185](https://arxiv.org/html/2608.13882#A4.E185)\) defines an elementF∈L2\(ν\)F\\in L^\{2\}\(\\nu\)\. Then
‖F‖L2\(ν\)\\displaystyle\\\|F\\\|\_\{L^\{2\}\(\\nu\)\}=\(a\)‖∫𝒰Lu𝑑μ\(u\)‖L2\(ν\)≤\(b\)∫𝒰L‖u‖L2\(ν\)d\|μ\|\(u\)≤\(c\)RL‖μ‖TV\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\left\\\|\\int\_\{\\mathcal\{U\}\_\{L\}\}u\\,\\mathrm\{d\}\\mu\(u\)\\right\\\|\_\{L^\{2\}\(\\nu\)\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}\\int\_\{\\mathcal\{U\}\_\{L\}\}\\\|u\\\|\_\{L^\{2\}\(\\nu\)\}\\,\\mathrm\{d\}\|\\mu\|\(u\)\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{\\leq\}\}R\_\{L\}\\\|\\mu\\\|\_\{\\mathrm\{TV\}\}\.\(188\)Here, \(a\) is the representation ofFF, \(b\) is the norm inequality for Bochner integrals with respect to finite signed measures, and \(c\) follows from the uniform bound \([184](https://arxiv.org/html/2608.13882#A4.E184)\)\. This proves \([186](https://arxiv.org/html/2608.13882#A4.E186)\)\. Every function appearing in the defining set for𝒱\(L\)\\mathcal\{V\}^\{\(L\)\}is therefore an element ofL2\(ν\)L^\{2\}\(\\nu\)\. Hence𝒱\(L\)\\mathcal\{V\}^\{\(L\)\}is a well\-defined subset ofL2\(ν\)L^\{2\}\(\\nu\)\. Finally, fixF∈𝒱\(L\)F\\in\\mathcal\{V\}^\{\(L\)\}\. By the definition of𝒱\(L\)\\mathcal\{V\}^\{\(L\)\}, there exists at least one finite signed measureμ\\muon𝒰L\\mathcal\{U\}\_\{L\}such thatF=∫𝒰Lu𝑑μ\(u\)F=\\int\_\{\\mathcal\{U\}\_\{L\}\}u\\mathrm\{d\}\\mu\(u\)and‖μ‖TV<∞\\\|\\mu\\\|\_\{\\mathrm\{TV\}\}<\\infty\. Consequently, the set
\{∥μ∥TV:F=∫𝒰Ludμ\(u\),μis a finite signed measure on𝒰L\}\\displaystyle\\left\\\{\\\|\\mu\\\|\_\{\\mathrm\{TV\}\}:F=\\int\_\{\\mathcal\{U\}\_\{L\}\}u\\mathrm\{d\}\\mu\(u\),\\ \\mu\\text\{ is a finite signed measure on \}\\mathcal\{U\}\_\{L\}\\right\\\}\(189\)is nonempty, is contained in\[0,∞\)\[0,\\infty\), and contains at least one finite value\. Its infimum therefore exists as a finite nonnegative real number\. Hence𝒞^var\(L\)\(F\)\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}\(F\)is well\-defined\. ∎
###### Lemma B0\.
\(Non\-degeneracy of the variation complexity\) Assume the hypotheses of[Lemma15](https://arxiv.org/html/2608.13882#Thmtheorem15), and letRLR\_\{L\}be the uniform atomic bound appearing in \([184](https://arxiv.org/html/2608.13882#A4.E184)\)\. Then, for everyF∈𝒱\(L\)F\\in\\mathcal\{V\}^\{\(L\)\},
‖F‖L2\(ν\)≤RL𝒞^var\(L\)\(F\)\.\\displaystyle\\left\\\|F\\right\\\|\_\{L^\{2\}\\left\(\\nu\\right\)\}\\leq R\_\{L\}\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}\\left\(F\\right\)\.\(190\)Consequently,
𝒞^var\(L\)\(F\)=0⟹F=0inL2\(ν\)\.\\displaystyle\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}\\left\(F\\right\)=0\\quad\\Longrightarrow\\quad F=0\\qquad\\text\{in \}L^\{2\}\\left\(\\nu\\right\)\.\(191\)Hence,𝒞^var\(L\)\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}is non\-degenerate on𝒱\(L\)\\mathcal\{V\}^\{\(L\)\}\.
###### Proof\.
FixF∈𝒱\(L\)F\\in\\mathcal\{V\}^\{\(L\)\}and letε\>0\\varepsilon\>0be arbitrary\. By the definition of𝒞^var\(L\)\(F\)\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}\(F\), there exists a finite signed measureμ\\muon𝒰L\\mathcal\{U\}\_\{L\}satisfyingF=∫𝒰Lu𝑑μ\(u\)F=\\int\_\{\\mathcal\{U\}\_\{L\}\}u\\,\\mathrm\{d\}\\mu\(u\)and‖μ‖TV≤𝒞^var\(L\)\(F\)\+ε\.\\\|\\mu\\\|\_\{\\mathrm\{TV\}\}\\leq\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}\(F\)\+\\varepsilon\.Applying[Lemma15](https://arxiv.org/html/2608.13882#Thmtheorem15)to this representation yields‖F‖L2\(ν\)≤RL‖μ‖TV≤RL\(𝒞^var\(L\)\(F\)\+ε\)\.\\\|F\\\|\_\{L^\{2\}\(\\nu\)\}\\leq R\_\{L\}\\\|\\mu\\\|\_\{\\mathrm\{TV\}\}\\leq R\_\{L\}\\left\(\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}\(F\)\+\\varepsilon\\right\)\.Since the above estimate holds for everyε\>0\\varepsilon\>0, lettingε↓0\\varepsilon\\downarrow 0gives‖F‖L2\(ν\)≤RL𝒞^var\(L\)\(F\)\\\|F\\\|\_\{L^\{2\}\(\\nu\)\}\\leq R\_\{L\}\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}\(F\), which proves \([190](https://arxiv.org/html/2608.13882#A4.E190)\)\. Now suppose that𝒞^var\(L\)\(F\)=0\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}\(F\)=0\. Then \([190](https://arxiv.org/html/2608.13882#A4.E190)\) immediately implies‖F‖L2\(ν\)=0\\\|F\\\|\_\{L^\{2\}\(\\nu\)\}=0\. SinceL2\(ν\)L^\{2\}\(\\nu\)is a normed vector space, its norm is positive definite\. ThereforeF=0F=0inL2\(ν\)L^\{2\}\(\\nu\), proving \([191](https://arxiv.org/html/2608.13882#A4.E191)\)\. The final assertion follows directly from this implication\.
∎
###### Lemma B0\.
\(Finite\-profile characterization and finite\-architecture bounds\)FixL≥2L\\geq 2,m≥1m\\geq 1, andG≥2G\\geq 2, and consider the finite lower\-support architecture introduced in[Section2\.2](https://arxiv.org/html/2608.13882#S2.SS2)\. Let−A𝒳=t0<t1<⋯<tj0=0<⋯<tG=A𝒳\-A\_\{\\mathcal\{X\}\}=t\_\{0\}<t\_\{1\}<\\cdots<t\_\{j\_\{0\}\}=0<\\cdots<t\_\{G\}=A\_\{\\mathcal\{X\}\}be the grid from \([12](https://arxiv.org/html/2608.13882#S2.E12)\)\. Then the following statements hold\.
1. 1\.A functionq:ℝ→ℝq:\\mathbb\{R\}\\rightarrow\\mathbb\{R\}belongs to𝒬G,A𝒳\\mathcal\{Q\}\_\{G,A\_\{\\mathcal\{X\}\}\}if and only if 1. \(a\)qqis continuous onℝ\\mathbb\{R\}; 2. \(b\)qqis affine on every interval\[tj,tj\+1\]\\left\[t\_\{j\},t\_\{j\+1\}\\right\], wherej∈\{0,…,G−1\}j\\in\\left\\\{0,\\ldots,G\-1\\right\\\}; 3. \(c\)qqis constant on\(−∞,−A𝒳\]\\left\(\-\\infty,\-A\_\{\\mathcal\{X\}\}\\right\]and on\[A𝒳,∞\)\\left\[A\_\{\\mathcal\{X\}\},\\infty\\right\); 4. \(d\)q\(0\)=0q\\left\(0\\right\)=0; 5. \(e\)the nodal coefficients satisfy ∑j=0G−1\(q\(tj\+1\)−q\(tj\)\)2tj\+1−tj≤1\.\\displaystyle\\sum\_\{j=0\}^\{G\-1\}\\frac\{\\left\(q\\left\(t\_\{j\+1\}\\right\)\-q\\left\(t\_\{j\}\\right\)\\right\)^\{2\}\}\{t\_\{j\+1\}\-t\_\{j\}\}\\leq 1\.\(192\) Moreover, every such function belongs toℋk\(B\)\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}, and its Brownian RKHS norm is given exactly by ‖q‖ℋk\(B\)2=∑j=0G−1\(q\(tj\+1\)−q\(tj\)\)2tj\+1−tj\.\\displaystyle\\left\\\|q\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}^\{2\}=\\sum\_\{j=0\}^\{G\-1\}\\frac\{\\left\(q\\left\(t\_\{j\+1\}\\right\)\-q\\left\(t\_\{j\}\\right\)\\right\)^\{2\}\}\{t\_\{j\+1\}\-t\_\{j\}\}\.\(193\)
2. 2\.Everya∈𝒜L−1m,Ga\\in\\mathcal\{A\}\_\{L\-1\}^\{m,G\}satisfies \|a\(𝐱\)\|≤A𝒳2−\(L−2\),𝐱∈𝒳\.\\displaystyle\\left\|a\\left\(\\mathbf\{x\}\\right\)\\right\|\\leq A\_\{\\mathcal\{X\}\}^\{\\,2^\{\-\(L\-2\)\}\},\\qquad\\mathbf\{x\}\\in\\mathcal\{X\}\.\(194\)Consequently, supa∈𝒜L−1m,G‖a‖L∞\(𝒳\)≤A𝒳2−\(L−2\)\.\\displaystyle\\sup\_\{a\\in\\mathcal\{A\}\_\{L\-1\}^\{m,G\}\}\\left\\\|a\\right\\\|\_\{L^\{\\infty\}\\left\(\\mathcal\{X\}\\right\)\}\\leq A\_\{\\mathcal\{X\}\}^\{\\,2^\{\-\(L\-2\)\}\}\.\(195\) Furthermore, everyf∈𝒟Lm,Gf\\in\\mathcal\{D\}\_\{L\}^\{m,G\}satisfies \|f\(𝐱\)\|≤A𝒳2−\(L−1\),𝐱∈𝒳\.\\displaystyle\\left\|f\\left\(\\mathbf\{x\}\\right\)\\right\|\\leq A\_\{\\mathcal\{X\}\}^\{\\,2^\{\-\(L\-1\)\}\},\\qquad\\mathbf\{x\}\\in\\mathcal\{X\}\.\(196\)
3. 3\.ForL≥3L\\geq 3, every element of𝒜L−1m,G\\mathcal\{A\}\_\{L\-1\}^\{m,G\}is represented by at most PL−1,m,G\\displaystyle P\_\{L\-1,m,G\}:=md\+\(L−2\)m\(G\+1\)\+\(L−3\)m2\+m\\displaystyle:=md\+\\left\(L\-2\\right\)m\\left\(G\+1\\right\)\+\\left\(L\-3\\right\)m^\{2\}\+m\(197\)real parameters\. ForL=2L=2, every element of𝒜1m,G=𝒰1\\mathcal\{A\}\_\{1\}^\{m,G\}=\\mathcal\{U\}\_\{1\}is represented by at most P1,m,G:=d\\displaystyle P\_\{1,m,G\}:=d\(198\)real parameters\.
###### Proof\.
We prove the three assertions separately\.
##### Proof of[1](https://arxiv.org/html/2608.13882#A4.I1.i1)\.
We first prove the forward implication\. Letq∈𝒬G,A𝒳q\\in\\mathcal\{Q\}\_\{G,A\_\{\\mathcal\{X\}\}\}\. By \([17](https://arxiv.org/html/2608.13882#S2.E17)\), there existsg∈ℋk\(B\)g\\in\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}such that
q=ΠGg,‖g‖ℋk\(B\)≤1\.\\displaystyle q=\\Pi\_\{G\}g,\\qquad\\left\\\|g\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}\\leq 1\.\(199\)By the construction ofΠG\\Pi\_\{G\}in \([13](https://arxiv.org/html/2608.13882#S2.E13)\)–\([16](https://arxiv.org/html/2608.13882#S2.E16)\), the functionqqis continuous onℝ\\mathbb\{R\}, is affine on every grid interval, and is constant outside\[−A𝒳,A𝒳\]\\left\[\-A\_\{\\mathcal\{X\}\},A\_\{\\mathcal\{X\}\}\\right\]\. Since the grid contains the origin and the Brownian RKHS is anchored at zero,
q\(0\)\\displaystyle q\\left\(0\\right\)=\(a\)\(ΠGg\)\(0\)=\(b\)g\(0\)=\(c\)0\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\left\(\\Pi\_\{G\}g\\right\)\\left\(0\\right\)\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}g\\left\(0\\right\)\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{=\}\}0\.\(200\)Here, \(a\) uses the representationq=ΠGgq=\\Pi\_\{G\}g, \(b\) follows from the nodal interpolation property attj0=0t\_\{j\_\{0\}\}=0in \([14](https://arxiv.org/html/2608.13882#S2.E14)\), and \(c\) uses the defining anchor condition ofℋk\(B\)\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}given in[Section2](https://arxiv.org/html/2608.13882#S2)\. Fixj∈\{0,…,G−1\}j\\in\\left\\\{0,\\ldots,G\-1\\right\\\}\. For almost everyt∈\(tj,tj\+1\)t\\in\\left\(t\_\{j\},t\_\{j\+1\}\\right\), differentiating the affine interpolation formula \([14](https://arxiv.org/html/2608.13882#S2.E14)\) gives
q′\(t\)\\displaystyle q^\{\\prime\}\\left\(t\\right\)=\(a\)g\(tj\+1\)−g\(tj\)tj\+1−tj=\(b\)q\(tj\+1\)−q\(tj\)tj\+1−tj\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\frac\{g\\left\(t\_\{j\+1\}\\right\)\-g\\left\(t\_\{j\}\\right\)\}\{t\_\{j\+1\}\-t\_\{j\}\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\frac\{q\\left\(t\_\{j\+1\}\\right\)\-q\\left\(t\_\{j\}\\right\)\}\{t\_\{j\+1\}\-t\_\{j\}\}\.\(201\)Here, \(a\) differentiates \([14](https://arxiv.org/html/2608.13882#S2.E14)\), while \(b\) uses the nodal interpolation identitiesq\(tj\)=g\(tj\)q\\left\(t\_\{j\}\\right\)=g\\left\(t\_\{j\}\\right\)andq\(tj\+1\)=g\(tj\+1\)q\\left\(t\_\{j\+1\}\\right\)=g\\left\(t\_\{j\+1\}\\right\)\. Moreover,q′\(t\)=0q^\{\\prime\}\\left\(t\\right\)=0for almost everyt∉\[−A𝒳,A𝒳\]t\\notin\\left\[\-A\_\{\\mathcal\{X\}\},A\_\{\\mathcal\{X\}\}\\right\]becauseqqis constant outside the interpolation interval\. Thus,qqis absolutely continuous and
∫ℝ\|q′\(t\)\|2𝑑t\\displaystyle\\int\_\{\\mathbb\{R\}\}\\left\|q^\{\\prime\}\\left\(t\\right\)\\right\|^\{2\}\\,\\mathrm\{d\}t=\(a\)∑j=0G−1∫tjtj\+1\|q′\(t\)\|2𝑑t=\(b\)∑j=0G−1∫tjtj\+1\|q\(tj\+1\)−q\(tj\)tj\+1−tj\|2𝑑t\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\sum\_\{j=0\}^\{G\-1\}\\int\_\{t\_\{j\}\}^\{t\_\{j\+1\}\}\\left\|q^\{\\prime\}\\left\(t\\right\)\\right\|^\{2\}\\,\\mathrm\{d\}t\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\sum\_\{j=0\}^\{G\-1\}\\int\_\{t\_\{j\}\}^\{t\_\{j\+1\}\}\\left\|\\frac\{q\\left\(t\_\{j\+1\}\\right\)\-q\\left\(t\_\{j\}\\right\)\}\{t\_\{j\+1\}\-t\_\{j\}\}\\right\|^\{2\}\\,\\mathrm\{d\}t\(202\)=\(c\)∑j=0G−1\(q\(tj\+1\)−q\(tj\)\)2tj\+1−tj\.\\displaystyle\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{=\}\}\\sum\_\{j=0\}^\{G\-1\}\\frac\{\\left\(q\\left\(t\_\{j\+1\}\\right\)\-q\\left\(t\_\{j\}\\right\)\\right\)^\{2\}\}\{t\_\{j\+1\}\-t\_\{j\}\}\.\(203\)Here, \(a\) uses the vanishing ofq′q^\{\\prime\}outside\[−A𝒳,A𝒳\]\\left\[\-A\_\{\\mathcal\{X\}\},A\_\{\\mathcal\{X\}\}\\right\]and the fact that the grid intervals partition this interval, \(b\) substitutes the derivative formula above, and \(c\) evaluates each integral over an interval of lengthtj\+1−tjt\_\{j\+1\}\-t\_\{j\}\. We next control this energy by the Brownian RKHS norm ofgg\. Sinceggis absolutely continuous, the fundamental theorem of calculus and the Cauchy–Schwarz inequality give
\|g\(tj\+1\)−g\(tj\)\|2\\displaystyle\\left\|g\\left\(t\_\{j\+1\}\\right\)\-g\\left\(t\_\{j\}\\right\)\\right\|^\{2\}=\(a\)\|∫tjtj\+1g′\(t\)𝑑t\|2≤\(b\)\(tj\+1−tj\)∫tjtj\+1\|g′\(t\)\|2𝑑t\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\left\|\\int\_\{t\_\{j\}\}^\{t\_\{j\+1\}\}g^\{\\prime\}\\left\(t\\right\)\\,\\mathrm\{d\}t\\right\|^\{2\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}\\left\(t\_\{j\+1\}\-t\_\{j\}\\right\)\\int\_\{t\_\{j\}\}^\{t\_\{j\+1\}\}\\left\|g^\{\\prime\}\\left\(t\\right\)\\right\|^\{2\}\\,\\mathrm\{d\}t\.\(204\)Here, \(a\) applies the fundamental theorem of calculus for absolutely continuous functions, while \(b\) applies the Cauchy–Schwarz inequality on\[tj,tj\+1\]\\left\[t\_\{j\},t\_\{j\+1\}\\right\]\. Dividing bytj\+1−tjt\_\{j\+1\}\-t\_\{j\}and summing over the grid intervals yields
∑j=0G−1\(g\(tj\+1\)−g\(tj\)\)2tj\+1−tj≤\(a\)∑j=0G−1∫tjtj\+1\|g′\(t\)\|2𝑑t=\(b\)∫−A𝒳A𝒳\|g′\(t\)\|2𝑑t\\displaystyle\\sum\_\{j=0\}^\{G\-1\}\\frac\{\\left\(g\\left\(t\_\{j\+1\}\\right\)\-g\\left\(t\_\{j\}\\right\)\\right\)^\{2\}\}\{t\_\{j\+1\}\-t\_\{j\}\}\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}\\sum\_\{j=0\}^\{G\-1\}\\int\_\{t\_\{j\}\}^\{t\_\{j\+1\}\}\\left\|g^\{\\prime\}\\left\(t\\right\)\\right\|^\{2\}\\,\\mathrm\{d\}t\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\int\_\{\-A\_\{\\mathcal\{X\}\}\}^\{A\_\{\\mathcal\{X\}\}\}\\left\|g^\{\\prime\}\\left\(t\\right\)\\right\|^\{2\}\\,\\mathrm\{d\}t\(205\)≤\(c\)∫ℝ\|g′\(t\)\|2𝑑t=\(d\)‖g‖ℋk\(B\)2≤\(e\)1\.\\displaystyle\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{\\leq\}\}\\int\_\{\\mathbb\{R\}\}\\left\|g^\{\\prime\}\\left\(t\\right\)\\right\|^\{2\}\\,\\mathrm\{d\}t\\stackrel\{\{\\scriptstyle\(d\)\}\}\{\{=\}\}\\left\\\|g\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}^\{2\}\\stackrel\{\{\\scriptstyle\(e\)\}\}\{\{\\leq\}\}1\.\(206\)Here, \(a\) sums the preceding intervalwise estimates, \(b\) uses the fact that the grid intervals partition\[−A𝒳,A𝒳\]\\left\[\-A\_\{\\mathcal\{X\}\},A\_\{\\mathcal\{X\}\}\\right\], \(c\) enlarges the integration domain, \(d\) applies the Brownian RKHS norm characterization in[Section2](https://arxiv.org/html/2608.13882#S2), and \(e\) uses‖g‖ℋk\(B\)≤1\\left\\\|g\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}\\leq 1\. Sinceq\(tj\)=g\(tj\)q\\left\(t\_\{j\}\\right\)=g\\left\(t\_\{j\}\\right\)at every grid node, the preceding estimate proves \([192](https://arxiv.org/html/2608.13882#A4.E192)\)\. Together with the anchor condition and the derivative\-energy identity established above, the Brownian RKHS characterization in[Section2](https://arxiv.org/html/2608.13882#S2)shows thatq∈ℋk\(B\)q\\in\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}and yields the exact norm formula \([193](https://arxiv.org/html/2608.13882#A4.E193)\)\. We now prove the converse implication\. Letq:ℝ→ℝq:\\mathbb\{R\}\\rightarrow\\mathbb\{R\}be continuous, affine on every grid interval, constant outside\[−A𝒳,A𝒳\]\\left\[\-A\_\{\\mathcal\{X\}\},A\_\{\\mathcal\{X\}\}\\right\], anchored at zero, and satisfying \([192](https://arxiv.org/html/2608.13882#A4.E192)\)\. Thenqqis absolutely continuous, and, for everyj∈\{0,…,G−1\}j\\in\\left\\\{0,\\ldots,G\-1\\right\\\},
q′\(t\)=q\(tj\+1\)−q\(tj\)tj\+1−tj\\displaystyle q^\{\\prime\}\\left\(t\\right\)=\\frac\{q\\left\(t\_\{j\+1\}\\right\)\-q\\left\(t\_\{j\}\\right\)\}\{t\_\{j\+1\}\-t\_\{j\}\}\(207\)for almost everyt∈\(tj,tj\+1\)t\\in\\left\(t\_\{j\},t\_\{j\+1\}\\right\), whereasq′\(t\)=0q^\{\\prime\}\\left\(t\\right\)=0for almost everyt∉\[−A𝒳,A𝒳\]t\\notin\\left\[\-A\_\{\\mathcal\{X\}\},A\_\{\\mathcal\{X\}\}\\right\]\. Consequently, the Brownian RKHS characterization in[Section2](https://arxiv.org/html/2608.13882#S2)givesq∈ℋk\(B\)q\\in\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}and
‖q‖ℋk\(B\)2\\displaystyle\\left\\\|q\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}^\{2\}=\(a\)∫ℝ\|q′\(t\)\|2𝑑t=\(b\)∑j=0G−1\(q\(tj\+1\)−q\(tj\)\)2tj\+1−tj≤\(c\)1\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\int\_\{\\mathbb\{R\}\}\\left\|q^\{\\prime\}\\left\(t\\right\)\\right\|^\{2\}\\,\\mathrm\{d\}t\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\sum\_\{j=0\}^\{G\-1\}\\frac\{\\left\(q\\left\(t\_\{j\+1\}\\right\)\-q\\left\(t\_\{j\}\\right\)\\right\)^\{2\}\}\{t\_\{j\+1\}\-t\_\{j\}\}\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{\\leq\}\}1\.\(208\)Here, \(a\) applies the Brownian RKHS norm formula in[Section2](https://arxiv.org/html/2608.13882#S2), \(b\) substitutes the derivative representation above and evaluates the integrals over the grid intervals, and \(c\) uses \([192](https://arxiv.org/html/2608.13882#A4.E192)\)\. This calculation also proves \([193](https://arxiv.org/html/2608.13882#A4.E193)\) for every function satisfying the stated conditions\. It remains to verify thatqqis reproduced by the interpolation operator\. For everyj∈\{0,…,G−1\}j\\in\\left\\\{0,\\ldots,G\-1\\right\\\}and everyt∈\[tj,tj\+1\]t\\in\\left\[t\_\{j\},t\_\{j\+1\}\\right\], the affinity ofqqgives
q\(t\)\\displaystyle q\\left\(t\\right\)=\(a\)tj\+1−ttj\+1−tjq\(tj\)\+t−tjtj\+1−tjq\(tj\+1\)=\(b\)\(ΠGq\)\(t\)\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\frac\{t\_\{j\+1\}\-t\}\{t\_\{j\+1\}\-t\_\{j\}\}q\\left\(t\_\{j\}\\right\)\+\\frac\{t\-t\_\{j\}\}\{t\_\{j\+1\}\-t\_\{j\}\}q\\left\(t\_\{j\+1\}\\right\)\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\left\(\\Pi\_\{G\}q\\right\)\\left\(t\\right\)\.\(209\)Here, \(a\) is the affine interpolation identity, while \(b\) applies \([14](https://arxiv.org/html/2608.13882#S2.E14)\) withg=qg=q\. Outside\[−A𝒳,A𝒳\]\\left\[\-A\_\{\\mathcal\{X\}\},A\_\{\\mathcal\{X\}\}\\right\], bothqqandΠGq\\Pi\_\{G\}qare constant with the same endpoint values by \([15](https://arxiv.org/html/2608.13882#S2.E15)\) and \([16](https://arxiv.org/html/2608.13882#S2.E16)\)\. Hence,ΠGq=q\\Pi\_\{G\}q=qonℝ\\mathbb\{R\}\. Sinceq∈ℋk\(B\)q\\in\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}and‖q‖ℋk\(B\)≤1\\left\\\|q\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}\\leq 1, the definition \([17](https://arxiv.org/html/2608.13882#S2.E17)\) impliesq∈𝒬G,A𝒳q\\in\\mathcal\{Q\}\_\{G,A\_\{\\mathcal\{X\}\}\}\. This proves[Item1](https://arxiv.org/html/2608.13882#A4.I1.i1)\.
##### Proof of[2](https://arxiv.org/html/2608.13882#A4.I1.i2)\.
We first derive the pointwise Brownian estimate for the finite profile class\. Letq∈𝒬G,A𝒳q\\in\\mathcal\{Q\}\_\{G,A\_\{\\mathcal\{X\}\}\}andt∈ℝt\\in\\mathbb\{R\}\. DefineIt:=\[min\{0,t\},max\{0,t\}\]I\_\{t\}:=\\left\[\\min\\left\\\{0,t\\right\\\},\\max\\left\\\{0,t\\right\\\}\\right\]\. Sinceq\(0\)=0q\\left\(0\\right\)=0,qqis absolutely continuous, and‖q‖ℋk\(B\)≤1\\left\\\|q\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}\\leq 1, we have
\|q\(t\)\|\\displaystyle\\left\|q\\left\(t\\right\)\\right\|=\(a\)\|q\(t\)−q\(0\)\|=\(b\)\|∫0tq′\(s\)𝑑s\|≤\(c\)∫It\|q′\(s\)\|𝑑s≤\(d\)\|It\|1/2\(∫It\|q′\(s\)\|2𝑑s\)1/2\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\left\|q\\left\(t\\right\)\-q\\left\(0\\right\)\\right\|\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\left\|\\int\_\{0\}^\{t\}q^\{\\prime\}\\left\(s\\right\)\\,\\mathrm\{d\}s\\right\|\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{\\leq\}\}\\int\_\{I\_\{t\}\}\\left\|q^\{\\prime\}\\left\(s\\right\)\\right\|\\,\\mathrm\{d\}s\\stackrel\{\{\\scriptstyle\(d\)\}\}\{\{\\leq\}\}\\left\|I\_\{t\}\\right\|^\{1/2\}\\left\(\\int\_\{I\_\{t\}\}\\left\|q^\{\\prime\}\\left\(s\\right\)\\right\|^\{2\}\\,\\mathrm\{d\}s\\right\)^\{1/2\}≤\(e\)\|t\|1/2‖q‖ℋk\(B\)≤\(f\)\|t\|1/2\.\\displaystyle\\stackrel\{\{\\scriptstyle\(e\)\}\}\{\{\\leq\}\}\\left\|t\\right\|^\{1/2\}\\left\\\|q\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}\\stackrel\{\{\\scriptstyle\(f\)\}\}\{\{\\leq\}\}\\left\|t\\right\|^\{1/2\}\.\(210\)Here, \(a\) uses the anchor condition, \(b\) applies the fundamental theorem of calculus, \(c\) accounts for both possible orientations of the integral, \(d\) applies the Cauchy–Schwarz inequality onItI\_\{t\}, \(e\) uses\|It\|=\|t\|\\left\|I\_\{t\}\\right\|=\\left\|t\\right\|and enlarges the derivative integral toℝ\\mathbb\{R\}, and \(f\) uses the unit\-norm property of𝒬G,A𝒳\\mathcal\{Q\}\_\{G,A\_\{\\mathcal\{X\}\}\}\. We first considerL=2L=2\. By \([21](https://arxiv.org/html/2608.13882#S2.E21)\), everya∈𝒜1m,Ga\\in\\mathcal\{A\}\_\{1\}^\{m,G\}has the forma\(𝐱\)=𝝎⊤𝐱a\\left\(\\mathbf\{x\}\\right\)=\\bm\{\\omega\}^\{\\top\}\\mathbf\{x\}for some𝝎∈Ω⊆𝕊d−1\\bm\{\\omega\}\\in\\Omega\\subseteq\\mathbb\{S\}^\{d\-1\}\. Therefore,
\|a\(𝐱\)\|\\displaystyle\\left\|a\\left\(\\mathbf\{x\}\\right\)\\right\|=\(a\)\|𝝎⊤𝐱\|≤\(b\)‖𝝎‖2‖𝐱‖2≤\(c\)R𝒳≤\(d\)A𝒳\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\left\|\\bm\{\\omega\}^\{\\top\}\\mathbf\{x\}\\right\|\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}\\left\\\|\\bm\{\\omega\}\\right\\\|\_\{2\}\\left\\\|\\mathbf\{x\}\\right\\\|\_\{2\}\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{\\leq\}\}R\_\{\\mathcal\{X\}\}\\stackrel\{\{\\scriptstyle\(d\)\}\}\{\{\\leq\}\}A\_\{\\mathcal\{X\}\}\.\(211\)Here, \(a\) uses the definition of the first\-layer support, \(b\) applies the Euclidean Cauchy–Schwarz inequality, \(c\) uses‖𝝎‖2=1\\left\\\|\\bm\{\\omega\}\\right\\\|\_\{2\}=1and the definition ofR𝒳R\_\{\\mathcal\{X\}\}, and \(d\) usesA𝒳=max\{1,R𝒳\}A\_\{\\mathcal\{X\}\}=\\max\\left\\\{1,R\_\{\\mathcal\{X\}\}\\right\\\}\. Since2−\(L−2\)=12^\{\-\(L\-2\)\}=1whenL=2L=2, this proves \([194](https://arxiv.org/html/2608.13882#A4.E194)\) in the base case\. Suppose now thatL≥3L\\geq 3, and setK:=L−2K:=L\-2\. Fix an admissible parameter tuple𝜽\\bm\{\\theta\}and𝐱∈𝒳\\mathbf\{x\}\\in\\mathcal\{X\}\. At the linear layer,
maxj∈\[m\]\|zj\(0\)\(𝐱\)\|\\displaystyle\\max\_\{j\\in\\left\[m\\right\]\}\\left\|z\_\{j\}^\{\(0\)\}\\left\(\\mathbf\{x\}\\right\)\\right\|=\(a\)maxj∈\[m\]\|𝝎j⊤𝐱\|≤\(b\)maxj∈\[m\]‖𝝎j‖2‖𝐱‖2≤\(c\)R𝒳≤\(d\)A𝒳\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\max\_\{j\\in\\left\[m\\right\]\}\\left\|\\bm\{\\omega\}\_\{j\}^\{\\top\}\\mathbf\{x\}\\right\|\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}\\max\_\{j\\in\\left\[m\\right\]\}\\left\\\|\\bm\{\\omega\}\_\{j\}\\right\\\|\_\{2\}\\left\\\|\\mathbf\{x\}\\right\\\|\_\{2\}\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{\\leq\}\}R\_\{\\mathcal\{X\}\}\\stackrel\{\{\\scriptstyle\(d\)\}\}\{\{\\leq\}\}A\_\{\\mathcal\{X\}\}\.\(212\)Here, \(a\) applies \([27](https://arxiv.org/html/2608.13882#S2.E27)\), \(b\) applies the Euclidean Cauchy–Schwarz inequality, \(c\) uses𝝎j∈Ω⊆𝕊d−1\\bm\{\\omega\}\_\{j\}\\in\\Omega\\subseteq\\mathbb\{S\}^\{d\-1\}, and \(d\) applies \([11](https://arxiv.org/html/2608.13882#S2.E11)\)\. At the first profile layer,
\|zj\(1\)\(𝐱\)\|\\displaystyle\\left\|z\_\{j\}^\{\(1\)\}\\left\(\\mathbf\{x\}\\right\)\\right\|=\(a\)\|qj\(1\)\(zj\(0\)\(𝐱\)\)\|≤\(b\)\|zj\(0\)\(𝐱\)\|1/2≤\(c\)A𝒳1/2,j∈\[m\]\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\left\|q\_\{j\}^\{\(1\)\}\\left\(z\_\{j\}^\{\(0\)\}\\left\(\\mathbf\{x\}\\right\)\\right\)\\right\|\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}\\left\|z\_\{j\}^\{\(0\)\}\\left\(\\mathbf\{x\}\\right\)\\right\|^\{1/2\}\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{\\leq\}\}A\_\{\\mathcal\{X\}\}^\{1/2\},\\qquad j\\in\\left\[m\\right\]\.\(213\)Here, \(a\) uses \([28](https://arxiv.org/html/2608.13882#S2.E28)\), \(b\) applies \([210](https://arxiv.org/html/2608.13882#A4.E210)\), and \(c\) uses \([212](https://arxiv.org/html/2608.13882#A4.E212)\)\. We now proceed recursively\. Fixr∈\{2,…,K\}r\\in\\left\\\{2,\\ldots,K\\right\\\}and assume that
maxk∈\[m\]\|zk\(r−1\)\(𝐱\)\|≤A𝒳2−\(r−1\)\.\\displaystyle\\max\_\{k\\in\\left\[m\\right\]\}\\left\|z\_\{k\}^\{\(r\-1\)\}\\left\(\\mathbf\{x\}\\right\)\\right\|\\leq A\_\{\\mathcal\{X\}\}^\{\\,2^\{\-\(r\-1\)\}\}\.\(214\)For everyj∈\[m\]j\\in\\left\[m\\right\],
\|sj\(r\)\(𝐱\)\|\\displaystyle\\left\|s\_\{j\}^\{\(r\)\}\\left\(\\mathbf\{x\}\\right\)\\right\|=\(a\)\|∑k=1mWjk\(r\)zk\(r−1\)\(𝐱\)\|≤\(b\)∑k=1m\|Wjk\(r\)\|\|zk\(r−1\)\(𝐱\)\|≤\(c\)A𝒳2−\(r−1\)∑k=1m\|Wjk\(r\)\|\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\left\|\\sum\_\{k=1\}^\{m\}W\_\{jk\}^\{\(r\)\}z\_\{k\}^\{\(r\-1\)\}\\left\(\\mathbf\{x\}\\right\)\\right\|\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}\\sum\_\{k=1\}^\{m\}\\left\|W\_\{jk\}^\{\(r\)\}\\right\|\\left\|z\_\{k\}^\{\(r\-1\)\}\\left\(\\mathbf\{x\}\\right\)\\right\|\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{\\leq\}\}A\_\{\\mathcal\{X\}\}^\{\\,2^\{\-\(r\-1\)\}\}\\sum\_\{k=1\}^\{m\}\\left\|W\_\{jk\}^\{\(r\)\}\\right\|≤\(d\)A𝒳2−\(r−1\)\.\\displaystyle\\stackrel\{\{\\scriptstyle\(d\)\}\}\{\{\\leq\}\}A\_\{\\mathcal\{X\}\}^\{\\,2^\{\-\(r\-1\)\}\}\.\(215\)Here, \(a\) applies \([29](https://arxiv.org/html/2608.13882#S2.E29)\), \(b\) applies the triangle inequality, \(c\) invokes \([214](https://arxiv.org/html/2608.13882#A4.E214)\), and \(d\) uses the row\-sum constraint \([19](https://arxiv.org/html/2608.13882#S2.E19)\)\. Applying the profile bound gives
\|zj\(r\)\(𝐱\)\|\\displaystyle\\left\|z\_\{j\}^\{\(r\)\}\\left\(\\mathbf\{x\}\\right\)\\right\|=\(a\)\|qj\(r\)\(sj\(r\)\(𝐱\)\)\|≤\(b\)\|sj\(r\)\(𝐱\)\|1/2≤\(c\)A𝒳2−r\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\left\|q\_\{j\}^\{\(r\)\}\\left\(s\_\{j\}^\{\(r\)\}\\left\(\\mathbf\{x\}\\right\)\\right\)\\right\|\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}\\left\|s\_\{j\}^\{\(r\)\}\\left\(\\mathbf\{x\}\\right\)\\right\|^\{1/2\}\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{\\leq\}\}A\_\{\\mathcal\{X\}\}^\{\\,2^\{\-r\}\}\.\(216\)Here, \(a\) applies \([30](https://arxiv.org/html/2608.13882#S2.E30)\), \(b\) applies \([210](https://arxiv.org/html/2608.13882#A4.E210)\), and \(c\) uses \([215](https://arxiv.org/html/2608.13882#A4.E215)\)\. By induction,
maxj∈\[m\]\|zj\(r\)\(𝐱\)\|≤A𝒳2−r,r∈\[K\]\.\\displaystyle\\max\_\{j\\in\\left\[m\\right\]\}\\left\|z\_\{j\}^\{\(r\)\}\\left\(\\mathbf\{x\}\\right\)\\right\|\\leq A\_\{\\mathcal\{X\}\}^\{\\,2^\{\-r\}\},\\qquad r\\in\\left\[K\\right\]\.\(217\)SinceA𝒳≥1A\_\{\\mathcal\{X\}\}\\geq 1, \([215](https://arxiv.org/html/2608.13882#A4.E215)\) also gives
\|sj\(r\)\(𝐱\)\|≤A𝒳,r∈\{2,…,K\},j∈\[m\]\.\\displaystyle\\left\|s\_\{j\}^\{\(r\)\}\\left\(\\mathbf\{x\}\\right\)\\right\|\\leq A\_\{\\mathcal\{X\}\},\\qquad r\\in\\left\\\{2,\\ldots,K\\right\\\},\\quad j\\in\\left\[m\\right\]\.\(218\)Thus, every profile is evaluated within the interpolation interval used in the finite architecture\. Finally, using the readout constraint,
\|a𝜽\(𝐱\)\|\\displaystyle\\left\|a\_\{\\bm\{\\theta\}\}\\left\(\\mathbf\{x\}\\right\)\\right\|=\(a\)\|∑j=1mβjzj\(K\)\(𝐱\)\|≤\(b\)∑j=1m\|βj\|\|zj\(K\)\(𝐱\)\|≤\(c\)A𝒳2−K∑j=1m\|βj\|≤\(d\)A𝒳2−K=\(e\)A𝒳2−\(L−2\)\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\left\|\\sum\_\{j=1\}^\{m\}\\beta\_\{j\}z\_\{j\}^\{\(K\)\}\\left\(\\mathbf\{x\}\\right\)\\right\|\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}\\sum\_\{j=1\}^\{m\}\\left\|\\beta\_\{j\}\\right\|\\left\|z\_\{j\}^\{\(K\)\}\\left\(\\mathbf\{x\}\\right\)\\right\|\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{\\leq\}\}A\_\{\\mathcal\{X\}\}^\{\\,2^\{\-K\}\}\\sum\_\{j=1\}^\{m\}\\left\|\\beta\_\{j\}\\right\|\\stackrel\{\{\\scriptstyle\(d\)\}\}\{\{\\leq\}\}A\_\{\\mathcal\{X\}\}^\{\\,2^\{\-K\}\}\\stackrel\{\{\\scriptstyle\(e\)\}\}\{\{=\}\}A\_\{\\mathcal\{X\}\}^\{\\,2^\{\-\(L\-2\)\}\}\.\(219\)Here, \(a\) applies \([31](https://arxiv.org/html/2608.13882#S2.E31)\), \(b\) applies the triangle inequality, \(c\) uses \([217](https://arxiv.org/html/2608.13882#A4.E217)\), \(d\) uses \([20](https://arxiv.org/html/2608.13882#S2.E20)\), and \(e\) substitutesK=L−2K=L\-2\. This proves \([194](https://arxiv.org/html/2608.13882#A4.E194)\) and \([195](https://arxiv.org/html/2608.13882#A4.E195)\)\. It remains to establish the range of the associated outer dictionary\. Fixf∈𝒟Lm,Gf\\in\\mathcal\{D\}\_\{L\}^\{m,G\}\. By \([37](https://arxiv.org/html/2608.13882#S2.E37)\), there existsa∈𝒜L−1m,Ga\\in\\mathcal\{A\}\_\{L\-1\}^\{m,G\}such thatf∈ℋka,‖f‖ℋka≤1f\\in\\mathcal\{H\}\_\{k\_\{a\}\},\\qquad\\left\\\|f\\right\\\|\_\{\\mathcal\{H\}\_\{k\_\{a\}\}\}\\leq 1\. For every𝐱∈𝒳\\mathbf\{x\}\\in\\mathcal\{X\}, the reproducing property gives
\|f\(𝐱\)\|\\displaystyle\\left\|f\\left\(\\mathbf\{x\}\\right\)\\right\|=\(a\)\|⟨f,ka\(𝐱,⋅\)⟩ℋka\|≤\(b\)‖f‖ℋka‖ka\(𝐱,⋅\)‖ℋka=\(c\)‖f‖ℋkaka\(𝐱,𝐱\)1/2\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\left\|\\left\\langle f,k\_\{a\}\\left\(\\mathbf\{x\},\\cdot\\right\)\\right\\rangle\_\{\\mathcal\{H\}\_\{k\_\{a\}\}\}\\right\|\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}\\left\\\|f\\right\\\|\_\{\\mathcal\{H\}\_\{k\_\{a\}\}\}\\left\\\|k\_\{a\}\\left\(\\mathbf\{x\},\\cdot\\right\)\\right\\\|\_\{\\mathcal\{H\}\_\{k\_\{a\}\}\}\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{=\}\}\\left\\\|f\\right\\\|\_\{\\mathcal\{H\}\_\{k\_\{a\}\}\}k\_\{a\}\\left\(\\mathbf\{x\},\\mathbf\{x\}\\right\)^\{1/2\}≤\(d\)\|a\(𝐱\)\|1/2≤\(e\)A𝒳2−\(L−1\)\.\\displaystyle\\stackrel\{\{\\scriptstyle\(d\)\}\}\{\{\\leq\}\}\\left\|a\\left\(\\mathbf\{x\}\\right\)\\right\|^\{1/2\}\\stackrel\{\{\\scriptstyle\(e\)\}\}\{\{\\leq\}\}A\_\{\\mathcal\{X\}\}^\{\\,2^\{\-\(L\-1\)\}\}\.\(220\)Here, \(a\) applies the reproducing property inℋka\\mathcal\{H\}\_\{k\_\{a\}\}, \(b\) applies the Cauchy–Schwarz inequality, \(c\) uses the standard RKHS section\-norm identity, \(d\) uses‖f‖ℋka≤1\\left\\\|f\\right\\\|\_\{\\mathcal\{H\}\_\{k\_\{a\}\}\}\\leq 1and
ka\(𝐱,𝐱\)=k\(B\)\(a\(𝐱\),a\(𝐱\)\)=\|a\(𝐱\)\|,\\displaystyle k\_\{a\}\\left\(\\mathbf\{x\},\\mathbf\{x\}\\right\)=k^\{\(\\mathrm\{B\}\)\}\\left\(a\\left\(\\mathbf\{x\}\\right\),a\\left\(\\mathbf\{x\}\\right\)\\right\)=\\left\|a\\left\(\\mathbf\{x\}\\right\)\\right\|,\(221\)and \(e\) applies \([194](https://arxiv.org/html/2608.13882#A4.E194)\)\. This proves \([196](https://arxiv.org/html/2608.13882#A4.E196)\)\.
##### Proof of[3](https://arxiv.org/html/2608.13882#A4.I1.i3)\.
We first considerL=2L=2\. In this case,
𝒜1m,G=𝒰1=\{𝐱↦𝝎⊤𝐱:𝝎∈Ω\}\.\\displaystyle\\mathcal\{A\}\_\{1\}^\{m,G\}=\\mathcal\{U\}\_\{1\}=\\left\\\{\\mathbf\{x\}\\mapsto\\bm\{\\omega\}^\{\\top\}\\mathbf\{x\}:\\bm\{\\omega\}\\in\\Omega\\right\\\}\.\(222\)Every such function is specified by theddcoordinates of𝝎\\bm\{\\omega\}\. Therefore,P1,m,G=dP\_\{1,m,G\}=d, which proves \([198](https://arxiv.org/html/2608.13882#A4.E198)\)\. Suppose now thatL≥3L\\geq 3, and setK:=L−2K:=L\-2\. We count the four parameter groups appearing in \([22](https://arxiv.org/html/2608.13882#S2.E22)\)\. First, the architecture containsmmfirst\-layer directions𝝎1,…,𝝎m∈ℝd\\bm\{\\omega\}\_\{1\},\\ldots,\\bm\{\\omega\}\_\{m\}\\in\\mathbb\{R\}^\{d\}\. Hence, the directions contribute at most
Pdir\\displaystyle P\_\{\\mathrm\{dir\}\}≤\(a\)md\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}md\.\(223\)Here, \(a\) countsddambient coordinates for each of themmdirections\. The restriction𝝎j∈Ω\\bm\{\\omega\}\_\{j\}\\in\\Omegadoes not increase this upper bound\. Second, each of theKKprofile layers containsmmfinite Brownian profiles\. Every profile is specified by at mostG\+1G\+1nodal coefficients\. Therefore, the profile parameters contribute at most
Pprof\\displaystyle P\_\{\\mathrm\{prof\}\}≤\(a\)Km\(G\+1\)\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}Km\\left\(G\+1\\right\)\.\(224\)Here, \(a\) multiplies the numberKmKmof profiles by the upper boundG\+1G\+1on the number of nodal coordinates per profile\. The anchor condition and the energy constraint restrict the admissible parameter set but do not increase its ambient dimension\. Third, mixing matrices occur at the levelsr∈\{2,…,K\}r\\in\\left\\\{2,\\ldots,K\\right\\\}\. WhenK=1K=1, this index set is empty\. WhenK≥2K\\geq 2, it containsK−1K\-1indices\. Each matrix belongs toℝm×m\\mathbb\{R\}^\{m\\times m\}and therefore contains at mostm2m^\{2\}real coordinates\. Consequently,
Pmix\\displaystyle P\_\{\\mathrm\{mix\}\}≤\(a\)\(K−1\)m2\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}\\left\(K\-1\\right\)m^\{2\}\.\(225\)Here, \(a\) multiplies the number of internal mixing levels by the number of entries in each matrix\. The row\-sum constraints restrict the admissible matrices but introduce no additional parameters\. Fourth, the final readout vector𝜷∈ℝm\\bm\{\\beta\}\\in\\mathbb\{R\}^\{m\}contributes at most
Pout\\displaystyle P\_\{\\mathrm\{out\}\}≤\(a\)m\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}m\(226\)real parameters\. Here, \(a\) counts the entries of the readout vector\. Again, theℓ1\\ell^\{1\}constraint restricts the admissible set without increasing its ambient dimension\. Combining \([223](https://arxiv.org/html/2608.13882#A4.E223)\), \([224](https://arxiv.org/html/2608.13882#A4.E224)\), \([225](https://arxiv.org/html/2608.13882#A4.E225)\), and \([226](https://arxiv.org/html/2608.13882#A4.E226)\) gives
PL−1,m,G\\displaystyle P\_\{L\-1,m,G\}≤\(a\)Pdir\+Pprof\+Pmix\+Pout\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}P\_\{\\mathrm\{dir\}\}\+P\_\{\\mathrm\{prof\}\}\+P\_\{\\mathrm\{mix\}\}\+P\_\{\\mathrm\{out\}\}≤\(b\)md\+Km\(G\+1\)\+\(K−1\)m2\+m\\displaystyle\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}md\+Km\\left\(G\+1\\right\)\+\\left\(K\-1\\right\)m^\{2\}\+m=\(c\)md\+\(L−2\)m\(G\+1\)\+\(L−3\)m2\+m\.\\displaystyle\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{=\}\}md\+\\left\(L\-2\\right\)m\\left\(G\+1\\right\)\+\\left\(L\-3\\right\)m^\{2\}\+m\.\(227\)Here, \(a\) separates the four parameter groups, \(b\) substitutes their individual upper bounds, and \(c\) usesK=L−2K=L\-2\. This proves \([197](https://arxiv.org/html/2608.13882#A4.E197)\) and completes the proof\. ∎
###### Lemma B0\.
\(Recursive RKHS structure of the Variation BKL dictionary\)Fixl≥2l\\geq 2andu∈𝒰l−1u\\in\\mathcal\{U\}\_\{l\-1\}\. Letkuk\_\{u\}be the Brownian pullback kernel defined in \([2](https://arxiv.org/html/2608.13882#S2.E2)\), and letℋku\\mathcal\{H\}\_\{k\_\{u\}\}denote its RKHS\. DefineTu:ℋk\(B\)⟶ℝ𝒳T\_\{u\}:\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\\longrightarrow\\mathbb\{R\}^\{\\mathcal\{X\}\},\(Tug\)\(𝐱\):=g\(u\(𝐱\)\),𝐱∈𝒳,\\left\(T\_\{u\}g\\right\)\\left\(\\mathbf\{x\}\\right\):=g\\left\(u\\left\(\\mathbf\{x\}\\right\)\\right\),\\quad\\mathbf\{x\}\\in\\mathcal\{X\},and𝒩u:=ker\(Tu\)=\{g∈ℋk\(B\):g\(u\(𝐱\)\)=0for every𝐱∈𝒳\}\.\\mathcal\{N\}\_\{u\}:=\\ker\\left\(T\_\{u\}\\right\)=\\left\\\{g\\in\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}:g\\left\(u\\left\(\\mathbf\{x\}\\right\)\\right\)=0\\text\{ for every \}\\mathbf\{x\}\\in\\mathcal\{X\}\\right\\\}\.Then𝒩u\\mathcal\{N\}\_\{u\}is a closed linear subspace ofℋk\(B\)\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\. Let𝒩u⟂\\mathcal\{N\}\_\{u\}^\{\\perp\}denote its orthogonal complement inℋk\(B\)\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\. Then
ℋku=Tu\(ℋk\(B\)\)=Tu\(𝒩u⟂\)\.\\displaystyle\\mathcal\{H\}\_\{k\_\{u\}\}=T\_\{u\}\\left\(\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\\right\)=T\_\{u\}\\left\(\\mathcal\{N\}\_\{u\}^\{\\perp\}\\right\)\.\(228\)as sets of functions on𝒳\\mathcal\{X\}\. More precisely, for everyf∈ℋkuf\\in\\mathcal\{H\}\_\{k\_\{u\}\}, there exists a uniquegf∈𝒩u⟂g\_\{f\}\\in\\mathcal\{N\}\_\{u\}^\{\\perp\}such that
Tugf=f\.\\displaystyle T\_\{u\}g\_\{f\}=f\.\(229\)This representative satisfies
‖f‖ℋku\\displaystyle\\left\\\|f\\right\\\|\_\{\\mathcal\{H\}\_\{k\_\{u\}\}\}=‖gf‖ℋk\(B\)\\displaystyle=\\left\\\|g\_\{f\}\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}=min\{‖g‖ℋk\(B\):g∈ℋk\(B\),Tug=f\}\.\\displaystyle=\\min\\left\\\{\\left\\\|g\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}:g\\in\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\},\\;T\_\{u\}g=f\\right\\\}\.\(230\)Consequently,
\{Tug:g∈ℋk\(B\),‖g‖ℋk\(B\)≤1\}=\{f∈ℋku:‖f‖ℋku≤1\}\.\\displaystyle\\left\\\{T\_\{u\}g:g\\in\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\},\\;\\left\\\|g\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}\\leq 1\\right\\\}=\\left\\\{f\\in\\mathcal\{H\}\_\{k\_\{u\}\}:\\left\\\|f\\right\\\|\_\{\\mathcal\{H\}\_\{k\_\{u\}\}\}\\leq 1\\right\\\}\.\(231\)Therefore, the depth\-llrecursive dictionary satisfies
𝒰l=⋃u∈𝒰l−1\{f∈ℋku:‖f‖ℋku≤1\},\\displaystyle\\mathcal\{U\}\_\{l\}=\\bigcup\_\{u\\in\\mathcal\{U\}\_\{l\-1\}\}\\left\\\{f\\in\\mathcal\{H\}\_\{k\_\{u\}\}:\\left\\\|f\\right\\\|\_\{\\mathcal\{H\}\_\{k\_\{u\}\}\}\\leq 1\\right\\\},\(232\)or equivalently,
𝒰l=\{𝐱↦g\(u\(𝐱\)\):u∈𝒰l−1,g∈ℋk\(B\),‖g‖ℋk\(B\)≤1\}\.\\displaystyle\\mathcal\{U\}\_\{l\}=\\left\\\{\\mathbf\{x\}\\mapsto g\\left\(u\\left\(\\mathbf\{x\}\\right\)\\right\):u\\in\\mathcal\{U\}\_\{l\-1\},\\;g\\in\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\},\\;\\left\\\|g\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}\\leq 1\\right\\\}\.\(233\)Let𝒱\(l\)\\mathcal\{V\}^\{\(l\)\}denote the depth\-llvariation space obtained by replacingLLwithllin \([7](https://arxiv.org/html/2608.13882#S2.E7)\)\. Then
𝒱\(l\)=\{F∈L2\(ν\):F=∫𝒰ladρ\(a\),ρ∈ℳ\(𝒰l\),‖ρ‖TV<∞\}\.\\displaystyle\\mathcal\{V\}^\{\(l\)\}=\\left\\\{F\\in L^\{2\}\\left\(\\nu\\right\):F=\\int\_\{\\mathcal\{U\}\_\{l\}\}a\\,\\mathrm\{d\}\\rho\\left\(a\\right\),\\;\\rho\\in\\mathcal\{M\}\\left\(\\mathcal\{U\}\_\{l\}\\right\),\\;\\left\\\|\\rho\\right\\\|\_\{\\mathrm\{TV\}\}<\\infty\\right\\\}\.\(234\)Thus,𝒱\(l\)\\mathcal\{V\}^\{\(l\)\}is the variation hull generated by the recursively constructed union of Brownian pullback RKHS unit balls\.
###### Proof\.
Fixl≥2l\\geq 2andu∈𝒰l−1u\\in\\mathcal\{U\}\_\{l\-1\}, and set
Zu:=u\(𝒳\)\.\\displaystyle Z\_\{u\}:=u\\left\(\\mathcal\{X\}\\right\)\.\(235\)Apply the kernel pullback and restriction result \([18](https://arxiv.org/html/2608.13882#bib.bib1), Lemma C4\(i\)–\(iii\)\) with
Z0\\displaystyle Z\_\{0\}=ℝ,\\displaystyle=\\mathbb\{R\},Z\\displaystyle Z=Zu,\\displaystyle=Z\_\{u\},k\\displaystyle k=k\(B\)\.\\displaystyle=k^\{\(\\mathrm\{B\}\)\}\.\(236\)It follows that, for everyg∈ℋk\(B\)g\\in\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\},
Tug\\displaystyle T\_\{u\}g∈ℋku,\\displaystyle\\in\\mathcal\{H\}\_\{k\_\{u\}\},‖Tug‖ℋku\\displaystyle\\left\\\|T\_\{u\}g\\right\\\|\_\{\\mathcal\{H\}\_\{k\_\{u\}\}\}≤‖g‖ℋk\(B\)\.\\displaystyle\\leq\\left\\\|g\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}\.\(237\)Moreover, for everyf∈ℋkuf\\in\\mathcal\{H\}\_\{k\_\{u\}\}, there exists a minimum\-norm representativegf∈ℋk\(B\)g\_\{f\}\\in\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}such that
Tugf\\displaystyle T\_\{u\}g\_\{f\}=f,\\displaystyle=f,‖gf‖ℋk\(B\)\\displaystyle\\left\\\|g\_\{f\}\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}=‖f‖ℋku\.\\displaystyle=\\left\\\|f\\right\\\|\_\{\\mathcal\{H\}\_\{k\_\{u\}\}\}\.\(238\)The operatorTu:ℋk\(B\)→ℋkuT\_\{u\}:\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\\rightarrow\\mathcal\{H\}\_\{k\_\{u\}\}is linear and continuous\. Therefore,
𝒩u=ker\(Tu\)\\displaystyle\\mathcal\{N\}\_\{u\}=\\ker\\left\(T\_\{u\}\\right\)\(239\)is a closed linear subspace ofℋk\(B\)\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\. LetPuP\_\{u\}denote the orthogonal projection ofℋk\(B\)\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}onto𝒩u\\mathcal\{N\}\_\{u\}\. SincePugf∈𝒩uP\_\{u\}g\_\{f\}\\in\\mathcal\{N\}\_\{u\}, one hasTu\(Pugf\)=0T\_\{u\}\\left\(P\_\{u\}g\_\{f\}\\right\)=0\. Therefore,
Tu\(gf−Pugf\)\\displaystyle T\_\{u\}\\left\(g\_\{f\}\-P\_\{u\}g\_\{f\}\\right\)=\(a\)Tugf−Tu\(Pugf\)=\(b\)f\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}T\_\{u\}g\_\{f\}\-T\_\{u\}\\left\(P\_\{u\}g\_\{f\}\\right\)\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}f\.\(240\)Here, \(a\) applies the linearity ofTuT\_\{u\}, while \(b\) usesTugf=fT\_\{u\}g\_\{f\}=fandTu\(Pugf\)=0T\_\{u\}\\left\(P\_\{u\}g\_\{f\}\\right\)=0\. Suppose thatPugf≠0P\_\{u\}g\_\{f\}\\neq 0\. The vectorsgf−Pugfg\_\{f\}\-P\_\{u\}g\_\{f\}andPugfP\_\{u\}g\_\{f\}are orthogonal\. Thus,
‖gf−Pugf‖ℋk\(B\)2\\displaystyle\\left\\\|g\_\{f\}\-P\_\{u\}g\_\{f\}\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}^\{2\}=\(a\)‖gf‖ℋk\(B\)2−‖Pugf‖ℋk\(B\)2<\(b\)‖gf‖ℋk\(B\)2\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\left\\\|g\_\{f\}\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}^\{2\}\-\\left\\\|P\_\{u\}g\_\{f\}\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}^\{2\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{<\}\}\\left\\\|g\_\{f\}\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}^\{2\}\.\(241\)Here, \(a\) applies the Pythagorean identity to the orthogonal decomposition
gf=\(gf−Pugf\)\+Pugf,\\displaystyle g\_\{f\}=\\left\(g\_\{f\}\-P\_\{u\}g\_\{f\}\\right\)\+P\_\{u\}g\_\{f\},\(242\)while \(b\) usesPugf≠0P\_\{u\}g\_\{f\}\\neq 0\. The preceding two displays produce another representative offfwith strictly smallerℋk\(B\)\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\-norm, contradicting the minimum\-norm property ofgfg\_\{f\}\. Hence,
Pugf\\displaystyle P\_\{u\}g\_\{f\}=0,\\displaystyle=0,gf\\displaystyle g\_\{f\}∈𝒩u⟂\.\\displaystyle\\in\\mathcal\{N\}\_\{u\}^\{\\perp\}\.\(243\)We next prove uniqueness in𝒩u⟂\\mathcal\{N\}\_\{u\}^\{\\perp\}\. Lethf∈𝒩u⟂h\_\{f\}\\in\\mathcal\{N\}\_\{u\}^\{\\perp\}satisfyTuhf=fT\_\{u\}h\_\{f\}=f\. Hence,
Tu\(gf−hf\)\\displaystyle T\_\{u\}\\left\(g\_\{f\}\-h\_\{f\}\\right\)=\(a\)Tugf−Tuhf=\(b\)0,\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}T\_\{u\}g\_\{f\}\-T\_\{u\}h\_\{f\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}0,gf−hf\\displaystyle g\_\{f\}\-h\_\{f\}∈\(c\)𝒩u∩𝒩u⟂=\(d\)\{0\}\.\\displaystyle\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{\\in\}\}\\mathcal\{N\}\_\{u\}\\cap\\mathcal\{N\}\_\{u\}^\{\\perp\}\\stackrel\{\{\\scriptstyle\(d\)\}\}\{\{=\}\}\\left\\\{0\\right\\\}\.\(244\)Here, \(a\) applies the linearity ofTuT\_\{u\}, \(b\) usesTugf=Tuhf=fT\_\{u\}g\_\{f\}=T\_\{u\}h\_\{f\}=f, \(c\) follows from the first line together withgf,hf∈𝒩u⟂g\_\{f\},h\_\{f\}\\in\\mathcal\{N\}\_\{u\}^\{\\perp\}, and \(d\) uses
𝒩u∩𝒩u⟂=\{0\}\.\\displaystyle\\mathcal\{N\}\_\{u\}\\cap\\mathcal\{N\}\_\{u\}^\{\\perp\}=\\left\\\{0\\right\\\}\.\(245\)Thus,gf=hfg\_\{f\}=h\_\{f\}, which proves \([229](https://arxiv.org/html/2608.13882#A4.E229)\)\. The preceding construction also identifies the range of the pullback operator\. Consequently,
ℋku\\displaystyle\\mathcal\{H\}\_\{k\_\{u\}\}⊆\(a\)Tu\(𝒩u⟂\)⊆\(b\)Tu\(ℋk\(B\)\)⊆\(c\)ℋku\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\subseteq\}\}T\_\{u\}\\left\(\\mathcal\{N\}\_\{u\}^\{\\perp\}\\right\)\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\subseteq\}\}T\_\{u\}\\left\(\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\\right\)\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{\\subseteq\}\}\\mathcal\{H\}\_\{k\_\{u\}\}\.\(246\)Here, \(a\) follows because everyf∈ℋkuf\\in\\mathcal\{H\}\_\{k\_\{u\}\}has a representativegf∈𝒩u⟂g\_\{f\}\\in\\mathcal\{N\}\_\{u\}^\{\\perp\}, \(b\) uses𝒩u⟂⊆ℋk\(B\)\\mathcal\{N\}\_\{u\}^\{\\perp\}\\subseteq\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}, and \(c\) applies the pullback contraction established above\. Therefore, all three sets are equal, which proves \([228](https://arxiv.org/html/2608.13882#A4.E228)\)\. Letg∈ℋk\(B\)g\\in\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}be any representative off∈ℋkuf\\in\\mathcal\{H\}\_\{k\_\{u\}\}, so thatTug=fT\_\{u\}g=f\. Therefore,
‖f‖ℋku\\displaystyle\\left\\\|f\\right\\\|\_\{\\mathcal\{H\}\_\{k\_\{u\}\}\}=\(a\)‖Tug‖ℋku≤\(b\)‖g‖ℋk\(B\)\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\left\\\|T\_\{u\}g\\right\\\|\_\{\\mathcal\{H\}\_\{k\_\{u\}\}\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}\\left\\\|g\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}\.\(247\)Here, \(a\) usesTug=fT\_\{u\}g=f, while \(b\) applies the pullback contraction\. The minimum\-norm representativegfg\_\{f\}is admissible and attains equality in this lower bound\. Consequently,
‖f‖ℋku\\displaystyle\\left\\\|f\\right\\\|\_\{\\mathcal\{H\}\_\{k\_\{u\}\}\}=\(a\)‖gf‖ℋk\(B\)=\(b\)min\{‖g‖ℋk\(B\):g∈ℋk\(B\),Tug=f\}\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\left\\\|g\_\{f\}\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\min\\left\\\{\\left\\\|g\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}:g\\in\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\},\\;T\_\{u\}g=f\\right\\\}\.\(248\)Here, \(a\) applies the minimum\-extension identity established above, while \(b\) combines the preceding lower bound with the admissibility ofgfg\_\{f\}\. This proves \([230](https://arxiv.org/html/2608.13882#A4.E230)\)\. We now prove the exact unit\-ball identity\. Letg∈ℋk\(B\)g\\in\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}satisfy‖g‖ℋk\(B\)≤1\\left\\\|g\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}\\leq 1\. Hence,
‖Tug‖ℋku\\displaystyle\\left\\\|T\_\{u\}g\\right\\\|\_\{\\mathcal\{H\}\_\{k\_\{u\}\}\}≤\(a\)‖g‖ℋk\(B\)≤\(b\)1\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}\\left\\\|g\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}1\.\(249\)Here, \(a\) applies the pullback contraction, while \(b\) uses the assumed norm bound ongg\. Therefore,
\{Tug:g∈ℋk\(B\),‖g‖ℋk\(B\)≤1\}⊆\{f∈ℋku:‖f‖ℋku≤1\}\.\\displaystyle\\left\\\{T\_\{u\}g:g\\in\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\},\\;\\left\\\|g\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}\\leq 1\\right\\\}\\subseteq\\left\\\{f\\in\\mathcal\{H\}\_\{k\_\{u\}\}:\\left\\\|f\\right\\\|\_\{\\mathcal\{H\}\_\{k\_\{u\}\}\}\\leq 1\\right\\\}\.\(250\)Conversely, letf∈ℋkuf\\in\\mathcal\{H\}\_\{k\_\{u\}\}satisfy‖f‖ℋku≤1\\left\\\|f\\right\\\|\_\{\\mathcal\{H\}\_\{k\_\{u\}\}\}\\leq 1\. Its minimum\-norm representative satisfiesTugf=fT\_\{u\}g\_\{f\}=f\. Hence,
‖gf‖ℋk\(B\)\\displaystyle\\left\\\|g\_\{f\}\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}=\(a\)‖f‖ℋku≤\(b\)1\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\left\\\|f\\right\\\|\_\{\\mathcal\{H\}\_\{k\_\{u\}\}\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}1\.\(251\)Here, \(a\) applies the minimum\-extension identity, while \(b\) uses the assumed norm bound onff\. Therefore,
\{f∈ℋku:‖f‖ℋku≤1\}⊆\{Tug:g∈ℋk\(B\),‖g‖ℋk\(B\)≤1\}\.\\displaystyle\\left\\\{f\\in\\mathcal\{H\}\_\{k\_\{u\}\}:\\left\\\|f\\right\\\|\_\{\\mathcal\{H\}\_\{k\_\{u\}\}\}\\leq 1\\right\\\}\\subseteq\\left\\\{T\_\{u\}g:g\\in\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\},\\;\\left\\\|g\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}\\leq 1\\right\\\}\.\(252\)Combining the two inclusions proves \([231](https://arxiv.org/html/2608.13882#A4.E231)\)\. Finally, apply the unit\-ball identity to the recursive definition of the atomic dictionary\. Therefore,
𝒰l\\displaystyle\\mathcal\{U\}\_\{l\}=\(a\)⋃u∈𝒰l−1\{Tug:g∈ℋk\(B\),‖g‖ℋk\(B\)≤1\}\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\bigcup\_\{u\\in\\mathcal\{U\}\_\{l\-1\}\}\\left\\\{T\_\{u\}g:g\\in\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\},\\;\\left\\\|g\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}\\leq 1\\right\\\}=\(b\)⋃u∈𝒰l−1\{f∈ℋku:‖f‖ℋku≤1\}\.\\displaystyle\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\bigcup\_\{u\\in\\mathcal\{U\}\_\{l\-1\}\}\\left\\\{f\\in\\mathcal\{H\}\_\{k\_\{u\}\}:\\left\\\|f\\right\\\|\_\{\\mathcal\{H\}\_\{k\_\{u\}\}\}\\leq 1\\right\\\}\.\(253\)Here, \(a\) applies the compositional definition of𝒰l\\mathcal\{U\}\_\{l\}, while \(b\) applies \([231](https://arxiv.org/html/2608.13882#A4.E231)\) for every fixedu∈𝒰l−1u\\in\\mathcal\{U\}\_\{l\-1\}\. This proves \([232](https://arxiv.org/html/2608.13882#A4.E232)\)\. By the definition ofTuT\_\{u\}, \([233](https://arxiv.org/html/2608.13882#A4.E233)\) is the equivalent compositional form of the same identity\. By the definition of the depth\-llVBKL space,
𝒱\(l\)=\{F∈L2\(ν\):F=∫𝒰ladρ\(a\),ρ∈ℳ\(𝒰l\),‖ρ‖TV<∞\}\.\\displaystyle\\mathcal\{V\}^\{\(l\)\}=\\left\\\{F\\in L^\{2\}\\left\(\\nu\\right\):F=\\int\_\{\\mathcal\{U\}\_\{l\}\}a\\,\\mathrm\{d\}\\rho\\left\(a\\right\),\\;\\rho\\in\\mathcal\{M\}\\left\(\\mathcal\{U\}\_\{l\}\\right\),\\;\\left\\\|\\rho\\right\\\|\_\{\\mathrm\{TV\}\}<\\infty\\right\\\}\.\(254\)This proves \([234](https://arxiv.org/html/2608.13882#A4.E234)\) and completes the proof\. ∎
###### Lemma B0\.
\(Exact reduction from finite variation hulls to symmetric dictionaries\)Letn∈ℕn\\in\\mathbb\{N\},n≥1n\\geq 1, and let𝐱1,…,𝐱n∈𝒳\\mathbf\{x\}\_\{1\},\\ldots,\\mathbf\{x\}\_\{n\}\\in\\mathcal\{X\}be fixed sample points\. Let𝒟\\mathcal\{D\}be a nonempty class of real\-valued functions on𝒳\\mathcal\{X\}, equipped with aσ\\sigma\-algebra for which every evaluation mapf⟼f\(𝐱\)f\\longmapsto f\\left\(\\mathbf\{x\}\\right\),𝐱∈𝒳,\\mathbf\{x\}\\in\\mathcal\{X\},is measurable\. Assume that𝒟\\mathcal\{D\}is symmetric:
𝒟=−𝒟:=\{−f:f∈𝒟\}\.\\displaystyle\\mathcal\{D\}=\-\\mathcal\{D\}:=\\left\\\{\-f:f\\in\\mathcal\{D\}\\right\\\}\.\(255\)
ForR\>0R\>0, define the signed\-measure variation hull generated by𝒟\\mathcal\{D\}as
𝒲R\(𝒟\):=\{Fμ:Fμ\(𝐱\)=∫𝒟f\(𝐱\)dμ\(f\),μ∈ℳ\(𝒟\),‖μ‖TV≤R\}\.\\displaystyle\\mathcal\{W\}\_\{R\}\\left\(\\mathcal\{D\}\\right\):=\\left\\\{F\_\{\\mu\}:F\_\{\\mu\}\\left\(\\mathbf\{x\}\\right\)=\\int\_\{\\mathcal\{D\}\}f\\left\(\\mathbf\{x\}\\right\)\\,\\mathrm\{d\}\\mu\\left\(f\\right\),\\;\\mu\\in\\mathcal\{M\}\\left\(\\mathcal\{D\}\\right\),\\;\\left\\\|\\mu\\right\\\|\_\{\\mathrm\{TV\}\}\\leq R\\right\\\}\.\(256\)Assume that every such integral is well defined at𝐱1,…,𝐱n\\mathbf\{x\}\_\{1\},\\ldots,\\mathbf\{x\}\_\{n\}, and define
M𝐱\(𝒟\):=supf∈𝒟maxi∈\[n\]\|f\(𝐱i\)\|<∞\.\\displaystyle M\_\{\\mathbf\{x\}\}\\left\(\\mathcal\{D\}\\right\):=\\sup\_\{f\\in\\mathcal\{D\}\}\\max\_\{i\\in\\left\[n\\right\]\}\\left\|f\\left\(\\mathbf\{x\}\_\{i\}\\right\)\\right\|<\\infty\.\(257\)Then
ℜ^n\(𝒲R\(𝒟\)\)\\displaystyle\\widehat\{\\mathfrak\{R\}\}\_\{n\}\\left\(\\mathcal\{W\}\_\{R\}\\left\(\\mathcal\{D\}\\right\)\\right\)=Rℜ^n\(𝒟\)\\displaystyle=R\\widehat\{\\mathfrak\{R\}\}\_\{n\}\\left\(\\mathcal\{D\}\\right\)≤RM𝐱\(𝒟\)\.\\displaystyle\\leq RM\_\{\\mathbf\{x\}\}\\left\(\\mathcal\{D\}\\right\)\.\(258\)In particular, fixL≥2L\\geq 2,m≥1m\\geq 1, andG≥2G\\geq 2\. The finite outer Brownian dictionary𝒟Lm,G\\mathcal\{D\}\_\{L\}^\{m,G\}defined in \([37](https://arxiv.org/html/2608.13882#S2.E37)\) is symmetric\. Define
B𝐱:=supa∈𝒜L−1m,Gmaxi∈\[n\]\|a\(𝐱i\)\|\.\\displaystyle B\_\{\\mathbf\{x\}\}:=\\sup\_\{a\\in\\mathcal\{A\}\_\{L\-1\}^\{m,G\}\}\\max\_\{i\\in\\left\[n\\right\]\}\\left\|a\\left\(\\mathbf\{x\}\_\{i\}\\right\)\\right\|\.\(259\)Then
M𝐱\(𝒟Lm,G\)≤B𝐱1/2\.\\displaystyle M\_\{\\mathbf\{x\}\}\\left\(\\mathcal\{D\}\_\{L\}^\{m,G\}\\right\)\\leq B\_\{\\mathbf\{x\}\}^\{1/2\}\.\(260\)Consequently,
ℜ^n\(𝒲RL;m,G\)\\displaystyle\\widehat\{\\mathfrak\{R\}\}\_\{n\}\\left\(\\mathcal\{W\}\_\{R\}^\{L;m,G\}\\right\)=Rℜ^n\(𝒟Lm,G\)\\displaystyle=R\\widehat\{\\mathfrak\{R\}\}\_\{n\}\\left\(\\mathcal\{D\}\_\{L\}^\{m,G\}\\right\)≤RB𝐱1/2\.\\displaystyle\\leq RB\_\{\\mathbf\{x\}\}^\{1/2\}\.\(261\)
###### Proof\.
Fix sample points𝐱1,…,𝐱n∈𝒳\\mathbf\{x\}\_\{1\},\\ldots,\\mathbf\{x\}\_\{n\}\\in\\mathcal\{X\}and fix a realizationε1,…,εn\\varepsilon\_\{1\},\\ldots,\\varepsilon\_\{n\}of the Rademacher variables\. Define the sample\-dependent linear functional
Φε\(f\):=1n∑i=1nεif\(𝐱i\),f∈𝒟\.\\displaystyle\\Phi\_\{\\varepsilon\}\\left\(f\\right\):=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\varepsilon\_\{i\}f\\left\(\\mathbf\{x\}\_\{i\}\\right\),\\qquad f\\in\\mathcal\{D\}\.\(262\)By \([257](https://arxiv.org/html/2608.13882#A4.E257)\), for everyf∈𝒟f\\in\\mathcal\{D\},
\|Φε\(f\)\|\\displaystyle\\left\|\\Phi\_\{\\varepsilon\}\\left\(f\\right\)\\right\|=\(a\)\|1n∑i=1nεif\(𝐱i\)\|≤\(b\)1n∑i=1n\|εi\|\|f\(𝐱i\)\|=\(c\)1n∑i=1n\|f\(𝐱i\)\|≤\(d\)M𝐱\(𝒟\)\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\left\|\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\varepsilon\_\{i\}f\\left\(\\mathbf\{x\}\_\{i\}\\right\)\\right\|\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\left\|\\varepsilon\_\{i\}\\right\|\\left\|f\\left\(\\mathbf\{x\}\_\{i\}\\right\)\\right\|\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{=\}\}\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\left\|f\\left\(\\mathbf\{x\}\_\{i\}\\right\)\\right\|\\stackrel\{\{\\scriptstyle\(d\)\}\}\{\{\\leq\}\}M\_\{\\mathbf\{x\}\}\\left\(\\mathcal\{D\}\\right\)\.\(263\)Here, \(a\) applies the definition ofΦε\\Phi\_\{\\varepsilon\}, \(b\) applies the triangle inequality, \(c\) uses\|εi\|=1\\left\|\\varepsilon\_\{i\}\\right\|=1for everyi∈\[n\]i\\in\\left\[n\\right\], and \(d\) applies \([257](https://arxiv.org/html/2608.13882#A4.E257)\)\. Thus,Φε\\Phi\_\{\\varepsilon\}is bounded on𝒟\\mathcal\{D\}\. We next use the symmetry of𝒟\\mathcal\{D\}\. For everyf∈𝒟f\\in\\mathcal\{D\},
\|Φε\(f\)\|\\displaystyle\\left\|\\Phi\_\{\\varepsilon\}\\left\(f\\right\)\\right\|=\(a\)max\{Φε\(f\),−Φε\(f\)\}=\(b\)max\{Φε\(f\),Φε\(−f\)\}≤\(c\)suph∈𝒟Φε\(h\)\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\max\\left\\\{\\Phi\_\{\\varepsilon\}\\left\(f\\right\),\-\\Phi\_\{\\varepsilon\}\\left\(f\\right\)\\right\\\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\max\\left\\\{\\Phi\_\{\\varepsilon\}\\left\(f\\right\),\\Phi\_\{\\varepsilon\}\\left\(\-f\\right\)\\right\\\}\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{\\leq\}\}\\sup\_\{h\\in\\mathcal\{D\}\}\\Phi\_\{\\varepsilon\}\\left\(h\\right\)\.\(264\)Here, \(a\) uses the elementary identity\|s\|=max\{s,−s\}\\left\|s\\right\|=\\max\\left\\\{s,\-s\\right\\\}, \(b\) uses the linearity ofΦε\\Phi\_\{\\varepsilon\}, and \(c\) usesf∈𝒟f\\in\\mathcal\{D\}together with−f∈𝒟\-f\\in\\mathcal\{D\}, which follows from \([255](https://arxiv.org/html/2608.13882#A4.E255)\)\. Taking the supremum overf∈𝒟f\\in\\mathcal\{D\}in \([264](https://arxiv.org/html/2608.13882#A4.E264)\) gives
supf∈𝒟\|Φε\(f\)\|≤supf∈𝒟Φε\(f\)\.\\displaystyle\\sup\_\{f\\in\\mathcal\{D\}\}\\left\|\\Phi\_\{\\varepsilon\}\\left\(f\\right\)\\right\|\\leq\\sup\_\{f\\in\\mathcal\{D\}\}\\Phi\_\{\\varepsilon\}\\left\(f\\right\)\.\(265\)Conversely,
supf∈𝒟Φε\(f\)\\displaystyle\\sup\_\{f\\in\\mathcal\{D\}\}\\Phi\_\{\\varepsilon\}\\left\(f\\right\)≤\(a\)supf∈𝒟\|Φε\(f\)\|\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}\\sup\_\{f\\in\\mathcal\{D\}\}\\left\|\\Phi\_\{\\varepsilon\}\\left\(f\\right\)\\right\|\.\(266\)Here, \(a\) usesΦε\(f\)≤\|Φε\(f\)\|\\Phi\_\{\\varepsilon\}\\left\(f\\right\)\\leq\\left\|\\Phi\_\{\\varepsilon\}\\left\(f\\right\)\\right\|for everyf∈𝒟f\\in\\mathcal\{D\}\. Combining \([265](https://arxiv.org/html/2608.13882#A4.E265)\) and \([266](https://arxiv.org/html/2608.13882#A4.E266)\) yields
supf∈𝒟\|Φε\(f\)\|=supf∈𝒟Φε\(f\)\.\\displaystyle\\sup\_\{f\\in\\mathcal\{D\}\}\\left\|\\Phi\_\{\\varepsilon\}\\left\(f\\right\)\\right\|=\\sup\_\{f\\in\\mathcal\{D\}\}\\Phi\_\{\\varepsilon\}\\left\(f\\right\)\.\(267\)We now establish the upper bound for the variation hull\. Using \([256](https://arxiv.org/html/2608.13882#A4.E256)\), we obtain
supF∈𝒲R\(𝒟\)1n∑i=1nεiF\(𝐱i\)=\(a\)supμ∈ℳ\(𝒟\)‖μ‖TV≤R1n∑i=1nεi∫𝒟f\(𝐱i\)𝑑μ\(f\)\\displaystyle\\sup\_\{F\\in\\mathcal\{W\}\_\{R\}\\left\(\\mathcal\{D\}\\right\)\}\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\varepsilon\_\{i\}F\\left\(\\mathbf\{x\}\_\{i\}\\right\)\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\sup\_\{\\begin\{subarray\}\{c\}\\mu\\in\\mathcal\{M\}\\left\(\\mathcal\{D\}\\right\)\\\\ \\left\\\|\\mu\\right\\\|\_\{\\mathrm\{TV\}\}\\leq R\\end\{subarray\}\}\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\varepsilon\_\{i\}\\int\_\{\\mathcal\{D\}\}f\\left\(\\mathbf\{x\}\_\{i\}\\right\)\\,\\mathrm\{d\}\\mu\\left\(f\\right\)=\(b\)supμ∈ℳ\(𝒟\)‖μ‖TV≤R∫𝒟\(1n∑i=1nεif\(𝐱i\)\)𝑑μ\(f\)\\displaystyle\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\sup\_\{\\begin\{subarray\}\{c\}\\mu\\in\\mathcal\{M\}\\left\(\\mathcal\{D\}\\right\)\\\\ \\left\\\|\\mu\\right\\\|\_\{\\mathrm\{TV\}\}\\leq R\\end\{subarray\}\}\\int\_\{\\mathcal\{D\}\}\\left\(\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\varepsilon\_\{i\}f\\left\(\\mathbf\{x\}\_\{i\}\\right\)\\right\)\\,\\mathrm\{d\}\\mu\\left\(f\\right\)=\(c\)supμ∈ℳ\(𝒟\)‖μ‖TV≤R∫𝒟Φε\(f\)𝑑μ\(f\)≤\(d\)supμ∈ℳ\(𝒟\)‖μ‖TV≤R∫𝒟\|Φε\(f\)\|d\|μ\|\(f\)\\displaystyle\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{=\}\}\\sup\_\{\\begin\{subarray\}\{c\}\\mu\\in\\mathcal\{M\}\\left\(\\mathcal\{D\}\\right\)\\\\ \\left\\\|\\mu\\right\\\|\_\{\\mathrm\{TV\}\}\\leq R\\end\{subarray\}\}\\int\_\{\\mathcal\{D\}\}\\Phi\_\{\\varepsilon\}\\left\(f\\right\)\\,\\mathrm\{d\}\\mu\\left\(f\\right\)\\stackrel\{\{\\scriptstyle\(d\)\}\}\{\{\\leq\}\}\\sup\_\{\\begin\{subarray\}\{c\}\\mu\\in\\mathcal\{M\}\\left\(\\mathcal\{D\}\\right\)\\\\ \\left\\\|\\mu\\right\\\|\_\{\\mathrm\{TV\}\}\\leq R\\end\{subarray\}\}\\int\_\{\\mathcal\{D\}\}\\left\|\\Phi\_\{\\varepsilon\}\\left\(f\\right\)\\right\|\\,\\mathrm\{d\}\\left\|\\mu\\right\|\\left\(f\\right\)≤\(e\)supμ∈ℳ\(𝒟\)‖μ‖TV≤R\[supf∈𝒟\|Φε\(f\)\|\]\|μ\|\(𝒟\)≤\(f\)Rsupf∈𝒟\|Φε\(f\)\|=\(g\)Rsupf∈𝒟Φε\(f\)\.\\displaystyle\\stackrel\{\{\\scriptstyle\(e\)\}\}\{\{\\leq\}\}\\sup\_\{\\begin\{subarray\}\{c\}\\mu\\in\\mathcal\{M\}\\left\(\\mathcal\{D\}\\right\)\\\\ \\left\\\|\\mu\\right\\\|\_\{\\mathrm\{TV\}\}\\leq R\\end\{subarray\}\}\\left\[\\sup\_\{f\\in\\mathcal\{D\}\}\\left\|\\Phi\_\{\\varepsilon\}\\left\(f\\right\)\\right\|\\right\]\\left\|\\mu\\right\|\\left\(\\mathcal\{D\}\\right\)\\stackrel\{\{\\scriptstyle\(f\)\}\}\{\{\\leq\}\}R\\sup\_\{f\\in\\mathcal\{D\}\}\\left\|\\Phi\_\{\\varepsilon\}\\left\(f\\right\)\\right\|\\stackrel\{\{\\scriptstyle\(g\)\}\}\{\{=\}\}R\\sup\_\{f\\in\\mathcal\{D\}\}\\Phi\_\{\\varepsilon\}\\left\(f\\right\)\.\(268\)Here, \(a\) substitutes the measure representation from \([256](https://arxiv.org/html/2608.13882#A4.E256)\), \(b\) interchanges a finite sum and the integral, \(c\) invokes \([262](https://arxiv.org/html/2608.13882#A4.E262)\), \(d\) applies the total\-variation inequality for finite signed measures, \(e\) bounds the integrand by its supremum over𝒟\\mathcal\{D\}, \(f\) uses
\|μ\|\(𝒟\)=‖μ‖TV≤R,\\displaystyle\\left\|\\mu\\right\|\\left\(\\mathcal\{D\}\\right\)=\\left\\\|\\mu\\right\\\|\_\{\\mathrm\{TV\}\}\\leq R,\(269\)and \(g\) applies \([267](https://arxiv.org/html/2608.13882#A4.E267)\)\. We next prove the reverse inequality\. SetSε:=supf∈𝒟Φε\(f\)\.S\_\{\\varepsilon\}:\\allowbreak=\\sup\_\{f\\in\\mathcal\{D\}\}\\Phi\_\{\\varepsilon\}\\left\(f\\right\)\.By \([263](https://arxiv.org/html/2608.13882#A4.E263)\), the quantitySεS\_\{\\varepsilon\}is finite\. Letη\>0\\eta\>0\. By the defining property of the supremum, there existsfη∈𝒟f\_\{\\eta\}\\in\\mathcal\{D\}such that
Φε\(fη\)≥Sε−η\.\\displaystyle\\Phi\_\{\\varepsilon\}\\left\(f\_\{\\eta\}\\right\)\\geq S\_\{\\varepsilon\}\-\\eta\.\(270\)Define the finite positive measureμη:=Rδfη,\\mu\_\{\\eta\}:\\allowbreak=R\\delta\_\{f\_\{\\eta\}\},whereδfη\\delta\_\{f\_\{\\eta\}\}denotes the Dirac measure atfηf\_\{\\eta\}\. Its total variation satisfies
‖μη‖TV\\displaystyle\\left\\\|\\mu\_\{\\eta\}\\right\\\|\_\{\\mathrm\{TV\}\}=\(a\)R‖δfη‖TV=\(b\)R\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}R\\left\\\|\\delta\_\{f\_\{\\eta\}\}\\right\\\|\_\{\\mathrm\{TV\}\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}R\.\(271\)Here, \(a\) uses the positive homogeneity of the total\-variation norm, while \(b\) uses the fact that a Dirac probability measure has total variation equal to one\. Therefore, the function
Fη\(𝐱\)\\displaystyle F\_\{\\eta\}\\left\(\\mathbf\{x\}\\right\):=∫𝒟f\(𝐱\)dμη\(f\)=\(a\)Rfη\(𝐱\)\\displaystyle:=\\int\_\{\\mathcal\{D\}\}f\\left\(\\mathbf\{x\}\\right\)\\,\\mathrm\{d\}\\mu\_\{\\eta\}\\left\(f\\right\)\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}Rf\_\{\\eta\}\\left\(\\mathbf\{x\}\\right\)\(272\)belongs to𝒲R\(𝒟\)\\mathcal\{W\}\_\{R\}\\left\(\\mathcal\{D\}\\right\)\. Here, \(a\) applies the defining property of the Dirac measure\. Consequently,
supF∈𝒲R\(𝒟\)1n∑i=1nεiF\(𝐱i\)≥\(a\)1n∑i=1nεiFη\(𝐱i\)=\(b\)RΦε\(fη\)≥\(c\)R\(Sε−η\)\.\\displaystyle\\sup\_\{F\\in\\mathcal\{W\}\_\{R\}\\left\(\\mathcal\{D\}\\right\)\}\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\varepsilon\_\{i\}F\\left\(\\mathbf\{x\}\_\{i\}\\right\)\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\geq\}\}\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\varepsilon\_\{i\}F\_\{\\eta\}\\left\(\\mathbf\{x\}\_\{i\}\\right\)\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}R\\Phi\_\{\\varepsilon\}\\left\(f\_\{\\eta\}\\right\)\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{\\geq\}\}R\\left\(S\_\{\\varepsilon\}\-\\eta\\right\)\.\(273\)Here, \(a\) evaluates the supremum at the admissible functionFηF\_\{\\eta\}, \(b\) applies \([272](https://arxiv.org/html/2608.13882#A4.E272)\) and \([262](https://arxiv.org/html/2608.13882#A4.E262)\), and \(c\) invokes \([270](https://arxiv.org/html/2608.13882#A4.E270)\)\. Since \([273](https://arxiv.org/html/2608.13882#A4.E273)\) holds for everyη\>0\\eta\>0, lettingη↓0\\eta\\downarrow 0gives
supF∈𝒲R\(𝒟\)1n∑i=1nεiF\(𝐱i\)≥RSε=Rsupf∈𝒟Φε\(f\)\.\\displaystyle\\sup\_\{F\\in\\mathcal\{W\}\_\{R\}\\left\(\\mathcal\{D\}\\right\)\}\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\varepsilon\_\{i\}F\\left\(\\mathbf\{x\}\_\{i\}\\right\)\\geq RS\_\{\\varepsilon\}=R\\sup\_\{f\\in\\mathcal\{D\}\}\\Phi\_\{\\varepsilon\}\\left\(f\\right\)\.\(274\)Combining \([268](https://arxiv.org/html/2608.13882#A4.E268)\) and \([274](https://arxiv.org/html/2608.13882#A4.E274)\), we obtain the fixed\-Rademacher identity
supF∈𝒲R\(𝒟\)1n∑i=1nεiF\(𝐱i\)=Rsupf∈𝒟1n∑i=1nεif\(𝐱i\)\.\\displaystyle\\sup\_\{F\\in\\mathcal\{W\}\_\{R\}\\left\(\\mathcal\{D\}\\right\)\}\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\varepsilon\_\{i\}F\\left\(\\mathbf\{x\}\_\{i\}\\right\)=R\\sup\_\{f\\in\\mathcal\{D\}\}\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\varepsilon\_\{i\}f\\left\(\\mathbf\{x\}\_\{i\}\\right\)\.\(275\)Taking expectation with respect toε1,…,εn\\varepsilon\_\{1\},\\ldots,\\varepsilon\_\{n\}in \([275](https://arxiv.org/html/2608.13882#A4.E275)\) yields
ℜ^n\(𝒲R\(𝒟\)\)\\displaystyle\\widehat\{\\mathfrak\{R\}\}\_\{n\}\\left\(\\mathcal\{W\}\_\{R\}\\left\(\\mathcal\{D\}\\right\)\\right\)=\(a\)Rℜ^n\(𝒟\)\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}R\\widehat\{\\mathfrak\{R\}\}\_\{n\}\\left\(\\mathcal\{D\}\\right\)\.\(276\)Here, \(a\) applies the definition of empirical Rademacher complexity to both classes\. It remains to prove the envelope bound\. For every realization of the Rademacher variables,
supf∈𝒟Φε\(f\)\\displaystyle\\sup\_\{f\\in\\mathcal\{D\}\}\\Phi\_\{\\varepsilon\}\\left\(f\\right\)≤\(a\)supf∈𝒟\|Φε\(f\)\|≤\(b\)M𝐱\(𝒟\)\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}\\sup\_\{f\\in\\mathcal\{D\}\}\\left\|\\Phi\_\{\\varepsilon\}\\left\(f\\right\)\\right\|\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}M\_\{\\mathbf\{x\}\}\\left\(\\mathcal\{D\}\\right\)\.\(277\)Here, \(a\) usess≤\|s\|s\\leq\\left\|s\\right\|, and \(b\) invokes \([263](https://arxiv.org/html/2608.13882#A4.E263)\)\. Taking expectation in \([277](https://arxiv.org/html/2608.13882#A4.E277)\) and combining the result with \([276](https://arxiv.org/html/2608.13882#A4.E276)\) gives
ℜ^n\(𝒲R\(𝒟\)\)≤RM𝐱\(𝒟\)\.\\displaystyle\\widehat\{\\mathfrak\{R\}\}\_\{n\}\\left\(\\mathcal\{W\}\_\{R\}\\left\(\\mathcal\{D\}\\right\)\\right\)\\leq RM\_\{\\mathbf\{x\}\}\\left\(\\mathcal\{D\}\\right\)\.\(278\)This proves \([258](https://arxiv.org/html/2608.13882#A4.E258)\)\. We finally specialize the result to the finite outer Brownian dictionary\. Letf∈𝒟Lm,Gf\\in\\mathcal\{D\}\_\{L\}^\{m,G\}\. By \([37](https://arxiv.org/html/2608.13882#S2.E37)\), there existsa∈𝒜L−1m,Ga\\in\\mathcal\{A\}\_\{L\-1\}^\{m,G\}such thatf∈ℋkaf\\in\\mathcal\{H\}\_\{k\_\{a\}\},‖f‖ℋka≤1\\left\\\|f\\right\\\|\_\{\\mathcal\{H\}\_\{k\_\{a\}\}\}\\leq 1\. Since an RKHS unit ball is symmetric,
‖−f‖ℋka\\displaystyle\\left\\\|\-f\\right\\\|\_\{\\mathcal\{H\}\_\{k\_\{a\}\}\}=\(a\)‖f‖ℋka≤\(b\)1\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\left\\\|f\\right\\\|\_\{\\mathcal\{H\}\_\{k\_\{a\}\}\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}1\.\(279\)Here, \(a\) uses the absolute homogeneity of the Hilbert\-space norm, and \(b\) applies‖f‖ℋka≤1\\left\\\|f\\right\\\|\_\{\\mathcal\{H\}\_\{k\_\{a\}\}\}\\leq 1\. Thus,−f\-fbelongs to the same RKHS unit ball and therefore−f∈𝒟Lm,G\-f\\in\\mathcal\{D\}\_\{L\}^\{m,G\}\. Sincef∈𝒟Lm,Gf\\in\\mathcal\{D\}\_\{L\}^\{m,G\}was arbitrary,𝒟Lm,G=−𝒟Lm,G\.\\mathcal\{D\}\_\{L\}^\{m,G\}\\allowbreak=\-\\mathcal\{D\}\_\{L\}^\{m,G\}\.For everyi∈\[n\]i\\in\\left\[n\\right\], the reproducing property in the fixed RKHSℋka\\mathcal\{H\}\_\{k\_\{a\}\}gives
\|f\(𝐱i\)\|\\displaystyle\\left\|f\\left\(\\mathbf\{x\}\_\{i\}\\right\)\\right\|=\(a\)\|⟨f,ka\(𝐱i,⋅\)⟩ℋka\|≤\(b\)‖f‖ℋka‖ka\(𝐱i,⋅\)‖ℋka=\(c\)‖f‖ℋkaka\(𝐱i,𝐱i\)1/2\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\left\|\\left\\langle f,k\_\{a\}\\left\(\\mathbf\{x\}\_\{i\},\\cdot\\right\)\\right\\rangle\_\{\\mathcal\{H\}\_\{k\_\{a\}\}\}\\right\|\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}\\left\\\|f\\right\\\|\_\{\\mathcal\{H\}\_\{k\_\{a\}\}\}\\left\\\|k\_\{a\}\\left\(\\mathbf\{x\}\_\{i\},\\cdot\\right\)\\right\\\|\_\{\\mathcal\{H\}\_\{k\_\{a\}\}\}\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{=\}\}\\left\\\|f\\right\\\|\_\{\\mathcal\{H\}\_\{k\_\{a\}\}\}k\_\{a\}\\left\(\\mathbf\{x\}\_\{i\},\\mathbf\{x\}\_\{i\}\\right\)^\{1/2\}≤\(d\)ka\(𝐱i,𝐱i\)1/2=\(e\)\|a\(𝐱i\)\|1/2≤\(f\)B𝐱1/2\.\\displaystyle\\stackrel\{\{\\scriptstyle\(d\)\}\}\{\{\\leq\}\}k\_\{a\}\\left\(\\mathbf\{x\}\_\{i\},\\mathbf\{x\}\_\{i\}\\right\)^\{1/2\}\\stackrel\{\{\\scriptstyle\(e\)\}\}\{\{=\}\}\\left\|a\\left\(\\mathbf\{x\}\_\{i\}\\right\)\\right\|^\{1/2\}\\stackrel\{\{\\scriptstyle\(f\)\}\}\{\{\\leq\}\}B\_\{\\mathbf\{x\}\}^\{1/2\}\.\(280\)Here, \(a\) applies the reproducing property, \(b\) applies the Cauchy–Schwarz inequality inℋka\\mathcal\{H\}\_\{k\_\{a\}\}, \(c\) applies the RKHS section\-norm identity, \(d\) uses‖f‖ℋka≤1\\left\\\|f\\right\\\|\_\{\\mathcal\{H\}\_\{k\_\{a\}\}\}\\leq 1, \(e\) uses
ka\(𝐱i,𝐱i\)\\displaystyle k\_\{a\}\\left\(\\mathbf\{x\}\_\{i\},\\mathbf\{x\}\_\{i\}\\right\)=\(a\)k\(B\)\(a\(𝐱i\),a\(𝐱i\)\)=\(b\)\|a\(𝐱i\)\|,\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}k^\{\(\\mathrm\{B\}\)\}\\left\(a\\left\(\\mathbf\{x\}\_\{i\}\\right\),a\\left\(\\mathbf\{x\}\_\{i\}\\right\)\\right\)\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\left\|a\\left\(\\mathbf\{x\}\_\{i\}\\right\)\\right\|,\(281\)where \(a\) is the definition of the Brownian pullback kernel and \(b\) usesk\(B\)\(s,s\)=\|s\|k^\{\(\\mathrm\{B\}\)\}\\left\(s,s\\right\)=\\left\|s\\right\|, and \(f\) invokes \([259](https://arxiv.org/html/2608.13882#A4.E259)\)\. Taking the maximum overi∈\[n\]i\\in\\left\[n\\right\]and then the supremum overf∈𝒟Lm,Gf\\in\\mathcal\{D\}\_\{L\}^\{m,G\}in \([280](https://arxiv.org/html/2608.13882#A4.E280)\) givesM𝐱\(𝒟Lm,G\)≤B𝐱1/2M\_\{\\mathbf\{x\}\}\\left\(\\mathcal\{D\}\_\{L\}^\{m,G\}\\right\)\\leq B\_\{\\mathbf\{x\}\}^\{1/2\}, which proves \([260](https://arxiv.org/html/2608.13882#A4.E260)\)\. Finally, \([38](https://arxiv.org/html/2608.13882#S2.E38)\) and \([256](https://arxiv.org/html/2608.13882#A4.E256)\) give𝒲RL;m,G=𝒲R\(𝒟Lm,G\)\\mathcal\{W\}\_\{R\}^\{L;m,G\}=\\mathcal\{W\}\_\{R\}\\left\(\\mathcal\{D\}\_\{L\}^\{m,G\}\\right\)\. Applying \([258](https://arxiv.org/html/2608.13882#A4.E258)\) with𝒟=𝒟Lm,G\\mathcal\{D\}=\\mathcal\{D\}\_\{L\}^\{m,G\}and using \([260](https://arxiv.org/html/2608.13882#A4.E260)\) proves \([261](https://arxiv.org/html/2608.13882#A4.E261)\)\. This completes the proof\.
∎
###### Lemma B0\(Rademacher bound for unions of Brownian pullback RKHS balls\)\.
Letn∈ℕn\\in\\mathbb\{N\},n≥1n\\geq 1, and let𝐱1,…,𝐱n∈𝒳\\mathbf\{x\}\_\{1\},\\ldots,\\mathbf\{x\}\_\{n\}\\in\\mathcal\{X\}be fixed sample points\. Let𝒜\\mathcal\{A\}be a nonempty class of real\-valued support functions on𝒳\\mathcal\{X\}\. For everya∈𝒜a\\in\\mathcal\{A\}, define
ka\(𝐱,𝐱′\):=k\(B\)\(a\(𝐱\),a\(𝐱′\)\),𝐱,𝐱′∈𝒳\.\\displaystyle k\_\{a\}\\left\(\\mathbf\{x\},\\mathbf\{x\}^\{\\prime\}\\right\):=k^\{\(\\mathrm\{B\}\)\}\\left\(a\\left\(\\mathbf\{x\}\\right\),a\\left\(\\mathbf\{x\}^\{\\prime\}\\right\)\\right\),\\qquad\\mathbf\{x\},\\mathbf\{x\}^\{\\prime\}\\in\\mathcal\{X\}\.\(282\)Assume that
B𝐱\(𝒜\):=supa∈𝒜maxi∈\[n\]ka\(𝐱i,𝐱i\)<∞\.\\displaystyle B\_\{\\mathbf\{x\}\}\\left\(\\mathcal\{A\}\\right\):=\\sup\_\{a\\in\\mathcal\{A\}\}\\max\_\{i\\in\\left\[n\\right\]\}k\_\{a\}\\left\(\\mathbf\{x\}\_\{i\},\\mathbf\{x\}\_\{i\}\\right\)<\\infty\.\(283\)Define
ℭ^n,𝐱\(B\)\(𝒜\):=𝔼ε\[supa∈𝒜\|1n∑1≤i<j≤nεiεjka\(𝐱i,𝐱j\)\|\],\\displaystyle\\widehat\{\\mathfrak\{C\}\}\_\{n,\\mathbf\{x\}\}^\{\(B\)\}\\left\(\\mathcal\{A\}\\right\):=\\mathbb\{E\}\_\{\\varepsilon\}\\left\[\\sup\_\{a\\in\\mathcal\{A\}\}\\left\|\\frac\{1\}\{n\}\\sum\_\{1\\leq i<j\\leq n\}\\varepsilon\_\{i\}\\varepsilon\_\{j\}k\_\{a\}\\left\(\\mathbf\{x\}\_\{i\},\\mathbf\{x\}\_\{j\}\\right\)\\right\|\\right\],\(284\)whereε1,…,εn\\varepsilon\_\{1\},\\ldots,\\varepsilon\_\{n\}are independent Rademacher random variables\. Forr\>0r\>0, let
𝒟r\(𝒜\):=⋃a∈𝒜\{f∈ℋka:‖f‖ℋka≤r\}\.\\displaystyle\\mathcal\{D\}\_\{r\}\\left\(\\mathcal\{A\}\\right\):=\\bigcup\_\{a\\in\\mathcal\{A\}\}\\left\\\{f\\in\\mathcal\{H\}\_\{k\_\{a\}\}:\\left\\\|f\\right\\\|\_\{\\mathcal\{H\}\_\{k\_\{a\}\}\}\\leq r\\right\\\}\.\(285\)Then𝒟r\(𝒜\)\\mathcal\{D\}\_\{r\}\\left\(\\mathcal\{A\}\\right\)is symmetric and
ℜ^n\(𝒟r\(𝒜\)\)\\displaystyle\\widehat\{\\mathfrak\{R\}\}\_\{n\}\\left\(\\mathcal\{D\}\_\{r\}\\left\(\\mathcal\{A\}\\right\)\\right\)=rn𝔼ε\[supa∈𝒜\(∑i,j=1nεiεjka\(𝐱i,𝐱j\)\)1/2\]\\displaystyle=\\frac\{r\}\{n\}\\mathbb\{E\}\_\{\\varepsilon\}\\left\[\\sup\_\{a\\in\\mathcal\{A\}\}\\left\(\\sum\_\{i,j=1\}^\{n\}\\varepsilon\_\{i\}\\varepsilon\_\{j\}k\_\{a\}\\left\(\\mathbf\{x\}\_\{i\},\\mathbf\{x\}\_\{j\}\\right\)\\right\)^\{1/2\}\\right\]\(286\)≤rn\(2ℭ^n,𝐱\(B\)\(𝒜\)\+B𝐱\(𝒜\)\)1/2\.\\displaystyle\\leq\\frac\{r\}\{\\sqrt\{n\}\}\\left\(2\\widehat\{\\mathfrak\{C\}\}\_\{n,\\mathbf\{x\}\}^\{\(B\)\}\\left\(\\mathcal\{A\}\\right\)\+B\_\{\\mathbf\{x\}\}\\left\(\\mathcal\{A\}\\right\)\\right\)^\{1/2\}\.\(287\)
In particular, fixL≥2L\\geq 2,m≥1m\\geq 1, andG≥2G\\geq 2, and set
B𝐱:=supa∈𝒜L−1m,Gmaxi∈\[n\]\|a\(𝐱i\)\|\.\\displaystyle B\_\{\\mathbf\{x\}\}:=\\sup\_\{a\\in\\mathcal\{A\}\_\{L\-1\}^\{m,G\}\}\\max\_\{i\\in\\left\[n\\right\]\}\\left\|a\\left\(\\mathbf\{x\}\_\{i\}\\right\)\\right\|\.\(288\)Then
𝒟1\(𝒜L−1m,G\)=𝒟Lm,G,\\displaystyle\\mathcal\{D\}\_\{1\}\\left\(\\mathcal\{A\}\_\{L\-1\}^\{m,G\}\\right\)=\\mathcal\{D\}\_\{L\}^\{m,G\},\(289\)and
B𝐱\(𝒜L−1m,G\)=B𝐱\.\\displaystyle B\_\{\\mathbf\{x\}\}\\left\(\\mathcal\{A\}\_\{L\-1\}^\{m,G\}\\right\)=B\_\{\\mathbf\{x\}\}\.\(290\)Consequently,
ℜ^n\(𝒟Lm,G\)≤1n\(2ℭ^n,𝐱\(B\)\(𝒜L−1m,G\)\+B𝐱\)1/2\.\\displaystyle\\widehat\{\\mathfrak\{R\}\}\_\{n\}\\left\(\\mathcal\{D\}\_\{L\}^\{m,G\}\\right\)\\leq\\frac\{1\}\{\\sqrt\{n\}\}\\left\(2\\widehat\{\\mathfrak\{C\}\}\_\{n,\\mathbf\{x\}\}^\{\(B\)\}\\left\(\\mathcal\{A\}\_\{L\-1\}^\{m,G\}\\right\)\+B\_\{\\mathbf\{x\}\}\\right\)^\{1/2\}\.\(291\)
###### Proof\.
We divide the proof into four steps\.
##### Step 1: symmetry of the union class\.
Fixf∈𝒟r\(𝒜\)f\\in\\mathcal\{D\}\_\{r\}\\left\(\\mathcal\{A\}\\right\)\. By \([285](https://arxiv.org/html/2608.13882#A4.E285)\), there existsa∈𝒜a\\in\\mathcal\{A\}such that
f\\displaystyle f∈ℋka,\\displaystyle\\in\\mathcal\{H\}\_\{k\_\{a\}\},‖f‖ℋka\\displaystyle\\left\\\|f\\right\\\|\_\{\\mathcal\{H\}\_\{k\_\{a\}\}\}≤r\.\\displaystyle\\leq r\.\(292\)Sinceℋka\\mathcal\{H\}\_\{k\_\{a\}\}is a vector space,−f∈ℋka\-f\\in\\mathcal\{H\}\_\{k\_\{a\}\}\. Moreover,
‖−f‖ℋka\\displaystyle\\left\\\|\-f\\right\\\|\_\{\\mathcal\{H\}\_\{k\_\{a\}\}\}=\(a\)\|−1\|‖f‖ℋka=\(b\)‖f‖ℋka≤\(c\)r\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\left\|\-1\\right\|\\left\\\|f\\right\\\|\_\{\\mathcal\{H\}\_\{k\_\{a\}\}\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\left\\\|f\\right\\\|\_\{\\mathcal\{H\}\_\{k\_\{a\}\}\}\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{\\leq\}\}r\.\(293\)Here, \(a\) applies the absolute homogeneity of the Hilbert\-space norm, \(b\) uses\|−1\|=1\\left\|\-1\\right\|=1, and \(c\) invokes \([292](https://arxiv.org/html/2608.13882#A4.E292)\)\. Therefore,−f∈𝒟r\(𝒜\)\-f\\in\\mathcal\{D\}\_\{r\}\\left\(\\mathcal\{A\}\\right\)\. Sincef∈𝒟r\(𝒜\)f\\in\\mathcal\{D\}\_\{r\}\\left\(\\mathcal\{A\}\\right\)was arbitrary,
𝒟r\(𝒜\)=−𝒟r\(𝒜\)\.\\displaystyle\\mathcal\{D\}\_\{r\}\\left\(\\mathcal\{A\}\\right\)=\-\\mathcal\{D\}\_\{r\}\\left\(\\mathcal\{A\}\\right\)\.\(294\)Fix a realizationε1,…,εn\\varepsilon\_\{1\},\\ldots,\\varepsilon\_\{n\}of the Rademacher variables and define
Λε\(f\):=1n∑i=1nεif\(𝐱i\),f∈𝒟r\(𝒜\)\.\\displaystyle\\Lambda\_\{\\varepsilon\}\\left\(f\\right\):=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\varepsilon\_\{i\}f\\left\(\\mathbf\{x\}\_\{i\}\\right\),\\qquad f\\in\\mathcal\{D\}\_\{r\}\\left\(\\mathcal\{A\}\\right\)\.\(295\)By \([294](https://arxiv.org/html/2608.13882#A4.E294)\),
supf∈𝒟r\(𝒜\)\|Λε\(f\)\|\\displaystyle\\sup\_\{f\\in\\mathcal\{D\}\_\{r\}\\left\(\\mathcal\{A\}\\right\)\}\\left\|\\Lambda\_\{\\varepsilon\}\\left\(f\\right\)\\right\|=\(a\)supf∈𝒟r\(𝒜\)max\{Λε\(f\),−Λε\(f\)\}=\(b\)supf∈𝒟r\(𝒜\)max\{Λε\(f\),Λε\(−f\)\}\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\sup\_\{f\\in\\mathcal\{D\}\_\{r\}\\left\(\\mathcal\{A\}\\right\)\}\\max\\left\\\{\\Lambda\_\{\\varepsilon\}\\left\(f\\right\),\-\\Lambda\_\{\\varepsilon\}\\left\(f\\right\)\\right\\\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\sup\_\{f\\in\\mathcal\{D\}\_\{r\}\\left\(\\mathcal\{A\}\\right\)\}\\max\\left\\\{\\Lambda\_\{\\varepsilon\}\\left\(f\\right\),\\Lambda\_\{\\varepsilon\}\\left\(\-f\\right\)\\right\\\}=\(c\)supf∈𝒟r\(𝒜\)Λε\(f\)\.\\displaystyle\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{=\}\}\\sup\_\{f\\in\\mathcal\{D\}\_\{r\}\\left\(\\mathcal\{A\}\\right\)\}\\Lambda\_\{\\varepsilon\}\\left\(f\\right\)\.\(296\)Here, \(a\) uses\|s\|=max\{s,−s\}\\left\|s\\right\|=\\max\\left\\\{s,\-s\\right\\\}, \(b\) uses the linearity ofΛε\\Lambda\_\{\\varepsilon\}, and \(c\) uses the symmetry \([294](https://arxiv.org/html/2608.13882#A4.E294)\)\. Thus, the absolute\-value form may be used without changing the empirical Rademacher complexity\.
##### Step 2: exact RKHS reduction for a fixed support\.
Fixa∈𝒜a\\in\\mathcal\{A\}and define
Va,ε:=∑i=1nεika\(𝐱i,⋅\)∈ℋka\.\\displaystyle V\_\{a,\\varepsilon\}:=\\sum\_\{i=1\}^\{n\}\\varepsilon\_\{i\}k\_\{a\}\\left\(\\mathbf\{x\}\_\{i\},\\cdot\\right\)\\in\\mathcal\{H\}\_\{k\_\{a\}\}\.\(297\)For everyf∈ℋkaf\\in\\mathcal\{H\}\_\{k\_\{a\}\}, the reproducing property gives
∑i=1nεif\(𝐱i\)\\displaystyle\\sum\_\{i=1\}^\{n\}\\varepsilon\_\{i\}f\\left\(\\mathbf\{x\}\_\{i\}\\right\)=\(a\)∑i=1nεi⟨f,ka\(𝐱i,⋅\)⟩ℋka=\(b\)⟨f,∑i=1nεika\(𝐱i,⋅\)⟩ℋka=\(c\)⟨f,Va,ε⟩ℋka\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\sum\_\{i=1\}^\{n\}\\varepsilon\_\{i\}\\left\\langle f,k\_\{a\}\\left\(\\mathbf\{x\}\_\{i\},\\cdot\\right\)\\right\\rangle\_\{\\mathcal\{H\}\_\{k\_\{a\}\}\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\left\\langle f,\\sum\_\{i=1\}^\{n\}\\varepsilon\_\{i\}k\_\{a\}\\left\(\\mathbf\{x\}\_\{i\},\\cdot\\right\)\\right\\rangle\_\{\\mathcal\{H\}\_\{k\_\{a\}\}\}\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{=\}\}\\left\\langle f,V\_\{a,\\varepsilon\}\\right\\rangle\_\{\\mathcal\{H\}\_\{k\_\{a\}\}\}\.\(298\)Here, \(a\) applies the reproducing property in the fixed RKHSℋka\\mathcal\{H\}\_\{k\_\{a\}\}, \(b\) uses the linearity of the inner product in its second argument, and \(c\) invokes \([297](https://arxiv.org/html/2608.13882#A4.E297)\)\. Consequently,
supf∈ℋka‖f‖ℋka≤r\|1n∑i=1nεif\(𝐱i\)\|=\(a\)1nsupf∈ℋka‖f‖ℋka≤r\|⟨f,Va,ε⟩ℋka\|\\displaystyle\\sup\_\{\\begin\{subarray\}\{c\}f\\in\\mathcal\{H\}\_\{k\_\{a\}\}\\\\ \\left\\\|f\\right\\\|\_\{\\mathcal\{H\}\_\{k\_\{a\}\}\}\\leq r\\end\{subarray\}\}\\left\|\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\varepsilon\_\{i\}f\\left\(\\mathbf\{x\}\_\{i\}\\right\)\\right\|\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\frac\{1\}\{n\}\\sup\_\{\\begin\{subarray\}\{c\}f\\in\\mathcal\{H\}\_\{k\_\{a\}\}\\\\ \\left\\\|f\\right\\\|\_\{\\mathcal\{H\}\_\{k\_\{a\}\}\}\\leq r\\end\{subarray\}\}\\left\|\\left\\langle f,V\_\{a,\\varepsilon\}\\right\\rangle\_\{\\mathcal\{H\}\_\{k\_\{a\}\}\}\\right\|=\(b\)rn‖Va,ε‖ℋka\.\\displaystyle\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\frac\{r\}\{n\}\\left\\\|V\_\{a,\\varepsilon\}\\right\\\|\_\{\\mathcal\{H\}\_\{k\_\{a\}\}\}\.\(299\)Here, \(a\) applies \([298](https://arxiv.org/html/2608.13882#A4.E298)\), and \(b\) applies the Hilbert\-space dual\-norm identity\. More explicitly, the Cauchy–Schwarz inequality gives the upper bound
\|⟨f,Va,ε⟩ℋka\|≤r‖Va,ε‖ℋka,\\displaystyle\\left\|\\left\\langle f,V\_\{a,\\varepsilon\}\\right\\rangle\_\{\\mathcal\{H\}\_\{k\_\{a\}\}\}\\right\|\\leq r\\left\\\|V\_\{a,\\varepsilon\}\\right\\\|\_\{\\mathcal\{H\}\_\{k\_\{a\}\}\},\(300\)and equality is attained, wheneverVa,ε≠0V\_\{a,\\varepsilon\}\\neq 0, by choosingf=rVa,ε‖Va,ε‖ℋkaf=r\\frac\{V\_\{a,\\varepsilon\}\}\{\\left\\\|V\_\{a,\\varepsilon\}\\right\\\|\_\{\\mathcal\{H\}\_\{k\_\{a\}\}\}\}\. WhenVa,ε=0V\_\{a,\\varepsilon\}=0, both sides of \([299](https://arxiv.org/html/2608.13882#A4.E299)\) are zero\. The squared norm in \([299](https://arxiv.org/html/2608.13882#A4.E299)\) satisfies
‖Va,ε‖ℋka2\\displaystyle\\left\\\|V\_\{a,\\varepsilon\}\\right\\\|\_\{\\mathcal\{H\}\_\{k\_\{a\}\}\}^\{2\}=\(a\)⟨∑i=1nεika\(𝐱i,⋅\),∑j=1nεjka\(𝐱j,⋅\)⟩ℋka=\(b\)∑i,j=1nεiεj⟨ka\(𝐱i,⋅\),ka\(𝐱j,⋅\)⟩ℋka\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\left\\langle\\sum\_\{i=1\}^\{n\}\\varepsilon\_\{i\}k\_\{a\}\\left\(\\mathbf\{x\}\_\{i\},\\cdot\\right\),\\sum\_\{j=1\}^\{n\}\\varepsilon\_\{j\}k\_\{a\}\\left\(\\mathbf\{x\}\_\{j\},\\cdot\\right\)\\right\\rangle\_\{\\mathcal\{H\}\_\{k\_\{a\}\}\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\sum\_\{i,j=1\}^\{n\}\\varepsilon\_\{i\}\\varepsilon\_\{j\}\\left\\langle k\_\{a\}\\left\(\\mathbf\{x\}\_\{i\},\\cdot\\right\),k\_\{a\}\\left\(\\mathbf\{x\}\_\{j\},\\cdot\\right\)\\right\\rangle\_\{\\mathcal\{H\}\_\{k\_\{a\}\}\}=\(c\)∑i,j=1nεiεjka\(𝐱i,𝐱j\)\.\\displaystyle\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{=\}\}\\sum\_\{i,j=1\}^\{n\}\\varepsilon\_\{i\}\\varepsilon\_\{j\}k\_\{a\}\\left\(\\mathbf\{x\}\_\{i\},\\mathbf\{x\}\_\{j\}\\right\)\.\(301\)Here, \(a\) substitutes \([297](https://arxiv.org/html/2608.13882#A4.E297)\), \(b\) expands the inner product by bilinearity, and \(c\) applies the RKHS kernel\-section identity\. In particular,
∑i,j=1nεiεjka\(𝐱i,𝐱j\)≥0,\\displaystyle\\sum\_\{i,j=1\}^\{n\}\\varepsilon\_\{i\}\\varepsilon\_\{j\}k\_\{a\}\\left\(\\mathbf\{x\}\_\{i\},\\mathbf\{x\}\_\{j\}\\right\)\\geq 0,\(302\)because it is the squared Hilbert\-space norm appearing on the left\-hand side of \([301](https://arxiv.org/html/2608.13882#A4.E301)\)\. Combining \([299](https://arxiv.org/html/2608.13882#A4.E299)\) and \([301](https://arxiv.org/html/2608.13882#A4.E301)\) gives
supf∈ℋka‖f‖ℋka≤r\|1n∑i=1nεif\(𝐱i\)\|=rn\(∑i,j=1nεiεjka\(𝐱i,𝐱j\)\)1/2\.\\displaystyle\\sup\_\{\\begin\{subarray\}\{c\}f\\in\\mathcal\{H\}\_\{k\_\{a\}\}\\\\ \\left\\\|f\\right\\\|\_\{\\mathcal\{H\}\_\{k\_\{a\}\}\}\\leq r\\end\{subarray\}\}\\left\|\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\varepsilon\_\{i\}f\\left\(\\mathbf\{x\}\_\{i\}\\right\)\\right\|=\\frac\{r\}\{n\}\\left\(\\sum\_\{i,j=1\}^\{n\}\\varepsilon\_\{i\}\\varepsilon\_\{j\}k\_\{a\}\\left\(\\mathbf\{x\}\_\{i\},\\mathbf\{x\}\_\{j\}\\right\)\\right\)^\{1/2\}\.\(303\)
##### Step 3: reduction of the union to Brownian quadratic chaos\.
Using \([285](https://arxiv.org/html/2608.13882#A4.E285)\), \([296](https://arxiv.org/html/2608.13882#A4.E296)\), and \([303](https://arxiv.org/html/2608.13882#A4.E303)\), we obtain
supf∈𝒟r\(𝒜\)1n∑i=1nεif\(𝐱i\)=\(a\)supf∈𝒟r\(𝒜\)\|1n∑i=1nεif\(𝐱i\)\|\\displaystyle\\sup\_\{f\\in\\mathcal\{D\}\_\{r\}\\left\(\\mathcal\{A\}\\right\)\}\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\varepsilon\_\{i\}f\\left\(\\mathbf\{x\}\_\{i\}\\right\)\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\sup\_\{f\\in\\mathcal\{D\}\_\{r\}\\left\(\\mathcal\{A\}\\right\)\}\\left\|\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\varepsilon\_\{i\}f\\left\(\\mathbf\{x\}\_\{i\}\\right\)\\right\|=\(b\)supa∈𝒜supf∈ℋka‖f‖ℋka≤r\|1n∑i=1nεif\(𝐱i\)\|=\(c\)rnsupa∈𝒜\(∑i,j=1nεiεjka\(𝐱i,𝐱j\)\)1/2\.\\displaystyle\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\sup\_\{a\\in\\mathcal\{A\}\}\\sup\_\{\\begin\{subarray\}\{c\}f\\in\\mathcal\{H\}\_\{k\_\{a\}\}\\\\ \\left\\\|f\\right\\\|\_\{\\mathcal\{H\}\_\{k\_\{a\}\}\}\\leq r\\end\{subarray\}\}\\left\|\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\varepsilon\_\{i\}f\\left\(\\mathbf\{x\}\_\{i\}\\right\)\\right\|\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{=\}\}\\frac\{r\}\{n\}\\sup\_\{a\\in\\mathcal\{A\}\}\\left\(\\sum\_\{i,j=1\}^\{n\}\\varepsilon\_\{i\}\\varepsilon\_\{j\}k\_\{a\}\\left\(\\mathbf\{x\}\_\{i\},\\mathbf\{x\}\_\{j\}\\right\)\\right\)^\{1/2\}\.\(304\)Here, \(a\) applies \([296](https://arxiv.org/html/2608.13882#A4.E296)\), \(b\) applies the union definition \([285](https://arxiv.org/html/2608.13882#A4.E285)\), and \(c\) applies \([303](https://arxiv.org/html/2608.13882#A4.E303)\) for each fixeda∈𝒜a\\in\\mathcal\{A\}\. Taking expectation with respect to the Rademacher variables in \([304](https://arxiv.org/html/2608.13882#A4.E304)\) proves \([286](https://arxiv.org/html/2608.13882#A4.E286)\)\. For notational convenience, define
Qa\(ε\):=∑i,j=1nεiεjka\(𝐱i,𝐱j\)\.\\displaystyle Q\_\{a\}\\left\(\\varepsilon\\right\):=\\sum\_\{i,j=1\}^\{n\}\\varepsilon\_\{i\}\\varepsilon\_\{j\}k\_\{a\}\\left\(\\mathbf\{x\}\_\{i\},\\mathbf\{x\}\_\{j\}\\right\)\.\(305\)By \([302](https://arxiv.org/html/2608.13882#A4.E302)\), one hasQa\(ε\)≥0Q\_\{a\}\\left\(\\varepsilon\\right\)\\geq 0,a∈𝒜a\\in\\mathcal\{A\}\. Since the square\-root function is increasing on\[0,∞\)\\left\[0,\\infty\\right\),
supa∈𝒜Qa\(ε\)1/2\\displaystyle\\sup\_\{a\\in\\mathcal\{A\}\}Q\_\{a\}\\left\(\\varepsilon\\right\)^\{1/2\}=\(a\)\(supa∈𝒜Qa\(ε\)\)1/2\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\left\(\\sup\_\{a\\in\\mathcal\{A\}\}Q\_\{a\}\\left\(\\varepsilon\\right\)\\right\)^\{1/2\}\.\(306\)Here, \(a\) uses the monotonicity and continuity of the square\-root function\. Applying Jensen’s inequality to the concave functiont↦t1/2t\\mapsto t^\{1/2\}gives
𝔼ε\[supa∈𝒜Qa\(ε\)1/2\]=\(a\)𝔼ε\[\(supa∈𝒜Qa\(ε\)\)1/2\]≤\(b\)\(𝔼ε\[supa∈𝒜Qa\(ε\)\]\)1/2\.\\displaystyle\\mathbb\{E\}\_\{\\varepsilon\}\\left\[\\sup\_\{a\\in\\mathcal\{A\}\}Q\_\{a\}\\left\(\\varepsilon\\right\)^\{1/2\}\\right\]\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\mathbb\{E\}\_\{\\varepsilon\}\\left\[\\left\(\\sup\_\{a\\in\\mathcal\{A\}\}Q\_\{a\}\\left\(\\varepsilon\\right\)\\right\)^\{1/2\}\\right\]\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}\\left\(\\mathbb\{E\}\_\{\\varepsilon\}\\left\[\\sup\_\{a\\in\\mathcal\{A\}\}Q\_\{a\}\\left\(\\varepsilon\\right\)\\right\]\\right\)^\{1/2\}\.\(307\)Here, \(a\) applies \([306](https://arxiv.org/html/2608.13882#A4.E306)\), and \(b\) applies Jensen’s inequality\.
##### Step 4: diagonal and off\-diagonal decomposition\.
For everya∈𝒜a\\in\\mathcal\{A\},
Qa\(ε\)\\displaystyle Q\_\{a\}\\left\(\\varepsilon\\right\)=\(a\)∑i=1nεi2ka\(𝐱i,𝐱i\)\+2∑1≤i<j≤nεiεjka\(𝐱i,𝐱j\)\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\sum\_\{i=1\}^\{n\}\\varepsilon\_\{i\}^\{2\}k\_\{a\}\\left\(\\mathbf\{x\}\_\{i\},\\mathbf\{x\}\_\{i\}\\right\)\+2\\sum\_\{1\\leq i<j\\leq n\}\\varepsilon\_\{i\}\\varepsilon\_\{j\}k\_\{a\}\\left\(\\mathbf\{x\}\_\{i\},\\mathbf\{x\}\_\{j\}\\right\)=\(b\)∑i=1nka\(𝐱i,𝐱i\)\+2∑1≤i<j≤nεiεjka\(𝐱i,𝐱j\)\.\\displaystyle\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\sum\_\{i=1\}^\{n\}k\_\{a\}\\left\(\\mathbf\{x\}\_\{i\},\\mathbf\{x\}\_\{i\}\\right\)\+2\\sum\_\{1\\leq i<j\\leq n\}\\varepsilon\_\{i\}\\varepsilon\_\{j\}k\_\{a\}\\left\(\\mathbf\{x\}\_\{i\},\\mathbf\{x\}\_\{j\}\\right\)\.\(308\)Here, \(a\) separates the diagonal and off\-diagonal terms in the double sum, and \(b\) usesεi2=1\\varepsilon\_\{i\}^\{2\}=1for everyi∈\[n\]i\\in\\left\[n\\right\]\. Consequently,
supa∈𝒜Qa\(ε\)\\displaystyle\\sup\_\{a\\in\\mathcal\{A\}\}Q\_\{a\}\\left\(\\varepsilon\\right\)≤\(a\)2supa∈𝒜\|∑1≤i<j≤nεiεjka\(𝐱i,𝐱j\)\|\+supa∈𝒜∑i=1nka\(𝐱i,𝐱i\)\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}2\\sup\_\{a\\in\\mathcal\{A\}\}\\left\|\\sum\_\{1\\leq i<j\\leq n\}\\varepsilon\_\{i\}\\varepsilon\_\{j\}k\_\{a\}\\left\(\\mathbf\{x\}\_\{i\},\\mathbf\{x\}\_\{j\}\\right\)\\right\|\+\\sup\_\{a\\in\\mathcal\{A\}\}\\sum\_\{i=1\}^\{n\}k\_\{a\}\\left\(\\mathbf\{x\}\_\{i\},\\mathbf\{x\}\_\{i\}\\right\)≤\(b\)2supa∈𝒜\|∑1≤i<j≤nεiεjka\(𝐱i,𝐱j\)\|\+nB𝐱\(𝒜\)\.\\displaystyle\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}2\\sup\_\{a\\in\\mathcal\{A\}\}\\left\|\\sum\_\{1\\leq i<j\\leq n\}\\varepsilon\_\{i\}\\varepsilon\_\{j\}k\_\{a\}\\left\(\\mathbf\{x\}\_\{i\},\\mathbf\{x\}\_\{j\}\\right\)\\right\|\+nB\_\{\\mathbf\{x\}\}\\left\(\\mathcal\{A\}\\right\)\.\(309\)Here, \(a\) applies \([308](https://arxiv.org/html/2608.13882#A4.E308)\) and bounds the off\-diagonal term by its absolute value, while \(b\) applies \([283](https://arxiv.org/html/2608.13882#A4.E283)\) to every diagonal kernel value\. Taking expectation in \([309](https://arxiv.org/html/2608.13882#A4.E309)\) gives
𝔼ε\[supa∈𝒜Qa\(ε\)\]≤\(a\)2𝔼ε\[supa∈𝒜\|∑1≤i<j≤nεiεjka\(𝐱i,𝐱j\)\|\]\+nB𝐱\(𝒜\)\\displaystyle\\mathbb\{E\}\_\{\\varepsilon\}\\left\[\\sup\_\{a\\in\\mathcal\{A\}\}Q\_\{a\}\\left\(\\varepsilon\\right\)\\right\]\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}2\\mathbb\{E\}\_\{\\varepsilon\}\\left\[\\sup\_\{a\\in\\mathcal\{A\}\}\\left\|\\sum\_\{1\\leq i<j\\leq n\}\\varepsilon\_\{i\}\\varepsilon\_\{j\}k\_\{a\}\\left\(\\mathbf\{x\}\_\{i\},\\mathbf\{x\}\_\{j\}\\right\)\\right\|\\right\]\+nB\_\{\\mathbf\{x\}\}\\left\(\\mathcal\{A\}\\right\)=\(b\)2nℭ^n,𝐱\(B\)\(𝒜\)\+nB𝐱\(𝒜\)=\(c\)n\(2ℭ^n,𝐱\(B\)\(𝒜\)\+B𝐱\(𝒜\)\)\.\\displaystyle\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}2n\\widehat\{\\mathfrak\{C\}\}\_\{n,\\mathbf\{x\}\}^\{\(B\)\}\\left\(\\mathcal\{A\}\\right\)\+nB\_\{\\mathbf\{x\}\}\\left\(\\mathcal\{A\}\\right\)\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{=\}\}n\\left\(2\\widehat\{\\mathfrak\{C\}\}\_\{n,\\mathbf\{x\}\}^\{\(B\)\}\\left\(\\mathcal\{A\}\\right\)\+B\_\{\\mathbf\{x\}\}\\left\(\\mathcal\{A\}\\right\)\\right\)\.\(310\)Here, \(a\) takes expectations in \([309](https://arxiv.org/html/2608.13882#A4.E309)\), \(b\) applies the normalization in \([284](https://arxiv.org/html/2608.13882#A4.E284)\), and \(c\) factors outnn\. Combining \([286](https://arxiv.org/html/2608.13882#A4.E286)\), \([307](https://arxiv.org/html/2608.13882#A4.E307)\), and \([310](https://arxiv.org/html/2608.13882#A4.E310)\), we obtain
ℜ^n\(𝒟r\(𝒜\)\)\\displaystyle\\widehat\{\\mathfrak\{R\}\}\_\{n\}\\left\(\\mathcal\{D\}\_\{r\}\\left\(\\mathcal\{A\}\\right\)\\right\)≤rn\[n\(2ℭ^n,𝐱\(B\)\(𝒜\)\+B𝐱\(𝒜\)\)\]1/2\\displaystyle\\leq\\frac\{r\}\{n\}\\left\[n\\left\(2\\widehat\{\\mathfrak\{C\}\}\_\{n,\\mathbf\{x\}\}^\{\(B\)\}\\left\(\\mathcal\{A\}\\right\)\+B\_\{\\mathbf\{x\}\}\\left\(\\mathcal\{A\}\\right\)\\right\)\\right\]^\{1/2\}=rn\(2ℭ^n,𝐱\(B\)\(𝒜\)\+B𝐱\(𝒜\)\)1/2\.\\displaystyle=\\frac\{r\}\{\\sqrt\{n\}\}\\left\(2\\widehat\{\\mathfrak\{C\}\}\_\{n,\\mathbf\{x\}\}^\{\(B\)\}\\left\(\\mathcal\{A\}\\right\)\+B\_\{\\mathbf\{x\}\}\\left\(\\mathcal\{A\}\\right\)\\right\)^\{1/2\}\.\(311\)Here, the inequality substitutes \([310](https://arxiv.org/html/2608.13882#A4.E310)\) into the Jensen estimate, and the equality simplifiesn1/2/n=n−1/2n^\{1/2\}/n=n^\{\-1/2\}\. This proves \([287](https://arxiv.org/html/2608.13882#A4.E287)\)\. We finally specialize the result to the finite lower\-support architecture\. By \([37](https://arxiv.org/html/2608.13882#S2.E37)\),
𝒟Lm,G\\displaystyle\\mathcal\{D\}\_\{L\}^\{m,G\}=\(a\)⋃a∈𝒜L−1m,G\{f∈ℋka:‖f‖ℋka≤1\}=\(b\)𝒟1\(𝒜L−1m,G\)\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\bigcup\_\{a\\in\\mathcal\{A\}\_\{L\-1\}^\{m,G\}\}\\left\\\{f\\in\\mathcal\{H\}\_\{k\_\{a\}\}:\\left\\\|f\\right\\\|\_\{\\mathcal\{H\}\_\{k\_\{a\}\}\}\\leq 1\\right\\\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\mathcal\{D\}\_\{1\}\\left\(\\mathcal\{A\}\_\{L\-1\}^\{m,G\}\\right\)\.\(312\)Here, \(a\) applies the definition of𝒟Lm,G\\mathcal\{D\}\_\{L\}^\{m,G\}, and \(b\) applies \([285](https://arxiv.org/html/2608.13882#A4.E285)\) withr=1r=1\. This proves \([289](https://arxiv.org/html/2608.13882#A4.E289)\)\. Moreover, for everya∈𝒜L−1m,Ga\\in\\mathcal\{A\}\_\{L\-1\}^\{m,G\}andi∈\[n\]i\\in\\left\[n\\right\],
ka\(𝐱i,𝐱i\)\\displaystyle k\_\{a\}\\left\(\\mathbf\{x\}\_\{i\},\\mathbf\{x\}\_\{i\}\\right\)=\(a\)k\(B\)\(a\(𝐱i\),a\(𝐱i\)\)=\(b\)\|a\(𝐱i\)\|\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}k^\{\(\\mathrm\{B\}\)\}\\left\(a\\left\(\\mathbf\{x\}\_\{i\}\\right\),a\\left\(\\mathbf\{x\}\_\{i\}\\right\)\\right\)\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\left\|a\\left\(\\mathbf\{x\}\_\{i\}\\right\)\\right\|\.\(313\)Here, \(a\) applies \([282](https://arxiv.org/html/2608.13882#A4.E282)\), and \(b\) usesk\(B\)\(s,s\)=\|s\|k^\{\(\\mathrm\{B\}\)\}\\left\(s,s\\right\)=\\left\|s\\right\|,s∈ℝs\\in\\mathbb\{R\}\. Taking the maximum overi∈\[n\]i\\in\\left\[n\\right\]and then the supremum overa∈𝒜L−1m,Ga\\in\\mathcal\{A\}\_\{L\-1\}^\{m,G\}in \([313](https://arxiv.org/html/2608.13882#A4.E313)\) proves \([290](https://arxiv.org/html/2608.13882#A4.E290)\)\. Finally, apply \([287](https://arxiv.org/html/2608.13882#A4.E287)\) with𝒜=𝒜L−1m,G,r=1\\mathcal\{A\}=\\mathcal\{A\}\_\{L\-1\}^\{m,G\},r=1\. Using \([289](https://arxiv.org/html/2608.13882#A4.E289)\) and \([290](https://arxiv.org/html/2608.13882#A4.E290)\) gives \([291](https://arxiv.org/html/2608.13882#A4.E291)\)\. This completes the proof\.
∎
###### Lemma B0\.
\(Finite threshold\-chaos bound\)Letn≥1n\\geq 1, and let𝒯\\mathcal\{T\}be a nonempty finite collection of subsets of\[n\]\\left\[n\\right\]\. SetN:=\|𝒯\|≥1\.N:\\allowbreak=\\left\|\\mathcal\{T\}\\right\|\\geq 1\.For everyT∈𝒯T\\in\\mathcal\{T\}, define
QT\(ε\):=1n∑1≤i<j≤nεiεj𝟏\{i∈T\}𝟏\{j∈T\},\\displaystyle Q\_\{T\}\\left\(\\varepsilon\\right\):=\\frac\{1\}\{n\}\\sum\_\{1\\leq i<j\\leq n\}\\varepsilon\_\{i\}\\varepsilon\_\{j\}\\mathbf\{1\}\_\{\\left\\\{i\\in T\\right\\\}\}\\mathbf\{1\}\_\{\\left\\\{j\\in T\\right\\\}\},\(314\)whereε1,…,εn\\varepsilon\_\{1\},\\ldots,\\varepsilon\_\{n\}are independent Rademacher random variables\. Then there exists a universal numerical constantC\>0C\>0such that
𝔼ε\[maxT∈𝒯\|QT\(ε\)\|\]≤Cln\(1\+N\)\.\\displaystyle\\mathbb\{E\}\_\{\\varepsilon\}\\left\[\\max\_\{T\\in\\mathcal\{T\}\}\\left\|Q\_\{T\}\\left\(\\varepsilon\\right\)\\right\|\\right\]\\leq C\\ln\\left\(1\+N\\right\)\.\(315\)
###### Proof\.
For everyT∈𝒯T\\in\\mathcal\{T\}, define
ST:=∑i∈Tεi=∑i=1nεi𝟏\{i∈T\}\.\\displaystyle S\_\{T\}:=\\sum\_\{i\\in T\}\\varepsilon\_\{i\}=\\sum\_\{i=1\}^\{n\}\\varepsilon\_\{i\}\\mathbf\{1\}\_\{\\left\\\{i\\in T\\right\\\}\}\.\(316\)Expanding the square gives
ST2\\displaystyle S\_\{T\}^\{2\}=\(a\)∑i=1nεi2𝟏\{i∈T\}\+2∑1≤i<j≤nεiεj𝟏\{i∈T\}𝟏\{j∈T\}=\(b\)\|T\|\+2nQT\(ε\)\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\sum\_\{i=1\}^\{n\}\\varepsilon\_\{i\}^\{2\}\\mathbf\{1\}\_\{\\left\\\{i\\in T\\right\\\}\}\+2\\sum\_\{1\\leq i<j\\leq n\}\\varepsilon\_\{i\}\\varepsilon\_\{j\}\\mathbf\{1\}\_\{\\left\\\{i\\in T\\right\\\}\}\\mathbf\{1\}\_\{\\left\\\{j\\in T\\right\\\}\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\left\|T\\right\|\+2nQ\_\{T\}\\left\(\\varepsilon\\right\)\.\(317\)Here, \(a\) expands the square of \([316](https://arxiv.org/html/2608.13882#A4.E316)\), while \(b\) usesεi2=1,i∈\[n\]\\varepsilon\_\{i\}^\{2\}=1,\\qquad i\\in\\left\[n\\right\], and applies the definition \([314](https://arxiv.org/html/2608.13882#A4.E314)\)\. Rearranging \([317](https://arxiv.org/html/2608.13882#A4.E317)\) yields
QT\(ε\)=ST2−\|T\|2n\.\\displaystyle Q\_\{T\}\\left\(\\varepsilon\\right\)=\\frac\{S\_\{T\}^\{2\}\-\\left\|T\\right\|\}\{2n\}\.\(318\)Consequently,
\|QT\(ε\)\|\\displaystyle\\left\|Q\_\{T\}\\left\(\\varepsilon\\right\)\\right\|=\(a\)12n\|ST2−\|T\|\|≤\(b\)ST22n\+\|T\|2n≤\(c\)ST22n\+12\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\frac\{1\}\{2n\}\\left\|S\_\{T\}^\{2\}\-\\left\|T\\right\|\\right\|\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}\\frac\{S\_\{T\}^\{2\}\}\{2n\}\+\\frac\{\\left\|T\\right\|\}\{2n\}\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{\\leq\}\}\\frac\{S\_\{T\}^\{2\}\}\{2n\}\+\\frac\{1\}\{2\}\.\(319\)Here, \(a\) applies \([318](https://arxiv.org/html/2608.13882#A4.E318)\), \(b\) applies the triangle inequality, and \(c\) uses\|T\|≤n\\left\|T\\right\|\\leq n\. Taking the maximum overT∈𝒯T\\in\\mathcal\{T\}in \([319](https://arxiv.org/html/2608.13882#A4.E319)\) gives
maxT∈𝒯\|QT\(ε\)\|≤12nmaxT∈𝒯ST2\+12\.\\displaystyle\\max\_\{T\\in\\mathcal\{T\}\}\\left\|Q\_\{T\}\\left\(\\varepsilon\\right\)\\right\|\\leq\\frac\{1\}\{2n\}\\max\_\{T\\in\\mathcal\{T\}\}S\_\{T\}^\{2\}\+\\frac\{1\}\{2\}\.\(320\)We next derive a uniform tail bound for the random variablesSTS\_\{T\}\. FixT∈𝒯T\\in\\mathcal\{T\}andλ∈ℝ\\lambda\\in\\mathbb\{R\}\. Independence of the Rademacher variables gives
𝔼ε\[exp\(λST\)\]\\displaystyle\\mathbb\{E\}\_\{\\varepsilon\}\\left\[\\exp\\left\(\\lambda S\_\{T\}\\right\)\\right\]=\(a\)∏i∈T𝔼εi\[exp\(λεi\)\]=\(b\)\(cosh\(λ\)\)\|T\|≤\(c\)exp\(λ2\|T\|2\)\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\prod\_\{i\\in T\}\\mathbb\{E\}\_\{\\varepsilon\_\{i\}\}\\left\[\\exp\\left\(\\lambda\\varepsilon\_\{i\}\\right\)\\right\]\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\left\(\\cosh\\left\(\\lambda\\right\)\\right\)^\{\\left\|T\\right\|\}\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{\\leq\}\}\\exp\\left\(\\frac\{\\lambda^\{2\}\\left\|T\\right\|\}\{2\}\\right\)\(321\)≤\(d\)exp\(λ2n2\)\.\\displaystyle\\stackrel\{\{\\scriptstyle\(d\)\}\}\{\{\\leq\}\}\\exp\\left\(\\frac\{\\lambda^\{2\}n\}\{2\}\\right\)\.\(322\)Here, \(a\) uses independence, \(b\) uses𝔼εi\[exp\(λεi\)\]=eλ\+e−λ2=cosh\(λ\)\\mathbb\{E\}\_\{\\varepsilon\_\{i\}\}\\left\[\\exp\\left\(\\lambda\\varepsilon\_\{i\}\\right\)\\right\]=\\frac\{e^\{\\lambda\}\+e^\{\-\\lambda\}\}\{2\}=\\cosh\\left\(\\lambda\\right\), \(c\) applies the elementary inequalitycosh\(λ\)≤exp\(λ22\)\\cosh\\left\(\\lambda\\right\)\\leq\\exp\\left\(\\frac\{\\lambda^\{2\}\}\{2\}\\right\),λ∈ℝ\\lambda\\in\\mathbb\{R\}, and \(d\) uses\|T\|≤n\\left\|T\\right\|\\leq n\. Lett\>0t\>0\. For everyλ\>0\\lambda\>0, Markov’s inequality and \([322](https://arxiv.org/html/2608.13882#A4.E322)\) give
ℙε\(ST≥t\)\\displaystyle\\mathbb\{P\}\_\{\\varepsilon\}\\left\(S\_\{T\}\\geq t\\right\)=\(a\)ℙε\(exp\(λST\)≥exp\(λt\)\)≤\(b\)exp\(−λt\)𝔼ε\[exp\(λST\)\]\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\mathbb\{P\}\_\{\\varepsilon\}\\left\(\\exp\\left\(\\lambda S\_\{T\}\\right\)\\geq\\exp\\left\(\\lambda t\\right\)\\right\)\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}\\exp\\left\(\-\\lambda t\\right\)\\mathbb\{E\}\_\{\\varepsilon\}\\left\[\\exp\\left\(\\lambda S\_\{T\}\\right\)\\right\]\(323\)≤\(c\)exp\(−λt\+λ2n2\)\.\\displaystyle\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{\\leq\}\}\\exp\\left\(\-\\lambda t\+\\frac\{\\lambda^\{2\}n\}\{2\}\\right\)\.\(324\)Here, \(a\) uses the strict monotonicity of the exponential function, \(b\) applies Markov’s inequality, and \(c\) applies \([322](https://arxiv.org/html/2608.13882#A4.E322)\)\. Choosingλ:=tn\\lambda:\\allowbreak=\\frac\{t\}\{n\}in \([324](https://arxiv.org/html/2608.13882#A4.E324)\) yields
ℙε\(ST≥t\)≤exp\(−t22n\)\.\\displaystyle\\mathbb\{P\}\_\{\\varepsilon\}\\left\(S\_\{T\}\\geq t\\right\)\\leq\\exp\\left\(\-\\frac\{t^\{2\}\}\{2n\}\\right\)\.\(325\)Applying the same argument to−ST\-S\_\{T\}gives
ℙε\(ST≤−t\)≤exp\(−t22n\)\.\\displaystyle\\mathbb\{P\}\_\{\\varepsilon\}\\left\(S\_\{T\}\\leq\-t\\right\)\\leq\\exp\\left\(\-\\frac\{t^\{2\}\}\{2n\}\\right\)\.\(326\)Therefore,
ℙε\(\|ST\|≥t\)\\displaystyle\\mathbb\{P\}\_\{\\varepsilon\}\\left\(\\left\|S\_\{T\}\\right\|\\geq t\\right\)≤\(a\)ℙε\(ST≥t\)\+ℙε\(ST≤−t\)≤\(b\)2exp\(−t22n\)\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}\\mathbb\{P\}\_\{\\varepsilon\}\\left\(S\_\{T\}\\geq t\\right\)\+\\mathbb\{P\}\_\{\\varepsilon\}\\left\(S\_\{T\}\\leq\-t\\right\)\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}2\\exp\\left\(\-\\frac\{t^\{2\}\}\{2n\}\\right\)\.\(327\)Here, \(a\) applies the union bound, and \(b\) combines \([325](https://arxiv.org/html/2608.13882#A4.E325)\) and \([326](https://arxiv.org/html/2608.13882#A4.E326)\)\. Applying the union bound over theNNsets in𝒯\\mathcal\{T\}gives
ℙε\(maxT∈𝒯\|ST\|≥t\)\\displaystyle\\mathbb\{P\}\_\{\\varepsilon\}\\left\(\\max\_\{T\\in\\mathcal\{T\}\}\\left\|S\_\{T\}\\right\|\\geq t\\right\)≤\(a\)∑T∈𝒯ℙε\(\|ST\|≥t\)≤\(b\)2Nexp\(−t22n\)\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}\\sum\_\{T\\in\\mathcal\{T\}\}\\mathbb\{P\}\_\{\\varepsilon\}\\left\(\\left\|S\_\{T\}\\right\|\\geq t\\right\)\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}2N\\exp\\left\(\-\\frac\{t^\{2\}\}\{2n\}\\right\)\.\(328\)Here, \(a\) applies the union bound, while \(b\) applies \([327](https://arxiv.org/html/2608.13882#A4.E327)\) to everyT∈𝒯T\\in\\mathcal\{T\}\. Define
Z:=maxT∈𝒯ST2\.\\displaystyle Z:=\\max\_\{T\\in\\mathcal\{T\}\}S\_\{T\}^\{2\}\.\(329\)SinceZ≥0Z\\geq 0, the tail\-integral representation gives
𝔼ε\[Z\]=∫0∞ℙε\(Z≥s\)𝑑s\.\\displaystyle\\mathbb\{E\}\_\{\\varepsilon\}\\left\[Z\\right\]=\\int\_\{0\}^\{\\infty\}\\mathbb\{P\}\_\{\\varepsilon\}\\left\(Z\\geq s\\right\)\\,\\mathrm\{d\}s\.\(330\)Moreover,
\{Z≥s\}\\displaystyle\\left\\\{Z\\geq s\\right\\\}=\(a\)\{maxT∈𝒯\|ST\|≥s1/2\}\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\left\\\{\\max\_\{T\\in\\mathcal\{T\}\}\\left\|S\_\{T\}\\right\|\\geq s^\{1/2\}\\right\\\}\.\(331\)Here, \(a\) follows from the definition \([329](https://arxiv.org/html/2608.13882#A4.E329)\)\. Consequently, \([328](https://arxiv.org/html/2608.13882#A4.E328)\) gives
ℙε\(Z≥s\)≤2Nexp\(−s2n\)\.\\displaystyle\\mathbb\{P\}\_\{\\varepsilon\}\\left\(Z\\geq s\\right\)\\leq 2N\\exp\\left\(\-\\frac\{s\}\{2n\}\\right\)\.\(332\)Sets0:=2nln\(2N\)\.s\_\{0\}:\\allowbreak=2n\\ln\\left\(2N\\right\)\.Since probabilities are bounded above by one,
𝔼ε\[Z\]\\displaystyle\\mathbb\{E\}\_\{\\varepsilon\}\\left\[Z\\right\]=\(a\)∫0s0ℙε\(Z≥s\)𝑑s\+∫s0∞ℙε\(Z≥s\)𝑑s\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\int\_\{0\}^\{s\_\{0\}\}\\mathbb\{P\}\_\{\\varepsilon\}\\left\(Z\\geq s\\right\)\\,\\mathrm\{d\}s\+\\int\_\{s\_\{0\}\}^\{\\infty\}\\mathbb\{P\}\_\{\\varepsilon\}\\left\(Z\\geq s\\right\)\\,\\mathrm\{d\}s≤\(b\)s0\+∫s0∞2Nexp\(−s2n\)𝑑s=\(c\)2nln\(2N\)\+4nNexp\(−s02n\)\\displaystyle\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}s\_\{0\}\+\\int\_\{s\_\{0\}\}^\{\\infty\}2N\\exp\\left\(\-\\frac\{s\}\{2n\}\\right\)\\,\\mathrm\{d\}s\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{=\}\}2n\\ln\\left\(2N\\right\)\+4nN\\exp\\left\(\-\\frac\{s\_\{0\}\}\{2n\}\\right\)=\(d\)2nln\(2N\)\+2n\.\\displaystyle\\stackrel\{\{\\scriptstyle\(d\)\}\}\{\{=\}\}2n\\ln\\left\(2N\\right\)\+2n\.\(333\)Here, \(a\) splits the integral in \([330](https://arxiv.org/html/2608.13882#A4.E330)\) ats0s\_\{0\}, \(b\) bounds the first probability by one and applies \([332](https://arxiv.org/html/2608.13882#A4.E332)\) to the second integral, \(c\) evaluates the exponential integral, and \(d\) usesexp\(−s02n\)=exp\(−ln\(2N\)\)=12N\\exp\\left\(\-\\frac\{s\_\{0\}\}\{2n\}\\right\)=\\exp\\left\(\-\\ln\\left\(2N\\right\)\\right\)=\\frac\{1\}\{2N\}\. Taking expectations in \([320](https://arxiv.org/html/2608.13882#A4.E320)\) and applying \([333](https://arxiv.org/html/2608.13882#A4.E333)\) gives
𝔼ε\[maxT∈𝒯\|QT\(ε\)\|\]\\displaystyle\\mathbb\{E\}\_\{\\varepsilon\}\\left\[\\max\_\{T\\in\\mathcal\{T\}\}\\left\|Q\_\{T\}\\left\(\\varepsilon\\right\)\\right\|\\right\]≤\(a\)12n𝔼ε\[Z\]\+12≤\(b\)ln\(2N\)\+32≤\(c\)Cln\(1\+N\)\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}\\frac\{1\}\{2n\}\\mathbb\{E\}\_\{\\varepsilon\}\\left\[Z\\right\]\+\\frac\{1\}\{2\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}\\ln\\left\(2N\\right\)\+\\frac\{3\}\{2\}\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{\\leq\}\}C\\ln\\left\(1\+N\\right\)\.\(334\)Here, \(a\) applies \([320](https://arxiv.org/html/2608.13882#A4.E320)\), \(b\) substitutes \([333](https://arxiv.org/html/2608.13882#A4.E333)\), and \(c\) usesN≥1N\\geq 1together with2N≤\(1\+N\)22N\\leq\\left\(1\+N\\right\)^\{2\},32≤32ln2ln\(1\+N\)\\frac\{3\}\{2\}\\leq\\frac\{3\}\{2\\ln 2\}\\ln\\left\(1\+N\\right\), and enlarges the universal numerical constant\. This proves \([315](https://arxiv.org/html/2608.13882#A4.E315)\) and completes the proof\. ∎
###### Lemma B0\(Brownian chaos controlled by signed threshold traces\)\.
Letn≥1n\\geq 1, let𝐱1,…,𝐱n∈𝒳\\mathbf\{x\}\_\{1\},\\ldots,\\mathbf\{x\}\_\{n\}\\in\\mathcal\{X\}be fixed sample points, and let𝒜\\mathcal\{A\}be a nonempty class of real\-valued support functions on𝒳\\mathcal\{X\}\. Assume thatB𝐱\(𝒜\)<∞B\_\{\\mathbf\{x\}\}\\left\(\\mathcal\{A\}\\right\)<\\infty, wherekak\_\{a\}andB𝐱\(𝒜\)B\_\{\\mathbf\{x\}\}\\left\(\\mathcal\{A\}\\right\)are defined in \([282](https://arxiv.org/html/2608.13882#A4.E282)\) and \([283](https://arxiv.org/html/2608.13882#A4.E283)\), respectively\. Define the signed threshold trace family by
𝒯n,𝐱±\(𝒜\):=\{\{i∈\[n\]:σa\(𝐱i\)≥τ\}:a∈𝒜,σ∈\{−1,1\},τ∈ℝ\}\.\\displaystyle\\mathcal\{T\}\_\{n,\\mathbf\{x\}\}^\{\\pm\}\\left\(\\mathcal\{A\}\\right\):=\\left\\\{\\left\\\{i\\in\\left\[n\\right\]:\\sigma a\\left\(\\mathbf\{x\}\_\{i\}\\right\)\\geq\\tau\\right\\\}:a\\in\\mathcal\{A\},\\;\\sigma\\in\\left\\\{\-1,1\\right\\\},\\;\\tau\\in\\mathbb\{R\}\\right\\\}\.\(335\)Then𝒯n,𝐱±\(𝒜\)\\mathcal\{T\}\_\{n,\\mathbf\{x\}\}^\{\\pm\}\\left\(\\mathcal\{A\}\\right\)is a nonempty finite collection of subsets of\[n\]\\left\[n\\right\], and there exists a universal numerical constantC\>0C\>0such that
ℭ^n,𝐱\(B\)\(𝒜\)≤CB𝐱\(𝒜\)ln\(1\+\|𝒯n,𝐱±\(𝒜\)\|\)\.\\displaystyle\\widehat\{\\mathfrak\{C\}\}\_\{n,\\mathbf\{x\}\}^\{\(B\)\}\\left\(\\mathcal\{A\}\\right\)\\leq CB\_\{\\mathbf\{x\}\}\\left\(\\mathcal\{A\}\\right\)\\ln\\left\(1\+\\left\|\\mathcal\{T\}\_\{n,\\mathbf\{x\}\}^\{\\pm\}\\left\(\\mathcal\{A\}\\right\)\\right\|\\right\)\.\(336\)Here,ℭ^n,𝐱\(B\)\\widehat\{\\mathfrak\{C\}\}\_\{n,\\mathbf\{x\}\}^\{\(B\)\}is the Brownian quadratic\-chaos complexity defined in \([284](https://arxiv.org/html/2608.13882#A4.E284)\)\. Consequently, ifE≥0E\\geq 0satisfies
ln\|𝒯n,𝐱±\(𝒜\)\|≤E,\\displaystyle\\ln\\left\|\\mathcal\{T\}\_\{n,\\mathbf\{x\}\}^\{\\pm\}\\left\(\\mathcal\{A\}\\right\)\\right\|\\leq E,\(337\)then
ℭ^n,𝐱\(B\)\(𝒜\)≤CB𝐱\(𝒜\)\(1\+E\)\.\\displaystyle\\widehat\{\\mathfrak\{C\}\}\_\{n,\\mathbf\{x\}\}^\{\(B\)\}\\left\(\\mathcal\{A\}\\right\)\\leq CB\_\{\\mathbf\{x\}\}\\left\(\\mathcal\{A\}\\right\)\\left\(1\+E\\right\)\.\(338\)
###### Proof\.
Set
B:=B𝐱\(𝒜\)\.\\displaystyle B:=B\_\{\\mathbf\{x\}\}\\left\(\\mathcal\{A\}\\right\)\.\(339\)We first consider the caseB=0B=0\. By \([283](https://arxiv.org/html/2608.13882#A4.E283)\) and the diagonal identityka\(𝐱i,𝐱i\)=\|a\(𝐱i\)\|k\_\{a\}\\left\(\\mathbf\{x\}\_\{i\},\\mathbf\{x\}\_\{i\}\\right\)=\\left\|a\\left\(\\mathbf\{x\}\_\{i\}\\right\)\\right\|, one has
a\(𝐱i\)=0,a∈𝒜,i∈\[n\]\.\\displaystyle a\\left\(\\mathbf\{x\}\_\{i\}\\right\)=0,\\qquad a\\in\\mathcal\{A\},\\quad i\\in\\left\[n\\right\]\.\(340\)Consequently, for everya∈𝒜a\\in\\mathcal\{A\}andi,j∈\[n\]i,j\\in\\left\[n\\right\],
ka\(𝐱i,𝐱j\)\\displaystyle k\_\{a\}\\left\(\\mathbf\{x\}\_\{i\},\\mathbf\{x\}\_\{j\}\\right\)=\(a\)k\(B\)\(0,0\)=\(b\)0\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}k^\{\(\\mathrm\{B\}\)\}\\left\(0,0\\right\)\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}0\.\(341\)Here, \(a\) uses \([340](https://arxiv.org/html/2608.13882#A4.E340)\), and \(b\) follows from the definition of the Brownian kernel\. Therefore,ℭ^n,𝐱\(B\)\(𝒜\)=0\\widehat\{\\mathfrak\{C\}\}\_\{n,\\mathbf\{x\}\}^\{\(B\)\}\\left\(\\mathcal\{A\}\\right\)=0, and \([336](https://arxiv.org/html/2608.13882#A4.E336)\) holds trivially\. We may therefore assume throughout the remainder of the proof thatB\>0B\>0\. We first establish the threshold representation of the full Brownian kernel on\[−B,B\]\\left\[\-B,B\\right\]\. Fixs,t∈\[−B,B\]s,t\\in\\left\[\-B,B\\right\]\. The definition of the Brownian kernel gives
k\(B\)\(s,t\)\\displaystyle k^\{\(\\mathrm\{B\}\)\}\\left\(s,t\\right\)=\(a\)𝟏\{s≥0\}𝟏\{t≥0\}min\{s,t\}\+𝟏\{s≤0\}𝟏\{t≤0\}min\{−s,−t\}\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\mathbf\{1\}\_\{\\left\\\{s\\geq 0\\right\\\}\}\\mathbf\{1\}\_\{\\left\\\{t\\geq 0\\right\\\}\}\\min\\left\\\{s,t\\right\\\}\+\\mathbf\{1\}\_\{\\left\\\{s\\leq 0\\right\\\}\}\\mathbf\{1\}\_\{\\left\\\{t\\leq 0\\right\\\}\}\\min\\left\\\{\-s,\-t\\right\\\}\.\(342\)Here, \(a\) follows by considering the three possible sign configurations\. Ifs,t≥0s,t\\geq 0, then
k\(B\)\(s,t\)=s\+t−\|s−t\|2=min\{s,t\}\.\\displaystyle k^\{\(\\mathrm\{B\}\)\}\\left\(s,t\\right\)=\\frac\{s\+t\-\\left\|s\-t\\right\|\}\{2\}=\\min\\left\\\{s,t\\right\\\}\.\(343\)Ifs,t≤0s,t\\leq 0, thenk\(B\)\(s,t\)=min\{−s,−t\}k^\{\(\\mathrm\{B\}\)\}\\left\(s,t\\right\)=\\min\\left\\\{\-s,\-t\\right\\\}\. Ifssandtthave opposite signs, then\|s−t\|=\|s\|\+\|t\|\\left\|s\-t\\right\|=\\left\|s\\right\|\+\\left\|t\\right\|, and hencek\(B\)\(s,t\)=0k^\{\(\\mathrm\{B\}\)\}\\left\(s,t\\right\)=0\. For the nonnegative branch,
∫0B𝟏\{s≥r\}𝟏\{t≥r\}dr=\(a\)𝟏\{s≥0\}𝟏\{t≥0\}∫0min\{s,t\}1dr=\(b\)𝟏\{s≥0\}𝟏\{t≥0\}min\{s,t\}\.\\displaystyle\\int\_\{0\}^\{B\}\\mathbf\{1\}\_\{\\left\\\{s\\geq r\\right\\\}\}\\mathbf\{1\}\_\{\\left\\\{t\\geq r\\right\\\}\}\\,\\mathrm\{d\}r\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\mathbf\{1\}\_\{\\left\\\{s\\geq 0\\right\\\}\}\\mathbf\{1\}\_\{\\left\\\{t\\geq 0\\right\\\}\}\\int\_\{0\}^\{\\min\\left\\\{s,t\\right\\\}\}1\\,\\mathrm\{d\}r\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\mathbf\{1\}\_\{\\left\\\{s\\geq 0\\right\\\}\}\\mathbf\{1\}\_\{\\left\\\{t\\geq 0\\right\\\}\}\\min\\left\\\{s,t\\right\\\}\.\(344\)Here, \(a\) uses the fact that both threshold inequalities hold exactly forr∈\[0,min\{s,t\}\]r\\in\\left\[0,\\min\\left\\\{s,t\\right\\\}\\right\]whens,t≥0s,t\\geq 0, and for no positive\-measure set ofrrotherwise\. Equality \(b\) evaluates the integral\. Similarly,
∫0B𝟏\{s≤−r\}𝟏\{t≤−r\}dr\\displaystyle\\int\_\{0\}^\{B\}\\mathbf\{1\}\_\{\\left\\\{s\\leq\-r\\right\\\}\}\\mathbf\{1\}\_\{\\left\\\{t\\leq\-r\\right\\\}\}\\,\\mathrm\{d\}r=\(a\)𝟏\{s≤0\}𝟏\{t≤0\}∫0min\{−s,−t\}1dr\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\mathbf\{1\}\_\{\\left\\\{s\\leq 0\\right\\\}\}\\mathbf\{1\}\_\{\\left\\\{t\\leq 0\\right\\\}\}\\int\_\{0\}^\{\\min\\left\\\{\-s,\-t\\right\\\}\}1\\,\\mathrm\{d\}r=\(b\)𝟏\{s≤0\}𝟏\{t≤0\}min\{−s,−t\}\.\\displaystyle\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\mathbf\{1\}\_\{\\left\\\{s\\leq 0\\right\\\}\}\\mathbf\{1\}\_\{\\left\\\{t\\leq 0\\right\\\}\}\\min\\left\\\{\-s,\-t\\right\\\}\.\(345\)Here, \(a\) rewritess≤−rs\\leq\-randt≤−rt\\leq\-rasr≤−sr\\leq\-sandr≤−tr\\leq\-t, while \(b\) evaluates the resulting integral\. Combining \([342](https://arxiv.org/html/2608.13882#A4.E342)\), \([344](https://arxiv.org/html/2608.13882#A4.E344)\), and \([345](https://arxiv.org/html/2608.13882#A4.E345)\) gives
k\(B\)\(s,t\)\\displaystyle k^\{\(\\mathrm\{B\}\)\}\\left\(s,t\\right\)=∫0B𝟏\{s≥r\}𝟏\{t≥r\}dr\+∫0B𝟏\{s≤−r\}𝟏\{t≤−r\}dr\.\\displaystyle=\\int\_\{0\}^\{B\}\\mathbf\{1\}\_\{\\left\\\{s\\geq r\\right\\\}\}\\mathbf\{1\}\_\{\\left\\\{t\\geq r\\right\\\}\}\\,\\mathrm\{d\}r\\qquad\+\\int\_\{0\}^\{B\}\\mathbf\{1\}\_\{\\left\\\{s\\leq\-r\\right\\\}\}\\mathbf\{1\}\_\{\\left\\\{t\\leq\-r\\right\\\}\}\\,\\mathrm\{d\}r\.\(346\)Fixa∈𝒜a\\in\\mathcal\{A\}andr∈\[0,B\]r\\in\\left\[0,B\\right\]\. Define
Ta,r\+\\displaystyle T\_\{a,r\}^\{\+\}:=\{i∈\[n\]:a\(𝐱i\)≥r\},\\displaystyle:=\\left\\\{i\\in\\left\[n\\right\]:a\\left\(\\mathbf\{x\}\_\{i\}\\right\)\\geq r\\right\\\},\(347\)Ta,r−\\displaystyle T\_\{a,r\}^\{\-\}:=\{i∈\[n\]:−a\(𝐱i\)≥r\}=\{i∈\[n\]:a\(𝐱i\)≤−r\}\.\\displaystyle:=\\left\\\{i\\in\\left\[n\\right\]:\-a\\left\(\\mathbf\{x\}\_\{i\}\\right\)\\geq r\\right\\\}=\\left\\\{i\\in\\left\[n\\right\]:a\\left\(\\mathbf\{x\}\_\{i\}\\right\)\\leq\-r\\right\\\}\.\(348\)By \([335](https://arxiv.org/html/2608.13882#A4.E335)\),
Ta,r\+\\displaystyle T\_\{a,r\}^\{\+\}∈𝒯n,𝐱±\(𝒜\),\\displaystyle\\in\\mathcal\{T\}\_\{n,\\mathbf\{x\}\}^\{\\pm\}\\left\(\\mathcal\{A\}\\right\),Ta,r−\\displaystyle T\_\{a,r\}^\{\-\}∈𝒯n,𝐱±\(𝒜\)\.\\displaystyle\\in\\mathcal\{T\}\_\{n,\\mathbf\{x\}\}^\{\\pm\}\\left\(\\mathcal\{A\}\\right\)\.\(349\)By \([283](https://arxiv.org/html/2608.13882#A4.E283)\) and the diagonal identityka\(𝐱,𝐱\)=\|a\(𝐱\)\|k\_\{a\}\\left\(\\mathbf\{x\},\\mathbf\{x\}\\right\)=\\left\|a\\left\(\\mathbf\{x\}\\right\)\\right\|, the valuesa\(𝐱i\)a\\left\(\\mathbf\{x\}\_\{i\}\\right\)anda\(𝐱j\)a\\left\(\\mathbf\{x\}\_\{j\}\\right\)belong to\[−B,B\]\\left\[\-B,B\\right\]\. Applying \([346](https://arxiv.org/html/2608.13882#A4.E346)\) withs=a\(𝐱i\)s=a\\left\(\\mathbf\{x\}\_\{i\}\\right\)andt=a\(𝐱j\)t=a\\left\(\\mathbf\{x\}\_\{j\}\\right\)gives
ka\(𝐱i,𝐱j\)\\displaystyle k\_\{a\}\\left\(\\mathbf\{x\}\_\{i\},\\mathbf\{x\}\_\{j\}\\right\)=\(a\)∫0B𝟏\{i∈Ta,r\+\}𝟏\{j∈Ta,r\+\}dr\+∫0B𝟏\{i∈Ta,r−\}𝟏\{j∈Ta,r−\}dr\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\int\_\{0\}^\{B\}\\mathbf\{1\}\_\{\\left\\\{i\\in T\_\{a,r\}^\{\+\}\\right\\\}\}\\mathbf\{1\}\_\{\\left\\\{j\\in T\_\{a,r\}^\{\+\}\\right\\\}\}\\,\\mathrm\{d\}r\\qquad\+\\int\_\{0\}^\{B\}\\mathbf\{1\}\_\{\\left\\\{i\\in T\_\{a,r\}^\{\-\}\\right\\\}\}\\mathbf\{1\}\_\{\\left\\\{j\\in T\_\{a,r\}^\{\-\}\\right\\\}\}\\,\\mathrm\{d\}r\.\(350\)Here, \(a\) applies \([346](https://arxiv.org/html/2608.13882#A4.E346)\) and then uses \([347](https://arxiv.org/html/2608.13882#A4.E347)\) and \([348](https://arxiv.org/html/2608.13882#A4.E348)\)\. For everyT⊆\[n\]T\\subseteq\\left\[n\\right\], define
QT\(ε\):=1n∑1≤i<j≤nεiεj𝟏\{i∈T\}𝟏\{j∈T\}\.\\displaystyle Q\_\{T\}\\left\(\\varepsilon\\right\):=\\frac\{1\}\{n\}\\sum\_\{1\\leq i<j\\leq n\}\\varepsilon\_\{i\}\\varepsilon\_\{j\}\\mathbf\{1\}\_\{\\left\\\{i\\in T\\right\\\}\}\\mathbf\{1\}\_\{\\left\\\{j\\in T\\right\\\}\}\.\(351\)Substituting \([350](https://arxiv.org/html/2608.13882#A4.E350)\) into the Brownian chaos sum gives
1n∑1≤i<j≤nεiεjka\(𝐱i,𝐱j\)=\(a\)∫0BQTa,r\+\(ε\)𝑑r\+∫0BQTa,r−\(ε\)𝑑r\.\\displaystyle\\frac\{1\}\{n\}\\sum\_\{1\\leq i<j\\leq n\}\\varepsilon\_\{i\}\\varepsilon\_\{j\}k\_\{a\}\\left\(\\mathbf\{x\}\_\{i\},\\mathbf\{x\}\_\{j\}\\right\)\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\int\_\{0\}^\{B\}Q\_\{T\_\{a,r\}^\{\+\}\}\\left\(\\varepsilon\\right\)\\,\\mathrm\{d\}r\+\\int\_\{0\}^\{B\}Q\_\{T\_\{a,r\}^\{\-\}\}\\left\(\\varepsilon\\right\)\\,\\mathrm\{d\}r\.\(352\)Here, \(a\) interchanges the finite sum overi<ji<jwith the two integrals and applies \([351](https://arxiv.org/html/2608.13882#A4.E351)\)\. Taking absolute values in \([352](https://arxiv.org/html/2608.13882#A4.E352)\) and applying the triangle inequality gives
\|1n∑1≤i<j≤nεiεjka\(𝐱i,𝐱j\)\|\\displaystyle\\left\|\\frac\{1\}\{n\}\\sum\_\{1\\leq i<j\\leq n\}\\varepsilon\_\{i\}\\varepsilon\_\{j\}k\_\{a\}\\left\(\\mathbf\{x\}\_\{i\},\\mathbf\{x\}\_\{j\}\\right\)\\right\|≤\(a\)∫0B\|QTa,r\+\(ε\)\|𝑑r\+∫0B\|QTa,r−\(ε\)\|𝑑r\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}\\int\_\{0\}^\{B\}\\left\|Q\_\{T\_\{a,r\}^\{\+\}\}\\left\(\\varepsilon\\right\)\\right\|\\,\\mathrm\{d\}r\+\\int\_\{0\}^\{B\}\\left\|Q\_\{T\_\{a,r\}^\{\-\}\}\\left\(\\varepsilon\\right\)\\right\|\\,\\mathrm\{d\}r≤\(b\)2BmaxT∈𝒯n,𝐱±\(𝒜\)\|QT\(ε\)\|\.\\displaystyle\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}2B\\max\_\{T\\in\\mathcal\{T\}\_\{n,\\mathbf\{x\}\}^\{\\pm\}\\left\(\\mathcal\{A\}\\right\)\}\\left\|Q\_\{T\}\\left\(\\varepsilon\\right\)\\right\|\.\(353\)Here, \(a\) applies the triangle inequality for integrals, while \(b\) uses \([349](https://arxiv.org/html/2608.13882#A4.E349)\) and the fact that both integration intervals have lengthBB\. The right\-hand side of \([353](https://arxiv.org/html/2608.13882#A4.E353)\) does not depend onaa\. Therefore,
supa∈𝒜\|1n∑1≤i<j≤nεiεjka\(𝐱i,𝐱j\)\|≤2BmaxT∈𝒯n,𝐱±\(𝒜\)\|QT\(ε\)\|\.\\displaystyle\\sup\_\{a\\in\\mathcal\{A\}\}\\left\|\\frac\{1\}\{n\}\\sum\_\{1\\leq i<j\\leq n\}\\varepsilon\_\{i\}\\varepsilon\_\{j\}k\_\{a\}\\left\(\\mathbf\{x\}\_\{i\},\\mathbf\{x\}\_\{j\}\\right\)\\right\|\\leq 2B\\max\_\{T\\in\\mathcal\{T\}\_\{n,\\mathbf\{x\}\}^\{\\pm\}\\left\(\\mathcal\{A\}\\right\)\}\\left\|Q\_\{T\}\\left\(\\varepsilon\\right\)\\right\|\.\(354\)Since every member of𝒯n,𝐱±\(𝒜\)\\mathcal\{T\}\_\{n,\\mathbf\{x\}\}^\{\\pm\}\\left\(\\mathcal\{A\}\\right\)is a subset of\[n\]\\left\[n\\right\],1≤\|𝒯n,𝐱±\(𝒜\)\|≤2n1\\leq\\left\|\\mathcal\{T\}\_\{n,\\mathbf\{x\}\}^\{\\pm\}\\left\(\\mathcal\{A\}\\right\)\\right\|\\leq 2^\{n\}\. Thus, the trace family is finite and nonempty\. Taking expectations in \([354](https://arxiv.org/html/2608.13882#A4.E354)\) and applying[Lemma21](https://arxiv.org/html/2608.13882#Thmtheorem21)gives
ℭ^n,𝐱\(B\)\(𝒜\)\\displaystyle\\widehat\{\\mathfrak\{C\}\}\_\{n,\\mathbf\{x\}\}^\{\(B\)\}\\left\(\\mathcal\{A\}\\right\)≤\(a\)2B𝔼ε\[maxT∈𝒯n,𝐱±\(𝒜\)\|QT\(ε\)\|\]≤\(b\)2CBln\(1\+\|𝒯n,𝐱±\(𝒜\)\|\)≤\(c\)CBln\(1\+\|𝒯n,𝐱±\(𝒜\)\|\)\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}2B\\mathbb\{E\}\_\{\\varepsilon\}\\left\[\\max\_\{T\\in\\mathcal\{T\}\_\{n,\\mathbf\{x\}\}^\{\\pm\}\\left\(\\mathcal\{A\}\\right\)\}\\left\|Q\_\{T\}\\left\(\\varepsilon\\right\)\\right\|\\right\]\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}2CB\\ln\\left\(1\+\\left\|\\mathcal\{T\}\_\{n,\\mathbf\{x\}\}^\{\\pm\}\\left\(\\mathcal\{A\}\\right\)\\right\|\\right\)\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{\\leq\}\}CB\\ln\\left\(1\+\\left\|\\mathcal\{T\}\_\{n,\\mathbf\{x\}\}^\{\\pm\}\\left\(\\mathcal\{A\}\\right\)\\right\|\\right\)\.\(355\)Here, \(a\) applies \([354](https://arxiv.org/html/2608.13882#A4.E354)\) and the definition \([284](https://arxiv.org/html/2608.13882#A4.E284)\), \(b\) applies[Lemma21](https://arxiv.org/html/2608.13882#Thmtheorem21), and \(c\) enlarges the universal numerical constant to absorb the factor22\. Substituting \([339](https://arxiv.org/html/2608.13882#A4.E339)\) proves \([336](https://arxiv.org/html/2608.13882#A4.E336)\)\. Finally, suppose that \([337](https://arxiv.org/html/2608.13882#A4.E337)\) holds\. SetN𝒯:=\|𝒯n,𝐱±\(𝒜\)\|N\_\{\\mathcal\{T\}\}:=\\left\|\\mathcal\{T\}\_\{n,\\mathbf\{x\}\}^\{\\pm\}\\left\(\\mathcal\{A\}\\right\)\\right\|\. SinceN𝒯≥1N\_\{\\mathcal\{T\}\}\\geq 1,
ln\(1\+N𝒯\)\\displaystyle\\ln\\left\(1\+N\_\{\\mathcal\{T\}\}\\right\)≤\(a\)ln\(2N𝒯\)=\(b\)ln2\+lnN𝒯≤\(c\)1\+E\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}\\ln\\left\(2N\_\{\\mathcal\{T\}\}\\right\)\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\ln 2\+\\ln N\_\{\\mathcal\{T\}\}\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{\\leq\}\}1\+E\.\(356\)Here, \(a\) uses1\+N𝒯≤2N𝒯1\+N\_\{\\mathcal\{T\}\}\\leq 2N\_\{\\mathcal\{T\}\}, \(b\) uses the logarithm of a product, and \(c\) usesln2≤1\\ln 2\\leq 1together with \([337](https://arxiv.org/html/2608.13882#A4.E337)\)\. Combining \([336](https://arxiv.org/html/2608.13882#A4.E336)\) and \([356](https://arxiv.org/html/2608.13882#A4.E356)\) proves \([338](https://arxiv.org/html/2608.13882#A4.E338)\)\. This completes the proof\.
∎
###### Lemma B0\.
\(VC dimension of signed finite\-architecture threshold classes\)LetL,m,G∈ℕL,m,G\\in\\mathbb\{N\}satisfyL≥2L\\geq 2,m≥1m\\geq 1, andG≥2G\\geq 2\. LetPL−1,m,GP\_\{L\-1,m,G\}be the parameter\-count bound specified in[Section2\.2](https://arxiv.org/html/2608.13882#S2.SS2)and established in[Lemma17](https://arxiv.org/html/2608.13882#Thmtheorem17), part[3](https://arxiv.org/html/2608.13882#A4.I1.i3), and setP:=PL−1,m,G\.P:=P\_\{L\-1,m,G\}\.LetΘL−1,m,G⊆ℝP\\Theta\_\{L\-1,m,G\}\\subseteq\\mathbb\{R\}^\{P\}denote the admissible parameter set of𝒜L−1m,G\\mathcal\{A\}\_\{L\-1\}^\{m,G\}\. If a particular representation uses fewer thanPPcoordinates, identify it with an element ofℝP\\mathbb\{R\}^\{P\}by padding the remaining coordinates with zeros\. For𝛉∈ΘL−1,m,G\\bm\{\\theta\}\\in\\Theta\_\{L\-1,m,G\}, leta𝛉a\_\{\\bm\{\\theta\}\}denote the associated lower\-support function\. Define
ℋL−1,m,G±:=\{h𝜽,τ,σ:𝜽∈ΘL−1,m,G,τ∈ℝ,σ∈\{−1,1\}\},\\displaystyle\\mathcal\{H\}\_\{L\-1,m,G\}^\{\\pm\}:=\\left\\\{h\_\{\\bm\{\\theta\},\\tau,\\sigma\}:\\bm\{\\theta\}\\in\\Theta\_\{L\-1,m,G\},\\;\\tau\\in\\mathbb\{R\},\\;\\sigma\\in\\left\\\{\-1,1\\right\\\}\\right\\\},\(357\)where
h𝜽,τ,σ\(𝐱\):=𝟏\{σa𝜽\(𝐱\)≥τ\},𝐱∈𝒳\.\\displaystyle h\_\{\\bm\{\\theta\},\\tau,\\sigma\}\\left\(\\mathbf\{x\}\\right\):=\\mathbf\{1\}\_\{\\left\\\{\\sigma a\_\{\\bm\{\\theta\}\}\\left\(\\mathbf\{x\}\\right\)\\geq\\tau\\right\\\}\},\\qquad\\mathbf\{x\}\\in\\mathcal\{X\}\.\(358\)Then there existsCL\>0C\_\{L\}\>0, depending only onLL, such that
VC\(ℋL−1,m,G±\)≤CL\(P\+1\)ln\(e\(P\+1\)\)\.\\displaystyle\\operatorname\{VC\}\\left\(\\mathcal\{H\}\_\{L\-1,m,G\}^\{\\pm\}\\right\)\\leq C\_\{L\}\\left\(P\+1\\right\)\\ln\\left\(e\\left\(P\+1\\right\)\\right\)\.\(359\)Equivalently,
VC\(ℋL−1,m,G±\)≤CL\(PL−1,m,G\+1\)ln\(e\(PL−1,m,G\+1\)\)\.\\displaystyle\\operatorname\{VC\}\\left\(\\mathcal\{H\}\_\{L\-1,m,G\}^\{\\pm\}\\right\)\\leq C\_\{L\}\\left\(P\_\{L\-1,m,G\}\+1\\right\)\\ln\\left\(e\\left\(P\_\{L\-1,m,G\}\+1\\right\)\\right\)\.\(360\)
###### Proof\.
SetW:=P\+1W:=P\+1\. The additional coordinate inWWis the variable thresholdτ\\tau\. For each fixedσ∈\{−1,1\}\\sigma\\in\\left\\\{\-1,1\\right\\\}, the corresponding classifiers are therefore parameterized by𝜼:=\(𝜽,τ\)∈ℝW\\bm\{\\eta\}:=\\left\(\\bm\{\\theta\},\\tau\\right\)\\in\\mathbb\{R\}^\{W\}\. We divide the proof into four steps\.
Step 1: enlargement of the admissible parameter set\.The admissible architecture parameters satisfy the direction, profile\-energy, mixing, and readout constraints introduced in[Section2\.2](https://arxiv.org/html/2608.13882#S2.SS2)\. For an upper bound on the number of possible labelings, we may remove all of these constraints\. More precisely, let𝒜~L−1m,G\\widetilde\{\\mathcal\{A\}\}\_\{L\-1\}^\{m,G\}denote the class obtained by allowing every coordinate of𝜽\\bm\{\\theta\}to vary freely inℝP\\mathbb\{R\}^\{P\}, while retaining the same computational graph, the same interpolation knots, and the same constant profile extensions\. Then𝒜L−1m,G⊆𝒜~L−1m,G\\mathcal\{A\}\_\{L\-1\}^\{m,G\}\\subseteq\\widetilde\{\\mathcal\{A\}\}\_\{L\-1\}^\{m,G\}\. Consequently,ℋL−1,m,G±⊆ℋ~L−1,m,G±\\mathcal\{H\}\_\{L\-1,m,G\}^\{\\pm\}\\subseteq\\widetilde\{\\mathcal\{H\}\}\_\{L\-1,m,G\}^\{\\pm\}, whereℋ~L−1,m,G±\\widetilde\{\\mathcal\{H\}\}\_\{L\-1,m,G\}^\{\\pm\}is defined by \([357](https://arxiv.org/html/2608.13882#A4.E357)\) with𝜽∈ℝP\\bm\{\\theta\}\\in\\mathbb\{R\}^\{P\}\. By the monotonicity of the VC dimension under class inclusion,
VC\(ℋL−1,m,G±\)\\displaystyle\\operatorname\{VC\}\\left\(\\mathcal\{H\}\_\{L\-1,m,G\}^\{\\pm\}\\right\)≤VC\(ℋ~L−1,m,G±\)\.\\displaystyle\\leq\\operatorname\{VC\}\\left\(\\widetilde\{\\mathcal\{H\}\}\_\{L\-1,m,G\}^\{\\pm\}\\right\)\.\(361\)We will therefore count the labelings generated by the enlarged class\.
Step 2: polynomial sign\-pattern estimate\.We use the following standard sign\-pattern estimate\. Letp1,…,pMp\_\{1\},\\ldots,p\_\{M\}be real polynomials inWWreal variables, each having degree at mostDD\. Then the number of realizable sign vectors
\(sgn\(p1\(𝜼\)\),…,sgn\(pM\(𝜼\)\)\)∈\{−1,0,1\}M,𝜼∈ℝW,\\displaystyle\\left\(\\operatorname\{sgn\}\\left\(p\_\{1\}\\left\(\\bm\{\\eta\}\\right\)\\right\),\\ldots,\\operatorname\{sgn\}\\left\(p\_\{M\}\\left\(\\bm\{\\eta\}\\right\)\\right\)\\right\)\\in\\left\\\{\-1,0,1\\right\\\}^\{M\},\\qquad\\bm\{\\eta\}\\in\\mathbb\{R\}^\{W\},\(362\)is bounded by
𝔖\(M,D,W\)≤\(C0D\(M\+W\)\)W,\\displaystyle\\mathfrak\{S\}\\left\(M,D,W\\right\)\\leq\\left\(C\_\{0\}D\\left\(M\+W\\right\)\\right\)^\{W\},\(363\)whereC0\>0C\_\{0\}\>0is a universal numerical constant\. This is the standard polynomial sign\-pattern estimate underlying semialgebraic VC\-dimension bounds; see, for example,[13](https://arxiv.org/html/2608.13882#bib.bib2)\. For a binary classℋ\\mathcal\{H\}, define its growth function by
Πℋ\(N\):=sup\(𝐱1,…,𝐱N\)∈𝒳N\|\{\(h\(𝐱1\),…,h\(𝐱N\)\):h∈ℋ\}\|\.\\displaystyle\\Pi\_\{\\mathcal\{H\}\}\\left\(N\\right\):=\\sup\_\{\\left\(\\mathbf\{x\}\_\{1\},\\ldots,\\mathbf\{x\}\_\{N\}\\right\)\\in\\mathcal\{X\}^\{N\}\}\\left\|\\left\\\{\\left\(h\\left\(\\mathbf\{x\}\_\{1\}\\right\),\\ldots,h\\left\(\\mathbf\{x\}\_\{N\}\\right\)\\right\):h\\in\\mathcal\{H\}\\right\\\}\\right\|\.\(364\)We will prove that there exists a constantCL,1\>0C\_\{L,1\}\>0, depending only onLL, such that
Πℋ~L−1,m,G±\(N\)≤2\(CL,1NP\)\(L−1\)W\.\\displaystyle\\Pi\_\{\\widetilde\{\\mathcal\{H\}\}\_\{L\-1,m,G\}^\{\\pm\}\}\\left\(N\\right\)\\leq 2\\left\(C\_\{L,1\}NP\\right\)^\{\\left\(L\-1\\right\)W\}\.\(365\)
Step 3: layerwise semialgebraic decomposition\. FixN≥1N\\geq 1and fixed points𝐱1,…,𝐱N∈𝒳\\mathbf\{x\}\_\{1\},\\ldots,\\mathbf\{x\}\_\{N\}\\in\\mathcal\{X\}\. SetK:=L−2K:=L\-2\. Thus,K=0K=0whenL=2L=2, whereasK≥1K\\geq 1whenL≥3L\\geq 3\. We first identify the polynomial degree generated by the recursive architecture\. WhenL=2L=2, one has
a𝜽\(𝐱i\)=𝝎⊤𝐱i,\\displaystyle a\_\{\\bm\{\\theta\}\}\\left\(\\mathbf\{x\}\_\{i\}\\right\)=\\bm\{\\omega\}^\{\\top\}\\mathbf\{x\}\_\{i\},\(366\)which is a polynomial of degree one in𝜽\\bm\{\\theta\}\. Suppose now thatL≥3L\\geq 3\. For each profileqj\(r\)q\_\{j\}^\{\(r\)\}, write its nodal parameters ascj,0\(r\),…,cj,G\(r\)c\_\{j,0\}^\{\(r\)\},\\ldots,c\_\{j,G\}^\{\(r\)\}\. Using the constant extensions from \([15](https://arxiv.org/html/2608.13882#S2.E15)\) and \([16](https://arxiv.org/html/2608.13882#S2.E16)\), the profile has the piecewise representation
qj\(r\)\(t\)\\displaystyle q\_\{j\}^\{\(r\)\}\\left\(t\\right\)=\{cj,0\(r\),t≤t0,tℓ\+1−ttℓ\+1−tℓcj,ℓ\(r\)\+t−tℓtℓ\+1−tℓcj,ℓ\+1\(r\),t∈\[tℓ,tℓ\+1\],ℓ∈\{0,…,G−1\},cj,G\(r\),t≥tG\.\\displaystyle=\\begin\{cases\}c\_\{j,0\}^\{\(r\)\},&t\\leq t\_\{0\},\\\\\[5\.69054pt\] \\displaystyle\\frac\{t\_\{\\ell\+1\}\-t\}\{t\_\{\\ell\+1\}\-t\_\{\\ell\}\}c\_\{j,\\ell\}^\{\(r\)\}\+\\frac\{t\-t\_\{\\ell\}\}\{t\_\{\\ell\+1\}\-t\_\{\\ell\}\}c\_\{j,\\ell\+1\}^\{\(r\)\},&t\\in\\left\[t\_\{\\ell\},t\_\{\\ell\+1\}\\right\],\\quad\\ell\\in\\left\\\{0,\\ldots,G\-1\\right\\\},\\\\\[8\.53581pt\] c\_\{j,G\}^\{\(r\)\},&t\\geq t\_\{G\}\.\\end\{cases\}\(367\)At a grid knot, the two adjacent affine formulas have the same value\. Therefore, any fixed convention for assigning a knot to one of its adjacent intervals yields the same profile value\. Fori∈\[N\]i\\in\\left\[N\\right\],j∈\[m\]j\\in\\left\[m\\right\], andr∈\[K\]r\\in\\left\[K\\right\], define the input to therr\-th profile by
vi,j\(1\)\\displaystyle v\_\{i,j\}^\{\(1\)\}:=zj\(0\)\(𝐱i\),\\displaystyle:=z\_\{j\}^\{\(0\)\}\\left\(\\mathbf\{x\}\_\{i\}\\right\),\(368\)vi,j\(r\)\\displaystyle v\_\{i,j\}^\{\(r\)\}:=sj\(r\)\(𝐱i\),r∈\{2,…,K\}\.\\displaystyle:=s\_\{j\}^\{\(r\)\}\\left\(\\mathbf\{x\}\_\{i\}\\right\),\\qquad r\\in\\left\\\{2,\\ldots,K\\right\\\}\.\(369\)The active region ofqj\(r\)q\_\{j\}^\{\(r\)\}is determined by the signs of theG\+1G\+1quantities
pi,j,ℓ\(r\)\(𝜼\):=vi,j\(r\)−tℓ,ℓ∈\{0,…,G\}\.\\displaystyle p\_\{i,j,\\ell\}^\{\(r\)\}\\left\(\\bm\{\\eta\}\\right\):=v\_\{i,j\}^\{\(r\)\}\-t\_\{\\ell\},\\qquad\\ell\\in\\left\\\{0,\\ldots,G\\right\\\}\.\(370\)We now prove by induction that, after fixing all profile regions up to levelrr, every valuezj\(r\)\(𝐱i\)z\_\{j\}^\{\(r\)\}\\left\(\\mathbf\{x\}\_\{i\}\\right\)is a polynomial in𝜼\\bm\{\\eta\}of degree at most2r2r\. At the first profile level,
vi,j\(1\)\\displaystyle v\_\{i,j\}^\{\(1\)\}=\(a\)zj\(0\)\(𝐱i\)=\(b\)𝝎j⊤𝐱i\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}z\_\{j\}^\{\(0\)\}\\left\(\\mathbf\{x\}\_\{i\}\\right\)\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\bm\{\\omega\}\_\{j\}^\{\\top\}\\mathbf\{x\}\_\{i\}\.\(371\)Here, \(a\) applies \([368](https://arxiv.org/html/2608.13882#A4.E368)\), while \(b\) applies \([27](https://arxiv.org/html/2608.13882#S2.E27)\)\. Since the sample point𝐱i\\mathbf\{x\}\_\{i\}is fixed, \([371](https://arxiv.org/html/2608.13882#A4.E371)\) is a polynomial of degree one in the architecture parameters\. Once the active region is fixed, \([367](https://arxiv.org/html/2608.13882#A4.E367)\) shows that
zj\(1\)\(𝐱i\)\\displaystyle z\_\{j\}^\{\(1\)\}\\left\(\\mathbf\{x\}\_\{i\}\\right\)=qj\(1\)\(vi,j\(1\)\)\\displaystyle=q\_\{j\}^\{\(1\)\}\\left\(v\_\{i,j\}^\{\(1\)\}\\right\)\(372\)is a polynomial of degree at most two\. Indeed, on an interior interval, \([367](https://arxiv.org/html/2608.13882#A4.E367)\) can be rewritten as
qj\(1\)\(vi,j\(1\)\)\\displaystyle q\_\{j\}^\{\(1\)\}\\left\(v\_\{i,j\}^\{\(1\)\}\\right\)=tℓ\+1cj,ℓ\(1\)−tℓcj,ℓ\+1\(1\)tℓ\+1−tℓ\+cj,ℓ\+1\(1\)−cj,ℓ\(1\)tℓ\+1−tℓvi,j\(1\)\.\\displaystyle=\\frac\{t\_\{\\ell\+1\}c\_\{j,\\ell\}^\{\(1\)\}\-t\_\{\\ell\}c\_\{j,\\ell\+1\}^\{\(1\)\}\}\{t\_\{\\ell\+1\}\-t\_\{\\ell\}\}\+\\frac\{c\_\{j,\\ell\+1\}^\{\(1\)\}\-c\_\{j,\\ell\}^\{\(1\)\}\}\{t\_\{\\ell\+1\}\-t\_\{\\ell\}\}v\_\{i,j\}^\{\(1\)\}\.\(373\)The first term in \([373](https://arxiv.org/html/2608.13882#A4.E373)\) has degree one in the nodal parameters\. The second term is the product of a degree\-one nodal expression and the degree\-one polynomialvi,j\(1\)v\_\{i,j\}^\{\(1\)\}, and therefore has degree at most two\. On either exterior region, the profile value is a single nodal parameter and has degree one\. This proves the degree bound atr=1r=1\. Fixr∈\{2,…,K\}r\\in\\left\\\{2,\\ldots,K\\right\\\}and assume that, after fixing the profile regions up to levelr−1r\-1, everyzk\(r−1\)\(𝐱i\)z\_\{k\}^\{\(r\-1\)\}\\left\(\\mathbf\{x\}\_\{i\}\\right\), withi∈\[N\]i\\in\\left\[N\\right\]andk∈\[m\]k\\in\\left\[m\\right\], is a polynomial in𝜼\\bm\{\\eta\}of degree at most2\(r−1\)2\\left\(r\-1\\right\)\. Then
vi,j\(r\)\\displaystyle v\_\{i,j\}^\{\(r\)\}=\(a\)sj\(r\)\(𝐱i\)=\(b\)∑k=1mWjk\(r\)zk\(r−1\)\(𝐱i\)\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}s\_\{j\}^\{\(r\)\}\\left\(\\mathbf\{x\}\_\{i\}\\right\)\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\sum\_\{k=1\}^\{m\}W\_\{jk\}^\{\(r\)\}z\_\{k\}^\{\(r\-1\)\}\\left\(\\mathbf\{x\}\_\{i\}\\right\)\.\(374\)Here, \(a\) applies \([369](https://arxiv.org/html/2608.13882#A4.E369)\), and \(b\) applies \([29](https://arxiv.org/html/2608.13882#S2.E29)\)\. Each mixing coefficientWjk\(r\)W\_\{jk\}^\{\(r\)\}has polynomial degree one in𝜼\\bm\{\\eta\}\. The induction hypothesis therefore implies that every summand in \([374](https://arxiv.org/html/2608.13882#A4.E374)\) has degree at most1\+2\(r−1\)=2r−11\+2\\left\(r\-1\\right\)=2r\-1\. Consequently,
deg\(vi,j\(r\)\)≤2r−1\.\\displaystyle\\deg\\left\(v\_\{i,j\}^\{\(r\)\}\\right\)\\leq 2r\-1\.\(375\)Once the active region at levelrris fixed, the same calculation as in \([373](https://arxiv.org/html/2608.13882#A4.E373)\) gives
deg\(zj\(r\)\(𝐱i\)\)\\displaystyle\\deg\\left\(z\_\{j\}^\{\(r\)\}\\left\(\\mathbf\{x\}\_\{i\}\\right\)\\right\)≤\(a\)1\+deg\(vi,j\(r\)\)≤\(b\)2r\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}1\+\\deg\\left\(v\_\{i,j\}^\{\(r\)\}\\right\)\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}2r\.\(376\)Here, \(a\) follows because an interior profile branch multiplies one nodal coefficient by the profile input, while \(b\) applies \([375](https://arxiv.org/html/2608.13882#A4.E375)\)\. This completes the induction\. We next count the possible profile\-region states\. Let
𝒫0:=\{ℝW\}\.\\displaystyle\\mathcal\{P\}\_\{0\}:=\\left\\\{\\mathbb\{R\}^\{W\}\\right\\\}\.\(377\)Forr∈\[K\]r\\in\\left\[K\\right\], let𝒫r\\mathcal\{P\}\_\{r\}denote the refinement of𝒫r−1\\mathcal\{P\}\_\{r\-1\}obtained by fixing the signs of all polynomials \([370](https://arxiv.org/html/2608.13882#A4.E370)\) at therr\-th profile level\.
For every fixed member of𝒫r−1\\mathcal\{P\}\_\{r\-1\}, there is one region polynomial for every sample point, every unit, and every grid knot\. Hence, the number of region polynomials at levelrris
Mr\\displaystyle M\_\{r\}=Nm\(G\+1\)\.\\displaystyle=Nm\\left\(G\+1\\right\)\.\(378\)Their degrees are bounded by2r−1≤2K\+12r\-1\\leq 2K\+1\. DefineDL:=2K\+1=2L−3D\_\{L\}:=2K\+1=2L\-3\. WhenL≥3L\\geq 3, the parameter count \([197](https://arxiv.org/html/2608.13882#A4.E197)\) contains the profile contributionKm\(G\+1\)Km\\left\(G\+1\\right\)\. SinceK≥1K\\geq 1, one obtains
m\(G\+1\)≤P\.\\displaystyle m\\left\(G\+1\\right\)\\leq P\.\(379\)Moreover, sinceP≥1P\\geq 1,
W=P\+1≤2P\.\\displaystyle W=P\+1\\leq 2P\.\(380\)Therefore,
Mr\+W\\displaystyle M\_\{r\}\+W≤\(a\)NP\+2P≤\(b\)3NP\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}NP\+2P\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}3NP\.\(381\)Here, \(a\) uses \([378](https://arxiv.org/html/2608.13882#A4.E378)\), \([379](https://arxiv.org/html/2608.13882#A4.E379)\), and \([380](https://arxiv.org/html/2608.13882#A4.E380)\), while \(b\) usesN≥1N\\geq 1\. Applying \([363](https://arxiv.org/html/2608.13882#A4.E363)\) inside each member of𝒫r−1\\mathcal\{P\}\_\{r\-1\}gives
\|𝒫r\|\\displaystyle\\left\|\\mathcal\{P\}\_\{r\}\\right\|≤\(a\)\|𝒫r−1\|\(C0DL\(Mr\+W\)\)W≤\(b\)\|𝒫r−1\|\(3C0DLNP\)W\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}\\left\|\\mathcal\{P\}\_\{r\-1\}\\right\|\\left\(C\_\{0\}D\_\{L\}\\left\(M\_\{r\}\+W\\right\)\\right\)^\{W\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}\\left\|\\mathcal\{P\}\_\{r\-1\}\\right\|\\left\(3C\_\{0\}D\_\{L\}NP\\right\)^\{W\}\.\(382\)Here, \(a\) applies the polynomial sign\-pattern estimate to the region polynomials on each preceding state, and \(b\) applies \([381](https://arxiv.org/html/2608.13882#A4.E381)\)\. Set
CL,0:=3C0DL\.\\displaystyle C\_\{L,0\}:=3C\_\{0\}D\_\{L\}\.\(383\)Iterating \([382](https://arxiv.org/html/2608.13882#A4.E382)\) fromr=1r=1tor=Kr=Kand using \([377](https://arxiv.org/html/2608.13882#A4.E377)\) gives
\|𝒫K\|\\displaystyle\\left\|\\mathcal\{P\}\_\{K\}\\right\|≤\(CL,0NP\)KW\.\\displaystyle\\leq\\left\(C\_\{L,0\}NP\\right\)^\{KW\}\.\(384\)WhenK=0K=0, the right\-hand side equals one and \([384](https://arxiv.org/html/2608.13882#A4.E384)\) reduces to\|𝒫0\|=1\\left\|\\mathcal\{P\}\_\{0\}\\right\|=1\. On every member of𝒫K\\mathcal\{P\}\_\{K\}, the architecture output is a polynomial in𝜼\\bm\{\\eta\}\. ForL≥3L\\geq 3, \([31](https://arxiv.org/html/2608.13882#S2.E31)\) gives
a𝜽\(𝐱i\)\\displaystyle a\_\{\\bm\{\\theta\}\}\\left\(\\mathbf\{x\}\_\{i\}\\right\)=∑j=1mβjzj\(K\)\(𝐱i\)\.\\displaystyle=\\sum\_\{j=1\}^\{m\}\\beta\_\{j\}z\_\{j\}^\{\(K\)\}\\left\(\\mathbf\{x\}\_\{i\}\\right\)\.\(385\)Since the readout coefficients have degree one and the final hidden coordinates have degree at most2K2K,
deg\(a𝜽\(𝐱i\)\)\\displaystyle\\deg\\left\(a\_\{\\bm\{\\theta\}\}\\left\(\\mathbf\{x\}\_\{i\}\\right\)\\right\)≤1\+2K=DL\.\\displaystyle\\leq 1\+2K=D\_\{L\}\.\(386\)ForL=2L=2, the same estimate follows from \([366](https://arxiv.org/html/2608.13882#A4.E366)\) andDL=1D\_\{L\}=1\. Fixσ∈\{−1,1\}\\sigma\\in\\left\\\{\-1,1\\right\\\}\. On every member of𝒫K\\mathcal\{P\}\_\{K\}, the labels of theNNsample points are determined by the signs of theNNpolynomials
ri,σ\(𝜼\):=σa𝜽\(𝐱i\)−τ,i∈\[N\]\.\\displaystyle r\_\{i,\\sigma\}\\left\(\\bm\{\\eta\}\\right\):=\\sigma a\_\{\\bm\{\\theta\}\}\\left\(\\mathbf\{x\}\_\{i\}\\right\)\-\\tau,\\qquad i\\in\\left\[N\\right\]\.\(387\)By \([386](https://arxiv.org/html/2608.13882#A4.E386)\), these polynomials have degree at mostDLD\_\{L\}\. Furthermore,
N\+W\\displaystyle N\+W≤\(a\)N\+2P≤\(b\)3NP\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}N\+2P\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}3NP\.\(388\)Here, \(a\) applies \([380](https://arxiv.org/html/2608.13882#A4.E380)\), and \(b\) usesN≥1N\\geq 1andP≥1P\\geq 1\. Thus, \([363](https://arxiv.org/html/2608.13882#A4.E363)\) implies that the number of label vectors generated on each member of𝒫K\\mathcal\{P\}\_\{K\}for one fixed signσ\\sigmais at most
\(C0DL\(N\+W\)\)W\\displaystyle\\left\(C\_\{0\}D\_\{L\}\\left\(N\+W\\right\)\\right\)^\{W\}≤\(CL,0NP\)W,\\displaystyle\\leq\\left\(C\_\{L,0\}NP\\right\)^\{W\},\(389\)where the inequality follows from \([388](https://arxiv.org/html/2608.13882#A4.E388)\) and \([383](https://arxiv.org/html/2608.13882#A4.E383)\)\. There are two possible choices ofσ\\sigma\. Combining \([384](https://arxiv.org/html/2608.13882#A4.E384)\) and \([389](https://arxiv.org/html/2608.13882#A4.E389)\) therefore yields
Πℋ~L−1,m,G±\(N\)\\displaystyle\\Pi\_\{\\widetilde\{\\mathcal\{H\}\}\_\{L\-1,m,G\}^\{\\pm\}\}\\left\(N\\right\)≤\(a\)2\|𝒫K\|\(CL,0NP\)W≤\(b\)2\(CL,0NP\)\(K\+1\)W=\(c\)2\(CL,0NP\)\(L−1\)W\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}2\\left\|\\mathcal\{P\}\_\{K\}\\right\|\\left\(C\_\{L,0\}NP\\right\)^\{W\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}2\\left\(C\_\{L,0\}NP\\right\)^\{\\left\(K\+1\\right\)W\}\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{=\}\}2\\left\(C\_\{L,0\}NP\\right\)^\{\\left\(L\-1\\right\)W\}\.\(390\)Here, \(a\) sums the bounds corresponding to the two signs, \(b\) substitutes \([384](https://arxiv.org/html/2608.13882#A4.E384)\), and \(c\) usesK=L−2K=L\-2\. This proves \([365](https://arxiv.org/html/2608.13882#A4.E365)\) after relabeling the depth\-dependent constant\.
Step 4: inversion of the growth bound\.Suppose thatℋ~L−1,m,G±\\widetilde\{\\mathcal\{H\}\}\_\{L\-1,m,G\}^\{\\pm\}shattersNNpoints\. Then
2N\\displaystyle 2^\{N\}≤\(a\)Πℋ~L−1,m,G±\(N\)≤\(b\)2\(CL,0NP\)\(L−1\)W\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}\\Pi\_\{\\widetilde\{\\mathcal\{H\}\}\_\{L\-1,m,G\}^\{\\pm\}\}\\left\(N\\right\)\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}2\\left\(C\_\{L,0\}NP\\right\)^\{\\left\(L\-1\\right\)W\}\.\(391\)Here, \(a\) follows from the definition of shattering, and \(b\) applies \([390](https://arxiv.org/html/2608.13882#A4.E390)\)\. Taking natural logarithms gives
Nln2≤ln2\+\(L−1\)Wln\(CL,0NP\)\.\\displaystyle N\\ln 2\\leq\\ln 2\+\\left\(L\-1\\right\)W\\ln\\left\(C\_\{L,0\}NP\\right\)\.\(392\)SetcL:=L−1c\_\{L\}:=L\-1andr:=N/Wr:=N/W\. SinceP≤WP\\leq WandN=rWN=rW,
ln\(CL,0NP\)\\displaystyle\\ln\\left\(C\_\{L,0\}NP\\right\)=\(a\)ln\(CL,0rWP\)≤\(b\)lnCL,0\+lnr\+2lnW\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\ln\\left\(C\_\{L,0\}rWP\\right\)\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}\\ln C\_\{L,0\}\+\\ln r\+2\\ln W\.\(393\)Here, \(a\) substitutesN=rWN=rW, and \(b\) usesP≤WP\\leq W\. Dividing \([392](https://arxiv.org/html/2608.13882#A4.E392)\) byWWand applying \([393](https://arxiv.org/html/2608.13882#A4.E393)\) gives
rln2\\displaystyle r\\ln 2≤\(a\)ln2W\+cL\(lnCL,0\+lnr\+2lnW\)≤\(b\)ln2\+cL\(lnCL,0\+lnr\+2lnW\)\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}\\frac\{\\ln 2\}\{W\}\+c\_\{L\}\\left\(\\ln C\_\{L,0\}\+\\ln r\+2\\ln W\\right\)\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}\\ln 2\+c\_\{L\}\\left\(\\ln C\_\{L,0\}\+\\ln r\+2\\ln W\\right\)\.\(394\)Here, \(a\) performs the division and substitution, while \(b\) usesW≥1W\\geq 1\. Define
ηL:=ln22cL\.\\displaystyle\\eta\_\{L\}:=\\frac\{\\ln 2\}\{2c\_\{L\}\}\.\(395\)The elementary inequalitylny≤y−1\\ln y\\leq y\-1fory\>0y\>0, applied withy=ηLry=\\eta\_\{L\}r, gives
lnr\\displaystyle\\ln r=\(a\)ln\(ηLr\)−lnηL≤\(b\)ηLr−1−lnηL\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\ln\\left\(\\eta\_\{L\}r\\right\)\-\\ln\\eta\_\{L\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}\\eta\_\{L\}r\-1\-\\ln\\eta\_\{L\}\.\(396\)Here, \(a\) uses the logarithm of a product, while \(b\) applieslny≤y−1\\ln y\\leq y\-1\. Multiplying \([396](https://arxiv.org/html/2608.13882#A4.E396)\) bycLc\_\{L\}and using \([395](https://arxiv.org/html/2608.13882#A4.E395)\) gives
cLlnr\\displaystyle c\_\{L\}\\ln r≤rln22\+cL\(−1−lnηL\)\.\\displaystyle\\leq\\frac\{r\\ln 2\}\{2\}\+c\_\{L\}\\left\(\-1\-\\ln\\eta\_\{L\}\\right\)\.\(397\)Substituting \([397](https://arxiv.org/html/2608.13882#A4.E397)\) into \([394](https://arxiv.org/html/2608.13882#A4.E394)\) and movingrln2/2r\\ln 2/2to the left\-hand side yields
rln22\\displaystyle\\frac\{r\\ln 2\}\{2\}≤ln2\+cL\(lnCL,0−1−lnηL\+2lnW\)\.\\displaystyle\\leq\\ln 2\+c\_\{L\}\\left\(\\ln C\_\{L,0\}\-1\-\\ln\\eta\_\{L\}\+2\\ln W\\right\)\.\(398\)SincecLc\_\{L\},CL,0C\_\{L,0\}, andηL\\eta\_\{L\}depend only onLL, there exists a constantCL,2\>0C\_\{L,2\}\>0, depending only onLL, such that
r≤CL,2ln\(eW\)\.\\displaystyle r\\leq C\_\{L,2\}\\ln\\left\(eW\\right\)\.\(399\)Multiplying \([399](https://arxiv.org/html/2608.13882#A4.E399)\) byWWand usingN=rWN=rWgives
N≤CL,2Wln\(eW\)\.\\displaystyle N\\leq C\_\{L,2\}W\\ln\\left\(eW\\right\)\.\(400\)Since every shattered set satisfies \([400](https://arxiv.org/html/2608.13882#A4.E400)\),
VC\(ℋ~L−1,m,G±\)\\displaystyle\\operatorname\{VC\}\\left\(\\widetilde\{\\mathcal\{H\}\}\_\{L\-1,m,G\}^\{\\pm\}\\right\)≤\(a\)CL,2Wln\(eW\)=\(b\)CL,2\(P\+1\)ln\(e\(P\+1\)\)\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}C\_\{L,2\}W\\ln\\left\(eW\\right\)\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}C\_\{L,2\}\\left\(P\+1\\right\)\\ln\\left\(e\\left\(P\+1\\right\)\\right\)\.\(401\)Here, \(a\) applies the definition of the VC dimension, while \(b\) substitutesW=P\+1W=P\+1\. Finally, \([361](https://arxiv.org/html/2608.13882#A4.E361)\) and \([401](https://arxiv.org/html/2608.13882#A4.E401)\) give
VC\(ℋL−1,m,G±\)≤CL\(P\+1\)ln\(e\(P\+1\)\),\\displaystyle\\operatorname\{VC\}\\left\(\\mathcal\{H\}\_\{L\-1,m,G\}^\{\\pm\}\\right\)\\leq C\_\{L\}\\left\(P\+1\\right\)\\ln\\left\(e\\left\(P\+1\\right\)\\right\),\(402\)after relabeling the depth\-dependent constant\. This proves \([359](https://arxiv.org/html/2608.13882#A4.E359)\)\. SubstitutingP=PL−1,m,GP=P\_\{L\-1,m,G\}proves \([360](https://arxiv.org/html/2608.13882#A4.E360)\) and completes the proof\. ∎
###### Lemma B0\.
\(Signed threshold entropy of finite lower\-support architectures\)FixL≥2L\\geq 2,m≥1m\\geq 1,G≥2G\\geq 2,n≥1n\\geq 1, and sample points𝐱1,…,𝐱n∈𝒳\\mathbf\{x\}\_\{1\},\\ldots,\\mathbf\{x\}\_\{n\}\\in\\mathcal\{X\}\. Let
vL−1,m,G:=VC\(ℋL−1,m,G±\),\\displaystyle v\_\{L\-1,m,G\}:=\\operatorname\{VC\}\\left\(\\mathcal\{H\}\_\{L\-1,m,G\}^\{\\pm\}\\right\),\(403\)whereℋL−1,m,G±\\mathcal\{H\}\_\{L\-1,m,G\}^\{\\pm\}is defined in \([357](https://arxiv.org/html/2608.13882#A4.E357)\)\. The signed trace family from \([335](https://arxiv.org/html/2608.13882#A4.E335)\) satisfies
𝒯n,𝐱±\(𝒜L−1m,G\)=\{\{i∈\[n\]:h\(𝐱i\)=1\}:h∈ℋL−1,m,G±\}\.\\displaystyle\\mathcal\{T\}\_\{n,\\mathbf\{x\}\}^\{\\pm\}\\left\(\\mathcal\{A\}\_\{L\-1\}^\{m,G\}\\right\)=\\left\\\{\\left\\\{i\\in\\left\[n\\right\]:h\\left\(\\mathbf\{x\}\_\{i\}\\right\)=1\\right\\\}:h\\in\\mathcal\{H\}\_\{L\-1,m,G\}^\{\\pm\}\\right\\\}\.\(404\)IfvL−1,m,G=0v\_\{L\-1,m,G\}=0, then
\|𝒯n,𝐱±\(𝒜L−1m,G\)\|=1\.\\displaystyle\\left\|\\mathcal\{T\}\_\{n,\\mathbf\{x\}\}^\{\\pm\}\\left\(\\mathcal\{A\}\_\{L\-1\}^\{m,G\}\\right\)\\right\|=1\.\(405\)IfvL−1,m,G≥1v\_\{L\-1,m,G\}\\geq 1, then
ln\|𝒯n,𝐱±\(𝒜L−1m,G\)\|\\displaystyle\\ln\\left\|\\mathcal\{T\}\_\{n,\\mathbf\{x\}\}^\{\\pm\}\\left\(\\mathcal\{A\}\_\{L\-1\}^\{m,G\}\\right\)\\right\|≤\(vL−1,m,G∧n\)ln\(envL−1,m,G∧n\)\.\\displaystyle\\leq\\left\(v\_\{L\-1,m,G\}\\wedge n\\right\)\\ln\\left\(\\frac\{en\}\{v\_\{L\-1,m,G\}\\wedge n\}\\right\)\.\(406\)Consequently, there exists a constantCL\>0C\_\{L\}\>0, depending only onLL, such that
ln\|𝒯n,𝐱±\(𝒜L−1m,G\)\|\\displaystyle\\ln\\left\|\\mathcal\{T\}\_\{n,\\mathbf\{x\}\}^\{\\pm\}\\left\(\\mathcal\{A\}\_\{L\-1\}^\{m,G\}\\right\)\\right\|≤CL\(PL−1,m,G\+1\)ln\(e\(PL−1,m,G\+1\)\)ln\(en\)\.\\displaystyle\\leq C\_\{L\}\\left\(P\_\{L\-1,m,G\}\+1\\right\)\\ln\\left\(e\\left\(P\_\{L\-1,m,G\}\+1\\right\)\\right\)\\ln\\left\(en\\right\)\.\(407\)
###### Proof\.
For notational convenience, set
𝒜\\displaystyle\\mathcal\{A\}:=𝒜L−1m,G,\\displaystyle:=\\mathcal\{A\}\_\{L\-1\}^\{m,G\},ℋ±\\displaystyle\\mathcal\{H\}^\{\\pm\}:=ℋL−1,m,G±,\\displaystyle:=\\mathcal\{H\}\_\{L\-1,m,G\}^\{\\pm\},v\\displaystyle v:=vL−1,m,G\.\\displaystyle:=v\_\{L\-1,m,G\}\.\(408\)We first prove the exact trace identity \([404](https://arxiv.org/html/2608.13882#A4.E404)\)\. LetT∈𝒯n,𝐱±\(𝒜\)T\\in\\mathcal\{T\}\_\{n,\\mathbf\{x\}\}^\{\\pm\}\\left\(\\mathcal\{A\}\\right\)\. By \([335](https://arxiv.org/html/2608.13882#A4.E335)\), there exista𝜽∈𝒜a\_\{\\bm\{\\theta\}\}\\in\\mathcal\{A\},σ∈\{−1,1\}\\sigma\\in\\left\\\{\-1,1\\right\\\}, andτ∈ℝ\\tau\\in\\mathbb\{R\}such that
T=\{i∈\[n\]:σa𝜽\(𝐱i\)≥τ\}\.\\displaystyle T=\\left\\\{i\\in\\left\[n\\right\]:\\sigma a\_\{\\bm\{\\theta\}\}\\left\(\\mathbf\{x\}\_\{i\}\\right\)\\geq\\tau\\right\\\}\.\(409\)Defineh:=h𝜽,τ,σh:=h\_\{\\bm\{\\theta\},\\tau,\\sigma\}\. By \([357](https://arxiv.org/html/2608.13882#A4.E357)\), one hash∈ℋ±h\\in\\mathcal\{H\}^\{\\pm\}\. Moreover,
\{i∈\[n\]:h\(𝐱i\)=1\}\\displaystyle\\left\\\{i\\in\\left\[n\\right\]:h\\left\(\\mathbf\{x\}\_\{i\}\\right\)=1\\right\\\}=\(a\)\{i∈\[n\]:𝟏\{σa𝜽\(𝐱i\)≥τ\}=1\}=\(b\)\{i∈\[n\]:σa𝜽\(𝐱i\)≥τ\}=\(c\)T\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\left\\\{i\\in\\left\[n\\right\]:\\mathbf\{1\}\_\{\\left\\\{\\sigma a\_\{\\bm\{\\theta\}\}\\left\(\\mathbf\{x\}\_\{i\}\\right\)\\geq\\tau\\right\\\}\}=1\\right\\\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\left\\\{i\\in\\left\[n\\right\]:\\sigma a\_\{\\bm\{\\theta\}\}\\left\(\\mathbf\{x\}\_\{i\}\\right\)\\geq\\tau\\right\\\}\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{=\}\}T\.\(410\)Here, \(a\) applies \([358](https://arxiv.org/html/2608.13882#A4.E358)\), \(b\) applies the defining property of an indicator function, and \(c\) invokes \([409](https://arxiv.org/html/2608.13882#A4.E409)\)\. Therefore,
𝒯n,𝐱±\(𝒜\)⊆\{\{i∈\[n\]:h\(𝐱i\)=1\}:h∈ℋ±\}\.\\displaystyle\\mathcal\{T\}\_\{n,\\mathbf\{x\}\}^\{\\pm\}\\left\(\\mathcal\{A\}\\right\)\\subseteq\\left\\\{\\left\\\{i\\in\\left\[n\\right\]:h\\left\(\\mathbf\{x\}\_\{i\}\\right\)=1\\right\\\}:h\\in\\mathcal\{H\}^\{\\pm\}\\right\\\}\.\(411\)Conversely, leth∈ℋ±h\\in\\mathcal\{H\}^\{\\pm\}\. By \([357](https://arxiv.org/html/2608.13882#A4.E357)\), there exist𝜽∈ΘL−1,m,G\\bm\{\\theta\}\\in\\Theta\_\{L\-1,m,G\},τ∈ℝ\\tau\\in\\mathbb\{R\}, andσ∈\{−1,1\}\\sigma\\in\\left\\\{\-1,1\\right\\\}such thath=h𝜽,τ,σh=h\_\{\\bm\{\\theta\},\\tau,\\sigma\}\. Therefore,
\{i∈\[n\]:h\(𝐱i\)=1\}\\displaystyle\\left\\\{i\\in\\left\[n\\right\]:h\\left\(\\mathbf\{x\}\_\{i\}\\right\)=1\\right\\\}=\(a\)\{i∈\[n\]:σa𝜽\(𝐱i\)≥τ\}∈\(b\)𝒯n,𝐱±\(𝒜\)\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\left\\\{i\\in\\left\[n\\right\]:\\sigma a\_\{\\bm\{\\theta\}\}\\left\(\\mathbf\{x\}\_\{i\}\\right\)\\geq\\tau\\right\\\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\in\}\}\\mathcal\{T\}\_\{n,\\mathbf\{x\}\}^\{\\pm\}\\left\(\\mathcal\{A\}\\right\)\.\(412\)Here, \(a\) applies \([358](https://arxiv.org/html/2608.13882#A4.E358)\), and \(b\) applies \([335](https://arxiv.org/html/2608.13882#A4.E335)\)\. Hence,
\{\{i∈\[n\]:h\(𝐱i\)=1\}:h∈ℋ±\}⊆𝒯n,𝐱±\(𝒜\)\.\\displaystyle\\left\\\{\\left\\\{i\\in\\left\[n\\right\]:h\\left\(\\mathbf\{x\}\_\{i\}\\right\)=1\\right\\\}:h\\in\\mathcal\{H\}^\{\\pm\}\\right\\\}\\subseteq\\mathcal\{T\}\_\{n,\\mathbf\{x\}\}^\{\\pm\}\\left\(\\mathcal\{A\}\\right\)\.\(413\)Combining \([411](https://arxiv.org/html/2608.13882#A4.E411)\) and \([413](https://arxiv.org/html/2608.13882#A4.E413)\) proves \([404](https://arxiv.org/html/2608.13882#A4.E404)\)\. We now count the traces\. By \([404](https://arxiv.org/html/2608.13882#A4.E404)\), identifying each subset of\[n\]\\left\[n\\right\]with its binary indicator vector shows that the cardinality of the signed trace family is exactly the number of binary label vectors induced byℋ±\\mathcal\{H\}^\{\\pm\}on the fixed sample:
\|𝒯n,𝐱±\(𝒜\)\|\\displaystyle\\left\|\\mathcal\{T\}\_\{n,\\mathbf\{x\}\}^\{\\pm\}\\left\(\\mathcal\{A\}\\right\)\\right\|=\|\{\(h\(𝐱1\),…,h\(𝐱n\)\):h∈ℋ±\}\|\.\\displaystyle=\\left\|\\left\\\{\\left\(h\\left\(\\mathbf\{x\}\_\{1\}\\right\),\\ldots,h\\left\(\\mathbf\{x\}\_\{n\}\\right\)\\right\):h\\in\\mathcal\{H\}^\{\\pm\}\\right\\\}\\right\|\.\(414\)Suppose first thatv=0v=0\. If two members ofℋ±\\mathcal\{H\}^\{\\pm\}gave different labels at some𝐱∈𝒳\\mathbf\{x\}\\in\\mathcal\{X\}, then the singleton\{𝐱\}\\left\\\{\\mathbf\{x\}\\right\\\}would be shattered\. This would implyv≥1v\\geq 1, contradictingv=0v=0\. Thus, every member ofℋ±\\mathcal\{H\}^\{\\pm\}has the same value at every point of𝒳\\mathcal\{X\}\. Consequently, only one label vector occurs on the sample, and\|𝒯n,𝐱±\(𝒜\)\|=1\\left\|\\mathcal\{T\}\_\{n,\\mathbf\{x\}\}^\{\\pm\}\\left\(\\mathcal\{A\}\\right\)\\right\|=1\. This proves \([405](https://arxiv.org/html/2608.13882#A4.E405)\)\. Assume now thatv≥1v\\geq 1\. We distinguish two cases\. Ifv<nv<n, the Sauer–Shelah lemma and \([414](https://arxiv.org/html/2608.13882#A4.E414)\) give
\|𝒯n,𝐱±\(𝒜\)\|\\displaystyle\\left\|\\mathcal\{T\}\_\{n,\\mathbf\{x\}\}^\{\\pm\}\\left\(\\mathcal\{A\}\\right\)\\right\|≤\(a\)∑j=0v\(nj\)≤\(b\)\(env\)v\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}\\sum\_\{j=0\}^\{v\}\\binom\{n\}\{j\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}\\left\(\\frac\{en\}\{v\}\\right\)^\{v\}\.\(415\)Here, \(a\) applies the Sauer–Shelah lemma to a class of VC dimensionvv, while \(b\) applies the standard binomial\-sum estimate for1≤v<n1\\leq v<n\. Ifv≥nv\\geq n, every trace is a subset of\[n\]\\left\[n\\right\], and therefore
\|𝒯n,𝐱±\(𝒜\)\|\\displaystyle\\left\|\\mathcal\{T\}\_\{n,\\mathbf\{x\}\}^\{\\pm\}\\left\(\\mathcal\{A\}\\right\)\\right\|≤\(a\)2n≤\(b\)en=\(c\)\(enn\)n\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}2^\{n\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}e^\{n\}\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{=\}\}\\left\(\\frac\{en\}\{n\}\\right\)^\{n\}\.\(416\)Here, \(a\) counts all subsets of\[n\]\\left\[n\\right\], \(b\) uses2≤e2\\leq e, and \(c\) simplifies the fraction\. Define
q:=v∧n\.\\displaystyle q:=v\\wedge n\.\(417\)Both \([415](https://arxiv.org/html/2608.13882#A4.E415)\) and \([416](https://arxiv.org/html/2608.13882#A4.E416)\) can then be written as
\|𝒯n,𝐱±\(𝒜\)\|≤\(enq\)q\.\\displaystyle\\left\|\\mathcal\{T\}\_\{n,\\mathbf\{x\}\}^\{\\pm\}\\left\(\\mathcal\{A\}\\right\)\\right\|\\leq\\left\(\\frac\{en\}\{q\}\\right\)^\{q\}\.\(418\)Taking natural logarithms in \([418](https://arxiv.org/html/2608.13882#A4.E418)\) gives
ln\|𝒯n,𝐱±\(𝒜\)\|\\displaystyle\\ln\\left\|\\mathcal\{T\}\_\{n,\\mathbf\{x\}\}^\{\\pm\}\\left\(\\mathcal\{A\}\\right\)\\right\|≤\(a\)qln\(enq\)=\(b\)\(v∧n\)ln\(env∧n\)\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}q\\ln\\left\(\\frac\{en\}\{q\}\\right\)\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\left\(v\\wedge n\\right\)\\ln\\left\(\\frac\{en\}\{v\\wedge n\}\\right\)\.\(419\)Here, \(a\) takes logarithms, and \(b\) substitutes \([417](https://arxiv.org/html/2608.13882#A4.E417)\)\. This proves \([406](https://arxiv.org/html/2608.13882#A4.E406)\)\. It remains to derive the simpler parameter\-dependent estimate\. Sinceq≥1q\\geq 1, the monotonicity of the logarithm gives
ln\(enq\)\\displaystyle\\ln\\left\(\\frac\{en\}\{q\}\\right\)≤ln\(en\)\.\\displaystyle\\leq\\ln\\left\(en\\right\)\.\(420\)Moreover,
q=v∧n≤v\.\\displaystyle q=v\\wedge n\\leq v\.\(421\)Combining \([419](https://arxiv.org/html/2608.13882#A4.E419)\), \([420](https://arxiv.org/html/2608.13882#A4.E420)\), and \([421](https://arxiv.org/html/2608.13882#A4.E421)\) gives
ln\|𝒯n,𝐱±\(𝒜\)\|≤vln\(en\)\.\\displaystyle\\ln\\left\|\\mathcal\{T\}\_\{n,\\mathbf\{x\}\}^\{\\pm\}\\left\(\\mathcal\{A\}\\right\)\\right\|\\leq v\\ln\\left\(en\\right\)\.\(422\)Finally,[Lemma23](https://arxiv.org/html/2608.13882#Thmtheorem23)gives
v\\displaystyle v=\(a\)VC\(ℋL−1,m,G±\)≤\(b\)CL\(PL−1,m,G\+1\)ln\(e\(PL−1,m,G\+1\)\)\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\operatorname\{VC\}\\left\(\\mathcal\{H\}\_\{L\-1,m,G\}^\{\\pm\}\\right\)\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}C\_\{L\}\\left\(P\_\{L\-1,m,G\}\+1\\right\)\\ln\\left\(e\\left\(P\_\{L\-1,m,G\}\+1\\right\)\\right\)\.\(423\)Here, \(a\) applies \([403](https://arxiv.org/html/2608.13882#A4.E403)\), and \(b\) applies \([360](https://arxiv.org/html/2608.13882#A4.E360)\)\. Substituting \([423](https://arxiv.org/html/2608.13882#A4.E423)\) into \([422](https://arxiv.org/html/2608.13882#A4.E422)\) gives
ln\|𝒯n,𝐱±\(𝒜L−1m,G\)\|\\displaystyle\\ln\\left\|\\mathcal\{T\}\_\{n,\\mathbf\{x\}\}^\{\\pm\}\\left\(\\mathcal\{A\}\_\{L\-1\}^\{m,G\}\\right\)\\right\|≤CL\(PL−1,m,G\+1\)ln\(e\(PL−1,m,G\+1\)\)ln\(en\),\\displaystyle\\leq C\_\{L\}\\left\(P\_\{L\-1,m,G\}\+1\\right\)\\ln\\left\(e\\left\(P\_\{L\-1,m,G\}\+1\\right\)\\right\)\\ln\\left\(en\\right\),\(424\)which proves \([407](https://arxiv.org/html/2608.13882#A4.E407)\) and completes the proof\. ∎
###### Lemma B0\.
\(Brownian chaos bound from finite\-architecture threshold entropy\) LetL,m,G,n∈ℕL,m,G,n\\in\\mathbb\{N\}satisfyL≥2L\\geq 2,m≥1m\\geq 1,G≥2G\\geq 2, andn≥1n\\geq 1, and let𝐱1,…,𝐱n∈𝒳\\mathbf\{x\}\_\{1\},\\ldots,\\mathbf\{x\}\_\{n\}\\in\\mathcal\{X\}be fixed sample points\. LetPL−1,m,GP\_\{L\-1,m,G\}be the parameter\-count bound specified in[Section2\.2](https://arxiv.org/html/2608.13882#S2.SS2)and established in[Lemma17](https://arxiv.org/html/2608.13882#Thmtheorem17), part[3](https://arxiv.org/html/2608.13882#A4.I1.i3)\. Set
P\\displaystyle P:=PL−1,m,G,\\displaystyle:=P\_\{L\-1,m,G\},\(425\)ΓL−1,m,G,n\\displaystyle\\Gamma\_\{L\-1,m,G,n\}:=1\+\(P\+1\)ln\(e\(P\+1\)\)ln\(en\),\\displaystyle:=1\+\\left\(P\+1\\right\)\\ln\\left\(e\\left\(P\+1\\right\)\\right\)\\ln\\left\(en\\right\),\(426\)B𝐱\\displaystyle B\_\{\\mathbf\{x\}\}:=supa∈𝒜L−1m,Gmaxi∈\[n\]\|a\(𝐱i\)\|\.\\displaystyle:=\\sup\_\{a\\in\\mathcal\{A\}\_\{L\-1\}^\{m,G\}\}\\max\_\{i\\in\\left\[n\\right\]\}\\left\|a\\left\(\\mathbf\{x\}\_\{i\}\\right\)\\right\|\.\(427\)ThenB𝐱<∞\.B\_\{\\mathbf\{x\}\}<\\infty\.Moreover, there existsCL\>0C\_\{L\}\>0, depending only onLL, such that
ℭ^n,𝐱\(B\)\(𝒜L−1m,G\)≤CLB𝐱ΓL−1,m,G,n\.\\displaystyle\\widehat\{\\mathfrak\{C\}\}\_\{n,\\mathbf\{x\}\}^\{\(B\)\}\\left\(\\mathcal\{A\}\_\{L\-1\}^\{m,G\}\\right\)\\leq C\_\{L\}B\_\{\\mathbf\{x\}\}\\Gamma\_\{L\-1,m,G,n\}\.\(428\)
###### Proof\.
For notational convenience, set
𝒜\\displaystyle\\mathcal\{A\}:=𝒜L−1m,G,\\displaystyle:=\\mathcal\{A\}\_\{L\-1\}^\{m,G\},\(429\)𝒯\\displaystyle\\mathcal\{T\}:=𝒯n,𝐱±\(𝒜\),\\displaystyle:=\\mathcal\{T\}\_\{n,\\mathbf\{x\}\}^\{\\pm\}\\left\(\\mathcal\{A\}\\right\),\(430\)ΛL−1,m,G,n\\displaystyle\\Lambda\_\{L\-1,m,G,n\}:=\(P\+1\)ln\(e\(P\+1\)\)ln\(en\)\.\\displaystyle:=\\left\(P\+1\\right\)\\ln\\left\(e\\left\(P\+1\\right\)\\right\)\\ln\\left\(en\\right\)\.\(431\)By \([426](https://arxiv.org/html/2608.13882#A4.E426)\), one has
ΓL−1,m,G,n=1\+ΛL−1,m,G,n\.\\displaystyle\\Gamma\_\{L\-1,m,G,n\}=1\+\\Lambda\_\{L\-1,m,G,n\}\.\(432\)We first verify that the support envelope is finite\. By[Lemma17](https://arxiv.org/html/2608.13882#Thmtheorem17)[2](https://arxiv.org/html/2608.13882#A4.I1.i2), everya∈𝒜a\\in\\mathcal\{A\}satisfies
\|a\(𝐱\)\|≤A𝒳2−\(L−2\),𝐱∈𝒳\.\\displaystyle\\left\|a\\left\(\\mathbf\{x\}\\right\)\\right\|\\leq A\_\{\\mathcal\{X\}\}^\{\\,2^\{\-\(L\-2\)\}\},\\qquad\\mathbf\{x\}\\in\\mathcal\{X\}\.\(433\)Consequently,
B𝐱\\displaystyle B\_\{\\mathbf\{x\}\}=\(a\)supa∈𝒜maxi∈\[n\]\|a\(𝐱i\)\|≤\(b\)A𝒳2−\(L−2\)<\(c\)∞\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\sup\_\{a\\in\\mathcal\{A\}\}\\max\_\{i\\in\\left\[n\\right\]\}\\left\|a\\left\(\\mathbf\{x\}\_\{i\}\\right\)\\right\|\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}A\_\{\\mathcal\{X\}\}^\{\\,2^\{\-\(L\-2\)\}\}\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{<\}\}\\infty\.\(434\)Here, \(a\) applies \([427](https://arxiv.org/html/2608.13882#A4.E427)\), \(b\) applies \([433](https://arxiv.org/html/2608.13882#A4.E433)\) at every sample point, and \(c\) uses the compactness of𝒳\\mathcal\{X\}and the definition ofA𝒳A\_\{\\mathcal\{X\}\}in \([11](https://arxiv.org/html/2608.13882#S2.E11)\)\. We next control the signed threshold entropy\. By[Lemma24](https://arxiv.org/html/2608.13882#Thmtheorem24), there exists a constantCent,L\>0C\_\{\\mathrm\{ent\},L\}\>0, depending only onLL, such that
ln\|𝒯\|\\displaystyle\\ln\\left\|\\mathcal\{T\}\\right\|≤\(a\)Cent,L\(P\+1\)ln\(e\(P\+1\)\)ln\(en\)=\(b\)Cent,LΛL−1,m,G,n\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}C\_\{\\mathrm\{ent\},L\}\\left\(P\+1\\right\)\\ln\\left\(e\\left\(P\+1\\right\)\\right\)\\ln\\left\(en\\right\)\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}C\_\{\\mathrm\{ent\},L\}\\Lambda\_\{L\-1,m,G,n\}\.\(435\)Here, \(a\) applies \([407](https://arxiv.org/html/2608.13882#A4.E407)\), while \(b\) applies \([431](https://arxiv.org/html/2608.13882#A4.E431)\)\. Define
EL−1,m,G,n:=Cent,LΛL−1,m,G,n\.\\displaystyle E\_\{L\-1,m,G,n\}:=C\_\{\\mathrm\{ent\},L\}\\Lambda\_\{L\-1,m,G,n\}\.\(436\)Then \([435](https://arxiv.org/html/2608.13882#A4.E435)\) gives
ln\|𝒯n,𝐱±\(𝒜\)\|≤EL−1,m,G,n\.\\displaystyle\\ln\\left\|\\mathcal\{T\}\_\{n,\\mathbf\{x\}\}^\{\\pm\}\\left\(\\mathcal\{A\}\\right\)\\right\|\\leq E\_\{L\-1,m,G,n\}\.\(437\)Applying[Lemma22](https://arxiv.org/html/2608.13882#Thmtheorem22)with the support class𝒜\\mathcal\{A\}and the envelopeB𝐱B\_\{\\mathbf\{x\}\}gives
ℭ^n,𝐱\(B\)\(𝒜\)\\displaystyle\\widehat\{\\mathfrak\{C\}\}\_\{n,\\mathbf\{x\}\}^\{\(B\)\}\\left\(\\mathcal\{A\}\\right\)≤\(a\)CtrB𝐱\(1\+EL−1,m,G,n\)=\(b\)CtrB𝐱\(1\+Cent,LΛL−1,m,G,n\),\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}C\_\{\\mathrm\{tr\}\}B\_\{\\mathbf\{x\}\}\\left\(1\+E\_\{L\-1,m,G,n\}\\right\)\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}C\_\{\\mathrm\{tr\}\}B\_\{\\mathbf\{x\}\}\\left\(1\+C\_\{\\mathrm\{ent\},L\}\\Lambda\_\{L\-1,m,G,n\}\\right\),\(438\)whereCtr\>0C\_\{\\mathrm\{tr\}\}\>0is a universal numerical constant\. Here, \(a\) applies \([338](https://arxiv.org/html/2608.13882#A4.E338)\) using \([437](https://arxiv.org/html/2608.13882#A4.E437)\), and \(b\) substitutes \([436](https://arxiv.org/html/2608.13882#A4.E436)\)\. SetCaux,L:=max\{1,Cent,L\}\.C\_\{\\mathrm\{aux\},L\}:\\allowbreak=\\max\\left\\\{1,C\_\{\\mathrm\{ent\},L\}\\right\\\}\.SinceΛL−1,m,G,n≥0\\Lambda\_\{L\-1,m,G,n\}\\geq 0, we obtain
1\+Cent,LΛL−1,m,G,n\\displaystyle 1\+C\_\{\\mathrm\{ent\},L\}\\Lambda\_\{L\-1,m,G,n\}≤\(a\)Caux,L\+Caux,LΛL−1,m,G,n\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}C\_\{\\mathrm\{aux\},L\}\+C\_\{\\mathrm\{aux\},L\}\\Lambda\_\{L\-1,m,G,n\}=\(b\)Caux,L\(1\+ΛL−1,m,G,n\)\\displaystyle\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}C\_\{\\mathrm\{aux\},L\}\\left\(1\+\\Lambda\_\{L\-1,m,G,n\}\\right\)=\(c\)Caux,LΓL−1,m,G,n\.\\displaystyle\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{=\}\}C\_\{\\mathrm\{aux\},L\}\\Gamma\_\{L\-1,m,G,n\}\.\(439\)Here, \(a\) uses1≤Caux,L,Cent,L≤Caux,L1\\leq C\_\{\\mathrm\{aux\},L\},C\_\{\\mathrm\{ent\},L\}\\leq C\_\{\\mathrm\{aux\},L\}, \(b\) factors outCaux,LC\_\{\\mathrm\{aux\},L\}, and \(c\) applies \([432](https://arxiv.org/html/2608.13882#A4.E432)\)\. Substituting \([439](https://arxiv.org/html/2608.13882#A4.E439)\) into \([438](https://arxiv.org/html/2608.13882#A4.E438)\) gives
ℭ^n,𝐱\(B\)\(𝒜\)\\displaystyle\\widehat\{\\mathfrak\{C\}\}\_\{n,\\mathbf\{x\}\}^\{\(B\)\}\\left\(\\mathcal\{A\}\\right\)≤\(a\)CtrCaux,LB𝐱ΓL−1,m,G,n=\(b\)CLB𝐱ΓL−1,m,G,n,\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}C\_\{\\mathrm\{tr\}\}C\_\{\\mathrm\{aux\},L\}B\_\{\\mathbf\{x\}\}\\Gamma\_\{L\-1,m,G,n\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}C\_\{L\}B\_\{\\mathbf\{x\}\}\\Gamma\_\{L\-1,m,G,n\},\(440\)whereCL:=CtrCaux,LC\_\{L\}:=C\_\{\\mathrm\{tr\}\}C\_\{\\mathrm\{aux\},L\}\. Here, \(a\) combines \([438](https://arxiv.org/html/2608.13882#A4.E438)\) and \([439](https://arxiv.org/html/2608.13882#A4.E439)\), while \(b\) defines the depth\-dependent constantCLC\_\{L\}\. Substituting𝒜=𝒜L−1m,G\\mathcal\{A\}=\\mathcal\{A\}\_\{L\-1\}^\{m,G\}proves \([428](https://arxiv.org/html/2608.13882#A4.E428)\) and completes the proof\. ∎
###### Lemma B0\.
\(Uniform Hölder, pointwise, andL2\(ν\)L^\{2\}\(\\nu\)bounds for atomic VBKL generators\)LetL≥2L\\geq 2and assume that‖𝛚‖2≤1\\\|\\bm\{\\omega\}\\\|\_\{2\}\\leq 1for every𝛚∈Ω\\bm\{\\omega\}\\in\\Omega\. Define
RL:=\(∫𝒳‖𝐱‖22−\(L−2\)𝑑ν\(𝐱\)\)1/2\.\\displaystyle R\_\{L\}:=\\left\(\\int\_\{\\mathcal\{X\}\}\\\|\\mathbf\{x\}\\\|\_\{2\}^\{\\,2^\{\-\(L\-2\)\}\}\\,\\mathrm\{d\}\\nu\(\\mathbf\{x\}\)\\right\)^\{1/2\}\.\(441\)Then everyf∈𝒰Lf\\in\\mathcal\{U\}\_\{L\}satisfies
\|f\(𝐱\)−f\(𝐱′\)\|\\displaystyle\|f\(\\mathbf\{x\}\)\-f\(\\mathbf\{x\}^\{\\prime\}\)\|≤‖𝐱−𝐱′‖22−\(L−1\),𝐱,𝐱′∈𝒳,\\displaystyle\\leq\\\|\\mathbf\{x\}\-\\mathbf\{x\}^\{\\prime\}\\\|\_\{2\}^\{\\,2^\{\-\(L\-1\)\}\},\\qquad\\mathbf\{x\},\\mathbf\{x\}^\{\\prime\}\\in\\mathcal\{X\},\|f\(𝐱\)\|\\displaystyle\|f\(\\mathbf\{x\}\)\|≤‖𝐱‖22−\(L−1\),𝐱∈𝒳,\\displaystyle\\leq\\\|\\mathbf\{x\}\\\|\_\{2\}^\{\\,2^\{\-\(L\-1\)\}\},\\qquad\\mathbf\{x\}\\in\\mathcal\{X\},\(442\)and consequently
‖f‖L2\(ν\)≤RL\.\\displaystyle\\\|f\\\|\_\{L^\{2\}\(\\nu\)\}\\leq R\_\{L\}\.\(443\)
###### Proof\.
Forg∈ℋk\(B\)g\\in\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}with‖g‖ℋk\(B\)≤1\\\|g\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}\\leq 1, the reproducing property and the Brownian kernel metric give
\|g\(s\)−g\(t\)\|\\displaystyle\|g\(s\)\-g\(t\)\|≤‖k\(B\)\(s,⋅\)−k\(B\)\(t,⋅\)‖ℋk\(B\)=\|s−t\|1/2,\\displaystyle\\leq\\\|k^\{\(\\mathrm\{B\}\)\}\(s,\\cdot\)\-k^\{\(\\mathrm\{B\}\)\}\(t,\\cdot\)\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}=\|s\-t\|^\{1/2\},\|g\(s\)\|\\displaystyle\|g\(s\)\|≤‖k\(B\)\(s,⋅\)‖ℋk\(B\)=\|s\|1/2\.\\displaystyle\\leq\\\|k^\{\(\\mathrm\{B\}\)\}\(s,\\cdot\)\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}=\|s\|^\{1/2\}\.\(444\)Consequently, for every supporta:𝒳→ℝa:\\mathcal\{X\}\\to\\mathbb\{R\},
\|g\(a\(𝐱\)\)−g\(a\(𝐱′\)\)\|\\displaystyle\|g\(a\(\\mathbf\{x\}\)\)\-g\(a\(\\mathbf\{x\}^\{\\prime\}\)\)\|≤\|a\(𝐱\)−a\(𝐱′\)\|1/2,\\displaystyle\\leq\|a\(\\mathbf\{x\}\)\-a\(\\mathbf\{x\}^\{\\prime\}\)\|^\{1/2\},\|g\(a\(𝐱\)\)\|\\displaystyle\|g\(a\(\\mathbf\{x\}\)\)\|≤\|a\(𝐱\)\|1/2\.\\displaystyle\\leq\|a\(\\mathbf\{x\}\)\|^\{1/2\}\.\(445\)At the first level,\|𝝎⊤\(𝐱−𝐱′\)\|≤‖𝐱−𝐱′‖2\|\\bm\{\\omega\}^\{\\top\}\(\\mathbf\{x\}\-\\mathbf\{x\}^\{\\prime\}\)\|\\leq\\\|\\mathbf\{x\}\-\\mathbf\{x\}^\{\\prime\}\\\|\_\{2\}and\|𝝎⊤𝐱\|≤‖𝐱‖2\|\\bm\{\\omega\}^\{\\top\}\\mathbf\{x\}\|\\leq\\\|\\mathbf\{x\}\\\|\_\{2\}\. Applying \([445](https://arxiv.org/html/2608.13882#A4.E445)\) at each of theL−1L\-1Brownian compositions proves the two atomic bounds\. Squaring the pointwise estimate and integrating yields
‖f‖L2\(ν\)2≤∫𝒳‖𝐱‖22−\(L−2\)𝑑ν\(𝐱\)=RL2,\\displaystyle\\\|f\\\|\_\{L^\{2\}\(\\nu\)\}^\{2\}\\leq\\int\_\{\\mathcal\{X\}\}\\\|\\mathbf\{x\}\\\|\_\{2\}^\{\\,2^\{\-\(L\-2\)\}\}\\,\\mathrm\{d\}\\nu\(\\mathbf\{x\}\)=R\_\{L\}^\{2\},\(446\)which proves the claim\. ∎
###### Lemma B0\.
\(Attainment of the variation complexity\)Assume that𝒰L\\mathcal\{U\}\_\{L\}is compact inL2\(ν\)L^\{2\}\\left\(\\nu\\right\)\. Then, for everyF∈𝒱\(L\)F\\in\\mathcal\{V\}^\{\(L\)\}, there exists a finite signed Radon measureμF\\mu\_\{F\}on𝒰L\\mathcal\{U\}\_\{L\}such that
F=∫𝒰LudμF\(u\)inL2\(ν\),\\displaystyle F=\\int\_\{\\mathcal\{U\}\_\{L\}\}u\\,\\mathrm\{d\}\\mu\_\{F\}\\left\(u\\right\)\\qquad\\text\{in \}L^\{2\}\\left\(\\nu\\right\),\(447\)and
‖μF‖TV=𝒞^var\(L\)\(F\)\.\\displaystyle\\left\\\|\\mu\_\{F\}\\right\\\|\_\{\\mathrm\{TV\}\}=\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}\\left\(F\\right\)\.\(448\)
###### Proof\.
FixF∈𝒱\(L\)F\\in\\mathcal\{V\}^\{\(L\)\}and setcF:=𝒞^var\(L\)\(F\)\.c\_\{F\}:\\allowbreak=\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}\\left\(F\\right\)\.SinceF∈𝒱\(L\)F\\in\\mathcal\{V\}^\{\(L\)\}, one hascF<∞c\_\{F\}<\\infty\. By the definition ofcFc\_\{F\}, for everyn≥1n\\geq 1there exists a finite signed measureμn\\mu\_\{n\}on𝒰L\\mathcal\{U\}\_\{L\}such that
F\\displaystyle F=∫𝒰Ludμn\(u\)inL2\(ν\),\\displaystyle=\\int\_\{\\mathcal\{U\}\_\{L\}\}u\\,\\mathrm\{d\}\\mu\_\{n\}\\left\(u\\right\)\\qquad\\text\{in \}L^\{2\}\\left\(\\nu\\right\),\(449\)‖μn‖TV\\displaystyle\\left\\\|\\mu\_\{n\}\\right\\\|\_\{\\mathrm\{TV\}\}≤cF\+1n\.\\displaystyle\\leq c\_\{F\}\+\\frac\{1\}\{n\}\.\(450\)Because𝒰L\\mathcal\{U\}\_\{L\}is a compact metric space, every finite signed Borel measure on𝒰L\\mathcal\{U\}\_\{L\}is a finite signed Radon measure\. Moreover, \([450](https://arxiv.org/html/2608.13882#A4.E450)\) impliessupn≥1‖μn‖TV≤cF\+1<∞\.\\sup\_\{n\\geq 1\}\\left\\\|\\mu\_\{n\}\\right\\\|\_\{\\mathrm\{TV\}\}\\leq c\_\{F\}\+1<\\infty\.Since𝒰L\\mathcal\{U\}\_\{L\}is compact and metrizable,C\(𝒰L\)C\\left\(\\mathcal\{U\}\_\{L\}\\right\)is separable\. By the Riesz–Markov representation theorem, the space of finite signed Radon measures on𝒰L\\mathcal\{U\}\_\{L\}is isometrically identified withC\(𝒰L\)∗C\\left\(\\mathcal\{U\}\_\{L\}\\right\)^\{\*\}, where the dual norm is the total\-variation norm\. The sequence\(μn\)n≥1\\left\(\\mu\_\{n\}\\right\)\_\{n\\geq 1\}therefore lies in the closed dual ball of radiuscF\+1c\_\{F\}\+1\. By the Banach–Alaoglu theorem, this ball is weak\-\* compact\. SinceC\(𝒰L\)C\\left\(\\mathcal\{U\}\_\{L\}\\right\)is separable, the weak\-\* topology on this ball is metrizable, and the ball is consequently weak\-\* sequentially compact\. Hence, after passing to a subsequence, not relabelled, there exists a finite signed Radon measureμF\\mu\_\{F\}on𝒰L\\mathcal\{U\}\_\{L\}such that
μn⇀∗μFinC\(𝒰L\)∗\.\\displaystyle\\mu\_\{n\}\\stackrel\{\{\\scriptstyle\*\}\}\{\{\\rightharpoonup\}\}\\mu\_\{F\}\\qquad\\text\{in \}C\\left\(\\mathcal\{U\}\_\{L\}\\right\)^\{\*\}\.\(451\)We next prove thatμF\\mu\_\{F\}representsFF\. Since𝒰L\\mathcal\{U\}\_\{L\}is compact inL2\(ν\)L^\{2\}\\left\(\\nu\\right\), the canonical map
ι:𝒰L⟶L2\(ν\),ι\(u\):=u,\\displaystyle\\iota:\\mathcal\{U\}\_\{L\}\\longrightarrow L^\{2\}\\left\(\\nu\\right\),\\qquad\\iota\\left\(u\\right\):=u,\(452\)is continuous and bounded\. It is therefore strongly measurable and Bochner integrable with respect to the finite signed measureμF\\mu\_\{F\}\. Consequently,∫𝒰LudμF\(u\)∈L2\(ν\)\\int\_\{\\mathcal\{U\}\_\{L\}\}u\\,\\mathrm\{d\}\\mu\_\{F\}\\left\(u\\right\)\\in L^\{2\}\\left\(\\nu\\right\)is well\-defined\. Fixφ∈L2\(ν\)\\varphi\\in L^\{2\}\\left\(\\nu\\right\)\. The scalar\-valued map
u⟼⟨u,φ⟩L2\(ν\)\\displaystyle u\\longmapsto\\left\\langle u,\\varphi\\right\\rangle\_\{L^\{2\}\\left\(\\nu\\right\)\}\(453\)is continuous on𝒰L\\mathcal\{U\}\_\{L\}\. Therefore,
⟨F,φ⟩L2\(ν\)\\displaystyle\\left\\langle F,\\varphi\\right\\rangle\_\{L^\{2\}\\left\(\\nu\\right\)\}=\(a\)⟨∫𝒰Ludμn\(u\),φ⟩L2\(ν\)=\(b\)∫𝒰L⟨u,φ⟩L2\(ν\)dμn\(u\)\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\left\\langle\\int\_\{\\mathcal\{U\}\_\{L\}\}u\\,\\mathrm\{d\}\\mu\_\{n\}\\left\(u\\right\),\\varphi\\right\\rangle\_\{L^\{2\}\\left\(\\nu\\right\)\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\int\_\{\\mathcal\{U\}\_\{L\}\}\\left\\langle u,\\varphi\\right\\rangle\_\{L^\{2\}\\left\(\\nu\\right\)\}\\,\\mathrm\{d\}\\mu\_\{n\}\\left\(u\\right\)⟶\(c\)∫𝒰L⟨u,φ⟩L2\(ν\)dμF\(u\)=\(d\)⟨∫𝒰LudμF\(u\),φ⟩L2\(ν\)\.\\displaystyle\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{\\longrightarrow\}\}\\int\_\{\\mathcal\{U\}\_\{L\}\}\\left\\langle u,\\varphi\\right\\rangle\_\{L^\{2\}\\left\(\\nu\\right\)\}\\,\\mathrm\{d\}\\mu\_\{F\}\\left\(u\\right\)\\stackrel\{\{\\scriptstyle\(d\)\}\}\{\{=\}\}\\left\\langle\\int\_\{\\mathcal\{U\}\_\{L\}\}u\\,\\mathrm\{d\}\\mu\_\{F\}\\left\(u\\right\),\\varphi\\right\\rangle\_\{L^\{2\}\\left\(\\nu\\right\)\}\.\(454\)Here, \(a\) applies \([449](https://arxiv.org/html/2608.13882#A4.E449)\), \(b\) applies the duality identity for the Bochner integral, \(c\) follows from \([451](https://arxiv.org/html/2608.13882#A4.E451)\) and the continuity of \([453](https://arxiv.org/html/2608.13882#A4.E453)\), and \(d\) applies the Bochner\-integral duality identity once more\. Since the left\-hand side of \([454](https://arxiv.org/html/2608.13882#A4.E454)\) does not depend onnn, the limit identity gives
⟨F,φ⟩L2\(ν\)=⟨∫𝒰LudμF\(u\),φ⟩L2\(ν\)\\displaystyle\\left\\langle F,\\varphi\\right\\rangle\_\{L^\{2\}\\left\(\\nu\\right\)\}=\\left\\langle\\int\_\{\\mathcal\{U\}\_\{L\}\}u\\,\\mathrm\{d\}\\mu\_\{F\}\\left\(u\\right\),\\varphi\\right\\rangle\_\{L^\{2\}\\left\(\\nu\\right\)\}\(455\)for everyφ∈L2\(ν\)\\varphi\\in L^\{2\}\\left\(\\nu\\right\)\. The non\-degeneracy of theL2\(ν\)L^\{2\}\\left\(\\nu\\right\)inner product therefore yields
F=∫𝒰LudμF\(u\)inL2\(ν\),\\displaystyle F=\\int\_\{\\mathcal\{U\}\_\{L\}\}u\\,\\mathrm\{d\}\\mu\_\{F\}\\left\(u\\right\)\\qquad\\text\{in \}L^\{2\}\\left\(\\nu\\right\),\(456\)which proves \([447](https://arxiv.org/html/2608.13882#A4.E447)\)\. It remains to prove attainment of the variation complexity\. The total\-variation norm is weak\-\* lower semicontinuous, and hence
‖μF‖TV\\displaystyle\\left\\\|\\mu\_\{F\}\\right\\\|\_\{\\mathrm\{TV\}\}≤\(a\)lim infn→∞‖μn‖TV≤\(b\)limn→∞\(cF\+1n\)=\(c\)cF≤\(d\)‖μF‖TV\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}\\liminf\_\{n\\rightarrow\\infty\}\\left\\\|\\mu\_\{n\}\\right\\\|\_\{\\mathrm\{TV\}\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}\\lim\_\{n\\rightarrow\\infty\}\\left\(c\_\{F\}\+\\frac\{1\}\{n\}\\right\)\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{=\}\}c\_\{F\}\\stackrel\{\{\\scriptstyle\(d\)\}\}\{\{\\leq\}\}\\left\\\|\\mu\_\{F\}\\right\\\|\_\{\\mathrm\{TV\}\}\.\(457\)Here, \(a\) applies weak\-\* lower semicontinuity of the dual norm, \(b\) applies \([450](https://arxiv.org/html/2608.13882#A4.E450)\), \(c\) evaluates the limit, and \(d\) follows from the definition ofcF=𝒞^var\(L\)\(F\)c\_\{F\}=\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}\\left\(F\\right\)becauseμF\\mu\_\{F\}representsFF\. Thus, every inequality in \([457](https://arxiv.org/html/2608.13882#A4.E457)\) is an equality, and therefore‖μF‖TV=cF=𝒞^var\(L\)\(F\)\\left\\\|\\mu\_\{F\}\\right\\\|\_\{\\mathrm\{TV\}\}=c\_\{F\}=\\widehat\{\\mathcal\{C\}\}\_\{\\mathrm\{var\}\}^\{\(L\)\}\\left\(F\\right\)\. This proves \([448](https://arxiv.org/html/2608.13882#A4.E448)\) and completes the proof\. ∎
###### Lemma B0\.
\(Compactness of restricted Brownian profiles\)LetA\>0A\>0, and define
𝒢A:=\{g\|\[−A,A\]:g∈ℋk\(B\),‖g‖ℋk\(B\)≤1\}\.\\displaystyle\\mathcal\{G\}\_\{A\}:=\\left\\\{g\|\_\{\\left\[\-A,A\\right\]\}:g\\in\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\},\\ \\left\\\|g\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}\\leq 1\\right\\\}\.\(458\)Then𝒢A\\mathcal\{G\}\_\{A\}is compact inC\(\[−A,A\]\)C\\left\(\\left\[\-A,A\\right\]\\right\)equipped with the supremum norm\.
###### Proof\.
SinceC\(\[−A,A\]\)C\\left\(\\left\[\-A,A\\right\]\\right\), equipped with the supremum norm, is a metric space, it is sufficient to prove sequential compactness\. Let\(gn\)n≥1\\left\(g\_\{n\}\\right\)\_\{n\\geq 1\}be an arbitrary sequence in𝒢A\\mathcal\{G\}\_\{A\}\. By the definition of𝒢A\\mathcal\{G\}\_\{A\}, for everyn≥1n\\geq 1there existsg~n∈ℋk\(B\)\\widetilde\{g\}\_\{n\}\\in\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}such that
gn\\displaystyle g\_\{n\}=g~n\|\[−A,A\],\\displaystyle=\\widetilde\{g\}\_\{n\}\|\_\{\\left\[\-A,A\\right\]\},\(459\)‖g~n‖ℋk\(B\)\\displaystyle\\left\\\|\\widetilde\{g\}\_\{n\}\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}≤1\.\\displaystyle\\leq 1\.\(460\)By the characterization of the Brownian RKHS,g~n\(0\)=0\\widetilde\{g\}\_\{n\}\\left\(0\\right\)=0, the functiong~n\\widetilde\{g\}\_\{n\}is absolutely continuous, andg~n′∈L2\(ℝ\)\\widetilde\{g\}\_\{n\}^\{\\prime\}\\in L^\{2\}\\left\(\\mathbb\{R\}\\right\)\. Moreover,
∫ℝ\|g~n′\(t\)\|2𝑑t\\displaystyle\\int\_\{\\mathbb\{R\}\}\\left\|\\widetilde\{g\}\_\{n\}^\{\\prime\}\\left\(t\\right\)\\right\|^\{2\}\\,\\mathrm\{d\}t=\(a\)‖g~n‖ℋk\(B\)2≤\(b\)1\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\left\\\|\\widetilde\{g\}\_\{n\}\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}^\{2\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}1\.\(461\)Here, \(a\) applies the norm characterization ofℋk\(B\)\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}, while \(b\) applies \([460](https://arxiv.org/html/2608.13882#A4.E460)\)\. We first establish equicontinuity\. Fixs,t∈\[−A,A\]s,t\\in\\left\[\-A,A\\right\]withs<ts<t\. Sinceg~n\\widetilde\{g\}\_\{n\}is absolutely continuous,
gn\(t\)−gn\(s\)\\displaystyle g\_\{n\}\\left\(t\\right\)\-g\_\{n\}\\left\(s\\right\)=\(a\)g~n\(t\)−g~n\(s\)=\(b\)∫stg~n′\(r\)𝑑r\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\widetilde\{g\}\_\{n\}\\left\(t\\right\)\-\\widetilde\{g\}\_\{n\}\\left\(s\\right\)\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\int\_\{s\}^\{t\}\\widetilde\{g\}\_\{n\}^\{\\prime\}\\left\(r\\right\)\\,\\mathrm\{d\}r\.\(462\)Here, \(a\) applies \([459](https://arxiv.org/html/2608.13882#A4.E459)\), and \(b\) applies the fundamental theorem of calculus for absolutely continuous functions\. Consequently,
\|gn\(t\)−gn\(s\)\|\\displaystyle\\left\|g\_\{n\}\\left\(t\\right\)\-g\_\{n\}\\left\(s\\right\)\\right\|≤\(a\)\(∫st\|g~n′\(r\)\|2𝑑r\)1/2\|t−s\|1/2\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}\\left\(\\int\_\{s\}^\{t\}\\left\|\\widetilde\{g\}\_\{n\}^\{\\prime\}\\left\(r\\right\)\\right\|^\{2\}\\,\\mathrm\{d\}r\\right\)^\{1/2\}\\left\|t\-s\\right\|^\{1/2\}≤\(b\)\(∫ℝ\|g~n′\(r\)\|2𝑑r\)1/2\|t−s\|1/2≤\(c\)\|t−s\|1/2\.\\displaystyle\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}\\left\(\\int\_\{\\mathbb\{R\}\}\\left\|\\widetilde\{g\}\_\{n\}^\{\\prime\}\\left\(r\\right\)\\right\|^\{2\}\\,\\mathrm\{d\}r\\right\)^\{1/2\}\\left\|t\-s\\right\|^\{1/2\}\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{\\leq\}\}\\left\|t\-s\\right\|^\{1/2\}\.\(463\)Here, \(a\) applies the Cauchy–Schwarz inequality on\[s,t\]\\left\[s,t\\right\], \(b\) enlarges the domain of integration, and \(c\) applies \([461](https://arxiv.org/html/2608.13882#A4.E461)\)\. Thus,\(gn\)n≥1\\left\(g\_\{n\}\\right\)\_\{n\\geq 1\}is uniformly1/21/2\-Hölder continuous, and hence equicontinuous, on\[−A,A\]\\left\[\-A,A\\right\]\. We next establish uniform boundedness\. Since
gn\(0\)\\displaystyle g\_\{n\}\\left\(0\\right\)=\(a\)g~n\(0\)=\(b\)0,\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\widetilde\{g\}\_\{n\}\\left\(0\\right\)\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}0,\(464\)where \(a\) applies \([459](https://arxiv.org/html/2608.13882#A4.E459)\) and \(b\) uses the Brownian RKHS anchor condition, the estimate \([463](https://arxiv.org/html/2608.13882#A4.E463)\) gives, for everyt∈\[−A,A\]t\\in\\left\[\-A,A\\right\],
\|gn\(t\)\|\\displaystyle\\left\|g\_\{n\}\\left\(t\\right\)\\right\|=\(a\)\|gn\(t\)−gn\(0\)\|≤\(b\)\|t\|1/2≤\(c\)A1/2\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\left\|g\_\{n\}\\left\(t\\right\)\-g\_\{n\}\\left\(0\\right\)\\right\|\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}\\left\|t\\right\|^\{1/2\}\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{\\leq\}\}A^\{1/2\}\.\(465\)Here, \(a\) applies \([464](https://arxiv.org/html/2608.13882#A4.E464)\), \(b\) applies \([463](https://arxiv.org/html/2608.13882#A4.E463)\) with one endpoint equal to zero, and \(c\) uses\|t\|≤A\\left\|t\\right\|\\leq A\. The sequence\(gn\)n≥1\\left\(g\_\{n\}\\right\)\_\{n\\geq 1\}is therefore equicontinuous and uniformly bounded inC\(\[−A,A\]\)C\\left\(\\left\[\-A,A\\right\]\\right\)\. By the Arzelà–Ascoli theorem, there exist a subsequence, not relabelled, and a functiong∈C\(\[−A,A\]\)g\\in C\\left\(\\left\[\-A,A\\right\]\\right\)such that
‖gn−g‖L∞\(\[−A,A\]\)⟶0\.\\displaystyle\\left\\\|g\_\{n\}\-g\\right\\\|\_\{L^\{\\infty\}\\left\(\\left\[\-A,A\\right\]\\right\)\}\\longrightarrow 0\.\(466\)Sincegn\(0\)=0g\_\{n\}\\left\(0\\right\)=0for everyn≥1n\\geq 1, the uniform convergence gives
\|g\(0\)\|\\displaystyle\\left\|g\\left\(0\\right\)\\right\|=\(a\)\|g\(0\)−gn\(0\)\|≤\(b\)‖g−gn‖L∞\(\[−A,A\]\)⟶\(c\)0\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\left\|g\\left\(0\\right\)\-g\_\{n\}\\left\(0\\right\)\\right\|\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}\\left\\\|g\-g\_\{n\}\\right\\\|\_\{L^\{\\infty\}\\left\(\\left\[\-A,A\\right\]\\right\)\}\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{\\longrightarrow\}\}0\.\(467\)Here, \(a\) usesgn\(0\)=0g\_\{n\}\\left\(0\\right\)=0, \(b\) bounds pointwise evaluation by the supremum norm, and \(c\) applies \([466](https://arxiv.org/html/2608.13882#A4.E466)\)\. Hence,
g\(0\)=0\.\\displaystyle g\\left\(0\\right\)=0\.\(468\)It remains to prove thatg∈𝒢Ag\\in\\mathcal\{G\}\_\{A\}\. For everyn≥1n\\geq 1, define
hn:=g~n′\|\[−A,A\]\.\\displaystyle h\_\{n\}:=\\widetilde\{g\}\_\{n\}^\{\\prime\}\|\_\{\\left\[\-A,A\\right\]\}\.\(469\)Then
‖hn‖L2\(\[−A,A\]\)2\\displaystyle\\left\\\|h\_\{n\}\\right\\\|\_\{L^\{2\}\\left\(\\left\[\-A,A\\right\]\\right\)\}^\{2\}=\(a\)∫−AA\|g~n′\(t\)\|2𝑑t≤\(b\)∫ℝ\|g~n′\(t\)\|2𝑑t≤\(c\)1\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\int\_\{\-A\}^\{A\}\\left\|\\widetilde\{g\}\_\{n\}^\{\\prime\}\\left\(t\\right\)\\right\|^\{2\}\\,\\mathrm\{d\}t\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}\\int\_\{\\mathbb\{R\}\}\\left\|\\widetilde\{g\}\_\{n\}^\{\\prime\}\\left\(t\\right\)\\right\|^\{2\}\\,\\mathrm\{d\}t\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{\\leq\}\}1\.\(470\)Here, \(a\) applies \([469](https://arxiv.org/html/2608.13882#A4.E469)\), \(b\) enlarges the domain of integration, and \(c\) applies \([461](https://arxiv.org/html/2608.13882#A4.E461)\)\. Thus,\(hn\)n≥1\\left\(h\_\{n\}\\right\)\_\{n\\geq 1\}is a bounded sequence in the Hilbert spaceL2\(\[−A,A\]\)L^\{2\}\\left\(\\left\[\-A,A\\right\]\\right\)\. Since Hilbert spaces are reflexive, bounded sequences are weakly sequentially relatively compact\. Therefore, after passing to a further subsequence, not relabelled, there existsh∈L2\(\[−A,A\]\)h\\in L^\{2\}\\left\(\\left\[\-A,A\\right\]\\right\)such that
hn⇀hweakly inL2\(\[−A,A\]\)\.\\displaystyle h\_\{n\}\\rightharpoonup h\\qquad\\text\{weakly in \}L^\{2\}\\left\(\\left\[\-A,A\\right\]\\right\)\.\(471\)The further subsequence continues to satisfy \([466](https://arxiv.org/html/2608.13882#A4.E466)\)\. Fixs,t∈\[−A,A\]s,t\\in\\left\[\-A,A\\right\]withs<ts<t\. By absolute continuity and \([469](https://arxiv.org/html/2608.13882#A4.E469)\),
gn\(t\)−gn\(s\)=∫sthn\(r\)𝑑r\.\\displaystyle g\_\{n\}\\left\(t\\right\)\-g\_\{n\}\\left\(s\\right\)=\\int\_\{s\}^\{t\}h\_\{n\}\\left\(r\\right\)\\,\\mathrm\{d\}r\.\(472\)The uniform convergence gives
gn\(t\)−gn\(s\)⟶g\(t\)−g\(s\)\.\\displaystyle g\_\{n\}\\left\(t\\right\)\-g\_\{n\}\\left\(s\\right\)\\longrightarrow g\\left\(t\\right\)\-g\\left\(s\\right\)\.\(473\)On the other hand,
∫sthn\(r\)𝑑r\\displaystyle\\int\_\{s\}^\{t\}h\_\{n\}\\left\(r\\right\)\\,\\mathrm\{d\}r=\(a\)⟨hn,𝟏\[s,t\]⟩L2\(\[−A,A\]\)⟶\(b\)⟨h,𝟏\[s,t\]⟩L2\(\[−A,A\]\)=\(c\)∫sth\(r\)𝑑r\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\left\\langle h\_\{n\},\\mathbf\{1\}\_\{\\left\[s,t\\right\]\}\\right\\rangle\_\{L^\{2\}\\left\(\\left\[\-A,A\\right\]\\right\)\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\longrightarrow\}\}\\left\\langle h,\\mathbf\{1\}\_\{\\left\[s,t\\right\]\}\\right\\rangle\_\{L^\{2\}\\left\(\\left\[\-A,A\\right\]\\right\)\}\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{=\}\}\\int\_\{s\}^\{t\}h\\left\(r\\right\)\\,\\mathrm\{d\}r\.\(474\)Here, \(a\) rewrites the integral as anL2\(\[−A,A\]\)L^\{2\}\\left\(\\left\[\-A,A\\right\]\\right\)inner product, \(b\) applies \([471](https://arxiv.org/html/2608.13882#A4.E471)\) with the fixed test function𝟏\[s,t\]\\mathbf\{1\}\_\{\\left\[s,t\\right\]\}, and \(c\) rewrites the limiting inner product as an integral\. Comparing \([473](https://arxiv.org/html/2608.13882#A4.E473)\) and \([474](https://arxiv.org/html/2608.13882#A4.E474)\) in \([472](https://arxiv.org/html/2608.13882#A4.E472)\) gives
g\(t\)−g\(s\)=∫sth\(r\)𝑑r,−A≤s<t≤A\.\\displaystyle g\\left\(t\\right\)\-g\\left\(s\\right\)=\\int\_\{s\}^\{t\}h\\left\(r\\right\)\\,\\mathrm\{d\}r,\\qquad\-A\\leq s<t\\leq A\.\(475\)Since the interval\[−A,A\]\\left\[\-A,A\\right\]has finite measure andh∈L2\(\[−A,A\]\)h\\in L^\{2\}\\left\(\\left\[\-A,A\\right\]\\right\), the Cauchy–Schwarz inequality gives
∫−AA\|h\(r\)\|𝑑r\\displaystyle\\int\_\{\-A\}^\{A\}\\left\|h\\left\(r\\right\)\\right\|\\,\\mathrm\{d\}r≤\(a\)\(2A\)1/2‖h‖L2\(\[−A,A\]\)<\(b\)∞\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}\\left\(2A\\right\)^\{1/2\}\\left\\\|h\\right\\\|\_\{L^\{2\}\\left\(\\left\[\-A,A\\right\]\\right\)\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{<\}\}\\infty\.\(476\)Here, \(a\) applies the Cauchy–Schwarz inequality, and \(b\) usesh∈L2\(\[−A,A\]\)h\\in L^\{2\}\\left\(\\left\[\-A,A\\right\]\\right\)\. Thus,h∈L1\(\[−A,A\]\)h\\in L^\{1\}\\left\(\\left\[\-A,A\\right\]\\right\)\. Equation \([475](https://arxiv.org/html/2608.13882#A4.E475)\) therefore shows thatggis absolutely continuous on\[−A,A\]\\left\[\-A,A\\right\]and thatg′=halmost everywhere on\[−A,A\]\.g^\{\\prime\}\\allowbreak=h\\qquad\\text\{almost everywhere on \}\\left\[\-A,A\\right\]\.By weak lower semicontinuity of the Hilbert\-space norm,
∫−AA\|h\(r\)\|2𝑑r\\displaystyle\\int\_\{\-A\}^\{A\}\\left\|h\\left\(r\\right\)\\right\|^\{2\}\\,\\mathrm\{d\}r=\(a\)‖h‖L2\(\[−A,A\]\)2≤\(b\)lim infn→∞‖hn‖L2\(\[−A,A\]\)2≤\(c\)1\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\left\\\|h\\right\\\|\_\{L^\{2\}\\left\(\\left\[\-A,A\\right\]\\right\)\}^\{2\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}\\liminf\_\{n\\rightarrow\\infty\}\\left\\\|h\_\{n\}\\right\\\|\_\{L^\{2\}\\left\(\\left\[\-A,A\\right\]\\right\)\}^\{2\}\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{\\leq\}\}1\.\(477\)Here, \(a\) applies the definition of theL2L^\{2\}norm, \(b\) applies weak lower semicontinuity under \([471](https://arxiv.org/html/2608.13882#A4.E471)\), and \(c\) applies \([470](https://arxiv.org/html/2608.13882#A4.E470)\)\. Define the zero extension ofhhby
h^\(t\):=h\(t\)𝟏\[−A,A\]\(t\),t∈ℝ,\\displaystyle\\widehat\{h\}\\left\(t\\right\):=h\\left\(t\\right\)\\mathbf\{1\}\_\{\\left\[\-A,A\\right\]\}\\left\(t\\right\),\\qquad t\\in\\mathbb\{R\},\(478\)and define
g^\(t\):=∫0th^\(r\)𝑑r,t∈ℝ\.\\displaystyle\\widehat\{g\}\\left\(t\\right\):=\\int\_\{0\}^\{t\}\\widehat\{h\}\\left\(r\\right\)\\,\\mathrm\{d\}r,\\qquad t\\in\\mathbb\{R\}\.\(479\)Becauseh^\\widehat\{h\}has bounded support and belongs toL2\(ℝ\)L^\{2\}\\left\(\\mathbb\{R\}\\right\), one also hash^∈L1\(ℝ\)\\widehat\{h\}\\in L^\{1\}\\left\(\\mathbb\{R\}\\right\)\. Therefore,g^\\widehat\{g\}is absolutely continuous onℝ\\mathbb\{R\}, satisfiesg^\(0\)=0,\\widehat\{g\}\\left\(0\\right\)\\allowbreak=0,and has weak derivative
g^′=h^almost everywhere onℝ\.\\displaystyle\\widehat\{g\}^\{\\prime\}=\\widehat\{h\}\\qquad\\text\{almost everywhere on \}\\mathbb\{R\}\.\(480\)We now verify thatg^\\widehat\{g\}extendsgg\. Ift∈\[0,A\]t\\in\\left\[0,A\\right\], then
g^\(t\)\\displaystyle\\widehat\{g\}\\left\(t\\right\)=\(a\)∫0th\(r\)𝑑r=\(b\)g\(t\)−g\(0\)=\(c\)g\(t\)\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\int\_\{0\}^\{t\}h\\left\(r\\right\)\\,\\mathrm\{d\}r\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}g\\left\(t\\right\)\-g\\left\(0\\right\)\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{=\}\}g\\left\(t\\right\)\.\(481\)Here, \(a\) applies \([478](https://arxiv.org/html/2608.13882#A4.E478)\), \(b\) applies \([475](https://arxiv.org/html/2608.13882#A4.E475)\) withs=0s=0, and \(c\) applies \([468](https://arxiv.org/html/2608.13882#A4.E468)\)\. Ift∈\[−A,0\]t\\in\\left\[\-A,0\\right\], then
g^\(t\)\\displaystyle\\widehat\{g\}\\left\(t\\right\)=\(a\)∫0th\(r\)dr=\(b\)−∫t0h\(r\)dr=\(c\)−\(g\(0\)−g\(t\)\)=\(d\)g\(t\)\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\int\_\{0\}^\{t\}h\\left\(r\\right\)\\,\\mathrm\{d\}r\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\-\\int\_\{t\}^\{0\}h\\left\(r\\right\)\\,\\mathrm\{d\}r\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{=\}\}\-\\left\(g\\left\(0\\right\)\-g\\left\(t\\right\)\\right\)\\stackrel\{\{\\scriptstyle\(d\)\}\}\{\{=\}\}g\\left\(t\\right\)\.\(482\)Here, \(a\) applies \([479](https://arxiv.org/html/2608.13882#A4.E479)\), \(b\) reverses the orientation of the integral, \(c\) applies \([475](https://arxiv.org/html/2608.13882#A4.E475)\), and \(d\) applies \([468](https://arxiv.org/html/2608.13882#A4.E468)\)\. Therefore,
g^\|\[−A,A\]=g\.\\displaystyle\\widehat\{g\}\|\_\{\\left\[\-A,A\\right\]\}=g\.\(483\)Finally,
‖g^‖ℋk\(B\)2\\displaystyle\\left\\\|\\widehat\{g\}\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}^\{2\}=\(a\)∫ℝ\|g^′\(t\)\|2𝑑t=\(b\)∫ℝ\|h^\(t\)\|2𝑑t=\(c\)∫−AA\|h\(t\)\|2𝑑t≤\(d\)1\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\int\_\{\\mathbb\{R\}\}\\left\|\\widehat\{g\}^\{\\prime\}\\left\(t\\right\)\\right\|^\{2\}\\,\\mathrm\{d\}t\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\int\_\{\\mathbb\{R\}\}\\left\|\\widehat\{h\}\\left\(t\\right\)\\right\|^\{2\}\\,\\mathrm\{d\}t\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{=\}\}\\int\_\{\-A\}^\{A\}\\left\|h\\left\(t\\right\)\\right\|^\{2\}\\,\\mathrm\{d\}t\\stackrel\{\{\\scriptstyle\(d\)\}\}\{\{\\leq\}\}1\.\(484\)Here, \(a\) applies the Brownian RKHS norm characterization, \(b\) applies \([480](https://arxiv.org/html/2608.13882#A4.E480)\), \(c\) applies the definition of the zero extension \([478](https://arxiv.org/html/2608.13882#A4.E478)\), and \(d\) applies \([477](https://arxiv.org/html/2608.13882#A4.E477)\)\. Hence,
g^∈ℋk\(B\),‖g^‖ℋk\(B\)≤1\.\\displaystyle\\widehat\{g\}\\in\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\},\\qquad\\left\\\|\\widehat\{g\}\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}\\leq 1\.\(485\)Together with \([483](https://arxiv.org/html/2608.13882#A4.E483)\), the definition of𝒢A\\mathcal\{G\}\_\{A\}therefore givesg∈𝒢Ag\\in\\mathcal\{G\}\_\{A\}\. We have shown that every sequence in𝒢A\\mathcal\{G\}\_\{A\}admits a subsequence converging uniformly on\[−A,A\]\\left\[\-A,A\\right\]to an element of𝒢A\\mathcal\{G\}\_\{A\}\. Thus,𝒢A\\mathcal\{G\}\_\{A\}is sequentially compact inC\(\[−A,A\]\)C\\left\(\\left\[\-A,A\\right\]\\right\)\. Since this is a metric space, sequential compactness is equivalent to compactness\. Therefore,𝒢A\\mathcal\{G\}\_\{A\}is compact inC\(\[−A,A\]\)C\\left\(\\left\[\-A,A\\right\]\\right\)equipped with the supremum norm\. ∎
###### Lemma B0\.
\(Compactness of the atomic VBKL dictionary\)Let𝒳⊆ℝd\\mathcal\{X\}\\subseteq\\mathbb\{R\}^\{d\}be compact, letν\\nube a Borel probability measure on𝒳\\mathcal\{X\}, and letΩ⊆𝕊d−1\\Omega\\subseteq\\mathbb\{S\}^\{d\-1\}be compact\. Let\(𝒰l\)l≥1\\left\(\\mathcal\{U\}\_\{l\}\\right\)\_\{l\\geq 1\}be the recursive atomic dictionaries introduced in[Section2\.1](https://arxiv.org/html/2608.13882#S2.SS1)\. Then, for everyL≥1L\\geq 1, the class𝒰L\\mathcal\{U\}\_\{L\}is compact inC\(𝒳\)C\\left\(\\mathcal\{X\}\\right\)equipped with the supremum norm\. Consequently, under the canonical continuous embeddingC\(𝒳\)⟶L2\(ν\),C\\left\(\\mathcal\{X\}\\right\)\\longrightarrow L^\{2\}\\left\(\\nu\\right\),the class𝒰L\\mathcal\{U\}\_\{L\}is also compact inL2\(ν\)L^\{2\}\\left\(\\nu\\right\)\.
###### Proof\.
We prove the first assertion by induction on the recursion level\. For the base case, define
T\\displaystyle T:Ω⟶C\(𝒳\),\\displaystyle:\\Omega\\longrightarrow C\\left\(\\mathcal\{X\}\\right\),\(T𝝎\)\(𝐱\)\\displaystyle\\left\(T\\bm\{\\omega\}\\right\)\\left\(\\mathbf\{x\}\\right\):=𝝎⊤𝐱\.\\displaystyle:=\\bm\{\\omega\}^\{\\top\}\\mathbf\{x\}\.\(486\)SetM𝒳:=sup𝐱∈𝒳‖𝐱‖2M\_\{\\mathcal\{X\}\}:=\\sup\_\{\\mathbf\{x\}\\in\\mathcal\{X\}\}\\left\\\|\\mathbf\{x\}\\right\\\|\_\{2\}\. Since𝒳\\mathcal\{X\}is compact and𝐱↦‖𝐱‖2\\mathbf\{x\}\\mapsto\\left\\\|\\mathbf\{x\}\\right\\\|\_\{2\}is continuous, one hasM𝒳<∞M\_\{\\mathcal\{X\}\}<\\infty\. For every𝝎,𝝎′∈Ω\\bm\{\\omega\},\\bm\{\\omega\}^\{\\prime\}\\in\\Omega,
‖T𝝎−T𝝎′‖∞\\displaystyle\\left\\\|T\\bm\{\\omega\}\-T\\bm\{\\omega\}^\{\\prime\}\\right\\\|\_\{\\infty\}=\(a\)sup𝐱∈𝒳\|\(𝝎−𝝎′\)⊤𝐱\|≤\(b\)sup𝐱∈𝒳‖𝝎−𝝎′‖2‖𝐱‖2=\(c\)M𝒳‖𝝎−𝝎′‖2\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\sup\_\{\\mathbf\{x\}\\in\\mathcal\{X\}\}\\left\|\\left\(\\bm\{\\omega\}\-\\bm\{\\omega\}^\{\\prime\}\\right\)^\{\\top\}\\mathbf\{x\}\\right\|\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}\\sup\_\{\\mathbf\{x\}\\in\\mathcal\{X\}\}\\left\\\|\\bm\{\\omega\}\-\\bm\{\\omega\}^\{\\prime\}\\right\\\|\_\{2\}\\left\\\|\\mathbf\{x\}\\right\\\|\_\{2\}\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{=\}\}M\_\{\\mathcal\{X\}\}\\left\\\|\\bm\{\\omega\}\-\\bm\{\\omega\}^\{\\prime\}\\right\\\|\_\{2\}\.\(487\)Here, \(a\) applies the definition ofTTand the supremum norm, \(b\) applies the Euclidean Cauchy–Schwarz inequality, and \(c\) factors out the quantity‖𝝎−𝝎′‖2\\left\\\|\\bm\{\\omega\}\-\\bm\{\\omega\}^\{\\prime\}\\right\\\|\_\{2\}and applies the definition ofM𝒳M\_\{\\mathcal\{X\}\}\. Thus,TTis Lipschitz continuous\. SinceΩ\\Omegais compact and𝒰1=T\(Ω\)\\mathcal\{U\}\_\{1\}=T\\left\(\\Omega\\right\), the continuity ofTTimplies that𝒰1\\mathcal\{U\}\_\{1\}is compact inC\(𝒳\)C\\left\(\\mathcal\{X\}\\right\)\.
Assume now that𝒰l\\mathcal\{U\}\_\{l\}is compact inC\(𝒳\)C\\left\(\\mathcal\{X\}\\right\)for somel≥1l\\geq 1\. We prove that𝒰l\+1\\mathcal\{U\}\_\{l\+1\}is compact inC\(𝒳\)C\\left\(\\mathcal\{X\}\\right\)\.
Let\(fn\)n≥1\\left\(f\_\{n\}\\right\)\_\{n\\geq 1\}be an arbitrary sequence in𝒰l\+1\\mathcal\{U\}\_\{l\+1\}\. By the recursive definition of the atomic dictionary, for everyn≥1n\\geq 1there existan∈𝒰l,gn∈ℋk\(B\)a\_\{n\}\\in\\mathcal\{U\}\_\{l\},g\_\{n\}\\in\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}such that
‖gn‖ℋk\(B\)\\displaystyle\\left\\\|g\_\{n\}\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}≤1,\\displaystyle\\leq 1,\(488\)fn\(𝐱\)\\displaystyle f\_\{n\}\\left\(\\mathbf\{x\}\\right\)=gn\(an\(𝐱\)\),𝐱∈𝒳\.\\displaystyle=g\_\{n\}\\left\(a\_\{n\}\\left\(\\mathbf\{x\}\\right\)\\right\),\\qquad\\mathbf\{x\}\\in\\mathcal\{X\}\.\(489\)By the induction hypothesis,𝒰l\\mathcal\{U\}\_\{l\}is compact inC\(𝒳\)C\\left\(\\mathcal\{X\}\\right\)\. Hence, after passing to a subsequence, not relabelled, there existsa∈𝒰la\\in\\mathcal\{U\}\_\{l\}such that
‖an−a‖∞⟶0\.\\displaystyle\\left\\\|a\_\{n\}\-a\\right\\\|\_\{\\infty\}\\longrightarrow 0\.\(490\)Since𝒰l\\mathcal\{U\}\_\{l\}is compact inC\(𝒳\)C\\left\(\\mathcal\{X\}\\right\)and the supremum norm is continuous, the class𝒰l\\mathcal\{U\}\_\{l\}is bounded\. DefineA:=max\{1,supb∈𝒰l‖b‖∞\}A:=\\max\\left\\\{1,\\sup\_\{b\\in\\mathcal\{U\}\_\{l\}\}\\left\\\|b\\right\\\|\_\{\\infty\}\\right\\\}\. ThenA<∞A<\\infty, and
a\(𝒳\)\\displaystyle a\\left\(\\mathcal\{X\}\\right\)⊆\[−A,A\],\\displaystyle\\subseteq\\left\[\-A,A\\right\],an\(𝒳\)\\displaystyle a\_\{n\}\\left\(\\mathcal\{X\}\\right\)⊆\[−A,A\],n≥1\.\\displaystyle\\subseteq\\left\[\-A,A\\right\],\\qquad n\\geq 1\.\(491\)For everyn≥1n\\geq 1, the restrictiongn\|\[−A,A\]g\_\{n\}\|\_\{\\left\[\-A,A\\right\]\}belongs to the compact profile class𝒢A\\mathcal\{G\}\_\{A\}from[Lemma28](https://arxiv.org/html/2608.13882#Thmtheorem28)\. Therefore, after passing to a further subsequence, not relabelled, there existsg∗∈𝒢Ag\_\{\\ast\}\\in\\mathcal\{G\}\_\{A\}such that
supt∈\[−A,A\]\|gn\(t\)−g∗\(t\)\|⟶0\.\\displaystyle\\sup\_\{t\\in\\left\[\-A,A\\right\]\}\\left\|g\_\{n\}\\left\(t\\right\)\-g\_\{\\ast\}\\left\(t\\right\)\\right\|\\longrightarrow 0\.\(492\)By the definition of𝒢A\\mathcal\{G\}\_\{A\}, there existsg~∈ℋk\(B\)\\widetilde\{g\}\\in\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}such that
g~\|\[−A,A\]\\displaystyle\\widetilde\{g\}\|\_\{\\left\[\-A,A\\right\]\}=g∗,\\displaystyle=g\_\{\\ast\},‖g~‖ℋk\(B\)\\displaystyle\\left\\\|\\widetilde\{g\}\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}≤1\.\\displaystyle\\leq 1\.\(493\)We now prove uniform convergence of the composed functions\. Fix𝐱∈𝒳\\mathbf\{x\}\\in\\mathcal\{X\}\. Sincean\(𝐱\)a\_\{n\}\\left\(\\mathbf\{x\}\\right\)anda\(𝐱\)a\\left\(\\mathbf\{x\}\\right\)belong to\[−A,A\]\\left\[\-A,A\\right\],
\|gn\(an\(𝐱\)\)−g~\(a\(𝐱\)\)\|≤\(a\)\|gn\(an\(𝐱\)\)−g~\(an\(𝐱\)\)\|\+\|g~\(an\(𝐱\)\)−g~\(a\(𝐱\)\)\|\\displaystyle\\left\|g\_\{n\}\\left\(a\_\{n\}\\left\(\\mathbf\{x\}\\right\)\\right\)\-\\widetilde\{g\}\\left\(a\\left\(\\mathbf\{x\}\\right\)\\right\)\\right\|\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}\\left\|g\_\{n\}\\left\(a\_\{n\}\\left\(\\mathbf\{x\}\\right\)\\right\)\-\\widetilde\{g\}\\left\(a\_\{n\}\\left\(\\mathbf\{x\}\\right\)\\right\)\\right\|\+\\left\|\\widetilde\{g\}\\left\(a\_\{n\}\\left\(\\mathbf\{x\}\\right\)\\right\)\-\\widetilde\{g\}\\left\(a\\left\(\\mathbf\{x\}\\right\)\\right\)\\right\|=\(b\)\|gn\(an\(𝐱\)\)−g∗\(an\(𝐱\)\)\|\+\|g~\(an\(𝐱\)\)−g~\(a\(𝐱\)\)\|\\displaystyle\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{=\}\}\\left\|g\_\{n\}\\left\(a\_\{n\}\\left\(\\mathbf\{x\}\\right\)\\right\)\-g\_\{\\ast\}\\left\(a\_\{n\}\\left\(\\mathbf\{x\}\\right\)\\right\)\\right\|\+\\left\|\\widetilde\{g\}\\left\(a\_\{n\}\\left\(\\mathbf\{x\}\\right\)\\right\)\-\\widetilde\{g\}\\left\(a\\left\(\\mathbf\{x\}\\right\)\\right\)\\right\|≤\(c\)supt∈\[−A,A\]\|gn\(t\)−g∗\(t\)\|\+\|g~\(an\(𝐱\)\)−g~\(a\(𝐱\)\)\|\.\\displaystyle\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{\\leq\}\}\\sup\_\{t\\in\\left\[\-A,A\\right\]\}\\left\|g\_\{n\}\\left\(t\\right\)\-g\_\{\\ast\}\\left\(t\\right\)\\right\|\+\\left\|\\widetilde\{g\}\\left\(a\_\{n\}\\left\(\\mathbf\{x\}\\right\)\\right\)\-\\widetilde\{g\}\\left\(a\\left\(\\mathbf\{x\}\\right\)\\right\)\\right\|\.\(494\)Here, \(a\) applies the triangle inequality, \(b\) usesg~=g∗\\widetilde\{g\}=g\_\{\\ast\}on\[−A,A\]\\left\[\-A,A\\right\], and \(c\) bounds the first term by the uniform profile distance\.
Define the modulus of continuity ofg~\\widetilde\{g\}on\[−A,A\]\\left\[\-A,A\\right\]byωg~\(δ\):=sup\\omega\_\{\\widetilde\{g\}\}\\left\(\\delta\\right\):=\\sup\{\|g~\(s\)−g~\(t\)\|:\\left\\\{\\left\|\\widetilde\{g\}\\left\(s\\right\)\-\\widetilde\{g\}\\left\(t\\right\)\\right\|:\\right\.s,t∈\[−A,A\],\|s−t\|≤δ\}\.\\left\.s,t\\in\\left\[\-A,A\\right\],\\;\\left\|s\-t\\right\|\\leq\\delta\\right\\\}\.Sinceg~\\widetilde\{g\}is continuous on the compact interval\[−A,A\]\\left\[\-A,A\\right\], it is uniformly continuous, and thereforeωg~\(δ\)⟶0\\omega\_\{\\widetilde\{g\}\}\\left\(\\delta\\right\)\\longrightarrow 0asδ↓0\.\\text\{as \}\\delta\\downarrow 0\.Taking the supremum over𝐱∈𝒳\\mathbf\{x\}\\in\\mathcal\{X\}in \([494](https://arxiv.org/html/2608.13882#A4.E494)\) gives
‖fn−g~∘a‖∞\\displaystyle\\left\\\|f\_\{n\}\-\\widetilde\{g\}\\circ a\\right\\\|\_\{\\infty\}≤\(a\)supt∈\[−A,A\]\|gn\(t\)−g∗\(t\)\|\+ωg~\(‖an−a‖∞\)\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}\\sup\_\{t\\in\\left\[\-A,A\\right\]\}\\left\|g\_\{n\}\\left\(t\\right\)\-g\_\{\\ast\}\\left\(t\\right\)\\right\|\+\\omega\_\{\\widetilde\{g\}\}\\left\(\\left\\\|a\_\{n\}\-a\\right\\\|\_\{\\infty\}\\right\)\.\(495\)Here, \(a\) also usesfn=gn∘anf\_\{n\}=g\_\{n\}\\circ a\_\{n\}and the definition of the modulus of continuity\. The first term on the right\-hand side converges to zero by \([492](https://arxiv.org/html/2608.13882#A4.E492)\)\. The second converges to zero by \([490](https://arxiv.org/html/2608.13882#A4.E490)\) and the continuity property ofωg~\\omega\_\{\\widetilde\{g\}\}\. Consequently,‖fn−g~∘a‖∞⟶0\.\\left\\\|f\_\{n\}\-\\widetilde\{g\}\\circ a\\right\\\|\_\{\\infty\}\\longrightarrow 0\.Sincea∈𝒰la\\in\\mathcal\{U\}\_\{l\},g~∈ℋk\(B\)\\widetilde\{g\}\\in\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}, and‖g~‖ℋk\(B\)≤1\\left\\\|\\widetilde\{g\}\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}\\leq 1, the recursive definition givesg~∘a∈𝒰l\+1\\widetilde\{g\}\\circ a\\in\\mathcal\{U\}\_\{l\+1\}\. Thus, every sequence in𝒰l\+1\\mathcal\{U\}\_\{l\+1\}admits a subsequence converging inC\(𝒳\)C\\left\(\\mathcal\{X\}\\right\)to an element of𝒰l\+1\\mathcal\{U\}\_\{l\+1\}\. Therefore,𝒰l\+1\\mathcal\{U\}\_\{l\+1\}is sequentially compact\. SinceC\(𝒳\)C\\left\(\\mathcal\{X\}\\right\)is a metric space, sequential compactness is equivalent to compactness\. The induction is complete, and𝒰L\\mathcal\{U\}\_\{L\}is compact inC\(𝒳\)C\\left\(\\mathcal\{X\}\\right\)for everyL≥1L\\geq 1\.
Finally, consider the canonical mapJ:C\(𝒳\)⟶L2\(ν\),Jf:=fJ:C\\left\(\\mathcal\{X\}\\right\)\\longrightarrow L^\{2\}\\left\(\\nu\\right\),\\qquad Jf:=f\. For everyf∈C\(𝒳\)f\\in C\\left\(\\mathcal\{X\}\\right\),
‖Jf‖L2\(ν\)2\\displaystyle\\left\\\|Jf\\right\\\|\_\{L^\{2\}\\left\(\\nu\\right\)\}^\{2\}=\(a\)∫𝒳\|f\(𝐱\)\|2𝑑ν\(𝐱\)≤\(b\)‖f‖∞2ν\(𝒳\)=\(c\)‖f‖∞2\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\int\_\{\\mathcal\{X\}\}\\left\|f\\left\(\\mathbf\{x\}\\right\)\\right\|^\{2\}\\,\\mathrm\{d\}\\nu\\left\(\\mathbf\{x\}\\right\)\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}\\left\\\|f\\right\\\|\_\{\\infty\}^\{2\}\\nu\\left\(\\mathcal\{X\}\\right\)\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{=\}\}\\left\\\|f\\right\\\|\_\{\\infty\}^\{2\}\.\(496\)Here, \(a\) applies the definition of theL2\(ν\)L^\{2\}\\left\(\\nu\\right\)norm, \(b\) uses\|f\(𝐱\)\|≤‖f‖∞\\left\|f\\left\(\\mathbf\{x\}\\right\)\\right\|\\leq\\left\\\|f\\right\\\|\_\{\\infty\}, and \(c\) usesν\(𝒳\)=1\\nu\\left\(\\mathcal\{X\}\\right\)=1\. Hence,JJis continuous\. Since𝒰L\\mathcal\{U\}\_\{L\}is compact inC\(𝒳\)C\\left\(\\mathcal\{X\}\\right\), its image underJJis compact inL2\(ν\)L^\{2\}\\left\(\\nu\\right\)\. Thus,𝒰L\\mathcal\{U\}\_\{L\}is compact when regarded as a subset ofL2\(ν\)L^\{2\}\\left\(\\nu\\right\)\. ∎
###### Lemma B0\.
\(Uniform Brownian profile interpolation estimate\)LetA\>0A\>0andm≥1m\\geq 1, let−A=t0<t1<⋯<tm=A\-A=t\_\{0\}<t\_\{1\}<\\cdots<t\_\{m\}=Abe the uniform grid on\[−A,A\]\\left\[\-A,A\\right\], and letΠmg\\Pi\_\{m\}gdenote the continuous piecewise\-linear interpolant ofggon this grid\. Then, for everyg∈ℋk\(B\)g\\in\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\},
sups∈\[−A,A\]\|g\(s\)−\(Πmg\)\(s\)\|\\displaystyle\\sup\_\{s\\in\\left\[\-A,A\\right\]\}\\left\|g\\left\(s\\right\)\-\\left\(\\Pi\_\{m\}g\\right\)\\left\(s\\right\)\\right\|≤\(A2m\)1/2‖g‖ℋk\(B\)\.\\displaystyle\\leq\\left\(\\frac\{A\}\{2m\}\\right\)^\{1/2\}\\left\\\|g\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}\.\(497\)
###### Proof\.
Seth:=2A/mh:=2A/m\. Fixs∈\[ti,ti\+1\]s\\in\\left\[t\_\{i\},t\_\{i\+1\}\\right\], and define
ϕs\(r\):=ti\+1−sh𝟏\[ti,s\]\(r\)−s−tih𝟏\[s,ti\+1\]\(r\)\.\\displaystyle\\phi\_\{s\}\\left\(r\\right\):=\\frac\{t\_\{i\+1\}\-s\}\{h\}\\mathbf\{1\}\_\{\\left\[t\_\{i\},s\\right\]\}\\left\(r\\right\)\-\\frac\{s\-t\_\{i\}\}\{h\}\\mathbf\{1\}\_\{\\left\[s,t\_\{i\+1\}\\right\]\}\\left\(r\\right\)\.\(498\)The interpolation formula and absolute continuity give
g\(s\)−\(Πmg\)\(s\)\\displaystyle g\\left\(s\\right\)\-\\left\(\\Pi\_\{m\}g\\right\)\\left\(s\\right\)=∫titi\+1g′\(r\)ϕs\(r\)𝑑r\.\\displaystyle=\\int\_\{t\_\{i\}\}^\{t\_\{i\+1\}\}g^\{\\prime\}\\left\(r\\right\)\\phi\_\{s\}\\left\(r\\right\)\\,\\mathrm\{d\}r\.\(499\)Moreover,
‖ϕs‖L2\(\[ti,ti\+1\]\)2\\displaystyle\\left\\\|\\phi\_\{s\}\\right\\\|\_\{L^\{2\}\\left\(\\left\[t\_\{i\},t\_\{i\+1\}\\right\]\\right\)\}^\{2\}=\(a\)\(s−ti\)\(ti\+1−s\)h≤\(b\)h4\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{=\}\}\\frac\{\\left\(s\-t\_\{i\}\\right\)\\left\(t\_\{i\+1\}\-s\\right\)\}\{h\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}\\frac\{h\}\{4\}\.\(500\)Here, \(a\) is direct integration and \(b\) uses\(s−ti\)\(ti\+1−s\)≤h2/4\\left\(s\-t\_\{i\}\\right\)\\left\(t\_\{i\+1\}\-s\\right\)\\leq h^\{2\}/4\. Therefore,
\|g\(s\)−\(Πmg\)\(s\)\|\\displaystyle\\left\|g\\left\(s\\right\)\-\\left\(\\Pi\_\{m\}g\\right\)\\left\(s\\right\)\\right\|≤\(a\)‖g′‖L2\(\[ti,ti\+1\]\)‖ϕs‖L2\(\[ti,ti\+1\]\)≤\(b\)h2‖g‖ℋk\(B\)=\(c\)\(A2m\)1/2‖g‖ℋk\(B\)\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}\\left\\\|g^\{\\prime\}\\right\\\|\_\{L^\{2\}\\left\(\\left\[t\_\{i\},t\_\{i\+1\}\\right\]\\right\)\}\\left\\\|\\phi\_\{s\}\\right\\\|\_\{L^\{2\}\\left\(\\left\[t\_\{i\},t\_\{i\+1\}\\right\]\\right\)\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}\\frac\{\\sqrt\{h\}\}\{2\}\\left\\\|g\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{=\}\}\\left\(\\frac\{A\}\{2m\}\\right\)^\{1/2\}\\left\\\|g\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}\.\(501\)Here, \(a\) is Cauchy–Schwarz, \(b\) uses the preceding norm bound and the Brownian RKHS norm characterization, and \(c\) substitutesh=2A/mh=2A/m\. Taking the supremum proves \([497](https://arxiv.org/html/2608.13882#A4.E497)\)\. ∎
###### Lemma B0\.
\(Uniform range bound for lower\-level atomic supports\)Assume that𝒳⊆ℝd\\mathcal\{X\}\\subseteq\\mathbb\{R\}^\{d\}is compact and that‖𝛚‖2≤1\\left\\\|\\bm\{\\omega\}\\right\\\|\_\{2\}\\leq 1for every𝛚∈Ω\\bm\{\\omega\}\\in\\Omega\. LetL≥2L\\geq 2, and setR𝒳:=sup𝐱∈𝒳‖𝐱‖2R\_\{\\mathcal\{X\}\}:=\\sup\_\{\\mathbf\{x\}\\in\\mathcal\{X\}\}\\left\\\|\\mathbf\{x\}\\right\\\|\_\{2\}\. Then everya∈𝒰L−1a\\in\\mathcal\{U\}\_\{L\-1\}satisfies\|a\(𝐱\)\|≤R𝒳2−\(L−2\),𝐱∈𝒳\\left\|a\\left\(\\mathbf\{x\}\\right\)\\right\|\\leq R\_\{\\mathcal\{X\}\}^\{\\,2^\{\-\(L\-2\)\}\},\\qquad\\mathbf\{x\}\\in\\mathcal\{X\}\. Consequently,a\(𝒳\)⊆\[−A0,A0\]a\\left\(\\mathcal\{X\}\\right\)\\subseteq\\left\[\-A\_\{0\},A\_\{0\}\\right\], whereA0:=R𝒳2−\(L−2\)A\_\{0\}:=R\_\{\\mathcal\{X\}\}^\{\\,2^\{\-\(L\-2\)\}\}\.
###### Proof\.
Compactness of𝒳\\mathcal\{X\}givesR𝒳<∞R\_\{\\mathcal\{X\}\}<\\infty\. Fixa∈𝒰L−1a\\in\\mathcal\{U\}\_\{L\-1\}\. IfL=2L=2, thena\(𝐱\)=𝝎⊤𝐱a\\left\(\\mathbf\{x\}\\right\)=\\bm\{\\omega\}^\{\\top\}\\mathbf\{x\}for some𝝎∈Ω\\bm\{\\omega\}\\in\\Omega, and
\|a\(𝐱\)\|\\displaystyle\\left\|a\\left\(\\mathbf\{x\}\\right\)\\right\|≤\(a\)‖𝝎‖2‖𝐱‖2≤\(b\)R𝒳\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}\\left\\\|\\bm\{\\omega\}\\right\\\|\_\{2\}\\left\\\|\\mathbf\{x\}\\right\\\|\_\{2\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}R\_\{\\mathcal\{X\}\}\.\(502\)Here, \(a\) is Cauchy–Schwarz and \(b\) uses‖𝝎‖2≤1\\left\\\|\\bm\{\\omega\}\\right\\\|\_\{2\}\\leq 1\. IfL≥3L\\geq 3,[Lemma26](https://arxiv.org/html/2608.13882#Thmtheorem26)applied at depthL−1L\-1gives
\|a\(𝐱\)\|\\displaystyle\\left\|a\\left\(\\mathbf\{x\}\\right\)\\right\|≤\(a\)‖𝐱‖22−\(L−2\)≤\(b\)R𝒳2−\(L−2\)\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}\\left\\\|\\mathbf\{x\}\\right\\\|\_\{2\}^\{\\,2^\{\-\(L\-2\)\}\}\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}R\_\{\\mathcal\{X\}\}^\{\\,2^\{\-\(L\-2\)\}\}\.\(503\)Here, \(a\) is the atomic pointwise estimate and \(b\) is the definition ofR𝒳R\_\{\\mathcal\{X\}\}\. Since2−\(L−2\)=12^\{\-\(L\-2\)\}=1whenL=2L=2, both cases prove the stated bound, and the range inclusion follows immediately\. ∎
###### Lemma B0\.
\(Stability of Brownian profile interpolation under composition\)LetA\>0A\>0andm∈ℕm\\in\\mathbb\{N\},m≥1m\\geq 1\. Let−A=t0<t1<⋯<tm=A\-A=t\_\{0\}<t\_\{1\}<\\cdots<t\_\{m\}=Abe the uniform grid on\[−A,A\]\\left\[\-A,A\\right\], and letΠm\\Pi\_\{m\}denote the associated continuous piecewise\-linear interpolation operator\. Leta:𝒳⟶\[−A,A\]a:\\mathcal\{X\}\\longrightarrow\\left\[\-A,A\\right\]be measurable, and letg∈ℋk\(B\)g\\in\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\. Defineu\(𝐱\):=g\(a\(𝐱\)\)u\\left\(\\mathbf\{x\}\\right\):=g\\left\(a\\left\(\\mathbf\{x\}\\right\)\\right\),um\(𝐱\):=\(Πmg\)\(a\(𝐱\)\)u\_\{m\}\\left\(\\mathbf\{x\}\\right\):=\\left\(\\Pi\_\{m\}g\\right\)\\left\(a\\left\(\\mathbf\{x\}\\right\)\\right\),𝐱∈𝒳\.\\mathbf\{x\}\\in\\mathcal\{X\}\.Then
‖u−um‖L2\(ν\)≤\(A2\)1/2m−1/2‖g‖ℋk\(B\)\.\\displaystyle\\left\\\|u\-u\_\{m\}\\right\\\|\_\{L^\{2\}\\left\(\\nu\\right\)\}\\leq\\left\(\\frac\{A\}\{2\}\\right\)^\{1/2\}m^\{\-1/2\}\\left\\\|g\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}\.\(504\)
###### Proof\.
For every𝐱∈𝒳\\mathbf\{x\}\\in\\mathcal\{X\}, the range assumption onaaand[Lemma30](https://arxiv.org/html/2608.13882#Thmtheorem30)give
\|u\(𝐱\)−um\(𝐱\)\|\\displaystyle\\left\|u\\left\(\\mathbf\{x\}\\right\)\-u\_\{m\}\\left\(\\mathbf\{x\}\\right\)\\right\|≤\(a\)\(A2m\)1/2‖g‖ℋk\(B\)\.\\displaystyle\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}\\left\(\\frac\{A\}\{2m\}\\right\)^\{1/2\}\\left\\\|g\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}\.\(505\)Here, \(a\) applies \([497](https://arxiv.org/html/2608.13882#A4.E497)\) ats=a\(𝐱\)s=a\\left\(\\mathbf\{x\}\\right\)\. Squaring, integrating, and usingν\(𝒳\)=1\\nu\\left\(\\mathcal\{X\}\\right\)=1yields
‖u−um‖L2\(ν\)\\displaystyle\\left\\\|u\-u\_\{m\}\\right\\\|\_\{L^\{2\}\\left\(\\nu\\right\)\}≤\(A2m\)1/2‖g‖ℋk\(B\)=\(A2\)1/2m−1/2‖g‖ℋk\(B\),\\displaystyle\\leq\\left\(\\frac\{A\}\{2m\}\\right\)^\{1/2\}\\left\\\|g\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\}=\\left\(\\frac\{A\}\{2\}\\right\)^\{1/2\}m^\{\-1/2\}\\left\\\|g\\right\\\|\_\{\\mathcal\{H\}\_\{k^\{\(\\mathrm\{B\}\)\}\}\},\(506\)which proves \([504](https://arxiv.org/html/2608.13882#A4.E504)\)\. ∎
## References
- N\. AronszajnTheory of reproducing kernels\.Transactions of the American Mathematical Society68\(3\),pp\. 337–404\.Cited by:[Remark 5](https://arxiv.org/html/2608.13882#Thmtheorem5.p1.1.1)\.
- Bach \(2017\)F\. BachBreaking the curse of dimensionality with convex neural networks\.Journal of Machine Learning Research18\(19\),pp\. 1–53\.Cited by:[§1](https://arxiv.org/html/2608.13882#S1.p2.1),[item 2](https://arxiv.org/html/2608.13882#S2.I1.i2.p1.1)\.
- Barron and Klusowski \(2019\)A\. R\. Barron and J\. M\. KlusowskiComplexity, statistical risk, and metric entropy of deep nets using total path variation\.arXiv preprint arXiv:1902\.00800\.Cited by:[§1](https://arxiv.org/html/2608.13882#S1.p1.1)\.
- Barron \(1993\)A\. R\. BarronUniversal approximation bounds for superpositions of a sigmoidal function\.IEEE Transactions on Information Theory39\(3\),pp\. 930–945\.Cited by:[§1](https://arxiv.org/html/2608.13882#S1.p2.1),[item 2](https://arxiv.org/html/2608.13882#S2.I1.i2.p1.1)\.
- Bartlett \(1996\)P\. L\. BartlettFor valid generalization the size of the weights is more important than the size of the network\.InAdvances in Neural Information Processing Systems,Vol\.9\.Cited by:[§1](https://arxiv.org/html/2608.13882#S1.p1.1)\.
- Bartolucciet al\.\(2023\)F\. Bartolucci, E\. D\. Vito, L\. Rosasco, and S\. VigognaUnderstanding neural networks with reproducing kernel banach spaces\.Applied and Computational Harmonic Analysis62,pp\. 194–236\.Cited by:[§1](https://arxiv.org/html/2608.13882#S1.p2.1)\.
- Bartolucciet al\.\(2024\)F\. Bartolucci, E\. D\. Vito, L\. Rosasco, and S\. VigognaNeural reproducing kernel banach spaces and representer theorems for deep networks\.arXiv preprint arXiv:2403\.08750\.Cited by:[§1](https://arxiv.org/html/2608.13882#S1.p2.1)\.
- Chen \(2024\)Z\. ChenNeural hilbert ladders: multi\-layer neural networks in function space\.Journal of Machine Learning Research25\(109\),pp\. 1–65\.Cited by:[§1](https://arxiv.org/html/2608.13882#S1.p2.1)\.
- DeVore \(1998\)R\. A\. DeVoreNonlinear approximation\.Acta Numerica7,pp\. 51–150\.Cited by:[§1](https://arxiv.org/html/2608.13882#S1.SS0.SSS0.Px4.p1.1)\.
- E and Wojtowytsch \(2020\)W\. E and S\. WojtowytschOn the banach spaces associated with multi\-layer ReLU networks: function representation, approximation theory and gradient descent dynamics\.CSIAM Transactions on Applied Mathematics1\(3\),pp\. 387–440\.Cited by:[§1](https://arxiv.org/html/2608.13882#S1.p2.1)\.
- Eldan and Shamir \(2016\)R\. Eldan and O\. ShamirThe power of depth for feedforward neural networks\.InConference on Learning Theory,pp\. 907–940\.Cited by:[§1](https://arxiv.org/html/2608.13882#S1.SS0.SSS0.Px4.p1.1)\.
- Follain and Bach \(2025\)B\. Follain and F\. BachEnhanced feature learning via regularisation: integrating neural networks and kernel methods\.Journal of Machine Learning Research26\(172\),pp\. 1–56\.Cited by:[§1](https://arxiv.org/html/2608.13882#S1.SS0.SSS0.Px2.p1.1)\.
- Goldberg and Jerrum \(1995\)P\. W\. Goldberg and M\. JerrumBounding the vapnik–chervonenkis dimension of concept classes parameterized by real numbers\.Machine Learning18\(2–3\),pp\. 131–148\.External Links:[Document](https://dx.doi.org/10.1007/BF00993408)Cited by:[Appendix D](https://arxiv.org/html/2608.13882#A4.SS0.SSS0.Px7.p8.3)\.
- Golowichet al\.\(2018\)N\. Golowich, A\. Rakhlin, and O\. ShamirSize\-independent sample complexity of neural networks\.InConference on Learning Theory,pp\. 297–299\.Cited by:[§1](https://arxiv.org/html/2608.13882#S1.p1.1)\.
- Heeringaet al\.\(2025\)T\. J\. Heeringa, L\. Spek, and C\. BruneDeep networks are reproducing kernel chains\.arXiv preprint arXiv:2501\.03697\.Cited by:[§1](https://arxiv.org/html/2608.13882#S1.p2.1)\.
- Kurková and Sanguineti \(2002\)V\. Kurková and M\. SanguinetiBounds on rates of variable\-basis and neural\-network approximation\.IEEE Transactions on Information Theory47\(6\),pp\. 2659–2666\.Cited by:[§1](https://arxiv.org/html/2608.13882#S1.SS0.SSS0.Px4.p1.1)\.
- Maet al\.\(2022\)C\. Ma L\. Wuet al\.The barron space and the flow\-induced function spaces for neural network models\.Constructive Approximation55\(1\),pp\. 369–406\.Cited by:[§1](https://arxiv.org/html/2608.13882#S1.p2.1)\.
- Mohammadigohariet al\.\(2026\)M\. Mohammadigohari, G\. Di Fatta, G\. Nicosia, and P\. M\. PardalosBrownian kernel ladders\.Note:Revised preprintExternal Links:2606\.15812,[Link](https://arxiv.org/abs/2606.15812)Cited by:[Appendix D](https://arxiv.org/html/2608.13882#A4.SS0.SSS0.Px3.p2.2.1),[§1](https://arxiv.org/html/2608.13882#S1.SS0.SSS0.Px2.p1.1),[§2\.1](https://arxiv.org/html/2608.13882#S2.SS1.p1.1),[§7\.3](https://arxiv.org/html/2608.13882#S7.SS3.p1.15),[§7\.3](https://arxiv.org/html/2608.13882#S7.SS3.p1.5)\.
- Nakhleh and Nowak \(2026\)J\. Nakhleh and R\. D\. NowakDeep neural variation spaces: a unifying perspective on depth and complexity\.External Links:2607\.05546,[Link](https://arxiv.org/abs/2607.05546)Cited by:[§1](https://arxiv.org/html/2608.13882#S1.SS0.SSS0.Px3.p1.1),[§1](https://arxiv.org/html/2608.13882#S1.p2.1),[item 2](https://arxiv.org/html/2608.13882#S2.I1.i2.p1.1)\.
- Neyshaburet al\.\(2015\)B\. Neyshabur, R\. Tomioka, and N\. SrebroNorm\-based capacity control in neural networks\.InConference on Learning Theory,pp\. 1376–1401\.Cited by:[§1](https://arxiv.org/html/2608.13882#S1.p1.1)\.
- Ongie and Parhi \(2026\)G\. Ongie and R\. ParhiRepresentation costs in data science: foundations and the quasi\-banach spaces of deep neural networks\.arXiv preprint arXiv:2606\.14954\.Cited by:[§1](https://arxiv.org/html/2608.13882#S1.p2.1)\.
- Ongieet al\.\(2020\)G\. Ongie, R\. Willett, D\. Soudry, and N\. SrebroA function space view of bounded norm infinite width relu nets: the multivariate case\.InInternational Conference on Learning Representations,Note:Originally available as arXiv:1907\.05681 \(2019\)Cited by:[§1](https://arxiv.org/html/2608.13882#S1.p2.1)\.
- Parhi and Nowak \(2021\)R\. Parhi and R\. D\. NowakBanach space representer theorems for neural networks and ridge splines\.Journal of Machine Learning Research22\(43\),pp\. 1–40\.Cited by:[§1](https://arxiv.org/html/2608.13882#S1.p2.1),[item 2](https://arxiv.org/html/2608.13882#S2.I1.i2.p1.1)\.
- Parhi and Nowak \(2022a\)R\. Parhi and R\. D\. NowakNear\-minimax optimal estimation with shallow relu neural networks\.IEEE Transactions on Information Theory69\(2\),pp\. 1125–1140\.Cited by:[§1](https://arxiv.org/html/2608.13882#S1.p2.1)\.
- Parhi and Nowak \(2022b\)R\. Parhi and R\. D\. NowakWhat kinds of functions do deep neural networks learn? insights from variational spline theory\.SIAM Journal on Mathematics of Data Science4\(2\),pp\. 464–489\.Cited by:[§1](https://arxiv.org/html/2608.13882#S1.p2.1)\.
- Parkinsonet al\.\(2024\)S\. Parkinson, G\. Ongie, R\. Willett, O\. Shamir, and N\. SrebroDepth separation in norm\-bounded infinite\-width neural networks\.InProceedings of the 37th Annual Conference on Learning Theory,pp\. 4082–4114\.Cited by:[§1](https://arxiv.org/html/2608.13882#S1.SS0.SSS0.Px4.p1.1)\.
- Savareseet al\.\(2019\)P\. Savarese, I\. Evron, D\. Soudry, and N\. SrebroHow do infinite width bounded norm networks look in function space?\.InConference on Learning Theory,pp\. 2667–2690\.Cited by:[§1](https://arxiv.org/html/2608.13882#S1.p2.1)\.
- Shenoudaet al\.\(2024\)J\. Shenouda, R\. Parhi, K\. Lee, and R\. D\. NowakVariation spaces for multi\-output neural networks: insights on multi\-task learning and network compression\.Journal of Machine Learning Research25\(231\),pp\. 1–40\.Cited by:[§1](https://arxiv.org/html/2608.13882#S1.p2.1)\.
- Siegel and Xu \(2023\)J\. W\. Siegel and J\. XuCharacterization of the variation spaces corresponding to shallow neural networks\.Constructive Approximation57\(3\),pp\. 1109–1132\.Cited by:[§1](https://arxiv.org/html/2608.13882#S1.p2.1),[item 2](https://arxiv.org/html/2608.13882#S2.I1.i2.p1.1)\.
- Siegel and Xu \(2024\)J\. W\. Siegel and J\. XuSharp bounds on the approximation rates, metric entropy, and n\-widths of shallow neural networks\.Foundations of Computational Mathematics24\(2\),pp\. 481–537\.Cited by:[§1](https://arxiv.org/html/2608.13882#S1.SS0.SSS0.Px4.p1.1),[§1](https://arxiv.org/html/2608.13882#S1.p2.1)\.
- Telgarsky \(2016\)M\. TelgarskyBenefits of depth in neural networks\.InConference on Learning Theory,pp\. 1517–1539\.Cited by:[§1](https://arxiv.org/html/2608.13882#S1.SS0.SSS0.Px4.p1.1)\.
- Temlyakov \(2011\)V\. N\. TemlyakovGreedy approximation\.Cambridge University Press\.Cited by:[§1](https://arxiv.org/html/2608.13882#S1.SS0.SSS0.Px4.p1.1)\.
- Vardi and Shamir \(2020\)G\. Vardi and O\. ShamirNeural networks with small weights and depth\-separation barriers\.InAdvances in Neural Information Processing Systems,Vol\.33,pp\. 19433–19442\.Cited by:[§1](https://arxiv.org/html/2608.13882#S1.SS0.SSS0.Px4.p1.1)\.
- Venturiet al\.\(2022\)L\. Venturi, S\. Jelassi, T\. Ozuch, and J\. BrunaDepth separation beyond radial functions\.Journal of Machine Learning Research23\(122\),pp\. 1–56\.Cited by:[§1](https://arxiv.org/html/2608.13882#S1.SS0.SSS0.Px4.p1.1)\.Similar Articles
The interesting BDH question: What if LLM memory lived in the network weights instead of the ever-growing KV cache?
This article analyzes Jan Chorowski's BDH architecture proposal, which explores embedding LLM memory directly into network weights using sparse high-dimensional key-query spaces as an alternative to traditional KV caches.
Data-Driven Variational Basis Learning Beyond Neural Networks: A Non-Neural Framework for Adaptive Basis Discovery
This paper introduces Data-Driven Variational Basis Learning (DVBL), a non-neural framework that learns basis functions directly from data through variational optimization, offering interpretability and mathematical transparency compared to neural networks.
Automated Kernel Discovery Towards Understanding High-dimensional Bayesian Optimization
The paper introduces Kernel Discovery, an LLM-driven evolutionary framework for high-dimensional Bayesian optimization that searches a broader kernel space and achieves state-of-the-art results on benchmarks.
BODHI: Do LLMs Branch Out and Discover Heterogeneous Inferences?
This paper investigates whether RLVR-trained LLMs branch out to discover heterogeneous inferences, using maze-solving experiments and BODHI-Trees to show that policy entropy collapse is accompanied by reduced semantic branching entropy, limiting rollout diversity.
ALAS: Additive Learnable Alpha-Stable Kernels for Flexible Bayesian Optimization
This paper introduces ALAS, a flexible Gaussian Process kernel family that learns the stability parameter from data to adapt smoothness, capturing both smooth trends and sharp irregularities, with a separable variant for higher dimensions and theoretical guarantees on information gain.