When Is the Sharp Covariance Envelope Tight? Feature-Only Geometry for Volume-Sampled Least Squares
Summary
This paper establishes a sharp Loewner envelope for coefficient covariance under volume-sampled least squares, using feature-only geometry to analyze tightness conditions.
View Cached Full Text
Cached at: 08/28/26, 09:45 AM
# When Is the Sharp Covariance Envelope Tight? Feature-Only Geometry for Volume-Sampled Least Squares
Source: [https://arxiv.org/html/2608.26877](https://arxiv.org/html/2608.26877)
###### Abstract
Prior analyses by Dereziński and Warmuth established all\-size sampling identities, selected\-OLS unbiasedness, and inverse moments for ordinary volume sampling, while their exact arbitrary\-fixed\-response loss and prediction\-covariance formulas are at the rank\-size endpoints=ds=d\. We establish a Loewner envelope for centered coefficient covariance for every full\-rank fixed pool, response, and legal budgetd≤s≤md\\leq s\\leq munder ordinary indexed fixed\-size volume sampling followed by selected unweighted least squares; its coefficient is globally sharp over the full\-rank class\. Global sharpness does not determine attainability on the pool in hand\. Under positive loss, strict\-interior budgets, and no coloops, a feature\-only marginνA\\nu\_\{A\}gives the exact fixed\-design spectral phase:νA\>0\\nu\_\{A\}\>0if and only if the normalized spectral envelope is strict for every compatible residual, whereasνA=0\\nu\_\{A\}=0if and only if some compatible residual is spectrally tight; the same zero\-margin residual is tight at every strict\-interior budget\. A residual\-augmented change of measure supplies the response\-aware mechanism and a one\-sided quantitative slack bound, while support saturation proves the attainment direction\. Critical equal\-leverage geometry interprets the boundary, and sound lower certificates yield conservative same\-primitive cardinality decisions\. Frozen\-feature examples show that the certificate is nonvacuous and measure the fixed\-pool cost of its authorized reduction\. The claims concern conditional centered, full\-Gram\-whitened coefficient covariance, not population generalization\.
## 1Introduction
##### Frozen\-feature readouts and directional refitting\.
A frozen\-feature linear readout is a standard downstream interface: representation\-learning papers evaluate a fixed encoder by fitting a linear classifier on top\[[8](https://arxiv.org/html/2608.26877#bib.bib8),[19](https://arxiv.org/html/2608.26877#bib.bib19)\], and few\-shot workflows fit that readout from a finite labelled support set\[[26](https://arxiv.org/html/2608.26877#bib.bib26)\]\. The composition of the support set can affect the resulting classifier, as reflected by the sample\-selection\-bias literature\[[28](https://arxiv.org/html/2608.26877#bib.bib28)\]\. Even after the representation, pool, and response are held fixed, randomized subset refitting creates variation in the fitted coefficients\. We isolate the least\-squares readout primitive in this interface; we do not analyze classification loss or representation quality\. A scalar same\-pool loss summarizes aggregate refitting variation but does not reveal whether that variation is diffuse or concentrated along a single coefficient direction\. We therefore study the most variable direction of a refitted linear head\.
##### The fixed\-pool, pre\-response question\.
We study a single deterministic finite regression pool\. For each fixed response, the chosen subset is the only source of randomness; we then maximize the conditional covariance over the compatible residual sphere\. The resulting worst\-residual margin depends only on the feature matrix, not on an observed response\. The controlled primitive is ordinary indexed fixed\-size volume sampling followed by selected unweighted least squares\. Its centered, full\-Gram\-whitened coefficient covariance retains the directional information hidden by a trace\.
##### From the rank\-size boundary to a sharp all\-budget envelope\.
Ordinary volume sampling already has determinant laws, all\-size selected\-OLS unbiasedness, and inverse\-Gram moment identities\. For arbitrary fixed responses, however, Dereziński and Warmuth’s exact loss and prediction\-covariance formulas are at the rank\-size endpoints=ds=d; their NeurIPS paper explicitly left expected\-loss bounds fors\>ds\>dopen, and the JMLR article reported no extension of its covariance formula beyonds=ds=d\[[11](https://arxiv.org/html/2608.26877#bib.bib11),[12](https://arxiv.org/html/2608.26877#bib.bib12)\]\. A later ordinary\-volume lower construction exhibits the same budget dependence on a particular nonuniform\-leverage family\[[13](https://arxiv.org/html/2608.26877#bib.bib13)\]\. Theorem 1 addresses this recorded larger\-budget boundary with a sharp Loewner envelope for every fixed pool, response, and legal budget\.
##### The fixed\-design phase left by a global envelope\.
Global sharpness is a class\-level statement: a coefficient can be unimprovable over all designs and responses while remaining unattainable on the particular feature pool in hand\. Our Theorem 2 exposes a contraction for one realized residual, but neither result settles the fixed\-design quantifier: is the ceiling strict for every compatible residual, or is it attained by some compatible residual? Theorem 3 answers this question from feature geometry alone\. Under the conditions below, two eligible feature pools obeying the same globally sharp benchmark can lie in different design\-specific covariance\-envelope attainability phases: a zero\-margin pool admits one compatible positive\-loss residual that attains the normalized spectral coefficient\-covariance envelope, whereas a positive\-margin pool keeps every compatible residual below it by a uniform positive gap\.
##### The answer: a geometric phase boundary\.
The marginνA\\nu\_\{A\}is the weakest contraction enforced by feature geometry over all compatible residual directions\. After whitening the design toA⊤A=IdA^\{\\top\}A=I\_\{d\}, let
𝒵A=\{z∈ker\(A⊤\):∥z∥2=1\},RA\(z\)=∑i=1m\(1−∥ai∥22−zi2\)aiai⊤,νA=minz∈𝒵Aλmin\(RA\(z\)\),\\mathcal\{Z\}\_\{A\}=\\\{z\\in\\ker\(A^\{\\top\}\):\\lVert z\\rVert\_\{2\}=1\\\},\\qquad R\_\{A\}\(z\)=\\sum\_\{i=1\}^\{m\}\\bigl\(1\-\\lVert a\_\{i\}\\rVert\_\{2\}^\{2\}\-z\_\{i\}^\{2\}\\bigr\)a\_\{i\}a\_\{i\}^\{\\top\},\\qquad\\nu\_\{A\}=\\min\_\{z\\in\\mathcal\{Z\}\_\{A\}\}\\lambda\_\{\\min\}\(R\_\{A\}\(z\)\),whereai⊤a\_\{i\}^\{\\top\}is rowiiofAA\. The margin is defined entirely by feature geometry and ranges over the compatible normalized residual sphere\. Let𝖳specRU\(A,s\)\\mathsf\{T\}\_\{\\rm spec\}^\{\\rm RU\}\(A,s\)denote the maximum normalized spectral coefficient covariance over that same sphere\. Under positive loss,d<s<md<s<m,m≥d\+2m\\geq d\+2, and no\-coloop leverage∥ai∥22<1\\lVert a\_\{i\}\\rVert\_\{2\}^\{2\}<1for every row, the robust phase result is
νA\>0⟺𝖳specRU\(A,s\)<αs,νA=0⟺𝖳specRU\(A,s\)=αs,αs=m−sm−d\.\\nu\_\{A\}\>0\\ \\Longleftrightarrow\\ \\mathsf\{T\}\_\{\\rm spec\}^\{\\rm RU\}\(A,s\)<\\alpha\_\{s\},\\qquad\\nu\_\{A\}=0\\ \\Longleftrightarrow\\ \\mathsf\{T\}\_\{\\rm spec\}^\{\\rm RU\}\(A,s\)=\\alpha\_\{s\},\\qquad\\alpha\_\{s\}=\\frac\{m\-s\}\{m\-d\}\.The positive branch follows from the response\-aware contraction, but the zero branch is not formal: a vanishing one\-sided slack bound need not imply equality\. Here zero margin forces support saturation under the ordinary volume law, producing one compatible residual that attains the ceiling for every strict\-interior subset size; the phase classifies uniform strictness versus existential tightness, not the positive slack magnitude\.
##### Result hierarchy\.
1. \(i\)*Headline fixed\-design covariance\-envelope phase \(Theorem 3\)\.*The feature\-only marginνA\\nu\_\{A\}exactly separates uniform strictness for every compatible residual from attainment by some compatible residual, with one common zero\-margin witness across strict\-interior budgets\. Its quantitative slack statement is one\-sided: it is lower\-bounded by a positive multiple ofνA\\nu\_\{A\}, not identified exactly by it\.
2. \(ii\)*Globally sharp all\-budget benchmark and response\-aware mechanism \(Theorems 1–2\)\.*Every fixed pool and response obeys M¯s⪯m−sm−dL∗Id,\\overline\{M\}\_\{s\}\\preceq\\frac\{m\-s\}\{m\-d\}L^\{\*\}I\_\{d\},\(1\)and the coefficient is sharp over the full\-rank class\. Residual augmentation gives an exact interior representation and a boundary\-valid one\-sided resolvent\.
3. \(iii\)*Critical interpretation and sound consequence \(Theorem 4 and Corollary 5\)\.*In the critical equal\-leverage class, the classical projective repeated\-pair boundary is the zero set of the continuous residual\-direction margin, with two\-sided coherence control\. Sound lower certificates give conservative cardinality decisions; the fixed\-profile instantiation is supplementary\.
We finally evaluate the certificate/action layer on frozen public feature pools, recording nonvacuous decisions and their measured fixed\-pool cost\. Figure[1](https://arxiv.org/html/2608.26877#S1.F1)summarizes the structural inputs, exact phase, and conservative action\.
Globally sharp all\-budget Loewner benchmark \(Theorem 1\)M¯s/L∗⪯αsId\\overline\{M\}\_\{s\}/L^\{\*\}\\preceq\\alpha\_\{s\}I\_\{d\}sharp benchmarkResponse\-aware mechanism and resolvent \(Theorem 2\)auxiliary lawPBP\_\{B\}; response\-aware slacksampling and selected OLS remain underPXP\_\{X\}
structural inputs
↘↙\\searrow\\hskip 110\.40253pt\\swarrow
Headline fixed\-design phase \(Theorem 3\)νA\>0⟺𝖳specRU<αs\\nu\_\{A\}\>0\\Longleftrightarrow\\mathsf\{T\}\_\{\\rm spec\}^\{\\rm RU\}<\\alpha\_\{s\}νA=0⟺𝖳specRU=αs\\nu\_\{A\}=0\\Longleftrightarrow\\mathsf\{T\}\_\{\\rm spec\}^\{\\rm RU\}=\\alpha\_\{s\}positive magnitude is one\-sided; zero witness is existential
↗\\nearrowcertifies positive branch
Computable lower certificatet\(A\)\>0⟹νA\>0t\(A\)\>0\\ \\Longrightarrow\\ \\nu\_\{A\}\>0sufficient only⇢\\dashrightarrowConservative cardinality actioncertified ceiling selectssssame primitive; cardinality changes
Figure 1:Theorem[3](https://arxiv.org/html/2608.26877#Thmtheorem3)is central: Theorems 1–2 supply its benchmark and response\-aware mechanism, while a computable lower certificate turns only the positive branch into a conservative cardinality action\.
## 2Setting, metric, and scope
LetX∈ℝm×dX\\in\\mathbb\{R\}^\{m\\times d\}have full column rank and lety∈ℝmy\\in\\mathbb\{R\}^\{m\}be an arbitrary deterministic fixed response\. Define
G=X⊤X,w∗=G−1X⊤y,e=y−Xw∗,L∗=∥e∥22\.G=X^\{\\top\}X,\\qquad w^\{\*\}=G^\{\-1\}X^\{\\top\}y,\\qquad e=y\-Xw^\{\*\},\\qquad L^\{\*\}=\\lVert e\\rVert\_\{2\}^\{2\}\.\(2\)For an unordered subset of indexed rowsS⊆\[m\]S\\subseteq\[m\],\|S\|=s\|S\|=s, write
GS=XS⊤XS,DS=det\(GS\)\.G\_\{S\}=X\_\{S\}^\{\\top\}X\_\{S\},\\qquad D\_\{S\}=\\det\(G\_\{S\}\)\.\(3\)Repeated rows at different indices remain different observations\. Form\>dm\>dandd≤s≤md\\leq s\\leq m, ordinary unrescaled fixed\-size volume sampling is
ZX,s:=∑\|U\|=sDU=\(m−ds−d\)det\(G\),ℙX\(S\)=DSZX,s\.Z\_\{X,s\}:=\\sum\_\{\|U\|=s\}D\_\{U\}=\\binom\{m\-d\}\{s\-d\}\\det\(G\),\\qquad\\mathbb\{P\}\_\{X\}\(S\)=\\frac\{D\_\{S\}\}\{Z\_\{X,s\}\}\.\(4\)This determinant law and its normalizer are standard components of volume sampling\[[11](https://arxiv.org/html/2608.26877#bib.bib11),[12](https://arxiv.org/html/2608.26877#bib.bib12),[3](https://arxiv.org/html/2608.26877#bib.bib3),[22](https://arxiv.org/html/2608.26877#bib.bib22)\]\. OnlyDS\>0D\_\{S\}\>0sets support an estimator\. On such a set, use unweighted selected least squares
wS=GS−1XS⊤yS,LS=∥XSwS−yS∥22\.w\_\{S\}=G\_\{S\}^\{\-1\}X\_\{S\}^\{\\top\}y\_\{S\},\\qquad L\_\{S\}=\\lVert X\_\{S\}w\_\{S\}\-y\_\{S\}\\rVert\_\{2\}^\{2\}\.\(5\)Zero\-volume sets have zero probability and receive no pseudoinverse, ridge term, rescaling, weighting, replacement rule, or selected fit\.
The centered second moment and its invariant form are
Ms=𝔼X\[\(wS−w∗\)\(wS−w∗\)⊤\],M¯s=G1/2MsG1/2\.M\_\{s\}=\\mathbb\{E\}\_\{X\}\[\(w\_\{S\}\-w^\{\*\}\)\(w\_\{S\}\-w^\{\*\}\)^\{\\top\}\],\\qquad\\overline\{M\}\_\{s\}=G^\{1/2\}M\_\{s\}G^\{1/2\}\.\(6\)Form\>dm\>d,
α=m−sm−d,β=s−dm−d=1−α\.\\alpha=\\frac\{m\-s\}\{m\-d\},\\qquad\\beta=\\frac\{s\-d\}\{m\-d\}=1\-\\alpha\.\(7\)WhenαL∗\>0\\alpha L^\{\*\}\>0, define the exact normalized directional quantityqs:=λmax\(M¯s\)/\(αL∗\)q\_\{s\}:=\\lambda\_\{\\max\}\(\\overline\{M\}\_\{s\}\)/\(\\alpha L^\{\*\}\)\. This is not the raw Euclidean eigenvalue ofMsM\_\{s\}; it measures only the conditional fixed\-pool quantity defined above\. No directional quotient is formed at a zero denominator\.
### 2\.1Endpoint and support conventions
These branches are applied before any normalized ratio or residual direction is formed\. Ifm=dm=d, thens=ms=m, the selected fit is the full fit and both losses and the covariance are zero; no expression containing\(m−d\)−1\(m\-d\)^\{\-1\}is evaluated\. Ifm\>dm\>dands=ms=m, the unique sample givesMs=0M\_\{s\}=0andα=0\\alpha=0, so no directional quotient is formed\. IfL∗=0L^\{\*\}=0, every supported selected fit equalsw∗w^\{\*\}andMs=0M\_\{s\}=0; no residual direction is formed\. Ifs=d<ms=d<m, supported square systems interpolate and the rank\-\(d\+1\)\(d\+1\)augmentation used later is not invoked\. Every expectation is conditional on the fixed pair\(X,y\)\(X,y\)\.
The full\-pool loss identity used below is
∥Xw−y∥22=L∗\+\(w−w∗\)⊤G\(w−w∗\)\(w∈ℝd\)\.\\lVert Xw\-y\\rVert\_\{2\}^\{2\}=L^\{\*\}\+\(w\-w^\{\*\}\)^\{\\top\}G\(w\-w^\{\*\}\)\\qquad\(w\\in\\mathbb\{R\}^\{d\}\)\.\(8\)
## 3Universal budget envelope
Theorem[1](https://arxiv.org/html/2608.26877#Thmtheorem1)establishes the design\-independent benchmark: for every fixed pool, response, and legal budget, it bounds the centered whitened coefficient covariance in Loewner order before taking any trace\. Its proof uses the indexed Cauchy–Binet normalizer, rank\-size basis moments, uniform padding, and the selected\-residual first moment\. Here and below,*row general position*means that every indexed set ofddrows is nonsingular\. A*full\-column\-rank boundary design*has full column rank but is outside row general position, so at least one indexeddd\-row subset is singular\. For symmetric matrices,A⪯BA\\preceq Bdenotes the Loewner \(positive\-semidefinite\) order:B−AB\-Ais positive semidefinite\.
###### Theorem 1\(Universal budget envelope\)\.
For every full\-column\-rankXX, every fixedyy,m\>dm\>d, andd≤s≤md\\leq s\\leq m, ordinary sampling \([4](https://arxiv.org/html/2608.26877#S2.E4)\) followed by \([5](https://arxiv.org/html/2608.26877#S2.E5)\) satisfies
𝔼XwS=w∗,Ms⪯αL∗G−1,M¯s⪯αL∗Id\.\\mathbb\{E\}\_\{X\}w\_\{S\}=w^\{\*\},\\qquad M\_\{s\}\\preceq\\alpha L^\{\*\}G^\{\-1\},\\qquad\\overline\{M\}\_\{s\}\\preceq\\alpha L^\{\*\}I\_\{d\}\.\(9\)Consequently,
𝔼X∥XwS−y∥22≤\(1\+dα\)L∗=\(1\+d\(m−s\)m−d\)L∗\.\\mathbb\{E\}\_\{X\}\\lVert Xw\_\{S\}\-y\\rVert\_\{2\}^\{2\}\\leq\(1\+d\\alpha\)L^\{\*\}=\\left\(1\+\\frac\{d\(m\-s\)\}\{m\-d\}\\right\)L^\{\*\}\.\(10\)The endpoint and zero\-volume support conventions in Section[2\.1](https://arxiv.org/html/2608.26877#S2.SS1)apply before any ratio or residual direction is formed\. The coefficient is attained by the full\-column\-rank core\-plus\-zero family in Proposition[14](https://arxiv.org/html/2608.26877#Thmtheorem14); this is a coefficient\-attainment statement, not an equality classification\. For positive\-loss row\-general\-position designs withd≥2d\\geq 2andd<s<md<s<m, the matrix inequality is strict, while the same normalized coefficient is a perturbative supremum in that interior\.
##### Proof overview: interior slack, then boundary\.
On the positive\-loss row\-general\-position interior withd<s<md<s<m, couple a size\-ssvolume sample to a volume\-sampled rank\-size basis plus uniform padding\. The matrix law of total covariance and the selected\-residual first moment give
αL∗G−1−Ms=𝔼X\[LS\(GS−1−G−1\)\]⪰0\.\\alpha L^\{\*\}G^\{\-1\}\-M\_\{s\}=\\mathbb\{E\}\_\{X\}\\\!\\left\[L\_\{S\}\\bigl\(G\_\{S\}^\{\-1\}\-G^\{\-1\}\\bigr\)\\right\]\\succeq 0\.\(11\)This exact slack identity is an interior statement\. For an arbitrary full\-column\-rank boundary design, a determinant\-weighted directional limit retains the one\-sided envelope without continuing this decomposition or assigning an inverse to a singular subset\. The attainment and strictness arguments are in Appendix[C](https://arxiv.org/html/2608.26877#A3)\.
For every coefficient contrastcc, \([9](https://arxiv.org/html/2608.26877#S3.E9)\) bounds its subset\-induced fluctuation byαL∗c⊤G−1c\\alpha L^\{\*\}c^\{\\top\}G^\{\-1\}c\. Taking the trace through the full\-pool Pythagoras identity yields \([10](https://arxiv.org/html/2608.26877#S3.E10)\); this is a conditional same\-pool statement\. Thus the matrix inequality controls every contrast simultaneously, before taking the trace\. The budget coefficientα=\(m−s\)/\(m−d\)\\alpha=\(m\-s\)/\(m\-d\)is attained on the stated boundary family, is strictly slack on the regular positive\-loss interior, and remains the perturbative interior supremum; “globally sharp” refers to this coefficient statement\. The ordinary law and its inverse\-Gram ingredients are standard components of volume sampling\[[11](https://arxiv.org/html/2608.26877#bib.bib11),[12](https://arxiv.org/html/2608.26877#bib.bib12),[3](https://arxiv.org/html/2608.26877#bib.bib3),[22](https://arxiv.org/html/2608.26877#bib.bib22)\]\.
## 4Residual\-augmented mechanism
Theorem[2](https://arxiv.org/html/2608.26877#Thmtheorem2)exposes the response\-dependent mechanism behind the universal ceiling by augmenting the whitened design with the normalized residual\. This augmentation is used only in the analysis; sampling and fitting remain under the original ordinary volume law and selected OLS primitive\. OnL∗\>0L^\{\*\}\>0andd<s<md<s<m, whiten the design and append the normalized residual:
A=XG−1/2,z=e/L∗,B=\[Az\]\.A=XG^\{\-1/2\},\\qquad z=e/\\sqrt\{L^\{\*\}\},\\qquad B=\[A\\ z\]\.\(12\)ThenB⊤B=Id\+1B^\{\\top\}B=I\_\{d\+1\}\. Writebi⊤=\[ai⊤zi\]b\_\{i\}^\{\\top\}=\[a\_\{i\}^\{\\top\}\\ z\_\{i\}\], putKS=AS⊤ASK\_\{S\}=A\_\{S\}^\{\\top\}A\_\{S\}, and define
hi=∥bi∥22,R=∑i=1m\(1−hi\)aiai⊤,γ=m−sm−d−1\.h\_\{i\}=\\lVert b\_\{i\}\\rVert\_\{2\}^\{2\},\\qquad R=\\sum\_\{i=1\}^\{m\}\(1\-h\_\{i\}\)a\_\{i\}a\_\{i\}^\{\\top\},\\quad\\gamma=\\frac\{m\-s\}\{m\-d\-1\}\.\(13\)
The linear consequence below is the response\-uniform slack input to Theorem[3](https://arxiv.org/html/2608.26877#Thmtheorem3): a feature\-only lower bound onRRrules out the universal envelope for every compatible residual\.
LetPBP\_\{B\}denote the response\-dependent auxiliary law induced by the full residual for the following identity and bound\. Its distinct rank\-\(d\+1\)\(d\+1\)support is
ZB,s=∑\|S\|=sdet\(BS⊤BS\)=\(m−d−1s−d−1\),PB\(S\)=det\(BS⊤BS\)ZB,s\.Z\_\{B,s\}=\\sum\_\{\|S\|=s\}\\det\(B\_\{S\}^\{\\top\}B\_\{S\}\)=\\binom\{m\-d\-1\}\{s\-d\-1\},\\qquad P\_\{B\}\(S\)=\\frac\{\\det\(B\_\{S\}^\{\\top\}B\_\{S\}\)\}\{Z\_\{B,s\}\}\.\(14\)Its support can be strictly smaller than the ordinary support in \([4](https://arxiv.org/html/2608.26877#S2.E4)\); no inverse is taken on a zero augmented volume\.
###### Theorem 2\(Residual\-augmented representation and response\-aware resolvent\)\.
SupposeL∗\>0L^\{\*\}\>0andd<s<md<s<m\. IfXXis in row general position, then the ordinary law and the auxiliary lawPBP\_\{B\}are related by the exact change of measure below; sampling and fitting remain underPXP\_\{X\}\.
PX\(S\)LS=βL∗PB\(S\),M¯sL∗=Id−β𝔼B\[KS−1\],𝔼BKS=Id−γR\.P\_\{X\}\(S\)L\_\{S\}=\\beta L^\{\*\}P\_\{B\}\(S\),\\qquad\\frac\{\\overline\{M\}\_\{s\}\}\{L^\{\*\}\}=I\_\{d\}\-\\beta\\mathbb\{E\}\_\{B\}\[K\_\{S\}^\{\-1\}\],\\qquad\\mathbb\{E\}\_\{B\}K\_\{S\}=I\_\{d\}\-\\gamma R\.\(15\)Classical operator Jensen for matrix inversion, with the minus sign in the middle identity reversing the order, gives the interior resolvent bound
M¯s⪯L∗\[Id−β\(Id−γR\)−1\]\\boxed\{\\overline\{M\}\_\{s\}\\preceq L^\{\*\}\\left\[I\_\{d\}\-\\beta\(I\_\{d\}\-\\gamma R\)^\{\-1\}\\right\]\}\(16\)and its linear consequence
M¯s⪯αL∗Id−βγL∗R\.\\overline\{M\}\_\{s\}\\preceq\\alpha L^\{\*\}I\_\{d\}\-\\beta\\gamma L^\{\*\}R\.\(17\)For an arbitrary full\-column\-rank boundary design, as defined before Theorem[1](https://arxiv.org/html/2608.26877#Thmtheorem1), only the congruent one\-sided inequality is asserted\. WithQ=∑i\(1−hi\)xixi⊤Q=\\sum\_\{i\}\(1\-h\_\{i\}\)x\_\{i\}x\_\{i\}^\{\\top\},
Ms⪯L∗\[G−1−β\(G−γQ\)−1\]\.M\_\{s\}\\preceq L^\{\*\}\\left\[G^\{\-1\}\-\\beta\(G\-\\gamma Q\)^\{\-1\}\\right\]\.\(18\)The exact inverse\-moment covariance identity in \([15](https://arxiv.org/html/2608.26877#S4.E15)\) does not extend through a rank\-changing boundary\.
##### Proof overview: residual change of measure\.
On the row\-general\-position interior withL∗\>0L^\{\*\}\>0andd<s<md<s<m, the auxiliary lawPBP\_\{B\}first gives the loss\-weighted relationPX\(S\)LS=βL∗PB\(S\)P\_\{X\}\(S\)L\_\{S\}=\\beta L^\{\*\}P\_\{B\}\(S\), and then the compact chain
PX\(S\)LS\\displaystyle P\_\{X\}\(S\)L\_\{S\}=βL∗PB\(S\),\\displaystyle=\\beta L^\{\*\}P\_\{B\}\(S\),M¯sL∗\\displaystyle\\frac\{\\overline\{M\}\_\{s\}\}\{L^\{\*\}\}=Id−β𝔼B\[KS−1\],\\displaystyle=I\_\{d\}\-\\beta\\mathbb\{E\}\_\{B\}\[K\_\{S\}^\{\-1\}\],\(19\)𝔼BKS\\displaystyle\\mathbb\{E\}\_\{B\}K\_\{S\}=Id−γR\\displaystyle=I\_\{d\}\-\\gamma R⟹classical Jensen\\displaystyle\\Longrightarrow^\{\\text\{classical Jensen\}\}M¯s⪯L∗\[Id−β\(Id−γR\)−1\]\.\\displaystyle\\overline\{M\}\_\{s\}\\preceq L^\{\*\}\[I\_\{d\}\-\\beta\(I\_\{d\}\-\\gamma R\)^\{\-1\}\]\.\(20\)The first line is a loss\-weighted change of measure and augmented first moment; Jensen is the classical inversion step\[[20](https://arxiv.org/html/2608.26877#bib.bib20)\]\. At an arbitrary full\-column\-rank boundary, only the one\-sided resolvent inequality remains\. The full support\-aware derivation and the boundary arithmetic fixture are in Appendix[D](https://arxiv.org/html/2608.26877#A4)\.
##### Second\-moment refinement\.
Under the same augmented law, every supported whitened sub\-Gram matrix is a positive contraction, so an exact resolvent identity gives a computable second\-centered\-Gram\-moment tightening of the response\-aware Jensen ceiling\. Proposition[16](https://arxiv.org/html/2608.26877#Thmtheorem16)states the interior Loewner hierarchy and its one\-sided full\-column\-rank extension\.
## 5Robust pre\-response geometry
The universal envelope is a statement about the worst direction over all feature pools and responses\. For a particular feature pool, the remaining question is whether any compatible residual can attain that envelope\. The answer below is determined before a response is observed\.
LetX∈ℝm×dX\\in\\mathbb\{R\}^\{m\\times d\}have full column rank and writeG=X⊤XG=X^\{\\top\}X\. In whitened coordinates, set
A=XG−1/2,A⊤A=Id,A=XG^\{\-1/2\},\\qquad A^\{\\top\}A=I\_\{d\},\(21\)and writeai⊤a\_\{i\}^\{\\top\}for rowiiofAAandℓi=∥ai∥22\\ell\_\{i\}=\\lVert a\_\{i\}\\rVert\_\{2\}^\{2\}for its leverage\. The admissible normalized residual sphere is
𝒵A=\{z∈ker\(A⊤\):∥z∥2=1\}\.\\mathcal\{Z\}\_\{A\}=\\\{z\\in\\ker\(A^\{\\top\}\):\\lVert z\\rVert\_\{2\}=1\\\}\.\(22\)Forz∈𝒵Az\\in\\mathcal\{Z\}\_\{A\}, define the feature–residual matrix and its robust pre\-response margin by
RA\(z\)=∑i=1m\(1−ℓi−zi2\)aiai⊤,νA=minz∈𝒵Aλmin\(RA\(z\)\)\.R\_\{A\}\(z\)=\\sum\_\{i=1\}^\{m\}\(1\-\\ell\_\{i\}\-z\_\{i\}^\{2\}\)a\_\{i\}a\_\{i\}^\{\\top\},\\qquad\\nu\_\{A\}=\\min\_\{z\\in\\mathcal\{Z\}\_\{A\}\}\\lambda\_\{\\min\}\(R\_\{A\}\(z\)\)\.\(23\)
To define the response\-uniform target, fixθ∈ℝd\\theta\\in\\mathbb\{R\}^\{d\}andL∗\>0L^\{\*\}\>0and associate the compatible responseyz=Aθ\+L∗zy\_\{z\}=A\\theta\+\\sqrt\{L^\{\*\}\}zwithzz\. Under ordinary unrescaled indexed size\-ssvolume sampling, followed by selected unweighted least squares on positive\-volume subsets, let
αs=m−sm−d,βs=s−dm−d,γs=m−sm−d−1\.\\alpha\_\{s\}=\\frac\{m\-s\}\{m\-d\},\\qquad\\beta\_\{s\}=\\frac\{s\-d\}\{m\-d\},\\qquad\\gamma\_\{s\}=\\frac\{m\-s\}\{m\-d\-1\}\.\(24\)Here and belowd<s<md<s<m\. LetM¯sA,z\\overline\{M\}\_\{s\}^\{A,z\}denote the centered, full\-Gram\-whitened coefficient second moment of this sampler and estimator for\(A,yz\)\(A,y\_\{z\}\)\. BecauseA⊤A=IdA^\{\\top\}A=I\_\{d\}, whitening is the identity\. Zero\-volume sets have zero probability and receive no fit\. The response\-uniform target is the worst spectral covariance over the unit residual sphere,
𝖳specRU\(A,s\)=maxz∈𝒵Aλmax\(M¯sA,z\)L∗\.\\mathsf\{T\}\_\{\\mathrm\{spec\}\}^\{\\mathrm\{RU\}\}\(A,s\)=\\max\_\{z\\in\\mathcal\{Z\}\_\{A\}\}\\frac\{\\lambda\_\{\\max\}\(\\overline\{M\}\_\{s\}^\{A,z\}\)\}\{L^\{\*\}\}\.\(25\)ThusνA\\nu\_\{A\}is formed from features alone, while the maximum in \([25](https://arxiv.org/html/2608.26877#S5.E25)\) quantifies what that feature geometry guarantees uniformly over compatible residual directions\.
###### Theorem 3\(Robust pre\-response phase boundary\)\.
Letd≥1d\\geq 1,m≥d\+2m\\geq d\+2, andd<s<md<s<m\. Suppose thatA∈ℝm×dA\\in\\mathbb\{R\}^\{m\\times d\}has orthonormal columns and no coloop, namely
ℓi<1for everyi∈\[m\]\.\\ell\_\{i\}<1\\qquad\\text\{for every \}i\\in\[m\]\.\(26\)Use ordinary unrescaled indexed size\-ssvolume sampling followed by selected unweighted least squares on positive\-volume subsets\. For everyθ∈ℝd\\theta\\in\\mathbb\{R\}^\{d\}andL∗\>0L^\{\*\}\>0, with the definitions above,νA≥0\\nu\_\{A\}\\geq 0and
αs−𝖳specRU\(A,s\)≥βsγsνA\.\\boxed\{\\alpha\_\{s\}\-\\mathsf\{T\}\_\{\\mathrm\{spec\}\}^\{\\mathrm\{RU\}\}\(A,s\)\\geq\\beta\_\{s\}\\gamma\_\{s\}\\nu\_\{A\}\.\}\(27\)Moreover, the feature margin exactly classifies strictness versus tightness of the universal spectral envelope:
νA\>0⟺𝖳specRU\(A,s\)<αs,\\boxed\{\\nu\_\{A\}\>0\\quad\\Longleftrightarrow\\quad\\mathsf\{T\}\_\{\\mathrm\{spec\}\}^\{\\mathrm\{RU\}\}\(A,s\)<\\alpha\_\{s\},\}\(28\)νA=0⟺𝖳specRU\(A,s\)=αs\.\\boxed\{\\nu\_\{A\}=0\\quad\\Longleftrightarrow\\quad\\mathsf\{T\}\_\{\\mathrm\{spec\}\}^\{\\mathrm\{RU\}\}\(A,s\)=\\alpha\_\{s\}\.\}\(29\)IfνA=0\\nu\_\{A\}=0, there exists one residual directionz∗∈𝒵Az\_\{\*\}\\in\\mathcal\{Z\}\_\{A\}such that, simultaneously for everys′∈\{d\+1,…,m−1\}s^\{\\prime\}\\in\\\{d\+1,\\ldots,m\-1\\\},
λmax\(M¯s′A,z∗\)L∗=m−s′m−d\.\\frac\{\\lambda\_\{\\max\}\(\\overline\{M\}\_\{s^\{\\prime\}\}^\{A,z\_\{\*\}\}\)\}\{L^\{\*\}\}=\\frac\{m\-s^\{\\prime\}\}\{m\-d\}\.\(30\)The right\-hand side of \([27](https://arxiv.org/html/2608.26877#S5.E27)\) is a guaranteed lower bound on the spectral slack\. The exact statement is the phase classification in \([28](https://arxiv.org/html/2608.26877#S5.E28)\)– \([29](https://arxiv.org/html/2608.26877#S5.E29)\), not an equality formula for the positive slack magnitude\.
Under the conditions above, two eligible feature pools obeying the same globally sharp benchmark can lie in different design\-specific covariance\-envelope attainability phases: a zero\-margin pool admits one compatible positive\-loss residual that attains the normalized spectral coefficient\-covariance envelope, whereas a positive\-margin pool keeps every compatible residual below it by a uniform positive gap\. A detailed ledger of result domains, response quantifiers, and logical status appears in Appendix[A](https://arxiv.org/html/2608.26877#A1)\.
##### Proof mechanism\.
Append the normalized residual to the whitened design and writeB=\[Az\]B=\[A\\ z\]\. The columns ofBBare orthonormal, so its projection diagonalhi=ℓi\+zi2h\_\{i\}=\\ell\_\{i\}\+z\_\{i\}^\{2\}is at most one and
RA\(z\)=∑i\(1−hi\)aiai⊤⪰0\.R\_\{A\}\(z\)=\\sum\_\{i\}\(1\-h\_\{i\}\)a\_\{i\}a\_\{i\}^\{\\top\}\\succeq 0\.\(31\)The response\-aware covariance bound of Theorem[2](https://arxiv.org/html/2608.26877#Thmtheorem2)gives, for every compatible residual,
M¯sA,zL∗⪯αsId−βsγsRA\(z\)\.\\frac\{\\overline\{M\}\_\{s\}^\{A,z\}\}\{L^\{\*\}\}\\preceq\\alpha\_\{s\}I\_\{d\}\-\\beta\_\{s\}\\gamma\_\{s\}R\_\{A\}\(z\)\.\(32\)Minimizing the smallest eigenvalue ofRA\(z\)R\_\{A\}\(z\)over the residual sphere therefore gives the uniform slack in \([27](https://arxiv.org/html/2608.26877#S5.E27)\)\. The zero\-margin direction uses a support\-geometric argument: a zero\-margin witness forces every row active in its null direction to have saturated augmented leverage\. Those saturated rows induce mutually exclusive omission events under ordinary volume sampling, whose exact exclusion probabilities sum to the universal envelope\. Conversely, equality in the envelope forces a zero Rayleigh quotient ofRA\(z\)R\_\{A\}\(z\), hence zero margin\. These are separate arguments: positive margin gives uniform slack through the response\-aware bound; zero margin gives support saturation and a common equality witness; and envelope attainment forces zero margin\. The determinant support calculation and the zero\-margin witness are given in Appendix[E](https://arxiv.org/html/2608.26877#A5)\.
## 6Critical equal\-leverage geometry
The phase variable in Theorem[3](https://arxiv.org/html/2608.26877#Thmtheorem3)has a particularly concrete interpretation at critical redundancy\. Supposem=2dm=2dand every whitened row has leverageℓi=1/2\\ell\_\{i\}=1/2\. WithP=AA⊤P=AA^\{\\top\}andP⟂=I2d−PP\_\{\\perp\}=I\_\{2d\}\-P, eachz∈𝒵Az\\in\\mathcal\{Z\}\_\{A\}determines
Qz=P⟂−zz⊤,RA\(z\)=A⊤Diag\(diagQz\)A\.Q\_\{z\}=P\_\{\\perp\}\-zz^\{\\top\},\\qquad R\_\{A\}\(z\)=A^\{\\top\}\\operatorname\{Diag\}\(\\operatorname\{diag\}Q\_\{z\}\)A\.\(33\)HereQzQ\_\{z\}is the rank\-\(d−1\)\(d\-1\)projection obtained by removing one unit direction from the*fixed*Naimark complementP⟂P\_\{\\perp\}\. Thus the weights
\(diagQz\)i=12−zi2\(\\operatorname\{diag\}Q\_\{z\}\)\_\{i\}=\\frac\{1\}\{2\}\-z\_\{i\}^\{2\}\(34\)are coupled fractional weights generated by an admissible residual direction; they are neither freely chosen frame weights nor a binary erasure mask\. This interpretation uses classical Parseval frame theory, Naimark\-complement geometry, and frame\-erasure results\[[5](https://arxiv.org/html/2608.26877#bib.bib5),[7](https://arxiv.org/html/2608.26877#bib.bib7)\]\.
Define the normalized coherence
ρA=maxi≠j2\|ai⊤aj\|\.\\rho\_\{A\}=\\max\_\{i\\neq j\}2\\lvert a\_\{i\}^\{\\top\}a\_\{j\}\\rvert\.\(35\)
###### Theorem 4\(Critical residual geometry\)\.
Letd≥2d\\geq 2and letA∈ℝ2d×dA\\in\\mathbb\{R\}^\{2d\\times d\}have orthonormal columns and∥ai∥22=1/2\\lVert a\_\{i\}\\rVert\_\{2\}^\{2\}=1/2for everyii\. Then
1−ρA4≤νA≤\(1−ρA\)\(3\+ρA\)8\.\\boxed\{\\frac\{1\-\\rho\_\{A\}\}\{4\}\\leq\\nu\_\{A\}\\leq\\frac\{\(1\-\\rho\_\{A\}\)\(3\+\\rho\_\{A\}\)\}\{8\}\.\}\(36\)Moreover, the following are equivalent:
1. \(i\)νA=0\\nu\_\{A\}=0;
2. \(ii\)there are distinct indicesi,ji,j, a unit vectorv∈ℝdv\\in\\mathbb\{R\}^\{d\}, and signsσi,σj∈\{±1\}\\sigma\_\{i\},\\sigma\_\{j\}\\in\\\{\\pm 1\\\}such that ai=σi2v,aj=σj2v,ak⊤v=0\(k∉\{i,j\}\)\.a\_\{i\}=\\frac\{\\sigma\_\{i\}\}\{\\sqrt\{2\}\}v,\\qquad a\_\{j\}=\\frac\{\\sigma\_\{j\}\}\{\\sqrt\{2\}\}v,\\qquad a\_\{k\}^\{\\top\}v=0\\quad\(k\\notin\\\{i,j\\\}\)\.\(37\)
Hence the zero boundary is the projectively repeated\-pair boundary with an orthogonal remainder, while \([36](https://arxiv.org/html/2608.26877#S6.E36)\) controls the magnitude of the continuous residual\-direction margin away from that boundary\.
The projective\-duplicate boundary itself is classical: for a uniform Parseval frame, the binary two\-erasure formula of[Bodmann and Paulsen \[5\]](https://arxiv.org/html/2608.26877#bib.bib5)gives the corresponding singular retained\-frame boundary\. Theorem[4](https://arxiv.org/html/2608.26877#Thmtheorem4)applies that ancestry to the different functional in \([33](https://arxiv.org/html/2608.26877#S6.E33)\): one fixed complement, one removed residual direction, and the induced lower frame bound\. In particular, coherence boundsνA\\nu\_\{A\}from two sides but does not generally compute it exactly\.
##### Mechanism\.
Because\[Az\]\[A\\ z\]has orthonormal columns,zi2≤1/2z\_\{i\}^\{2\}\\leq 1/2and every summand inRA\(z\)=∑i\(1/2−zi2\)aiai⊤R\_\{A\}\(z\)=\\sum\_\{i\}\(1/2\-z\_\{i\}^\{2\}\)a\_\{i\}a\_\{i\}^\{\\top\}is positive semidefinite\. A zero Rayleigh quotient therefore forces every row visible in the null direction to saturatezi2=1/2z\_\{i\}^\{2\}=1/2\. Unit norm permits exactly two such coordinates, and Parsevalness forces the two corresponding rows to be the pair in \([37](https://arxiv.org/html/2608.26877#S6.E37)\)\. For the quantitative statement, rescale to the unit\-norm tight frameui=2aiu\_\{i\}=\\sqrt\{2\}a\_\{i\}\. The residual squares form a capped simplex, while the sum of its two largest directional frame energies is at most1\+ρA1\+\\rho\_\{A\}\. This gives the lower bound; projecting the signed difference of a coherence\-maximizing pair into the fixed complement gives the upper bound\. Appendix[F](https://arxiv.org/html/2608.26877#A6)gives the complete argument and the source boundary\.
##### Sharp critical witness\.
AtνA=0\\nu\_\{A\}=0, there exists one compatible residual, fixed before sampling and independent of the budget, that attains the universal spectral envelope at everys∈\{d\+1,…,2d−1\}s\\in\\\{d\+1,\\ldots,2d\-1\\\}; its exact rank\-one covariance, proof, and qualified source contrast are in Appendix[F](https://arxiv.org/html/2608.26877#A6)\.
\(a\) Shared ancestry, different adversaryClassical binary erasurewi∈\{0,1\}w\_\{i\}\\in\\\{0,1\\\}from a chosen erasure maskshared Parseval/Naimark geometryResidual\-induced continuous weightswi\(z\)=12−zi2=\[diag\(P⟂−zz⊤\)\]iw\_\{i\}\(z\)=\\dfrac\{1\}\{2\}\-z\_\{i\}^\{2\}=\\bigl\[\\operatorname\{diag\}\(P\_\{\\perp\}\-zz^\{\\top\}\)\\bigr\]\_\{i\}One unitz∈ker\(A⊤\)z\\in\\ker\(A^\{\\top\}\)couples the weights: they are neither freely chosen frame weights nor a binary mask\.\(b\) Critical equal\-leverage phaseRepeated projective pair with orthogonal remainderai=σi2v,aj=σj2v,ak⊤v=0a\_\{i\}=\\dfrac\{\\sigma\_\{i\}\}\{\\sqrt\{2\}\}v,\\hskip 9\.24994pta\_\{j\}=\\dfrac\{\\sigma\_\{j\}\}\{\\sqrt\{2\}\}v,\\hskip 9\.24994pta\_\{k\}^\{\\top\}v=0⟺νA=0⟹∃z∗:envelope attainedfor everyd<s<2d\\Longleftrightarrow\\ \\nu\_\{A\}=0\\hskip 9\.24994pt\\Longrightarrow\\hskip 9\.24994pt\\begin\{subarray\}\{c\}\\exists\\,z\_\{\*\}:\\ \\text\{envelope attained\}\\\\ \\text\{for every \}d<s<2d\\end\{subarray\}Projectively separated geometry\(ρA<1\)\(\\rho\_\{A\}<1\)⟺νA\>0⟺strict slack for everycompatible residual\\Longleftrightarrow\\ \\nu\_\{A\}\>0\\hskip 9\.24994pt\\Longleftrightarrow\\hskip 9\.24994pt\\begin\{subarray\}\{c\}\\text\{strict slack for every\}\\\\ \\text\{compatible residual\}\\end\{subarray\}The zero branch is existential in the residual; the positive branch is uniform\.Figure 2:Critical\-class schematic \(m=2dm=2d, equal leverage\)\. Panel \(a\) records the classical binary two\-erasure and Naimark\-complement ancestry\[[5](https://arxiv.org/html/2608.26877#bib.bib5),[7](https://arxiv.org/html/2608.26877#bib.bib7)\]and contrasts it with the residual\-coupled continuous weights here\. Panel \(b\) shows the existential zero branch and the uniform positive branch of theνA\\nu\_\{A\}phase\.
## 7Sound certificates and same\-primitive action
Theorem[3](https://arxiv.org/html/2608.26877#Thmtheorem3)uses the exact feature marginνA\\nu\_\{A\}, but computing that minimum over the residual sphere is not required to certify its positive branch\. A*sound pre\-response certificate*is any feature\-only scalart\(A\)t\(A\)for which
0≤t\(A\)≤νA0\\leq t\(A\)\\leq\\nu\_\{A\}\(38\)has been proved\. This is deliberately one\-sided: a positive value certifies robust strict slack, whereas a zero value, failed sufficient test, or abstention is inconclusive\.
###### Corollary 5\(Sound computable pre\-response certificates\)\.
For every full\-column\-rankXXwithm\>dm\>d, the general spectral certificatecXc\_\{X\}defined in Appendix[G](https://arxiv.org/html/2608.26877#A7)satisfies0≤cX≤νA0\\leq c\_\{X\}\\leq\\nu\_\{A\}; under the assumptions of Theorem[4](https://arxiv.org/html/2608.26877#Thmtheorem4), the critical\-coherence certificate defined there satisfies0≤tρ\(A\)≤νA0\\leq t\_\{\\rho\}\(A\)\\leq\\nu\_\{A\}\. More generally, every verifiedt\(A\)t\(A\)satisfying \([38](https://arxiv.org/html/2608.26877#S7.E38)\) obeys, on the no\-coloop positive\-loss strict\-interior domain of Theorem[3](https://arxiv.org/html/2608.26877#Thmtheorem3),
αs−𝖳specRU\(A,s\)≥βsγst\(A\)\.\\boxed\{\\alpha\_\{s\}\-\\mathsf\{T\}\_\{\\mathrm\{spec\}\}^\{\\mathrm\{RU\}\}\(A,s\)\\geq\\beta\_\{s\}\\gamma\_\{s\}t\(A\)\.\}\(39\)Thus any pass witht\(A\)\>0t\(A\)\>0certifies strictness uniformly over every compatible residual direction\. Failure or abstention does not implyνA=0\\nu\_\{A\}=0, envelope attainment, or the absence of another sound route\.
For the predeclared\(m,d,ξ\)=\(512,64,1/3\)\(m,d,\\xi\)=\(512,64,1/3\)profile, the supplementary fixed\-threshold instantiation authorizess=338s=338on a verified pass and otherwise falls back tos=363s=363; both branches retain the same primitive and give the same one\-sided13L∗Id\\tfrac\{1\}\{3\}L^\{\*\}I\_\{d\}covariance ceiling\. The predicate, integer arithmetic, guards, and nonoptimality qualification are in Appendix[G](https://arxiv.org/html/2608.26877#A7)\.
## 8Pre\-response decisions and their measured cost
Frozen features determine whether the certificate authorizes the smaller subset or invokes the universal\-envelope fallback before any response is released\. Across 72 fixed frozen\-encoding cells, the sound sufficient rule produced action/fallback counts of20/420/4,4/204/20, and5/195/19in the three released blocks, so it both authorized the smaller cardinality and abstained\. In all nine released action cells, the post\-action lower\-budget excess\-risk difference was positive, quantifying a fixed\-pool cost of the authorized reduction\. These descriptive cells do not estimate𝖳specRU\\mathsf\{T\}\_\{\\rm spec\}^\{\\rm RU\}, test Theorem[3](https://arxiv.org/html/2608.26877#Thmtheorem3), or establish population performance or selector dominance; a fallback is inconclusive rather than evidence of the zero phase\. Exact costs, covariance and accuracy summaries, matched\-budget selector context, draws, and selection notes are in Appendix[B](https://arxiv.org/html/2608.26877#A2)\. The fresh Fashion–MNIST rows are disjoint from the initial row block but share its task and frozen encoder inventory\.
## 9Related work
##### Nearest theorem\-level comparisons\.
Table[1](https://arxiv.org/html/2608.26877#S9.T1)separates the closest audited results by sampling law, estimator, budget\-specific object, response model, covariance target, and equality quantifier\. The all\-size determinant law, selected\-OLS unbiasedness, and inverse moments are prior results; the rank\-size boundary in the table concerns the narrower arbitrary\-fixed\-response loss and prediction\-covariance formulas\.
Table 1:Nearest theorem\-level comparison under the audited source versions; this is not a systematic literature census\. Budget labels refer only to the object named in the same cell\. Relations are to the present fixed\-pool coefficient\-covariance envelope and strict/tight phase, not to volume sampling as a whole\.Other audited neighbors change the law, estimator, response model, or target: volume\-rescaled and random\-design regression, generic\-sketch covariance, with\-replacement debiased fits, and active or dependent\-leverage designs treat different primitives or statistics\[[15](https://arxiv.org/html/2608.26877#bib.bib15),[16](https://arxiv.org/html/2608.26877#bib.bib16),[10](https://arxiv.org/html/2608.26877#bib.bib10),[23](https://arxiv.org/html/2608.26877#bib.bib23),[9](https://arxiv.org/html/2608.26877#bib.bib9),[18](https://arxiv.org/html/2608.26877#bib.bib18),[24](https://arxiv.org/html/2608.26877#bib.bib24)\]\. Their component overlap is discussed below; none is used here as a worldwide absence claim\.
##### Critical frame and erasure geometry\.
Parseval frames, Naimark complements, projection diagonals, and weighted\-frame operators provide geometric ancestry\[[5](https://arxiv.org/html/2608.26877#bib.bib5),[7](https://arxiv.org/html/2608.26877#bib.bib7),[21](https://arxiv.org/html/2608.26877#bib.bib21),[1](https://arxiv.org/html/2608.26877#bib.bib1),[4](https://arxiv.org/html/2608.26877#bib.bib4)\]\. Bodmann and Paulsen’s Section 3, specifically Definition 3\.1, Remark 3\.2, and the formula before Definition 3\.5, identifies the classical repeated\-pair boundary for binary erasures\[[5](https://arxiv.org/html/2608.26877#bib.bib5)\]\. Here that boundary is translated to a continuous constrained margin, withQz=P⟂−zz⊤Q\_\{z\}=P\_\{\\perp\}\-zz^\{\\top\}and coupled weights1/2−zi21/2\-z\_\{i\}^\{2\}\. Casazza et al\.’s Theorems 2\.4–2\.5 and Proposition 3\.1 provide the Naimark\-complement and equal\-norm projection machinery\[[7](https://arxiv.org/html/2608.26877#bib.bib7)\]; the overlap is partial because the residual\-coupled functional and its bounds differ \(relation: partial overlap\)\.
##### Pre\-response selection and certificates\.
Subset design for fixed\-pool regression has also been studied under other primitives\. Chen and Price’s Theorems 1 and 7 use weighted sparsification and weighted empirical risk minimization\[[9](https://arxiv.org/html/2608.26877#bib.bib9)\]\. Gittens and Magdon\-Ismail’s Lemma 2\.1 and equations \(12\) and \(21\) give the residual/leverage deletion object, while Theorem 3\.1 targets scalar full\-pool loss\[[18](https://arxiv.org/html/2608.26877#bib.bib18)\]\. Shimizu et al\.’s Theorem 1\.1 uses fixed\-cardinality dependent leverage sampling and an inverse\-inclusion rescaled fit\[[24](https://arxiv.org/html/2608.26877#bib.bib24)\]\. These works address related fixed\-pool design questions but alter the fitted objective, law, estimator, or statistic\. The overlap is partial: our sufficient test changes only cardinality when it passes \(relation: partial overlap\)\.
Finally, Dereziński and Warmuth’s Theorems 5–6 average matched selected\-Gram inverse or pseudoinverse objects under the ordinary feature law\[[11](https://arxiv.org/html/2608.26877#bib.bib11),[12](https://arxiv.org/html/2608.26877#bib.bib12)\]\. In the residual\-augmented calculation,PBP\_\{B\}is induced byB=\[Az\]B=\[A\\ z\], whereas the inverse target remainsAS⊤ASA\_\{S\}^\{\\top\}A\_\{S\}\. This response\-induced law–target mismatch is confined to the supplementary contraction correction; operational sampling remains underPXP\_\{X\}\. Its first\-order step is classical operator Jensen \(Theorem 2\.1\)\[[20](https://arxiv.org/html/2608.26877#bib.bib20)\]; the ordered remainder is a specialized local calculation \(relation: partial overlap\)\.
## 10Limitations and conclusion
##### Scope\.
We analyze conditional centered, full\-Gram\-whitened coefficient covariance for one indexed fixed feature pool under ordinary unrescaled fixed\-size volume sampling and selected unweighted OLS; the subset draw is the only randomness\. The exact phase theorem requires positive full\-fit loss, a strict\-interior budget, and no coloops\. The residual\-augmented identity additionally requires row general position, while its boundary form is one\-sided\. Feature\-only screens are sufficient: a pass certifies positive margin, whereas failure or abstention is inconclusive\. The cardinality rule and finite\-pool evidence do not claim population generalization, selector dominance, or a new sampler, and we do not establish the computational complexity of evaluatingνA\\nu\_\{A\}exactly for a general design\.
##### Conclusion\.
A globally sharp covariance envelope need not be tight for a fixed feature pool\. Under the stated conditions, feature geometry alone decides the phase:νA\>0\\nu\_\{A\}\>0means uniform strictness, whileνA=0\\nu\_\{A\}=0means one compatible residual attains the envelope at every strict\-interior budget\. Critical geometry makes the boundary concrete, and sound lower certificates convert verified positive lower bounds into conservative cardinality decisions without changing the sampler or fit\.
### Use of generative AI
Generative\-AI tools assisted with manuscript drafting and revision,LaTeXsource organization, compilation diagnostics, and research\-verification workflows\. AI\-generated outputs, including audit reports, are not evidence for mathematical correctness, novelty, or empirical claims\. The author takes responsibility for all mathematical statements, proofs, citations, calculations, experimental results, and final source artifacts\. No AI system is listed as an author\.
## References
- \[1\]Jorge Antezana, Pedro Massey, Mariano Ruiz, and Demetrio Stojanoff\.The Schur–Horn theorem for operators and frames with prescribed norms and frame operator\.*Illinois Journal of Mathematics*, 51\(2\):537–560, 2007\.URL[https://arxiv\.org/abs/math/0508646](https://arxiv.org/abs/math/0508646)\.
- \[2\]Ljiljana Arambašić and Damir Bakić\.Expansions from frame coefficients with erasures, 2016\.URL[https://arxiv\.org/abs/1602\.01656](https://arxiv.org/abs/1602.01656)\.arXiv:1602\.01656v1\.
- \[3\]Haim Avron and Christos Boutsidis\.Faster subset selection for matrices and applications\.*SIAM Journal on Matrix Analysis and Applications*, 34\(4\):1464–1499, 2013\.doi:10\.1137/120867287\.URL[https://epubs\.siam\.org/doi/10\.1137/120867287](https://epubs.siam.org/doi/10.1137/120867287)\.
- \[4\]Peter Balazs, Jean\-Pierre Antoine, and Anna Gryboś\.Weighted and controlled frames\.*International Journal of Wavelets, Multiresolution and Information Processing*, 8\(1\):109–132, 2010\.doi:10\.1142/S0219691310003377\.URL[https://doi\.org/10\.1142/S0219691310003377](https://doi.org/10.1142/S0219691310003377)\.
- \[5\]Bernhard G\. Bodmann and Vern I\. Paulsen\.Frames, graphs and erasures\.*Linear Algebra and its Applications*, 404:118–146, 2005\.doi:10\.1016/j\.laa\.2005\.02\.016\.URL[https://doi\.org/10\.1016/j\.laa\.2005\.02\.016](https://doi.org/10.1016/j.laa.2005.02.016)\.
- \[6\]Thomas Brooks, D\. Pope, and Michael Marcolini\.Airfoil Self\-Noise\.UCI Machine Learning Repository, 1989\.URL[https://doi\.org/10\.24432/C5VW2C](https://doi.org/10.24432/C5VW2C)\.
- \[7\]Peter G\. Casazza, Matthew Fickus, Dustin G\. Mixon, Jesse Peterson, and Ihar Smalyanau\.Every hilbert space frame has a naimark complement\.*Journal of Mathematical Analysis and Applications*, 406\(1\):111–119, 2013\.doi:10\.1016/j\.jmaa\.2013\.04\.047\.URL[https://doi\.org/10\.1016/j\.jmaa\.2013\.04\.047](https://doi.org/10.1016/j.jmaa.2013.04.047)\.
- \[8\]Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton\.A simple framework for contrastive learning of visual representations\.In*Proceedings of the 37th International Conference on Machine Learning*, volume 119 of*Proceedings of Machine Learning Research*, pages 1597–1607\. PMLR, 2020\.URL[https://proceedings\.mlr\.press/v119/chen20j\.html](https://proceedings.mlr.press/v119/chen20j.html)\.
- \[9\]Xue Chen and Eric Price\.Active regression via linear\-sample sparsification\.In*Proceedings of the Thirty\-Second Conference on Learning Theory*, volume 99 of*Proceedings of Machine Learning Research*, pages 663–695\. PMLR, 2019\.URL[https://proceedings\.mlr\.press/v99/chen19a\.html](https://proceedings.mlr.press/v99/chen19a.html)\.
- \[10\]Jocelyn T\. Chi and Ilse C\. F\. Ipsen\.A projector\-based approach to quantifying total and excess uncertainties for sketched linear regression\.*Information and Inference: A Journal of the IMA*, 11\(3\):1055–1077, 2022\.doi:10\.1093/imaiai/iaab016\.URL[https://doi\.org/10\.1093/imaiai/iaab016](https://doi.org/10.1093/imaiai/iaab016)\.
- \[11\]Michał Dereziński and Manfred K\. Warmuth\.Unbiased estimates for linear regression via volume sampling\.In*Advances in Neural Information Processing Systems 30*, pages 3084–3093, 2017\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2017/file/54e36c5ff5f6a1802925ca009f3ebb68\-Paper\.pdf](https://proceedings.neurips.cc/paper_files/paper/2017/file/54e36c5ff5f6a1802925ca009f3ebb68-Paper.pdf)\.
- \[12\]Michał Dereziński and Manfred K\. Warmuth\.Reverse iterative volume sampling for linear regression\.*Journal of Machine Learning Research*, 19\(23\):1–39, 2018\.URL[https://www\.jmlr\.org/papers/volume19/17\-781/17\-781\.pdf](https://www.jmlr.org/papers/volume19/17-781/17-781.pdf)\.
- \[13\]Michał Dereziński, Manfred K\. Warmuth, and Daniel J\. Hsu\.Leveraged volume sampling for linear regression\.In*Advances in Neural Information Processing Systems 31*, pages 2505–2514, 2018\.URL[https://papers\.neurips\.cc/paper\_files/paper/2018/file/2ba8698b79439589fdd2b0f7218d8b07\-Paper\.pdf](https://papers.neurips.cc/paper_files/paper/2018/file/2ba8698b79439589fdd2b0f7218d8b07-Paper.pdf)\.
- \[14\]Michał Dereziński, Kenneth L\. Clarkson, Michael W\. Mahoney, and Manfred K\. Warmuth\.Minimax experimental design: Bridging the gap between statistical and worst\-case approaches to least squares regression\.In*Proceedings of the 32nd Conference on Learning Theory*, volume 99 of*Proceedings of Machine Learning Research*, pages 1050–1069, 2019a\.URL[https://proceedings\.mlr\.press/v99/derezinski19b/derezinski19b\.pdf](https://proceedings.mlr.press/v99/derezinski19b/derezinski19b.pdf)\.
- \[15\]Michał Dereziński, Manfred K\. Warmuth, and Daniel Hsu\.Correcting the bias in least squares regression with volume\-rescaled sampling\.In*Proceedings of the Twenty\-Second International Conference on Artificial Intelligence and Statistics*, volume 89 of*Proceedings of Machine Learning Research*, pages 944–953, 2019b\.URL[https://proceedings\.mlr\.press/v89/derezinski19a/derezinski19a\.pdf](https://proceedings.mlr.press/v89/derezinski19a/derezinski19a.pdf)\.
- \[16\]Michał Dereziński, Manfred K\. Warmuth, and Daniel Hsu\.Unbiased estimators for random design regression\.*Journal of Machine Learning Research*, 23:1–46, 2022\.URL[https://www\.jmlr\.org/papers/volume23/19\-571/19\-571\.pdf](https://www.jmlr.org/papers/volume23/19-571/19-571.pdf)\.
- \[17\]Ethan N\. Epperly\.Adaptive randomized pivoting and volume sampling, 2026\.URL[https://arxiv\.org/pdf/2510\.02513v2](https://arxiv.org/pdf/2510.02513v2)\.arXiv:2510\.02513v2\.
- \[18\]Alex Gittens and Malik Magdon\-Ismail\.Reduced label complexity for tightℓ2\\ell\_\{2\}regression, 2023\.URL[https://arxiv\.org/abs/2305\.07486](https://arxiv.org/abs/2305.07486)\.arXiv:2305\.07486v1\.
- \[19\]Jean\-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre H\. Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, Bilal Piot, Koray Kavukcuoglu, Remi Munos, and Michal Valko\.Bootstrap your own latent: A new approach to self\-supervised learning\.In*Advances in Neural Information Processing Systems*, volume 33, pages 21271–21284, 2020\.URL[https://proceedings\.neurips\.cc/paper/2020/hash/f3ada80d5c4ee70142b17b8192b2958e\-Abstract\.html](https://proceedings.neurips.cc/paper/2020/hash/f3ada80d5c4ee70142b17b8192b2958e-Abstract.html)\.
- \[20\]Frank Hansen and Gert K\. Pedersen\.Jensen’s operator inequality\.*Bulletin of the London Mathematical Society*, 35\(4\):553–564, 2003\.doi:10\.1112/S0024609303002200\.URL[https://doi\.org/10\.1112/S0024609303002200](https://doi.org/10.1112/S0024609303002200)\.
- \[21\]Richard V\. Kadison\.The pythagorean theorem: I\. the finite case\.*Proceedings of the National Academy of Sciences*, 99\(7\):4178–4184, 2002\.doi:10\.1073/pnas\.032677199\.URL[https://doi\.org/10\.1073/pnas\.032677199](https://doi.org/10.1073/pnas.032677199)\.
- \[22\]Chengtao Li, Stefanie Jegelka, and Suvrit Sra\.Polynomial time algorithms for dual volume sampling\.In*Advances in Neural Information Processing Systems 30*, pages 5038–5047, 2017\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2017/file/18bb68e2b38e4a8ce7cf4f6b2625768c\-Paper\.pdf](https://proceedings.neurips.cc/paper_files/paper/2017/file/18bb68e2b38e4a8ce7cf4f6b2625768c-Paper.pdf)\.
- \[23\]Chengmei Niu, Sachin Garg, Michał Dereziński, and Zhenyu Liao\.Debiasing random oblique projections for subsampled ols and fast cur in high dimensions, 2026\.URL[https://arxiv\.org/abs/2605\.24955](https://arxiv.org/abs/2605.24955)\.arXiv:2605\.24955v1\.
- \[24\]Atsushi Shimizu, Xiaoou Cheng, Christopher Musco, and Jonathan Weare\.Improved active learning via dependent leverage score sampling\.In*International Conference on Learning Representations*, pages 29585–29608, 2024\.URL[https://proceedings\.iclr\.cc/paper\_files/paper/2024/file/7f05193e5487287a890df7fbc3554427\-Paper\-Conference\.pdf](https://proceedings.iclr.cc/paper_files/paper/2024/file/7f05193e5487287a890df7fbc3554427-Paper-Conference.pdf)\.
- \[25\]Thomas Strohmer and Robert W\. Heath Jr\.Grassmannian frames with applications to coding and communication\.*Applied and Computational Harmonic Analysis*, 14\(3\):257–275, 2003\.doi:10\.1016/S1063\-5203\(03\)00023\-X\.URL[https://doi\.org/10\.1016/S1063\-5203\(03\)00023\-X](https://doi.org/10.1016/S1063-5203(03)00023-X)\.
- \[26\]Yonglong Tian, Yue Wang, Dilip Krishnan, Joshua B\. Tenenbaum, and Phillip Isola\.Rethinking few\-shot image classification: A good embedding is all you need?In*Computer Vision – ECCV 2020*, pages 266–282\. Springer, 2020\.doi:10\.1007/978\-3\-030\-58568\-6\_16\.URL[https://www\.ecva\.net/papers/eccv\_2020/papers\_ECCV/html/2118\_ECCV\_2020\_paper\.php](https://www.ecva.net/papers/eccv_2020/papers_ECCV/html/2118_ECCV_2020_paper.php)\.
- \[27\]Athanasios Tsanas and Angeliki Xifara\.Energy Efficiency\.UCI Machine Learning Repository, 2012\.URL[https://doi\.org/10\.24432/C51307](https://doi.org/10.24432/C51307)\.
- \[28\]Jing Xu, Xu Luo, Xinglin Pan, Wenjie Pei, Yanan Li, and Zenglin Xu\.Alleviating the sample selection bias in few\-shot learning by removing projection to the centroid\.In*Advances in Neural Information Processing Systems*, volume 35, pages 21073–21086, 2022\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2022/hash/84b686f7cc7b7751e9aaac0da74f755a\-Abstract\.html](https://proceedings.neurips.cc/paper_files/paper/2022/hash/84b686f7cc7b7751e9aaac0da74f755a-Abstract.html)\.
- \[29\]I\-Cheng Yeh\.Concrete Compressive Strength\.UCI Machine Learning Repository, 1998a\.URL[https://doi\.org/10\.24432/C5PK67](https://doi.org/10.24432/C5PK67)\.
- \[30\]I\-Cheng Yeh\.Modeling of strength of high\-performance concrete using artificial neural networks\.*Cement and Concrete Research*, 28\(12\):1797–1808, 1998b\.doi:10\.1016/S0008\-8846\(98\)00165\-3\.
## Proof support and supplementary material
This appendix supplies the derivations used by the main result chain\. All subsets are unordered subsets of indexed rows\. A zero\-volume set has zero probability and receives no selected inverse or fit\. The proof order follows the result dependencies: ordinary law and covariance envelope, residual augmentation, robust pre\-response geometry, critical geometry, and the response\-uniform feature certificate\. Additional balanced and finite diagnostic material is retained afterward as supplementary context only\.
### Supplementary fixed\-profile instantiation
For the fixed\-threshold route, define in the original feature coordinates
H0=X⊤diag\(1−ℓ\)X,ℓ−=miniℓi,ℓ\+=maxiℓi,c⋆=1−ℓ−,K=⌈1c⋆⌉,Φ13\(X\)=H0−\[1360\+c⋆ℓ\+K\]G\.\\begin\{split\}H\_\{0\}&=X^\{\\top\}\\operatorname\{diag\}\(1\-\\ell\)X,\\qquad\\ell\_\{\-\}=\\min\_\{i\}\\ell\_\{i\},\\quad\\ell\_\{\+\}=\\max\_\{i\}\\ell\_\{i\},\\\\ c\_\{\\star\}&=1\-\\ell\_\{\-\},\\qquad K=\\left\\lceil\\frac\{1\}\{c\_\{\\star\}\}\\right\\rceil,\\\\ \\Phi\_\{13\}\(X\)&=H\_\{0\}\-\\left\[\\frac\{13\}\{60\}\+c\_\{\\star\}\\ell\_\{\+\}K\\right\]G\.\\end\{split\}\(40\)The strict guards0<ℓi<10<\\ell\_\{i\}<1for everyii, together withΦ13\(X\)⪰0\\Phi\_\{13\}\(X\)\\succeq 0, certify13/60≤νA13/60\\leq\\nu\_\{A\}by the derivation in Appendix[G](https://arxiv.org/html/2608.26877#A7)\.
###### Corollary 6\(Fixed\-profile certified instantiation\)\.
Fix the paper profile
\(m,d,ξ\)=\(512,64,1/3\),n=m−d=448\.\(m,d,\\xi\)=\(512,64,1/3\),\\qquad n=m\-d=448\.\(41\)If the strict guards and fixed\-threshold predicate \([40](https://arxiv.org/html/2608.26877#Ax1.E40)\) pass, then the certified count
q13=⌊30n\(n−1\)77n−90⌋=174q\_\{13\}=\\left\\lfloor\\frac\{30n\(n\-1\)\}\{77n\-90\}\\right\\rfloor=174\(42\)givess13=m−q13=338s\_\{13\}=m\-q\_\{13\}=338and, for every fixed positive\-loss response,
M¯338⪯13L∗Id\.\\boxed\{\\overline\{M\}\_\{338\}\\preceq\\frac\{1\}\{3\}L^\{\*\}I\_\{d\}\.\}\(43\)If that sufficient route does not pass or abstains, the universal\-envelope fallbackqH=⌊n/3⌋=149q\_\{H\}=\\lfloor n/3\\rfloor=149givessH=363s\_\{H\}=363and the same one\-sided tolerance certificate atsHs\_\{H\}\. Both branches retain the ordinary indexed fixed\-size volume sampler and selected unweighted OLS; only the cardinality differs\. This sufficient rule makes no minimality, optimality, label\-saving, or predictive\-utility claim\.
## Theorem\-to\-proof map and scope
The main statements map to the appendix as follows\.
- •Universal normalizer and support:Proposition[7](https://arxiv.org/html/2608.26877#Thmtheorem7); the endpoint and support conventions precede it\.
- •Universal unbiasedness and covariance envelope: Theorem[8](https://arxiv.org/html/2608.26877#Thmtheorem8), Proposition[10](https://arxiv.org/html/2608.26877#Thmtheorem10), Lemma[11](https://arxiv.org/html/2608.26877#Thmtheorem11), and Proposition[13](https://arxiv.org/html/2608.26877#Thmtheorem13)\.
- •Full\-pool loss consequence:Equation \([67](https://arxiv.org/html/2608.26877#A3.E67)\), using the full\-pool Pythagoras identity\.
- •Attainment, strict interior slack, and supremum:Propositions[14](https://arxiv.org/html/2608.26877#Thmtheorem14)and[15](https://arxiv.org/html/2608.26877#Thmtheorem15)\.
- •Residual\-augmented exact transform:Equations \([94](https://arxiv.org/html/2608.26877#A4.E94)\) and \([98](https://arxiv.org/html/2608.26877#A4.E98)\), with their row\-general\-position/interior domains\.
- •Resolvent and boundary inequality:Equations \([103](https://arxiv.org/html/2608.26877#A4.E103)\)– \([121](https://arxiv.org/html/2608.26877#A4.E121)\); the final boundary remark states the identity–inequality distinction\.
- •Interior identity versus boundary inequality:Proposition[17](https://arxiv.org/html/2608.26877#Thmtheorem17)gives an exact boundary enumeration where the inverse\-moment identity fails and the boundary resolvent inequality has positive slack\.
- •Robust pre\-response phase boundary:Theorem[3](https://arxiv.org/html/2608.26877#Thmtheorem3), with compactness, positive semidefiniteness, the uniform slack implication, the zero\-margin support calculation, and the converse equality argument in Appendix[E](https://arxiv.org/html/2608.26877#A5)\.
- •Critical equal\-leverage geometry and sharp witness:Theorem[4](https://arxiv.org/html/2608.26877#Thmtheorem4)and the unnumbered sharp critical witness in Section[6](https://arxiv.org/html/2608.26877#S6), proved in Appendix[F](https://arxiv.org/html/2608.26877#A6)\. The repeated\-projective\-pair boundary is stated with its classical binary\-erasure ancestry\.
- •Response\-uniform feature certificate:Corollary[5](https://arxiv.org/html/2608.26877#Thmtheorem5), with geometry proof, coefficient ledger, and structural calibration in Appendix[G](https://arxiv.org/html/2608.26877#A7)and Subsection[G\.5](https://arxiv.org/html/2608.26877#A7.SS5); Theorem[2](https://arxiv.org/html/2608.26877#Thmtheorem2)supplies only the one\-sided resolvent inequality at arbitrary full\-column\-rank boundaries\.
- •Supplementary additional results:the balanced structural\-gap statement and proof, followed by finite fixed\-pool diagnostics and their reproducibility record\. These items are not part of the main claim chain and are not evidence for the robust phase theorem\.
The exact augmented identity is never used outside its stated interior domain\. Its witness hash1=1h\_\{1\}=1, and the ordinary class definition does not assume full spark\. The only asymptotic interpretation retained for the supplementary class quantity is its statedr=o\(d\)r=o\(d\)regime\.
### Scope summary
## Appendix ADetailed claim and assumption ledger
Table 2:Detailed claim, quantifier, and assumption ledger for the theorem and consequence chain\. “Exact phase” in Theorem[3](https://arxiv.org/html/2608.26877#Thmtheorem3)refers to the strict\-versus\-tight equivalence; the slack magnitude has only a one\-sided lower bound\. Certificate passes are sufficient, while failure or abstention is inconclusive\.
## Appendix BSupplementary evidence ledger
The full fixed\-pool evidence ledger retains metric definitions, every numerical cell, draw record, and selection notes\. It is supplementary to the compact descriptive summary in Section[8](https://arxiv.org/html/2608.26877#S8)\. For each fixed pool, letw∗=argminw∥Xw−y∥22w^\{\*\}=\\arg\\min\_\{w\}\\lVert Xw\-y\\rVert\_\{2\}^\{2\}andL∗=∥Xw∗−y∥22\>0L^\{\*\}=\\lVert Xw^\{\*\}\-y\\rVert\_\{2\}^\{2\}\>0\. For drawrr, with selected\-OLS fitw^r\\widehat\{w\}\_\{r\}, the reported normalized excess risk is exactly
er=∥Xw^r−y∥22−L∗dL∗\.e\_\{r\}=\\frac\{\\lVert X\\widehat\{w\}\_\{r\}\-y\\rVert\_\{2\}^\{2\}\-L^\{\*\}\}\{dL^\{\*\}\}\.Writinge¯s\\bar\{e\}\_\{s\}for its mean over the 100 draws at subset sizess, the lower\-budget difference isD=e¯338−e¯363D=\\bar\{e\}\_\{338\}\-\\bar\{e\}\_\{363\}\. The covariance columns useX=QRX=QRand the empirical covariance aboutw∗w^\{\*\}
C^s=1100∑r=1100\[R\(w^r−w∗\)\]\[R\(w^r−w∗\)\]⊤,R⊤R=X⊤X\.\\widehat\{C\}\_\{s\}=\\frac\{1\}\{100\}\\sum\_\{r=1\}^\{100\}\\bigl\[R\(\\widehat\{w\}\_\{r\}\-w^\{\*\}\)\\bigr\]\\bigl\[R\(\\widehat\{w\}\_\{r\}\-w^\{\*\}\)\\bigr\]^\{\\top\},\\qquad R^\{\\top\}R=X^\{\\top\}X\.Thus the displayedtr\(C^s\)/d\\operatorname\{tr\}\(\\widehat\{C\}\_\{s\}\)/dandλmax\(C^s\)\\lambda\_\{\\max\}\(\\widehat\{C\}\_\{s\}\)are raww∗w^\{\*\}\-centered, full\-Gram\-whitened covariance summaries: unlikeere\_\{r\}, they are not divided byL∗L^\{\*\}, and they are not estimates of𝖳specRU\(A,s\)\\mathsf\{T\}\_\{\\mathrm\{spec\}\}^\{\\mathrm\{RU\}\}\(A,s\)\.
Table 3:Compact descriptive fixed\-pool certificate/action evidence\. The counts are feature\-only actions or fallbacks\. PositiveDDis the observed lower\-budget differencee¯338−e¯363\\bar\{e\}\_\{338\}\-\\bar\{e\}\_\{363\}\. The matched\-budget directions are the observed per\-cell normalized same\-pool excess riske¯r\\bar\{e\}\_\{r\}ats=338s=338on the five fresh cells selected by the feature\-only action; they do not establish selector dominance\. Full metric definitions, cells, draws, and selection notes are in Appendix[B](https://arxiv.org/html/2608.26877#A2)\.Table 4:Evidence ladder for a feature\-only cardinality decision and its descriptive fixed\-pool consequences\. Block A records decisions made from frozen features before labels or responses were accessed\. Block B reports the locked lower\-budget comparison with the unchanged ordinary indexed fixed\-size volume\-sampling and selected unweighted OLS primitive: each arm uses 100 domain\-separated draws, withs=338s=338versuss=363s=363andD=e¯338−e¯363D=\\bar\{e\}\_\{338\}\-\\bar\{e\}\_\{363\}\. Block C reports the sames=338s=338budget on five action\-selected fresh\-row cells for volume, uniform, and leverage sampling\. The fresh\-row cells use rows disjoint from the initial Fashion\-MNIST block but share its task and frozen encoder inventory\. All entries are descriptive fixed\-pool estimates; the table measures the decision\-layer cost and context rather than a population\-performance claim or the worst\-residual phase quantity\.A\. Pre\-response action and fallback
B\. Locked lower\-budget cost–risk \(s=338s=338versuss=363s=363\)
Cell \(local public label\)e¯338\\bar\{e\}\_\{338\}e¯363\\bar\{e\}\_\{363\}DDtrace/dd338 / 363λmax\\lambda\_\{\\max\}338 / 363accuracy338 / 363Initial rows, cell 10\.0015669150\.001294128\+×10−4\+2\.72787\\\!\\times\\\!10^\{\-4\}0\.166320 / 0\.1373650\.743613 / 0\.7625120\.951484 / 0\.952188Initial rows, cell 20\.0014970210\.001260033\+×10−4\+2\.36987\\\!\\times\\\!10^\{\-4\}0\.116934 / 0\.0984230\.641793 / 0\.5945790\.969941 / 0\.970801Initial rows, cell 30\.0014695290\.001171797\+×10−4\+2\.97733\\\!\\times\\\!10^\{\-4\}0\.118658 / 0\.0946170\.578329 / 0\.5298200\.968047 / 0\.968066Initial rows, cell 40\.0015716350\.001219075\+×10−4\+3\.52561\\\!\\times\\\!10^\{\-4\}0\.134503 / 0\.1043300\.847321 / 0\.5416390\.956816 / 0\.957539Fresh rows, cell 10\.001508710\.00118513\+×10−4\+3\.2358\\\!\\times\\\!10^\{\-4\}0\.151272 / 0\.1188280\.689214 / 0\.7632640\.945410 / 0\.946777Fresh rows, cell 20\.001578980\.00124190\+×10−4\+3\.3709\\\!\\times\\\!10^\{\-4\}0\.101830 / 0\.0800910\.570481 / 0\.4736260\.970684 / 0\.971172Fresh rows, cell 30\.001522800\.00120305\+×10−4\+3\.1975\\\!\\times\\\!10^\{\-4\}0\.111436 / 0\.0880370\.500476 / 0\.4192260\.981621 / 0\.982500Fresh rows, cell 40\.001618540\.00132181\+×10−4\+2\.9674\\\!\\times\\\!10^\{\-4\}0\.124914 / 0\.1020130\.765129 / 0\.6326680\.969473 / 0\.969648Fresh rows, cell 50\.001484330\.00116875\+×10−4\+3\.1557\\\!\\times\\\!10^\{\-4\}0\.101266 / 0\.0797360\.699423 / 0\.5541910\.972285 / 0\.972813Initial\-row block:DDrange\[2\.36987,3\.52561\]×10−4\[2\.36987,3\.52561\]\\\!\\times\\\!10^\{\-4\}; mean\+×10−4\+2\.90017\\\!\\times\\\!10^\{\-4\}\.Fresh\-row block:DDrange\[2\.9674,3\.3709\]×10−4\[2\.9674,3\.3709\]\\\!\\times\\\!10^\{\-4\}; mean\+×10−4\+3\.1855\\\!\\times\\\!10^\{\-4\}\.
C\. Matched\-budget context \(s=338s=338; five action\-selected fresh\-row cells\)
*Notes\.*Block A is feature\-only and precedes response/label access\. Blocks B and C use the same fixed\-pool refitting target\. Their trace/ddand top\-eigenvalue columns are the raww∗w^\{\*\}\-centered, full\-Gram\-whitened covariance summaries defined above; they are not normalized byL∗L^\{\*\}and are not estimates of𝖳specRU\(A,s\)\\mathsf\{T\}\_\{\\mathrm\{spec\}\}^\{\\mathrm\{RU\}\}\(A,s\)\. Accuracy is binary accuracy\. In Block C, “five action\-selected” means the cells were chosen by the feature\-only decision before the matched\-budget release; it is not a random\-cell or task\-level replication\. The fresh\-row study is a descriptive row\-block replication with the shared Fashion–MNIST task and encoder inventory stated in the caption, and the selector directions do not establish general dominance\.
## Appendix CThe universal fixed\-pool covariance envelope
This section treats a fixed design and a fixed response\. LetX∈ℝm×dX\\in\\mathbb\{R\}^\{m\\times d\}have full column rank and lety∈ℝmy\\in\\mathbb\{R\}^\{m\}be arbitrary\. Write
G=X⊤X,w∗=G−1X⊤y,e=y−Xw∗,L∗=∥e∥22\.G=X^\{\\top\}X,\\qquad w^\{\*\}=G^\{\-1\}X^\{\\top\}y,\\qquad e=y\-Xw^\{\*\},\\qquad L^\{\*\}=\\lVert e\\rVert\_\{2\}^\{2\}\.\(44\)For an unordered set of row indicesS⊆\[m\]S\\subseteq\[m\],\|S\|=s\|S\|=s, put
GS=XS⊤XS,DS=det\(GS\)\.G\_\{S\}=X\_\{S\}^\{\\top\}X\_\{S\},\\qquad D\_\{S\}=\\det\(G\_\{S\}\)\.\(45\)The sampler below is ordinary, unrescaled, indexed, fixed\-size, and without replacement; it is followed by unweighted least squares\. Thus rows having the same numerical value remain different indexed observations\. This is the ordinary fixed\-size volume law studied, in equivalent row or transposed column notation, in\[[12](https://arxiv.org/html/2608.26877#bib.bib12),[3](https://arxiv.org/html/2608.26877#bib.bib3),[22](https://arxiv.org/html/2608.26877#bib.bib22)\]\. Only sets withDS\>0D\_\{S\}\>0support an estimator, and on such sets
wS=GS−1XS⊤yS,LS=∥XSwS−yS∥22\.w\_\{S\}=G\_\{S\}^\{\-1\}X\_\{S\}^\{\\top\}y\_\{S\},\\qquad L\_\{S\}=\\lVert X\_\{S\}w\_\{S\}\-y\_\{S\}\\rVert\_\{2\}^\{2\}\.\(46\)No fit is assigned to a zero\-volume set\. There is no rescaling, ridge term, importance weight, replacement, or random\-response expectation in these definitions; every expectation below is conditional on the displayed fixedXXandyy\.
The endpoint branches will be used without taking a quotient\. Ifm=dm=d, thens=ms=m,XXis invertible,wS=w∗=X−1yw\_\{S\}=w^\{\*\}=X^\{\-1\}y, and the covariance and both residual losses are zero; no expression containing\(m−d\)−1\(m\-d\)^\{\-1\}is evaluated\. Ifm\>dm\>dands=ms=m, the unique set is the full pool and its centered second moment is zero\. IfL∗=0L^\{\*\}=0, theny=Xw∗y=Xw^\{\*\}and every supported selected fit equalsw∗w^\{\*\}, so again the centered second moment is zero\. Ifs=d<ms=d<m, every supported square selected system interpolates; no rank\-\(d\+1\)\(d\+1\)representation is used here\. These conventions also cover repeated rows and zero\-volume sets\.
### C\.1Fixed\-cardinality normalizer and theorem statement
###### Proposition 7\(Fixed\-cardinality normalizer\)\.
Form\>dm\>dandd≤s≤md\\leq s\\leq m,
ZX,s:=∑\|S\|=sDS=\(m−ds−d\)det\(G\)\>0,ℙX\(S\)=DSZX,s\.Z\_\{X,s\}:=\\sum\_\{\|S\|=s\}D\_\{S\}=\\binom\{m\-d\}\{s\-d\}\\det\(G\)\>0,\\qquad\\mathbb\{P\}\_\{X\}\(S\)=\\frac\{D\_\{S\}\}\{Z\_\{X,s\}\}\.\(47\)
###### Proof\.
For each indexedSS, Cauchy–Binet gives
DS=∑T⊆S\|T\|=ddet\(XT\)2\.D\_\{S\}=\\sum\_\{\\begin\{subarray\}\{c\}T\\subseteq S\\\\ \|T\|=d\\end\{subarray\}\}\\det\(X\_\{T\}\)^\{2\}\.\(48\)Each indexeddd\-setTToccurs in exactly\(m−ds−d\)\\binom\{m\-d\}\{s\-d\}indexed size\-sssupersets\. Summing \([48](https://arxiv.org/html/2608.26877#A3.E48)\) and applying Cauchy–Binet once more yields
∑\|S\|=sDS=\(m−ds−d\)∑\|T\|=ddet\(XT\)2=\(m−ds−d\)det\(X⊤X\)\.\\sum\_\{\|S\|=s\}D\_\{S\}=\\binom\{m\-d\}\{s\-d\}\\sum\_\{\|T\|=d\}\\det\(X\_\{T\}\)^\{2\}=\\binom\{m\-d\}\{s\-d\}\\det\(X^\{\\top\}X\)\.Full column rank makes the final determinant positive\. The count is over indexed subsets, so it is unchanged when some rows coincide; terms with zero determinant simply contribute zero\. ∎
Until unbiasedness has been established, define only the centered second moment
Cs:=𝔼X\[\(wS−w∗\)\(wS−w∗\)⊤\]\.C\_\{s\}:=\\mathbb\{E\}\_\{X\}\\bigl\[\(w\_\{S\}\-w^\{\*\}\)\(w\_\{S\}\-w^\{\*\}\)^\{\\top\}\\bigr\]\.\(49\)Form\>dm\>d, let
α=m−sm−d,β=s−dm−d=1−α\.\\alpha=\\frac\{m\-s\}\{m\-d\},\\qquad\\beta=\\frac\{s\-d\}\{m\-d\}=1\-\\alpha\.\(50\)
###### Theorem 8\(Universal covariance envelope for fixed\-pool volume\-sampled least squares\)\.
For every full\-column\-rankXX, every fixedyy,m\>dm\>d, andd≤s≤md\\leq s\\leq m,
𝔼XwS=w∗\.\\mathbb\{E\}\_\{X\}w\_\{S\}=w^\{\*\}\.\(51\)ConsequentlyMs:=CsM\_\{s\}:=C\_\{s\}is the covariance ofwSw\_\{S\}, and
Ms⪯αL∗G−1,M¯s:=G1/2MsG1/2⪯αL∗Id\.M\_\{s\}\\preceq\\alpha L^\{\*\}G^\{\-1\},\\qquad\\overline\{M\}\_\{s\}:=G^\{1/2\}M\_\{s\}G^\{1/2\}\\preceq\\alpha L^\{\*\}I\_\{d\}\.\(52\)Moreover,
𝔼X∥XwS−y∥22≤\(1\+dα\)L∗=\(1\+d\(m−s\)m−d\)L∗\.\\mathbb\{E\}\_\{X\}\\lVert Xw\_\{S\}\-y\\rVert\_\{2\}^\{2\}\\leq\(1\+d\\alpha\)L^\{\*\}=\\left\(1\+\\frac\{d\(m\-s\)\}\{m\-d\}\\right\)L^\{\*\}\.\(53\)The coefficient in \([52](https://arxiv.org/html/2608.26877#A3.E52)\) is attained by the family in Proposition[14](https://arxiv.org/html/2608.26877#Thmtheorem14)\. For positive\-loss row\-general\-position designs withd≥2d\\geq 2it is strictly unattained whend<s<md<s<m, but it remains their operator\-norm supremum as described in Proposition[15](https://arxiv.org/html/2608.26877#Thmtheorem15)\.
The directional ratio associated with \([52](https://arxiv.org/html/2608.26877#A3.E52)\) isλmax\(M¯s\)/\(αL∗\)\\lambda\_\{\\max\}\(\\overline\{M\}\_\{s\}\)/\(\\alpha L^\{\*\}\)only whenαL∗\>0\\alpha L^\{\*\}\>0\. It is not the corresponding raw Euclidean eigenvalue ofMsM\_\{s\}, and no normalized ratio is formed ats=ms=morL∗=0L^\{\*\}=0\.
### C\.2Rank\-size basis moments and uniform padding
The following rank\-size calculation includes the off\-diagonal second moments\. Rank\-size unbiasedness and volume\-sampling moments are known ingredients\[[11](https://arxiv.org/html/2608.26877#bib.bib11),[12](https://arxiv.org/html/2608.26877#bib.bib12)\]; the direct derivation fixes the signs and support needed below\.
###### Lemma 9\(Rank\-size basis moments\)\.
LetA∈ℝq×dA\\in\\mathbb\{R\}^\{q\\times d\}have full column rank, suppose every indexeddd\-row submatrix is nonsingular, and fixb∈ℝqb\\in\\mathbb\{R\}^\{q\}\. Define
K=A⊤A,v=K−1A⊤b,r=b−Av\.K=A^\{\\top\}A,\\qquad v=K^\{\-1\}A^\{\\top\}b,\\qquad r=b\-Av\.Draw add\-setTTwith probabilitydet\(AT\)2/det\(K\)\\det\(A\_\{T\}\)^\{2\}/\\det\(K\)and putvT=AT−1bTv\_\{T\}=A\_\{T\}^\{\-1\}b\_\{T\}\. Then
𝔼vT=v,𝔼\[\(vT−v\)\(vT−v\)⊤\]=∥r∥22K−1\.\\mathbb\{E\}v\_\{T\}=v,\\qquad\\mathbb\{E\}\[\(v\_\{T\}\-v\)\(v\_\{T\}\-v\)^\{\\top\}\]=\\lVert r\\rVert\_\{2\}^\{2\}K^\{\-1\}\.\(54\)
###### Proof\.
The normal equations giveA⊤r=0A^\{\\top\}r=0\. Forj∈\[d\]j\\in\[d\], letBjB\_\{j\}beAAwith columnjjreplaced byrr\. Cramer’s rule, using increasing row order throughout, gives
det\(AT\)\(vT−v\)j=det\(\(Bj\)T\)\.\\det\(A\_\{T\}\)\(v\_\{T\}\-v\)\_\{j\}=\\det\(\(B\_\{j\}\)\_\{T\}\)\.\(55\)The polarized Cauchy–Binet identity
∑\|T\|=ddet\(CT\)det\(DT\)=det\(C⊤D\)\\sum\_\{\|T\|=d\}\\det\(C\_\{T\}\)\\det\(D\_\{T\}\)=\\det\(C^\{\\top\}D\)\(56\)therefore gives
det\(K\)𝔼\[\(vT−v\)j\]=det\(A⊤Bj\)=0,\\det\(K\)\\mathbb\{E\}\[\(v\_\{T\}\-v\)\_\{j\}\]=\\det\(A^\{\\top\}B\_\{j\}\)=0,because columnjjofA⊤BjA^\{\\top\}B\_\{j\}isA⊤r=0A^\{\\top\}r=0\.
For the full second moment, letρ=∥r∥22\\rho=\\lVert r\\rVert\_\{2\}^\{2\}\. Equations \([55](https://arxiv.org/html/2608.26877#A3.E55)\)–\([56](https://arxiv.org/html/2608.26877#A3.E56)\) give
det\(K\)𝔼\[\(vT−v\)j\(vT−v\)k\]=det\(Bj⊤Bk\)\.\\det\(K\)\\mathbb\{E\}\[\(v\_\{T\}\-v\)\_\{j\}\(v\_\{T\}\-v\)\_\{k\}\]=\\det\(B\_\{j\}^\{\\top\}B\_\{k\}\)\.\(57\)Ifj=kj=k, expansion through the replaced row and column gives
det\(Bj⊤Bj\)=ρdet\(K−j,−j\)=ρdet\(K\)\(K−1\)jj\.\\det\(B\_\{j\}^\{\\top\}B\_\{j\}\)=\\rho\\det\(K\_\{\-j,\-j\}\)=\\rho\\det\(K\)\(K^\{\-1\}\)\_\{jj\}\.Ifj≠kj\\neq k, rowjjofBj⊤BkB\_\{j\}^\{\\top\}B\_\{k\}has the sole nonzero entryρ\\rhoin columnkk, and hence
det\(Bj⊤Bk\)=\(−1\)j\+kρdet\(K−j,−k\)=ρdet\(K\)\(K−1\)kj\.\\det\(B\_\{j\}^\{\\top\}B\_\{k\}\)=\(\-1\)^\{j\+k\}\\rho\\det\(K\_\{\-j,\-k\}\)=\\rho\\det\(K\)\(K^\{\-1\}\)\_\{kj\}\.Symmetry ofK−1K^\{\-1\}completes every entry of \([54](https://arxiv.org/html/2608.26877#A3.E54)\)\. ∎
Say thatXXis in*row general position*when every indexed set ofddrows is nonsingular\. Assume this condition temporarily\. First sampleSSfrom \([47](https://arxiv.org/html/2608.26877#A3.E47)\); conditional onSS, sample add\-setT⊆ST\\subseteq Sby the rank\-size volume law withinXSX\_\{S\}\. The joint law is
ℙX\(S,T\)=DSZX,sdet\(XT\)2DS=det\(XT\)2\(m−ds−d\)det\(G\)\.\\mathbb\{P\}\_\{X\}\(S,T\)=\\frac\{D\_\{S\}\}\{Z\_\{X,s\}\}\\frac\{\\det\(X\_\{T\}\)^\{2\}\}\{D\_\{S\}\}=\\frac\{\\det\(X\_\{T\}\)^\{2\}\}\{\\binom\{m\-d\}\{s\-d\}\\det\(G\)\}\.\(58\)Consequently
ℙX\(T\)=det\(XT\)2det\(G\),ℙX\(S∣T\)=\(m−ds−d\)−1\(T⊆S\)\.\\mathbb\{P\}\_\{X\}\(T\)=\\frac\{\\det\(X\_\{T\}\)^\{2\}\}\{\\det\(G\)\},\\qquad\\mathbb\{P\}\_\{X\}\(S\\mid T\)=\\binom\{m\-d\}\{s\-d\}^\{\-1\}\\quad\(T\\subseteq S\)\.\(59\)ThusS\|TS\\mid Tis uniform over the size\-sssupersets of the sampled basis\. This basis\-plus\-uniform\-padding representation is also used in the fixed\-size volume\-sampling literature\[[15](https://arxiv.org/html/2608.26877#bib.bib15)\]\.
###### Proposition 10\(Exact row\-general\-position covariance decomposition\)\.
IfXXis in row general position, then
𝔼XwS=w∗,Ms=L∗G−1−𝔼X\[LSGS−1\]\.\\mathbb\{E\}\_\{X\}w\_\{S\}=w^\{\*\},\\qquad M\_\{s\}=L^\{\*\}G^\{\-1\}\-\\mathbb\{E\}\_\{X\}\[L\_\{S\}G\_\{S\}^\{\-1\}\]\.\(60\)
###### Proof\.
Apply Lemma[9](https://arxiv.org/html/2608.26877#Thmtheorem9)conditionally with\(A,b\)=\(XS,yS\)\(A,b\)=\(X\_\{S\},y\_\{S\}\)and then unconditionally with\(A,b\)=\(X,y\)\(A,b\)=\(X,y\)\. The coupling gives
𝔼\[wT∣S\]=wS,Cov\(wT∣S\)=LSGS−1,\\mathbb\{E\}\[w\_\{T\}\\mid S\]=w\_\{S\},\\qquad\\operatorname\{Cov\}\(w\_\{T\}\\mid S\)=L\_\{S\}G\_\{S\}^\{\-1\},\(61\)and
𝔼wT=w∗,Cov\(wT\)=L∗G−1\.\\mathbb\{E\}w\_\{T\}=w^\{\*\},\\qquad\\operatorname\{Cov\}\(w\_\{T\}\)=L^\{\*\}G^\{\-1\}\.\(62\)Iterated expectation proves𝔼wS=w∗\\mathbb\{E\}w\_\{S\}=w^\{\*\}\. The matrix law of total covariance applied to \([61](https://arxiv.org/html/2608.26877#A3.E61)\) and \([62](https://arxiv.org/html/2608.26877#A3.E62)\) gives
L∗G−1=𝔼X\[LSGS−1\]\+Cov\(wS\),L^\{\*\}G^\{\-1\}=\\mathbb\{E\}\_\{X\}\[L\_\{S\}G\_\{S\}^\{\-1\}\]\+\\operatorname\{Cov\}\(w\_\{S\}\),which is \([60](https://arxiv.org/html/2608.26877#A3.E60)\)\. No trace has been taken\. ∎
### C\.3The selected residual first moment
###### Lemma 11\(Selected residual mean\)\.
For every full\-column\-rankXX,m\>dm\>d, andd≤s≤md\\leq s\\leq m,
𝔼XLS=βL∗\.\\mathbb\{E\}\_\{X\}L\_\{S\}=\\beta L^\{\*\}\.\(63\)This identity does not requireL∗\>0L^\{\*\}\>0\.
###### Proof\.
LetH=\[Xy\]∈ℝm×\(d\+1\)H=\[X\\ \\ y\]\\in\\mathbb\{R\}^\{m\\times\(d\+1\)\}\. On a supported set, the Schur complement gives
det\(HS⊤HS\)=DSLS,det\(H⊤H\)=det\(G\)L∗\.\\det\(H\_\{S\}^\{\\top\}H\_\{S\}\)=D\_\{S\}L\_\{S\},\\qquad\\det\(H^\{\\top\}H\)=\\det\(G\)L^\{\*\}\.\(64\)IfDS=0D\_\{S\}=0, thenrank\(XS\)<d\\operatorname\{rank\}\(X\_\{S\}\)<d, sorank\(HS\)≤d\\operatorname\{rank\}\(H\_\{S\}\)\\leq dand the determinant on the left is also zero\. Thus, with unsupported terms understood as zero determinant weights, the first equality may be summed over allSS\.
Fors≥d\+1s\\geq d\+1, fixed\-cardinality Cauchy–Binet applied to thed\+1d\+1columns ofHHyields
∑\|S\|=sdet\(HS⊤HS\)=\(m−d−1s−d−1\)det\(H⊤H\)\.\\sum\_\{\|S\|=s\}\\det\(H\_\{S\}^\{\\top\}H\_\{S\}\)=\\binom\{m\-d\-1\}\{s\-d\-1\}\\det\(H^\{\\top\}H\)\.Ifrank\(H\)≤d\\operatorname\{rank\}\(H\)\\leq d, both sides are zero, so this step does not divide byL∗L^\{\*\}\. Dividing only by the positive ordinary normalizer from Proposition[7](https://arxiv.org/html/2608.26877#Thmtheorem7)gives
𝔼XLS=\(m−d−1s−d−1\)\(m−ds−d\)L∗=s−dm−dL∗\.\\mathbb\{E\}\_\{X\}L\_\{S\}=\\frac\{\\binom\{m\-d\-1\}\{s\-d\-1\}\}\{\\binom\{m\-d\}\{s\-d\}\}L^\{\*\}=\\frac\{s\-d\}\{m\-d\}L^\{\*\}\.Whens=ds=d, every supported square system interpolates and both sides of \([63](https://arxiv.org/html/2608.26877#A3.E63)\) are zero\. ∎
### C\.4The Loewner envelope in row general position
For row\-general\-positionXX, combine Proposition[10](https://arxiv.org/html/2608.26877#Thmtheorem10)and Lemma[11](https://arxiv.org/html/2608.26877#Thmtheorem11)\. SinceGS⪯GG\_\{S\}\\preceq G, inverse order givesGS−1⪰G−1G\_\{S\}^\{\-1\}\\succeq G^\{\-1\}, and hence
αL∗G−1−Ms\\displaystyle\\alpha L^\{\*\}G^\{\-1\}\-M\_\{s\}=𝔼X\[LSGS−1\]−βL∗G−1\\displaystyle=\\mathbb\{E\}\_\{X\}\[L\_\{S\}G\_\{S\}^\{\-1\}\]\-\\beta L^\{\*\}G^\{\-1\}\(65\)=𝔼X\[LS\(GS−1−G−1\)\]⪰0\.\\displaystyle=\\mathbb\{E\}\_\{X\}\\\!\\left\[L\_\{S\}\(G\_\{S\}^\{\-1\}\-G^\{\-1\}\)\\right\]\\succeq 0\.Congruence byG1/2G^\{1/2\}proves both inequalities in \([52](https://arxiv.org/html/2608.26877#A3.E52)\) on this interior class\.
The full\-pool loss consequence is also a matrix corollary\. The normal equationsX⊤e=0X^\{\\top\}e=0imply, for everyww,
∥Xw−y∥22=L∗\+∥w−w∗∥G2\.\\lVert Xw\-y\\rVert\_\{2\}^\{2\}=L^\{\*\}\+\\lVert w\-w^\{\*\}\\rVert\_\{G\}^\{2\}\.\(66\)Therefore
𝔼X∥XwS−y∥22\\displaystyle\\mathbb\{E\}\_\{X\}\\lVert Xw\_\{S\}\-y\\rVert\_\{2\}^\{2\}=L∗\+tr\(GMs\)\\displaystyle=L^\{\*\}\+\\tr\(GM\_\{s\}\)\(67\)=L∗\+tr\(M¯s\)≤\(1\+dα\)L∗,\\displaystyle=L^\{\*\}\+\\tr\(\\overline\{M\}\_\{s\}\)\\leq\(1\+d\\alpha\)L^\{\*\},which proves \([53](https://arxiv.org/html/2608.26877#A3.E53)\) in row general position\.
### C\.5Arbitrary full\-rank boundary designs
Unbiasedness and the covariance inequality require different boundary arguments\. In particular, the exact decomposition \([60](https://arxiv.org/html/2608.26877#A3.E60)\) is not asserted when some selected Gram matrices are singular\.
###### Proposition 12\(Determinant\-weighted continuation of the mean\)\.
Equation \([51](https://arxiv.org/html/2608.26877#A3.E51)\) holds for every full\-column\-rankXX, including designs outside row general position\.
###### Proof\.
For a supported set,
DSwS=adj\(GS\)XS⊤yS\.D\_\{S\}w\_\{S\}=\\operatorname\{adj\}\(G\_\{S\}\)X\_\{S\}^\{\\top\}y\_\{S\}\.\(68\)The right side is polynomial in the entries ofXXandyy\. It vanishes whenGSG\_\{S\}is singular: indeed, withAS=adj\(GS\)XS⊤A\_\{S\}=\\operatorname\{adj\}\(G\_\{S\}\)X\_\{S\}^\{\\top\},
ASAS⊤=adj\(GS\)GSadj\(GS\)=0,A\_\{S\}A\_\{S\}^\{\\top\}=\\operatorname\{adj\}\(G\_\{S\}\)G\_\{S\}\\operatorname\{adj\}\(G\_\{S\}\)=0,soAS=0A\_\{S\}=0\. Thus singular subsets contribute zero to the determinant\-weighted numerator without receiving a selected fit\.
On the dense set of row\-general\-position matrices, Proposition[10](https://arxiv.org/html/2608.26877#Thmtheorem10)and Proposition[7](https://arxiv.org/html/2608.26877#Thmtheorem7)imply
∑\|S\|=sadj\(GS\)XS⊤yS=\(m−ds−d\)adj\(G\)X⊤y\.\\sum\_\{\|S\|=s\}\\operatorname\{adj\}\(G\_\{S\}\)X\_\{S\}^\{\\top\}y\_\{S\}=\\binom\{m\-d\}\{s\-d\}\\operatorname\{adj\}\(G\)X^\{\\top\}y\.\(69\)Both sides are polynomial, so the identity holds for everyXXby polynomial continuation\. For full\-column\-rankXX, divide \([69](https://arxiv.org/html/2608.26877#A3.E69)\) byZX,s=\(m−ds−d\)det\(G\)Z\_\{X,s\}=\\binom\{m\-d\}\{s\-d\}\\det\(G\)and use the preceding zero\-numerator fact\. The result is
𝔼XwS=G−1X⊤y=w∗\.\\mathbb\{E\}\_\{X\}w\_\{S\}=G^\{\-1\}X^\{\\top\}y=w^\{\*\}\.∎
###### Proposition 13\(One\-sided boundary passage for the covariance\)\.
The Loewner inequalities \([52](https://arxiv.org/html/2608.26877#A3.E52)\), and consequently the loss bound \([53](https://arxiv.org/html/2608.26877#A3.E53)\), hold for every full\-column\-rankXX\.
###### Proof\.
LetV∈ℝm×dV\\in\\mathbb\{R\}^\{m\\times d\}be the Vandermonde matrixVij=ij−1V\_\{ij\}=i^\{j\-1\}\. For eachdd\-setTT,det\(XT\+tVT\)\\det\(X\_\{T\}\+tV\_\{T\}\)is a nonzero polynomial intt, since its leading coefficient isdet\(VT\)≠0\\det\(V\_\{T\}\)\\neq 0\. Choose a sequenceta↓0t\_\{a\}\\downarrow 0avoiding the finitely many roots of all these polynomials and setX\(a\)=X\+taVX^\{\(a\)\}=X\+t\_\{a\}V\. Then everyX\(a\)X^\{\(a\)\}is in row general position andX\(a\)→XX^\{\(a\)\}\\to X\.
Use a subscriptaafor quantities formed fromX\(a\)X^\{\(a\)\}, keepingyyfixed\. Fixu∈ℝdu\\in\\mathbb\{R\}^\{d\}, and let𝒫=\{S:\|S\|=s,DS\>0\}\\mathcal\{P\}=\\\{S:\|S\|=s,\\ D\_\{S\}\>0\\\}\. ForS∈𝒫S\\in\\mathcal\{P\}, ordinary inverse continuity gives
DS,a\{u⊤\(wS,a−wa∗\)\}2⟶DS\{u⊤\(wS−w∗\)\}2\.D\_\{S,a\}\\\{u^\{\\top\}\(w\_\{S,a\}\-w\_\{a\}^\{\*\}\)\\\}^\{2\}\\longrightarrow D\_\{S\}\\\{u^\{\\top\}\(w\_\{S\}\-w^\{\*\}\)\\\}^\{2\}\.No selected inverse is continued forS∉𝒫S\\notin\\mathcal\{P\}\. All such omitted perturbed terms are nonnegative, so the finite sum obeys
ZX,su⊤Msu\\displaystyle Z\_\{X,s\}\\,u^\{\\top\}M\_\{s\}u=lima∑S∈𝒫DS,a\{u⊤\(wS,a−wa∗\)\}2\\displaystyle=\\lim\_\{a\}\\sum\_\{S\\in\\mathcal\{P\}\}D\_\{S,a\}\\\{u^\{\\top\}\(w\_\{S,a\}\-w\_\{a\}^\{\*\}\)\\\}^\{2\}\(70\)≤lim infa∑\|S\|=sDS,a\{u⊤\(wS,a−wa∗\)\}2\.\\displaystyle\\leq\\liminf\_\{a\}\\sum\_\{\|S\|=s\}D\_\{S,a\}\\\{u^\{\\top\}\(w\_\{S,a\}\-w\_\{a\}^\{\*\}\)\\\}^\{2\}\.For everyaa, the row\-general\-position envelope bounds the last sum by
ZX\(a\),sαLa∗u⊤Ga−1u\.Z\_\{X^\{\(a\)\},s\}\\,\\alpha L^\{\*\}\_\{a\}\\,u^\{\\top\}G\_\{a\}^\{\-1\}u\.Full column rank persists nearXX, and the normalizer, full fit, residual loss, Gram inverse, and right side all converge ordinarily\. Hence \([70](https://arxiv.org/html/2608.26877#A3.E70)\) gives
u⊤Msu≤αL∗u⊤G−1u\.u^\{\\top\}M\_\{s\}u\\leq\\alpha L^\{\*\}u^\{\\top\}G^\{\-1\}u\.Because this holds for every fixeduu, it is the first inequality in \([52](https://arxiv.org/html/2608.26877#A3.E52)\); congruence gives the second\. Finally, Proposition[12](https://arxiv.org/html/2608.26877#Thmtheorem12)makesMsM\_\{s\}a covariance, and \([66](https://arxiv.org/html/2608.26877#A3.E66)\)–\([67](https://arxiv.org/html/2608.26877#A3.E67)\) give the loss corollary\. The argument retains limiting positive\-volume terms and drops only newly supported nonnegative terms; it does not continue \([60](https://arxiv.org/html/2608.26877#A3.E60)\) to the boundary\. ∎
### C\.6Exact matrix attainment, strict interior slack, and supremum
###### Proposition 14\(Core\-plus\-zero exact matrix attainment\)\.
Fixm\>d≥1m\>d\\geq 1, letq=m−d−1q=m\-d\-1, and define
X0=\[Id𝟏d⊤0q×d\],y=\[𝟏d−10q\]\.X\_\{0\}=\\begin\{bmatrix\}I\_\{d\}\\\\ \\mathbf\{1\}\_\{d\}^\{\\top\}\\\\ 0\_\{q\\times d\}\\end\{bmatrix\},\\qquad y=\\begin\{bmatrix\}\\mathbf\{1\}\_\{d\}\\\\ \-1\\\\ 0\_\{q\}\\end\{bmatrix\}\.\(71\)For everyd≤s≤md\\leq s\\leq m, this family satisfies
Ms=αL∗G−1\.M\_\{s\}=\\alpha L^\{\*\}G^\{\-1\}\.\(72\)
###### Proof\.
The firstd\+1d\+1rows are the core\. Direct calculation gives
G=Id\+𝟏d𝟏d⊤,G−1=Id−𝟏d𝟏d⊤d\+1,w∗=0,L∗=d\+1\.G=I\_\{d\}\+\\mathbf\{1\}\_\{d\}\\mathbf\{1\}\_\{d\}^\{\\top\},\\qquad G^\{\-1\}=I\_\{d\}\-\\frac\{\\mathbf\{1\}\_\{d\}\\mathbf\{1\}\_\{d\}^\{\\top\}\}\{d\+1\},\\qquad w^\{\*\}=0,\\qquad L^\{\*\}=d\+1\.\(73\)A supported size\-ssset contains either exactlyddcore rows or alld\+1d\+1\. Everydd\-core minor has squared determinant one, whereas the all\-core Gram determinant isd\+1d\+1\. With the convention that an out\-of\-range binomial coefficient is zero, the aggregate determinant weights are
Wd=\(d\+1\)\(qs−d\),Wd\+1=\(d\+1\)\(qs−d−1\)\.W\_\{d\}=\(d\+1\)\\binom\{q\}\{s\-d\},\\qquad W\_\{d\+1\}=\(d\+1\)\\binom\{q\}\{s\-d\-1\}\.\(74\)Pascal’s identity therefore gives
ℙX\(exactlydcore rows\)=WdWd\+Wd\+1=m−sm−d=α\.\\mathbb\{P\}\_\{X\}\(\\text\{exactly $d$ core rows\}\)=\\frac\{W\_\{d\}\}\{W\_\{d\}\+W\_\{d\+1\}\}=\\frac\{m\-s\}\{m\-d\}=\\alpha\.\(75\)
On the all\-core event, the normal equations givewS=0w\_\{S\}=0\. Conditional on exactlyddcore rows, the omitted core row is uniform among thed\+1d\+1possibilities and the corresponding interpolating fits are
a0=𝟏d,ai=𝟏d−\(d\+1\)ei,i=1,…,d,a\_\{0\}=\\mathbf\{1\}\_\{d\},\\qquad a\_\{i\}=\\mathbf\{1\}\_\{d\}\-\(d\+1\)e\_\{i\},\\quad i=1,\\ldots,d,\(76\)whereeie\_\{i\}is theiith coordinate vector inℝd\\mathbb\{R\}^\{d\}\. Their mean is zero, and their complete matrix second moment is
1d\+1\(a0a0⊤\+∑i=1daiai⊤\)\\displaystyle\\frac\{1\}\{d\+1\}\\left\(a\_\{0\}a\_\{0\}^\{\\top\}\+\\sum\_\{i=1\}^\{d\}a\_\{i\}a\_\{i\}^\{\\top\}\\right\)=\(d\+1\)Id−𝟏d𝟏d⊤\\displaystyle=\(d\+1\)I\_\{d\}\-\\mathbf\{1\}\_\{d\}\\mathbf\{1\}\_\{d\}^\{\\top\}\(77\)=L∗G−1\.\\displaystyle=L^\{\*\}G^\{\-1\}\.Indeed, expanding the sum gives\(d\+1\)2Id−\(d\+1\)𝟏d𝟏d⊤\(d\+1\)^\{2\}I\_\{d\}\-\(d\+1\)\\mathbf\{1\}\_\{d\}\\mathbf\{1\}\_\{d\}^\{\\top\}before division byd\+1d\+1\. Multiplying \([77](https://arxiv.org/html/2608.26877#A3.E77)\) by the event probability \([75](https://arxiv.org/html/2608.26877#A3.E75)\) proves \([72](https://arxiv.org/html/2608.26877#A3.E72)\), including all off\-diagonal entries\. The same formulas covers=ds=dands=ms=mthroughα=1\\alpha=1andα=0\\alpha=0, respectively\. This proposition supplies an attaining family only\. ∎
###### Proposition 15\(Strict interior slack and row\-general\-position supremum\)\.
IfXXis in row general position,L∗\>0L^\{\*\}\>0,d≥2d\\geq 2, andd<s<md<s<m, then
Ms≺αL∗G−1\.M\_\{s\}\\prec\\alpha L^\{\*\}G^\{\-1\}\.\(78\)For every fixed\(m,d,s\)\(m,d,s\)withm\>dm\>d,d≥2d\\geq 2, andd≤s<md\\leq s<m, row\-general\-position perturbationsX\(a\)→X0X^\{\(a\)\}\\to X\_\{0\}of the family \([71](https://arxiv.org/html/2608.26877#A3.E71)\) can be chosen so that, with the same fixedyy,
‖Ga1/2Ms,aGa1/2La∗−αId‖op⟶0\.\\left\\lVert\\frac\{G\_\{a\}^\{1/2\}M\_\{s,a\}G\_\{a\}^\{1/2\}\}\{L^\{\*\}\_\{a\}\}\-\\alpha I\_\{d\}\\right\\rVert\_\{\\mathrm\{op\}\}\\longrightarrow 0\.\(79\)Thus, ford≥2d\\geq 2andd<s<md<s<m, the coefficient is not attained in row general position but is its supremum\. At the rank\-size endpoints=ds=d, the basis identity already gives equality for every positive\-loss row\-general\-position design; ats=ms=mno normalized ratio is formed\.
###### Proof\.
For strictness, the slack is the positive\-semidefinite expectation in \([65](https://arxiv.org/html/2608.26877#A3.E65)\)\. We show that it has no nonzero null direction\. LetH=\[Xe\]H=\[X\\ \\ e\]\. SinceX⊤e=0X^\{\\top\}e=0andL∗\>0L^\{\*\}\>0,rank\(H\)=d\+1\\operatorname\{rank\}\(H\)=d\+1; sinced≥2d\\geq 2andd<s<md<s<m, alsom≥d\+2m\\geq d\+2\. The rank\-\(d\+1\)\(d\+1\)row matroid ofHHhas a circuit\. No set of at mostddrows ofHHcan be dependent, because projection of such a dependence onto the firstddcoordinates would contradict row general position ofXX\. Hence the circuit contains at leastd\+1d\+1rows\. Every circuit element is a non\-coloop, soHHhas at leastd\+1d\+1non\-coloops\.
Fixu≠0u\\neq 0and putv=G−1u≠0v=G^\{\-1\}u\\neq 0\. At mostd−1d\-1rows can satisfyxi⊤v=0x\_\{i\}^\{\\top\}v=0: anyddsuch rows would form a singulardd\-row submatrix\. Thus some non\-coloopiisatisfiesxi⊤v≠0x\_\{i\}^\{\\top\}v\\neq 0\. Becauseiiis not a coloop,HHhas a row basisRRavoidingii\. Extend it to a size\-sssetR⊆S⊆\[m\]∖\{i\}R\\subseteq S\\subseteq\[m\]\\setminus\\\{i\\\}, possible becaused\+1≤s≤m−1d\+1\\leq s\\leq m\-1\. Thenrank\(HS\)=d\+1\\operatorname\{rank\}\(H\_\{S\}\)=d\+1, and the Schur identity \([64](https://arxiv.org/html/2608.26877#A3.E64)\) together with row general position givesLS\>0L\_\{S\}\>0andDS\>0D\_\{S\}\>0\.
LetCS=G−GS=XSc⊤XScC\_\{S\}=G\-G\_\{S\}=X\_\{S^\{c\}\}^\{\\top\}X\_\{S^\{c\}\}\. Sinceu=Gvu=Gv, direct expansion gives
u⊤\(GS−1−G−1\)u\\displaystyle u^\{\\top\}\(G\_\{S\}^\{\-1\}\-G^\{\-1\}\)u=v⊤\{CS\+CSGS−1CS\}v\\displaystyle=v^\{\\top\}\\\{C\_\{S\}\+C\_\{S\}G\_\{S\}^\{\-1\}C\_\{S\}\\\}v\(80\)≥v⊤CSv=∥XScv∥22\>0,\\displaystyle\\geq v^\{\\top\}C\_\{S\}v=\\lVert X\_\{S^\{c\}\}v\\rVert\_\{2\}^\{2\}\>0,where the last inequality usesi∉Si\\notin Sandxi⊤v≠0x\_\{i\}^\{\\top\}v\\neq 0\. ThisSShas positive probability and positiveLSL\_\{S\}\. Every other summand in \([65](https://arxiv.org/html/2608.26877#A3.E65)\) is positive semidefinite, so the slack is positive in every nonzero directionuu, proving \([78](https://arxiv.org/html/2608.26877#A3.E78)\)\.
For the supremum, use the Vandermonde perturbations from Proposition[13](https://arxiv.org/html/2608.26877#Thmtheorem13)withX=X0X=X\_\{0\}\. The directional liminf \([70](https://arxiv.org/html/2608.26877#A3.E70)\), applied to a fixed basis of coefficient directions and summed, together with exact attainment at the limit, gives
lim infatr\(GaMs,a\)≥dαL∗\.\\liminf\_\{a\}\\tr\(G\_\{a\}M\_\{s,a\}\)\\geq d\\alpha L^\{\*\}\.\(81\)For completeness, one may take the fixed directionsuj=G1/2eju\_\{j\}=G^\{1/2\}e\_\{j\}: their limiting sum istr\(GMs\)=dαL∗\\tr\(GM\_\{s\}\)=d\\alpha L^\{\*\}\. ReplacingGGbyGaG\_\{a\}in the trace changes the sum byo\(1\)o\(1\), becauseGa→GG\_\{a\}\\to GandMs,a⪯αLa∗Ga−1M\_\{s,a\}\\preceq\\alpha L^\{\*\}\_\{a\}G\_\{a\}^\{\-1\}keepsMs,aM\_\{s,a\}bounded\. The envelope also gives the matching upper bound
tr\(GaMs,a\)≤dαLa∗⟶dαL∗\.\\tr\(G\_\{a\}M\_\{s,a\}\)\\leq d\\alpha L^\{\*\}\_\{a\}\\longrightarrow d\\alpha L^\{\*\}\.\(82\)Consequently the matrices
Ra:=αLa∗Id−Ga1/2Ms,aGa1/2R\_\{a\}:=\\alpha L^\{\*\}\_\{a\}I\_\{d\}\-G\_\{a\}^\{1/2\}M\_\{s,a\}G\_\{a\}^\{1/2\}are positive semidefinite and satisfytr\(Ra\)→0\\tr\(R\_\{a\}\)\\to 0\. For a positive\-semidefinite matrix,∥Ra∥op≤tr\(Ra\)\\lVert R\_\{a\}\\rVert\_\{\\mathrm\{op\}\}\\leq\\tr\(R\_\{a\}\)\. SinceLa∗→L∗=d\+1\>0L^\{\*\}\_\{a\}\\to L^\{\*\}=d\+1\>0, division byLa∗L^\{\*\}\_\{a\}proves \([79](https://arxiv.org/html/2608.26877#A3.E79)\)\. Whend≥2d\\geq 2andd<s<md<s<m, combining this convergence with \([78](https://arxiv.org/html/2608.26877#A3.E78)\) distinguishes supremum from attainment\. ∎
## Appendix DResidual\-augmented representation and response\-aware resolvent envelope
This section concerns ordinary, unrescaled, fixed\-size volume sampling of unordered sets of row indices, without replacement, followed by unweighted least squares\. Equal rows at different indices remain different observations\. The response is fixed throughout, and every expectation is conditional on the displayed pair\.
LetX∈ℝm×dX\\in\\mathbb\{R\}^\{m\\times d\}have full column rank and lety∈ℝmy\\in\\mathbb\{R\}^\{m\}\. Write
G=X⊤X,w∗=G−1X⊤y,e=y−Xw∗,L∗=∥e∥22\.G=X^\{\\top\}X,\\qquad w^\{\*\}=G^\{\-1\}X^\{\\top\}y,\\qquad e=y\-Xw^\{\*\},\\qquad L^\{\*\}=\\lVert e\\rVert\_\{2\}^\{2\}\.\(83\)For an indexedss\-setS⊆\[m\]S\\subseteq\[m\], putGS=XS⊤XSG\_\{S\}=X\_\{S\}^\{\\top\}X\_\{S\}andDS=det\(GS\)D\_\{S\}=\\det\(G\_\{S\}\)\. The ordinary law and its fixed\-cardinality Cauchy–Binet normalizer are
ZX,s=∑\|S\|=sDS=\(m−ds−d\)det\(G\),ℙX\(S\)=DSZX,s\.Z\_\{X,s\}=\\sum\_\{\|S\|=s\}D\_\{S\}=\\binom\{m\-d\}\{s\-d\}\\det\(G\),\\qquad\\mathbb\{P\}\_\{X\}\(S\)=\\frac\{D\_\{S\}\}\{Z\_\{X,s\}\}\.\(84\)This ordinary fixed\-size law and its inverse\-Gram machinery are standard in the volume\-sampling literature\[[12](https://arxiv.org/html/2608.26877#bib.bib12),[3](https://arxiv.org/html/2608.26877#bib.bib3),[22](https://arxiv.org/html/2608.26877#bib.bib22)\]; the support and normalizers are displayed here because they are essential to the measure change below\. Only sets withDS\>0D\_\{S\}\>0are estimator\-supported\. On such sets, and only on such sets, define
wS=GS−1XS⊤yS,LS=∥XSwS−yS∥22\.w\_\{S\}=G\_\{S\}^\{\-1\}X\_\{S\}^\{\\top\}y\_\{S\},\\qquad L\_\{S\}=\\lVert X\_\{S\}w\_\{S\}\-y\_\{S\}\\rVert\_\{2\}^\{2\}\.\(85\)There is no pseudoinverse or assigned fit on a zero\-volume set\. For later use, define on every full\-column\-rank design the supported centered second moment
Ms=∑\|S\|=sDS\>0ℙX\(S\)\(wS−w∗\)\(wS−w∗\)⊤\.M\_\{s\}=\\sum\_\{\\begin\{subarray\}\{c\}\|S\|=s\\\\ D\_\{S\}\>0\\end\{subarray\}\}\\mathbb\{P\}\_\{X\}\(S\)\(w\_\{S\}\-w^\{\*\}\)\(w\_\{S\}\-w^\{\*\}\)^\{\\top\}\.\(86\)
The nontrivial domain in this section is
L∗\>0,d<s<m\.L^\{\*\}\>0,\\qquad d<s<m\.\(87\)Thusm≥d\+2m\\geq d\+2, and the constants
α=m−sm−d,β=s−dm−d=1−α,γ=m−sm−d−1\\alpha=\\frac\{m\-s\}\{m\-d\},\\qquad\\beta=\\frac\{s\-d\}\{m\-d\}=1\-\\alpha,\\qquad\\gamma=\\frac\{m\-s\}\{m\-d\-1\}\(88\)are well defined\. The endpoint branches are separate: ifm=dm=d, necessarilys=ms=mand the selected and full fits coincide; ifs=ms=m, the unique sample has zero centered second moment; ifL∗=0L^\{\*\}=0, every supported selected fit equalsw∗w^\{\*\}and no residual direction is formed; and ifs=d<ms=d<m, supported square systems interpolate but the rank\-\(d\+1\)\(d\+1\)augmented law below is not invoked\. In particular, no quotient in \([87](https://arxiv.org/html/2608.26877#A4.E87)\) is silently evaluated at an endpoint\.
### D\.1Whitening and the exact change of measure
Assume first thatXXis in row general position, meaning that every indexeddd\-row submatrix is nonsingular, in addition to \([87](https://arxiv.org/html/2608.26877#A4.E87)\)\. Define
A=XG−1/2,z=eL∗,B=\[Az\]\.A=XG^\{\-1/2\},\\qquad z=\\frac\{e\}\{\\sqrt\{L^\{\*\}\}\},\\qquad B=\[A\\ \\ z\]\.\(89\)The full normal equations give
A⊤A=Id,A⊤z=0,∥z∥2=1,B⊤B=Id\+1\.A^\{\\top\}A=I\_\{d\},\\qquad A^\{\\top\}z=0,\\qquad\\lVert z\\rVert\_\{2\}=1,\\qquad B^\{\\top\}B=I\_\{d\+1\}\.\(90\)For eachSS, letKS=AS⊤ASK\_\{S\}=A\_\{S\}^\{\\top\}A\_\{S\}\. The augmented volume law is a different law, of rankd\+1d\+1, with its own support and normalizer:
ZB,s=∑\|S\|=sdet\(BS⊤BS\)=\(m−d−1s−d−1\),ℙB\(S\)=det\(BS⊤BS\)ZB,s\.Z\_\{B,s\}=\\sum\_\{\|S\|=s\}\\det\(B\_\{S\}^\{\\top\}B\_\{S\}\)=\\binom\{m\-d\-1\}\{s\-d\-1\},\\qquad\\mathbb\{P\}\_\{B\}\(S\)=\\frac\{\\det\(B\_\{S\}^\{\\top\}B\_\{S\}\)\}\{Z\_\{B,s\}\}\.\(91\)HerePBP\_\{B\}is a response\-dependent analysis law induced by the full residual, not the sampler forwSw\_\{S\}; the estimator continues to sample and fit under the ordinary feature\-only lawPXP\_\{X\}\. An augmented\-supported set hasrank\(BS\)=d\+1\\operatorname\{rank\}\(B\_\{S\}\)=d\+1, hencerank\(AS\)=d\\operatorname\{rank\}\(A\_\{S\}\)=dandKS≻0K\_\{S\}\\succ 0\. Thus every inverse under𝔼B\\mathbb\{E\}\_\{B\}below is taken only on augmented support\. No inverse is assigned to a zero\-augmented\-volume set\. Augmented support is contained in, and can be strictly smaller than, ordinary support\.
For an ordinary\-supported set, the Schur determinant formula andyS=XSw∗\+L∗zSy\_\{S\}=X\_\{S\}w^\{\*\}\+\\sqrt\{L^\{\*\}\}z\_\{S\}give
det\(BS⊤BS\)=det\(KS\)\(∥zS∥22−zS⊤ASKS−1AS⊤zS\)=det\(KS\)LSL∗\.\\det\(B\_\{S\}^\{\\top\}B\_\{S\}\)=\\det\(K\_\{S\}\)\\left\(\\lVert z\_\{S\}\\rVert\_\{2\}^\{2\}\-z\_\{S\}^\{\\top\}A\_\{S\}K\_\{S\}^\{\-1\}A\_\{S\}^\{\\top\}z\_\{S\}\\right\)=\\det\(K\_\{S\}\)\\frac\{L\_\{S\}\}\{L^\{\*\}\}\.\(92\)IfSSis not ordinary\-supported, thenrank\(AS\)<d\\operatorname\{rank\}\(A\_\{S\}\)<dandrank\(BS\)≤d\\operatorname\{rank\}\(B\_\{S\}\)\\leq d, so both determinants in the corresponding zero\-equals\-zero rank statement vanish; neitherwSw\_\{S\}norLSL\_\{S\}is defined there\. Alsodet\(KS\)=DS/det\(G\)\\det\(K\_\{S\}\)=D\_\{S\}/\\det\(G\)\. Consequently, if the finite measureμX\\mu\_\{X\}is defined by
μX\(\{S\}\)=\{ℙX\(S\)LS,DS\>0,0,DS=0,\\mu\_\{X\}\(\\\{S\\\}\)=\\begin\{cases\}\\mathbb\{P\}\_\{X\}\(S\)L\_\{S\},&D\_\{S\}\>0,\\\\ 0,&D\_\{S\}=0,\\end\{cases\}\(93\)then the two distinct normalizers in \([84](https://arxiv.org/html/2608.26877#A4.E84)\) and \([91](https://arxiv.org/html/2608.26877#A4.E91)\) yield the exact setwise change of measure
ℙX\(S\)LS=βL∗ℙB\(S\)\(DS\>0\),μX\(\{S\}\)=βL∗ℙB\(S\)\(allS\),ZB,s\(m−ds−d\)=β\.\\begin\{gathered\}\\boxed\{\\mathbb\{P\}\_\{X\}\(S\)L\_\{S\}=\\beta L^\{\*\}\\mathbb\{P\}\_\{B\}\(S\)\\quad\(D\_\{S\}\>0\),\\qquad\\mu\_\{X\}\(\\\{S\\\}\)=\\beta L^\{\*\}\\mathbb\{P\}\_\{B\}\(S\)\\quad\(\\text\{all \}S\)\},\\\\ \\dfrac\{Z\_\{B,s\}\}\{\\binom\{m\-d\}\{s\-d\}\}=\\beta\.\\end\{gathered\}\(94\)The zero branch in \([93](https://arxiv.org/html/2608.26877#A4.E93)\) extends only a measure, not an inverse fit\.
### D\.2Exact row\-general\-position covariance transform
For completeness, the ordinary row\-general\-position basis coupling gives the matrix identity needed here; related basis\-moment and basis\-padding representations appear in\[[11](https://arxiv.org/html/2608.26877#bib.bib11),[12](https://arxiv.org/html/2608.26877#bib.bib12),[15](https://arxiv.org/html/2608.26877#bib.bib15)\]\. Couple a size\-ssordinary volume sampleSSto a size\-ddvolume\-sampled basisT⊆ST\\subseteq S\. The polarized Cauchy–Binet and Cramer calculation for a rank\-size basis gives
𝔼\[wT∣S\]=wS,Cov\(wT∣S\)=LSGS−1,\\mathbb\{E\}\[w\_\{T\}\\mid S\]=w\_\{S\},\\qquad\\operatorname\{Cov\}\(w\_\{T\}\\mid S\)=L\_\{S\}G\_\{S\}^\{\-1\},\(95\)whereas its full\-design version gives
𝔼wT=w∗,Cov\(wT\)=L∗G−1\.\\mathbb\{E\}w\_\{T\}=w^\{\*\},\\qquad\\operatorname\{Cov\}\(w\_\{T\}\)=L^\{\*\}G^\{\-1\}\.\(96\)Iterated expectation proves𝔼XwS=w∗\\mathbb\{E\}\_\{X\}w\_\{S\}=w^\{\*\}, and total covariance gives
Ms=CovX\(wS\)=L∗G−1−𝔼X\[LSGS−1\]\.M\_\{s\}=\\operatorname\{Cov\}\_\{X\}\(w\_\{S\}\)=L^\{\*\}G^\{\-1\}\-\\mathbb\{E\}\_\{X\}\[L\_\{S\}G\_\{S\}^\{\-1\}\]\.\(97\)SinceKS−1=G1/2GS−1G1/2K\_\{S\}^\{\-1\}=G^\{1/2\}G\_\{S\}^\{\-1\}G^\{1/2\}, congruence of \([97](https://arxiv.org/html/2608.26877#A4.E97)\) and then \([94](https://arxiv.org/html/2608.26877#A4.E94)\) prove the exact transform
M¯sL∗=Id−β𝔼B\[KS−1\],M¯s=G1/2MsG1/2\.\\boxed\{\\frac\{\\overline\{M\}\_\{s\}\}\{L^\{\*\}\}=I\_\{d\}\-\\beta\\,\\mathbb\{E\}\_\{B\}\[K\_\{S\}^\{\-1\}\]\},\\qquad\\overline\{M\}\_\{s\}=G^\{1/2\}M\_\{s\}G^\{1/2\}\.\(98\)Equation \([98](https://arxiv.org/html/2608.26877#A4.E98)\) is an exact covariance identity only for row\-general\-positionXXwithL∗\>0L^\{\*\}\>0andd<s<md<s<m\. This restriction is part of the statement, not merely a proof convenience\.
### D\.3Augmented exclusion marginal and the augmented Gram first moment
Letai⊤a\_\{i\}^\{\\top\}andbi⊤b\_\{i\}^\{\\top\}denote rowiiofAAandBB, respectively, and define
hi=∥bi∥22,R=∑i=1m\(1−hi\)aiai⊤\.h\_\{i\}=\\lVert b\_\{i\}\\rVert\_\{2\}^\{2\},\\qquad R=\\sum\_\{i=1\}^\{m\}\(1\-h\_\{i\}\)a\_\{i\}a\_\{i\}^\{\\top\}\.\(99\)The fixed\-cardinality Cauchy–Binet step is part of the ordinary volume\-law machinery\[[12](https://arxiv.org/html/2608.26877#bib.bib12)\]; applying it toB−iB\_\{\-i\}gives the augmented exclusion marginal
ℙB\(i∉S\)\\displaystyle\\mathbb\{P\}\_\{B\}\(i\\notin S\)=\(m−d−2s−d−1\)det\(B−i⊤B−i\)\(m−d−1s−d−1\)\\displaystyle=\\frac\{\\binom\{m\-d\-2\}\{s\-d\-1\}\\det\(B\_\{\-i\}^\{\\top\}B\_\{\-i\}\)\}\{\\binom\{m\-d\-1\}\{s\-d\-1\}\}=m−sm−d−1det\(Id\+1−bibi⊤\)=γ\(1−hi\)\.\\displaystyle=\\frac\{m\-s\}\{m\-d\-1\}\\det\(I\_\{d\+1\}\-b\_\{i\}b\_\{i\}^\{\\top\}\)=\\gamma\(1\-h\_\{i\}\)\.\(100\)The last equality is the rank\-one determinant lemma\. In particular, the augmented exclusion factor isγ\(1−hi\)\\gamma\(1\-h\_\{i\}\), not a reciprocal factor\. Since∑iaiai⊤=Id\\sum\_\{i\}a\_\{i\}a\_\{i\}^\{\\top\}=I\_\{d\}, linearity now gives
𝔼BKS=∑iℙB\(i∈S\)aiai⊤=Id−γR≻0\.\\boxed\{\\mathbb\{E\}\_\{B\}K\_\{S\}=\\sum\_\{i\}\\mathbb\{P\}\_\{B\}\(i\\in S\)a\_\{i\}a\_\{i\}^\{\\top\}=I\_\{d\}\-\\gamma R\\succ 0\.\}\(101\)Strict positivity follows because every augmented\-supportedSShasKS≻0K\_\{S\}\\succ 0and the augmented law has nonempty support\.
### D\.4Classical operator Jensen and two resolvent bounds
The classical operator Jensen inequality for the inversion map on the positive\-definite cone\[[20](https://arxiv.org/html/2608.26877#bib.bib20)\]gives
𝔼B\[KS−1\]⪰\(𝔼BKS\)−1=\(Id−γR\)−1\.\\mathbb\{E\}\_\{B\}\[K\_\{S\}^\{\-1\}\]\\succeq\(\\mathbb\{E\}\_\{B\}K\_\{S\}\)^\{\-1\}=\(I\_\{d\}\-\\gamma R\)^\{\-1\}\.\(102\)This generic Jensen step is classical; the volume\-specific inputs are the change of measure and augmented exclusion first moment above\. Substitution into \([98](https://arxiv.org/html/2608.26877#A4.E98)\) reverses the final order because of its minus sign and proves
M¯s⪯L∗\[Id−β\(Id−γR\)−1\]\.\\boxed\{\\overline\{M\}\_\{s\}\\preceq L^\{\*\}\\left\[I\_\{d\}\-\\beta\(I\_\{d\}\-\\gamma R\)^\{\-1\}\\right\]\.\}\(103\)Moreover,R⪰0R\\succeq 0andId−γR≻0I\_\{d\}\-\\gamma R\\succ 0, so scalar functional calculus gives\(Id−γR\)−1⪰Id\+γR\(I\_\{d\}\-\\gamma R\)^\{\-1\}\\succeq I\_\{d\}\+\\gamma R\. Hence the linear corollary used as the response\-uniform slack input to Theorem[3](https://arxiv.org/html/2608.26877#Thmtheorem3)is
M¯s⪯αL∗Id−βγL∗R\.\\boxed\{\\overline\{M\}\_\{s\}\\preceq\\alpha L^\{\*\}I\_\{d\}\-\\beta\\gamma L^\{\*\}R\.\}\(104\)The preceding change\-of\-measure and augmented exclusion\-moment identities use the volume\-law structure; no high\-dimensional sharpness statement is made for \([103](https://arxiv.org/html/2608.26877#A4.E103)\)\.
### D\.5Second\-centered\-Gram refinement of the response\-aware resolvent
###### Proposition 16\(Contraction\-specific second\-moment refinement\)\.
Assume thatXXhas full column rank and thatL∗\>0L^\{\*\}\>0andd<s<md<s<m\. LetA,z,BA,z,BandℙB\\mathbb\{P\}\_\{B\}be as in \([89](https://arxiv.org/html/2608.26877#A4.E89)\)–\([91](https://arxiv.org/html/2608.26877#A4.E91)\), and write
ΩB=\{S⊆\[m\]:\|S\|=s,det\(BS⊤BS\)\>0\}\.\\Omega\_\{B\}=\\\{S\\subseteq\[m\]:\|S\|=s,\\ \\det\(B\_\{S\}^\{\\top\}B\_\{S\}\)\>0\\\}\.\(105\)The lawℙB\\mathbb\{P\}\_\{B\}is analysis\-only, and every selected inverse below is restricted toΩB\\Omega\_\{B\}\. Define
Hs=𝔼BKS,ΔS=KS−Hs,Vs=𝔼B\[ΔS2\],𝒬s=Hs−1VsHs−1\.H\_\{s\}=\\mathbb\{E\}\_\{B\}K\_\{S\},\\qquad\\Delta\_\{S\}=K\_\{S\}\-H\_\{s\},\\qquad V\_\{s\}=\\mathbb\{E\}\_\{B\}\[\\Delta\_\{S\}^\{2\}\],\\qquad\\mathcal\{Q\}\_\{s\}=H\_\{s\}^\{\-1\}V\_\{s\}H\_\{s\}^\{\-1\}\.\(106\)Then0≺KS⪯Id0\\prec K\_\{S\}\\preceq I\_\{d\}onΩB\\Omega\_\{B\}, and
𝔼B\[KS−1\]−Hs−1=𝒬s\+Hs−1𝔼B\[ΔS\(KS−1−Id\)ΔS\]Hs−1⪰𝒬s⪰0\.\\mathbb\{E\}\_\{B\}\[K\_\{S\}^\{\-1\}\]\-H\_\{s\}^\{\-1\}=\\mathcal\{Q\}\_\{s\}\+H\_\{s\}^\{\-1\}\\mathbb\{E\}\_\{B\}\\\!\\left\[\\Delta\_\{S\}\(K\_\{S\}^\{\-1\}\-I\_\{d\}\)\\Delta\_\{S\}\\right\]H\_\{s\}^\{\-1\}\\succeq\\mathcal\{Q\}\_\{s\}\\succeq 0\.\(107\)The coefficient ofVsV\_\{s\}is one: this is an exact ordered inverse remainder, not a Taylor approximation\. IfXXis in row general position, the exact covariance transform gives
M¯s⪯L∗\[Id−βHs−1−β𝒬s\]⪯L∗\[Id−βHs−1\]\.\\overline\{M\}\_\{s\}\\preceq L^\{\*\}\\left\[I\_\{d\}\-\\beta H\_\{s\}^\{\-1\}\-\\beta\\mathcal\{Q\}\_\{s\}\\right\]\\preceq L^\{\*\}\\left\[I\_\{d\}\-\\beta H\_\{s\}^\{\-1\}\\right\]\.\(108\)
The correction is computable from singleton and pair exclusions\. PutT=ScT=S^\{c\},k=m−sk=m\-s,q=m−d−1q=m\-d\-1,Di=aiai⊤D\_\{i\}=a\_\{i\}a\_\{i\}^\{\\top\}, and retainhi=∥bi∥22h\_\{i\}=\\lVert b\_\{i\}\\rVert\_\{2\}^\{2\}\. The singleton exclusion probability is
πi\\displaystyle\\pi\_\{i\}=ℙB\(i∈T\)=kq\(1−hi\),\\displaystyle=\\mathbb\{P\}\_\{B\}\(i\\in T\)=\\frac\{k\}\{q\}\(1\-h\_\{i\}\),\(109\)and, forq≥2q\\geq 2, the pair exclusion probability is
πij\\displaystyle\\pi\_\{ij\}=ℙB\(i,j∈T\)=k\(k−1\)q\(q−1\)\[\(1−hi\)\(1−hj\)−\(bi⊤bj\)2\]\.\\displaystyle=\\mathbb\{P\}\_\{B\}\(i,j\\in T\)=\\frac\{k\(k\-1\)\}\{q\(q\-1\)\}\\left\[\(1\-h\_\{i\}\)\(1\-h\_\{j\}\)\-\(b\_\{i\}^\{\\top\}b\_\{j\}\)^\{2\}\\right\]\.\(110\)Whenm=d\+2m=d\+2, the strict domain forcesq=k=1q=k=1; in that branch setπij=0\\pi\_\{ij\}=0rather than evaluating \([110](https://arxiv.org/html/2608.26877#A4.E110)\)\. In either branch,
Vs=∑iπi\(1−πi\)Di2\+∑i<j\(πij−πiπj\)\(DiDj\+DjDi\)\.V\_\{s\}=\\sum\_\{i\}\\pi\_\{i\}\(1\-\\pi\_\{i\}\)D\_\{i\}^\{2\}\+\\sum\_\{i<j\}\(\\pi\_\{ij\}\-\\pi\_\{i\}\\pi\_\{j\}\)\(D\_\{i\}D\_\{j\}\+D\_\{j\}D\_\{i\}\)\.\(111\)
For every full\-column\-rank design in the same strict domain, without a row general position assumption, define
Fs=𝔼BGS,Ws=𝔼B\[\(GS−Fs\)G−1\(GS−Fs\)\]⪰0\.F\_\{s\}=\\mathbb\{E\}\_\{B\}G\_\{S\},\\qquad W\_\{s\}=\\mathbb\{E\}\_\{B\}\[\(G\_\{S\}\-F\_\{s\}\)G^\{\-1\}\(G\_\{S\}\-F\_\{s\}\)\]\\succeq 0\.\(112\)For the ordinary\-law supported, full\-fit\-centered second momentMsM\_\{s\}in \([86](https://arxiv.org/html/2608.26877#A4.E86)\), only the one\-sided hierarchy
Ms⪯L∗\[G−1−βFs−1−βFs−1WsFs−1\]⪯L∗\[G−1−βFs−1\]M\_\{s\}\\preceq L^\{\*\}\\left\[G^\{\-1\}\-\\beta F\_\{s\}^\{\-1\}\-\\beta F\_\{s\}^\{\-1\}W\_\{s\}F\_\{s\}^\{\-1\}\\right\]\\preceq L^\{\*\}\\left\[G^\{\-1\}\-\\beta F\_\{s\}^\{\-1\}\\right\]\(113\)is asserted\. Outside the row\-general\-position interior this is not an exact inverse\-moment covariance identity or an exact covariance\-gap decomposition\. The proposition is not invoked at the existing zero branchesL∗=0L^\{\*\}=0,s=ms=m, orm=dm=d, nor ats=d<ms=d<m; whenm=d\+1m=d\+1the strict domain is empty\.
###### Proof\.
ForS∈ΩBS\\in\\Omega\_\{B\}, augmented support impliesKS≻0K\_\{S\}\\succ 0, whileId−KS=ASc⊤ASc⪰0I\_\{d\}\-K\_\{S\}=A\_\{S^\{c\}\}^\{\\top\}A\_\{S^\{c\}\}\\succeq 0\. ThusKSK\_\{S\}is a positive contraction andHs≻0H\_\{s\}\\succ 0\. WithK=KSK=K\_\{S\},H=HsH=H\_\{s\}, andΔ=K−H\\Delta=K\-H, direct multiplication, without commuting any factors, gives
K−1=H−1−H−1ΔH−1\+H−1ΔK−1ΔH−1\.K^\{\-1\}=H^\{\-1\}\-H^\{\-1\}\\Delta H^\{\-1\}\+H^\{\-1\}\\Delta K^\{\-1\}\\Delta H^\{\-1\}\.\(114\)Taking expectations removes the centered linear term\. SplittingKS−1=Id\+\(KS−1−Id\)K\_\{S\}^\{\-1\}=I\_\{d\}\+\(K\_\{S\}^\{\-1\}\-I\_\{d\}\)in the last term yields \([107](https://arxiv.org/html/2608.26877#A4.E107)\)\. Its remainder is positive semidefinite becauseKS−1−Id⪰0K\_\{S\}^\{\-1\}\-I\_\{d\}\\succeq 0\. Thus the first\-order comparison is the classical operator\-Jensen step for inversion\[[20](https://arxiv.org/html/2608.26877#bib.bib20)\], whereas \([107](https://arxiv.org/html/2608.26877#A4.E107)\) is the support\-restricted resolvent calculation for these contractions\. Combining it with \([98](https://arxiv.org/html/2608.26877#A4.E98)\) proves \([108](https://arxiv.org/html/2608.26877#A4.E108)\)\.
For the marginal formulas, fixed\-cardinality Cauchy–Binet applied after deleting one or two rows gives the support\-aware specialization of the generic dual\-volume inclusion\-marginal mechanism\[[22](https://arxiv.org/html/2608.26877#bib.bib22)\]\. In particular,
ℙB\(i∈T\)\\displaystyle\\mathbb\{P\}\_\{B\}\(i\\in T\)=\(q−1s−d−1\)\(qs−d−1\)det\(Id\+1−bibi⊤\),\\displaystyle=\\frac\{\\binom\{q\-1\}\{s\-d\-1\}\}\{\\binom\{q\}\{s\-d\-1\}\}\\det\(I\_\{d\+1\}\-b\_\{i\}b\_\{i\}^\{\\top\}\),ℙB\(i,j∈T\)\\displaystyle\\mathbb\{P\}\_\{B\}\(i,j\\in T\)=\(q−2s−d−1\)\(qs−d−1\)det\(Id\+1−bibi⊤−bjbj⊤\),\\displaystyle=\\frac\{\\binom\{q\-2\}\{s\-d\-1\}\}\{\\binom\{q\}\{s\-d\-1\}\}\\det\(I\_\{d\+1\}\-b\_\{i\}b\_\{i\}^\{\\top\}\-b\_\{j\}b\_\{j\}^\{\\top\}\),\(115\)where the second line is used only forq≥2q\\geq 2\. The rank\-one and rank\-two determinant lemmas give \([109](https://arxiv.org/html/2608.26877#A4.E109)\) and \([110](https://arxiv.org/html/2608.26877#A4.E110)\)\. Ifq=k=1q=k=1, the excluded set has size one, so its pair\-exclusion probability is zero directly\. Finally,KS=Id−∑i∈TDiK\_\{S\}=I\_\{d\}\-\\sum\_\{i\\in T\}D\_\{i\}; expanding the square of its centered version retains both product orders and gives the anticommutator in \([111](https://arxiv.org/html/2608.26877#A4.E111)\)\.
On a row\-general\-position design,Fs=G1/2HsG1/2F\_\{s\}=G^\{1/2\}H\_\{s\}G^\{1/2\}andWs=G1/2VsG1/2W\_\{s\}=G^\{1/2\}V\_\{s\}G^\{1/2\}, so congruence of \([108](https://arxiv.org/html/2608.26877#A4.E108)\) gives the raw\-coordinate hierarchy\. The following subsection supplies the directional\-liminf passage that retains only this one\-sided hierarchy on an arbitrary full\-column\-rank boundary, without continuing any selected inverse through zero support\. ∎
### D\.6One\-sided extension to arbitrary full\-column\-rank boundary designs
Now letXXbe an arbitrary full\-column\-rank design, still withL∗\>0L^\{\*\}\>0andd<s<md<s<m; row general position is no longer assumed\. The quantitiesA,z,B,hiA,z,B,h\_\{i\}remain defined by \([89](https://arxiv.org/html/2608.26877#A4.E89)\) and \([99](https://arxiv.org/html/2608.26877#A4.E99)\)\. In original coordinates put
Q=∑i=1m\(1−hi\)xixi⊤\.Q=\\sum\_\{i=1\}^\{m\}\(1\-h\_\{i\}\)x\_\{i\}x\_\{i\}^\{\\top\}\.\(116\)The augmented exclusion calculation itself did not require row general position\. Because augmented support impliesGS≻0G\_\{S\}\\succ 0, it gives the exact first moment
Fs=𝔼BGS=G−γQ≻0\.F\_\{s\}=\\mathbb\{E\}\_\{B\}G\_\{S\}=G\-\\gamma Q\\succ 0\.\(117\)The first\-order member of Proposition[16](https://arxiv.org/html/2608.26877#Thmtheorem16)is therefore the response\-aware upper inequality
Ms⪯L∗\[G−1−β\(G−γQ\)−1\]\.\\boxed\{M\_\{s\}\\preceq L^\{\*\}\\left\[G^\{\-1\}\-\\beta\(G\-\\gamma Q\)^\{\-1\}\\right\]\.\}\(118\)
To justify the extension without continuing a singular selected inverse, choose row\-general\-position matricesXt→XX\_\{t\}\\to X\. Withyyfixed, all full quantitiesGt,wt∗,Lt∗,hi,t,QtG\_\{t\},w\_\{t\}^\{\*\},L\_\{t\}^\{\*\},h\_\{i,t\},Q\_\{t\}converge to their unperturbed counterparts;Lt∗\>0L\_\{t\}^\{\*\}\>0for all sufficiently largett\. For a fixed coefficient directionuu, write the ordinary centered\-second\-moment numerator as the finite sum
∑\|S\|=sDS\>0DS\{u⊤\(wS−w∗\)\}2\.\\sum\_\{\\begin\{subarray\}\{c\}\|S\|=s\\\\ D\_\{S\}\>0\\end\{subarray\}\}D\_\{S\}\\\{u^\{\\top\}\(w\_\{S\}\-w^\{\*\}\)\\\}^\{2\}\.\(119\)Every summand displayed in \([119](https://arxiv.org/html/2608.26877#A4.E119)\) is the limit of its perturbed counterpart\. Sets withDS=0D\_\{S\}=0contribute no boundary estimator and their perturbed contributions are nonnegative\. Dropping precisely those newly supported terms therefore gives the one\-sided comparison
u⊤Msu≤lim inft→0u⊤Ms,tu\.u^\{\\top\}M\_\{s\}u\\leq\\liminf\_\{t\\to 0\}u^\{\\top\}M\_\{s,t\}u\.\(120\)The ordinary normalizer converges\. The finite\-sum definitions ofFsF\_\{s\}andWsW\_\{s\}in \([112](https://arxiv.org/html/2608.26877#A4.E112)\) are continuous under this perturbation, and augmented support givesFs≻0F\_\{s\}\\succ 0; hence both right\-hand sides of \([113](https://arxiv.org/html/2608.26877#A4.E113)\) are continuous\. Applying the row\-general\-position hierarchy toXtX\_\{t\}and passing to the directional liminf proves \([113](https://arxiv.org/html/2608.26877#A4.E113)\), and in particular \([118](https://arxiv.org/html/2608.26877#A4.E118)\)\. The same passage applied to \([104](https://arxiv.org/html/2608.26877#A4.E104)\), or the elementary inequality\(G−γQ\)−1⪰G−1\+γG−1QG−1\(G\-\\gamma Q\)^\{\-1\}\\succeq G^\{\-1\}\+\\gamma G^\{\-1\}QG^\{\-1\}, gives its original\-coordinate linear form
Ms⪯αL∗G−1−βγL∗G−1QG−1\.M\_\{s\}\\preceq\\alpha L^\{\*\}G^\{\-1\}\-\\beta\\gamma L^\{\*\}G^\{\-1\}QG^\{\-1\}\.\(121\)
###### Proposition 17\(Boundary failure of the exact inverse\-moment identity\)\.
Let
X=\[1001100100\],y=\(1,−1,2,0,3\)⊤,s=3\.X=\\begin\{bmatrix\}1&0\\\\ 0&1\\\\ 1&0\\\\ 0&1\\\\ 0&0\\end\{bmatrix\},\\qquad y=\(1,\-1,2,0,3\)^\{\\top\},\\qquad s=3\.\(122\)ThenG=2I2G=2I\_\{2\},L∗=10L^\{\*\}=10, andMs=I2/6M\_\{s\}=I\_\{2\}/6, whereas
𝔼B\[GS−1\]=3940I2,Ms−L∗\{G−1−β𝔼B\[GS−1\]\}=−1912I2≠0\.\\mathbb\{E\}\_\{B\}\[G\_\{S\}^\{\-1\}\]=\\frac\{39\}\{40\}I\_\{2\},\\qquad M\_\{s\}\-L^\{\*\}\\left\\\{G^\{\-1\}\-\\beta\\mathbb\{E\}\_\{B\}\[G\_\{S\}^\{\-1\}\]\\right\\\}=\-\\frac\{19\}\{12\}I\_\{2\}\\neq 0\.\(123\)The boundary resolvent inequality nevertheless remains strict, with slack\(209/126\)I2≻0\(209/126\)I\_\{2\}\\succ 0\.
###### Proof\.
Herew∗=\(3/2,−1/2\)⊤w^\{\*\}=\(3/2,\-1/2\)^\{\\top\}ande=\(−1/2,−1/2,1/2,1/2,3\)⊤e=\(\-1/2,\-1/2,1/2,1/2,3\)^\{\\top\}, which gives the statedGGandL∗L^\{\*\}\. PutC1=\{1,3\}C\_\{1\}=\\\{1,3\\\}andC2=\{2,4\}C\_\{2\}=\\\{2,4\\\}\. The eight ordinary\-supported triples, whose normalizer isZX,3=12Z\_\{X,3\}=12, split as follows:
All four sign combinations occur in the last row\. Weighting byℙX\(S\)=DS/12\\mathbb\{P\}\_\{X\}\(S\)=D\_\{S\}/12, the three rows contribute respectivelydiag\(0,1/12\)\\operatorname\{diag\}\(0,1/12\),diag\(1/12,0\)\\operatorname\{diag\}\(1/12,0\), andI2/12I\_\{2\}/12to the centered second moment\. HenceMs=I2/6M\_\{s\}=I\_\{2\}/6\.
For the augmented law,ZB,3=1Z\_\{B,3\}=1\. The setwise Schur determinant gives
ℙB\(S\)=det\(BS⊤BS\)=DSLSdet\(G\)L∗=DSLS40\.\\mathbb\{P\}\_\{B\}\(S\)=\\det\(B\_\{S\}^\{\\top\}B\_\{S\}\)=\\frac\{D\_\{S\}L\_\{S\}\}\{\\det\(G\)L^\{\*\}\}=\\frac\{D\_\{S\}L\_\{S\}\}\{40\}\.\(124\)Thus the per\-triple probabilities in the three rows are, respectively,1/401/40,1/401/40, and9/409/40\. Their inverse Gramians arediag\(1/2,1\)\\operatorname\{diag\}\(1/2,1\),diag\(1,1/2\)\\operatorname\{diag\}\(1,1/2\), andI2I\_\{2\}\. Exact summation therefore gives
𝔼B\[GS−1\]=240diag\(1/2,1\)\+240diag\(1,1/2\)\+3640I2=3940I2\.\\mathbb\{E\}\_\{B\}\[G\_\{S\}^\{\-1\}\]=\\frac\{2\}\{40\}\\operatorname\{diag\}\(1/2,1\)\+\\frac\{2\}\{40\}\\operatorname\{diag\}\(1,1/2\)\+\\frac\{36\}\{40\}I\_\{2\}=\\frac\{39\}\{40\}I\_\{2\}\.\(125\)Sinceβ=1/3\\beta=1/3, substitution yields the defect in \([123](https://arxiv.org/html/2608.26877#A4.E123)\)\.
Finally,hi=xi⊤G−1xi\+ei2/L∗=21/40h\_\{i\}=x\_\{i\}^\{\\top\}G^\{\-1\}x\_\{i\}\+e\_\{i\}^\{2\}/L^\{\*\}=21/40fori≤4i\\leq 4, whileh5=9/10h\_\{5\}=9/10\. Henceγ=1\\gamma=1,Q=\(19/20\)I2Q=\(19/20\)I\_\{2\}, andG−γQ=\(21/20\)I2G\-\\gamma Q=\(21/20\)I\_\{2\}\. The exact boundary slack is therefore
L∗\[G−1−β\(G−γQ\)−1\]−Ms=10\(12−132021\)I2−16I2=209126I2≻0\.L^\{\*\}\\left\[G^\{\-1\}\-\\beta\(G\-\\gamma Q\)^\{\-1\}\\right\]\-M\_\{s\}=10\\left\(\\frac\{1\}\{2\}\-\\frac\{1\}\{3\}\\frac\{20\}\{21\}\\right\)I\_\{2\}\-\\frac\{1\}\{6\}I\_\{2\}=\\frac\{209\}\{126\}I\_\{2\}\\succ 0\.\(126\)∎
## Appendix EProof of the robust pre\-response phase boundary
This section gives the support calculation behind Theorem[3](https://arxiv.org/html/2608.26877#Thmtheorem3)\. It uses ordinary unrescaled indexed fixed\-size volume sampling and selected unweighted least squares\. Zero\-volume sets have zero probability and receive no selected fit\. The proof has four steps: compactness and positive semidefiniteness; the response\-uniform one\-sided slack bound; attainment from a zero\-margin witness; and the converse implication from envelope equality\.
attainment and PSD⟶uniform slack,zero\-margin support saturation⟶converse equality argument\.\\begin\{gathered\}\\text\{attainment and PSD\}\\longrightarrow\\text\{uniform slack\},\\\\\[\-2\.0pt\] \\text\{zero\-margin support saturation\}\\longrightarrow\\text\{converse equality argument\}\.\\end\{gathered\}\(127\)
### E\.1Domain, support, and attainment
Throughout this proof,A∈ℝm×dA\\in\\mathbb\{R\}^\{m\\times d\}has orthonormal columns,d≥1d\\geq 1,m≥d\+2m\\geq d\+2, andd<s<md<s<m\. For anyθ∈ℝd\\theta\\in\\mathbb\{R\}^\{d\}, let
𝒵A=\{z∈ker\(A⊤\):∥z∥2=1\},yz=Aθ\+L∗z,L∗\>0\.\\mathcal\{Z\}\_\{A\}=\\\{z\\in\\ker\(A^\{\\top\}\):\\lVert z\\rVert\_\{2\}=1\\\},\\qquad y\_\{z\}=A\\theta\+\\sqrt\{L^\{\*\}\}z,\\qquad L^\{\*\}\>0\.\(128\)The residual sphere is nonempty and compact becausem\>dm\>d\. For a positive\- volume indexed setSS, define
pA\(S\)=det\(AS⊤AS\)\(m−ds−d\),wS\(z\)=\(AS⊤AS\)−1AS⊤\(yz\)S\.p\_\{A\}\(S\)=\\frac\{\\det\(A\_\{S\}^\{\\top\}A\_\{S\}\)\}\{\\binom\{m\-d\}\{s\-d\}\},\\qquad w\_\{S\}\(z\)=\(A\_\{S\}^\{\\top\}A\_\{S\}\)^\{\-1\}A\_\{S\}^\{\\top\}\(y\_\{z\}\)\_\{S\}\.\(129\)The fixed\-cardinality Cauchy–Binet identity makes the displayed probabilities sum to one over all indexedss\-sets; zero determinants simply contribute zero\. This is the ordinary volume law and its support convention\. For fixedAAandSS,wS\(z\)w\_\{S\}\(z\)is affine inzz, so the finite sum
M¯sA,z=∑\|S\|=sdet\(AS⊤AS\)\>0pA\(S\)\(wS\(z\)−θ\)\(wS\(z\)−θ\)⊤\\overline\{M\}\_\{s\}^\{A,z\}=\\sum\_\{\\begin\{subarray\}\{c\}\|S\|=s\\\\ \\det\(A\_\{S\}^\{\\top\}A\_\{S\}\)\>0\\end\{subarray\}\}p\_\{A\}\(S\)\(w\_\{S\}\(z\)\-\\theta\)\(w\_\{S\}\(z\)\-\\theta\)^\{\\top\}\(130\)is continuous inzz\. Hence both the minimum definingνA\\nu\_\{A\}and the maximum defining𝖳specRU\(A,s\)\\mathsf\{T\}\_\{\\mathrm\{spec\}\}^\{\\mathrm\{RU\}\}\(A,s\)are attained\.
Forz∈𝒵Az\\in\\mathcal\{Z\}\_\{A\}, the matrixB=\[Az\]B=\[A\\ z\]has orthonormal columns\. ConsequentlyBB⊤BB^\{\\top\}is an orthogonal projection and
hi:=\(BB⊤\)ii=ℓi\+zi2≤1\.h\_\{i\}:=\(BB^\{\\top\}\)\_\{ii\}=\\ell\_\{i\}\+z\_\{i\}^\{2\}\\leq 1\.\(131\)It follows that
RA\(z\)=∑i=1m\(1−hi\)aiai⊤⪰0,νA≥0\.R\_\{A\}\(z\)=\\sum\_\{i=1\}^\{m\}\(1\-h\_\{i\}\)a\_\{i\}a\_\{i\}^\{\\top\}\\succeq 0,\\qquad\\nu\_\{A\}\\geq 0\.\(132\)
### E\.2Uniform slack and the strict direction
The response\-aware covariance relation, in its boundary\-valid one\-sided form, gives for everyz∈𝒵Az\\in\\mathcal\{Z\}\_\{A\}
M¯sA,zL∗⪯αsId−βsγsRA\(z\),αs=m−sm−d,βs=s−dm−d,γs=m−sm−d−1\.\\frac\{\\overline\{M\}\_\{s\}^\{A,z\}\}\{L^\{\*\}\}\\preceq\\alpha\_\{s\}I\_\{d\}\-\\beta\_\{s\}\\gamma\_\{s\}R\_\{A\}\(z\),\\qquad\\alpha\_\{s\}=\\frac\{m\-s\}\{m\-d\},\\quad\\beta\_\{s\}=\\frac\{s\-d\}\{m\-d\},\\quad\\gamma\_\{s\}=\\frac\{m\-s\}\{m\-d\-1\}\.\(133\)This step is valid for arbitrary full\-column\-rank designs after whitening; it does not continue an exact inverse\-moment identity through a rank\-changing support boundary\. SinceRA\(z\)⪰νAIdR\_\{A\}\(z\)\\succeq\\nu\_\{A\}I\_\{d\}for every admissiblezz, \([133](https://arxiv.org/html/2608.26877#A5.E133)\) implies
M¯sA,zL∗⪯\(αs−βsγsνA\)Id\.\\frac\{\\overline\{M\}\_\{s\}^\{A,z\}\}\{L^\{\*\}\}\\preceq\(\\alpha\_\{s\}\-\\beta\_\{s\}\\gamma\_\{s\}\\nu\_\{A\}\)I\_\{d\}\.\(134\)Taking the largest eigenvalue and then the maximum overzzproves \([27](https://arxiv.org/html/2608.26877#S5.E27)\)\. Becauseβsγs\>0\\beta\_\{s\}\\gamma\_\{s\}\>0, it also provesνA\>0⇒𝖳specRU\(A,s\)<αs\\nu\_\{A\}\>0\\Rightarrow\\mathsf\{T\}\_\{\\mathrm\{spec\}\}^\{\\mathrm\{RU\}\}\(A,s\)<\\alpha\_\{s\}\.
### E\.3A zero\-margin witness attains the envelope
Assume now thatνA=0\\nu\_\{A\}=0\. By attainment and positive semidefiniteness, there arez∈𝒵Az\\in\\mathcal\{Z\}\_\{A\}and a unit vectorv∈ℝdv\\in\\mathbb\{R\}^\{d\}such that
v⊤RA\(z\)v=0\.v^\{\\top\}R\_\{A\}\(z\)v=0\.\(135\)Set
ci=ai⊤v,J=\{i:ci≠0\}\.c\_\{i\}=a\_\{i\}^\{\\top\}v,\\qquad J=\\\{i:c\_\{i\}\\neq 0\\\}\.\(136\)Using \([131](https://arxiv.org/html/2608.26877#A5.E131)\), every term in
0=v⊤RA\(z\)v=∑i\(1−hi\)ci20=v^\{\\top\}R\_\{A\}\(z\)v=\\sum\_\{i\}\(1\-h\_\{i\}\)c\_\{i\}^\{2\}\(137\)is nonnegative\. Therefore
hi=1\(i∈J\),∑i∈Jci2=∥Av∥22=1\.h\_\{i\}=1\\quad\(i\\in J\),\\qquad\\sum\_\{i\\in J\}c\_\{i\}^\{2\}=\\lVert Av\\rVert\_\{2\}^\{2\}=1\.\(138\)The no\-coloop condition gives
zi2=1−ℓi\>0\(i∈J\)\.z\_\{i\}^\{2\}=1\-\\ell\_\{i\}\>0\\quad\(i\\in J\)\.\(139\)
For a saturated row,BB⊤ei=eiBB^\{\\top\}e\_\{i\}=e\_\{i\}, hence
ei=Aai\+zzi\.e\_\{i\}=Aa\_\{i\}\+zz\_\{i\}\.\(140\)LetSSbe a positive\-volume set and suppose it omitsi∈Ji\\in J\. Restricting \([140](https://arxiv.org/html/2608.26877#A5.E140)\) toSSgiveszS=−\(1/zi\)ASaiz\_\{S\}=\-\(1/z\_\{i\}\)A\_\{S\}a\_\{i\}\. The selected least\-squares equations then yield
wS\(z\)−θL∗=\(AS⊤AS\)−1AS⊤zS=−aizi\.\\frac\{w\_\{S\}\(z\)\-\\theta\}\{\\sqrt\{L^\{\*\}\}\}=\(A\_\{S\}^\{\\top\}A\_\{S\}\)^\{\-1\}A\_\{S\}^\{\\top\}z\_\{S\}=\-\\frac\{a\_\{i\}\}\{z\_\{i\}\}\.\(141\)
No positive\-volume set can omit two distinct membersi,j∈Ji,j\\in J\. Otherwise \([141](https://arxiv.org/html/2608.26877#A5.E141)\) would giveai/zi=aj/zja\_\{i\}/z\_\{i\}=a\_\{j\}/z\_\{j\}\. At the same time, two distinct saturated rows are orthogonal in the augmented space, so
ai⊤aj\+zizj=0\.a\_\{i\}^\{\\top\}a\_\{j\}\+z\_\{i\}z\_\{j\}=0\.\(142\)Writingai=ziqa\_\{i\}=z\_\{i\}qandaj=zjqa\_\{j\}=z\_\{j\}qfor the common vectorq=ai/zi=aj/zjq=a\_\{i\}/z\_\{i\}=a\_\{j\}/z\_\{j\}would turn the left\-hand side intozizj\(∥q∥22\+1\)z\_\{i\}z\_\{j\}\(\\lVert q\\rVert\_\{2\}^\{2\}\+1\), which is nonzero by \([139](https://arxiv.org/html/2608.26877#A5.E139)\)\. This is a contradiction\.
IfSScontains every member ofJJ, all omitted rows are orthogonal tovv\. Thus
\(AS⊤AS\)v=v\.\(A\_\{S\}^\{\\top\}A\_\{S\}\)v=v\.\(143\)By symmetry,v⊤\(AS⊤AS\)−1=v⊤v^\{\\top\}\(A\_\{S\}^\{\\top\}A\_\{S\}\)^\{\-1\}=v^\{\\top\}, and therefore
v⊤wS\(z\)−θL∗=v⊤AS⊤zS=∑i∈Jcizi=v⊤A⊤z=0\.v^\{\\top\}\\frac\{w\_\{S\}\(z\)\-\\theta\}\{\\sqrt\{L^\{\*\}\}\}=v^\{\\top\}A\_\{S\}^\{\\top\}z\_\{S\}=\\sum\_\{i\\in J\}c\_\{i\}z\_\{i\}=v^\{\\top\}A^\{\\top\}z=0\.\(144\)The two cases above exhaust the positive\-volume support\.
Fori∈Ji\\in J, fixed\-cardinality Cauchy–Binet gives the exact exclusion marginal
ℙA\(i∉S\)\\displaystyle\\mathbb\{P\}\_\{A\}\(i\\notin S\)=\(m−d−1s−d\)det\(A−i⊤A−i\)\(m−ds−d\)\\displaystyle=\\frac\{\\binom\{m\-d\-1\}\{s\-d\}\\det\(A\_\{\-i\}^\{\\top\}A\_\{\-i\}\)\}\{\\binom\{m\-d\}\{s\-d\}\}=m−sm−d\(1−ℓi\)=αszi2\.\\displaystyle=\\frac\{m\-s\}\{m\-d\}\(1\-\\ell\_\{i\}\)=\\alpha\_\{s\}z\_\{i\}^\{2\}\.\(145\)HereA−i⊤A−i=Id−aiai⊤≻0A\_\{\-i\}^\{\\top\}A\_\{\-i\}=I\_\{d\}\-a\_\{i\}a\_\{i\}^\{\\top\}\\succ 0by the no\-coloop condition, and the rank\-one determinant lemma givesdet\(Id−aiai⊤\)=1−ℓi\\det\(I\_\{d\}\-a\_\{i\}a\_\{i\}^\{\\top\}\)=1\-\\ell\_\{i\}\. The omission events for members ofJJare mutually exclusive, as shown above, and the no\-omission branch has zero directional error\. Consequently,
v⊤M¯sA,zvL∗\\displaystyle\\frac\{v^\{\\top\}\\overline\{M\}\_\{s\}^\{A,z\}v\}\{L^\{\*\}\}=∑i∈JℙA\(i∉S\)ci2zi2\\displaystyle=\\sum\_\{i\\in J\}\\mathbb\{P\}\_\{A\}\(i\\notin S\)\\frac\{c\_\{i\}^\{2\}\}\{z\_\{i\}^\{2\}\}=∑i∈Jαszi2ci2zi2=αs∑i∈Jci2=αs\.\\displaystyle=\\sum\_\{i\\in J\}\\alpha\_\{s\}z\_\{i\}^\{2\}\\frac\{c\_\{i\}^\{2\}\}\{z\_\{i\}^\{2\}\}=\\alpha\_\{s\}\\sum\_\{i\\in J\}c\_\{i\}^\{2\}=\\alpha\_\{s\}\.\(146\)The universal spectral envelope gives the reverse inequalityλmax\(M¯sA,z\)/L∗≤αs\\lambda\_\{\\max\}\(\\overline\{M\}\_\{s\}^\{A,z\}\)/L^\{\*\}\\leq\\alpha\_\{s\}\. Hence thiszzattains the envelope\. The witnesszzand the directionvvwere selected from the feature geometry and do not depend onss\. Applying the same calculation at anys′∈\{d\+1,…,m−1\}s^\{\\prime\}\\in\\\{d\+1,\\ldots,m\-1\\\}proves the simultaneous witness statement in \([30](https://arxiv.org/html/2608.26877#S5.E30)\)\.
### E\.4Envelope equality forces zero margin
Conversely, suppose𝖳specRU\(A,s\)=αs\\mathsf\{T\}\_\{\\mathrm\{spec\}\}^\{\\mathrm\{RU\}\}\(A,s\)=\\alpha\_\{s\}\. Compactness gives a maximizingz∈𝒵Az\\in\\mathcal\{Z\}\_\{A\}; choose a unit top eigenvectoruuofM¯sA,z\\overline\{M\}\_\{s\}^\{A,z\}\. Evaluating \([133](https://arxiv.org/html/2608.26877#A5.E133)\) inuugives
αs\\displaystyle\\alpha\_\{s\}=u⊤M¯sA,zuL∗\\displaystyle=\\frac\{u^\{\\top\}\\overline\{M\}\_\{s\}^\{A,z\}u\}\{L^\{\*\}\}\(147\)≤αs−βsγsu⊤RA\(z\)u\.\\displaystyle\\leq\\alpha\_\{s\}\-\\beta\_\{s\}\\gamma\_\{s\}u^\{\\top\}R\_\{A\}\(z\)u\.SinceRA\(z\)⪰0R\_\{A\}\(z\)\\succeq 0andβsγs\>0\\beta\_\{s\}\\gamma\_\{s\}\>0, equality forces
u⊤RA\(z\)u=0\.u^\{\\top\}R\_\{A\}\(z\)u=0\.\(148\)For a positive\-semidefinite matrix this impliesλmin\(RA\(z\)\)=0\\lambda\_\{\\min\}\(R\_\{A\}\(z\)\)=0\. The nonnegativity already proved then givesνA=0\\nu\_\{A\}=0\. This converse does not need the no\-coloop condition; that condition is retained in the theorem because it is used by the forward support argument\.
Combining the strict implication, the zero\-margin attainment argument, and this converse proves both equivalences in Theorem[3](https://arxiv.org/html/2608.26877#Thmtheorem3)\. ∎
### E\.5Scope boundaries
The no\-coloop condition is sufficient, not asserted to be necessary\. It cannot be removed wholesale: for
d=1,m=3,s=2,A=\(1,0,0\)⊤,d=1,\\qquad m=3,\\qquad s=2,\\qquad A=\(1,0,0\)^\{\\top\},\(149\)the margin is zero, but every positive\-volume sample contains the leverage\-one row and the selected coefficient is deterministic, so𝖳specRU\(A,2\)=0<1/2=α2\\mathsf\{T\}\_\{\\mathrm\{spec\}\}^\{\\mathrm\{RU\}\}\(A,2\)=0<1/2=\\alpha\_\{2\}\. The theorem also requiresL∗\>0L^\{\*\}\>0and the strict\-interior regimed<s<md<s<m; its normalized residual and covariance ratio are not evaluated atL∗=0L^\{\*\}=0,s=ds=d, ors=ms=m\. The ordinary sampler, selected unweighted estimator, and centered full\-Gram\-whitened spectral covariance are part of the statement\. No rescaling, replacement, reweighting, ridge term, alternative estimator, or different covariance or loss metric is used\.
## Appendix FCritical geometry: proofs and source boundary
This section proves Theorem[4](https://arxiv.org/html/2608.26877#Thmtheorem4)and the unnumbered sharp critical witness in Section[6](https://arxiv.org/html/2608.26877#S6)\. Throughout,A∈ℝ2d×dA\\in\\mathbb\{R\}^\{2d\\times d\}has orthonormal columns, each row has squared norm1/21/2,P=AA⊤P=AA^\{\\top\}, andP⟂=I2d−PP\_\{\\perp\}=I\_\{2d\}\-P\.
### F\.1Fixed\-complement residual weighting
Forz∈𝒵Az\\in\\mathcal\{Z\}\_\{A\}, one hasP⟂z=zP\_\{\\perp\}z=z\. Therefore
Qz=P⟂−zz⊤Q\_\{z\}=P\_\{\\perp\}\-zz^\{\\top\}\(150\)is an orthogonal projection:Qz2=QzQ\_\{z\}^\{2\}=Q\_\{z\},Qz⪰0Q\_\{z\}\\succeq 0, andrank\(Qz\)=d−1\\operatorname\{rank\}\(Q\_\{z\}\)=d\-1\. Its diagonal is
\(Qz\)ii=\(P⟂\)ii−zi2=1−ℓi−zi2=12−zi2\.\(Q\_\{z\}\)\_\{ii\}=\(P\_\{\\perp\}\)\_\{ii\}\-z\_\{i\}^\{2\}=1\-\\ell\_\{i\}\-z\_\{i\}^\{2\}=\\frac\{1\}\{2\}\-z\_\{i\}^\{2\}\.\(151\)Consequently,
RA\(z\)=∑i\(12−zi2\)aiai⊤=A⊤Diag\(diagQz\)A\.R\_\{A\}\(z\)=\\sum\_\{i\}\\left\(\\frac\{1\}\{2\}\-z\_\{i\}^\{2\}\\right\)a\_\{i\}a\_\{i\}^\{\\top\}=A^\{\\top\}\\operatorname\{Diag\}\(\\operatorname\{diag\}Q\_\{z\}\)A\.\(152\)This proves the fixed\-complement representation used in the main text\. It also giveszi2≤1/2z\_\{i\}^\{2\}\\leq 1/2and henceRA\(z\)⪰0R\_\{A\}\(z\)\\succeq 0\.
The ingredients surrounding \([152](https://arxiv.org/html/2608.26877#A6.E152)\) are established frame theory\. Naimark complementation supplies the complementary projection and Gram geometry\[[7](https://arxiv.org/html/2608.26877#bib.bib7)\]; finite projection diagonals and Schur–Horn frame admissibility describe broader diagonal realization problems\[[21](https://arxiv.org/html/2608.26877#bib.bib21),[1](https://arxiv.org/html/2608.26877#bib.bib1)\]; and weighted\-frame operators treat exogenously specified weights\[[4](https://arxiv.org/html/2608.26877#bib.bib4)\]\. The restriction here is narrower:P⟂P\_\{\\perp\}is fixed,Qz⪯P⟂Q\_\{z\}\\preceq P\_\{\\perp\}has corank one in that fixed range, and the weights are then minimized overz∈ran\(P⟂\)∩S2d−1z\\in\\operatorname\{ran\}\(P\_\{\\perp\}\)\\cap S^\{2d\-1\}\. This attribution distinguishes the residual functional from a claim to Naimark complementation, projection diagonals, or weighted frames\.
### F\.2Zero locus
Suppose first thatνA=0\\nu\_\{A\}=0\. Compactness of𝒵A\\mathcal\{Z\}\_\{A\}and of the unit sphere inℝd\\mathbb\{R\}^\{d\}gives unit vectorsz∈𝒵Az\\in\\mathcal\{Z\}\_\{A\}andv∈ℝdv\\in\\mathbb\{R\}^\{d\}for which
0=v⊤RA\(z\)v=∑i\(12−zi2\)\(ai⊤v\)2\.0=v^\{\\top\}R\_\{A\}\(z\)v=\\sum\_\{i\}\\left\(\\frac\{1\}\{2\}\-z\_\{i\}^\{2\}\\right\)\(a\_\{i\}^\{\\top\}v\)^\{2\}\.\(153\)Every summand is nonnegative\. Hence every row nonorthogonal tovvmust havezi2=1/2z\_\{i\}^\{2\}=1/2\. Parsevalness gives
∑i\(ai⊤v\)2=1,\(ai⊤v\)2≤∥ai∥22=12\.\\sum\_\{i\}\(a\_\{i\}^\{\\top\}v\)^\{2\}=1,\\qquad\(a\_\{i\}^\{\\top\}v\)^\{2\}\\leq\\lVert a\_\{i\}\\rVert\_\{2\}^\{2\}=\\frac\{1\}\{2\}\.\(154\)At least two rows are therefore nonorthogonal tovv\. Since∑izi2=1\\sum\_\{i\}z\_\{i\}^\{2\}=1, exactly two coordinates, sayiiandjj, can be saturated\. The two nonzero terms in \([154](https://arxiv.org/html/2608.26877#A6.E154)\) must both equal1/21/2\. Equality in Cauchy–Schwarz yields signsσi,σj\\sigma\_\{i\},\\sigma\_\{j\}such that
ai=σi2v,aj=σj2v,ak⊤v=0\(k∉\{i,j\}\)\.a\_\{i\}=\\frac\{\\sigma\_\{i\}\}\{\\sqrt\{2\}\}v,\\qquad a\_\{j\}=\\frac\{\\sigma\_\{j\}\}\{\\sqrt\{2\}\}v,\\qquad a\_\{k\}^\{\\top\}v=0\\quad\(k\\notin\\\{i,j\\\}\)\.\(155\)
Conversely, assume \([155](https://arxiv.org/html/2608.26877#A6.E155)\) and define
zi=σi2,zj=−σj2,zk=0\(k∉\{i,j\}\)\.z\_\{i\}=\\frac\{\\sigma\_\{i\}\}\{\\sqrt\{2\}\},\\qquad z\_\{j\}=\-\\frac\{\\sigma\_\{j\}\}\{\\sqrt\{2\}\},\\qquad z\_\{k\}=0\\quad\(k\\notin\\\{i,j\\\}\)\.\(156\)ThenA⊤z=0A^\{\\top\}z=0,∥z∥2=1\\lVert z\\rVert\_\{2\}=1, andv⊤RA\(z\)v=0v^\{\\top\}R\_\{A\}\(z\)v=0\. Since everyRA\(z\)R\_\{A\}\(z\)is positive semidefinite, this proves the equivalence in Theorem[4](https://arxiv.org/html/2608.26877#Thmtheorem4)\.
This zero boundary is not claimed as a newly identified erasure geometry\. For a uniform Parseval frame, the exact binary two\-erasure formula of[Bodmann and Paulsen \[5\]](https://arxiv.org/html/2608.26877#bib.bib5)gives, in the present normalization, binary error\(1\+ρA\)/2\(1\+\\rho\_\{A\}\)/2and retained binary lower bound\(1−ρA\)/2\(1\-\\rho\_\{A\}\)/2\. It therefore has the same projective\-duplicate singular boundary\. Binary masking is a different functional from \([152](https://arxiv.org/html/2608.26877#A6.E152)\): the latter uses the coupled fractional weights1/2−zi21/2\-z\_\{i\}^\{2\}generated by one direction inside a fixed complement\.
### F\.3Coherence sandwich
Putui=2aiu\_\{i\}=\\sqrt\{2\}a\_\{i\}\. Then∥ui∥2=1\\lVert u\_\{i\}\\rVert\_\{2\}=1,∑iuiui⊤=2Id\\sum\_\{i\}u\_\{i\}u\_\{i\}^\{\\top\}=2I\_\{d\}, andρA=maxi≠j\|ui⊤uj\|\\rho\_\{A\}=\\max\_\{i\\neq j\}\|u\_\{i\}^\{\\top\}u\_\{j\}\|\. For unitz∈𝒵Az\\in\\mathcal\{Z\}\_\{A\}and unitv∈ℝdv\\in\\mathbb\{R\}^\{d\}, set
pi=zi2,xi=\(ui⊤v\)2\.p\_\{i\}=z\_\{i\}^\{2\},\\qquad x\_\{i\}=\(u\_\{i\}^\{\\top\}v\)^\{2\}\.\(157\)Then0≤pi≤1/20\\leq p\_\{i\}\\leq 1/2,∑ipi=1\\sum\_\{i\}p\_\{i\}=1,∑ixi=2\\sum\_\{i\}x\_\{i\}=2, and
v⊤RA\(z\)v=12−12∑ipixi\.v^\{\\top\}R\_\{A\}\(z\)v=\\frac\{1\}\{2\}\-\\frac\{1\}\{2\}\\sum\_\{i\}p\_\{i\}x\_\{i\}\.\(158\)The maximum of∑ipixi\\sum\_\{i\}p\_\{i\}x\_\{i\}over the capped simplex0≤pi≤1/20\\leq p\_\{i\}\\leq 1/2,∑ipi=1\\sum\_\{i\}p\_\{i\}=1, places mass1/21/2on the two largestxix\_\{i\}\. For any pair,
xi\+xj≤λmax\(uiui⊤\+ujuj⊤\)=1\+\|ui⊤uj\|≤1\+ρA\.x\_\{i\}\+x\_\{j\}\\leq\\lambda\_\{\\max\}\(u\_\{i\}u\_\{i\}^\{\\top\}\+u\_\{j\}u\_\{j\}^\{\\top\}\)=1\+\|u\_\{i\}^\{\\top\}u\_\{j\}\|\\leq 1\+\\rho\_\{A\}\.\(159\)Thus∑ipixi≤\(1\+ρA\)/2\\sum\_\{i\}p\_\{i\}x\_\{i\}\\leq\(1\+\\rho\_\{A\}\)/2, and \([158](https://arxiv.org/html/2608.26877#A6.E158)\) gives
νA≥1−ρA4\.\\nu\_\{A\}\\geq\\frac\{1\-\\rho\_\{A\}\}\{4\}\.\(160\)
For the other direction, choosei,ji,jattainingρA\\rho\_\{A\}and chooseϵ∈\{±1\}\\epsilon\\in\\\{\\pm 1\\\}so thatui⊤\(ϵuj\)=ρAu\_\{i\}^\{\\top\}\(\\epsilon u\_\{j\}\)=\\rho\_\{A\}\. Define
b=ei−ϵej2,h=b⊤P⟂b=1\+ρA2,z=P⟂bh\.b=\\frac\{e\_\{i\}\-\\epsilon e\_\{j\}\}\{\\sqrt\{2\}\},\\qquad h=b^\{\\top\}P\_\{\\perp\}b=\\frac\{1\+\\rho\_\{A\}\}\{2\},\\qquad z=\\frac\{P\_\{\\perp\}b\}\{\\sqrt\{h\}\}\.\(161\)The vectorzzis a unit residual, and a direct coordinate calculation gives
zi2=zj2=1\+ρA4\.z\_\{i\}^\{2\}=z\_\{j\}^\{2\}=\\frac\{1\+\\rho\_\{A\}\}\{4\}\.\(162\)For
v=ui\+ϵuj2\(1\+ρA\),v=\\frac\{u\_\{i\}\+\\epsilon u\_\{j\}\}\{\\sqrt\{2\(1\+\\rho\_\{A\}\)\}\},\(163\)one hasxi=xj=\(1\+ρA\)/2x\_\{i\}=x\_\{j\}=\(1\+\\rho\_\{A\}\)/2\. Keeping only these two nonnegative terms in \([158](https://arxiv.org/html/2608.26877#A6.E158)\) yields
νA≤12−\(1\+ρA\)28=\(1−ρA\)\(3\+ρA\)8\.\\nu\_\{A\}\\leq\\frac\{1\}\{2\}\-\\frac\{\(1\+\\rho\_\{A\}\)^\{2\}\}\{8\}=\\frac\{\(1\-\\rho\_\{A\}\)\(3\+\\rho\_\{A\}\)\}\{8\}\.\(164\)Together, \([160](https://arxiv.org/html/2608.26877#A6.E160)\) and \([164](https://arxiv.org/html/2608.26877#A6.E164)\) prove \([36](https://arxiv.org/html/2608.26877#S6.E36)\)\.
### F\.4Ordinary\-volume sharp witness
Assume the zero geometry \([155](https://arxiv.org/html/2608.26877#A6.E155)\), fixθ∈ℝd\\theta\\in\\mathbb\{R\}^\{d\}andL∗\>0L^\{\*\}\>0, and use the single residualz∗z\_\{\*\}in \([156](https://arxiv.org/html/2608.26877#A6.E156)\), fixed before sampling and independently of the budget\. Fixs=d\+rs=d\+rwith1≤r≤d−11\\leq r\\leq d\-1\. Choose an orthonormal basis\[v,V\]\[v,V\]ofℝd\\mathbb\{R\}^\{d\}, and fork∉\{i,j\}k\\notin\\\{i,j\\\}letck=V⊤akc\_\{k\}=V^\{\\top\}a\_\{k\}\. The matrixCCwith these2d−22d\-2rows satisfies
C⊤C=Id−1\.C^\{\\top\}C=I\_\{d\-1\}\.\(165\)A positive\-volumess\-set contains exactly one or both ofi,ji,j\. IfTTcollects its remaining rows, the two determinant branches are
det\(AS⊤AS\)=12det\(CT⊤CT\)\(exactly one\),det\(AS⊤AS\)=det\(CT⊤CT\)\(both\)\.\\det\(A\_\{S\}^\{\\top\}A\_\{S\}\)=\\frac\{1\}\{2\}\\det\(C\_\{T\}^\{\\top\}C\_\{T\}\)\\quad\\text\{\(exactly one\)\},\\qquad\\det\(A\_\{S\}^\{\\top\}A\_\{S\}\)=\\det\(C\_\{T\}^\{\\top\}C\_\{T\}\)\\quad\\text\{\(both\)\}\.\(166\)The fixed\-cardinality Cauchy–Binet identity gives their total weights
Wone=\(d−1r\),Wboth=\(d−1r−1\),Wone\+Wboth=\(dr\)\.W\_\{\\mathrm\{one\}\}=\\binom\{d\-1\}\{r\},\\qquad W\_\{\\mathrm\{both\}\}=\\binom\{d\-1\}\{r\-1\},\\qquad W\_\{\\mathrm\{one\}\}\+W\_\{\\mathrm\{both\}\}=\\binom\{d\}\{r\}\.\(167\)Thus
ℙA\(exactly one ofi,jis selected\)=\(d−1r\)\(dr\)=d−rd=2d−sd\.\\mathbb\{P\}\_\{A\}\(\\text\{exactly one of $i,j$ is selected\}\)=\\frac\{\\binom\{d\-1\}\{r\}\}\{\\binom\{d\}\{r\}\}=\\frac\{d\-r\}\{d\}=\\frac\{2d\-s\}\{d\}\.\(168\)
Foryz∗=Aθ\+L∗z∗y\_\{z\_\{\*\}\}=A\\theta\+\\sqrt\{L^\{\*\}\}z\_\{\*\}, the selected coefficient error on the one\-member branch is±L∗v\\pm\\sqrt\{L^\{\*\}\}v; the two signs have equal determinant weight and therefore cancel in the mean\. On the both\-member branch, the two residual contributions cancel and the coefficient error is zero\. Direct summation over the ordinary\-volume support gives
M¯sA,z∗=d−rdL∗vv⊤=2d−sdL∗vv⊤,λmax\(M¯sA,z∗\)L∗=2d−sd\.\\overline\{M\}\_\{s\}^\{A,z\_\{\*\}\}=\\frac\{d\-r\}\{d\}L^\{\*\}vv^\{\\top\}=\\frac\{2d\-s\}\{d\}L^\{\*\}vv^\{\\top\},\\qquad\\frac\{\\lambda\_\{\\max\}\(\\overline\{M\}\_\{s\}^\{A,z\_\{\*\}\}\)\}\{L^\{\*\}\}=\\frac\{2d\-s\}\{d\}\.\(169\)The samez∗z\_\{\*\}works for everyr∈\{1,…,d−1\}r\\in\\\{1,\\ldots,d\-1\\\}, proving the unnumbered sharp critical witness in Section[6](https://arxiv.org/html/2608.26877#S6)\. This direct support enumeration is valid even when the repeated\-pair design is outside row general position; no exact interior inverse\-moment identity is extended to that boundary\.
The scope is deliberately narrow\. The response is fixed before the random subset is drawn; the law is ordinary unrescaled indexed fixed\-size volume sampling; the estimator is selected unweighted least squares on positive\-volume sets;d<s<2dd<s<2dandL∗\>0L^\{\*\}\>0; and the statistic is the centered, full\-Gram\-whitened spectral coefficient covariance\. Equation \([169](https://arxiv.org/html/2608.26877#A6.E169)\) is rank one, so it does not saturate the trace\-derived scalar\-loss envelope\.
[Dereziński and Warmuth \[12\]](https://arxiv.org/html/2608.26877#bib.bib12)establish ordinary\-volume sampling identities, including fixed\-response control at rank size and all\-size unbiasedness and inverse\-Gram moments\. The scalar\-loss obstruction of[Dereziński et al\. \[13\]](https://arxiv.org/html/2608.26877#bib.bib13)uses the same ordinary law and selected OLS and has the same numerical excess factor when its dimensions are specialized, but its theorem takes a highly nonuniform\-leverage limit in a two\-block family\. At equal block scale that family becomes the duplicated\-pair geometry after whitening, but the published obstruction is a scalar expected loss statement, not the fixed one\-coordinate residual and centered rank\-one covariance equality in \([169](https://arxiv.org/html/2608.26877#A6.E169)\)\. The witness here is therefore presented only as a critical\-class corollary of the phase geometry, not as a separate volume\-sampling obstruction or a claim to the factor or duplicated\-pair construction\.
## Appendix GSound pre\-response certificates: proofs and boundaries
This section proves Corollaries[5](https://arxiv.org/html/2608.26877#Thmtheorem5)and[6](https://arxiv.org/html/2608.26877#Thmtheorem6)\. It separates the exact phase variableνA\\nu\_\{A\}from several sufficient lower certificates\. None of the arguments changes the ordinary indexed fixed\-size volume law or the selected unweighted OLS estimator\.
### G\.1The general spectral certificate
LetN∈ℝm×\(m−d\)N\\in\\mathbb\{R\}^\{m\\times\(m\-d\)\}have orthonormal columns spanningker\(A⊤\)\\ker\(A^\{\\top\}\), letDℓ=diag\(ℓ\)D\_\{\\ell\}=\\operatorname\{diag\}\(\\ell\), and setP⟂=Im−AA⊤=NN⊤P\_\{\\perp\}=I\_\{m\}\-AA^\{\\top\}=NN^\{\\top\}\. Define
R0=∑i=1m\(1−ℓi\)aiai⊤,τX=λmax\(N⊤DℓN\),cX=\[λmin\(R0\)−τX\]\+\.R\_\{0\}=\\sum\_\{i=1\}^\{m\}\(1\-\\ell\_\{i\}\)a\_\{i\}a\_\{i\}^\{\\top\},\\quad\\tau\_\{X\}=\\lambda\_\{\\max\}\\\!\\left\(N^\{\\top\}D\_\{\\ell\}N\\right\),\\quad c\_\{X\}=\\bigl\[\\lambda\_\{\\min\}\(R\_\{0\}\)\-\\tau\_\{X\}\\bigr\]\_\{\+\}\.\(170\)For a unitz∈ker\(A⊤\)z\\in\\ker\(A^\{\\top\}\), write
T\(z\)=∑izi2aiai⊤,RA\(z\)=R0−T\(z\)\.T\(z\)=\\sum\_\{i\}z\_\{i\}^\{2\}a\_\{i\}a\_\{i\}^\{\\top\},\\qquad R\_\{A\}\(z\)=R\_\{0\}\-T\(z\)\.\(171\)Because\[Az\]\[A\\ z\]has orthonormal columns, its row leverages satisfyℓi\+zi2≤1\\ell\_\{i\}\+z\_\{i\}^\{2\}\\leq 1, and hence
RA\(z\)=∑i\(1−ℓi−zi2\)aiai⊤⪰0\.R\_\{A\}\(z\)=\\sum\_\{i\}\(1\-\\ell\_\{i\}\-z\_\{i\}^\{2\}\)a\_\{i\}a\_\{i\}^\{\\top\}\\succeq 0\.\(172\)For everyv∈ℝdv\\in\\mathbb\{R\}^\{d\},
v⊤T\(z\)v\\displaystyle v^\{\\top\}T\(z\)v=∑izi2\(ai⊤v\)2≤∥v∥22∑iℓizi2\\displaystyle=\\sum\_\{i\}z\_\{i\}^\{2\}\(a\_\{i\}^\{\\top\}v\)^\{2\}\\leq\\lVert v\\rVert\_\{2\}^\{2\}\\sum\_\{i\}\\ell\_\{i\}z\_\{i\}^\{2\}=∥v∥22z⊤P⟂DℓP⟂z≤τX∥v∥22\.\\displaystyle=\\lVert v\\rVert\_\{2\}^\{2\}z^\{\\top\}P\_\{\\perp\}D\_\{\\ell\}P\_\{\\perp\}z\\leq\\tau\_\{X\}\\lVert v\\rVert\_\{2\}^\{2\}\.\(173\)ThusT\(z\)⪯τXIdT\(z\)\\preceq\\tau\_\{X\}I\_\{d\}and
RA\(z\)⪰\[λmin\(R0\)−τX\]Id\.R\_\{A\}\(z\)\\succeq\[\\lambda\_\{\\min\}\(R\_\{0\}\)\-\\tau\_\{X\}\]I\_\{d\}\.\(174\)Combining this with \([172](https://arxiv.org/html/2608.26877#A7.E172)\) givesRA\(z\)⪰cXIdR\_\{A\}\(z\)\\succeq c\_\{X\}I\_\{d\}for every admissiblezz\. Minimizing overzzproves
0≤cX≤νA\.0\\leq c\_\{X\}\\leq\\nu\_\{A\}\.\(175\)
The computation can use the residual\-space compression
τX=λmax\(N⊤DℓN\),\\tau\_\{X\}=\\lambda\_\{\\max\}\(N^\{\\top\}D\_\{\\ell\}N\),\(176\)without forming a densem×mm\\times mprojector\. The value is invariant under an invertible feature\-coordinate changeX↦XHX\\mapsto XH: the corresponding whitened matrices differ by a right orthogonal factor, soR0R\_\{0\}changes by orthogonal congruence whileP⟂,Dℓ,τXP\_\{\\perp\},D\_\{\\ell\},\\tau\_\{X\}, andcXc\_\{X\}are unchanged\.
### G\.2Capped, top\-K, coherence, and fixed\-threshold routes
The following hierarchy records the other sufficient routes used in the main text\. Setci=1−ℓic\_\{i\}=1\-\\ell\_\{i\}and
𝒫A=\{p∈ℝm:0≤pi≤ci,∑ipi=1\},ν¯cap\(A\)=minp∈𝒫Aλmin\(A⊤diag\(c−p\)A\)\.\\mathcal\{P\}\_\{A\}=\\left\\\{p\\in\\mathbb\{R\}^\{m\}:0\\leq p\_\{i\}\\leq c\_\{i\},\\ \\sum\_\{i\}p\_\{i\}=1\\right\\\},\\qquad\\underline\{\\nu\}\_\{\\rm cap\}\(A\)=\\min\_\{p\\in\\mathcal\{P\}\_\{A\}\}\\lambda\_\{\\min\}\\\!\\left\(A^\{\\top\}\\operatorname\{diag\}\(c\-p\)A\\right\)\.\(177\)For every residual direction,p=z⊙2p=z^\{\\odot 2\}belongs to𝒫A\\mathcal\{P\}\_\{A\}: the augmented projection gives0≤zi2≤1−ℓi0\\leq z\_\{i\}^\{2\}\\leq 1\-\\ell\_\{i\}and unit norm gives∑izi2=1\\sum\_\{i\}z\_\{i\}^\{2\}=1\. Enlarging the realizable residual\-square set therefore gives
0≤ν¯cap\(A\)≤νA\.0\\leq\\underline\{\\nu\}\_\{\\rm cap\}\(A\)\\leq\\nu\_\{A\}\.\(178\)
Letc⋆=1−ℓ−c\_\{\\star\}=1\-\\ell\_\{\-\},K=⌈1/c⋆⌉K=\\lceil 1/c\_\{\\star\}\\rceil, and
ΛA,K=max\|J\|=Kλmax\(∑i∈Jaiai⊤\)\.\\Lambda\_\{A,K\}=\\max\_\{\|J\|=K\}\\lambda\_\{\\max\}\\\!\\left\(\\sum\_\{i\\in J\}a\_\{i\}a\_\{i\}^\{\\top\}\\right\)\.\(179\)For a unitvv, uniform\-cap fractional knapsack bounds∑ipi\(ai⊤v\)2\\sum\_\{i\}p\_\{i\}\(a\_\{i\}^\{\\top\}v\)^\{2\}byc⋆c\_\{\\star\}times the sum of theKKlargest directional row energies\. Consequently
ttop\(A\):=\[λmin\(R0\)−c⋆ΛA,K\]\+≤ν¯cap\(A\)\.t\_\{\\rm top\}\(A\):=\[\\lambda\_\{\\min\}\(R\_\{0\}\)\-c\_\{\\star\}\\Lambda\_\{A,K\}\]\_\{\+\}\\leq\\underline\{\\nu\}\_\{\\rm cap\}\(A\)\.\(180\)If every row is nonzero, define the normalized\-row coherence
ρArow=maxi≠j\|ai⊤aj\|ℓiℓj\.\\rho\_\{A\}^\{\\rm row\}=\\max\_\{i\\neq j\}\\frac\{\|a\_\{i\}^\{\\top\}a\_\{j\}\|\}\{\\sqrt\{\\ell\_\{i\}\\ell\_\{j\}\}\}\.\(181\)Gershgorin’s theorem applied to each normalized\-rowKK\-Gram matrix gives
ΛA,K≤ℓ\+\{1\+\(K−1\)ρArow\}\.\\Lambda\_\{A,K\}\\leq\\ell\_\{\+\}\\\{1\+\(K\-1\)\\rho\_\{A\}^\{\\rm row\}\\\}\.\(182\)It follows that
tcoh\(A\):=\[λmin\(R0\)−c⋆ℓ\+\{1\+\(K−1\)ρArow\}\]\+≤ν¯cap\(A\)≤νA\.t\_\{\\rm coh\}\(A\):=\[\\lambda\_\{\\min\}\(R\_\{0\}\)\-c\_\{\\star\}\\ell\_\{\+\}\\\{1\+\(K\-1\)\\rho\_\{A\}^\{\\rm row\}\\\}\]\_\{\+\}\\leq\\underline\{\\nu\}\_\{\\rm cap\}\(A\)\\leq\\nu\_\{A\}\.\(183\)Atm=2dm=2dwith equal leverage, the sharper specialized statementtρ=\(1−ρA\)/4≤νAt\_\{\\rho\}=\(1\-\\rho\_\{A\}\)/4\\leq\\nu\_\{A\}is precisely the lower side of Theorem[4](https://arxiv.org/html/2608.26877#Thmtheorem4); itsρA\\rho\_\{A\}is the critical normalized coherence in \([35](https://arxiv.org/html/2608.26877#S6.E35)\)\.
tρ\(A\)=1−ρA4\.t\_\{\\rho\}\(A\)=\\frac\{1\-\\rho\_\{A\}\}\{4\}\.\(184\)
For the fixed\-threshold predicate, congruence ofΦ13\(X\)⪰0\\Phi\_\{13\}\(X\)\\succeq 0byG−1/2G^\{\-1/2\}gives
λmin\(R0\)≥1360\+c⋆ℓ\+K\.\\lambda\_\{\\min\}\(R\_\{0\}\)\\geq\\frac\{13\}\{60\}\+c\_\{\\star\}\\ell\_\{\+\}K\.\(185\)The strict lower\-leverage guard makesρArow\\rho\_\{A\}^\{\\rm row\}well defined, and Cauchy–Schwarz givesρArow≤1\\rho\_\{A\}^\{\\rm row\}\\leq 1\. Hence
λmin\(R0\)−c⋆ℓ\+\{1\+\(K−1\)ρArow\}\\displaystyle\\lambda\_\{\\min\}\(R\_\{0\}\)\-c\_\{\\star\}\\ell\_\{\+\}\\\{1\+\(K\-1\)\\rho\_\{A\}^\{\\rm row\}\\\}≥1360\+c⋆ℓ\+\(K−1\)\(1−ρArow\)≥1360\.\\displaystyle\\qquad\\geq\\frac\{13\}\{60\}\+c\_\{\\star\}\\ell\_\{\+\}\(K\-1\)\(1\-\\rho\_\{A\}^\{\\rm row\}\)\\geq\\frac\{13\}\{60\}\.\(186\)Together with \([183](https://arxiv.org/html/2608.26877#A7.E183)\), this proves that a completed fixed\-threshold pass certifies13/60≤νA13/60\\leq\\nu\_\{A\}\. A zero row cannot be silently admitted: normalized\-row coherence would be undefined even when the matrix predicate happens to be positive\. A leverage\-one row is likewise outside the no\-coloop route\. These are eligibility failures, not evidence thatνA=0\\nu\_\{A\}=0\.
### G\.3Propagation to robust slack
Lett\(A\)t\(A\)be any proved lower certificate in \([38](https://arxiv.org/html/2608.26877#S7.E38)\)\. On the domain of Theorem[3](https://arxiv.org/html/2608.26877#Thmtheorem3), its quantitative statement gives
αs−𝖳specRU\(A,s\)≥βsγsνA≥βsγst\(A\),\\alpha\_\{s\}\-\\mathsf\{T\}\_\{\\mathrm\{spec\}\}^\{\\mathrm\{RU\}\}\(A,s\)\\geq\\beta\_\{s\}\\gamma\_\{s\}\\nu\_\{A\}\\geq\\beta\_\{s\}\\gamma\_\{s\}t\(A\),\(187\)which proves \([39](https://arxiv.org/html/2608.26877#S7.E39)\)\. Only the first inequality is tied to the exact phase theorem\. ReplacingνA\\nu\_\{A\}by a lower certificate is a sufficient implication and supplies no converse\.
For a fixed response withL∗\>0L^\{\*\}\>0, putz=e/L∗z=e/\\sqrt\{L^\{\*\}\},n=m−dn=m\-d, andq=m−sq=m\-s\. The boundary\-valid one\-sided residual resolvent gives
M¯sL∗⪯Id−n−qn\(Id−qn−1RA\(z\)\)−1\.\\frac\{\\overline\{M\}\_\{s\}\}\{L^\{\*\}\}\\preceq I\_\{d\}\-\\frac\{n\-q\}\{n\}\\left\(I\_\{d\}\-\\frac\{q\}\{n\-1\}R\_\{A\}\(z\)\\right\)^\{\-1\}\.\(188\)IfRA\(z\)⪰tIdR\_\{A\}\(z\)\\succeq tI\_\{d\}, inversion reverses Loewner order and yields
M¯sL∗\\displaystyle\\frac\{\\overline\{M\}\_\{s\}\}\{L^\{\*\}\}⪯\(1−\(n−q\)/n1−qt/\(n−1\)\)Id\\displaystyle\\preceq\\left\(1\-\\frac\{\(n\-q\)/n\}\{1\-qt/\(n\-1\)\}\\right\)I\_\{d\}=q\(n−1−nt\)n\(n−1−qt\)Id=bn,q\(t\)Id\.\\displaystyle=\\frac\{q\(n\-1\-nt\)\}\{n\(n\-1\-qt\)\}I\_\{d\}=b\_\{n,q\}\(t\)I\_\{d\}\.\(189\)This proves \([G\.3](https://arxiv.org/html/2608.26877#A7.EGx10)\)\. It uses only the one\-sided resolvent at arbitrary full\-rank support boundaries, not the exact row\-general\-position inverse\-moment covariance identity\. The expectation inM¯s\\overline\{M\}\_\{s\}is conditional on one fixed indexed pool and one fixed response; only the ordinary subset draw is random\.
### G\.4Fixed\-profile cardinality arithmetic
Att=13/60t=13/60and target toleranceξ=1/3\\xi=1/3, the denominator in \([189](https://arxiv.org/html/2608.26877#A7.E189)\) is positive for1≤q≤n−11\\leq q\\leq n\-1, because
n−1−1360q≥4760\(n−1\)\>0\.n\-1\-\\frac\{13\}\{60\}q\\geq\\frac\{47\}\{60\}\(n\-1\)\>0\.\(190\)Exact cross multiplication gives
bn,q\(13/60\)≤13⟺q\(77n−90\)≤30n\(n−1\)\.b\_\{n,q\}\(13/60\)\\leq\\frac\{1\}\{3\}\\quad\\Longleftrightarrow\\quad q\(77n\-90\)\\leq 30n\(n\-1\)\.\(191\)Thus the largest integer admitted by this particular scalar certificate is
q13\(n\)=⌊30n\(n−1\)77n−90⌋\.q\_\{13\}\(n\)=\\left\\lfloor\\frac\{30n\(n\-1\)\}\{77n\-90\}\\right\\rfloor\.\(192\)The word “largest” here concerns only inequality \([191](https://arxiv.org/html/2608.26877#A7.E191)\); it is not a minimal\-data or optimal\-cardinality claim\. The universal\-envelope certificate is thet=0t=0case, for whichbn,q\(0\)=q/nb\_\{n,q\}\(0\)=q/nand hence
qH\(n\)=⌊n3⌋\.q\_\{H\}\(n\)=\\left\\lfloor\\frac\{n\}\{3\}\\right\\rfloor\.\(193\)For everyn≥2n\\geq 2,
30n\(n−1\)77n−90≥n3,30n\(n−1\)77n−90<n−1,\\frac\{30n\(n\-1\)\}\{77n\-90\}\\geq\\frac\{n\}\{3\},\\qquad\\frac\{30n\(n\-1\)\}\{77n\-90\}<n\-1,\(194\)soq13≥qHq\_\{13\}\\geq q\_\{H\}\. Atn=2n=2, both counts are zero and route to the full\-pool endpoint without evaluating the strict\-interior formula\. Forn≥3n\\geq 3,1≤qH≤q13≤n−21\\leq q\_\{H\}\\leq q\_\{13\}\\leq n\-2\.
For the frozen paper profile,n=512−64=448n=512\-64=448, and exact integer evaluation gives
q13=⌊300384017203⌋=174,qH=⌊4483⌋=149\.q\_\{13\}=\\left\\lfloor\\frac\{3003840\}\{17203\}\\right\\rfloor=174,\\qquad q\_\{H\}=\\left\\lfloor\\frac\{448\}\{3\}\\right\\rfloor=149\.\(195\)Therefores13=512−174=338s\_\{13\}=512\-174=338andsH=512−149=363s\_\{H\}=512\-149=363\. Substitution in \([189](https://arxiv.org/html/2608.26877#A7.E189)\) proves Corollary[6](https://arxiv.org/html/2608.26877#Thmtheorem6)\. The fixed\-threshold branch and fallback change onlyss; neither changes subset weights at thatss, the selected fit, or the covariance target\.
### G\.5Structural calibration and abstention boundaries
The inexpensivecXc\_\{X\}route is intentionally conservative\. Since
R0=Id−A⊤DℓA,tr\(P⟂DℓP⟂\)=∑iℓi\(1−ℓi\),R\_\{0\}=I\_\{d\}\-A^\{\\top\}D\_\{\\ell\}A,\\qquad\\operatorname\{tr\}\(P\_\{\\perp\}D\_\{\\ell\}P\_\{\\perp\}\)=\\sum\_\{i\}\\ell\_\{i\}\(1\-\\ell\_\{i\}\),\(196\)we have
λmin\(R0\)≤∑iℓi\(1−ℓi\)d,τX≥∑iℓi\(1−ℓi\)m−d\.\\lambda\_\{\\min\}\(R\_\{0\}\)\\leq\\frac\{\\sum\_\{i\}\\ell\_\{i\}\(1\-\\ell\_\{i\}\)\}\{d\},\\qquad\\tau\_\{X\}\\geq\\frac\{\\sum\_\{i\}\\ell\_\{i\}\(1\-\\ell\_\{i\}\)\}\{m\-d\}\.\(197\)HencecX=0c\_\{X\}=0wheneverd<m≤2dd<m\\leq 2d\. For an equal\-leverage Parseval design,
cX=\[1−2dm\]\+,c\_\{X\}=\\left\[1\-\\frac\{2d\}\{m\}\\right\]\_\{\+\},\(198\)which is positive whenm\>2dm\>2d\. That redundancy condition alone is not sufficient: the full\-rank Parseval core\-plus\-zero designA=\[Id;0\(m−d\)×d\]A=\[I\_\{d\};0\_\{\(m\-d\)\\times d\}\]hascX=0c\_\{X\}=0for arbitrarily largemm\. At the critical boundarym=2dm=2d,cXc\_\{X\}is therefore silent even though the coherence certificate in \([184](https://arxiv.org/html/2608.26877#A7.E184)\) can be positive\.
These zeros demonstrate abstention, not impossibility\. In particular,cX=0c\_\{X\}=0,ttop=0t\_\{\\rm top\}=0,tcoh=0t\_\{\\rm coh\}=0, a failed fixed\-threshold predicate, or an ineligible guard does not showνA=0\\nu\_\{A\}=0and does not invoke the tight branch of Theorem[3](https://arxiv.org/html/2608.26877#Thmtheorem3)\. The positive\-loss and strict\-interior formulas are not evaluated atL∗=0L^\{\*\}=0,s=ds=d,s=ms=m, orm=dm=d; the endpoint statements in the universal\-envelope and residual\-mechanism sections remain controlling\.
## Appendix HSupplementary balanced result
The following retained result is supplementary and is not a premise of the main robust\-phase or critical\-geometry claims\.
### H\.1Balanced ordinary geometry and the worst direction
VanishingcXc\_\{X\}limits one scalar relaxation, not the feature geometry\. Returning directly to Theorem[2](https://arxiv.org/html/2608.26877#Thmtheorem2), its response\-aware linear consequence improves the universal ceiling under ordinary geometric conditions\. Setm=2dm=2dands=d\+rs=d\+r\. The balanced statement works directly in whitened coordinates:A=XG−1/2A=XG^\{\-1/2\}, soA⊤A=IdA^\{\\top\}A=I\_\{d\}\. The quantityΔ\(A,y,r\)\\Delta\(A,y;r\)is the relative shortfall of the worst\-direction covariance from the universal ceiling\(1−r/d\)L∗\(A,y\)\(1\-r/d\)L^\{\*\}\(A,y\)\. Atm=2dm=2d, the response\-uniform feature certificate hascX=0c\_\{X\}=0, while the response\-aware route here can still certify a gap; the balanced theorem is independent of that certificate\. Ford≥2d\\geq 2, define
𝒞d\(σ0\)=\{A∈ℝ2d×d:A⊤A=Id,∥ai∥22=12∀i,mini≠jsin∠proj\(ai,aj\)≥σ0\},σ0=398\.\\mathcal\{C\}\_\{d\}\(\\sigma\_\{0\}\)=\\left\\\{\\begin\{array\}\[\]\{l\}A\\in\\mathbb\{R\}^\{2d\\times d\}:A^\{\\top\}A=I\_\{d\},\\quad\\lVert a\_\{i\}\\rVert\_\{2\}^\{2\}=\\tfrac\{1\}\{2\}\\ \\forall i,\\\\\[\-1\.0pt\] \\min\_\{i\\neq j\}\\sin\\angle\_\{\\mathrm\{proj\}\}\(a\_\{i\},a\_\{j\}\)\\geq\\sigma\_\{0\}\\end\{array\}\\right\\\},\\qquad\\sigma\_\{0\}=\\frac\{\\sqrt\{39\}\}\{8\}\.\(199\)where, for nonzero vectorsu,vu,v,∠proj\(u,v\)\\angle\_\{\\mathrm\{proj\}\}\(u,v\)is the angle between their one\-dimensional spans, and hence
sin∠proj\(u,v\)=1−\(u⊤v\)2/\(∥u∥22∥v∥22\)\.\\sin\\angle\_\{\\mathrm\{proj\}\}\(u,v\)=\\sqrt\{1\-\(u^\{\\top\}v\)^\{2\}/\(\\lVert u\\rVert\_\{2\}^\{2\}\\lVert v\\rVert\_\{2\}^\{2\}\)\}\.
##### Small\-dimension nonvacuity\.
Forσ0=39/8\\sigma\_\{0\}=\\sqrt\{39\}/8, the class𝒞2\(σ0\)\\mathcal\{C\}\_\{2\}\(\\sigma\_\{0\}\)is empty: four projective lines inℝ2\\mathbb\{R\}^\{2\}contain a pair at projective angle at mostπ/4\\pi/4, whereasσ0\>sin\(π/4\)\\sigma\_\{0\}\>\\sin\(\\pi/4\)\. Thus the universald=2d=2clause below is formally valid but vacuous\. The explicit witness below separately proves nonemptiness ford=4nd=4^\{n\},n≥1n\\geq 1; no claim is made here about the complete nonemptiness range in other dimensions\.
ForA∈𝒞d\(σ0\)A\\in\\mathcal\{C\}\_\{d\}\(\\sigma\_\{0\}\)and fixedyywithL∗\(A,y\)\>0L^\{\*\}\(A,y\)\>0, define
Δ\(A,y,r\)=1−λmax\(M¯d\+r\)\(1−r/d\)L∗\(A,y\),Δd,r∗\(σ0\)=infA∈𝒞d\(σ0\)infy∈ℝ2dL∗\(A,y\)\>0Δ\(A,y,r\)\.\\Delta\(A,y;r\)=1\-\\frac\{\\lambda\_\{\\max\}\(\\overline\{M\}\_\{d\+r\}\)\}\{\(1\-r/d\)L^\{\*\}\(A,y\)\},\\qquad\\Delta^\{\*\}\_\{d,r\}\(\\sigma\_\{0\}\)=\\inf\_\{A\\in\\mathcal\{C\}\_\{d\}\(\\sigma\_\{0\}\)\}\\inf\_\{\\begin\{subarray\}\{c\}y\\in\\mathbb\{R\}^\{2d\}\\\\ L^\{\*\}\(A,y\)\>0\\end\{subarray\}\}\\Delta\(A,y;r\)\.\(200\)
The next balanced structural gap has two distinct branches\. Its universal lower side \(for every design and every response in the stated class\) follows from Theorem[2](https://arxiv.org/html/2608.26877#Thmtheorem2)after the angle condition boundsRR; its single\-pair upper side is a separate deterministic witness\. For terminology used in that witness,*full spark*means every indexeddd\-row minor is nonzero; an augmented*coloop*is a row withhi=1h\_\{i\}=1\.
###### Theorem 19\(Balanced structural gap and a restricted matching\-order witness\)\.
Universal lower clause\.For every integerd≥2d\\geq 2, every integer1≤r≤d−11\\leq r\\leq d\-1, everyA∈𝒞d\(σ0\)A\\in\\mathcal\{C\}\_\{d\}\(\\sigma\_\{0\}\), and every fixedy∈ℝ2dy\\in\\mathbb\{R\}^\{2d\}withL∗\(A,y\)\>0L^\{\*\}\(A,y\)\>0,
Δ\(A,y,r\)≥3r38\(d−1\)\.\\Delta\(A,y;r\)\\geq\\frac\{3r\}\{38\(d\-1\)\}\.\(201\)
Single\-pair witness clause\.For everyd=4nd=4^\{n\},n≥1n\\geq 1, there exists one deterministicAd∈𝒞d\(σ0\)A\_\{d\}\\in\\mathcal\{C\}\_\{d\}\(\\sigma\_\{0\}\)and one fixedyd∈ℝ2dy\_\{d\}\\in\\mathbb\{R\}^\{2d\}withL∗\(Ad,yd\)\>0L^\{\*\}\(A\_\{d\},y\_\{d\}\)\>0such that the same pair works simultaneously for every integer1≤r<d/21\\leq r<d/2:
Δ\(Ad,yd,r\)<rd−r\.\\Delta\(A\_\{d\},y\_\{d\};r\)<\\frac\{r\}\{d\-r\}\.\(202\)
Only on exactlyd=4nd=4^\{n\}and1≤r<d/21\\leq r<d/2does the class\-infimum squeeze follow:
3r38\(d−1\)≤Δd,r∗\(σ0\)<rd−r\.\\boxed\{\\frac\{3r\}\{38\(d\-1\)\}\\leq\\Delta^\{\*\}\_\{d,r\}\(\\sigma\_\{0\}\)<\\frac\{r\}\{d\-r\}\}\.\(203\)Thus the universal clause gives anΩ\(r/d\)\\Omega\(r/d\)gap on its full stated class, whereas theΘ\(r/d\)\\Theta\(r/d\)class\-infimum squeeze holds only alongd=4nd=4^\{n\}and1≤r<d/21\\leq r<d/2\. Equivalently, the normalized worst\-direction ratio1−Δ\(Ad,yd,r\)1\-\\Delta\(A\_\{d\},y\_\{d\};r\)approaches the universal envelope only forr=o\(d\)r=o\(d\)alongd=4nd=4^\{n\}, not for a fixed positive budget fraction or for all budgets\. The class definition assumes isotropy, equal ordinary row norm, and pairwise projective separation only\. The particular witness is ordinary full spark and angle separated but has augmented colooph1=1h\_\{1\}=1; no augmented\-margin claim follows\.
The universal lower side follows from the response\-aware linear consequence after pairwise separation givesR⪰\(3/38\)IdR\\succeq\(3/38\)I\_\{d\}\. The single\-pair side is a separate deterministic witness calculation; it does not turn the witness into a statement about every design or response\. The full\-spark construction, Schur calculation, and quantifier bookkeeping are in Appendix[H\.2](https://arxiv.org/html/2608.26877#A8.SS2); the frame context is credited at its point of use\[[25](https://arxiv.org/html/2608.26877#bib.bib25),[2](https://arxiv.org/html/2608.26877#bib.bib2)\]\.
### H\.2Balanced projective separation and the worst direction
This section uses ordinary, unrescaled, fixed\-size volume sampling on unordered subsets of indexed rows, without replacement, followed by unweighted least squares\. Thus equal rows at different indices remain different observations\. IfX∈ℝm×dX\\in\\mathbb\{R\}^\{m\\times d\}has full column rank andy∈ℝmy\\in\\mathbb\{R\}^\{m\}is fixed, write
G=X⊤X,w∗=G−1X⊤y,e=y−Xw∗,L∗=∥e∥22\.G=X^\{\\top\}X,\\qquad w^\{\*\}=G^\{\-1\}X^\{\\top\}y,\\qquad e=y\-Xw^\{\*\},\\qquad L^\{\*\}=\\lVert e\\rVert\_\{2\}^\{2\}\.\(204\)For an indexedss\-setSS, putGS=XS⊤XSG\_\{S\}=X\_\{S\}^\{\\top\}X\_\{S\}andDS=det\(GS\)D\_\{S\}=\\det\(G\_\{S\}\)\. Fixed\-cardinality Cauchy–Binet gives
ZX,s:=∑\|S\|=sDS=\(m−ds−d\)det\(G\),ℙX\(S\)=DSZX,s\.Z\_\{X,s\}:=\\sum\_\{\|S\|=s\}D\_\{S\}=\\binom\{m\-d\}\{s\-d\}\\det\(G\),\\qquad\\mathbb\{P\}\_\{X\}\(S\)=\\frac\{D\_\{S\}\}\{Z\_\{X,s\}\}\.\(205\)Only sets withDS\>0D\_\{S\}\>0support an estimator, and on such sets
wS=GS−1XS⊤yS,Ms=𝔼X\[\(wS−w∗\)\(wS−w∗\)⊤\],M¯s=G1/2MsG1/2\.w\_\{S\}=G\_\{S\}^\{\-1\}X\_\{S\}^\{\\top\}y\_\{S\},\\qquad M\_\{s\}=\\mathbb\{E\}\_\{X\}\[\(w\_\{S\}\-w^\{\*\}\)\(w\_\{S\}\-w^\{\*\}\)^\{\\top\}\],\\qquad\\overline\{M\}\_\{s\}=G^\{1/2\}M\_\{s\}G^\{1/2\}\.\(206\)There is no estimator or pseudoinverse on a zero\-volume set\. In particular, there is no rescaling, importance weight, ridge term, replacement, or stochastic\-response interpretation in this section\. Formula \([205](https://arxiv.org/html/2608.26877#A8.E205)\) follows by expanding eachDSD\_\{S\}over its indexeddd\-subsets and observing that each indexeddd\-set has exactly\(m−ds−d\)\\binom\{m\-d\}\{s\-d\}size\-sssupersets\.
We first record the response\-aware inequality used below\. Its proof is included to make clear which identity requires row general position and which inequality remains valid without it\.
###### Proposition 20\(Linear response\-aware covariance envelope\)\.
Assumem\>dm\>d,L∗\>0L^\{\*\}\>0, andd<s<md<s<m\. Define
A=XG−1/2,z=eL∗,B=\[Az\],bi⊤=\[ai⊤zi\]\.A=XG^\{\-1/2\},\\qquad z=\\frac\{e\}\{\\sqrt\{L^\{\*\}\}\},\\qquad B=\[A\\ z\],\\qquad b\_\{i\}^\{\\top\}=\[a\_\{i\}^\{\\top\}\\ z\_\{i\}\]\.\(207\)ThenB⊤B=Id\+1B^\{\\top\}B=I\_\{d\+1\}\. With
hi=∥bi∥22,qi=1−hi,R=∑i=1mqiaiai⊤,h\_\{i\}=\\lVert b\_\{i\}\\rVert\_\{2\}^\{2\},\\qquad q\_\{i\}=1\-h\_\{i\},\\qquad R=\\sum\_\{i=1\}^\{m\}q\_\{i\}a\_\{i\}a\_\{i\}^\{\\top\},\(208\)and
α=m−sm−d,β=s−dm−d,γ=m−sm−d−1,\\alpha=\\frac\{m\-s\}\{m\-d\},\\qquad\\beta=\\frac\{s\-d\}\{m\-d\},\\qquad\\gamma=\\frac\{m\-s\}\{m\-d\-1\},\(209\)one has, for every full\-column\-rankXX,
M¯s⪯αL∗Id−βγL∗R\.\\overline\{M\}\_\{s\}\\preceq\\alpha L^\{\*\}I\_\{d\}\-\\beta\\gamma L^\{\*\}R\.\(210\)
###### Proof\.
The diagonal of the orthogonal projectorBB⊤BB^\{\\top\}lies in\[0,1\]\[0,1\], soqi≥0q\_\{i\}\\geq 0andR⪰0R\\succeq 0\. First suppose thatXXis in row general position\. For a sampled set write
KS=AS⊤AS,uS=KS−1AS⊤zS,ℓS=∥zS−ASuS∥22\.K\_\{S\}=A\_\{S\}^\{\\top\}A\_\{S\},\\qquad u\_\{S\}=K\_\{S\}^\{\-1\}A\_\{S\}^\{\\top\}z\_\{S\},\\qquad\\ell\_\{S\}=\\lVert z\_\{S\}\-A\_\{S\}u\_\{S\}\\rVert\_\{2\}^\{2\}\.\(211\)Fort∈ℝdt\\in\\mathbb\{R\}^\{d\}, orthogonally decomposezS=ASuS\+\(zS−ASuS\)z\_\{S\}=A\_\{S\}u\_\{S\}\+\(z\_\{S\}\-A\_\{S\}u\_\{S\}\)and apply the matrix determinant lemma\. This gives
det\(\(AS\+zSt⊤\)⊤\(AS\+zSt⊤\)\)=det\(KS\)\{\(1\+t⊤uS\)2\+ℓSt⊤KS−1t\}\.\\begin\{split\}&\\det\\\!\\left\(\(A\_\{S\}\+z\_\{S\}t^\{\\top\}\)^\{\\top\}\(A\_\{S\}\+z\_\{S\}t^\{\\top\}\)\\right\)\\\\ &\\qquad=\\det\(K\_\{S\}\)\\left\\\{\(1\+t^\{\\top\}u\_\{S\}\)^\{2\}\+\\ell\_\{S\}t^\{\\top\}K\_\{S\}^\{\-1\}t\\right\\\}\.\\end\{split\}\(212\)On the other hand,\(A\+zt⊤\)⊤\(A\+zt⊤\)=Id\+tt⊤\(A\+zt^\{\\top\}\)^\{\\top\}\(A\+zt^\{\\top\}\)=I\_\{d\}\+tt^\{\\top\}\. Summing \([212](https://arxiv.org/html/2608.26877#A8.E212)\) over allss\-sets and using fixed\-cardinality Cauchy–Binet, then comparing the linear and quadratic coefficients intt, yields
𝔼XuS=0,𝔼X\[uSuS⊤\+ℓSKS−1\]=Id\.\\mathbb\{E\}\_\{X\}u\_\{S\}=0,\\qquad\\mathbb\{E\}\_\{X\}\\\!\\left\[u\_\{S\}u\_\{S\}^\{\\top\}\+\\ell\_\{S\}K\_\{S\}^\{\-1\}\\right\]=I\_\{d\}\.\(213\)The coefficient error satisfiesG1/2\(wS−w∗\)=L∗uSG^\{1/2\}\(w\_\{S\}\-w^\{\*\}\)=\\sqrt\{L^\{\*\}\}\\,u\_\{S\}\. Hence
M¯sL∗=Id−𝔼X\[ℓSKS−1\]\.\\frac\{\\overline\{M\}\_\{s\}\}\{L^\{\*\}\}=I\_\{d\}\-\\mathbb\{E\}\_\{X\}\[\\ell\_\{S\}K\_\{S\}^\{\-1\}\]\.\(214\)
For analysis only, letPBP\_\{B\}denote the response\-dependent augmented volume law induced by the full residual\. It reweights the proof distribution; the estimator continues to sample and fit under the ordinary feature\-only lawPXP\_\{X\}\. On this analysis law, rank\-\(d\+1\)\(d\+1\)sets have augmented volume probabilities
ZB,s=\(m−d−1s−d−1\),ℙB\(S\)=det\(BS⊤BS\)ZB,s\.Z\_\{B,s\}=\\binom\{m\-d\-1\}\{s\-d\-1\},\\qquad\\mathbb\{P\}\_\{B\}\(S\)=\\frac\{\\det\(B\_\{S\}^\{\\top\}B\_\{S\}\)\}\{Z\_\{B,s\}\}\.\(215\)Augmented support impliesKS≻0K\_\{S\}\\succ 0, and the Schur complement gives, including the zero\-equals\-zero unsupported case,
det\(BS⊤BS\)=det\(KS\)ℓS\.\\det\(B\_\{S\}^\{\\top\}B\_\{S\}\)=\\det\(K\_\{S\}\)\\ell\_\{S\}\.\(216\)The ratio of the two normalizers isZB,s/ZX,s=βZ\_\{B,s\}/Z\_\{X,s\}=\\betaafter whitening\. Therefore
ℙX\(S\)ℓS=βℙB\(S\),M¯sL∗=Id−β𝔼B\[KS−1\]\.\\mathbb\{P\}\_\{X\}\(S\)\\ell\_\{S\}=\\beta\\mathbb\{P\}\_\{B\}\(S\),\\qquad\\frac\{\\overline\{M\}\_\{s\}\}\{L^\{\*\}\}=I\_\{d\}\-\\beta\\mathbb\{E\}\_\{B\}\[K\_\{S\}^\{\-1\}\]\.\(217\)The identity in \([217](https://arxiv.org/html/2608.26877#A8.E217)\) is being asserted here only under row general position,L∗\>0L^\{\*\}\>0, andd<s<md<s<m\.
The fixed\-cardinality Cauchy–Binet step is part of the ordinary volume\-law machinery\[[12](https://arxiv.org/html/2608.26877#bib.bib12)\]; applying it here and usingdet\(Id\+1−bibi⊤\)=1−hi\\det\(I\_\{d\+1\}\-b\_\{i\}b\_\{i\}^\{\\top\}\)=1\-h\_\{i\}gives the augmented exclusion marginal
ℙB\(i∉S\)=\(m−d−2s−d−1\)\(m−d−1s−d−1\)\(1−hi\)=γqi\.\\mathbb\{P\}\_\{B\}\(i\\notin S\)=\\frac\{\\binom\{m\-d\-2\}\{s\-d\-1\}\}\{\\binom\{m\-d\-1\}\{s\-d\-1\}\}\(1\-h\_\{i\}\)=\\gamma q\_\{i\}\.\(218\)Consequently
𝔼BKS=Id−∑iℙB\(i∉S\)aiai⊤=Id−γR≻0\.\\mathbb\{E\}\_\{B\}K\_\{S\}=I\_\{d\}\-\\sum\_\{i\}\\mathbb\{P\}\_\{B\}\(i\\notin S\)a\_\{i\}a\_\{i\}^\{\\top\}=I\_\{d\}\-\\gamma R\\succ 0\.\(219\)For completeness, the inverse\-moment inequality used next follows by averaging the positive semidefinite block matrices
\[KSIdIdKS−1\]⪰0\\begin\{bmatrix\}K\_\{S\}&I\_\{d\}\\\\ I\_\{d\}&K\_\{S\}^\{\-1\}\\end\{bmatrix\}\\succeq 0\(220\)and taking a Schur complement:
𝔼B\[KS−1\]⪰\(𝔼BKS\)−1\.\\mathbb\{E\}\_\{B\}\[K\_\{S\}^\{\-1\}\]\\succeq\(\\mathbb\{E\}\_\{B\}K\_\{S\}\)^\{\-1\}\.\(221\)Combining \([217](https://arxiv.org/html/2608.26877#A8.E217)\)– \([221](https://arxiv.org/html/2608.26877#A8.E221)\), with attention to the minus sign, and using\(Id−γR\)−1⪰Id\+γR\(I\_\{d\}\-\\gamma R\)^\{\-1\}\\succeq I\_\{d\}\+\\gamma R, gives
M¯sL∗⪯Id−β\(Id−γR\)−1⪯αId−βγR\.\\frac\{\\overline\{M\}\_\{s\}\}\{L^\{\*\}\}\\preceq I\_\{d\}\-\\beta\(I\_\{d\}\-\\gamma R\)^\{\-1\}\\preceq\\alpha I\_\{d\}\-\\beta\\gamma R\.\(222\)
It remains to justify the last inequality for a design that is not in row general position\. PerturbXXthrough row\-general\-position matricesXt→XX\_\{t\}\\to X, keepingyyfixed\. For smalltt, full column rank andLt∗\>0L\_\{t\}^\{\*\}\>0persist\. For every fixed coefficient directionvv, write the left side of the corresponding original\-coordinate inequality as the finite determinant\-weighted sum
1ZXt,s∑\|S\|=sDS,t\{v⊤\(wS,t−wt∗\)\}2\.\\frac\{1\}\{Z\_\{X\_\{t\},s\}\}\\sum\_\{\|S\|=s\}D\_\{S,t\}\\\{v^\{\\top\}\(w\_\{S,t\}\-w\_\{t\}^\{\*\}\)\\\}^\{2\}\.\(223\)Every term whose limiting determinant is positive converges to the corresponding supported term forXX\. The remaining terms are nonnegative and may be discarded in the liminf\. The normalizer,wt∗w\_\{t\}^\{\*\},Lt∗L\_\{t\}^\{\*\},Gt−1G\_\{t\}^\{\-1\}, and
Qt=∑i\(1−hi,t\)xi,txi,t⊤Q\_\{t\}=\\sum\_\{i\}\(1\-h\_\{i,t\}\)x\_\{i,t\}x\_\{i,t\}^\{\\top\}\(224\)all converge ordinarily\. Thus the row\-general\-position inequality passes one\-sidedly to
Ms⪯αL∗G−1−βγL∗G−1QG−1\.M\_\{s\}\\preceq\\alpha L^\{\*\}G^\{\-1\}\-\\beta\\gamma L^\{\*\}G^\{\-1\}QG^\{\-1\}\.\(225\)Congruence byG1/2G^\{1/2\}is exactly \([210](https://arxiv.org/html/2608.26877#A8.E210)\)\. This argument never continues a singular selected inverse\. Finally, the determinant\-weighted mean has a polynomial continuation:DSwS=adj\(GS\)XS⊤ySD\_\{S\}w\_\{S\}=\\operatorname\{adj\}\(G\_\{S\}\)X\_\{S\}^\{\\top\}y\_\{S\}, and this numerator vanishes whenGSG\_\{S\}is singular\. Hence𝔼XwS=w∗\\mathbb\{E\}\_\{X\}w\_\{S\}=w^\{\*\}also on the boundary, so the centered second moment in \([206](https://arxiv.org/html/2608.26877#A8.E206)\) is indeed the covariance\. ∎
We now specialize to the balanced class\. Ford≥2d\\geq 2and0<σ0≤10<\\sigma\_\{0\}\\leq 1, define
𝒞d\(σ0\)=\{A∈ℝ2d×d:A⊤A=Id,∥ai∥22=12for everyi,mini≠jsin∠proj\(ai,aj\)≥σ0\},\\mathcal\{C\}\_\{d\}\(\\sigma\_\{0\}\)=\\left\\\{A\\in\\mathbb\{R\}^\{2d\\times d\}:\\begin\{aligned\} &A^\{\\top\}A=I\_\{d\},\\quad\\lVert a\_\{i\}\\rVert\_\{2\}^\{2\}=\\frac\{1\}\{2\}\\ \\text\{for every \}i,\\\\ &\\min\_\{i\\neq j\}\\sin\\angle\_\{\\mathrm\{proj\}\}\(a\_\{i\},a\_\{j\}\)\\geq\\sigma\_\{0\}\\end\{aligned\}\\right\\\},\(226\)where
sin∠proj\(ai,aj\)=1−\(ai⊤aj\)2∥ai∥22∥aj∥22\.\\sin\\angle\_\{\\mathrm\{proj\}\}\(a\_\{i\},a\_\{j\}\)=\\sqrt\{1\-\\frac\{\(a\_\{i\}^\{\\top\}a\_\{j\}\)^\{2\}\}\{\\lVert a\_\{i\}\\rVert\_\{2\}^\{2\}\\lVert a\_\{j\}\\rVert\_\{2\}^\{2\}\}\}\.\(227\)Full spark is not a premise in \([226](https://arxiv.org/html/2608.26877#A8.E226)\)\.
ForA∈𝒞d\(σ0\)A\\in\\mathcal\{C\}\_\{d\}\(\\sigma\_\{0\}\), useX=AX=Ain \([205](https://arxiv.org/html/2608.26877#A8.E205)\)–\([206](https://arxiv.org/html/2608.26877#A8.E206)\)\. For a fixed responsey∈ℝ2dy\\in\\mathbb\{R\}^\{2d\}satisfyingL∗\(A,y\)\>0L^\{\*\}\(A,y\)\>0and an integer1≤r≤d−11\\leq r\\leq d\-1, define
Δ\(A,y,r\)=1−λmax\(M¯d\+r\)\(1−r/d\)L∗\(A,y\)\\Delta\(A,y;r\)=1\-\\frac\{\\lambda\_\{\\max\}\(\\overline\{M\}\_\{d\+r\}\)\}\{\(1\-r/d\)L^\{\*\}\(A,y\)\}\(228\)and, with the response domain written explicitly,
Δd,r∗\(σ0\)=infA∈𝒞d\(σ0\)infy∈ℝ2dL∗\(A,y\)\>0Δ\(A,y,r\)\.\\Delta^\{\*\}\_\{d,r\}\(\\sigma\_\{0\}\)=\\inf\_\{A\\in\\mathcal\{C\}\_\{d\}\(\\sigma\_\{0\}\)\}\\inf\_\{\\begin\{subarray\}\{c\}y\\in\\mathbb\{R\}^\{2d\}\\\\ L^\{\*\}\(A,y\)\>0\\end\{subarray\}\}\\Delta\(A,y;r\)\.\(229\)
###### Proposition 21\(Uniform gap from pairwise projective separation\)\.
Letd≥2d\\geq 2,1≤r≤d−11\\leq r\\leq d\-1, and0<σ≤10<\\sigma\\leq 1\. For everyA∈𝒞d\(σ\)A\\in\\mathcal\{C\}\_\{d\}\(\\sigma\)and every fixedy∈ℝ2dy\\in\\mathbb\{R\}^\{2d\}withL∗\(A,y\)\>0L^\{\*\}\(A,y\)\>0,
Δ\(A,y,r\)≥rd−1ψ\(σ\),ψ\(σ\)=1−1−σ22\(3−1−σ2\)\.\\Delta\(A,y;r\)\\geq\\frac\{r\}\{d\-1\}\\psi\(\\sigma\),\\qquad\\psi\(\\sigma\)=\\frac\{1\-\\sqrt\{1\-\\sigma^\{2\}\}\}\{2\(3\-\\sqrt\{1\-\\sigma^\{2\}\}\)\}\.\(230\)
###### Proof\.
BecauseA⊤A=IdA^\{\\top\}A=I\_\{d\}, the full residual obeysA⊤z=0A^\{\\top\}z=0and∥z∥2=1\\lVert z\\rVert\_\{2\}=1, wherez=e/L∗z=e/\\sqrt\{L^\{\*\}\}\. In the notation of \([208](https://arxiv.org/html/2608.26877#A8.E208)\), equal row norm gives
qi=12−zi2\.q\_\{i\}=\\frac\{1\}\{2\}\-z\_\{i\}^\{2\}\.\(231\)These numbers are nonnegative because\[Az\]\[A\\ z\]has orthonormal columns\. For an arbitrary unit vectoru∈ℝdu\\in\\mathbb\{R\}^\{d\}, putti=\(ai⊤u\)2t\_\{i\}=\(a\_\{i\}^\{\\top\}u\)^\{2\}\. Parseval and the row\-norm constraint give
∑i=12dti=1,0≤ti≤12\.\\sum\_\{i=1\}^\{2d\}t\_\{i\}=1,\\qquad 0\\leq t\_\{i\}\\leq\\frac\{1\}\{2\}\.\(232\)Ifq\(1\)≤q\(2\)≤⋯q\_\{\(1\)\}\\leq q\_\{\(2\)\}\\leq\\cdotsare the orderedqiq\_\{i\}, minimizing a linear functional over this capped simplex places mass1/21/2on each of its two smallest coefficients\. Therefore
u⊤Ru=∑iqiti≥q\(1\)\+q\(2\)2\.u^\{\\top\}Ru=\\sum\_\{i\}q\_\{i\}t\_\{i\}\\geq\\frac\{q\_\{\(1\)\}\+q\_\{\(2\)\}\}\{2\}\.\(233\)
Leti,ji,jattain the two smallest coefficients, putQ=qi\+qjQ=q\_\{i\}\+q\_\{j\}, and setc=1−σ2c=\\sqrt\{1\-\\sigma^\{2\}\}\. Equation \([231](https://arxiv.org/html/2608.26877#A8.E231)\) and∥z∥2=1\\lVert z\\rVert\_\{2\}=1imply
zi2\+zj2=1−Q,∑k≠i,jzk2=Q\.z\_\{i\}^\{2\}\+z\_\{j\}^\{2\}=1\-Q,\\qquad\\sum\_\{k\\neq i,j\}z\_\{k\}^\{2\}=Q\.\(234\)The residual normal equations give
ziai\+zjaj=−∑k≠i,jzkak\.z\_\{i\}a\_\{i\}\+z\_\{j\}a\_\{j\}=\-\\sum\_\{k\\neq i,j\}z\_\{k\}a\_\{k\}\.\(235\)The squared norm of the right side is at mostQQ, since∑k≠i,jakak⊤⪯Id\\sum\_\{k\\neq i,j\}a\_\{k\}a\_\{k\}^\{\\top\}\\preceq I\_\{d\}\. Pairwise projective separation is used only on the left side: it gives\|ai⊤aj\|≤c/2\|a\_\{i\}^\{\\top\}a\_\{j\}\|\\leq c/2, and hence
∥ziai\+zjaj∥22≥12\(zi2\+zj2\)−c\|zizj\|≥1−c2\(1−Q\)\.\\begin\{split\}\\lVert z\_\{i\}a\_\{i\}\+z\_\{j\}a\_\{j\}\\rVert\_\{2\}^\{2\}&\\geq\\frac\{1\}\{2\}\(z\_\{i\}^\{2\}\+z\_\{j\}^\{2\}\)\-c\|z\_\{i\}z\_\{j\}\|\\\\ &\\geq\\frac\{1\-c\}\{2\}\(1\-Q\)\.\\end\{split\}\(236\)Comparison with the right side of \([235](https://arxiv.org/html/2608.26877#A8.E235)\) yields
qi\+qj=Q≥1−c3−c\.q\_\{i\}\+q\_\{j\}=Q\\geq\\frac\{1\-c\}\{3\-c\}\.\(237\)Together with \([233](https://arxiv.org/html/2608.26877#A8.E233)\), this proves
R⪰ψ\(σ\)Id\.R\\succeq\\psi\(\\sigma\)I\_\{d\}\.\(238\)
Proposition[20](https://arxiv.org/html/2608.26877#Thmtheorem20)applies even if this member of𝒞d\(σ\)\\mathcal\{C\}\_\{d\}\(\\sigma\)is not in row general position\. Form=2dm=2dands=d\+rs=d\+r,
α=1−rd,β=rd,γ=d−rd−1,βγα=rd−1\.\\alpha=1\-\\frac\{r\}\{d\},\\qquad\\beta=\\frac\{r\}\{d\},\\qquad\\gamma=\\frac\{d\-r\}\{d\-1\},\\qquad\\frac\{\\beta\\gamma\}\{\\alpha\}=\\frac\{r\}\{d\-1\}\.\(239\)Thus \([210](https://arxiv.org/html/2608.26877#A8.E210)\) and \([238](https://arxiv.org/html/2608.26877#A8.E238)\) give
λmax\(M¯d\+r\)\(1−r/d\)L∗≤1−rd−1ψ\(σ\),\\frac\{\\lambda\_\{\\max\}\(\\overline\{M\}\_\{d\+r\}\)\}\{\(1\-r/d\)L^\{\*\}\}\\leq 1\-\\frac\{r\}\{d\-1\}\\psi\(\\sigma\),\(240\)which is equivalent to \([230](https://arxiv.org/html/2608.26877#A8.E230)\)\. ∎
Pairwise projective separation entered only the two\-row estimate \([236](https://arxiv.org/html/2608.26877#A8.E236)\)\. It supplies neither a common margin fordd\-row minors nor a bound away from one on the augmented leverageshih\_\{i\}\.
We next construct the design and response used for the opposite side of the gap\. The deterministic choice in the construction is formal rather than numerical\.
###### Lemma 22\(Deterministic rational orthogonal perturbation\)\.
For everyd=4nd=4^\{n\},n≥1n\\geq 1, there is a deterministically specified rational matrixQd∈O\(d\)Q\_\{d\}\\in O\(d\)such that every square minor ofQdQ\_\{d\}is nonzero and
∥Qd−Hd∥max<18,\\lVert Q\_\{d\}\-H\_\{d\}\\rVert\_\{\\max\}<\\frac\{1\}\{8\},\(241\)whereHdH\_\{d\}is the normalized Sylvester matrix\.
###### Proof\.
The entries ofHdH\_\{d\}are±d−1/2=±2−n\\pm d^\{\-1/2\}=\\pm 2^\{\-n\}, soHdH\_\{d\}is rational\. We now declare the enumeration that definesQdQ\_\{d\}\. Enumerateℚ\\mathbb\{Q\}by writing each rational uniquely asp/qp/q, withq≥1q\\geq 1andgcd\(\|p\|,q\)=1\\gcd\(\|p\|,q\)=1, in increasing order of\|p\|\+q\|p\|\+q, breaking ties byppand thenqq\. For1≤i<j≤d1\\leq i<j\\leq dandt∈ℚt\\in\\mathbb\{Q\}, letGij\(t\)G\_\{ij\}\(t\)be the coordinate\-plane rotation whose nontrivial block is
11\+t2\[1−t2−2t2t1−t2\]\.\\frac\{1\}\{1\+t^\{2\}\}\\begin\{bmatrix\}1\-t^\{2\}&\-2t\\\\ 2t&1\-t^\{2\}\\end\{bmatrix\}\.\(242\)Order the triples\(i,j,t\)\(i,j,t\)first by the index ofttin the preceding enumeration and then lexicographically by\(i,j\)\(i,j\)\. Enumerate all finite words in these triples by stages: at stageNN, list lexicographically every previously unlisted word of length at mostNNwhose letter indices are at mostNN\. Include the empty word\. IfR1,R2,…R\_\{1\},R\_\{2\},\\ldotsare the corresponding products, thenRk∈SO\(d\)∩ℚd×dR\_\{k\}\\in SO\(d\)\\cap\\mathbb\{Q\}^\{d\\times d\}andRkHdR\_\{k\}H\_\{d\}lies in the connected component ofO\(d\)O\(d\)containingHdH\_\{d\}\. DefineQdQ\_\{d\}to be the firstRkHdR\_\{k\}H\_\{d\}satisfying, in exact rational arithmetic, both \([241](https://arxiv.org/html/2608.26877#A8.E241)\) and the nonvanishing of every square minor\.
It remains to show that the search terminates\. Rational points of the form in \([242](https://arxiv.org/html/2608.26877#A8.E242)\) are dense on the unit circle, and coordinate\-plane rotations generateSO\(d\)SO\(d\)\. Thus the enumerated products are dense inSO\(d\)SO\(d\), and\{RkHd\}\\\{R\_\{k\}H\_\{d\}\\\}is dense in the component containingHdH\_\{d\}\. For any fixed proper square minor, its determinant polynomial is not identically zero on either component ofO\(d\)O\(d\): a signed permutation can make that minor nonzero, and a sign on an unused coordinate can be chosen to fix the component\. The full minor is always nonzero onO\(d\)O\(d\)\. Hence the zero set of each minor has empty interior in either component\. There are only finitely many square minors, so the totally nonsingular locus is dense and open in each component\. Its intersection with the open entrywise1/81/8\-ball aboutHdH\_\{d\}is nonempty and open\. Density of the enumerated family therefore places someRkHdR\_\{k\}H\_\{d\}in that intersection, proving termination and making the first eligible matrix a deterministic choice\. ∎
###### Proposition 23\(Separated full\-spark witness with a large covariance direction\)\.
Letd=4nd=4^\{n\},n≥1n\\geq 1, and useQdQ\_\{d\}from Lemma[22](https://arxiv.org/html/2608.26877#Thmtheorem22)\. Set
Ad=12\[IdQd\],zd=12\[e1−Qde1\],yd=zd\.A\_\{d\}=\\frac\{1\}\{\\sqrt\{2\}\}\\begin\{bmatrix\}I\_\{d\}\\\\ Q\_\{d\}\\end\{bmatrix\},\\qquad z\_\{d\}=\\frac\{1\}\{\\sqrt\{2\}\}\\begin\{bmatrix\}e\_\{1\}\\\\ \-Q\_\{d\}e\_\{1\}\\end\{bmatrix\},\\qquad y\_\{d\}=z\_\{d\}\.\(243\)ThenAd∈𝒞d\(39/8\)A\_\{d\}\\in\\mathcal\{C\}\_\{d\}\(\\sqrt\{39\}/8\)\. In addition, this particular witness is full spark\. The same design\-response pair satisfies, for every integer1≤r<d/21\\leq r<d/2,
Δ\(Ad,yd,r\)<rd−r\.\\Delta\(A\_\{d\},y\_\{d\};r\)<\\frac\{r\}\{d\-r\}\.\(244\)
###### Proof\.
Orthogonality ofQdQ\_\{d\}gives
Ad⊤Ad=12\(Id\+Qd⊤Qd\)=Id,∥ai∥22=12\(1≤i≤2d\),A\_\{d\}^\{\\top\}A\_\{d\}=\\frac\{1\}\{2\}\(I\_\{d\}\+Q\_\{d\}^\{\\top\}Q\_\{d\}\)=I\_\{d\},\\qquad\\lVert a\_\{i\}\\rVert\_\{2\}^\{2\}=\\frac\{1\}\{2\}\\quad\(1\\leq i\\leq 2d\),\(245\)and
Ad⊤zd=12\(e1−Qd⊤Qde1\)=0,∥zd∥22=1\.A\_\{d\}^\{\\top\}z\_\{d\}=\\frac\{1\}\{2\}\(e\_\{1\}\-Q\_\{d\}^\{\\top\}Q\_\{d\}e\_\{1\}\)=0,\\qquad\\lVert z\_\{d\}\\rVert\_\{2\}^\{2\}=1\.\(246\)Thus the full least\-squares coefficient forydy\_\{d\}isw∗=0w^\{\*\}=0, its residual iszdz\_\{d\}, andL∗=1L^\{\*\}=1\.
Select top\-block indicesI⊆\[d\]I\\subseteq\[d\]and bottom\-block indicesJ⊆\[d\]J\\subseteq\[d\]with\|I\|\+\|J\|=d\|I\|\+\|J\|=d\. Expansion along the selected identity rows gives
det\(Ad\)I∪\(d\+J\)=±2−d/2det\(Qd\)J,Ic\.\\det\(A\_\{d\}\)\_\{I\\cup\(d\+J\)\}=\\mathord\{\\pm\}2^\{\-d/2\}\\det\(Q\_\{d\}\)\_\{J,I^\{c\}\}\.\(247\)The empty determinant is one\. Total nonsingularity ofQdQ\_\{d\}makes every quantity in \([247](https://arxiv.org/html/2608.26877#A8.E247)\) nonzero, so everyddrows of this witness are independent\. This verifies full spark for the witness; it does not add full spark to the definition of𝒞d\\mathcal\{C\}\_\{d\}\.
Rows in the same block are orthogonal\. The normalized absolute inner product of top rowiiand bottom rowjjis\|\(Qd\)ji\|\|\(Q\_\{d\}\)\_\{ji\}\|\. Sinced≥4d\\geq 4,
\|\(Qd\)ji\|<d−1/2\+18≤58,\|\(Q\_\{d\}\)\_\{ji\}\|<d^\{\-1/2\}\+\\frac\{1\}\{8\}\\leq\\frac\{5\}\{8\},\(248\)and therefore every cross\-block projective sine is strictly larger than
1−\(5/8\)2=398\.\\sqrt\{1\-\(5/8\)^\{2\}\}=\\frac\{\\sqrt\{39\}\}\{8\}\.\(249\)Together with \([245](https://arxiv.org/html/2608.26877#A8.E245)\), this proves the asserted class membership\.
Fixs=d\+rs=d\+r\. Cauchy–Binet gives the exact ordinary normalizer
ZA=\(2d−dd\+r−d\)det\(Ad⊤Ad\)=\(dr\)\.Z\_\{A\}=\\binom\{2d\-d\}\{d\+r\-d\}\\det\(A\_\{d\}^\{\\top\}A\_\{d\}\)=\\binom\{d\}\{r\}\.\(250\)ForB=\[Adzd\]B=\[A\_\{d\}\\ z\_\{d\}\], equations \([245](https://arxiv.org/html/2608.26877#A8.E245)\)–\([246](https://arxiv.org/html/2608.26877#A8.E246)\) giveB⊤B=Id\+1B^\{\\top\}B=I\_\{d\+1\}and hence
ZB=\(2d−\(d\+1\)d\+r−\(d\+1\)\)=\(d−1r−1\),ZBZA=rd=β\.Z\_\{B\}=\\binom\{2d\-\(d\+1\)\}\{d\+r\-\(d\+1\)\}=\\binom\{d\-1\}\{r\-1\},\\qquad\\frac\{Z\_\{B\}\}\{Z\_\{A\}\}=\\frac\{r\}\{d\}=\\beta\.\(251\)
The augmented leverage of the first top row is, at first use,
h1=∥a1∥22\+z12=12\+12=1\.h\_\{1\}=\\lVert a\_\{1\}\\rVert\_\{2\}^\{2\}\+z\_\{1\}^\{2\}=\\frac\{1\}\{2\}\+\\frac\{1\}\{2\}=1\.\(252\)Rotate the first coefficient column and the residual column orthogonally:
c\+=Ade1\+zd2=\[e10\],c−=Ade1−zd2=\[0Qde1\]\.c\_\{\+\}=\\frac\{A\_\{d\}e\_\{1\}\+z\_\{d\}\}\{\\sqrt\{2\}\}=\\begin\{bmatrix\}e\_\{1\}\\\\ 0\\end\{bmatrix\},\\qquad c\_\{\-\}=\\frac\{A\_\{d\}e\_\{1\}\-z\_\{d\}\}\{\\sqrt\{2\}\}=\\begin\{bmatrix\}0\\\\ Q\_\{d\}e\_\{1\}\\end\{bmatrix\}\.\(253\)The columnc\+c\_\{\+\}is supported only on row11\. Thus every augmented\-supported set contains row11, in agreement with \([252](https://arxiv.org/html/2608.26877#A8.E252)\)\. Removing that row andc\+c\_\{\+\}from the rotatedBBleaves a matrixC∈ℝ\(2d−1\)×dC\\in\\mathbb\{R\}^\{\(2d\-1\)\\times d\}with orthonormal columns\. Every augmented\-supported set has the form
S=\{1\}∪T,\|T\|=d\+r−1,D=CT⊤CT≻0\.S=\\\{1\\\}\\cup T,\\qquad\|T\|=d\+r\-1,\\qquad D=C\_\{T\}^\{\\top\}C\_\{T\}\\succ 0\.\(254\)
Order the columns ofCCasc−c\_\{\-\}followed by the columnsAde2,…,AdedA\_\{d\}e\_\{2\},\\ldots,A\_\{d\}e\_\{d\}, and partitionDDaccordingly\. SinceAde1=\(c\+\+c−\)/2A\_\{d\}e\_\{1\}=\(c\_\{\+\}\+c\_\{\-\}\)/\\sqrt\{2\}, the ordinary selected Gram matrix is
KS=\[\(1\+D11\)/2D1R/2DR1/2DRR\]\.K\_\{S\}=\\begin\{bmatrix\}\(1\+D\_\{11\}\)/2&D\_\{1R\}/\\sqrt\{2\}\\\\ D\_\{R1\}/\\sqrt\{2\}&D\_\{RR\}\\end\{bmatrix\}\.\(255\)Positive definiteness ofDDgives
δ=D11−D1RDRR−1DR1\>0\.\\delta=D\_\{11\}\-D\_\{1R\}D\_\{RR\}^\{\-1\}D\_\{R1\}\>0\.\(256\)Taking the Schur complement ofDRRD\_\{RR\}in \([255](https://arxiv.org/html/2608.26877#A8.E255)\) proves the pointwise identity
e1⊤KS−1e1=21\+δ<2e\_\{1\}^\{\\top\}K\_\{S\}^\{\-1\}e\_\{1\}=\\frac\{2\}\{1\+\\delta\}<2\(257\)on every augmented\-supported set\.
The ordinary design is row general position by \([247](https://arxiv.org/html/2608.26877#A8.E247)\), so the exact identity \([217](https://arxiv.org/html/2608.26877#A8.E217)\) applies\. HereG=IdG=I\_\{d\}andL∗=1L^\{\*\}=1; averaging the strict pointwise bound in \([257](https://arxiv.org/html/2608.26877#A8.E257)\) therefore yields
e1⊤Md\+re1=1−rd𝔼B\[e1⊤KS−1e1\]\>1−2rd\.e\_\{1\}^\{\\top\}M\_\{d\+r\}e\_\{1\}=1\-\\frac\{r\}\{d\}\\,\\mathbb\{E\}\_\{B\}\[e\_\{1\}^\{\\top\}K\_\{S\}^\{\-1\}e\_\{1\}\]\>1\-\\frac\{2r\}\{d\}\.\(258\)Forr<d/2r<d/2, division by1−r/d\>01\-r/d\>0gives
λmax\(M¯d\+r\)\(1−r/d\)L∗\>1−2r/d1−r/d\.\\frac\{\\lambda\_\{\\max\}\(\\overline\{M\}\_\{d\+r\}\)\}\{\(1\-r/d\)L^\{\*\}\}\>\\frac\{1\-2r/d\}\{1\-r/d\}\.\(259\)Subtracting from one proves \([244](https://arxiv.org/html/2608.26877#A8.E244)\)\. NeitherAdA\_\{d\}norydy\_\{d\}depends onrr, so the conclusion holds simultaneously for all integers1≤r<d/21\\leq r<d/2in each fixed dimension\. ∎
###### Theorem 24\(Worst\-direction gap under balanced projective separation\)\.
Let
σ0=398\.\\sigma\_\{0\}=\\frac\{\\sqrt\{39\}\}\{8\}\.\(260\)For everyd=4nd=4^\{n\},n≥1n\\geq 1, and every integer1≤r<d/21\\leq r<d/2,
3r38\(d−1\)≤Δd,r∗\(σ0\)<rd−r\.\\boxed\{\\frac\{3r\}\{38\(d\-1\)\}\\leq\\Delta^\{\*\}\_\{d,r\}\(\\sigma\_\{0\}\)<\\frac\{r\}\{d\-r\}\.\}\(261\)The lower inequality is uniform over every design in𝒞d\(σ0\)\\mathcal\{C\}\_\{d\}\(\\sigma\_\{0\}\)and every fixed response in the domainL∗\(A,y\)\>0L^\{\*\}\(A,y\)\>0\. The strict upper inequality is supplied by the single design\-response pair in Proposition[23](https://arxiv.org/html/2608.26877#Thmtheorem23)\.
###### Proof\.
Since1−σ02=5/8\\sqrt\{1\-\\sigma\_\{0\}^\{2\}\}=5/8,
ψ\(σ0\)=1−5/82\(3−5/8\)=338\.\\psi\(\\sigma\_\{0\}\)=\\frac\{1\-5/8\}\{2\(3\-5/8\)\}=\\frac\{3\}\{38\}\.\(262\)Proposition[21](https://arxiv.org/html/2608.26877#Thmtheorem21), with its all\-design and all\-response quantifiers, gives the lower bound after taking the two infima in \([229](https://arxiv.org/html/2608.26877#A8.E229)\)\. Proposition[23](https://arxiv.org/html/2608.26877#Thmtheorem23)provides one member of the class and one positive\-loss fixed response, so its strict gap bound gives the upper bound on the same infimum\. ∎
On the stated subsequence and budget range, the two sides of \([261](https://arxiv.org/html/2608.26877#A8.E261)\) are constant multiples ofr/dr/d; more explicitly, they lie between\(3/38\)\(r/d\)\(3/38\)\(r/d\)and2r/d2r/d\. Hence the class quantityΔd,r∗\(σ0\)\\Delta^\{\*\}\_\{d,r\}\(\\sigma\_\{0\}\)has orderr/dr/don exactly this domain\. For the explicit witness, \([258](https://arxiv.org/html/2608.26877#A8.E258)\) and the universal covariance envelope squeeze the normalized worst\-direction ratio to one wheneverr=rd=o\(d\)r=r\_\{d\}=o\(d\)alongd=4nd=4^\{n\}\. This last near\-attainment interpretation is not asserted for a positive limiting budget fraction or for all budgets\. Nothing here gives a trace near\-attainment statement, a commondd\-minor margin from pairwise separation, or an augmented\-leverage margin; indeed the witness has the augmented coloop \([252](https://arxiv.org/html/2608.26877#A8.E252)\)\.
## Appendix ISupplementary finite diagnostics
The following finite diagnostics are descriptive supplementary records and do not test the robust phase theorem\.
### I\.1Controlled residual localization
To isolate the response\-geometry term while holding the design, ordinary sampler, estimator, and subset budget fixed, letHdH\_\{d\}be the orthogonal Walsh–Hadamard matrix in Sylvester order and use the construction
Ad\\displaystyle A\_\{d\}=12\[IdHd\],\\displaystyle=\\frac\{1\}\{\\sqrt\{2\}\}\\begin\{bmatrix\}I\_\{d\}\\\\ H\_\{d\}\\end\{bmatrix\},zλ\\displaystyle z\_\{\\lambda\}=12\[vλ−Hdvλ\],\\displaystyle=\\frac\{1\}\{\\sqrt\{2\}\}\\begin\{bmatrix\}v\_\{\\lambda\}\\\\ \-H\_\{d\}v\_\{\\lambda\}\\end\{bmatrix\},\(263\)vλ\\displaystyle v\_\{\\lambda\}=sin\(\(1−λ\)θ\)sinθvloc\+sin\(λθ\)sinθvflat,\\displaystyle=\\frac\{\\sin\(\(1\-\\lambda\)\\theta\)\}\{\\sin\\theta\}v\_\{\\rm loc\}\+\\frac\{\\sin\(\\lambda\\theta\)\}\{\\sin\\theta\}v\_\{\\rm flat\},θ\\displaystyle\\theta=arccos\(vloc⊤vflat\)\.\\displaystyle=\\arccos\(v\_\{\\rm loc\}^\{\\top\}v\_\{\\rm flat\}\)\.whereH1=\[1\]H\_\{1\}=\[1\]andH2d=2−1/2\[HdHdHd−Hd\]H\_\{2d\}=2^\{\-1/2\}\\begin\{bmatrix\}H\_\{d\}&H\_\{d\}\\\\ H\_\{d\}&\-H\_\{d\}\\end\{bmatrix\}\. Ford=4kd=4^\{k\}, writei=∑ℓ=02k−1bℓ2ℓi=\\sum\_\{\\ell=0\}^\{2k\-1\}b\_\{\\ell\}2^\{\\ell\}withbℓ∈\{0,1\}b\_\{\\ell\}\\in\\\{0,1\\\}, and set
\(vflat\)i\+1=d−1/2\(−1\)∑j=0k−1b2jb2j\+1,i=0,…,d−1\.\(v\_\{\\rm flat\}\)\_\{i\+1\}=d^\{\-1/2\}\(\-1\)^\{\\sum\_\{j=0\}^\{k\-1\}b\_\{2j\}b\_\{2j\+1\}\},\\qquad i=0,\\ldots,d\-1\.Thusvloc=e1v\_\{\\rm loc\}=e\_\{1\}andvflatv\_\{\\rm flat\}are fully specified unit directions;vflatv\_\{\\rm flat\}has flat Walsh spectrum for the displayed dimensions\. The path usesλ∈\{0,1/4,1/2,3/4,1\}\\lambda\\in\\\{0,1/4,1/2,3/4,1\\\}\. Form=2dm=2dands=d\+rs=d\+r, writebsb\_\{s\}for the response\-aware ceiling normalized by the universal ceiling,τres=1−bs\\tau\_\{\\rm res\}=1\-b\_\{s\}, andΔ=1−qs\\Delta=1\-q\_\{s\}for the corresponding normalized covariance shortfall, with hats denoting sampled plug\-in quantities\. Theorem[2](https://arxiv.org/html/2608.26877#Thmtheorem2)gives the direct bridge
qs≤bs⟹Δ=1−qs≥1−bs=τres\.q\_\{s\}\\leq b\_\{s\}\\qquad\\Longrightarrow\\qquad\\Delta=1\-q\_\{s\}\\geq 1\-b\_\{s\}=\\tau\_\{\\rm res\}\.\(264\)
Figure 3:Controlled residual\-localization path atr/d=1/2r/d=1/2\. Within each panel the Walsh–Hadamard design and ordinary subset law are fixed; only the residual follows this geodesic\. Curves show deterministic ceiling contractionτres=1−bs\\tau\_\{\\rm res\}=1\-b\_\{s\}and exact \(d=4d=4\) or seeded Monte\-Carlo \(d∈\{16,64\}d\\in\\\{16,64\\\}\) covariance shortfall\. Error bars are eight\-batch delete\-one jackknife diagnostics, not confidence intervals\. This is a finite fixed\-pool mechanism illustration, not theorem validation or a generic effect\.The quantityτres\\tau\_\{\\rm res\}increased at every grid step on all 11 design–budget trajectories, with exact endpointsτres\(0\)=r/\(3d\+r\)\\tau\_\{\\rm res\}\(0\)=r/\(3d\+r\)andτres\(1\)=r/\(d\+r\)\\tau\_\{\\rm res\}\(1\)=r/\(d\+r\)\. Dimension four uses exact determinant enumeration; dimensions 16 and 64 use 16,384 seeded draws per cell and eight\-batch delete\-one jackknife diagnostics\. Every sampled cell’s upper Monte\-Carlo diagnostic endpoint forq^s\\widehat\{q\}\_\{s\}remained belowbsb\_\{s\}at the reported resolution\. Across the five displayed grid points, this records a monotone ceiling contraction with design, sampler, estimator, and budget fixed—a finite mechanism illustration, not proof evidence or a generic response\-direction effect\. Appendix[I\.3\.1](https://arxiv.org/html/2608.26877#A9.SS3.SSS1)and the ancillary cell record provide the corresponding finite computational record\.
### I\.2Public fixed\-pool diagnostic
On these three complete pools, the boundary\-valid one\-sided response\-aware ceiling contracts the universal ceiling strongly, but the plug\-in occupies only11\.3%11\.3\\%–15\.3%15\.3\\%of that ceiling\. Thus the contraction is nonvacuous but substantial slack remains; this is not practical tightness\.
We fix the complete indexed Airfoil Self\-Noise\[[6](https://arxiv.org/html/2608.26877#bib.bib6)\], Concrete Compressive Strength\[[29](https://arxiv.org/html/2608.26877#bib.bib29),[30](https://arxiv.org/html/2608.26877#bib.bib30)\], and Energy Efficiency\[[27](https://arxiv.org/html/2608.26877#bib.bib27)\]pools, their full\-rank representations, and five budgets each\. Energy uses heating load in its frozen 14\-column encoding\. The deterministic same\-pool comparison makes no population, predictive, performance, or inference claim\.
For each listed budget, this diagnostic evaluates the response\-aware ceiling of Theorem[2](https://arxiv.org/html/2608.26877#Thmtheorem2)and does not computecXc\_\{X\}orηX\(s\)\\eta\_\{X\}\(s\)\. In QR\-whitened coordinates, withRRandγ\\gammaas in Section[4](https://arxiv.org/html/2608.26877#S4), define
Us=L∗\{Id−β\(Id−γR\)−1\},bs=λmax\(Us\)αL∗,τs=1−bs\.U\_\{s\}=L^\{\*\}\\\{I\_\{d\}\-\\beta\(I\_\{d\}\-\\gamma R\)^\{\-1\}\\\},\\qquad b\_\{s\}=\\frac\{\\lambda\_\{\\max\}\(U\_\{s\}\)\}\{\\alpha L^\{\*\}\},\\qquad\\tau\_\{s\}=1\-b\_\{s\}\.\(265\)Hereτs\\tau\_\{s\}is a bound contraction, not an actual covariance reduction, andq^s\\widehat\{q\}\_\{s\}is the normalized Monte\-Carlo plug\-in forqsq\_\{s\}\. Across all 15 cells,q^s/bs=0\.1133\\widehat\{q\}\_\{s\}/b\_\{s\}=0\.1133–0\.15340\.1534, or6\.526\.52–8\.828\.82\-fold ceiling\-to\-plug\-in slack\. These finite ratios are not exact values ofqsq\_\{s\}; nor do they provide confidence intervals or evidence of theorem validation, typicality, or population behavior\.
Figure 4:Response\-aware diagnostic for all 15 frozen cells\. Each panel shows five budgets on a common normalized log scale: universal ceiling11, deterministicbsb\_\{s\}, and Monte\-Carlo plug\-inq^s\\widehat\{q\}\_\{s\}\. This is a descriptive same\-pool hierarchy, not exact covariance or a predictive or population comparison\.Appendix[J](https://arxiv.org/html/2608.26877#A10)gives the pool summary \(Table[1](https://arxiv.org/html/2608.26877#A10.T1)\), estimator construction, complete 15\-cell record, recomputation endpoints, and negative companion screen\. The finite Monte\-Carlo output lies below the proved one\-sided ceiling at reported resolution; the endpoints diagnose computation only\.
### I\.3Calculation record
The following table describes the finite diagnostic calculations and what they report\. The complete\-pool calculation is instantiated by the descriptive record in Appendix[J](https://arxiv.org/html/2608.26877#A10)\. None of these computations extends a theorem domain or serves as theorem evidence\.
Table E1:Finite diagnostic calculations and their descriptive outputs\.#### I\.3\.1Controlled residual\-localization record
The controlled path uses the construction in \([263](https://arxiv.org/html/2608.26877#A9.E263)\) atd∈\{4,16,64\}d\\in\\\{4,16,64\\\},m=2dm=2d, and
d=4:r∈\{1,2,3\},d=16:r∈\{1,4,8,12\},d=64:r∈\{1,16,32,48\},λ∈\{0,1/4,1/2,3/4,1\}\.\\begin\{split\}d=4&:\\quad r\\in\\\{1,2,3\\\},\\\\ d=16&:\\quad r\\in\\\{1,4,8,12\\\},\\\\ d=64&:\\quad r\\in\\\{1,16,32,48\\\},\\end\{split\}\\qquad\\lambda\\in\\\{0,1/4,1/2,3/4,1\\\}\.\(266\)The design, ordinary volume law, unweighted subset OLS estimator, and each budget are fixed along a trajectory; only the residual follows the displayed geodesic\. Each dimension\-four cell is an exact determinant sum over the ordinary\-volume subsets\. Each larger cell is a Monte Carlo plug\-in from eight fixed seeded batches of 2,048 draws\. Its symmetric delete\-one\-batch jackknife endpoints diagnose Monte Carlo variation only: they are neither confidence intervals nor exact inequality certificates\.
For the sampled cells, each ordinary\-volume draw samples a rank\-ddprojection\-DPP basis from thin\-QR coordinates and then uniformly pads it withs−ds\-ddistinct rows from the complement\. PCG64DXSM used eight fixed streams per sampled cell\. The ancillary machine\-readable seed manifest and provenance record give the exact 64\-bit integer supplied to the generator for each of the 320 streams, indexed by\(d,r,λ,j\)\(d,r,\\lambda,j\); the 15 dimension\-four cells are exact enumerations and use no random stream\. The records use neutral public aliases; the listed integers, rather than those aliases, define the retained streams\. The retained calculation used CPython 3\.13\.5 and NumPy 2\.2\.6 on Linux 5\.15 x86\_64 with glibc 2\.35\. An emitted subset whose numerical rank fails NumPy’s default relative SVD criterion aborts the calculation; it is never discarded and redrawn\. Together with the ancillary seed manifest and full cell record, the construction, sampler description, and software environment document this finite illustration\. Under the stated NumPy version, the manifest supports replay of the raw PCG64DXSM streams and independent reimplementation; it is not a claim of bitwise end\-to\-end numerical reproduction from the manuscript alone\. No executable code or generated logs are released with this preprint\.
Table[E2](https://arxiv.org/html/2608.26877#A9.T2)summarizes the ancillary 55\-cell record by trajectory\. “Increases” counts the four successive steps on the five\-point grid; it makes no claim between grid points\. For exact rows,qsq\_\{s\}andΔ=1−qs\\Delta=1\-q\_\{s\}are enumerated quantities\. For sampled rows, the table usesq^s\\widehat\{q\}\_\{s\}andΔ^=1−q^s\\widehat\{\\Delta\}=1\-\\widehat\{q\}\_\{s\}, and the final gap column is correspondinglybs−qsb\_\{s\}\-q\_\{s\}orbs−q^sb\_\{s\}\-\\widehat\{q\}\_\{s\}\. Every sampled cell’s upper Monte Carlo diagnostic endpoint forq^s\\widehat\{q\}\_\{s\}remained belowbsb\_\{s\}at the reported computation resolution; the smallest such difference was0\.005550\.00555\. The ancillary record reports every one of the 55 displayed\-grid cells\.
Table E2:Trajectory\-level summary of the 55\-cell controlled residual\-localization record\. The ranges cover all five displayed path points\. The maximum half\-width is a batch\-jackknife Monte Carlo diagnostic; exact rows have no error bar\.This targeted finite path does not establish prevalence, typicality, tightness, a population effect, or a result for the response\-uniform feature certificate\. It does not revise the separate generic\-direction failure below\.
#### I\.3\.2Companion fixed\-design response\-direction diagnostic
This companion calculation uses one fixed synthetic design at each listed dimension\. At each of three listed subset sizes, the reported value is the range ofτs\\tau\_\{s\}across four specified generic residual directions\. For this descriptive screen, a budget is labeled as qualifying when this range is at least0\.050\.05\. No listed budget meets that criterion\. Four additional targeted stress directions are excluded by definition and do not enter the count\. This is a finite configuration\-level diagnostic failure, not a falsification of any theorem\.
Table E3:Fixed\-design generic response\-direction diagnostic\. A qualifying budget requires a reported range of at least0\.050\.05\.
## Appendix JFixed\-pool diagnostic record
This appendix gives the complete record for the three complete indexed fixed pools in Section[I\.2](https://arxiv.org/html/2608.26877#A9.SS2)\. Every representation has a literal intercept\. All non\-intercept encoded columns were centered and population\-RMS scaled once on the complete pool, never separately within a subset; responses were untransformed\. The pools use fixed full\-rank representations and five fixed budgets each: Airfoil Self\-Noise\(m,d\)=\(1503,6\)\(m,d\)=\(1503,6\)withs∈\{156,380,755,1129,1353\}s\\in\\\{156,380,755,1129,1353\\\}, Concrete Compressive Strength\(m,d\)=\(1030,9\)\(m,d\)=\(1030,9\)withs∈\{111,264,520,775,928\}s\\in\\\{111,264,520,775,928\\\}, and Energy Efficiency with heating load in its fixed encoded 14\-column representation,\(m,d\)=\(768,14\)\(m,d\)=\(768,14\)withs∈\{89,203,391,580,693\}s\\in\\\{89,203,391,580,693\\\}\. For Airfoil, the five predictors are frequency, angle of attack, chord length, free\-stream velocity, and suction\-side displacement thickness; the response is scaled sound pressure\. For Concrete, the eight predictors are cement, blast\-furnace slag, fly ash, water, superplasticizer, coarse aggregate, fine aggregate, and age; the response is compressive strength, and physical indexed duplicate rows are retained\. For Energy, the 14 columns are the intercept, five continuous predictors, threeX6X6indicators \(reference level 2\), and fiveX8X8indicators \(reference level 0\)\. The design omitsX2X2under the exact relationX2=X3\+2X4X2=X3\+2X4;Y1Y1\(Heating Load\) is the response andY2Y2\(Cooling Load\) is excluded\.
For the retained plug\-in construction, let the reduced QR factorization beX=QqrRqrX=Q\_\{\\mathrm\{qr\}\}R\_\{\\mathrm\{qr\}\}, with positive diagonal inRqrR\_\{\\mathrm\{qr\}\}, so thatRqr⊤Rqr=GR\_\{\\mathrm\{qr\}\}^\{\\top\}R\_\{\\mathrm\{qr\}\}=G\. For drawℓ\\ellin batchjjof 2,048 ordinary\-volume draws, defineδs,j,ℓ=Rqr\(wSj,ℓ−w∗\)\\delta\_\{s,j,\\ell\}=R\_\{\\mathrm\{qr\}\}\(w\_\{S\_\{j,\\ell\}\}\-w^\{\*\}\)\. This is the retained QR\-whitened coefficient error, an orthogonal\-coordinate version ofG1/2G^\{1/2\}whitening consistent with the normalized directional metric\. The within\-batch second\-moment matrix and aggregate statistic are
Ms,j=12,048∑ℓ=12,048δs,j,ℓδs,j,ℓ⊤,M¯^s=116∑j=116Ms,j,q^s=λmax\(M¯^s\)αL∗\.M\_\{s,j\}=\\frac\{1\}\{2\{,\}048\}\\sum\_\{\\ell=1\}^\{2\{,\}048\}\\delta\_\{s,j,\\ell\}\\delta\_\{s,j,\\ell\}^\{\\top\},\\qquad\\widehat\{\\overline\{M\}\}\_\{s\}=\\frac\{1\}\{16\}\\sum\_\{j=1\}^\{16\}M\_\{s,j\},\\qquad\\widehat\{q\}\_\{s\}=\\frac\{\\lambda\_\{\\max\}\(\\widehat\{\\overline\{M\}\}\_\{s\}\)\}\{\\alpha L^\{\*\}\}\.\(267\)
For every cell,bsb\_\{s\},τs\\tau\_\{s\}, andq^s\\widehat\{q\}\_\{s\}have the operational definitions in \([265](https://arxiv.org/html/2608.26877#A9.E265)\)–\([267](https://arxiv.org/html/2608.26877#A10.E267)\)\. In particular,bsb\_\{s\}is the arbitrary\-full\-rank, boundary\-valid one\-sided response\-aware resolvent ceiling,τs=1−bs\\tau\_\{s\}=1\-b\_\{s\}is its deterministic contraction from the normalized universal ceiling of one, andq^s\\widehat\{q\}\_\{s\}is a retained Monte\-Carlo plug\-in covariance statistic\. The ratiosq^s/bs\\widehat\{q\}\_\{s\}/b\_\{s\}andbs/q^sb\_\{s\}/\\widehat\{q\}\_\{s\}are descriptive only and should not be interpreted as exactqsq\_\{s\}values, confidence intervals, inferential quantities, or observed covariance reductions\. We reportbs/q^sb\_\{s\}/\\widehat\{q\}\_\{s\}to make the 6\.52–8\.82\-fold slack explicit\. Because public row general position was not exhaustively certified, this record does not use the interior exact identity\.
Each cell uses a fixed schedule of 16 batches of 2,048 subset draws \(32,768 draws total\) and 4,096 whole\-batch resamples\. The endpoint column is the 0\.5%–99\.5% linear\-percentile range from those resamples after recomputing the top eigenvalue\. It diagnoses Monte\-Carlo computation only: it is neither a confidence interval nor inference for a population or dataset\-sampling quantity\.
##### Sampling and numerical details\.
Each ordinary\-volume draw samples a rank\-ddprojection\-DPP basis from reduced QR coordinates and then uniformly pads it withs−ds\-ddistinct complement rows\. Each row is indexed by public pool, budget, and subset\-size aliases in the ancillary cell record\. PCG64DXSM used 16 fixed sampling streams and one fixed whole\-batch\-resampling stream per cell\. The ancillary machine\-readable seed manifest and provenance record give the exact 128\-bit integer supplied to the generator for each of the 255 streams, indexed by pool, subset size, batch, and stream purpose\. The listed integers, rather than the aliases, define the retained streams used for the reported calculations\. The retained environment was CPython 3\.13\.5, NumPy 2\.2\.6, and mpmath 1\.3\.0 on Linux 5\.15 x86\_64 with glibc 2\.35\. Before public\-pool materialization, exact rational toy instances checked equality of the determinant law, complete basis\-plus\-padding path aggregation, and an independently implemented reverse\-deletion law; fixed PCG64DXSM streams then smoke\-checked production sampler frequencies\. The full pool is checked with NumPy’s default relative SVD rank criterion, and every sampled basis is checked against the exact encoded row data\. Selected OLS first uses unweighted binary64 Cholesky; a failed factor or backward\-error check is replayed on the same subset at 160 and then 256 bits from the exact encodings\. A support or numerical failure terminates the calculation rather than discarding or replacing a draw\. Together with the fixed preprocessing, ancillary seed manifest, and full cell record, these details document the sampling and numerical protocol for the public\-pool diagnostic\. Under the stated NumPy version, the manifest supports replay of the raw PCG64DXSM streams and independent reimplementation; it is not a claim of bitwise end\-to\-end numerical reproduction from the manuscript alone\. No executable code or generated logs are released with this preprint\.
Table 1:Public fixed\-pool response\-aware summary\. Hereτs=1−bs\\tau\_\{s\}=1\-b\_\{s\}is the deterministic universal\-to\-response\-aware ceiling contraction, whileq^s/bs\\widehat\{q\}\_\{s\}/b\_\{s\}is descriptive Monte\-Carlo plug\-in occupancy only\. Across all 15 cells,q^s/bs=0\.1133\\widehat\{q\}\_\{s\}/b\_\{s\}=0\.1133–0\.15340\.1534, equivalentlybs/q^s=6\.52b\_\{s\}/\\widehat\{q\}\_\{s\}=6\.52–8\.828\.82\.Table 2:Complete fixed\-pool diagnostic record\. The final column gives Monte\-Carlo\-only 0\.5%–99\.5% whole\-batch\-resampling endpoints forq^s\\widehat\{q\}\_\{s\}; it is not inferential uncertainty\.For reference, all three poolwise medians in Table[1](https://arxiv.org/html/2608.26877#A10.T1)are greater than0\.100\.10:
mediansτs\>0\.10\.\\operatorname\{median\}\_\{s\}\\tau\_\{s\}\>0\.10\.This cutoff is an uncalibrated display convention; the comparison is neither hypothesis testing nor theorem validation\. The negative companion diagnostic in Appendix[I\.3\.2](https://arxiv.org/html/2608.26877#A9.SS3.SSS2)remains unchanged by this complete\-pool record\.Similar Articles
From Perturbation Correction to Geometry-Aware Sampling: Sharpness-Guided Equilibrium Sampling for Balanced Flat Minima in Long-Tailed Learning
Introduces Sharpness-Guided Equilibrium Sampling (SGS) that dynamically adjusts sampling probabilities using sharpness estimates to achieve balanced flat minima in long-tailed learning, achieving significant gains on CIFAR-100 LT and ImageNet-LT.
Exact Limits of Random Projections for Preserving Geometry: Distance Recovery, Nearest-Neighbor Rankings, and Covariance Shape in Gaussian Models
This paper shows that the Johnson–Lindenstrauss lemma can be uninformative about retained geometry in high dimensions and derives exact limits for features like distance recovery, nearest-neighbor rankings, and covariance shape in Gaussian models.
Equivariant Covariance Tensors: Guaranteed SPD Uncertainty for Tensor-Valued Geometric Learning
This paper introduces a framework for E3-equivariant uncertainty quantification in tensor-valued geometric learning, ensuring symmetric positive-definite covariances via matrix exponentiation and proposing a robust Log-Euclidean Equivariant Scoring Objective.
Restricted Eigenvalues Beyond Gaussian Width: Threshold Occupancy under Heavy Tails
The paper addresses restricted eigenvalue bounds for heavy-tailed designs, showing that the Gaussian width law fails and introducing threshold occupancy as a critical factor. It derives sharp sample complexity results under isotropic heavy-tailed distributions.
Feature Geometry of LoRA Adapters: A Sparse Autoencoder Analysis of Representational Divergence in Fine-Tuned Language Models
This paper uses Sparse Autoencoders to analyze the geometry of LoRA-induced representations in language models, finding that LoRA updates occupy partially distinct feature structures not fully captured by pretrained interpretability dictionaries.