Low-Interaction-Rank Learning: Unifying Multiplicative Dual-Encoder Heads
Summary
This paper introduces low interaction rank as a unified theoretical framework for multiplicative dual-encoder networks, covering approximation, sample complexity, normalization, and identifiability, with experiments on operator learning and CLIP models.
View Cached Full Text
Cached at: 08/13/26, 03:37 PM
# Low-Interaction-Rank Learning: Unifying Multiplicative Dual-Encoder Heads
Source: [https://arxiv.org/html/2608.11661](https://arxiv.org/html/2608.11661)
Sen LiThanks:Corresponding Author: Sen LiAffiliation:The Hong Kong University of Science and TechnologyAffiliation:The Hong Kong University of Science and Technology \(Guangzhou\)
###### Abstract
A multiplicative dual\-encoder network computes a real\-valued output for a pair of inputs as the inner product of their separate encodings\. This architecture has been developed independently in operator learning, bipartite matching, contrastive vision\-language models, retrieval, and other areas, yet no unified theory guides the basic design decisions: how many interaction modes to represent, how to normalize the encoders, and when the architecture should be avoided\. We provide such a foundation by introducing the class of functions of low interaction rank, a class whose intrinsic complexity is measured by its interaction spectrum\. Within this framework, approximation error decomposes into a spectral truncation term and an encoder\-realization term; sample complexity is governed by the sum of the two encoder complexities rather than their product; and a usability criterion based on spectral decay determines when the architecture can succeed\. The same framework exposes a central identifiability problem: the encoders are defined only up to a linear gauge symmetry that leaves the learned coordinates arbitrary\. We show that normalization is gauge fixing and that whitening pins the interaction modes up to permutation and sign, thereby explaining the uninterpretability of contrastive dimensions and providing a constructive remedy\. Experiments on synthetic kernels, operator learning, and CLIP models validate the theoretical predictions: spectral decay rates match the predicted scaling, whitening recovers the true modes, and independently trained CLIP models are related by a single rotation which, after removal by whitening, exposes interpretable concept axes\. The code of this paper is provided at[https://github\.com/RS2002/Mul\-Net](https://github.com/RS2002/Mul-Net)\.
## 1Introduction
A common architectural pattern has emerged independently across several machine learning communities: computing a real\-valued output for a pair of inputs as the inner product of their separately encoded representations\. Contrastive models use it for cross\-modal similarity\([17](https://arxiv.org/html/2608.11661#bib.bib1);[27](https://arxiv.org/html/2608.11661#bib.bib15)\); operator networks evaluate learned maps via branch\-trunk inner products\([14](https://arxiv.org/html/2608.11661#bib.bib2)\); retrieval systems rank queries against documents with two\-tower encoders\([11](https://arxiv.org/html/2608.11661#bib.bib3)\); goal\-conditioned reinforcement learning and successor features represent values as inner products of encodings\([21](https://arxiv.org/html/2608.11661#bib.bib4);[9](https://arxiv.org/html/2608.11661#bib.bib5);[3](https://arxiv.org/html/2608.11661#bib.bib6)\); and knowledge\-graph completion, linear attention, factorization machines, and multiplicative matching networks all share the same underlying structure\([24](https://arxiv.org/html/2608.11661#bib.bib7);[23](https://arxiv.org/html/2608.11661#bib.bib8);[12](https://arxiv.org/html/2608.11661#bib.bib9);[19](https://arxiv.org/html/2608.11661#bib.bib10);[30](https://arxiv.org/html/2608.11661#bib.bib13);[29](https://arxiv.org/html/2608.11661#bib.bib14)\)\.
In each domain, this structure provides tangible computational and representational benefits\. Yet the corresponding methods have been developed largely in isolation, even though they are all instances of the same fundamental architecture, differing only in the nature of the two inputs and how the output is consumed\. Consequently, each community independently confronts the same design questions: how many interaction modes to retain, how to normalize the encoders to ensure stable training and meaningful representations, and when the architecture is fundamentally ill\-suited\.
We now formalize this shared architecture: a multiplicative dual\-encoder head computes
F\(u,v\)≈⟨fθ\(u\),gφ\(v\)⟩=∑k=1dfk\(u\)gk\(v\)\\boxed\{F\(u,v\)\\approx\\langle f\_\{\\theta\}\(u\),g\_\{\\varphi\}\(v\)\\rangle=\\sum\_\{k=1\}^\{d\}f\_\{k\}\(u\)\\,g\_\{k\}\(v\)\}\(1\)wherefθf\_\{\\theta\}andgφg\_\{\\varphi\}are learned encoders mapping the two inputs into a common spaceℝd\\mathbb\{R\}^\{d\}\. We study the class
ℳd=\{F:F\(u,v\)=⟨f\(u\),g\(v\)⟩,f:U→ℝd,g:V→ℝd\}\\mathcal\{M\}\_\{d\}=\\\{\\,F:\\,F\(u,v\)=\\langle f\(u\),g\(v\)\\rangle,\\;f:U\\to\\mathbb\{R\}^\{d\},\\;g:V\\to\\mathbb\{R\}^\{d\}\\,\\\}\(2\)of functions representable by a rank\-ddhead\. The complexity of a targetFFwithin this class is measured by its interaction spectrum\{σk\}k≥1\\\{\\sigma\_\{k\}\\\}\_\{k\\geq 1\}, the singular values of the integral operator whose kernel isFF; its interaction rank is the number of nonzero singular values\.
Within this framework, we answer four questions about approximation error, encoder identifiability, sample complexity, and fundamental architectural limitations\.
- •Approximation:The error of any rank\-ddhead decomposes into a spectral truncation term, unavoidable by any encoder, and an encoder\-realization term; target smoothness controls the decay rate\. \(Section[2](https://arxiv.org/html/2608.11661#S2)\)
- •Identifiability:The representation \([1](https://arxiv.org/html/2608.11661#S1.E1)\) is invariant under a linear gauge symmetry acting on the pair of encoders, so the encoders of any trained model are defined only up to this symmetry\. We show that the normalizations used across these domains are precisely gauge\-fixing choices, and we prove that whitening pins the interaction modes up to permutation and sign\. This explains why the individual dimensions of contrastive models lack inherent semantic meaning and provides a constructive remedy\. \(Section[3](https://arxiv.org/html/2608.11661#S3)\)
- •Estimation:Sample complexity is governed by the sum of the two encoder complexities rather than their product; for smooth targets, the optimal rank grows only logarithmically in the sample budget\. \(Section[4](https://arxiv.org/html/2608.11661#S4)\)
- •Usability:A flat interaction spectrum forces every rank\-ddhead to a relative error floor of1−d/N1\-d/N, whereNNis the size of the discrete input domain; this barrier is escaped by early\-interaction networks and yields a practical criterion based on the measured spectrum\. \(Section[4](https://arxiv.org/html/2608.11661#S4)\)
To further validate our conclusions, we conducted experiments on synthetic kernels, operator learning, and CLIP models, confirming the predicted spectral behavior, demonstrating unique mode recovery under whitening, and verifying the flat\-spectrum error floor\. \(Section[5](https://arxiv.org/html/2608.11661#S5)\)
## 2The Low\-Interaction\-Rank Class: Unification and Approximation
Section[1](https://arxiv.org/html/2608.11661#S1)introduced the classℳd\\mathcal\{M\}\_\{d\}of rank\-ddheads\. We now associate with every target a spectral object that measures its intrinsic complexity within this class, show that ten existing method families are instances of the same architecture, and develop an approximation theory that quantifies how well a rank\-ddhead can represent a given target\.
### 2\.1The class and the interaction spectrum
Let𝒰\\mathcal\{U\}and𝒱\\mathcal\{V\}be compact metric spaces with probability measuresμU\\mu\_\{U\},μV\\mu\_\{V\}, and letF∈L2\(μU⊗μV\)F\\in L^\{2\}\(\\mu\_\{U\}\\otimes\\mu\_\{V\}\)be the target function of two arguments\. A rank\-ddhead is a pair of encodersf:𝒰→ℝdf:\\mathcal\{U\}\\to\\mathbb\{R\}^\{d\},g:𝒱→ℝdg:\\mathcal\{V\}\\to\\mathbb\{R\}^\{d\}with coordinate functionsfk∈L2\(μU\)f\_\{k\}\\in L^\{2\}\(\\mu\_\{U\}\),gk∈L2\(μV\)g\_\{k\}\\in L^\{2\}\(\\mu\_\{V\}\)\.
###### Definition 1\(Interaction spectrum\)\.
ForF∈L2\(μU⊗μV\)F\\in L^\{2\}\(\\mu\_\{U\}\\otimes\\mu\_\{V\}\), the*interaction operator*is the integral operatorTF:L2\(μV\)→L2\(μU\)T\_\{F\}:L^\{2\}\(\\mu\_\{V\}\)\\to L^\{2\}\(\\mu\_\{U\}\)defined by
\(TFh\)\(u\)=∫𝒱F\(u,v\)h\(v\)dμV\(v\)\.\(T\_\{F\}h\)\(u\)=\\int\_\{\\mathcal\{V\}\}F\(u,v\)\\,h\(v\)\\,d\\mu\_\{V\}\(v\)\.\(3\)SinceF∈L2\(μU⊗μV\)F\\in L^\{2\}\(\\mu\_\{U\}\\otimes\\mu\_\{V\}\), this operator is Hilbert\-Schmidt, andFFadmits the Schmidt decomposition
F=∑k≥1σkak⊗bk,\(ak⊗bk\)\(u,v\)=ak\(u\)bk\(v\),F=\\sum\_\{k\\geq 1\}\\sigma\_\{k\}\\,a\_\{k\}\\otimes b\_\{k\},\\qquad\(a\_\{k\}\\otimes b\_\{k\}\)\(u,v\)=a\_\{k\}\(u\)\\,b\_\{k\}\(v\),\(4\)with singular valuesσ1≥σ2≥⋯≥0\\sigma\_\{1\}\\geq\\sigma\_\{2\}\\geq\\cdots\\geq 0and orthonormal systems\{ak\}⊂L2\(μU\)\\\{a\_\{k\}\\\}\\subset L^\{2\}\(\\mu\_\{U\}\),\{bk\}⊂L2\(μV\)\\\{b\_\{k\}\\\}\\subset L^\{2\}\(\\mu\_\{V\}\)\. The sequence\{σk\}\\\{\\sigma\_\{k\}\\\}is the*interaction spectrum*ofFF; the numberi\-rank\(F\)=\#\{k:σk\>0\}\\mathrm\{i\}\\text\{\-\}\\mathrm\{rank\}\(F\)=\\\#\\\{k:\\sigma\_\{k\}\>0\\\}is its*interaction rank*; and the functionsaka\_\{k\},bkb\_\{k\}are its*interaction modes*\. The class introduced in Section[1](https://arxiv.org/html/2608.11661#S1)satisfiesℳd=\{F:i\-rank\(F\)≤d\}\\mathcal\{M\}\_\{d\}=\\\{F:\\mathrm\{i\}\\text\{\-\}\\mathrm\{rank\}\(F\)\\leq d\\\}, and the spectrum depends on the reference measures\.
By the isometry betweenL2\(μU⊗μV\)L^\{2\}\(\\mu\_\{U\}\\otimes\\mu\_\{V\}\)and the Hilbert\-Schmidt class, a function is representable by a rank\-ddhead if and only ifrank\(TS\)≤d\\rank\(T\_\{S\}\)\\leq d; hence the parametric class \([2](https://arxiv.org/html/2608.11661#S1.E2)\) coincides withℳd\\mathcal\{M\}\_\{d\}\(Lemma[1](https://arxiv.org/html/2608.11661#Thmlemma1), Appendix[B](https://arxiv.org/html/2608.11661#A2)\)\. The spectral truncation term∑k\>dσk2\\sum\_\{k\>d\}\\sigma\_\{k\}^\{2\}depends only on the target and the reference measures, not on the encoders: no head, regardless of its capacity, can representFFwith error below this quantity\. The bound is attained inL2L^\{2\}by the truncated Schmidt series, and is unique whenσd\>σd\+1\\sigma\_\{d\}\>\\sigma\_\{d\+1\}\(Corollary[1](https://arxiv.org/html/2608.11661#Thmcorollary1), Appendix[B](https://arxiv.org/html/2608.11661#A2)\)\. The decay rate of this tail, which is controlled by the smoothness ofFF\(Theorem[3](https://arxiv.org/html/2608.11661#Thmtheorem3)\), determines how many interaction modesdda good head must retain\.
This is the operator\-theoretic analogue of classical low\-rank matrix approximation: the Schmidt decomposition is the infinite\-dimensional Eckart\-Young theorem\([22](https://arxiv.org/html/2608.11661#bib.bib28);[7](https://arxiv.org/html/2608.11661#bib.bib29)\), and the truncation term is exactly the Hilbert\-Schmidt error of the best rank\-ddapproximation ofTFT\_\{F\}\. The spectrum is relative to the reference measuresμU,μV\\mu\_\{U\},\\mu\_\{V\}, which fix the inner product in which the modes are orthonormal; the same target may have different spectra under different measures\. All subsequent claims therefore hold for the measures under which a head is trained and evaluated\.
### 2\.2Ten families as instances of the head
Appendix[A](https://arxiv.org/html/2608.11661#A1)collects ten method families from different communities in Table[5](https://arxiv.org/html/2608.11661#A1.T5)\. Each computes a real\-valued output for a pair of inputs as the inner product of two separately learned encodings\. The questions of how many interaction modesddto keep, how to normalize the two encoders \(Section[3](https://arxiv.org/html/2608.11661#S3)\), and how many samples the head requires \(Section[4](https://arxiv.org/html/2608.11661#S4)\) are answered independently by each community; within the present framework, they are the same questions\.
### 2\.3Approximation and separation from single\-tower models
We now study how wellℳd\\mathcal\{M\}\_\{d\}approximates a fixed target\. Corollary[1](https://arxiv.org/html/2608.11661#Thmcorollary1)isolates the part of the error that no encoder can avoid\. The remaining part is attributable to the encoders themselves\. We assume that the encoder classes contain functions approaching the scaled modesσkak\\sigma\_\{k\}a\_\{k\}andbkb\_\{k\}with errorsηf,k,ηg,k\\eta\_\{f,k\},\\eta\_\{g,k\}as in \([20](https://arxiv.org/html/2608.11661#A2.E20)\) \(Assumption[1](https://arxiv.org/html/2608.11661#Thmassumption1), Appendix[B](https://arxiv.org/html/2608.11661#A2)\); for neural encoders, these errors follow standard ReLU approximation rates\([25](https://arxiv.org/html/2608.11661#bib.bib21)\)\.
###### Theorem 1\(Approximation error decomposition\)\.
LetF∈L2\(μU⊗μV\)F\\in L^\{2\}\(\\mu\_\{U\}\\otimes\\mu\_\{V\}\)have interaction spectrum\{σk\}\\\{\\sigma\_\{k\}\\\}\. Under Assumption[1](https://arxiv.org/html/2608.11661#Thmassumption1), withBf=maxk‖f^k‖L2\(μU\)B\_\{f\}=\\max\_\{k\}\\\|\\hat\{f\}\_\{k\}\\\|\_\{L^\{2\}\(\\mu\_\{U\}\)\}andBg=maxk‖g^k‖L2\(μV\)B\_\{g\}=\\max\_\{k\}\\\|\\hat\{g\}\_\{k\}\\\|\_\{L^\{2\}\(\\mu\_\{V\}\)\},
inff∈ℱg∈𝒢‖F−⟨f,g⟩‖L2\(μU⊗μV\)2≤∑k\>dσk2⏟truncation term\+C∑k≤d\(σk2ηg,k2\+Bg2ηf,k2\)⏟realization term,\\inf\_\{\\begin\{subarray\}\{c\}f\\in\\mathcal\{F\}\\\\ g\\in\\mathcal\{G\}\\end\{subarray\}\}\\\|F\-\\langle f,g\\rangle\\\|\_\{L^\{2\}\(\\mu\_\{U\}\\otimes\\mu\_\{V\}\)\}^\{2\}\\;\\leq\\;\\underbrace\{\\sum\_\{k\>d\}\\sigma\_\{k\}^\{2\}\}\_\{\\text\{truncation term\}\}\\;\+\\;\\underbrace\{C\\sum\_\{k\\leq d\}\\bigl\(\\sigma\_\{k\}^\{2\}\\,\\eta\_\{g,k\}^\{2\}\+B\_\{g\}^\{2\}\\,\\eta\_\{f,k\}^\{2\}\\bigr\)\}\_\{\\text\{realization term\}\},\(5\)whereCCis an absolute constant\.
###### Proof idea\.
The error is split around the truncated Schmidt series, and the tensor\-product product rule is applied term by term\. Orthonormality of the ideal modes diagonalizes the two leading sums and avoids a factor ofdd; the remaining term is bounded by Cauchy\-Schwarz\. The full calculation, including the explicit constant, is deferred to Appendix[B](https://arxiv.org/html/2608.11661#A2)\. ∎
The decomposition is tight up to the constant: the truncation term is intrinsic to the target, while the realization term depends on the model\. Dense encoder classes therefore approach the lower bound of Corollary[1](https://arxiv.org/html/2608.11661#Thmcorollary1)\. Whether training actually finds the coordinate alignment required by \([20](https://arxiv.org/html/2608.11661#A2.E20)\) is precisely the identifiability question addressed in Section[3](https://arxiv.org/html/2608.11661#S3)\.
Target smoothness controls the truncation term: anss\-smooth target has spectrumσk=O\(k−s/mV\)\\sigma\_\{k\}=O\(k^\{\-s/m\_\{V\}\}\), while real\-analytic targets decay exponentially\. The precise tail bounds are given in Theorem[3](https://arxiv.org/html/2608.11661#Thmtheorem3)\(Appendix[B](https://arxiv.org/html/2608.11661#A2)\)\.
The dual\-encoder head separates from single\-tower models on both the approximation and the computational axes\. On the approximation axis, it avoids the curse of the joint dimension: approximating a rank\-ddtarget withinε\\varepsilonwith ReLU encoders requires
Ndual=O\(d\(ε−mU/s\+ε−mV/s\)\)N\_\{\\mathrm\{dual\}\}=O\\bigl\(d\(\\varepsilon^\{\-m\_\{U\}/s\}\+\\varepsilon^\{\-m\_\{V\}/s\}\)\\bigr\)\(6\)parameters, whereas any single\-tower model that treatsFFas a general\(mU\+mV\)\(m\_\{U\}\+m\_\{V\}\)\-dimensionalHsH^\{s\}function requiresΩ\(ε−\(mU\+mV\)/s\)\\Omega\(\\varepsilon^\{\-\(m\_\{U\}\+m\_\{V\}\)/s\}\)parameters in the worst case; henceNdual/Nsingle→0N\_\{\\mathrm\{dual\}\}/N\_\{\\mathrm\{single\}\}\\to 0asε→0\\varepsilon\\to 0\(Proposition[2](https://arxiv.org/html/2608.11661#Thmproposition2)\)\. On the compute axis, evaluating the head on all pairs costsO\(\(n\+m\)Cnet\+nmd\)O\(\(n\+m\)\\,C\_\{\\mathrm\{net\}\}\+nm\\,d\)against quadratic joint processing, recovering the complexity arguments of QK\-attention matching\([29](https://arxiv.org/html/2608.11661#bib.bib14)\), linear attention\([12](https://arxiv.org/html/2608.11661#bib.bib9)\), and DeepONet\([14](https://arxiv.org/html/2608.11661#bib.bib2)\)\(Proposition[1](https://arxiv.org/html/2608.11661#Thmproposition1), Appendix[B](https://arxiv.org/html/2608.11661#A2)\)\.
## 3Identifiability as Gauge\-Fixing
The theory of Section[2](https://arxiv.org/html/2608.11661#S2)characterizes what a rank\-ddhead can represent, but it does not specify which representation a training algorithm will actually find\. The representation is far from unique: for any invertible matrixAA, the pair\(Af,A−⊤g\)\(Af,\\,A^\{\-\\top\}g\)represents exactly the same function\. Consequently, the encoders of any trained model are defined only up to a symmetry of dimensiond2d^\{2\}\. This redundancy is not purely formal\. It renders the population risk constant along large orbits, so the Hessian at a minimizer is highly degenerate and optimization becomes ill\-posed\. At the same time, it leaves the learned coordinates essentially arbitrary\. Normalizations were introduced precisely to control these instabilities\([17](https://arxiv.org/html/2608.11661#bib.bib1);[30](https://arxiv.org/html/2608.11661#bib.bib13)\); a normalization makes the encoders meaningful exactly to the extent that it removes the symmetry\.
We now formalize this view\. Each normalization is a gauge\-fixing choice, a rule for selecting one representative from each equivalence class\. The residual symmetry that survives a constraint is precisely what remains arbitrary under it\. The interaction spectrum constitutes the invariant content of the symmetry, and the gaps in that spectrum control how much identifiability any normalization can provide\.
###### Definition 2\(Gauge action\)\.
ForA∈GLdA\\in\\mathrm\{GL\}\_\{d\}, the gauge transformation acts on the encoders by
ΦA\(f,g\)=\(Af,A−⊤g\),\\Phi\_\{A\}\(f,g\)=\\bigl\(Af,\\;A^\{\-\\top\}g\\bigr\),\(7\)where\(Af\)\(u\)=Af\(u\)\(Af\)\(u\)=A\\,f\(u\)and\(A−⊤g\)\(v\)=A−⊤g\(v\)\\bigl\(A^\{\-\\top\}g\\bigr\)\(v\)=A^\{\-\\top\}g\(v\)\. The set of all such transformations is the gauge group of the head\.
Gauge transformations leave the represented function, and therefore the population risk, unchanged:
⟨Af\(u\),A−⊤g\(v\)⟩=⟨f\(u\),g\(v\)⟩\.\\langle Af\(u\),A^\{\-\\top\}g\(v\)\\rangle=\\langle f\(u\),g\(v\)\\rangle\.\(8\)The second momentsΣf=𝔼u\[ff⊤\]\\Sigma\_\{f\}=\\mathbb\{E\}\_\{u\}\[ff^\{\\top\}\]andΣg=𝔼v\[gg⊤\]\\Sigma\_\{g\}=\\mathbb\{E\}\_\{v\}\[gg^\{\\top\}\]transform by congruence, so the eigenvalues of the productΣfΣg\\Sigma\_\{f\}\\Sigma\_\{g\}are gauge invariants \(Lemma[2](https://arxiv.org/html/2608.11661#Thmlemma2), Appendix[C](https://arxiv.org/html/2608.11661#A3)\)\. Unfixed gauge also makes optimization ill\-posed\. Along the orbit of a minimizer the risk is flat, so the Hessian has at leastd2−dimSd^\{2\}\-\\dim Szero directions; without normalization this is genericallyd2d^\{2\}\(Proposition[4](https://arxiv.org/html/2608.11661#Thmproposition4)\)\. A normalization pins exactly the subgroup whose residual freedom it eliminates; a smaller residual group therefore leaves fewer degenerate directions and yields a better\-conditioned problem\.
#### Normalizations as sections of the orbit\.
Two levels of normalization occur in practice, and only one affects identifiability\. Embedding\-level normalization constrains the encoders themselves and determines the residual gauge\. Output\-level normalization, such as the contrastive softmax in CLIP or top\-kkselection in retrieval, acts on the score matrix after the inner product; it does not constrain the encoders and leaves the embedding\-level gauge untouched\. Table[6](https://arxiv.org/html/2608.11661#A3.T6)in Appendix[C](https://arxiv.org/html/2608.11661#A3)classifies the embedding\-level schemes used in practice\.
#### What each normalization identifies\.
The residual symmetry of a normalization indicates which transformations can act on a trained model without changing its output\. The normalizations used in practice were adopted on heuristic grounds, for training stability or empirical feature decorrelation, rather than from an analysis of which symmetry they remove; the gauge view supplies the missing analysis and lets us rank them by residual symmetry\. We characterize the two\-sided schemes used in practice, in order of increasing identifiability, and find that only whitening removes the symmetry completely\.
Cosine normalization is the weakest two\-sided scheme\([17](https://arxiv.org/html/2608.11661#bib.bib1)\)\. Two\-sided unit normalizationf~=f/‖f‖\\tilde\{f\}=f/\\\|f\\\|,g~=g/‖g‖\\tilde\{g\}=g/\\\|g\\\|with scorescos∠\(f\(u\),g\(v\)\)\\cos\\angle\(f\(u\),g\(v\)\), as employed by contrastive models such as CLIP\([17](https://arxiv.org/html/2608.11661#bib.bib1)\), leaves a residual gauge exactly equal to the orthogonal group: for every rotationR∈OR\\in\\mathrm\{O\}, the encodingsRfRfandRgRggive pointwise identical cosine scores, and no other transformation does \(Theorem[4](https://arxiv.org/html/2608.11661#Thmtheorem4)\)\. Consequently, the individual coordinates of a cosine\-normalized embedding carry no intrinsic meaning; only rotation\-invariant quantities \(norms, pairwise inner products, and the eigenvalues ofΣfΣg\\Sigma\_\{f\}\\Sigma\_\{g\}\) are well\-defined\. This provides a theorem\-level explanation of the widely reported uninterpretability of contrastive embedding dimensions, and it identifies the whitening scheme of Theorem[2](https://arxiv.org/html/2608.11661#Thmtheorem2)as the remedy \(Corollary[2](https://arxiv.org/html/2608.11661#Thmcorollary2), Appendix[C](https://arxiv.org/html/2608.11661#A3)\)\.
Two\-sided nonnegativity,f\(u\)⪰0f\(u\)\\succeq 0andg\(v\)⪰0g\(v\)\\succeq 0\(realized by a softplus or exponential output layer\([30](https://arxiv.org/html/2608.11661#bib.bib13)\)\), shrinks the residual group to the monomial groupPerm⋉D\+\\mathrm\{Perm\}\\ltimes\\mathrm\{D\}\_\{\+\}of coordinate permutations composed with positive diagonal scalings\. Under separability conditions\([6](https://arxiv.org/html/2608.11661#bib.bib11);[1](https://arxiv.org/html/2608.11661#bib.bib12)\), the encoders are identified up to permutation and scaling\. The cost is an expressivity limitation: negatively correlated interaction modes are excluded \(Theorem[5](https://arxiv.org/html/2608.11661#Thmtheorem5), Appendix[C](https://arxiv.org/html/2608.11661#A3)\)\.
Whitening is the unique scheme among those used in practice that removes the continuous gauge symmetry entirely\. Cosine and nonnegativity each leave a nontrivial residual group and were adopted largely heuristically; by contrast, we prove that whitening admits a full identifiability guarantee: under a spectral gap it pins the interaction modes up to permutation and sign, recovering both the modes and the spectrum itself\.
###### Theorem 2\(Whitening fixes the gauge up to permutation and sign\)\.
Let the interaction singular values be distinct,σ1\>⋯\>σd\>σd\+1\\sigma\_\{1\}\>\\cdots\>\\sigma\_\{d\}\>\\sigma\_\{d\+1\}, and impose the whitening constraints
Σg=Id,Σf=Λ=diag\(λ1,…,λd\),λ1≥⋯≥λd≥0,\\Sigma\_\{g\}=I\_\{d\},\\qquad\\Sigma\_\{f\}=\\Lambda=\\operatorname\{diag\}\(\\lambda\_\{1\},\\dots,\\lambda\_\{d\}\),\\quad\\lambda\_\{1\}\\geq\\cdots\\geq\\lambda\_\{d\}\\geq 0,\(9\)withΛ\\Lambdadiagonal and nonincreasing\. Constraints of this form are standard in self\-supervised representation learning, where they are used to decorrelate features\([8](https://arxiv.org/html/2608.11661#bib.bib16);[26](https://arxiv.org/html/2608.11661#bib.bib17);[2](https://arxiv.org/html/2608.11661#bib.bib18)\)\. Then every global minimizer of the population risk satisfies
fk=±σkak,gk=±bk,λk=σk2,k=1,…,d,f\_\{k\}=\\pm\\,\\sigma\_\{k\}\\,a\_\{k\},\\qquad g\_\{k\}=\\pm\\,b\_\{k\},\\qquad\\lambda\_\{k\}=\\sigma\_\{k\}^\{2\},\\qquad k=1,\\dots,d,\(10\)with signs paired so that the productfkgk=σkakbkf\_\{k\}g\_\{k\}=\\sigma\_\{k\}a\_\{k\}b\_\{k\}is uniquely determined, and with the coordinate labeling fixed by the ordering convention in \([9](https://arxiv.org/html/2608.11661#S3.E9)\)\. Thekk\-th coordinate of the whitened encoders recovers thekk\-th interaction mode of the target, and the diagonal ofΛ\\Lambdarecovers the interaction spectrum\. Equivalently, the encoders are identified up to permutation and sign\.
###### Proof idea\.
The rank\-truncation bound of Corollary[1](https://arxiv.org/html/2608.11661#Thmcorollary1)implies that every minimizer attains the minimal risk\. Whenσd\>σd\+1\\sigma\_\{d\}\>\\sigma\_\{d\+1\}, this forces the represented function to equal the unique truncationFdF\_\{d\}\. The whitening constraints then require the coefficient matrices to satisfyCC⊤=ICC^\{\\top\}=IandCS2C⊤=ΛCS^\{2\}C^\{\\top\}=\\Lambda, whereS=diag\(σ1,…,σd\)S=\\operatorname\{diag\}\(\\sigma\_\{1\},\\dots,\\sigma\_\{d\}\)has distinct eigenvalues; this forcesCCto be a signed permutation\. The full argument is in Appendix[C](https://arxiv.org/html/2608.11661#A3)\. ∎
#### Implementing the whitening gauge\.
Hard constraints \([9](https://arxiv.org/html/2608.11661#S3.E9)\) can be enforced by minibatch whitening\. Alternatively, a quadratic penalty inherits the identifiability of the hard constraint and improves conditioning monotonically with the penalty weight \(Theorem[6](https://arxiv.org/html/2608.11661#Thmtheorem6), Appendix[C](https://arxiv.org/html/2608.11661#A3)\)\.
In practice, the second moments are estimated from a finite sample\. With probability1−δ1\-\\delta, the identified modes carry error of orderr/Δr/\\Delta, wherer=O\(B2log\(d/δ\)/n\)r=O\\bigl\(B^\{2\}\\sqrt\{\\log\(d/\\delta\)/n\}\\bigr\),BBbounds the encoder outputs, andΔ=mink\(σk2−σk\+12\)\\Delta=\\min\_\{k\}\(\\sigma\_\{k\}^\{2\}\-\\sigma\_\{k\+1\}^\{2\}\)is the spectral gap\. Achieving accuracyε\\varepsilontherefore requires
n≳B4log\(d/δ\)Δ2ε2n\\gtrsim\\frac\{B^\{4\}\\log\(d/\\delta\)\}\{\\Delta^\{2\}\\varepsilon^\{2\}\}\(11\)samples \(Appendix[E](https://arxiv.org/html/2608.11661#A5)\)\. The same gap thus governs optimization conditioning, identifiability, and estimation simultaneously \(Corollary[3](https://arxiv.org/html/2608.11661#Thmcorollary3)\)\.
Section[5](https://arxiv.org/html/2608.11661#S5)measures the gauge, conditioning, and expressivity consequences of the various schemes on controlled targets; the two real systems, branch\-trunk operator learning\([14](https://arxiv.org/html/2608.11661#bib.bib2)\)and independently trained CLIP models, are treated in Appendix[F](https://arxiv.org/html/2608.11661#A6)\(Theorem[4](https://arxiv.org/html/2608.11661#Thmtheorem4)\)\. Appendix[C](https://arxiv.org/html/2608.11661#A3)lists the measurable quantities on which the schemes are compared\.
## 4Sample Complexity and When Not to Use the Head
The interaction spectrum controls both estimation error and fundamental limitations\. The excess risk of a head fitted fromnnsamples decomposes into an approximation term that decreases as the rankddgrows, and an estimation term that increases withddthrough the complexity of the encoder classes\. Balancing these two terms yields a rank\-selection rule\. When the spectrum is flat, however, every rank\-ddhead suffers a relative\-error floor that no amount of capacity or data can overcome; this yields a practical criterion for abandoning the architecture\.
#### Setup and notation\.
We observennindependent samples\(ui,vi,yi\)\(u\_\{i\},v\_\{i\},y\_\{i\}\)withyi=F\(ui,vi\)\+εiy\_\{i\}=F\(u\_\{i\},v\_\{i\}\)\+\\varepsilon\_\{i\}and fit the head by empirical risk minimization over the class
ℋd=\{\(u,v\)↦⟨f\(u\),g\(v\)⟩:f∈ℱ,g∈𝒢\}\\mathcal\{H\}\_\{d\}=\\bigl\\\{\\,\(u,v\)\\mapsto\\langle f\(u\),g\(v\)\\rangle:\\ f\\in\\mathcal\{F\},\\ g\\in\\mathcal\{G\}\\,\\bigr\\\}\(12\)of rank\-ddheads\. The encoders are bounded,‖f\(u\)‖≤Bf\\\|f\(u\)\\\|\\leq B\_\{f\}and‖g\(v\)‖≤Bg\\\|g\(v\)\\\|\\leq B\_\{g\}, so the scores are bounded byB=BfBgB=B\_\{f\}B\_\{g\}\. Letℜn\\mathfrak\{R\}\_\{n\}denote the Rademacher complexity, extended to vector\-valued classes coordinatewise\. The class complexity collapses to the sum of the marginal complexities:
ℜn\(ℋd\)≤2d\(Bgℜn\(ℱ\)\+Bfℜn\(𝒢\)\),\\mathfrak\{R\}\_\{n\}\(\\mathcal\{H\}\_\{d\}\)\\leq 2d\\bigl\(B\_\{g\}\\,\\mathfrak\{R\}\_\{n\}\(\\mathcal\{F\}\)\+B\_\{f\}\\,\\mathfrak\{R\}\_\{n\}\(\\mathcal\{G\}\)\\bigr\),\(13\)for shared per\-coordinate classes \(Lemma[3](https://arxiv.org/html/2608.11661#Thmlemma3), Appendix[D](https://arxiv.org/html/2608.11661#A4)\)\.
The excess risk of the fitted head therefore decomposes into the approximation terms of Theorem[1](https://arxiv.org/html/2608.11661#Thmtheorem1)plus an estimation term controlled by this complexity:
‖F^−F‖L2\(μU⊗μV\)2≲∑k\>dσk2\+ℰreal\+Bd\(ℜn\(ℱ\)\+ℜn\(𝒢\)\)\+B2log\(1/δ\)n,\\\|\\hat\{F\}\-F\\\|\_\{L^\{2\}\(\\mu\_\{U\}\\otimes\\mu\_\{V\}\)\}^\{2\}\\;\\lesssim\\;\\sum\_\{k\>d\}\\sigma\_\{k\}^\{2\}\+\\mathcal\{E\}\_\{\\mathrm\{real\}\}\+B\\,d\\bigl\(\\mathfrak\{R\}\_\{n\}\(\\mathcal\{F\}\)\+\\mathfrak\{R\}\_\{n\}\(\\mathcal\{G\}\)\\bigr\)\+B^\{2\}\\sqrt\{\\frac\{\\log\(1/\\delta\)\}\{n\}\},\(14\)whereℰreal=C∑k≤d\(σk2ηg,k2\+Bg2ηf,k2\)\\mathcal\{E\}\_\{\\mathrm\{real\}\}=C\\sum\_\{k\\leq d\}\(\\sigma\_\{k\}^\{2\}\\eta\_\{g,k\}^\{2\}\+B\_\{g\}^\{2\}\\eta\_\{f,k\}^\{2\}\)and≲\\lesssimhides absolute constants, with probability1−δ1\-\\delta\(Corollary[4](https://arxiv.org/html/2608.11661#Thmcorollary4), Appendix[D](https://arxiv.org/html/2608.11661#A4)\)\.
#### Linear encoders anchor the rate; rank selection\.
With linear encodersf\(u\)=U⊤uf\(u\)=U^\{\\top\}u,g\(v\)=V⊤vg\(v\)=V^\{\\top\}von inputsu∈ℝpu\\in\\mathbb\{R\}^\{p\},v∈ℝqv\\in\\mathbb\{R\}^\{q\}, the head computesu⊤Θvu^\{\\top\}\\Theta vwithΘ=UV⊤\\Theta=UV^\{\\top\}of rank at mostdd\. Estimation reduces to low\-rank matrix sensing of the ground\-truth matrixΘ⋆\\Theta^\{\\star\}, whose minimax rate, withσε2\\sigma\_\{\\varepsilon\}^\{2\}the noise variance ofεi\\varepsilon\_\{i\}, is known exactly:
𝔼‖Θ^−Θ⋆‖F2≍σε2d\(p\+q\)/n\.\\mathbb\{E\}\\\|\\hat\{\\Theta\}\-\\Theta^\{\\star\}\\\|\_\{F\}^\{2\}\\asymp\\sigma\_\{\\varepsilon\}^\{2\}\\,d\(p\+q\)/n\.\(15\)This rate provides the baseline for the general bound \(Theorem[7](https://arxiv.org/html/2608.11661#Thmtheorem7), Appendix[D](https://arxiv.org/html/2608.11661#A4)\)\.
The rankddis the primary design choice for the head\. It faces two opposing forces: the truncation term∑k\>dσk2\\sum\_\{k\>d\}\\sigma\_\{k\}^\{2\}decreases withdd, while the estimation term grows linearly withdd\. The optimal rank balances these two contributions\. When the spectrum decays exponentially,σk2=O\(e−2ck\)\\sigma\_\{k\}^\{2\}=O\(e^\{\-2ck\}\), balancing gives
d⋆≍12clognp\+q,d^\{\\star\}\\asymp\\frac\{1\}\{2c\}\\log\\frac\{n\}\{p\+q\},\(16\)and yields a near\-parametric rate
‖F^−F‖L22≲\(p\+q\)log\(n/\(p\+q\)\)n\.\\\|\\hat\{F\}\-F\\\|\_\{L^\{2\}\}^\{2\}\\lesssim\\frac\{\(p\+q\)\\log\(n/\(p\+q\)\)\}\{n\}\.\(17\)Fast spectral decay thus eliminates the curse of dimensionality \(Corollary[5](https://arxiv.org/html/2608.11661#Thmcorollary5), Appendix[D](https://arxiv.org/html/2608.11661#A4)\)\. The three terms in \([14](https://arxiv.org/html/2608.11661#S4.E14)\) form the complete error ledger: fast decay keepsd⋆d^\{\\star\}small; slow decay makes the head expensive in every respect\. The rest of this section analyzes the regime in which even the optimald⋆d^\{\\star\}is inadequate\.
#### Identifiability has a sample cost\.
Identifiability, as characterized in Section[3](https://arxiv.org/html/2608.11661#S3), is a population\-level statement\. Realizing it from finite samples incurs an additional cost, detailed in Appendix[E](https://arxiv.org/html/2608.11661#A5)\(Corollary[6](https://arxiv.org/html/2608.11661#Thmcorollary6)\)\. The total mode error of a fitted and subsequently whitened head \(Proposition[6](https://arxiv.org/html/2608.11661#Thmproposition6)\) decomposes into the excess risk of Corollary[4](https://arxiv.org/html/2608.11661#Thmcorollary4)plus the empirical whitening error\. Both terms are divided by the spectral gapΔ\\Deltaand decay asn−1/2n^\{\-1/2\}\.
#### Late versus early interaction\.
We now compare the multiplicative head, which we refer to as a late\-interaction model, with early\-interaction models in which the two inputs are combined during encoding\. By Corollary[1](https://arxiv.org/html/2608.11661#Thmcorollary1), every rank\-ddhead pays at least the spectral tail, so the spectrum determines the outcome\.
Define the*effective interaction rank*
dε\(F\)=min\{d:∑k\>dσk2∑k≥1σk2≤ε\}d\_\{\\varepsilon\}\(F\)=\\min\\Bigl\\\{d:\\frac\{\\sum\_\{k\>d\}\\sigma\_\{k\}^\{2\}\}\{\\sum\_\{k\\geq 1\}\\sigma\_\{k\}^\{2\}\}\\leq\\varepsilon\\Bigr\\\}\(18\)as the smallest embedding dimension for which a rank\-ddhead achieves relative errorε\\varepsilon\. The spectrum is called*flat*at scaleNNwhendε\(F\)=Ω\(\(1−ε\)N\)d\_\{\\varepsilon\}\(F\)=\\Omega\(\(1\-\\varepsilon\)N\)\.
The equality function realizes this worst case exactly\. On𝒰=𝒱=\{±1\}m\\mathcal\{U\}=\\mathcal\{V\}=\\\{\\pm 1\\\}^\{m\}withN=2mN=2^\{m\}, the interaction operator satisfiesTFeq=1NIdT\_\{F\_\{\\mathrm\{eq\}\}\}=\\frac\{1\}\{N\}\\mathrm\{Id\}\. The singular values are all equal to1/N1/\\sqrt\{N\}, the interaction rank isNN, and every rank\-ddhead achieves relative error exactly1−d/N1\-d/N\. Thus, attaining relative errorε\\varepsilonrequiresd≥\(1−ε\)2md\\geq\(1\-\\varepsilon\)2^\{m\}embedding dimensions, exponential in the input size \(Theorem[8](https://arxiv.org/html/2608.11661#Thmtheorem8), Appendix[D](https://arxiv.org/html/2608.11661#A4)\)\.
By contrast, an early\-interaction predictor withO\(m\)O\(m\)parameters representsFeqF\_\{\\mathrm\{eq\}\}exactly\. The identity⟨u,v⟩=m−2Ham\(u,v\)\\langle u,v\\rangle=m\-2\\,\\mathrm\{Ham\}\(u,v\)allows a simple threshold on the inner product to recover the equality indicator, whether implemented by a cross\-attention layer with identity projections or by a joint MLP withO\(m\)O\(m\)ReLUs\. Combined with Theorem[8](https://arxiv.org/html/2608.11661#Thmtheorem8), this yields an exponential separation:O\(m\)O\(m\)parameters suffice under early interaction, while any late\-interaction head requiresΩ\(2m\)\\Omega\(2^\{m\}\)dimensions to reach constant relative error \(Theorem[9](https://arxiv.org/html/2608.11661#Thmtheorem9), Appendix[D](https://arxiv.org/html/2608.11661#A4)\)\.
These observations lead to a concrete usability criterion \(Theorem[10](https://arxiv.org/html/2608.11661#Thmtheorem10), Appendix[D](https://arxiv.org/html/2608.11661#A4)\)\. When the effective interaction rankdε\(F\)d\_\{\\varepsilon\}\(F\)is small, a late\-interaction head is appropriate: its approximation and estimation errors are controllable, and its computational and statistical advantages are preserved\. When the spectrum is flat, the truncation floor forces every rank\-ddhead to relative error at leastε\\varepsilon, while early\-interaction models can still reachε\\varepsilonwith polynomial resources when the target has low\-dimensional joint structure\. The deciding quantity, the spectral decay rate, can be estimated from data by a weighted singular value decomposition of the empirical kernel matrix, making the criterion actionable before either model is trained\.
The criterion is a boundary, not a verdict\. The early\-interaction upper bounds assume low\-dimensional joint structure; a target that is both high\-rank and unstructured is difficult for both architectures\. Moreover, the parameter counts in Theorem[9](https://arxiv.org/html/2608.11661#Thmtheorem9)are statements about representational existence, not about learnability\. Section[5](https://arxiv.org/html/2608.11661#S5)instantiates both sides of the dichotomy: on gapped spectra the head identifies and estimates well, while flat and narrow\-band targets drive it to the exponential\-rank floor\.
## 5Experiments
Our experiments are designed for mechanism validation rather than benchmark chasing: each study isolates a quantity that the theory predicts exactly\. The controlled synthetic setting \(Section[5\.1](https://arxiv.org/html/2608.11661#S5.SS1)\) supplies exact ground truth and is reported in full, with the tables and figures supporting each conclusion placed alongside the corresponding paragraph\. Two real architectures instantiate the multiplicative head directly, DeepONet and CLIP, and are reported in full in Appendix[F](https://arxiv.org/html/2608.11661#A6)due to page limitation; here we state only the conclusions they establish\. On DeepONet, post\-hoc whitening recovers the analytic operator eigenbasis from an end\-to\-end\-trained model\. On two pretrained CLIP pairs, the interaction spectra are shared across models while per\-dimension concept probes are not, and fitting the single residual rotation reconciles them\.
### 5\.1Controlled synthetic study
#### Setup\.
This study validates the identifiability and usability claims of Sections[3](https://arxiv.org/html/2608.11661#S3)and[4](https://arxiv.org/html/2608.11661#S4)in a setting with exact ground truth\. We fix a targetF\(u,v\)=∑kσkak\(u\)bk\(v\)F\(u,v\)=\\sum\_\{k\}\\sigma\_\{k\}a\_\{k\}\(u\)b\_\{k\}\(v\)with shifted Legendre modes and a prescribed spectrum, fit linear dual encoders over a fixed Legendre feature map, and record four quantities per run: the fit error, the alignment of the recovered modes with the trueak,bka\_\{k\},b\_\{k\}, the encoder\-covariance conditioning, and the cross\-seed gauge distance\. The encoder\-realization error of Assumption[1](https://arxiv.org/html/2608.11661#Thmassumption1)is zero, so normalization is the only variable across runs\. The four paragraphs below use these quantities to test, in turn, the identifiability ranking of the gauges \(Theorem[2](https://arxiv.org/html/2608.11661#Thmtheorem2)\), the gap law of estimation \(Corollaries[3](https://arxiv.org/html/2608.11661#Thmcorollary3)and[6](https://arxiv.org/html/2608.11661#Thmcorollary6)\), the sample\-complexity rate \(Theorem[7](https://arxiv.org/html/2608.11661#Thmtheorem7)\), and the flat\-spectrum floor \(Theorem[8](https://arxiv.org/html/2608.11661#Thmtheorem8)\)\.
#### Normalization as gauge\-fixing\.
The seven schemes of Table[1](https://arxiv.org/html/2608.11661#S5.T1)separate cleanly by residual group\. Only whitening attains perfect mode alignment \(1\.0001\.000\) with zero cross\-seed gauge distance and drives the conditioning diagnosticmax\(condΣf,condΣg\)\\max\(\\mathrm\{cond\}\\,\\Sigma\_\{f\},\\mathrm\{cond\}\\,\\Sigma\_\{g\}\)to its theoretical floorσ12/σd2≈9\.5\\sigma\_\{1\}^\{2\}/\\sigma\_\{d\}^\{2\}\\approx 9\.5\(Corollary[3](https://arxiv.org/html/2608.11661#Thmcorollary3)\); every looser gauge remains stuck at alignment0\.530\.53to0\.650\.65, its residual group leaving the modes rotationally entangled \(Theorem[4](https://arxiv.org/html/2608.11661#Thmtheorem4)\)\. The gauge\-invariant diagnosticcond\(ΣfΣg\)\\mathrm\{cond\}\(\\Sigma\_\{f\}\\Sigma\_\{g\}\)cannot separate the schemes \(Lemma[2](https://arxiv.org/html/2608.11661#Thmlemma2)\)\. The two\-sided nonnegative gauge incurs error4\.254\.25, the predicted out\-of\-class cost of excluding negatively correlated modes \(Theorem[5](https://arxiv.org/html/2608.11661#Thmtheorem5)\)\. Normalization is therefore gauge fixing, and whitening is the only scheme that removes the symmetry completely, exactly as Theorem[2](https://arxiv.org/html/2608.11661#Thmtheorem2)predicts\.
Table 1:The seven normalization schemes, four seeds each\. Residual gauge: the subgroup that survives the constraint \(Theorem[2](https://arxiv.org/html/2608.11661#Thmtheorem2)predictsPerm⋉D±\\mathrm\{Perm\}\\ltimes\\mathrm\{D\}\_\{\\pm\}for whitening,O\\mathrm\{O\}for cosine,Perm⋉D\+\\mathrm\{Perm\}\\ltimes\\mathrm\{D\}\_\{\+\}for two\-sided nonnegativity\)\. Fit MSE: squared regression error\. Alignment: mean mode alignment of the two encoders after optimal permutation and sign matching\. Cond:max\(condΣf,condΣg\)\\max\(\\mathrm\{cond\}\\,\\Sigma\_\{f\},\\mathrm\{cond\}\\,\\Sigma\_\{g\}\), gauge dependent\. Cross\-seed: mean pairwise gauge distance between seeds\.
#### The gap as the single unifying quantity\.
The spectral gap is the single quantity governing identification\. With the finite\-sample whitening estimator of Appendix[E](https://arxiv.org/html/2608.11661#A5), we sweep the gap at fixed sample size and the sample size at fixed gap, eight configurations in all \(Table[2](https://arxiv.org/html/2608.11661#S5.T2); Figure[3](https://arxiv.org/html/2608.11661#A6.F3)in Appendix[F](https://arxiv.org/html/2608.11661#A6)plots the collapse\)\. The mode\-identification error scales as1/Δ1/\\Deltaat fixed sample size and asn−1/2n^\{\-1/2\}at fixed gap, whereΔ=mink\(σk2−σk\+12\)\\Delta=\\min\_\{k\}\(\\sigma\_\{k\}^\{2\}\-\\sigma\_\{k\+1\}^\{2\}\), and the two sweeps collapse onto the single relation
err⋅Δ⋅n≈2\.48\\mathrm\{err\}\\cdot\\Delta\\cdot\\sqrt\{n\}\\approx 2\.48\(19\)with coefficient of variation0\.120\.12: gap and sample size enter the identification error only through their product, the concrete face of Corollaries[3](https://arxiv.org/html/2608.11661#Thmcorollary3)and[6](https://arxiv.org/html/2608.11661#Thmcorollary6)\.
Table 2:Gap and sample\-size sweeps of the finite\-sample whitening estimator \(Appendix[E](https://arxiv.org/html/2608.11661#A5)\)\. Left group: the spectral gapΔ=mink\(σk2−σk\+12\)\\Delta=\\min\_\{k\}\(\\sigma\_\{k\}^\{2\}\-\\sigma\_\{k\+1\}^\{2\}\)swept by scaling a base spectrum at fixedn=3200n=3200\. Right group: sample size swept at fixedΔ=0\.900\\Delta=0\.900\. Mode error is the mean mode identification error; the producterr⋅Δ⋅n\\mathrm\{err\}\\cdot\\Delta\\cdot\\sqrt\{n\}is near\-constant across both groups,≈2\.48\\approx 2\.48\.
#### The sample\-complexity rate law\.
In the linear regime the head is exactly a rank\-ddmatrix sensing problem, and Theorem[7](https://arxiv.org/html/2608.11661#Thmtheorem7)predicts𝔼‖Θ^−Θ⋆‖F2≍σε2d\(p\+q\)/n\\mathbb\{E\}\\\|\\hat\{\\Theta\}\-\\Theta^\{\\star\}\\\|\_\{F\}^\{2\}\\asymp\\sigma\_\{\\varepsilon\}^\{2\}d\(p\+q\)/n\. We sweep each of the three variables with the others held fixed and fit log\-log slopes, the three sweeps shown in Figure[1](https://arxiv.org/html/2608.11661#S5.F1)\. The measured slopes are−1\.03\-1\.03innn,\+0\.89\+0\.89indd, and\+1\.24\+1\.24inp\+qp\+q, against the theoretical−1,\+1,\+1\-1,\+1,\+1; the excess in the last slope is the expected Marchenko\-Pastur finite\-sample correction, and gradient descent on the factored pair matches the least\-squares slope\. These sweeps reject the loose general bound of Remark[6](https://arxiv.org/html/2608.11661#Thmremark6), which would predict slope−1/2\-1/2: estimation is governed by the sum of the marginal encoder complexities, and the1/n1/ndependence is genuine\.
Figure 1:The sample\-complexity law of Theorem[7](https://arxiv.org/html/2608.11661#Thmtheorem7)\. Log\-log sweeps of the Frobenius excess risk against sample sizenn\(left\), interaction rankdd\(center\), and ambient dimensionp\+qp\+q\(right\)\. Points are the measured risk; the dotted line is the theoretical rateσ2d\(2p−d\)/n\\sigma^\{2\}d\(2p\-d\)/n, with the finite\-nncorrected form \(dashed\) in the right panel\.
#### The usability criterion\.
A flat spectrum floors the head regardless of capacity\. We test two flat and near\-flat regimes, both tabulated in Tables[4](https://arxiv.org/html/2608.11661#S5.T4)and[4](https://arxiv.org/html/2608.11661#S5.T4)\. On the Boolean equality kernelF=𝟏\[u=v\]F=\\mathbf\{1\}\[u=v\]on\{±1\}m\\\{\\pm 1\\\}^\{m\}, which has a completely flat spectrum, a rank\-ddhead incurs relative error exactly1−d/N1\-d/Nand requiresd≈0\.9⋅2md\\approx 0\.9\\cdot 2^\{m\}dimensions to reach10%10\\%error, while an early\-interaction threshold on⟨u,v⟩\\langle u,v\\rangleis exact with2m\+12m\+1parameters \(Theorems[8](https://arxiv.org/html/2608.11661#Thmtheorem8)and[9](https://arxiv.org/html/2608.11661#Thmtheorem9)\)\. Smooth narrow\-band kernels give the polynomial counterpartdε=Θ\(1/h\)d\_\{\\varepsilon\}=\\Theta\(1/h\)\(Remark[7](https://arxiv.org/html/2608.11661#Thmremark7)\)\. A trained\-model anchor on a block\-boxcar kernel realizes the same floor \(Table[10](https://arxiv.org/html/2608.11661#A6.T10)and Figure[5](https://arxiv.org/html/2608.11661#A6.F5), Appendix[F](https://arxiv.org/html/2608.11661#A6)\)\.
Table 3:The boolean equality kernelF=𝟏\[u=v\]F=\\mathbf\{1\}\[u=v\]on\{±1\}m\\\{\\pm 1\\\}^\{m\},N=2mN=2^\{m\}\. Rankdd: embedding dimension at which the head first reaches10%10\\%relative error\. Early: parameters2m\+12m\+1and error of the early\-interaction comparator\.
Table 4:Narrowband periodic\-Gaussian kernel of bandwidthhhon the circle: interaction rankdεd\_\{\\varepsilon\}forε=10%\\varepsilon=10\\%error\. The rank grows asΘ\(1/h\)\\Theta\(1/h\)\(Remark[7](https://arxiv.org/html/2608.11661#Thmremark7)\)\.
## 6Conclusion
This paper has examined a single architectural pattern, the multiplicative dual\-encoder head, through a unifying lens: the interaction spectrum of the target\. Through theoretical analysis and numerical experiments, we answer four central questions\. \(i\) Approximation error decomposes into a truncation term intrinsic to the target and a realization term intrinsic to the encoders, with target smoothness controlling the decay rate\. \(ii\) Sample complexity is governed by the sum of the two marginal encoder complexities, with the linear regime anchoring the minimax rate\. \(iii\) Usability is determined by the decay rate itself: a flat spectrum imposes an error floor that no rank\-ddhead can escape, while early\-interaction models can, and the deciding quantity is measurable from data before either model is trained\. \(iv\) Identifiability, the problem that the framework brings into focus, is gauge fixing: every normalization used in practice is a section of the gauge orbit; whitening pins the interaction modes up to permutation and sign; and the spectral gap simultaneously controls optimization conditioning, identifiability, and estimation\.
## Ethics Statement
*This work adheres to the principles outlined in the ICLR Code of Ethics\.*
## References
- Aroraet al\.\(2012\)S\. Arora, R\. Ge, R\. Kannan, and A\. MoitraComputing a nonnegative matrix factorization–provably\.InProceedings of the forty\-fourth annual ACM symposium on Theory of computing,pp\. 145–162\.Cited by:[Appendix A](https://arxiv.org/html/2608.11661#A1.SS0.SSS0.Px11.p1.1),[Appendix C](https://arxiv.org/html/2608.11661#A3.SS0.SSS0.Px6.p1.1),[§E\.2](https://arxiv.org/html/2608.11661#A5.SS2.p1.1.1),[§3](https://arxiv.org/html/2608.11661#S3.SS0.SSS0.Px2.p3.1),[Definition 3](https://arxiv.org/html/2608.11661#Thmdefinition3.p1.2.1),[Theorem 5](https://arxiv.org/html/2608.11661#Thmtheorem5.p1.2.1)\.
- Bardeset al\.\(2022\)A\. Bardes, J\. Ponce, and Y\. LeCunVICReg: variance\-invariance\-covariance regularization for self\-supervised learning\.InInternational Conference on Learning Representations,Cited by:[Appendix A](https://arxiv.org/html/2608.11661#A1.SS0.SSS0.Px11.p1.1),[Theorem 2](https://arxiv.org/html/2608.11661#Thmtheorem2.p1.2.1)\.
- Barretoet al\.\(2017\)A\. Barreto, W\. Dabney, R\. Munos, J\. J\. Hunt, T\. Schaul, H\. Van Hasselt, and D\. SilverSuccessor features for transfer in reinforcement learning\.Advances in neural information processing systems30\.Cited by:[Appendix A](https://arxiv.org/html/2608.11661#A1.SS0.SSS0.Px6.p1.1),[Table 5](https://arxiv.org/html/2608.11661#A1.T5.2.7.1.1.1),[§1](https://arxiv.org/html/2608.11661#S1.p1.1)\.
- Bartlett and Mendelson \(2002\)P\. L\. Bartlett and S\. MendelsonRademacher and gaussian complexities: risk bounds and structural results\.Journal of machine learning research3\(Nov\),pp\. 463–482\.Cited by:[Appendix D](https://arxiv.org/html/2608.11661#A4.SS0.SSS0.Px2.p1.2)\.
- Candes and Plan \(2011\)E\. J\. Candes and Y\. PlanTight oracle inequalities for low\-rank matrix recovery from a minimal number of noisy random measurements\.IEEE Transactions on Information Theory57\(4\),pp\. 2342–2359\.Cited by:[Appendix A](https://arxiv.org/html/2608.11661#A1.SS0.SSS0.Px11.p1.1),[Appendix D](https://arxiv.org/html/2608.11661#A4.SS0.SSS0.Px3.p1.1)\.
- Donoho and Stodden \(2003\)D\. Donoho and V\. StoddenWhen does non\-negative matrix factorization give a correct decomposition into parts?\.Advances in neural information processing systems16\.Cited by:[Appendix A](https://arxiv.org/html/2608.11661#A1.SS0.SSS0.Px11.p1.1),[Appendix C](https://arxiv.org/html/2608.11661#A3.SS0.SSS0.Px6.p1.1),[§3](https://arxiv.org/html/2608.11661#S3.SS0.SSS0.Px2.p3.1),[Definition 3](https://arxiv.org/html/2608.11661#Thmdefinition3.p1.2.1),[Theorem 5](https://arxiv.org/html/2608.11661#Thmtheorem5.p1.2.1)\.
- Eckart and Young \(1936\)C\. Eckart and G\. YoungThe approximation of one matrix by another of lower rank\.Psychometrika1\(3\),pp\. 211–218\.Cited by:[§2\.1](https://arxiv.org/html/2608.11661#S2.SS1.p3.1)\.
- Ermolovet al\.\(2021\)A\. Ermolov, A\. Siarohin, E\. Sangineto, and N\. SebeWhitening for self\-supervised representation learning\.InInternational conference on machine learning,pp\. 3015–3024\.Cited by:[Appendix A](https://arxiv.org/html/2608.11661#A1.SS0.SSS0.Px11.p1.1),[Theorem 2](https://arxiv.org/html/2608.11661#Thmtheorem2.p1.2.1)\.
- Honget al\.\(2021\)Z\. Hong, G\. Yang, and P\. AgrawalBi\-linear value networks for multi\-goal reinforcement learning\.InInternational Conference on Learning Representations,Cited by:[Appendix A](https://arxiv.org/html/2608.11661#A1.SS0.SSS0.Px5.p1.1),[Table 5](https://arxiv.org/html/2608.11661#A1.T5.2.6.1.1.1),[§1](https://arxiv.org/html/2608.11661#S1.p1.1)\.
- Junet al\.\(2019\)K\. Jun, R\. Willett, S\. Wright, and R\. NowakBilinear bandits with low\-rank structure\.InInternational Conference on Machine Learning,pp\. 3163–3172\.Cited by:[Appendix A](https://arxiv.org/html/2608.11661#A1.SS0.SSS0.Px11.p1.1),[Remark 6](https://arxiv.org/html/2608.11661#Thmremark6.p1.1.1)\.
- Karpukhinet al\.\(2020\)V\. Karpukhin, B\. Oguz, S\. Min, P\. Lewis, L\. Wu, S\. Edunov, D\. Chen, and W\. YihDense passage retrieval for open\-domain question answering\.InProceedings of the 2020 conference on empirical methods in natural language processing \(EMNLP\),pp\. 6769–6781\.Cited by:[Appendix A](https://arxiv.org/html/2608.11661#A1.SS0.SSS0.Px3.p1.1),[Table 5](https://arxiv.org/html/2608.11661#A1.T5.2.4.1.1.1),[§1](https://arxiv.org/html/2608.11661#S1.p1.1)\.
- Katharopouloset al\.\(2020\)A\. Katharopoulos, A\. Vyas, N\. Pappas, and F\. FleuretTransformers are rnns: fast autoregressive transformers with linear attention\.InInternational conference on machine learning,pp\. 5156–5165\.Cited by:[Appendix A](https://arxiv.org/html/2608.11661#A1.SS0.SSS0.Px8.p1.1),[Table 5](https://arxiv.org/html/2608.11661#A1.T5.2.9.1.1.1),[§1](https://arxiv.org/html/2608.11661#S1.p1.1),[§2\.3](https://arxiv.org/html/2608.11661#S2.SS3.p5.2),[Remark 2](https://arxiv.org/html/2608.11661#Thmremark2.p1.1.1)\.
- Little and Reade \(1984\)G\. Little and J\. B\. ReadeEigenvalues of analytic kernels\.SIAM journal on mathematical analysis15\(1\),pp\. 133–136\.Cited by:[Appendix B](https://arxiv.org/html/2608.11661#A2.SS0.SSS0.Px4.p1.2)\.
- Luet al\.\(2021\)L\. Lu, P\. Jin, G\. Pang, Z\. Zhang, and G\. E\. KarniadakisLearning nonlinear operators via deeponet based on the universal approximation theorem of operators\.Nature machine intelligence3\(3\),pp\. 218–229\.Cited by:[Appendix A](https://arxiv.org/html/2608.11661#A1.SS0.SSS0.Px2.p1.1),[Table 5](https://arxiv.org/html/2608.11661#A1.T5.2.3.1.1.1),[§1](https://arxiv.org/html/2608.11661#S1.p1.1),[§2\.3](https://arxiv.org/html/2608.11661#S2.SS3.p5.2),[§3](https://arxiv.org/html/2608.11661#S3.SS0.SSS0.Px3.p3.1),[Remark 2](https://arxiv.org/html/2608.11661#Thmremark2.p1.1.1)\.
- Maurer \(2016\)A\. MaurerA vector\-contraction inequality for rademacher complexities\.InInternational Conference on Algorithmic Learning Theory,pp\. 3–17\.Cited by:[Appendix D](https://arxiv.org/html/2608.11661#A4.SS0.SSS0.Px1.p1.2)\.
- Negahban and Wainwright \(2009\)S\. N\. Negahban and M\. J\. WainwrightEstimation of \(near\) low\-rank matrices with noise and high\-dimensional scaling\.InInternational Conference on Machine Learning,External Links:[Link](https://api.semanticscholar.org/CorpusID:1004801)Cited by:[Appendix A](https://arxiv.org/html/2608.11661#A1.SS0.SSS0.Px11.p1.1),[Appendix D](https://arxiv.org/html/2608.11661#A4.SS0.SSS0.Px3.p1.1)\.
- Radfordet al\.\(2021\)A\. Radford, J\. W\. Kim, C\. Hallacy, A\. Ramesh, G\. Goh, S\. Agarwal, G\. Sastry, A\. Askell, P\. Mishkin, J\. Clark,et al\.Learning transferable visual models from natural language supervision\.InInternational conference on machine learning,pp\. 8748–8763\.Cited by:[Table 5](https://arxiv.org/html/2608.11661#A1.T5.2.2.1.1.1),[Table 6](https://arxiv.org/html/2608.11661#A3.T6.2.5.1.1.1),[§1](https://arxiv.org/html/2608.11661#S1.p1.1),[§3](https://arxiv.org/html/2608.11661#S3.SS0.SSS0.Px2.p2.1),[§3](https://arxiv.org/html/2608.11661#S3.p1.1),[Theorem 4](https://arxiv.org/html/2608.11661#Thmtheorem4.p1.1.1)\.
- Reade \(1983\)J\. B\. ReadeEigenvalues of positive definite kernels\.SIAM Journal on Mathematical Analysis14\(1\),pp\. 152–157\.Cited by:[Appendix B](https://arxiv.org/html/2608.11661#A2.SS0.SSS0.Px4.p1.2)\.
- Rendle \(2010\)S\. RendleFactorization machines\.In2010 IEEE International conference on data mining,pp\. 995–1000\.Cited by:[Appendix A](https://arxiv.org/html/2608.11661#A1.SS0.SSS0.Px9.p1.1),[Table 5](https://arxiv.org/html/2608.11661#A1.T5.2.10.1.1.1),[§1](https://arxiv.org/html/2608.11661#S1.p1.1)\.
- Rohde and Tsybakov \(2009\)A\. Rohde and A\. TsybakovEstimation of high\-dimensional low\-rank matrices\.Annals of Statistics39,pp\. 887–930\.External Links:[Link](https://api.semanticscholar.org/CorpusID:88512332)Cited by:[Appendix A](https://arxiv.org/html/2608.11661#A1.SS0.SSS0.Px11.p1.1),[Appendix D](https://arxiv.org/html/2608.11661#A4.SS0.SSS0.Px3.p1.1)\.
- Schaulet al\.\(2015\)T\. Schaul, D\. Horgan, K\. Gregor, and D\. SilverUniversal value function approximators\.InInternational conference on machine learning,pp\. 1312–1320\.Cited by:[Appendix A](https://arxiv.org/html/2608.11661#A1.SS0.SSS0.Px4.p1.1),[Table 5](https://arxiv.org/html/2608.11661#A1.T5.2.5.1.1.1),[§1](https://arxiv.org/html/2608.11661#S1.p1.1)\.
- Schmidt \(1907\)E\. SchmidtZur theorie der linearen und nichtlinearen integralgleichungen: i\. teil: entwicklung willkürlicher funktionen nach systemen vorgeschriebener\.Mathematische Annalen63\(4\),pp\. 433–476\.Cited by:[§2\.1](https://arxiv.org/html/2608.11661#S2.SS1.p3.1)\.
- Trouillonet al\.\(2016\)T\. Trouillon, J\. Welbl, S\. Riedel, É\. Gaussier, and G\. BouchardComplex embeddings for simple link prediction\.InInternational conference on machine learning,pp\. 2071–2080\.Cited by:[Appendix A](https://arxiv.org/html/2608.11661#A1.SS0.SSS0.Px7.p1.1),[Table 5](https://arxiv.org/html/2608.11661#A1.T5.2.8.1.1.1),[§1](https://arxiv.org/html/2608.11661#S1.p1.1)\.
- Yanget al\.\(2014\)B\. Yang, W\. Yih, X\. He, J\. Gao, and L\. DengEmbedding entities and relations for learning and inference in knowledge bases\.arXiv preprint arXiv:1412\.6575\.Cited by:[Appendix A](https://arxiv.org/html/2608.11661#A1.SS0.SSS0.Px7.p1.1),[Table 5](https://arxiv.org/html/2608.11661#A1.T5.2.8.1.1.1),[§1](https://arxiv.org/html/2608.11661#S1.p1.1)\.
- Yarotsky \(2017\)D\. YarotskyError bounds for approximations with deep relu networks\.Neural networks94,pp\. 103–114\.Cited by:[Appendix B](https://arxiv.org/html/2608.11661#A2.SS0.SSS0.Px5.p1.1),[§2\.3](https://arxiv.org/html/2608.11661#S2.SS3.p1.1),[Remark 7](https://arxiv.org/html/2608.11661#Thmremark7.p1.1.1)\.
- Zbontaret al\.\(2021\)J\. Zbontar, L\. Jing, I\. Misra, Y\. LeCun, and S\. DenyBarlow twins: self\-supervised learning via redundancy reduction\.InInternational conference on machine learning,pp\. 12310–12320\.Cited by:[Appendix A](https://arxiv.org/html/2608.11661#A1.SS0.SSS0.Px11.p1.1),[Theorem 2](https://arxiv.org/html/2608.11661#Thmtheorem2.p1.2.1)\.
- Zhaoet al\.\(2025\)Z\. Zhao, T\. Chen, Z\. Cai, X\. Li, H\. Li, Q\. Chen, and G\. ZhuCrossfi: a cross domain wi\-fi sensing framework based on siamese network\.IEEE Internet of Things Journal12\(12\),pp\. 20138–20155\.Cited by:[§1](https://arxiv.org/html/2608.11661#S1.p1.1)\.
- Zhao and Li \(2025\)Z\. Zhao and S\. LiThe impacts of data privacy regulations on food\-delivery platforms\.Transportation Research Part C: Emerging Technologies181,pp\. 105364\.Cited by:[Appendix A](https://arxiv.org/html/2608.11661#A1.SS0.SSS0.Px10.p1.1)\.
- Zhao and Li \(2026a\)Z\. Zhao and S\. LiDiscriminatory order assignment and payment\-setting of on\-demand food\-delivery platforms: a multi\-action and multi\-agent reinforcement learning framework\.Transportation Research Part E: Logistics and Transportation Review208,pp\. 104653\.Cited by:[Appendix A](https://arxiv.org/html/2608.11661#A1.SS0.SSS0.Px10.p1.1),[Table 5](https://arxiv.org/html/2608.11661#A1.T5.2.11.1.1.1),[§1](https://arxiv.org/html/2608.11661#S1.p1.1),[§2\.3](https://arxiv.org/html/2608.11661#S2.SS3.p5.2)\.
- Zhao and Li \(2026b\)Z\. Zhao and S\. LiTriple\-bert: do we really need marl for order dispatch on ride\-sharing platforms?\.InInternational Conference on Learning Representations,Vol\.2026,pp\. 153780–153808\.Cited by:[Appendix A](https://arxiv.org/html/2608.11661#A1.SS0.SSS0.Px10.p1.1),[Table 5](https://arxiv.org/html/2608.11661#A1.T5.2.11.1.1.1),[Table 6](https://arxiv.org/html/2608.11661#A3.T6.2.3.1.1.1),[§1](https://arxiv.org/html/2608.11661#S1.p1.1),[§3](https://arxiv.org/html/2608.11661#S3.SS0.SSS0.Px2.p3.1),[§3](https://arxiv.org/html/2608.11661#S3.p1.1),[Remark 2](https://arxiv.org/html/2608.11661#Thmremark2.p1.1.1),[Remark 4](https://arxiv.org/html/2608.11661#Thmremark4.p1.1.1)\.
## Appendix Contents
## Appendix ADerivations for the Unification Table
This appendix spells out, for each of the ten families of Table[5](https://arxiv.org/html/2608.11661#A1.T5), the identification that makes it an instance of the multiplicative dual\-encoder head \([1](https://arxiv.org/html/2608.11661#S1.E1)\): the two input spaces, the encoder maps, the score, the reference measures that define the interaction spectrum, and the boundary cases where an additional output\-level normalization acts after the inner product and therefore leaves the gauge untouched\. Table[5](https://arxiv.org/html/2608.11661#A1.T5)collects the ten families, each computing a real\-valued output for a pair of inputs as the inner product of two separately learned encodings; reading across the rows makes visible what is shared, and the paragraphs below identify the interaction operator and its spectrum for each family, including the boundary cases\.
#### CLIP\.
Inputs are an imageIIand a textTT; the encoders are the image towerffand the text towergg, bothℓ2\\ell\_\{2\}\-normalized; the score issim\(I,T\)=⟨f\(I\),g\(T\)⟩\\mathrm\{sim\}\(I,T\)=\\langle f\(I\),g\(T\)\\rangle, rescaled by a temperature and fed into a softmax over the batch\. The reference measures are the empirical distributions over images and over texts in the training batch\. The temperature and the softmax act on the similarity matrix after the inner product, so they are output\-level and do not affect the embedding\-level gauge; the unit\-norm constraints are embedding\-level and leave exactly the residual rotationO\\mathrm\{O\}of Theorem[4](https://arxiv.org/html/2608.11661#Thmtheorem4)\.
#### DeepONet\.
Inputs are an input functionaa\(on a sensor grid\) and a query pointyy; the encoders are the branch netbband the trunk nettt; the score is𝒢\(a\)\(y\)≈⟨b\(a\),t\(y\)⟩\\mathcal\{G\}\(a\)\(y\)\\approx\\langle b\(a\),t\(y\)\\rangle\([14](https://arxiv.org/html/2608.11661#bib.bib2)\)\. The reference measures are the distribution over input functions and the measure over query points\. For a linear operator with kernelκ=∑kσkβk\(y\)γk\(x\)\\kappa=\\sum\_\{k\}\\sigma\_\{k\}\\beta\_\{k\}\(y\)\\gamma\_\{k\}\(x\)and a white input measure, the interaction spectrum coincides with the operator singular values and the trunk’s ideal basis with the left singular functions, which is the ground truth used in Appendix[F\.2](https://arxiv.org/html/2608.11661#A6.SS2)\. The boundary case is the scale imbalance between branch and trunk outputs, familiar in practice; it is the scale redundancy of the gauge group, diagnosed in Section[3](https://arxiv.org/html/2608.11661#S3)\. POD\-DeepONet substitutes a precomputed orthogonal POD basis for the learned trunk\([14](https://arxiv.org/html/2608.11661#bib.bib2)\), motivated by the same nonuniqueness; our treatment is complementary, characterizing when the learned basis becomes identifiable instead of replacing it \(Appendix[F\.2](https://arxiv.org/html/2608.11661#A6.SS2)\)\.
#### Two\-tower retrieval\.
Inputs are a queryqqand a documentdd; the encoders are two towersφ,ψ\\varphi,\\psi; the score is a relevance measurerel\(q,d\)=⟨φ\(q\),ψ\(d\)⟩\\mathrm\{rel\}\(q,d\)=\\langle\\varphi\(q\),\\psi\(d\)\\rangle\([11](https://arxiv.org/html/2608.11661#bib.bib3)\)\. The reference measures are the query distribution and the document distribution, and the interaction spectrum is the relevance spectrum of the underlying ranker\. The top\-kknearest\-neighbor search is output\-level\.
#### UVFA \(universal value function approximators\)\.
Inputs are a statessand a goalgg; the encoders are the state encoderϕ\(s\)\\phi\(s\), a learned map from states to features, and the goal encoderψ\(g\)\\psi\(g\), a learned map from goals to features; the score is the goal\-conditioned valueV\(s,g\)=⟨ϕ\(s\),ψ\(g\)⟩V\(s,g\)=\\langle\\phi\(s\),\\psi\(g\)\\rangle\([21](https://arxiv.org/html/2608.11661#bib.bib4)\)\. The reference measures are the state\-visitation distribution and the goal distribution\. Goal\-conditioned value functions are naturally low\-rank when goals and states share a low\-dimensional feature structure, which is exactly the case in which the representation \([1](https://arxiv.org/html/2608.11661#S1.E1)\) is sample\-efficient \(Section[4](https://arxiv.org/html/2608.11661#S4)\)\. The interaction spectrum also gives these methods a rank\-selection rule they currently lack: the number of features in a successor\-feature representation or the width of a value\-network embedding is typically a tuned hyperparameter, whereas Corollary[5](https://arxiv.org/html/2608.11661#Thmcorollary5)ties it to the decay of the spectrum\.
#### BVN \(bilinear value networks\)\.
Inputs are a state\-action pair\(s,a\)\(s,a\)and a state\-goal pair\(s,g\)\(s,g\); the encoders are the MLPf\(s,a\)f\(s,a\)encoding the state\-action pair and the MLPφ\(s,g\)\\varphi\(s,g\)encoding the state\-goal pair; the score isQ\(s,a,g\)=⟨f\(s,a\),φ\(s,g\)⟩Q\(s,a,g\)=\\langle f\(s,a\),\\varphi\(s,g\)\\rangle\([9](https://arxiv.org/html/2608.11661#bib.bib5)\)\. The reference measures are the joint distributions over\(s,a\)\(s,a\)and\(s,g\)\(s,g\)\. Note that the two arguments share the statess: the head’s theory applies to the functionQQof the pair of arguments, and the shared state is exactly the coupling that makesQQnon\-separable in\(s,a,g\)\(s,a,g\)directly, which is why the two\-input factorization matters\.
#### Successor features\.
Inputs are a statessand a reward weight vectorww; the encoder on the state side is the successor\-feature mapϕ\(s\)\\phi\(s\)and the encoder on the weight side is the identityww\(a singleton encoder class, not learned\); the score isVw\(s\)=⟨ϕ\(s\),w⟩V\_\{w\}\(s\)=\\langle\\phi\(s\),w\\rangle\([3](https://arxiv.org/html/2608.11661#bib.bib6)\)\. This is the head with𝒢=\{id\}\\mathcal\{G\}=\\\{\\mathrm\{id\}\\\}, whose Rademacher complexity is zero, so the sample\-complexity bounds of Section[4](https://arxiv.org/html/2608.11661#S4)specialize to the standard guarantee for successor features: only the state\-side class matters\.
#### DistMult and ComplEx\.
Inputs are a head entityhhand a tail entitytt; the encoders are entity embeddingseh,et∈ℝde\_\{h\},e\_\{t\}\\in\\mathbb\{R\}^\{d\}; the score is the bilinear form⟨eh,r∘et⟩\\langle e\_\{h\},r\\circ e\_\{t\}\\rangleof DistMult or the real part of the complex producteh⊤diag\(r\)et¯e\_\{h\}^\{\\top\}\\,\\mathrm\{diag\}\(r\)\\,\\overline\{e\_\{t\}\}of ComplEx, whererris the relation embedding\([24](https://arxiv.org/html/2608.11661#bib.bib7);[23](https://arxiv.org/html/2608.11661#bib.bib8)\)\. Both are rank\-ddbilinear forms on the entity embeddings, i\.e\., heads with linear encoders and a fixed middle matrix; the interaction spectrum is governed by the Gram structure of the entity embeddings under the relation\. The reference measure is the empirical distribution over observed triples\.
#### Linear attention\.
Inputs are a queryqqand a keykkin a common space; the encoders are the feature mapϕ\\phiapplied to each; the attention weight is⟨ϕ\(q\),ϕ\(k\)⟩\\langle\\phi\(q\),\\phi\(k\)\\rangle\([12](https://arxiv.org/html/2608.11661#bib.bib9)\)\. The aggregation step∑jϕ\(kj\)vj⊤\\sum\_\{j\}\\phi\(k\_\{j\}\)v\_\{j\}^\{\\top\}makes the compute separation of Proposition[1](https://arxiv.org/html/2608.11661#Thmproposition1)exact: the pairwise inner products are never formed, giving theO\(n2\)→O\(n\)O\(n^\{2\}\)\\to O\(n\)complexity\. The reference measure is the distribution over keys \(or key\-value pairs\)\. Finite\-dimensional feature maps realize the head exactly; softmax attention with random\-feature maps is the limiting case of an infinite\-dimensionalϕ\\phi\.
#### Factorization machines\.
Inputs are a feature indexiiand a feature indexjj; the encoders are the embedding mapsvi,vjv\_\{i\},v\_\{j\}; the pairwise term isxixj⟨vi,vj⟩x\_\{i\}x\_\{j\}\\langle v\_\{i\},v\_\{j\}\\rangle\([19](https://arxiv.org/html/2608.11661#bib.bib10)\)\. The reference measure is the empirical distribution over feature pairs weighted by the data\. The interaction spectrum of the pairwise term is the spectrum of the Gram matrix of the embeddings, and the rank selection of Section[4](https://arxiv.org/html/2608.11661#S4)is the standard FM capacity control viewed through the interaction spectrum\.
#### QK\-attention matching\.
Inputs are a workerwwand an orderoo; the encoders are a BERT\-style encoderffand an MLPgg; the score is the match value⟨f\(w\),g\(o\)⟩\\langle f\(w\),g\(o\)\\rangleforming theQQ\-matrix of a bipartite assignment problem, with a one\-sided softplus and unit\-norm constraint on the order encoder\([30](https://arxiv.org/html/2608.11661#bib.bib13);[29](https://arxiv.org/html/2608.11661#bib.bib14);[28](https://arxiv.org/html/2608.11661#bib.bib30)\)\. The reference measure is the empirical distribution over worker\-order pairs\. The matching layer that consumes the score matrix is output\-level; the one\-sided constraint is embedding\-level but, as Remark[4](https://arxiv.org/html/2608.11661#Thmremark4)explains, it is a stability fix rather than an identifiability scheme\. The two papers differ only in the downstream use of the same head, which is why they share the identifiability properties derived in this paper\.
Table 5:Ten families of methods as instances of the multiplicative dual\-encoder head \([1](https://arxiv.org/html/2608.11661#S1.E1)\): for each family, the two inputs, the encoder maps, the output, and the downstream use\.
#### Positioning against prior work\.
The paragraphs above identify each family and its interaction operator; we now position the paper’s claims against the closest prior work, none of which treats the families as instances of a shared architecture or formulates the identifiability question\. On the estimation side, the linear case of the head is rank\-ddmatrix sensing, and the paper relies on the classical minimax theory, upper bounds via restricted\-isometry\-type analyses\([5](https://arxiv.org/html/2608.11661#bib.bib24);[16](https://arxiv.org/html/2608.11661#bib.bib25)\)and matching lower bounds\([20](https://arxiv.org/html/2608.11661#bib.bib26)\); bilinear bandits\([10](https://arxiv.org/html/2608.11661#bib.bib27)\)study the online counterpart\. The contribution relative to this line is a translation rather than a new bound: the estimation problem of a trained head is identified with matrix sensing, and the minimax rate is connected to the rank\-selection rule of Section[4](https://arxiv.org/html/2608.11661#S4)through the decay of the interaction spectrum, which the matrix\-sensing literature does not formulate because it treats the rank as given\. On the identifiability side, the direct precursor of our remedy is self\-supervised whitening: W\-MSE applies hard whitening to the embeddings\([8](https://arxiv.org/html/2608.11661#bib.bib16)\), and Barlow Twins and VICReg implement its soft relaxations, pushing the cross\-correlation matrix toward the identity and regularizing variance and covariance\([26](https://arxiv.org/html/2608.11661#bib.bib17);[2](https://arxiv.org/html/2608.11661#bib.bib18)\); it is also known that whitening reduces a general linear gauge to an orthogonal ambiguity, a polar\-decomposition observation\. Against this background, our contribution is not the whitening operation itself, which we do not claim, but three things: the interpretation of normalization as a section of the gauge orbit and the classification of the residual subgroups across the normalizations used in practice \(Table[6](https://arxiv.org/html/2608.11661#A3.T6), Appendix[C](https://arxiv.org/html/2608.11661#A3)\); the theorem that whitening together with a spectral gap pins the modes down to permutation and sign, strictly beyond the orthogonal ambiguity that decorrelation alone leaves \(Theorem[2](https://arxiv.org/html/2608.11661#Thmtheorem2)\); and the statement that the interaction spectrum is the unique gauge\-invariant content of a trained model, an interpretability claim with a theorem behind it rather than an empirical observation\. The nonnegative variant of the bilinear form has its own classical identifiability theory, separability\-type conditions for uniqueness up to permutation and scale\([6](https://arxiv.org/html/2608.11661#bib.bib11);[1](https://arxiv.org/html/2608.11661#bib.bib12)\), which the gauge analysis of Section[3](https://arxiv.org/html/2608.11661#S3)subsumes as a special case \(Theorem[5](https://arxiv.org/html/2608.11661#Thmtheorem5)\)\.
## Appendix BProofs for Section 2
We prove the results of Section[2](https://arxiv.org/html/2608.11661#S2)in order\. Throughout,∥⋅∥\\\|\\cdot\\\|denotes theL2\(μU⊗μV\)L^\{2\}\(\\mu\_\{U\}\\otimes\\mu\_\{V\}\)norm, and∥⋅∥L2\(μU\)\\\|\\cdot\\\|\_\{L^\{2\}\(\\mu\_\{U\}\)\},∥⋅∥L2\(μV\)\\\|\\cdot\\\|\_\{L^\{2\}\(\\mu\_\{V\}\)\}the marginal norms\.
###### Assumption 1\(Encoder realizability\)\.
The encoder classesℱ\\mathcal\{F\}and𝒢\\mathcal\{G\}contain functionsf^,g^\\hat\{f\},\\hat\{g\}such that, for some nonnegative errorsηf,k,ηg,k\\eta\_\{f,k\},\\eta\_\{g,k\},
∥f^k−σkak∥L2\(μU\)≤ηf,k,∥g^k−bk∥L2\(μV\)≤ηg,k,k=1,…,d\.\\\|\\hat\{f\}\_\{k\}\-\\sigma\_\{k\}a\_\{k\}\\\|\_\{L^\{2\}\(\\mu\_\{U\}\)\}\\leq\\eta\_\{f,k\},\\qquad\\\|\\hat\{g\}\_\{k\}\-b\_\{k\}\\\|\_\{L^\{2\}\(\\mu\_\{V\}\)\}\\leq\\eta\_\{g,k\},\\qquad k=1,\\dots,d\.\(20\)
###### Proposition 1\(Compute separation\)\.
Evaluating a rank\-ddhead on alln×mn\\times minput pairs costs
O\(\(n\+m\)Cnet\+nmd\),O\\bigl\(\(n\+m\)\\,C\_\{\\mathrm\{net\}\}\+nm\\,d\\bigr\),\(21\)whereCnetC\_\{\\mathrm\{net\}\}is the cost of one encoder forward pass, whereas a model that processes each pair jointly costsO\(nmCnet\)O\(nm\\,C\_\{\\mathrm\{net\}\}\)\. When the encoder forward pass dominates the inner product, the head is cheaper by a factorCnet/dC\_\{\\mathrm\{net\}\}/d\.
###### Lemma 1\(Rank equivalence\)\.
A functionS∈L2\(μU⊗μV\)S\\in L^\{2\}\(\\mu\_\{U\}\\otimes\\mu\_\{V\}\)can be written asS\(u,v\)=⟨f\(u\),g\(v\)⟩S\(u,v\)=\\langle f\(u\),g\(v\)\\ranglefor some encodersf:𝒰→ℝdf:\\mathcal\{U\}\\to\\mathbb\{R\}^\{d\}andg:𝒱→ℝdg:\\mathcal\{V\}\\to\\mathbb\{R\}^\{d\}if and only ifrank\(TS\)≤d\\rank\(T\_\{S\}\)\\leq d\. In particular, the parametric class of eq\. \([2](https://arxiv.org/html/2608.11661#S1.E2)\) coincides withℳd=\{F:i\-rank\(F\)≤d\}\\mathcal\{M\}\_\{d\}=\\\{F:\\mathrm\{i\}\\text\{\-\}\\mathrm\{rank\}\(F\)\\leq d\\\}, and we use the two descriptions interchangeably\.
#### Proof of Lemma[1](https://arxiv.org/html/2608.11661#Thmlemma1)\(Rank equivalence\)\.
The forward direction is immediate: ifS=∑k≤dfk⊗gkS=\\sum\_\{k\\leq d\}f\_\{k\}\\otimes g\_\{k\}, then the interaction operator is a sum of at mostddrank\-one operators,
TS=∑k≤d⟨⋅,gk⟩L2\(μV\)fk,T\_\{S\}=\\sum\_\{k\\leq d\}\\langle\\cdot,g\_\{k\}\\rangle\_\{L^\{2\}\(\\mu\_\{V\}\)\}\\,f\_\{k\},\(22\)sorank\(TS\)≤d\\rank\(T\_\{S\}\)\\leq d\. Conversely, ifrank\(TS\)=r≤d\\rank\(T\_\{S\}\)=r\\leq d, the Schmidt decomposition ofSSreadsS=∑k≤rσkak⊗bkS=\\sum\_\{k\\leq r\}\\sigma\_\{k\}\\,a\_\{k\}\\otimes b\_\{k\}withσk\>0\\sigma\_\{k\}\>0, and the encodersf=\(f1,…,fd\)f=\(f\_\{1\},\\dots,f\_\{d\}\),g=\(g1,…,gd\)g=\(g\_\{1\},\\dots,g\_\{d\}\)defined byfk=σkakf\_\{k\}=\\sigma\_\{k\}a\_\{k\}andgk=bkg\_\{k\}=b\_\{k\}fork≤rk\\leq r, andfk=gk=0f\_\{k\}=g\_\{k\}=0fork\>rk\>r, satisfyS=⟨f,g⟩S=\\langle f,g\\rangle\.
###### Corollary 1\(Spectral truncation lower bound\)\.
For every rank\-ddhead\(f,g\)\(f,g\),
‖F−⟨f,g⟩‖L2\(μU⊗μV\)2≥∑k\>dσk2,\\\|F\-\\langle f,g\\rangle\\\|\_\{L^\{2\}\(\\mu\_\{U\}\\otimes\\mu\_\{V\}\)\}^\{2\}\\;\\geq\\;\\sum\_\{k\>d\}\\sigma\_\{k\}^\{2\},\(23\)and the bound is attained inL2L^\{2\}by the truncated Schmidt seriesFd=∑k≤dσkak⊗bkF\_\{d\}=\\sum\_\{k\\leq d\}\\sigma\_\{k\}\\,a\_\{k\}\\otimes b\_\{k\}, which is the unique best rank\-ddapproximation wheneverσd\>σd\+1\\sigma\_\{d\}\>\\sigma\_\{d\+1\}\.
#### Proof of Corollary[1](https://arxiv.org/html/2608.11661#Thmcorollary1)\(Spectral truncation lower bound\)\.
The mapS↦TSS\\mapsto T\_\{S\}is an isometric isomorphism betweenL2\(μU⊗μV\)L^\{2\}\(\\mu\_\{U\}\\otimes\\mu\_\{V\}\)and the Hilbert\-Schmidt classHS\(L2\(μV\),L2\(μU\)\)\\mathrm\{HS\}\(L^\{2\}\(\\mu\_\{V\}\),L^\{2\}\(\\mu\_\{U\}\)\)equipped with the inner product⟨TS,TS′⟩=tr\(TS∗TS′\)\\langle T\_\{S\},T\_\{S^\{\\prime\}\}\\rangle=\\tr\(T\_\{S\}^\{\*\}T\_\{S^\{\\prime\}\}\); indeed‖TS‖HS2=∑k≥1σk2=‖S‖L22\\\|T\_\{S\}\\\|\_\{\\mathrm\{HS\}\}^\{2\}=\\sum\_\{k\\geq 1\}\\sigma\_\{k\}^\{2\}=\\\|S\\\|\_\{L^\{2\}\}^\{2\}\. Under this isometry, the rank constraint of Lemma[1](https://arxiv.org/html/2608.11661#Thmlemma1)becomes the operator rank constraint, and the Eckart\-Young\-Mirsky theorem for compact operators \(see e\.g\. Gohberg and Krein\) gives that the best rank\-ddapproximation ofTFT\_\{F\}is the truncation to its leadingddsingular triples, with squared error∑k\>dσk2\\sum\_\{k\>d\}\\sigma\_\{k\}^\{2\}\. Uniqueness: whenσd\>σd\+1\\sigma\_\{d\}\>\\sigma\_\{d\+1\}, the leadingdd\-dimensional singular subspace ofTFT\_\{F\}is unique, henceFd=∑k≤dσkak⊗bkF\_\{d\}=\\sum\_\{k\\leq d\}\\sigma\_\{k\}a\_\{k\}\\otimes b\_\{k\}is the unique minimizer\.
#### Proof of Theorem[1](https://arxiv.org/html/2608.11661#Thmtheorem1)\(Approximation error decomposition\)\.
Fix realizationsf^∈ℱ\\hat\{f\}\\in\\mathcal\{F\},g^∈𝒢\\hat\{g\}\\in\\mathcal\{G\}satisfying Assumption[1](https://arxiv.org/html/2608.11661#Thmassumption1), and writefk⋆=σkakf\_\{k\}^\{\\star\}=\\sigma\_\{k\}a\_\{k\},gk⋆=bkg\_\{k\}^\{\\star\}=b\_\{k\}for the ideal modes, so thatFd=∑k≤dfk⋆⊗gk⋆=⟨f⋆,g⋆⟩F\_\{d\}=\\sum\_\{k\\leq d\}f\_\{k\}^\{\\star\}\\otimes g\_\{k\}^\{\\star\}=\\langle f^\{\\star\},g^\{\\star\}\\rangle\. Setεk=fk⋆−f^k\\varepsilon\_\{k\}=f\_\{k\}^\{\\star\}\-\\hat\{f\}\_\{k\}andδk=gk⋆−g^k\\delta\_\{k\}=g\_\{k\}^\{\\star\}\-\\hat\{g\}\_\{k\}, with‖εk‖L2\(μU\)≤ηf,k\\\|\\varepsilon\_\{k\}\\\|\_\{L^\{2\}\(\\mu\_\{U\}\)\}\\leq\\eta\_\{f,k\}and‖δk‖L2\(μV\)≤ηg,k\\\|\\delta\_\{k\}\\\|\_\{L^\{2\}\(\\mu\_\{V\}\)\}\\leq\\eta\_\{g,k\}\. Split the error aroundFdF\_\{d\}:
‖F−⟨f^,g^⟩‖≤‖F−Fd‖⏟=ℰtrunc\+R,R:=‖Fd−⟨f^,g^⟩‖,\\\|F\-\\langle\\hat\{f\},\\hat\{g\}\\rangle\\\|\\leq\\underbrace\{\\\|F\-F\_\{d\}\\\|\}\_\{=\\sqrt\{\\mathcal\{E\}\_\{\\mathrm\{trunc\}\}\}\}\+R,\\qquad R:=\\\|F\_\{d\}\-\\langle\\hat\{f\},\\hat\{g\}\\rangle\\\|,\(24\)whereℰtrunc:=∑k\>dσk2\\mathcal\{E\}\_\{\\mathrm\{trunc\}\}:=\\sum\_\{k\>d\}\\sigma\_\{k\}^\{2\}by Corollary[1](https://arxiv.org/html/2608.11661#Thmcorollary1)\. For each coordinate, the product rule for tensor products gives
fk⋆⊗gk⋆−f^k⊗g^k=εk⊗bk\+σkak⊗δk−εk⊗δk\.f\_\{k\}^\{\\star\}\\otimes g\_\{k\}^\{\\star\}\-\\hat\{f\}\_\{k\}\\otimes\\hat\{g\}\_\{k\}=\\varepsilon\_\{k\}\\otimes b\_\{k\}\+\\sigma\_\{k\}a\_\{k\}\\otimes\\delta\_\{k\}\-\\varepsilon\_\{k\}\\otimes\\delta\_\{k\}\.\(25\)Sum overk≤dk\\leq dand take norms\. The first sum is diagonalized by the orthonormality of\{bk\}\\\{b\_\{k\}\\\}:
‖∑k≤dεk⊗bk‖2=∑k≤d‖εk‖L2\(μU\)2=∑k≤dηf,k2,\\Bigl\\\|\\sum\_\{k\\leq d\}\\varepsilon\_\{k\}\\otimes b\_\{k\}\\Bigr\\\|^\{2\}=\\sum\_\{k\\leq d\}\\\|\\varepsilon\_\{k\}\\\|\_\{L^\{2\}\(\\mu\_\{U\}\)\}^\{2\}=\\sum\_\{k\\leq d\}\\eta\_\{f,k\}^\{2\},\(26\)and likewise the orthonormality of\{ak\}\\\{a\_\{k\}\\\}diagonalizes the second sum:
‖∑k≤dσkak⊗δk‖2=∑k≤dσk2‖δk‖L2\(μV\)2≤∑k≤dσk2ηg,k2\.\\Bigl\\\|\\sum\_\{k\\leq d\}\\sigma\_\{k\}a\_\{k\}\\otimes\\delta\_\{k\}\\Bigr\\\|^\{2\}=\\sum\_\{k\\leq d\}\\sigma\_\{k\}^\{2\}\\\|\\delta\_\{k\}\\\|\_\{L^\{2\}\(\\mu\_\{V\}\)\}^\{2\}\\leq\\sum\_\{k\\leq d\}\\sigma\_\{k\}^\{2\}\\eta\_\{g,k\}^\{2\}\.\(27\)This is the step that avoids a factor ofdd: bounding the sums coordinatewise and applying Cauchy\-Schwarz over theddcoordinates would pay a factordd, whereas the orthonormality of the ideal modes makes the two sums orthogonal\. The third sum is bounded by Cauchy\-Schwarz:
‖∑k≤dεk⊗δk‖≤∑k≤d‖εk‖‖δk‖≤\(∑k≤dηf,k2\)1/2\(∑k≤dηg,k2\)1/2\.\\Bigl\\\|\\sum\_\{k\\leq d\}\\varepsilon\_\{k\}\\otimes\\delta\_\{k\}\\Bigr\\\|\\leq\\sum\_\{k\\leq d\}\\\|\\varepsilon\_\{k\}\\\|\\,\\\|\\delta\_\{k\}\\\|\\leq\\Bigl\(\\sum\_\{k\\leq d\}\\eta\_\{f,k\}^\{2\}\\Bigr\)^\{1/2\}\\Bigl\(\\sum\_\{k\\leq d\}\\eta\_\{g,k\}^\{2\}\\Bigr\)^\{1/2\}\.\(28\)WithX=\(∑kηf,k2\)1/2X=\(\\sum\_\{k\}\\eta\_\{f,k\}^\{2\}\)^\{1/2\},Y=\(∑kσk2ηg,k2\)1/2Y=\(\\sum\_\{k\}\\sigma\_\{k\}^\{2\}\\eta\_\{g,k\}^\{2\}\)^\{1/2\},Z=\(∑kηg,k2\)1/2Z=\(\\sum\_\{k\}\\eta\_\{g,k\}^\{2\}\)^\{1/2\}, the three bounds giveR≤X\+Y\+XZR\\leq X\+Y\+XZ, hence by\(x\+y\+z\)2≤3\(x2\+y2\+z2\)\(x\+y\+z\)^\{2\}\\leq 3\(x^\{2\}\+y^\{2\}\+z^\{2\}\),
R2≤3∑k≤dηf,k2\+3∑k≤dσk2ηg,k2\+3\(∑k≤dηf,k2\)\(∑k≤dηg,k2\)\.R^\{2\}\\leq 3\\sum\_\{k\\leq d\}\\eta\_\{f,k\}^\{2\}\+3\\sum\_\{k\\leq d\}\\sigma\_\{k\}^\{2\}\\eta\_\{g,k\}^\{2\}\+3\\Bigl\(\\sum\_\{k\\leq d\}\\eta\_\{f,k\}^\{2\}\\Bigr\)\\Bigl\(\\sum\_\{k\\leq d\}\\eta\_\{g,k\}^\{2\}\\Bigr\)\.\(29\)In the informative regime the mode errors are small: assuming∑k≤dηg,k2≤1\\sum\_\{k\\leq d\}\\eta\_\{g,k\}^\{2\}\\leq 1\(the modes are realized with bounded total error, the only regime in which the bound is informative\) andηg,k≤1/2\\eta\_\{g,k\}\\leq 1/2for allkk\(so thatBg=maxk‖g^k‖L2\(μV\)≥mink\(‖bk‖−‖δk‖\)≥1/2B\_\{g\}=\\max\_\{k\}\\\|\\hat\{g\}\_\{k\}\\\|\_\{L^\{2\}\(\\mu\_\{V\}\)\}\\geq\\min\_\{k\}\(\\\|b\_\{k\}\\\|\-\\\|\\delta\_\{k\}\\\|\)\\geq 1/2\), the cross term absorbs into the first term and∑kηf,k2≤4Bg2∑kηf,k2\\sum\_\{k\}\\eta\_\{f,k\}^\{2\}\\leq 4B\_\{g\}^\{2\}\\sum\_\{k\}\\eta\_\{f,k\}^\{2\}, giving
R2≤3∑k≤dσk2ηg,k2\+24Bg2∑k≤dηf,k2\.R^\{2\}\\leq 3\\sum\_\{k\\leq d\}\\sigma\_\{k\}^\{2\}\\eta\_\{g,k\}^\{2\}\+24B\_\{g\}^\{2\}\\sum\_\{k\\leq d\}\\eta\_\{f,k\}^\{2\}\.\(30\)Finally, since‖F−⟨f^,g^⟩‖≤ℰtrunc\+R\\\|F\-\\langle\\hat\{f\},\\hat\{g\}\\rangle\\\|\\leq\\sqrt\{\\mathcal\{E\}\_\{\\mathrm\{trunc\}\}\}\+R,
‖F−⟨f^,g^⟩‖2≤2ℰtrunc\+2R2≤2ℰtrunc\+6∑k≤dσk2ηg,k2\+48Bg2∑k≤dηf,k2\.\\\|F\-\\langle\\hat\{f\},\\hat\{g\}\\rangle\\\|^\{2\}\\leq 2\\mathcal\{E\}\_\{\\mathrm\{trunc\}\}\+2R^\{2\}\\leq 2\\mathcal\{E\}\_\{\\mathrm\{trunc\}\}\+6\\sum\_\{k\\leq d\}\\sigma\_\{k\}^\{2\}\\eta\_\{g,k\}^\{2\}\+48B\_\{g\}^\{2\}\\sum\_\{k\\leq d\}\\eta\_\{f,k\}^\{2\}\.\(31\)As2ℰtrunc≤48ℰtrunc2\\mathcal\{E\}\_\{\\mathrm\{trunc\}\}\\leq 48\\mathcal\{E\}\_\{\\mathrm\{trunc\}\}trivially, the right\-hand side is at most4848times the sum of the truncation term and the realization term of the theorem\. Taking the infimum overf∈ℱf\\in\\mathcal\{F\},g∈𝒢g\\in\\mathcal\{G\}, the bound holds withC≤48≤144C\\leq 48\\leq 144; we do not optimize the constant, which is what the statement of the theorem records\.
###### Theorem 3\(Smoothness controls the truncation term\)\.
Let𝒰⊂ℝmU\\mathcal\{U\}\\subset\\mathbb\{R\}^\{m\_\{U\}\}and𝒱⊂ℝmV\\mathcal\{V\}\\subset\\mathbb\{R\}^\{m\_\{V\}\}be bounded domains, and letμU,μV\\mu\_\{U\},\\mu\_\{V\}have densities bounded above and below with respect to Lebesgue measure\. IfFFisss\-smooth in its second argument,F∈Hs\(𝒰×𝒱\)F\\in H^\{s\}\(\\mathcal\{U\}\\times\\mathcal\{V\}\), then
σk=O\(k−s/mV\),∑k\>dσk2=O\(d1−2s/mV\)for2s\>mV\.\\sigma\_\{k\}=O\\bigl\(k^\{\-s/m\_\{V\}\}\\bigr\),\\qquad\\sum\_\{k\>d\}\\sigma\_\{k\}^\{2\}=O\\bigl\(d^\{1\-2s/m\_\{V\}\}\\bigr\)\\quad\\text\{for \}2s\>m\_\{V\}\.\(32\)IfFFis real analytic on𝒰×𝒱\\mathcal\{U\}\\times\\mathcal\{V\}, thenσk=O\(e−ck\)\\sigma\_\{k\}=O\(e^\{\-ck\}\)for somec\>0c\>0and∑k\>dσk2=O\(e−2cd\)\\sum\_\{k\>d\}\\sigma\_\{k\}^\{2\}=O\(e^\{\-2cd\}\)\.
#### Proof of Theorem[3](https://arxiv.org/html/2608.11661#Thmtheorem3)\(Smoothness controls the truncation term\)\.
The squared singular values ofTFT\_\{F\}are the eigenvalues of the self\-adjoint, positive semidefinite integral operatorTF∗TFT\_\{F\}^\{\*\}T\_\{F\}onL2\(μV\)L^\{2\}\(\\mu\_\{V\}\)with kernel
K\(v,v′\)=∫𝒰F\(u,v\)F\(u,v′\)dμU\(u\),K\(v,v^\{\\prime\}\)=\\int\_\{\\mathcal\{U\}\}F\(u,v\)\\,F\(u,v^\{\\prime\}\)\\,d\\mu\_\{U\}\(u\),\(33\)which is positive definite and inherits the smoothness ofFFin its second argument: differentiating under the integral, anss\-smooth kernel on a boundedmVm\_\{V\}\-dimensional domain has eigenvalues decaying asO\(k−2s/mV\)O\(k^\{\-2s/m\_\{V\}\}\); this is the classical eigenvalue decay for positive definite kernels\([18](https://arxiv.org/html/2608.11661#bib.bib19)\), and the analytic case decays exponentially\([13](https://arxiv.org/html/2608.11661#bib.bib20)\)\. Taking square roots givesσk=O\(k−s/mV\)\\sigma\_\{k\}=O\(k^\{\-s/m\_\{V\}\}\), and summing the tail,
∑k\>dσk2=O\(∫d∞x−2s/mVdx\)=O\(d1−2s/mV\)for2s\>mV,\\sum\_\{k\>d\}\\sigma\_\{k\}^\{2\}=O\\Bigl\(\\int\_\{d\}^\{\\infty\}x^\{\-2s/m\_\{V\}\}\\,dx\\Bigr\)=O\\bigl\(d^\{1\-2s/m\_\{V\}\}\\bigr\)\\quad\\text\{for \}2s\>m\_\{V\},\(34\)which is finite exactly when2s\>mV2s\>m\_\{V\}\. In the analytic case,σk=O\(e−ck\)\\sigma\_\{k\}=O\(e^\{\-ck\}\)gives∑k\>dσk2=O\(e−2cd\)\\sum\_\{k\>d\}\\sigma\_\{k\}^\{2\}=O\(e^\{\-2cd\}\)by the geometric series\.
###### Proposition 2\(Approximation separation\)\.
LetF∈ℳdF\\in\\mathcal\{M\}\_\{d\}with interaction modesak∈Hs\(𝒰\)a\_\{k\}\\in H^\{s\}\(\\mathcal\{U\}\),bk∈Hs\(𝒱\)b\_\{k\}\\in H^\{s\}\(\\mathcal\{V\}\)\. To approximateFFwithinε\\varepsiloninL2\(μU⊗μV\)L^\{2\}\(\\mu\_\{U\}\\otimes\\mu\_\{V\}\), a dual\-encoder head whose encoders are ReLU networks requires
Ndual=O\(d\(ε−mU/s\+ε−mV/s\)\)N\_\{\\mathrm\{dual\}\}=O\\Bigl\(d\\bigl\(\\varepsilon^\{\-m\_\{U\}/s\}\+\\varepsilon^\{\-m\_\{V\}/s\}\\bigr\)\\Bigr\)\(35\)parameters, whereas any single\-tower model that approximatesFFas a general\(mU\+mV\)\(m\_\{U\}\+m\_\{V\}\)\-dimensionalHsH^\{s\}function requiresΩ\(ε−\(mU\+mV\)/s\)\\Omega\(\\varepsilon^\{\-\(m\_\{U\}\+m\_\{V\}\)/s\}\)parameters in the worst case\. HenceNdual/Nsingle→0N\_\{\\mathrm\{dual\}\}/N\_\{\\mathrm\{single\}\}\\to 0asε→0\\varepsilon\\to 0\.
#### Proof of Proposition[2](https://arxiv.org/html/2608.11661#Thmproposition2)\(Approximation separation\)\.
SinceF∈ℳdF\\in\\mathcal\{M\}\_\{d\}with modesak∈Hs\(𝒰\)a\_\{k\}\\in H^\{s\}\(\\mathcal\{U\}\),bk∈Hs\(𝒱\)b\_\{k\}\\in H^\{s\}\(\\mathcal\{V\}\), the truncation term of Theorem[1](https://arxiv.org/html/2608.11661#Thmtheorem1)vanishes for the rank\-ddhead, and the entire approximation error is the realization term\. Approximate each of the2d2dmodes by ReLU networks with the standardHsH^\{s\}rates\([25](https://arxiv.org/html/2608.11661#bib.bib21)\): to reach mode errorε\\varepsilonon a marginal space of dimensionmmrequiresO\(ε−m/s\)O\(\\varepsilon^\{\-m/s\}\)parameters\. Setting the mode errors toε\\varepsilonmakes the realization term of Theorem[1](https://arxiv.org/html/2608.11661#Thmtheorem1)of orderO\(ε2\)O\(\\varepsilon^\{2\}\), so a total of
Ndual=O\(d\(ε−mU/s\+ε−mV/s\)\)N\_\{\\mathrm\{dual\}\}=O\\Bigl\(d\\bigl\(\\varepsilon^\{\-m\_\{U\}/s\}\+\\varepsilon^\{\-m\_\{V\}/s\}\\bigr\)\\Bigr\)\(36\)parameters suffice\. For the lower bound, a single\-tower model that treatsFFas a general\(mU\+mV\)\(m\_\{U\}\+m\_\{V\}\)\-dimensionalHsH^\{s\}function must contend with the Kolmogorov width of the Sobolev class, whoseL2L^\{2\}approximation is bounded below byΩ\(ε−\(mU\+mV\)/s\)\\Omega\(\\varepsilon^\{\-\(m\_\{U\}\+m\_\{V\}\)/s\}\)parameters in the worst case \(classical Sobolev width results\)\. SincemU\+mV\>mU,mVm\_\{U\}\+m\_\{V\}\>m\_\{U\},m\_\{V\}, the exponent is additive in the dimensions, soNdual/Nsingle=O\(εmmin/s\)→0N\_\{\\mathrm\{dual\}\}/N\_\{\\mathrm\{single\}\}=O\\bigl\(\\varepsilon^\{m\_\{\\min\}/s\}\\bigr\)\\to 0asε→0\\varepsilon\\to 0, wheremmin=min\{mU,mV\}m\_\{\\min\}=\\min\\\{m\_\{U\},m\_\{V\}\\\}\.
#### Proof of Proposition[1](https://arxiv.org/html/2608.11661#Thmproposition1)\(Compute separation\)\.
Counting operations: the head requiresnnforward passes offf,mmforward passes ofgg, andnmnminner products of dimensiondd, i\.e\.,nmdnmdmultiply\-adds, for a total ofO\(\(n\+m\)Cnet\+nmd\)O\(\(n\+m\)C\_\{\\mathrm\{net\}\}\+nmd\)\. A joint model processes each of thenmnmpairs separately, costingO\(nmCnet\)O\(nm\\,C\_\{\\mathrm\{net\}\}\)\. WhenCnet≫dC\_\{\\mathrm\{net\}\}\\gg d, the head is cheaper by a factor of orderCnet/dC\_\{\\mathrm\{net\}\}/din the regime wherenm≫n\+mnm\\gg n\+m\.
## Appendix CProofs for Section 3
We prove the results of Section[3](https://arxiv.org/html/2608.11661#S3)in order\. Throughout,ℒ\(f,g\)=𝔼u,v\[\(F\(u,v\)−⟨f\(u\),g\(v\)⟩\)2\]\\mathcal\{L\}\(f,g\)=\\mathbb\{E\}\_\{u,v\}\[\(F\(u,v\)\-\\langle f\(u\),g\(v\)\\rangle\)^\{2\}\]is the population risk\. Table[6](https://arxiv.org/html/2608.11661#A3.T6)collects the normalization taxonomy moved from Section[3](https://arxiv.org/html/2608.11661#S3); Proposition[3](https://arxiv.org/html/2608.11661#Thmproposition3)and Remark[4](https://arxiv.org/html/2608.11661#Thmremark4)complete the picture\.
Table 6:Embedding\-level normalizations as sections of the gauge orbit\. Each row shows a constraint on the encoders, the part of the gauge it pins, the residual group that survives, and what the scheme identifies\. See Proposition[3](https://arxiv.org/html/2608.11661#Thmproposition3)and Theorems[4](https://arxiv.org/html/2608.11661#Thmtheorem4)to[2](https://arxiv.org/html/2608.11661#Thmtheorem2)for the statements, and Appendix[C](https://arxiv.org/html/2608.11661#A3)for proofs\.The gauge groupGLd\\mathrm\{GL\}\_\{d\}stratifies into per\-coordinate scalings \(D\+\\mathrm\{D\}\_\{\+\},dddimensions\), rotations \(O\\mathrm\{O\},d\(d−1\)/2d\(d\-1\)/2\), permutations and signs \(discrete\), and the remaining shears\. The guarantees of Section[3](https://arxiv.org/html/2608.11661#S3)come from two\-sided constraints, which act on both encoders and can pin the joint freedom of the pair; one\-sided constraints cannot do this in general, and the schemes that use them are best understood as stability fixes\. Normalization schemes are therefore compared on four measurable quantities: the residual gauge subgroup of Table[6](https://arxiv.org/html/2608.11661#A3.T6), the Hessian condition number on the constrained manifold, related to the gap by Corollary[3](https://arxiv.org/html/2608.11661#Thmcorollary3), the expressivity cost of the constraint, for example the bounded range\[−1,1\]\[\-1,1\]of cosine normalization, and the output semantics, which determines which decoder and loss can be attached\.
###### Proposition 3\(Per\-coordinate standardization pins the scales\)\.
Consider the constraintdiag\(Σg\)=I\\operatorname\{diag\}\(\\Sigma\_\{g\}\)=I, which standardizes each coordinate of one encoder to unit variance\. It pins the positive diagonal subgroupD\+\\mathrm\{D\}\_\{\+\}of dimensiondd; generically the residual gauge has dimensiond2−dd^\{2\}\-dand contains the shears and rotations that mix coordinates\.
###### Lemma 2\(Gauge invariance\)\.
For everyA∈GLdA\\in\\mathrm\{GL\}\_\{d\}and every pair of inputs\(u,v\)\(u,v\),⟨Af\(u\),A−⊤g\(v\)⟩=⟨f\(u\),g\(v\)⟩\\langle Af\(u\),A^\{\-\\top\}g\(v\)\\rangle=\\langle f\(u\),g\(v\)\\rangle, soΦA\\Phi\_\{A\}leaves the represented function and hence the population risk unchanged\. With the second momentsΣf=𝔼u\[f\(u\)f\(u\)⊤\]\\Sigma\_\{f\}=\\mathbb\{E\}\_\{u\}\[f\(u\)f\(u\)^\{\\top\}\]andΣg=𝔼v\[g\(v\)g\(v\)⊤\]\\Sigma\_\{g\}=\\mathbb\{E\}\_\{v\}\[g\(v\)g\(v\)^\{\\top\}\], the action transforms them byΣAf=AΣfA⊤\\Sigma\_\{Af\}=A\\Sigma\_\{f\}A^\{\\top\}andΣA−⊤g=A−⊤ΣgA−1\\Sigma\_\{A^\{\-\\top\}g\}=A^\{\-\\top\}\\Sigma\_\{g\}A^\{\-1\}, so the eigenvalues of the productΣfΣg\\Sigma\_\{f\}\\Sigma\_\{g\}are gauge invariants\.
#### Proof of Lemma[2](https://arxiv.org/html/2608.11661#Thmlemma2)\(Gauge invariance\)\.
The identity is the cancellation
⟨Af\(u\),A−⊤g\(v\)⟩=f\(u\)⊤A⊤A−⊤g\(v\)=f\(u\)⊤g\(v\)=⟨f\(u\),g\(v\)⟩,\\langle Af\(u\),A^\{\-\\top\}g\(v\)\\rangle=f\(u\)^\{\\top\}A^\{\\top\}A^\{\-\\top\}g\(v\)=f\(u\)^\{\\top\}g\(v\)=\\langle f\(u\),g\(v\)\\rangle,\(37\)so the represented function and hence the risk are unchanged\. For the second moments, substitution gives
ΣAf=𝔼u\[Af\(u\)f\(u\)⊤A⊤\]=AΣfA⊤,ΣA−⊤g=A−⊤ΣgA−1,\\Sigma\_\{Af\}=\\mathbb\{E\}\_\{u\}\[Af\(u\)f\(u\)^\{\\top\}A^\{\\top\}\]=A\\Sigma\_\{f\}A^\{\\top\},\\qquad\\Sigma\_\{A^\{\-\\top\}g\}=A^\{\-\\top\}\\Sigma\_\{g\}A^\{\-1\},\(38\)and the product transforms by the similarityΣAfΣA−⊤g=AΣfΣgA−1\\Sigma\_\{Af\}\\Sigma\_\{A^\{\-\\top\}g\}=A\\,\\Sigma\_\{f\}\\Sigma\_\{g\}\\,A^\{\-1\}, whose eigenvalues are invariant\. At a global minimizer, the represented function equals the rank\-ddtruncationFdF\_\{d\}of the target whenσd\>σd\+1\\sigma\_\{d\}\>\\sigma\_\{d\+1\}\(Corollary[1](https://arxiv.org/html/2608.11661#Thmcorollary1)\), and the coefficient computation of the proof of Theorem[2](https://arxiv.org/html/2608.11661#Thmtheorem2)below shows that the eigenvalues ofΣfΣg\\Sigma\_\{f\}\\Sigma\_\{g\}are then the squared singular valuesσk2\\sigma\_\{k\}^\{2\}, which is the claim of Remark[5](https://arxiv.org/html/2608.11661#Thmremark5)\.
###### Proposition 4\(Ill\-posedness from unfixed gauge\)\.
Let\(f⋆,g⋆\)\(f^\{\\star\},g^\{\\star\}\)be a global minimizer of the risk and letSSbe its stabilizer, the set ofA∈GLdA\\in\\mathrm\{GL\}\_\{d\}withΦA\(f⋆,g⋆\)=\(f⋆,g⋆\)\\Phi\_\{A\}\(f^\{\\star\},g^\{\\star\}\)=\(f^\{\\star\},g^\{\\star\}\)\. Then the Hessian of the risk at\(f⋆,g⋆\)\(f^\{\\star\},g^\{\\star\}\)has at least
dimker∇2ℒ\(f⋆,g⋆\)≥d2−dimS\\dim\\ker\\nabla^\{2\}\\mathcal\{L\}\(f^\{\\star\},g^\{\\star\}\)\\;\\geq\\;d^\{2\}\-\\dim S\(39\)zero directions, coming from the tangent space of the gauge orbit\. Without normalization this is genericallyd2d^\{2\}degenerate directions, so the second\-order problem is rank\-deficient and the optimization problem is ill\-posed\.
#### Proof of Proposition[4](https://arxiv.org/html/2608.11661#Thmproposition4)\(Ill\-posedness from unfixed gauge\)\.
Let\(f⋆,g⋆\)\(f^\{\\star\},g^\{\\star\}\)be a global minimizer andSSits stabilizer\. By Lemma[2](https://arxiv.org/html/2608.11661#Thmlemma2), the risk is constant along the orbit𝒪=\{ΦA\(f⋆,g⋆\):A∈GLd\}\\mathcal\{O\}=\\\{\\Phi\_\{A\}\(f^\{\\star\},g^\{\\star\}\):A\\in\\mathrm\{GL\}\_\{d\}\\\}, so for everyX∈𝔤𝔩dX\\in\\mathfrak\{gl\}\_\{d\}the curveA\(t\)=exp\(tX\)A\(t\)=\\exp\(tX\)gives a pathγX\(t\)=ΦA\(t\)\(f⋆,g⋆\)\\gamma\_\{X\}\(t\)=\\Phi\_\{A\(t\)\}\(f^\{\\star\},g^\{\\star\}\)along whichℒ∘γX≡ℒ\(f⋆,g⋆\)\\mathcal\{L\}\\circ\\gamma\_\{X\}\\equiv\\mathcal\{L\}\(f^\{\\star\},g^\{\\star\}\)\. The velocity is
γ˙X\(0\)=\(Xf⋆,−X⊤g⋆\)∈T\(f⋆,g⋆\),\\dot\{\\gamma\}\_\{X\}\(0\)=\(Xf^\{\\star\},\\,\-X^\{\\top\}g^\{\\star\}\)\\in T\_\{\(f^\{\\star\},g^\{\\star\}\)\},\(40\)and since the risk is constant along the path, its second derivative in this direction vanishes\. Because\(f⋆,g⋆\)\(f^\{\\star\},g^\{\\star\}\)is a global minimizer,∇2ℒ⪰0\\nabla^\{2\}\\mathcal\{L\}\\succeq 0, and a semidefinite quadratic form vanishes in a direction if and only if that direction lies in the kernel, soγ˙X\(0\)∈ker∇2ℒ\\dot\{\\gamma\}\_\{X\}\(0\)\\in\\ker\\nabla^\{2\}\\mathcal\{L\}for everyXX\. The mapX↦γ˙X\(0\)X\\mapsto\\dot\{\\gamma\}\_\{X\}\(0\)has kernel equal to the Lie algebra of the stabilizerSS, hence its image, which lies entirely in the kernel of the Hessian, has dimensiond2−dimSd^\{2\}\-\\dim S\. This proves the bound\. Without normalization the stabilizer at a generic minimizer is trivial, so the Hessian has at leastd2d^\{2\}zero directions\.
#### Proof of Proposition[3](https://arxiv.org/html/2608.11661#Thmproposition3)\(Per\-coordinate standardization pins the scales\)\.
The constraintdiag\(Σg\)=I\\operatorname\{diag\}\(\\Sigma\_\{g\}\)=Iisddscalar conditions that fix theddcoordinate scales ofgg: for the transformationA−1=diag\(s1,…,sd\)∈D\+A^\{\-1\}=\\operatorname\{diag\}\(s\_\{1\},\\dots,s\_\{d\}\)\\in\\mathrm\{D\}\_\{\+\}acting ongg, the condition forcessk2\[Σg\]kk=1s\_\{k\}^\{2\}\[\\Sigma\_\{g\}\]\_\{kk\}=1, pinning eachsks\_\{k\}\. TransformationsAAthat mix coordinates generically changediag\(A−⊤ΣgA−1\)\\operatorname\{diag\}\(A^\{\-\\top\}\\Sigma\_\{g\}A^\{\-1\}\)unless they also rotateΣg\\Sigma\_\{g\}to a new matrix with the same diagonal; the residual group is the set ofAApreserving the diagonal condition, which contains the orthogonal group conjugated by any diagonal scale and the shears that preserve the diagonal ofΣg\\Sigma\_\{g\}\. Counting dimensions, the constraint fixes exactly thedd\-dimensional subgroupD\+\\mathrm\{D\}\_\{\+\}, and the generic residual dimension isd2−dd^\{2\}\-d\.
###### Theorem 4\(Cosine normalization leaves a full rotation\)\.
Consider two\-sided unit normalization,f~=f/‖f‖\\tilde\{f\}=f/\\\|f\\\|andg~=g/‖g‖\\tilde\{g\}=g/\\\|g\\\|, with scorescos∠\(f\(u\),g\(v\)\)\\cos\\angle\(f\(u\),g\(v\)\), as in the contrastive models of Table[5](https://arxiv.org/html/2608.11661#A1.T5)\([17](https://arxiv.org/html/2608.11661#bib.bib1)\); a temperature that rescales the logits changes no gauge\. If the encodings are nondegenerate, in that the ranges offfandggcontain open subsets ofℝd\\mathbb\{R\}^\{d\}, the residual gauge group is exactly the orthogonal group,
R=\{A∈GLd:A⊤A=I\}=O\.R=\\bigl\\\{\\,A\\in\\mathrm\{GL\}\_\{d\}:\\;A^\{\\top\}A=I\\,\\bigr\\\}=\\mathrm\{O\}\.\(41\)For every rotationR∈OR\\in\\mathrm\{O\}, the encodingsRfRfandRgRggive pointwise identical cosine scores, and no transformation outsideO\\mathrm\{O\}does\.
###### Corollary 2\(Contrastive dimensions are uninterpretable\)\.
Under the residual symmetry of Theorem[4](https://arxiv.org/html/2608.11661#Thmtheorem4), the individual coordinates of a cosine\-normalized embedding carry no intrinsic meaning: replacing\(f,g\)\(f,g\)by\(Rf,Rg\)\(Rf,Rg\)for any rotationRRyields the same model and the same loss, and only rotation\-invariant quantities are well defined, among them norms, pairwise inner products, and the eigenvalues ofΣfΣg\\Sigma\_\{f\}\\Sigma\_\{g\}, which by Remark[5](https://arxiv.org/html/2608.11661#Thmremark5)are the squared interaction singular values at any optimum\. This is a theorem\-level explanation of the widely reported observation that contrastive embedding dimensions are uninterpretable; the remedy is the whitening scheme of Theorem[2](https://arxiv.org/html/2608.11661#Thmtheorem2)\.
#### Proof of Theorem[4](https://arxiv.org/html/2608.11661#Thmtheorem4)\(Cosine normalization leaves a full rotation\)\.
LetRRbe the residual group of the two\-sided unit normalization, i\.e\., the set ofA∈GLdA\\in\\mathrm\{GL\}\_\{d\}such that, after renormalizing, the cosine scores are unchanged\.
\(⊇\)\(\\supseteq\)ForR∈OR\\in\\mathrm\{O\},‖Rf\(u\)‖=‖f\(u\)‖\\\|Rf\(u\)\\\|=\\\|f\(u\)\\\|and‖Rg\(v\)‖=‖g\(v\)‖\\\|Rg\(v\)\\\|=\\\|g\(v\)\\\|by orthogonality, and⟨Rf\(u\),Rg\(v\)⟩=⟨f\(u\),g\(v\)⟩\\langle Rf\(u\),Rg\(v\)\\rangle=\\langle f\(u\),g\(v\)\\rangle, hence
cos∠\(Rf\(u\),Rg\(v\)\)=⟨Rf\(u\),Rg\(v\)⟩‖Rf\(u\)‖‖Rg\(v\)‖=⟨f\(u\),g\(v\)⟩‖f\(u\)‖‖g\(v\)‖\\cos\\angle\(Rf\(u\),Rg\(v\)\)=\\frac\{\\langle Rf\(u\),Rg\(v\)\\rangle\}\{\\\|Rf\(u\)\\\|\\,\\\|Rg\(v\)\\\|\}=\\frac\{\\langle f\(u\),g\(v\)\\rangle\}\{\\\|f\(u\)\\\|\\,\\\|g\(v\)\\\|\}\(42\)pointwise, soO⊆R\\mathrm\{O\}\\subseteq R\.
\(⊆\)\(\\subseteq\)SupposeAApreserves every cosine between the encodings\. Since the ranges offfandggcontain open subsets ofℝd\\mathbb\{R\}^\{d\}, the encodings realize a full\-dimensional family of directions, soAApreserves all angles between vectors\. A linear map that preserves every angle is a similarity: by the classical characterization of conformal linear maps,A=cRA=cRfor somec\>0c\>0andR∈OR\\in\\mathrm\{O\}\. The scalarccis absorbed by the unit renormalization of both encoders, so within the cosine scoreAAacts identically to the rotationRR, givingR⊆OR\\subseteq\\mathrm\{O\}\.
#### Corollary[2](https://arxiv.org/html/2608.11661#Thmcorollary2)follows immediately\.
For any rotationR∈OR\\in\\mathrm\{O\}the pair\(Rf,Rg\)\(Rf,Rg\)is a different encoder pair with pointwise identical cosine scores and identical loss, so the individual coordinates carry no intrinsic meaning\. The rotation\-invariant quantities are the norms, the inner products, and the eigenvalues ofΣfΣg\\Sigma\_\{f\}\\Sigma\_\{g\}; by Remark[5](https://arxiv.org/html/2608.11661#Thmremark5)these are the squared interaction singular values at any optimum\. Corollary[2](https://arxiv.org/html/2608.11661#Thmcorollary2)is the statement of these two sentences at the theorem level\.
###### Theorem 5\(Two\-sided nonnegativity gives NMF identifiability\)\.
Suppose both encoders are constrained to nonnegative outputs,f\(u\)⪰0f\(u\)\\succeq 0andg\(v\)⪰0g\(v\)\\succeq 0, as obtained by applying a softplus or exponential nonlinearity to raw encodings\. On a finite set of inputs the model is the nonnegative matrix factorizationM=FG⊤M=FG^\{\\top\}of a nonnegative matrixMM\. If the encodings are nondegenerate, in that their ranges generate the positive orthant, the residual gauge group shrinks to the monomial group,
R=\{A∈GLd:A⪰0,A−⊤⪰0\}=Perm⋉D\+,R=\\bigl\\\{\\,A\\in\\mathrm\{GL\}\_\{d\}:\\;A\\succeq 0,\\;A^\{\-\\top\}\\succeq 0\\,\\bigr\\\}=\\mathrm\{Perm\}\\ltimes\\mathrm\{D\}\_\{\+\},\(43\)the group of coordinate permutations composed with positive diagonal scalings\. If the factorization is separable in the sense of[6](https://arxiv.org/html/2608.11661#bib.bib11), with an anchor row for each latent factor as in[1](https://arxiv.org/html/2608.11661#bib.bib12), then it is essentially unique and the encoders are identified up to permutation and scaling\.
#### Proof of Theorem[5](https://arxiv.org/html/2608.11661#Thmtheorem5)\(Two\-sided nonnegativity gives NMF identifiability\)\.
On a finite set of inputs, the nonnegative encoders form matricesF∈ℝ≥0n×dF\\in\\mathbb\{R\}\_\{\\geq 0\}^\{n\\times d\},G∈ℝ≥0m×dG\\in\\mathbb\{R\}\_\{\\geq 0\}^\{m\\times d\}, and the model is the nonnegative factorizationM=FG⊤M=FG^\{\\top\}\. The residual group consists ofA∈GLdA\\in\\mathrm\{GL\}\_\{d\}such thatAf⪰0Af\\succeq 0andA−⊤g⪰0A^\{\-\\top\}g\\succeq 0for all nonnegative encodings\. Since the ranges generate the positive orthant,Af⪰0Af\\succeq 0for all nonnegativeffforcesA⪰0A\\succeq 0entrywise, and likewiseA−⊤⪰0A^\{\-\\top\}\\succeq 0\. A real matrix that is nonnegative and whose inverse transpose is also nonnegative is a generalized permutation matrix: its inverse transpose being nonnegative forces each row and column ofAAto have exactly one nonzero entry, and nonnegativity makes that entry positive\. HenceR⊆Perm⋉D\+R\\subseteq\\mathrm\{Perm\}\\ltimes\\mathrm\{D\}\_\{\+\}, and the reverse inclusion is immediate, soR=Perm⋉D\+R=\\mathrm\{Perm\}\\ltimes\\mathrm\{D\}\_\{\+\}\. Under separability in the sense of[6](https://arxiv.org/html/2608.11661#bib.bib11), each latent factor has an anchor row whose encoding is nonzero in exactly one coordinate, which identifies the extreme rays of the latent cone and hence the factor directions; the constructive algorithm of[1](https://arxiv.org/html/2608.11661#bib.bib12)then yields the decomposition uniquely up to permutation and scaling\. The degradation of this guarantee when separability holds only approximately is analyzed in Appendix[E](https://arxiv.org/html/2608.11661#A5)\.
#### Proof of Theorem[2](https://arxiv.org/html/2608.11661#Thmtheorem2)\(Whitening fixes the gauge up to permutation and sign\)\.
The proof has four steps\.
\(1\) Every global minimizer represents the rank\-ddtruncation\. By Corollary[1](https://arxiv.org/html/2608.11661#Thmcorollary1)the minimal risk is∑k\>dσk2\\sum\_\{k\>d\}\\sigma\_\{k\}^\{2\}, attained exactly by the truncated seriesFd=∑k≤dσkak⊗bkF\_\{d\}=\\sum\_\{k\\leq d\}\\sigma\_\{k\}a\_\{k\}\\otimes b\_\{k\}, which is unique sinceσd\>σd\+1\\sigma\_\{d\}\>\\sigma\_\{d\+1\}\. Hence any minimizer satisfies⟨f\(u\),g\(v\)⟩=Fd\(u,v\)\\langle f\(u\),g\(v\)\\rangle=F\_\{d\}\(u,v\)almost everywhere, i\.e\.,
∑k=1dfk\(u\)gk\(v\)=∑k=1dσkak\(u\)bk\(v\)\.\\sum\_\{k=1\}^\{d\}f\_\{k\}\(u\)\\,g\_\{k\}\(v\)=\\sum\_\{k=1\}^\{d\}\\sigma\_\{k\}\\,a\_\{k\}\(u\)\\,b\_\{k\}\(v\)\.\(44\)
\(2\) The span of the encoders is pinned\. As rank\-ddrepresentations of the same tensor, the left factor spaces and right factor spaces coincide:span\{f1,…,fd\}=span\{a1,…,ad\}\\mathrm\{span\}\\\{f\_\{1\},\\dots,f\_\{d\}\\\}=\\mathrm\{span\}\\\{a\_\{1\},\\dots,a\_\{d\}\\\}andspan\{g1,…,gd\}=span\{b1,…,bd\}\\mathrm\{span\}\\\{g\_\{1\},\\dots,g\_\{d\}\\\}=\\mathrm\{span\}\\\{b\_\{1\},\\dots,b\_\{d\}\\\}\. So there exist coefficient matricesB,C∈ℝd×dB,C\\in\\mathbb\{R\}^\{d\\times d\}withf=Ba\[d\]f=B\\,a\_\{\[d\]\}andg=Cb\[d\]g=C\\,b\_\{\[d\]\}, wherea\[d\]=\(a1,…,ad\)⊤a\_\{\[d\]\}=\(a\_\{1\},\\dots,a\_\{d\}\)^\{\\top\}and likewise forbb\. Substituting into the identity above and using the orthonormality of the modes gives
B⊤C=diag\(σ1,…,σd\)=:S\.B^\{\\top\}C=\\operatorname\{diag\}\(\\sigma\_\{1\},\\dots,\\sigma\_\{d\}\)=:S\.\(45\)
\(3\) The whitening constraints forceCCto be a signed permutation\. By the orthonormality of\{ak\}\\\{a\_\{k\}\\\},
Σf=𝔼u\[f\(u\)f\(u\)⊤\]=BB⊤,Σg=CC⊤\.\\Sigma\_\{f\}=\\mathbb\{E\}\_\{u\}\[f\(u\)f\(u\)^\{\\top\}\]=B\\,B^\{\\top\},\\qquad\\Sigma\_\{g\}=C\\,C^\{\\top\}\.\(46\)The constraintΣg=I\\Sigma\_\{g\}=IgivesCC⊤=ICC^\{\\top\}=I, soC∈OC\\in\\mathrm\{O\}; the constraintΣf=Λ\\Sigma\_\{f\}=\\Lambdadiagonal givesBB⊤=ΛBB^\{\\top\}=\\Lambda\. FromB⊤C=SB^\{\\top\}C=SandC∈OC\\in\\mathrm\{O\}we getB⊤=SC−1=SC⊤B^\{\\top\}=SC^\{\-1\}=SC^\{\\top\}, soB=CSB=CS, and therefore
BB⊤=CS2C⊤=Λ\.BB^\{\\top\}=C\\,S^\{2\}\\,C^\{\\top\}=\\Lambda\.\(47\)Thus the orthogonal matrixCCdiagonalizes the diagonal matrixS2=diag\(σ12,…,σd2\)S^\{2\}=\\operatorname\{diag\}\(\\sigma\_\{1\}^\{2\},\\dots,\\sigma\_\{d\}^\{2\}\)\. Since theσk\\sigma\_\{k\}are distinct, the eigenspaces ofS2S^\{2\}are the coordinate axes, and any orthogonal matrix diagonalizingS2S^\{2\}maps each axis to itself up to sign, soC∈Perm⋉D±C\\in\\mathrm\{Perm\}\\ltimes\\mathrm\{D\}\_\{\\pm\}andΛ=S2\\Lambda=S^\{2\}, i\.e\.,λk=σk2\\lambda\_\{k\}=\\sigma\_\{k\}^\{2\}with the ordering convention matching the decreasingσk\\sigma\_\{k\}\. With the ordering fixed,C=diag\(±1,…,±1\)C=\\operatorname\{diag\}\(\\pm 1,\\dots,\\pm 1\)andB=CS=diag\(±σ1,…,±σd\)B=CS=\\operatorname\{diag\}\(\\pm\\sigma\_\{1\},\\dots,\\pm\\sigma\_\{d\}\)\.
\(4\) Substitute back\. With the coordinate order fixed,gk=±bkg\_\{k\}=\\pm b\_\{k\}andfk=±σkakf\_\{k\}=\\pm\\sigma\_\{k\}a\_\{k\}, the signs being paired so that the productfkgk=σkakbkf\_\{k\}g\_\{k\}=\\sigma\_\{k\}a\_\{k\}b\_\{k\}is unique\. This is exactly \([10](https://arxiv.org/html/2608.11661#S3.E10)\), and the encoders are identified up to permutation and sign\.
###### Corollary 3\(The spectral gap controls conditioning and convergence\)\.
On the manifold defined by the constraints \([9](https://arxiv.org/html/2608.11661#S3.E9)\), every global minimizer is isolated up to the discrete group of signs and permutations, and the Hessian of the risk is nondegenerate in the tangent directions of the manifold, with smallest nonzero curvature bounded below by a constant multiple of the spectral gapΔ=min1≤k<d\(σk2−σk\+12\)\\Delta=\\min\_\{1\\leq k<d\}\(\\sigma\_\{k\}^\{2\}\-\\sigma\_\{k\+1\}^\{2\}\); a projected gradient method that maintains the constraints therefore converges locally at a linear rate1−μ/L1\-\\mu/L, whereμ\\muis bounded below by a constant multiple ofΔ\\DeltaandLLis the local smoothness constant\. Wider gaps make the problem both better conditioned and faster to converge\.
#### Proof of Corollary[3](https://arxiv.org/html/2608.11661#Thmcorollary3)\(The spectral gap controls conditioning and convergence\)\.
Isolation: by Theorem[2](https://arxiv.org/html/2608.11661#Thmtheorem2), on the manifold defined by \([9](https://arxiv.org/html/2608.11661#S3.E9)\) every global minimizer is unique up to the discrete group of signs and the permutation, and the represented functionFdF\_\{d\}is a strict minimum of the risk by the strict inequalityσd\>σd\+1\\sigma\_\{d\}\>\\sigma\_\{d\+1\}; the Hessian of the risk is therefore nondegenerate in the tangent directions of the manifold, and its smallest nonzero curvature is positive\. The quantitative bound: the Hessian curvature in the tangent directions is controlled by the second derivative of the risk as the encoders move within the constraint manifold, and by the Davis\-KahansinΘ\\sin\\Thetatheorem the perturbation of a leading singular subspace under a perturbation of sizeϵ\\epsilonis at mostϵ/Δ\\epsilon/\\Delta, whereΔ=mink\(σk2−σk\+12\)\\Delta=\\min\_\{k\}\(\\sigma\_\{k\}^\{2\}\-\\sigma\_\{k\+1\}^\{2\}\)\. Inverting this stability relation, the curvature of the risk in directions that move the encoders off the ideal modes is bounded below by a constant multiple ofΔ\\Delta: a displacement of sizettin function space changes the risk by at leastcΔt2c\\,\\Delta\\,t^\{2\}for the truncation\-optimal function\. A projected gradient method that maintains the constraints therefore converges locally at a linear rate1−μ/L1\-\\mu/Lwithμ≥cΔ\\mu\\geq c\\,\\DeltaandLLthe local smoothness constant, as stated\.
###### Theorem 6\(Soft whitening recovers the hard solution\)\.
Forρ\>0\\rho\>0, consider the penalized risk
ℒρ\(f,g\)=ℒ\(f,g\)\+ρ\(‖Σg−Id‖F2\+‖offdiag\(Σf\)‖F2\)\.\\mathcal\{L\}\_\{\\rho\}\(f,g\)=\\mathcal\{L\}\(f,g\)\+\\rho\\Bigl\(\\bigl\\\|\\Sigma\_\{g\}\-I\_\{d\}\\bigr\\\|\_\{F\}^\{2\}\+\\bigl\\\|\\operatorname\{offdiag\}\(\\Sigma\_\{f\}\)\\bigr\\\|\_\{F\}^\{2\}\\Bigr\)\.\(48\)Asρ→∞\\rho\\to\\infty, any convergent sequence of minimizers ofℒρ\\mathcal\{L\}\_\{\\rho\}converges to a solution satisfying the constraints of Theorem[2](https://arxiv.org/html/2608.11661#Thmtheorem2)up to the ordering convention, and hence to the recovered modes \([10](https://arxiv.org/html/2608.11661#S3.E10)\)\. For finiteρ\\rho, the residual gauge directions that break the constraintΣg=I\\Sigma\_\{g\}=Icarry curvature at least2ρc2\\rho\\,cfor a constantc\>0c\>0that depends on the second moments, so conditioning improves monotonically withρ\\rho\.
#### Proof of Theorem[6](https://arxiv.org/html/2608.11661#Thmtheorem6)\(Soft whitening recovers the hard solution\)\.
\(i\) Convergence of minimizers\. The zero set of the penalty
ρ\(‖Σg−I‖F2\+‖offdiag\(Σf\)‖F2\)\\rho\\Bigl\(\\\|\\Sigma\_\{g\}\-I\\\|\_\{F\}^\{2\}\+\\\|\\operatorname\{offdiag\}\(\\Sigma\_\{f\}\)\\\|\_\{F\}^\{2\}\\Bigr\)\(49\)is exactly the whitening constraint \([9](https://arxiv.org/html/2608.11661#S3.E9)\) up to the ordering convention\. Sinceℒρ≥ℒ\\mathcal\{L\}\_\{\\rho\}\\geq\\mathcal\{L\}, any limit point of a convergent sequence of minimizers satisfies the constraint \(otherwise the penalty, and with itℒρ\\mathcal\{L\}\_\{\\rho\}, would diverge\) and is a global minimizer ofℒ\\mathcal\{L\}subject to the constraint, to which Theorem[2](https://arxiv.org/html/2608.11661#Thmtheorem2)applies\.
\(ii\) Curvature of the penalty\. Consider a gauge generatorX∈𝔤𝔩dX\\in\\mathfrak\{gl\}\_\{d\}acting on thegg\-side asg\(t\)=exp\(−tX⊤\)gg\(t\)=\\exp\(\-tX^\{\\top\}\)g\. The second moment evolves asΣg\(t\)=exp\(−tX⊤\)Σgexp\(−tX\)\\Sigma\_\{g\(t\)\}=\\exp\(\-tX^\{\\top\}\)\\Sigma\_\{g\}\\exp\(\-tX\)\. At a constraint\-satisfying pointΣg=I\\Sigma\_\{g\}=I, the second derivative of‖Σg\(t\)−I‖F2\\\|\\Sigma\_\{g\(t\)\}\-I\\\|\_\{F\}^\{2\}att=0t=0is2‖X\+X⊤‖F22\\\|X\+X^\{\\top\}\\\|\_\{F\}^\{2\}: expanding the exponential givesΣg\(t\)=I−t\(X⊤\+X\)\+t2X⊤X\+O\(t3\)\\Sigma\_\{g\(t\)\}=I\-t\(X^\{\\top\}\+X\)\+t^\{2\}X^\{\\top\}X\+O\(t^\{3\}\), so the squared deviation has leading termt2‖X\+X⊤‖F2t^\{2\}\\\|X\+X^\{\\top\}\\\|\_\{F\}^\{2\}multiplied by22\. Directions that breakΣg=I\\Sigma\_\{g\}=Ihave nonzero symmetric part, hence carry penalty curvature at least2ρc2\\rho\\,cwithc\>0c\>0determined by the second moments through the contribution ofoffdiag\(Σf\)\\operatorname\{offdiag\}\(\\Sigma\_\{f\}\); directions withXXantisymmetric are rotations, which preserveΣg=I\\Sigma\_\{g\}=Iand are exactly the directions whose residual freedom is fixed by the ordering convention\. Conditioning therefore improves monotonically withρ\\rho, as stated\.
## Appendix DProofs for Section 4
We prove the results of Section[4](https://arxiv.org/html/2608.11661#S4)\. Throughout,ℛ\(h\)=𝔼u,v\[\(F\(u,v\)−h\(u,v\)\)2\]\\mathcal\{R\}\(h\)=\\mathbb\{E\}\_\{u,v\}\[\(F\(u,v\)\-h\(u,v\)\)^\{2\}\]is the population risk andℛ^\\hat\{\\mathcal\{R\}\}its empirical version over thennsamples\. Corollary[6](https://arxiv.org/html/2608.11661#Thmcorollary6)and Proposition[6](https://arxiv.org/html/2608.11661#Thmproposition6)are proved in Appendix[E](https://arxiv.org/html/2608.11661#A5), as indicated in the main text\.
###### Lemma 3\(Rademacher collapse of the inner\-product class\)\.
For the class \([12](https://arxiv.org/html/2608.11661#S4.E12)\) with per\-coordinate classesℱ\(k\),𝒢\(k\)\\mathcal\{F\}^\{\(k\)\},\\mathcal\{G\}^\{\(k\)\},
ℜn\(ℋd\)≤∑k=1dℜn\(ℱ\(k\)⋅𝒢\(k\)\)≤2∑k=1d\(Bgℜn\(ℱ\(k\)\)\+Bfℜn\(𝒢\(k\)\)\),\\mathfrak\{R\}\_\{n\}\(\\mathcal\{H\}\_\{d\}\)\\;\\leq\\;\\sum\_\{k=1\}^\{d\}\\mathfrak\{R\}\_\{n\}\\bigl\(\\mathcal\{F\}^\{\(k\)\}\\cdot\\mathcal\{G\}^\{\(k\)\}\\bigr\)\\;\\leq\\;2\\sum\_\{k=1\}^\{d\}\\Bigl\(B\_\{g\}\\,\\mathfrak\{R\}\_\{n\}\(\\mathcal\{F\}^\{\(k\)\}\)\+B\_\{f\}\\,\\mathfrak\{R\}\_\{n\}\(\\mathcal\{G\}^\{\(k\)\}\)\\Bigr\),\(50\)so in particularℜn\(ℋd\)≤2d\(Bgℜn\(ℱ\)\+Bfℜn\(𝒢\)\)\\mathfrak\{R\}\_\{n\}\(\\mathcal\{H\}\_\{d\}\)\\leq 2d\\bigl\(B\_\{g\}\\,\\mathfrak\{R\}\_\{n\}\(\\mathcal\{F\}\)\+B\_\{f\}\\,\\mathfrak\{R\}\_\{n\}\(\\mathcal\{G\}\)\\bigr\)for shared per\-coordinate classes: the complexity of the head is a sum of marginal complexities, not a complexity of the joint space\.
#### Proof of Lemma[3](https://arxiv.org/html/2608.11661#Thmlemma3)\(Rademacher collapse of the inner\-product class\)\.
Decompose the score as⟨f\(u\),g\(v\)⟩=∑k=1dfk\(u\)gk\(v\)\\langle f\(u\),g\(v\)\\rangle=\\sum\_\{k=1\}^\{d\}f\_\{k\}\(u\)g\_\{k\}\(v\)\. Rademacher complexity is subadditive over sums of function classes, so
ℜn\(ℋd\)≤∑k=1dℜn\(ℱ\(k\)⋅𝒢\(k\)\),\\mathfrak\{R\}\_\{n\}\(\\mathcal\{H\}\_\{d\}\)\\leq\\sum\_\{k=1\}^\{d\}\\mathfrak\{R\}\_\{n\}\\bigl\(\\mathcal\{F\}^\{\(k\)\}\\cdot\\mathcal\{G\}^\{\(k\)\}\\bigr\),\(51\)whereℱ\(k\)⋅𝒢\(k\)=\{\(u,v\)↦fk\(u\)gk\(v\)\}\\mathcal\{F\}^\{\(k\)\}\\cdot\\mathcal\{G\}^\{\(k\)\}=\\\{\(u,v\)\\mapsto f\_\{k\}\(u\)g\_\{k\}\(v\)\\\}is the pointwise product class\. For each coordinate, the product map\(fk,gk\)↦fkgk\(f\_\{k\},g\_\{k\}\)\\mapsto f\_\{k\}g\_\{k\}is Lipschitz in its first argument with constantBgB\_\{g\}and in its second with constantBfB\_\{f\}on the bounded range; the vector contraction inequality of[15](https://arxiv.org/html/2608.11661#bib.bib22)then bounds the complexity of the product class by the sum of the two marginal complexities,
ℜn\(ℱ\(k\)⋅𝒢\(k\)\)≤2\(Bgℜn\(ℱ\(k\)\)\+Bfℜn\(𝒢\(k\)\)\),\\mathfrak\{R\}\_\{n\}\\bigl\(\\mathcal\{F\}^\{\(k\)\}\\cdot\\mathcal\{G\}^\{\(k\)\}\\bigr\)\\leq 2\\Bigl\(B\_\{g\}\\,\\mathfrak\{R\}\_\{n\}\(\\mathcal\{F\}^\{\(k\)\}\)\+B\_\{f\}\\,\\mathfrak\{R\}\_\{n\}\(\\mathcal\{G\}^\{\(k\)\}\)\\Bigr\),\(52\)the factor22absorbing the contraction constants\. Summing over theddcoordinates and usingℜn\(ℱ\)=maxkℜn\(ℱ\(k\)\)\\mathfrak\{R\}\_\{n\}\(\\mathcal\{F\}\)=\\max\_\{k\}\\mathfrak\{R\}\_\{n\}\(\\mathcal\{F\}^\{\(k\)\}\)gives the displayed bound\. The key structural point is that no complexity of the joint space𝒰×𝒱\\mathcal\{U\}\\times\\mathcal\{V\}appears: the head is paid for by its two marginal classes\.
###### Corollary 4\(Excess risk of empirical risk minimization\)\.
LetF^\\hat\{F\}be the empirical risk minimizer overℋd\\mathcal\{H\}\_\{d\}\. With probability at least1−δ1\-\\deltaover the sample, the excess risk decomposes into an approximation part, the truncation and realization terms of Theorem[1](https://arxiv.org/html/2608.11661#Thmtheorem1), and an estimation part through the class complexity:
‖F^−F‖L2\(μU⊗μV\)2≲∑k\>dσk2\+ℰreal\+Bd\(ℜn\(ℱ\)\+ℜn\(𝒢\)\)\+B2log\(1/δ\)n,\\\|\\hat\{F\}\-F\\\|\_\{L^\{2\}\(\\mu\_\{U\}\\otimes\\mu\_\{V\}\)\}^\{2\}\\;\\lesssim\\;\\sum\_\{k\>d\}\\sigma\_\{k\}^\{2\}\+\\mathcal\{E\}\_\{\\mathrm\{real\}\}\\;\+\\;B\\,d\\bigl\(\\mathfrak\{R\}\_\{n\}\(\\mathcal\{F\}\)\+\\mathfrak\{R\}\_\{n\}\(\\mathcal\{G\}\)\\bigr\)\\;\+\\;B^\{2\}\\sqrt\{\\frac\{\\log\(1/\\delta\)\}\{n\}\},\(53\)whereℰreal=C∑k≤d\(σk2ηg,k2\+Bg2ηf,k2\)\\mathcal\{E\}\_\{\\mathrm\{real\}\}=C\\sum\_\{k\\leq d\}\(\\sigma\_\{k\}^\{2\}\\eta\_\{g,k\}^\{2\}\+B\_\{g\}^\{2\}\\eta\_\{f,k\}^\{2\}\)is the realization term and≲\\lesssimhides absolute constants\.
#### Proof of Corollary[4](https://arxiv.org/html/2608.11661#Thmcorollary4)\(Excess risk of empirical risk minimization\)\.
The regression oracle decomposition: for anyh∈ℋdh\\in\\mathcal\{H\}\_\{d\}, writingF^\\hat\{F\}for the empirical risk minimizer,
‖F−F^‖L22≤‖F−h‖L22\+2suph′∈ℋd\|ℛ^\(h′\)−ℛ\(h′\)\|\.\\\|F\-\\hat\{F\}\\\|\_\{L^\{2\}\}^\{2\}\\leq\\\|F\-h\\\|\_\{L^\{2\}\}^\{2\}\+2\\sup\_\{h^\{\\prime\}\\in\\mathcal\{H\}\_\{d\}\}\\bigl\|\\hat\{\\mathcal\{R\}\}\(h^\{\\prime\}\)\-\\mathcal\{R\}\(h^\{\\prime\}\)\\bigr\|\.\(54\)The uniform deviation is controlled as follows\. The squared lossℓ\(y,h\)=\(y−h\)2\\ell\(y,h\)=\(y\-h\)^\{2\}is4B4B\-Lipschitz inhhon\|y\|,\|h\|≤B\|y\|,\|h\|\\leq B, so by the Talagrand contraction inequality the Rademacher complexity of the loss class is at most4Bℜn\(ℋd\)4B\\,\\mathfrak\{R\}\_\{n\}\(\\mathcal\{H\}\_\{d\}\)\([4](https://arxiv.org/html/2608.11661#bib.bib23)\)\. Symmetrization bounds the expected uniform deviation by twice the Rademacher complexity of the loss class, and McDiarmid’s inequality, with the bounded loss\|y−h\|2≤4B2\|y\-h\|^\{2\}\\leq 4B^\{2\}, gives the concentration term:
suph′∈ℋd\|ℛ^\(h′\)−ℛ\(h′\)\|≲Bℜn\(ℋd\)\+B2log\(1/δ\)n\\sup\_\{h^\{\\prime\}\\in\\mathcal\{H\}\_\{d\}\}\\bigl\|\\hat\{\\mathcal\{R\}\}\(h^\{\\prime\}\)\-\\mathcal\{R\}\(h^\{\\prime\}\)\\bigr\|\\lesssim B\\,\\mathfrak\{R\}\_\{n\}\(\\mathcal\{H\}\_\{d\}\)\+B^\{2\}\\sqrt\{\\frac\{\\log\(1/\\delta\)\}\{n\}\}\(55\)with probability at least1−δ1\-\\delta\. Choosingh=h⋆h=h^\{\\star\}, the population optimum, makes‖F−h⋆‖2\\\|F\-h^\{\\star\}\\\|^\{2\}exactly the approximation term of Theorem[1](https://arxiv.org/html/2608.11661#Thmtheorem1), and Lemma[3](https://arxiv.org/html/2608.11661#Thmlemma3)boundsℜn\(ℋd\)≤2d\(Bgℜn\(ℱ\)\+Bfℜn\(𝒢\)\)\\mathfrak\{R\}\_\{n\}\(\\mathcal\{H\}\_\{d\}\)\\leq 2d\(B\_\{g\}\\mathfrak\{R\}\_\{n\}\(\\mathcal\{F\}\)\+B\_\{f\}\\mathfrak\{R\}\_\{n\}\(\\mathcal\{G\}\)\); absorbing the Lipschitz and contraction constants into the absolute constant gives \([14](https://arxiv.org/html/2608.11661#S4.E14)\)\. For parameterized encoder classes withpfp\_\{f\}andpgp\_\{g\}parameters under norm control, the standard Rademacher bound for parameterized classes givesℜn\(ℱ\)=O\(pf/n\)\\mathfrak\{R\}\_\{n\}\(\\mathcal\{F\}\)=O\(\\sqrt\{p\_\{f\}/n\}\)andℜn\(𝒢\)=O\(pg/n\)\\mathfrak\{R\}\_\{n\}\(\\mathcal\{G\}\)=O\(\\sqrt\{p\_\{g\}/n\}\), and the estimation term becomesO\(Bd\(pf\+pg\)/n\)O\(Bd\\sqrt\{\(p\_\{f\}\+p\_\{g\}\)/n\}\)\.
###### Theorem 7\(Linear heads are low\-rank matrix sensing\)\.
LetΘ⋆∈ℝp×q\\Theta^\{\\star\}\\in\\mathbb\{R\}^\{p\\times q\}have rankddand‖Θ⋆‖F≤B\\\|\\Theta^\{\\star\}\\\|\_\{F\}\\leq B, and suppose the sensing matricesuivi⊤u\_\{i\}v\_\{i\}^\{\\top\}are drawn from an isotropic sub\-Gaussian family and the noise is sub\-Gaussian with varianceσε2\\sigma\_\{\\varepsilon\}^\{2\}\. The rank\-constrained estimator attains
𝔼‖Θ^−Θ⋆‖F2≍σε2d\(p\+q\)n,\\mathbb\{E\}\\bigl\\\|\\hat\{\\Theta\}\-\\Theta^\{\\star\}\\bigr\\\|\_\{F\}^\{2\}\\;\\asymp\\;\\sigma\_\{\\varepsilon\}^\{2\}\\,\\frac\{d\(p\+q\)\}\{n\},\(56\)and this rate is minimax optimal\. Equivalently,𝔼‖F^−F‖L22≍σε2d\(p\+q\)/n\\mathbb\{E\}\\\|\\hat\{F\}\-F\\\|\_\{L^\{2\}\}^\{2\}\\asymp\\sigma\_\{\\varepsilon\}^\{2\}\\,d\(p\+q\)/n\.
#### Proof of Theorem[7](https://arxiv.org/html/2608.11661#Thmtheorem7)\(Linear heads are low\-rank matrix sensing\)\.
The head computesu⊤Θvu^\{\\top\}\\Theta vwithΘ=UV⊤\\Theta=UV^\{\\top\}, and the observations areyi=⟨Θ⋆,uivi⊤⟩F\+εiy\_\{i\}=\\langle\\Theta^\{\\star\},\\,u\_\{i\}v\_\{i\}^\{\\top\}\\rangle\_\{F\}\+\\varepsilon\_\{i\}, the standard form of low\-rank matrix sensing with sensing matricesAi=uivi⊤A\_\{i\}=u\_\{i\}v\_\{i\}^\{\\top\}and rank constraintrank\(Θ\)≤d\\rank\(\\Theta\)\\leq d\. Upper bound: the rank\-ddmanifold hasd\(p\+q−d\)d\(p\+q\-d\)free parameters, and under restricted strong convexity of the sensing operator, which holds for isotropic sub\-Gaussian sensing matrices, the constrained estimator satisfies𝔼‖Θ^−Θ⋆‖F2≲σε2d\(p\+q\)/n\\mathbb\{E\}\\\|\\hat\{\\Theta\}\-\\Theta^\{\\star\}\\\|\_\{F\}^\{2\}\\lesssim\\sigma\_\{\\varepsilon\}^\{2\}\\,d\(p\+q\)/n; this is the robust error bound of[5](https://arxiv.org/html/2608.11661#bib.bib24)together with the high\-dimensional rate of[16](https://arxiv.org/html/2608.11661#bib.bib25)\. Lower bound: a Fano argument over a packing of the rank\-ddmanifold gives the matching minimax lower boundΩ\(σε2d\(p\+q\)/n\)\\Omega\(\\sigma\_\{\\varepsilon\}^\{2\}\\,d\(p\+q\)/n\)\([20](https://arxiv.org/html/2608.11661#bib.bib26)\)\. The equivalence with the function\-space rate follows from the isotropy ofuuandvv:𝔼u,v\[\(u⊤Δv\)2\]≍‖Δ‖F2\\mathbb\{E\}\_\{u,v\}\[\(u^\{\\top\}\\Delta v\)^\{2\}\]\\asymp\\\|\\Delta\\\|\_\{F\}^\{2\}for sub\-Gaussian isotropic arguments, so𝔼‖F^−F‖L22≍σε2d\(p\+q\)/n\\mathbb\{E\}\\\|\\hat\{F\}\-F\\\|\_\{L^\{2\}\}^\{2\}\\asymp\\sigma\_\{\\varepsilon\}^\{2\}\\,d\(p\+q\)/n\.
###### Corollary 5\(Optimal rank selection\)\.
If the interaction spectrum decays exponentially,σk2=O\(e−2ck\)\\sigma\_\{k\}^\{2\}=O\(e^\{\-2ck\}\), and the encoders are linear, the total error is minimized at
d⋆≍12clognp\+q,d^\{\\star\}\\;\\asymp\\;\\frac\{1\}\{2c\}\\log\\frac\{n\}\{p\+q\},\(57\)and the achieved error is‖F^−F‖L22≲\(p\+q\)log\(n/\(p\+q\)\)/n\\\|\\hat\{F\}\-F\\\|\_\{L^\{2\}\}^\{2\}\\lesssim\(p\+q\)\\log\(n/\(p\+q\)\)/n, a near\-parametric rate in which fast spectral decay eliminates the curse of dimensionality\.
#### Proof of Corollary[5](https://arxiv.org/html/2608.11661#Thmcorollary5)\(Optimal rank selection\)\.
With exponential decay and linear encoders the total error of Theorem[7](https://arxiv.org/html/2608.11661#Thmtheorem7)and the truncation term balance as
err\(d\)≲e−2cd\+d\(p\+q\)n\.\\mathrm\{err\}\(d\)\\;\\lesssim\\;e^\{\-2cd\}\+\\frac\{d\(p\+q\)\}\{n\}\.\(58\)Setting the two terms equal,e−2cd⋆≍d⋆\(p\+q\)/ne^\{\-2cd^\{\\star\}\}\\asymp d^\{\\star\}\(p\+q\)/n, and solving givesd⋆≍\(1/2c\)log\(n/\(p\+q\)\)d^\{\\star\}\\asymp\(1/2c\)\\log\(n/\(p\+q\)\), where the logarithmic factor inddis absorbed by the constant\. Substituting back,e−2cd⋆≍\(p\+q\)/ne^\{\-2cd^\{\\star\}\}\\asymp\(p\+q\)/n, so the achieved error is
‖F^−F‖L22≲\(p\+q\)n\(1\+lognp\+q\)≍\(p\+q\)log\(n/\(p\+q\)\)n,\\\|\\hat\{F\}\-F\\\|\_\{L^\{2\}\}^\{2\}\\lesssim\\frac\{\(p\+q\)\}\{n\}\\Bigl\(1\+\\log\\frac\{n\}\{p\+q\}\\Bigr\)\\asymp\\frac\{\(p\+q\)\\log\(n/\(p\+q\)\)\}\{n\},\(59\)the near\-parametric rate stated in the corollary\.
###### Theorem 8\(Flat spectrum forces an exponential embedding\)\.
Let𝒰=𝒱=\{±1\}m\\mathcal\{U\}=\\mathcal\{V\}=\\\{\\pm 1\\\}^\{m\}with uniform measure andN=2mN=2^\{m\}, and consider the equality functionFeq\(u,v\)=𝟏\[u=v\]F\_\{\\mathrm\{eq\}\}\(u,v\)=\\mathbf\{1\}\[u=v\]\. Its interaction operator isTFeq=1NIdT\_\{F\_\{\\mathrm\{eq\}\}\}=\\tfrac\{1\}\{N\}\\mathrm\{Id\}, so allNNsingular values equal1/N1/Nand the interaction rank isNN\. Every rank\-ddhead therefore satisfies
inff,g‖Feq−⟨f,g⟩‖L22‖Feq‖L22=1−dN,\\frac\{\\inf\_\{f,g\}\\bigl\\\|F\_\{\\mathrm\{eq\}\}\-\\langle f,g\\rangle\\bigr\\\|\_\{L^\{2\}\}^\{2\}\}\{\\bigl\\\|F\_\{\\mathrm\{eq\}\}\\bigr\\\|\_\{L^\{2\}\}^\{2\}\}\\;=\\;1\-\\frac\{d\}\{N\},\(60\)so achieving relative errorε\\varepsilonrequiresd≥\(1−ε\)2md\\geq\(1\-\\varepsilon\)2^\{m\}, an embedding dimension exponential in the input size\.
#### Proof of Theorem[8](https://arxiv.org/html/2608.11661#Thmtheorem8)\(Flat spectrum forces an exponential embedding\)\.
Compute the interaction operator ofFeqF\_\{\\mathrm\{eq\}\}on the uniform measure over\{±1\}m\\\{\\pm 1\\\}^\{m\}withN=2mN=2^\{m\}:
\(TFeqh\)\(u\)=∑v∈\{±1\}m𝟏\[u=v\]h\(v\)1N=h\(u\)N,\(T\_\{F\_\{\\mathrm\{eq\}\}\}h\)\(u\)=\\sum\_\{v\\in\\\{\\pm 1\\\}^\{m\}\}\\mathbf\{1\}\[u=v\]\\,h\(v\)\\,\\frac\{1\}\{N\}=\\frac\{h\(u\)\}\{N\},\(61\)soTFeq=1NIdT\_\{F\_\{\\mathrm\{eq\}\}\}=\\tfrac\{1\}\{N\}\\mathrm\{Id\}and allNNsingular values equal1/N1/N\. Hence‖Feq‖L22=∑k≥1σk2=N/N2=1/N\\\|F\_\{\\mathrm\{eq\}\}\\\|\_\{L^\{2\}\}^\{2\}=\\sum\_\{k\\geq 1\}\\sigma\_\{k\}^\{2\}=N/N^\{2\}=1/Nand∑k\>dσk2=\(N−d\)/N2\\sum\_\{k\>d\}\\sigma\_\{k\}^\{2\}=\(N\-d\)/N^\{2\}ford<Nd<N\. By Corollary[1](https://arxiv.org/html/2608.11661#Thmcorollary1), the relative squared error of the best rank\-ddhead is
inff,g‖Feq−⟨f,g⟩‖L22‖Feq‖L22=\(N−d\)/N21/N=1−dN,\\frac\{\\inf\_\{f,g\}\\\|F\_\{\\mathrm\{eq\}\}\-\\langle f,g\\rangle\\\|\_\{L^\{2\}\}^\{2\}\}\{\\\|F\_\{\\mathrm\{eq\}\}\\\|\_\{L^\{2\}\}^\{2\}\}=\\frac\{\(N\-d\)/N^\{2\}\}\{1/N\}=1\-\\frac\{d\}\{N\},\(62\)which is at mostε\\varepsilononly ifd≥\(1−ε\)N=\(1−ε\)2md\\geq\(1\-\\varepsilon\)N=\(1\-\\varepsilon\)2^\{m\}\.
###### Theorem 9\(Early interaction escapes\)\.
For the equality functionFeqF\_\{\\mathrm\{eq\}\}on\{±1\}m×\{±1\}m\\\{\\pm 1\\\}^\{m\}\\times\\\{\\pm 1\\\}^\{m\}, an early\-interaction predictor withO\(m\)O\(m\)parameters representsFeqF\_\{\\mathrm\{eq\}\}exactly, with zero error\. Together with Theorem[8](https://arxiv.org/html/2608.11661#Thmtheorem8)this is an exponential separation:O\(m\)O\(m\)parameters suffice under early interaction, while any late\-interaction head needsΩ\(2m\)\\Omega\(2^\{m\}\)embedding dimensions to reach constant relative error\.
#### Proof of Theorem[9](https://arxiv.org/html/2608.11661#Thmtheorem9)\(Early interaction escapes\)\.
The construction uses the identityuivi=1u\_\{i\}v\_\{i\}=1if and only ifui=viu\_\{i\}=v\_\{i\}, so that
⟨u,v⟩=∑i=1muivi=m−2Ham\(u,v\),\\langle u,v\\rangle=\\sum\_\{i=1\}^\{m\}u\_\{i\}v\_\{i\}=m\-2\\,\\mathrm\{Ham\}\(u,v\),\(63\)and⟨u,v⟩=m\\langle u,v\\rangle=mif and only ifu=vu=v, with the next possible valuem−2m\-2\. HenceFeq\(u,v\)=𝟏\[⟨u,v⟩≥m−1\]F\_\{\\mathrm\{eq\}\}\(u,v\)=\\mathbf\{1\}\[\\langle u,v\\rangle\\geq m\-1\]\. A cross\-attention layer with identity projections computes the query\-key inner product⟨u,v⟩\\langle u,v\\ranglewithO\(m\)O\(m\)parameters, and a threshold unitϕ\(s\)=ReLU\(s−\(m−1\)\)−ReLU\(s−\(m\+1\)\)\\phi\(s\)=\\mathrm\{ReLU\}\(s\-\(m\-1\)\)\-\\mathrm\{ReLU\}\(s\-\(m\+1\)\)evaluates to11exactly ons=ms=mand to00elsewhere on the discrete range, usingO\(1\)O\(1\)parameters\. A joint MLP computes the same quantity bit by bit: each productuivi=\(\(ui\+vi\)2−2\)/2u\_\{i\}v\_\{i\}=\(\(u\_\{i\}\+v\_\{i\}\)^\{2\}\-2\)/2costsO\(1\)O\(1\)ReLU units, the sum costsO\(m\)O\(m\)additions, and the same threshold completes the network, againO\(m\)O\(m\)parameters in total\. A cross\-encoder, being a strict superset of cross\-attention, inherits the construction\. Together with Theorem[8](https://arxiv.org/html/2608.11661#Thmtheorem8), this is the exponential separationO\(m\)O\(m\)versusΩ\(2m\)\\Omega\(2^\{m\}\)\.
###### Theorem 10\(Usability criterion\)\.
Given a problem\(F,μU,μV\)\(F,\\mu\_\{U\},\\mu\_\{V\}\), a target relative errorε\\varepsilon, and a sample budgetnn: if the effective interaction rankdε\(F\)d\_\{\\varepsilon\}\(F\)is small, a late\-interaction head is appropriate, since its approximation and estimation errors are then small and it enjoys the compute and statistical separations of Proposition[1](https://arxiv.org/html/2608.11661#Thmproposition1)and Lemma[3](https://arxiv.org/html/2608.11661#Thmlemma3)\. If the spectrum is flat,dε\(F\)=Ω\(N\)d\_\{\\varepsilon\}\(F\)=\\Omega\(N\), the truncation floor of Corollary[1](https://arxiv.org/html/2608.11661#Thmcorollary1)locks every rank\-ddhead to relative error at leastε\\varepsilonregardless of capacity, samples, or training, while early\-interaction models reachε\\varepsilonpolynomially when the target admits a low\-dimensional joint structure\. The deciding quantity, the spectral decay rate, is estimable from data by a weighted singular value decomposition of the empirical kernel matrix, so the criterion is actionable before training either model\.
#### Proof of Theorem[10](https://arxiv.org/html/2608.11661#Thmtheorem10)\(Usability criterion\)\.
First case: ifdε\(F\)d\_\{\\varepsilon\}\(F\)is small, taked≳dε\(F\)d\\gtrsim d\_\{\\varepsilon\}\(F\)\. The approximation error of Theorem[1](https://arxiv.org/html/2608.11661#Thmtheorem1)is at mostε‖F‖2\\varepsilon\\\|F\\\|^\{2\}plus the realization term, and the estimation error of Corollary[4](https://arxiv.org/html/2608.11661#Thmcorollary4)scales asd\(ℜn\(ℱ\)\+ℜn\(𝒢\)\)d\(\\mathfrak\{R\}\_\{n\}\(\\mathcal\{F\}\)\+\\mathfrak\{R\}\_\{n\}\(\\mathcal\{G\}\)\), small indd; the compute separation of Proposition[1](https://arxiv.org/html/2608.11661#Thmproposition1)and the marginal\-complexity collapse of Lemma[3](https://arxiv.org/html/2608.11661#Thmlemma3)apply without further conditions\. Second case: if the spectrum is flat withdε\(F\)=Ω\(N\)d\_\{\\varepsilon\}\(F\)=\\Omega\(N\), then for everyd<dε\(F\)d<d\_\{\\varepsilon\}\(F\)the truncation floor of Corollary[1](https://arxiv.org/html/2608.11661#Thmcorollary1), in the explicit form of Theorem[8](https://arxiv.org/html/2608.11661#Thmtheorem8), locks the relative error of every rank\-ddhead aboveε\\varepsilonregardless of capacity, samples, or training, while the constructive bounds of Theorem[9](https://arxiv.org/html/2608.11661#Thmtheorem9)and Remark[7](https://arxiv.org/html/2608.11661#Thmremark7)show that early\-interaction models reachε\\varepsilonwith a polynomial parameter budget when the target depends on its arguments through a low\-dimensional joint statistic\. Actionability: the empirical kernel matrixM^ij=y\(ui,vj\)\\hat\{M\}\_\{ij\}=y\(u\_\{i\},v\_\{j\}\)on a grid of input pairs, weighted by the reference measures, has a weighted singular value decomposition whose values estimate\{σk\}\\\{\\sigma\_\{k\}\\\}; a log\-linear fit of the decay ofσ^k\\hat\{\\sigma\}\_\{k\}againstkk\(exponential\) orlogk\\log k\(polynomial\) decides the decay type and reads offdεd\_\{\\varepsilon\}before either model is trained\.
## Appendix EFinite\-Sample Technical Lemmas
This appendix collects the finite\-sample lemmas used in the main text\. Section A makes the whitening identification of Theorem[2](https://arxiv.org/html/2608.11661#Thmtheorem2)quantitative: concentration of empirical second moments \(Lemma[4](https://arxiv.org/html/2608.11661#Thmlemma4)\), stability of the whitening map \(Lemma[5](https://arxiv.org/html/2608.11661#Thmlemma5)\), and a Davis\-Kahan propagation step \(Theorem[11](https://arxiv.org/html/2608.11661#Thmtheorem11)\) deliver the sample complexity of Corollary[6](https://arxiv.org/html/2608.11661#Thmcorollary6)\. Section B quantifies the degradation of the NMF route when separability holds only approximately \(Theorem[12](https://arxiv.org/html/2608.11661#Thmtheorem12)\)\. Section C couples identification with estimation in one ledger \(Proposition[5](https://arxiv.org/html/2608.11661#Thmproposition5)\), which is the proof of Proposition[6](https://arxiv.org/html/2608.11661#Thmproposition6)\. Throughout,∥⋅∥\\\|\\cdot\\\|on matrices is the spectral norm and∥⋅∥F\\\|\\cdot\\\|\_\{F\}the Frobenius norm\.
### E\.1Empirical whitening and mode identification
###### Lemma 4\(Matrix Bernstein concentration of empirical moments\)\.
Let\{ui\}i=1n\\\{u\_\{i\}\\\}\_\{i=1\}^\{n\}be drawn i\.i\.d\. fromμU\\mu\_\{U\}and letffhave bounded outputs,‖f\(u\)‖≤Bf\\\|f\(u\)\\\|\\leq B\_\{f\}\. The empirical second momentΣ^f=1n∑if\(ui\)f\(ui\)⊤\\hat\{\\Sigma\}\_\{f\}=\\tfrac\{1\}\{n\}\\sum\_\{i\}f\(u\_\{i\}\)f\(u\_\{i\}\)^\{\\top\}satisfies, for everyt\>0t\>0,
Pr\[∥Σ^f−Σf∥≥t\]≤2dexp\(−nt2/2νf\+Bf2t/3\),νf:=Bf2∥Σf∥,\\Pr\\bigl\[\\\|\\hat\{\\Sigma\}\_\{f\}\-\\Sigma\_\{f\}\\\|\\geq t\\bigr\]\\leq 2d\\,\\exp\\\!\\Bigl\(\\frac\{\-nt^\{2\}/2\}\{\\nu\_\{f\}\+B\_\{f\}^\{2\}t/3\}\\Bigr\),\\qquad\\nu\_\{f\}:=B\_\{f\}^\{2\}\\\|\\Sigma\_\{f\}\\\|,\(64\)and hence, with probability at least1−δ1\-\\delta,
‖Σ^f−Σf‖≤2νflog\(2d/δ\)n\+2Bf2log\(2d/δ\)3n=:rf\(n,δ\),\\\|\\hat\{\\Sigma\}\_\{f\}\-\\Sigma\_\{f\}\\\|\\leq\\sqrt\{\\frac\{2\\nu\_\{f\}\\log\(2d/\\delta\)\}\{n\}\}\+\\frac\{2B\_\{f\}^\{2\}\\log\(2d/\\delta\)\}\{3n\}=:r\_\{f\}\(n,\\delta\),\(65\)which isO\(Bf2log\(d/δ\)/n\)O\\bigl\(B\_\{f\}^\{2\}\\sqrt\{\\log\(d/\\delta\)/n\}\\bigr\)for largenn\. The symmetric statement holds forggwith radiusrg\(n,δ\)r\_\{g\}\(n,\\delta\)\.
###### Proof\.
The matricesXi=f\(ui\)f\(ui\)⊤−ΣfX\_\{i\}=f\(u\_\{i\}\)f\(u\_\{i\}\)^\{\\top\}\-\\Sigma\_\{f\}are independent, zero mean, and symmetric, with‖Xi‖≤‖f\(ui\)f\(ui\)⊤‖\+‖Σf‖≤2Bf2\\\|X\_\{i\}\\\|\\leq\\\|f\(u\_\{i\}\)f\(u\_\{i\}\)^\{\\top\}\\\|\+\\\|\\Sigma\_\{f\}\\\|\\leq 2B\_\{f\}^\{2\}\. For the variance parameter, since𝔼\[Xi2\]⪯𝔼‖f\(ui\)‖2f\(ui\)f\(ui\)⊤⪯Bf2Σf\\mathbb\{E\}\[X\_\{i\}^\{2\}\]\\preceq\\mathbb\{E\}\\\|f\(u\_\{i\}\)\\\|^\{2\}f\(u\_\{i\}\)f\(u\_\{i\}\)^\{\\top\}\\preceq B\_\{f\}^\{2\}\\Sigma\_\{f\}, we have‖∑i𝔼\[Xi2\]‖≤nBf2‖Σf‖=nνf\\bigl\\\|\\sum\_\{i\}\\mathbb\{E\}\[X\_\{i\}^\{2\}\]\\bigr\\\|\\leq nB\_\{f\}^\{2\}\\\|\\Sigma\_\{f\}\\\|=n\\nu\_\{f\}\. The matrix Bernstein inequality \(Tropp, An Introduction to Matrix Concentration Inequalities\) applied to∑iXi\\sum\_\{i\}X\_\{i\}gives the tail bound, and solving the tail forttat the levelδ\\deltagives the displayed high\-probability bound\. ∎
###### Lemma 5\(Stability of the whitening map\)\.
LetΣg⪰γI\\Sigma\_\{g\}\\succeq\\gamma Iforγ\>0\\gamma\>0, and suppose‖Σ^g−Σg‖≤rg≤γ/2\\\|\\hat\{\\Sigma\}\_\{g\}\-\\Sigma\_\{g\}\\\|\\leq r\_\{g\}\\leq\\gamma/2\. Then
∥Σ^g−1/2−Σg−1/2∥≤C0rgγ3/2,\\bigl\\\|\\hat\{\\Sigma\}\_\{g\}^\{\-1/2\}\-\\Sigma\_\{g\}^\{\-1/2\}\\bigr\\\|\\leq\\frac\{C\_\{0\}\\,r\_\{g\}\}\{\\gamma^\{3/2\}\},\(66\)for an absolute constantC0C\_\{0\}\.
###### Proof\.
The mapA↦A−1/2A\\mapsto A^\{\-1/2\}is Fréchet differentiable on the coneA⪰γIA\\succeq\\gamma Iwith derivative bounded by12γ−3/2\\tfrac\{1\}\{2\}\\gamma^\{\-3/2\}: using the integral representationA−1/2=1π∫0∞\(A\+sI\)−1s−1/2dsA^\{\-1/2\}=\\tfrac\{1\}\{\\pi\}\\int\_\{0\}^\{\\infty\}\(A\+sI\)^\{\-1\}s^\{\-1/2\}\\,ds, the derivative atAAapplied toHHis−1π∫0∞\(A\+sI\)−1H\(A\+sI\)−1s−1/2ds\-\\tfrac\{1\}\{\\pi\}\\int\_\{0\}^\{\\infty\}\(A\+sI\)^\{\-1\}H\(A\+sI\)^\{\-1\}s^\{\-1/2\}\\,ds, whose norm is at most∥H∥⋅1π∫0∞\(γ\+s\)−2s−1/2ds=∥H∥12γ−3/2\\\|H\\\|\\cdot\\tfrac\{1\}\{\\pi\}\\int\_\{0\}^\{\\infty\}\(\\gamma\+s\)^\{\-2\}s^\{\-1/2\}\\,ds=\\\|H\\\|\\tfrac\{1\}\{2\}\\gamma^\{\-3/2\}\. The conditionrg≤γ/2r\_\{g\}\\leq\\gamma/2keepsΣ^g⪰γ2I\\hat\{\\Sigma\}\_\{g\}\\succeq\\tfrac\{\\gamma\}\{2\}Iby Weyl’s inequality, so the mean value theorem applies along the segment and gives the bound withC0C\_\{0\}absorbing the integration constant\. ∎
###### Theorem 11\(Finite\-sample mode identification\)\.
Under the assumptions of Theorem[2](https://arxiv.org/html/2608.11661#Thmtheorem2), suppose the encoders represent the population optimum exactly,⟨f,g⟩=Fd\\langle f,g\\rangle=F\_\{d\}, and whitening is performed on the empirical moments of a sample of sizenn, with encoder outputs bounded byBB\. Letr=max\(rf,rg\)r=\\max\(r\_\{f\},r\_\{g\}\)be the moment radii of Lemma[4](https://arxiv.org/html/2608.11661#Thmlemma4), letΔ=min1≤k<d\(σk2−σk\+12\)\\Delta=\\min\_\{1\\leq k<d\}\(\\sigma\_\{k\}^\{2\}\-\\sigma\_\{k\+1\}^\{2\}\)be the spectral gap, and assumer≤c1min\(γ,Δ\)r\\leq c\_\{1\}\\min\(\\gamma,\\Delta\)for a small absolute constantc1c\_\{1\}\. Then with probability at least1−δ1\-\\delta, the empirically whitened encoders\(f^,g^\)\(\\hat\{f\},\\hat\{g\}\)recover the modes up to
\|sin∠\(g^k,bk\)\|≤CrΔ,‖g^k−\(±bk\)‖≤C′rΔ,\|Λ^kk−σk2\|≤C′′r,\\bigl\|\\sin\\angle\(\\hat\{g\}\_\{k\},b\_\{k\}\)\\bigr\|\\leq\\frac\{C\\,r\}\{\\Delta\},\\qquad\\\|\\hat\{g\}\_\{k\}\-\(\\pm b\_\{k\}\)\\\|\\leq\\frac\{C^\{\\prime\}\\,r\}\{\\Delta\},\\qquad\|\\hat\{\\Lambda\}\_\{kk\}\-\\sigma\_\{k\}^\{2\}\|\\leq C^\{\\prime\\prime\}\\,r,\(67\)whereΛ^kk\\hat\{\\Lambda\}\_\{kk\}is the estimatedkk\-th interaction strength\. In particular the identification error isO\(B2log\(d/δ\)/n/Δ\)O\\bigl\(B^\{2\}\\sqrt\{\\log\(d/\\delta\)/n\}/\\Delta\\bigr\), vanishing asn−1/2n^\{\-1/2\}and amplified by the inverse spectral gap\.
###### Proof\.
\(1\) Empirical whitening followed by an eigendecomposition of the whitened second moment is equivalent to a singular value decomposition of the empirical interaction operator\. Since the encoders representFd=∑k≤dσkak⊗bkF\_\{d\}=\\sum\_\{k\\leq d\}\\sigma\_\{k\}a\_\{k\}\\otimes b\_\{k\}exactly, the population singular subspaces are spanned by\{bk\}\\\{b\_\{k\}\\\}on thegg\-side and\{ak\}\\\{a\_\{k\}\\\}on theff\-side\. \(2\) The perturbation of the empirical operator relative to its population version is bounded byC3r/γ3/2≤C4rC\_\{3\}\\,r/\\gamma^\{3/2\}\\leq C\_\{4\}\\,r: the moment error‖Σ^f−Σf‖≤rf\\\|\\hat\{\\Sigma\}\_\{f\}\-\\Sigma\_\{f\}\\\|\\leq r\_\{f\}from Lemma[4](https://arxiv.org/html/2608.11661#Thmlemma4), the whitening error from Lemma[5](https://arxiv.org/html/2608.11661#Thmlemma5), and the remaining factors are absorbed into the constant, which for brevity absorbs theγ\\gammadependence as well\. \(3\) Thekk\-th empirical mode is thekk\-th singular vector of the perturbed operator\. By the Davis\-KahansinΘ\\sin\\Thetatheorem, the perturbation of a singular vector is bounded by the operator perturbation divided by the gap between the corresponding singular value squared and its neighbors, which is at leastΔ\\Delta:
sin∠\(g^k,bk\)≤‖E‖min\(σk−12−σk2,σk2−σk\+12\)≤CrΔ\.\\sin\\angle\(\\hat\{g\}\_\{k\},b\_\{k\}\)\\leq\\frac\{\\\|E\\\|\}\{\\min\(\\sigma\_\{k\-1\}^\{2\}\-\\sigma\_\{k\}^\{2\},\\ \\sigma\_\{k\}^\{2\}\-\\sigma\_\{k\+1\}^\{2\}\)\}\\leq\\frac\{C\\,r\}\{\\Delta\}\.\(68\)Choosing the sign so that the inner product is nonnegative,‖g^k−\(±bk\)‖≤2sin∠≤C′r/Δ\\\|\\hat\{g\}\_\{k\}\-\(\\pm b\_\{k\}\)\\\|\\leq\\sqrt\{2\}\\,\\sin\\angle\\leq C^\{\\prime\}r/\\Delta\. \(4\) For the strengths, Weyl’s inequality gives\|Λ^kk−σk2\|≤‖E‖≤C′′r\|\\hat\{\\Lambda\}\_\{kk\}\-\\sigma\_\{k\}^\{2\}\|\\leq\\\|E\\\|\\leq C^\{\\prime\\prime\}r\. ∎
###### Corollary 6\(Sample complexity of whitening\)\.
Under the assumptions of Theorem[11](https://arxiv.org/html/2608.11661#Thmtheorem11), whitening the empirical second moments of a sample of sizennidentifies the interaction modes to accuracyε\\varepsilonwith probability at least1−δ1\-\\deltaprovided
n≳B4log\(d/δ\)Δ2ε2,n\\;\\gtrsim\\;\\frac\{B^\{4\}\\,\\log\(d/\\delta\)\}\{\\Delta^\{2\}\\,\\varepsilon^\{2\}\},\(69\)whereΔ=min1≤k<d\(σk2−σk\+12\)\\Delta=\\min\_\{1\\leq k<d\}\(\\sigma\_\{k\}^\{2\}\-\\sigma\_\{k\+1\}^\{2\}\)is the spectral gap\.
#### Proof of Corollary[6](https://arxiv.org/html/2608.11661#Thmcorollary6)\(Sample complexity of whitening\)\.
SetCr/Δ≤εCr/\\Delta\\leq\\varepsilonwithr=O\(B2log\(d/δ\)/n\)r=O\\bigl\(B^\{2\}\\sqrt\{\\log\(d/\\delta\)/n\}\\bigr\)from Lemma[4](https://arxiv.org/html/2608.11661#Thmlemma4)and solve fornn:
n≳B4log\(d/δ\)Δ2ε2\.n\\;\\gtrsim\\;\\frac\{B^\{4\}\\,\\log\(d/\\delta\)\}\{\\Delta^\{2\}\\,\\varepsilon^\{2\}\}\.\(70\)The sample budget scales as1/Δ21/\\Delta^\{2\}in the gap and1/ε21/\\varepsilon^\{2\}in the target accuracy, matching \([69](https://arxiv.org/html/2608.11661#A5.E69)\)\.
###### Corollary 7\(Block stability under small gaps\)\.
When a pair of singular values nearly coincides,Δk=σk2−σk\+12→0\\Delta\_\{k\}=\\sigma\_\{k\}^\{2\}\-\\sigma\_\{k\+1\}^\{2\}\\to 0, the single\-mode bound of Theorem[11](https://arxiv.org/html/2608.11661#Thmtheorem11)diverges: individual coordinates become unidentifiable\. The corresponding eigenspace remains stable: treating\{k,k\+1\}\\\{k,k\+1\\\}as a block, the projection error of the block is at mostO\(r/Δ~\)O\(r/\\tilde\{\\Delta\}\), whereΔ~\\tilde\{\\Delta\}is the gap between the block and the rest of the spectrum\. Under repeated singular values the identification therefore degrades to uniqueness up to rotation within the block, the quantitative version of the degenerate\-case remark of Section[3](https://arxiv.org/html/2608.11661#S3)\.
### E\.2Approximate separability in the NMF route
###### Definition 3\(α\\alpha\-approximate separability\)\.
A nonnegative factor matrixF∈ℝ≥0n×dF\\in\\mathbb\{R\}\_\{\\geq 0\}^\{n\\times d\}with rowsf\(ui\)f\(u\_\{i\}\)isα\\alpha\-separable if for every latent factorkkthere is a rowi\(k\)i\(k\)with
Fi\(k\),k≥1,∑ℓ≠kFi\(k\),ℓ≤α,F\_\{i\(k\),k\}\\geq 1,\\qquad\\sum\_\{\\ell\\neq k\}F\_\{i\(k\),\\ell\}\\leq\\alpha,\(71\)i\.e\., a near\-anchor row dominated by coordinatekkwith residual mass at mostα\\alpha\. The caseα=0\\alpha=0is exact separability\([6](https://arxiv.org/html/2608.11661#bib.bib11);[1](https://arxiv.org/html/2608.11661#bib.bib12)\)\.
###### Theorem 12\(Degradation of approximately separable NMF\)\.
LetM=FG⊤M=FG^\{\\top\}be a population nonnegative rank\-ddfactorization in whichFFisα\\alpha\-separable and the true factor cone is well conditioned, with the smallest angle between its extreme rays at leastθ0\>0\\theta\_\{0\}\>0\. Then every nonnegative optimal factorization\(F^,G^\)\(\\hat\{F\},\\hat\{G\}\)satisfies, in a suitable permutation and scaling gauge,
minP∈Perm,D∈D\+‖F^PD−F‖≤Cαsinθ0‖F‖,\\min\_\{P\\in\\mathrm\{Perm\},\\ D\\in\\mathrm\{D\}\_\{\+\}\}\\\|\\hat\{F\}\\,P\\,D\-F\\\|\\leq\\frac\{C\\,\\alpha\}\{\\sin\\theta\_\{0\}\}\\,\\\|F\\\|,\(72\)so the identification error is linear in the separability deficitα\\alphaand amplified by the cone condition number1/sinθ01/\\sin\\theta\_\{0\}\. Asα→0\\alpha\\to 0the bound recovers the exact uniqueness of Theorem[5](https://arxiv.org/html/2608.11661#Thmtheorem5)\.
###### Proof sketch\.
Identification in separable NMF is recovery of the extreme rays of the latent cone from the data points: under exact separability the anchor rows are exactly the extreme rays, and the constructive algorithm of[1](https://arxiv.org/html/2608.11661#bib.bib12)identifies them uniquely up to permutation and scaling\. Underα\\alpha\-approximate separability the near\-anchor rows deviate from the true extreme rays by at mostα\\alpha, and by the stability of vertex recovery in convex cones, a perturbation of the vertex set byα\\alphamoves the recovered rays byO\(α/sinθ0\)O\(\\alpha/\\sin\\theta\_\{0\}\)when the adjacent rays are separated by angle at leastθ0\\theta\_\{0\}; this is the robust version of separable NMF\([1](https://arxiv.org/html/2608.11661#bib.bib12)\)following the noise\-robustness analysis of Gillis and Vavasis\. The two measurable quantitiesα\\alphaandθ0\\theta\_\{0\}therefore provide an identifiability budget for the NMF route, which is what Section[5](https://arxiv.org/html/2608.11661#S5)reports for its nonnegativity condition\. ∎
### E\.3Coupling identification with estimation
###### Proposition 5\(Identification and estimation in one ledger\)\.
LetF^\\hat\{F\}be the empirical risk minimizer of Corollary[4](https://arxiv.org/html/2608.11661#Thmcorollary4)with encoder capacity fixed and whitening applied on the empirical moments\. Then the total identification error of thekk\-th mode decomposes as
‖g^k−\(±bk\)‖≤O\(‖F^−F‖L22Δ\)\+O\(rΔ\),\\\|\\hat\{g\}\_\{k\}\-\(\\pm b\_\{k\}\)\\\|\\leq O\\\!\\Bigl\(\\frac\{\\sqrt\{\\\|\\hat\{F\}\-F\\\|\_\{L^\{2\}\}^\{2\}\}\}\{\\Delta\}\\Bigr\)\+O\\\!\\Bigl\(\\frac\{r\}\{\\Delta\}\\Bigr\),\(73\)where the first term converts the excess risk of the estimate into subspace error via Davis\-Kahan and the second is the empirical whitening error of Theorem[11](https://arxiv.org/html/2608.11661#Thmtheorem11)\. Both terms scale as1/Δ1/\\Deltaand decay asn−1/2n^\{\-1/2\}\.
###### Proof\.
Split the error around the population whitened solution by the triangle inequality\. The second piece vanishes: at the population optimum the whitening theorem \(Theorem[2](https://arxiv.org/html/2608.11661#Thmtheorem2)\) identifies the modes exactly\. The first piece is the deviation of the estimate:‖F^−Fd‖L22≤‖F^−F‖L22\\\|\\hat\{F\}\-F\_\{d\}\\\|\_\{L^\{2\}\}^\{2\}\\leq\\\|\\hat\{F\}\-F\\\|\_\{L^\{2\}\}^\{2\}by optimality, and the Davis\-Kahan step of Theorem[11](https://arxiv.org/html/2608.11661#Thmtheorem11)converts anL2L^\{2\}deviation of the represented function into a singular\-vector deviation divided by the gap, giving the first term\. ∎
###### Proposition 6\(Mode error ledger\)\.
LetF^\\hat\{F\}be the empirical risk minimizer of Corollary[4](https://arxiv.org/html/2608.11661#Thmcorollary4)and letf^,g^\\hat\{f\},\\hat\{g\}be its whitened encoders\. The error in recovering thekk\-th interaction mode satisfies
max\{‖f^k−fk‖,‖g^k−gk‖\}=O\(‖F^−F‖L22\+rΔ\),\\max\\bigl\\\{\\\|\\hat\{f\}\_\{k\}\-f\_\{k\}\\\|,\\ \\\|\\hat\{g\}\_\{k\}\-g\_\{k\}\\\|\\bigr\\\}\\;=\\;O\\\!\\left\(\\frac\{\\sqrt\{\\\|\\hat\{F\}\-F\\\|\_\{L^\{2\}\}^\{2\}\}\+r\}\{\\Delta\}\\right\),\(74\)withr=O\(B2log\(d/δ\)/n\)r=O\\bigl\(B^\{2\}\\sqrt\{\\log\(d/\\delta\)/n\}\\bigr\); both terms are inversely proportional to the gap and decay asn−1/2n^\{\-1/2\}\.
#### Proof of Proposition[6](https://arxiv.org/html/2608.11661#Thmproposition6)\(Mode error ledger\)\.
Substitute the excess\-risk bound of Corollary[4](https://arxiv.org/html/2608.11661#Thmcorollary4)for‖F^−F‖L22\\\|\\hat\{F\}\-F\\\|\_\{L^\{2\}\}^\{2\}into the first term of Proposition[5](https://arxiv.org/html/2608.11661#Thmproposition5), and the moment radiusr=O\(B2log\(d/δ\)/n\)r=O\\bigl\(B^\{2\}\\sqrt\{\\log\(d/\\delta\)/n\}\\bigr\)of Lemma[4](https://arxiv.org/html/2608.11661#Thmlemma4)into the second\. The result is exactly \([74](https://arxiv.org/html/2608.11661#A5.E74)\):
max\{‖f^k−fk‖,‖g^k−gk‖\}=O\(‖F^−F‖L22\+rΔ\)\.\\max\\bigl\\\{\\\|\\hat\{f\}\_\{k\}\-f\_\{k\}\\\|,\\ \\\|\\hat\{g\}\_\{k\}\-g\_\{k\}\\\|\\bigr\\\}=O\\\!\\Bigl\(\\frac\{\\sqrt\{\\\|\\hat\{F\}\-F\\\|\_\{L^\{2\}\}^\{2\}\}\+r\}\{\\Delta\}\\Bigr\)\.\(75\)Both terms are inversely proportional to the gapΔ\\Deltaand decay asn−1/2n^\{\-1/2\}, which is the statement that identification and estimation are governed by the same spectral gap\.
## Appendix FExperimental Details and Additional Results
This appendix records the complete setups of the experiments of Section[5](https://arxiv.org/html/2608.11661#S5)and the additional results that the main text summarizes\. All numbers are produced by the scripts in the companion repository, whose data files are grouped per study\. We state the metrics alongside the tables, because a recurring theme is that gauge\-invariant diagnostics cannot separate normalizations while gauge\-dependent ones can\.
### F\.1Controlled synthetic study
#### Targets and training\.
We fixF\(u,v\)=∑k≤dσkak\(u\)bk\(v\)F\(u,v\)=\\sum\_\{k\\leq d\}\\sigma\_\{k\}\\,a\_\{k\}\(u\)b\_\{k\}\(v\)withak,bka\_\{k\},b\_\{k\}shifted Legendre polynomials, orthonormal on\[0,1\]\[0,1\], and a prescribed spectrum\{σk\}\\\{\\sigma\_\{k\}\\\}chosen distinct and separated\. Encoders are linear maps over a fixed Legendre feature map,f\(u\)=U⊤φ\(u\)f\(u\)=U^\{\\top\}\\varphi\(u\)withU∈ℝp×dU\\in\\mathbb\{R\}^\{p\\times d\}, so the encoder realization error of Assumption[1](https://arxiv.org/html/2608.11661#Thmassumption1)is zero and the normalization is the only variable across runs\. Pointsu,vu,vare drawn on a common grid ofnnsites per marginal\. Training uses Adam with global\-norm gradient clipping; the bare bilinear objective is nonconvex and stiff, and plain gradient descent diverges\. Analytic gradients are checked against finite differences to about10−910^\{\-9\}\. Every scheme runs four seeds\.
#### The seven normalization schemes\.
The main text reports the headline of this study; Table[1](https://arxiv.org/html/2608.11661#S5.T1)in Section[5\.1](https://arxiv.org/html/2608.11661#S5.SS1)gives the full panel, shown graphically in Figure[2](https://arxiv.org/html/2608.11661#A6.F2)\. Fit MSE is the squared regression error\. Alignment is the mean over modes of\|⟨f^k,ak⟩\|\|\\langle\\hat\{f\}\_\{k\},a\_\{k\}\\rangle\|after optimally matching permutation and sign, reported for the two encoders\. Cond is the gauge\-dependent diagnosticmax\(condΣf,condΣg\)\\max\(\\mathrm\{cond\}\\,\\Sigma\_\{f\},\\mathrm\{cond\}\\,\\Sigma\_\{g\}\)\. Cross\-seed is the gauge distance between solutions from different seeds, the mean over pairs of the Procrustes\-optimal embedding distance\. Only whitening reaches alignment1\.0001\.000with cross\-seed distance0\.0000\.000, matching Theorem[2](https://arxiv.org/html/2608.11661#Thmtheorem2); every looser gauge stalls at alignment0\.530\.53to0\.650\.65with cross\-seed drift0\.280\.28to0\.350\.35, the residualO\\mathrm\{O\}of cosine and the fullGLd\\mathrm\{GL\}\_\{d\}of none leaving the modes rotationally entangled \(Theorem[4](https://arxiv.org/html/2608.11661#Thmtheorem4)\)\. Whitening drives Cond to its floorσ12/σd2≈9\.5\\sigma\_\{1\}^\{2\}/\\sigma\_\{d\}^\{2\}\\approx 9\.5\(Corollary[3](https://arxiv.org/html/2608.11661#Thmcorollary3)\)\. The two\-sided nonnegative gauge stays at Fit MSE4\.254\.25: positivity forces⟨f,g⟩≥0\\langle f,g\\rangle\\geq 0, the sign\-changing target is out of class \(Theorem[5](https://arxiv.org/html/2608.11661#Thmtheorem5)\), and identifiability is bought at the cost of excluding negative interactions\. One measurement subtlety, exploited throughout:cond\(ΣfΣg\)\\mathrm\{cond\}\(\\Sigma\_\{f\}\\Sigma\_\{g\}\)is gauge invariant, its eigenvalues equal theσk2\\sigma\_\{k\}^\{2\}by Lemma[2](https://arxiv.org/html/2608.11661#Thmlemma2), so it cannot separate normalizations and is not reported as a diagnostic\.
Figure 2:The seven normalization schemes of Table[1](https://arxiv.org/html/2608.11661#S5.T1), means over four seeds\. Left: mode alignment with the true basis \(dashed line:11\)\. Middle: the conditioning diagnosticmax\(condΣf,condΣg\)\\max\(\\mathrm\{cond\}\\,\\Sigma\_\{f\},\\mathrm\{cond\}\\,\\Sigma\_\{g\}\)on a log scale, whose floor isσ12/σd2≈9\.5\\sigma\_\{1\}^\{2\}/\\sigma\_\{d\}^\{2\}\\approx 9\.5\. Right: cross\-seed gauge distance \(the short tick marks the exact zero of hard whitening\)\. Abbreviations: l2 = L2 normalization, coord = per\-coordinate standardization, cos = cosine normalization, nonneg = two\-sided nonnegativity, w\-soft = soft whitening, w\-hard = hard whitening\.
#### The gap collapse\.
The gap sweeps use the finite\-sample whitening estimator of Appendix[E](https://arxiv.org/html/2608.11661#A5)directly, empirical second\-moment whitening followed by eigendecomposition, rather than a trained optimizer: with Legendre features the target is represented exactly, and the training identification error has a sharp threshold aroundn≈300n\\approx 300that would hide the1/Δ1/\\Deltalaw, whereas the perturbation mechanism of the estimator exposes it cleanly\. Table[2](https://arxiv.org/html/2608.11661#S5.T2)in Section[5\.1](https://arxiv.org/html/2608.11661#S5.SS1)reports the two sweeps\. The eight points collapse onto a single constant,err⋅Δ⋅n≈2\.48\\mathrm\{err\}\\cdot\\Delta\\cdot\\sqrt\{n\}\\approx 2\.48with coefficient of variation0\.120\.12: gap and sample size enter the identification error only through the product, the concrete face of Corollaries[6](https://arxiv.org/html/2608.11661#Thmcorollary6)and[3](https://arxiv.org/html/2608.11661#Thmcorollary3)\. Figure[3](https://arxiv.org/html/2608.11661#A6.F3)shows the collapse\.
Figure 3:The gap collapse\. Mode identification error of the finite\-sample whitening estimator: against the inverse spectral gap1/Δ1/\\Deltaat fixednn\(left\) and against sample sizennat fixedΔ\\Delta\(center\)\. Right: the product error⋅Δ⋅n\\,\\cdot\\Delta\\cdot\\sqrt\{n\}per run across both sweeps, near\-constant at≈2\.48\\approx 2\.48\.
#### The sample\-complexity rate law\.
The linear setting is rank\-ddmatrix sensing,yi=ui⊤Θ⋆vi\+εiy\_\{i\}=u\_\{i\}^\{\\top\}\\Theta^\{\\star\}v\_\{i\}\+\\varepsilon\_\{i\}withrankΘ⋆=d\\rank\\Theta^\{\\star\}=d, and Theorem[7](https://arxiv.org/html/2608.11661#Thmtheorem7)predicts𝔼‖Θ^−Θ⋆‖F2≍σε2d\(p\+q\)/n\\mathbb\{E\}\\\|\\hat\{\\Theta\}\-\\Theta^\{\\star\}\\\|\_\{F\}^\{2\}\\asymp\\sigma\_\{\\varepsilon\}^\{2\}d\(p\+q\)/n\. We sweep each variable with the others fixed and fit log\-log slopes:nnatp=q=4p=q=4,d=2d=2gives−1\.03\-1\.03\(theory−1\-1\);ddatp=q=16p=q=16,n=2048n=2048gives\+0\.89\+0\.89\(theory\+1\+1\);p\+qp\+qatn=104n=10^\{4\},d=4d=4gives\+1\.24\+1\.24\(theory\+1\+1\)\. The excess of the last slope is accounted for by the Marchenko\-Pastur finite\-sample correction of the Gaussian design: measured ratios toσε2d\(2p−d\)/n\\sigma\_\{\\varepsilon\}^\{2\}d\(2p\-d\)/nare1\.00,1\.05,0\.98,1\.09,1\.131\.00,1\.05,0\.98,1\.09,1\.13acrossp=q=8,…,32p=q=8,\\dots,32, tracking the correction factorpq/\(n−pq\)pq/\(n\-pq\), which grows from1\.011\.01to1\.111\.11over the same range\. Gradient descent on the factored pair\(U,V\)\(U,V\), the estimator class of this paper, is warm\-started from the truncated least\-squares factors and reproduces the least\-squares slope,−1\.01\-1\.01, on the same grid\. The sweep rejects the loose general bound of Remark[6](https://arxiv.org/html/2608.11661#Thmremark6), which scales asd3/2\(p\+q\)/nd^\{3/2\}\\sqrt\{\(p\+q\)/n\}and predicts slope−1/2\-1/2innn: the1/n1/ndependence on samples is real \(Figure[1](https://arxiv.org/html/2608.11661#S5.F1)\)\.
#### The negative results\.
Tables[4](https://arxiv.org/html/2608.11661#S5.T4)and[4](https://arxiv.org/html/2608.11661#S5.T4)in Section[5\.1](https://arxiv.org/html/2608.11661#S5.SS1)give the two regimes; Figure[4](https://arxiv.org/html/2608.11661#A6.F4)shows both graphically\. On the boolean equality kernelF=𝟏\[u=v\]F=\\mathbf\{1\}\[u=v\]on\{±1\}m\\\{\\pm 1\\\}^\{m\}, the embedding rank needed for10%10\\%error tracksd≈0\.9⋅2md\\approx 0\.9\\cdot 2^\{m\}, exponential inmm, while the early\-interaction threshold on⟨u,v⟩\\langle u,v\\ranglerepresentsFFexactly with2m\+12m\+1parameters \(Theorems[8](https://arxiv.org/html/2608.11661#Thmtheorem8)and[9](https://arxiv.org/html/2608.11661#Thmtheorem9)\)\. On the narrowband periodic\-Gaussian kernelFh\(u,v\)=ρh\(u−v\)F\_\{h\}\(u,v\)=\\rho\_\{h\}\(u\-v\)of bandwidthhhon the circle, the required rankdεd\_\{\\varepsilon\}grows asΘ\(1/h\)\\Theta\(1/h\)as the kernel approaches the diagonal \(Remark[7](https://arxiv.org/html/2608.11661#Thmremark7)\)\.
Figure 4:The negative results\. Left: on the boolean equality kernel, parameters needed to reach10%10\\%relative error against the number of input bitsmm, for the dual\-encoder head \(rankdd\) and the early\-interaction threshold\. Center: for the narrowband periodic\-Gaussian kernel, the dual\-encoder rank needed for10%10\\%error against the inverse bandwidth1/h1/h, with a linear fit\. Right: relative squared error of the dual\-encoder head against its rankdd, one curve per bandwidthhh, with the10%10\\%line marked\.Figure 5:The flat\-spectrum error floor on a trained DeepONet head \(Theorems[8](https://arxiv.org/html/2608.11661#Thmtheorem8)and[9](https://arxiv.org/html/2608.11661#Thmtheorem9)\)\. Left: normalized interaction spectraσk\\sigma\_\{k\}for three operators, from exponentially decaying \(heat\) to polynomially decaying \(antiderivative\) to exactly flat \(boxcar block operator\)\. Right: test relative error of a rank\-dddual\-encoder head versusdd, from an SVD warm start and from random initialization, against the predicted floor1−d/m1\-d/m; the dashed line is an early\-interaction MLP baseline\.
### F\.2DeepONet
This study carries the identifiability results of Section[3](https://arxiv.org/html/2608.11661#S3)onto a real dual\-encoder architecture\. We train branch\-trunk DeepONet heads end to end on four operators and ask three questions, one per paragraph below: whether the rank\-selection error curve tracks the decay of the operator spectrum \(Corollary[5](https://arxiv.org/html/2608.11661#Thmcorollary5)\); whether post\-hoc whitening, applied after training without touching the loss, lifts the learned trunk basis to the analytic operator eigenmodes and collapses the cross\-seed gauge distance \(Theorem[2](https://arxiv.org/html/2608.11661#Thmtheorem2)\); and whether the whitened gauge improves out\-of\-distribution error \(Corollary[3](https://arxiv.org/html/2608.11661#Thmcorollary3)\)\. Two operators, the heat semigroup and the Volterra antiderivative, have closed\-form spectra and provide exact ground truth for the recovered basis; two, viscous Burgers and Darcy flow, have only empirical spectral proxies and are used for the qualitative claims\. The main\-text Section[5](https://arxiv.org/html/2608.11661#S5)reports the conclusions; the tables and figures below give the full numbers\.
#### Architecture and training\.
The branch and trunk networks are MLPs of width128128and depth33, outputting embeddings of dimensiond=32d=32; the rank\-selection sweep of Table[7](https://arxiv.org/html/2608.11661#A6.T7)variesddfrom 2 to 64\. Training runs 300 epochs with Adam at learning rate10−310^\{\-3\}and cosine annealing, batch size 256, and five seeds per configuration\. Inputs are Gaussian random fields\. The four operators are the heat semigroupetΔe^\{t\\Delta\}\(analytic singular valuesσk=e−\(kπ\)2t\\sigma\_\{k\}=e^\{\-\(k\\pi\)^\{2\}t\}, exponential\), the antiderivative or Volterra operator \(σk=1/\(\(k−12\)π\)\\sigma\_\{k\}=1/\(\(k\-\\tfrac\{1\}\{2\}\)\\pi\), polynomial\), and viscous Burgers and Darcy flow, whose spectra are empirical proxies, labeled as such\. Whitening is realized in two stages: the stable soft\-whitening penalty during training, the relaxation of Theorem[6](https://arxiv.org/html/2608.11661#Thmtheorem6)and closely related to the Barlow Twins objective, and hard whitening post hoc on the active rank\-rrsubspace\. The two\-stage design is forced by the data: the interaction spectra decay fast, leaving an effective rank of 3 to 8 against an embedding width of 32, and in\-forward hard whitening inverts about 25 near\-zero covariance eigenvalues and backpropagates through a degenerate eigendecomposition, exactly the small\-gap failure predicted by the block\-stability analysis of Corollary[7](https://arxiv.org/html/2608.11661#Thmcorollary7)in Appendix[E](https://arxiv.org/html/2608.11661#A5)\. Because whitening leaves⟨b,t⟩\\langle b,t\\rangleunchanged, applying it after training is exact, and projecting to the active subspace first avoids inverting dead directions\. The empirical interaction spectrum is estimated as the square roots of the eigenvalues ofΣb1/2ΣtΣb1/2\\Sigma\_\{b\}^\{1/2\}\\Sigma\_\{t\}\\Sigma\_\{b\}^\{1/2\}, computed witheigvalsh\\mathrm\{eigvalsh\}; this avoids the numerically unstable SVD of the ill\-conditioned productΣb1/2Σt1/2\\Sigma\_\{b\}^\{1/2\}\\Sigma\_\{t\}^\{1/2\}and agrees with the direct computation to10−910^\{\-9\}on well\-conditioned instances\.
#### Spectra and rank selection\.
The normalized empirical spectra reproduce the dichotomy qualitatively, heat1,0\.22,0\.028,…1,0\.22,0\.028,\\dotsand antiderivative1,0\.14,0\.042,…1,0\.14,0\.042,\\dots\. Because the network allocates no capacity to near\-null modes, the learned spectra decay faster than the analytic ones, so we claim recovery of the dominant modes and of the decay type, not a mode\-by\-mode match\. Table[7](https://arxiv.org/html/2608.11661#A6.T7)gives the rank sweep\. The curves are flat pastd≈8d\\approx 8and the argmin is dominated by noise, so we do not report a single optimal rank; the honest quantity is the drop fromd=4d=4tod=8d=8,5%5\\%for heat \(early saturation, small optimal rank\) versus31%31\\%for antiderivative \(more modes needed\), consistent with Corollary[5](https://arxiv.org/html/2608.11661#Thmcorollary5)\. This is qualitative support, and the figure is labeled accordingly\.
Table 7:Rank selection, test relativeL2L^\{2\}error versus trunk widthdd\. The drop fromd=4d=4tod=8d=8is the claimed quantity:5%5\\%for heat versus31%31\\%for antiderivative \(Corollary[5](https://arxiv.org/html/2608.11661#Thmcorollary5)\)\.
#### Basis recovery\.
Table[8](https://arxiv.org/html/2608.11661#A6.T8)gives the full per\-gauge numbers behind Figure[6](https://arxiv.org/html/2608.11661#A6.F6)\. Alignment is the mean over modes of\|⟨t^k,βk⟩\|\|\\langle\\hat\{t\}\_\{k\},\\beta\_\{k\}\\rangle\|after optimal permutation and sign matching against the analytic left singular functions, before and after the post\-hoc whitening gauge fix; cross\-seed is the mean pairwise gauge distance between seeds\. Whitening lifts the loose gauges from0\.50\.5to0\.60\.6to0\.780\.78to0\.940\.94and collapses the cross\-seed distance by roughly an order of magnitude; for heat the recovered modes are the analytic Fourier basis2sin\(kπy\)\\sqrt\{2\}\\sin\(k\\pi y\), a hard ground truth\. On the nonlinear benchmarks, where no analytic basis exists, whitening still tightens the cross\-seed distance for Burgers,0\.30→0\.080\.30\\to 0\.08; Darcy is the weakest case,0\.56→0\.330\.56\\to 0\.33, and its proxy spectrum is the noisiest, which the figure notes\.
Figure 6:Trunk\-encoder basis quality before and after post\-hoc whitening \(Theorem[2](https://arxiv.org/html/2608.11661#Thmtheorem2)\)\. Each group on the horizontal axis corresponds to one operator \(heat, anti = antiderivative, Burg = Burgers, Darcy\) trained under one gauge constraint \(none, cosine, or whiten = soft whitening\)\. Left: alignment of the learned trunk basis with the analytic operator eigenmodes \(11is exact\)\. Right: distance between bases learned from different random seeds \(lower indicates greater stability\)\.Table 8:Basis recovery per gauge, five seeds\. Alignment: mean mode alignment against the analytic left singular functions, raw and after post\-hoc whitening\. Cross\-seed: mean pairwise gauge distance, raw and after whitening\. The effect is Theorem[2](https://arxiv.org/html/2608.11661#Thmtheorem2)on a trained model\.
#### The payoff of gauge\-fixing\.
Table[9](https://arxiv.org/html/2608.11661#A6.T9)gives the in\-distribution and out\-of\-distribution relativeL2L^\{2\}errors of the none and soft\-whitening gauges; the out\-of\-distribution test set draws inputs from a different smoothness regime of the same operator family\. Soft whitening attains the lowest error in all eight cells, with the out\-of\-distribution gain largest,2\.4×2\.4\\timesfor heat\. Two caveats are recorded with the table\. First, soft whitening does not reduce seed\-to\-seed error variance; if anything it increases it, heat standard deviation0\.0005→0\.00360\.0005\\to 0\.0036and Darcy0\.0017→0\.01810\.0017\\to 0\.0181, so the stability it confers is in the identified basis of Table[8](https://arxiv.org/html/2608.11661#A6.T8), not in the error\. Second, the soft\-whitening training loss carries the decorrelation penalty and its curve is not directly comparable to the unpenalized gauges, so we make no claim about convergence speed\.
Table 9:The payoff of gauge\-fixing, five seeds: relativeL2L^\{2\}error in distribution and out of distribution, none versus soft whitening\. Soft whitening is best in all cells; the out\-of\-distribution gain is the headline \(Corollary[3](https://arxiv.org/html/2608.11661#Thmcorollary3)\)\.
#### The flat\-spectrum anchor\.
The block\-boxcar operatorG\(a\)\(y\)=cj\(y\)G\(a\)\(y\)=c\_\{j\(y\)\}, the output is the block mean of the input field at the block containingyy, withm=16m=16blocks, has an exactly flat interaction spectrum: sixteen equal singular values, the seventeenth at10−1610^\{\-16\}\. Table[10](https://arxiv.org/html/2608.11661#A6.T10)gives the trained relative error of the dual\-encoder head with linear encoders, from both the SVD warm start and random initialization, against the floor1−d/m1\-d/mof Theorem[8](https://arxiv.org/html/2608.11661#Thmtheorem8); the worst deviation is0\.0140\.014, the finite\-sample excess, and gradient descent reaches the floor from scratch\. The early\-interaction comparator, a ReLU MLP onconcat\(a,onehot\(y\)\)\\mathrm\{concat\}\(a,\\mathrm\{onehot\}\(y\)\)with12\.412\.4k parameters, four times the budget of thed=16d=16head, reaches0\.013±0\.0010\.013\\pm 0\.001on the same test set \(Theorem[9](https://arxiv.org/html/2608.11661#Thmtheorem9)\)\. One caveat belongs in the setup: with only 8k training samples the comparator memorizes, training MSE4×10−54\\times 10^\{\-5\}and test error1\.71\.7; at 60k samples, where parameters are far fewer than samples, memorization is impossible and the escape shown in Figure[5](https://arxiv.org/html/2608.11661#A6.F5)is an honest generalization result\. The comparison deliberately pits the floor against a well\-trained early model\.
Table 10:The flat\-spectrum anchor: trained relative error of a linear\-encoder head on the block\-boxcar operator versus the floor1−d/m1\-d/mof Theorem[8](https://arxiv.org/html/2608.11661#Thmtheorem8),[9](https://arxiv.org/html/2608.11661#Thmtheorem9)\. The two initializations agree, and the error hugs the floor from either start\.
### F\.3CLIP
This study tests the gauge view on heterogeneous modalities and on pretrained models that never saw each other\. It proceeds in two stages\. The white\-box synthetic kernel supplies exact ground truth for the mechanism: under linear encoders and white inputs the interaction spectrum equals the singular values ofWf⊤WgW\_\{f\}^\{\\top\}W\_\{g\}, which is gauge invariant, and two independently trained encoders are two random gauges of the same target; on this kernel we verify that the spectrum is invariant and equals the truth, that per\-dimension concept probes do not transfer while the canonical subspace does, and that whitening plus gap recovers the text modes exactly\. The real CLIP pairs then answer whether the same statements hold for contrastive models trained in the wild: is the interaction spectrum shared across independently pretrained models, are per\-dimension probes not, and is the residual relationship a single rotation that whitening removes, exposing interpretable concept axes \(Theorem[4](https://arxiv.org/html/2608.11661#Thmtheorem4)\)? The main\-text Section[5](https://arxiv.org/html/2608.11661#S5)reports the conclusions; the tables and figures below give the full numbers\.
#### The synthetic multimodal kernel\.
The white\-box counterpart isF⋆\(u,v\)=∑kρk\(pk⊤u\)\(qk⊤v\)F^\{\\star\}\(u,v\)=\\sum\_\{k\}\\rho\_\{k\}\\,\(p\_\{k\}^\{\\top\}u\)\(q\_\{k\}^\{\\top\}v\)withu∼𝒩\(0,Ip\)u\\sim\\mathcal\{N\}\(0,I\_\{p\}\),v∼𝒩\(0,Iq\)v\\sim\\mathcal\{N\}\(0,I\_\{q\}\), orthogonal canonical directionsP,QP,Q, and distinct, gapped canonical correlationsρ=\(0\.9,0\.75,0\.6,0\.45,0\.3,0\.15\)\\rho=\(0\.9,0\.75,0\.6,0\.45,0\.3,0\.15\)\. The key identity: under linear encoders and white inputs the interaction spectrum equals the singular values ofB=Wf⊤WgB=W\_\{f\}^\{\\top\}W\_\{g\}, which is explicitly gauge invariant,f→Aff\\to Af,g→A−⊤gg\\to A^\{\-\\top\}gleavesWf⊤WgW\_\{f\}^\{\\top\}W\_\{g\}unchanged, and two independently trained encoders are two random gaugesA1,A2A\_\{1\},A\_\{2\}\. Table[11](https://arxiv.org/html/2608.11661#A6.T11)collects the numbers: the spectrum is invariant across gauges to3×10−163\\times 10^\{\-16\}and equals the trueρ\\rhoto4×10−164\\times 10^\{\-16\}; per\-dimension concept probes transfer at0\.430\.43, the best\-matched axis reaching only0\.650\.65, while the canonical subspace aligns to1\.00001\.0000; whitening plus gap recovers the true text modes at\|cos\|=1\.000\|\\cos\|=1\.000, against0\.5540\.554raw and0\.5810\.581under cosine, which is stuck at the residualO\\mathrm\{O\}; and cross\-encoder frame stability rises from0\.6840\.684to1\.0001\.000\. The distance between the0\.580\.58of cosine and the1\.001\.00of whitening is, numerically, the contribution of whitening beyond the decorrelation that self\-supervised methods already perform, precisely the quantity Theorem[2](https://arxiv.org/html/2608.11661#Thmtheorem2)isolates\.
Table 11:Synthetic multimodal kernel, ground truth known\. The interaction spectrum is the unique gauge\-invariant content; whitening plus gap pins the modes, cosine does not\.
#### Real CLIP\.
Two backbone pairs, each two independent gauges of one similarity function: ViT\-B/32 \(OpenAI\) versus ViT\-B/32 \(LAION\-2B\), the same architecture with different pretraining data, and ViT\-B/32 \(OpenAI\) versus RN50 \(OpenAI\), different architectures\. Image features come from CIFAR\-100 test images, 49 fine classes spanning animals, vehicles, plants, and household objects, 64 images per class; text features are prompt\-ensembled class embeddings plus attribute contrast directions; both towers are projected to a common 32\-dimensional working space, the image PCA top\-dddirections\. Whitening is realized via canonical correlation analysis with shrinkage covariance, the finite\-sample counterpart of the Theorem[2](https://arxiv.org/html/2608.11661#Thmtheorem2)gauge fix\.
Table[12](https://arxiv.org/html/2608.11661#A6.T12)and Figure[7](https://arxiv.org/html/2608.11661#A6.F7)give the spectrum comparison across backbone pairs: near\-identical spectra, agreement best at the top\. Per\-dimension probe transfer is0\.280\.28without alignment and0\.920\.92after fitting a single global orthogonal map: the two embeddings are related by oneO\\mathrm\{O\}rotation, the residual group Theorem[4](https://arxiv.org/html/2608.11661#Thmtheorem4)predicts for the cosine gauge, and no coordinatewise repair exists short of discovering that rotation\. Table[13](https://arxiv.org/html/2608.11661#A6.T13)gives the group\-signal regression of true labels on the top\-6 mode activations: the whitened span contains the animate and natural axes, the cosine span contains only the animate axis, and per\-mode statistics show why, whitening concentrates the animate signal in a single mode \(\|t\|=3\.1\|t\|=3\.1\) while cosine spreads it across several\. Figure[8](https://arxiv.org/html/2608.11661#A6.F8)shows the canonical text modes as concept axes: mode 1 separates animals \(leopard, bear, wolf, camel\) from flowers and fruit, mode 2 is the vehicle axis \(bus, streetcar, pickup truck, train\), mode 3 separates plants from household objects, and mode 5 the food axis \(orange, apple, pear, sweet pepper\); the cosine principal coordinates are semantic mixtures by comparison\. Cross\-backbone frame stability, matching modes by their backbone\-independent class\-signature profiles, favors whitening in both pairs,0\.770\.77versus0\.720\.72for the same\-architecture pair and0\.840\.84versus0\.620\.62across architectures\.
Table 12:Interaction spectra across independently pretrained CLIP models\. Pearson correlation and normalized L1 difference of the two spectra; per\-dimension probe transfer raw and after fitting the single residualO\\mathrm\{O\}rotation \(Theorem[4](https://arxiv.org/html/2608.11661#Thmtheorem4)\)\.Figure 7:Shared structure across independently trained CLIP models \(Theorem[4](https://arxiv.org/html/2608.11661#Thmtheorem4)\)\. Left: normalized interaction spectraσk\\sigma\_\{k\}for three CLIP encoders \(a ViT\-B/32 reference against a LAION\-trained ViT\-B/32 and an RN50\), with Pearson correlations to the reference in the legend\. Right: accuracy of a linear concept probe trained on one model and applied to the other, using raw coordinates and after fitting a single shared rotation, for same\-architecture and cross\-architecture pairs\.Table 13:Group\-signal concentration on the reference backbone:R2R^\{2\}of regressing true group labels on the top\-6 mode activations, whitened canonical modes versus cosine principal coordinates\. Cross\-backbone frame stability is reported in the text\.Figure 8:Whitened canonical text modes \(rows\) over the 49 CIFAR\-100 classes \(columns\) of the reference backbone: modes read as concept axes \(animals, vehicles, plants, food\), while the cosine principal coordinates are semantic mixtures\.
#### What whitening buys for CLIP in practice\.
None of the steps above requires retraining: whitening is a post\-hoc second\-moment transform, it leaves the inner\-product scores unchanged, and at inference time it costs one fixed linear map\. It can therefore be applied as an analysis layer on any deployed CLIP model, and the results above suggest three concrete payoffs\. First, model alignment: because independently pretrained pairs are related by a single residual rotation, whitening both models into the shared canonical frame makes their coordinates comparable dimension by dimension, a prerequisite for model merging, distillation, and cross\-model probing that raw cosine coordinates do not offer\. Second, interpretability with known limits: the whitened axes read as concept axes \(animals, vehicles, plants, and food in the CIFAR\-100 probe\), and the spectral gap quantifies in advance how much interpretability any basis can carry, so an analyst knows where the axes are trustworthy and where the small\-gap failure of Appendix[E](https://arxiv.org/html/2608.11661#A5)leaves them partially mixed\. Third, a spectral summary of the model itself: the interaction spectrum is a gauge\-invariant description of what a CLIP pair can discriminate, and its decay predicts the effective interaction rank \(Corollary[5](https://arxiv.org/html/2608.11661#Thmcorollary5)\) and hence how many concept axes are worth reading out at all\. A natural next step is to use the whitened canonical frame as the initialization or constraint for fine\-tuning, so that new axes are learned in coordinates that are already interpretable and transferable across models\.
#### Summary\.
The synthetic study establishes the three claims where ground truth is exact: whitening uniquely fixes the gauge, the gap controls identification through the productΔn\\Delta\\sqrt\{n\}\(collapse constant2\.482\.48\), and flat spectra force exponential late\-interaction cost\. DeepONet carries identifiability onto a real architecture, recovering the analytic eigenbasis with0\.920\.92to0\.940\.94alignment and an out\-of\-distribution payoff up to2\.4×2\.4\\times; CLIP extends the claim to heterogeneous modalities, the spectrum shared across independently pretrained models while per\-dimension probes are not\. On a flat spectrum the head is floored at1−d/m1\-d/mwhile an early\-interaction model reaches0\.0130\.013: the interaction spectrum decides whether a multiplicative dual\-encoder head is the right tool\.
## Appendix GIndex of results
Table[14](https://arxiv.org/html/2608.11661#A7.T14)collects the paper’s numbered results with a one\-line statement of what each establishes and where it is stated or proved\. The results are grouped by the four questions the paper answers, with the finite\-sample lemmas of Appendix[E](https://arxiv.org/html/2608.11661#A5)last; within each group the numbering follows the order of appearance, and the table doubles as a reading map for the appendices\. Remarks and the technical lemmas internal to proofs are not indexed separately; they are cited where they are used\.
Table 14:Index of the paper’s numbered results\.ResultWhat it establishesWhereUnification and approximation \(Section[2](https://arxiv.org/html/2608.11661#S2)\)Definition[1](https://arxiv.org/html/2608.11661#Thmdefinition1)Interaction spectrum, rank, and modes as the singular data of the targetSection[2](https://arxiv.org/html/2608.11661#S2)Theorem[1](https://arxiv.org/html/2608.11661#Thmtheorem1)Approximation error decomposes into a truncation term and an encoder\-realization termSection[2](https://arxiv.org/html/2608.11661#S2)Assumption[1](https://arxiv.org/html/2608.11661#Thmassumption1)Encoder classes contain approximations of the scaled interaction modesAppendix[B](https://arxiv.org/html/2608.11661#A2)Proposition[1](https://arxiv.org/html/2608.11661#Thmproposition1)Evaluating a rank\-ddhead on all pairs costsO\(\(n\+m\)Cnet\+nmd\)O\(\(n\+m\)\\,C\_\{\\mathrm\{net\}\}\+nm\\,d\)Appendix[B](https://arxiv.org/html/2608.11661#A2)Lemma[1](https://arxiv.org/html/2608.11661#Thmlemma1)Representability ofFFby a rank\-ddhead iffrank\(TF\)≤d\\rank\(T\_\{F\}\)\\leq dAppendix[B](https://arxiv.org/html/2608.11661#A2)Corollary[1](https://arxiv.org/html/2608.11661#Thmcorollary1)The truncation term is the best possible error of any rank\-ddheadAppendix[B](https://arxiv.org/html/2608.11661#A2)Theorem[3](https://arxiv.org/html/2608.11661#Thmtheorem3)Target smoothness controls the decay of the truncation tailAppendix[B](https://arxiv.org/html/2608.11661#A2)Proposition[2](https://arxiv.org/html/2608.11661#Thmproposition2)The dual head escapes the curse of the joint input dimensionAppendix[B](https://arxiv.org/html/2608.11661#A2)Identifiability as gauge\-fixing \(Section[3](https://arxiv.org/html/2608.11661#S3)\)Definition[2](https://arxiv.org/html/2608.11661#Thmdefinition2)Gauge action ofGLd\\mathrm\{GL\}\_\{d\}on encoder pairs leaves the represented function unchangedSection[3](https://arxiv.org/html/2608.11661#S3)Theorem[2](https://arxiv.org/html/2608.11661#Thmtheorem2)Under a spectral gap, whitening identifies the modes up to permutation and signSection[3](https://arxiv.org/html/2608.11661#S3)Proposition[3](https://arxiv.org/html/2608.11661#Thmproposition3)Per\-coordinate standardization pins the coordinate scales but no moreAppendix[C](https://arxiv.org/html/2608.11661#A3)Lemma[2](https://arxiv.org/html/2608.11661#Thmlemma2)The eigenvalues ofΣfΣg\\Sigma\_\{f\}\\Sigma\_\{g\}are gauge invariants, equal to theσk2\\sigma\_\{k\}^\{2\}at any optimumAppendix[C](https://arxiv.org/html/2608.11661#A3)Proposition[4](https://arxiv.org/html/2608.11661#Thmproposition4)Unfixed gauge renders the risk flat along the orbit and optimization ill\-posedAppendix[C](https://arxiv.org/html/2608.11661#A3)Theorem[4](https://arxiv.org/html/2608.11661#Thmtheorem4)Cosine normalization leaves the full orthogonal group as residual gaugeAppendix[C](https://arxiv.org/html/2608.11661#A3)Corollary[2](https://arxiv.org/html/2608.11661#Thmcorollary2)Individual coordinates of a cosine\-normalized embedding carry no intrinsic meaningAppendix[C](https://arxiv.org/html/2608.11661#A3)Theorem[5](https://arxiv.org/html/2608.11661#Thmtheorem5)Two\-sided nonnegativity gives NMF\-style identifiability, at the cost of excluding negative modesAppendix[C](https://arxiv.org/html/2608.11661#A3)Corollary[3](https://arxiv.org/html/2608.11661#Thmcorollary3)The spectral gap controls conditioning, identifiability, and convergence simultaneouslyAppendix[C](https://arxiv.org/html/2608.11661#A3)Theorem[6](https://arxiv.org/html/2608.11661#Thmtheorem6)A quadratic whitening penalty inherits the identifiability of the hard constraintAppendix[C](https://arxiv.org/html/2608.11661#A3)Estimation and rank selection \(Section[4](https://arxiv.org/html/2608.11661#S4)\)Lemma[3](https://arxiv.org/html/2608.11661#Thmlemma3)The Rademacher complexity of the inner\-product class collapses to the sum of the encoder complexitiesAppendix[D](https://arxiv.org/html/2608.11661#A4)Corollary[4](https://arxiv.org/html/2608.11661#Thmcorollary4)Excess risk of empirical risk minimization on the head classAppendix[D](https://arxiv.org/html/2608.11661#A4)Theorem[7](https://arxiv.org/html/2608.11661#Thmtheorem7)Linear heads are rank\-ddmatrix sensing with minimax rateσε2d\(p\+q\)/n\\sigma\_\{\\varepsilon\}^\{2\}d\(p\+q\)/nAppendix[D](https://arxiv.org/html/2608.11661#A4)Corollary[5](https://arxiv.org/html/2608.11661#Thmcorollary5)Optimal rank grows logarithmically in the sample budget for smooth targetsAppendix[D](https://arxiv.org/html/2608.11661#A4)The usability criterion \(Section[4](https://arxiv.org/html/2608.11661#S4)\)Theorem[8](https://arxiv.org/html/2608.11661#Thmtheorem8)A flat spectrum forces every rank\-ddhead to relative error1−d/N1\-d/NAppendix[D](https://arxiv.org/html/2608.11661#A4)Theorem[9](https://arxiv.org/html/2608.11661#Thmtheorem9)Early\-interaction models escape the flat\-spectrum floorAppendix[D](https://arxiv.org/html/2608.11661#A4)Theorem[10](https://arxiv.org/html/2608.11661#Thmtheorem10)Usability criterion: decide by the measured spectrum whether the head can succeedAppendix[D](https://arxiv.org/html/2608.11661#A4)Finite\-sample identification \(Appendix[E](https://arxiv.org/html/2608.11661#A5)\)Lemma[4](https://arxiv.org/html/2608.11661#Thmlemma4)Matrix Bernstein concentration for empirical second momentsAppendix[E](https://arxiv.org/html/2608.11661#A5)Lemma[5](https://arxiv.org/html/2608.11661#Thmlemma5)Stability of the empirical whitening map under moment perturbationAppendix[E](https://arxiv.org/html/2608.11661#A5)Theorem[11](https://arxiv.org/html/2608.11661#Thmtheorem11)Finite\-sample mode identification with error of orderr/Δr/\\DeltaAppendix[E](https://arxiv.org/html/2608.11661#A5)Corollary[6](https://arxiv.org/html/2608.11661#Thmcorollary6)Sample complexity of the whitening estimator:n≳B4log\(d/δ\)/\(Δ2ε2\)n\\gtrsim B^\{4\}\\log\(d/\\delta\)/\(\\Delta^\{2\}\\varepsilon^\{2\}\)Appendix[E](https://arxiv.org/html/2608.11661#A5)Corollary[7](https://arxiv.org/html/2608.11661#Thmcorollary7)Block stability of mode recovery under small gapsAppendix[E](https://arxiv.org/html/2608.11661#A5)Definition[3](https://arxiv.org/html/2608.11661#Thmdefinition3)α\\alpha\-approximate separability as a quantitative relaxation of NMF conditionsAppendix[E](https://arxiv.org/html/2608.11661#A5)Theorem[12](https://arxiv.org/html/2608.11661#Thmtheorem12)Approximately separable targets degrade identifiability gracefullyAppendix[E](https://arxiv.org/html/2608.11661#A5)Proposition[5](https://arxiv.org/html/2608.11661#Thmproposition5)Identification and estimation errors combine into one ledgerAppendix[E](https://arxiv.org/html/2608.11661#A5)Proposition[6](https://arxiv.org/html/2608.11661#Thmproposition6)Mode error as the bridge between risk and identified basisAppendix[E](https://arxiv.org/html/2608.11661#A5)Similar Articles
Information-Theoretic Decomposition for Multimodal Interaction Learning
This paper presents an information-theoretic analysis of multimodal learning, revealing the need to capture sample-specific interactions, and proposes DMIL, a paradigm that explicitly models and learns from these interactions via variational decomposition and fine-tuning, achieving superior performance.
UMER: Unifying Embedding and Ranking via Pair-Aware Discriminative Reasoning for Universal Multimodal Retrieval
UMER introduces a unified framework for multimodal retrieval that combines embedding and ranking via pair-aware discriminative reasoning, achieving state-of-the-art performance on the MMEB-V2 benchmark.
DualOptim+: Bridging Shared and Decoupled Optimizer States for Better Machine Unlearning in Large Language Models
Introduces DualOptim+, an optimization framework for LLM unlearning that uses shared base states and decoupled delta states to balance forgetting and retaining objectives, with a quantized variant for reduced memory.
The Implicit Bias of Depth: From Neural Collapse to Softmax Codes
This paper studies how depth alone induces an implicit low-rank bias in deep unconstrained feature models trained without regularization, shifting the optimal solution from neural collapse to softmax codes, and provides the first asymptotic and dynamic characterization of this bias under gradient descent with cross-entropy loss.
Transforming LLMs into Efficient Cross-Encoders via Knowledge Distillation for RAG Reranking
This paper presents a method to fine-tune LLaMA 3 8B as an efficient reranker for Retrieval-Augmented Generation using knowledge distillation and 4-bit quantization, achieving 14-21% gains in retrieval metrics over cross-encoder baselines with reduced inference cost.