Which Tasks Survive Self-Supervised Learning?

arXiv cs.LG Papers

Summary

This Texas A&M paper introduces "semantic recoverability" to explain which downstream tasks survive same-instance self-supervised learning, linking recoverability to class-distance-normalized variance, few-shot nearest-centroid transfer, and the spectral subspace spanned by cross-view-stable modes of the SSL objective.

arXiv:2609.38393v1 Announce Type: new Abstract: Same-instance self-supervised learning (SSL) learns representations by enforcing consistency across two views of the same underlying instance. This principle alone, however, does not determine which downstream tasks remain recoverable from the learned representation. We study this question through \emph{semantic recoverability}, defined as the amount of a task's posterior score captured by the represented function space. We show that, for centered and whitened representations, recoverability exactly determines directional class-distance-normalized variance (CDNV), controls few-shot nearest-centroid classification, and governs the strength of task-relevant semantic directions. The population linear probe and centroid axis coincide, and multiple well-recovered tasks approach a factorial centroid geometry. We then analyze a canonical two-view SSL objective and show that its population optimum spans the leading cross-view-stable modes of the associated two-view operator. This yields a closed-form spectral characterization of semantic recoverability: a downstream task is preserved to the extent that its posterior lies in the selected spectral subspace. We validate these predictions on synthetic and real datasets across several SSL methods, testing the predicted relationships among recoverability, directional geometry, spectral structure, and few-shot transfer. Together, these results give a task-level account of what information survives same-instance SSL and how the retained information appears in downstream geometry and transfer.
Original Article
View Cached Full Text

Cached at: 10/02/26, 09:49 AM

# Which Tasks Survive Self-Supervised Learning?
Source: [https://arxiv.org/html/2609.38393](https://arxiv.org/html/2609.38393)
Lucas BryantTracy ZhuTomer Galanti††thanks:Correspondence togalanti@tamu\.eduAffiliation:Department of Computer Science and EngineeringAffiliation:Texas A&M University

###### Abstract

Same\-instance self\-supervised learning \(SSL\) learns representations by enforcing consistency across two views of the same underlying instance\. This principle alone, however, does not determine which downstream tasks remain recoverable from the learned representation\. We study this question through*semantic recoverability*, defined as the amount of a task’s posterior score captured by the represented function space\. We show that, for centered and whitened representations, recoverability exactly determines directional class\-distance\-normalized variance \(CDNV\), controls few\-shot nearest\-centroid classification, and governs the strength of task\-relevant semantic directions\. The population linear probe and centroid axis coincide, and multiple well\-recovered tasks approach a factorial centroid geometry\. We then analyze a canonical two\-view SSL objective and show that its population optimum spans the leading cross\-view\-stable modes of the associated two\-view operator\. This yields a closed\-form spectral characterization of semantic recoverability: a downstream task is preserved to the extent that its posterior lies in the selected spectral subspace\. We validate these predictions on synthetic and real datasets across several SSL methods, testing the predicted relationships among recoverability, directional geometry, spectral structure, and few\-shot transfer\. Together, these results give a task\-level account of what information survives same\-instance SSL and how the retained information appears in downstream geometry and transfer\.

## 1Introduction

Self\-supervised learning \(SSL\)\([Balestriero et al\., 2023](https://arxiv.org/html/2609.38393#bib.bib42)\)learns representations without labels and performs well across vision, language, speech, and multimodal learning\([Chen et al\., 2020](https://arxiv.org/html/2609.38393#bib.bib9);[He et al\., 2020](https://arxiv.org/html/2609.38393#bib.bib22);[Zbontar et al\., 2021](https://arxiv.org/html/2609.38393#bib.bib53);[He et al\., 2022](https://arxiv.org/html/2609.38393#bib.bib54);[Oquab et al\., 2024](https://arxiv.org/html/2609.38393#bib.bib52);[Gao et al\., 2021](https://arxiv.org/html/2609.38393#bib.bib27);[Reimers and Gurevych, 2019](https://arxiv.org/html/2609.38393#bib.bib25);[Schneider et al\., 2019](https://arxiv.org/html/2609.38393#bib.bib66);[Baevski et al\., 2020](https://arxiv.org/html/2609.38393#bib.bib64);[Hsu et al\., 2021](https://arxiv.org/html/2609.38393#bib.bib65);[Baevski et al\., 2022](https://arxiv.org/html/2609.38393#bib.bib67);[Radford et al\., 2021](https://arxiv.org/html/2609.38393#bib.bib30);[Jia et al\., 2021](https://arxiv.org/html/2609.38393#bib.bib28);[Zhai et al\., 2023](https://arxiv.org/html/2609.38393#bib.bib68);[Tschannen et al\., 2025](https://arxiv.org/html/2609.38393#bib.bib29)\)\. A common principle is simple: two views of the same instance should receive compatible representations\. Yet two views can share many kinds of information: semantic content, style, background, nuisance structure, or accidental regularities\. Agreement alone therefore does not reveal which downstream distinctions survive\.

This leads to a basic task\-level question: given a representation learned only from paired views, which semantic tasks can still be recovered from it? For example, a representation learned from color\-augmented face images might preserve whether a person is smiling while losing information about hair color\. Existing theory provides two pieces of the answer\. Operator and spectral analyses characterize which cross\-view\-stable directions are preferred by idealized SSL objectives\([Zhai et al\., 2025](https://arxiv.org/html/2609.38393#bib.bib6);[Zhai, 2025](https://arxiv.org/html/2609.38393#bib.bib7);[Ermolov et al\., 2021](https://arxiv.org/html/2609.38393#bib.bib75);[Weng et al\., 2022](https://arxiv.org/html/2609.38393#bib.bib74)\), while geometric analyses show that SSL representations can develop strong label\-aligned directional structure without global class collapse\([Luthra et al\., 2025](https://arxiv.org/html/2609.38393#bib.bib63);[Wang and Isola, 2020](https://arxiv.org/html/2609.38393#bib.bib34);[Garrido et al\., 2023](https://arxiv.org/html/2609.38393#bib.bib15);[Balestriero and LeCun, 2022](https://arxiv.org/html/2609.38393#bib.bib58);[Qiu et al\., 2023](https://arxiv.org/html/2609.38393#bib.bib8)\)\. We therefore ask:

Which downstream tasks survive same\-instance SSL, and what geometry and few\-shot behavior follow from their recoverability?

Our starting point is that this question is task\-dependent\. A representation can preserve one semantic partition while discarding another\. For a fixed downstream task, the relevant object is the part of its posterior score that remains visible in the represented function space, as illustrated in Fig\.[2](https://arxiv.org/html/2609.38393#S4.F2)\. We call this quantity*captured posterior energy*,B⁡\(F\)B\(F\), and use it to measure semantic recoverability independently of how the representation was learned\.

Our theory addresses two questions: why semantic recoverability matters, and when SSL makes a task recoverable\. Our contributions are threefold\.\(i\)We formulate task survival through captured posterior energy, separating recoverability from any particular SSL objective\.\(ii\)For centered and whitened representations, we derive an exact relation between recoverability and directional class\-distance\-normalized variance \(CDNV\)\([Luthra et al\., 2025](https://arxiv.org/html/2609.38393#bib.bib63)\), establish probe–centroid alignment, obtain a direct finite\-sample guarantee for few\-shot nearest\-centroid classification, and characterize the joint centroid geometry of multiple well\-recovered tasks\.\(iii\)For a canonical two\-view population SSL objective, we show that the learned subspace is spanned by the leading cross\-view\-stable modes of the associated two\-view operator\. Consequently, semantic recoverability is exactly the spectral overlap between the task posterior and the selected subspace, yielding a criterion for which downstream tasks survive SSL\. We further extend this subspace perspective to covariance\-regularized and prediction\-based SSL surrogates in the appendix\.

### 1\.1Related Work

Our work connects two lines of SSL theory that are usually studied separately: task\-relevant representation geometry and the population subspaces selected by two\-view objectives\.

Directional geometry and downstream transfer\.SSL representations can support strong downstream classification and retrieval, while remaining globally anisotropic and retaining substantial variance in nuisance directions\([Chen et al\., 2020](https://arxiv.org/html/2609.38393#bib.bib9);[Caron et al\., 2021](https://arxiv.org/html/2609.38393#bib.bib47);[Ben\-Shaul et al\., 2023](https://arxiv.org/html/2609.38393#bib.bib39);[Weng et al\., 2025](https://arxiv.org/html/2609.38393#bib.bib49);[Oquab et al\., 2024](https://arxiv.org/html/2609.38393#bib.bib52);[Wang and Isola, 2020](https://arxiv.org/html/2609.38393#bib.bib34);[Wang and Liu, 2021](https://arxiv.org/html/2609.38393#bib.bib33);[Chen et al\., 2021](https://arxiv.org/html/2609.38393#bib.bib37);[HaoChen et al\., 2021](https://arxiv.org/html/2609.38393#bib.bib18);[Saunshi et al\., 2019](https://arxiv.org/html/2609.38393#bib.bib26)\)\. This motivates task\-aware geometric measures rather than global collapse\. In particular, directional CDNV isolates within\-class variation along the semantic decision direction and has been shown to track linear\-probe and few\-shot transfer across SSL methods\([Luthra et al\., 2025](https://arxiv.org/html/2609.38393#bib.bib63)\)\. Related work on neural collapse and few\-shot transfer connects class geometry to generalization\([Papyan et al\., 2020](https://arxiv.org/html/2609.38393#bib.bib43);[Han et al\., 2022](https://arxiv.org/html/2609.38393#bib.bib44);[Galanti et al\., 2022c](https://arxiv.org/html/2609.38393#bib.bib48);[Galanti et al\., 2022b](https://arxiv.org/html/2609.38393#bib.bib46);[Galanti et al\., 2022a](https://arxiv.org/html/2609.38393#bib.bib45);[Wang et al\., 2023](https://arxiv.org/html/2609.38393#bib.bib41);[Goldblum et al\., 2020](https://arxiv.org/html/2609.38393#bib.bib40)\)\. Our goal is complementary: we ask what property of a label\-free learned subspace predicts whether this favorable task geometry will appear\.

Spectral and operator views of SSL\.Many SSL methods combine cross\-view agreement with mechanisms that prevent representational collapse\([Wang and Isola, 2020](https://arxiv.org/html/2609.38393#bib.bib34);[Zbontar et al\., 2021](https://arxiv.org/html/2609.38393#bib.bib53);[Bardes et al\., 2022](https://arxiv.org/html/2609.38393#bib.bib5);[Garrido et al\., 2023](https://arxiv.org/html/2609.38393#bib.bib15);[Balestriero and LeCun, 2022](https://arxiv.org/html/2609.38393#bib.bib58);[Ermolov et al\., 2021](https://arxiv.org/html/2609.38393#bib.bib75)\)\. At the population level, whitening\-style formulations connect naturally to CCA, maximal correlation, and conditional expectation operators\([Hotelling, 1936](https://arxiv.org/html/2609.38393#bib.bib77);[Hannan, 1961](https://arxiv.org/html/2609.38393#bib.bib79);[Breiman and Friedman, 1985](https://arxiv.org/html/2609.38393#bib.bib73);[Lai and Fyfe, 2000](https://arxiv.org/html/2609.38393#bib.bib72);[Hardoon et al\., 2004](https://arxiv.org/html/2609.38393#bib.bib71);[Fukumizu et al\., 2007](https://arxiv.org/html/2609.38393#bib.bib78);[Andrew et al\., 2013](https://arxiv.org/html/2609.38393#bib.bib80);[Michaeli et al\., 2016](https://arxiv.org/html/2609.38393#bib.bib70);[Asoodeh et al\., 2015](https://arxiv.org/html/2609.38393#bib.bib76)\)\. Recent analyses characterize the subspace selected by such objectives through the leading spectrum of an induced two\-view operator\([Zhai et al\., 2025](https://arxiv.org/html/2609.38393#bib.bib6);[Zhai, 2025](https://arxiv.org/html/2609.38393#bib.bib7)\)\. We use this characterization for computing which task posteriors are recoverable from the selected subspace\.

Relation to broader SSL theory\.A broader literature studies why SSL captures useful structure through mutual\-information views\([Bachman et al\., 2019](https://arxiv.org/html/2609.38393#bib.bib23);[McAllester and Stratos, 2020](https://arxiv.org/html/2609.38393#bib.bib55);[Tschannen et al\., 2020](https://arxiv.org/html/2609.38393#bib.bib56)\), alignment and uniformity\([Wang and Isola, 2020](https://arxiv.org/html/2609.38393#bib.bib34);[Wang and Liu, 2021](https://arxiv.org/html/2609.38393#bib.bib33);[Chen et al\., 2021](https://arxiv.org/html/2609.38393#bib.bib37)\), latent\-variable recovery and augmentation structure\([Saunshi et al\., 2019](https://arxiv.org/html/2609.38393#bib.bib26);[Tosh et al\., 2021](https://arxiv.org/html/2609.38393#bib.bib35);[Zimmermann et al\., 2021](https://arxiv.org/html/2609.38393#bib.bib32);[Ash et al\., 2022](https://arxiv.org/html/2609.38393#bib.bib38);[Nozawa and Sato, 2021](https://arxiv.org/html/2609.38393#bib.bib19);[HaoChen et al\., 2021](https://arxiv.org/html/2609.38393#bib.bib18);[Shen et al\., 2022](https://arxiv.org/html/2609.38393#bib.bib17);[Awasthi et al\., 2022](https://arxiv.org/html/2609.38393#bib.bib11);[Saunshi et al\., 2022](https://arxiv.org/html/2609.38393#bib.bib36)\), and optimization, projection heads, and sample complexity\([HaoChen and Ma, 2023](https://arxiv.org/html/2609.38393#bib.bib16);[Parulekar et al\., 2023](https://arxiv.org/html/2609.38393#bib.bib31);[Wen and Li, 2021](https://arxiv.org/html/2609.38393#bib.bib13);[Tian, 2023](https://arxiv.org/html/2609.38393#bib.bib12);[Tian et al\., 2020](https://arxiv.org/html/2609.38393#bib.bib24);[Gupta et al\., 2022](https://arxiv.org/html/2609.38393#bib.bib60);[Gui et al\., 2023](https://arxiv.org/html/2609.38393#bib.bib57);[Alon et al\., 2024](https://arxiv.org/html/2609.38393#bib.bib59)\)\. Our question is narrower and task\-conditional: given the information retained by a learned representation, which downstream posterior scores survive, and what geometric and few\-shot consequences follow?

## 2Problem Setup

We consider a latent instance that generates an observed sample and two conditionally independent views, together with a semantic task defined at the instance level\. A representation is learned from paired views without access to the task labels and is then evaluated by how well the task can be recovered from a small labeled support set\.

### 2\.1Latent Instance Model and Two\-View Generation

LetCCdenote the*latent instance*\(for example, the underlying object, scene, or document\) taking values in a measurable space𝒞\\mathcal\{C\}, with distributionPCP\_\{C\}\. The semantic label is determined at the instance level:Y=y⁡\(C\)∈\{±1\}Y~=~y\(C\)\\in\\\{\\pm 1\\\},ℙ⁡\(Y=\+1\)=ℙ⁡\(Y=−1\)=12\\mathbb\{P\}\(Y=\+1\)~=~\\mathbb\{P\}\(Y=\-1\)~=~\\tfrac\{1\}\{2\}\. ThusCCmay contain much more information thanYY: many latent instances can share the same label\.

We letX\(0\)X^\{\(0\)\}denote the sample associated withCC, drawn asX\(0\)∼P0\(⋅∣C\)X^\{\(0\)\}\\sim P\_\{0\}\(\\cdot\\mid C\)\. A positive pair is formed by drawing two independent views,X\(1\),X\(2\)∼i\.i\.d\.P\(⋅∣X\(0\)\)X^\{\(1\)\},X^\{\(2\)\}\\stackrel\{\{\\scriptstyle\\mathrm\{i\.i\.d\.\}\}\}\{\{\\sim\}\}P\(\\cdot\\mid X^\{\(0\)\}\)\. In vision, for example,CCmay encode image content, while each view is produced by a random augmentation such as cropping, color jitter, or blur\. LetPXP\_\{X\}denote the marginal law of a single view\.

### 2\.2Downstream Few\-Shot Learning

LetF:𝒳→ℝrF:\\mathcal\{X\}\\to\\mathbb\{R\}^\{r\}be a representation\. We evaluate its downstream quality through the expectedmm\-shot error of a classifier built on top ofFF\. Given a support set withmmlabeled examples per class, with all support examples sampled independently,Sm=\{\(Xs,i,s\):s∈\{±1\}S\_\{m\}=\\\{\(X\_\{s,i\},s\):s\\in\\\{\\pm 1\\\},Xs,i∼P\(⋅∣Y=s\)X\_\{s,i\}\\sim P\(\\cdot\\mid Y=s\), \(i∈\[m\]i\\in\[m\]\)\}\\\}, letg^SmA\\hat\{g\}^\{\\mathrm\{A\}\}\_\{S\_\{m\}\}denote the classifier returned by a learning ruleA\\mathrm\{A\}from the embedded support set\. We defineerrmA​\(F\):=𝔼Sm​\[ℙ\(X,Y\)​\(g^SmA​\(F⁡\(X\)\)≠Y\)\]\\textnormal\{err\}^\{\\mathrm\{A\}\}\_\{m\}\(F\):=\\mathbb\{E\}\_\{S\_\{m\}\}\\\!\\left\[\\mathbb\{P\}\_\{\(X,Y\)\}\\\!\\big\(\\hat\{g\}^\{\\mathrm\{A\}\}\_\{S\_\{m\}\}\(F\(X\)\)\\neq Y\\big\)\\right\], where\(X,Y\)\(X,Y\)is drawn from the single\-view distribution induced by the latent\-instance model\.

We focus on nearest\-class\-centroid \(NCC\) classification:gSmNCC​\(z\):=arg​mins∈\{±1\}⁡‖z−μ^s‖2g^\{\\mathrm\{NCC\}\}\_\{S\_\{m\}\}\(z\):=\\argmin\_\{s\\in\\\{\\pm 1\\\}\}\\\|z\-\\widehat\{\\mu\}\_\{s\}\\\|\_\{2\}, whereμ^s:=1m​∑i=1mF⁡\(Xs,i\)\\widehat\{\\mu\}\_\{s\}:=\\frac\{1\}\{m\}\\sum\_\{i=1\}^\{m\}F\(X\_\{s,i\}\)\. This gives an operational notion of task survival: a useful representation should make the task learnable from only a small labeled support set\. The next section identifies a population quantity that controls when this happens\.

## 3Semantic Recoverability and Its Consequences

Task stableacross viewsRecoverabilityBr=∑j≤rβj2B\_\{r\}=\\sum\_\{j\\leq r\}\\beta\_\{j\}^\{2\}\(i\)low directional CDNVV~F=1−Br2​Br\\widetilde\{V\}\_\{F\}=\\frac\{1\-B\_\{r\}\}\{2B\_\{r\}\}\(ii\)few\-shot NCCerrmNCC≲1−Br\+rm\\mathrm\{err\}^\{\\mathrm\{NCC\}\}\_\{m\}\\lesssim 1\-B\_\{r\}\+\\frac\{r\}\{m\}\(iii\)multitask geometrymY≈\(Br\(t\)​Yt\)t≤km\_\{Y\}\\approx\\bigl\(\\sqrt\{B\_\{r\}^\{\(t\)\}\}\\,Y\_\{t\}\\bigr\)\_\{t\\leq k\}top\-rrmodesofTTFigure 1:Conceptual flow of the theory\. Same\-instance SSL retains the most cross\-view\-stable modes of the two\-view operatorTT\(Prop\.[4\.1](https://arxiv.org/html/2609.38393#S4.Thmtheorem1)\), so a task survives to the extent that its posterior overlaps that subspace \(Cor\.[4\.2](https://arxiv.org/html/2609.38393#S4.Thmtheorem2)\)\. Under centered whitening,BrB\_\{r\}fixes the directional collapse law \(Prop\.[3\.1](https://arxiv.org/html/2609.38393#S3.Thmtheorem1)\) and few\-shot nearest\-centroid guarantee \(Thm\.[3\.4](https://arxiv.org/html/2609.38393#S3.Thmtheorem4)\); the task\-specificBr\(t\)B\_\{r\}^\{\(t\)\}determine the multitask centroid geometry \(Thm\.[3\.5](https://arxiv.org/html/2609.38393#S3.Thmtheorem5)\)\.### 3\.1Captured Posterior Energy

For the balanced task, define the posterior scoreη⁡\(x\):=𝔼⁡\[Y∣X=x\]\\eta\(x\):=\\mathbb\{E\}\[Y\\mid X=x\]\. The posterior contains the information aboutYYavailable from a single view\. For a representationF=\(F1,…,Fr\)F=\(F\_\{1\},\\dots,F\_\{r\}\), let𝒮F:=span⁡\{F1,…,Fr\}⊂L2​\(PX\)\\mathcal\{S\}\_\{F\}:=\\mathrm\{span\}\\\{F\_\{1\},\\dots,F\_\{r\}\\\}\\subset L^\{2\}\(P\_\{X\}\)denote its represented function space\. We define the*captured posterior energy*B⁡\(F\):=‖Π𝒮F​η‖L2​\(PX\)2B\(F\):=\\\|\\Pi\_\{\\mathcal\{S\}\_\{F\}\}\\eta\\\|\_\{L^\{2\}\(P\_\{X\}\)\}^\{2\}, whereΠ𝒮F\\Pi\_\{\\mathcal\{S\}\_\{F\}\}is the orthogonal projection onto𝒮F\\mathcal\{S\}\_\{F\}\.

This definition does not require SSL or whitening\. It asks directly how much of the Bayes\-relevant posterior score remains in the function space represented byFF\. The dependence on the downstream task is intentional: the same representation may preserve one semantic partition and discard another\. If one wants a literal fraction of the single\-view task signal that is retained, it isB⁡\(F\)/‖η‖L2​\(PX\)2B\(F\)/\\\|\\eta\\\|\_\{L^\{2\}\(P\_\{X\}\)\}^\{2\}\. In particular, largeB⁡\(F\)B\(F\)means that most of the posterior is expressible through the representation, while smallB⁡\(F\)B\(F\)means that the task is largely absent from the represented subspace\.

The quantity is invariant to invertible changes of coordinates in feature space because such changes leave𝒮F\\mathcal\{S\}\_\{F\}unchanged\. Whitening enters below for another reason: it converts this intrinsic subspace quantity into exact Euclidean statements about feature geometry and few\-shot classification\.

### 3\.2Why Recoverability Matters Under Whitening

Assume for the remainder of this section thatFFis centered and whitened, so that𝔼⁡\[F⁡\(X\)\]=0\\mathbb\{E\}\[F\(X\)\]=0and𝔼⁡\[F⁡\(X\)​F​\(X\)⊤\]=Ir\\mathbb\{E\}\[F\(X\)F\(X\)^\{\\top\}\]=I\_\{r\}\. The coordinate functions ofFFare then orthonormal inL2​\(PX\)L^\{2\}\(P\_\{X\}\), and the intrinsic recoverability quantity has the simple feature\-space formB⁡\(F\)=‖𝔼⁡\[Y​F​\(X\)\]‖22B\(F\)=\\\|\\mathbb\{E\}\[YF\(X\)\]\\\|\_\{2\}^\{2\}\.

Letμs:=𝔼⁡\[F⁡\(X\)∣Y=s\]\\mu\_\{s\}:=\\mathbb\{E\}\[F\(X\)\\mid Y=s\],Σs:=Cov⁡\(F⁡\(X\)∣Y=s\)\\Sigma\_\{s\}:=\\mathrm\{Cov\}\(F\(X\)\\mid Y=s\), andΔ:=μ\+−μ−\\Delta:=\\mu\_\{\+\}\-\\mu\_\{\-\}\. WhenΔ≠0\\Delta\\neq 0, letu:=Δ/‖Δ‖2u:=\\Delta/\\\|\\Delta\\\|\_\{2\}\. The CDNV \(VFV\_\{F\}\) and directional CDNV \(V~F\\tilde\{V\}\_\{F\}\) are

VF:=\(Tr⁡\(Σ\+\)\+Tr⁡\(Σ−\)\)/‖Δ‖22,V~F:=\(u⊤​Σ\+​u\+u⊤​Σ−​u\)/‖Δ‖22\.V\_\{F\}:=\(\\mathrm\{Tr\}\(\\Sigma\_\{\+\}\)\+\\mathrm\{Tr\}\(\\Sigma\_\{\-\}\)\)/\\\|\\Delta\\\|\_\{2\}^\{2\},\\qquad\\tilde\{V\}\_\{F\}:=\(u^\{\\top\}\\Sigma\_\{\+\}u\+u^\{\\top\}\\Sigma\_\{\-\}u\)/\\\|\\Delta\\\|\_\{2\}^\{2\}\.\(1\)WhenΔ=0\\Delta=0, we setVF=V~F=\+∞V\_\{F\}=\\tilde\{V\}\_\{F\}=\+\\infty\. Under whitening, total within\-class variation can remain large even when variation along the semantic direction is small\.

#### 3\.2\.1Exact directional\-collapse law

###### Proposition 3\.1\.

LetF:𝒳→ℝrF:\\mathcal\{X\}\\to\\mathbb\{R\}^\{r\}be centered and whitened, and letY∈\{±1\}Y\\in\\\{\\pm 1\\\}be balanced\. ThenVF=\(r−B⁡\(F\)\)/\(2​B​\(F\)\)V\_\{F\}=\(r\-B\(F\)\)/\(2B\(F\)\)andV~F=\(1−B⁡\(F\)\)/\(2​B​\(F\)\)\\tilde\{V\}\_\{F\}=\(1\-B\(F\)\)/\(2B\(F\)\), with both quantities equal to\+∞\+\\inftywhenB⁡\(F\)=0B\(F\)=0\.

Prop\.[3\.1](https://arxiv.org/html/2609.38393#S3.Thmtheorem1)gives an exact population law: directional CDNV vanishes as recoverability approaches one and diverges as recoverability vanishes\. Thus the representation need not collapse globally for the task to become geometrically simple; it only needs to collapse along the semantic direction that survives in the represented subspace\.

#### 3\.2\.2Probe–centroid alignment

A second consequence of whitening is that the population squared\-loss probe and the centroid decision axis coincide\.

###### Corollary 3\.3\(MSE probe and decision\-axis alignment\)\.

LetF:𝒳→ℝrF:\\mathcal\{X\}\\to\\mathbb\{R\}^\{r\}be centered and whitened, and letY∈\{±1\}Y\\in\\\{\\pm 1\\\}be balanced\. Defineμs:=𝔼⁡\[F⁡\(X\)∣Y=s\]\\mu\_\{s\}:=\\mathbb\{E\}\[F\(X\)\\mid Y=s\]fors∈\{±1\}s\\in\\\{\\pm 1\\\}, and letΔ:=μ\+−μ−\\Delta:=\\mu\_\{\+\}\-\\mu\_\{\-\}\. Then the unique population MSE probe,\(w⋆,b⋆\)∈arg​minw∈ℝr,b∈ℝ⁡𝔼​\[\(Y−\(w⊤​F​\(X\)\+b\)\)2\]\(w^\{\\star\},b^\{\\star\}\)\\in\\argmin\_\{w\\in\\mathbb\{R\}^\{r\},b\\in\\mathbb\{R\}\}\\mathbb\{E\}\[\(Y\-\(w^\{\\top\}F\(X\)\+b\)\)^\{2\}\], satisfiesb⋆=0b^\{\\star\}=0andw⋆=𝔼⁡\[Y​F​\(X\)\]=12​Δw^\{\\star\}=\\mathbb\{E\}\[YF\(X\)\]=\\tfrac\{1\}\{2\}\\Delta\. In particular,μ−=−μ\+\\mu\_\{\-\}=\-\\mu\_\{\+\}, andΔ=2​μ\+=2​w⋆\\Delta=2\\mu\_\{\+\}=2w^\{\\star\}\.

Cor\.[3\.3](https://arxiv.org/html/2609.38393#S3.Thmtheorem3)shows that linear probing and NCC use the same separating direction\. Moreover,‖w⋆‖22=B⁡\(F\)\\\|w^\{\\star\}\\\|\_\{2\}^\{2\}=B\(F\), so the strength of the population probe is itself a measurement of semantic recoverability\.

#### 3\.2\.3Few\-shot transfer guarantees

The geometric identities above translate recoverability into an operational downstream guarantee\. In the whitened setting, the same scalar that controls directional geometry also controls the expected error of anmm\-shot nearest\-centroid classifier\.

###### Theorem 3\.4\(Few\-shot NCC bound via captured posterior energy\)\.

LetF:𝒳→ℝrF:\\mathcal\{X\}\\to\\mathbb\{R\}^\{r\}be centered and whitened, and letY∈\{±1\}Y\\in\\\{\\pm 1\\\}be balanced\. Then, for everym≥1m\\geq 1,

errmNCC​\(F\)≤1−B⁡\(F\)\+r−B⁡\(F\)m\+1−B⁡\(F\)1−B⁡\(F\)\+2​m​B​\(F\)=2​V~F1\+2​V~F\+\(r−1\)\+2​r​V~Fm⁡\(1\+2​V~F\)\+V~FV~F\+m,\\mathrm\{err\}^\{\\mathrm\{NCC\}\}\_\{m\}\(F\)\\leq 1\-B\(F\)\+\\frac\{r\-B\(F\)\}\{m\}\+\\frac\{1\-B\(F\)\}\{1\-B\(F\)\+2mB\(F\)\}=\\frac\{2\\tilde\{V\}\_\{F\}\}\{1\+2\\tilde\{V\}\_\{F\}\}\+\\frac\{\(r\-1\)\+2r\\tilde\{V\}\_\{F\}\}\{m\(1\+2\\tilde\{V\}\_\{F\}\)\}\+\\frac\{\\tilde\{V\}\_\{F\}\}\{\\tilde\{V\}\_\{F\}\+m\},where the second expression is understood by continuity whenB⁡\(F\)=0B\(F\)=0\.

The bound separates two effects\. The leading term reflects task information not captured by the representation, while the remaining terms arise from estimating class centroids from finitely many examples\. Whitening is what makes this a one\-scalar Euclidean statement: without it, coordinate rescaling could change‖𝔼⁡\[Y​F\]‖22\\\|\\mathbb\{E\}\[YF\]\\\|\_\{2\}^\{2\}and Euclidean NCC would no longer coincide with the natural population geometry\. The coordinate\-invariant analogue is obtained by rewhitening, or equivalently by using a Mahalanobis metric in the original feature space\.

#### 3\.2\.4Multitask semantic geometry

Recoverability also determines how several semantic tasks are organized jointly\. We use a superscript to index tasks, reserving the subscriptrrfor representation rank\. Considerkkbalanced binary tasksY1,…,Yk∈\{±1\}Y\_\{1\},\\dots,Y\_\{k\}\\in\\\{\\pm 1\\\}, and writeY:=\(Y1,…,Yk\)Y:=\(Y\_\{1\},\\dots,Y\_\{k\}\)\. For each tasktt, letηt​\(x\)=𝔼⁡\[Yt∣X=x\]\\eta\_\{t\}\(x\)=\\mathbb\{E\}\[Y\_\{t\}\\mid X=x\],B\(t\)​\(F\)=‖Π𝒮F​ηt‖L2​\(PX\)2B^\{\(t\)\}\(F\)=\\\|\\Pi\_\{\\mathcal\{S\}\_\{F\}\}\\eta\_\{t\}\\\|\_\{L^\{2\}\(P\_\{X\}\)\}^\{2\}, andεt=𝔼⁡\[Var⁡\(Yt∣X\)\]=1−‖ηt‖L2​\(PX\)2\\varepsilon\_\{t\}=\\mathbb\{E\}\[\\mathrm\{Var\}\(Y\_\{t\}\\mid X\)\]=1\-\\\|\\eta\_\{t\}\\\|\_\{L^\{2\}\(P\_\{X\}\)\}^\{2\}\. Letwt=𝔼⁡\[Yt​F​\(X\)\]w\_\{t\}=\\mathbb\{E\}\[Y\_\{t\}F\(X\)\]\. AssumingB\(t\)​\(F\)\>0B^\{\(t\)\}\(F\)\>0, defineut=wt/‖wt‖2u\_\{t\}=w\_\{t\}/\\\|w\_\{t\}\\\|\_\{2\}andZt​\(X\)=⟨ut,F⁡\(X\)⟩Z\_\{t\}\(X\)=\\langle u\_\{t\},F\(X\)\\rangle\.

###### Theorem 3\.5\(Near\-orthogonality and centroid geometry\)\.

LetF:𝒳→ℝrF:\\mathcal\{X\}\\to\\mathbb\{R\}^\{r\}be centered and whitened\. Assume eachYt∈\{±1\}Y\_\{t\}\\in\\\{\\pm 1\\\}is balanced andB\(t\)​\(F\)\>0B^\{\(t\)\}\(F\)\>0for alltt; arbitrary dependence amongY1,…,YkY\_\{1\},\\dots,Y\_\{k\}is allowed\. Fori≠ji\\neq j, defineρi​j:=𝔼⁡\[Yi​Yj\]\\rho\_\{ij\}:=\\mathbb\{E\}\[Y\_\{i\}Y\_\{j\}\]\. Then, for anyi≠ji\\neq j,

\|ui⊤​uj\|≤\(\|ρi​j\|\+εi​εj\+\(‖ηi‖L22−B\(i\)​\(F\)\)​\(‖ηj‖L22−B\(j\)​\(F\)\)\)/B\(i\)​\(F\)​B\(j\)​\(F\)\.\|u\_\{i\}^\{\\top\}u\_\{j\}\|~\\leq~\\left\(\|\\rho\_\{ij\}\|\+\\sqrt\{\\varepsilon\_\{i\}\\varepsilon\_\{j\}\}\+\\sqrt\{\\big\(\\\|\\eta\_\{i\}\\\|\_\{L^\{2\}\}^\{2\}\-B^\{\(i\)\}\(F\)\\big\)\\big\(\\\|\\eta\_\{j\}\\\|\_\{L^\{2\}\}^\{2\}\-B^\{\(j\)\}\(F\)\\big\)\}\\right\)/\\sqrt\{B^\{\(i\)\}\(F\)B^\{\(j\)\}\(F\)\}\.Moreover, lettingmY:=𝔼⁡\[Z⁡\(X\)∣Y\]∈ℝkm\_\{Y\}:=\\mathbb\{E\}\[Z\(X\)\\mid Y\]\\in\\mathbb\{R\}^\{k\},

𝔼⁡\[‖mY−\(B\(1\)​\(F\)​Y1,…,B\(k\)​\(F\)​Yk\)‖22\]≤∑t=1k\(1−B\(t\)​\(F\)\)\.\\mathbb\{E\}\\Big\[\\big\\\|m\_\{Y\}\-\\big\(\\sqrt\{B^\{\(1\)\}\(F\)\}Y\_\{1\},\\dots,\\sqrt\{B^\{\(k\)\}\(F\)\}Y\_\{k\}\\big\)\\big\\\|\_\{2\}^\{2\}\\Big\]~\\leq~\\sum\_\{t=1\}^\{k\}\\big\(1\-B^\{\(t\)\}\(F\)\\big\)\.

Thm\.[3\.5](https://arxiv.org/html/2609.38393#S3.Thmtheorem5)shows that well\-recovered tasks induce strong semantic axes that become nearly orthogonal when task correlations and residual posterior uncertainty are small\. Their joint centroids approach\(B\(t\)​\(F\)​Yt\)t≤k\\big\(\\sqrt\{B^\{\(t\)\}\(F\)\}Y\_\{t\}\\big\)\_\{t\\leq k\}, yielding an approximately axis\-aligned hyperrectangle whose side lengths are set by the task\-specific capture values\. Thus recoverability controls the strength of each semantic direction, while cross\-task dependence determines how closely the overall geometry approaches a factorial product\.

The results so far deliberately make no claim about howFFwas learned\. They say that*if*a representation captures a task, thenB⁡\(F\)B\(F\)determines its directional geometry, its few\-shot behavior, and its role in multitask centroid structure\. We now turn to the complementary question:why should a same\-instance SSL objective produce a large value ofB⁡\(F\)B\(F\)for some tasks and not others?

## 4Recoverability Under Same\-Instance SSL

The preceding results show that the recoverability quantityB⁡\(F\)B\(F\)characterizes the task\-relevant geometry of a representation and controls its few\-shot behavior\. The next question is whether, for a representation learned by same\-instance SSL, this quantity can itself be identified explicitly from the SSL objective\. We show that this is indeed the case: under the whitening\-based population objective, the learned representation spans the leading eigenspace of the two\-view operator, andB⁡\(F\)B\(F\)becomes the spectral overlap between the downstream posterior and these view\-stable modes\.

### 4\.1A Canonical Two\-View Population Objective

Many SSL methods\([Grill et al\., 2020](https://arxiv.org/html/2609.38393#bib.bib10);[Zbontar et al\., 2021](https://arxiv.org/html/2609.38393#bib.bib53);[Bardes et al\., 2022](https://arxiv.org/html/2609.38393#bib.bib5);[Chen et al\., 2020](https://arxiv.org/html/2609.38393#bib.bib9);[Caron et al\., 2021](https://arxiv.org/html/2609.38393#bib.bib47);[Ermolov et al\., 2021](https://arxiv.org/html/2609.38393#bib.bib75);[Tao et al\., 2022](https://arxiv.org/html/2609.38393#bib.bib51)\)combine instance\-level agreement with an anti\-collapse mechanism that promotes global spread, decorrelation, or isotropy\([Wang and Isola, 2020](https://arxiv.org/html/2609.38393#bib.bib34);[Garrido et al\., 2023](https://arxiv.org/html/2609.38393#bib.bib15);[Balestriero and LeCun, 2022](https://arxiv.org/html/2609.38393#bib.bib58)\)\. To obtain a clean population characterization of the represented subspace, we study the canonical whitening\-based surrogate \(W\-MSE\)\([Ermolov et al\., 2021](https://arxiv.org/html/2609.38393#bib.bib75);[Weng et al\., 2022](https://arxiv.org/html/2609.38393#bib.bib74)\)

maxF:𝒳→ℝr𝔼\[⟨F\(X\(1\)\),F\(X\(2\)\)⟩\]subject to𝔼\[F\(X\)\]=0,𝔼\[F\(X\)F\(X\)⊤\]=Ir\.\\max\_\{F:\\mathcal\{X\}\\to\\mathbb\{R\}^\{r\}\}\\mathbb\{E\}\\\!\\left\[\\langle F\(X^\{\(1\)\}\),F\(X^\{\(2\)\}\)\\rangle\\right\]\\quad\\text\{subject to\}\\quad\\mathbb\{E\}\[F\(X\)\]=0,\\qquad\\mathbb\{E\}\[F\(X\)F\(X\)^\{\\top\}\]=I\_\{r\}\.\(2\)The objective rewards agreement across two views of the same instance, while whitening rules out collapse and fixes the feature metric\. Its role in our theory is specific: it gives an analyzable rule for whichrr\-dimensional function space is selected from the two\-view distribution\.

Strict whitening is the cleanest setting for this identification and for the exact geometry above\. App\.[C](https://arxiv.org/html/2609.38393#A3)studies a covariance\-regularized surrogate related to Barlow Twins\([Zbontar et al\., 2021](https://arxiv.org/html/2609.38393#bib.bib53)\)and VICReg\([Bardes et al\., 2022](https://arxiv.org/html/2609.38393#bib.bib5)\), while App\.[D](https://arxiv.org/html/2609.38393#A4)gives a corresponding reduced\-rank regression view of prediction\-based objectives such as BYOL\([Grill et al\., 2020](https://arxiv.org/html/2609.38393#bib.bib10)\), SimSiam\([Tao et al\., 2022](https://arxiv.org/html/2609.38393#bib.bib51)\), and JEPA\([Assran et al\., 2023](https://arxiv.org/html/2609.38393#bib.bib50)\)\.

Figure 2:Geometric interpretation of semantic recoverability\.The shaded plane is the SSL\-selected subspace𝒰r=span⁡\{ψ1,ψ2\}\\mathcal\{U\}\_\{r\}=\\operatorname\{span\}\\\{\\psi\_\{1\},\\psi\_\{2\}\\\}of leading two\-view eigenfunctions\. The task posteriorη\\etadecomposes into its projectionΠr​η\\Pi\_\{r\}\\etaonto this subspace and an orthogonal residual\. From left to right, the posterior is fully captured, partially captured, or discarded\. The captured posterior energyBr=‖Πr​η‖L2​\(PX\)2B\_\{r\}=\\\|\\Pi\_\{r\}\\eta\\\|\_\{L^\{2\}\(P\_\{X\}\)\}^\{2\}measures the retained task\-relevant information\.
### 4\.2The Two\-View Operator and the Selected Subspace

Define the two\-view conditional expectation operatorT:L2​\(PX\)→L2​\(PX\)T:L^\{2\}\(P\_\{X\}\)\\to L^\{2\}\(P\_\{X\}\)by\(T​f\)​\(x\):=𝔼⁡\[f⁡\(X\(2\)\)∣X\(1\)=x\]\(Tf\)\(x\):=\\mathbb\{E\}\[f\(X^\{\(2\)\}\)\\mid X^\{\(1\)\}=x\]\. ThusTTmaps a function of one view to its conditional expectation from another view of the same instance\. In the conditionally i\.i\.d\. two\-view model, exchangeability gives⟨f,T​g⟩=⟨T​f,g⟩\\langle f,Tg\\rangle=\\langle Tf,g\\rangle, while conditional independence givenX\(0\)X^\{\(0\)\}gives⟨f,T​f⟩=𝔼⁡\[\(𝔼⁡\[f⁡\(X\)∣X\(0\)\]\)2\]≥0\\langle f,Tf\\rangle=\\mathbb\{E\}\[\(\\mathbb\{E\}\[f\(X\)\\mid X^\{\(0\)\}\]\)^\{2\}\]\\geq 0\. HenceTTis self\-adjoint and positive semidefinite, and an eigenvalue measures the cross\-view stability of the corresponding mode\.

###### Proposition 4\.1\.

LetL02​\(PX\):=\{f∈L2​\(PX\):𝔼⁡\[f⁡\(X\)\]=0\}L\_\{0\}^\{2\}\(P\_\{X\}\):=\\\{f\\in L^\{2\}\(P\_\{X\}\):\\mathbb\{E\}\[f\(X\)\]=0\\\}, and assume thatTTrestricted toL02​\(PX\)L\_\{0\}^\{2\}\(P\_\{X\}\)admits an orthonormal eigenbasis\(ψj\)j≥1\(\\psi\_\{j\}\)\_\{j\\geq 1\}, with eigenvaluesλ1≥λ2≥⋯≥0\\lambda\_\{1\}\\geq\\lambda\_\{2\}\\geq\\cdots\\geq 0\. Assumeλr\>0\\lambda\_\{r\}\>0andλr\>λr\+1\\lambda\_\{r\}\>\\lambda\_\{r\+1\}\. Thenψ1,…,ψr\\psi\_\{1\},\\dots,\\psi\_\{r\}are feasible for \([2](https://arxiv.org/html/2609.38393#S4.E2)\) and achieve the optimal value∑j=1rλj\\sum\_\{j=1\}^\{r\}\\lambda\_\{j\}\. Moreover, ifF:𝒳→ℝrF:\\mathcal\{X\}\\to\\mathbb\{R\}^\{r\}is any optimal solution of \([2](https://arxiv.org/html/2609.38393#S4.E2)\), then its coordinates span the samerr\-dimensional subspace asψ1,…,ψr\\psi\_\{1\},\\dots,\\psi\_\{r\}\.

Prop\.[4\.1](https://arxiv.org/html/2609.38393#S4.Thmtheorem1)says that the population objective keeps therrmost cross\-view\-stable modes\. The representation is determined only up to an orthogonal change of basis within this selected subspace\. This result identifies what SSL preserves without referring to any downstream labels; the task enters only when we ask how much of its posterior lies in that subspace\.

### 4\.3Closed\-Form Recoverability at the SSL Optimum

LetF⋆F^\{\\star\}be any population\-optimal solution of \([2](https://arxiv.org/html/2609.38393#S4.E2)\), and define𝒰r:=span⁡\{ψ1,…,ψr\}⊂L2​\(PX\)\\mathcal\{U\}\_\{r\}:=\\mathrm\{span\}\\\{\\psi\_\{1\},\\dots,\\psi\_\{r\}\\\}\\subset L^\{2\}\(P\_\{X\}\), withΠr\\Pi\_\{r\}denoting the orthogonal projection onto𝒰r\\mathcal\{U\}\_\{r\}\. For a balanced binary label with posterior scoreη⁡\(x\)=𝔼⁡\[Y∣X=x\]\\eta\(x\)=\\mathbb\{E\}\[Y\\mid X=x\], writeβj:=⟨η,ψj⟩L2​\(PX\)\\beta\_\{j\}:=\\langle\\eta,\\psi\_\{j\}\\rangle\_\{L^\{2\}\(P\_\{X\}\)\}\.

###### Corollary 4\.2\.

Under the assumptions of Prop\.[4\.1](https://arxiv.org/html/2609.38393#S4.Thmtheorem1),Πr​η=∑j=1rβj​ψj\\Pi\_\{r\}\\eta=\\sum\_\{j=1\}^\{r\}\\beta\_\{j\}\\psi\_\{j\}, andB⁡\(F⋆\)=‖Πr​η‖L2​\(PX\)2=∑j=1rβj2=:BrB\(F^\{\\star\}\)=\\\|\\Pi\_\{r\}\\eta\\\|\_\{L^\{2\}\(P\_\{X\}\)\}^\{2\}=\\sum\_\{j=1\}^\{r\}\\beta\_\{j\}^\{2\}=:B\_\{r\}\.

For multiple downstream tasks, we writeBr\(t\)B\_\{r\}^\{\(t\)\}for the rank\-rrcapture of tasktt; equivalently,Br\(t\)=B\(t\)​\(F⋆\)B\_\{r\}^\{\(t\)\}=B^\{\(t\)\}\(F^\{\\star\}\)for an optimal rank\-rrrepresentation\.

This is the closed\-form answer to when recoverability is large under the SSL model\. The task is preserved when its posterior score is concentrated on the modes that are most stable across the two views, and it is discarded when its posterior lies primarily in view\-variant directions outside the selected subspace\. The two\-view objective itself remains label\-free; labels enter only through the overlap used to evaluate a particular downstream task\.

Combining this spectral identity with the previous section immediately turnsBrB\_\{r\}into task\-level predictions\. A large spectral overlap implies low directional CDNV, strong few\-shot NCC performance, and pronounced semantic axes; across several tasks, the valuesBr\(t\)B\_\{r\}^\{\(t\)\}determine the side lengths of the approximate hyperrectangle\. In the recoverable regime where the augmentations preserve the task and the increasing view\-stable subspaces capture its posterior,BrB\_\{r\}approaches one as the representation dimension grows\. Thus the overall mechanism is: SSL selects view\-stable structure, a task survives according to its posterior overlap with that structure, and the amount that survives determines the downstream geometry and transfer behavior\.

## 5Experiments

\(a\) W\-MSE\-CelebA\(b\) VICReg\-CelebA\(c\) Barlow\-IM1K\(d\) I\-JEPA\-IM1KFigure 3:Captured posterior energy on CelebA\.MeanBr\(t\)B\_\{r\}^\{\(t\)\}over the 1, 5, and 10 most recoverable CelebA attributes and for all 40 attributes, shown across spectral ranksrr\. Shaded regions denote the corresponding interquartile ranges\. The dashed curve is the mean for a matched randomly initialized encoder, with dotted curves showing its interquartile range\.\(a\) W\-MSE\-CelebA\(b\) VICReg\-CelebA\(c\) Barlow\-IM1K\(d\) I\-JEPA\-IM1KFigure 4:Predicted versus observed directional CDNV on CelebA\.Each point is one CelebA attribute at a given rankrr\. The predictionV~^pred=\(1−B^A\(t\)\)/\(2​B^A\(t\)\)\\widehat\{\\tilde\{V\}\}\_\{\\mathrm\{pred\}\}=\(1\-\\widehat\{B\}\_\{A\}^\{\(t\)\}\)/\(2\\widehat\{B\}\_\{A\}^\{\(t\)\}\)is computed on split A\. Class centroids, the semantic direction, and observed directional CDNV are estimated independently on held\-out split B\. The dashed line isy=xy=x, and each panel contains all 40 attributes\.\(a\) W\-MSE\-CelebA\(b\) I\-JEPA\-IM1KFigure 5:Captured posterior energy controls few\-shot NCC error\.Empiricalmm\-shot NCC error versus task\-specificBr\(t\)B\_\{r\}^\{\(t\)\}across CelebA attributes, for several ranksrr\(panels\) and support sizesmm\(colors\)\. Each point corresponds to a downstream attribute, and dashed curves are the corresponding upper bounds from Thm\.[3\.4](https://arxiv.org/html/2609.38393#S3.Thmtheorem4)\.\(a\) W\-MSE\-CelebA\(b\) I\-JEPA\-IM1K\(c\) Centroid RMSE\(d\) Task\-axis overlapFigure 6:Multitask centroid geometry on CelebA\.\(a–b\) Illustrative held\-out joint centroids for three attributes, compared with the hyperrectangle predicted by Thm\.[3\.5](https://arxiv.org/html/2609.38393#S3.Thmtheorem5); triplets are selected on the training split\. \(c–d\) Results over all 1646 eligible triplets per encoder, without recoverability or orthogonality screening\. Larger mean task recoverability is associated with lower normalized centroid RMSE and smaller maximum pairwise task\-axis overlap\. Curves show bin means and shaded regions show interquartile ranges; selection and normalization details are in App\.[B\.1](https://arxiv.org/html/2609.38393#A2.SS1)\.\(a\) W\-MSE\-CelebA\(b\) VICReg\-CelebA\(c\) Barlow\-IM1K\(d\) I\-JEPA\-IM1KFigure 7:Recoverability from the paired\-view spectral subspace\.Directly measured task recoverabilityB^\(t\)​\(F\)\\widehat\{B\}^\{\(t\)\}\(F\)versus its reconstructionB^r⋆\(t\)\\widehat\{B\}\_\{r^\{\\star\}\}^\{\(t\)\}from the label\-free paired\-view spectral modes\. Each point is one CelebA attribute; the dashed line is perfect agreement\. For ImageNet\-pretrained encoders, the retained subspace is reduced\-rank and the median absolute discrepancy is below0\.050\.05across models; additional encoders are shown in App\.[B\.1](https://arxiv.org/html/2609.38393#A2.SS1)\.### 5\.1Settings

We test the main empirical consequences of the theory on natural and synthetic data\. Full dataset, preprocessing, and training details are given in App\.[A](https://arxiv.org/html/2609.38393#A1)\.

Datasets\.The main experiments use CelebA\([Liu et al\., 2015](https://arxiv.org/html/2609.38393#bib.bib2)\), whose 40 binary attributes provide a collection of downstream semantic tasks\. For each attribute, we balance the two classes by subsampling so that the empirical problem matches the binary setting of Sec\.[2](https://arxiv.org/html/2609.38393#S2)\. Controlled experiments on dSprites\([Matthey et al\., 2017](https://arxiv.org/html/2609.38393#bib.bib3)\)and 3DShapes\([Burgess and Kim, 2018](https://arxiv.org/html/2609.38393#bib.bib4)\)are reported in App\.[B\.2](https://arxiv.org/html/2609.38393#A2.SS2)\.

SSL methods and rank\-rrrepresentations\.We train W\-MSE\([Ermolov et al\., 2021](https://arxiv.org/html/2609.38393#bib.bib75)\)and VICReg\([Bardes et al\., 2022](https://arxiv.org/html/2609.38393#bib.bib5)\)from scratch on CelebA with ResNet\-50\([He et al\., 2016](https://arxiv.org/html/2609.38393#bib.bib21)\)backbones\. We also evaluate publicly available ImageNet\-1K\([Deng et al\., 2009](https://arxiv.org/html/2609.38393#bib.bib20)\)\-pretrained VICReg and Barlow Twins\([Zbontar et al\., 2021](https://arxiv.org/html/2609.38393#bib.bib53)\)ResNet\-50 encoders and an I\-JEPA\([Assran et al\., 2023](https://arxiv.org/html/2609.38393#bib.bib50)\)ViT\-H/14 encoder\. To vary representation rank, we estimate a paired\-view spectral basis from unlabeled data and retain its leadingrrdirections\. The retained features are then centered and whitened before the label\-dependent measurements ofBrB\_\{r\}below\.

Estimating semantic recoverability\.For a balanced downstream task, letβ^=𝔼^​\[Y​F​\(X\)\]\\hat\{\\beta\}=\\widehat\{\\mathbb\{E\}\}\[YF\(X\)\]andG^=𝔼^​\[F⁡\(X\)​F​\(X\)⊤\]\\widehat\{G\}=\\widehat\{\\mathbb\{E\}\}\[F\(X\)F\(X\)^\{\\top\}\]\. We estimate captured posterior energy byB^​\(F\)=β^⊤​G^−1​β^\\widehat\{B\}\(F\)=\\hat\{\\beta\}^\{\\top\}\\widehat\{G\}^\{\-1\}\\hat\{\\beta\}\. This form accounts for finite\-sample deviations from exact whitening\. WhenG^=Ir\\widehat\{G\}=I\_\{r\}, it reduces toB^​\(F\)=‖𝔼^​\[Y​F​\(X\)\]‖22\\widehat\{B\}\(F\)=\\\|\\widehat\{\\mathbb\{E\}\}\[YF\(X\)\]\\\|\_\{2\}^\{2\}\. Unless noted otherwise, all reported recoverability values use this estimator\.

### 5\.2Results

Recoverability is task\-dependent\.Because the rank\-rrsubspaces are nested,Br\(t\)B\_\{r\}^\{\(t\)\}is nondecreasing inrr\. Fig\.[3](https://arxiv.org/html/2609.38393#S5.F3)nevertheless shows substantial variation across CelebA attributes: the most recoverable tasks approachBr\(t\)≈1B\_\{r\}^\{\(t\)\}\\approx 1at moderate ranks, while the average over all 40 attributes remains lower\. SSL representations also exceed random\-initialization baselines across a broad range of ranks\. Tasks differ markedly in both saturation rate and attainable recoverability\. The same pattern appears for additional pretrained encoders in App\.[B\.1](https://arxiv.org/html/2609.38393#A2.SS1)and on dSprites and 3DShapes in App\.[B\.2](https://arxiv.org/html/2609.38393#A2.SS2)\.

Recoverability predicts directional geometry\.Prop\.[3\.1](https://arxiv.org/html/2609.38393#S3.Thmtheorem1)gives the exact relationV~F=\(1−B⁡\(F\)\)/\(2​B​\(F\)\)\\tilde\{V\}\_\{F\}=\(1\-B\(F\)\)/\(2B\(F\)\)for centered, whitened representations\. We test it with a split\-sample protocol: recoverability is estimated on split A and directional CDNV independently on split B\. Fig\.[4](https://arxiv.org/html/2609.38393#S5.F4)shows close agreement over more than two orders of magnitude, across ranks and SSL representations\. Additional encoder results are given in App\.[B\.1](https://arxiv.org/html/2609.38393#A2.SS1)\.

Recoverability controls few\-shot transfer\.Thm\.[3\.4](https://arxiv.org/html/2609.38393#S3.Thmtheorem4)predicts that, at fixed rank and shot count, largerBr\(t\)B\_\{r\}^\{\(t\)\}yields lower NCC error, while increasing the support sizemmreduces centroid\-estimation error\. Fig\.[5](https://arxiv.org/html/2609.38393#S5.F5)shows both trends across CelebA attributes: empirical error decreases with recoverability, and the bound tightens asmmgrows\. The bound remains above the observed error throughout the evaluated configurations and becomes more informative for well\-recovered tasks\. App\.[B\.1](https://arxiv.org/html/2609.38393#A2.SS1)reports the remaining encoders and Tab\.[1](https://arxiv.org/html/2609.38393#A2.T1)shows comparisons with earlier neural\-collapse\-based transfer bounds\([Galanti et al\., 2023](https://arxiv.org/html/2609.38393#bib.bib14);[Luthra et al\., 2025](https://arxiv.org/html/2609.38393#bib.bib63);[Luthra et al\., 2026](https://arxiv.org/html/2609.38393#bib.bib1)\)\.

Multiple semantic factors\.Thm\.[3\.5](https://arxiv.org/html/2609.38393#S3.Thmtheorem5)predicts two related effects: well\-recovered tasks should have nearly orthogonal semantic axes when cross\-task correlations and residual uncertainty are small, and their joint centroids should approach the corresponding factorial hyperrectangle\. Fig\.[6](https://arxiv.org/html/2609.38393#S5.F6)\(a–b\) shows two illustrative triplets selected on the training split using fixed recoverability and near\-orthogonality thresholds \(see App\.[B\.1](https://arxiv.org/html/2609.38393#A2.SS1)\); the displayed centroids are then measured on held\-out test images\. Panels \(c–d\) use every eligible triplet, with no additional recoverability or orthogonality screening\. Across encoders, higher mean task recoverability is associated with both lower centroid RMSE and smaller task\-axis overlap, consistent with the geometry predicted by the theorem\.

The spectral subspace recovers semantic energy\.Cor\.[4\.2](https://arxiv.org/html/2609.38393#S4.Thmtheorem2)identifies recoverability at the population optimum with posterior energy in the leading two\-view spectral subspace\. We test this empirically by comparingB^\(t\)​\(F\)\\widehat\{B\}^\{\(t\)\}\(F\), measured directly from the learned representation, withB^r⋆\(t\)\\widehat\{B\}\_\{r^\{\\star\}\}^\{\(t\)\}, reconstructed from the estimated paired\-view spectral modes at the label\-free rankr⋆r^\{\\star\}defined in App\.[B\.1](https://arxiv.org/html/2609.38393#A2.SS1)\. Fig\.[7](https://arxiv.org/html/2609.38393#S5.F7)shows close agreement\. For CelebA\-trained W\-MSE and VICReg, the retained basis spans nearly the full projector space, so near\-exact agreement is expected\. For ImageNet\-pretrained encoders, the lower\-dimensional spectral subspace still recovers most of the full\-representation semantic energy, with median absolute discrepancies below0\.050\.05across models\. The subspace is estimated without downstream labels\. For I\-JEPA, the paired\-view operator is a surrogate analysis kernel rather than the masking distribution used in pretraining; more details can be found in App\.[A](https://arxiv.org/html/2609.38393#A1)\.

## 6Discussion and Limitations

Our results identifysemantic recoverabilityas the key quantity linking same\-instance SSL to downstream geometry and transfer\. The task\-specific captured posterior energyBr\(t\)B\_\{r\}^\{\(t\)\}provides a unified account of directional neural collapse, probe–centroid alignment, multitask centroid structure, and few\-shot performance\. These results are derived primarily for population\-optimal, centered and whitened representations\. We also analyze a covariance\-regularized surrogate, but extending the theory to finite\-sample training and imperfect optimization remains an important direction for future work\. Our main analysis also focuses on balanced binary tasks and conditionally i\.i\.d\. two\-view generation\. Extending the framework to multiclass tasks and more general view dependencies is another natural direction\.

## AI use statement

In this work, we used generative AI tools \(ChatGPT and Claude Code\) to assist with developing and refining theoretical frameworks, formulating mathematical claims, discussing proof strategies, and providing feedback on experimental methodology\. These tools also assisted with experimental design and implementation, including code generation and refinement, as well as research brainstorming, literature exploration, manuscript organization, drafting and editing, and LaTeX presentation\.

All AI\-assisted theoretical arguments, mathematical derivations, experimental designs, and generated or modified code were reviewed by the authors for correctness and consistency\. Relevant literature was independently examined, and all experimental results and their interpretation were evaluated by the authors\. The authors take full responsibility for the final content of the paper, including its text, claims, code, results, and other artifacts produced with the assistance of generative AI\.

## Ethics statement

This work studies theoretical and empirical properties of representations learned through self\-supervised learning\. It does not involve new human\-subject studies or the collection of new personal data\. The experiments use existing datasets and pretrained models for research purposes\. The work does not develop or evaluate a high\-risk application\.

## Reproducibility statement

We state the assumptions and formal results in the main text and provide complete proofs in the appendix\. The experimental sections and supplementary material document the datasets, pretrained representations, evaluation protocols, preprocessing, training procedures, and implementation details needed to reproduce our empirical results\. We also specify the procedures used to estimate semantic recoverability, directional CDNV, spectral structure, and few\-shot transfer performance\.

## References

- Alonet al\.\(2024\)N\. Alon, D\. Avdiukhin, D\. Elboim, O\. Fischer, and G\. YaroslavtsevOptimal sample complexity of contrastive learning\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=NU9AYHJvYe)Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p4.1)\.
- Andrewet al\.\(2013\)G\. Andrew, R\. Arora, J\. Bilmes, and K\. LivescuDeep canonical correlation analysis\.InProceedings of the 30th International Conference on Machine Learning,S\. Dasgupta and D\. McAllester \(Eds\.\),Proceedings of Machine Learning Research, Vol\.28,Atlanta, Georgia, USA,pp\. 1247–1255\.External Links:[Link](https://proceedings.mlr.press/v28/andrew13.html)Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p3.1)\.
- Ashet al\.\(2022\)J\. Ash, S\. Goel, A\. Krishnamurthy, and D\. MisraInvestigating the role of negatives in contrastive representation learning\.InProceedings of The 25th International Conference on Artificial Intelligence and Statistics,G\. Camps\-Valls, F\. J\. R\. Ruiz, and I\. Valera \(Eds\.\),Proceedings of Machine Learning Research, Vol\.151,pp\. 7187–7209\.External Links:[Link](https://proceedings.mlr.press/v151/ash22a.html)Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p4.1)\.
- Asoodehet al\.\(2015\)S\. Asoodeh, F\. Alajaji, and T\. LinderOn maximal correlation, mutual information and data privacy\.In2015 IEEE 14th Canadian Workshop on Information Theory \(CWIT\),pp\. 27–31\.External Links:[Document](https://dx.doi.org/10.1109/CWIT.2015.7255145)Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p3.1)\.
- Assranet al\.\(2023\)M\. Assran, Q\. Duval, I\. Misra, P\. Bojanowski, P\. Vincent, M\. G\. Rabbat, Y\. LeCun, and N\. BallasSelf\-supervised learning from images with a joint\-embedding predictive architecture\.InCVPR,pp\. 15619–15629\.External Links:[Link](https://doi.org/10.1109/CVPR52729.2023.01499)Cited by:[Appendix A](https://arxiv.org/html/2609.38393#A1.p12.1),[§4\.1](https://arxiv.org/html/2609.38393#S4.SS1.p2.1),[§5\.1](https://arxiv.org/html/2609.38393#S5.SS1.p3.1)\.
- Awasthiet al\.\(2022\)P\. Awasthi, N\. Dikkala, and P\. KamathDo more negative samples necessarily hurt in contrastive learning?\.InProceedings of the 39th International Conference on Machine Learning,K\. Chaudhuri, S\. Jegelka, L\. Song, C\. Szepesvari, G\. Niu, and S\. Sabato \(Eds\.\),Proceedings of Machine Learning Research, Vol\.162,pp\. 1101–1116\.External Links:[Link](https://proceedings.mlr.press/v162/awasthi22b.html)Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p4.1)\.
- Bachmanet al\.\(2019\)P\. Bachman, R\. D\. Hjelm, and W\. BuchwalterLearning representations by maximizing mutual information across views\.InAdvances in Neural Information Processing Systems,Vol\.32,pp\. 15509–15519\.External Links:[Link](https://proceedings.neurips.cc/paper/2019/hash/ddf354219aac374f1d40b7e760ee5bb7-Abstract.html)Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p4.1)\.
- Baevskiet al\.\(2022\)A\. Baevski, W\. Hsu, Q\. Xu, A\. Babu, J\. Gu, and M\. AuliData2vec: a general framework for self\-supervised learning in speech, vision and language\.InProceedings of the 39th International Conference on Machine Learning,K\. Chaudhuri, S\. Jegelka, L\. Song, C\. Szepesvari, G\. Niu, and S\. Sabato \(Eds\.\),Proceedings of Machine Learning Research, Vol\.162,pp\. 1298–1312\.External Links:[Link](https://proceedings.mlr.press/v162/baevski22a.html)Cited by:[§1](https://arxiv.org/html/2609.38393#S1.p1.1)\.
- Baevskiet al\.\(2020\)A\. Baevski, H\. Zhou, A\. Mohamed, and M\. AuliWav2vec 2\.0: a framework for self\-supervised learning of speech representations\.InProceedings of the 34th International Conference on Neural Information Processing Systems,NIPS ’20,Red Hook, NY, USA\.External Links:ISBN 9781713829546Cited by:[§1](https://arxiv.org/html/2609.38393#S1.p1.1)\.
- Balestrieroet al\.\(2023\)R\. Balestriero, M\. Ibrahim, V\. Sobal, A\. Morcos, S\. Shekhar, T\. Goldstein, F\. Bordes, A\. Bardes, G\. Mialon, Y\. Tian, A\. Schwarzschild, A\. G\. Wilson, J\. Geiping, Q\. Garrido, P\. Fernandez, A\. Bar, H\. Pirsiavash, Y\. LeCun, and M\. GoldblumA cookbook of self\-supervised learning\.External Links:2304\.12210,[Link](https://arxiv.org/abs/2304.12210)Cited by:[§1](https://arxiv.org/html/2609.38393#S1.p1.1)\.
- Balestriero and LeCun \(2022\)R\. Balestriero and Y\. LeCunContrastive and non\-contrastive self\-supervised learning recover global and local spectral embedding methods\.InAdvances in Neural Information Processing Systems,A\. H\. Oh, A\. Agarwal, D\. Belgrave, and K\. Cho \(Eds\.\),External Links:[Link](https://openreview.net/forum?id=jQgsZDspz5h)Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p3.1),[§1](https://arxiv.org/html/2609.38393#S1.p2.1),[§4\.1](https://arxiv.org/html/2609.38393#S4.SS1.p1.1)\.
- Bardeset al\.\(2022\)A\. Bardes, J\. Ponce, and Y\. LeCunVICReg: variance\-invariance\-covariance regularization for self\-supervised learning\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=xm6YD62D1Ub)Cited by:[Appendix A](https://arxiv.org/html/2609.38393#A1.p6.1),[Appendix A](https://arxiv.org/html/2609.38393#A1.p8.1),[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p3.1),[§4\.1](https://arxiv.org/html/2609.38393#S4.SS1.p1.1),[§4\.1](https://arxiv.org/html/2609.38393#S4.SS1.p2.1),[§5\.1](https://arxiv.org/html/2609.38393#S5.SS1.p3.1)\.
- Ben\-Shaulet al\.\(2023\)I\. Ben\-Shaul, R\. Shwartz\-Ziv, T\. Galanti, S\. Dekel, and Y\. LeCunReverse engineering self\-supervised learning\.InAdvances in Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=NsVEjx6YPd)Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p2.1)\.
- Breiman and Friedman \(1985\)L\. Breiman and J\. H\. FriedmanEstimating optimal transformations for multiple regression and correlation\.Journal of the American Statistical Association80\(391\),pp\. 580–598\.External Links:[Document](https://dx.doi.org/10.1080/01621459.1985.10478157),[Link](https://doi.org/10.1080/01621459.1985.10478157)Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p3.1)\.
- Burgess and Kim \(2018\)C\. Burgess and H\. Kim3D shapes dataset\.Note:https://github\.com/deepmind/3d\-shapes/Cited by:[§5\.1](https://arxiv.org/html/2609.38393#S5.SS1.p2.1)\.
- Caronet al\.\(2021\)M\. Caron, H\. Touvron, I\. Misra, H\. Jégou, J\. Mairal, P\. Bojanowski, and A\. JoulinEmerging properties in self\-supervised vision transformers\.InProceedings of the IEEE/CVF International Conference on Computer Vision \(ICCV\),pp\. 9650–9660\.Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p2.1),[§4\.1](https://arxiv.org/html/2609.38393#S4.SS1.p1.1)\.
- Chenet al\.\(2020\)T\. Chen, S\. Kornblith, M\. Norouzi, and G\. HintonA simple framework for contrastive learning of visual representations\.InProceedings of the 37th International Conference on Machine Learning,H\. D\. III and A\. Singh \(Eds\.\),Proceedings of Machine Learning Research, Vol\.119,pp\. 1597–1607\.External Links:[Link](https://proceedings.mlr.press/v119/chen20j.html)Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p2.1),[§1](https://arxiv.org/html/2609.38393#S1.p1.1),[§4\.1](https://arxiv.org/html/2609.38393#S4.SS1.p1.1)\.
- Chenet al\.\(2021\)T\. Chen, C\. Luo, and L\. LiIntriguing properties of contrastive losses\.Advances in Neural Information Processing Systems34,pp\. 11834–11845\.Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p2.1),[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p4.1)\.
- Denget al\.\(2009\)J\. Deng, W\. Dong, R\. Socher, L\. Li, K\. Li, and L\. Fei\-FeiImageNet: a large\-scale hierarchical image database\.In2009 IEEE Conference on Computer Vision and Pattern Recognition,pp\. 248–255\.External Links:[Document](https://dx.doi.org/10.1109/CVPR.2009.5206848)Cited by:[§5\.1](https://arxiv.org/html/2609.38393#S5.SS1.p3.1)\.
- Ermolovet al\.\(2021\)A\. Ermolov, A\. Siarohin, E\. Sangineto, and N\. SebeWhitening for self\-supervised representation learning\.InProceedings of the 38th International Conference on Machine Learning,M\. Meila and T\. Zhang \(Eds\.\),Proceedings of Machine Learning Research, Vol\.139,pp\. 3015–3024\.External Links:[Link](https://proceedings.mlr.press/v139/ermolov21a.html)Cited by:[Appendix A](https://arxiv.org/html/2609.38393#A1.p10.1),[Appendix A](https://arxiv.org/html/2609.38393#A1.p8.1),[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p3.1),[§1](https://arxiv.org/html/2609.38393#S1.p2.1),[§4\.1](https://arxiv.org/html/2609.38393#S4.SS1.p1.1),[§5\.1](https://arxiv.org/html/2609.38393#S5.SS1.p3.1)\.
- Fukumizuet al\.\(2007\)K\. Fukumizu, F\. R\. Bach, and A\. GrettonStatistical consistency of kernel canonical correlation analysis\.Journal of Machine Learning Research8\(14\),pp\. 361–383\.External Links:[Link](http://jmlr.org/papers/v8/fukumizu07a.html)Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p3.1)\.
- Galantiet al\.\(2023\)T\. Galanti, L\. Galanti, and I\. Ben\-ShaulComparative generalization bounds for deep neural networks\.Transactions on Machine Learning Research\.External Links:ISSN 2835\-8856,[Link](https://openreview.net/forum?id=162TqkUNPO)Cited by:[§B\.1](https://arxiv.org/html/2609.38393#A2.SS1.p8.1),[Table 1](https://arxiv.org/html/2609.38393#A2.T1.4.2.1),[§5\.2](https://arxiv.org/html/2609.38393#S5.SS2.p3.1)\.
- Galantiet al\.\(2022a\)T\. Galanti, A\. György, and M\. HutterGeneralization bounds for few\-shot transfer learning with pretrained classifiers\.External Links:2212\.12532,[Link](https://arxiv.org/abs/2212.12532)Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p2.1)\.
- Galantiet al\.\(2022b\)T\. Galanti, A\. György, and M\. HutterImproved generalization bounds for transfer learning via neural collapse\.InFirst Workshop on Pre\-training: Perspectives, Pitfalls, and Paths Forward at ICML 2022,External Links:[Link](https://openreview.net/forum?id=VrK7pKwOhT_)Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p2.1)\.
- Galantiet al\.\(2022c\)T\. Galanti, A\. György, and M\. HutterOn the role of neural collapse in transfer learning\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=SwIp410B6aQ)Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p2.1)\.
- Gaoet al\.\(2021\)T\. Gao, X\. Yao, and D\. ChenSimCSE: simple contrastive learning of sentence embeddings\.InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing,M\. Moens, X\. Huang, L\. Specia, and S\. W\. Yih \(Eds\.\),Online and Punta Cana, Dominican Republic,pp\. 6894–6910\.External Links:[Link](https://aclanthology.org/2021.emnlp-main.552/),[Document](https://dx.doi.org/10.18653/v1/2021.emnlp-main.552)Cited by:[§1](https://arxiv.org/html/2609.38393#S1.p1.1)\.
- Garridoet al\.\(2023\)Q\. Garrido, Y\. Chen, A\. Bardes, L\. Najman, and Y\. LeCunOn the duality between contrastive and non\-contrastive self\-supervised learning\.InThe Eleventh International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=kDEL91Dufpa)Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p3.1),[§1](https://arxiv.org/html/2609.38393#S1.p2.1),[§4\.1](https://arxiv.org/html/2609.38393#S4.SS1.p1.1)\.
- Goldblumet al\.\(2020\)M\. Goldblum, S\. Reich, L\. Fowl, R\. Ni, V\. Cherepanova, and T\. GoldsteinUnraveling meta\-learning: understanding feature representations for few\-shot tasks\.InProceedings of the 37th International Conference on Machine Learning,H\. D\. III and A\. Singh \(Eds\.\),Proceedings of Machine Learning Research, Vol\.119,pp\. 3607–3616\.External Links:[Link](https://proceedings.mlr.press/v119/goldblum20a.html)Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p2.1)\.
- Goyalet al\.\(2017\)P\. Goyal, P\. Dollár, R\. Girshick, P\. Noordhuis, L\. Wesolowski, A\. Kyrola, A\. Tulloch, Y\. Jia, and K\. HeAccurate, large minibatch SGD: training ImageNet in 1 hour\.External Links:1706\.02677,[Link](https://arxiv.org/abs/1706.02677)Cited by:[Appendix A](https://arxiv.org/html/2609.38393#A1.p8.1)\.
- Grillet al\.\(2020\)J\. Grill, F\. Strub, F\. Altché, C\. Tallec, P\. Richemond, E\. Buchatskaya, C\. Doersch, B\. Avila Pires, Z\. Guo, M\. Gheshlaghi Azar, B\. Piot, k\. kavukcuoglu, R\. Munos, and M\. ValkoBootstrap your own latent \- a new approach to self\-supervised learning\.InAdvances in Neural Information Processing Systems,H\. Larochelle, M\. Ranzato, R\. Hadsell, M\.F\. Balcan, and H\. Lin \(Eds\.\),Vol\.33,pp\. 21271–21284\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2020/file/f3ada80d5c4ee70142b17b8192b2958e-Paper.pdf)Cited by:[§4\.1](https://arxiv.org/html/2609.38393#S4.SS1.p1.1),[§4\.1](https://arxiv.org/html/2609.38393#S4.SS1.p2.1)\.
- Guiet al\.\(2023\)Y\. Gui, C\. Ma, and Y\. ZhongUnraveling projection heads in contrastive learning: insights from expansion and shrinkage\.External Links:2306\.03335,[Link](https://arxiv.org/abs/2306.03335)Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p4.1)\.
- Guptaet al\.\(2022\)K\. Gupta, T\. Ajanthan, A\. van den Hengel, and S\. GouldUnderstanding and improving the role of projection head in self\-supervised learning\.InNeurIPS 2022 Workshop: Self\-Supervised Learning – Theory and Practice,External Links:2212\.11491,[Link](https://arxiv.org/abs/2212.11491)Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p4.1)\.
- Hanet al\.\(2022\)X\.Y\. Han, V\. Papyan, and D\. L\. DonohoNeural collapse under MSE loss: proximity to and dynamics on the central path\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=w1UbdvWH_R3)Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p2.1)\.
- Hannan \(1961\)E\. J\. HannanThe general theory of canonical correlation and its relation to functional analysis\.Journal of the Australian Mathematical Society2\(2\),pp\. 229–242\.External Links:[Document](https://dx.doi.org/10.1017/S1446788700026707)Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p3.1)\.
- HaoChen and Ma \(2023\)J\. Z\. HaoChen and T\. MaA theoretical study of inductive biases in contrastive learning\.InThe Eleventh International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=AuEgNlEAmed)Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p4.1)\.
- HaoChenet al\.\(2021\)J\. Z\. HaoChen, C\. Wei, A\. Gaidon, and T\. MaProvable guarantees for self\-supervised deep learning with spectral contrastive loss\.InAdvances in Neural Information Processing Systems,M\. Ranzato, A\. Beygelzimer, Y\. Dauphin, P\.S\. Liang, and J\. W\. Vaughan \(Eds\.\),Vol\.34,pp\. 5000–5011\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2021/file/27debb435021eb68b3965290b5e24c49-Paper.pdf)Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p2.1),[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p4.1)\.
- Hardoonet al\.\(2004\)D\. R\. Hardoon, S\. Szedmak, and J\. Shawe\-TaylorCanonical correlation analysis: an overview with application to learning methods\.Neural Computation16\(12\),pp\. 2639–2664\.External Links:[Document](https://dx.doi.org/10.1162/0899766042321814)Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p3.1)\.
- Heet al\.\(2022\)K\. He, X\. Chen, S\. Xie, Y\. Li, P\. Dollár, and R\. GirshickMasked autoencoders are scalable vision learners\.InProceedings of the IEEE/CVF conference on computer vision and pattern recognition,pp\. 16000–16009\.Cited by:[§1](https://arxiv.org/html/2609.38393#S1.p1.1)\.
- Heet al\.\(2020\)K\. He, H\. Fan, Y\. Wu, S\. Xie, and R\. GirshickMomentum contrast for unsupervised visual representation learning\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 9729–9738\.External Links:[Link](https://openaccess.thecvf.com/content_CVPR_2020/html/He_Momentum_Contrast_for_Unsupervised_Visual_Representation_Learning_CVPR_2020_paper.html)Cited by:[§1](https://arxiv.org/html/2609.38393#S1.p1.1)\.
- Heet al\.\(2016\)K\. He, X\. Zhang, S\. Ren, and J\. SunDeep residual learning for image recognition\.InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 770–778\.External Links:[Link](https://openaccess.thecvf.com/content_cvpr_2016/html/He_Deep_Residual_Learning_CVPR_2016_paper.html)Cited by:[Appendix A](https://arxiv.org/html/2609.38393#A1.p11.1),[§5\.1](https://arxiv.org/html/2609.38393#S5.SS1.p3.1)\.
- Hotelling \(1936\)H\. HotellingRelations between two sets of variates\.Biometrika28\(3/4\),pp\. 321–377\.External Links:[Document](https://dx.doi.org/10.2307/2333955),[Link](https://doi.org/10.2307/2333955)Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p3.1)\.
- Hsuet al\.\(2021\)W\. Hsu, B\. Bolte, Y\. H\. Tsai, K\. Lakhotia, R\. Salakhutdinov, and A\. MohamedHuBERT: self\-supervised speech representation learning by masked prediction of hidden units\.IEEE/ACM Trans\. Audio, Speech and Lang\. Proc\.29,pp\. 3451–3460\.External Links:ISSN 2329\-9290,[Link](https://doi.org/10.1109/TASLP.2021.3122291),[Document](https://dx.doi.org/10.1109/TASLP.2021.3122291)Cited by:[§1](https://arxiv.org/html/2609.38393#S1.p1.1)\.
- Jiaet al\.\(2021\)C\. Jia, Y\. Yang, Y\. Xia, Y\. Chen, Z\. Parekh, H\. Pham, Q\. Le, Y\. Sung, Z\. Li, and T\. DuerigScaling up visual and vision\-language representation learning with noisy text supervision\.InProceedings of the 38th International Conference on Machine Learning,M\. Meila and T\. Zhang \(Eds\.\),Proceedings of Machine Learning Research, Vol\.139,pp\. 4904–4916\.External Links:[Link](https://proceedings.mlr.press/v139/jia21b.html)Cited by:[§1](https://arxiv.org/html/2609.38393#S1.p1.1)\.
- Lai and Fyfe \(2000\)P\. L\. Lai and C\. FyfeKernel and nonlinear canonical correlation analysis\.International Journal of Neural Systems10\(5\),pp\. 365–377\.External Links:[Document](https://dx.doi.org/10.1142/S012906570000034X),[Link](https://doi.org/10.1142/S012906570000034X)Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p3.1)\.
- Liuet al\.\(2015\)Z\. Liu, P\. Luo, X\. Wang, and X\. TangDeep learning face attributes in the wild\.InProceedings of the IEEE International Conference on Computer Vision \(ICCV\),pp\. 3730–3738\.External Links:[Document](https://dx.doi.org/10.1109/ICCV.2015.425),[Link](https://openaccess.thecvf.com/content_iccv_2015/html/Liu_Deep_Learning_Face_ICCV_2015_paper.html)Cited by:[Appendix A](https://arxiv.org/html/2609.38393#A1.p1.1),[§5\.1](https://arxiv.org/html/2609.38393#S5.SS1.p2.1)\.
- Loshchilov and Hutter \(2017\)I\. Loshchilov and F\. HutterSGDR: stochastic gradient descent with warm restarts\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=Skq89Scxx)Cited by:[Appendix A](https://arxiv.org/html/2609.38393#A1.p8.1)\.
- Loshchilov and Hutter \(2019\)I\. Loshchilov and F\. HutterDecoupled weight decay regularization\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=Bkg6RiCqY7)Cited by:[Appendix A](https://arxiv.org/html/2609.38393#A1.p8.1)\.
- Luthraet al\.\(2026\)A\. Luthra, Y\. Salunkhe, and T\. GalantiDirectional neural collapse explains few\-shot transfer in self\-supervised learning\.InInternational Conference on Machine Learning,External Links:2603\.03530,[Link](https://arxiv.org/abs/2603.03530)Cited by:[§B\.1](https://arxiv.org/html/2609.38393#A2.SS1.p8.1),[Table 1](https://arxiv.org/html/2609.38393#A2.T1.4.4.1),[§5\.2](https://arxiv.org/html/2609.38393#S5.SS2.p3.1)\.
- Luthraet al\.\(2025\)A\. Luthra, T\. Yang, and T\. GalantiSelf\-supervised contrastive learning is approximately supervised contrastive learning\.InAdvances in Neural Information Processing Systems,Vol\.38,pp\. 46671–46708\.External Links:[Document](https://dx.doi.org/10.52202/085713-1392),[Link](https://proceedings.neurips.cc/paper_files/paper/2025/hash/3b85cbbaab0008dfc52250a857725f70-Abstract-Conference.html)Cited by:[§B\.1](https://arxiv.org/html/2609.38393#A2.SS1.p8.1),[Table 1](https://arxiv.org/html/2609.38393#A2.T1.4.3.1),[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p2.1),[§1](https://arxiv.org/html/2609.38393#S1.p2.1),[§1](https://arxiv.org/html/2609.38393#S1.p5.1),[§5\.2](https://arxiv.org/html/2609.38393#S5.SS2.p3.1)\.
- Mattheyet al\.\(2017\)L\. Matthey, I\. Higgins, D\. Hassabis, and A\. LerchnerDSprites: disentanglement testing sprites dataset\.Note:https://github\.com/deepmind/dsprites\-dataset/Cited by:[§5\.1](https://arxiv.org/html/2609.38393#S5.SS1.p2.1)\.
- McAllester and Stratos \(2020\)D\. McAllester and K\. StratosFormal limitations on the measurement of mutual information\.InProceedings of the Twenty Third International Conference on Artificial Intelligence and Statistics,S\. Chiappa and R\. Calandra \(Eds\.\),Proceedings of Machine Learning Research, Vol\.108,pp\. 875–884\.External Links:[Link](https://proceedings.mlr.press/v108/mcallester20a.html)Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p4.1)\.
- Michaeliet al\.\(2016\)T\. Michaeli, W\. Wang, and K\. LivescuNonparametric canonical correlation analysis\.InProceedings of The 33rd International Conference on Machine Learning,M\. F\. Balcan and K\. Q\. Weinberger \(Eds\.\),Proceedings of Machine Learning Research, Vol\.48,New York, New York, USA,pp\. 1967–1976\.External Links:[Link](https://proceedings.mlr.press/v48/michaeli16.html)Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p3.1)\.
- Nozawa and Sato \(2021\)K\. Nozawa and I\. SatoUnderstanding negative samples in instance discriminative self\-supervised representation learning\.InAdvances in Neural Information Processing Systems,M\. Ranzato, A\. Beygelzimer, Y\. Dauphin, P\.S\. Liang, and J\. W\. Vaughan \(Eds\.\),Vol\.34,pp\. 5784–5797\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2021/file/2dace78f80bc92e6d7493423d729448e-Paper.pdf)Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p4.1)\.
- Oquabet al\.\(2024\)M\. Oquab, T\. Darcet, T\. Moutakanni, H\. V\. Vo, M\. Szafraniec, V\. Khalidov, P\. Fernandez, D\. Haziza, F\. Massa, A\. El\-Nouby, M\. Assran, N\. Ballas, W\. Galuba, R\. Howes, P\. Huang, S\. Li, I\. Misra, M\. Rabbat, V\. Sharma, G\. Synnaeve, H\. Xu, H\. Jégou, J\. Mairal, P\. Labatut, A\. Joulin, and P\. BojanowskiDINOv2: learning robust visual features without supervision\.Transactions on Machine Learning Research\.External Links:[Link](https://openreview.net/forum?id=a68SUt6zFt)Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p2.1),[§1](https://arxiv.org/html/2609.38393#S1.p1.1)\.
- Papyanet al\.\(2020\)V\. Papyan, X\. Y\. Han, and D\. L\. DonohoPrevalence of neural collapse during the terminal phase of deep learning training\.Proceedings of the National Academy of Sciences117\(40\),pp\. 24652–24663\.External Links:[Document](https://dx.doi.org/10.1073/pnas.2015509117),[Link](https://www.pnas.org/doi/abs/10.1073/pnas.2015509117),https://www\.pnas\.org/doi/pdf/10\.1073/pnas\.2015509117Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p2.1)\.
- Parulekaret al\.\(2023\)A\. Parulekar, L\. Collins, K\. Shanmugam, A\. Mokhtari, and S\. ShakkottaiInfoNCE loss provably learns cluster\-preserving representations\.InProceedings of Thirty Sixth Conference on Learning Theory,G\. Neu and L\. Rosasco \(Eds\.\),Proceedings of Machine Learning Research, Vol\.195,pp\. 1914–1961\.External Links:[Link](https://proceedings.mlr.press/v195/parulekar23a.html)Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p4.1)\.
- Qiuet al\.\(2023\)Z\. Qiu, Q\. Hu, Z\. Yuan, D\. Zhou, L\. Zhang, and T\. YangNot all semantics are created equal: contrastive self\-supervised learning with automatic temperature individualization\.InProceedings of the 40th International Conference on Machine Learning,A\. Krause, E\. Brunskill, K\. Cho, B\. Engelhardt, S\. Sabato, and J\. Scarlett \(Eds\.\),Proceedings of Machine Learning Research, Vol\.202,pp\. 28389–28421\.External Links:[Link](https://proceedings.mlr.press/v202/qiu23a.html)Cited by:[§1](https://arxiv.org/html/2609.38393#S1.p2.1)\.
- Radfordet al\.\(2021\)A\. Radford, J\. W\. Kim, C\. Hallacy, A\. Ramesh, G\. Goh, S\. Agarwal, G\. Sastry, A\. Askell, P\. Mishkin, J\. Clark, G\. Krueger, and I\. SutskeverLearning transferable visual models from natural language supervision\.InProceedings of the 38th International Conference on Machine Learning,M\. Meila and T\. Zhang \(Eds\.\),Proceedings of Machine Learning Research, Vol\.139,pp\. 8748–8763\.External Links:[Link](https://proceedings.mlr.press/v139/radford21a.html)Cited by:[§1](https://arxiv.org/html/2609.38393#S1.p1.1)\.
- Reimers and Gurevych \(2019\)N\. Reimers and I\. GurevychSentence\-BERT: sentence embeddings using Siamese BERT\-networks\.InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing \(EMNLP\-IJCNLP\),K\. Inui, J\. Jiang, V\. Ng, and X\. Wan \(Eds\.\),Hong Kong, China,pp\. 3982–3992\.External Links:[Link](https://aclanthology.org/D19-1410/),[Document](https://dx.doi.org/10.18653/v1/D19-1410)Cited by:[§1](https://arxiv.org/html/2609.38393#S1.p1.1)\.
- Saunshiet al\.\(2022\)N\. Saunshi, J\. Ash, S\. Goel, D\. Misra, C\. Zhang, S\. Arora, S\. Kakade, and A\. KrishnamurthyUnderstanding contrastive learning requires incorporating inductive biases\.InProceedings of the 39th International Conference on Machine Learning,K\. Chaudhuri, S\. Jegelka, L\. Song, C\. Szepesvari, G\. Niu, and S\. Sabato \(Eds\.\),Proceedings of Machine Learning Research, Vol\.162,pp\. 19250–19286\.External Links:[Link](https://proceedings.mlr.press/v162/saunshi22a.html)Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p4.1)\.
- Saunshiet al\.\(2019\)N\. Saunshi, O\. Plevrakis, S\. Arora, M\. Khodak, and H\. KhandeparkarA theoretical analysis of contrastive unsupervised representation learning\.InProceedings of the 36th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.97,pp\. 5628–5637\.External Links:[Link](https://proceedings.mlr.press/v97/saunshi19a.html)Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p2.1),[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p4.1)\.
- Schneideret al\.\(2019\)S\. Schneider, A\. Baevski, R\. Collobert, and M\. AuliWav2vec: unsupervised pre\-training for speech recognition\.InInterspeech 2019,pp\. 3465–3469\.External Links:[Document](https://dx.doi.org/10.21437/Interspeech.2019-1873),[Link](https://www.isca-archive.org/interspeech_2019/schneider19_interspeech.html)Cited by:[§1](https://arxiv.org/html/2609.38393#S1.p1.1)\.
- Shenet al\.\(2022\)K\. Shen, R\. M\. Jones, A\. Kumar, S\. M\. Xie, J\. Z\. Haochen, T\. Ma, and P\. LiangConnect, not collapse: explaining contrastive learning for unsupervised domain adaptation\.InProceedings of the 39th International Conference on Machine Learning,K\. Chaudhuri, S\. Jegelka, L\. Song, C\. Szepesvari, G\. Niu, and S\. Sabato \(Eds\.\),Proceedings of Machine Learning Research, Vol\.162,pp\. 19847–19878\.External Links:[Link](https://proceedings.mlr.press/v162/shen22d.html)Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p4.1)\.
- Taoet al\.\(2022\)C\. Tao, H\. Wang, X\. Zhu, J\. Dong, S\. Song, G\. Huang, and J\. DaiExploring the equivalence of siamese self\-supervised learning via a unified gradient framework\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 14431–14440\.Cited by:[§4\.1](https://arxiv.org/html/2609.38393#S4.SS1.p1.1),[§4\.1](https://arxiv.org/html/2609.38393#S4.SS1.p2.1)\.
- Tianet al\.\(2020\)Y\. Tian, C\. Sun, B\. Poole, D\. Krishnan, C\. Schmid, and P\. IsolaWhat makes for good views for contrastive learning?\.InAdvances in Neural Information Processing Systems,H\. Larochelle, M\. Ranzato, R\. Hadsell, M\.F\. Balcan, and H\. Lin \(Eds\.\),Vol\.33,pp\. 6827–6839\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2020/file/4c2e5eaae9152079b9e95845750bb9ab-Paper.pdf)Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p4.1)\.
- Tian \(2023\)Y\. TianUnderstanding the role of nonlinearity in training dynamics of contrastive learning\.InThe Eleventh International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=s130rTE3U_X)Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p4.1)\.
- Toshet al\.\(2021\)C\. Tosh, A\. Krishnamurthy, and D\. HsuContrastive learning, multi\-view redundancy, and linear models\.InAlgorithmic Learning Theory,pp\. 1179–1206\.Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p4.1)\.
- Tschannenet al\.\(2020\)M\. Tschannen, J\. Djolonga, P\. K\. Rubenstein, S\. Gelly, and M\. LucicOn mutual information maximization for representation learning\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=rkxoh24FPH)Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p4.1)\.
- Tschannenet al\.\(2025\)M\. Tschannen, A\. Gritsenko, X\. Wang, M\. F\. Naeem, I\. Alabdulmohsin, N\. Parthasarathy, T\. Evans, L\. Beyer, Y\. Xia, B\. Mustafa, O\. Hénaff, J\. Harmsen, A\. Steiner, and X\. ZhaiSigLIP 2: multilingual vision\-language encoders with improved semantic understanding, localization, and dense features\.External Links:2502\.14786,[Link](https://arxiv.org/abs/2502.14786)Cited by:[§1](https://arxiv.org/html/2609.38393#S1.p1.1)\.
- Wang and Liu \(2021\)F\. Wang and H\. LiuUnderstanding the behaviour of contrastive loss\.InProceedings of the IEEE/CVF conference on computer vision and pattern recognition,pp\. 2495–2504\.Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p2.1),[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p4.1)\.
- Wang and Isola \(2020\)T\. Wang and P\. IsolaUnderstanding contrastive representation learning through alignment and uniformity on the hypersphere\.InInternational Conference on Machine Learning,pp\. 9929–9939\.Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p2.1),[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p3.1),[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p4.1),[§1](https://arxiv.org/html/2609.38393#S1.p2.1),[§4\.1](https://arxiv.org/html/2609.38393#S4.SS1.p1.1)\.
- Wanget al\.\(2023\)Z\. Wang, Y\. Luo, L\. Zheng, Z\. Huang, and M\. BaktashmotlaghHow far pre\-trained models are from neural collapse on the target dataset informs their transferability\.In2023 IEEE/CVF International Conference on Computer Vision \(ICCV\),pp\. 5526–5535\.External Links:[Document](https://dx.doi.org/10.1109/ICCV51070.2023.00511)Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p2.1)\.
- Wen and Li \(2021\)Z\. Wen and Y\. LiToward understanding the feature learning process of self\-supervised contrastive learning\.InProceedings of the 38th International Conference on Machine Learning,M\. Meila and T\. Zhang \(Eds\.\),Proceedings of Machine Learning Research, Vol\.139,pp\. 11112–11122\.External Links:[Link](https://proceedings.mlr.press/v139/wen21c.html)Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p4.1)\.
- Wenget al\.\(2025\)X\. Weng, J\. An, X\. Ma, B\. Qi, J\. Luo, X\. Yang, J\. S\. Dong, and L\. HuangClustering properties of self\-supervised learning\.InProceedings of the 42nd International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.267,pp\. 66597–66616\.External Links:[Link](https://proceedings.mlr.press/v267/weng25a.html)Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p2.1)\.
- Wenget al\.\(2022\)X\. Weng, L\. Huang, L\. Zhao, R\. M\. Anwer, S\. Khan, and F\. KhanAn investigation into whitening loss for self\-supervised learning\.InAdvances in Neural Information Processing Systems,A\. H\. Oh, A\. Agarwal, D\. Belgrave, and K\. Cho \(Eds\.\),External Links:[Link](https://openreview.net/forum?id=BbUxkmrstyk)Cited by:[§1](https://arxiv.org/html/2609.38393#S1.p2.1),[§4\.1](https://arxiv.org/html/2609.38393#S4.SS1.p1.1)\.
- Zbontaret al\.\(2021\)J\. Zbontar, L\. Jing, I\. Misra, Y\. LeCun, and S\. DenyBarlow twins: self\-supervised learning via redundancy reduction\.InInternational conference on machine learning,pp\. 12310–12320\.Cited by:[Appendix A](https://arxiv.org/html/2609.38393#A1.p12.1),[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p3.1),[§1](https://arxiv.org/html/2609.38393#S1.p1.1),[§4\.1](https://arxiv.org/html/2609.38393#S4.SS1.p1.1),[§4\.1](https://arxiv.org/html/2609.38393#S4.SS1.p2.1),[§5\.1](https://arxiv.org/html/2609.38393#S5.SS1.p3.1)\.
- Zhaiet al\.\(2025\)R\. Zhai, K\. Yang, B\. Varıcı, C\. Tsai, J\. Z\. Kolter, and P\. K\. RavikumarContextures: representations from contexts\.InProceedings of the 42nd International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.267,pp\. 74318–74347\.External Links:[Link](https://proceedings.mlr.press/v267/zhai25c.html)Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p3.1),[§1](https://arxiv.org/html/2609.38393#S1.p2.1)\.
- Zhai \(2025\)R\. ZhaiContextures: the mechanism of representation learning\.Ph\.D\. Thesis,Carnegie Mellon University\.Note:PhD dissertation, CMU\-CS\-25\-104External Links:2504\.19792,[Link](https://arxiv.org/abs/2504.19792)Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p3.1),[§1](https://arxiv.org/html/2609.38393#S1.p2.1)\.
- Zhaiet al\.\(2023\)X\. Zhai, B\. Mustafa, A\. Kolesnikov, and L\. BeyerSigmoid loss for language image pre\-training\.InProceedings of the IEEE/CVF International Conference on Computer Vision \(ICCV\),pp\. 11975–11986\.Cited by:[§1](https://arxiv.org/html/2609.38393#S1.p1.1)\.
- Zimmermannet al\.\(2021\)R\. S\. Zimmermann, Y\. Sharma, S\. Schneider, M\. Bethge, and W\. BrendelContrastive learning inverts the data generating process\.InInternational Conference on Machine Learning,pp\. 12979–12990\.Cited by:[§1\.1](https://arxiv.org/html/2609.38393#S1.SS1.p4.1)\.

## Appendix AImplementation Details

Datasets\.CelebA\([Liu et al\., 2015](https://arxiv.org/html/2609.38393#bib.bib2)\)contains 202,599 RGB face images at resolution178×218178\\times 218, covering 10,177 identities and 40 binary attributes\. We use the standard identity\-disjoint split of 162,770 training, 19,867 validation, and 19,962 test images\.

The dSprites dataset contains 737,280 binary64×6464\\times 64images generated from six latent factors\. We retain two of the three shapes, the four most extreme scale levels, and the eight most extreme values of each position coordinate, leaving orientation unrestricted\. This gives 20,480 images, split randomly into 18,432 training and 2,048 held\-out test images\. Binary downstream tasks are formed by thresholding selected generative factors\.

3DShapes contains 480,000 RGB64×6464\\times 64images generated from six latent factors\. We keep six of the ten object\-hue levels and the four most extreme scale levels, giving 144,000 images\. We use 129,600 images for training and 14,400 for held\-out testing\. As with dSprites, binary downstream tasks are obtained by thresholding selected latent factors; the factors used in each analysis are indicated in the corresponding figure\.

Evaluation protocol\.The held\-out test split is used for the saturation curves and the centroid\-geometry evaluations\. For the recoverability and directional\-CDNV analyses, we construct disjoint subsets A and B from the training data\. The subsets contain 81,385 images each for CelebA, 2,048 images each for dSprites, and 14,400 images each for 3DShapes\. The paired\-view spectral basis, whitening transform, and recoverability estimates are computed on A; directional CDNV is then measured independently on B\.

Evaluation preprocessing\.All label\-dependent quantities \(recoverability, directional CDNV, NCC error, and linear\-probe quantities\) are measured on clean, single\-view images\. CelebA images are center\-cropped to160×160160\\times 160and resized to128×128128\\times 128for the models trained from scratch\. dSprites and 3DShapes are kept at their native64×6464\\times 64resolution and are bilinearly upsampled to the input resolution of externally pretrained backbones when needed\. Inputs are normalized with ImageNet\-1K channel statistics\. On CelebA, augmented pairs are used only to estimate the paired\-view spectral basis\. On dSprites and 3DShapes, the same deterministic resize\-and\-normalize transform is used both for paired\-view spectral estimation and for clean feature extraction\.

VICReg and W\-MSE augmentations\.We use the same view distribution for the CelebA\-trained VICReg and W\-MSE models\. Each view is generated by a random resized crop with scale\(0\.2,1\.0\)\(0\.2,1\.0\)and aspect ratio\(3/4,4/3\)\(3/4,4/3\), resized to128×128128\\times 128, followed by horizontal flipping with probability0\.50\.5\. We apply color jitter with brightness0\.40\.4, contrast0\.40\.4, saturation0\.20\.2, and hue0\.050\.05with probability0\.80\.8, grayscale conversion with probability0\.10\.1, and Gaussian blur with probability0\.50\.5using a7×77\\times 7kernel andσ∼𝒰⁡\[0\.1,2\.0\]\\sigma\\sim\\mathcal\{U\}\[0\.1,2\.0\]\. We use a minimum crop scale of0\.20\.2, rather than the ImageNet value0\.080\.08, to preserve more facial structure\. Unlike the original VICReg pipeline\([Bardes et al\., 2022](https://arxiv.org/html/2609.38393#bib.bib5)\), we do not use solarization and use the same blur probability for both views\.

Pretrained encoders\.For ImageNet\-1K\-pretrained VICReg, Barlow Twins, and I\-JEPA, CelebA images are resized to a shorter side of 256 pixels and center\-cropped to224×224224\\times 224\. To estimate the paired\-view spectral basis, we use an ImageNet\-scale version of the augmentation pipeline above: random resized crops with scale\(0\.08,1\.0\)\(0\.08,1\.0\), color\-jitter hue0\.10\.1, grayscale probability0\.20\.2, and a23×2323\\times 23Gaussian blur kernel\. For I\-JEPA, these augmented pairs define a surrogate two\-view kernel for analysis; they are not the masking\-based views used during I\-JEPA pretraining\. For dSprites and 3DShapes, externally pretrained encoders receive bilinearly upsampled224×224224\\times 224inputs\.

Training details\.We train VICReg\([Bardes et al\., 2022](https://arxiv.org/html/2609.38393#bib.bib5)\)and W\-MSE\([Ermolov et al\., 2021](https://arxiv.org/html/2609.38393#bib.bib75)\)from scratch on the CelebA training split for 1000 epochs\. Both use AdamW\([Loshchilov and Hutter, 2019](https://arxiv.org/html/2609.38393#bib.bib69)\)withβ1=0\.9\\beta\_\{1\}=0\.9,β2=0\.999\\beta\_\{2\}=0\.999,ϵ=10−8\\epsilon=10^\{\-8\}, and weight decay10−610^\{\-6\}\. The learning rate is scaled linearly with global batch size,η=ηbase​B/256\\eta=\\eta\_\{\\mathrm\{base\}\}B/256, with a 10\-epoch linear warm\-up\([Goyal et al\., 2017](https://arxiv.org/html/2609.38393#bib.bib62)\)followed by cosine decay\([Loshchilov and Hutter, 2017](https://arxiv.org/html/2609.38393#bib.bib61)\)\.

For VICReg, the projection head isLinear\(2048→\\rightarrow2048, bias=False\)→\\rightarrowBN→\\rightarrowReLU→\\rightarrowLinear\(2048→\\rightarrow2048\)\. The invariance, variance, and covariance coefficients areλ=25\\lambda=25,μ=25\\mu=25, andν=1\\nu=1\. We use a global batch size of 1024 across two NVIDIA H100 GPUs andηbase=1\.5×10−4\\eta\_\{\\mathrm\{base\}\}=1\.5\\times 10^\{\-4\}\.

For W\-MSE, the projection head isLinear\(2048→\\rightarrow2048, bias=False\)→\\rightarrowBN→\\rightarrowReLU→\\rightarrowLinear\(2048→\\rightarrow128\)\. Following\([Ermolov et al\., 2021](https://arxiv.org/html/2609.38393#bib.bib75)\), embeddings from each view are split into sub\-batches of 256 and whitened independently using the inverse Cholesky factor of the empirical covariance, with no covariance regularization \(ϵ=0\\epsilon=0\)\. The loss is normalized mean\-squared error between corresponding whitened embeddings\. The two views are processed in one concatenated forward pass\. We use a global batch size of 512 on one NVIDIA H100 andηbase=10−3\\eta\_\{\\mathrm\{base\}\}=10^\{\-3\}\.

Architectures\.Both CelebA\-trained models use a standard ResNet\-50\([He et al\., 2016](https://arxiv.org/html/2609.38393#bib.bib21)\)\. At128×128128\\times 128resolution, the final residual stage produces a4×44\\times 4feature map, which is globally averaged to a 2048\-dimensional representation\. Unless stated otherwise, representation analyses use these pooled backbone features rather than projection\-head outputs\.

We also evaluate public ImageNet\-1K\-pretrained VICReg and Barlow Twins\([Zbontar et al\., 2021](https://arxiv.org/html/2609.38393#bib.bib53)\)ResNet\-50 encoders and an I\-JEPA\([Assran et al\., 2023](https://arxiv.org/html/2609.38393#bib.bib50)\)ViT\-H/14 encoder, without fine\-tuning\. The ResNet\-50 encoders produce 2048\-dimensional pooled features\. I\-JEPA has 32 transformer layers, 16 attention heads, hidden dimension 1280, and MLP dimension 5120\. At224×224224\\times 224, it produces 256 patch tokens and no classification token; we average the final\-layer\-normalized patch tokens to obtain a 1280\-dimensional representation\.

## Appendix BAdditional Results

### B\.1Additional CelebA Results

This section reports the remaining CelebA results using the same protocol as Sec\.[5](https://arxiv.org/html/2609.38393#S5)\.

Semantic recoverability\.Fig\.[8](https://arxiv.org/html/2609.38393#A2.F8)\(a\) shows the rank\-dependent recoverability curves for ImageNet\-1K\-pretrained VICReg\. It reproduces the main\-text pattern: recoverability varies substantially across attributes, the strongest tasks saturate at lower rank, and the learned representations remain above their random\-initialization baselines\.

\(a\) Captured posterior energy\(b\) Predicted vs\. observedV~\\widetilde\{V\}\(c\) Spectral identificationFigure 8:Additional recoverability results for VICReg\-IM1K on CelebA\.\(a\) Captured posterior energy across representation ranks, using the same attribute groups and random\-initialization reference as Fig\.[3](https://arxiv.org/html/2609.38393#S5.F3)\. \(b\) Directional CDNV predicted from recoverability estimated on split A versus directional CDNV measured independently on held\-out split B; the dashed line isy=xy=x\. \(c\) Recoverability measured directly from the learned representation versus its reconstruction from the paired\-view spectral subspace; the dashed line denotes perfect agreement\.Directional geometry\.Fig\.[8](https://arxiv.org/html/2609.38393#A2.F8)\(b\) gives the held\-out test of Prop\.[3\.1](https://arxiv.org/html/2609.38393#S3.Thmtheorem1)\. Predictions are computed from split A and observed directional CDNV from split B\. The agreement withy=xy=xremains tight across attributes and ranks\.

Few\-shot transfer\.Fig\.[9](https://arxiv.org/html/2609.38393#A2.F9)reports the remaining few\-shot results\. Across models, ranks, and shot counts, empirical NCC error decreases with recoverability and the bound tracks the same rank–shot tradeoff seen in the main text\.

\(a\) VICReg\-CelebA\(b\) VICReg\-IM1K\(c\) Barlow\-IM1KFigure 9:Captured posterior energy and few\-shot NCC error\.Results for VICReg\-CelebA, VICReg\-IM1K, and Barlow Twins\-IM1K across ranksrrand support sizesmm\. Dashed curves are the upper bound from Thm\.[3\.4](https://arxiv.org/html/2609.38393#S3.Thmtheorem4)\.Multitask geometry and spectral identification\.Each encoder is evaluated at a label\-free rankr⋆r^\{\\star\}, defined as the number of paired\-view spectral directions whose eigenvalue is at least10−310^\{\-3\}times the largest eigenvalue\. This givesr⋆=946r^\{\\star\}=946for W\-MSE, 2046 for VICReg\-CelebA, 763 for VICReg\-IM1K, 806 for Barlow\-IM1K, and 263 for I\-JEPA\-IM1K\.

For each attribute triplet, the task\-specific capture valuesBr⋆\(t\)B\_\{r^\{\\star\}\}^\{\(t\)\}and the task axes are estimated on the training split, weighting the eight joint\-label cells equally; the centroids are then measured on the held\-out test split\. We writeB¯r⋆:=13​∑t=13Br⋆\(t\)\\overline\{B\}\_\{r^\{\\star\}\}:=\\frac\{1\}\{3\}\\sum\_\{t=1\}^\{3\}B\_\{r^\{\\star\}\}^\{\(t\)\}for the triplet mean recoverability used in Fig\.[6](https://arxiv.org/html/2609.38393#S5.F6)\(c–d\)\. Normalized centroid RMSE is the root\-mean\-square distance between observed and predicted centroids divided by the root\-mean\-square radius of the predicted box\. Fig\.[6](https://arxiv.org/html/2609.38393#S5.F6)\(c–d\) use all triplets drawn from the 26 attributes for which each binary class contains at least 10% of the data and every joint\-label cell contains at least 1000 training and 100 test examples\. This yields 1646 triplets per encoder\. No recoverability or orthogonality screening is applied to these quantitative panels\. The curves are means in bins of width 0\.05 containing at least 12 triplets; shaded regions show interquartile ranges\. The illustrative cubes in Figs\.[6](https://arxiv.org/html/2609.38393#S5.F6)\(a–b\) and[10](https://arxiv.org/html/2609.38393#A2.F10)\(a–b\) are selected on the training split only from triplets satisfyingmint⁡Br⋆\(t\)≥0\.10\\min\_\{t\}B\_\{r^\{\\star\}\}^\{\(t\)\}\\geq 0\.10and pairwise\|cos⁡\(ui,uj\)\|≤0\.12\|\\cos\(u\_\{i\},u\_\{j\}\)\|\\leq 0\.12\.

Fig\.[10](https://arxiv.org/html/2609.38393#A2.F10)adds held\-out centroid visualizations for VICReg\-IM1K and Barlow Twins\-IM1K, and Fig\.[8](https://arxiv.org/html/2609.38393#A2.F8)\(c\) and Fig\.[7](https://arxiv.org/html/2609.38393#S5.F7)\(c\) reports the corresponding spectral\-identification plots for VICReg\-IM1K and Barlow Twins\-IM1K respectively\. The latter compare recoverability measured in the full feature space with recoverability reconstructed from the label\-free paired\-view spectral subspace\.

\(a\) VICReg\-IM1K\(b\) Barlow\-IM1KFigure 10:Multitask centroid geometry on CelebA\.Held\-out joint centroids and the hyperrectangles predicted by Thm\.[3\.5](https://arxiv.org/html/2609.38393#S3.Thmtheorem5)for ImageNet\-1K\-pretrained VICReg and Barlow Twins\. Experiment setup is same as Fig\.[6](https://arxiv.org/html/2609.38393#S5.F6)Comparison with prior transfer bounds\.Table[1](https://arxiv.org/html/2609.38393#A2.T1)compares Thm\.[3\.4](https://arxiv.org/html/2609.38393#S3.Thmtheorem4)with earlier neural\-collapse\-based transfer bounds\([Galanti et al\., 2023](https://arxiv.org/html/2609.38393#bib.bib14);[Luthra et al\., 2025](https://arxiv.org/html/2609.38393#bib.bib63);[Luthra et al\., 2026](https://arxiv.org/html/2609.38393#bib.bib1)\)\. On W\-MSE trained on CelebA, the earlier bounds remain above0\.50\.5for all 40 attributes and all five shot counts shown\. Our bound is non\-vacuous in several rank–shot configurations and falls below0\.50\.5already atm=1m=1\. The minimizing rank increases as more support examples become available, matching the rank–shot tradeoff in Thm\.[3\.4](https://arxiv.org/html/2609.38393#S3.Thmtheorem4)\.

Table 1:Comparison of few\-shot transfer bounds on CelebA\.Bounds on NCC classification error for W\-MSE trained on CelebA \(“Male/Female” attribute\) \. Our bound is minimized over the evaluated ranks separately for each shot countmm\. Prior bounds useℓ2\\ell\_\{2\}\-normalized full\-dimensional features, whereas ours uses whitened top\-rrfeatures\. Values below0\.50\.5are non\-vacuous relative to chance classification\.
### B\.2Results on Synthetic Datasets

We use dSprites and 3DShapes to check that the same recoverability–geometry relationships are visible in controlled factorized data\. Figs\.[11](https://arxiv.org/html/2609.38393#A2.F11)–[13](https://arxiv.org/html/2609.38393#A2.F13)report rank\-dependent recoverability, the directional\-CDNV identity, and held\-out multitask centroid geometry\. The displayed encoders are ImageNet\-1K\-pretrained VICReg, I\-JEPA, and Barlow Twins\. Across both datasets, the qualitative behavior matches the CelebA results\.

\(a\) VICReg \(dSprites\)\(b\) I\-JEPA \(dSprites\)\(c\) Barlow Twins \(dSprites\)\(d\) VICReg \(3DShapes\)\(e\) I\-JEPA \(3DShapes\)\(f\) Barlow Twins \(3DShapes\)Figure 11:Semantic recoverability on dSprites and 3DShapes\.Task\-specific captured posterior energyBr\(t\)B\_\{r\}^\{\(t\)\}across representation ranks for ImageNet\-1K\-pretrained VICReg, I\-JEPA, and Barlow Twins\.\(a\) VICReg \(dSprites\)\(b\) I\-JEPA \(dSprites\)\(c\) Barlow Twins \(dSprites\)\(d\) VICReg \(3DShapes\)\(e\) I\-JEPA \(3DShapes\)\(f\) Barlow Twins \(3DShapes\)Figure 12:Predicted versus observed directional CDNV on synthetic datasets\.Results for VICReg, I\-JEPA, and Barlow Twins on dSprites and 3DShapes\. Each point is one attribute at one rankrr; the prediction is computed on split A and the observed directional CDNV on held\-out split B\. Dashed lines show the theoretical predictiony=xy=x\.\(a\) VICReg \(dSprites\)\(b\) I\-JEPA \(dSprites\)\(c\) Barlow Twins \(dSprites\)\(d\) VICReg \(3DShapes\)\(e\) I\-JEPA \(3DShapes\)\(f\) Barlow Twins \(3DShapes\)Figure 13:Multitask centroid geometry on synthetic datasets\.Held\-out joint centroids compared with the hyperrectangles predicted by Thm\.[3\.5](https://arxiv.org/html/2609.38393#S3.Thmtheorem5)\.

## Appendix CCovariance\-Regularized SSL Objectives

In the main text we studied the population surrogate \([2](https://arxiv.org/html/2609.38393#S4.E2)\), which enforces strict whitening𝔼⁡\[F⁡\(X\)​F​\(X\)⊤\]=Ir\\mathbb\{E\}\[F\(X\)F\(X\)^\{\\top\}\]=I\_\{r\}\. This yields an exact strict\-whitening theory in which the main directional geometry is governed by the single semantic capture parameterBr=‖Πr​η‖L2​\(PX\)2B\_\{r\}=\\\|\\Pi\_\{r\}\\eta\\\|\_\{L^\{2\}\(P\_\{X\}\)\}^\{2\}\. Modern SSL objectives typically do not enforce exact whitening, but instead penalize covariance deviations through variance/covariance regularizers, as in Barlow Twins and VICReg\. In this appendix we show that a natural covariance\-regularized surrogate preserves the same spectral subspace identified by the two\-view operatorTTin Prop\.[4\.1](https://arxiv.org/html/2609.38393#S4.Thmtheorem1), but the resulting geometry is no longer governed by a single scalar\. Instead, three explicit quantities\(Bγ,Cγ,Sγ\)\(B\_\{\\gamma\},C\_\{\\gamma\},S\_\{\\gamma\}\)appear, with the strict\-whitening formulas recovered asγ→∞\\gamma\\to\\infty\.

### C\.1Population covariance\-regularized objective

We consider the following population surrogate:

maxF:𝒳→ℝr𝔼\[⟨F\(X\(1\)\),F\(X\(2\)\)⟩\]−γ2‖𝔼\[F\(X\)F\(X\)⊤\]−Ir‖F2s\.t\.𝔼\[F\(X\)\]=0,\\max\_\{F:\\mathcal\{X\}\\to\\mathbb\{R\}^\{r\}\}\\ \\mathbb\{E\}\\\!\\left\[\\langle F\(X^\{\(1\)\}\),F\(X^\{\(2\)\}\)\\rangle\\right\]\-\\frac\{\\gamma\}\{2\}\\left\\\|\\mathbb\{E\}\[F\(X\)F\(X\)^\{\\top\}\]\-I\_\{r\}\\right\\\|\_\{F\}^\{2\}\\quad\\text\{s\.t\.\}\\quad\\mathbb\{E\}\[F\(X\)\]=0,\(3\)whereγ\>0\\gamma\>0controls the strength of the covariance penalty\. Asγ→∞\\gamma\\to\\infty, the penalty enforces𝔼⁡\[F⁡\(X\)​F​\(X\)⊤\]→Ir\\mathbb\{E\}\[F\(X\)F\(X\)^\{\\top\}\]\\to I\_\{r\}, recovering the strict whitening setting\. The centering constraint𝔼⁡\[F⁡\(X\)\]=0\\mathbb\{E\}\[F\(X\)\]=0models the fact that practical implementations explicitly center features and, in a population analysis, can be imposed without loss of generality by working with centered featuresF¯​\(X\):=F⁡\(X\)−𝔼⁡\[F⁡\(X\)\]\\bar\{F\}\(X\):=F\(X\)\-\\mathbb\{E\}\[F\(X\)\]and rewriting the objective in terms ofF¯\\bar\{F\}\.

### C\.2Operator form and canonical modes

Recall the two\-view conditional expectation operator\(T​f\)​\(x\):=𝔼⁡\[f⁡\(X\(2\)\)∣X\(1\)=x\]\(Tf\)\(x\):=\\mathbb\{E\}\[f\(X^\{\(2\)\}\)\\mid X^\{\(1\)\}=x\],T:L2​\(PX\)→L2​\(PX\)T:L^\{2\}\(P\_\{X\}\)\\to L^\{2\}\(P\_\{X\}\)\. For anyf,g∈L2​\(PX\)f,g\\in L^\{2\}\(P\_\{X\}\), we have

𝔼⁡\[f⁡\(X\(1\)\)​g​\(X\(2\)\)\]=⟨f,T​g⟩L2​\(PX\)\.\\mathbb\{E\}\\big\[f\(X^\{\(1\)\}\)g\(X^\{\(2\)\}\)\\big\]=\\langle f,Tg\\rangle\_\{L^\{2\}\(P\_\{X\}\)\}\.\(4\)In particular,

𝔼⁡\[f⁡\(X\(1\)\)​f​\(X\(2\)\)\]=⟨f,T​f⟩L2​\(PX\)\.\\mathbb\{E\}\\big\[f\(X^\{\(1\)\}\)f\(X^\{\(2\)\}\)\\big\]=\\langle f,Tf\\rangle\_\{L^\{2\}\(P\_\{X\}\)\}\.
Let\(λj,ψj\)j≥1\(\\lambda\_\{j\},\\psi\_\{j\}\)\_\{j\\geq 1\}be the orthonormal eigenbasis ofTTonL02​\(PX\)L\_\{0\}^\{2\}\(P\_\{X\}\)from Prop\.[4\.1](https://arxiv.org/html/2609.38393#S4.Thmtheorem1), ordered as

λ1≥λ2≥⋯≥0\.\\lambda\_\{1\}\\geq\\lambda\_\{2\}\\geq\\cdots\\geq 0\.Thus

T​ψj=λj​ψj,⟨ψi,ψj⟩L2​\(PX\)=δi​j,𝔼⁡\[ψj​\(X\)\]=0\.T\\psi\_\{j\}=\\lambda\_\{j\}\\psi\_\{j\},\\qquad\\langle\\psi\_\{i\},\\psi\_\{j\}\\rangle\_\{L^\{2\}\(P\_\{X\}\)\}=\\delta\_\{ij\},\\qquad\\mathbb\{E\}\[\\psi\_\{j\}\(X\)\]=0\.These are exactly the canonical observable modes identified in Prop\.[4\.1](https://arxiv.org/html/2609.38393#S4.Thmtheorem1)\. This appendix keeps the same operator and asks how covariance regularization rescales its top\-rrTT\-eigenspace\.

### C\.3Structure of the population optimum

We show that covariance regularization preserves theTT\-selected spectral subspace and only introduces mode\-dependent scalings within that subspace\.

###### Proposition C\.1\(Structure of the regularized optimum\)\.

Assumeλr\>0\\lambda\_\{r\}\>0andλr\>λr\+1\\lambda\_\{r\}\>\\lambda\_\{r\+1\}\. Then any population maximizerFγ⋆F^\{\\star\}\_\{\\gamma\}of \([3](https://arxiv.org/html/2609.38393#A3.E3)\) can be written as

Fγ⋆​\(x\)=R⁡\(s1​ψ1​\(x\),…,sr​ψr​\(x\)\),F^\{\\star\}\_\{\\gamma\}\(x\)=R\\,\\big\(s\_\{1\}\\psi\_\{1\}\(x\),\\dots,s\_\{r\}\\psi\_\{r\}\(x\)\\big\),\(5\)whereR∈ℝr×rR\\in\\mathbb\{R\}^\{r\\times r\}is orthogonal ands1,…,sr≥0s\_\{1\},\\dots,s\_\{r\}\\geq 0are scalars\. Moreover, for the quadratic penalty in \([3](https://arxiv.org/html/2609.38393#A3.E3)\), the optimal scalings satisfy

sj2=1\+λjγ,j=1,…,r\.s\_\{j\}^\{2\}=1\+\\frac\{\\lambda\_\{j\}\}\{\\gamma\},\\qquad j=1,\\dots,r\.

Prop\.[C\.1](https://arxiv.org/html/2609.38393#A3.Thmtheorem1)implies that the regularized optimum can be reduced to the whitened setting by an explicit population rewhitening step, and therefore inherits the same multitask geometric structure established in the main text\. LetFγ⋆F^\{\\star\}\_\{\\gamma\}be a population maximizer of \([3](https://arxiv.org/html/2609.38393#A3.E3)\), and define

Σγ:=Cov\(Fγ⋆\(X\)\),F~γ\(x\):=Σγ−1/2Fγ⋆\(x\)\.\\Sigma\_\{\\gamma\}:=\\mathrm\{Cov\}\(F^\{\\star\}\_\{\\gamma\}\(X\)\),\\qquad\\widetilde\{F\}\_\{\\gamma\}\(x\):=\\Sigma\_\{\\gamma\}^\{\-1/2\}F^\{\\star\}\_\{\\gamma\}\(x\)\.For the maximizer in Prop\.[C\.1](https://arxiv.org/html/2609.38393#A3.Thmtheorem1), we havesj2=1\+λj/γ\>0s\_\{j\}^\{2\}=1\+\\lambda\_\{j\}/\\gamma\>0, henceΣγ≻0\\Sigma\_\{\\gamma\}\\succ 0andΣγ−1/2\\Sigma\_\{\\gamma\}^\{\-1/2\}is well\-defined\. ThenF~γ\\widetilde\{F\}\_\{\\gamma\}is feasible for the strictly whitened objective \([2](https://arxiv.org/html/2609.38393#S4.E2)\) and is population\-optimal\. Indeed, by Prop\.[C\.1](https://arxiv.org/html/2609.38393#A3.Thmtheorem1)we may write

Fγ⋆​\(x\)=R​diag​\(s1,…,sr\)​\(ψ1​\(x\),…,ψr​\(x\)\),F^\{\\star\}\_\{\\gamma\}\(x\)=R\\,\\mathrm\{diag\}\(s\_\{1\},\\dots,s\_\{r\}\)\\,\\big\(\\psi\_\{1\}\(x\),\\dots,\\psi\_\{r\}\(x\)\\big\),so

Σγ=Rdiag\(s12,…,sr2\)R⊤,Σγ−1/2=Rdiag\(s1−1,…,sr−1\)R⊤,\\Sigma\_\{\\gamma\}=R\\,\\mathrm\{diag\}\(s\_\{1\}^\{2\},\\dots,s\_\{r\}^\{2\}\)\\,R^\{\\top\},\\qquad\\Sigma\_\{\\gamma\}^\{\-1/2\}=R\\,\\mathrm\{diag\}\(s\_\{1\}^\{\-1\},\\dots,s\_\{r\}^\{\-1\}\)\\,R^\{\\top\},and therefore

F~γ​\(x\)=R⁡\(ψ1​\(x\),…,ψr​\(x\)\)\.\\widetilde\{F\}\_\{\\gamma\}\(x\)=R\\,\\big\(\\psi\_\{1\}\(x\),\\dots,\\psi\_\{r\}\(x\)\\big\)\.By Prop\.[4\.1](https://arxiv.org/html/2609.38393#S4.Thmtheorem1), any representation of the formR⁡\(ψ1,…,ψr\)R\(\\psi\_\{1\},\\dots,\\psi\_\{r\}\)is feasible for \([2](https://arxiv.org/html/2609.38393#S4.E2)\) and achieves the optimal value∑j=1rλj\\sum\_\{j=1\}^\{r\}\\lambda\_\{j\}, henceF~γ\\widetilde\{F\}\_\{\\gamma\}is a whitened population optimum\.

Consequently, all main\-text results that rely only on the whitened\-optimum structure, including the multitask near\-orthogonality and factorial\-centroid geometry in Thm\.[3\.5](https://arxiv.org/html/2609.38393#S3.Thmtheorem5), apply verbatim toF~γ\\widetilde\{F\}\_\{\\gamma\}\. In the original feature coordinates ofFγ⋆F^\{\\star\}\_\{\\gamma\}, this same structure is obtained by applying the linear mapΣγ1/2\\Sigma\_\{\\gamma\}^\{1/2\}, so the hyperrectangle in rewhitened probe coordinates becomes a parallelepiped in the rawFγ⋆F^\{\\star\}\_\{\\gamma\}coordinates\. This highlights that covariance regularization preserves the operator\-selected semantic organization, but can distort Euclidean geometry unless one explicitly rewhitens features, or equivalently uses a Mahalanobis nearest\-centroid rule with metricΣγ−1\\Sigma\_\{\\gamma\}^\{\-1\}\.

### C\.4How the CDNV formulas change

LetY=y⁡\(C\)∈\{±1\}Y=y\(C\)\\in\\\{\\pm 1\\\}be balanced and define the posterior score

η⁡\(x\)=𝔼⁡\[Y∣X=x\]\.\\eta\(x\)=\\mathbb\{E\}\[Y\\mid X=x\]\.Write the posterior directly in theTT\-eigenbasis:

η=∑j≥1βj​ψj,βj:=⟨η,ψj⟩L2​\(PX\)\.\\eta=\\sum\_\{j\\geq 1\}\\beta\_\{j\}\\psi\_\{j\},\\qquad\\beta\_\{j\}:=\\langle\\eta,\\psi\_\{j\}\\rangle\_\{L^\{2\}\(P\_\{X\}\)\}\.Then, by Cor\.[4\.2](https://arxiv.org/html/2609.38393#S4.Thmtheorem2),

Πr​η=∑j=1rβj​ψj,Br=‖Πr​η‖L2​\(PX\)2=∑j=1rβj2\.\\Pi\_\{r\}\\eta=\\sum\_\{j=1\}^\{r\}\\beta\_\{j\}\\psi\_\{j\},\\qquad B\_\{r\}=\\\|\\Pi\_\{r\}\\eta\\\|\_\{L^\{2\}\(P\_\{X\}\)\}^\{2\}=\\sum\_\{j=1\}^\{r\}\\beta\_\{j\}^\{2\}\.
Define the scalars

Bγ\\displaystyle B\_\{\\gamma\}:=∑j=1rβj2​sj2,\\displaystyle:=\\sum\_\{j=1\}^\{r\}\\beta\_\{j\}^\{2\}s\_\{j\}^\{2\},Cγ\\displaystyle C\_\{\\gamma\}:=∑j=1rβj2​sj4,\\displaystyle:=\\sum\_\{j=1\}^\{r\}\\beta\_\{j\}^\{2\}s\_\{j\}^\{4\},Sγ\\displaystyle S\_\{\\gamma\}:=∑j=1rsj2\.\\displaystyle:=\\sum\_\{j=1\}^\{r\}s\_\{j\}^\{2\}\.\(6\)
###### Proposition C\.2\.

LetFγ⋆F^\{\\star\}\_\{\\gamma\}be any population maximizer of \([3](https://arxiv.org/html/2609.38393#A3.E3)\), written as in \([5](https://arxiv.org/html/2609.38393#A3.E5)\)\. Letμ±\\mu\_\{\\pm\}andΣ±\\Sigma\_\{\\pm\}be the class means and class covariances defined as in the main text\. Thenμ−=−μ\+\\mu\_\{\-\}=\-\\mu\_\{\+\}andΔ=μ\+−μ−=2​μ\+\\Delta=\\mu\_\{\+\}\-\\mu\_\{\-\}=2\\mu\_\{\+\}\. Moreover, whenBγ\>0B\_\{\\gamma\}\>0, the CDNV and directional CDNV defined in \([1](https://arxiv.org/html/2609.38393#S3.E1)\) satisfy

V~Fγ⋆=12​\(CγBγ2−1\),VFγ⋆=12​\(SγBγ−1\)\.\\tilde\{V\}\_\{F^\{\\star\}\_\{\\gamma\}\}=\\frac\{1\}\{2\}\\left\(\\frac\{C\_\{\\gamma\}\}\{B\_\{\\gamma\}^\{2\}\}\-1\\right\),\\qquad V\_\{F^\{\\star\}\_\{\\gamma\}\}=\\frac\{1\}\{2\}\\left\(\\frac\{S\_\{\\gamma\}\}\{B\_\{\\gamma\}\}\-1\\right\)\.\(7\)WhenBγ=0B\_\{\\gamma\}=0, we set both quantities to\+∞\+\\inftyby the same convention as in the main text\.

ForBγ\>0B\_\{\\gamma\}\>0, Prop\.[C\.2](https://arxiv.org/html/2609.38393#A3.Thmtheorem2)can be plugged directly into a general CDNV\-based few\-shot guarantee\. Namely, withF=Fγ⋆F=F^\{\\star\}\_\{\\gamma\}, substituting \([7](https://arxiv.org/html/2609.38393#A3.E7)\) yields the explicitγ\\gamma\-dependent bound

errmNCC​\(Fγ⋆\)≤4​\(CγBγ2−1\)\+8m​12​\(SγBγ−1\)\+\(8m\+4m\)​12​\(SγBγ−1\)\.\\mathrm\{err\}^\{\\mathrm\{NCC\}\}\_\{m\}\(F^\{\\star\}\_\{\\gamma\}\)\\leq 4\\Big\(\\frac\{C\_\{\\gamma\}\}\{B\_\{\\gamma\}^\{2\}\}\-1\\Big\)\+\\frac\{8\}\{\\sqrt\{m\}\}\\sqrt\{\\frac\{1\}\{2\}\\Big\(\\frac\{S\_\{\\gamma\}\}\{B\_\{\\gamma\}\}\-1\\Big\)\}\+\\Big\(\\frac\{8\}\{\\sqrt\{m\}\}\+\\frac\{4\}\{m\}\\Big\)\\frac\{1\}\{2\}\\Big\(\\frac\{S\_\{\\gamma\}\}\{B\_\{\\gamma\}\}\-1\\Big\)\.\(8\)In particular, this bound depends onγ\\gammaonly through the three scalarsBγ,Cγ,SγB\_\{\\gamma\},C\_\{\\gamma\},S\_\{\\gamma\}defined in \([6](https://arxiv.org/html/2609.38393#A3.E6)\)\.

Thusγ\\gammacontrols a concrete tradeoff in \([8](https://arxiv.org/html/2609.38393#A3.E8)\)\. The leading directional term depends onCγ/Bγ2C\_\{\\gamma\}/B\_\{\\gamma\}^\{2\}, while the finite\-mmterms depend onSγ/BγS\_\{\\gamma\}/B\_\{\\gamma\}\. Under the quadratic penaltysj2=1\+λj/γs\_\{j\}^\{2\}=1\+\\lambda\_\{j\}/\\gamma, decreasingγ\\gammainflates all selected modes, with larger inflation on more stable modes\. This increases total variance throughSγS\_\{\\gamma\}, but it also increases signal strength along the task\-relevant direction throughBγB\_\{\\gamma\}andCγC\_\{\\gamma\}\. The resulting few\-shot behavior is therefore governed by how the posterior coefficientsβj\\beta\_\{j\}overlap with the modes that are most amplified by the regularizer\.

In the strict\-whitening limitsj→1s\_\{j\}\\to 1, we recover

Bγ→Br,Cγ→Br,Sγ→r,B\_\{\\gamma\}\\to B\_\{r\},\\qquad C\_\{\\gamma\}\\to B\_\{r\},\\qquad S\_\{\\gamma\}\\to r,and therefore

V~Fγ⋆→1−Br2​Br,VFγ⋆→r−Br2​Br,\\tilde\{V\}\_\{F^\{\\star\}\_\{\\gamma\}\}\\to\\frac\{1\-B\_\{r\}\}\{2B\_\{r\}\},\\qquad V\_\{F^\{\\star\}\_\{\\gamma\}\}\\to\\frac\{r\-B\_\{r\}\}\{2B\_\{r\}\},exactly as in Prop\.[3\.1](https://arxiv.org/html/2609.38393#S3.Thmtheorem1)\.

## Appendix DPrediction\-Based SSL Objectives as Reduced\-Rank Regression

This appendix gives a simple population\-level connection between prediction\-based SSL objectives, such as BYOL, SimSiam, and JEPA, and classical reduced\-rank regression / partial least squares\. We do not aim to model algorithmic details such as stop\-gradient or EMA teachers\. Instead, we study a natural whitened population surrogate and show that its optimizer selects the sameTT\-spectral subspace as \([2](https://arxiv.org/html/2609.38393#S4.E2)\)\. As a result, the main\-text conclusions that depend only on the selected subspace, and therefore on the induced capture valueBrB\_\{r\}, carry over directly to this prediction\-based setting\.

### D\.1A stylized whitened prediction surrogate

LetF:𝒳→ℝrF:\\mathcal\{X\}\\to\\mathbb\{R\}^\{r\}be a shared representation used on both views, and letW∈ℝr×rW\\in\\mathbb\{R\}^\{r\\times r\}be a linear predictor\. Consider the population objective

minF:𝒳→ℝrW∈ℝr×r𝔼\[∥F\(X\(2\)\)−WF\(X\(1\)\)∥22\]s\.t\.𝔼\[F\(X\)\]=0,𝔼\[F\(X\)F\(X\)⊤\]=Ir,\\min\_\{\\begin\{subarray\}\{c\}F:\\mathcal\{X\}\\to\\mathbb\{R\}^\{r\}\\\\ W\\in\\mathbb\{R\}^\{r\\times r\}\\end\{subarray\}\}\\ \\mathbb\{E\}\\Big\[\\big\\\|F\(X^\{\(2\)\}\)\-WF\(X^\{\(1\)\}\)\\big\\\|\_\{2\}^\{2\}\\Big\]\\quad\\text\{s\.t\.\}\\quad\\mathbb\{E\}\[F\(X\)\]=0,\\ \\ \\mathbb\{E\}\[F\(X\)F\(X\)^\{\\top\}\]=I\_\{r\},\(9\)where\(X\(1\),X\(2\)\)\(X^\{\(1\)\},X^\{\(2\)\}\)is a same\-instance pair with conditionally i\.i\.d\. views as in Sec\.[2\.1](https://arxiv.org/html/2609.38393#S2.SS1)\.

Define

C12​\(F\):=𝔼⁡\[F⁡\(X\(1\)\)​F​\(X\(2\)\)⊤\],C21​\(F\):=C12​\(F\)⊤\.C\_\{12\}\(F\):=\\mathbb\{E\}\\big\[F\(X^\{\(1\)\}\)F\(X^\{\(2\)\}\)^\{\\top\}\\big\],\\qquad C\_\{21\}\(F\):=C\_\{12\}\(F\)^\{\\top\}\.\(10\)
###### Proposition D\.1\(RRR reduction to a spectral objective\)\.

Fix any feasibleFFin \([9](https://arxiv.org/html/2609.38393#A4.E9)\)\. The unique minimizer overWWis

W⋆​\(F\)=C21​\(F\)=C12​\(F\)⊤,W^\{\\star\}\(F\)=C\_\{21\}\(F\)=C\_\{12\}\(F\)^\{\\top\},and the resulting minimum equals

minW⁡𝔼⁡\[‖F⁡\(X\(2\)\)−W​F​\(X\(1\)\)‖22\]=r−‖C21​\(F\)‖F2\.\\min\_\{W\}\\ \\mathbb\{E\}\\Big\[\\big\\\|F\(X^\{\(2\)\}\)\-WF\(X^\{\(1\)\}\)\\big\\\|\_\{2\}^\{2\}\\Big\]=r\-\\\|C\_\{21\}\(F\)\\\|\_\{F\}^\{2\}\.\(11\)Hence minimizing \([9](https://arxiv.org/html/2609.38393#A4.E9)\) is equivalent to maximizing‖C21​\(F\)‖F2\\\|C\_\{21\}\(F\)\\\|\_\{F\}^\{2\}over feasibleFF\.

###### Proof\.

Using𝔼⁡\[F⁡\(X\)​F​\(X\)⊤\]=Ir\\mathbb\{E\}\[F\(X\)F\(X\)^\{\\top\}\]=I\_\{r\},

𝔼​‖F⁡\(X\(2\)\)−W​F​\(X\(1\)\)‖22=r−2​Tr​\(W​C12​\(F\)\)\+‖W‖F2=r\+‖W−C21​\(F\)‖F2−‖C21​\(F\)‖F2\.\\mathbb\{E\}\\\|F\(X^\{\(2\)\}\)\-WF\(X^\{\(1\)\}\)\\\|\_\{2\}^\{2\}=r\-2\\mathrm\{Tr\}\(WC\_\{12\}\(F\)\)\+\\\|W\\\|\_\{F\}^\{2\}=r\+\\\|W\-C\_\{21\}\(F\)\\\|\_\{F\}^\{2\}\-\\\|C\_\{21\}\(F\)\\\|\_\{F\}^\{2\}\.Thus the unique minimizer isW⋆​\(F\)=C21​\(F\)=C12​\(F\)⊤W^\{\\star\}\(F\)=C\_\{21\}\(F\)=C\_\{12\}\(F\)^\{\\top\}, and the minimum isr−‖C21​\(F\)‖F2r\-\\\|C\_\{21\}\(F\)\\\|\_\{F\}^\{2\}\. ∎

### D\.2Selected subspace and relationship to the alignment objective

For scalarf,g∈L2​\(PX\)f,g\\in L^\{2\}\(P\_\{X\}\)with𝔼⁡\[f⁡\(X\)\]=𝔼⁡\[g⁡\(X\)\]=0\\mathbb\{E\}\[f\(X\)\]=\\mathbb\{E\}\[g\(X\)\]=0, the same\-instance model gives

𝔼⁡\[f⁡\(X\(2\)\)​g​\(X\(1\)\)\]=⟨f,T​g⟩L2​\(PX\)\.\\mathbb\{E\}\[f\(X^\{\(2\)\}\)g\(X^\{\(1\)\}\)\]=\\langle f,Tg\\rangle\_\{L^\{2\}\(P\_\{X\}\)\}\.Thus, ifF=\(f1,…,fr\)F=\(f\_\{1\},\\dots,f\_\{r\}\)is feasible and the coordinates\{fk\}\\\{f\_\{k\}\\\}are orthonormal inL2​\(PX\)L^\{2\}\(P\_\{X\}\), then

\(C21​\(F\)\)i​j=𝔼⁡\[fi​\(X\(2\)\)​fj​\(X\(1\)\)\]=⟨fi,T​fj⟩L2​\(PX\)\.\\big\(C\_\{21\}\(F\)\\big\)\_\{ij\}=\\mathbb\{E\}\[f\_\{i\}\(X^\{\(2\)\}\)f\_\{j\}\(X^\{\(1\)\}\)\]=\\langle f\_\{i\},Tf\_\{j\}\\rangle\_\{L^\{2\}\(P\_\{X\}\)\}\.\(12\)Consequently,‖C21​\(F\)‖F2\\\|C\_\{21\}\(F\)\\\|\_\{F\}^\{2\}is the squared Hilbert–Schmidt norm of the compressionPU​T​PUP\_\{U\}TP\_\{U\}, whereU=span⁡\{f1,…,fr\}⊂L02​\(PX\)U=\\mathrm\{span\}\\\{f\_\{1\},\\dots,f\_\{r\}\\\}\\subset L\_\{0\}^\{2\}\(P\_\{X\}\)\.

Let\(λj,ψj\)j≥1\(\\lambda\_\{j\},\\psi\_\{j\}\)\_\{j\\geq 1\}be the orthonormal eigenbasis ofTTfrom Prop\.[4\.1](https://arxiv.org/html/2609.38393#S4.Thmtheorem1), withλ1≥λ2≥⋯≥0\\lambda\_\{1\}\\geq\\lambda\_\{2\}\\geq\\cdots\\geq 0\.

###### Proposition D\.2\(Prediction objective selects the top\-rroperator subspace\)\.

Assumeλr\>0\\lambda\_\{r\}\>0andλr\>λr\+1\\lambda\_\{r\}\>\\lambda\_\{r\+1\}\. Then the maximum of‖C21​\(F\)‖F2\\\|C\_\{21\}\(F\)\\\|\_\{F\}^\{2\}over feasibleFFequals∑j=1rλj2\\sum\_\{j=1\}^\{r\}\\lambda\_\{j\}^\{2\}, and any maximizerF⋆F^\{\\star\}has coordinates spanningspan⁡\{ψ1,…,ψr\}\\mathrm\{span\}\\\{\\psi\_\{1\},\\dots,\\psi\_\{r\}\\\}\. In particular, there exists an orthogonalR∈ℝr×rR\\in\\mathbb\{R\}^\{r\\times r\}such that

F⋆​\(x\)=R⁡\(ψ1​\(x\),…,ψr​\(x\)\)\.F^\{\\star\}\(x\)=R\\,\(\\psi\_\{1\}\(x\),\\dots,\\psi\_\{r\}\(x\)\)\.\(13\)

###### Proof\.

LetU=span⁡\{f1,…,fr\}U=\\mathrm\{span\}\\\{f\_\{1\},\\dots,f\_\{r\}\\\}and letPUP\_\{U\}be the orthogonal projector ontoUU\. In the orthonormal basis\{fj\}\\\{f\_\{j\}\\\},C21​\(F\)C\_\{21\}\(F\)is the matrix ofPU​T​PUP\_\{U\}TP\_\{U\}, and hence

‖C21​\(F\)‖F2=‖PU​T​PU‖HS2≤‖T​PU‖HS2=Tr⁡\(PU​T2\)≤∑j=1rλj2\.\\\|C\_\{21\}\(F\)\\\|\_\{F\}^\{2\}=\\\|P\_\{U\}TP\_\{U\}\\\|\_\{\\mathrm\{HS\}\}^\{2\}\\leq\\\|TP\_\{U\}\\\|\_\{\\mathrm\{HS\}\}^\{2\}=\\mathrm\{Tr\}\(P\_\{U\}T^\{2\}\)\\leq\\sum\_\{j=1\}^\{r\}\\lambda\_\{j\}^\{2\}\.The last inequality is Ky Fan’s variational principle applied to the positive semidefinite operatorT2T^\{2\}\. Equality is attained forU=span⁡\{ψ1,…,ψr\}U=\\mathrm\{span\}\\\{\\psi\_\{1\},\\dots,\\psi\_\{r\}\\\}, and the eigengapλr\>λr\+1\\lambda\_\{r\}\>\\lambda\_\{r\+1\}makes this maximizing subspace unique\. ∎

### D\.3Consequences for recoverability, CDNV, and downstream guarantees

Prop\.[D\.2](https://arxiv.org/html/2609.38393#A4.Thmtheorem2)shows that the whitened prediction surrogate \([9](https://arxiv.org/html/2609.38393#A4.E9)\) selects the samerr\-dimensionalTT\-eigenspace as \([2](https://arxiv.org/html/2609.38393#S4.E2)\)\. In particular, any maximizerF⋆F^\{\\star\}of \([9](https://arxiv.org/html/2609.38393#A4.E9)\) has coordinates spanningspan⁡\{ψ1,…,ψr\}\\mathrm\{span\}\\\{\\psi\_\{1\},\\dots,\\psi\_\{r\}\\\}, soF⋆F^\{\\star\}is feasible for \([2](https://arxiv.org/html/2609.38393#S4.E2)\)\. Since the strictly whitened alignment objective \([2](https://arxiv.org/html/2609.38393#S4.E2)\) depends only on the chosenrr\-dimensional subspace and is maximized by any orthonormal basis of the top\-rreigenspace, as shown in Prop\.[4\.1](https://arxiv.org/html/2609.38393#S4.Thmtheorem1),F⋆F^\{\\star\}is also population\-optimal for \([2](https://arxiv.org/html/2609.38393#S4.E2)\) and achieves the optimal value∑j=1rλj\\sum\_\{j=1\}^\{r\}\\lambda\_\{j\}\. Therefore, all main\-text results that apply to whitened population optima of \([2](https://arxiv.org/html/2609.38393#S4.E2)\) apply directly toF⋆F^\{\\star\}\.

In particular, for any downstream binary labelY=y⁡\(C\)Y=y\(C\)with posterior scoreη⁡\(x\)=𝔼⁡\[Y∣X=x\]\\eta\(x\)=\\mathbb\{E\}\[Y\\mid X=x\], writing

η=∑j≥1βj​ψj,\\eta=\\sum\_\{j\\geq 1\}\\beta\_\{j\}\\psi\_\{j\},the capture value remains

Br=∑j=1rβj2,B\_\{r\}=\\sum\_\{j=1\}^\{r\}\\beta\_\{j\}^\{2\},and the closed\-form directional CDNV and CDNV laws for whitened optima from Prop\.[3\.1](https://arxiv.org/html/2609.38393#S3.Thmtheorem1)hold unchanged forF⋆F^\{\\star\}\. Consequently, the same few\-shot NCC guarantees from Thm\.[3\.4](https://arxiv.org/html/2609.38393#S3.Thmtheorem4)and the same multitask geometry conclusions from Thm\.[3\.5](https://arxiv.org/html/2609.38393#S3.Thmtheorem5)follow immediately for this stylized prediction\-based surrogate\.

## Appendix EProofs

### E\.1Proofs of the Results in the Main Text

See[3\.1](https://arxiv.org/html/2609.38393#S3.Thmtheorem1)

###### Proof\.

Letμs:=𝔼⁡\[F⁡\(X\)∣Y=s\]\\mu\_\{s\}:=\\mathbb\{E\}\[F\(X\)\\mid Y=s\]andΣs:=Cov⁡\(F⁡\(X\)∣Y=s\)\\Sigma\_\{s\}:=\\mathrm\{Cov\}\(F\(X\)\\mid Y=s\)fors∈\{±1\}s\\in\\\{\\pm 1\\\}, and setΔ:=μ\+−μ−\\Delta:=\\mu\_\{\+\}\-\\mu\_\{\-\}withu:=Δ/‖Δ‖2u:=\\Delta/\\\|\\Delta\\\|\_\{2\}whenΔ≠0\\Delta\\neq 0\. Balancedness ofYYand𝔼⁡\[F⁡\(X\)\]=0\\mathbb\{E\}\[F\(X\)\]=0give12​μ\+\+12​μ−=0\\tfrac\{1\}\{2\}\\mu\_\{\+\}\+\\tfrac\{1\}\{2\}\\mu\_\{\-\}=0, soμ−=−μ\+\\mu\_\{\-\}=\-\\mu\_\{\+\}andΔ=2​μ\+\\Delta=2\\mu\_\{\+\}\. Moreover,𝔼⁡\[Y​F​\(X\)\]=12​μ\+−12​μ−=μ\+\\mathbb\{E\}\[YF\(X\)\]=\\tfrac\{1\}\{2\}\\mu\_\{\+\}\-\\tfrac\{1\}\{2\}\\mu\_\{\-\}=\\mu\_\{\+\}, and by iterated expectation this also equals𝔼⁡\[η⁡\(X\)​F​\(X\)\]\\mathbb\{E\}\[\\eta\(X\)F\(X\)\]\.

WriteF=\(f1,…,fr\)F=\(f\_\{1\},\\dots,f\_\{r\}\)\. Centering and whitening imply that\{f1,…,fr\}\\\{f\_\{1\},\\dots,f\_\{r\}\\\}is orthonormal inL2​\(PX\)L^\{2\}\(P\_\{X\}\), so𝒮F=span⁡\{f1,…,fr\}\\mathcal\{S\}\_\{F\}=\\mathrm\{span\}\\\{f\_\{1\},\\dots,f\_\{r\}\\\}is anrr\-dimensional subspace and\(μ\+\)i=𝔼⁡\[η⁡\(X\)​fi​\(X\)\]=⟨η,fi⟩\(\\mu\_\{\+\}\)\_\{i\}=\\mathbb\{E\}\[\\eta\(X\)f\_\{i\}\(X\)\]=\\langle\\eta,f\_\{i\}\\rangle\. HenceΠ𝒮F​η=∑i\(μ\+\)i​fi\\Pi\_\{\\mathcal\{S\}\_\{F\}\}\\eta=\\sum\_\{i\}\(\\mu\_\{\+\}\)\_\{i\}f\_\{i\}, and by orthonormality‖μ\+‖22=‖Π𝒮F​η‖L22=B⁡\(F\)\\\|\\mu\_\{\+\}\\\|\_\{2\}^\{2\}=\\\|\\Pi\_\{\\mathcal\{S\}\_\{F\}\}\\eta\\\|\_\{L^\{2\}\}^\{2\}=B\(F\), giving‖Δ‖22=4​B​\(F\)\\\|\\Delta\\\|\_\{2\}^\{2\}=4B\(F\)\. IfB⁡\(F\)=0B\(F\)=0thenΔ=0\\Delta=0and both ratios are\+∞\+\\infty; assumeB⁡\(F\)\>0B\(F\)\>0below\.

By the law of total covariance and whitening,Ir=12​\(Σ\+\+Σ−\)\+Cov⁡\(μY\)I\_\{r\}=\\tfrac\{1\}\{2\}\(\\Sigma\_\{\+\}\+\\Sigma\_\{\-\}\)\+\\mathrm\{Cov\}\(\\mu\_\{Y\}\), whereμY∈\{μ\+,μ−\}\\mu\_\{Y\}\\in\\\{\\mu\_\{\+\},\\mu\_\{\-\}\\\}uniformly\. Sinceμ−=−μ\+\\mu\_\{\-\}=\-\\mu\_\{\+\},Cov⁡\(μY\)=μ\+​μ\+⊤\\mathrm\{Cov\}\(\\mu\_\{Y\}\)=\\mu\_\{\+\}\\mu\_\{\+\}^\{\\top\}, yielding

Σ\+\+Σ−=2​\(Ir−μ\+​μ\+⊤\)\.\\Sigma\_\{\+\}\+\\Sigma\_\{\-\}=2\(I\_\{r\}\-\\mu\_\{\+\}\\mu\_\{\+\}^\{\\top\}\)\.\(14\)Taking traces in \([14](https://arxiv.org/html/2609.38393#A5.E14)\) givesTr⁡\(Σ\+\)\+Tr⁡\(Σ−\)=2​\(r−B⁡\(F\)\)\\mathrm\{Tr\}\(\\Sigma\_\{\+\}\)\+\\mathrm\{Tr\}\(\\Sigma\_\{\-\}\)=2\(r\-B\(F\)\), soVF=\(r−B⁡\(F\)\)/\(2​B​\(F\)\)V\_\{F\}=\(r\-B\(F\)\)/\(2B\(F\)\)after dividing by‖Δ‖22=4​B​\(F\)\\\|\\Delta\\\|\_\{2\}^\{2\}=4B\(F\)\. For the directional ratio,u=μ\+/‖μ\+‖2u=\\mu\_\{\+\}/\\\|\\mu\_\{\+\}\\\|\_\{2\}givesu⊤​\(Σ\+\+Σ−\)​u=2​\(1−‖μ\+‖22\)=2​\(1−B⁡\(F\)\)u^\{\\top\}\(\\Sigma\_\{\+\}\+\\Sigma\_\{\-\}\)u=2\(1\-\\\|\\mu\_\{\+\}\\\|\_\{2\}^\{2\}\)=2\(1\-B\(F\)\)by \([14](https://arxiv.org/html/2609.38393#A5.E14)\), soV~F=\(1−B⁡\(F\)\)/\(2​B​\(F\)\)\\tilde\{V\}\_\{F\}=\(1\-B\(F\)\)/\(2B\(F\)\)\. ∎

See[3\.3](https://arxiv.org/html/2609.38393#S3.Thmtheorem3)

###### Proof\.

LetZ:=F⁡\(X\)Z:=F\(X\), so that𝔼⁡\[Z\]=0\\mathbb\{E\}\[Z\]=0,𝔼⁡\[Z​Z⊤\]=Ir\\mathbb\{E\}\[ZZ^\{\\top\}\]=I\_\{r\}, and𝔼⁡\[Y\]=0\\mathbb\{E\}\[Y\]=0\. Expanding the squared loss and using these identities gives

ℛ⁡\(w,b\):=𝔼⁡\[\(Y−\(w⊤​Z\+b\)\)2\]=1−2​w⊤​𝔼​\[Y​Z\]\+‖w‖22\+b2,\\mathcal\{R\}\(w,b\):=\\mathbb\{E\}\\big\[\(Y\-\(w^\{\\top\}Z\+b\)\)^\{2\}\\big\]=1\-2w^\{\\top\}\\mathbb\{E\}\[YZ\]\+\\\|w\\\|\_\{2\}^\{2\}\+b^\{2\},a strictly convex quadratic in\(w,b\)\(w,b\)whose unique minimizer isw⋆=𝔼⁡\[Y​Z\]w^\{\\star\}=\\mathbb\{E\}\[YZ\]andb⋆=0b^\{\\star\}=0\. Letμs:=𝔼⁡\[Z∣Y=s\]\\mu\_\{s\}:=\\mathbb\{E\}\[Z\\mid Y=s\]\. Balancedness gives𝔼⁡\[Y​Z\]=12​\(μ\+−μ−\)=12​Δ\\mathbb\{E\}\[YZ\]=\\tfrac\{1\}\{2\}\(\\mu\_\{\+\}\-\\mu\_\{\-\}\)=\\tfrac\{1\}\{2\}\\Deltaand𝔼⁡\[Z\]=12​μ\+\+12​μ−=0\\mathbb\{E\}\[Z\]=\\tfrac\{1\}\{2\}\\mu\_\{\+\}\+\\tfrac\{1\}\{2\}\\mu\_\{\-\}=0, soμ−=−μ\+\\mu\_\{\-\}=\-\\mu\_\{\+\}and henceΔ=2​μ\+=2​w⋆\\Delta=2\\mu\_\{\+\}=2w^\{\\star\}\. ∎

See[3\.4](https://arxiv.org/html/2609.38393#S3.Thmtheorem4)

###### Proof\.

IfB⁡\(F\)=0B\(F\)=0, the stated bound is trivial since its right\-hand side is at least11\. Assume henceforth thatB⁡\(F\)\>0B\(F\)\>0\. LetZ:=F⁡\(X\)Z:=F\(X\)andw:=𝔼⁡\[Y​Z\]w:=\\mathbb\{E\}\[YZ\], so𝔼⁡\[Z\]=0\\mathbb\{E\}\[Z\]=0,𝔼⁡\[Z​Z⊤\]=Ir\\mathbb\{E\}\[ZZ^\{\\top\}\]=I\_\{r\}, and‖w‖22=B⁡\(F\)\\\|w\\\|\_\{2\}^\{2\}=B\(F\)\. Balancedness and𝔼⁡\[Z\]=0\\mathbb\{E\}\[Z\]=0giveμ±:=𝔼⁡\[Z∣Y=±1\]=±w\\mu\_\{\\pm\}:=\\mathbb\{E\}\[Z\\mid Y=\\pm 1\]=\\pm w, henceΔ=2​w\\Delta=2wand‖Δ‖22=4​B​\(F\)\\\|\\Delta\\\|\_\{2\}^\{2\}=4B\(F\)\. As in the proof of Prop\.[3\.1](https://arxiv.org/html/2609.38393#S3.Thmtheorem1), the law of total covariance yields

Σ\+\+Σ−=2​\(Ir−w​w⊤\),\\Sigma\_\{\+\}\+\\Sigma\_\{\-\}=2\(I\_\{r\}\-ww^\{\\top\}\),\(15\)whereΣs:=Cov⁡\(Z∣Y=s\)\\Sigma\_\{s\}:=\\mathrm\{Cov\}\(Z\\mid Y=s\); taking traces givesTr⁡\(Σ\+\)\+Tr⁡\(Σ−\)=2​\(r−B⁡\(F\)\)\\mathrm\{Tr\}\(\\Sigma\_\{\+\}\)\+\\mathrm\{Tr\}\(\\Sigma\_\{\-\}\)=2\(r\-B\(F\)\), and contracting withu:=w/‖w‖2u:=w/\\\|w\\\|\_\{2\}givesu⊤​\(Σ\+\+Σ−\)​u=2​\(1−B⁡\(F\)\)u^\{\\top\}\(\\Sigma\_\{\+\}\+\\Sigma\_\{\-\}\)u=2\(1\-B\(F\)\)\.

NCC rule as a margin\.Draw anmm\-shot support set with empirical class meansμ^±:=1m​∑i=1mZ±,i\\hat\{\\mu\}\_\{\\pm\}:=\\tfrac\{1\}\{m\}\\sum\_\{i=1\}^\{m\}Z\_\{\\pm,i\}, and setw^:=\(μ^\+−μ^−\)/2\\hat\{w\}:=\(\\hat\{\\mu\}\_\{\+\}\-\\hat\{\\mu\}\_\{\-\}\)/2andc:=\(μ^\+\+μ^−\)/2c:=\(\\hat\{\\mu\}\_\{\+\}\+\\hat\{\\mu\}\_\{\-\}\)/2\. The NCC rule isgSmNCC​\(z\)=sign​\(⟨w^,z−c⟩\)g\_\{S\_\{m\}\}^\{\\mathrm\{NCC\}\}\(z\)=\\textnormal\{sign\}\(\\langle\\hat\{w\},z\-c\\rangle\), since its decision boundary is the perpendicular bisector ofμ^±\\hat\{\\mu\}\_\{\\pm\}\. Define the signed marginMSm:=2​Y​⟨w^,Z−c⟩M\_\{S\_\{m\}\}:=2Y\\langle\\hat\{w\},Z\-c\\rangle\. Regardless of how ties are broken,\{gSmNCC\(Z\)≠Y\}⊆\{MSm≤0\}\\\{g\_\{S\_\{m\}\}^\{\\mathrm\{NCC\}\}\(Z\)\\neq Y\\\}\\subseteq\\\{M\_\{S\_\{m\}\}\\leq 0\\\}\.

Cantelli bound conditional on the support set\.Since\(Z,Y\)\(Z,Y\)is independent ofSmS\_\{m\}with𝔼⁡\[Y​Z\]=w\\mathbb\{E\}\[YZ\]=w,𝔼⁡\[Z\]=0\\mathbb\{E\}\[Z\]=0,𝔼⁡\[Z​Z⊤\]=Ir\\mathbb\{E\}\[ZZ^\{\\top\}\]=I\_\{r\},

𝔼⁡\[MSm∣Sm\]=2​⟨w^,w⟩,Var⁡\(MSm∣Sm\)=4​\(‖w^‖22\+⟨w^,c⟩2−⟨w^,w⟩2\)\.\\mathbb\{E\}\[M\_\{S\_\{m\}\}\\mid S\_\{m\}\]=2\\langle\\hat\{w\},w\\rangle,\\qquad\\mathrm\{Var\}\(M\_\{S\_\{m\}\}\\mid S\_\{m\}\)=4\\big\(\\\|\\hat\{w\}\\\|\_\{2\}^\{2\}\+\\langle\\hat\{w\},c\\rangle^\{2\}\-\\langle\\hat\{w\},w\\rangle^\{2\}\\big\)\.On the eventG:=\{⟨w^,w⟩\>0\}G:=\\\{\\langle\\hat\{w\},w\\rangle\>0\\\}, Cantelli’s inequality givesℙ⁡\(MSm≤0∣Sm\)≤1−f⁡\(Sm\)\\mathbb\{P\}\(M\_\{S\_\{m\}\}\\leq 0\\mid S\_\{m\}\)\\leq 1\-f\(S\_\{m\}\), where

f⁡\(Sm\):=\{⟨w^,w⟩2‖w^‖22\+⟨w^,c⟩2,w^≠0,0,w^=0\.f\(S\_\{m\}\):=\\begin\{cases\}\\dfrac\{\\langle\\hat\{w\},w\\rangle^\{2\}\}\{\\\|\\hat\{w\}\\\|\_\{2\}^\{2\}\+\\langle\\hat\{w\},c\\rangle^\{2\}\},&\\hat\{w\}\\neq 0,\\\\\[2\.84526pt\] 0,&\\hat\{w\}=0\.\\end\{cases\}Since0≤f⁡\(Sm\)≤10\\leq f\(S\_\{m\}\)\\leq 1, boundingf⁡\(Sm\)​𝟏G≥f⁡\(Sm\)−𝟏Gcf\(S\_\{m\}\)\\mathbf\{1\}\_\{G\}\\geq f\(S\_\{m\}\)\-\\mathbf\{1\}\_\{G^\{c\}\}and taking expectations gives

errmNCC​\(F\)≤1−𝔼⁡\[f⁡\(Sm\)\]\+ℙ⁡\(Gc\)\.\\mathrm\{err\}\_\{m\}^\{\\mathrm\{NCC\}\}\(F\)\\leq 1\-\\mathbb\{E\}\[f\(S\_\{m\}\)\]\+\\mathbb\{P\}\(G^\{c\}\)\.\(16\)
Lower bound on𝔼⁡\[f⁡\(Sm\)\]\\mathbb\{E\}\[f\(S\_\{m\}\)\]\.Whenw^≠0\\hat\{w\}\\neq 0, setp:=⟨w^,w⟩2/‖w^‖22=‖Pspan⁡\(w^\)​w‖22p:=\\langle\\hat\{w\},w\\rangle^\{2\}/\\\|\\hat\{w\}\\\|\_\{2\}^\{2\}=\\\|P\_\{\\mathrm\{span\}\(\\hat\{w\}\)\}w\\\|\_\{2\}^\{2\}; thenf⁡\(Sm\)=p/\(1\+⟨w^/‖w^‖2,c⟩2\)f\(S\_\{m\}\)=p/\(1\+\\langle\\hat\{w\}/\\\|\\hat\{w\}\\\|\_\{2\},c\\rangle^\{2\}\)\. The inequalityx/\(1\+y\)≥x−yx/\(1\+y\)\\geq x\-yfor0≤x≤10\\leq x\\leq 1,y≥0y\\geq 0, combined withp≤‖w‖22=B⁡\(F\)≤1p\\leq\\\|w\\\|\_\{2\}^\{2\}=B\(F\)\\leq 1, yieldsf⁡\(Sm\)≥p−‖c‖22f\(S\_\{m\}\)\\geq p\-\\\|c\\\|\_\{2\}^\{2\}\. Bounding the distance fromwwtospan⁡\(w^\)\\mathrm\{span\}\(\\hat\{w\}\)by‖w−w^‖2\\\|w\-\\hat\{w\}\\\|\_\{2\}givesp≥B⁡\(F\)−‖w^−w‖22p\\geq B\(F\)\-\\\|\\hat\{w\}\-w\\\|\_\{2\}^\{2\}; whenw^=0\\hat\{w\}=0the inequalityf⁡\(Sm\)≥B⁡\(F\)−‖w^−w‖22−‖c‖22f\(S\_\{m\}\)\\geq B\(F\)\-\\\|\\hat\{w\}\-w\\\|\_\{2\}^\{2\}\-\\\|c\\\|\_\{2\}^\{2\}still holds since then‖w−w^‖22=B⁡\(F\)\\\|w\-\\hat\{w\}\\\|\_\{2\}^\{2\}=B\(F\)\. Writingw^−w=\(ε\+−ε−\)/2\\hat\{w\}\-w=\(\\varepsilon\_\{\+\}\-\\varepsilon\_\{\-\}\)/2andc=\(ε\+\+ε−\)/2c=\(\\varepsilon\_\{\+\}\+\\varepsilon\_\{\-\}\)/2withε±:=μ^±−μ±\\varepsilon\_\{\\pm\}:=\\hat\{\\mu\}\_\{\\pm\}\-\\mu\_\{\\pm\}, class independence gives

𝔼​‖w^−w‖22=𝔼​‖c‖22=Tr⁡\(Σ\+\)\+Tr⁡\(Σ−\)4​m=r−B⁡\(F\)2​m,\\mathbb\{E\}\\\|\\hat\{w\}\-w\\\|\_\{2\}^\{2\}=\\mathbb\{E\}\\\|c\\\|\_\{2\}^\{2\}=\\frac\{\\mathrm\{Tr\}\(\\Sigma\_\{\+\}\)\+\\mathrm\{Tr\}\(\\Sigma\_\{\-\}\)\}\{4m\}=\\frac\{r\-B\(F\)\}\{2m\},so𝔼⁡\[f⁡\(Sm\)\]≥B⁡\(F\)−\(r−B⁡\(F\)\)/m\\mathbb\{E\}\[f\(S\_\{m\}\)\]\\geq B\(F\)\-\(r\-B\(F\)\)/m\.

*Bound onℙ⁡\(Gc\)\\mathbb\{P\}\(G^\{c\}\)\.*Since⟨w^,w⟩=B⁡\(F\)\+⟨w^−w,w⟩\\langle\\hat\{w\},w\\rangle=B\(F\)\+\\langle\\hat\{w\}\-w,w\\rangle, we have𝔼⁡\[⟨w^,w⟩\]=B⁡\(F\)\\mathbb\{E\}\[\\langle\\hat\{w\},w\\rangle\]=B\(F\), and

Var⁡\(⟨w^,w⟩\)=w⊤​Cov​\(w^−w\)​w=B⁡\(F\)4​m​u⊤​\(Σ\+\+Σ−\)​u=B​\(F\)​\(1−B​\(F\)\)2​m\\mathrm\{Var\}\(\\langle\\hat\{w\},w\\rangle\)=w^\{\\top\}\\mathrm\{Cov\}\(\\hat\{w\}\-w\)w=\\frac\{B\(F\)\}\{4m\}\\,u^\{\\top\}\(\\Sigma\_\{\+\}\+\\Sigma\_\{\-\}\)u=\\frac\{B\(F\)\(1\-B\(F\)\)\}\{2m\}using \([15](https://arxiv.org/html/2609.38393#A5.E15)\)\. Cantelli givesℙ⁡\(Gc\)≤\(1−B⁡\(F\)\)/\(1−B⁡\(F\)\+2​m​B​\(F\)\)\\mathbb\{P\}\(G^\{c\}\)\\leq\(1\-B\(F\)\)/\(1\-B\(F\)\+2mB\(F\)\)\.

Combining this with \([16](https://arxiv.org/html/2609.38393#A5.E16)\) and the lower bound on𝔼⁡\[f⁡\(Sm\)\]\\mathbb\{E\}\[f\(S\_\{m\}\)\]yieldserrmNCC​\(F\)≤1−B⁡\(F\)\+\(r−B⁡\(F\)\)/m\+\(1−B⁡\(F\)\)/\(1−B⁡\(F\)\+2​m​B​\(F\)\)\\mathrm\{err\}\_\{m\}^\{\\mathrm\{NCC\}\}\(F\)\\leq 1\-B\(F\)\+\(r\-B\(F\)\)/m\+\(1\-B\(F\)\)/\(1\-B\(F\)\+2mB\(F\)\)\. ∎

See[3\.5](https://arxiv.org/html/2609.38393#S3.Thmtheorem5)

###### Proof\.

WriteF=\(f1,…,fr\)F=\(f\_\{1\},\\dots,f\_\{r\}\)\. Centering and whitening make\{f1,…,fr\}\\\{f\_\{1\},\\dots,f\_\{r\}\\\}orthonormal inL2​\(PX\)L^\{2\}\(P\_\{X\}\), so𝒮F:=span⁡\{f1,…,fr\}\\mathcal\{S\}\_\{F\}:=\\mathrm\{span\}\\\{f\_\{1\},\\dots,f\_\{r\}\\\}isrr\-dimensional with orthogonal projectorΠ𝒮F\\Pi\_\{\\mathcal\{S\}\_\{F\}\}\. For each tasktt, letwt:=𝔼⁡\[Yt​F​\(X\)\]∈ℝrw\_\{t\}:=\\mathbb\{E\}\[Y\_\{t\}F\(X\)\]\\in\\mathbb\{R\}^\{r\}with coordinates\(wt\)ℓ=𝔼⁡\[ηt​\(X\)​fℓ​\(X\)\]=⟨ηt,fℓ⟩\(w\_\{t\}\)\_\{\\ell\}=\\mathbb\{E\}\[\\eta\_\{t\}\(X\)f\_\{\\ell\}\(X\)\]=\\langle\\eta\_\{t\},f\_\{\\ell\}\\rangle, soΠ𝒮F​ηt=∑ℓ\(wt\)ℓ​fℓ\\Pi\_\{\\mathcal\{S\}\_\{F\}\}\\eta\_\{t\}=\\sum\_\{\\ell\}\(w\_\{t\}\)\_\{\\ell\}f\_\{\\ell\}and hence

‖wt‖22=‖Π𝒮F​ηt‖L22=B\(t\)​\(F\)\.\\\|w\_\{t\}\\\|\_\{2\}^\{2\}=\\\|\\Pi\_\{\\mathcal\{S\}\_\{F\}\}\\eta\_\{t\}\\\|\_\{L^\{2\}\}^\{2\}=B^\{\(t\)\}\(F\)\.\(17\)Writeηt⟂:=\(I−Π𝒮F\)​ηt\\eta\_\{t\}^\{\\perp\}:=\(I\-\\Pi\_\{\\mathcal\{S\}\_\{F\}\}\)\\eta\_\{t\}, so‖ηt⟂‖L22=‖ηt‖L22−B\(t\)​\(F\)\\\|\\eta\_\{t\}^\{\\perp\}\\\|\_\{L^\{2\}\}^\{2\}=\\\|\\eta\_\{t\}\\\|\_\{L^\{2\}\}^\{2\}\-B^\{\(t\)\}\(F\)\.

Near\-orthogonality\.By \([17](https://arxiv.org/html/2609.38393#A5.E17)\),ui⊤​uj=⟨Π𝒮F​ηi,Π𝒮F​ηj⟩/B\(i\)​\(F\)​B\(j\)​\(F\)u\_\{i\}^\{\\top\}u\_\{j\}=\\langle\\Pi\_\{\\mathcal\{S\}\_\{F\}\}\\eta\_\{i\},\\Pi\_\{\\mathcal\{S\}\_\{F\}\}\\eta\_\{j\}\\rangle/\\sqrt\{B^\{\(i\)\}\(F\)B^\{\(j\)\}\(F\)\}\. Since⟨Π𝒮F​ηi,Π𝒮F​ηj⟩=⟨ηi,ηj⟩−⟨ηi⟂,ηj⟂⟩\\langle\\Pi\_\{\\mathcal\{S\}\_\{F\}\}\\eta\_\{i\},\\Pi\_\{\\mathcal\{S\}\_\{F\}\}\\eta\_\{j\}\\rangle=\\langle\\eta\_\{i\},\\eta\_\{j\}\\rangle\-\\langle\\eta\_\{i\}^\{\\perp\},\\eta\_\{j\}^\{\\perp\}\\rangle, Cauchy–Schwarz on the residual term gives

\|⟨Π𝒮F​ηi,Π𝒮F​ηj⟩\|≤\|⟨ηi,ηj⟩\|\+\(‖ηi‖L22−B\(i\)​\(F\)\)​\(‖ηj‖L22−B\(j\)​\(F\)\)\.\\big\|\\langle\\Pi\_\{\\mathcal\{S\}\_\{F\}\}\\eta\_\{i\},\\Pi\_\{\\mathcal\{S\}\_\{F\}\}\\eta\_\{j\}\\rangle\\big\|\\leq\|\\langle\\eta\_\{i\},\\eta\_\{j\}\\rangle\|\+\\sqrt\{\(\\\|\\eta\_\{i\}\\\|\_\{L^\{2\}\}^\{2\}\-B^\{\(i\)\}\(F\)\)\(\\\|\\eta\_\{j\}\\\|\_\{L^\{2\}\}^\{2\}\-B^\{\(j\)\}\(F\)\)\}\.\(18\)To bound the first term, setξt:=Yt−ηt​\(X\)\\xi\_\{t\}:=Y\_\{t\}\-\\eta\_\{t\}\(X\); then𝔼⁡\[ξt∣X\]=0\\mathbb\{E\}\[\\xi\_\{t\}\\mid X\]=0and𝔼⁡\[ξt2\]=εt\\mathbb\{E\}\[\\xi\_\{t\}^\{2\}\]=\\varepsilon\_\{t\}\. ExpandingYi​Yj=\(ηi\+ξi\)​\(ηj\+ξj\)Y\_\{i\}Y\_\{j\}=\(\\eta\_\{i\}\+\\xi\_\{i\}\)\(\\eta\_\{j\}\+\\xi\_\{j\}\)and using conditional mean zero on the cross terms,ρi​j=𝔼⁡\[ηi​\(X\)​ηj​\(X\)\]\+𝔼⁡\[ξi​ξj\]\\rho\_\{ij\}=\\mathbb\{E\}\[\\eta\_\{i\}\(X\)\\eta\_\{j\}\(X\)\]\+\\mathbb\{E\}\[\\xi\_\{i\}\\xi\_\{j\}\], so\|⟨ηi,ηj⟩\|≤\|ρi​j\|\+εi​εj\|\\langle\\eta\_\{i\},\\eta\_\{j\}\\rangle\|\\leq\|\\rho\_\{ij\}\|\+\\sqrt\{\\varepsilon\_\{i\}\\varepsilon\_\{j\}\}by Cauchy–Schwarz\. Substituting into \([18](https://arxiv.org/html/2609.38393#A5.E18)\) yields the stated bound on\|ui⊤​uj\|\|u\_\{i\}^\{\\top\}u\_\{j\}\|; no independence among theYtY\_\{t\}was used\.

Centroid geometry\.LetZ⁡\(X\)=\(Z1​\(X\),…,Zk​\(X\)\)Z\(X\)=\(Z\_\{1\}\(X\),\\dots,Z\_\{k\}\(X\)\)andmY\(t\):=𝔼⁡\[Zt​\(X\)∣Y\]m\_\{Y\}^\{\(t\)\}:=\\mathbb\{E\}\[Z\_\{t\}\(X\)\\mid Y\]\. By \([17](https://arxiv.org/html/2609.38393#A5.E17)\),𝔼⁡\[Yt​Zt​\(X\)\]=ut⊤​wt=‖wt‖2=B\(t\)​\(F\)\\mathbb\{E\}\[Y\_\{t\}Z\_\{t\}\(X\)\]=u\_\{t\}^\{\\top\}w\_\{t\}=\\\|w\_\{t\}\\\|\_\{2\}=\\sqrt\{B^\{\(t\)\}\(F\)\}, and iterated expectation gives𝔼⁡\[Yt​mY\(t\)\]=B\(t\)​\(F\)\\mathbb\{E\}\[Y\_\{t\}m\_\{Y\}^\{\(t\)\}\]=\\sqrt\{B^\{\(t\)\}\(F\)\}\. Expanding the square and usingYt2=1Y\_\{t\}^\{2\}=1,

𝔼⁡\[\(mY\(t\)−B\(t\)​\(F\)​Yt\)2\]=𝔼⁡\[\(mY\(t\)\)2\]−B\(t\)​\(F\)≤𝔼⁡\[Zt2\]−B\(t\)​\(F\)=1−B\(t\)​\(F\),\\mathbb\{E\}\\big\[\(m\_\{Y\}^\{\(t\)\}\-\\sqrt\{B^\{\(t\)\}\(F\)\}\\,Y\_\{t\}\)^\{2\}\\big\]=\\mathbb\{E\}\[\(m\_\{Y\}^\{\(t\)\}\)^\{2\}\]\-B^\{\(t\)\}\(F\)\\leq\\mathbb\{E\}\[Z\_\{t\}^\{2\}\]\-B^\{\(t\)\}\(F\)=1\-B^\{\(t\)\}\(F\),where Jensen gives the inequality and whitening gives𝔼⁡\[Zt2\]=ut⊤​Ir​ut=1\\mathbb\{E\}\[Z\_\{t\}^\{2\}\]=u\_\{t\}^\{\\top\}I\_\{r\}u\_\{t\}=1; this step also used no joint structure of theYtY\_\{t\}\. Summing overttyields the claimed centroid bound\. ∎

See[4\.1](https://arxiv.org/html/2609.38393#S4.Thmtheorem1)

###### Proof\.

WriteF=\(f1,…,fr\)F=\(f\_\{1\},\\dots,f\_\{r\}\)\. The constraints𝔼⁡\[F⁡\(X\)\]=0\\mathbb\{E\}\[F\(X\)\]=0and𝔼⁡\[F⁡\(X\)​F​\(X\)⊤\]=Ir\\mathbb\{E\}\[F\(X\)F\(X\)^\{\\top\}\]=I\_\{r\}are equivalent to requiring\{f1,…,fr\}\\\{f\_\{1\},\\dots,f\_\{r\}\\\}to be an orthonormal family inL02​\(PX\)L\_\{0\}^\{2\}\(P\_\{X\}\)\. Using \([4](https://arxiv.org/html/2609.38393#A3.E4)\), the objective is therefore

𝔼⁡\[⟨F⁡\(X\(1\)\),F⁡\(X\(2\)\)⟩\]=∑ℓ=1r⟨fℓ,T​fℓ⟩L2​\(PX\)\.\\mathbb\{E\}\\\!\\left\[\\langle F\(X^\{\(1\)\}\),F\(X^\{\(2\)\}\)\\rangle\\right\]=\\sum\_\{\\ell=1\}^\{r\}\\langle f\_\{\\ell\},Tf\_\{\\ell\}\\rangle\_\{L^\{2\}\(P\_\{X\}\)\}\.
Expandfℓ=∑j≥1aℓ​j​ψjf\_\{\\ell\}=\\sum\_\{j\\geq 1\}a\_\{\\ell j\}\\psi\_\{j\}in the orthonormal eigenbasis ofTT, and definesj:=∑ℓ=1raℓ​j2s\_\{j\}:=\\sum\_\{\\ell=1\}^\{r\}a\_\{\\ell j\}^\{2\}\. SinceT​ψj=λj​ψjT\\psi\_\{j\}=\\lambda\_\{j\}\\psi\_\{j\},

∑ℓ=1r⟨fℓ,T​fℓ⟩=∑j≥1λj​sj,0≤sj≤1,∑j≥1sj=r\.\\sum\_\{\\ell=1\}^\{r\}\\langle f\_\{\\ell\},Tf\_\{\\ell\}\\rangle=\\sum\_\{j\\geq 1\}\\lambda\_\{j\}s\_\{j\},\\qquad 0\\leq s\_\{j\}\\leq 1,\\qquad\\sum\_\{j\\geq 1\}s\_\{j\}=r\.The inequalities follow because the coefficient matrixA=\(aℓ​j\)A=\(a\_\{\\ell j\}\)has orthonormal rows\. Sinceλ1≥λ2≥⋯\\lambda\_\{1\}\\geq\\lambda\_\{2\}\\geq\\cdots, the maximum is attained bys1=⋯=sr=1s\_\{1\}=\\cdots=s\_\{r\}=1andsj=0s\_\{j\}=0forj\>rj\>r, with optimal value∑j=1rλj\\sum\_\{j=1\}^\{r\}\\lambda\_\{j\}\.

TakingF0=\(ψ1,…,ψr\)F\_\{0\}=\(\\psi\_\{1\},\\dots,\\psi\_\{r\}\)attains this value, soψ1,…,ψr\\psi\_\{1\},\\dots,\\psi\_\{r\}are feasible and optimal\. Finally, the eigengapλr\>λr\+1\\lambda\_\{r\}\>\\lambda\_\{r\+1\}makes the maximizing allocation unique\. Hence any optimalFFhas coordinate spanspan⁡\{ψ1,…,ψr\}\\mathrm\{span\}\\\{\\psi\_\{1\},\\dots,\\psi\_\{r\}\\\}\. ∎

See[4\.2](https://arxiv.org/html/2609.38393#S4.Thmtheorem2)

###### Proof\.

By Prop\.[4\.1](https://arxiv.org/html/2609.38393#S4.Thmtheorem1),𝒮F⋆=𝒰r\\mathcal\{S\}\_\{F^\{\\star\}\}=\\mathcal\{U\}\_\{r\}\. Since\(ψj\)j≥1\(\\psi\_\{j\}\)\_\{j\\geq 1\}is an orthonormal eigenbasis ofTTonL02​\(PX\)L\_\{0\}^\{2\}\(P\_\{X\}\)andη∈L02​\(PX\)\\eta\\in L\_\{0\}^\{2\}\(P\_\{X\}\), writeη=∑j≥1βj​ψj\\eta=\\sum\_\{j\\geq 1\}\\beta\_\{j\}\\psi\_\{j\}, whereβj=⟨η,ψj⟩L2​\(PX\)\\beta\_\{j\}=\\langle\\eta,\\psi\_\{j\}\\rangle\_\{L^\{2\}\(P\_\{X\}\)\}\. HenceΠr​η=∑j=1rβj​ψj\\Pi\_\{r\}\\eta=\\sum\_\{j=1\}^\{r\}\\beta\_\{j\}\\psi\_\{j\}, and orthonormality givesB⁡\(F⋆\)=‖Πr​η‖L2​\(PX\)2=∑j=1rβj2B\(F^\{\\star\}\)=\\\|\\Pi\_\{r\}\\eta\\\|\_\{L^\{2\}\(P\_\{X\}\)\}^\{2\}=\\sum\_\{j=1\}^\{r\}\\beta\_\{j\}^\{2\}\. ∎

### E\.2Proofs for Appendix[C](https://arxiv.org/html/2609.38393#A3)

See[C\.1](https://arxiv.org/html/2609.38393#A3.Thmtheorem1)

###### Proof\.

LetG=𝔼⁡\[F⁡\(X\)​F​\(X\)⊤\]G=\\mathbb\{E\}\[F\(X\)F\(X\)^\{\\top\}\]\. Both terms in \([3](https://arxiv.org/html/2609.38393#A3.E3)\) are invariant under orthogonal changes of feature coordinates, so diagonalizeGGand write the resulting coordinates asgj=tj​hjg\_\{j\}=\\sqrt\{t\_\{j\}\}\\,h\_\{j\}, wheretj≥0t\_\{j\}\\geq 0and the nonzerohjh\_\{j\}’s are orthonormal inL02​\(PX\)L\_\{0\}^\{2\}\(P\_\{X\}\); complete them arbitrarily to an orthonormal family when sometj=0t\_\{j\}=0\. Reorder so thatt1≥⋯≥trt\_\{1\}\\geq\\cdots\\geq t\_\{r\}\. The objective becomes

∑j=1rtj​⟨hj,T​hj⟩−γ2​∑j=1r\(tj−1\)2\.\\sum\_\{j=1\}^\{r\}t\_\{j\}\\langle h\_\{j\},Th\_\{j\}\\rangle\-\\frac\{\\gamma\}\{2\}\\sum\_\{j=1\}^\{r\}\(t\_\{j\}\-1\)^\{2\}\.By the weighted Ky Fan principle,∑jtj​⟨hj,T​hj⟩≤∑jtj​λj\\sum\_\{j\}t\_\{j\}\\langle h\_\{j\},Th\_\{j\}\\rangle\\leq\\sum\_\{j\}t\_\{j\}\\lambda\_\{j\}\. Hence the objective is at most

∑j=1r\[tj​λj−γ2​\(tj−1\)2\],\\sum\_\{j=1\}^\{r\}\\left\[t\_\{j\}\\lambda\_\{j\}\-\\frac\{\\gamma\}\{2\}\(t\_\{j\}\-1\)^\{2\}\\right\],whosejj\-th term is uniquely maximized attj=1\+λj/γt\_\{j\}=1\+\\lambda\_\{j\}/\\gamma\. Equality is attained by takinghj=ψjh\_\{j\}=\\psi\_\{j\}\. The eigengap fixes the maximizingrr\-dimensional subspace, while rotations within eigenspaces of repeated eigenvalues are absorbed into an orthogonal output matrix\. Thus every maximizer has the stated form withsj2=tj=1\+λj/γs\_\{j\}^\{2\}=t\_\{j\}=1\+\\lambda\_\{j\}/\\gamma\. ∎

See[C\.2](https://arxiv.org/html/2609.38393#A3.Thmtheorem2)

###### Proof\.

LetD=diag⁡\(s12,…,sr2\)D=\\mathrm\{diag\}\(s\_\{1\}^\{2\},\\dots,s\_\{r\}^\{2\}\)andb=\(s1​β1,…,sr​βr\)⊤b=\(s\_\{1\}\\beta\_\{1\},\\dots,s\_\{r\}\\beta\_\{r\}\)^\{\\top\}\. ThenCov⁡\(Fγ⋆​\(X\)\)=R​D​R⊤\\mathrm\{Cov\}\(F^\{\\star\}\_\{\\gamma\}\(X\)\)=RDR^\{\\top\}, while balancedness and centering giveμ\+=R​b\\mu\_\{\+\}=Rb,μ−=−μ\+\\mu\_\{\-\}=\-\\mu\_\{\+\}, and‖μ\+‖22=Bγ\\\|\\mu\_\{\+\}\\\|\_\{2\}^\{2\}=B\_\{\\gamma\}\. By total covariance,

Σ\+\+Σ−=2​\(R​D​R⊤−μ\+​μ\+⊤\)\.\\Sigma\_\{\+\}\+\\Sigma\_\{\-\}=2\\big\(RDR^\{\\top\}\-\\mu\_\{\+\}\\mu\_\{\+\}^\{\\top\}\\big\)\.Taking traces and using‖Δ‖22=4​Bγ\\\|\\Delta\\\|\_\{2\}^\{2\}=4B\_\{\\gamma\}yieldsVFγ⋆=12​\(Sγ/Bγ−1\)V\_\{F^\{\\star\}\_\{\\gamma\}\}=\\tfrac\{1\}\{2\}\(S\_\{\\gamma\}/B\_\{\\gamma\}\-1\)\. Foru=μ\+/Bγu=\\mu\_\{\+\}/\\sqrt\{B\_\{\\gamma\}\}, we haveu⊤​R​D​R⊤​u=Cγ/Bγu^\{\\top\}RDR^\{\\top\}u=C\_\{\\gamma\}/B\_\{\\gamma\}, so contraction alonguugivesV~Fγ⋆=12​\(Cγ/Bγ2−1\)\\tilde\{V\}\_\{F^\{\\star\}\_\{\\gamma\}\}=\\tfrac\{1\}\{2\}\(C\_\{\\gamma\}/B\_\{\\gamma\}^\{2\}\-1\)\. The caseBγ=0B\_\{\\gamma\}=0follows from the stated convention\. ∎

Similar Articles

Self-Distillation as a Performance Recovery Mechanism for LLMs: Counteracting Compression and Catastrophic Forgetting

arXiv cs.CL

This paper introduces Self-Distillation Fine-Tuning (SDFT) as a recovery mechanism for LLMs suffering from performance degradation due to catastrophic forgetting, quantization, and pruning. The authors provide theoretical justification using Centered Kernel Alignment (CKA) to demonstrate that self-distillation aligns the student model's high-dimensional manifold with the teacher's optimal structure, effectively recovering lost capabilities.

Similarity-Aware Machine Unlearning

arXiv cs.LG

This paper proposes a retain-aware localization method for machine unlearning that reduces collateral damage to semantically similar retained examples, and introduces a retain-similar evaluation set. Experiments on CIFAR-10 with ResNet18 show reduced collateral damage and improved unlearning metrics.

UniSD: Towards a Unified Self-Distillation Framework for Large Language Models

Hugging Face Daily Papers

This paper introduces UniSD, a unified self-distillation framework for adapting large language models that integrates mechanisms for supervision reliability, representation alignment, and training stability. Experimental results show that UniSD improves performance over base models and existing baselines across multiple benchmarks.