When Graph-JEPA Learns the Wrong Thing: Diagnosing and Repairing Category-Conditional Collapse
Summary
This paper diagnoses category-conditional collapse in Graph-JEPA models, where standard metrics indicate healthy representations but usable instance information is absent, and proposes a repair method while proving structural reducibility in the retrieval target.
View Cached Full Text
Cached at: 08/24/26, 04:30 AM
# When Graph-JEPA Learns the Wrong Thing: Diagnosing and Repairing Category-Conditional Collapse
Source: [https://arxiv.org/html/2608.20516](https://arxiv.org/html/2608.20516)
Gollam Rabbygollam\.rabby@l3s\.deSören Auerauer@tib\.euAffiliation:TIB Leibniz Information Centre for Science and Technology,Affiliation:Hannover, Germany
###### Abstract
Joint\-embedding predictive architectures are selected almost universally by linear probing and by effective rank\. We report a case in which both read healthily while the representation carries zero usable instance information\. We then repair it, and a second failure appears: the repaired metric saturates on a target that provably carries no structural information\. Our corpus is a heterogeneous scientific\-reasoning graph over57,90357\{,\}903scientific articles, each article forming a subgraph\. A Graph\-JEPA is trained to predict one masked aspect from a subgraph’s remaining aspects\. It attains linear\-probe accuracy0\.8710\.871and node\-level effective rank1818–4747\. Masked\-aspect retrieval against the full corpus nevertheless recovers0\.000\.00of14\.414\.4recoverable bits \(MRR=1\.9×10−4\\mathrm\{MRR\}=1\.9\\times 10^\{\-4\}against an exact chance level of1\.99×10−41\.99\\times 10^\{\-4\}, permutationp=0\.98p=0\.98\)\. Three upper bounds on the same pool, features, and scoring code recover almost everything: a parameter\-free average of the visible aspects reaches\+14\.281\+14\.281bits, Okapi BM25\+14\.335\+14\.335, and a same\-evaluation positive control\+14\.220\+14\.220\. This excludes the corpus, the masking, the pool, and the metric as causes\. We locate the mechanism in a measured variance allocation, reported across encoders and within the single encoder that produced the trained latents\. The frozen inputs place86\.05%86\.05\\%of their variance on subgraph identity and0\.40%0\.40\\%on aspect identity, with the trained latents placing0\.39%0\.39\\%and99\.61%99\.61\\%\. The deficit is a property of the objective’s optimum rather than of its optimisation: we prove that the category\-measurable solution is a global minimum of the coupled predictor and EMA\-target objective\. We also measure the allocation at every checkpoint, because our own rank trajectory shows the pooled\-rank signature is already present at initialisation\. A repaired configuration reaches14\.37714\.377of14\.37914\.379bits, above the13\.86513\.865\-bit training\-free oracle\. Changing only the loss back to regression returns it to0\.3070\.307bits\. This14\.0514\.05bit swing on one variable confirms the mechanism directly\. The repair nonetheless licenses nothing about reasoning\. We prove that the retrieval target is structurally reducible\. The intra\-subgraph edge set is a deterministic function of the node census, so the subgraph supplies no information beyond each subgraph’s own features\. The training\-free oracle therefore already reaches96\.4%96\.4\\%of the ceiling, and our largest positive effect is the learning\-rate schedule \(\+1\.337\+1\.337bits\) rather than any architectural factor\. Across ten converged cells, bits \(11\.808−14\.37911\.808\-14\.379\) and a held\-out reasoning probe \(0\.711−0\.9850\.711\-0\.985\) show no reliable relationship\. Replacing the target with a data\-derived one then fails a pre\-registered data\-quality gate:25\.96%25\.96\\%of contradicting\-evidence nodes are exact duplicate placeholder strings, and the remainder is\+0\.2377\+0\.2377more generic than supporting evidence\. Rank, probes, and the task metric can all saturate on an evaluation that cannot support the claim\. We release a pre\-registered harness that adds a reducibility audit and a data\-derived target gate to the standard toolkit\.
## 1Introduction
Self\-supervised learning through joint\-embedding prediction is now a dominant paradigm\([Assran et al\. 2023](https://arxiv.org/html/2608.20516#bib.bib1);[Dawid & LeCun 2023](https://arxiv.org/html/2608.20516#bib.bib5)\)\. Rather than reconstructing inputs, a JEPA predicts the latent representation of a masked target from a visible context\. An exponential moving average target encoder supplies the signal\. Graph\-JEPA\([Skenderi et al\. 2025](https://arxiv.org/html/2608.20516#bib.bib20)\)adapts this to graphs, predicting a low\-dimensional coordinate of each masked subgraph from its context\.
Latent\-predictive objectives admit degenerate solutions, so the field has converged on two health checks: a linear probe on a downstream label, and the effective rank of the representation\([Roy & Vetterli 2007](https://arxiv.org/html/2608.20516#bib.bib18);[Jing et al\. 2022](https://arxiv.org/html/2608.20516#bib.bib15)\)\. A model with high probe accuracy and effective rank well above one is generally taken to have avoided collapse\. RankMe\([Garrido et al\. 2023](https://arxiv.org/html/2608.20516#bib.bib10)\)elevates rank to a model\-selection criterion outright\.
We report an analysis in which both checks pass while the representation is worthless for its intended use\. A second, sharper failure then appears once the first is fixed\. We build a heterogeneous reasoning graph over57,90357\{,\}903scientific papers\. Each scientific paper is decomposed into a subgraph of typed reasoning nodes:claim,method,result,evidenceandimplication\. We train a Graph\-JEPA to predict one masked aspect from the others \(Sec\.[3](https://arxiv.org/html/2608.20516#S3), App\.[A7](https://arxiv.org/html/2608.20516#A7)\)\. The probe reaches0\.8710\.871, the node\-level effective rank is1818–4747of128128, and the loss converges\. Retrieval of the masked aspect from the full57,90357\{,\}903\-item corpus nevertheless recovers0\.000\.00of14\.37914\.379recoverable bits, at permutationp=0\.98p=0\.98\(Fig\.[1](https://arxiv.org/html/2608.20516#S1.F1)\)\.
Figure 1:Neither standard health check predicts retrieval, and the null is exact\.\(left\)Pooled effective rank spans1\.21\.2–300300across all trained configurations and the ceilings; every trained point lies on the chance line and the ceilings lie∼14\{\\sim\}14bits above it\.\(centre\)Probe accuracy spans0\.570\.57–0\.970\.97with no corresponding movement; the vertical spread is sampling noise around chance\.\(right\)Cumulative distribution of the gold rank: baseline, centred baseline, best encoder and best fix are indistinguishable from the uniform\-random curve \(dotted\) over the entire range, while the ceiling reaches≈0\.6\{\\approx\}0\.6at rank11\. Not a shifted distribution, the same distribution\.Three controls decide the first result\.All three use the identical pool, masking, and scoring code\.*\(i\)*A parameter\-free average of the visible aspects’ frozen features recovers\+14\.281\+14\.281bits \(MRR=0\.9719\\mathrm\{MRR\}=0\.9719, CI\[0\.9674,0\.9761\]\[0\.9674,0\.9761\]\)\.*\(ii\)*Okapi BM25 over raw text, using no learned representation, recovers\+14\.335\+14\.335bits\.*\(iii\)*A positive control, namely a same\-architecture predictor trained on the same frozen features under the same evaluation, recovers\+14\.220\+14\.220bits\. The information is present, and the harness extracts it\. No standard diagnostic detects that the trained model does not\.
The repair is the more interesting result\.A corrected objective takes the same pipeline from0\.000\.00to14\.37714\.377of14\.37914\.379bits, above the13\.86513\.865\-bit training\-free oracle\. Swapping only the loss back to regression returns it to0\.3070\.307bits: a14\.0514\.05\-bit swing on one variable\. Yet the repaired score licenses nothing\. We prove that the target is structurally reducible\. And across ten converged cells, the optimised metric and a held\-out reasoning probe show no reliable relationship \(Spearman−0\.24\-0\.24,n=10n=10\)\.
##### What is new relative to prior probe–task mismatch reports\.
The first dissociation is complete rather than partial: zero of14\.37914\.379recoverable bits, with a permutation null and a positive control that excludes every non\-model explanation we could construct\. The mechanism is a measured allocation, together with a proof that the degenerate configuration is a global optimum rather than an unfavourable property of the data\. The same corpus that supplies86\.05%86\.05\\%of its variance to paper identity yields a representation that supplies0\.39%0\.39\\%\(Sec\.[7](https://arxiv.org/html/2608.20516#S7)\)\. The failure also localises to the query bank\. That places it outside the reach of the entire candidate\-side toolkit—whitening, Mahalanobis metrics, hyperbolic embeddings—and it explains why every such remedy bought0\.4bits0\.4\\text\{ bits\}against a14\.4bits14\.4\\text\{ bits\}deficit\. The repair then exposes a failure one level up\. We prove that the retrieval target is structurally reducible \(Prop\.[5](https://arxiv.org/html/2608.20516#Thmproposition5)\), and we show that the data\-derived replacement fails a data\-quality gate\. No representation\-level metric on this corpus currently licenses a reasoning claim \(Sec\.[9](https://arxiv.org/html/2608.20516#S9)\)\.
##### Contributions\.
All retrieval numbers are bits recovered relative to a uniform random ranker,ℬ=1Nlog2N\!−𝔼\[log2r\]\\mathcal\{B\}=\\frac\{1\}\{N\}\\log\_\{2\}N\!\-\\mathbb\{E\}\[\\log\_\{2\}r\], reported with bootstrap intervals and an exact chance baseline\. We report no ratios against chance \(Sec\.[3](https://arxiv.org/html/2608.20516#S3), App\.[A8](https://arxiv.org/html/2608.20516#A8)\)\.\(1\) A complete dissociation with a positive control\(Secs\.[5](https://arxiv.org/html/2608.20516#S5)–[6](https://arxiv.org/html/2608.20516#S6)\): probe0\.8710\.871against0\.000\.00bits \(Fig\.[1](https://arxiv.org/html/2608.20516#S1.F1)\), while a control through the identical harness recovers\+14\.220\+14\.220\(98\.9%98\.9\\%of ceiling\)\.\(2\) The mechanism as a measured variance allocation, endpoint and trajectory\(Sec\.[7](https://arxiv.org/html/2608.20516#S7)\): paper\-identity share86\.05%→0\.39%86\.05\\%\\\!\\to\\\!0\.39\\%, and aspect share0\.40%→99\.61%0\.40\\%\\\!\\to\\\!99\.61\\%\. We report this within one encoder as well as across two\. A per\-checkpoint measurement separates what the converged representation encodes from when it came to encode it \(Sec\.[7\.1](https://arxiv.org/html/2608.20516#S7.SS1)\)\. A phase analysis at the measuredκ^,ε^\\hat\{\\kappa\},\\hat\{\\varepsilon\}attributes11\.011\.0of12\.84512\.845lost bits to the allocation\.\(3\) Query\-side localisation\(Sec\.[5](https://arxiv.org/html/2608.20516#S5)\): the candidate bank hasrcand=18\.1r\_\{\\mathrm\{cand\}\}=18\.1and DC ratio37\.637\.6\. The query bank hasrquery=1\.9r\_\{\\mathrm\{query\}\}=1\.9, DC energy0\.99990\.9999and self\-similarity0\.0040\.004, consistent with the bound of Prop\.[1](https://arxiv.org/html/2608.20516#Thmproposition1)\.\(4\) One non\-identifiability result, two invariance results and one reducibility result\(Sec\.[4](https://arxiv.org/html/2608.20516#S4)\)\. The coupled predictor and EMA\-target objective admits global minima that differ by the entire recoverable budget, so the loss does not identify whether instance identity is encoded\. A degenerate query bank induces the same candidate permutation under any injective transform\. The minimiser of a squared\-distance objective is a Fréchet mean in any metric space\. And a subgraph whose edge set is determined by its node census supplies no information beyond node features\.\(5\) A one\-variable confirmation of the mechanism\(Sec\.[9](https://arxiv.org/html/2608.20516#S9)\): holding frame, cue, budget, schedule and seed fixed, changing only the loss moves the pipeline→0\.30714\.359\\\!\\to\\\!0\.307bits, with the reasoning probe falling to the majority class\.\(6\) A data\-derived\-target gate that fires\(Sec\.[9](https://arxiv.org/html/2608.20516#S9)\):25\.96%25\.96\\%of contradicting\-evidence nodes are exact duplicates, and the genericness gap is\+0\.2377\+0\.2377\. The paper\-identity leak we expected to find is measured at\+0\.0003\+0\.0003, i\.e\. absent\.\(7\) A characterisation of the instrument\(Sec\.[8](https://arxiv.org/html/2608.20516#S8)\): BM25 recovers\+14\.335\+14\.335bits, and a function\-word\-only classifier identifies aspect type at0\.9370\.937against chance0\.3330\.333\.\(8\) A pre\-registered harnesscosting0\.770\.77GPU\-hours, with a2020\-assertion instrument self\-test that must pass before any measurement is taken \(App\.[A9](https://arxiv.org/html/2608.20516#A9),[A10](https://arxiv.org/html/2608.20516#A10)\)\.
Figure 2:Pipeline and diagnosis\.\(1\)Graph construction over57,90357\{,\}903papers with typed reasoning nodes and edges\.\(2\)Patch extraction;methodandresultare exact singletons\.\(3\)Frozen sentence features feed a heterogeneous GNN trained with a JEPA objective; we sweep target geometry, aggregator, loss, target frame, structural cue, and regularisation\.\(4\)Both standard health checks pass while retrieval recovers zero bits\. Three ceilings show the signal is present, the variance allocation shows where it goes, and Sec\.[9](https://arxiv.org/html/2608.20516#S9)shows that repairing it saturates a reducible target\.Figure 3:From the global reasoning graph to a paper\-local target patch\.\(a\)The heterogeneous graph; an extraction window selects one paper and its incident reasoning nodes\.\(b\)The extracted subgraph, a paper node joined by typed edges to its claim, method and result nodes\. Note that every edge shown is present for*every*paper possessing the endpoint node types, which is the content of Prop\.[5](https://arxiv.org/html/2608.20516#Thmproposition5)\.
## 2Related Work
An extended discussion is in App\.[A6](https://arxiv.org/html/2608.20516#A6)\.
##### JEPAs and non\-contrastive collapse\.
JEPAs replace reconstruction with latent prediction\([Assran et al\. 2023](https://arxiv.org/html/2608.20516#bib.bib1);[Dawid & LeCun 2023](https://arxiv.org/html/2608.20516#bib.bib5)\)\. The same philosophy underlies BYOL\([Grill et al\. 2020](https://arxiv.org/html/2608.20516#bib.bib12)\), VICReg\([Bardes et al\. 2022](https://arxiv.org/html/2608.20516#bib.bib3)\)and Barlow Twins\([Zbontar et al\. 2021](https://arxiv.org/html/2608.20516#bib.bib24)\)\. Existing theory targets complete collapse, in whichffbecomes constant on the input domain\. The failure we report is conditional\. The representation is high\-rank and well spread across the candidate bank, and collapses only within each latent category\. No existing criterion flags it\. The converged latents allocate their variance to the category rather than to the instance \(Table[3](https://arxiv.org/html/2608.20516#S7.T3)\)\. Prop\.[1](https://arxiv.org/html/2608.20516#Thmproposition1)then shows why that configuration is a global optimum of the coupled objective rather than a failure of its optimisation\.
##### Graph\-JEPA and evaluation\-target reducibility\.
Graph\-JEPA\([Skenderi et al\. 2025](https://arxiv.org/html/2608.20516#bib.bib20)\)predicts a low\-dimensional hyperbolic coordinate of each masked patch\. It builds on masked reconstruction\([Hou et al\. 2023](https://arxiv.org/html/2608.20516#bib.bib13)\)and on contrastive multi\-view objectives\([Yao et al\. 2024](https://arxiv.org/html/2608.20516#bib.bib23)\)\. It reports strong subgraph classification, which is precisely the regime in which a category\-measurable solution is not merely adequate but optimal\. Prop\.[5](https://arxiv.org/html/2608.20516#Thmproposition5)raises a different question: is the subgraph structure a deterministic function of the node census, and therefore information\-free? To our knowledge, the graph\-SSL evaluation literature has not asked that question\. It is a one\-line check, and it invalidates the structural interpretation of any gain on such a task\.
##### Class collapse and effective rank\.
[Graf et al\. 2021](https://arxiv.org/html/2608.20516#bib.bib11)show that the minimiser of the supervised contrastive loss maps every member of a class to a single point, with the classes arranged on a regular simplex\.[Papyan et al\. 2020](https://arxiv.org/html/2608.20516#bib.bib16)document the same terminal\-phase geometry under cross\-entropy, as neural collapse\.[Chen et al\. 2022](https://arxiv.org/html/2608.20516#bib.bib4)isolate class collapse as a distinct failure mode of supervised contrastive representations, and separate it from feature suppression\. All three are supervised: the collapsing partition is the label set\. Ours is the unsupervised analogue, with latent categories supplied by the masking scheme rather than by labels\. It is also total rather than partial\. By Prop\.[1](https://arxiv.org/html/2608.20516#Thmproposition1), it arises as a global optimum of a self\-referential objective rather than of a labeled one\. Dimensional collapse\([Jing et al\. 2022](https://arxiv.org/html/2608.20516#bib.bib15)\), effective rank\([Roy & Vetterli 2007](https://arxiv.org/html/2608.20516#bib.bib18)\)and RankMe\([Garrido et al\. 2023](https://arxiv.org/html/2608.20516#bib.bib10)\)are the standard summaries\. We supply an extreme counterexample: two configurations clear the conventionalrpool\>10r\_\{\\mathrm\{pool\}\}\>10threshold, are among the worst by probe accuracy, and do not retrieve \(App\.[A24](https://arxiv.org/html/2608.20516#A24)\)\.
##### Anisotropy, whitening, and machine\-extracted corpora\.
Frozen sentence embeddings are strongly anisotropic\([Ethayarajh 2019](https://arxiv.org/html/2608.20516#bib.bib7);[Gao et al\. 2019](https://arxiv.org/html/2608.20516#bib.bib8)\), and whitening is the standard remedy\([Su et al\. 2021](https://arxiv.org/html/2608.20516#bib.bib21)\)\. That literature instruments the candidate bank\. We prove that candidate\-side corrections cannot address a degenerate query bank \(Prop\.[3](https://arxiv.org/html/2608.20516#Thmproposition3)\)\. Separately, parameter\-free baselines are indispensable for correct attribution\([Gao et al\. 2021](https://arxiv.org/html/2608.20516#bib.bib9)\)\. We adopt the strongest form we could construct, together with a lexical bound\([Robertson & Zaragoza 2009](https://arxiv.org/html/2608.20516#bib.bib17)\)\. Representing scientific articles as typed reasoning graphs connects to scholarly knowledge graphs\([Auer et al\. 2020](https://arxiv.org/html/2608.20516#bib.bib2)\)\. Sec\.[9](https://arxiv.org/html/2608.20516#S9)adds a caution specific to LLM\-extracted scholarly corpora\.
## 3Setup and Protocol
##### Corpus and graph\.
We build one heterogeneous graph from58,14958\{,\}149scientific article records\.111[https://laion\.ai/notes/summaries/](https://laion.ai/notes/summaries/)\([Schuhmann et al\. 2025](https://arxiv.org/html/2608.20516#bib.bib19)\)Each scientific article forms one subgraph\. Every scientific article is decomposed into five typed reasoning nodes:*claim*,*method*,*result*,*evidence*and*implication*\. These are joined by typed edges that mirror scientific structure, for example,method→\\toresult→\\toclaimandclaim→\\toevidence\. Subgraphs are linked to one another through sharedfieldhubs and intra\-corpus citations\. App\.[A7](https://arxiv.org/html/2608.20516#A7)contains the full schema and all cardinalities\. Claims average4\.354\.35per scientific article, whereasmethodandresultare exact singletons \(m=1\.00m\{=\}1\.00\)\. That singleton property makes the pooling control in Sec\.[5](https://arxiv.org/html/2608.20516#S5)decisive: with one member per patch, mean pooling is provably the identity map\. Every node is embedded with a frozen sentence encoder and is never fine\-tuned\. Eachscientific articlenode carries one of488488field labels, which are used only by the probe and never as an input\.
Corpus sizes, deliberately distinguished\.The graph build requires full coverage and a valid subgraph, which gives57,90357\{,\}903maskable scientific articles\. The diagnostic harness requires only non\-empty aspect text, which gives58,14558\{,\}145of58,14958\{,\}149\. Chance MRR is1\.99×10−41\.99\\times 10^\{\-4\}for both\. The recoverable budgets agree to0\.0060\.006bits \(ℬmax=14\.379\\mathcal\{B\}^\{\\max\}=14\.379versus14\.38514\.385; Eq\.[4](https://arxiv.org/html/2608.20516#A8.E4)\)\. The two sets are therefore comparable in bits, rather than merely assumed to be\. Every table names the set and the sentence encoder it used \(App\.[A5](https://arxiv.org/html/2608.20516#A5)\)\.
##### Task and configurations\.
For each maskable subgraph, we hold out one present aspect and predict its patch embedding from the remaining aspects\. A heterogeneous GIN encoder\([Wang & Zhang 2022](https://arxiv.org/html/2608.20516#bib.bib22);[Hu et al\. 2020](https://arxiv.org/html/2608.20516#bib.bib14)\)produces node latents\. A pooling operator aggregates the target aspect’s nodes\. An EMA target encoder provides the target\. A transformer context mixer with a random\-walk structural encoding\([Dwivedi et al\. 2022](https://arxiv.org/html/2608.20516#bib.bib6)\)feeds a latent predictor\. Masking is leak\-safe: holding out an aspect removes every relation incident to the held\-out nodes, not only itshas\_aaedges\. This distinction turns out to be decisive in Sec\.[9](https://arxiv.org/html/2608.20516#S9), where retaining a single typed relation inflates a probe from0\.85350\.8535to1\.00001\.0000AUC\. We sweep six encoder configurations, four pooling operators, two losses, two target frames, three structural cues, and target\-branch variance regularisation\. We hold everything else fixed and use five seeds unless stated\.
Protocol R\(identical for every retrieval number in this paper\)\.*Queries:*4,0004\{,\}000papers, sampled once with seed00and held fixed across all models\. Lexical systems are∼500×\{\\sim\}500\\timescostlier per query, so they use a nested2,0002\{,\}000subsample that agrees to0\.010\.01bits \(App\.[A20](https://arxiv.org/html/2608.20516#A20)\)\.*Pool:*all masked\-aspect representations of the corpus, aspect\-matched and never in\-batch\. Hard pools of theKKnearest within\-category neighbours are also reported\.*Score:*cosine in the model’s own retrieval space\. Any retrieval frame is fitted on the candidate bank and applied to both sides\.*Metric:*bits recoveredℬ=1Nlog2N\!−𝔼\[log2r\]\\mathcal\{B\}=\\frac\{1\}\{N\}\\log\_\{2\}N\!\-\\mathbb\{E\}\[\\log\_\{2\}r\], with MRR and Hits@kkalongside\. Hereℬ=0\\mathcal\{B\}=0is*exactly*chance, andℬmax=1Nlog2N\!\\mathcal\{B\}^\{\\max\}=\\frac\{1\}\{N\}\\log\_\{2\}N\!is the achievable ceiling \(App\.[A8](https://arxiv.org/html/2608.20516#A8)\)\.*Uncertainty:*percentile bootstrap over queries \(B=2000B\{=\}2000\)\. Systems sharing queries are compared with apairedbootstrap on per\-queryΔ\(1/r\)\\Delta\(1/r\)\.*Null:*𝔼\[MRR\]=HN/N=1\.99×10−4\\mathbb\{E\}\[\\mathrm\{MRR\}\]=H\_\{N\}/N=1\.99\\times 10^\{\-4\}\. We run a permutation test with200200permutations, soppis bounded below by1/\(200\+1\)1/\(200\{\+\}1\)and is quoted asp<0\.005p<0\.005at that floor\.*No ratios:*atMRR≈1\.99×10−4\\mathrm\{MRR\}\\approx 1\.99\\times 10^\{\-4\}, ratios against chance are dominated by sampling noise\. We therefore report magnitudes as bit deficits against a measured ceiling\.*Instrument first:*no measurement is taken until a2020\-assertion self\-test passes\. It verifies that a perfect ranker returnsℬmax\\mathcal\{B\}^\{\\max\}, that a uniform ranker returns00, that a constant scorer returnsAUC=0\.500\\mathrm\{AUC\}=0\.500exactly, and that every masking condition genuinely isolates its targets \(App\.[A9](https://arxiv.org/html/2608.20516#A9)\)\.
##### Probe and pre\-registration\.
A linear classifier on the frozen pooled representation predicts the field label\. Splits are disjoint and label\-stratified, per seed, and majority\-class accuracy is below0\.050\.05\. Every threshold that converts a measurement into a verdict is fixed in a configuration file\. That file was written to disk before any result was computed \(App\.[A10](https://arxiv.org/html/2608.20516#A10)\), and was not revised afterward\.
## 4Why a Constant\-per\-Category Query Is a Global Optimum
Decompose each target into a category effect and an instance residual,
zp,a=μ\+αa⏟3options\+δp,a⏟57,903options\.z\_\{p,a\}\\;=\\;\\underbrace\{\\mu\+\\alpha\_\{a\}\}\_\{3\\ \\text\{options\}\}\\;\+\\;\\underbrace\{\\delta\_\{p,a\}\}\_\{57\{,\}903\\ \\text\{options\}\}\.\(1\)Retrieval requires recoveringδp,a\\delta\_\{p,a\}\. Whether the objective requires it turns on a distinction the standard analysis of latent\-predictive losses elides: in a JEPA the target is not exogenous data but the output of a trailing copy of the encoder being trained\. We therefore state the fixed\-point form first \(Prop\.[1](https://arxiv.org/html/2608.20516#Thmproposition1)\) and recover the familiar conditional\-mean statement as the special case in which the target is frozen \(Prop\.[2](https://arxiv.org/html/2608.20516#Thmproposition2)\)\. Proofs for all five results are in App\.[A11](https://arxiv.org/html/2608.20516#A11)\.
##### The coupled objective\.
Letfθf\_\{\\theta\}denote the context branch \(encoder, mixer, predictor\) andgθ¯g\_\{\\bar\{\\theta\}\}the target branch, whose parameters are an exponential moving average ofθ\\theta, soθ¯=θ\\bar\{\\theta\}=\\thetaat any stationary point\. Writeccfor the visible context together with the target designator \(the aspect cue of Sec\.[3](https://arxiv.org/html/2608.20516#S3)\) andttfor the held\-out subgraph\. The population objective is
R\(θ\)=𝔼‖fθ\(c\)−gθ¯\(t\)‖2,θ¯=EMA\(θ\)\.R\(\\theta\)\\;=\\;\\mathbb\{E\}\\,\\big\\\|f\_\{\\theta\}\(c\)\-g\_\{\\bar\{\\theta\}\}\(t\)\\big\\\|^\{2\},\\qquad\\bar\{\\theta\}=\\mathrm\{EMA\}\(\\theta\)\.\(2\)Both arguments are learned\. Eq\.[1](https://arxiv.org/html/2608.20516#S4.E1)treatszzas exogenous and the next proposition does not\.
###### Proposition 1\(Category\-measurable fixed points are global minima\.RRdoes not identifyδ\\delta\)\.
Leta\(⋅\)a\(\\cdot\)denote aspect identity, with\|𝒜\|=3\|\\mathcal\{A\}\|=3, and suppose the designator inccdeterminesa\(t\)a\(t\), as it does in our pipeline\. Then*\(i\)*for a*fixed*target branchgg, the risk\-minimising context branch isf⋆\(c\)=𝔼\[g\(t\)∣c\]f^\{\\star\}\(c\)=\\mathbb\{E\}\[g\(t\)\\mid c\]\.*\(ii\)*for*any*mapγ\\gammaon the finite category set𝒜\\mathcal\{A\}, the pairg=γ∘ag=\\gamma\\circ a,f=γ∘af=\\gamma\\circ aattainsR=0R=0and is therefore a*global*minimiser of Eq\.[2](https://arxiv.org/html/2608.20516#S4.E2); and*\(iii\)*wheneverccdeterminestt, the set of global minimisers also contains pairs in whichggis injective on instances\. ConsequentlyRRalone does not identify whetherδ\\deltais encoded: two global optima of the same objective differ byℬmax\\mathcal\{B\}^\{\\max\}bits of retrievable information\. The branch in*\(ii\)*yields a query bank of effective rank at most\|𝒜\|\|\\mathcal\{A\}\|while remaining*non\-constant*, so it is invisible to any collapse criterion that tests for constancy or for rank11\.
###### Proposition 2\(Conditional\-mean collapse: the frozen\-target special case\)\.
If the targetzzis exogenous and fixed, the population minimiser of𝔼‖f\(c\)−z‖2\\mathbb\{E\}\\\|f\(c\)\-z\\\|^\{2\}isf⋆\(c\)=𝔼\[z∣c\]f^\{\\star\}\(c\)=\\mathbb\{E\}\[z\\mid c\]; if in additionccdetermines the categoryaabut is*weakly*informative aboutδ\\delta, thenf⋆\(c\)=μ\+αa\+𝔼\[δ∣c\]→μ\+αaf^\{\\star\}\(c\)=\\mu\+\\alpha\_\{a\}\+\\mathbb\{E\}\[\\delta\\mid c\]\\to\\mu\+\\alpha\_\{a\}, the category centroid, with achieved loss𝔼‖δ‖2\\mathbb\{E\}\\\|\\delta\\\|^\{2\}\.
Our own oracle refutes the precondition of Prop\.[2](https://arxiv.org/html/2608.20516#Thmproposition2), which is why Prop\.[1](https://arxiv.org/html/2608.20516#Thmproposition1)is the operative statement\.A parameter\-free average of the visible aspects’ frozen features recovers\+14\.281\+14\.281of14\.38514\.385bits from exactly the contextccthe model is given \(Sec\.[6](https://arxiv.org/html/2608.20516#S6)\)\. On this corpusccis therefore strongly informative aboutδ\\delta—𝔼\[δ∣c\]≈δ\\mathbb\{E\}\[\\delta\\mid c\]\\approx\\delta, so against a frozen informative target the centroid would not be optimal and0\.000\.00bits would indicate mis\-optimisation\. What we observe is instead the category\-measurable branch of Prop\.[1](https://arxiv.org/html/2608.20516#Thmproposition1)*\(ii\)*, which is reachable only because the target is itself learned and can discardδ\\deltain step with the predictor\.
The inversion to internalise\.The model is neither underfitting nor mis\-optimising: it sits at a global minimum of Eq\.[2](https://arxiv.org/html/2608.20516#S4.E2), one of many, and the one that carries no instance information\. Our measurements agree: the achieved loss sits within a small constant of the predicted floor \(ℒ∞/𝔼‖δ‖2=3\.296\\mathcal\{L\}\_\{\\infty\}/\\mathbb\{E\}\\\|\\delta\\\|^\{2\}=3\.296\), the median predictor gradient norm is4\.61×10−14\.61\\times 10^\{\-1\}, and the regression cell of Sec\.[9](https://arxiv.org/html/2608.20516#S9)reaches1\.9×10−251\.9\\times 10^\{\-25\}while recovering2\.1%2\.1\\%of the ceiling loss value compatible only with the category\-measurable branch\.
##### Why this explains rung 4, which Prop\.[2](https://arxiv.org/html/2608.20516#Thmproposition2)cannot\.
Whitening the inputs raises the instance\-variance share25×25\\times\(Sec\.[5](https://arxiv.org/html/2608.20516#S5)\) and moves bits by00\. Under Prop\.[2](https://arxiv.org/html/2608.20516#Thmproposition2)that is anomalous: a largerδ\\deltasignal should raise𝔼\[δ∣c\]\\mathbb\{E\}\[\\delta\\mid c\]and hence the encoded identity\. Under Prop\.[1](https://arxiv.org/html/2608.20516#Thmproposition1)it is predicted: the zero\-risk family of*\(ii\)*depends on the inputs only througha\(c\)a\(c\), so it is invariant to any feature transformation that preserves category decodability\. Whitening preserves the probe still reads0\.7550\.755, so the fixed point survives\. What destroys the family is making the category answer unavailable in the target, which is exactly the repair of Sec\.[9](https://arxiv.org/html/2608.20516#S9)\.
###### Proposition 3\(Frame invariance of a degenerate query bank\)\.
Ifqi≡qq\_\{i\}\\equiv qfor allii, then for any score functionssand any injective transformTT,argsortjs\(T\(qi\),T\(cj\)\)\\operatorname\{argsort\}\_\{j\}s\(T\(q\_\{i\}\),T\(c\_\{j\}\)\)is the same permutation for everyii\. With gold items assigned uniformly at random,𝔼\[MRR\]=HN/N\\mathbb\{E\}\[\\mathrm\{MRR\}\]=H\_\{N\}/Nandℬ=0\\mathcal\{B\}=0exactly, independent ofTT\.
###### Proposition 4\(Geometry cannot help\)\.
For any metric space admitting a Fréchet mean, the minimiser of𝔼d\(f\(c\),z\)2\\mathbb\{E\}\\,d\(f\(c\),z\)^\{2\}\. The Fréchet mean ofp\(z∣c\)p\(z\\mid c\)is the arithmetic mean underℓ2\\ell\_\{2\}and the mean direction under cosine\. The Karcher mean inℍn\\mathbb\{H\}^\{n\}with the unique Fréchet mean on any Hadamard manifold\.
###### Proposition 5\(Structural reducibility of a census\-determined graph\)\.
LetGGhave edge setEEand let𝒩\\mathcal\{N\}be its node\-type census\. If the sub\-edge set incident to every subgraph is a deterministic function of𝒩\\mathcal\{N\}alone\. So thatEEis constant across all subgraphs with the same census\. Then for any message\-passing encoderΦ\\Phiof any depth, the patch representationΦ\(G\)p\\Phi\(G\)\_\{p\}is a fixed function of the feature vectors ofpp’s own nodes\. No scoring rule built onΦ\\Phican access information absent from those features\.*The bound is on information, not performance:*a trained encoder may still exceed a fixed training\-free statistic of the same features by learning a better metric over them, and such a gain is a re\-metrisation rather than structural learning\.
Hyperbolic space is the first failure mode: as a Hadamard manifold, its Fréchet mean is guaranteed to exist and be unique, whereas Euclidean space has conditional distributions with ill\-defined means\. Curving the space renames the centroid \(App\.[A12](https://arxiv.org/html/2608.20516#A12)\)\. Jointly, the five results predict that retrieval fails under any regression\-style latent objective regardless of target geometry\. The failures under a contrastive objective whose negatives do not demand within\-category resolution \(Sec\.[8](https://arxiv.org/html/2608.20516#S8)\) and cannot be recovered post hoc by any frame, metric or manifold once the query bank is degenerate\.
Table 1:The diagnostic ladder\(57,90357\{,\}903set, Protocol R, full pool\)\.ℬ\\mathcal\{B\}is bits recovered ofℬmax=14\.379\\mathcal\{B\}^\{\\max\}=14\.379\(Eq\.[4](https://arxiv.org/html/2608.20516#A8.E4)\);00is exactly chance\.rcand,rqueryr\_\{\\mathrm\{cand\}\},r\_\{\\mathrm\{query\}\}are the effective ranks of the candidate bank and of the model’s*predictions*\. Rung 7 is the decisive control; rungs 8–10 are the ceilings of Sec\.[6](https://arxiv.org/html/2608.20516#S6); rung 11 is the repaired configuration of Sec\.[9](https://arxiv.org/html/2608.20516#S9), included here so the whole trajectory is in one place\. Per\-rung detail: App\.[A13](https://arxiv.org/html/2608.20516#A13)\.\#Interventionrnoder\_\{\\mathrm\{node\}\}rcandr\_\{\\mathrm\{cand\}\}rqueryr\_\{\\mathrm\{query\}\}ℬ\\mathcal\{B\}Diagnosis0mean pool, cosine target*\(baseline\)*18\.318\.318\.118\.11\.91\.90\.000\.00p=0\.98p=0\.98: chance1\+\+target VICReg———0\.000\.00no effect2infonce, in\-batch negatives———0\.000\.00no effect3sum/deepsets/attnpooling———0\.000\.00no effect4\+\+input whitening*\(instance var\.×25\\times 25\)*———0\.000\.00rank↑\\uparrow,ℬ\\mathcal\{B\}flat5retrieval frame: centre / rm\-topkk/ ZCA———0\.40\.40\.4bits0\.4\\text\{ bits\}on14\.4bits14\.4\\text\{ bits\}6encoder depth0→30\\to 3→18\.054\.9\\\!\\to\\\!18\.0——0\.000\.00depth not binding7*singleton control*\(method,m=1\.00m\{=\}1\.00\)47\.447\.447\.347\.31\.97\\mathbf\{1\.97\}0\.000\.00pooling is identity8p1*positive control*———\+14\.220\+14\.220harness can learn9oracle*\(no encoder, no training\)*—305\.0305\.0—\+14\.281\+14\.281task is solvable10bm25*\(no embeddings at all\)*———\+14\.335\+14\.335task is easy11repaired objective, 20k, 3 seeds———\+14\.377\\mathbf\{\+14\.377\}100\.0%100\.0\\%; see §[9](https://arxiv.org/html/2608.20516#S9)*Chance under Protocol R*0\.000\.00*Training\-free oracle, block A*13\.86513\.865*Ceiling*ℬmax=1Nlog2N\!\\mathcal\{B\}^\{\\max\}=\\frac\{1\}\{N\}\\log\_\{2\}N\!14\.37914\.379
## 5The Diagnostic Ladder
##### The symptom is invariant to objective, aggregator and depth \(rungs 0–3, 6\)\.
The baseline attains probe0\.8710\.871andℬ=0\.00\\mathcal\{B\}=0\.00atp=0\.98p=0\.98\. Target\-branch variance regularisation, InfoNCE over standard in\-batch negatives, and sum or deepsets or attention pooling each change nothing\. Sum pooling lowers pooled rank to1\.251\.25while raising the probe to0\.9630\.963: rank and probe move in opposite directions and retrieval ignores both\. Across depths00–33, node\-level rank falls monotonically→18\.054\.9\\\!\\to\\\!18\.0message passing does compress while bits do not move\. Also, a depth\-00linear encoder fails identically to a depth\-33GNN\. Over\-smoothing is real and irrelevant\.
##### Rank restoration is decoupled from task success \(rung 4\)\.
Whitening the inputs raises the instance\-variance share from0\.39%0\.39\\%to9\.86%9\.86\\%\. A25×25\\timesincrease in the gradient incentive to encode identity raises input effective rank to310\.9310\.9of384384\. Bits stay at00\. The probe meanwhile falls from0\.8710\.871to0\.7550\.755, because whitening removes precisely the shared\-mean component it was exploiting\. This is the rung Prop\.[1](https://arxiv.org/html/2608.20516#Thmproposition1)predicts and Prop\.[2](https://arxiv.org/html/2608.20516#Thmproposition2)cannot\.
##### The failure is query\-side, and re\-metrisation confirms it \(rung 5\)\.
Five frames fitted on the candidate bank and applied to both sides span1\.7×10−41\.7\\times 10^\{\-4\}\(raw\) to2\.6×10−42\.6\\times 10^\{\-4\}\(all\-but\-the\-top\-11\):0\.4bits0\.4\\text\{ bits\}against a14\.4bits14\.4\\text\{ bits\}deficit, as Prop\.[3](https://arxiv.org/html/2608.20516#Thmproposition3)predicts\. The accompanying bank statistics are the point of the experiment: the candidate bank has DC ratio37\.637\.6atrcand=18\.1r\_\{\\mathrm\{cand\}\}=18\.1, while the query bank has DC ratio162\.2162\.2, DC energy0\.99990\.9999and mean pairwise self\-similarity0\.0040\.004atrquery=1\.9r\_\{\\mathrm\{query\}\}=1\.9of128128\. Four thousand distinct questions produce essentially one vector, and that number is a quantitative check on the theory rather than a coincidence: Prop\.[1](https://arxiv.org/html/2608.20516#Thmproposition1)*\(ii\)*bounds the query bank of the degenerate branch aterank≤\|𝒜\|=3\\operatorname\{erank\}\\leq\|\\mathcal\{A\}\|=3, and we measurerquery=1\.9r\_\{\\mathrm\{query\}\}=1\.9\. No transformation of a near\-constant query can rank anything, and changing the objective so the query is not constant can, and does \(Sec\.[9](https://arxiv.org/html/2608.20516#S9)\)\.
##### Pooling is exonerated by a provable control \(rung 7\)\.
Method patches havem=1\.00m\{=\}1\.00, so mean pooling is the identity map: it carriesrnode=47\.42r\_\{\\mathrm\{node\}\}=47\.42torpool=47\.31r\_\{\\mathrm\{pool\}\}=47\.31, rank preserved to two decimals\. The model’s predictions for those same patches nevertheless occupy effective rank1\.971\.97, and the gold item’s standardised similarity margin isz=−0\.001z=\-0\.001\. The correct answer is indistinguishable from the pool mean\. Across aspects,rpoolr\_\{\\mathrm\{pool\}\}spans2\.052\.05to47\.3147\.31\(23×23\\times\) with no movement in bits \(App\.[A18](https://arxiv.org/html/2608.20516#A18)\)\.
## 6Three Upper Bounds: Task, Harness and Objective Are Separable
The ladder establishes what collapses but not whether the task is hard, whether the harness works, or whether the objective is at fault\. We separate these with three controls under Protocol R on the58,14558\{,\}145set\.oracleforms the query as the mean of the raw frozen embeddings of the subgraph’s other aspects and ranks by cosine without any encoder, objective, or training\.bm25is an Okapi BM25 index over the raw aspect texts\([Robertson & Zaragoza 2009](https://arxiv.org/html/2608.20516#bib.bib17)\), withbm25\-no\-ovdeleting every55\-gram shared between query and gold\.p1is a same\-architecture latent predictor trained on the same frozen features with the same masking, pool and scoring code, whose purpose is to establish that the harness can recover identity\.
Table 2:Bits recovered against pool difficulty\(58,14558\{,\}145set\)\. Hard pools of sizeKKare theKKnearest within\-category neighbours of the gold item, so difficulty increases downward; the pool hasN=K\+1N=K\{\+\}1items andℬmax=1Nlog2N\!\\mathcal\{B\}^\{\\max\}=\\frac\{1\}\{N\}\\log\_\{2\}N\!\(Eq\.[4](https://arxiv.org/html/2608.20516#A8.E4)\), which lies1\.44271\.4427bits below the index entropylog2N\\log\_\{2\}N\. Vector systems use4,0004\{,\}000queries, lexical systems a nested2,0002\{,\}000subsample\. All four systems are within1\.2%1\.2\\%of the ceiling at every rung; the baseline Graph\-JEPA recovers0\.0%0\.0\\%and the repaired configuration of Sec\.[9](https://arxiv.org/html/2608.20516#S9)recovers100\.0%100\.0\\%\.KKℬmax\\mathcal\{B\}^\{\\max\}bm25bm25\-no\-ovoraclep1220\.8620\.862\+0\.860\+0\.860\+0\.850\+0\.850\+0\.840\+0\.840\+0\.830\+0\.83010102\.2952\.295\+2\.284\+2\.284\+2\.276\+2\.276\+2\.260\+2\.260\+2\.240\+2\.2401001005\.2625\.262\+5\.240\+5\.240\+5\.213\+5\.213\+5\.198\+5\.198\+5\.167\+5\.167100010008\.5318\.531\+8\.501\+8\.501\+8\.445\+8\.445\+8\.455\+8\.455\+8\.421\+8\.421full14\.38514\.385\+14\.335\+14\.335\+14\.250\+14\.250\+14\.281\+14\.281\+14\.220\+14\.220full pool, % ofℬmax\\mathcal\{B\}^\{\\max\}99\.7%\\mathbf\{99\.7\\%\}99\.1%99\.1\\%99\.3%99\.3\\%98\.9%98\.9\\%full\-pool MRR / R@10\.99100\.9910/0\.98850\.98850\.96970\.9697/0\.96150\.96150\.97190\.9719/0\.95930\.95930\.95500\.9550/0\.93800\.938095%95\\%CI on MRR\[0\.9868,0\.9944\]\[0\.9868,0\.9944\]\[0\.9624,0\.9763\]\[0\.9624,0\.9763\]\[0\.9674,0\.9761\]\[0\.9674,0\.9761\]\[0\.9497,0\.9604\]\[0\.9497,0\.9604\]*Graph\-JEPA, baseline*0\.000\.00bits=0\.0%=\\mathbf\{0\.0\\%\}ofℬmax\\mathcal\{B\}^\{\\max\}\(p=0\.98p=0\.98; Table[1](https://arxiv.org/html/2608.20516#S4.T1)\)*Graph\-JEPA, repaired*\+14\.377\+14\.377bits=100\.0%=\\mathbf\{100\.0\\%\}ofℬmax\\mathcal\{B\}^\{\\max\}\(Table[6](https://arxiv.org/html/2608.20516#S8.T6)\)##### Result 1: The task is solvable, and the harness works\.
All three controls recover between98\.9%98\.9\\%and99\.7%99\.7\\%of the14\.38514\.385recoverable bits at the full pool, with a total spread of0\.1150\.115bits\. In particular,p1recovers\+14\.220\+14\.220bits through the identical masking, pool, and scoring code\. This excludes the corpus, the task definition, the pool construction, the masking, and the metric as explanations of the null\. It also establishes the premise of Prop\.[1](https://arxiv.org/html/2608.20516#Thmproposition1)*\(iii\)*: the context determines the target well enough that an identity\-preserving global optimum exists\.
Whatp1does and does not license\.p1differs from the model under test in two factors at once: \(i\) frozen text features rather than hetero node features, and \(ii\) a reference predictor rather than the full stack\. It therefore bounds the harness and does not by itself attribute the gap\. The pre\-registered2×22\\times 2factorial that separates the two factors: \(i\) in which a cell is unable to obtain the trainer or \(ii\) features it requested emits no verdict, is in App\.[A21](https://arxiv.org/html/2608.20516#A21)\.
##### Result 2: Training costs measurable identity, monotonically in difficulty\.
p1sits beloworacleat every rung, and the paired bootstrap on per\-queryΔ\(1/r\)\\Delta\(1/r\)separates them with non\-overlapping intervals at all four pool sizes:−0\.020\-0\.020bits atK=10K\{=\}10,−0\.031\-0\.031atK=100K\{=\}100,−0\.034\-0\.034atK=1000K\{=\}1000and−0\.061\-0\.061\(the full pool\)\. Also the R@1 falling is→0\.93800\.9593\\\!\\to\\\!0\.9380\. Even a predictor that mostly works for training, and more as the pool hardens\.
##### Result 3: Difficulty is nearly flat in pool size, and this is a warning\.
FromK=10K\{=\}10to the full58,14558\{,\}145item pool and a5,800×5\{,\}800\\timeschange\. Also, the BM25’s MRR falls only→0\.99100\.9960\\\!\\to\\\!0\.9910\. A near\-constant bit deficit across that range is the signature of a task in which the gold item is separated by a large margin\. Sec\.[9](https://arxiv.org/html/2608.20516#S9)shows why that matters beyond instrument calibration: \(i\) the same property makes the task reducible, and \(ii\) therefore makes a near\-ceiling score uninformative about structure\.
## 7The Mechanism: Which Variance Is Encoded, and When
Prop\.[1](https://arxiv.org/html/2608.20516#Thmproposition1)says a category\-measurable solution is a global optimum\. Prop\.[2](https://arxiv.org/html/2608.20516#Thmproposition2)’s precondition is refuted by our oracle approach\. What remains to be established empirically is \(i\) which variance component the converged representation carries, and \(ii\) whether the allocation is produced by training or is already present at initialisation\.
Table 3:The variance allocation\.Share of total variance attributable to aspect identity and to paper identity\. Columns 1 and 4 are block B \(MPNet\-768768\); columns 2 and 3 are block A \(MiniLM\-384384\), i\.e\. the*same*encoder that produced the trained latents, so the central comparison is*within*one feature space and the cross\-encoder pair is retained only for continuity with the three ceilings\. The two groupings are separate one\-way decompositions and need not sum to100%100\\%\.Variance attributable tofrozen inputs\(block B\)frozen inputs\(block A\)trained latents\(block A\)inputs, de\-templated\(block B\)aspect identity \(ρasp\\rho\_\{\\mathrm\{asp\}\}\)0\.40%0\.40\\%\-99\.61%\\mathbf\{99\.61\\%\}0\.00000\.0000paper identity \(ρpap\\rho\_\{\\mathrm\{pap\}\}\)86\.05%\\mathbf\{86\.05\\%\}\-0\.39%0\.39\\%—bits recovered on this representation\+14\.281\+14\.281\(oracle\)13\.86513\.865\(oracle\)0\.000\.00\(trained\)\+14\.28\+14\.28% of the matchingℬmax\\mathcal\{B\}^\{\\max\}99\.3%99\.3\\%96\.4%96\.4\\%0\.0%0\.0\\%99\.3%99\.3\\%The finding, stated as an endpoint comparison\.The frozen inputs allocate86\.05%86\.05\\%of their variance to subgraph identity and0\.40%0\.40\\%to aspect identity\. The trained latents allocate0\.39%0\.39\\%and99\.61%99\.61\\%\. The identity is present at input, as three independent controls confirm, and absent from the trained representation, by a factor of roughly250×250\\timesin each direction\. This is a comparison of two endpoints and by itself says nothing about the trajectory between them\. Sec\.[7\.1](https://arxiv.org/html/2608.20516#S7.SS1)measures them separately, because the two readings support different claims\.
##### How much of the collapse does the aspect share explain?
We simulate the retrieval problem at the measured operating point\. Generating targets aszp,a=ραa\+1−ρS1/2vp\+εηz\_\{p,a\}=\\sqrt\{\\rho\}\\,\\alpha\_\{a\}\+\\sqrt\{1\-\\rho\}\\,S^\{1/2\}v\_\{p\}\+\\varepsilon\\etawithSjj∝j−κS\_\{jj\}\\propto j^\{\-\\kappa\}, whereρ\\rhois the aspect share,κ\\kappathe measured power\-law decay of the identity covariance \(κ^=0\.942\\hat\{\\kappa\}=0\.942\) andε\\varepsilonthe predictor’s irreducible error estimated from the measured loss floor \(ε^=1\.515\\hat\{\\varepsilon\}=1\.515\) \(bank size20,00020\{,\}000, soℬmax=12\.845\\mathcal\{B\}^\{\\max\}=12\.845\)\. Full sweep in App\.[A22](https://arxiv.org/html/2608.20516#A22)\.
Table 4:Phase analysis at the measured operating point, atκ^=0\.942\\hat\{\\kappa\}=0\.942andε^=1\.515\\hat\{\\varepsilon\}=1\.515, bank size20,00020\{,\}000soℬmax=12\.845\\mathcal\{B\}^\{\\max\}=12\.845\(Eq\.[4](https://arxiv.org/html/2608.20516#A8.E4)\)\. The trained representation’s measured share \(99\.61%99\.61\\%\) predicts the loss of11\.04511\.045of12\.84512\.845bits; the inputs’ share \(0\.40%0\.40\\%\) predicts no loss, which is what the three controls observe\. Bits lost is12\.845−ℬ12\.845\-\\mathcal\{B\}and is monotone inρ\\rhoby construction\. Critical shareρ⋆=0\.99999\\rho^\{\\star\}=0\.99999\(1−ρ⋆=5\.1×10−61\-\\rho^\{\\star\}=5\.1\\times 10^\{\-6\}\); removingκ^\\hat\{\\kappa\}orε^\\hat\{\\varepsilon\}atρ=0\.9961\\rho=0\.9961moves MRR by<0\.004<0\.004\.aspect shareρ\\rho0\.00400\.00400\.90\.90\.990\.990\.9961\\mathbf\{0\.9961\}0\.9990\.9990\.99990\.9999MRR1\.0001\.0000\.8930\.8930\.1070\.1070\.0171\\mathbf\{0\.0171\}0\.00430\.00430\.00140\.0014ℬ\\mathcal\{B\}\+12\.845\+12\.845\+12\.200\+12\.200\+3\.833\+3\.833\+1\.80\\mathbf\{\+1\.80\}\+0\.843\+0\.843\+0\.266\+0\.266bits lost0\.0000\.0000\.6450\.6459\.0129\.01211\.045\\mathbf\{11\.045\}12\.00212\.00212\.57912\.579% ofℬmax\\mathcal\{B\}^\{\\max\}lost0\.0%0\.0\\%5\.0%5\.0\\%70\.2%70\.2\\%𝟖𝟔%\\mathbf\{86\\%\}93\.4%93\.4\\%97\.9%97\.9\\%
##### The allocation is dominant, and still not sufficient\.
Three statements in decreasing strength:*\(i\)*The aspect share is the dominant term, atρ=0\.0040\\rho=0\.0040retrieval is perfect and atρ=0\.9961\\rho=0\.9961it loses11\.04511\.045of12\.84512\.845bits, such as the allocation alone destroys86%86\\%of the recoverable identity\.*\(ii\)*It is not the only term: the baseline pipeline loses100%100\\%, and total collapse in simulation requiresρ≥ρ⋆=0\.99999\\rho\\geq\\rho^\{\\star\}=0\.99999, above the measured value\. The residual14%14\\%is unexplained\.*\(iii\)*Anisotropy is not the residual factor, removingκ^\\hat\{\\kappa\}orε^\\hat\{\\varepsilon\}at the measuredρ\\rhochanges MRR by less than0\.0040\.004, so leave\-one\-out rejects the anisotropy hypothesis we initially favoured\. Sec\.[9](https://arxiv.org/html/2608.20516#S9)then confirms the direction of*\(i\)*in the real feature space by intervening on the objective rather than simulating it\.
### 7\.1Is the allocation performed by training, or present at initialisation?
The endpoint comparison above is compatible with two mechanisms, and Fig\.[10](https://arxiv.org/html/2608.20516#A32.F10)\(b\) shows pooled effective rank at its final value at epoch00with flat thereafter\. App\.[A29](https://arxiv.org/html/2608.20516#A29)shows the same compression at initialisation across all node types\. We therefore measure the variance allocation itself at every checkpoint with one forward pass each, on the block\-A encoder rather than inferring the dynamics from the endpoints\. The protocol is in App\.[A15](https://arxiv.org/html/2608.20516#A15)\.
Table 5:Variance allocation over training\(block A, within one encoder,ℬ\\mathcal\{B\}againstℬmax=14\.379\\mathcal\{B\}^\{\\max\}=14\.379\)\. One forward pass per checkpoint;ρpap\\rho\_\{\\mathrm\{pap\}\}andρasp\\rho\_\{\\mathrm\{asp\}\}are the same one\-way decompositions as Table[3](https://arxiv.org/html/2608.20516#S7.T3), andrqueryr\_\{\\mathrm\{query\}\}is the effective rank of the model’s predictions\. Step00is the randomly initialised encoder before any gradient step\.stepρpap\\rho\_\{\\mathrm\{pap\}\}ρasp\\rho\_\{\\mathrm\{asp\}\}rqueryr\_\{\\mathrm\{query\}\}ℬ\\mathcal\{B\}2020k \(final\)0\.39%0\.39\\%99\.61%99\.61\\%1\.91\.90\.000\.00##### Reading \(A\): The allocation shifts during training\.
Ifρpap\\rho\_\{\\mathrm\{pap\}\}falls over the run, the endpoint comparison is a genuine re\-allocation\. Also, the pooled\-rank flatness of Fig\.[10](https://arxiv.org/html/2608.20516#A32.F10)\(b\) is a separate, weaker statement: rank is a spectral\-entropy summary that a category\-measurable solution can satisfy at initialisation while the allocation still moves\. After that, we re\-allocate the training variance away from instance identity, and Table[5](https://arxiv.org/html/2608.20516#S7.T5)is the evidence for it\.
##### Reading \(B\): The degeneracy precedes training, and the objective does not remove it\.
Ifρpap\\rho\_\{\\mathrm\{pap\}\}is already at its final level at step00, then training does not perform an inversion: the category measurable configuration is where the architecture starts\. The objective supplies no gradient that would leave it, because by Prop\.[1](https://arxiv.org/html/2608.20516#Thmproposition1)*\(ii\)*the configuration is already a global minimum\. This is the weaker claim, and it is the one our rank trajectory independently supports\. It does not weaken the analysis because the endpoint comparison, the three ceilings and the single\-variable loss ablation are unaffected, but it changes the mechanism from "training destroys identity" to "the objective has no incentive to build it," and only the latter is licensed by a flat rank curve\.
##### Either way, the loss is still the lever\.
Both readings are consistent with the intervention of Sec\.[9](https://arxiv.org/html/2608.20516#S9), changing only the loss moves the pipeline→0\.30714\.359\\\!\\to\\\!0\.307bits, and the repaired configuration’s bits rise from−0\.07\-0\.07at step11to\+14\.24\+14\.24by step500500\(App\.[A14](https://arxiv.org/html/2608.20516#A14)\)\. An objective that admits the degenerate fixed point retains it, and another objective that removes the category answer from the target does not\.
## 8What the Task Measures and What We Tried
##### The task is largely lexical, and strongly templated\.
Deleting every55\-gram shared between query and gold costs BM25 only0\.0850\.085bits \(\+→\+14\.250\+14\.335\\\!\\to\\\!\+14\.250\)\.bm25\-no\-ovstill ties the embedding oracle: the redundancy among a subgraph’s claims, methods and results is distributional and topical, not copy\-paste\. A logistic classifier restricted to function words only identifies which aspect a text is at0\.9370\.937against chance0\.3330\.333\(templating index0\.9060\.906;0\.9860\.986with full TF\-IDF\)\. The most frequent44\-grams cover25\.725\.7–39\.0%39\.0\\%of subgraph per aspect\. That figure matters twice over: it is a caveat on the instrument, and it is the empirical reason the designator\-determines\-category premise of Prop\.[1](https://arxiv.org/html/2608.20516#Thmproposition1)holds so strongly here\. The category answer is available from surface form alone\. Crucially, after removing the per\-aspect mean, the aspect share falls→0\.00000\.0040\\\!\\to\\\!0\.0000while the training\-free oracle still recovers\+14\.28\+14\.28bits atMRR=0\.972\\mathrm\{MRR\}=0\.972\. Templating inflates aspect\-type separability, not subgraph identity\. Full audit in App\.[A27](https://arxiv.org/html/2608.20516#A27)\.
A good instrument and a poor benchmark\.Masked\-aspect retrieval is redundancy\-based subgraph identification\. But BM25 deliberately recovers99\.7%99\.7\\%of the ceiling, and its easiness is exactly what makes0\.000\.00bits informative\. We make no claim that performance on it measures scientific reasoning\. Sec\.[9](https://arxiv.org/html/2608.20516#S9)strengthens this from a caveat into a proof\.
##### Seven interventions, one flat line\.
Props\.[1](https://arxiv.org/html/2608.20516#Thmproposition1)–[2](https://arxiv.org/html/2608.20516#Thmproposition2)identify two levers: \(i\) the set aggregator and \(ii\) the loss\. We test the full grid at0\.360\.36additional GPU\-hours\. Across all seven cellsrpoolr\_\{\\mathrm\{pool\}\}spans1\.251\.25–2\.032\.03and the probe spans0\.6040\.604–0\.9640\.964, while bits recovered stay at00in every cell\. The best cell \(attention/InfoNCE,MRR=2\.3×10−4\\mathrm\{MRR\}=2\.3\\times 10^\{\-4\}\) is within sampling noise of the baseline against a14\.4bits14\.4\\text\{ bits\}deficit \(App\.[A23](https://arxiv.org/html/2608.20516#A23)\)\. The reason InfoNCE alone does not help is instructive, and is exactly Prop\.[1](https://arxiv.org/html/2608.20516#Thmproposition1)*\(ii\)*in operational form\. Also, with the aspect\-mixed in\-batch negatives and winning the contrastive game requires only answering \(such as "claim, method, or result?"\)\. A category\-measurable branch therefore still attains the optimum\. Negatives must be drawn from within the dominant categorical partition of the target space, or the objective is solved by a category classifier\. Sec\.[9](https://arxiv.org/html/2608.20516#S9)tests the corollary and pairs InfoNCE with a target frame that removes the per\-aspect mean, so that the category answer is no longer available in the target, and the collapse resolves\.
##### Competing explanations and remaining limitations\.
Each explanation below is answered by a control that shares the pipeline under evaluation\.*The frozen encoder is not responsible\.*The oracle consumes exactly these embeddings and recovers\+14\.281\+14\.281bits, and BM25 consumes none and recovers\+14\.335\+14\.335\. Also, whitening raises effective rank and lowers probe accuracy while leaving bits at zero\.*The harness and the metric are sound\.*p1recovers\+14\.220\+14\.220bits through the identical masking, pooling, and scoring code with a cheat query returnsMRR=1\.000\\mathrm\{MRR\}=1\.000\. The4,0004\{,\}000query subsample reproduces the full\-corpus baseline to0\.010\.01bits, and a2020\-assertion self\-test verifies every estimator before it is used \(App\.[A9](https://arxiv.org/html/2608.20516#A9)\)\.*The ceiling is correctly specified\.*Monte Carlo verifies the chance term of Eq\.[3](https://arxiv.org/html/2608.20516#A8.E3)at every pool size \(App\.[A8](https://arxiv.org/html/2608.20516#A8)\)\.*Corpus redundancy is not the limiting factor\.*A controlled bank with DC ratio up to100100remains perfectly rankable given an informative query \(App\.[A26](https://arxiv.org/html/2608.20516#A26)\)\.*Pooling does not destroy identity\.*On them=1m\{=\}1singleton patches mean pooling is provably the identity map, and the collapse is unchanged\.*The retrieval geometry is not the operative variable\.*Post\-hoc re\-metrisation buys0\.4bits0\.4\\text\{ bits\}, as Props\.[3](https://arxiv.org/html/2608.20516#Thmproposition3)–[4](https://arxiv.org/html/2608.20516#Thmproposition4)predict, whereas changing the objective moves the pipeline by the entire recoverable budget \(Sec\.[9](https://arxiv.org/html/2608.20516#S9)\)\.*The account is not constructed after the fact\.*Prop\.[1](https://arxiv.org/html/2608.20516#Thmproposition1)*\(ii\)*bounds the degenerate query bank aterank≤3\\operatorname\{erank\}\\leq 3independently of any measurement, and we observerquery=1\.9r\_\{\\mathrm\{query\}\}=1\.9\. It also predicts the rung\-4 whitening null, which Prop\.[2](https://arxiv.org/html/2608.20516#Thmproposition2)does not\. Two limitations remain to be addressed\. Our re\-implementation does not reproduce the original method’s published numbers0\.6860\.686on MUTAG against0\.8740\.874, and0\.7390\.739on PROTEINS against0\.7500\.750\(App\.[A28](https://arxiv.org/html/2608.20516#A28)\)\. All results come from a single corpus, which is the most serious remaining threat to generality\.
Table 6:Eleven cells at matched schedule\(block C; all share the graph cache, seed and evaluation;ℬ\\mathcal\{B\}againstℬmax=14\.379\\mathcal\{B\}^\{\\max\}=14\.379; A1 is a held\-out reasoning probe with a pre\-registered void line at0\.700\.70, defined in App\.[A16](https://arxiv.org/html/2608.20516#A16)\)\. Theregrow is the decisive single\-variable result: only the loss differs from thecenter:0/nce/33k row\. Bits span11\.808−14\.37911\.808\-14\.379and A1 spans0\.711−0\.9850\.711\-0\.985at Spearmanρ=−0\.24\\rho=\-0\.24\(n=10n=10, not significant\), so the two metrics show no reliable relationship—this paper’s thesis restated on the repaired model\. We claim no trade\-off and no direction\.Configlossstepsℬ\\mathcal\{B\}A1Readingraw\-skip gatence6k14\.37914\.3790\.8750\.875gate→0\.3886\\to 0\.3886: encoder usedcenter:1nce3k14\.37914\.3790\.9630\.963center:0, mainnce20k14\.377±0\.00214\.377\\pm 0\.0020\.753±0\.0100\.753\\pm 0\.01033seedscenter:0, aspect cuence6k14\.37614\.3760\.8000\.800center:0, inductivence6k14\.37614\.3760\.7690\.769transduction gap\+0\.001\+0\.001center:0nce3k14\.359\\mathbf\{14\.359\}0\.8620\.862raw:0*\(baseline\)*nce3k14\.28914\.2890\.9640\.96499\.4%99\.4\\%of ceilingcenter:0, rwse cuence6k14\.20914\.2090\.7110\.711center:0, no cuence6k13\.52913\.5290\.9070\.907worst bits, 2nd\-best A1raw:1nce3k11\.80811\.8080\.985\\mathbf\{0\.985\}worst bits, best A1center:0reg3k0\.307\\mathbf\{0\.307\}0\.567\\mathbf\{0\.567\}collapse;14\.0514\.05\-bit swing*training\-free oracle*13\.86513\.865—96\.4%96\.4\\%of ceiling*ceiling*14\.37914\.379—0\.5140\.514bits headroomTable 7:The data\-derived target and its pre\-registered gates\(block C,33seeds, paper\-grouped splits\)\.*Upper:*the gates, with thresholds fixed before measurement\.*Lower:*the probe conditions, which we report even though the gates failed, because two of them are independently informative\. Every threshold appears in App\.[A10](https://arxiv.org/html/2608.20516#A10)\.Gate / conditionMeasuredThresholdVerdictG1 exact\-duplicate share ofchallenged\_by25\.96%25\.96\\%≤20%\\leq 20\\%FAILG1b genericness gap \(challenge−\-support\)\+0\.2377\+0\.2377≤0\.15\\leq 0\.15FAILG0 census,challenged\_byedges121,516121\{,\}516≥2,000\\geq 2\{,\}000passimpliesduplicate share0\.00%0\.00\\%≤20%\\leq 20\\%passpaper\-identity leak \(claim\- vs paper\-grouped\)\+0\.0003\+0\.0003—*absent*shuffled\-label control0\.49750\.4975≈0\.500\\approx 0\.500passcosine\-only floor0\.85390\.8539—similarity aloneraw pair features, paper\-grouped0\.95610\.9561—the barmodel, target edgespresent1\.00001\.0000—void:\+0\.1465\+0\.1465edge\-type leakmodel, target edgesmasked0\.85350\.8535—−0\.0004\-0\.0004vs\. the cosine floor
## 9The Repair Succeeds, and the Metric Stops Meaning Anything
Sec\.[8](https://arxiv.org/html/2608.20516#S8)ends with a falsifiable prediction\. Remove the per\-aspect mean from the target during training, so that the category answer is unavailable, and pair this with in\-batch negatives\. This section reports the confirmation, then explains why it licenses far less than it appears to\.
##### Result 1: The objective is the lever, and one variable proves it\.
We held the target frame, structural budget, schedule, and seed fixed\. We then changed only the loss, from InfoNCE to regression\. Bits fall from14\.35914\.359to0\.3070\.307, or2\.1%2\.1\\%of14\.37914\.379\. The reasoning probe falls to0\.5670\.567, the majority class, andMRR=0\.0009\\mathrm\{MRR\}=0\.0009\. That is a14\.0514\.05\-bit swing on a single variable\. The regression run is the informative part\. It reached a training loss of1\.9×10−251\.9\\times 10^\{\-25\}while recovering only2\.1%2\.1\\%of the ceiling\. Near\-zero risk is therefore attained on a branch that carries no instance information, as Prop\.[1](https://arxiv.org/html/2608.20516#Thmproposition1)*\(ii\)*permits\. The loss value alone does not identify what was learned\. One consequence is worth stating separately\. A regression loss near zero indicates collapse; an InfoNCE loss near zero indicates discrimination\. The two magnitudes are not comparable\.
##### Result 2: The architectural factors we swept are small, and the schedule is not\.
The factors we varied inside the architecture matter less than the one we had treated as a nuisance\. The matched\-budget2×22\\times 2\(App\.[A14](https://arxiv.org/html/2608.20516#A14)\) gives a target\-frame main effect of\+0\.071\+0\.071bits\. The shared\-basis effect is−2\.481\-2\.481, and the interaction is\+2\.500\+2\.500\. The structural pattern is worth\+0\.167\+0\.167\(aspect vs\. RWSE\) and\+0\.680\+0\.680\(RWSE vs\. none\)\. The frame effect is small because the task is nearly exhausted\. The baseline cell already reaches14\.28914\.289bits,99\.4%99\.4\\%of ceiling, leaving only0\.0900\.090bits available\. The frame captures79%79\\%of that\. Against this, changing only the learning\-rate schedule at a fixed step count moves the baseline cell by\+1\.337\+1\.337bits \(→14\.28912\.952\\\!\\to\\\!14\.289\)\.
##### Result 3: Bits and the reasoning probe are decoupled\.
Across1010converged cells, bits span11\.808−14\.37911\.808\-14\.379and A1 spans0\.711−0\.9850\.711\-0\.985, at Spearmanρ=−0\.24\\rho=\-0\.24\. This is not significant at this sample size\. The cells are also neither independent nor randomly selected, so we claim no trade\-off and no direction\. Table[6](https://arxiv.org/html/2608.20516#S8.T6)does show the practical consequence\. A bits\-driven selection and an A1\-driven selection choose different cells\.raw:1 has the lowest bits of any converged run \(11\.80811\.808\) and the highest A1 \(0\.9850\.985\)\. Dropping the structural cue costs\+0\.680\+0\.680bits while raising A1 to0\.9070\.907\. A within\-configuration version points the same way: from33k to2020k steps, bits move\+0\.226\+0\.226and A1 moves−0\.225\-0\.225\. Those two points differ in budget and in schedule length, so we quote them as a bound on a joint effect only\. A1’s own construction and its floor are audited in App\.[A16](https://arxiv.org/html/2608.20516#A16)\. That audit is a caveat on this result, not a footnote to it\. Optimising either number guarantees nothing about the other\. This is why Sec\.[10](https://arxiv.org/html/2608.20516#S10)recommends watching a second metric rather than predicting its direction\.
##### Result 4: The target is structurally reducible, so a near\-ceiling score is uninformative\.
Prop\.[5](https://arxiv.org/html/2608.20516#Thmproposition5)applies literally to this graph\. The intra\-subgraph relations are added for every subgraph possessing the relevant node types\. Their cardinalities are therefore exactly the node census:produceshas57,90357\{,\}903edges \(11per paper\),groundshas251,938251\{,\}938\(11per claim\), andhas claimhas251,938251\{,\}938\. The edge set is a constant function of the census and carries zero information\. Message passing over it can only mix a subgraph’s own aspect vectors\. That is precisely what the training\-free oracle does by hand\. Hence the oracle reaches13\.86513\.865of14\.37914\.379bits with no training at all:96\.4%96\.4\\%of the ceiling, leaving0\.5140\.514bits of headroom for any model\.citesis the only non\-census\-determined subgraph\-level relation\. It retains11,79111\{,\}791of2,211,1182\{,\}211\{,\}118references after restriction to the corpus \(0\.53%0\.53\\%,0\.200\.20edges per paper\)\. A learned gate between raw features and encoder output settles atσ=0\.3886\\sigma=0\.3886, such as toward the encoder\. The encoder does therefore earn its place\. By Prop\.[5](https://arxiv.org/html/2608.20516#Thmproposition5), however, it earns it as a better metric over census\-determined features, not as a structure learner\.
The reducibility check should be standard\.Compute each relation’s cardinality and compare it to the node census\. If a relation’s count equals the count of one of its endpoint types, that relation is a deterministic function of node existence and contributes nothing\. On our graph,55of99relations fail this check\. Together they constitute the entire intra\-subgraph structure\. No rank statistic, probe or retrieval score detects this, because the failure is in the target, not in the representation\.
##### Result 5: The data\-derived replacement fails a data\-quality gate\.
Three relations are data\-derived:supported by\(251,922251\{,\}922\),challenged by\(121,516121\{,\}516\) andimplies\(251,938251\{,\}938\)\. We made them prediction targets, and gated the attempt on data quality before training\. Both gates fired \(Table[7](https://arxiv.org/html/2608.20516#S8.T7)\)\.25\.96%25\.96\\%of contradicting\-evidence nodes are exact duplicate rows\. The largest single group contains3,0313\{,\}031identical strings\. The duplication is also almost perfectly asymmetric\.31,54931\{,\}549evidence nodes lie in duplicate groups, and31,54531\{,\}545challenge edges point at duplicated rows\. Essentially every boilerplate node is a challenge node, and supporting evidence is nearly free of exact duplication\. A duplicate filter alone would not repair the target: the surviving text is\+0\.2377\+0\.2377more self\-similar than supporting evidence\. This asymmetry is the apparent polarity signal\.cos\(claim,support\)=0\.6607\\cos\(\\text\{claim\},\\text\{support\}\)=0\.6607againstcos\(claim,challenge\)=0\.3061\\cos\(\\text\{claim\},\\text\{challenge\}\)=0\.3061, atd=\+1\.638d=\+1\.638\. A cosine\-only classifier already reaches0\.85390\.8539AUC\. A model trained on this target learns to detect fillers\.
##### Two by\-products of the failed attempt that we would otherwise have got wrong\.
The first is a confound that does not exist\. Polarity is unevenly distributed across papers\.35,65035\{,\}650papers hold at least one challenge edge,61\.6%61\.6\\%of the corpus, against94\.3%94\.3\\%under independent per\-claim assignment\. We therefore predicted that a claim\-grouped split would leak paper identity, and pre\-registered a paper\-grouped one\. We reran the identical probe with only the grouping changed\. The leak is\+0\.0003\+0\.0003AUC: absent\. We report this because the pre\-registration rested on arithmetic we had not verified against the graph, and the experiment refuted it\. The second is a leak we did not anticipate, and it is large\. With the target relations left in the graph, the probe reads1\.00001\.0000AUC\.to\_heterogives each relation its own weights, and evidence nodes are leaves \(373,438=251,922\+121,516373\{,\}438=251\{,\}922\+121\{,\}516exactly\)\. An evidence embedding is therefore one of two linear maps of its own features, and the label is trivially decodable\. Masking every target relation in both directions drops the probe to0\.85350\.8535\. That is indistinguishable from the cosine floor, at−0\.0004\-0\.0004\. Typed relations make edge\-level label leakage a property of the schema, not of the split\.
What this section licenses, and what it does not\.It licenses four claims\. The collapse is repairable, and the loss is the operative variable\. The repaired metric is near\-saturated on a target that is provably reducible\. The optimised metric and a reasoning\-relevant probe move independently across the cells we ran\. And the data\-derived alternative is currently unusable on this corpus\. It does*not*license any claim that the repaired representation is or is not reasoning\-relevant\. The measurement that would settle that does not yet exist on this corpus, which is precisely the finding\. Nor do we present the two cells at exactly14\.37914\.379bits as headline results\. AnMRR\\mathrm\{MRR\}of1\.00001\.0000is the same signature our own harness flags as degenerate, and App\.[A14](https://arxiv.org/html/2608.20516#A14)records the leak check we ran on them\. The required fix is at the extraction layer, not the model: a length\-and\-pattern guard where the evidence string is accepted, then one rebuild, then a re\-measurement of14\.37914\.379,13\.86513\.865and every derived reference point in the same pass\.
## 10Discussion
A representation can pass every standard health check and carry zero usable instance information; the configuration that does so is a*global*optimum of the coupled predictor/EMA\-target objective rather than a failure of its optimisation, and the converged latents allocate their variance to the category rather than to the instance; the failure lives on the query side, where the standard toolkit does not look and provably cannot reach; and once repaired, the metric saturates on a target that carries no structural information\. The reusable methodological result is a set of five instruments: a*matched training\-free oracle*separating task difficulty from pipeline capability, a*guarded positive control*separating harness from model,*query\-bank instrumentation*separating candidate\-side from query\-side pathology, a*reducibility audit*of the evaluation target, and a*data\-quality gate*on any target derived from machine extraction\.
##### The concrete open problems\.
Three, stated rather than claimed\.*\(i\)*The allocation accounts for11\.011\.0of12\.84512\.845lost bits and anisotropy is excluded as the residual by leave\-one\-out; localising the remaining14%14\\%needs a query\-side ablation comparing the full predictor against a structural\-encoding\-only query, and the2×22\\times 2factorial of App\.[A21](https://arxiv.org/html/2608.20516#A21)run to completion\.*\(ii\)*Whether the repaired representation carries reasoning\-relevant structure is*unmeasurable on this corpus today*, because the only data\-derived targets fail Table[7](https://arxiv.org/html/2608.20516#S8.T7)\. That is an extraction problem with a known fix and a known cost—one rebuild plus re\-measurement of every ceiling—and it is the first item of future work rather than a caveat\.*\(iii\)*Every claim here is made on one corpus whose evaluation target we prove to be census\-determined\. The decisive next experiment is therefore a graph whose edge set is*not*a function of the census: if the collapse survives, it is a property of the objective alone; if it does not, Prop\.[5](https://arxiv.org/html/2608.20516#Thmproposition5)becomes a scope condition rather than a caveat\. Either outcome is informative, and the experiment costs about one GPU\-hour on the harness we release\.
##### Recommendations\.
\(1\)Report*bits recovered*with a bootstrap CI and an exact chance baseline, against the*recoverable*ceiling1Nlog2N\!\\frac\{1\}\{N\}\\log\_\{2\}N\!rather than the index entropylog2N\\log\_\{2\}N; the two differ by1\.44271\.4427bits and only the former assigns a uniform ranker exactly zero\. Never report a ratio against chance\.\(2\)Run a*matched training\-free oracle*—if no\-training is close to ceiling, your headroom is the story, not your score\. On our task, the oracle takes96\.4%96\.4\\%and leaves0\.5140\.514bits for every model ever trained on it\.\(3\)*Audit the target for reducibility*: compare every relation’s cardinality to the node census; a relation whose count equals an endpoint type’s count is information\-free\.\(4\)Run a*positive control*sharing the harness but not the model, and verify it is not numerically identical to the system it controls for\.\(5\)*Instrument the query bank*—effective rank and mean pairwise cosine of the model’s*predictions*; a rank at or below the number of latent categories means the experiment is over\.\(6\)Compute the*variance decomposition at every checkpoint*, not only at the ends; one forward pass each, and it is the only way to distinguish a degeneracy the objective*created*from one it merely*failed to remove*\.\(7\)With typed relations,*mask both directions of every target relation*and verify by perturbation that target\-side representations are invariant to context features; a single retained relation moved our probe→1\.00000\.8535\\\!\\to\\\!1\.0000\.\(8\)*Gate machine\-extracted targets on data quality*before training—exact\-duplicate census and a genericness statistic—and treat a fired gate as a result\.\(9\)*Watch a second metric, and report the nuisance factors\.*Ours moved independently of the optimised one, and the largest positive effect in our entire sweep was the learning\-rate schedule at\+1\.337\+1\.337bits, which no convention requires anyone to disclose\.
## 11Limitations and Conclusion
##### Limitations\.
One real corpus, and an evaluation target we prove to be census\-determined, so the scope of the empirical claim is this pipeline on this corpus rather than JEPAs in general; leakage controls bound but do not eliminate surface overlap \(bm25\-no\-ovretains\+14\.250\+14\.250bits\); the probe uses488488coarse labels; the pooled\-rank signature is present at initialisation \(Fig\.[10](https://arxiv.org/html/2608.20516#A32.F10)\(b\)\), so the endpoint comparison of Table[3](https://arxiv.org/html/2608.20516#S7.T3)is a statement about*what*the converged representation encodes and Table[5](https://arxiv.org/html/2608.20516#S7.T5)is the only evidence we offer about*when*; the33k versus2020k comparison differs in both budget and schedule length and therefore bounds a joint effect only; the bits/A1 result isn=10n=10at Spearman−0\.24\-0\.24over non\-independent cells and we draw no directional conclusion from it, and A1’s own floor is audited in App\.[A16](https://arxiv.org/html/2608.20516#A16); two cells reportMRR=1\.0000\\mathrm\{MRR\}=1\.0000, which we treat as requiring the leak check of App\.[A14](https://arxiv.org/html/2608.20516#A14)rather than as a headline; a large permutationpp\-value is a failure to reject rather than positive evidence of equality, so we describe the baseline as indistinguishable from chance at the resolution of4,0004\{,\}000queries rather than as exactly chance; our re\-implementation does not reproduce the original method’s MUTAG number;14%14\\%of the baseline collapse is unexplained by the allocation alone; and our corpus’s contradicting\-evidence field is25\.96%25\.96\\%placeholder text, which we discovered only after building a task on it\.
##### Conclusion\.
We trained a Graph\-JEPA on a reasoning graph over57,90357\{,\}903papers\. Its loss converged, its probe reached0\.8710\.871, its effective rank stayed healthy, and it recovered0\.000\.00of14\.37914\.379recoverable bits \(p=0\.98p=0\.98\)—while a parameter\-free average of the same frozen features recovered\+14\.281\+14\.281, BM25\+14\.335\+14\.335, and a positive control through the identical harness\+14\.220\+14\.220, all within1\.2%1\.2\\%of the ceiling\. A controlled ladder eliminated the encoder, its depth, the pooling operator, the target regularisation and the retrieval geometry, and localised the failure to a query bank of effective rank1\.91\.9—at or below the33\-category bound our theory predicts\. Repairing the objective took it to14\.37714\.377bits, above the13\.86513\.865\-bit oracle, and reverting only the loss returned it to0\.3070\.307—a14\.0514\.05\-bit swing that confirms the mechanism on one variable\. And then the interesting part: the retrieval target is provably reducible, so a near\-ceiling score cannot evidence learned structure; the optimised metric and a reasoning\-relevant probe move independently; the largest positive effect we measured was the learning\-rate schedule; and the only data\-derived alternative is25\.96%25\.96\\%placeholder text\. Class collapse is known in supervised metric learning, mean\-prediction in JEPA, and probe–task mismatch in vision; we show these are one phenomenon, that it can be complete rather than partial, that it is*a global optimum*of the coupled predictor/EMA\-target objective rather than a bug in its optimisation, that it is repairable—and that repairing it exposes a failure the standard toolkit is not even pointed at, because that failure is in the evaluation target rather than the representation\.
#### Reproducibility Statement
Graph construction, all configurations, Protocol R, the three ceilings and their leakage controls, the difficulty ladder, the extraction audit, the phase analysis, the checkpoint allocation measurement of Sec\.[7\.1](https://arxiv.org/html/2608.20516#S7.SS1), the repair sweep of Sec\.[9](https://arxiv.org/html/2608.20516#S9), the reducibility audit, the polarity data gates and the pre\-registered factorial are released with the per\-seed logs underlying every table\. Decision thresholds are written to disk before results \(App\.[A10](https://arxiv.org/html/2608.20516#A10)\)\. The2020\-assertion instrument self\-test \(App\.[A9](https://arxiv.org/html/2608.20516#A9)\) runs standalone in∼20\{\\sim\}20s and must pass before any measurement is taken; it is the reason we caught the\+0\.1465\+0\.1465\-AUC edge\-type leak of Sec\.[9](https://arxiv.org/html/2608.20516#S9)before reporting rather than after\. The bits accounting of App\.[A8](https://arxiv.org/html/2608.20516#A8)is independently reproducible from the released module\. The diagnosis costs0\.770\.77GPU\-hours on one L40S \(App\.[A33](https://arxiv.org/html/2608.20516#A33)\)\. App\.[A1](https://arxiv.org/html/2608.20516#A1)maps every quantitative claim in the abstract and introduction to the table that establishes it, and App\.[A5](https://arxiv.org/html/2608.20516#A5)records which run, corpus subset and sentence encoder produced each number\.
#### Use of Large Language Models
An LLM was used for two purposes\. The first was language editing of author\-written text and LaTeX formatting\. The second was as an aid in reviewing experimental design\. In that second role, LLM critique led to the positive control, the paired bootstrap, the pre\-registration, the correction to the bits ceiling documented in App\.[A8](https://arxiv.org/html/2608.20516#A8), and the reducibility and data\-quality gates of Sec\.[9](https://arxiv.org/html/2608.20516#S9)\. The same critique also generated four claims that this paper reports as refuted or corrected\.*\(i\)*A predicted paper\-identity leak, which we measure at\+0\.0003\+0\.0003\.*\(ii\)*A predicted identity between the reasoning probe and a paper\-level label that the graph does not support\.*\(iii\)*A target\-frame main effect quoted as\+1\.199\+1\.199bits, which is\+0\.071\+0\.071at matched budget\.*\(iv\)*A conditional\-mean account of the collapse whose precondition our own training\-free oracle refutes, corrected in Sec\.[4](https://arxiv.org/html/2608.20516#S4)\. We record all four for one reason\. A confident wrong number, or a confident wrong theorem, that survives into a draft is the characteristic failure mode of this workflow\. Item*\(iii\)*was caught only by reading a log stage we had previously skipped\. All experiments were designed, implemented and executed by the authors\.
#### Ethics Statement
The corpus consists of metadata and machine\-extracted content from publicly available scholarly records; no human subjects, personal data, or sensitive attributes are involved\. We report a negative result concerning an existing published method; our intent is diagnostic rather than dismissive; we disclose in App\.[A28](https://arxiv.org/html/2608.20516#A28)that our re\-implementation does not reproduce the original method’s benchmark numbers, and we release the harness so the diagnosis can be contested\. We additionally disclose that our own corpus contains25\.96%25\.96\\%placeholder text in one extracted field, since publishing a scholarly\-reasoning resource without that disclosure would propagate the artifact\.
## References
- Assran et al\. \(2023\)Mahmoud Assran, Quentin Duval, Ishan Misra, Piotr Bojanowski, Pascal Vincent, Michael G\. Rabbat, Yann LeCun, and Nicolas Ballas\.Self\-supervised learning from images with a joint\-embedding predictive architecture\.In*IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17\-24, 2023*, pp\. 15619–15629\. IEEE, 2023\.doi:10\.1109/CVPR52729\.2023\.01499\.URL[https://doi\.org/10\.1109/CVPR52729\.2023\.01499](https://doi.org/10.1109/CVPR52729.2023.01499)\.
- Auer et al\. \(2020\)Sören Auer, Allard Oelen, Muhammad Haris, Markus Stocker, Jennifer D’Souza, Kheir Eddine Farfar, Lars Vogt, Manuel Prinz, Vitalis Wiens, and Mohamad Yaser Jaradeh\.Improving access to scientific literature with knowledge graphs\.*Bibliothek Forschung und Praxis*, 44\(3\):516–529, 2020\.
- Bardes et al\. \(2022\)Adrien Bardes, Jean Ponce, and Yann LeCun\.Vicreg: Variance\-invariance\-covariance regularization for self\-supervised learning\.In*The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25\-29, 2022*\. OpenReview\.net, 2022\.URL[https://openreview\.net/forum?id=xm6YD62D1Ub](https://openreview.net/forum?id=xm6YD62D1Ub)\.
- Chen et al\. \(2022\)Mayee F\. Chen, Daniel Y\. Fu, Avanika Narayan, Michael Zhang, Zhao Song, Kayvon Fatahalian, and Christopher Ré\.Perfectly balanced: Improving transfer and robustness of supervised contrastive learning\.In Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csaba Szepesvári, Gang Niu, and Sivan Sabato \(eds\.\),*International Conference on Machine Learning, ICML 2022, 17\-23 July 2022, Baltimore, Maryland, USA*, volume 162 of*Proceedings of Machine Learning Research*, pp\. 3090–3122\. PMLR, 2022\.URL[https://proceedings\.mlr\.press/v162/chen22d\.html](https://proceedings.mlr.press/v162/chen22d.html)\.
- Dawid & LeCun \(2023\)Anna Dawid and Yann LeCun\.Introduction to latent variable energy\-based models: A path towards autonomous machine intelligence\.*CoRR*, abs/2306\.02572, 2023\.doi:10\.48550/ARXIV\.2306\.02572\.URL[https://doi\.org/10\.48550/arXiv\.2306\.02572](https://doi.org/10.48550/arXiv.2306.02572)\.
- Dwivedi et al\. \(2022\)Vijay Prakash Dwivedi, Anh Tuan Luu, Thomas Laurent, Yoshua Bengio, and Xavier Bresson\.Graph neural networks with learnable structural and positional representations\.In*The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25\-29, 2022*\. OpenReview\.net, 2022\.URL[https://openreview\.net/forum?id=wTTjnvGphYj](https://openreview.net/forum?id=wTTjnvGphYj)\.
- Ethayarajh \(2019\)Kawin Ethayarajh\.How contextual are contextualized word representations? comparing the geometry of bert, elmo, and GPT\-2 embeddings\.In Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan \(eds\.\),*Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP\-IJCNLP 2019, Hong Kong, China, November 3\-7, 2019*, pp\. 55–65\. Association for Computational Linguistics, 2019\.doi:10\.18653/V1/D19\-1006\.URL[https://doi\.org/10\.18653/v1/D19\-1006](https://doi.org/10.18653/v1/D19-1006)\.
- Gao et al\. \(2019\)Jun Gao, Di He, Xu Tan, Tao Qin, Liwei Wang, and Tie\-Yan Liu\.Representation degeneration problem in training natural language generation models\.In*7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6\-9, 2019*\. OpenReview\.net, 2019\.URL[https://openreview\.net/forum?id=SkEYojRqtm](https://openreview.net/forum?id=SkEYojRqtm)\.
- Gao et al\. \(2021\)Tianyu Gao, Xingcheng Yao, and Danqi Chen\.Simcse: Simple contrastive learning of sentence embeddings\.In Marie\-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen\-tau Yih \(eds\.\),*Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7\-11 November, 2021*, pp\. 6894–6910\. Association for Computational Linguistics, 2021\.doi:10\.18653/V1/2021\.EMNLP\-MAIN\.552\.URL[https://doi\.org/10\.18653/v1/2021\.emnlp\-main\.552](https://doi.org/10.18653/v1/2021.emnlp-main.552)\.
- Garrido et al\. \(2023\)Quentin Garrido, Randall Balestriero, Laurent Najman, and Yann LeCun\.Rankme: Assessing the downstream performance of pretrained self\-supervised representations by their rank\.In Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett \(eds\.\),*International Conference on Machine Learning, ICML 2023, 23\-29 July 2023, Honolulu, Hawaii, USA*, volume 202 of*Proceedings of Machine Learning Research*, pp\. 10929–10974\. PMLR, 2023\.URL[https://proceedings\.mlr\.press/v202/garrido23a\.html](https://proceedings.mlr.press/v202/garrido23a.html)\.
- Graf et al\. \(2021\)Florian Graf, Christoph D\. Hofer, Marc Niethammer, and Roland Kwitt\.Dissecting supervised constrastive learning\.In Marina Meila and Tong Zhang \(eds\.\),*Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18\-24 July 2021, Virtual Event*, volume 139 of*Proceedings of Machine Learning Research*, pp\. 3821–3830\. PMLR, 2021\.URL[http://proceedings\.mlr\.press/v139/graf21a\.html](http://proceedings.mlr.press/v139/graf21a.html)\.
- Grill et al\. \(2020\)Jean\-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre H\. Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Ávila Pires, Zhaohan Guo, Mohammad Gheshlaghi Azar, Bilal Piot, Koray Kavukcuoglu, Rémi Munos, and Michal Valko\.Bootstrap your own latent \- A new approach to self\-supervised learning\.In Hugo Larochelle, Marc’Aurelio Ranzato, Raia Hadsell, Maria\-Florina Balcan, and Hsuan\-Tien Lin \(eds\.\),*Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6\-12, 2020, virtual*, 2020\.URL[https://proceedings\.neurips\.cc/paper/2020/hash/f3ada80d5c4ee70142b17b8192b2958e\-Abstract\.html](https://proceedings.neurips.cc/paper/2020/hash/f3ada80d5c4ee70142b17b8192b2958e-Abstract.html)\.
- Hou et al\. \(2023\)Zhenyu Hou, Yufei He, Yukuo Cen, Xiao Liu, Yuxiao Dong, Evgeny Kharlamov, and Jie Tang\.Graphmae2: A decoding\-enhanced masked self\-supervised graph learner\.In Ying Ding, Jie Tang, Juan F\. Sequeda, Lora Aroyo, Carlos Castillo, and Geert\-Jan Houben \(eds\.\),*Proceedings of the ACM Web Conference 2023, WWW 2023, Austin, TX, USA, 30 April 2023 \- 4 May 2023*, pp\. 737–746\. ACM, 2023\.doi:10\.1145/3543507\.3583379\.URL[https://doi\.org/10\.1145/3543507\.3583379](https://doi.org/10.1145/3543507.3583379)\.
- Hu et al\. \(2020\)Ziniu Hu, Yuxiao Dong, Kuansan Wang, and Yizhou Sun\.Heterogeneous graph transformer\.In Yennun Huang, Irwin King, Tie\-Yan Liu, and Maarten van Steen \(eds\.\),*WWW ’20: The Web Conference 2020, Taipei, Taiwan, April 20\-24, 2020*, pp\. 2704–2710\. ACM / IW3C2, 2020\.doi:10\.1145/3366423\.3380027\.URL[https://doi\.org/10\.1145/3366423\.3380027](https://doi.org/10.1145/3366423.3380027)\.
- Jing et al\. \(2022\)Li Jing, Pascal Vincent, Yann LeCun, and Yuandong Tian\.Understanding dimensional collapse in contrastive self\-supervised learning\.In*The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25\-29, 2022*\. OpenReview\.net, 2022\.URL[https://openreview\.net/forum?id=YevsQ05DEN7](https://openreview.net/forum?id=YevsQ05DEN7)\.
- Papyan et al\. \(2020\)Vardan Papyan, X\. Y\. Han, and David L\. Donoho\.Prevalence of neural collapse during the terminal phase of deep learning training\.*CoRR*, abs/2008\.08186, 2020\.URL[https://arxiv\.org/abs/2008\.08186](https://arxiv.org/abs/2008.08186)\.
- Robertson & Zaragoza \(2009\)Stephen E\. Robertson and Hugo Zaragoza\.The probabilistic relevance framework: BM25 and beyond\.*Found\. Trends Inf\. Retr\.*, 3\(4\):333–389, 2009\.doi:10\.1561/1500000019\.URL[https://doi\.org/10\.1561/1500000019](https://doi.org/10.1561/1500000019)\.
- Roy & Vetterli \(2007\)Olivier Roy and Martin Vetterli\.The effective rank: A measure of effective dimensionality\.In*15th European Signal Processing Conference, EUSIPCO 2007, Poznan, Poland, September 3\-7, 2007*, pp\. 606–610\. IEEE, 2007\.URL[https://ieeexplore\.ieee\.org/document/7098875/](https://ieeexplore.ieee.org/document/7098875/)\.
- Schuhmann et al\. \(2025\)Christoph Schuhmann, Gollam Rabby, Ameya Prabhu, Tawsif Ahmed, Andreas Hochlehnert, Huu Nguyen, Nick Akinci Heidrich, Ludwig Schmidt, Robert Kaczmarczyk, Sören Auer, Jenia Jitsev, and Matthias Bethge\.Project alexandria: Towards freeing scientific knowledge from copyright burdens via llms\.*CoRR*, abs/2502\.19413, 2025\.doi:10\.48550/ARXIV\.2502\.19413\.URL[https://doi\.org/10\.48550/arXiv\.2502\.19413](https://doi.org/10.48550/arXiv.2502.19413)\.
- Skenderi et al\. \(2025\)Geri Skenderi, Hang Li, Jiliang Tang, and Marco Cristani\.Graph\-level representation learning with joint\-embedding predictive architectures\.*Trans\. Mach\. Learn\. Res\.*, 2025, 2025\.URL[https://openreview\.net/forum?id=v47f4DwYZb](https://openreview.net/forum?id=v47f4DwYZb)\.
- Su et al\. \(2021\)Jianlin Su, Jiarun Cao, Weijie Liu, and Yangyiwen Ou\.Whitening sentence representations for better semantics and faster retrieval\.*CoRR*, abs/2103\.15316, 2021\.URL[https://arxiv\.org/abs/2103\.15316](https://arxiv.org/abs/2103.15316)\.
- Wang & Zhang \(2022\)Xiyuan Wang and Muhan Zhang\.How powerful are spectral graph neural networks\.In Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csaba Szepesvári, Gang Niu, and Sivan Sabato \(eds\.\),*International Conference on Machine Learning, ICML 2022, 17\-23 July 2022, Baltimore, Maryland, USA*, volume 162 of*Proceedings of Machine Learning Research*, pp\. 23341–23362\. PMLR, 2022\.URL[https://proceedings\.mlr\.press/v162/wang22am\.html](https://proceedings.mlr.press/v162/wang22am.html)\.
- Yao et al\. \(2024\)Tianjun Yao, Yongqiang Chen, Zhenhao Chen, Kai Hu, Zhiqiang Shen, and Kun Zhang\.Empowering graph invariance learning with deep spurious infomax\.In Ruslan Salakhutdinov, Zico Kolter, Katherine A\. Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, and Felix Berkenkamp \(eds\.\),*Forty\-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21\-27, 2024*, volume 235 of*Proceedings of Machine Learning Research*, pp\. 56860–56884\. PMLR / OpenReview\.net, 2024\.URL[https://proceedings\.mlr\.press/v235/yao24a\.html](https://proceedings.mlr.press/v235/yao24a.html)\.
- Zbontar et al\. \(2021\)Jure Zbontar, Li Jing, Ishan Misra, Yann LeCun, and Stéphane Deny\.Barlow twins: Self\-supervised learning via redundancy reduction\.In Marina Meila and Tong Zhang \(eds\.\),*Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18\-24 July 2021, Virtual Event*, volume 139 of*Proceedings of Machine Learning Research*, pp\. 12310–12320\. PMLR, 2021\.URL[http://proceedings\.mlr\.press/v139/zbontar21a\.html](http://proceedings.mlr.press/v139/zbontar21a.html)\.
## Appendix
## Appendix A1Claim–evidence map
Table[8](https://arxiv.org/html/2608.20516#A1.T8)lists every quantitative claim made in the abstract and introduction together with the table that establishes it\. Each value is a single macro in the source, so prose and tables cannot diverge\.
Table 8:Each quantitative claim in the abstract and introduction, its value, and the table or figure that establishes it\. Every value appears exactly once as a macro in the source, so prose and tables cannot diverge\.ClaimValueSourceTrained retrieval is indistinguishable from chance0\.000\.00bits,p=0\.98p=0\.98Table[1](https://arxiv.org/html/2608.20516#S4.T1), rung 0Chance level1\.99×10−41\.99\\times 10^\{\-4\}\(MRR\),00bitsSec\.[3](https://arxiv.org/html/2608.20516#S3), Eq\.[3](https://arxiv.org/html/2608.20516#A8.E3)Recoverable ceiling, block A / B14\.37914\.379/14\.38514\.385bitsEq\.[4](https://arxiv.org/html/2608.20516#A8.E4)Linear probe succeeds0\.8710\.871Table[28](https://arxiv.org/html/2608.20516#A23.T28)Candidate bank is healthy18\.118\.1, DC37\.637\.6Table[1](https://arxiv.org/html/2608.20516#S4.T1)Query bank is degenerate1\.91\.9, DC162\.2162\.2Table[1](https://arxiv.org/html/2608.20516#S4.T1)Predicted query\-bank bounderank≤3\\operatorname\{erank\}\\leq 3Prop\.[1](https://arxiv.org/html/2608.20516#Thmproposition1)Pooling is the identity \(m=1m\{=\}1\)→47\.3147\.42\\\!\\to\\\!47\.31Table[23](https://arxiv.org/html/2608.20516#A18.T23)Oracle ceiling \(block B\)\+14\.281\+14\.281bits \(99\.3%99\.3\\%\)Table[2](https://arxiv.org/html/2608.20516#S6.T2)Lexical ceiling\+14\.335\+14\.335bits \(99\.7%99\.7\\%\)Table[2](https://arxiv.org/html/2608.20516#S6.T2)Positive control\+14\.220\+14\.220bits \(98\.9%98\.9\\%\)Table[2](https://arxiv.org/html/2608.20516#S6.T2)Non\-lexical signal survives\+14\.250\+14\.250bits \(99\.1%99\.1\\%\)Table[2](https://arxiv.org/html/2608.20516#S6.T2)Allocation, inputs \(block B\)0\.40%0\.40\\%/86\.05%86\.05\\%Table[3](https://arxiv.org/html/2608.20516#S7.T3)Allocation, inputs \(block A, within\-encoder\)see col\. 2Table[3](https://arxiv.org/html/2608.20516#S7.T3)Allocation, trained latents99\.61%99\.61\\%/0\.39%0\.39\\%Table[3](https://arxiv.org/html/2608.20516#S7.T3)Allocation over trainingper checkpointTable[5](https://arxiv.org/html/2608.20516#S7.T5)Allocation explains86%86\\%of loss11\.04511\.045of12\.84512\.845bitsTable[4](https://arxiv.org/html/2608.20516#S7.T4)Critical shareρ⋆=0\.99999\\rho^\{\\star\}=0\.99999Table[4](https://arxiv.org/html/2608.20516#S7.T4)Anisotropy not the residual factor<0\.004<0\.004MRRTable[27](https://arxiv.org/html/2608.20516#A22.T27)Post\-hoc frame sweep gain0\.4bits0\.4\\text\{ bits\}Table[19](https://arxiv.org/html/2608.20516#A13.T19)Depth is not binding→18\.054\.9\\\!\\to\\\!18\.0, bits flatTable[1](https://arxiv.org/html/2608.20516#S4.T1), rung 6Templating index0\.9370\.937vs0\.3330\.333Table[31](https://arxiv.org/html/2608.20516#A27.T31)De\-templated oracle\+14\.28\+14\.28bitsTable[31](https://arxiv.org/html/2608.20516#A27.T31)No post\-hoc intervention clears the null0\.000\.00bits, 7 cellsTable[28](https://arxiv.org/html/2608.20516#A23.T28)Dissociation across 20 cellsprobe0\.7550\.755–0\.9600\.960,rpoolr\_\{\\mathrm\{pool\}\}1\.81\.8–13\.913\.9Table[29](https://arxiv.org/html/2608.20516#A24.T29)Anchor dim\.→1282\\\!\\to\\\!128,rtgtr\_\{\\mathrm\{tgt\}\}flat1\.61\.6–2\.02\.0Table[30](https://arxiv.org/html/2608.20516#A25.T30)Repair reaches near\-ceiling14\.37714\.377bits \(100\.0%100\.0\\%\)Table[6](https://arxiv.org/html/2608.20516#S8.T6)Loss is the lever \(one variable\)→0\.30714\.359\\\!\\to\\\!0\.307\(14\.0514\.05\)Table[6](https://arxiv.org/html/2608.20516#S8.T6)Target\-frame effect, matched budget\+0\.071\+0\.071bitsTable[20](https://arxiv.org/html/2608.20516#A14.T20)Schedule effect, matched steps\+1\.337\+1\.337bitsTable[20](https://arxiv.org/html/2608.20516#A14.T20)Bits and A1 move independentlyρ=−0\.24\\rho=\-0\.24,n=10n=10Table[6](https://arxiv.org/html/2608.20516#S8.T6)Cue ablation14\.37614\.376/14\.20914\.209/13\.52913\.529Table[6](https://arxiv.org/html/2608.20516#S8.T6)Transduction gap\+0\.001\+0\.001bitsTable[6](https://arxiv.org/html/2608.20516#S8.T6)Raw\-skip gateσ=0\.3886\\sigma=0\.3886Table[6](https://arxiv.org/html/2608.20516#S8.T6)Training\-free oracle, block A13\.86513\.865bits \(96\.4%96\.4\\%\)Table[6](https://arxiv.org/html/2608.20516#S8.T6)Headroom for any model0\.5140\.514bitsTable[6](https://arxiv.org/html/2608.20516#S8.T6)Census\-determined relations57,90357\{,\}903,251,938251\{,\}938Table[15](https://arxiv.org/html/2608.20516#A7.T15)citesis near\-empty11,79111\{,\}791of2,211,1182\{,\}211\{,\}118\(0\.53%0\.53\\%\)Table[15](https://arxiv.org/html/2608.20516#A7.T15)Duplicate placeholder share25\.96%25\.96\\%Table[7](https://arxiv.org/html/2608.20516#S8.T7)Genericness gap\+0\.2377\+0\.2377Table[7](https://arxiv.org/html/2608.20516#S8.T7)Edge\-type label leak→1\.00000\.8535\\\!\\to\\\!1\.0000Table[7](https://arxiv.org/html/2608.20516#S8.T7)Predicted leak is absent\+0\.0003\+0\.0003Table[7](https://arxiv.org/html/2608.20516#S8.T7)Faithfulness gap disclosed0\.6860\.686vs0\.8740\.874Table[32](https://arxiv.org/html/2608.20516#A28.T32)
## Appendix A2Number provenance
Table[13](https://arxiv.org/html/2608.20516#A5.T13)records which run, corpus subset and sentence encoder produced each block; the two caveats we do not smooth over follow it\.
Table 9:Which run, corpus subset and sentence encoder produced each block\. We keep these separate rather than pooling them, because the blocks measure different objects and the distinction carries the central results of Secs\.[7](https://arxiv.org/html/2608.20516#S7)and[9](https://arxiv.org/html/2608.20516#S9)\.BlockPapersEncoderUsed forA \(trained pipeline\)57,90357\{,\}903MiniLM\-L6,384384\-dTables[1](https://arxiv.org/html/2608.20516#S4.T1),[3](https://arxiv.org/html/2608.20516#S7.T3)\(cols\. 2–3\),[5](https://arxiv.org/html/2608.20516#S7.T5),[23](https://arxiv.org/html/2608.20516#A18.T23),[19](https://arxiv.org/html/2608.20516#A13.T19),[25](https://arxiv.org/html/2608.20516#A19.T25),[28](https://arxiv.org/html/2608.20516#A23.T28),[29](https://arxiv.org/html/2608.20516#A24.T29),[30](https://arxiv.org/html/2608.20516#A25.T30)B \(diagnostic harness\)58,14558\{,\}145MPNet\-base,768768\-dTables[2](https://arxiv.org/html/2608.20516#S6.T2),[3](https://arxiv.org/html/2608.20516#S7.T3)\(cols\. 1, 4\),[4](https://arxiv.org/html/2608.20516#S7.T4),[27](https://arxiv.org/html/2608.20516#A22.T27),[31](https://arxiv.org/html/2608.20516#A27.T31)C \(repair\+\+target audit\)57,90357\{,\}903MiniLM\-L6,384384\-dTables[6](https://arxiv.org/html/2608.20516#S8.T6),[20](https://arxiv.org/html/2608.20516#A14.T20),[7](https://arxiv.org/html/2608.20516#S8.T7),[15](https://arxiv.org/html/2608.20516#A7.T15),[22](https://arxiv.org/html/2608.20516#A17.T22)Blocks A and C share the*same graph cache byte\-for\-byte*: no rebuild occurred between them, which is what makes13\.86513\.865,14\.28914\.289,14\.37714\.377and14\.37914\.379directly comparable and is why the extraction fix of Sec\.[9](https://arxiv.org/html/2608.20516#S9)must be accompanied by re\-measuring all four\. All blocks use the same raw records, the same aspect fields \(claims,methodological\_details,key\_results\), the same masking convention and the same bits measure; chance levels agree to two significant figures and recoverable ceilings to0\.0060\.006bits\.
##### Two provenance caveats we do not smooth over\.
*\(i\)*Table[3](https://arxiv.org/html/2608.20516#S7.T3)now reports the input\-side decomposition in*both*feature spaces: column 2 is measured on the block\-A encoder that produced the trained latents in column 3, so the central comparison is within one encoder, and column 1 is retained only because the three ceilings of Sec\.[6](https://arxiv.org/html/2608.20516#S6)were computed on block\-B features\. The block\-A pair is a one\-forward\-pass measurement on the cached graph and required no retraining\. Where the two input columns disagree in magnitude, the text quotes the within\-encoder pair\.*\(ii\)*Our structural\-encoding schema dump prints17,63117\{,\}631citesedges where the stored graph object reports11,79111\{,\}791\. The difference is not a factor of two and is therefore not explained by reverse\-edge addition; we quote11,79111\{,\}791throughout, flag the discrepancy rather than choosing silently, and note that it does not affect Prop\.[5](https://arxiv.org/html/2608.20516#Thmproposition5), under whichcitesis negligible at either count \(0\.53%0\.53\\%or0\.80%0\.80\\%of2,211,1182\{,\}211\{,\}118\)\.
## Appendix A3Extended related work
##### Non\-contrastive collapse theory\.
Collapse\-avoidance analyses of BYOL\-style methods\([Grill et al\. 2020](https://arxiv.org/html/2608.20516#bib.bib12)\), VICReg\([Bardes et al\. 2022](https://arxiv.org/html/2608.20516#bib.bib3)\)and Barlow Twins\([Zbontar et al\. 2021](https://arxiv.org/html/2608.20516#bib.bib24)\)characterise conditions under which the encoder does not become constant on its input domain, and dimensional\-collapse analyses\([Jing et al\. 2022](https://arxiv.org/html/2608.20516#bib.bib15)\)characterise when the representation occupies a low\-dimensional subspace\. Both are*global*statements\. The failure we report is invisible to both: the representation occupies many dimensions and varies substantially over inputs, but within each latent category it is constant, and it is the within\-category variation that the downstream task requires\. Prop\.[1](https://arxiv.org/html/2608.20516#Thmproposition1)adds the reason no loss\-based criterion can flag it—the risk is identical on branches that differ by the entire recoverable budget\.
##### Class collapse in the supervised setting\.
[Graf et al\. 2021](https://arxiv.org/html/2608.20516#bib.bib11)characterise the minimiser of the supervised contrastive loss as a class\-collapsed simplex configuration;[Papyan et al\. 2020](https://arxiv.org/html/2608.20516#bib.bib16)report the same terminal\-phase geometry under cross\-entropy;[Chen et al\. 2022](https://arxiv.org/html/2608.20516#bib.bib4)separate class collapse from feature suppression and show it degrades transfer\. The partition that collapses in all three is the*label*set, and the remedy available there, modifying the supervision, is unavailable to a self\-supervised objective whose categories are latent\. Our setting differs in exactly that respect, and the repair of Sec\.[9](https://arxiv.org/html/2608.20516#S9)is the self\-supervised analogue: remove the category answer from the target rather than from the labels\.
##### Probe\-based evaluation\.
A probe measures whatever partition its labels induce, and if that partition coincides with the shortcut the objective took, the probe reports the shortcut as success\. In our case, the probe decodes field labels, which correlate with aspect\-level structure, and it reads0\.600\.60–0\.960\.96across configurations whose retrieval performance is uniformly zero\. Sec\.[9](https://arxiv.org/html/2608.20516#S9)adds the converse hazard: a*second*probe chosen to be reasoning\-relevant moves independently of the optimised metric across ten converged cells, so reporting one number from either family is insufficient in both directions\.
##### Retrieval evaluation practice and target reducibility\.
In\-batch scoring remains common\. Our difficulty ladder \(Table[2](https://arxiv.org/html/2608.20516#S6.T2)\) shows why this matters quantitatively: bits recovered rise from\+0\.860\+0\.860atK=2K\{=\}2to\+14\.335\+14\.335at the full pool for the same system, while MRR is nearly flat, so a pool\-size\-dependent metric reported without the pool size is close to uninterpretable\. Separately, we are not aware of prior work in graph SSL that audits whether the*evaluation target*’s structure is a deterministic function of the node census\. Prop\.[5](https://arxiv.org/html/2608.20516#Thmproposition5)is elementary, but on our graph it retires the structural interpretation of an otherwise near\-ceiling result, and the check costs one pass over the edge\-type cardinalities\.
## Appendix A4Setup, schema and corpus statistics
Table 10:All effective ranks use Eq\.[5](https://arxiv.org/html/2608.20516#A8.E5)on mean\-centred matrices\. The distinction betweenrcandr\_\{\\mathrm\{cand\}\}andrqueryr\_\{\\mathrm\{query\}\}is central: the former is what standard practice measures, the latter is where the failure lives\.SymbolCode nameMeaningrnoder\_\{\\mathrm\{node\}\}\(/128/128\)node\_rkeffective rank of node latentsrpoolr\_\{\\mathrm\{pool\}\}\(/128/128\)pool\_rkeffective rank of pooled patch vectorsrcandr\_\{\\mathrm\{cand\}\}\(/128/128\)cand\_erankeffective rank of the*candidate bank*rqueryr\_\{\\mathrm\{query\}\}\(/128/128\)query\_eff\_rankeffective rank of the*model’s predictions*DCratio\\mathrm\{DC\}\_\{\\mathrm\{ratio\}\}dc\_ratio‖μ‖/𝔼‖x−μ‖\\\|\\mu\\\|/\\mathbb\{E\}\\\|x\-\\mu\\\|DCenergy\\mathrm\{DC\}\_\{\\mathrm\{energy\}\}dc\_energy‖μ‖2/\(‖μ‖2\+𝔼‖x−μ‖2\)\\\|\\mu\\\|^\{2\}/\(\\\|\\mu\\\|^\{2\}\+\\mathbb\{E\}\\\|x\-\\mu\\\|^\{2\}\)ρasp,ρpap\\rho\_\{\\mathrm\{asp\}\},\\rho\_\{\\mathrm\{pap\}\}rho\_between\_\*variance share between aspects / papersκ^\\hat\{\\kappa\}target\_kappapower\-law decay of the candidate spectrumε^\\hat\{\\varepsilon\}from loss floorpredictor error relative to the identity signal\|𝒜\|\|\\mathcal\{A\}\|—number of latent categories;33hereA1a1\_accheld\-out reasoning probe \(App\.[A16](https://arxiv.org/html/2608.20516#A16)\)ℬ\\mathcal\{B\}bits\_recovered1Nlog2N\!−𝔼\[log2r\]\\frac\{1\}\{N\}\\log\_\{2\}N\!\-\\mathbb\{E\}\[\\log\_\{2\}r\];00is chanceℬmax\\mathcal\{B\}^\{\\max\}bits\_max1Nlog2N\!\\frac\{1\}\{N\}\\log\_\{2\}N\!; the achievable ceilingTable 11:Heterogeneous graph statistics \(57,90357\{,\}903maskable papers\)\. The right\-hand column is the reducibility audit of Sec\.[9](https://arxiv.org/html/2608.20516#S9): a relation whose cardinality equals an endpoint type’s node count is a deterministic function of the census and carries zero information\.Node typecountEdge typecountcensus\-determined?paper57,903\(paper, has\_claim, claim\)251,938yes \(=\|=\|claim\|\|\)claim251,938\(paper, has\_method, method\)57,903yes \(=\|=\|paper\|\|\)method57,903\(paper, has\_result, result\)57,903yes \(=\|=\|paper\|\|\)result57,903\(method, produces, result\)57,903yes \(=\|=\|paper\|\|\)evidence373,438\(result, grounds, claim\)251,938yes \(=\|=\|claim\|\|\)implication251,938\(claim, supported\_by, evidence\)251,922no \(data\-derived\)field488\(claim, challenged\_by, evidence\)121,516no \(data\-derived\)\(claim, implies, implication\)251,938yes \(=\|=\|claim\|\|\)\(paper, cites, paper\)11,791no;0\.53%0\.53\\%of2,211,1182\{,\}211\{,\}118Table 12:Per\-aspect membership and untrained in\-degree\. The singleton rows are what make the pooling control decisive\.Dissociation across 20 cellsprobe0\.7550\.755–0\.9600\.960,rpoolr\_\{\\mathrm\{pool\}\}1\.81\.8–13\.913\.9Table[29](https://arxiv.org/html/2608.20516#A24.T29)Anchor dim\.→1282\\\!\\to\\\!128,rtgtr\_\{\\mathrm\{tgt\}\}flat1\.61\.6–2\.02\.0Table[30](https://arxiv.org/html/2608.20516#A25.T30)Repair reaches near\-ceiling14\.37714\.377bits \(100\.0%100\.0\\%\)Table[6](https://arxiv.org/html/2608.20516#S8.T6)Loss is the lever \(one variable\)→0\.30714\.359\\\!\\to\\\!0\.307\(14\.0514\.05\)Table[6](https://arxiv.org/html/2608.20516#S8.T6)Target\-frame effect, matched budget\+0\.071\+0\.071bitsTable[20](https://arxiv.org/html/2608.20516#A14.T20)Schedule effect, matched steps\+1\.337\+1\.337bitsTable[20](https://arxiv.org/html/2608.20516#A14.T20)Bits and A1 show no reliable relationρ=−0\.24\\rho=\-0\.24,n=10n=10Table[6](https://arxiv.org/html/2608.20516#S8.T6)Cue ablation14\.37614\.376/14\.20914\.209/13\.52913\.529Table[6](https://arxiv.org/html/2608.20516#S8.T6)Transduction gap\+0\.001\+0\.001bitsTable[6](https://arxiv.org/html/2608.20516#S8.T6)Raw\-skip gateσ=0\.3886\\sigma=0\.3886Table[6](https://arxiv.org/html/2608.20516#S8.T6)Training\-free oracle, block A13\.86513\.865bits \(96\.4%96\.4\\%\)Table[6](https://arxiv.org/html/2608.20516#S8.T6)Headroom for any model0\.5140\.514bitsTable[6](https://arxiv.org/html/2608.20516#S8.T6)Census\-determined relations57,90357\{,\}903,251,938251\{,\}938Table[15](https://arxiv.org/html/2608.20516#A7.T15)citesis near\-empty11,79111\{,\}791of2,211,1182\{,\}211\{,\}118\(0\.53%0\.53\\%\)Table[15](https://arxiv.org/html/2608.20516#A7.T15)Duplicate placeholder share25\.96%25\.96\\%Table[7](https://arxiv.org/html/2608.20516#S8.T7)Genericness gap\+0\.2377\+0\.2377Table[7](https://arxiv.org/html/2608.20516#S8.T7)Edge\-type label leak→1\.00000\.8535\\\!\\to\\\!1\.0000Table[7](https://arxiv.org/html/2608.20516#S8.T7)Predicted leak is absent\+0\.0003\+0\.0003Table[7](https://arxiv.org/html/2608.20516#S8.T7)Faithfulness gap disclosed0\.6860\.686vs0\.8740\.874Table[32](https://arxiv.org/html/2608.20516#A28.T32)
## Appendix A5Number provenance
Table[13](https://arxiv.org/html/2608.20516#A5.T13)records which run, corpus subset and sentence encoder produced each block; the two caveats we do not smooth over follow it\.
Table 13:Which run, corpus subset and sentence encoder produced each block\. We keep these separate rather than pooling them, because the blocks measure different objects and the distinction carries the central results of Secs\.[7](https://arxiv.org/html/2608.20516#S7)and[9](https://arxiv.org/html/2608.20516#S9)\.BlockPapersEncoderUsed forA \(trained pipeline\)57,90357\{,\}903MiniLM\-L6,384384\-dTables[1](https://arxiv.org/html/2608.20516#S4.T1),[5](https://arxiv.org/html/2608.20516#S7.T5),[23](https://arxiv.org/html/2608.20516#A18.T23),[19](https://arxiv.org/html/2608.20516#A13.T19),[25](https://arxiv.org/html/2608.20516#A19.T25),[28](https://arxiv.org/html/2608.20516#A23.T28),[29](https://arxiv.org/html/2608.20516#A24.T29),[30](https://arxiv.org/html/2608.20516#A25.T30); cols\. 2–3 of Table[3](https://arxiv.org/html/2608.20516#S7.T3)B \(diagnostic harness\)58,14558\{,\}145MPNet\-base,768768\-dTables[2](https://arxiv.org/html/2608.20516#S6.T2),[4](https://arxiv.org/html/2608.20516#S7.T4),[27](https://arxiv.org/html/2608.20516#A22.T27),[31](https://arxiv.org/html/2608.20516#A27.T31); cols\. 1 and 4 of Table[3](https://arxiv.org/html/2608.20516#S7.T3)C \(repair\+\+target audit\)57,90357\{,\}903MiniLM\-L6,384384\-dTables[6](https://arxiv.org/html/2608.20516#S8.T6),[20](https://arxiv.org/html/2608.20516#A14.T20),[7](https://arxiv.org/html/2608.20516#S8.T7),[15](https://arxiv.org/html/2608.20516#A7.T15),[22](https://arxiv.org/html/2608.20516#A17.T22),[21](https://arxiv.org/html/2608.20516#A16.T21)Blocks A and C share the*same graph cache byte\-for\-byte*: no rebuild occurred between them, which is what makes13\.86513\.865,14\.28914\.289,14\.37714\.377and14\.37914\.379directly comparable and is why the extraction fix of Sec\.[9](https://arxiv.org/html/2608.20516#S9)must be accompanied by re\-measuring all four\. All blocks use the same raw records, the same aspect fields \(claims,methodological\_details,key\_results\), the same masking convention and the same bits measure; chance levels agree to two significant figures and recoverable ceilings to0\.0060\.006bits\.
##### Two provenance caveats we do not smooth over\.
*\(i\)*Table[3](https://arxiv.org/html/2608.20516#S7.T3)reports the input\-side decomposition in*both*feature spaces: column 2 is measured on the block\-A encoder that produced the trained latents in column 3, so the central comparison is within one encoder, and column 1 is retained only because the three ceilings of Sec\.[6](https://arxiv.org/html/2608.20516#S6)were computed on block\-B features\. The block\-A pair is a one\-forward\-pass measurement on the cached graph and required no retraining\. Where the two input columns disagree in magnitude we quote the within\-encoder pair in the text\.*\(ii\)*Our structural\-encoding schema dump prints17,63117\{,\}631citesedges where the stored graph object reports11,79111\{,\}791\. The difference is not a factor of two and is therefore not explained by reverse\-edge addition; we quote11,79111\{,\}791throughout, flag the discrepancy rather than choosing silently, and note that it does not affect Prop\.[5](https://arxiv.org/html/2608.20516#Thmproposition5), under whichcitesis negligible at either count \(0\.53%0\.53\\%or0\.80%0\.80\\%of2,211,1182\{,\}211\{,\}118\)\.
## Appendix A6Extended related work
This appendix expands the four paragraphs of Sec\.[2](https://arxiv.org/html/2608.20516#S2): collapse theory, probe\-based evaluation, retrieval practice, and the reducibility question we are not aware of being asked elsewhere in graph SSL\.
##### Non\-contrastive collapse theory\.
Collapse\-avoidance analyses of BYOL\-style methods\([Grill et al\. 2020](https://arxiv.org/html/2608.20516#bib.bib12)\), VICReg\([Bardes et al\. 2022](https://arxiv.org/html/2608.20516#bib.bib3)\)and Barlow Twins\([Zbontar et al\. 2021](https://arxiv.org/html/2608.20516#bib.bib24)\)characterise conditions under which the encoder does not become constant on its input domain, and dimensional\-collapse analyses\([Jing et al\. 2022](https://arxiv.org/html/2608.20516#bib.bib15)\)characterise when the representation occupies a low\-dimensional subspace\. Both are*global*statements\. The failure we report is invisible to both: the representation occupies many dimensions and varies substantially over inputs, but within each latent category it is constant, and it is the within\-category variation that the downstream task requires\. Prop\.[1](https://arxiv.org/html/2608.20516#Thmproposition1)adds the reason this is not a defect of optimisation: the category\-measurable branch is a global minimum of the coupled objective and the reason no variance\-based regulariser on the*candidate*side removes it: the branch is already non\-constant and already high\-rank across categories\.
##### Supervised class collapse\.
The supervised analogues are well characterised\.[Graf et al\. 2021](https://arxiv.org/html/2608.20516#bib.bib11)derive the optimum of the supervised contrastive loss and show each class maps to a point;[Papyan et al\. 2020](https://arxiv.org/html/2608.20516#bib.bib16)describe the same terminal\-phase geometry under cross\-entropy;[Chen et al\. 2022](https://arxiv.org/html/2608.20516#bib.bib4)separate class collapse from feature suppression and show the former is not a prerequisite for good linear\-probe transfer\. The gap we fill is that all three take the collapsing partition to be given by labels\. In a masked\-prediction JEPA, the partition is supplied by the masking scheme itself and is never named anywhere in the objective, which is exactly why no criterion in the selection toolkit is watching for it\.
##### Probe\-based evaluation\.
A probe measures whatever partition its labels induce, and if that partition coincides with the shortcut the objective took, the probe reports the shortcut as success\. In our case the probe decodes field labels, which correlate with aspect\-level structure, and it reads0\.600\.60–0\.960\.96across configurations whose retrieval performance is uniformly zero\. Sec\.[9](https://arxiv.org/html/2608.20516#S9)adds the converse hazard: a*second*probe chosen to be reasoning\-relevant moves independently of the optimised metric across ten converged cells, so reporting one number from either family is insufficient in both directions\. App\.[A16](https://arxiv.org/html/2608.20516#A16)audits that second probe rather than treating it as a fixed point of reference\.
##### Retrieval evaluation practice and target reducibility\.
In\-batch scoring remains common\. Our difficulty ladder \(Table[2](https://arxiv.org/html/2608.20516#S6.T2)\) shows why this matters quantitatively: bits recovered rise from\+0\.860\+0\.860atK=2K\{=\}2to\+14\.335\+14\.335at the full pool for the same system, while MRR is nearly flat, so a pool\-size\-dependent metric reported without the pool size is close to uninterpretable\. Separately, we are not aware of prior work in graph SSL that audits whether the*evaluation target*’s structure is a deterministic function of the node census\. Prop\.[5](https://arxiv.org/html/2608.20516#Thmproposition5)is elementary, but on our graph it retires the structural interpretation of an otherwise near\-ceiling result, and the check costs one pass over the edge\-type cardinalities\.
## Appendix A7Setup, schema and corpus statistics
This appendix gives the pipeline diagram \(Fig\.[2](https://arxiv.org/html/2608.20516#S1.F2)\), the patch extraction view \(Fig\.[3](https://arxiv.org/html/2608.20516#S1.F3)\), the notation table \(Table[14](https://arxiv.org/html/2608.20516#A7.T14)\), the full edge census with its reducibility audit \(Table[15](https://arxiv.org/html/2608.20516#A7.T15)\), per\-aspect cardinalities \(Table[16](https://arxiv.org/html/2608.20516#A7.T16)\) and the corpus statistics\.
Table 14:Notation\. All effective ranks use Eq\.[5](https://arxiv.org/html/2608.20516#A8.E5)on mean\-centred matrices\. The distinction betweenrcandr\_\{\\mathrm\{cand\}\}andrqueryr\_\{\\mathrm\{query\}\}is central: the former is what standard practice measures, the latter is where the failure lives\.SymbolCode nameMeaningrnoder\_\{\\mathrm\{node\}\}\(/128/128\)node\_rkeffective rank of node latentsrpoolr\_\{\\mathrm\{pool\}\}\(/128/128\)pool\_rkeffective rank of pooled patch vectorsrcandr\_\{\\mathrm\{cand\}\}\(/128/128\)cand\_erankeffective rank of the*candidate bank*rqueryr\_\{\\mathrm\{query\}\}\(/128/128\)query\_eff\_rankeffective rank of the*model’s predictions*DCratio\\mathrm\{DC\}\_\{\\mathrm\{ratio\}\}dc\_ratio‖μ‖/𝔼‖x−μ‖\\\|\\mu\\\|/\\mathbb\{E\}\\\|x\-\\mu\\\|DCenergy\\mathrm\{DC\}\_\{\\mathrm\{energy\}\}dc\_energy‖μ‖2/\(‖μ‖2\+𝔼‖x−μ‖2\)\\\|\\mu\\\|^\{2\}/\(\\\|\\mu\\\|^\{2\}\+\\mathbb\{E\}\\\|x\-\\mu\\\|^\{2\}\)ρasp,ρpap\\rho\_\{\\mathrm\{asp\}\},\\rho\_\{\\mathrm\{pap\}\}rho\_between\_\*variance share between aspects / papers\|𝒜\|\|\\mathcal\{A\}\|—number of latent categories;33hereκ^\\hat\{\\kappa\}target\_kappapower\-law decay of the candidate spectrumε^\\hat\{\\varepsilon\}from loss floorpredictor error relative to the identity signalA1a1\_accheld\-out reasoning probe \(App\.[A16](https://arxiv.org/html/2608.20516#A16)\)ℬ\\mathcal\{B\}bits\_recovered1Nlog2N\!−𝔼\[log2r\]\\frac\{1\}\{N\}\\log\_\{2\}N\!\-\\mathbb\{E\}\[\\log\_\{2\}r\];00is chanceℬmax\\mathcal\{B\}^\{\\max\}bits\_max1Nlog2N\!\\frac\{1\}\{N\}\\log\_\{2\}N\!; the achievable ceilingTable 15:Heterogeneous graph statistics \(57,90357\{,\}903maskable papers\)\. The right\-hand column is the reducibility audit of Sec\.[9](https://arxiv.org/html/2608.20516#S9): a relation whose cardinality equals an endpoint type’s node count is a deterministic function of the census and carries zero information\.Node typecountEdge typecountcensus\-determined?paper57,903\(paper, has\_claim, claim\)251,938yes \(=\|=\|claim\|\|\)claim251,938\(paper, has\_method, method\)57,903yes \(=\|=\|paper\|\|\)method57,903\(paper, has\_result, result\)57,903yes \(=\|=\|paper\|\|\)result57,903\(method, produces, result\)57,903yes \(=\|=\|paper\|\|\)evidence373,438\(result, grounds, claim\)251,938yes \(=\|=\|claim\|\|\)implication251,938\(claim, supported\_by, evidence\)251,922no \(data\-derived\)field488\(claim, challenged\_by, evidence\)121,516no \(data\-derived\)\(claim, implies, implication\)251,938yes \(=\|=\|claim\|\|\)\(paper, cites, paper\)11,791no;0\.53%0\.53\\%of2,211,1182\{,\}211\{,\}118Table 16:Per\-aspect membership and untrained in\-degree\. The singleton rows are what make the pooling control decisive\.Aspectmembers/patch \(mean\)maxpatches with≥2\\geq 2mean in\-degreeroleclaim4\.354\.35101099\.5%99\.5\\%4\.484\.48maskable, multi\-membermethod1\.001\.00110\.0%0\.0\\%2\.002\.00maskable, singletonresult1\.001\.00110\.0%0\.0\\%6\.356\.35maskable, singleton##### Corpus statistics\.
The graph is built from58,14958\{,\}149readable records;246246lack a required aspect and are excluded, yielding57,90357\{,\}903papers with full coverage, while the diagnostic harness applies the weaker non\-empty\-text criterion and retains58,14558\{,\}145\. Empty fields are rare \(claim33,method33,result44\)\. Mean text lengths in the diagnostic set are19551955,16061606and12191219characters for claim, method and result\. The corpus spans2626coarse subject categories, the largest being Medicine \(16,99816\{,\}998\), Agricultural \(5,1745\{,\}174\) and Environmental \(4,5624\{,\}562\)\. Citation edges are restricted to the intra\-corpus subset:11,79111\{,\}791references resolve inside the corpus and2,199,3272\{,\}199\{,\}327dangling references are dropped; coverage at source is high \(52,056/58,14952\{,\}056/58\{,\}149files carry non\-empty referenced works\), so the sparsity ofcitesreflects corpus closure rather than extraction failure—but the consequence for Prop\.[5](https://arxiv.org/html/2608.20516#Thmproposition5)is the same either way\.
## Appendix A8Bits recovered and effective rank
This appendix defines the measure, explains why the ceiling is1Nlog2N\!\\frac\{1\}\{N\}\\log\_\{2\}N\!rather thanlog2N\\log\_\{2\}N, records the correction relative to an earlier version, and gives the effective\-rank definition used throughout\.
##### The measure\.
For a query whose gold item attains rankrrin a pool ofNNcandidates, define
ℬ=1Nlog2N\!⏟𝔼chance\[log2r\]−𝔼\[log2r\]\.\\mathcal\{B\}\\;=\\;\\underbrace\{\\tfrac\{1\}\{N\}\\log\_\{2\}N\!\}\_\{\\textstyle\\mathbb\{E\}\_\{\\mathrm\{chance\}\}\[\\log\_\{2\}r\]\}\\;\-\\;\\mathbb\{E\}\[\\log\_\{2\}r\]\.\(3\)The subtracted term is exactly the expected log\-rank of a uniform random ranker: ifrris uniform on\{1,…,N\}\\\{1,\\dots,N\\\}then𝔼\[log2r\]=1N∑k=1Nlog2k=1Nlog2N\!\\mathbb\{E\}\[\\log\_\{2\}r\]=\\frac\{1\}\{N\}\\sum\_\{k=1\}^\{N\}\\log\_\{2\}k=\\frac\{1\}\{N\}\\log\_\{2\}N\!\. Henceℬ=0\\mathcal\{B\}=0at chance*identically*, not asymptotically, and a ranker that always returns the gold item first \(r≡1r\\equiv 1\) attains
ℬmax=1Nlog2N\!=log2\(\(N\!\)1/N\)=log2N−log2e\+log2\(2πN\)2N\+O\(N−2\)\.\\mathcal\{B\}^\{\\max\}\\;=\\;\\tfrac\{1\}\{N\}\\log\_\{2\}N\!\\;=\\;\\log\_\{2\}\\\!\\big\(\(N\!\)^\{1/N\}\\big\)\\;=\\;\\log\_\{2\}N\-\\log\_\{2\}e\+\\tfrac\{\\log\_\{2\}\(2\\pi N\)\}\{2N\}\+O\(N^\{\-2\}\)\.\(4\)
##### Why the ceiling is notlog2N\\log\_\{2\}N\.
The index entropylog2N\\log\_\{2\}Nis the cost of*naming*one ofNNitems; it is not attainable as a bits\-recovered value, because the reference point is the geometric rather than arithmetic mean of the ranks:\(N\!\)1/N≈N/e\(N\!\)^\{1/N\}\\approx N/e, notNN\. The gap islog2e=1\.4427\\log\_\{2\}e=1\.4427bits and is essentially constant aboveN≈103N\\approx 10^\{3\}\. Usinglog2N\\log\_\{2\}Nassigns a uniform random ranker\+1\.4427\+1\.4427bits instead of00and understates every system’s fraction of the recoverable budget—by a factor that grows as the pool shrinks\. Table[2](https://arxiv.org/html/2608.20516#S6.T2)makes this visible: atK=2K\{=\}2the two ceilings are0\.8620\.862and1\.5851\.585, so the same measurement reads99\.8%99\.8\\%or54\.3%54\.3\\%depending on which is used\.
##### Correction relative to an earlier version\.
An earlier version of this manuscript reportedℬmax=log2N\\mathcal\{B\}^\{\\max\}=\\log\_\{2\}N\. Every*measured*value is unchanged: the harness always subtracted the chance term of Eq\.[3](https://arxiv.org/html/2608.20516#A8.E3), so the recordedℬ\\mathcal\{B\}column was correct throughout\. What was wrong was the ceiling those values were compared against, the percentages derived from it, and one cell of the phase table \(Table[4](https://arxiv.org/html/2608.20516#S7.T4),ρ=0\.9961\\rho=0\.9961\), whose “bits lost” entry alone had been computed against14\.314\.3; that inconsistency also made the printed row non\-monotone inρ\\rho, which is impossible\. All ceilings and percentages here use Eq\.[4](https://arxiv.org/html/2608.20516#A8.E4), and the allocation’s share of the loss is consequently86%86\\%rather than87%87\\%\.
##### Verification\.
The released module prints Eq\.[4](https://arxiv.org/html/2608.20516#A8.E4)for every pool size used in this paper, checks the exact summation against a log\-gamma evaluation, and confirms by Monte Carlo over4×1054\\times 10^\{5\}synthetic queries per pool that a uniform random ranker recoversℬ=0\\mathcal\{B\}=0to within0\.0020\.002bits while a perfect ranker recovers exactlyℬmax\\mathcal\{B\}^\{\\max\}\. It also reports what the discarded formula would have given at chance, namely\+1\.4427\+1\.4427bits at every pool size\. The self\-test of App\.[A9](https://arxiv.org/html/2608.20516#A9)re\-verifies this against the two ceilings actually printed by the training code,ℬmax\(3000\)=10\.108\\mathcal\{B\}^\{\\max\}\(3000\)=10\.108andℬmax\(57,903\)=14\.379\\mathcal\{B\}^\{\\max\}\(57\{,\}903\)=14\.379\.
##### Why bits rather than MRR\.
Eq\.[3](https://arxiv.org/html/2608.20516#A8.E3)is additive in the information\-theoretic sense, comparable across pool sizes, and does not saturate: BM25’s MRR moves by0\.0050\.005fromK=10K\{=\}10to the full pool while its bits move by12\.05112\.051\. The same insensitivity is why we treatMRR=1\.0000\\mathrm\{MRR\}=1\.0000as a warning rather than a result \(App\.[A14](https://arxiv.org/html/2608.20516#A14)\)\.
##### Effective rank\.
ForX∈ℝn×dX\\in\\mathbb\{R\}^\{n\\times d\}we take the singular values of the mean\-centred matrix, normalise topi=σi/∑jσjp\_\{i\}=\\sigma\_\{i\}/\\sum\_\{j\}\\sigma\_\{j\}, and define
erank\(X\)=exp\(−∑ipilogpi\),\\operatorname\{erank\}\(X\)=\\exp\\\!\\Big\(\-\\sum\_\{i\}p\_\{i\}\\log p\_\{i\}\\Big\),\(5\)following[Roy & Vetterli 2007](https://arxiv.org/html/2608.20516#bib.bib18), interpolating between11anddd; intervals are bootstrapped over10310^\{3\}row resamples\. Because this is a spectral\-entropy measure rather than an algebraic rank, values*below*an algebraic ceiling are expected when one singular direction dominates, which is why several target ranks in Table[30](https://arxiv.org/html/2608.20516#A25.T30)fall below2\.02\.0, and why theerank≤\|𝒜\|\\operatorname\{erank\}\\leq\|\\mathcal\{A\}\|bound of Prop\.[1](https://arxiv.org/html/2608.20516#Thmproposition1)is consistent with the measuredrquery=1\.9r\_\{\\mathrm\{query\}\}=1\.9rather than requiring exactly33\. This paper computes Eq\.[5](https://arxiv.org/html/2608.20516#A8.E5)on four distinct matrices—node latents, pooled patch vectors, the candidate bank and the query bank—and only the last detects the failure\.
## Appendix A9The instrument self\-test
Table[17](https://arxiv.org/html/2608.20516#A9.T17)lists the2020assertions the harness runs before any measurement is taken, and the two paragraphs that follow record which of them were added because an earlier version produced a wrong number, and one way in which a self\-test can itself conceal a refutation\.
Table 17:The2020assertions that must pass before any measurement is taken\. The harness refuses to proceed on failure\. Items marked†\\daggerwere added after an earlier version of the pipeline produced a confidently wrong number that the assertion would have caught; we list them in that spirit rather than as boilerplate\.EstimatorAssertionExpectedAUCperfect / inverted / constant / random1\.01\.0/0\.00\.0/0\.5000\.500/≈0\.5\{\\approx\}0\.5AUC ties†constant scorer, ties averagedexactly0\.5000\.500balanced accuracyalways\-majority on95:595\{:\}50\.5000\.500ℬmax\\mathcal\{B\}^\{\\max\}†matches the training code’s printed ceilings10\.10810\.108,14\.37914\.379retrievalperfect / random bitsℬmax\\mathcal\{B\}^\{\\max\}/≈0\{\\approx\}0linear probeseparable / random labels\>0\.90\>0\.90/≈0\.5\{\\approx\}0\.5group split†folds share no groupdisjointgenericnessrandom / tight cluster<0\.2<0\.2/\>0\.9\>0\.9duplicate census300300planted duplicates of10001000exactly300300target masking†all target relations, both directionsemptyunrelated relationsuntouched by maskingpreservedtarget frameunfittedapply\(\)raisesRuntimeErrorInfoNCEaligned loss<<shuffled losstrueto\_heteroprecondition†every node type is a destinationrejects otherwiselazy parameters†none at constructionall materialisedEMA target†identical to encoder at step00bitwise equalrelation cue33distinct vectors at initnon\-degenerateforward pass, masked†survives empty target relationsfiniteleaf isolation†masked target invariant to*all*context featuresbitwisecensus identity373,438=251,922\+121,516373\{,\}438=251\{,\}922\+121\{,\}516exactThe last assertion is the one that matters\. Under target\-edge masking an evidence node is a leaf, so its representation must be a function of its own features alone; we verify this by perturbing*every*context feature by a constant and requiring the masked output to be bitwise unchanged\. Incomplete masking is the single failure mode capable of manufacturing a false positive, and it is exactly what produced the1\.00001\.0000\-AUC reading in Table[7](https://arxiv.org/html/2608.20516#S8.T7)before masking was applied\.
##### A self\-test can also conceal\.
An earlier version of this table asserted a paper\-level count taken from a*different*label \(24,98324\{,\}983rather than the measured35,65035\{,\}650\)\. The assertion passed, because the literal was internally consistent, and it thereby concealed the refutation reported in Sec\.[9](https://arxiv.org/html/2608.20516#S9)\. A test that validates a hard\-coded constant rather than a measured quantity is worse than no test; that assertion now takes the measured census as an argument\. The same principle applies to the ceilings: theℬmax\\mathcal\{B\}^\{\\max\}assertion compares against the value the training code*prints*, not against a number typed into the test\.
## Appendix A10Pre\-registered thresholds
Table[18](https://arxiv.org/html/2608.20516#A10.T18)lists every threshold that converts a measurement into a verdict, all of them written to disk before any result was computed\.
Table 18:Decision thresholds, fixed and written to disk before any result was computed and not revised afterwards\. Each converts a measurement into a verdict; publishing them is what makes the verdicts auditable\. The last four are the target\-side gates introduced for Sec\.[9](https://arxiv.org/html/2608.20516#S9)\.NameMeaningValuecollapse\_bitsbelow this many bits, a system is “collapsed”1\.01\.0probe\_encodes\_bitsabove this, the context linearly encodes identity4\.04\.0control\_capable\_bitsabove this, a control demonstrates capability1\.01\.0paired\_alphalevel for paired bootstrap intervals0\.050\.05degenerate\_bits\_epsbelow this gap, a control is degenerate with the model10−310^\{\-3\}sim\_match\_bitstolerance for “simulation reproduces observation”1\.01\.0kappa\_material\_bitsabove this, anisotropy is called load\-bearing0\.50\.5a1\_voidbelow this, the reasoning probe carries no signal0\.700\.70dup\_max\_fracmax exact\-duplicate share of a target relation0\.200\.20generic\_max\_gapmax genericness gap between target classes0\.150\.15min\_target\_edgesmin edges for a data\-derived target2,0002\{,\}000Three consequences\. The collapse criterion is bits\-based rather than rank\-based, which is why Table[4](https://arxiv.org/html/2608.20516#S7.T4)reportsρ=0\.9961\\rho=0\.9961as a partial collapse and why the regression cell of Table[6](https://arxiv.org/html/2608.20516#S8.T6), at0\.3070\.307bits, is labelled collapsed without appeal to its rank statistics\.degenerate\_bits\_epsexists because an earlier version of the harness reported a “positive control” agreeing with the model under test to four decimal places—the same configuration compared to itself; the guard now returns*vacuous*rather than a verdict if nothing eligible survives\. Anddup\_max\_frac/generic\_max\_gapwere fixed before the polarity census was computed, which is the only reason Table[7](https://arxiv.org/html/2608.20516#S8.T7)can be read as a result rather than as a post\-hoc excuse for a null\.
## Appendix A11Proofs
Proofs of the five results of Sec\.[4](https://arxiv.org/html/2608.20516#S4), in the order stated\.
###### Proof of Prop\.[1](https://arxiv.org/html/2608.20516#Thmproposition1)\.
*\(i\)*For fixedccand fixedgg,argminu𝔼\[‖u−g\(t\)‖2∣c\]=𝔼\[g\(t\)∣c\]\\arg\\min\_\{u\}\\mathbb\{E\}\[\\\|u\-g\(t\)\\\|^\{2\}\\mid c\]=\\mathbb\{E\}\[g\(t\)\\mid c\]by the first\-order condition, and the minimising function is this conditional mean pointwise\.
*\(ii\)*By hypothesis there is a measurablea^\\hat\{a\}witha^\(c\)=a\(t\)\\hat\{a\}\(c\)=a\(t\)almost surely—in our pipelinea^\\hat\{a\}is a lookup on the aspect cue\. Setf=γ∘a^f=\\gamma\\circ\\hat\{a\}andg=γ∘ag=\\gamma\\circ a\. Thenf\(c\)−g\(t\)=γ\(a\(t\)\)−γ\(a\(t\)\)=0f\(c\)\-g\(t\)=\\gamma\(a\(t\)\)\-\\gamma\(a\(t\)\)=0almost surely, soR=0R=0; sinceR≥0R\\geq 0everywhere, this is a global minimum\. The pair is realisable in the parametric class: a message\-passing encoder attains a per\-type constant by driving the feature\-dependent path to zero, and the EMA constraint is satisfied becauseθ¯=θ\\bar\{\\theta\}=\\thetaat the stationary point, so the two branches may shareγ\\gamma\. The same argument holds verbatim for the cosine loss, replacing∥⋅∥2\\\|\\cdot\\\|^\{2\}by1−cos1\-\\cos\.
*\(iii\)*Supposeccdeterminestt, i\.e\. there isΨ\\Psiwitht=Ψ\(c\)t=\\Psi\(c\)almost surely; on our corpus the training\-free oracle shows this holds to within99\.3%99\.3\\%of the recoverable budget\. For any injectivegg, takef=g∘Ψf=g\\circ\\Psi; thenR=0R=0andggretainsδ\\delta\. Both branches are global minima and they differ by the full recoverable budget, which is the claimed unidentifiability\.
Finally, in*\(ii\)*the image ofγ\\gammahas at most\|𝒜\|\|\\mathcal\{A\}\|points, so the mean\-centred query bank has at most\|𝒜\|−1\|\\mathcal\{A\}\|\-1non\-zero singular values anderank≤\|𝒜\|\\operatorname\{erank\}\\leq\|\\mathcal\{A\}\|by Eq\.[5](https://arxiv.org/html/2608.20516#A8.E5), whileffis non\-constant wheneverγ\\gammais injective on𝒜\\mathcal\{A\}\. ∎
###### Proof of Prop\.[2](https://arxiv.org/html/2608.20516#Thmproposition2)\.
For fixedcc,argminu𝔼\[‖u−z‖2∣c\]=𝔼\[z∣c\]\\arg\\min\_\{u\}\\mathbb\{E\}\[\\\|u\-z\\\|^\{2\}\\mid c\]=\\mathbb\{E\}\[z\\mid c\]by the first\-order condition2\(u−𝔼\[z∣c\]\)=02\(u\-\\mathbb\{E\}\[z\\mid c\]\)=0; the minimising function is thereforef⋆\(c\)=𝔼\[z∣c\]f^\{\\star\}\(c\)=\\mathbb\{E\}\[z\\mid c\]pointwise\. Substituting Eq\.[1](https://arxiv.org/html/2608.20516#S4.E1)gives𝔼\[z∣c\]=μ\+αa\(c\)\+𝔼\[δ∣c\]\\mathbb\{E\}\[z\\mid c\]=\\mu\+\\alpha\_\{a\(c\)\}\+\\mathbb\{E\}\[\\delta\\mid c\], and the achieved loss is𝔼‖z−𝔼\[z∣c\]‖2=𝔼‖δ−𝔼\[δ∣c\]‖2≤𝔼‖δ‖2\\mathbb\{E\}\\\|z\-\\mathbb\{E\}\[z\\mid c\]\\\|^\{2\}=\\mathbb\{E\}\\\|\\delta\-\\mathbb\{E\}\[\\delta\\mid c\]\\\|^\{2\}\\leq\\mathbb\{E\}\\\|\\delta\\\|^\{2\}, with equality whenccis independent ofδ\\delta\. The hypothesis thatccis weakly informative aboutδ\\deltais what drives𝔼\[δ∣c\]→0\\mathbb\{E\}\[\\delta\\mid c\]\\to 0; as Sec\.[4](https://arxiv.org/html/2608.20516#S4)notes, that hypothesis is false on our corpus, which is why this proposition is stated as the frozen\-target special case and not as the account of our observations\. ∎
###### Proof of Prop\.[3](https://arxiv.org/html/2608.20516#Thmproposition3)\.
Sinceqi≡qq\_\{i\}\\equiv q, the score vector\(s\(T\(q\),T\(cj\)\)\)j=1N\\big\(s\(T\(q\),T\(c\_\{j\}\)\)\\big\)\_\{j=1\}^\{N\}does not depend onii, so its argsort is a single fixed permutationπ\\piof\[N\]\[N\]and the gold item’s rank isri=π−1\(gi\)r\_\{i\}=\\pi^\{\-1\}\(g\_\{i\}\)\. Becauseπ−1\\pi^\{\-1\}is a bijection of\[N\]\[N\]andgig\_\{i\}is uniform on\[N\]\[N\], the rankrir\_\{i\}is also uniform on\[N\]\[N\]; this is the step that makes the conclusion exact rather than approximate\. Therefore
𝔼\[1r\]=1N∑k=1N1k=HNN,𝔼\[log2r\]=1N∑k=1Nlog2k=1Nlog2N\!,\\mathbb\{E\}\\big\[\\tfrac\{1\}\{r\}\\big\]=\\tfrac\{1\}\{N\}\\sum\_\{k=1\}^\{N\}\\tfrac\{1\}\{k\}=\\tfrac\{H\_\{N\}\}\{N\},\\qquad\\mathbb\{E\}\[\\log\_\{2\}r\]=\\tfrac\{1\}\{N\}\\sum\_\{k=1\}^\{N\}\\log\_\{2\}k=\\tfrac\{1\}\{N\}\\log\_\{2\}N\!,and substituting the second identity into Eq\.[3](https://arxiv.org/html/2608.20516#A8.E3)givesℬ=0\\mathcal\{B\}=0exactly\. None of these quantities depends onTT, so no injective re\-metrisation can alter them\. Injectivity ofTTis used only to guarantee thatTTdoes not merge candidates and thereby changeNN\. ∎
The proposition is exact forqi≡qq\_\{i\}\\equiv q, whereas we measurerquery=1\.9r\_\{\\mathrm\{query\}\}=1\.9and self\-similarity0\.0040\.004: nearly, not exactly, constant\. A perturbation bound expressingℬ\\mathcal\{B\}as a function of query\-bank dispersion would close that gap, and we flag its absence rather than treating the exact statement as if it covered the measured regime\. What the measurement does establish directly is the empirical version, rung 5 of Table[1](https://arxiv.org/html/2608.20516#S4.T1): five frames buy0\.4bits0\.4\\text\{ bits\}against a14\.4bits14\.4\\text\{ bits\}deficit\.
###### Proof of Prop\.[4](https://arxiv.org/html/2608.20516#Thmproposition4)\.
By definition the Fréchet mean of a distributionPPon\(ℳ,d\)\(\\mathcal\{M\},d\)isargmin∫u∈ℳd\(u,z\)2𝑑P\(z\)\\arg\\min\_\{u\\in\\mathcal\{M\}\}\\int d\(u,z\)^\{2\}\\,dP\(z\)\. Conditioning onccand minimising pointwise inu=f\(c\)u=f\(c\)gives exactly this variational problem forP=p\(⋅∣c\)P=p\(\\cdot\\mid c\)\. Existence and uniqueness hold on Hadamard manifolds by convexity ofu↦d\(u,z\)2u\\mapsto d\(u,z\)^\{2\}along geodesics\. ∎
###### Proof of Prop\.[5](https://arxiv.org/html/2608.20516#Thmproposition5)\.
Write theℓ\\ell\-layer message\-passing update for nodevvashv\(ℓ\)=ϕ\(ℓ\)\(hv\(ℓ−1\),\{\{hu\(ℓ−1\):u∈𝒩\(v\)\}\}\)h\_\{v\}^\{\(\\ell\)\}=\\phi^\{\(\\ell\)\}\\big\(h\_\{v\}^\{\(\\ell\-1\)\},\\\{\\\!\\\{h\_\{u\}^\{\(\\ell\-1\)\}:u\\in\\mathcal\{N\}\(v\)\\\}\\\!\\\}\\big\)\. The receptive field ofvvafterℓ\\elllayers is determined byEE, and by hypothesisEEis a fixed function of the census𝒩\\mathcal\{N\}, identical across all graphs sharing that census\. Hence for any two graphsG,G′G,G^\{\\prime\}with the same census, the map from the multiset of node features invv’sℓ\\ell\-hop neighbourhood tohv\(ℓ\)h\_\{v\}^\{\(\\ell\)\}is the same function\. Restricting to patchpp: because every relation incident toppis present for all patches withpp’s node types,pp’sℓ\\ell\-hop neighbourhood contains exactlypp’s own nodes plus nodes reachable only through census\-determined relations, whose membership is likewise constant\. ThereforeΦ\(G\)p=Ψ\(\{xu:u∈p\}\)\\Phi\(G\)\_\{p\}=\\Psi\(\\\{x\_\{u\}:u\\in p\\\}\)for a fixedΨ\\Psidepending only on the parameters and the census, and any scores\(Φ\(G\)p,Φ\(G\)p′\)s\(\\Phi\(G\)\_\{p\},\\Phi\(G\)\_\{p^\{\\prime\}\}\)factors through\(\{xu\}u∈p,\{xu\}u∈p′\)\(\\\{x\_\{u\}\\\}\_\{u\\in p\},\\\{x\_\{u\}\\\}\_\{u\\in p^\{\\prime\}\}\)\. The data\-processing inequality then givesI\(gold,score\)≤I\(gold,\{xu\}u∈p\)I\(\\text\{gold\};\\,\\text\{score\}\)\\leq I\(\\text\{gold\};\\,\\\{x\_\{u\}\\\}\_\{u\\in p\}\), which is the information available to a training\-free statistic of the same features\. ∎
Two remarks, both restated in the proposition itself because they are the two most likely misreadings\. The proposition bounds*information*, not*performance*: a trained encoder may still exceed a fixed training\-free statistic by choosing a better metric over identical information, which is what our\+0\.514\+0\.514\-bit margin over the oracle is and why a reader should not treat that margin as a contradiction\. And the hypothesis is checkable in one pass—compare each relation’s cardinality to its endpoint types’ node counts \(Table[15](https://arxiv.org/html/2608.20516#A7.T15)\)\.
## Appendix A12Two mechanisms we eliminated
Both statements below are mathematically correct and were for a substantial period our leading hypotheses\. Both are refuted as explanations by controls in the main text\. We record them because the elimination is part of the contribution\.
###### Lemma 1\(Scalarised targets are rank\-22\)\.
If each patch is summarised by a scalar mean angleαi=1\|Si\|∑j∈Siϕ\(zj\)\\alpha\_\{i\}=\\frac\{1\}\{\|S\_\{i\}\|\}\\sum\_\{j\\in S\_\{i\}\}\\phi\(z\_\{j\}\)and mapped toti=\(coshαi,sinhαi\)t\_\{i\}=\(\\cosh\\alpha\_\{i\},\\sinh\\alpha\_\{i\}\), thenrank\(T\)≤2\\operatorname\{rank\}\(T\)\\leq 2anderank\(T\)≤2\\operatorname\{erank\}\(T\)\\leq 2for any encoder and any pooling operator\.
###### Proof\.
Every row ofTTlies on the image of the one\-parameter curveα↦\(coshα,sinhα\)⊂ℝ2\\alpha\\mapsto\(\\cosh\\alpha,\\sinh\\alpha\)\\subset\\mathbb\{R\}^\{2\}, so the row space lies in a two\-dimensional subspace\. ∎
Why it is not the explanation\.A Euclidean\-cosine pipeline with no hyperbolic map anywhere fails identically \(Table[1](https://arxiv.org/html/2608.20516#S4.T1), rungs 0–3\), and the pathology is present at initialisation \(Fig\.[10](https://arxiv.org/html/2608.20516#A32.F10)\(b\)\)\. Raising the algebraic ceiling from22to128128with orthonormal multi\-anchor targets leaves the*measured*target rank between1\.61\.6and2\.02\.0\(Table[30](https://arxiv.org/html/2608.20516#A25.T30)\)\.
###### Proposition 6\(Mean pooling contracts effective rank\)\.
Let patch members bezi\(p\)=c\+δp\+ηi\(p\)z^\{\(p\)\}\_\{i\}=c\+\\delta\_\{p\}\+\\eta^\{\(p\)\}\_\{i\}withδp,ηi\(p\)\\delta\_\{p\},\\eta^\{\(p\)\}\_\{i\}independent and zero\-mean, covariancesΣδ,Ση\\Sigma\_\{\\delta\},\\Sigma\_\{\\eta\}\. Mean pooling overmmmembers yieldsΣpool\(m\)=Σδ\+1mΣη\\Sigma\_\{\\mathrm\{pool\}\}\(m\)=\\Sigma\_\{\\delta\}\+\\frac\{1\}\{m\}\\Sigma\_\{\\eta\}, soerank\(Σpool\)\\operatorname\{erank\}\(\\Sigma\_\{\\mathrm\{pool\}\}\)is non\-increasing inmmwhenevererank\(Ση\)\>erank\(Σδ\)\\operatorname\{erank\}\(\\Sigma\_\{\\eta\}\)\>\\operatorname\{erank\}\(\\Sigma\_\{\\delta\}\)\.
Why it is not the explanation\.Formethodpatchesm=1\.00m\{=\}1\.00exactly, so mean pooling is the identity map; the measurement confirms it,rnode=47\.42→rpool=47\.31r\_\{\\mathrm\{node\}\}=47\.42\\to r\_\{\\mathrm\{pool\}\}=47\.31, and the query bank still collapses to1\.971\.97\.
Both eliminated hypotheses arecandidate\-sidestories\. Every instrument we reached for—effective rank, anisotropy, DC ratio, spectra, whitening—is a candidate\-side instrument\. We ran the same measurement on the query bank only after every candidate\-side account was exhausted\. It took one line to settle the question\.Instrument the queries\.The analogous lesson in Sec\.[9](https://arxiv.org/html/2608.20516#S9)is one level further out: every instrument in this paper, including the query\-side ones, is a*representation*\-side instrument, and the reducibility of the target is invisible to all of them\.Instrument the target\.
## Appendix A13Per\-rung ladder detail
Table[19](https://arxiv.org/html/2608.20516#A13.T19)gives the post\-hoc frame sweep of rung 5 with the bank statistics that make it interpretable; the two paragraphs after it cover rungs 6 and 4\.
Table 19:Post\-hoc retrieval\-frame sweep\(three seeds, block A\)\. All frames are fitted on the candidate bank and applied to both sides\. The best frame buys0\.4bits0\.4\\text\{ bits\}against a14\.4bits14\.4\\text\{ bits\}deficit, as Prop\.[3](https://arxiv.org/html/2608.20516#Thmproposition3)predicts\. This is a different experiment from the*trained*target\-frame effect of App\.[A14](https://arxiv.org/html/2608.20516#A14)and the two must not be conflated: this sweep re\-metrises an already\-collapsed query bank, whereas that one changes what the objective asks for\.FrameMRRnoteraw1\.7×10−41\.7\\times 10^\{\-4\}—centre2\.1×10−42\.1\\times 10^\{\-4\}shared mean removedall\-but\-the\-top\-112\.6×10−42\.6\\times 10^\{\-4\}bestall\-but\-the\-top\-222\.1×10−42\.1\\times 10^\{\-4\}—ZCA \(fitted on candidates\)2\.2×10−42\.2\\times 10^\{\-4\}full whiteningcandidate bankDC ratio37\.637\.6rcand=18\.1r\_\{\\mathrm\{cand\}\}=18\.1query bankDC ratio162\.2162\.2rquery=1\.9r\_\{\\mathrm\{query\}\}=1\.9, DC energy0\.99990\.9999query self\-similarity0\.0040\.004mean pairwise cosine##### Encoder depth\.
Across depths00–33, node\-level effective rank falls monotonically from54\.954\.9to18\.018\.0while MRR moves from1\.9×10−41\.9\\times 10^\{\-4\}to2\.0×10−42\.0\\times 10^\{\-4\}\. A depth\-00linear encoder with no message passing whatsoever fails identically to a depth\-33GNN; see Fig\.[9](https://arxiv.org/html/2608.20516#A32.F9)\.
##### Whitening\.
Whitening the inputs raises the instance\-variance share of the trained representation from0\.39%0\.39\\%to9\.86%9\.86\\%and the input effective rank to310\.9310\.9of384384\(Table[33](https://arxiv.org/html/2608.20516#A30.T33)\), while MRR stays at chance and the probe falls from0\.8710\.871to0\.7550\.755\. Sec\.[4](https://arxiv.org/html/2608.20516#S4)explains why this null is predicted rather than anomalous: the zero\-risk family of Prop\.[1](https://arxiv.org/html/2608.20516#Thmproposition1)*\(ii\)*depends on the inputs only through category decodability, which whitening preserves\.
## Appendix A14The repair sweep in full
Table[20](https://arxiv.org/html/2608.20516#A14.T20)gives the matched\-budget factorial; the paragraphs after it record a correction, the three unconfounded single\-factor ablations, our treatment of the two saturated cells, and the dose–response experiment we have not run\.
Table 20:The2×22\\times 2over target frame and shared basis at matched budget\(33k steps, one seed, block C, ceiling14\.37914\.379;centerremoves the per\-aspect target mean,*shared*ties the input projection across node types\)\. Read the factorial, not the best cell\. The baseline cell already reaches99\.4%99\.4\\%of ceiling, so only0\.0900\.090bits are available to any factor here; this is the reducibility of Prop\.[5](https://arxiv.org/html/2608.20516#Thmproposition5)showing up as a compressed dynamic range\.Framesharedℬ\\mathcal\{B\}% ofℬmax\\mathcal\{B\}^\{\\max\}A1 proberaw00*\(baseline\)*14\.28914\.28999\.4%99\.4\\%0\.9640\.964center0014\.35914\.35999\.9%99\.9\\%0\.8620\.862raw1111\.80811\.80882\.1%82\.1\\%0\.9850\.985center1114\.37914\.379100\.0%100\.0\\%0\.9630\.963target\-frame main effect\+0\.071\+0\.07179%79\\%of the0\.0900\.090availableshared\-basis main effect−2\.481\-2\.481both\+0\.090\+0\.090interaction\+2\.500\+2\.500learning\-rate schedule, same33k steps\+1\.337\\mathbf\{\+1\.337\}→14\.28912\.952\\\!\\to\\\!14\.289*training\-free oracle*13\.86513\.86596\.4%96\.4\\%—##### A correction we owe the reader\.
An earlier draft reported the target\-frame main effect as\+1\.199\+1\.199bits\. That figure was computed against a baseline of12\.95212\.952bits obtained under an older learning\-rate schedule; at matched schedule the baseline is14\.28914\.289and the frame effect is\+0\.071\+0\.071\. The\+1\.199\+1\.199therefore conflated the frame with the schedule, and the schedule is the larger term\. We correct it here, and we note that the error was invisible until we read a log stage we had previously skipped—which is an argument for the claim–evidence map of App\.[A1](https://arxiv.org/html/2608.20516#A1)rather than against it\.
##### The unconfounded ablations\.
Three further cells share budget, schedule, frame, loss and seed, so each isolates one factor\.*Structural cue*\(66k steps\): aspect embedding14\.37614\.376bits, RWSE14\.20914\.209, none13\.52913\.529—so a33\-entry learned embedding beats a1616\-dimensional RWSE by\+0\.167\+0\.167bits and removing the cue entirely costs\+0\.680\+0\.680\. This is consistent with App\.[A31](https://arxiv.org/html/2608.20516#A31): the RWSE has effective rank1\.091\.09–1\.281\.28of1616because the graph is census\-determined\.*Transduction*\(66k steps\): fitting the target frame on a disjoint half rather than the candidate bank changes bits by\+0\.001\+0\.001, retiring the concern that the frame is a transduction artefact\.*Raw\-skip*\(66k steps\): a learned gate between raw features and encoder output moves fromσ=0\.5\\sigma=0\.5at initialisation toσ=0\.3886\\sigma=0\.3886, i\.e\. toward the encoder, so the encoder is doing work that raw features alone do not\.
Two cells reportMRR=1\.0000\\mathrm\{MRR\}=1\.0000and we do not headline them\.The raw\-skip cell andcenter:1 both reach14\.37914\.379bits,100\.0%100\.0\\%of ceiling\. Our own harness flagsMRR=1\.0000±0\.0000\\mathrm\{MRR\}=1\.0000\\pm 0\.0000as a degeneracy signature, and in the polarity probe that exact reading was the\+0\.1465\+0\.1465\-AUC edge\-type leak \(Table[7](https://arxiv.org/html/2608.20516#S8.T7)\)\. Given Prop\.[5](https://arxiv.org/html/2608.20516#Thmproposition5)and an oracle already at96\.4%96\.4\\%, near\-ceiling is*expected*rather than suspicious—but expected is not the same as verified\. We therefore report the2020k,33\-seed cell \(14\.377±0\.00214\.377\\pm 0\.002,MRR=0\.9996\\mathrm\{MRR\}=0\.9996\) as the paper’s repaired number, and treat the two saturated cells as requiring the candidate\-perturbation leak check before any stronger claim\.
##### The dose–response experiment that would settle the mechanism\.
The harness supports testing the allocation by*intervention*rather than simulation: the per\-aspect mean of the real features is rescaled by a factorα\\alpha, sweeping the achieved aspect shareρ\\rho; the model is retrained at each dose; and bits recovered are measured against the*measured*ρ\\rho\. This yields an empirical critical share to compare against the simulatedρ⋆=0\.99999\\rho^\{\\star\}=0\.99999in the actual feature space\. The design, dose grid and stopping rule are in the pre\-registration file\. It is not reported here, and together with the non\-census\-determined control graph of Sec\.[10](https://arxiv.org/html/2608.20516#S10)it is the measurement we would run first with more compute\.
## Appendix A15Protocol for the checkpoint allocation measurement
Table[5](https://arxiv.org/html/2608.20516#S7.T5)is produced as follows\. Checkpoints are the ones the training loop already writes; no rerun is required\. At each checkpoint we load the encoder, run one forward pass over the block\-A graph cache with the same masking convention as Protocol R, and compute three quantities on the resulting patch representations: the one\-way variance share between aspects \(ρasp\\rho\_\{\\mathrm\{asp\}\}\), the one\-way share between papers \(ρpap\\rho\_\{\\mathrm\{pap\}\}\), and the effective rank of the model’s predictions \(rqueryr\_\{\\mathrm\{query\}\}\) by Eq\.[5](https://arxiv.org/html/2608.20516#A8.E5)\. Bits recovered are then measured under the unmodified Protocol R on the same4,0004\{,\}000fixed queries, so the column is directly comparable to Table[1](https://arxiv.org/html/2608.20516#S4.T1)\. Both variance shares are computed in the block\-A feature space, so the trajectory is internally consistent and does not inherit the cross\-encoder caveat of App\.[A5](https://arxiv.org/html/2608.20516#A5)\.
##### Why the measurement is necessary rather than decorative\.
An endpoint comparison cannot distinguish a degeneracy the objective*created*from one it merely*failed to remove*, and those two readings support different claims—the first about training dynamics, the second about the optimum\. Our own rank trajectory \(Fig\.[10](https://arxiv.org/html/2608.20516#A32.F10)\(b\)\) and untrained rank measurements \(App\.[A29](https://arxiv.org/html/2608.20516#A29)\) already point toward the second\. Because the distinction costs one forward pass per checkpoint and changes what the paper is entitled to say, we regard it as the minimum standard for any claim of the form “training re\-allocates variance,” and Recommendation \(6\) of Sec\.[10](https://arxiv.org/html/2608.20516#S10)states it as such\.
## Appendix A16The held\-out reasoning probe A1
A1 is the second metric of Sec\.[9](https://arxiv.org/html/2608.20516#S9), and because a decoupling claim is only as strong as the metric it decouples from, we specify its construction and audit its floor rather than treating it as a fixed reference point\.
Table 21:The A1 probe: construction, references and floors\. The void line0\.700\.70is pre\-registered \(App\.[A10](https://arxiv.org/html/2608.20516#A10)\): a reading below it is treated as carrying no signal, and no cell in Table[6](https://arxiv.org/html/2608.20516#S8.T6)falls below it\.PropertyValueNotepre\-registered void line0\.700\.70App\.[A10](https://arxiv.org/html/2608.20516#A10)observed range across cells0\.711−0\.9850\.711\-0\.985Table[6](https://arxiv.org/html/2608.20516#S8.T6)##### The construct\-validity caveat, stated plainly\.
If A1 draws on thechallenged\_byrelation, then by Table[7](https://arxiv.org/html/2608.20516#S8.T7)its target is25\.96%25\.96\\%exact\-duplicate placeholder text with a cosine\-only floor of0\.85390\.8539, and the decoupling result of Sec\.[9](https://arxiv.org/html/2608.20516#S9)would then be uninterpretable in*both*directions: neither the presence nor the absence of a relationship with bits recovered could be attributed to reasoning content\. If A1 does not draw on that relation, this paragraph does not apply and the table above says so explicitly\. We separate the two cases rather than leaving the reader to infer which holds, because the same corpus audit that produced Table[7](https://arxiv.org/html/2608.20516#S8.T7)is the reason the question arises\.
## Appendix A17The polarity target: census, gates and probe conditions
Table[22](https://arxiv.org/html/2608.20516#A17.T22)gives the census behind the two failed gates of Table[7](https://arxiv.org/html/2608.20516#S8.T7); the paragraphs after it establish that the duplication is asymmetric and identify the one data\-derived relation that survives\.
Table 22:Census of the data\-derived relations\(block C\)\. Every evidence node is a leaf—373,438=251,922\+121,516373\{,\}438=251\{,\}922\+121\{,\}516exactly—because the source fields are plain strings, so each claim yields one supporting and at most one contradicting evidence node\. This is the structural fact behind the\+0\.1465\+0\.1465\-AUC edge\-type leak\.QuantityValueNoteclaimnodes251,938251\{,\}938supported\_byedges251,922251\{,\}922≈1\{\\approx\}1per claimchallenged\_byedges121,516121\{,\}51648\.2%48\.2\\%of claimsevidencenodes373,438373\{,\}438=251,922\+121,516=251\{,\}922\+121\{,\}516exactlyimpliesedges251,938251\{,\}93811per claim \(census\-determined\)papers with≥1\\geq 1challenge35,65035\{,\}65061\.6%61\.6\\%of corpusexpected under per\-claim assignment—94\.3%94\.3\\%challenges per affected paper3\.413\.41vs\.4\.354\.35claims/paperduplicateevidencerows \(groups≥50\\geq 50\)31,54931\{,\}5498\.45%8\.45\\%of all evidencechallenge edges onto duplicated rows31,54531\{,\}54525\.96%25\.96\\%of challengeslargest single duplicate group3,0313\{,\}031identical stringsuniqueimplicationrows250,528250\{,\}528of251,938251\{,\}938;0\.00%0\.00\\%duplicated##### The duplication is almost perfectly asymmetric\.
31,54931\{,\}549evidence nodes lie in duplicate groups of size≥50\\geq 50, and31,54531\{,\}545challenge edges point at duplicated rows\. These agree to four nodes, so essentially every boilerplate evidence node is a*challenge*node and supporting evidence is nearly free of exact duplication\. That asymmetry is a complete account of the apparent polarity signal without invoking any semantic difference:cos\(claim,support\)=0\.6607\\cos\(\\text\{claim\},\\text\{support\}\)=0\.6607againstcos\(claim,challenge\)=0\.3061\\cos\(\\text\{claim\},\\text\{challenge\}\)=0\.3061withd=\+1\.638d=\+1\.638, and a cosine\-only classifier at0\.85390\.8539AUC\. The genericness gap of\+0\.2377\+0\.2377shows the non\-duplicated remainder is also formulaic, so a duplicate filter alone is insufficient; the fix is a length\-and\-pattern guard at the point where the extracted string is accepted, followed by one rebuild and a re\-measurement of14\.37914\.379,13\.86513\.865and every derived reference point\.
##### What survives\.
impliesis clean:250,528250\{,\}528of251,938251\{,\}938rows unique,0\.00%0\.00\\%in duplicate groups\. It is also census\-determined \(11per claim, Table[15](https://arxiv.org/html/2608.20516#A7.T15)\), so it cannot serve as a*structural*target—but as a*content*target its text is usable, and it is the one data\-derived signal in this corpus that passes the quality gate\. Restricting the target set toimpliesis a three\-line change and is the cheapest way to obtain a data\-derived target that clears our own gates, rather than leaving Sec\.[9](https://arxiv.org/html/2608.20516#S9)to end on a null\.
## Appendix A18Per\-aspect breakdown
Table[23](https://arxiv.org/html/2608.20516#A18.T23)and Fig\.[4](https://arxiv.org/html/2608.20516#A18.F4)give the per\-aspect view that makes the singleton control decisive, and Table[24](https://arxiv.org/html/2608.20516#A18.T24)records the aggregate that concealed it\.
Table 23:Per\-aspect breakdown\(block A, five seeds\)\.methodandresultare exact singletons, so mean pooling is the identity map:rnode≈rpoolr\_\{\\mathrm\{node\}\}\\approx r\_\{\\mathrm\{pool\}\}to two decimals\.zzis the gold item’s standardised similarity margin;z≈0z\\approx 0means the gold candidate is indistinguishable from the pool mean\.Aspectmmrnoder\_\{\\mathrm\{node\}\}rpoolr\_\{\\mathrm\{pool\}\}rqueryr\_\{\\mathrm\{query\}\}MRRMRR \(centred\)marginzzclaim4\.354\.355\.405\.404\.914\.911\.871\.871\.5×10−41\.5\\times 10^\{\-4\}2\.2×10−42\.2\\times 10^\{\-4\}−0\.006\-0\.006method1\.001\.0047\.4247\.4247\.3147\.311\.971\.972\.1×10−42\.1\\times 10^\{\-4\}1\.5×10−41\.5\\times 10^\{\-4\}−0\.001\-0\.001result1\.001\.002\.042\.042\.052\.051\.621\.622\.0×10−42\.0\\times 10^\{\-4\}2\.0×10−42\.0\\times 10^\{\-4\}\+0\.033\+0\.033*Chance*1\.99×10−41\.99\\times 10^\{\-4\}1\.99×10−41\.99\\times 10^\{\-4\}00Figure 4:The control that exonerates pooling and localises the failure to the query side\.\(a\)Formethod,m=1\.00m\{=\}1\.00, so mean pooling is provably the identity map: node rank47\.447\.4carries through to pooled rank47\.347\.3\. The model’s*predictions*for those patches nevertheless occupy effective rank1\.971\.97\(black diamonds\), at or below the33\-category bound of Prop\.[1](https://arxiv.org/html/2608.20516#Thmproposition1)\.\(b\)A23×23\\timesrange in pooled rank produces no movement in bits recovered; dotted line is chance, dashed line the14\.37914\.379\-bit ceiling\. Five seeds\.Table 24:Isolated pre\-target rank measurement\.A mean\-pool JEPA with a Euclidean\-cosine target, no hyperbolic map and no anchors\. Mean±\\pmstd over five seeds\. The per\-aspect breakdown above shows this aggregate conceals a23×23\\timesspread—which is why the aggregate, alone, misled us\.QuantityValue \(/128/128\)Node\-latent effective rankrnoder\_\{\\mathrm\{node\}\}18\.3±0\.918\.3\\pm 0\.9healthyPooled\-latent effective rankrpoolr\_\{\\mathrm\{pool\}\}1\.97±0\.041\.97\\pm 0\.04low, but see Table[23](https://arxiv.org/html/2608.20516#A18.T23)
## Appendix A19Oracle leakage controls
Table[25](https://arxiv.org/html/2608.20516#A19.T25)reports the training\-free oracle per aspect together with three leakage controls:no\-summaryremoves the title/abstract node from the context,bm25\-no\-ovdeletes every55\-gram shared between query and gold, and acheatquery equal to the target’s own embedding validates the metric atMRR=1\.000\\mathrm\{MRR\}=1\.000\. Fig\.[5](https://arxiv.org/html/2608.20516#A19.F5)plots the same numbers\.
Table 25:Oracle retrieval and leakage controls, per aspect, block A\.bm25\-no\-ovis a lower bound on the non\-lexical signal;bm25uses no learned representation\.no\-summaryremoves the paper title/abstract node from the context\. Acheatquery equal to the target’s own embedding validates the metric\.FeaturesAspectctx MRRno\-summarybm25bm25\-no\-ovwhitenedclaim0\.973\\mathbf\{0\.973\}0\.9530\.9530\.9790\.9790\.7320\.732method0\.908\\mathbf\{0\.908\}0\.8340\.8340\.9740\.9740\.8020\.802result0\.938\\mathbf\{0\.938\}0\.9210\.9210\.9890\.9890\.8540\.854rawclaim0\.9430\.9430\.9160\.916——method0\.8780\.8780\.7960\.796——result0\.9250\.9250\.9160\.916——cheat\(metric sanity\)1\.0001\.000———*Trained pipeline, baseline*1\.9×10−41\.9\\times 10^\{\-4\}———*Chance*1\.99×10−41\.99\\times 10^\{\-4\}———Figure 5:Four training\-free or positive\-control retrievers, all near ceiling; the baseline trained model at chance \(dotted\)\.bm25is the strongest evidence because it uses no learned representation whatsoever, andbm25\-no\-ovstill recovers\+14\.250\+14\.250bits—99\.1%99\.1\\%of the14\.38514\.385\-bit ceiling—after deleting every shared55\-gram\.
## Appendix A20Query\-subsample validation
Vector systems are evaluated on4,0004\{,\}000queries and lexical systems on a nested2,0002\{,\}000subsample, because BM25 costs roughly500×500\\timesmore per query than a dot product\. The subsample is validated rather than assumed: the positive control on4,0004\{,\}000queries recovers\+14\.220\+14\.220bits against\+14\.211\+14\.211from the full evaluation, agreement of0\.010\.01bits, and the nested lexical subsample agrees with the4,0004\{,\}000\-query set within the bootstrap interval at every rung of Table[2](https://arxiv.org/html/2608.20516#S6.T2)\. All reported intervals are over queries, so subsampling is reflected in the stated uncertainty\.
## Appendix A21The2×22\\times 2factorial separating feature space from trainer
p1differs from the model under test in two factors at once, so it bounds the harness without attributing the gap\. We therefore pre\-registered a2×22\\times 2design over\{\\\{reference predictor, full pipeline\}×\{\\\}\\times\\\{frozen text features, hetero node features\}\\\}in which exactly one factor differs between comparable cells, and in which a cell that cannot obtain the trainer or features it requested is recorded asprovenance\_mismatchand emits*no verdict*\. Silence is the correct output when a premise is unmet; the alternative—silently substituting a fallback—is what produced the degenerate control described in App\.[A10](https://arxiv.org/html/2608.20516#A10)\.
Table 26:Factorial design\. The two diagonal cells are the systems reported in the main text; the off\-diagonal cells isolate each factor and are not yet reported\. We list them explicitly rather than omitting them, so that the limits of the attribution are visible: on the evidence in this paper,p1bounds the harness and does not separate feature space from trainer\. Both missing cells are33k\-step runs on the existing cache\.frozen text featureshetero node featuresreference predictor\+14\.220\+14\.220bits \(p1\)*not yet reported*full pipeline*not yet reported*0\.000\.00bits \(Table[1](https://arxiv.org/html/2608.20516#S4.T1)\)
## Appendix A22Full phase sweep
Table[27](https://arxiv.org/html/2608.20516#A22.T27)gives the full simulation sweep summarised in Table[4](https://arxiv.org/html/2608.20516#S7.T4), including the leave\-one\-out block that retires anisotropy as the residual factor\.
Table 27:Phase analysis, full sweep\(bank size20,00020\{,\}000, soℬmax=12\.845\\mathcal\{B\}^\{\\max\}=12\.845by Eq\.[4](https://arxiv.org/html/2608.20516#A8.E4)\)\. Row 1 variesρ\\rhowithκ=ε=0\\kappa=\\varepsilon=0; row 2 uses the measuredκ^=0\.942,ε^=1\.515\\hat\{\\kappa\}=0\.942,\\hat\{\\varepsilon\}=1\.515\. The leave\-one\-out block shows that neither nuisance parameter is load\-bearing\.ρ\\rho0\.00400\.00400\.50\.50\.90\.90\.990\.990\.9961\\mathbf\{0\.9961\}0\.9990\.9990\.99990\.9999MRR,ρ\\rhoonly1\.0001\.0001\.0001\.0000\.89280\.89280\.10680\.10680\.02040\.02040\.00430\.00430\.00140\.0014ℬ\\mathcal\{B\},ρ\\rhoonly\+12\.845\+12\.845\+12\.845\+12\.845\+12\.200\+12\.200\+3\.833\+3\.833\+1\.911\+1\.911\+0\.843\+0\.843\+0\.266\+0\.266MRR, matchedκ^,ε^\\hat\{\\kappa\},\\hat\{\\varepsilon\}1\.0001\.000———0\.01710\.0171——ℬ\\mathcal\{B\}, matched\+12\.845\+12\.845———\+1\.80\+1\.80——bits lost, matched0\.0000\.000———11\.04511\.045——*Leave\-one\-out at the measured operating point*dropρ\\rho\(set0\.50\.5\): MRR0\.97400\.9740; dropκ^\\hat\{\\kappa\}:1\.00001\.0000; dropε^\\hat\{\\varepsilon\}:1\.00001\.0000; drop all:1\.00001\.0000critical shareρ⋆=0\.99999\\rho^\{\\star\}=0\.99999, i\.e\.1−ρ⋆=5\.1×10−61\-\\rho^\{\\star\}=5\.1\\times 10^\{\-6\}##### Interpretation\.
Theρ\\rho\-only row reproduces the qualitative shape of a collapse driven purely by the variance share\. Adding the measured nuisance parameters atρ=0\.9961\\rho=0\.9961moves MRR from0\.02040\.0204to0\.01710\.0171—far below the pre\-registered materiality threshold—so we reject anisotropy as the residual factor\. Because the criterion for “total collapse” is bits\-based \(App\.[A10](https://arxiv.org/html/2608.20516#A10)\), the critical share is located where recovered bits fall below1\.01\.0, givingρ⋆=0\.99999\\rho^\{\\star\}=0\.99999\. Note that “bits lost” is12\.845−ℬ12\.845\-\\mathcal\{B\}throughout and is therefore monotone inρ\\rho; a non\-monotone loss column is a sign that two different ceilings have been mixed, which is the error corrected in App\.[A8](https://arxiv.org/html/2608.20516#A8)\.
Figure 6:The objective optimises the wrong variance component\.Left: the frozen inputs,86\.05%86\.05\\%of variance between papers\. Right: the trained latents,0\.39%0\.39\\%\. Whitening the inputs raises the trained instance share to9\.86%9\.86\\%and bits recovered still do not move—the invariance Prop\.[1](https://arxiv.org/html/2608.20516#Thmproposition1)predicts; changing the*objective*does \(Table[6](https://arxiv.org/html/2608.20516#S8.T6)\)\.
## Appendix A23The seven\-cell intervention grid
Table[28](https://arxiv.org/html/2608.20516#A23.T28)gives the full aggregator×\\timesloss grid referenced in Sec\.[8](https://arxiv.org/html/2608.20516#S8): seven cells, five seeds each,rpoolr\_\{\\mathrm\{pool\}\}spanning1\.251\.25–2\.032\.03and probe accuracy spanning0\.6040\.604–0\.9640\.964, with bits recovered at00throughout\. Fig\.[7](https://arxiv.org/html/2608.20516#A23.F7)plots it\.
Table 28:Seven interventions, one flat line\(block A, five seeds, Protocol R\)\.rpoolr\_\{\\mathrm\{pool\}\}spans1\.251\.25–2\.032\.03and the probe spans0\.6040\.604–0\.9640\.964; bits recovered stay at00in every cell, against a ceiling of14\.37914\.379\. What this grid does*not*vary is the target frame and the learning\-rate schedule, both of which App\.[A14](https://arxiv.org/html/2608.20516#A14)shows to matter; we therefore present it as a negative result about the aggregator and the in\-batch\-negative form of the loss, not about the pipeline\.PoolingLossrpoolr\_\{\\mathrm\{pool\}\}probeMRRℬ\\mathcal\{B\}mean1−cos1\-\\cos*\(baseline\)*1\.971\.970\.8710\.8711\.9×10−41\.9\\times 10^\{\-4\}0\.000\.00meaninfonce1\.961\.960\.8670\.8672\.1×10−42\.1\\times 10^\{\-4\}0\.000\.00sum1−cos1\-\\cos1\.251\.250\.9630\.9631\.8×10−41\.8\\times 10^\{\-4\}0\.000\.00deepsets1−cos1\-\\cos1\.701\.700\.6040\.6041\.9×10−41\.9\\times 10^\{\-4\}0\.000\.00deepsetsinfonce1\.811\.810\.7420\.7421\.7×10−41\.7\\times 10^\{\-4\}0\.000\.00attninfonce2\.032\.030\.9640\.9642\.3×𝟏𝟎−𝟒\\mathbf\{2\.3\\times 10^\{\-4\}\}0\.000\.00mean1−cos1\-\\cos\+\+tgt\-vic1\.971\.970\.8710\.8711\.9×10−41\.9\\times 10^\{\-4\}0\.000\.00oracleceiling \(block B\)0\.97190\.9719\+14\.281\+14\.281*repaired objective*\(Table[6](https://arxiv.org/html/2608.20516#S8.T6)\)0\.99960\.9996\+14\.377\+14\.377ℬmax\\mathcal\{B\}^\{\\max\}—14\.37914\.379Chance1\.99×10−41\.99\\times 10^\{\-4\}0\.000\.00Figure 7:Seven interventions, one flat line\.Raw and centred cosine, mean±\\pmstd over five seeds\. Dotted line is chance \(ℬ=0\\mathcal\{B\}=0\), dashed line the14\.37914\.379\-bit ceiling\.
## Appendix A24Breadth of the dissociation
Table[29](https://arxiv.org/html/2608.20516#A24.T29)extends the dissociation across twenty objective×\\timestarget\-construction cells, including two that a rank\-based selection criterion would prefer\.
Protocol note\.Every retrieval number in the main text is measured against the full pool\. Retrieval columns are omitted from Tables[29](https://arxiv.org/html/2608.20516#A24.T29)and[30](https://arxiv.org/html/2608.20516#A25.T30)because those runs predate Protocol R and were scored in\-batch; in\-batch MRR over a few hundred candidates overstates full\-pool MRR by roughly two orders of magnitude and is not rescalable\. Probe accuracy and effective rank are protocol\-independent properties of the representation\. This omission is itself one of the paper’s recommendations\.
Table 29:Twenty objective×\\timestarget\-construction cells; the dissociation holds across all of them\.Probe accuracy spans0\.7550\.755–0\.9600\.960and pooled effective rank1\.81\.8–13\.913\.9—a7\.7×7\.7\\timesrange including two cells that clear the conventionalrpool\>10r\_\{\\mathrm\{pool\}\}\>10threshold—with no accompanying change in retrieval\. Mean over five seeds\.ObjectiveTarget constructionProbe accrpool/128r\_\{\\mathrm\{pool\}\}/128Rank verdicthyperbolic, 2\-Dmean \(own only\)0\.8400\.8401\.91\.9collapsedrel\-hetero\+\+white0\.8220\.8223\.13\.1collapsedrel\-hetero \(raw\)0\.9400\.9401\.81\.8collapsedhyperbolic\+\+vicmean \(own only\)0\.8590\.8592\.02\.0collapsedrel\-hetero\+\+white0\.8120\.81213\.9\\mathbf\{13\.9\}rank\-recoveredrel\-hetero \(raw\)0\.9540\.9542\.22\.2collapsedhyperbolic\+\+vic\+mean \(own only\)0\.8580\.8582\.02\.0collapsedrel\-hetero\+\+white0\.8210\.82112\.1\\mathbf\{12\.1\}rank\-recoveredrel\-hetero \(raw\)0\.9600\.9602\.22\.2collapsedhyperbolic, full reg\.mean \(own only\)0\.8710\.8711\.91\.9collapsedrel\-hetero\+\+white0\.8080\.8083\.93\.9collapsedrel\-hetero \(raw\)0\.9470\.9471\.81\.8collapsedLorentzmean \(own only\)0\.8720\.8722\.02\.0collapsedrel\-hetero\+\+white0\.8080\.8084\.44\.4collapsedrel\-hetero \(raw\)0\.9480\.9481\.81\.8collapsedEuclidean cosinemean, raw0\.8710\.8712\.02\.0collapsedmean, whitened0\.7550\.7553\.83\.8collapsedrel\-hetero, raw0\.9490\.9491\.81\.8collapsedrel\-hetero, whitened0\.8030\.8034\.24\.2collapsedrel\-hetero, wh\., claim\-only0\.8030\.8034\.24\.2collapsed*Range across all twenty cells*0\.7550\.755–0\.9600\.9601\.81\.8–13\.913\.9##### What this table adds\.
The dissociation is not specific to the Euclidean\-cosine objective: it holds for a unit\-hyperbola target, a Lorentz\-model target, and two strengths of VICReg\. It survives a target construction that deliberately injects paper\-specific structural variance from cross\-type reasoning and citation neighbours\. And two cells clear the conventionalrpool\>10r\_\{\\mathrm\{pool\}\}\>10threshold and would be*preferred*by a rank\-based selection criterion such as RankMe\([Garrido et al\. 2023](https://arxiv.org/html/2608.20516#bib.bib10)\); both are among the worst cells by probe accuracy, and neither retrieves\. Rank and probe accuracy are not merely insufficient here—they disagree with each other\. Note also that none of these twenty cells varies the loss form or the learning\-rate schedule, which are the two axes App\.[A14](https://arxiv.org/html/2608.20516#A14)identifies as load\-bearing: a2020\-cell sweep can miss the operative variable entirely\.
## Appendix A25The anchor\-dimension study
Table[30](https://arxiv.org/html/2608.20516#A25.T30)raises the target’s algebraic rank ceiling from22to128128and shows the measured target rank does not follow\.
Table 30:Raising the target’s algebraic rank ceiling from22to2k2kdoes not raise the measured target rank\.Each patch angle is replaced bykkangles against QR\-orthonormalised anchors\. The ceiling rises to128128atk=64k\{=\}64and the measured target effective rankrtgtr\_\{\\mathrm\{tgt\}\}does not move\.ObjectiveTargetProbe accrpool/128r\_\{\\mathrm\{pool\}\}/128rtgtr\_\{\\mathrm\{tgt\}\}hyperbolic, 2\-Dmean \(2\-D\)0\.8400\.8401\.91\.92\.02\.0\(ceiling\)k=4k\{=\}4\(→ℝ8\\to\\mathbb\{R\}^\{8\}\)0\.8690\.8691\.91\.91\.61\.6k=16k\{=\}16\(→ℝ32\\to\\mathbb\{R\}^\{32\}\)0\.8690\.8691\.91\.91\.91\.9k=64k\{=\}64\(→ℝ128\\to\\mathbb\{R\}^\{128\}\)0\.8670\.8671\.91\.91\.91\.9hyperbolic\+\+vicmean \(2\-D\)0\.8590\.8592\.02\.02\.02\.0\(ceiling\)k=4k\{=\}40\.8590\.8592\.02\.01\.71\.7k=16k\{=\}160\.8570\.8572\.02\.01\.91\.9k=64k\{=\}640\.8600\.8602\.02\.02\.02\.0hyperbolic, full reg\.mean \(2\-D\)0\.8710\.8711\.91\.92\.02\.0\(ceiling\)k=4k\{=\}40\.8690\.8691\.91\.91\.61\.6k=16k\{=\}160\.8690\.8691\.91\.91\.91\.9k=64k\{=\}640\.8670\.8671\.91\.91\.91\.9Lorentzmean \(2\-D\)0\.8720\.8722\.02\.02\.02\.0\(ceiling\)k=4k\{=\}40\.8690\.8691\.91\.91\.61\.6k=16k\{=\}160\.8690\.8691\.91\.91\.91\.9k=64k\{=\}640\.8670\.8671\.91\.91\.91\.9*Algebraic ceiling atk=64k\{=\}64*——128128*Measured range*——1\.61\.6–2\.02\.0##### Why the anchor result matters for the query\-side account\.
The anchor map is injective by construction and its output dimension is swept over a16×16\\timesrange, yetrtgtr\_\{\\mathrm\{tgt\}\}stays pinned between1\.61\.6and2\.02\.0\. A map cannot manufacture variation its input does not contain: the patch latents reaching the anchors already lie on a near\-one\-dimensional set\. In hindsight, the flatrtgtr\_\{\\mathrm\{tgt\}\}column was the query\-side signal showing through a candidate\-side measurement, and Prop\.[1](https://arxiv.org/html/2608.20516#Thmproposition1)says why no injective re\-parameterisation of the target could have changed it\.
## Appendix A26Synthetic separability control
To test whether high candidate\-bank redundancy alone impairs rankability, we generate banks with controlled shared\-mean dominance and query them with the*true*context vector, so only bank separability is measured\. Retrieval stays atMRR=1\.000\\mathrm\{MRR\}=1\.000across the entire range, up to a DC ratio of100100, far beyond anything our corpus exhibits \(candidate DC ratio37\.637\.6\)\. A bank with mean pairwise cosine near11remains perfectly rankable provided the residual is consistent between context and target\. This retires “the data are too similar” quantitatively rather than rhetorically, and distinguishes*similarity*from*indistinguishability*\.
## Appendix A27Extraction and templating audit
Table[31](https://arxiv.org/html/2608.20516#A27.T31)collects the templating audit of Sec\.[8](https://arxiv.org/html/2608.20516#S8)and the placeholder\-text audit of Sec\.[9](https://arxiv.org/html/2608.20516#S9)in one place, because the two measure different pathologies and only one of them is benign\.
Table 31:Templating audit on the58,14558\{,\}145set; text probes on an8,0008\{,\}000\-document subsample, variance probes on all records\. The function\-word classifier uses a fixed stopword vocabulary and cannot access content\. The final block is the placeholder\-text audit of Sec\.[9](https://arxiv.org/html/2608.20516#S9), on the block\-C relations\.ProbeDetailValueaspect classification, function words onlyfixed stopword vocabulary0\.9370\.937aspect classification, full TF\-IDF5050k features0\.9860\.986chancethree aspects0\.3330\.333templating index\(acc−chance\)/\(1−chance\)\(\\text\{acc\}\-\\text\{chance\}\)/\(1\-\\text\{chance\}\)0\.9060\.906top44\-gram coverage,claim*“this suggests that the”*31\.8%31\.8\\%top44\-gram coverage,method*“the study employed a”*39\.0%39\.0\\%top44\-gram coverage,result*“the study found that”*25\.7%25\.7\\%ρasp\\rho\_\{\\mathrm\{asp\}\}, raw inputsone\-way, aspect grouping0\.00400\.0040ρasp\\rho\_\{\\mathrm\{asp\}\}, per\-aspect mean removedde\-templated0\.00000\.0000ρpap\\rho\_\{\\mathrm\{pap\}\}, raw inputsone\-way, paper grouping0\.86050\.8605oracle after de\-templatingMRR / bits0\.9720\.972/\+14\.28\+14\.28ridge probe, context→\\tooracle spaceMRR / bits0\.85340\.8534/\+13\.868\+13\.868loss floor ratioℒ∞/𝔼‖δ‖2\\mathcal\{L\}\_\{\\infty\}/\\mathbb\{E\}\\\|\\delta\\\|^\{2\}—3\.2963\.296median predictor gradient norm—4\.61×10−14\.61\\times 10^\{\-1\}*Placeholder\-text audit, block C relations \(Sec\.[9](https://arxiv.org/html/2608.20516#S9)\)*exact\-duplicate share,challenged\_bygroups of≥50\\geq 50identical rows25\.96%25\.96\\%exact\-duplicate share,impliessame criterion0\.00%0\.00\\%largest duplicate groupidentical strings3,0313\{,\}031genericness, supporting evidencemean cos to own centroid0\.29210\.2921genericness, contradicting evidencemean cos to own centroid0\.52980\.5298genericness gapchallenge−\-support\+0\.2377\+0\.2377cos\(claim,support\)\\cos\(\\text\{claim\},\\text\{support\}\)raw features0\.66070\.6607cos\(claim,challenge\)\\cos\(\\text\{claim\},\\text\{challenge\}\)raw features0\.30610\.3061Cohen’sddsupport vs\. challenge\+1\.638\+1\.638Interpretation of the upper blocks is in Sec\.[8](https://arxiv.org/html/2608.20516#S8): templating is severe and inflates*aspect\-type*separability, but paper identity survives its removal almost intact—the de\-templated oracle still recovers99\.3%99\.3\\%of the14\.38514\.385\-bit ceiling\. We report the0\.9370\.937figure prominently rather than in passing, because a reader who discovers it independently would reasonably treat its absence as concealment, and because it is the empirical reason the designator\-determines\-category premise of Prop\.[1](https://arxiv.org/html/2608.20516#Thmproposition1)is so easily satisfied on this corpus\.
##### The two audits measure different pathologies, and only one is benign\.
The templating audit says the*surface form*of an aspect is predictable; the placeholder audit says the*content*of one relation is frequently absent while the field is nominally populated\. The first inflates a nuisance dimension and is neutralised by removing the per\-aspect mean, after which the recoverable identity is essentially unchanged \(\+14\.28\+14\.28bits\)\. The second cannot be neutralised by any transformation of the features, because the information was never extracted:3,0313\{,\}031identical strings carry one string’s worth of information regardless of how they are embedded\. This is the distinction we would have missed had we run only the templating audit, and it is why we recommend both\.
##### Why exact\-duplicate hashing was not sufficient on its own\.
Hashing detects only identical rows\. The genericness statistic—mean cosine of each class’s members to their own centroid—detects the formulaic\-but\-unique remainder, and it is the larger effect here: after the25\.96%25\.96\\%of exactly duplicated contradicting\-evidence rows are set aside, the surviving text is still\+0\.2377\+0\.2377more self\-similar than supporting evidence\. A pipeline that filtered duplicates and declared the target clean would have trained on filler and reported a0\.850\.85\-AUC “polarity” result\. Both statistics are one pass over the embedding matrix and cost under a minute at this corpus size \(App\.[A33](https://arxiv.org/html/2608.20516#A33)\)\.
## Appendix A28Faithfulness of the reference implementation
Table[32](https://arxiv.org/html/2608.20516#A28.T32)reports our re\-implementation on the original method’s graph\-classification benchmarks\. We do not clear MUTAG and therefore do not claim implementation faithfulness\.
Table 32:Our reference re\-implementation on the original method’s graph\-classification benchmarks\. We do not clear MUTAG and therefore do not claim implementation faithfulness\.DatasetoursreportedverdictMUTAG0\.6860\.6860\.8740\.874not reproducedPROTEINS0\.7390\.7390\.7500\.750within toleranceTraining is additionally unstable on PROTEINS, with the objective increasing over the run and predictor gradient norms reachingO\(103\)O\(10^\{3\}\)\. Three consequences\. We scope every claim to the pipeline we describe and document, not to the original method’s published configuration\. All conclusions rest on*internal*controls—the oracle, the lexical bound, the positive control and the repair sweep all share the pipeline under test and the identical evaluation—so the diagnosis does not depend on having matched an external benchmark\. And the two structural findings of Sec\.[9](https://arxiv.org/html/2608.20516#S9)are, if anything,*independent*of implementation fidelity: Prop\.[5](https://arxiv.org/html/2608.20516#Thmproposition5)is a statement about the graph’s edge cardinalities, which no choice of encoder can change, and the placeholder\-text census is a statement about the corpus, which no choice of objective can change\. We regard disclosing the gap as a precondition for the central claim being taken seriously, and flag closing it as required work before the method itself is characterised\.
## Appendix A29Untrained rank and degree
This appendix reports rank and DC ratio per node type*before*any training, and is one of the two measurements bearing on the trajectory question of Sec\.[7\.1](https://arxiv.org/html/2608.20516#S7.SS1)\.
Before any training we measure effective rank and DC ratio per node type at encoder depths00–33\. At depth00all node types have effective rank4242–5959and DC ratio3\.13\.1–3\.63\.6\. Each additional message\-passing layer lowers rank and raises DC ratio for every node type: at depth33, ranks fall to1010–2929and DC ratios rise to2\.42\.4–12\.612\.6\. The compression is therefore a property of message passing at initialisation, not of the objective consistent with Fig\.[10](https://arxiv.org/html/2608.20516#A32.F10)\(b\), where pooled rank is at its final value at epoch00\. Together these two measurements are why Sec\.[7\.1](https://arxiv.org/html/2608.20516#S7.SS1)measures the variance allocation per checkpoint instead of inferring a dynamic re\-allocation from the endpoints\.
The relation between rank and in\-degree is*not*monotoneevidencehas mean in\-degree1\.001\.00and the lowest rank of any type, whilefieldhas in\-degree118\.65118\.65and rank2626–3434so we make no claim of a degree rank law\. Theevidencerow deserves a second look in light of Sec\.[9](https://arxiv.org/html/2608.20516#S9): its in\-degree of exactly1\.001\.00is the leaf property \(373,438=251,922\+121,516373\{,\}438=251\{,\}922\+121\{,\}516\), which is simultaneously why it has the lowest rank, why mean pooling over it is trivial, and why leaving a single target relation unmasked produces the\+0\.1465\+0\.1465\-AUC leak of Table[7](https://arxiv.org/html/2608.20516#S8.T7)\. One structural fact, three separate symptoms, which we did not connect until the self\-test’s leaf\-isolation assertion \(App\.[A9](https://arxiv.org/html/2608.20516#A9)\) forced the question\.
## Appendix A30Input\-feature anisotropy
Table[33](https://arxiv.org/html/2608.20516#A30.T33)records the whitened input ranks that rung 4’s rank\-restoration claim rests on, together with a logging gap we do not paper over\.
Table 33:Effective rank of the frozen input features \(of384384\) per node type after PCA/ZCA whitening, applied per node type before the encoder\. The per\-node\-type*raw*input ranks were not written to the released logs for this run and we therefore do not tabulate them; raw input anisotropy is instead characterised by the depth\-00DC ratios of App\.[A29](https://arxiv.org/html/2608.20516#A29)\(3\.13\.1–3\.63\.6\), and the aggregate raw\-to\-whitened change is the310\.9310\.9\-of\-384384figure quoted for rung 4 of Table[1](https://arxiv.org/html/2608.20516#S4.T1)\. The whitened values below are the ones the ladder’s rank\-restoration claim rests on\.claimmethodresultpaperwhitened320\.8320\.8309\.2309\.2313\.1313\.1313\.1313\.1The missing raw column is a genuine gap in our logging rather than a selective omission, and we prefer to say so than to recompute it from a rerun and present it as if it were the original measurement\. It affects only the precision of rung 4’s description, not its conclusion: whitening raises rank and*lowers*the probe while leaving bits at zero, and both of those are measured in the same run\.
## Appendix A31Random\-walk structural encoding audit
This appendix records the RWSE statistics, and then a correction to how we originally read them\.
The context mixer consumes a random\-walk structural encoding on the intra\-paper reasoning subgraph\. Every paper has at least one intra\-paper reasoning edge \(100%100\\%coverage\); the mean number of reasoning edges per subgraph is32\.3032\.30; the global patch\-RWSE standard deviation is0\.19590\.1959; the fraction of patches with non\-zero RWSE is100%100\\%; and the encoding uses1616random\-walk steps\. The encoding is therefore non\-degenerate at input, and the collapse cannot be attributed to an all\-zero structural encoding\.
A correction to how we previously read this audit\.We originally cited100%100\\%coverage and non\-zero variance as evidence that the structural encoding is*informative*\. Prop\.[5](https://arxiv.org/html/2608.20516#Thmproposition5)shows that inference was wrong\. The intra\-paper edge set is a deterministic function of the node census \(Table[15](https://arxiv.org/html/2608.20516#A7.T15)\), so the RWSE is a deterministic function of the census too: it varies across*aspect types*, and across papers only through the claim count, and it carries no paper\-identifying information beyond that\. Non\-zero variance is necessary but not sufficient for informativeness, and the effective\-rank measurement makes the point quantitatively—per\-aspect RWSE effective rank is1\.091\.09–1\.281\.28of1616, i\.e\. the1616\-dimensional encoding spans barely more than one direction\. A structural encoding on a census\-determined graph is close to a constant by construction\. The check that would have caught this is the cardinality comparison of Table[15](https://arxiv.org/html/2608.20516#A7.T15), not a variance statistic\.
The cue ablation of App\.[A14](https://arxiv.org/html/2608.20516#A14)closes the loop empirically: a*three\-entry*learned aspect embedding outperforms the1616\-dimensional RWSE by\+0\.167\+0\.167bits \(14\.37614\.376versus14\.20914\.209\), which is what one expects if the RWSE is carrying little beyond aspect identity\. Removing the cue entirely costs\+0\.680\+0\.680bits \(13\.52913\.529\), so the predictor does need*some*target designator—it simply does not need a structural one\. Whether the*trained*query path preserves what little the RWSE does carry is a separate question and is the first of the open measurements in Sec\.[10](https://arxiv.org/html/2608.20516#S10)\.
## Appendix A32Supporting figures
Figures[8](https://arxiv.org/html/2608.20516#A32.F8)–[13](https://arxiv.org/html/2608.20516#A32.F13)support Secs\.[5](https://arxiv.org/html/2608.20516#S5)through[9](https://arxiv.org/html/2608.20516#S9): the embedding geometry, the depth sweep, the training dynamics that bear on Sec\.[7\.1](https://arxiv.org/html/2608.20516#S7.SS1), the difficulty ladder, the DC\-ratio comparison and the post\-hoc frame sweep\.
Figure 8:Trained patches form one point\-mass per aspect\.First two principal components\. Trained representations \(left, centre\) occupy two or three tight clusters corresponding to aspect type; the frozen features \(right\) form a diffuse cloud in which individual papers are separable\. The representation encodes*which aspect*, not*which paper*—the visual form of Table[3](https://arxiv.org/html/2608.20516#S7.T3)and of theerank≤\|𝒜\|\\operatorname\{erank\}\\leq\|\\mathcal\{A\}\|bound of Prop\.[1](https://arxiv.org/html/2608.20516#Thmproposition1)\.Figure 9:Depth is not the binding factor\.\(a\)Bits recovered across six encoder configurations from a depth\-00linear map to a depth\-33GNN; all at chance, the14\.37914\.379\-bit ceiling dashed\.\(b\)Message passing compresses node rank monotonically \(54\.9→18\.054\.9\\to 18\.0\) and the pooled rank is flat regardless\.Figure 10:Nothing is learned away\.\(a\)Training loss converges within∼15\{\\sim\}15epochs, on a twin axis because the cosine and InfoNCE losses differ by an order of magnitude—and note that this is exactly the comparison Sec\.[9](https://arxiv.org/html/2608.20516#S9)warns against making on magnitude alone, since a regression loss near zero is collapse while an InfoNCE loss near zero is discrimination\.\(b\)Pooled effective rank is already at its final value at epoch00and stays flat: the baseline pathology is present at initialisation rather than induced by optimisation, which is why Sec\.[7\.1](https://arxiv.org/html/2608.20516#S7.SS1)measures the variance allocation per checkpoint rather than asserting a re\-allocation\. Contrast the repaired configuration, whose bits rise from−0\.07\-0\.07at step11to\+14\.24\+14\.24by step500500\(App\.[A14](https://arxiv.org/html/2608.20516#A14)\); the difference between the two trajectories is the*objective*, as the single\-variable regression control establishes\.Figure 11:Bits recovered against pool difficultyfor the three ceilings and the lexical no\-overlap control \(Table[2](https://arxiv.org/html/2608.20516#S6.T2)\)\. All four series are ceilings or controls—bm25,bm25\-no\-ov,oracleand thep1positive control—and all four track the1Nlog2N\!\\frac\{1\}\{N\}\\log\_\{2\}N\!envelope with a near\-constant deficit across a5,800×5\{,\}800\\timeschange in pool size\. The baseline Graph\-JEPA is the single point at00bits at the right\-hand end\. This figure replaces every ratio\-to\-chance statement in earlier drafts\.Figure 12:Shared\-mean dominance by pipeline stage\.DC ratio‖μ‖/𝔼‖x−μ‖\\\|\\mu\\\|/\\mathbb\{E\}\\\|x\-\\mu\\\|on a log axis\. The frozen features sit near10−210^\{\-2\}; the trained query bank sits at162\.2162\.2—a factor of∼104\{\\sim\}10^\{4\}\.Figure 13:A constant query cannot be rescued by any frame applied after training\.Five retrieval frames fitted on the candidate bank; the best buys0\.4bits0\.4\\text\{ bits\}against a14\.4bits14\.4\\text\{ bits\}deficit to the ceiling \(dashed\)\. This is a quantitative confirmation of Prop\.[3](https://arxiv.org/html/2608.20516#Thmproposition3)rather than a failed ablation: the way out is to change what the objective asks for, and the single\-variable demonstration of that is the loss ablation of Sec\.[9](https://arxiv.org/html/2608.20516#S9)\(→0\.30714\.359\\\!\\to\\\!0\.307bits\), not any transformation of the retrieval space\.
## Appendix A33Compute environment and budget
All experiments ran on a single node with2×2\\timesNVIDIA L40S GPUs \(4646GB each\) or one NVIDIA H100 NVL \(9595GB\) for the2020k\-step runs, PyTorch2\.4\.0\+2\.4\.0\{\+\}cu121, PyTorch Geometric2\.8\.02\.8\.0, scikit\-learn1\.7\.21\.7\.2, Python3\.103\.10\. Graph construction and RWSE precomputation are cached and reused across every block; blocks A and C share that cache byte\-for\-byte \(App\.[A5](https://arxiv.org/html/2608.20516#A5)\)\.
Table 34:Measured compute budget\.The Protocol R diagnosis costs0\.770\.77GPU\-hours; the additional diagnostic harness of Secs\.[6](https://arxiv.org/html/2608.20516#S6)–[7](https://arxiv.org/html/2608.20516#S7)costs4\.74\.7minutes on the same node\. The low cost is deliberate: every check recommended in Sec\.[10](https://arxiv.org/html/2608.20516#S10)is cheap enough to run*before*committing to a training budget—and the two checks that changed this paper’s conclusions, the reducibility audit and the duplicate census, are the two cheapest lines in the table\.StageWall\-clockPeak GPUGPU\-hoursPreprocess: graph build\+\+sentence embed \(one\-time\)∼7\{\\sim\}7min†——Preprocess: RWSE cache \(one\-time\)∼20\{\\sim\}20s——Instrument self\-test \(2020assertions\)1818s1\.21\.2GB<0\.01<0\.01Untrained rank / DC vs\. degree11s4\.04\.0GB<0\.01<0\.01Protocol R, raw\+\+whitened \(10 runs\)66m 45 s18\.818\.8GB0\.110\.11Per\-aspect breakdown \(5 seeds\)33m 16 s18\.118\.1GB0\.050\.05Retrieval\-frame sweep \(3 seeds\)11m 57 s18\.018\.0GB0\.030\.03Encoder sweep \(6 configs×\\times3 seeds\)44m 56 s22\.622\.6GB0\.080\.08Oracle\+\+bm25\+\+leakage controls66m 59 s3\.23\.2GB0\.120\.12Pooling×\\timesloss grid \(7 cells×\\times5 seeds\)2121m 32 s18\.918\.9GB0\.360\.36Difficulty ladder\+\+hard pools \(BM25 index reused\)4646s6\.16\.1GB0\.010\.01Extraction / templating audit7575s2\.02\.0GB0\.020\.02Phase analysis \(6060simulations\+\+bisection\)33s2\.42\.4GB<0\.01<0\.01Bits\-accounting verification \(CPU only\)99s—00Checkpoint allocation pass\(App\.[A15](https://arxiv.org/html/2608.20516#A15)\)<𝟏\\mathbf\{<1\}min2\.02\.0GB<0\.01\\mathbf\{<0\.01\}Reducibility audit\(edge\-cardinality pass\)<𝟏\\mathbf\{<1\}s—𝟎\\mathbf\{0\}Duplicate\+\+genericness census𝟒𝟏\\mathbf\{41\}s2\.12\.1GB0\.01\\mathbf\{0\.01\}Polarity probe, both groupings \(3 seeds×\\times2\)22m 29 s9\.49\.4GB0\.040\.04Matched\-budget2×22\\times 2,33k steps \(1 seed\)3232m 08 s31\.431\.4GB0\.540\.54Loss ablation, regression control,33k \(1 seed\)77m 51 s30\.830\.8GB0\.130\.13Cue ablation,3×3\\times66k steps \(1 seed\)4747m 12 s32\.032\.0GB0\.790\.79Transduction\+\+raw\-skip,2×2\\times66k \(1 seed\)3232m 07 s32\.632\.6GB0\.540\.54Main run,2020k steps \(3 seeds\)22h 37 m34\.134\.1GB2\.622\.62Total, diagnosis only𝟒𝟔\\mathbf\{46\}min22\.6\\mathbf\{22\.6\}GB0\.77\\mathbf\{0\.77\}Total, including the repair sweep6\.1\\mathbf\{6\.1\}h34\.1\\mathbf\{34\.1\}GB5\.44\\mathbf\{5\.44\}
†One\-time; dominated by sentence encoding of∼1\.05\{\\sim\}1\.05M nodes\.
##### The cost asymmetry is the practical message\.
The two measurements that changed this paper’s conclusions cost under a minute combined: the reducibility audit is a pass over nine integers, and the duplicate and genericness census is a single pass over an embedding matrix\. The sweep they retrospectively reinterpreted cost4\.64\.6GPU\-hours, four orders of magnitude more\. We had the budget to train for2020k steps on three seeds long before we had the discipline to compare an edge count against a node count\. Any reader who takes one thing from this paper should take that ordering: audit the target before you train on it, because the audit is free and the training is not\. The checkpoint allocation pass of App\.[A15](https://arxiv.org/html/2608.20516#A15)belongs in the same category, and so does the non\-census\-determined control graph named in Sec\.[10](https://arxiv.org/html/2608.20516#S10): about one GPU\-hour, against a claim that the rest of the paper cannot make without it\.
##### And one measurement that cost nothing but attention\.
The correction recorded in App\.[A14](https://arxiv.org/html/2608.20516#A14)—that the target\-frame main effect is\+0\.071\+0\.071bits at matched budget rather than the\+1\.199\+1\.199an earlier draft reported, and that the learning\-rate schedule is worth\+1\.337\+1\.337bits—came from reading a log stage we had already paid for and previously skipped\. No new compute was required\. We mention it because the marginal value of re\-reading a completed run is easy to underestimate relative to launching another one\.Similar Articles
NodeJEPA: Structure-Conditioned Latent Prediction for Node-Level Graph Self-Supervised Learning
This paper introduces NodeJEPA, a joint-embedding predictive architecture for node-level graph self-supervised learning that predicts latent representations of masked structure-aware ego-subgraphs, avoiding reconstruction and hand-crafted augmentations. The method is evaluated on node classification benchmarks and shows competitive performance.
Diagnosing JEPA World Models with Action-Conditioned Predictive Consistency
Introduces Action-Conditioned Predictive Consistency (ACPC), a diagnostic for JEPA world models that measures how clean and perturbed observations diverge under action-conditioned rollouts, with theoretical bounds on prediction error and planner cost. Experiments on visual control tasks validate the diagnostic across models like LeWM and PLDM.
HP-JEPA: Hierarchical Partitioning for Multi-Resolution Graph Joint-Embedding Predictive Learning
This paper introduces HP-JEPA, a hierarchical partitioning framework for multi-resolution graph joint-embedding predictive learning, which outperforms the fixed-resolution Graph-JEPA baseline on most graph classification and regression benchmarks.
SJEPA: Learning Elegant Latent Dynamics with Hybrid Symbolic-Neural Predictors
SJEPA introduces a reconstruction-free JEPA framework that learns hybrid symbolic-neural latent dynamics, aiming for the simplest adequate predictive representation. Experiments show it discovers simpler symbolic dynamics with lower rollout error than post-hoc fitting, while controlling symbolic-neural allocation under grammar misspecification.
Flow-JEPA: Flow Matching for Robust Latent Dynamics in JEPA World Models
Flow-JEPA introduces a conditional flow matching approach to JEPA world models, improving robustness and accuracy in predicting future latent states under noisy conditions.