Phases in a class of associative memories via hidden neurons

arXiv cs.LG Papers

Summary

The paper analyzes a class of associative memories with hidden neurons, deriving phase diagrams and storage capacities using replica methods and linking to softmax attention in transformers.

arXiv:2609.10976v1 Announce Type: new Abstract: Associative memory in the Hopfield network is attractor dynamics in a disordered many-body system, and higher-order and exponential extensions turn its retrieval update into softmax attention. The polynomial and exponential regimes have been analyzed by different methods, with no common architecture in which to ask what fixes the storage scale. In this paper we study the bipartite architecture of Krotov and Hopfield, which we call the class $H$, whose model is fixed by a Lagrangian for each layer, taking the hidden neurons as the order parameter of retrieval. At polynomial load the replica method yields the replica-symmetric phase diagrams and closed-form capacities, and the crosstalk moment is common to Ising and spherical visible neurons, so their differences come from the visible entropy. With a softmax hidden layer the load is exponential, and a copy representation maps the thermodynamics onto random-energy-model counting, with paramagnetic, condensed, and frozen phases. Heating destabilizes retrieval by quantized reassignments of attention, and typical Gaussian patterns remain metastable at every load. The regimes differ in their crosstalk statistics, central-limit at polynomial load and large-deviation at exponential load, and the class $H$ splits retrieval into two roles, the visible Lagrangian fixing stability and the hidden one the storage scale, two axes that may also guide the design of new Lagrangians.
Original Article
View Cached Full Text

Cached at: 09/11/26, 08:24 AM

# Phases in a class of associative memories via hidden neurons
Source: [https://arxiv.org/html/2609.10976](https://arxiv.org/html/2609.10976)
Toshihiro OtaEmail:[ota\_toshihiro@cyberagent\.co\.jp](mailto:[email protected])Affiliation:CyberAgent AI Lab, Shibuya, Tokyo 150–0002, JapanAffiliation:RIKEN iTHEMS, Wako, Saitama 351–0198, JapanMasato TakiEmail:[taki\_m@rikkyo\.ac\.jp](mailto:[email protected])Affiliation:Graduate School of Artificial Intelligence and Science, Rikkyo University, Toshima, Tokyo 171–8501, JapanAffiliation:RIKEN iTHEMS, Wako, Saitama 351–0198, Japan

###### Abstract

Associative memory in the Hopfield network is attractor dynamics in a disordered many\-body system, and higher\-order and exponential extensions turn its retrieval update into softmax attention\. The polynomial and exponential regimes have been analyzed by different methods, with no common architecture in which to ask what fixes the storage scale\. In this paper we study the bipartite architecture of Krotov and Hopfield, which we call the classℋ\\mathcal\{H\}, whose model is fixed by a Lagrangian for each layer, taking the hidden neurons as the order parameter of retrieval\. At polynomial load the replica method yields the replica\-symmetric phase diagrams and closed\-form capacities, and the crosstalk moment is common to Ising and spherical visible neurons, so their differences come from the visible entropy\. With a softmax hidden layer the load is exponential, and a copy representation maps the thermodynamics onto random\-energy\-model counting, with paramagnetic, condensed, and frozen phases\. Heating destabilizes retrieval by quantized reassignments of attention, and typical Gaussian patterns remain metastable at every load\. The regimes differ in their crosstalk statistics, central\-limit at polynomial load and large\-deviation at exponential load, and the classℋ\\mathcal\{H\}splits retrieval into two roles, the visible Lagrangian fixing stability and the hidden one the storage scale, two axes that may also guide the design of new Lagrangians\.

## IIntroduction

Associative memory is a content\-addressable mechanism that retrieves a whole memory from a partial cue\. The Hopfield network formalized this retrieval as the collective dynamics of many interacting neurons, with the stored patterns realized as attractors of an energy landscape\[[1](https://arxiv.org/html/2609.10976#bib.bib1)\]\. The network is thereby a disordered many\-body system, and the replica method of spin\-glass statistical mechanics quantified the competition between the retrieval and the spin\-glass phases and the storage capacity of the model\[[2](https://arxiv.org/html/2609.10976#bib.bib2),[3](https://arxiv.org/html/2609.10976#bib.bib3),[4](https://arxiv.org/html/2609.10976#bib.bib4)\]\.

This framework has since been extended considerably\[[5](https://arxiv.org/html/2609.10976#bib.bib5),[6](https://arxiv.org/html/2609.10976#bib.bib6)\]\. Replacing the pairwise interactions by higher\-order ones raises the number of retrievable patterns from linear in the number of neurons to polynomial, with a degree that grows with the order of the interaction\[[7](https://arxiv.org/html/2609.10976#bib.bib7),[8](https://arxiv.org/html/2609.10976#bib.bib8),[9](https://arxiv.org/html/2609.10976#bib.bib9),[10](https://arxiv.org/html/2609.10976#bib.bib10)\], and exponential interactions make it exponentially large in the number of neurons\[[11](https://arxiv.org/html/2609.10976#bib.bib11)\]\. With a log\-sum\-exp energy the retrieval update takes the form of softmax attention, placing associative memory in direct correspondence with the attention mechanism of the transformer\[[12](https://arxiv.org/html/2609.10976#bib.bib12),[13](https://arxiv.org/html/2609.10976#bib.bib13),[14](https://arxiv.org/html/2609.10976#bib.bib14)\]\.

The two regimes have been analyzed with different tools: the replica method at polynomial load, and large\-deviation and extreme\-value statistics at exponential load, which control both the retrieval thresholds\[[15](https://arxiv.org/html/2609.10976#bib.bib15)\]and the finite\-temperature transitions\[[16](https://arxiv.org/html/2609.10976#bib.bib16),[17](https://arxiv.org/html/2609.10976#bib.bib17),[18](https://arxiv.org/html/2609.10976#bib.bib18)\]\. What is missing is a setting in which one can ask, within a single architecture, what fixes the storage scale and how the crosstalk statistics change with it\. The bipartite architecture of Krotov and Hopfield provides one\[[19](https://arxiv.org/html/2609.10976#bib.bib19)\]: a visible and a hidden layer coupled only by pairwise interactions, in which the choice of a Lagrangian for each layer alone determines the model\. We call this two\-layer family the*classℋ\\mathcal\{H\}*, thek=2k=2member of a hierarchical classℋk\\mathcal\{H\}\_\{k\}ofkkcoupled layers, and describe it in Sec\.[II](https://arxiv.org/html/2609.10976#S2)\. Its three representatives singled out in\[[19](https://arxiv.org/html/2609.10976#bib.bib19)\], studied so far as separate models with different effective energies, are Model A, with Ising visible neurons, Model B, whose hidden layer is a softmax, and Model C, with spherical visible neurons\. We take a further step and treat the hidden neurons themselves as the order parameter of memory retrieval, which describes different visible geometries and hidden nonlinearities in one language\.

Sec\.[III](https://arxiv.org/html/2609.10976#S3)treats the statistical mechanics of Models A and C at polynomial load, where the replica method yields the replica\-symmetric phase diagrams, with the zero\-temperature capacities in closed form\. The moment that measures the crosstalk of the non\-retrieved patterns is common to the two models, while their retrieval phases differ sharply: the quadratic spherical model is marginal and stores nothing\[[20](https://arxiv.org/html/2609.10976#bib.bib20)\], whereas the higher\-order interaction restores a retrieval phase\. In Sec\.[IV](https://arxiv.org/html/2609.10976#S4), we study Model B, whose hidden sector reduces to the attention weights assigned to the stored patterns, so that retrieval is the concentration of attention\. Its natural load is exponential, and a representation of the thermodynamics in terms of a temperature\-dependent number of copies, each selecting one stored pattern, turns the problem into one of counting with the structure of the random energy model\[[21](https://arxiv.org/html/2609.10976#bib.bib21),[22](https://arxiv.org/html/2609.10976#bib.bib22)\]\. This yields the paramagnetic, condensed, and frozen phases, and shows that heating destabilizes retrieval not by a smooth erosion of the overlap but by quantized reassignments of attention, the retrieval of a typical Gaussian pattern remaining metastable at every load\.

What separates the two regimes is the character of the crosstalk statistics, central\-limit and insensitive to the pattern ensemble at polynomial load, large\-deviation and ensemble\-dependent at exponential load\. One consequence is that at exponential load the rare Gaussian patterns of atypically large norm lie below typical retrieval in free energy, so that a typical memory is never the equilibrium phase and the network operates as a metastable device\. Underlying both statements is the picture the classℋ\\mathcal\{H\}supplies, in which the retrieval problem separates into two roles carried by the two Lagrangians\. The visible Lagrangian fixes, through the visible entropy, the stability of retrieval under a given crosstalk, which is what isolates the difference between Models A and C\. The hidden Lagrangian fixes the storage scale and the character of the disorder statistics, a polynomial nonlinearity yielding polynomial load and central\-limit crosstalk, and the log\-sum\-exp yielding exponential load and large deviations\. In these terms the sequence from the classical Hopfield network through dense associative memory to attention is not a succession of separate theories, but a set of positions on these two axes, compared in the common order\-parameter language that the hidden neurons provide\.111The code for our numerical experiments is available at[https://github\.com/Toshihiro\-Ota/classh](https://github.com/Toshihiro-Ota/classh)\.

## IIPreliminaries

To fix notation, in this section we give an overview of the classℋ\\mathcal\{H\}, originally proposed in\[[19](https://arxiv.org/html/2609.10976#bib.bib19)\], and provide the statistical mechanical setup for our main discussions in the subsequent sections\. Details of the classℋ\\mathcal\{H\}and of the more general class\-ℋk\\mathcal\{H\}\_\{k\}associative memories are presented in Appendix[A](https://arxiv.org/html/2609.10976#A1)\.

### II\.1Overview of the classℋ\\mathcal\{H\}

The dynamical variables in this system consist ofNvN\_\{v\}visible neuronsv⁡\(t\)∈ℝNvv\(t\)\\in\\mathbb\{R\}^\{N\_\{v\}\}andNhN\_\{h\}hidden neuronsh⁡\(t\)∈ℝNhh\(t\)\\in\\mathbb\{R\}^\{N\_\{h\}\}, and their interactions are represented byξ\(h,v\)∈ℝNh×Nv\\xi^\{\(h,v\)\}\\in\\mathbb\{R\}^\{N\_\{h\}\\times N\_\{v\}\}andξ\(v,h\)∈ℝNv×Nh\\xi^\{\(v,h\)\}\\in\\mathbb\{R\}^\{N\_\{v\}\\times N\_\{h\}\}, with the constraintξ\(v,h\)=\(ξ\(h,v\)\)⊤\\xi^\{\(v,h\)\}=\(\\xi^\{\(h,v\)\}\)^\{\\top\}, see Fig\.[1](https://arxiv.org/html/2609.10976#S2.F1)\. The dynamics of the system is governed by the “Lagrangians”Lv:ℝNv→ℝ\{L\_\{v\}\}\\colon\\mathbb\{R\}^\{N\_\{v\}\}\\to\\mathbb\{R\}andLh:ℝNh→ℝ\{L\_\{h\}\}\\colon\\mathbb\{R\}^\{N\_\{h\}\}\\to\\mathbb\{R\}, which determine the activation functions of the neurons as their gradients,

f=∇Lh,g=∇Lv\.f=\\nabla\{L\_\{h\}\},\\qquad g=\\nabla\{L\_\{v\}\}\.\(1\)
The dynamical equations of the system and the energy function are given by

τv​d​v​\(t\)d​t=λτh​ξ\(v,h\)​f​\(h⁡\(t\)\)−v⁡\(t\),\\displaystyle\{\\tau\_\{v\}\}\\frac\{dv\(t\)\}\{dt\}=\\frac\{\\lambda\}\{\{\\tau\_\{h\}\}\}\\xi^\{\(v,h\)\}f\(h\(t\)\)\-v\(t\),\(2\)τh​d​h​\(t\)d​t=λτv​ξ\(h,v\)​g​\(v⁡\(t\)\)−h⁡\(t\),\\displaystyle\{\\tau\_\{h\}\}\\frac\{dh\(t\)\}\{dt\}=\\frac\{\\lambda\}\{\{\\tau\_\{v\}\}\}\\xi^\{\(h,v\)\}g\(v\(t\)\)\-h\(t\),\(3\)and

Eξ​\(v,h\)\\displaystyle E\_\{\\xi\}\(v,h\)=1τv​\(v⊤​g​\(v\)−Lv​\(v\)\)\\displaystyle=\\frac\{1\}\{\{\\tau\_\{v\}\}\}\\left\(v^\{\\top\}g\(v\)\-\{L\_\{v\}\}\(v\)\\right\)\(4\)\+1τh​\(h⊤​f​\(h\)−Lh​\(h\)\)−λτh​τv​f​\(h\)⊤​ξ\(h,v\)​g​\(v\)\\displaystyle\+\\frac\{1\}\{\{\\tau\_\{h\}\}\}\\left\(h^\{\\top\}f\(h\)\-\{L\_\{h\}\}\(h\)\\right\)\-\\frac\{\\lambda\}\{\{\\tau\_\{h\}\}\{\\tau\_\{v\}\}\}f\(h\)^\{\\top\}\\xi^\{\(h,v\)\}g\(v\)≕1τv​Ev​\(v\)\+1τh​Eh​\(h\)\+λτh​τv​Eint​\(v,h\),\\displaystyle\\eqcolon\\frac\{1\}\{\{\\tau\_\{v\}\}\}\{E\_\{v\}\}\(v\)\+\\frac\{1\}\{\{\\tau\_\{h\}\}\}\{E\_\{h\}\}\(h\)\+\\frac\{\\lambda\}\{\{\\tau\_\{h\}\}\{\\tau\_\{v\}\}\}E\_\{\\mathrm\{int\}\}\(v,h\),whereτv\{\\tau\_\{v\}\}andτh\{\\tau\_\{h\}\}are the relaxation time constants of the visible and hidden neurons, respectively, andλ\\lambdais a coupling constant\. The combinationsv⊤​g​\(v\)−Lv​\(v\)v^\{\\top\}g\(v\)\-\{L\_\{v\}\}\(v\)andh⊤​f​\(h\)−Lh​\(h\)h^\{\\top\}f\(h\)\-\{L\_\{h\}\}\(h\)are the Legendre transforms of the Lagrangians evaluated at the activationsg⁡\(v\)g\(v\)andf⁡\(h\)f\(h\)\. In this sense,EξE\_\{\\xi\}is the “Hamiltonian” associated with the pair of Lagrangians\. In fact, the dynamical equations above can be viewed as one half of Hamilton’s canonical equations restricted to the Legendre constraint surface, which makes the dynamics dissipative, see Appendix[A](https://arxiv.org/html/2609.10976#A1)\.

![Refer to caption](https://arxiv.org/html/2609.10976v1/neurons.png)Figure 1:The class\-ℋ\\mathcal\{H\}associative memories\. TheNvN\_\{v\}visible neurons and theNhN\_\{h\}hidden neurons form a bipartite network with no intralayer connections\.Provided that the Hessians of the Lagrangians are positive \(semi\-\)definite, this energy function monotonically decreases along the solution trajectory of the dynamical equations,

d​Eξ​\(v⁡\(t\),h⁡\(t\)\)d​t≤0\.\\frac\{dE\_\{\\xi\}\(v\(t\),h\(t\)\)\}\{dt\}\\leq 0\.\(5\)If, in addition, the overall energy function is bounded from below, the trajectory is guaranteed to converge to a fixed\-point attractor state, which corresponds to one of the local minima of the energy function\. Such fixed points may be identified with the stored memories, and the convergence toward them with memory retrieval\.

In the adiabatic limit,τv≫τh\{\\tau\_\{v\}\}\\gg\{\\tau\_\{h\}\}, the hidden neurons relax much faster than the visible ones and can be adiabatically eliminated: takingτh→0\{\\tau\_\{h\}\}\\to 0in the dynamical equations, we obtain

h⁡\(t\)=λτv​ξ\(h,v\)​g​\(v⁡\(t\)\),h\(t\)=\\frac\{\\lambda\}\{\{\\tau\_\{v\}\}\}\\xi^\{\(h,v\)\}g\(v\(t\)\),\(6\)so that the hidden neurons instantaneously follow the visible configuration, and their stateshμ​\(t\)h\_\{\\mu\}\(t\)measure the overlap between the activation of the visible neurons and the patternsξμ\(h,v\)\\xi^\{\(h,v\)\}\_\{\\mu\}\. In this sense, each hidden neuron acts as a feature detector for the corresponding pattern, and the hidden neurons serve as the order parameter of memory retrieval in the system\[[19](https://arxiv.org/html/2609.10976#bib.bib19)\]\.

### II\.2Partition function

To consider the statistical mechanics of the classℋ\\mathcal\{H\}with the energy function Eq\. \([4](https://arxiv.org/html/2609.10976#S2.E4)\), here we introduce a formal partition function in a general setup:

Zξ​\(β\)=∫d​v​𝑑h​exp⁡\(−β​Eξ​\(v,h\)\),Z\_\{\\xi\}\(\\beta\)=\\int dvdh\\exp\\left\(\-\\beta E\_\{\\xi\}\(v,h\)\\right\),\(7\)whereβ\\betais the inverse temperature\. This expression is formal: for certain choices of Lagrangians, the energy function has flat directions along which the integral diverges \(e\.g\., forLv\{L\_\{v\}\}homogeneous of degree one,Ev\{E\_\{v\}\}vanishes identically and the energy is independent of the overall scale ofvv\), and thus the precise partition function, including the integration domain and measure, is defined for each model in the corresponding sections below\.

We can write this partition function as

Zξ​\(β\)=∫d​v​e−βτv​Ev​\(v\)×∫d​h​exp⁡\{−βτh​\(Eh​\(h\)\+λτv​Eint​\(v,h\)\)\}\.Z\_\{\\xi\}\(\\beta\)=\\int dv\\,e^\{\-\\frac\{\\beta\}\{\{\\tau\_\{v\}\}\}\{E\_\{v\}\}\(v\)\}\\\\ \\times\\int dh\\exp\\left\\\{\-\\frac\{\\beta\}\{\{\\tau\_\{h\}\}\}\\left\(\{E\_\{h\}\}\(h\)\+\\frac\{\\lambda\}\{\{\\tau\_\{v\}\}\}E\_\{\\mathrm\{int\}\}\(v,h\)\\right\)\\right\\\}\.\(8\)In this form, the adiabatic limit is rephrased asβ/τh→∞\\beta/\{\\tau\_\{h\}\}\\to\\infty, in which the thermal fluctuations of the hidden neurons are suppressed\. By the saddle\-point approximation, thehhintegral then localizes at the stationary point of the integrand, which is given by

h∗=λτv​ξ\(h,v\)​g​\(v\)\.h\_\{\\ast\}=\\frac\{\\lambda\}\{\{\\tau\_\{v\}\}\}\\xi^\{\(h,v\)\}g\(v\)\.\(9\)Thus, in the adiabatic limit, the partition function becomes

Zξ​\(β\)\\displaystyle Z\_\{\\xi\}\(\\beta\)\(10\)≈∫d​v​exp⁡\{−β⁡\(1τv​Ev​\(v\)−1τh​Lh​\(h∗\)\)\}​\(2​πβ/τh\)Nh/2\\displaystyle\\approx\\int dv\\exp\\left\\\{\-\\beta\\left\(\\frac\{1\}\{\{\\tau\_\{v\}\}\}\{E\_\{v\}\}\(v\)\-\\frac\{1\}\{\{\\tau\_\{h\}\}\}\{L\_\{h\}\}\(h\_\{\\ast\}\)\\right\)\\right\\\}\\left\(\\frac\{2\\pi\}\{\\beta/\{\\tau\_\{h\}\}\}\\right\)^\{N\_\{h\}/2\}=\(2​πβ/τh\)Nh/2​∫d​v​exp⁡\(−β​Eξ​\(v,h∗\)\)\.\\displaystyle=\\left\(\\frac\{2\\pi\}\{\\beta/\{\\tau\_\{h\}\}\}\\right\)^\{N\_\{h\}/2\}\\int dv\\exp\\left\(\-\\beta E\_\{\\xi\}\(v,h\_\{\\ast\}\)\\right\)\.Strictly, the Gaussian prefactor also carries the factor\(detHessLh\(h∗\)\)−1/2\(\\det\\operatorname\{Hess\}\{L\_\{h\}\}\(h\_\{\\ast\}\)\)^\{\-1/2\}, which we suppress at this leading order\. Its role is examined for Model B in Appendix[C\.2](https://arxiv.org/html/2609.10976#A3.SS2)\.

As discussed in this section, the hidden neurons play the role of the order parameter of memory retrieval in the system\. In the following sections, we investigate the properties of Models A, B, and C from the viewpoint of the role of the hidden neurons\.

## IIIModels A and C

Models A and C are members of a family of models within the classℋ\\mathcal\{H\}\. Let us consider a family of Lagrangians,222These models can be further generalized, see Appendix[B\.1](https://arxiv.org/html/2609.10976#A2.SS1)\. For the purpose of the present paper, it suffices to consider theℓp\\ell\_\{p\}\-norm\.

Lv​\(v\)=‖v‖p,Lh​\(h\)=∑μF⁡\(hμ\),\{L\_\{v\}\}\(v\)=\\left\\lVert v\\right\\rVert\_\{p\},\\qquad\{L\_\{h\}\}\(h\)=\\sum\_\{\\mu\}F\(h\_\{\\mu\}\),\(11\)where1≤p≤∞1\\leq p\\leq\\infty\. The corresponding activation functions are

gi​\(v\)=sgn⁡\(vi\)​\|vi\|p−1‖v‖pp−1,fμ​\(h\)=F′​\(hμ\)\.\\displaystyle g\_\{i\}\(v\)=\\frac\{\\sgn\(v\_\{i\}\)\\left\\lvert v\_\{i\}\\right\\rvert^\{p\-1\}\}\{\\left\\lVert v\\right\\rVert\_\{p\}^\{p\-1\}\},\\qquad f\_\{\\mu\}\(h\)=F^\{\\prime\}\(h\_\{\\mu\}\)\.\(12\)Sincev⊤​g​\(v\)−Lv=0v^\{\\top\}g\(v\)\-\{L\_\{v\}\}=0for these Lagrangians, the bare visible neurons vanish in the energy function, that is, the energy depends onvvonly through the activationgg:

Eξ​\(v,h\)\\displaystyle E\_\{\\xi\}\(v,h\)=1τh​∑μ\(hμ​F′​\(hμ\)−F⁡\(hμ\)\)\\displaystyle=\\frac\{1\}\{\{\\tau\_\{h\}\}\}\\sum\_\{\\mu\}\\left\(h\_\{\\mu\}F^\{\\prime\}\(h\_\{\\mu\}\)\-F\(h\_\{\\mu\}\)\\right\)\(13\)−λτh​τv∑μ,iF′\(hμ\)ξ\(h,v\)μ​igi\(v\)\.\\displaystyle\-\\frac\{\\lambda\}\{\{\\tau\_\{h\}\}\{\\tau\_\{v\}\}\}\\sum\_\{\\mu,i\}F^\{\\prime\}\(h\_\{\\mu\}\)\\xi^\{\(h,v\)\}\_\{\\mu i\}g\_\{i\}\(v\)\.
The activation of the visible neurons satisfies‖g⁡\(v\)‖p′=1\\left\\lVert g\(v\)\\right\\rVert\_\{p^\{\\prime\}\}=1, wherep′p^\{\\prime\}is the Hölder conjugate ofppdefined by1/p\+1/p′=11/p\+1/p^\{\\prime\}=1\. Since the energy is independent of the overall scale ofvv\(cf\. Sec\.[II\.2](https://arxiv.org/html/2609.10976#S2.SS2)\), the partition function for this family should be defined with the visible integral restricted to a sphere:

Zξ\(β\)=∫ℝNhdhe−βτh∑μ\(hμF′\(hμ\)−F\(hμ\)\)×∫Sd​Ω​\(x\)​exp⁡\{β​λτh​τv​∑μ,iF′​\(hμ\)​ξμ​i\(h,v\)​xi\},Z\_\{\\xi\}\(\\beta\)=\\int\_\{\\mathbb\{R\}^\{N\_\{h\}\}\}dh\\,e^\{\-\\frac\{\\beta\}\{\{\\tau\_\{h\}\}\}\\sum\_\{\\mu\}\\left\(h\_\{\\mu\}F^\{\\prime\}\(h\_\{\\mu\}\)\-F\(h\_\{\\mu\}\)\\right\)\}\\\\ \\times\\int\_\{S\}d\\Omega\(\\mathrm\{x\}\)\\exp\\left\\\{\\frac\{\\beta\\lambda\}\{\{\\tau\_\{h\}\}\{\\tau\_\{v\}\}\}\\sum\_\{\\mu,i\}F^\{\\prime\}\(h\_\{\\mu\}\)\\xi^\{\(h,v\)\}\_\{\\mu i\}\\mathrm\{x\}\_\{i\}\\right\\\},\(14\)where

S=𝕊p′Nv−1​\(Nv1/p′\)≔\{x∈ℝNv\|‖x‖p′=Nv1/p′\},S=\\mathbb\{S\}\_\{p^\{\\prime\}\}^\{N\_\{v\}\-1\}\(N\_\{v\}^\{1/p^\{\\prime\}\}\)\\coloneq\\left\\\{\\mathrm\{x\}\\in\\mathbb\{R\}^\{N\_\{v\}\}\\,\\middle\|\\,\\left\\lVert\\mathrm\{x\}\\right\\rVert\_\{p^\{\\prime\}\}=N\_\{v\}^\{1/p^\{\\prime\}\}\\right\\\},\(15\)andd​Ωd\\Omegais the standard measure on this sphere for1<p<∞1<p<\\infty, while forp=1,∞p=1,\\inftythe integral reduces to a discrete sum\. The radius ofSSis chosen such that the components are normalized asxi=O⁡\(1\)\\mathrm\{x\}\_\{i\}=O\(1\): in particular,SSis the sphere of radiusNv\\sqrt\{N\_\{v\}\}forp=2p=2, and the integral reduces to the sum overx∈\{±1\}Nv\\mathrm\{x\}\\in\\left\\\{\\pm 1\\right\\\}^\{N\_\{v\}\}forp=1p=1\.

Models A and C correspond to the casesp=1p=1andp=2p=2, respectively\. In what follows, we study this family with the hidden Lagrangian specified byF⁡\(x\)=xk/kF\(x\)=x^\{k\}/kwith a positive even integerkk\.333The exponentkkhere is unrelated to the indexkkof the classℋk\\mathcal\{H\}\_\{k\}\. Several further notational conflicts occur in what follows, but the meaning should always be clear from the context\.The replica method provides a powerful tool for this purpose\. Since the hidden neurons represent the order parameters of memory retrieval in these systems, we first integrate out the visible neurons and then perform the quenched average over the random patterns\. We focus on the replica symmetric \(RS\) solutions of these models\.

### III\.1Model A

The model\-A energy function is given by

EξA​\(v,h\)=Nv​k−1τh​k​∑μhμk−λτh​τv​∑μ,ihμk−1​ξμ​i\(h,v\)​sgn⁡\(vi\),\{E\_\{\\xi\}^\{\\mathrm\{A\}\}\}\(v,h\)=N\_\{v\}\\frac\{k\-1\}\{\{\\tau\_\{h\}\}k\}\\sum\_\{\\mu\}h\_\{\\mu\}^\{k\}\-\\frac\{\\lambda\}\{\{\\tau\_\{h\}\}\{\\tau\_\{v\}\}\}\\sum\_\{\\mu,i\}h\_\{\\mu\}^\{k\-1\}\\xi^\{\(h,v\)\}\_\{\\mu i\}\\sgn\(v\_\{i\}\),\(16\)where the factorNvN\_\{v\}in the first term is introduced for the extensivity of the energy function\. This modification is equivalent to the renormalizationsτh→τh/Nv\{\\tau\_\{h\}\}\\to\{\\tau\_\{h\}\}/N\_\{v\}andλ→λ/Nv\\lambda\\to\\lambda/N\_\{v\}, which fix the correct normalization ofhhand do not affect any other property of Model A as an associative memory\. From Eq\. \([14](https://arxiv.org/html/2609.10976#S3.E14)\), the partition function for this model reads

ZξA\(β\)=∫dhexp\{−Nvγk∑μmμk\+∑i=1Nvlog\(2cosh\[βk∑μξ\(v,h\)i​μmμk−1\]\)\},\{Z\_\{\\xi\}^\{\\mathrm\{A\}\}\}\(\\beta\)=\\int dh\\exp\\bigg\\\{\-N\_\{v\}\\gamma\_\{k\}\\sum\_\{\\mu\}m\_\{\\mu\}^\{k\}\\\\ \+\\sum\_\{i=1\}^\{N\_\{v\}\}\\log\\bigg\(2\\cosh\\bigg\[\\beta\_\{k\}\\sum\_\{\\mu\}\\xi^\{\(v,h\)\}\_\{i\\mu\}m\_\{\\mu\}^\{k\-1\}\\bigg\]\\bigg\)\\bigg\\\},\(17\)where

γk=k−1k​βk,βk=βτh​\(λτv\)k,mμ=hμλ/τv,\\gamma\_\{k\}=\\frac\{k\-1\}\{k\}\\beta\_\{k\},\\quad\\beta\_\{k\}=\\frac\{\\beta\}\{\{\\tau\_\{h\}\}\}\\left\(\\frac\{\\lambda\}\{\{\\tau\_\{v\}\}\}\\right\)^\{k\},\\quad m\_\{\\mu\}=\\frac\{h\_\{\\mu\}\}\{\\lambda/\{\\tau\_\{v\}\}\},\(18\)and we dropped an irrelevant overall constant\. Note that the stationarity of Eq\. \([17](https://arxiv.org/html/2609.10976#S3.E17)\) with respect tomμm\_\{\\mu\}constrainsmμm\_\{\\mu\}to be the thermal average of the overlap between the patternξμ\(v,h\)\\xi^\{\(v,h\)\}\_\{\\mu\}and the visible spinssgn⁡\(vi\)\\sgn\(v\_\{i\}\), in accordance with the general role of the hidden neurons discussed in Sec\.[II\.1](https://arxiv.org/html/2609.10976#S2.SS1)\.

We compute the quenched free energy per visible neuron,

fA\(β\)=−limNv→∞1β​Nv𝔼ξ\[logZξA\(β\)\],f^\{\\mathrm\{A\}\}\(\\beta\)=\-\\lim\_\{N\_\{v\}\\to\\infty\}\\frac\{1\}\{\\beta N\_\{v\}\}\\mathbb\{E\}\_\{\\xi\}\\left\[\\log\{Z\_\{\\xi\}^\{\\mathrm\{A\}\}\}\(\\beta\)\\right\],\(19\)by the replica method, with the patterns drawn independently asξμ​i\(h,v\)=±1\\xi^\{\(h,v\)\}\_\{\\mu i\}=\\pm 1with equal probability\. We work in the high\-load regime of dense associative memories\[[7](https://arxiv.org/html/2609.10976#bib.bib7),[8](https://arxiv.org/html/2609.10976#bib.bib8),[9](https://arxiv.org/html/2609.10976#bib.bib9)\],

Nh=αk​Nvk−1,N\_\{h\}=\\alpha\_\{k\}N\_\{v\}^\{k\-1\},\(20\)and consider retrieval states in which a single pattern is condensed,m1=m=O⁡\(1\)m\_\{1\}=m=O\(1\), while the remaining overlaps are of orderNv−1/2N\_\{v\}^\{\-1/2\}\. Fork\>2k\>2, the analysis is performed in the adiabatic limitβ/τh→∞\\beta/\{\\tau\_\{h\}\}\\to\\inftyat fixedβk\\beta\_\{k\}, in which the hidden neurons are enslaved to the visible configuration\. The details of the replica computation are presented in Appendix[B\.2](https://arxiv.org/html/2609.10976#A2.SS2)\.

Under the RS ansatz, we eventually obtain the free energy, up to an additive constant independent of the order parameters,

β​fA=γk​mk\+αk​βk22​r​\(1−q\)\+αk2​Ψk​\(q\)−∫Dzlog2coshβk\(mk−1\+αk​rz\),\\beta f^\{\\mathrm\{A\}\}=\\gamma\_\{k\}m^\{k\}\+\\frac\{\\alpha\_\{k\}\\beta\_\{k\}^\{2\}\}\{2\}r\(1\-q\)\+\\frac\{\\alpha\_\{k\}\}\{2\}\\Psi\_\{k\}\(q\)\\\\ \-\\int Dz\\log 2\\cosh\\beta\_\{k\}\\left\(m^\{k\-1\}\+\\sqrt\{\\alpha\_\{k\}r\}\\,z\\right\),\(21\)whereDz≔dze−z2/2/2​πDz\\coloneq dz\\,e^\{\-z^\{2\}/2\}/\\sqrt\{2\\pi\}is the standard Gaussian measure,qqis the Edwards–Anderson order parameter of the visible spins,rrmeasures the strength of the crosstalk noise from the non\-condensed patterns, and the noise entropic term reads

Ψk​\(q\)=\{log⁡\[1−βk​\(1−q\)\]−βk​q1−βk​\(1−q\),k=2,βk2​∫0qℳk​\(s\)​ds,k\>2,\\Psi\_\{k\}\(q\)=\\begin\{cases\}\\log\\left\[1\-\\beta\_\{k\}\(1\-q\)\\right\]\-\\dfrac\{\\beta\_\{k\}q\}\{1\-\\beta\_\{k\}\(1\-q\)\},&k=2,\\\\\[5\.69054pt\] \\beta\_\{k\}^\{2\}\\displaystyle\\int\_\{0\}^\{q\}\\mathcal\{M\}\_\{k\}\(s\)\\,ds,&k\>2,\\end\{cases\}\(22\)withℳk\\mathcal\{M\}\_\{k\}defined in Eq\. \([26](https://arxiv.org/html/2609.10976#S3.E26)\) below\. The equations of state for the order parameters follow from the stationarity of the free energy\. We then obtain

m\\displaystyle m=∫Dztanhβk\(mk−1\+αk​rz\),\\displaystyle=\\int Dz\\tanh\\beta\_\{k\}\\left\(m^\{k\-1\}\+\\sqrt\{\\alpha\_\{k\}r\}\\,z\\right\),\(23\)q\\displaystyle q=∫D​z​tanh2⁡βk​\(mk−1\+αk​r​z\),\\displaystyle=\\int Dz\\tanh^\{2\}\\beta\_\{k\}\\left\(m^\{k\-1\}\+\\sqrt\{\\alpha\_\{k\}r\}\\,z\\right\),\(24\)r\\displaystyle r=ℳk​\(q\),\\displaystyle=\\mathcal\{M\}\_\{k\}\(q\),\(25\)where

ℳk​\(q\)=\{q\(1−βk​\(1−q\)\)2,k=2,𝔼⁡\[Xk−1​Yk−1\],k\>2,\\mathcal\{M\}\_\{k\}\(q\)=\\begin\{cases\}\\dfrac\{q\}\{\\left\(1\-\\beta\_\{k\}\(1\-q\)\\right\)^\{2\}\},&k=2,\\\\\[11\.38109pt\] \\mathbb\{E\}\\left\[X^\{k\-1\}Y^\{k\-1\}\\right\],&k\>2,\\end\{cases\}\(26\)and\(X,Y\)\(X,Y\)in the second line denotes a pair of standard Gaussian variables with correlation𝔼⁡\[X​Y\]=q\\mathbb\{E\}\\left\[XY\\right\]=q\. Fork\>2k\>2,ℳk​\(q\)\\mathcal\{M\}\_\{k\}\(q\)is a polynomial inqqwith positive coefficients, e\.g\.,ℳ4​\(q\)=9​q\+6​q3\\mathcal\{M\}\_\{4\}\(q\)=9q\+6q^\{3\}, andℳk​\(1\)=\(2​k−3\)\!\!\\mathcal\{M\}\_\{k\}\(1\)=\(2k\-3\)\!\!\.

The two cases in Eq\. \([26](https://arxiv.org/html/2609.10976#S3.E26)\) reflect the fate of the Onsager reaction field\. Relative to its bare central\-limit scale, the feedback of a non\-condensed mode onto itself is of orderβk\(1−q\)Nv−\(k−2\)/2\\beta\_\{k\}\(1\-q\)N\_\{v\}^\{\-\(k\-2\)/2\}\. Fork=2k=2this feedback is marginal and must be resummed to all orders, which produces the denominator1−βk​\(1−q\)1\-\\beta\_\{k\}\(1\-q\)of the classical result by Amit, Gutfreund, and Sompolinsky \(AGS\)\[[2](https://arxiv.org/html/2609.10976#bib.bib2),[3](https://arxiv.org/html/2609.10976#bib.bib3)\]\. Fork\>2k\>2it vanishes in the thermodynamic limit, so that the non\-condensed overlaps behave as bare Gaussian variables\. A cavity derivation of this dichotomy, together with the closed polynomial form ofℳk\\mathcal\{M\}\_\{k\}, is given in Appendix[B\.2\.5](https://arxiv.org/html/2609.10976#A2.SS2.SSS5)\.

Zero\-temperature limit and capacity condition\.Sendingβk→∞\\beta\_\{k\}\\to\\infty, the overlap satisfiesq→1q\\to 1while the combinationC≔βk​\(1−q\)C\\coloneq\\beta\_\{k\}\(1\-q\)remains finite, and the equations of state close in the pair\(m,C\)\(m,C\)\. The elementary evaluation is given in Appendix[B\.2\.5](https://arxiv.org/html/2609.10976#A2.SS2.SSS5)\. The retrieval overlap obeys

m=\{erf\}⁡\(mk−12​αk​r\),m=\\erf\\left\(\\frac\{m^\{k\-1\}\}\{\\sqrt\{2\\alpha\_\{k\}r\}\}\\right\),\(27\)the frozen response is

C=2π​αk​r​exp⁡\(−m2​k−22​αk​r\),C=\\sqrt\{\\frac\{2\}\{\\pi\\alpha\_\{k\}r\}\}\\,\\exp\\left\(\-\\frac\{m^\{2k\-2\}\}\{2\\alpha\_\{k\}r\}\\right\),\(28\)and the crosstalk moment Eq\. \([26](https://arxiv.org/html/2609.10976#S3.E26)\) becomes

r=\{\(1−C\)−2,k=2,\(2​k−3\)\!\!,k\>2,r=\\begin\{cases\}\(1\-C\)^\{\-2\},&k=2,\\\\ \(2k\-3\)\!\!,&k\>2,\\end\{cases\}\(29\)where the former reproduces the classical AGS zero\-temperature equations\[[2](https://arxiv.org/html/2609.10976#bib.bib2),[3](https://arxiv.org/html/2609.10976#bib.bib3)\], and the latter is the2​\(k−1\)2\(k\-1\)\-th moment of the standard Gaussian\.

In terms of the signal\-to\-noise ratiot≔mk−1/2​αk​rt\\coloneq m^\{k\-1\}/\\sqrt\{2\\alpha\_\{k\}r\}, any retrieval solution withm\>0m\>0lies on the parametric curve

αk​\(t\)=m​\(t\)2​\(k−1\)2​t2​r​\(t\),m⁡\(t\)=\{erf\}⁡\(t\),\\alpha\_\{k\}\(t\)=\\frac\{m\(t\)^\{2\(k\-1\)\}\}\{2t^\{2\}\\,r\(t\)\},\\qquad m\(t\)=\\erf\(t\),\(30\)wherer⁡\(t\)=\(2​k−3\)\!\!r\(t\)=\(2k\-3\)\!\!fork\>2k\>2, while fork=2k=2the reaction field enters throughr⁡\(t\)=\[1−C⁡\(t\)\]−2r\(t\)=\[1\-C\(t\)\]^\{\-2\}withC⁡\(t\)=2​t​e−t2/\[π​m​\(t\)\]C\(t\)=2te^\{\-t^\{2\}\}/\[\\sqrt\{\\pi\}\\,m\(t\)\]\. Sinceαk​\(t\)→0\\alpha\_\{k\}\(t\)\\to 0both ast→0t\\to 0and ast→∞t\\to\\infty, the curve attains an interior maximum, and retrieval solutions exist if and only if

αk≤αc≔maxt\>0⁡αk​\(t\),\\alpha\_\{k\}\\leq\\alpha\_\{c\}\\coloneq\\max\_\{t\>0\}\\alpha\_\{k\}\(t\),\(31\)where the critical loadαc\\alpha\_\{c\}marks a spinodal at which the stable and unstable retrieval branches merge\. Fork=2k=2, the maximum occurs attc≃1\.51t\_\{c\}\\simeq 1\.51, yieldingαc≃0\.138\\alpha\_\{c\}\\simeq 0\.138andmc≃0\.97m\_\{c\}\\simeq 0\.97, in agreement with the classical Hopfield result\[[2](https://arxiv.org/html/2609.10976#bib.bib2),[3](https://arxiv.org/html/2609.10976#bib.bib3)\]\. Fork\>2k\>2, the capacity takes the closed form

αc=maxt\>0⁡\[\{erf\}⁡\(t\)\]2​\(k−1\)2​t2​\(2​k−3\)\!\!,\\alpha\_\{c\}=\\max\_\{t\>0\}\\frac\{\\left\[\\erf\(t\)\\right\]^\{2\(k\-1\)\}\}\{2t^\{2\}\\,\(2k\-3\)\!\!\},\(32\)which givesαc≃1\.32×10−2\\alpha\_\{c\}\\simeq 1\.32\\times 10^\{\-2\}fork=4k=4andαc≃1\.67×10−4\\alpha\_\{c\}\\simeq 1\.67\\times 10^\{\-4\}fork=6k=6\. Note that the roles of the reaction field are opposite in the two cases: fork=2k=2, neglecting it \(C→0C\\to 0,r→1r\\to 1\) would overestimate the capacity by more than a factor of four, giving2/π≃0\.642/\\pi\\simeq 0\.64, whereas fork\>2k\>2the reaction\-free expression Eq\. \([32](https://arxiv.org/html/2609.10976#S3.E32)\) is not an approximation but the leading\-order RS result\. The double factorial here is precisely the combinatorial factor that controls the error\-free capacity estimate of the dense associative memory\[[10](https://arxiv.org/html/2609.10976#bib.bib10)\]\. For largekk, the maximum is attained attc≃ln⁡kt\_\{c\}\\simeq\\sqrt\{\\ln k\}, so that

αc∼O⁡\(1\)2​tc2​\(2​k−3\)\!\!\.\\alpha\_\{c\}\\sim\\frac\{O\(1\)\}\{2t\_\{c\}^\{2\}\\,\(2k\-3\)\!\!\}\.\(33\)The signal factormc2​\(k−1\)m\_\{c\}^\{2\(k\-1\)\}remainsO⁡\(1\)O\(1\), and the collapse of the capacity is driven by the factorial growth of the crosstalk variance\.

Critical temperature atαk=0\\alpha\_\{k\}=0and steepening of the boundary\.Atαk=0\\alpha\_\{k\}=0the crosstalk noise vanishes and the retrieval overlap satisfies

m=tanh⁡\(βk​mk−1\),m=\\tanh\\left\(\\beta\_\{k\}m^\{k\-1\}\\right\),\(34\)independently ofrr\. The corresponding critical point is determined by the spinodal condition

1=βk​\(k−1\)​mk−2​\{sech\}2⁡\(βk​mk−1\),1=\\beta\_\{k\}\(k\-1\)\\,m^\{k\-2\}\\sech^\{2\}\\left\(\\beta\_\{k\}m^\{k\-1\}\\right\),\(35\)and the critical valueTcT\_\{c\}of the effective temperatureT≔βk−1T\\coloneq\\beta\_\{k\}^\{\-1\}decreases withkkonly slowly \(see Table[1](https://arxiv.org/html/2609.10976#S3.T1)\)\. By contrast, the zero\-temperature capacity collapses factorially withkk, so the retrieval boundary connecting\(αk,T\)=\(0,Tc\)\(\\alpha\_\{k\},T\)=\(0,T\_\{c\}\)to\(αc,0\)\(\\alpha\_\{c\},0\)steepens rapidly:

\|d​Td​αk\|∼Tcαc∼\(2​k−3\)\!\!\\left\|\\frac\{dT\}\{d\\alpha\_\{k\}\}\\right\|\\sim\\frac\{T\_\{c\}\}\{\\alpha\_\{c\}\}\\sim\(2k\-3\)\!\!\(36\)up to factors varying slowly withkk, and the boundary becomes nearly vertical askkgrows\.

Physically, thekk\-body interaction sharpens the retrieval landscape through the signal termmk−1m^\{k\-1\}, while the crosstalk from the non\-condensed patterns enters through the2​\(k−1\)2\(k\-1\)\-th moment of their Gaussian overlaps\. Rare large fluctuations dominate this moment, so the cost of storing one additional pattern grows factorially withkk\. The model is thus robust against thermal noise atαk=0\\alpha\_\{k\}=0but fragile against pattern loading atαk\>0\\alpha\_\{k\}\>0, an asymmetry that becomes extreme for largekk\.

Table 1:Critical values of Model A obtained from the RS equations of state: the critical temperatureTcT\_\{c\}atαk=0\\alpha\_\{k\}=0, the zero\-temperature critical loadαc\\alpha\_\{c\}, and the resulting steepnessTc/αcT\_\{c\}/\\alpha\_\{c\}of the retrieval boundary\.![Refer to caption](https://arxiv.org/html/2609.10976v1/fig/modelA_phase_diagram_k2_k4_k6.png)Figure 2:RS phase diagram of Model A in the\(αk,T\)\(\\alpha\_\{k\},T\)plane for \(a\)k=2k=2, \(b\)k=4k=4, and \(c\)k=6k=6, whereT=1/βkT=1/\\beta\_\{k\}is the effective temperature\. In the retrieval phase \(R\) the retrieval states are the global minima of the free energy\. In the region M they persist only as metastable states\. The red solid line is the retrieval spinodalTR​\(αk\)T\_\{R\}\(\\alpha\_\{k\}\), beyond which no retrieval solution of Eq\. \([23](https://arxiv.org/html/2609.10976#S3.E23)\)–Eq\. \([25](https://arxiv.org/html/2609.10976#S3.E25)\) exists\. The black solid line is the first\-order boundaryTM​\(αk\)T\_\{M\}\(\\alpha\_\{k\}\), on which the retrieval free energy crosses that of them=0m=0branch\. The dotted line is the continuous spin\-glass transitionTg​\(αk\)T\_\{g\}\(\\alpha\_\{k\}\)of Eq\. \([37](https://arxiv.org/html/2609.10976#S3.E37)\), separating the paramagnet \(P\) from the RS spin glass \(SG\)\. Fork\>2k\>2the first\-order boundary is reentrant, extending to larger loads at intermediate temperatures than atT=0T=0\.Phase diagram\.The complete RS phase diagram in the\(αk,T\)\(\\alpha\_\{k\},T\)plane is assembled in Fig\.[2](https://arxiv.org/html/2609.10976#S3.F2)\. Besides the retrieval states, the equations of state admit anm=0m=0solution withq\>0q\>0, a spin glass sustained by the crosstalk noise alone\. Linearizing Eq\. \([24](https://arxiv.org/html/2609.10976#S3.E24)\) and Eq\. \([25](https://arxiv.org/html/2609.10976#S3.E25)\) atm=0m=0and smallqq, whereq≃βk2​αk​ℳk​\(q\)q\\simeq\\beta\_\{k\}^\{2\}\\alpha\_\{k\}\\mathcal\{M\}\_\{k\}\(q\)andℳk​\(q\)≃\[\(k−1\)\!\!\]2​q\\mathcal\{M\}\_\{k\}\(q\)\\simeq\[\(k\-1\)\!\!\]^\{2\}\\,qfork\>2k\>2, shows that this solution bifurcates continuously from the paramagnet at

Tg​\(αk\)=\{1\+αk,k=2,\(k−1\)\!\!​αk,k\>2,T\_\{g\}\(\\alpha\_\{k\}\)=\\begin\{cases\}1\+\\sqrt\{\\alpha\_\{k\}\},&k=2,\\\\\[2\.84526pt\] \(k\-1\)\!\!\\,\\sqrt\{\\alpha\_\{k\}\},&k\>2,\\end\{cases\}\(37\)where thek=2k=2expression follows from the same expansion with the resummed momentℳ2\\mathcal\{M\}\_\{2\}and reproduces the AGS spin\-glass line\[[4](https://arxiv.org/html/2609.10976#bib.bib4)\]\. Retrieval solutions exist below the spinodalTR​\(αk\)T\_\{R\}\(\\alpha\_\{k\}\), which connects\(0,Tc\)\(0,T\_\{c\}\)to\(αc,0\)\(\\alpha\_\{c\},0\)\. They are the global minima of the free energy only below the first\-order boundaryTM​\(αk\)T\_\{M\}\(\\alpha\_\{k\}\), obtained by equating the retrieval free energy Eq\. \([21](https://arxiv.org/html/2609.10976#S3.E21)\) with that of them=0m=0branch \(the spin glass forT<TgT<T\_\{g\}and the paramagnet above\)\. Between the two lines retrieval survives as a metastable state\. Fork=2k=2the transition atαk=0\\alpha\_\{k\}=0is continuous, so the metastable band opens only at finite load and closes atT=0T=0betweenαM≃0\.051\\alpha\_\{M\}\\simeq 0\.051andαc≃0\.138\\alpha\_\{c\}\\simeq 0\.138\[[4](https://arxiv.org/html/2609.10976#bib.bib4)\]\. Fork\>2k\>2the transition is first order already atαk=0\\alpha\_\{k\}=0, and the band persists down to zero load, where retrieval is metastable against the paramagnet forTM​\(0\)<T<TcT\_\{M\}\(0\)<T<T\_\{c\}\. The distinct competitor at small load reflects a hierarchy of local stability: akk\-body coupling exerts no mean field on a disordered configuration \(the local field involves a product ofk−1k\-1spins\), so fork\>2k\>2the paramagnet remains locally stable at every temperature, as in the purely ferromagnetickk\-spin model, and only the accumulated crosstalk of many patterns, with the small\-qqvarianceαk​\[\(k−1\)\!\!\]2​q\\alpha\_\{k\}\[\(k\-1\)\!\!\]^\{2\}qunderlying Eq\. \([37](https://arxiv.org/html/2609.10976#S3.E37)\), can freeze them=0m=0sector\. SinceTg→0T\_\{g\}\\to 0asαk→0\\alpha\_\{k\}\\to 0, at small load retrieval competes directly with the paramagnet on a glass\-free background, and the boundaries admit simple estimates: balancing the retrieval energy density−1/k\-1/kagainst the paramagnetic entropylog⁡2\\log 2, and against the zero\-temperature spin\-glass energy density−2​αk​\(2​k−3\)\!\!/π\-\\sqrt\{2\\alpha\_\{k\}\(2k\-3\)\!\!/\\pi\}, yields

TM​\(0\)=1k​log⁡2,αM​\(0\)≃π2​k2​\(2​k−3\)\!\!,T\_\{M\}\(0\)=\\frac\{1\}\{k\\log 2\},\\qquad\\alpha\_\{M\}\(0\)\\simeq\\frac\{\\pi\}\{2k^\{2\}\\,\(2k\-3\)\!\!\},\(38\)where the former holds up to corrections exponentially small inkk, and both reproduce the numerical boundaries of Fig\.[2](https://arxiv.org/html/2609.10976#S3.F2)within a few percent\. The factorial steepening seen in Table[1](https://arxiv.org/html/2609.10976#S3.T1)is directly visible in the figure: askkgrows, the retrieval and metastable regions collapse toward the temperature axis, and the spin\-glass phase comes to dominate the diagram\.

Reentrance of the first\-order boundary\.Fork\>2k\>2the first\-order boundary in Fig\.[2](https://arxiv.org/html/2609.10976#S3.F2)is visibly reentrant:αM​\(T\)\\alpha\_\{M\}\(T\)exceeds its zero\-temperature value at intermediate temperatures, by about20%20\\%fork=4k=4and54%54\\%fork=6k=6\. The mechanism is the entropic asymmetry between the competing states\. In the reentrant window of loads the spin glass has the lower free energy atT=0T=0, but upon heating them=0m=0background melts into the paramagnet already atTg∝αkT\_\{g\}\\propto\\sqrt\{\\alpha\_\{k\}\}, whereas the retrieval state, protected by the large local fieldmk−1m^\{k\-1\}, retains an exponentially small entropy and remains effectively frozen\. Retrieval thereby recovers the global minimum in an intermediate window of temperatures, which closes on the scaleTM​\(0\)T\_\{M\}\(0\)of Eq\. \([38](https://arxiv.org/html/2609.10976#S3.E38)\)\. Consistently, the reentrance is nearly absent fork=2k=2, whereTg=1\+αkT\_\{g\}=1\+\\sqrt\{\\alpha\_\{k\}\}holds the background frozen throughout the retrieval region\. All lines are computed within the RS ansatz\. The spin\-glass phase and the low\-temperature boundaries acquire corrections from replica symmetry breaking \(RSB\)\[[23](https://arxiv.org/html/2609.10976#bib.bib23)\], which are known to be small for the retrieval boundaries atk=2k=2\[[4](https://arxiv.org/html/2609.10976#bib.bib4)\]and are expected to remain so fork\>2k\>2\[[7](https://arxiv.org/html/2609.10976#bib.bib7)\]\. In particular, the far weaker low\-temperature reentrance of the spinodalTRT\_\{R\}\(below four percent for the values ofkkshown\) is the familiar pathology of the RS solution, which RSB removes fork=2k=2while raising the capacity slightly toαc≃0\.144\\alpha\_\{c\}\\simeq 0\.144\[[24](https://arxiv.org/html/2609.10976#bib.bib24)\]\. The reentrance ofTMT\_\{M\}instead operates at intermediate temperatures through the melting of the background, although its magnitude may shift under RSB\.

### III\.2Model C

Model C is thep=2p=2member of the family, called the spherical memory model in\[[19](https://arxiv.org/html/2609.10976#bib.bib19)\]: the visible degrees of freedom are continuous variables on the sphereS=𝕊Nv−1​\(Nv\)S=\\mathbb\{S\}^\{N\_\{v\}\-1\}\(\\sqrt\{N\_\{v\}\}\)\. The model\-C energy function is

EξC​\(v,h\)=Nv​k−1τh​k​∑μhμk−λτh​τv​∑μ,ihμk−1​ξμ​i\(h,v\)​xi,\{E\_\{\\xi\}^\{\\mathrm\{C\}\}\}\(v,h\)=N\_\{v\}\\frac\{k\-1\}\{\{\\tau\_\{h\}\}k\}\\sum\_\{\\mu\}h\_\{\\mu\}^\{k\}\-\\frac\{\\lambda\}\{\{\\tau\_\{h\}\}\{\\tau\_\{v\}\}\}\\sum\_\{\\mu,i\}h\_\{\\mu\}^\{k\-1\}\\xi^\{\(h,v\)\}\_\{\\mu i\}\\mathrm\{x\}\_\{i\},\(39\)where the factorNvN\_\{v\}in the first term is the same extensivity insertion as in Model A\. The partition function then reads

ZξC\(β\)=∫dhe−Nvγk∑μmμk×∫Sd​Ω​\(x\)​exp⁡\{βk​∑μ,imμk−1​ξμ​i\(h,v\)​xi\},\{Z\_\{\\xi\}^\{\\mathrm\{C\}\}\}\(\\beta\)=\\int dh\\,e^\{\-N\_\{v\}\\gamma\_\{k\}\\sum\_\{\\mu\}m\_\{\\mu\}^\{k\}\}\\\\ \\times\\int\_\{S\}d\\Omega\(\\mathrm\{x\}\)\\exp\\left\\\{\\beta\_\{k\}\\sum\_\{\\mu,i\}m\_\{\\mu\}^\{k\-1\}\\xi^\{\(h,v\)\}\_\{\\mu i\}\\mathrm\{x\}\_\{i\}\\right\\\},\(40\)with the sameγk\\gamma\_\{k\},βk\\beta\_\{k\}, andmμm\_\{\\mu\}as in Model A\. The only change relative to Eq\. \([17](https://arxiv.org/html/2609.10976#S3.E17)\) is that the visible trace runs over the sphereSSinstead of the hypercube\{±1\}Nv\\left\\\{\\pm 1\\right\\\}^\{N\_\{v\}\}\. In particular, the stationarity with respect tomμm\_\{\\mu\}again identifiesmμm\_\{\\mu\}with the overlap between the patternξμ\(h,v\)\\xi^\{\(h,v\)\}\_\{\\mu\}and the visible configurationx\\mathrm\{x\}\.

The patterns are now drawn independently and uniformly from the same sphere,ξμ\(h,v\)∼Unif⁡\(S\)\\xi^\{\(h,v\)\}\_\{\\mu\}\\sim\\mathrm\{Unif\}\(S\), and we work in the same high\-load regimeNh=αk​Nvk−1N\_\{h\}=\\alpha\_\{k\}N\_\{v\}^\{k\-1\}with a single condensed pattern\. By the rotational invariance of the pattern ensemble and of the visible measure, the condensed pattern can be rotated toξ1\(h,v\)=\(1,…,1\)\\xi^\{\(h,v\)\}\_\{1\}=\(1,\\ldots,1\)\(the continuous analogue of the gauge choice for binary patterns\) exactly at finiteNvN\_\{v\}, while the remaining patterns are asymptotically Gaussian, with fixed\-norm corrections that do not affect the RS equations\. Fork\>2k\>2, the analysis is again performed in the adiabatic limitβ/τh→∞\\beta/\{\\tau\_\{h\}\}\\to\\inftyat fixedβk\\beta\_\{k\}\. The details of the replica computation are presented in Appendix[B\.3](https://arxiv.org/html/2609.10976#A2.SS3)\.

In contrast to Model A, the visible trace is a Gaussian integral on the sphere and can be carried out exactly, so no single\-site integral survives in the final expressions\. Under the RS ansatz, after eliminating the condensed hidden mode and the Lagrange multiplier of the spherical constraint at their saddle points, we obtain the free energy, up to an additive constant independent of the order parameters,

β​fC=\\displaystyle\\beta f^\{\\mathrm\{C\}\}=−βkk​mk\+αk2​Ψk​\(q\)\\displaystyle\-\\frac\{\\beta\_\{k\}\}\{k\}m^\{k\}\+\\frac\{\\alpha\_\{k\}\}\{2\}\\Psi\_\{k\}\(q\)\(41\)−12​\{log⁡\(1−q\)\+q−m21−q\},\\displaystyle\-\\frac\{1\}\{2\}\\left\\\{\\log\(1\-q\)\+\\frac\{q\-m^\{2\}\}\{1\-q\}\\right\\\},wheremmis the overlap between the visible configuration and the condensed pattern,qqis the Edwards–Anderson order parameter of the spherical visible state, andΨk\\Psi\_\{k\}is the same noise entropic term as in Eq\. \([21](https://arxiv.org/html/2609.10976#S3.E21)\)\. The equations of state follow from the stationarity of the free energy:

0\\displaystyle 0=\(1−βk​\(1−q\)​mk−2\)​m,\\displaystyle=\\left\(1\-\\beta\_\{k\}\(1\-q\)\\,m^\{k\-2\}\\right\)m,\(42\)q−m2\(1−q\)2\\displaystyle\\frac\{q\-m^\{2\}\}\{\(1\-q\)^\{2\}\}=αk​βk2​ℳk​\(q\),\\displaystyle=\\alpha\_\{k\}\\beta\_\{k\}^\{2\}\\,\\mathcal\{M\}\_\{k\}\(q\),\(43\)where we usedΨk′​\(q\)=βk2​ℳk​\(q\)\\Psi\_\{k\}^\{\\prime\}\(q\)=\\beta\_\{k\}^\{2\}\\mathcal\{M\}\_\{k\}\(q\), which holds uniformly inkk, with the same crosstalk momentℳk\\mathcal\{M\}\_\{k\}as in Eq\. \([26](https://arxiv.org/html/2609.10976#S3.E26)\)\. The crosstalk strengthr=ℳk​\(q\)r=\\mathcal\{M\}\_\{k\}\(q\)of Eq\. \([25](https://arxiv.org/html/2609.10976#S3.E25)\) thus carries over unchanged\. In the gaugeξ1\(h,v\)=\(1,…,1\)\\xi^\{\(h,v\)\}\_\{1\}=\(1,\\ldots,1\), the RS saddle point gives⟨xi⟩=m\\left\\langle\\mathrm\{x\}\_\{i\}\\right\\rangle=mand1Nv​∑i⟨xi⟩2=q\\frac\{1\}\{N\_\{v\}\}\\sum\_\{i\}\\left\\langle\\mathrm\{x\}\_\{i\}\\right\\rangle^\{2\}=q\. Since the spherical constraint fixes1Nv​∑i⟨xi2⟩=1\\frac\{1\}\{N\_\{v\}\}\\sum\_\{i\}\\left\\langle\\mathrm\{x\}\_\{i\}^\{2\}\\right\\rangle=1, the combinationβk​\(1−q\)\\beta\_\{k\}\(1\-q\)is the static susceptibility of the spherical state\. The crosstalk moment is the same as in Model A because the non\-condensed overlapsNv−1/2∑iξ\(h,v\)μ​ixiN\_\{v\}^\{\-1/2\}\\sum\_\{i\}\\xi^\{\(h,v\)\}\_\{\\mu i\}\\mathrm\{x\}\_\{i\}obey the same central\-limit statistics, governed solely byqq: the crosstalk is universal across the visible ensembles, and the dichotomy of the Onsager reaction field discussed below Eq\. \([26](https://arxiv.org/html/2609.10976#S3.E26)\) applies without modification\. The signal equation Eq\. \([42](https://arxiv.org/html/2609.10976#S3.E42)\), in turn, is of the form familiar from the ferromagnetically biased sphericalpp\-spin model\[[25](https://arxiv.org/html/2609.10976#bib.bib25)\], thekk\-body signal competing with the entropy of the sphere\.

![Refer to caption](https://arxiv.org/html/2609.10976v1/fig/modelC_phase_diagram_k2_k4_k6.png)Figure 3:RS phase diagram of Model C in the\(αk,T\)\(\\alpha\_\{k\},T\)plane for \(a\)k=2k=2, \(b\)k=4k=4, and \(c\)k=6k=6, in the same conventions as Fig\.[2](https://arxiv.org/html/2609.10976#S3.F2)\. Fork=2k=2retrieval survives only on the segmentαk=0\\alpha\_\{k\}=0,T≤1T\\leq 1\(red line on the vertical axis\), reflecting the zero capacity of the spherical model\. Fork\>2k\>2the diagram has the same topology as that of Model A, with uniformly smaller retrieval regions\.The casek=2k=2\.Fork=2k=2the signal equation degenerates: a retrieval solutionm\>0m\>0requiresβk​\(1−q\)=1\\beta\_\{k\}\(1\-q\)=1, which is precisely the marginality condition at which the AGS denominator1−βk​\(1−q\)1\-\\beta\_\{k\}\(1\-q\)ofℳ2\\mathcal\{M\}\_\{2\}vanishes\. The right\-hand side of Eq\. \([43](https://arxiv.org/html/2609.10976#S3.E43)\) then diverges, so that no RS retrieval solution exists at anyαk\>0\\alpha\_\{k\}\>0, includingT=0T=0: the quadratic spherical model has zero storage capacity\. This is the known marginality of the spherical Hopfield model\[[20](https://arxiv.org/html/2609.10976#bib.bib20)\], in which a retrieval phase is recovered only after the energy function is stabilized by an additional quartic term\. Atαk=0\\alpha\_\{k\}=0, Eq\. \([43](https://arxiv.org/html/2609.10976#S3.E43)\) givesq=m2q=m^\{2\}, and the signal equation yieldsm2=1−Tm^\{2\}=1\-T: retrieval sets in continuously belowTc=1T\_\{c\}=1\.

Zero\-temperature capacity fork\>2k\>2\.Sendingβk→∞\\beta\_\{k\}\\to\\inftywithC=βk​\(1−q\)C=\\beta\_\{k\}\(1\-q\)finite, as in Model A, the signal equation givesC=m2−kC=m^\{2\-k\}, whileq→1q\\to 1andℳk​\(1\)=\(2​k−3\)\!\!\\mathcal\{M\}\_\{k\}\(1\)=\(2k\-3\)\!\!\. The noise equation Eq\. \([43](https://arxiv.org/html/2609.10976#S3.E43)\) then closes algebraically \(no error function appears, since the visible integral is Gaussian\), and any retrieval solution lies on the curve

αk​\(m\)=\(1−m2\)​m2​\(k−2\)\(2​k−3\)\!\!,\\alpha\_\{k\}\(m\)=\\frac\{\(1\-m^\{2\}\)\\,m^\{2\(k\-2\)\}\}\{\(2k\-3\)\!\!\},\(44\)whose maximum over0<m<10<m<1is attained atmc2=\(k−2\)/\(k−1\)m\_\{c\}^\{2\}=\(k\-2\)/\(k\-1\)and yields the capacity in closed form:

αc=\(k−2\)k−2\(k−1\)k−1​\(2​k−3\)\!\!,\\alpha\_\{c\}=\\frac\{\(k\-2\)^\{k\-2\}\}\{\(k\-1\)^\{k\-1\}\\,\(2k\-3\)\!\!\},\(45\)again a spinodal at which the stable and unstable retrieval branches merge\. This givesαc=4/405≃9\.88×10−3\\alpha\_\{c\}=4/405\\simeq 9\.88\\times 10^\{\-3\}fork=4k=4andαc≃8\.67×10−5\\alpha\_\{c\}\\simeq 8\.67\\times 10^\{\-5\}fork=6k=6\. Note that the retrieval overlap at capacity,mc≃0\.82m\_\{c\}\\simeq 0\.82and0\.890\.89respectively, lies visibly below the corresponding values0\.920\.92and0\.960\.96of Model A: even at zero temperature, the soft spherical spins trade retrieval quality against the crosstalk\.

Critical temperature and retrieval boundary\.Atαk=0\\alpha\_\{k\}=0the equations of state giveq=m2q=m^\{2\}and1=βk​\(1−m2\)​mk−21=\\beta\_\{k\}\(1\-m^\{2\}\)\\,m^\{k\-2\}, whose solvability condition determines, fork\>2k\>2,

Tc=2k​\(k−2k\)\(k−2\)/2\.T\_\{c\}=\\frac\{2\}\{k\}\\left\(\\frac\{k\-2\}\{k\}\\right\)^\{\(k\-2\)/2\}\.\(46\)The transition is again of the spinodal type, with the overlap jumping tomc2=\(k−2\)/km\_\{c\}^\{2\}=\(k\-2\)/katTcT\_\{c\}\. The resulting critical values are collected in Table[2](https://arxiv.org/html/2609.10976#S3.T2)\. As for Model A, the retrieval boundary steepens asTc/αc∼\(2​k−3\)\!\!T\_\{c\}/\\alpha\_\{c\}\\sim\(2k\-3\)\!\!, driven by the same factorial growth of the crosstalk variance, while every entry is uniformly smaller than its model\-A counterpart in Table[1](https://arxiv.org/html/2609.10976#S3.T1)\.

Table 2:Critical values of Model C obtained from the RS equations of state, in the same conventions as Table[1](https://arxiv.org/html/2609.10976#S3.T1)\. Fork=2k=2retrieval survives only atαk=0\\alpha\_\{k\}=0, so thatαc=0\\alpha\_\{c\}=0and the steepness is not defined\.Phase diagram\.The resulting RS phase diagram is shown in Fig\.[3](https://arxiv.org/html/2609.10976#S3.F3)\. Since the crosstalk momentℳk\\mathcal\{M\}\_\{k\}is universal, them=0m=0sector is governed by the same equations as in Model A, and the continuous spin\-glass lineTg​\(αk\)T\_\{g\}\(\\alpha\_\{k\}\)of Eq\. \([37](https://arxiv.org/html/2609.10976#S3.E37)\) carries over unchanged\. The first\-order boundaryTM​\(αk\)T\_\{M\}\(\\alpha\_\{k\}\)again follows by equating the free energy Eq\. \([41](https://arxiv.org/html/2609.10976#S3.E41)\) of the retrieval branch with that of them=0m=0branch and, in contrast to Model A, its zero\-temperature endpoint is available in closed form: along the capacity curve Eq\. \([44](https://arxiv.org/html/2609.10976#S3.E44)\), the energy balance between the retrieval and spin\-glass states reduces to the condition1−m2=1/\(k−1\)\\sqrt\{1\-m^\{2\}\}=1/\(k\-1\), which yields

αM​\(0\)=kk−2​\(k−2\)k−2\(k−1\)2​k−2​\(2​k−3\)\!\!,\\alpha\_\{M\}\(0\)=\\frac\{k^\{k\-2\}\\,\(k\-2\)^\{k\-2\}\}\{\(k\-1\)^\{2k\-2\}\\,\(2k\-3\)\!\!\},\(47\)i\.e\.,αM​\(0\)≃5\.85×10−3\\alpha\_\{M\}\(0\)\\simeq 5\.85\\times 10^\{\-3\}fork=4k=4and3\.60×10−53\.60\\times 10^\{\-5\}fork=6k=6\. In contrast to Model A, neither boundary of Model C is reentrant: the zero\-temperature slope of the spinodal is strictly negative,

∂TαR\|T=0∝\(1−mc2\)​ℳk′​\(1\)−ℳk​\(1\)<0,\\partial\_\{T\}\\alpha\_\{R\}\|\_\{T=0\}\\propto\(1\-m\_\{c\}^\{2\}\)\\,\\mathcal\{M\}\_\{k\}^\{\\prime\}\(1\)\-\\mathcal\{M\}\_\{k\}\(1\)<0,\(48\)and the first\-order boundary is likewise monotone\. The reentrance mechanism of Model A is suppressed by the visible entropy: the spherical retrieval state pays the confinement entropy−12​log⁡\(1−q\)≃12​log⁡βk\-\\frac\{1\}\{2\}\\log\(1\-q\)\\simeq\\frac\{1\}\{2\}\\log\\beta\_\{k\}of the condensate, which grows without bound at low temperature, so it melts together with them=0m=0background instead of remaining frozen against it\. Fork=2k=2the retrieval phase degenerates to the segmentαk=0\\alpha\_\{k\}=0,T≤1T\\leq 1, the graphical expression of the marginality discussed above, while the spin\-glass line, being common to the two models, is unaffected by the collapse of the retrieval sector\.

The comparison between Models A and C, displayed side by side in Fig\.[2](https://arxiv.org/html/2609.10976#S3.F2)and Fig\.[3](https://arxiv.org/html/2609.10976#S3.F3), thus isolates the role of the visible degrees of freedom: the crosstalk momentℳk\\mathcal\{M\}\_\{k\}is universal, so all differences, from the collapse of thek=2k=2capacity to the absence of the reentrant boundary, originate from the visible entropy\. Fork=2k=2, the saturating sign activation of the Ising spins is essential for retrieval: replacing it by the linear spherical activation leaves the retrieval state only marginally confined and destroys the capacity entirely\. Fork\>2k\>2, thekk\-body signal restores a genuine retrieval phase, though with capacity and critical temperature reduced byO⁡\(1\)O\(1\)factors relative to Model A at eachkk\. The retrieval landscape of this family is therefore shaped jointly by the sharpness of the visible activation and by the order of the hidden nonlinearity, with the latter dominating at largekkthrough the common factorial collapse of Eq\. \([32](https://arxiv.org/html/2609.10976#S3.E32)\) and Eq\. \([45](https://arxiv.org/html/2609.10976#S3.E45)\)\.

## IVModel B

Model B is the attention model of the classℋ\\mathcal\{H\}\[[19](https://arxiv.org/html/2609.10976#bib.bib19)\], defined by the Lagrangians

Lv\(v\)=12‖v‖2,Lh\(h\)=log∑μehμ,\{L\_\{v\}\}\(v\)=\\frac\{1\}\{2\}\\left\\lVert v\\right\\rVert^\{2\},\\qquad\{L\_\{h\}\}\(h\)=\\log\\sum\_\{\\mu\}e^\{h\_\{\\mu\}\},\(49\)with the activation functionsg⁡\(v\)=vg\(v\)=vandf⁡\(h\)=softmax⁡\(h\)f\(h\)=\\mathrm\{softmax\}\(h\), wheresoftmax​\(h\)μ≔ehμ/∑νehν\\mathrm\{softmax\}\(h\)\_\{\\mu\}\\coloneq e^\{h\_\{\\mu\}\}/\\sum\_\{\\nu\}e^\{h\_\{\\nu\}\}\. From Eq\. \([4](https://arxiv.org/html/2609.10976#S2.E4)\), the model\-B energy function reads

EξB​\(v,h\)\\displaystyle\{E\_\{\\xi\}^\{\\mathrm\{B\}\}\}\(v,h\)=12​τv‖v‖2\+1τh\(h⊤softmax\(h\)−log∑μehμ\)\\displaystyle=\\frac\{1\}\{2\{\\tau\_\{v\}\}\}\\left\\lVert v\\right\\rVert^\{2\}\+\\frac\{1\}\{\{\\tau\_\{h\}\}\}\\left\(h^\{\\top\}\\mathrm\{softmax\}\(h\)\-\\log\\sum\_\{\\mu\}e^\{h\_\{\\mu\}\}\\right\)\(50\)−λτh​τv​softmax​\(h\)⊤​ξ\(h,v\)​v\.\\displaystyle\-\\frac\{\\lambda\}\{\{\\tau\_\{h\}\}\{\\tau\_\{v\}\}\}\\mathrm\{softmax\}\(h\)^\{\\top\}\\xi^\{\(h,v\)\}v\.The hidden Legendre term has an information\-theoretic meaning: in terms of the activationf=softmax⁡\(h\)f=\\mathrm\{softmax\}\(h\)it equals∑μfμ​log⁡fμ\\sum\_\{\\mu\}f\_\{\\mu\}\\log f\_\{\\mu\}, the negative Shannon entropy of the probability vectorff\. Since the energy depends onhhonly throughff\(the softmax is invariant under the uniform shifth→h\+c​1h\\to h\+c\\,\\bm\{1\}\), we may change variables and regard the energy as a function onℝNv×ΔNh\\mathbb\{R\}^\{N\_\{v\}\}\\times\\Delta^\{N\_\{h\}\},

EξB​\(v,f\)=‖v‖22​τv\+1τh​∑μfμ​log⁡fμ−λτh​τv​f⊤​ξ\(h,v\)​v,\{E\_\{\\xi\}^\{\\mathrm\{B\}\}\}\(v,f\)=\\frac\{\\left\\lVert v\\right\\rVert^\{2\}\}\{2\{\\tau\_\{v\}\}\}\+\\frac\{1\}\{\{\\tau\_\{h\}\}\}\\sum\_\{\\mu\}f\_\{\\mu\}\\log f\_\{\\mu\}\-\\frac\{\\lambda\}\{\{\\tau\_\{h\}\}\{\\tau\_\{v\}\}\}f^\{\\top\}\\xi^\{\(h,v\)\}v,\(51\)whereΔNh≔\{f∈ℝNh\|fμ\>0,∑μfμ=1\}\\Delta^\{N\_\{h\}\}\\coloneq\\left\\\{f\\in\\mathbb\{R\}^\{N\_\{h\}\}\\,\\middle\|\\,f\_\{\\mu\}\>0,~\\sum\_\{\\mu\}f\_\{\\mu\}=1\\right\\\}is the open simplex\. In these variables the hidden sector reduces to the normalized weightsffthat the network assigns to the stored patterns\. These are the attention weights of the modern Hopfield network and of the transformer attention mechanism\[[12](https://arxiv.org/html/2609.10976#bib.bib12),[13](https://arxiv.org/html/2609.10976#bib.bib13),[14](https://arxiv.org/html/2609.10976#bib.bib14)\], and, in accordance with the general discussion of Sec\.[II\.1](https://arxiv.org/html/2609.10976#S2.SS1), they constitute the order parameter of memory retrieval: retrieval of the patternμ\\mucorresponds to the concentration offfat the vertexeμe\_\{\\mu\}of the simplex\.

Following Sec\.[II\.2](https://arxiv.org/html/2609.10976#S2.SS2), the partition function of Model B is defined by444The measured​fdfdescends from the flat measured​hdhafter the zero mode along the uniform shift is factored out\. This procedure affects only subexponential prefactors and is detailed in Appendix[C\.2](https://arxiv.org/html/2609.10976#A3.SS2)\.

ZξB​\(β\)=∫ℝNv×ΔNhd​v​𝑑f​exp⁡\(−β​EξB​\(v,f\)\)\.\{Z\_\{\\xi\}^\{\\mathrm\{B\}\}\}\(\\beta\)=\\int\_\{\\mathbb\{R\}^\{N\_\{v\}\}\\times\\Delta^\{N\_\{h\}\}\}dv\\,df\\exp\\left\(\-\\beta\{E\_\{\\xi\}^\{\\mathrm\{B\}\}\}\(v,f\)\\right\)\.\(52\)In the adiabatic limitβ/τh→∞\\beta/\{\\tau\_\{h\}\}\\to\\infty, theffintegral localizes at the stationary point

f∗=softmax⁡\(λτv​ξ\(h,v\)​v\),f^\{\\ast\}=\\mathrm\{softmax\}\\left\(\\frac\{\\lambda\}\{\{\\tau\_\{v\}\}\}\\xi^\{\(h,v\)\}v\\right\),\(53\)which is preciselyf⁡\(h∗\)f\(h\_\{\\ast\}\)evaluated at the adiabatic saddle pointh∗h\_\{\\ast\}of Sec\.[II\.2](https://arxiv.org/html/2609.10976#S2.SS2), and the partition function reduces to the visible integral Eq\. \([10](https://arxiv.org/html/2609.10976#S2.E10)\) with the effective energy

EξB​\(v,h∗\)=‖v‖22​τv−1τh​log​∑μeλτv​\(ξ\(h,v\)​v\)μ\.\{E\_\{\\xi\}^\{\\mathrm\{B\}\}\}\(v,h\_\{\\ast\}\)=\\frac\{\\left\\lVert v\\right\\rVert^\{2\}\}\{2\{\\tau\_\{v\}\}\}\-\\frac\{1\}\{\{\\tau\_\{h\}\}\}\\log\\sum\_\{\\mu\}e^\{\\frac\{\\lambda\}\{\{\\tau\_\{v\}\}\}\(\\xi^\{\(h,v\)\}v\)\_\{\\mu\}\}\.\(54\)The analysis below depends on the parameters only through the two combinationsβ/τh\\beta/\{\\tau\_\{h\}\}andβ~≔\(β/τv\)​\(λ/τh\)2\\tilde\{\\beta\}\\coloneq\(\\beta/\{\\tau\_\{v\}\}\)\(\\lambda/\{\\tau\_\{h\}\}\)^\{2\}\. In parallel with the reduction of the couplings toβk\\beta\_\{k\}for Models A and C, we fix the normalizationτv=1\{\\tau\_\{v\}\}=1andτh=λ\{\\tau\_\{h\}\}=\\lambda, for whichβ~=β\\tilde\{\\beta\}=\\beta: the effective visible energy in Eq\. \([54](https://arxiv.org/html/2609.10976#S4.E54)\) becomes12​‖v‖2−1λ​log​∑μeλ​\(ξ\(h,v\)​v\)μ\\frac\{1\}\{2\}\\left\\lVert v\\right\\rVert^\{2\}\-\\frac\{1\}\{\\lambda\}\\log\\sum\_\{\\mu\}e^\{\\lambda\(\\xi^\{\(h,v\)\}v\)\_\{\\mu\}\}, precisely the energy function of the modern Hopfield network\[[13](https://arxiv.org/html/2609.10976#bib.bib13)\]in the form whose exponential storage was analyzed in\[[15](https://arxiv.org/html/2609.10976#bib.bib15)\]\. In this normalizationλ\\lambdacontrols the sharpness of the attention, whileβ\\betaremains the genuine inverse temperature\.

As for Models A and C, for Model B the visible sector can be integrated out exactly: the visible Lagrangian is quadratic, so thevvintegral in Eq\. \([52](https://arxiv.org/html/2609.10976#S4.E52)\) is Gaussian\. Carrying it out yields, up to an overall constant,

ZξB​\(β\)\\displaystyle\{Z\_\{\\xi\}^\{\\mathrm\{B\}\}\}\(\\beta\)∝∫ΔNhd​f​eΦ⁡\(f\),\\displaystyle\\propto\\int\_\{\\Delta^\{N\_\{h\}\}\}df\\,e^\{\\Phi\(f\)\},\(55\)Φ⁡\(f\)\\displaystyle\\Phi\(f\)=−βλ∑μfμlogfμ\+β2f⊤Gf,\\displaystyle=\-\\frac\{\\beta\}\{\\lambda\}\\sum\_\{\\mu\}f\_\{\\mu\}\\log f\_\{\\mu\}\+\\frac\{\\beta\}\{2\}\\,f^\{\\top\}Gf,\(56\)whereG≔ξ\(h,v\)​ξ\(v,h\)G\\coloneq\\xi^\{\(h,v\)\}\\xi^\{\(v,h\)\}is the Gram matrix of the patterns,Gμ​ν=ξμ\(h,v\)⋅ξν\(h,v\)G\_\{\\mu\\nu\}=\\xi^\{\(h,v\)\}\_\{\\mu\}\\cdot\\xi^\{\(h,v\)\}\_\{\\nu\}\. We refer to Eq\. \([56](https://arxiv.org/html/2609.10976#S4.E56)\) as the*ffrepresentation*of Model B: the entire thermodynamics is expressed by the attention weights alone, as a competition between the attention entropy, which favors delocalized attention, and the positive semi\-definite energyf⊤​G​f=‖ξ\(v,h\)​f‖2f^\{\\top\}Gf=\\left\\lVert\\xi^\{\(v,h\)\}f\\right\\rVert^\{2\}, which favors concentration\.

We draw the patterns independently as Gaussian variables,ξμ​i\(h,v\)∼𝒩⁡\(0,1\)\\xi^\{\(h,v\)\}\_\{\\mu i\}\\sim\\mathcal\{N\}\(0,1\), for whichGμ​μ=Nv\(1\+O\(Nv−1/2\)\)G\_\{\\mu\\mu\}=N\_\{v\}\(1\+O\(N\_\{v\}^\{\-1/2\}\)\)whileGμ​ν=O⁡\(Nv\)G\_\{\\mu\\nu\}=O\(\\sqrt\{N\_\{v\}\}\)forμ≠ν\\mu\\neq\\nu\. The two terms ofΦ\\Phiare then of widely different orders\. The energy is extensive: attention concentrated on a single pattern already gainsβ2​Gμ​μ≈β2​Nv\\frac\{\\beta\}\{2\}G\_\{\\mu\\mu\}\\approx\\frac\{\\beta\}\{2\}N\_\{v\}\. The entropy, by contrast, is bounded by its value at uniform attention,βλ​log⁡Nh\\frac\{\\beta\}\{\\lambda\}\\log N\_\{h\}: it is of order one per stored pattern rather than per neuron\. A genuine competition between the two therefore requireslog⁡Nh∼Nv\\log N\_\{h\}\\sim N\_\{v\}, and the natural high\-load regime of Model B is the exponential load

Nh=eα​Nv,N\_\{h\}=e^\{\\alpha N\_\{v\}\},\(57\)in contrast with the polynomial loadsNh=αk​Nvk−1N\_\{h\}=\\alpha\_\{k\}N\_\{v\}^\{k\-1\}of Models A and C\. A second contrast concerns the method: the disorder now enters only through the Gram matrix, whose relevant statistics are large deviations of pattern norms and overlaps, so the analysis below proceeds by direct counting arguments rather than by the replica method\. The relation between the two approaches is clarified in Sec\.[IV\.1\.1](https://arxiv.org/html/2609.10976#S4.SS1.SSS1)\.

### IV\.1Capacity at zero temperature

Vertex condensation\.The geometry of Eq\. \([56](https://arxiv.org/html/2609.10976#S4.E56)\) dictates the structure of the low\-temperature states\. The energyf⊤​G​f=‖∑μfμ​ξμ\(h,v\)‖2f^\{\\top\}Gf=\\left\\lVert\\sum\_\{\\mu\}f\_\{\\mu\}\\xi^\{\(h,v\)\}\_\{\\mu\}\\right\\rVert^\{2\}is a convex function offf, and a convex function on a compact convex set attains its maximum at an extreme point \(Bauer’s maximum principle\)\. On the simplex the extreme points are the vertices\. Explicitly,

‖∑μfμ​ξμ\(h,v\)‖≤∑μfμ​‖ξμ\(h,v\)‖≤maxμ⁡‖ξμ\(h,v\)‖,\\left\\lVert\\sum\_\{\\mu\}f\_\{\\mu\}\\,\\xi^\{\(h,v\)\}\_\{\\mu\}\\right\\rVert\\leq\\sum\_\{\\mu\}f\_\{\\mu\}\\left\\lVert\\xi^\{\(h,v\)\}\_\{\\mu\}\\right\\rVert\\leq\\max\_\{\\mu\}\\left\\lVert\\xi^\{\(h,v\)\}\_\{\\mu\}\\right\\rVert,\(58\)with equality at the vertex of the maximizing pattern\. The entropy opposes this concentration only weakly: spreading the attention uniformly overMMpatterns reduces the energy gain fromβ2​Nv\\frac\{\\beta\}\{2\}N\_\{v\}toβ2​Nv/M\\frac\{\\beta\}\{2\}N\_\{v\}/Mat leading order, an extensive loss, whereas the entropy gain is merelyβλ​log⁡M\\frac\{\\beta\}\{\\lambda\}\\log M\. For any subexponential load the entropy is thus negligible at leading order, and the Gibbs measure condenses onto the vertices of the simplex, each vertex describing the retrieval state of one pattern withvvlocalized near it \(cf\. Eq\. \([53](https://arxiv.org/html/2609.10976#S4.E53)\)\)\. At the exponential load Eq\. \([57](https://arxiv.org/html/2609.10976#S4.E57)\) the competition becomes genuine, but not through interior attention distributions: the measure instead spreads over exponentially many nearly pure vertex states,ZξB≈∑μZμ\{Z\_\{\\xi\}^\{\\mathrm\{B\}\}\}\\approx\\sum\_\{\\mu\}Z\_\{\\mu\}, and the entropy is the counting entropy of this decomposition\. The capacity question is whether the state condensed on a given typical pattern survives against this exponentially large background\.

Retrieval capacity\.Fix the retrieved patternξ1\(h,v\)\\xi^\{\(h,v\)\}\_\{1\}, of typical norm‖ξ1\(h,v\)‖2≈Nv\\left\\lVert\\xi^\{\(h,v\)\}\_\{1\}\\right\\rVert^\{2\}\\approx N\_\{v\}\. The stationarity of the effective energy in Eq\. \([54](https://arxiv.org/html/2609.10976#S4.E54)\) is the fixed\-point conditionv=ξ\(v,h\)​f∗​\(v\)v=\\xi^\{\(v,h\)\}f^\{\\ast\}\(v\): the visible state is the attention\-weighted superposition of the stored patterns\. At zero temperature the retrieval state lies atv≈ξ1\(h,v\)v\\approx\\xi^\{\(h,v\)\}\_\{1\}, up to corrections controlled by the leak of attention computed below, and each pattern feels the attention fieldaμ=λ​ξμ\(h,v\)⋅va\_\{\\mu\}=\\lambda\\,\\xi^\{\(h,v\)\}\_\{\\mu\}\\cdot v\. Atv=ξ1\(h,v\)v=\\xi^\{\(h,v\)\}\_\{1\}one hasa1=λ​‖ξ1\(h,v\)‖2≈λ​Nva\_\{1\}=\\lambda\\left\\lVert\\xi^\{\(h,v\)\}\_\{1\}\\right\\rVert^\{2\}\\approx\\lambda N\_\{v\}, while forμ≥2\\mu\\geq 2, conditioned onξ1\(h,v\)\\xi^\{\(h,v\)\}\_\{1\}, the overlapsξμ\(h,v\)⋅ξ1\(h,v\)\\xi^\{\(h,v\)\}\_\{\\mu\}\\cdot\\xi^\{\(h,v\)\}\_\{1\}are independent centered Gaussians of variance‖ξ1\(h,v\)‖2\\left\\lVert\\xi^\{\(h,v\)\}\_\{1\}\\right\\rVert^\{2\}, so thataμ=λ​Nv​ωμa\_\{\\mu\}=\\lambda\\sqrt\{N\_\{v\}\}\\,\\omega\_\{\\mu\}with independent standard Gaussian variablesωμ\\omega\_\{\\mu\}\. The condensed attention weightp≔f1∗p\\coloneq f^\{\\ast\}\_\{1\}, the model\-B analogue of the retrieval overlapmmof Models A and C, then obeys

1−p≤e−a1​L,L≔∑μ≥2eλ​Nv​ωμ\.1\-p\\leq e^\{\-a\_\{1\}\}L,\\qquad L\\coloneq\\sum\_\{\\mu\\geq 2\}e^\{\\lambda\\sqrt\{N\_\{v\}\}\\,\\omega\_\{\\mu\}\}\.\(59\)The sumLLis evaluated by the counting argument that underlies extreme value statistics: the number of patterns whose variableωμ\\omega\_\{\\mu\}reaches the levelx​Nvx\\sqrt\{N\_\{v\}\}iseNv​\(α−x2/2\)e^\{N\_\{v\}\(\\alpha\-x^\{2\}/2\)\}to leading exponential order, so the levels up toxmax=2​αx\_\{\\max\}=\\sqrt\{2\\alpha\}are populated while higher levels are empty with high probability\. Retaining the maximal term,

log⁡LNv=max0≤x≤2​α⁡\[α−x22\+λ​x\]=\{α\+λ22,λ≤2​α,λ​2​α,λ≥2​α\.\\frac\{\\log L\}\{N\_\{v\}\}=\\max\_\{0\\leq x\\leq\\sqrt\{2\\alpha\}\}\\big\[\\alpha\-\\tfrac\{x^\{2\}\}\{2\}\+\\lambda x\\big\]=\\begin\{cases\}\\alpha\+\\dfrac\{\\lambda^\{2\}\}\{2\},&\\lambda\\leq\\sqrt\{2\\alpha\},\\\\\[5\.69054pt\] \\lambda\\sqrt\{2\\alpha\},&\\lambda\\geq\\sqrt\{2\\alpha\}\.\\end\{cases\}\(60\)In the first branch the maximum is attained at the interior pointx∗=λx^\{\\ast\}=\\lambda: the leak is carried by exponentially many patterns of moderate overlap, andLLis self\-averaging and coincides with its annealed average\. In the second branch the maximum is pinned at the boundaryxmaxx\_\{\\max\}: the leak is dominated by theO⁡\(1\)O\(1\)most aligned patterns, and the sum is frozen\. The retrieval state is self\-consistent precisely when the leak vanishes,−λ\+Nv−1​log⁡L<0\-\\lambda\+N\_\{v\}^\{\-1\}\\log L<0\. Combining the two branches yields the zero\-temperature capacity

αc​\(λ\)=\{λ−λ22,λ≤1,12,λ≥1,\\alpha\_\{c\}\(\\lambda\)=\\begin\{cases\}\\lambda\-\\dfrac\{\\lambda^\{2\}\}\{2\},&\\lambda\\leq 1,\\\\\[5\.69054pt\] \\dfrac\{1\}\{2\},&\\lambda\\geq 1,\\end\{cases\}\(61\)the two branches matching continuously atλ=1\\lambda=1\. Details of these estimates are collected in Appendix[C\.1](https://arxiv.org/html/2609.10976#A3.SS1)\.

Retrieval fails in physically distinct ways in the two regimes of Eq\. \([61](https://arxiv.org/html/2609.10976#S4.E61)\)\. Forλ<1\\lambda<1the attention is too soft: retrieval is destroyed by the aggregate crosstalk of exponentially many weakly correlated patterns, and sharpening the attention raises the capacity\. Forλ\>1\\lambda\>1the crosstalk is dominated by the single most aligned competitor, whose overlapmaxμ≥2⁡ξμ\(h,v\)⋅ξ1\(h,v\)≈2​α​Nv\\max\_\{\\mu\\geq 2\}\\xi^\{\(h,v\)\}\_\{\\mu\}\\cdot\\xi^\{\(h,v\)\}\_\{1\}\\approx\\sqrt\{2\\alpha\}\\,N\_\{v\}matches the signal‖ξ1\(h,v\)‖2≈Nv\\left\\lVert\\xi^\{\(h,v\)\}\_\{1\}\\right\\rVert^\{2\}\\approx N\_\{v\}atα=1/2\\alpha=1/2\. Sinceλ\\lambdamultiplies the signal and the competitor field alike, it cancels from this comparison: no attention sharpness overcomes a competitor as aligned as the signal itself, hence the ceiling\. The capacity Eq\. \([61](https://arxiv.org/html/2609.10976#S4.E61)\) is the threshold for retrieving a typical pattern\. Requiring that alleα​Nve^\{\\alpha N\_\{v\}\}patterns be retrievable simultaneously is a stricter demand, met only at lower loads\[[15](https://arxiv.org/html/2609.10976#bib.bib15)\]\. The value of the ceiling is also specific to the Gaussian ensemble, which enters through the counting ratex2/2x^\{2\}/2: for patterns drawn on the sphere the ceiling disappears and the capacity continues to grow logarithmically inλ\\lambda\[[15](https://arxiv.org/html/2609.10976#bib.bib15)\], while for binary patterns exponential capacities were established rigorously in\[[11](https://arxiv.org/html/2609.10976#bib.bib11)\], and the capacity has been computed for more general ensembles, including patterns drawn from a hidden manifold\[[26](https://arxiv.org/html/2609.10976#bib.bib26)\]\. This sensitivity is characteristic of the exponential regime: at polynomial load the crosstalk is governed by central\-limit statistics and is largely insensitive to the pattern ensemble \(Sec\.[III](https://arxiv.org/html/2609.10976#S3)\), whereas at exponential load it is governed by large deviations, which are not universal\.

Relation to the random energy model\.The leak sum in Eq\. \([59](https://arxiv.org/html/2609.10976#S4.E59)\) is precisely the partition function of a random energy model \(REM\) witheα​Nve^\{\\alpha N\_\{v\}\}independent Gaussian energy levels at inverse temperatureλ\\lambda\[[21](https://arxiv.org/html/2609.10976#bib.bib21),[27](https://arxiv.org/html/2609.10976#bib.bib27)\], and Eq\. \([60](https://arxiv.org/html/2609.10976#S4.E60)\) is the standard large\-deviation evaluation of its free energy\[[28](https://arxiv.org/html/2609.10976#bib.bib28)\]\. The two branches are the two phases of the REM: the entropy\-dominated phase, in which annealed and quenched averages agree, and the condensed phase below the freezing transitionλ=2​α\\lambda=\\sqrt\{2\\alpha\}, in which the measure concentrates on finitely many levels\. This identification makes the agreement with\[[15](https://arxiv.org/html/2609.10976#bib.bib15)\]structural rather than accidental: there, the zero\-temperature energy landscape is split into the signal of the retrieved pattern and a noise term recognized as the free energy of an auxiliary REM, and the resulting typical\-pattern capacityα1​\(λ\)\\alpha\_\{1\}\(\\lambda\)coincides exactly with Eq\. \([61](https://arxiv.org/html/2609.10976#S4.E61)\)\. The vertex condensation of the Gibbs measure and the landscape analysis are two routes to the same variational problem\.

The same extreme value statistics also fixes the status of the retrieval states in the equilibrium ensemble\. At exponential load there exist patterns of atypically large norm, up to‖ξμ∗\(h,v\)‖2=\(1\+εmax\)​Nv\\left\\lVert\\xi^\{\(h,v\)\}\_\{\\mu^\{\\ast\}\}\\right\\rVert^\{2\}=\(1\+\\varepsilon\_\{\\max\}\)N\_\{v\}withεmax\>0\\varepsilon\_\{\\max\}\>0determined by the large deviations of the norm, and the state retrieving such a pattern has energy density−\(1\+εmax\)/2\-\(1\+\\varepsilon\_\{\\max\}\)/2, strictly below the value−1/2\-1/2of typical retrieval\. Typical retrieval is therefore metastable at any exponential load, and Eq\. \([61](https://arxiv.org/html/2609.10976#S4.E61)\) is a spinodal, in the same sense as the capacities of Models A and C\. The equilibrium phase structure built on the condensed extreme patterns is the subject of Sec\.[IV\.2](https://arxiv.org/html/2609.10976#S4.SS2)\.

#### IV\.1\.1Copy representation

Theffrepresentation admits an equivalent discrete formulation, which both fixes its integration measure and clarifies its relation to the replica method\. Consider the adiabatic partition function, the visible integral of the effective energy Eq\. \([54](https://arxiv.org/html/2609.10976#S4.E54)\), at the discrete temperatures for whichn≔β/λn\\coloneq\\beta/\\lambdais a positive integer\. The logarithm in the exponent then exponentiates into thenn\-th power of∑μeλ​ξμ\(h,v\)⋅v\\sum\_\{\\mu\}e^\{\\lambda\\,\\xi^\{\(h,v\)\}\_\{\\mu\}\\cdot v\}, the power expands by the multinomial theorem into a sum overnn\-tuples of pattern indices, and the Gaussianvvintegral gives, up to an overall constant,

ZξB​\(β\)∝∑μ1,…,μn=1Nhexp⁡\{λ2​n​‖∑j=1nξμj\(h,v\)‖2\}\.\{Z\_\{\\xi\}^\{\\mathrm\{B\}\}\}\(\\beta\)\\propto\\sum\_\{\\mu\_\{1\},\\dots,\\mu\_\{n\}=1\}^\{N\_\{h\}\}\\exp\\left\\\{\\frac\{\\lambda\}\{2n\}\\left\\lVert\\,\\sum\_\{j=1\}^\{n\}\\xi^\{\(h,v\)\}\_\{\\mu\_\{j\}\}\\right\\rVert^\{2\}\\right\\\}\.\(62\)An integer\-power representation of this type was introduced in\[[14](https://arxiv.org/html/2609.10976#bib.bib14)\]\. The hidden\-sector fluctuations around the adiabatic saddle point renormalize the exponent asβ/λ→β/λ\+Nh/2\\beta/\\lambda\\to\\beta/\\lambda\+N\_\{h\}/2, together with a shift of the summed patterns, and we relegate these corrections to Appendix[C\.2](https://arxiv.org/html/2609.10976#A3.SS2)\.

The expansion Eq\. \([62](https://arxiv.org/html/2609.10976#S4.E62)\) describesnn“copies”, each selecting one stored pattern, which interact through the Gram matrix of the selected patterns\. Coarse\-graining a copy configuration by its empirical measuref^μ≔nμ/n\\hat\{f\}\_\{\\mu\}\\coloneq n\_\{\\mu\}/n, withnμn\_\{\\mu\}the number of copies selecting the patternμ\\mu, the number of configurations in a class is the multinomial coefficientn\!/∏μnμ\!n\!/\\prod\_\{\\mu\}n\_\{\\mu\}\!, and Stirling’s formula converts Eq\. \([62](https://arxiv.org/html/2609.10976#S4.E62)\) into

ZξB\(β\)∝∑f^exp\{−βλ∑μf^μlogf^μ\+β2f^⊤Gf^\}\{Z\_\{\\xi\}^\{\\mathrm\{B\}\}\}\(\\beta\)\\propto\\sum\_\{\\hat\{f\}\}\\exp\\left\\\{\-\\frac\{\\beta\}\{\\lambda\}\\sum\_\{\\mu\}\\hat\{f\}\_\{\\mu\}\\log\\hat\{f\}\_\{\\mu\}\+\\frac\{\\beta\}\{2\}\\,\\hat\{f\}^\{\\top\}G\\hat\{f\}\\right\\\}\(63\)up to subexponential factors, where the sum runs over the gridf^∈ΔNh∩\(ℤ≥0/n\)Nh\\hat\{f\}\\in\\Delta^\{N\_\{h\}\}\\cap\(\\mathbb\{Z\}\_\{\\geq 0\}/n\)^\{N\_\{h\}\}\. This is exactly the functionalΦ\\Phiof Eq\. \([56](https://arxiv.org/html/2609.10976#S4.E56)\): the copy representation is theffrepresentation with its measure made explicit as a counting measure\. This distinction is not innocuous at exponential load, where the simplex has dimensioneα​Nv−1e^\{\\alpha N\_\{v\}\}\-1and the volume factors of a continuum measure are a priori uncontrolled\. The counting measure supplied by the model itself carries no such factors, its only extensive entropy being the choice of the selected patterns\.

The physical picture behind Eq\. \([62](https://arxiv.org/html/2609.10976#S4.E62)\) is transparent\. Reversing the Gaussian integration shows that, given a copy configuration, the visible state is Gaussian with mean the centroid1n​∑jξμj\(h,v\)\\frac\{1\}\{n\}\\sum\_\{j\}\\xi^\{\(h,v\)\}\_\{\\mu\_\{j\}\}and variance1/β1/\\betaper component: thenncopies are parallel retrieval queries sharing the single contextvv, and the temperature enters only through the number of queries,n=β/λn=\\beta/\\lambda\. Retrieval is the fully aligned configuration in which all copies select the same pattern, and the elementary excitation above it is a single copy defecting to a competitor, which carries the attention quantum1/n=λ​T1/n=\\lambda T\. AsT→0T\\to 0the number of copies diverges, the aligned configurations reproduce the vertex condensation described above, and the stability of alignment against single\-copy defection reproduces the capacity Eq\. \([61](https://arxiv.org/html/2609.10976#S4.E61)\) \(see Appendix[C\.3](https://arxiv.org/html/2609.10976#A3.SS3)\)\.

The copy representation also locates the present analysis relative to the replica method: in the replica approach the patterns are averaged first, at a formal number of replicas continued to zero, whereas here the thermal variablevvis integrated out first, at a physical, integer number of copies, so the thermal and quenched averages exchange roles\. The effectivelogdet\\log\\detinteraction that the pattern average generates among replicas reappears here as the counting entropy of the pattern choices, its Legendre transform \(Appendix[C\.2](https://arxiv.org/html/2609.10976#A3.SS2)\)\. Likewise, the symmetric structure of the order parameter, an ansatz in the replica computation, arises here as the exact maximizer within each class of copy configurations, and the frozen branch of Eq\. \([60](https://arxiv.org/html/2609.10976#S4.E60)\) plays the role that one\-step RSB \(1RSB\) plays in the REM\. The construction realizes the clone method of Monasson\[[22](https://arxiv.org/html/2609.10976#bib.bib22)\], with the number of clones set by the physical temperature rather than introduced as an auxiliary parameter\.

### IV\.2Phase diagram at finite temperature

The copy representation reduces the finite\-temperature problem to a combinatorial one\. The temperature enters only through the copy numbern=β/λn=\\beta/\\lambda, and a configuration is specified by the partition of thenncopies into groups occupying distinct patterns\. Since the weight in Eq\. \([62](https://arxiv.org/html/2609.10976#S4.E62)\) depends on the occupied patterns only through their Gram matrix, the disorder average reduces to a counting problem: the number of choices ofMMdistinct patterns with prescribed normalized Gram matrixQj​l=Gμj​μl/NvQ\_\{jl\}=G\_\{\\mu\_\{j\}\\mu\_\{l\}\}/N\_\{v\}iseNv​\(M​α−IM​\(Q\)\)e^\{N\_\{v\}\(M\\alpha\-I\_\{M\}\(Q\)\)\}to leading exponential order, where

IM\(Q\)=12\[TrQ−logdetQ−M\]I\_\{M\}\(Q\)=\\frac\{1\}\{2\}\\left\[\\Tr Q\-\\log\\det Q\-M\\right\]\(64\)is the large\-deviation rate function of the Gram matrix, the relative entropy between the centered Gaussian ensembles of covarianceQQand of unit covariance\[[28](https://arxiv.org/html/2609.10976#bib.bib28)\]\. ForM​α<IM​\(Q\)M\\alpha<I\_\{M\}\(Q\)no such choice exists with high probability\. Maximizing the counting factor times the weight under this existence constraint extends the two branches of Eq\. \([60](https://arxiv.org/html/2609.10976#S4.E60)\) to every configuration class: an annealed branch, attained at the typical Gram matrix of an exponentially tilted ensemble, and a frozen branch, pinned to the most extreme configurations actually present\. The optimization over the partition classes is elementary and is carried out in Appendix[C\.3](https://arxiv.org/html/2609.10976#A3.SS3)\. Here we summarize the results\.

![Refer to caption](https://arxiv.org/html/2609.10976v1/fig/modelB_phase_diagram_lam05_lam15.png)Figure 4:Finite\-temperature phase diagram of Model B in the\(α,T\)\(\\alpha,T\)plane at \(a\)λ=0\.5\\lambda=0\.5and \(b\)λ=1\.5\\lambda=1\.5\. Black solid lines are the first\-order boundaries between the paramagnetic phase \(P\) and the condensed sector, obtained by equating the free energies Eq\. \([65](https://arxiv.org/html/2609.10976#S4.E65)\)\. Forλ≥1\\lambda\\geq 1the paramagnetic phase is absent, and in \(a\) the P–F boundary meetsT=0T=0atαth​\(0\)\\alpha\_\{\\mathrm\{th\}\}\(0\)determined by2​α/λ=1\+εmax​\(α\)2\\alpha/\\lambda=1\+\\varepsilon\_\{\\max\}\(\\alpha\), see Appendix[C\.1](https://arxiv.org/html/2609.10976#A3.SS1)\. The dotted line is the freezing line Eq\. \([66](https://arxiv.org/html/2609.10976#S4.E66)\), the continuous transition on which the entropy of the condensed phase \(C\) vanishes and the measure freezes onto the maximum\-norm patterns \(F\)\. It is independent ofλ\\lambdaand hence identical in the two panels\. The thick red line is the retrieval spinodal, evaluated numerically from the one\-quantum criterion of Appendix[C\.4](https://arxiv.org/html/2609.10976#A3.SS4), whose annealed and frozen branches are Eq\. \([68](https://arxiv.org/html/2609.10976#S4.E68)\) and Eq\. \([69](https://arxiv.org/html/2609.10976#S4.E69)\)\. The dashed line shows its closed\-form asymptotics, and the dash\-dotted line is the switch lineI∗​\(T\)I^\{\\ast\}\(T\)between the annealed and extreme\-value evaluations of the defection destination\. The hatched region is the metastable retrieval region, which terminates atT=1/λT=1/\\lambda, where the attention quantumλ​T\\lambda Treaches the full attention weight \(n=1n=1\)\.Equilibrium phases\.Only the two extreme partitions survive the optimization, all copies on separate patterns or all copies on a single pattern\. Measuring the free energy per visible neuron relative to the Gaussian reference∫dve−β‖v‖2/2\\int dv\\,e^\{\-\\beta\\left\\lVert v\\right\\rVert^\{2\}/2\}, the resulting branches are

fP\\displaystyle f\_\{\\mathrm\{P\}\}=−αλ\+T2​log⁡\(1−λ\),\\displaystyle=\-\\frac\{\\alpha\}\{\\lambda\}\+\\frac\{T\}\{2\}\\log\(1\-\\lambda\),\(65\)fC\\displaystyle f\_\{\\mathrm\{C\}\}=−T​α\+T2​log⁡\(1−β\),\\displaystyle=\-T\\alpha\+\\frac\{T\}\{2\}\\log\(1\-\\beta\),fF\\displaystyle f\_\{\\mathrm\{F\}\}=−1\+εmax​\(α\)2,\\displaystyle=\-\\frac\{1\+\\varepsilon\_\{\\max\}\(\\alpha\)\}\{2\},and the equilibrium phase at given\(α,λ,T\)\(\\alpha,\\lambda,T\)is the branch of lowest free energy\. In the paramagnetic phase \(P\) each copy selects its own pattern and contributes the selection entropyα\\alpha\. The visible state is a weak thermal condensate,⟨‖v‖2⟩/Nv=T/\(1−λ\)\\left\\langle\\left\\lVert v\\right\\rVert^\{2\}\\right\\rangle/N\_\{v\}=T/\(1\-\\lambda\), which vanishes asT→0T\\to 0\. The selection entropy itself, however, does not vanish: it enters in proportion to the copy numbern=β/λn=\\beta/\\lambdaand leaves the temperature\-independent term−α/λ\-\\alpha/\\lambdainfPf\_\{\\mathrm\{P\}\}, so that the paramagnet survives down toT=0T=0and remains the equilibrium phase at loads beyond the zero\-temperature interceptαth​\(0\)\\alpha\_\{\\mathrm\{th\}\}\(0\)of the P–F boundary in Fig\.[4](https://arxiv.org/html/2609.10976#S4.F4)\(a\): at exponential load, the entropy of pattern selection acts as an energy\. In the condensed phase \(C\) allnncopies align on a single pattern, thermally tilted toward the norm\(1−β\)−1​Nv\(1\-\\beta\)^\{\-1\}N\_\{v\}, and the measure is an ergodic mixture over the exponentially many retrieval lumps of Sec\.[IV\.1](https://arxiv.org/html/2609.10976#S4.SS1), with entropy densityα−κ⁡\(β\)\\alpha\-\\kappa\(\\beta\), whereκ⁡\(β\)≔I1​\(\(1−β\)−1\)=12​\[β/\(1−β\)\+log⁡\(1−β\)\]\\kappa\(\\beta\)\\coloneq I\_\{1\}\(\(1\-\\beta\)^\{\-1\}\)=\\frac\{1\}\{2\}\[\\beta/\(1\-\\beta\)\+\\log\(1\-\\beta\)\]is the counting cost of the tilted norm\. This entropy vanishes on the freezing line

Tf​\(α\)=1\+1εmax​\(α\),T\_\{f\}\(\\alpha\)=1\+\\frac\{1\}\{\\varepsilon\_\{\\max\}\(\\alpha\)\},\(66\)whereεmax​\(α\)\\varepsilon\_\{\\max\}\(\\alpha\), defined byI1​\(1\+εmax\)=αI\_\{1\}\(1\+\\varepsilon\_\{\\max\}\)=\\alpha, makes precise the maximal norm excess introduced in Sec\.[IV\.1](https://arxiv.org/html/2609.10976#S4.SS1)\. BelowTfT\_\{f\}the mixture freezes onto theO⁡\(1\)O\(1\)patterns of maximal norm: this frozen phase \(F\) is the retrieval state of the maximum\-norm pattern, its free energy is temperature independent, and the transition atTfT\_\{f\}is the continuous freezing transition of the REM\[[21](https://arxiv.org/html/2609.10976#bib.bib21),[27](https://arxiv.org/html/2609.10976#bib.bib27)\]\. The existence conditions of the two annealed branches,λ<1\\lambda<1for P andT\>1T\>1for C, are two instances of a single criterion: a state with attention spread overn/sn/spatterns has inverse participation ratio‖f‖2=s/n\\left\\lVert f\\right\\rVert^\{2\}=s/n, and its visible Gaussian fluctuations are stable only whileβ​‖f‖2<1\\beta\\left\\lVert f\\right\\rVert^\{2\}<1, the instability atβ​‖f‖2=1\\beta\\left\\lVert f\\right\\rVert^\{2\}=1being the divergence of the visible integral in theffrepresentation Eq\. \([56](https://arxiv.org/html/2609.10976#S4.E56)\)\. In particular, forλ≥1\\lambda\\geq 1the paramagnetic phase is absent altogether\. The first\-order boundaries between P and the condensed sector follow by equating the branches of Eq\. \([65](https://arxiv.org/html/2609.10976#S4.E65)\), and the resulting phase diagram is shown in Fig\.[4](https://arxiv.org/html/2609.10976#S4.F4)\. Forλ<1\\lambda<1the paramagnetic and the two condensed phases meet at a triple point, located at\(α,T\)≃\(0\.66,1\.38\)\(\\alpha,T\)\\simeq\(0\.66,1\.38\)forλ=0\.5\\lambda=0\.5\.

Retrieval branch\.Typical retrieval appears in this landscape as a constrained branch\. Fixing the attention weightp=f^1p=\\hat\{f\}\_\{1\}condensed on a typical pattern as a reaction coordinate, and optimizing over the destinations and the Gram matrix of then⁡\(1−p\)n\(1\-p\)remaining copies, yields the Landau function

fR​\(p\)=−αλ​\(1−p\)\+T2​log⁡A−p22​A,f\_\{\\mathrm\{R\}\}\(p\)=\-\\frac\{\\alpha\}\{\\lambda\}\(1\-p\)\+\\frac\{T\}\{2\}\\log A\-\\frac\{p^\{2\}\}\{2A\},\(67\)whereA≔1−λ⁡\(1−p\)A\\coloneq 1\-\\lambda\(1\-p\)\. The three terms are the selection entropy of the leaked attention, the determinant of the visible fluctuations softened by the leak, and the signal energy\. The leaked copies respond linearly to the retrieval field and align weakly with the retrieved pattern, which amplifies the visible overlap tom=p/A≥pm=p/A\\geq p\. The branch interpolates between the paramagnet,fR​\(0\)=fPf\_\{\\mathrm\{R\}\}\(0\)=f\_\{\\mathrm\{P\}\}, and pure retrieval,fR​\(1\)=−12f\_\{\\mathrm\{R\}\}\(1\)=\-\\frac\{1\}\{2\}, the zero\-temperature retrieval energy density of Sec\.[IV\.1](https://arxiv.org/html/2609.10976#S4.SS1)\. Its interior stationary point is a maximum: a barrier between the two endpoints, not a phase\. Interior attention distributions therefore never become states at any temperature, which is the finite\-temperature form of vertex condensation\. The construction is a Franz–Parisi potential\[[29](https://arxiv.org/html/2609.10976#bib.bib29)\]whose reference configuration is the stored pattern itself, and the retrieval state atp=1p=1is a metastable state in the restricted\-ensemble sense\[[30](https://arxiv.org/html/2609.10976#bib.bib30)\]\.

Finite\-temperature capacity\.The stability of the retrieval endpoint is not governed by the slope of Eq\. \([67](https://arxiv.org/html/2609.10976#S4.E67)\) atp=1p=1\. The attention weight moves on the latticep=1−k​λ​Tp=1\-k\\lambda Twithk∈ℤ≥0k\\in\\mathbb\{Z\}\_\{\\geq 0\}: the elementary excitation is the defection of a single copy, which carries the attention quantumλ​T\\lambda T, and retrieval survives as long as a single defection raises the free energy\. Optimizing the destination of the defecting copy over the patterns present in the ensemble yields the finite\-temperature spinodal

αc​\(λ,T\)=λ⁡\(2−λ−λ​T\)2​\(1−λ2​T\)\+12​log⁡\(1−λ2​T\)\\alpha\_\{c\}\(\\lambda,T\)=\\frac\{\\lambda\(2\-\\lambda\-\\lambda T\)\}\{2\(1\-\\lambda^\{2\}T\)\}\+\\frac\{1\}\{2\}\\log\\left\(1\-\\lambda^\{2\}T\\right\)\(68\)in the annealed regime, in which destinations of the optimal overlap and norm exist in exponential number\. For sharp attention,λ2≳2​α\\lambda^\{2\}\\gtrsim 2\\alpha, they do not, and the defection freezes onto the single most aligned pattern present, as in the frozen branch of Eq\. \([60](https://arxiv.org/html/2609.10976#S4.E60)\)\. Two effects of orderTTthen favor the defection: the condensate left behind is thinner by one copy, and an atypically large norm of the destination pattern becomes an energy gain in its own right\. The zero\-temperature ceiling accordingly bends downward,

αceil​\(λ,T\)=12​\(1−λ​T2\)2\+O⁡\(T2\),\\alpha\_\{\\mathrm\{ceil\}\}\(\\lambda,T\)=\\frac\{1\}\{2\}\\left\(1\-\\frac\{\\lambda T\}\{2\}\\right\)^\{2\}\+O\(T^\{2\}\),\(69\)and the finite\-temperature capacity follows Eq\. \([68](https://arxiv.org/html/2609.10976#S4.E68)\) in the annealed regime and Eq\. \([69](https://arxiv.org/html/2609.10976#S4.E69)\) in the frozen regime, with the branch switch atλ2≈2​α\\lambda^\{2\}\\approx 2\\alpha\. AsT→0T\\to 0the two branches reduce to the zero\-temperature capacity Eq\. \([61](https://arxiv.org/html/2609.10976#S4.E61)\)\. In the opposite direction, the metastable region terminates atT=1/λT=1/\\lambda, where the attention quantumλ​T\\lambda Treaches the full attention weight and a single defection already empties the condensate \(see Fig\.[4](https://arxiv.org/html/2609.10976#S4.F4)\)\. The approach to this endpoint is controlled by the frozen branch: at smallα\\alphadestinations of the optimal overlap cease to exist, so the annealed expression Eq\. \([68](https://arxiv.org/html/2609.10976#S4.E68)\), whose zero lies below1/λ1/\\lambda, underestimates the stability of retrieval, and the defection onto the most favorable pattern actually present sustains a narrow metastable strip up toT=1/λT=1/\\lambda\. A continuum treatment ofppwould overestimate the capacity at anyT\>0T\>0: it resolves barriers thinner than one attention quantum and counts the selection entropy per infinitesimal attention weight rather than per copy\. The thermal destabilization of retrieval thus proceeds by discrete reassignments of attention, not by a smooth erosion of the overlap\.

The results above are derived at the integer temperature pointsn∈ℤ\>0n\\in\\mathbb\{Z\}\_\{\>0\}, but they are not tied to this lattice: replacing the multinomial expansion of Sec\.[IV\.1\.1](https://arxiv.org/html/2609.10976#S4.SS1.SSS1)by the generalized binomial expansion of the same integrand extends the sector decomposition, the attention quantumλ​T\\lambda T, and all the formulas above unchanged to arbitrary realβ\\beta\. The complete real\-temperature construction of the phase diagram is presented in Appendix[C\.4](https://arxiv.org/html/2609.10976#A3.SS4)\.

The phase diagram fixes the status of retrieval at exponential load\. Sinceεmax\>0\\varepsilon\_\{\\max\}\>0at anyα\>0\\alpha\>0, the frozen phase lies below pure retrieval,fF<fR​\(1\)=−12f\_\{\\mathrm\{F\}\}<f\_\{\\mathrm\{R\}\}\(1\)=\-\\frac\{1\}\{2\}, at all temperatures: for Gaussian patterns, typical retrieval never becomes the equilibrium phase, and the associative memory operates throughout the hatched region of Fig\.[4](https://arxiv.org/html/2609.10976#S4.F4)as a metastable state, with an escape time exponentially large inNvN\_\{v\}, governed in the Arrhenius sense by the one\-quantum barrier of Appendix[C\.4](https://arxiv.org/html/2609.10976#A3.SS4), mapped over the metastable region in Fig\.[5](https://arxiv.org/html/2609.10976#A3.F5)\. This sharpens the zero\-temperature statement of Sec\.[IV\.1](https://arxiv.org/html/2609.10976#S4.SS1)and contrasts with Models A and C in two respects\. First, the mechanism of thermal destabilization is different: there the retrieval overlap erodes smoothly through the equations of state, while here the temperature acts solely through the number of copies, and retrieval is undone by quantized reassignments of attention\. Second, the scale is different: the retrieval boundaries of Models A and C steepen factorially withkk\(Table[1](https://arxiv.org/html/2609.10976#S3.T1)and Table[2](https://arxiv.org/html/2609.10976#S3.T2)\), whereas the model\-B retrieval region retains an extent of order unity in bothα\\alphaandTT\. Finally, the equilibrium dominance of the frozen phase is a norm\-fluctuation effect of the Gaussian ensemble, driven by the same large deviations that set the capacity ceiling: for patterns drawn on the sphere,εmax≡0\\varepsilon\_\{\\max\}\\equiv 0, the frozen phase loses its advantage, and typical retrieval can compete as a genuine equilibrium phase\. The ensemble sensitivity noted below Eq\. \([61](https://arxiv.org/html/2609.10976#S4.E61)\) thus extends from the capacity to the entire equilibrium structure\.

## VDiscussion

We have investigated the statistical mechanics of the class\-ℋ\\mathcal\{H\}associative memories\[[19](https://arxiv.org/html/2609.10976#bib.bib19)\]through the hidden neurons, the order parameter of memory retrieval\. For Models A and C \(Sec\.[III](https://arxiv.org/html/2609.10976#S3)\), the replica method at polynomial load yields the RS phase diagrams and closed\-form capacities\. Since the crosstalk moment is universal, the model dependence, foremost the zero capacity of the quadratic spherical model, rests on the visible entropy, while the factorial growth of the crosstalk variance collapses both capacities at largekk\. For Model B \(Sec\.[IV](https://arxiv.org/html/2609.10976#S4)\), the hidden sector reduces to the attention weights, the load is exponential, and the copy representation maps the thermodynamics onto REM counting, with paramagnetic, condensed, and frozen phases and typical retrieval metastable, undone thermally by quantized reassignments of attention\. The two regimes differ in their crosstalk statistics, central\-limit and largely insensitive to the pattern ensemble at polynomial load, large\-deviation and ensemble\-dependent at exponential load\. Retrieval thus separates into the two roles carried by the Lagrangians of the classℋ\\mathcal\{H\}: the visible Lagrangian fixes, through the visible entropy, the stability of retrieval under a given crosstalk, while the hidden Lagrangian fixes the storage scale and with it the character of the crosstalk statistics\. The models analyzed here are positions on these two axes rather than separate theories\.

Two limitations deserve mention\. First, all results for Models A and C are derived within the RS ansatz, the capacities being the spinodals of the RS free energy\. The de Almeida–Thouless \(AT\) stability of the RS saddle points\[[23](https://arxiv.org/html/2609.10976#bib.bib23)\]has not been examined, and the spin\-glass phases and the low\-temperature boundaries acquire corrections from RSB, known to be small for the retrieval boundaries atk=2k=2\[[4](https://arxiv.org/html/2609.10976#bib.bib4),[24](https://arxiv.org/html/2609.10976#bib.bib24)\]and expected to remain so atk\>2k\>2\[[7](https://arxiv.org/html/2609.10976#bib.bib7)\]\. Second, fork\>2k\>2and for Model B the hidden sector is treated in the adiabatic limitβ/τh→∞\\beta/\{\\tau\_\{h\}\}\\to\\infty, with the hidden neurons enslaved to the visible configuration\. At finiteβ/τh\\beta/\{\\tau\_\{h\}\}the classℋ\\mathcal\{H\}is a genuinely two\-temperature system, whose phase diagrams away from the adiabatic limit we do not address\. The equal\-temperature end of this interpolation is the bipartite Gibbs measure of a restricted Boltzmann machine, whose equivalence with the Hopfield model\[[31](https://arxiv.org/html/2609.10976#bib.bib31)\]and whose phase diagram for general hidden priors\[[32](https://arxiv.org/html/2609.10976#bib.bib32)\]are known\.

These limitations mark the first direction for future work, completing the phase diagrams under RSB\. For Models A and C this amounts to locating the AT lines and computing the one\-step corrections to the spin\-glass phase and the retrieval boundaries, for which the techniques developed for the Hopfield model\[[4](https://arxiv.org/html/2609.10976#bib.bib4),[24](https://arxiv.org/html/2609.10976#bib.bib24)\], dense networks\[[33](https://arxiv.org/html/2609.10976#bib.bib33),[34](https://arxiv.org/html/2609.10976#bib.bib34)\], and bipartite spin glasses\[[35](https://arxiv.org/html/2609.10976#bib.bib35)\]apply directly\. For Model B the corresponding step is the stability of the sector decomposition against fluctuations between sectors, together with a test of the predicted Ruelle statistics of the attention weights in the frozen phase\[[36](https://arxiv.org/html/2609.10976#bib.bib36)\]\(Appendix[C\.4](https://arxiv.org/html/2609.10976#A3.SS4)\)\.

A second direction concerns the hidden\-sector corrections of Model B\. The Gaussian fluctuations of the hidden neurons around the adiabatic saddle point produce the exponent shiftβ/λ→β/λ\+Nh/2\\beta/\\lambda\\to\\beta/\\lambda\+N\_\{h\}/2and an additive shift of the summed patterns \(Appendix[C\.2](https://arxiv.org/html/2609.10976#A3.SS2)\), neither of which is small at exponential load\. Whether they can be absorbed into a redefinition of the reference measure, leaving the rate\-level phase diagram intact, is the main structural question left open by our analysis \(Appendix[C\.4](https://arxiv.org/html/2609.10976#A3.SS4)\)\. Since both originate from the zero mode of the softmax, their resolution would clarify how the normalization of the attention weights, shared by the transformer attention and the modern Hopfield network\[[12](https://arxiv.org/html/2609.10976#bib.bib12),[13](https://arxiv.org/html/2609.10976#bib.bib13),[14](https://arxiv.org/html/2609.10976#bib.bib14)\], affects the thermodynamics beyond the leading order\.

The third direction is more open\-ended\. Regarded as design variables rather than properties of given models, the two axes suggest constructing new Lagrangians, or selecting them with prescribed retrieval properties, in the spirit of recent programs on the design of energy\-based architectures\[[5](https://arxiv.org/html/2609.10976#bib.bib5),[37](https://arxiv.org/html/2609.10976#bib.bib37),[6](https://arxiv.org/html/2609.10976#bib.bib6)\]\. The family generating Models A and C \(Appendix[B\.1](https://arxiv.org/html/2609.10976#A2.SS1)\) is the smallest setting in which both axes can be varied, and the same question extends to the classℋk\\mathcal\{H\}\_\{k\}\(Appendix[A](https://arxiv.org/html/2609.10976#A1)\), where additional layers of hidden neurons implement the hierarchical associative memory\[[38](https://arxiv.org/html/2609.10976#bib.bib38),[39](https://arxiv.org/html/2609.10976#bib.bib39)\]\. Identifying the order parameters of the deeper hidden layers, and the phases they support, would extend the present analysis to genuinely hierarchical models, for which exponential capacities from distributed hidden representations\[[40](https://arxiv.org/html/2609.10976#bib.bib40)\]and emergent computations in assemblies of networks\[[41](https://arxiv.org/html/2609.10976#bib.bib41)\]have recently been reported\.

## Appendix AOverview of the class\-ℋk\\mathcal\{H\}\_\{k\}associative memories

The class\-ℋk\\mathcal\{H\}\_\{k\}associative memories are a slight modification of Krotov’s hierarchical associative memory\[[38](https://arxiv.org/html/2609.10976#bib.bib38)\],555ℋk\\mathcal\{H\}\_\{k\}stands for “Hierarchical Associative Memory with𝒌\\bm\{k\}layers” \(HAMk, cf\.\[[42](https://arxiv.org/html/2609.10976#bib.bib42)\]\), namely the class of associative memories withkkdefining Lagrangians\. One can also think of it as an extendedKrotov–Hopfield model\[[19](https://arxiv.org/html/2609.10976#bib.bib19)\], which in this paper we call the classℋ\\mathcal\{H\}, a special case ofℋk\\mathcal\{H\}\_\{k\}withk=2k=2\.the modification being that the energy function carries explicit coupling constantsλA\+1,A\\lambda\_\{A\+1,A\}and layer\-wise weights1/τA1/\\tau\_\{A\}\. The classℋk\\mathcal\{H\}\_\{k\}consists ofk\(≥2\)k~\(\\geq 2\)layers, withNAN\_\{A\}neurons in layerAA\(A=1,…,kA=1,\\dots,k\)\. This system is a generalization of the model discussed in\[[19](https://arxiv.org/html/2609.10976#bib.bib19)\]and is specified by the following components: the neuron statesxA​\(t\)∈ℝNAx^\{A\}\(t\)\\in\\mathbb\{R\}^\{N\_\{A\}\}in each layer, interaction matrices between the adjacent layers,ξ\(A\+1,A\)∈ℝNA\+1×NA\\xi^\{\(A\+1,A\)\}\\in\\mathbb\{R\}^\{N\_\{A\+1\}\\times N\_\{A\}\}, and activation functions

gA:ℝNA→ℝNA,A=1,…,k,g^\{A\}:\\mathbb\{R\}^\{N\_\{A\}\}\\to\\mathbb\{R\}^\{N\_\{A\}\},\\quad A=1,\\dots,k,\(70\)determined through LagrangiansLA:ℝNA→ℝL^\{A\}:\\mathbb\{R\}^\{N\_\{A\}\}\\to\\mathbb\{R\}such thatgA=∇LAg^\{A\}=\\nabla L^\{A\}\. Note that the interaction matrices are required to be “symmetric”:ξ\(A,A\+1\)≔\(ξ\(A\+1,A\)\)⊤\\xi^\{\(A,A\+1\)\}\\coloneq\(\\xi^\{\(A\+1,A\)\}\)^\{\\top\}\. \(For more details, see\[[38](https://arxiv.org/html/2609.10976#bib.bib38), Sec\. 3\]\.\)

Let us now define an energy function onT∗​\(ℝ∑A=1kNA\)≃∏A=1k\(ℝNA×ℝNA\)T^\{\\ast\}\(\\mathbb\{R\}^\{\\sum\_\{A=1\}^\{k\}N\_\{A\}\}\)\\simeq\\prod\_\{A=1\}^\{k\}\(\\mathbb\{R\}^\{N\_\{A\}\}\\times\\mathbb\{R\}^\{N\_\{A\}\}\)by

E⁡\(\{yA\},\{xA\}\)=∑A=1k1τA​\(\(xA\)⊤​yA−LA​\(xA\)\)−∑A=1k−1λA\+1,AτA\+1​τA​\(yA\+1\)⊤​ξ\(A\+1,A\)​yA,E\(\\left\\\{y^\{A\}\\right\\\},\\left\\\{x^\{A\}\\right\\\}\)=\\sum\_\{A=1\}^\{k\}\\frac\{1\}\{\\tau\_\{A\}\}\\left\(\(x^\{A\}\)^\{\\top\}y^\{A\}\-L^\{A\}\(x^\{A\}\)\\right\)\-\\sum\_\{A=1\}^\{k\-1\}\\frac\{\\lambda\_\{A\+1,A\}\}\{\\tau\_\{A\+1\}\\tau\_\{A\}\}\(y^\{A\+1\}\)^\{\\top\}\\xi^\{\(A\+1,A\)\}y^\{A\},\(71\)whereλA\+1,A\\lambda\_\{A\+1,A\}are coupling constants between the\(A\+1\)\(A\+1\)\- andAA\-th layers and are assumed to be symmetric,λA\+1,A=λA,A\+1\\lambda\_\{A\+1,A\}=\\lambda\_\{A,A\+1\}\. This energy function may be regarded as a Hamiltonian function on the phase space, with the variablesyAy^\{A\}playing the role of the positions and the neuron statesxAx^\{A\}that of the conjugate momenta\. The first sum inEEis then the Legendre transform of the Lagrangians\. The physical configurations lie on the Legendre constraint surfaceΣ≔\{yA=gA\(xA\)\}A=1k\\Sigma\\coloneq\\left\\\{y^\{A\}=g^\{A\}\(x^\{A\}\)\\right\\\}\_\{A=1\}^\{k\}, which is precisely the locus where the gradients∇xAE=\(yA−gA​\(xA\)\)/τA\\nabla\_\{x^\{A\}\}E=\(y^\{A\}\-g^\{A\}\(x^\{A\}\)\)/\\tau\_\{A\}vanish, namely, where the position sector of Hamilton’s canonical equations,d​yA/d​t=∇xAEdy^\{A\}/dt=\\nabla\_\{x^\{A\}\}E, becomes stationary\. We then define the dynamical equations of the system as the momentum sector of the canonical equations restricted to the constraint surfaceΣ\\Sigma, with the conventionλ1,0≡0\\lambda\_\{1,0\}\\equiv 0andλk\+1,k≡0\\lambda\_\{k\+1,k\}\\equiv 0:

d​xA​\(t\)d​t\\displaystyle\\frac\{dx^\{A\}\(t\)\}\{dt\}≔−∇yAE\(\{yA\},\{xA\}\)\|\{\(yA,xA\)=\(gA\(xA\(t\)\),xA\(t\)\)\}A=1k\\displaystyle\\coloneq\-\\nabla\_\{y^\{A\}\}E\(\\left\\\{y^\{A\}\\right\\\},\\left\\\{x^\{A\}\\right\\\}\)\\big\|\_\{\\left\\\{\(y^\{A\},x^\{A\}\)=\(g^\{A\}\(x^\{A\}\(t\)\),x^\{A\}\(t\)\)\\right\\\}\_\{A=1\}^\{k\}\}\(72\)=1τA​\(λA,A−1τA−1​ξ\(A,A−1\)​gA−1​\(xA−1​\(t\)\)\+λA,A\+1τA\+1​ξ\(A,A\+1\)​gA\+1​\(xA\+1​\(t\)\)−xA​\(t\)\),\\displaystyle=\\frac\{1\}\{\\tau\_\{A\}\}\\left\(\\frac\{\\lambda\_\{A,A\-1\}\}\{\\tau\_\{A\-1\}\}\\xi^\{\(A,A\-1\)\}g^\{A\-1\}\\left\(x^\{A\-1\}\(t\)\\right\)\+\\frac\{\\lambda\_\{A,A\+1\}\}\{\\tau\_\{A\+1\}\}\\xi^\{\(A,A\+1\)\}g^\{A\+1\}\\left\(x^\{A\+1\}\(t\)\\right\)\-x^\{A\}\(t\)\\right\),\(73\)where∇yA\\nabla\_\{y^\{A\}\}are the gradient operators with respect toyAy^\{A\}, and we used the fact that\(ξ\(A\+1,A\)\)⊤=ξ\(A,A\+1\)\(\\xi^\{\(A\+1,A\)\}\)^\{\\top\}=\\xi^\{\(A,A\+1\)\}\. The energy function of the classℋk\\mathcal\{H\}\_\{k\}, which serves as a Lyapunov function of the dynamical system, is defined by

Eℋk​\(\{xA\}\)\\displaystyle E\_\{\\mathcal\{H\}\_\{k\}\}\(\\left\\\{x^\{A\}\\right\\\}\)≔E\(\{yA\},\{xA\}\)\|\{yA=gA\(xA\)\}A=1k\\displaystyle\\coloneq E\(\\left\\\{y^\{A\}\\right\\\},\\left\\\{x^\{A\}\\right\\\}\)\\big\|\_\{\\left\\\{y^\{A\}=g^\{A\}\(x^\{A\}\)\\right\\\}\_\{A=1\}^\{k\}\}\(74\)=∑A=1k1τA​\(\(xA\)⊤​gA​\(xA\)−LA​\(xA\)\)−∑A=1k−1λA\+1,AτA\+1​τA​\(gA\+1​\(xA\+1\)\)⊤​ξ\(A\+1,A\)​gA​\(xA\)\.\\displaystyle=\\sum\_\{A=1\}^\{k\}\\frac\{1\}\{\\tau\_\{A\}\}\\left\(\(x^\{A\}\)^\{\\top\}g^\{A\}\(x^\{A\}\)\-L^\{A\}\(x^\{A\}\)\\right\)\-\\sum\_\{A=1\}^\{k\-1\}\\frac\{\\lambda\_\{A\+1,A\}\}\{\\tau\_\{A\+1\}\\tau\_\{A\}\}\\big\(g^\{A\+1\}\(x^\{A\+1\}\)\\big\)^\{\\top\}\\xi^\{\(A\+1,A\)\}g^\{A\}\(x^\{A\}\)\.
Since the position sector of the canonical equations is replaced by the constraintyA=gA​\(xA\)y^\{A\}=g^\{A\}\(x^\{A\}\), the flow is not symplectic, and the energy is not conserved along the trajectory\. Instead, provided that the Hessians of the Lagrangians are positive \(semi\-\)definite, this energy function monotonically decreases along the solution trajectory of the dynamical equations,

d​Eℋk​\(\{xA​\(t\)\}\)d​t=−∑A=1k\(d​xA​\(t\)d​t\)⊤HessLA\(xA\(t\)\)d​xA​\(t\)d​t≤0,\\frac\{dE\_\{\\mathcal\{H\}\_\{k\}\}\(\\left\\\{x^\{A\}\(t\)\\right\\\}\)\}\{dt\}=\-\\sum\_\{A=1\}^\{k\}\\left\(\\frac\{dx^\{A\}\(t\)\}\{dt\}\\right\)^\{\\top\}\\operatorname\{Hess\}L^\{A\}\(x^\{A\}\(t\)\)\\,\\frac\{dx^\{A\}\(t\)\}\{dt\}\\leq 0,\(75\)which exhibits the dissipative nature of the system\. If, in addition, the overall energy function is bounded from below, the trajectory is guaranteed to converge to a fixed\-point attractor state, which corresponds to one of the local minima of the energy function\. Such fixed points may be identified with the stored memories, and the convergence toward them with memory retrieval\. Settingk=2k=2and identifyingx1=vx^\{1\}=vandx2=hx^\{2\}=hreproduces the system discussed in the main text\. For the two\-layer \(k=2k=2\) case, as in the main text, we drop the subscriptkkand refer to the model as the class\-ℋ\\mathcal\{H\}associative memories\.

## Appendix BDetails on Models A and C

### B\.1A family of associative memories that generates Models A and C

In this appendix we characterize the family of visible Lagrangians underlying the construction of Sec\.[III](https://arxiv.org/html/2609.10976#S3)and explain in what sense Models A and C are its distinguished members\. Throughout, the hidden sector is kept general, and it reenters only in the final remark\.

Euler’s identity and scale invariance\.In the energy function Eq\. \([4](https://arxiv.org/html/2609.10976#S2.E4)\), the visible neurons enter both through the Legendre termEv​\(v\)=v⊤​g​\(v\)−Lv​\(v\)\{E\_\{v\}\}\(v\)=v^\{\\top\}g\(v\)\-\{L\_\{v\}\}\(v\)and through the activationg⁡\(v\)g\(v\)in the interaction term\. The bare visible neurons drop out of the energy precisely when the Legendre term vanishes identically,

v⊤∇Lv\(v\)−Lv\(v\)=0,v^\{\\top\}\\nabla\{L\_\{v\}\}\(v\)\-\{L\_\{v\}\}\(v\)=0,\(76\)in which case the energy depends onvvonly throughg⁡\(v\)g\(v\), as observed in Sec\.[III](https://arxiv.org/html/2609.10976#S3)\. Equation \([76](https://arxiv.org/html/2609.10976#A2.E76)\) is Euler’s identity, and by Euler’s theorem on homogeneous functions it holds on an open cone on whichLv\{L\_\{v\}\}is differentiable if and only ifLv\{L\_\{v\}\}is positively homogeneous of degree one there\[[43](https://arxiv.org/html/2609.10976#bib.bib43)\],

Lv​\(t​v\)=t​Lv​\(v\),t\>0\.\{L\_\{v\}\}\(tv\)=t\\,\{L\_\{v\}\}\(v\),\\qquad t\>0\.\(77\)Differentiating Eq\. \([77](https://arxiv.org/html/2609.10976#A2.E77)\) with respect tovvshows that the activation is then homogeneous of degree zero,g⁡\(t​v\)=g⁡\(v\)g\(tv\)=g\(v\): it is a pure readout of the direction ofvv, insensitive to its scale\. This is the family\-wide origin of the scale invariance of the energy noted in Sec\.[III](https://arxiv.org/html/2609.10976#S3)\.

Degree\-one homogeneity fixesLv\{L\_\{v\}\}along each ray from the origin once its value on a single cross\-section is given\. Choosing the Euclidean unit sphere as the cross\-section yields the representation

Lv​\(v\)=‖v‖​ϕ​\(v‖v‖\),ϕ:𝕊Nv−1→ℝ,\{L\_\{v\}\}\(v\)=\\left\\lVert v\\right\\rVert\\,\\phi\\left\(\\frac\{v\}\{\\left\\lVert v\\right\\rVert\}\\right\),\\qquad\\phi\\colon\\mathbb\{S\}^\{N\_\{v\}\-1\}\\to\\mathbb\{R\},\(78\)where‖⋅‖\\left\\lVert\\cdot\\right\\rVertis the Euclidean norm\. Conversely, every functionϕ\\phion the unit sphere defines a positively homogeneousLv\{L\_\{v\}\}through Eq\. \([78](https://arxiv.org/html/2609.10976#A2.E78)\)\. No differentiability is needed for the representation itself\. Smoothness enters only through the correspondence thatLv\{L\_\{v\}\}isC1C^\{1\}away from the origin exactly whenϕ∈C1​\(𝕊Nv−1\)\\phi\\in C^\{1\}\(\\mathbb\{S\}^\{N\_\{v\}\-1\}\), which is what defines the activationg=∇Lvg=\\nabla\{L\_\{v\}\}there\. In particular,ϕ\\phineed not be an elementary function, so the family is genuinely infinite\-dimensional\. In two dimensions, Eq\. \([78](https://arxiv.org/html/2609.10976#A2.E78)\) is simplyLv=r​ϕ​\(θ\)\{L\_\{v\}\}=r\\phi\(\\theta\)in polar coordinates, and the convexity requirement introduced next takes the classical formϕ⁡\(θ\)\+ϕ′′​\(θ\)≥0\\phi\(\\theta\)\+\\phi^\{\\prime\\prime\}\(\\theta\)\\geq 0of the support\-function condition\[[44](https://arxiv.org/html/2609.10976#bib.bib44)\]\.

Convexity and support functions\.Within the classℋ\\mathcal\{H\}, the Lagrangians are not arbitrary: the monotonic decrease of the energy \(the defining property of an associative memory\) requires positive semi\-definite Hessians, i\.e\., the Lagrangians must be convex, though not necessarily strictly\. For the nondifferentiable members below, convexity itself is the appropriate formulation of the same requirement\. A convex, everywhere finite, positively degree\-one homogeneous function is precisely a sublinear function, and every sublinear function is the support function of a uniquely determined compact convex setK⊂ℝNvK\\subset\\mathbb\{R\}^\{N\_\{v\}\}\[[45](https://arxiv.org/html/2609.10976#bib.bib45), Sec\. 13\],

Lv​\(v\)=hK​\(v\)≔maxu∈K⁡u⊤​v,K=∂Lv​\(0\),\{L\_\{v\}\}\(v\)=h\_\{K\}\(v\)\\coloneq\\max\_\{u\\in K\}u^\{\\top\}v,\\qquad K=\\partial\{L\_\{v\}\}\(0\),\(79\)where∂Lv​\(0\)\\partial\{L\_\{v\}\}\(0\)denotes the subdifferential at the origin\. The family underlying Models A and C is thus classified by a convex body: the choice ofKKdetermines the model\. Moreover, whereverLv\{L\_\{v\}\}is differentiable, the gradient is the unique point ofKKat which the linear functionu↦u⊤​vu\\mapsto u^\{\\top\}vattains its maximum\[[45](https://arxiv.org/html/2609.10976#bib.bib45), Sec\. 25\], so that

g⁡\(v\)=∇hK​\(v\)∈∂K\.g\(v\)=\\nabla h\_\{K\}\(v\)\\in\\partial K\.\(80\)The effective visible degree of freedom is therefore the activationu=g⁡\(v\)u=g\(v\), which lives on the boundary ofKK: the direction ofvvdetermines the point of∂K\\partial Kat which the hyperplane with normalvvsupportsKK\.

Since the energy is exactly constant along each open ray\{t​v:t\>0\}\\left\\\{tv:t\>0\\right\\\}, the radial direction is a flat direction \(an exact zero mode\) ofβ​Eξ\\beta E\_\{\\xi\}for every member of the family, and the naive visible integral in Eq\. \([7](https://arxiv.org/html/2609.10976#S2.E7)\) diverges in proportion to the infinite radial volume\. This divergence is an overall factor, independent of the patterns and of all order parameters, so it factors out of every observable\. Removing it in the standard collective\-coordinate manner leaves an integral over the space of rays\. Any cross\-section of the rays represents this quotient, and the natural choice is provided by the activation itself: the visible trace becomes an integral overu=g⁡\(v\)∈∂Ku=g\(v\)\\in\\partial K, rescaled so that the components areO⁡\(1\)O\(1\)as in the main text\. This is the general form of the prescription anticipated in Sec\.[II\.2](https://arxiv.org/html/2609.10976#S2.SS2)and adopted for the family in Eq\. \([14](https://arxiv.org/html/2609.10976#S3.E14)\)\. One point deserves emphasis: the quotient fixes the domain of the visible trace but not its measure, which must be supplied as part of the model definition\. For polytopes \(includingp=1,∞p=1,\\inftybelow\) the activation takes finitely many values and the counting measure is canonical, and forp=2p=2the rotation\-invariant measure on the sphere is singled out by symmetry\. For the other members inequivalent natural choices coexist \(e\.g\., the surface measure and the cone measure on the dual sphere differ by an explicit density\[[46](https://arxiv.org/html/2609.10976#bib.bib46)\]\), and in Eq\. \([14](https://arxiv.org/html/2609.10976#S3.E14)\) the standard surface measure is understood\.

Theℓp\\ell\_\{p\}family\.ForLv=‖v‖p\{L\_\{v\}\}=\\left\\lVert v\\right\\rVert\_\{p\}with1≤p≤∞1\\leq p\\leq\\infty, the maximum in Eq\. \([79](https://arxiv.org/html/2609.10976#A2.E79)\) is Hölder’s inequalityu⊤​v≤‖u‖p′​‖v‖pu^\{\\top\}v\\leq\\left\\lVert u\\right\\rVert\_\{p^\{\\prime\}\}\\left\\lVert v\\right\\rVert\_\{p\}together with its equality condition, attained precisely at the activation given in Sec\.[III](https://arxiv.org/html/2609.10976#S3), and the body is the unit ball of the dual norm,

K=Bp′≔\{u∈ℝNv\|‖u‖p′≤1\},1p\+1p′=1,K=B\_\{p^\{\\prime\}\}\\coloneq\\left\\\{u\\in\\mathbb\{R\}^\{N\_\{v\}\}\\,\\middle\|\\,\\left\\lVert u\\right\\rVert\_\{p^\{\\prime\}\}\\leq 1\\right\\\},\\qquad\\frac\{1\}\{p\}\+\\frac\{1\}\{p^\{\\prime\}\}=1,\(81\)whose boundary carries the constraint‖g‖p′=1\\left\\lVert g\\right\\rVert\_\{p^\{\\prime\}\}=1of Sec\.[III](https://arxiv.org/html/2609.10976#S3)\. Three cases deserve mention\. \(i\)p=2p=2:KKis the Euclidean ball, which is self\-dual\. The activationg⁡\(v\)=v/‖v‖g\(v\)=v/\\left\\lVert v\\right\\rVertsweeps the whole unit sphere, the visible degrees of freedom are spherical spins, and the member is Model C\. \(ii\)p=1p=1:K=\[−1,1\]NvK=\[\-1,1\]^\{N\_\{v\}\}is the hypercube, the dual ball ofℓ∞\\ell\_\{\\infty\}\. Off the coordinate hyperplanes, the activation isg⁡\(v\)=sgn⁡\(v\)g\(v\)=\\sgn\(v\)componentwise, whose image is the finite vertex set\{±1\}Nv\\left\\\{\\pm 1\\right\\\}^\{N\_\{v\}\}\. The visible degrees of freedom are Ising spins, and the member is Model A, a discrete spin model arising from continuous neurons through a piecewise linear Lagrangian\. \(iii\)p=∞p=\\infty:K=conv⁡\{±e1,…,±eNv\}K=\\mathrm\{conv\}\\left\\\{\\pm e\_\{1\},\\ldots,\\pm e\_\{N\_\{v\}\}\\right\\\}is the cross\-polytope, witheie\_\{i\}the standard basis vectors\. Wherever the largest component\|vi∗\|\\left\\lvert v\_\{i^\{\\ast\}\}\\right\\rvertis unique,g⁡\(v\)=sgn⁡\(vi∗\)​ei∗g\(v\)=\\sgn\(v\_\{i^\{\\ast\}\}\)\\,e\_\{i^\{\\ast\}\}, so the visible layer collapses to a single winner\-take\-all \(one\-hot\) degree of freedom with2​Nv2N\_\{v\}states\.

Cases \(ii\) and \(iii\) illustrate a general rule\. Abbreviating the local field acting on the visible layer in Eq\. \([14](https://arxiv.org/html/2609.10976#S3.E14)\) byci≔β​λτh​τv​∑μF′​\(hμ\)​ξμ​i\(h,v\)c\_\{i\}\\coloneq\\frac\{\\beta\\lambda\}\{\{\\tau\_\{h\}\}\{\\tau\_\{v\}\}\}\\sum\_\{\\mu\}F^\{\\prime\}\(h\_\{\\mu\}\)\\,\\xi^\{\(h,v\)\}\_\{\\mu i\}, the visible trace takes the form∫Sd​Ω​\(x\)​ec⊤​x\\int\_\{S\}d\\Omega\(\\mathrm\{x\}\)\\,e^\{c^\{\\top\}\\mathrm\{x\}\}: a Laplace\-type transform of the chosen measure on∂K\\partial K\. IfKKis a polytope with verticesa1,…,aMa\_\{1\},\\ldots,a\_\{M\}, thenhK​\(v\)=max1≤l≤M⁡al⊤​vh\_\{K\}\(v\)=\\max\_\{1\\leq l\\leq M\}a\_\{l\}^\{\\top\}v, the activation is piecewise constant with values in the vertex set, and the visible trace is a finite exponential sum,

∫Sd​Ω​\(x\)​ec⊤​x⟶∑l=1Mwl​ec⊤​al,\\int\_\{S\}d\\Omega\(\\mathrm\{x\}\)\\,e^\{c^\{\\top\}\\mathrm\{x\}\}\\;\\longrightarrow\\;\\sum\_\{l=1\}^\{M\}w\_\{l\}\\,e^\{c^\{\\top\}a\_\{l\}\},\(82\)with weightswlw\_\{l\}given by the counting measure and the vertices rescaled according to the radius convention of Eq\. \([14](https://arxiv.org/html/2609.10976#S3.E14)\)\. Discrete spin models thus arise within the classℋ\\mathcal\{H\}as the polyhedral shapesKK, with no need to postulate discrete variables at the outset\. For the hypercube the sum factorizes over the sites,∑x∈\{±1\}Nvec⊤​x=∏i2coshci\\sum\_\{\\mathrm\{x\}\\in\\left\\\{\\pm 1\\right\\\}^\{N\_\{v\}\}\}e^\{c^\{\\top\}\\mathrm\{x\}\}=\\prod\_\{i\}2\\cosh c\_\{i\}, which is the elementary identity behind Eq\. \([17](https://arxiv.org/html/2609.10976#S3.E17)\)\. For the cross\-polytope it is a single sum over sites,∑i\(eci\+e−ci\)\\sum\_\{i\}\(e^\{c\_\{i\}\}\+e^\{\-c\_\{i\}\}\)up to the radius normalization\. Forp=2p=2, on the other hand, the trace is rotation invariant and is evaluated exactly by Gaussian methods in Appendix[B\.3](https://arxiv.org/html/2609.10976#A2.SS3)\. Within theℓp\\ell\_\{p\}family,p=1p=1andp=2p=2are therefore the two members whose visible trace closes in elementary terms at everyNvN\_\{v\}\. This is the quantitative content of the footnote in Sec\.[III](https://arxiv.org/html/2609.10976#S3)and the reason the main text focuses on Models A and C\.

For the remaining members, general1<p<∞1<p<\\infty, no elementary closed form of the visible trace is known, but three standard representations organize the available results\. \(i\)*Probabilistic representation\.*IfY=\(Y1,…,YNv\)Y=\(Y\_\{1\},\\ldots,Y\_\{N\_\{v\}\}\)has i\.i\.d\. components with density proportional toe−\|t\|p′e^\{\-\\left\\lvert t\\right\\rvert^\{p^\{\\prime\}\}\}, thenY/‖Y‖p′Y/\\left\\lVert Y\\right\\rVert\_\{p^\{\\prime\}\}is distributed according to the cone measure on the unit dual sphere∂Bp′\\partial B\_\{p^\{\\prime\}\}and is independent of‖Y‖p′\\left\\lVert Y\\right\\rVert\_\{p^\{\\prime\}\}\[[47](https://arxiv.org/html/2609.10976#bib.bib47)\]:

u​=d​Y‖Y‖p′,Yi∼e−\|t\|p′2​Γ​\(1\+1/p′\)​d​ti\.i\.d\.,u\\overset\{\\mathrm\{d\}\}\{=\}\\frac\{Y\}\{\\left\\lVert Y\\right\\rVert\_\{p^\{\\prime\}\}\},\\qquad Y\_\{i\}\\sim\\frac\{e^\{\-\\left\\lvert t\\right\\rvert^\{p^\{\\prime\}\}\}\}\{2\\Gamma\(1\+1/p^\{\\prime\}\)\}\\,dt\\quad\\text\{i\.i\.d\.\},\(83\)and the same holds for the surface measure after inserting an explicit density\[[46](https://arxiv.org/html/2609.10976#bib.bib46)\]\. This trades the hard constraint forNvN\_\{v\}independent single\-site variables coupled only through the norm‖Y‖p′\\left\\lVert Y\\right\\rVert\_\{p^\{\\prime\}\}\(theℓp′\\ell\_\{p^\{\\prime\}\}analogue of representing the uniform measure on the sphere by a Gaussian vector conditioned on its radius\), and is the natural starting point for a saddle\-point evaluation of the trace at largeNvN\_\{v\}\. \(ii\)*Mellin–Barnes representation\.*Resolving the residual radial constraint in Eq\. \([83](https://arxiv.org/html/2609.10976#A2.E83)\) by a Mellin–Barnes contour integral expresses the trace as a single contour integral over products of Gamma functions, i\.e\., a representation of Fox\-HHtype, from which asymptotics can be extracted by shifting the contour\[[48](https://arxiv.org/html/2609.10976#bib.bib48)\]\. \(iii\)*Confluent\-hypergeometric special case\.*When the local field has a single non\-vanishing component,c\|eic\\parallel e\_\{i\}, the trace collapses to the one\-dimensional Beta\-family integral∫−11ec​t​\(1−\|t\|p′\)\(Nv−1\)/p′−1​𝑑t\\int\_\{\-1\}^\{1\}e^\{ct\}\(1\-\\left\\lvert t\\right\\rvert^\{p^\{\\prime\}\}\)^\{\(N\_\{v\}\-1\)/p^\{\\prime\}\-1\}\\,dt\(the one\-coordinate marginal of the cone measure\), whose term\-by\-term expansion is a confluent series of Kummer type\. Forp′=2p^\{\\prime\}=2it is precisely a confluent hypergeometric functionF11\{\}\_\{1\}F\_\{1\}, equivalently a modified Bessel function\[[49](https://arxiv.org/html/2609.10976#bib.bib49), Ch\. 13\], and for rationalp′p^\{\\prime\}it reduces to finite combinations of generalized hypergeometric functions\[[48](https://arxiv.org/html/2609.10976#bib.bib48)\]\. None of these yields an elementary closed form at a generic field: Models A and C are the two exactly solvable members of the family, between which the otherℓp\\ell\_\{p\}members interpolate\.

Remarks\.Three structural comments conclude this subsection\. \(i\)*The origin is necessarily singular\.*IfLv\{L\_\{v\}\}is differentiable at the origin, the subdifferentialK=∂Lv​\(0\)K=\\partial\{L\_\{v\}\}\(0\)reduces to the single point∇Lv​\(0\)\\nabla\{L\_\{v\}\}\(0\)\[[45](https://arxiv.org/html/2609.10976#bib.bib45), Sec\. 25\], and Eq\. \([79](https://arxiv.org/html/2609.10976#A2.E79)\) givesLv\(v\)=∇Lv\(0\)⊤v\{L\_\{v\}\}\(v\)=\\nabla\{L\_\{v\}\}\(0\)^\{\\top\}v: the Lagrangian is linear, the activation is a constant vector, and the interaction energy no longer depends on the visible configuration: no memory is stored\. Every nontrivial member of the family is therefore non\-differentiable atv=0v=0: the kinks of‖⋅‖1\\left\\lVert\\cdot\\right\\rVert\_\{1\}on the coordinate hyperplanes and the conical point of‖⋅‖2\\left\\lVert\\cdot\\right\\rVert\_\{2\}at the origin are structural features of scale\-free activations rather than accidental features of the two examples\. Physically,v=0v=0carries no direction information, and a pure direction readout cannot be continuously defined there\. \(ii\)*Norms and beyond\.*The support functionhKh\_\{K\}is a norm exactly whenKKis centrally symmetric with the origin in its interior\[[44](https://arxiv.org/html/2609.10976#bib.bib44)\]\. This is the case relevant to the±\\pm\-symmetric patterns of the main text\. General members allow asymmetric bodies, e\.g\.,hK​\(v\)=maxl⁡al⊤​vh\_\{K\}\(v\)=\\max\_\{l\}a\_\{l\}^\{\\top\}vwith a generic vertex set, for whichLv​\(−v\)≠Lv​\(v\)\{L\_\{v\}\}\(\-v\)\\neq\{L\_\{v\}\}\(v\)and patterns and anti\-patterns are treated asymmetrically\. The family is in one\-to\-one correspondence with compact convex bodies, of which theℓp\\ell\_\{p\}balls form a one\-parameter slice\. This is the precise content of the footnote in Sec\.[III](https://arxiv.org/html/2609.10976#S3)\. \(iii\)*The hidden layer\.*The criterion Eq\. \([76](https://arxiv.org/html/2609.10976#A2.E76)\) applies equally toLh\{L\_\{h\}\}: a degree\-one homogeneous hidden Lagrangian would remove the bare hidden neurons from Eq\. \([4](https://arxiv.org/html/2609.10976#S2.E4)\) as well\. The choiceLh=∑μF⁡\(hμ\)\{L\_\{h\}\}=\\sum\_\{\\mu\}F\(h\_\{\\mu\}\)withF⁡\(x\)=xk/kF\(x\)=x^\{k\}/kis homogeneous of degreek≠1k\\neq 1, so the hidden Legendre term survives \(it is precisely the confining termNv​γk​∑μmμkN\_\{v\}\\gamma\_\{k\}\\sum\_\{\\mu\}m\_\{\\mu\}^\{k\}of Eq\. \([17](https://arxiv.org/html/2609.10976#S3.E17)\)\), and the bare hidden neurons remain in the energy, which allows them to act as the order parameters of memory retrieval\.

### B\.2Model A replica analysis

In this appendix we derive the RS free energy Eq\. \([21](https://arxiv.org/html/2609.10976#S3.E21)\) and the equations of state Eq\. \([23](https://arxiv.org/html/2609.10976#S3.E23)\)–Eq\. \([25](https://arxiv.org/html/2609.10976#S3.E25)\) of Model A\. We follow the standard procedure of the replica method\[[50](https://arxiv.org/html/2609.10976#bib.bib50),[51](https://arxiv.org/html/2609.10976#bib.bib51)\]: after averaging the replicated partition function over the patterns, the crosstalk noise is characterized by an overlap matrixRRand its conjugateR^\\hat\{R\}, and the hidden\-sector integral factorizes over the non\-condensed modes\. Fork=2k=2the hidden sector is Gaussian, the computation closes exactly, and it reproduces the AGS theory of the Hopfield model\[[2](https://arxiv.org/html/2609.10976#bib.bib2),[3](https://arxiv.org/html/2609.10976#bib.bib3),[4](https://arxiv.org/html/2609.10976#bib.bib4)\]\. Fork\>2k\>2, in contrast, we show that the non\-condensed hidden integral admits no consistent evaluation at finiteβk\\beta\_\{k\}\(a genuine pathology of the continuous hidden variables at the loadNh=αk​Nvk−1N\_\{h\}=\\alpha\_\{k\}N\_\{v\}^\{k\-1\}, not a technical artifact\), so that the analysis must be based on the adiabatic reduction of Sec\.[II\.2](https://arxiv.org/html/2609.10976#S2.SS2), for which we then carry out the corrected computation\.

#### B\.2\.1Replica setup and order parameters

Throughout we work with the normalized hidden variablesmμ=hμ/\(λ/τv\)m\_\{\\mu\}=h\_\{\\mu\}/\(\\lambda/\{\\tau\_\{v\}\}\)introduced in the main text, and the Jacobian of this change of variables only contributes an irrelevant constant\. Introducing the replica indexa=1,…,na=1,\\ldots,nand averaging Eq\. \([17](https://arxiv.org/html/2609.10976#S3.E17)\) over the patternsξμ​i\(h,v\)∼Unif⁡\(\{±1\}\)\\xi^\{\(h,v\)\}\_\{\\mu i\}\\sim\\mathrm\{Unif\}\(\\left\\\{\\pm 1\\right\\\}\), which are independent across the sitesii, we have

𝔼ξ\[\(ZξA\)n\]=∫∏a=1ndmae−Nvγk∑μ,a\(maμ\)k∏i=1Nv𝔼ξi\[∏a=1n2coshβk\(∑μξi​μ\(v,h\)\(mμa\)k−1\)\],\\mathbb\{E\}\_\{\\xi\}\\left\[\(\{Z\_\{\\xi\}^\{\\mathrm\{A\}\}\}\)^\{n\}\\right\]=\\int\\prod\_\{a=1\}^\{n\}dm^\{a\}\\,e^\{\-N\_\{v\}\\gamma\_\{k\}\\sum\_\{\\mu,a\}\(m^\{a\}\_\{\\mu\}\)^\{k\}\}\\prod\_\{i=1\}^\{N\_\{v\}\}\\mathbb\{E\}\_\{\\xi\_\{i\}\}\\left\[\\prod\_\{a=1\}^\{n\}2\\cosh\\beta\_\{k\}\\bigg\(\\sum\_\{\\mu\}\\xi^\{\(v,h\)\}\_\{i\\mu\}\(m^\{a\}\_\{\\mu\}\)^\{k\-1\}\\bigg\)\\right\],\(84\)whereξi\\xi\_\{i\}denotes theii\-th row of the pattern matrix\. Since the pattern distribution is invariant underξi​μ\(v,h\)→ξi​1\(v,h\)​ξi​μ\(v,h\)\\xi^\{\(v,h\)\}\_\{i\\mu\}\\to\\xi^\{\(v,h\)\}\_\{i1\}\\xi^\{\(v,h\)\}\_\{i\\mu\}for eachii, we may setξi​1\(v,h\)=1\\xi^\{\(v,h\)\}\_\{i1\}=1for alliiwithout loss of generality\. To describe the retrieval phase, we assume that only the first pattern is condensed,

ma≔m1a=O\(1\),mμa=O\(Nv−1/2\)\(μ≥2\),m^\{a\}\\coloneq m^\{a\}\_\{1\}=O\(1\),\\qquad m^\{a\}\_\{\\mu\}=O\(N\_\{v\}^\{\-1/2\}\)\\quad\(\\mu\\geq 2\),\(85\)and collect the non\-condensed contributions to the local field into the crosstalk noise

uia≔∑μ≥2ξi​μ\(v,h\)​\(mμa\)k−1\.u\_\{i\}^\{a\}\\coloneq\\sum\_\{\\mu\\geq 2\}\\xi^\{\(v,h\)\}\_\{i\\mu\}\(m^\{a\}\_\{\\mu\}\)^\{k\-1\}\.\(86\)For fixed\{mμa\}\\left\\\{m^\{a\}\_\{\\mu\}\\right\\\}, the vectorsui=\(uia\)a=1nu\_\{i\}=\(u\_\{i\}^\{a\}\)\_\{a=1\}^\{n\}are i\.i\.d\. across the sites with zero mean, and by the central limit theorem they become Gaussian at largeNvN\_\{v\},

ui→d𝒩⁡\(0,R\),Ra​b≔∑μ≥2\(mμa\)k−1​\(mμb\)k−1,u\_\{i\}\\xrightarrow\{\\mathrm\{d\}\}\\mathcal\{N\}\(0,R\),\\qquad R^\{ab\}\\coloneq\\sum\_\{\\mu\\geq 2\}\(m^\{a\}\_\{\\mu\}\)^\{k\-1\}\(m^\{b\}\_\{\\mu\}\)^\{k\-1\},\(87\)where the covariance matrixRRis of orderNh×Nv−\(k−1\)=αk=O⁡\(1\)N\_\{h\}\\times N\_\{v\}^\{\-\(k\-1\)\}=\\alpha\_\{k\}=O\(1\)on the assumed scaling of the non\-condensed modes\. The site average in Eq\. \([84](https://arxiv.org/html/2609.10976#A2.E84)\) thus becomes

𝔼ξi\[∏a2coshβk\(\(ma\)k−1\+uia\)\]→Ψ\(m,R\)≔𝔼z∼𝒩⁡\(0,R\)\[∏a2coshβk\(\(ma\)k−1\+za\)\]\.\\mathbb\{E\}\_\{\\xi\_\{i\}\}\\left\[\\prod\_\{a\}2\\cosh\\beta\_\{k\}\\big\(\(m^\{a\}\)^\{k\-1\}\+u\_\{i\}^\{a\}\\big\)\\right\]\\to\\Psi\(m,R\)\\coloneq\\mathbb\{E\}\_\{z\\sim\\mathcal\{N\}\(0,R\)\}\\left\[\\prod\_\{a\}2\\cosh\\beta\_\{k\}\\big\(\(m^\{a\}\)^\{k\-1\}\+z^\{a\}\\big\)\\right\]\.\(88\)
We next promoteRa​bR^\{ab\}to independent integration variables by inserting the identity

1\\displaystyle 1=∫∏a,bd​Ra​b​δ​\(Ra​b−∑μ≥2\(mμa\)k−1​\(mμb\)k−1\)\\displaystyle=\\int\\prod\_\{a,b\}dR^\{ab\}\\,\\delta\\bigg\(R^\{ab\}\-\\sum\_\{\\mu\\geq 2\}\(m^\{a\}\_\{\\mu\}\)^\{k\-1\}\(m^\{b\}\_\{\\mu\}\)^\{k\-1\}\\bigg\)\(89\)∝∫∏a,bdRa​bdR^a​bexp\{−βk2​Nv2∑a,bR^a​b\(Ra​b−∑μ≥2\(maμ\)k−1\(mbμ\)k−1\)\},\\displaystyle\\propto\\int\\prod\_\{a,b\}dR^\{ab\}d\\hat\{R\}^\{ab\}\\exp\\bigg\\\{\-\\frac\{\\beta\_\{k\}^\{2\}N\_\{v\}\}\{2\}\\sum\_\{a,b\}\\hat\{R\}^\{ab\}\\bigg\(R^\{ab\}\-\\sum\_\{\\mu\\geq 2\}\(m^\{a\}\_\{\\mu\}\)^\{k\-1\}\(m^\{b\}\_\{\\mu\}\)^\{k\-1\}\\bigg\)\\bigg\\\},where the conjugate variablesR^a​b\\hat\{R\}^\{ab\}run along the imaginary axis and take real values at the saddle point, and the prefactorβk2​Nv/2\\beta\_\{k\}^\{2\}N\_\{v\}/2is chosen for later convenience\. The integrals over the non\-condensed modes then factorize,

∏μ=2Nh∫∏admμaexp\[−Nvγk∑a\(mμa\)k\+βk2​Nv2∑a,bR^a​b\(mμa\)k−1\(mμb\)k−1\]=ℐk\(R^\)Nh−1,\\prod\_\{\\mu=2\}^\{N\_\{h\}\}\\int\\prod\_\{a\}dm^\{a\}\_\{\\mu\}\\exp\\bigg\[\-N\_\{v\}\\gamma\_\{k\}\\sum\_\{a\}\(m^\{a\}\_\{\\mu\}\)^\{k\}\+\\frac\{\\beta\_\{k\}^\{2\}N\_\{v\}\}\{2\}\\sum\_\{a,b\}\\hat\{R\}^\{ab\}\(m^\{a\}\_\{\\mu\}\)^\{k\-1\}\(m^\{b\}\_\{\\mu\}\)^\{k\-1\}\\bigg\]=\\mathcal\{I\}\_\{k\}\(\\hat\{R\}\)^\{N\_\{h\}\-1\},\(90\)with the single\-mode integral

ℐk\(R^\)≔∫ℝn∏adxaexp\[−Nvγk∑a\(xa\)k\+βk2​Nv2∑a,bR^a​b\(xa\)k−1\(xb\)k−1\],\\mathcal\{I\}\_\{k\}\(\\hat\{R\}\)\\coloneq\\int\_\{\\mathbb\{R\}^\{n\}\}\\prod\_\{a\}dx^\{a\}\\exp\\bigg\[\-N\_\{v\}\\gamma\_\{k\}\\sum\_\{a\}\(x^\{a\}\)^\{k\}\+\\frac\{\\beta\_\{k\}^\{2\}N\_\{v\}\}\{2\}\\sum\_\{a,b\}\\hat\{R\}^\{ab\}\(x^\{a\}\)^\{k\-1\}\(x^\{b\}\)^\{k\-1\}\\bigg\],\(91\)and the replicated partition function reads

𝔼ξ\[\(ZξA\)n\]=∫∏adma∫∏a,bdRa​bdR^a​bexp\[−Nvγk∑a\(ma\)k−βk2​Nv2∑a,bR^a​bRa​b\+NvlogΨ\(m,R\)\]ℐk\(R^\)Nh−1\.\\mathbb\{E\}\_\{\\xi\}\\left\[\(\{Z\_\{\\xi\}^\{\\mathrm\{A\}\}\}\)^\{n\}\\right\]=\\int\\prod\_\{a\}dm^\{a\}\\int\\prod\_\{a,b\}dR^\{ab\}d\\hat\{R\}^\{ab\}\\exp\\bigg\[\-N\_\{v\}\\gamma\_\{k\}\\sum\_\{a\}\(m^\{a\}\)^\{k\}\-\\frac\{\\beta\_\{k\}^\{2\}N\_\{v\}\}\{2\}\\sum\_\{a,b\}\\hat\{R\}^\{ab\}R^\{ab\}\+N\_\{v\}\\log\\Psi\(m,R\)\\bigg\]\\,\\mathcal\{I\}\_\{k\}\(\\hat\{R\}\)^\{N\_\{h\}\-1\}\.\(92\)
We now impose the RS ansatz,

ma=m,Ra​a=αk​rd,Ra​b=αk​r​\(a≠b\),R^a​a=r^d,R^a​b=r^​\(a≠b\),m^\{a\}=m,\\qquad R^\{aa\}=\\alpha\_\{k\}r\_\{d\},\\quad R^\{ab\}=\\alpha\_\{k\}r~\(a\\neq b\),\\qquad\\hat\{R\}^\{aa\}=\\hat\{r\}\_\{d\},\\quad\\hat\{R\}^\{ab\}=\\hat\{r\}~\(a\\neq b\),\(93\)where the factorsαk\\alpha\_\{k\}are inserted so thatrrmatches the noise parameter of the main text\. To evaluateΨ\\Psi, we realize the Gaussian vectorz∼𝒩⁡\(0,R\)z\\sim\\mathcal\{N\}\(0,R\)as

za=αk​r​z\+αk​\(rd−r\)​wa,z,wa∼𝒩⁡\(0,1\)​i\.i\.d\.,z^\{a\}=\\sqrt\{\\alpha\_\{k\}r\}\\,z\+\\sqrt\{\\alpha\_\{k\}\(r\_\{d\}\-r\)\}\\,w^\{a\},\\qquad z,w^\{a\}\\sim\\mathcal\{N\}\(0,1\)~\\text\{i\.i\.d\.\},\(94\)which decomposes the noise into a componentzzfrozen across the replicas and thermal componentswaw^\{a\}independent between the replicas\. Using∫D​w​2​cosh⁡\(A\+B​w\)=2​eB2/2​cosh⁡A\\int Dw\\,2\\cosh\(A\+Bw\)=2e^\{B^\{2\}/2\}\\cosh Aand expanding to first order innn, we obtain

logΨ\(m,R\)=n\[βk2​αk2\(rd−r\)\+∫Dzlog2coshβk\(mk−1\+αk​rz\)\]\+O\(n2\)\.\\log\\Psi\(m,R\)=n\\bigg\[\\frac\{\\beta\_\{k\}^\{2\}\\alpha\_\{k\}\}\{2\}\(r\_\{d\}\-r\)\+\\int Dz\\log 2\\cosh\\beta\_\{k\}\\big\(m^\{k\-1\}\+\\sqrt\{\\alpha\_\{k\}r\}\\,z\\big\)\\bigg\]\+O\(n^\{2\}\)\.\(95\)Similarly, the conjugate term in Eq\. \([92](https://arxiv.org/html/2609.10976#A2.E92)\) becomes

−βk2​Nv2∑a,bR^a​bRa​b=−βk2​Nv​αk2\[nr^drd\+n\(n−1\)r^r\]\.\-\\frac\{\\beta\_\{k\}^\{2\}N\_\{v\}\}\{2\}\\sum\_\{a,b\}\\hat\{R\}^\{ab\}R^\{ab\}=\-\\frac\{\\beta\_\{k\}^\{2\}N\_\{v\}\\alpha\_\{k\}\}\{2\}\\big\[n\\hat\{r\}\_\{d\}r\_\{d\}\+n\(n\-1\)\\hat\{r\}r\\big\]\.\(96\)
Two of the saddle\-point equations can already be read off, becauseℐk\\mathcal\{I\}\_\{k\}depends only on the conjugate variables\. Stationarity of the exponent with respect tordr\_\{d\}gives, from Eq\. \([95](https://arxiv.org/html/2609.10976#A2.E95)\) and the conjugate term,

βk2​αk2−βk2​αk2​r^d=0⟹r^d=1\.\\frac\{\\beta\_\{k\}^\{2\}\\alpha\_\{k\}\}\{2\}\-\\frac\{\\beta\_\{k\}^\{2\}\\alpha\_\{k\}\}\{2\}\\hat\{r\}\_\{d\}=0\\qquad\\Longrightarrow\\qquad\\hat\{r\}\_\{d\}=1\.\(97\)Substitutingr^d=1\\hat\{r\}\_\{d\}=1back, the twordr\_\{d\}\-dependent terms cancel identically, so thatrdr\_\{d\}andr^d\\hat\{r\}\_\{d\}disappear from the free energy altogether\. The value ofrdr\_\{d\}itself is fixed by stationarity with respect tor^d\\hat\{r\}\_\{d\}, which identifiesαk​rd\\alpha\_\{k\}r\_\{d\}with the average of∑μ≥2\(mμa\)2​\(k−1\)\\sum\_\{\\mu\\geq 2\}\(m^\{a\}\_\{\\mu\}\)^\{2\(k\-1\)\}in the hidden sector, but it plays no further role\. Stationarity with respect torrgives, using Gaussian integration by parts∫Dzztanhβk\(mk−1\+αk​rz\)=βkαk​r∫Dz\{sech\}2βk\(mk−1\+αk​rz\)\\int Dz\\,z\\tanh\\beta\_\{k\}\(m^\{k\-1\}\+\\sqrt\{\\alpha\_\{k\}r\}z\)=\\beta\_\{k\}\\sqrt\{\\alpha\_\{k\}r\}\\int Dz\\sech^\{2\}\\beta\_\{k\}\(m^\{k\-1\}\+\\sqrt\{\\alpha\_\{k\}r\}z\),

r^=∫D​z​tanh2⁡βk​\(mk−1\+αk​r​z\)=q\.\\hat\{r\}=\\int Dz\\tanh^\{2\}\\beta\_\{k\}\\big\(m^\{k\-1\}\+\\sqrt\{\\alpha\_\{k\}r\}\\,z\\big\)=q\.\(98\)The conjugate variabler^\\hat\{r\}is therefore precisely the Edwards–Anderson order parameterqqof the visible spinssi=sgn⁡\(vi\)s\_\{i\}=\\sgn\(v\_\{i\}\): the integrand of Eq\. \([98](https://arxiv.org/html/2609.10976#A2.E98)\) is the squared thermal average⟨si⟩2\\left\\langle s\_\{i\}\\right\\rangle^\{2\}in the effective single\-site measure\. Note that Eq\. \([97](https://arxiv.org/html/2609.10976#A2.E97)\) and Eq\. \([98](https://arxiv.org/html/2609.10976#A2.E98)\) hold for everykk\. What distinguishesk=2k=2fromk\>2k\>2is the remaining hidden\-sector factorℐk​\(R^\)\\mathcal\{I\}\_\{k\}\(\\hat\{R\}\), to which we now turn\.

#### B\.2\.2The casek=2k=2: Gaussian hidden sector and the AGS theory

Fork=2k=2we haveγ2=βk/2\\gamma\_\{2\}=\\beta\_\{k\}/2, and the single\-mode integral Eq\. \([91](https://arxiv.org/html/2609.10976#A2.E91)\) is Gaussian:

ℐ2\(R^\)=∫ℝn∏adxaexp\[−βk​Nv2∑a,bxa\(δa​b−βkR^a​b\)xb\]∝det\(In−βkR^\)−1/2,\\mathcal\{I\}\_\{2\}\(\\hat\{R\}\)=\\int\_\{\\mathbb\{R\}^\{n\}\}\\prod\_\{a\}dx^\{a\}\\exp\\bigg\[\-\\frac\{\\beta\_\{k\}N\_\{v\}\}\{2\}\\sum\_\{a,b\}x^\{a\}\\big\(\\delta^\{ab\}\-\\beta\_\{k\}\\hat\{R\}^\{ab\}\\big\)x^\{b\}\\bigg\]\\propto\\det\\big\(I\_\{n\}\-\\beta\_\{k\}\\hat\{R\}\\big\)^\{\-1/2\},\(99\)up to an irrelevant constant\. Under the RS ansatz,R^\\hat\{R\}has the eigenvaluer^d\+\(n−1\)​r^\\hat\{r\}\_\{d\}\+\(n\-1\)\\hat\{r\}\(non\-degenerate\) andr^d−r^\\hat\{r\}\_\{d\}\-\\hat\{r\}with degeneracyn−1n\-1, so that

logdet\(In−βkR^\)=\(n−1\)log\(1−βk\(r^d−r^\)\)\+log\(1−βk\(r^d−r^\)−nβkr^\)\.\\log\\det\\big\(I\_\{n\}\-\\beta\_\{k\}\\hat\{R\}\\big\)=\(n\-1\)\\log\\big\(1\-\\beta\_\{k\}\(\\hat\{r\}\_\{d\}\-\\hat\{r\}\)\\big\)\+\\log\\big\(1\-\\beta\_\{k\}\(\\hat\{r\}\_\{d\}\-\\hat\{r\}\)\-n\\beta\_\{k\}\\hat\{r\}\\big\)\.\(100\)Expanding to first order innnand insertingr^d=1\\hat\{r\}\_\{d\}=1,r^=q\\hat\{r\}=q, we obtain

log⁡ℐ2​\(R^\)=−n2​\[log⁡\(1−βk​\(1−q\)\)−βk​q1−βk​\(1−q\)\]\+O⁡\(n2\)=−n2​Ψ2​\(q\)\+O⁡\(n2\),\\log\\mathcal\{I\}\_\{2\}\(\\hat\{R\}\)=\-\\frac\{n\}\{2\}\\bigg\[\\log\\big\(1\-\\beta\_\{k\}\(1\-q\)\\big\)\-\\frac\{\\beta\_\{k\}q\}\{1\-\\beta\_\{k\}\(1\-q\)\}\\bigg\]\+O\(n^\{2\}\)=\-\\frac\{n\}\{2\}\\Psi\_\{2\}\(q\)\+O\(n^\{2\}\),\(101\)which, multiplied byNh−1≃αk​NvN\_\{h\}\-1\\simeq\\alpha\_\{k\}N\_\{v\}, yields precisely the noise entropic term−n​Nv​αk2​Ψ2​\(q\)\-\\frac\{nN\_\{v\}\\alpha\_\{k\}\}\{2\}\\Psi\_\{2\}\(q\)of the main text\. Collecting Eq\. \([95](https://arxiv.org/html/2609.10976#A2.E95)\)–Eq\. \([101](https://arxiv.org/html/2609.10976#A2.E101)\) in Eq\. \([92](https://arxiv.org/html/2609.10976#A2.E92)\) and using the replica trick

fA\(β\)=−1β​Nvlimn→01n\(𝔼ξ\[\(ZξA\)n\]−1\),f^\{\\mathrm\{A\}\}\(\\beta\)=\-\\frac\{1\}\{\\beta N\_\{v\}\}\\lim\_\{n\\to 0\}\\frac\{1\}\{n\}\\Big\(\\mathbb\{E\}\_\{\\xi\}\\left\[\(\{Z\_\{\\xi\}^\{\\mathrm\{A\}\}\}\)^\{n\}\\right\]\-1\\Big\),\(102\)we arrive at the RS free energy Eq\. \([21](https://arxiv.org/html/2609.10976#S3.E21)\) withΨk=Ψ2\\Psi\_\{k\}=\\Psi\_\{2\}\. In the assembly, the diagonal terms combine asβk2​αk2​\(rd−r^d​rd\)=0\\frac\{\\beta\_\{k\}^\{2\}\\alpha\_\{k\}\}\{2\}\(r\_\{d\}\-\\hat\{r\}\_\{d\}r\_\{d\}\)=0, which is the cancellation ofrdr\_\{d\}anticipated above, while the off\-diagonal terms produceαk​βk22​r​\(1−q\)\\frac\{\\alpha\_\{k\}\\beta\_\{k\}^\{2\}\}\{2\}r\(1\-q\)upon usingr^=q\\hat\{r\}=q\. Finally, stationarity with respect tor^\\hat\{r\}picks up contributions from the conjugate term and fromlog⁡ℐ2\\log\\mathcal\{I\}\_\{2\}:

βk2​αk2r=−αk∂∂r^limn→01nlogℐ2=αk2βk2​r^\(1−βk​\(r^d−r^\)\)2⟹r=q\(1−βk​\(1−q\)\)2=ℳ2\(q\),\\frac\{\\beta\_\{k\}^\{2\}\\alpha\_\{k\}\}\{2\}r=\-\\alpha\_\{k\}\\frac\{\\partial\}\{\\partial\\hat\{r\}\}\\lim\_\{n\\to 0\}\\frac\{1\}\{n\}\\log\\mathcal\{I\}\_\{2\}=\\frac\{\\alpha\_\{k\}\}\{2\}\\frac\{\\beta\_\{k\}^\{2\}\\hat\{r\}\}\{\\big\(1\-\\beta\_\{k\}\(\\hat\{r\}\_\{d\}\-\\hat\{r\}\)\\big\)^\{2\}\}\\qquad\\Longrightarrow\\qquad r=\\frac\{q\}\{\\big\(1\-\\beta\_\{k\}\(1\-q\)\\big\)^\{2\}\}=\\mathcal\{M\}\_\{2\}\(q\),\(103\)reproducing the AGS equations of state of the Hopfield model near saturation\[[2](https://arxiv.org/html/2609.10976#bib.bib2),[3](https://arxiv.org/html/2609.10976#bib.bib3),[4](https://arxiv.org/html/2609.10976#bib.bib4)\]\. We emphasize that fork=2k=2no adiabatic assumption has been made: the hidden variables enter quadratically, their Hessian is field\-independent, and the Gaussian integration is exact at anyβ/τh\\beta/\{\\tau\_\{h\}\}\.

#### B\.2\.3The casek\>2k\>2: breakdown of the hidden\-sector integral and the adiabatic reduction

Fork\>2k\>2the integral Eq\. \([91](https://arxiv.org/html/2609.10976#A2.E91)\) does not admit an analogous evaluation, for a reason that is structural rather than technical\. Rescalingxa=Nvc​yax^\{a\}=N\_\{v\}^\{c\}y^\{a\}, the two terms in the exponent ofℐk\\mathcal\{I\}\_\{k\}scale as

Nv​γk​∑a\(xa\)k∼Nv1\+k​c,βk2​Nv2​∑a,bR^a​b​\(xa\)k−1​\(xb\)k−1∼Nv1\+2​\(k−1\)​c\.N\_\{v\}\\gamma\_\{k\}\\sum\_\{a\}\(x^\{a\}\)^\{k\}\\sim N\_\{v\}^\{1\+kc\},\\qquad\\frac\{\\beta\_\{k\}^\{2\}N\_\{v\}\}\{2\}\\sum\_\{a,b\}\\hat\{R\}^\{ab\}\(x^\{a\}\)^\{k\-1\}\(x^\{b\}\)^\{k\-1\}\\sim N\_\{v\}^\{1\+2\(k\-1\)c\}\.\(104\)Fork=2k=2the two exponents coincide for anycc, and the choicec=−1/2c=\-1/2makes bothO⁡\(1\)O\(1\)\. TheNh−1≃αk​NvN\_\{h\}\-1\\simeq\\alpha\_\{k\}N\_\{v\}modes then sum up to an extensive contribution, which is the calculation of Appendix[B\.2\.2](https://arxiv.org/html/2609.10976#A2.SS2.SSS2)\. Fork\>2k\>2, however, no choice ofccbalances the two terms at a nontrivial order: at the central\-limit scalec=−1/2c=\-1/2, which underlies the ansatzR=O⁡\(1\)R=O\(1\)in Eq\. \([87](https://arxiv.org/html/2609.10976#A2.E87)\), both exponents are negative and the integrand ofℐk\\mathcal\{I\}\_\{k\}becomes flat, so the integral is not confined to this scale\. Instead,ℐk\\mathcal\{I\}\_\{k\}is dominated by the bare confinement scalec=−1/kc=\-1/k, on which the coupling term still vanishes asNv\(2−k\)/kN\_\{v\}^\{\(2\-k\)/k\}\. This has two consequences\. First, the typical amplitude of a non\-condensed mode is\|mμ\|∼Nv−1/k≫Nv−1/2\|m\_\{\\mu\}\|\\sim N\_\{v\}^\{\-1/k\}\\gg N\_\{v\}^\{\-1/2\}, so that the diagonal noise covariance evaluates to

Ra​a=∑μ≥2\(mμa\)2​\(k−1\)∼αkNvk−1⋅Nv−2\(k−1\)/k=αkNv\(k−1\)​\(k−2\)/k⟶∞,R^\{aa\}=\\sum\_\{\\mu\\geq 2\}\(m^\{a\}\_\{\\mu\}\)^\{2\(k\-1\)\}\\sim\\alpha\_\{k\}N\_\{v\}^\{k\-1\}\\cdot N\_\{v\}^\{\-2\(k\-1\)/k\}=\\alpha\_\{k\}N\_\{v\}^\{\(k\-1\)\(k\-2\)/k\}\\longrightarrow\\infty,\(105\)in contradiction withR=O⁡\(1\)R=O\(1\): the crosstalk variance diverges, and the continuous hidden variables cannot sustain the loadNh=αk​Nvk−1N\_\{h\}=\\alpha\_\{k\}N\_\{v\}^\{k\-1\}at any finiteβk\\beta\_\{k\}\. Second, and equivalently, theR^\\hat\{R\}\-dependent part of\(Nh−1\)​log⁡ℐk\(N\_\{h\}\-1\)\\log\\mathcal\{I\}\_\{k\}is of orderNv\(k2−2​k\+2\)/kN\_\{v\}^\{\(k^\{2\}\-2k\+2\)/k\}, which is superextensive fork\>2k\>2, so no extensive saddle\-point structure exists for the conjugate variables\. We also note that this pathology cannot be removed by any limit of the time\-scale parameters: restoring the original variables viah=\(λ/τv\)​mh=\(\\lambda/\{\\tau\_\{v\}\}\)mshows that the equilibrium partition function Eq\. \([14](https://arxiv.org/html/2609.10976#S3.E14)\) depends onβ/τh\\beta/\{\\tau\_\{h\}\}andλ/τv\\lambda/\{\\tau\_\{v\}\}only through the combinationβk=\(β/τh\)​\(λ/τv\)k\\beta\_\{k\}=\(\\beta/\{\\tau\_\{h\}\}\)\(\\lambda/\{\\tau\_\{v\}\}\)^\{k\}, so that “β/τh→∞\\beta/\{\\tau\_\{h\}\}\\to\\inftyat fixedβk\\beta\_\{k\}” leaves the equilibrium measure unchanged\.

Fork\>2k\>2we therefore*define*Model A through the adiabatic reduction of Sec\.[II\.2](https://arxiv.org/html/2609.10976#S2.SS2), in which the hidden neurons are replaced by their stationary values given the visible configuration, as in Eq\. \([10](https://arxiv.org/html/2609.10976#S2.E10)\)\. This definition is motivated by the dynamics: in the adiabatic regimeτv≫τh\{\\tau\_\{v\}\}\\gg\{\\tau\_\{h\}\}the hidden neurons deterministically track the visible configuration, and their thermal fluctuations \(the source of the runaway Eq\. \([105](https://arxiv.org/html/2609.10976#A2.E105)\)\) are switched off\. Concretely, writing the partition function with the visible spinssi∈\{±1\}s\_\{i\}\\in\\left\\\{\\pm 1\\right\\\}unintegrated and eliminating each modemμm\_\{\\mu\}by its stationarity condition,k​γk​mμk−1=\(k−1\)​βk​mμk−2​m^μk\\gamma\_\{k\}m\_\{\\mu\}^\{k\-1\}=\(k\-1\)\\beta\_\{k\}m\_\{\\mu\}^\{k\-2\}\\hat\{m\}\_\{\\mu\}withm^μ​\(s\)≔1Nv​∑iξμ​i\(h,v\)​si\\hat\{m\}\_\{\\mu\}\(s\)\\coloneq\\frac\{1\}\{N\_\{v\}\}\\sum\_\{i\}\\xi^\{\(h,v\)\}\_\{\\mu i\}s\_\{i\}, we obtainmμ=m^μ​\(s\)m\_\{\\mu\}=\\hat\{m\}\_\{\\mu\}\(s\)and

−Nv​γk​mμk\+βk​Nv​mμk−1​m^μ\|mμ=m^μ=βkk​Nv​m^μk,ZξA,ad≔∑s∈\{±1\}Nvexp⁡\[βkk​Nv​∑μm^μ​\(s\)k\],\-N\_\{v\}\\gamma\_\{k\}m\_\{\\mu\}^\{k\}\+\\beta\_\{k\}N\_\{v\}m\_\{\\mu\}^\{k\-1\}\\hat\{m\}\_\{\\mu\}\\Big\|\_\{m\_\{\\mu\}=\\hat\{m\}\_\{\\mu\}\}=\\frac\{\\beta\_\{k\}\}\{k\}N\_\{v\}\\hat\{m\}\_\{\\mu\}^\{k\},\\qquad Z^\{\\mathrm\{A,ad\}\}\_\{\\xi\}\\coloneq\\sum\_\{s\\in\\left\\\{\\pm 1\\right\\\}^\{N\_\{v\}\}\}\\exp\\bigg\[\\frac\{\\beta\_\{k\}\}\{k\}N\_\{v\}\\sum\_\{\\mu\}\\hat\{m\}\_\{\\mu\}\(s\)^\{k\}\\bigg\],\(106\)which is the dense associative memory with polynomial energy\[[10](https://arxiv.org/html/2609.10976#bib.bib10)\], also known as the multiconnected network\[[7](https://arxiv.org/html/2609.10976#bib.bib7)\]\. For the condensed mode this substitution coincides with the exact Laplace evaluation of its integral, whose exponent isO⁡\(Nv\)O\(N\_\{v\}\)with anO⁡\(1\)O\(1\)saddle\. Fork=2k=2it agrees with the exact Gaussian integration up to a constant, as noted above\. The definition is therefore consistent acrosskk, and the pathology is confined to the thermal fluctuations of the non\-condensed continuous modes atk\>2k\>2\.

#### B\.2\.4The casek\>2k\>2: noise sector from the cumulant expansion

We now carry out the replica analysis of Eq\. \([106](https://arxiv.org/html/2609.10976#A2.E106)\)\. Replicating and averaging over the patterns, the condensed and non\-condensed sectors factorize:

𝔼ξ​\[\(ZξA,ad\)n\]=∑\{sa\}exp⁡\[βkk​Nv​∑am^1​\(sa\)k\]​∏μ≥2𝔼ξμ​\[exp⁡\(ε​∑a\(y^μa\)k\)\],ε≔βkk​Nv1−k/2,\\mathbb\{E\}\_\{\\xi\}\\left\[\(Z^\{\\mathrm\{A,ad\}\}\_\{\\xi\}\)^\{n\}\\right\]=\\sum\_\{\\left\\\{s^\{a\}\\right\\\}\}\\exp\\bigg\[\\frac\{\\beta\_\{k\}\}\{k\}N\_\{v\}\\sum\_\{a\}\\hat\{m\}\_\{1\}\(s^\{a\}\)^\{k\}\\bigg\]\\prod\_\{\\mu\\geq 2\}\\mathbb\{E\}\_\{\\xi\_\{\\mu\}\}\\left\[\\exp\\bigg\(\\varepsilon\\sum\_\{a\}\(\\hat\{y\}^\{a\}\_\{\\mu\}\)^\{k\}\\bigg\)\\right\],\\qquad\\varepsilon\\coloneq\\frac\{\\beta\_\{k\}\}\{k\}N\_\{v\}^\{1\-k/2\},\(107\)where we used the gaugeξi​1\(v,h\)=1\\xi^\{\(v,h\)\}\_\{i1\}=1and introduced the rescaled non\-condensed overlaps

y^μa≔1Nv​∑iξμ​i\(h,v\)​sia=Nv​m^μ​\(sa\)\.\\hat\{y\}^\{a\}\_\{\\mu\}\\coloneq\\frac\{1\}\{\\sqrt\{N\_\{v\}\}\}\\sum\_\{i\}\\xi^\{\(h,v\)\}\_\{\\mu i\}s\_\{i\}^\{a\}=\\sqrt\{N\_\{v\}\}\\,\\hat\{m\}\_\{\\mu\}\(s^\{a\}\)\.\(108\)For fixed spin configurations, the vector\(y^μa\)a=1n\(\\hat\{y\}^\{a\}\_\{\\mu\}\)\_\{a=1\}^\{n\}is a normalized sum of i\.i\.d\. bounded random variables and becomes jointly Gaussian at largeNvN\_\{v\}by the central limit theorem,

\(y^μa\)a→d𝒩⁡\(0,Q\),Qa​b=qa​b≔1Nv​∑isia​sib,qa​a=1,\(\\hat\{y\}^\{a\}\_\{\\mu\}\)\_\{a\}\\xrightarrow\{\\mathrm\{d\}\}\\mathcal\{N\}\(0,Q\),\\qquad Q\_\{ab\}=q\_\{ab\}\\coloneq\\frac\{1\}\{N\_\{v\}\}\\sum\_\{i\}s\_\{i\}^\{a\}s\_\{i\}^\{b\},\\qquad q\_\{aa\}=1,\(109\)with the replica\-overlap matrix of the spins as its covariance\. Sinceε→0\\varepsilon\\to 0fork\>2k\>2, the per\-pattern average in Eq\. \([107](https://arxiv.org/html/2609.10976#A2.E107)\) is evaluated by the cumulant expansion

log⁡𝔼ξμ​\[eε​∑a\(y^a\)k\]=ε​∑a𝔼⁡\[\(y^a\)k\]\+ε22​∑a,bCov⁡\(\(y^a\)k,\(y^b\)k\)\+O⁡\(ε3\)\.\\log\\mathbb\{E\}\_\{\\xi\_\{\\mu\}\}\\left\[e^\{\\varepsilon\\sum\_\{a\}\(\\hat\{y\}^\{a\}\)^\{k\}\}\\right\]=\\varepsilon\\sum\_\{a\}\\mathbb\{E\}\\left\[\(\\hat\{y\}^\{a\}\)^\{k\}\\right\]\+\\frac\{\\varepsilon^\{2\}\}\{2\}\\sum\_\{a,b\}\\mathrm\{Cov\}\\big\(\(\\hat\{y\}^\{a\}\)^\{k\},\(\\hat\{y\}^\{b\}\)^\{k\}\\big\)\+O\(\\varepsilon^\{3\}\)\.\(110\)The first term is independent of the spin configurations: all single\-replica cumulants ofy^a\\hat\{y\}^\{a\}are spin\-independent, becauseκj\(y^a\)=Nv−j/2κj\(ξ\)∑i\(sia\)j\\kappa\_\{j\}\(\\hat\{y\}^\{a\}\)=N\_\{v\}^\{\-j/2\}\\kappa\_\{j\}\(\\xi\)\\sum\_\{i\}\(s\_\{i\}^\{a\}\)^\{j\}vanishes for oddjj\(symmetry ofξ\\xi\) and reduces toNv1−j/2​κj​\(ξ\)N\_\{v\}^\{1\-j/2\}\\kappa\_\{j\}\(\\xi\)for evenjj\(si2=1s\_\{i\}^\{2\}=1\)\. It therefore contributes a state\-independent constant \(superextensive, of orderNvk/2N\_\{v\}^\{k/2\}after summing over the patterns, but a pure constant\) relative to which we define the free energy \(see also Remark 2 below\)\. The second term is governed by the covariance kernel

Φk​\(q\)≔Cov⁡\(Xk,Yk\),\(X,Y\)​standard Gaussian with correlation​𝔼​\[X​Y\]=q,\\Phi\_\{k\}\(q\)\\coloneq\\mathrm\{Cov\}\\big\(X^\{k\},Y^\{k\}\\big\),\\qquad\(X,Y\)~\\text\{standard Gaussian with correlation \}\\mathbb\{E\}\\left\[XY\\right\]=q,\(111\)evaluated atq=qa​bq=q\_\{ab\}\. Note thatΦk​\(0\)=0\\Phi\_\{k\}\(0\)=0, and the diagonal termsa=ba=bcontribute theqq\-independent constantn​Φk​\(1\)n\\Phi\_\{k\}\(1\), which we drop\. Summing over theNh−1≃αk​Nvk−1N\_\{h\}\-1\\simeq\\alpha\_\{k\}N\_\{v\}^\{k\-1\}patterns, the noise sector contributes the extensive action

\(Nh−1\)​ε22​∑a≠bΦk​\(qa​b\)=αk​βk22​k2​Nv​∑a≠bΦk​\(qa​b\),\(N\_\{h\}\-1\)\\,\\frac\{\\varepsilon^\{2\}\}\{2\}\\sum\_\{a\\neq b\}\\Phi\_\{k\}\(q\_\{ab\}\)=\\frac\{\\alpha\_\{k\}\\beta\_\{k\}^\{2\}\}\{2k^\{2\}\}N\_\{v\}\\sum\_\{a\\neq b\}\\Phi\_\{k\}\(q\_\{ab\}\),\(112\)while the third and higher cumulants are subextensive for evenk≥4k\\geq 4\(Remark 2\)\. In this representation the collapse of the noise\-sector order parameters is manifest: the covarianceRa​bR^\{ab\}of Eq\. \([87](https://arxiv.org/html/2609.10976#A2.E87)\), evaluated on the enslaved modesmμ=m^μ​\(s\)m\_\{\\mu\}=\\hat\{m\}\_\{\\mu\}\(s\), self\-averages by the law of large numbers to

Ra​b=Nv−\(k−1\)​∑μ≥2\(y^μa\)k−1​\(y^μb\)k−1⟶αk​𝔼​\[Xk−1​Yk−1\]\|q=qa​b=αk​ℳk​\(qa​b\),R^\{ab\}=N\_\{v\}^\{\-\(k\-1\)\}\\sum\_\{\\mu\\geq 2\}\(\\hat\{y\}^\{a\}\_\{\\mu\}\)^\{k\-1\}\(\\hat\{y\}^\{b\}\_\{\\mu\}\)^\{k\-1\}\\longrightarrow\\alpha\_\{k\}\\mathbb\{E\}\\left\[X^\{k\-1\}Y^\{k\-1\}\\right\]\\Big\|\_\{q=q\_\{ab\}\}=\\alpha\_\{k\}\\mathcal\{M\}\_\{k\}\(q\_\{ab\}\),\(113\)so that, in contrast to thek=2k=2computation, the pair\(R,R^\)\(R,\\hat\{R\}\)carries no independent degrees of freedom: the crosstalk noise is entirely slaved to the spin overlapqa​bq\_\{ab\}\. In particular the diagonal element is finite,Ra​a→αk​ℳk​\(1\)=αk​\(2​k−3\)\!\!R^\{aa\}\\to\\alpha\_\{k\}\\mathcal\{M\}\_\{k\}\(1\)=\\alpha\_\{k\}\(2k\-3\)\!\!, in contrast with the divergence Eq\. \([105](https://arxiv.org/html/2609.10976#A2.E105)\) of the continuous model\.

The remaining steps are standard\. We introduce the condensed overlap and the spin overlaps with their conjugates,

1=∫∏ad​ma​δ​\(Nv​ma−∑isia\),1=∫∏a<bd​qa​b​δ​\(Nv​qa​b−∑isia​sib\),1=\\int\\prod\_\{a\}dm^\{a\}\\,\\delta\\Big\(N\_\{v\}m^\{a\}\-\\sum\_\{i\}s\_\{i\}^\{a\}\\Big\),\\quad 1=\\int\\prod\_\{a<b\}dq\_\{ab\}\\,\\delta\\Big\(N\_\{v\}q\_\{ab\}\-\\sum\_\{i\}s\_\{i\}^\{a\}s\_\{i\}^\{b\}\\Big\),\(114\)represented with conjugate variablesm~a\\tilde\{m\}^\{a\}andq^a​b\\hat\{q\}\_\{ab\}as in Eq\. \([89](https://arxiv.org/html/2609.10976#A2.E89)\), upon which the spin trace factorizes over the sites\. Under the RS ansatzma=mm^\{a\}=m,qa​b=qq\_\{ab\}=q,q^a​b=q^\\hat\{q\}\_\{ab\}=\\hat\{q\}\(a<ba<b\), the conjugatem~a=m~\\tilde\{m\}^\{a\}=\\tilde\{m\}is eliminated by its saddlem~=βk​mk−1\\tilde\{m\}=\\beta\_\{k\}m^\{k\-1\}, and the single\-site trace gives the familiar

limn→01n​log⁡Trs​exp​\[q^2​\(∑asa\)2−n​q^2\+m~​∑asa\]=−q^2\+∫D​z​log⁡2​cosh⁡\(βk​mk−1\+q^​z\)\.\\lim\_\{n\\to 0\}\\frac\{1\}\{n\}\\log\\mathrm\{Tr\}\_\{s\}\\exp\\bigg\[\\frac\{\\hat\{q\}\}\{2\}\\Big\(\\sum\_\{a\}s^\{a\}\\Big\)^\{2\}\-\\frac\{n\\hat\{q\}\}\{2\}\+\\tilde\{m\}\\sum\_\{a\}s^\{a\}\\bigg\]=\-\\frac\{\\hat\{q\}\}\{2\}\+\\int Dz\\log 2\\cosh\\big\(\\beta\_\{k\}m^\{k\-1\}\+\\sqrt\{\\hat\{q\}\}\\,z\\big\)\.\(115\)Stationarity with respect toq^\\hat\{q\}returnsq=∫D​z​tanh2⁡\(βk​mk−1\+q^​z\)q=\\int Dz\\tanh^\{2\}\(\\beta\_\{k\}m^\{k\-1\}\+\\sqrt\{\\hat\{q\}\}z\), while stationarity with respect toqqties the conjugate to the noise kernel:

q^=αk​βk2k2​Φk′​\(q\)\.\\hat\{q\}=\\frac\{\\alpha\_\{k\}\\beta\_\{k\}^\{2\}\}\{k^\{2\}\}\\Phi\_\{k\}^\{\\prime\}\(q\)\.\(116\)The derivative of the kernel is evaluated by Price’s theorem\[[52](https://arxiv.org/html/2609.10976#bib.bib52)\]: for jointly Gaussian\(X,Y\)\(X,Y\)with correlationqq,

∂∂q​𝔼​\[Xk​Yk\]=k2​𝔼​\[Xk−1​Yk−1\]\.\\frac\{\\partial\}\{\\partial q\}\\mathbb\{E\}\\left\[X^\{k\}Y^\{k\}\\right\]=k^\{2\}\\,\\mathbb\{E\}\\left\[X^\{k\-1\}Y^\{k\-1\}\\right\]\.\(117\)A short proof follows from the Hermite expansion: writingxm=∑jcj\(m\)​Hej​\(x\)x^\{m\}=\\sum\_\{j\}c^\{\(m\)\}\_\{j\}\\mathrm\{He\}\_\{j\}\(x\)with the probabilists’ Hermite polynomialsHej\\mathrm\{He\}\_\{j\}and using the orthogonality𝔼⁡\[Hei​\(X\)​Hej​\(Y\)\]=δi​j​j\!​qj\\mathbb\{E\}\\left\[\\mathrm\{He\}\_\{i\}\(X\)\\mathrm\{He\}\_\{j\}\(Y\)\\right\]=\\delta\_\{ij\}\\,j\!\\,q^\{j\}, one has𝔼⁡\[Xm​Ym\]=∑j\(cj\(m\)\)2​j\!​qj\\mathbb\{E\}\\left\[X^\{m\}Y^\{m\}\\right\]=\\sum\_\{j\}\(c^\{\(m\)\}\_\{j\}\)^\{2\}j\!\\,q^\{j\}\. Differentiating term by term and usingk​cj−1\(k−1\)=j​cj\(k\)kc^\{\(k\-1\)\}\_\{j\-1\}=jc^\{\(k\)\}\_\{j\}, which is the Hermite\-coefficient form ofdd​x​xk=k​xk−1\\frac\{d\}\{dx\}x^\{k\}=kx^\{k\-1\}together withHej′=j​Hej−1\\mathrm\{He\}\_\{j\}^\{\\prime\}=j\\mathrm\{He\}\_\{j\-1\}, yields Eq\. \([117](https://arxiv.org/html/2609.10976#A2.E117)\)\. Combining Eq\. \([116](https://arxiv.org/html/2609.10976#A2.E116)\) and Eq\. \([117](https://arxiv.org/html/2609.10976#A2.E117)\) and defining the noise parameter as in Eq\. \([113](https://arxiv.org/html/2609.10976#A2.E113)\),

r≔ℳk​\(q\)=𝔼⁡\[Xk−1​Yk−1\],q^=αk​βk2​ℳk​\(q\)=αk​βk2​r,r\\coloneq\\mathcal\{M\}\_\{k\}\(q\)=\\mathbb\{E\}\\left\[X^\{k\-1\}Y^\{k\-1\}\\right\],\\qquad\\hat\{q\}=\\alpha\_\{k\}\\beta\_\{k\}^\{2\}\\,\\mathcal\{M\}\_\{k\}\(q\)=\\alpha\_\{k\}\\beta\_\{k\}^\{2\}r,\(118\)the local field becomesβk​mk−1\+q^​z=βk​\(mk−1\+αk​r​z\)\\beta\_\{k\}m^\{k\-1\}\+\\sqrt\{\\hat\{q\}\}z=\\beta\_\{k\}\(m^\{k\-1\}\+\\sqrt\{\\alpha\_\{k\}r\}z\), matching the main text\. Assembling all the terms atn→0n\\to 0, theΦk\\Phi\_\{k\}contribution enters the free energy asαk​βk22​k2​Φk​\(q\)\\frac\{\\alpha\_\{k\}\\beta\_\{k\}^\{2\}\}\{2k^\{2\}\}\\Phi\_\{k\}\(q\), and sinceΦk​\(0\)=0\\Phi\_\{k\}\(0\)=0andΦk′=k2​ℳk\\Phi\_\{k\}^\{\\prime\}=k^\{2\}\\mathcal\{M\}\_\{k\},

αk​βk22​k2​Φk​\(q\)=αk​βk22​∫0qℳk​\(s\)​𝑑s=αk2​Ψk​\(q\),\\frac\{\\alpha\_\{k\}\\beta\_\{k\}^\{2\}\}\{2k^\{2\}\}\\Phi\_\{k\}\(q\)=\\frac\{\\alpha\_\{k\}\\beta\_\{k\}^\{2\}\}\{2\}\\int\_\{0\}^\{q\}\\mathcal\{M\}\_\{k\}\(s\)\\,ds=\\frac\{\\alpha\_\{k\}\}\{2\}\\Psi\_\{k\}\(q\),\(119\)which is precisely the noise entropic term of Eq\. \([21](https://arxiv.org/html/2609.10976#S3.E21)\) fork\>2k\>2\. The termsγk​mk\\gamma\_\{k\}m^\{k\}\(fromβkk​mk−m~​m\\frac\{\\beta\_\{k\}\}\{k\}m^\{k\}\-\\tilde\{m\}matm~=βk​mk−1\\tilde\{m\}=\\beta\_\{k\}m^\{k\-1\}\) andαk​βk22​r​\(1−q\)\\frac\{\\alpha\_\{k\}\\beta\_\{k\}^\{2\}\}\{2\}r\(1\-q\)\(fromq^2​\(1−q\)\\frac\{\\hat\{q\}\}\{2\}\(1\-q\)\) complete the RS free energy, up to the additive constant−αk​βk22​k2​Φk​\(1\)\-\\frac\{\\alpha\_\{k\}\\beta\_\{k\}^\{2\}\}\{2k^\{2\}\}\\Phi\_\{k\}\(1\)and the state\-independent constants discussed above\. The equations of state Eq\. \([23](https://arxiv.org/html/2609.10976#S3.E23)\)–Eq\. \([25](https://arxiv.org/html/2609.10976#S3.E25)\) then follow from stationarity, as verified directly in the main text\.

#### B\.2\.5Consistency checks and remarks

Remark 1: the Onsager reaction field\.The two cases ofℳk\\mathcal\{M\}\_\{k\}can be cross\-checked by a cavity argument that does not rely on replicas\. In the effective spin model Eq\. \([106](https://arxiv.org/html/2609.10976#A2.E106)\), adding a single non\-condensed modeμ\\mushifts the local field on the spinsis\_\{i\}byβk​ξi​μ\(v,h\)​m^μk−1\\beta\_\{k\}\\xi^\{\(v,h\)\}\_\{i\\mu\}\\hat\{m\}\_\{\\mu\}^\{k\-1\}, and hence its thermal average byδ⁡⟨si⟩=βk​\(1−⟨si⟩2\)​ξi​μ\(v,h\)​m^μk−1\\delta\\left\\langle s\_\{i\}\\right\\rangle=\\beta\_\{k\}\(1\-\\left\\langle s\_\{i\}\\right\\rangle^\{2\}\)\\,\\xi^\{\(v,h\)\}\_\{i\\mu\}\\hat\{m\}\_\{\\mu\}^\{k\-1\}in linear response\. Feeding this back into the overlap of the same mode gives the self\-feedback \(Onsager reaction\)

δm^μ=1Nv∑iξμ​i\(h,v\)δ⟨si⟩=βk\(1−q\)m^μk−1,i\.e\.δy^μ=βk\(1−q\)Nv−\(k−2\)/2y^μk−1\.\\delta\\hat\{m\}\_\{\\mu\}=\\frac\{1\}\{N\_\{v\}\}\\sum\_\{i\}\\xi^\{\(h,v\)\}\_\{\\mu i\}\\,\\delta\\left\\langle s\_\{i\}\\right\\rangle=\\beta\_\{k\}\(1\-q\)\\,\\hat\{m\}\_\{\\mu\}^\{k\-1\},\\qquad\\text\{i\.e\.\}\\qquad\\delta\\hat\{y\}\_\{\\mu\}=\\beta\_\{k\}\(1\-q\)\\,N\_\{v\}^\{\-\(k\-2\)/2\}\\,\\hat\{y\}\_\{\\mu\}^\{k\-1\}\.\(120\)Fork=2k=2the feedback is marginal,O⁡\(1\)O\(1\), and must be resummed to all orders: with the frozen cavity fieldq​z\\sqrt\{q\}z, the self\-consistent solution is⟨y^⟩z=q​z/\(1−βk​\(1−q\)\)\\left\\langle\\hat\{y\}\\right\\rangle\_\{z\}=\\sqrt\{q\}z/\(1\-\\beta\_\{k\}\(1\-q\)\), so thatr=∫D​z​⟨y^⟩z2=q/\(1−βk​\(1−q\)\)2=ℳ2​\(q\)r=\\int Dz\\left\\langle\\hat\{y\}\\right\\rangle\_\{z\}^\{2\}=q/\(1\-\\beta\_\{k\}\(1\-q\)\)^\{2\}=\\mathcal\{M\}\_\{2\}\(q\)\. The geometric resummation is the origin of the AGS denominator, in agreement with Appendix[B\.2\.2](https://arxiv.org/html/2609.10976#A2.SS2.SSS2)\. Fork\>2k\>2the feedback vanishes in the thermodynamic limit, the non\-condensed overlaps remain bare Gaussian variables, andr=ℳk​\(q\)r=\\mathcal\{M\}\_\{k\}\(q\)is their bare moment, in agreement with Appendix[B\.2\.4](https://arxiv.org/html/2609.10976#A2.SS2.SSS4)\. This is the content of the scaling statement in the main text\.

Remark 2: validity of the cumulant expansion\.The truncation of Eq\. \([110](https://arxiv.org/html/2609.10976#A2.E110)\) at second order is controlled as follows\. \(i\) The third cumulant contributesO⁡\(ε3\)O\(\\varepsilon^\{3\}\)per pattern, henceO⁡\(Nv2−k/2\)O\(N\_\{v\}^\{2\-k/2\}\)in total after multiplication byNhN\_\{h\}, which is subextensive fork≥4k\\geq 4\. Sincekkis even in our setting, this covers allk\>2k\>2, and higher cumulants are smaller still\. \(ii\) The corrections to the joint Gaussianity of\(y^a\)a\(\\hat\{y\}^\{a\}\)\_\{a\}are of relative orderNv−1N\_\{v\}^\{\-1\}\(Edgeworth\), and, for±1\\pm 1spins and symmetrically distributed patterns, the single\-replica corrections are exactly state\-independent, as noted below Eq\. \([110](https://arxiv.org/html/2609.10976#A2.E110)\)\. The state\-dependent cross\-replica corrections contribute only atO⁡\(1\)O\(1\)in total\. \(iii\) The first cumulant produces a state\-independent constant of orderNvk/2N\_\{v\}^\{k/2\}, which is superextensive but common to all configurations and all phases, and drops out of every order\-parameter equation and free energy difference\. \(iv\) The evenness ofkkis used twice: it bounds the hidden potential from below, and it guarantees the vanishing of the odd pattern cumulants used in \(i\) and \(ii\)\.

Remark 3: closed form of the noise kernel\.The Hermite expansion provides the explicit polynomial form ofℳk\\mathcal\{M\}\_\{k\}\. Writingxk−1=∑jck,j​Hej​\(x\)x^\{k\-1\}=\\sum\_\{j\}c\_\{k,j\}\\mathrm\{He\}\_\{j\}\(x\)with

ck,j=\(k−1\)\!j\!​2\(k−1−j\)/2​\(k−1−j2\)\!,j≡k−1​\(mod​2\),0≤j≤k−1,c\_\{k,j\}=\\frac\{\(k\-1\)\!\}\{j\!\\,2^\{\(k\-1\-j\)/2\}\\left\(\\frac\{k\-1\-j\}\{2\}\\right\)\!\},\\qquad j\\equiv k\-1~\(\\mathrm\{mod\}~2\),\\quad 0\\leq j\\leq k\-1,\(121\)the orthogonality𝔼⁡\[Hei​\(X\)​Hej​\(Y\)\]=δi​j​j\!​qj\\mathbb\{E\}\\left\[\\mathrm\{He\}\_\{i\}\(X\)\\mathrm\{He\}\_\{j\}\(Y\)\\right\]=\\delta\_\{ij\}\\,j\!\\,q^\{j\}gives

ℳk​\(q\)=∑jck,j2​j\!​qj,ℳk​\(1\)=𝔼⁡\[X2​\(k−1\)\]=\(2​k−3\)\!\!,\\mathcal\{M\}\_\{k\}\(q\)=\\sum\_\{j\}c\_\{k,j\}^\{2\}\\,j\!\\,q^\{j\},\\qquad\\mathcal\{M\}\_\{k\}\(1\)=\\mathbb\{E\}\\left\[X^\{2\(k\-1\)\}\\right\]=\(2k\-3\)\!\!,\(122\)a polynomial with positive coefficients, for instanceℳ4​\(q\)=9​q\+6​q3\\mathcal\{M\}\_\{4\}\(q\)=9q\+6q^\{3\}andℳ6​\(q\)=225​q\+600​q3\+120​q5\\mathcal\{M\}\_\{6\}\(q\)=225q\+600q^\{3\}\+120q^\{5\}\. The same polynomial kernel appears as the noise covariance in the dynamical mean\-field theory of dense associative memories with the polynomial \(Krotov–Hopfield\-type\) energy\[[53](https://arxiv.org/html/2609.10976#bib.bib53),[54](https://arxiv.org/html/2609.10976#bib.bib54)\], with the static correspondence between the equal\-time correlation and the replica overlap\. The monomial kernel∝qk−1\\propto q^\{k\-1\}familiar frompp\-spin models arises instead for the diagonal\-free \(Abbott–Arian\-type\) variant of the interaction\[[8](https://arxiv.org/html/2609.10976#bib.bib8),[54](https://arxiv.org/html/2609.10976#bib.bib54)\], which is a different model from ours\.

Remark 4: the zero\-temperature limit\.The zero\-temperature equations quoted in Sec\.[III\.1](https://arxiv.org/html/2609.10976#S3.SS1)follow from Eq\. \([23](https://arxiv.org/html/2609.10976#S3.E23)\)–Eq\. \([25](https://arxiv.org/html/2609.10976#S3.E25)\) by standard manipulations\. Asβk→∞\\beta\_\{k\}\\to\\inftyat fixedC=βk​\(1−q\)C=\\beta\_\{k\}\(1\-q\), the hyperbolic tangent in Eq\. \([23](https://arxiv.org/html/2609.10976#S3.E23)\) reduces to the sign of its argument, and∫D​z​sgn⁡\(mk−1\+αk​r​z\)=\{erf\}⁡\(mk−1/2​αk​r\)\\int Dz\\,\\sgn\(m^\{k\-1\}\+\\sqrt\{\\alpha\_\{k\}r\}\\,z\)=\\erf\(m^\{k\-1\}/\\sqrt\{2\\alpha\_\{k\}r\}\)yields Eq\. \([27](https://arxiv.org/html/2609.10976#S3.E27)\)\. For the frozen response, one combines1−q=∫D​z​\{sech\}2​βk​\(mk−1\+αk​r​z\)1\-q=\\int Dz\\sech^\{2\}\\beta\_\{k\}\(m^\{k\-1\}\+\\sqrt\{\\alpha\_\{k\}r\}\\,z\)withβk​\{sech\}2⁡\(βk​x\)→2​δ​\(x\)\\beta\_\{k\}\\sech^\{2\}\(\\beta\_\{k\}x\)\\to 2\\delta\(x\)\. The delta function picks up the Gaussian density at the zeroz0=−mk−1/αk​rz\_\{0\}=\-m^\{k\-1\}/\\sqrt\{\\alpha\_\{k\}r\}of the local field, which gives Eq\. \([28](https://arxiv.org/html/2609.10976#S3.E28)\)\. Finally,q→1q\\to 1at fixedCCreduces the crosstalk moment Eq\. \([26](https://arxiv.org/html/2609.10976#S3.E26)\) to Eq\. \([29](https://arxiv.org/html/2609.10976#S3.E29)\): fork=2k=2the denominator survives as1−βk​\(1−q\)→1−C1\-\\beta\_\{k\}\(1\-q\)\\to 1\-C, while fork\>2k\>2the polynomial is simply evaluated atq=1q=1\.

### B\.3Model C replica analysis

In this appendix we derive the RS free energy Eq\. \([41](https://arxiv.org/html/2609.10976#S3.E41)\) and the equations of state Eq\. \([42](https://arxiv.org/html/2609.10976#S3.E42)\)–Eq\. \([43](https://arxiv.org/html/2609.10976#S3.E43)\) of Model C\. The hidden sector of Model C coincides with that of Model A: after the conjugate insertion, the non\-condensed hidden modes factorize into the same single\-mode integral Eq\. \([91](https://arxiv.org/html/2609.10976#A2.E91)\), so the dichotomy established in Appendix[B\.2](https://arxiv.org/html/2609.10976#A2.SS2)carries over: the hidden sector closes exactly fork=2k=2and admits no consistent evaluation fork\>2k\>2, where the model is defined by the adiabatic reduction\. We exploit this from the outset by working with the formulation in which the hidden variables are eliminated in favor of the pattern overlaps\. This formulation is*exact*fork=2k=2and*definitional*fork\>2k\>2, and it leads directly to the two\-parameter free energy quoted in the main text\. The computation that retains the hidden variables explicitly is presented in Appendix[B\.3\.5](https://arxiv.org/html/2609.10976#A2.SS3.SSS5)for completeness\. What is genuinely new compared with Model A is the visible sector: the spherical trace is Gaussian and is carried out exactly, so that no single\-site integral remains in the final expressions\.

#### B\.3\.1Overlap formulation and replica setup

We start from the partition function Eq\. \([40](https://arxiv.org/html/2609.10976#S3.E40)\)\. The flat radial direction of the visible variables has already been removed by the sphere\-restricted definition of Eq\. \([14](https://arxiv.org/html/2609.10976#S3.E14)\), and no further regularization is needed\. The basic objects are the pattern overlaps of a visible configurationx∈S\\mathrm\{x\}\\in S,

m^μ​\(x\)≔1Nv​∑iξμ​i\(h,v\)​xi\.\\hat\{m\}\_\{\\mu\}\(\\mathrm\{x\}\)\\coloneq\\frac\{1\}\{N\_\{v\}\}\\sum\_\{i\}\\xi^\{\(h,v\)\}\_\{\\mu i\}\\mathrm\{x\}\_\{i\}\.\(123\)Fork=2k=2the hidden variables can be eliminated exactly:γ2=βk/2\\gamma\_\{2\}=\\beta\_\{k\}/2, eachmμm\_\{\\mu\}appears quadratically in Eq\. \([40](https://arxiv.org/html/2609.10976#S3.E40)\), and the Gaussian integral gives, up to an overall constant,

∫d​mμ​exp⁡\(−βk2​Nv​mμ2\+βk​Nv​m^μ​\(x\)​mμ\)∝exp⁡\(βk2​Nv​m^μ​\(x\)2\)\.\\int dm\_\{\\mu\}\\exp\\Big\(\-\\frac\{\\beta\_\{k\}\}\{2\}N\_\{v\}m\_\{\\mu\}^\{2\}\+\\beta\_\{k\}N\_\{v\}\\hat\{m\}\_\{\\mu\}\(\\mathrm\{x\}\)\\,m\_\{\\mu\}\\Big\)\\propto\\exp\\Big\(\\frac\{\\beta\_\{k\}\}\{2\}N\_\{v\}\\hat\{m\}\_\{\\mu\}\(\\mathrm\{x\}\)^\{2\}\\Big\)\.\(124\)Fork\>2k\>2the direct evaluation of the hidden sector fails in exactly the manner analyzed in Appendix[B\.2\.3](https://arxiv.org/html/2609.10976#A2.SS2.SSS3): the non\-condensed modes factorize into the single\-mode integralℐk​\(R^\)\\mathcal\{I\}\_\{k\}\(\\hat\{R\}\)of Eq\. \([91](https://arxiv.org/html/2609.10976#A2.E91)\), which admits no extensive saddle\-point structure, and the crosstalk variance of the continuous modes suffers the runaway Eq\. \([105](https://arxiv.org/html/2609.10976#A2.E105)\)\. We return to this formulation in Appendix[B\.3\.5](https://arxiv.org/html/2609.10976#A2.SS3.SSS5)\. As for Model A, fork\>2k\>2we therefore*define*Model C through the adiabatic reduction of Sec\.[II\.2](https://arxiv.org/html/2609.10976#S2.SS2)\. The enslaved valuemμ=m^μ​\(x\)m\_\{\\mu\}=\\hat\{m\}\_\{\\mu\}\(\\mathrm\{x\}\)follows from the same per\-mode stationarity algebra as in Eq\. \([106](https://arxiv.org/html/2609.10976#A2.E106)\), and we work with

ZξC,ad≔∫Sd​Ω​\(x\)​exp⁡\[βkk​Nv​∑μm^μ​\(x\)k\],Z^\{\\mathrm\{C,ad\}\}\_\{\\xi\}\\coloneq\\int\_\{S\}d\\Omega\(\\mathrm\{x\}\)\\,\\exp\\bigg\[\\frac\{\\beta\_\{k\}\}\{k\}N\_\{v\}\\sum\_\{\\mu\}\\hat\{m\}\_\{\\mu\}\(\\mathrm\{x\}\)^\{k\}\\bigg\],\(125\)which is exact fork=2k=2by Eq\. \([124](https://arxiv.org/html/2609.10976#A2.E124)\) and is the adiabatic definition fork\>2k\>2, so that the analysis below is uniform inkk\.

Replicating Eq\. \([125](https://arxiv.org/html/2609.10976#A2.E125)\) and averaging over the patterns, we use the rotational invariance of the spherical ensemble to rotate the condensed pattern toξ1\(h,v\)=\(1,…,1\)\\xi^\{\(h,v\)\}\_\{1\}=\(1,\\ldots,1\)\. This gauge choice is exact at finiteNvN\_\{v\}\(Remark 1 below\), and the condensed term becomes a function of the visible magnetizations,

m^1​\(xa\)=1Nv​∑ixia≕ma\.\\hat\{m\}\_\{1\}\(\\mathrm\{x\}^\{a\}\)=\\frac\{1\}\{N\_\{v\}\}\\sum\_\{i\}\\mathrm\{x\}^\{a\}\_\{i\}\\eqcolon m^\{a\}\.\(126\)For the non\-condensed patterns we introduce the rescaled overlaps

y^μa≔1Nv​∑iξμ​i\(h,v\)​xia=Nv​m^μ​\(xa\)\.\\hat\{y\}^\{a\}\_\{\\mu\}\\coloneq\\frac\{1\}\{\\sqrt\{N\_\{v\}\}\}\\sum\_\{i\}\\xi^\{\(h,v\)\}\_\{\\mu i\}\\mathrm\{x\}^\{a\}\_\{i\}=\\sqrt\{N\_\{v\}\}\\,\\hat\{m\}\_\{\\mu\}\(\\mathrm\{x\}^\{a\}\)\.\(127\)Since the spherical ensemble has𝔼⁡\[ξi​μ\(v,h\)​ξj​μ\(v,h\)\]=δi​j\\mathbb\{E\}\\left\[\\xi^\{\(v,h\)\}\_\{i\\mu\}\\xi^\{\(v,h\)\}\_\{j\\mu\}\\right\]=\\delta\_\{ij\}, the covariance of the overlaps is exactly the replica\-overlap matrix of the visible variables,

𝔼ξμ​\[y^μa​y^μb\]=1Nv​∑ixia​xib≕qa​b,qa​a=1,\\mathbb\{E\}\_\{\\xi\_\{\\mu\}\}\\left\[\\hat\{y\}^\{a\}\_\{\\mu\}\\hat\{y\}^\{b\}\_\{\\mu\}\\right\]=\\frac\{1\}\{N\_\{v\}\}\\sum\_\{i\}\\mathrm\{x\}^\{a\}\_\{i\}\\mathrm\{x\}^\{b\}\_\{i\}\\eqcolon q\_\{ab\},\\qquad q\_\{aa\}=1,\(128\)where the diagonal is fixed by the spherical constraint, the counterpart for continuous spins ofsi2=1s\_\{i\}^\{2\}=1in Model A\. At largeNvN\_\{v\}the vector\(y^μa\)a\(\\hat\{y\}^\{a\}\_\{\\mu\}\)\_\{a\}becomes jointly Gaussian with covariance matrixQ=\(qa​b\)Q=\(q\_\{ab\}\)\. Moreover, by rotational invariance the exact per\-pattern average depends on the visible configurations only throughQQ, so that all finite\-NvN\_\{v\}corrections are automatically functions of the order parameters \(Remark 1\)\.

#### B\.3\.2Noise sector

The average over one non\-condensed pattern produces the noise factor

𝔼ξμ​\[exp⁡\(ε​∑a\(y^μa\)k\)\],ε=βkk​Nv1−k/2,\\mathbb\{E\}\_\{\\xi\_\{\\mu\}\}\\left\[\\exp\\bigg\(\\varepsilon\\sum\_\{a\}\(\\hat\{y\}^\{a\}\_\{\\mu\}\)^\{k\}\\bigg\)\\right\],\\qquad\\varepsilon=\\frac\{\\beta\_\{k\}\}\{k\}N\_\{v\}^\{1\-k/2\},\(129\)of the same form as in Eq\. \([107](https://arxiv.org/html/2609.10976#A2.E107)\)\.

Fork=2k=2we haveε=βk/2=O⁡\(1\)\\varepsilon=\\beta\_\{k\}/2=O\(1\), and the Gaussian asymptotics of\(y^μa\)a\(\\hat\{y\}^\{a\}\_\{\\mu\}\)\_\{a\}gives the closed form

𝔼ξμ\[eβk2​∑a\(y^μa\)2\]→det\(In−βkQ\)−1/2,\\mathbb\{E\}\_\{\\xi\_\{\\mu\}\}\\left\[e^\{\\frac\{\\beta\_\{k\}\}\{2\}\\sum\_\{a\}\(\\hat\{y\}^\{a\}\_\{\\mu\}\)^\{2\}\}\\right\]\\to\\det\\big\(I\_\{n\}\-\\beta\_\{k\}Q\\big\)^\{\-1/2\},\(130\)convergent forβk​λmax​\(Q\)<1\\beta\_\{k\}\\lambda\_\{\\max\}\(Q\)<1, withλmax\\lambda\_\{\\max\}the largest eigenvalue ofQQ\. Under the RS ansatzqa​b=qq\_\{ab\}=q\(a≠ba\\neq b\), the eigenvalues ofQQare1−q1\-qwith degeneracyn−1n\-1and1\+\(n−1\)​q1\+\(n\-1\)q, both of which tend to1−q1\-qin the replica limit, so that

logdet\(In−βkQ\)=\(n−1\)log\(1−βk\(1−q\)\)\+log\(1−βk\(1−q\)−nβkq\)=nΨ2\(q\)\+O\(n2\),\\log\\det\\big\(I\_\{n\}\-\\beta\_\{k\}Q\\big\)=\(n\-1\)\\log\\big\(1\-\\beta\_\{k\}\(1\-q\)\\big\)\+\\log\\big\(1\-\\beta\_\{k\}\(1\-q\)\-n\\beta\_\{k\}q\\big\)=n\\,\\Psi\_\{2\}\(q\)\+O\(n^\{2\}\),\(131\)which is the same algebra as Eq\. \([101](https://arxiv.org/html/2609.10976#A2.E101)\) evaluated at\(r^d,r^\)=\(1,q\)\(\\hat\{r\}\_\{d\},\\hat\{r\}\)=\(1,q\)\. Multiplying by−12​\(Nh−1\)≃−αk​Nv2\-\\frac\{1\}\{2\}\(N\_\{h\}\-1\)\\simeq\-\\frac\{\\alpha\_\{k\}N\_\{v\}\}\{2\}, the noise sector contributes−n​αk​Nv2​Ψ2​\(q\)\-\\frac\{n\\alpha\_\{k\}N\_\{v\}\}\{2\}\\Psi\_\{2\}\(q\)to the replicated exponent, i\.e\.,\+αk2​Ψ2​\(q\)\+\\frac\{\\alpha\_\{k\}\}\{2\}\\Psi\_\{2\}\(q\)toβ​fC\\beta f^\{\\mathrm\{C\}\}\.

Fork\>2k\>2we haveε→0\\varepsilon\\to 0, and Eq\. \([129](https://arxiv.org/html/2609.10976#A2.E129)\) is evaluated by the cumulant expansion Eq\. \([110](https://arxiv.org/html/2609.10976#A2.E110)\), whose structure carries over intact\. The first cumulant is state\-independent because it depends only on the diagonalqa​a=1q\_\{aa\}=1, now enforced by the spherical constraint\. The second cumulant is governed by the same covariance kernelΦk​\(qa​b\)\\Phi\_\{k\}\(q\_\{ab\}\)of Eq\. \([111](https://arxiv.org/html/2609.10976#A2.E111)\)\. The third and higher cumulants are subextensive for evenk≥4k\\geq 4, with the Model C\-specific aspects of the power counting collected in Remark 2\. Summing over theNh−1N\_\{h\}\-1patterns, the noise sector contributes the extensive action Eq\. \([112](https://arxiv.org/html/2609.10976#A2.E112)\), which, by Price’s theorem Eq\. \([117](https://arxiv.org/html/2609.10976#A2.E117)\), integrates toαk2​Ψk​\(q\)\\frac\{\\alpha\_\{k\}\}\{2\}\\Psi\_\{k\}\(q\)in the free energy, exactly as in Appendix[B\.2\.4](https://arxiv.org/html/2609.10976#A2.SS2.SSS4)\.

In both cases, therefore, the noise sector entersβ​fC\\beta f^\{\\mathrm\{C\}\}asαk2​Ψk​\(q\)\\frac\{\\alpha\_\{k\}\}\{2\}\\Psi\_\{k\}\(q\), with the uniform derivativeΨk′​\(q\)=βk2​ℳk​\(q\)\\Psi\_\{k\}^\{\\prime\}\(q\)=\\beta\_\{k\}^\{2\}\\mathcal\{M\}\_\{k\}\(q\)quoted in the main text \(fork=2k=2directly from Eq\. \([131](https://arxiv.org/html/2609.10976#A2.E131)\), fork\>2k\>2from Price’s theorem\)\. Note also that the noise factors depend on the visible configurations only through the overlap matrixQQ, which is precisely the quantity fixed by the conjugate insertions of the next subsection\.

#### B\.3\.3Visible Gaussian sector and the reduced free energy

The remaining trace over the visible variables is organized by introducing the condensed overlaps and the replica overlaps with their conjugates,

1=∫∏ad​ma​δ​\(Nv​ma−∑ixia\),1=∫∏a<bd​qa​b​δ​\(Nv​qa​b−∑ixia​xib\),1=\\int\\prod\_\{a\}dm^\{a\}\\,\\delta\\Big\(N\_\{v\}m^\{a\}\-\\sum\_\{i\}\\mathrm\{x\}^\{a\}\_\{i\}\\Big\),\\qquad 1=\\int\\prod\_\{a<b\}dq\_\{ab\}\\,\\delta\\Big\(N\_\{v\}q\_\{ab\}\-\\sum\_\{i\}\\mathrm\{x\}^\{a\}\_\{i\}\\mathrm\{x\}^\{b\}\_\{i\}\\Big\),\(132\)represented with conjugate variablesm~a\\tilde\{m\}^\{a\}andq^a​b\\hat\{q\}\_\{ab\}as in Eq\. \([89](https://arxiv.org/html/2609.10976#A2.E89)\), together with the spherical constraints

1=∫c−i​∞c\+i​∞∏ad​ua4​π​i​exp⁡\[−ua2​\(∑i\(xia\)2−Nv\)\],c\>0\.1=\\int\_\{c\-\\mathrm\{i\}\\infty\}^\{c\+\\mathrm\{i\}\\infty\}\\prod\_\{a\}\\frac\{du^\{a\}\}\{4\\pi\\mathrm\{i\}\}\\exp\\bigg\[\-\\frac\{u^\{a\}\}\{2\}\\Big\(\\sum\_\{i\}\(\\mathrm\{x\}^\{a\}\_\{i\}\)^\{2\}\-N\_\{v\}\\Big\)\\bigg\],\\qquad c\>0\.\(133\)The multiplieruau^\{a\}is precisely the conjugate of the diagonal overlapqa​aq\_\{aa\}: it absorbs the entire diagonal sector, playing the role of the pair\(rd,r^d\)\(r\_\{d\},\\hat\{r\}\_\{d\}\)in Model A, where we foundr^d=1\\hat\{r\}\_\{d\}=1and a complete cancellation ofrdr\_\{d\}\(Eq\. \([97](https://arxiv.org/html/2609.10976#A2.E97)\) and below\)\. After these insertions the visible integral factorizes over the sites\. Under the RS ansatzma=mm^\{a\}=m,m~a=m~\\tilde\{m\}^\{a\}=\\tilde\{m\},ua=uu^\{a\}=u,qa​b=qq\_\{ab\}=q,q^a​b=q^\\hat\{q\}\_\{ab\}=\\hat\{q\}, the single\-site measure is annn\-dimensional Gaussian\. Decoupling the replica coupling with a frozen Gaussian field,eq^​∑a<bxa​xb=e−n​q^2​∫D​z​eq^​z​∑axae^\{\\hat\{q\}\\sum\_\{a<b\}\\mathrm\{x\}^\{a\}\\mathrm\{x\}^\{b\}\}=e^\{\-\\frac\{n\\hat\{q\}\}\{2\}\}\\int Dz\\,e^\{\\sqrt\{\\hat\{q\}\}z\\sum\_\{a\}\\mathrm\{x\}^\{a\}\}, each site contributes

∫D​z​∏a∫d​xa​exp⁡\[−D2​\(xa\)2\+\(m~\+q^​z\)​xa\],D≔u\+q^,\\int Dz\\prod\_\{a\}\\int d\\mathrm\{x\}^\{a\}\\exp\\Big\[\-\\frac\{D\}\{2\}\(\\mathrm\{x\}^\{a\}\)^\{2\}\+\\big\(\\tilde\{m\}\+\\sqrt\{\\hat\{q\}\}\\,z\\big\)\\mathrm\{x\}^\{a\}\\Big\],\\qquad D\\coloneq u\+\\hat\{q\},\(134\)whose logarithm is, expanding to first order innnand dropping constants,

limn→01n​log⁡\(site factor\)=−12​log⁡D\+m~2\+q^2​D\.\\lim\_\{n\\to 0\}\\frac\{1\}\{n\}\\log\(\\text\{site factor\}\)=\-\\frac\{1\}\{2\}\\log D\+\\frac\{\\tilde\{m\}^\{2\}\+\\hat\{q\}\}\{2D\}\.\(135\)Collecting the condensed term, the conjugate terms, the constraint terms, the noise sector of Appendix[B\.3\.2](https://arxiv.org/html/2609.10976#A2.SS3.SSS2), and Eq\. \([135](https://arxiv.org/html/2609.10976#A2.E135)\), we obtain the variational free energy, up to additive constants,

β​fC=−βkk​mk\+m~​m−q^​q2\+αk2​Ψk​\(q\)−D−q^2\+12​log⁡D−m~2\+q^2​D,\\beta f^\{\\mathrm\{C\}\}=\-\\frac\{\\beta\_\{k\}\}\{k\}m^\{k\}\+\\tilde\{m\}m\-\\frac\{\\hat\{q\}q\}\{2\}\+\\frac\{\\alpha\_\{k\}\}\{2\}\\Psi\_\{k\}\(q\)\-\\frac\{D\-\\hat\{q\}\}\{2\}\+\\frac\{1\}\{2\}\\log D\-\\frac\{\\tilde\{m\}^\{2\}\+\\hat\{q\}\}\{2D\},\(136\)where we traded the multiplieru=D−q^u=D\-\\hat\{q\}forDD\.

The stationarity conditions of Eq\. \([136](https://arxiv.org/html/2609.10976#A2.E136)\) are elementary\. Variations with respect tom~\\tilde\{m\}andDDgive

m~=D​m,1D\+m~2\+q^D2=1,\\tilde\{m\}=Dm,\\qquad\\frac\{1\}\{D\}\+\\frac\{\\tilde\{m\}^\{2\}\+\\hat\{q\}\}\{D^\{2\}\}=1,\(137\)the latter being the spherical constraint1Nv​∑i⟨xi2⟩=1\\frac\{1\}\{N\_\{v\}\}\\sum\_\{i\}\\left\\langle\\mathrm\{x\}\_\{i\}^\{2\}\\right\\rangle=1, while variation with respect toq^\\hat\{q\}identifies

q=m~2\+q^D2=m2\+q^D2\.q=\\frac\{\\tilde\{m\}^\{2\}\+\\hat\{q\}\}\{D^\{2\}\}=m^\{2\}\+\\frac\{\\hat\{q\}\}\{D^\{2\}\}\.\(138\)This exhibitsqqas the Edwards–Anderson order parameter: in the single\-site measure,⟨x⟩=\(m~\+q^​z\)/D\\left\\langle\\mathrm\{x\}\\right\\rangle=\(\\tilde\{m\}\+\\sqrt\{\\hat\{q\}\}z\)/D, so thatq=∫D​z​⟨x⟩2q=\\int Dz\\left\\langle\\mathrm\{x\}\\right\\rangle^\{2\}decomposes into the condensed partm2m^\{2\}and the glassy partq^/D2\\hat\{q\}/D^\{2\}\. Solving Eq\. \([137](https://arxiv.org/html/2609.10976#A2.E137)\)–Eq\. \([138](https://arxiv.org/html/2609.10976#A2.E138)\),

D=11−q,m~=m1−q,q^=q−m2\(1−q\)2,D=\\frac\{1\}\{1\-q\},\\qquad\\tilde\{m\}=\\frac\{m\}\{1\-q\},\\qquad\\hat\{q\}=\\frac\{q\-m^\{2\}\}\{\(1\-q\)^\{2\}\},\(139\)so that1−q=1/D1\-q=1/Dis the single\-site thermal variance left after the condensed and glassy components are removed, andβk​\(1−q\)\\beta\_\{k\}\(1\-q\)is the static susceptibility of the spherical state\. The remaining two conditions are the equations of state\. Variation with respect toqqgives, usingΨk′=βk2​ℳk\\Psi\_\{k\}^\{\\prime\}=\\beta\_\{k\}^\{2\}\\mathcal\{M\}\_\{k\},

q^=αk​βk2​ℳk​\(q\)⟹q−m2\(1−q\)2=αk​βk2​ℳk​\(q\),\\hat\{q\}=\\alpha\_\{k\}\\beta\_\{k\}^\{2\}\\,\\mathcal\{M\}\_\{k\}\(q\)\\qquad\\Longrightarrow\\qquad\\frac\{q\-m^\{2\}\}\{\(1\-q\)^\{2\}\}=\\alpha\_\{k\}\\beta\_\{k\}^\{2\}\\,\\mathcal\{M\}\_\{k\}\(q\),\(140\)which is Eq\. \([43](https://arxiv.org/html/2609.10976#S3.E43)\), and variation with respect tommgivesm~=βk​mk−1\\tilde\{m\}=\\beta\_\{k\}m^\{k\-1\}, which combined withm~=m/\(1−q\)\\tilde\{m\}=m/\(1\-q\)yields the signal equation Eq\. \([42](https://arxiv.org/html/2609.10976#S3.E42)\)\.

Finally, we eliminate the auxiliary variables\. SubstitutingDDandm~\\tilde\{m\}from Eq\. \([139](https://arxiv.org/html/2609.10976#A2.E139)\) into Eq\. \([136](https://arxiv.org/html/2609.10976#A2.E136)\), theq^\\hat\{q\}\-dependent terms cancel identically,

−q^​q2\+q^2−q^2​\(1−q\)=0,\-\\frac\{\\hat\{q\}q\}\{2\}\+\\frac\{\\hat\{q\}\}\{2\}\-\\frac\{\\hat\{q\}\}\{2\}\(1\-q\)=0,\(141\)\(the three terms coming from the conjugate, constraint, and single\-site terms, respectively, so thatq^\\hat\{q\}need not even be substituted\), whilem~​m−m~2/\(2​D\)=m2/\(2​\(1−q\)\)\\tilde\{m\}m\-\\tilde\{m\}^\{2\}/\(2D\)=m^\{2\}/\(2\(1\-q\)\), and we arrive at

β​fC=−βkk​mk\+αk2​Ψk​\(q\)−12​\[log⁡\(1−q\)\+q−m21−q\]−12,\\beta f^\{\\mathrm\{C\}\}=\-\\frac\{\\beta\_\{k\}\}\{k\}m^\{k\}\+\\frac\{\\alpha\_\{k\}\}\{2\}\\Psi\_\{k\}\(q\)\-\\frac\{1\}\{2\}\\left\[\\log\(1\-q\)\+\\frac\{q\-m^\{2\}\}\{1\-q\}\\right\]\-\\frac\{1\}\{2\},\(142\)which is the free energy Eq\. \([41](https://arxiv.org/html/2609.10976#S3.E41)\) of the main text, up to the additive constant−1/2\-1/2\. Because the eliminated variables were removed at their exact stationary points, the stationarity of Eq\. \([142](https://arxiv.org/html/2609.10976#A2.E142)\) in\(m,q\)\(m,q\)reproduces Eq\. \([42](https://arxiv.org/html/2609.10976#S3.E42)\)–Eq\. \([43](https://arxiv.org/html/2609.10976#S3.E43)\), as stated in the main text\.

#### B\.3\.4The casek=2k=2: marginality and zero capacity

Fork=2k=2the two expressions for the conjugate field obtained above,m~=βk​m\\tilde\{m\}=\\beta\_\{k\}mandm~=m/\(1−q\)\\tilde\{m\}=m/\(1\-q\), are compatible only if

\[1−βk​\(1−q\)\]​m=0\.\\big\[1\-\\beta\_\{k\}\(1\-q\)\\big\]\\,m=0\.\(143\)A retrieval state thus requires the marginality conditionβk​\(1−q\)=1\\beta\_\{k\}\(1\-q\)=1, and when it holds the magnitude ofmmis left undetermined at this order: the condensed direction is a flat direction \(zero mode\) of the quadratic spherical energy rather than a genuine minimum\. Atαk=0\\alpha\_\{k\}=0the flatness is lifted by the spherical constraint itself: Eq\. \([140](https://arxiv.org/html/2609.10976#A2.E140)\) givesq=m2q=m^\{2\}, andβk​\(1−m2\)=1\\beta\_\{k\}\(1\-m^\{2\}\)=1determinesm2=1−Tm^\{2\}=1\-TbelowTc=1T\_\{c\}=1, as quoted in the main text\.

The noise sector exhibits the same marginality from a complementary viewpoint\. The exactk=2k=2noise factor Eq\. \([130](https://arxiv.org/html/2609.10976#A2.E130)\) is convergent only forβk​λmax​\(Q\)<1\\beta\_\{k\}\\lambda\_\{\\max\}\(Q\)<1, which under the RS ansatz becomesβk​\(1−q\)<1\\beta\_\{k\}\(1\-q\)<1in the replica limit\. A retrieval state lies precisely at the boundary of this normalizability region, wheredet\(In−βk​Q\)→0\\det\(I\_\{n\}\-\\beta\_\{k\}Q\)\\to 0: the fluctuations of the non\-condensed overlaps soften and condense, which is the standard condensation mechanism of spherical models\[[55](https://arxiv.org/html/2609.10976#bib.bib55),[56](https://arxiv.org/html/2609.10976#bib.bib56)\], and the divergence ofℳ2​\(q\)=q/\(1−βk​\(1−q\)\)2\\mathcal\{M\}\_\{2\}\(q\)=q/\(1\-\\beta\_\{k\}\(1\-q\)\)^\{2\}is its precursor\. Consequently the noise equation Eq\. \([43](https://arxiv.org/html/2609.10976#S3.E43)\) admits no solution withm\>0m\>0at anyαk\>0\\alpha\_\{k\}\>0, includingT=0T=0, and the quadratic spherical model has zero storage capacity, in agreement with the marginal behavior of the spherical Hopfield model\[[20](https://arxiv.org/html/2609.10976#bib.bib20)\]\. There, a retrieval phase is restored by augmenting the Hamiltonian with a quartic term\. Within the classℋ\\mathcal\{H\}, the same lifting of the flat direction is provided by the hidden nonlinearity withk\>2k\>2\.

#### B\.3\.5Direct computation with continuous hidden variables

For completeness, we sketch the computation that retains the hidden variables explicitly, parallel to Appendix[B\.2\.1](https://arxiv.org/html/2609.10976#A2.SS2.SSS1), and show where it connects to the overlap formulation above\. Averaging the replicated Eq\. \([40](https://arxiv.org/html/2609.10976#S3.E40)\) over the non\-condensed patterns gives the noise couplingβk22​∑a,bRa​b​∑ixia​xib\\frac\{\\beta\_\{k\}^\{2\}\}\{2\}\\sum\_\{a,b\}R^\{ab\}\\sum\_\{i\}\\mathrm\{x\}^\{a\}\_\{i\}\\mathrm\{x\}^\{b\}\_\{i\}with the same covariance matrixRa​b=∑μ≥2\(mμa\)k−1​\(mμb\)k−1R^\{ab\}=\\sum\_\{\\mu\\geq 2\}\(m^\{a\}\_\{\\mu\}\)^\{k\-1\}\(m^\{b\}\_\{\\mu\}\)^\{k\-1\}as in Eq\. \([87](https://arxiv.org/html/2609.10976#A2.E87)\)\. Inserting the conjugate pair\(R,R^\)\(R,\\hat\{R\}\)exactly as in Eq\. \([89](https://arxiv.org/html/2609.10976#A2.E89)\), the non\-condensed hidden modes factorize into the single\-mode integralℐk​\(R^\)\\mathcal\{I\}\_\{k\}\(\\hat\{R\}\)of Eq\. \([91](https://arxiv.org/html/2609.10976#A2.E91)\)\. With the spherical constraints Eq\. \([133](https://arxiv.org/html/2609.10976#A2.E133)\) inserted, the visible trace is an unconstrained Gaussian integral,

∫∏a,idxiaexp\[−12∑i𝐱i⊤S𝐱i\+βk∑a,ihak−1xia\]=\(2π\)n​Nv/2\(detS\)−Nv/2exp\(Nv​βk22𝒃⊤S−1𝒃\),\\int\\prod\_\{a,i\}d\\mathrm\{x\}^\{a\}\_\{i\}\\exp\\bigg\[\-\\frac\{1\}\{2\}\\sum\_\{i\}\\bm\{\\mathrm\{x\}\}\_\{i\}^\{\\top\}S\\,\\bm\{\\mathrm\{x\}\}\_\{i\}\+\\beta\_\{k\}\\sum\_\{a,i\}h\_\{a\}^\{k\-1\}\\mathrm\{x\}^\{a\}\_\{i\}\\bigg\]=\(2\\pi\)^\{nN\_\{v\}/2\}\(\\det S\)^\{\-N\_\{v\}/2\}\\exp\\bigg\(\\frac\{N\_\{v\}\\beta\_\{k\}^\{2\}\}\{2\}\\bm\{b\}^\{\\top\}S^\{\-1\}\\bm\{b\}\\bigg\),\(144\)where𝐱i=\(xia\)a\\bm\{\\mathrm\{x\}\}\_\{i\}=\(\\mathrm\{x\}^\{a\}\_\{i\}\)\_\{a\},ha≔m1ah\_\{a\}\\coloneq m^\{a\}\_\{1\}is the condensed hidden mode,𝒃=\(hak−1\)a\\bm\{b\}=\(h\_\{a\}^\{k\-1\}\)\_\{a\}, andS≔diag⁡\(ua\)−βk2​RS\\coloneq\\mathrm\{diag\}\(u^\{a\}\)\-\\beta\_\{k\}^\{2\}R\. Under the RS ansatz \(ha=hh\_\{a\}=h,ua=uu^\{a\}=u,Ra​a=αk​rdR^\{aa\}=\\alpha\_\{k\}r\_\{d\},Ra​b=αk​rR^\{ab\}=\\alpha\_\{k\}r,R^a​a=r^d\\hat\{R\}^\{aa\}=\\hat\{r\}\_\{d\},R^a​b=r^\\hat\{R\}^\{ab\}=\\hat\{r\}\),

S=D~​In−αk​βk2​r​11⊤,D~≔u−αk​βk2​\(rd−r\),S=\\tilde\{D\}\\,I\_\{n\}\-\\alpha\_\{k\}\\beta\_\{k\}^\{2\}r\\,\\bm\{1\}\\bm\{1\}^\{\\top\},\\qquad\\tilde\{D\}\\coloneq u\-\\alpha\_\{k\}\\beta\_\{k\}^\{2\}\(r\_\{d\}\-r\),\(145\)whose eigenvalues areD~\\tilde\{D\}\(with degeneracyn−1n\-1\) andD~−n​αk​βk2​r\\tilde\{D\}\-n\\alpha\_\{k\}\\beta\_\{k\}^\{2\}r, and the Sherman–Morrison formula gives

limn→01n​log​detS=log⁡D~−αk​βk2​rD~,limn→01n​𝒃⊤​S−1​𝒃=h2​k−2D~\.\\lim\_\{n\\to 0\}\\frac\{1\}\{n\}\\log\\det S=\\log\\tilde\{D\}\-\\frac\{\\alpha\_\{k\}\\beta\_\{k\}^\{2\}r\}\{\\tilde\{D\}\},\\qquad\\lim\_\{n\\to 0\}\\frac\{1\}\{n\}\\bm\{b\}^\{\\top\}S^\{\-1\}\\bm\{b\}=\\frac\{h^\{2k\-2\}\}\{\\tilde\{D\}\}\.\(146\)The stationarity conditions with respect tordr\_\{d\}anduuread

r^d=1D~\+βk2​\(αk​r\+h2​k−2\)D~2,1D~\+βk2​\(αk​r\+h2​k−2\)D~2=1,\\hat\{r\}\_\{d\}=\\frac\{1\}\{\\tilde\{D\}\}\+\\frac\{\\beta\_\{k\}^\{2\}\(\\alpha\_\{k\}r\+h^\{2k\-2\}\)\}\{\\tilde\{D\}^\{2\}\},\\qquad\\frac\{1\}\{\\tilde\{D\}\}\+\\frac\{\\beta\_\{k\}^\{2\}\(\\alpha\_\{k\}r\+h^\{2k\-2\}\)\}\{\\tilde\{D\}^\{2\}\}=1,\(147\)so thatr^d=1\\hat\{r\}\_\{d\}=1, and all remainingrdr\_\{d\}\-dependence then cancels via the shiftu=D~\+αk​βk2​\(rd−r\)u=\\tilde\{D\}\+\\alpha\_\{k\}\\beta\_\{k\}^\{2\}\(r\_\{d\}\-r\), in complete parallel with Model A \(Eq\. \([97](https://arxiv.org/html/2609.10976#A2.E97)\) and below\)\. The stationarity condition with respect torrgives

r^=βk2​\(αk​r\+h2​k−2\)D~2=m2\+αk​βk2​rD~2=q,m≔βk​hk−1D~,\\hat\{r\}=\\frac\{\\beta\_\{k\}^\{2\}\(\\alpha\_\{k\}r\+h^\{2k\-2\}\)\}\{\\tilde\{D\}^\{2\}\}=m^\{2\}\+\\frac\{\\alpha\_\{k\}\\beta\_\{k\}^\{2\}r\}\{\\tilde\{D\}^\{2\}\}=q,\\qquad m\\coloneq\\frac\{\\beta\_\{k\}h^\{k\-1\}\}\{\\tilde\{D\}\},\(148\)the Edwards–Anderson overlap of the visible variables, as in Eq\. \([98](https://arxiv.org/html/2609.10976#A2.E98)\): per site,m=⟨xi⟩m=\\left\\langle\\mathrm\{x\}\_\{i\}\\right\\rangleis the mean of the visible Gaussian andαk​βk2​r/D~2=\(S−1\)a​b\\alpha\_\{k\}\\beta\_\{k\}^\{2\}r/\\tilde\{D\}^\{2\}=\(S^\{\-1\}\)\_\{ab\}\(a≠ba\\neq b\) is its frozen covariance, so thatr^=⟨xia⟩​⟨xib⟩\+\(S−1\)a​b\\hat\{r\}=\\left\\langle\\mathrm\{x\}^\{a\}\_\{i\}\\right\\rangle\\left\\langle\\mathrm\{x\}^\{b\}\_\{i\}\\right\\rangle\+\(S^\{\-1\}\)\_\{ab\}\. Fork=2k=2, the hidden factorℐ2​\(R^\)\\mathcal\{I\}\_\{2\}\(\\hat\{R\}\)is evaluated exactly as in Appendix[B\.2\.2](https://arxiv.org/html/2609.10976#A2.SS2.SSS2)and yields the noise termαk2​Ψ2​\(q\)\\frac\{\\alpha\_\{k\}\}\{2\}\\Psi\_\{2\}\(q\)together withr=ℳ2​\(q\)r=\\mathcal\{M\}\_\{2\}\(q\)\(Eq\. \([101](https://arxiv.org/html/2609.10976#A2.E101)\) and Eq\. \([103](https://arxiv.org/html/2609.10976#A2.E103)\)\)\. Assembling all terms and eliminating\(h,u\)\(h,u\)reproduces Eq\. \([142](https://arxiv.org/html/2609.10976#A2.E142)\), with the condensed hidden mode tied to the visible overlap bym=βk​\(1−q\)​hk−1m=\\beta\_\{k\}\(1\-q\)h^\{k\-1\}\. Combined with thehh\-saddle, this relation givesm=hm=hon the retrieval branch: the condensed hidden neuron equals the overlap it detects, in accordance with the general role of the hidden neurons discussed in Sec\.[II\.1](https://arxiv.org/html/2609.10976#S2.SS1)\. Fork\>2k\>2, this computation terminates at the same obstruction as in Model A:ℐk\\mathcal\{I\}\_\{k\}admits no extensive evaluation \(Appendix[B\.2\.3](https://arxiv.org/html/2609.10976#A2.SS2.SSS3)\), and one must pass to the adiabatic formulation Eq\. \([125](https://arxiv.org/html/2609.10976#A2.E125)\), upon which the pair\(R,R^\)\(R,\\hat\{R\}\)collapses onto the visible overlap,Ra​b→αk​ℳk​\(qa​b\)R^\{ab\}\\to\\alpha\_\{k\}\\mathcal\{M\}\_\{k\}\(q\_\{ab\}\)as in Eq\. \([113](https://arxiv.org/html/2609.10976#A2.E113)\), and the computation reduces to the one performed in Appendix[B\.3\.1](https://arxiv.org/html/2609.10976#A2.SS3.SSS1)–Appendix[B\.3\.3](https://arxiv.org/html/2609.10976#A2.SS3.SSS3)\.

#### B\.3\.6Remarks

Remark 1: exactness of the gauge rotation and the Gaussian replacement\.The rotation that maps the condensed pattern to\(1,…,1\)\(1,\\ldots,1\)is an orthogonal transformation ofℝNv\\mathbb\{R\}^\{N\_\{v\}\}under which both the spherical pattern ensemble and the visible measured​Ω​\(x\)d\\Omega\(\\mathrm\{x\}\)are invariant\. The gauge choice is therefore exact at any finiteNvN\_\{v\}, just as the site\-wise sign gauge of Model A\. The only asymptotic step in the pattern average is the replacement of the remaining spherical patterns by Gaussian vectors\. By rotational invariance, the exact average of any function of the overlaps\(y^μa\)a\(\\hat\{y\}^\{a\}\_\{\\mu\}\)\_\{a\}depends on the visible configurations only through the Gram matrixQQ\. The finite\-NvN\_\{v\}corrections are therefore automatically functions of the order parameters, and they are suppressed byO⁡\(Nv−1\)O\(N\_\{v\}^\{\-1\}\)relative to the Gaussian leading term, leaving the RS equations unaffected\.

Remark 2: validity of the cumulant expansion\.The power counting of Remark 2 in Appendix[B\.2\.5](https://arxiv.org/html/2609.10976#A2.SS2.SSS5)applies with two modifications\. First, the overlaps are bounded,\|y^μa\|≤Nv\|\\hat\{y\}^\{a\}\_\{\\mu\}\|\\leq\\sqrt\{N\_\{v\}\}by the Cauchy–Schwarz inequality with‖ξμ\(h,v\)‖2=‖xa‖2=Nv\\left\\lVert\\xi^\{\(h,v\)\}\_\{\\mu\}\\right\\rVert\_\{2\}=\\left\\lVert\\mathrm\{x\}^\{a\}\\right\\rVert\_\{2\}=\\sqrt\{N\_\{v\}\}, so the per\-pattern average Eq\. \([129](https://arxiv.org/html/2609.10976#A2.E129)\) and all its cumulants are finite\. Note that naively replacingy^\\hat\{y\}by an exactly Gaussian variable*inside*the exponential would produce a divergent average fork\>2k\>2\. The correct order of operations is to expand in cumulants first and to evaluate each cumulant by the Gaussian asymptotics\. The divergence of the naive replacement is precisely theℐk\\mathcal\{I\}\_\{k\}pathology of Appendix[B\.2\.3](https://arxiv.org/html/2609.10976#A2.SS2.SSS3)in another guise\. Second, the role played bysi2=1s\_\{i\}^\{2\}=1in Model A \(the state\-independence of the first cumulant\) is now played by the spherical constraint, which fixesqa​a=1q\_\{aa\}=1exactly\. The third cumulant isO⁡\(Nv2−k/2\)O\(N\_\{v\}^\{2\-k/2\}\)by the same counting as in Model A, subextensive for evenk≥4k\\geq 4\.

Remark 3: the Onsager reaction field\.The cavity argument of Remark 1 in Appendix[B\.2\.5](https://arxiv.org/html/2609.10976#A2.SS2.SSS5)carries over withsi→xis\_\{i\}\\to\\mathrm\{x\}\_\{i\}: the linear response of a soft spin to the field shift of one non\-condensed mode is controlled by its single\-site thermal variance, whose average1Nv​∑i\(⟨xi2⟩−⟨xi⟩2\)=1−q\\frac\{1\}\{N\_\{v\}\}\\sum\_\{i\}\(\\left\\langle\\mathrm\{x\}\_\{i\}^\{2\}\\right\\rangle\-\\left\\langle\\mathrm\{x\}\_\{i\}\\right\\rangle^\{2\}\)=1\-qfollows from the spherical constraint\. The self\-feedback of a non\-condensed overlap is again of relative orderβk\(1−q\)Nv−\(k−2\)/2\\beta\_\{k\}\(1\-q\)N\_\{v\}^\{\-\(k\-2\)/2\}: marginal atk=2k=2, where its geometric resummation produces the denominator ofℳ2\\mathcal\{M\}\_\{2\}\(equivalently, the determinant Eq\. \([130](https://arxiv.org/html/2609.10976#A2.E130)\) resums it exactly\), and vanishing fork\>2k\>2, wherer=ℳk​\(q\)r=\\mathcal\{M\}\_\{k\}\(q\)is the bare Gaussian moment\. This confirms, by an argument independent of replicas, that the crosstalk moment of Model C coincides with that of Model A, as stated in the main text\.

## Appendix CDetails on Model B

This appendix collects the computations behind Sec\.[IV](https://arxiv.org/html/2609.10976#S4)\. Appendix[C\.1](https://arxiv.org/html/2609.10976#A3.SS1)carries out the zero\-temperature estimates underlying the capacity Eq\. \([61](https://arxiv.org/html/2609.10976#S4.E61)\)\. Appendix[C\.2](https://arxiv.org/html/2609.10976#A3.SS2)derives the copy representation exactly, including the hidden\-sector zero mode and the shifts quoted in Sec\.[IV\.1\.1](https://arxiv.org/html/2609.10976#S4.SS1.SSS1), and makes the duality with the replica method quantitative\. Appendix[C\.3](https://arxiv.org/html/2609.10976#A3.SS3)performs the optimization over copy configurations behind the finite\-temperature phases, the retrieval branch, and the capacity of Sec\.[IV\.2](https://arxiv.org/html/2609.10976#S4.SS2)\. Appendix[C\.4](https://arxiv.org/html/2609.10976#A3.SS4)reconstructs the phase diagram at arbitrary real temperature and settles the analytic continuation off the integer\-temperature lattice\. Throughout we use the normalization of Sec\.[IV](https://arxiv.org/html/2609.10976#S4),τv=1\{\\tau\_\{v\}\}=1andτh=λ\{\\tau\_\{h\}\}=\\lambda, for whichβ~=β\\tilde\{\\beta\}=\\beta, with Gaussian patterns at the exponential load Eq\. \([57](https://arxiv.org/html/2609.10976#S4.E57)\)\. Rates of partition functions are denotedφ≔limNv→∞Nv−1​log⁡ZξB\\varphi\\coloneq\\lim\_\{N\_\{v\}\\to\\infty\}N\_\{v\}^\{\-1\}\\log\{Z\_\{\\xi\}^\{\\mathrm\{B\}\}\}, measured relative to the Gaussian reference∫dve−β‖v‖2/2\\int dv\\,e^\{\-\\beta\\left\\lVert v\\right\\rVert^\{2\}/2\}as in Sec\.[IV\.2](https://arxiv.org/html/2609.10976#S4.SS2), and are related to the free energy densities of the main text byf=−T​φf=\-T\\varphi\.

### C\.1Vertex condensation and the zero\-temperature capacity

Vertex condensation\.The Gram matrix is positive semi\-definite,x⊤​G​x=‖ξ\(v,h\)​x‖2≥0x^\{\\top\}Gx=\\left\\lVert\\xi^\{\(v,h\)\}x\\right\\rVert^\{2\}\\geq 0, so the energyQ⁡\(f\)≔f⊤​G​fQ\(f\)\\coloneq f^\{\\top\}Gfis convex on the simplex, and a convex function on a compact convex set attains its maximum at an extreme point\[[45](https://arxiv.org/html/2609.10976#bib.bib45), Sec\. 32\]\. Together with the elementary bound Eq\. \([58](https://arxiv.org/html/2609.10976#S4.E58)\), this places the maximum at the vertex of the pattern of maximal norm, a deterministic fact independent of the pattern ensemble\. The entropic competition quoted in Sec\.[IV\.1](https://arxiv.org/html/2609.10976#S4.SS1)is quantified as follows\. For the uniform mixturefmix=1M​∑μ∈Seμf\_\{\\mathrm\{mix\}\}=\\frac\{1\}\{M\}\\sum\_\{\\mu\\in S\}e\_\{\\mu\}over a setSSofMMpatterns,

Q⁡\(fmix\)=1M2​\[∑μ∈S‖ξμ\(h,v\)‖2\+∑μ≠ν∈Sξμ\(h,v\)⋅ξν\(h,v\)\]=NvM\+O⁡\(NvM\),Q\(f\_\{\\mathrm\{mix\}\}\)=\\frac\{1\}\{M^\{2\}\}\\Big\[\\sum\_\{\\mu\\in S\}\\left\\lVert\\xi^\{\(h,v\)\}\_\{\\mu\}\\right\\rVert^\{2\}\+\\sum\_\{\\begin\{subarray\}\{c\}\\mu\\neq\\nu\\in S\\end\{subarray\}\}\\xi^\{\(h,v\)\}\_\{\\mu\}\\cdot\\xi^\{\(h,v\)\}\_\{\\nu\}\\Big\]=\\frac\{N\_\{v\}\}\{M\}\+O\\big\(\\tfrac\{\\sqrt\{N\_\{v\}\}\}\{M\}\\big\),\(149\)since theM⁡\(M−1\)M\(M\-1\)zero\-mean cross terms, each of sizeO⁡\(Nv\)O\(\\sqrt\{N\_\{v\}\}\), add up toO⁡\(M​Nv\)O\(M\\sqrt\{N\_\{v\}\}\)typically\. Hence

Φ⁡\(eμ\)−Φ⁡\(fmix\)≈β2​Nv​\(1−1M\)−βλ​log⁡M,\\Phi\(e\_\{\\mu\}\)\-\\Phi\(f\_\{\\mathrm\{mix\}\}\)\\approx\\frac\{\\beta\}\{2\}N\_\{v\}\\left\(1\-\\frac\{1\}\{M\}\\right\)\-\\frac\{\\beta\}\{\\lambda\}\\log M,\(150\)which is positive unlesslog⁡M≳Nv\\log M\\gtrsim N\_\{v\}: the entropy of an interior mixture can never offset its extensive energy cost at subexponential load, and the competition becomes genuine only at the load Eq\. \([57](https://arxiv.org/html/2609.10976#S4.E57)\)\. Even then the competitor is not an interior point of the simplex\. The Gibbs measure decomposes into lumps attached to the vertices,

ZξB≈∑μ=1NhZμ,Zμ≍eβ2​‖ξμ\(h,v\)‖2,\{Z\_\{\\xi\}^\{\\mathrm\{B\}\}\}\\approx\\sum\_\{\\mu=1\}^\{N\_\{h\}\}Z\_\{\\mu\},\\qquad Z\_\{\\mu\}\\asymp e^\{\\frac\{\\beta\}\{2\}\\left\\lVert\\xi^\{\(h,v\)\}\_\{\\mu\}\\right\\rVert^\{2\}\},\(151\)and the entropy is gained by spreading the measure over exponentially many, individually almost pure, lumps\. The distinction between a mixed configuration and a mixture of pure states is the same as in spin\-glass theory, and the counting of the lumps is precisely the entropy evaluated below\. A final caution concerns the continuum measure: the simplex has dimensionNh−1=eα​Nv−1N\_\{h\}\-1=e^\{\\alpha N\_\{v\}\}\-1, so the flat measure carries volume factors of orderlog⁡Vol⁡\(ΔNh\)∼−Nh​log⁡Nh\\log\\mathrm\{Vol\}\(\\Delta^\{N\_\{h\}\}\)\\sim\-N\_\{h\}\\log N\_\{h\}, doubly exponential inNvN\_\{v\}, which would overwhelm anyeO⁡\(Nv\)e^\{O\(N\_\{v\}\)\}energy unless they cancel exactly\. The statements above are therefore formulated either through the adiabatic saddle point Eq\. \([53](https://arxiv.org/html/2609.10976#S4.E53)\), for which vertex condensation is the single\-term dominance of the log\-sum\-exp, or through the discrete copy representation of Appendix[C\.2](https://arxiv.org/html/2609.10976#A3.SS2), whose counting measure carries no volume factors\.

Norm statistics and freezing\.The single\-pattern rate function quoted throughout the main text follows from Cramér’s theorem\[[28](https://arxiv.org/html/2609.10976#bib.bib28),[57](https://arxiv.org/html/2609.10976#bib.bib57)\]\. For‖ξμ\(h,v\)‖2=∑i\(ξμ​i\(h,v\)\)2\\left\\lVert\\xi^\{\(h,v\)\}\_\{\\mu\}\\right\\rVert^\{2\}=\\sum\_\{i\}\(\\xi^\{\(h,v\)\}\_\{\\mu i\}\)^\{2\}a sum ofNvN\_\{v\}i\.i\.d\. squared Gaussians, the cumulant generating function per component is

Λ⁡\(t\)≔log⁡𝔼​et​ξ2=−12​log⁡\(1−2​t\),t<12,\\Lambda\(t\)\\coloneq\\log\\mathbb\{E\}\\,e^\{t\\xi^\{2\}\}=\-\\frac\{1\}\{2\}\\log\(1\-2t\),\\qquad t<\\frac\{1\}\{2\},\(152\)and the Legendre transformsupt\[t​x−Λ⁡\(t\)\]\\sup\_\{t\}\[tx\-\\Lambda\(t\)\], whose stationary point ist∗=12​\(1−1/x\)t^\{\\ast\}=\\frac\{1\}\{2\}\(1\-1/x\), yields

ℙ⁡\(‖ξμ\(h,v\)‖2Nv≈x\)≍e−Nv​I1​\(x\),I1​\(x\)=x−1−log⁡x2,\\mathbb\{P\}\\left\(\\frac\{\\left\\lVert\\xi^\{\(h,v\)\}\_\{\\mu\}\\right\\rVert^\{2\}\}\{N\_\{v\}\}\\approx x\\right\)\\asymp e^\{\-N\_\{v\}I\_\{1\}\(x\)\},\\quad I\_\{1\}\(x\)=\\frac\{x\-1\-\\log x\}\{2\},\(153\)theM=1M=1case of Eq\. \([64](https://arxiv.org/html/2609.10976#S4.E64)\)\. The rateI1I\_\{1\}is convex, vanishes atx=1x=1, and has the bounded derivativeI1′​\(x\)=12​\(1−1/x\)<12I\_\{1\}^\{\\prime\}\(x\)=\\frac\{1\}\{2\}\(1\-1/x\)<\\frac\{1\}\{2\}\. Writing‖ξμ\(h,v\)‖2=\(1\+εμ\)​Nv\\left\\lVert\\xi^\{\(h,v\)\}\_\{\\mu\}\\right\\rVert^\{2\}=\(1\+\\varepsilon\_\{\\mu\}\)N\_\{v\}, the number of patterns at norm excessε\\varepsilonis

𝒩⁡\(ε\)≍Nh​e−Nv​I1​\(1\+ε\)=eNv​\[α−I1​\(1\+ε\)\]\.\\mathcal\{N\}\(\\varepsilon\)\\asymp N\_\{h\}\\,e^\{\-N\_\{v\}I\_\{1\}\(1\+\\varepsilon\)\}=e^\{N\_\{v\}\[\\alpha\-I\_\{1\}\(1\+\\varepsilon\)\]\}\.\(154\)Forα\>I1​\(1\+ε\)\\alpha\>I\_\{1\}\(1\+\\varepsilon\)the occupation numbers of independent patterns concentrate on this value \(Var/Mean2≍e−Nv​\[α−I1\]→0\\mathrm\{Var\}/\\mathrm\{Mean\}^\{2\}\\asymp e^\{\-N\_\{v\}\[\\alpha\-I\_\{1\}\]\}\\to 0by the second\-moment method\), while forα<I1​\(1\+ε\)\\alpha<I\_\{1\}\(1\+\\varepsilon\)the level is empty with high probability by Markov’s inequality\. The maximal norm excess is thereforeεmax​\(α\)\\varepsilon\_\{\\max\}\(\\alpha\)withI1​\(1\+εmax\)=αI\_\{1\}\(1\+\\varepsilon\_\{\\max\}\)=\\alpha, as stated in Sec\.[IV\.2](https://arxiv.org/html/2609.10976#S4.SS2), andεmax≈2​α\\varepsilon\_\{\\max\}\\approx 2\\sqrt\{\\alpha\}for smallα\\alpha\. The lump sum Eq\. \([151](https://arxiv.org/html/2609.10976#A3.E151)\) then follows by the maximal\-term principle,

1Nv​log​∑μeβ2​‖ξμ\(h,v\)‖2=β2\+max0≤ε≤εmax⁡\[α−I1​\(1\+ε\)\+β2​ε\]\.\\frac\{1\}\{N\_\{v\}\}\\log\\sum\_\{\\mu\}e^\{\\frac\{\\beta\}\{2\}\\left\\lVert\\xi^\{\(h,v\)\}\_\{\\mu\}\\right\\rVert^\{2\}\}=\\frac\{\\beta\}\{2\}\+\\max\_\{0\\leq\\varepsilon\\leq\\varepsilon\_\{\\max\}\}\\left\[\\alpha\-I\_\{1\}\(1\+\\varepsilon\)\+\\frac\{\\beta\}\{2\}\\varepsilon\\right\]\.\(155\)The interior stationary pointI1′​\(1\+ε∗\)=β/2I\_\{1\}^\{\\prime\}\(1\+\\varepsilon^\{\\ast\}\)=\\beta/2givesε∗=β/\(1−β\)\\varepsilon^\{\\ast\}=\\beta/\(1\-\\beta\), which exists forβ<1\\beta<1and lies in the populated range whileI1​\(1\+ε∗\)=κ⁡\(β\)≤αI\_\{1\}\(1\+\\varepsilon^\{\\ast\}\)=\\kappa\(\\beta\)\\leq\\alpha\. Substitution collapses Eq\. \([155](https://arxiv.org/html/2609.10976#A3.E155)\) toα−12​log⁡\(1−β\)\\alpha\-\\frac\{1\}\{2\}\\log\(1\-\\beta\), in exact agreement with the annealed average𝔼eβ2​χNv2=\(1−β\)−Nv/2\\mathbb\{E\}\\,e^\{\\frac\{\\beta\}\{2\}\\chi^\{2\}\_\{N\_\{v\}\}\}=\(1\-\\beta\)^\{\-N\_\{v\}/2\}per pattern: the sum is carried by exponentially many typical terms and self\-averages\. SinceI1′<12I\_\{1\}^\{\\prime\}<\\frac\{1\}\{2\}, forβ≥1\\beta\\geq 1the bracket in Eq\. \([155](https://arxiv.org/html/2609.10976#A3.E155)\) increases monotonically, and forκ⁡\(β\)\>α\\kappa\(\\beta\)\>\\alphathe stationary point leaves the populated range\. In either case the maximum is pinned at the boundaryεmax\\varepsilon\_\{\\max\}, where the counting exponent vanishes and

1Nv​log​∑μeβ2​‖ξμ\(h,v\)‖2=β2​\(1\+εmax​\(α\)\):\\frac\{1\}\{N\_\{v\}\}\\log\\sum\_\{\\mu\}e^\{\\frac\{\\beta\}\{2\}\\left\\lVert\\xi^\{\(h,v\)\}\_\{\\mu\}\\right\\rVert^\{2\}\}=\\frac\{\\beta\}\{2\}\\left\(1\+\\varepsilon\_\{\\max\}\(\\alpha\)\\right\):\(156\)the sum is dominated by theO⁡\(1\)O\(1\)patterns of maximal norm, and the quenched value falls below the annealed one, which diverges altogether forβ≥1\\beta\\geq 1\. This is the freezing mechanism of the REM, for whose rigorous treatment see\[[58](https://arxiv.org/html/2609.10976#bib.bib58)\]\. These two branches are the free energiesfCf\_\{\\mathrm\{C\}\}andfFf\_\{\\mathrm\{F\}\}of Eq\. \([65](https://arxiv.org/html/2609.10976#S4.E65)\), and the thresholdβ=1\\beta=1reappears in Appendix[C\.3](https://arxiv.org/html/2609.10976#A3.SS3)as the stability criterion of the visible fluctuations\.

Leak sum and capacity\.The stationary points of the effective energy Eq\. \([54](https://arxiv.org/html/2609.10976#S4.E54)\) obey the fixed\-point equationv=ξ\(v,h\)​f∗​\(v\)v=\\xi^\{\(v,h\)\}f^\{\\ast\}\(v\)of Sec\.[IV\.1](https://arxiv.org/html/2609.10976#S4.SS1)\. For the retrieval ansatz we evaluate the fields atv=ξ1\(h,v\)v=\\xi^\{\(h,v\)\}\_\{1\}and verify self\-consistency afterwards\. Conditioned onξ1\(h,v\)\\xi^\{\(h,v\)\}\_\{1\}, the overlapsξμ\(h,v\)⋅ξ1\(h,v\)\\xi^\{\(h,v\)\}\_\{\\mu\}\\cdot\\xi^\{\(h,v\)\}\_\{1\}forμ≥2\\mu\\geq 2are exactly i\.i\.d\. centered Gaussians of variance‖ξ1\(h,v\)‖2≈Nv\\left\\lVert\\xi^\{\(h,v\)\}\_\{1\}\\right\\rVert^\{2\}\\approx N\_\{v\}, which is the statementaμ=λ​Nv​ωμa\_\{\\mu\}=\\lambda\\sqrt\{N\_\{v\}\}\\,\\omega\_\{\\mu\}used in Eq\. \([59](https://arxiv.org/html/2609.10976#S4.E59)\)\. The leak sumLLis a Boltzmann sum over the linear random energiesλ​Nv​ωμ\\lambda\\sqrt\{N\_\{v\}\}\\,\\omega\_\{\\mu\}, so the counting dichotomy above applies with the Gaussian ratex2/2x^\{2\}/2in place ofI1I\_\{1\}: levelsωμ≈x​Nv\\omega\_\{\\mu\}\\approx x\\sqrt\{N\_\{v\}\}are populated byeNv​\(α−x2/2\)e^\{N\_\{v\}\(\\alpha\-x^\{2\}/2\)\}patterns up toxmax=2​αx\_\{\\max\}=\\sqrt\{2\\alpha\}and empty beyond, which is the content of Eq\. \([60](https://arxiv.org/html/2609.10976#S4.E60)\), with the interior branch again matching the annealed average ofLLand the boundary branch frozen on theO⁡\(1\)O\(1\)most aligned patterns\. Retrieval is self\-consistent when1−p≤e−λ​Nv​L→01\-p\\leq e^\{\-\\lambda N\_\{v\}\}L\\to 0, i\.e\.Nv−1​log⁡L<λN\_\{v\}^\{\-1\}\\log L<\\lambda\. On the frozen branch this readsλ​2​α<λ\\lambda\\sqrt\{2\\alpha\}<\\lambda, i\.e\.α<12\\alpha<\\frac\{1\}\{2\}, while on the annealed branch it readsα\+λ2/2<λ\\alpha\+\\lambda^\{2\}/2<\\lambda, i\.e\.α<λ−λ2/2\\alpha<\\lambda\-\\lambda^\{2\}/2\. The branch applicable at the threshold is determined by comparingλ\\lambdawith2​αc\\sqrt\{2\\alpha\_\{c\}\}, which reduces toλ≷1\\lambda\\gtrless 1, and the two conditions combine into the capacity Eq\. \([61](https://arxiv.org/html/2609.10976#S4.E61)\), continuous atλ=1\\lambda=1\. The correction to the ansatz is controlled by the same leak:δ≔v¯−ξ1\(h,v\)=∑μ≥2fμ∗​ξμ\(h,v\)\\delta\\coloneq\\bar\{v\}\-\\xi^\{\(h,v\)\}\_\{1\}=\\sum\_\{\\mu\\geq 2\}f^\{\\ast\}\_\{\\mu\}\\xi^\{\(h,v\)\}\_\{\\mu\}obeys‖δ‖≤\(1−p\)​maxμ​‖ξμ\(h,v\)‖\\left\\lVert\\delta\\right\\rVert\\leq\(1\-p\)\\max\_\{\\mu\}\\left\\lVert\\xi^\{\(h,v\)\}\_\{\\mu\}\\right\\rVert, exponentially small whenever the leak exponent is negative, so the fixed point survives in a neighborhood ofξ1\(h,v\)\\xi^\{\(h,v\)\}\_\{1\}throughout the retrieval region\.

Zero\-temperature thermodynamics\.At the retrieval saddle the effective energy Eq\. \([54](https://arxiv.org/html/2609.10976#S4.E54)\) isE⁡\(ξ1\(h,v\)\)≈Nv2−1λ⋅λ​Nv=−Nv2E\(\\xi^\{\(h,v\)\}\_\{1\}\)\\approx\\frac\{N\_\{v\}\}\{2\}\-\\frac\{1\}\{\\lambda\}\\cdot\\lambda N\_\{v\}=\-\\frac\{N\_\{v\}\}\{2\}, the leak contributing onlye−O⁡\(Nv\)e^\{\-O\(N\_\{v\}\)\}corrections\. The pointv=0v=0is also stationary,∇E\|0=−Nh−1∑μξμ\(h,v\)=O\(Nv/Nh\)\\nabla E\|\_\{0\}=\-N\_\{h\}^\{\-1\}\\sum\_\{\\mu\}\\xi^\{\(h,v\)\}\_\{\\mu\}=O\(\\sqrt\{N\_\{v\}/N\_\{h\}\}\), with energyE⁡\(0\)=−1λ​log⁡Nh=−αλ​NvE\(0\)=\-\\frac\{1\}\{\\lambda\}\\log N\_\{h\}=\-\\frac\{\\alpha\}\{\\lambda\}N\_\{v\}\. Typical retrieval therefore lies below this paramagnetic point only forα<λ/2\\alpha<\\lambda/2, and forλ≤1\\lambda\\leq 1the windowλ/2<α<λ−λ2/2\\lambda/2<\\alpha<\\lambda\-\\lambda^\{2\}/2supports retrieval only as a metastable state\. Neither of these is the true zero\-temperature equilibrium, however: the retrieval state of the maximal\-norm pattern has energy density−\(1\+εmax\)/2\-\(1\+\\varepsilon\_\{\\max\}\)/2, below both, and its leak condition is satisfied wherever the comparison is relevant because the signal is enhanced by the factor1\+εmax1\+\\varepsilon\_\{\\max\}\. Equating it with the paramagnetic value yields the zero\-temperature equilibrium boundary2​α/λ=1\+εmax​\(α\)2\\alpha/\\lambda=1\+\\varepsilon\_\{\\max\}\(\\alpha\), whose small\-λ\\lambdasolution isα≈λ/2\+λ3/2/2\\alpha\\approx\\lambda/2\+\\lambda^\{3/2\}/\\sqrt\{2\}\. This is theT→0T\\to 0limit of the P–F boundary derived in Appendix[C\.3](https://arxiv.org/html/2609.10976#A3.SS3), and the extreme\-value enhancement is precisely the amount by which the equilibrium boundary exceeds the naive comparisonα=λ/2\\alpha=\\lambda/2\.

### C\.2Copy representation and hidden\-sector corrections

Zero mode of the hidden Hessian and the shifts\.The adiabatic evaluation Eq\. \([10](https://arxiv.org/html/2609.10976#S2.E10)\) presumes a positive\-definite Hessian of the hidden Lagrangian\. For Model B this fails in exactly one direction\. At any point of the hidden space,

Hess⁡\(Lh​\(h\)\)=diag\(f\)−f​f⊤,f=softmax⁡\(h\),\\operatorname\{Hess\}\\big\(\{L\_\{h\}\}\(h\)\\big\)=\\mathop\{\\mathrm\{diag\}\}\\nolimits\(f\)\-ff^\{\\top\},\\qquad f=\\mathrm\{softmax\}\(h\),\(157\)and the normalization∑μfμ=1\\sum\_\{\\mu\}f\_\{\\mu\}=1makes the all\-one vector an exact zero mode,

\(diag\(f\)−f​f⊤\)​1=f−f​∑μfμ=0\.\\big\(\\mathop\{\\mathrm\{diag\}\}\\nolimits\(f\)\-ff^\{\\top\}\\big\)\\,\\bm\{1\}=f\-f\\sum\_\{\\mu\}f\_\{\\mu\}=0\.\(158\)This is not an accident of the saddle point but the gauge direction of the softmax, which is invariant under the uniform shifth→h\+c​1h\\to h\+c\\,\\bm\{1\}noted in Sec\.[IV](https://arxiv.org/html/2609.10976#S4)\. The energy is exactly constant along𝟏\\bm\{1\}, not merely flat to quadratic order\. The remaining spectrum is well behaved: Eq\. \([157](https://arxiv.org/html/2609.10976#A3.E157)\) is a rank\-one downdate of a positive diagonal matrix, its eigenvalues interlace the softmax weights, and exactly one eigenvalue vanishes while the others remain positive\. Its pseudo\-determinant \(the product of the nonzero eigenvalues, denoteddet′\{\\det\}^\{\\prime\}\) follows from the matrix determinant lemma with anϵ\\epsilonregularization,

det\(Hess\+ϵ​I\)\\displaystyle\\det\\big\(\\operatorname\{Hess\}\+\\epsilon I\\big\)=\(1−∑μfμ2fμ\+ϵ\)​∏μ\(fμ\+ϵ\)\\displaystyle=\\Big\(1\-\\sum\_\{\\mu\}\\frac\{f\_\{\\mu\}^\{2\}\}\{f\_\{\\mu\}\+\\epsilon\}\\Big\)\\prod\_\{\\mu\}\(f\_\{\\mu\}\+\\epsilon\)\(159\)=\[Nh​ϵ\+O⁡\(ϵ2\)\]​∏μfμ​\(1\+O⁡\(ϵ\)\),\\displaystyle=\\big\[N\_\{h\}\\,\\epsilon\+O\(\\epsilon^\{2\}\)\\big\]\\prod\_\{\\mu\}f\_\{\\mu\}\\,\\big\(1\+O\(\\epsilon\)\\big\),so that

det′Hess⁡\(Lh​\(h\)\)=limϵ→0det\(Hess\+ϵ​I\)ϵ=Nh​∏μ=1Nhfμ\.\{\\det\}^\{\\prime\}\\operatorname\{Hess\}\\big\(\{L\_\{h\}\}\(h\)\\big\)=\\lim\_\{\\epsilon\\to 0\}\\frac\{\\det\(\\operatorname\{Hess\}\+\\epsilon I\)\}\{\\epsilon\}=N\_\{h\}\\prod\_\{\\mu=1\}^\{N\_\{h\}\}f\_\{\\mu\}\.\(160\)Expanding thehhintegral around the adiabatic saddle point along the Hessian eigenbasis, the flat direction contributes a divergent volumeV0V\_\{0\}that is independent ofvvand of the patterns\. Dividing it out and performing the remainingNh−1N\_\{h\}\-1Gaussian modes yields the corrected version of Eq\. \([10](https://arxiv.org/html/2609.10976#S2.E10)\),

ZξB\(β\)/V0=\(2​πβ/λ\)\(Nh−1\)/2∫ℝNvdve−β​EξB​\(v,h∗\)\(det′Hess\(Lh\(h∗\)\)\)−1/2\.\{Z\_\{\\xi\}^\{\\mathrm\{B\}\}\}\(\\beta\)/V\_\{0\}=\\left\(\\frac\{2\\pi\}\{\\beta/\\lambda\}\\right\)^\{\\\!\(N\_\{h\}\-1\)/2\}\\int\_\{\\mathbb\{R\}^\{N\_\{v\}\}\}dv\\;e^\{\-\\beta\{E\_\{\\xi\}^\{\\mathrm\{B\}\}\}\(v,h\_\{\\ast\}\)\}\\left\(\{\\det\}^\{\\prime\}\\operatorname\{Hess\}\(\{L\_\{h\}\}\(h\_\{\\ast\}\)\)\\right\)^\{\-1/2\}\.\(161\)The exponent\(Nh−1\)/2\(N\_\{h\}\-1\)/2in place ofNh/2N\_\{h\}/2, together with the quotient byV0V\_\{0\}, is the measure prescription stated in the footnote of Sec\.[IV](https://arxiv.org/html/2609.10976#S4): the flat measured​fdfon the simplex is the flat measured​hdhmodulo the gauge orbit\.

The determinant factor in Eq\. \([161](https://arxiv.org/html/2609.10976#A3.E161)\) is itself an exponential of the fields\. Witha≔λ​ξ\(h,v\)​va\\coloneq\\lambda\\,\\xi^\{\(h,v\)\}vandfν∗=softmax​\(a\)ν=eaν/∑μeaμf^\{\\ast\}\_\{\\nu\}=\\mathrm\{softmax\}\(a\)\_\{\\nu\}=e^\{a\_\{\\nu\}\}/\\sum\_\{\\mu\}e^\{a\_\{\\mu\}\},

\(det′Hess\)−1/2=Nh−1/2e−12∑νaν\(∑μeaμ\)Nh/2,\\left\(\{\\det\}^\{\\prime\}\\operatorname\{Hess\}\\right\)^\{\-1/2\}=N\_\{h\}^\{\-1/2\}\\,e^\{\-\\frac\{1\}\{2\}\\sum\_\{\\nu\}a\_\{\\nu\}\}\\Big\(\\sum\_\{\\mu\}e^\{a\_\{\\mu\}\}\\Big\)^\{N\_\{h\}/2\},\(162\)so the integrand of Eq\. \([161](https://arxiv.org/html/2609.10976#A3.E161)\) combines with the power\(∑μeaμ\)β/λ\(\\sum\_\{\\mu\}e^\{a\_\{\\mu\}\}\)^\{\\beta/\\lambda\}of the effective energy Eq\. \([54](https://arxiv.org/html/2609.10976#S4.E54)\) into

exp⁡\[−β2​‖v‖2−12​∑νaν\]​\(∑μeaμ\)β/λ\+Nh/2\.\\exp\\Big\[\-\\frac\{\\beta\}\{2\}\\left\\lVert v\\right\\rVert^\{2\}\-\\frac\{1\}\{2\}\\sum\_\{\\nu\}a\_\{\\nu\}\\Big\]\\Big\(\\sum\_\{\\mu\}e^\{a\_\{\\mu\}\}\\Big\)^\{\\beta/\\lambda\+N\_\{h\}/2\}\.\(163\)Whenever the combined exponentn¯≔β/λ\+Nh/2\\bar\{n\}\\coloneq\\beta/\\lambda\+N\_\{h\}/2is a positive integer, the multinomial theorem linearizes the power into a sum overn¯\\bar\{n\}\-tuples, each term is Gaussian invv, and completing the square component by component gives, up to the same prefactors,

ZξB\(β\)/V0∝∑μ1,…,μn¯exp\[λ22​β‖ξ~\{μ\}‖2\],ξ~\{μ\}≔−12∑ν=1Nhξν\(h,v\)\+∑j=1n¯ξμj\(h,v\)\.\{Z\_\{\\xi\}^\{\\mathrm\{B\}\}\}\(\\beta\)/V\_\{0\}\\propto\\sum\_\{\\mu\_\{1\},\\dots,\\mu\_\{\\bar\{n\}\}\}\\exp\\Big\[\\frac\{\\lambda^\{2\}\}\{2\\beta\}\\left\\lVert\\tilde\{\\xi\}\_\{\\left\\\{\\mu\\right\\\}\}\\right\\rVert^\{2\}\\Big\],\\quad\\tilde\{\\xi\}\_\{\\left\\\{\\mu\\right\\\}\}\\coloneq\-\\frac\{1\}\{2\}\\sum\_\{\\nu=1\}^\{N\_\{h\}\}\\xi^\{\(h,v\)\}\_\{\\nu\}\+\\sum\_\{j=1\}^\{\\bar\{n\}\}\\xi^\{\(h,v\)\}\_\{\\mu\_\{j\}\}\.\(164\)Relative to the representation Eq\. \([62](https://arxiv.org/html/2609.10976#S4.E62)\), which coincides with the visible\-only computation of\[[14](https://arxiv.org/html/2609.10976#bib.bib14)\], the hidden\-sector fluctuations thus produce exactly the two corrections quoted in Sec\.[IV\.1\.1](https://arxiv.org/html/2609.10976#S4.SS1.SSS1): the exponent shiftβ/λ→β/λ\+Nh/2\\beta/\\lambda\\to\\beta/\\lambda\+N\_\{h\}/2and the additive pattern shift−12∑νξ\(h,v\)ν\-\\frac\{1\}\{2\}\\sum\_\{\\nu\}\\xi^\{\(h,v\)\}\_\{\\nu\}\. Two features of Eq\. \([164](https://arxiv.org/html/2609.10976#A3.E164)\) deserve emphasis at exponential load\. First, positivity of the temperature imposesβ/λ=n¯−Nh/2\>0\\beta/\\lambda=\\bar\{n\}\-N\_\{h\}/2\>0: the admissible integers satisfyn¯\>Nh/2\\bar\{n\}\>N\_\{h\}/2, so the small copy numbers, in particularn¯=1\\bar\{n\}=1, are excluded onceNhN\_\{h\}is large\. Second, the shift is not a perturbation: the components of12​∑νξν\(h,v\)\\frac\{1\}\{2\}\\sum\_\{\\nu\}\\xi^\{\(h,v\)\}\_\{\\nu\}are of orderNh\\sqrt\{N\_\{h\}\}, so atNh=eα​NvN\_\{h\}=e^\{\\alpha N\_\{v\}\}the shifted sum is dominated by the correction term itself\. The analysis of the main text is therefore formulated for the adiabatic partition function proper, the leading saddle\-point term for which the unshifted expansion Eq\. \([62](https://arxiv.org/html/2609.10976#S4.E62)\) is an exact identity atn=β/λ∈ℤ\>0n=\\beta/\\lambda\\in\\mathbb\{Z\}\_\{\>0\}\. The corrections Eq\. \([164](https://arxiv.org/html/2609.10976#A3.E164)\) constitute the complete Gaussian fluctuation content of the hidden sector around that limit, and their consistent treatment at exponential load remains open \(see the remarks closing Appendix[C\.4](https://arxiv.org/html/2609.10976#A3.SS4)\)\.

Alignment versus dispersion\.Reversing the Gaussian integration that produced Eq\. \([62](https://arxiv.org/html/2609.10976#S4.E62)\) shows that, conditioned on a copy configuration, the visible state is Gaussian around the centroid of the selected patterns,

v\|\{μj\}∼𝒩⁡\(1n​∑j=1nξμj\(h,v\),β−1​INv\),v\\,\\big\|\\,\\left\\\{\\mu\_\{j\}\\right\\\}\\;\\sim\\;\\mathcal\{N\}\\Big\(\\frac\{1\}\{n\}\\sum\_\{j=1\}^\{n\}\\xi^\{\(h,v\)\}\_\{\\mu\_\{j\}\},\\;\\beta^\{\-1\}I\_\{N\_\{v\}\}\\Big\),\(165\)which is the parallel\-query picture of Sec\.[IV\.1\.1](https://arxiv.org/html/2609.10976#S4.SS1.SSS1)and will be used repeatedly below to read off visible observables\. Expanding the energy of Eq\. \([62](https://arxiv.org/html/2609.10976#S4.E62)\),

‖∑jξμj\(h,v\)‖2=∑j‖ξμj\(h,v\)‖2\+∑j≠l\[Nv​δμj​μl\+O⁡\(Nv\)\],\\left\\lVert\\sum\_\{j\}\\xi^\{\(h,v\)\}\_\{\\mu\_\{j\}\}\\right\\rVert^\{2\}=\\sum\_\{j\}\\left\\lVert\\xi^\{\(h,v\)\}\_\{\\mu\_\{j\}\}\\right\\rVert^\{2\}\+\\sum\_\{j\\neq l\}\\Big\[N\_\{v\}\\,\\delta\_\{\\mu\_\{j\}\\mu\_\{l\}\}\+O\(\\sqrt\{N\_\{v\}\}\)\\Big\],\(166\)the copies form annn\-site system withNhN\_\{h\}states per site, a ferromagnetic gain of orderNvN\_\{v\}for every pair of aligned copies, and random couplings of orderNv\\sqrt\{N\_\{v\}\}otherwise\. The zero\-temperature equilibrium of Appendix[C\.1](https://arxiv.org/html/2609.10976#A3.SS1)can be recovered from this picture by pure counting\. The fully aligned configurations \(μj≡μ\\mu\_\{j\}\\equiv\\mu, withNhN\_\{h\}choices\) carry the exponentλ2​n​n2​Nv\+α​Nv=\(β2\+α\)​Nv\\frac\{\\lambda\}\{2n\}n^\{2\}N\_\{v\}\+\\alpha N\_\{v\}=\(\\frac\{\\beta\}\{2\}\+\\alpha\)N\_\{v\}, while the fully dispersed ones \(nμ∈\{0,1\}n\_\{\\mu\}\\in\\left\\\{0,1\\right\\\}, with≈en​α​Nv\\approx e^\{n\\alpha N\_\{v\}\}choices\) retain only the diagonal energy,\(λ2\+n​α\)​Nv\(\\frac\{\\lambda\}\{2\}\+n\\alpha\)N\_\{v\}\. The difference,

\(β2\+α\)−\(λ2\+n​α\)=\(n−1\)​\(λ2−α\),\\Big\(\\frac\{\\beta\}\{2\}\+\\alpha\\Big\)\-\\Big\(\\frac\{\\lambda\}\{2\}\+n\\alpha\\Big\)=\(n\-1\)\\Big\(\\frac\{\\lambda\}\{2\}\-\\alpha\\Big\),\(167\)shows that alignment is favored precisely forα<λ/2\\alpha<\\lambda/2, reproducing the zero\-temperature equilibrium threshold of Appendix[C\.1](https://arxiv.org/html/2609.10976#A3.SS1)without any reference to theffrepresentation\. The frozen phase corresponds, in the same language, to the co\-condensation of all copies onto the few patterns of extreme norm\.

Legendre duality and one\-step RSB\.The qualitative dictionary of Sec\.[IV\.1\.1](https://arxiv.org/html/2609.10976#S4.SS1.SSS1)becomes a precise convex duality\. On the replica side one computes𝔼ξ​\[\(ZξB\)r\]\\mathbb\{E\}\_\{\\xi\}\\left\[\(\{Z\_\{\\xi\}^\{\\mathrm\{B\}\}\}\)^\{r\}\\right\]at a formal replica numberrr\. The patterns enter the replicated exponent through∑iξi⊤​Θ​ξi\\sum\_\{i\}\\xi\_\{i\}^\{\\top\}\\Theta\\,\\xi\_\{i\}with the tilting matrixΘ≔β2​∑a=1rfa​\(fa\)⊤\\Theta\\coloneq\\frac\{\\beta\}\{2\}\\sum\_\{a=1\}^\{r\}f^\{a\}\(f^\{a\}\)^\{\\top\}, whereξi∈ℝNh\\xi\_\{i\}\\in\\mathbb\{R\}^\{N\_\{h\}\}collects theii\-th components of all patterns, and the Gaussian average per visible component is the matrix version of Eq\. \([152](https://arxiv.org/html/2609.10976#A3.E152)\),

𝔼eξ⊤​Θ​ξ=det\(I−2Θ\)−1/2≕eΛ⁡\(Θ\),\\mathbb\{E\}\\,e^\{\\xi^\{\\top\}\\Theta\\xi\}=\\det\(I\-2\\Theta\)^\{\-1/2\}\\eqcolon e^\{\\Lambda\(\\Theta\)\},\(168\)valid forI−2​Θ≻0I\-2\\Theta\\succ 0\. By Sylvester’s identity,det\(INh−β​∑afa​\(fa\)⊤\)=det\(Ir−β​R\)\\det\(I\_\{N\_\{h\}\}\-\\beta\\sum\_\{a\}f^\{a\}\(f^\{a\}\)^\{\\top\}\)=\\det\(I\_\{r\}\-\\beta R\)with the replica overlap matrixRa​b=fa⋅fbR^\{ab\}=f^\{a\}\\cdot f^\{b\}, so the disorder average generates the effective replica coupling−Nv2logdet\(Ir−βR\)\-\\frac\{N\_\{v\}\}\{2\}\\log\\det\(I\_\{r\}\-\\beta R\)\. On the copy side the disorder statistics appear instead as the counting rateIM​\(Q\)I\_\{M\}\(Q\)of Eq\. \([64](https://arxiv.org/html/2609.10976#S4.E64)\)\. The two are a Legendre–Fenchel pair\[[45](https://arxiv.org/html/2609.10976#bib.bib45),[28](https://arxiv.org/html/2609.10976#bib.bib28)\],

IM​\(Q\)\\displaystyle I\_\{M\}\(Q\)=supΘ\[Tr⁡\(Θ​Q\)−Λ⁡\(Θ\)\],\\displaystyle=\\sup\_\{\\Theta\}\\big\[\\Tr\(\\Theta Q\)\-\\Lambda\(\\Theta\)\\big\],\(169\)Λ⁡\(Θ\)\\displaystyle\\Lambda\(\\Theta\)=supQ\[Tr⁡\(Θ​Q\)−IM​\(Q\)\],\\displaystyle=\\sup\_\{Q\}\\big\[\\Tr\(\\Theta Q\)\-I\_\{M\}\(Q\)\\big\],with stationarityQ∗=\(I−2​Θ\)−1Q^\{\\ast\}=\(I\-2\\Theta\)^\{\-1\}: the optimal Gram matrix is the covariance of the exponentially tilted Gaussian ensemble\. The annealed branches of Appendix[C\.3](https://arxiv.org/html/2609.10976#A3.SS3)realize the second line of Eq\. \([169](https://arxiv.org/html/2609.10976#A3.E169)\) at the rank\-one tilts generated by the copy energy, and the conjugate order parameter that a replica computation introduces through the Fourier representation ofδ\\deltafunctions is precisely the tilting matrixΘ\\Theta: the field of the exponential change of measure that selects which distortion of the pattern statistics dominates\. Freezing is the point where the two routes diverge\. The Gaussian average Eq\. \([168](https://arxiv.org/html/2609.10976#A3.E168)\) runs over an unbounded ensemble, whereas onlyeα​Nve^\{\\alpha N\_\{v\}\}patterns exist\. The correct variational problem is the Fenchel transform constrained to the populated set,

supQ\[Tr⁡\(Θ​Q\)−IM​\(Q\)\]subject toIM​\(Q\)≤M​α\.\\sup\_\{Q\}\\;\\big\[\\Tr\(\\Theta Q\)\-I\_\{M\}\(Q\)\\big\]\\quad\\text\{subject to\}\\quad I\_\{M\}\(Q\)\\leq M\\alpha\.\(170\)When the constraint is active, the Karush–Kuhn–Tucker condition with multiplierη≥0\\eta\\geq 0readsΘ=\(1\+η\)∇IM\(Q\)\\Theta=\(1\+\\eta\)\\nabla I\_\{M\}\(Q\), i\.e\.

Q=\(I−2​Θ1\+η\)−1:Q=\\Big\(I\-\\frac\{2\\Theta\}\{1\+\\eta\}\\Big\)^\{\-1\}:\(171\)the tilt is renormalized by the factor\(1\+η\)−1∈\(0,1\]\(1\+\\eta\)^\{\-1\}\\in\(0,1\]\. In the scalar case relevant to the frozen phase \(tiltβ/2\\beta/2, boundary conditionI1​\(Q\)=αI\_\{1\}\(Q\)=\\alpha, i\.e\.Q=1\+εmaxQ=1\+\\varepsilon\_\{\\max\}\), Eq\. \([171](https://arxiv.org/html/2609.10976#A3.E171)\) givesβ/\(1\+η\)=εmax/\(1\+εmax\)=βf\\beta/\(1\+\\eta\)=\\varepsilon\_\{\\max\}/\(1\+\\varepsilon\_\{\\max\}\)=\\beta\_\{f\}, so that

11\+η=βfβ=TTf,\\frac\{1\}\{1\+\\eta\}=\\frac\{\\beta\_\{f\}\}\{\\beta\}=\\frac\{T\}\{T\_\{f\}\},\(172\)precisely the temperature dependence of the one\-step Parisi parameter of the REM\[[21](https://arxiv.org/html/2609.10976#bib.bib21),[27](https://arxiv.org/html/2609.10976#bib.bib27),[58](https://arxiv.org/html/2609.10976#bib.bib58)\]: the frozen phase is pinned at the edge of the populated set and behaves as if held at the freezing temperature, which is the origin of the temperature independence offFf\_\{\\mathrm\{F\}\}\.

The identification of the freezing prescription with one\-step RSB can be checked exactly on the lump sum Eq\. \([155](https://arxiv.org/html/2609.10976#A3.E155)\), which is a REM in its own right\. Computing𝔼​Z~r\\mathbb\{E\}\\tilde\{Z\}^\{r\}forZ~=∑μeβ2​‖ξμ\(h,v\)‖2\\tilde\{Z\}=\\sum\_\{\\mu\}e^\{\\frac\{\\beta\}\{2\}\\left\\lVert\\xi^\{\(h,v\)\}\_\{\\mu\}\\right\\rVert^\{2\}\}with the one\-step ansatz,r/xr/xblocks ofxxreplicas co\-condensing on distinct patterns, each block contributeseα​Nv​𝔼​ex​β2​χNv2=eNv​\[α−12​log⁡\(1−x​β\)\]e^\{\\alpha N\_\{v\}\}\\,\\mathbb\{E\}\\,e^\{\\frac\{x\\beta\}\{2\}\\chi^\{2\}\_\{N\_\{v\}\}\}=e^\{N\_\{v\}\[\\alpha\-\\frac\{1\}\{2\}\\log\(1\-x\\beta\)\]\}, so per replica

φ1​R​S​B​\(x\)=1x​\[α−12​log⁡\(1−x​β\)\]\.\\varphi\_\{\\mathrm\{1RSB\}\}\(x\)=\\frac\{1\}\{x\}\\Big\[\\alpha\-\\frac\{1\}\{2\}\\log\(1\-x\\beta\)\\Big\]\.\(173\)Extremizing over the block size continued tox∈\(0,1\]x\\in\(0,1\]yieldsα=κ⁡\(x​β\)\\alpha=\\kappa\(x\\beta\), whose solution isx∗​β=βfx^\{\\ast\}\\beta=\\beta\_\{f\}, i\.e\.x∗=T/Tfx^\{\\ast\}=T/T\_\{f\}, in agreement with the multiplier Eq\. \([172](https://arxiv.org/html/2609.10976#A3.E172)\)\. Substituting back,α−12​log⁡\(1−βf\)=α\+12​log⁡\(1\+εmax\)=εmax/2\\alpha\-\\frac\{1\}\{2\}\\log\(1\-\\beta\_\{f\}\)=\\alpha\+\\frac\{1\}\{2\}\\log\(1\+\\varepsilon\_\{\\max\}\)=\\varepsilon\_\{\\max\}/2and

φ1​R​S​B​\(x∗\)=ββf⋅εmax2=β2​\(1\+εmax\),\\varphi\_\{\\mathrm\{1RSB\}\}\(x^\{\\ast\}\)=\\frac\{\\beta\}\{\\beta\_\{f\}\}\\cdot\\frac\{\\varepsilon\_\{\\max\}\}\{2\}=\\frac\{\\beta\}\{2\}\\big\(1\+\\varepsilon\_\{\\max\}\\big\),\(174\)which is the frozen counting value Eq\. \([156](https://arxiv.org/html/2609.10976#A3.E156)\) exactly\. Forβ<βf\\beta<\\beta\_\{f\}the stationary point falls atx\>1x\>1, the extremum over the physical interval is the endpointx=1x=1, and the RS evaluation returns the annealed branch, again in agreement\. For the sign conventions of ther→0r\\to 0extremization see\[[27](https://arxiv.org/html/2609.10976#bib.bib27)\]\. The one\-step structure also fixes the weight statistics of the frozen phase: the lump weightswμ=eβ2​‖ξμ\(h,v\)‖2/Z~w\_\{\\mu\}=e^\{\\frac\{\\beta\}\{2\}\\left\\lVert\\xi^\{\(h,v\)\}\_\{\\mu\}\\right\\rVert^\{2\}\}/\\tilde\{Z\}follow a Ruelle point process of parameterx∗=T/Tfx^\{\\ast\}=T/T\_\{f\}\[[36](https://arxiv.org/html/2609.10976#bib.bib36),[50](https://arxiv.org/html/2609.10976#bib.bib50)\], with participation ratio𝔼​∑μwμ2=1−T/Tf\\mathbb\{E\}\\sum\_\{\\mu\}w\_\{\\mu\}^\{2\}=1\-T/T\_\{f\}\. Since thewμw\_\{\\mu\}are the attention weights of theffrepresentation, this is a direct quantitative prediction for the attention statistics in the F phase\.

Clone method\.Finally, the copy representation realizes the clone method of Monasson\[[22](https://arxiv.org/html/2609.10976#bib.bib22)\]with a physical clone number\. The single\-group sector of Appendix[C\.3](https://arxiv.org/html/2609.10976#A3.SS3)\(allnncopies on one pattern\) has the rate

φ=maxε⁡\[α−I1​\(1\+ε\)⏟complexity​Σ​\(ε\)\+n⋅λ2​\(1\+ε\)⏟per\-clone gain\],\\varphi=\\max\_\{\\varepsilon\}\\Big\[\\underbrace\{\\alpha\-I\_\{1\}\(1\+\\varepsilon\)\}\_\{\\text\{complexity \}\\Sigma\(\\varepsilon\)\}\+n\\cdot\\underbrace\{\\frac\{\\lambda\}\{2\}\(1\+\\varepsilon\)\}\_\{\\text\{per\-clone gain\}\}\\Big\],\(175\)which is exactly the clone free energy: the choice of the occupied state \(pattern\) is counted once for the whole clone collective, while the energy is paid once per clone\. By the envelope theorem,∂φ/∂n=λ2​\(1\+ε∗\)\\partial\\varphi/\\partial n=\\frac\{\\lambda\}\{2\}\(1\+\\varepsilon^\{\\ast\}\)reads off the occupied level andφ−n​∂φ/∂n=Σ⁡\(ε∗\)\\varphi\-n\\,\\partial\\varphi/\\partial n=\\Sigma\(\\varepsilon^\{\\ast\}\)the complexity, so the clone\-number derivative extracts the configurational entropy exactly as in the structural\-glass application\. The distinctive feature of Model B is that the clone numbern=β/λn=\\beta/\\lambdais not an auxiliary parameter but the physical temperature itself\. Scanningnnthrough real values is scanning the temperature, and the analytic control of this continuation is supplied by the generalized binomial expansion of Appendix[C\.4](https://arxiv.org/html/2609.10976#A3.SS4)\.

### C\.3Finite\-temperature analysis in the copy representation

Counting lemma\.We first derive the counting statement of Sec\.[IV\.2](https://arxiv.org/html/2609.10976#S4.SS2)\. FixMMdistinct indicesμ1,…,μM\\mu\_\{1\},\\dots,\\mu\_\{M\}\. For each visible componentiithe vectorxi≔\(ξμ1​i\(h,v\),…,ξμM​i\(h,v\)\)⊤x\_\{i\}\\coloneq\(\\xi^\{\(h,v\)\}\_\{\\mu\_\{1\}i\},\\dots,\\xi^\{\(h,v\)\}\_\{\\mu\_\{M\}i\}\)^\{\\top\}is standard Gaussian inℝM\\mathbb\{R\}^\{M\}, i\.i\.d\. overii, and the normalized Gram matrix is the empirical covarianceQ^=1Nv​∑ixi​xi⊤\\hat\{Q\}=\\frac\{1\}\{N\_\{v\}\}\\sum\_\{i\}x\_\{i\}x\_\{i\}^\{\\top\}\. Its cumulant generating function per component is the matrix average Eq\. \([168](https://arxiv.org/html/2609.10976#A3.E168)\), and Cramér’s theorem givesℙ⁡\(Q^≈Q\)≍e−Nv​IM​\(Q\)\\mathbb\{P\}\(\\hat\{Q\}\\approx Q\)\\asymp e^\{\-N\_\{v\}I\_\{M\}\(Q\)\}with the Legendre transform of Eq\. \([169](https://arxiv.org/html/2609.10976#A3.E169)\): the stationary pointΘ∗=12​\(I−Q−1\)\\Theta^\{\\ast\}=\\frac\{1\}\{2\}\(I\-Q^\{\-1\}\)yieldsTr⁡\(Θ∗​Q\)=12​Tr⁡\(Q−I\)\\Tr\(\\Theta^\{\\ast\}Q\)=\\frac\{1\}\{2\}\\Tr\(Q\-I\)andΛ⁡\(Θ∗\)=12​log​detQ\\Lambda\(\\Theta^\{\\ast\}\)=\\frac\{1\}\{2\}\\log\\det Q, which yields the rate Eq\. \([64](https://arxiv.org/html/2609.10976#S4.E64)\)\. Multiplying by theeM​α​Nve^\{M\\alpha N\_\{v\}\}choices of the indices, and repeating the second\-moment and Markov arguments of Appendix[C\.1](https://arxiv.org/html/2609.10976#A3.SS1), the number ofMM\-tuples with Gram matrixQQconcentrates oneNv​\[M​α−IM​\(Q\)\]e^\{N\_\{v\}\[M\\alpha\-I\_\{M\}\(Q\)\]\}when the exponent is positive and vanishes with high probability otherwise\. In the eigenbasis ofQQthe rate decomposes as

IM​\(Q\)=∑a=1MI1​\(Λa\),I\_\{M\}\(Q\)=\\sum\_\{a=1\}^\{M\}I\_\{1\}\(\\Lambda\_\{a\}\),\(176\)withΛa\\Lambda\_\{a\}the eigenvalues ofQQ, and this decomposition organizes every computation below\.

Partition classes and the bulk branches\.Following Sec\.[IV\.2](https://arxiv.org/html/2609.10976#S4.SS2), classify the copy configurations by the partition of thenncopies over distinct patterns\. Consider first the symmetric classes:MMgroups of equal sizen/Mn/M, each occupying one ofMMdistinct patterns \(the mixed classes relevant for retrieval are treated below, and the remaining channels in Appendix[C\.4](https://arxiv.org/html/2609.10976#A3.SS4)\)\. The number of ways to distribute the copies is subexponential inNvN\_\{v\}, and the weight in Eq\. \([62](https://arxiv.org/html/2609.10976#S4.E62)\) depends on the configuration only through the Gram matrixQQof the occupied patterns, with‖∑jξμj\(h,v\)‖2=\(n/M\)2​1⊤​Q​1​Nv\\left\\lVert\\sum\_\{j\}\\xi^\{\(h,v\)\}\_\{\\mu\_\{j\}\}\\right\\rVert^\{2\}=\(n/M\)^\{2\}\\,\\bm\{1\}^\{\\top\}Q\\,\\bm\{1\}\\,N\_\{v\}, so the class exponent per visible neuron is

φ⁡\[Q\]=M​α−IM​\(Q\)\+β2​M2​1⊤​Q​1\.\\varphi\[Q\]=M\\alpha\-I\_\{M\}\(Q\)\+\\frac\{\\beta\}\{2M^\{2\}\}\\,\\bm\{1\}^\{\\top\}Q\\,\\bm\{1\}\.\(177\)The energy couples only to the coherent quadratic form𝟏⊤​Q​1\\bm\{1\}^\{\\top\}Q\\,\\bm\{1\}\. Writingwa≔\(𝟏⋅ea\)2/Mw\_\{a\}\\coloneq\(\\bm\{1\}\\cdot e\_\{a\}\)^\{2\}/Mfor the weights of the eigenvectors ofQQin the coherent direction \(∑awa=1\\sum\_\{a\}w\_\{a\}=1\), so that𝟏⊤​Q​𝟏=M​∑awa​Λa\\bm\{1\}^\{\\top\}Q\\bm\{1\}=M\\sum\_\{a\}w\_\{a\}\\Lambda\_\{a\}, the convexity and positivity ofI1I\_\{1\}give

IM​\(Q\)=∑aI1​\(Λa\)≥∑awa​I1​\(Λa\)≥I1​\(∑awa​Λa\),I\_\{M\}\(Q\)=\\sum\_\{a\}I\_\{1\}\(\\Lambda\_\{a\}\)\\geq\\sum\_\{a\}w\_\{a\}I\_\{1\}\(\\Lambda\_\{a\}\)\\geq I\_\{1\}\\Big\(\\sum\_\{a\}w\_\{a\}\\Lambda\_\{a\}\\Big\),\(178\)with equality if and only if𝟏/M\\bm\{1\}/\\sqrt\{M\}is an eigenvector and all remaining eigenvalues lie at the minimum ofI1I\_\{1\}, i\.e\. at11\. The symmetric formQ=\(qd−q\)​I\+q​11⊤Q=\(q\_\{d\}\-q\)I\+q\\,\\bm\{1\}\\bm\{1\}^\{\\top\}with transverse eigenvalueΛ2=qd−q=1\\Lambda\_\{2\}=q\_\{d\}\-q=1is therefore not an ansatz but the exact maximizer within each class, and the exponent reduces to one variable, the coherent eigenvalueΛ1\\Lambda\_\{1\},

φ⁡\(M,Λ1\)=M​α−I1​\(Λ1\)\+β2​M​Λ1\.\\varphi\(M,\\Lambda\_\{1\}\)=M\\alpha\-I\_\{1\}\(\\Lambda\_\{1\}\)\+\\frac\{\\beta\}\{2M\}\\Lambda\_\{1\}\.\(179\)ForM=1M=1this is precisely the lump sum Eq\. \([155](https://arxiv.org/html/2609.10976#A3.E155)\), of which Eq\. \([179](https://arxiv.org/html/2609.10976#A3.E179)\) is the generalization to collectively occupied patterns\. The maximization overΛ1\\Lambda\_\{1\}repeats the annealed/frozen dichotomy of Appendix[C\.1](https://arxiv.org/html/2609.10976#A3.SS1)\. The interior stationary pointI1′​\(Λ1∗\)=β/2​MI\_\{1\}^\{\\prime\}\(\\Lambda\_\{1\}^\{\\ast\}\)=\\beta/2MgivesΛ1∗=\(1−β/M\)−1\\Lambda\_\{1\}^\{\\ast\}=\(1\-\\beta/M\)^\{\-1\}, which exists forβ<M\\beta<M\. At the stationary point the identityβ2​M​Λ1∗−12​\(Λ1∗−1\)=0\\frac\{\\beta\}\{2M\}\\Lambda\_\{1\}^\{\\ast\}\-\\frac\{1\}\{2\}\(\\Lambda\_\{1\}^\{\\ast\}\-1\)=0collapses the exponent to

φann​\(M\)=M​α−12​log⁡\(1−βM\),M​α≥κ⁡\(βM\),\\varphi\_\{\\mathrm\{ann\}\}\(M\)=M\\alpha\-\\frac\{1\}\{2\}\\log\\Big\(1\-\\frac\{\\beta\}\{M\}\\Big\),\\qquad M\\alpha\\geq\\kappa\\Big\(\\frac\{\\beta\}\{M\}\\Big\),\(180\)in exact agreement with the annealed average of the class partition function \(the rank\-one tilt has the single nontrivial eigenvalueβ/2​M\\beta/2Malong𝟏\\bm\{1\}, so the determinant in Eq\. \([168](https://arxiv.org/html/2609.10976#A3.E168)\) is\(1−β/M\)−1/2\(1\-\\beta/M\)^\{\-1/2\}\): the annealed branch is the self\-averaging regime\. When the stationary point violates the counting constraint, or whenβ≥M\\beta\\geq M, the maximum is pinned at the boundary of the populated set, where the counting exponent vanishes\. SinceI1I\_\{1\}is convex with its minimum atΛ=1\\Lambda=1, the counting conditionI1​\(Λ\)=M​αI\_\{1\}\(\\Lambda\)=M\\alphahas one root below and one above unity, and the exponent Eq\. \([179](https://arxiv.org/html/2609.10976#A3.E179)\) increases withΛ1\\Lambda\_\{1\}throughout the populated range, so the boundary is the upper rootΛ\+\>1\\Lambda\_\{\+\}\>1, the largest coherent eigenvalue carried by a populated class:

φfroz​\(M\)=β2​M​Λ\+,I1​\(Λ\+\)=M​α,Λ\+\>1\.\\varphi\_\{\\mathrm\{froz\}\}\(M\)=\\frac\{\\beta\}\{2M\}\\Lambda\_\{\+\},\\qquad I\_\{1\}\(\\Lambda\_\{\+\}\)=M\\alpha,\\quad\\Lambda\_\{\+\}\>1\.\(181\)AtM=1M=1one hasΛ\+=1\+εmax​\(α\)\\Lambda\_\{\+\}=1\+\\varepsilon\_\{\\max\}\(\\alpha\), and Eq\. \([181](https://arxiv.org/html/2609.10976#A3.E181)\) reduces to Eq\. \([156](https://arxiv.org/html/2609.10976#A3.E156)\)\. The optimization overMMis now elementary\. On the annealed branch,d​φann/d​M=α−β/\[2​M​\(M−β\)\]d\\varphi\_\{\\mathrm\{ann\}\}/dM=\\alpha\-\\beta/\[2M\(M\-\\beta\)\]tends to−∞\-\\inftyasM→β\+M\\to\\beta^\{\+\}and toα\>0\\alpha\>0at largeMM: the branch has an interior minimum inMM, so its maximum is attained at the endpoints,M=nM=n\(all copies dispersed\) orM=1M=1\(all copies aligned\)\. On the frozen branch, differentiating the constraint in Eq\. \([181](https://arxiv.org/html/2609.10976#A3.E181)\) gives

d​φfrozd​M=−β2​M2​Λ\+​log⁡Λ\+Λ\+−1<0,\\frac\{d\\varphi\_\{\\mathrm\{froz\}\}\}\{dM\}=\-\\frac\{\\beta\}\{2M^\{2\}\}\\,\\frac\{\\Lambda\_\{\+\}\\log\\Lambda\_\{\+\}\}\{\\Lambda\_\{\+\}\-1\}<0,\(182\)so freezing always prefers full alignment,M=1M=1\. Three bulk branches survive\. ForM=nM=n\(usingβ/n=λ\\beta/n=\\lambda\),φP=n​α−12​log⁡\(1−λ\)\\varphi\_\{\\mathrm\{P\}\}=n\\alpha\-\\frac\{1\}\{2\}\\log\(1\-\\lambda\)\. ForM=1M=1annealed,φC=α−12​log⁡\(1−β\)\\varphi\_\{\\mathrm\{C\}\}=\\alpha\-\\frac\{1\}\{2\}\\log\(1\-\\beta\)with validityβ<1\\beta<1andα≥κ⁡\(β\)\\alpha\\geq\\kappa\(\\beta\)\. ForM=1M=1frozen,φF=β2​\(1\+εmax\)\\varphi\_\{\\mathrm\{F\}\}=\\frac\{\\beta\}\{2\}\(1\+\\varepsilon\_\{\\max\}\)\. Throughf=−T​φf=\-T\\varphithese are the free energies Eq\. \([65](https://arxiv.org/html/2609.10976#S4.E65)\)\. The entropy of the C branch issC=−∂fC/∂T=α−κ\(β\)s\_\{\\mathrm\{C\}\}=\-\\partial f\_\{\\mathrm\{C\}\}/\\partial T=\\alpha\-\\kappa\(\\beta\), which vanishes exactly on the freezing line Eq\. \([66](https://arxiv.org/html/2609.10976#S4.E66)\)\. Thereβf=εmax/\(1\+εmax\)\\beta\_\{f\}=\\varepsilon\_\{\\max\}/\(1\+\\varepsilon\_\{\\max\}\), and substituting into either branch gives the common valueφ=εmax/2\\varphi=\\varepsilon\_\{\\max\}/2, so C and F connect continuously, the REM freezing scenario in the form used in Appendix[C\.2](https://arxiv.org/html/2609.10976#A3.SS2)\. By the conditional Gaussian Eq\. \([165](https://arxiv.org/html/2609.10976#A3.E165)\), the F state is the complete retrieval state of the maximal\-norm pattern, as asserted in the main text\. The P state, by contrast, is a weak thermal condensate: the tilted Gram matrix of the dispersed patterns follows from the stationary covarianceQ∗=\(I−2​Θ\)−1Q^\{\\ast\}=\(I\-2\\Theta\)^\{\-1\}by the Sherman–Morrison formula,q∗=λ2​T/\(1−λ\)q^\{\\ast\}=\\lambda^\{2\}T/\(1\-\\lambda\)off the diagonal, so the centroid in Eq\. \([165](https://arxiv.org/html/2609.10976#A3.E165)\) carries‖1n​∑jξμj\(h,v\)‖2/Nv=Λ1∗/n=λ​T/\(1−λ\)\\left\\lVert\\frac\{1\}\{n\}\\sum\_\{j\}\\xi^\{\(h,v\)\}\_\{\\mu\_\{j\}\}\\right\\rVert^\{2\}/N\_\{v\}=\\Lambda\_\{1\}^\{\\ast\}/n=\\lambda T/\(1\-\\lambda\), and adding the thermal varianceTTof the Gaussian fluctuations,

⟨‖v‖2⟩Nv=λ​T1−λ\+T=T1−λ,\\frac\{\\left\\langle\\left\\lVert v\\right\\rVert^\{2\}\\right\\rangle\}\{N\_\{v\}\}=\\frac\{\\lambda T\}\{1\-\\lambda\}\+T=\\frac\{T\}\{1\-\\lambda\},\(183\)the value quoted in Sec\.[IV\.2](https://arxiv.org/html/2609.10976#S4.SS2): even the paramagnet weakly aligns the selected patterns \(q∗\>0q^\{\\ast\}\>0\), a linear\-response enhancement that vanishes asT→0T\\to 0\. Finally, the existence conditions unify: a symmetric class has attention weights1/M1/MonMMpatterns, hence inverse participation ratio‖f‖2=1/M\\left\\lVert f\\right\\rVert^\{2\}=1/M, and the annealed conditionβ<M\\beta<Misβ​‖f‖2<1\\beta\\left\\lVert f\\right\\rVert^\{2\}<1\. For P this readsλ<1\\lambda<1and for C it readsβ<1\\beta<1, and the instability atβ​‖f‖2=1\\beta\\left\\lVert f\\right\\rVert^\{2\}=1is the divergence of the visible Gaussian integral in Eq\. \([56](https://arxiv.org/html/2609.10976#S4.E56)\), i\.e\. the softening of the visible fluctuations by concentrated attention\. Forλ≥1\\lambda\\geq 1the paramagnetic branch does not exist\.

Retrieval branch\.Condition on the retrieved patternξ1\(h,v\)\\xi^\{\(h,v\)\}\_\{1\}, of typical norm, and consider the mixed classes of Sec\.[IV\.2](https://arxiv.org/html/2609.10976#S4.SS2):n​pnpcopies onξ1\(h,v\)\\xi^\{\(h,v\)\}\_\{1\}andN′≔n⁡\(1−p\)N^\{\\prime\}\\coloneq n\(1\-p\)copies dispersed on distinct patternsμ≠1\\mu\\neq 1, with overlapscj=ξμj\(h,v\)⋅ξ1\(h,v\)/Nvc\_\{j\}=\\xi^\{\(h,v\)\}\_\{\\mu\_\{j\}\}\\cdot\\xi^\{\(h,v\)\}\_\{1\}/N\_\{v\}and mutual Gram matrixQ′Q^\{\\prime\}\. The energy depends only on∑jcj\\sum\_\{j\}c\_\{j\}and𝟏⊤​Q′​𝟏\\bm\{1\}^\{\\top\}Q^\{\\prime\}\\bm\{1\}, so the Jensen argument Eq\. \([178](https://arxiv.org/html/2609.10976#A3.E178)\), supplemented by the Cauchy–Schwarz inequality for thecjc\_\{j\}, again makes the uniform and symmetric form \(cj=cc\_\{j\}=c,Q′=\(qd−q\)​I\+q​𝟏𝟏⊤Q^\{\\prime\}=\(q\_\{d\}\-q\)I\+q\\bm\{1\}\\bm\{1\}^\{\\top\}\) exact within the family\. The counting rate conditioned onξ1\(h,v\)\\xi^\{\(h,v\)\}\_\{1\}follows from the longitudinal–transverse decompositionξμj\(h,v\)=ℓj​e1\+ξj⟂\\xi^\{\(h,v\)\}\_\{\\mu\_\{j\}\}=\\ell\_\{j\}\\,e\_\{1\}\+\\xi^\{\\perp\}\_\{j\}withe1=ξ1\(h,v\)/‖ξ1\(h,v\)‖e\_\{1\}=\\xi^\{\(h,v\)\}\_\{1\}/\\left\\lVert\\xi^\{\(h,v\)\}\_\{1\}\\right\\rVert\. The longitudinal components are scalar unit Gaussians constrained toℓj=c​Nv\\ell\_\{j\}=c\\sqrt\{N\_\{v\}\}, at ratec2/2c^\{2\}/2each, while the transverse vectors are i\.i\.d\. standard Gaussian in the orthogonal complement with mutual overlapsξj⟂⋅ξl⟂/Nv=Qj​l′−c2\\xi\_\{j\}^\{\\perp\}\\cdot\\xi\_\{l\}^\{\\perp\}/N\_\{v\}=Q^\{\\prime\}\_\{jl\}\-c^\{2\}, at rateIN′​\(Q′−c2​𝟏𝟏⊤\)I\_\{N^\{\\prime\}\}\(Q^\{\\prime\}\-c^\{2\}\\bm\{1\}\\bm\{1\}^\{\\top\}\)\. Adding the two, the longitudinal cost cancels against the trace shift,

Icond\\displaystyle I\_\{\\mathrm\{cond\}\}=N′​c22\+IN′​\(Q′−c2​𝟏𝟏⊤\)\\displaystyle=\\frac\{N^\{\\prime\}c^\{2\}\}\{2\}\+I\_\{N^\{\\prime\}\}\\big\(Q^\{\\prime\}\-c^\{2\}\\bm\{1\}\\bm\{1\}^\{\\top\}\\big\)\(184\)=12\[TrQ′−logdet\(Q′−c2𝟏𝟏⊤\)−N′\],\\displaystyle=\\frac\{1\}\{2\}\\Big\[\\Tr Q^\{\\prime\}\-\\log\\det\\big\(Q^\{\\prime\}\-c^\{2\}\\bm\{1\}\\bm\{1\}^\{\\top\}\\big\)\-N^\{\\prime\}\\Big\],so only the Schur complement survives in the log\-determinant\. Equivalently, apply Eq\. \([64](https://arxiv.org/html/2609.10976#S4.E64)\) to the full\(N′\+1\)\(N^\{\\prime\}\+1\)\-tuple includingξ1\(h,v\)\\xi^\{\(h,v\)\}\_\{1\}and usedetQfull=det\(Q′−c2​𝟏𝟏⊤\)\\det Q\_\{\\mathrm\{full\}\}=\\det\(Q^\{\\prime\}\-c^\{2\}\\bm\{1\}\\bm\{1\}^\{\\top\}\)together with the vanishing conditioning costI1​\(1\)=0I\_\{1\}\(1\)=0\. In the symmetric form the eigenvalues ofQ′−c2​𝟏𝟏⊤Q^\{\\prime\}\-c^\{2\}\\bm\{1\}\\bm\{1\}^\{\\top\}areΛ1′−N′​c2\\Lambda\_\{1\}^\{\\prime\}\-N^\{\\prime\}c^\{2\}\(coherent,Λ1′=qd\+\(N′−1\)​q\\Lambda\_\{1\}^\{\\prime\}=q\_\{d\}\+\(N^\{\\prime\}\-1\)q\) andΛ2′=qd−q\\Lambda\_\{2\}^\{\\prime\}=q\_\{d\}\-q\(transverse\), so

Icond=\(N′−1\)​I1​\(Λ2′\)\+12​\[Λ1′−log⁡\(Λ1′−N′​c2\)−1\],I\_\{\\mathrm\{cond\}\}=\(N^\{\\prime\}\-1\)\\,I\_\{1\}\(\\Lambda\_\{2\}^\{\\prime\}\)\+\\frac\{1\}\{2\}\\Big\[\\Lambda\_\{1\}^\{\\prime\}\-\\log\\big\(\\Lambda\_\{1\}^\{\\prime\}\-N^\{\\prime\}c^\{2\}\\big\)\-1\\Big\],\(185\)where the trace carriesΛ1′\\Lambda\_\{1\}^\{\\prime\}but the logarithm its Schur shift: the mismatch is the imprint of the longitudinal cost\. The energy exponent follows from‖n​p​ξ1\(h,v\)\+∑jξμj\(h,v\)‖2/Nv=\(n​p\)2\+2​n​p​N′​c\+N′​Λ1′\\left\\lVert np\\,\\xi^\{\(h,v\)\}\_\{1\}\+\\sum\_\{j\}\\xi^\{\(h,v\)\}\_\{\\mu\_\{j\}\}\\right\\rVert^\{2\}/N\_\{v\}=\(np\)^\{2\}\+2npN^\{\\prime\}c\+N^\{\\prime\}\\Lambda\_\{1\}^\{\\prime\}, multiplied byλ/2​n\\lambda/2n:

β2​p2\+β​p​\(1−p\)​c\+λ⁡\(1−p\)2​Λ1′\.\\frac\{\\beta\}\{2\}p^\{2\}\+\\beta p\(1\-p\)\\,c\+\\frac\{\\lambda\(1\-p\)\}\{2\}\\Lambda\_\{1\}^\{\\prime\}\.\(186\)The coefficient of the last term isλ\\lambda, notβ\\beta: the coherent enhancement of a dispersed copy is that of a single copy, and this asymmetry between the condensed group and the leak drives the entire structure below\. Assembling the selection entropyN′​αN^\{\\prime\}\\alpha, the rate Eq\. \([185](https://arxiv.org/html/2609.10976#A3.E185)\), and the energy Eq\. \([186](https://arxiv.org/html/2609.10976#A3.E186)\), and introducingA≔\(Λ1′−N′​c2\)−1A\\coloneq\(\\Lambda\_\{1\}^\{\\prime\}\-N^\{\\prime\}c^\{2\}\)^\{\-1\}, the stationarity conditions are elementary\. The transverse eigenvalue carries no energy, soI1′​\(Λ2′\)=0I\_\{1\}^\{\\prime\}\(\\Lambda\_\{2\}^\{\\prime\}\)=0givesΛ2′=1\\Lambda\_\{2\}^\{\\prime\}=1\. Theccderivative balances−N′​c​A\-N^\{\\prime\}cAfrom the logarithm againstβ​p​\(1−p\)=λ​p​N′\\beta p\(1\-p\)=\\lambda pN^\{\\prime\}from the cross term, and theΛ1′\\Lambda\_\{1\}^\{\\prime\}derivative balances−12\+A2\-\\frac\{1\}\{2\}\+\\frac\{A\}\{2\}againstλ⁡\(1−p\)/2\\lambda\(1\-p\)/2\. Hence

Λ2′=1,A=1−λ⁡\(1−p\),c=λ​pA\.\\Lambda\_\{2\}^\{\\prime\}=1,\\qquad A=1\-\\lambda\(1\-p\),\\qquad c=\\frac\{\\lambda p\}\{A\}\.\(187\)Substituting back, theΛ1′\\Lambda\_\{1\}^\{\\prime\}terms combine into−A2​Λ1′=−12−A2​N′​c2\-\\frac\{A\}\{2\}\\Lambda\_\{1\}^\{\\prime\}=\-\\frac\{1\}\{2\}\-\\frac\{A\}\{2\}N^\{\\prime\}c^\{2\}, the constant cancels, the twocc\-dependent terms add to\+n​λ22p2\(1−p\)/A\+\\frac\{n\\lambda^\{2\}\}\{2\}p^\{2\}\(1\-p\)/A, and the identityA\+λ⁡\(1−p\)=1A\+\\lambda\(1\-p\)=1collapses the sum with the condensate self\-energy into a single term:

φR​\(p\)=βλ​\(1−p\)​α−12​log⁡A\+β2​p2A,\\varphi\_\{\\mathrm\{R\}\}\(p\)=\\frac\{\\beta\}\{\\lambda\}\(1\-p\)\\,\\alpha\-\\frac\{1\}\{2\}\\log A\+\\frac\{\\beta\}\{2\}\\frac\{p^\{2\}\}\{A\},\(188\)which is the Landau function Eq\. \([67](https://arxiv.org/html/2609.10976#S4.E67)\) throughfR=−T​φRf\_\{\\mathrm\{R\}\}=\-T\\varphi\_\{\\mathrm\{R\}\}, with the endpoint valuesφR​\(0\)=φP\\varphi\_\{\\mathrm\{R\}\}\(0\)=\\varphi\_\{\\mathrm\{P\}\}andφR​\(1\)=β/2\\varphi\_\{\\mathrm\{R\}\}\(1\)=\\beta/2quoted in the main text\. The visible overlap follows from the centroid Eq\. \([165](https://arxiv.org/html/2609.10976#A3.E165)\),

m=v⋅ξ1\(h,v\)Nv=p\+\(1−p\)​c=pA,c=λ​m,m=\\frac\{v\\cdot\\xi^\{\(h,v\)\}\_\{1\}\}\{N\_\{v\}\}=p\+\(1\-p\)\\,c=\\frac\{p\}\{A\},\\qquad c=\\lambda m,\(189\)where the second form usesA\+λ⁡\(1−p\)=1A\+\\lambda\(1\-p\)=1again\. The relationc=λ​mc=\\lambda midentifies the stationary overlap of the destinations as a linear response to the retrieval field, and the first form of Eq\. \([189](https://arxiv.org/html/2609.10976#A3.E189)\) is a self\-consistent loop,m=p\+λ⁡\(1−p\)​mm=p\+\\lambda\(1\-p\)\\,m: the leak amplifies the visible overlap by the geometric series of the loop gainλ⁡\(1−p\)\\lambda\(1\-p\), and the positivityA\>0A\>0of the loop stiffness is once more the criterionβ​‖f‖2<1\\beta\\left\\lVert f\\right\\rVert^\{2\}<1, evaluated on the leak weights‖fleak‖2=\(1−p\)/n\\left\\lVert f\_\{\\mathrm\{leak\}\}\\right\\rVert^\{2\}=\(1\-p\)/n\. For later use we record the stationarity of Eq\. \([188](https://arxiv.org/html/2609.10976#A3.E188)\) inpp\. Using1/A=\(1−λ​m\)/\(1−λ\)1/A=\(1\-\\lambda m\)/\(1\-\\lambda\), the condition∂pφR=0\\partial\_\{p\}\\varphi\_\{\\mathrm\{R\}\}=0closes in the visible overlap alone,

α=FT​\(m\)≔λ​m​\(1−λ​m2\)−λ2​T2​1−λ​m1−λ,\\alpha=F\_\{T\}\(m\)\\coloneq\\lambda m\\Big\(1\-\\frac\{\\lambda m\}\{2\}\\Big\)\-\\frac\{\\lambda^\{2\}T\}\{2\}\\,\\frac\{1\-\\lambda m\}\{1\-\\lambda\},\(190\)andFTF\_\{T\}is strictly increasing onm∈\[0,1\]m\\in\[0,1\]forλ<1\\lambda<1, so the interior stationary pointm†m^\{\\dagger\}is unique\. It is the barrier of Sec\.[IV\.2](https://arxiv.org/html/2609.10976#S4.SS2), with the zero\-temperature rootλ​m†=1−1−2​α\\lambda m^\{\\dagger\}=1\-\\sqrt\{1\-2\\alpha\}\. The continuum endpoint criterion,α<FT​\(1\)=λ−λ22​\(1\+T\)\\alpha<F\_\{T\}\(1\)=\\lambda\-\\frac\{\\lambda^\{2\}\}\{2\}\(1\+T\), is what a smooth treatment ofppwould predict for the spinodal, and the next paragraph corrects it\.

One\-copy defection and the spinodal\.The attention weight moves on the latticep=1−k/np=1\-k/n, with spacingλ​T\\lambda Tthat is not small at finite temperature, so the local stability of retrieval is decided by the nearest sector,φR​\(1\)≷φR​\(1−1/n\)\\varphi\_\{\\mathrm\{R\}\}\(1\)\\gtrless\\varphi\_\{\\mathrm\{R\}\}\(1\-1/n\), not by the endpoint slope\. We evaluate the one\-copy defection directly, which also serves as an independent check of Eq\. \([188](https://arxiv.org/html/2609.10976#A3.E188)\)\. Move one copy fromξ1\(h,v\)\\xi^\{\(h,v\)\}\_\{1\}to a patternμ\\muwith overlapx≔ξμ\(h,v\)⋅ξ1\(h,v\)/Nvx\\coloneq\\xi^\{\(h,v\)\}\_\{\\mu\}\\cdot\\xi^\{\(h,v\)\}\_\{1\}/N\_\{v\}and norm‖ξμ\(h,v\)‖2/Nv=1\+ε\\left\\lVert\\xi^\{\(h,v\)\}\_\{\\mu\}\\right\\rVert^\{2\}/N\_\{v\}=1\+\\varepsilon\. Conditioned onξ1\(h,v\)\\xi^\{\(h,v\)\}\_\{1\}, the pair rate function is theM=2M=2case of Eq\. \([64](https://arxiv.org/html/2609.10976#S4.E64)\),

I⁡\(x,ε\)=x22\+I1​\(1\+ε−x2\),I\(x,\\varepsilon\)=\\frac\{x^\{2\}\}\{2\}\+I\_\{1\}\\big\(1\+\\varepsilon\-x^\{2\}\\big\),\(191\)the longitudinal Gaussian cost plus the transverse norm cost\. Note that conditioning on the overlapxxmakes the typical norm excessx2x^\{2\}, at no additional cost\. The energy gain of the defection is

Δ⁡\(x,ε\)\\displaystyle\\Delta\(x,\\varepsilon\)=λ2​n​\[\(n−1\)2\+2​\(n−1\)​x\+1\+ε−n2\]\\displaystyle=\\frac\{\\lambda\}\{2n\}\\Big\[\(n\-1\)^\{2\}\+2\(n\-1\)x\+1\+\\varepsilon\-n^\{2\}\\Big\]\(192\)=λ~\(x−1\)\+λ2​T2ε,λ~≔λ\(1−λT\),\\displaystyle=\\tilde\{\\lambda\}\\,\(x\-1\)\+\\frac\{\\lambda^\{2\}T\}\{2\}\\,\\varepsilon,\\qquad\\tilde\{\\lambda\}\\coloneq\\lambda\(1\-\\lambda T\),an exchange term, the loss of alignment with then−1n\-1copies left behind \(the defector does not interact with itself, hence the thinning factor1−λ​T1\-\\lambda T\), and a norm term, an energy channel of orderTTthat favors defection toward patterns of large norm\. The defection sector carries the exponentsupI≤α\[α−I⁡\(x,ε\)\+Δ⁡\(x,ε\)\]\\sup\_\{I\\leq\\alpha\}\[\\alpha\-I\(x,\\varepsilon\)\+\\Delta\(x,\\varepsilon\)\]\. In the annealed regime the interior stationary point is, in the variabley≔1\+ε−x2y\\coloneq 1\+\\varepsilon\-x^\{2\},

y∗=11−λ2​T,x∗=λ~1−λ2​T=λ⁡\(1−λ​T\)1−λ2​T,y^\{\\ast\}=\\frac\{1\}\{1\-\\lambda^\{2\}T\},\\qquad x^\{\\ast\}=\\frac\{\\tilde\{\\lambda\}\}\{1\-\\lambda^\{2\}T\}=\\frac\{\\lambda\(1\-\\lambda T\)\}\{1\-\\lambda^\{2\}T\},\(193\)and collecting the terms \(the1/n21/n^\{2\}contributions cancel exactly\) gives

Δ​φdef=α−λ⁡\(2−λ−λ​T\)2​\(1−λ2​T\)−12​log⁡\(1−λ2​T\),\\Delta\\varphi\_\{\\mathrm\{def\}\}=\\alpha\-\\frac\{\\lambda\(2\-\\lambda\-\\lambda T\)\}\{2\(1\-\\lambda^\{2\}T\)\}\-\\frac\{1\}\{2\}\\log\\big\(1\-\\lambda^\{2\}T\\big\),\(194\)which coincides exactly withφR​\(1−1/n\)−φR​\(1\)\\varphi\_\{\\mathrm\{R\}\}\(1\-1/n\)\-\\varphi\_\{\\mathrm\{R\}\}\(1\)computed from Eq\. \([188](https://arxiv.org/html/2609.10976#A3.E188)\): the collective derivation and the single\-defection computation, which optimizes the destination geometry independently, agree term by term\. Retrieval is locally stable whileΔ​φdef<0\\Delta\\varphi\_\{\\mathrm\{def\}\}<0, which is the spinodal Eq\. \([68](https://arxiv.org/html/2609.10976#S4.E68)\)\. Its small\-TTexpansion,αc=λ−λ22−λ2​T2​\[1\+\(1−λ\)2\]\+O⁡\(T2\)\\alpha\_\{c\}=\\lambda\-\\frac\{\\lambda^\{2\}\}\{2\}\-\\frac\{\\lambda^\{2\}T\}\{2\}\[1\+\(1\-\\lambda\)^\{2\}\]\+O\(T^\{2\}\), lies strictly below the continuum criterion below Eq\. \([190](https://arxiv.org/html/2609.10976#A3.E190)\) except atλ=1\\lambda=1: the continuum treatment resolves barriers thinner than one attention quantum and therefore overestimates the capacity at anyT\>0T\>0, as stated in the main text\. The defection saddle Eq\. \([193](https://arxiv.org/html/2609.10976#A3.E193)\) also encodes the physics: the optimal destination is slightly aligned with the signal \(x∗\>0x^\{\\ast\}\>0, the one\-copy version of the linear responsec=λ​mc=\\lambda m\) and slightly long \(y∗\>1y^\{\\ast\}\>1, the norm channel opened by the temperature\)\.

Frozen ceiling\.The annealed defection assumed that destinations with the optimal\(x∗,ε∗\)\(x^\{\\ast\},\\varepsilon^\{\\ast\}\)exist, i\.e\.I⁡\(x∗,ε∗\)≤αI\(x^\{\\ast\},\\varepsilon^\{\\ast\}\)\\leq\\alpha\. AsT→0T\\to 0,I⁡\(x∗,ε∗\)→λ2/2I\(x^\{\\ast\},\\varepsilon^\{\\ast\}\)\\to\\lambda^\{2\}/2, so forα≲λ2/2\\alpha\\lesssim\\lambda^\{2\}/2\(sharp attention\) they do not, and the supremum is attained on the boundaryI=αI=\\alpha: the defection freezes onto the most favorable pattern actually present, the one\-copy analogue of the frozen branch of Eq\. \([60](https://arxiv.org/html/2609.10976#S4.E60)\)\. On the boundary the counting term vanishes and one maximizesΔ⁡\(x,ε\)\\Delta\(x,\\varepsilon\)alone subject toI⁡\(x,ε\)=αI\(x,\\varepsilon\)=\\alpha\. The Lagrange conditions,∂x\\partial\_\{x\}:λ~=η​x/y\\tilde\{\\lambda\}=\\eta\\,x/yand∂ε\\partial\_\{\\varepsilon\}:λ2​T2=η2​\(1−1/y\)\\frac\{\\lambda^\{2\}T\}\{2\}=\\frac\{\\eta\}\{2\}\(1\-1/y\), combine into

y−1=λ​T​x1−λ​T:y\-1=\\frac\{\\lambda T\\,x\}\{1\-\\lambda T\}:\(195\)the optimal extreme destination shifts toward norm excess asTTgrows, and returns to the conditionally typical norm \(y→1y\\to 1\) atT=0T=0\. SinceI1​\(y\)=O⁡\(\(y−1\)2\)I\_\{1\}\(y\)=O\(\(y\-1\)^\{2\}\), the transverse cost enters the constraint only atO⁡\(T2\)O\(T^\{2\}\), sox=2​α​\(1\+O⁡\(T2\)\)x=\\sqrt\{2\\alpha\}\\,\(1\+O\(T^\{2\}\)\): the extreme overlap itself is temperature independent at this order, and the temperature enters only through the gain,

Δ=λ⁡\(x−1\)\+λ2​T​\[\(1−x\)\+x22\]\+O⁡\(T2\)\.\\Delta=\\lambda\(x\-1\)\+\\lambda^\{2\}T\\Big\[\(1\-x\)\+\\frac\{x^\{2\}\}\{2\}\\Big\]\+O\(T^\{2\}\)\.\(196\)Near the thresholdx≈1x\\approx 1the bracket equals12\\frac\{1\}\{2\}, so stabilityΔ<0\\Delta<0requiresx<xc=1−λ​T/2\+O⁡\(T2\)x<x\_\{c\}=1\-\\lambda T/2\+O\(T^\{2\}\), i\.e\.α<xc2/2\\alpha<x\_\{c\}^\{2\}/2, which is the ceiling Eq\. \([69](https://arxiv.org/html/2609.10976#S4.E69)\)\. The two order\-TTmechanisms quoted in the main text are the two terms of Eq\. \([192](https://arxiv.org/html/2609.10976#A3.E192)\), the exchange thinning of the condensate and the norm channel\. The boundary optimization ranges over all patterns actually present, so it automatically includes the defection toward the maximal\-norm pattern \(x≈0x\\approx 0,ε=εmax\\varepsilon=\\varepsilon\_\{\\max\}\), the first step of nucleation toward the F phase\. The solution of Eq\. \([195](https://arxiv.org/html/2609.10976#A3.E195)\) dominates it at low temperature, so the one\-copy instability channels are exhausted by this computation\. The finite\-temperature capacity thus follows Eq\. \([68](https://arxiv.org/html/2609.10976#S4.E68)\) on the annealed side and Eq\. \([69](https://arxiv.org/html/2609.10976#S4.E69)\) on the frozen side of the branch switch atα≈λ2/2\+O⁡\(T\)\\alpha\\approx\\lambda^\{2\}/2\+O\(T\), as in Sec\.[IV\.2](https://arxiv.org/html/2609.10976#S4.SS2)\. The switch is the finite\-temperature continuation of the branch pointλ=2​α\\lambda=\\sqrt\{2\\alpha\}of Eq\. \([60](https://arxiv.org/html/2609.10976#S4.E60)\)\.

### C\.4Complete phase diagram for continuous temperature

The results of Appendix[C\.3](https://arxiv.org/html/2609.10976#A3.SS3)are exact at the integer temperature pointsn=β/λ∈ℤ\>0n=\\beta/\\lambda\\in\\mathbb\{Z\}\_\{\>0\}\. In this appendix we reconstruct the phase diagram from the visible representation, which is defined at every real temperature from the outset, and show that all formulas of the main text hold as stated for arbitrary realβ\\beta\.666The results in this subsection are obtained with the help of Claude Fable 5\.Along the way the fate of the fluctuation determinant and of the hidden\-sector zero mode is clarified, and the analytic continuation off the integer lattice is proved unique\. Throughout,s≔β/λs\\coloneq\\beta/\\lambdadenotes the real \(in the uniqueness argument, complex\) continuation of the copy number, withs=ns=non the lattice\.

Constrained partition function\.Fix the visible overlapm=v⋅ξ1\(h,v\)/Nvm=v\\cdot\\xi^\{\(h,v\)\}\_\{1\}/N\_\{v\}with a typical patternξ1\(h,v\)\\xi^\{\(h,v\)\}\_\{1\}as a reaction coordinate and defineZ⁡\(m\)≔∫𝒟​v​δ​\(v⋅ξ1\(h,v\)/Nv−m\)​e−β​E​\(v\)Z\(m\)\\coloneq\\int\\mathcal\{D\}v\\,\\delta\(v\\cdot\\xi^\{\(h,v\)\}\_\{1\}/N\_\{v\}\-m\)\\,e^\{\-\\beta E\(v\)\}, withEEthe effective energy Eq\. \([54](https://arxiv.org/html/2609.10976#S4.E54)\) and𝒟​v\\mathcal\{D\}vthe Gaussian reference measure\. The orthogonal decompositionv=m​ξ1\(h,v\)\+v⟂v=m\\,\\xi^\{\(h,v\)\}\_\{1\}\+v\_\{\\perp\}fixes the longitudinal component with unit Jacobian, and the fields become

a1=λ​Nv​m,aμ=λ⁡\(Nv​m​xμ\+ξμ⟂⋅v⟂\),μ≥2,a\_\{1\}=\\lambda N\_\{v\}m,\\qquad a\_\{\\mu\}=\\lambda\\big\(N\_\{v\}m\\,x\_\{\\mu\}\+\\xi^\{\\perp\}\_\{\\mu\}\\cdot v\_\{\\perp\}\\big\),\\quad\\mu\\geq 2,\(197\)withxμ≔ξμ\(h,v\)⋅ξ1\(h,v\)/Nvx\_\{\\mu\}\\coloneq\\xi^\{\(h,v\)\}\_\{\\mu\}\\cdot\\xi^\{\(h,v\)\}\_\{1\}/N\_\{v\},ξμ⟂≔ξμ\(h,v\)−xμ​ξ1\(h,v\)\\xi^\{\\perp\}\_\{\\mu\}\\coloneq\\xi^\{\(h,v\)\}\_\{\\mu\}\-x\_\{\\mu\}\\xi^\{\(h,v\)\}\_\{1\}, andyμ≔‖ξμ⟂‖2/Nvy\_\{\\mu\}\\coloneq\\left\\lVert\\xi^\{\\perp\}\_\{\\mu\}\\right\\rVert^\{2\}/N\_\{v\}\. Splitting off the signal and writing the leak sum Eq\. \([59](https://arxiv.org/html/2609.10976#S4.E59)\) in the present variables,L=∑μ≥2eaμL=\\sum\_\{\\mu\\geq 2\}e^\{a\_\{\\mu\}\}, together withu≔e−a1​Lu\\coloneq e^\{\-a\_\{1\}\}L,

Z⁡\(m\)=eNv​β​\(m−m2/2\)​⟨\(1\+u\)s⟩⟂,Z\(m\)=e^\{N\_\{v\}\\beta\(m\-m^\{2\}/2\)\}\\,\\big\\langle\(1\+u\)^\{s\}\\big\\rangle\_\{\\perp\},\(198\)where⟨⋅⟩⟂\\langle\\cdot\\rangle\_\{\\perp\}averages overv⟂∼𝒩⁡\(0,β−1​INv−1\)v\_\{\\perp\}\\sim\\mathcal\{N\}\(0,\\beta^\{\-1\}I\_\{N\_\{v\}\-1\}\)and is normalized so that the termu0u^\{0\}contributes11\. Conditioned onξ1\(h,v\)\\xi^\{\(h,v\)\}\_\{1\}, the variables\(xμ,yμ\)\(x\_\{\\mu\},y\_\{\\mu\}\)are independent acrossμ\\mu, with joint rateI⁡\(x,y\)=x2/2\+I1​\(y\)I\(x,y\)=x^\{2\}/2\+I\_\{1\}\(y\), which is Eq\. \([191](https://arxiv.org/html/2609.10976#A3.E191)\) in the variablesy=1\+ε−x2y=1\+\\varepsilon\-x^\{2\}\. The counting dichotomy of Appendix[C\.1](https://arxiv.org/html/2609.10976#A3.SS1)applies to them unchanged\.

Sector decomposition of the retrieval basin\.By Newton’s generalized binomial theorem, for every real \(indeed complex\) exponentss,

\(1\+u\)s=∑k=0∞\(sk\)​uk,\(sk\)=s\(s−1\)⋯\(s−k\+1\)k\!,\(1\+u\)^\{s\}=\\sum\_\{k=0\}^\{\\infty\}\\binom\{s\}\{k\}\\,u^\{k\},\\qquad\\binom\{s\}\{k\}=\\frac\{s\(s\-1\)\\cdots\(s\-k\+1\)\}\{k\!\},\(199\)convergent foru<1u<1\. In the retrieval basinuuis exponentially small, term\-by\-term evaluation is justified, and the expansion decomposesZ⁡\(m\)Z\(m\)into sectors labeled by the integer defection numberkk, of condensed weightpk=1−k​λ​Tp\_\{k\}=1\-k\\lambda T: the attention lattice and its quantumλ​T\\lambda Tare properties of the model at every real temperature, not artifacts of integernn\. Only the combinatorial coefficients\(sk\)\\binom\{s\}\{k\}are continued, and they are subexponential inNvN\_\{v\}\. The convergence conditionu<1u<1is the retrieval condition itself, so the expansion is precisely the excitation theory of the retrieval state\. Thek=0k=0sector givesφ0​\(m\)=β⁡\(m−m2/2\)\\varphi\_\{0\}\(m\)=\\beta\(m\-m^\{2\}/2\), maximal atm=1m=1with valueβ/2\\beta/2\. Fork=1k=1, using the Gaussian generating function⟨eλ​ξμ⟂⋅v⟂⟩⟂=eNv​\(λ2​T/2\)​yμ\\langle e^\{\\lambda\\,\\xi^\{\\perp\}\_\{\\mu\}\\cdot v\_\{\\perp\}\}\\rangle\_\{\\perp\}=e^\{N\_\{v\}\(\\lambda^\{2\}T/2\)\\,y\_\{\\mu\}\}and the counting dichotomy,

Δ1​\(m\)=supI⁡\(x,y\)≤α\[α−I⁡\(x,y\)\+λ​m​\(x−1\)\+λ2​T2​y\]\.\\Delta\_\{1\}\(m\)=\\sup\_\{I\(x,y\)\\leq\\alpha\}\\Big\[\\alpha\-I\(x,y\)\+\\lambda m\(x\-1\)\+\\frac\{\\lambda^\{2\}T\}\{2\}y\\Big\]\.\(200\)The annealed stationary point inyyis governed by the identity

−I1​\(y∗\)\+t2​y∗=−12​log⁡\(1−t\),y∗=11−t,\-I\_\{1\}\(y\_\{\\ast\}\)\+\\frac\{t\}\{2\}\\,y\_\{\\ast\}=\-\\frac\{1\}\{2\}\\log\(1\-t\),\\qquad y\_\{\\ast\}=\\frac\{1\}\{1\-t\},\(201\)here att=λ2​Tt=\\lambda^\{2\}T, while the stationary overlap isx∗=λ​mx\_\{\\ast\}=\\lambda m, so that

Δ1​\(m\)=α−λ​m​\(1−λ​m2\)−12​log⁡\(1−λ2​T\),\\Delta\_\{1\}\(m\)=\\alpha\-\\lambda m\\Big\(1\-\\frac\{\\lambda m\}\{2\}\\Big\)\-\\frac\{1\}\{2\}\\log\\big\(1\-\\lambda^\{2\}T\\big\),\(202\)valid whileI⁡\(x∗,y∗\)=λ2​m2/2\+κ⁡\(λ2​T\)≤αI\(x\_\{\\ast\},y\_\{\\ast\}\)=\\lambda^\{2\}m^\{2\}/2\+\\kappa\(\\lambda^\{2\}T\)\\leq\\alpha\. Maximizingφ0\+Δ1\\varphi\_\{0\}\+\\Delta\_\{1\}overmmgivesm1∗=\(1−λ​T\)/\(1−λ2​T\)m^\{\\ast\}\_\{1\}=\(1\-\\lambda T\)/\(1\-\\lambda^\{2\}T\), which is the overlap amplificationp1/A1p\_\{1\}/A\_\{1\}of Eq\. \([189](https://arxiv.org/html/2609.10976#A3.E189)\) evaluated in this sector, and the valueφR​\(1−λ​T\)\\varphi\_\{\\mathrm\{R\}\}\(1\-\\lambda T\), i\.e\. exactly the copy result Eq\. \([188](https://arxiv.org/html/2609.10976#A3.E188)\) atp=1−1/np=1\-1/n, now for arbitrary realβ\\beta\. In particular the one\-quantum stability criterion, hence the spinodal Eq\. \([68](https://arxiv.org/html/2609.10976#S4.E68)\), holds at every real temperature\. For generalkkwith distinct destinations, the transverse average factorizes over theNv−1N\_\{v\}\-1transverse coordinates\. Per coordinate the destinations contributew=∑j=1kzj∼𝒩⁡\(0,k\)w=\\sum\_\{j=1\}^\{k\}z\_\{j\}\\sim\\mathcal\{N\}\(0,k\)and

𝔼eλ2​T2​w2=\(1−kλ2T\)−1/2,\\mathbb\{E\}\\,e^\{\\frac\{\\lambda^\{2\}T\}\{2\}w^\{2\}\}=\\big\(1\-k\\lambda^\{2\}T\\big\)^\{\-1/2\},\(203\)which resums the destination norms and their mutual response to all orders\. The longitudinal tilts contributeeNv​λ2​m2/2e^\{N\_\{v\}\\lambda^\{2\}m^\{2\}/2\}each, independently, sincemmis held fixed\. With the signal losse−k​λ​m​Nve^\{\-k\\lambda mN\_\{v\}\}and the countingek​α​Nve^\{k\\alpha N\_\{v\}\},

Δk​\(m\)=k⁡\[α−λ​m​\(1−λ​m2\)\]−12​log⁡\(1−k​λ2​T\),\\Delta\_\{k\}\(m\)=k\\Big\[\\alpha\-\\lambda m\\Big\(1\-\\frac\{\\lambda m\}\{2\}\\Big\)\\Big\]\-\\frac\{1\}\{2\}\\log\\big\(1\-k\\lambda^\{2\}T\\big\),\(204\)and in terms ofpk=1−k​λ​Tp\_\{k\}=1\-k\\lambda TandAk=1−λ⁡\(1−pk\)=1−k​λ2​TA\_\{k\}=1\-\\lambda\(1\-p\_\{k\}\)=1\-k\\lambda^\{2\}T,

φk​\(m\)\\displaystyle\\varphi\_\{k\}\(m\)=φ0​\(m\)\+Δk​\(m\)\\displaystyle=\\varphi\_\{0\}\(m\)\+\\Delta\_\{k\}\(m\)\(205\)=βλ​\(1−pk\)​α−12​log⁡Ak\+β⁡\(pk​m−Ak2​m2\)\.\\displaystyle=\\frac\{\\beta\}\{\\lambda\}\(1\-p\_\{k\}\)\\,\\alpha\-\\frac\{1\}\{2\}\\log A\_\{k\}\+\\beta\\Big\(p\_\{k\}m\-\\frac\{A\_\{k\}\}\{2\}m^\{2\}\\Big\)\.This is the two\-variable Landau surfaceφ⁡\(m,p\)\\varphi\(m,p\)of the model, derived exactly on the latticepk=1−k​λ​Tp\_\{k\}=1\-k\\lambda Tat arbitrary real temperature\. Maximizing overmmreturns Eq\. \([188](https://arxiv.org/html/2609.10976#A3.E188)\), so the entire retrieval branch of Appendix[C\.3](https://arxiv.org/html/2609.10976#A3.SS3)is recovered without any integer assumption\. Themmdirection is exactly continuous and smooth, while theppdirection is quantized: the discreteness of the finite\-temperature physics resides solely in the attention coordinate\. The validity conditions are the annealed counting condition per quantum,α≥λ2​m2/2\+κ⁡\(λ2​T\)\\alpha\\geq\\lambda^\{2\}m^\{2\}/2\+\\kappa\(\\lambda^\{2\}T\), and the positivityAk\>0A\_\{k\}\>0, i\.e\.k<1/\(λ2​T\)k<1/\(\\lambda^\{2\}T\)\. The physical lattice ends atpk≥0p\_\{k\}\\geq 0, i\.e\.k≤sk\\leq s, beyond which the binomial coefficients alternate in sign and the bulk is described by the dual expansion below\.

Co\-defection channels\.Theuku^\{k\}terms with coinciding destination indices describejjcopies defecting to the same pattern\. The coherent coupling is enhanced byj2j^\{2\}, giving

Δjco​\(m\)=α−j​λ​m\+j2​λ2​m22−12​log⁡\(1−j2​λ2​T\),\\Delta\_\{j\}^\{\\mathrm\{co\}\}\(m\)=\\alpha\-j\\lambda m\+\\frac\{j^\{2\}\\lambda^\{2\}m^\{2\}\}\{2\}\-\\frac\{1\}\{2\}\\log\\big\(1\-j^\{2\}\\lambda^\{2\}T\\big\),\(206\)with tilted overlapx∗=j​λ​mx\_\{\\ast\}=j\\lambda m, which requires the existence conditionI=j2​λ2​m2/2≤αI=j^\{2\}\\lambda^\{2\}m^\{2\}/2\\leq\\alpha\. Formally Eq\. \([206](https://arxiv.org/html/2609.10976#A3.E206)\) would overtakejjseparate defections forα<λ2​m2\\alpha<\\lambda^\{2\}m^\{2\}, but that region violates the existence condition already atj=2j=2\(which demandsα≥2​λ2​m2\\alpha\\geq 2\\lambda^\{2\}m^\{2\}\), so the annealed co\-defection channel is never dominant where it is valid\. The frozen co\-defection,jjquanta onto the most favorable pattern actually present, gains onlyjjtimes the single\-quantum exchange term plus theO⁡\(T\)O\(T\)norm enhancement, and therefore does not destabilize retrieval before the one\-copy channel does\. The spinodal thus remains the one\-copy criterion,min\\minof Eq\. \([68](https://arxiv.org/html/2609.10976#S4.E68)\) and Eq\. \([69](https://arxiv.org/html/2609.10976#S4.E69)\), at every real temperature\. The role of the co\-defection ladder is instead to provide the nucleation path toward the condensed phases, as shown below\. Similarly, the frozen\-leak evaluation of Appendix[C\.3](https://arxiv.org/html/2609.10976#A3.SS3)carries over at fixedmm: the boundary optimization givesΔ1froz​\(m\)=λ​m​\(2​α−1\)\+λ2​T/2\+O⁡\(T2\)\\Delta\_\{1\}^\{\\mathrm\{froz\}\}\(m\)=\\lambda m\(\\sqrt\{2\\alpha\}\-1\)\+\\lambda^\{2\}T/2\+O\(T^\{2\}\), and relaxingmmto its optimumm∗=1−λ​T​\(1−2​α\)m\_\{\\ast\}=1\-\\lambda T\(1\-\\sqrt\{2\\alpha\}\)yields

Δfroz=λ⁡\(2​α−1\)\+λ2​T​\[1−2​α\+α\]\+O⁡\(T2\),\\Delta^\{\\mathrm\{froz\}\}=\\lambda\(\\sqrt\{2\\alpha\}\-1\)\+\\lambda^\{2\}T\\big\[1\-\\sqrt\{2\\alpha\}\+\\alpha\\big\]\+O\(T^\{2\}\),\(207\)which agrees term by term with Eq\. \([196](https://arxiv.org/html/2609.10976#A3.E196)\) atx=2​αx=\\sqrt\{2\\alpha\}\(the exchange factor1−λ​T1\-\\lambda Tof the copy computation reappears here as themm\-relaxation term: the grouping of terms differs, the total does not\), and reproduces the ceiling Eq\. \([69](https://arxiv.org/html/2609.10976#S4.E69)\) at real temperature\.

Bulk branches and endpoint closure\.Outside the basin \(u\>1u\>1\) one expands dually,\(ea1\+L\)s=Ls​\(1\+u−1\)s\(e^\{a\_\{1\}\}\+L\)^\{s\}=L^\{s\}\(1\+u^\{\-1\}\)^\{s\}: leak\-dominated states into which signal is injected\. For the paramagnetic branch the annealed leak exponentNv−1​log⁡L=α\+λ2​ρ2/2N\_\{v\}^\{\-1\}\\log L=\\alpha\+\\lambda^\{2\}\\rho^\{2\}/2is linear inρ2=‖v‖2/Nv\\rho^\{2\}=\\left\\lVert v\\right\\rVert^\{2\}/N\_\{v\}, so the visible integral is exactly Gaussian and

φP​\(m\)=s​α−12​log⁡\(1−λ\)−β⁡\(1−λ\)2​m2,\\varphi\_\{\\mathrm\{P\}\}\(m\)=s\\alpha\-\\frac\{1\}\{2\}\\log\(1\-\\lambda\)\-\\frac\{\\beta\(1\-\\lambda\)\}\{2\}m^\{2\},\(208\)maximal atm=0m=0where it reproducesφP\\varphi\_\{\\mathrm\{P\}\}of Appendix[C\.3](https://arxiv.org/html/2609.10976#A3.SS3)withn→sn\\to sreal\. The longitudinal stiffnessβ⁡\(1−λ\)\\beta\(1\-\\lambda\)is the inverse of the condensate size Eq\. \([183](https://arxiv.org/html/2609.10976#A3.E183)\)\. For the condensed sector, a statev≈ξμ\(h,v\)v\\approx\\xi^\{\(h,v\)\}\_\{\\mu\}with overlapx=mx=mand norm1\+ε=m2\+y1\+\\varepsilon=m^\{2\}\+ycarries

φcond\(m\)=supy:m22\+I1​\(y\)≤α\[α−m22−I1\(y\)\+β2\(m2\+y\)\]\.\\varphi\_\{\\mathrm\{cond\}\}\(m\)=\\sup\_\{y:\\,\\frac\{m^\{2\}\}\{2\}\+I\_\{1\}\(y\)\\leq\\alpha\}\\Big\[\\alpha\-\\frac\{m^\{2\}\}\{2\}\-I\_\{1\}\(y\)\+\\frac\{\\beta\}\{2\}\\big\(m^\{2\}\+y\\big\)\\Big\]\.\(209\)The interior stationary point,y∗=\(1−β\)−1y\_\{\\ast\}=\(1\-\\beta\)^\{\-1\}forβ<1\\beta<1by Eq\. \([201](https://arxiv.org/html/2609.10976#A3.E201)\) att=βt=\\beta, gives

φC​\(m\)=α−12​log⁡\(1−β\)−1−β2​m2,\\varphi\_\{\\mathrm\{C\}\}\(m\)=\\alpha\-\\frac\{1\}\{2\}\\log\(1\-\\beta\)\-\\frac\{1\-\\beta\}\{2\}m^\{2\},\(210\)whose existence conditionβ<1\\beta<1appears automatically\. The boundary evaluation,I1​\(y\)=α−m2/2I\_\{1\}\(y\)=\\alpha\-m^\{2\}/2, gives

φF​\(m\)=β2​\[m2\+1\+εmax​\(α−m22\)\],m2≤2​α,\\varphi\_\{\\mathrm\{F\}\}\(m\)=\\frac\{\\beta\}\{2\}\\Big\[m^\{2\}\+1\+\\varepsilon\_\{\\max\}\\Big\(\\alpha\-\\frac\{m^\{2\}\}\{2\}\\Big\)\\Big\],\\qquad m^\{2\}\\leq 2\\alpha,\(211\)whereεmax​\(a\)\\varepsilon\_\{\\max\}\(a\)solvesI1​\(1\+ε\)=aI\_\{1\}\(1\+\\varepsilon\)=a\. Since∂mφF=−βm/εmax<0\\partial\_\{m\}\\varphi\_\{\\mathrm\{F\}\}=\-\\beta m/\\varepsilon\_\{\\max\}<0, the frozen branch is maximal atm=0m=0: the extreme\-norm pattern is typically orthogonal toξ1\(h,v\)\\xi^\{\(h,v\)\}\_\{1\}\. The freezing line at fixedmmisκ⁡\(β\)=α−m2/2\\kappa\(\\beta\)=\\alpha\-m^\{2\}/2, reducing atm=0m=0to Eq\. \([66](https://arxiv.org/html/2609.10976#S4.E66)\)\. The decisive structural fact is that the defection lattice closes exactly onto these bulk branches\. Substituting the terminusk=nk=nof the distinct\-defection ladder \(at integernn, withp=0p=0,k​λ=βk\\lambda=\\beta, andk​λ2​T=λk\\lambda^\{2\}T=\\lambda\) into Eq\. \([205](https://arxiv.org/html/2609.10976#A3.E205)\),

φn​\(m\)=n​α−12​log⁡\(1−λ\)−β⁡\(1−λ\)2​m2=φP​\(m\),\\varphi\_\{n\}\(m\)=n\\alpha\-\\frac\{1\}\{2\}\\log\(1\-\\lambda\)\-\\frac\{\\beta\(1\-\\lambda\)\}\{2\}m^\{2\}=\\varphi\_\{\\mathrm\{P\}\}\(m\),\(212\)and the terminusj=nj=nof the co\-defection ladder, whose gain−β2​m2\+β​m​x\+β2​y\-\\frac\{\\beta\}\{2\}m^\{2\}\+\\beta mx\+\\frac\{\\beta\}\{2\}ycompletes the square asβ2​\[1\+ε−\(m−x\)2\]\\frac\{\\beta\}\{2\}\[1\+\\varepsilon\-\(m\-x\)^\{2\}\], is optimized on the boundary atx=mx=mand closes onto Eq\. \([211](https://arxiv.org/html/2609.10976#A3.E211)\) \(interior stationarity giving Eq\. \([210](https://arxiv.org/html/2609.10976#A3.E210)\) instead\)\. The retrieval state is thus connected to the bulk phases by two quantized nucleation ladders living on a single surface, distinct defections leading to P and co\-defections to C or F\. At non\-integerssthe terminusk=sk=sis not a lattice point, but the dual expansion supplies the bulk values independently, and they agree with thep→0p\\to 0values of Eq\. \([205](https://arxiv.org/html/2609.10976#A3.E205)\), so the two expansions connect without a gap\.

Fluctuation determinant, zero mode, and uniqueness\.Three structural questions raised in Appendix[C\.2](https://arxiv.org/html/2609.10976#A3.SS2)are settled by this construction\.

\(a\)*One determinant, two interpretations\.*The only Gaussian fluctuation determinant surviving in the entire construction is the transverse visible one, and it admits two equivalent evaluations\. Averaging over the destination patterns first,𝔼ξ⟂​exp⁡\[λ​∑jξj⟂⋅v⟂\]=exp⁡\[k​λ22​‖v⟂‖2\]\\mathbb\{E\}\_\{\\xi^\{\\perp\}\}\\exp\[\\lambda\\sum\_\{j\}\\xi^\{\\perp\}\_\{j\}\\cdot v\_\{\\perp\}\]=\\exp\[\\frac\{k\\lambda^\{2\}\}\{2\}\\left\\lVert v\_\{\\perp\}\\right\\rVert^\{2\}\], which renormalizes the transverse stiffnessβ→β−k​λ2=β​Ak\\beta\\to\\beta\-k\\lambda^\{2\}=\\beta A\_\{k\}\. The ratio to the reference measure is

\(ββ−k​λ2\)\(Nv−1\)/2=Ak−\(Nv−1\)/2⟹−12logAk,\\Big\(\\frac\{\\beta\}\{\\beta\-k\\lambda^\{2\}\}\\Big\)^\{\\\!\(N\_\{v\}\-1\)/2\}=A\_\{k\}^\{\-\(N\_\{v\}\-1\)/2\}\\;\\Longrightarrow\\;\-\\frac\{1\}\{2\}\\log A\_\{k\},\(213\)so the logarithm in Eq\. \([204](https://arxiv.org/html/2609.10976#A3.E204)\) is the determinant of the visible transverse fluctuations softened by the attention leak,Ak\>0A\_\{k\}\>0is their stability condition, andAk→0A\_\{k\}\\to 0atk=sk=sis the pointλ→1\\lambda\\to 1: the paramagnetic instability\. Averaging overv⟂v\_\{\\perp\}first instead gives the Gram\-fluctuation evaluation Eq\. \([203](https://arxiv.org/html/2609.10976#A3.E203)\), whose large\-deviation form is the variational problemsupQ⟂\[−Ik​\(Q⟂\)\+λ2​T2​𝟏⊤​Q⟂​𝟏\]\\sup\_\{Q^\{\\perp\}\}\[\-I\_\{k\}\(Q^\{\\perp\}\)\+\\frac\{\\lambda^\{2\}T\}\{2\}\\bm\{1\}^\{\\top\}Q^\{\\perp\}\\bm\{1\}\]\. The symmetric decomposition puts the transverse eigenvalues at11and the coherent one atΛ∗=\(1−k​λ2​T\)−1\\Lambda\_\{\\ast\}=\(1\-k\\lambda^\{2\}T\)^\{\-1\}, with the same value by Eq\. \([201](https://arxiv.org/html/2609.10976#A3.E201)\)\. The two interpretations are the two sides of the Legendre pair Eq\. \([169](https://arxiv.org/html/2609.10976#A3.E169)\) realized sector by sector: the energy\-sidelogdet\\log\\detof a replica computation and the entropy\-side rate function of the counting derivation are one object viewed from its two convex\-dual sides\.

\(b\)*The fate of the zero mode\.*No pseudo\-determinant appears above, and the reason is instructive\. The visible representation is written in terms of the log\-sum\-exp, i\.e\. after the softmax gauge orbit of Appendix[C\.2](https://arxiv.org/html/2609.10976#A3.SS2)has been fixed\. The divergent zero\-mode volumeV0V\_\{0\}of Eq\. \([161](https://arxiv.org/html/2609.10976#A3.E161)\) is a configuration\-independent constant absorbed into the reference measure\. Conversely, a Gaussian fluctuation expansion of theffrepresentation around a retrieval vertex is structurally degenerate: by Eq\. \([160](https://arxiv.org/html/2609.10976#A3.E160)\),det′Hess=Nh​∏μfμ→0\{\\det\}^\{\\prime\}\\operatorname\{Hess\}=N\_\{h\}\\prod\_\{\\mu\}f\_\{\\mu\}\\to 0asf→e1f\\to e\_\{1\}, since every non\-condensed weight is exponentially small, reflecting the divergence of the entropy curvature∂f2\(f​log⁡f\)=1/f\\partial^\{2\}\_\{f\}\(f\\log f\)=1/fat the boundary of the simplex\. A Gaussian theory of fluctuations around the vertex therefore does not exist\. The correct fluctuation theory is the discrete sector decomposition itself, whose elementary excitations are the quantized defections of attention weightλ​T\\lambda Tand whose exact generating function is the binomial series Eq\. \([199](https://arxiv.org/html/2609.10976#A3.E199)\)\. This also sharpens the distinction, noted in Sec\.[IV\.1\.1](https://arxiv.org/html/2609.10976#S4.SS1.SSS1), from a variant model in which the attention entropy is rescaled by hand to be extensive: such a model equilibrates at an interior saddle point with macroscopically spread attention, where1/f1/fremains finite and the Gaussian expansion is legitimate\. The two models differ already at the level of their fluctuation theories, discrete versus Gaussian\.

\(c\)*Uniqueness of the continuation\.*At finiteNvN\_\{v\}and fixed disorder,Z⁡\(m,β\)Z\(m;\\beta\)is defined by the visible integral for every realβ\>0\\beta\>0and is analytic inβ\\beta\. The multinomial expansion at integernnand the binomial expansion Eq\. \([199](https://arxiv.org/html/2609.10976#A3.E199)\) are exact rewritings of this single function on their respective domains, so within the retrieval basin there is no continuation to choose\. The only question is the exchange of the thermodynamic limit with the continuation, and in the basin the series Eq\. \([199](https://arxiv.org/html/2609.10976#A3.E199)\) converges locally uniformly inβ\\beta, the dominant sector index is finite, and eachΔk\\Delta\_\{k\}is analytic by Eq\. \([204](https://arxiv.org/html/2609.10976#A3.E204)\), soφ⁡\(m,β\)\\varphi\(m;\\beta\)is analytic up to branch switches and continuous across them\. Agreement with the copy values at the integer points was verified at Eq\. \([202](https://arxiv.org/html/2609.10976#A3.E202)\)\. A stronger statement holds: the integer\-point data alone already determine the answer\. By Carlson’s theorem\[[59](https://arxiv.org/html/2609.10976#bib.bib59), Sec\. 9\], a function analytic inℜ⁡s≥0\\Re s\\geq 0, of exponential type with\|F⁡\(s\)\|≤C​eτ​\|s\|\|F\(s\)\|\\leq Ce^\{\\tau\|s\|\}for someτ<π\\tau<\\pi, and vanishing on the non\-negative integers, vanishes identically\. In the basin,\|\(1\+u\)s\|=\(1\+u\)ℜ⁡s≤2ℜ⁡s\|\(1\+u\)^\{s\}\|=\(1\+u\)^\{\\Re s\}\\leq 2^\{\\Re s\}, so the candidate functions have type at mostlog⁡2<π\\log 2<\\pi, and the difference of any two continuations that agree on the lattice is identically zero\. The growth condition is what excludes the spurious continuationssin⁡\(π​s\)\\sin\(\\pi s\), of type exactlyπ\\pi\. With it, the caveat usually attached to integer\-power representations is removed for the entire retrieval basin\.

Assembly of the phase diagram\.Table[3](https://arxiv.org/html/2609.10976#A3.T3)collects the branches constructed above with their validity conditions\. All live on the singlemmaxis and are connected through the defection lattice by the endpoint closure\.

Table 3:Branches of the constrained rateφ⁡\(m\)\\varphi\(m\)at real temperature\. The retrieval sectors live on the attention latticepk=1−k​λ​Tp\_\{k\}=1\-k\\lambda T\. P, C, and F are maximal atm=0m=0, and R atp=1p=1\.![Refer to caption](https://arxiv.org/html/2609.10976v1/fig/modelB_phase_diagram_lam05_lam15_full.png)Figure 5:Same as Fig\.[4](https://arxiv.org/html/2609.10976#S4.F4), with the dynamical content of the real\-temperature construction overlaid\. The blue dashed curves are contours of the escape barrier Eq\. \([215](https://arxiv.org/html/2609.10976#A3.E215)\), evaluated along the lower of the two escape ladders, at the levelsΔ​φb=0\.05\\Delta\\varphi\_\{\\mathrm\{b\}\}=0\.05,0\.10\.1,0\.20\.2,0\.50\.5, and11\. The barrier vanishes on the spinodal and grows toward smallα\\alphaand low temperature\. The insets show the constrained rateφ\\varphialong the two escape ladders at the starred reference points,\(α,T\)=\(0\.15,0\.40\)\(\\alpha,T\)=\(0\.15,0\.40\)in \(a\) and\(0\.14,0\.27\)\(0\.14,0\.27\)in \(b\), against the transferred attention weight1−p1\-p: the distinct\-defection ladder \(blue, annealed curve with rungs atpk=1−k​λ​Tp\_\{k\}=1\-k\\lambda T\) ends at the paramagnetic value, and the co\-defection ladder \(green\) at the condensed value, whose frozen terminus reproducesφF\\varphi\_\{\\mathrm\{F\}\}exactly\. In \(b\) the distinct ladder is cut off by the transverse stability conditionAk\>0A\_\{k\}\>0before reaching its terminus, and only the co\-defection ladder connects retrieval to the bulk\.The spinodal read off the table is the one\-quantum criterion established above, with the annealed and frozen branches Eq\. \([68](https://arxiv.org/html/2609.10976#S4.E68)\) and Eq\. \([69](https://arxiv.org/html/2609.10976#S4.E69)\), now as a statement at every real temperature\. The two branches follow from a single expression: eliminatingmmat its saddle,1−m=λ​T​\(1−x\)1\-m=\\lambda T\(1\-x\), reduces the one\-quantum criterion to the variational problem

Δ⁡\(α,λ,T\)=maxI⁡\(x,y\)≤α⁡\[α−I⁡\(x,y\)−λ⁡\(1−x\)\+λ2​T2​\(\(1−x\)2\+y\)\],\\Delta\(\\alpha;\\lambda,T\)=\\max\_\{I\(x,y\)\\leq\\alpha\}\\Big\[\\alpha\-I\(x,y\)\-\\lambda\(1\-x\)\+\\frac\{\\lambda^\{2\}T\}\{2\}\\big\(\(1\-x\)^\{2\}\+y\\big\)\\Big\],\(214\)with retrieval locally stable whileΔ<0\\Delta<0\. The interior stationary point of Eq\. \([214](https://arxiv.org/html/2609.10976#A3.E214)\) reproduces Eq\. \([68](https://arxiv.org/html/2609.10976#S4.E68)\), the boundary evaluation reproduces Eq\. \([69](https://arxiv.org/html/2609.10976#S4.E69)\), the switch between the two is the existence line of the annealed destination,α=I∗​\(T\)≔\(x∗\)2/2\+κ⁡\(λ2​T\)\\alpha=I^\{\\ast\}\(T\)\\coloneq\(x^\{\\ast\}\)^\{2\}/2\+\\kappa\(\\lambda^\{2\}T\)withx∗x^\{\\ast\}the annealed saddle Eq\. \([193](https://arxiv.org/html/2609.10976#A3.E193)\) \(the dash\-dotted curve of the figures\), and the rootΔ=0\\Delta=0is the spinodal drawn in Fig\.[4](https://arxiv.org/html/2609.10976#S4.F4)and Fig\.[5](https://arxiv.org/html/2609.10976#A3.F5)\. The same surface carries the barrier\. The increments of Eq\. \([204](https://arxiv.org/html/2609.10976#A3.E204)\) inkkstart atα−λ​m​\(1−λ​m/2\)\\alpha\-\\lambda m\(1\-\\lambda m/2\)and grow convexly through the logarithm, so forα<αc\\alpha<\\alpha\_\{c\}the sector weights first decrease and then increase along the ladder\. The continuum stationary point is Eq\. \([190](https://arxiv.org/html/2609.10976#A3.E190)\), and the barrier height of the retrieval state is

Δ​φb=β2−φR​\(p†\),\\Delta\\varphi\_\{\\mathrm\{b\}\}=\\frac\{\\beta\}\{2\}\-\\varphi\_\{\\mathrm\{R\}\}\(p^\{\\dagger\}\),\(215\)withp†p^\{\\dagger\}the sector corresponding to the rootm†m^\{\\dagger\}of Eq\. \([190](https://arxiv.org/html/2609.10976#A3.E190)\) throughm†=p†/A†m^\{\\dagger\}=p^\{\\dagger\}/A^\{\\dagger\}\. AtT=0T=0,λ​m†=1−1−2​α\\lambda m^\{\\dagger\}=1\-\\sqrt\{1\-2\\alpha\}gives a closed form, and at finiteTTa single numerical root suffices\. In the Arrhenius senseτesc∼eNv​Δ​φb\\tau\_\{\\mathrm\{esc\}\}\\sim e^\{N\_\{v\}\\Delta\\varphi\_\{\\mathrm\{b\}\}\}, the barrier turns the equilibrium construction into the lifetime map of Fig\.[5](https://arxiv.org/html/2609.10976#A3.F5): the contours ofΔ​φb\\Delta\\varphi\_\{\\mathrm\{b\}\}fan out from the spinodal, on which the barrier vanishes, and grow toward smallα\\alphaand low temperature, so that retrieval is protected by an order\-one rate barrier already at moderate depth inside the metastable region\. The insets resolve the competition between the two ladders of the endpoint closure at a fixed reference point\. Forλ=0\.5\\lambda=0\.5the co\-defection ladder has both the lower barrier \(0\.190\.19against0\.340\.34at the starred point\) and a terminus, the frozen valueφF\\varphi\_\{\\mathrm\{F\}\}, far above the retrieval valueβ/2\\beta/2, while the distinct ladder terminates at a paramagnetic value belowβ/2\\beta/2: escape toward P is closed there and opens only at largerα\\alpha, whereφP\\varphi\_\{\\mathrm\{P\}\}exceedsφR\\varphi\_\{\\mathrm\{R\}\}, so distinct defections dominate the escape at large load and co\-defections at small load and low temperature\. Forλ=1\.5\\lambda=1\.5the inset shows the distinct ladder running into the divergence of the determinant−12​log⁡Ak\-\\frac\{1\}\{2\}\\log A\_\{k\}asAk→0A\_\{k\}\\to 0, so the escape proceeds through the co\-defection ladder alone\. The general statement is given below\. The equilibrium boundaries follow by comparing themm\-optimized branches\. Equating Eq\. \([208](https://arxiv.org/html/2609.10976#A3.E208)\) and Eq\. \([211](https://arxiv.org/html/2609.10976#A3.E211)\) atm=0m=0gives the P–F line

αλ\+T2​\|log⁡\(1−λ\)\|=1\+εmax​\(α\)2,\\frac\{\\alpha\}\{\\lambda\}\+\\frac\{T\}\{2\}\\big\|\\log\(1\-\\lambda\)\\big\|=\\frac\{1\+\\varepsilon\_\{\\max\}\(\\alpha\)\}\{2\},\(216\)whoseT→0T\\to 0limit is the zero\-temperature equilibrium boundary of Appendix[C\.1](https://arxiv.org/html/2609.10976#A3.SS1)\. Equating Eq\. \([208](https://arxiv.org/html/2609.10976#A3.E208)\) and Eq\. \([210](https://arxiv.org/html/2609.10976#A3.E210)\) gives the P–C line

\(s−1\)​α=12​log⁡1−λ1−β,\(s\-1\)\\,\\alpha=\\frac\{1\}\{2\}\\log\\frac\{1\-\\lambda\}\{1\-\\beta\},\(217\)both first order, while C–F is the continuous freezing line Eq\. \([66](https://arxiv.org/html/2609.10976#S4.E66)\)\. The status of retrieval is unchanged from the main text:φF−φR​\(1\)=β2​εmax\>0\\varphi\_\{\\mathrm\{F\}\}\-\\varphi\_\{\\mathrm\{R\}\}\(1\)=\\frac\{\\beta\}\{2\}\\varepsilon\_\{\\max\}\>0at everyα\>0\\alpha\>0and every temperature, so typical retrieval never becomes the equilibrium phase for Gaussian patterns\. Forλ≥1\\lambda\\geq 1the transverse stabilityAk\>0A\_\{k\}\>0fails atk=1/\(λ2​T\)≤sk=1/\(\\lambda^\{2\}T\)\\leq s: the distinct\-defection ladder terminates before reaching its P terminus, which is the ladder\-level manifestation of the absence of the paramagnetic phase\. The equilibrium is then C or F alone, with retrieval metastable below the ceiling Eq\. \([69](https://arxiv.org/html/2609.10976#S4.E69)\)\. This is the structure displayed in Fig\.[4](https://arxiv.org/html/2609.10976#S4.F4)and, with the barrier contours and escape ladders overlaid, in Fig\.[5](https://arxiv.org/html/2609.10976#A3.F5), now established at all real temperatures: relative to the integer\-lattice derivation, every curve is promoted from an interpolation to an exact solution, the restriction that the lattice caps the accessible temperatures atT=1/λT=1/\\lambdadisappears \(so the C region stands for everyλ\\lambda\), and the barrier Eq\. \([215](https://arxiv.org/html/2609.10976#A3.E215)\) adds dynamical information not visible in the equilibrium diagram\.

Remarks and open problems\.We close with the points left open by the present analysis\. \(i\)*Hidden\-sector corrections\.*The shifts of Eq\. \([164](https://arxiv.org/html/2609.10976#A3.E164)\), the exponentβ/λ→β/λ\+Nh/2\\beta/\\lambda\\to\\beta/\\lambda\+N\_\{h\}/2with its constraintn¯\>Nh/2\\bar\{n\}\>N\_\{h\}/2and the pattern shift−12∑νξ\(h,v\)ν\-\\frac\{1\}\{2\}\\sum\_\{\\nu\}\\xi^\{\(h,v\)\}\_\{\\nu\}, are not small at exponential load\. Whether they can be absorbed as a change of the reference measure within the sector decomposition, leaving the rate\-level phase diagram intact, is the main structural question left open\. \(ii\)*The basin edge\.*Atu→1u\\to 1the binomial coefficients alternate in sign and the series requires resummation\. A uniform treatment of the basin edge, e\.g\. through the Mellin–Barnes representation of\(1\+u\)s\(1\+u\)^\{s\}\[[60](https://arxiv.org/html/2609.10976#bib.bib60)\], would give the precise structure on the spinodal line itself, though the location of the spinodal, fixed by thek=0k=0versusk=1k=1comparison, is not affected\. \(iii\)*Subexponential corrections\.*The prefactors of\(sk\)\\binom\{s\}\{k\}, the fluctuation determinants collected here only at rate level, and the resulting finite\-size shifts ofαc\\alpha\_\{c\}, of orderlog⁡Nv/Nv\\log N\_\{v\}/N\_\{v\}, are relevant to numerical verification of theTTslopes in Eq\. \([68](https://arxiv.org/html/2609.10976#S4.E68)\) and Eq\. \([69](https://arxiv.org/html/2609.10976#S4.E69)\)\. \(iv\)*Stability beyond one step\.*Within each sector the freezing prescription is exact \(Appendix[C\.2](https://arxiv.org/html/2609.10976#A3.SS2)\), but fluctuations between sectors \(an analogue of the AT condition\) and the predicted Ruelle statistics of the F\-phase attention weights\[[36](https://arxiv.org/html/2609.10976#bib.bib36)\]remain unexamined\. \(v\)*Ensemble dependence\.*For spherical patternsεmax≡0\\varepsilon\_\{\\max\}\\equiv 0and the F phase loses its advantage, so typical retrieval can equilibrate, while for binary patterns the Gaussian overlap ratex2/2x^\{2\}/2is replaced by a binary entropy\. Redoing the construction for these ensembles and matching\[[15](https://arxiv.org/html/2609.10976#bib.bib15),[11](https://arxiv.org/html/2609.10976#bib.bib11)\]quantitatively would separate what is universal in the phase diagram from what is a norm\-fluctuation effect\. \(vi\)*Dynamics\.*The barrier Eq\. \([215](https://arxiv.org/html/2609.10976#A3.E215)\) and the two escape ladders are equilibrium statements about a reaction coordinate\. Connecting them to genuine relaxation dynamics, nucleation prefactors, and the multi\-step retrieval dynamics of attention networks is left for future work\.

## References

- \[1\]J\. J\. Hopfield, Neural networks and physical systems with emergent collective computational abilities\., Proceedings of the National Academy of Sciences79, 2554 \(1982\)\.
- \[2\]D\. J\. Amit, H\. Gutfreund, and H\. Sompolinsky, Spin\-glass models of neural networks, Physical Review A32, 1007 \(1985a\)\.
- \[3\]D\. J\. Amit, H\. Gutfreund, and H\. Sompolinsky, Storing infinite numbers of patterns in a spin\-glass model of neural networks, Physical Review Letters55, 1530 \(1985b\)\.
- \[4\]D\. J\. Amit, H\. Gutfreund, and H\. Sompolinsky, Statistical mechanics of neural networks near saturation, Annals of Physics173, 30 \(1987\)\.
- \[5\]D\. Krotov, A new frontier for hopfield networks, Nature Reviews Physics5, 366 \(2023\)\.
- \[6\]D\. Krotov, B\. Hoover, P\. Ram, and B\. Pham, Modern methods in associative memory, arXiv preprint arXiv:2507\.06211 \(2025\)\.
- \[7\]E\. Gardner, Multiconnected neural network models, Journal of Physics A: Mathematical and General20, 3453 \(1987\)\.
- \[8\]L\. F\. Abbott and Y\. Arian, Storage capacity of generalized networks, Physical Review A36, 5091 \(1987\)\.
- \[9\]P\. Baldi and S\. S\. Venkatesh, Number of stable points for spin\-glasses and neural networks of higher orders, Physical Review Letters58, 913 \(1987\)\.
- \[10\]D\. Krotov and J\. J\. Hopfield, Dense associative memory for pattern recognition, in*Advances in Neural Information Processing Systems*, Vol\. 29 \(2016\)\.
- \[11\]M\. Demircigil, J\. Heusel, M\. Löwe, S\. Upgang, and F\. Vermet, On a model of associative memory with huge storage capacity, Journal of Statistical Physics168, 288 \(2017\)\.
- \[12\]A\. Vaswani, N\. Shazeer, N\. Parmar, J\. Uszkoreit, L\. Jones, A\. N\. Gomez, L\. u\. Kaiser, and I\. Polosukhin, Attention is all you need, in*Advances in Neural Information Processing Systems*, Vol\. 30 \(2017\)\.
- \[13\]H\. Ramsauer, B\. Schäfl, J\. Lehner, P\. Seidl, M\. Widrich, L\. Gruber, M\. Holzleitner, T\. Adler, D\. Kreil, M\. K\. Kopp, G\. Klambauer, J\. Brandstetter, and S\. Hochreiter, Hopfield networks is all you need, in*International Conference on Learning Representations*\(2021\)\.
- \[14\]T\. Ota and R\. Karakida, Attention in a family of Boltzmann machines emerging from modern Hopfield networks, Neural Computation35, 1463 \(2023\)\.
- \[15\]C\. Lucibello and M\. Mézard, Exponential capacity of dense associative memories, Physical Review Letters132, 077301 \(2024\)\.
- \[16\]F\. Koulischer, C\. Goemaere, T\. van der Meersch, J\. Deleu, and T\. Demeester, Exploring the temperature\-dependent phase transition in modern Hopfield networks, arXiv preprint arXiv:2311\.18434 \(2023\)\.
- \[17\]J\. Y\.\-C\. Hu, D\. Wu, and H\. Liu, Provably optimal memory capacity for modern Hopfield models: Transformer\-compatible dense associative memories as spherical codes, in*Advances in Neural Information Processing Systems*, Vol\. 37 \(2024\)\.
- \[18\]T\. Petrova, E\. Polyachenko, and R\. State, Geometric entropy and retrieval phase transitions in continuous thermal dense associative memory, in*International Conference on Machine Learning*\(PMLR, 2026\)\.
- \[19\]D\. Krotov and J\. J\. Hopfield, Large associative memory problem in neurobiology and machine learning, in*International Conference on Learning Representations*\(2021\)\.
- \[20\]D\. Bollé, T\. M\. Nieuwenhuizen, I\. Pérez Castillo, and T\. Verbeiren, A spherical Hopfield model, Journal of Physics A: Mathematical and General36, 10269 \(2003\)\.
- \[21\]B\. Derrida, Random\-energy model: An exactly solvable model of disordered systems, Physical Review B24, 2613 \(1981\)\.
- \[22\]R\. Monasson, Structural glass transition and the entropy of the metastable states, Physical Review Letters75, 2847 \(1995\)\.
- \[23\]J\. R\. L\. de Almeida and D\. J\. Thouless, Stability of the Sherrington\-Kirkpatrick solution of a spin glass model, Journal of Physics A: Mathematical and General11, 983 \(1978\)\.
- \[24\]A\. Crisanti, D\. J\. Amit, and H\. Gutfreund, Saturation level of the Hopfield model for neural network, Europhysics Letters2, 337 \(1986\)\.
- \[25\]A\. Crisanti and H\.\-J\. Sommers, The sphericalpp\-spin interaction spin glass model: the statics, Zeitschrift für Physik B Condensed Matter87, 341 \(1992\)\.
- \[26\]B\. Achilli, L\. Ambrogioni, C\. Lucibello, M\. Mezard, and E\. Ventura, The capacity of modern Hopfield networks under the data manifold hypothesis, in*Proceedings of ICLR 2025 workshop: New Frontiers in Associative Memories*\(2025\)\.
- \[27\]M\. Mezard and A\. Montanari,*Information, physics, and computation*\(Oxford University Press, 2009\)\.
- \[28\]H\. Touchette, The large deviation approach to statistical mechanics, Physics Reports478, 1 \(2009\)\.
- \[29\]S\. Franz and G\. Parisi, Recipes for metastable states in spin glasses, Journal de Physique I5, 1401 \(1995\)\.
- \[30\]O\. Penrose and J\. L\. Lebowitz, Rigorous treatment of metastable states in the van der Waals\-Maxwell theory, Journal of Statistical Physics3, 211 \(1971\)\.
- \[31\]A\. Barra, A\. Bernacchia, E\. Santucci, and P\. Contucci, On the equivalence of Hopfield networks and Boltzmann machines, Neural Networks34, 1 \(2012\)\.
- \[32\]A\. Barra, G\. Genovese, P\. Sollich, and D\. Tantari, Phase diagram of restricted Boltzmann machines and generalized Hopfield networks with arbitrary priors, Physical Review E97, 022310 \(2018\)\.
- \[33\]L\. Albanese, F\. Alemanno, A\. Alessandrelli, and A\. Barra, Replica symmetry breaking in dense Hebbian neural networks, Journal of Statistical Physics189, 24 \(2022\)\.
- \[34\]L\. Albanese, A\. Alessandrelli, A\. Annibale, and A\. Barra, Replica symmetry breaking in supervised and unsupervised Hebbian networks, Journal of Physics A: Mathematical and Theoretical57, 165003 \(2024\)\.
- \[35\]G\. S\. Hartnett, E\. Parker, and E\. Geist, Replica symmetry breaking in bipartite spin glasses and neural networks, Physical Review E98, 022116 \(2018\)\.
- \[36\]D\. Ruelle, A mathematical reformulation of Derrida’s REM and GREM, Communications in Mathematical Physics108, 225 \(1987\)\.
- \[37\]B\. Hoover, Y\. Liang, B\. Pham, R\. Panda, H\. Strobelt, D\. H\. Chau, M\. Zaki, and D\. Krotov, Energy transformer, in*Advances in Neural Information Processing Systems*, Vol\. 36 \(2023\)\.
- \[38\]D\. Krotov, Hierarchical associative memory, arXiv preprint arXiv:2107\.06446 \(2021\)\.
- \[39\]R\. Karakida, T\. Ota, and M\. Taki, Hierarchical associative memory, parallelized MLP\-Mixer, and symmetry breaking, arXiv preprint arXiv:2406\.12220 \(2024\)\.
- \[40\]M\. Shafiei Kafraj, D\. Krotov, and P\. E\. Latham, A biologically plausible dense associative memory with exponential capacity, arXiv preprint arXiv:2601\.00984 \(2026\)\.
- \[41\]E\. Agliari, A\. Alessandrelli, A\. Barra, M\. S\. Centonze, and F\. Ricci\-Tersenghi, Networks of neural networks: more is different, Neural Networks194, 108181 \(2026\)\.
- \[42\]B\. Hoover, D\. H\. Chau, H\. Strobelt, and D\. Krotov, A universal abstraction for hierarchical hopfield networks, in*The Symbiosis of Deep Learning and Differential Equations II*\(2022\)\.
- \[43\]R\. Courant and F\. John,*Introduction to Calculus and Analysis: Volume II*\(Springer, 1989\)\.
- \[44\]R\. Schneider,*Convex Bodies: The Brunn–Minkowski Theory*, 2nd ed\., Encyclopedia of Mathematics and its Applications, Vol\. 151 \(Cambridge University Press, 2014\)\.
- \[45\]R\. T\. Rockafellar,*Convex Analysis*, Princeton Mathematical Series, Vol\. 28 \(Princeton University Press, 1970\)\.
- \[46\]F\. Barthe, O\. Guédon, S\. Mendelson, and A\. Naor, A probabilistic approach to the geometry of theℓpn\\ell\_\{p\}^\{n\}\-ball, The Annals of Probability33, 480 \(2005\)\.
- \[47\]G\. Schechtman and J\. Zinn, On the volume of the intersection of twoLpnL\_\{p\}^\{n\}balls, Proceedings of the American Mathematical Society110, 217 \(1990\)\.
- \[48\]A\. P\. Prudnikov, Y\. A\. Brychkov, and O\. I\. Marichev,*Integrals and Series, Vol\. 3: More Special Functions*\(Gordon and Breach Science Publishers, 1990\)\.
- \[49\]F\. W\. J\. Olver, D\. W\. Lozier, R\. F\. Boisvert, and C\. W\. Clark, eds\.,*NIST Handbook of Mathematical Functions*\(Cambridge University Press, 2010\)\.
- \[50\]M\. Mezard, G\. Parisi, and M\. A\. Virasoro,*Spin Glass Theory and Beyond: An Introduction to the Replica Method and Its Applications*, Vol\. 9 \(World Scientific Publishing Company, 1987\)\.
- \[51\]H\. Nishimori,*Statistical Physics of Spin Glasses and Information Processing: An Introduction*\(Oxford University Press, 2001\)\.
- \[52\]R\. Price, A useful theorem for nonlinear devices having Gaussian inputs, IRE Transactions on Information Theory4, 69 \(1958\)\.
- \[53\]K\. Mimura, J\. Takeuchi, Y\. Sumikawa, Y\. Kabashima, and A\. C\. C\. Coolen, Dynamical properties of dense associative memory, arXiv preprint arXiv:2506\.00851 \(2025\)\.
- \[54\]Y\. Sumikawa and Y\. Kabashima, Testing the role of diagonal interactions in high\-order Hopfield models via dynamical mean\-field theory, arXiv preprint arXiv:2604\.03115 \(2026\)\.
- \[55\]T\. H\. Berlin and M\. Kac, The spherical model of a ferromagnet, Physical Review86, 821 \(1952\)\.
- \[56\]J\. M\. Kosterlitz, D\. J\. Thouless, and R\. C\. Jones, Spherical model of a spin\-glass, Physical Review Letters36, 1217 \(1976\)\.
- \[57\]A\. Dembo and O\. Zeitouni,*Large Deviations Techniques and Applications*, 2nd ed\., Stochastic Modelling and Applied Probability, Vol\. 38 \(Springer, 1998\)\.
- \[58\]A\. Bovier,*Statistical Mechanics of Disordered Systems: A Mathematical Perspective*, Cambridge Series in Statistical and Probabilistic Mathematics, Vol\. 18 \(Cambridge University Press, 2006\)\.
- \[59\]R\. P\. Boas,*Entire Functions*, Pure and Applied Mathematics, Vol\. 5 \(Academic Press, 1954\)\.
- \[60\]R\. B\. Paris and D\. Kaminski,*Asymptotics and Mellin\-Barnes Integrals*, Encyclopedia of Mathematics and its Applications, Vol\. 85 \(Cambridge University Press, 2001\)\.

Similar Articles

The Spectral Geometry of Thought: Phase Transitions, Instruction Reversal, Token-Level Dynamics, and Perfect Correctness Prediction in How Transformers Reason

arXiv cs.LG

A comprehensive spectral analysis across 11 LLMs revealing that transformers exhibit phase transitions in hidden activation spaces during reasoning versus factual recall, with seven fundamental phenomena including spectral compression, instruction-tuning reversal, and perfect correctness prediction (AUC=1.0) based solely on spectral properties.

RNNs vs Transformers vs SSMs: where should AI memory live for continual learning?

Reddit r/artificial

A technical analysis comparing memory designs in RNNs, Transformers, and SSMs, arguing that the key question is where to store sequence state rather than which architecture is better. Discusses trade-offs between compressed hidden states, growing KV caches, and synaptic-like memory in model connectivity.

Variational Linear Attention: Stable Associative Memory for Long-Context Transformers

arXiv cs.LG

This paper introduces Variational Linear Attention (VLA), a method that stabilizes memory states in linear attention mechanisms for long-context transformers. VLA reframes memory updates as an online regularized least-squares problem, proving bounded state norms and demonstrating significant speedups and improved retrieval accuracy over standard linear attention and DeltaNet.