Representational Capacity: Geometric Limits on Feature Representation in Transformer Language Models
Summary
This paper introduces a quantitative framework for estimating how many near-orthogonal directions a transformer language model's latent space can support, based on the linear representation and superposition hypotheses. The authors define representational capacity as an upper bound on distinguishable features and show it is exponentially sensitive to the allowed deviation from orthogonality, with larger models favoring tighter constraints.
View Cached Full Text
Cached at: 06/03/26, 09:39 AM
# Representational Capacity: Geometric Limits on Feature Representation in Transformer Language Models
Source: [https://arxiv.org/html/2606.02765](https://arxiv.org/html/2606.02765)
###### Abstract
Model dimension \(dmodeld\_\{model\}\) is a fundamental hyperparameter in transformer\-based language models, yet its role in determining the geometric limits of feature representation remains under\-explored\. Grounded in the Linear Representation and Superposition Hypotheses—which together propose that models encode features as near\-orthogonal directions in the latent space—we develop a quantitative framework for estimating how many such directions a model’s latent space can support\. We first establish the embedding matrix as a measurable proxy for the near\-orthogonality constraints operating across the latent space, proposing that the boundary between meaningful token relationships and incidental similarity in the pairwise cosine similarity distribution provides a concrete estimate of the model’s accepted deviationε\\varepsilonfrom perfect orthogonality\. Applying this metric across dozens of open\-source language models reveals two distinct classes: models with highε\\varepsilonwhose embeddings lack near\-orthogonal structure, and models with lowε\\varepsilonthat maintain strict near\-orthogonality constraints\. We then show that the standard Johnson\-Lindenstrauss lemma dramatically underestimates the packing efficiency of trained representations and derive an adjusted capacity formula in which the number of near\-orthogonal directions depends on the ratio of vectors to dimensions \(k/dk/d\) rather than the raw count alone—a single modification that reduces prediction error by two orders of magnitude with no additional free parameters\. Combining these results, we define*representational capacity*as a quantitative upper bound on the number of distinguishable directions available for features and embeddings within a model’s latent space\. The analysis reveals that capacity is exponentially sensitive toε\\varepsilon, and that larger models tend to favor tighter orthogonality constraints over maximizing raw capacity—a pattern compatible with several explanations \(a stability–capacity trade\-off, a ceiling on usable concepts, or confounds with overall model scale\) that we leave open for future work\.
††footnotetext:Code:[https://github\.com/Alex\-Guha/representational\-capacity](https://github.com/Alex-Guha/representational-capacity)## 1Introduction
Model dimension \(dmodeld\_\{model\}\) is one of the hyperparameters that controls a transformer\-based language model’s parameter count, and primarily determines the size of the embedding[G](https://arxiv.org/html/2606.02765#A1.I1.ix1)and latent space[G](https://arxiv.org/html/2606.02765#A1.I1.ix3)within the model111Definitions for terms marked with a superscript ‘G’ can be found in the Glossary \(Appendix[A](https://arxiv.org/html/2606.02765#A1)\)\.\. In practice,dmodeld\_\{model\}is heuristically chosen to scale with the other hyperparameters, generally in powers of 2 to facilitate efficient GPU computation\.
Naively, one might expect feature vectors of lengthdmodeld\_\{model\}to use each basis vector for a distinct feature\. For example, ifdmodel=3d\_\{model\}=3, the vector\[1,0,0\]\[1,0,0\]might represent “Cat”, while\[0,1,0\]\[0,1,0\]represents “Dog”—anddmodeld\_\{model\}would directly bound the number of features the model could work with\. In practice, neural networks have long been understood to use*distributed representations*[G](https://arxiv.org/html/2606.02765#A1.I1.ix11), in which each feature is represented by multiple basis dimensions being active simultaneously, and each basis dimension participates in representing many different features \(Bengioet al\.\([2014](https://arxiv.org/html/2606.02765#bib.bib17)\)\)\. This leads to*polysemanticity*[G](https://arxiv.org/html/2606.02765#A1.I1.ix9), the phenomenon where individual neurons activate for multiple, seemingly unrelated inputs \(Olahet al\.\([2020](https://arxiv.org/html/2606.02765#bib.bib18)\)\)\.
A specific instantiation of distributed representations that has gained significant traction is the*Linear Representation Hypothesis*[G](https://arxiv.org/html/2606.02765#A1.I1.ix5)\(LRH\)\. Following from the introduction of the idea of linguistic regularities in the context of word embeddings \(*i\.e\.*“King−\-Man\+\+Woman≈\\approxQueen”\) \(Mikolovet al\.\([2013](https://arxiv.org/html/2606.02765#bib.bib2)\)\), this idea purports that neural language models broadly tend to represent concepts and features as directions in the latent space\. Recent works have provided strong evidence for the existence of such linear feature directions in trained models\.Cunninghamet al\.\([2023](https://arxiv.org/html/2606.02765#bib.bib12)\)employed Sparse Autoencoders \(SAEs\) to decompose language model activations into sparse, interpretable components, recovering feature directions including “parts of individual names, especially last names” and “legal terms and court case references”\. Building on this,Templetonet al\.\([2024](https://arxiv.org/html/2606.02765#bib.bib13)\)scaled the approach to Claude 3 Sonnet and extracted millions of monosemantic features—specific entities, code syntax, abstract concepts—including a feature direction for “The Golden Gate Bridge”, and demonstrated that amplifying activations along such directions predictably steers model behavior\.Parket al\.\([2024](https://arxiv.org/html/2606.02765#bib.bib14)\)complement these empirical findings with a theoretical analysis showing that causal interventions along specific directions can predictably manipulate model behavior, further solidifying the link between linear directions and conceptual representations\.
Building on Linear Representations, the*Superposition Hypothesis*[G](https://arxiv.org/html/2606.02765#A1.I1.ix6)accounts for polysemanticity by suggesting that neural networks leverage near\-orthogonality[G](https://arxiv.org/html/2606.02765#A1.I1.ix7)to represent more concepts than the number of available dimensions \(Elhageet al\.\([2022](https://arxiv.org/html/2606.02765#bib.bib1)\)\)\. This idea is grounded in the Johnson\-Lindenstrauss \(JL\) lemma[G](https://arxiv.org/html/2606.02765#A1.I1.ix8)\(Johnson and Lindenstrauss \([1984](https://arxiv.org/html/2606.02765#bib.bib4)\)\)\. As detailed in Appendix[B](https://arxiv.org/html/2606.02765#A2), the lemma’s guarantee of distance preservation implies that inner products between unit vectors are also preserved within a small accepted deviationε\\varepsilon, allowing exponentially many near\-orthogonal directions to exist in high\-dimensional space\. According to the Superposition Hypothesis, neural networks exploit this property by representing features as near\-orthogonal directions inℝdmodel\\mathbb\{R\}^\{d\_\{model\}\}, allowing the number of representable concepts to grow exponentially relative todmodeld\_\{model\}\.
Crucially, the SAE\-based studies above not only demonstrate the existence of linear feature representations but also provide direct empirical evidence for superposition in the latent space\.Templetonet al\.\([2024](https://arxiv.org/html/2606.02765#bib.bib13)\)extracted millions of interpretable features from Claude 3 Sonnet, a model whose latent dimension is orders of magnitude smaller than the number of recovered features\. Since these features are represented as directions in a space with far fewer dimensions than features, they must necessarily be arranged near\-orthogonally—the geometric hallmark of superposition\. This establishes that near\-orthogonality in the latent space is not merely a theoretical possibility but an observed property of trained models\.
#### Contributions\.
This paper investigates the geometric constraints that govern how many features a transformer can represent within itsdmodeld\_\{model\}\-dimensional latent space\. If models encode features as near\-orthogonal directions as the Superposition Hypothesis proposes and SAE studies empirically support, then the number of such directions is bounded by geometric properties of the space: specifically, its dimension and the tolerance for deviation from perfect orthogonality\. Motivated by the relationship between tokenization and embedding structure \(discussed in Section[2](https://arxiv.org/html/2606.02765#S2)\), we analyze the similarity distribution of the embedding matrix as a method to estimate the accepted deviationε\\varepsilonfor near\-orthogonality within a model’s latent space\. At initialization, the embedding matrix maps orthogonal one\-hot vectors intoℝdmodel\\mathbb\{R\}^\{d\_\{model\}\}via random weights, and because random vectors in high\-dimensional space are near\-orthogonal with high probability—the geometric property underlying the Johnson\-Lindenstrauss lemma—the initial embeddings inherit this near\-orthogonal structure\. Training modifies but largely preserves it: the trained distributions develop an extended right tail corresponding to lexical relationships \(morphological variants of the same token\) and semantic relationships \(conceptually related tokens\), while the bulk of unrelated token pairs remains tightly clustered near zero similarity\. The boundary between these meaningful relationships and incidental similarity—estimated asμ\+2σ\\mu\+2\\sigmaof the distribution—provides a concrete, if heuristic, threshold forε\\varepsilon\. Applied across dozens of open\-source models, this estimator reveals two distinct classes: high\-ε\\varepsilonmodels that lack near\-orthogonal embedding structure, and low\-ε\\varepsilonmodels that maintain it\. We then show that the standard Johnson\-Lindenstrauss bound dramatically underestimates the packing achieved by trained representations, and derive an empirically adjusted formula in which capacity depends on the ratiok/dk/drather thankkalone—a single modification that reduces prediction error by two orders of magnitude with no additional free parameters\. Combining these, we define*representational capacity*[G](https://arxiv.org/html/2606.02765#A1.I1.ix10)as a quantitative upper bound on the number of distinguishable directions available within the latent space, revealing that the available near\-orthogonal directions constitute a shared resource among embeddings, unembeddings, and features, that capacity is exponentially sensitive toε\\varepsilon, and that larger models tend to favor tighter orthogonality over raw capacity\.
## 2Embeddings
This section establishes embeddings as a measurable proxy for estimating the accepted deviationε\\varepsilonfor near\-orthogonality[G](https://arxiv.org/html/2606.02765#A1.I1.ix7)within a model’s latent space[G](https://arxiv.org/html/2606.02765#A1.I1.ix3)\. We characterize the post\-training similarity distribution of the embedding matrix \(including the lexical and semantic relationships that form its tail\), propose an estimator forε\\varepsilon, and apply it across dozens of models to reveal two distinct classes\.
### 2\.1Tokenization and the Embedding Space
Tokenization maps each ofVVvocabulary tokens to a learneddmodeld\_\{model\}\-dimensional vector via an embedding matrix𝑬∈ℝV×dmodel\\boldsymbol\{E\}\\in\\mathbb\{R\}^\{V\\times d\_\{model\}\}[G](https://arxiv.org/html/2606.02765#A1.I1.ix1), equivalent to multiplying a one\-hot vector𝒆i∈ℝV\\boldsymbol\{e\}\_\{i\}\\in\\mathbb\{R\}^\{V\}by𝑬\\boldsymbol\{E\}to obtain𝒙i=𝑬⊤𝒆i\\boldsymbol\{x\}\_\{i\}=\\boldsymbol\{E\}^\{\\top\}\\boldsymbol\{e\}\_\{i\}\. By construction, the input one\-hot vectors are perfectly orthogonal:⟨𝒆i,𝒆j⟩=0\\langle\\boldsymbol\{e\}\_\{i\},\\boldsymbol\{e\}\_\{j\}\\rangle=0for alli≠ji\\neq j\. Perfect orthogonality, however, requires a space with dimensionality at least equal to the number of vectors—VVdimensions forVVtokens—and sincedmodel≪Vd\_\{model\}\\ll Vin practice \(typicalVVis 30,000–130,000 whiledmodeld\_\{model\}ranges from 768–8,192\),𝑬\\boldsymbol\{E\}necessarily projects these one\-hot vectors into a much smaller space where strict orthogonality is no longer achievable\. At random initialization the resulting embeddings are near\-orthogonal with high probability via the Johnson\-Lindenstrauss lemma[G](https://arxiv.org/html/2606.02765#A1.I1.ix8), so the embedding matrix can be understood as a compressed representation of vocabulary space; as we will see, training largely preserves this structure\.
These embeddings reside inℝdmodel\\mathbb\{R\}^\{d\_\{model\}\}, the same space occupied by all subsequent latent representations: the residual stream means embeddings, intermediate latents[G](https://arxiv.org/html/2606.02765#A1.I1.ix2), and features all coexist inℝdmodel\\mathbb\{R\}^\{d\_\{model\}\}and are subject to the same geometric constraints\. We therefore hypothesize that the near\-orthogonality observed in embeddings reflects properties of the broader latent space, with one qualification: as discussed in Appendix[C](https://arxiv.org/html/2606.02765#A3), models with tied embedding and unembedding matrices exhibit different structural properties, suggesting their embeddings may not be representative\.
### 2\.2Post\-Training Embedding Structure
For each model, we compute the pairwise cosine similaritysim\(i,j\)=⟨𝒙i,𝒙j⟩/\(‖𝒙i‖‖𝒙j‖\)\\text\{sim\}\(i,j\)=\\langle\\boldsymbol\{x\}\_\{i\},\\boldsymbol\{x\}\_\{j\}\\rangle/\(\\\|\\boldsymbol\{x\}\_\{i\}\\\|\\\|\\boldsymbol\{x\}\_\{j\}\\\|\)between all token embeddings\. Across many trained models, the resulting distributions are tightly centered near zero \(Figure[1](https://arxiv.org/html/2606.02765#S2.F1)\), closely resembling their random initialization; near\-orthogonality is preserved through training\.
Figure 1:Distribution of pairwise cosine similarity between token embeddings across various models, demonstrating near\-orthogonality\.This preservation is not coincidental: while next\-token prediction does not explicitly incentivize near\-orthogonality, it also does not require embeddings to collapse onto each other\. We hypothesize that the model preserves near\-orthogonality as a means of maintaining a structured representational space: if embeddings*lean toward*[G](https://arxiv.org/html/2606.02765#A1.I1.ix12)hidden feature directions, then maintaining near\-orthogonality helps avoid interference between features \(as suggested by the Linear Representation Hypothesis\)\.
Two systematic deviations from random initialization are nonetheless visible on closer inspection \(Figure[2](https://arxiv.org/html/2606.02765#S2.F2)\): the distributions are shifted slightly to the right of zero \(an aversion to negative similarities, plausibly because negative pairwise similarities make the QKV projections work harder to produce useful query\-key interactions\), and they exhibit slightly asymmetric tails—the right tail extends further than the left, reflecting meaningful relationships between certain token pairs, which we examine next\.
Figure 2:Zoomed view of embedding similarity distributions, revealing the right shift from zero and slightly asymmetric tails\.
### 2\.3Lexical and Semantic Token Relationships
The extended right tail in the similarity distributions corresponds to token pairs with genuine lexical or semantic relationships\. Distinguishing these from incidental similarity is essential: meaningful similarity should not count against orthogonality, while incidental similarity represents the model’s tolerance for feature interference\.
\(a\)Lexical relationships
\(b\)Semantic relationships
Figure 3:Examples of lexical and semantic relationships in token embeddings\. Lexical relationships \(a\) show tokens with shared surface forms; semantic relationships \(b\) show conceptually related tokens without lexical overlap\. The final pair \(quick–un\) serves as an unrelated baseline\.Prominent lexical relationships can be identified by examining nearest\-neighbor structure\. For each tokenii, we compute its nearest neighbornn\(i\)=argmaxj≠isim\(i,j\)\\text\{nn\}\(i\)=\\operatorname\*\{argmax\}\_\{j\\neq i\}\\text\{sim\}\(i,j\), and call a tokenkka*primary*token if it appears frequently as the nearest neighbor of others:\|\{i:nn\(i\)=k\}\|\>τ\|\\\{i:\\text\{nn\}\(i\)=k\\\}\|\>\\taufor some thresholdτ\\tau\. These primary tokens \(Figure[3](https://arxiv.org/html/2606.02765#S2.F3)a\) typically have upwards of 10 secondary tokens closer to them than to any other embedding, and they anchor clusters of capitalization or morphological variants \(*e\.g\.*“cat”/“Cat”/“cats”\) that account for much of the right tail\. Semantic relationships \(Figure[3](https://arxiv.org/html/2606.02765#S2.F3)b\) are more subtle: rather than vectors leaning toward each other directly, conceptually related tokens \(*e\.g\.*“cat” and “dog”\) appear to lean independently toward shared feature directions[G](https://arxiv.org/html/2606.02765#A1.I1.ix4)\(*e\.g\.*“animal”\), yielding more limited direct similarity\. This distinction remains speculative\.
Figure 4:The top 40 most similar embeddings to the vector𝒙king−𝒙man\+𝒙woman\\boldsymbol\{x\}\_\{\\text\{king\}\}\-\\boldsymbol\{x\}\_\{\\text\{man\}\}\+\\boldsymbol\{x\}\_\{\\text\{woman\}\}, excluding tokens containing “king”, “woman”, or “women”\. The presence of “queen” and “girl” demonstrates semantic structure encoded as directions in the embedding space\.The classic analogy “King−\-Man\+\+Woman≈\\approxQueen” \(Mikolovet al\.\([2013](https://arxiv.org/html/2606.02765#bib.bib2)\)\) provides further evidence for this semantic structure\. Constructing a query𝒒=𝒙king−𝒙man\+𝒙woman\\boldsymbol\{q\}=\\boldsymbol\{x\}\_\{\\text\{king\}\}\-\\boldsymbol\{x\}\_\{\\text\{man\}\}\+\\boldsymbol\{x\}\_\{\\text\{woman\}\}and retrieving its nearest neighbors among the embedding matrix surfaces both “queen” and “girl” within the top 40 \(Figure[4](https://arxiv.org/html/2606.02765#S2.F4)\)\. That “queen” does not rank first likely reflects the imprecision of vector arithmetic as a navigation tool through semantic feature directions; the mere presence of these semantically appropriate tokens corroborates the LRH for decoder\-based models—semantic relationships are encoded as directions that can be meaningfully combined\.
### 2\.4Estimating and Generalizing Accepted Deviationε\\varepsilon
Token pairs beyond the boundary of meaningful similarity are near\-orthogonal largely as a consequence of high\-dimensional geometry, and the boundary itself provides a natural estimate for the model’s tolerance of incidental similarity between unrelated directions\. We propose estimatingε\\varepsilonas approximatelyμ\+2σ\\mu\+2\\sigmaof the pairwise similarity distribution, whereμ\\muandσ\\sigmaare its mean and standard deviation\. This threshold is a motivated heuristic rather than a formally justified bound, and is intended to mark the transition from the bulk of unrelated similarities to the relationship\-driven tail\.
This estimator is grounded in embedding geometry, and extending it to features warrants care: the embedding matrix is a fixed set of learned vectors, while features are directions that may be activated to varying degrees during inference\. The connection between the two is nonetheless stronger than mere shared dimensionality\. The weight matrices in attention and MLP layers—which perform the only direct interaction with latent vectors outside of normalization—function as collections of feature probes\. Each weight vector𝒘i\\boldsymbol\{w\}\_\{i\}extracts information from a latent𝒉\\boldsymbol\{h\}via the inner product
𝒉⋅𝒘i=‖𝒉‖‖𝒘i‖cosθ,\\boldsymbol\{h\}\\cdot\\boldsymbol\{w\}\_\{i\}=\\\|\\boldsymbol\{h\}\\\|\\\|\\boldsymbol\{w\}\_\{i\}\\\|\\cos\\theta,\(1\)which is fundamentally a scaled similarity measurement on the sameℝdmodel\\mathbb\{R\}^\{d\_\{model\}\}space, meaning weight vectors must learn directions that correspond to the features they probe\. Individual weight vectors need not align perfectly with any single feature direction—they may represent linear combinations of features, making them difficult to interpret in isolation even when the underlying feature space is well\-structured\. The outputs of these transformations are then combined with the residual stream by addition, preserving the underlying geometric structure\. Nonlinear activations applied after linear transformations can be understood as response functions operating on these similarity scores rather than as modifications to the underlying geometry: directions must first be linearly distinguishable via dot product before nonlinear transformations can selectively act on them\. This framing does not preclude features that emerge from multi\-layer nonlinear composition, but it does establish that each individual layer’s interaction with the latent space is geometrically constrained by the same near\-orthogonality measurement that governs embeddings\. Embeddings additionally offer a practical advantage as a fixed, measurable quantity, unlike intermediate representations which vary with input\. This generalization nonetheless remains an empirical hypothesis: layer normalization and the specific learned weight configurations could still impose different effective constraints on intermediate representations\.
### 2\.5Two Classes of Models
To validateε\\varepsilonas a meaningful metric, we apply it across dozens of language models and compare againstdmodeld\_\{model\}\. The analysis reveals two distinct classes \(Figure[5](https://arxiv.org/html/2606.02765#S2.F5); per\-model values in Appendix[D](https://arxiv.org/html/2606.02765#A4), TableLABEL:tab:model\_repr\_cap\), suggestingε\\varepsiloncaptures a genuine structural property\.
\(a\)Two classes of models
\(b\)The second class, zoomed in
Figure 5:Model dimension compared to estimatedε\\varepsilonacross various language models, revealing two distinct classes\.The first class consists of models with generally lowerdmodeld\_\{model\}and wide similarity distributions, yieldingε\>0\.2\\varepsilon\>0\.2\.222An extreme outlier is Gemma 7B, whoseε\\varepsilonapproaches 1\. We do not investigate the cause here, but it appears more consistent with idiosyncratic training or architectural factors than with a deliberate design choice\.These models exhibit similarity\-distribution means far from zero \(typically 0\.1–0\.9\), the clearest evidence that they do not leverage superposition at the embedding level: when average pairwise similarity is substantially positive, embeddings cannot serve as distinguishable, near\-orthogonal directions\. Nearly all models in this class have tied embedding and unembedding matrices \(see Appendix[C](https://arxiv.org/html/2606.02765#A3)\)\.
The second class consists of models with generally higherdmodeld\_\{model\}and tight distributions clustered around zero \(ε<0\.1\\varepsilon<0\.1, most around 0\.09 or less\), consistent with active use of superposition[G](https://arxiv.org/html/2606.02765#A1.I1.ix6)to pack many distinguishable directions into the latent space\. The cause of the division is not entirely clear: tying correlates with the first class but does not fully explain it, since a few tied models appear in the second class and some untied ones in the first\. The remainder of the paper focuses on the second class, whereε\\varepsilon\-based capacity bounds are meaningful\.
### 2\.6Section Summary
This section established embeddings as a measurable proxy for the near\-orthogonality constraints operating in a model’s latent space\. At initialization, embedding matrices map orthogonal one\-hot vectors into a near\-orthogonal configuration inℝdmodel\\mathbb\{R\}^\{d\_\{model\}\}, since random vectors in high\-dimensional space are approximately orthogonal with high probability\. Training modifies this initial structure only modestly: the distributions shift right \(avoiding negative similarities\) and develop an extended right tail \(encoding lexical and semantic relationships\), but the fundamental near\-orthogonality persists\. The boundary between the main distribution and the relationship\-driven tail—estimated as approximatelyμ\+2σ\\mu\+2\\sigma—provides a concrete threshold for the accepted deviationε\\varepsilon\. Under the assumption that embeddings and features coexist in the samedmodeld\_\{model\}\-dimensional space and are subject to similar geometric constraints, this tolerance plausibly extends to the feature directions that the model learns to use, though this remains a hypothesis rather than a proven relationship\.
Applying this metric across dozens of models reveals two distinct classes: high\-ε\\varepsilonmodels that do not appear to use superposition at the embedding level, and low\-ε\\varepsilonmodels that maintain tight near\-orthogonality\. Embeddings are therefore not merely a lookup table for token representations\. They are a compressed representation of vocabulary space that*leans toward*[G](https://arxiv.org/html/2606.02765#A1.I1.ix12)hidden feature directions, placing tokens in superposition[G](https://arxiv.org/html/2606.02765#A1.I1.ix6)alongside whatever features the model has learned\. Both embeddings and features draw from the same pool of near\-orthogonal directions inℝdmodel\\mathbb\{R\}^\{d\_\{model\}\}, subject to the same geometric constraints quantified byε\\varepsilon\. The following section uses thisε\\varepsilonestimate to derive a quantitative bound on representational capacity\.
## 3Representational Capacity
We now convertε\\varepsiloninto a quantitative bound on*representational capacity*[G](https://arxiv.org/html/2606.02765#A1.I1.ix10)—an upper bound on the number of near\-orthogonal directions available for features, embeddings, and other learned representations within a model’s latent space\. The Johnson\-Lindenstrauss \(JL\) lemma[G](https://arxiv.org/html/2606.02765#A1.I1.ix8)provides a natural starting point, but as we show, its random\-projection assumption dramatically underestimates the packing achieved by trained models\. We derive an empirically adjusted relationship that better matches the geometry of optimized representations and use it to define and compute representational capacity\.
### 3\.1Near\-Orthogonal Directions as a Shared Resource
Throughout this section we usekkfor the number of near\-orthogonal directions andd=dmodeld=d\_\{model\}for the latent dimension, consistent with the JL formulation in Appendix[B](https://arxiv.org/html/2606.02765#A2)\.
Near\-orthogonal directions inℝdmodel\\mathbb\{R\}^\{d\_\{model\}\}serve multiple purposes within a transformer model\. Features—the conceptual units the model has learned to represent—occupy directions in this space, as described by the Linear Representation Hypothesis[G](https://arxiv.org/html/2606.02765#A1.I1.ix5)\. Token embeddings and unembeddings each also require directions: the embedding matrix mapsVVtokens intoℝdmodel\\mathbb\{R\}^\{d\_\{model\}\}, and the unembedding matrix \(when untied\) maps latent representations back to logits over the vocabulary\. For models with separate embedding and unembedding matrices, a rough estimate of the total near\-orthogonal directions required is
k≈2V\+kfeatures,k\\approx 2V\+k\_\{\\text\{features\}\},\(2\)whereVVis the vocabulary size andkfeaturesk\_\{\\text\{features\}\}is the number of feature directions\. For tied models embeddings and unembeddings share directions, reducing this to approximatelyk≈V\+kfeaturesk\\approx V\+k\_\{\\text\{features\}\}\. As noted in Section[2](https://arxiv.org/html/2606.02765#S2), however, most tied models fall into the high\-ε\\varepsilonclass and may not leverage near\-orthogonality at the embedding level; the few tied models in the low\-ε\\varepsilonclass \(*e\.g\.*Gemma\) should be considered with this adjustment\.
While we cannot directly measurekfeaturesk\_\{\\text\{features\}\}, we can work in reverse: givendmodeld\_\{model\}andε\\varepsilon, estimate the totalkkavailable, and subtract the vocabulary contribution to bound the remaining capacity for features\.
### 3\.2The Johnson\-Lindenstrauss Framework and Its Limits
The JL lemma \(Appendix[B](https://arxiv.org/html/2606.02765#A2)\) guarantees that forkkunit vectors there exists a linear map toℝd\\mathbb\{R\}^\{d\}preserving inner products within±ε\\pm\\varepsilon, provided
d≥C⋅lnkε2,equivalentlyk≤exp\(d⋅ε2C\),d\\geq C\\cdot\\frac\{\\ln k\}\{\\varepsilon^\{2\}\},\\qquad\\text\{equivalently\}\\qquad k\\leq\\exp\\\!\\left\(\\frac\{d\\cdot\\varepsilon^\{2\}\}\{C\}\\right\),\(3\)whereCCdepends on the construction\. This implies exponential growth ofkkwithddand quadratic growth withε\\varepsilon—the geometric basis for the Superposition Hypothesis[G](https://arxiv.org/html/2606.02765#A1.I1.ix6)\.
#### The problem with random vectors\.
Applied to trained embeddings, however, this bound is wildly off\. Using the best provenC=8C=8for Llama 2 7B \(dmodel=4096d\_\{model\}=4096,ε=0\.0645\\varepsilon=0\.0645\) yields
k≤exp\(4096×0\.064528\)=exp\(2\.130\)≈8\.4k\\leq\\exp\\\!\\left\(\\frac\{4096\\times 0\.0645^\{2\}\}\{8\}\\right\)=\\exp\(2\.130\)\\approx 8\.4\(4\)near\-orthogonal directions, against an actual vocabulary of 32,000\. To check whetherC=8C=8is simply too conservative, we empirically tighten it\. For each\(k,d\)\(k,d\)we generateT=1000T=1000independent trials ofkkrandom unit vectors and record the best achievable worst\-case similarity:
εrandom∗\(k,d\)=mint∈\{1,…,T\}maxi≠j\|sim\(𝒗i\(t\),𝒗j\(t\)\)\|\.\\varepsilon^\{\*\}\_\{\\text\{random\}\}\(k,d\)=\\min\_\{t\\in\\\{1,\\ldots,T\\\}\}\\max\_\{i\\neq j\}\\lvert\\text\{sim\}\(\\boldsymbol\{v\}\_\{i\}^\{\(t\)\},\\boldsymbol\{v\}\_\{j\}^\{\(t\)\}\)\\rvert\.\(5\)Fitting \([3](https://arxiv.org/html/2606.02765#S3.E3)\) to this data yieldsC≈3\.029C\\approx 3\.029with excellent fit \(R2=0\.9985R^\{2\}=0\.9985, MAPE1\.3%1\.3\\%, NRMSE0\.9%0\.9\\%; Figure[6](https://arxiv.org/html/2606.02765#S3.F6)\)\. Even with this tightened constant, Llama 2 7B is bounded at onlyexp\(4096×0\.06452/3\.029\)≈277\\exp\(4096\\times 0\.0645^\{2\}/3\.029\)\\approx 277directions—still nearly two orders of magnitude short of its 32,000\-token vocabulary\. The issue is the assumption rather than the constant: trained embeddings are not random projections, but the result of gradient\-based optimization that has discovered arrangements packing far more near\-orthogonal directions than random chance allows\.
Figure 6:The standard JL relationshipε=C⋅ln\(k\)/d\\varepsilon=\\sqrt\{C\\cdot\\ln\(k\)/d\}fitted to empirically generated random vector data, yieldingC≈3\.029C\\approx 3\.029withR2=0\.9985R^\{2\}=0\.9985\. Even this tightened constant cannot account for the packing achieved by trained embeddings\.
### 3\.3An Adjusted Relationship for Optimized Vectors
To characterize what optimized vectors achieve, we initializekkrandom unit vectors inℝd\\mathbb\{R\}^\{d\}and minimize a high\-exponent penaltyℒ=∑i≠j\|Gij\|p\\mathcal\{L\}=\\sum\_\{i\\neq j\}\|G\_\{ij\}\|^\{p\}on the off\-diagonal of the Gram matrix𝑮=𝑽⊤𝑽\\boldsymbol\{G\}=\\boldsymbol\{V\}^\{\\top\}\\boldsymbol\{V\}\(p∈\[40,60\]p\\in\[40,60\], selected per scale; Adam optimizer \(Kingma and Ba \([2017](https://arxiv.org/html/2606.02765#bib.bib15)\)\); 5,000 steps\)\. This implicitly minimizes the maximum pairwise similarity, mimicking the tight packing that gradient descent produces during training\. We sweepd∈\{32,64,128,256,512,768,1024,1536,2048,2560,3072,3584,4096\}d\\in\\\{32,64,128,256,512,768,1024,1536,2048,2560,3072,3584,4096\\\}andkkfrom 2,000 to 32,000, incrementing by 2,000 up to 8,000 for all dimensions and by 4,000 from 8,000 to 32,000 ford≥1536d\\geq 1536\(the coarser step reflecting the increased cost of optimizing larger configurations\), and recordε∗=maxi≠j\|Gij\|\\varepsilon^\{\*\}=\\max\_\{i\\neq j\}\|G\_\{ij\}\|\.
Refitting only the constantCCin \([3](https://arxiv.org/html/2606.02765#S3.E3)\) to this optimized data gives a poor fit \(MAPE≈975%\\approx 975\\%\): the standard JL surface is too flat to track the empirical curve \(Figure[7](https://arxiv.org/html/2606.02765#S3.F7), left\)\. Among many functional\-form modifications we tested, replacingln\(k\)\\ln\(k\)withln\(k/d\)\\ln\(k/d\)produced by far the best alignment:
ε=C⋅ln\(k/d\)d,equivalentlyk≤d⋅exp\(d⋅ε2C\)\.\\varepsilon=\\sqrt\{\\frac\{C\\cdot\\ln\(k/d\)\}\{d\}\},\\qquad\\text\{equivalently\}\\qquad k\\leq d\\cdot\\exp\\\!\\left\(\\frac\{d\\cdot\\varepsilon^\{2\}\}\{C\}\\right\)\.\(6\)Capacity now depends on the*ratio*of vectors to dimensions, not just the raw count\. With a single free parameter \(C≈1\.293C\\approx 1\.293\), this fit yieldsR2=0\.9984R^\{2\}=0\.9984, MAPE7\.9%7\.9\\%—a123×123\\timesMAPE reduction over standard JL on the same data, with no extra parameters\. The factor ofddmultiplying the exponential is what allows optimized packing to dwarf what random projections can achieve\.
Figure 7:Standard JL formula \(left\) vs\. the adjusted relationship \(right\) fitted to optimized vector data\. Both have one free parameter; the adjusted form fits dramatically better\. Red points: empirical data\.A fully parameterized variantε=C⋅ln\(ka/d\)b/dc\\varepsilon=\\sqrt\{C\\cdot\\ln\(k^\{a\}/d\)^\{b\}/d^\{c\}\}confirms the structure: fitting yieldsa≈1\.07a\\approx 1\.07,c≈0\.97c\\approx 0\.97\(validatingk/dk/dand1/d1/dscaling\), with only modest gains in fit quality \(R2=0\.9998R^\{2\}=0\.9998, MAPE5\.4%5\.4\\%; Figure[8](https://arxiv.org/html/2606.02765#S3.F8)\)\. Rearranging forkkgives
k≤\(d⋅exp\(\(dc⋅ε2C\)1/b\)\)1/a\.k\\leq\\left\(d\\cdot\\exp\\\!\\left\(\\left\(\\frac\{d^\{c\}\\cdot\\varepsilon^\{2\}\}\{C\}\\right\)^\{1/b\}\\right\)\\right\)^\{1/a\}\.\(7\)We use this fully parameterized form \(C=0\.458C=0\.458,a=1\.067a=1\.067,b=1\.447b=1\.447,c=0\.972c=0\.972\) for the per\-model estimates in Appendix[D](https://arxiv.org/html/2606.02765#A4), but the single\-parameter \([6](https://arxiv.org/html/2606.02765#S3.E6)\) captures the essential structure\.
Figure 8:The fully parameterized formulaε=C⋅ln\(ka/d\)b/dc\\varepsilon=\\sqrt\{C\\cdot\\ln\(k^\{a\}/d\)^\{b\}/d^\{c\}\}fitted to optimized vector data, yieldingR2=0\.9998R^\{2\}=0\.9998\. The modest improvement over the single\-parameter form \([6](https://arxiv.org/html/2606.02765#S3.E6)\) confirms that thek/dk/dratio is the essential structural insight rather than the extra degrees of freedom\.
### 3\.4Defining Representational Capacity
We define the model’s*representational capacity*as the upper bound onkkgivendmodeld\_\{model\}andε\\varepsilonvia the adjusted relationship above\. For Llama 2 7B \(dmodel=4096d\_\{model\}=4096,ε=0\.0645\\varepsilon=0\.0645\) the fully parameterized form givesk≤3\.99×107k\\leq 3\.99\\times 10^\{7\}—roughly 40 million near\-orthogonal directions, dramatically higher than the JL estimate of∼277\\sim 277, and comfortably accommodating both the 32,000\-token vocabulary \(consuming∼64,000\\sim 64\{,\}000directions across embeddings and unembeddings\) and the millions of features SAE studies have begun to extract\.
Two caveats are essential\. First, capacity bounds geometric possibility, not realized usage: that 40 million directions*can*exist does not mean the model has learned that many features, nor that its computational architecture \(depth, attention heads, MLP width\) could effectively utilize them\. Second, theε\\varepsilonestimate originates from embedding geometry, and its extension to feature geometry remains a hypothesis \(cf\. Section[2](https://arxiv.org/html/2606.02765#S2)\)\. Representational capacity is therefore most useful as a*relative*metric: a model with capacity10810^\{8\}has fundamentally more room for features than one with capacity10610^\{6\}, regardless of how many features either actually uses\.
### 3\.5Capacity Across Model Classes
The two\-class split from Section[2](https://arxiv.org/html/2606.02765#S2)extends to capacity\. For Class 1 models \(ε\>0\.2\\varepsilon\>0\.2\) the framework does not meaningfully apply: their embeddings lack near\-orthogonal structure in the first place, so the formula yields vacuous bounds \(*e\.g\.*k≤2×1020k\\leq 2\\times 10^\{20\}atdmodel=2048,ε=0\.25d\_\{model\}=2048,\\varepsilon=0\.25\)\. It is possible Class 1 models leverage near\-orthogonality at the feature level despite their embeddings showing no such structure, but our embedding\-based analysis cannot resolve this\. We focus the remaining analysis on Class 2\.
For Class 2 models \(ε<0\.1\\varepsilon<0\.1\), capacities range from∼105\\sim 10^\{5\}up to∼1010\\sim 10^\{10\}–101110^\{11\}across the models analyzed \(Appendix[D](https://arxiv.org/html/2606.02765#A4), TableLABEL:tab:model\_repr\_cap\), comfortably accommodating their vocabularies\.333Llama 2 70B is a notable exception, withε=0\.09962\\varepsilon=0\.09962landing it at∼1015\\sim 10^\{15\}; this likely reflects idiosyncratic training or architectural factors\.The highestε\\varepsilonwe observe within Class 2 is 0\.09, with no models approaching the 0\.1–0\.2 gap from below; this clustering suggests that models may not benefit from, or perhaps cannot sustain, deviations much larger than this threshold while maintaining the structure required for effective superposition\. Within this regime,ε\\varepsilonexerts a stronger pull on capacity thandmodeld\_\{model\}: at fixedε=0\.09\\varepsilon=0\.09, increasingdmodeld\_\{model\}from 2048 to 3072 gives a 30×\\timescapacity boost \(2\.0×107→6\.0×1082\.0\\\!\\times\\\!10^\{7\}\\to 6\.0\\\!\\times\\\!10^\{8\}\), but at fixeddmodel=8192d\_\{model\}=8192, raisingε\\varepsilonfrom 0\.05 to 0\.09 gives a six\-orders\-of\-magnitude jump \(2\.4×108→2\.0×10142\.4\\\!\\times\\\!10^\{8\}\\to 2\.0\\\!\\times\\\!10^\{14\}\)\. The exponential dependence onε2\\varepsilon^\{2\}makes overlap tolerance, not dimension, the dominant lever\.
Despite this leverage, observed models do not fully exploit it\. Figure[5](https://arxiv.org/html/2606.02765#S2.F5)b in Section[2](https://arxiv.org/html/2606.02765#S2)shows a negative correlation betweendmodeld\_\{model\}andε\\varepsilonwithin Class 2: larger models maintain tighter near\-orthogonality\.444We do not present a separate plot of capacity againstdmodeld\_\{model\}orε\\varepsilon: since capacity is a deterministic function of the two, such a figure would carry no information beyond Figure[5](https://arxiv.org/html/2606.02765#S2.F5)b\.This implies a trade\-off between raw capacity and internal stability—largerε\\varepsilonadmits more features but more interference, while smallerε\\varepsilonsacrifices capacity for more reliable feature retrieval—and suggests that as models scale, they favor stability over capacity maximization\. The trend rests on limited data—only a handful of analyzed models havedmodel\>6000d\_\{model\}\>6000, making it difficult to assess whether it continues at larger scales—and the causal mechanism remains unclear\. Several explanations are plausible: there may be a practical ceiling on useful capacity beyond which additional feature directions provide diminishing returns; tighter orthogonality may be*necessary*at higherdmodeld\_\{model\}to maintain internal stability regardless of capacity benefits; or tighterε\\varepsilonmay simply correlate with overall model size, which tends to increase alongsidedmodeld\_\{model\}in practice\. Disentangling these factors would require controlled experiments that varydmodeld\_\{model\}while holding other architectural choices constant, an expensive undertaking that remains beyond the scope of this work\.
### 3\.6Summary
The standard JL lemma, designed for random projections, dramatically underestimates the packing efficiency of learned representations: even with an empirically\-tightened constant it bounds Llama 2 7B to fewer than 300 directions, far short of its 32,000\-token vocabulary\. Replacingln\(k\)\\ln\(k\)withln\(k/d\)\\ln\(k/d\)in the JL relationship—a single\-parameter modification motivated by direct optimization of vector arrangements—reduces prediction error by two orders of magnitude \(MAPE975%→8%975\\%\\to 8\\%\) and yields capacity bounds consistent with what is observed in trained models\. Applied across models, representational capacity is best understood as a*relative*metric bounding geometric possibility rather than learned utilization: it reveals that Class 1 models fall outside the framework, that Class 2 capacities span ten orders of magnitude, and that overlap toleranceε\\varepsilonis the dominant lever, while empirically larger models trade raw capacity for tighter near\-orthogonality\.
## 4Conclusion
#### Summary\.
This paper investigated the geometric constraints governing how many features a transformer\-based language model can represent within itsdmodeld\_\{model\}\-dimensional latent space, proceeding in four phases across two chapters\.*In Section[2](https://arxiv.org/html/2606.02765#S2)*, we first established the embedding matrix as a measurable proxy for the near\-orthogonality constraints operating across the latent space, and used the boundary between the bulk of the cosine\-similarity distribution and its right tail—estimated asμ\+2σ\\mu\+2\\sigma—as a concrete threshold for the accepted deviationε\\varepsilon\. Applying this metric across dozens of models then revealed two distinct classes with no intermediate cases: high\-ε\\varepsilonmodels whose embeddings are not approximately orthogonal, and low\-ε\\varepsilonmodels that maintain strict near\-orthogonality\.*In Section[3](https://arxiv.org/html/2606.02765#S3)*, we showed that the standard Johnson–Lindenstrauss framework, designed for random projections, dramatically underestimates the packing efficiency of trained representations—predicting fewer than 300 directions for a model with a 32,000\-token vocabulary\. By optimizing sets of vectors to mimic what gradient descent implicitly achieves, we derived an adjusted relationship in which capacity depends on the ratiok/dk/drather thankkalone, reducing prediction error by two orders of magnitude with no additional free parameters; we then combined this with the per\-modelε\\varepsilonestimates to define and compute*representational capacity*as a quantitative upper bound on the distinguishable directions available in the latent space\. The resulting picture is one in which embeddings, unembeddings, and features draw from a shared,ε\\varepsilon\-bounded pool of near\-orthogonal directions, and in which larger models tend to favor tighter orthogonality constraints over maximizing raw capacity—suggesting that representational stability may matter more than sheer geometric room at scale\.
#### Limitations\.
Several caveats apply\. \(i\) The assumption that embedding geometry reflects the broader latent space is a motivated hypothesis, not a proven correspondence: embeddings are a fixed set of learned vectors, while features are dynamically activated directions, and layer normalization plus learned weights could impose different effective constraints over successive layers\. \(ii\) Theμ\+2σ\\mu\+2\\sigmathreshold forε\\varepsilonis heuristic, and capacity is exponentially sensitive toε2\\varepsilon^\{2\}—a shift fromε=0\.06\\varepsilon=0\.06to0\.090\.09moves estimated capacity by orders of magnitude—making absolute capacity values unreliable even though relative comparisons remain informative\. \(iii\) Architectural details we did not model \(rotary positional encodings on QK subspaces, layer normalization rescaling\) may impose distinct constraints on intermediate representations\. \(iv\) Capacity bounds geometric possibility, not realized utilization: the gap between how many directions*can*exist and how many a given architecture can effectively learn and process remains unquantified\.
#### Future work\.
Several directions follow naturally\. The most direct test of our central assumption would measure near\-orthogonality in intermediate latent representations across inputs and layers, confirming or refuting whether embedding\-derivedε\\varepsilongeneralizes to the residual stream\. Correlating capacity against benchmark performance could reveal whether a “representational scaling law” exists, complementing existing scaling heuristics for choosingdmodeld\_\{model\}\. A complete scaling theory will likely need to account for both representational capacity \(how many features can be stored\) and computational capacity \(how many can be effectively processed by the network’s depth, width, and attention budget\)\. Finally, the bifurcation into high\- and low\-ε\\varepsilonclasses raises open questions: what architectural or training factors cause the divide, is there a critical scale at which models transition into superposition, and do Class 1 models leverage superposition internally despite their embeddings showing no evidence of it?
#### Implications for model design\.
Representational capacity currently serves as a diagnostic property of trained models, sinceε\\varepsilonis not a directly controllable hyperparameter—it emerges from training dynamics and architectural choices such as embedding/unembedding tying\. The clustering of Class 2ε\\varepsilonvalues below 0\.09 and the observed tendency of larger\-dmodeld\_\{model\}models to maintain tighter orthogonality together hint at a possible ceiling on useful capacity, beyond which additional geometric room provides diminishing returns\. If such a ceiling exists,dmodeld\_\{model\}could in principle be chosen so that the resulting capacity atε≲0\.09\\varepsilon\\lesssim 0\.09approximately matches it, avoiding wasted dimensions; conversely, models that have not yet saturated this ceiling may benefit more from increaseddmodeld\_\{model\}than from other forms of scaling\. Either possibility would convert the current heuristic of choosingdmodeld\_\{model\}in powers of two into a principled, capacity\-targeted choice—though distinguishing a true ceiling from a stability requirement at scale will require controlled experiments that varydmodeld\_\{model\}while holding other architectural choices fixed\.
#### Closing remarks\.
We started from a simple question: given a model’s latent dimension, how many features can it represent? The answer depends not just on dimension but on how tightly the model constrains the overlap between its representations—a quantity that emerges from training rather than being set by design\. The framework here reframesdmodeld\_\{model\}as the determinant of a finite geometric resource that embeddings, unembeddings, and features all draw from, with the model’s tolerance for overlap governing how far that resource stretches\. Whether this geometric perspective ultimately connects to model capability—whether capacity predicts what a model can learn, not just what it can store—remains the most compelling open question\.
## References
- M\. Abdin, S\. A\. Jacobs, A\. A\. Awan, J\. Aneja, A\. Awadallah, H\. Awadalla, N\. Bach, A\. Bahree, A\. Bakhtiari, H\. Behl,et al\.\(2024\)Phi\-3 technical report: a highly capable language model locally on your phone\.arXiv preprint arXiv:2404\.14219\.External Links:[Link](https://arxiv.org/abs/2404.14219)Cited by:[Appendix D](https://arxiv.org/html/2606.02765#A4.p2.1)\.
- M\. AI \(2024\)Moonshot ai: kimi intelligent assistant\.Note:[https://www\.moonshot\.cn/](https://www.moonshot.cn/)Accessed: 2026Cited by:[Appendix D](https://arxiv.org/html/2606.02765#A4.p2.1)\.
- Y\. Bengio, A\. Courville, and P\. Vincent \(2014\)Representation learning: a review and new perspectives\.External Links:1206\.5538,[Link](https://arxiv.org/abs/1206.5538)Cited by:[item Distributed Representations](https://arxiv.org/html/2606.02765#A1.I1.ix11.p1.1),[§1](https://arxiv.org/html/2606.02765#S1.p2.5)\.
- H\. Cunningham, A\. Ewart, L\. Riggs, R\. Huben, and L\. Sharkey \(2023\)Sparse autoencoders find highly interpretable features in language models\.External Links:2309\.08600,[Link](https://arxiv.org/abs/2309.08600)Cited by:[§1](https://arxiv.org/html/2606.02765#S1.p3.3)\.
- DeepSeek\-AIet al\.\(2024\)DeepSeek\-v3 technical report\.External Links:2412\.19437,[Link](https://arxiv.org/abs/2412.19437)Cited by:[Appendix D](https://arxiv.org/html/2606.02765#A4.p2.1)\.
- A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Yang, A\. Fan,et al\.\(2024\)The llama 3 herd of models\.arXiv preprint arXiv:2407\.21783\.External Links:[Link](https://arxiv.org/abs/2407.21783)Cited by:[Appendix D](https://arxiv.org/html/2606.02765#A4.p2.1)\.
- N\. Elhage, T\. Hume, C\. Olsson, N\. Schiefer, T\. Henighan, S\. Kravec, Z\. Hatfield\-Dodds, R\. Lasenby, D\. Drain, C\. Chen, R\. Grosse, S\. McCandlish, J\. Kaplan, D\. Amodei, M\. Wattenberg, and C\. Olah \(2022\)Toy models of superposition\.Transformer Circuits Thread\.External Links:[Link](https://transformer-circuits.pub/2022/toy_model/index.html)Cited by:[§1](https://arxiv.org/html/2606.02765#S1.p4.3)\.
- A\. Q\. Jiang, A\. Sablayrolles, A\. Mensch, C\. Bamford, D\. S\. Chaplot, D\. de Las Casas, F\. Bressand, G\. Lengyel, G\. Lample, L\. Saulnier, L\. R\. Lavaud, M\. Lachaux, P\. Stock, T\. L\. Scao, T\. Lavril, T\. Wang, T\. Lacroix, and W\. E\. Sayed \(2023\)Mistral 7b\.arXiv preprint arXiv:2310\.06825\.External Links:[Link](https://arxiv.org/abs/2310.06825)Cited by:[Appendix D](https://arxiv.org/html/2606.02765#A4.p2.1)\.
- W\. B\. Johnson and J\. Lindenstrauss \(1984\)Extensions of Lipschitz mappings into a Hilbert space\.InConference in modern analysis and probability \(New Haven, Conn\., 1982\),R\. Beals, A\. Beck, and A\. Bellow \(Eds\.\),Contemporary Mathematics, Vol\.26,Providence, RI,pp\. 189–206\.External Links:ISBN 0\-8218\-5030\-X,[Document](https://dx.doi.org/10.1090/conm/026/737400),[Link](https://archive.org/details/conferenceinmode0000conf/page/189),[MathReview Entry](https://www.ams.org/mathscinet-getitem?mr=737400)Cited by:[§1](https://arxiv.org/html/2606.02765#S1.p4.3)\.
- D\. P\. Kingma and J\. Ba \(2017\)Adam: a method for stochastic optimization\.External Links:1412\.6980,[Link](https://arxiv.org/abs/1412.6980)Cited by:[§3\.3](https://arxiv.org/html/2606.02765#S3.SS3.p1.9)\.
- T\. Mikolov, W\. Yih, and G\. Zweig \(2013\)Linguistic regularities in continuous space word representations\.InProceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,L\. Vanderwende, H\. Daumé III, and K\. Kirchhoff \(Eds\.\),Atlanta, Georgia,pp\. 746–751\.External Links:[Link](https://aclanthology.org/N13-1090/)Cited by:[§1](https://arxiv.org/html/2606.02765#S1.p3.3),[§2\.3](https://arxiv.org/html/2606.02765#S2.SS3.p3.4)\.
- MiniMax \(2025\)MiniMax m2\.Note:[https://www\.minimaxi\.com/news/minimax\-m2](https://www.minimaxi.com/news/minimax-m2)Proprietary Model ReleaseCited by:[Appendix D](https://arxiv.org/html/2606.02765#A4.p2.1)\.
- C\. Olah, N\. Cammarata, L\. Schubert, G\. Goh, M\. Petrov, and S\. Carter \(2020\)Zoom in: an introduction to circuits\.Distill\.Note:https://distill\.pub/2020/circuits/zoom\-inExternal Links:[Document](https://dx.doi.org/10.23915/distill.00024.001)Cited by:[item Polysemanticity](https://arxiv.org/html/2606.02765#A1.I1.ix9.p1.1),[§1](https://arxiv.org/html/2606.02765#S1.p2.5)\.
- OpenAI \(2025\)GPT oss: open source generative pre\-trained transformers\.Note:[https://huggingface\.co/openai](https://huggingface.co/openai)Accessed: 2026Cited by:[Appendix D](https://arxiv.org/html/2606.02765#A4.p2.1)\.
- K\. Park, Y\. J\. Choe, and V\. Veitch \(2024\)The linear representation hypothesis and the geometry of large language models\.External Links:2311\.03658,[Link](https://arxiv.org/abs/2311.03658)Cited by:[§1](https://arxiv.org/html/2606.02765#S1.p3.3)\.
- G\. Team and G\. DeepMind \(2024\)Gemma 2: improving open language models at a practical size\.arXiv preprint arXiv:2408\.00118\.External Links:[Link](https://arxiv.org/abs/2408.00118)Cited by:[Appendix D](https://arxiv.org/html/2606.02765#A4.p2.1)\.
- G\. Team, A\. Zeng, B\. Xu, B\. Wang, C\. Zhang, D\. Yin, D\. Rojas, G\. Feng, H\. Cao, H\. Zhao,et al\.\(2024\)GLM\-4: all tools integrated\.arXiv preprint arXiv:2406\.12793\.External Links:[Link](https://arxiv.org/abs/2406.12793)Cited by:[Appendix D](https://arxiv.org/html/2606.02765#A4.p2.1)\.
- A\. Templeton, T\. Conerly, J\. Marcus, J\. Lindsey, T\. Bricken, B\. Chen, A\. Pearce, C\. Citro, E\. Ameisen, A\. Jones, H\. Cunningham, N\. L\. Turner, C\. McDougall, M\. MacDiarmid, C\. D\. Freeman, T\. R\. Sumers, E\. Rees, J\. Batson, A\. Jermyn, S\. Carter, C\. Olah, and T\. Henighan \(2024\)Scaling monosemanticity: extracting interpretable features from claude 3 sonnet\.Transformer Circuits Thread\.External Links:[Link](https://transformer-circuits.pub/2024/scaling-monosemanticity/index.html)Cited by:[§1](https://arxiv.org/html/2606.02765#S1.p3.3),[§1](https://arxiv.org/html/2606.02765#S1.p5.1)\.
- A\. Yang, B\. Yang, B\. Hui, B\. Zheng, B\. Yu, C\. Zhou, C\. Li, C\. Li, D\. Liu, F\. Huang, G\. Dong, H\. Zhao, H\. Wang, H\. Wei, H\. Yin,et al\.\(2024\)Qwen2 technical report\.arXiv preprint arXiv:2407\.10671\.Note:Qwen 2\.5 is built upon the architecture detailed in the Qwen2 reportExternal Links:[Link](https://arxiv.org/abs/2407.10671)Cited by:[Appendix D](https://arxiv.org/html/2606.02765#A4.p2.1)\.
- P\. Zhang, G\. Zeng, T\. Wang, and W\. Lu \(2024\)TinyLlama: an open\-source small language model\.arXiv preprint arXiv:2401\.02385\.External Links:[Link](https://arxiv.org/abs/2401.02385)Cited by:[Appendix D](https://arxiv.org/html/2606.02765#A4.p2.1)\.
## Appendix AGlossary
The following glossary provides definitions for key terms and concepts used throughout this paper, particularly those relevant to the analysis of model dimension and representational capacity\.
EmbeddingsThe initial vector representations of tokens, produced by the embedding matrix\. These are the starting points for the model’s internal processing\. Once they pass through the first decoder block, they are no longer considered embeddings and instead become latents\.
LatentsThe vector outputs of the decoder blocks, representing the evolving representations of tokens as they pass through the network\. Unlike embeddings, which are static lookups from a learned matrix, latents are dynamically computed based on context\.
Note on terminology:Related works often use the term “activations” interchangeably with what this paper refers to as latents\. Here, “activations” specifically denotes the output vectors of individual neural layers \(e\.g\., the query vectors resulting fromWQXW\_\{Q\}X\), whereas “latents” refers to the full residual stream representations after each decoder block\.
Latent SpaceThe high\-dimensional vector space where the model’s internal representations \(embeddings and latents\) reside\. The dimensionality of this space is determined by the model dimensiondmodeld\_\{model\}\.
Feature/ConceptAn interpretable unit of information or functionality within the model, such as a specific entity, grammatical rule, or abstract idea\. Under the Linear Representation Hypothesis, features are represented as directions in the latent space\.
Linear Representation HypothesisA hypothesis purporting that neural language models represent concepts and features as linear directions in the latent space\. Formally, for any conceptcc, there exists a direction vector𝒗c∈ℝdmodel\\boldsymbol\{v\}\_\{c\}\\in\\mathbb\{R\}^\{d\_\{model\}\}such that the activation or presence of conceptccin a latent vector𝒙∈ℝdmodel\\boldsymbol\{x\}\\in\\mathbb\{R\}^\{d\_\{model\}\}is given by the inner product⟨𝒙,𝒗c⟩\\langle\\boldsymbol\{x\},\\boldsymbol\{v\}\_\{c\}\\rangle\.
Superposition HypothesisA hypothesis proposing that neural networks leverage the properties of high\-dimensional geometry—specifically near\-orthogonality—to represent more concepts than the number of available dimensions\. This allows the number of encoded concepts to grow exponentially withdmodeld\_\{model\}\.
Near\-orthogonalityA geometric property describing a set of vectors that are approximately orthogonal to each other\. Formally, a set of unit vectors𝒱=\{𝒗1,…,𝒗k\}⊂ℝd\\mathcal\{V\}=\\\{\\boldsymbol\{v\}\_\{1\},\\dots,\\boldsymbol\{v\}\_\{k\}\\\}\\subset\\mathbb\{R\}^\{d\}isε\\varepsilon\-nearly orthogonal if the magnitude of the inner product between any distinct pair is bounded by an accepted deviationε\\varepsilon, meaning\|⟨𝒗i,𝒗j⟩\|≤ε\|\\langle\\boldsymbol\{v\}\_\{i\},\\boldsymbol\{v\}\_\{j\}\\rangle\|\\leq\\varepsilonfor alli≠ji\\neq j, given0<ε<10<\\varepsilon<1\.
Johnson\-Lindenstrauss \(JL\) LemmaA fundamental result in high\-dimensional geometry establishing that points in high\-dimensional space can be embedded into lower dimensions while approximately preserving pairwise distances and inner products\. Critically for the Superposition Hypothesis, the lemma’s distance preservation guarantee implies that orthogonal vectors remain near\-orthogonal after projection, enabling high\-dimensional spaces to accommodate exponentially many near\-orthogonal directions:k≤exp\(ε2dC\)k\\leq\\exp\(\\frac\{\\varepsilon^\{2\}d\}\{C\}\)for some constantC\>0C\>0\(see full derivation in Appendix[B](https://arxiv.org/html/2606.02765#A2)\)\.
PolysemanticityThe phenomenon observed in neural networks where individual neurons activate for multiple, distinct and often unrelated inputs or concepts \(Olahet al\.\[[2020](https://arxiv.org/html/2606.02765#bib.bib18)\]\)\.
Representational CapacityA quantitative upper bound on the number of near\-orthogonal directions available within a model’s latent space, given its model dimensiondmodeld\_\{model\}and the accepted deviationε\\varepsilon\. Embeddings, unembeddings, and features all draw from this shared geometric resource\.
Distributed RepresentationsA representation paradigm where each feature is encoded as a pattern of activation across multiple basis dimensions, and each basis dimension participates in representing multiple features \(Bengioet al\.\[[2014](https://arxiv.org/html/2606.02765#bib.bib17)\]\)\.
Note on terminology:In the original distributed representations literature, the term “feature” referred to individual dimensions or neurons in the representation space\. Under this paper’s terminology—where “feature” is synonymous with “concept” \(see Feature/Concept above\)—distributed representations describe how each interpretable concept is spread across many basis dimensions, and conversely, each basis dimension contributes to encoding many concepts\.
Leaning TowardA geometric relationship describing when a vector has a higher\-than\-expected projection onto a particular direction, where the expectation is set by the accepted deviationε\\varepsilonfor near\-orthogonality\. When a vector𝒗\\boldsymbol\{v\}leans toward a direction𝒅\\boldsymbol\{d\}, the inner product⟨𝒗,𝒅⟩\>ε\\langle\\boldsymbol\{v\},\\boldsymbol\{d\}\\rangle\>\\varepsilon, indicating meaningful alignment rather than incidental similarity\. This concept is central to understanding how embeddings encode information about features while remaining near\-orthogonal to unrelated directions\.
## Appendix BJohnson\-Lindenstrauss Lemma
###### Theorem B\.1\(Johnson\-Lindenstrauss Lemma, Angle Form\)\.
Let0<δ<10<\\delta<1and letX=\{𝐱1,…,𝐱k\}X=\\\{\\boldsymbol\{x\}\_\{1\},\\dots,\\boldsymbol\{x\}\_\{k\}\\\}be a set of unit vectors in a \(possibly infinite\-dimensional\) Hilbert space\. There exists a linear map
f:ℝN→ℝdwithd=O\(δ−2logk\),f:\\mathbb\{R\}^\{N\}\\to\\mathbb\{R\}^\{d\}\\quad\\text\{with\}\\quad d=O\(\\delta^\{\-2\}\\log k\),such that for alli,ji,j,
\|⟨f\(𝒙i\),f\(𝒙j\)⟩−⟨𝒙i,𝒙j⟩\|≤3δ\.\\bigl\|\\langle f\(\\boldsymbol\{x\}\_\{i\}\),f\(\\boldsymbol\{x\}\_\{j\}\)\\rangle\-\\langle\\boldsymbol\{x\}\_\{i\},\\boldsymbol\{x\}\_\{j\}\\rangle\\bigr\|\\leq 3\\delta\.In particular, if𝐱i⟂𝐱j\\boldsymbol\{x\}\_\{i\}\\perp\\boldsymbol\{x\}\_\{j\}, then
\|⟨f\(𝒙i\),f\(𝒙j\)⟩\|≤2δ,\|\\langle f\(\\boldsymbol\{x\}\_\{i\}\),f\(\\boldsymbol\{x\}\_\{j\}\)\\rangle\|\\leq 2\\delta,so the images are near\-orthogonal\.
###### Proof\.
We begin with the standard Johnson\-Lindenstrauss lemma\. There exists a linear mapffsuch that for all𝒖,𝒗∈X∪\{𝟎\}\\boldsymbol\{u\},\\boldsymbol\{v\}\\in X\\cup\\\{\\boldsymbol\{0\}\\\},
\(1−δ\)‖𝒖−𝒗‖2≤‖f\(𝒖\)−f\(𝒗\)‖2≤\(1\+δ\)‖𝒖−𝒗‖2\.\(1\-\\delta\)\\\|\\boldsymbol\{u\}\-\\boldsymbol\{v\}\\\|^\{2\}\\leq\\\|f\(\\boldsymbol\{u\}\)\-f\(\\boldsymbol\{v\}\)\\\|^\{2\}\\leq\(1\+\\delta\)\\\|\\boldsymbol\{u\}\-\\boldsymbol\{v\}\\\|^\{2\}\.\(8\)
Step 1: Norm preservation\.Setting𝒗=𝟎\\boldsymbol\{v\}=\\boldsymbol\{0\}in \([8](https://arxiv.org/html/2606.02765#A2.E8)\) yields
\(1−δ\)‖𝒖‖2≤‖f\(𝒖\)‖2≤\(1\+δ\)‖𝒖‖2\.\(1\-\\delta\)\\\|\\boldsymbol\{u\}\\\|^\{2\}\\leq\\\|f\(\\boldsymbol\{u\}\)\\\|^\{2\}\\leq\(1\+\\delta\)\\\|\\boldsymbol\{u\}\\\|^\{2\}\.Since each𝒙i\\boldsymbol\{x\}\_\{i\}is a unit vector,
1−δ≤‖f\(𝒙i\)‖2≤1\+δ\.1\-\\delta\\leq\\\|f\(\\boldsymbol\{x\}\_\{i\}\)\\\|^\{2\}\\leq 1\+\\delta\.\(9\)
Step 2: Distance between two unit vectors\.For any unit vectors𝒖,𝒗\\boldsymbol\{u\},\\boldsymbol\{v\},
‖𝒖−𝒗‖2=‖𝒖‖2\+‖𝒗‖2−2⟨𝒖,𝒗⟩=2−2⟨𝒖,𝒗⟩\.\\\|\\boldsymbol\{u\}\-\\boldsymbol\{v\}\\\|^\{2\}=\\\|\\boldsymbol\{u\}\\\|^\{2\}\+\\\|\\boldsymbol\{v\}\\\|^\{2\}\-2\\langle\\boldsymbol\{u\},\\boldsymbol\{v\}\\rangle=2\-2\\langle\\boldsymbol\{u\},\\boldsymbol\{v\}\\rangle\.Applying \([8](https://arxiv.org/html/2606.02765#A2.E8)\),
2\(1−δ\)\(1−⟨𝒖,𝒗⟩\)≤‖f\(𝒖\)−f\(𝒗\)‖2≤2\(1\+δ\)\(1−⟨𝒖,𝒗⟩\)\.2\(1\-\\delta\)\(1\-\\langle\\boldsymbol\{u\},\\boldsymbol\{v\}\\rangle\)\\leq\\\|f\(\\boldsymbol\{u\}\)\-f\(\\boldsymbol\{v\}\)\\\|^\{2\}\\leq 2\(1\+\\delta\)\(1\-\\langle\\boldsymbol\{u\},\\boldsymbol\{v\}\\rangle\)\.\(10\)
Step 3: Recover inner products using polarization\.For any vectors𝒂,𝒃\\boldsymbol\{a\},\\boldsymbol\{b\},
⟨𝒂,𝒃⟩=‖𝒂‖2\+‖𝒃‖2−‖𝒂−𝒃‖22\.\\langle\\boldsymbol\{a\},\\boldsymbol\{b\}\\rangle=\\frac\{\\\|\\boldsymbol\{a\}\\\|^\{2\}\+\\\|\\boldsymbol\{b\}\\\|^\{2\}\-\\\|\\boldsymbol\{a\}\-\\boldsymbol\{b\}\\\|^\{2\}\}\{2\}\.Applying this tof\(𝒖\),f\(𝒗\)f\(\\boldsymbol\{u\}\),f\(\\boldsymbol\{v\}\),
⟨f\(𝒖\),f\(𝒗\)⟩=‖f\(𝒖\)‖2\+‖f\(𝒗\)‖2−‖f\(𝒖\)−f\(𝒗\)‖22\.\\langle f\(\\boldsymbol\{u\}\),f\(\\boldsymbol\{v\}\)\\rangle=\\frac\{\\\|f\(\\boldsymbol\{u\}\)\\\|^\{2\}\+\\\|f\(\\boldsymbol\{v\}\)\\\|^\{2\}\-\\\|f\(\\boldsymbol\{u\}\)\-f\(\\boldsymbol\{v\}\)\\\|^\{2\}\}\{2\}\.\(11\)
Step 4: Bounding the inner product distortion\.Using \([9](https://arxiv.org/html/2606.02765#A2.E9)\) and the upper bound from \([10](https://arxiv.org/html/2606.02765#A2.E10)\),
⟨f\(𝒖\),f\(𝒗\)⟩\\displaystyle\\langle f\(\\boldsymbol\{u\}\),f\(\\boldsymbol\{v\}\)\\rangle≥2\(1−δ\)−2\(1\+δ\)\(1−⟨𝒖,𝒗⟩\)2\\displaystyle\\geq\\frac\{2\(1\-\\delta\)\-2\(1\+\\delta\)\(1\-\\langle\\boldsymbol\{u\},\\boldsymbol\{v\}\\rangle\)\}\{2\}=\(1−δ\)−\(1\+δ\)\+\(1\+δ\)⟨𝒖,𝒗⟩\\displaystyle=\(1\-\\delta\)\-\(1\+\\delta\)\+\(1\+\\delta\)\\langle\\boldsymbol\{u\},\\boldsymbol\{v\}\\rangle=\(1\+δ\)⟨𝒖,𝒗⟩−2δ\.\\displaystyle=\(1\+\\delta\)\\langle\\boldsymbol\{u\},\\boldsymbol\{v\}\\rangle\-2\\delta\.Similarly, using the lower bound from \([10](https://arxiv.org/html/2606.02765#A2.E10)\),
⟨f\(𝒖\),f\(𝒗\)⟩\\displaystyle\\langle f\(\\boldsymbol\{u\}\),f\(\\boldsymbol\{v\}\)\\rangle≤2\(1\+δ\)−2\(1−δ\)\(1−⟨𝒖,𝒗⟩\)2\\displaystyle\\leq\\frac\{2\(1\+\\delta\)\-2\(1\-\\delta\)\(1\-\\langle\\boldsymbol\{u\},\\boldsymbol\{v\}\\rangle\)\}\{2\}=\(1\+δ\)−\(1−δ\)\+\(1−δ\)⟨𝒖,𝒗⟩\\displaystyle=\(1\+\\delta\)\-\(1\-\\delta\)\+\(1\-\\delta\)\\langle\\boldsymbol\{u\},\\boldsymbol\{v\}\\rangle=\(1−δ\)⟨𝒖,𝒗⟩\+2δ\.\\displaystyle=\(1\-\\delta\)\\langle\\boldsymbol\{u\},\\boldsymbol\{v\}\\rangle\+2\\delta\.Thus,
\|⟨f\(𝒖\),f\(𝒗\)⟩−⟨𝒖,𝒗⟩\|≤2δ\+δ\|⟨𝒖,𝒗⟩\|≤3δ,\|\\langle f\(\\boldsymbol\{u\}\),f\(\\boldsymbol\{v\}\)\\rangle\-\\langle\\boldsymbol\{u\},\\boldsymbol\{v\}\\rangle\|\\leq 2\\delta\+\\delta\|\\langle\\boldsymbol\{u\},\\boldsymbol\{v\}\\rangle\|\\leq 3\\delta,where the last inequality uses\|⟨𝒖,𝒗⟩\|≤1\|\\langle\\boldsymbol\{u\},\\boldsymbol\{v\}\\rangle\|\\leq 1for unit vectors\.
Step 5: Near\-orthogonality\.If𝒖⟂𝒗\\boldsymbol\{u\}\\perp\\boldsymbol\{v\}, then⟨𝒖,𝒗⟩=0\\langle\\boldsymbol\{u\},\\boldsymbol\{v\}\\rangle=0, and the bounds in Step 4 give
\|⟨f\(𝒖\),f\(𝒗\)⟩\|≤2δ\.\|\\langle f\(\\boldsymbol\{u\}\),f\(\\boldsymbol\{v\}\)\\rangle\|\\leq 2\\delta\.Hence the images are near\-orthogonal\.
Step 6: Angles\.Since‖f\(𝒖\)‖,‖f\(𝒗\)‖∈\[1−δ,1\+δ\]\\\|f\(\\boldsymbol\{u\}\)\\\|,\\\|f\(\\boldsymbol\{v\}\)\\\|\\in\[\\sqrt\{1\-\\delta\},\\sqrt\{1\+\\delta\}\], the cosine of the angleθ′\\theta^\{\\prime\}betweenf\(𝒖\)f\(\\boldsymbol\{u\}\)andf\(𝒗\)f\(\\boldsymbol\{v\}\)satisfies
\|cosθ′\|=\|⟨f\(𝒖\),f\(𝒗\)⟩\|‖f\(𝒖\)‖‖f\(𝒗\)‖≤3δ1−δ=O\(δ\)\.\|\\cos\\theta^\{\\prime\}\|=\\frac\{\|\\langle f\(\\boldsymbol\{u\}\),f\(\\boldsymbol\{v\}\)\\rangle\|\}\{\\\|f\(\\boldsymbol\{u\}\)\\\|\\\|f\(\\boldsymbol\{v\}\)\\\|\}\\leq\\frac\{3\\delta\}\{1\-\\delta\}=O\(\\delta\)\.
This completes the proof\. ∎
## Appendix CEmbedding and Unembedding Relationship
This appendix examines the relationship between embedding and unembedding matrices, which informs the interpretation of embedding\-basedε\\varepsilonestimates in Section[2](https://arxiv.org/html/2606.02765#S2)and motivates the qualification noted there for tied models\.
At the output layer, the final latent representations are converted to logits via the unembedding matrix𝑼∈ℝV×dmodel\\boldsymbol\{U\}\\in\\mathbb\{R\}^\{V\\times d\_\{model\}\}, with the logit for tokeniigiven by the inner product⟨𝒉,𝒖i⟩\\langle\\boldsymbol\{h\},\\boldsymbol\{u\}\_\{i\}\\ranglebetween the final latent𝒉\\boldsymbol\{h\}and the corresponding row𝒖i\\boldsymbol\{u\}\_\{i\}\. In*untied*models,𝑬\\boldsymbol\{E\}and𝑼\\boldsymbol\{U\}are learned independently and may differ significantly; in*tied*models the same matrix serves both roles\.
Figure 9:Cosine similarity between corresponding token embeddings and unembeddings for several models with untied matrices\.Comparing the embedding and unembedding vectors for each token in untied models \(Figure[9](https://arxiv.org/html/2606.02765#A3.F9)\) reveals that they are largely orthogonal: the model transforms representations from an input embedding space into a distinct output unembedding space over the course of the residual stream\. The similarity distributions of embeddings \(Section[2](https://arxiv.org/html/2606.02765#S2)\) and unembeddings \(Figure[10](https://arxiv.org/html/2606.02765#A3.F10)\) also differ noticeably, likely reflecting the different roles these matrices play—embeddings provide a useful starting representation for contextualization, while unembeddings must support accurate next\-token prediction\.
Figure 10:Distribution of cosine similarity between each token unembedding and all others\.In tied models the embedding similarity distributions follow the unembedding pattern, since the two are the same matrix\. This is relevant to the two\-class split in Section[2](https://arxiv.org/html/2606.02765#S2): nearly all Class 1 \(high\-ε\\varepsilon\) models are tied, suggesting that forcing one matrix to satisfy both objectives discourages the near\-orthogonal structure that untied embeddings exhibit\. Tying does not fully account for the divide—a few tied models maintain lowε\\varepsilonand some untied ones do not—so the structural cause remains open\.
## Appendix DModels Analyzed
Per\-model values ofdmodeld\_\{model\}, the estimated accepted deviationε\\varepsilon\(μ\+2σ\\mu\+2\\sigmaof the pairwise embedding cosine similarity distribution\), and the resulting representational capacity computed via the fully parameterized form of \([6](https://arxiv.org/html/2606.02765#S3.E6)\) \(Section[3\.3](https://arxiv.org/html/2606.02765#S3.SS3)\)\. Class 1 models \(ε\>0\.2\\varepsilon\>0\.2\) are marked indeterminate: their embeddings lack near\-orthogonal structure, so the bound is vacuous \(cf\. Section[3\.5](https://arxiv.org/html/2606.02765#S3.SS5)\)\. Across the Class 2 models, the bulk of capacities top out near∼1010\\sim 10^\{10\}–101110^\{11\}; the lone exception is Llama 2 70B, whoseε=0\.09962\\varepsilon=0\.09962sits at the very edge of the Class 1 boundary and lands its capacity at∼1015\\sim 10^\{15\}\.
Table 1:Per\-modeldmodeld\_\{model\}, accepted deviationε\\varepsilon, and representational capacity\.Model Namedmodeld\_\{model\}ε\\varepsilonRep\. CapacityDeepSeek V3\.2 Exp71680\.059171\.16e91\.16\\mathrm\{e\}\{9\}DeepSeek R171680\.059171\.16e91\.16\\mathrm\{e\}\{9\}Gemma 7B30720\.99918indeterminateGemma 2 2B23040\.35084indeterminateGemma 2 9B35840\.39785indeterminateGemma 2 27B46080\.090134\.77e104\.77\\mathrm\{e\}\{10\}Gemma 3 270M6400\.41554indeterminateGemma 3 1B11520\.20164indeterminateGemma 3 12B38400\.086352\.51e92\.51\\mathrm\{e\}\{9\}Gemma 3 27B53760\.074214\.36e94\.36\\mathrm\{e\}\{9\}GLM 4\.651200\.058826\.14e76\.14\\mathrm\{e\}\{7\}GPT OSS 120B28800\.22984indeterminateGPT OSS 20B28800\.28120indeterminateKimi K271680\.067631\.47e101\.47\\mathrm\{e\}\{10\}Llama 2 7B40960\.064503\.99e73\.99\\mathrm\{e\}\{7\}Llama 2 13B51200\.067665\.11e85\.11\\mathrm\{e\}\{8\}Llama 2 70B81920\.099628\.15e158\.15\\mathrm\{e\}\{15\}Llama 3 8B40960\.059451\.42e71\.42\\mathrm\{e\}\{7\}Llama 3\.1 8B40960\.055426\.37e66\.37\\mathrm\{e\}\{6\}Llama 3\.1 70B81920\.044134\.39e74\.39\\mathrm\{e\}\{7\}Llama 3\.1 405B163840\.048131\.22e111\.22\\mathrm\{e\}\{11\}Llama 3\.2 1B20480\.28223indeterminateLlama 3\.2 3B30720\.29395indeterminateMiniMax M230720\.076184\.39e74\.39\\mathrm\{e\}\{7\}Mistral Small 3\.2 24B51200\.075203\.39e93\.39\\mathrm\{e\}\{9\}Mistral 7B Instruct v0\.340960\.070191\.33e81\.33\\mathrm\{e\}\{8\}Phi 225600\.054974\.57e54\.57\\mathrm\{e\}\{5\}Phi 3 mini 128k30720\.055361\.21e61\.21\\mathrm\{e\}\{6\}Phi 451200\.058655\.91e75\.91\\mathrm\{e\}\{7\}Qwen 2\.5 0\.5B8960\.36035indeterminateQwen 2\.5 1\.5B15360\.25998indeterminateQwen 2\.5 3B20480\.21014indeterminateQwen 2\.5 7B35840\.075011\.20e81\.20\\mathrm\{e\}\{8\}Qwen 2\.5 14B51200\.052691\.51e71\.51\\mathrm\{e\}\{7\}Qwen 2\.5 32B51200\.055542\.88e72\.88\\mathrm\{e\}\{7\}Qwen 2\.5 72B81920\.061318\.46e98\.46\\mathrm\{e\}\{9\}Qwen 3 0\.6B10240\.25204indeterminateQwen 3 4B25600\.20834indeterminateQwen 3 8B40960\.059691\.49e71\.49\\mathrm\{e\}\{7\}Qwen 3 32B51200\.072631\.77e91\.77\\mathrm\{e\}\{9\}Qwen 3 30B A3B20480\.064565\.67e55\.67\\mathrm\{e\}\{5\}Qwen 3 Next 80B A3B20480\.074222\.07e62\.07\\mathrm\{e\}\{6\}Qwen 3 235B A22B40960\.048741\.77e61\.77\\mathrm\{e\}\{6\}Qwen 3 Coder 480B61440\.046601\.21e71\.21\\mathrm\{e\}\{7\}TinyLlama 1\.1B20480\.063294\.81e54\.81\\mathrm\{e\}\{5\}Similar Articles
Characterizing the Representational Capacity of Neural Processes
This paper theoretically characterizes the representational capacity of Neural Process (NP) architectures, proving a strict hierarchy among Conditional, Attentive, Convolutional, and Transformer NPs, and showing that finite-dimensional latent variables do not expand representational capacity beyond the encoder.
The Geometry of Inference in Transformer Residual Streams
This paper studies how transformer intermediate residual states become specialized to their own final output states, finding that directional alignment and endpoint rank improve even when Euclidean distance barely changes across six pretrained language models. It offers a high-dimensional model separating norm, alignment, and endpoint geometry, proving that straight-path convergence cannot introduce new competitors, and linking residual geometry to output token rankings.
Canonical-basis realignment for Transformer LLMs: every hidden axis becomes independently measurable and controllable
This research presents a canonical basis transformation for Transformer language models, enabling lossless, independent measurement and control of hidden space axes while maintaining model behavior. It includes demonstrations on multiple models and provides causal evidence for functional geometry in LLMs.
On the Expressive Power of Transformers
A survey paper examining the expressive power of transformers as language recognizers, using concepts and methods from circuit complexity to compare them with classical models of computation.
Disentangling Representation Evolution in Transformers through Directional Decomposition
This paper decomposes transformer representation updates into parallel and perpendicular components to study evolution geometry, linking it to editing robustness, compression diagnosis, and training improvements.