Capsule Lens: Locating and Tracking Concept Geometry in Model Representations

arXiv cs.LG Papers

Summary

Capsule Lens is a framework for mechanistic interpretability that uses geometric capsules to locate and track how concepts are encoded in neural network representations, enabling analysis of both static and dynamic model internals.

arXiv:2609.05575v1 Announce Type: new Abstract: Understanding how concepts are encoded in the internal representations of machine learning models is a central problem in mechanistic interpretability, essential both for the science of deep learning and for the trustworthy deployment of increasingly capable models. Existing approaches to interpret model representations mainly map representations onto more interpretable spaces and do not directly characterize how concepts occupy representation space; various hypotheses have been proposed, but often lack of rigorous validation and largely focus on static representations. In this work, we introduce Capsule Lens, a framework that matches the region a concept occupies with a simple, trackable geometric form, a capsule, defined by several interpretable parameters, fitted in closed form to each concept's geometry and validated on held-out samples. We apply Capsule Lens in two major settings: static and dynamic representations. On static representations, we demonstrate how to locate concept geometry across various models, and how the span and norm curves uncover important geometric characteristics. On dynamic representations, we present three case studies tracking representation drifts induced by distinct training settings, CLIP pretraining, RL post-training on visual question answering, and RL post-training on mathematical reasoning. These analyses reveal qualitatively different geometric dynamics, ranging from broad network-wide restructuring in CLIP pretraining to localized and concept-specific changes in RL post-training. Our results include findings aligned with existing literature as well as novel observations. We believe Capsule Lens stands as a promising tool for locating, analyzing, and tracking concept geometry in both static and dynamic representations.
Original Article
View Cached Full Text

Cached at: 09/10/26, 08:22 AM

# Capsule Lens: Locating and Tracking Concept Geometry in Model Representations
Source: [https://arxiv.org/html/2609.05575](https://arxiv.org/html/2609.05575)
Yiming Tang††thanks:Equal contribution\. Email:yiming@nus\.edu\.sgHarshvardhan Saini11footnotemark:1Affiliation:Indian Institute of Technology \(ISM\) DhanbadSamyak JhaAffiliation:Indian Institute of Technology \(ISM\) DhanbadHuaming ChenAffiliation:The University of SydneyXufeng DuanAffiliation:The Chinese University of Hong KongDianbo Liu††thanks:Corresponding author\. Email:dianbo@nus\.edu\.sgAffiliation:National University of Singapore

###### Abstract

Understanding how concepts are encoded in the internal representations of machine learning models is a central problem in mechanistic interpretability, essential both for the science of deep learning and for the trustworthy deployment of increasingly capable models\. Existing approaches to interpret model representations mainly map representations onto more interpretable spaces and do not directly characterize how concepts occupy representation space; various hypotheses have been proposed, but often lack of rigorous validation and largely focus on static representations\. In this work, we introduceCapsule Lens, a framework that matches the region a concept occupies with a simple, trackable geometric form, a*capsule*, defined by several interpretable parameters, fitted in closed form to each concept’s geometry and validated on held\-out samples\. We apply Capsule Lens in two major settings: static and dynamic representations\. On static representations, we demonstrate how to locate concept geometry across various models, and how the span and norm curves uncover important geometric characteristics\. On dynamic representations, we present three case studies tracking representation drifts induced by distinct training settings, CLIP pretraining, RL post\-training on visual question answering, and RL post\-training on mathematical reasoning\. These analyses reveal qualitatively different geometric dynamics, ranging from broad network\-wide restructuring in CLIP pretraining to localized and concept\-specific changes in RL post\-training\. Our results include findings aligned with existing literature as well as novel observations\. We believe Capsule Lens stands as a promising tool for locating, analyzing, and tracking concept geometry in both static and dynamic representations\.

## 1Introduction

Understanding how concepts are encoded in the internal representations of neural networks is a central goal of mechanistic interpretability\([Sharkey et al\., 2025](https://arxiv.org/html/2609.05575#bib.bib5)\)\. As models achieve remarkable capabilities across diverse domains\([Brown et al\., 2020](https://arxiv.org/html/2609.05575#bib.bib42);[Guo et al\., 2025](https://arxiv.org/html/2609.05575#bib.bib67);[Grattafiori et al\., 2024](https://arxiv.org/html/2609.05575#bib.bib41)\), deciphering what they represent—and how—is important for both the science of deep learning and the trustworthy deployment of increasingly capable systems\([Bereska and Gavves, 2024](https://arxiv.org/html/2609.05575#bib.bib47)\)\. Beyond this static picture, an equally important question is how these internal representations change during training\. Post\-training methods such as supervised fine\-tuning and reinforcement learning are now deployed at scale\([Ouyang et al\., 2022](https://arxiv.org/html/2609.05575#bib.bib6);[Shao et al\., 2024](https://arxiv.org/html/2609.05575#bib.bib11)\), yet their effects on model internals remain poorly understood\([Wang et al\., 2025b](https://arxiv.org/html/2609.05575#bib.bib66);[Xu et al\., 2024](https://arxiv.org/html/2609.05575#bib.bib27)\)\. This gap is consequential, as improper fine\-tuning can induce broad behavioral changes, including misalignment\([Betley et al\., 2026](https://arxiv.org/html/2609.05575#bib.bib68);[Dong et al\., 2025](https://arxiv.org/html/2609.05575#bib.bib13)\)\. Understanding such dynamics first requires locating concept geometry in representation space and then tracking how it changes during training\.

To address these questions, researchers have developed interpretability lenses and neural geometry approaches, each addressing a different part of the problem\. Interpretability lenses map internal activations onto more interpretable spaces: the Logit Lens and Tuned Lens map hidden states onto tokens\([nostalgebraist, 2020](https://arxiv.org/html/2609.05575#bib.bib29);[Belrose et al\., 2023](https://arxiv.org/html/2609.05575#bib.bib20)\), Patchscopes and Rep2Text decode representations into natural language\([Ghandeharioun et al\., 2024](https://arxiv.org/html/2609.05575#bib.bib21);[Zhao et al\., 2025](https://arxiv.org/html/2609.05575#bib.bib1)\), and sparse autoencoders decompose polysemantic activations into more interpretable features\([Cunningham et al\., 2023](https://arxiv.org/html/2609.05575#bib.bib2);[Gao et al\., 2024](https://arxiv.org/html/2609.05575#bib.bib24);[Tang et al\., 2025a](https://arxiv.org/html/2609.05575#bib.bib30)\)\. These lenses reveal*what*information is represented, but provide limited insight into*how*a concept occupies representation space\. Neural geometry instead studies the geometric structure of representations, proposing that concepts may be encoded as linear directions\([Park et al\., 2023](https://arxiv.org/html/2609.05575#bib.bib33);[Elhage et al\., 2022](https://arxiv.org/html/2609.05575#bib.bib32)\), polytopes and hierarchies\([Park et al\., 2024](https://arxiv.org/html/2609.05575#bib.bib38)\), or irreducibly multidimensional manifolds\([Engels et al\., 2024](https://arxiv.org/html/2609.05575#bib.bib35)\)\. However, these accounts typically posit or probe geometric structures rather than directly estimating and validating the region occupied by a concept\([Tang et al\., 2025b](https://arxiv.org/html/2609.05575#bib.bib4);[Bhalla et al\., 2026](https://arxiv.org/html/2609.05575#bib.bib48)\)\. Moreover, both lines of work largely focus on static representations, providing limited insight into how concept geometry changes during training\.

Following recent work\([Xiong, 2026](https://arxiv.org/html/2609.05575#bib.bib49);[Modell et al\., 2025](https://arxiv.org/html/2609.05575#bib.bib50)\), we formulate a concept as a binary indicator on a dataset: a sample either exhibits the concept or does not\. Under this formulation, concept geometry is the region occupied by the concept’s positive samples in representation space, approximated by a set of points\([Sorscher et al\., 2022](https://arxiv.org/html/2609.05575#bib.bib45);[Chung and Abbott, 2021](https://arxiv.org/html/2609.05575#bib.bib43);[Bhalla et al\., 2026](https://arxiv.org/html/2609.05575#bib.bib48);[Wollschläger et al\., 2025](https://arxiv.org/html/2609.05575#bib.bib51)\)\. Our key observation is that understanding how a concept is encoded reduces to characterizing this geometry inℝd\\mathbb\{R\}^\{d\}, while understanding how training reshapes the encoding requires tracking how the geometry changes across checkpoints\. Drawing inspiration from diverse geometric views of concept representations\([Engels et al\., 2024](https://arxiv.org/html/2609.05575#bib.bib35);[Modell et al\., 2025](https://arxiv.org/html/2609.05575#bib.bib50);[Sorscher et al\., 2022](https://arxiv.org/html/2609.05575#bib.bib45)\), we approximate each concept geometry with a simple parametric shape—a capsule—whose parameters can be directly compared across representations and model states\.

In this work, we introduce Capsule Lens, a framework that approximates the region occupied by a concept with a simple, trackable geometric form—a*capsule*, defined by a unit axis, an axis span, and a norm band—fitted in closed form to each concept’s representation cloud\. We also introduce two coverage curves, the span curve and the norm curve, which reveal how representations populate the fitted region beyond its boundary parameters, together with a validation protocol that evaluates fitted capsules on held\-out samples\. We apply Capsule Lens in two complementary settings\. For static representations, we characterize concept geometry across various model architectures, examining its internal structure, disentanglement, and variation across model depth\. For dynamic representations, we use three case studies to track how concept geometry changes under distinct training settings, spanning CLIP pretraining and RL post\-training on visual question answering and mathematical reasoning\. These studies reveal qualitatively different geometric dynamics: broad network\-wide restructuring in CLIP pretraining and more localized, concept\-specific changes in RL post\-training\. In particular, we observe a concept\-specific concentration effect under RLVR, where the geometry of a small subset of concepts sharpens at particular layers while most concept regions remain nearly unchanged\. We believe Capsule Lens provides a practical tool for locating, characterizing, and tracking concept geometry across both static representations and training dynamics\.

Our contributions are listed as follows:

- •We propose Capsule Lens, a parametric, closed\-form, and trackable lens on concept geometry, equipped with span and norm coverage curves\.
- •We locate concept geometry across diverse model architectures, characterizing its internal structure, disentanglement, and layer\-wise evolution\.
- •Across three case studies, we uncover qualitatively different geometric dynamics across distinct training settings: broad network\-wide restructuring in CLIP pretraining, and localized, concept\-specific changes in RL post\-training\.
- •We further identify a concept\-specific concentration effect under RLVR, where a small subset of concept geometries sharpen at particular layers while most concepts remain nearly unchanged, revealing an underexplored form of representation change\.

## 2Method

In this section, we introduce our proposed framework, Capsule Lens\. Section[2\.1](https://arxiv.org/html/2609.05575#S2.SS1)introduces the formal definitions of concepts and capsules\. Section[2\.2](https://arxiv.org/html/2609.05575#S2.SS2)describes capsule matching, the closed\-form procedure that fits a capsule to a concept’s representation geometry, along with the span and norm curves that describe how the concept’s representations distribute within the fitted region\. Section[2\.3](https://arxiv.org/html/2609.05575#S2.SS3)presents the held\-out validation protocol used to evaluate fitted capsules\. Section[2\.4](https://arxiv.org/html/2609.05575#S2.SS4)explains how we track concept geometry by comparing matched capsules across training checkpoints\.

![Refer to caption](https://arxiv.org/html/2609.05575v1/Figure1.png)Figure 1:Overview of the Capsule Lens framework\.Given a model and a dataset, the framework extracts representations for a set of predefined concepts\. Capsule matching fits a capsule𝒞⁡\(a,sa,nlow,nhigh\)\\mathcal\{C\}\(a,s\_\{a\},n\_\{\\mathrm\{low\}\},n\_\{\\mathrm\{high\}\}\)to each concept’s representation cloud in closed form\. The fitted capsule parameters and the associated span curveFs​\(t\)F\_\{s\}\(t\)and norm curveFn​\(t\)F\_\{n\}\(t\)are used to analyze the static geometry of concepts\. The drift metrics—axis rotationΔ​a\\Delta a, span changeΔ​sa\\Delta s\_\{a\}, and norm\-band changesΔ​nlow,Δ​nhigh,Δ​Nb\\Delta n\_\{\\mathrm\{low\}\},\\Delta n\_\{\\mathrm\{high\}\},\\Delta N\_\{b\}—are used to track how this geometry shifts across training checkpoints\.### 2\.1Preliminaries: Concepts and Capsules

We work in the representation spaceℝd\\mathbb\{R\}^\{d\}\. A modelϕθ:𝒟→ℝd\\phi\_\{\\theta\}:\\mathcal\{D\}\\to\\mathbb\{R\}^\{d\}maps each sample to a representation, e\.g\., a token’s activations at a given layer of a large language model, or the visual embedding of an image encoder\. For a nonzerox∈ℝdx\\in\\mathbb\{R\}^\{d\}and a unit vectora∈ℝda\\in\\mathbb\{R\}^\{d\}\(‖a‖=1\\\|a\\\|=1\), the*cosine distance*betweenxxandaaisdcos​\(x,a\)=1−⟨x,a⟩‖x‖∈\[0,2\]\.d\_\{\\cos\}\(x,a\)=1\-\\frac\{\\langle x,a\\rangle\}\{\\\|x\\\|\}\\in\[0,2\]\.

###### Definition 1\(Concept Geometry\)\.

A conceptccis a binary indicatorIc:𝒟→\{0,1\}I\_\{c\}:\\mathcal\{D\}\\to\\\{0,1\\\}on a dataset𝒟\\mathcal\{D\}, whose positive set is𝒟c=\{z∈𝒟:Ic​\(z\)=1\}\\mathcal\{D\}\_\{c\}=\\\{z\\in\\mathcal\{D\}:I\_\{c\}\(z\)=1\\\}\. A modelϕθ\\phi\_\{\\theta\}induces the*concept geometry*

Γcθ=ϕθ​\(𝒟c\)⊂ℝd\.\\Gamma\_\{c\}^\{\\theta\}=\\phi\_\{\\theta\}\(\\mathcal\{D\}\_\{c\}\)\\subset\\mathbb\{R\}^\{d\}\.

Fundamentally, a concept geometry is a set of points in the representation space, and characterizing the geometric structure of this point set is a central question in concept\-based mechanistic interpretability\. We approach this question by fitting a simple parametric shape to the point cloud, so that studying the concept geometry reduces to reading and tracking the fitted parameters\.

###### Definition 2\(Capsule\)\.

Given a unit axisa∈ℝda\\in\\mathbb\{R\}^\{d\}with‖a‖=1\\\|a\\\|=1, an axis spansa∈\[0,2\]s\_\{a\}\\in\[0,2\], and norm bounds0≤nlow≤nhigh0\\leq n\_\{\\mathrm\{low\}\}\\leq n\_\{\\mathrm\{high\}\}, the corresponding*capsule*is the region in the representation space:

𝒞\(a,sa,nlow,nhigh\)=\{x∈ℝd\|dcos\(x,a\)≤sa,nlow≤∥x∥≤nhigh\}\.\\mathcal\{C\}\(a,s\_\{a\},n\_\{\\mathrm\{low\}\},n\_\{\\mathrm\{high\}\}\)=\\left\\\{x\\in\\mathbb\{R\}^\{d\}\\;\\middle\|\\;d\_\{\\cos\}\(x,a\)\\leq s\_\{a\},\\;n\_\{\\mathrm\{low\}\}\\leq\\\|x\\\|\\leq n\_\{\\mathrm\{high\}\}\\right\\\}\.\(1\)

Geometrically, a capsule is an angular cone about the axisaaintersected with a norm shell\. The motivation for using this form is that the aspects of concept geometry we care about can be directly read from the parameters of𝒞⁡\(a,sa,nlow,nhigh\)\\mathcal\{C\}\(a,s\_\{a\},n\_\{\\mathrm\{low\}\},n\_\{\\mathrm\{high\}\}\)\. Specifically, the axisaagives the central direction of the concept’s representations; the axis spansas\_\{a\}gives their angular spread around that direction; and the norm boundsnlow,nhighn\_\{\\mathrm\{low\}\},n\_\{\\mathrm\{high\}\}give the magnitude interval they occupy, whose width we write as the*norm band*Nb=nhigh−nlowN\_\{b\}=n\_\{\\mathrm\{high\}\}\-n\_\{\\mathrm\{low\}\}\.

### 2\.2Capsule Matching and Coverage Curves

A concept geometryΓcθ\\Gamma\_\{c\}^\{\\theta\}is in general an intractable point set\. Prior work approximates such geometries as high\-dimensional ellipsoids\([Sorscher et al\., 2022](https://arxiv.org/html/2609.05575#bib.bib45)\)or tracks their topology via persistent homology\([Malhotra et al\., 2026](https://arxiv.org/html/2609.05575#bib.bib44)\); we instead approximateΓcθ\\Gamma\_\{c\}^\{\\theta\}with a finite sample of points,Γ¯cθ=\{x1,…,xm\}⊂ℝd\\bar\{\\Gamma\}\_\{c\}^\{\\theta\}=\\\{x\_\{1\},\\ldots,x\_\{m\}\\\}\\subset\\mathbb\{R\}^\{d\}, extracted fromϕθ\\phi\_\{\\theta\}on the positive set𝒟c\\mathcal\{D\}\_\{c\}, and match a capsule to these points via the closed\-form procedure of Algorithm[1](https://arxiv.org/html/2609.05575#alg1)\. Writingx^i=xi/‖xi‖\\hat\{x\}\_\{i\}=x\_\{i\}/\\\|x\_\{i\}\\\|for the unit\-normalized samples, the axis is the normalized mean direction,

a=r¯/‖r¯‖,r¯=1m​∑i=1mx^i,a=\\bar\{r\}/\\\|\\bar\{r\}\\\|,\\qquad\\bar\{r\}=\\frac\{1\}\{m\}\\sum\_\{i=1\}^\{m\}\\hat\{x\}\_\{i\},\(2\)the axis span is the maximum cosine distance from the samples to the axis,sa=maxi⁡dcos​\(xi,a\)s\_\{a\}=\\max\_\{i\}d\_\{\\cos\}\(x\_\{i\},a\), and the norm bounds are the minimum and maximum sample norms,nlow=mini⁡‖xi‖n\_\{\\mathrm\{low\}\}=\\min\_\{i\}\\\|x\_\{i\}\\\|andnhigh=maxi⁡‖xi‖n\_\{\\mathrm\{high\}\}=\\max\_\{i\}\\\|x\_\{i\}\\\|\. The resulting capsule𝒞⁡\(a,sa,nlow,nhigh\)\\mathcal\{C\}\(a,s\_\{a\},n\_\{\\mathrm\{low\}\},n\_\{\\mathrm\{high\}\}\)covers all fitting samples by construction and provides a compact, closed\-form summary of the concept’s geometry\.

Algorithm 1Capsule Matching0:samples

Γ¯cθ=\{x1,…,xm\}\\bar\{\\Gamma\}\_\{c\}^\{\\theta\}=\\\{x\_\{1\},\\ldots,x\_\{m\}\\\}
0:capsule

𝒞⁡\(a,sa,nlow,nhigh\)\\mathcal\{C\}\(a,s\_\{a\},n\_\{\\mathrm\{low\}\},n\_\{\\mathrm\{high\}\}\)
1:

x^i←xi/‖xi‖\\hat\{x\}\_\{i\}\\leftarrow x\_\{i\}/\\\|x\_\{i\}\\\|for

i=1,…,mi=1,\\ldots,m
2:

r¯←1m​∑i=1mx^i\\bar\{r\}\\leftarrow\\frac\{1\}\{m\}\\sum\_\{i=1\}^\{m\}\\hat\{x\}\_\{i\};

a←r¯/‖r¯‖a\\leftarrow\\bar\{r\}/\\\|\\bar\{r\}\\\|⊳\\trianglerightaxis

3:

sa←maxi⁡dcos​\(xi,a\)s\_\{a\}\\leftarrow\\max\_\{i\}d\_\{\\cos\}\(x\_\{i\},a\)⊳\\trianglerightaxis span

4:

nlow←mini⁡‖xi‖n\_\{\\mathrm\{low\}\}\\leftarrow\\min\_\{i\}\\\|x\_\{i\}\\\|;

nhigh←maxi⁡‖xi‖n\_\{\\mathrm\{high\}\}\\leftarrow\\max\_\{i\}\\\|x\_\{i\}\\\|⊳\\trianglerightnorm bounds

5:return

𝒞⁡\(a,sa,nlow,nhigh\)\\mathcal\{C\}\(a,s\_\{a\},n\_\{\\mathrm\{low\}\},n\_\{\\mathrm\{high\}\}\)

#### Coverage curves\.

The capsule parameters summarize the boundary of the region a concept occupies, but miss how the concept’s representations are distributed within it\. We therefore introduce two coverage curves for the span parameter and the norm parameters: the span curveFs​\(t\)F\_\{s\}\(t\)and the norm curveFn​\(t\)F\_\{n\}\(t\)\. The*span curve*of a capsule is the function

Fs​\(t\)=1\|Γ¯cθ\|​\|\{xi∈Γ¯cθ\|0≤dcos​\(xi,a\)≤t\}\|,F\_\{s\}\(t\)=\\frac\{1\}\{\\left\|\\bar\{\\Gamma\}\_\{c\}^\{\\theta\}\\right\|\}\\left\|\\left\\\{x\_\{i\}\\in\\bar\{\\Gamma\}\_\{c\}^\{\\theta\}\\;\\middle\|\\;0\\leq d\_\{\\cos\}\(x\_\{i\},a\)\\leq t\\right\\\}\\right\|,\(3\)defined fort∈\[0,sa\]t\\in\[0,s\_\{a\}\], reporting the fraction of the concept’s samples covered by the cone of axisaaand spanttasttgrows from00to the fitted span\. The*norm curve*is defined analogously,

Fn​\(t\)=1\|Γ¯cθ\|​\|\{xi∈Γ¯cθ\|nlow≤‖xi‖≤t\}\|,F\_\{n\}\(t\)=\\frac\{1\}\{\\left\|\\bar\{\\Gamma\}\_\{c\}^\{\\theta\}\\right\|\}\\left\|\\left\\\{x\_\{i\}\\in\\bar\{\\Gamma\}\_\{c\}^\{\\theta\}\\;\\middle\|\\;n\_\{\\mathrm\{low\}\}\\leq\\\|x\_\{i\}\\\|\\leq t\\right\\\}\\right\|,\(4\)defined fort∈\[nlow,nhigh\]t\\in\[n\_\{\\mathrm\{low\}\},n\_\{\\mathrm\{high\}\}\], reporting the fraction of samples whose norms fall belowttwithin the fitted norm bounds\. Both curves are empirical cumulative distributions of the quantities bounded by the capsule, increasing the resolution of Capsule Lens beyond its boundary parameters: a span curve that rises steeply near00indicates representations concentrated tightly around the axis with the fitted span determined by a few outlying samples; a curve that rises late indicates a genuinely broad angular distribution; and flat segments separating steep rises suggest multiple angular clusters\.

### 2\.3Capsule Validation

We validate a fitted capsule along two complementary dimensions:*coverage*and*disentanglement*\. We split the positive set𝒟c\\mathcal\{D\}\_\{c\}into disjoint fitting and validation sets𝒟cfit\\mathcal\{D\}\_\{c\}^\{\\mathrm\{fit\}\}and𝒟cval\\mathcal\{D\}\_\{c\}^\{\\mathrm\{val\}\}, fit a capsule𝒞c\\mathcal\{C\}\_\{c\}on𝒟cfit\\mathcal\{D\}\_\{c\}^\{\\mathrm\{fit\}\}, and define its held\-out coverage as

Cov⁡\(𝒞c\)=P⁡\(ϕ⁡\(x\)∈𝒞c∣Ic​\(x\)=1\)=1\|𝒟cval\|​\|\{z∈𝒟cval\|ϕ⁡\(z\)∈𝒞c\}\|\.\\mathrm\{Cov\}\(\\mathcal\{C\}\_\{c\}\)=P\(\\phi\(x\)\\in\\mathcal\{C\}\_\{c\}\\mid I\_\{c\}\(x\)=1\)=\\frac\{1\}\{\\left\|\\mathcal\{D\}\_\{c\}^\{\\mathrm\{val\}\}\\right\|\}\\left\|\\left\\\{z\\in\\mathcal\{D\}\_\{c\}^\{\\mathrm\{val\}\}\\;\\middle\|\\;\\phi\(z\)\\in\\mathcal\{C\}\_\{c\}\\right\\\}\\right\|\.\(5\)This measures whether the fitted region generalizes to unseen samples of the same concept\. To characterize its overlap with other concepts, we construct a negative validation set𝒟not​\-​cval\\mathcal\{D\}\_\{\\mathrm\{not\}\\text\{\-\}c\}^\{\\mathrm\{val\}\}by sampling examples that do not belong to conceptcc, and define the*disentanglement score*

Dis⁡\(c\)=P⁡\(Ic​\(x\)=1∣ϕ⁡\(x\)∈𝒞c\)=ρ\+ρ\+\+ρ−,\\mathrm\{Dis\}\(c\)=P\(I\_\{c\}\(x\)=1\\mid\\phi\(x\)\\in\\mathcal\{C\}\_\{c\}\)=\\frac\{\\rho^\{\+\}\}\{\\rho^\{\+\}\+\\rho^\{\-\}\},\(6\)whereρ\+=1\|𝒟cval\|​\|\{z∈𝒟cval\|ϕ⁡\(z\)∈𝒞c\}\|\\rho^\{\+\}=\\frac\{1\}\{\|\\mathcal\{D\}\_\{c\}^\{\\mathrm\{val\}\}\|\}\\left\|\\left\\\{z\\in\\mathcal\{D\}\_\{c\}^\{\\mathrm\{val\}\}\\;\\middle\|\\;\\phi\(z\)\\in\\mathcal\{C\}\_\{c\}\\right\\\}\\right\|andρ−=1\|𝒟not​\-​cval\|​\|\{z∈𝒟not​\-​cval\|ϕ⁡\(z\)∈𝒞c\}\|\\rho^\{\-\}=\\frac\{1\}\{\|\\mathcal\{D\}\_\{\\mathrm\{not\}\\text\{\-\}c\}^\{\\mathrm\{val\}\}\|\}\\left\|\\left\\\{z\\in\\mathcal\{D\}\_\{\\mathrm\{not\}\\text\{\-\}c\}^\{\\mathrm\{val\}\}\\;\\middle\|\\;\\phi\(z\)\\in\\mathcal\{C\}\_\{c\}\\right\\\}\\right\|\. HigherDis⁡\(c\)\\mathrm\{Dis\}\(c\)indicates stronger separation of the target concept from others within the region𝒞c\\mathcal\{C\}\_\{c\}\.

### 2\.4Tracking Concept Geometry

Because capsules are fitted independently at every model state, comparing the capsule𝒞c\(t\)\\mathcal\{C\}\_\{c\}^\{\(t\)\}fitted at training checkpointttagainst the base capsule𝒞c\(0\)\\mathcal\{C\}\_\{c\}^\{\(0\)\}provides a direct summary of how concept geometries change during training\. Table[1](https://arxiv.org/html/2609.05575#S2.T1)summarizes the drift metrics we track per concept and checkpoint\. Aggregating these signals across concepts and layers turns a single capsule fit into a trajectory that localizes where training reshapes representations \(by layer\), and identifies which concepts it affects most \(by concept\)\. Section[5](https://arxiv.org/html/2609.05575#S5)demonstrates these can track representation drifts\.

Table 1:Capsule drift metrics for tracking concept geometry across checkpoints\.Each metric compares the capsule fitted at checkpointttagainst the base capsule at checkpoint00\. The axis driftΔ​a\\Delta ameasures the rotation of the concept axis; the span changeΔ​sa\\Delta s\_\{a\}measures the widening \(\>0\>0\) or narrowing \(<0<0\) of the concept’s angular spread; and the norm\-band changesΔ​Nb\\Delta N\_\{b\},Δ​nhigh\\Delta n\_\{\\mathrm\{high\}\},Δ​nlow\\Delta n\_\{\\mathrm\{low\}\}measure how the magnitude interval the concept occupies expands, contracts, or shifts\.

## 3Validating Capsule Lens

We validate the reliability of Capsule Lens for locating and tracking concept geometry\. Section[3\.1](https://arxiv.org/html/2609.05575#S3.SS1)describes the construction of visual and textual concepts used throughout our experiments\. Section[3\.2](https://arxiv.org/html/2609.05575#S3.SS2)evaluates fitted capsules using held\-out coverage and disentanglement metrics\. Section[3\.3](https://arxiv.org/html/2609.05575#S3.SS3)compares the capsule with alternative geometric approximations\. Further robustness analyses on sampling and boundary estimation are provided in Appendix[D](https://arxiv.org/html/2609.05575#A4)\.

### 3\.1Concept Construction

We construct both visual and textual concepts, where each concept is represented by a collection of samples exhibiting certain property\. These collections of data samples can be built with human annotation or as token groups in context\.

#### Visual concepts\.

We construct visual concepts from Visual Genome\([Krishna et al\., 2016](https://arxiv.org/html/2609.05575#bib.bib52)\), where human annotators provide named bounding boxes for objects and regions\. An image is considered positive for conceptccif it contains a bounding box labelledcccovering at least1%1\\%of the image\. For each concept, we retain them=100m=100images in which the corresponding bounding box occupies the largest fraction\. Across the resulting1,0001\{,\}000concepts, the median bounding\-box area is30%30\\%of the image\. The samples for each concept are then divided into disjoint fitting and validation sets\.

#### Textual concepts\.

We construct textual concepts in two complementary ways, either stated astoken groups in contextorlabeled data samples with human annotations\. A*lexical*concept is defined by grouping occurrences of one or more related tokens across different contexts\. We construct1818lexical concepts from GSM8K\([Cobbe et al\., 2021](https://arxiv.org/html/2609.05575#bib.bib58)\), covering arithmetic operations, units of measurement, and recurring entities, and1,0261\{,\}026lexical concepts from the Brown\([Francis and Kučera, 1964](https://arxiv.org/html/2609.05575#bib.bib60)\), NLTK Gutenberg\([Bird et al\., 2009](https://arxiv.org/html/2609.05575#bib.bib62)\), and Reuters RCV1\([Lewis et al\., 2004](https://arxiv.org/html/2609.05575#bib.bib61)\)corpora, spanning general\-domain language\. A*labelled*concept is defined by a human\-annotated semantic category\. We construct4343labelled concepts from SemCor\([Miller et al\., 1993](https://arxiv.org/html/2609.05575#bib.bib59)\)\. Each sample is represented by the token sequence cropped at the corresponding token position\.

### 3\.2Held\-Out Coverage and Disentanglement of Capsules

We evaluate held\-out coverage and disentanglement across all model components and concept types introduced above\. For each concept, we use5050samples for capsule matching and a disjoint set of5050samples for held\-out coverage\. Disentanglement is evaluated using balanced positive and negative samples,\|𝒟cval\|=\|𝒟not​\-​cval\|\|\\mathcal\{D\}\_\{c\}^\{\\mathrm\{val\}\}\|=\|\\mathcal\{D\}\_\{\\mathrm\{not\}\\text\{\-\}c\}^\{\\mathrm\{val\}\}\|\. Table[2](https://arxiv.org/html/2609.05575#S3.T2)summarizes the results\.

Fitted capsules achieve consistently high held\-out coverage across models, components, and concept types, with an overall mean of0\.9430\.943\. Disentanglement exhibits substantially greater variation: visual concepts remain close to the balanced baseline across most components, whereas textual concepts are more strongly separated, particularly lexical concepts\.

Table 2:Held\-out coverage and disentanglement across models and components\.We report the mean and standard deviation ofCov\\mathrm\{Cov\}andDis\\mathrm\{Dis\}across fitted capsules\.
### 3\.3Comparison with Alternative Geometric Approximations

We compare the capsule with four alternative geometric approximations—a cone, ball, norm band, and axis\-aligned box—fitted to the same concept representations\. We evaluate each choice using held\-out coverage and disentanglement, together with the geometric information that can be directly extracted from its parameters, including angular structure, norm structure, and whether the two can be separated\. Table[3](https://arxiv.org/html/2609.05575#S3.T3)summarizes these comparisons under the original min/max boundary\.

Capsules achieve comparable held\-out coverage to cones and balls while attaining the highest disentanglement among the geometric approximations with meaningful coverage\. Moreover, the capsule provides a more informative description of concept geometry by separating angular extent from representation norm\. We therefore adopt it as a compact approximation whose parameters directly characterize distinct geometric properties of the concept region\.

Table 3:Comparison with alternative geometric approximations\.Held\-out coverage and disentanglementDis\\mathrm\{Dis\}under the original min/max boundary \(q=100q=100\), averaged across concept types, model components, and representative layers\.∗The box exhibits extremely low held\-out coverage because coordinate\-wise bounds generalize poorly in high\-dimensional anisotropic representation spaces, a manifestation of the curse of dimensionality\([Köppen, 2000](https://arxiv.org/html/2609.05575#bib.bib71)\)\.

## 4Analyzing Concept Geometry in Static Representations

We next use Capsule Lens to analyze concept geometry within static representations\. Section[4\.1](https://arxiv.org/html/2609.05575#S4.SS1)examines the internal structure of concept geometry\. Section[4\.2](https://arxiv.org/html/2609.05575#S4.SS2)shows how Capsule Lens can uncover geometric differences across model layers\.

### 4\.1Internal Structure of Concept Geometry

The capsule parameters describe the boundary of a concept’s representation region, while the span curveFs​\(t\)F\_\{s\}\(t\)and norm curveFn​\(t\)F\_\{n\}\(t\)reveal how representations are distributed inside that boundary\. Figure[2](https://arxiv.org/html/2609.05575#S4.F2)illustrates this distinction for the concept*stripes*at layer 12 of CLIP ViT\-B/32\. Its span curve exhibits two sharp rises separated by a plateau, revealing two angular clusters that cannot be inferred from the fitted spansa=0\.655s\_\{a\}=0\.655alone\. In contrast, the norm curve rises smoothly, indicating a comparatively unimodal distribution in representation magnitude\. Thus, concepts with similar capsule boundaries can exhibit substantially different internal geometric structure, which the coverage curves make explicit\.

![Refer to caption](https://arxiv.org/html/2609.05575v1/Figure2.png)Figure 2:Coverage curves reveal internal concept geometry beyond capsule boundaries\.For*stripes*at layer 12 of CLIP ViT\-B/32, the span distribution andFs​\(t\)F\_\{s\}\(t\)reveal two angular clusters, while the norm distribution andFn​\(t\)F\_\{n\}\(t\)vary smoothly\. The fitted spansas\_\{a\}captures the outer extent of the concept region, whereas the coverage curves reveal how representations populate that region\.
### 4\.2Layer\-wise Evolution of Concept Disentanglement

We examine how concept disentanglement changes across model depth\. For visual concepts, we compute the mean disentanglement scoreDis\\mathrm\{Dis\}at each relative layer depth across CLIP, Qwen2\-VL, and LLaVA components; for textual concepts, we report the corresponding layer\-wise scores for lexical, math, and labelled concepts in Qwen2\.5\-1\.5B\.

Figure[3](https://arxiv.org/html/2609.05575#S4.F3)reveals contrasting depth profiles across modalities\. Several visual encoders show a pronounced increase in disentanglement toward their final layers, consistent with geometric separation phenomena such as neural collapse\([Papyan et al\., 2020](https://arxiv.org/html/2609.05575#bib.bib69)\)\. In contrast, textual concepts become progressively more entangled in later LLM layers, particularly for lexical concepts, consistent with prior findings that lexical identity and lexical semantics are strongest in earlier representations and become less explicit with depth\([Liu et al\., 2024](https://arxiv.org/html/2609.05575#bib.bib63);[Li and Subramani, 2025](https://arxiv.org/html/2609.05575#bib.bib64)\)\.

![Refer to caption](https://arxiv.org/html/2609.05575v1/Figure3.png)Figure 3:Concept disentanglement across model depth\.Left:visual concepts across CLIP, Qwen2\-VL, and LLaVA components, showing pronounced late\-layer increases in disentanglement for several visual representations\.Right:textual concepts in Qwen2\.5\-1\.5B, where disentanglement generally decreases toward later layers, most prominently for lexical concepts\.

## 5Tracking Concept Geometry in Dynamic Representations

We demonstrate how Capsule Lens tracks concept geometry drifts induced by different training pipelines\. Section[5\.1](https://arxiv.org/html/2609.05575#S5.SS1)studies CLIP pretraining, Section[5\.2](https://arxiv.org/html/2609.05575#S5.SS2)examines RL post\-training on visual question answering, and Section[5\.3](https://arxiv.org/html/2609.05575#S5.SS3)analyzes RL post\-training on mathematical reasoning\.

### 5\.1Case Study I: CLIP Pretraining

We first study how concept geometry forms during pretraining from scratch\. We train a standard CLIP model\([Radford et al\., 2021](https://arxiv.org/html/2609.05575#bib.bib53)\)with a symmetric InfoNCE objective on MSCOCO\([Lin et al\., 2014](https://arxiv.org/html/2609.05575#bib.bib54)\), CC3M\([Sharma et al\., 2018](https://arxiv.org/html/2609.05575#bib.bib55)\), and CC12M\([Changpinyo et al\., 2021](https://arxiv.org/html/2609.05575#bib.bib56)\)\. We track the visual concepts of Section[3\.1](https://arxiv.org/html/2609.05575#S3.SS1)at five ViT blocks and six training checkpoints, using epoch1010as the reference because the randomly initialized model has no meaningful concept axes\.

Pretraining substantially restructures concept geometry throughout the network \(Figure[4](https://arxiv.org/html/2609.05575#S5.F4)\)\. Concept axes rotate at every probed block, while span dynamics differ sharply with depth: blocks22–88widen whereas the final block contracts fromsa≈0\.61s\_\{a\}\\approx 0\.61to0\.280\.28\. Norm bands also contract steadily across depth\. Per\-concept drift is comparatively uniform, indicating that pretraining produces broad rather than strongly concept\-selective geometric restructuring\. Appendix[E](https://arxiv.org/html/2609.05575#A5)shows additional analyses\.

![Refer to caption](https://arxiv.org/html/2609.05575v1/Figure4.png)Figure 4:CLIP pretraining restructures concept geometry throughout the network\.a\)Axis spans widen in blocks22–88but contract strongly in the final block\.b\)Concept axes rotate substantially at all probed depth\.c–d\)Norm\-band midpoint and width contract steadily throughout training\.
### 5\.2Case Study II: RL Post\-Training on Visual Question Answering

We next study RL post\-training on VLMs \(Appendix[F](https://arxiv.org/html/2609.05575#A6)\)\. We fine\-tune Qwen2\-VL\([Yang et al\., 2024](https://arxiv.org/html/2609.05575#bib.bib46)\)with PPO and LoRA on OKVQA\([Marino et al\., 2019](https://arxiv.org/html/2609.05575#bib.bib57)\), tracking the visual concepts of Section[3\.1](https://arxiv.org/html/2609.05575#S3.SS1)at five layers and five checkpoints\. Capsules from the base model serve as the reference for all drifts\.

Unlike pretraining, PPO\-induced geometric change is highly localized in depth \(Figure[5](https://arxiv.org/html/2609.05575#S5.F5)\)\. Axis spans remain nearly unchanged, while axis drift concentrates almost entirely in the final layer, reachingΔ​a≈0\.026\\Delta a\\approx 0\.026as earlier layers remain close to their base geometry\. The final\-layer norm band contracts rapidly within the first5050steps and then stabilizes even as its axis continues to rotate\. PPO therefore induces substantially smaller and more localized geometric restructuring than pretraining\.

![Refer to caption](https://arxiv.org/html/2609.05575v1/Figure5.png)Figure 5:PPO post\-training concentrates geometric change in the final layer\.a\)Axis spans remain nearly unchanged\.b\)Axis drift occurs predominantly at layer2727\.c–d\)The final\-layer norm band contracts rapidly and then stabilizes while axis rotation continues\.
### 5\.3Case Study III: RL Post\-Training on Mathematical Reasoning

Finally, we study RLVR for LLM reasoning \(Appendix[G](https://arxiv.org/html/2609.05575#A7)\)\. We train Qwen2\.5\-1\.5B with GRPO\([Shao et al\., 2024](https://arxiv.org/html/2609.05575#bib.bib11)\)on500500GSM8K problems\([Cobbe et al\., 2021](https://arxiv.org/html/2609.05575#bib.bib58)\)for750750optimizer steps, using answer correctness with an additional formatting bonus\. After training, held\-out GSM8K accuracy have moderate improvement from0\.5120\.512to0\.5360\.536\. We track the1818mathematical concepts of Section[3\.1](https://arxiv.org/html/2609.05575#S3.SS1)at five layers and four checkpoints relative to the base model\.

RLVR produces a third geometric pattern: changes are sparse across both concepts and layers\. As shown in Figure[6](https://arxiv.org/html/2609.05575#S5.F6), most concept–layer pairs remain close to their base geometry, while a small subset exhibits pronounced span contractions or norm\-bound shifts\. The clearest example isunit\_time\_minute, whose layer\-2727span contracts fromsa=0\.711s\_\{a\}=0\.711to0\.5940\.594late in training while its axis orientation and norm geometry remain comparatively stable\. Thus, unlike the broad restructuring of pretraining, RLVR selectively concentrates particular concept geometries at particular depths while producing a moderate improvement in held\-out task accuracy\.

![Refer to caption](https://arxiv.org/html/2609.05575v1/Figure6.png)Figure 6:RLVR induces sparse, concept\-specific concentration of representation geometry\.Left:at the final checkpoint, most concept–layer pairs remain close to their base geometry, while a small number exhibit pronounced span contractions or norm\-bound shifts\.Right:for the representative conceptunit\_time\_minute, the final\-layer axis span contracts sharply late in training while norm geometry remain comparatively stable\. These results show that RLVR selectively sharpens particular concept geometries at particular layers rather than comprehensively change the model\.

## 6Conclusion

In this work, we introduce Capsule Lens, a novel framework that can locate and characterize concept geometry in static representations and can track how that geometry changes in dynamic representations\. We validate Capsule Lens systematically and show how it can be used to analyze concept geometry across a range of models\. We further demonstrate, through three case studies spanning distinct training pipelines, how Capsule Lens can track concept geometry as it shifts under training\.

### AI Use Statement

In this work, we used generative AI tools, including ChatGPT, Claude, and Claude Code to implement methods, write and edit code, and assist in drafting the manuscript\. We did not use generative AI tools to generate synthetic data, formulate or prove mathematical claims, or propose our core hypotheses, methodology, or experimental design; these were carried out by the authors\. We have reviewed all AI\-assisted work: code was checked against expected behavior and the underlying data, and all AI\-drafted text and result interpretations were verified against experimental outputs and revised where necessary\. We take responsibility for the final content of this work, including text, claims, and artifacts produced with the aid of generative AI\.

### Ethics Statement

This work uses only publicly available, non\-sensitive data and involves no human\-subjects research, personal data, or dual\-use capability beyond that of the underlying pretrained models we analyze\. We do not foresee ethics concerns specific to this submission\.

### Reproducibility Statement

Capsule Lens is fully specified in closed form: the capsule\-fitting procedure and the associated span and norm coverage curves are detailed in Sec\.[2](https://arxiv.org/html/2609.05575#S2), and the validation protocol \(held\-out coverage, concept\-specificity, and sampling\-stability checks\) is described in Sec\.[2](https://arxiv.org/html/2609.05575#S2)and Sec\.[3](https://arxiv.org/html/2609.05575#S3)\. All concept construction, model choices, and training details for the case studies are provided in the main text and appendices\. Code implementing capsule fitting, validation, and all reported experiments will be released upon publication\.

## References

- Alain and Bengio \(2016\)G\. Alain and Y\. BengioUnderstanding intermediate layers using linear classifier probes\.ArXiv preprintabs/1610\.01644\.External Links:[Link](https://arxiv.org/abs/1610.01644)Cited by:[§A\.1](https://arxiv.org/html/2609.05575#A1.SS1.p1.1)\.
- Arditiet al\.\(2024\)A\. Arditi, O\. Obeso, A\. Syed, D\. Paleka, N\. Panickssery, W\. Gurnee, and N\. NandaRefusal in language models is mediated by a single direction\.ArXiv preprintabs/2406\.11717\.External Links:[Link](https://arxiv.org/abs/2406.11717)Cited by:[§A\.2](https://arxiv.org/html/2609.05575#A1.SS2.p1.1),[Appendix B](https://arxiv.org/html/2609.05575#A2.SS0.SSS0.Px3.p1.1)\.
- Belroseet al\.\(2023\)N\. Belrose, I\. Ostrovsky, L\. McKinney, Z\. Furman, L\. Smith, D\. Halawi, S\. Biderman, and J\. SteinhardtEliciting latent predictions from transformers with the tuned lens\.ArXiv preprintabs/2303\.08112\.External Links:[Link](https://arxiv.org/abs/2303.08112)Cited by:[§A\.1](https://arxiv.org/html/2609.05575#A1.SS1.p1.1),[§1](https://arxiv.org/html/2609.05575#S1.p2.1)\.
- Bereska and Gavves \(2024\)L\. Bereska and E\. GavvesMechanistic interpretability for ai safety – a review\.ArXiv preprintabs/2404\.14082\.External Links:[Link](https://arxiv.org/abs/2404.14082)Cited by:[§1](https://arxiv.org/html/2609.05575#S1.p1.1)\.
- Betleyet al\.\(2026\)J\. Betley, N\. Warncke, A\. Sztyber\-Betley, D\. Tan, X\. Bao, M\. Soto, M\. Srivastava, N\. Labenz, and O\. EvansTraining large language models on narrow tasks can lead to broad misalignment\.Nature649\(8097\),pp\. 584–589\.External Links:ISSN 1476\-4687,[Link](http://dx.doi.org/10.1038/s41586-025-09937-5),[Document](https://dx.doi.org/10.1038/s41586-025-09937-5)Cited by:[§1](https://arxiv.org/html/2609.05575#S1.p1.1)\.
- Bhallaet al\.\(2026\)U\. Bhalla, T\. Fel, C\. Rager, S\. Feucht, T\. Haklay, D\. Wurgaft, S\. Boppana, M\. Kowal, V\. Shyam, J\. Merullo, A\. Geiger, and E\. S\. LubanaDo sparse autoencoders capture concept manifolds?\.ArXiv preprintabs/2604\.28119\.External Links:[Link](https://arxiv.org/abs/2604.28119)Cited by:[Appendix B](https://arxiv.org/html/2609.05575#A2.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.05575#S1.p2.1),[§1](https://arxiv.org/html/2609.05575#S1.p3.1)\.
- Birdet al\.\(2009\)S\. Bird, E\. Klein, and E\. LoperNatural language processing with python: analyzing text with the natural language toolkit\.O’Reilly Media,Sebastopol, CA\.Cited by:[§3\.1](https://arxiv.org/html/2609.05575#S3.SS1.SSS0.Px2.p1.1)\.
- Brownet al\.\(2020\)T\. B\. Brown, B\. Mann, N\. Ryder, M\. Subbiah, J\. Kaplan, P\. Dhariwal, A\. Neelakantan, P\. Shyam, G\. Sastry, A\. Askell, S\. Agarwal, A\. Herbert\-Voss, G\. Krueger, T\. Henighan, R\. Child, A\. Ramesh, D\. M\. Ziegler, J\. Wu, C\. Winter, C\. Hesse, M\. Chen, E\. Sigler, M\. Litwin, S\. Gray, B\. Chess, J\. Clark, C\. Berner, S\. McCandlish, A\. Radford, I\. Sutskever, and D\. AmodeiLanguage models are few\-shot learners\.InAdvances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6\-12, 2020, virtual,H\. Larochelle, M\. Ranzato, R\. Hadsell, M\. Balcan, and H\. Lin \(Eds\.\),External Links:[Link](https://proceedings.neurips.cc/paper/2020/hash/1457c0d6bfcb4967418bfb8ac142f64a-Abstract.html)Cited by:[§1](https://arxiv.org/html/2609.05575#S1.p1.1)\.
- Bussmannet al\.\(2025\)B\. Bussmann, N\. Nabeshima, A\. Karvonen, and N\. NandaLearning multi\-level features with matryoshka sparse autoencoders\.ArXiv preprintabs/2503\.17547\.External Links:[Link](https://arxiv.org/abs/2503.17547)Cited by:[§A\.1](https://arxiv.org/html/2609.05575#A1.SS1.p1.1)\.
- Changpinyoet al\.\(2021\)S\. Changpinyo, P\. Sharma, N\. Ding, and R\. SoricutConceptual 12m: pushing web\-scale image\-text pre\-training to recognize long\-tail visual concepts\.ArXiv preprintabs/2102\.08981\.External Links:[Link](https://arxiv.org/abs/2102.08981)Cited by:[§5\.1](https://arxiv.org/html/2609.05575#S5.SS1.p1.1)\.
- Chuet al\.\(2025\)T\. Chu, Y\. Zhai, J\. Yang, S\. Tong, S\. Xie, D\. Schuurmans, Q\. V\. Le, S\. Levine, and Y\. MaSFT memorizes, rl generalizes: a comparative study of foundation model post\-training\.ArXiv preprintabs/2501\.17161\.External Links:[Link](https://arxiv.org/abs/2501.17161)Cited by:[§A\.3](https://arxiv.org/html/2609.05575#A1.SS3.p1.1),[Appendix B](https://arxiv.org/html/2609.05575#A2.SS0.SSS0.Px1.p1.1),[Appendix B](https://arxiv.org/html/2609.05575#A2.SS0.SSS0.Px4.p1.1)\.
- Chung and Abbott \(2021\)S\. Chung and L\. F\. AbbottNeural population geometry: an approach for understanding biological and artificial neural networks\.Current opinion in neurobiology70,pp\. 137–144\.Cited by:[§1](https://arxiv.org/html/2609.05575#S1.p3.1)\.
- Cobbeet al\.\(2021\)K\. Cobbe, V\. Kosaraju, M\. Bavarian, M\. Chen, H\. Jun, L\. Kaiser, M\. Plappert, J\. Tworek, J\. Hilton, R\. Nakano, C\. Hesse, and J\. SchulmanTraining verifiers to solve math word problems\.External Links:2110\.14168Cited by:[§3\.1](https://arxiv.org/html/2609.05575#S3.SS1.SSS0.Px2.p1.1),[§5\.3](https://arxiv.org/html/2609.05575#S5.SS3.p1.1)\.
- Cunninghamet al\.\(2023\)H\. Cunningham, A\. Ewart, L\. Riggs, R\. Huben, and L\. SharkeySparse autoencoders find highly interpretable features in language models\.ArXiv preprintabs/2309\.08600\.External Links:[Link](https://arxiv.org/abs/2309.08600)Cited by:[§A\.1](https://arxiv.org/html/2609.05575#A1.SS1.p1.1),[§1](https://arxiv.org/html/2609.05575#S1.p2.1)\.
- Donget al\.\(2025\)Y\. Dong, X\. Jiang, Y\. Tao, H\. Liu, K\. Zhang, L\. Mou, R\. Cao, Y\. Ma, J\. Chen, B\. Li, Z\. Jin, F\. Huang, Y\. Li, and G\. LiRL\-plus: countering capability boundary collapse of llms in reinforcement learning with hybrid\-policy optimization\.ArXiv preprintabs/2508\.00222\.External Links:[Link](https://arxiv.org/abs/2508.00222)Cited by:[§A\.3](https://arxiv.org/html/2609.05575#A1.SS3.p1.1),[§1](https://arxiv.org/html/2609.05575#S1.p1.1)\.
- Duanet al\.\(2026\)X\. Duan, Z\. Yao, X\. Zhou, Z\. Fu, B\. Kan, Y\. Wang, B\. Xiao, H\. Zhang, X\. Tang, E\. Nie, and Z\. G\. CaiMechanistic interpretability for understanding language abilities in large language models: a survey\.Preprints\.External Links:[Document](https://dx.doi.org/10.20944/preprints202607.1441.v1),[Link](https://doi.org/10.20944/preprints202607.1441.v1)Cited by:[Appendix B](https://arxiv.org/html/2609.05575#A2.SS0.SSS0.Px1.p1.1)\.
- Elhageet al\.\(2022\)N\. Elhage, T\. Hume, C\. Olsson, N\. Schiefer, T\. Henighan, S\. Kravec, Z\. Hatfield\-Dodds, R\. Lasenby, D\. Drain, C\. Chen, R\. Grosse, S\. McCandlish, J\. Kaplan, D\. Amodei, M\. Wattenberg, and C\. OlahToy models of superposition\.ArXiv preprintabs/2209\.10652\.External Links:[Link](https://arxiv.org/abs/2209.10652)Cited by:[§A\.2](https://arxiv.org/html/2609.05575#A1.SS2.p1.1),[§1](https://arxiv.org/html/2609.05575#S1.p2.1)\.
- Engelset al\.\(2024\)J\. Engels, E\. J\. Michaud, I\. Liao, W\. Gurnee, and M\. TegmarkNot all language model features are one\-dimensionally linear\.ArXiv preprintabs/2405\.14860\.External Links:[Link](https://arxiv.org/abs/2405.14860)Cited by:[§A\.2](https://arxiv.org/html/2609.05575#A1.SS2.p1.1),[Appendix B](https://arxiv.org/html/2609.05575#A2.SS0.SSS0.Px1.p1.1),[Appendix B](https://arxiv.org/html/2609.05575#A2.SS0.SSS0.Px2.p1.1),[Appendix B](https://arxiv.org/html/2609.05575#A2.SS0.SSS0.Px3.p1.1),[§1](https://arxiv.org/html/2609.05575#S1.p2.1),[§1](https://arxiv.org/html/2609.05575#S1.p3.1)\.
- Francis and Kučera \(1964\)W\. N\. Francis and H\. KučeraManual of information to accompany a standard corpus of present\-day edited american english, for use with digital computers\.Technical reportDepartment of Linguistics, Brown University,Providence, Rhode Island\.Cited by:[§3\.1](https://arxiv.org/html/2609.05575#S3.SS1.SSS0.Px2.p1.1)\.
- Fraser\-Talienteet al\.\(2026\)K\. Fraser\-Taliente, S\. Kantamneni, E\. Ong, D\. Mossing, C\. Lu, P\. C\. Bogdan, E\. Ameisen, J\. Chen, D\. Kishylau, A\. Pearce,et al\.Natural language autoencoders produce unsupervised explanations of llm activations\.Transformer Circuits Thread\.Cited by:[§A\.1](https://arxiv.org/html/2609.05575#A1.SS1.p1.1)\.
- Gaoet al\.\(2024\)L\. Gao, T\. D\. la Tour, H\. Tillman, G\. Goh, R\. Troll, A\. Radford, I\. Sutskever, J\. Leike, and J\. WuScaling and evaluating sparse autoencoders\.ArXiv preprintabs/2406\.04093\.External Links:[Link](https://arxiv.org/abs/2406.04093)Cited by:[§A\.1](https://arxiv.org/html/2609.05575#A1.SS1.p1.1),[§1](https://arxiv.org/html/2609.05575#S1.p2.1)\.
- Ghandehariounet al\.\(2024\)A\. Ghandeharioun, A\. Caciularu, A\. Pearce, L\. Dixon, and M\. GevaPatchscopes: a unifying framework for inspecting hidden representations of language models\.ArXiv preprintabs/2401\.06102\.External Links:[Link](https://arxiv.org/abs/2401.06102)Cited by:[§A\.1](https://arxiv.org/html/2609.05575#A1.SS1.p1.1),[§1](https://arxiv.org/html/2609.05575#S1.p2.1)\.
- Grattafioriet al\.\(2024\)A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan, A\. Yang, A\. Fan, A\. Goyal, A\. Hartshorn, A\. Yang, A\. Mitra, A\. Sravankumar, A\. Korenev, A\. Hinsvark, A\. Rao, A\. Zhang, A\. Rodriguez, A\. Gregerson, A\. Spataru, B\. Roziere, B\. Biron, B\. Tang, B\. Chern, C\. Caucheteux, C\. Nayak, C\. Bi, C\. Marra, C\. McConnell, C\. Keller, C\. Touret, C\. Wu, C\. Wong, C\. C\. Ferrer, C\. Nikolaidis, D\. Allonsius, D\. Song, D\. Pintz, D\. Livshits, D\. Wyatt, D\. Esiobu, D\. Choudhary, D\. Mahajan, D\. Garcia\-Olano, D\. Perino, D\. Hupkes, E\. Lakomkin, E\. AlBadawy, E\. Lobanova, E\. Dinan, E\. M\. Smith, F\. Radenovic, F\. Guzmán, F\. Zhang, G\. Synnaeve, G\. Lee, G\. L\. Anderson, G\. Thattai, G\. Nail, G\. Mialon, G\. Pang, G\. Cucurell, H\. Nguyen, H\. Korevaar, H\. Xu, H\. Touvron, I\. Zarov, I\. A\. Ibarra, I\. Kloumann, I\. Misra, I\. Evtimov, J\. Zhang, J\. Copet, J\. Lee, J\. Geffert, J\. Vranes, J\. Park, J\. Mahadeokar, J\. Shah, J\. van der Linde, J\. Billock, J\. Hong, J\. Lee, J\. Fu, J\. Chi, J\. Huang, J\. Liu, J\. Wang, J\. Yu, J\. Bitton, J\. Spisak, J\. Park, J\. Rocca, J\. Johnstun, J\. Saxe, J\. Jia, K\. V\. Alwala, K\. Prasad, K\. Upasani, K\. Plawiak, K\. Li, K\. Heafield, K\. Stone, K\. El\-Arini, K\. Iyer, K\. Malik, K\. Chiu, K\. Bhalla, K\. Lakhotia, L\. Rantala\-Yeary, L\. van der Maaten, L\. Chen, L\. Tan, L\. Jenkins, L\. Martin, L\. Madaan, L\. Malo, L\. Blecher, L\. Landzaat, L\. de Oliveira, M\. Muzzi, M\. Pasupuleti, M\. Singh, M\. Paluri, M\. Kardas, M\. Tsimpoukelli, M\. Oldham, M\. Rita, M\. Pavlova, M\. Kambadur, M\. Lewis, M\. Si, M\. K\. Singh, M\. Hassan, N\. Goyal, N\. Torabi, N\. Bashlykov, N\. Bogoychev, N\. Chatterji, N\. Zhang, O\. Duchenne, O\. Çelebi, P\. Alrassy, P\. Zhang, P\. Li, P\. Vasic, P\. Weng, P\. Bhargava, P\. Dubal, P\. Krishnan, P\. S\. Koura, P\. Xu, Q\. He, Q\. Dong, R\. Srinivasan, R\. Ganapathy, R\. Calderer, R\. S\. Cabral, R\. Stojnic, R\. Raileanu, R\. Maheswari, R\. Girdhar, R\. Patel, R\. Sauvestre, R\. Polidoro, R\. Sumbaly, R\. Taylor, R\. Silva, R\. Hou, R\. Wang, S\. Hosseini, S\. Chennabasappa, S\. Singh, S\. Bell, S\. S\. Kim, S\. Edunov, S\. Nie, S\. Narang, S\. Raparthy, S\. Shen, S\. Wan, S\. Bhosale, S\. Zhang, S\. Vandenhende, S\. Batra, S\. Whitman, S\. Sootla, S\. Collot, S\. Gururangan, S\. Borodinsky, T\. Herman, T\. Fowler, T\. Sheasha, T\. Georgiou, T\. Scialom, T\. Speckbacher, T\. Mihaylov, T\. Xiao, U\. Karn, V\. Goswami, V\. Gupta, V\. Ramanathan, V\. Kerkez, V\. Gonguet, V\. Do, V\. Vogeti, V\. Albiero, V\. Petrovic, W\. Chu, W\. Xiong, W\. Fu, W\. Meers, X\. Martinet, X\. Wang, X\. Wang, X\. E\. Tan, X\. Xia, X\. Xie, X\. Jia, X\. Wang, Y\. Goldschlag, Y\. Gaur, Y\. Babaei, Y\. Wen, Y\. Song, Y\. Zhang, Y\. Li, Y\. Mao, Z\. D\. Coudert, Z\. Yan, Z\. Chen, Z\. Papakipos, A\. Singh, A\. Srivastava, A\. Jain, A\. Kelsey, A\. Shajnfeld, A\. Gangidi, A\. Victoria, A\. Goldstand, A\. Menon, A\. Sharma, A\. Boesenberg, A\. Baevski, A\. Feinstein, A\. Kallet, A\. Sangani, A\. Teo, A\. Yunus, A\. Lupu, A\. Alvarado, A\. Caples, A\. Gu, A\. Ho, A\. Poulton, A\. Ryan, A\. Ramchandani, A\. Dong, A\. Franco, A\. Goyal, A\. Saraf, A\. Chowdhury, A\. Gabriel, A\. Bharambe, A\. Eisenman, A\. Yazdan, B\. James, B\. Maurer, B\. Leonhardi, B\. Huang, B\. Loyd, B\. D\. Paola, B\. Paranjape, B\. Liu, B\. Wu, B\. Ni, B\. Hancock, B\. Wasti, B\. Spence, B\. Stojkovic, B\. Gamido, B\. Montalvo, C\. Parker, C\. Burton, C\. Mejia, C\. Liu, C\. Wang, C\. Kim, C\. Zhou, C\. Hu, C\. Chu, C\. Cai, C\. Tindal, C\. Feichtenhofer, C\. Gao, D\. Civin, D\. Beaty, D\. Kreymer, D\. Li, D\. Adkins, D\. Xu, D\. Testuggine, D\. David, D\. Parikh, D\. Liskovich, D\. Foss, D\. Wang, D\. Le, D\. Holland, E\. Dowling, E\. Jamil, E\. Montgomery, E\. Presani, E\. Hahn, E\. Wood, E\. Le, E\. Brinkman, E\. Arcaute, E\. Dunbar, E\. Smothers, F\. Sun, F\. Kreuk, F\. Tian, F\. Kokkinos, F\. Ozgenel, F\. Caggioni, F\. Kanayet, F\. Seide, G\. M\. Florez, G\. Schwarz, G\. Badeer, G\. Swee, G\. Halpern, G\. Herman, G\. Sizov, Guangyi, Zhang, G\. Lakshminarayanan, H\. Inan, H\. Shojanazeri, H\. Zou, H\. Wang, H\. Zha, H\. Habeeb, H\. Rudolph, H\. Suk, H\. Aspegren, H\. Goldman, H\. Zhan, I\. Damlaj, I\. Molybog, I\. Tufanov, I\. Leontiadis, I\. Veliche, I\. Gat, J\. Weissman, J\. Geboski, J\. Kohli, J\. Lam, J\. Asher, J\. Gaya, J\. Marcus, J\. Tang, J\. Chan, J\. Zhen, J\. Reizenstein, J\. Teboul, J\. Zhong, J\. Jin, J\. Yang, J\. Cummings, J\. Carvill, J\. Shepard, J\. McPhie, J\. Torres, J\. Ginsburg, J\. Wang, K\. Wu, K\. H\. U, K\. Saxena, K\. Khandelwal, K\. Zand, K\. Matosich, K\. Veeraraghavan, K\. Michelena, K\. Li, K\. Jagadeesh, K\. Huang, K\. Chawla, K\. Huang, L\. Chen, L\. Garg, L\. A, L\. Silva, L\. Bell, L\. Zhang, L\. Guo, L\. Yu, L\. Moshkovich, L\. Wehrstedt, M\. Khabsa, M\. Avalani, M\. Bhatt, M\. Mankus, M\. Hasson, M\. Lennie, M\. Reso, M\. Groshev, M\. Naumov, M\. Lathi, M\. Keneally, M\. Liu, M\. L\. Seltzer, M\. Valko, M\. Restrepo, M\. Patel, M\. Vyatskov, M\. Samvelyan, M\. Clark, M\. Macey, M\. Wang, M\. J\. Hermoso, M\. Metanat, M\. Rastegari, M\. Bansal, N\. Santhanam, N\. Parks, N\. White, N\. Bawa, N\. Singhal, N\. Egebo, N\. Usunier, N\. Mehta, N\. P\. Laptev, N\. Dong, N\. Cheng, O\. Chernoguz, O\. Hart, O\. Salpekar, O\. Kalinli, P\. Kent, P\. Parekh, P\. Saab, P\. Balaji, P\. Rittner, P\. Bontrager, P\. Roux, P\. Dollar, P\. Zvyagina, P\. Ratanchandani, P\. Yuvraj, Q\. Liang, R\. Alao, R\. Rodriguez, R\. Ayub, R\. Murthy, R\. Nayani, R\. Mitra, R\. Parthasarathy, R\. Li, R\. Hogan, R\. Battey, R\. Wang, R\. Howes, R\. Rinott, S\. Mehta, S\. Siby, S\. J\. Bondu, S\. Datta, S\. Chugh, S\. Hunt, S\. Dhillon, S\. Sidorov, S\. Pan, S\. Mahajan, S\. Verma, S\. Yamamoto, S\. Ramaswamy, S\. Lindsay, S\. Lindsay, S\. Feng, S\. Lin, S\. C\. Zha, S\. Patil, S\. Shankar, S\. Zhang, S\. Zhang, S\. Wang, S\. Agarwal, S\. Sajuyigbe, S\. Chintala, S\. Max, S\. Chen, S\. Kehoe, S\. Satterfield, S\. Govindaprasad, S\. Gupta, S\. Deng, S\. Cho, S\. Virk, S\. Subramanian, S\. Choudhury, S\. Goldman, T\. Remez, T\. Glaser, T\. Best, T\. Koehler, T\. Robinson, T\. Li, T\. Zhang, T\. Matthews, T\. Chou, T\. Shaked, V\. Vontimitta, V\. Ajayi, V\. Montanez, V\. Mohan, V\. S\. Kumar, V\. Mangla, V\. Ionescu, V\. Poenaru, V\. T\. Mihailescu, V\. Ivanov, W\. Li, W\. Wang, W\. Jiang, W\. Bouaziz, W\. Constable, X\. Tang, X\. Wu, X\. Wang, X\. Wu, X\. Gao, Y\. Kleinman, Y\. Chen, Y\. Hu, Y\. Jia, Y\. Qi, Y\. Li, Y\. Zhang, Y\. Zhang, Y\. Adi, Y\. Nam, Yu, Wang, Y\. Zhao, Y\. Hao, Y\. Qian, Y\. Li, Y\. He, Z\. Rait, Z\. DeVito, Z\. Rosnbrick, Z\. Wen, Z\. Yang, Z\. Zhao, and Z\. MaThe llama 3 herd of models\.ArXiv preprintabs/2407\.21783\.External Links:[Link](https://arxiv.org/abs/2407.21783)Cited by:[§1](https://arxiv.org/html/2609.05575#S1.p1.1)\.
- Guoet al\.\(2025\)D\. Guo, D\. Yang, H\. Zhang, J\. Song, P\. Wang, Q\. Zhu, R\. Xu, R\. Zhang, S\. Ma, X\. Bi, X\. Zhang, X\. Yu, Y\. Wu, Z\. F\. Wu, Z\. Gou, Z\. Shao, Z\. Li, Z\. Gao, A\. Liu, B\. Xue, B\. Wang, B\. Wu, B\. Feng, C\. Lu, C\. Zhao, C\. Deng, C\. Ruan, D\. Dai, D\. Chen, D\. Ji, E\. Li, F\. Lin, F\. Dai, F\. Luo, G\. Hao, G\. Chen, G\. Li, H\. Zhang, H\. Xu, H\. Ding, H\. Gao, H\. Qu, H\. Li, J\. Guo, J\. Li, J\. Chen, J\. Yuan, J\. Tu, J\. Qiu, J\. Li, J\. L\. Cai, J\. Ni, J\. Liang, J\. Chen, K\. Dong, K\. Hu, K\. You, K\. Gao, K\. Guan, K\. Huang, K\. Yu, L\. Wang, L\. Zhang, L\. Zhao, L\. Wang, L\. Zhang, L\. Xu, L\. Xia, M\. Zhang, M\. Zhang, M\. Tang, M\. Zhou, M\. Li, M\. Wang, M\. Li, N\. Tian, P\. Huang, P\. Zhang, Q\. Wang, Q\. Chen, Q\. Du, R\. Ge, R\. Zhang, R\. Pan, R\. Wang, R\. J\. Chen, R\. L\. Jin, R\. Chen, S\. Lu, S\. Zhou, S\. Chen, S\. Ye, S\. Wang, S\. Yu, S\. Zhou, S\. Pan, S\. S\. Li, S\. Zhou, S\. Wu, T\. Yun, T\. Pei, T\. Sun, T\. Wang, W\. Zeng, W\. Liu, W\. Liang, W\. Gao, W\. Yu, W\. Zhang, W\. L\. Xiao, W\. An, X\. Liu, X\. Wang, X\. Chen, X\. Nie, X\. Cheng, X\. Liu, X\. Xie, X\. Liu, X\. Yang, X\. Li, X\. Su, X\. Lin, X\. Q\. Li, X\. Jin, X\. Shen, X\. Chen, X\. Sun, X\. Wang, X\. Song, X\. Zhou, X\. Wang, X\. Shan, Y\. K\. Li, Y\. Q\. Wang, Y\. X\. Wei, Y\. Zhang, Y\. Xu, Y\. Li, Y\. Zhao, Y\. Sun, Y\. Wang, Y\. Yu, Y\. Zhang, Y\. Shi, Y\. Xiong, Y\. He, Y\. Piao, Y\. Wang, Y\. Tan, Y\. Ma, Y\. Liu, Y\. Guo, Y\. Ou, Y\. Wang, Y\. Gong, Y\. Zou, Y\. He, Y\. Xiong, Y\. Luo, Y\. You, Y\. Liu, Y\. Zhou, Y\. X\. Zhu, Y\. Huang, Y\. Li, Y\. Zheng, Y\. Zhu, Y\. Ma, Y\. Tang, Y\. Zha, Y\. Yan, Z\. Z\. Ren, Z\. Ren, Z\. Sha, Z\. Fu, Z\. Xu, Z\. Xie, Z\. Zhang, Z\. Hao, Z\. Ma, Z\. Yan, Z\. Wu, Z\. Gu, Z\. Zhu, Z\. Liu, Z\. Li, Z\. Xie, Z\. Song, Z\. Pan, Z\. Huang, Z\. Xu, Z\. Zhang, and Z\. ZhangDeepSeek\-r1 incentivizes reasoning in llms through reinforcement learning\.Nature645\(8081\),pp\. 633–638\.External Links:ISSN 1476\-4687,[Link](http://dx.doi.org/10.1038/s41586-025-09422-z),[Document](https://dx.doi.org/10.1038/s41586-025-09422-z)Cited by:[§1](https://arxiv.org/html/2609.05575#S1.p1.1)\.
- Gurnee and Tegmark \(2024\)W\. Gurnee and M\. TegmarkLanguage models represent space and time\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 2483–2503\.Cited by:[§A\.2](https://arxiv.org/html/2609.05575#A1.SS2.p1.1)\.
- Hewitt and Manning \(2019\)J\. Hewitt and C\. D\. ManningA structural probe for finding syntax in word representations\.InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 \(Long and Short Papers\),Minneapolis, Minnesota,pp\. 4129–4138\.External Links:[Document](https://dx.doi.org/10.18653/v1/N19-1419),[Link](https://aclanthology.org/N19-1419)Cited by:[§A\.1](https://arxiv.org/html/2609.05575#A1.SS1.p1.1)\.
- Huhet al\.\(2024\)M\. Huh, B\. Cheung, T\. Wang, and P\. IsolaThe platonic representation hypothesis\.ArXiv preprintabs/2405\.07987\.External Links:[Link](https://arxiv.org/abs/2405.07987)Cited by:[§A\.2](https://arxiv.org/html/2609.05575#A1.SS2.p1.1)\.
- Jiralerspong and Bricken \(2026\)T\. Jiralerspong and T\. BrickenCross\-architecture model diffing with crosscoders: unsupervised discovery of differences between llms\.ArXiv preprintabs/2602\.11729\.External Links:[Link](https://arxiv.org/abs/2602.11729)Cited by:[§A\.1](https://arxiv.org/html/2609.05575#A1.SS1.p1.1),[Appendix B](https://arxiv.org/html/2609.05575#A2.SS0.SSS0.Px1.p1.1)\.
- Köppen \(2000\)M\. KöppenThe curse of dimensionality\.In5th online world conference on soft computing in industrial applications \(WSC5\),Vol\.1,pp\. 4–8\.Cited by:[Table 3](https://arxiv.org/html/2609.05575#S3.T3)\.
- Korchinskiet al\.\(2025\)D\. J\. Korchinski, D\. Karkada, Y\. Bahri, and M\. WyartOn the emergence of linear analogies in word embeddings\.ArXiv preprintabs/2505\.18651\.External Links:[Link](https://arxiv.org/abs/2505.18651)Cited by:[§A\.2](https://arxiv.org/html/2609.05575#A1.SS2.p1.1)\.
- Krishnaet al\.\(2016\)R\. Krishna, Y\. Zhu, O\. Groth, J\. Johnson, K\. Hata, J\. Kravitz, S\. Chen, Y\. Kalantidis, L\. Li, D\. A\. Shamma, M\. S\. Bernstein, and F\. LiVisual genome: connecting language and vision using crowdsourced dense image annotations\.ArXiv preprintabs/1602\.07332\.External Links:[Link](https://arxiv.org/abs/1602.07332)Cited by:[§3\.1](https://arxiv.org/html/2609.05575#S3.SS1.SSS0.Px1.p1.1)\.
- Leeet al\.\(2023\)H\. Lee, S\. Phatale, H\. Mansoor, T\. Mesnard, J\. Ferret, K\. Lu, C\. Bishop, E\. Hall, V\. Carbune, A\. Rastogi, and S\. PrakashRLAIF vs\. rlhf: scaling reinforcement learning from human feedback with ai feedback\.ArXiv preprintabs/2309\.00267\.External Links:[Link](https://arxiv.org/abs/2309.00267)Cited by:[§A\.3](https://arxiv.org/html/2609.05575#A1.SS3.p1.1)\.
- Lewiset al\.\(2004\)D\. D\. Lewis, Y\. Yang, T\. G\. Rose, and F\. LiRCV1: a new benchmark collection for text categorization research\.Journal of Machine Learning Research5,pp\. 361–397\.Cited by:[§3\.1](https://arxiv.org/html/2609.05575#S3.SS1.SSS0.Px2.p1.1)\.
- Liet al\.\(2022\)K\. Li, A\. K\. Hopkins, D\. Bau, F\. Viégas, H\. Pfister, and M\. WattenbergEmergent world representations: exploring a sequence model trained on a synthetic task\.ArXiv preprintabs/2210\.13382\.External Links:[Link](https://arxiv.org/abs/2210.13382)Cited by:[§A\.2](https://arxiv.org/html/2609.05575#A1.SS2.p1.1)\.
- Li and Subramani \(2025\)M\. Li and N\. SubramaniModel internal sleuthing: finding lexical identity and inflectional features in modern language models\.ArXiv preprintabs/2506\.02132\.External Links:[Link](https://arxiv.org/abs/2506.02132)Cited by:[§4\.2](https://arxiv.org/html/2609.05575#S4.SS2.p2.1)\.
- Liet al\.\(2025\)Y\. Li, E\. J\. Michaud, D\. D\. Baek, J\. Engels, X\. Sun, and M\. TegmarkThe geometry of concepts: sparse autoencoder feature structure\.Entropy27\(4\),pp\. 344\.External Links:ISSN 1099\-4300,[Link](http://dx.doi.org/10.3390/e27040344),[Document](https://dx.doi.org/10.3390/e27040344)Cited by:[§A\.2](https://arxiv.org/html/2609.05575#A1.SS2.p1.1)\.
- Linet al\.\(2014\)T\. Lin, M\. Maire, S\. Belongie, L\. Bourdev, R\. Girshick, J\. Hays, P\. Perona, D\. Ramanan, C\. L\. Zitnick, and P\. DollárMicrosoft coco: common objects in context\.ArXiv preprintabs/1405\.0312\.External Links:[Link](https://arxiv.org/abs/1405.0312)Cited by:[§5\.1](https://arxiv.org/html/2609.05575#S5.SS1.p1.1)\.
- Lindseyet al\.\(2024\)J\. Lindsey, A\. Templeton, J\. Marcus, T\. Conerly, J\. Batson, and C\. OlahSparse crosscoders for cross\-layer features and model diffing\.Transformer Circuits Thread,pp\. 3982–3992\.Cited by:[§A\.1](https://arxiv.org/html/2609.05575#A1.SS1.p1.1),[Appendix B](https://arxiv.org/html/2609.05575#A2.SS0.SSS0.Px1.p1.1)\.
- Liuet al\.\(2025\)M\. Liu, S\. Diao, X\. Lu, J\. Hu, X\. Dong, Y\. Choi, J\. Kautz, and Y\. DongProRL: prolonged reinforcement learning expands reasoning boundaries in large language models\.ArXiv preprintabs/2505\.24864\.External Links:[Link](https://arxiv.org/abs/2505.24864)Cited by:[§A\.3](https://arxiv.org/html/2609.05575#A1.SS3.p1.1),[Appendix B](https://arxiv.org/html/2609.05575#A2.SS0.SSS0.Px4.p1.1)\.
- Liuet al\.\(2024\)Z\. Liu, C\. Kong, Y\. Liu, and M\. SunFantastic semantics and where to find them: investigating which layers of generative llms reflect lexical semantics\.ArXiv preprintabs/2403\.01509\.External Links:[Link](https://arxiv.org/abs/2403.01509)Cited by:[§4\.2](https://arxiv.org/html/2609.05575#S4.SS2.p2.1)\.
- Malhotraet al\.\(2026\)N\. Malhotra, J\. Ambadkar, A\. Gupta, K\. Kasivel, A\. Schwarz, K\. Ferry, and A\. MonodTracking representation dynamics in large language models with persistent homology\.ArXiv preprintabs/2606\.19542\.External Links:[Link](https://arxiv.org/abs/2606.19542)Cited by:[Appendix B](https://arxiv.org/html/2609.05575#A2.SS0.SSS0.Px3.p1.1),[§2\.2](https://arxiv.org/html/2609.05575#S2.SS2.p1.1)\.
- Marinoet al\.\(2019\)K\. Marino, M\. Rastegari, A\. Farhadi, and R\. MottaghiOK\-VQA: A visual question answering benchmark requiring external knowledge\.InIEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16\-20, 2019,pp\. 3195–3204\.External Links:[Document](https://dx.doi.org/10.1109/CVPR.2019.00331),[Link](http://openaccess.thecvf.com/content/_CVPR/_2019/html/Marino/_OK-VQA/_A/_Visual/_Question/_Answering/_Benchmark/_Requiring/_External/_Knowledge/_CVPR/_2019/_paper.html)Cited by:[§5\.2](https://arxiv.org/html/2609.05575#S5.SS2.p1.1)\.
- Milleret al\.\(1993\)G\. A\. Miller, C\. Leacock, R\. Tengi, and R\. T\. BunkerA semantic concordance\.InHuman Language Technology: Proceedings of a Workshop Held at Plainsboro, New Jersey, March 21\-24, 1993,External Links:[Link](https://aclanthology.org/H93-1061)Cited by:[§3\.1](https://arxiv.org/html/2609.05575#S3.SS1.SSS0.Px2.p1.1)\.
- Modellet al\.\(2025\)A\. Modell, P\. Rubin\-Delanchy, and N\. WhiteleyThe origins of representation manifolds in large language models\.ArXiv preprintabs/2505\.18235\.External Links:[Link](https://arxiv.org/abs/2505.18235)Cited by:[Appendix B](https://arxiv.org/html/2609.05575#A2.SS0.SSS0.Px2.p1.1),[Appendix B](https://arxiv.org/html/2609.05575#A2.SS0.SSS0.Px3.p1.1),[§1](https://arxiv.org/html/2609.05575#S1.p3.1)\.
- nostalgebraist \(2020\)nostalgebraistInterpreting gpt: the logit lens\.Note:[https://www\.lesswrong\.com/posts/AcKRB8wDpdaN6v6ru/interpreting\-gpt\-the\-logit\-lens](https://www.lesswrong.com/posts/AcKRB8wDpdaN6v6ru/interpreting-gpt-the-logit-lens)LessWrong blog postCited by:[§A\.1](https://arxiv.org/html/2609.05575#A1.SS1.p1.1),[§1](https://arxiv.org/html/2609.05575#S1.p2.1)\.
- Ouyanget al\.\(2022\)L\. Ouyang, J\. Wu, X\. Jiang, D\. Almeida, C\. L\. Wainwright, P\. Mishkin, C\. Zhang, S\. Agarwal, K\. Slama, A\. Ray, J\. Schulman, J\. Hilton, F\. Kelton, L\. Miller, M\. Simens, A\. Askell, P\. Welinder, P\. Christiano, J\. Leike, and R\. LoweTraining language models to follow instructions with human feedback\.ArXiv preprintabs/2203\.02155\.External Links:[Link](https://arxiv.org/abs/2203.02155)Cited by:[§A\.3](https://arxiv.org/html/2609.05575#A1.SS3.p1.1),[§1](https://arxiv.org/html/2609.05575#S1.p1.1)\.
- Papyanet al\.\(2020\)V\. Papyan, X\. Y\. Han, and D\. L\. DonohoPrevalence of neural collapse during the terminal phase of deep learning training\.Proceedings of the National Academy of Sciences117\(40\),pp\. 24652–24663\.External Links:ISSN 1091\-6490,[Link](http://dx.doi.org/10.1073/pnas.2015509117),[Document](https://dx.doi.org/10.1073/pnas.2015509117)Cited by:[§A\.2](https://arxiv.org/html/2609.05575#A1.SS2.p1.1),[§4\.2](https://arxiv.org/html/2609.05575#S4.SS2.p2.1)\.
- Parket al\.\(2024\)K\. Park, Y\. J\. Choe, Y\. Jiang, and V\. VeitchThe geometry of categorical and hierarchical concepts in large language models\.ArXiv preprintabs/2406\.01506\.External Links:[Link](https://arxiv.org/abs/2406.01506)Cited by:[§A\.2](https://arxiv.org/html/2609.05575#A1.SS2.p1.1),[Appendix B](https://arxiv.org/html/2609.05575#A2.SS0.SSS0.Px3.p1.1),[§1](https://arxiv.org/html/2609.05575#S1.p2.1)\.
- Parket al\.\(2023\)K\. Park, Y\. J\. Choe, and V\. VeitchThe linear representation hypothesis and the geometry of large language models\.ArXiv preprintabs/2311\.03658\.External Links:[Link](https://arxiv.org/abs/2311.03658)Cited by:[§A\.2](https://arxiv.org/html/2609.05575#A1.SS2.p1.1),[Appendix B](https://arxiv.org/html/2609.05575#A2.SS0.SSS0.Px1.p1.1),[Appendix B](https://arxiv.org/html/2609.05575#A2.SS0.SSS0.Px3.p1.1),[§1](https://arxiv.org/html/2609.05575#S1.p2.1)\.
- Radfordet al\.\(2021\)A\. Radford, J\. W\. Kim, C\. Hallacy, A\. Ramesh, G\. Goh, S\. Agarwal, G\. Sastry, A\. Askell, P\. Mishkin, J\. Clark, G\. Krueger, and I\. SutskeverLearning transferable visual models from natural language supervision\.ArXiv preprintabs/2103\.00020\.External Links:[Link](https://arxiv.org/abs/2103.00020)Cited by:[§5\.1](https://arxiv.org/html/2609.05575#S5.SS1.p1.1)\.
- Renet al\.\(2026\)Q\. Ren, P\. Wang, R\. Cai, S\. Shao, D\. Guo, Y\. Xie, Y\. Li, Q\. Zhang, X\. Hu, J\. Shao, and D\. LiuRethinking generalization in reasoning sft: a conditional analysis on optimization, data, and model capability\.ArXiv preprintabs/2604\.06628\.External Links:[Link](https://arxiv.org/abs/2604.06628)Cited by:[§A\.3](https://arxiv.org/html/2609.05575#A1.SS3.p1.1)\.
- Shaoet al\.\(2025\)R\. Shao, S\. S\. Li, R\. Xin, S\. Geng, Y\. Wang, S\. Oh, S\. S\. Du, N\. Lambert, S\. Min, R\. Krishna, Y\. Tsvetkov, H\. Hajishirzi, P\. W\. Koh, and L\. ZettlemoyerSpurious rewards: rethinking training signals in rlvr\.ArXiv preprintabs/2506\.10947\.External Links:[Link](https://arxiv.org/abs/2506.10947)Cited by:[§A\.3](https://arxiv.org/html/2609.05575#A1.SS3.p1.1)\.
- Shaoet al\.\(2024\)Z\. Shao, P\. Wang, Q\. Zhu, R\. Xu, J\. Song, X\. Bi, H\. Zhang, M\. Zhang, Y\. K\. Li, Y\. Wu, and D\. GuoDeepSeekMath: pushing the limits of mathematical reasoning in open language models\.ArXiv preprintabs/2402\.03300\.External Links:[Link](https://arxiv.org/abs/2402.03300)Cited by:[§A\.3](https://arxiv.org/html/2609.05575#A1.SS3.p1.1),[§1](https://arxiv.org/html/2609.05575#S1.p1.1),[§5\.3](https://arxiv.org/html/2609.05575#S5.SS3.p1.1)\.
- Sharkeyet al\.\(2025\)L\. Sharkey, B\. Chughtai, J\. Batson, J\. Lindsey, J\. Wu, L\. Bushnaq, N\. Goldowsky\-Dill, S\. Heimersheim, A\. Ortega, J\. Bloom,et al\.Open problems in mechanistic interpretability\.ArXiv preprintabs/2501\.16496\.External Links:[Link](https://arxiv.org/abs/2501.16496)Cited by:[§1](https://arxiv.org/html/2609.05575#S1.p1.1)\.
- Sharmaet al\.\(2018\)P\. Sharma, N\. Ding, S\. Goodman, and R\. SoricutConceptual captions: a cleaned, hypernymed, image alt\-text dataset for automatic image captioning\.InProceedings of the 56th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),Melbourne, Australia,pp\. 2556–2565\.External Links:[Document](https://dx.doi.org/10.18653/v1/P18-1238),[Link](https://aclanthology.org/P18-1238)Cited by:[§5\.1](https://arxiv.org/html/2609.05575#S5.SS1.p1.1)\.
- Shenet al\.\(2025\)M\. Shen, Z\. Zhi, C\. Liu, S\. Xing, Z\. Tu, and C\. LiuDoes rlvr extend reasoning boundaries? investigating capability expansion in vision\-language models\.ArXiv preprintabs/2511\.00710\.External Links:[Link](https://arxiv.org/abs/2511.00710)Cited by:[§A\.3](https://arxiv.org/html/2609.05575#A1.SS3.p1.1)\.
- Sorscheret al\.\(2022\)B\. Sorscher, S\. Ganguli, and H\. SompolinskyNeural representational geometry underlies few\-shot concept learning\.Proceedings of the National Academy of Sciences119\(43\),pp\. e2200800119\.Cited by:[Appendix B](https://arxiv.org/html/2609.05575#A2.SS0.SSS0.Px2.p1.1),[Appendix B](https://arxiv.org/html/2609.05575#A2.SS0.SSS0.Px3.p1.1),[§1](https://arxiv.org/html/2609.05575#S1.p3.1),[§2\.2](https://arxiv.org/html/2609.05575#S2.SS2.p1.1)\.
- Tanget al\.\(2025a\)Y\. Tang, A\. Lagzian, S\. Anumasa, Q\. Zou, Y\. Zhu, Y\. Zhang, T\. Nguyen, Y\. Tham, E\. Adeli, C\. Cheng,et al\.Human\-like content analysis for generative ai with language\-grounded sparse encoders\.ArXiv preprintabs/2508\.18236\.External Links:[Link](https://arxiv.org/abs/2508.18236)Cited by:[§A\.1](https://arxiv.org/html/2609.05575#A1.SS1.p1.1),[§1](https://arxiv.org/html/2609.05575#S1.p2.1)\.
- Tanget al\.\(2025b\)Y\. Tang, H\. Saini, Z\. Yao, Z\. Lin, Y\. Liao, J\. Cui, Y\. Wang, M\. Du, and D\. LiuA unified theory of sparse dictionary learning in mechanistic interpretability: piecewise biconvexity and spurious minima\.ArXiv preprintabs/2512\.05534\.External Links:[Link](https://arxiv.org/abs/2512.05534)Cited by:[§1](https://arxiv.org/html/2609.05575#S1.p2.1)\.
- Wanget al\.\(2025a\)S\. Wang, L\. Yu, C\. Gao, C\. Zheng, S\. Liu, R\. Lu, K\. Dang, X\. Chen, J\. Yang, Z\. Zhang, Y\. Liu, A\. Yang, A\. Zhao, Y\. Yue, S\. Song, B\. Yu, G\. Huang, and J\. LinBeyond the 80/20 rule: high\-entropy minority tokens drive effective reinforcement learning for llm reasoning\.ArXiv preprintabs/2506\.01939\.External Links:[Link](https://arxiv.org/abs/2506.01939)Cited by:[§A\.3](https://arxiv.org/html/2609.05575#A1.SS3.p1.1),[Appendix B](https://arxiv.org/html/2609.05575#A2.SS0.SSS0.Px4.p1.1)\.
- Wanget al\.\(2025b\)X\. Wang, Y\. Hu, W\. Du, R\. Cheng, B\. Wang, and D\. ZouTowards understanding fine\-tuning mechanisms of LLMs via circuit analysis\.InProceedings of the 42nd International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.267\.External Links:[Link](https://icml.cc/virtual/2025/poster/46507)Cited by:[Appendix B](https://arxiv.org/html/2609.05575#A2.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.05575#S1.p1.1)\.
- Wanget al\.\(2025c\)Y\. Wang, Q\. Yang, Z\. Zeng, L\. Ren, L\. Liu, B\. Peng, H\. Cheng, X\. He, K\. Wang, J\. Gao, W\. Chen, S\. Wang, S\. S\. Du, and Y\. ShenReinforcement learning for reasoning in large language models with one training example\.ArXiv preprintabs/2504\.20571\.External Links:[Link](https://arxiv.org/abs/2504.20571)Cited by:[§A\.3](https://arxiv.org/html/2609.05575#A1.SS3.p1.1)\.
- Wenet al\.\(2025\)X\. Wen, Z\. Liu, S\. Zheng, S\. Ye, Z\. Wu, Y\. Wang, Z\. Xu, X\. Liang, J\. Li, Z\. Miao, J\. Bian, and M\. YangReinforcement learning with verifiable rewards implicitly incentivizes correct reasoning in base llms\.ArXiv preprintabs/2506\.14245\.External Links:[Link](https://arxiv.org/abs/2506.14245)Cited by:[§A\.3](https://arxiv.org/html/2609.05575#A1.SS3.p1.1),[Appendix B](https://arxiv.org/html/2609.05575#A2.SS0.SSS0.Px1.p1.1),[Appendix B](https://arxiv.org/html/2609.05575#A2.SS0.SSS0.Px4.p1.1)\.
- Wollschlägeret al\.\(2025\)T\. Wollschläger, J\. Elstner, S\. Geisler, V\. Cohen\-Addad, S\. Günnemann, and J\. GasteigerThe geometry of refusal in large language models: concept cones and representational independence\.ArXiv preprintabs/2502\.17420\.External Links:[Link](https://arxiv.org/abs/2502.17420)Cited by:[Appendix B](https://arxiv.org/html/2609.05575#A2.SS0.SSS0.Px3.p1.1),[§1](https://arxiv.org/html/2609.05575#S1.p3.1)\.
- Wurgaftet al\.\(2026\)D\. Wurgaft, C\. Rager, M\. Kowal, V\. Shyam, S\. Feucht, U\. Bhalla, T\. Haklay, E\. Bigelow, R\. Sarfati, T\. McGrath, O\. Lewis, J\. Merullo, N\. Goodman, T\. Fel, A\. Geiger, and E\. S\. LubanaManifold steering reveals the shared geometry of neural network representation and behavior\.ArXiv preprintabs/2605\.05115\.External Links:[Link](https://arxiv.org/abs/2605.05115)Cited by:[§A\.2](https://arxiv.org/html/2609.05575#A1.SS2.p1.1)\.
- Xiong \(2026\)B\. XiongThe lattice representation hypothesis of large language models\.ArXiv preprintabs/2603\.01227\.External Links:[Link](https://arxiv.org/abs/2603.01227)Cited by:[§1](https://arxiv.org/html/2609.05575#S1.p3.1)\.
- Xuet al\.\(2024\)Y\. Xu, Y\. Wang, H\. Huang, and H\. WangTracking the feature dynamics in llm training: a mechanistic study\.ArXiv preprintabs/2412\.17626\.External Links:[Link](https://arxiv.org/abs/2412.17626)Cited by:[§A\.1](https://arxiv.org/html/2609.05575#A1.SS1.p1.1),[Appendix B](https://arxiv.org/html/2609.05575#A2.SS0.SSS0.Px1.p1.1),[Appendix B](https://arxiv.org/html/2609.05575#A2.SS0.SSS0.Px4.p1.1),[§1](https://arxiv.org/html/2609.05575#S1.p1.1)\.
- Yanget al\.\(2024\)A\. Yang, B\. Yang, B\. Hui, B\. Zheng, B\. Yu, C\. Zhou, C\. Li, C\. Li, D\. Liu, F\. Huang, G\. Dong, H\. Wei, H\. Lin, J\. Tang, J\. Wang, J\. Yang, J\. Tu, J\. Zhang, J\. Ma, J\. Xu, J\. Zhou, J\. Bai, J\. He, J\. Lin, K\. Dang, K\. Lu, K\. Chen, K\. Yang, M\. Li, M\. Xue, N\. Ni, P\. Zhang, P\. Wang, R\. Peng, R\. Men, R\. Gao, R\. Lin, S\. Wang, S\. Bai, S\. Tan, T\. Zhu, T\. Li, T\. Liu, W\. Ge, X\. Deng, X\. Zhou, X\. Ren, X\. Zhang, X\. Wei, X\. Ren, Y\. Fan, Y\. Yao, Y\. Zhang, Y\. Wan, Y\. Chu, Y\. Liu, Z\. Cui, Z\. Zhang, and Z\. FanQwen2 technical report\.ArXiv preprintabs/2407\.10671\.External Links:[Link](https://arxiv.org/abs/2407.10671)Cited by:[§5\.2](https://arxiv.org/html/2609.05575#S5.SS2.p1.1)\.
- Yueet al\.\(2025\)Y\. Yue, Z\. Chen, R\. Lu, A\. Zhao, Z\. Wang, Y\. Yue, S\. Song, and G\. HuangDoes reinforcement learning really incentivize reasoning capacity in llms beyond the base model?\.ArXiv preprintabs/2504\.13837\.External Links:[Link](https://arxiv.org/abs/2504.13837)Cited by:[§A\.3](https://arxiv.org/html/2609.05575#A1.SS3.p1.1),[Appendix B](https://arxiv.org/html/2609.05575#A2.SS0.SSS0.Px1.p1.1),[Appendix B](https://arxiv.org/html/2609.05575#A2.SS0.SSS0.Px4.p1.1)\.
- Zhaoet al\.\(2025\)H\. Zhao, Z\. He, Y\. Tang, F\. Yang, A\. Payani, D\. Liu, and M\. DuRep2Text: decoding full text from a single llm token representation\.ArXiv preprintabs/2511\.06571\.External Links:[Link](https://arxiv.org/abs/2511.06571)Cited by:[§A\.1](https://arxiv.org/html/2609.05575#A1.SS1.p1.1),[§1](https://arxiv.org/html/2609.05575#S1.p2.1)\.
- Zhuet al\.\(2025\)X\. Zhu, M\. Xia, Z\. Wei, W\. Chen, D\. Chen, and Y\. MengThe surprising effectiveness of negative reinforcement in llm reasoning\.ArXiv preprintabs/2506\.01347\.External Links:[Link](https://arxiv.org/abs/2506.01347)Cited by:[§A\.3](https://arxiv.org/html/2609.05575#A1.SS3.p1.1)\.

## Appendix ARelated Work

### A\.1Lens to Interpret Model Internals

Diverse methods have been proposed to interpret a model’s internal representation\. The Logit Lens\([nostalgebraist, 2020](https://arxiv.org/html/2609.05575#bib.bib29)\)projects intermediate hidden states through the unembedding to reveal a layerwise prediction trajectory, and the Tuned Lens\([Belrose et al\., 2023](https://arxiv.org/html/2609.05575#bib.bib20)\)corrects its basis\-mismatch failures with learned per\-layer translators\. Patchscopes\([Ghandeharioun et al\., 2024](https://arxiv.org/html/2609.05575#bib.bib21)\)patches hidden states into inspection prompts to decode them in natural language, and Rep2Text\([Zhao et al\., 2025](https://arxiv.org/html/2609.05575#bib.bib1)\)decodes full text from a single token representation\. Probing classifiers\([Alain and Bengio, 2016](https://arxiv.org/html/2609.05575#bib.bib22);[Hewitt and Manning, 2019](https://arxiv.org/html/2609.05575#bib.bib23)\)test whether predefined concepts are linearly decodable\. Sparse autoencoders decompose polysemantic activations into dictionaries of monosemantic features\([Cunningham et al\., 2023](https://arxiv.org/html/2609.05575#bib.bib2);[Gao et al\., 2024](https://arxiv.org/html/2609.05575#bib.bib24);[Bussmann et al\., 2025](https://arxiv.org/html/2609.05575#bib.bib3)\), with LanSE grounding features directly in natural language\([Tang et al\., 2025a](https://arxiv.org/html/2609.05575#bib.bib30)\)and natural language autoencoders producing unsupervised explanations of activations\([Fraser\-Taliente et al\., 2026](https://arxiv.org/html/2609.05575#bib.bib28)\)\. For comparison across models and training, crosscoders diff models through shared dictionaries\([Lindsey et al\., 2024](https://arxiv.org/html/2609.05575#bib.bib25);[Jiralerspong and Bricken, 2026](https://arxiv.org/html/2609.05575#bib.bib26)\)and SAE Track follows feature dynamics along a training trajectory\([Xu et al\., 2024](https://arxiv.org/html/2609.05575#bib.bib27)\)\. Despite their diversity, all of these lenses map activations onto tokens, language, or predefined and learned features, spaces that are easier to interpret, but none directly reports*how concepts are encoded*in representation space—the spread, anisotropy, and norm structure of the region a concept occupies\.

### A\.2Neural Geometry

Neural geometry refers to the study of how information is spatially organized in the representation spaces of neural networks, often through overarching hypotheses supported by suggestive yet partial evidence\. The linear representation hypothesis holds that high\-level concepts are encoded as directions\([Park et al\., 2023](https://arxiv.org/html/2609.05575#bib.bib33)\), with roots in the linear analogies of word embeddings\([Korchinski et al\., 2025](https://arxiv.org/html/2609.05575#bib.bib39)\); superposition further explains how more features than dimensions can coexist as almost\-orthogonal directions\([Elhage et al\., 2022](https://arxiv.org/html/2609.05575#bib.bib32)\), and diverse evidence supports this linear view, from refusal directions to linear representations of space and time and emergent world models\([Arditi et al\., 2024](https://arxiv.org/html/2609.05575#bib.bib34);[Gurnee and Tegmark, 2024](https://arxiv.org/html/2609.05575#bib.bib36);[Li et al\., 2022](https://arxiv.org/html/2609.05575#bib.bib40)\)\. Beyond directions, various works uncover richer, nonlinear structures: categorical and hierarchical concepts form polytopes and orthogonal hierarchies\([Park et al\., 2024](https://arxiv.org/html/2609.05575#bib.bib38)\), some features are irreducibly multi\-dimensional, forming circles for days and months\([Engels et al\., 2024](https://arxiv.org/html/2609.05575#bib.bib35)\), SAE feature clouds exhibit multi\-scale structure\([Li et al\., 2025](https://arxiv.org/html/2609.05575#bib.bib70)\), and manifold steering reveals shared geometry between representation and behavior\([Wurgaft et al\., 2026](https://arxiv.org/html/2609.05575#bib.bib37)\)\. At a global level, neural collapse suggests that class representations might collapse to a simplex structure in the terminal phase of training\([Papyan et al\., 2020](https://arxiv.org/html/2609.05575#bib.bib69)\), while the platonic representation hypothesis suggests that representations across models and modalities might converge\([Huh et al\., 2024](https://arxiv.org/html/2609.05575#bib.bib31)\)\. However, we currently lack a method to track how training reshapes neural geometry; such changes remain characterized mostly by hypotheses that are not yet sufficiently supported\.

### A\.3Understanding How Post\-Training \(SFT & RL\) Works

Supervised fine\-tuning \(SFT\) and reinforcement learning \(RL\) are the prevailing paradigms of model post\-training, widely utilized for human alignment via RLHF and RLAIF\([Ouyang et al\., 2022](https://arxiv.org/html/2609.05575#bib.bib6);[Lee et al\., 2023](https://arxiv.org/html/2609.05575#bib.bib8)\)and for enhancing reasoning via RLVR\([Shao et al\., 2024](https://arxiv.org/html/2609.05575#bib.bib11);[Wang et al\., 2025c](https://arxiv.org/html/2609.05575#bib.bib7)\)\. Yet how these methods actually work remains poorly understood, and existing works hold seemingly contradictory views\. Some argue RLVR merely sharpens the output distribution over solutions the base model can already produce\([Yue et al\., 2025](https://arxiv.org/html/2609.05575#bib.bib10);[Shao et al\., 2025](https://arxiv.org/html/2609.05575#bib.bib12)\)and may even collapse the model’s capability boundary\([Dong et al\., 2025](https://arxiv.org/html/2609.05575#bib.bib13)\); others find RL genuinely incentivizes correct reasoning\([Wen et al\., 2025](https://arxiv.org/html/2609.05575#bib.bib14)\)and that prolonged RL expands reasoning boundaries toward novel capabilities\([Liu et al\., 2025](https://arxiv.org/html/2609.05575#bib.bib9);[Shen et al\., 2025](https://arxiv.org/html/2609.05575#bib.bib15)\)\. Comparative studies report that SFT tends to memorize training data while RL generalizes to unseen rule and visual variants, in both LLMs and VLMs\([Chu et al\., 2025](https://arxiv.org/html/2609.05575#bib.bib19)\)\. However, others show this narrative is conditional rather than universal, with SFT’s generalization depending on optimization choices, data composition, and model capability\([Ren et al\., 2026](https://arxiv.org/html/2609.05575#bib.bib18)\)\. Finer\-grained analyses identify key components of RL, such as high\-entropy forking tokens\([Wang et al\., 2025a](https://arxiv.org/html/2609.05575#bib.bib16)\), and show negative reinforcement mitigates diversity loss\([Zhu et al\., 2025](https://arxiv.org/html/2609.05575#bib.bib17)\)\. Crucially, almost all of these works provide perspectives at the behavior level, measuring shifts in the output distributions, and do not answer what RL and SFT do to the representation space itself\.

## Appendix BDiscussion

#### Tracking representation geometry\.

Mechanistic interpretability has largely focused on characterizing representations at a fixed model checkpoint, through directions, features, probes, or geometric structure\([Park et al\., 2023](https://arxiv.org/html/2609.05575#bib.bib33);[Engels et al\., 2024](https://arxiv.org/html/2609.05575#bib.bib35);[Lindsey et al\., 2024](https://arxiv.org/html/2609.05575#bib.bib25);[Jiralerspong and Bricken, 2026](https://arxiv.org/html/2609.05575#bib.bib26);[Duan et al\., 2026](https://arxiv.org/html/2609.05575#bib.bib65)\)\. Comparatively less attention has been paid to how the geometry of those representations evolves throughout training, although recent work has begun to study feature and circuit dynamics across checkpoints\([Xu et al\., 2024](https://arxiv.org/html/2609.05575#bib.bib27);[Wang et al\., 2025b](https://arxiv.org/html/2609.05575#bib.bib66)\)\. This question is increasingly important as modern models undergo multiple stages of pretraining and post\-training whose effects on internal representations remain only partially understood\([Chu et al\., 2025](https://arxiv.org/html/2609.05575#bib.bib19);[Yue et al\., 2025](https://arxiv.org/html/2609.05575#bib.bib10);[Wen et al\., 2025](https://arxiv.org/html/2609.05575#bib.bib14)\)\. We therefore view tracking representation geometry as a complementary problem to interpreting a static representation: rather than asking only what a model represents, we also ask where and how that representation changes as the model is trained\. Capsule Lens provides a simple way to make such changes measurable by matching the same interpretable geometric parameterization across model states\.

#### Coverage and disentanglement as recall and precision\.

The two metrics used in this work admit a natural interpretation analogous to recall and precision\. Held\-out coverage measures the fraction of unseen positive samples captured by a fitted capsule and can therefore be viewed as the*recall*of the capsule in locating a concept’s geometry\. Conversely, the disentanglement score measures the probability that a representation admitted by a capsule belongs to its target concept and can be viewed as its*precision*\. This distinction is important because prior work suggests that concept representations need not be perfectly isolated or one\-dimensional, but may instead occupy overlapping, multidimensional, or manifold\-like regions\([Sorscher et al\., 2022](https://arxiv.org/html/2609.05575#bib.bib45);[Engels et al\., 2024](https://arxiv.org/html/2609.05575#bib.bib35);[Modell et al\., 2025](https://arxiv.org/html/2609.05575#bib.bib50);[Bhalla et al\., 2026](https://arxiv.org/html/2609.05575#bib.bib48)\)\. Since a capsule is an intentionally simple geometric family fitted in closed form, high held\-out coverage indicates that the fitted region generalizes to unseen instances of the concept, whereas high disentanglement indicates that this simple region is relatively specific to that concept\. We therefore view disentanglement not only as a property of the fitting procedure, but also as a characteristic of the underlying concept geometry: concepts whose representations are compactly approximated by a capsule should naturally achieve higher precision than concepts whose geometry is strongly intertwined with other concepts\.

#### Why capsules?

Our choice of capsules is motivated primarily by simplicity and interpretability rather than by the claim that concept geometry must take this form\. Prior work has described concept representations using directions\([Park et al\., 2023](https://arxiv.org/html/2609.05575#bib.bib33);[Arditi et al\., 2024](https://arxiv.org/html/2609.05575#bib.bib34)\), higher\-dimensional manifolds\([Sorscher et al\., 2022](https://arxiv.org/html/2609.05575#bib.bib45);[Engels et al\., 2024](https://arxiv.org/html/2609.05575#bib.bib35);[Modell et al\., 2025](https://arxiv.org/html/2609.05575#bib.bib50)\), categorical and hierarchical geometric structures\([Park et al\., 2024](https://arxiv.org/html/2609.05575#bib.bib38)\), and cone\-like regions\([Wollschläger et al\., 2025](https://arxiv.org/html/2609.05575#bib.bib51)\)\. Capsules draw inspiration from these geometric views while retaining only a small number of directly interpretable parameters: the axis describes orientation, the axis span describes angular spread, and the norm bounds describe radial extent\. These parameters can be computed in closed form and compared directly across layers, models, and training checkpoints\. More expressive geometric descriptions, including ellipsoidal or topological characterizations\([Sorscher et al\., 2022](https://arxiv.org/html/2609.05575#bib.bib45);[Malhotra et al\., 2026](https://arxiv.org/html/2609.05575#bib.bib44)\), may capture richer structure but are less immediately reducible to a small set of quantities whose changes can be interpreted across model states\. Capsule Lens therefore deliberately trades geometric flexibility for a compact and trackable description of representation\.

#### From case studies to systematic hypotheses\.

The three dynamic case studies reveal several geometric patterns that merit substantially more systematic investigation\. CLIP pretraining produces broad restructuring across the network, PPO post\-training on visual question answering produces much smaller and strongly depth\-localized changes, and RLVR on mathematical reasoning reveals sparse concept\- and layer\-specific concentration\. These observations connect to a growing literature showing that pretraining, fine\-tuning, and reinforcement learning can affect model representations in qualitatively different ways\([Chu et al\., 2025](https://arxiv.org/html/2609.05575#bib.bib19);[Xu et al\., 2024](https://arxiv.org/html/2609.05575#bib.bib27);[Liu et al\., 2025](https://arxiv.org/html/2609.05575#bib.bib9);[Yue et al\., 2025](https://arxiv.org/html/2609.05575#bib.bib10);[Wen et al\., 2025](https://arxiv.org/html/2609.05575#bib.bib14);[Wang et al\., 2025a](https://arxiv.org/html/2609.05575#bib.bib16)\)\. They also raise broader questions about which geometric changes are associated with behavioral improvements, whether different training objectives produce characteristic geometric signatures, and whether these patterns generalize across models, datasets, optimization settings, and random seeds\. Establishing such general laws would require dedicated experimental studies for each training regime and is beyond the scope of the present work\. Our goal here is instead to demonstrate that Capsule Lens provides a practical instrument for posing and studying these questions\. We therefore treat the findings of our case studies as empirical hypotheses suggested by the framework and leave their systematic validation to future work\.

## Appendix CFuture Work

We highlight four promising directions for future work:

- •Systematically tracking specific training pipelines\.Capsule Lens can be used to study in detail how particular training procedures reshape concept geometry\. Extending the case studies in this work into dedicated analyses across objectives, datasets, checkpoints, and model families can provide key insights to better understand different training pipelines\.
- •Characterizing geometric differences across model layers\.Capsule Lens can be used to systematically compare concept geometry across depth, helping reveal how concepts are formed, transformed, concentrated, or disentangled between different layers of a model\.
- •Studying hierarchical concepts through capsule intersections\.Capsule Lens may offer a geometric view of concept hierarchy by analyzing how capsules corresponding to related concepts intersect, overlap, or contain one another\. This could provide a direct way to investigate how hierarchical semantic structure is organized in representation space\.
- •Steering representations with capsules\.Beyond analysis, the geometric structure identified by Capsule Lens may provide a basis for representation steering\. Interventions along a capsule’s axis, within its angular span, or toward particular regions of its norm band could offer a geometrically grounded way to manipulate concept\-related representations and study their causal effects on model behavior\.

## Appendix DSystematically Validating Capsule Lens

We provide systematic experiments further validating the reliability and design choices of Capsule Lens\. Section[D\.1](https://arxiv.org/html/2609.05575#A4.SS1)examines the effect of the number of samples used for capsule matching and the stability of the estimated geometry under sampling\. Section[D\.2](https://arxiv.org/html/2609.05575#A4.SS2)evaluates percentile\-based capsule boundaries in place of the minimum and maximum fitting samples, examining the resulting trade\-off between held\-out coverage and disentanglement\. Section[D\.3](https://arxiv.org/html/2609.05575#A4.SS3)further compares capsules with alternative geometric approximations across different boundary choices\.

### D\.1Effect of the Number of Samples per Concept

We examine how the number of samples used for capsule matching affects the estimated geometry\. We varymmfrom55to100100while evaluating each fitted capsule on the same set of100100held\-out samples\. The experiment covers796796visual concepts across eight model components, using three representative layers from each component and three independent draws per concept\.

Figure 7:Effect of the number of samplesmmon fitted capsules\.Results are averaged over796796visual concepts, eight model components, three representative layers per component, and three independent draws\.a\)Held\-out coverage increases withmmand gradually saturates as more samples are used for capsule matching\.b\)Axis spansas\_\{a\}and norm bandNbN\_\{b\}, normalized by their values atm=100m=100, increase withmmas the fitted capsule progressively covers the concept geometry\.Held\-out coverage increases rapidly withmm, reaching0\.8780\.878atm=25m=25,0\.9410\.941atm=50m=50, and0\.9710\.971atm=100m=100\(Figure[7](https://arxiv.org/html/2609.05575#A4.F7)\)\. The fitted axis span and norm band also increase withmm, as larger samples progressively capture more of the concept’s angular and norm extent\. In particular, them=50m=50fitting size used in our main experiments already achieves a mean held\-out coverage of0\.9410\.941, while increasing the sample size further provides diminishing gains\.

We further evaluate the sampling stability of the fitted axis by independently matching two capsules to disjoint sets ofmmsamples from the same concept and measuring the cosine distanceΔ​a\\Delta abetween their axes\. The estimated axes become increasingly consistent asmmgrows, with meanΔ​a\\Delta adecreasing from0\.11500\.1150atm=5m=5to0\.01410\.0141atm=50m=50and0\.00710\.0071atm=100m=100\. Table[4](https://arxiv.org/html/2609.05575#A4.T4)further shows that this trend is consistent across model components\. These values provide a reference scale for interpreting the axis drifts measured across training checkpoints\.

Table 4:Held\-out coverage, axis span, and axis dispersion across model components\.Results are computed over796796visual concepts and three representative layers per component\. Axis dispersionΔ​a\\Delta ais the cosine distance between axes independently fitted from two disjoint sample sets of the same concept\.
### D\.2Robustness to Capsule Boundary Estimation

Capsule matching defines its boundary using the extrema of the fitting samples:sa=maxi⁡dcos​\(xi,a\)s\_\{a\}=\\max\_\{i\}d\_\{\\cos\}\(x\_\{i\},a\),nlow=mini⁡‖xi‖n\_\{\\mathrm\{low\}\}=\\min\_\{i\}\\\|x\_\{i\}\\\|, andnhigh=maxi⁡‖xi‖n\_\{\\mathrm\{high\}\}=\\max\_\{i\}\\\|x\_\{i\}\\\|\. To examine whether our results depend on this choice, we replace these extrema with percentile\-based boundaries parameterized byqq\. We setsas\_\{a\}to theqq\-th percentile of the angular distances and\(nlow,nhigh\)\(n\_\{\\mathrm\{low\}\},n\_\{\\mathrm\{high\}\}\)to the centralq%q\\%interval of the sample norms, while keeping the fitted axis unchanged\. We evaluateq∈\{100,95,90,85,80\}q\\in\\\{100,95,90,85,80\\\}, whereq=100q=100recovers the original capsule\. We perform this analysis across different types of concepts and evaluate each fitted capsule using held\-out coverage and disentanglementDis\\mathrm\{Dis\}\.

Figure 8:Capsule parameters under percentile\-based boundaries\.Axis spansas\_\{a\}and norm bandNbN\_\{b\}are normalized by their values atq=100q=100and averaged across concepts\. Shaded regions indicate±1\\pm 1standard deviation\. Lower percentiles produce progressively tighter capsule boundaries across all four concept types\.As shown in Figure[8](https://arxiv.org/html/2609.05575#A4.F8), both the axis span and norm band decrease consistently as the percentile is lowered\. The largest contraction occurs betweenq=100q=100andq=95q=95: atq=95q=95, the mean axis span is0\.7080\.708–0\.8260\.826of its original value across concept types, while the mean norm band is0\.7400\.740–0\.7830\.783of its original value\. Thus, the extreme samples have a substantial effect on the fitted boundary, but the same qualitative contraction is observed across modalities and concept constructions\.

Figure 9:Held\-out coverage and disentanglement under percentile\-based boundaries\.Held\-out coverage \(left\) and disentanglementDis\\mathrm\{Dis\}\(right\) as the percentileqqdecreases from the original min/max boundary \(q=100q=100\) to increasingly tighter capsules\. Results are averaged across concepts, with shaded regions indicating±1\\pm 1standard deviation\.Tighter boundaries induce a consistent trade\-off between coverage and disentanglement \(Figure[9](https://arxiv.org/html/2609.05575#A4.F9)\)\. Reducingqqfrom100100to8080decreases held\-out coverage by approximately0\.320\.32–0\.340\.34across all four concept types, while increasing disentanglement by0\.0610\.061for visual,0\.1400\.140for lexical,0\.2010\.201for math, and0\.1700\.170for labelled concepts\. These results show that percentile\-based boundaries can substantially increase the selectivity of fitted capsules, but only at the cost of excluding a considerable fraction of unseen positive samples\. We therefore retain the min/max boundary in Capsule Lens, prioritizing generalization to unseen samples while reporting disentanglement separately to characterize overlap with other concepts\.

### D\.3Alternative Geometric Approximations under Different Boundaries

We further examine whether the comparison with alternative geometric approximations in Section[3\.3](https://arxiv.org/html/2609.05575#S3.SS3)persists under percentile\-based boundaries\. A*cone*retains the capsule axisaaand angular spansas\_\{a\}without constraining the norm, while a*band*retains only the norm interval\[nlow,nhigh\]\[n\_\{\\mathrm\{low\}\},n\_\{\\mathrm\{high\}\}\]\. A*ball*is centered at the mean of the fitting samples with radiusr=maxi⁡‖xi−c‖r=\\max\_\{i\}\\\|x\_\{i\}\-c\\\|, and a*box*uses the coordinate\-wise minima and maxima of the fitting samples\. All shapes are fitted to the same5050samples and evaluated using the same concepts, layers, held\-out samples, and negatives\.

We apply the percentile\-based boundary construction of Section[D\.2](https://arxiv.org/html/2609.05575#A4.SS2)to each geometric approximation\. Table[5](https://arxiv.org/html/2609.05575#A4.T5)reports held\-out coverage and disentanglement acrossq∈\{100,95,90,85,80\}q\\in\\\{100,95,90,85,80\\\}\. Tightening the boundaries produces the expected coverage–disentanglement trade\-off for the capsule, cone, and ball\. Across all five percentiles, the capsule consistently achieves higher disentanglement than the cone and ball, while the cone and ball retain higher coverage at the sameqq\. The band remains substantially less disentangled across all boundary choices, withDis\\mathrm\{Dis\}increasing only from0\.5140\.514to0\.5320\.532, further indicating that norm information alone provides limited separation\. The box continues to exhibit extremely low held\-out coverage across all settings\. Overall, the results reveal a consistent pattern across boundary choices: directional information accounts for most of the observed concept separation, while incorporating the norm constraint yields a modest additional increase in disentanglement\. The capsule therefore provides a compact parameterization that jointly captures both aspects while maintaining high held\-out coverage under its original boundary\.

Table 5:Alternative geometric approximations under percentile\-based boundaries\.Held\-out coverage \(Cov\.\) and disentanglementDis\\mathrm\{Dis\}, averaged across concept types, model components, and representative layers\. Lowerqqcorresponds to a tighter boundary;q=100q=100recovers the original min/max construction\.

## Appendix ECase Study I: Tracking Concept Geometry in CLIP Pretraining

This appendix extends the main\-text discussion of Section[5\.1](https://arxiv.org/html/2609.05575#S5.SS1)and Figure[4](https://arxiv.org/html/2609.05575#S5.F4), giving additional views on the same training run and checkpoints — the within\-concept distribution behind the block\-11 span contraction, and the concept\-level analysis of the drift signal\.

### E\.1Axis\-span contraction is population\-wide

Figure[10](https://arxiv.org/html/2609.05575#A5.F10)shows that the block\-11 span contraction reported in the main text holds across the concept population, rather than being driven by a subset of concepts\.

![Refer to caption](https://arxiv.org/html/2609.05575v1/images/experiment1/sa_violin_block11.png)Figure 10:Block\-11 span contraction is visible across the concept distribution\.Distribution of axis\-spansas\_\{a\}across concepts at block 11, by training epoch\. The distribution shifts downward \(mean≈0\.60→0\.28\\approx 0\.60\\to 0\.28\) and narrows over training, without a separate non\-contracting subpopulation appearing at any checkpoint\.
### E\.2Concept\-level structure

Figures[11](https://arxiv.org/html/2609.05575#A5.F11)–[12](https://arxiv.org/html/2609.05575#A5.F12)break down the block\-11 rotation reported in the main text by concept, illustrating that Capsule Lens can rank individual concepts by drift magnitude within a single probed block\.

![Refer to caption](https://arxiv.org/html/2609.05575v1/images/experiment1/top_drifted_concepts.png)Figure 11:Top\-20 most drifted concepts, block 11\.Total axis drift \(epoch\-10→\\tofinal\) summed across the five probed blocks \(stacked bars, one color per block\), for the 20 concepts with the largest totals\. Totals span a narrow range \(≈2\.88\\approx 2\.88–2\.922\.92\)\.![Refer to caption](https://arxiv.org/html/2609.05575v1/images/experiment1/least_drifted_concepts.png)Figure 12:Top\-20 most stable concepts, block 11\.Total axis drift \(epoch\-10→\\tofinal\) summed across the five probed blocks, for the 20 concepts with the smallest totals\. Totals \(≈2\.36\\approx 2\.36–2\.462\.46\) are only≈20%\\approx 20\\%below the most\-drifted group in Figure[11](https://arxiv.org/html/2609.05575#A5.F11), indicating that the gap between the two extremes is present but modest relative to the overall drift scale\.
### E\.3Axis drift vs\. span change

Figure[13](https://arxiv.org/html/2609.05575#A5.F13)relates the block\-11 span contraction to the per\-concept drift of Section[E\.2](https://arxiv.org/html/2609.05575#A5.SS2), jointly plotting both quantities for individual concepts\.

![Refer to caption](https://arxiv.org/html/2609.05575v1/images/experiment1/drift_sa_quadrants.png)Figure 13:Axis drift and span change at block 11, final checkpoint\.Axis drift vs\. span changeΔ​sa\\Delta s\_\{a\}for individual concepts, with five example concepts labeled at each of three regions of the plot: lowest span change in magnitude \(blue:*square*,*meeting*,*coffee*,*human*\), highest axis drift \(red:*field*,*grazing*,*zebra*,*zebras*,*enclosure*\), and largest span change in magnitude \(green:*arms*,*goal*,*salad*,*batter*,*award*\)\. All plotted concepts fall in a narrow drift band \(≈0\.65\\approx 0\.65–0\.870\.87\) and haveΔ​sa<0\\Delta s\_\{a\}<0\(Δ​sa∈\[−0\.57,−0\.14\]\\Delta s\_\{a\}\\in\[\-0\.57,\-0\.14\]\)\.

## Appendix FCase Study II: PPO for VLM on Visual Question Answering

This appendix extends the main\-text discussion of PPO post\-training on Qwen2\-VL, giving additional views on the same training run and checkpoints — the concept\-level breakdown of the layer\-27 drift signal, and the joint relationship between axis drift and span change\.

### F\.1Concept\-level heterogeneity

Figures[15](https://arxiv.org/html/2609.05575#A6.F15)–[17](https://arxiv.org/html/2609.05575#A6.F17)rank individual concepts by total axis drift at layer 27, the layer where drift concentrates\.

![Refer to caption](https://arxiv.org/html/2609.05575v1/images/experiment2/sa_violin_layer27.png)Figure 14:Axis\-span distribution at layer 27 across checkpoints\.Distribution of axis\-spansas\_\{a\}across concepts at layer 27, by PPO training step\. The distribution shape is similar across all five checkpoints; the mean varies between≈0\.71\\approx 0\.71and≈0\.74\\approx 0\.74\.![Refer to caption](https://arxiv.org/html/2609.05575v1/images/experiment2/top_drifted_concepts.png)Figure 15:Top\-20 most drifted concepts, layer 27\.Total axis drift at step 200, summed across the five probed layers \(stacked bars, one color per layer\), for the 20 concepts with the largest totals\. Layer 27 \(top segment\) accounts for most of each total\.
### F\.2Axis drift vs\. span change

Figure[16](https://arxiv.org/html/2609.05575#A6.F16)jointly plots axis drift and span changeΔ​sa\\Delta s\_\{a\}for individual concepts at layer 27, the final checkpoint\.

![Refer to caption](https://arxiv.org/html/2609.05575v1/images/experiment2/drift_sa_quadrants.png)Figure 16:Axis drift and span change at layer 27, final checkpoint\.Axis drift vs\. span changeΔ​sa\\Delta s\_\{a\}\(step 200 minus base\) for individual concepts, with five example concepts labeled at each of three regions of the plot: largest positive span change \(blue:*greeting*,*funny*,*running*,*breakfast*,*coach*\), lowest span change in magnitude \(red:*bunch*,*open*,*behind*,*ready*,*red*\), and largest negative span change \(green:*fans*,*electric*,*west*,*hall*,*police*\)\. Axis drift for all plotted concepts falls between0\.0180\.018and0\.0330\.033\.![Refer to caption](https://arxiv.org/html/2609.05575v1/images/experiment2/least_drifted_concepts.png)Figure 17:Top\-20 most stable concepts, layer 27\.Total axis drift at step 200, summed across the five probed layers, for the 20 concepts with the smallest totals\. The y\-axis maximum \(≈0\.027\\approx 0\.027\) is close to Figure[15](https://arxiv.org/html/2609.05575#A6.F15)’s \(≈0\.041\\approx 0\.041\), indicating the gap between most\- and least\-drifted concepts is present but modest\.![Refer to caption](https://arxiv.org/html/2609.05575v1/images/experiment2/drift_heatmap_layer27.png)Figure 18:Per\-concept axis drift at layer 27 across training\.Rows are the top 60 concepts by mean drift, columns are PPO training steps\. Color encodes cosine distance from the base axis\.

## Appendix GCase Study III: RLVR on Mathematical Reasoning

This appendix provides additional analyses for the RLVR experiment in Section[5\.3](https://arxiv.org/html/2609.05575#S5.SS3)\. We examine both the population\-level evolution of concept geometry and the individual trajectories of all1818mathematical concepts under the same GRPO training run\. Capsules are fitted using5050samples per concept and evaluated on a disjoint set of5050samples; held\-out coverage averages0\.9420\.942across training, ranging from0\.800\.80to1\.001\.00\.

### G\.1Population\-Averaged Geometry

We first average each capsule parameter across the1818mathematical concepts at every probed layer\. Figure[19](https://arxiv.org/html/2609.05575#A7.F19)shows that population\-level geometry remains highly stable under RLVR\. Mean axis drift stays below7×10−67\\times 10^\{\-6\}at every layer, and axis spans remain nearly unchanged at layers33–2020; only layer2727exhibits a small non\-monotonic change followed by a modest contraction\. Mean norm\-band midpoint and width are similarly stable throughout training\. These averages therefore indicate that RLVR induces little global restructuring of concept geometry\.

Figure 19:Population\-averaged concept geometry during RLVR\.\(a\)Mean axis span across the1818mathematical concepts remains nearly constant at layers33–2020, while layer2727shows a small late contraction\.\(b\)Mean axis drift from the base model remains below7×10−67\\times 10^\{\-6\}at every layer\.\(c,d\)Mean norm\-band midpoint and width remain largely unchanged throughout training\.The population average can nevertheless conceal changes concentrated in a small number of concepts\. We therefore inspect the trajectory of each mathematical concept separately\.

### G\.2Per\-Concept Trajectories

Per\-concept trajectories reveal a strongly heterogeneous response to RLVR\. At layer2727, only77of the1818concepts change their axis span by more than0\.0050\.005, whileunit\_time\_minuteexhibits by far the largest contraction and is analyzed in Figure[6](https://arxiv.org/html/2609.05575#S5.F6)of the main text\. The following figures report axis drift, axis span, norm\-band midpoint, and norm\-band width for the remaining concepts at all five probed layers\.

Figure 20:RLVR trajectories of arithmetic\-operation concepts \(I\)\.The conceptsop\_addition,op\_subtraction,op\_multiplication, andop\_divisionremain largely stable across training, with only minor changes in axis span and norm geometry\.Figure 21:RLVR trajectories of arithmetic\-operation concepts \(II\)\.The conceptsop\_percentage,op\_fraction\_ratio, andop\_comparisonalso exhibit limited geometric change, althoughop\_percentageshows somewhat larger variation than most other operation concepts\. Overall, arithmetic\-operation concepts remain geometrically stable under RLVR\.Taken together, these trajectories explain why the population averages remain nearly unchanged despite several visible outliers: RLVR does not broadly restructure the mathematical concept space, but selectively modifies the geometry of a small number of concepts at particular layers\. This provides additional evidence for the concept\- and layer\-specific concentration effect highlighted in Section[5\.3](https://arxiv.org/html/2609.05575#S5.SS3)\.

Figure 22:RLVR trajectories of entity concepts\.entity\_peopleexhibits one of the clearest geometric changes outsideunit\_time\_minute, including a layer\-2727span contraction and substantial norm\-bound changes at intermediate layers\. In contrast,entity\_countable\_objectremains comparatively stable, further illustrating the concept\-specific nature of RLVR\-induced geometric change\.Figure 23:RLVR trajectories of unit concepts \(I\)\.The conceptsunit\_dollar,unit\_distance,unit\_weight, andunit\_time\_hourremain largely stable throughout training, with only small changes in axis span and norm geometry\.Figure 24:RLVR trajectories of unit concepts \(II\)\.The conceptsunit\_time\_day,unit\_time\_week,unit\_time\_month, andunit\_time\_yearexhibit heterogeneous but generally limited changes\. Among them,unit\_time\_monthandunit\_time\_yearshow the largest remaining layer\-2727span contractions, but both are substantially smaller than the contraction observed forunit\_time\_minutein the main text\. This further indicates that RLVR\-induced concentration is concept\-specific rather than shared uniformly across the time\-unit family\.

## Appendix HGeometry\-Guided Representation Steering with Capsule Lens

Beyond locating and tracking concept geometry, we test whether the geometry identified by Capsule Lens can also provide useful directions for representation steering\. For a concept whose span curve exhibits multimodal structure, we apply sphericalkk\-means to its unit\-normalized representations to obtain mode centroids\. Given source and target centroidsμi\\mu\_\{i\}andμj\\mu\_\{j\}, we define the unit steering directiondi→j=\(μj−μi\)/‖μj−μi‖d\_\{i\\rightarrow j\}=\(\\mu\_\{j\}\-\\mu\_\{i\}\)/\\\|\\mu\_\{j\}\-\\mu\_\{i\}\\\|\. We use the fitted norm shell to calibrate the intervention magnitude, settingm=\(nlow\+nhigh\)/2m=\(n\_\{\\mathrm\{low\}\}\+n\_\{\\mathrm\{high\}\}\)/2andλ=k​m\\lambda=kmwithk∈\[0\.6,0\.9\]k\\in\[0\.6,0\.9\], and intervene ash′=h\+λ​di→jh^\{\\prime\}=h\+\\lambda d\_\{i\\rightarrow j\}\. When the two centroids correspond to different modes within one concept, we call this*intra\-concept steering*; when they correspond to different concepts, we call it*inter\-concept steering*\.

#### Intra\-concept steering\.

Table[6](https://arxiv.org/html/2609.05575#A8.T6)shows representative examples for polysemous textual concepts\. Steering toward opposite modes can selectively elicit different meanings of the same concept, suggesting that the modes identified within a capsule correspond to behaviorally meaningful semantic structure\.

Table 6:Intra\-concept steering between semantic modes\.Representative generations from Qwen2\.5\-1\.5B after steering toward different modes within the same concept geometry\.
#### Inter\-concept steering\.

We further construct directions between centroids belonging to different concepts\. Table[7](https://arxiv.org/html/2609.05575#A8.T7)shows that the same geometric construction can shift generations toward distinct textual and visual concepts\.

Table 7:Inter\-concept steering across textual and visual representations\.Each example is shown with its base generation followed by the generation obtained after steering toward the target concept\.Together, these qualitative results suggest that the geometry recovered by Capsule Lens can provide actionable directions for representation intervention: internal modes support steering between different meanings of the same concept, while relative concept locations support steering across concepts\. We leave systematic evaluation of this capability to future work\.

Similar Articles

One Lens, Many Worlds : A Capability-Typed Interface for World-Model Interpretability

arXiv cs.LG

This paper introduces WorldModelLens, an open-source substrate for interpretability of world models, using a capability-typed adapter interface that generalizes across diverse architectures like PlaNet, Dreamer, IRIS, and I-JEPA. The framework provides a unified hook-and-cache layer for activation analysis and adds only ~12% overhead when inactive.