Finding Usable Weight Mechanisms with Tiled SVD

arXiv cs.AI Papers

Summary

This paper proposes extracting mechanism mounts directly from linear weight sites via column-tiled SVD, offering an alternative to proxy dictionaries like sparse autoencoders for mechanistic interpretability. Evaluated on Gemma-2-2B, the method passes all 182 site-layer checks.

arXiv:2608.06969v1 Announce Type: new Abstract: The dominant approach to mechanistic interpretability trains proxy dictionaries such as sparse autoencoders and labels features from max-activating text. The best such atlases identify con- cepts, but that identity lives in the learned dictionary rather than in the network weights them- selves. We propose extracting mechanism mounts directly from linear sites by column-tiled SVD: each mount is a triple (v,u,{\sigma}) read as trigger, write, and strength. Identity is the weight rule. We evaluate mounts with a pre-registered suite judged on full-write energy lift rather than tile-local lift. On Gemma-2-2B with WikiText-2 (16,384-token subsample), all seven linear maps are scored: residual writes (mlp.down, attn.o) receive full A/B/C with steer after post-sublayer RMSNorm and pass 52/52 site-layers; other maps receive A/B only (mlp.gate/attn.q/attn.k/effective mlp.up/attn.v 26/26 each). Aggregate: 182/182 GO. We release library code, the corpus builder, the experiment entrypoint, and unit tests.
Original Article
View Cached Full Text

Cached at: 08/10/26, 08:00 AM

# 1 Introduction
Source: [https://arxiv.org/html/2608.06969](https://arxiv.org/html/2608.06969)
Finding Usable Weight Mechanisms with Tiled SVD

Ash Manvi

Aquin Labs

ash@aquin\.app

Samreena Tajreen

Aquin Labs

samreena@aquin\.app

Abstract

> The dominant approach to mechanistic interpretability trains proxy dictionaries such as sparse autoencoders and labels features from max\-activating text\. The best such atlases identify concepts, but that identity lives in the learned dictionary rather than in the network weights themselves\. We propose extracting*mechanism mounts*directly from linear sites by column\-tiled SVD: each mount is a triple\(v,u,σ\)\(v,u,\\sigma\)read as trigger, write, and strength\. Identity is the weight rule\. We evaluate mounts with a pre\-registered suite judged on full\-write energy lift rather than tile\-local lift\. On Gemma\-2\-2B with WikiText\-2 \(16,384\-token subsample\), all seven linear maps are scored: residual writes \(mlp\.down,attn\.o\) receive full A/B/C with steer after post\-sublayer RMSNorm and pass52/52site\-layers; other maps receive A/B only \(mlp\.gate/attn\.q/attn\.k/effectivemlp\.up/attn\.v26/26 each\)\. Aggregate:182/182GO\. We release library code, the corpus builder, the experiment entrypoint, and unit tests\.

Sparse autoencoders and related proxy dictionaries have become standard tools in mechanistic interpretability for naming directions in neural networks\[[1](https://arxiv.org/html/2608.06969#bib.bib1),[2](https://arxiv.org/html/2608.06969#bib.bib2)\]\. Trained on activations and labeled from max\-activating text, they produce concept atlases that are useful for describing what a model represents\. Numerous efforts have since refined dictionary learning, scaling, and automatic labeling, while also improving coverage of polysemantic neurons\.

These atlases typically identify a feature by the learned dictionary atom and its verbal label\. Aligning interpretation to a separately trained codebook, rather than to a particular weight matrix inside the network, means that “what this direction means” is answered in the proxy space\. The weight rule that actually writes into the residual stream is left implicit\. Recent work has shown that singular vectors of MLP and attention matrices, read out through the unembedding, often form interpretable token clusters and can be edited\[[3](https://arxiv.org/html/2608.06969#bib.bib3)\], and has framed singular modes as detector\-effector units\[[4](https://arxiv.org/html/2608.06969#bib.bib4)\]\. In most such settings, however, SVD is used as a lens or circuit primitive\[[5](https://arxiv.org/html/2608.06969#bib.bib5)\], not as a fair test of which chunking ofWWyields usable on\-distribution mechanisms\.

In this work we propose extracting*mechanism mounts*directly from linear sites by column\-tiled SVD\. Each mount is a triple\(v,u,σ\)\(v,u,\\sigma\)read as trigger, write, and strength; identity is the weight rule itself\. We evaluate mounts with a measurement stack judged on full\-write energy lift rather than tile\-local lift, which favors one\-column tiles tautologically, together with coverage saturation and a depth\-conditioned steer check against the final unembedding on residual writes\. On Gemma\-2\-2B, with a WikiText\-2 subsample, the suite covers all seven linear maps per layer and passes on all 182 site\-layers oncemlp\.upandattn\.vuse effective\-path mounts\.

## 2Method

### 2\.1Model and sites

We evaluategoogle/gemma\-2\-2b\(26 layers, indices0,…,250,\\ldots,25\)\. Each layer exposes seven linear maps\. Two are*residual writes*:mlp\.down\(W∈ℝ2304×9216W\\in\\mathbb\{R\}^\{2304\\times 9216\}\) andattn\.o\(W∈ℝ2304×2048W\\in\\mathbb\{R\}^\{2304\\times 2048\}\)\. These receive the full suite of chunking, coverage, and causal checks \(A/B/C\), with Experiment C injecting after Gemma\-2 post\-sublayer RMSNorm so the steered direction lands in the residual stream\. The remaining five maps \(mlp\.gate,mlp\.up,attn\.q,attn\.k,attn\.v\) receive A/B only: their outputs are not residual writes, so unembed alignment is not the right causal metric\. The same full\-write energy lift and coverage criteria apply everywhere; only C is site\-conditional\.

Corpus text comes from WikiText\-2 \(raw train\), built byscripts/build\_corpus\.py\. Forwards collect all tokens, then subsample to16,38416\{,\}384tokens with seed0for memory\.

![Refer to caption](https://arxiv.org/html/2608.06969v1/figures/vt_block.png)Figure 1:Linear sites in one Gemma\-style decoder block \(scaled stand\-in for VisualTorch\)\. OrangeABOnlyLinear:attn\.q/k/v,mlp\.gate/up\(A/B only\)\. GreenResidualWriteLinear:attn\.oandmlp\.down\(A/B/C residual writes\)\.
### 2\.2Tile\-SVD mounts

LetTTbe the tile width andkkthe number of modes per tile\. Defaults are site\-aware:T=512T\{=\}512for residual and MLP maps,T=256T\{=\}256forattn\.k/attn\.v, andT=128T\{=\}128forattn\.q\(with a6464fallback on hard layers\)\. Experiments A and C usek=2k\{=\}2; Experiment B sweepsk∈\{1,2,4,8,16\}k\\in\\\{1,2,4,8,16\\\}\.

Partition the input columns ofWWinto tiles\[st,et\)\[s\_\{t\},e\_\{t\}\)and factor each tile

Bt=W:,st:et=Ut​Σt​Vt⊤\.B\_\{t\}=W\_\{:,s\_\{t\}:e\_\{t\}\}=U\_\{t\}\\,\\Sigma\_\{t\}\\,V\_\{t\}^\{\\top\}\.For modei<ki<k,

u=Ut​\[:,i\]‖Ut​\[:,i\]‖2,σ=Σt​\[i\],v=\(Vt⊤\)​\[i,:\]\.u=\\frac\{U\_\{t\}\[:,i\]\}\{\\\|U\_\{t\}\[:,i\]\\\|\_\{2\}\},\\qquad\\sigma=\\Sigma\_\{t\}\[i\],\\qquad v=\(V\_\{t\}^\{\\top\}\)\[i,:\]\.A mount is the triple\(v,u,σ\)\(v,u,\\sigma\)with site and layer metadata\. Identity is this weight rule, not a verbal label\. Mount id:tile:\{t\}:sv\{i\}\.

Under a matched mount budget we compare four constructions:

![Refer to caption](https://arxiv.org/html/2608.06969v1/figures/tile_svd.png)Figure 2:Column\-tiled SVD of a linear weight\. Each tile yields modes\(v,u,σ\)\(v,u,\\sigma\)read as trigger, write, and strength\.
### 2\.3Triggering and full\-write energy lift

On real forwards with site inputsxx, trigger coefficients are

at,j=xt,sj:ej⋅vj\.a\_\{t,j\}=x\_\{t,\\,s\_\{j\}:e\_\{j\}\}\\cdot v\_\{j\}\.Tile writes use the corresponding column block:Δ​htile=x:,s:e​B⊤\\Delta h^\{\\mathrm\{tile\}\}=x\_\{:,s:e\}B^\{\\top\}\.

SVD identity \(sanity, not proof\) requires correlation ofaawithΔ​htile​u\\Delta h^\{\\mathrm\{tile\}\}uabove0\.990\.99and relative slope error below0\.050\.05\. This identity holds almost always by construction, so it cannot separate usable mounts from unused ones\.

Tile\-local energy lift

Ltile=𝔼t​\[\(at​σ\)2‖Δ​httile‖22\+ε\]−𝔼t​\[\(Δ​httile⋅u~\)2‖Δ​httile‖22\+ε\]L\_\{\\mathrm\{tile\}\}=\\mathbb\{E\}\_\{t\}\\\!\\left\[\\frac\{\(a\_\{t\}\\sigma\)^\{2\}\}\{\\\|\\Delta h^\{\\mathrm\{tile\}\}\_\{t\}\\\|\_\{2\}^\{2\}\+\\varepsilon\}\\right\]\-\\mathbb\{E\}\_\{t\}\\\!\\left\[\\frac\{\(\\Delta h^\{\\mathrm\{tile\}\}\_\{t\}\\cdot\\tilde\{u\}\)^\{2\}\}\{\\\|\\Delta h^\{\\mathrm\{tile\}\}\_\{t\}\\\|\_\{2\}^\{2\}\+\\varepsilon\}\\right\]favors one\-column tiles tautologically \(Ltile≈1L\_\{\\mathrm\{tile\}\}\\approx 1for column sampling\)\. Pass criteria therefore use*full\-write*energy lift against the site’s write tensorΔ​h\\Delta h:

Lfull​\(u\)=𝔼t​\[\(Δ​ht⋅u\)2‖Δ​ht‖22\+ε\]−𝔼t​\[\(Δ​ht⋅u~\)2‖Δ​ht‖22\+ε\],L\_\{\\mathrm\{full\}\}\(u\)=\\mathbb\{E\}\_\{t\}\\\!\\left\[\\frac\{\(\\Delta h\_\{t\}\\cdot u\)^\{2\}\}\{\\\|\\Delta h\_\{t\}\\\|\_\{2\}^\{2\}\+\\varepsilon\}\\right\]\-\\mathbb\{E\}\_\{t\}\\\!\\left\[\\frac\{\(\\Delta h\_\{t\}\\cdot\\tilde\{u\}\)^\{2\}\}\{\\\|\\Delta h\_\{t\}\\\|\_\{2\}^\{2\}\+\\varepsilon\}\\right\],with random directionsu~\\tilde\{u\}seeded by10007\+j⋅99710007\+j\\cdot 997\.

### 2\.4Coverage saturation

Weight coverage is the fraction of‖W‖F2\\\|W\\\|\_\{F\}^\{2\}kept by per\-tile rank\-kkreconstructions\. Sparse write coverage builds a dictionary of unit mount directions, selects the top\-kactivek\_\{\\mathrm\{active\}\}\(default88\) mounts per token, reconstructsΔ​h\\Delta hby least squares, and reports explained energy\. Coverage lift is the gap versus a random dictionary of matched size\. Saturation \(B1\) requires the lift peak to meet a site\-dependent floor \(0\.250\.25residual,0\.150\.15other maps,0\.080\.08effective up/v paths\), the early sweep window to lie within0\.080\.08of the peak \(m∈\{1,2\}m\\in\\\{1,2\\\}residual;m∈\{1,2,4\}m\\in\\\{1,2,4\\\}otherwise\), and the final point not to collapse more than0\.100\.10below the peak\.

### 2\.5Causal steer versus unembed

For residual writes only, we read the final\-logit geometry of a write directionuuby an unembed lens: approximate final RMSNorm, thent=Wlm​ut=W\_\{\\mathrm\{lm\}\}u\. Steering addsα​u\\alpha uwithα=2\\alpha\{=\}2on the post\-attention or post\-FF RMSNorm module \(not ono\_proj/down\_projalone\), averages last\-tokenΔ\\Deltalogits over up to eight texts, and reports Spearmanρ​\(Δ¯,t\)\\rho\(\\overline\{\\Delta\},t\)and top\-20 Jaccard\. C1 passes ifρ≥0\.05\\rho\\geq 0\.05orJ20≥0\.05J\_\{20\}\\geq 0\.05, and is*required*only when the site is a residual write and layerℓ≥6\\ell\\geq 6\. Early residual layers still report C; they do not fail when alignment is weak\. Non\-residual sites skip C\.

### 2\.6Effective\-path mounts

Rawmlp\.upandattn\.vmodule weights fail residual\-shaped A/B: the map that is used on\-distribution is not the raw matrix\. Defaults therefore mount from effective maps\. Formlp\.up,W⋆=diag​\(g¯\)​WupW^\{\\star\}=\\mathrm\{diag\}\(\\bar\{g\}\)\\,W\_\{\\mathrm\{up\}\}with corpus\-mean gate activationsg¯\\bar\{g\}\(fallbacks: gate\-mixture tile SVD; compose\-through\-downW↓​diag​\(g¯\)​WupW\_\{\\downarrow\}\\mathrm\{diag\}\(\\bar\{g\}\)W\_\{\\mathrm\{up\}\}\)\. Forattn\.v, ridge least\-squaresx→x\\tomixed\-vv\(fallback:W†=Wo​expandGQA​\(Wv\)W^\{\\dagger\}=W\_\{o\}\\,\\mathrm\{expand\}\_\{\\mathrm\{GQA\}\}\(W\_\{v\}\)\)\. Scoring uses the corresponding write tensors \(gated product or mixed\-vv/oowrites\)\.

![Refer to caption](https://arxiv.org/html/2608.06969v1/figures/vt_effective.png)Figure 3:Effective\-path mounts\. Left: mean\-gateW⋆W^\{\\star\}formlp\.up\. Right: ridge lstsqx→x\\tomixed\-vvforattn\.v\(compose\-o fallback\)\.
### 2\.7Pass criteria

A site\-layer is accepted when all applicable checks pass:

The judge isjudge\_paper\_goinsrc/atlas/mount/paper\_eval\.py\. The runner searches site\-aware tile sizes and effective\-path pools, keeps the best passing trial, and aggregates across all seven sites×\\times26 layers\.

## 3Experiments

### 3\.1Setup

We build a WikiText\-2 corpus and run the full suite once the model is loaded:

```
python scripts/build_corpus.py --out data/corpus/train.jsonl
python scripts/run_paper_experiments.py --layers all --sites all \
  --device cuda --texts data/corpus/train.jsonl \
  --out-dir data/eval/paper_experiments_all
```

Defaults: site\-aware tile sizes,k=2k\{=\}2modes per tile for A/C, modes sweep\{1,2,4,8,16\}\\\{1,2,4,8,16\\\}for B,nsteer=8n\_\{\\mathrm\{steer\}\}\{=\}8,α=2\\alpha\{=\}2,16,38416\{,\}384tokens subsampled from86,10986\{,\}109collected\. Outputs land under\{site\_slug\}/L\{n\}/with aggregatesites\.csv\.

Reported run\.All seven sites×\\times26 layers pass:182/182\(residual A/B/C52/52; other A/B130/130\)\.

### 3\.2Experiment A: chunking

On both residual\-write sites, tile full\-write lift exceeds whole\-matrix SVD, column sampling, and random at every depth exceptmlp\.downlayer 25, where tile≈\\approxwhole and A4 still passes on ratio\.attn\.ouses fewer mounts \(din=2048d\_\{\\mathrm\{in\}\}\{=\}2048\) thanmlp\.downand still wins A everywhere\. Columntile\_lift≈0\.999\\approx 0\.999is ignored; A1–A4 use full\-write lift only\.

![Refer to caption](https://arxiv.org/html/2608.06969v1/figures/lift_depth.png)Figure 4:Experiment A\. Mean full\-write energy lift versus depth for residual writes\. Tile SVD exceeds whole\-matrix SVD, column sampling, and random\.Local column structure inWWis not well summarized by a single global SVD for on\-distribution write energy\. Tiling recovers higher\-energy write directions under a fixed mount count on both residual writes\.

### 3\.3Experiment B: coverage versus modes

On residual writes, coverage lift is already high atm=1m\{=\}1orm=2m\{=\}2and then flat: B1 passes all 52 residual site\-layers under the early\-saturation rule\. Extra modes mostly add redundancy\.

![Refer to caption](https://arxiv.org/html/2608.06969v1/figures/coverage_modes.png)Figure 5:Experiment B\. Coverage lift versus modes per tile\. Lift is high bym=1m\{=\}1orm=2m\{=\}2and then flat \(saturation\)\.
### 3\.4Experiment C: causal depth

With post\-norm residual injection, mean Spearman versus unembed\(uu\) rises with depth on both residual sites\. Mid\-depthattn\.ono longer fails C1\. Final layers reachρ≈0\.91\\rho\\approx 0\.91\(mlp\.down\) andρ≈0\.75\\rho\\approx 0\.75\(attn\.o\)\. The C1 waiver forℓ<6\\ell<6remains motivated by this curve\.

![Refer to caption](https://arxiv.org/html/2608.06969v1/figures/steer_depth.png)Figure 6:Experiment C\. Mean Spearman of steeredΔ\\Deltalogits versus unembed\(uu\) after post\-norm injection\. Gray band:L<6L<6\(C1 waived\)\.
### 3\.5Aggregate verdict

Tile\-SVD mounts with full\-write energy lift, early coverage saturation, and depth\-aware post\-norm steer vs unembed pass on every site\-layer of all seven linear maps oncemlp\.upandattn\.vuse effective write maps rather than raw module weights\.

## 4Discussion

The measurement stack is the durable part of this work\. Full\-write energy lift separates tile SVD from whole\-matrix SVD, column sampling, and random under a matched budget, without rewarding the one\-column tautology that inflates tile\-local lift\. Coverage saturates early on residual writes\. With post\-norm injection, steer versus unembed alignment is a clear depth curve on bothmlp\.downandattn\.o, and the same judge accepts all seven linear maps oncemlp\.upandattn\.vare mounted from their effective write maps\. That is an honest multi\-site protocol, not an MLP\-only demo\.

Novelty is thinner\. Singular vectors of transformer weights, detector\-effector units, and unembed readouts already exist\[[3](https://arxiv.org/html/2608.06969#bib.bib3),[4](https://arxiv.org/html/2608.06969#bib.bib4),[5](https://arxiv.org/html/2608.06969#bib.bib5)\]\. We do not claim human\-readable concept names, and we do not claim to replace sparse autoencoders for concept discovery\[[1](https://arxiv.org/html/2608.06969#bib.bib1),[2](https://arxiv.org/html/2608.06969#bib.bib2)\]\. The wedge is fair chunking, a negative result about tile\-local metrics, coverage saturation, and a depth\-conditioned causal check packaged as a reproducible suite\.

Scope stays sharp\. Experiment C applies only whereuuis a residual direction\. Rawmlp\.up/attn\.vmodule weights fail residual\-shaped A/B by design; the supported object is the effective path\. All numbers are for Gemma\-2\-2B on a WikiText\-2 subsample\.

## 5Limitations

This study uses a single model family and size \(Gemma\-2\-2B\)\. Experiment C applies only to residual\-write sites; gate, up, q, k, and v are judged on A/B alone\. The WikiText\-2 subsample \(16,38416\{,\}384of86,10986\{,\}109tokens\) may bias which mounts look strong\. Energy lift is not human meaning: mounts carry no semantic labels\. Steering uses short texts, fixedα=2\\alpha\{=\}2, and last\-token logits after post\-sublayer RMSNorm\. The C1 waiver forℓ<6\\ell<6is principled from the depth curve but remains a design choice; earlyρ\\rhoshould always be reported\. Rawmlp\.upandattn\.vmodule weights fail residual\-shaped A/B and require effective\-path mounts\.mlp\.downlayer 25 is marginal on A4 \(tile≈\\approxwhole\) yet still passes;attn\.oruns with fewer mounts thanmlp\.down\.

## 6Conclusion

We extract tile\-SVD mechanism mounts\(v,u,σ\)\(v,u,\\sigma\)from linear sites of Gemma\-2\-2B and score them with full\-write energy lift, coverage saturation, and a depth\-conditioned steer check against the final unembedding on residual writes\. Under a matched mount budget, tiling beats whole\-matrix SVD, column sampling, and random; write coverage saturates by one to two modes per tile; and post\-norm residual steers agree with unembed\(uu\) increasingly with depth\. Across all seven linear maps and 26 layers the suite passes182/182site\-layers oncemlp\.upandattn\.vuse effective\-path mounts\. We release the library, corpus builder, experiment entrypoint, and unit tests, with mount identity defined as the weight rule itself\.

## Acknowledgments

This work was produced at Aquin Labs\. Gemma\-2\-2B is released by Google under its model license; accept that license before downloading weights\.

## References

- \[1\]Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nicholas L Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Alex Tamkin, Karina Nguyen, Brayden McLean, Josiah E Burke, Tristan Hume, Shan Carter, Tom Henighan, and Chris Olah\. Towards monosemanticity: Decomposing language models with dictionary learning\.*Transformer Circuits Thread*, 2023\.
- \[2\]Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey\. Sparse autoencoders find highly interpretable features in language models\. arXiv preprint arXiv:2309\.08600, 2023\.
- \[3\]Beren Millidge and Sid Black\. The singular value decompositions of transformer weight matrices are highly interpretable\.*Alignment Forum*, 2022\.
- \[4\]Min Xue and Artur Andrzejak\. SVD as a fast interpretability method for transformers\. In*Proceedings of the 43rd International Conference on Machine Learning \(ICML\)*, 2026\.
- \[5\]Areeb Ahmad, Abhinav Joshi, and Ashutosh Modi\. Beyond components: Singular vector\-based interpretability of transformer circuits\. arXiv preprint arXiv:2511\.20273, 2025\.

Similar Articles

Targeted Recovery of Weight-Space Mechanisms From Neural Networks

arXiv cs.LG

This paper introduces Targeted Parameter Decomposition (tPD), a method that selectively recovers interpretable weight-space mechanisms from neural networks for specific inputs, reducing compute requirements compared to full decomposition. It validates tPD on toy models and transformer language models, demonstrating faithful circuit recovery and surgical ablation with minimal side effects.

Experimenting with hypersurface-constrained dynamic weight updating [P]

Reddit r/MachineLearning

An experimental language model architecture uses hypersurfaces for dynamic weight updating to reduce training parameters, achieving better performance than unrolled baselines while using only 16% of the parameters. The approach is tested on the FineWeb-Edu dataset and includes a GitHub implementation.

A Mesoscopic View of Transformer Weights Through Row and Column Scale Fields

arXiv cs.LG

This paper introduces 'row and column scale fields'—median-centred log-RMS profiles of Transformer weight matrices over their channels—to mesoscopically analyze how weight magnitude is distributed across functional channels, showing how these fields evolve through training, align across projections, relate to AdamW's second-moment structure, and how edits to them affect model loss.

Decompose Sparsely Where You Should, Absorb Densely Where You Should No

arXiv cs.LG

The paper hypothesizes that language model activations contain a low-rank dense component that is inefficiently represented by sparse autoencoders (SAEs). By adding a linear bottleneck to absorb dense structure, the authors reduce dense latents and improve sparse probing performance on Gemma-2-2B.