Canonical-basis realignment for Transformer LLMs: every hidden axis becomes independently measurable and controllable

Lobsters Hottest Papers

Summary

This research presents a canonical basis transformation for Transformer language models, enabling lossless, independent measurement and control of hidden space axes while maintaining model behavior. It includes demonstrations on multiple models and provides causal evidence for functional geometry in LLMs.

<p>The code essentially gives you a way to rotate a Transformer's internal coordinate system into a canonical basis that aligns with its own weight matrices in a lossless way. By absorbing the normalization gains directly into the adjacent weights and using orthogonal matrices built from the singular vectors of the model, you can transform architectures like Qwen or Pythia without altering their outputs or perplexity scores.</p> <p>Applying this transform reveals the actual hidden geometric structures operating inside the network. Once the model is rotated into this new perspective, you can see its internal mechanisms that were previously opaque. The authors found things like a bipolar oscillator where specific axes form inhibitory pairs that fire against each other in perfect opposition. They also observed a kind of rhythmic respiration across layers where the model alternates between absorbing knowledge and filtering it. On top of that, it exposed a homeostatic defense mechanism that aggressively erases any localized perturbations within just a couple of layers.</p> <p>Practically speaking, researchers now have a powerful lens for mapping out how models actually do reasoning. For example, it turns out that the effective rank of the correlation matrix in a half billion parameter model might be as low as eleven independent patterns. Reframing how we look at the internal activations of language models provides a standardized way to study their underlying architecture.</p> <p><a href="https://lobste.rs/s/wg65qn/canonical_basis_realignment_for">Comments</a></p>
Original Article
View Cached Full Text

Cached at: 08/29/26, 09:53 PM

todotge/canonical-basis

Source: https://github.com/todotge/canonical-basis

A Canonical Basis for Interpreting Transformer Language Models

The Canonical Basis for Language Models (CBLL): a lossless coordinate transformation that opens the black box of Transformer LLMs and makes every axis of the hidden space independently measurable.

This repository is the reproducible companion to the Zenodo paper:

Gernone, G. (2026). The Hidden Geometry of Transformer Weights: A Journey Inside the Black Box. Zenodo. DOI: 10.5281/zenodo.20520986.

paper/The Hidden Geometry of Transformer Weights: A Canonical Basis for Interpreting Transformer Language Models_EN.md is the updated version 2 of the paper, extended with the causal ablation, the multi-architecture measurements, the LayerNorm bridge, and the MoE analysis. Every number in the paper comes from a script in scripts/ and pre-computed data in data/.


What this repository demonstrates

  1. Affine realignment is lossless — absorbing RMSNorm gains into the adjacent weight matrices and rotating with a Householder matrix changes the model’s coordinates without changing its behavior. Qwen 2.5 0.5B: PPL 25.38 = 25.38, MMLU 47.50% = 47.50%. SmolLM2 1.7B: PPL 6.6018 → 6.6045, top-5 overlap 5/5.

  2. Cross-layer U-alignment — the left singular vectors of the FFN down-projection are strongly aligned across layers on Qwen 2.5 0.5B (mean 0.651, max 0.928 over all 276 layer pairs; adjacent pairs align more strongly). The model has a shared set of preferred directions that no one imposed.

  3. Rich Club and bipolar oscillator — in the canonical basis, the 896 axes split into 309 positive-pole and 292 negative-pole axes; 83% of the positive-pole axes have a dedicated inhibitory partner. Master pair axis 62 ↔ axis 570, ρ = −0.97.

  4. Respiration — the POS/NEG ratio oscillates across the 24 layers in four phases (Encode 1.36 → Process 0.42–0.88 → Decode 1.28 → Output 0.54), invariant to input content. The same five layers [21, 3, 23, 2, 22] are the top activators for every prompt tested.

  5. Homeostasis — any intermediate perturbation of the residual stream is erased within two layers (5× → 1.4× → 1.0×), by the combined action of RMSNorm, attention softmax, and the SiLU operating range. This is architectural, not learned.

  6. Spectral collapse — the singular value magnitudes are nearly identical across layers (rank 1/24, ratio 23.2×). Layer identity lives in the geometry (U, V), not in the spectrum.

  7. Six spectral indices — cohesive, torsional, informational, dimensional, rhythmic, and vorticity indices quantify the structure per layer (means 0.588 / 0.917 / 0.709 / 0.378 / 0.595 / 0.650), with effective dimensionality 4.77/6 and a 41× isotropic collapse between the weight spectrum (k90/d 0.71) and the activation spectrum (k90/d 0.017).

  8. Causal evidence: single-axis ablation — zeroing axis 62 alone (0.11% of the model) collapses MMLU from 47.50% to 21.25% and destroys output coherence (PPL 4.24 → 23858). Zeroing its anti-correlated partner axis 570 degrades facts while keeping fluency. Five control axes show no effect (≤ ±1.25 pp). The geometry is functional, not decorative.

  9. Architectural generality — the realignment is lossless on RMSNorm families (Qwen, SmolLM2) and, via a DC-preserving rotation, on LayerNorm families (Pythia 1.4B: PPL 9.2286 → 9.2359, greedy generation identical 3/3). Native alignment measured on six families spans 0.02 (OLMo2) to 0.94 (Qwen); normalization type alone does not determine it.

  10. MoE structure — on OLMoE-1B-7B (64 experts), per-expert cross-layer alignment is 0.086 and cross-expert alignment is 0.111: the expert structure is per-expert, not shared across layers.


Why the canonical basis

The hidden state of a Transformer is a vector in R^d, but the basis in which it lives is arbitrary — whatever the training converged to. In the standard basis, nothing about dimension i is meaningful, and nothing can be compared across layers.

The canonical basis rotates the model so that axis k of the hidden state corresponds to a specific spectral direction of the model’s own weight matrices. After the rotation:

  • each axis can be measured independently (energy, correlation, entropy);
  • each axis can be manipulated independently (zeroed, amplified, traced from layer 0 to layer 23);
  • phenomena that are smeared across all 896 dimensions in the standard basis — the bipolar oscillator, the respiration, the critical axis — are localized onto single axes.

The rotation is lossless: the realigned model produces the same outputs as the original. It is a microscope, not a modification.

Why it is not a free operation

RMSNorm uses learned per-channel gains, and (g ⊙ h)R^T ≠ g ⊙ (hR^T) — the gains break rotational symmetry. The gains must first be absorbed into the adjacent weight matrices (W' = W @ diag(g)), making the normalizations uniform. LayerNorm additionally subtracts the mean, which is equivariant under rotation only for rotations that fix the ones-vector (DC-preserving rotations); its γ and β parameters are absorbed into weight columns and projection biases. Both procedures are lossless and are implemented in scripts/.


How to use

Quick verification (pre-computed data, no GPU)

All paper numbers are in data/ as JSON:

# Single-axis ablation (Section 4 of the paper)
python3 -c "import json; d=json.load(open('data/single_axis_ablation.json')); \
print('baseline', d['baseline']['mmlu']['accuracy']); \
[print(k, v['mmlu']['accuracy'], v['delta_mmlu']) for k,v in d.items() if k.startswith('axis_')]"

# Multi-architecture alignment (Section 5.3)
python3 -c "import json; d=json.load(open('data/gguf_cbll_multiarch.json')); \
[print(k, v['u_alignment_mean']) for k,v in d.items()]"

# Pythia DC bridge (Section 5.2)
python3 -c "import json; d=json.load(open('data/pythia_dc_bridge_results.json')); \
print(d['ppl_original'], '->', d['ppl_canonical'], 'lossless:', d['lossless'])"

Full reproduction (GPU with 6 GB VRAM tested)

pip install -r requirements.txt
bash reproduce.sh        # 7 steps, ~30 min: realign → collect → diagnose → ablate

Individual steps

python scripts/save_model.py --model Qwen/Qwen2.5-0.5B-Instruct \
    --output compressed_models/realigned_qwen05b     # realignment (lossless)

python scripts/diagnose_shared_sigma.py              # U-alignment, homeostasis, sigma
python scripts/investigate_rich_club.py              # POS/NEG poles
python scripts/analyze_correlations.py               # 896×896 correlation matrix
python scripts/benchmark_mmlu.py                     # MMLU 12×20 baseline
python scripts/ablate_single_axis.py \
    --test-axes 62 570 0 400 50 100 200 500 800      # causal ablation (9 axes)

Other models

  • SmolLM2 1.7B (RMSNorm, Llama-style): scripts/realign_smollm_cbll.py then scripts/smollm_cbll_continue.py for canonical-basis access.
  • Pythia 1.4B (LayerNorm, GPT-NeoX): scripts/pythia_dc_bridge.py — DC-preserving rotation, LayerNorm kept, lossless.
  • OLMoE-1B-7B (MoE, RMSNorm): scripts/analyze_olmoe_cbll.py — reads the rotated fp16 weights and computes dense/expert/cross-expert alignments and spectral indices.
  • GGUF models (Falcon3, StarCoder, Nemotron, DeepSeek-Coder, OLMo2): scripts/analyze_gguf_cbll.py — dequantizes GGUF tensors from Ollama blobs and measures native U-alignment and k90/d.

Configuration: where the paths live

There are no hardcoded machine paths in the published scripts. Everything is configured through environment variables or falls back to standard locations:

VariableUsed byDefaultWhat to set
HF_HOMEall scripts that download models~/.cache/huggingfaceWhere HuggingFace stores model checkpoints. Set once if you keep them elsewhere.
HF_HUB_CACHEHF hub downloads$HF_HOME/hubSub-cache for hub files. Normally not needed.
OLLAMA_BLOBSscripts/analyze_gguf_cbll.py~/.ollama/models/blobsDirectory containing Ollama GGUF blobs (files named sha256-...). Set to your Ollama models directory if it is not the default.
OLMOE_ROTATED_DIRscripts/analyze_olmoe_cbll.py~/.cache/huggingface/models--allenai--OLMoE-1B-7B-0924/rotated_fp16Directory with the rotated fp16 OLMoE shards.
OLMOE_R_PATHscripts/analyze_olmoe_cbll.pynext to the rotated weightsPath to the olmoe_R.npy rotation matrix.

Example:

export HF_HOME=/data/models
export OLLAMA_BLOBS=/data/ollama/blobs
export OLMOE_ROTATED_DIR=/data/olmoe/rotated_fp16
export OLMOE_R_PATH=/data/olmoe/olmoe_R.npy
bash reproduce.sh

Scripts write their outputs to runs/ (gitignored) and read the realigned model from compressed_models/ (gitignored, produced by scripts/save_model.py). Pre-computed results live in data/ and are never overwritten by a run — regenerate, then compare against data/.


Debugging and troubleshooting

Expected values (sanity checks)

StepExpectedIf not, check
save_model.pylogit diff < 1e-3, PPL identicalrotation built in fp32? gains all absorbed? hooks registered?
benchmark_mmlu.pybaseline 47.50% (114/240)tokenizer: Qwen chat template, choices encoded without special tokens
ablate_single_axis.pyaxis 62 → 21.25%, axis 500 → 47.50%realigned state_dict loaded with strict=True? hooks: embed @ R^T, last layer @ R
realign_smollm_cbll.pyPPL 6.6018 → 6.6045, top-5 5/5fp16 noise of 10⁻² in logits is expected and harmless
pythia_dc_bridge.pyPPL 9.2286 → 9.2359, greedy 3/3 identicalTO-side biases rotated (b @ R^T)? R fixes the ones-vector?
analyze_gguf_cbll.pytable matches data/gguf_cbll_multiarch.jsonOLLAMA_BLOBS correct? pip install gguf?

Common failure modes

  1. “Realigned model produces garbage” — almost always a missed absorption or a wrong hook direction. Check in order: (a) are all RMSNorm gains 1.0 after absorption? (b) embed hook rotates with R^T and the last-layer hook with R (row-vector convention)? (c) are the TO-side biases rotated (b @ R^T)? Qwen-class models have no biases; GPT-NeoX has them on every projection.

  2. “PPL differs by a small amount” — expected. fp16 forward passes introduce logit differences of ~10⁻²; behaviorally invisible (top-5 5/5, greedy generation identical). Bit-exactness requires fp32.

  3. “k90/d values look wrong” — use full SVD (np.linalg.svd(W, compute_uv=False)), not randomized SVD with a small number of components. Randomized SVD underestimates k90 badly. The OLMoE script includes both; trust d_k90_fullsvd.json.

  4. “Axis zeroing does nothing” — check the axis is zeroed in canonical space (after the embed rotation, before the unrotation), and that the intervention is at the weight level (zero the W_down row), not the activation level: activation-level perturbations are erased by homeostasis within two layers — that is the point of Section 3.3.

  5. “CUDA out of memory” — the GPU holds one model at a time (6 GB VRAM). Compare logits by saving the reference logits, freeing the model, then loading the canonical one (the pattern used in realign_smollm_cbll.py).

  6. “LayerNorm realignment degrades” — make sure you use the DC-preserving rotation, not the RMSNorm replacement. The condition is R^T 𝟙 = 𝟙; verify with np.abs(R @ np.ones(d) - np.ones(d)).max() < 1e-8.

Verifying the canonical basis is accessible

The quickest check, on any realigned model:

# zero canonical axis k at the last layer output, before unrotation
# PPL axis 0 (Qwen) -> catastrophic; axis 500 -> unchanged
python scripts/ablate_single_axis.py --test-axes 62 500 --skip-chat

Or interactively with demo/chat.py: /ablate 62, then /restore.


Models and architectures

FamilyNormalizationStatusWhat we have
Qwen 2.5 0.5B InstructRMSNormFull CBLL pipelinealignment, Rich Club, respiration, homeostasis, collapse, ablation
Qwen 2.5 1.5BRMSNormRealignedlosslessness, alignment
Qwen 3.5 4BRMSNorm (hybrid Mamba-FFN)Native measurementsU-alignment 0.606, k90 0.706
SmolLM2 1.7BRMSNormFull pipelinelossless realignment, axis access, indices
Pythia 1.4BLayerNormDC bridgelossless realignment, axis access, indices
Falcon3 3BLayerNormNative measurementsU-alignment 0.474, k90 0.670
StarCoder 1B/3BLayerNormNative measurementsU-alignment 0.45–0.47
Nemotron-Mini 4BRMSNorm + biasNative measurementsU-alignment 0.249
DeepSeek-Coder 6.7BRMSNormNative measurementsU-alignment 0.157
OLMo2 7Bnon-parametric LNNative measurementsU-alignment 0.020
OLMoE-1B-7BRMSNorm (MoE, 64 experts)Realigned + weight-level CBLLdense/expert/cross-expert alignment, sigma redundancy

On request we extend the pipeline to further architectures: the DC-preserving rotation covers every LayerNorm family, and the absorption procedure covers every parametric normalization.

Realigned models we release: the OLMo MoE family (OLMoE-1B-7B) in canonical basis — rotated fp16 weights plus the rotation matrix — so that anyone can run the measurements without re-running the realignment.


Interactive demo: per-axis control chat

demo/chat.py is a self-contained interactive chat with surgical per-axis control — the fastest way to feel the causal result of Section 4.

pip install torch transformers
python demo/chat.py                    # normal chat
python demo/chat.py --ablate 62        # chat with axis 62 removed
python demo/chat.py --ablate 62 --compare   # side-by-side baseline vs ablated
python demo/chat.py --prompt "Capital of Italy?"   # non-interactive single prompt

First run downloads Qwen 2.5 0.5B (~1 GB), runs the affine realignment (~30 s GPU, ~60 s CPU), and caches the realigned model. Subsequent runs load the cache instantly. Works on CPU (GPU optional, ~3× faster).

Interactive commands inside the chat:

/ablate 62              # zero axis 62 at every layer's FFN output
/ablate 62 570          # zero multiple axes
/restore                # restore all axes
/quit                   # exit

What you see (from demo/README.md):

═══ BASELINE (all axes active) ═══
  Explain quantum computing in one sentence.
  Quantum computing uses qubits that can exist in multiple states simultaneously,
  enabling faster computation for certain problems.

═══ AXIS [62] ZEROED ═══
  Explain quantum computing in one sentence.
  The the the the the the the the the the...

═══ RESTORED ═══
  Explain quantum computing in one sentence.
  Quantum computing uses qubits that can exist in multiple states simultaneously,
  enabling faster computation for certain problems.

Axis 62 has near-zero static activation (energy rank 895/896) but maximum betweenness centrality (3702) in the axis correlation network — a dynamic controller, not a static feature. Zeroing it at every layer removes its contribution from the residual stream, and the model loses coordination across layers. /restore brings it back — the intervention is reversible, the effect is immediate, and it reproduces the causal ablation of Section 4 interactively.


Repository layout

canonical-basis/
├── README.md                  # this file
├── reproduce.sh               # full pipeline, 7 steps
├── requirements.txt           # torch, transformers, datasets, numpy, scipy
├── CITATION.cff               # citation metadata
├── LICENSE
├── paper/
│   ├── The Hidden Geometry of Transformer Weights:A Canonical Basis for Interpreting Transformer Language Models_EN.md     # paper v2 (English)
│   └── The Hidden Geometry of Transformer Weights:A Canonical Basis for Interpreting Transformer Language Models_IT.md  # paper v2 (Italian)
├── demo/
│   ├── chat.py                # interactive per-axis control chat
│   └── README.md              # demo usage
├── assets/                    # figures and demo generators
├── validation/                # validation notes
├── scripts/                   # one script per paper claim
└── data/                      # pre-computed results (all paper numbers)

What is published and what is not. data/ contains the pre-computed results — the exact numbers cited in the paper. scripts/ contains the code that produces them. The regenerated outputs (runs/) and the realigned model weights (compressed_models/) are not committed: they are large, and they are regenerated deterministically by reproduce.sh from a public HuggingFace checkpoint. data/ is the ground truth of record; a fresh run writes to runs/ and must match data/ within fp16 noise.


License

Code: MIT. Data: CC-BY 4.0.

Patent notice

This repository is released for research purposes only. For commercial use, and for data on other models not cited in this work, contact [email protected].

Citation

@misc{gernone2026canonical,
  title={The Hidden Geometry of Transformer Weights: A Journey Inside the Black Box},
  author={Gianluca Gernone},
  year={2026},
  note={Version 2, with reproducible scripts and data},
  howpublished={Zenodo DOI: 10.5281/zenodo.21935673}
}

Similar Articles

See it, Say it, Sorted: Mechanistic Diagnosis and Parameter-Space Mitigation of Emergent Misalignment in LLMs

Hugging Face Daily Papers

A mechanistic study of emergent misalignment in LLMs finds that directional Hessian curvature concentrates on semantic pivot tokens and that harmful-safesubspace divergence drives failures; the authors propose a geometric mitigation framework that orthogonally projects harmful gradient subspaces, suppressing emergent misalignment by up to 80% on Qwen2.5-14B-IT.

The Geometry of Inference in Transformer Residual Streams

Hugging Face Daily Papers

This paper studies how transformer intermediate residual states become specialized to their own final output states, finding that directional alignment and endpoint rank improve even when Euclidean distance barely changes across six pretrained language models. It offers a high-dimensional model separating norm, alignment, and endpoint geometry, proving that straight-path convergence cannot introduce new competitors, and linking residual geometry to output token rankings.