@Tiberiu_Musat_: What really happens during LLM training? Can we ever disentangle those billions of interacting parameters? In our lates…

X AI KOLs Timeline Papers

Summary

This preprint reveals that under certain data symmetries, LLM training dynamics can be reduced to a low-dimensional subspace, making analysis more tractable and interpretable, with each coordinate corresponding to a clear mechanism.

What really happens during LLM training? Can we ever disentangle those billions of interacting parameters? In our latest preprint, we present a promising direction. Under certain data symmetries, the parameter space turns out to have a low-dimensional subspace with self-contained dynamics during training. If training begins in this subspace, it never leaves. The subspace has a low dimensionality, which means that we can reduce the training dynamics to a small number of pseudo-parameters. This makes theoretical and experimental analysis much more tractable. Moreover, the subspace is highly interpretable: each coordinate corresponds to a clear mechanism. For example, an induction head is composed of 3 directions in this subspace. Big thanks to my co-authors @tpimentelms @NicolasZucchet Paper link in the first reply
Original Article
View Cached Full Text

Cached at: 07/17/26, 12:24 AM

What really happens during LLM training? Can we ever disentangle those billions of interacting parameters? In our latest preprint, we present a promising direction.

Under certain data symmetries, the parameter space turns out to have a low-dimensional subspace with self-contained dynamics during training. If training begins in this subspace, it never leaves.

The subspace has a low dimensionality, which means that we can reduce the training dynamics to a small number of pseudo-parameters. This makes theoretical and experimental analysis much more tractable.

Moreover, the subspace is highly interpretable: each coordinate corresponds to a clear mechanism. For example, an induction head is composed of 3 directions in this subspace.

Big thanks to my co-authors @tpimentelms @NicolasZucchet

Paper link in the first reply

Similar Articles

LLMs are not the black box you were promised

Hacker News Top

An article summarizing Anthropic's 2025 paper on mechanistic interpretability, showing that LLMs are not black boxes and that circuit tracing can reveal multi-step reasoning and human-identifiable concepts.