Tag
This preprint reveals that under certain data symmetries, LLM training dynamics can be reduced to a low-dimensional subspace, making analysis more tractable and interpretable, with each coordinate corresponding to a clear mechanism.