induction-head

Tag

Cards List
#induction-head

@Tiberiu_Musat_: What really happens during LLM training? Can we ever disentangle those billions of interacting parameters? In our lates…

X AI KOLs Timeline · 5d ago Cached

This preprint reveals that under certain data symmetries, LLM training dynamics can be reduced to a low-dimensional subspace, making analysis more tractable and interpretable, with each coordinate corresponding to a clear mechanism.

0 favorites 0 likes
← Back to home

Submit Feedback