Tag
This paper proposes using AstroPT, a transformer trained on galaxy images, as a testbed for studying concept emergence during training, finding that galaxy properties emerge in a fixed difficulty-based sequence, which can inform mechanistic interpretability methods for large language models.
This paper introduces a bifurcation theory of representation dynamics to detect when neural networks acquire structured representations during training, using a Hessian analysis of a GMM probe. The resulting ratio β/β_c serves as a label-free phase coordinate that predicts the onset of usable structure and can forecast feature interpretability in sparse autoencoders early in training.