Tag
Proposes Topologically Regularized Side-Path (TRSP) to mitigate representation collapse in LLMs by balancing spectral trade-offs between mixing efficiency and information capacity, achieving significant gains on long-context benchmarks.
This paper demonstrates that attention sinks, representation collapse, and norm stratification are not unique to attention mechanisms but are general consequences of content-based routing under a norm-blind similarity metric, as shown across multiple architectures including transformers, graph attention, state-space models, and recurrent mixers.
This paper studies representation collapse in sequential post-training of large language models, showing that repeated adaptation stages compress internal representations, reducing plasticity and out-of-domain generalization. The authors propose lightweight interventions to preserve future learnability without sacrificing behavioral gains.