Tag
The paper examines why pretrained language models fail to preserve continuous mixtures of token embeddings in reasoning, identifying architectural distortions and dynamical system effects as primary failure sources.