Tag
This paper studies how transformer intermediate residual states become specialized to their own final output states, finding that directional alignment and endpoint rank improve even when Euclidean distance barely changes across six pretrained language models. It offers a high-dimensional model separating norm, alignment, and endpoint geometry, proving that straight-path convergence cannot introduce new competitors, and linking residual geometry to output token rankings.