The Geometry of Inference in Transformer Residual Streams
Summary
This paper studies how transformer intermediate residual states become specialized to their own final output states, finding that directional alignment and endpoint rank improve even when Euclidean distance barely changes across six pretrained language models. It offers a high-dimensional model separating norm, alignment, and endpoint geometry, proving that straight-path convergence cannot introduce new competitors, and linking residual geometry to output token rankings.
View Cached Full Text
Cached at: 10/01/26, 08:23 AM
Paper page - The Geometry of Inference in Transformer Residual Streams
Source: https://huggingface.co/papers/2609.37824
Abstract
Transformerlanguagemodelsbuildpredictionsthroughsuccessiveresidualupdates,buthowtheirrepresentationsbecomespecifictoaneventualoutcomeremainsunclear.Westudythisprocessbycomparingintermediateresidualstateswiththeirownfinalstatesandanempiricalbankoffinalstatesfromothercontexts.Acrosssixpretrainedlanguagemodels,theownendpointbecomespreferabletotheaveragealternativeearly,whilemanyindividualendpointsremaincloser.Thesecompetingsetsgenerallyshrinkwithdepth,buttheirmembershipchangesandtheirsurvivingendpointsneednotbecomemoresimilartooneanother.DirectionalalignmentandendpointrankcanthereforeimprovewhileEuclideandistancetothefinalstatechangeslittle.Wedevelopasimplehigh-dimensionalmodelthatseparatestherolesofnorm,alignment,andendpointgeometry,showinghowgradualdirectionalchangescanproducesharpreductionsincompetition.WealsoprovethatastraightpathtowardtheownendpointcannotintroducenewcompetitorsundereitherEuclideanorcosinedistance;observedentriesthusestablishdeparturesfromstraight-lineconvergence.Finally,endpointsassociatedwithlower-rankedoutputtokenstendtoliefartherawayincosinedistanceacrossallstudiedmodels,connectingresidualgeometrytooutputorganization.Together,thesefindingscharacterizeincreasinggeometricspecificityduringtransformerinferenceandexplainwhydistance,competitorcount,andconcentrationofthesurvivingendpointsprovidedistinctviewsofthatprocess.
View arXiv pageView PDFAdd to collection
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.37824 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.37824 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.37824 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Geometric and Behavioral Stratification in Transformer Residual Streams
This paper investigates the residual stream geometry of trained transformers, showing that the prediction direction acts as a privileged anchor that stratifies residual-stream variation into narrow, readout-relevant and broad, computational regions across 18 models. The findings have implications for interpretability and evaluation.
An Analysis of Residual-Stream Geometry Across Transformer Depth
This paper proposes a geometric analysis of transformer residual streams across depth, using relative displacement and orthogonal Procrustes analysis to reveal structured regularities in six instruction-tuned models on code generation and translation tasks.
The Changing Geometry of Grammar: Dimensionality and Neighborhood Reorganization across Transformer Layers
This paper investigates how the grammatical role of tokens shapes the geometry of transformer representations across layers, finding distinct evolution patterns in encoder versus decoder models.
The Geometry of Semantic Space: A Continuous Geometric Framework for the Transformer Architecture
Presents a continuous geometric framework modeling Transformer operations as integro-differential equations on a semantic fiber bundle, validated across multiple architectures.
I Found a Hidden Ratio in Transformers That Predicts Geometric Stability [R]
The article presents a discovered spectral ratio between MLP and attention norms that predicts geometric stability in transformer models, with an optimal range of 0.5–2 to prevent rank collapse.