The Geometry of Inference in Transformer Residual Streams

Hugging Face Daily Papers Papers

Summary

This paper studies how transformer intermediate residual states become specialized to their own final output states, finding that directional alignment and endpoint rank improve even when Euclidean distance barely changes across six pretrained language models. It offers a high-dimensional model separating norm, alignment, and endpoint geometry, proving that straight-path convergence cannot introduce new competitors, and linking residual geometry to output token rankings.

Transformer language models build predictions through successive residual updates, but how their representations become specific to an eventual outcome remains unclear. We study this process by comparing intermediate residual states with their own final states and an empirical bank of final states from other contexts. Across six pretrained language models, the own endpoint becomes preferable to the average alternative early, while many individual endpoints remain closer. These competing sets generally shrink with depth, but their membership changes and their surviving endpoints need not become more similar to one another. Directional alignment and endpoint rank can therefore improve while Euclidean distance to the final state changes little. We develop a simple high-dimensional model that separates the roles of norm, alignment, and endpoint geometry, showing how gradual directional changes can produce sharp reductions in competition. We also prove that a straight path toward the own endpoint cannot introduce new competitors under either Euclidean or cosine distance; observed entries thus establish departures from straight-line convergence. Finally, endpoints associated with lower-ranked output tokens tend to lie farther away in cosine distance across all studied models, connecting residual geometry to output organization. Together, these findings characterize increasing geometric specificity during transformer inference and explain why distance, competitor count, and concentration of the surviving endpoints provide distinct views of that process.
Original Article
View Cached Full Text

Cached at: 10/01/26, 08:23 AM

Paper page - The Geometry of Inference in Transformer Residual Streams

Source: https://huggingface.co/papers/2609.37824

Abstract

Transformerlanguagemodelsbuildpredictionsthroughsuccessiveresidualupdates,buthowtheirrepresentationsbecomespecifictoaneventualoutcomeremainsunclear.Westudythisprocessbycomparingintermediateresidualstateswiththeirownfinalstatesandanempiricalbankoffinalstatesfromothercontexts.Acrosssixpretrainedlanguagemodels,theownendpointbecomespreferabletotheaveragealternativeearly,whilemanyindividualendpointsremaincloser.Thesecompetingsetsgenerallyshrinkwithdepth,buttheirmembershipchangesandtheirsurvivingendpointsneednotbecomemoresimilartooneanother.DirectionalalignmentandendpointrankcanthereforeimprovewhileEuclideandistancetothefinalstatechangeslittle.Wedevelopasimplehigh-dimensionalmodelthatseparatestherolesofnorm,alignment,andendpointgeometry,showinghowgradualdirectionalchangescanproducesharpreductionsincompetition.WealsoprovethatastraightpathtowardtheownendpointcannotintroducenewcompetitorsundereitherEuclideanorcosinedistance;observedentriesthusestablishdeparturesfromstraight-lineconvergence.Finally,endpointsassociatedwithlower-rankedoutputtokenstendtoliefartherawayincosinedistanceacrossallstudiedmodels,connectingresidualgeometrytooutputorganization.Together,thesefindingscharacterizeincreasinggeometricspecificityduringtransformerinferenceandexplainwhydistance,competitorcount,andconcentrationofthesurvivingendpointsprovidedistinctviewsofthatprocess.

View arXiv pageView PDFAdd to collection

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.37824 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.37824 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.37824 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Geometric and Behavioral Stratification in Transformer Residual Streams

arXiv cs.LG

This paper investigates the residual stream geometry of trained transformers, showing that the prediction direction acts as a privileged anchor that stratifies residual-stream variation into narrow, readout-relevant and broad, computational regions across 18 models. The findings have implications for interpretability and evaluation.

An Analysis of Residual-Stream Geometry Across Transformer Depth

arXiv cs.LG

This paper proposes a geometric analysis of transformer residual streams across depth, using relative displacement and orthogonal Procrustes analysis to reveal structured regularities in six instruction-tuned models on code generation and translation tasks.