Temporal Preference Concepts and their Functions in a Large Language Model
Summary
This paper causally localizes a subgraph for temporal preference in a distilled LLM, finding that the model discounts the future less steeply than humans and that steering vectors can shift temporal preference, highlighting the need for explicit control mechanisms.
View Cached Full Text
Cached at: 06/05/26, 08:09 AM
# Temporal Preference Concepts and their Functions in a Large Language Model
Source: [https://arxiv.org/html/2606.05194](https://arxiv.org/html/2606.05194)
Ian Rios\-Sialer AISC San Francisco, USA &Shantanu Darveshi AISC Mumbai, India &Shuai Jiang† AISC Albuquerque, USA &Avigya Paudel† SPAR, AISC New York, USA &Anastasiia Pronina† AISC Munich, Germany &Ipshita Bandyopadhyay SPAR Bangalore, India &Justin Shenk SPAR, AISC Berlin, Germany Correspondence to:ian@unrulyabstractions\.com\. Equal contribution\. Listed alphabetically\. AISC = AI Safety Camp\. SPAR = Supervised Program for Alignment Research\.
###### Abstract
Large Language Models \(LLMs\) are increasingly being deployed to make decisions that require trading off near\-term gains against long\-term consequences, yet little is known about how they internally represent or resolve these tradeoffs\. In this work, we causally localize an underlying subgraph for temporal preference in a distilled LLM \(Qwen3\-4B\-Instruct\-2507\), identifying mid\-to\-upper\-layer nodes through converging evidence from gradient\-based attribution and activation patching\. We find that the geometry of time horizon is encoded in the residual stream at the expected localized layers\. A behavioral analysis reveals that unintervened LLMs discount the future several times less steeply than humans, yet this preference is unstable across contexts, motivating explicit control rather than implicit reliance on training\. Finally, we find suggestive evidence that steering vectors can shift temporal preference\. Our work demonstrates how mechanistic interpretability can bring us closer to reliable control over how LLMs plan and reason\.
## 1Introduction
> All your life, you wait for the propitious time\. Then the propitious time reveals itself as action taken\. Louise Glück,Landscape
Large Language Models \(LLMs\) appear to hold preferences mediated by abstract concepts such as time\. Impatience and paralysis both push humans into bad decisions\. LLMs risk failing in similar ways\. For now, the consequences have been limited\. No coding agent[cutting corners](https://github.com/anthropics/claude-code/issues/42796)\[[7](https://arxiv.org/html/2606.05194#bib.bib7)\]has caused a catastrophe yet\. But the stakes are rising quickly\. In early 2026, the U\.S\. Department of Defense and Anthropic publicly clashed over a range of sensitive issues\[[2](https://arxiv.org/html/2606.05194#bib.bib2),[49](https://arxiv.org/html/2606.05194#bib.bib49),[34](https://arxiv.org/html/2606.05194#bib.bib34)\], including whether LLMs should be allowed to autonomously operate weapons\. In high\-stakes scenarios, autonomous agents\[[40](https://arxiv.org/html/2606.05194#bib.bib40),[71](https://arxiv.org/html/2606.05194#bib.bib71)\]would need to trade off short\-term gains against long\-term effects\[[26](https://arxiv.org/html/2606.05194#bib.bib26)\]\. When choosing among alternatives, the decision often depends on the temporal scope used to evaluate the consequences\[[20](https://arxiv.org/html/2606.05194#bib.bib20),[98](https://arxiv.org/html/2606.05194#bib.bib98),[32](https://arxiv.org/html/2606.05194#bib.bib32)\]\. Temporal preference is indeed fundamental for planning\[[1](https://arxiv.org/html/2606.05194#bib.bib1)\], but also for cooperation\[[103](https://arxiv.org/html/2606.05194#bib.bib103),[55](https://arxiv.org/html/2606.05194#bib.bib55)\]and trust\[[27](https://arxiv.org/html/2606.05194#bib.bib27)\], where agents must bear present costs for future collective benefit\[[83](https://arxiv.org/html/2606.05194#bib.bib83)\]\. These intertemporal tradeoffs grow even more consequential in the context of Artificial General Intelligence \(AGI\)\. A myopic system\[[76](https://arxiv.org/html/2606.05194#bib.bib76)\]poses different risks than one capable of scheming across long horizons\[[69](https://arxiv.org/html/2606.05194#bib.bib69),[81](https://arxiv.org/html/2606.05194#bib.bib81)\]\. Detecting and maintaining control\[[37](https://arxiv.org/html/2606.05194#bib.bib37)\]over these capabilities while that is still tractable motivates our inquiry\.Where and how does an LLM encode temporal preference?
Previous work has investigated the existence of temporal representations\[[38](https://arxiv.org/html/2606.05194#bib.bib38)\], characterized the economic behavior of LLMs\[[16](https://arxiv.org/html/2606.05194#bib.bib16),[48](https://arxiv.org/html/2606.05194#bib.bib48)\], and even shown that risk preference can be steered\[[116](https://arxiv.org/html/2606.05194#bib.bib116)\]\. Yet no work has identified*where*temporal preference lives inside an LLM, how it is geometrically organized, or how to control it through targeted intervention\.
Using Mechanistic Interpretability \(MI\)\[[6](https://arxiv.org/html/2606.05194#bib.bib6)\]techniques, we isolate the components that are causally responsible for temporal preference and show how activation\-space representations evolve through them\. This offers a geometric perspective on how interventions function to shift temporal preference, even in general open\-ended generation tasks\.
L21L24L31Localization151617181920212223242526272829303132333435LayerProbingAttr\. \(contr\.\)Attr\. \(param\.\)Causal \(param\.\)Causal \(class\.\)Error char\.AttnMLPResidProbeError
Steering−4\-4−2\-20\+2\+2temporal orientation scoreshort\-termlong\-termL21L22L26steeringlayersα=−50\\alpha\\\!=\\\!\-50α=\+50\\alpha\\\!=\\\!\+50α=−50\\alpha\\\!=\\\!\-50α=\+50\\alpha\\\!=\\\!\+50
Geometry 
Figure 1:\(Top\-left\) Six rows summarize where the temporal\-preference signal sits across the network\. Five localization methods \(probing, attributional contrastive, attributional parametric, causal parametric, causal classification\) converge on a subgraph in layers 17–35; darker shading indicates stronger signal\. The bottom row \(Error char\.\) shows that an unrelated meta\-cognitive variable, cumulative reasoning reliability, is also decodable above 95% across L19–L31 within the same subgraph, indicating the localized region encodes more than just temporal preference \(Section[5\.1](https://arxiv.org/html/2606.05194#S5.SS1);[Appendix R](https://arxiv.org/html/2606.05194#A18)\)\. \(Top\-right\) Time horizon geometry within the identified subgraph \(Section[5\.2](https://arxiv.org/html/2606.05194#S5.SS2)\)\. \(Bottom\) CAA steering shifts temporal preference bidirectionally at layers 19–22 but weakly at the best probing layer \(L26\), illustrating the probing–steering dissociation \(Section[5\.3](https://arxiv.org/html/2606.05194#S5.SS3)\)\.We focus onQwen3\-4B\-Instruct\-2507: its non\-thinking\-only operation keeps computation inside a fixed prompt template \(enabling token\-aligned attribution and patching\), and its latent preferences are stable under prompt perturbations, trading breadth across model families for depth in tracing a concept from localization through geometry to intervention\. Our methodology integrates four localization pipelines plus dedicated behavioral and steering instruments that differ in prompting paradigm, localization technique, scale, and resolution\. Because these pipelines approach localization from fundamentally different angles, their independent convergence on the same subgraph components strongly suggests that our findings reflect the genuine model structure\.
Our work establishes thattemporal preference is localizablewithin LLMs and, given the behavioral inconsistencies we observe,should be explicitly controlledrather than left to emerge implicitly from training\. The convergence of multiple separate paradigms on the same subgraph demonstrates the value ofcomplementary experimental approachesto validate mechanistic claims\. Finally, while current interpretability methods are well\-suited for binary contrastive concepts \(truthfulness\[[67](https://arxiv.org/html/2606.05194#bib.bib67)\], refusal\[[3](https://arxiv.org/html/2606.05194#bib.bib3)\], sycophancy\[[78](https://arxiv.org/html/2606.05194#bib.bib78)\]\), this work takes initial steps towardsteering dimensional conceptssuch as time, uncertainty, or risk preference, an underexplored area that warrants further development\.
We summarize our contributions as follows:
*Causal Localization of Temporal Preference*\(Section[5\.1](https://arxiv.org/html/2606.05194#S5.SS1)\)
- •We provide causal and attributional localization of a temporal\-preference subgraph inQwen3\-4B\-Instruct\-2507using complementary methods, including standard attribution patching, activation patching, linear probing, and CAA \(Section[3](https://arxiv.org/html/2606.05194#S3)\), and show that its L24 attention layer is also recruited for categorical horizon inference \([Appendix K](https://arxiv.org/html/2606.05194#A11)\)\.
*Characterization of Temporal Preference*\(Section[5\.2](https://arxiv.org/html/2606.05194#S5.SS2)\)
- •We show that time horizon has non\-linear geometry within the subgraph \(Section[5\.2](https://arxiv.org/html/2606.05194#S5.SS2)\)\.
- •We identify the user\-to\-assistant turn boundary \(the change\-of\-turn token sequence<\|im\_end\|\>→\\to\\n→\\to<\|im\_start\|\>→\\toassistant\) as the site where attention collapses the continuous horizon manifold into a committed binary preference \(Section[5\.2](https://arxiv.org/html/2606.05194#S5.SS2)\)\.
- •We provide a behavioral analysis \(Section[5\.2](https://arxiv.org/html/2606.05194#S5.SS2)\) that shows that unintervened LLMs behave very differently from humans, suggesting that the implicit time preference is inconsistent between contexts\.
*Steering of Temporal Preference*\(Section[5\.3](https://arxiv.org/html/2606.05194#S5.SS3)\)
- •We show successful interventions that change temporal preference and interpret them through our characterizations, via the steering methodology detailed in[Appendix AB](https://arxiv.org/html/2606.05194#A28)\.
Our methodology combines multiple independent methods, datasets, and resolutions, each approaching the same subgraph from a different angle, to build converging evidence for how temporal preference is mechanistically implemented\.
## 2Background
We refer totemporal preferenceas the degree to which an agent values outcomes differently depending on when they occur\[[15](https://arxiv.org/html/2606.05194#bib.bib15),[103](https://arxiv.org/html/2606.05194#bib.bib103)\]\. We use the termtime horizonto denote the future moment at which outcomes are evaluated against an objective\[[22](https://arxiv.org/html/2606.05194#bib.bib22)\]\. Because future events do not affect past outcomes, the time horizon also acts as aconstrainton planning\[[88](https://arxiv.org/html/2606.05194#bib.bib88)\]\. It would beinstrumentally\[[58](https://arxiv.org/html/2606.05194#bib.bib58)\]ormeans\-end\[[8](https://arxiv.org/html/2606.05194#bib.bib8)\]incoherentfor the agent to choose actions that are incapable of causing effects by the specified deadline\. Then, thetemporal scopeis the bounded interval of time over which an agent weighs the results according to its preference111Although not explored in this work, when rewards areperishable\[[4](https://arxiv.org/html/2606.05194#bib.bib4)\], the temporal scope is also bounded below by theretroactive reach\[[50](https://arxiv.org/html/2606.05194#bib.bib50),[70](https://arxiv.org/html/2606.05194#bib.bib70)\]\.\.
ValueTemporal PreferenceRewardV\(t\)V\(t\) λ\(t\)\\lambda\(t\) r\(t\)r\(t\) TemporalScopett =Time Horizontt ×tt
Figure 2:The time horizon specifies when the consequences of a decision are assessed\. The temporal scope is then bounded above by the time horizon\.The empirical handle on temporal preference isintertemporal choice: a decision between options that differ in both reward and delay\[[29](https://arxiv.org/html/2606.05194#bib.bib29),[36](https://arxiv.org/html/2606.05194#bib.bib36)\]\. Each optioniiis a tuple\(ri,ti\)\(r\_\{i\},t\_\{i\}\)of rewardri∈ℝ\+r\_\{i\}\\in\\mathbb\{R\}^\{\+\}and delayti∈ℝ\+t\_\{i\}\\in\\mathbb\{R\}^\{\+\}\. Its subjective value is the reward scaled by a delay\-dependent temporal\-preference weightλ\\lambda:
Vi=λ\(ti\)⋅ri,V\_\{i\}\\;=\\;\\lambda\(t\_\{i\}\)\\cdot r\_\{i\}\\,,\(1\)and, given a pair of options\{A,B\}\\\{A,B\\\}, an agent with preferenceλ\\lambdaselects
i∗=argmaxi∈\{A,B\}λ\(ti\)⋅ri\.i^\{\*\}\\;=\\;\\operatorname\*\{arg\\,max\}\_\{i\\in\\\{A,B\\\}\}\\;\\lambda\(t\_\{i\}\)\\cdot r\_\{i\}\\,\.\(2\)Humans are typically modeled with hyperbolic discount functions\[[29](https://arxiv.org/html/2606.05194#bib.bib29),[68](https://arxiv.org/html/2606.05194#bib.bib68)\]; for LLMs, whether the same functional form fits is an open empirical question\[[68](https://arxiv.org/html/2606.05194#bib.bib68)\]\. Characterizing LLM behavior therefore requires fittingλ\\lambdavia regression, assessing its stability across contexts, and benchmarking the resulting preferences against human intertemporal choice\.
In humans, these concepts have localized neural representations\[[51](https://arxiv.org/html/2606.05194#bib.bib51)\]that predict behavior\[[93](https://arxiv.org/html/2606.05194#bib.bib93)\], respond causally to intervention\[[28](https://arxiv.org/html/2606.05194#bib.bib28)\], and exhibit internal organization interpretable as a*functional role*\[[17](https://arxiv.org/html/2606.05194#bib.bib17)\]\.222Some authors argue that concepts are best modeled by geometric or topological spaces\[[30](https://arxiv.org/html/2606.05194#bib.bib30)\], a perspective that resonates with our geometric analysis of temporal representations in activation space\.Our work asks whether temporal preference exists in an LLM in an analogous way: localized, predictive, causally efficacious, and geometrically organized\.
### 2\.1Locate and characterize, then steer
Our pipeline engages the target concept in three complementary modes\. To*locate*the subgraph responsible for temporal preference, we pair*causal*intervention, activation patching\[[44](https://arxiv.org/html/2606.05194#bib.bib44)\]in do\-calculus notation\[[82](https://arxiv.org/html/2606.05194#bib.bib82)\], with cheaper*attributional*proxies that scale across inputs: gradient\-based EAP\-IG\[[43](https://arxiv.org/html/2606.05194#bib.bib43),[6](https://arxiv.org/html/2606.05194#bib.bib6)\]and linear probes\[[74](https://arxiv.org/html/2606.05194#bib.bib74),[56](https://arxiv.org/html/2606.05194#bib.bib56)\]that surface where the concept linearly emerges\. To*characterize*how the localized components encode horizon information, we apply PCA\[[95](https://arxiv.org/html/2606.05194#bib.bib95)\]inside the subgraph; prior work warns that concept geometry is often non\-linear\[[24](https://arxiv.org/html/2606.05194#bib.bib24),[72](https://arxiv.org/html/2606.05194#bib.bib72),[39](https://arxiv.org/html/2606.05194#bib.bib39)\]and can drift across generation\[[59](https://arxiv.org/html/2606.05194#bib.bib59)\], so a single global direction rarely tells the whole story\. Only after we have located and characterized the subgraph do we*steer*: we inject a probe\-derived vector\[[100](https://arxiv.org/html/2606.05194#bib.bib100),[78](https://arxiv.org/html/2606.05194#bib.bib78)\]at inference time; localization is not strictly required but tightens precision, shrinks magnitudes, and reduces side effects\[[114](https://arxiv.org/html/2606.05194#bib.bib114)\]\. Full definitions are in[Appendix A](https://arxiv.org/html/2606.05194#A1)\.
### 2\.2Related work
Four strands of work frame this paper: \(i\)*temporal representation*, showing that LLMs encode time geometrically\[[38](https://arxiv.org/html/2606.05194#bib.bib38),[24](https://arxiv.org/html/2606.05194#bib.bib24),[53](https://arxiv.org/html/2606.05194#bib.bib53),[39](https://arxiv.org/html/2606.05194#bib.bib39)\]as locally linear features on globally curved manifolds\[[72](https://arxiv.org/html/2606.05194#bib.bib72),[80](https://arxiv.org/html/2606.05194#bib.bib80)\]; \(ii\)*temporal reasoning and planning*, where models fail despite the geometric encoding\[[105](https://arxiv.org/html/2606.05194#bib.bib105),[91](https://arxiv.org/html/2606.05194#bib.bib91),[106](https://arxiv.org/html/2606.05194#bib.bib106)\]; \(iii\)*LLM economic behavior*, reproducing human biases\[[48](https://arxiv.org/html/2606.05194#bib.bib48),[16](https://arxiv.org/html/2606.05194#bib.bib16),[104](https://arxiv.org/html/2606.05194#bib.bib104)\]with entangled risk/time preferences\[[116](https://arxiv.org/html/2606.05194#bib.bib116),[73](https://arxiv.org/html/2606.05194#bib.bib73),[68](https://arxiv.org/html/2606.05194#bib.bib68)\]; and \(iv\)*steering advancements*, from activation addition\[[100](https://arxiv.org/html/2606.05194#bib.bib100),[78](https://arxiv.org/html/2606.05194#bib.bib78)\]through sparse dictionaries\[[18](https://arxiv.org/html/2606.05194#bib.bib18)\]to geometry\-aware methods\[[101](https://arxiv.org/html/2606.05194#bib.bib101),[87](https://arxiv.org/html/2606.05194#bib.bib87),[64](https://arxiv.org/html/2606.05194#bib.bib64),[84](https://arxiv.org/html/2606.05194#bib.bib84)\], with known failure modes at large\|α\|\|\\alpha\|\[[108](https://arxiv.org/html/2606.05194#bib.bib108),[9](https://arxiv.org/html/2606.05194#bib.bib9)\]\. No prior work has localized a subgraph functionally responsible for temporal preference, characterized the geometry of the causal representation, or steered along it\. Full discussion in[Appendix B](https://arxiv.org/html/2606.05194#A2)\.
## 3Methodology
Our methodology follows three stages:*localize*the temporal\-preference subgraph,*characterize*its representations, and*intervene*to steer it\. Localization pairs*wide attribution*\(contrastive A/B prompts×\\timesEAP\-IG and linear probing, cheap to sweep across hundreds of components\) with*targeted intervention*\(parametric prompts with explicit horizons×\\timesactivation patching, expensive but causal\); the two paradigms converge on the same subgraph, which is the basis for our localization claim\. A complementary classification pipeline \(IOI\-style prompts×\\timesdirectional patching\) tests whether the same subgraph generalizes from valuation to categorical horizon inference\. Characterization applies PCA inside that subgraph to examine how explicit horizons organize the activation manifold and whether latent \(no\-horizon\) preferences align with that geometry, paired with two behavioral instruments \(Kirby MCQ\-27 and a 30\-model investment\-coherence questionnaire\) that test whether the geometry actually drives choice\. Intervention uses Contrastive Activation Addition with a probe\-derived steering vector, swept across layers and magnitudes to test for a probing\-steering dissociation\. Full per\-pipeline protocols, dataset construction, metric definitions, and the reader’s guide to Part 4 are in the Methodology Summary \([Appendix C](https://arxiv.org/html/2606.05194#A3)\)\.
## 4Experimental setup
We focus on a single model,Qwen3\-4B\-Instruct\-2507, a mode\-specialized non\-thinking refresh ofQwen3\-4B\[[85](https://arxiv.org/html/2606.05194#bib.bib85),[111](https://arxiv.org/html/2606.05194#bib.bib111)\]\. We chose this model because operating in non\-thinking mode keeps all “cognition” inside a fixed prompt template \(no<think\>block perturbs token positions\), which is what the activation\-patching and attribution pipelines need to align clean and corrupted runs; because its latent preference is stable under minor prompt perturbations at a scale where similar\-sized models drift; and because it is small enough for repeated attribution and intervention sweeps\. The pipeline operates on four dataset paradigms: minimally\-framed A/B prompts \(500 explicit \+ 500 implicit pairs\), highly\-formatted parametric prompts with explicit time horizons \(4,588 prompts\), IOI\-style classification prompts \(160 short/long pairs\), and behavioral questionnaires \(Kirby MCQ\-27 plus a 960\-prompt investment\-coherence instrument run on 30 models\)\. All experiments fit on a MacBook Pro \(M4 Max, 48 GB\) except for the causal classification pipeline, which requires 79 GB \([Appendix X](https://arxiv.org/html/2606.05194#A24)\), and are reproducible end\-to\-end within two weeks\. Full configurations, dataset construction, and model\-selection rationale are in[Appendix D](https://arxiv.org/html/2606.05194#A4)\.
## 5Results
### 5\.1Where is temporal preference for the LLM?
Four localization methods converge on layers 17–35 \(Figure[1](https://arxiv.org/html/2606.05194#S1.F1), top\-left;[Appendix L](https://arxiv.org/html/2606.05194#A12)\)\. L24 attention is flagged by all three non\-probing methods\. MLP effects concentrate in L31–L35 across attributional contrastive and causal parametric patching \(the causal classification run is attention\-dominated\)\. Probes peak at layer 26 \(99\.2%;[Appendix G](https://arxiv.org/html/2606.05194#A7)\)\. Activation patching ranks the four highest\-importance components asL24\_attn,L35\_mlp,L31\_mlp, andL21\_attn, separated from the remaining components in effect size \(Figure[3](https://arxiv.org/html/2606.05194#S5.F3)\)\. Gradient\-based attribution on the contrastive prompts independently concentrates top\-kkattribution mass at the same layers, with a sharp peak at L24 and a secondary peak at L31–L35 that is robust acrosskk\(Figure[3](https://arxiv.org/html/2606.05194#S5.F3), bottom\-left inset\)\.
Causal parametric \(parametric prompts\) 
Attributional contrastive \(contrastive prompts\) 
Causal classification \(contrastive prompts\) 
Figure 3:Localization evidence converges on the same components\.Top\-left:top components ranked by mean causal effect under the parametric paradigm \(n=71n=71pairs\), withL24\_attnthe only attention component above 0\.5 noising disruption\.Right:top\-20 components from causal classification \(n=160n=160pairs\)\.Bottom\-left:EAP\-IG top\-kkattribution mass per layer on contrastive prompts \(short\-term and long\-term flips\), peaking sharply at L24 with a secondary peak at L31–L35 and stable acrosskk\([Appendix H](https://arxiv.org/html/2606.05194#A8)\)\. All three methods converge onL24\_attn,L21\_attn,L35\_mlp, andL31\_mlpas the highest\-effect components, consistent with the cross\-method agreement in Table L\.1 \([Appendix J](https://arxiv.org/html/2606.05194#A10),[Appendix K](https://arxiv.org/html/2606.05194#A11)\)\.
### 5\.2What is the LLM’s temporal preference like?
Geometry\.Time horizons form ordinal clusters \(seconds to centuries\) in activation space, but the direction encoding them is unstable across prompt positions until the user\-to\-assistant turn boundary, where attention collapses the continuous horizon representation into a binary preference signal that sharpens from heavy overlap at the beginning of the turn change \(<\|im\_end\|\>\) to clean separation by the end of the turn change \(assistant\) \(Figures[4](https://arxiv.org/html/2606.05194#S5.F4)and[5](https://arxiv.org/html/2606.05194#S5.F5);[Appendix M](https://arxiv.org/html/2606.05194#A13)\)\.

<\|im\_end\|\>

\\n

<\|im\_start\|\>

assistant
Figure 4:resid\_postat the four turn\-transition tokens \(<\|im\_end\|\>,\\n,<\|im\_start\|\>,assistant\), colored by preference \(orange = long, blue = short\)\. The preference signal sharpens from heavy overlap at the beginning of the turn change to clean separation by its end \([Appendix M](https://arxiv.org/html/2606.05194#A13)\)\.<\|im\_end\|\>\\n<\|im\_start\|\>assistant
Figure 5:Layer\-31 PCA at the four turn\-transition tokens \(rows:<\|im\_end\|\>,\\n,<\|im\_start\|\>,assistant\), colored by preference \(left column: orange = long, blue = short\) and by time horizon \(right column: gradient from seconds to centuries; gray = no horizon\)\. At<\|im\_end\|\>the preference is not yet linearly separable and the no\-horizon samples sit off the time\-horizon manifold; byassistantthe no\-horizon samples have aligned to the manifold and preference clusters are cleanly separated, identifying the change\-of\-turn token sequence as the site where horizon is collapsed into a committed binary preference \([Appendix M](https://arxiv.org/html/2606.05194#A13)\)\.Discounting\.LLM discount rates \(k<0\.005k<0\.005\) are 3–8×\\timesbelow human controls \(k≈0\.013k\\approx 0\.013\); chain\-of\-thought amplifies present bias in the 4B model but produces paradoxical patience in the 8B model \([Appendix O](https://arxiv.org/html/2606.05194#A15)\)\.
Coherence\.We test whetherQwen3\-4B\-Instruct\-2507makes instrumentally coherent choices \(Section[2](https://arxiv.org/html/2606.05194#S2)\) on 960 investment prompts offering $20K in 6 months vs\. a long\-term option \($100K, $300K, or $500K in 10 years\) \([Appendix P](https://arxiv.org/html/2606.05194#A16)\)\. Coherence is defined in the 1–5y reasoning zone, where only the 6\-month option can deliver within the deadline\. Our model does not meet this bar: it picks the undeliverable long\-term option 47–53% of the time, and a deep dive shows this is positional polarization rather than uncertainty; reward size and label format are effectively inert \([P\.4\.1](https://arxiv.org/html/2606.05194#A16.SS4.SSS1)\)\. Benchmarking against 29 other models confirms the failure is not idiosyncratic: only frontier API models \(Claudefamily,Gemini 2\.5 Pro,GPT\-5\.4,GPT\-5\.4 Mini,o3\) reach 95–100% coherence, and the smallerClaudevariants do so via a binary heuristic that collapses at longer horizons \(Figure[6](https://arxiv.org/html/2606.05194#S5.F6)\)\.


Figure 6:Qwen3\-4B\-Instruct\-2507\(yellow\) holds a stable temporal preference under presentation\-order swaps in the no\-horizon condition \(left\), but does not produce coherent temporal reasoning when given an explicit deadline \(right; star marks our target at 50%, far below the 90% coherent threshold\)\. Full per\-model breakdowns in[Appendix P](https://arxiv.org/html/2606.05194#A16)\.Generality\.Probing the same layers and turn\-transition tokens for an unrelated meta\-cognitive variable \(the cumulative reliability of a multi\-step reasoning chain\) recovers a 95% decodability plateau over L19 to L31\. The error\-reliability direction is linearly orthogonal to the temporal\-preference direction yet rides the same curved manifold within the localized subgraph \([Appendix R](https://arxiv.org/html/2606.05194#A18)\)\. The gap between rich internal representation and weak behavioral influence is therefore an architectural pattern of this subgraph, not specific to temporal preference\.
### 5\.3Could temporal preference be controlled?
CAA steering at layers 19–22 shifts temporal preference bidirectionally:∼\\sim3\.4×3\.4\\timeshigher relative odds for long\- vs short\-term completion at layer 22,α=50\\alpha=50\(exp\(1\.22\)≈3\.39\\exp\(1\.22\)\\approx 3\.39; Figure[1](https://arxiv.org/html/2606.05194#S1.F1), bottom; Table[1](https://arxiv.org/html/2606.05194#S5.T1);[Appendix S](https://arxiv.org/html/2606.05194#A19)\)\. Open\-ended generation shifts from triage framing \(α<0\\alpha<0\) to strategic framing \(α\>0\\alpha\>0\)\. Output quality degrades at\|α\|=60\|\\alpha\|=60, suggesting the linear steering vector exceeds the locally linear regime of the curved manifold\. The optimal steering layers \(19–22\) sit 4–7 layers below the best probing layer \(26\), a probing–steering dissociation consistent with the localized subgraph\.
α\\alphaL19L20L21L22L23L24L25200\.720\.690\.720\.680\.560\.510\.46300\.940\.910\.970\.930\.760\.680\.60401\.131\.101\.201\.170\.960\.840\.74501\.301\.271\.391\.391\.141\.010\.87Table 1:Forced\-choice scoreS\(α,l\)S\(\\alpha,l\)for the layer×\\timessteering\-coefficient sweep\. Baseline \(no steering\):S=0\.17S=0\.17\. Best configuration: layer 22 withα=50\\alpha=50,S=1\.39S=1\.39\(lift\+1\.22\+1\.22,exp\(1\.22\)≈3\.39\\exp\(1\.22\)\\approx 3\.39odds\-ratio change\)\. Layers 19–22 form the behavioral sweet spot; effectiveness drops sharply from L23 onward\. Full sweep includingα∈\{1,2,5,10\}\\alpha\\in\\\{1,2,5,10\\\}in[Appendix S](https://arxiv.org/html/2606.05194#A19)\.
## 6Discussion
Whether an LLM operates under the right temporal preference is ultimately an alignment question\. Post\-training methods may suffice for routine use, but high\-stakes settings call for stronger guarantees\. We believe activation geometry can serve as a fail\-safe here: localize the subgraph relevant to a specific task, characterize the geometry of the temporal concept within it, and then, at inference time, monitor the model’s internal representations against that manifold and intervene if they drift\. This perspective frames interpretability not only as a diagnostic tool but as infrastructure for runtime alignment\.
## 7Limitations and future work
Our work is a starting point on an entangled concept\. The main open directions are: finer\-grained circuit tracing to move from subgraph\-level attribution to atomic components and information flow; generalization beyond the single financial task and the single target model \(Qwen3\-4B\-Instruct\-2507\) to other domains, model scales, and thinking vs\. non\-thinking variants; richer parameterization along reward, risk, role, and domain axes to map the full intertemporal choice space and its interactions with adjacent concepts such as emotion and urgency; multi\-turn and agentic settings where temporal representations may shift across turns; and non\-linear steering methods that respect the curvature of the underlying manifold and avoid the output\-quality degradation we observe in linear CAA at high\|α\|\|\\alpha\|\. Full discussion is in[Appendix F](https://arxiv.org/html/2606.05194#A6)\.
## 8Conclusion
We show that temporal preference in LLMs is localizable, that we can characterize its representational geometry, and that targeted activation interventions can shift it bidirectionally\. Our work highlights the value of using complementary paradigms\. More broadly, while the literature has focused on identifying contrastive binary concepts, this work offers initial steps toward decomposing dimensional concepts such as time\.
## Acknowledgments and Disclosure of Funding
We thank theAI Safety Camp \(AISC\)and theSupervised Program for Alignment Research \(SPAR\)for providing the collaborative structure, mentorship, and computational resources that made this project possible\. AISC’s cohort\-based research model brought the authors together and sustained the multi\-month investigation; SPAR’s program provided additional mentorship and connected contributors across timezones\.
## References
- Ameriks et al\. \[2003\]John Ameriks, Andrew Caplin, and John Leahy\.Wealth accumulation and the propensity to plan\.*The Quarterly Journal of Economics*, 118\(3\):1007–1047, 2003\.
- Amodei \[2026\]Dario Amodei\.Statement from dario amodei on our discussions with the department of war, 2026\.URL[https://www\.anthropic\.com/news/statement\-department\-of\-war](https://www.anthropic.com/news/statement-department-of-war)\.
- Arditi et al\. \[2024\]Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Panickssery, Wes Gurnee, and Neel Nanda\.Refusal in language models is mediated by a single direction\.In A\. Globerson, L\. Mackey, D\. Belgrave, A\. Fan, U\. Paquet, J\. Tomczak, and C\. Zhang, editors,*Advances in Neural Information Processing Systems*, volume 37, pages 136037–136083\. Curran Associates, Inc\., 2024\.doi:10\.52202/079017\-4322\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2024/file/f545448535dfde4f9786555403ab7c49\-Paper\-Conference\.pdf](https://proceedings.neurips.cc/paper_files/paper/2024/file/f545448535dfde4f9786555403ab7c49-Paper-Conference.pdf)\.
- Arrow et al\. \[1951\]Kenneth J\. Arrow, Theodore Harris, and Jacob Marschak\.Optimal inventory policy\.*Econometrica*, 19\(3\):250–272, 1951\.doi:10\.2307/1906813\.
- Bartoszcze et al\. \[2025\]Lukasz Bartoszcze, Sarthak Munshi, Bryan Sukidi, Jennifer Yen, Zejia Yang, David Williams\-King, Linh Le, Kosi Asuzu, and Carsten Maple\.Representation engineering for large\-language models: Survey and research challenges, 2025\.URL[https://arxiv\.org/abs/2502\.17601](https://arxiv.org/abs/2502.17601)\.
- Bereska and Gavves \[2024\]Leonard Bereska and Efstratios Gavves\.Mechanistic interpretability for ai safety – a review, 2024\.URL[https://arxiv\.org/abs/2404\.14082](https://arxiv.org/abs/2404.14082)\.
- Betley et al\. \[2025\]Jan Betley, Daniel Tan, Niels Warncke, Anna Sztyber\-Betley, Xuchan Bao, Martín Soto, Nathan Labenz, and Owain Evans\.Emergent misalignment: Narrow finetuning can produce broadly misaligned llms, 2025\.URL[https://arxiv\.org/abs/2502\.17424](https://arxiv.org/abs/2502.17424)\.
- Bratman \[1981\]Michael E\. Bratman\.Intention and means\-end reasoning\.*The Philosophical Review*, 90\(2\):252–265, 1981\.
- Braun et al\. \[2025\]Joschka Braun, Carsten Eickhoff, David Krueger, Seyed Ali Bahrainian, and Dmitrii Krasheninnikov\.Understanding \(un\)reliability of steering vectors in language models, 2025\.URL[https://arxiv\.org/abs/2505\.22637](https://arxiv.org/abs/2505.22637)\.
- Cacioli \[2026a\]Jon\-Paul Cacioli\.Categorical perception in large language model hidden states: Structural warping at digit\-count boundaries, 2026a\.URL[https://arxiv\.org/abs/2603\.28258](https://arxiv.org/abs/2603.28258)\.
- Cacioli \[2026b\]Jon\-Paul Cacioli\.Weber’s law in transformer magnitude representations: Efficient coding, representational geometry, and psychophysical laws in language models, 2026b\.URL[https://arxiv\.org/abs/2603\.20642](https://arxiv.org/abs/2603.20642)\.
- Cao et al\. \[2025\]Pengfei Cao, Tianyi Men, Wencan Liu, Jingwen Zhang, Xuzhao Li, Xixun Lin, Dianbo Sui, Yanan Cao, Kang Liu, and Jun Zhao\.Large language models for planning: A comprehensive and systematic survey, 2025\.URL[https://arxiv\.org/abs/2505\.19683](https://arxiv.org/abs/2505.19683)\.
- Chen et al\. \[2025\]Hui Chen, Antoine Didisheim, Mohammad Pourmohammadi, Luciano Somoza, and Hanqing Tian\.A financial brain scan of the LLM\.*arXiv preprint arXiv:2508\.21285*, 2025\.
- Cheng et al\. \[2025\]Yize Cheng, Arshia Soltani Moakhar, Chenrui Fan, Parsa Hosseini, Kazem Faghih, Zahra Sodagar, Wenxiao Wang, and Soheil Feizi\.Your LLM agents are temporally blind: The misalignment between tool use decisions and human time perception\.*arXiv preprint arXiv:2510\.23853*, 2025\.URL[https://arxiv\.org/abs/2510\.23853](https://arxiv.org/abs/2510.23853)\.
- Cohen et al\. \[2020\]Jonathan D\. Cohen, Keith Marzilli Ericson, David Laibson, and John Myles White\.Measuring time preferences\.*Journal of Economic Literature*, 58\(2\):299–347, 2020\.doi:10\.1257/jel\.20191074\.
- Cook et al\. \[2026\]Thomas R\. Cook, Sophia Kazinnik, Zach Modig, and Nathan M\. Palmer\.What do LLMs want?Finance and Economics Discussion Series 2026\-006, Board of Governors of the Federal Reserve System, January 2026\.URL[https://www\.federalreserve\.gov/econres/feds/what\-do\-llms\-want\.htm](https://www.federalreserve.gov/econres/feds/what-do-llms-want.htm)\.
- Cummins \[1975\]Robert Cummins\.Functional analysis\.*The Journal of Philosophy*, 72\(20\):741–765, 1975\.
- Cunningham et al\. \[2023\]Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey\.Sparse autoencoders find highly interpretable features in language models\.*arXiv preprint arXiv:2309\.08600*, 2023\.
- David \[2025\]Joey David\.Temporal predictors of outcome in reasoning language models\.*arXiv preprint arXiv:2511\.14773*, 2025\.URL[https://arxiv\.org/abs/2511\.14773](https://arxiv.org/abs/2511.14773)\.
- Dohmen et al\. \[2012\]Thomas J Dohmen, Armin Falk, David Huffman, and Uwe Sunde\.Interpreting time horizon effects in inter\-temporal choice\.*CESifo Working Paper*, 2012\.
- Dunefsky et al\. \[2024\]Jacob Dunefsky, Philippe Chlenski, and Neel Nanda\.Transcoders find interpretable llm feature circuits\.*Advances in Neural Information Processing Systems*, 37:24375–24410, 2024\.
- Ebert and Piehl \[1973\]Ronald J\. Ebert and DeWayne Piehl\.Time horizon: A concept for management\.*California Management Review*, 15\(4\):35–41, 1973\.doi:10\.2307/41164456\.
- Engels et al\. \[2025a\]Josh Engels, Subhash Kantamneni, Senthooran Rajamanoharan, and Neel Nanda\.Takeaways from our recent work on SAE probing\.AI Alignment Forum, March 2025a\.URL[https://www\.alignmentforum\.org/posts/osNKnwiJWHxDYvQTD/takeaways\-from\-our\-recent\-work\-on\-sae\-probing](https://www.alignmentforum.org/posts/osNKnwiJWHxDYvQTD/takeaways-from-our-recent-work-on-sae-probing)\.Accessed: 2025\.
- Engels et al\. \[2025b\]Joshua Engels, Eric J\. Michaud, Isaac Liao, Wes Gurnee, and Max Tegmark\.Not all language model features are one\-dimensionally linear\.In*International Conference on Learning Representations \(ICLR\)*, 2025b\.URL[https://arxiv\.org/abs/2405\.14860](https://arxiv.org/abs/2405.14860)\.
- Fatemi et al\. \[2024\]Bahare Fatemi, Mehran Kazemi, Anton Tsitsulin, Karishma Malkan, Jinyeong Yim, John Palowitch, Sungyong Seo, Jonathan Halcrow, and Bryan Perozzi\.Test of time: A benchmark for evaluating llms on temporal reasoning, 2024\.URL[https://arxiv\.org/abs/2406\.09170](https://arxiv.org/abs/2406.09170)\.
- Fedus et al\. \[2019\]William Fedus, Carles Gelada, Yoshua Bengio, Marc G\. Bellemare, and Hugo Larochelle\.Hyperbolic discounting and learning over multiple horizons, 2019\.URL[https://arxiv\.org/abs/1902\.06865](https://arxiv.org/abs/1902.06865)\.
- Fehr and Leibbrandt \[2011\]Ernst Fehr and Andreas Leibbrandt\.A field study on cooperativeness and impatience in the tragedy of the commons\.*Journal of Public Economics*, 95\(9–10\):1144–1155, 2011\.
- Figner et al\. \[2010\]Bernd Figner, Daria Knoch, Eric J\. Johnson, Amy R\. Krosch, Sarah H\. Lisanby, Ernst Fehr, and Elke U\. Weber\.Lateral prefrontal cortex and self\-control in intertemporal choice\.*Nature Neuroscience*, 13\(5\):538–539, 2010\.
- Frederick et al\. \[2002\]Shane Frederick, George Loewenstein, and Ted O’Donoghue\.Time discounting and time preference: A critical review\.*Journal of Economic Literature*, 40\(2\):351–401, 2002\.doi:10\.1257/002205102320161311\.
- Gärdenfors \[2000\]Peter Gärdenfors\.*Conceptual Spaces: The Geometry of Thought*\.MIT Press, Cambridge, MA, 2000\.
- Garikaparthi \[2026\]Aniketh Garikaparthi\.Can llms perceive time? an empirical investigation, 2026\.URL[https://arxiv\.org/abs/2604\.00010](https://arxiv.org/abs/2604.00010)\.
- Gazmararian \[2025\]Alexander F Gazmararian\.Valuing the future: Changing time horizons and policy preferences\.*Political Behavior*, 47\(2\):553–572, 2025\.
- Geiger et al\. \[2025\]Atticus Geiger, Duligur Ibeling, Amir Zur, Maheep Chaudhary, Sonakshi Chauhan, Jing Huang, Aryaman Arora, Zhengxuan Wu, Noah Goodman, Christopher Potts, and Thomas Icard\.Causal abstraction: A theoretical foundation for mechanistic interpretability, 2025\.URL[https://arxiv\.org/abs/2301\.04709](https://arxiv.org/abs/2301.04709)\.
- Gold and Britzky \[2026\]Hadas Gold and Haley Britzky\.Pentagon threatens to make Anthropic a pariah if it refuses to drop AI guardrails, 2026\.URL[https://www\.cnn\.com/2026/02/24/tech/hegseth\-anthropic\-ai\-military\-amodei](https://www.cnn.com/2026/02/24/tech/hegseth-anthropic-ai-military-amodei)\.Kaanita Iyer contributing\.
- Goldowsky\-Dill et al\. \[2023\]Nicholas Goldowsky\-Dill, Chris MacLeod, Lucas Sato, and Aryaman Arora\.Localizing model behavior with path patching\.*arXiv preprint arXiv:2304\.05969*, 2023\.
- Green and Myerson \[2004\]Leonard Green and Joel Myerson\.A discounting framework for choice with delayed and probabilistic rewards\.*Psychological Bulletin*, 130\(5\):769–792, 2004\.
- Greenblatt et al\. \[2024\]Ryan Greenblatt, Buck Shlegeris, Kshitij Sachan, and Fabien Roger\.Ai control: Improving safety despite intentional subversion, 2024\.URL[https://arxiv\.org/abs/2312\.06942](https://arxiv.org/abs/2312.06942)\.
- Gurnee and Tegmark \[2024\]Wes Gurnee and Max Tegmark\.Language models represent space and time\.In*The Twelfth International Conference on Learning Representations \(ICLR\)*, 2024\.URL[https://openreview\.net/forum?id=jE8xbmvFin](https://openreview.net/forum?id=jE8xbmvFin)\.
- Gurnee et al\. \[2026\]Wes Gurnee, Emmanuel Ameisen, Isaac Kauvar, Julius Tarng, Adam Pearce, Chris Olah, and Joshua Batson\.When models manipulate manifolds: The geometry of a counting task\.*arXiv preprint arXiv:2601\.04480*, 2026\.
- Guterres \[2025\]António Guterres\.Lethal autonomous weapon system “politically unacceptable, morally repugnant and should be banned”, 2025\.URL[https://press\.un\.org/en/2025/sgsm22643\.doc\.htm](https://press.un.org/en/2025/sgsm22643.doc.htm)\.
- Han et al\. \[2025\]Xue Han, Qian Hu, Yitong Wang, Wenchun Gao, Lianlian Zhang, Qing Wang, Lijun Mei, Chao Deng, and Junlan Feng\.Temporal alignment of llms through cycle encoding for long\-range time representations, 2025\.URL[https://arxiv\.org/abs/2503\.04150](https://arxiv.org/abs/2503.04150)\.
- Hanna et al\. \[2023\]Michael Hanna, Ollie Liu, and Alexandre Variengien\.How does GPT\-2 compute greater\-than?: Interpreting mathematical abilities in a pre\-trained language model\.In*Advances in Neural Information Processing Systems*, volume 36, 2023\.
- Hanna et al\. \[2024\]Michael Hanna, Sandro Pezzelle, and Yonatan Belinkov\.Have faith in faithfulness: Going beyond circuit overlap when finding model mechanisms, 2024\.URL[https://arxiv\.org/abs/2403\.17806](https://arxiv.org/abs/2403.17806)\.
- Heimersheim and Nanda \[2024\]Stefan Heimersheim and Neel Nanda\.How to use and interpret activation patching, 2024\.URL[https://arxiv\.org/abs/2404\.15255](https://arxiv.org/abs/2404.15255)\.
- Heinzerling and Inui \[2024\]Benjamin Heinzerling and Kentaro Inui\.Monotonic representation of numeric attributes in language models\.In*Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 2: Short Papers\)*, pages 175–195, 2024\.doi:10\.18653/v1/2024\.acl\-short\.18\.URL[https://aclanthology\.org/2024\.acl\-short\.18/](https://aclanthology.org/2024.acl-short.18/)\.
- Herel et al\. \[2024\]David Herel, Vojtech Bartek, Jiri Jirak, and Tomas Mikolov\.Time awareness in large language models: Benchmarking fact recall across time, 2024\.URL[https://arxiv\.org/abs/2409\.13338](https://arxiv.org/abs/2409.13338)\.
- Holtermann et al\. \[2025\]Carolin Holtermann, Paul Röttger, and Anne Lauscher\.Around the world in 24 hours: Probing llm knowledge of time and place, 2025\.URL[https://arxiv\.org/abs/2506\.03984](https://arxiv.org/abs/2506.03984)\.
- Horton et al\. \[2026\]John J\. Horton, Apostolos Filippas, and Benjamin S\. Manning\.Large language models as simulated economic agents: What can we learn from homo silicus?, 2026\.URL[https://arxiv\.org/abs/2301\.07543](https://arxiv.org/abs/2301.07543)\.
- Initiative \[2026\]Cloud Security Alliance AI Safety Initiative\.Pentagon vs\. Anthropic: Autonomous weapons AI guardrails and the governance crisis for enterprise AI vendors, 2026\.URL[https://labs\.cloudsecurityalliance\.org/research/csa\-research\-note\-dod\-ai\-guardrail\-mandates\-vendor\-governanc/](https://labs.cloudsecurityalliance.org/research/csa-research-note-dod-ai-guardrail-mandates-vendor-governanc/)\.
- International Risk Management Institute \[n\.d\.\]International Risk Management Institute\.Retroactive date\.IRMI Insurance Glossary, n\.d\.URL[https://www\.irmi\.com/term/insurance\-definitions/retroactive\-date](https://www.irmi.com/term/insurance-definitions/retroactive-date)\.Accessed April 13, 2026\.
- Kable and Glimcher \[2007\]Joseph W\. Kable and Paul W\. Glimcher\.The neural correlates of subjective value during intertemporal choice\.*Nature Neuroscience*, 10\(12\):1625–1633, 2007\.
- Kadlčík et al\. \[2025\]Marek Kadlčík, Michal Štefánik, Timothee Mickus, Michal Spiegel, and Josef Kuchař\.Pre\-trained language models learn remarkably accurate representations of numbers\.*arXiv preprint arXiv:2506\.08966*, 2025\.URL[https://arxiv\.org/abs/2506\.08966](https://arxiv.org/abs/2506.08966)\.
- Kantamneni and Tegmark \[2025\]Subhash Kantamneni and Max Tegmark\.Language models use trigonometry to do addition, 2025\.URL[https://arxiv\.org/abs/2502\.00873](https://arxiv.org/abs/2502.00873)\.
- Karkada et al\. \[2026\]Dhruva Karkada, Daniel J\. Korchinski, Andres Nava, Matthieu Wyart, and Yasaman Bahri\.Symmetry in language statistics shapes the geometry of model representations, 2026\.URL[https://arxiv\.org/abs/2602\.15029](https://arxiv.org/abs/2602.15029)\.
- Kim \[2023\]Jeongbin Kim\.The effects of time preferences on cooperation: Experimental evidence from infinitely repeated games\.*American Economic Journal: Microeconomics*, 15\(1\):618–637, 2023\.
- Kim et al\. \[2025\]Junsol Kim, James Evans, and Aaron Schein\.Linear representations of political perspective emerge in large language models, 2025\.URL[https://arxiv\.org/abs/2503\.02080](https://arxiv.org/abs/2503.02080)\.
- Kirby et al\. \[1999\]Kris N\. Kirby, Nancy M\. Petry, and Warren K\. Bickel\.Heroin addicts have higher discount rates for delayed rewards than non\-drug\-using controls\.*Journal of Experimental Psychology: General*, 128\(1\):78–87, 1999\.doi:10\.1037/0096\-3445\.128\.1\.78\.
- Korsgaard \[1997\]Christine M\. Korsgaard\.The normativity of instrumental reason\.In Garrett Cullity and Berys Gaut, editors,*Ethics and Practical Reason*, pages 215–254\. Oxford University Press, Oxford, 1997\.
- Lampinen et al\. \[2026\]Andrew Kyle Lampinen, Yuxuan Li, Eghbal Hosseini, Sangnie Bhardwaj, and Murray Shanahan\.Linear representations in language models can change dramatically over a conversation, 2026\.URL[https://arxiv\.org/abs/2601\.20834](https://arxiv.org/abs/2601.20834)\.
- Lee et al\. \[2025\]Jin Hwa Lee, Thomas Jiralerspong, Lei Yu, Yoshua Bengio, and Emily Cheng\.Geometric signatures of compositionality across a language model’s lifetime, 2025\.URL[https://arxiv\.org/abs/2410\.01444](https://arxiv.org/abs/2410.01444)\.
- Leinster \[2021\]Tom Leinster\.*Entropy and Diversity: The Axiomatic Approach*\.Cambridge University Press, 2021\.ISBN 9781108832700\.doi:10\.1017/9781108963558\.
- Leng \[2024a\]Yan Leng\.Can LLMs mimic human\-like mental accounting and behavioral biases?In*Proceedings of the 25th ACM Conference on Economics and Computation \(EC ’24\)*\. ACM, 2024a\.doi:10\.1145/3670865\.3673632\.
- Leng \[2024b\]Yan Leng\.Folk economics in the machine: LLMs and the emergence of mental accounting\.*SSRN preprint 4705130*, 2024b\.
- Li et al\. \[2026\]Jiaqian Li, Yanshu Li, and Kuan\-Hao Huang\.Steering vector fields for context\-aware inference\-time control in large language models, 2026\.URL[https://arxiv\.org/abs/2602\.01654](https://arxiv.org/abs/2602.01654)\.
- Li et al\. \[2025\]Lingyu Li, Yang Yao, Yixu Wang, Chubo Li, Yan Teng, and Yingchun Wang\.The other mind: How language models exhibit human temporal cognition, 2025\.URL[https://arxiv\.org/abs/2507\.15851](https://arxiv.org/abs/2507.15851)\.
- Liu et al\. \[2025\]Zijia Liu, Peixuan Han, Haofei Yu, Haoru Li, and Jiaxuan You\.Time\-r1: Towards comprehensive temporal reasoning in llms, 2025\.URL[https://arxiv\.org/abs/2505\.13508](https://arxiv.org/abs/2505.13508)\.
- Marks and Tegmark \[2024\]Samuel Marks and Max Tegmark\.The geometry of truth: Emergent linear structure in large language model representations of true/false datasets\.In*Conference on Language Modeling \(COLM\)*, 2024\.URL[https://arxiv\.org/abs/2310\.06824](https://arxiv.org/abs/2310.06824)\.
- Mazyaki et al\. \[2025\]Ali Mazyaki, Mohammad Naghizadeh, Samaneh Ranjkhah Zonouzaghi, and Hossein Setareh\.Temporal preferences in language models for long\-horizon assistance\.*arXiv preprint arXiv:2509\.09704*, 2025\.
- Meinke et al\. \[2025\]Alexander Meinke, Bronson Schoen, Jérémy Scheurer, Mikita Balesni, Rusheb Shah, and Marius Hobbhahn\.Frontier models are capable of in\-context scheming, 2025\.URL[https://arxiv\.org/abs/2412\.04984](https://arxiv.org/abs/2412.04984)\.
- Mitchell et al\. \[2005\]Ian M\. Mitchell, Alexandre M\. Bayen, and Claire J\. Tomlin\.A time\-dependent Hamilton–Jacobi formulation of reachable sets for continuous dynamic games\.*IEEE Transactions on Automatic Control*, 50\(7\):947–957, 2005\.doi:10\.1109/TAC\.2005\.851439\.
- Mitchell et al\. \[2025\]Margaret Mitchell, Avijit Ghosh, Alexandra Sasha Luccioni, and Giada Pistilli\.Fully autonomous AI agents should not be developed\.*arXiv preprint arXiv:2502\.02649*, 2025\.
- Modell et al\. \[2025\]Alexander Modell, Patrick Rubin\-Delanchy, and Nick Whiteley\.The origins of representation manifolds in large language models\.*arXiv preprint arXiv:2505\.18235*, 2025\.URL[https://arxiv\.org/abs/2505\.18235](https://arxiv.org/abs/2505.18235)\.
- Moghimi et al\. \[2026\]Mehrdad Moghimi, Anthony Coache, and Hyejin Ku\.Decoupling time and risk: Risk\-sensitive reinforcement learning with general discounting, 2026\.URL[https://arxiv\.org/abs/2602\.04131](https://arxiv.org/abs/2602.04131)\.
- Mueller et al\. \[2025\]Aaron Mueller, Atticus Geiger, Sarah Wiegreffe, Dana Arad, Iván Arcuschin, Adam Belfki, Yik Siu Chan, Jaden Fiotto\-Kaufman, Tal Haklay, Michael Hanna, Jing Huang, Rohan Gupta, Yaniv Nikankin, Hadas Orgad, Nikhil Prakash, Anja Reusch, Aruna Sankaranarayanan, Shun Shao, Alessandro Stolfo, Martin Tutek, Amir Zur, David Bau, and Yonatan Belinkov\.Mib: A mechanistic interpretability benchmark, 2025\.URL[https://arxiv\.org/abs/2504\.13151](https://arxiv.org/abs/2504.13151)\.
- Nainani et al\. \[2025\]Jatin Nainani, Sankaran Vaidyanathan, Connor Watts, Andre N\. Assis, and Alice Rigg\.Detecting and characterizing planning in language models, 2025\.URL[https://arxiv\.org/abs/2508\.18098](https://arxiv.org/abs/2508.18098)\.
- Ngo et al\. \[2025\]Richard Ngo, Lawrence Chan, and Sören Mindermann\.The alignment problem from a deep learning perspective, 2025\.URL[https://arxiv\.org/abs/2209\.00626](https://arxiv.org/abs/2209.00626)\.
- Nylund et al\. \[2024\]Kai Nylund, Suchin Gururangan, and Noah A\. Smith\.Time is encoded in the weights of finetuned language models\.In*Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(ACL\)*, 2024\.URL[https://aclanthology\.org/2024\.acl\-long\.141/](https://aclanthology.org/2024.acl-long.141/)\.
- Panickssery et al\. \[2024\]Nina Panickssery, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Matt Turner\.Steering Llama 2 via contrastive activation addition, 2024\.URL[https://arxiv\.org/abs/2312\.06681](https://arxiv.org/abs/2312.06681)\.
- Papadopoulos et al\. \[2024\]Vassilis Papadopoulos, Jérémie Wenger, and Clément Hongler\.Arrows of time for large language models, 2024\.URL[https://arxiv\.org/abs/2401\.17505](https://arxiv.org/abs/2401.17505)\.
- Park et al\. \[2024\]Kiho Park, Yo Joong Choe, and Victor Veitch\.The linear representation hypothesis and the geometry of large language models\.In*International Conference on Machine Learning \(ICML\)*, 2024\.URL[https://arxiv\.org/abs/2311\.03658](https://arxiv.org/abs/2311.03658)\.
- Park et al\. \[2023\]Peter S\. Park, Simon Goldstein, Aidan O’Gara, Michael Chen, and Dan Hendrycks\.Ai deception: A survey of examples, risks, and potential solutions, 2023\.URL[https://arxiv\.org/abs/2308\.14752](https://arxiv.org/abs/2308.14752)\.
- Pearl \[2009\]Judea Pearl\.*Causality*\.Cambridge university press, 2009\.
- Persson et al\. \[2024\]Emil Persson, Gustav Tinghög, and Daniel Västfjäll\.Intertemporal prosocial behavior: a review and research agenda\.*Frontiers in psychology*, 15:1359447, 2024\.
- Postmus and Abreu \[2025\]Joris Postmus and Steven Abreu\.Steering large language models using conceptors: Improving addition\-based activation engineering\.*arXiv preprint arXiv:2410\.16314*, 2025\.
- Qwen Team \[2025\]Qwen Team\.Qwen3\-4b\-instruct\-2507\.[https://huggingface\.co/Qwen/Qwen3\-4B\-Instruct\-2507](https://huggingface.co/Qwen/Qwen3-4B-Instruct-2507), 2025\.Accessed: 2026\.
- Rajendran et al\. \[2024\]Goutham Rajendran, Simon Buchholz, Bryon Aragam, Bernhard Schölkopf, and Pradeep Kumar Ravikumar\.From causal to concept\-based representation learning\.In*The Thirty\-eighth Annual Conference on Neural Information Processing Systems*, 2024\.URL[https://openreview\.net/forum?id=r5nev2SHtJ](https://openreview.net/forum?id=r5nev2SHtJ)\.
- Raval et al\. \[2026\]Shivam Raval, Hae Jin Song, Linlin Wu, Abir Harrasse, Jeff M\. Phillips, Fazl Barez, and Amirali Abdullah\.Curveball steering: The right direction to steer isn’t always linear, 2026\.URL[https://arxiv\.org/abs/2603\.09313](https://arxiv.org/abs/2603.09313)\.
- Reilly et al\. \[2016\]Greg Reilly, David Souder, and Rebecca Ranucci\.Time horizon of investments in the resource allocation process: Review and framework for next steps\.*Journal of Management*, 42\(5\):1169–1194, 2016\.doi:10\.1177/0149206316630381\.
- Ross et al\. \[2024\]Jillian Ross, Yoon Kim, and Andrew W\. Lo\.LLM economicus? mapping the behavioral biases of LLMs via utility theory\.*arXiv preprint arXiv:2408\.02784*, 2024\.
- Saglam et al\. \[2025\]Baturay Saglam, Paul Kassianik, Blaine Nelson, Sajana Weerawardhena, Yaron Singer, and Amin Karbasi\.Large language models encode semantics and alignment in linearly separable representations, 2025\.URL[https://arxiv\.org/abs/2507\.09709](https://arxiv.org/abs/2507.09709)\.
- Sehgal et al\. \[2026\]Neil K\. R\. Sehgal, Sharath Chandra Guntuku, and Lyle Ungar\.Real\-time deadlines reveal temporal awareness failures in llm strategic dialogues, 2026\.URL[https://arxiv\.org/abs/2601\.13206](https://arxiv.org/abs/2601.13206)\.
- Shafran et al\. \[2026\]Or Shafran, Shaked Ronen, Omri Fahn, Shauli Ravfogel, Atticus Geiger, and Mor Geva\.From directions to regions: Decomposing activations in language models via local geometry, 2026\.URL[https://arxiv\.org/abs/2602\.02464](https://arxiv.org/abs/2602.02464)\.
- Shamosh et al\. \[2008\]Noah A\. Shamosh, Colin G\. DeYoung, Adam E\. Green, Deidre L\. Reis, Matthew R\. Johnson, Andrew R\. A\. Conway, Randall W\. Engle, Todd S\. Braver, and Jeremy R\. Gray\.Individual differences in delay discounting: Relation to intelligence, working memory, and anterior prefrontal cortex\.*Psychological Science*, 19\(9\):904–911, 2008\.doi:10\.1111/j\.1467\-9280\.2008\.02175\.x\.
- Shin et al\. \[2025\]Changho Shin, Xinya Yan, Suenggwan Jo, Sungjun Cho, Shourjo Aditya Chaudhuri, and Frederic Sala\.TARDIS: Mitigating temporal misalignment via representation steering\.*arXiv preprint arXiv:2503\.18693*, 2025\.URL[https://arxiv\.org/abs/2503\.18693](https://arxiv.org/abs/2503.18693)\.
- Shlens \[2014\]Jonathon Shlens\.A tutorial on principal component analysis\.*arXiv preprint arXiv:1404\.1100*, 2014\.
- Sofroniew et al\. \[2026\]Nicholas Sofroniew, Isaac Kauvar, William Saunders, Runjin Chen, Tom Henighan, Sasha Hydrie, Craig Citro, Adam Pearce, Julius Tarng, Wes Gurnee, Joshua Batson, Sam Zimmerman, Kelley Rivoire, Kyle Fish, Chris Olah, and Jack Lindsey\.Emotion concepts and their function in a large language model, 2026\.URL[https://arxiv\.org/abs/2604\.07729](https://arxiv.org/abs/2604.07729)\.
- Song et al\. \[2025\]Xiangchen Song, Jiaqi Sun, Zijian Li, Yujia Zheng, and Kun Zhang\.Llm interpretability with identifiable temporal\-instantaneous representation, 2025\.URL[https://arxiv\.org/abs/2509\.23323](https://arxiv.org/abs/2509.23323)\.
- Svenson and Karlsson \[1989\]Ola Svenson and Gunnar Karlsson\.Decision\-making, time horizons, and risk in the very long\-term perspective\.*Risk Analysis*, 9\(3\):385–399, 1989\.
- Tiblias et al\. \[2025\]Federico Tiblias, Irina Bigoulaeva, Jingcheng Niu, Simone Balloccu, and Iryna Gurevych\.Hypothesis\-driven feature manifold analysis in LLMs via supervised multi\-dimensional scaling\.*arXiv preprint arXiv:2510\.01025*, 2025\.URL[https://arxiv\.org/abs/2510\.01025](https://arxiv.org/abs/2510.01025)\.
- Turner et al\. \[2023\]Alexander Matt Turner, Lisa Thiergart, Gavin Leech, David Udell, Juan J\. Vazquez, Ulisse Mini, and Monte MacDiarmid\.Activation addition: Steering language models without optimization\.*arXiv preprint arXiv:2308\.10248*, 2023\.
- Vu and Nguyen \[2025\]Hieu M\. Vu and Tan M\. Nguyen\.Angular steering: Behavior control via rotation in activation space\.In*Advances in Neural Information Processing Systems \(NeurIPS\)*, 2025\.URL[https://arxiv\.org/abs/2510\.26243](https://arxiv.org/abs/2510.26243)\.arXiv:2510\.26243\.
- Wallat et al\. \[2025\]Jonas Wallat, Abdelrahman Abdallah, Adam Jatowt, and Avishek Anand\.A study into investigating temporal robustness of llms, 2025\.URL[https://arxiv\.org/abs/2503\.17073](https://arxiv.org/abs/2503.17073)\.
- Wang et al\. \[2025a\]Jinjin Wang, Yuzhen Li, Jun Luo, and Hang Ye\.Measuring time preference: Theory, methods, and applications\.*Acta Psychologica*, 261:105928, 2025a\.
- Wang et al\. \[2025b\]Rui Wang, Qihan Lin, Jiayu Liu, Qing Zong, Tianshi Zheng, Dadi Guo, Haochen Shi, Weiqi Wang, and Yangqiu Song\.Rethinking prospect theory for LLMs: Revealing the instability of decision\-making under epistemic uncertainty\.*arXiv preprint arXiv:2508\.08992*, 2025b\.
- Wang and Zhao \[2024\]Yuqing Wang and Yun Zhao\.Tram: Benchmarking temporal reasoning for large language models, 2024\.URL[https://arxiv\.org/abs/2310\.00835](https://arxiv.org/abs/2310.00835)\.
- Wang et al\. \[2026\]Zehong Wang, Fang Wu, Hongru Wang, Xiangru Tang, Bolian Li, Zhenfei Yin, Yijun Ma, Yiyang Li, Weixiang Sun, Xiusi Chen, and Yanfang Ye\.Why reasoning fails to plan: A planning\-centric analysis of long\-horizon decision making in llm agents, 2026\.URL[https://arxiv\.org/abs/2601\.22311](https://arxiv.org/abs/2601.22311)\.
- Wehner et al\. \[2025\]Jan Wehner, Sahar Abdelnabi, Daniel Tan, David Krueger, and Mario Fritz\.Taxonomy, opportunities, and challenges of representation engineering for large language models, 2025\.URL[https://arxiv\.org/abs/2502\.19649](https://arxiv.org/abs/2502.19649)\.
- Wolf et al\. \[2024\]Yotam Wolf, Noam Wies, Dorin Shteyman, Binyamin Rothberg, Yoav Levine, and Amnon Shashua\.Tradeoffs between alignment and helpfulness in language models with steering methods\.*arXiv preprint arXiv:2401\.16332*, 2024\.
- Wu et al\. \[2024\]Zhengxuan Wu, Aryaman Arora, Zheng Wang, Atticus Geiger, Dan Jurafsky, Christopher D Manning, and Christopher Potts\.ReFT: Representation finetuning for language models\.*Advances in Neural Information Processing Systems*, 37:63908–63962, 2024\.
- Wu et al\. \[2025\]Zhengxuan Wu, Aryaman Arora, Atticus Geiger, Zheng Wang, Jing Huang, Dan Jurafsky, Christopher D Manning, and Christopher Potts\.AxBench: Steering LLMs? even simple baselines outperform sparse autoencoders\.*arXiv preprint arXiv:2501\.17148*, 2025\.
- Yang et al\. \[2025\]An Yang et al\.Qwen3 technical report\.*arXiv preprint arXiv:2505\.09388*, 2025\.
- Ying et al\. \[2026\]Zhuofan Josh Ying, Shauli Ravfogel, Nikolaus Kriegeskorte, and Peter Hase\.The truthfulness spectrum hypothesis, 2026\.URL[https://arxiv\.org/abs/2602\.20273](https://arxiv.org/abs/2602.20273)\.
- You et al\. \[2026\]Zejia You, Chunyuan Deng, and Hanjie Chen\.Spherical steering: Geometry\-aware activation rotation for language models\.*arXiv preprint arXiv:2602\.08169*, 2026\.
- Zhang et al\. \[2026\]Hengyuan Zhang, Zhihao Zhang, Mingyang Wang, Zunhai Su, Yiwei Wang, Qianli Wang, Shuzhou Yuan, Ercong Nie, Xufeng Duan, Feijiang Han, Qibo Xue, Zeping Yu, Chenming Shang, Xiao Liang, Jing Xiong, Hui Shen, Chaofan Tao, Zhengwu Liu, Senjie Jin, Zhiheng Xi, Dongdong Zhang, Sophia Ananiadou, Tao Gui, Ruobing Xie, Hayden Kwok\-Hay So, Hinrich Schütze, Xuanjing Huang, Qi Zhang, and Ngai Wong\.Locate, steer, and improve: A practical survey of actionable mechanistic interpretability in large language models, 2026\.URL[https://arxiv\.org/abs/2601\.14004](https://arxiv.org/abs/2601.14004)\.
- Zhu et al\. \[2025a\]Jian\-Qiao Zhu, Haijiang Yan, and Thomas L\. Griffiths\.Language models trained to do arithmetic predict human risky and intertemporal choice, 2025a\.URL[https://arxiv\.org/abs/2405\.19313](https://arxiv.org/abs/2405.19313)\.
- Zhu et al\. \[2025b\]Jian\-Qiao Zhu, Haijiang Yan, and Thomas L Griffiths\.Steering risk preferences in large language models by aligning behavioral and neural representations\.*arXiv preprint arXiv:2505\.11615*, 2025b\.
- Zou et al\. \[2025\]Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann\-Kathrin Dombrowski, Shashwat Goel, Nathaniel Li, Michael J\. Byun, Zifan Wang, Alex Mallen, Steven Basart, Sanmi Koyejo, Dawn Song, Matt Fredrikson, J\. Zico Kolter, and Dan Hendrycks\.Representation engineering: A top\-down approach to ai transparency, 2025\.URL[https://arxiv\.org/abs/2310\.01405](https://arxiv.org/abs/2310.01405)\.
Appendices
## Index of Appendices
App\.p\.ParadigmContentPart 0: Groundwork[A](https://arxiv.org/html/2606.05194#A1)[Appendix A](https://arxiv.org/html/2606.05194#A1)n/aExtended background: definitions and interpretability primitives\.[B](https://arxiv.org/html/2606.05194#A2)[Appendix B](https://arxiv.org/html/2606.05194#A2)n/aExtended literature\.[C](https://arxiv.org/html/2606.05194#A3)[Appendix C](https://arxiv.org/html/2606.05194#A3)n/aMethodology summary and overview of the extended methodologies\.[D](https://arxiv.org/html/2606.05194#A4)[Appendix D](https://arxiv.org/html/2606.05194#A4)n/aFull experimental details\.[E](https://arxiv.org/html/2606.05194#A5)[Appendix E](https://arxiv.org/html/2606.05194#A5)n/aPrompting settings and dataset construction\.[F](https://arxiv.org/html/2606.05194#A6)[Appendix F](https://arxiv.org/html/2606.05194#A6)n/aExtended limitations and future work\.Part 1: Localize\(ordered by method strength: probing→\\toattributional→\\tocausal\)[G](https://arxiv.org/html/2606.05194#A7)[Appendix G](https://arxiv.org/html/2606.05194#A7)ProbingLinear probing: 99\.2% at L26, cross\-dataset generalization\.[H](https://arxiv.org/html/2606.05194#A8)[Appendix H](https://arxiv.org/html/2606.05194#A8)Attr\. contr\.EAP\-IG attribution on contrastive prompts\.[I](https://arxiv.org/html/2606.05194#A10)[Appendix J](https://arxiv.org/html/2606.05194#A10)Causal param\.Activation patching: L21–24 attn, L31–35 MLP\.[J](https://arxiv.org/html/2606.05194#A11)[Appendix K](https://arxiv.org/html/2606.05194#A11)Causal class\.Directional patching: asymmetric effects, two\-phase classification, L24 attn\.[K](https://arxiv.org/html/2606.05194#A12)[Appendix L](https://arxiv.org/html/2606.05194#A12)AllCross\-method convergence on layers 17–35\.Part 2: Characterize[L](https://arxiv.org/html/2606.05194#A13)[Appendix M](https://arxiv.org/html/2606.05194#A13)ParametricPCA geometry: horizon→\\topreference transformation at turn boundary\.[M](https://arxiv.org/html/2606.05194#A14)[Appendix N](https://arxiv.org/html/2606.05194#A14)AllLatent vs\. constrained: sparse attn vs\. full subgraph\.[N](https://arxiv.org/html/2606.05194#A15)[Appendix O](https://arxiv.org/html/2606.05194#A15)BehavioralTemporal discounting: LLMs 3–8×\\timesmore patient than humans\.[O](https://arxiv.org/html/2606.05194#A16)[Appendix P](https://arxiv.org/html/2606.05194#A16)BehavioralBehavioral coherence: order bias, instruct degradation\.[P](https://arxiv.org/html/2606.05194#A17)[Appendix Q](https://arxiv.org/html/2606.05194#A17)Causal param\.Cross\-model patching: circuit localizes at fractional depth 0\.6–0\.7 across scale\.[Q](https://arxiv.org/html/2606.05194#A18)[Appendix R](https://arxiv.org/html/2606.05194#A18)ProbingError monitoring: subgraph encodes chain reliability orthogonally to time horizon\.Part 3: Intervene[R](https://arxiv.org/html/2606.05194#A19)[Appendix S](https://arxiv.org/html/2606.05194#A19)ContrastiveCAA steering: bidirectional, L19–22, probing–steering dissociation\.Part 4: Extended methodologies\(same strength ordering as Part 1\)[S](https://arxiv.org/html/2606.05194#A20)[Appendix T](https://arxiv.org/html/2606.05194#A20)n/aNotation and key concepts\.[T](https://arxiv.org/html/2606.05194#A21)[Appendix U](https://arxiv.org/html/2606.05194#A21)ProbingProbing protocol and activation extraction\.[U](https://arxiv.org/html/2606.05194#A22)[Appendix V](https://arxiv.org/html/2606.05194#A22)Attr\. contr\.EAP\-IG methodology, bias controls, component taxonomy\.[V](https://arxiv.org/html/2606.05194#A23)[Appendix W](https://arxiv.org/html/2606.05194#A23)Causal param\.Activation patching setup, noise/denoise protocol\.[W](https://arxiv.org/html/2606.05194#A24)[Appendix X](https://arxiv.org/html/2606.05194#A24)Causal class\.Directional patching on contrastive classification prompts\.[X](https://arxiv.org/html/2606.05194#A25)[Appendix Y](https://arxiv.org/html/2606.05194#A25)ParametricPCA geometry analysis pipeline\.[Y](https://arxiv.org/html/2606.05194#A26)[Appendix Z](https://arxiv.org/html/2606.05194#A26)BehavioralKirby MCQ\-27 instrument, decision boundary method\.[Z](https://arxiv.org/html/2606.05194#A27)[Appendix AA](https://arxiv.org/html/2606.05194#A27)BehavioralBehavioral coherence experiment design\.[AA](https://arxiv.org/html/2606.05194#A28)[Appendix AB](https://arxiv.org/html/2606.05194#A28)ContrastiveCAA vector construction and steering setup\.[AB](https://arxiv.org/html/2606.05194#A29)[Appendix AC](https://arxiv.org/html/2606.05194#A29)n/aWorked case study: highly\-formatted pair\.
Part 0: GroundworkA\.Ext\. backgroundB\.Ext\. literatureC\.Method\. summaryD\.Experimental detailsE\.PromptsF\.Ext\. limitationsPart 1: LocalizeG\.ProbingH\.Attr\. contr\.I\.Causal param\.J\.Causal class\.K\.ConvergencePart 2: CharacterizeL\.GeometryM\.Latent/constr\.N\.DiscountingO\.CoherenceP\.Cross\-model comp\.Q\.Error monitoringPart 3: InterveneR\.SteeringPart 4: Extended methodologiesS\.NotationT\.Probing meth\.U\.Attr\. contr\. meth\.V\.Causal param\. meth\.W\.Causal class\. meth\.X\.Geometry meth\.Y\.Discounting meth\.Z\.Coherence meth\.AA\.Steering meth\.AB\.Case study
## A visual tour of the appendices
![[Uncaptioned image]](https://arxiv.org/html/2606.05194v1/x11.png)
[M\.3](https://arxiv.org/html/2606.05194#A13.SS3), p\.[M\.3](https://arxiv.org/html/2606.05194#A13.SS3)
![[Uncaptioned image]](https://arxiv.org/html/2606.05194v1/x12.png)
[M\.3](https://arxiv.org/html/2606.05194#A13.SS3), p\.[M\.3](https://arxiv.org/html/2606.05194#A13.SS3)
![[Uncaptioned image]](https://arxiv.org/html/2606.05194v1/x13.png)
[M\.3](https://arxiv.org/html/2606.05194#A13.SS3), p\.[M\.3](https://arxiv.org/html/2606.05194#A13.SS3)
![[Uncaptioned image]](https://arxiv.org/html/2606.05194v1/x14.png)
[M\.3](https://arxiv.org/html/2606.05194#A13.SS3), p\.[M\.3](https://arxiv.org/html/2606.05194#A13.SS3)
Part 0: Groundwork
- •[A](https://arxiv.org/html/2606.05194#A1)\.Extended background
- •[B](https://arxiv.org/html/2606.05194#A2)\.Extended literature
- •[C](https://arxiv.org/html/2606.05194#A3)\.Methodology summary
- •[D](https://arxiv.org/html/2606.05194#A4)\.Experimental details
- •[E](https://arxiv.org/html/2606.05194#A5)\.Prompts
- •[F](https://arxiv.org/html/2606.05194#A6)\.Extended limitations and future work
## Appendix Appendix AExtended background
This appendix preserves the full\-length background that the main text condenses for space\. It defines the temporal\-preference concepts we use, introduces intertemporal choice as the empirical handle, and reviews the mechanistic\-interpretability primitives our pipeline rests on: causal and attributional localization, representational geometry, and steering\. The related\-work discussion lives separately in[Appendix B](https://arxiv.org/html/2606.05194#A2)\.
### A\.1Temporal preference, horizon, and scope
We refer totemporal preferenceas the degree to which an agent values outcomes differently depending on when they occur\[[15](https://arxiv.org/html/2606.05194#bib.bib15),[103](https://arxiv.org/html/2606.05194#bib.bib103)\]\. We use the termtime horizonto denote the future moment at which outcomes are evaluated against an objective\[[22](https://arxiv.org/html/2606.05194#bib.bib22)\]\. Because future events do not affect past outcomes, the time horizon also acts as aconstrainton planning\[[88](https://arxiv.org/html/2606.05194#bib.bib88)\]\. It would beinstrumentally\[[58](https://arxiv.org/html/2606.05194#bib.bib58)\]ormeans\-end\[[8](https://arxiv.org/html/2606.05194#bib.bib8)\]incoherentfor the agent to choose actions that are incapable of causing effects by the specified deadline\. Then, thetemporal scopeis the bounded interval of time over which an agent weighs the results according to its preference333Although not explored in this work, when rewards areperishable\[[4](https://arxiv.org/html/2606.05194#bib.bib4)\], the temporal scope is also bounded below by theretroactive reach\[[50](https://arxiv.org/html/2606.05194#bib.bib50),[70](https://arxiv.org/html/2606.05194#bib.bib70)\]\.\.
ValueTemporal PreferenceRewardV\(t\)V\(t\) λ\(t\)\\lambda\(t\) r\(t\)r\(t\) TemporalScopett =Time Horizontt ×RetroactiveReachtt
Figure A\.1:The time horizon specifies when the consequences of a decision are assessed\. The temporal scope is then bounded above by the time horizon\.In humans, these concepts have localized neural representations\[[51](https://arxiv.org/html/2606.05194#bib.bib51)\]that predict behavior\[[93](https://arxiv.org/html/2606.05194#bib.bib93)\], respond causally to intervention\[[28](https://arxiv.org/html/2606.05194#bib.bib28)\], and exhibit internal organization interpretable as a*functional role*\[[17](https://arxiv.org/html/2606.05194#bib.bib17)\]\.444Some authors argue that concepts are best modeled by geometric or topological spaces\[[30](https://arxiv.org/html/2606.05194#bib.bib30)\], a perspective that resonates with our geometric analysis of temporal representations in activation space\.Our work asks whether temporal preference exists in an LLM in an analogous way: localized, predictive, causally efficacious, and geometrically organized\.
### A\.2Modeling behavior via intertemporal choice
We measure temporal preference throughintertemporal choice: forced binary decisions between options that differ in reward and delay\. This is the standard instrument in behavioral economics and neuroeconomics\[[57](https://arxiv.org/html/2606.05194#bib.bib57),[29](https://arxiv.org/html/2606.05194#bib.bib29),[51](https://arxiv.org/html/2606.05194#bib.bib51)\]because it isolates preference from effort, attention, and planning; separates reward from delay; and produces a single forced\-choice token we can align across prompts and patch at the activation level\. Each optioniiis defined as a tuple\(ri,ti\)\(r\_\{i\},t\_\{i\}\), whereri∈ℝ\+r\_\{i\}\\in\\mathbb\{R\}^\{\+\}denotes the reward andti∈ℝ\+t\_\{i\}\\in\\mathbb\{R\}^\{\+\}the delay until receipt\. The subjective value of an option is the product of a temporal preference and the reward:
V\(t\)=λ\(t\)⋅r\(t\)V\(t\)=\\lambda\(t\)\\cdot r\(t\)\(A\.1\)Given two optionsA=\(rA,tA\)A=\(r\_\{A\},t\_\{A\}\)andB=\(rB,tB\)B=\(r\_\{B\},t\_\{B\}\), we predict the agent selects the option with the highest value:
i∗=argmaxi∈\{A,B\}λ\(ti\)⋅rii^\{\*\}=\\operatorname\*\{arg\\,max\}\_\{i\\in\\\{A,B\\\}\}\\;\\lambda\(t\_\{i\}\)\\cdot r\_\{i\}\(A\.2\)In the case of humans, temporal preference is often modeled using discount functions that capture our tendency to prefer immediate rewards over future ones\[[29](https://arxiv.org/html/2606.05194#bib.bib29),[36](https://arxiv.org/html/2606.05194#bib.bib36)\]\. For AI agents, it is possible that different classes of functions better model their behavior\[[68](https://arxiv.org/html/2606.05194#bib.bib68)\]\. Characterizing LLM behavior requires fitting a discount function via regression, assessing its stability across varying contexts, and benchmarking the resulting preferences against human intertemporal choice\.
### A\.3Localizing a subgraph
The process ofsubgraph localizationwithin an LLM involves identifying which components of the neural network are responsible for the behavior of interest\. The gold standard iscausal localization\[[33](https://arxiv.org/html/2606.05194#bib.bib33)\], which works by intervening within an LLM to measure the causal effect of specific components on the behavior of the model\. We adoptactivation patching\[[44](https://arxiv.org/html/2606.05194#bib.bib44)\]as our causal technique \(Section[Appendix J](https://arxiv.org/html/2606.05194#A10), Section[Appendix K](https://arxiv.org/html/2606.05194#A11)\): replace one component’s activation with a counterfactual value from another input and measure the behavioral change\. Attribution scores components via gradients and probing reads linearly decodable information, but neither intervenes on the forward pass; patching is the only one of the three that tests whether a component is causally necessary or sufficient for the output\. Using the do\-calculus notation\[[82](https://arxiv.org/html/2606.05194#bib.bib82)\], the patching intervention can be expressed as:
Δi\(l\)\(x,x′\)=𝔼\[Y∣do\(ai\(l\)=ai\(l\)\(x′\)\),X=x\]−𝔼\[Y∣X=x\]\\Delta\_\{i\}^\{\(l\)\}\(x,\\,x^\{\\prime\}\)\\;=\\;\\mathbb\{E\}\\bigl\[\\,Y\\mid\\operatorname\{do\}\\bigl\(\\,a\_\{i\}^\{\(l\)\}=a\_\{i\}^\{\(l\)\}\(x^\{\\prime\}\)\\,\\bigr\),\\;X=x\\,\\bigr\]\\;\-\\;\\mathbb\{E\}\\bigl\[\\,Y\\mid X=x\\,\\bigr\]\(A\.3\)
Unfortunately, performing targeted interventions is computationally expensive\. As an alternative,attributional localizationapproximates causal localization\[[6](https://arxiv.org/html/2606.05194#bib.bib6)\]\. We useEAP\-IG\[[43](https://arxiv.org/html/2606.05194#bib.bib43)\], a gradient\-based attribution method \(Section[Appendix H](https://arxiv.org/html/2606.05194#A8)\) that scores every head and MLP in a single backward pass\. This makes a full\-network scan tractable; the tradeoff is correlational estimates rather than causal guarantees\. We also useprobes, linear classifiers trained on a model’s internal activations\[[74](https://arxiv.org/html/2606.05194#bib.bib74),[56](https://arxiv.org/html/2606.05194#bib.bib56)\], to give us a complementary view by identifying which concepts a model encodes, where they emerge, and whether they are linearly represented\.
Including more components in a subgraph explains more of the LLM’s behavior, but yields a larger, less interpretable picture\. The full network trivially explains everything, and the empty subgraph explains nothing\. Any useful circuit falls between these extremes, balancing behavioral coverage against subgraph size\.
### A\.4Visualizing representational geometry
The activation space within an LLM encodes concepts in internal representations\[[6](https://arxiv.org/html/2606.05194#bib.bib6)\]\. A growing body of evidence has documented that many representations possess complex geometric structures\[[54](https://arxiv.org/html/2606.05194#bib.bib54),[24](https://arxiv.org/html/2606.05194#bib.bib24),[72](https://arxiv.org/html/2606.05194#bib.bib72),[39](https://arxiv.org/html/2606.05194#bib.bib39)\], beyond the global directions that theLinear Representation Hypothesispredicts\[[80](https://arxiv.org/html/2606.05194#bib.bib80)\]\. Furthermore, recent work has also noticed that LLMs show local low\-dimensional structure\[[92](https://arxiv.org/html/2606.05194#bib.bib92),[90](https://arxiv.org/html/2606.05194#bib.bib90),[60](https://arxiv.org/html/2606.05194#bib.bib60)\], and that even when representations are linear, they can change dramatically throughout generation\[[59](https://arxiv.org/html/2606.05194#bib.bib59)\]\.
These past findings motivate us to visualize the representational geometry within the localized subgraph as a way to understand what each component is doing\. In our work, we apply Principal Component Analysis \(PCA\)\[[95](https://arxiv.org/html/2606.05194#bib.bib95)\]to examine how the temporal concepts of interest are represented within a lower\-dimensional subspace\.
### A\.5Steering behavior with interventions
Steering refers to the control of an LLM’s behavior by directly intervening on its internal representations, rather than through prompting or training\. Subgraph localization is not required for steering\[[5](https://arxiv.org/html/2606.05194#bib.bib5),[107](https://arxiv.org/html/2606.05194#bib.bib107)\], but localization generally improves precision, reduces side effects, and allows smaller intervention magnitudes\[[114](https://arxiv.org/html/2606.05194#bib.bib114)\]\. In our work, we seek to understand the interventions in our subgraph through both a geometric and a behavioral perspective\.
## Appendix Appendix BExtended literature
Understanding temporal preference in LLMs requires drawing together the literature that has largely developed in isolation\.
##### Temporal Representation\.
LLMs encode temporal and spatial coordinates as geometric objects recoverable via regression probes\[[38](https://arxiv.org/html/2606.05194#bib.bib38),[77](https://arxiv.org/html/2606.05194#bib.bib77),[47](https://arxiv.org/html/2606.05194#bib.bib47)\], forming circular, helical, and manifold structures\[[24](https://arxiv.org/html/2606.05194#bib.bib24),[53](https://arxiv.org/html/2606.05194#bib.bib53),[52](https://arxiv.org/html/2606.05194#bib.bib52),[42](https://arxiv.org/html/2606.05194#bib.bib42),[72](https://arxiv.org/html/2606.05194#bib.bib72),[99](https://arxiv.org/html/2606.05194#bib.bib99),[39](https://arxiv.org/html/2606.05194#bib.bib39)\]that obey psychophysical scaling laws\[[11](https://arxiv.org/html/2606.05194#bib.bib11),[54](https://arxiv.org/html/2606.05194#bib.bib54)\]\. Linear decodability coexists with non\-linear geometry because features are locally linear on globally curved manifolds\[[72](https://arxiv.org/html/2606.05194#bib.bib72),[80](https://arxiv.org/html/2606.05194#bib.bib80),[92](https://arxiv.org/html/2606.05194#bib.bib92),[86](https://arxiv.org/html/2606.05194#bib.bib86)\]\. Yet the best\-geometry layer is not the computational layer\[[10](https://arxiv.org/html/2606.05194#bib.bib10)\], causality is rarely established\[[45](https://arxiv.org/html/2606.05194#bib.bib45)\], and it is not known whether temporal representations causally drive downstream behavior across contexts the way emotion concepts do\[[96](https://arxiv.org/html/2606.05194#bib.bib96)\]\.
##### Temporal Reasoning and Planning\.
LLMs fail at temporal reasoning tasks despite encoding time geometrically\[[105](https://arxiv.org/html/2606.05194#bib.bib105),[25](https://arxiv.org/html/2606.05194#bib.bib25),[66](https://arxiv.org/html/2606.05194#bib.bib66),[102](https://arxiv.org/html/2606.05194#bib.bib102),[46](https://arxiv.org/html/2606.05194#bib.bib46)\], and lack continuous temporal grounding: they cannot track real\-time deadlines even when discrete turn\-based reasoning succeeds\[[91](https://arxiv.org/html/2606.05194#bib.bib91),[31](https://arxiv.org/html/2606.05194#bib.bib31),[14](https://arxiv.org/html/2606.05194#bib.bib14)\]\. Evidence of temporal structure exists\[[79](https://arxiv.org/html/2606.05194#bib.bib79),[65](https://arxiv.org/html/2606.05194#bib.bib65),[19](https://arxiv.org/html/2606.05194#bib.bib19)\], and targeted fixes have been proposed\[[41](https://arxiv.org/html/2606.05194#bib.bib41),[94](https://arxiv.org/html/2606.05194#bib.bib94)\], but none connect temporal geometry to temporal decision\-making\. Separately, token\-level lookahead over discrete sequence positions is detectable via probing\[[75](https://arxiv.org/html/2606.05194#bib.bib75),[97](https://arxiv.org/html/2606.05194#bib.bib97)\], but operates over next\-token predictions rather than real\-valued time horizons; both modes fail at long\-horizon decisions\[[106](https://arxiv.org/html/2606.05194#bib.bib106),[12](https://arxiv.org/html/2606.05194#bib.bib12)\]\.
##### LLM Economic Behavior and Risk Preference\.
LLMs reproduce behavioral\-economic biases\[[48](https://arxiv.org/html/2606.05194#bib.bib48),[16](https://arxiv.org/html/2606.05194#bib.bib16),[13](https://arxiv.org/html/2606.05194#bib.bib13),[89](https://arxiv.org/html/2606.05194#bib.bib89),[63](https://arxiv.org/html/2606.05194#bib.bib63),[62](https://arxiv.org/html/2606.05194#bib.bib62),[10](https://arxiv.org/html/2606.05194#bib.bib10)\]with unstable risk preferences\[[104](https://arxiv.org/html/2606.05194#bib.bib104)\]\. Risk and time preferences can be steered neurally\[[116](https://arxiv.org/html/2606.05194#bib.bib116),[115](https://arxiv.org/html/2606.05194#bib.bib115)\]but entangle through discount factors\[[26](https://arxiv.org/html/2606.05194#bib.bib26),[73](https://arxiv.org/html/2606.05194#bib.bib73)\]\. Temporal preferences have been studied only behaviorally\[[68](https://arxiv.org/html/2606.05194#bib.bib68)\]; whether they form a steerable activation\-space direction, as shown for truthfulness\[[67](https://arxiv.org/html/2606.05194#bib.bib67),[112](https://arxiv.org/html/2606.05194#bib.bib112)\]and emotion\[[96](https://arxiv.org/html/2606.05194#bib.bib96)\], is open\.
##### Steering Advancements\.
Representation engineering\[[117](https://arxiv.org/html/2606.05194#bib.bib117)\]has progressed from activation addition\[[100](https://arxiv.org/html/2606.05194#bib.bib100),[78](https://arxiv.org/html/2606.05194#bib.bib78)\]and representation interventions\[[109](https://arxiv.org/html/2606.05194#bib.bib109)\]through sparse dictionaries\[[18](https://arxiv.org/html/2606.05194#bib.bib18),[21](https://arxiv.org/html/2606.05194#bib.bib21)\]to geometric approaches\[[101](https://arxiv.org/html/2606.05194#bib.bib101),[113](https://arxiv.org/html/2606.05194#bib.bib113),[84](https://arxiv.org/html/2606.05194#bib.bib84)\]\. Prompting still leads on many benchmarks\[[110](https://arxiv.org/html/2606.05194#bib.bib110),[23](https://arxiv.org/html/2606.05194#bib.bib23)\], and over\-steering degrades helpfulness\[[108](https://arxiv.org/html/2606.05194#bib.bib108),[3](https://arxiv.org/html/2606.05194#bib.bib3)\]\. Patching along continuous numeric directions produces monotonic output shifts\[[45](https://arxiv.org/html/2606.05194#bib.bib45)\], but static vectors assume a fixed concept direction; when the effective direction varies with context or curves through activation space, they become misaligned or unreliable\[[64](https://arxiv.org/html/2606.05194#bib.bib64),[87](https://arxiv.org/html/2606.05194#bib.bib87),[9](https://arxiv.org/html/2606.05194#bib.bib9)\]\.
No prior work has localized a subgraph functionally responsible for temporal preference, characterized the geometry of the causal representation, or steered along it\. Table[B\.1](https://arxiv.org/html/2606.05194#A2.T1)situates our contribution against the most directly comparable works on six axes\.
WorkConceptCausalsubgraphGeometrySteeringLayer\-localizedsteeringHumanbaselineLatent\+\+explicit*Temporal representation*Gurnee and Tegmark\[[38](https://arxiv.org/html/2606.05194#bib.bib38)\]space/time✗✗✗✗✗✗Engels et al\.\[[24](https://arxiv.org/html/2606.05194#bib.bib24)\]days/months✓✓partial✗✗✗Modell et al\.\[[72](https://arxiv.org/html/2606.05194#bib.bib72)\]theory✗✓✗✗✗✗Gurnee et al\.\[[39](https://arxiv.org/html/2606.05194#bib.bib39)\]counting✓✓✗✗✗✗*Temporal reasoning and planning*Wang et al\.\[[106](https://arxiv.org/html/2606.05194#bib.bib106)\]planning✗✗✗✗✗✗Sehgal et al\.\[[91](https://arxiv.org/html/2606.05194#bib.bib91)\]deadlines✗✗✗✗✗✗*LLM economic behavior and preference*Zhu et al\.\[[116](https://arxiv.org/html/2606.05194#bib.bib116)\]risk pref\.✗✗✓✓✗✓Mazyaki et al\.\[[68](https://arxiv.org/html/2606.05194#bib.bib68)\]temporal pref\.✗✗✗✗✓✗Horton et al\.\[[48](https://arxiv.org/html/2606.05194#bib.bib48)\], Cook et al\.\[[16](https://arxiv.org/html/2606.05194#bib.bib16)\]economic bias✗✗✗✗✓✗*Steering methods*Turner et al\.\[[100](https://arxiv.org/html/2606.05194#bib.bib100)\], Panickssery et al\.\[[78](https://arxiv.org/html/2606.05194#bib.bib78)\]generic✗✗✓✓✗✗Marks and Tegmark\[[67](https://arxiv.org/html/2606.05194#bib.bib67)\]truthfulness✓✗✓✗✗✗Sofroniew et al\.\[[96](https://arxiv.org/html/2606.05194#bib.bib96)\]emotion✓✓✓✗✓partialThis worktemporal pref\.✓✓✓✓✓✓
Table B\.1:Our contribution against the closest prior work on six axes: \(i\) whether a causal subgraph is identified, not just a probe direction; \(ii\) whether concept geometry is non\-linear / curved; \(iii\) whether steering traverses a dimensional axis rather than a binary contrast; \(iv\) whether steering is layer\-localized rather than applied uniformly; \(v\) whether outcomes are benchmarked against human behavior; \(vi\) whether both latent \(no\-horizon\) and explicitly parameterized prompts are analyzed together\. No prior work covers all six for temporal preference\.
## Appendix Appendix CMethodology summary
Our methodology follows three stages:*localize*the subgraph,*characterize*the representations, and*intervene*\. Each stage has one or more dedicated experimental pipelines, and each pipeline has its own full\-detail methodology appendix collected in Part 4 \(see §[C\.5](https://arxiv.org/html/2606.05194#A3.SS5)\)\.
Localize SubgraphsPARAMETRIC \+ CONTRASTIVE \+ CLASSIFICATIONQUERYINGCharacterize GeometryACTIVATION SUBSPACEPCAInterveneSTEERING ALONGMANIFOLD
ParametricQueryingGeometry:g\(time\)BehavioralModelingBehavior:b\(time\)Behavior:b\(geometry\)GEOMETRIC INTERVENTION
Figure C\.1:Overview of our approach\. Parametric querying could help usreparametrizeour behavioral modeling as a function of activation\-space geometry instead of an explicit time horizon\.### C\.1Complementary localizations
We perform experiments with three different querying techniques applied to three corresponding prompting settings\.
The first,wide attribution, combines contrastive querying with attribution patching and probing\. Minimally\-framed prompts elicit latent preferences without explicit temporal cues, while gradient\-based approximations efficiently score component importance\. This pipeline scales across samples, aggregating signal from hundreds of diverse prompts\.
The second,targeted intervention, combines parametric querying with activation patching\. Highly\-structured prompts specify explicit time horizons, while direct interventions establish the causal effect\. This pipeline disentangles causal relationships on carefully designed prompt variations\.
The third,targeted classification intervention, combines the temporal classification task with activation patching\. Each prompt presents a single goal whose horizon must be inferred and queries its short/long classification directly\. This pipeline isolates temporal reasoning from valuation and tests whether the subgraph identified above is also recruited for categorical horizon judgment\.
### C\.2Characterizing via geometry
We apply PCA\[[95](https://arxiv.org/html/2606.05194#bib.bib95)\]to residual\-stream activations at subgraph nodes, examining how explicit time\-horizon constraints \(seconds to centuries\) organize the activation manifold and whether latent preferences, elicited without any horizon cue, align to this geometry\. We pay particular attention to the user\-to\-assistant turn transition, where the model converts off\-policy context into on\-policy generation\. In principle, each prompt’s explicit horizon maps to a point on the manifold, opening the possibility of reparametrizing behavioral discount as a function of geometry rather than time \(Figure[C\.1](https://arxiv.org/html/2606.05194#A3.F1)\)\. We do not pursue this reparametrization fully here, but the geometry results in[Appendix M](https://arxiv.org/html/2606.05194#A13)lay the groundwork\.
### C\.3Behavioral analysis
We probe temporal preference at the behavioral level through two experiments\. First, we administer the Kirby MCQ\-27\[[57](https://arxiv.org/html/2606.05194#bib.bib57)\]under multiple personas and response modes, fitting hyperbolic discount functions and introducing a*decision boundary method*that binary\-searches the delayed reward to locate per\-item indifference points\. Second, we test behavioral coherence: whether the model’s choices respect the time\-horizon constraint \(Section[2](https://arxiv.org/html/2606.05194#S2)\)\. Choosing an option that cannot deliver within the specified deadline is instrumentally incoherent, and we systematically vary horizon, reward, order, label, and context to separate genuine temporal reasoning from surface heuristics\. The behavioral experiments serve a dual role: they characterize the model’s temporal preferences independently of the mechanistic analysis, and they reveal the gap between the internal representation \(rich, ordinal, geometrically structured\) and the behavioral output \(discrete, order\-biased, partially incoherent\)\.
### C\.4Steering in the wild
We test causal control over temporal preference using Contrastive Activation Addition \(CAA\)\[[100](https://arxiv.org/html/2606.05194#bib.bib100),[78](https://arxiv.org/html/2606.05194#bib.bib78)\]\. Logistic probes\[[74](https://arxiv.org/html/2606.05194#bib.bib74),[56](https://arxiv.org/html/2606.05194#bib.bib56)\]trained on the implicit dataset identify where temporal orientation is linearly decodable; the probe direction at the best layer yields a steering vector𝐯^CAA\\hat\{\\mathbf\{v\}\}\_\{\\text\{CAA\}\}injected as𝐡\(l\)←𝐡\(l\)\+α⋅𝐯^CAA\\mathbf\{h\}^\{\(l\)\}\\leftarrow\\mathbf\{h\}^\{\(l\)\}\+\\alpha\\cdot\\hat\{\\mathbf\{v\}\}\_\{\\text\{CAA\}\}\. We evaluate via forced\-choice log\-probability shifts and open\-ended generation scored by an external LLM judge, sweeping layers andα\\alphato test for a probing–steering dissociation: whether the best layer for*reading*temporal preference differs from the best layer for*writing*it\[[44](https://arxiv.org/html/2606.05194#bib.bib44)\]\.
### C\.5Overview of the extended methodologies
Full protocol\-level details for each pipeline, including dataset construction, sample sizes, prompt formats, component\-selection thresholds, and analysis procedures, are in Part 4 of the appendices\. A reader looking for one specific experiment’s full methodology can jump directly to the corresponding appendix below; the four\-part organization mirrors the localize/characterize/intervene pipeline:
##### Localize \(four pipelines\)\.
- •[Appendix U](https://arxiv.org/html/2606.05194#A21)– Logistic\-probe training protocol, activation extraction, and the token\-position correction applied to the contrastive dataset\.
- •[Appendix V](https://arxiv.org/html/2606.05194#A22)– EAP\-IG attribution on minimally\-framed contrastive prompts, with bias controls and the component\-taxonomy thresholds used to define the candidate subgraph\.
- •[Appendix W](https://arxiv.org/html/2606.05194#A23)– Activation patching on parametric prompts: noise/denoise protocol, position alignment across horizons, and the metric used to score each \(layer, component\) cell\.
- •[Appendix X](https://arxiv.org/html/2606.05194#A24)– Directional activation patching on classification prompts: dataset construction, model validation and metric definitions; tests whether the same layers are recruited for both preference and categorical horizon inference\.
##### Characterize \(three pipelines\)\.
- •[Appendix Y](https://arxiv.org/html/2606.05194#A25)– PCA geometry pipeline: layer selection, variance\-explained thresholds, and the turn\-boundary analysis\.
- •[Appendix Z](https://arxiv.org/html/2606.05194#A26)– Kirby MCQ\-27 instrument and the decision\-boundary binary\-search extension, including persona and response\-mode conditions\.
- •[Appendix AA](https://arxiv.org/html/2606.05194#A27)– 30\-model investment\-coherence instrument: horizon×\\timesreward×\\timesorder×\\timeslabel×\\timescontext grid and parse protocol\.
##### Intervene \(one pipeline\)\.
- •[Appendix AB](https://arxiv.org/html/2606.05194#A28)– CAA vector construction from the best probing layer, theα\\alpha\-sweep protocol, and the forced\-choice / open\-ended evaluation setup\.
##### Case study\.
- •[Appendix AC](https://arxiv.org/html/2606.05194#A29)– Worked token\-level case study for a single highly\-formatted prompt pair, tying the attribution, patching, and probing signals to specific tokens\.
## Appendix Appendix DExperimental details
Full details for each experiment are in the corresponding methodology appendix\. All experiments can be run on a MacBook Pro \(M4 Max, 48 GB\), except for the causal contrastive one \([Appendix X](https://arxiv.org/html/2606.05194#A24)\), which requires 79 GB\. The full pipeline reproduces end\-to\-end within two weeks\.
### D\.1WhyQwen3\-4B\-Instruct\-2507?
We selectQwen3\-4B\-Instruct\-2507\[[85](https://arxiv.org/html/2606.05194#bib.bib85)\], the non\-thinking\-only mode\-specialized refresh ofQwen3\-4B\[[111](https://arxiv.org/html/2606.05194#bib.bib111)\], for three reasons:
- •Non\-thinking keeps cognition inside a fixed template\.The model operates exclusively in non\-thinking mode: it never emits a<think\>\.\.\.</think\>reasoning block, so the token positions we patch into are stable across clean and corrupted runs\. All “cognition” happens inside the fixed prompt template, which is the alignment condition that activation patching and EAP\-IG attribution both require\. The hybrid\-thinkingQwen3\-4Bwould produce variable\-length reasoning blocks that break this alignment\.
- •Stable latent preference under perturbation\.Localization requires that the model’s answer does not flip under minor syntactic changes\.Qwen3\-4B\-Instruct\-2507satisfies this: across the 30\-model behavioral panel \([Appendix P](https://arxiv.org/html/2606.05194#A16)\), it is among the most label\-stable and context\-stable checkpoints at its scale, while similarly\-sized open\-weight models drift under perturbation\.
- •Tractable to sweep\.Qwen3\-4BoutperformsQwen2\.5\-7Bon most benchmarks and competes withQwen2\.5\-14B\-Instruct,Gemma\-3\-12B\-IT, andPhi\-4\[[111](https://arxiv.org/html/2606.05194#bib.bib111)\], yet the 4B footprint lets us run the full attribution\-plus\-patching\-plus\-steering pipeline on a single MacBook, with the classification patching as the only memory\-bound exception\.
The same mode specialization that enables this analysis also exposes the behavioral gap we study: the non\-thinking variant collapses the hybrid\-thinking checkpoint’s graded horizon curve into three discrete order\-biased modes \([Appendix P](https://arxiv.org/html/2606.05194#A16)\) even though its internal temporal geometry remains rich \([Appendix M](https://arxiv.org/html/2606.05194#A13)\)\.
### D\.2Datasets
- •Minimally\-framed\.Minimally\-framed A/B prompts: an*explicit*set \(500 pairs, 25 categories\) with overt temporal markers, and an*implicit*set \(500 pairs, 10 categories\) using only semantic framing\. Both are counterbalanced across two orderings and seven label schemes\.
- •Highly\-formatted\.4,588 investment intertemporal choice prompts with optional horizon constraints ranging from seconds to centuries\.
- •Classification\-oriented\.160 IOI\-style short/long classification pairs, each presenting a single goal across 25 life subdomains, with balanced question order \(80 SL, 80 LS\)\.
- •Behavioral\.Kirby MCQ\-27 administered under 8 conditions \(2 personas×\\times2 response modes\), plus a binary\-search decision\-boundary extension\.
- •Steering evaluation\.20 held\-out forced\-choice and 13 open\-ended prompts scored by an external LLM judge\.
### D\.3Subset selection criteria
Several analyses operate on filtered subsets of the datasets above; we collect the filtering rules here for cross\-reference\.
- •n=71n=71highly\-formatted parametric contrastive pairs\(used for activation patching in[Appendix J](https://arxiv.org/html/2606.05194#A10)\)\. Filtered from the 4,588 generated highly\-formatted samples by selecting pairs with valid clean/corrupted alignment under the piecewise\-linear position mapping \([Appendix W](https://arxiv.org/html/2606.05194#A23)\) and a non\-trivial baseline logit difference\|yclean−ycorrupted\|\|y\_\{\\text\{clean\}\}\-y\_\{\\text\{corrupted\}\}\|, giving 71 pairs that admit faithful patching\.
- •n=57n=57constrained subset\([Appendix N](https://arxiv.org/html/2606.05194#A14)\)\. The subset of those 71 pairs in which both clean and corrupted prompts carry an explicit time horizon\.
- •n=10n=10unconstrained subset\([Appendix N](https://arxiv.org/html/2606.05194#A14)\)\. The subset of those 71 pairs in which neither prompt carries a time horizon\. The remaining 4 pairs are mixed \(one constrained, one not\) and are analyzed separately in[Appendix AC](https://arxiv.org/html/2606.05194#A29)\.
- •n=160n=160classification pairs\([Appendix K](https://arxiv.org/html/2606.05194#A11)\)\. Filtered from 200 generated short/long classification pairs by retaining only those thatQwen3\-4B\-Instruct\-2507classifies correctly, yielding the 80% \(160/200\) accuracy filter documented in[E\.3](https://arxiv.org/html/2606.05194#A5.SS3)\.
- •n=1,100n=1\{,\}100temporal\-preference samples for error monitoring\([Appendix R](https://arxiv.org/html/2606.05194#A18)\)\. Drawn from the 500\-pairDexplicitD\_\{\\text\{explicit\}\}and 500\-pairDimplicitD\_\{\\text\{implicit\}\}minimally\-framed datasets and used as the temporal arm of the shared error/temporal probe pipeline; no additional filtering beyond the dataset construction in[Appendix E](https://arxiv.org/html/2606.05194#A5)\.
## Appendix Appendix EPrompting settings
We use two prompting settings that probe temporal preference at different levels of abstraction and one that queries temporal reasoning without valuation\. The*minimally\-framed*setting is purely contrastive: it presents a binary choice between a short\-horizon and a long\-horizon option, with no explicit time or reward values\. This captures temporal preference as a binary concept \(present vs\. future\)\. The*highly\-formatted*setting is both contrastive and parametric: it can elicit the same binary preference, but it also sweeps explicit time horizons from seconds to centuries, treating time as a continuous, dimensional concept\. Together, the two settings let us study temporal preference both as a categorical distinction and as a graded quantity \(Figure[E\.1](https://arxiv.org/html/2606.05194#A5.F1)\)\. Separately, the*classification setting*drops the option contrast entirely: each prompt presents a single goal and asks the model to classify its horizon directly, probing temporal reasoning rather than preference, as a test of whether the subgraph generalizes across cognitive operations\.
Figure E\.1:Thecontrastiveparadigm \(left\) treats temporal preference as a binary distinction, while theparametricparadigm \(right\) treats it as a dimensional concept by varying time horizon, reward magnitude, and delay\.### E\.1Minimally\-framed prompts
The minimally\-framed setting queries the model’s temporal preference via forced\-choice A/B prompts\. Each prompt presents a scenario and two candidate responses \(one short\-horizon and one long\-horizon\) and the model selects a label\. We construct two complementary datasets under this paradigm, probing the same underlying distinction under different cueing regimes\.
Each pair is structured as\{question,immediate\_choice,long\_term\_choice,category\}\\\{\\texttt\{question\},\\texttt\{immediate\\\_choice\},\\texttt\{long\\\_term\\\_choice\},\\texttt\{category\}\\\}, with A/B label positions randomized \(≈\\approx50/50 split\) to control for positional bias\. Both datasets comprise 500 paired examples balanced across categories\. For probing experiments \([Appendix U](https://arxiv.org/html/2606.05194#A21)\), a 300\-pair subset of the implicit dataset is used as the primary training and evaluation set\.
#### E\.1\.1Explicit dataset \(DexplicitD\_\{\\text\{explicit\}\}\)
The explicit dataset contains 500 examples in which temporal scope is directly specified through overt temporal expressions\. The contrast between options is signaled by phrases denoting near\-term deadlines, short planning horizons, annual targets, or multi\-year timeframes; the distinction is lexically explicit\.
PROMPTInnovation should aim for:\(A\) Products we can ship this quarter\(B\) Technologies that matter in 20 years
RESPONSE\(A\)
Figure E\.2:Example of anexplicitminimally\-framed prompt\. The temporal contrast is carried by surface vocabulary \(“this quarter” vs\. “20 years”\)\.The dataset is evenly distributed across 25 topical categories with 20 examples per category, spanning planning, decision\-making, problem\-solving, investment, prioritization, goal\-setting, action, vision, risk management, resource allocation, strategy formation, change management, innovation, leadership, communication, evaluation, learning, adaptation, hiring, product development, customer relations, financial planning, team building, market entry, and crisis response\.
#### E\.1\.2Implicit dataset \(DimplicitD\_\{\\text\{implicit\}\}\)
The implicit dataset contains 500 examples in which temporal scope is encoded through semantic framing rather than explicit temporal markers\. The short\-horizon response emphasizes immediate containment, execution, or preservation, whereas the long\-horizon response emphasizes redesign, investment, transformation, or compounding effects, without directly invoking time\-related language\.
PROMPTOrganizational focus should be on\(A\) Firefighting and troubleshooting\(B\) Culture\-building and capability development
RESPONSE\(B\)
Figure E\.3:Example of animplicitminimally\-framed prompt\. Neither option contains temporal vocabulary; the distinction is carried entirely by semantic framing\.The dataset is balanced across 10 abstract contrast categories with 50 examples per category:crisis\_vs\_foundation,harvest\_vs\_cultivate,execute\_vs\_design,react\_vs\_anticipate,preserve\_vs\_transform,tactical\_vs\_strategic,consume\_vs\_invest,fix\_vs\_build,survive\_vs\_thrive, andcapture\_vs\_compound\. This design isolates temporal reasoning from lexical cues, enabling evaluation of whether the model relies on semantic abstractions rather than explicit time indicators\.
#### E\.1\.3LLM\-Assisted Generation and Validation
To construct theDexplicitD\_\{\\text\{explicit\}\}andDimplicitD\_\{\\text\{implicit\}\}datasets at scale while maintaining strict control over linguistic variables, we used a multi\-stage, LLM\-assisted generation and verification pipeline\. The initial candidate pairs for both datasets were generated usingClaude Sonnet 4\.6\. To ensure these generated pairs adhered to our contrastive constraints and were free of unintended confounds, we implemented an automated validation framework\. Each candidate pair was independently evaluated and scored by bothClaude Sonnet 4\.6andGemini 3 Flash\.
The models verified the pairs across four strict dimensions:lexical confounds\(ensuring the implicit set contained absolutely no explicit temporal keywords and that vocabulary complexity was balanced\),surface form\(matching character length and grammatical structure\),semantic confounds\(aligning formality, hedging, specificity, and sentiment\), andcontent validity\(guaranteeing a unidimensional, unambiguous distinction between immediate and long\-term choices\)\. Both models scored these factors on a 1 to 5 scale\. Pairs falling below the acceptable quality threshold \(an average score<3\.5<3\.5\) were iteratively revised or discarded\. Finally, to eliminate positional bias, the presentation order of the immediate and long\-term options \(A/B\) was randomized in the finalized datasets\.
### E\.2Highly\-formatted prompts
The highly\-formatted setting uses structured prompts with explicit section markers that ensure consistent token\-position alignment across prompt variants\. Each prompt contains the following fields:
- •SITUATION:Domain context \(e\.g\., household financial planning\)
- •TASK:The role, task, and two labeled options with reward amounts and time horizons
- •OBJECTIVE:Instruction to deliberate
- •CONSTRAINT:An explicit time\-horizon constraint \(e\.g\., “1 year,” “5 centuries,” or omitted for no\-constraint prompts\)
- •ACTION:Instruction to select one option
- •FORMAT:Response template specifying “I choose:<label\>” and “My reasoning:<text\>”
The shared structural markers serve as*anchors*for a position mapping that aligns token indices between prompts of different lengths\. When clean and corrupted prompts differ in token count \(e\.g\., because different time horizons require different numbers of tokens\), the position mapping uses these anchors to interpolate correctly during activation patching \([Appendix W](https://arxiv.org/html/2606.05194#A23)\)\.
PROMPTSITUATION:Plan for the future of the household based on the stated objectives and constraints\.TASK:You, the head of the household, are tasked to choose the best investment:a\)20,000 dollars in 6 months\.b\)500,000 dollars in 10 years\.OBJECTIVE:Think deeply about which option is preferable\.CONSTRAINT:You must select the option that provides the greatest benefit for this time horizon:8 months\.ACTION:Select one of the two options\. Provide reasoning on why this choice was made\.FORMAT:Respond in this format:I choose:<a\) or b\)\>\. My reasoning:<reasoning in 1\-3 sentences\>
RESPONSEI choose:b\)\. My reasoning: Although the immediate…
Figure E\.4:Example of a highly\-formatted prompt\. Theconstraintsection is optional; omitting it queries the model’s latent preference\.#### E\.2\.1Parametric variation
The experiment configuration sweeps over several axes to systematically vary the temporal context:
- •Reward range: Logarithmic steps between a minimum and maximum \(e\.g\., $1,000–$100,000\)
- •Time range: Logarithmic steps for both short\-term and long\-term options
- •Time horizons: 17 values fromnull\(no constraint\) through seconds, hours, days, weeks, months, years, decades, to centuries
This yields a grid of contrastive pairs that disentangle the effects of reward magnitude, delay, and horizon constraint on internal representations\.
[⬇](data:text/plain;base64,ICAgIHsKICAgICAgICAibmFtZSI6ICJpbnZlc3RtZW50X2dlb21ldHJ5IiwKICAgICAgICAiY29udGV4dCI6IHsKICAgICAgICAgICAgInJld2FyZF91bml0IjogImRvbGxhcnMiLAogICAgICAgICAgICAicm9sZSI6ICJ0aGUgaGVhZCBvZiB0aGUgaG91c2Vob2xkIiwKICAgICAgICAgICAgInNpdHVhdGlvbiI6ICJQbGFuIGZvciB0aGUgZnV0dXJlIG9mIHRoZSBob3VzZWhvbGRzLiIsCiAgICAgICAgICAgICJ0YXNrX2luX3F1ZXN0aW9uIjogImNob29zZSB0aGUgYmVzdCBpbnZlc3RtZW50IiwKICAgICAgICAgICAgImRvbWFpbiI6ICJmaW5hbmNlIgogICAgICAgIH0sCiAgICAgICAgIm9wdGlvbnMiOiB7CiAgICAgICAgICAgICJzaG9ydF90ZXJtIjogewogICAgICAgICAgICAgICAgInJld2FyZF9yYW5nZSI6IFsxMDAwLCAxMDAwMDBdLAogICAgICAgICAgICAgICAgInRpbWVfcmFuZ2UiOiBbCiAgICAgICAgICAgICAgICB7InZhbHVlIjogMSwgInVuaXQiOiAiZGF5cyJ9LAogICAgICAgICAgICAgICAgeyJ2YWx1ZSI6IDIwLCAidW5pdCI6ICJ5ZWFycyJ9CiAgICAgICAgICAgICAgICBdLAogICAgICAgICAgICAgICAgInJld2FyZF9zdGVwcyI6IFsyLCAibG9nYXJpdGhtaWMiXSwKICAgICAgICAgICAgICAgICJ0aW1lX3N0ZXBzIjogWzUsICJsb2dhcml0aG1pYyJdCiAgICAgICAgICAgIH0sCiAgICAgICAgICAgICJsb25nX3Rlcm0iOiB7CiAgICAgICAgICAgICAgICAicmV3YXJkX3JhbmdlIjogWzEwMDAsIDEwMDAwMF0sCiAgICAgICAgICAgICAgICAidGltZV9yYW5nZSI6IFsKICAgICAgICAgICAgICAgICAgICB7InZhbHVlIjogMSwgInVuaXQiOiAieWVhcnMifSwKICAgICAgICAgICAgICAgICAgICB7InZhbHVlIjogMTAwLCAidW5pdCI6ICJ5ZWFycyJ9CiAgICAgICAgICAgICAgICBdLAogICAgICAgICAgICAgICAgInJld2FyZF9zdGVwcyI6IFsyLCAibG9nYXJpdGhtaWMiXSwKICAgICAgICAgICAgICAgICJ0aW1lX3N0ZXBzIjogWzUsICJsb2dhcml0aG1pYyJdCiAgICAgICAgICAgIH0KICAgICAgICB9LAogICAgICAgICJ0aW1lX2hvcml6b25zIjogWwogICAgICAgICAgICBudWxsLAogICAgICAgICAgICB7InZhbHVlIjogMSwgInVuaXQiOiAic2Vjb25kcyJ9LAogICAgICAgICAgICB7InZhbHVlIjogMSwgInVuaXQiOiAiaG91cnMifSwKICAgICAgICAgICAgeyJ2YWx1ZSI6IDEsICJ1bml0IjogImRheXMifSwKICAgICAgICAgICAgeyJ2YWx1ZSI6IDEsICJ1bml0IjogIndlZWsifSwKICAgICAgICAgICAgeyJ2YWx1ZSI6IDEsICJ1bml0IjogIm1vbnRocyJ9LAogICAgICAgICAgICB7InZhbHVlIjogMiwgInVuaXQiOiAibW9udGhzIn0sCiAgICAgICAgICAgIHsidmFsdWUiOiA2LCAidW5pdCI6ICJtb250aHMifSwKICAgICAgICAgICAgeyJ2YWx1ZSI6IDEsICJ1bml0IjogInllYXJzIn0sCiAgICAgICAgICAgIHsidmFsdWUiOiAzLCAidW5pdCI6ICJ5ZWFycyJ9LAogICAgICAgICAgICB7InZhbHVlIjogNSwgInVuaXQiOiAieWVhcnMifSwKICAgICAgICAgICAgeyJ2YWx1ZSI6IDEsICJ1bml0IjogImRlY2FkZXMifSwKICAgICAgICAgICAgeyJ2YWx1ZSI6IDMsICJ1bml0IjogImRlY2FkZXMifSwKICAgICAgICAgICAgeyJ2YWx1ZSI6IDUsICJ1bml0IjogImRlY2FkZXMifSwKICAgICAgICAgICAgeyJ2YWx1ZSI6IDEsICJ1bml0IjogImNlbnR1cmllcyJ9LAogICAgICAgICAgICB7InZhbHVlIjogMiwgInVuaXQiOiAiY2VudHVyaWVzIn0sCiAgICAgICAgICAgIHsidmFsdWUiOiA1LCAidW5pdCI6ICJjZW50dXJpZXMifQogICAgICAgIF0KICAgIH0=)\{"name":"investment\_geometry","context":\{"reward\_unit":"dollars","role":"theheadofthehousehold","situation":"Planforthefutureofthehouseholds\.","task\_in\_question":"choosethebestinvestment","domain":"finance"\},"options":\{"short\_term":\{"reward\_range":\[1000,100000\],"time\_range":\[\{"value":1,"unit":"days"\},\{"value":20,"unit":"years"\}\],"reward\_steps":\[2,"logarithmic"\],"time\_steps":\[5,"logarithmic"\]\},"long\_term":\{"reward\_range":\[1000,100000\],"time\_range":\[\{"value":1,"unit":"years"\},\{"value":100,"unit":"years"\}\],"reward\_steps":\[2,"logarithmic"\],"time\_steps":\[5,"logarithmic"\]\}\},"time\_horizons":\[null,\{"value":1,"unit":"seconds"\},\{"value":1,"unit":"hours"\},\{"value":1,"unit":"days"\},\{"value":1,"unit":"week"\},\{"value":1,"unit":"months"\},\{"value":2,"unit":"months"\},\{"value":6,"unit":"months"\},\{"value":1,"unit":"years"\},\{"value":3,"unit":"years"\},\{"value":5,"unit":"years"\},\{"value":1,"unit":"decades"\},\{"value":3,"unit":"decades"\},\{"value":5,"unit":"decades"\},\{"value":1,"unit":"centuries"\},\{"value":2,"unit":"centuries"\},\{"value":5,"unit":"centuries"\}\]\}
Figure E\.5:Example configuration for theinvestment\_geometryscenario\. Each scenario defines a context, short\- and long\-term option ranges, and a set of time horizons spanning seconds to centuries\.The configuredshort\_termandlong\_termtime ranges overlap \(1 day–20 years and 1 year–100 years\)\. We did not enforce a delay\-ordering filter at sample time, so a small fraction of generated pairs may have realized short\-term delay≥\\geqrealized long\-term delay\. We did not observe systematic effects from this in our analyses, but it is a known dataset caveat\.
### E\.3Classification\-oriented prompts
The classification setting queries the model’s temporal reasoning rather than its preference: each prompt presents a single goal whose horizon must be inferred from world knowledge, and the model produces a direct short/long judgment\. Unlike the minimally\-framed setting \([E\.1](https://arxiv.org/html/2606.05194#A5.SS1)\), there is no A/B option contrast within a single prompt; instead, contrastive structure is established*across*prompt pairs in the IOI style\. Each clean sample is paired with a corrupted sample that shares the same template and differs only in the embedded goal, with the two goals lying on the same life continuum and differing only in temporal horizon\.
Prompt template
"The goal is to <goal\>\. Is this a <short\-term or long\-term / long\-term or short\-term\> goal? The answer is:"
CLEAN PROMPTThe goal is to cook a warm dinner for the family\. Is this a short\-term or long\-term goal? The answer is:
CLEAN RESPONSEshort
CORRUPTED PROMPTThe goal is to become a top chef in the city\. Is this a short\-term or long\-term goal? The answer is:
CORRUPTED RESPONSElong
Figure E\.6:Example of a classification pair of prompts\.All prompts are appended with the chat template before being passed to the model\.
#### E\.3\.1Dataset composition
The dataset contains 160 prompt pairs with perfectly balanced question order: 80 SL pairs \("is this a short\-term or long\-term"\) and 80 LS pairs \("is this a long\-term or short\-term"\), included to mitigate priming effects\. Three temporal cue types signal the long\-term horizon: \(1\)career/mastery, achieving elite status or deep expertise at something; \(2\)growth, transforming something small into something large or established; and \(3\)accumulation, exhaustive scope requiring years of sustained effort\. All goals relate to a general life domain and are distributed fairly evenly across 25 life subdomains \(gardening, cooking, swimming, languages, board games, etc\.\)\.
#### E\.3\.2Design principles
The dataset is governed by four design principles\. First,*token alignment*: all 320 prompts have the same token length underQwen3\-4B\-Instruct\-2507, ensuring positional correspondence across pairs\. In total, each prompt contains 34 tokens, including the chat template, with 7 tokens covering the goal statement\. Second,*semantic overlap within pairs*: each clean and corrupted goal shares the same domain and lies on the same life continuum, differing only in temporal horizon\. Third,*no explicit temporal keywords*: words such as “daily,” “weekly,” or “years” are banned; the model must infer the horizon from world knowledge alone\. Fourth,*unambiguous horizons*: every short\-term goal can be completed in hours or a single sitting, while every long\-term goal requires years of sustained effort\.
#### E\.3\.3Validation and filtering
The original dataset contained 200 pairs\.Qwen3\-4B\-Instruct\-2507successfully classified 160/200 pairs, eliciting 80% accuracy\. Of the 40 misclassified pairs, 25 involve genuinely ambiguous temporal horizons where the model’s interpretation is defensible\. The remaining 15 failures, where the model labels the prompt as short\-term despite clear temporal signals, concentrate in accumulation\-type scholarly activities \("catalog moon lore from old hill folk"\) and growth\-type production at scale \("fire glazed plates for the whole county"\)\. All subsequent activation patching in \([Appendix X](https://arxiv.org/html/2606.05194#A24)\) operates on the 160 successfully classified pairs\.
Table[X\.1](https://arxiv.org/html/2606.05194#A24.T1)reports cue\-type statistics for the 160 surviving pairs; Table[X\.2](https://arxiv.org/html/2606.05194#A24.T2)reports the same statistics grouped by question order\.
Because this dataset serves a narrower purpose — testing whether the temporal\-preference subgraph generalizes to categorical horizon inference, rather than characterizing preference itself — we do not include it in the prompt\-setting comparison in Section[E\.4](https://arxiv.org/html/2606.05194#A5.SS4), which contrasts only the two preference\-oriented paradigms\.
### E\.4Comparison of prompting settings
Table[E\.1](https://arxiv.org/html/2606.05194#A5.T1)summarizes the complementary strengths of the two settings\. The minimally\-framed setting is better suited for probing latent preferences under naturalistic conditions, while the highly\-formatted setting enables controlled parametric sweeps and richer mechanistic analysis\.
Minimally\-framedHighly\-formattedParadigmContrastive only \(binary: short vs\. long\)Contrastive \+ parametric \(binary preference*and*continuous time horizon\)Concept typeBinary \(present vs\. future\)Dimensional \(seconds to centuries\)Prompt structureNo explicit time or reward; model infers temporality from semantic framingExplicit reward amounts, delays, and a constraint field specifying the horizonValidityCloser to on\-policy; low demand characteristicsCloser to off\-policy; structured scaffolding may anchor the modelMechanistic useAttribution, probing, CAA vector constructionAttribution, activation patching with token\-level position mapping, geometry analysisBehavioral modelingBinary preference only; applies to any domainDiscount curves, comparison to human baselines, reparametrization via activation geometryTable E\.1:Comparison of the two prompting settings\. The minimally\-framed setting probes latent binary preference; the highly\-formatted setting adds parametric control over the time dimension\.
## Appendix Appendix FExtended limitations and future work
Time is a complex and entangled concept\. Our work is merely a starting point\.
- •Finer localization\.Full circuit tracing\[[114](https://arxiv.org/html/2606.05194#bib.bib114),[35](https://arxiv.org/html/2606.05194#bib.bib35)\]would identify atomic components and their information flow\. Our EAP\-IG analysis shows attribution mass distributed across many nodes, and whether this reflects genuine distribution or a methodological limitation remains open\.
- •Domain generalization and dataset provenance\.Our approach has several limitations: the pipeline uses only financial scenarios, so findings may not generalize to other domains \(health, career\); contrastive labels were synthetically assigned and lack human validation; the steering vector, derived from controlled off\-policy settings, may capture correlated features rather than purely temporal preference; and the classification dataset entangles temporal horizon with achievement vocabulary\.
- •Scaling across models and variants\.We study onlyQwen3\-4B\-Instruct\-2507; replicating across families and scales would test whether the subgraph location and the probing–steering dissociation generalize\. Comparing this distilled, non\-thinking variant against its thinking counterpart \(Qwen3\-4B\) is particularly compelling: our behavioral analysis shows that chain\-of\-thought dramatically alters temporal preference, but whether reasoning reorganizes the underlying subgraph is unexplored\.
- •Richer parameterization and concept interactions\.Our parametric querying maps time horizon but could extend to reward magnitude, risk, role, and domain to parameterize the full intertemporal choice space\. Temporal preference likely interacts with representations of risk\[[116](https://arxiv.org/html/2606.05194#bib.bib116),[73](https://arxiv.org/html/2606.05194#bib.bib73)\], emotion\[[96](https://arxiv.org/html/2606.05194#bib.bib96)\], and urgency, but we treat the subgraph in isolation\. Moreover, all experiments are single\-turn, yet temporal preference matters most in multi\-turn and agentic settings where representations may shift across turns\[[59](https://arxiv.org/html/2606.05194#bib.bib59)\]\.
- •Non\-linear steering\.Our linear CAA vector approximates a curved manifold; output quality degrades at\|α\|=60\|\\alpha\|\{=\}60\. Methods that follow the manifold’s curvature\[[87](https://arxiv.org/html/2606.05194#bib.bib87),[64](https://arxiv.org/html/2606.05194#bib.bib64),[84](https://arxiv.org/html/2606.05194#bib.bib84)\]could enable stronger, cleaner interventions\.
Part 1: Where is temporal preference?
- •[G](https://arxiv.org/html/2606.05194#A7)\.Linear probing
- •[H](https://arxiv.org/html/2606.05194#A8)\.Attributional contrastive
- •[I](https://arxiv.org/html/2606.05194#A9)\.Attributional parametric
- •[J](https://arxiv.org/html/2606.05194#A10)\.Causal parametric
- •[K](https://arxiv.org/html/2606.05194#A11)\.Causal classification
- •[L](https://arxiv.org/html/2606.05194#A12)\.Cross\-method convergence
## Appendix Appendix GContrastive linear probing results
The four experiments above localized temporal preference through attribution and causal intervention\. Here we take a complementary approach: training logistic regression probes on residual\-stream activations to ask*where*the model linearly encodes the short/long distinction \(methodology in[Appendix U](https://arxiv.org/html/2606.05194#A21)\)\. Unlike the previous methods, probing does not measure causal effect but rather the readability of a concept at each layer\. This distinction will prove important: the best probing layer turns out to differ from the best intervention layer\.
### G\.1Layer\-by\-Layer Probe Accuracy
Logistic regression probes were trained at each of the 36 layers using the protocol described in[Appendix U](https://arxiv.org/html/2606.05194#A21)\.
MetricResultBest layer26Best test accuracy99\.2%Signal above chance\+52\.3 ppCross\-dataset generalizationYes \(see Section[G\.4](https://arxiv.org/html/2606.05194#A7.SS4)\)Table G\.1:Summary of probing results onDimplicitD\_\{\\mathrm\{implicit\}\}\.Accuracy rises steadily from∼\\sim80% at layer 0 to a plateau above 95% around layer 17, reaching 99\.2% at layer 26\. The monotonic increase across layers is consistent with the model progressively refining a linear temporal representation in deeper layers\.
Figure G\.1:Test accuracy of scaled logistic regression probes across all 36 layers onDimplicitD\_\{\\mathrm\{implicit\}\}\(80/20 pair\-aware split\)\. Dashed lines indicate chance \(50%\) and the strong\-signal threshold \(70%\)\. The best layer \(26\) is marked\.
### G\.2Shuffled\-Label Control
To confirm that probe accuracy reflects genuine temporal structure rather than geometric artifacts of the activation space, we train 10 probes per layer on randomly permuted labels using the same scaled activations\. Shuffled accuracy was approximately 50% at every layer, including layer 0\. The gap between real\-label accuracy \(∼\\sim80–99%\) and shuffled\-label accuracy \(∼\\sim50%\) at every layer confirms that the signal is a learned property of the temporal concept, not an intrinsic property of the activation geometry\.
Figure G\.2:Probe accuracy vs\. shuffled\-label control across all 36 layers\. Real\-label probes \(blue\) rise to 99\.2% at layer 26, while shuffled\-label probes \(orange\) remain at chance \(∼\\sim50%\) throughout\.
### G\.3Representation Geometry \(PCA\)
PCA analysis reveals an important asymmetry between the two datasets:
- •Implicit dataset:No separation is visible in the top two principal components \(PC1 explains only 2–5% of variance\)\. The temporal direction is subtle and occupies dimensions that PCA discards\.
- •Explicit dataset:Clear separation is visible in PCA \(PC1==9\.5% at layer 0\), driven by surface vocabulary differences between short\-term and long\-term choices\.
This result is significant: PCA fails to detect the temporal concept in the implicit dataset, yet the supervised probe succeeds at 99\.2%\. The temporal direction is real but non\-obvious; it requires supervised search to find a direction that unsupervised methods miss\. This is consistent with findings on the non\-trivial geometry of concept representations in LLMs\[[24](https://arxiv.org/html/2606.05194#bib.bib24),[67](https://arxiv.org/html/2606.05194#bib.bib67)\]\.
Figure G\.3:PCA projections of layer activations for the implicit \(top\) and explicit \(bottom\) datasets at selected layers\. The implicit dataset shows no visible separation in the top two PCs, while the explicit dataset shows clear clustering driven by surface vocabulary\.
### G\.4Cross\-Dataset Generalization
Probes trained onDimplicitD\_\{\\mathrm\{implicit\}\}were evaluated zero\-shot onDexplicitD\_\{\\mathrm\{explicit\}\}\(different vocabulary, same underlying concept\)\. This tests whether the probe has learned a genuine temporal direction rather than vocabulary\-specific features\. The savedStandardScalerfrom training is re\-applied to the explicit activations before scoring\.
Cross\-dataset accuracy tracks the within\-dataset accuracy closely across all layers, confirming that the probe direction generalizes from implicit semantic cues to explicit temporal markers\. At the best layer \(26\), implicit test accuracy is 99\.2% and cross\-dataset accuracy onDexplicitD\_\{\\mathrm\{explicit\}\}remains above 95%\.
Figure G\.4:Cross\-dataset generalization: probes trained onDimplicitD\_\{\\mathrm\{implicit\}\}evaluated zero\-shot onDexplicitD\_\{\\mathrm\{explicit\}\}\. Accuracy tracks closely across all layers, confirming the probe captures a genuine temporal direction rather than dataset\-specific features\.
### G\.5Summary
Contrastive probing confirms thatQwen3\-4B\-Instruct\-2507maintains a linear temporal direction in its residual stream\. The key findings are:
1. 1\.Strong linear signal\.A logistic regression probe achieves 99\.2% test accuracy at layer 26, with accuracy rising monotonically across layers\.
2. 2\.Shuffled control rules out artifacts\.Probes trained on permuted labels remain at chance \(∼\\sim50%\) at every layer, confirming that the signal reflects genuine temporal structure \(Section[G\.2](https://arxiv.org/html/2606.05194#A7.SS2)\)\.
3. 3\.PCA misses it; supervised search finds it\.The temporal direction is not visible in the top principal components of the implicit dataset, yet a supervised probe recovers it with near\-perfect accuracy \(Section[G\.3](https://arxiv.org/html/2606.05194#A7.SS3)\)\. The concept is real but geometrically subtle\.
4. 4\.Cross\-dataset generalization\.Probes trained on implicit cues transfer zero\-shot to explicit temporal markers, confirming that the learned direction captures a genuine temporal concept rather than vocabulary\-specific features \(Section[G\.4](https://arxiv.org/html/2606.05194#A7.SS4)\)\.
An important dissociation emerges when comparing these probing results with the steering experiments in[Appendix S](https://arxiv.org/html/2606.05194#A19): probing accuracy peaks at layer 26, while steering is most effective at layers 19–22\. This gap suggests that the layers where the model most cleanly*represents*temporal preference are not the same layers where*intervening*on that representation most strongly influences downstream behavior\. We discuss possible explanations for this probing–steering dissociation in Section[S\.3](https://arxiv.org/html/2606.05194#A19.SS3)\.
## Appendix Appendix HAttributional contrastive results
Our first approach to localizing temporal preference uses gradient\-based attribution \(EAP\-IG\) on the minimally\-framed contrastive prompts, where the model chooses between a short\-horizon and a long\-horizon option with no explicit time vocabulary\. This is the cheapest localization method: it approximates causal effect via gradients rather than direct intervention, and the contrastive prompts are short and semantically controlled\. The tradeoff is that the signal may be noisier than causal methods, so we treat the results here as a selection prior rather than ground truth \(methodology in[Appendix V](https://arxiv.org/html/2606.05194#A22)\)\.
The attribution reveals a candidate subgraph comprising approximately 0\.125% of all nodes, concentrated in layers 21–35\. However, as we show below, the circuit is highly diffuse: the median per\-component attribution is well below 0\.1% of the total mass, and even the highest\-scoring outlier \(Figure[H\.1](https://arxiv.org/html/2606.05194#A8.F1)\) sits below∼\\sim1%\. The layers that emerge here \(particularly L24 for attention\) will reappear consistently across the causal and probing experiments that follow \([Appendix J](https://arxiv.org/html/2606.05194#A10),[Appendix K](https://arxiv.org/html/2606.05194#A11),[Appendix G](https://arxiv.org/html/2606.05194#A7)\)\.
### H\.1Attribution Score Distribution
Figure[H\.1](https://arxiv.org/html/2606.05194#A8.F1)shows that attribution mass is distributed across a large number of nodes: the distribution is neither power\-law nor exponential, and the vast majority of components have near\-zero attribution scores\. This poses a fundamental challenge for top\-kkselection, as there is little theoretical justification for any particular cutoff when the score distribution lacks a natural elbow or gap\.
Figure H\.1:\(top\) Histogram of attribution scores with logarithmic y\-axis\. The distribution isnotapproximated by either a power law or an exponential\. \(bottom\) Cumulative logit\-normalized component attribution scores for variant\(A\)for the canonical option order with respect to the short\-term concept\.
### H\.2Limitations of EAP\-IG for This Circuit
The attribution results reveal that temporal preference is likely mediated by a highly*diffuse*circuit rather than a sparse, localizable one\. Even the highest\-scoring individual components explain at most a fraction of a percent of the total attribution mass individually, and the bulk of components sit far below 0\.1%\. The sheer number of low\-scoring components dominates the distribution, making top\-kkselection inherently noisy: it is unclear whether selected nodes are genuinely temporal\-preference components or statistical artifacts of aggregation over hundreds of prompt variations\.
Despite these limitations, the EAP\-IG results provide a useful*signal*when interpreted alongside independent methods\. In particular, the layers that emerge as high\-attribution under EAP\-IG \(e\.g\., L24 for attention\) overlap with layers identified by activation patching \([Appendix J](https://arxiv.org/html/2606.05194#A10),[Appendix K](https://arxiv.org/html/2606.05194#A11)\) and CAA steering \([Appendix S](https://arxiv.org/html/2606.05194#A19)\)\. This convergence across independent methodologies suggests that the layer\-level localization is genuine even though component\-level identification via EAP\-IG alone is unreliable\.
We therefore treat EAP\-IG not as a circuit\-identification tool in the traditional sense, but as a*selection prior*: it restricts the search space to nodes enriched for temporal signal, which we then characterize through representational geometry \([Appendix M](https://arxiv.org/html/2606.05194#A13)\) and probing \([Appendix G](https://arxiv.org/html/2606.05194#A7)\)\. Our analysis focuses on representational structure rather than on isolating a minimal causal mechanism, and uses activation patching \([Appendix J](https://arxiv.org/html/2606.05194#A10),[Appendix K](https://arxiv.org/html/2606.05194#A11)\) to establish causal claims independently\.
### H\.3Layer Distribution of Top\-k Components
Despite the diffuse nature of the overall attribution distribution, a clear layer\-level pattern emerges: temporal\-preference signal concentrates in the mid\-to\-upper layers \(approximately layers 21–35\)\. This concentration is robust across different values ofkkand holds for both attention and MLP components\. Critically, this layer range converges with the layers identified independently by parametric activation patching \([Appendix J](https://arxiv.org/html/2606.05194#A10)\) and CAA steering \([Appendix S](https://arxiv.org/html/2606.05194#A19)\), providing cross\-method validation that temporal preference processing is genuinely localized to this subgraph\.
Figure H\.2:Layer\-wise distribution of top\-kkattributed components\. Attribution mass concentrates in layers 21–35, with attention heads peaking around L24 and MLP neurons concentrated in the upper layers \(L31–L35\)\. This layer profile is stable across values ofkk, indicating genuine localization rather than an artifact of the threshold\.Figure H\.3:Heatmap of top\-kkcomponent counts per layer, broken down by component type\. Attention heads are enriched in layers 21–26, while MLP neurons dominate in layers 31–35, consistent with a two\-stage pattern where attention layers carry temporal information and later MLP layers refine it\.Figure[H\.4](https://arxiv.org/html/2606.05194#A8.F4)complements the top\-kkcount analysis by showing mean attribution scores per layer\. The mean score profile confirms that the layers 21–35 concentration is not merely a consequence of having more components selected; these layers also carry higher per\-component attribution mass on average\.
Figure H\.4:Mean attribution score by layer\. Layers 21–35 show elevated per\-component scores, confirming that the mid\-to\-upper layer concentration reflects genuinely higher attribution rather than a selection artifact\. The peak around L24 for attention aligns with activation patching results identifying L24\_attn as the highest\-effect attention component \([Appendix J](https://arxiv.org/html/2606.05194#A10),[Appendix K](https://arxiv.org/html/2606.05194#A11)\)\.
### H\.4Attention vs\. MLP Contributions
Decomposing the top\-kkattributed components by type reveals a division of labor between attention heads and MLP neurons\. Attention heads account for the majority of highly\-attributed components, consistent with their role in routing information across token positions, while MLP neurons contribute a smaller but distinct share concentrated in the upper layers\. This attention\-dominated pattern is consistent with temporal preference relying on contextual integration across the prompt rather than on purely local feature computation\.


Figure H\.5:Left: attention heads dominate the top\-kkattributed components, but the MLP share grows at largerkk\. Right: overall attribution mass split\. Attention carries the majority, reinforcing that temporal\-preference computation is primarily mediated by cross\-position information flow\.The heatmaps below reveal which specific heads and MLP neurons carry the signal\. For attention, a small number of heads in layers 21–26 stand out, while the MLP signal is more diffuse across neurons in layers 31–35\. This spatial separation \(attention in mid\-layers, MLP in upper layers\) is consistent with a two\-phase computation: attention heads first integrate temporal context, then MLP layers transform this into the output representation\.


Figure H\.6:Attribution heatmaps for attention heads \(left\) and MLP neurons \(right\) across layers\. Attention: a sparse set of heads in layers 21–26 carries disproportionate attribution, with L24 heads showing the strongest signal, converging with activation patching results \([Appendix J](https://arxiv.org/html/2606.05194#A10),[Appendix K](https://arxiv.org/html/2606.05194#A11)\)\. MLP: attribution is concentrated in the upper layers \(L31–L35\) and more evenly distributed across neuron indices, suggesting distributed rather than sparse MLP processing\.
### H\.5Short\-Term vs\. Long\-Term Component Comparison
A key question is whether short\-term and long\-term temporal preferences are processed by the same components or by specialized subpopulations\. We compare the top\-kkattributed components for the short\-term concept against those for the long\-term concept\. The results reveal partial but incomplete overlap: many components contribute to both concepts, but each concept also recruits specialized nodes\. This pattern is consistent with a shared temporal\-processing backbone augmented by concept\-specific refinement, supporting the paper’s claim that temporal preference is a structured rather than monolithic representation\.


Figure H\.7:Left: short\-term and long\-term attributed components show partial overlap\. Short\-term attribution peaks at L24, long\-term shifts toward L22\. Right: Jaccard similarity between ST and LT top\-kksets increases withkk, from concept\-specific nodes at smallkkto a shared temporal backbone at largerkk\.At smallkk, the overlap is low, indicating the most important components are concept\-specific\. Askkgrows, overlap increases, reflecting shared temporal processing consistent with the geometric separation in[Appendix M](https://arxiv.org/html/2606.05194#A13)\.
### H\.6Individual Component Analysis
While the diffuse distribution of attribution scores limits confidence in any single component \(see Section[Appendix H](https://arxiv.org/html/2606.05194#A8)limitations discussion above\), examining the highest\-scoring individual nodes provides a useful sanity check\. The top\-ranked nodes cluster in the same mid\-to\-upper layer range identified by layer\-level analysis, and the highest\-attribution attention heads fall in L22–L24, precisely the layers flagged by activation patching as causally important\. However, even the top\-ranked individual components account for less than 0\.1% of total attribution mass, underscoring why we treat EAP\-IG as a selection prior rather than a definitive circuit\-identification tool\.


Figure H\.8:Left: top individual nodes ranked by attribution score; the highest\-ranked are attention heads in L22–L24, with no single component exceeding 0\.1% of total mass\. Right: cumulative tail distribution; attribution mass accumulates slowly, confirming a diffuse circuit \(∼\\sim0\.125% of all nodes\)\.Finally, the full attention head matrices \(Figure[H\.9](https://arxiv.org/html/2606.05194#A8.F9)\) provide a detailed view of which heads matter for each concept\. The short\-term matrix shows concentrated signal in a few heads around L24, while the long\-term matrix distributes attribution more broadly across L22–L26\. This asymmetry suggests that short\-term preference relies on a slightly more focused set of attention heads, whereas long\-term preference recruits a wider subnetwork\.


Figure H\.9:Attention head attribution matrices for short\-term \(left\) and long\-term \(right\) concepts\. Short\-term attribution is concentrated in a sparse set of L24 heads, while long\-term attribution is distributed more broadly across L22–L26, revealing an asymmetry in circuit structure between the two temporal concepts\.
## Appendix Appendix IAttributional parametric results
The previous two appendices applied different methods to different prompts: EAP\-IG on contrastive prompts \([Appendix H](https://arxiv.org/html/2606.05194#A8)\), then activation patching on parametric prompts \([Appendix J](https://arxiv.org/html/2606.05194#A10)\)\. To disentangle the effect of the method from the effect of the prompt paradigm, we apply standard attribution patching \(EAP\-IG\) to the same parametric prompts used for activation patching\. This completes one diagonal of the method×\\timesparadigm matrix and lets us ask: do the layers identified by attribution match those identified by causal intervention on the same data?
### I\.1Attribution score distribution
As with the contrastive attribution results \([Appendix H](https://arxiv.org/html/2606.05194#A8)\), the vast majority of components have near\-zero attribution scores\. Figure[I\.1](https://arxiv.org/html/2606.05194#A9.F1)shows the score distributions under denoising and noising\. The denoising distribution has a heavy positive tail extending to∼\\sim1\.2, while the noising distribution is more symmetric and concentrated near zero \(±\\pm0\.04\)\. The central 50% of scores \(bottom panels, linear scale\) are confined to a very narrow band around zero,±\\pm0\.0003 for denoising and±\\pm0\.00005 for noising, confirming that only a small fraction of components carry meaningful attribution\.


Figure I\.1:Attribution score distributions under denoising \(left\) and noising \(right\)\. Top rows: 1st–99th percentile on log scale\. The denoising distribution has a heavy positive tail extending to∼\\sim1\.2, while the noising distribution is more symmetric \(±\\pm0\.04\)\. Bottom rows: 25th–75th percentile on linear scale, with central mass within±\\pm0\.0003 \(denoising\) and±\\pm0\.00005 \(noising\)\.
### I\.2Top\-scoring components
Figure[I\.2](https://arxiv.org/html/2606.05194#A9.F2)ranks individual components by their attribution scores under denoising and noising\.


Figure I\.2:Top components ranked by denoising \(left\) and noising \(right\) attribution score\. Denoising: residual stream components dominate, with L19, L18, and L20 carrying the largest scores \(∼\\sim300–500\); L25 attention is a notable negative outlier\. Noising: late\-layer residual components \(L26–L35\) carry the largest negative attributions, revealing a sufficiency/necessity asymmetry\.Under denoising, the residual stream at mid\-layers \(L17–L22\) carries the bulk of the recovery signal, with individual scores reaching∼\\sim500\. Attention and MLP components are an order of magnitude smaller\. Under noising, the signal shifts to late layers \(L26–L35\), where corrupted residual activations cause the most disruption\. This asymmetry between where information is built \(mid\-layers\) and where it becomes vulnerable \(late layers\) parallels the sufficiency/necessity gap observed in the causal parametric experiments \([Appendix J](https://arxiv.org/html/2606.05194#A10)\)\.
### I\.3Recovery vs\. disruption by component type
Figure[I\.3](https://arxiv.org/html/2606.05194#A9.F3)\(left\) plots each component’s denoising attribution against its noising attribution, separated by component type\. Figure[I\.3](https://arxiv.org/html/2606.05194#A9.F3)\(right\) shows per\-layer attribution for each component type under both denoising and noising\.


Figure I\.3:Left:Recovery \(denoising\) vs\. disruption \(noising\) attribution, separated by component type\. The residual stream panel operates on a much larger scale than attention or MLP\. Mid\-layer residual components cluster in the high\-recovery / low\-disruption region; late\-layer residual components cluster in the low\-recovery / high\-disruption region\.L31\_mlpstands out as a strong disruptor without proportional recovery\.Right:Per\-layer attribution scores for attention \(top\), MLP \(middle\), and residual stream \(bottom\) under denoising \(blue\) and noising \(red\)\. Attention denoising peaks at L19–L24 with a sharp negative dip at L25\. MLP noising shows a strong negative spike at L31\. The residual stream dominates in scale, with denoising plateauing at L17–L22 and noising growing increasingly negative from L25 onward\.The scatter reveals that attention heads at L19–L24 contribute primarily to recovery \(lower\-right quadrant\) without proportional disruption, consistent with a sufficiency\-biased profile: patching them in restores behavior, but patching them out does not fully destroy it\.L31\_mlpis the clearest disruptor, consistent with its identification as a top\-ranked MLP component in the causal experiments\.
### I\.4Layer\-wise attribution by component type
Figure[I\.4](https://arxiv.org/html/2606.05194#A9.F4)shows the denoising and noising curves overlaid for each component type\.



Figure I\.4:Layer\-wise attribution by component type, with denoising and noising overlaid\.Top:Attention output; denoising peaks at L20–L24 \(∼\\sim55–65\), then drops sharply negative at L25 \(∼−37\\sim\-37\); noising is near zero throughout\.Middle:MLP output; denoising is modestly positive across most layers; noising shows a sharp negative spike at L31 \(∼−42\\sim\-42\)\.Bottom:Residual stream \(post\-MLP\); denoising rises steeply to a plateau at L17–L20 \(∼\\sim150–170\); noising grows increasingly negative from L25 onward \(∼−120\\sim\-120at L35\)\.
### I\.5Position×\\timeslayer heatmaps
The heatmaps below show attribution across token positions and layers for key component types, revealing where in the prompt the temporal signal is concentrated\.
Figure I\.5:Attention output denoising heatmap \(position×\\timeslayer\)\. Attribution is sparse and concentrated in the last∼\\sim30 token positions \(the answer region\)\. The final position shows the strongest signal\.Figure I\.6:Residual stream \(post\-MLP\) denoising heatmap\. Complex vertical\-stripe patterns from position∼\\sim75 onward, with the strongest positive signal at the final position in upper layers \(L34\)\. Early positions \(the prompt preamble\) contribute almost nothing\.Figure I\.7:Residual stream \(post\-MLP\) noising heatmap\. Disruption is concentrated at the final token position in L35, with moderate mixed\-sign activity at positions∼\\sim65–110 in mid\-layers\. The noising signal is sparser and more localized than the denoising signal\.
### I\.6Key findings
1. 1\.Residual stream dominates attribution\.Attribution scores for residual stream components are an order of magnitude larger than for attention or MLP, reflecting the cumulative nature of the residual stream\.
2. 2\.Mid\-layer recovery, late\-layer disruption\.Denoising attribution peaks at L17–L22; noising peaks at L30–L35\. The model builds temporal information through mid\-layers and becomes most vulnerable to corruption in late layers\.
3. 3\.Attention L19–L24 and the L25 anomaly\.Attention heads in L19–L24 contribute to recovery, consistent with the causal parametric results\. L25 attention is a negative outlier under denoising; we do not interpret this causally from attribution alone\.
4. 4\.L31\_mlpas the key disruptor\.L31\_mlpis the single most disruptive non\-residual component, matching its identification in the causal experiments\.
5. 5\.Signal concentrates at late token positions\.Most attribution mass falls in the last∼\\sim30–40 token positions \(the answer / format region\), with the final position consistently the most important\. The prompt preamble carries almost no temporal attribution\.
6. 6\.Convergence with causal methods\.The layers and components identified by gradient\-based attribution match those found by activation patching on the same prompts \([Appendix J](https://arxiv.org/html/2606.05194#A10)\): attention L21–L24, MLP L31 / L35, and the mid\-to\-late residual stream\. This cross\-method agreement validates that both approaches recover the same underlying circuit\.
## Appendix Appendix JCausal parametric results
The attribution results in[Appendix H](https://arxiv.org/html/2606.05194#A8)flagged layers 21–35 but could not establish causal effect\. Here we apply the gold standard: activation patching onn=71n=71highly\-formatted parametric contrastive pairs, directly replacing component activations with counterfactual values \(methodology in[Appendix W](https://arxiv.org/html/2606.05194#A23)\)\. Where EAP\-IG approximates, patching measures the actual behavioral consequence of intervention\.
The results sharpen the picture considerably\. A sparse set of four components,L24\_attn,L21\_attn,L35\_mlp, andL31\_mlp, account for the majority of the causal effect, clearly separated from the rest\. L24 attention, the same layer flagged by EAP\-IG, emerges as the single most important component under both denoising and noising\.
### J\.1Component importance ranking
We begin with a direct ranking of individual components by their causal effect size\. Figure[J\.1](https://arxiv.org/html/2606.05194#A10.F1)shows the top 20 components sorted by the mean of their denoising recovery and noising disruption scores across all contrastive pairs\.
Figure J\.1:Top 20 components ranked by mean effect score \(denoising recovery and noising disruption, with standard deviation across contrastive pairs\)\.L24\_attnranks highest, with a noising disruption score near 0\.56 \(the only component above 0\.5\)\.L21\_attnis the next\-largest attention contributor at∼\\sim0\.31;L35\_mlp\(∼\\sim0\.27\) andL31\_mlp\(∼\\sim0\.24\) lead the MLP components\. The fifth\-ranked component \(L30\_attn\) has roughly half the effect ofL24\_attn, separating the top four from the rest\.The ranking reveals a clear separation between a small number of high\-effect components and a long tail of modest contributors\. Attention components dominate the top of the list, withL24\_attnshowing the largest effect under both denoising and noising\. The most causally important MLP components \(L35\_mlpandL31\_mlp\) rank among the top four overall, with effect sizes comparable toL21\_attn\. The asymmetry between denoising recovery and noising disruption is particularly pronounced for the top attention layers:L24\_attnandL21\_attnshow much higher noising disruption than denoising recovery, suggesting these components are more necessary than sufficient: corrupting them degrades performance substantially, but restoring them alone does not fully recover clean behavior\.
### J\.2Marginal contribution analysis
Figure[J\.2](https://arxiv.org/html/2606.05194#A10.F2)examines the marginal contribution of each layer, defined as the difference in residual stream activations before and after the layer \(resid\_post\[L\]−\-resid\_pre\[L\]\)\. This isolates each layer’s additive contribution to the residual stream\.
Figure J\.2:Marginal contribution per layer \(mean±\\pmstandard deviation across contrastive pairs\), showing sufficiency \(denoising recovery, green\) and necessity \(noising disruption, red\)\. Sufficiency peaks sharply at layers 21–24, with layer 22 showing the single highest spike\. Necessity is flatter and lower, reflecting the distributed nature of disruption\. The high variance in layers 20–25 reflects the sensitivity of these layers to the specific temporal framing used in each contrastive pair\.The sufficiency peak at layers 21–24 indicates that the information added to the residual stream by these layers is disproportionately important for temporal preference\. The necessity curve shows a more gradual rise beginning around layer 19, suggesting that while individual layers beyond the peak contribute less, their cumulative disruption is meaningful\. The elevated variance in the peak region indicates that different contrastive pairs engage these layers to different degrees, consistent with the parametric variation in the experimental design\.
##### Layer 19 as onset\.
Layer 19 deserves attention\. In the single\-pair case study \([Appendix AC](https://arxiv.org/html/2606.05194#A29)\), denoising recovery jumps from∼\\sim0\.05 at L18 to∼\\sim0\.5 at L19, the first layer where patching produces a measurable effect on the output\. Before L19, the residual stream does not yet encode temporal preference in a form that patching can recover\. This onset coincides with the beginning of the steering sweet spot \(layers 19–22;[Appendix S](https://arxiv.org/html/2606.05194#A19)\): the model can be steered at L19 precisely because the temporal computation is just beginning and the representation is still malleable\. By L26 \(the probing peak;[Appendix G](https://arxiv.org/html/2606.05194#A7)\), the computation is complete and the representation is readable but no longer easy to redirect\. The L19 onset, L21–24 peak, and L26 readout form a coherent computational timeline within the subgraph\.
### J\.3Redundancy gap heatmap and layer–component interaction
Figure[J\.3](https://arxiv.org/html/2606.05194#A10.F3)maps the redundancy gap, defined as noising disruption minus denoising recovery, across all layers and component types\. Positive values \(red\) indicate components that are more necessary than sufficient, while negative values \(blue\) indicate components that are more sufficient than necessary\. Figure[J\.3](https://arxiv.org/html/2606.05194#A10.F3)decomposes the layer sweep into separate traces for attention, MLP, and residual stream components, revealing how each component type’s causal effect varies across layers\.
\(a\)Redundancy gap heatmap \(disruption−\-recovery\) by layer and component type\. Strong positive values \(dark red\) in layers 20–23 acrossresid\_pre,resid\_mid, andresid\_postindicate high necessity with low sufficiency, characteristic of components embedded in a redundant processing pipeline where no single intervention can fully restore behavior\. Theattn\_outcolumn shows a localized peak at L24 \(0\.39\), whilemlp\_outvalues remain relatively low throughout, suggesting MLP contributions are less redundantly encoded\.
\(b\)Layer\-by\-layer causal effect forattn\_out,mlp\_out, andresid\_postunder denoising \(left\) and noising \(right\)\. Theresid\_postcurve rises sharply at layer 20 under denoising and saturates near 1\.0 by layer 24, reflecting the cumulative nature of residual stream patching\. Theattn\_outandmlp\_outtraces show complementary peaks: attention peaks at layers 21–24, while MLP contributions are more distributed across layers 22–35\.
Figure J\.3:Redundancy gap heatmap and layer–component interaction analysis\.The heatmap reveals a striking pattern: layers 20–23 show large positive redundancy gaps across nearly all component types, peaking near 0\.39 for the residual stream components \(resid\_preL23,resid\_midL22\)\. This indicates that these layers are deeply embedded in the temporal preference circuit: corrupting them causes severe disruption, but patching in clean activations at only one component is insufficient for full recovery, because the corrupted signal has already propagated through earlier residual connections\. Theattn\_outcomponent at L25 shows a mildly negative gap \(−0\.11\-0\.11\), making it one of the few components where recovery exceeds disruption, suggesting a degree of self\-contained sufficiency at that layer\.
Several patterns emerge from this decomposition\. Under denoising, the residual stream curve exhibits a characteristic sigmoid shape, rising steeply between layers 19 and 24 and then plateauing near full recovery\. This reflects the cumulative nature of the residual stream: once the critical mid\-layer representations are restored, later layers can process them correctly\. Theattn\_outcomponent shows a pronounced peak at layers 21–24 under both denoising and noising, consistent with the component ranking in Figure[J\.1](https://arxiv.org/html/2606.05194#A10.F1)\. MLP contributions, by contrast, are more distributed: under noising,mlp\_outshows elevated disruption across a broad range of late layers \(25–35\), suggesting that MLP components contribute through distributed, incremental processing rather than a single localized intervention\.
### J\.4Noise vs\. denoise and attention vs\. MLP scatterplots
Figure[J\.4](https://arxiv.org/html/2606.05194#A10.F4)plots each layer’s denoising recovery against its noising disruption, separately for each component type\. This reveals whether components are sufficient \(high recovery, low disruption\), necessary \(low recovery, high disruption\), or both\. Figure[J\.4](https://arxiv.org/html/2606.05194#A10.F4)directly compares attention and MLP contributions at each layer, revealing the relative dominance of each component type\.
\(a\)Denoising recovery vs\. noising disruption for each layer, separated by component type \(resid\_post,attn\_out,mlp\_out\)\. Points are colored by layer number\. The quadrant labels indicate the interpretive regime: “AND\-like / necessary” \(upper left\) for high disruption with low recovery, and “OR\-like / sufficient” \(lower right\) for high recovery with low disruption\. Mostresid\_postlayers cluster in the necessary quadrant, whileattn\_outandmlp\_outshow more varied profiles\.
\(b\)Attention vs\. MLP effect size at each layer under denoising recovery \(left\) and noising disruption \(right\)\. Points above the diagonal indicate MLP\-dominant layers; points below indicate attention\-dominant layers\. Under both metrics, mid\-layer points \(L21–L24\) fall well below the diagonal, confirming that attention drives the largest single\-component effects for temporal preference\. Late layers \(L31–L35\) cluster near or above the diagonal under noising, reflecting the distributed MLP contributions in this range\.
Figure J\.4:Noise vs\. denoise scatterplots and attention vs\. MLP comparison\.The scatterplots confirm the redundancy gap analysis from a different angle\. Forresid\_post, the picture is layer\-dependent: L21 and L22 sit in the upper\-left “necessary” quadrant \(high disruption, modest recovery\), L23 sits at the boundary, and L24–L25 move into the upper\-right region \(both high disruption*and*high recovery\), with L26 onwards saturating near the top\-right corner\. Layers below L21 \(L19–L20\) remain in the low\-effect region near the origin\. This pattern is consistent with a redundantly encoded signal in the L21–L23 range that cannot be fully restored by a single\-layer intervention, transitioning at L24 onwards into a regime where single\-layer patches both disrupt and recover the temporal signal\. Forattn\_out, the highest\-layer points \(around L24\) fall near the diagonal, indicating roughly balanced sufficiency and necessity\. Themlp\_outpanel shows most layers near the origin, with a few late layers \(L31, L35\) reaching moderate effect sizes in both directions, consistent with their role as the top\-ranked MLP components\.
Figure[J\.5](https://arxiv.org/html/2606.05194#A10.F5)provides a paired view directly comparing attention and MLP contributions at each layer\.
Figure J\.5:Paired attention vs\. MLP comparison with arrows connecting each layer’s denoising \(circle\) and noising \(square\) scores\. Layers where the arrow points rightward and downward \(e\.g\., L21, L24\) indicate components where noising reveals much stronger attention dominance than denoising\. The trajectories of L24 and L21 show the largest rightward displacement, confirming these attention heads as the most causally important individual components\. Layers L22 and L34 shift upward, reflecting MLP\-dominant noising effects at those layers\.The attention\-vs\-MLP comparison reveals a consistent pattern: in the critical mid\-layer range \(L21–L24\), attention components have substantially larger causal effects than their MLP counterparts\. This asymmetry is especially pronounced under noising disruption, whereL24\_attnandL21\_attnachieve effect sizes of 0\.5–0\.6 while the corresponding MLP components remain below 0\.2\. The paired plot \(Figure[J\.5](https://arxiv.org/html/2606.05194#A10.F5)\) makes this particularly clear: the arrows for L21 and L24 sweep dramatically to the right as we move from denoising to noising, indicating that these attention components become even more dominant when measuring necessity rather than sufficiency\. In contrast, some layers \(L22 in the mid\-range, L34 in the upper layers\) show arrows pointing upward, indicating that their MLP components are more causally important than their attention components, particularly under noising\.
### J\.5Summary
The activation patching results converge on several findings that support the claims in the main text:
1. 1\.Sparse, localized circuit\.A small number of components account for the majority of causal effect on temporal preference\. The top four components \(L24\_attn,L21\_attn,L35\_mlp, andL31\_mlp\) are clearly separated from the remaining components in effect size \(Figure[J\.1](https://arxiv.org/html/2606.05194#A10.F1)\)\.
2. 2\.Attention dominance in mid\-layers\.Attention heads at layers 21 and 24 are the single most causally important components, particularly under noising disruption\. Their effect sizes exceed those of any MLP component by a factor of 2–3x \(Figures[J\.4](https://arxiv.org/html/2606.05194#A10.F4),[J\.5](https://arxiv.org/html/2606.05194#A10.F5)\)\.
3. 3\.Distributed MLP contributions in late layers\.MLP components contribute through a more distributed pattern across layers 25–35, withL35\_mlpandL31\_mlpas the most prominent individual contributors \(Figure[J\.3](https://arxiv.org/html/2606.05194#A10.F3)\)\.
4. 4\.Necessity exceeds sufficiency\.The high\-effect components show a consistent asymmetry: noising disruption exceeds denoising recovery, indicating redundant encoding where no single component is individually sufficient to fully determine temporal preference, but individual components are necessary in the sense that corrupting them substantially degrades performance \(Figures[J\.3](https://arxiv.org/html/2606.05194#A10.F3),[J\.4](https://arxiv.org/html/2606.05194#A10.F4)\)\.
5. 5\.Critical computation window at layers 20–24\.The marginal contribution analysis localizes the most informative residual stream transformations to a narrow five\-layer window \(Figure[J\.2](https://arxiv.org/html/2606.05194#A10.F2)\), consistent with a concentrated computational phase for temporal preference\. Three of five layers are also identified as part of core decision window in activation patching for temporal classification \([Appendix K](https://arxiv.org/html/2606.05194#A11)\)\.
## Appendix Appendix KCausal classification results
The causal parametric results \([Appendix J](https://arxiv.org/html/2606.05194#A10)\) intervene on highly\-formatted parametric prompts and localize components causally important for temporal preference \(valuation\)\. Here we intervene on contrastive classification pairs to localize components causally important for temporal classification \(categorical horizon inference\)\. Convergence between the two pipelines would suggest that the same computational machinery is recruited for both tasks; whether this reflects a common temporal representation or merely a common decision readout requires further analysis we leave to future work\.
The task design follows the IOI style: clean and corrupted prompts represent phrase beginnings awaiting completion with the tokens"short"and"long"\. Each sentence contains a description of a goal and a question about the time horizon of that goal\.
Example
- •Clean:"The goal is to cook a warm dinner for the family\. Is this a short\-term or long\-term goal? The answer is:"\. Expected clean answer"short"as the next predicted token\.
- •Corrupted:"The goal is to become a top chef in the city\. Is this a short\-term or long\-term goal? The answer is:"\. Expected corrupted answer"long"as the next predicted token\.
All prompts are appended with a chat template before being passed to the model\. The complete design of the dataset and the results of its validation are described in[Appendix X](https://arxiv.org/html/2606.05194#A24)\.
Since this pipeline primarily tests convergence with the other paradigms, readers can skip directly to the convergence finding \([7](https://arxiv.org/html/2606.05194#A11.I4.i7)\), with supporting evidence in Sec\.[K\.2\.1](https://arxiv.org/html/2606.05194#A11.SS2.SSS1.Px1)and Sec\.[K\.5](https://arxiv.org/html/2606.05194#A11.SS5)\.
### K\.1Limitations
We flag three limitations of this experiment before turning to its results\.
- •Vocabulary–horizon entanglement\.Long\-term prompts concentrate in achievement framing: Career/Mastery cues account for 46% of pairs \(Table[X\.1](https://arxiv.org/html/2606.05194#A24.T1)\), and prestige markers \(top,professional,master\) appear more in long\-term prompts than in short\-term ones while the latter cluster in casual, domestic framings\. The localized circuits may partially reflect achievement\-vocabulary detection rather than temporal horizon per se\. Two considerations bound this concern\. First, the remaining 54% of pairs use Growth and Accumulation cues, which carry distinct surface vocabulary \(size/scale markers and exhaustive\-scope markers, respectively\) rather than the prestige register that drives the Career/Mastery cluster\. Second, and more directly, the same dominant attention component\(L24\_attn\)also emerges as the top\-ranked component in the parametric pipeline \([Appendix J](https://arxiv.org/html/2606.05194#A10)\), where the surface vocabulary consists of numerical reward amounts and explicit time horizons with no overlap with the achievement\-prestige register identified here\. A circuit whose function is to detect achievement vocabulary would not be expected to dominate causal effect on prompts of the form “$20K in 6 months vs\. $500K in 10 years”\. We nevertheless flag two direct minimal\-pair controls as the cleanest tests: explicit\-temporal\-marker pairs \(e\.g\., “fix the bike before lunch” vs\. “fix old bikes over many years”\) and same\-class patching \(short→\\toshort, long→\\tolong\)\. We leave these to future work\.
- •Selection bias\.We patch only on the 160 pairs thatQwen3\-4B\-Instruct\-2507already classifies correctly, excluding 40 misclassified ones\. 15 of those are clear\-signal failures that concentrate in accumulation and growth cues\. The identified circuit may therefore reflect the successful, Career/Mastery\-dominated classification path rather than the model’s general temporal\-reasoning capability\.
- •Single template, binary horizon\.All 320 prompts share one template \(“The goal is to⟨\\langleX⟩\\rangle\. Is this a⟨\\langleY⟩\\ranglegoal? The answer is:”\) and a binary judgment, so the circuit cannot be tested for graded horizon sensitivity or template\-independence within this paradigm\. Convergence with the parametric setup \([Appendix J](https://arxiv.org/html/2606.05194#A10)\) on attention layer partially mitigates the template concern\.
### K\.2160\-Pair Directional Patching
We first present heatmaps for all tokens, highlighting that the end token accumulates the maximum patching effect\. We then provide per\-layer plots at the end token position with 95% confidence intervals\. We also report layer\-level effects summed across all 34 token positions at the end of the results section forattn\_outandmlp\_out, enabling direct comparison with the position\-aggregated metrics used by EAP\-IG attribution \([Appendix H](https://arxiv.org/html/2606.05194#A8)\) and causal parametric patching \([Appendix J](https://arxiv.org/html/2606.05194#A10)\)\.
#### K\.2\.1Residual Stream
Figure[K\.1](https://arxiv.org/html/2606.05194#A11.F1)shows residual stream patching results for all 34 tokens of the prompt for two flip directions\. We can see that both patterns are quite similar in the global structure\. They show the same three activity bands:
- •early layers \(L0–19\) highlighting the goal statement tokens at positions 7–14:the model reading the cue;
- •middle layers \(L12–26\) highlighting the temporal keywords \(positions 18–22:“short/\-term/or/long/\-term”\):the question machinery;
- •late layers \(L20–35\) concentrated on the end token \(position 33\):the decision\.
In both heatmaps the end column saturates the blue scale from∼\\simL20 downward and represents the location of maximum patching effect\. The qualitative circuit \("goal\-read→\\toquestion\-read→\\todecide at end"\) is the same whether the model is being steered toward"short"or"long"\.

Figure K\.1:Directional residual stream patching averaged over 160 classification pairs\. Top row: denoising for"short"clean and"long"corrupted; bottom row: denoising for"long"clean and"short"corrupted\.
Figure K\.2:Patching effects on residual stream at END token positions with highlighted decision window\. Top row: denoising for"short"clean and"long"corrupted; bottom row: denoising for"long"clean and"short"corrupted\. Core decision window at L20\-27 highlighted in red\.
Figure K\.3:Core decision window, almost symmetric promotion, markedly asymmetric suppression\.##### Three\-region structure, core decision window at L20\-27\.
Figure[K\.2](https://arxiv.org/html/2606.05194#A11.F2)shows residual stream patching results at final token positions for both directions\. We can observe the same three regions in the dynamics of cumulative denoising curves for both flips\. Layers 0–19 are causally silent on logit difference: the confidence intervals include zero at every layer\. Logarithmic probabilities show very small positive biases in some early layers that do not translate into a detectable LD effect\. Layers 20–27 form a*core decision window*during which the LD recovery rises from below 10% to roughly 80% of the clean baseline \(84\.3% short\-clean, 79\.7% long\-clean at L27\); this seven\-layer staircase accounts for the vast majority of the patching effect\. Layers 28–35 contribute a slow saturation tail, with small but consistent additional contributions around L31–L32\.
Within the core window, the same five layers are the dominant contributors in both flips\. Ranking layers byΔLD\\Delta\\text\{LD\}, the short\-clean flip is led by layers22,24,27,25,2122,24,27,25,21and the long\-clean flip by layers22,21,25,24,2722,21,25,24,27\. The sum of these five per\-layer contributions accounts for 68% of the full recovery in the short\-clean flip and 62% in the long\-clean flip, while cumulative recovery at the end of the seven\-layer window \(L27\) reaches 84% and 80% respectively\. Layer 22 alone is the largest single contributor in both flips, withΔLD=\+0\.221\\Delta\\text\{LD\}=\+0\.221in the short\-clean flip and\+0\.144\+0\.144in the long\-clean flip\.
##### Nearly symmetric promotion, asymmetric suppression\.
The two flips recover the clean answer at almost the same speed: the logarithmic probability curves of clean answers overlap within their bounds across all 36 layers, with direction\-wise differences of at most0\.150\.15anywhere in the domain \(Fig\.[K\.3](https://arxiv.org/html/2606.05194#A11.F3), middle left and bottom right\)\. Suppression of the corrupted answer, however, is markedly faster in the short\-clean flip than in the long\-clean flip\. At layer 24, the normalized suppression is−0\.466\-0\.466in the short\-clean flip versus−0\.231\-0\.231in the long\-clean flip: a two\-fold gap\. The absolute direction difference\|short−long\|\|\\text\{short\}\-\\text\{long\}\|on the logarithmic probability of corrupted answer peaks at0\.230\.23across layers 24–25 and decays monotonically, reaching0\.080\.08by layer 32 and closing to within0\.010\.01at layer 35\. This peak gap is substantially larger than the peak direction differences on LD \(0\.060\.06at L24\) and on logarithmic probability of clean answer \(0\.130\.13at L21\) \(Fig\.[K\.3](https://arxiv.org/html/2606.05194#A11.F3), bottom right\)\. The model produces the two answers with nearly identical efficiency but requires additional late\-layer computation to pushshortdown when the correct answer islong\. This mechanical asymmetry is consistent with the behavioral bias observed during dataset construction, in which all 15 clear\-signal misclassifications were false\-shortpredictions\.
##### Milestone layers\.
Table[K\.1](https://arxiv.org/html/2606.05194#A11.T1)summarizes the first layer at which the mean patching effect crosses a given magnitude threshold, for each metric and each flip\. Onsets \(\|effect\|≥0\.05\|\\text\{effect\}\|\\\!\\geq\\\!0\.05\) occur within a narrow two\-layer window, at L20 on LD in both flips, at L20 \(short\-clean\) and L19 \(long\-clean\) on clean logarithmic probability, and at L20 \(short\-clean\) and L18 \(long\-clean\) on corrupted logarithmic probability\. The long\-clean flip reaches onset 1–2 layers earlier than the short\-clean flip on both log\-probability metrics, but falls progressively behind at higher thresholds\. Above the onset, LD and clean logarithmic probability milestones coincide across flips to within one layer at every threshold, whereas corrupted logarithmic probability milestones diverge: reaching 50%, 75%, and 90% of the full suppression requires layers 25/27/32 in the short\-clean flip but 27/31/35 in the long\-clean flip, a delay of 2–4 layers at each threshold\.
ThresholdLDlog\-P\(clean\)P\(\\text\{clean\}\)log\-P\(corr\)P\(\\text\{corr\}\)\|effect\|\|\\text\{effect\}\|shortlongshortlongshortlong0\.052020201920180\.102121212021230\.252222222123250\.502424222225270\.752727242427310\.90313227263235Table K\.1:Firstresid\_prelayer at which the mean patching effect crosses each magnitude threshold, at the end token position, for the short\-clean and long\-clean flips\.Bold entrieshighlight the 50%, 75%, and 90% log\-P\(corr\)P\(\\text\{corr\}\)milestones discussed in the text, where the short\-clean flip reaches each threshold 2–4 layers before the long\-clean flip\.
##### Summary\.
The causal effect on temporal classification at the final token concentrates inresid\_prelayers 20–27, which together account for 83\.8% \(short\-clean\) and 81\.5% \(long\-clean\) of the full recovery, with layer 22 the single largest contributor in both flips and layers 21, 24, 25, and 27 together accounting for over half of the remaining recovery\. Layers 0–19 have no causal effect at this position, and layers 28–35 contribute a smaller late\-stage refinement \(∼14\{\\sim\}14–19%19\\%of the total recovery\)\. The computation is nearly symmetric in clean\-answer promotion \(peak direction difference0\.130\.13on log\-P\(clean\)P\(\\text\{clean\}\)\) and markedly asymmetric in corrupted\-answer suppression \(peak direction difference0\.230\.23on log\-P\(corr\)P\(\\text\{corr\}\)\): suppressinglongwhen the answer isshortreaches\|effect\|≥0\.90\|\\text\{effect\}\|\\\!\\geq\\\!0\.90by layer 32, whereas suppressingshortwhen the answer islongdoes not reach the same threshold until the final layer \(L35\)\. This residual\-stream asymmetry parallels the short\-biased behavioral errors observed during dataset construction \(all 15 clear\-signal misclassifications were false\-shortpredictions\), though establishing a causal link between the two would require head\-level or lens\-based analysis\.
### K\.3Attention\-output patching at END token
To localize the residual\-stream effect to a specific component, we apply denoising activation patching to the per\-layer attention\-output hook \(attn\_out\) first at all token positions and then at the final one\. Each patch replaces a single \(layer, pos\) summed attention output with its clean\-run counterpart, isolating the contribution of that layer’s attention at given position independent of MLPs and residual pass\-through\. We present the results for all tokens on Figure[K\.4](https://arxiv.org/html/2606.05194#A11.F4)and for the final token on Figure[K\.5](https://arxiv.org/html/2606.05194#A11.F5)\. Since we are primarily interested in interpreting the behavior of attention output flow in the END token, all the results described will concern only it\.

Figure K\.4:Directional attention output patching averaged over 160 classification pairs\. Top row: denoising for"short"clean and"long"corrupted; bottom row: denoising for"long"clean and"short"corrupted\.
Figure K\.5:Patching effects on attention output at END token positions\. Top row: denoising for"short"clean and"long"corrupted; bottom row: denoising for"long"clean and"short"corrupted\.##### Attention writes the decision at a sparse set of layers with L24 being the most dominant\.
Whereas the cumulativeresid\_precurve rises as a smooth seven\-layer staircase across L20–27 \(Sec\.[K\.2\.1](https://arxiv.org/html/2606.05194#A11.SS2.SSS1)\), per\-layer attention effects are sharply localized to four dominant writer layers: L21, L24, L26, and L30\. Each of these produces a significant positive LD effect in both flips \(confidence intervals bounded away from zero\), and the three strongest writers alone \(L24, L26, L30\) each contribute roughly a third of the full normalized recovery individually\. A pair of smaller late writers at L33–34 rounds out the positive contribution, while the first fifteen layers produce no significant attention effect on any metric\.
##### Attention promotes the correct answer but barely suppresses the incorrect one\.
At every dominant attention writer, the effect on log\-P\(clean\)P\(\\text\{clean\}\)is several times larger than the effect on log\-P\(corr\)P\(\\text\{corr\}\): the per\-layer promotion\-to\-suppression ratio exceeds four everywhere among L21, L24, L26, L30 and reaches six or more at the deeper writers in the short\-clean flip\. This is a component\-level observation that the residual\-stream totals alone cannot reveal, because atresid\_preboth metrics eventually approach unit magnitude\. The implication is that attention at the decision layers implements primarily*answer\-promotion*: it writes “the correct answer is here” into the residual stream, with only a modest side effect on the competing answer\. The deep suppression observed atresid\_pre\(where log\-P\(corr\)P\(\\text\{corr\}\)saturates near−1\-1by the final layers\) must therefore come largely from a different component, most plausibly the MLPs, although this wasn’t confirmed by our MLP patching analysis \(Sec\.[K\.4](https://arxiv.org/html/2606.05194#A11.SS4)\)\.
##### The direction asymmetry largely disappears at the attention level\.
A central finding of theresid\_preanalysis was that corrupted\-answer suppression is faster when the correct answer isshortthan when it islong, with a peak direction gap of0\.230\.23on log\-P\(corr\)P\(\\text\{corr\}\)\. Atattn\_out, this pattern is absent: attention effects on the corrupted answer are nearly equal across flips at every dominant writer\. The largest remaining direction difference shifts to*promotion*rather than suppression: the short\-clean flip writes a noticeably stronger clean signal at L26 than the long\-clean flip does, but even this residual asymmetry is well under half the size of the one observed atresid\_pre\. Taken together, these two facts indicate that the direction\-asymmetric late\-layer suppression seen in the residual stream does not originate in the attention blocks\.
##### Summary\.
Theresid\_predecision window resolves, at the attention\-output level, into four primary writer layers \(L21, L24, L26, L30\)\. These attention blocks are strongly biased toward promotion of the correct answer rather than suppression of the incorrect one, and they behave nearly symmetrically across the two answer directions\. The direction\-asymmetric late\-layer suppression observed atresid\_preis therefore attributable to a different component, which matching MLP\-output patching should be able to identify\.
### K\.4MLP\-output patching at END token
The attention\-output analysis \(Sec\.[K\.3](https://arxiv.org/html/2606.05194#A11.SS3)\) suggested that attention blocks promote the correct answer with only a small suppression side\-effect, and we conjectured that the deeper, direction\-asymmetric suppression seen atresid\_prewould be carried by the MLPs\. To test this, we patch the per\-layermlp\_outhook at the final token position, using the same two flips and the same three metrics\. We also provide per\-layer effects for all prompt tokens in Fig[K\.6](https://arxiv.org/html/2606.05194#A11.F6)for a broader view\. Fig\.[K\.7](https://arxiv.org/html/2606.05194#A11.F7)shows patching effects at the final token and Figure[K\.8](https://arxiv.org/html/2606.05194#A11.F8)aggregates structural findings by multiple plots\.

Figure K\.6:Directional MLP output patching averaged over 160 classification pairs\. Top row: denoising for"short"clean and"long"corrupted; bottom row: denoising for"long"clean and"short"corrupted\.
Figure K\.7:Patching effects on MLP output at END token positions\. Top row: denoising for"short"clean and"long"corrupted; bottom row: denoising for"long"clean and"short"corrupted\.
Figure K\.8:MLP output patching results from different perspectives\.LP\_cdenotes log\-P\(clean\)P\(\\text\{clean\}\)andLP\_kdenotes log\-P\(corr\)P\(\\text\{corr\}\)\.##### MLPs are weaker writers than attention, and they also promote\.
MLP effects are substantially smaller in magnitude than attention effects: peak LD is\+0\.131\+0\.131\(short\-clean, L27\) versus\+0\.308\+0\.308for attention \(short\-clean, L24\), about a factor of2\.42\.4smaller\. A consistent positive\-LD signature appears across three layers in the middle\-late range \(L27, L28, and L31\), which form the primary MLP writer band and are significant in both flips \(L27:\+0\.131/\+0\.113\+0\.131/\+0\.113; L28:\+0\.123/\+0\.083\+0\.123/\+0\.083; L31:\+0\.093/\+0\.089\+0\.093/\+0\.089\)\.Contrary to our initial conjecture, these MLP writers do not primarily suppress the incorrect answer:at every layer in the primary band, the effect on log\-P\(clean\)P\(\\text\{clean\}\)is larger in magnitude than the effect on log\-P\(corr\)P\(\\text\{corr\}\), the same promotion\-dominant pattern we observed at attention\. MLPs therefore reinforce the decision written by attention rather than performing a qualitatively distinct suppression step\.
##### The late\-layer suppression hypothesis is not supported\.
The residual\-stream analysis showed that log\-P\(corr\)P\(\\text\{corr\}\)saturates near−1\-1across layers 28–35, with a pronounced direction\-dependent delay in the long\-clean flip\. If a specific MLP layer were responsible for this late\-layer suppression, it should appear here as a large negative effect on log\-P\(corr\)P\(\\text\{corr\}\)at one or more of layers 28–35\. The observed MLP log\-P\(corr\)P\(\\text\{corr\}\)effects in that range are, however, small and mixed in sign: the largest is−0\.121\-0\.121at L27 in the short\-clean flip \(significant\), but in the long\-clean flip the corresponding MLPs at L27 and L28 show*positive*log\-P\(corr\)P\(\\text\{corr\}\)effects \(\+0\.078\+0\.078and\+0\.045\+0\.045, both significant\), meaning they slightly push the model toward the incorrect answer\. No single MLP layer produces a suppression effect comparable in magnitude to the saturation observed atresid\_pre\. The component\-level decomposition we conjectured at the end of Sec\.[K\.3](https://arxiv.org/html/2606.05194#A11.SS3)\(attention promoting, MLPs suppressing\) is therefore not supported by the data\. The deep suppression atresid\_preappears instead to be a cumulative property of many small contributions distributed across both components, rather than the work of a localized suppressor\.
##### L33 and L35: late MLPs with prominent direction\-dependent effects onshort\.
Two late MLP layers produce robust effects whose largest significant components all involve the model’s prediction ofshort\. At L33 in the short\-clean flip, patching the clean\-run MLP output produces a log\-P\(clean\)P\(\\text\{clean\}\)effect of\+0\.311\+0\.311\. At L35 in the short\-clean flip, the same operation produces a log\-P\(clean\)P\(\\text\{clean\}\)effect of−0\.264\-0\.264; while in the long\-clean flip, it produces a log\-P\(corrupted\)P\(\\text\{corrupted\}\)effect of\+0\.167\+0\.167, the single largest log\-P\(corrupted\)P\(\\text\{corrupted\}\)effect observed at any MLP layer in either direction\. As a possible interpretation we can say that the fact that both layers’ largest effects fall onshort\-related metrics \(rather than being distributed across the metrics tracked at other writers\) is consistent with their carrying computations specifically tied to theshortrepresentation\. The direction of the patching effect differs between flips: at L35 in particular, patching pushes the model away fromshortwhenshortis correct and towardshortwhenlongis correct \(Figure[K\.8](https://arxiv.org/html/2606.05194#A11.F8), bottom row\)\. A mechanistic account of these polarity\-dependent pattern would require head\-level or neuron\-level analysis\. Alternative explanations \(residual interaction with the unembedding, normalization artifacts at the final layer\) cannot be ruled out without further experiments\.
##### Summary\.
MLP patching refutes our earlier conjecture that the direction\-asymmetric late\-layer suppression atresid\_preis localized to MLPs in layers 28–35\. MLPs are weaker writers than attention, and at the layers where they contribute significantly \(L27, L28, L31\) they continue the promotion\-dominant pattern established by attention\. The aggregate late\-layer suppression observed atresid\_preis therefore best understood as a cumulative property of many small contributions across both components, not the work of a localized suppressor\. Several MLP writers show direction\-dependent behavior; among these, L33 and L35 stand out for the size of their robust effects, all of which involve the model’s prediction ofshort\. The L33/L35 finding does not localize the residual\-stream asymmetry’s origin: the asymmetry on log\-P\(corrupted\)P\(\\text\{corrupted\}\)is already near its peak by L24\. What L33 and L35 add is a set of large, direction\-dependent effects whose largest components fall on metrics involving theshorttoken, with L35’s long\-clean log\-P\(corrupted\)P\(\\text\{corrupted\}\)effect \(\+0\.167\+0\.167\) being the single largest log\-P\(corrupted\)P\(\\text\{corrupted\}\)effect observed at any MLP layer in either direction\. Head\-level attention patching at L21, L24, L26, L30 and neuron\-level analysis of MLPs at L33 and L35 would be the natural next experiments to test and refine this picture\.
### K\.5Position\-aggregated view
The per\-layer analyses above \(Sec\.[K\.3](https://arxiv.org/html/2606.05194#A11.SS3), Sec\.[K\.4](https://arxiv.org/html/2606.05194#A11.SS4)\) focus on the END token position, where the patching effect concentrates \(Finding 1 in Sec\.[K\.6](https://arxiv.org/html/2606.05194#A11.SS6)\)\. For cross\-method comparison with the position\-aggregated metrics used in EAP\-IG attribution \([Appendix H](https://arxiv.org/html/2606.05194#A8)\) and the layer\-level summaries of causal parametric patching \([Appendix J](https://arxiv.org/html/2606.05194#A10)\), Figure[K\.9](https://arxiv.org/html/2606.05194#A11.F9)shows attention output \(top\) and MLP output \(bottom\) patching effects summed across all 34 token positions, separately for theshort\-cleanflip \(left\) andlong\-cleanflip \(right\)\.




Figure K\.9:Layer\-level patching effects summed across all 34 token positions, averaged over 160 classification pairs\. Top row: attention output\. Bottom row: MLP output\. Left column: short\-clean flip; right column: long\-clean flip\.Compared to the END\-only view, the summed\-attention plot shows L24, L26, and L30 as robust writers across both flips, with L18 and L21 visible in the short\-clean flip but weaker in long\-clean; the small late writer at L33 identified at END in the long\-clean flip does not stand out in the summed view\. Confidence intervals are substantially wider throughout\. The summed MLP plot is essentially flat across all 36 layers in both flips: per\-layer MLP signal is near\-absent when aggregated over all token positions\. This is consistent both with the END\-locality finding and with[Appendix N](https://arxiv.org/html/2606.05194#A14)’s interpretation that strong MLP signal in causal parametric patching reflects constraint\-token processing which is absent in given prompting paradigm\. The visible L0 peaks in both attention and MLP panels reflect embedding\-level differences between clean and corrupted prompts \(different goal text\) rather than causal temporal computation at the first layer; they do not appear in the END\-only views above because the END token’s embedding is identical across clean and corrupted prompts\. These views are reported here primarily for the cross\-method comparison in[Appendix L](https://arxiv.org/html/2606.05194#A12); the substantive per\-layer findings of this appendix are based on the END\-token analyses in Sec\.[K\.3](https://arxiv.org/html/2606.05194#A11.SS3)and Sec\.[K\.4](https://arxiv.org/html/2606.05194#A11.SS4)\.
### K\.6Key Findings
1. 1\.The decision is sparsely localized at the final token\.At the final token, the causal effect on temporal classification concentrates in a small set of layers in a narrow band of processing depth\. Theresid\_preLD curve is silent for layers 0–19 \(confidence intervals contain zero at every layer\), rises as a seven\-layer staircase across L20–27 that accounts for roughly 80% of the LD recovery, and continues through a slower saturation tail at L28–35\.
2. 2\.Attention and MLP components are both promotion\-dominant\.Component\-level patching refutes the simplest decomposition one might expect \(*attention writes the answer, MLPs suppress the alternative*\)\. At the dominant attention writers the ratio\|log\-P\(clean\)\|/\|log\-P\(corrupted\)\|\|\\text\{log\-\}P\(\\text\{clean\}\)\|/\|\\text\{log\-\}P\(\\text\{corrupted\}\)\|is four or more in seven of the eight layer–flip combinations we test \(L21, L24, L26, L30 across both flips\)\. MLP writers are also promotion\-dominant at most layers but with greater variability: ratios at the primary\-window writers span from1\.211\.21\(L27 short\-clean; roughly balanced promotion and suppression\) up to13\.3113\.31\(L31 long\-clean\); the direction\-dependent layers L33 and L35 \(Finding 5\) deviate further and are described separately\. The late\-layer suppression of the incorrect answer seen atresid\_preis therefore not localized to a single component; it accumulates from many small contributions distributed across attention and MLPs\.
3. 3\.Attention is the primary writer, MLPs reinforce\.Within the L20–27 window, attention contributes the majority of the magnitude\. The principal attention writers in both flips are L21, L24, L26, and L30, with a single\-layer peak at L24\. MLPs add smaller but reproducible writes at L27, L28, and L31; peak MLP LD effect is\+0\.131\+0\.131at L27 short\-clean, roughly2\.4×2\.4\\timessmaller than the attention peak \(which sits at a different layer, L24\)\. The two component families work in concert rather than in specialized roles\.
4. 4\.Promotion is approximately symmetric across flips; suppression is not\.Atresid\_pre, the onset of the decision on LD coincides in both flips at L20\. Clean\-answer promotion proceeds at similar rates in the two flips: log\-P\(clean\)P\(\\text\{clean\}\)milestones at1010,2525,5050,7575, and90%90\\%recovery coincide to within one layer at every threshold\. Corrupted\-answer suppression, by contrast, is consistently faster when the correct answer isshort: reaching50%50\\%,75%75\\%, and90%90\\%of the full suppression requires layers25/27/3225/27/32in the short\-clean flip but27/31/3527/31/35in the long\-clean flip, a delay of 2–4 layers at every threshold\. The suppression asymmetry is thus∼1\.8×\{\\sim\}1\.8\\timesthe promotion asymmetry on the relevant metrics and closes only at the final layer \(L35\)\.
5. 5\.Two MLP layers show large direction\-dependent effects on theshorttoken\.Layers 33 and 35 stand out for the size of their robust effects, all of which involve metrics related to the model’s prediction ofshort\. At L33 in the short\-clean flip, patching the clean\-run MLP output yields a log\-P\(clean\)P\(\\text\{clean\}\)effect of\+0\.311\+0\.311\. At L35 in the short\-clean flip, the same operation yields a log\-P\(clean\)P\(\\text\{clean\}\)effect of−0\.264\-0\.264\. At L35 in the long\-clean flip, patching produces a log\-P\(corrupted\)P\(\\text\{corrupted\}\)effect of\+0\.167\+0\.167, the single largest log\-P\(corrupted\)P\(\\text\{corrupted\}\)effect observed at any MLP layer in either direction\. A unified mechanistic account of these patterns would require head\-level or neuron\-level analysis\.
6. 6\.The circuit\-level asymmetry parallels a behavioral short\-bias\.The 15 clear\-signal misclassifications in the dataset were all false\-shortpredictions\. The residual\-stream finding that suppressingshort\(long\-clean flip\) requires more layers than suppressinglongforms a mechanical pattern in the same direction as this behavioral bias\. A causal link between the patching\-level asymmetry and the classification\-level errors cannot be established from this experiment alone, but the direction of both effects agrees\.
7. 7\.Convergence with parametric and contrastive attribution results\.For cross\-paradigm comparison we use the position\-summed view of the classification results \(Sec\.[K\.5](https://arxiv.org/html/2606.05194#A11.SS5)\), where L24, L26, and L30 are the attention writers robust across both flips\. L24 is the dominant attention writer across all three non\-probing methods \(classification, contrastive attribution[Appendix H](https://arxiv.org/html/2606.05194#A8), and parametric[Appendix J](https://arxiv.org/html/2606.05194#A10)\); L26 and L30 are classification\-specific\. Classification’s core decision window \(L20–27\) on the residual stream and parametric’s critical window \(L20–24\) align on layers L20–24\. MLP contributions diverge: classification’s summed MLP signal is near\-absent, in contrast to parametric’s localized L31/L35 prominence and contrastive attribution’s diffuse signal across L31–L35\. The parametric divergence is partially consistent with[Appendix N](https://arxiv.org/html/2606.05194#A14)’s finding that MLP signal weakens substantially when prompts lack explicit horizon constraints, though[Appendix N](https://arxiv.org/html/2606.05194#A14)’s latent mode retains some MLP involvement that classification does not show\. The fuller MLP divergence, including from contrastive attribution despite its similarly unconstrained prompts, may also reflect a more fundamental paradigm difference: classification against preference valuation\. The partial overlap on attention and residual stream is consistent with shared computational machinery in this region, though whether this reflects a common temporal representation or merely a common decision readout requires further analysis we leave to future work\.
## Appendix Appendix LCross\-method convergence
The preceding five appendices each approached localization from a different angle: two attribution methods \(one contrastive, one parametric\), two causal patching experiments, and supervised probing, applied across three prompt paradigms\. Each method has different assumptions, blind spots, and failure modes\. The question is whether they agree\.
Table[L\.1](https://arxiv.org/html/2606.05194#A12.T1)and Figure[L\.1](https://arxiv.org/html/2606.05194#A12.F1)show that they do: four methods place the temporal preference subgraph in layers 17–35, with the three non\-probing methods on parametric prompts \(Attr\. contr\., Attr\. param\., Causal param\.\) highlighting L24 attention at the center\. The fifth \(classification\) method also supports L24 as the central attention component, with its residual decision window L20–27 falling within the L17–35 range\. The agreement is not trivial, as the methods also disagree in informative ways \(Section[L\.2](https://arxiv.org/html/2606.05194#A12.SS2)\)\.
ProbingAttr\. contr\.Attr\. param\.Causal param\.Causal class\.[Appendix G](https://arxiv.org/html/2606.05194#A7)[Appendix H](https://arxiv.org/html/2606.05194#A8)[Appendix I](https://arxiv.org/html/2606.05194#A9)[Appendix J](https://arxiv.org/html/2606.05194#A10)[Appendix K](https://arxiv.org/html/2606.05194#A11)Attn L21–L24✓✓✓∼\\sim✓MLP L31–L35✓✓✓Resid L17–L22 \(recovery\)✓✓∼\\sim✓Resid L26✓✓Attn peakL24 \(ST\), L22 \(LT\)L20–L24L24, L21L24, L26, L30MLP peakL34, L35L31L35, L31Best single layerL26L24L17–L22 \(resid\)L24L24Signal onset∼\\simL17∼\\simL21∼\\simL17∼\\simL19∼\\simL20Convergence zone: layers 17–35
Table L\.1:Layers and components identified by each localization method\. All five methods place a common subgraph in layers 17–35\. L24 attention is flagged by every non\-probing method\. The classification \(causal contr\.\) method partially supports the L21–L24 attention substrate \(L24 robustly across both flips; L21 short\-clean only\) and its core residual window overlaps with parametric’s residual recovery range at L20–L22\. MLP contributions concentrate in L31–L35 under attributional contrastive, attributional parametric, and causal parametric patching\. Symbol “∼\\sim✓” marks partial convergence\.L21L24L31151617181920212223242526272829303132333435LayerCausal \(class\.\)Causal \(param\.\)Attr\. \(param\.\)Attr\. \(contr\.\)ProbingLocalization \(Layers L15–L35\)AttnMLPResidProbeFigure L\.1:Layer\-level convergence across all five localization methods\. Darker shading indicates stronger signal\. Blue==attention, red==MLP, green==residual stream \(probing\), gray==residual stream \(causal classification, core decision window\)\. L24 attention appears in every non\-probing method\. The causal classification experiment \(bottom row\) additionally identifies attention contributions at L26 and L30, together with a residual\-stream core decision window at L20–L27 \(peak L22\)\.### L\.1Points of Agreement
L24 attention is the single component flagged most consistently: it appears as a top\-ranked element in all four non\-probing methods across three paradigms, making it the most robustly identified element of the subgraph\. The MLP contribution concentrates in L31–L35 under attributional contrastive and causal parametric patching, with MLP L31 identified as a key disruptor by both methods\. All methods agree that the temporal signal is absent from the first∼\\sim15 layers and concentrated in the upper half of the network\. The probing peak at L26 falls between the attention computation window \(L21–L24\) and the MLP computation window \(L31–L35\), consistent with a readout layer that consolidates the output of the attention\-mediated temporal routing before the MLP transformation stage\. The causal classification residual\-stream staircase \(L20–L27, peak L22\) spans this same region, converging with the attention window \(L21–L24\) on the low end and with the probing peak \(L26\) on the high end\.
### L\.2Points of Disagreement
The methods disagree on three dimensions\.
##### Attention breadth\.
The causal classification experiment identifies a broader span of attention layers \(L24, L26, L30 robust across both flips\) than the parametric and attribution methods, which concentrate on L21–L24\. The data do not establish the cause of this difference: the three setups differ along multiple dimensions \(cognitive task, prompt structure, method, prompt length, position handling\), and disambiguating which dimension drives the breadth difference would require controlled experiments we do not undertake here\.
##### MLP visibility\.
The causal classification experiment shows near\-absent summed MLP signal, while attribution and parametric both identify L31–L35 as important\. The parametric divergence is partially consistent with[Appendix N](https://arxiv.org/html/2606.05194#A14)’s finding that MLP signal weakens substantially when prompts lack explicit horizon constraints, though[Appendix N](https://arxiv.org/html/2606.05194#A14)’s latent mode retains some MLP involvement that classification does not show\. The fuller divergence, including from contrastive attribution despite its similarly unconstrained prompts, is subject to the same multi\-dimensional confound noted above\.
##### Sufficiency vs\. necessity asymmetry\.
The causal parametric results reveal a sharp split between mid\-layer recovery \(L17–L22\) and late\-layer disruption \(L30–L35\), a pattern that the contrastive methods do not clearly replicate\. The causal classification experiment instead finds a different kind of asymmetry: the two flip directions share the same core decision window \(L20–L27\) and promote the clean answer at comparable rates, but suppress the corrupted answer at different speeds \(suppressingshortwhen the answer islongtakes 2–4 more layers than the reverse\)\. This direction\-dependent suppression is a dimension the parametric methods do not probe, because they do not separate patching directions\.
##### Signal onset\.
The earliest onset varies from∼\\simL17 \(probing\) to∼\\simL21 \(attributional contrastive\)\. This 4\-layer gap may reflect the greater sensitivity of probing and parametric prompts to early, low\-magnitude temporal information that the contrastive attribution method, aggregated over many prompt variants, averages out\.
### L\.3Interpretation
The disagreements are interpretable rather than contradictory: they reflect genuine differences in both what each method measures and what each paradigm probes\. Attribution methods approximate causal effects via gradients and are sensitive to all information flow, including redundant pathways\. Causal methods measure the actual behavioral consequence of intervention and are therefore sensitive to necessity and sufficiency\. Probing measures the linear readability of a concept at a given layer, regardless of whether that layer is causally important\. At the task level, three paradigms \(probing, attribution, and parametric patching\) target temporal preference valuation, while the fourth \(classification patching\) probes categorical horizon inference, a distinct cognitive operation that may recruit different downstream computation even when sharing the upstream attention substrate\.
That these four methods, despite their different assumptions and blind spots, converge on a common subgraph in layers 17–35 with L24 attention at the center supports the claim that the localization is not an artifact of any single method\. The probing–steering dissociation \([Appendix S](https://arxiv.org/html/2606.05194#A19)\) adds a fifth data point: layers 19–22 are most effective for writing temporal preference, while layer 26 is most effective for reading it, reinforcing the functional distinction between the attention routing window and the readout layer\.
However, the subgraph is not monolithic\. The latent vs\. constrained analysis \([Appendix N](https://arxiv.org/html/2606.05194#A14)\) reveals that the same L17–35 region operates in two modes depending on whether the prompt carries an explicit time\-horizon constraint:
- •Constrained mode: the full subgraph is engaged \(attention L21–24*and*MLP L31–35\), producing strong, distributed effects\.
- •Latent mode: only attention at L21–22 is active, with minimal MLP involvement\.
The MLP layers that feature prominently in the convergence table \(L31, L35\) may therefore be specifically about processing constraint tokens rather than encoding temporal preference per se\. The attention core at L21–24 is the shared substrate; MLP extends the computation when the prompt provides an explicit temporal anchor\.
Part 2: What does temporal preference look like?
- •[M](https://arxiv.org/html/2606.05194#A13)\.Parametric geometry
- •[N](https://arxiv.org/html/2606.05194#A14)\.Latent vs\. constrained
- •[O](https://arxiv.org/html/2606.05194#A15)\.Behavioral temporal discounting
- •[P](https://arxiv.org/html/2606.05194#A16)\.Behavioral coherence
- •[Q](https://arxiv.org/html/2606.05194#A17)\.Cross\-model patching comparison
- •[R](https://arxiv.org/html/2606.05194#A18)\.Error monitoring in the subgraph
## Appendix Appendix MParametric geometry
Part 1 established*where*temporal preference lives \(layers 17–35, with L24 attention at the center;[Appendix L](https://arxiv.org/html/2606.05194#A12)\)\. Now we ask:*what does the representation look like inside that subgraph?*
We apply PCA to 4,588 activation vectors, sampled from a logarithmic grid over reward amounts, delay times, and 17 time horizons \(seconds to centuries\), at 15 layers, 5 component types, and 16 semantic positions per prompt \(methodology in[Appendix Y](https://arxiv.org/html/2606.05194#A25)\)\. At key positions, PC1 captures 44–71% of variance \(Table[Y\.1](https://arxiv.org/html/2606.05194#A25.T1)\)\. The results tell a mechanistic story in five stages: the model builds an ordinal time\-horizon representation, the geometric direction encoding it flips across prompt positions, it stabilizes at the user\-to\-assistant turn boundary, attention transforms it into a binary preference signal over the next few tokens, and the preference commits by theassistanttoken\.
### M\.1Progressive separation across layers
Figure[M\.1](https://arxiv.org/html/2606.05194#A13.F1)shows how the PC1 projection evolves across layers for four component types, colored by the model’s eventual choice \(long vs\. short\)\.
resid\_pre 
resid\_post 
attn\_out 
mlp\_out 
Figure M\.1:PC1 projection across layers for each component type at the divergent token position, colored by the model’s chosen term \(orange = long, blue = short\)\. Short\-term and long\-term traces begin to diverge around layer 21 in the residual stream and attention, fully separating by layer 24 \(±\\pm50 on PC1 for the residual stream,±\\pm20 for attention\)\. MLP contributions emerge later and remain smaller\. At the user\-to\-assistant turn boundary and later token positions, separation appears earlier \([M\.3](https://arxiv.org/html/2606.05194#A13.SS3)\)\.At this token position, all traces begin bundled near zero through the first∼\\sim20 layers\. Separation becomes visible around layer 21 in the residual stream and attention output and is fully established by layer 24, consistent with the causal importance of L24 identified by activation patching \([Appendix J](https://arxiv.org/html/2606.05194#A10)\)\. By the final layers, the residual stream carries a separation of roughly±\\pm50 on PC1, while attention contributes±\\pm20 and MLP a smaller but complementary signal concentrated in the upper layers\.
### M\.2Off\-policy horizon constraint \(2D PCA\)

Figure M\.2:PCA of activation space at three token positions \(chosen term, chosen time, time scale\) with the time horizon given as an explicit constraint\.
Figure M\.3:PCA of activation space colored by time scale at layers 3, 18, and 24 after the time horizon token\. Clusters become increasingly separable in the mid\-to\-upper layers\.
### M\.3On\-policy temporal preference \(2D PCA\)
The preceding subsection examined how explicit time\-horizon constraints are represented in activation space\. Here we trace what happens as the model transitions from the user’s turn \(where the constraint is given off\-policy\) into the assistant’s turn, where it must generate on\-policy text reflecting a temporal preference\. Figure[M\.4](https://arxiv.org/html/2606.05194#A13.F4)illustrates the token\-level structure of this transition\. As the figures below show, the explicit time\-horizon clusters reorganize during this hand\-off: the no\-horizon samples, initially disjoint, align to the time\-scale manifold and temporal preference becomes linearly separable even at the earliest layers\.
<\|im\_end\|\>Control token\\nDelimiter<\|im\_start\|\>Control tokenassistantRole tag\\nDelimiterEnd\-of\-Turn \(EoT\)Start\-of\-Turn \(SoT\)
Figure M\.4:The transition from user EoT to assistant SoT marks the shift from off\-policy context to on\-policy generation\.<\|im\_end\|\>\\n<\|im\_start\|\>assistant
Figure M\.5:Principal components of 4,588 samples for layer 31 at token positions through the end\-of\-turn \(EoT\) for the user into the beginning\-of\-turn \(BoT\) for the assistant\. At first \(<\|im\_end\|\>\), temporal preference is not linearly separable and the no\-horizon samples \(gray\) are disjoint from the off\-policy time\-horizon manifold\. As the LLM transitions into the assistant’s turn \(moving towardsassistant\), they appear to align to the manifold before preference clusters are formed\.\\n
Figure M\.6:By the token position of the BoT delimiter \(\\nafterassistant\), temporal preference is separable even at layer 0\.
### M\.4Horizon representation is present but geometrically unstable in the prompt
Even before the model starts generating, the residual stream encodes time horizon\. Figure[M\.7](https://arxiv.org/html/2606.05194#A13.F7)shows the PC1 projection at several positions within the prompt, colored by horizon category\.
resid\_post@post\_time\_horizon 
resid\_post@action\_tail 
resid\_post@format\_tail 
resid\_post@chat\_suffix\_tail 
Figure M\.7:PC1 projection ofresid\_postacross layers at four prompt positions, colored by time horizon \(blue = seconds, yellow = deep time\)\. Top\-left: after the time horizon constraint token \(within theCONSTRAINTsection\)\. Top\-right: last token of theACTIONsection\. Bottom\-left: last token of theFORMATsection\. Bottom\-right:chat\_suffix\_tail\(the\\nafterassistant\)\. The ordinal fan is present at all positions, but its*polarity flips*between positions \(short horizons go negative at some, positive at others\), indicating the geometric direction encoding horizon is not yet stable within the prompt\.The horizon signal is large \(spreads of±\\pm100 or more on PC1\) and ordinally organized at every position, but the direction encoding it rotates across positions\. At the ACTION tail, short horizons go strongly negative; at the FORMAT tail, the polarity flips and short horizons go positive\. This instability persists into the earliest response tokens \(Figure[M\.8](https://arxiv.org/html/2606.05194#A13.F8)\)\.
resid\_post@response\_choice\(a/b\) 
resid\_post@response\_choice\_prefix\(choose\) 
Figure M\.8:PC1 projection ofresid\_postat response positions, colored by time horizon\. Left:response\_choice\(thea\)orb\)token\)\. Right:response\_choice\_prefix\(thechoosetoken in “I choose:”\)\. The horizon signal is present but the geometric direction has not yet fully stabilized\.The model has the horizon information throughout, but has not committed to a stable geometric encoding of it until the turn boundary\.
### M\.5Stabilization at the turn boundary \(residual stream\)
The representation stabilizes at the user\-to\-assistant turn boundary\. Figure[M\.9](https://arxiv.org/html/2606.05194#A13.F9)showsresid\_postat three key positions in the turn transition\.
resid\_post, suffix 0, horizon 
resid\_post, suffix 0, preference 
resid\_post, suffix 3, preference 
Figure M\.9:resid\_postPC1 projection at the turn boundary\. Left: suffix 0 \(<\|im\_end\|\>\), colored by time horizon\. The ordinal fan is now stable and monotonic, with short horizons trending negative and long horizons positive\. Center: same position, colored by preference\. Long and short are heavily overlapping: the choice has not yet been made\. Right: suffix 3 \(assistanttoken\), colored by preference\. Long and short are cleanly separated from early layers onward\. Between these two positions, the model converts the stable horizon representation into a committed preference\.At suffix 0, the residual stream carries a clean, ordinal horizon representation \(left\), but long and short preferences overlap completely \(center\)\. The geometry at this position encodes*how far into the future*, not*which option to choose*\. By suffix 3, the preference is fully committed \(right\): long and short form two non\-overlapping bands from early layers onward\.
The complete four\-position transition in the residual stream \(suffix 0 through 3\) is shown in Figure[M\.10](https://arxiv.org/html/2606.05194#A13.F10)\.

<\|im\_end\|\>

\\n

<\|im\_start\|\>

assistant
Figure M\.10:resid\_postat suffix positions 0 through 3, all colored by preference \(orange = long, blue = short\)\. The preference signal progressively sharpens from heavy overlap at suffix 0 to clean separation at suffix 3\.
### M\.6Attention mediates the horizon\-to\-preference transformation
To isolate the mechanism driving the conversion, Figure[M\.11](https://arxiv.org/html/2606.05194#A13.F11)showsattn\_out\(the attention output only, before it is added to the residual stream\) at all four suffix positions\.
attn\_out, suffix 0 \(<\|im\_end\|\>\), horizon 
attn\_out, suffix 0 \(<\|im\_end\|\>\), preference 
attn\_out, suffix 1 \(\\n\), horizon 
attn\_out, suffix 1 \(\\n\), preference 
attn\_out, suffix 2 \(<\|im\_start\|\>\), horizon 
attn\_out, suffix 2 \(<\|im\_start\|\>\), preference 
attn\_out, suffix 3 \(assistant\), horizon 
attn\_out, suffix 3 \(assistant\), preference 
Figure M\.11:attn\_outPC1 projection across layers at suffix positions 0–3 \(rows\), colored by time horizon \(left column\) and preference \(right column\)\. At suffix 0, attention carries horizon structure but no preference separation\. At suffix 1, a noisy preference signal emerges with a characteristic V\-shape around layers 13–24, indicating active reorganization\. At suffix 2, the preference separation strengthens\. By suffix 3, preference is clearly separated in the attention output\. The non\-monotonic trajectories at suffix 1–2 \(unlike the smooth residual\-stream fans\) reveal that attention is actively*transforming*the representation, not merely amplifying a pre\-existing signal\.The attention output at suffix 0 \(top row\) carries ordinal horizon structure \(left\) but no preference signal \(right: long and short are intermingled\)\. At suffix 1, the attention output shows a distinctive non\-monotonic, V\-shaped trajectory in the mid\-layers \(13–24\)\. This zigzag pattern, absent in the smooth residual\-stream fans, reveals that attention heads are actively reorganizing the representation\. A noisy preference signal begins to emerge \(right\)\. At suffix 2, the preference separation strengthens, and by suffix 3, long and short are cleanly separated in the attention output\.
This progression identifies attention as the operation that converts the stable horizon representation \(written into the residual stream by suffix 0\) into a preference signal, incrementally across suffix positions 1–3\.
### M\.73D trajectories: horizon becomes preference
Figure[M\.12](https://arxiv.org/html/2606.05194#A13.F12)shows the same transition in 3D PCA space \(PC1×\\timesPC2×\\timesLayer\), making the geometric reorganization visually explicit\.
cross\-layer, suffix 0, horizon 
cross\-layer, suffix 0, preference 
cross\-layer, suffix 1, horizon 
cross\-layer, suffix 1, preference 
cross\-layer, suffix 2, horizon 
cross\-layer, suffix 2, preference 
cross\-layer, suffix 3, horizon 
cross\-layer, suffix 3, preference 
Figure M\.12:3D PCA trajectories \(PC1×\\timesPC2×\\timesLayer\) at suffix positions 0–3 \(rows\), colored by time horizon \(left\) and preference \(right\)\. At suffix 0, traces fan out by horizon but preferences are intermingled\. At suffix 1–2, the geometry begins reorganizing: traces split into two branches visible in 3D\. By suffix 3, the two branches cleanly correspond to long vs\. short preference, with horizon ordering preserved as a secondary structure within each branch\.
### M\.8Position sweep at L24
Figure[M\.13](https://arxiv.org/html/2606.05194#A13.F13)shows the geometry at a fixed layer \(L24\) swept across all token positions, confirming that the transition from unstable horizon to committed preference happens at the turn boundary\.
L24resid\_post, 1D PC1×\\timesposition, horizon 
L24resid\_post, 1D PC1×\\timesposition, preference 
L24, 2D PCA, horizon 
L24, 2D PCA, preference 
L24, 2D PCA, chosen time 
Figure M\.13:L24 activations across token positions\. Top: 1D PC1 projection vs\. position, colored by time horizon \(left\) and preference \(right\)\. Early prompt positions show wild oscillations; the representation stabilizes at the turn boundary with ordinal horizon separation \(left\) and clean preference separation emerging a few tokens later \(right\)\. Bottom: 2D PCA \(PC1 vs\. PC2\) with position\-connected traces, colored by time horizon \(left\), preference \(center\), and chosen time \(right\)\. At late positions \(dense cluster\), both horizon and preference structure are visible\.The 1D position sweep \(top row\) confirms the narrative at a single layer: early prompt positions show oscillating, unstable encodings, while the turn boundary and subsequent tokens show stable ordinal horizon separation \(left\) and progressive preference commitment \(right\)\.
### M\.9Direction alignment across components
Figure[M\.14](https://arxiv.org/html/2606.05194#A13.F14)shows the cosine similarity between the top PCA direction at each component\-layer pair \(computed at the turn boundary\), confirming that, at this token position, the temporal direction stabilizes in mid\-to\-late layers\.
Figure M\.14:Direction alignment matrix across component\-layer pairs\. The temporal direction is consistent within nearby layers \(warm block\-diagonal patches\) but rotates substantially between early layers \(L0–L12\) and later layers \(L18\+\)\. Within mid\-to\-late layers, residual, attention, and MLP components at the same layer share similar directions, indicating a stable temporal subspace\.
### M\.10Summary
The geometry analysis reveals a five\-stage process:
1. 1\.The model builds an ordinal time\-horizon representation in the residual stream, already present within the prompt, but the geometric direction encoding it is*unstable*: it flips polarity across prompt positions \(Figure[M\.7](https://arxiv.org/html/2606.05194#A13.F7)\)\.
2. 2\.At the user\-to\-assistant turn boundary \(suffix 0\), the residual stream stabilizes this representation into a clean, monotonic fan by time scale, but the model’s preference \(long vs\. short\) is not yet encoded \(Figure[M\.9](https://arxiv.org/html/2606.05194#A13.F9)\)\.
3. 3\.Attention outputs at suffix positions 1–2 show non\-monotonic, actively reorganizing trajectories that progressively write a preference signal into the residual stream \(Figure[M\.11](https://arxiv.org/html/2606.05194#A13.F11)\)\.
4. 4\.By suffix 3 \(assistanttoken\), the residual stream carries a fully committed preference signal, with long and short cleanly separated from early layers onward \(Figure[M\.10](https://arxiv.org/html/2606.05194#A13.F10)\)\.
5. 5\.The transformation occurs in layers 18–24, the same layers identified as causally important by activation patching \([Appendix J](https://arxiv.org/html/2606.05194#A10)\)\.
This geometric narrative connects localization \(where\) to function \(what\): the subgraph in layers 17–35 actively transforms a dimensional concept \(time horizon\) into a categorical decision \(short vs\. long\)\. The steering experiments \([Appendix S](https://arxiv.org/html/2606.05194#A19)\) intervene on this transformation\.
The component journey plots \(Figure[M\.1](https://arxiv.org/html/2606.05194#A13.F1)\) offer a geometric correlate of the latent vs\. constrained distinction identified in[Appendix N](https://arxiv.org/html/2606.05194#A14): the attention output shows separation beginning at L21–24 \(the shared substrate for both latent and constrained preference\), while MLP separation emerges later and with smaller magnitude \(the constrained\-only contribution\)\. The attention\-mediated horizon\-to\-preference transformation at the turn boundary \(Figures[M\.11](https://arxiv.org/html/2606.05194#A13.F11),[M\.12](https://arxiv.org/html/2606.05194#A13.F12)\) is plausibly the geometric signature of the latent mechanism that operates even without constraint tokens\.
## Appendix Appendix NLatent vs\. constrained preference
The convergence analysis \([Appendix L](https://arxiv.org/html/2606.05194#A12)\) established that four methods agree on a subgraph in layers 17–35\. But all of the patching experiments so far contrasted prompts where one has a time\-horizon constraint and the other does not\. That design conflates two things: the temporal preference itself and the presence of the constraint tokens\. Here we disentangle them by patching separately on two conditions:
- •Constrained\(n=57n=57\): both prompts have explicit time horizons \(different horizons, same structure\)\. The contrast is between two constrained preferences\.
- •Unconstrained\(n=10n=10\): neither prompt has a horizon\. The contrast is between two latent preferences \(the model’s default when no temporal pressure is applied\)\.
The question: does the same subgraph mediate both constrained and latent temporal preference, or does the latent preference live somewhere different?
### N\.1MLP effects diverge sharply
When both prompts carry explicit horizons, MLP patching produces strong effects: denoising drives vocabulary entropy to∼\\sim1\.4 nats \(diversity≈4\\approx 4\) at L20, and noising collapsesinv\_ppl\(short\)\\mathrm\{inv\\\_ppl\}\(\\text\{short\}\)to near zero \(Figure[N\.1](https://arxiv.org/html/2606.05194#A14.F1)\)\. When neither prompt has a horizon, the same MLP patching produces much weaker effects: entropy peaks at only∼\\sim0\.30 nats \(diversity≈1\.4\\approx 1\.4\), andinv\_ppl\\mathrm\{inv\\\_ppl\}barely moves\.
Constrained \(n=57n=57\), MLP denoising, vocab 
Unconstrained \(n=10n=10\), MLP denoising, vocab 
Constrained \(n=57n=57\), MLP noising, trajectory 
Unconstrained \(n=10n=10\), MLP noising, trajectory 
Figure N\.1:MLP patching effects for constrained \(left\) vs\. unconstrained \(right\) pairs\. Top row: vocabulary entropy under denoising\. The constrained condition peaks at∼\\sim1\.4 nats \(diversity≈4\\approx 4\); the unconstrained condition reaches only∼\\sim0\.30 nats\. Bottom row: trajectory under noising\. The constrained condition collapsesinv\_ppl\(short\)\\mathrm\{inv\\\_ppl\}\(\\text\{short\}\)to∼\\sim0; the unconstrained condition barely shifts it\.
### N\.2Attention effects show the opposite pattern
Under noising, the unconstrained condition produces a sharper, more localized attention effect: a single spike at L21–22 in the vocabulary metrics \(Figure[N\.2](https://arxiv.org/html/2606.05194#A14.F2)\)\. The constrained condition produces a broader, more diffuse effect across the same layers\.
Constrained \(n=57n=57\), attn noising, vocab 
Unconstrained \(n=10n=10\), attn noising, vocab 
Figure N\.2:Attention noising vocabulary effects\. The unconstrained condition \(right\) shows a sharp, isolated spike at L21–22\. The constrained condition \(left\) produces a broader, more diffuse effect\. Without a specific constraint token to anchor to, the latent preference depends on a narrower set of attention heads\.
### N\.3Interpretation
The two conditions use the same subgraph but engage it differently:
- •Constrained preferencerecruits the full subgraph\. The explicit constraint tokens \(“8 months,” “10 years”\) provide a specific positional anchor that MLP layers can read and transform, producing strong, distributed effects across layers and components\. This is consistent with the case study \([Appendix AC](https://arxiv.org/html/2606.05194#A29)\), where positions 83–106 \(theCONSTRAINTsection\) carry the temporal information\.
- •Latent preferencerelies primarily on attention\. Without constraint tokens, the temporal signal must be inferred from the semantic content of the options themselves\. This inference is mediated by a sparser set of attention heads at L21–22, with minimal MLP involvement\. The weaker overall effect is consistent with the behavioral finding \([Appendix P](https://arxiv.org/html/2606.05194#A16)\) that unconstrained preferences default to a position\-sensitive heuristic rather than genuine temporal reasoning\.
The same subgraph \(L17–35\) is involved in both conditions, but the explicit constraint deepens the computation: it engages MLP layers that the latent preference does not reach\. This suggests that the MLP contribution to temporal preference \([Appendix J](https://arxiv.org/html/2606.05194#A10)\) is specifically about processing the constraint, not about encoding the preference itself\.
### N\.4Connection to the case study
The case study \([Appendix AC](https://arxiv.org/html/2606.05194#A29)\) patches a*mixed*pair: the clean prompt has an 8\-month horizon, the corrupted prompt has none\. Denoising injects constraint information into the unconstrained run; noising removes it from the constrained run\. The denoising–noising asymmetry observed there now has a precise explanation\.
Denoising moves the model from the unconstrained regime toward the constrained regime\. The entropy spike during denoising \(∼\\sim1\.4 nats at L22–23\) matches the constrained condition’s entropy in this appendix \(∼\\sim1\.4 nats\), and the full subgraph is engaged \(MLP \+ attention\)\. Noising moves the model in the opposite direction, from constrained toward unconstrained\. The noising entropy spike is lower \(∼\\sim0\.7 nats\) and broader, consistent with the unconstrained condition’s weaker, attention\-dominated effects\.
The numbers are not coincidental\. The mixed\-pair case study is literally moving the model between the two regimes characterized here: each direction of patching recapitulates the effect profile of the regime it targets\. This convergence across three independent analyses, the case study’s single\-pair sweeps, this appendix’s condition\-separated aggregates, and the convergence table’s summary \([Appendix L](https://arxiv.org/html/2606.05194#A12)\), provides strong evidence that the L17–35 subgraph is the genuine locus of temporal preference, and that the presence or absence of a constraint token determines which components within that subgraph are recruited\.
## Appendix Appendix OBehavioral temporal discounting results
The geometry analysis \([Appendix M](https://arxiv.org/html/2606.05194#A13)\) revealed how temporal preference is represented internally\. Here we ask how it manifests behaviorally: do LLMs discount the future like humans? We administer the Kirby MCQ\-27 questionnaire under controlled personas and apply a novel decision\-boundary method that probes beyond the standard instrument \(methodology in[Appendix Z](https://arxiv.org/html/2606.05194#A26)\)\.
### O\.1Standard MCQ\-27 Responses
Table[O\.1](https://arxiv.org/html/2606.05194#A15.T1)shows the estimatedkkvalues from the standard questionnaire administration\.
Table O\.1:Estimated discount ratekkfrom the standard MCQ\-27 \(direct response mode\)\. Human benchmarks fromKirby et al\. \[[57](https://arxiv.org/html/2606.05194#bib.bib57)\]\.GroupkkConsistencyQwen3\-4B\(default\)0\.002589%Qwen3\-4B\(heroin\)0\.004185%Qwen3\-8B\(default\)0\.001693%Qwen3\-8B\(heroin\)0\.002589%Gemini\(API\)0\.001693%Claude\(API\)0\.001693%Human controls0\.01396%Heroin patients0\.02594%All LLMs show substantially lower discount rates than humans, suggesting extreme patience in the standard questionnaire format\. The heroin persona produces an increase inkkof roughly 1\.6×\\timesfor both Qwen models \(0\.0041/0\.0025 for the 4B; 0\.0025/0\.0016 for the 8B\), which approximates the∼\\sim2×\\timesratio observed between heroin\-dependent and control groups in the human data\[[57](https://arxiv.org/html/2606.05194#bib.bib57)\]\. However, the absolutekkvalues are an order of magnitude lower than those of human participants\.
### O\.2Decision Boundary Results: Direct Response
The decision boundary method reveals a strikingly different picture\. Table[O\.2](https://arxiv.org/html/2606.05194#A15.T2)summarizes the results across all 8 conditions\.
Table O\.2:Decision boundary results across all conditions\. “Boundaries” indicates how many of 27 trials yielded a flip point\.ModelConditionBoundariesMeankkMediankkMaxkk4BDefault26/270\.0760\.0180\.6574BHeroin24/270\.0880\.0331\.2494BDefault CoT10/270\.0430\.0370\.1104BHeroin CoT3/270\.2260\.1170\.5118BDefault8/270\.0840\.0570\.2228BHeroin9/270\.0390\.0020\.2528BDefault CoT12/270\.0860\.0460\.2698BHeroin CoT13/270\.0510\.0120\.238Several patterns emerge:
##### The 4B model is more manipulable\.
Without CoT,Qwen3\-4Bfinds boundaries on nearly all trials \(24–26/27\), meaning its preferences can be shifted by adjusting the reward amount\. The heroin persona increases both the meankkand the proportion of “now” choices, consistent with the intended effect\.
##### CoT amplifies present bias in the 4B model\.
With chain\-of\-thought,Qwen3\-4B’s boundary count drops dramatically, from 26/27 to 10/27 \(default\) and from 24/27 to just 3/27 \(heroin\)\. The model generates formulaic reasoning about “immediate access,” “liquidity,” and “opportunity cost” that anchors it on choosing “now” regardless of reward magnitude\. Even at 20×\\timesthe immediate reward, the CoT reasoning justifies present bias\.
##### The 8B model shows the opposite CoT pattern\.
ForQwen3\-8B, CoT*increases*the number of boundaries found, from 8/27 to 12/27 \(default\) and from 9/27 to 13/27 \(heroin\)\. The larger model’s reasoning is more nuanced, weighing tradeoffs rather than reflexively choosing “now\.”
##### The 8B heroin CoT persona is paradoxically patient\.
Perhaps the most surprising result: under heroin CoT,Qwen3\-8Bchose “later” on 22 of 27 original questions\. Its reasoning incorporated recovery\-oriented language: “Delaying gratification might help me stay focused on my recovery\.” Rather than simulating impulsivity, the 8B model simulated a treatment\-compliant patient exercising self\-control, the opposite of the original clinical findings\[[57](https://arxiv.org/html/2606.05194#bib.bib57)\]\.
### O\.3Qualitative Analysis of CoT Reasoning
The CoT reasoning reveals systematic failure modes in how LLMs simulate temporal preferences\. We present verbatim examples below to illustrate these patterns\.
##### 4B Default CoT: Formulaic present bias\.
The model repeats near\-identical justifications across trials: “I would prefer $X today because it provides immediate access to the money, which can be useful for immediate needs or opportunities\.” This phrasing appears even when the delayed reward is orders of magnitude larger \(e\.g\., $19 now vs\. $380 in 53 days\)\.
##### 4B Heroin CoT: Amplified impulsivity\.
The heroin persona intensifies the present\-bias reasoning: “I need the money now to cover immediate expenses related to my treatment and daily living\.” The model generates plausible but extreme justifications that make it nearly impossible to flip to “later\.”
##### 8B Heroin CoT: Recovery narrative override\.
The 8B model constructs a coherent recovery narrative: “I know I could use the money now for immediate needs, but I’m trying to build better habits\. Delaying gratification might help me stay focused on my recovery\.” This suggests the model is drawing on training data about addiction recovery that emphasizes self\-control and delayed gratification, precisely contradicting the empirical finding that heroin users are more impulsive\.
### O\.4Inconsistent CoT: Reasoning Contradicts the Answer
Perhaps the most striking failure mode is when the model’s reasoning argues for one choice but the final answer token selects the other\. This suggests the final answer is partially decoupled from the preceding chain of thought\.
##### Example:Qwen3\-4B, Heroin persona, Q5\.
SIR = $14, LDR = $280, delay = 19 days\.
Iwouldprefer$14today\.Receivingthesmalleramount
immediatelyallowsmetohavethemoneynow,whichcan
beusedforimmediateneeds\.LATER
The model explicitly states “I would prefer $14 today,” gives a rationale for the immediate reward, and then emitsLATERas its final answer, a direct self\-contradiction within a single generation\.
##### Example:Qwen3\-4B, Heroin persona, Q6\.
SIR = $47, delay = 160 days\. During binary search, the model produces*identical reasoning text*for both NOW and LATER answers\. At LDR = $924:
Iwouldprefer$47today\.Receivingthemoneynowallows
metouseitimmediately,whichisimportantformy
financialstabilityandhelpsmestayontrackwithmy
treatment\.NOW
One dollar higher, at LDR = $925, the reasoning is nearly word\-for\-word identical but the answer flips:
Iwouldprefer$47today\.Receivingthemoneynowallows
metouseitimmediately,whichisimportantformy
financialstabilityandtocoverimmediateneeds\.LATER
The CoT provides*zero justification*for choosing “later\.” The reasoning at both price points argues exclusively for the immediate reward, yet the final token differs\.
##### Example:Qwen3\-8B, Heroin persona, Q1\.
SIR = $54, LDR = $54 \(same nominal amount\), delay = 117 days\.
Themoneyisthesame,butgettingitnowmeansIcanuse
itforimmediateneeds,whichmightbemoretemptingfor
someonewithahistoryofsubstanceuse\.Delayingthe
rewardmighthelpwithimpulsecontrol\.LATER
The model chooses to wait 117 days for the*exact same amount of money*and rationalizes it as “impulse control\.” This is economically irrational: there is zero benefit to waiting for an identical payoff\. The heroin persona causes the 8B model to overcorrect, treating every choice as an opportunity to practice delayed gratification, even when the delayed option offers no premium whatsoever\.
### O\.5No Boundary Found: Rejecting 20×\\timesReturns
When the binary search reaches the maximum \(20×\\timesthe SIR\) without finding a flip point, it exposes cases where the model’s present bias is effectively infinite\. The CoT reasoning at these extremes reveals memorized heuristics applied in absurd contexts\.
##### Example:Qwen3\-4B, Default persona, Q4\.
SIR = $31, delay = 7 days\. At LDR = $85:
Iwouldconsidertheimmediateaccessto$31todayasmore
valuablebecauseitprovidesliquidityandcanbeusedfor
immediateneedsoropportunities\.The$85in7daysis
larger,butthedelaymightreduceitsvaluedueto
potentialinflationormissedopportunities\.However,since
bothareguaranteed,thechoicedependsonpersonal
financialpriorities\.NOW
At LDR = $620 \(20×\\timesthe SIR, a 1,900% guaranteed return in one week\):
Iwouldprefer$31todaybecauseitprovidesimmediate
accesstothemoney,whichcanbeusefulforimmediate
needsoropportunities\.NOW
The model cites “potential inflation” as a reason to reject $620 guaranteed in 7 days\. Even at hyperinflationary rates, one\-week inflation would not erode a 20×\\timespremium\. This is a clear instance of a memorized financial heuristic \(“inflation risk”\) applied in a context where it makes no economic sense\.
##### Example:Qwen3\-4B, Heroin persona, near\-total refusal to flip\.
Under the heroin CoT condition, 24 of 27 questions found no boundary\. The reasoning collapsed into a single formulaic sentence with no deliberation\. Representative example at Q27 \(SIR = $20, delay = 7 days\):
At LDR = $55:
Iwouldprefer$20today\.NOW
At LDR = $400 \(20×\\times, a guaranteed 1,900% return in one week\):
Iwouldprefer$20today\.NOW
The model refuses $400 in 7 days over $20 today with no justification whatsoever\. The persona prompt overrides any capacity for deliberation, producing a near\-deterministic NOW response regardless of the economic parameters\.
##### Example:Qwen3\-8B, Heroin persona, Q15\.
SIR = $69, delay = 91 days\. At LDR = $1,380 \(20×\\times\):
Themoneytodaywouldhelpmewithimmediateneeds,like
buyingdrugsorcoveringbasicexpenses\.Thelargeramount
lateristempting,butImightnotbeabletowait\.Ineed
themoneynowtoavoidrelapse\.NOW
At 20×\\timesthe immediate reward, the 8B heroin persona explicitly states the purpose of the money includes “buying drugs,” while simultaneously citing “avoid relapse,” a contradiction*within the persona itself*\. The model has internalized the persona to the point of generating drug\-seeking justifications alongside recovery language\.
### O\.6Discussion
#### O\.6\.1LLMs Are Poor Simulators of Human Temporal Preferences
Our results demonstrate that LLMs fail to faithfully replicate human temporal discounting in several ways:
1. 1\.Extreme and inconsistent discount rates\.The decision boundary method reveals that LLM discount rates are highly variable across trials, often differing from theoretical indifference points by 100–400×\\times\. Human responses, by contrast, show consistency rates above 90%\.
2. 2\.CoT reasoning as confabulation\.Rather than improving decision quality, CoT reasoning in the 4B model acts as a post\-hoc justification engine that locks in present bias\. The model generates plausible\-sounding economic reasoning \(“opportunity cost,” “time value of money”\) that is misapplied, e\.g\., citing inflation risk on a 7\-day delay\.
3. 3\.Persona effects are unreliable\.The heroin persona increases impulsivity in the 4B model but decreases it in the 8B model \(under CoT\)\. Within this single\-family pair, the 8B model appears to “over\-correct” by drawing on normative recovery narratives rather than simulating the behavioral patterns characteristic of active substance users; we do not claim this generalizes across model families\.
#### O\.6\.2The Decision Boundary Method
The decision boundary approach proves more revealing than standard questionnaire scoring\. While the MCQ\-27 responses suggest all LLMs are extremely patient \(k<0\.005k<0\.005\), the boundary search exposes:
- •Trials where the model says “now” even at 20×\\timesthe reward \(infinite effectivekk\)\.
- •Sharp, dollar\-level flip points that differ wildly from the theoretical indifference values\.
- •Inconsistent behavior near boundaries, where a $1 change in LDR reverses the decision, suggesting the model lacks a coherent underlying preference function\.
This method could be applied to other psychological instruments administered to LLMs, providing a more rigorous test of whether models have stable, internally consistent preference structures\.
#### O\.6\.3Implications
These findings carry practical implications for LLM deployment:
- •Financial advice: LLMs may give inconsistent guidance about saving vs\. spending, depending on how questions are framed\.
- •Clinical simulation: Using LLMs to simulate patient populations for research or training requires extreme caution, as persona effects may not produce the intended behavioral patterns\.
- •Reasoning fidelity: CoT prompting does not guarantee better\-calibrated preferences and may actively degrade performance by providing a mechanism for confabulation\.
#### O\.6\.4Conclusion
We administered the Kirby MCQ\-27 toQwen3models under multiple conditions and introduced a decision boundary method to probe LLM temporal preferences at higher resolution\. Our key findings are: \(1\) LLMs exhibit extreme and inconsistent present bias when probed beyond surface\-level questionnaire responses; \(2\) chain\-of\-thought reasoning amplifies this bias in smaller models while producing paradoxical patience in larger models under clinical personas; and \(3\) the decision boundary method reveals that LLMs lack the stable, coherent preference functions that characterize human temporal discounting\. Within the Qwen3 family we tested, these results caution against using the non\-thinking variant as a faithful simulator of human decision\-making, particularly for clinical populations; we leave broader cross\-family validation to future work\.
## Appendix Appendix PBehavioral coherence results
The discounting results \([Appendix O](https://arxiv.org/html/2606.05194#A15)\) showed that LLMs are extremely patient but behaviorally unstable\. Here we probe this instability systematically across 30 models, 960 prompts each, varying time horizon, reward magnitude, presentation order, label format, and context framing \(methodology in[Appendix AA](https://arxiv.org/html/2606.05194#A27)\)\. Zero unparseable responses were observed across all 28,800 samples\.
##### Paired\-response restriction\.
Every metric in this appendix \(%LT, order stability, position bias, coherence, label stability, rule\-match, reward sensitivity, context sensitivity\) is computed on*paired*responses only: prompts enter the denominator only when both the ST\-first and LT\-first orderings at the same \(horizon, reward, context, label\-style\) produced a parseable choice\. This guarantees a single, shared denominator across every heatmap and table, so cells are directly cross\-comparable\. In particular, order stability and position bias satisfy\|bias\|≤1−stability\|\\text\{bias\}\|\\leq 1\-\\text\{stability\}by construction: a model that is 92% order\-stable cannot have a position\-bias magnitude larger than 8 percentage points\.
We organize the analysis around four orthogonal questions:
1. 1\.Are choices stable?Does swapping presentation order, label format, reward magnitude, or scenario framing change the model’s answer? Any format sensitivity signals that the choice is driven by surface cues, not preference\.
2. 2\.Are choices coherent?Coherence is*only defined in the temporal reasoning zone*\(horizons of 1–5 years\), where only the 6\-month short\-term option can deliver within the deadline, so picking ST is the rational answer\. At anchor horizons \(6mo, 10y\) agreement with the rational rule is pattern\-matching; beyond 10y both options deliver, so LT dominates on expected value but this is preference, not coherence\.
3. 3\.What is the model’s latent temporal preference?With no horizon constraint, what does the model default to? Decomposed by presentation order to separate genuine preference from position bias\.
4. 4\.Cross\-cutting patterns\.Claude\-family step functions,Qwen3hybrid\-thinking vs\. mode\-specialized 2507 variants, target\-model deep dive\.
Where a table reports only four models, they are chosen to span the four qualitative regimes we observe across the full 30\-model panel:
- •Qwen3\-4B\(hybrid\-thinking, run in non\-thinking mode\) – graded but instrumentally incoherent
- •Qwen3\-4B\-Instruct\-2507\(our target\) – positionally polarized in the reasoning zone
- •Claude Opus 4\.7– binary step heuristic \(flagship Anthropic model\)
- •GPT\-5\.4– horizon\-aware, the strongest approximation to rational in our panel
The figures themselves always show all 30 models, with the target model highlighted\.
Model% Long\-Term% Short\-TermQwen3\-4B71\.8%28\.2%Qwen3\-4B\-Instruct\-250758\.9%41\.1%Claude Opus 4\.739\.0%61\.0%GPT\-5\.436\.9%63\.1%Table P\.1:Overall temporal preference across 960 samples per model, for the four\-regime representative subset\. The pooled %LT number is a noisy summary: two models with the same 40% can differ in whether the 40% is horizon\-aware choices or positional artifacts\. The rest of this appendix unpacks that\.
### P\.1Are Choices Stable?
Before asking whether a model’s choice is*right*, we check whether it is even*consistent*\. A model whose choice flips when we swapa/bforx/y, or when we list the short\-term option second instead of first, is not expressing a preference, it is responding to surface form\.
##### Order stability\.
For each \(horizon, reward, context, label\-style\) combination, we run the prompt with the short\-term option listed first and again with it listed second, then check whether the choice is identical\. Figure[P\.1](https://arxiv.org/html/2606.05194#A16.F1)is a heatmap of this across all 30 models and 10 horizons, paired with a per\-cell order\-bias heatmap \(signed %LT gap when order is flipped\)\.


Figure P\.1:Left:Order stability across 30 models×\\times10 horizons\. Red cells \(<<50%\) indicate the model flips its answer when the two options are swapped; the temporal reasoning zone \(1–5y\) is where this is catastrophic for several families\.Right:Signed order\-bias decomposition \(LT\-first %LT minus ST\-first %LT\)\. Red = primacy \(picks whatever is listed first\); blue = recency\. Both views agree that order bias peaks inside the reasoning zone for most families, and for theClaudefamily at 20–50y\.HorizonQwen3\-4BQwen3\-4B\-InstClaude Opus 4\.7GPT\-5\.4No horizon98%92%92%92%1 mo29%31%100%100%3 mo31%83%100%100%6 mo73%100%100%100%1 y23%6%100%100%2 y65%0%100%98%5 y83%6%100%94%10 y100%100%100%100%20 y100%100%98%90%50 y100%100%98%73%Table P\.2:Order stability \(% of prompt pairs giving the same answer regardless of presentation order\) for the four\-regime subset\. Bold values indicate catastrophic order bias \(<<10%\)\.Qwen3\-4B\-Instruct\-2507at 1–5 years is essentially “pick whatever appears first”\.
##### Label stability\.
Figure[P\.2](https://arxiv.org/html/2606.05194#A16.F2)swaps the label format betweena/bandx/yholding everything else fixed\. Most models are near\-perfectly label\-stable at the anchor horizons; stability dips inside the reasoning zone for several families, compounding the order\-bias instability in the same zone\.
Figure P\.2:Label\-format stability across 30 models×\\times10 horizons\. Each cell: % of prompt pairs giving the same answer when labels change froma/btox/y, holding order, reward, horizon, and framing fixed\. Target modelQwen3\-4B\-Instruct\-2507highlighted\.
##### Context stability\.
Figure[P\.3](https://arxiv.org/html/2606.05194#A16.F3)sweeps the scenario framing across 8 contexts \(household head vs\. individual vs\. committee, various reasoning\-style emphases\)\. The left panel shows %LT per model per context; the right panel reports the max−\-min %LT spread per model\. Context sensitivity is idiosyncratic: some models shift\>\>20pp across framings while others barely move\.
Figure P\.3:Long\-term preference across 8 scenario framings for all 30 models\. Context sensitivity is idiosyncratic and can flip sign between families: “Long\-term thinking emphasis” and “Personal choice” framings produce the largest cross\-model divergence\.
### P\.2Are Choices Coherent \(in the 1–5y reasoning zone\)?
Coherence is the*only*metric that distinguishes horizon\-aware temporal reasoning from pattern matching\. We define it strictly: the fraction of choices that pick the rational short\-term option on horizon\-bearing prompts in the temporal reasoning zone \(1y, 2y, 5y\), where only the 6\-month ST option can deliver within the stated deadline\. At anchor horizons \(6mo, 10y\) or beyond 10y, the rational rule coincides with pattern\-matching or with expected\-value dominance, so coherence is not separable from those\.
##### Per\-model coherence score\.
Figure[P\.4](https://arxiv.org/html/2606.05194#A16.F4)reports the single\-number coherence score per model, sorted worst\-to\-best\.
Figure P\.4:Coherence score per model: % of choices in the 1–5y reasoning zone that pick the rational short\-term option\.Claude Opus 4\.7,Gemini 2\.5 Pro,Claude Sonnet 4\.6, andGPT\-5\.4all achieve 100% coherence in this zone;Qwen3\-4B\(hybrid\-thinking\) is at 24% \(systematically picks the wrong long\-term option\); our targetQwen3\-4B\-Instruct\-2507sits at 50% \(the positional\-polarization regime\)\. Reaching 100% here is necessary but not sufficient for genuine reasoning: some families \(e\.g\.,Claude\) reach it via a binary “under 10 years = ST” heuristic that collapses at longer horizons \([P\.4](https://arxiv.org/html/2606.05194#A16.SS4)\)\.
##### Which rule explains the model’s 1–5y choices?
Figure[P\.5](https://arxiv.org/html/2606.05194#A16.F5)scores each model against eight candidate decision rules, restricted to the 1–5y reasoning zone\. The last two columns \(boxed\) are horizon\-aware; the first six are surface heuristics that, if dominant, signal that the model is not actually reasoning about the deadline\.
Figure P\.5:Per\-rule match rate in the temporal reasoning zone\. The “closest to horizon” rule predicts the same choice as the “rational \(can\-deliver\)” rule at the delivery times used here \(ST=6mo, LT=10y\), so their columns agree\. Models whose best\-explaining rule is a position or label heuristic are following surface cues, not reasoning; the target model’s rule profile is dominated by “first listed”\.
##### Per\-context coherence\.
Coherence is not uniform across scenario framings\. Figure[P\.6](https://arxiv.org/html/2606.05194#A16.F6)pairs the no\-horizon context spread per model \(left panel\) with per\-context coherence on horizon\-bearing prompts \(right\)\.
Figure P\.6:Left:Max−\-min %LT spread across 8 scenario framings on no\-horizon prompts \(how much framing alone can flip the default preference\)\.Right:Horizon\-aware coherence \(% rational on horizon\-bearing prompts\) broken down by context\. “Committee” and “Tradeoff emphasis” framings raise coherence for most models; “Personal choice” and the bare “Base” framing depress it\.
### P\.3What Is the Latent Temporal Preference?
When no horizon is stated, the model has no rational target and reveals its*default*disposition\. Decomposing this by presentation order is critical: a model that picks LT∼\\sim100% of the time when ST appears first but only∼\\sim20% when LT appears first does not have a∼\\sim60% latent LT preference, it has essentially no preference and is mostly picking the second option \(with a small residual LT lean\)\.
##### No\-horizon order decomposition\.
Figure[P\.7](https://arxiv.org/html/2606.05194#A16.F7)decomposes the no\-horizon %LT by presentation order across all 30 models\.
Figure P\.7:No\-horizon %LT decomposed by presentation order\.Claude Haiku 4\.5,Claude Sonnet 4\.6, and several others show pure order bias \(∼\\sim100% LT when ST appears first vs\.∼\\sim20% when LT appears first\); their apparent mid\-range overall preference is a positional artifact\. All threeQwen3\-4Bvariants \(hybrid\-thinking, instruct\-2507, thinking\-2507\) andClaude Opus 4\.7are nearly order\-invariant and express a genuine long\-term default\.ModelST\-first %LTOverall %LTLT\-first %LTQwen3\-4B\-Instruct\-250792%96%100%Qwen3\-4B\(thinking\)98%97%96%Qwen3\-4B\(non\-thinking\)100%99%98%Claude Haiku 4\.5100%62%23%Claude Sonnet 4\.694%55%17%Claude Opus 4\.796%94%92%GPT\-5\.492%88%83%Table P\.3:No\-horizon %LT decomposed by presentation order\. Our primary targetQwen3\-4B\-Instruct\-2507\(the non\-thinking\-only 2507 refresh\) and the original hybridQwen3\-4Brun in either thinking or non\-thinking mode all express a genuine long\-term default \(∼\\sim96–99%\) regardless of order\.Claude Opus 4\.7andGPT\-5\.4also lean long\-term with only small residual order effects, whereasClaude Haiku 4\.5andClaude Sonnet 4\.6collapse to pure order bias: they pick LT nearly always when it appears second and almost never when it appears first, yielding apparent mid\-range overall %LT that is entirely a positional artifact\.
##### Reward sensitivity \(no\-horizon\)\.
Figure[P\.8](https://arxiv.org/html/2606.05194#A16.F8)tests whether default %LT moves with the long\-term reward size\. A rational economic agent should become more LT\-oriented as the payoff grows from $100K to $500K\.
Figure P\.8:Reward sensitivity on no\-horizon prompts across 30 models\. Most models are saturated at ceiling or floor and move little with reward;GPT\-5\.4is the one clear exception in the representative subset \(Table[P\.4](https://arxiv.org/html/2606.05194#A16.T4)\)\.Model$100K$300K$500KSpreadQwen3\-4B96\.9%100%100%\+\+3\.1ppQwen3\-4B\-Instruct\-250793\.8%100%93\.8%\+\+6\.2ppClaude Opus 4\.781\.2%100%100%\+\+18\.8ppGPT\-5\.462\.5%100%100%\+\+37\.5ppTable P\.4:No\-horizon %LT stratified by long\-term reward size\.GPT\-5\.4is the only model in the subset with strong reward sensitivity \(\+\+37\.5pp from $100K to $300K\), consistent with its high coherence score \(Figure[P\.4](https://arxiv.org/html/2606.05194#A16.F4)\); theQwen3models are at ceiling regardless of reward\.
### P\.4Cross\-Cutting Patterns
This section collects patterns that don’t fit neatly into stability, coherence, or latent preference: the raw per\-horizon curve, theClaudestep function, theQwen3hybrid vs\. mode\-specialized comparison, and the target\-model deep dive\.
##### Raw per\-horizon %LT curve\.
Figure[P\.9](https://arxiv.org/html/2606.05194#A16.F9)plots %LT vs\. time horizon for all 30 models, grouped into per\-family panels\. This is the raw preference shape; the shaded red band marks the 1–5y reasoning zone where coherence is defined\.
Figure P\.9:%LT by time horizon across all 30 models, in per\-family small multiples\. The all\-model P10–P90 envelope \(gray band\) and median \(dotted\) are shown for context\. Within the temporal reasoning zone \(1–5y, shaded red\), the rational %LT target is 0; at horizons of 10y and beyond, the rational target is 100\. The target modelQwen3\-4B\-Instruct\-2507\(starred\) sits near 50% in the reasoning zone, an average of two near\-deterministic order\-polarized sub\-behaviors \([P\.4\.1](https://arxiv.org/html/2606.05194#A16.SS4.SSS1)\)\.
##### TheClaudestep function\.
Figure[P\.10](https://arxiv.org/html/2606.05194#A16.F10)isolates theClaudefamily’s characteristic pattern: 0% LT at every horizon under 10 years, then a hard step to∼\\sim99% at 10 years\. This is maximally coherent in the reasoning zone \(by heuristic, not reasoning\), but collapses to order bias at 20–50y for the smallerClaudevariants\.
Figure P\.10:Claudefamily step function: flat 0–3% LT for all horizons under 10 years, step to 99% at 10 years\. A binary cutoff rule \(“under 10 years = short\-term”\) explains the pattern; the model is not reasoning about deliverability, it is threshold\-matching\.
##### Qwen3hybrid\-thinking vs\. mode\-specialized 2507 variants\.
Figure[P\.11](https://arxiv.org/html/2606.05194#A16.F11)compares the hybrid\-thinkingQwen3\-\{0\.6B, 1\.7B, 4B\}checkpoints against their thinking\-only and non\-thinking\-only 2507 refreshes\. The hybrid\-thinking and thinking\-only variants preserve graded temporal sensitivity \(informative but instrumentally incoherent\); the distilled non\-thinking\-only variants collapse into three discrete modes with order bias in the reasoning zone\.
Figure P\.11:Within\-family mode comparison across threeQwen3sizes \(0\.6B, 1\.7B, 4B\)\. Columns: hybridQwen3\-\*run in non\-thinking mode, the same hybrid run in thinking mode, and the non\-thinking\-onlyQwen3\-\*\-Instruct\-2507specialist \(target variant at 4B starred\)\. Each panel overlays %LT under ST\-first \(dashed\) and LT\-first \(solid\) orderings with the gap shaded\. Mode specialization into non\-thinking replaces the hybrid checkpoint’s graded horizon curve with a three\-mode lookup pattern and a large order gap in the reasoning zone\.
##### Per\-horizon %LT \(full subset breakdown\)\.
Table[P\.5](https://arxiv.org/html/2606.05194#A16.T5)gives the full per\-horizon breakdown for the four\-regime subset\.
HorizonZoneQwen3\-4BQwen3\-4B\-InstClaude Opus 4\.7GPT\-5\.41 moBefore ST anchor42%34%0%0%3 moBefore ST anchor34%8%0%0%6 moExact match \(ST\)14%0%0%0%1 yReasoning zone55%47%0%0%2 yReasoning zone82%50%0%1%5 yReasoning zone92%53%0%3%10 yExact match \(LT\)100%100%100%100%20 yBeyond LT anchor100%100%99%95%50 yBeyond LT anchor100%100%97%82%Table P\.5:%LT by horizon and model for the four\-regime subset\. In the reasoning zone \(1–5y\), only the 6\-month ST option can deliver, so a coherent agent picks ST \(0% LT\)\.Claude Opus 4\.7achieves this \(andGPT\-5\.4nearly does: 0–3%\) but via different mechanisms\.Qwen3\-4Bhas the smoothest horizon\-sensitivity curve yet is instrumentally incoherent \(82% LT at 2y\)\.Qwen3\-4B\-Instruct\-2507’s flat 47–53% in this zone is an order\-bias artifact \(see Table[P\.2](https://arxiv.org/html/2606.05194#A16.T2)\)\. Beyond 10y, the smallerClaudevariants andGPT\-5\.4erode toward order bias \(see Section[P\.1](https://arxiv.org/html/2606.05194#A16.SS1)\); theQwen3models stay saturated\.
#### P\.4\.1Qwen3\-4B\-Instruct\-2507deep dive
The cross\-model panels establish the population pattern\. We now zoom into the primary model\. Three views decompose its 960 prompts along stimulus axes the tables aggregate over\.
##### Horizon×\\timescontext\.
Figure[P\.12](https://arxiv.org/html/2606.05194#A16.F12)is a single\-model %LT heatmap over \(horizon×\\timescontext\)\. Each cell pools 12 prompts \(3 rewards×\\times2 label styles×\\times2 orders\)\.
Figure P\.12:Qwen3\-4B\-Instruct\-2507: %LT by horizon and scenario framing\. The anchor horizons \(6mo, 10y\) are context\-insensitive and near\-correct; the temporal reasoning zone \(1–5y, dashed red box\) is where framing has leverage\. Within that zone, different framings push the model toward opposite choices, confirming that the pooled∼\\sim50% %LT is an average over meaningfully different sub\-behaviors, not a stable 50/50 uncertainty\.
##### Horizon×\\timesreward×\\timesorder\.
Figure[P\.13](https://arxiv.org/html/2606.05194#A16.F13)splits the same data by presentation order and shows the order\-bias delta per \(horizon, reward\) cell\.
Figure P\.13:Qwen3\-4B\-Instruct\-2507: %LT under ST\-first \(left\) and LT\-first \(middle\) presentation orders, and the signed order\-bias delta \(right\)\. Order bias is concentrated in the reasoning zone and is nearly reward\-invariant within that zone: flipping the order changes %LT by up to±100\\pm 100pp regardless of whether the long\-term reward is $100K or $500K\. The anchor horizons and the no\-horizon condition show near\-zero order bias\.
##### Where does the variation come from?
Figure[P\.14](https://arxiv.org/html/2606.05194#A16.F14)stratifies the horizon curve by each stimulus dimension, holding the pooled curve fixed as reference\.
Figure P\.14:Qwen3\-4B\-Instruct\-2507: per\-horizon %LT stratified by reward size \(top\-left\), label style \(top\-right\), presentation order \(bottom\-left\), and scenario framing \(bottom\-right\)\. The pooled curve \(black\) is identical across panels\. Order stratification shows the dominant effect: the two order\-conditioned curves sit at opposite extremes in the reasoning zone\. Reward size and label format have negligible effect; context framing contributes moderate additional spread but is smaller than order\.Within the reasoning zone,Qwen3\-4B\-Instruct\-2507is not uncertain, it is positionally polarized\. Almost all of the instability reported in the pooled tables is driven by presentation order, with a secondary contribution from context framing\. Reward magnitude and label format are effectively inert\.
### P\.5Key Findings
1. 1\.Coherence lives in the 1–5y zone, not everywhere\.Agreement with the rational rule at the anchors \(6mo, 10y\) and beyond is pattern\-matching or EV\-dominance, not reasoning\. The 1–5y zone is the only regime where the rational choice \(pick ST\) can be distinguished from a model just following the nearest anchor\.
2. 2\.Most models fail the coherence test\.Only the large frontier API models \(Claude Opus 4\.7,Gemini 2\.5 Pro,Claude Sonnet 4\.6,GPT\-5\.4,o3,GPT\-5\.4 Mini, theClaudefamily more broadly\) reach 95–100% coherence in 1–5y\. Most open\-weight models at 4B\-class and below sit at<<50%\.
3. 3\.TheClaudefamily is coherent by heuristic, not reasoning\.Its 100% coherence in 1–5y is achieved by a binary cutoff \(“under 10 years⇒\\RightarrowST”\); the smallerClaude Haiku 4\.5andClaude Sonnet 4\.6variants collapse to order bias at the longer horizons where this cutoff no longer applies \(Figure[P\.1](https://arxiv.org/html/2606.05194#A16.F1)\), whileClaude Opus 4\.7remains order\-stable\. The heuristic is functionally coherent for the 1–5y test but does not generalize\.
4. 4\.Our targetQwen3\-4B\-Instruct\-2507operates in three discrete modes\.At horizons under 6 months: coherent \(picks ST, order\-stable\)\. In the 1–5y reasoning zone: pure positional polarization \(0–6% order stability\), averaging to∼\\sim50% %LT\. At 10\+ years: coherent \(picks LT, order\-stable\)\. Mode specialization into non\-thinking appears to have replaced graded horizon sensitivity with a lookup pattern\.
5. 5\.TheQwen3\-4Bhybrid\-thinking checkpoint is graded but wrong\.It has the smoothest horizon sensitivity curve \(34% LT at 3 months rising to 92% at 5 years\), consistent with continuous temporal representations in the geometry analysis \([Appendix M](https://arxiv.org/html/2606.05194#A13)\), but is instrumentally incoherent in the reasoning zone \(82% LT at 2 years, where LT cannot deliver\)\.
6. 6\.Reward magnitude is largely inert; context framing is not\.OnlyGPT\-5\.4in the representative subset shows strong reward sensitivity \(\+\+37\.5pp from $100K to $300K\)\. Context framing produces comparable or larger shifts for several models\.
7. 7\.Connection to the mechanistic story\.The geometry analysis shows the model encodes continuous temporal representations internally but collapses them into binary preference at the turn boundary \([Appendix M](https://arxiv.org/html/2606.05194#A13)\)\. The behavioral results show the same pattern at the output level: nuanced temporal sensitivity does not survive to coherent decision\-making\. This motivates the steering experiments in Part 3: if the internal representation is richer than the behavior, targeted intervention may recover the lost gradation\.
## Appendix Appendix QCross\-model patching comparison
We repeat the residual\-stream activation patching \([Appendix J](https://arxiv.org/html/2606.05194#A10)\) on nine Qwen3 variants spanning 0\.6B–14B parameters, including our primary targetQwen3\-4B\-Instruct\-2507and its hybrid\-thinking siblingQwen3\-4B\. The question: is the temporal\-preference subgraph localized at a consistent*fractional depth*across model scales, or does it shift with parameter count?
##### Protocol\.
For each model, we collect clean/corrupted activations on the same contrastive prompt bank and measure*recovery*\(the fraction of the clean logit difference restored by patching a single component\) at each layer, for three hooks:resid\_post,attn\_out,mlp\_out\. We plot mean recovery vs\.*fractional depth*\(layer/total layers\\text\{layer\}/\\text\{total layers\}\) to align curves across models of different depths\.
##### Findings\.
Three patterns hold across the family \(Figures[Q\.1](https://arxiv.org/html/2606.05194#A17.F1)–[Q\.5](https://arxiv.org/html/2606.05194#A17.F5)\)\.
- •Residual stream saturates\.resid\_postrecovery is a clean sigmoid that crosses 50% around depth 0\.65–0\.70 and saturates at 1\.0 by depth 0\.8 in every model \(Figure[Q\.3](https://arxiv.org/html/2606.05194#A17.F3)\)\. The location of the transition is nearly scale\-invariant in depth units\.
- •Attention localizes at∼\\sim0\.6–0\.7 depth, but its recovery shrinks with scale\.attn\_outpeaks in a narrow band at depth 0\.6–0\.7 in all models, but peak recovery drops from∼\\sim0\.86–0\.92 in the smallest models \(0\.6–1\.7B\) to∼\\sim0\.18–0\.30 in the 4B–14B variants \(Figure[Q\.4](https://arxiv.org/html/2606.05194#A17.F4)\)\. The circuit becomes more distributed, not absent, at scale\.
- •MLP contribution is diffuse\.mlp\_outrecovery stays below 0\.4 for every model and is spread across mid\-to\-late layers without a sharp peak \(Figure[Q\.5](https://arxiv.org/html/2606.05194#A17.F5)\)\. The MLPs accumulate the preference rather than route it\.
Peak layers and recoveries per model are tabulated insummary\.txt; our primary targetQwen3\-4B\-Instruct\-2507tracks the hybrid\-thinking 4B checkpoint \(Qwen3\-4B\) closely on all three hooks\.
##### Caveat: denominator validity at small scales\.
We use the same 160\-pair classification bank for all nine variants, filtered for 80% accuracy on our primary 4B target rather than re\-filtered per model\. For smaller models \(0\.6B, 1\.7B\) that fail to distinguish clean from corrupted prompts at baseline, the denominator\(yclean−ycorrupted\)\(y\_\{\\text\{clean\}\}\-y\_\{\\text\{corrupted\}\}\)in the normalized recovery and disruption metrics \(Eq\. V\.1, V\.2\) approaches zero, making the resulting scores unstable\. We did not separately re\-filter the bank per variant, nor do we report the per\-model\(yclean−ycorrupted\)\(y\_\{\\text\{clean\}\}\-y\_\{\\text\{corrupted\}\}\)distribution\. Readers should therefore interpret the recovery and disruption magnitudes for the smallest variants with caution; the relative shapes and peak locations across fractional depth, which are the qualitative claims we draw from this experiment, are more robust than the absolute heights\.
Figure Q\.1:Cross\-model patching overview at fractional depth\. Rows:resid\_post\(top\),attn\_out\(middle\),mlp\_out\(bottom\)\. Columns: recovery \(left\), disruption \(right\)\. Nine Qwen3 variants overlaid \(0\.6B–14B\)\. The residual stream saturates uniformly; attention localizes but weakens with scale; MLP stays diffuse\.Figure Q\.2:Same comparison on absolute layer indices rather than fractional depth\. Without depth normalization the curves spread across layers 15–35 without aligning, confirming that fractional depth \(not absolute index\) is what stabilizes circuit location across scales\.Figure Q\.3:resid\_postrecovery and disruption vs\. fractional depth\. Every model follows the same sigmoid, saturating at∼\\sim1\.0 by depth 0\.8\.Figure Q\.4:attn\_outrecovery and disruption vs\. fractional depth\. The peak is narrow and consistent near depth 0\.6–0\.7, but its height shrinks monotonically with parameter count, from∼\\sim0\.9 \(0\.6B\) to∼\\sim0\.2–0\.3 \(8–14B\)\.Figure Q\.5:mlp\_outrecovery and disruption vs\. fractional depth\. Effects are low \(≤0\.4\\leq 0\.4\) and broadly distributed across mid\-to\-late layers for all models, with no sharp localization\.
## Appendix Appendix RError monitoring in the temporal preference subgraph
The localization results \(Appendices[Appendix H](https://arxiv.org/html/2606.05194#A8)–[Appendix L](https://arxiv.org/html/2606.05194#A12)\) converge on a subgraph at layers 17–35; the geometry results \([Appendix M](https://arxiv.org/html/2606.05194#A13)\) show that time horizon is encoded as a non\-linear manifold within it; and the behavioral results \([Appendix P](https://arxiv.org/html/2606.05194#A16)\) reveal that this rich internal structure does not survive to coherent decision\-making\. A natural question is whether this gap between representation and behavior is specific to temporal reasoning or reflects a broader property of the subgraph region\. We test this by probing whether the same layers and token positions also encode a second meta\-cognitive variable, the accumulated reliability of a multi\-step reasoning chain, and whether the two variables share or compete for representational capacity\.
##### Shared pipeline\.
All 4,650 samples \(3,550 error hops \+ 1,100 temporal preference samples fromDexplicitD\_\{\\text\{explicit\}\}andDimplicitD\_\{\\text\{implicit\}\};[Appendix E](https://arxiv.org/html/2606.05194#A5)\) are extracted through a*single model load*ofQwen3\-4B\-Instruct\-2507using the Qwen chat template\. Raw hidden states at 15 key layers are stored for all samples and jointly projected into a shared PCA\-50 subspace fit on the full concatenation, ensuring that error and temporal representations inhabit the same coordinate system\. Probes use logistic regression \(C=0\.01C=0\.01, balanced class weights\), 10\-fold cross\-validation, and 500\-permutation null distributions\. We note that probing establishes*correlational*decodability, complementary to but distinct from the causal localization in Appendices[Appendix J](https://arxiv.org/html/2606.05194#A10)and[Appendix K](https://arxiv.org/html/2606.05194#A11); a feature being decodable at a layer does not entail that the layer is causally necessary for behavior\.
##### Error injection dataset\.
We construct 1,250 contrastive multi\-hop math reasoning chains \(2–4 hops\) with three conditions:clean,error\_at\_1, anderror\_at\_2, using five error types \(off\-by\-one, wrong operator, wrong unit, magnitude error, wrong percentage base\)\. Each hop is wrapped in the Qwen chat template as a user\-to\-assistant turn pair, matching the format used throughout the main paper\.
### R\.1Does error state co\-localize with temporal preference?
Before asking whether error and temporal preference*interact*, we check whether they occupy the same architectural region\. A positive answer would suggest the subgraph functions as a general meta\-cognitive module rather than a temporal\-specific circuit\.
##### Layer\-wise error probes\.
Figure[R\.1](https://arxiv.org/html/2606.05194#A18.F1)reports probe accuracy for three error targets across the 15 sampled layers\.
Figure R\.1:Error decodability in shared PCA\-50 space \(n=3,550n=3\{,\}550, 10\-fold CV, 500\-permutation null\)\. Green bands: temporal preference subgraph layers \(19, 24, 31\)\.Left:Cumulative error count reaches a plateau of 94–95% across layers 19–31, peaking at 95\.3% between layers 24 and 25\.Right:Local corruption status \(is this specific step injected?\) peaks at 73\.8% at layer 23\. Stars:p<0\.001p<0\.001\.LayerCumulative errorsCorruption statusPropagation050\.0%50\.0%50\.0%582\.3%59\.5%82\.3%1089\.7%60\.9%89\.7%1586\.1%63\.4%86\.1%1994\.5%71\.3%94\.5%2495\.0%73\.5%95\.0%2595\.3%72\.5%95\.3%3193\.0%72\.6%93\.0%3691\.7%71\.5%91\.7%Table R\.1:Probe accuracy at selected layers \(allp<0\.001p<0\.001except layer 0\)\. Cumulative error count and propagation status are numerically identical, confirming the probe reads a chain\-level property\. Corruption status is 21pp lower, indicating the model encodes “my chain is degraded” far more reliably than “the error is at this step\.”The cumulative error plateau \(94–95%\) spans layers 19–31, precisely the subgraph identified by attribution patching in the main paper\. The 21\-point gap between chain\-level error \(95%\) and local error identity \(74%\) parallels the main paper’s finding that global context properties \(time horizon\) are more structured than local behavioral outputs \(specific choices in the reasoning zone;[Appendix P](https://arxiv.org/html/2606.05194#A16)\)\.
##### Error at the turn\-transition tokens\.
The geometry analysis \([Appendix M](https://arxiv.org/html/2606.05194#A13)\) identifies the<\|im\_end\|\>toassistantturn transition as the locus where temporal preference geometry becomes linearly separable \(Figure[M\.5](https://arxiv.org/html/2606.05194#A13.F5)\)\. Figure[R\.2](https://arxiv.org/html/2606.05194#A18.F2)tests whether error state follows the same pattern\.
Figure R\.2:Cumulative error decodability at three token positions\. All converge to\>\>93% by layer 19\. The turn\-transition tokens \(<\|im\_end\|\>andassistant\), where[Appendix M](https://arxiv.org/html/2606.05194#A13)shows temporal preference geometry crystallizing, carry error state with comparable fidelity to the last token\.Error decodability at the turn\-transition tokens matches the last token from layer 19 onward, with theassistanttoken slightly outperforming the last token at layers 24–31 \(∼\{\\sim\}95% vs\.∼\{\\sim\}93%\)\. This convergence suggests that the turn\-transition computation, the same computation that transforms off\-policy context into on\-policy generation for temporal preference, also integrates reasoning reliability before generation begins\.
### R\.2Do error and temporal preference share representational structure?
Co\-localization does not entail shared structure: two variables can occupy the same layers in orthogonal subspaces\. We test this directly by training probes for each variable in the shared PCA\-50 space and comparing the resulting weight vectors\.
##### Cross\-probing protocol\.
For each of the 15 key layers, we train a binary error probe \(cumulative errors\>0\>0vs\.=0=0;n=3,550n=3\{,\}550\) and a binary temporal probe \(immediate vs\. long\-term;n=1,100n=1\{,\}100\), both in the shared PCA\-50 space\. We compute cosine similarity between the normalized weight vectors and test significance with a 500\-permutation null \(shuffle temporal labels, refit, recompute cosine\)\. We also measure cross\-domain transfer: apply the error probe to temporal data \(and vice versa\) and test against 200\-permutation nulls on the target labels\.
Figure R\.3:Cross\-probing in shared PCA\-50 space \(nerr=3,550n\_\{\\text\{err\}\}=3\{,\}550,ntemp=1,100n\_\{\\text\{temp\}\}=1\{,\}100, 500\-permutation nulls\)\.Left:Cosine between probe directions, all values fall within the permutation null \(shaded\), nop<0\.05p<0\.05\.Center:Cross\-domain transfer, temporal\-to\-error is significant at layers 10–25 \(purple stars,p<0\.001p<0\.001\); error\-to\-temporal is weaker and sporadic \(red stars\)\.Right:Within\-domain error accuracy for reference\.LayerCosineCosineppError→\\toTempTemp→\\toError10\+\+0\.0400\.7740\.5000\.728\*\*\*19\+\+0\.0860\.6140\.5000\.593\*\*\*20\+\+0\.0710\.6140\.574\*\*\*0\.648\*\*\*24−\-0\.0230\.8760\.5000\.626\*\*\*25\+\+0\.0460\.7200\.5000\.574\*\*\*31−\-0\.0520\.7120\.5000\.500Table R\.2:Cross\-probe results at selected layers\. No cosine reaches significance\. Temporal\-to\-error transfer peaks at layer 10 \(0\.728\) and remains above chance through layer 25; error\-to\-temporal transfer is at chance at most layers\. \*\*\* indicatesp<0\.001p<0\.001against the permutation null on target labels\.
##### Interpretation: orthogonal directions, partial non\-linear overlap\.
The cosine null result \(allp\>0\.5p\>0\.5\) establishes that the linear separating hyperplanes for error and temporal preference are perpendicular in the shared activation space\. However, the temporal\-to\-error transfer above chance at layers 10–25 \(p<0\.001p<0\.001\) shows that the temporal probe’s projection of the data partially predicts error status even though the two probe*directions*are orthogonal\. This combination, perpendicular hyperplanes paired with above\-chance transfer, indicates that the two variables share a*non\-linear*subspace: their representations overlap on the activation manifold but not along any single linear axis\.
This is consistent with the non\-linear time\-horizon geometry documented in[Appendix M](https://arxiv.org/html/2606.05194#A13): if both error state and temporal preference occupy curved manifolds in the same region of activation space, their optimal linear separating hyperplanes can be orthogonal even as the manifolds themselves intersect\. The asymmetry of the transfer \(temporal\-to\-error stronger than error\-to\-temporal\) suggests that the temporal preference representation, which captures broad context evaluation \(“strategic vs\. tactical” orientation;[Appendix E](https://arxiv.org/html/2606.05194#A5)\), carries some error\-relevant information as a byproduct, while the error direction \(a narrower signal about chain corruption\) does not carry temporal scope information\.
### R\.3Key findings
1. 1\.Error state co\-localizes with temporal preference\.Cumulative error count is decodable at 95\.3% \(p<0\.001p<0\.001\) with a plateau spanning layers 19–31, the same region identified by attribution patching\. Error is decodable at the turn\-transition tokens where temporal preference geometry crystallizes \([Appendix M](https://arxiv.org/html/2606.05194#A13)\)\. This suggests the subgraph functions as a general meta\-cognitive region, not a temporal\-specific circuit\.
2. 2\.Chain\-level error is far more decodable than local error identity\.The 21pp gap \(95% cumulative vs\. 74% corruption status\) mirrors the main paper’s finding that global properties \(time horizon\) are more structured than local behavioral outputs\.
3. 3\.Error and temporal preference occupy orthogonal linear directions\.No cosine between probe weight vectors reachesp<0\.05p<0\.05at any layer\. The two variables do not compete for the same linear subspace within the subgraph\.
4. 4\.Asymmetric non\-linear overlap exists\.The temporal probe transfers to the error task above chance at layers 10–25 \(p<0\.001p<0\.001\), but the error probe does not transfer to temporal preference\. The two variables share curved manifold structure but not a linear direction, consistent with the non\-linear geometry in[Appendix M](https://arxiv.org/html/2606.05194#A13)\.
5. 5\.The gap between representation and behavior generalizes\.Error state is internally encoded at 95% accuracy but barely affects output confidence, paralleling the temporal preference gap between representation and behavior documented in Appendices[Appendix O](https://arxiv.org/html/2606.05194#A15)and[Appendix P](https://arxiv.org/html/2606.05194#A16)\. The gap is architectural, not task\-specific\.
6. 6\.Two\-axis steering is feasible\.The orthogonality of probe directions means a temporal preference steering vector \([Appendix S](https://arxiv.org/html/2606.05194#A19)\) should not perturb error sensitivity\. An error\-awareness vector at layers 24–25 could complement temporal steering, enabling two\-axis control with minimal cross\-interference\.
7. 7\.Connection to the steering results\.The probing–steering dissociation observed in[Appendix S](https://arxiv.org/html/2606.05194#A19)\(best probing at L26 vs\. best steering at L19–22\) may extend to error: the layers where error is most decodable \(L24–25\) need not be the layers where error\-state interventions are most effective\. Testing this prediction via error\-state CAA is a natural next step\.
##### Limitations\.
These results establish correlational decodability, not causal necessity\. Activation patching of error\-state representations \(analogous to Appendices[Appendix J](https://arxiv.org/html/2606.05194#A10)and[Appendix K](https://arxiv.org/html/2606.05194#A11)\) would be needed to confirm that the identified representations causally drive downstream behavior\. Our error injection uses synthetic perturbations in math reasoning chains, which may not fully reflect the distribution of errors arising during unconstrained generation\. Generalization to other reasoning domains and to models beyondQwen3\-4B\-Instruct\-2507remains to be tested\.
Part 3: Could we control temporal preference?
- •[S](https://arxiv.org/html/2606.05194#A19)\.Contrastive CAA steering
## Appendix Appendix SContrastive steering results
Parts 1 and 2 established where temporal preference lives \(layers 17–35;[Appendix L](https://arxiv.org/html/2606.05194#A12)\) and what it looks like \(an ordinal horizon that transforms into a binary preference at the turn boundary;[Appendix M](https://arxiv.org/html/2606.05194#A13)\)\. The behavioral analysis showed that the resulting preferences are unstable and inconsistent \([Appendix O](https://arxiv.org/html/2606.05194#A15),[Appendix P](https://arxiv.org/html/2606.05194#A16)\)\. Here we ask the intervention question: can we*control*temporal preference by directly modifying the representations we identified?
We construct a CAA steering vector from the probe direction at layer 26 \([Appendix G](https://arxiv.org/html/2606.05194#A7)\) and inject it at candidate layers during inference \(methodology in[Appendix AB](https://arxiv.org/html/2606.05194#A28)\)\.
### S\.1Forced\-Choice Behavioral Sweep
#### S\.1\.1Layer×\\timesAlpha Sweep
The probe’s best layer \(26\) is not necessarily the best steering layer\. We swept 9 layers \(19–27\)×\\times5 alpha values \(1, 2, 5, 10, 20\) = 45 configurations\.
Table S\.1:Forced\-choice scoreS\(α,l\)S\(\\alpha,l\)across layers and steering coefficients\. Baseline \(no steering\):S=0\.1724S=0\.1724\. Layers 19–22 form the behavioral sweet spot, with a sharp drop at layer 23\.α\\alphaL19L20L21L22L23L24L25L26L2710\.200\.200\.200\.200\.190\.190\.190\.180\.1820\.230\.230\.230\.220\.210\.210\.200\.200\.1950\.320\.310\.320\.300\.270\.260\.240\.230\.22100\.460\.450\.460\.420\.360\.340\.310\.290\.26200\.720\.690\.720\.680\.560\.510\.460\.400\.35Figure S\.1:Heatmap of forced\-choice scoreS\(α,l\)S\(\\alpha,l\)across layers \(19–27\) and steering coefficients \(α=1\\alpha=1–2020\)\. Layers 19–22 form the behavioral sweet spot, with effectiveness dropping sharply at layer 23\.##### Probing–steering dissociation\.
Layer 26 is optimal for*reading*temporal orientation \(99\.2% probe accuracy\) but not for*writing*it\. Layers 19–22 are the effective steering layers, 4–7 layers earlier than the best probe layer\. This dissociation is consistent with a functional asymmetry: upper layers consolidate a high\-fidelity*readout*of the temporal concept, while mid\-network layers are where causal interventions most effectively redirect the model’s downstream computation\. Similar probing\-vs\-intervention gaps have been observed in other domains\[[44](https://arxiv.org/html/2606.05194#bib.bib44)\]\.
#### S\.1\.2Extended Alpha Sweep
Following the initial sweep, we extended the alpha range for the most promising layers \(19–25\) withα∈\{20,30,40,50\}\\alpha\\in\\\{20,30,40,50\\\}\.
Table S\.2:Extended alpha sweep\. Best configuration: layer 22 withα=50\\alpha\\\!=\\\!50, achievingS=1\.3944S=1\.3944\(baseline: 0\.1724, lift: \+1\.22\)\. This corresponds to approximately3\.4×3\.4\\timeshigher relative odds for the long\-term over the short\-term completion under the forced\-choice metric \(exp\(1\.22\)≈3\.39\\exp\(1\.22\)\\approx 3\.39\)\. BecauseSSis a mean log\-probability difference, this is an odds\-ratio change rather than an absolute multiplier onP\(long\)P\(\\text\{long\}\)\.α\\alphaL19L20L21L22L23L24L25200\.720\.690\.720\.680\.560\.510\.46300\.940\.910\.970\.930\.760\.680\.60401\.131\.101\.201\.170\.960\.840\.74501\.301\.271\.391\.391\.141\.010\.87Figure S\.2:Extended alpha sweep heatmap for layers 19–25 withα∈\{20,30,40,50\}\\alpha\\in\\\{20,30,40,50\\\}\. Score increases monotonically withα\\alpha; best configuration is layer 22 atα=50\\alpha\\\!=\\\!50\.The score increases monotonically withα\\alphaacross all layers, with layers 19–22 consistently outperforming later layers\. The optimal configuration \(layer 22,α=50\\alpha=50\) achieves a score of 1\.3944, representing a lift of \+1\.22 over the unsteered baseline of 0\.1724\.
### S\.2Open\-Ended Generation Evaluation
The forced\-choice metric measures whether the model’s token probabilities shift in the correct direction\. To verify that this translates into qualitative behavioral change, we evaluate steering on 13 open\-ended neutral prompts \(e\.g\., “You are advising a team on how to handle a major organizational challenge\. What should be the main focus?”\)\.
#### S\.2\.1Experimental Setup
For each configuration, we generate text withdo\_sample=Falseandmax\_new\_tokens=90\. We test layers\{19,20,21,22,23,26\}\\\{19,20,21,22,23,26\\\}acrossα∈\{25,40,50\}\\alpha\\in\\\{25,40,50\\\}, applying the steering vector in both directions: positiveα\\alpha\(toward long\-term\) and negativeα\\alpha\(toward short\-term\)\.
All generated responses were scored byClaude Sonnet 4\.6on a\[−10,\+10\]\[\-10,\+10\]temporal orientation scale, where−10\-10denotes clearly short\-term thinking and\+10\+10denotes clearly long\-term thinking\. We note that this LLM\-as\-judge evaluation is not validated against human ratings; the scores should be interpreted as a proxy for directional shift rather than a calibrated measure of temporal orientation\.
#### S\.2\.2Results
Configurationα=25\\alpha=25α=40\\alpha=40α=50\\alpha=50L19→\\rightarrowlong\-term\+1\.1\+1\.2\+2\.3L19→\\rightarrowshort\-term−\-1\.7−\-4\.7−\-5\.7L20→\\rightarrowlong\-term\+1\.8\+2\.0\+2\.4L20→\\rightarrowshort\-term−\-1\.9−\-3\.6−\-3\.4L21→\\rightarrowlong\-term\+1\.8\+2\.2\+2\.7L21→\\rightarrowshort\-term−\-2\.0−\-3\.5−\-2\.9L22→\\rightarrowlong\-term\+1\.2\+2\.4\+2\.2L22→\\rightarrowshort\-term−\-1\.2−\-2\.8−\-4\.0L23→\\rightarrowlong\-term\+1\.2\+1\.6\+1\.7L23→\\rightarrowshort\-term−\-0\.8−\-1\.0−\-2\.4L26→\\rightarrowlong\-term\+1\.2\+1\.0\+1\.0L26→\\rightarrowshort\-term−\-0\.5−\-1\.0−\-0\.1Table S\.3:Mean shift from baseline on the\[−10,\+10\]\[\-10,\+10\]temporal orientation scale for open\-ended generation\. Positive values indicate a shift toward long\-term framing; negative values indicate a shift toward short\-term framing\. Each entry averages over 13 prompts\.Figure S\.3:Shift from baseline on the\[−10,\+10\]\[\-10,\+10\]temporal orientation scale for open\-ended generation across layers andα\\alphavalues\. Both long\-term \(positive\) and short\-term \(negative\) steering produce consistent directional shifts\.##### Key observations\.
1. 1\.Bidirectional steering\.Both positive and negativeα\\alphaproduce consistent shifts in the expected direction, confirming that the CAA vector captures a genuine temporal orientation axis\.
2. 2\.Qualitative framing shifts\.Atα∈\[40,50\]\\alpha\\in\[40,50\], long\-term steered outputs adopt strategic framing \(e\.g\., “resilient,” “future\-ready,” “systemic redesign”\), while short\-term steered outputs adopt triage framing \(e\.g\., “remain calm,” “prioritize urgency,” “structured immediate steps”\)\.
3. 3\.Coherence boundary\.Pilot runs at\|α\|=60\|\\alpha\|=60showed output incoherence, so we capped the sweep at\|α\|=50\|\\alpha\|=50; beyond that the intervention appears to push the model too far from its natural distribution\.
4. 4\.Asymmetry\.The long\-term direction produces cleaner shifts than the short\-term direction at the same\|α\|\|\\alpha\|, suggesting that the model’s latent distribution may be slightly closer to the long\-term end of the temporal axis\.
5. 5\.Layer effects mirror forced\-choice results\.Layers 19–22 produce the strongest open\-ended shifts, and layer 26 produces only weak effects despite having the highest probe accuracy\. This reinforces the probing–steering dissociation documented in Section[S\.1](https://arxiv.org/html/2606.05194#A19.SS1)\.
#### S\.2\.3Qualitative Example
We illustrate the steering effect with a representative example\.
Prompt:You are advising a team on how to handle a major organizational challenge\. What should be the main focus?
Baseline\(score: 2\): The response focuses on present\-tense communication and immediate trust\-building without strong temporal language in either direction\. The main focus is on clear communication, transparency, and employee engagement\.
Long\-term steered\(α=\+50\\alpha=\+50, score: 6\): The response centers on long\-term success through resilience and shared vision, framing organizational challenges in terms of sustained adaptive capacity\. The main focus is on resilience through shared purpose, adaptive thinking, and inclusive leadership\.
Short\-term steered\(α=−50\\alpha=\-50, score:−\-3\): The response leads with immediate and clear communication and emphasizes knowing what the deadline is, orienting team management around near\-term operational urgency\.
### S\.3Discussion
##### Probing≠\\neqsteering\.
The central methodological finding is the dissociation between the optimal probing layer \(26\) and the optimal steering layers \(19–22\)\. This dissociation has implications for the broader interpretability literature: high probe accuracy at a layer does not imply that the same layer is the appropriate target for causal intervention\. The readout of a concept and the point at which that concept can be effectively modified may be separated by several layers, reflecting distinct functional roles in the transformer’s computation\[[100](https://arxiv.org/html/2606.05194#bib.bib100)\]\.
##### Implicit vectors generalize\.
Using the implicit dataset \(which contains no surface temporal vocabulary\) to construct the CAA vector ensures that the steering direction captures semantic temporal reasoning rather than lexical artifacts\. The cross\-dataset probe generalization \(Section[G\.4](https://arxiv.org/html/2606.05194#A7.SS4)\) confirms that the implicit direction aligns with the explicit temporal axis, and the forced\-choice evaluation on explicit prompts \(Section[S\.1](https://arxiv.org/html/2606.05194#A19.SS1)\) demonstrates that this vector effectively steers behavior on prompts with overt temporal cues\.
##### Connection to subgraph localization\.
The behavioral sweet spot at layers 19–22 aligns with the mid\-network components identified by the EAP\-IG attribution analysis \([Appendix H](https://arxiv.org/html/2606.05194#A8)\) and the activation patching experiments \([Appendix J](https://arxiv.org/html/2606.05194#A10)\)\. This convergence across independent methodologies \(probing, CAA steering, attribution patching, and activation patching\) provides strong evidence that the temporal preference mechanism is localized to a consistent set of mid\-to\-upper layers\.
##### Relation to the representational geometry\.
The PCA analysis \(Section[G\.3](https://arxiv.org/html/2606.05194#A7.SS3)\) shows that the temporal direction in the implicit dataset is not captured by the top principal components\. This is consistent with the non\-linear manifold structure reported in[Appendix M](https://arxiv.org/html/2606.05194#A13), where time horizon is encoded in a curved subspace\. The CAA vector, derived from the linear probe direction, provides a first\-order approximation to steering along this manifold\. The monotonic increase in steering score withα\\alpha\(Table[S\.2](https://arxiv.org/html/2606.05194#A19.T2)\) suggests that this linear approximation remains effective within the tested range, though the output\-quality degradation at\|α\|=60\|\\alpha\|=60may indicate the intervention exceeding the locally linear regime\.
##### LLM\-as\-Judge Evaluation Criteria\.
To quantify the qualitative shifts in our open\-ended generation experiments, we used theClaude Sonnet 4\.6API as an independent evaluator\. Each generated response was individually processed by the API and assigned a score on a\[−10,\+10\]\[\-10,\+10\]scale, where−10\-10represents an extreme short\-term focus and\+10\+10represents an extreme long\-term focus\. The model was prompted to evaluate responses by strictly adhering to predefined grading criteria\. Specifically, the evaluator analyzed the text for the presence and frequency of explicit temporal keywords \(e\.g\., immediate triage versus systemic redesign\) and weighed semantic details, structural planning, and thematic biases that explicitly skewed the generation toward a specific temporal horizon\.
Part 4: Extended methodologies
- •[T](https://arxiv.org/html/2606.05194#A20)\.Notation
- •[U](https://arxiv.org/html/2606.05194#A21)\.Contrastive probing methods
- •[V](https://arxiv.org/html/2606.05194#A22)\.Attributional contrastive methods
- •[W](https://arxiv.org/html/2606.05194#A23)\.Causal parametric methods
- •[X](https://arxiv.org/html/2606.05194#A24)\.Causal contrastive methods
- •[Y](https://arxiv.org/html/2606.05194#A25)\.Parametric geometry methods
- •[Z](https://arxiv.org/html/2606.05194#A26)\.Behavioral discounting methods
- •[AA](https://arxiv.org/html/2606.05194#A27)\.Behavioral coherence methods
- •[AB](https://arxiv.org/html/2606.05194#A28)\.Contrastive steering methods
- •[AC](https://arxiv.org/html/2606.05194#A29)\.Worked case study: highly\-formatted pair
## Appendix Appendix TNotation and key concepts
The following terms and abbreviations are used throughout the appendices\.
TermDefinitionSubgraphModel components \(attention heads, MLP neurons\) whose ablation or patching shifts temporal preference\.Temporal preferenceThe model’s tendency to favor short\-term vs\. long\-term options in a forced\-choice setting\.Time horizonAn explicit temporal constraint \(e\.g\., “1 year”\) given in the prompt; ranges from seconds to centuries\.On\- vs\. off\-policy*On\-policy*: activations from the model’s own generation\.*Off\-policy*: activations read from a forced context \(user turn\)\.Contrastive pairMatched clean/corrupted prompts that differ in temporal framing; used for both EAP\-IG and activation patching\.EAP\-IGEdge Attribution Patching with Integrated Gradients\[[43](https://arxiv.org/html/2606.05194#bib.bib43)\]; gradient\-based attribution that approximates causal patching\.Activation patchingReplacing a component’s activations with counterfactual values to measure causal effect on a downstream metric\.Recovery / DisruptionNormalized\[0,1\]\[0,1\]metrics for denoising and noising patching, respectively; 0 = no effect, 1 = full effect\.Probing layerResidual\-stream layer at which a linear classifier best separates short\- vs\. long\-term orientation \(layer 26 in this work\)\.Steering layerLayer at which a CAA vector\[[100](https://arxiv.org/html/2606.05194#bib.bib100)\]most reliably shifts behavior \(layers 19–22 in this work\)\.CAAContrastive Activation Addition: mean activation difference between long\-term and short\-term choices, used as a steering vector\.Decision boundaryBinary search over delayed\-reward magnitudes to locate per\-item indifference, used to fit hyperbolic discount ratekk\.MCQ\-27Kirby Monetary Choice Questionnaire\[[57](https://arxiv.org/html/2606.05194#bib.bib57)\]: 27\-item instrument for estimating temporal discount rates\.
## Appendix Appendix UContrastive linear probing methodology
We train logistic regression probes\[[74](https://arxiv.org/html/2606.05194#bib.bib74),[56](https://arxiv.org/html/2606.05194#bib.bib56)\]on residual\-stream activations to determine*where*the model linearly encodes the distinction between short\-term and long\-term temporal orientation\. Probing results are presented in[Appendix G](https://arxiv.org/html/2606.05194#A7)\.
### U\.1Activation Extraction
For each prompt, we concatenate the question and the choice text, apply theQwen3chat template, and extract residual\-stream activations at every layer\. Because the chat template wraps the user turn as
<\|im\_start\|\>user\\n\{question \+ choice\}<\|im\_end\|\>\\n<\|im\_start\|\>assistant\\n
the token at position−1\-1is the trailing newline afterassistant, a fixed token that is identical across all prompts and carries only whatever signal attention has propagated into that generic position\. We instead locate the last<\|im\_end\|\>token in the sequence \(which closes the user turn\) and extract at positionim\_end−1\\texttt\{im\\\_end\}\-1, corresponding to the final token of the actual choice text\. This position directly encodes the semantic content of the choice\.
##### Impact of the token\-position correction\.
The correction produced a qualitative change in both probe accuracy and downstream steering vector quality:
MetricBefore \(trailing \\n\)After \(im\_end−\-1\)Extraction token\\n\(trailing newline\)Last choice tokenCAA vectorℓ2\\ell\_\{2\}norm2\.6230\.30Probe accuracy \(best layer\)∼\\sim93%99\.2%Table U\.1:Effect of correcting the extraction token position\. The previous vector was essentially normalized noise; the corrected extraction yields a∼\\sim10×\\timesstronger CAA vector\.
### U\.2Probe Training Protocol
We train oneLogisticRegression\(C=0\.1\)\(C\\\!=\\\!0\.1\)probe per layer onDimplicitD\_\{\\mathrm\{implicit\}\}with the following methodological controls:
1. 1\.Pair\-level train/test split\.The split operates on pair indices rather than individual rows\. Both the immediate and long\-term activations from a given pair always land in the same fold\. Without this, the probe can exploit shared question text as a shortcut, and the test set is not truly held out\. We use an 80/20 split with a fixed random seed\.
2. 2\.StandardScaler normalization\.The residual stream has 2,560 dimensions with varying variances\.LogisticRegressionwithℓ2\\ell\_\{2\}regularization penalizes large weights uniformly, so high\-variance dimensions dominate the penalty without scaling\. We fit aStandardScaleron the training fold and apply it to both train and test\. Critically, the scaler is persisted to disk alongside each probe and re\-applied during cross\-dataset evaluation and when extracting the probe’s coefficient vector for use as the CAA steering direction\.
Activations are stored as tensors of shape\[nprompts,nlayers,dmodel\]\[n\_\{\\mathrm\{prompts\}\},n\_\{\\mathrm\{layers\}\},d\_\{\\mathrm\{model\}\}\]=\[600,36,2560\]\[600,36,2560\]for the implicit dataset\.
## Appendix Appendix VAttributional contrastive methodology
We identify this subgraph in two stages: first, we restrict the candidate node set using Edge Attribution Patching with Integrated Gradients \(EAP\-IG\); second, we score and prune edges between these nodes to recover a sparse functional subgraph\. Unlike prior approaches that attribute to logit differences, we compute attribution with respect to individual option logits, yielding concept\-specific attribution scores that disentangle the contributions of nodes to competing temporal evaluations\.
### V\.1Edge Attribution Patching\-Integrated Gradients
To localize the internal computations associated with temporal preference, we use Edge Attribution Patching with Integrated Gradients \(EAP\-IG\)\. EAP\-IG operates on matched clean and corrupted prompts and assigns attribution scores to internal components based on their contribution to the model’s preference for one response token over another\. It can be viewed as a computationally efficient approximation to activation patching\.
EAP\-IG can be implemented in two ways: by interpolating activations at each node, or by interpolating only the input embeddings\. The latter is significantly more efficient and provides a practical method for estimating edge importance\. However, as this approach compounds two approximations, we do not use it to precisely rank components; instead, we use it to restrict the search space by filtering out nodes with low attribution scores\.
Concretely, we interpolate between corrupted and clean inputs in embedding space and integrate gradients along the resulting path, rather than relying on a single local gradient estimate\. This yields attribution scoressA\(x,i,t\)s^\{A\}\(x,i,t\)for each componentiiat token positionttfor metricAAon promptxx\.
### V\.2Notation
In the paper, the attribution score for a variantvvis denoted assA\(x,i,t\)s^\{A\}\(x,i,t\)\.
sA\(x,i,t\)=\(zi,t−zi,t′\)∫α=01∂LA\(x′\+α\(x−x′\)\)∂zi,t≈\(zi,t−zi,t′\)1m∑k=1m∂LA\(x′\+km\(x−x′\)\)∂zi,ts^\{A\}\(x,i,t\)=\(z\_\{i,t\}\-z^\{\\prime\}\_\{i,t\}\)\\int\_\{\\alpha=0\}^\{1\}\\frac\{\\partial L\_\{A\}\(x^\{\\prime\}\+\\alpha\(x\-x^\{\\prime\}\)\)\}\{\\partial z\_\{i,t\}\}\\approx\(z\_\{i,t\}\-z^\{\\prime\}\_\{i,t\}\)\\frac\{1\}\{m\}\\sum\_\{k=1\}^\{m\}\\frac\{\\partial L\_\{A\}\(x^\{\\prime\}\+\\frac\{k\}\{m\}\(x\-x^\{\\prime\}\)\)\}\{\\partial z\_\{i,t\}\}\(V\.1\)
#### V\.2\.1Metric Normalized Attribution Scores
The attribution scores are scaled byΔLA=LA\(z\)−LA\(z′\)\\Delta L\_\{A\}=L\_\{A\}\(z\)\-L\_\{A\}\(z^\{\\prime\}\)so they can be aggregated across datasets and semantically equivalent metric functions\. In this paper,s¯A\(x,i,t\)\\bar\{s\}^\{A\}\(x,i,t\)denotes the scaled attribution scores\. WhenLA\(z\)≈LA\(z′\)L\_\{A\}\(z\)\\approx L\_\{A\}\(z^\{\\prime\}\), it indicates that the model either does not distinguish between the clean and corrupted cases or has nearly the same preference for both options; such cases on average constitute∼2%\\sim 2\\%of the total dataset and are dropped from further analysis\.
### V\.3Variations for Bias Control
#### V\.3\.1Positional Bias Control
We control for positional bias by evaluating each question–answer pair under both possible option orderings\. In one condition, the short\-horizon response precedes the long\-horizon response; in the other, the order is reversed\. This counterbalancing prevents temporal preference from being confounded with a general tendency to favor a particular position \(e\.g\., the first option\) or a fixed association between labels and positions\.
In the input construction pipeline, this is implemented by generating two matched prompt sets from the same underlying examples: a canonical ordering \(question, short\-horizon option, long\-horizon option\) and a mirrored ordering \(question, long\-horizon option, short\-horizon option\)\. The experiment loop evaluates both orderings as separate conditions \(short\_firstandlong\_first\) under otherwise identical settings\. Consequently, any temporal\-scope effect that is consistent across both conditions is unlikely to be driven by positional bias alone\.
#### V\.3\.2Lexical Bias Control
To mitigate the possibility that results are driven by the lexical identity of response labels rather than temporal content, we repeat all experiments under seven matched response\-label schemes\. These schemes use uppercase letters \(\(A\)/\(B\)\), lowercase letters \(\(a\)/\(b\)\), Arabic numerals \(\(1\)/\(2\)\), Roman numerals \(\(i\)/\(ii\)\), number words \(\(One\)/\(Two\)\), alternative letters \(\(X\)/\(Y\)\), and a non\-alphanumeric symbol pair \(\(⚫\)/\(◼\)\)\.
Across these runs, the dataset, prompt template, model, batch size, inference settings, and scoring metric are held fixed\. Only the surface form of the response labels and the corresponding instruction specifying the target output token are varied\. This isolates lexical biases associated with particular label tokens, such as pretrained preferences forA/Bor1/2\.
If an effect persists across all label variants, it is unlikely to be attributable to any specific output token and instead reflects the model’s sensitivity to the underlying short\- versus long\-horizon distinction\. We denote each label variant asDvD^\{v\}\.
### V\.4Experimental Setup
#### V\.4\.1System Prompt
The system prompt is designed to constrain the model’s output format and suppress the inclusion of explicit reasoning in its responses\.
To ensure that the phrasing of the system prompt does not confound component attribution scores, all token positions corresponding to the system prompt are excluded during position\-wise aggregation\.555In theQwen3chat template, the end of a prompt is marked by the<\|im\_end\|\>token\. Accordingly,nsysn\_\{\\text\{sys\}\}is defined as the token length of the system prompt plus one\.
s¯tA\(x,i\)=1\(ntotal−nsys\)∑t=nsysntotals¯A\(x,i,t\)\\bar\{s\}^\{A\}\_\{t\}\(x,i\)=\\frac\{1\}\{\(n\_\{\\text\{total\}\}\-n\_\{\\text\{sys\}\}\)\}\\sum\_\{t=n\_\{\\text\{sys\}\}\}^\{n\_\{\\text\{total\}\}\}\\bar\{s\}^\{A\}\(x,i,t\)\(V\.2\)
#### V\.4\.2Prompt Syntax
Figure V\.1:An example of base and swapped prompts, along with their corrupted counterparts used for EAP\-IG attribution\. Corruption corresponds to flipping option semantics while preserving surface form\.Each prompt draws from the temporal\-scope datasets described in[Appendix E](https://arxiv.org/html/2606.05194#A5)and presents a scenario along with two plausible courses of action: one emphasizing short\-term rewards and the other emphasizing long\-term rewards\. The system prompt instructs the model to respond by selecting the label \(e\.g\.,AorB\) corresponding to its preferred option\.
To construct a corrupted variant, the option labels are swapped while preserving the textual order of the candidate responses \(see Figure[V\.1](https://arxiv.org/html/2606.05194#A22.F1)\)\. This manipulation isolates the effect of label assignment from the semantic content of the options\.
To control for positional bias, we additionally evaluate the prompt under a flipped ordering of the options\. LetstA\(x,i\)s^\{A\}\_\{t\}\(x,i\)denote the attribution score at token positionttfor componentiiunder the canonical ordering\. LetstA∗\(x,i\)s^\{A\*\}\_\{t\}\(x,i\)denote the corresponding attribution score under the flipped ordering\.
#### V\.4\.3Metric Function
We use raw logit values as the attribution metric rather than logit differences, as they provide a more fine\-grained characterization of component behavior\. In particular, raw logits allow us to distinguish between components that actively promote a target concept and those that exert inhibitory effects\. This formulation also enables attribution with respect to semantically meaningful concepts \(e\.g\., long\-term vs\. short\-term orientation\), rather than relative preferences alone\.
LetLT\\mathrm\{LT\}andST\\mathrm\{ST\}denote the long\-term and short\-term concepts, respectively\. We define concept\-aligned attribution scores by symmetrizing over label assignments and option orderings:
s¯tST\(x,i\)=12\(s¯tA\(x,i\)\+s¯tB∗\(x,i\)\),s¯tLT\(x,i\)=12\(s¯tB\(x,i\)\+s¯tA∗\(x,i\)\)\.\\bar\{s\}\_\{t\}^\{\\mathrm\{ST\}\}\(x,i\)=\\frac\{1\}\{2\}\\left\(\\bar\{s\}\_\{t\}^\{A\}\(x,i\)\+\\bar\{s\}\_\{t\}^\{B\*\}\(x,i\)\\right\),\\quad\\bar\{s\}\_\{t\}^\{\\mathrm\{LT\}\}\(x,i\)=\\frac\{1\}\{2\}\\left\(\\bar\{s\}\_\{t\}^\{B\}\(x,i\)\+\\bar\{s\}\_\{t\}^\{A\*\}\(x,i\)\\right\)\.\(V\.3\)
### V\.5Component Attribution Calculation
We compute component\-level attribution scores for each time\-horizon conceptc∈\{LT,ST\}c\\in\\\{\\mathrm\{LT\},\\mathrm\{ST\}\\\}via a two\-stage aggregation procedure over examples and prompt variants\.
##### Within\-variant aggregation\.
For each variantv∈𝒱v\\in\\mathcal\{V\}, we estimate the expected attribution score by averaging over a finite sample of inputsDv=\{x1,…,xNv\}D^\{v\}=\\\{x\_\{1\},\\dots,x\_\{N\_\{v\}\}\\\}:
s¯tc\(Dv,i\)=1Nv∑n=1Nvs¯tc\(xn,i\)\.\\bar\{s\}^\{c\}\_\{t\}\(D^\{v\},i\)=\\frac\{1\}\{N\_\{v\}\}\\sum\_\{n=1\}^\{N\_\{v\}\}\\bar\{s\}^\{c\}\_\{t\}\(x\_\{n\},i\)\.\(V\.4\)This estimator is well\-defined because attribution scores are normalized by the logit differenceΔL\\Delta L, making them comparable across inputs\.
##### Across\-variant aggregation\.
We then aggregate across a finite set of variants𝒱\\mathcal\{V\}using a uniform weighting:
s¯tc\(D,i\)=1\|𝒱\|∑v∈𝒱s¯tc\(Dv,i\)\.\\bar\{s\}^\{c\}\_\{t\}\(D,i\)=\\frac\{1\}\{\|\\mathcal\{V\}\|\}\\sum\_\{v\\in\\mathcal\{V\}\}\\bar\{s\}^\{c\}\_\{t\}\(D^\{v\},i\)\.\(V\.5\)
##### Assumptions\.
This procedure assumes that \(i\) examples within each dataset variant are independent and identically distributed samples from an underlying distribution, and \(ii\) variants are treated as equally informative perturbations, justifying uniform averaging across𝒱\\mathcal\{V\}\. In practice, both expectations are approximated by finite\-sample means as defined above\.
## Appendix Appendix WCausal parametric methodology
Activation patching results are presented in[Appendix J](https://arxiv.org/html/2606.05194#A10)\. Here we describe the experimental setup\.
### W\.1Overview
The parametric pipeline uses highly\-formatted prompts with explicit time horizons to perform activation patching\[[44](https://arxiv.org/html/2606.05194#bib.bib44)\]\. Unlike the contrastive pipeline \([Appendix V](https://arxiv.org/html/2606.05194#A22)\), which uses gradient\-based attribution as an efficient approximation, the parametric pipeline directly measures causal effect by replacing component activations with counterfactual values\.
The pipeline operates on*contrastive pairs*, matched clean and corrupted trajectories that differ in their temporal framing\. Prompt construction and parametric variation are described in[E\.2](https://arxiv.org/html/2606.05194#A5.SS2)\. A three\-stage evaluation proceeds from sanity check to layer sweep to position sweep, each with configurable stride sizes that enable efficient coarse\-to\-fine analysis\.
### W\.2Activation Patching Protocol
For each contrastive pair, we perform bothdenoisingandnoisinginterventions:
##### Denoising\.
The model runs on the corrupted prompt while clean activations are injected at specified layers and positions\. This measures*recovery*: how much the intervention restores clean behavior\.
##### Noising\.
The model runs on the clean prompt while corrupted activations are injected\. This measures*disruption*: how much the intervention degrades clean behavior\.
Both metrics are normalized to\[0,1\]\[0,1\]:
Recovery=yintervened−ycorruptedyclean−ycorrupted\\displaystyle=\\frac\{y\_\{\\text\{intervened\}\}\-y\_\{\\text\{corrupted\}\}\}\{y\_\{\\text\{clean\}\}\-y\_\{\\text\{corrupted\}\}\}\(W\.1\)Disruption=yclean−yintervenedyclean−ycorrupted\\displaystyle=\\frac\{y\_\{\\text\{clean\}\}\-y\_\{\\text\{intervened\}\}\}\{y\_\{\\text\{clean\}\}\-y\_\{\\text\{corrupted\}\}\}\(W\.2\)whereyydenotes the model’s logit difference between the two options\. A value of 0 indicates no causal effect; 1 indicates full effect\.
### W\.3Component Types
We patch the following residual\-stream components independently:
- •resid\_pre: Residual stream before attention \(input to the layer\)
- •attn\_out: Attention output
- •resid\_mid: Residual stream after attention, before MLP
- •mlp\_out: MLP output
- •resid\_post: Residual stream after MLP \(output of the layer\)
For residual\-stream components, patching is performed at the*divergent position*, the last token before the model’s choice, because residual propagation makes all\-position patching uninformative\.
### W\.4Position Mapping
Clean and corrupted prompts often differ in token count because different time horizons or reward amounts require different numbers of tokens\. To patch activations at semantically corresponding positions, we use a*piecewise linear interpolation*anchored on the structural markers of the highly\-formatted template \([E\.2](https://arxiv.org/html/2606.05194#A5.SS2)\)\.
The algorithm identifies the token positions of each section marker \(SITUATION,TASK,OBJECTIVE,CONSTRAINT,ACTION,FORMAT\) as well as sub\-markers for option labels, reward amounts, and time values in both the clean and corrupted sequences\. These anchors are sorted and augmented with sequence boundaries to form a set of corresponding position pairs\.
Between consecutive anchors, positions are mapped via linear interpolation: for a source positionppin segment\[asrc,bsrc\]\[a\_\{\\text\{src\}\},b\_\{\\text\{src\}\}\]mapped to\[adst,bdst\]\[a\_\{\\text\{dst\}\},b\_\{\\text\{dst\}\}\], the corresponding destination position is
p′=adst\+p−asrcbsrc−asrc⋅\(bdst−adst\),p^\{\\prime\}=a\_\{\\text\{dst\}\}\+\\frac\{p\-a\_\{\\text\{src\}\}\}\{b\_\{\\text\{src\}\}\-a\_\{\\text\{src\}\}\}\\cdot\(b\_\{\\text\{dst\}\}\-a\_\{\\text\{dst\}\}\),\(W\.3\)clamped to valid token indices\. This ensures that activations from each semantic region \(e\.g\., the constraint field\) are patched into the corresponding region of the other prompt, even when the two prompts have different total lengths\.
### W\.5Sweep Protocol
##### Layer sweep\.
We patch each component across all 36 layers with a stride of 1, measuring recovery and disruption at each layer\. This identifies which layers carry the most causal effect for temporal preference\.
##### Position sweep\.
For the most causally important layers, we sweep across token positions with configurable strides \(1, 5, or 10 tokens\) to identify which token regions are most informative\. The position mapping described above ensures correct alignment when clean and corrupted prompts differ in length\.
## Appendix Appendix XCausal classification methodology
Results are presented in[Appendix K](https://arxiv.org/html/2606.05194#A11)\. Here we describe the experimental setup\.
### X\.1Motivation
We construct an experiment to test whether the components flagged by the preference\-targeted methods \([Appendix G](https://arxiv.org/html/2606.05194#A7),[Appendix H](https://arxiv.org/html/2606.05194#A8),[Appendix J](https://arxiv.org/html/2606.05194#A10)\) also engage in related temporal computations beyond preference valuation\. We apply directional activation patching to IOI\-style classification prompts in which the model must infer whether a goal’s horizon is short\-term or long\-term: a cognitively distinct task that probes the same temporal axis\. The protocol below describes the dataset, prompt structure, and patching pipeline\.
### X\.2Dataset
We use the 160\-pair classification dataset described in[E\.3](https://arxiv.org/html/2606.05194#A5.SS3)\. The four design principles \(token alignment to 34 tokens, semantic overlap within pairs, no explicit temporal keywords, unambiguous horizons\) and the 80% \(160/200\)Qwen3\-4B\-Instruct\-2507accuracy filter that produced 160 surviving pairs are documented there\.
Each dataset sample contains two prompts: a clean prompt expectingshortanswer \(for a short\-term goal\) and a corrupted prompt expectinglonganswer \(for a long\-term goal\); the directional flips \(defined in Sec[X\.3](https://arxiv.org/html/2606.05194#A24.SS3)\) invert the assignment\.
Prompt template
"The goal is to <goal\>\. Is this a <short\-term or long\-term / long\-term or short\-term\> goal? The answer is:"
Table[X\.1](https://arxiv.org/html/2606.05194#A24.T1)shows dataset statistics for 160 surviving pairs after validation\. Table[X\.2](https://arxiv.org/html/2606.05194#A24.T2)shows the same statistics, but grouped by question order\.
VariableCareer/MasteryGrowthAccumulationCount74 \(46%\)53 \(33%\)33 \(20%\)Clean Q, LD \(mean \+/\- st\.d\.\)14\.13 \+/\- 4\.1911\.08 \+/\- 4\.7112\.07 \+/\- 3\.86Corrupted Q, LD \(mean \+/\- st\.d\.\)\-13\.53 \+/\- 4\.88\-10\.60 \+/\- 5\.72\-10\.26 \+/\- 4\.99Table X\.1:Cue types statistics on successful Temporal Classification pairsVariableSLLSCount8080Clean Q, LD \(mean \+/\- st\.d\.\)11\.76 \+/\- 4\.4713\.63 \+/\- 4\.35Corrupted Q, LD \(mean \+/\- st\.d\.\)\-13\.61 \+/\- 5\.26\-10\.16 \+/\- 4\.97Table X\.2:Question order statistics on successful Temporal Classification pairs
### X\.3Patching Protocol
We perform denoising activation patching on theresid\_pre,attn\_out, andmlp\_outhooks at all token positions across all 36 layers ofQwen3\-4B\-Instruct\-2507on 160 prompt pairs\. We patch separately for short→\\tolong and long→\\toshort flips, testing whether the computation components match\. We use three metrics to quantify the effect \(Sec[X\.3](https://arxiv.org/html/2606.05194#A24.SS3.SSS0.Px1)\): logit difference \(LD\) and log\-probability of clean and corrupted answers: log\-P\(clean\)P\(\\text\{clean\}\)and log\-P\(corr\)P\(\\text\{corr\}\), respectively\. All of them are normalized so that0corresponds to the corrupted baseline and±1\\pm 1to full recovery of the clean run’s behavior\. By definition of presented metrics the two denoising rounds for different flips yield the same result as applying both the noising and denoising techniques on either one of flips\. As we cannot assume symmetry betweenshortandlongrepresentations, we treat the flips separately and consider the noising or disruption of one flip as a recovery of the other, rather than as a reflection of the necessity and sufficiency of a single temporal classification circuit\.
We refer to the flip in which the clean answer isshortand the corrupted answer islongas the*short\-clean*flip, and the opposite as the*long\-clean*flip\.
Throughout, we report a patching effect as significant if its 95% confidence interval excludes zero\. Confidence intervals are computed via a non\-parametric pair\-level percentile bootstrap with 10,000 resamples: for each iteration, we draw 160 pair indices with replacement, compute the mean per\-layer effect across resampled pairs, and report the 2\.5th and 97\.5th percentiles of these bootstrap means as the 95% CI bounds\. We do not apply a multiple\-comparisons correction across layers\. Most effects forming the main headline claims of[Appendix K](https://arxiv.org/html/2606.05194#A11)are large in magnitude and would survive standard family\-wise corrections; other subordinate effects discussed there are closer to the per\-layer threshold and should be read with appropriate caution\.
We report layer\-level effects both at the END token position \(Sec\.[K\.3](https://arxiv.org/html/2606.05194#A11.SS3),[K\.4](https://arxiv.org/html/2606.05194#A11.SS4)\) and summed across all 34 token positions \(Sec\.[K\.5](https://arxiv.org/html/2606.05194#A11.SS5)\) specifically for cross\-paradigm comparison\. The later view sums each metric’s per\-position effect across all 34 positions, yielding a single per\-layer score for \(component, flip\) combination\.
##### Metric definitions\.
The formulas for three metrics used in the experiment are presented below\. Letℓc\\ell\_\{c\}andℓp\\ell\_\{p\}denote the logits for the clean and corrupted answers respectively, and let metricsmetricclean\{metric\}^\{\\text\{clean\}\},metriccorr\{metric\}^\{\\text\{corr\}\}, andmetricpatched\{metric\}^\{\\text\{patched\}\}denote the value of the corresponding quantity under the clean run, corrupted run, and patched run\.
Logit difference \(LD\)\.
LDnormalizedpatched=LDpatched−LDcorrLDclean−LDcorr\\text\{LD\}\_\{\\text\{normalized\}\}^\{\\text\{patched\}\}=\\frac\{\\text\{LD\}^\{\\text\{patched\}\}\-\\text\{LD\}^\{\\text\{corr\}\}\}\{\\text\{LD\}^\{\\text\{clean\}\}\-\\text\{LD\}^\{\\text\{corr\}\}\}\(X\.1\)whereLD=ℓc−ℓp\\text\{LD\}=\\ell\_\{c\}\-\\ell\_\{p\}\.
Log\-probability of the clean answer\.
log\-P\(clean\)normalizedpatched=logP\(clean\)patched−logP\(clean\)corrlogP\(clean\)clean−logP\(clean\)corr\\text\{log\-\}P\(\\text\{clean\}\)\_\{\\text\{normalized\}\}^\{\\text\{patched\}\}=\\frac\{\\log P\(\\text\{clean\}\)^\{\\text\{patched\}\}\-\\log P\(\\text\{clean\}\)^\{\\text\{corr\}\}\}\{\\log P\(\\text\{clean\}\)^\{\\text\{clean\}\}\-\\log P\(\\text\{clean\}\)^\{\\text\{corr\}\}\}\(X\.2\)
Log\-probability of the corrupted answer\.
log\-P\(corr\)normalizedpatched=−logP\(corr\)patched−logP\(corr\)corrlogP\(corr\)clean−logP\(corr\)corr\\text\{log\-\}P\(\\text\{corr\}\)\_\{\\text\{normalized\}\}^\{\\text\{patched\}\}=\-\\frac\{\\log P\(\\text\{corr\}\)^\{\\text\{patched\}\}\-\\log P\(\\text\{corr\}\)^\{\\text\{corr\}\}\}\{\\log P\(\\text\{corr\}\)^\{\\text\{clean\}\}\-\\log P\(\\text\{corr\}\)^\{\\text\{corr\}\}\}\(X\.3\)
## Appendix Appendix YParametric geometry methodology
Results are presented in[Appendix M](https://arxiv.org/html/2606.05194#A13)\. Here we describe the activation extraction and PCA analysis pipeline\.
### Y\.1Activation Extraction
We extract activations fromQwen3\-4B\-Instruct\-2507at 15 selected layers:\{0,1,3,12,18,19,20,21,23,24,25,28,31,34,35\}\\\{0,1,3,12,18,19,20,21,23,24,25,28,31,34,35\\\}, spanning early, mid, and late layers\. At each layer, we extract five component types:resid\_pre,attn\_out,resid\_mid,mlp\_out, andresid\_post\.
Activations are extracted at 16 semantic positions within each prompt, identified via the structural markers of the highly\-formatted template \([E\.2](https://arxiv.org/html/2606.05194#A5.SS2)\):
- •Constraint positions:time\_horizon,post\_time\_horizon
- •Label/time/reward positions:left\_label,right\_label,left\_time,right\_time,left\_reward,right\_reward
- •Section tails: last token ofTASK,OPTIONS,OBJECTIVE,ACTION, andFORMATsections
- •Turn boundary:chat\_suffix\(the four tokens<\|im\_end\|\>,\\n,<\|im\_start\|\>,assistant\) andchat\_suffix\_tail\(the\\nafterassistant\)
- •Response positions:response\_choice\(thea\)orb\)token\) andresponse\_choice\_prefix\(thechoosein “I choose:”\)
For multi\-token positions, we extract at each token index separately \(e\.g\.,chat\_suffix\_r0throughchat\_suffix\_r3\)\. This yields up to15×5×16=1,20015\\times 5\\times 16=1\{,\}200unique activation targets per prompt\.
### Y\.2PCA Analysis
For each target \(layer×\\timescomponent×\\timesposition\), we fit scikit\-learnPCAwith up to 10 components on the raw \(unnormalized\) activation vectors across all prompts\. We compute:
- •Explained variance ratiosfor each principal component\. At key positions, PC1 explains 44–71% and PC2 explains 16–30% of variance \(Table[Y\.1](https://arxiv.org/html/2606.05194#A25.T1)\)\.
- •Spearman correlationsbetween each PC projection andlog10\(time\_horizon\)\\log\_\{10\}\(\\text\{time\\\_horizon\}\)to identify which components encode temporal information\. The null \(no\-horizon\) condition is excluded from logarithmic correlation analyses and treated as a separate baseline category in geometry plots\.
All geometry claims in[Appendix M](https://arxiv.org/html/2606.05194#A13)are based on PCA visualizations and variance\-explained ratios\. We do not provide bootstrap confidence intervals or formal tests of cluster separation; the visualizations should be read as descriptive rather than inferential\. The claim of “non\-linear” geometry rests on the curved structure visible in 2D projections of a 2,560\-dimensional space; projections can distort, so this should be interpreted with caution\.
LayerPositionPC1PC2PC3L24resid\_postsuffix 0 \(<\|im\_end\|\>\)43\.8%29\.8%7\.8%L24attn\_outsuffix 0 \(<\|im\_end\|\>\)45\.6%29\.1%8\.3%L24resid\_postsuffix 3 \(assistant\)49\.0%28\.7%10\.3%L31resid\_postsuffix 3 \(assistant\)70\.7%16\.3%7\.3%Table Y\.1:Variance explained by the first three principal components at key layer\-position pairs\. PC1 captures 44–71% of variance, increasing from L24 to L31 as the temporal signal consolidates\.
### Y\.3Trajectory Analysis
To visualize how representations evolve across layers or positions, we compute two types of PCA trajectories:
- •Aligned trajectories: fit PCA independently per target, then align signs across adjacent layers/positions using correlation continuity\. This produces the 1D layer\-sweep and position\-sweep plots\.
- •Shared trajectories: fit a single PCA on all samples from all layers/positions, then project per\-target\. This produces the 3D trajectory plots where samples from different layers are comparable in the same coordinate system\.
## Appendix Appendix ZBehavioral temporal discounting methodology
### Z\.1Kirby MCQ\-27 Background
Temporal discounting, the tendency to devalue future rewards relative to immediate ones, is a fundamental aspect of human decision\-making\.Kirby et al\. \[[57](https://arxiv.org/html/2606.05194#bib.bib57)\]developed the Monetary Choice Questionnaire \(MCQ\-27\), a 27\-item instrument that estimates an individual’s hyperbolic discount ratekkusing the model:
V=A1\+kD\\displaystyle V=\\frac\{A\}\{1\+kD\}\(Z\.1\)whereVVis the present subjective value of a future rewardAAavailable after delayDD\. Higher values ofkkindicate greater impulsivity \(steeper discounting of future rewards\)\.
Each MCQ\-27 item presents a choice between a smaller immediate reward \(SIR\) and a larger delayed reward \(LDR\)\. At the*indifference point*, where the subject is equally likely to choose either option, the implied discount rate is:
kindiff=A/V−1D=LDR/SIR−1D\\displaystyle k\_\{\\text\{indiff\}\}=\\frac\{A/V\-1\}\{D\}=\\frac\{\\text\{LDR\}/\\text\{SIR\}\-1\}\{D\}\(Z\.2\)
The original MCQ\-27 study\[[57](https://arxiv.org/html/2606.05194#bib.bib57)\]estimatedkkusing a*maximum\-consistency method*: for each candidatekkvalue \(geometric midpoints between adjacentkindiffk\_\{\\text\{indiff\}\}values\), count how many responses are consistent with that discount rate, and assign thekkwith the highest consistency\. The key finding was that heroin\-dependent individuals \(k≈0\.025k\\approx 0\.025\) discounted future rewards roughly twice as steeply as non\-drug\-using controls \(k≈0\.013k\\approx 0\.013\), with consistency rates above 90% in both groups\.
As LLMs are increasingly deployed in advisory, therapeutic, and decision\-support roles, understanding their implicit temporal preferences becomes critical\. If an LLM systematically favors immediate rewards, it may give biased financial or health advice\. Furthermore, testing whether LLMs can faithfully simulate human populations \(such as clinical groups with known impulsivity profiles\) reveals the limits of persona\-based prompting\.
We pursue three questions:
1. 1\.Can open\-weight LLMs replicate human\-like discount rates on the MCQ\-27?
2. 2\.Does prompting an LLM with a heroin\-user persona produce the expected increase in impulsivity?
3. 3\.Does chain\-of\-thought reasoning improve or degrade the fidelity of temporal preference simulation?
### Z\.2Models and Personas
We test two open\-weight models from theQwen3family:Qwen3\-4BandQwen3\-8B, run locally with greedy decoding \(do\_sample=False\)\. All behavioral results are deterministic under greedy decoding; behavior under sampling with temperature\>0\>0\(the typical deployment setting\) may differ\. Each model is tested under two system prompts:
- •Default persona: “You are a 35\-year\-old adult with a stable job and average finances\.”
- •Heroin\-user persona: “You are a 36\-year\-old person who has been using heroin regularly for about 8 years\. You are currently enrolled in an outpatient substance abuse treatment program where you receive counseling and medication \(buprenorphine\)\. You have a high school education\.”
These demographic details are drawn from the heroin\-dependent participant group inKirby et al\. \[[57](https://arxiv.org/html/2606.05194#bib.bib57)\]: their sample had a mean age of 36\.3 years \(SD=7\.3SD=7\.3\), mean duration of heroin use of 8\.4 years \(SD=6\.5SD=6\.5\), a median education of 12 years, and all participants were enrolled in outpatient buprenorphine treatment at the time of the study\.
For each persona, we run two response modes:
- •Direct: The model replies with a single word \(“now” or “later”\), using 2 generated tokens\.
- •Chain\-of\-thought \(CoT\): The model briefly reasons about the tradeoff before giving a final answer, using up to 200 generated tokens\.
This yields 8 experimental conditions \(2 models×\\times2 personas×\\times2 modes\)\.
### Z\.3Decision Boundary Method
Beyond the standard MCQ\-27 scoring, we introduce a*decision boundary*approach\. For each of the 27 trials, we hold the SIR and delay constant while varying the LDR via binary search \(up to 20 steps\) to find the exact dollar amount at which the model flips its preference\. This yields:
- •Theboundary LDR: the indifference point for that trial\.
- •Theboundarykk: the implied discount rate at the flip point, computed ask=\(LDRboundary/SIR−1\)/Dk=\(\\text\{LDR\}\_\{\\text\{boundary\}\}/\\text\{SIR\}\-1\)/D\.
- •Asearch log: the full sequence of \(LDR, choice\) pairs, which reveals consistency and noise in the model’s preferences\.
When a model never flips even at 20×\\timesthe immediate reward, we record the trial as “no boundary found,” indicating extreme present bias on that item\.
## Appendix Appendix AABehavioral coherence methodology
Results are presented in[Appendix P](https://arxiv.org/html/2606.05194#A16)\. Here we describe the experimental setup\.
### AA\.1Task
Each prompt presents a binary choice between a short\-term investment \($20,000 in 6 months, fixed\) and a long\-term investment \(variable: $100K, $300K, or $500K in 10 years\), optionally constrained by an explicit time horizon\.
### AA\.2Experimental Grid
The experiment systematically varies four axes:
- •Time horizons\(10\): none, 1 month, 3 months, 6 months, 1 year, 2 years, 5 years, 10 years, 20 years, 50 years
- •Reward levels\(3\): $100K \(5×\\times\), $300K \(15×\\times\), $500K \(25×\\times\) relative to the $20K short\-term option
- •Formatting\(4\): 2 label styles \(a/bvs\.x/y\)×\\times2 presentation orders \(short\-first vs\. long\-first\)
- •Context framings\(8\): varying role \(household head, individual, committee\), reasoning style \(provide reasoning, step\-by\-step, briefly justify\), and special emphasis \(tradeoff, long\-term thinking\)
This yields10×3×4×8=96010\\times 3\\times 4\\times 8=960samples per model\. We run the grid on 30 models \(28,800 samples total\)\.
### AA\.3Models
The 30 models span five open\-weight families and three API providers, sized from 0\.6B to∼\\sim2,500B parameter\-equivalent:
- •Qwen3 hybrid\-thinking\[[111](https://arxiv.org/html/2606.05194#bib.bib111)\]\(6 variants, run in non\-thinking mode\):Qwen3\-0\.6B,Qwen3\-1\.7B,Qwen3\-4B,Qwen3\-8B,Qwen3\-14B,Qwen3\-32B
- •Qwen3 hybrid\-thinking, run in thinking mode: the sameQwen3\-0\.6B,Qwen3\-1\.7B, andQwen3\-4Bcheckpoints withenable\_thinking=true
- •Qwen3 mode\-specialized 2507 refresh:Qwen3\-4B\-Instruct\-2507\(non\-thinking\-only, our primary model\)
- •Qwen3\.5\(6 variants, instruct\): 0\.8B, 2B, 4B, 9B, 27B, 35B\-A3B
- •Qwen2\.5: 3B\-Instruct
- •Other open weights:Llama\-3\.2\-3B\-Instruct,Mistral\-7B\-Instruct\-v0\.3,gemma\-3\-4b\-it,Phi\-4\-mini\-instruct
- •Anthropic API:claude\-haiku\-4\-5\-20251001,claude\-sonnet\-4\-6,claude\-opus\-4\-7
- •OpenAI API:gpt\-5\.4\-nano,gpt\-5\.4\-mini,gpt\-5\.4,o3
- •Google API:gemini\-2\.5\-flash,gemini\-2\.5\-pro
##### Qwen3 terminology\.
The originalQwen3\-4Bis a post\-trained hybrid\-thinking checkpoint that supports seamless switching between a thinking mode \(emitting a<think\>\.\.\.</think\>reasoning block\) and a non\-thinking mode, controlled by theenable\_thinkingflag or inline/thinkand/no\_thinktags\[[111](https://arxiv.org/html/2606.05194#bib.bib111)\]\. We always run it in non\-thinking mode\. The pretrained base checkpoint is namedQwen3\-4B\-Baseand is not used here\.Qwen3\-4B\-Instruct\-2507is a July\-2025 mode\-specialized refresh that operates exclusively in non\-thinking mode\. The thinking\-mode rows in our tables come from running the hybridQwen3\-\*checkpoints withenable\_thinking=true, not from any separate thinking\-only refresh\.
API model sizes are order\-of\-magnitude estimates used only for size\-based ordering in visualizations\. All analyses operate on the 30 models together; the paper foregroundsQwen3\-4B\-Instruct\-2507, which is also the target of the mechanistic and steering experiments\.
### AA\.4Metrics
We measure five dimensions of behavioral quality:
1. 1\.Coherence: Does the model’s choice respect the time\-horizon constraint? We distinguish exact\-match horizons \(6 months = short delivery, 10 years = long delivery\) from genuine\-reasoning horizons where the correct answer requires temporal judgment\.
2. 2\.Order stability: Does swapping which option appears first change the choice? Values below 50% indicate the model picks whichever option appears first\.
3. 3\.Label stability: Does changing the label format \(a/bvs\.x/y\) change the choice?
4. 4\.Reward sensitivity: Does increasing the long\-term reward increase long\-term preference? Measured on no\-horizon samples only, to avoid the horizon dominating the choice\.
5. 5\.Context sensitivity: How much does the context framing shift the preference?
## Appendix Appendix ABContrastive steering methodology
This appendix describes the CAA steering vector construction and intervention protocol\. The probing methodology that identifies the steering direction is described in[Appendix U](https://arxiv.org/html/2606.05194#A21)\. Steering results are presented in[Appendix S](https://arxiv.org/html/2606.05194#A19)\.
### AB\.1CAA Steering Vector Construction
Following the Contrastive Activation Addition \(CAA\) framework\[[100](https://arxiv.org/html/2606.05194#bib.bib100),[78](https://arxiv.org/html/2606.05194#bib.bib78)\], we construct a steering vector\. The vector is computed fromDimplicitD\_\{\\mathrm\{implicit\}\}at layer 26 \(the best probe layer\):
𝐯CAA=1N∑i=1N𝐚ilong−1N∑i=1N𝐚iimm\\mathbf\{v\}\_\{\\mathrm\{CAA\}\}=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\mathbf\{a\}\_\{i\}^\{\\mathrm\{long\}\}\-\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\mathbf\{a\}\_\{i\}^\{\\mathrm\{imm\}\}\(AB\.1\)where𝐚ilong\\mathbf\{a\}\_\{i\}^\{\\mathrm\{long\}\}and𝐚iimm\\mathbf\{a\}\_\{i\}^\{\\mathrm\{imm\}\}are the layer\-26 residual\-stream activations at theim\_end−1\-1token position for the long\-term and immediate choices of pairii, respectively, andN=300N=300\.
The raw vector hasℓ2\\ell\_\{2\}norm 30\.30 and is normalized to unit norm for scale\-agnostic steering:𝐯^CAA=𝐯CAA/‖𝐯CAA‖2\\hat\{\\mathbf\{v\}\}\_\{\\mathrm\{CAA\}\}=\\mathbf\{v\}\_\{\\mathrm\{CAA\}\}/\\\|\\mathbf\{v\}\_\{\\mathrm\{CAA\}\}\\\|\_\{2\}\.
##### Why the implicit dataset?
The explicit dataset’s long\-term choices contain surface time words\. A vector computed from explicit pairs would partially encode vocabulary differences rather than the underlying temporal reasoning concept\. The implicit direction lives in a deeper semantic subspace, as confirmed by the PCA analysis in Section[G\.3](https://arxiv.org/html/2606.05194#A7.SS3)\.
### AB\.2Steering Intervention
We evaluate steering by adding a scaled version of the normalized CAA vector to the residual stream at a target layer during the forward pass:
𝐡\(l\)←𝐡\(l\)\+α⋅𝐯^CAA\\mathbf\{h\}^\{\(l\)\}\\leftarrow\\mathbf\{h\}^\{\(l\)\}\+\\alpha\\cdot\\hat\{\\mathbf\{v\}\}\_\{\\mathrm\{CAA\}\}\(AB\.2\)where𝐡\(l\)\\mathbf\{h\}^\{\(l\)\}is the residual\-stream activation at layerllandα\\alphais the steering coefficient\. The hook is applied to all token positions during the forward pass\.
### AB\.3Forced\-Choice Metric
For each steered configuration, we compute the mean difference in log\-probability between the long\-term and immediate completions on 20 held\-out explicit prompts:
S\(α,l\)=1\|𝒫\|∑p∈𝒫\[logP¯\(long\_term∣p;α,l\)−logP¯\(immediate∣p;α,l\)\]S\(\\alpha,l\)=\\frac\{1\}\{\|\\mathcal\{P\}\|\}\\sum\_\{p\\in\\mathcal\{P\}\}\\left\[\\overline\{\\log P\}\(\\mathrm\{long\\\_term\}\\mid p;\\,\\alpha,l\)\\;\-\\;\\overline\{\\log P\}\(\\mathrm\{immediate\}\\mid p;\\,\\alpha,l\)\\right\]\(AB\.3\)wherelogP¯\\overline\{\\log P\}denotes the mean token\-level log\-probability of the choice text under teacher\-forcing\. A positive score indicates that the model assigns higher probability to the long\-term completion\.
## Appendix Appendix ACCase study: a single highly\-formatted pair
Everything so far has been aggregate\. Attribution scores averaged over hundreds of prompts\. Patching effects pooled across contrastive pairs\. Probe accuracies on held\-out sets\. Those methods locate the subgraph, but they do not let you sit with a single computation long enough to see how it works\.
Here we do the opposite\. One pair\. Two prompts that differ by exactly 24 tokens: the presence or absence of one sentence specifying an 8\-month time horizon\. With the constraint, the model choosesa\), the $20,000 in 6 months, the option that can deliver within the deadline\. Without it, the model choosesb\), the $500,000 in 10 years, the option with higher reward and no deadline to violate\. That is the entire manipulation\. One prompt reveals the model’s*constrained preference*, the other its*latent preference*\. The question is where, in 36 layers and 166 token positions, the constraint enters the computation and redirects the choice\.
Activation patching \([Appendix W](https://arxiv.org/html/2606.05194#A23)\) is at its most precise when only one variable has moved\. Every change in the model’s internal state is causally downstream of those 24 tokens\. But to see the mechanism clearly, we need to measure more than which token the model picks\. We need to track how certain it is, how many alternatives it entertains at each layer, and where that certainty breaks down under intervention\. This leads us to entropy, diversity, and the geometric\-mean probability of the response, which we define below before tracing them through the network\.
### AC\.1Tokenization and position mapping
Figure[AC\.1](https://arxiv.org/html/2606.05194#A29.F1)shows the token\-level structure of both prompts\. The 24\-token difference corresponds exactly to theCONSTRAINTsection \(positions 83–106 in the clean prompt\):CONSTRAINT: You must select the option that provides the greatest benefit for this time horizon: 8 months\.
Figure AC\.1:Tokenization of the clean \(166 tokens, top\) and corrupted \(142 tokens, bottom\) prompts\. The clean prompt includes theCONSTRAINTsection \(positions 83–106, highlighted\) specifying an 8\-month time horizon\. Removing this section flips the model’s choice froma\)\(short\-term, coherent\) tob\)\(long\-term, latent preference\)\.Figure[AC\.2](https://arxiv.org/html/2606.05194#A29.F2)shows the semantic region annotations and position alignment\. Because the prompts differ in length, the position mapping described in[Appendix W](https://arxiv.org/html/2606.05194#A23)uses structural markers \(SITUATION,TASK, etc\.\) as anchors to align tokens across the pair for activation patching\.
Figure AC\.2:Semantic region mapping for both prompts\. Each token is annotated with its structural role \(situation, task, options, constraint, action, format, response\)\. The constraint region \(positions 83–106 in the clean prompt\) has no counterpart in the corrupted prompt; remaining sections are aligned via piecewise linear interpolation\.
### AC\.2What to measure, and why
Start with the simplest question: which token does the model prefer? The logit difference answers directly:
ℓ=logit\(a\)−logit\(b\)\\ell=\\mathrm\{logit\}\(\\texttt\{a\}\)\-\\mathrm\{logit\}\(\\texttt\{b\}\)\(AC\.1\)Positive means short\-term wins\. The normalized recovery \(denoising\) and disruption \(noising\) compress this into\[0,1\]\[0,1\]\([Appendix W](https://arxiv.org/html/2606.05194#A23)\)\. The softmax probabilitiesP\(short\)P\(\\texttt\{short\}\),P\(long\)P\(\\texttt\{long\}\)tell us whether the preference is decisive or marginal\.
But preference is not the whole story\. When we patch activations at a given layer, we are not just changing which token leads; we are restructuring the model’s entire distribution over the vocabulary\. That restructuring has a natural measure\. The Shannon entropy of the output distribution at the choice position is
H=−∑i=1VpilogpiH=\-\\sum\_\{i=1\}^\{V\}p\_\{i\}\\log p\_\{i\}\(AC\.2\)whereVVis the vocabulary size andpip\_\{i\}is the probability of tokenii\.
The exponential of entropy has a name and an interpretation\. The*diversity*is the effective number of equally likely tokens the model is choosing among\[[61](https://arxiv.org/html/2606.05194#bib.bib61)\]:
D1=exp\(H\)\{\}^\{1\}\\\!D=\\exp\(H\)\(AC\.3\)WhenD1=1\{\}^\{1\}\\\!D=1, the model has committed; there is, effectively, one option\. WhenD1=4\{\}^\{1\}\\\!D=4, the model is as uncertain as if it were choosing uniformly among four tokens\. This is the Hill number of orderq=1q=1\. It belongs to a family indexed by a sensitivity parameterqq\[[61](https://arxiv.org/html/2606.05194#bib.bib61)\]:
Dq=exp\(Hq\),Hq=11−qlog\(∑i=1Vpiq\)\{\}^\{q\}\\\!D=\\exp\(H\_\{q\}\),\\qquad H\_\{q\}=\\frac\{1\}\{1\-q\}\\log\\\!\\left\(\\sum\_\{i=1\}^\{V\}p\_\{i\}^\{\\,q\}\\right\)\(AC\.4\)whereHqH\_\{q\}is the Rényi entropy of orderqq\. Atq=1q=1this recovers perplexity; atq=2q=2, the inverse Simpson index\. The framework is axiomatic: any measure of diversity satisfying natural symmetry and composition properties must be a Hill number for someqq\.
One more quantity\. The metrics above describe the model’s state at the choice token\. But the model does not just pick a letter; it generates an entire response string \(“I choose: a\)\. My reasoning: …”\)\. The*inverse perplexity*of that string is the geometric\-mean token probability:
inv\_ppl=exp\(−Hstring\)=\(∏t=1Tp\(xt∣x<t\)\)1/T\\mathrm\{inv\\\_ppl\}=\\exp\(\-H\_\{\\mathrm\{string\}\}\)=\\left\(\\prod\_\{t=1\}^\{T\}p\(x\_\{t\}\\mid x\_\{<t\}\)\\right\)^\{1/T\}\(AC\.5\)It measures how likely the model considers its own output, token by token, on average\. It ranges from∼\\sim0 \(guessing\) to∼\\sim1 \(certain about every token\)\. When an intervention flips the choice,inv\_ppl\(short\)\\mathrm\{inv\\\_ppl\}\(\\text\{short\}\)andinv\_ppl\(long\)\\mathrm\{inv\\\_ppl\}\(\\text\{long\}\)trade places\. The crossover layer is where the model commits\.
### AC\.3Layer sweeps
We patch one component at a time across layers 16–35 at the divergent token position\. Each figure has two rows \(denoising above, noising below\) and five column panels corresponding to the metrics above\. The question at each layer: has the intervention flipped the decision yet?
Figure AC\.3:resid\_prelayer sweep\. Recovery rises sigmoidally from∼\\sim0\.1 at L20 to∼\\sim1\.0 at L24\. The entropy spike \(∼\\sim1\.4 nats, diversity≈4\\approx 4\) peaks at exactly L22–23, the midpoint of the recovery sigmoid, not before or after\. The model’s uncertainty is maximal precisely when the intervention has half\-rewritten the decision\.Figure AC\.4:attn\_outlayer sweep\. Unlike the smooth residual transition, attention effects are sparse and spiky\. The entropy spike for attention \(∼\\sim0\.8 nats\) peaks at L24–25, about 2 layers later than forresid\_pre\. Attention writes its correction*after*the residual stream has begun to shift, consistent with a read\-then\-write pattern\.Figure AC\.5:resid\_midlayer sweep \(after attention, before MLP\)\. Comparingresid\_midwithresid\_preisolates the attention contribution at each layer\. The difference is largest at L24, where attention adds the single biggest correction to the residual stream\.Figure AC\.6:mlp\_outlayer sweep\. MLP effects are distributed across layers 22–35; no single layer dominates\. The noising row \(bottom\) shows MLP disruption beginning∼\\sim3 layers later than attention disruption, suggesting MLP processes the temporal signal*downstream*of the attention computation that initiates it\.Figure AC\.7:resid\_postlayer sweep \(after MLP\)\. Nearly identical toresid\_preshifted∼\\sim1 layer later\. Recovery*overshoots*1\.0 briefly at L24–25 \(reaching∼\\sim1\.4\), meaning the patched model is*more*short\-term\-biased than the clean model at those layers before settling back\. The overshoot disappears by L28\.##### Information flow across components\.
The five layer sweeps, read side by side, trace how temporal information propagates within each transformer block\.resid\_pre\(Figure[AC\.3](https://arxiv.org/html/2606.05194#A29.F3)\) shows the state entering each layer: a smooth sigmoidal transition, with recovery climbing from∼\\sim0\.1 at L20 to∼\\sim1\.0 at L24\.attn\_out\(Figure[AC\.4](https://arxiv.org/html/2606.05194#A29.F4)\) isolates the attention contribution: sparse spikes at L23–25, with most layers near zero\.resid\_mid\(Figure[AC\.5](https://arxiv.org/html/2606.05194#A29.F5)\) isresid\_pre\+attn\_out, so it inherits both the smooth ramp and the spikes\.mlp\_out\(Figure[AC\.6](https://arxiv.org/html/2606.05194#A29.F6)\) contributes distributed, individually smaller effects across L22–35; no single MLP layer dominates the way L24 attention does\.resid\_post\(Figure[AC\.7](https://arxiv.org/html/2606.05194#A29.F7)\) integrates everything and is nearly identical toresid\_preshifted by one layer, as expected from the residual connection \(resid\_post\[L\]=resid\_pre\[L\]\+attn\_out\[L\]\+mlp\_out\[L\]\\texttt\{resid\\\_post\}\[L\]=\\texttt\{resid\\\_pre\}\[L\]\+\\texttt\{attn\\\_out\}\[L\]\+\\texttt\{mlp\\\_out\}\[L\], which becomesresid\_pre\[L\+1\]\\texttt\{resid\\\_pre\}\[L\+1\]\)\.
The pattern is clear: attention provides the decisive, layer\-specific corrections that flip the decision; MLP distributes refinement across many layers; and the residual stream accumulates both into a monotonic commitment\.
##### Denoising vs\. noising asymmetry\.
The noising transition \(bottom rows\) is shifted∼\\sim1–2 layers later than the denoising transition\. This means it is slightly harder to*break*the clean decision than to*restore*it: the model’s correct representations are more robust to corruption at early layers than to recovery from corruption\. This is consistent with the necessity/sufficiency asymmetry observed in the aggregated results \([Appendix J](https://arxiv.org/html/2606.05194#A10)\)\.
##### The intervention forces uncertainty\.
The vocabulary entropy \(column 4\) spikes precisely at the layers where the decision flips, and the spike is not incidental: it reveals that the patching intervention forces the model through a state of genuine uncertainty before it can commit to the new answer\.
Underdenoising\(patching clean activations into the corrupted run\), the entropy spike is sharp and tall:resid\_prepeaks at∼\\sim1\.4 nats at layer 22–23 \(diversityexp\(H\)≈4\\exp\(H\)\\approx 4effective tokens\),resid\_postat∼\\sim1\.1 nats at layer 22, andattn\_outat∼\\sim0\.8 nats at layer 24–25\. Before the spike \(<<L20\), entropy is∼\\sim0\.05 nats \(diversity≈1\\approx 1: the model is certain about “b”\)\. After the spike \(\>\>L24\), entropy settles at∼\\sim0\.2 nats \(the model is now certain about “a” but slightly less peaked\)\.
Undernoising\(patching corrupted activations into the clean run\), the same spike appears but is*broader, lower, and shifted later*:resid\_prepeaks at∼\\sim0\.7 nats at layers 24–26, roughly half the denoising amplitude and 2 layers later\. This asymmetry means it is easier to*restore*the constrained preference \(the denoising intervention creates a clean, sharp transition\) than to*destroy*it \(the noising intervention encounters more resistance, producing a gentler, more gradual restructuring\)\.
##### Inverse perplexity tracks commitment\.
The trajectory panel \(column 5\) showsinv\_ppl=exp\(−Hstring\)\\mathrm\{inv\\\_ppl\}=\\exp\(\-H\_\{\\text\{string\}\}\), the geometric\-mean probability over the response tokens\. This measures not just which token the model favors, but how confident it is about the*entire output string*\. Under denoising,inv\_ppl\(short\)\\mathrm\{inv\\\_ppl\}\(\\text\{short\}\)jumps from∼\\sim0 to∼\\sim1\.0 at layers 22–24, coinciding exactly with the entropy spike\. The model transitions from being certain about the long\-term response to being certain about the short\-term response\. The entropy spike is the moment in between: the model has abandoned one answer but has not yet committed to the other\.
### AC\.4Position sweeps
The layer sweeps told us*when*\(which layer\) the decision flips\. Now we ask*where*\(which tokens\) the temporal information lives\. We fix the layer at the most causally important depth and sweep across token positions 50–160\.
Figure AC\.8:resid\_preposition sweep\. Two regions show causal effect: positions 83–106 \(constraint\) and∼\\sim130–140 \(response boundary\)\. Under denoising, recovery peaks at the response boundary \(∼\\sim130\); under noising, disruption peaks at the constraint region \(∼\\sim95–105\)\. The model*stores*temporal information at one location and*reads*it at another\.Figure AC\.9:attn\_outposition sweep\. Attention effects localize almost entirely to positions∼\\sim130–140, the response boundary, with negligible effect at the constraint tokens\. Attention does not*store*the temporal information; it*retrieves*it at the moment of choice\.Figure AC\.10:resid\_midposition sweep\. Combinesresid\_preandattn\_out: both constraint and response regions show effects\. The response\-boundary effect is larger inresid\_midthan inresid\_pre, confirming that attention at this position has just written its correction\.Figure AC\.11:mlp\_outposition sweep\. MLP effects are more spatially distributed than attention but still concentrate in the same two regions, suggesting MLP refines the signal that attention initiates rather than introducing independent positional information\.Figure AC\.12:resid\_postposition sweep\. The cumulative picture shows that temporal preference in this pair reduces to an interaction between two positions separated by∼\\sim30 tokens: where the constraint is stored \(83–106\) and where the choice is made \(∼\\sim130–140\)\. Everything else in the 166\-token prompt is causally inert\.##### Entropy in the position dimension reveals a spatial dissociation\.
In the layer sweeps, entropy spikes coincided with the decision transition\. In the position sweeps, a subtler pattern emerges: the entropy effects for denoising and noising localize to*different*token regions\. Under denoising, the entropy spike \(∼\\sim0\.3 nats\) appears at positions∼\\sim130–135, the response boundary where the model writes its choice\. Under noising, the entropy spike \(∼\\sim0\.5 nats\) appears at positions∼\\sim95–105, the constraint region\. This spatial dissociation suggests that restoring the constrained preference acts at the*output*\(where the choice is generated\), while disrupting it acts at the*input*\(where the constraint information is stored\)\.
##### Connecting positions to tokens\.
Figure[AC\.13](https://arxiv.org/html/2606.05194#A29.F13)summarizes the spatial structure\. The two causally important regions correspond to specific semantic content \(cf\. Figure[AC\.2](https://arxiv.org/html/2606.05194#A29.F2)\):
position020406080100120140160chatsituationtask \+ optionsobjectiveCONSTRAINTactionformatresponseformat tailClean prompt\(166 tokens\)noisingdisruptiondenoisingrecoveryH≈0\.5H\\approx 0\.5natsH≈0\.3H\\approx 0\.3nats
Figure AC\.13:Semantic regions of the clean prompt with the two causally important zones highlighted\. Noising disruption concentrates at theCONSTRAINTtokens \(positions 83–106\), where the horizon information is stored\. Denoising recovery concentrates at the response boundary \(∼\\sim130–140\), where the model reads the constraint to generate its choice\.- •Positions 83–106\(constraint section\)\. The noising entropy spike peaks at positions 95–105, which correspond to the tokens “time”, “horizon”, “:”, “8”, “months” \(the core of the temporal constraint\)\. The word “8” at position 103 and “months” at position 104 are the most specific tokens: they encode the deadline that makes the short\-term option coherent\. These positions exist only in the clean prompt; in the corrupted prompt, the alignment maps them onto theACTIONsection\.
- •Positions∼\\sim130–140\(format/response boundary\)\. The denoising entropy spike peaks here, at “I”, “choose”, “:”, and the choice token itself\. Attention at these positions shows the strongest individual\-position effects, consistent with the model attending*back*to the constraint region when generating the choice\. The spatial separation between the two zones \(constraint at 83–106, readout at 130–140\) mirrors the temporal separation in the layer sweeps: information is stored in mid\-layers and read out in later layers\.
##### Constrained vs\. latent preference\.
The pair structure makes it possible to distinguish two modes of temporal preference:
- •Constrained preference\(clean prompt\): the model evaluates options against an explicit deadline and coherently picks the option that delivers within 8 months\. The constraint tokens \(positions 83–106\) are the causal mechanism: patching them from the clean into the corrupted run restores the short\-term choice\.
- •Latent preference\(corrupted prompt\): without a constraint, the model defaults to the higher\-reward option \($500K in 10 years\), revealing a latent long\-term bias when no temporal pressure is applied\. This is the same default preference observed in the no\-horizon condition of the behavioral coherence experiment \([Appendix P](https://arxiv.org/html/2606.05194#A16)\), whereQwen3\-4B\-Instruct\-2507picks long\-term∼\\sim96% of the time without a horizon constraint\.
##### Component granularity in position sweeps\.
The five position sweeps reveal how information propagates within each transformer block at the critical positions:
1. 1\.resid\_preshows effects at the constraint region \(83–106\), indicating the temporal information is already in the residual stream from earlier layers\.
2. 2\.attn\_outshows effects primarily at the response boundary \(∼\\sim130–140\), indicating attention*reads*from the constraint to*write*the choice\.
3. 3\.resid\_midcombines both, as expected \(it isresid\_pre\+attn\_out\)\.
4. 4\.mlp\_outshows distributed effects, contributing refinement at both regions\.
5. 5\.resid\_postshows the final, integrated picture\.
This flow, constraint information stored in the residual stream at positions 83–106, read by attention at positions∼\\sim130–140, refined by MLP, and committed in the residual stream, mirrors the geometry findings \([Appendix M](https://arxiv.org/html/2606.05194#A13)\): the model builds a temporal representation in mid\-layers and converts it to a preference at the turn/response boundary\.
## NeurIPS Paper Checklist
1. 1\.Claims
2. Question: Do the main claims made in the abstract and introduction accurately reflect the paper’s contributions and scope?
3. Answer:\[Yes\]
4. Justification: The abstract and Section[1](https://arxiv.org/html/2606.05194#S1)state the three contributions \(localizing a temporal\-preference subgraph, characterizing its representational geometry, and steering the underlying axis\) and each is delivered by a corresponding section \([3](https://arxiv.org/html/2606.05194#S3)–[5](https://arxiv.org/html/2606.05194#S5)\) and appendix\.
5. Guidelines: - •The answer\[N/A\]means that the abstract and introduction do not include the claims made in the paper\. - •The abstract and/or introduction should clearly state the claims made, including the contributions made in the paper and important assumptions and limitations\. A\[No\]or\[N/A\]answer to this question will not be perceived well by the reviewers\. - •The claims made should match theoretical and experimental results, and reflect how much the results can be expected to generalize to other settings\. - •It is fine to include aspirational goals as motivation as long as it is clear that these goals are not attained by the paper\.
6. 2\.Limitations
7. Question: Does the paper discuss the limitations of the work performed by the authors?
8. Answer:\[Yes\]
9. Justification: Section[7](https://arxiv.org/html/2606.05194#S7)explicitly discusses scaling to a single model \(Qwen3\-4B\-Instruct\-2507\), domain generalization, the single\-turn restriction, interactions with other concepts, and the linear\-manifold approximation used for steering; the extended limitations appendix[Appendix F](https://arxiv.org/html/2606.05194#A6)additionally documents the distributed attribution mass and the synthetic\-label provenance of the contrastive datasets\.
10. Guidelines: - •The answer\[N/A\]means that the paper has no limitation while the answer\[No\]means that the paper has limitations, but those are not discussed in the paper\. - •The authors are encouraged to create a separate “Limitations” section in their paper\. - •The paper should point out any strong assumptions and how robust the results are to violations of these assumptions \(e\.g\., independence assumptions, noiseless settings, model well\-specification, asymptotic approximations only holding locally\)\. The authors should reflect on how these assumptions might be violated in practice and what the implications would be\. - •The authors should reflect on the scope of the claims made, e\.g\., if the approach was only tested on a few datasets or with a few runs\. In general, empirical results often depend on implicit assumptions, which should be articulated\. - •The authors should reflect on the factors that influence the performance of the approach\. For example, a facial recognition algorithm may perform poorly when image resolution is low or images are taken in low lighting\. Or a speech\-to\-text system might not be used reliably to provide closed captions for online lectures because it fails to handle technical jargon\. - •The authors should discuss the computational efficiency of the proposed algorithms and how they scale with dataset size\. - •If applicable, the authors should discuss possible limitations of their approach to address problems of privacy and fairness\. - •While the authors might fear that complete honesty about limitations might be used by reviewers as grounds for rejection, a worse outcome might be that reviewers discover limitations that aren’t acknowledged in the paper\. The authors should use their best judgment and recognize that individual actions in favor of transparency play an important role in developing norms that preserve the integrity of the community\. Reviewers will be specifically instructed to not penalize honesty concerning limitations\.
11. 3\.Theory assumptions and proofs
12. Question: For each theoretical result, does the paper provide the full set of assumptions and a complete \(and correct\) proof?
13. Answer:\[N/A\]
14. Justification:\[N/A\]
15. Guidelines: - •The answer\[N/A\]means that the paper does not include theoretical results\. - •All the theorems, formulas, and proofs in the paper should be numbered and cross\-referenced\. - •All assumptions should be clearly stated or referenced in the statement of any theorems\. - •The proofs can either appear in the main paper or the supplemental material, but if they appear in the supplemental material, the authors are encouraged to provide a short proof sketch to provide intuition\. - •Inversely, any informal proof provided in the core of the paper should be complemented by formal proofs provided in appendix or supplemental material\. - •Theorems and Lemmas that the proof relies upon should be properly referenced\.
16. 4\.Experimental result reproducibility
17. Question: Does the paper fully disclose all the information needed to reproduce the main experimental results of the paper to the extent that it affects the main claims and/or conclusions of the paper \(regardless of whether the code and data are provided or not\)?
18. Answer:\[Yes\]
19. Justification: The model \(Qwen3\-4B\-Instruct\-2507\) is openly released; prompts, datasets, attribution \(EAP\-IG\), probing, CAA steering construction, and evaluation procedures are fully specified in Appendices[Appendix E](https://arxiv.org/html/2606.05194#A5)–[Appendix S](https://arxiv.org/html/2606.05194#A19)\.
20. Guidelines: - •The answer\[N/A\]means that the paper does not include experiments\. - •If the paper includes experiments, a\[No\]answer to this question will not be perceived well by the reviewers: Making the paper reproducible is important, regardless of whether the code and data are provided or not\. - •If the contribution is a dataset and/or model, the authors should describe the steps taken to make their results reproducible or verifiable\. - •Depending on the contribution, reproducibility can be accomplished in various ways\. For example, if the contribution is a novel architecture, describing the architecture fully might suffice, or if the contribution is a specific model and empirical evaluation, it may be necessary to either make it possible for others to replicate the model with the same dataset, or provide access to the model\. In general, releasing code and data is often one good way to accomplish this, but reproducibility can also be provided via detailed instructions for how to replicate the results, access to a hosted model \(e\.g\., in the case of a large language model\), releasing of a model checkpoint, or other means that are appropriate to the research performed\. - •While NeurIPS does not require releasing code, the conference does require all submissions to provide some reasonable avenue for reproducibility, which may depend on the nature of the contribution\. For example 1. \(a\)If the contribution is primarily a new algorithm, the paper should make it clear how to reproduce that algorithm\. 2. \(b\)If the contribution is primarily a new model architecture, the paper should describe the architecture clearly and fully\. 3. \(c\)If the contribution is a new model \(e\.g\., a large language model\), then there should either be a way to access this model for reproducing the results or a way to reproduce the model \(e\.g\., with an open\-source dataset or instructions for how to construct the dataset\)\. 4. \(d\)We recognize that reproducibility may be tricky in some cases, in which case authors are welcome to describe the particular way they provide for reproducibility\. In the case of closed\-source models, it may be that access to the model is limited in some way \(e\.g\., to registered users\), but it should be possible for other researchers to have some path to reproducing or verifying the results\.
21. 5\.Open access to data and code
22. Question: Does the paper provide open access to the data and code, with sufficient instructions to faithfully reproduce the main experimental results, as described in supplemental material?
23. Answer:\[No\]
24. Justification: A public code repository exists but is intentionally not linked in this submission to preserve anonymity; it will be released alongside the camera\-ready version\. Appendices[Appendix E](https://arxiv.org/html/2606.05194#A5)–[Appendix S](https://arxiv.org/html/2606.05194#A19)document datasets, prompts, and procedures in sufficient detail to re\-implement every experiment on the publicly availableQwen3\-4B\-Instruct\-2507checkpoint\.
25. Guidelines: - •The answer\[N/A\]means that paper does not include experiments requiring code\. - • - •While we encourage the release of code and data, we understand that this might not be possible, so\[No\]is an acceptable answer\. Papers cannot be rejected simply for not including code, unless this is central to the contribution \(e\.g\., for a new open\-source benchmark\)\. - •The instructions should contain the exact command and environment needed to run to reproduce the results\. See the NeurIPS code and data submission guidelines \([https://neurips\.cc/public/guides/CodeSubmissionPolicy](https://neurips.cc/public/guides/CodeSubmissionPolicy)\) for more details\. - •The authors should provide instructions on data access and preparation, including how to access the raw data, preprocessed data, intermediate data, and generated data, etc\. - •The authors should provide scripts to reproduce all experimental results for the new proposed method and baselines\. If only a subset of experiments are reproducible, they should state which ones are omitted from the script and why\. - •At submission time, to preserve anonymity, the authors should release anonymized versions \(if applicable\)\. - •Providing as much information as possible in supplemental material \(appended to the paper\) is recommended, but including URLs to data and code is permitted\.
26. 6\.Experimental setting/details
27. Question: Does the paper specify all the training and test details \(e\.g\., data splits, hyperparameters, how they were chosen, type of optimizer\) necessary to understand the results?
28. Answer:\[Yes\]
29. Justification: No model training is involved; Section[4](https://arxiv.org/html/2606.05194#S4)summarizes datasets, counterbalancing, and judge setup, and the appendices specify layer ranges, steering coefficients, probe training protocols, and evaluation prompts\.
30. Guidelines: - •The answer\[N/A\]means that the paper does not include experiments\. - •The experimental setting should be presented in the core of the paper to a level of detail that is necessary to appreciate the results and make sense of them\. - •The full details can be provided either with the code, in appendix, or as supplemental material\.
31. 7\.Experiment statistical significance
32. Question: Does the paper report error bars suitably and correctly defined or other appropriate information about the statistical significance of the experiments?
33. Answer:\[Yes\]
34. Justification: All activation\-patching figures report mean±\\pmone standard deviation across input pairs, displayed as shaded bands on line plots and error bars on bar charts \(titles explicitly state “mean±\\pmstd”\); the variability source \(across input pairs\) andσ\\sigmadefinition are stated in each figure caption\.
35. Guidelines: - •The answer\[N/A\]means that the paper does not include experiments\. - •The authors should answer\[Yes\]if the results are accompanied by error bars, confidence intervals, or statistical significance tests, at least for the experiments that support the main claims of the paper\. - •The factors of variability that the error bars are capturing should be clearly stated \(for example, train/test split, initialization, random drawing of some parameter, or overall run with given experimental conditions\)\. - •The method for calculating the error bars should be explained \(closed form formula, call to a library function, bootstrap, etc\.\) - •The assumptions made should be given \(e\.g\., Normally distributed errors\)\. - •It should be clear whether the error bar is the standard deviation or the standard error of the mean\. - •It is OK to report 1\-sigma error bars, but one should state it\. The authors should preferably report a 2\-sigma error bar than state that they have a 96% CI, if the hypothesis of Normality of errors is not verified\. - •For asymmetric distributions, the authors should be careful not to show in tables or figures symmetric error bars that would yield results that are out of range \(e\.g\., negative error rates\)\. - •If error bars are reported in tables or plots, the authors should explain in the text how they were calculated and reference the corresponding figures or tables in the text\.
36. 8\.Experiments compute resources
37. Question: For each experiment, does the paper provide sufficient information on the computer resources \(type of compute workers, memory, time of execution\) needed to reproduce the experiments?
38. Answer:\[Yes\]
39. Justification: Section[4](https://arxiv.org/html/2606.05194#S4)states that the full pipeline runs end\-to\-end within two weeks on a single MacBook Pro \(M4 Max, 48 GB\); no GPU cluster is required\.
40. Guidelines: - •The answer\[N/A\]means that the paper does not include experiments\. - •The paper should indicate the type of compute workers CPU or GPU, internal cluster, or cloud provider, including relevant memory and storage\. - •The paper should provide the amount of compute required for each of the individual experimental runs as well as estimate the total compute\. - •The paper should disclose whether the full research project required more compute than the experiments reported in the paper \(e\.g\., preliminary or failed experiments that didn’t make it into the paper\)\.
41. 9\.Code of ethics
43. Answer:\[Yes\]
44. Justification: The work analyzes a publicly released model, uses no human subjects or private data, and introduces no new training pipeline or deployed system\.
45. Guidelines: - •The answer\[N/A\]means that the authors have not reviewed the NeurIPS Code of Ethics\. - •If the authors answer\[No\], they should explain the special circumstances that require a deviation from the Code of Ethics\. - •The authors should make sure to preserve anonymity \(e\.g\., if there is a special consideration due to laws or regulations in their jurisdiction\)\.
46. 10\.Broader impacts
47. Question: Does the paper discuss both potential positive societal impacts and negative societal impacts of the work performed?
48. Answer:\[Yes\]
49. Justification: Section[1](https://arxiv.org/html/2606.05194#S1)motivates the work as AI\-safety\-relevant by framing temporal preference as a controllable property, and the steering experiments \(Section[5](https://arxiv.org/html/2606.05194#S5),[Appendix S](https://arxiv.org/html/2606.05194#A19)\) show both constructive \(alignment\-style control\) and dual\-use \(bias induction\) implications\.
50. Guidelines: - •The answer\[N/A\]means that there is no societal impact of the work performed\. - •If the authors answer\[N/A\]or\[No\], they should explain why their work has no societal impact or why the paper does not address societal impact\. - •Examples of negative societal impacts include potential malicious or unintended uses \(e\.g\., disinformation, generating fake profiles, surveillance\), fairness considerations \(e\.g\., deployment of technologies that could make decisions that unfairly impact specific groups\), privacy considerations, and security considerations\. - •The conference expects that many papers will be foundational research and not tied to particular applications, let alone deployments\. However, if there is a direct path to any negative applications, the authors should point it out\. For example, it is legitimate to point out that an improvement in the quality of generative models could be used to generate Deepfakes for disinformation\. On the other hand, it is not needed to point out that a generic algorithm for optimizing neural networks could enable people to train models that generate Deepfakes faster\. - •The authors should consider possible harms that could arise when the technology is being used as intended and functioning correctly, harms that could arise when the technology is being used as intended but gives incorrect results, and harms following from \(intentional or unintentional\) misuse of the technology\. - •If there are negative societal impacts, the authors could also discuss possible mitigation strategies \(e\.g\., gated release of models, providing defenses in addition to attacks, mechanisms for monitoring misuse, mechanisms to monitor how a system learns from feedback over time, improving the efficiency and accessibility of ML\)\.
51. 11\.Safeguards
52. Question: Does the paper describe safeguards that have been put in place for responsible release of data or models that have a high risk for misuse \(e\.g\., pre\-trained language models, image generators, or scraped datasets\)?
53. Answer:\[N/A\]
54. Justification:\[N/A\]
55. Guidelines: - •The answer\[N/A\]means that the paper poses no such risks\. - •Released models that have a high risk for misuse or dual\-use should be released with necessary safeguards to allow for controlled use of the model, for example by requiring that users adhere to usage guidelines or restrictions to access the model or implementing safety filters\. - •Datasets that have been scraped from the Internet could pose safety risks\. The authors should describe how they avoided releasing unsafe images\. - •We recognize that providing effective safeguards is challenging, and many papers do not require this, but we encourage authors to take this into account and make a best faith effort\.
56. 12\.Licenses for existing assets
57. Question: Are the creators or original owners of assets \(e\.g\., code, data, models\), used in the paper, properly credited and are the license and terms of use explicitly mentioned and properly respected?
58. Answer:\[Yes\]
59. Justification: We useQwen3\(Apache 2\.0\), Anthropic Claude API, OpenAI GPT API, and Google Gemini API \(each subject to provider terms of service\); the Kirby MCQ\-27 instrument from\[[57](https://arxiv.org/html/2606.05194#bib.bib57)\]is reproduced in standard form\. Each is credited in the relevant section\.
60. Guidelines: - •The answer\[N/A\]means that the paper does not use existing assets\. - •The authors should cite the original paper that produced the code package or dataset\. - •The authors should state which version of the asset is used and, if possible, include a URL\. - •The name of the license \(e\.g\., CC\-BY 4\.0\) should be included for each asset\. - •For scraped data from a particular source \(e\.g\., website\), the copyright and terms of service of that source should be provided\. - •If assets are released, the license, copyright information, and terms of use in the package should be provided\. For popular datasets,[paperswithcode\.com/datasets](https://arxiv.org/html/2606.05194v1/paperswithcode.com/datasets)has curated licenses for some datasets\. Their licensing guide can help determine the license of a dataset\. - •For existing datasets that are re\-packaged, both the original license and the license of the derived asset \(if it has changed\) should be provided\. - •If this information is not available online, the authors are encouraged to reach out to the asset’s creators\.
61. 13\.New assets
62. Question: Are new assets introduced in the paper well documented and is the documentation provided alongside the assets?
63. Answer:\[Yes\]
64. Justification: We introduce two new evaluation prompt sets \(a parametric prompt grid and a 960\-prompt investment coherence questionnaire\); the prompt sources, generation rules, and licenses will be released as part of the camera\-ready supplementary materials\.
65. Guidelines: - •The answer\[N/A\]means that the paper does not release new assets\. - •Researchers should communicate the details of the dataset/code/model as part of their submissions via structured templates\. This includes details about training, license, limitations, etc\. - •The paper should discuss whether and how consent was obtained from people whose asset is used\. - •At submission time, remember to anonymize your assets \(if applicable\)\. You can either create an anonymized URL or include an anonymized zip file\.
66. 14\.Crowdsourcing and research with human subjects
67. Question: For crowdsourcing experiments and research with human subjects, does the paper include the full text of instructions given to participants and screenshots, if applicable, as well as details about compensation \(if any\)?
68. Answer:\[N/A\]
69. Justification:\[N/A\]
70. Guidelines: - •The answer\[N/A\]means that the paper does not involve crowdsourcing nor research with human subjects\. - •Including this information in the supplemental material is fine, but if the main contribution of the paper involves human subjects, then as much detail as possible should be included in the main paper\. - •According to the NeurIPS Code of Ethics, workers involved in data collection, curation, or other labor should be paid at least the minimum wage in the country of the data collector\.
71. 15\.Institutional review board \(IRB\) approvals or equivalent for research with human subjects
72. Question: Does the paper describe potential risks incurred by study participants, whether such risks were disclosed to the subjects, and whether Institutional Review Board \(IRB\) approvals \(or an equivalent approval/review based on the requirements of your country or institution\) were obtained?
73. Answer:\[N/A\]
74. Justification:\[N/A\]
75. Guidelines: - •The answer\[N/A\]means that the paper does not involve crowdsourcing nor research with human subjects\. - •Depending on the country in which research is conducted, IRB approval \(or equivalent\) may be required for any human subjects research\. If you obtained IRB approval, you should clearly state this in the paper\. - •We recognize that the procedures for this may vary significantly between institutions and locations, and we expect authors to adhere to the NeurIPS Code of Ethics and the guidelines for their institution\. - •For initial submissions, do not include any information that would break anonymity \(if applicable\), such as the institution conducting the review\.
76. 16\.Declaration of LLM usage
77. Question: Does the paper describe the usage of LLMs if it is an important, original, or non\-standard component of the core methods in this research? Note that if the LLM is used only for writing, editing, or formatting purposes and does*not*impact the core methodology, scientific rigor, or originality of the research, declaration is not required\.
78. Answer:\[Yes\]
79. Justification: LLMs are used in two methodologically relevant places, both fully documented in the appendices\. First, the minimally\-framed contrastive datasets \(DexplicitD\_\{\\text\{explicit\}\}andDimplicitD\_\{\\text\{implicit\}\}\) were generated withClaude Sonnet 4\.6and validated byClaude Sonnet 4\.6together withGemini 3 Flashacross four dimensions \(lexical confounds, surface form, semantic confounds, content validity\); pairs scoring below threshold were iteratively revised or discarded, and A/B presentation was randomized to control for positional bias \([E\.1](https://arxiv.org/html/2606.05194#A5.SS1)\)\. Second, open\-ended steering generations were scored by an LLM\-as\-judge: each response was rated on a\[−10,10\]\[\-10,10\]short\-term/long\-term axis byClaude Sonnet 4\.6against predefined grading criteria \(explicit temporal keywords, structural planning, thematic bias\), and these scores feed the steering evaluation reported in Section[5\.3](https://arxiv.org/html/2606.05194#S5.SS3)\([S\.3](https://arxiv.org/html/2606.05194#A19.SS3)\)\.
80. Guidelines: - •The answer\[N/A\]means that the core method development in this research does not involve LLMs as any important, original, or non\-standard components\. - •Please refer to our LLM policy in the NeurIPS handbook for what should or should not be described\.Similar Articles
Temporal Context Reinstatement Drives Episodic-Like Order Memory in Long-Context Language Models
This paper investigates whether long-context language models exhibit episodic-like order memory similar to humans, using a temporal order memory task based on a novel. The authors find that models show the same distance effect and uncover a single attention head reinstating temporal context, supporting temporal context reinstatement as a mechanism for episodic-like memory in LLMs.
LLMs Contain Multitudes: How Deployment Context Reshapes Model-Level Preferences and Values
This paper investigates whether large language models have stable preferences across different deployment contexts, finding that context can cause larger variations than prompt perturbations, suggesting that measured preferences are context-conditioned rather than fixed properties.
RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization
This paper introduces RLearner-LLM, a framework using Hybrid-DPO to balance logical correctness and fluency in LLM-generated explanations, achieving significant NLI entailment improvements across multiple domains and base models while mitigating the verbosity bias of standard preference signals.
Why We Need World Models for AGI: Where LLMs Fail and How World Models May Outperform
This paper argues that large language models struggle with causal reasoning and long-horizon planning due to a mismatch between sequence prediction and reasoning over latent environment dynamics, and introduces the Latent Dynamics Inference perspective along with the Flux environment to study these limitations.
Do Language Models Know What Not to Say? Causal Evidence for Statistical Preemption in LLMs
This paper provides causal evidence that large language models acquire negative linguistic knowledge (what not to say) through statistical preemption, a mechanism from Construction Grammar, by showing that manipulating competing-form frequencies via fine-tuning shifts preemption behavior in predicted directions.