Predictive Memory Localization: Forecasting Selective Intervention Paths from Internal Signals
Summary
This paper introduces Predictive Memory Localization (PML), a method that forecasts selective intervention outcomes from internal model signals, distinguishing calibrated target movement from semantic-neighbor and capability damage. Experiments across 3,000 records and nine datasets show improved selective-path prediction and risk-aware intervention decisions.
View Cached Full Text
Cached at: 08/14/26, 09:28 AM
# Predictive Memory Localization: Forecasting Selective Intervention Paths from Internal Signals
Source: [https://arxiv.org/html/2608.12892](https://arxiv.org/html/2608.12892)
Tian ZeyuLucas Qingyang FangZhisheng ChenShuang ChenYuhao LuoQiannian Zhao\\corresponding
###### Abstract
Activation steering turns localized representations into control directions, but localization alone does not reveal whether a direction has a selective operating regime\. We introduce*Predictive Memory Localization*\(PML\), which treats the measured\-grid intervention path as the predictive object of memory localization\. PML separates random\-calibrated target movement from semantic\-neighbor and capability damage, and compares static localization and supervised geometry with a strength\-disjoint low\-dose causal response\. Our frozen study covers 3,000 records from nine datasets and fourteen domains, yielding 30,000 distinct record–direction–layer paths and 210,000 distinct path–strength evaluations\. At layer 7, the geometry\-derived RFM/AGOP direction reaches 13\.1% target\-any and 12\.3% clean\-any, exceeding random by 3\.6 and 3\.4 percentage points under a record\-paired bootstrap\. Across record\-, dataset\-, and domain\-grouped splits, responses at\|α\|=0\.1\|\\alpha\|=0\.1are the strongest signal for outcomes at disjoint strengths\|α\|∈\{0\.25,0\.5\}\|\\alpha\|\\in\\\{0\.25,0\.5\\\}\. On held\-out records, a predictor\-driven selector chooses a coefficient or abstains, improves utility and reduces semantic\-neighbor damage relative to a train\-tuned fixed\-strength policy, and avoids most evaluations in a dense scan\. Across three residual\-norm\-matched base models, learned directions retain selective\-path gains and low\-dose responses yield0\.8010\.801–0\.8280\.828record\-held\-out macro AUROC\. PML therefore turns memory localization into a falsifiable forecast of margin\-level selective outcomes and a risk\-aware intervention decision\.
## 1Introduction
Language\-model internals have been associated with feed\-forward memories, knowledge neurons, and causally important states\([7](https://arxiv.org/html/2608.12892#bib.bib1);[3](https://arxiv.org/html/2608.12892#bib.bib2);[14](https://arxiv.org/html/2608.12892#bib.bib4)\)\. This literature suggests an operational hypothesis: once information is localized, the corresponding representation should provide a useful control point\. Activation steering tests that hypothesis without changing model parameters\([26](https://arxiv.org/html/2608.12892#bib.bib10);[22](https://arxiv.org/html/2608.12892#bib.bib11);[37](https://arxiv.org/html/2608.12892#bib.bib14);[23](https://arxiv.org/html/2608.12892#bib.bib19)\)\. Yet behavior varies across prompts, layers, models, and coefficients, and stronger interventions can damage specificity or unrelated capabilities\([12](https://arxiv.org/html/2608.12892#bib.bib12);[24](https://arxiv.org/html/2608.12892#bib.bib13);[25](https://arxiv.org/html/2608.12892#bib.bib20);[8](https://arxiv.org/html/2608.12892#bib.bib18)\)\. Steering is thus a multidimensional control problem\([31](https://arxiv.org/html/2608.12892#bib.bib16);[19](https://arxiv.org/html/2608.12892#bib.bib21);[33](https://arxiv.org/html/2608.12892#bib.bib22)\)\.
Localization describes a representation, whereas controllability is revealed across a*measured intervention path*\. Target leverage can coexist with semantic\-neighbor or capability damage, and a clean effect at one coefficient need not persist at another\. Predicting calibrated target, damage, and clean outcomes over a pre\-specified grid asks which localized directions admit a usable measured coefficient\. A single endpoint cannot answer this question: the same direction may show leverage at one measured strength, damage at another, or no coefficient where the two separate cleanly\.
We introduce*Predictive Memory Localization*\(PML\), which maps each record, direction, and layer to those random\-calibrated outcomes\. It organizes predictive evidence from baseline belief and metadata, static localization, and supervised representation geometry to a strength\-disjoint low\-dose causal response\. This design directly tests whether static evidence forecasts later behavior or whether an inexpensive causal measurement is required\. A collateral\-aware policy then selects a coefficient or abstains\. Unlike generation\-level concept prediction\([5](https://arxiv.org/html/2608.12892#bib.bib24)\), PML audits localized factual and reasoning directions with explicit collateral probes and strength\-disjoint labels\.
Figure 1:PML from representation to selective control\. \(a\) Direction estimators use desired and contrast activations at a selected block, where the resulting direction is injected\. \(b\) Static evidence and a disjoint low\-dose response forecast target benefit, collateral risk, and utility over candidate coefficients for selection or abstention\. Curves are schematic\.On 3,000 frozen records from nine datasets and fourteen domains, learned directions improve target and clean\-path incidence over random\. Static localization adds little predictive value, whereas a disjoint low\-dose response dominates later\-path prediction across grouped transfer\. Both findings replicate across three residual\-norm\-matched base models \(0\.801–0\.828 record\-held\-out macro AUROC\)\. A held\-out selector improves utility and reduces neighbor damage relative to a fixed intervention\. The central result is thata small causal response is the most useful forecast of later selective behavior\.
This work makes three contributions:
- •We define a random\-calibrated measured\-grid path that jointly records target leverage, semantic\-neighbor damage, capability damage, and clean measured coefficients for each record, direction, and layer\.
- •We separate static localization and supervised geometry from a strength\-disjoint causal probe, showing that the former are weak predictive priors while the latter supplies the dominant signal across grouped and cross\-model evaluation\.
- •We connect forecasting to action with a held\-out collateral\-aware policy that selects one coefficient or abstains, providing a proof of concept for reducing dense intervention evaluation\.
## 2Related Work
#### Knowledge representation, localization, and editing\.
Transformer internals have been associated with feed\-forward memories, knowledge neurons, and causally important states\([7](https://arxiv.org/html/2608.12892#bib.bib1);[3](https://arxiv.org/html/2608.12892#bib.bib2);[14](https://arxiv.org/html/2608.12892#bib.bib4)\)\. Editing uses learned editors, external memory, or targeted parameter updates\([4](https://arxiv.org/html/2608.12892#bib.bib3);[17](https://arxiv.org/html/2608.12892#bib.bib5);[18](https://arxiv.org/html/2608.12892#bib.bib6);[15](https://arxiv.org/html/2608.12892#bib.bib7)\), with efficacy, generalization, and locality organized by surveys and toolkits\([35](https://arxiv.org/html/2608.12892#bib.bib25);[27](https://arxiv.org/html/2608.12892#bib.bib26)\)\. Localization need not identify components that support editing or unlearning\([9](https://arxiv.org/html/2608.12892#bib.bib8);[11](https://arxiv.org/html/2608.12892#bib.bib9)\); PML tests its predictive connection to intervention outcomes\.
#### Activation steering and control tradeoffs\.
Activation engineering intervenes on intermediate representations\([26](https://arxiv.org/html/2608.12892#bib.bib10)\); contrastive activation addition, representation engineering, and representation surgery construct or analyze steering directions from examples and population structure\([22](https://arxiv.org/html/2608.12892#bib.bib11);[37](https://arxiv.org/html/2608.12892#bib.bib14);[23](https://arxiv.org/html/2608.12892#bib.bib19)\)\. Other work studies truthfulness components, coefficient scaling, and preference–utility tradeoffs\([12](https://arxiv.org/html/2608.12892#bib.bib12);[24](https://arxiv.org/html/2608.12892#bib.bib13);[32](https://arxiv.org/html/2608.12892#bib.bib23)\)\. AxBench compares concept detection and steering methods\([31](https://arxiv.org/html/2608.12892#bib.bib16)\), while broader evaluations expose sensitivity to instruction phrasing, prompt distribution, layer choice, and model scale\([25](https://arxiv.org/html/2608.12892#bib.bib20);[2](https://arxiv.org/html/2608.12892#bib.bib17);[8](https://arxiv.org/html/2608.12892#bib.bib18)\)\. These studies motivate joint leverage and side\-effect measurement; PML additionally asks whether localization evidence forecasts their coexistence for each record–direction–layer path\.
#### Evaluation and prediction of steerability\.
Reliable steering evaluations increasingly separate behavioral success from coherence, specificity, and unintended change\([19](https://arxiv.org/html/2608.12892#bib.bib21);[33](https://arxiv.org/html/2608.12892#bib.bib22)\)\. Most closely related,[5](https://arxiv.org/html/2608.12892#bib.bib24)predict under\-, successful, or over\-steering from first\-token dynamics and rank strengths to reduce rollouts\. PML instead predicts a record–direction–layer path on a pre\-specified coefficient grid, calibrates target, semantic\-neighbor, and capability events against random directions, and optimizes collateral\-aware margin utility\. Its weak feature is measured at a coefficient excluded from the stronger\-coefficient labels, after which a policy chooses one coefficient or abstains\. SteerBoost targets efficient generation\- level concept alignment; PML tests whether localization and a fixed low\-dose diagnostic forecast selective margin outcomes\.
#### Supervised geometry as a diagnostic\.
Recursive feature machines use the average gradient outer product to learn task\-adapted geometry\([21](https://arxiv.org/html/2608.12892#bib.bib15)\)\. PML uses its leading direction, spectral concentration, and alignment as candidate signals, while explicitly testing rather than assuming that supervised sensitivity implies selective control\.
## 3Predictive Memory Localization
PML connects a localized residual representation to a measured strength\-indexed intervention path and a held\-out decision \(Figure[1](https://arxiv.org/html/2608.12892#S1.F1)\)\. Here, “memory localization” denotes a record\-conditioned internal signal associated with an answer relation, not a claim that the relation resides in one unique physical component\.
### 3\.1Records, Probes, and Interventions
Recordiicontains target prompts𝒯i\\mathcal\{T\}\_\{i\}, semantic\-neighbor prompts𝒩i\\mathcal\{N\}\_\{i\}, and general\-capability prompts𝒞i\\mathcal\{C\}\_\{i\}\. The first set expresses the behavior to suppress or enhance; the latter two test whether nearby knowledge or unrelated abilities are preserved\. For promptxxwith desired answery\+y^\{\+\}and contrasty−y^\{\-\}, the answer margin is
m\(x\)=logp\(y\+∣x\)−logp\(y−∣x\)\.m\(x\)=\\log p\(y^\{\+\}\\mid x\)\-\\log p\(y^\{\-\}\\mid x\)\.\(1\)All outcomes are changes from the unperturbed margin, making paths comparable across records with different baseline confidence\.
At layerℓ\\ell, methodssconstructs a unit\-norm record\-specific direction𝐯iℓs\\mathbf\{v\}\_\{i\\ell s\}\. Intervention strengthα\\alphamodifies the residual activation as
𝐡ℓ′=𝐡ℓ\+α𝐯iℓs\.\\mathbf\{h\}^\{\\prime\}\_\{\\ell\}=\\mathbf\{h\}\_\{\\ell\}\+\\alpha\\mathbf\{v\}\_\{i\\ell s\}\.\(2\)Negative and positive coefficients test suppression and enhancement with the same direction\. We writeΔiT\(α\)\\Delta^\{T\}\_\{i\}\(\\alpha\)for signed target\-margin movement andΔiN\(α\),ΔiC\(α\)\\Delta^\{N\}\_\{i\}\(\\alpha\),\\Delta^\{C\}\_\{i\}\(\\alpha\)for neighbor and capability movement; a sufficiently negative collateral change is damage\.
### 3\.2Random\-Calibrated Intervention Paths
Target, neighbor, and capability responses have different null scales, so random directions define separate 95th\-percentile thresholdsτT\\tau\_\{T\},τN\\tau\_\{N\}, andτC\\tau\_\{C\}\. A target effect crossesτT\\tau\_\{T\}in the intended direction, while neighbor or capability damage crosses the corresponding negative threshold\. A*clean strength*produces a target effect while crossing neither damage threshold at that same coefficient\.
Across an ordered, finite strength set, these events form a measured\-grid intervention path\. Target and collateral onset identify the first observed crossing, and adjacent clean strengths form an observed clean region\. These are descriptive grid statistics, not estimates of an unmeasured continuous window\. The confirmatory prediction task focuses on four directly measured events at held\-out stronger coefficients:*Target\-any*, semantic\-neighbor damage, capability damage, and*Clean\-any*\. Percentages count record–method–layer paths rather than prompts or individual strengths\.
### 3\.3Prediction Task
PML forecasts later path outcomes from progressively stronger evidence\. Baseline beliefBBand metadataMMare augmented with static localization featuresLLand, for RFM, geometry featuresGG\. Low\-dose response featuresRRmeasure target and collateral movement atα=±0\.1\\alpha=\\pm 0\.1\. BecauseRRrequires intervention, it is a cheap dynamic diagnostic rather than static localization\. The central comparisons test whetherLLorGGimproves onB\+MB\+M, and whether the strength\-disjoint responseRRforecasts outcomes at\|α\|∈\{0\.25,0\.5\}\|\\alpha\|\\in\\\{0\.25,0\.5\\\}\. A downstream policy then uses these forecasts to select a coefficient or abstain\.
## 4Signals and Predictive Models
PML evaluates a common path\-prediction and decision pipeline over several direction families, separating the quality of a direction from the evidence used to forecast its later behavior\.
### 4\.1Direction Families
#### Random control\.
A seeded unit\-norm Gaussian direction defines response thresholds and the paired null baseline\. The loggedmatched\_norm\_randomentry is numerically identical and retained only for auditability\.
#### Mean difference\.
For positive and contrast activation setsℋ\+\\mathcal\{H\}^\{\+\}andℋ−\\mathcal\{H\}^\{\-\}, we use
𝐯mean=𝝁\+−𝝁−‖𝝁\+−𝝁−‖2,\\mathbf\{v\}\_\{\\mathrm\{mean\}\}=\\frac\{\\boldsymbol\{\\mu\}^\{\+\}\-\\boldsymbol\{\\mu\}^\{\-\}\}\{\\\|\\boldsymbol\{\\mu\}^\{\+\}\-\\boldsymbol\{\\mu\}^\{\-\}\\\|\_\{2\}\},\(3\)the multi\-example analogue of activation addition\([26](https://arxiv.org/html/2608.12892#bib.bib10);[22](https://arxiv.org/html/2608.12892#bib.bib11)\)\.
#### Linear and logistic probes\.
Normalized classifier weights provide supervised discriminative directions without nonlinear feature learning\.
#### RFM/AGOP direction\.
A recursive feature machine estimates task\-adapted geometry through the average gradient outer product\([21](https://arxiv.org/html/2608.12892#bib.bib15)\)\. Its leading eigenvector provides a low\-rank direction, while spectrum, concentration, and alignment statistics become geometry features\. A top\-kkboundary study tests whether broader supervised subspaces trade selectivity for leverage\.
All directions are fitted independently per record and layer and injected at the selected residual block during evaluation\. Detailed activation extraction, fitting hyperparameters, inference settings, and continuation scoring are included in the released protocol\.
### 4\.2Predictive Evidence and Grouped Evaluation
Static localizationLLsummarizes class separation, projection, saliency, and direction agreement; RFM geometryGGsummarizes spectral concentration and alignment\. These features are nested with baseline beliefBB, metadataMM, and the low\-dose responseRRso that method identity or an observed response cannot be misattributed to static localization\.
We fit class\-balanced logistic regression and random forests for target, damage, and clean outcomes\. Five\-fold evaluation groups all paths from the same record, dataset, or domain, preventing related trajectories from crossing a train–test boundary\. We report prevalence, AUROC, and average precision in the main paper; additional classification and calibration metrics are in the released artifacts\.
### 4\.3Risk\-Aware Strength Selection
For each nonzero candidate coefficient, the selector predicts clean\-effect probability and continuous utility\. Withq\(α\)∈\{−1,\+1\}q\(\\alpha\)\\in\\\{\-1,\+1\\\}denoting the requested suppression or enhancement sign, measured utility is
ui\(α\)=\\displaystyle u\_\{i\}\(\\alpha\)=\{\}q\(α\)ΔiT\(α\)−max\{0,−ΔiN\(α\)\}\\displaystyle q\(\\alpha\)\\Delta^\{T\}\_\{i\}\(\\alpha\)\-\\max\\\{0,\-\\Delta^\{N\}\_\{i\}\(\\alpha\)\\\}\(4\)−max\{0,−ΔiC\(α\)\}\.\\displaystyle\-\\max\\\{0,\-\\Delta^\{C\}\_\{i\}\(\\alpha\)\\\}\.Target movement is rewarded and collateral margin decreases receive unit penalties\. The decision score is
u^i\(α\)\+0\.1p^i\(clean∣α\)−0\.01\|α\|\.\\widehat\{u\}\_\{i\}\(\\alpha\)\+0\.1\\widehat\{p\}\_\{i\}\(\\mathrm\{clean\}\\mid\\alpha\)\-0\.01\|\\alpha\|\.\(5\)The policy selects the highest\-scoring coefficient and abstains when its score is nonpositive\. This utility is defined on teacher\-forced answer\-margin changes, not free\-generation correctness\. We evaluate it against both no intervention, whose utility is zero by construction, and a train\-tuned fixed coefficient\.
## 5Experiments
We organize the evidence around three questions: whether learned directions create selective intervention paths, which internal signals forecast those paths, and whether those forecasts improve strength decisions\. We report the frozen confirmatory study here; the Supplementary Material adds protocol details, uncertainty estimates, and analyses not shown in the main paper\.
### 5\.1Frozen Multidomain Study
The benchmark contains 3,000 records from nine public sources: MMLU\-Pro, MMLU\-Redux 2\.0, AI2 ARC, OpenBookQA, SciQ, LiveBench reasoning and math, HellaSwag, and QASC\([28](https://arxiv.org/html/2608.12892#bib.bib27);[6](https://arxiv.org/html/2608.12892#bib.bib28);[1](https://arxiv.org/html/2608.12892#bib.bib29);[16](https://arxiv.org/html/2608.12892#bib.bib30);[29](https://arxiv.org/html/2608.12892#bib.bib31);[30](https://arxiv.org/html/2608.12892#bib.bib32);[36](https://arxiv.org/html/2608.12892#bib.bib33);[10](https://arxiv.org/html/2608.12892#bib.bib34)\)\. The collection spans fourteen academic, scientific, commonsense, mathematical, and reasoning domains\. A schema\-constrained generation step converts each source item into disjoint direction\-fitting statements, three record\-specific target probes, and three semantic\-neighbor probes; four globally balanced capability probes are assigned per record\. The worked example below traces one source item from fitting evidence to evaluation probes and its resulting path label\. Dataset composition, validation, and additional examples appear in the Supplementary Material\.
Worked example\.One frozen record from construction to path label; direction\-fitting statements and evaluation probes are disjoint\.
We evaluate Qwen3\-1\.7B\-Base\([34](https://arxiv.org/html/2608.12892#bib.bib35)\)at two middle\-depth Transformer blocks, reported as blocks 7 and 11 by the model implementation and selected using a 500\-record layer\-selection subset \(Table[2](https://arxiv.org/html/2608.12892#S5.T2)\)\. We compare five direction constructions over a signed coefficient sweep\. Random directions calibrate target and collateral thresholds, while the weak response at\|α\|=0\.1\|\\alpha\|=0\.1is disjoint from the stronger coefficients used to define confirmatory outcomes\. Table[1](https://arxiv.org/html/2608.12892#S5.T1)defines the reported path rates, and prediction folds hold out complete records, datasets, or domains\.
For cross\-model confirmation, the same frozen 500\-record subset is evaluated on Qwen3\-1\.7B, Qwen3\.5\-2B\-Base\([20](https://arxiv.org/html/2608.12892#bib.bib36)\), and Ministral\-3\-3B\-Base\([13](https://arxiv.org/html/2608.12892#bib.bib37)\)\. We align relative intervention budget‖αv‖2/‖h‖2\\\|\\alpha v\\\|\_\{2\}/\\\|h\\\|\_\{2\}using
sm,ℓ=RMS\(hm,ℓ\)RMS\(href,ℓ\)dmdref,αm,ℓ=sm,ℓαref\.s\_\{m,\\ell\}=\\frac\{\\operatorname\{RMS\}\(h\_\{m,\\ell\}\)\}\{\\operatorname\{RMS\}\(h\_\{\\mathrm\{ref\},\\ell\}\)\}\\sqrt\{\\frac\{d\_\{m\}\}\{d\_\{\\mathrm\{ref\}\}\}\},\\qquad\\alpha\_\{m,\\ell\}=s\_\{m,\\ell\}\\alpha\_\{\\mathrm\{ref\}\}\.\(6\)Scales are fixed on a disjoint 100\-record calibration set, after which each model recalibrates its random\-response thresholds on the formal cohort\. The confirmation retains random, mean\-difference, logistic, and RFM/AGOP directions at two pre\-specified blocks per model\. We omit linear because it is not strongest at either primary\-study block, which preserves a symmetric comparison across architectures\.
Table 2:Layer selection on a 500\-record subset\. Target\-margin span across relative depth identifies blocks 7 and 11 as the strongest distinct middle\-depth blocks; random spans remain near zero\.Table 1:Random\-calibrated path incidence \(%\) for the 3,000\-record primary study and three 500\-record confirmations\. N/C\-dmg\. are collateral damage;Δ\\Deltacolumns are learned\-minus\-random points\. Bold marks each block’s strongest learned result\.
### 5\.2Learned Directions Improve Selective Paths
Table[1](https://arxiv.org/html/2608.12892#S5.T1)reports both the powered 3,000\-record outcome estimate and the three pre\-specified 500\-record confirmations\. Its entries are path\-incidence percentages;Δ\\DeltaT andΔ\\DeltaC are percentage\-point differences from the within\-model random control, not AUROC\. In the primary study, layer\-7 RFM/AGOP reaches13\.1% Target and 12\.3% Clean, the strongest selective\-path result\. More broadly, learned directions improve target leverage and clean\-path incidence over random controls, although the best construction depends on the block\.
Record\-paired bootstrap intervals confirm the learned\-over\-random Target and Clean gains for the strongest primary\-study directions\. Collateral\-damage intervals include zero, so the evidence supports more usable intervention paths rather than a universal reduction in every form of collateral movement\. Full intervals are reported in the supplement\.
The residual\-norm\-matched confirmations preserve the same qualitative pattern at model\-dependent magnitudes\. Figure[2](https://arxiv.org/html/2608.12892#S5.F2)shows that learned directions generally move upward from their random controls in target leverage, but not uniformly leftward toward lower collateral incidence\. The Qwen3\-1\.7B subset also closely tracks the powered estimate, separating cohort variation from the cross\-model scale alignment\.
### 5\.3Low\-Dose Responses Forecast Later Outcomes
Prediction is the central test of PML\. Static localization is a weak prior: Figure[3](https://arxiv.org/html/2608.12892#S5.F3)establishes a consistent diagnostic hierarchy\. Static localizationLLadds little beyond base and metadata features, and supervised geometryGGdoes not change that conclusion\. In contrast, the strength\-disjoint weak responseRRproduces the dominant gain across outcomes and models\. Table[3](https://arxiv.org/html/2608.12892#S6.T3)further separates this gain from ordinary scalar weak\-to\-strong correlation\. Positive\-label prevalence is only 8–11%, so the table reports AP with AUROC\. For Target\-any and Clean\-any, a multivariateRR\-only random forest substantially improves on a single signed response, with the complete predictor adding a smaller final gain\. Final macro AUROC remains around0\.80–0\.85under record\-, dataset\-, and domain\-held\-out evaluation\. Removing all 500 records used for layer selection leaves the four principal full\-predictor AUROCs within 0\.01 of the 3,000\-record estimates\. Thus, observing a structured low\-dose causal response is substantially more informative about the later intervention path than static localization geometry alone, and this conclusion is not explained by the layer\-selection overlap\.
Figure 2:Target leverage and collateral incidence on the common 500\-record cohort\. Each panel shows one base model; colors identify direction construction and marker shapes identify the shallower or deeper pre\-specified block\. The vertical axis is Target\-any path incidence, while the horizontal axis averages semantic\-neighbor and capability\-damage incidence\. Points are not connected because the two blocks are separately pre\-specified evaluations rather than a continuous trajectory\. Axes are panel\-specific so that within\-model trade\-offs remain visible\.Figure 3:Cross\-model path prediction\. Per\-outcome record\-held\-out AUROC on the common 500\-record cohort\. The first three random\-forest columns cumulatively add base margins and record descriptors \(BB\), method/block metadata \(MM\), and static localization \(LL\)\. The final two columns use the completeB\+M\+L\+RB\{\+\}M\{\+\}L\{\+\}Rfeatures with a linear predictor or random forest, providing a predictor\-class ablation\. The weak response produces the dominant feature gain under both predictors, while the nonlinear model gives the strongest final performance\.
### 5\.4Forecasts Improve Strength Decisions
We train candidate\-strength models from complementary dense and sparse trajectories, then evaluate on disjoint dense\-grid records\. Before a candidate outcome is revealed, each policy selects one coefficient or abstains\. Table[4](https://arxiv.org/html/2608.12892#S6.T4)gives a 100\-record proof of concept\. Relative to a train\-tuned fixed coefficient, utility improves by0\.055and0\.034, mainly as neighbor damage falls from 6\.0% to 1\.8% and 5\.2% to 2\.2%\. Suppression exceeds no intervention; enhancement does not significantly do so\. The policy averages 2\.55/2\.61 coefficient evaluations—two weak probes plus a final action—versus 26 for a dense scan; shared direction fitting is excluded\. Weight sensitivity and full outcomes are supplementary\.
### 5\.5Endpoint Scope
Free\-generation stress tests are reported only in the Supplementary Material\. They show that margin movement can alter text but does not yield reliable wrong\-to\-right correction; PML’s primary evidence is therefore margin\-level, not a claim of stable generated\-answer control\.
## 6Discussion
#### Measured paths separate leverage from selectivity\.
Learned directions increase Target and Clean incidence, yet their neighbor\- and capability\-damage differences remain statistically unresolved\. A direction can therefore create more usable measured coefficients without becoming uniformly safer\. Layer 7 similarly offers greater leverage together with more collateral movement than layer 11\. The relevant object is not maximal sensitivity but the coexistence of target and damage responses on the same pre\-specified grid\. Measured clean regions identify where leverage and selectivity coincide, but they should not be read as broad continuous operating windows: most observed clean paths contain only one measured clean coefficient, and strict target\-first or damage\-first orderings are rare\. These topology statistics are therefore diagnostics that motivate multi\-strength evaluation rather than the primary prediction labels\.
Table 3:Record\-held\-out prediction on Qwen3\-1\.7B\-Base\. Cells are AUROC/AP; Prev\. is label prevalence\. Scalar is training\-free,RR\-only uses the multivariate weak profile, and Full usesB\+M\+L\+RB\+M\+L\+R\.Table 4:Held\-out decisions on 100 records and 900 paths per objective\. Utility differences use a paired record bootstrap; N\-dmg\. is fixed→\\toselector\.
#### Budget matching preserves the diagnostic hierarchy\.
Residual\-norm matching aligns relative intervention magnitude, not response rates: Qwen3\.5 shows the largest gains and Ministral is intermediate\. All three models nevertheless reproduce more learned clean paths and a much larger predictive contribution from low\-dose response than static localization\.
#### A small causal response is more actionable than static localization\.
Supervised geometry is valuable for constructing high\-leverage directions and describing their concentration and alignment, but these static descriptors add little forecasting power by themselves\. In contrast, a strength\-disjoint low\-dose response strongly predicts later outcomes across record, dataset, and domain transfer\. This suggests that localization becomes actionable when it is paired with a cheap causal measurement of the specific path, rather than when static separation is treated as sufficient evidence of control\. PML therefore complements concept detection and average steering evaluation\([31](https://arxiv.org/html/2608.12892#bib.bib16);[2](https://arxiv.org/html/2608.12892#bib.bib17);[8](https://arxiv.org/html/2608.12892#bib.bib18)\)by testing whether an internal signal supports effective and selective margin movement at a chosen strength\.
The weak, outcome\-specific transfer of activation\-path features to ROME marks a second boundary\. A path can diagnose a fragile or promising activation intervention without identifying the best persistent parameter edit\([9](https://arxiv.org/html/2608.12892#bib.bib8)\); PML does not equate activation controllability with editability\.
## 7Limitations
The primary study uses Qwen3\-1\.7B and two blocks selected on a 500\-record subset\. Width\-corrected residual\-norm confirmations add Qwen3\.5\-2B and Ministral\-3\-3B at two pre\-specified aligned blocks, but they support claims about those relative depths rather than global layer optimality\. Because every model recalibrates its own random null, cross\-model evidence establishes within\-model contrasts and predictor ordering, not absolute incidence comparisons across architectures\. The Qwen3\-1\.7B subset differs slightly from the 3,000\-record estimate because of sampling and threshold recalibration\.
The primary outcomes are teacher\-forced margins at two held\-out stronger coefficients\. Finite\-grid topology and free\-generation stress tests are supplementary; the latter do not establish reliable wrong\-to\-right control\. Schema\-constrained probes receive an independent expert audit and exclusion sensitivity, but independently authored validation remains future work\. The paired differences in neighbor and capability damage remain unresolved, and Static localization and AGOP geometry also provide limited incremental prediction once a weak response is observed\. Their present value is therefore structural—direction construction, concentration, and alignment diagnostics—rather than a standalone guarantee of path quality\.
The response thresholds are frozen global percentiles from random\-direction controls\. This supplies a common within\-study null, but it does not establish that the same numeric thresholds transport to a new model or domain\. Likewise, the strength policy optimizes one declared utility with unit collateral penalties, a 0\.1 clean bonus, and a 0\.01 magnitude penalty\. Weight sensitivity preserves gains over the fixed policy, not universally over no intervention\.
Finally, the 100\-record selector is a proof of concept\. Its cost reduction counts coefficient evaluations—two weak probes and a selected action—while excluding shared direction fitting and feature extraction\. Neighbor and capability probes operationalize two collateral channels but cannot exhaust downstream side effects; the small ROME transfer study is also insufficient for claims about general editing success or model\-wide safety\.
## 8Conclusion
Predictive Memory Localization forecasts measured\-grid target, neighbor, capability, and clean margin outcomes\. On 3,000 records, learned directions improve Target and Clean incidence over random; at block 7, RFM/AGOP reaches 13\.1% Target and 12\.3% Clean\. A strength\-disjoint low\-dose response dominates static localization for later\-outcome prediction, and the diagnostic hierarchy replicates across three matched base models\. A held\-out selector then improves utility over a fixed coefficient, reduces neighbor damage, and replaces a dense scan with two weak probes and a selected action or abstention\. PML thus converts a static localization claim into a falsifiable forecast and a risk\-aware decision while making its margin\-level and finite\-grid scope explicit\.
The broader implication is that representation evidence and intervention quality are related but distinct\. A direction can be well localized without providing a selective operating regime, whereas a small causal response can reveal leverage and collateral risk before a stronger action\. Random\-direction calibration makes this distinction measurable and exposes cases where abstention is appropriate\.
## References
- Clarket al\.\(2018\)P\. Clark, I\. Cowhey, O\. Etzioni, T\. Khot, A\. Sabharwal, C\. Schoenick, and O\. TafjordThink you have solved question answering? try ARC, the AI2 reasoning challenge\.Note:arXiv preprint arXiv:1803\.05457External Links:1803\.05457,[Document](https://dx.doi.org/10.48550/arXiv.1803.05457)Cited by:[Appendix S1](https://arxiv.org/html/2608.12892#A1.p2.1),[§5\.1](https://arxiv.org/html/2608.12892#S5.SS1.p1.1)\.
- Da Silvaet al\.\(2025\)P\. Q\. Da Silva, H\. Sethuraman, D\. Rajagopal, H\. Hajishirzi, and S\. KumarSteering off course: reliability challenges in steering language models\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 19856–19882\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.974),[Link](https://aclanthology.org/2025.acl-long.974/)Cited by:[§2](https://arxiv.org/html/2608.12892#S2.SS0.SSS0.Px2.p1.1),[§6](https://arxiv.org/html/2608.12892#S6.SS0.SSS0.Px3.p1.1)\.
- Daiet al\.\(2022\)D\. Dai, L\. Dong, Y\. Hao, Z\. Sui, B\. Chang, and F\. WeiKnowledge neurons in pretrained transformers\.InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 8493–8502\.External Links:[Document](https://dx.doi.org/10.18653/v1/2022.acl-long.581),[Link](https://aclanthology.org/2022.acl-long.581/)Cited by:[§1](https://arxiv.org/html/2608.12892#S1.p1.1),[§2](https://arxiv.org/html/2608.12892#S2.SS0.SSS0.Px1.p1.1)\.
- De Caoet al\.\(2021\)N\. De Cao, W\. Aziz, and I\. TitovEditing factual knowledge in language models\.InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing,pp\. 6491–6506\.External Links:[Document](https://dx.doi.org/10.18653/v1/2021.emnlp-main.522),[Link](https://aclanthology.org/2021.emnlp-main.522/)Cited by:[§2](https://arxiv.org/html/2608.12892#S2.SS0.SSS0.Px1.p1.1)\.
- Fanet al\.\(2026\)Z\. Fan, Z\. Zhang, Q\. Xu, Y\. Cai, J\. Wang, F\. Wei, D\. He, Y\. Tang, Y\. Sun, and D\. TaoWhen is your LLM steerable?\.Note:arXiv preprint arXiv:2606\.11599External Links:2606\.11599,[Document](https://dx.doi.org/10.48550/arXiv.2606.11599),[Link](https://arxiv.org/abs/2606.11599)Cited by:[§1](https://arxiv.org/html/2608.12892#S1.p3.1),[§2](https://arxiv.org/html/2608.12892#S2.SS0.SSS0.Px3.p1.1)\.
- Gemaet al\.\(2024\)A\. P\. Gema, J\. O\. J\. Leang, G\. Hong, A\. Devoto, A\. C\. M\. Mancino, R\. Saxena, X\. He, Y\. Zhao, X\. Du, A\. Madotto, J\. Z\. K\. Lai, T\. Kocmi, A\. F\. Aji, K\. Heafield, T\. Baldwin, and A\. BirchAre we done with MMLU?\.Note:arXiv preprint arXiv:2406\.04127External Links:2406\.04127,[Document](https://dx.doi.org/10.48550/arXiv.2406.04127)Cited by:[Appendix S1](https://arxiv.org/html/2608.12892#A1.p2.1),[§5\.1](https://arxiv.org/html/2608.12892#S5.SS1.p1.1)\.
- Gevaet al\.\(2021\)M\. Geva, R\. Schuster, J\. Berant, and O\. LevyTransformer feed\-forward layers are key\-value memories\.InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing,pp\. 5484–5495\.External Links:[Document](https://dx.doi.org/10.18653/v1/2021.emnlp-main.446),[Link](https://aclanthology.org/2021.emnlp-main.446/)Cited by:[§1](https://arxiv.org/html/2608.12892#S1.p1.1),[§2](https://arxiv.org/html/2608.12892#S2.SS0.SSS0.Px1.p1.1)\.
- Goyal and Daumé III \(2026\)N\. Goyal and H\. Daumé IIISteering safely or off a cliff? rethinking specificity and robustness in inference\-time interventions\.InProceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 5723–5738\.External Links:[Document](https://dx.doi.org/10.18653/v1/2026.eacl-long.268),[Link](https://aclanthology.org/2026.eacl-long.268/)Cited by:[§1](https://arxiv.org/html/2608.12892#S1.p1.1),[§2](https://arxiv.org/html/2608.12892#S2.SS0.SSS0.Px2.p1.1),[§6](https://arxiv.org/html/2608.12892#S6.SS0.SSS0.Px3.p1.1)\.
- Haseet al\.\(2023\)P\. Hase, M\. Bansal, B\. Kim, and A\. GhandehariounDoes localization inform editing? surprising differences in causality\-based localization vs\. knowledge editing in language models\.InAdvances in Neural Information Processing Systems,Vol\.36,pp\. 17643–17668\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2023/hash/3927bbdcf0e8d1fa8aa23c26f358a281-Abstract-Conference.html)Cited by:[§2](https://arxiv.org/html/2608.12892#S2.SS0.SSS0.Px1.p1.1),[§6](https://arxiv.org/html/2608.12892#S6.SS0.SSS0.Px3.p2.1)\.
- Khotet al\.\(2020\)T\. Khot, P\. Clark, M\. Guerquin, P\. Jansen, and A\. SabharwalQASC: a dataset for question answering via sentence composition\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.34,pp\. 8082–8090\.External Links:[Document](https://dx.doi.org/10.1609/aaai.v34i05.6319)Cited by:[Appendix S1](https://arxiv.org/html/2608.12892#A1.p2.1),[§5\.1](https://arxiv.org/html/2608.12892#S5.SS1.p1.1)\.
- Leeet al\.\(2025\)H\. Lee, U\. Hwang, and G\. KimDoes localization inform unlearning? a rigorous examination of local parameter attribution for knowledge unlearning in language models\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 21809–21830\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.1109),[Link](https://aclanthology.org/2025.emnlp-main.1109/)Cited by:[§2](https://arxiv.org/html/2608.12892#S2.SS0.SSS0.Px1.p1.1)\.
- Liet al\.\(2023\)K\. Li, O\. Patel, F\. Viégas, H\. Pfister, and M\. WattenbergInference\-time intervention: eliciting truthful answers from a language model\.InAdvances in Neural Information Processing Systems,Vol\.36,pp\. 41451–41530\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2023/hash/81b8390039b7302c909cb769f8b6cd93-Abstract-Conference.html)Cited by:[§1](https://arxiv.org/html/2608.12892#S1.p1.1),[§2](https://arxiv.org/html/2608.12892#S2.SS0.SSS0.Px2.p1.1)\.
- Liuet al\.\(2026\)A\. H\. Liu, B\. Barbier, B\. Bose, A\. Cohen, R\. Cohendet, E\. Dupont, M\. K\. Eddine, L\. Fresson, L\. Grinsztajn,et al\.Ministral 3\.arXiv preprint arXiv:2601\.08584\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2601.08584)Cited by:[Appendix S3](https://arxiv.org/html/2608.12892#A3.p1.1),[§5\.1](https://arxiv.org/html/2608.12892#S5.SS1.p4.1)\.
- Menget al\.\(2022\)K\. Meng, D\. Bau, A\. Andonian, and Y\. BelinkovLocating and editing factual associations in GPT\.InAdvances in Neural Information Processing Systems,Vol\.35,pp\. 17359–17372\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2022/hash/6f1d43d5a82a37e89b0665b33bf3a182-Abstract-Conference.html)Cited by:[§1](https://arxiv.org/html/2608.12892#S1.p1.1),[§2](https://arxiv.org/html/2608.12892#S2.SS0.SSS0.Px1.p1.1)\.
- Menget al\.\(2023\)K\. Meng, A\. S\. Sharma, A\. Andonian, Y\. Belinkov, and D\. BauMass\-editing memory in a transformer\.InThe Eleventh International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=MkbcAHIYgyS)Cited by:[§2](https://arxiv.org/html/2608.12892#S2.SS0.SSS0.Px1.p1.1)\.
- Mihaylovet al\.\(2018\)T\. Mihaylov, P\. Clark, T\. Khot, and A\. SabharwalCan a suit of armor conduct electricity? a new dataset for open book question answering\.InProceedings of the 2018 Conference on Empirical Methods in Natural Language Processing,pp\. 2381–2391\.External Links:[Document](https://dx.doi.org/10.18653/v1/D18-1260)Cited by:[Appendix S1](https://arxiv.org/html/2608.12892#A1.p2.1),[§5\.1](https://arxiv.org/html/2608.12892#S5.SS1.p1.1)\.
- Mitchellet al\.\(2022a\)E\. Mitchell, C\. Lin, A\. Bosselut, C\. Finn, and C\. D\. ManningFast model editing at scale\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=0DcZxeWfOPt)Cited by:[§2](https://arxiv.org/html/2608.12892#S2.SS0.SSS0.Px1.p1.1)\.
- Mitchellet al\.\(2022b\)E\. Mitchell, C\. Lin, A\. Bosselut, C\. Finn, and C\. D\. ManningMemory\-based model editing at scale\.InProceedings of the 39th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.162,pp\. 15817–15831\.External Links:[Link](https://proceedings.mlr.press/v162/mitchell22a.html)Cited by:[§2](https://arxiv.org/html/2608.12892#S2.SS0.SSS0.Px1.p1.1)\.
- Preset al\.\(2024\)I\. Pres, L\. Ruis, E\. S\. Lubana, and D\. KruegerTowards reliable evaluation of behavior steering interventions in LLMs\.Note:arXiv preprint arXiv:2410\.17245External Links:2410\.17245,[Document](https://dx.doi.org/10.48550/arXiv.2410.17245),[Link](https://arxiv.org/abs/2410.17245)Cited by:[§1](https://arxiv.org/html/2608.12892#S1.p1.1),[§2](https://arxiv.org/html/2608.12892#S2.SS0.SSS0.Px3.p1.1)\.
- Qwen Team \(2026\)Qwen TeamQwen3\.5\-2B\-Base\.Note:Hugging Face model cardOfficial model release and architecture specificationExternal Links:[Link](https://huggingface.co/Qwen/Qwen3.5-2B-Base)Cited by:[Appendix S3](https://arxiv.org/html/2608.12892#A3.p1.1),[§5\.1](https://arxiv.org/html/2608.12892#S5.SS1.p4.1)\.
- Radhakrishnanet al\.\(2022\)A\. Radhakrishnan, D\. Beaglehole, P\. Pandit, and M\. BelkinMechanism for feature learning in neural networks and backpropagation\-free machine learning models\.Note:arXiv preprint arXiv:2212\.13881External Links:2212\.13881,[Document](https://dx.doi.org/10.48550/arXiv.2212.13881),[Link](https://arxiv.org/abs/2212.13881)Cited by:[§2](https://arxiv.org/html/2608.12892#S2.SS0.SSS0.Px4.p1.1),[§4\.1](https://arxiv.org/html/2608.12892#S4.SS1.SSS0.Px4.p1.1)\.
- Rimskyet al\.\(2024\)N\. Rimsky, N\. Gabrieli, J\. Schulz, M\. Tong, E\. Hubinger, and A\. TurnerSteering llama 2 via contrastive activation addition\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 15504–15522\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.828),[Link](https://aclanthology.org/2024.acl-long.828/)Cited by:[§1](https://arxiv.org/html/2608.12892#S1.p1.1),[§2](https://arxiv.org/html/2608.12892#S2.SS0.SSS0.Px2.p1.1),[§4\.1](https://arxiv.org/html/2608.12892#S4.SS1.SSS0.Px2.p1.2)\.
- Singhet al\.\(2024\)A\. Singh, I\. Padhi, J\. Shen, A\. Dutta, P\. Jain, and J\. SunRepresentation surgery: theory and practice of affine steering\.InProceedings of The 27th International Conference on Artificial Intelligence and Statistics,Proceedings of Machine Learning Research, Vol\.238,pp\. 35328–35351\.External Links:[Link](https://proceedings.mlr.press/v238/singh24d.html)Cited by:[§1](https://arxiv.org/html/2608.12892#S1.p1.1),[§2](https://arxiv.org/html/2608.12892#S2.SS0.SSS0.Px2.p1.1)\.
- Stoehret al\.\(2024\)N\. Stoehr, K\. Du, V\. Snæbjarnarson, R\. West, R\. Cotterell, and A\. ScheinActivation scaling for steering and interpreting language models\.InFindings of the Association for Computational Linguistics: EMNLP 2024,pp\. 8189–8200\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.findings-emnlp.479),[Link](https://aclanthology.org/2024.findings-emnlp.479/)Cited by:[§1](https://arxiv.org/html/2608.12892#S1.p1.1),[§2](https://arxiv.org/html/2608.12892#S2.SS0.SSS0.Px2.p1.1)\.
- Tanet al\.\(2024\)D\. C\. H\. Tan, D\. Chanin, A\. Lynch, D\. Kanoulas, B\. Paige, A\. Garriga\-Alonso, and R\. KirkAnalysing the generalisation and reliability of steering vectors\.InAdvances in Neural Information Processing Systems 37,External Links:[Document](https://dx.doi.org/10.52202/079017-4417),[Link](https://proceedings.neurips.cc/paper_files/paper/2024/hash/fb3ad59a84799bfb8d700e56d19c231b-Abstract-Conference.html)Cited by:[§1](https://arxiv.org/html/2608.12892#S1.p1.1),[§2](https://arxiv.org/html/2608.12892#S2.SS0.SSS0.Px2.p1.1)\.
- Turneret al\.\(2023\)A\. M\. Turner, L\. Thiergart, D\. Leech, D\. Udell, J\. J\. Vazquez, U\. Mini, and M\. MacDiarmidSteering language models with activation engineering\.Note:arXiv preprint arXiv:2308\.10248External Links:2308\.10248,[Document](https://dx.doi.org/10.48550/arXiv.2308.10248),[Link](https://arxiv.org/abs/2308.10248)Cited by:[§1](https://arxiv.org/html/2608.12892#S1.p1.1),[§2](https://arxiv.org/html/2608.12892#S2.SS0.SSS0.Px2.p1.1),[§4\.1](https://arxiv.org/html/2608.12892#S4.SS1.SSS0.Px2.p1.2)\.
- Wanget al\.\(2024a\)P\. Wang, N\. Zhang, B\. Tian, Z\. Xi, Y\. Yao, Z\. Xu, M\. Wang, S\. Mao, X\. Wang, S\. Cheng, K\. Liu, Y\. Ni, G\. Zheng, and H\. ChenEasyEdit: an easy\-to\-use knowledge editing framework for large language models\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 3: System Demonstrations\),pp\. 82–93\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.acl-demos.9),[Link](https://aclanthology.org/2024.acl-demos.9/)Cited by:[§2](https://arxiv.org/html/2608.12892#S2.SS0.SSS0.Px1.p1.1)\.
- Wanget al\.\(2024b\)Y\. Wang, X\. Ma, G\. Zhang, Y\. Ni, A\. Chandra, S\. Guo, W\. Ren, A\. Arulraj, X\. He, Z\. Jiang, T\. Li, M\. Ku, K\. Wang, A\. Zhuang, R\. Fan, X\. Yue, and W\. ChenMMLU\-Pro: a more robust and challenging multi\-task language understanding benchmark\.Note:arXiv preprint arXiv:2406\.01574External Links:2406\.01574,[Document](https://dx.doi.org/10.48550/arXiv.2406.01574)Cited by:[Appendix S1](https://arxiv.org/html/2608.12892#A1.p2.1),[§5\.1](https://arxiv.org/html/2608.12892#S5.SS1.p1.1)\.
- Welblet al\.\(2017\)J\. Welbl, N\. F\. Liu, and M\. GardnerCrowdsourcing multiple choice science questions\.InProceedings of the 3rd Workshop on Noisy User\-generated Text,pp\. 94–106\.External Links:[Document](https://dx.doi.org/10.18653/v1/W17-4413)Cited by:[Appendix S1](https://arxiv.org/html/2608.12892#A1.p2.1),[§5\.1](https://arxiv.org/html/2608.12892#S5.SS1.p1.1)\.
- Whiteet al\.\(2025\)C\. White, S\. Dooley, M\. Roberts, A\. Pal, B\. Feuer, S\. Jain, R\. Shwartz\-Ziv, N\. Jain, K\. Saifullah, S\. Naidu, C\. Hegde, Y\. LeCun, T\. Goldstein, W\. Neiswanger, and M\. GoldblumLiveBench: a challenging, contamination\-limited LLM benchmark\.InThe Thirteenth International Conference on Learning Representations,Cited by:[Appendix S1](https://arxiv.org/html/2608.12892#A1.p2.1),[§5\.1](https://arxiv.org/html/2608.12892#S5.SS1.p1.1)\.
- Wuet al\.\(2025\)Z\. Wu, A\. Geiger, T\. Icard, C\. Potts, and N\. D\. GoodmanAxBench: steering LLMs? even simple baselines outperform sparse autoencoders\.InProceedings of the 42nd International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.267,pp\. 67735–67768\.External Links:[Link](https://proceedings.mlr.press/v267/wu25a.html)Cited by:[§1](https://arxiv.org/html/2608.12892#S1.p1.1),[§2](https://arxiv.org/html/2608.12892#S2.SS0.SSS0.Px2.p1.1),[§6](https://arxiv.org/html/2608.12892#S6.SS0.SSS0.Px3.p1.1)\.
- Xuet al\.\(2026a\)Y\. Xu, T\. Fang, C\. Chen, and L\. YangWhy steering works: toward a unified view of language model parameter dynamics\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 14333–14354\.External Links:[Document](https://dx.doi.org/10.18653/v1/2026.acl-long.704),[Link](https://aclanthology.org/2026.acl-long.704/)Cited by:[§2](https://arxiv.org/html/2608.12892#S2.SS0.SSS0.Px2.p1.1)\.
- Xuet al\.\(2026b\)Z\. Xu, Y\. Xu, Y\. Shen, H\. Liu, S\. Guo, and Y\. SunHow controllable are large language models? a unified evaluation across behavioral granularities\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 25155–25188\.External Links:[Document](https://dx.doi.org/10.18653/v1/2026.acl-long.1225),[Link](https://aclanthology.org/2026.acl-long.1225/)Cited by:[§1](https://arxiv.org/html/2608.12892#S1.p1.1),[§2](https://arxiv.org/html/2608.12892#S2.SS0.SSS0.Px3.p1.1)\.
- Yanget al\.\(2025\)A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv,et al\.Qwen3 Technical Report\.arXiv preprint arXiv:2505\.09388\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2505.09388)Cited by:[Appendix S2](https://arxiv.org/html/2608.12892#A2.p1.1),[§5\.1](https://arxiv.org/html/2608.12892#S5.SS1.p3.1)\.
- Yaoet al\.\(2023\)Y\. Yao, P\. Wang, B\. Tian, S\. Cheng, Z\. Li, S\. Deng, H\. Chen, and N\. ZhangEditing large language models: problems, methods, and opportunities\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,pp\. 10222–10240\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.632),[Link](https://aclanthology.org/2023.emnlp-main.632/)Cited by:[§2](https://arxiv.org/html/2608.12892#S2.SS0.SSS0.Px1.p1.1)\.
- Zellerset al\.\(2019\)R\. Zellers, A\. Holtzman, Y\. Bisk, A\. Farhadi, and Y\. ChoiHellaSwag: can a machine really finish your sentence?\.InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics,pp\. 4791–4800\.External Links:[Document](https://dx.doi.org/10.18653/v1/P19-1472)Cited by:[Appendix S1](https://arxiv.org/html/2608.12892#A1.p2.1),[§5\.1](https://arxiv.org/html/2608.12892#S5.SS1.p1.1)\.
- Zouet al\.\(2023\)A\. Zou, L\. Phan, S\. Chen, J\. Campbell, P\. Guo, R\. Ren, A\. Pan, X\. Yin, M\. Mazeika, A\. Dombrowski, S\. Goel, N\. Li, M\. J\. Byun, Z\. Wang, A\. Mallen, S\. Basart, S\. Koyejo, D\. Song, M\. Fredrikson, J\. Z\. Kolter, and D\. HendrycksRepresentation engineering: a top\-down approach to ai transparency\.Note:arXiv preprint arXiv:2310\.01405External Links:2310\.01405,[Document](https://dx.doi.org/10.48550/arXiv.2310.01405),[Link](https://arxiv.org/abs/2310.01405)Cited by:[§1](https://arxiv.org/html/2608.12892#S1.p1.1),[§2](https://arxiv.org/html/2608.12892#S2.SS0.SSS0.Px2.p1.1)\.
## Appendix
## Appendix S1Frozen Benchmark Construction
The frozen benchmark is constructed from nine public multiple\-choice and reasoning sources \(Table[S5](https://arxiv.org/html/2608.12892#A1.T5)\)\. Records retain their source dataset, domain, release year, and freshness group\. The final collection contains 1,950 records from 2024 sources and 1,050 records from pre\-2024 sources\. Fourteen domains range from situated commonsense and elementary science to recent corrected facts, academic questions, mathematics, and reasoning\.
The source benchmarks are MMLU\-Pro and MMLU\-Redux 2\.0\([28](https://arxiv.org/html/2608.12892#bib.bib27);[6](https://arxiv.org/html/2608.12892#bib.bib28)\), AI2 ARC, OpenBookQA, and SciQ\([1](https://arxiv.org/html/2608.12892#bib.bib29);[16](https://arxiv.org/html/2608.12892#bib.bib30);[29](https://arxiv.org/html/2608.12892#bib.bib31)\), LiveBench\([30](https://arxiv.org/html/2608.12892#bib.bib32)\), HellaSwag\([36](https://arxiv.org/html/2608.12892#bib.bib33)\), and QASC\([10](https://arxiv.org/html/2608.12892#bib.bib34)\)\.
Each record contains a direction\-fitting set and an evaluation assignment\. The direction\-fitting statements never contain the target, neighbor, or capability evaluation fields\. Evaluation uses three target probes, three semantic\-neighbor probes, and four capability probes per record\. Capability assignments are round\-robin balanced over 2,834 unique probes; every capability probe is used four or five times\. This prevents a small set of easy generic questions from dominating the capability\-damage estimate\.
### S1\.1Independent Expert Semantic Audit
Multiple human experts independently reviewed a dataset\-stratified sample of 108 complete records, 12 from each source\. They checked the desired and contrast answers, direction\-fitting statements, three target probes, three semantic neighbors, and four capability probes\. Sixty\-five records pass without issue, 32 have a minor issue that does not reverse the intended relation, and 11 fail at least one semantic criterion\. This yields an 89\.8% acceptable rate on the audited sample\. Because the sample is diagnostic rather than a population error estimate, we additionally remove the known failed records and then all known non\-pass records from the full CPU analyses\. Table[S6](https://arxiv.org/html/2608.12892#A2.T6)shows that principal predictor AUROC changes by at most 0\.004\. At block 7, all learned Target and Clean gains over random also retain positive paired 95% intervals under both exclusions; for RFM/AGOP, the most conservative exclusion gives\+0\.0369\+0\.0369Target and\+0\.0352\+0\.0352Clean\. Thus the reported effects are not driven by the audited exceptions, while independently authored validation remains important future work\.
Table S5:Source composition of the frozen 3,000\-record benchmark\.
## Appendix S2Confirmatory Experimental Details
We evaluate Qwen3\-1\.7B\-Base\([34](https://arxiv.org/html/2608.12892#bib.bib35)\)at 0\-indexed layers 7 and 11, selected from a dense pilot over layers 7, 11, 15, 19, 23, and 27\. The main grid is\{−0\.5,−0\.25,−0\.1,0,0\.1,0\.25,0\.5\}\\\{\-0\.5,\-0\.25,\-0\.1,0,0\.1,0\.25,0\.5\\\}\. Five distinct direction families plus a matched\-random audit entry, which is numerically identical to the random control, and two layers produce 30,000 distinct record–method–layer paths and 36,000 logged rows\. Seven coefficients and ten probes per record produce 210,000 distinct and 252,000 logged path–strength rows, plus 2\.52 million logged probe\-level rows\. Equivalently, the audit artifacts contain 360,000 seven\-point probe trajectories, of which 300,000 are distinct after removing the duplicate entry\. GPU inference uses one NVIDIA GeForce RTX 5090 with bfloat16, batch size two, and seed 113\.
#### Direction construction and inference\.
Each direction is fitted independently for one record and layer using 6–13 positive and 6–13 contrast prompts\. Prompts are tokenized without added special tokens and left padded; the direction\-fitting activation is the final non\-padding prompt token at the selected layer\. Mean difference, ridge linear, and logistic use the same activation matrix and labels\. Ridge regularization is10−310^\{\-3\}, logistic regression runs for at most 1,000 iterations, and RFM uses three iterations with a Laplace kernel of bandwidth 10 and regularization10−310^\{\-3\}\. During evaluation, the unit direction is added to the selected block output at every token position\. Correct and contrast continuations are teacher\-forced separately, and path labels use their mean per\-token log\-probability margin to reduce continuation\-length effects\.
Random\-direction 95th percentiles define separate response thresholds:τT=0\.1753\\tau\_\{T\}=0\.1753,τN=0\.1567\\tau\_\{N\}=0\.1567, andτC=0\.1248\\tau\_\{C\}=0\.1248\. The confirmatory labels use only\|α\|∈\{0\.25,0\.5\}\|\\alpha\|\\in\\\{0\.25,0\.5\\\}\. Weak\-response features use only\|α\|=0\.1\|\\alpha\|=0\.1, so the observed probe strength is disjoint from all label strengths\.
Feature groupBBcontains unperturbed target, neighbor, and capability margins\. GroupMMcontains method, layer, dataset, domain, freshness, and release year\. GroupLLcontains separation, saliency, threshold\-accuracy, and direction\-agreement statistics, whileGGcontains RFM/AGOP spectrum and alignment statistics\. GroupRRcontains signed target, neighbor, and capability responses atα=±0\.1\\alpha=\\pm 0\.1\. Prediction uses class\-balanced logistic regression and random forests with five grouped folds by record, dataset, or domain\. We report AUROC, average precision, balanced accuracy, F1, and Brier score in the released result artifacts; the tables below retain AUROC and AP, the two threshold\-independent ranking metrics\.
#### Computational environment\.
Experiments and paper\-facing analyses run on Ubuntu 22\.04\.5 LTS with an Intel Xeon Gold 6530 CPU \(14 allocated cores\), 117 GiB RAM, and one NVIDIA GeForce RTX 5090 GPU \(driver 580\.65\.06; CUDA toolkit 13\.0\)\. The software environment uses Python 3\.10\.20, PyTorch 2\.12\.0\+cu132, Transformers 5\.8\.1, scikit\-learn 1\.7\.2, pandas 2\.3\.3, NumPy 2\.2\.5, and SciPy 1\.15\.3\.
Table S6:Sensitivity to the independent expert semantic audit\. The stratified audit covers 108 records \(12 per source\): 65 pass, 32 minor issue, and 11 fail\. Cells are complete\-predictor AUROC/AP after retaining all records, excluding the 11 fails, or excluding all 43 non\-pass records\.
## Appendix S3Dimension\-Corrected Cross\-Model Protocol
The confirmation uses one seed\-113 stratified subset of 500 frozen records for Qwen3\-1\.7B\-Base, Qwen3\.5\-2B\-Base\([20](https://arxiv.org/html/2608.12892#bib.bib36)\), and Ministral\-3\-3B\-Base\([13](https://arxiv.org/html/2608.12892#bib.bib37)\)\. A separate seed\-127 calibration set contains 100 records and has zero overlap with the formal subset\. At each aligned layer, 1,941 direction\-construction prompt states are used to estimate the median residual RMS\. The calibration set is used only to fix model–layer scales, not to select records, methods, or outcomes\.
Directions have unitℓ2\\ell\_\{2\}norm\. Consequently, matching only the numerical residual RMS would not align the intervention relative to the residual\-stateℓ2\\ell\_\{2\}norm when hidden widths differ\. We match‖αv‖2/‖h‖2\\\|\\alpha v\\\|\_\{2\}/\\\|h\\\|\_\{2\}with
sm,ℓ=RMS\(hm,ℓ\)RMS\(href,ℓ\)dmdref,αm,ℓ=sm,ℓαref\.s\_\{m,\\ell\}=\\frac\{\\operatorname\{RMS\}\(h\_\{m,\\ell\}\)\}\{\\operatorname\{RMS\}\(h\_\{\\mathrm\{ref\},\\ell\}\)\}\\sqrt\{\\frac\{d\_\{m\}\}\{d\_\{\\mathrm\{ref\}\}\}\},\\qquad\\alpha\_\{m,\\ell\}=s\_\{m,\\ell\}\\alpha\_\{\\mathrm\{ref\}\}\.\(S7\)Qwen3\-1\.7B and Qwen3\.5 have hidden width 2,048; Ministral has width 3,072 and therefore includes a3,072/2,048\\sqrt\{3\{,\}072/2\{,\}048\}correction\. The resulting layerwise scales are 1\.000/1\.000 for Qwen3\-1\.7B layers 7/11, 0\.088/0\.041 for Qwen3\.5 layers 6/9, and 0\.091/0\.052 for Ministral layers 6/10\. The manifest stores the RMS ratio, hidden\-width ratio, scale, record hash, and source\-summary hashes\.
All models retain the same target, neighbor, and capability probes; reference coefficient grid; and confirmatory outcome definitions\. At both aligned layers we evaluate random, mean difference, logistic, and RFM/AGOP, for 24 model–layer–method configurations, 12,000 record\-level paths, and 84,000 path–strength evaluations\. Linear remains in the complete 3,000\-record primary study but is omitted here because logistic represents the same supervised discriminative direction family\. Each model recalibrates target, neighbor, and capability thresholds from its own random directions on the frozen 500\-record cohort\. Qwen3\-1\.7B has scale one, so its 500\-record outcomes are exact subsets of the completed primary run; only the cohort\-level random thresholds are recomputed\.
The corresponding path\-incidence outcomes are reported in the main paper\. This section records the calibration and normalization choices needed to reproduce that comparison without duplicating the main result table\.
## Appendix S4Uncertainty and Worked Examples
For Table[S7](https://arxiv.org/html/2608.12892#A4.T7), each learned row is paired with the random row for the same record and layer\. We resample the 3,000 record identifiers with replacement 10,000 times and recompute the mean paired difference; all method\-specific measurements from a sampled record remain together\. The table reports percentile intervals from the deterministic bootstrap implemented in the paper asset script\.
Table S7:Record\-paired percentage\-point differences from random in the frozen 3,000\-record study, with 95% intervals from 10,000 record bootstrap resamples\. Positive target/clean differences are favorable; positive damage differences are unfavorable\.### S4\.1Worked Frozen Record and Path Labels
The main paper traces one unchanged record from direction\-fitting statements to target, neighbor, capability, and path labels\. Table[S8](https://arxiv.org/html/2608.12892#A4.T8)adds clean\-only, collateral\-without\-clean, no\-effect, and mixed\-sign examples\. The mixed example also illustrates why Clean and damage\-any are not complements: one sign can contain a clean operating coefficient while collateral movement occurs elsewhere on the full path\.
Table S8:Additional frozen\-record path examples\. Onsets list the first suppression/enhancement/damage threshold crossings; a dash denotes no crossing\. The examples illustrate label construction rather than quantitative evidence\.
## Appendix S5Additional Prediction Analysis
Across all cohorts, split types, targets, and both prediction models, addingLLtoB\+MB\+Mchanges AUROC by\+0\.0021\+0\.0021and AP by\+0\.0005\+0\.0005on average\. Adding the disjoint weak probeRRtoB\+MB\+Mchanges AUROC by\+0\.1774\+0\.1774and AP by\+0\.1821\+0\.1821\. GeometryGGis outcome\-specific within the RFM cohort: it improves later enhancement and target\-any prediction but reduces capability\-damage AUROC, so we do not treat it as a uniformly beneficial feature family\. Table[S13](https://arxiv.org/html/2608.12892#A5.T13)reports the corresponding RFM\-cohort differences explicitly\.
The main paper reports the cross\-model summary and the principal CPU controls\. Here we provide the full layer\-selection exclusion, weak\-response baseline, grouped\-transfer, and RFM\-specific geometry results\.
Table S9:Macro AUROC across the eight outcomes visualized in the main paper\.ΔL\\Delta LandΔR\\Delta Rare cumulative gains; Final isB\+M\+L\+RB\+M\+L\+R\. Dataset and Domain use the final predictor\. The Qwen3\-1\.7B subset comes from the primary cohort\.### S5\.1Measured\-Grid Path Topology
We recompute descriptive topology on all nonzero measured strengths\|α\|∈\{0\.1,0\.25,0\.5\}\|\\alpha\|\\in\\\{0\.1,0\.25,0\.5\\\}\. Across learned directions, a target crossing exists on 7\.34% of paths, a clean coefficient on 6\.94%, and a nonmonotonic target\-or\-damage indicator on 10\.90%; the corresponding random rates are 5\.98%, 5\.55%, and 10\.65%\. Strict target\-first and damage\-first patterns are rare \(0\.37% and 0\.46% for learned directions\), so they are descriptive rather than primary prediction targets\. Among learned paths with any clean coefficient, 76\.3% contain exactly one measured clean coefficient\. These results justify evaluating multiple strengths, but not a claim that broad continuous clean windows are common\.
### S5\.2Layer\-Selection Exclusion and Weak\-Response Controls
The 500 records used to select the two primary blocks are a subset of the 3,000\-record cohort\. Removing them leaves 2,500 records and 30,000 method–block paths\. At block 11, all four learned direction families retain positive paired Target\-any and Clean\-any differences from random \(Table[S10](https://arxiv.org/html/2608.12892#A5.T10)\)\. The complete record\-held\-out random forest also remains within 0\.01 AUROC of its full\-cohort estimate on each principal outcome \(Table[S11](https://arxiv.org/html/2608.12892#A5.T11)\)\. Thus neither the outcome nor prediction result is driven by reuse of the layer\-selection subset\.
The training\-free scalar control uses only the signed\|α\|=0\.1\|\\alpha\|=0\.1response matched to each later outcome\. TheRR\-only random forest instead uses the multivariate low\-dose response profile\. For target and clean outcomes, this profile substantially improves over the scalar score; adding base, method/block, and static\-localization features provides a smaller final AUROC gain\. We report prevalence and AP alongside AUROC because the positive path labels are sparse\. The complete feature set is not uniformly best in AP, so our conclusion concerns the additional ranking information in the structured weak response rather than universal dominance on every metric\.
Table S10:Sensitivity after excluding the 500\-record layer\-selection subset\. Entries are learned\-minus\-random path\-incidence differences at block 11 with record\-paired 95% bootstrap intervals on the remaining 2,500 records\.
Table S11:Record\-held\-out CPU prediction controls on Qwen3\-1\.7B\-Base\. In panel \(a\), each cell is AUROC/AP on all 3,000 records\. Scalar is the training\-free signed weak\-response score;RR\-only and Full are random forests, where Full usesB\+M\+L\+RB\+M\+L\+R\. Panel \(b\) reports prevalence, AUROC, and AP with 95% intervals after excluding the 500\-record layer\-selection subset\.
### S5\.3Grouped Transfer and Geometry
Table S12:Macro performance over the eight confirmatory later\-strength outcomes\. Values are random\-forest AUROC/AP in the all\-method cohort\.
Table S13:RFM\-cohort geometry ablation, averaged over record\-, dataset\-, and domain\-grouped splits\.GGcontains AGOP spectrum and alignment features\. Deltas compareB\+M\+L\+GB\+M\+L\+Gwith theB\+M\+LB\+M\+Lbase\.
Figure S4:Grouped\-transfer robustness\. Static localization leaves the baseline nearly unchanged, whereas the disjoint weak response produces a large AUROC gain under record\-, dataset\-, and domain\-held\-out evaluation\.
## Appendix S6Multi\-Fidelity Strength Selection
The dense training study contains 500 records, three methods, six layers, and 27 coefficients, yielding 243,000 raw record\-strength rows\. A disjoint 100\-record validation study contributes 24,300 raw dense rows\. Excluding the zero coefficient leaves 234,000 dense training candidates and 23,400 validation candidates\. Multi\-fidelity training adds 2,400 non\-overlapping records with sparse trajectories, for 2,900 training records and 320,400 nonzero candidate rows; validation overlap is zero\.
We evaluate selection by held\-out policy replay\. The global fixed coefficient is chosen only from mean utility on the 500\-record dense training set, which selects−1\-1for suppression and\+1\+1for enhancement\. For every unseen validation path, a learned policy observes metadata, the candidate coefficient, and the two\|α\|=0\.1\|\\alpha\|=0\.1weak responses, then either chooses one coefficient or abstains\. Only after this decision do we reveal the measured dense\-grid outcome at the selected coefficient\. Thus the validation response curve is not available to the selector or to fixed\-baseline tuning\.
Without a weak response, the multi\-fidelityB\+M\+AB\+M\+Agradient\-boosted decision tree \(GBDT\) reaches clean\-alpha AUROC0\.6100\.610and AP0\.0870\.087\. AddingRRraises theM\+A\+RM\+A\+Rmodel to AUROC0\.8490\.849and AP0\.3230\.323; addingBBorLLafterRRdoes not improve these ranking metrics consistently\. For the sameM\+A\+RM\+A\+Rmodel, multi\-fidelity training increases held\-out clean rate over dense\-only training by 1\.0 point for suppression and 0\.8 points for enhancement, while utility changes by less than0\.0010\.001\.
The practical gain comes from risk\-aware selection rather than uniformly higher clean rates \(Table[S14](https://arxiv.org/html/2608.12892#A6.T14)\)\. No intervention has zero utility by construction, while the train\-tuned fixed\|α\|=1\|\\alpha\|=1policy has negative utility in both directions\. The multi\-fidelityM\+A\+RM\+A\+Rselector abstains on 45\.4% of suppression paths and 39\.1% of enhancement paths, reducing neighbor damage from 6\.0% to 1\.8% and from 5\.2% to 2\.2%, respectively\. Both paired utility gains over the fixed coefficient are significant\. Suppression utility is also significantly positive relative to no intervention, whereas the enhancement interval against zero overlaps zero \(Table[S16](https://arxiv.org/html/2608.12892#A6.T16)\)\. Its expected cost is 2\.55 and 2\.61 intervention evaluations per path, about 90% fewer than the 26\-point nonzero dense scan \(Table[S15](https://arxiv.org/html/2608.12892#A6.T15)\)\. A substantial oracle gap remains, so the selector is a decision aid rather than a replacement for dense evaluation\.
Table S14:Held\-out dense\-grid strength selection on 100 records\.AAdenotes candidate\-strength features; MF denotes multi\-fidelity training\.Table S15:Expected intervention evaluations per held\-out path\. Learned policies use two weak probes and execute the selected coefficient only when they do not abstain\. Reduction is relative to evaluating all 26 nonzero dense\-grid coefficients\.Table S16:Record\-paired uncertainty for the multi\-fidelityM\+A\+RM\+A\+Rselector on 100 held\-out records\. Brackets are 95% intervals from 10,000 record bootstrap resamples\. Outcome changes are selector\-minus\-fixed in percentage points\.Table S17:Selector utility sensitivity on the same 100 held\-out records\. Each row retrains the utility head after changing one declared weight\. Entries are paired utility gains for suppression/enhancement\.Selector uncertainty also uses records, not paths, as the sampling unit\. We first average the nine method–layer paths within each of the 100 validation records, then resample records 10,000 times\. Utility and damage comparisons are paired against the train\-tuned fixed policy on the same validation record\. Table[S17](https://arxiv.org/html/2608.12892#A6.T17)varies one utility weight at a time and retrains the utility head\. All seven configurations retain positive paired gains over the fixed policy and reduce neighbor damage, whereas gains over no intervention are not universal\. The declared weights are therefore not the only configuration that supports risk reduction, but the experiment does not imply deployment\-independent optimality\.
Table S18:Path\-conditioned free\-generation endpoint stress tests on 100 records per model\. G/D counts unique target prompts with correctness gain/damage at any nonzero coefficient\. Text and clause changes are measured againstα=0\\alpha=0\. Qwen3\.5 uses the unmatched raw\-alpha stress protocol\.
## Appendix S7Free\-Generation Endpoint Stress Tests
We keep all free\-generation evidence in the supplement because the primary claims concern random\-calibrated answer\-margin paths\. Two path\-conditioned 100\-record studies cover four directions, two aligned layers, and coefficients from−1\-1to11\. On Qwen3\-1\.7B, decoded text often changes but learned directions do not produce a target correction; one isolated random correction is not evidence of controllability\. On Qwen3\.5, the unmatched raw\-alpha stress test changes roughly half of nonzero\-strength target generations and produces several learned wrong\-to\-right transitions, but target damage balances or exceeds these corrections\. The experiment therefore establishes endpoint reachability at strong dose, not reliable endpoint improvement\.
## Appendix S8Asset Attribution
Main\-paper Figure 1\(a,c\) contains explanatory imagery rather than measured PML activations\. Panel \(a\) adapts Activation Atlas by Carter et al\. \(CC BY 4\.0\), and panel \(c\) adapts Deep Learning Visuals by David V\. Godoy \(CC BY 4\.0\)\. Panel \(b\) and all quantitative figures are generated from the paper workflow\.Similar Articles
PraMem: Practice-derived Experiential Memory for Long-horizon Behavior Prediction
PraMem proposes a paradigm shift for long-horizon behavior prediction by using a lengthy historical sequence as a resource to build experiential memory via practice, improving LLM performance on behavior prediction tasks.
Personalizing Embodied Multimodal Large Language Model Agents over Long-term User Interactions
This paper proposes Polar, a multimodal memory-augmented framework for personalizing embodied MLLM agents over long-term user interactions, using a knowledge graph and episodic memory to ground user-intended instances from accumulated context.
Navigating by Old Maps: The Pitfalls of Static Mechanistic Localization in LLM Post-Training
This paper challenges the 'Locate-then-Update' paradigm in LLM post-training by demonstrating that static mechanistic localization is insufficient due to the dynamic evolution of neural circuits during fine-tuning. It introduces new metrics to analyze circuit stability and proposes the need for predictive frameworks in mechanistic localization.
Early Prediction of Future Behavioral Strategy from Process Traces
A new process-level latent variable model (PLVM) predicts future behavioral strategies from partial process traces across tasks, demonstrated in PowerWash Simulator gameplay data.
Progressive Point Matching (8 minute read)
Progressive Point Matching (PPM) is a framework proposed to assign partial credit in reinforcement learning for long-horizon tasks in LLMs, addressing the inefficiency of sparse outcome rewards by treating reasoning as paths through a Markovian state space.