Tag
The paper introduces SAKIKO, a mechanistic auditing framework showing that activation steering interventions in tool-using LLMs may displace probability mass without genuine repair, finding that one +55 net-gain intervention corrupts over half of the baseline-correct decisions it touches.
A paper introducing PersonaDose, a calibrated activation-steering approach that controls language model persona expression via trait descriptions and requested mean intensity, evaluated across Llama-3.1-8B, Qwen3-8B, and Gemma-3-4B.
This paper introduces AIMES, a framework for adaptive multi-value activation steering in large language models that uses online observer feedback for state-aware adaptation without additional training, showing improved controllability over fixed methods.
This paper introduces a guarded gradient-based activation steering method to detect and correct shutdown-avoidance responses in the Qwen3.5-0.8B model, using a classifier and minimum-step policy for selective and limited effectiveness.
LocUS is a novel method for targeted activation steering in large language models that restricts interventions to specific attention heads and a subspace of the unembedding matrix, improving performance while preserving general capabilities.
The study reveals that chat templates control whether language models adopt a disclaimer voice (e.g., 'I'm just an AI') or an experiential voice (e.g., 'I feel'), and identifies an activation direction that can steer this behavior, impacting AI safety research.
This paper reveals that the optimal layer for linear probing to read concepts differs from the optimal layer for activation steering in omni-modal large language models, challenging common heuristics in representation engineering.
This study finds that activation steering in latent chain-of-thought reasoning is less effective than in explicit CoT, highlighting a transition gap where interventions in latent space fail to transfer to language generation.
The study explores steering LLMs towards human moral foundations using the Norwegian MFQ-30 questionnaire, evaluating prompt-level persona steering and activation-level interventions.
This paper explores the role of fine-grained category-specific harmfulness representations in LLMs for safety. It shows that category residuals, orthogonal to general harm, vary in encoding harmfulness and can induce refusal, with implications for understanding LLM safety mechanisms.
This study validates activation-steering claims for sycophancy in language models, finding no usable linear capitulation direction in two small LLMs and highlighting measurement hazards that underestimate sycophancy.
GeoSteer is an optimization-based method for norm-preserving activation steering in large language models, using geodesic updates on the representation manifold to improve control over model behavior while preserving activation norms.
This paper proposes a spillover-aware method for multi-value activation steering to achieve pluralistic alignment in LLMs, improving control over multiple value dimensions without fine-tuning or reward models.
Research article exploring the use of Jacobian space to extract steering vectors from concept tokens for steering LLM behaviors, with experiments on Qwen3-1.7B showing promise for simple tasks but limitations in complex scenarios.
This paper introduces a benchmark to validate that distribution-driven methods for activation steering in LLMs recover human value topologies aligned with Schwartz's theory, improving with model size but declining post-instruction tuning.
GAPS introduces dimension-level gating for conditional activation steering in language models, combining static and dynamic gates to selectively intervene and improve behavior-capability trade-off, with significant gains in toxicity mitigation and concept removal tasks.
This paper investigates whether fine-tuning undoes activation steering in language models, finding that while behavioral effects like refusal suppression can degrade under optimization pressure, the underlying weight modifications remain mechanistically stable.
This paper introduces Gated Activation Steering, a method to reduce sycophancy and hallucination in large language models for medical question answering using inference-time interventions. Evaluated on clinical data, it demonstrates improved robustness under user pressure.
This paper proposes the Defactualize-Steer-Rehydrate (DSR) framework, integrating knowledge graphs with activation steering to enhance style-controllable and fact-preserving generation in agentic conversational AI, evaluated on LLaMA models with significant improvements in factual recovery.
This paper introduces Predictive Memory Localization (PML), a method that forecasts selective intervention outcomes from internal model signals, distinguishing calibrated target movement from semantic-neighbor and capability damage. Experiments across 3,000 records and nine datasets show improved selective-path prediction and risk-aware intervention decisions.