activation-steering

Tag

Cards List
#activation-steering

When Does Correction Become Repair? Mechanistic Auditing of Internal Interventions in Tool-Using LLMs

arXiv cs.CL ↗ · yesterday Cached

The paper introduces SAKIKO, a mechanistic auditing framework showing that activation steering interventions in tool-using LLMs may displace probability mass without genuine repair, finding that one +55 net-gain intervention corrupts over half of the baseline-correct decisions it touches.

0 favorites 0 likes
#activation-steering

Persona Dosing: Calibrated Activation Steering for Graded Trait Control

arXiv cs.AI ↗ · yesterday Cached

A paper introducing PersonaDose, a calibrated activation-steering approach that controls language model persona expression via trait descriptions and requested mean intensity, evaluated across Llama-3.1-8B, Qwen3-8B, and Gemma-3-4B.

0 favorites 0 likes
#activation-steering

Adaptive Multi-Value Control in LLMs via Causal Activation Steering

arXiv cs.LG ↗ · 2d ago Cached

This paper introduces AIMES, a framework for adaptive multi-value activation steering in large language models that uses online observer feedback for state-aware adaptation without additional training, showing improved controllability over fixed methods.

0 favorites 0 likes
#activation-steering

Guarded Gradient-Based Activation Steering of Shutdown Responses in Qwen3.5-0.8B: A Minimum-Step Policy

arXiv cs.LG ↗ · 2d ago Cached

This paper introduces a guarded gradient-based activation steering method to detect and correct shutdown-avoidance responses in the Qwen3.5-0.8B model, using a classifier and minimum-step policy for selective and limited effectiveness.

0 favorites 0 likes
#activation-steering

LocUS: Head Selection and Subspace Projection for Targeted Activation Steering

arXiv cs.CL ↗ · 3d ago Cached

LocUS is a novel method for targeted activation steering in large language models that restricts interventions to specific attention heads and a subspace of the unembedding matrix, improving performance while preserving general capabilities.

0 favorites 0 likes
#activation-steering

"As a Language Model...": Chat Template Switches LLM Self-Referential Voice and Activation Steering Reproduces It

arXiv cs.LG ↗ · 2026-09-23 Cached

The study reveals that chat templates control whether language models adopt a disclaimer voice (e.g., 'I'm just an AI') or an experiential voice (e.g., 'I feel'), and identifies an activation direction that can steer this behavior, impacting AI safety research.

0 favorites 0 likes
#activation-steering

Read-Best Is Not Steer-Best: A Probing--Steering Layer Dissociation in Omni-Modal Large Language Models

arXiv cs.CL ↗ · 2026-09-22 Cached

This paper reveals that the optimal layer for linear probing to read concepts differs from the optimal layer for activation steering in omni-modal large language models, challenging common heuristics in representation engineering.

0 favorites 0 likes
#activation-steering

When Steering Fails in Latent Reasoning: A Latent-to-Language Transition Gap

arXiv cs.CL ↗ · 2026-09-21 Cached

This study finds that activation steering in latent chain-of-thought reasoning is less effective than in explicit CoT, highlighting a transition gap where interventions in latent space fail to transfer to language generation.

0 favorites 0 likes
#activation-steering

Steering LLMs Responses Towards Moral Foundations on the Norwegian MFQ-30

arXiv cs.CL ↗ · 2026-09-21 Cached

The study explores steering LLMs towards human moral foundations using the Norwegian MFQ-30 questionnaire, evaluating prompt-level persona steering and activation-level interventions.

0 favorites 0 likes
#activation-steering

The Role of Fine-grained Harm Signals in LLM Safety

arXiv cs.CL ↗ · 2026-09-18 Cached

This paper explores the role of fine-grained category-specific harmfulness representations in LLMs for safety. It shows that category residuals, orthogonal to general harm, vary in encoding harmfulness and can induce refusal, with implications for understanding LLM safety mechanisms.

0 favorites 0 likes
#activation-steering

No Usable Linear "Capitulation Direction" in Two Small LLMs: A Validation Protocol for Activation-Steering Claims, and a Cross-Family Behavioral Study of Sycophancy Under Pushback

arXiv cs.CL ↗ · 2026-09-17 Cached

This study validates activation-steering claims for sycophancy in language models, finding no usable linear capitulation direction in two small LLMs and highlighting measurement hazards that underestimate sycophancy.

0 favorites 0 likes
#activation-steering

GEOSTEER: Geodesic Optimization for Activation Steering in Large Language Models

arXiv cs.LG ↗ · 2026-09-11 Cached

GeoSteer is an optimization-based method for norm-preserving activation steering in large language models, using geodesic updates on the representation manifold to improve control over model behavior while preserving activation norms.

0 favorites 0 likes
#activation-steering

Spillover-Aware Multi-Value Steering for Pluralistic LLM Alignment

arXiv cs.AI ↗ · 2026-09-10 Cached

This paper proposes a spillover-aware method for multi-value activation steering to achieve pluralistic alignment in LLMs, improving control over multiple value dimensions without fine-tuning or reward models.

0 favorites 0 likes
#activation-steering

Extracting Steering Vectors from J space

Hacker News Top ↗ · 2026-09-06 Cached

Research article exploring the use of Jacobian space to extract steering vectors from concept tokens for steering LLM behaviors, with experiments on Qwen3-1.7B showing promise for simple tasks but limitations in complex scenarios.

0 favorites 0 likes
#activation-steering

Steering Geometry: Validating Human Value Geometry in LLM Steering Space

Hugging Face Daily Papers ↗ · 2026-09-05 Cached

This paper introduces a benchmark to validate that distribution-driven methods for activation steering in LLMs recover human value topologies aligned with Schwartz's theory, improving with model size but declining post-instruction tuning.

0 favorites 0 likes
#activation-steering

GAPS: Dimension-Level Gates for Conditional Activation Steering

arXiv cs.CL ↗ · 2026-09-03 Cached

GAPS introduces dimension-level gating for conditional activation steering in language models, combining static and dynamic gates to selectively intervene and improve behavior-capability trade-off, with significant gains in toxicity mitigation and concept removal tasks.

0 favorites 0 likes
#activation-steering

Does Fine-Tuning Undo Activation Steering? Behavioural Recovery Without Weight-Edit Reversal

arXiv cs.CL ↗ · 2026-08-27 Cached

This paper investigates whether fine-tuning undoes activation steering in language models, finding that while behavioral effects like refusal suppression can degrade under optimization pressure, the underlying weight modifications remain mechanistically stable.

0 favorites 0 likes
#activation-steering

Gated Activation Steering for Reducing Sycophancy & Hallucination in Medical Question Answering

arXiv cs.AI ↗ · 2026-08-26 Cached

This paper introduces Gated Activation Steering, a method to reduce sycophancy and hallucination in large language models for medical question answering using inference-time interventions. Evaluated on clinical data, it demonstrates improved robustness under user pressure.

0 favorites 0 likes
#activation-steering

Knowledge-Graph-Gated Defactualization for Style-Controllable and Fact-Preserving Generation in Agentic Conversational AI

arXiv cs.CL ↗ · 2026-08-24 Cached

This paper proposes the Defactualize-Steer-Rehydrate (DSR) framework, integrating knowledge graphs with activation steering to enhance style-controllable and fact-preserving generation in agentic conversational AI, evaluated on LLaMA models with significant improvements in factual recovery.

0 favorites 0 likes
#activation-steering

Predictive Memory Localization: Forecasting Selective Intervention Paths from Internal Signals

arXiv cs.AI ↗ · 2026-08-14 Cached

This paper introduces Predictive Memory Localization (PML), a method that forecasts selective intervention outcomes from internal model signals, distinguishing calibrated target movement from semantic-neighbor and capability damage. Experiments across 3,000 records and nine datasets show improved selective-path prediction and risk-aware intervention decisions.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback