activation-steering

Tag

Cards List
#activation-steering

Located but Not Releasable: Silent Gate Inversion and Bounded Linear Release

arXiv cs.CL ↗ · 2026-08-13 Cached

A preregistered stress test on a small transformer shows that while latent causal structure can be localized, releasing it into behavior fails: the gate detector inverts out-of-distribution and linear release directions are bounded below sufficiency, dissociating localization from behavioral release.

0 favorites 0 likes
#activation-steering

Forecasting Side Effects of Activation Steering

arXiv cs.AI ↗ · 2026-08-13 Cached

This paper investigates whether side effects of activation steering in language models can be predicted before intervention, constructing a cross-effect matrix across 67 behaviors and finding that side effects are systematic and forecastable from unsteered representations.

0 favorites 0 likes
#activation-steering

Deployable Per-Instance Multi-Layer Activation Steering for Large Language Models

arXiv cs.CL ↗ · 2026-08-11 Cached

This paper introduces a deployable per-instance, multi-layer activation steering technique for large language models, showing that optimal layer selection varies per input and can be predicted from the prompt embedding without gold labels at inference.

0 favorites 0 likes
#activation-steering

PoolBench: A Benchmark for Pooling Strategies in Concept Representation Evaluation for Decoder-Only LLMs

arXiv cs.CL ↗ · 2026-08-07 Cached

Introduces PoolBench, a benchmark that isolates pooling strategies in decoder-only LLM concept representation, evaluating 19 strategies across 17 concepts and 3 models to provide a standardized protocol for comparing pooling choices.

0 favorites 0 likes
#activation-steering

A Few Neurons Reveal When LLMs Misuse Tools: Sparse Detection and Selective Steering for Reliable Tool Use

arXiv cs.CL ↗ · 2026-08-04 Cached

This paper introduces PRISMS, a framework that uses a small set of failure-specific MLP neurons to detect and steer LLM tool-use errors (over-calling, missing calls, invalid arguments) with sparse readouts, improving reliability across multiple model families.

0 favorites 0 likes
#activation-steering

Role Steering of Language Models for Social Simulations

arXiv cs.CL ↗ · 2026-08-04 Cached

Introduces an activation-steering screening workflow for role-conditioned LLM agents in social simulations, evaluated on OLMo-3-7B-Instruct across a 275-role inventory and showing role-specific directions outperform assistant-axis control.

0 favorites 0 likes
#activation-steering

Token-Level Diagnosis of Sycophancy in LLMs with Attribution-Guided Steering

arXiv cs.CL ↗ · 2026-08-03 Cached

This paper introduces an Integrated Gradients-based token attribution method to diagnose sycophancy in LLMs at the token level, and proposes attribution-guided contrastive activation steering to reduce sycophantic behavior during inference without retraining.

0 favorites 0 likes
#activation-steering

On the Generalization of Steering Vectors for Chain-of-Thought Faithfulness

arXiv cs.AI ↗ · 2026-08-03 Cached

This paper investigates how well activation steering for improving chain-of-thought faithfulness generalizes across cue types, datasets, and steering vector construction methods across several Gemma and Qwen models. The authors find that when steering is effective, it generalizes broadly, and the evaluation setting matters more than the training setting or vector construction method.

0 favorites 0 likes
#activation-steering

Policy Gradient Steering: Interventions from Behavioral Objectives

arXiv cs.LG ↗ · 2026-07-31 Cached

Introduces Policy Gradient Steering (PGS), a method that formulates activation steering as a reinforcement learning problem, using policy gradients to construct removable, composable steering vectors from behavioral objectives. Validated in gridworld, chess puzzle, and football environments.

0 favorites 0 likes
#activation-steering

Latent-IM: Latent Interaction Management for Speech LLMs

arXiv cs.CL ↗ · 2026-07-30 Cached

Introduces Latent-IM, a framework for recovering interaction management from frozen speech LLMs using activation-based selection and steering for conversational moves. It improves end-to-end move accuracy by 12.5 points over the unsteered backbone.

0 favorites 0 likes
#activation-steering

Where Steering Signals Come From: Activation Source Selection in Activation Steering

arXiv cs.CL ↗ · 2026-07-29 Cached

This paper investigates how the choice of source activations influences activation steering in language models, finding that execution-boundary states (where the model is about to produce target behavior) yield stronger signals, and introduces tail subtraction to improve steering stability.

0 favorites 0 likes
#activation-steering

LoRA for Gender-Inclusive Rewriting and Activation Steering for Counter-Narrative Generation

arXiv cs.CL ↗ · 2026-07-28 Cached

This paper presents the IHLC submission to the LT-EDI 2026 shared task, using LoRA fine-tuning for gender-neutral rewriting (Rank 3) and activation steering for counter-narrative generation (Rank 6), highlighting both promise and limitations.

0 favorites 0 likes
#activation-steering

The Geometry of Personality: Activation Steering with Jungian Cognitive Functions

arXiv cs.CL ↗ · 2026-07-24 Cached

This paper introduces a framework using Jungian cognitive functions for activation steering in LLMs, demonstrating effective monotonic control over eight functions and revealing structured geometric relationships in activation space.

0 favorites 0 likes
#activation-steering

DecodeShare: Tracing the Shared Subspace of LLM Decode-Time Decisions

arXiv cs.AI ↗ · 2026-07-24 Cached

DecodeShare proposes a method to identify a low-dimensional subspace consistently shared across tasks in LLM decode-time hidden states and shows that disturbing this subspace degrades decision performance more than random or prefill-derived subspaces, with implications for activation steering.

0 favorites 0 likes
#activation-steering

GrAInS: Gradient-based Attribution for Inference-Time Steering of LLMs and VLMs

arXiv cs.CL ↗ · 2026-07-13 Cached

GrAInS is a contrastive gradient-based method that uses Integrated Gradients to identify influential tokens and construct steering vectors for inference-time steering of LLMs and VLMs, improving truthfulness and reducing hallucinations without degrading fluency.

0 favorites 0 likes
#activation-steering

Present but Rescaled: Chat-to-Agent Transfer of Additive Activation Steering

arXiv cs.LG ↗ · 2026-07-13 Cached

This paper presents the first systematic study of how additive activation steering, calibrated in single-turn chat, transfers to tool-using ReAct agents. It finds that while the injected direction reaches late layers at near-full strength across settings, the behavioral coupling varies unpredictably between models and contexts, with amplification up to 2x or attenuation, posing immediate safety concerns.

0 favorites 0 likes
#activation-steering

LLMs Silently Correct African American English: Auditing and Mitigating Dialect Bias via Activation Steering

arXiv cs.CL ↗ · 2026-07-09 Cached

This paper audits and mitigates dialect bias in large language models, showing they systematically prefer Standard American English over African American English. The authors introduce activation steering, a training-free method that reduces bias significantly while preserving fluency, and release the largest real-AAE parallel corpus to date.

0 favorites 0 likes
#activation-steering

A Quantized Native Runtime for On-Device Semantic Audio Generation

Hugging Face Daily Papers ↗ · 2026-07-09 Cached

This paper presents a dependency-free native runtime that enables efficient text-to-music generation on embedded devices through quantization and activation steering, achieving no measurable quality loss at 8-bit precision and enabling the 1.2B parameter model to run on a Raspberry Pi 5 at 4-bit precision.

0 favorites 0 likes
#activation-steering

A Coin Flip Per Token: Bernoulli Sparse Steering of Large Language Models

arXiv cs.LG ↗ · 2026-07-08 Cached

Introduces Stochastic Token Steering (STS) and Stochastic Block Steering (SBS) for LLM activation steering, which probabilistically gate steering signals per token or per sequence. Shows that steering only 50% of tokens recovers most of the dense-steering effect while preserving fluency, and that the behavioral outcome is rate-limited by cumulative signal dosage.

0 favorites 0 likes
#activation-steering

Controlling Tool Use with Heading-Specific Activation Steering

arXiv cs.AI ↗ · 2026-07-08 Cached

This paper investigates whether tool-use decisions in large language models have stable internal representations that can be extracted and manipulated via activation steering, demonstrating that heading-specific steering vectors can suppress unnecessary tool use across five open-source models and three domains. The geometric analysis reveals that tool-invocation steps exhibit diffuse, bimodal alignment rather than the clean linear structure expected for parametrically grounded concepts.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback