steering-vectors

Tag

Cards List
#steering-vectors

Safety Cost of Steering Vectors Is Separable and Reducible

arXiv cs.CL · 2026-08-11 Cached

This paper shows that steering vectors' safety degradation is separable and reducible, proposing a post-hoc correction via constrained optimization that restores model safety while preserving steering effectiveness.

0 favorites 0 likes
#steering-vectors

Cross-Architecture Steering Transfer in Language Models: A Systematic Empirical Study

arXiv cs.CL · 2026-08-07 Cached

A systematic empirical study showing that concept directions extracted from one language model can steer other independently trained models when sufficient scale (≥1.7B parameters) is reached, providing functional evidence for the Platonic Representation Hypothesis and highlighting scale thresholds for cross-model interpretability tools.

0 favorites 0 likes
#steering-vectors

Inverted Detection and Control in Steering Vectors

arXiv cs.LG · 2026-08-05 Cached

This paper identifies an 'inverted detection-control' phenomenon where some discriminative steering vectors, despite aligning with positive concept representations, consistently promote the opposite behavior. The authors propose a method to detect such inverted steering vectors without generation, enabling sign flips that improve a detection-based steering pipeline across multiple LLMs and concepts.

0 favorites 0 likes
#steering-vectors

On the Generalization of Steering Vectors for Chain-of-Thought Faithfulness

arXiv cs.AI · 2026-08-03 Cached

This paper investigates how well activation steering for improving chain-of-thought faithfulness generalizes across cue types, datasets, and steering vector construction methods across several Gemma and Qwen models. The authors find that when steering is effective, it generalizes broadly, and the evaluation setting matters more than the training setting or vector construction method.

0 favorites 0 likes
#steering-vectors

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models

Hugging Face Daily Papers · 2026-07-28 Cached

This paper investigates why multimodal LLMs fail on vision-centric tasks when visual evidence conflicts with language priors, introducing the WhatIfVis benchmark and showing that models often encode visual evidence but cannot reliably control their reliance on it.

0 favorites 0 likes
#steering-vectors

Probabilistic Concept-Aware Steering for Trustworthy LLM Inference

arXiv cs.AI · 2026-07-22 Cached

This paper introduces the Probabilistic Concept-Aware Steering (PCS) framework for LLM inference, which uses concept-driven steering vector retrieval and probabilistic strength calibration to improve interpretability, optimality, and generalizability, achieving over 30% higher direction accuracy and over 89% steering accuracy on multiple datasets.

0 favorites 0 likes
#steering-vectors

J-space comparisons across open models

Hacker News Top · 2026-07-15 Cached

This blog post extends Anthropic's verbalizable workspace paper by measuring how far steering directions from middle layers reach, when the structure forms during training, whether it transfers between models, and how it scales, all on open models.

0 favorites 0 likes
#steering-vectors

On the Limits of Steering Vectors for Preference-Aligned Generation

arXiv cs.CL · 2026-07-03 Cached

This paper systematically studies the limitations of steering vectors for controlled text generation, finding that their effectiveness varies across traits, degrades on task transfer, and suffers from composition tradeoffs.

0 favorites 0 likes
#steering-vectors

Harnessing the Latent Space: From Steering Vectors to Model Calibrators for Control and Trust

arXiv cs.CL · 2026-07-02 Cached

This paper proposes using steering vectors for control over language model behavior and latent space-based calibrators to assess trustworthiness, aiming to demystify internal representations and build more reliable AI systems.

0 favorites 0 likes
#steering-vectors

Search for Truth from Reasoning: A Dynamic Representation Editing Framework for Steering LLM Trajectories

arXiv cs.AI · 2026-06-30 Cached

This paper investigates the geometry of truth in LLM reasoning chains and proposes DynaSteer, a dynamic representation editing framework that uses pattern clustering and Fisher-LDA to steer trajectories towards truth while avoiding noise. Experiments show effectiveness on MATH benchmarks and generalization to coding tasks.

0 favorites 0 likes
#steering-vectors

SALSA: Speech Aware LLM Adaptation via Learned Steering Activation Vectors

arXiv cs.CL · 2026-06-02 Cached

SALSA introduces a lightweight adaptation method for speech-aware LLMs that learns layer-wise steering vectors via supervised objective, achieving significant improvements (up to 46.8% relative) on out-of-domain speech benchmarks, and shows that steering the encoder layers is more effective than modifying the LLM backbone.

0 favorites 0 likes
#steering-vectors

Predicting Where Steering Vectors Succeed

arXiv cs.CL · 2026-04-20 Cached

This paper introduces the Linear Accessibility Profile (LAP), a diagnostic method using logit lens to predict steering vector effectiveness across model layers, achieving ρ=+0.86 to +0.91 correlation on 24 concept families across five models. The work provides a systematic framework to determine which layers and concepts are suitable for steering interventions, replacing ad-hoc trial-and-error approaches.

0 favorites 0 likes
← Back to home

Submit Feedback