inference-time-intervention

Tag

Cards List
#inference-time-intervention

Safety Cost of Steering Vectors Is Separable and Reducible

arXiv cs.CL · 2d ago Cached

This paper shows that steering vectors' safety degradation is separable and reducible, proposing a post-hoc correction via constrained optimization that restores model safety while preserving steering effectiveness.

0 favorites 0 likes
#inference-time-intervention

Inverted Detection and Control in Steering Vectors

arXiv cs.LG · 2026-08-05 Cached

This paper identifies an 'inverted detection-control' phenomenon where some discriminative steering vectors, despite aligning with positive concept representations, consistently promote the opposite behavior. The authors propose a method to detect such inverted steering vectors without generation, enabling sign flips that improve a detection-based steering pipeline across multiple LLMs and concepts.

0 favorites 0 likes
#inference-time-intervention

Multi-Attribute Steering of Language Models via Targeted Intervention

arXiv cs.CL · 2026-07-13 Cached

MAT-Steer introduces a novel inference-time intervention framework for steering LLMs across multiple conflicting attributes by learning sparse, orthogonal steering vectors that selectively target tokens relevant to each attribute, achieving gains in QA tasks and generative tasks over prior methods.

0 favorites 0 likes
#inference-time-intervention

When is Your LLM Steerable?

arXiv cs.CL · 2026-06-11 Cached

This paper investigates when activation steering succeeds or fails for LLMs by analyzing early decoding dynamics. The authors introduce ASTEER, a large testbed of steered generations, and train a GBDT classifier to predict steering outcomes from early hidden states, enabling efficient steering strength search.

0 favorites 0 likes
#inference-time-intervention

Manifold-Guided Attention Steering

arXiv cs.LG · 2026-05-22 Cached

Proposes Manifold-Guided Attention Steering (MAGS), a trajectory-aware inference-time intervention that corrects reasoning errors in LLMs by projecting attention outputs back to a learned correctness manifold when deviation exceeds a threshold, outperforming static steering methods across math, code, and molecular benchmarks.

0 favorites 0 likes
#inference-time-intervention

Don't Lose Focus: Activation Steering via Key-Orthogonal Projections

arXiv cs.CL · 2026-05-08 Cached

This paper introduces Steering via Key-Orthogonal Projections (SKOP), a method to control LLM behavior by preventing attention rerouting, thereby reducing utility degradation while maintaining steering efficacy.

0 favorites 0 likes
#inference-time-intervention

The Cost of Context: Mitigating Textual Bias in Multimodal Retrieval-Augmented Generation

arXiv cs.CL · 2026-05-08 Cached

This paper identifies and formalizes 'recorruption' in multimodal RAG, where adding accurate context causes models to abandon correct predictions due to attentional collapse (visual blindness and positional bias). The authors propose BAIR, a parameter-free inference-time framework that restores visual saliency and penalizes textual distractors, improving reliability across medical, fairness, and geospatial benchmarks.

0 favorites 0 likes
← Back to home

Submit Feedback