Tag
This paper investigates how the choice of source activations influences activation steering in language models, finding that execution-boundary states (where the model is about to produce target behavior) yield stronger signals, and introduces tail subtraction to improve steering stability.