Tag
Topological Steering is a new framework for controlling large language model behavior using topological data analysis to capture global structures in activation spaces, enabling more robust behavioral control.
This paper proposes attribute-based activation steering to tailor LLM explanations to specific groups, achieving better specificity and factuality compared to prompting and state-of-the-art baselines.
This paper tests whether decodable empathy directions in LLMs can reliably shift automated empathy scores, finding that affective facet control is partial and cognitive steering is inconsistent, highlighting that detection does not imply control.
CircuitSteer is a novel framework that uses sparse autoencoders to identify and manipulate multi-layer semantic circuits in LLMs, enabling more robust and fluency-preserving behavioral steering compared to single-layer methods like CAA.
UniSteer introduces a text-guided activation flow matching method to learn a universal conditional velocity field in activation space, enabling versatile LLM behavior control and classification tasks without task-specific intervention modules.
This paper introduces Prototype-Based Sparse Steering, a method that applies sparse autoencoders to attention query activations in LLMs, then uses gradient-based optimization during inference to steer generation toward target behaviors. The approach is validated in both a logical planning task and a stylistic educational domain, demonstrating interpretable and disentangled control.
NousResearch releases Contrastive Neuron Attribution (CNA), a method to steer LLM behavior by ablating sparse MLP circuits without training autoencoders or degrading benchmarks, validated on refusal circuits across models up to 70B parameters.