Tag
This paper proposes attribute-based activation steering to tailor LLM explanations to specific groups, achieving better specificity and factuality compared to prompting and state-of-the-art baselines.
This paper investigates how large language models implicitly personalize outputs based on demographic cues, locating an internal activation signal that tracks these shifts and showing that removing this signal can suppress the behavior.
This paper introduces a misinformation detection framework using activation engineering, projecting last-token activations onto a learned 'falsehood direction' in LLM residual streams. It evaluates across Gemma, Llama, and Qwen models on fact-checking benchmarks, showing truthfulness is linearly separable in latent space.
This paper proposes compressing instruction prompts into a single activation vector via learned weighted sums of intermediate layer activations, achieving under 2% accuracy drop and revealing insights into LLM activation space structure.
UniSteer introduces a text-guided activation flow matching method to learn a universal conditional velocity field in activation space, enabling versatile LLM behavior control and classification tasks without task-specific intervention modules.