llm-steering

Tag

Cards List
#llm-steering

Topological Steering

arXiv cs.LG · 5d ago Cached

Topological Steering is a new framework for controlling large language model behavior using topological data analysis to capture global structures in activation spaces, enabling more robust behavioral control.

0 favorites 0 likes
#llm-steering

Attribute-Based Activation Steering of LLMs for Group-Specific Explanation Generation

arXiv cs.CL · 6d ago Cached

This paper proposes attribute-based activation steering to tailor LLM explanations to specific groups, achieving better specificity and factuality compared to prompting and state-of-the-art baselines.

0 favorites 0 likes
#llm-steering

Detection != Reliable Control: Decodable Empathy Directions Yield at Most Partial Shifts in Automated Empathy Scores

arXiv cs.CL · 2026-08-27 Cached

This paper tests whether decodable empathy directions in LLMs can reliably shift automated empathy scores, finding that affective facet control is partial and cognitive steering is inconsistent, highlighting that detection does not imply control.

0 favorites 0 likes
#llm-steering

CircuitSteer: Geometrically Aligned Multi-Layer Steering via Sparse Autoencoder Circuits

arXiv cs.LG · 2026-08-07 Cached

CircuitSteer is a novel framework that uses sparse autoencoders to identify and manipulate multi-layer semantic circuits in LLMs, enabling more robust and fluency-preserving behavioral steering compared to single-layer methods like CAA.

0 favorites 0 likes
#llm-steering

UniSteer: Text-Guided Flow Matching in Activation Space for Versatile LLM Steering

Hugging Face Daily Papers · 2026-05-28 Cached

UniSteer introduces a text-guided activation flow matching method to learn a universal conditional velocity field in activation space, enabling versatile LLM behavior control and classification tasks without task-specific intervention modules.

0 favorites 0 likes
#llm-steering

Steered Generation via Gradient-Based Optimization on Sparse Query Features

arXiv cs.LG · 2026-05-25 Cached

This paper introduces Prototype-Based Sparse Steering, a method that applies sparse autoencoders to attention query activations in LLMs, then uses gradient-based optimization during inference to steer generation toward target behaviors. The approach is validated in both a logical planning task and a stylistic educational domain, demonstrating interpretable and disentangled control.

0 favorites 0 likes
#llm-steering

@NousResearch: Today we release Contrastive Neuron Attribution (CNA), a method for steering LLM behavior by identifying and ablating s…

X AI KOLs Following · 2026-05-19 Cached

NousResearch releases Contrastive Neuron Attribution (CNA), a method to steer LLM behavior by ablating sparse MLP circuits without training autoencoders or degrading benchmarks, validated on refusal circuits across models up to 70B parameters.

0 favorites 0 likes
← Back to home

Submit Feedback