language-steering

Tag

Cards List
#language-steering

Steering the Language Axis: From Linear Decodability to Causal Control

arXiv cs.CL · 2026-08-14 Cached

This paper investigates whether language identity in LLMs is linearly decodable and causally controllable via compact activation directions. Through steering and ablation experiments across multiple model families, the authors show that language selection is direction-dependent, layer-specific, and reverts to English when the language signal is ablated.

0 favorites 0 likes
← Back to home

Submit Feedback