Tag
MANCE proposes a manifold-aware method for erasing concepts like gender or safety from model activations while minimizing collateral damage to other concepts, achieving state-of-the-art across 119 settings.
This paper investigates the geometry of truth in LLM reasoning chains and proposes DynaSteer, a dynamic representation editing framework that uses pattern clustering and Fisher-LDA to steer trajectories towards truth while avoiding noise. Experiments show effectiveness on MATH benchmarks and generalization to coding tasks.