Tag
This paper critiques the Beckmann-Butlin framework for LLM individuation, arguing that persona vectors are regime-dependent rather than substrate-identical, and provides empirical experiments on Qwen3 and Mistral models showing cross-regime asymmetries. It proposes a (vehicle, regime) pairing as the unit of representational content.
This paper investigates whether off-the-shelf persona steering vectors can reduce sycophancy in large language models, finding they achieve 68-98% of the effect of targeted Contrastive Activation Addition (CAA) without requiring sycophancy-specific training data, and that sycophancy is better understood as a persona-level property.