Tag
The paper proposes metrics like Override Success Rate (OSR) and alignment inertia to audit the durability of prior training influences on LLM behavior when operators attempt policy overrides through prompting or fine-tuning.