Tag
This paper introduces a mechanistic approach to improve LLM safety by characterizing a circuit for refusal behavior and using circuit-guided weight scaling, enhancing safety rates by 26.5% under attacks with minimal accuracy loss.
This paper proposes computation-efficient strategies for latent adversarial training (LAT) of LLMs, using low-rank representation fine-tuning and circuit-guided surrogate models to reduce per-step FLOPs by 48.1% while requiring only 0.0118% trainable parameters.