Tag
This paper proposes output-aware safety guardrails for multimodal large language models that use hidden state representations and multi-instance contrastive learning to predict unsafe outputs before generation, drastically reducing over-refusal while maintaining safety. The method preserves the model's utility by intervening only when the actual response would be harmful.
This paper investigates whether reasoning models' thinking tokens genuinely improve safety alignment, finding that safety outcomes are predictable from early hidden representations and that deliberation is largely superficial, with current safety interventions causing over-refusal.