safety-guards

Tag

Cards List
#safety-guards

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models

arXiv cs.AI · 2026-08-05 Cached

This paper identifies a refusal-cue shortcut in safety guard models, where inserting refusal expressions into harmful responses can flip their harmless classification. The authors audit datasets like WildGuardMix and GR-Train, show the issue persists in official models such as LlamaGuard3 and Qwen3Guard, and propose a post-hoc intervention to suppress shortcut-associated components.

0 favorites 0 likes
← Back to home

Submit Feedback