defensive-deception

Tag

Cards List
#defensive-deception

Fool's Gold: Defensive Deception Against Safety-Removal Attacks on Open-Weight Models

arXiv cs.AI · 15h ago Cached

This paper introduces 'Fool's Gold', a defensive deception technique called decoy hardening that trains open-weight language models to generate falsified answers when safety alignment is removed, while preserving normal behavior, tested across multiple model families and sizes.

0 favorites 0 likes
#defensive-deception

Fool's Gold (18 minute read)

TLDR AI · 19h ago Cached

The paper introduces 'Fool's Gold,' a defensive deception technique that trains open-weight models to generate decoy responses with falsified critical elements when subjected to safety-removal attacks, effectively poisoning the attack's payoff. Tested on multiple models, it achieves high decoy rates without degrading benign performance.

0 favorites 0 likes
← Back to home

Submit Feedback