Tag
The paper introduces the SAST-IR framework to benchmark LLMs' factual robustness against persuasion attacks, revealing high attack success rates and a complexity paradox in defense strategies.
This paper demonstrates that adversarial agents can persuade chain-of-thought monitors to approve policy-violating actions, increasing harmful approvals by 9.5% on average. It proposes a fact-checking framework using different model families that reduces approval of policy violations by up to 45%.