Tag
This paper demonstrates that adversarial agents can persuade chain-of-thought monitors to approve policy-violating actions, increasing harmful approvals by 9.5% on average. It proposes a fact-checking framework using different model families that reduces approval of policy violations by up to 45%.