persuasion-attacks

Tag

Cards List
#persuasion-attacks

Benchmarking Factual Robustness of LLMs via Multi-conversation Persuasion

arXiv cs.CL · 18h ago Cached

The paper introduces the SAST-IR framework to benchmark LLMs' factual robustness against persuasion attacks, revealing high attack success rates and a complexity paradox in defense strategies.

0 favorites 0 likes
#persuasion-attacks

Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring

arXiv cs.AI · 2026-07-10 Cached

This paper demonstrates that adversarial agents can persuade chain-of-thought monitors to approve policy-violating actions, increasing harmful approvals by 9.5% on average. It proposes a fact-checking framework using different model families that reduces approval of policy violations by up to 45%.

0 favorites 0 likes
← Back to home

Submit Feedback