attack-selection

Tag

Cards List
#attack-selection

Attack Selection in Agentic AI Control Evaluations Meaningfully Decreases Safety

arXiv cs.AI · 2026-06-08 Cached

This paper demonstrates that allowing attackers to strategically choose when to attack (attack selection) in agentic AI control evaluations significantly reduces measured safety, suggesting that current evaluations may overestimate safety against selective attackers.

0 favorites 0 likes
← Back to home

Submit Feedback