capability-confound

Tag

Cards List
#capability-confound

Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks

arXiv cs.AI ↗ · 2026-08-03 Cached

This paper audits four agent-safety benchmarks (R-Judge, InjecAgent, AgentHarm, AgentDojo) across many models, showing that their scores are confounded by capability, metrics have artifacts, and rankings disagree across benchmarks, undermining interchangeable safety claims.

0 favorites 0 likes
← Back to home

Submit Feedback