Tag
This paper introduces a cost-aware evaluation framework for security AI agents, measuring not just success rate but also inference and tool costs. It finds that offensive CTF performance scales with compute, while defensive SOC tasks depend more on disciplined tool use than raw reasoning budget.