jailbreak-evaluation

Tag

Cards List
#jailbreak-evaluation

Risk Under Pressure: Compute-Aware Evaluation of Adversarial Robustness in Language Models

arXiv cs.LG · 2026-06-11 Cached

This paper introduces a compute-aware evaluation framework for adversarial robustness of LLMs, proposing risk-compute curves and metrics based on FLOPs to better assess attack costs, finding that alignment training has non-monotonic effects and compute costs vary across models and harm categories.

0 favorites 0 likes
← Back to home

Submit Feedback