Tag
This paper proposes SEAV, a verification-centric framework for evaluating jailbreak robustness in large language models by assessing response validity and correctness, significantly reducing false-positive rates in safety assessments.
The paper introduces SCEval, a diagnostic evaluation protocol that applies structural corruptions to test the fragility of omni-modal large language models, revealing that clean accuracy does not ensure reliable cross-modal reasoning.
TestNav is a Pareto-guided framework for compositional robustness testing in deep learning models, optimizing for both performance degradation and input fidelity to identify severe yet realistic failures.