robustness-testing

Tag

Cards List
#robustness-testing

Validity-Aware Jailbreak Evaluation for Large Language Models

arXiv cs.AI · yesterday Cached

This paper proposes SEAV, a verification-centric framework for evaluating jailbreak robustness in large language models by assessing response validity and correctness, significantly reducing false-positive rates in safety assessments.

0 favorites 0 likes
#robustness-testing

Modality Fault Lines: Structural Corruptions Reveal Fragile Omni-Modal Reasoning

arXiv cs.CL · 2d ago Cached

The paper introduces SCEval, a diagnostic evaluation protocol that applies structural corruptions to test the fragility of omni-modal large language models, revealing that clean accuracy does not ensure reliable cross-modal reasoning.

0 favorites 0 likes
#robustness-testing

TESTNAV: Pareto-Guided Search for Compositional Robustness Testing

arXiv cs.AI · 2026-08-21 Cached

TestNav is a Pareto-guided framework for compositional robustness testing in deep learning models, optimizing for both performance degradation and input fidelity to identify severe yet realistic failures.

0 favorites 0 likes
← Back to home

Submit Feedback