Tag
This paper proposes a robustness evaluation framework for low-resource multilingual text-to-speech systems, focusing on complex text inputs like numbers and named entities, and introduces a Text Risk Score for pre-synthesis risk diagnosis.
This paper proposes counterfactual marginalisation as a test-time evaluation framework for assessing the robustness of machine learning models to nuisance variables like demographics in medical image analysis, using counterfactual image generation and prediction averaging.
The paper introduces a controlled benchmark for evaluating Large Language Models' robustness in step-level mathematical verification, revealing significant performance degradation on perturbed solution traces.
This paper investigates the safety of large language models (LLMs) beyond text inputs by examining emoji-augmented prompts, revealing gaps in current safety evaluations and model-dependent vulnerabilities.