Tag
The paper presents TeleAntiFraud 2.0, an audio-based benchmark for evaluating telecom fraud detection models using a mixed-tree generation pipeline and frozen monthly sets to address evolving fraud scripts and near-domain negatives.
SEA-SpeechBench is the first large-scale multitask benchmark for evaluating speech understanding in 11 Southeast Asian languages, highlighting performance gaps in current models.