Tag
This paper introduces StreamFraudNet, a weakly supervised incremental model for detecting phone scams from raw speech, achieving a ROC-AUC of 0.9953 and operating in real-time to provide early warnings during calls.
This paper presents a case study using unsupervised articulatory probing to examine how self-supervised speech models encode phonetic features across Mandarin sub-dialects, finding that salient features like labiality remain stable while finer spectral distinctions show dialect-dependent variation.