标签
This preprint introduces a generation-aligned diagnostic ladder that separates decision-rule misalignment from readout-coverage limitations in speech language models, showing that state decoding far exceeds generated accuracy in emotion recognition tasks.
本文介绍了MERaLiON-GR,一个面向英语和东南亚语言的语音性别识别模型,基于MERaLiON-SpeechEncoder-2,使用LoRA和ECAPA-TDNN头部进行微调,在多语言基准上取得了最先进的性能。