预训练ASR伪标签在嘈杂警察音频中的应用
摘要
本文系统评估了伪标签技术在适配预训练ASR模型(如Whisper和Qwen3-ASR)到嘈杂警察音频中的应用,引入了一种LLM-as-a-judge过滤方法和跨模型范式,以降低词错误率。
arXiv:2609.30469v1 Announce Type: new
Abstract: Pretrained ASR systems perform poorly on noisy Broadcast Police Communication (BPC), hindering efforts to understand police decision-making. Pseudo-labeling offers an unsupervised path to improve ASR without expensive human labels, but the efficacy of this approach on very noisy domains is not known. In this work, we systematically assess the opportunities and limits of pseudo-labeling to adapt foundation ASR models (Whisper and Qwen3-ASR) to noisy BPC domain corpora from Baltimore and Chicago. We demonstrate that existing internal confidence metrics (log-probabilities and STAR scores) fail to distinguish between high and low quality BPC pseudo-labels, and we introduce an external LLM-as-a-judge filtering paradigm that leverages parametric knowledge to discard contextually implausible transcripts. Our LLM-judging filters more aggressively than internal metrics and significantly reduces WER of the pseudo-labeled training sets across the Baltimore and Chicago BPC corpora, though a substantial gap remains relative to an oracle filter. We also introduce a new cross-model pseudo-labeling paradigm where one model is finetuned with pseudo-labels from the other, and we identify this method as a promising direction for future pseudo-labeling work.
查看缓存全文
缓存时间: 2026/09/28 09:36
# Pretrained ASR Pseudo-labeling for Noisy Police Audio Source: [https://arxiv.org/abs/2609.30469](https://arxiv.org/abs/2609.30469) [View PDF](https://arxiv.org/pdf/2609.30469) > Abstract:Pretrained ASR systems perform poorly on noisy Broadcast Police Communication \(BPC\), hindering efforts to understand police decision\-making\. Pseudo\-labeling offers an unsupervised path to improve ASR without expensive human labels, but the efficacy of this approach on very noisy domains is not known\. In this work, we systematically assess the opportunities and limits of pseudo\-labeling to adapt foundation ASR models \(Whisper and Qwen3\-ASR\) to noisy BPC domain corpora from Baltimore and Chicago\. We demonstrate that existing internal confidence metrics \(log\-probabilities and STAR scores\) fail to distinguish between high and low quality BPC pseudo\-labels, and we introduce an external LLM\-as\-a\-judge filtering paradigm that leverages parametric knowledge to discard contextually implausible transcripts\. Our LLM\-judging filters more aggressively than internal metrics and significantly reduces WER of the pseudo\-labeled training sets across the Baltimore and Chicago BPC corpora, though a substantial gap remains relative to an oracle filter\. We also introduce a new cross\-model pseudo\-labeling paradigm where one model is finetuned with pseudo\-labels from the other, and we identify this method as a promising direction for future pseudo\-labeling work\. ## Submission history From: Kaavya Chaparala \[[view email](https://arxiv.org/show-email/fd056099/2609.30469)\] **\[v1\]**Thu, 24 Sep 2026 19:05:53 UTC \(8,461 KB\)
相似文章
减少六层:Whisper编码器剪枝与无标签恢复
本文介绍了一种剪枝Whisper ASR模型编码器层的方法,将编码器大小减少18.5%,并通过无标签数据蒸馏恢复性能,同时发布了代码和预训练模型供采用。
半监督联邦ASR的实用方案:在线伪标签与服务器更新稳定化
本文提出了一种半监督联邦ASR的实用方案,采用在线伪标签和服务器更新稳定化,在域内和跨域设置中均展示了相对于先前方法的显著改进。
@FeitengLi: 其实这些问题都能很好的解决了 1. 扔掉 whisper,换 ASR 模型,Qwen3-ASR 就很不错幻觉很少、也有一些别的ASR选择,whisper 幻觉多也要求 30s片段,Qwen3-ASR 塞更长的音频识别越准确,最大支持 20…
推荐使用Qwen3-ASR替代Whisper以减少幻觉,使用LattifAI工具进行精确的音文本对齐和字幕生成,并介绍自己的OmniVAD-Kit项目用于语音活动检测。
转录儿童语音:ASR性能与获取可靠的正字法转写
这篇论文评估了九种ASR模型(Whisper、Parakeet、Wav2Vec2)在荷兰语儿童语音数据集JASMIN和DART上的表现,发现微调后的Whisper-medium取得了最佳性能(在JASMIN上WER为5.54%,在DART上为70.37%)。它还提出了一种选择方法,能够以高精度自动识别发音正确的录音片段,从而减少人工验证的需求。
嘈杂环境中的语音代理
一家语音公司训练了一个模型,该模型能消除噪声并识别主要说话者,在嘈杂环境中,领先的ASR模型的词错误率降低了50%。