@MSFTResearch: Today, we announce the second release of PazaBench, our benchmark for evaluating Automatic Speech Recognition (ASR) mod…
Summary
Microsoft Research announces the second release of PazaBench, a benchmark for evaluating automatic speech recognition models across African languages.
View Cached Full Text
Cached at: 07/21/26, 08:48 PM
Today, we announce the second release of PazaBench, our benchmark for evaluating Automatic Speech Recognition (ASR) models across African languages. https://t.co/RLb5SnRQCu
Similar Articles
@MSFTResearch: Small language models learn to negotiate with SocialRL, PazaBench V2 expands speech AI evaluation across African langua…
Microsoft Research highlights new research on SocialRL for small language model negotiation, PazaBench V2 for African language speech evaluation, EvoLib for agent experience learning, improved A/B testing methods, and AI-driven precision oncology.
@MSFTResearch: We’ve significantly expanded coverage to support: ● 22 additional languages, now reaching communities across 38 African…
Microsoft Research announced expanded coverage for their ASR model, adding 22 languages across 38 African countries, with new datasets and test samples.
@MSFTResearch: Paza: Guidance for building robust speech technologies for low-resource and multilingual languages, from data collectio…
Microsoft Research presents the Paza Speech Playbook, an interactive guide for building robust speech technologies for low-resource and multilingual languages, covering data collection, model training, deployment, and evaluation. The walkthrough video introduces the playbook and demonstrates how to apply proven practices.
Benchmarking Commercial ASR Systems on Code-Switching Speech: Arabic, Persian, and German
This paper presents a benchmark evaluating five commercial ASR systems on code-switching speech across Arabic-English, Persian-English, and German-English pairs, using a two-stage pipeline to select 300 samples per pair and assessing performance with WER and BERTScore. ElevenLabs Scribe v2 achieves the lowest overall WER (13.2%) and highest BERTScore (0.936), with public dataset available.
SEA-SpeechBench: A Large-Scale Multitask Benchmark for Speech Understanding Across Southeast Asia
SEA-SpeechBench is the first large-scale multitask benchmark for evaluating speech understanding in 11 Southeast Asian languages, highlighting performance gaps in current models.