@MSFTResearch: Today, we announce the second release of PazaBench, our benchmark for evaluating Automatic Speech Recognition (ASR) mod…
Summary
Microsoft Research announces the second release of PazaBench, a benchmark for evaluating automatic speech recognition models across African languages.
View Cached Full Text
Cached at: 07/21/26, 08:48 PM
Today, we announce the second release of PazaBench, our benchmark for evaluating Automatic Speech Recognition (ASR) models across African languages. https://t.co/RLb5SnRQCu
Similar Articles
@MSFTResearch: We’ve significantly expanded coverage to support: ● 22 additional languages, now reaching communities across 38 African…
Microsoft Research announced expanded coverage for their ASR model, adding 22 languages across 38 African countries, with new datasets and test samples.
Benchmarking Commercial ASR Systems on Code-Switching Speech: Arabic, Persian, and German
This paper presents a benchmark evaluating five commercial ASR systems on code-switching speech across Arabic-English, Persian-English, and German-English pairs, using a two-stage pipeline to select 300 samples per pair and assessing performance with WER and BERTScore. ElevenLabs Scribe v2 achieves the lowest overall WER (13.2%) and highest BERTScore (0.936), with public dataset available.
BlasBench: An Open Benchmark for Irish Speech Recognition
BlasBench introduces an open evaluation benchmark for Irish speech recognition with Irish-aware text normalization that preserves linguistic features like fadas, lenition, and eclipsis. The paper benchmarks 12 ASR systems across four architecture families, revealing significant generalization gaps and showing that existing multilingual systems struggle with Irish due to inadequate normalization.
Introducing the FFASR Leaderboard: Benchmarking ASR in the Real World
Introduces the FFASR Leaderboard, an open, community-driven benchmark for evaluating automatic speech recognition models under realistic far-field acoustic conditions, highlighting the significant performance gap between near-field and far-field scenarios.
Benchmarking Human and Automatic Speech Recognition of Diverse Speech: Initial Results
This paper compares state-of-the-art ASR systems to human listeners on recognizing diverse Dutch speech, finding that ASR systems match or exceed human performance in some cases, with Google Telephony leading. It highlights the impact of speaker age, regional accents, and test set selection on benchmarking conclusions.