Benchmarking Human and Automatic Speech Recognition of Diverse Speech: Initial Results
Summary
This paper compares state-of-the-art ASR systems to human listeners on recognizing diverse Dutch speech, finding that ASR systems match or exceed human performance in some cases, with Google Telephony leading. It highlights the impact of speaker age, regional accents, and test set selection on benchmarking conclusions.
View Cached Full Text
Cached at: 07/22/26, 08:24 AM
# Benchmarking Human and Automatic Speech Recognition of Diverse Speech: Initial Results
Source: [https://arxiv.org/html/2607.19049](https://arxiv.org/html/2607.19049)
###### Abstract
Humans are often considered to be the best listeners and seen as the upper\-bound performance of automatic speech recognition \(ASR\) systems\. We present a preliminary comparison of the performances of state\-of\-the\-art ASR systems and Dutch native listeners on the recognition of \`\`diverse'' speech, specifically Dutch child and older adults' speech and Flemish\. Google Telephony outperformed the other ASR systems\. Importantly, the ASR systems showed similar performance to the listeners, and in specific cases even outperformed them\. Slight performance differences between the listeners and ASR systems were found related to speaker’s age and regional accents and utterance length\. Future research should focus on making ASR systems more robust to acoustic variability related to aging and regional accents\. A comparison of ASR recognition performances on the test stimuli and the full Jasmin\-CGN test sets showed the influence of the specific test sets on the conclusions regarding benchmarking human and ASR performance\.
11footnotetext:These authors contributed equally\.## IIntroduction
Automatic speech recognition \(ASR\) systems have come a long way from being a magnitude worse in their performance than human listeners as Lippmann showed in his seminal work comparing human listeners and ASR systems from 1997\[[21](https://arxiv.org/html/2607.19049#bib.bib2)\]\. He concluded that ASR performance could be improved if research would focus on improving acoustic\-phonetic modelling, increasing robustness against noise and channel variability, and on modelling spontaneous speech\. ASR systems have since then regularly been benchmarked against human listeners to investigate performance differences in the recognition of phonemes\[[6](https://arxiv.org/html/2607.19049#bib.bib31)\], logotomes\[[25](https://arxiv.org/html/2607.19049#bib.bib30),[26](https://arxiv.org/html/2607.19049#bib.bib32)\], and words in quiet and background noise\[[6](https://arxiv.org/html/2607.19049#bib.bib31),[25](https://arxiv.org/html/2607.19049#bib.bib30),[26](https://arxiv.org/html/2607.19049#bib.bib32),[35](https://arxiv.org/html/2607.19049#bib.bib28),[2](https://arxiv.org/html/2607.19049#bib.bib12)\]in order to identify areas of potential improvement of current ASR systems\.
In 2017, ASR performance was shown to be on par with human transcription performance on two conversational English databases\[[42](https://arxiv.org/html/2607.19049#bib.bib35)\]; however, human performance was obtained from professional transcribers who have been found to have a much better performance than quick transcription, which is closer to everyday listening conditions\[[13](https://arxiv.org/html/2607.19049#bib.bib13)\]\. More recently, in 2024, Patman and colleagues\[[30](https://arxiv.org/html/2607.19049#bib.bib33)\]showed that human listeners consistently outperformed Wav2vec 2\.0\[[3](https://arxiv.org/html/2607.19049#bib.bib19)\]for speech produced by speakers with and without wearing a face mask in noise, while Whisper\-large\[[28](https://arxiv.org/html/2607.19049#bib.bib27)\]was found to outperform the human listeners in all conditions except pub noise\. Comparisons of the recognition errors of ASR systems and human listeners have highlighted both similarities and differences\[[42](https://arxiv.org/html/2607.19049#bib.bib35),[23](https://arxiv.org/html/2607.19049#bib.bib29),[24](https://arxiv.org/html/2607.19049#bib.bib34)\]\.
Human listeners are often seen as the upper\-bound performance of ASR systems\. However, the above results show that current state\-of\-the\-art \(SotA\) ASR systems not only reach very good performance for different listening conditions, they in fact in specific conditions outperform human listeners on standard speech\. At the same time, increasingly more research has shown that ASR performance is significantly worse for speakers who deviate from the \`\`standard'' native, adult speaker of a language without a strong accent and speech impediment\. For instance, child speech is recognised worse than adult speech \(e\.g\.,\[[31](https://arxiv.org/html/2607.19049#bib.bib3),[10](https://arxiv.org/html/2607.19049#bib.bib8),[18](https://arxiv.org/html/2607.19049#bib.bib11),[15](https://arxiv.org/html/2607.19049#bib.bib9)\]\) and atypical speech, e\.g\., due to dysarthria\[[32](https://arxiv.org/html/2607.19049#bib.bib54),[41](https://arxiv.org/html/2607.19049#bib.bib57),[22](https://arxiv.org/html/2607.19049#bib.bib55),[17](https://arxiv.org/html/2607.19049#bib.bib56),[1](https://arxiv.org/html/2607.19049#bib.bib58),[9](https://arxiv.org/html/2607.19049#bib.bib59)\], oral cancer\[[16](https://arxiv.org/html/2607.19049#bib.bib10)\]or cleft lip and palate\[[33](https://arxiv.org/html/2607.19049#bib.bib24)\], is worse recognised than typical speech\. Moreover, non\-native\[[40](https://arxiv.org/html/2607.19049#bib.bib14),[29](https://arxiv.org/html/2607.19049#bib.bib15),[44](https://arxiv.org/html/2607.19049#bib.bib6)\]and regionally\-accented\[[10](https://arxiv.org/html/2607.19049#bib.bib8),[18](https://arxiv.org/html/2607.19049#bib.bib11),[19](https://arxiv.org/html/2607.19049#bib.bib23),[37](https://arxiv.org/html/2607.19049#bib.bib16),[36](https://arxiv.org/html/2607.19049#bib.bib21),[34](https://arxiv.org/html/2607.19049#bib.bib17)\]speech are recognised worse than \`\`standard'' speech\.
This paper presents preliminary results investigating the question how SotA ASR systems compare to human listeners on the recognition of non\-standard or \`\`diverse'' conversational speech\. We compare recognition performance of native Dutch listeners to that of three off\-the\-shelf SotA ASR systems on Dutch speech from three diverse speaker groups: child speech, speech from older adults, and Dutch as spoken in Flanders \(i\.e\., Flemish\)\. We aim to identify areas for potential improvement of current ASR system's ability to deal with the acoustic variability of diverse speech to make them more inclusive, i\.e\., accessible to everyone irrespective of how one speaks or the language one speaks\.
## IIMethodology
We ran three separate listening experiments \(see Section[II\-B](https://arxiv.org/html/2607.19049#S2.SS2)\) to investigate the recognition performance of human listeners on child speech, older adults' speech, and Flemish teenager and older adults' speech and compared their recognition results with the recognition performances \(see Section[II\-D](https://arxiv.org/html/2607.19049#S2.SS4)\) of Google Telephony\[[14](https://arxiv.org/html/2607.19049#bib.bib25)\], Whisper\-large\-v3\[[28](https://arxiv.org/html/2607.19049#bib.bib27)\], and a custom Dutch Conformer model \(see Section[II\-C](https://arxiv.org/html/2607.19049#S2.SS3)\), on the exact same stimuli \(see Section[II\-A](https://arxiv.org/html/2607.19049#S2.SS1)\)\.
### II\-AStimuli
We selected 120 stimuli from the child, older adults and Flemish speakers from the Jasmin\-CGN corpus\[[7](https://arxiv.org/html/2607.19049#bib.bib20)\], which is a corpus of spoken Dutch from people living in the Netherlands and the Flanders region\. For each speaker group, we selected 40 stimuli from the human\-computer interaction \(HMI\) test set of the corpus, as this type of speech most closely resembles conversational speech\. For each experiment, we balanced the speech samples as much as possible over different demographic labels\. In all three experiments, 20 utterances were spoken by female speakers and 20 by male speakers\. Additionally:
Child speech \(7\-11 years\)\.For each age, four stimuli were selected of different utterance length, with 1 utterance of 5\-6 words, 1 of 7\-8 words, 1 of 9\-11 words, and 1 utterance of 12\-16 words \(excluding non\-speech sounds; mean word count: 8\.0, SD: 2\.4\)\. Regional accents were balanced as much as possible\.
Older adults' \(OA\) speech \(59\-96 years\)\.To capture the wide variability of regional accents, 5 male and 5 female speakers were selected from each of the 4 dialect regions in Jasmin\-CGN\. Age was distributed as good as possible over the regions and genders\. Each utterance consisted of at least 4 words \(excluding non\-speech; mean word count: 6\.3, SD: 1\.8\)\.
Flemish speech \(13\-17 and 65\-84 years\)\.We tested both teenagers and older adults\. Ten stimuli were selected from each of the 4 dialect regions, within each region balanced for age group and gender\. Each utterance consisted of 6\-16 words \(excluding non\-speech; mean word count: 7\.3, SD: 1\.2\)\.
### II\-BHuman listening experiments
#### II\-B1Participants
Forty\-five adult native listeners of Dutch were recruited from the social and work circles of the experimenters\. Twenty listeners \(10 female and 10 male participants, average age: 39\.5, SD: 9\.5\) participated in the child speech experiment111Originally the participants were split into two groups: ten listeners with and 10 listeners without experience with child speech, however, no performance difference was found between the two groups, hence, we report the aggregated results here\.; 14 participants in the experiment with older adult speech \(7 female and 7 male, average age: 37\.2, SD: 16\.7\), and 11 in the experiment with Flemish speech \(5 female and 6 male, average age 40\.1, SD: 13\.6\)\. No participants reported hearing problems\. Participants did not receive payment for their participation\. Each listener participated only in one experiment, except for two participants who participated in both the child speech and older adult speech experiments\. The study was approved by the ethical committee of our university\.
#### II\-B2Experimental set\-up
All stimuli were normalized to the same volume using the loudnorm filter of FFmpeg\[[38](https://arxiv.org/html/2607.19049#bib.bib18)\]\. The experiments were hosted in Qualtrics, ran on the same laptop, and used the same headphones \(Sennheiser HD 200 Pro\) within an experiment\. All experiments were held in a quiet room\. Prior to the experiments, listeners received instructions, signed a consent form, and had a short familiarisation with the experiment of one stimulus which was also used to adjust the volume to a comfortable level\. During the experiment, one stimulus was shown per page\. The order of the stimuli was randomized per participant to prevent order effects\. Listeners were allowed to listen to each stimulus only once to mimic everyday listening conditions; however, in the Flemish experiment, listeners were allowed to listen twice due to a different setting of the experiment\.
Participants were asked to type what they heard\. After every 10 stimuli, listeners were allowed to take a break for as long as they liked\. Correct transcriptions were not shown during the experiment because this could induce learning due to the pop\-out effect\[[8](https://arxiv.org/html/2607.19049#bib.bib26)\], which could influence the results\. Participants were allowed to compare their answers to the correct transcriptions after conclusion of the experiment, which many did\.
### II\-CAutomatic speech recognition systems
The three models used in this study were selected from the best performing systems in a recent study which compared 12 SotA ASR systems for Dutch on the diverse speech of Jasmin\-CGN\[[43](https://arxiv.org/html/2607.19049#bib.bib1)\]\. The first model is Google Telephony, which, in line with the HMI speech used in this study, is optimised for conversational, telephone speech\. Google Telephony was the second best model \(after Google Chirp\)\. We performed synchronous speech recognition using the Google Cloud API via the 2\.33\.0 version of the Speech\-to\-text V2 python package\. Whisper\-large\-v3 \(Whisper\) is an often used baseline model\. It was the 4th\-to\-6th best model \(depending on the speech type\) in\[[43](https://arxiv.org/html/2607.19049#bib.bib1)\]\. Whisper was downloaded from HuggingFace\. During testing, we employed beam search decoding with a beam size of 10\. We set the task to \`\`transcribe'', the language to \`\`Dutch'', and the temperature parameter to 0\. The third model is a custom Conformer model using XLSR\-53 features\[[5](https://arxiv.org/html/2607.19049#bib.bib22)\]trained with∼700\{\\sim\}700hours of Dutch standard speech from the Corpus Gesproken Nederlands \(CGN;\[[27](https://arxiv.org/html/2607.19049#bib.bib5)\]\), which obtained similar results to Whisper in\[[43](https://arxiv.org/html/2607.19049#bib.bib1)\]\.
### II\-DEvaluation
To make the human transcriptions in line with those of the ASR systems, all transcriptions were converted to lowercase, punctuation was removed \(except for the apostrophe\), any additional spaces were removed, digits/numbers were written out in full, and all non\-linguistic symbols, filler words, and other transcriptions of non\-lexical sounds were removed from the transcriptions of both the human listeners and the ASR systems\. Obvious typing errors and spelling mistakes were corrected when the intended word was fully unambiguous \(e\.g\., \`\`hius''→\\rightarrow\`\`huis'' \(house\); \`\`afentoe''→\\rightarrow\`\`af en toe'' \(sometimes\)\)\.
Recognition performances of the human listeners and the ASR systems were measured in Word Error Rate \(WER\)\. For the comparison of the performance of the human listeners and the three ASR systems, a paired bootstrap was used with 10000 speaker\-based resamples and 95% confidence intervals \(CI\)\[[4](https://arxiv.org/html/2607.19049#bib.bib51),[11](https://arxiv.org/html/2607.19049#bib.bib52)\]\. A difference was considered statistically significant if CI excludes zero\. P\-values are only reported when the results are significant\. Furthermore, we compared the type of errors made by the human listeners and the ASR systems\. Since\[[24](https://arxiv.org/html/2607.19049#bib.bib34)\]found that human errors are highly correlated with the speaker, and large performance disparities have been observed by ASR systems for different speaker groups, we also investigated the effect of the speakers' age, reported gender, and regional accent on the human listeners' and ASR results\. For statistical significance testing, a multi\-factor linear model
WER\_spk∼model×\(gender\+regional\_accents\)\\text\{WER\}\\\_\{\\text\{spk\}\}\\sim\\text\{model\}\\times\(\\text\{gender\}\+\\text\{regional\}\\\_\\text\{accents\}\)using the estimated marginal means package \(emmeans\)\[[20](https://arxiv.org/html/2607.19049#bib.bib53)\]was employed in R to explore the effect of gender and regional accents\. Moreover, we investigated the effect of the number of words in an utterance on the WERs\.
## IIIResults
### III\-ARecognition performance for the diverse speech
Table[I](https://arxiv.org/html/2607.19049#S3.T1)shows the WERs of the human listeners and the three ASR models for the three experiments on child speech, older adults' speech, and Flemish \(split for teen\(ager\) and older adults' speech and averaged over the full stimuli set\)\. For the human listeners, the standard deviations \(SDs\) per speaker group are provided\. These show that while for child speech the human listeners were fairly consistent in their responses \(and errors\), for the other speaker groups, the SD is quite a bit higher, indicating larger variability in the WERs for the individual participants\. Regarding the ASR systems, no significant differences were found between the three models\.
Comparing the performances of the human listeners and the ASR models, for child speech, there are no significant differences in the performance of the human listeners and the ASR systems\. For older adults' speech, there was a significant difference between the human listeners and the ASR systems: both the Google Telephony model \(95% CI \[\+3\.426%, \+16\.148%\],p= \.0031\) and Whisper \(95% CI \[\+2\.454%, \+12\.610%\],p= \.0041\) outperformed the human listeners\. For Flemish, on average, all three ASR systems outperformed the human listeners\. Splitting the results into teenager \(20 stimuli\) and older adults' speech \(20 stimuli\) for easier comparison with the Dutch child and older adults' results, for teenager speech, the Google Telephony model significantly outperformed not only Whisper \(95% CI \[\+2\.381%, \+17\.901%\],p= 0\.0067\) but also the human listeners \(95% CI \[\+6\.935%, \+13\.228%\],p<0\.001\\textit\{p\}<0\.001\)\. For Flemish older adults' speech, there was no significant difference between the human listeners and the ASR systems\.
TABLE I:WER \(%\) of the human listeners and three ASR models on the 40 stimuli of child speech, older adults' \(OA\) speech, and Flemish\. For the human listeners, the standard deviation of the WER is also given\. Bold indicates the best results for a speaker group\.
### III\-BAnalysis of the error types
Analysis of the insertions, deletions, and substitution patterns showed that for all three speaker groups, both the humans and the ASR systems showed very few insertions\. In all cases, the majority of the errors were substitutions, accounting for 8\.9%\-13\.2% of the WER for child speech, 11\.3\-14\.2% for older adults' speech, and 7\.1\-13\.0% of the WER for Flemish\. The Conformer model had the highest substitution rate for child and older adults' speech, while Whisper had the highest rate for Flemish, closely followed by the human listeners \(11\.3%\)\. For child speech and older adults' speech, the human listeners showed a relatively high deletion rate \(5\.3% and 9\.5%\)\. This suggests that human listeners do not always write something down when they do not understand what has been said, while ASR systems will recognise something when there is audio\. These results are somewhat counter those of\[[23](https://arxiv.org/html/2607.19049#bib.bib29)\]who observed far more deletions for the ASR systems, but these concerned primarily discourse markers, while our stimuli were selected to not contain these\. Note however that\[[24](https://arxiv.org/html/2607.19049#bib.bib34)\]found the opposite pattern, with human listeners being likelier to miss discourse markers\.
Analysis of the most common substitutions showed that both the human listeners and ASR systems made errors involvingme\(English:me\),mijn, and its reduced form'm\(English:mine\), which are not only acoustically very similar, but also semantically highly related\. Similar substitutions were found forikand its reduced form'k\(English:I\) andhetand its reduced form't\(English:it\)\. Normalising these transcriptions, reduced the WERs for the human listeners and ASR systems by between 6\-8% but did not change the patterns presented in Table[I](https://arxiv.org/html/2607.19049#S3.T1)\.
### III\-CThe effect of demographic variability
#### III\-C1The effect of the age of the speaker
Figure[1](https://arxiv.org/html/2607.19049#S3.F1)shows the WER of the human listeners and ASR systems split for the child speakers' ages\. Since each bin only contains 10 stimuli, no statistical tests were carried out and no hard conclusions can be drawn; nevertheless, a trend can be observed, which seems to suggest that recognition performance does not improve nor deteriorate with age for the human listeners and the ASR systems\. Further analysis of the high WER for the 9\-year old speakers, particularly for the human listeners, showed that this is primarily due to a single harder\-to\-recognise speaker\.
For the older adults' speech, Figure[2](https://arxiv.org/html/2607.19049#S3.F2)suggests an age trend, with increasing WER for the human listeners and the ASR systems with increasing age \(age is binned into 4 bins\), and particularly so for the human listeners and the Conformer model\. Google Telephony and Whisper seem to be slightly more robust against acoustic variability due to aging\. However, also here the number of stimuli per bin is low\. Future research will need to investigate these trends further\.
Figure 1:WER \(%\) of the human listeners and ASR systems split for the age of the child speakers\.Figure 2:WER \(%\) of the human listeners and ASR systems split per age bin for the older adult speakers\.
#### III\-C2The effect of the gender of the speaker
Table[II](https://arxiv.org/html/2607.19049#S3.T2)shows the WER differences between male and female speakers for the human listeners and the ASR systems for the child, older adults, and Flemish speakers\. For the child and older adults' speakers, the observed performance disparities for the human listeners and the ASR systems were non\-significant\. For Flemish, while the human listeners and Google Telephony again did not show a significant gender gap, Whisper \(t\(40\)=\-2\.149,p=0\.0378\) and the Conformer model \(t\(40\)=\-3\.131,p=0\.0033\) did\.
TABLE II:WER differences \(%\) for the gender of the speaker \(male\-female\) and regional accents \(highest and lowest WER\)\.
#### III\-C3The effect of regional accent of the speaker
Table[II](https://arxiv.org/html/2607.19049#S3.T2)shows the WER difference between the region with the highest and the lowest WER to investigate the effect of regional accent on recognition performance\. For the child speech experiment, speech from three regions was used, while for the older adults and Flemish speech experiments speech from four regions was used\. For the child speech, none of the \(small\) performance difference between the regions with the highest and the lowest WER were significant\. Similarly, despite the much larger WER differences between the regions for Flemish speech, none of these differences were significant\. Only for the older adults, the performance difference for the human listeners was found to be significant \(F\(3, 140\)=3\.260,p=0\.0235\), most likely due to a stronger regional accent for the older adults\. A post hoc Tukey test showed that the speakers from the West region were significantly better recognised than those from the South region \(t\(140\)=\-2\.855,p= 0\.0252\), which is not surprising as most listeners were from the West region\. The ASR systems did not show an effect of regional accent\.
### III\-DThe influence of utterance length
Figures[3](https://arxiv.org/html/2607.19049#S3.F3)and[4](https://arxiv.org/html/2607.19049#S3.F4)show the WERs split by the number of words in the utterance for the child speech and older adults' speech, respectively\. For the child speech, human listeners did not seem to show an effect of utterance length on WER; however, for the ASR systems, there is a downward trend for longer utterances\. This performance improvement for the ASR systems is potentially due to the longer\-range context, which better leverages the language modelling capabilities of large models\. This was also observed for Whisper in\[[43](https://arxiv.org/html/2607.19049#bib.bib1)\]\. For the older adult speakers, we observe a downward trend for the Whisper and the Conformer models, but not so for the human listeners and the Google Telephony model\. Note however that the number of stimuli with the higher number of words is low \(four utterances with length 9 and one with length 12\)\. Overall, there is no \(strong\) effect of utterance length\.
Figure 3:WER per utterance length for the child speech\.Figure 4:WER per utterance length for the older adults' speech\.
### III\-EASR results on the full Jasmin\-CGN test sets
To investigate the generalisability of the ASR results beyond the used stimuli, Table[III](https://arxiv.org/html/2607.19049#S3.T3)shows the WER for the three models on the full test sets of the child speech \(3526 utterances, 1\.45 hours\), older adults' speech \(8292 utterances, 3\.77 hours\), and Flemish speech \(5209 utterances, 2\.79 hours\) of Jasmin\-CGN\. Comparison with the results for the 40 stimuli \(Table[I](https://arxiv.org/html/2607.19049#S3.T1)\) shows that the order of the models' performance remains highly similar\. Nevertheless, we see that the WERs on the full test sets are on average 8\.4% higher than those for the experimental stimuli\. Comparing these results with those of the human listeners in Table[I](https://arxiv.org/html/2607.19049#S3.T1)on the smaller set of only 40 stimuli shows no significant differences between the performances of the human listeners and the ASR systems\.
TABLE III:WER \(%\) of the ASR models on the full child speech, older adults' speech, and Flemish speech test sets of Jasmin\-CGN\. Bold indicates the best results for a speaker group\.ASRDutchFlemishChildOATeenOAGoogle Telephony20\.324\.119\.922\.5Whisper27\.029\.329\.624\.9Conformer27\.829\.127\.122\.8
## IVGeneral discussion and conclusion
We presented preliminary results addressing the question how state\-of\-the\-art ASR systems compare to human listeners on the recognition of non\-standard or \`\`diverse'' conversational Dutch speech\. Specifically, we compared recognition performance of native Dutch listeners with those of Google Telephony and Whisper and a custom\-trained Conformer model for Dutch child speech and older adults' speech, and Flemish\. Overall Google Telephony obtained the best results of the three ASR systems222Google Telephony significantly outperformed Whisper and the custom Conformer model for several speaker groups\. Since the training data for both Google Telephony and Whisper are unknown, we cannot exclude the possibility that Google Telephony was trained with CGN\-Jasmin data\. However, the CGN\-Jasmin data provider requests disclosure explicitly, which was not made for Google Telephony\.and, surprisingly, significantly outperformed the human listeners for older adults' speech and Flemish teenager speech; moreover, Whisper outperformed the human listeners for older adults' speech\. We thus extend the results of\[[42](https://arxiv.org/html/2607.19049#bib.bib35)\]showing parity of ASR and human recognition performance of standard speech to diverse speech\. We show that ASR systems reach parity with human performance for child speech and some systems even exceed human performance for older adults' and Flemish speech\. So while human listeners are often seen as the upper\-bound performance of ASR systems, our results show that ASR systems in fact in specific conditions not only can outperform human listeners on standard\[[30](https://arxiv.org/html/2607.19049#bib.bib33)\]speech but also diverse speech\.
Analysis of the effect of demographic variability of the speakers on recognition performance showed a possible effect of age for older speakers, with increasing WER for both the human listeners and the ASR systems with increasing age, in line with\[[39](https://arxiv.org/html/2607.19049#bib.bib4)\]\. No effect of age was found for child speech in our small sample\. Since the stimuli and listeners used in the three experiments were different, we cannot draw strong conclusions from a comparison across different speaker groups; nevertheless, it is somewhat surprising to see that generally speaking the speech of the older adults was recognised worst by the ASR systems \(except for Flemish older adults' speech\) and the human listeners, while child speech, which acoustically deviates quite substantially from adult speech was overall recognised best\. These results are opposite those reported in the literature for Dutch for a hybrid TDNNF\-HMM model\[[10](https://arxiv.org/html/2607.19049#bib.bib8)\], Wav2Vec 2\.0\[[12](https://arxiv.org/html/2607.19049#bib.bib7)\]and Whisper\-large\-v2\[[12](https://arxiv.org/html/2607.19049#bib.bib7)\], which all had higher WERs for child speech than for older adults' speech\. Future research will need to show whether this is due to the chosen stimuli or whether this is a general effect\. Nevertheless, these results show that more recent ASR models have improved on the recognition of particularly child speech, making them more inclusive for this speaker group\.
The human listeners did not show an effect of gender, which was also the case for the ASR systems, except for Whisper and the Conformer model for the Flemish speech\. These results are largely in contrast to previous findings in the ASR literature \(e\.g\.,\[[10](https://arxiv.org/html/2607.19049#bib.bib8),[18](https://arxiv.org/html/2607.19049#bib.bib11),[19](https://arxiv.org/html/2607.19049#bib.bib23),[37](https://arxiv.org/html/2607.19049#bib.bib16),[36](https://arxiv.org/html/2607.19049#bib.bib21),[34](https://arxiv.org/html/2607.19049#bib.bib17),[12](https://arxiv.org/html/2607.19049#bib.bib7)\]\), which found gender effects for ASR systems, but in line with recent results from SotA ASR systems and models\[[43](https://arxiv.org/html/2607.19049#bib.bib1)\], which only found significant gender\-related performance disparities for Whisper\. An effect of regional accent was only found for human listeners for older adults' speech, not for the ASR systems, which is in contrast to earlier findings\[[10](https://arxiv.org/html/2607.19049#bib.bib8),[18](https://arxiv.org/html/2607.19049#bib.bib11)\]\. These results show that recent ASR systems have become more robust against acoustic variation due to regional accents, potentially due to larger variation in the large amounts of training data that are used for training the SotA systems\.
A comparison of the WER of the ASR systems on the experimental sets and the full Jasmin\-CGN diverse speech tests showed that the test stimuli are an important factor: different or more stimuli lead to different conclusions on the performance gap between human listeners and ASR systems\. Future research will need to focus on different stimuli and larger test set sizes to further investigate the performance differences and similarities between human listeners and ASR systems\. Moreover, for better comparison of the results for the different speaker groups, a similar number of participants from the general population, and an identical experimental set\-up for each speaker group will be used in future research\.
In conclusion, while human listeners are often seen as the best listeners, this does not mean they make no recognition errors\. In fact, their error rates on the diverse speech in this study are similar to and in some cases significantly higher than those of SotA ASR systems\. Nevertheless, overall performance of the ASR systems on diverse speech falls short compared to typical speech\. Although ASR systems have improved a lot since Lippmann's seminal paper\[[21](https://arxiv.org/html/2607.19049#bib.bib2)\], acoustic\-phonetic modelling still requires attention\. Particular areas for further improvement of ASR systems – to make them more inclusive – should focus on improving the modelling of the acoustic variability due to demographic variation related to, specifically, age and regional accents\.
## VAcknowledgements
We are grateful for the contributions of Ansen Weng to the listening experiments regarding the speech of the older adults\.
No generative AI was used for the writing of this paper\. ChatGPT was used for drafting debugging Python scripts for preprocessing the speech files and selecting the stimuli sets\.
## References
- \[1\]A\. Alsayegh and T\. Masood\(2025\)Zero\-shot recognition of dysarthric speech using commercial automatic speech recognition and multimodal large language models\.arXiv preprint arXiv:2512\.17474\.Cited by:[§I](https://arxiv.org/html/2607.19049#S1.p3.1)\.
- \[2\]Amit Juneja\(2012\)A comparison of automatic and human speech recognition in null grammar\.Journal of the Acoustical Society of America3\(131\),pp\. EL256–61\.Cited by:[§I](https://arxiv.org/html/2607.19049#S1.p1.1)\.
- \[3\]A\. Baevski, H\. Zhou, A\. Mohamed, and M\. Auli\(2020\)Wav2Vec 2\.0: a framework for self\-supervised learning of speech representations\.Neural Inf\. Process\. Syst\.\(33\),pp\. 12449–12460\.Cited by:[§I](https://arxiv.org/html/2607.19049#S1.p2.1)\.
- \[4\]M\. Bisani and H\. Ney\(2004\)Bootstrap estimates for confidence intervals in ASR performance evaluation\.InIEEE ICASSP,Vol\.1,pp\. I–409\.Cited by:[§II\-D](https://arxiv.org/html/2607.19049#S2.SS4.p2.1)\.
- \[5\]A\. Conneau, A\. Baevski, R\. Collobert, A\. Mohamed, and M\. Auli\(2021\)Unsupervised cross\-lingual representation learning for speech recognition\.\.InProceedings of Interspeech, Brno, Czechia,pp\. 2426–2430\.Cited by:[§II\-C](https://arxiv.org/html/2607.19049#S2.SS3.p1.1)\.
- \[6\]M\. Cooke and O\. Scharenborg\(2008\)The Interspeech 2008 Consonant Challenge\.\.InProceedings of Interspeech,Brisbane, Australia,pp\. 1765–1768\.Cited by:[§I](https://arxiv.org/html/2607.19049#S1.p1.1)\.
- \[7\]C\. Cucchiarini, H\. Van hamme, O\. Herwijnen, and F\. Smits\(2006\)JASMIN\-CGN: Extension of the Spoken Dutch Corpus with speech of elderly people, children and non\-natives in the human\-machine interaction modality\.InLREC 2006,Genoa, Italy\.Cited by:[§II\-A](https://arxiv.org/html/2607.19049#S2.SS1.p1.1)\.
- \[8\]M\. H\. Davis, I\. S\. Johnsrude, A\. Hervais\-Adelman, K\. Taylor, and C\. McGettigan\(2005\)Lexical information drives perceptual learning of distorted speech: evidence from the comprehension of noise\-vocoded sentences\.\.Journal of Experimental Psychology: General134\(2\),pp\. 222\.Cited by:[§II\-B2](https://arxiv.org/html/2607.19049#S2.SS2.SSS2.p2.1)\.
- \[9\]L\. De Russis and F\. Corno\(2019\)On the impact of dysarthric speech on contemporary ASR cloud platforms\.Journal of Reliable Intelligent Environments5\(3\),pp\. 163–172\.Cited by:[§I](https://arxiv.org/html/2607.19049#S1.p3.1)\.
- \[10\]S\. Feng, B\. M\. Halpern, O\. Kudina, and O\. Scharenborg\(2024\)Towards inclusive automatic speech recognition\.Computer Speech & Language84,pp\. 101567\.Cited by:[§I](https://arxiv.org/html/2607.19049#S1.p3.1),[§IV](https://arxiv.org/html/2607.19049#S4.p2.1),[§IV](https://arxiv.org/html/2607.19049#S4.p3.1)\.
- \[11\]L\. Ferrer, O\. Scharenborg, and T\. Bäckström\(2024\)Good practices for evaluation of machine learning systems\.arXiv preprint arXiv:2412\.03700\.Cited by:[§II\-D](https://arxiv.org/html/2607.19049#S2.SS4.p2.1)\.
- \[12\]M\. Fuckner, S\. Horsman, P\. Wiggers, and I\. Janssen\(2023\)Uncovering bias in asr systems: Evaluating Wav2vec2 and Whisper for dutch speakers\.In2023 International Conference on Speech Technology and Human\-Computer Dialogue \(SpeD\),pp\. 146–151\.Cited by:[§IV](https://arxiv.org/html/2607.19049#S4.p2.1),[§IV](https://arxiv.org/html/2607.19049#S4.p3.1)\.
- \[13\]M\. L\. Glenn, S\. M\. Strassel, H\. Lee, K\. Maeda, R\. Zakhary, and X\. Li\(2010\-05\)Transcription methods for consistency, volume and efficiency\.InProceedings of the Seventh International Conference on Language Resources and Evaluation \(LREC'10\),Valletta, Malta\.Cited by:[§I](https://arxiv.org/html/2607.19049#S1.p2.1)\.
- \[14\]Google Cloud\(2025\)Speech\-to\-text: transcription models\.Note:[https://cloud\.google\.com/speech\-to\-text/docs/transcription\-model](https://cloud.google.com/speech-to-text/docs/transcription-model)Accessed: 2025\-09\-08Cited by:[§II](https://arxiv.org/html/2607.19049#S2.p1.1)\.
- \[15\]P\. Gurunath Shivakumar and S\. Narayanan\(2022\)End\-to\-end neural systems for automatic children speech recognition: An empirical study\.Computer Speech & Language72,pp\. 101289\.Cited by:[§I](https://arxiv.org/html/2607.19049#S1.p3.1)\.
- \[16\]B\. M\. Halpern, S\. Feng, R\. van Son, M\. van den Brekel, and O\. Scharenborg\(2022\)Low\-resource automatic speech recognition and error analyses of oral cancer speech\.Speech Communication141,pp\. 14–27\.Cited by:[§I](https://arxiv.org/html/2607.19049#S1.p3.1)\.
- \[17\]A\. Hernandez, P\. A\. Pérez\-Toro, E\. Noeth, J\. R\. Orozco\-Arroyave, A\. Maier, and S\. H\. Yang\(2022\)Cross\-lingual self\-supervised speech representations for improved dysarthric speech recognition\.InInterspeech,pp\. 51–55\.External Links:[Document](https://dx.doi.org/10.21437/Interspeech.2022-10674),ISSN 2958\-1796Cited by:[§I](https://arxiv.org/html/2607.19049#S1.p3.1)\.
- \[18\]A\. Herygers, V\. Verkhodanova, M\. Coler, O\. Scharenborg, and M\. Georges\(2023\)Bias in Flemish automatic speech recognition\.InESSV Konferenz Elektronische Sprachsignalverarbeitung, Germany,Cited by:[§I](https://arxiv.org/html/2607.19049#S1.p3.1),[§IV](https://arxiv.org/html/2607.19049#S4.p3.1)\.
- \[19\]A\. Koenecke, A\. Nam, E\. Lake, J\. Nudell, M\. Quartey, Z\. Mengesha, C\. Toups, J\. R\. Rickford, D\. Jurafsky, and S\. Goel\(2020\)Racial disparities in automated speech recognition\.Proceedings of the National Academy of Sciences117\(14\),pp\. 7684–7689\.Cited by:[§I](https://arxiv.org/html/2607.19049#S1.p3.1),[§IV](https://arxiv.org/html/2607.19049#S4.p3.1)\.
- \[20\]R\. Lenth\(2023\)Emmeans: estimated marginal means, aka least\-squares means\_\.\.R package version 2\.0\. 1\.Cited by:[§II\-D](https://arxiv.org/html/2607.19049#S2.SS4.p2.2)\.
- \[21\]R\. P\. Lippmann\(1997\)Speech recognition by machines and humans\.Speech Communication\(22\),pp\. 1–15\.Cited by:[§I](https://arxiv.org/html/2607.19049#S1.p1.1),[§IV](https://arxiv.org/html/2607.19049#S4.p5.1)\.
- \[22\]S\. Liu, M\. Geng, S\. Hu, X\. Xie, M\. Cui, J\. Yu, X\. Liu, and H\. Meng\(2021\)Recent progress in the CUHK dysarthric speech recognition system\.IEEE/ACM Transactions on Audio, Speech, and Language Processing29,pp\. 2267–2281\.Cited by:[§I](https://arxiv.org/html/2607.19049#S1.p3.1)\.
- \[23\]A\. Lopez, A\. Liesenfeld, and M\. Dingemanse\(2022\-12–15September\)Evaluation of automatic speech recognition for conversational speech in Dutch, English and German: what goes missing?\.InProceedings of the 18th Conference on Natural Language Processing \(KONVENS 2022\),Potsdam, Germany,pp\. 135–143\.Cited by:[§I](https://arxiv.org/html/2607.19049#S1.p2.1),[§III\-B](https://arxiv.org/html/2607.19049#S3.SS2.p1.1)\.
- \[24\]C\. Mansfield, S\. Ng, G\. Levow, R\. A\. Wright, and M\. Ostendorf\(2021\)Revisiting parity of human vs\. machine conversational speech transcription\.InProceedings Interspeech 2021 – Annual Conference of the International Speech Communication Association,Brno, Czechia,pp\. 1997–2001\.Cited by:[§I](https://arxiv.org/html/2607.19049#S1.p2.1),[§II\-D](https://arxiv.org/html/2607.19049#S2.SS4.p2.1),[§III\-B](https://arxiv.org/html/2607.19049#S3.SS2.p1.1)\.
- \[25\]B\. T\. Meyer and B\. Kollmeier\(2011\)Robustness of spectro\-temporal features against intrinsic and extrinsic variations in automatic speech recognition\.Speech Communication53,pp\. 753–767\.Cited by:[§I](https://arxiv.org/html/2607.19049#S1.p1.1)\.
- \[26\]B\. T\. Meyer\(2013\)What's the difference? Comparing humans and machines on the aurora2 speech recognition task\.InProceedings of Interspeech,Lyon, France,pp\. 2634–2638\.Cited by:[§I](https://arxiv.org/html/2607.19049#S1.p1.1)\.
- \[27\]N\. Oostdijk\(2000\-05\)The Spoken Dutch Corpus\. Overview and first evaluation\.InProceedings of the 2nd International Conference on Language Resources and Evaluation \(LREC'00\),Athens, Greece\.Cited by:[§II\-C](https://arxiv.org/html/2607.19049#S2.SS3.p1.1)\.
- \[28\]OpenAI\(2023\)Whisper\-large\-v3\.External Links:[Link](https://huggingface.co/openai/whisper-large-v3)Cited by:[§I](https://arxiv.org/html/2607.19049#S1.p2.1),[§II](https://arxiv.org/html/2607.19049#S2.p1.1)\.
- \[29\]A\. Palanica, A\. Thommandram, A\. Lee, M\. Li, and Y\. Fossat\(2019\)Do you understand the words that are comin outta my mouth? Voice assistant comprehension of medication names\.NPJ Digital Medicine\(2\(1\)\),pp\. 1–6\.Cited by:[§I](https://arxiv.org/html/2607.19049#S1.p3.1)\.
- \[30\]C\. Patman and E\. Chodroff\(2024\)Speech recognition in adverse conditions by humans and machines\.\.Journal of the Acoustical Society of America \- Express Letters4 11\.Cited by:[§I](https://arxiv.org/html/2607.19049#S1.p2.1),[§IV](https://arxiv.org/html/2607.19049#S4.p1.1)\.
- \[31\]Y\. Qian, K\. Evanini, X\. Wang, C\. M\. Lee, and M\. Mulholland\(2017\)Bidirectional LSTM\-RNN for improving automated assessment of non\-native children’s speech\.InProceedings of Interspeech, Stockholm, Sweden,Cited by:[§I](https://arxiv.org/html/2607.19049#S1.p3.1)\.
- \[32\]E\. Sanders, M\. B\. Ruiter, L\. Beijer, and H\. Strik\(2002\)Automatic recognition of Dutch dysarthric speech, a pilot study\.InProceedings of the International Conference on Spoken Language Processing \(ICSLP\),pp\. 661–664\.Cited by:[§I](https://arxiv.org/html/2607.19049#S1.p3.1)\.
- \[33\]M\. Schuster, A\. Maier, T\. Haderlein, E\. Nkenke, U\. Wohlleben, F\. Rosanowski, U\. Eysholdt, and E\. Nöth\(2006\)Evaluation of speech intelligibility for children with cleft lip and palate by means of automatic speech recognition\.International Journal of Pediatric Otorhinolaryngology70\(10\),pp\. 1741–1747\.Cited by:[§I](https://arxiv.org/html/2607.19049#S1.p3.1)\.
- \[34\]M\. A\. Shariah and M\. Sawalha\(2013\)The effects of speakers’ gender, age, and region on overall performance of Arabic automatic speech recognition systems using the phonetically rich and balanced modern standard Arabic speech corpus\.InProceedings of the 2nd Workshop of Arabic Corpus Linguistics WACL\-2,Cited by:[§I](https://arxiv.org/html/2607.19049#S1.p3.1),[§IV](https://arxiv.org/html/2607.19049#S4.p3.1)\.
- \[35\]C\. Spille, B\. Kollmeier, and B\. T\. Meyer\(2018\)Comparing human and automatic speech recognition in simple and complex acoustic scenes\.Computer Speech & Language52,pp\. 123–140\.Cited by:[§I](https://arxiv.org/html/2607.19049#S1.p1.1)\.
- \[36\]R\. Tatman and C\. Kasten\(2017\)Effects of talker dialect, gender & race on accuracy of Bing speech and YouTube automatic captions\.\.InProceedings of Interspeech, Stockholm, Sweden,pp\. 934–938\.Cited by:[§I](https://arxiv.org/html/2607.19049#S1.p3.1),[§IV](https://arxiv.org/html/2607.19049#S4.p3.1)\.
- \[37\]R\. Tatman\(2017\-04\)Gender and dialect bias in YouTube's automatic captions\.InProceedings of the First ACL Workshop on Ethics in Natural Language Processing,D\. Hovy, S\. Spruit, M\. Mitchell, E\. M\. Bender, M\. Strube, and H\. Wallach \(Eds\.\),Valencia, Spain,pp\. 53–59\.Cited by:[§I](https://arxiv.org/html/2607.19049#S1.p3.1),[§IV](https://arxiv.org/html/2607.19049#S4.p3.1)\.
- \[38\]The FFmpeg developersFFmpeg loudnorm\.External Links:[Link](https://ffmpeg.org/ffmpeg-filters.html)Cited by:[§II\-B2](https://arxiv.org/html/2607.19049#S2.SS2.SSS2.p1.1)\.
- \[39\]R\. Vipperla, S\. Renals, and J\. Frankel\(2010\)Ageing voices: The effect of changes in voice parameters on ASR Performance\.EURASIP Journal on Audio, Speech, and Music Processing2010,pp\. 1–10\.Cited by:[§IV](https://arxiv.org/html/2607.19049#S4.p2.1)\.
- \[40\]Y\. Wu, D\. Rough, A\. Bleakley, J\. Edwards, O\. Cooney, P\. R\. Doyle, L\. Clark, and B\. R\. Cowan\(2020\)See what I’m saying? Comparing intelligent personal assistant use for native and non\-native language speakers\.In22nd International Conference on Human Computer Interaction with Mobile Devices and Services,pp\. 1–9\.Cited by:[§I](https://arxiv.org/html/2607.19049#S1.p3.1)\.
- \[41\]F\. Xiong, J\. Barker, and H\. Christensen\(2019\)Phonetic analysis of dysarthric speech tempo and applications to robust personalised dysarthric speech recognition\.InInternational Conference on Acoustics, Speech and Signal Processing \(ICASSP\),pp\. 5836–5840\.Cited by:[§I](https://arxiv.org/html/2607.19049#S1.p3.1)\.
- \[42\]W\. Xiong, J\. Droppo, X\. Huang, F\. Seide, M\. L\. Seltzer, A\. Stolcke, D\. Yu, and G\. Zweig\(2017\)Toward human parity in conversational speech recognition\.IEEE/ACM Transactions on Audio, Speech, and Language Processing25\(12\),pp\. 2410–2423\.External Links:[Document](https://dx.doi.org/10.1109/TASLP.2017.2756440)Cited by:[§I](https://arxiv.org/html/2607.19049#S1.p2.1),[§IV](https://arxiv.org/html/2607.19049#S4.p1.1)\.
- \[43\]Y\. Zhang, T\. Valck, and O\. Scharenborg\(2026\)Speech recognition performance disparities between Dutch diverse speaker groups\.\.External Links:[Link](https://doi.org/10.1515/phon-2025-0061)Cited by:[§II\-C](https://arxiv.org/html/2607.19049#S2.SS3.p1.1),[§III\-D](https://arxiv.org/html/2607.19049#S3.SS4.p1.1),[§IV](https://arxiv.org/html/2607.19049#S4.p3.1)\.
- \[44\]Y\. Zhang, Y\. Zhang, B\. M\. Halpern, T\. Patel, and O\. Scharenborg\(2022\)Mitigating bias against non\-native accents\.InProceedings of Interspeech, Incheon, Korea,pp\. 3168–3172\.Cited by:[§I](https://arxiv.org/html/2607.19049#S1.p3.1)\.Similar Articles
Benchmarking Commercial ASR Systems on Code-Switching Speech: Arabic, Persian, and German
This paper presents a benchmark evaluating five commercial ASR systems on code-switching speech across Arabic-English, Persian-English, and German-English pairs, using a two-stage pipeline to select 300 samples per pair and assessing performance with WER and BERTScore. ElevenLabs Scribe v2 achieves the lowest overall WER (13.2%) and highest BERTScore (0.936), with public dataset available.
Transcribing Children's Speech: ASR Performance and Obtaining Reliable Orthographic Transcriptions
This paper evaluates nine ASR models (Whisper, Parakeet, Wav2Vec2) on Dutch child speech datasets JASMIN and DART, finding that fine-tuned Whisper-medium achieves the best performance (WER 5.54% on JASMIN, 70.37% on DART). It also proposes a selection method to automatically identify correctly pronounced utterances with high precision, reducing the need for manual verification.
What Counts as an Error? Dual-Reference Benchmarking for Atypical ASR
This paper proposes a dual-reference benchmarking approach for atypical ASR, using both verbatim and intended transcriptions to evaluate 11 ASR models on stuttered speech, highlighting the importance of selecting the appropriate reference depending on the use case.
Can Voice Agents Handle Bilingual Customers? Benchmarking Frontier ASR on Code-Switched Speech
ServiceNow AI releases a benchmark and dataset for evaluating automatic speech recognition (ASR) on code-switched speech across four language pairs (Spanish-English, French-English, Canadian French-English, German-English) in enterprise HR and IT scenarios, finding that current frontier ASR models still struggle with code-switching, leading to higher error rates.
Introducing the FFASR Leaderboard: Benchmarking ASR in the Real World
Introduces the FFASR Leaderboard, an open, community-driven benchmark for evaluating automatic speech recognition models under realistic far-field acoustic conditions, highlighting the significant performance gap between near-field and far-field scenarios.