@svpino: Cutting down noise before sending the audio to a speech-to-text model makes a huge improvement. Voice isolation is the …

X AI KOLs Timeline News

Summary

Krisp released an open benchmark and dataset showing that voice isolation reduces word error rates in speech-to-text models by 73%, with significant improvements across workplace and call-center recordings.

Cutting down noise before sending the audio to a speech-to-text model makes a huge improvement. Voice isolation is the next big thing you want to pay attention to. Krisp just released an open benchmark and a dataset to measure the impact of voice isolation on STT accuracy. The results are eye-watering: 73% reduction in WER across 265 real recordings and 11 speech-to-text configurations. This compares an original recording with a voice-isolated version. Here are the results with voice isolation: • Overall word error rate: Down from 23.24% to 6.24%. • Workplace recordings: Down from 31.84% to 6.30%. • Call-center recordings: Down from 23.83% to 6.86%. If you have a speech-to-text setup, you can reproduce these experiments using your audio and see whether voice isolation helps. Here is the link: https://partner.krisp.ai/svpino-x-bm Thanks to the Krisp team for partnering with me on this post.
Original Article
View Cached Full Text

Cached at: 09/24/26, 10:30 PM

Cutting down noise before sending the audio to a speech-to-text model makes a huge improvement.

Voice isolation is the next big thing you want to pay attention to.

Krisp just released an open benchmark and a dataset to measure the impact of voice isolation on STT accuracy.

The results are eye-watering:

73% reduction in WER across 265 real recordings and 11 speech-to-text configurations.

This compares an original recording with a voice-isolated version.

Here are the results with voice isolation:

• Overall word error rate: Down from 23.24% to 6.24%. • Workplace recordings: Down from 31.84% to 6.30%. • Call-center recordings: Down from 23.83% to 6.86%.

If you have a speech-to-text setup, you can reproduce these experiments using your audio and see whether voice isolation helps.

Here is the link: https://partner.krisp.ai/svpino-x-bm

Thanks to the Krisp team for partnering with me on this post.


Voice Isolation STT Benchmark | Krisp

Source: https://krisp.ai/benchmarks/voice-isolation-benchmark/?utm_source=svpino&utm_medium=twitter&utm_campaign=vi_benchmark_2026&utm_content=benchmark This page requires JavaScript to display.

Unpacking...

Similar Articles

Room reverberation and low SNR hurt STT accuracy far more than model size

Reddit r/ArtificialInteligence

Room reverberation and low-frequency noise from the environment hurt speech-to-text accuracy far more than the choice of model size; front-end audio preprocessing like adaptive spectral subtraction can recover masked phonemes and reduce word error rate more effectively than upgrading the model backend.

Voice agents in noisy environments

Reddit r/AI_Agents

A speech company trained a model that cancels noise and identifies the primary speaker, achieving 50% lower word error rate on leading ASR models in noisy environments.