Tag
Orukeet is a 25-language speech recognizer built from NVIDIA Parakeet, using fitted Gabor kernels to replace encoder filters, outperforming the base model on most benchmarks.
The paper introduces α-split, a two-pool allocation method for differential privacy in federated learning, addressing cross-component budget collapse in speech-LLMs and improving utility and security against attacks.
This paper details the Eloquence team's approaches for Task 2 of the Interspeech 2026 MLC-SLM challenge, which involves multilingual multiple-choice question answering using speech LLMs with fine-tuning, in-context learning, and retrieval systems.
This paper evaluates the reusability of publicly available speech resources for low-resource languages, focusing on Central Kurdish as a case study.
This paper details the development of Sophea, a production bilingual Greek-English automatic speech recognition system, using iterative training, data filtering, and model ensembling to meet quality gates and achieve competitive benchmark results.
Confucius4-R2T2 is a low-latency, high-accuracy real-time speech recognition model developed by NetEase Youdao, featuring configurable chunking and stable output for applications like live captioning.
BuzzASR is a collection of language-specialized Whisper models for automatic speech recognition in 102 languages, outperforming Whisper-large-v3 on 77 languages with significant improvements in error rates and compression efficiency.
Desert Ant Labs offers small, specialized AI models for speech, text, and vision that run offline on devices, with an SDK for easy integration and no per-use costs.
This paper introduces Hybrid Search, a method to enhance automatic speech recognition in large audio language models by leveraging hidden-state interactions between the ASR-LLM and base LLM for targeted token correction, improving performance beyond global LLM-correction strategies.
The tweet discusses how AI hardware could reduce screen dependency by leveraging contextual understanding, and highlights HojoAI's open multilingual ASR model as a significant release.
This paper introduces SpeakPay and a Nepali financial speech dataset, showing that LoRA fine-tuning of Whisper reduces Word Error Rate by 67.2% and improves transaction success rates for low-resource language accessibility.
Microsoft has released VibeVoice-ASR-Streaming, a unified streaming ASR model that transcribes who said what with support for customized hotwords and 10 languages.
AirCaps Audio Research Lab launches with streaming speech and audio models designed for complex real-world environments, claiming superior performance on edge hardware compared to leading cloud models.
This paper benchmarks topic matching methods on real-world ASR transcripts from contact centers, finding that lightweight LLM matchers with natural language descriptions outperform regex and embedding approaches.
VibeVoice-ASR-Streaming is an LLM-based end-to-end model for streaming speaker-attributed speech recognition, achieving state-of-the-art performance with released 1.5B and 7B model weights.
Vellium v1.1.0 introduces live voice chat with local STT/TTS using Whisper and TeraTTSv2, and simplifies llama.cpp setup for easier integration with local AI models.
The paper proposes SAMA-ASR, a multimodal adapter that improves automatic speech recognition for low-resource languages by using semantic anchors from translations and acoustic anchors from speech, with experiments on Taiwanese Hokkien and Hakka showing effectiveness over baselines.
VoiceCodeBench is a new benchmark for evaluating exact structured-token recovery in automatic speech recognition, showing that traditional WER metrics are insufficient for production voice workflows.
This preregistered ablation study tests prompt-level context in a production speech transcription tool and finds no detectable change in side-level word error rate, contradicting earlier reports of gains from prompt conditioning.
This paper introduces PromptKWS, a novel prompt-guided open-vocabulary keyword spotting framework that uses prompt embeddings and cross-attention to improve accuracy, achieving over 10% improvement in wakeup rate and over 15% in accuracy compared to baseline systems.