speech-recognition

Tag

Cards List
#speech-recognition

Voice AI Architecture Discussion

Reddit r/AI_Agents · 2026-09-14

The article discusses the tradeoff between latency and control in voice AI architectures, comparing traditional cascaded systems with end-to-end models, and seeks community input on current practices.

0 favorites 0 likes
#speech-recognition

Orukeet, new ASR model based on Parakeet

Reddit r/LocalLLaMA · 2026-09-11 Cached

Orukeet is a 25-language speech recognizer built from NVIDIA Parakeet, using fitted Gabor kernels to replace encoder filters, outperforming the base model on most benchmarks.

0 favorites 0 likes
#speech-recognition

Component-Aware Differential Privacy for Federated Multilingual Speech-LLMs

arXiv cs.CL · 2026-09-11 Cached

The paper introduces α-split, a two-pool allocation method for differential privacy in federated learning, addressing cross-component budget collapse in speech-LLMs and improving utility and security against attacks.

0 favorites 0 likes
#speech-recognition

The Eloquence submission for Task 2 of the Interspeech 2026 MLC-SLM challenge

arXiv cs.CL · 2026-09-11 Cached

This paper details the Eloquence team's approaches for Task 2 of the Interspeech 2026 MLC-SLM challenge, which involves multilingual multiple-choice question answering using speech LLMs with fine-tuning, in-context learning, and retrieval systems.

0 favorites 0 likes
#speech-recognition

Assessing the Reusability of Public Speech Resources for Low-Resource Languages: A Central Kurdish Case Study

arXiv cs.CL · 2026-09-11 Cached

This paper evaluates the reusability of publicly available speech resources for low-resource languages, focusing on Central Kurdish as a case study.

0 favorites 0 likes
#speech-recognition

Building a Production Greek-English Speech Recognizer

Hugging Face Daily Papers · 2026-09-11 Cached

This paper details the development of Sophea, a production bilingual Greek-English automatic speech recognition system, using iterative training, data filtering, and model ensembling to meet quality gates and achieve competitive benchmark results.

0 favorites 0 likes
#speech-recognition

netease-youdao/Confucius4-R2T2

Hugging Face Models Trending · 2026-09-10 Cached

Confucius4-R2T2 is a low-latency, high-accuracy real-time speech recognition model developed by NetEase Youdao, featuring configurable chunking and stable output for applications like live captioning.

0 favorites 0 likes
#speech-recognition

BuzzASR: A Swarm of 100+ Monolingual Speech Recognition Models

arXiv cs.CL · 2026-09-10 Cached

BuzzASR is a collection of language-specialized Whisper models for automatic speech recognition in 102 languages, outperforming Whisper-large-v3 on 77 languages with significant improvements in error rates and compression efficiency.

0 favorites 0 likes
#speech-recognition

Desert Ant Labs

Product Hunt · 2026-09-09 Cached

Desert Ant Labs offers small, specialized AI models for speech, text, and vision that run offline on devices, with an SDK for easy integration and no per-use costs.

0 favorites 0 likes
#speech-recognition

Listen to the Latents: Self-Correcting Speech Recognition in Large Audio Language Models Through Hidden-State Interactions

arXiv cs.CL · 2026-09-04 Cached

This paper introduces Hybrid Search, a method to enhance automatic speech recognition in large audio language models by leveraging hidden-state interactions between the ASR-LLM and base LLM for targeted token correction, improving performance beyond global LLM-correction strategies.

0 favorites 0 likes
#speech-recognition

@rohanpaul_ai: My criteria for buying tech have changed a lot over the years. The biggest limitation of the smartphone might be that i…

X AI KOLs Following · 2026-09-03 Cached

The tweet discusses how AI hardware could reduce screen dependency by leveraging contextual understanding, and highlights HojoAI's open multilingual ASR model as a significant release.

0 favorites 0 likes
#speech-recognition

SpeakPay: Domain-Adaptive LoRA Fine-Tuning of Whisper for Low-Resource Nepali Financial Speech Recognition

arXiv cs.CL · 2026-09-03 Cached

This paper introduces SpeakPay and a Nepali financial speech dataset, showing that LoRA fine-tuning of Whisper reduces Word Error Rate by 67.2% and improves transaction success rates for low-resource language accessibility.

0 favorites 0 likes
#speech-recognition

Microsoft VibeVoice-ASR-Streaming Released

Reddit r/LocalLLaMA · 2026-09-03 Cached

Microsoft has released VibeVoice-ASR-Streaming, a unified streaming ASR model that transcribes who said what with support for customized hotwords and 10 languages.

0 favorites 0 likes
#speech-recognition

@tryaircaps: Today, we’re publicly launching the AirCaps Audio Research Lab. Real-world audio is 10x harder to understand than the d…

X AI KOLs Following · 2026-09-02 Cached

AirCaps Audio Research Lab launches with streaming speech and audio models designed for complex real-world environments, claiming superior performance on edge hardware compared to leading cloud models.

0 favorites 0 likes
#speech-recognition

Topic Matching in the Wild: Benchmark and Lessons from Real-World ASR Transcripts

arXiv cs.CL · 2026-09-02 Cached

This paper benchmarks topic matching methods on real-world ASR transcripts from contact centers, finding that lightweight LLM matchers with natural language descriptions outperform regex and embedding approaches.

0 favorites 0 likes
#speech-recognition

VibeVoice-ASR-Streaming Technical Report

Hugging Face Daily Papers · 2026-09-02 Cached

VibeVoice-ASR-Streaming is an LLM-based end-to-end model for streaming speaker-attributed speech recognition, achieving state-of-the-art performance with released 1.5B and 7B model weights.

0 favorites 0 likes
#speech-recognition

Vellium v1.1.0 — Live voice, local STT/TTS and easier llama.cpp setup

Reddit r/LocalLLaMA · 2026-09-01

Vellium v1.1.0 introduces live voice chat with local STT/TTS using Whisper and TeraTTSv2, and simplifies llama.cpp setup for easier integration with local AI models.

0 favorites 0 likes
#speech-recognition

Anchoring Speech with Semantics: A Multimodal Adapter Mechanism for Automatic Speech Recognition in Low-Resource Languages

arXiv cs.CL · 2026-09-01 Cached

The paper proposes SAMA-ASR, a multimodal adapter that improves automatic speech recognition for low-resource languages by using semantic anchors from translations and acoustic anchors from speech, with experiments on Taiwanese Hokkien and Hakka showing effectiveness over baselines.

0 favorites 0 likes
#speech-recognition

VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition

arXiv cs.CL · 2026-09-01 Cached

VoiceCodeBench is a new benchmark for evaluating exact structured-token recovery in automatic speech recognition, showing that traditional WER metrics are insufficient for production voice workflows.

0 favorites 0 likes
#speech-recognition

No Detectable Change in Side-Level WER from Prompt-Level Context: A Preregistered Ablation on a Production Oral-History Corpus

arXiv cs.CL · 2026-09-01 Cached

This preregistered ablation study tests prompt-level context in a production speech transcription tool and finds no detectable change in side-level word error rate, contradicting earlier reports of gains from prompt conditioning.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback