speech-recognition

Tag

Cards List
#speech-recognition

Brain-to-Language Decoding: Tasks, Signals, Methods, Evaluation, Practical Use and Beyond

arXiv cs.CL · 9h ago Cached

This survey paper reviews developments in brain-to-language decoding, translating neural activity into linguistic outputs for communication restoration and scientific study, covering tasks, methods, evaluation, and future directions.

0 favorites 0 likes
#speech-recognition

Ruby-ASR: Evidence-Preserving Supervision for Joint Orthographic and Lexical-Reading Recognition

arXiv cs.CL · 9h ago Cached

Ruby-ASR presents a new supervision method for Japanese automatic speech recognition that binds orthographic spans to their lexical readings, improving reading recognition while maintaining transcription accuracy.

0 favorites 0 likes
#speech-recognition

@SamuelZengML: Speech recognition is easy—until you ask it to listen forever. Today we’re open-sourcing Audio8 ASR Infinite: Ultra-low…

X AI KOLs Following · yesterday Cached

Open-sourcing Audio8 ASR Infinite, a speech recognition tool with ultra-low latency, unlimited audio support, 24/7 transcription, and built-in semantic turn detection, claimed to be new state-of-the-art for streaming ASR.

0 favorites 0 likes
#speech-recognition

A Practical Recipe for Semi-Supervised Federated ASR: Online Pseudo-Labels with Server Update Stabilization

arXiv cs.LG · yesterday Cached

This paper proposes a practical recipe for semi-supervised federated ASR using online pseudo-labels with server update stabilization, demonstrating significant improvements over prior methods in both in-domain and cross-domain settings.

0 favorites 0 likes
#speech-recognition

@xiangxiang103: That's insanely impressive! Hugging Face unleashed a dark horse this week—one that's dominating two tracks at once. Spe…

X AI KOLs Timeline · yesterday Cached

NetEase Youdao's open-source AI models R2T2 and T3PO have topped Hugging Face leaderboards for speech recognition and translation, outperforming major competitors with impressive real-time performance and stability.

0 favorites 0 likes
#speech-recognition

Speechka

Product Hunt · yesterday Cached

Speechka is a real-time voice translation tool that mimics the user's voice across 44 languages, available on macOS, Windows, and browsers.

0 favorites 0 likes
#speech-recognition

Aggregate WER is a useless metric for voice agents in production

Reddit r/AI_Agents · 3d ago

The article highlights how aggregate Word Error Rate (WER) metric fails to capture critical errors in voice agents for production, such as misrecognized bank codes and multilingual speech, necessitating per-field tracking for real-world applications like banking and collections.

0 favorites 0 likes
#speech-recognition

Scaling Forced Alignment to End-User Devices

arXiv cs.CL · 3d ago Cached

The paper proposes optimizations to the Viterbi algorithm using the Hirschberg algorithm and constrained random walk, reducing memory usage from 140 GB to 5 MB and improving speed, enabling forced alignment to run on end-user devices for better scalability in speech processing.

0 favorites 0 likes
#speech-recognition

Beyond WER: Entity and Disfluency Recall in Accented Conversational ASR

arXiv cs.CL · 3d ago Cached

This paper proposes a three-stage pipeline for accented conversational ASR that improves entity and disfluency recall, achieving 80–85% entity recall and outperforming baseline systems with fewer parameters.

0 favorites 0 likes
#speech-recognition

@svpino: This new model does something really cool: It turns speech into text as you speak. This is different from every other a…

X AI KOLs Timeline · 3d ago Cached

R2T2 is a low-latency and high-accuracy real-time speech recognition model that processes audio in small chunks and commits text without revision, suitable for applications like live captioning and translation.

0 favorites 0 likes
#speech-recognition

Stability vs. speed: Rethinking ASR for the age of voice agents

Reddit r/artificial · 6d ago

The article reflects on the need for stability over speed in ASR for voice agents, discussing the 'Confucius r2t2' model that integrates wait/commit mechanisms to enhance reliability.

0 favorites 0 likes
#speech-recognition

TypeDash

Product Hunt · 6d ago Cached

TypeDash is a desktop application that lets users control their computer and dictate text via voice commands, featuring app launching, web search, and optional ChatGPT-powered cleanup for speech recognition.

0 favorites 0 likes
#speech-recognition

T-SANDHI: Tone Sandhi-aware Adaptive Network with Decoupled Hybrid Injection for Low-resource Taiwanese Hokkien Speech Recognition

arXiv cs.CL · 2026-09-17 Cached

The paper proposes T-SANDHI, a Tone Sandhi-aware Adaptive Network for low-resource Taiwanese Hokkien speech recognition, which explicitly decouples tonal variations to improve accuracy on top of a Whisper backbone.

0 favorites 0 likes
#speech-recognition

@NetEaseYouDaoAI: Most streaming ASR gives you text fast. R2T2 gives you text you can act on. We’re open-sourcing Confucius4-R2T2, a 1.7B…

X AI KOLs Following · 2026-09-17 Cached

NetEase Youdao AI is open-sourcing Confucius4-R2T2, a 1.7B frontier real-time streaming ASR model designed for voice agents, featuring low latency, high accuracy, and configurable decoding chunks.

0 favorites 0 likes
#speech-recognition

DiaWhisper-DPO: Role-Attributed Transcription of Clinical Interviews via Failure-Mined Preference Optimization

arXiv cs.CL · 2026-09-16 Cached

The paper proposes DiaWhisper-DPO, an end-to-end model for transcription and role attribution in clinical interviews using failure-mined preference optimization, achieving high accuracy and reducing errors compared to cascaded baselines.

0 favorites 0 likes
#speech-recognition

@FinanceYF5: Voice AI is starting to make its way into phones. Open-source Audio8, which packages speech recognition and speech synt…

X AI KOLs Timeline · 2026-09-16 Cached

The post introduces the open-source Audio8 models, which enable on-device speech recognition and synthesis on phones and PCs, including an offline transcription version for iPhone.

0 favorites 0 likes
#speech-recognition

AI for everyone in every language

Google AI Blog · 2026-09-15 Cached

Google advances AI for language support with new Gemini models for real-time translation and transcription, aiming to cover 1,000 languages through initiatives like the Universal Speech Model.

0 favorites 0 likes
#speech-recognition

Show HN: Jexxa: High Speed on Device Dictation

Hacker News Top · 2026-09-15 Cached

Jexxa is a high-speed dictation tool for macOS that runs entirely on-device, ensuring privacy and fast performance with no data uploads.

0 favorites 0 likes
#speech-recognition

Token Merging for Multilingual Speech Recognition: A Systematic Study Across Model Scale and Fine-Tuning

arXiv cs.CL · 2026-09-15 Cached

This paper systematically evaluates token merging for multilingual speech recognition on the Whisper model family, demonstrating improved computational efficiency with minimal accuracy loss across low-resource languages and fine-tuned models.

0 favorites 0 likes
#speech-recognition

Voice AI Architecture Discussion

Reddit r/AI_Agents · 2026-09-14

The article discusses the tradeoff between latency and control in voice AI architectures, comparing traditional cascaded systems with end-to-end models, and seeks community input on current practices.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback