whisper

Tag

Cards List
#whisper

Oído: speech recognition that beats Whisper-tiny, running on a $5 microcontroller (open source)

Reddit r/LocalLLaMA ↗ · 18h ago

Oído is an open-source speech recognition system from Lokutor that runs NVIDIA's Conformer-CTC Small (13M params, int8) on a $5 ESP32-S3 microcontroller, outperforming Whisper tiny.en on LibriSpeech and noisy benchmarks with no GPU or NPU.

0 favorites 0 likes
#whisper

BanglaTurn: A Benchmark and Whisper-Based Model for End-of-Turn Detection in Bangla Speech

arXiv cs.CL ↗ · 6d ago Cached

The paper introduces BanglaTurn, a benchmark corpus and Whisper-based model for end-of-turn detection in Bangla speech, achieving 84.33% accuracy compared to a 69.28% baseline.

0 favorites 0 likes
#whisper

Six Layers Less: Encoder Pruning for Whisper with Label-Free Recovery

Hugging Face Daily Papers ↗ · 2026-09-23 Cached

The paper introduces a method to prune encoder layers in the Whisper ASR model, reducing encoder size by 18.5% and recovering performance through unlabeled data distillation, with code and a pre-trained model released for adoption.

0 favorites 0 likes
#whisper

Token Merging for Multilingual Speech Recognition: A Systematic Study Across Model Scale and Fine-Tuning

arXiv cs.CL ↗ · 2026-09-15 Cached

This paper systematically evaluates token merging for multilingual speech recognition on the Whisper model family, demonstrating improved computational efficiency with minimal accuracy loss across low-resource languages and fine-tuned models.

0 favorites 0 likes
#whisper

Automatic Lyric Transcription for Greek Songs: Scaling and Task Composition Effects in Whisper Adaptation

arXiv cs.CL ↗ · 2026-09-11 Cached

This paper introduces the first benchmark for automatic lyric transcription in Greek songs, demonstrating that adapting Whisper through multitask learning and two-stage training achieves a 27.2% Word Error Rate, significantly improving over zero-shot baselines.

0 favorites 0 likes
#whisper

BuzzASR: A Swarm of 100+ Monolingual Speech Recognition Models

arXiv cs.CL ↗ · 2026-09-10 Cached

BuzzASR is a collection of language-specialized Whisper models for automatic speech recognition in 102 languages, outperforming Whisper-large-v3 on 77 languages with significant improvements in error rates and compression efficiency.

0 favorites 0 likes
#whisper

SpeakPay: Domain-Adaptive LoRA Fine-Tuning of Whisper for Low-Resource Nepali Financial Speech Recognition

arXiv cs.CL ↗ · 2026-09-03 Cached

This paper introduces SpeakPay and a Nepali financial speech dataset, showing that LoRA fine-tuning of Whisper reduces Word Error Rate by 67.2% and improves transaction success rates for low-resource language accessibility.

0 favorites 0 likes
#whisper

A Speech Corpus for Mizo Automatic Speech Recognition: Whisper and SraVaani 1.0 Fine-Tuning with Morphology-Aware Evaluation

arXiv cs.CL ↗ · 2026-08-21 Cached

This study develops an Automatic Speech Recognition system for Mizo, a low-resource language, by fine-tuning Whisper and SraVaani 1.0 models, achieving a morphology-aware WER of 7.22% with Whisper-large-v3.

0 favorites 0 likes
#whisper

Wow: Generate subtitles and cut out filler words in the browser – 1.7K stars on GitHub. The most tedious parts of editing talking-head videos are typing subtitles and cutting out pauses, breaths, and repetitive filler. These tasks can account for more than half of the editing time. FlyCut…

X AI KOLs Timeline ↗ · 2026-08-16 Cached

FlyCut Caption is a browser-based AI video editing tool that automatically generates subtitles and cuts out filler words in videos, using Whisper and FunASR models for local processing without uploading footage.

0 favorites 0 likes
#whisper

@Ryrenz: Damn, you can find the clip you want from hours of footage with just one sentence — MIT License, Whisper transcription plus LLM analysis. The most painful part of editing long videos is finding footage. With hours of recordings, to find every segment about a topic, you can only drag the timeline and listen bit by bit, burning an entire afternoon. PreenCut ...

X AI KOLs Timeline ↗ · 2026-08-11 Cached

PreenCut is an open-source tool (MIT License) based on Whisper transcription and LLM analysis, allowing users to search for segments in long videos using natural language, with support for batch processing, export/merging, and a REST API.

0 favorites 0 likes
#whisper

Show HN: Vocal Slice – Cut audio by selecting text, fully on-device

Hacker News Top ↗ · 2026-08-10 Cached

Vocal Slice is a desktop application for audio editing that uses on-device Whisper transcription to allow users to select and export clips by highlighting text in the transcript, designed for podcasters and voice professionals.

0 favorites 0 likes
#whisper

@hank_aibtc: Whoa, this thing really blew my mind. Dug up a local tool called KrillinAI, free, video translation, precise subtitles, ultra-natural voiceover, voice cloning, the whole pipeline in one go. Chinese-English translation is especially stable, Whisper recognition accuracy is ridiculously high, LLM translates segment by segment without losing context, CosyV…

X AI KOLs Timeline ↗ · 2026-08-10 Cached

Introducing KrillinAI, a free locally-run video translation tool that supports precise subtitles, natural voiceover, and voice cloning. It integrates Whisper, LLM, and CosyVoice, and supports Win/Mac and yt-dlp.

0 favorites 0 likes
#whisper

How to Recognize New Words: A Comparison Between Context Biasing Methods and Speech LLMs

arXiv cs.CL ↗ · 2026-08-07 Cached

This paper compares context biasing methods and speech LLMs for recognizing rare and new words in automatic speech recognition, reporting trade-offs in word error rate across read and non-read speech.

0 favorites 0 likes
#whisper

A Study of ASR Adaptation and Representation Dimensionality Reduction in Persian Speech Emotion Recognition Using Whisper

arXiv cs.CL ↗ · 2026-08-07 Cached

This paper studies using Whisper for Persian speech emotion recognition, showing PCA-based dimensionality reduction improves performance and efficiency, while ASR fine-tuning offers only modest gains.

0 favorites 0 likes
#whisper

Room reverberation and low SNR hurt STT accuracy far more than model size

Reddit r/ArtificialInteligence ↗ · 2026-07-30

Room reverberation and low-frequency noise from the environment hurt speech-to-text accuracy far more than the choice of model size; front-end audio preprocessing like adaptive spectral subtraction can recover masked phonemes and reduce word error rate more effectively than upgrading the model backend.

0 favorites 0 likes
#whisper

Whisper Live - A nearly-live implementation of Open AI's Whisper, free & open-source

Reddit r/artificial ↗ · 2026-07-25 Cached

WhisperLive is an open-source real-time transcription tool using OpenAI's Whisper, supporting multiple backends like faster-whisper and TensorRT for live speech-to-text.

0 favorites 0 likes
#whisper

@xiangxiang103: A bro from Bilibili made a cyber girlfriend, so immersive! A fully local AI girlfriend, no internet, no API key. VAD: Silero VAD v5 STT: Whisper LLM: local llama.cpp TTS: Qwen3-TTS All four models packed into 1…

X AI KOLs Timeline ↗ · 2026-07-25 Cached

Introduces a fully local AI girlfriend project made by a Bilibili developer, integrating four models: Silero VAD, Whisper, llama.cpp, and Qwen3-TTS, all packed into 15G VRAM with hot-swapping capability.

0 favorites 0 likes
#whisper

Open Source, Free Tier Capable Whispr Using Cloudflare AI

Hacker News Top ↗ · 2026-07-16 Cached

VoiceBox is an open-source desktop voice-to-text tool that captures speech, transcribes it via Whisper on Cloudflare AI, and formats output with an LLM, auto-pasting the result into the active application.

0 favorites 0 likes
#whisper

Audio perception layer for LLM agents, with a memory that grows through use

Reddit r/LocalLLaMA ↗ · 2026-07-15

An experimental open-source framework that enables LLM agents to perceive and recognize non-speech audio events using local models (CLAP, Whisper, Silero VAD) and a growing concept memory. The system uses event-gated recognition, fingerprinting, and symbol-based reasoning, with no formal benchmarks yet.

0 favorites 0 likes
#whisper

Apple's new SpeechAnalyzer API, benchmarked against Whisper and its predecessor

Hacker News Top ↗ · 2026-07-13 Cached

Apple's new SpeechAnalyzer API significantly outperforms both its predecessor SFSpeechRecognizer and OpenAI's Whisper models in accuracy and speed for English on-device transcription, as benchmarked on an M2 Pro machine. The new API achieves a 2.12% word error rate on clean speech, compared to 3.74% for Whisper Small, and runs three times faster.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback