whisper

Tag

Cards List
#whisper

BanglaTurn: A Benchmark and Whisper-Based Model for End-of-Turn Detection in Bangla Speech

arXiv cs.CL ↗ · 10h ago Cached

The paper introduces BanglaTurn, a benchmark corpus and Whisper-based model for end-of-turn detection in Bangla speech, achieving 84.33% accuracy compared to a 69.28% baseline.

0 favorites 0 likes
#whisper

Six Layers Less: Encoder Pruning for Whisper with Label-Free Recovery

Hugging Face Daily Papers ↗ · 2d ago Cached

The paper introduces a method to prune encoder layers in the Whisper ASR model, reducing encoder size by 18.5% and recovering performance through unlabeled data distillation, with code and a pre-trained model released for adoption.

0 favorites 0 likes
#whisper

Token Merging for Multilingual Speech Recognition: A Systematic Study Across Model Scale and Fine-Tuning

arXiv cs.CL ↗ · 2026-09-15 Cached

This paper systematically evaluates token merging for multilingual speech recognition on the Whisper model family, demonstrating improved computational efficiency with minimal accuracy loss across low-resource languages and fine-tuned models.

0 favorites 0 likes
#whisper

Automatic Lyric Transcription for Greek Songs: Scaling and Task Composition Effects in Whisper Adaptation

arXiv cs.CL ↗ · 2026-09-11 Cached

This paper introduces the first benchmark for automatic lyric transcription in Greek songs, demonstrating that adapting Whisper through multitask learning and two-stage training achieves a 27.2% Word Error Rate, significantly improving over zero-shot baselines.

0 favorites 0 likes
#whisper

BuzzASR: A Swarm of 100+ Monolingual Speech Recognition Models

arXiv cs.CL ↗ · 2026-09-10 Cached

BuzzASR is a collection of language-specialized Whisper models for automatic speech recognition in 102 languages, outperforming Whisper-large-v3 on 77 languages with significant improvements in error rates and compression efficiency.

0 favorites 0 likes
#whisper

SpeakPay: Domain-Adaptive LoRA Fine-Tuning of Whisper for Low-Resource Nepali Financial Speech Recognition

arXiv cs.CL ↗ · 2026-09-03 Cached

This paper introduces SpeakPay and a Nepali financial speech dataset, showing that LoRA fine-tuning of Whisper reduces Word Error Rate by 67.2% and improves transaction success rates for low-resource language accessibility.

0 favorites 0 likes
#whisper

A Speech Corpus for Mizo Automatic Speech Recognition: Whisper and SraVaani 1.0 Fine-Tuning with Morphology-Aware Evaluation

arXiv cs.CL ↗ · 2026-08-21 Cached

This study develops an Automatic Speech Recognition system for Mizo, a low-resource language, by fine-tuning Whisper and SraVaani 1.0 models, achieving a morphology-aware WER of 7.22% with Whisper-large-v3.

0 favorites 0 likes
#whisper

Wow: Generate subtitles and cut out filler words in the browser – 1.7K stars on GitHub. The most tedious parts of editing talking-head videos are typing subtitles and cutting out pauses, breaths, and repetitive filler. These tasks can account for more than half of the editing time. FlyCut…

X AI KOLs Timeline ↗ · 2026-08-16 Cached

FlyCut Caption is a browser-based AI video editing tool that automatically generates subtitles and cuts out filler words in videos, using Whisper and FunASR models for local processing without uploading footage.

0 favorites 0 likes
#whisper

@Ryrenz: Damn, you can find the clip you want from hours of footage with just one sentence — MIT License, Whisper transcription plus LLM analysis. The most painful part of editing long videos is finding footage. With hours of recordings, to find every segment about a topic, you can only drag the timeline and listen bit by bit, burning an entire afternoon. PreenCut ...

X AI KOLs Timeline ↗ · 2026-08-11 Cached

PreenCut is an open-source tool (MIT License) based on Whisper transcription and LLM analysis, allowing users to search for segments in long videos using natural language, with support for batch processing, export/merging, and a REST API.

0 favorites 0 likes
#whisper

Show HN: Vocal Slice – Cut audio by selecting text, fully on-device

Hacker News Top ↗ · 2026-08-10 Cached

Vocal Slice is a desktop application for audio editing that uses on-device Whisper transcription to allow users to select and export clips by highlighting text in the transcript, designed for podcasters and voice professionals.

0 favorites 0 likes
#whisper

@hank_aibtc: Whoa, this thing really blew my mind. Dug up a local tool called KrillinAI, free, video translation, precise subtitles, ultra-natural voiceover, voice cloning, the whole pipeline in one go. Chinese-English translation is especially stable, Whisper recognition accuracy is ridiculously high, LLM translates segment by segment without losing context, CosyV…

X AI KOLs Timeline ↗ · 2026-08-10 Cached

Introducing KrillinAI, a free locally-run video translation tool that supports precise subtitles, natural voiceover, and voice cloning. It integrates Whisper, LLM, and CosyVoice, and supports Win/Mac and yt-dlp.

0 favorites 0 likes
#whisper

How to Recognize New Words: A Comparison Between Context Biasing Methods and Speech LLMs

arXiv cs.CL ↗ · 2026-08-07 Cached

This paper compares context biasing methods and speech LLMs for recognizing rare and new words in automatic speech recognition, reporting trade-offs in word error rate across read and non-read speech.

0 favorites 0 likes
#whisper

A Study of ASR Adaptation and Representation Dimensionality Reduction in Persian Speech Emotion Recognition Using Whisper

arXiv cs.CL ↗ · 2026-08-07 Cached

This paper studies using Whisper for Persian speech emotion recognition, showing PCA-based dimensionality reduction improves performance and efficiency, while ASR fine-tuning offers only modest gains.

0 favorites 0 likes
#whisper

Room reverberation and low SNR hurt STT accuracy far more than model size

Reddit r/ArtificialInteligence ↗ · 2026-07-30

Room reverberation and low-frequency noise from the environment hurt speech-to-text accuracy far more than the choice of model size; front-end audio preprocessing like adaptive spectral subtraction can recover masked phonemes and reduce word error rate more effectively than upgrading the model backend.

0 favorites 0 likes
#whisper

Whisper Live - A nearly-live implementation of Open AI's Whisper, free & open-source

Reddit r/artificial ↗ · 2026-07-25 Cached

WhisperLive is an open-source real-time transcription tool using OpenAI's Whisper, supporting multiple backends like faster-whisper and TensorRT for live speech-to-text.

0 favorites 0 likes
#whisper

@xiangxiang103: A bro from Bilibili made a cyber girlfriend, so immersive! A fully local AI girlfriend, no internet, no API key. VAD: Silero VAD v5 STT: Whisper LLM: local llama.cpp TTS: Qwen3-TTS All four models packed into 1…

X AI KOLs Timeline ↗ · 2026-07-25 Cached

Introduces a fully local AI girlfriend project made by a Bilibili developer, integrating four models: Silero VAD, Whisper, llama.cpp, and Qwen3-TTS, all packed into 15G VRAM with hot-swapping capability.

0 favorites 0 likes
#whisper

Open Source, Free Tier Capable Whispr Using Cloudflare AI

Hacker News Top ↗ · 2026-07-16 Cached

VoiceBox is an open-source desktop voice-to-text tool that captures speech, transcribes it via Whisper on Cloudflare AI, and formats output with an LLM, auto-pasting the result into the active application.

0 favorites 0 likes
#whisper

Audio perception layer for LLM agents, with a memory that grows through use

Reddit r/LocalLLaMA ↗ · 2026-07-15

An experimental open-source framework that enables LLM agents to perceive and recognize non-speech audio events using local models (CLAP, Whisper, Silero VAD) and a growing concept memory. The system uses event-gated recognition, fingerprinting, and symbol-based reasoning, with no formal benchmarks yet.

0 favorites 0 likes
#whisper

Apple's new SpeechAnalyzer API, benchmarked against Whisper and its predecessor

Hacker News Top ↗ · 2026-07-13 Cached

Apple's new SpeechAnalyzer API significantly outperforms both its predecessor SFSpeechRecognizer and OpenAI's Whisper models in accuracy and speed for English on-device transcription, as benchmarked on an M2 Pro machine. The new API achieves a 2.12% word error rate on clean speech, compared to 3.74% for Whisper Small, and runs three times faster.

0 favorites 0 likes
#whisper

Is there any AI with extremely high sensitivity to impaired speech?

Reddit r/ArtificialInteligence ↗ · 2026-07-12

A user seeks advice on AI speech recognition models that can accurately understand severely impaired speech, such as that of their minimally verbal brother with Down syndrome and autism, noting that current systems like Whisper fail to recognize it.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback