Tag
Google Gemini now supports sign language transcription from video, a notable accessibility advancement for AI models.
PreenCut is an open-source tool (MIT License) based on Whisper transcription and LLM analysis, allowing users to search for segments in long videos using natural language, with support for batch processing, export/merging, and a REST API.
Discusses how voice agents lose paralinguistic signals like tone, hesitation, and speaker identity when transcribing to text, and questions whether and how these features are captured and used downstream.
Wired highlights Meetily, a free open-source app that transcribes and summarizes meetings locally without subscription fees or cloud uploads, working across major video conferencing platforms.
Wired's roundup of the best AI notetakers in 2026, reviewing devices like Plaud NotePin S, HiDock P1, and SpeakON, with notes on pricing, battery, and transcription quality.
Santiago Perez notes that 1 in 3 online class participants use AI note-takers, then highlights Wispr Flow's new Notetaker feature which can listen from your computer, offering accurate transcription and concise summaries without the AI joining the meeting.
LiveTranscriber is an open-source iOS app that runs Whisper, Qwen3-ASR, Nemotron, MOSS, and Qwen3 fully offline on iPhone, offering speech transcription, multi-speaker support, summaries, and real-time translation. The developer shares the engineering challenges and invites feedback from ASR and on-device AI communities.
Wispr Flow announces Notetaker, an AI meeting notes tool that promises accurate, speaker-attributed transcription and the ability to search across meetings, messages, and emails.
A guide on evaluating speech-to-text APIs by ranking transcript mistakes based on their actual impact on users, rather than raw accuracy metrics.
OpenAI and Microsoft released GPT-transcribe and GPT-live-transcribe in Microsoft Foundry, offering high-accuracy asynchronous transcription and low-latency streaming transcription for recorded and live audio, with features like background noise handling, accent robustness, and alphanumeric perception.
OpenAI releases two new transcription models: GPT Live Transcribe for low-latency and GPT Transcribe for batch workloads, with up to 41% lower error rates and improved semantic accuracy using context.
Vexa is an open-source, self-hosted meeting bot and transcription API that enables developers to capture live conversations and turn them into owned knowledge.
An individual open-sourced a system of 10 AI agents that automate YouTube Shorts creation from podcast clips, including transcription, editing, captioning, and scheduling, with a quality-checking agent.
Demonstrates that audio-input LLMs can generate speech by optimizing noise towards a desired transcription, similar to DeepDream for speech. The resulting audio sounds like a scary demon.
Laxis is a tool that makes meeting notes awesome by enabling 4x faster typing and live translation.
Microsoft released vibevoice, a 7B model that transcribes up to an hour of audio in one shot with built-in speaker diarization and timestamps, supporting 50+ languages and running locally without API costs.
The article highlights the growing ubiquity of AI note-taking apps that automatically record meetings and conversations, leading to pushback such as a VC changing his Zoom name to explicitly deny consent. It explores the social, legal, and practical implications of always-on recording.
MOSS-Transcribe-Diarize, an open-source ASR model with multi-speaker diarization and hotword biasing, trends on Hugging Face after release.
NVIDIA open-sourced a 600M parameter model that transcribes 40 languages in real-time with 80ms latency, supporting multiple languages from a single checkpoint with built-in punctuation and capitalization.
VoiceBox is an open-source desktop voice-to-text tool that captures speech, transcribes it via Whisper on Cloudflare AI, and formats output with an LLM, auto-pasting the result into the active application.