@leeoxiang: Really need a precise timestamp alignment service
Summary
Feiteng Li announces the release of EdgeSpeak, a local-first, privacy-preserving accurate transcription service. It supports drag-and-drop audio/video or microphone recording transcription, semantic segmentation, word-level timestamps, and export to JSON/SRT/MD. It is compatible with OpenAI Audio API, CLI, SKILL, and MCP workflows.
View Cached Full Text
Cached at: 07/02/26, 04:26 PM
Really need a precise timestamp alignment service 👍
Feiteng (@FeitengLi): Today EdgeSpeak is officially released: local-first, no privacy leaks, accurate transcription
Drag and drop audio/video or microphone recording to transcribe:
Supports semantic segmentation, word-level timestamps, export to JSON / SRT / MD Compatible with OpenAI Audio API, CLI, SKILL, and MCP workflows.
EdgeSpeak is developed by Fable5 from start to finish,
Similar Articles
@FeitengLi: Next week, after adding speaker labeling and speech generation, it won't be this cheap early bird price anymore.
EdgeSpeak officially launched, a local-first, privacy-preserving accurate transcription tool, supporting semantic segmentation and timestamps, compatible with OpenAI Audio API, etc. It will later add speaker labeling and speech generation features.
@FeitengLi: Led by Fable 5 (just half a day), Codex relay development took a week. #EdgeSpeak is now live. Friends who shared, contact me to receive an invite code https://edgespeak.com/zh
EdgeSpeak desktop voice transcription tool is now live, featuring the local Lattice-2 voice model. It supports offline audio/video transcription, multiple languages and accents, and provides a local API for developers to integrate.
@mogician301: https://x.com/mogician301/status/2072250774332285073
Introduces an open-source tool called audio-to-text that uses AI and local tools (like faster-whisper and ffmpeg) to help users automatically generate, proofread, and burn subtitles, solving problems such as high costs, frequent errors, and cumbersome workflows in software like CapCut.
@uniswap12: Microsoft open-sourced a voice AI that can transcribe 60 minutes of long audio in one go, handling 4 people speaking simultaneously. VibeVoice, open-sourced by Microsoft, 24.8k stars, I only found out about it today. For converting recordings to text, I've been using Whisper, but it often times out on long meeting recordings and struggles with multi-speaker recognition...
Microsoft open-sourced the VibeVoice speech AI framework, which supports one-shot transcription of 60-minute long audio, multi-speaker diarization and timestamp labeling, and also provides multi-role TTS synthesis capabilities. It is based on Qwen2.5 and comes with a 0.5B lightweight real-time version. It has received 24.8k stars on GitHub.
@Huahuazo: The most annoying part of reposting overseas videos? Manually cutting subtitles, eye-straining translation alignment, and guessing audio sync—this workflow is so inefficient it makes you want to smash your keyboard. VideoLingo connects the entire process into an automated pipeline. It has 7k+ stars on GitHub and an MIT license, so it's reliable to use. …
VideoLingo is an open-source video translation, localization, and dubbing tool that uses WhisperX and AI to deliver Netflix-level subtitles and multilingual dubbing. It supports downloading via yt-dlp and multiple TTS options, helping video reposters automate the entire workflow.