@Honcia13: Highly recommend an open-source speech-to-subtitle tool! Incredible speed and top-notch quality! Supports multiple languages including Chinese, Japanese, Korean, English, etc., with specially optimized formatting rules for natural and professional subtitles. It's a desktop tool based on PySide6 + ElevenLabs API that can convert audio/video files or JSON…

X AI KOLs Timeline Tools

Summary

Recommend Scribe2SRT, an open-source speech-to-subtitle tool based on PySide6 and ElevenLabs API, supporting multiple languages with optimized formatting for fast generation of high-quality SRT subtitles.

Highly recommend an open-source speech-to-subtitle tool! Incredible speed and top-notch quality! Supports multiple languages including Chinese, Japanese, Korean, English, etc., with specially optimized formatting rules for natural and professional subtitles. It's a desktop tool based on PySide6 + ElevenLabs API that can intelligently convert audio/video files or JSON transcripts into high-quality SRT subtitles, especially suitable for the formatting conventions of Chinese, Japanese, Korean, and English. For those who make videos, edit, create courseware, or produce subtitles, this is really awesome! https://github.com/cylind/scribe2srt…
Original Article
View Cached Full Text

Cached at: 05/31/26, 07:03 AM

Strongly recommend an open-source speech-to-subtitle tool! Extremely fast with top-notch quality! Supports multiple languages including Chinese, Japanese, Korean, English, and more, with specially optimized formatting rules for natural and professional subtitles. This is a desktop tool based on PySide6 + ElevenLabs API that intelligently converts audio/video files or JSON transcripts into high-quality SRT subtitles, especially suited for CJK and English formatting conventions. Perfect for video creators, editors, course makers, and subtitle producers! https://github.com/cylind/scribe2srt… — # cylind/scribe2srt Source: https://github.com/cylind/scribe2srt # Scribe2SRT Scribe2SRT is a professional audio/video to subtitle tool. By integrating ElevenLabs speech recognition technology and intelligent subtitle segmentation algorithms, subtitle production becomes simple and efficient. ## 🚀 Key Features - 🎯 High-quality transcription: Based on ElevenLabs’ advanced speech recognition technology - 🌍 Multi-language support: Supports Chinese, English, Japanese, Korean, and more - 📝 Professional subtitle standards: Follows industry standards like Netflix for subtitle production - ⚡ Smart segmentation algorithm: Semantic segmentation based on punctuation priority, maintaining sentence integrity - 🔄 Smart retry mechanism: Automatically saves temporary files on failure, quick recovery on retry - 🎨 User-friendly interface: Clean and intuitive graphical user interface with drag-and-drop support - 📊 Real-time progress feedback: Clear progress display and status prompts ## 💻 Installation & Usage ### Quick Start 1. Go to the Releases page (https://github.com/cylind/scribe2srt/releases) and download the latest version 2. Extract and run the program directly 3. It is recommended to install FFmpeg: For video file processing, improving compatibility and efficiency Run from source (click to expand) #### Installation Steps 1. Download the project bash git clone https://github.com/your-username/scribe2srt.git cd scribe2srt 2. Install dependencies bash pip install -r requirements.txt 3. Run the program bash python app.py ## 📖 Usage Instructions ### Basic Workflow 1. Select input file - Click the “Select File” button or drag and drop a file into the program window - Supports three input types: - Audio files: All common audio formats (MP3, WAV, FLAC, M4A, AAC, OGG, etc.) - Video files: All common video formats (MP4, MOV, MKV, AVI, FLV, WEBM, etc.) - JSON transcript files: ElevenLabs format transcription data 2. Configure processing options - Language selection: Select the source language or use “Auto Detect” - Audio event marking: Choose whether to mark non-speech events (e.g., laughter, applause) 3. Start processing - Click the “Generate Subtitles” button to begin transcription - The program will display detailed processing progress 4. Get results - After processing, the SRT subtitle file is automatically saved in the same directory as the source file - The program shows the output file path ### Subtitle Quality Standards This tool follows professional subtitle production standards: - Duration control: Minimum 0.83 seconds, maximum 7.0 seconds - Character density: CJK languages up to 11 characters per second, Latin languages up to 15 characters per second - Line length limit: CJK languages max 25 characters per line, Latin languages max 42 characters per line - Semantic integrity: Prioritizes maintaining sentence completeness, segmentation based on punctuation priority ## ⚙️ Advanced Settings ### Subtitle Parameter Adjustment Via the “Subtitle Settings” menu, you can adjust: - Subtitle display duration and gap - Character density limits - Characters per line limit ### Large File Processing - Automatic segmentation for long files (90+ minutes) - Supports concurrent processing for faster speed - Smart retry mechanism ensures processing success ## 🔧 Technical Highlights ### Smart Segmentation Algorithm - Two-stage processing: Sentence pre-segmentation + intelligent merging - Punctuation priority: Segmentation strategy based on linguistic rules - Semantic integrity: Avoids breaking sentence structure - Multi-language optimization: Differentiated processing for different languages ## 📄 License This project is licensed under the MIT license. — If this project helps you, please give us a ⭐ Star!

Similar Articles

@GitHub_Daily: Downloaded Japanese videos with no subtitles, and whenever I search for subtitle files, they never sync with the timeline — it ruins the viewing experience. So I found WhisperSubTranslate, an open-source desktop app: drag in a video and it generates SRT subtitles, and can even translate them into Chinese. Speech recognition uses OpenAI's open-source Wh…

X AI KOLs Timeline

WhisperSubTranslate is an open-source desktop app that uses OpenAI's Whisper and Tencent's Hy-MT2 model for local video subtitle generation and translation, no internet or registration required.

@mogician301: https://x.com/mogician301/status/2072250774332285073

X AI KOLs Timeline

Introduces an open-source tool called audio-to-text that uses AI and local tools (like faster-whisper and ffmpeg) to help users automatically generate, proofread, and burn subtitles, solving problems such as high costs, frequent errors, and cumbersome workflows in software like CapCut.

@yhslgg: Bro, sharing another open-source video translation tool—pyVideoTrans, with 17,700 stars on GitHub, a must-have for video repurposing and localization! In a nutshell: drop a video in, and it automatically runs through the entire pipeline of speech recognition → subtitle translation → AI dubbing → video synthesis, outputting a complete video in another language. Core...

X AI KOLs Timeline

pyVideoTrans is an open-source video translation tool that supports automatic speech recognition, subtitle translation, AI dubbing, and video synthesis. It integrates multiple ASR, translation, and TTS engines, making it suitable for cross-language video production and localization.

@wsl8297: Want to turn ebooks or documents into audiobooks? Many tools sound too robotic or lack subtitle sync, leaving you frustrated. Then I found the open-source project Abogen: it supports ePub, PDF, plain text, etc., one-click conversion to high-quality audio with auto-generated synchronized subtitles. It uses Kokoro voice at its core…

X AI KOLs Timeline

Abogen is an open-source tool that can convert documents like ePub and PDF into high-quality audio with one click, automatically generating synchronized subtitles. It supports a voice mixer and multiple deployment methods.

@hank_aibtc: Whoa, this thing really blew my mind. Dug up a local tool called KrillinAI, free, video translation, precise subtitles, ultra-natural voiceover, voice cloning, the whole pipeline in one go. Chinese-English translation is especially stable, Whisper recognition accuracy is ridiculously high, LLM translates segment by segment without losing context, CosyV…

X AI KOLs Timeline

Introducing KrillinAI, a free locally-run video translation tool that supports precise subtitles, natural voiceover, and voice cloning. It integrates Whisper, LLM, and CosyVoice, and supports Win/Mac and yt-dlp.