Building a Fully Local PDF Read-Aloud & PDF-to-Audiobook Desktop App with Kokoro 82M, Qwen, and llama.cpp
Summary
Speechfony is a new fully local desktop app that reads PDFs and EPUBs aloud with sentence highlighting, semantic search, and MP3 audiobook export, using Kokoro TTS and on-device embeddings. It prioritizes privacy and offline use, with an open-source MIT license.
Similar Articles
Local, CPU-Friendly, High-Quality TTS (Text-to-Speech) with Kokoro
This article introduces Kokoro, a lightweight 82M-parameter text-to-speech model that runs locally on CPU, providing high-quality speech synthesis across multiple languages while preserving privacy. It explains how to set up Kokoro via a Docker container with an OpenAI-compatible API for easy integration.
@0x0SojalSec: Stop paying for ElevenLabs or cloud TTS. Free Clone voice in just 3 seconds fully Locally on Laptop Turns your docs int…
A free, fully local voice cloning and TTS tool powered by Qwen3-TTS and Kokoro runs on Apple Silicon via MLX, enabling studio-quality audiobook generation from PDFs without cloud services.
@Honcia13: The open-source tool that turns ebooks into audiobooks in seconds is here—Audiblez! Just drop in an EPUB and within minutes it outputs a high-quality M4B audiobook! It uses the Kokoro voice model with only 82M parameters, but the listening experience is incredibly natural. Highlights: Running Animal Farm on a T4 GPU takes only 5 minutes. Supports Chinese, English, and more…
Audiblez is an open-source tool that quickly converts EPUB ebooks into high-quality M4B audiobooks. It uses the Kokoro-82M voice model, supports multiple languages and a graphical interface, and can be installed with a single pip command.
@wsl8297: Want to turn ebooks or documents into audiobooks? Many tools sound too robotic or lack subtitle sync, leaving you frustrated. Then I found the open-source project Abogen: it supports ePub, PDF, plain text, etc., one-click conversion to high-quality audio with auto-generated synchronized subtitles. It uses Kokoro voice at its core…
Abogen is an open-source tool that can convert documents like ePub and PDF into high-quality audio with one click, automatically generating synchronized subtitles. It supports a voice mixer and multiple deployment methods.
Serving Local AI on my Jetson through Durable Streams
A developer documents building a self-hosted text-to-speech app on an NVIDIA Jetson Orin Nano using Kokoro-82M and durable streams, enabling reliable local AI inference with shareable audio outputs.