@GitHub_Daily: File format conversion, document recognition, audio-to-text – I've collected a bunch of various tools, but finding them is a hassle. filewizard wraps these tools into a self-hostable web page. Just drag files in and convert with a few clicks in the browser. Under the hood, it bundles FFmpeg, LibreOffice, Pandoc …

X AI KOLs Timeline Tools

Summary

filewizard is a self-hosted web tool that integrates mature tools like FFmpeg, LibreOffice, and Pandoc, supporting file format conversion, OCR, and audio transcription. Users can deploy it with one click via Docker, and file processing is done locally.

File format conversion, document recognition, audio-to-text – I've collected a bunch of various tools, but finding them is a hassle. filewizard wraps these tools into a self-hostable web page. Just drag files in and convert with a few clicks in the browser. Under the hood, it bundles mature tools like FFmpeg, LibreOffice, and Pandoc, covering a wide range of format conversions. It can perform OCR on images and PDFs, and transcribe audio into text using Whisper, with background tasks and real-time progress. GitHub: http://github.com/LoredCast/filewizard… It provides a one-click local deployment via Docker, so the project runs on your own machine and files are never sent to third parties. If you frequently deal with various file formats, set one up for yourself as an online toolbox.
Original Article
View Cached Full Text

Cached at: 07/16/26, 06:06 AM

File format conversion, document OCR, and audio-to-text — I’ve collected a bunch of different tools, but finding the right one is a hassle.

FileWizard bundles these tools into a self-hosted web page; just drag in a file to convert, with only a few clicks in your browser.

Under the hood, it wraps mature tools like FFmpeg, LibreOffice, and Pandoc, covering a broad set of format conversions.

You can run OCR on images and PDFs, use Whisper to transcribe audio into text, and it includes background tasks and real-time progress.

GitHub: http://github.com/LoredCast/filewizard…

It provides a one-click local deployment via Docker — the project runs on your own machine, so files don’t go to any third party.

If you frequently juggle various file formats, set one up yourself as an online toolbox.


LoredCast/filewizard

Source: https://github.com/LoredCast/filewizard

File Wizard

PayPal (https://www.paypal.me/unterrikermanu) Docker Pulls (https://hub.docker.com/r/loredcast/filewizard) Docker Image Version (https://hub.docker.com/r/loredcast/filewizard)

A self-hosted, browser-based utility for file conversion, OCR and audio transcription. It wraps common CLI and Python converters (FFmpeg, LibreOffice, Pandoc, ImageMagick, etc.), plus faster-whisper and Tesseract OCR.

Screenshot

Features

  • Convert between many file formats; extendable via settings.yml to add any CLI tool.
  • OCR for PDFs and images (tesseract / ocrmypdf).
  • Audio transcription using Whisper models.
  • Simple, responsive dark UI with drag-and-drop and file picker.
  • Background job processing with real-time status updates and persistent history.
  • /settings page for configuring conversion tools and OAuth (runs without auth in local mode).
  • CPU-only by default; a -cuda image is available for GPU use.

Security

Warning: exposing this app publicly without authentication risks arbitrary code execution. Intended for local use or behind a properly configured OAuth/OIDC provider.

Tech stack

FastAPI, vanilla HTML/JS/CSS frontend.

Installation

Recommended — Docker (pull from Docker Hub)

Images available:

  • loredcast/filewizard:latest (newest full release without cuda)
  • loredcast/filewizard:0.3-small (omits TeX and other large tools)
  • loredcast/filewizard:0.3-cuda (CUDA-enabled)

``

docker-compose.yml

version: “3.9” services: web: image: loredcast/filewizard:latest environment: - LOCAL_ONLY=True # False for Auth - SECRET_KEY= # set if using auth - UPLOADS_DIR=/app/uploads # inside the container - PROCESSED_DIR=/app/processed # inside the container - OMP_NUM_THREADS=1 - DOWNLOAD_KOKORO_ON_STARTUP=true ports: - “6969:8000” volumes: - ./config:/app/config # settings.yml will be here - ./uploads_data:/app/uploads - ./processed_data:/app/processed volumes: uploads_data: {} processed_data: {} ``

Copy docker-compose.yml from the repo or the above, adjust as needed, then:

bash docker compose up -d FileWizard will be available at localhost:6969

Build locally with Docker (new build types)

For different build configurations, use the BUILD_TYPE argument:

``bash

Full build (includes all dependencies but no CUDA)

docker build –build-arg BUILD_TYPE=full -t filewizard:full .

Small build (excludes TeX and markitdown dependencies for smaller image)

docker build –build-arg BUILD_TYPE=small -t filewizard:small .

CUDA build (includes CUDA support for GPU acceleration)

docker build –build-arg BUILD_TYPE=cuda -t filewizard:cuda . ``

Or with docker-compose:

``bash

For full build

docker compose build –build-arg BUILD_TYPE=full

For small build

docker compose build –build-arg BUILD_TYPE=small

For CUDA build

docker compose build –build-arg BUILD_TYPE=cuda ``

For CUDA builds, ensure you have:

  • NVIDIA Docker runtime installed (nvidia-docker2 package)
  • Compatible GPU with appropriate drivers
  • Add the GPU configuration to docker-compose.yml if building with compose: yaml deploy: resources: reservations: devices: - driver: nvidia count: all capabilities: [gpu]

For troubleshooting GPU issues, make sure:

  1. Your GPU drivers support the CUDA version (12.1)
  2. cuDNN libraries are properly installed in the container
  3. The nvidia-container-toolkit is properly configured
  4. Test NVIDIA setup with: docker run --rm --gpus all nvidia/cuda:12.1-base-ubuntu22.04 nvidia-smi

bash git clone https://github.com/LoredCast/filewizard.git cd filewizard docker compose up --build Note: building can be slow (TeX and other dependencies).

Manual (no Docker)

bash git clone https://github.com/LoredCast/filewizard.git cd filewizard python -m venv venv source venv/bin/activate # Windows: venv\\Scripts\\activate pip install -r requirements.txt chmod +x run.sh ./run.sh

Dependencies include fastapi, uvicorn, sqlalchemy, huey, faster-whisper, ocrmypdf, pytesseract, python-multipart, pyyaml, etc.

Configuration & docs

See the project Wiki for details and examples:
https://github.com/LoredCast/filewizard/wiki

Usage

  1. Open http://127.0.0.1:8000.
  2. Drag & drop or choose files.
  3. Select action: Convert, OCR, or Transcribe.
  4. Track job progress in the History table (updates automatically).

Tools Table

ToolCommon inputs (extensions / format names)Common outputs (extensions / format names)Notes
LibreOffice (soffice).odt, .fodt, .ott, .doc, .docx, .docm, .dot, .dotx, .rtf, .txt, .html/.htm/.xhtml, .xml, .sxw, .wps, .wpd, .abw, .pdb, .epub, .fb2, .lit, .lrf, .pages, .csv, .tsv, .xls, .xlsx, .xlsm, .ods, .sxc, .123, .dbf, .fb2.pdf, .pdfa, .odt, .fodt, .doc, .docx, .rtf, .txt, .html/.htm, .xhtml, .epub, .svg, .png, .jpg/.jpeg, .pptx, .ppt, .odp, .xls, .xlsx, .ods, .csv, .dbf, .pdb, .fb2Good for office/document conversions; fidelity varies with complex layouts.
PandocMarkdown flavors (.md, .markdown), .html/.htm, LaTeX (.tex), .rst, .docx, .odt, .epub, .ipynb, .opml, .adoc/asciidoc, .tex, .bib/citation inputs.html/.html5, .xhtml, .latex/.tex, .pdf (via LaTeX engine), .docx, .odt, .epub, .md (flavors), .gfm, .rst, .pptx, .man, .mediawiki, .docbookHighly configurable via templates/filters; requires LaTeX for PDF output.
Ghostscript (gs).ps, .eps, .pdf, PostScript streams.pdf (various compat levels incl PDF/A), .ps, .eps, raster images (.png, .jpg, .tiff, .bmp, .pnm)Useful for PDF manipulations, rasterization, and producing PDF/A.
Calibre (ebook-convert).epub, .mobi, .azw3, .azw, .fb2, .html, .docx, .doc, .rtf, .txt, .pdb, .lit, .tcr, .cbz, .cbr, .odt, .pdf (input with caveats).epub, .mobi (legacy), .azw3, .pdf, .docx, .rtf, .txt, .fb2, .htmlz, .pdb, .lrf, .lit, .tcr, .cbz, .cbrExcellent for ebook format conversions and metadata handling; PDF input/output fidelity varies.
FFmpegContainers & codecs: .mp4, .mkv, .mov, .avi, .webm, .flv, .wmv, .mpg/.mpeg, .ts, .m2ts, .3gp, audio: .mp3, .wav, .aac/.m4a, .flac, .ogg, .opus, image sequences (.png, .jpg, .tiff), HLS (.m3u8)Wide set: .mp4, .mkv, .mov, .webm, .avi, .flv, .mp3, .aac/.m4a, .wav, .flac, .ogg, .opus, .gif (animated), .ts, elementary streams, many codec/container combosExtremely versatile — audio/video transcoding, extraction, container changes, filters. Supported formats depend on build flags and linked libraries.
libvips (vips).jpg/.jpeg, .png, .tif/.tiff, .webp, .avif, .heif/.heic, .jp2, .gif (frames), .pnm, .fits, .exr, PDF (via poppler delegate).jpg/.jpeg, .png, .tif/.tiff, .webp, .avif, .heif, .jp2, .pnm, .v (VIPS native), .fits, .exrFast, memory-efficient image processing; great for large images and tiling.
GraphicsMagick (gm).jpg/.jpeg, .png, .gif, .tif/.tiff, .bmp, .ico, .eps, .pdf (via Ghostscript/poppler), .dpx, .pnm, .svg (if delegate), .webp (if built), .exr.jpg/.jpeg, .png, .webp (if enabled), .tif/.tiff, .gif, .bmp, .pdf (from images), .eps, .ico, .xpm, .dpxSimilar to ImageMagick but with different performance/behavior; supported formats depend on build/delegates.
ImageMagick (convert / magick)Same as GraphicsMagick (large set; many delegates)Same as GraphicsMagickOften used interchangeably; watch for security considerations when processing untrusted images.
Inkscape.svg/.svgz, .pdf, .eps, .ps, .ai (legacy imports), .dxf, raster images (.png, .jpg, .jpeg, .gif, .tiff, .bmp).svg, .pdf, .ps, .eps, .png, .emf, .wmf, .xaml, .dxf, .epsVector editing and export; CLI useful for batch SVG → PNG/PDF conversions.
libjxl (cjxl / djxl)Raster inputs: .png, .jpg/.jpeg, .ppm/.pbm/.pgm, .gif, etc..jxl (JPEG XL)Encoder/decoder for JPEG XL; availability depends on build.
resvg.svg/.svgz.png (raster)Fast, accurate SVG renderer — good for SVG→PNG conversion.
PotraceBitmaps: .pbm, .pgm, .ppm (PNM family), .bmp (via conversion)Vector: .svg, .pdf, .eps, .ps, .dxf, .geojsonTraces bitmaps to vector paths; often used with pre-conversion steps.
Potrace GUI / autotrace alternativesNot included but sometimes available in toolchains; behavior varies.
MarkItDown / markitdown.pdf, .docx, .doc, .pptx, .ppt, .xlsx, .xls, .html, .eml, .msg, .md, .txt, images, .epub.md (Markdown)Utility to extract/produce Markdown from various formats; implementation details vary.
pngquant.png (truecolor/rgba).png (quantized palette PNG)Lossy PNG quantization for smaller PNGs.
MozJPEG (cjpeg, jpegtran).ppm/.pbm/.pgm (PNM), .bmp, existing .jpg.jpg/.jpeg (MozJPEG-optimized)Produces smaller JPEGs with improved compression; good for recompression.
SoX (Sound eXchange).wav, .aiff, .mp3 (if libmp3lame), .flac, .ogg/.oga, .raw, .au, .voc, .w64, .gsm, .amr, .m4a (if libs present).wav, .aiff, .flac, .mp3, .ogg, .raw, .w64, .opus, .amr, .m4aAudio processing, normalization, effects; exact formats depend on linked libraries.
Tesseract OCR / ocrmypdfImage formats (.png, .jpg, .jpeg, .tiff), PDFs (image PDFs)Plain text (.txt), searchable PDF (PDF with text layer), HOCR, ALTO XMLOCR engine; language/training data required for best accuracy. ocrmypdf is a wrapper for PDF workflows.
faster-whisper / OpenAI Whisper (local)Audio: .mp3, .wav, .m4a, .flac, .ogg, .opus, .aacPlain text transcripts (.txt), .srt, .vtt, other subtitle formatsLocal Whisper implementations for speech-to-text. Models and speed depend on CPU/GPU and model variant.

Consider spending 30 seconds on the » usage survey « (https://app.youform.com/forms/sdlb6mto) to help improve the app and suggest changes!

ko-fi (https://ko-fi.com/loredcast)

Similar Articles

@yhslgg: Folks, today I'd like to introduce you to a hidden gem with 21,000 stars on GitHub. I tried it myself and was blown away! It's called SingleFile. It does one thing: completely packages any web page into a single HTML file and saves it locally. Images, CSS, fonts, styles — all bundled into one…

X AI KOLs Timeline

SingleFile is a free and open-source browser extension and CLI tool that can fully package any web page into a single HTML file, preserving images, CSS, and more. It works permanently offline and supports auto-save, batch save, and cloud storage.

@0xQiYan: Brothers, have you ever encountered situations where various format conversions require a membership, and still worry about not having one? Discovered an open-source project for format conversion that Microsoft and Google couldn't achieve, but a philosophy professor managed to do in his spare time. Pandoc—the document conversion artifact, one command, a few seconds, over 50 formats freely converted. Word to PDF, ...

X AI KOLs Timeline

Introducing the open-source document conversion tool Pandoc, developed in spare time by philosophy professor John MacFarlane, supporting conversion between over 50 formats, free, open source, and fully local.

@GitHub_Daily: Found another handy Skill — one sentence turns any content into a podcast, PPT, mind map, and more. It supports over 15 content sources, including WeChat official accounts, podcasts, YouTube videos, PDFs, ebooks, etc. It can also automatically detect and attempt to bypass paywalls, covering 300+ sites like The New York Times, The Wall Street Journal, etc.

X AI KOLs Timeline

An open-source Claude Code Skill that converts over 15 content sources (including WeChat articles, YouTube, and paywalled news) into podcasts, PPTs, mind maps, and more in one click, with automatic paywall bypass attempts.

@GitHub_Daily: Trying to feed webpage content to AI, but ending up with a bunch of navigation bars, ads, and garbled text, wasting most of the context window, and AI still can't understand it. So I found this open-source project PullMD, which can extract any webpage content and convert it into clean Markdown files. Just provide a URL, auto-detect page type, layer by layer...

X AI KOLs Timeline

PullMD is an open-source URL to Markdown service that automatically extracts the main content of a webpage, removing navigation, ads, and other clutter. It supports headless browsers and multiple interfaces (web, REST API, MCP), making it easy for AI tools and users to obtain clean webpage text.