@GitHub_Daily: When feeding documents to LLMs, Word, PPT, and Excel formats are all different, and the Markdown quality after conversion varies. The Firecrawl team wrote anydoc in Rust, which supports converting 14 office formats to Markdown with a median conversion speed under 5 milliseconds, …

X AI KOLs Timeline Tools

Summary

The Firecrawl team developed anydoc in Rust, an open-source library that quickly converts 14 office formats such as Word, PPT, and Excel into unified Markdown. The median conversion speed is under 5 milliseconds, and it supports Node.js, Python, and browser (WASM) usage.

Feeding documents to large models — Word, PPT, Excel all have different formats, and the Markdown quality produced by conversion can be inconsistent. The Firecrawl team wrote anydoc in Rust, which supports converting 14 office formats to Markdown with a median conversion speed under 5 milliseconds. It has already gained 12,000+ Stars! The Markdown structure is consistent across all formats, preserving tables, footnotes, and nested lists, so results won't change just because you switch formats. GitHub: http://github.com/firecrawl/anydoc… It can be used with Node.js, Python, and in the browser. The browser version converts files locally without uploading to a server. If you regularly need to convert documents in various formats for LLMs to process, this library is quite handy.
Original Article
View Cached Full Text

Cached at: 08/10/26, 03:32 PM

When feeding documents to large language models, Word, PPT, and Excel each have different formats, and the quality of the Markdown converted from them varies. The Firecrawl team built anydoc in Rust, supporting conversion of 14 office formats to Markdown with a median conversion speed of under 5 milliseconds—and it has already earned 12,000+ Stars! Every format converts to Markdown with a consistent structure; tables, footnotes, and nested lists are preserved, so the result doesn’t change just because you switched formats. GitHub: http://github.com/firecrawl/anydoc… It has three usage modes: Node.js, Python, and the browser; the browser version converts files locally without sending them to a server. For anyone who regularly needs to convert documents in various formats for LLM processing, this library is quite convenient.

firecrawl/anydoc Source: https://github.com/firecrawl/anydoc

anydoc Crates.io (https://crates.io/crates/anydoc) npm (https://www.npmjs.com/package/@firecrawl/anydoc) PyPI (https://pypi.org/project/firecrawl-anydoc/) License: MIT skills.sh (https://skills.sh/firecrawl/anydoc)

Fast Rust library that converts documents (Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, and PDF) into clean GitHub-Flavored Markdown. Includes bindings for Node.js, Python, and the browser (WebAssembly). Built by Firecrawl (https://firecrawl.dev) to turn any office document into LLM-ready Markdown in single-digit milliseconds, with one consistent output no matter which format goes in. It powers Firecrawl Parse (https://firecrawl.dev/parse), so if you’d rather not run it yourself, the hosted API gives you the same conversion plus our OCR models for the scanned pages anydoc can’t read on its own.

Try it in your browser (https://firecrawl.github.io/anydoc/): the demo page runs the library as WebAssembly, so files are converted locally and never leave your machine.

Quick start

Agent skill

anydoc ships as an Agent Skill (https://agentskills.io), so your agent can read any document it runs into:

npx skills add firecrawl/anydoc

The skill teaches the agent to convert documents with the anydoc CLI. Works with Claude Code (https://claude.ai/code), Codex (https://openai.com/codex/), Cursor (https://cursor.com), OpenCode (https://opencode.ai), and any other compatible agent (https://agentskills.io/clients).

CLI

npx @firecrawl/anydoc report.docx # Markdown to stdout
npx @firecrawl/anydoc slides.pptx -o slides.md # or to a file
npx @firecrawl/anydoc - --format csv < data.csv # read stdin

npx downloads the prebuilt binary for your platform on first run. For a permanent anydoc command, install globally with npm install -g @firecrawl/anydoc. Run anydoc --help for all options.

Node.js

npm install @firecrawl/anydoc
import { toDocument, toMarkdown, toMarkdownBytes } from '@firecrawl/anydoc';
// From a file path:
const markdown = await toMarkdown('report.docx');
// From bytes, with the format detected from the content:
const fromBytes = await toMarkdownBytes(bytes);
// Or name it, which signature-less formats (CSV) need:
const fromCsv = await toMarkdownBytes(bytes, 'csv');
// Or stop at the document model, which also carries embedded assets:
const document = await toDocument(bytes);

Full API reference: node/README.md

Python

pip install firecrawl-anydoc
import anydoc
# From a file path:
markdown = anydoc.to_markdown("report.docx")
# From bytes, with the format detected from the content:
markdown = anydoc.to_markdown_bytes(data)
# Or name it, which signature-less formats (CSV) need:
markdown = anydoc.to_markdown_bytes(data, "csv")
# Or stop at the document model, which also carries embedded assets:
document = anydoc.to_document(data)

Full API reference: python/README.md

Browser (WebAssembly)

npm install @firecrawl/anydoc-wasm
import init, { toMarkdownBytes, toDocument } from '@firecrawl/anydoc-wasm';
await init();
// From bytes, with the format detected from the content:
const markdown = toMarkdownBytes(bytes);
// Or name it, which signature-less formats (CSV) need:
const fromCsv = toMarkdownBytes(bytes, 'csv');
// Or stop at the document model, which also carries embedded assets:
const document = toDocument(bytes);

Full API reference: wasm/README.md

Rust

cargo add anydoc
// From a file path:
let markdown = anydoc::to_markdown("report.docx")?;
// From bytes, with the format detected from the content:
let markdown = anydoc::to_markdown_bytes(&bytes, None)?;
// Or name it, which signature-less formats (CSV) need:
let markdown = anydoc::to_markdown_bytes(&bytes, anydoc::Format::Csv)?;
// Or stop at the document model, which also carries embedded assets:
let document = anydoc::to_document(&bytes, None)?;

Features

  • One output for every format. Each format parses into a shared document model and renders through a single Markdown serializer, so escaping, tables, heading anchors, and footnotes behave identically whether the input was a .doc from 2003 or a .pptx from yesterday.
  • Full document structure. Headings with anchors, bold/italic/strikethrough, inline code and code blocks, links and internal cross-references, bulleted/numbered/nested/task lists with the source’s own numbering, tables with merged cells and header rows, block quotes, footnotes and endnotes, and speaker notes.
  • Embedded assets. Images and embedded objects render as their alt text in the Markdown, and the raw bytes stay available on the document model, tagged with their media type. Images with an external URL become ordinary Markdown images.
  • Content-based format detection. The format is read from the bytes themselves (PDF header, RTF open group, OLE stream names, ZIP package mimetype), so mislabeled files still convert correctly.
  • Fast. Pure Rust, no ML models, no external services. Median conversion time is under 5ms per document.
  • Bindings that stay out of the way. Node.js conversion runs on the libuv thread pool and never blocks the event loop; Python releases the GIL so other threads keep running. TypeScript types and Python stubs ship with the packages.
  • PDF support built in. Text-based PDFs convert locally through pdf-inspector (https://github.com/firecrawl/pdf-inspector), no OCR service required.
  • Agent ready. Ships as an Agent Skill: one npx skills add firecrawl/anydoc and any agent can read office documents.

Supported formats

FormatExtensions
Word.doc, .docx, .docm
PowerPoint.ppt, .pps, .pot, .pptx, .pptm, .ppsx, .ppsm
Excel.xls, .xlsx, .xlsm, .xlsb
OpenDocument.odt, .ods, .odp
Rich Text Format.rtf
EPUB.epub
CSV.csv
PDF.pdf

Benchmark

anydoc is measured against six other converters on 100 real-world documents spanning fourteen formats. Scores run from 0 to 100, higher is better; speed is the median time to convert one document.

toolformatsmedian msdocs judgedscorecompletenessstructureformattingcleanliness
anydoc14/144.4948187797881
libreoffice12/141129.5874059424024
unstructured8/14572.9586376595163
markitdown6/14134.8336578666052
pandoc5/14102.1345674575638
docling4/14513.6215760605751
mammoth1/1452.587084717551

Per format, like for like:

formatanydoclibreofficeunstructuredmarkitdownpandocdoclingmammoth
doc875767----
docm8448-----
docx88565371687170
epub77-727252--
odp8623-----
ods8238-----
odt805168-60--
ppt8026-----
pptx7424-66-52-
rtf885346-45--
xls80386662---
xlsm7632-----
xlsx72306655-47-

How quality was scored: an LLM judge (Claude Sonnet 5) compares two tools’ outputs blind against ground truth: the document’s first six pages, rendered to images by LibreOffice. Each output is scored on completeness, structure, formatting, and cleanliness. Every pair is judged twice with the outputs swapped to cancel position bias, for 482 verdicts in total. Each tool’s score averages its per-format scores over the formats it supports, so a corpus heavy in one format can’t skew it. It also means each row averages a different set of formats (mammoth’s 69 is docx alone, while anydoc’s 81 spans all fourteen), so the per-format table is the fair comparison. Speed is one warm conversion per document on a Ryzen 9 9950X3D (Windows 11, 64 GB DDR5-6400). anydoc and the Python libraries are timed with process spawn excluded; the CLI tools include it, since that is how they are used. The harness lives in bench/; the corpus is not redistributable and is not in the repo.

Best fit: pipelines that receive a mixed bag of office documents and need one consistent, structured Markdown output. In this comparison, anydoc was the only tool to cover all fourteen formats, scored highest on every judged format, and converted documents an order of magnitude faster than the next-fastest tool.

Format detection

The format is read from the file content, using the marker its specification designates: the PDF header, the RTF open group, OLE stream names, the ZIP package mimetype and content types. CSV has no such marker, so the extension or an explicit format names it instead.

Format::from_bytes(&bytes); // Some(Format::Docx), or None when nothing matches
Format::from_extension("pptm"); // Some(Format::Pptx)
Format::from_path(Path::new("report.odt")); // Some(Format::Odt)

The same three functions exist in Node (formatFromBytes, …) and Python (anydoc.format_from_bytes, …).

Errors

A conversion returns Err only when no meaningful Markdown could come out of the file. ConvertError names what went wrong:

match anydoc::to_markdown(path) {
    Ok(markdown) => Some(markdown),
    // No document comes out of these, so record the file and take the next one.
    Err(error @ (ConvertError::Encrypted | ConvertError::Unsupported(_))) => {
        unconverted.push((path, error));
        None
    }
    Err(error) => return Err(error),
}
VariantMeaning
UnsupportedUnknown format, or one that cannot be converted (an image-only PDF)
MalformedStructurally unusable: no meaningful content could be extracted
EncryptedEncrypted or password-protected
ResourceLimitCrossed a fixed safety limit (decompression, nesting, node count)
MissingPartA part required for any meaningful output is absent
IoThe file could not be read, from to_markdown only

Node and wasm publish the variant name on error.code; Python raises one anydoc.ConvertError subclass per variant, or OSError when the file cannot be read.

How it works

document bytes
│
├─► format detection → content markers, not the extension
│
├─► format parser → one per format (doc, docx, ppt, pptx, xls,
│   xlsx, odt/ods/odp, rtf, epub, csv)
│   │
│   └─► Document → shared model: blocks, inlines, tables,
│       footnotes, assets
│   │
│   └─► GFM serializer → Markdown
│
└─► PDF → pdf-inspector → Markdown directly

Because every format funnels through the same document model and serializer, output quirks get fixed once. A table-escaping fix for docx is automatically a table-escaping fix for rtf, odt, and everything else.

Development

cargo test
cd node && npm install && npm run build && npm test
cd python && pip install maturin && maturin develop && python -m unittest discover -s tests
wasm-pack build wasm --release --target web --scope firecrawl && node --test wasm/test.mjs # see wasm/README.md

A committed fixture corpus under tests/fixtures/ is snapshot-tested, tests/robustness.rs mutation-tests every fixture, and fuzz/ carries cargo-fuzz targets per format. The speed and quality benchmark lives in bench/.

Releases are tagged v, which publishes the crate, the npm package, and the PyPI wheels from .github/workflows/release.yml. The version lives in three places, bumped together for a release:

License

MIT

Similar Articles

@GitHub_Daily: Feeding PDF documents to large language models, a common practice is to run OCR first and then extract text, but in reality more than half of PDFs are already text-based and don't need OCR at all. The pdf-inspector recently open-sourced by the Firecrawl team can automatically determine whether a PDF is text-based or scanned, and for text-based ones...

X AI KOLs Timeline

The Firecrawl team open-sourced pdf-inspector, a fast Rust library that automatically determines whether a PDF is text-based or scanned, directly extracts text and converts it to Markdown, avoiding unnecessary OCR, supports table and multi-column layout detection, and provides Python/Node.js/Rust/WASM interfaces.

@Chenzeze777: Microsoft open-sourced a document tool with 140k stars — I compiled its 5 most practical use cases. MarkItDown, a Python tool, converts PDF/Word/PPT/Excel/HTML/images into clean Markdown text with one click. What you can do with it: · P…

X AI KOLs Timeline

Microsoft open-sourced MarkItDown, a lightweight Python tool that converts PDF, Word, PPT, Excel, HTML, and images into clean, structured Markdown text in one go, ideal for AI summarization, data analysis, knowledge base construction, and more.

@AIExplorerTim: Someone just released a tool that converts PDFs into clean, structured Markdown at speeds up to 100 pages/second. No GPU required. No API costs. No messy parsing. Just raw, usable data. It handles with ease: • Tables → Perfectly ex…

X AI KOLs Timeline

OpenDataLoader is an open-source tool that converts PDFs into structured Markdown and JSON, supporting local processing speeds of up to 100 pages/second without requiring a GPU or incurring API costs, designed specifically for RAG pipelines and PDF accessibility automation.