@GitHub_Daily: 给大模型喂 PDF 文档,常见做法是先跑 OCR 再提文字,但实际上有一半多的 PDF 本来就是文字版,根本不需要 OCR。 Firecrawl 团队最近开源的 pdf-inspector,能自动判断 PDF 是文字版还是扫描版,文字版的…

X AI KOLs Timeline 工具

摘要

Firecrawl 团队开源了 pdf-inspector,一个快速 Rust 库,可自动判断 PDF 是文字版还是扫描版,并直接提取文字转成 Markdown,免去不必要的 OCR,支持表格、多栏排版识别,提供 Python/Node.js/Rust/WASM 接口。

给大模型喂 PDF 文档,常见做法是先跑 OCR 再提文字,但实际上有一半多的 PDF 本来就是文字版,根本不需要 OCR。 Firecrawl 团队最近开源的 pdf-inspector,能自动判断 PDF 是文字版还是扫描版,文字版的直接提取并转成 Markdown。 最快 200 毫秒内提取完成,支持表格检测、多栏排版识别、标题层级还原,输出的 Markdown 结构干净。 GitHub:http://github.com/firecrawl/pdf-inspector… 提供 Python、Node.js、Rust 和浏览器端四套接口,浏览器里通过 WebAssembly 直接跑,不用后端服务。 做文档处理相关开发的朋友,拿来做 PDF 前置分类和提取挺合适的。
查看原文
查看缓存全文

缓存时间: 2026/08/06 04:33

给大模型喂 PDF 文档,常见做法是先跑 OCR 再提文字,但实际上有一半多的 PDF 本来就是文字版,根本不需要 OCR。

Firecrawl 团队最近开源的 pdf-inspector,能自动判断 PDF 是文字版还是扫描版,文字版的直接提取并转成 Markdown。

最快 200 毫秒内提取完成,支持表格检测、多栏排版识别、标题层级还原,输出的 Markdown 结构干净。

GitHub:http://github.com/firecrawl/pdf-inspector…

提供 Python、Node.js、Rust 和浏览器端四套接口,浏览器里通过 WebAssembly 直接跑,不用后端服务。

做文档处理相关开发的朋友,拿来做 PDF 前置分类和提取挺合适的。


firecrawl/pdf-inspector

Source: https://github.com/firecrawl/pdf-inspector

pdf-inspector

Crates.io npm PyPI License: MIT

Fast Rust library for PDF classification and text extraction. Detects whether a PDF is text-based or scanned, extracts text with position awareness, and converts to clean Markdown — all without OCR. Includes bindings for Python, Node.js, and browser WebAssembly.

Built by Firecrawl to handle text-based PDFs locally in under 200ms, skipping expensive OCR services for the ~54% of PDFs that don’t need them.

Features

  • Smart classification — Detect TextBased, Scanned, ImageBased, or Mixed PDFs in ~10-50ms by sampling content streams. Returns a confidence score (0.0-1.0) and per-page OCR routing.
  • Text extraction — Position-aware extraction with font info, X/Y coordinates, and automatic multi-column reading order.
  • Markdown conversion — Headings (H1-H4 via font size ratios), bullet/numbered/letter lists, code blocks (monospace font detection), tables (rectangle-based and heuristic), bold/italic formatting, URL linking, and page breaks.
  • Table detection — Dual-mode: rectangle-based detection from PDF drawing ops, plus heuristic detection from text alignment. Handles financial tables, footnotes, and continuation tables across pages.
  • CID font support — ToUnicode CMap decoding for Type0/Identity-H fonts, UTF-16BE, UTF-8, and Latin-1 encodings.
  • Multi-column layout — Automatic detection of newspaper-style columns, sequential reading order, and RTL text support.
  • Encoding issue detection — Automatically flags broken font encodings so callers can fall back to OCR.
  • Single document load — The document is parsed once and shared between detection and extraction, avoiding redundant I/O.
  • Browser WebAssembly — Run the same Rust parser locally in browsers and Web Workers, with embedded CMaps and no server round trip.
  • Lightweight — Pure Rust, no ML models, no external services. Single dependency on lopdf for PDF parsing.

Benchmark

Evaluated on the opendataloader-bench corpus (200 PDFs). Only local engines without model-based PDF parsing are shown; OCR was disabled. Scores are 0-1, higher is better.

EngineOverallReading Order (NID)Tables (TEDS)Headings (MHS)Speed (200 docs)
pdf-inspector0.8750.9150.8140.7880.470s
liteparse0.8730.9130.6930.8110.750s
opendataloader0.8310.9020.4890.7392.569s
pymupdf4llm0.7350.8860.4010.42417.117s
markitdown0.5890.8440.2730.00016.165s

Results were refreshed on July 31, 2026, on an Apple M4 Pro. Engine versions were pdf-inspector 0.2.6, LiteParse 2.10.1, OpenDataLoader 2.2.1, PyMuPDF4LLM 0.2.0, and MarkItDown 0.1.5. Speed is the median of five alternating or rotating complete corpus runs after an excluded warm-up run, with each parser processing documents sequentially in a single process.

The complete parser configuration, per-document predictions, evaluator output, and generated charts are available in the reproducible results branch.

Best fit: Native-text PDFs where speed, reading order, and table structure matter. In this comparison, pdf-inspector delivered the higher overall, reading-order, and table scores, along with the fastest complete run. That makes it a strong local default for reports, research papers, financial documents, invoices, and legal PDFs that need clean, structured Markdown without adding OCR latency or infrastructure.

Use the paired benchmark harness to compare two local builds against the exact same corpus and evaluator revision.

Quick start

Python

pip install maturin
maturin develop --release
import pdf_inspector

result = pdf_inspector.process_pdf("document.pdf")
print(result.pdf_type)   # "text_based", "scanned", "image_based", "mixed"
print(result.markdown)   # Markdown string or None

Full API reference: docs/python.md

Node.js

npm install @firecrawl/pdf-inspector
import { readFileSync } from 'fs';
import { processPdf, classifyPdf } from '@firecrawl/pdf-inspector';

const result = processPdf(readFileSync('document.pdf'));
console.log(result.pdfType);   // "TextBased", "Scanned", "ImageBased", "Mixed"
console.log(result.markdown);  // Markdown string or null

Full API reference: napi/README.md

Browser WebAssembly

npm install @firecrawl/pdf-inspector-wasm
import init, { processPdf } from '@firecrawl/pdf-inspector-wasm';

await init();
const response = await fetch('/document.pdf');
const pdf = new Uint8Array(await response.arrayBuffer());
const result = processPdf(pdf);

console.log(result.pdfType);
console.log(result.markdown);

Full API reference: wasm/README.md

Rust

Install from crates.io:

cargo add pdf-inspector

Or add it manually:

[dependencies]
pdf-inspector = "0.1"
use pdf_inspector::process_pdf;

let result = process_pdf("document.pdf")?;
println!("Type: {:?}", result.pdf_type);
if let Some(markdown) = &result.markdown {
    println!("{}", markdown);
}

Full API reference: docs/rust-api.md

CLI

# Install the CLI tools
cargo install pdf-inspector

# Convert PDF to Markdown
pdf2md document.pdf

# JSON output (for piping)
pdf2md document.pdf --json

# Positioned TextItem JSON, including is_underline metadata
pdf2md document.pdf --items-json

# Raw markdown only (no headers)
pdf2md document.pdf --raw

# Token-efficient output (collapses long dot leaders and similar source padding)
pdf2md document.pdf --compact

# Insert page break markers (<!-- Page N -->)
pdf2md document.pdf --pages

# Process only specific pages
pdf2md document.pdf --select-pages 1,3,5-10

# Detection only (no extraction)
detect-pdf document.pdf
detect-pdf document.pdf --json

# Detection + layout analysis (tables, columns)
detect-pdf document.pdf --analyze --json

From a source checkout, use cargo run --bin pdf2md -- document.pdf or cargo run --bin detect-pdf -- document.pdf instead.

Architecture

PDF bytes
  │
  ├─► detector         → PdfType (TextBased / Scanned / ImageBased / Mixed)
  │
  └─► extractor
        ├─ fonts        → font widths, encodings
        ├─ content_stream → walk PDF operators → TextItems + PdfRects
        ├─ xobjects     → Form XObject text, image placeholders
        ├─ links        → hyperlinks, AcroForm fields
        └─ layout       → column detection → line grouping → reading order
              │
              ├─► tables
              │     ├─ detect_rects      → rectangle-based tables (union-find)
              │     ├─ detect_heuristic  → alignment-based tables
              │     ├─ grid              → column/row assignment → cells
              │     └─ format            → cells → Markdown table
              │
              └─► markdown
                    ├─ analysis     → font stats, heading tiers
                    ├─ preprocess   → merge headings, drop caps
                    ├─ convert      → line loop + table/image insertion
                    ├─ classify     → captions, lists, code
                    └─ postprocess  → cleanup → final Markdown

The document is loaded once via load_document_from_path / load_document_from_mem and shared between the detection and extraction stages, so there’s no redundant parsing.

Project structure

src/
  lib.rs                — Public API, PdfOptions builder, convenience functions
  python.rs             — PyO3 Python bindings
  types.rs              — Shared types: TextItem, TextLine, PdfRect, ItemType
  text_utils.rs         — Character/text helpers (CJK, RTL, ligatures, bold/italic)
  process_mode.rs       — ProcessMode enum (DetectOnly, Analyze, Full)
  detector.rs           — Fast PDF type detection without full document load
  glyph_names.rs        — Adobe Glyph List → Unicode mapping
  tounicode.rs          — ToUnicode CMap parsing for CID-encoded text
  extractor/            — Text extraction pipeline
  tables/               — Table detection and formatting
  markdown/             — Markdown conversion and structure detection
  bin/                  — CLI tools (pdf2md, detect_pdf)
napi/                   — Node.js/Bun bindings (napi-rs)
wasm/                   — Browser bindings (wasm-bindgen)

How classification works

  1. Parse the xref table and page tree (no full object load)
  2. Select pages based on ScanStrategy (default: all pages with early exit)
  3. Look for Tj/TJ (text operators) and Do (image operators) in content streams
  4. Classify based on text operator presence across sampled pages

This detects 300+ page PDFs in milliseconds. The result includes pages_needing_ocr — a list of specific page numbers that lack text, enabling per-page OCR routing instead of all-or-nothing.

Scan strategies

StrategyBehaviorBest for
EarlyExit (default)Scan all pages, stop on first non-text pagePipelines routing TextBased PDFs to fast extraction
FullScan all pages, no early exitAccurate Mixed vs Scanned classification
Sample(n)Sample n evenly distributed pages (first, last, middle)Very large PDFs where speed matters more than precision
Pages(vec)Only scan specific 1-indexed page numbersWhen the caller knows which pages to check

Markdown output

The converter handles:

ElementHow it’s detected
Headings (H1-H4)Font size tiers relative to body text, with 0.5pt clustering
Bold/italicFont name patterns (Bold, Italic, Oblique)
Bullet lists, -, *, , , prefixes
Numbered lists1., 1), (1) patterns
Letter listsa., a), (a) patterns
Code blocksMonospace fonts (Courier, Consolas, Monaco, Menlo, Fira Code, JetBrains Mono) and keyword detection
TablesRectangle-based detection from PDF drawing ops + heuristic detection from text alignment
Financial tablesToken splitting for consolidated numeric values
Captions“Figure”, “Table”, “Source:” prefix detection
Sub/superscriptFont size and Y-offset relative to baseline
URLsConverted to Markdown links
HyphenationRejoins words broken across lines
Page numbersFiltered from output
Drop capsLarge initial letters merged with following text
Dot leadersTOC-style dots collapsed to “ … “

Use case: smart PDF routing

pdf-inspector was built for pipelines that process PDFs at scale. Instead of sending every PDF through OCR:

PDF arrives
  → pdf-inspector classifies it (~20ms)
  → TextBased + high confidence?
      YES → extract locally (~150ms), done
      NO  → send to OCR service (2-10s)

This saves cost and latency for the majority of PDFs that are already text-based (reports, papers, invoices, legal docs).

Debugging

See docs/debugging.md for RUST_LOG environment variable usage.

License

MIT

相似文章

@knowledgefxg: 实用开源小工具推荐:pdf-inspector 解决的是一个很实际的问题:并不是所有 PDF 都需要 OCR。 比方说你扔给它一个 PDF,它先判断这个 PDF 到底是什么类型——是正常的文字版(比如用 Word 导出的)、还是扫描版(图…

X AI KOLs Timeline

pdf-inspector 是一个开源的 Rust 库,用于智能分类 PDF 类型(文字版或扫描版),并提取文本和转换为 Markdown,避免不必要的 OCR,提高速度和节省成本。

firecrawl/pdf-inspector

GitHub Trending (daily)

Firecrawl 发布了 pdf-inspector,这是一个用于 PDF 分类和文本提取的快速 Rust 库,无需 OCR 即可将基于文本的 PDF 转换为 Markdown,并提供 Python、Node.js 和 WebAssembly 的绑定。

@Ryrenz: PDF 直接丢给大模型,表格乱码、公式崩溃、扫描件还认不出字。 这 5 个工具帮你把文档转成 AI 能读的干净格式。 1、Stirling-PDF — 本地 PDF 全能工具箱,85k star 合并、拆分、压缩、加水印、转格式,几十种操…

X AI KOLs Timeline

本文介绍了5个将文档转换为AI可读格式的开源工具:Stirling-PDF、MinerU、docling、marker和OCRmyPDF,分别针对PDF编辑、复杂文档转Markdown、文档结构化输出、高精度PDF转Markdown以及扫描件OCR处理。

@AIExplorerTim: 有人刚刚开发了一个工具,可以将 PDF 转换为 干净、结构化的 Markdown 速度达到 100 页/秒 不需要 GPU。 不需要 API 成本。 没有混乱的解析。 只有原始的、可用的数据。 它可以轻松处理的内容: • 表格 → 完美提…

X AI KOLs Timeline

OpenDataLoader 是一个开源工具,可将 PDF 转换为结构化的 Markdown 和 JSON,支持 100 页/秒的本地处理速度,无需 GPU 或 API 成本,专为 RAG 管道和 PDF 无障碍自动化设计。