pdf-processing

Tag

Cards List
#pdf-processing

@GYLQ520: PixelRAG is an open-source project from UC Berkeley SkyLab and other teams, ready to play. Its approach is straightforward, no longer relying on HTML or text parsing. Instead, it directly captures web pages and PDFs as screenshots and uses visual indexing for RAG retrieval. Everything that HTML parsing loses, such as tables, charts, and layouts, is preserved…

X AI KOLs Timeline · 2026-08-21 Cached

PixelRAG is an open-source project from UC Berkeley SkyLab and other teams. By capturing web pages and PDFs as screenshots and using visual indexing for retrieval, it improves RAG accuracy and significantly reduces token costs for AI Agents.

0 favorites 0 likes
#pdf-processing

@googledevs: Static PDF interactive web app, all in one prompt! We used Gemini 3.7 Flash in @GoogleAIStudio to transform a dense PDF…

X AI KOLs Timeline · 2026-08-18 Cached

Google Developers showcased using Gemini 3.7 Flash in Google AI Studio to transform a PDF annual report into an interactive web app, highlighting the model's data extraction and UI generation capabilities in a single prompt.

0 favorites 0 likes
#pdf-processing

@googleaidevs: We fed Gemini 3.7 Flash hundreds of PDFs from a Victorian book of botanical illustrations and tasked it with extracting…

X AI KOLs Timeline · 2026-08-17 Cached

Google AI Devs demonstrated Gemini 3.7 Flash by using it to extract and classify plants from historical botanical PDFs with an interactive visualization.

0 favorites 0 likes
#pdf-processing

@sunmer575399: Stirling-PDF really has automation figured out — 89.0k stars, no joke. 50+ PDF tools packed into one open-source platform: merge, split, compress, OCR, watermarks — basically covers all everyday needs. The killer feature is the automation pipeline: drag and drop nodes to batch-process PDFs without writing a single...

X AI KOLs Timeline · 2026-08-08 Cached

Stirling-PDF is an open-source PDF processing platform offering 50+ tools (merge, split, OCR, compression, etc.), supporting browser use, Docker self-hosting, and a REST API, as well as a no-code automation pipeline. The project has 89k+ stars on GitHub and can be fully deployed locally to protect privacy.

0 favorites 0 likes
#pdf-processing

@sitinme: Not just "have AI summarize a book", but go further: turning a book or a document package into a Skill that an AI Agent can repeatedly call. This idea is worth discussing. Previously, after buying and reading a book, when I later wanted to find a certain knowledge point, I couldn't find it after flipping through for a long time; asking AI might make things up; throwing the entire PD…

X AI KOLs Timeline · 2026-06-05 Cached

Introduces a tool called book-to-skill that converts books or document packages into AI Agent callable Skills. It supports PDF and other formats, generates SKILL.md and chapter indexes, avoiding loading the full context at once.

0 favorites 0 likes
#pdf-processing

Vision-capable LLMs vs. OCR for long-document (including charts, images, tables, etc.) QA

Reddit r/artificial · 2026-05-24

A benchmark comparing vision-capable LLMs (native PDF reading) against OCR-based pipelines on 30 long, image-heavy PDFs finds that OCR with layout extraction still outperforms vision models on chart/table-heavy pages and has a 0% failure rate vs. 7% for native PDF, though the sample size is small and many gaps are within noise.

0 favorites 0 likes
#pdf-processing

@wsl8297: If you have a bunch of PDFs, documents, project materials to feed to AI, Synthadoc is a direction worth looking at. GitHub: https://github.com/axoviq-ai/synthadoc… It compiles raw materials into a structured wiki at ingestion time, automatically...

X AI KOLs Timeline · 2026-05-23 Cached

Synthadoc is an open-source tool that compiles PDFs, documents, and other project materials into a structured local Markdown wiki, automatically establishing cross-references and detecting contradictions. It is suitable for personal or small teams for offline knowledge management.

0 favorites 0 likes
#pdf-processing

@tom_doerr: Turns technical books into Claude Code skills https://github.com/virgiliojr94/book-to-skill…

X AI KOLs Timeline · 2026-05-22 Cached

book-to-skill converts technical books into structured skills for Claude Code, allowing on-demand reference and eliminating hallucination.

0 favorites 0 likes
#pdf-processing

olmOCR: Unlocking Trillions of Tokens in PDFs with Vision Language Models

Papers with Code Trending · 2025-02-25 Cached

olmOCR is an open-source toolkit using a fine-tuned vision language model to extract clean text from PDFs while preserving structure, optimized for large-scale batch processing.

0 favorites 0 likes
#pdf-processing

opendatalab/MinerU

GitHub Trending (daily) · 2026-06-25 Cached

MinerU is an open-source tool by OpenDataLab for extracting data from PDFs and documents.

0 favorites 0 likes
← Back to home

Submit Feedback