document-processing

Tag

Cards List
#document-processing

We Used AI to Reduce a 50–70 Hour Manual Document Sorting Process to Around 3–5 Hours

Reddit r/ArtificialInteligence ↗ · 2026-07-28

A team used AI to automate a manual document sorting process, reducing labor from 50-70 hours to 3-5 hours per month by grouping scanned pages into documents and generating PDFs.

0 favorites 0 likes
#document-processing

@vista8: Baidu's technical strength is still impressive. Recently, this Unlimited OCR has even caught the attention of Yann LeCun! Unlimited OCR proposes Reference Sliding Window Attention (R-SWA)...

X AI KOLs Timeline ↗ · 2026-07-22 Cached

Baidu's open-source Unlimited OCR model proposes the Reference Sliding Window Attention (R-SWA) mechanism, achieving continuous parsing of dozens of pages with 3 billion parameters, gaining high attention on GitHub and HuggingFace.

0 favorites 0 likes
#document-processing

On-Device Field Extraction by Verify

Product Hunt ↗ · 2026-07-07

Verify announces on-device field extraction, enabling secure document data extraction even when offline.

0 favorites 0 likes
#document-processing

New PyMuPDF release, supports Markdown [N]

Reddit r/MachineLearning ↗ · 2026-07-01

PyMuPDF 1.28 adds first-class Markdown support, allowing PDF creation from Markdown text with CSS control.

0 favorites 0 likes
#document-processing

Launch HN: Parsewise (YC P25) – Reason Across Documents with an API

Hacker News Top ↗ · 2026-07-01 Cached

Parsewise is a YC-backed API that transforms unstructured documents into schema-compliant data with traceable lineage, enabling AI reasoning across documents.

0 favorites 0 likes
#document-processing

TurboOCR v3 — high-speed document OCR server (C++/CUDA), ~520 img/s on RTX 5090

Reddit r/LocalLLaMA ↗ · 2026-06-30

TurboOCR v3 is a self-hosted, high-speed OCR server that achieves ~520 images per second on an RTX 5090 using PP-OCRv6 models, with new structured parsing for tables and formulas.

0 favorites 0 likes
#document-processing

@MistralDevs: https://x.com/MistralDevs/status/2071625939444744521

X AI KOLs Timeline ↗ · 2026-06-29 Cached

Mistral Devs published a tutorial on building a medical document processing workflow using Mistral OCR, Agents, and Workflows, including a human-in-the-loop step for low-confidence classifications.

0 favorites 0 likes
#document-processing

@hasantoxr: I found the OCR tool built for the LLM era. It is called olmOCR. olmOCR takes PDFs, scans, PNGs, and JPEGs and turns th…

X AI KOLs Timeline ↗ · 2026-06-29 Cached

olmOCR is an open-source OCR tool from Ai2 that converts PDFs, scans, and images into clean Markdown, designed to prepare documents for LLM pipelines by preserving reading order and handling complex layouts.

0 favorites 0 likes
#document-processing

@itsclelia: The @n8n_io node for the LlamaParse Platform is now an officially verified community node Out of the box, you get acces…

X AI KOLs Following ↗ · 2026-06-26 Cached

LlamaIndex has released v5 and v6 of the LlamaParse Platform community node for n8n, now officially verified, providing document parsing, classification, splitting, extraction, and retrieval capabilities that can be used as tools for AI agents.

0 favorites 0 likes
#document-processing

@so_ainsight: Messy documents transform into structured knowledge with a single command. When feeding documents to AI, one quietly to…

X AI KOLs Timeline ↗ · 2026-06-26 Cached

Hyper-Extract is an Apache 2.0 open-source tool that converts unstructured documents into structured knowledge bases, supporting knowledge graphs, time-series data, and spatial information, enabling high-accuracy AI queries.

0 favorites 0 likes
#document-processing

@Ryrenz: Papers, contracts, PDFs — these open-source tools cover all document work: 1. opendatalab/MinerU (68.9k) — from Shanghai AI Lab, one-click PDF/document to markdown, excellent academic paper layout restoration. https://github.c…

X AI KOLs Timeline ↗ · 2026-06-25 Cached

This tweet summarizes 6 open-source tools covering PDF to markdown, document understanding, OCR, paper translation, and automatic literature review, aiming to streamline document workflows.

0 favorites 0 likes
#document-processing

@heynavtoor: A lawyer in Manhattan gets a 500-page contract. Every clause needs to be searchable. By hand: one week. An accountant i…

X AI KOLs Timeline ↗ · 2026-06-24 Cached

MinerU is a free, open-source tool that extracts text, tables, and equations from PDFs and scanned documents, supporting 109 languages and batch processing, saving hours of manual work.

0 favorites 0 likes
#document-processing

Docling vs Liteparse vs Mineru vs Unstructured for on-prem document processing for a university

Reddit r/LocalLLaMA ↗ · 2026-06-23

A comparison of on-prem document processing tools—Docling, Liteparse, Mineru, and Unstructured—for university use, evaluating their suitability for local deployment.

0 favorites 0 likes
#document-processing

@ErickSky: Baidu has just broken one of the biggest limitations of current OCR. Unlimited-OCR processes entire documents in a sing…

X AI KOLs Timeline ↗ · 2026-06-23 Cached

Baidu has released Unlimited-OCR, which processes entire documents in a single pass without chunking, overcoming a major limitation of current OCR technology.

0 favorites 0 likes
#document-processing

@VikParuchuri: We're open sourcing a 9B model that extracts structured data from documents at near-frontier performance. - 90.2% on ou…

X AI KOLs Following ↗ · 2026-06-19 Cached

Vik Paruchuri is open-sourcing a 9B model that extracts structured data from documents with near-frontier performance (90.2% on their benchmark, vs Gemini 3.5 Flash at 91.3%).

0 favorites 0 likes
#document-processing

@DataChaz: Messy documents in. Complex knowledge graphs out. One command line. If your pipeline simply compiles data into generic …

X AI KOLs Timeline ↗ · 2026-06-17 Cached

Hyper-Extract is an open-source framework that converts messy documents into typed knowledge structures, supporting multiple graph architectures like GraphRAG, LightRAG, and KG-Gen, with 10+ extraction engines and 80+ YAML templates for various domains.

0 favorites 0 likes
#document-processing

Typst 0.15 contains multitudes

Lobsters Hottest ↗ · 2026-06-15 Cached

Typst 0.15, a major release of the open-source typesetting system, introduces support for variable fonts, MathML export, multi-file output, multiple bibliographies, and multiple PDF standards, along with improved documentation and diagnostics.

0 favorites 0 likes
#document-processing

@TeksEdge: Need to OCR documents? PP-OCRv6 dropped — currently the best open-source OCR models you can download ◆︎ Fully Open Sour…

X AI KOLs Timeline ↗ · 2026-06-12 Cached

PP-OCRv6 is a new open-source OCR model series from Baidu's PaddleOCR, available in Tiny/Small/Medium sizes with excellent accuracy and speed, beating several commercial models.

0 favorites 0 likes
#document-processing

@DailyDoseOfDS_: Fine-tune DeepSeek-OCR on your own language! (100% local) Most vision models treat documents as massive sequences of to…

X AI KOLs Timeline ↗ · 2026-06-08 Cached

DeepSeek-OCR is a 3B vision model using context optical compression for efficient document processing. Fine-tuning it on Persian text using Unsloth achieved an 88.26% improvement in character error rate, all open-source and runnable on a single GPU.

0 favorites 0 likes
#document-processing

I made a small local model (llama3.2 3B) reliably extract structured JSON from documents - the hard part wasn't the model, it was everything around it

Reddit r/AI_Agents ↗ · 2026-06-05

A developer shares lessons from building a local document-to-JSON extractor using llama3.2 3B on Ollama, highlighting that deterministic post-processing and schema-constrained outputs matter more than model size, while seeking feedback on hallucination and context truncation issues with long documents.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback