Tag
Datalab's Marker v2 is an open-source document parsing tool that handles PDFs, images, DOCX, and PPTX, outputting clean Markdown and supporting over 90 languages for batch organization, knowledge bases, and AI document processing.
A comprehensive evaluation of 16 frontier vision language models on document parsing shows that Opus 5.5 offers the best performance relative to its price, especially for tables, while GPT-6 Luna is good for cost-effective parsing. For large-scale use, dedicated tools like LlamaParse are recommended, but Opus 5.5 leads for in-app parsing.
An evaluation of over 100 models for document parsing, covering frontier VLMs, open-weight VLMs, OCR tools, and open-source parsers, with results shared via parsebench.
WeVisDoc is an end-to-end document parser fine-tuned from Qwen3-VL models that converts page images to structured Markdown with LaTeX and HTML, achieving top performance on benchmarks like OmniDocBench v1.6.
WeVisDoc is a two-stage data-centric framework for robust end-to-end document parsing that expands coverage and uses targeted diagnostics to improve performance, achieving state-of-the-art results on benchmarks.
The post highlights the difficulty of processing complex real-life documents and promotes LlamaParse as a tool that uses multiple AI agents for efficient, scalable, and multimodal document parsing.
LlamaParse introduces high-effort confidence scores to enhance AI document parsing accuracy, enabling human-in-the-loop review and automated fallback for sensitive processes.
LlamaIndex has trained new AI models for parsing checkboxes from forms of any size and shape, converting them into structured JSON for programmatic use in applications like insurance and tax processing.
LlamaIndex has significantly improved LlamaParse over the past three months, achieving 10-20% better accuracy on complex tables, charts, and grounding while maintaining costs below 0.4 cents per page, as measured against their ParseBench benchmark.
Jina-OCR-v1 is an efficient end-to-end document parsing model that uses speculative decoding and dense verifiable rewards to achieve high accuracy and speed on low-budget GPUs, scoring 91.14 on OmniDocBench v1.6 and 83.4 on olmOCR-Bench.
The official LlamaParse connector has been launched on Claude, providing enhanced document parsing for complex tables, forms, and visual elements with high accuracy and cost-efficiency at scale.
Cohere introduces Parse, a cost-effective vision language model for processing large volumes of enterprise documents into structured data, offering high performance and scalability for use cases like RAG and document indexing.
LlamaParse has introduced native agentic spreadsheet extraction using a tuned model and harness for schema-guided extraction, enabling conversion of dense sheets like balance sheets into clean structured fields.
Datalab has released Marker v2, an open-source document parsing pipeline that efficiently converts PDFs, images, DOCX, and PPTX files to markdown with support for over 90 languages and high performance on GPUs.
LlamaParse has added revision tracking to extract Word-style tracked changes and reviewer comments as structured metadata, enabling AI agents to access a document's full revision history for enhanced collaboration in industries like legal and finance.
LlamaParse now handles revision tracking in documents, providing clean markdown of the final state and structured data for edits, deletions, and comments, addressing issues where parsers misinterpret tracked changes.
TeleOCR is a lightweight open-source Vision-Language Model for document parsing, achieving state-of-the-art performance on both digital and camera-captured documents using techniques like Multi-node Consensus Voting.
TeleOCR is a lightweight open-source Vision-Language Model designed for document parsing, unifying digital and camera-captured documents with state-of-the-art performance on benchmarks.
NaviDC-OCR is a unified vision-language framework that improves document parsing accuracy through deformation-aware learning and adaptive sampling, achieving state-of-the-art results on multiple benchmarks.
LlamaIndex introduces a new LlamaParse feature that automatically extracts complex form fields into structured JSON without requiring a predefined schema, simplifying document form processing.