Tag
TRACE introduces target-aware retrieval, attributed evidence localization, and contract-constrained extraction to close the 'grounding contract gap' in scientific LitTraceQA, scoring 0.760613 on the official 71-question test set with strong paper F1 (0.9728) but weaker table extraction performance.
The paper introduces GAD-RL, a method that adaptively gates and attenuates on-policy distillation supervision during joint post-training of vision-language models to improve OCR transcription faithfulness, achieving 59.92% Micro Recall on CHAOS-Bench on Qwen3.5-2B, surpassing GRPO by 8.45 points.
该研究评估了 Llama-3 8B、Qwen-2.5 14B 和 Llama-3.3 70B 等 LLM 在从 SEC 10-K 文件中提取财务数据时的表现,发现将模型规模与文档结构复杂度对齐可在零样本设置下获得高准确率,从而缓解量化金融数据集中超过 70% 公司面临的缺失数据偏差问题。
The article compares AI models like Jev and Quinn for document parsing tasks such as language detection and classification, evaluating their accuracy, speed, and applicability.
Datalab's Marker v2 is an open-source document parsing tool that handles PDFs, images, DOCX, and PPTX, outputting clean Markdown and supporting over 90 languages for batch organization, knowledge bases, and AI document processing.
A comprehensive evaluation of 16 frontier vision language models on document parsing shows that Opus 5.5 offers the best performance relative to its price, especially for tables, while GPT-6 Luna is good for cost-effective parsing. For large-scale use, dedicated tools like LlamaParse are recommended, but Opus 5.5 leads for in-app parsing.
An evaluation of over 100 models for document parsing, covering frontier VLMs, open-weight VLMs, OCR tools, and open-source parsers, with results shared via parsebench.
WeVisDoc is an end-to-end document parser fine-tuned from Qwen3-VL models that converts page images to structured Markdown with LaTeX and HTML, achieving top performance on benchmarks like OmniDocBench v1.6.
WeVisDoc is a two-stage data-centric framework for robust end-to-end document parsing that expands coverage and uses targeted diagnostics to improve performance, achieving state-of-the-art results on benchmarks.
The post highlights the difficulty of processing complex real-life documents and promotes LlamaParse as a tool that uses multiple AI agents for efficient, scalable, and multimodal document parsing.
LlamaParse introduces high-effort confidence scores to enhance AI document parsing accuracy, enabling human-in-the-loop review and automated fallback for sensitive processes.
LlamaIndex has trained new AI models for parsing checkboxes from forms of any size and shape, converting them into structured JSON for programmatic use in applications like insurance and tax processing.
LlamaIndex has significantly improved LlamaParse over the past three months, achieving 10-20% better accuracy on complex tables, charts, and grounding while maintaining costs below 0.4 cents per page, as measured against their ParseBench benchmark.
Jina-OCR-v1 is an efficient end-to-end document parsing model that uses speculative decoding and dense verifiable rewards to achieve high accuracy and speed on low-budget GPUs, scoring 91.14 on OmniDocBench v1.6 and 83.4 on olmOCR-Bench.
The official LlamaParse connector has been launched on Claude, providing enhanced document parsing for complex tables, forms, and visual elements with high accuracy and cost-efficiency at scale.
Cohere introduces Parse, a cost-effective vision language model for processing large volumes of enterprise documents into structured data, offering high performance and scalability for use cases like RAG and document indexing.
LlamaParse has introduced native agentic spreadsheet extraction using a tuned model and harness for schema-guided extraction, enabling conversion of dense sheets like balance sheets into clean structured fields.
Datalab has released Marker v2, an open-source document parsing pipeline that efficiently converts PDFs, images, DOCX, and PPTX files to markdown with support for over 90 languages and high performance on GPUs.
LlamaParse has added revision tracking to extract Word-style tracked changes and reviewer comments as structured metadata, enabling AI agents to access a document's full revision history for enhanced collaboration in industries like legal and finance.
LlamaParse now handles revision tracking in documents, providing clean markdown of the final state and structured data for edits, deletions, and comments, addressing issues where parsers misinterpret tracked changes.