document-parsing

Tag

Cards List
#document-parsing

TRACE: Target-Aware Retrieval, Attributed Evidence, and Contract-Constrained Extraction for LitTraceQA

arXiv cs.CL ↗ · 8h ago Cached

TRACE introduces target-aware retrieval, attributed evidence localization, and contract-constrained extraction to close the 'grounding contract gap' in scientific LitTraceQA, scoring 0.760613 on the official 71-question test set with strong paper F1 (0.9728) but weaker table extraction performance.

0 favorites 0 likes
#document-parsing

Improving OCR Faithfulness via Gated and Attenuated On-Policy Distillation

arXiv cs.AI ↗ · 8h ago Cached

The paper introduces GAD-RL, a method that adaptively gates and attenuates on-policy distillation supervision during joint post-training of vision-language models to improve OCR transcription faithfulness, achieving 59.92% Micro Recall on CHAOS-Bench on Qwen3.5-2B, surpassing GRPO by 8.45 points.

0 favorites 0 likes
#document-parsing

Resolving the Missing Financial Data Crisis: A Generative AI Pipeline for SEC 10-K Extraction

arXiv cs.CL ↗ · yesterday Cached

该研究评估了 Llama-3 8B、Qwen-2.5 14B 和 Llama-3.3 70B 等 LLM 在从 SEC 10-K 文件中提取财务数据时的表现,发现将模型规模与文档结构复杂度对齐可在零样本设置下获得高准确率,从而缓解量化金融数据集中超过 70% 公司面临的缺失数据偏差问题。

0 favorites 0 likes
#document-parsing

@llama_index: Document parsing requires a lot of on-the-fly decision making. We explored the early approaches to Jev and Jev-like mod…

X AI KOLs Timeline ↗ · 2d ago Cached

The article compares AI models like Jev and Quinn for document parsing tasks such as language detection and classification, evaluating their accuracy, speed, and applicability.

0 favorites 0 likes
#document-parsing

@x5cnhp: Spending hundreds of pages gnawing through PDFs until your wrist aches? This kind of work can actually be completely ha…

X AI KOLs Timeline ↗ · 4d ago Cached

Datalab's Marker v2 is an open-source document parsing tool that handles PDFs, images, DOCX, and PPTX, outputting clean Markdown and supporting over 90 languages for batch organization, knowledge bases, and AI document processing.

0 favorites 0 likes
#document-parsing

@jerryjliu0: We comprehensively evaluated 16 recent frontier VLMs - including Opus 5.5 and GPT-6 Sol/Luna - on whether higher effort…

X AI KOLs Timeline ↗ · 5d ago Cached

A comprehensive evaluation of 16 frontier vision language models on document parsing shows that Opus 5.5 offers the best performance relative to its price, especially for tables, while GPT-6 Luna is good for cost-effective parsing. For large-scale use, dedicated tools like LlamaParse are recommended, but Opus 5.5 leads for in-app parsing.

0 favorites 0 likes
#document-parsing

@jerryjliu0: we've evaluated 100+ models on doc parsing from frontier VLMs, open-weight VLMs, OCR tools, and OSS parsers check out p…

X AI KOLs Timeline ↗ · 2026-09-19 Cached

An evaluation of over 100 models for document parsing, covering frontier VLMs, open-weight VLMs, OCR tools, and open-source parsers, with results shared via parsebench.

0 favorites 0 likes
#document-parsing

WeVisDoc from tencent

Reddit r/LocalLLaMA ↗ · 2026-09-17

WeVisDoc is an end-to-end document parser fine-tuned from Qwen3-VL models that converts page images to structured Markdown with LaTeX and HTML, achieving top performance on benchmarks like OmniDocBench v1.6.

0 favorites 0 likes
#document-parsing

WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing

Hugging Face Daily Papers ↗ · 2026-09-17 Cached

WeVisDoc is a two-stage data-centric framework for robust end-to-end document parsing that expands coverage and uses targeted diagnostics to improve performance, achieving state-of-the-art results on benchmarks.

0 favorites 0 likes
#document-parsing

@svpino: Building an agent that works with real-life documents is still crazy hard. All of you are talking about AGI, and this i…

X AI KOLs Timeline ↗ · 2026-09-14 Cached

The post highlights the difficulty of processing complex real-life documents and promotes LlamaParse as a tool that uses multiple AI agents for efficient, scalable, and multimodal document parsing.

0 favorites 0 likes
#document-parsing

@jerryjliu0: One of the main issues with AI document parsing is that because no solution is 100% accuracy, it's hard to tell if a gi…

X AI KOLs Timeline ↗ · 2026-09-11 Cached

LlamaParse introduces high-effort confidence scores to enhance AI document parsing accuracy, enabling human-in-the-loop review and automated fallback for sensitive processes.

0 favorites 0 likes
#document-parsing

@jerryjliu0: We've trained new models for checkbox parsing This lets you convert checkboxes of any size and shape, filled or unfille…

X AI KOLs Timeline ↗ · 2026-09-10 Cached

LlamaIndex has trained new AI models for parsing checkboxes from forms of any size and shape, converting them into structured JSON for programmatic use in applications like insurance and tax processing.

0 favorites 0 likes
#document-parsing

@jerryjliu0: We've massively improved our document parsing capabilities across the board in the past ~3 months. Our LlamaParse cost-…

X AI KOLs Timeline ↗ · 2026-09-04 Cached

LlamaIndex has significantly improved LlamaParse over the past three months, achieving 10-20% better accuracy on complex tables, charts, and grounding while maintaining costs below 0.4 cents per page, as measured against their ParseBench benchmark.

0 favorites 0 likes
#document-parsing

Jina-OCR-v1: Efficient Document Parsing with Speculative Decoding and Dense Verifiable Rewards

arXiv cs.CL ↗ · 2026-09-04 Cached

Jina-OCR-v1 is an efficient end-to-end document parsing model that uses speculative decoding and dense verifiable rewards to achieve high accuracy and speed on low-budget GPUs, scoring 91.14 on OmniDocBench v1.6 and 83.4 on olmOCR-Bench.

0 favorites 0 likes
#document-parsing

@jerryjliu0: We’re excited to launch the official LlamaParse connector on Claude! Yes Claude can already read PDFs, and it probably …

X AI KOLs Timeline ↗ · 2026-08-31 Cached

The official LlamaParse connector has been launched on Claude, providing enhanced document parsing for complex tables, forms, and visual elements with high accuracy and cost-efficiency at scale.

0 favorites 0 likes
#document-parsing

Introducing Parse: Enterprise document intelligence at scale (5 minute read)

TLDR AI ↗ · 2026-08-28 Cached

Cohere introduces Parse, a cost-effective vision language model for processing large volumes of enterprise documents into structured data, offering high performance and scalability for use cases like RAG and document indexing.

0 favorites 0 likes
#document-parsing

@jerryjliu0: We've introduced native, agentic spreadsheet extraction into LlamaParse. Spreadsheets are a wildly different format fro…

X AI KOLs Timeline ↗ · 2026-08-27 Cached

LlamaParse has introduced native agentic spreadsheet extraction using a tuned model and harness for schema-guided extraction, enabling conversion of dense sheets like balance sheets into clean structured fields.

0 favorites 0 likes
#document-parsing

@akshay_pachaar: Turn any PDF, image, DOCX, and PPTX into clean markdown. Parsing one PDF is easy. Parsing millions of them is where it …

X AI KOLs Timeline ↗ · 2026-08-23 Cached

Datalab has released Marker v2, an open-source document parsing pipeline that efficiently converts PDFs, images, DOCX, and PPTX files to markdown with support for over 90 languages and high performance on GPUs.

0 favorites 0 likes
#document-parsing

@jerryjliu0: We built revision tracking in LlamaParse You can now extract Word-style tracked changes and reviewer comments as additi…

X AI KOLs Timeline ↗ · 2026-08-19 Cached

LlamaParse has added revision tracking to extract Word-style tracked changes and reviewer comments as structured metadata, enabling AI agents to access a document's full revision history for enhanced collaboration in industries like legal and finance.

0 favorites 0 likes
#document-parsing

@llama_index: Most parsers treat tracked changes as noise. Result: a deleted clause comes back as live text, and your pipeline reads …

X AI KOLs Timeline ↗ · 2026-08-18 Cached

LlamaParse now handles revision tracking in documents, providing clean markdown of the final state and structured data for edits, deletions, and comments, addressing issues where parsers misinterpret tracked changes.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback