Tag
This paper introduces Multi-Split Boundary Decision (MSBD) to reduce inference costs in zero-shot page stream segmentation using large language models, demonstrating improved efficiency while maintaining accuracy for appropriate window sizes.
PaperPod is a service that converts PDF documents into podcast episodes, generating a 20-minute two-host episode in about one minute without requiring GPU or studio equipment.
The learn-from-materials Agent Skill converts PDFs, EPUBs, Word docs, and PPTs into interactive learning websites with source tracing and knowledge base building.
Typst, a typesetting system positioned as a LaTeX replacement, has released version 0.15 with new features including support for variable fonts, MathML, and multiple bibliographies.
This paper introduces a comprehensive benchmark for evaluating LLMs in key-value extraction from documents under OCR noise, revealing substantial performance degradation and emphasizing the need for joint optimization of OCR quality and LLM reasoning.
The paper introduces evidence-aligned local composition, a method for restoring corrupted documents by inferring soft weightings over frozen domain experts from the marginal evidence of the corrupted observation, achieving high accuracy in tracking expert regions without labeled data.
TextGen is a tool that enables users to run powerful AI models locally on their computers with full privacy, offline access, and support for text, images, and documents through a simple download and setup.
Weaviate introduces a method to search PDFs without text extraction by embedding each page as an image using late-interaction multi-vector retrieval, demonstrated on NVIDIA investor decks.
The article outlines a multi-gate architecture for a high-accuracy semantic evidence and RAG system for financial documents, emphasizing traceability, reconciliation, and hybrid retrieval, and seeks feedback on its design.
The Connected Stack Conference in San Francisco will feature AI industry leaders from companies like Anthropic and LlamaIndex, focusing on enterprise AI and document infrastructure for AI agents.
The article discusses a two-pass document processing trend for AI agents, where a fast OSS pass enables efficient retrieval and a VLM-based pass ensures accuracy, promoting Llama Index's LiteParse and LlamaParse tools to enhance cost and performance.
Jerry Liu introduces ExtractBench and highlights the 'agentic plus' extractor in LlamaParse for handling massive volumes of fields in long documents, with benchmark results available on ExtractBench.
Jerry Liu promotes LlamaParse and LlamaAgents for large-scale document extraction, emphasizing LLM evals and hillclimbing for accuracy and cost. He also connects FDE work with evals and RL environments.
RAG-Anything is a multimodal document-processing RAG system built on LightRAG that parses documents, constructs a multimodal knowledge graph, and uses hybrid vector-graph retrieval to answer queries.
Xberg v1 is released as the successor to Kreuzberg, a high-performance content intelligence framework supporting 101 document formats, 367 code/data types, audio/video transcription, and URL ingestion, with pure-Rust PDF and OCR backends, multiple language bindings, and mobile/WASM support.
This paper presents IDP AutoOpt, an autonomous LLM agent that optimizes intelligent document processing pipeline configurations, matching or exceeding human-expert accuracy at lower cost and reducing configuration time from weeks to under two hours.
LlamaIndex announced an official batch parsing experience for LlamaParse, allowing users to parse up to 10,000 files at once through a dedicated UI with batch auditing and failure inspection, removing the need for custom async scripts.
A fully offline AI reads over 4,000 pages of declassified UFO files using OCR and vector database, providing cited answers locally without cloud or API keys.
Introduces DocAnnot, a framework that uses a large vision-language model, OCR, and a spatially informed contextual matching algorithm to automatically generate training datasets for key information extraction from documents, reducing manual annotation effort. Evaluated on CORD and SROIE benchmarks, it achieves reasonable F1-scores and enables efficient human verification.
A team used AI to automate a manual document sorting process, reducing labor from 50-70 hours to 3-5 hours per month by grouping scanned pages into documents and generating PDFs.