Tag
Opus 5.5 scores 93.9% on ParseBench for table parsing in PDFs, outperforming Opus 5 and other models, but is expensive compared to LlamaParse.
DocJev is a free, open-source Python tool for classifying and splitting complex document packets, supporting PDF, DOCX, and PPTX with optional OCR.
DocJev is a lightning-fast open-source library for document classification and splitting using Jev, with OCR backends like liteparse and LlamaParse, claiming 6x speed improvements over GPT-5.6-luna with equivalent accuracy.
An evaluation of over 100 models for document parsing, covering frontier VLMs, open-weight VLMs, OCR tools, and open-source parsers, with results shared via parsebench.
Baidu's open-source Unlimited-OCR model processes multiple PDF pages simultaneously, outperforming baselines like DeepSeek-OCR and supporting local execution with community integrations.
Jina-ocr-v1 is a multimodal vision language model designed for advanced document intelligence, reading text from images in multiple languages.
PeekPaste is a native clipboard manager for Mac that provides easy access to clipboard history with features like search, OCR, and organization, all stored locally for privacy.
Thoughts for Mac is a menubar application that enables quick note-taking using text, images, and voice, with additional AI-powered features for audio transcription and text transformations.
Harbor is a privacy-focused note-taking app that serves as an Evernote alternative, featuring OCR and handwriting search across multiple devices with a price-locked subscription model.
The author seeks advice on choosing OCR models for a multi-document ERP system that handles mixed languages, handwritten text, and complex scan formats.
The article describes the process of republishing a book using AI tools for OCR and contract drafting, highlighting challenges with formatting and unique text elements.
LlamaIndex has significantly improved LlamaParse over the past three months, achieving 10-20% better accuracy on complex tables, charts, and grounding while maintaining costs below 0.4 cents per page, as measured against their ParseBench benchmark.
Jina-OCR-v1 is an efficient end-to-end document parsing model that uses speculative decoding and dense verifiable rewards to achieve high accuracy and speed on low-budget GPUs, scoring 91.14 on OmniDocBench v1.6 and 83.4 on olmOCR-Bench.
This paper evaluates the performance of open multimodal large language models on Khmer document VQA, finding that while models handle English and numeric content well, understanding native Khmer script remains challenging. The study uses a diagnostic subset from the KH-FUNSD collection and compares different model configurations.
HSK Manga is a free tool that simplifies manga dialogue to HSK levels for Chinese learners, using AI and OCR to make reading more accessible for beginners.
Cohere introduces Parse, a cost-effective vision language model for processing large volumes of enterprise documents into structured data, offering high performance and scalability for use cases like RAG and document indexing.
AnyDoc now includes OCR capabilities for reading scanned documents, offering fast processing times and a free hosted option via Firecrawl.
OpenCode Senses is a local vision plugin for OpenCode that provides advanced image understanding capabilities such as OCR, object detection, and color analysis, running privately on local hardware without API keys.
A Claude Code plugin that recovers export-blocked Kindle highlights by extracting them verbatim from Amazon's notebook and Kindle app data on macOS.
OCR It is a Chrome extension that extracts text from un-copyable documents using local OCR, designed to help users prepare content for LLMs like Claude or ChatGPT.