ocr

Tag

Cards List
#ocr

@jerryjliu0: Opus 5.5 is the best frontier model for parsing tables in PDFs. We ran it through ParseBench, and it scored 93.9%, a 7%…

X AI KOLs Timeline ↗ · yesterday Cached

Opus 5.5 scores 93.9% on ParseBench for table parsing in PDFs, outperforming Opus 5 and other models, but is expensive compared to LlamaParse.

0 favorites 0 likes
#ocr

@jerryjliu0: DocJev is the fastest way to classify and split complex document packets . (and the default is fully free and OSS!) I m…

X AI KOLs Timeline ↗ · 2d ago Cached

DocJev is a free, open-source Python tool for classifying and splitting complex document packets, supporting PDF, DOCX, and PPTX with optional OCR.

0 favorites 0 likes
#ocr

@jerryjliu0: Introducing DocJev - a lightning-fast OSS library for document classification and splitting with jev Give a document al…

X AI KOLs Timeline ↗ · 4d ago Cached

DocJev is a lightning-fast open-source library for document classification and splitting using Jev, with OCR backends like liteparse and LlamaParse, claiming 6x speed improvements over GPT-5.6-luna with equivalent accuracy.

0 favorites 0 likes
#ocr

@jerryjliu0: we've evaluated 100+ models on doc parsing from frontier VLMs, open-weight VLMs, OCR tools, and OSS parsers check out p…

X AI KOLs Timeline ↗ · 5d ago Cached

An evaluation of over 100 models for document parsing, covering frontier VLMs, open-weight VLMs, OCR tools, and open-source parsers, with results shared via parsebench.

0 favorites 0 likes
#ocr

@XAMTO_AI: Long PDFs are cut into single pages and then stitched back together, where cross-page tables and reading order are most…

X AI KOLs Timeline ↗ · 5d ago Cached

Baidu's open-source Unlimited-OCR model processes multiple PDF pages simultaneously, outperforming baselines like DeepSeek-OCR and supporting local execution with community integrations.

0 favorites 0 likes
#ocr

@HuggingModels: OCR just got a major upgrade. jina-ocr-v1 is a multimodal vision language model built for document intelligence. It rea…

X AI KOLs Timeline ↗ · 5d ago Cached

Jina-ocr-v1 is a multimodal vision language model designed for advanced document intelligence, reading text from images in multiple languages.

0 favorites 0 likes
#ocr

PeekPaste

Product Hunt ↗ · 2026-09-15 Cached

PeekPaste is a native clipboard manager for Mac that provides easy access to clipboard history with features like search, OCR, and organization, all stored locally for privacy.

0 favorites 0 likes
#ocr

Thoughts for Mac

Product Hunt ↗ · 2026-09-14 Cached

Thoughts for Mac is a menubar application that enables quick note-taking using text, images, and voice, with additional AI-powered features for audio transcription and text transformations.

0 favorites 0 likes
#ocr

Harbor

Product Hunt ↗ · 2026-09-14 Cached

Harbor is a privacy-focused note-taking app that serves as an Evernote alternative, featuring OCR and handwriting search across multiple devices with a price-locked subscription model.

0 favorites 0 likes
#ocr

Looking for advice on OCR model selection for a multi-document ERP system

Reddit r/AI_Agents ↗ · 2026-09-08

The author seeks advice on choosing OCR models for a multi-document ERP system that handles mixed languages, handwritten text, and complex scan formats.

0 favorites 0 likes
#ocr

Picolibrary: A Small Press

Hacker News Top ↗ · 2026-09-05 Cached

The article describes the process of republishing a book using AI tools for OCR and contract drafting, highlighting challenges with formatting and unique text elements.

0 favorites 0 likes
#ocr

@jerryjliu0: We've massively improved our document parsing capabilities across the board in the past ~3 months. Our LlamaParse cost-…

X AI KOLs Timeline ↗ · 2026-09-04 Cached

LlamaIndex has significantly improved LlamaParse over the past three months, achieving 10-20% better accuracy on complex tables, charts, and grounding while maintaining costs below 0.4 cents per page, as measured against their ParseBench benchmark.

0 favorites 0 likes
#ocr

Jina-OCR-v1: Efficient Document Parsing with Speculative Decoding and Dense Verifiable Rewards

arXiv cs.CL ↗ · 2026-09-04 Cached

Jina-OCR-v1 is an efficient end-to-end document parsing model that uses speculative decoding and dense verifiable rewards to achieve high accuracy and speed on low-budget GPUs, scoring 91.14 on OmniDocBench v1.6 and 83.4 on olmOCR-Bench.

0 favorites 0 likes
#ocr

Do MLLMs Really Understand Low-Resource Khmer Documents? A Pilot Study on Khmer Document VQA

arXiv cs.CL ↗ · 2026-09-01 Cached

This paper evaluates the performance of open multimodal large language models on Khmer document VQA, finding that while models handle English and numeric content well, understanding native Khmer script remains challenging. The study uses a diagnostic subset from the KH-FUNSD collection and compares different model configurations.

0 favorites 0 likes
#ocr

HSK Manga: Turning Manga into Beginner-Level Chinese

Lobsters Hottest ↗ · 2026-08-30 Cached

HSK Manga is a free tool that simplifies manga dialogue to HSK levels for Chinese learners, using AI and OCR to make reading more accessible for beginners.

0 favorites 0 likes
#ocr

Introducing Parse: Enterprise document intelligence at scale (5 minute read)

TLDR AI ↗ · 2026-08-28 Cached

Cohere introduces Parse, a cost-effective vision language model for processing large volumes of enterprise documents into structured data, offering high performance and scalability for use cases like RAG and document indexing.

0 favorites 0 likes
#ocr

@nickscamara_: introducing ocr in anydoc now your agents can read scanned docs for free → sub-5ms for non ocr → 190ms median per ocr p…

X AI KOLs Timeline ↗ · 2026-08-27 Cached

AnyDoc now includes OCR capabilities for reading scanned documents, offering fast processing times and a free hosted option via Firecrawl.

0 favorites 0 likes
#ocr

OpenCode Senses: The most advanced Local Vision Plugin for OpenCode That Actually Understands Images

Reddit r/LocalLLaMA ↗ · 2026-08-27

OpenCode Senses is a local vision plugin for OpenCode that provides advanced image understanding capabilities such as OCR, object detection, and color analysis, running privately on local hardware without API keys.

0 favorites 0 likes
#ocr

A Claude Code skill that recovers export-blocked Kindle highlights

Hacker News Top ↗ · 2026-08-24 Cached

A Claude Code plugin that recovers export-blocked Kindle highlights by extracting them verbatim from Amazon's notebook and Kindle app data on macOS.

0 favorites 0 likes
#ocr

OCR It – pull text out of un-copyable documents for your LLM

Hacker News Top ↗ · 2026-08-24 Cached

OCR It is a Chrome extension that extracts text from un-copyable documents using local OCR, designed to help users prepare content for LLMs like Claude or ChatGPT.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback