ocr

Tag

Cards List
#ocr

@MaximeRivest: This device is very impressive. its got microphone, speaker support, text to speech not, ocr, a chat gpt 5.2 chat and i…

X AI KOLs Following · yesterday Cached

A user shares hands-on impressions of a thin, light Android-based device featuring microphone, speaker, OCR, text-to-speech, and a ChatGPT 5.2 chat, noting the only clear drawback is its lack of color.

0 favorites 0 likes
#ocr

nvidia/NVIDIA-Nemotron-Parse-2.0 · Hugging Face

Reddit r/LocalLLaMA · 2d ago Cached

NVIDIA releases Nemotron Parse 2.0, a document image parsing model that converts scanned PDFs and images into structured text with layout, bounding boxes, and reading order, adding multilingual OCR improvements and chart-aware parsing.

0 favorites 0 likes
#ocr

I compared even more parsers on 14 PDF-parsing capabilities using different types

Reddit r/LocalLLaMA · 2d ago

A benchmark comparing 8 PDF parsers across 14 capabilities, finding Chandra the most accurate while noting trade-offs like speed; LightOnOCR-1B impresses for its size but hallucinates on illegible text.

0 favorites 0 likes
#ocr

Xberg: a local "read any document" tool for agents

Reddit r/AI_Agents · 5d ago

Xberg is a local content intelligence framework (Rust core, MIT) that extracts text from 101 document formats via MCP server or CLI, reconstructs reading order and tables, and chunks for context windows—all on-device for AI agents.

0 favorites 0 likes
#ocr

I compared MinerU, Granite-Docling, and PaddleOCR-VL on 12 PDF-parsing capabilities using 6 document types

Reddit r/LocalLLaMA · 6d ago

A developer benchmarks three PDF-parsing models (MinerU, Granite-Docling, PaddleOCR-VL) across six document types and twelve capabilities, finding MinerU drops footers unless its markdown is rebuilt while Granite-Docling outputs cleaner native markdown tables.

0 favorites 0 likes
#ocr

@DanKornas: Dicklesworthstone/llm_aided_ocr OCR for scanned PDFs with LLM-powered text correction GitHub: Archive:

X AI KOLs Timeline · 6d ago Cached

A GitHub project, llm_aided_ocr, performs OCR on scanned PDFs and uses LLMs to correct and improve the extracted text.

0 favorites 0 likes
#ocr

Make Everything to Markdown

Hacker News Top · 2026-07-31 Cached

Klartext is a free, privacy-focused web tool that converts PDFs, scans, photos, Word, Excel, and PowerPoint files into Markdown (plus JSON) entirely on its own server, with automatic deletion after 24 hours and no external AI/OCR services.

0 favorites 0 likes
#ocr

space ocr

Product Hunt · 2026-07-31

Space OCR is an OCR tool that checks its own answers, available as an app or an API.

0 favorites 0 likes
#ocr

IDP AutoOpt: Agent-Driven Optimization of Document Processing Pipeline Configurations

arXiv cs.AI · 2026-07-31 Cached

This paper presents IDP AutoOpt, an autonomous LLM agent that optimizes intelligent document processing pipeline configurations, matching or exceeding human-expert accuracy at lower cost and reducing configuration time from weeks to under two hours.

0 favorites 0 likes
#ocr

@PrajwalTomar_: A fully offline AI just read over 4,000 pages of declassified UFO files and answered questions about them with citation…

X AI KOLs Timeline · 2026-07-30 Cached

A fully offline AI reads over 4,000 pages of declassified UFO files using OCR and vector database, providing cited answers locally without cloud or API keys.

0 favorites 0 likes
#ocr

If your PDF extractor returns an empty string, the file probably isn't broken

Reddit r/AI_Agents · 2026-07-30

A developer describes building a PDF triage tool that inspects the PDF structure to determine if it contains a text layer or is made of scanned images, preventing silent empty extraction failures, and shares pitfalls around false text detection and phantom tables.

0 favorites 0 likes
#ocr

Do VLMs Read or Rewrite? On Transcription Faithfulness in Vision-Language Models

arXiv cs.CL · 2026-07-27 Cached

This paper reveals that Vision-Language Models often rewrite rather than faithfully transcribe text when encountering perturbations like typos or visual artifacts, introducing the FaithC4 benchmark to evaluate this behavior across multiple models and languages.

0 favorites 0 likes
#ocr

@jerryjliu0: We just released a LiteParse feature that allows image-to-pdf conversion to be natively handled in Rust, removing depen…

X AI KOLs Following · 2026-07-25 Cached

LiteParse is a fast, lightweight, open-source PDF parsing tool written in Rust, supporting image-to-PDF conversion natively and providing spatial text parsing with bounding boxes, available via multiple packages (Rust, Node.js, Python, WASM).

0 favorites 0 likes
#ocr

@vista8: Baidu's technical strength is still impressive. Recently, this Unlimited OCR has even caught the attention of Yann LeCun! Unlimited OCR proposes Reference Sliding Window Attention (R-SWA)...

X AI KOLs Timeline · 2026-07-22 Cached

Baidu's open-source Unlimited OCR model proposes the Reference Sliding Window Attention (R-SWA) mechanism, achieving continuous parsing of dozens of pages with 3 billion parameters, gaining high attention on GitHub and HuggingFace.

0 favorites 0 likes
#ocr

@VikParuchuri: Marker is powered by the new Surya OCR 2 model (https://github.com/datalab-to/surya…), which is the highest-scoring mod…

X AI KOLs Timeline · 2026-07-21 Cached

Datalab announces Surya OCR 2, a 650M parameter OCR model that achieves top accuracy under 1B parameters on olmOCR benchmarks, with high speed and multilingual support.

0 favorites 0 likes
#ocr

@VikParuchuri: Get marker here - https://github.com/datalab-to/marker… . Quickstart: pip install marker-pdf, then marker /path/to/fold…

X AI KOLs Timeline · 2026-07-21 Cached

Marker is an open-source tool that converts PDFs, images, and other document formats to markdown, JSON, chunks, and HTML quickly and accurately, with optional LLM enhancement.

0 favorites 0 likes
#ocr

DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth

arXiv cs.LG · 2026-07-21 Cached

DocOCR-Eval proposes an annotation-free framework that uses a correction and ranking strategy to evaluate and select OCR tools without ground truth labels, showing that aggregating multiple multimodal large language models improves alignment with human rankings.

0 favorites 0 likes
#ocr

@denziideng: Say goodbye to manual screenshot translation! overlay-translator: one-tap floating window real-time translation, boosting efficiency in games and manga. Stuck by language when playing Japanese games or reading manga? Want real-time understanding of dialogue and UI without rooting your phone? Manual screenshot translation is cumbersome and breaks immersion. Today I'm sharing a game-changer: overlay-t…

X AI KOLs Timeline · 2026-07-21 Cached

Overlay-translator is an open-source Android real-time screen translation tool that requires no root. It overlays Chinese translations onto the original screen via a floating window, supporting games, manga, visual novels, and more. It integrates multiple OCR and translation engines and works offline.

0 favorites 0 likes
#ocr

@DataScienceDojo: A Chinese company just open-sourced an 𝐎𝐂𝐑 that fixes something most AI-powered OCR tools quietly struggle with: the…

X AI KOLs Timeline · 2026-07-20 Cached

Unlimited-OCR, a new open-source OCR model from a Chinese company, solves the memory growth issue common in AI OCR tools by keeping memory usage flat regardless of document length, enabling single-pass reading of dozens of pages at 32K context. It's MIT-licensed, 3B parameters, multilingual, and already popular on GitHub.

0 favorites 0 likes
#ocr

@S0N_IA: | CHINA HAS ENDED THE OCR BUSINESS A 3 billion parameter model the size of a peanut can read a full 100-page PDF in one…

X AI KOLs Timeline · 2026-07-20 Cached

Baidu releases Unlimited-OCR, a 3B parameter open-source model that reads full 100-page PDFs in one go with a 32K context window, achieving 93% accuracy and running locally. It has 1.9 million Hugging Face downloads but little mainstream attention.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback