document-ai

Tag

Cards List
#document-ai

@DataScienceDojo: A Chinese company just open-sourced an πŽπ‚π‘ that fixes something most AI-powered OCR tools quietly struggle with: the…

X AI KOLs Timeline β†— Β· 3d ago Cached

Unlimited-OCR, a new open-source OCR model from a Chinese company, solves the memory growth issue common in AI OCR tools by keeping memory usage flat regardless of document length, enabling single-pass reading of dozens of pages at 32K context. It's MIT-licensed, 3B parameters, multilingual, and already popular on GitHub.

0 favorites 0 likes
#document-ai

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI

Hugging Face Daily Papers β†— Β· 2026-07-13 Cached

MonkeyOCRv2 is a visual-text foundation model for document AI, pretrained on a large corpus of 113 million images across 17 languages using joint image-to-text generation and pixel-level document reconstruction. It achieves state-of-the-art results on document parsing and understanding tasks, outperforming previous models with a smaller vision encoder.

0 favorites 0 likes
#document-ai

datalab-to/lift

Hugging Face Models Trending β†— Β· 2026-06-19 Cached

Datalab releases lift, a model that extracts structured JSON from PDFs and images using schema-constrained decoding, with local and hosted inference options.

0 favorites 0 likes
#document-ai

Operationalizing Document AI: A Microservice Architecture for OCR and LLM Pipelines in Production

arXiv cs.AI β†— Β· 2026-05-20 Cached

This paper presents a microservice architecture for production document AI pipelines that combine classification, OCR, and LLM extraction, sharing design decisions and batch profiling insights that reveal OCR, not LLM parsing, dominates latency.

0 favorites 0 likes
#document-ai

PaddlePaddle/PaddleOCR

GitHub Trending (daily) β†— Β· 2026-06-05

PaddleOCR is a powerful, lightweight OCR toolkit that converts PDFs and images into structured data for AI applications, supporting 100+ languages and designed to bridge documents with LLMs.

0 favorites 0 likes
← Back to home

Submit Feedback