HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better
Summary
HunyuanOCR-1.5 is a lightweight end-to-end OCR vision-language model that improves efficiency via DFlash (6.37x inference speedup) and capability via Agentic Data Flow, achieving top-tier performance on document parsing, OCR, and multilingual tasks.
View Cached Full Text
Cached at: 07/09/26, 07:53 AM
Paper page - HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better
Source: https://huggingface.co/papers/2607.04884 Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
HunyuanOCR-1.5 is a lightweight end-to-end vision-language model that enhances OCR capabilities through improved efficiency via DFlash and enhanced capability via Agentic Data Flow, achieving fast inference and broad task coverage.
We present HunyuanOCR-1.5, a lightweight end-to-endOCR-specializedvision-language model. HunyuanOCRunifies document parsing, text spotting, information extraction, text-image translation, and multi-image document understanding within a single end-to-end VLM. Building upon the lightweight architecture of HunyuanOCR-1.0, HunyuanOCR-1.5 does not redesign the backbone, but systematically improves both efficiency and capability. For efficiency, we adaptDFlashtoOCRdecoding, significantly reducing the latency of long structured outputs such as dense documents, tables, and formulas while preserving output distribution. Powered byDFlash, HunyuanOCR-1.5 achieves a 6.37x Transformerinference speedupand a 2.14x speedup under vLLM, delivering the fastest inference among lightweightOCRVLMs. For capability, we proposeAgentic Data Flow, an agent-driven data construction system that transforms model weaknesses into executable data requirements and autonomously performs material search, quality verification, and pipeline development. It substantially improves long-tail capabilities in ancient-scriptOCR, fine-grained chart and table parsing, multi-image text-centric QA, low-resource multilingual parsing, and document hallucination evaluation. HunyuanOCR-1.5 ranks among the top-tier end-to-endOCRsolutions on OmniDocBench v1.6 while achieving new performance milestones across these long-tail tasks. Combined with an upgradedpretrainingandpost-trainingrecipe, HunyuanOCR-1.5 further extends its capability in high-resolution, long-context, and multi-task scenarios. Experiments demonstrate faster inference, broaderOCRcapability coverage, and the deployment advantages of a lightweightend-to-end model. We will release the model weights and training code to support future research and real-worldOCRapplications.
View arXiv pageView PDFAdd to collection
Get this paper in your agent:
hf papers read 2607\.04884
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper2
#### tencent/HunyuanOCR Image-Text-to-Text• 1B• Updated1 day ago • 505k • 764
#### Evan-613/HunyuanOCR Image-Text-to-Text• 1B• Updatedabout 5 hours ago
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2607.04884 in a dataset README.md to link it from this page.
Spaces citing this paper14
Browse 14 spaces citing this paper## Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
@techNmak: A lightweight VLM that beats the giants at OCR. (1.7B parameters, SOTA on OmniDocBench) dots. ocr is a new multilingual…
dots.ocr is a new lightweight 1.7B parameter multilingual vision-language model that achieves state-of-the-art performance on OmniDocBench, outperforming much larger models (72B+) at document parsing and OCR tasks.
@PaddlePaddle: PP-OCRv6 Tech Deep Dive Ep.1: In the Era of Large Models, Why Does Lightweight OCR Still Have Irreplaceable Value? — PP…
PP-OCRv6 is a lightweight OCR model (34.5M parameters) that challenges large VLMs with its MetaFormer architecture, offering efficient text detection and recognition across multiple deployment scenarios.
PP-OCRv6 on Hugging Face: 50-Language OCR from 1.5M to 34.5M Parameters
PP-OCRv6 is the latest generation of PaddleOCR's universal OCR model family, offering three tiers from 1.5M to 34.5M parameters, supporting 50 languages, and achieving significant accuracy improvements over previous versions.
baidu/Unlimited-OCR
Baidu releases Unlimited-OCR, a new model for one-shot long-horizon document parsing, building on Deepseek-OCR. It supports single image and multi-page/PDF parsing via Hugging Face Transformers and SGLang.
MonkeyOCRv2: A Visual-Text Foundation Model for Document AI
MonkeyOCRv2 is a visual-text foundation model for document AI, pretrained on a large corpus of 113 million images across 17 languages using joint image-to-text generation and pixel-level document reconstruction. It achieves state-of-the-art results on document parsing and understanding tasks, outperforming previous models with a smaller vision encoder.