document-layout-parsing

Tag

Cards List
#document-layout-parsing

dots.ocr: Multilingual Document Layout Parsing in a Single Vision-Language Model

Papers with Code Trending · 2025-12-02 Cached

This paper presents dots.ocr, a unified Vision-Language Model that jointly learns layout detection, text recognition, and relational understanding for multilingual document layout parsing. It achieves state-of-the-art results on OmniDocBench and introduces the XDocParse benchmark spanning 126 languages.

0 favorites 0 likes
← Back to home

Submit Feedback