@GoSailGlobal: Current OCR processes multi-page documents page by page. Every time you turn a page, memory is reset. Today, Baidu quietly open-sourced a model on GitHub and HuggingFace called Unlimited OCR, inspired by how humans copy books: - When copying a book, you don't reread hundreds of pages every time you write a word...

X AI KOLs Timeline Models

Summary

Baidu has open-sourced the Unlimited OCR model, which uses a Reference Sliding Window Attention (R-SWA) mechanism to parse documents up to 32K context in a single pass, eliminating the need for page-by-page inference.

Current OCR processes dozens of document pages page by page. Every page turn resets memory. Today, Baidu quietly open-sourced a model on GitHub and HuggingFace called Unlimited OCR, inspired by how humans copy books: - When copying a book, you don't reread the previous hundreds of pages every time you write a word. - You just glance at where you are nearby, and the rest is a kind of "soft forgetting". - Humans rely on this low-load continuous cognitive state to handle long-range tasks. Baidu turned this intuition into an attention mechanism called R-SWA (Reference Sliding Window Attention): Each token sees the entire image, but the output only looks back at the last 128 tokens. KV Cache stays constant, not growing with page count. The result: under 32K context, a single forward pass transcribes dozens of pages, turning a page-by-page for-loop into one continuous reading. A page-by-page for-loop is an engineering workaround. Continuous cognitive states are more like what intelligence should be. Baidu's approach lately has indeed been different. http://github.com/baidu/Unlimited-OCR… https://huggingface.co/baidu/Unlimited-OCR…
Original Article
View Cached Full Text

Cached at: 06/22/26, 09:41 AM

Unlimited OCR Works

Welcome the Era of One-shot Long-horizon Parsing.

Similar Articles

@geekbb: Baidu's open-source visual language model OCR project, upgraded from DeepSeek-OCR, focuses on one-shot parsing of extremely long documents. The model has two inference modes: 'gundam' mode for dense text in a single image, and 'base' mode for multi-page or PDF processing. https://github…

X AI KOLs Timeline

Baidu has open-sourced the visual language model Unlimited-OCR, upgraded from DeepSeek-OCR, supporting one-shot parsing of extremely long documents, offering two inference modes: gundam (dense text in a single image) and base (multi-page/PDF).

@CoderDaMing: China has open-sourced a peanut-sized OCR that can parse an entire 100-page PDF in one go. It's called 'Unlimited-OCR'. Only 3B parameters. Runs locally. Other OCR tools cut documents page by page, easily losing context. This one reads the entire document at once. → Single 'long-range' parse (32K context window...

X AI KOLs Timeline

China has open-sourced the OCR model Unlimited-OCR with only 3B parameters, which can parse an entire 100-page PDF in one go, supports local execution, achieves 93% accuracy, and is completely free and open-source.

@berryxia: Wow, this move directly poached DeepSeek's talent! Last night I saw this interesting OCR open-source model on HuggingFace and the fascinating story behind it. This OCR model is completely different from traditional ones! Its speed and accuracy are absolutely unbeatable~~ Let me start with some background, for those who are familiar…

X AI KOLs Timeline

Baidu has open-sourced the Unlimited OCR model, which uses the R-SWA attention mechanism to process hundreds of pages in a single pass without page splitting, with a constant KV Cache. The model innovatively mimics the attention pattern of humans copying books by hand and shares technical lineage with DeepSeek OCR, sparking discussions about talent mobility.

@Fenng: HuggingFace and GitHub charts hit top four, stars surpass 10k in just 5 days — Baidu Unlimited OCR becomes one of the fastest growing open source projects. I've seen many people mentioning Baidu's Unlimited-OCR in my timeline lately. Actually, OCR has always been a traditional strength of Baidu…

X AI KOLs Following

Baidu's open source project Unlimited-OCR tops four charts on HuggingFace and GitHub, with stars exceeding 10k in five days. The model uses a MoE architecture (3B total parameters, 570M activated parameters) and excels at continuous recognition of long documents. Inspired by how humans copy books, it also offers new ideas for long-term memory management in large models.