baidu

Tag

Cards List
#baidu

@vista8: Baidu's technical strength is still impressive. Recently, this Unlimited OCR has even caught the attention of Yann LeCun! Unlimited OCR proposes Reference Sliding Window Attention (R-SWA)...

X AI KOLs Timeline · 3d ago Cached

Baidu's open-source Unlimited OCR model proposes the Reference Sliding Window Attention (R-SWA) mechanism, achieving continuous parsing of dozens of pages with 3 billion parameters, gaining high attention on GitHub and HuggingFace.

0 favorites 0 likes
#baidu

@DataScienceDojo: A Chinese company just open-sourced an 𝐎𝐂𝐑 that fixes something most AI-powered OCR tools quietly struggle with: the…

X AI KOLs Timeline · 5d ago Cached

Unlimited-OCR, a new open-source OCR model from a Chinese company, solves the memory growth issue common in AI OCR tools by keeping memory usage flat regardless of document length, enabling single-pass reading of dozens of pages at 32K context. It's MIT-licensed, 3B parameters, multilingual, and already popular on GitHub.

0 favorites 0 likes
#baidu

@S0N_IA: | CHINA HAS ENDED THE OCR BUSINESS A 3 billion parameter model the size of a peanut can read a full 100-page PDF in one…

X AI KOLs Timeline · 5d ago Cached

Baidu releases Unlimited-OCR, a 3B parameter open-source model that reads full 100-page PDFs in one go with a 32K context window, achieving 93% accuracy and running locally. It has 1.9 million Hugging Face downloads but little mainstream attention.

0 favorites 0 likes
#baidu

@Fenng: HuggingFace and GitHub charts hit top four, stars surpass 10k in just 5 days — Baidu Unlimited OCR becomes one of the fastest growing open source projects. I've seen many people mentioning Baidu's Unlimited-OCR in my timeline lately. Actually, OCR has always been a traditional strength of Baidu…

X AI KOLs Following · 2026-06-29 Cached

Baidu's open source project Unlimited-OCR tops four charts on HuggingFace and GitHub, with stars exceeding 10k in five days. The model uses a MoE architecture (3B total parameters, 570M activated parameters) and excels at continuous recognition of long documents. Inspired by how humans copy books, it also offers new ideas for long-term memory management in large models.

0 favorites 0 likes
#baidu

@KaichaoYou: a great ocr model from baidu! you would be surprisied if you see how popular ocr models are, they are sometimes even mo…

X AI KOLs Timeline · 2026-06-28 Cached

Baidu's Unlimited-OCR model, using Reference Sliding Window Attention, is now supported in vLLM, enabling efficient one-shot parsing of entire books with constant memory usage.

0 favorites 0 likes
#baidu

@_akhaliq: Baidu just released Unlimited-OCR

X AI KOLs Following · 2026-06-23 Cached

Baidu has released Unlimited-OCR, an optical character recognition service with no usage limits.

0 favorites 0 likes
#baidu

@vanstriendaniel: It's raining OCR models again! @Baidu_Inc's Unlimited-OCR is one of the more interesting. You can try it without much e…

X AI KOLs Following · 2026-06-23 Cached

This post shows how to serve Baidu's Unlimited-OCR model as a temporary, OpenAI-compatible endpoint on Hugging Face Jobs, enabling multi-page document parsing with features like table-to-HTML and equation-to-LaTeX extraction.

0 favorites 0 likes
#baidu

Unlimited OCR: One-Shot Long-Horizon Parsing

Hacker News Top · 2026-06-23 Cached

Baidu releases Unlimited-OCR, an open-source model for one-shot long-horizon document parsing, building upon Deepseek-OCR with support for single images, multi-page documents, and PDFs.

0 favorites 0 likes
#baidu

@BaiduAI_News: We’re open-sourcing Unlimited OCR — built to read long documents in one pass. With 3B total parameters and only 500M ac…

X AI KOLs Timeline · 2026-06-23 Cached

Baidu open-sources Unlimited OCR, a 3B parameter model (500M activated) that reads long documents in a single pass using Reference Sliding Window Attention (R-SWA), achieving state-of-the-art results on OmniDocBench.

0 favorites 0 likes
#baidu

@ErickSky: Baidu has just broken one of the biggest limitations of current OCR. Unlimited-OCR processes entire documents in a sing…

X AI KOLs Timeline · 2026-06-23 Cached

Baidu has released Unlimited-OCR, which processes entire documents in a single pass without chunking, overcoming a major limitation of current OCR technology.

0 favorites 0 likes
#baidu

@geekbb: Baidu's open-source visual language model OCR project, upgraded from DeepSeek-OCR, focuses on one-shot parsing of extremely long documents. The model has two inference modes: 'gundam' mode for dense text in a single image, and 'base' mode for multi-page or PDF processing. https://github…

X AI KOLs Timeline · 2026-06-23 Cached

Baidu has open-sourced the visual language model Unlimited-OCR, upgraded from DeepSeek-OCR, supporting one-shot parsing of extremely long documents, offering two inference modes: gundam (dense text in a single image) and base (multi-page/PDF).

0 favorites 0 likes
#baidu

@berryxia: Wow, this move directly poached DeepSeek's talent! Last night I saw this interesting OCR open-source model on HuggingFace and the fascinating story behind it. This OCR model is completely different from traditional ones! Its speed and accuracy are absolutely unbeatable~~ Let me start with some background, for those who are familiar…

X AI KOLs Timeline · 2026-06-23 Cached

Baidu has open-sourced the Unlimited OCR model, which uses the R-SWA attention mechanism to process hundreds of pages in a single pass without page splitting, with a constant KV Cache. The model innovatively mimics the attention pattern of humans copying books by hand and shares technical lineage with DeepSeek OCR, sparking discussions about talent mobility.

0 favorites 0 likes
#baidu

@GoSailGlobal: Current OCR processes multi-page documents page by page. Every time you turn a page, memory is reset. Today, Baidu quietly open-sourced a model on GitHub and HuggingFace called Unlimited OCR, inspired by how humans copy books: - When copying a book, you don't reread hundreds of pages every time you write a word...

X AI KOLs Timeline · 2026-06-22 Cached

Baidu has open-sourced the Unlimited OCR model, which uses a Reference Sliding Window Attention (R-SWA) mechanism to parse documents up to 32K context in a single pass, eliminating the need for page-by-page inference.

0 favorites 0 likes
#baidu

@paulwalker99318: This LatePost interview is packed with information about Baidu US R&D, Scaling Laws, OpenAI, Anthropic, and Cerebras. > "Dario joining Baidu was a very important step in his career. He was recruited by Greg Diamos. And before joining Baidu, Dario didn't have a computer science or AI background — he came from math, physics, and biology. Greg Diamos saw his intuition for AI and ability to train models."

X AI KOLs Timeline · 2026-06-22 Cached

A summary of the LatePost interview, reviewing Baidu US R&D's early AI布局, including investing in Cerebras, nearly investing in OpenAI and Anthropic, and the flow of talent from Baidu to these companies.

0 favorites 0 likes
#baidu

Unlimited OCR Works

Hugging Face Daily Papers · 2026-06-22 Cached

Unlimited OCR introduces Reference Sliding Window Attention to eliminate growing memory consumption in long-sequence OCR tasks, enabling efficient transcription of multiple pages in a single forward pass.

0 favorites 0 likes
#baidu

baidu/Unlimited-OCR

Hugging Face Models Trending · 2026-06-19 Cached

Baidu releases Unlimited-OCR, a new model for one-shot long-horizon document parsing, building on Deepseek-OCR. It supports single image and multi-page/PDF parsing via Hugging Face Transformers and SGLang.

0 favorites 0 likes
#baidu

@smithandai: Baidu quietly released such a great product. This is really a necessity! Every time I translate a paper, the dense numbers give me a headache!

X AI KOLs Timeline · 2026-06-16

The user praised a product launched by Baidu that effectively solves the pain point of dealing with numbers when translating papers, considering it a necessity.

0 favorites 0 likes
#baidu

@rionaifantasy: Unbelievable! How Can a 34.5M Parameter OCR Beat a 235B Large Model? Let me tell you something ridiculous: I used to believe the future of OCR would inevitably be devoured by ever-larger multimodal large models. But after seeing PP-OCRv6 released by Baidu Wenxin, I've changed my mind. Because it doesn't follow the path of "continuing to pile on parameters..."

X AI KOLs Timeline · 2026-06-16 Cached

Baidu Wenxin releases PP-OCRv6, offering three model tiers: Tiny, Small, and Medium, supporting over 50 languages. The Tiny version is only 1.5MB and can run locally in a browser, with the fastest single-image inference at 97ms, proving that small specialized models can outperform large models on OCR tasks.

0 favorites 0 likes
#baidu

@AdinaYakup: PP-OCRv6 just released by Baidu @PaddlePaddle tiny 1.5M / small 7.7M / medium 34.5M 48+ languages Supports handwritten/…

X AI KOLs Following · 2026-06-11 Cached

Baidu's PaddlePaddle released PP-OCRv6, an OCR model supporting 48+ languages with tiny (1.5M), small (7.7M), and medium (34.5M) sizes, optimized for edge deployment and handwritten/printed/industrial/screen/card text.

0 favorites 0 likes
#baidu

@10xmylife: Compared to overseas, is the independent site ecosystem in China really that bad? If so, what are the reasons? When people search for a brand, they look on WeChat, Weibo, Douyin, Xiaohongshu, Taobao, etc., rarely on Baidu. Merchants also mostly rely on platforms rather than building their own official websites. I think Baidu itself being terrible is one reason, but there must be deeper causes…

X AI KOLs Following · 2026-05-21 Cached

Discusses why China's independent site (brand self-built official website) ecosystem is weaker than abroad, arguing that users mainly search for brands on platforms like WeChat, Weibo, and Douyin rather than Baidu, and that merchants rely on platforms instead of building their own official websites.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback