@DailyDoseOfDS_: Fine-tune DeepSeek-OCR on your own language! (100% local) Most vision models treat documents as massive sequences of to…

X AI KOLs Timeline Models

Summary

DeepSeek-OCR is a 3B vision model using context optical compression for efficient document processing. Fine-tuning it on Persian text using Unsloth achieved an 88.26% improvement in character error rate, all open-source and runnable on a single GPU.

Fine-tune DeepSeek-OCR on your own language! (100% local) Most vision models treat documents as massive sequences of tokens, making long-context processing expensive and slow. DeepSeek-OCR uses context optical compression to convert 2D layouts into vision tokens, enabling efficient processing of complex documents. It is a 3B-parameter vision model that achieves 97% precision while using 10x fewer vision tokens than text-based LLMs. In fact, you can easily fine-tune it for your specific use case on a single GPU. We used Unsloth to run this experiment on Persian text and saw an 88.26% improvement in character error rate. ↳ Base model: 149% character error rate (CER) ↳ Fine-tuned model: 60% CER (57% more accurate) ↳ Training time: 60 steps on a single GPU Persian was just the test case. You can swap in your own dataset for any language, document type, or specific domain you're working with. We've shared the complete guide in the next tweet, which includes the code, notebooks, and environment setup ready to run with a single click. Everything is 100% open-source!
Original Article
View Cached Full Text

Cached at: 06/08/26, 03:26 PM

Fine-tune DeepSeek-OCR on your own language!

(100% local)

Most vision models treat documents as massive sequences of tokens, making long-context processing expensive and slow.

DeepSeek-OCR uses context optical compression to convert 2D layouts into vision tokens, enabling efficient processing of complex documents.

It is a 3B-parameter vision model that achieves 97% precision while using 10x fewer vision tokens than text-based LLMs.

In fact, you can easily fine-tune it for your specific use case on a single GPU.

We used Unsloth to run this experiment on Persian text and saw an 88.26% improvement in character error rate.

↳ Base model: 149% character error rate (CER) ↳ Fine-tuned model: 60% CER (57% more accurate) ↳ Training time: 60 steps on a single GPU

Persian was just the test case. You can swap in your own dataset for any language, document type, or specific domain you’re working with.

We’ve shared the complete guide in the next tweet, which includes the code, notebooks, and environment setup ready to run with a single click.

Everything is 100% open-source!

Tech Stack:

  • @UnslothAI to run and fine-tune the model
  • @LightningAI environments for hosting and deployment

Find the code and environment setup here:

Similar Articles

@geekbb: Baidu's open-source visual language model OCR project, upgraded from DeepSeek-OCR, focuses on one-shot parsing of extremely long documents. The model has two inference modes: 'gundam' mode for dense text in a single image, and 'base' mode for multi-page or PDF processing. https://github…

X AI KOLs Timeline

Baidu has open-sourced the visual language model Unlimited-OCR, upgraded from DeepSeek-OCR, supporting one-shot parsing of extremely long documents, offering two inference modes: gundam (dense text in a single image) and base (multi-page/PDF).

baidu/Unlimited-OCR

Hugging Face Models Trending

Baidu releases Unlimited-OCR, a new model for one-shot long-horizon document parsing, building on Deepseek-OCR. It supports single image and multi-page/PDF parsing via Hugging Face Transformers and SGLang.

Cross-Temporal Sinhala OCR: Page-Level Adaptation and Diachronic Analysis

arXiv cs.CL

This paper introduces sinhala-ocr-lk-acts-1010, the first publicly available real-world page-level dataset for Sinhala OCR, and fine-tunes three vision language models (DeepSeek-OCR V1, DeepSeek-OCR V2, LightOnOCR-2-1B) using QLoRA. LightOnOCR-2-1B achieves a CER of 1.05%, outperforming both open-source and commercial OCR models, and maintains consistent performance across degraded documents from different time periods.