All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts

Hugging Face Daily Papers Papers

Summary

The paper introduces ScriptMoE, a script-aware mixture-of-experts architecture for all-in-one multilingual scene text recognition, along with the TextMuSS-10M synthetic dataset, achieving state-of-the-art accuracy on benchmarks.

Multilingual scene text recognition (STR) remains challenging due to the scarcity of training data for most languages and the difficulty of serving diverse scripts within a single model. Existing solutions either deploy one recognizer per language, inflating cost and introducing error accumulation, or rely on massive vision-language models (VLMs) that are expensive and still inaccurate on many scripts. In this work, we pursue an all-in-one multilingual recognizer that is simpler than per-language experts, lighter than VLMs, and more accurate than both. First, we construct TextMuSS-10M, a large-scale synthetic scene text dataset spanning 10 scripts and 229 languages. It provides balanced and sufficient supervision where real data is unavailable. Second, we propose ScriptMoE, a script-aware Mixture-of-Experts (MoE) architecture. It shares a single visual encoder and replaces the dense decoder with a sparse MoE block, which consists of an image-level router dispatches each image to the top-2 script-aligned experts and a shared expert absorbs cross-script knowledge. Extensive experiments on our assembled TextMuSS-Bench (10 scripts, 10,899 images) show that ScriptMoE achieves the highest accuracy of 82.06%, outperforming the strongest STR baseline by 1.31%. On the CC-OCR end-to-end multilingual task, replacing only the recognizer in PP-OCRv5 with ScriptMoE lifts F1 score from 65.71% to 80.89%, slightly surpassing the best VLM (80.73%) at a fraction of the parameter count.
Original Article
View Cached Full Text

Cached at: 09/23/26, 07:32 AM

Paper page - All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts

Source: https://huggingface.co/papers/2609.24058

Abstract

Multilingualscenetextrecognition(STR)remainschallengingduetothescarcityoftrainingdataformostlanguagesandthedifficultyofservingdiversescriptswithinasinglemodel.Existingsolutionseitherdeployonerecognizerperlanguage,inflatingcostandintroducingerroraccumulation,orrelyonmassivevision-languagemodels(VLMs)thatareexpensiveandstillinaccurateonmanyscripts.Inthiswork,wepursueanall-in-onemultilingualrecognizerthatissimplerthanper-languageexperts,lighterthanVLMs,andmoreaccuratethanboth.First,weconstructTextMuSS-10M,alarge-scalesyntheticscenetextdatasetspanning10scriptsand229languages.Itprovidesbalancedandsufficientsupervisionwhererealdataisunavailable.Second,weproposeScriptMoE,ascript-awareMixture-of-Experts(MoE)architecture.ItsharesasinglevisualencoderandreplacesthedensedecoderwithasparseMoEblock,whichconsistsofanimage-levelrouterdispatcheseachimagetothetop-2script-alignedexpertsandasharedexpertabsorbscross-scriptknowledge.ExtensiveexperimentsonourassembledTextMuSS-Bench(10scripts,10,899images)showthatScriptMoEachievesthehighestaccuracyof82.06%,outperformingthestrongestSTRbaselineby1.31%.OntheCC-OCRend-to-endmultilingualtask,replacingonlytherecognizerinPP-OCRv5withScriptMoEliftsF1scorefrom65.71%to80.89%,slightlysurpassingthebestVLM(80.73%)atafractionoftheparametercount.

View arXiv pageView PDFProject pageGitHub2Add to collection

Get this paper in your agent:

hf papers read 2609\.24058

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.24058 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.24058 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.24058 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

MobileMoE: Scaling On-Device Mixture of Experts

Hugging Face Daily Papers

MobileMoE introduces efficient on-device mixture-of-experts language models with sub-billion parameters, achieving better performance and efficiency than dense baselines and existing MoE models. The models are trained on open-source datasets and demonstrate significant speedups on commodity smartphones.

LongMoE: Longitudinal Multimodal Learning via Trajectory-Aware Mixture-of-Experts

arXiv cs.LG

LongMoE proposes a unified framework that jointly addresses modality missingness and longitudinal dynamics in multimodal clinical learning, using context-aware imputation, attentional tokenization, trajectory-aware encoding, and sparse mixture-of-experts routing. Experiments on ADNI, OASIS-3, and MIMIC-IV demonstrate improved robustness under missing modalities while remaining competitive in full-modality settings.