On the Limitations of Cross-Lingual Consistency in Multilingual Text-to-image Generation

arXiv cs.CL Papers

Summary

This paper introduces LingT2I, a 10-language, 33K-prompt benchmark for evaluating cross-lingual consistency in text-to-image generation, revealing linguistic inequality and language-dependent trade-offs across content generation and text rendering.

arXiv:2608.11002v1 Announce Type: new Abstract: Text-to-image (T2I) generation has achieved remarkable progress in recent years. However, existing research has largely focused on English-only settings, leaving cross-lingual performance gaps and language-specific effects insufficiently explored. To fill this gap, we introduce LingT2I, a benchmark covering 10 widely used languages with 33K prompts, designed to evaluate cross-lingual effects in both content generation and text rendering. Building on this benchmark, we conduct a comprehensive cross-lingual analysis, uncovering linguistic inequality and language-dependent trade-offs across evaluation dimensions. Beyond quantitative evaluation, we further reveal a range of language-dependent generation patterns, highlighting how linguistic factors and their corresponding cultural contexts systematically impact model outputs. Our benchmark and analysis provide a foundation for studying cross-lingual behavior in T2I generation and facilitate the development of more robust and inclusive models. Code and dataset are available at https://github.com/RISys-Lab/LingT2I.
Original Article
View Cached Full Text

Cached at: 08/12/26, 08:38 AM

# On the Limitations of Cross-Lingual Consistencyin Multilingual Text-to-image Generation
Source: [https://arxiv.org/html/2608.11002](https://arxiv.org/html/2608.11002)
## On the Limitations of Cross\-Lingual Consistency in Multilingual Text\-to\-image GenerationConference:Proceedings of the 34th ACM International Conference on Multimedia; November 10–14, 2026; Rio de Janeiro, BrazilProceedings of the 34th ACM International Conference on Multimedia \(MM ’26\), November 10–14, 2026, Rio de Janeiro, BrazilDOI:[10\.1145/3767308\.3835058](https://doi.org/10.1145/3767308.3835058)ISBN:979\-8\-4007\-2213\-4/2026/11CCS:Computing methodologies Computer visionCCS:Computing methodologies Machine translationCCS:General and reference Evaluation

,Zhonghao YanAffiliation:Queen Mary University of London,London,UKemail:[yanzhonghao531@gmail\.com](mailto:[email protected]),Binzhu XieAffiliation:The Chinese University of Hong Kong,Hong Kong,Chinaemail:[bzxie@cse\.cuhk\.edu\.hk](mailto:[email protected]),Shi QiuAffiliation:The Chinese University of Hong Kong,Hong Kong,Chinaemail:[shiqiu@cse\.cuhk\.edu\.hk](mailto:[email protected]),Muzammal NaseerNote:Corresponding author\.Affiliation:Khalifa University,Abu Dhabi,UAE ,The University of Western Australia,Perth,Australiaemail:[muhammadmuzammal\.naseer@ku\.ac\.ae](mailto:[email protected]),Naveed AkhtarAffiliation:The University of Melbourne,Melbourne,Australiaemail:[naveed\.akhtar1@unimelb\.edu\.au](mailto:[email protected])andMubarak ShahAffiliation:University of Central Florida,Florida,USAemail:[shah@crcv\.ucf\.edu](mailto:[email protected])

2026; © cc

###### Keywords:

Multilingual T2I, Cross\-lingual Benchmark, Language Fairness

††cc\-license:by\-nc\-nd![[Uncaptioned image]](https://arxiv.org/html/2608.11002v1/fig/risys-lab.png)On the Limitations of Cross\-Lingual Consistency in Multilingual Text\-to\-image GenerationSicheng Zhang1, Zhonghao Yan2, Binzhu Xie3, Shi Qiu3, Muzammal Naseer1,4,∗, Naveed Akhtar5, Mubarak Shah61Khalifa University,2Queen Mary University of London,3The Chinese University of Hong Kong,4The University of Western Australia,5The University of Melbourne,6University of Central Florida∗Corresponding authorText\-to\-image \(T2I\) generation has achieved remarkable progress in recent years\. However, existing research has largely focused on English\-only settings, leaving cross\-lingual performance gaps and language\-specific effects insufficiently explored\. To fill this gap, we introduceLingT2I, a benchmark covering 10 widely used languages with 33K prompts, designed to evaluate cross\-lingual effects in both content generation and text rendering\. Building on this benchmark, we conduct a comprehensive cross\-lingual analysis, uncovering linguistic inequality and language\-dependent trade\-offs across evaluation dimensions\. Beyond quantitative evaluation, we further reveal a range of language\-dependent generation patterns, highlighting how linguistic factors and their corresponding cultural contexts systematically impact model outputs\. Our benchmark and analysis provide a foundation for studying cross\-lingual behavior in T2I generation and facilitate the development of more robust and inclusive models\.GitHub:[https://github\.com/RISys\-Lab/LingT2I](https://github.com/RISys-Lab/LingT2I)![[Uncaptioned image]](https://arxiv.org/html/2608.11002v1/fig/huggingface_logo.png)Dataset:[https://huggingface\.co/datasets/RISys\-Lab/LingT2I](https://huggingface.co/datasets/RISys-Lab/LingT2I)

![Refer to caption](https://arxiv.org/html/2608.11002v1/teaser.png)Figure 1\.Challenges of multilingual T2I\.\(i\) Linguistic inequality: the Hindi version has an incorrect number of objects; \(ii\) Text rendering failures: misarrangements of letters and structural errors in characters; \(iii\) Multi\-dimensional trade\-offs across languages: the Korean results appear in oil\-painting style, inconsistent with the intended photograph style\. \(iv\) Language\-dependent generation patterns: the same prompt yields textiles with distinct cultural characteristics across languages\.## 1\.Introduction

Language is a primary interface between humans and artificial intelligence, playing a decisive role in shaping multimodal generative content\. Text\-to\-image \(T2I\) models exemplify this trend, achieving remarkable success in English through large\-scale diffusion frameworks\([58](https://arxiv.org/html/2608.11002#bib.bib38);[70](https://arxiv.org/html/2608.11002#bib.bib23)\)\. However, these advances largely rely on English\-dominant datasets like COCO Captions\([9](https://arxiv.org/html/2608.11002#bib.bib41)\)and LAION\-5B\([51](https://arxiv.org/html/2608.11002#bib.bib21)\), leaving their capabilities in multilingual settings largely underexplored\. In parallel, research in multilingual NLP has emphasized that linguistic diversity and inclusivity are crucial for developing equitable and culturally aware AI\([25](https://arxiv.org/html/2608.11002#bib.bib43);[32](https://arxiv.org/html/2608.11002#bib.bib42)\), underscoring the importance of extending T2I evaluation beyond English\.

Multilingual T2I generation faces two fundamental challenges: generating visual content must account for the unique cultural attributes embedded in different languages; the Text Rendering task—that is, generating images with specific textual content—requires handling diverse writing systems\. Previous work has made efforts to extend T2I models to multilingual settings, either by incorporating existing encoders with limited multilingual foundations\([29](https://arxiv.org/html/2608.11002#bib.bib1);[64](https://arxiv.org/html/2608.11002#bib.bib6);[47](https://arxiv.org/html/2608.11002#bib.bib5)\)or by leveraging large generative LLMs for stronger prompt interpretation\([70](https://arxiv.org/html/2608.11002#bib.bib23);[61](https://arxiv.org/html/2608.11002#bib.bib39)\)\. However, existing studies on multilingual T2I remain limited\([50](https://arxiv.org/html/2608.11002#bib.bib53);[18](https://arxiv.org/html/2608.11002#bib.bib22);[67](https://arxiv.org/html/2608.11002#bib.bib68);[23](https://arxiv.org/html/2608.11002#bib.bib58)\), focusing mainly on image quality while overlooking linguistic aspects and text rendering evaluation\.

This raises fundamental questions:do T2I models truly possess multilingual competence?More importantly,what factors underlie the performance disparities across languages?Figure[1](https://arxiv.org/html/2608.11002#S0.F1)highlights several representative phenomena: \(i\) prompts in low\-resource languages like Hindi consistently underperform compared to high\-resource ones, reflecting clear*linguistic inequality*; \(ii\) non\-Latin scripts in the Text Rendering task often appear broken, unreadable, or hallucinated, underscoring the difficulty of handling diversewriting systems; \(iii\) even for semantically identical prompts, different languages exhibit divergenttrade\-offsacross dimensions such as realism, semantic faithfulness, and style; these interactions may appear ascoupled improvements,conflicting trends, orbalanced compromises, highlighting the instability of cross\-lingual generalization; and \(iv\) generation behavior varies systematically across languages, indicating that T2I models are influenced not only by textual semantics but also by language\-specific priors\. The causes, including data distribution, linguistic morphology, and writing systems, remain underexplored, underscoring the need for a systematic framework for cross\-lingual analysis\.

To investigate these challenges, we introduceLingT2I, a benchmark specifically designed for analyzing cross\-lingual effects, which covers 10 widely used languages and evaluates bothContent GenerationandText Renderingtasks\. This unified dataset forms a foundation for large\-scale analysis of multilingual T2I generation\.

We benchmark several state\-of\-the\-art T2I models—including Nano Banana\([58](https://arxiv.org/html/2608.11002#bib.bib38)\), Z\-Image\([62](https://arxiv.org/html/2608.11002#bib.bib47)\), and EasyText\([35](https://arxiv.org/html/2608.11002#bib.bib28)\)—on LingT2I and present a comprehensive large\-scale cross\-lingual analysis\. Our results reveal threekey findings: \(i\) general\-purpose models exhibit severe linguistic inequality, with performance skewed toward high\-resource Indo\-European languages; \(ii\) non\-Latin writing systems remain a major bottleneck, leading to broken or unreadable text rendering; and \(iii\) language\-specific cultural and typological factors systematically impact generation behavior, reshaping trade\-offs across evaluation dimensions\. These findings expose fundamental limitations of current multilingual T2I systems and provide guidance for developing fairer and more culturally inclusive generative models\. Ourkey contributionsare as follows:

- •We present LingT2I, a new dataset covering 10 widely used languages with 33K prompts, designed to analyze cross\-lingual effects in both general Content Generation and Text Rendering\.
- •We provide the first comprehensive cross\-lingual analysis, revealing linguistic inequality and language\-specific trade\-offs across dimensions in T2I models\.
- •Our analysis reveals various language\-dependent generative patterns, providing valuable insights for model design\.

## 2\.Related Work

Multilingual Text\-to\-image Generation\.Recent works have endowed text\-to\-image models with multilingual abilities\. Models\([53](https://arxiv.org/html/2608.11002#bib.bib7);[19](https://arxiv.org/html/2608.11002#bib.bib48);[71](https://arxiv.org/html/2608.11002#bib.bib18);[7](https://arxiv.org/html/2608.11002#bib.bib4)\)such as SD 3\.5\([54](https://arxiv.org/html/2608.11002#bib.bib2)\), FLUX\([29](https://arxiv.org/html/2608.11002#bib.bib1)\), and Z\-Image\([62](https://arxiv.org/html/2608.11002#bib.bib47)\), adopt diffusion or diffusion transformer \(DiT\) architectures, where language understanding is primarily handled by pretrained text encoders\([64](https://arxiv.org/html/2608.11002#bib.bib6);[47](https://arxiv.org/html/2608.11002#bib.bib5);[57](https://arxiv.org/html/2608.11002#bib.bib49)\)\. Recent approaches such as HunyuanImage\-3\.0\([61](https://arxiv.org/html/2608.11002#bib.bib39)\), Janus\-Pro\([8](https://arxiv.org/html/2608.11002#bib.bib3)\)and NextStep\-1\([59](https://arxiv.org/html/2608.11002#bib.bib72)\)directly model text and image tokens within an autoregressive Transformer, where multilingual capability is intrinsic to the pretrained LLM backbone\([63](https://arxiv.org/html/2608.11002#bib.bib50);[14](https://arxiv.org/html/2608.11002#bib.bib51);[74](https://arxiv.org/html/2608.11002#bib.bib73)\)\. Advanced methods like Qwen\-Image\([70](https://arxiv.org/html/2608.11002#bib.bib23)\)and Omni\-Diffusion\([56](https://arxiv.org/html/2608.11002#bib.bib74)\)move beyond conventional pipelines by unifying language and visual modeling, where multilingual capability arises from the shared modeling space and training data\.

To specifically enhance multilingual capability, one direction leverages strong multilingual encoders such as AltDiffusion\([75](https://arxiv.org/html/2608.11002#bib.bib10)\)with AltCLIP\([10](https://arxiv.org/html/2608.11002#bib.bib8)\), another focuses on encoder\-generator alignment with lightweight adapters \(GlueGen\([43](https://arxiv.org/html/2608.11002#bib.bib12)\), MuLan\([72](https://arxiv.org/html/2608.11002#bib.bib11)\)\), and a third exploits parameter\-efficient distillation from English teachers \(PEA\-Diffusion\([36](https://arxiv.org/html/2608.11002#bib.bib14)\), X2I\([37](https://arxiv.org/html/2608.11002#bib.bib13)\)\)\.

Multilingual Text Rendering\.The ability to generate specified text within images serves as a key indicator of a T2I model’s linguistic competence\. Recent advances such as Glyph\-ByT5\([33](https://arxiv.org/html/2608.11002#bib.bib34);[34](https://arxiv.org/html/2608.11002#bib.bib35)\), AnyText\([66](https://arxiv.org/html/2608.11002#bib.bib26);[65](https://arxiv.org/html/2608.11002#bib.bib27)\), and EasyText\([35](https://arxiv.org/html/2608.11002#bib.bib28)\)have introduced specialized approaches that incorporate glyph\-aware encoders, OCR\-guided features, or DiT to improve multilingual text rendering\. Meanwhile, general\-purpose models\([29](https://arxiv.org/html/2608.11002#bib.bib1);[58](https://arxiv.org/html/2608.11002#bib.bib38);[70](https://arxiv.org/html/2608.11002#bib.bib23)\)have begun to emphasize text generation\. Nevertheless, multilingual text rendering remains limited in both capability and systematic evaluation\.

Language\-related Bias and Cross\-lingual Effects\.In NLP, cross\-lingual behavior has been extensively analyzed\([44](https://arxiv.org/html/2608.11002#bib.bib61);[48](https://arxiv.org/html/2608.11002#bib.bib63);[41](https://arxiv.org/html/2608.11002#bib.bib64);[24](https://arxiv.org/html/2608.11002#bib.bib65);[52](https://arxiv.org/html/2608.11002#bib.bib66)\), with studies showing significant linguistic inequality across languages\([25](https://arxiv.org/html/2608.11002#bib.bib43);[4](https://arxiv.org/html/2608.11002#bib.bib60);[49](https://arxiv.org/html/2608.11002#bib.bib62);[45](https://arxiv.org/html/2608.11002#bib.bib45);[78](https://arxiv.org/html/2608.11002#bib.bib46)\)\. Building on this, recent work has begun to investigate biases in T2I models more broadly\([11](https://arxiv.org/html/2608.11002#bib.bib69);[68](https://arxiv.org/html/2608.11002#bib.bib75);[17](https://arxiv.org/html/2608.11002#bib.bib76)\), such as social\([3](https://arxiv.org/html/2608.11002#bib.bib54);[28](https://arxiv.org/html/2608.11002#bib.bib57)\), cultural\([39](https://arxiv.org/html/2608.11002#bib.bib59);[27](https://arxiv.org/html/2608.11002#bib.bib37);[76](https://arxiv.org/html/2608.11002#bib.bib40)\), and geographic biases\([2](https://arxiv.org/html/2608.11002#bib.bib55);[20](https://arxiv.org/html/2608.11002#bib.bib15)\)\. However, these studies are still largely conducted with English prompts, making it difficult todisentangle intrinsic model biases from language\-dependent generation patterns\.

Despite these efforts, research on cross\-lingual effects in T2I models remains limited and has mostly focused on isolated specific phenomena or narrow technical aspects\([23](https://arxiv.org/html/2608.11002#bib.bib58);[18](https://arxiv.org/html/2608.11002#bib.bib22);[26](https://arxiv.org/html/2608.11002#bib.bib67);[67](https://arxiv.org/html/2608.11002#bib.bib68)\), such as differences in concept coverage\([50](https://arxiv.org/html/2608.11002#bib.bib53);[75](https://arxiv.org/html/2608.11002#bib.bib10)\), the effect of non\-Latin characters\([55](https://arxiv.org/html/2608.11002#bib.bib56)\), or case studies targeting individual languages\([38](https://arxiv.org/html/2608.11002#bib.bib52)\)\. Moreover, NeoBabel\([15](https://arxiv.org/html/2608.11002#bib.bib78)\)studies native multilingual generation and evaluates cross\-lingual consistency and code\-switching robustness\. However, systematic investigations into inherent linguistic inequality, multi\-dimensional trade\-offs, and latent language\-dependent generation patterns remain largely underexplored\.

## 3\.Cross\-lingual Benchmark: LingT2I

### 3\.1\.Benchmark Coverage

Task Selection\.The multilingual setting brings two fundamental challenges\. First, models must understand prompts in different languages and still generate images that are semantically accurate and culturally coherent\. Second, they must be able to render text faithfully across diverse writing systems, each with its own glyph complexity, layout, and formatting rules\. To capture these challenges in a structured way, LingT2I defines two evaluation tasks:Content GenerationandText Rendering\.

![Refer to caption](https://arxiv.org/html/2608.11002v1/dim_def_cg.png)Figure 2\.Evaluation Dimensions ofContent GenerationTask\.For each dimension, we provide its definition, multilingual examples, and representative examples of both high\-quality and failure cases in generated images\.Language Coverage\.To align with both Content Generation and Text Rendering, our language set balances cultural and semantic diversity and writing\-system variety\.*Linguistic branches*ground prompts in distinct cultural and semantic contexts that shape interpretation, whereas*writing systems*\(e\.g\., glyph complexity, reading direction, segmentation, and character composition\) directly determine the difficulty of text rendering\. Guided by this dual perspective, LingT2I covers 10 languages spanning diverse scripts and families \(Table[1](https://arxiv.org/html/2608.11002#S3.T1)\)\. The selection balances population size\([16](https://arxiv.org/html/2608.11002#bib.bib30)\), global coverage\([5](https://arxiv.org/html/2608.11002#bib.bib16)\), and the Power Language Index \(PLI\)\([6](https://arxiv.org/html/2608.11002#bib.bib36)\), ensuring representativeness and practical relevance\. For each language, we annotate its script type and linguistic branch111Classification of Japanese and Korean remains debated\., providing structured background for subsequent cross\-lingual and cultural analyses\.

Table 1\.Statistics and classification of the 10 languages in LingT2I, including speaker population \(Spk\., billion\)\([16](https://arxiv.org/html/2608.11002#bib.bib30)\), global coverage \(Cov\., %\)\([5](https://arxiv.org/html/2608.11002#bib.bib16)\), Power Language Index \(PLI\)\([6](https://arxiv.org/html/2608.11002#bib.bib36)\), script type, and language branch\.
### 3\.2\.Content Generation Subset

Evaluation Dimensions\.Inspired by existing benchmarks\([31](https://arxiv.org/html/2608.11002#bib.bib25);[77](https://arxiv.org/html/2608.11002#bib.bib9)\), we systematically organize a set of 10 evaluation dimensions that cover four fundamental aspects of image generation:Image Quality,Task Alignment,Diversity, andRobustness\. As shown in Figure[2](https://arxiv.org/html/2608.11002#S3.F2), these dimensions enable a systematic characterization of multi\-dimensional trade\-offs in multilingual generation\. Annotation Pipeline\.We construct the Content Generation subset based on the DOCCI dataset\([40](https://arxiv.org/html/2608.11002#bib.bib17)\)\. In our setting, we only utilize the textual component as the source corpus\. For each captioncc, its official annotation includes multiple aspects of the image, such as objects, attributes, spatial relationships, and scene descriptions\. To align with the predefined evaluation dimension set𝒟\\mathcal\{D\}, we design a dimension\-aware annotation and prompt construction pipeline\.

Specifically, we employ designed prompts to guide Gemini 2\.5 Flash\([13](https://arxiv.org/html/2608.11002#bib.bib19)\)to extract dimension\-relevant semantic information from the original captioncc, denoted asℐd=ℳ⁡\(c,d\)\\mathcal\{I\}\_\{d\}=\\mathcal\{M\}\(c,d\)for each target dimensiond∈𝒟d\\in\\mathcal\{D\}\. This process emphasizes the semantic components most relevant to the target dimension\. For dimensions with explicit information in the caption \(e\.g\., Content Alignment, Realism\),ℐd\\mathcal\{I\}\_\{d\}is further fed to the annotation model to generate concise and dimension\-focused promptspd=ℳ⁡\(ℐd\)p\_\{d\}=\\mathcal\{M\}\(\\mathcal\{I\}\_\{d\}\)\. For dimensions that are not explicitly reflected in the original caption \(e\.g\., Style, Bias\), we first instruct the model to compress the descriptioncc, and then perform conditional expansion\. For instance, we append control phrases such as “inssstyle” to explicitly guide the T2I model toward generating outputs that satisfy the target dimension\. Finally, for the Toxicity dimension, we directly adopt the existing Toxigen\([21](https://arxiv.org/html/2608.11002#bib.bib70)\)dataset to avoid introducing additional harmful content\.

All generated English promptspd\{p\_\{d\}\}are then translated into nine additional languages using Gemini 2\.5 Pro\([13](https://arxiv.org/html/2608.11002#bib.bib19)\), with constraints to preserve semantic consistency, cultural appropriateness, and stylistic fidelity\. Details can be found in Section[3\.4](https://arxiv.org/html/2608.11002#S3.SS4)\.

Data Statistics\.In total, this subset comprises 30K prompts, distributed evenly across 10 dimensions and 10 languages \(300 prompts per dimension per language\)\. As shown in Appendix[B\.2](https://arxiv.org/html/2608.11002#A2.SS2), the English subset averages 21\.9 words, while all the multilingual prompts average 43\.1 tokens with the mT5 tokenizer\([73](https://arxiv.org/html/2608.11002#bib.bib44)\)\.

Evaluation\.i\)General Evaluation:CLIPScore\([22](https://arxiv.org/html/2608.11002#bib.bib20)\)is a widely used metric that measures image\-text alignment by computing the cosine similarity between the generated image and its prompt using CLIP embeddings\. However, the original CLIP\([46](https://arxiv.org/html/2608.11002#bib.bib32)\)exhibits much stronger performance in English than in other languages\([69](https://arxiv.org/html/2608.11002#bib.bib71)\)\. To address this, we replace CLIP with the multilingual encoder MetaCLIP2\([12](https://arxiv.org/html/2608.11002#bib.bib31)\), which provides a fairer measure across languages\.

ii\)Dimensional Evaluation:TRIGScore\([77](https://arxiv.org/html/2608.11002#bib.bib9)\)is an MLLM\-based evaluation metric that leverages log\-probabilities to produce fine\-grained scores across multiple quality dimensions\. We adapt the Qwen\-2\.5\-VL\([60](https://arxiv.org/html/2608.11002#bib.bib33)\)Model and redesign the evaluation prompts to explicitly instruct the model to consider language\-specific factors, enabling it to directly account for cross\-linguistic understanding\. Details can be found in Appendix[C](https://arxiv.org/html/2608.11002#A3)\.

### 3\.3\.Text Rendering Subset

Evaluation Dimensions\.In theText Renderingtask, we shift the focus of analysis to the text itself, usingTextual QualityandHarmony with the Backgroundas the two primary dimensions\. Figure[3](https://arxiv.org/html/2608.11002#S3.F3)shows the detailed dimension definitions and examples\.

![Refer to caption](https://arxiv.org/html/2608.11002v1/dim_def_tr.png)Figure 3\.Evaluation Dimensions ofText RenderingTask\.Table 2\.Overall cross\-linguistic performance of Content Generation \(CG\) and Text Rendering \(TR\) models\.InCGtask, results are measured byCLIPScore↑\\uparrow; InTRtask, results are measured byAverage Precision↑\\uparrow\. For each model we report the average \(Avg\.↑\\uparrow\) and standard deviation \(Std\.↓\\downarrow\) across languages, where the variance indicates model\-level linguistic inequality\. We also provide per\-language averages by model category, highlighting the language\-level disparities\.Annotation Pipeline\.We use English samples from EasyText\([35](https://arxiv.org/html/2608.11002#bib.bib28)\)as the source of raw prompts\. We keep the background promptccfixed in English and only translate the rendered texttt, i\.e\.,\(c,ten\)→\(c,tℓ\)\(c,t\_\{\\text\{en\}\}\)\\rightarrow\(c,t\_\{\\ell\}\), thereby isolating language variation to the text rendering component\. This design allows us to focus specifically on rendering performance, while also aligning with the fact that most Text Rendering models are primarily optimized for English prompts\. The translations into nine additional languages are also performed using Gemini 2\.5 Pro\([13](https://arxiv.org/html/2608.11002#bib.bib19)\)\. Data Statistics\.TheText Renderingsubset contains 3K samples, with 300 prompts per language\. As shown in Appendix[B\.2](https://arxiv.org/html/2608.11002#A2.SS2), each prompt specifies a multilingual text string to be rendered, which averages 3\.0 tokens with mT5 tokenizer, accompanied by an English background description averaging 82\.3 words and 120\.2 tokens\. Evaluation\.i\)General Evaluation:Precision is the primary metric for text rendering, reflecting the correctness of the generated text\. We evaluate text rendering using standard precision metrics, including character\-level NED\([30](https://arxiv.org/html/2608.11002#bib.bib29)\), token\-level NED, and sentence\-level accuracy\. We report the average of these metrics as the final score\.

ii\)Dimensional Evaluation:We follow the MLLM\-as\-judge framework in EasyText\([35](https://arxiv.org/html/2608.11002#bib.bib28)\)and implement it using Gemini 2\.5 Flash as the evaluation model\. The prompts are adapted to specify the target language and explicitly guide the model to account for language\-specific characteristics across different writing systems\.

### 3\.4\.Quality Control

We adopt a three\-part quality control process for dataset construction:Automatic Processing and Verification, where all automatic processing steps for bothContent GenerationandText Renderingare performed using Gemini 2\.5 Pro and verified through back\-translation and GPT\-5 cross\-checking, with problematic cases manually corrected;Error Analysis and Iterative Refinement, where pilot experiments are conducted on a randomly sampled 5% subset to identify common data issues and refine prompt construction and filtering before large\-scale generation; andHuman Quality Check, where native speakers evaluate another randomly sampled 5% subset, with 98% of the samples judged to be semantically consistent across languages\. More details of this section can be found in Appendix[B](https://arxiv.org/html/2608.11002#A2)\.

## 4\.Experiments

Implementation Details\.All the experiments are conducted on 4 NVIDIA A100 64G GPUs\. We evaluate 17 recent text\-to\-image models for the two tasks, including general\-purpose models widely used for English prompts, models specifically trained or adapted for multilingual generation and text rendering models, all deployed with default settings\. \(see Appendix[D\.1](https://arxiv.org/html/2608.11002#A4.SS1)\)\. During Evaluation, we use metaclip\-2\-worldwide\-huge\-quickgelu\([12](https://arxiv.org/html/2608.11002#bib.bib31)\)for CLIPScore, and Qwen\-2\.5\-VL 72B\([60](https://arxiv.org/html/2608.11002#bib.bib33)\)for TRIGScore, and Gemini 2\.5 Flash\([13](https://arxiv.org/html/2608.11002#bib.bib19)\)and mT5\-base\([73](https://arxiv.org/html/2608.11002#bib.bib44)\)for text rendering average precision\.

![Refer to caption](https://arxiv.org/html/2608.11002v1/dimension.png)Figure 4\.Cross\-lingual dimension analysis\. Language\-dimension correlations inContent Generation\(a, b\) andText Rendering\(c, d\)\. Language\-dependent trade\-offs between key dimension pairs inContent Generation\(e\) andText Rendering\(f\)\. This analysis is based on fine\-grained results in Table[6](https://arxiv.org/html/2608.11002#A4.T6)and Table[7](https://arxiv.org/html/2608.11002#A4.T7)\(in Appendix\), derived from models with strong multilingual fairness\.### 4\.1\.Cross\-lingual Inequality Analysis

#### 4\.1\.1\.Content Generation Task

General\-purpose models exhibit substantially higher linguistic inequality than multilingual enhanced models\.As shown in Table[2](https://arxiv.org/html/2608.11002#S3.T2), we report both the average performance and the variance across languages for each model, with the variance indicating linguistic inequality\. Results indicate that multilingual\-enhanced models exhibit much lower variance, suggesting more balanced cross\-lingual performance, while most general\-purpose models suffer from severe linguistic inequality, with Qwen\-Image, Z\-Image, and Lumina\-Next as notable exceptions\. Native multilingual architectures achieve better fairness than post\-hoc adaptations\.As shown in Table[2](https://arxiv.org/html/2608.11002#S3.T2), multilingual\-enhanced variants yield higher fairness \(lower variance\) than their base models, but this often comes at the cost of reduced performance in privileged languages such as English and French\. In contrast, Qwen\-Image \(built upon Qwen\-2\.5\-VL\) achieves comparably low variance \(0\.03\) while maintaining superior overall quality, suggesting that native multilingual architectures offer a more effective path toward fairness than adapter\- or distillation\-based post\-hoc methods\. Even with reduced inequality, performance remains stratified across language families and cultures\.Using the classification in Table[1](https://arxiv.org/html/2608.11002#S3.T1), we analyze results from language branch and cultural perspectives\. Under general\-purpose models, Germanic and Romance language branches lead \(EN=0\.78; FR=0\.73; ES=0\.71; PT=0\.69\), while Slavic \(RU=0\.52\) and East Asian languages \(CJK\) trail; Indo\-Iranian and Semitic are lowest \(HI/AR=0\.38\)\. With multilingual\-enhanced variants, branch means narrow but persist: Chinese joins the top, Romance and Slavic converge around 0\.66\-0\.68, while other groups remain lower despite notable gains \(e\.g\., JA=0\.64, KO=0\.62, HI=0\.58, AR=0\.62\)\. Thus, even with improved fairness, language family and cultural stratification endures\.

#### 4\.1\.2\.Text Rendering Task

All models exhibit strong linguistic inequality\.As shown in theText Renderingsection of Table[2](https://arxiv.org/html/2608.11002#S3.T2), large variances remain across all models, indicating that linguistic inequality persists regardless of model category\. Overall text\-rendering ability is weak—even models specialized for rendering struggle\. Among them, EasyText achieves a more balanced trade\-off between overall performance \(0\.67\) and fairness \(0\.14\), yet language disparities remain pronounced\. Performance across writing systems is particularly uneven\.From the perspective of writing systems, languages using the Latin alphabet \(EN, ES, PT, FR\) consistently perform best, maintaining leading results in both general\-purpose and rendering\-oriented models\. Chinese shows clear improvement in rendering\-oriented models, while other non\-Latin scripts remain consistently weaker\. TheContent Generationand theText Renderingtasks demand different multilingual capabilities\.While Qwen\-Image achieves a strong cross\-lingual average and high fairness inContent Generationtask, it shows pronounced linguistic inequality inText Renderingtask: English and Chinese remain relatively strong, whereas Arabic, Hindi, and Korean lag substantially, indicating that the multilingual capabilities required are not interchangeable\.

### 4\.2\.Cross\-lingual Multi\-dimensional Analysis

Language\-Dimension Correlation\.Figure[4](https://arxiv.org/html/2608.11002#S4.F4)\(a\-d\) shows how linguistic and typological variations affect fine\-grained model behavior\. In theContent Generationtask, models reveal a strongbias: high\-resource Indo\-European languages \(e\.g\., English, Germanic, Romance\) favor white and male characters, reflecting social skew in English\-centric corpora\. Non\-Indo\-European languages \(e\.g\., Indo\-Aryan, Semitic, Koreanic, Slavic\) yield numerically lower bias and more diverse depictions \(Figure[4](https://arxiv.org/html/2608.11002#S4.F4)\(a\)\(b\)\), though this largely results from weaker semantic grounding rather than genuine fairness\. Language background also impactstoxicity\. Indo\-Aryan, Semitic, Koreanic, and Slavic achieve higherToxicityscores—meaning fewer harmful elements—than Sinitic and Germanic\. This may stem from lower data exposure, causing models to generate safer yet generic content, and from moderation pipelines tuned for English, which may over\-filter other languages\. In theText Renderingtask, Sinitic and Semitic languages show the lowestPrecisionandQuality, with frequent broken or malformed glyphs \(Figure[4](https://arxiv.org/html/2608.11002#S4.F4)\(c\)\(d\)\)\. By contrast, Germanic and Romance languages perform best, benefiting from Latin\-script familiarity\. These trends expose structural weaknesses in handling non\-Latin scripts\. Language\-dependent Trade\-offs\.Beyond individual metrics, languages also reshape how models balance dimensions \(Figure[4](https://arxiv.org/html/2608.11002#S4.F4)\(e\)\(f\)\)\. ForQwen\-ImageinContent Generation,Toxicity\-Styletrade\-offs vary by language: Hindi and Arabic produce safer but less stylistically consistent images, while English emphasizes coherent aesthetics at the cost of higher cultural bias\. Romance languages maintain a better balance, likely due to closer linguistic and cultural proximity to English\. ForNano BananainText Rendering,theAlignment\-Precisionrelation is language\-dependent\. Germanic and Romance maintain stable precision even at high alignment, whereas Sinitic, Slavic, and Indo\-Aryan degrade sharply—reflecting the complex and dense structure of their scripts\. Overall, these results highlight persistent limitations in multilingual T2I systems’ ability to achieve robustvisual\-linguistic groundingacross diverse writing systems\.

![Refer to caption](https://arxiv.org/html/2608.11002v1/Demographic-Bias.png)Figure 5\.Distribution of race and gender categories across ten languages, computed from Qwen\-Image outputs under the Bias dimension\.
### 4\.3\.Language\-dependent Generation Patterns

Detailed analysis procedures, including automated analysis and statistical estimation methods, are provided in Appendix[E](https://arxiv.org/html/2608.11002#A5)\. Demographic Bias\.As shown in Figure[5](https://arxiv.org/html/2608.11002#S4.F5), our analysis reveals a pronounced demographic bias across all languages, with consistent over\-representation of male subjects and specific racial groups\. Notably, a strong language\-demographic alignment is observed: generated images tend to reflect the dominant ethnic characteristics of each language’s primary regions\. For example, Hindi prompts predominantly yield Indian subjects \(94\.6%\), while Japanese and Korean prompts produce a high proportion of Asian subjects \(over 65%\)\. In contrast, Western languages such as Russian, French, and English show a strong bias toward White\-presenting subjects \(84\.9%–97\.1%\)\. These results suggest that model outputs are shaped by demographic distributions embedded in the training data\.

![Refer to caption](https://arxiv.org/html/2608.11002v1/samples-v2.png)Figure 6\.Language\-dependent cultural tendencies, showing how identical prompts produce culturally specific visual interpretations across languages, reflecting implicit cultural priors associated with each language\.![Refer to caption](https://arxiv.org/html/2608.11002v1/rendering-error.png)Figure 7\.Language\-dependent text rendering errors, showing variations in character correctness and structural fidelity across different writing systems, with distinct error patterns emerging for alphabetic and non\-alphabetic scripts\.![Refer to caption](https://arxiv.org/html/2608.11002v1/fig/tr_error_character.png)

![Refer to caption](https://arxiv.org/html/2608.11002v1/fig/tr_error_position.png)

Figure 8\.Analysis of text rendering errors across writing systems\. \(Top\) Character\-level error rates, where Latin scripts are grouped by character category \(uppercase/lowercase\), and CJK scripts are grouped by stroke count \(character complexity\)\. \(Bottom\) Error rates grouped by relative position in the sequence, where position denotes the normalized character position from start to end, enabling comparison of error patterns across Latin and CJK scripts\.Cultural Tendency\.As illustrated in the representative example in Figure[6](https://arxiv.org/html/2608.11002#S4.F6), the prompt “woman statue with seahorses sitting” shows clear cross\-lingual variation in cultural style: the Hindi version reflects traditional Indian sculptural aesthetics, while the Japanese version aligns with East Asian visual conventions, including regionally suggestive elements such as a carp\.

More importantly, this is not an isolated case\. Based on statistics from Qwen\-Image outputs, among all valid samples, 23\.2% of generated images contain identifiable culture\-specific visual elements\. Once explicit cultural cues appear, they tend to align strongly with the cultural region associated with the prompt language: 79\.6% have a primary culture tag that matches the prompt language, and this proportion further rises to 90\.2% when considering only samples assigned to a specific known culture\.

While prior studies\([68](https://arxiv.org/html/2608.11002#bib.bib75);[1](https://arxiv.org/html/2608.11002#bib.bib77);[17](https://arxiv.org/html/2608.11002#bib.bib76)\)report “Westernization” bias in T2I models under English settings, our multilingual results do not support this\. Western cultural tags account for only 3\.7% of valid samples, and only 1\.2% of non\-Western prompts shift toward Western culture\. Instead, multilingual prompting steers generation toward language\-specific cultural aesthetics, expressed through cues such as writing systems, architecture, clothing, and symbolic objects\.

Rendering Errors\.As shown in Figure[7](https://arxiv.org/html/2608.11002#S4.F7), rendering errors vary substantially across writing systems, with fundamentally different failure modes in Latin \(English, French, Spanish\) and CJK \(Chinese, Japanese, Korean\) scripts due to their distinct linguistic and structural properties\. In the top panels of Figure[8](https://arxiv.org/html/2608.11002#S4.F8), character\-level errors exhibit clear category\-specific patterns\. In Latin, errors are concentrated in the lowercase bucket\-especially for Nano Banana\-indicating unstable case consistency despite largely preserved character identity\. In contrast, CJK error rates correlate strongly with stroke count: EasyText remains stable until high\-complexity thresholds, whereas Nano Banana shows consistently high error rates across all stroke levels, suggesting sensitivity to glyph complexity\. The bottom panels of Figure[8](https://arxiv.org/html/2608.11002#S4.F8)further reveal positional differences\. Latin rendering follows a “stable prefix, fragile suffix” pattern, with errors accumulating toward the end of the sequence\. By contrast, CJK shows a breakdown of sequence integrity: EasyText degrades after initial positions, while NanoBanana exhibits high error rates from the outset\. Overall, Latin errors reflect gradual positional drift, whereas CJK errors indicate structural collapse of the sequence\.

### 4\.4\.Causal Analysis

Table 3\.Text\-rendering precision across languages for EasyText and Nano Banana\.![Refer to caption](https://arxiv.org/html/2608.11002v1/re_failure.png)Figure 9\.Representative failure patterns in multilingual text rendering\.Failure Pattern Analysis\.For the text rendering task, beyond the language\-level precision scores in Table[3](https://arxiv.org/html/2608.11002#S4.T3), we identify three failure patterns and use GPT\-5 to estimate their frequencies \(Figure[9](https://arxiv.org/html/2608.11002#S4.F9)\)\. At the model level, the glyph\-conditioned pipeline EasyText is slightly dominated by script\-specific structural errors \(39\.2% of erroneous outputs; 36\.9% semantic substitution\), whereas the semantic\-prior model NanoBanana is dominated by semantic substitution \(50\.0%; 40\.5% structural errors\), with script\-selection or romanization failure less frequent overall \(23\.9% vs\. 9\.5%\)\. Together, these patterns indicate that exact\-string rendering is particularly challenging for non\-Latin scripts: glyph\-conditioned models are more susceptible to structural errors, whereas semantic\-prior models tend to preserve meaning while failing to reproduce the requested string\.

Transliteration Control\.To determine whether non\-Latin scripts drive the cross\-lingual alignment gap, we replace the original scripts with Latin transliterations for 450 Arabic, Hindi, and Chinese samples from the Content Alignment \(TA\-C\) dimension\. Transliteration did not improve content alignment; scores dropped from 0\.78 to 0\.49 on average \(AR: 0\.80→\\rightarrow0\.30, HI: 0\.68→\\rightarrow0\.60, ZH: 0\.87→\\rightarrow0\.57\)\. These results indicate that cross\-lingual alignment depends on more than the surface form of the writing system\. The degradation after transliteration is consistent with limitations in language\-specific text representations and uneven multilingual training coverage\.

![Refer to caption](https://arxiv.org/html/2608.11002v1/alignment_bucket_bias_heatmap.png)Figure 10\.Bias scores across content\-alignment buckets\.Alignment\-conditioned Bias\.To separate genuine demographic bias from errors caused by weak semantic alignment, we group images from the Bias dimension into shared CLIPScore intervals and recompute the bias score within each alignment range\. As shown in Figure[10](https://arxiv.org/html/2608.11002#S4.F10), lower bias scores are concentrated in poorly aligned samples, indicating stronger demographic imbalance when generated content fails to reflect the prompt faithfully\. Although bias scores increase with alignment, demographic imbalance remains evident in highly aligned samples\. These results show that weak alignment amplifies demographic imbalance, while cross\-lingual bias persists after controlling for alignment\.

Culture\-tag Analysis\.To separate the effect of prompt language from explicit cultural conditioning, we compare three versions of the same English source prompts: translated non\-English prompts, English prompts with an explicit culture tag \(e\.g\., “in Hindi style”\), and culture\-neutral English prompts\. Across 600 Qwen\-Image samples, the target\-culture rates are 32\.5%, 75\.0%, and 0\.0%, respectively\. These results show that prompt language directly steers cultural visual tendencies, while explicit culture tags impose a stronger cultural prior\.

Figure 11\.Prompt fragmentation and generation quality across languages for Qwen\-Image\.Tokenization Analysis\.To examine the relationship between text\-side representation efficiency and multilingual generation, we compute the mean prompt\-fragmentation score for each language using the Qwen\-Image tokenizer and compare it with the corresponding mean CLIPScore\. As shown in Figure[11](https://arxiv.org/html/2608.11002#S4.F11), prompt fragmentation exhibits a strong negative Spearman rank correlation with CLIPScore across the ten languages \(ρ=−0\.89\\rho=\-0\.89\)\. Languages represented by more fragmented token sequences consistently achieve weaker image–text alignment\. This result identifies inefficient tokenization as a systematic text\-side bottleneck underlying the performance gap of non\-Latin languages\.

## 5\.Limitation and Insight

Limitation\.The proposed LingT2I benchmark inevitably involves translation, which can introduce bias; however, we apply strict verification and human checks to minimize such effects\. Similarly, while existing metrics are not fully language\-agnostic, we adopt multilingual encoders and adapt protocols to improve fairness\. Importantly, the observed performance gaps are large and consistent, and are therefore unlikely to be explained by these factors\.

Insight\.Our findings suggest several directions for future research\.First,multilingual capability should be achieved through native architectural design rather than post\-hoc adaptation\.Second,training data should be organized by language family and curated with cultural grounding\. Third, models should maintain balanced performance across evaluation dimensions, avoiding over\-optimization toward a single aspect of quality\.In addition,bias\-aware data curation and translation\-based augmentation may help mitigate cultural and demographic biases and improve cross\-lingual fairness\.Finally,given the challenges across writing systems, models could benefit from script\-specific rendering modules or training strategies tailored to their structural characteristics\.

###### Acknowledgements\.

This research was funded by Khalifa University of Science and Technology through the Faculty Start\-Ups under Project ID: KU\-INT\-FSU\-2005\-8474000775\.

## References

- Barveet al\.\(2025\)S\. Barve, A\. Mao, J\. M\. Shi, P\. Juneja, and K\. SahaCan we debias social stereotypes in ai\-generated images? examining text\-to\-image outputs and user perceptions\.arXiv preprint arXiv:2505\.20692\.Cited by:[§4\.3](https://arxiv.org/html/2608.11002#S4.SS3.p4.1)\.
- Basuet al\.\(2023\)A\. Basu, R\. V\. Babu, and D\. PruthiInspecting the geographical representativeness of images from text\-to\-image models\.InProceedings of the IEEE/CVF International Conference on Computer Vision,pp\. 5136–5147\.Cited by:[§2](https://arxiv.org/html/2608.11002#S2.p4.1)\.
- Bianchiet al\.\(2023\)F\. Bianchi, P\. Kalluri, E\. Durmus, F\. Ladhak, M\. Cheng, D\. Nozza, T\. Hashimoto, D\. Jurafsky, J\. Zou, and A\. CaliskanEasily accessible text\-to\-image generation amplifies demographic stereotypes at large scale\.InProceedings of the 2023 ACM conference on fairness, accountability, and transparency,pp\. 1493–1504\.Cited by:[§2](https://arxiv.org/html/2608.11002#S2.p4.1)\.
- Blasiet al\.\(2022\)D\. Blasi, A\. Anastasopoulos, and G\. NeubigSystematic inequalities in language technology performance across the world’s languages\.InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 5486–5505\.Cited by:[§2](https://arxiv.org/html/2608.11002#S2.p4.1)\.
- Central Intelligence Agency \(2025\)Central Intelligence AgencyThe world factbook\.Note:[https://www\.cia\.gov/the\-world\-factbook/](https://www.cia.gov/the-world-factbook/)Accessed: 2025\-09\-08Cited by:[§3\.1](https://arxiv.org/html/2608.11002#S3.SS1.p2.1),[Table 1](https://arxiv.org/html/2608.11002#S3.T1)\.
- Chan \(2016\)K\. L\. ChanPower language index\.Which are the world’s most influential languages\.Cited by:[§3\.1](https://arxiv.org/html/2608.11002#S3.SS1.p2.1),[Table 1](https://arxiv.org/html/2608.11002#S3.T1)\.
- Chenet al\.\(2024\)J\. Chen, C\. Ge, E\. Xie, Y\. Wu, L\. Yao, X\. Ren, Z\. Wang, P\. Luo, H\. Lu, and Z\. LiPixart\-σ\\sigma: weak\-to\-strong training of diffusion transformer for 4k text\-to\-image generation\.InEuropean Conference on Computer Vision,pp\. 74–91\.Cited by:[§D\.1\.1](https://arxiv.org/html/2608.11002#A4.SS1.SSS1.p1.1.20),[Table 9](https://arxiv.org/html/2608.11002#A4.T9.4.1.14.1),[§2](https://arxiv.org/html/2608.11002#S2.p1.1),[Table 2](https://arxiv.org/html/2608.11002#S3.T2.12.1.7.1)\.
- Chenet al\.\(2025\)X\. Chen, Z\. Wu, X\. Liu, Z\. Pan, W\. Liu, Z\. Xie, X\. Yu, and C\. RuanJanus\-pro: unified multimodal understanding and generation with data and model scaling\.arXiv preprint arXiv:2501\.17811\.Cited by:[§D\.1\.1](https://arxiv.org/html/2608.11002#A4.SS1.SSS1.p1.1.10),[Table 8](https://arxiv.org/html/2608.11002#A4.T8.4.1.36.1),[§2](https://arxiv.org/html/2608.11002#S2.p1.1),[Table 2](https://arxiv.org/html/2608.11002#S3.T2.12.1.8.1)\.
- Chenet al\.\(2015\)X\. Chen, H\. Fang, T\. Lin, R\. Vedantam, S\. Gupta, P\. Dollár, and C\. L\. ZitnickMicrosoft coco captions: data collection and evaluation server\.arXiv preprint arXiv:1504\.00325\.Cited by:[§1](https://arxiv.org/html/2608.11002#S1.p1.1)\.
- Chenet al\.\(2023\)Z\. Chen, G\. Liu, B\. Zhang, Q\. Yang, and L\. WuAltclip: altering the language encoder in clip for extended language capabilities\.InFindings of the Association for Computational Linguistics: ACL 2023,pp\. 8666–8682\.Cited by:[§2](https://arxiv.org/html/2608.11002#S2.p2.1)\.
- Chinchureet al\.\(2024\)A\. Chinchure, P\. Shukla, G\. Bhatt, K\. Salij, K\. Hosanagar, L\. Sigal, and M\. TurkTibet: identifying and evaluating biases in text\-to\-image generative models\.InEuropean Conference on Computer Vision,pp\. 429–446\.Cited by:[§2](https://arxiv.org/html/2608.11002#S2.p4.1)\.
- Chuanget al\.\(2025\)Y\. Chuang, Y\. Li, D\. Wang, C\. Yeh, K\. Lyu, R\. Raghavendra, J\. Glass, L\. Huang, J\. Weston, L\. Zettlemoyer,et al\.Meta clip 2: a worldwide scaling recipe\.arXiv preprint arXiv:2507\.22062\.Cited by:[§3\.2](https://arxiv.org/html/2608.11002#S3.SS2.p5.1),[§4](https://arxiv.org/html/2608.11002#S4.p1.1)\.
- Comaniciet al\.\(2025\)G\. Comanici, E\. Bieber, M\. Schaekermann, I\. Pasupat, N\. Sachdeva, I\. Dhillon, M\. Blistein, O\. Ram, D\. Zhang, E\. Rosen,et al\.Gemini 2\.5: pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities\.arXiv preprint arXiv:2507\.06261\.Cited by:[§3\.2](https://arxiv.org/html/2608.11002#S3.SS2.p2.1),[§3\.2](https://arxiv.org/html/2608.11002#S3.SS2.p3.1),[§3\.3](https://arxiv.org/html/2608.11002#S3.SS3.p2.1),[§4](https://arxiv.org/html/2608.11002#S4.p1.1)\.
- DeepSeek\-AI \(2024\)DeepSeek\-AIDeepSeek llm: scaling open\-source language models with longtermism\.arXiv preprint arXiv:2401\.02954\.External Links:[Link](https://github.com/deepseek-ai/DeepSeek-LLM)Cited by:[§2](https://arxiv.org/html/2608.11002#S2.p1.1)\.
- Derakhshaniet al\.\(2025\)M\. M\. Derakhshani, D\. Varghese, M\. Fadaee, and C\. G\. SnoekNeoBabel: a multilingual open tower for visual generation\.arXiv preprint arXiv:2507\.06137\.Cited by:[§2](https://arxiv.org/html/2608.11002#S2.p5.1)\.
- Eberhardet al\.\(2025\)D\. M\. Eberhard, G\. F\. Simons, and C\. D\. FennigEthnologue: languages of the world\(Website\)SIL International\.External Links:[Link](https://www.ethnologue.com/)Cited by:[§3\.1](https://arxiv.org/html/2608.11002#S3.SS1.p2.1),[Table 1](https://arxiv.org/html/2608.11002#S3.T1)\.
- Elsharifet al\.\(2025\)W\. Elsharif, M\. Alzubaidi, and M\. AgusCultural bias in text\-to\-image models: a systematic review of bias identification, evaluation, and mitigation strategies\.IEEE Access13\(\),pp\. 122636–122659\.External Links:[Document](https://dx.doi.org/10.1109/ACCESS.2025.3585745)Cited by:[§2](https://arxiv.org/html/2608.11002#S2.p4.1),[§4\.3](https://arxiv.org/html/2608.11002#S4.SS3.p4.1)\.
- Friedrichet al\.\(2025\)F\. Friedrich, K\. Hämmerl, P\. Schramowski, M\. Brack, J\. Libovickỳ, A\. Fraser, and K\. KerstingMultilingual text\-to\-image generation magnifies gender stereotypes\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 19656–19679\.Cited by:[§1](https://arxiv.org/html/2608.11002#S1.p2.1),[§2](https://arxiv.org/html/2608.11002#S2.p5.1)\.
- Gaoet al\.\(2024\)P\. Gao, L\. Zhuo, D\. Liu, R\. Du, X\. Luo, L\. Qiu, Y\. Zhang, C\. Lin, R\. Huang, S\. Geng,et al\.Lumina\-t2x: transforming text into any modality, resolution, and duration via flow\-based large diffusion transformers\.arXiv preprint arXiv:2405\.05945\.Cited by:[§D\.1\.1](https://arxiv.org/html/2608.11002#A4.SS1.SSS1.p1.1.35),[§2](https://arxiv.org/html/2608.11002#S2.p1.1),[Table 2](https://arxiv.org/html/2608.11002#S3.T2.12.1.9.1)\.
- Hallet al\.\(2023\)M\. Hall, C\. Ross, A\. Williams, N\. Carion, M\. Drozdzal, and A\. R\. SorianoDig in: evaluating disparities in image generations with indicators for geographic diversity\.arXiv preprint arXiv:2308\.06198\.Cited by:[§2](https://arxiv.org/html/2608.11002#S2.p4.1)\.
- Hartvigsenet al\.\(2022\)T\. Hartvigsen, S\. Gabriel, H\. Palangi, M\. Sap, D\. Ray, and E\. KamarToxigen: a large\-scale machine\-generated dataset for adversarial and implicit hate speech detection\.InProceedings of the 60th annual meeting of the association for computational linguistics \(volume 1: Long papers\),pp\. 3309–3326\.Cited by:[§3\.2](https://arxiv.org/html/2608.11002#S3.SS2.p2.1)\.
- Hesselet al\.\(2021\)J\. Hessel, A\. Holtzman, M\. Forbes, R\. L\. Bras, and Y\. ChoiCLIPScore: a reference\-free evaluation metric for image captioning\.InEMNLP,Cited by:[§3\.2](https://arxiv.org/html/2608.11002#S3.SS2.p5.1)\.
- Holtermannet al\.\(2026\)C\. Holtermann, F\. Schneider, and A\. LauscherSoS: analysis of surface over semantics in multilingual text\-to\-image generation\.InProceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 3955–3995\.Cited by:[§1](https://arxiv.org/html/2608.11002#S1.p2.1),[§2](https://arxiv.org/html/2608.11002#S2.p5.1)\.
- Huet al\.\(2020\)J\. Hu, S\. Ruder, A\. Siddhant, G\. Neubig, O\. Firat, and M\. JohnsonXtreme: a massively multilingual multi\-task benchmark for evaluating cross\-lingual generalisation\.InInternational conference on machine learning,pp\. 4411–4421\.Cited by:[§2](https://arxiv.org/html/2608.11002#S2.p4.1)\.
- Joshiet al\.\(2020\)P\. Joshi, S\. Santy, A\. Budhiraja, K\. Bali, and M\. ChoudhuryThe state and fate of linguistic diversity and inclusion in the nlp world\.InProceedings of the 58th annual meeting of the association for computational linguistics,pp\. 6282–6293\.Cited by:[§1](https://arxiv.org/html/2608.11002#S1.p1.1),[§2](https://arxiv.org/html/2608.11002#S2.p4.1)\.
- Kakebayashi and Mori \(2026\)R\. Kakebayashi and T\. MoriPoster: why do non\-english languages exhibit higher vulnerability to data poisoning attacks against text\-to\-image models?\.The Network and Distributed System Security \(NDSS\) Symposium\.Cited by:[§2](https://arxiv.org/html/2608.11002#S2.p5.1)\.
- Kannenet al\.\(2024\)N\. Kannen, A\. Ahmad, M\. Andreetto, V\. Prabhakaran, U\. Prabhu, A\. B\. Dieng, P\. Bhattacharyya, and S\. DaveBeyond aesthetics: cultural competence in text\-to\-image models\.Advances in Neural Information Processing Systems37,pp\. 13716–13747\.Cited by:[§2](https://arxiv.org/html/2608.11002#S2.p4.1)\.
- Klassertet al\.\(2026\)T\. Klassert, A\. Ulges, and B\. FuBAFIS: dataset\+ framework to assess occupational bias and human preference in modern text\-to\-image models\.InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision,pp\. 2168–2177\.Cited by:[§2](https://arxiv.org/html/2608.11002#S2.p4.1)\.
- Labs \(2024\)B\. F\. LabsFLUX\.Note:[https://github\.com/black\-forest\-labs/flux](https://github.com/black-forest-labs/flux)Cited by:[§D\.1\.1](https://arxiv.org/html/2608.11002#A4.SS1.SSS1.p1.1.17),[§D\.1\.2](https://arxiv.org/html/2608.11002#A4.SS1.SSS2.p1.1.5),[Table 10](https://arxiv.org/html/2608.11002#A4.T10.4.1.14.1),[Table 9](https://arxiv.org/html/2608.11002#A4.T9.4.1.3.1),[§1](https://arxiv.org/html/2608.11002#S1.p2.1),[§2](https://arxiv.org/html/2608.11002#S2.p1.1),[§2](https://arxiv.org/html/2608.11002#S2.p3.1),[Table 2](https://arxiv.org/html/2608.11002#S3.T2.12.1.22.1),[Table 2](https://arxiv.org/html/2608.11002#S3.T2.12.1.5.1)\.
- Lcvenshtcin \(1966\)V\. LcvenshtcinBinary coors capable or ‘correcting deletions, insertions, and reversals\.InSoviet physics\-doklady,Vol\.10\.Cited by:[§3\.3](https://arxiv.org/html/2608.11002#S3.SS3.p2.1)\.
- Leeet al\.\(2023\)T\. Lee, M\. Yasunaga, C\. Meng, Y\. Mai, J\. S\. Park, A\. Gupta, Y\. Zhang, D\. Narayanan, H\. Teufel, M\. Bellagente,et al\.Holistic evaluation of text\-to\-image models\.Advances in Neural Information Processing Systems36,pp\. 69981–70011\.Cited by:[§B\.1](https://arxiv.org/html/2608.11002#A2.SS1.p3.1),[§3\.2](https://arxiv.org/html/2608.11002#S3.SS2.p1.1)\.
- Liuet al\.\(2025\)C\. C\. Liu, I\. Gurevych, and A\. KorhonenCulturally aware and adapted nlp: a taxonomy and a survey of the state of the art\.Transactions of the Association for Computational Linguistics13,pp\. 652–689\.Cited by:[§1](https://arxiv.org/html/2608.11002#S1.p1.1)\.
- Liuet al\.\(2024a\)Z\. Liu, W\. Liang, Z\. Liang, C\. Luo, J\. Li, G\. Huang, and Y\. YuanGlyph\-byt5: a customized text encoder for accurate visual text rendering\.InEuropean Conference on Computer Vision,pp\. 361–377\.Cited by:[§2](https://arxiv.org/html/2608.11002#S2.p3.1)\.
- Liuet al\.\(2024b\)Z\. Liu, W\. Liang, Y\. Zhao, B\. Chen, L\. Liang, L\. Wang, J\. Li, and Y\. YuanGlyph\-byt5\-v2: a strong aesthetic baseline for accurate multilingual visual text rendering\.arXiv preprint arXiv:2406\.10208\.Cited by:[§2](https://arxiv.org/html/2608.11002#S2.p3.1)\.
- Luet al\.\(2026\)R\. Lu, Y\. Zhang, J\. Liu, H\. Wang, and Y\. SongEasytext: controllable diffusion transformer for multilingual text rendering\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.40,pp\. 7565–7573\.Cited by:[§D\.1\.2](https://arxiv.org/html/2608.11002#A4.SS1.SSS2.p1.1.13),[§1](https://arxiv.org/html/2608.11002#S1.p5.1),[§2](https://arxiv.org/html/2608.11002#S2.p3.1),[§3\.3](https://arxiv.org/html/2608.11002#S3.SS3.p2.1),[§3\.3](https://arxiv.org/html/2608.11002#S3.SS3.p3.1),[Table 2](https://arxiv.org/html/2608.11002#S3.T2.12.1.27.1)\.
- Maet al\.\(2024\)J\. Ma, C\. Chen, Q\. Xie, and H\. LuPea\-diffusion: parameter\-efficient adapter with knowledge distillation in non\-english text\-to\-image generation\.InEuropean Conference on Computer Vision,pp\. 89–105\.Cited by:[§D\.1\.1](https://arxiv.org/html/2608.11002#A4.SS1.SSS1.p1.1.23),[Table 9](https://arxiv.org/html/2608.11002#A4.T9.4.1.25.1),[§2](https://arxiv.org/html/2608.11002#S2.p2.1),[Table 2](https://arxiv.org/html/2608.11002#S3.T2.12.1.15.1)\.
- Maet al\.\(2025\)J\. Ma, Q\. Peng, X\. Guo, C\. Chen, H\. Lu, and Z\. YangX2i: seamless integration of multimodal understanding into diffusion transformer via attention distillation\.InProceedings of the IEEE/CVF International Conference on Computer Vision,pp\. 16733–16744\.Cited by:[§D\.1\.1](https://arxiv.org/html/2608.11002#A4.SS1.SSS1.p1.1.27),[Table 9](https://arxiv.org/html/2608.11002#A4.T9.4.1.36.1),[§2](https://arxiv.org/html/2608.11002#S2.p2.1),[Table 2](https://arxiv.org/html/2608.11002#S3.T2.12.1.16.1)\.
- Mittalet al\.\(2024\)S\. Mittal, A\. Sudan, M\. Vatsa, R\. Singh, T\. Glaser, and T\. HassnerNavigating text\-to\-image generative bias across indic languages\.InEuropean Conference on Computer Vision,pp\. 53–67\.Cited by:[§2](https://arxiv.org/html/2608.11002#S2.p5.1)\.
- Nayaket al\.\(2025\)S\. Nayak, M\. Bhatia, X\. Zhang, V\. Rieser, L\. A\. Hendricks, S\. Van Steenkiste, Y\. Goyal, K\. Stańczak, and A\. AgrawalCulturalframes: assessing cultural expectation alignment in text\-to\-image models and evaluation metrics\.InFindings of the Association for Computational Linguistics: EMNLP 2025,pp\. 20918–20953\.Cited by:[§2](https://arxiv.org/html/2608.11002#S2.p4.1)\.
- Onoeet al\.\(2024\)Y\. Onoe, S\. Rane, Z\. Berger, Y\. Bitton, J\. Cho, R\. Garg, A\. Ku, Z\. Parekh, J\. Pont\-Tuset, G\. Tanzer,et al\.Docci: descriptions of connected and contrasting images\.InEuropean Conference on Computer Vision,pp\. 291–309\.Cited by:[§3\.2](https://arxiv.org/html/2608.11002#S3.SS2.p1.1)\.
- Philippyet al\.\(2023\)F\. Philippy, S\. Guo, and S\. HaddadanTowards a common understanding of contributing factors for cross\-lingual transfer in multilingual language models: a review\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 5877–5891\.Cited by:[§2](https://arxiv.org/html/2608.11002#S2.p4.1)\.
- Podellet al\.\(2024\)D\. Podell, Z\. English, K\. Lacey, A\. Blattmann, T\. Dockhorn, J\. Müller, J\. Penna, and R\. RombachSdxl: improving latent diffusion models for high\-resolution image synthesis\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 1862–1874\.Cited by:[§D\.1\.1](https://arxiv.org/html/2608.11002#A4.SS1.SSS1.p1.1.4),[Table 8](https://arxiv.org/html/2608.11002#A4.T8.4.1.14.1.1),[Table 2](https://arxiv.org/html/2608.11002#S3.T2.12.1.4.1)\.
- Qinet al\.\(2023\)C\. Qin, N\. Yu, C\. Xing, S\. Zhang, Z\. Chen, S\. Ermon, Y\. Fu, C\. Xiong, and R\. XuGluegen: plug and play multi\-modal encoders for x\-to\-image generation\.InProceedings of the IEEE/CVF international conference on computer vision,pp\. 23085–23096\.Cited by:[§2](https://arxiv.org/html/2608.11002#S2.p2.1)\.
- Qinet al\.\(2025\)L\. Qin, Q\. Chen, Y\. Zhou, Z\. Chen, Y\. Li, L\. Liao, M\. Li, W\. Che, and P\. S\. YuA survey of multilingual large language models\.Patterns6\(1\)\.Cited by:[§2](https://arxiv.org/html/2608.11002#S2.p4.1)\.
- Qiuet al\.\(2022\)C\. Qiu, D\. Oneață, E\. Bugliarello, S\. Frank, and D\. ElliottMultilingual multimodal learning with machine translated text\.InFindings of the Association for Computational Linguistics: EMNLP 2022,Y\. Goldberg, Z\. Kozareva, and Y\. Zhang \(Eds\.\),Abu Dhabi, United Arab Emirates,pp\. 4178–4193\.External Links:[Link](https://aclanthology.org/2022.findings-emnlp.308/),[Document](https://dx.doi.org/10.18653/v1/2022.findings-emnlp.308)Cited by:[§2](https://arxiv.org/html/2608.11002#S2.p4.1)\.
- Radfordet al\.\(2021\)A\. Radford, J\. W\. Kim, C\. Hallacy, A\. Ramesh, G\. Goh, S\. Agarwal, G\. Sastry, A\. Askell, P\. Mishkin, J\. Clark,et al\.Learning transferable visual models from natural language supervision\.InInternational conference on machine learning,pp\. 8748–8763\.Cited by:[§3\.2](https://arxiv.org/html/2608.11002#S3.SS2.p5.1)\.
- Raffelet al\.\(2020\)C\. Raffel, N\. Shazeer, A\. Roberts, K\. Lee, S\. Narang, M\. Matena, Y\. Zhou, W\. Li, and P\. J\. LiuExploring the limits of transfer learning with a unified text\-to\-text transformer\.Journal of machine learning research21\(140\),pp\. 1–67\.Cited by:[§1](https://arxiv.org/html/2608.11002#S1.p2.1),[§2](https://arxiv.org/html/2608.11002#S2.p1.1)\.
- Rajaee and Monz \(2024\)S\. Rajaee and C\. MonzAnalyzing the evaluation of cross\-lingual knowledge transfer in multilingual language models\.InProceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 2895–2914\.Cited by:[§2](https://arxiv.org/html/2608.11002#S2.p4.1)\.
- Ranathunga and De Silva \(2022\)S\. Ranathunga and N\. De SilvaSome languages are more equal than others: probing deeper into the linguistic disparity in the nlp world\.InProceedings of the 2nd Conference of the Asia\-Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing \(Volume 1: Long Papers\),pp\. 823–848\.Cited by:[§2](https://arxiv.org/html/2608.11002#S2.p4.1)\.
- Saxon and Wang \(2023\)M\. Saxon and W\. Y\. WangMultilingual conceptual coverage in text\-to\-image models\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 4831–4848\.Cited by:[§1](https://arxiv.org/html/2608.11002#S1.p2.1),[§2](https://arxiv.org/html/2608.11002#S2.p5.1)\.
- Schuhmannet al\.\(2022\)C\. Schuhmann, R\. Beaumont, R\. Vencu, C\. Gordon, R\. Wightman, M\. Cherti, T\. Coombes, A\. Katta, C\. Mullis, M\. Wortsman,et al\.Laion\-5b: an open large\-scale dataset for training next generation image\-text models\.Advances in neural information processing systems35,pp\. 25278–25294\.Cited by:[§1](https://arxiv.org/html/2608.11002#S1.p1.1)\.
- Shaniet al\.\(2026\)C\. Shani, Y\. Reif, N\. Roll, D\. Jurafsky, and E\. ShutovaThe roots of performance disparity in multilingual language models: intrinsic modeling difficulty or design choices?\.arXiv preprint arXiv:2601\.07220\.Cited by:[§2](https://arxiv.org/html/2608.11002#S2.p4.1)\.
- Shiet al\.\(2020\)Z\. Shi, X\. Zhou, X\. Qiu, and X\. ZhuImproving image captioning with better use of caption\.InProceedings of the 58th annual meeting of the association for computational linguistics,pp\. 7454–7464\.Cited by:[§2](https://arxiv.org/html/2608.11002#S2.p1.1)\.
- Stability AI \(2024\)Stability AIStable diffusion 3\.5\.External Links:[Link](https://github.com/Stability-AI/sd3.5)Cited by:[§D\.1\.1](https://arxiv.org/html/2608.11002#A4.SS1.SSS1.p1.1.1),[Table 8](https://arxiv.org/html/2608.11002#A4.T8.4.1.3.1.1),[§2](https://arxiv.org/html/2608.11002#S2.p1.1),[Table 2](https://arxiv.org/html/2608.11002#S3.T2.12.1.3.1)\.
- Struppeket al\.\(2023\)L\. Struppek, D\. Hintersdorf, F\. Friedrich, P\. Schramowski, K\. Kersting,et al\.Exploiting cultural biases via homoglyphs in text\-to\-image synthesis\.Journal of Artificial Intelligence Research78,pp\. 1017–1068\.Cited by:[§2](https://arxiv.org/html/2608.11002#S2.p5.1)\.
- Tanet al\.\(2024\)Z\. Tan, M\. Yang, L\. Qin, H\. Yang, Y\. Qian, Q\. Zhou, C\. Zhang, and H\. LiAn empirical study and analysis of text\-to\-image generation using large language model\-powered textual representation\.InEuropean Conference on Computer Vision,pp\. 472–489\.Cited by:[§D\.1\.1](https://arxiv.org/html/2608.11002#A4.SS1.SSS1.p1.1.41),[§2](https://arxiv.org/html/2608.11002#S2.p1.1),[Table 2](https://arxiv.org/html/2608.11002#S3.T2.12.1.12.1)\.
- Teamet al\.\(2024\)G\. Team, T\. Mesnard, C\. Hardin, R\. Dadashi, S\. Bhupatiraju, S\. Pathak, L\. Sifre, M\. Rivière, M\. S\. Kale, J\. Love,et al\.Gemma: open models based on gemini research and technology\.arXiv preprint arXiv:2403\.08295\.Cited by:[§2](https://arxiv.org/html/2608.11002#S2.p1.1)\.
- Team \(2025a\)G\. G\. TeamNano banana: gemini ai image generator & photo editor\.Note:[https://gemini\.google/overview/image\-generation/](https://gemini.google/overview/image-generation/)Cited by:[§1](https://arxiv.org/html/2608.11002#S1.p1.1),[§1](https://arxiv.org/html/2608.11002#S1.p5.1),[§2](https://arxiv.org/html/2608.11002#S2.p3.1),[Table 2](https://arxiv.org/html/2608.11002#S3.T2.12.1.20.1)\.
- Teamet al\.\(2025\)N\. Team, C\. Han, G\. Li, J\. Wu, Q\. Sun, Y\. Cai, Y\. Peng, Z\. Ge, D\. Zhou, H\. Tang, H\. Zhou, K\. Liu, A\. Huang, B\. Wang, C\. Miao, D\. Sun, E\. Yu, F\. Yin, G\. Yu, H\. Nie, H\. Lv, H\. Hu, J\. Wang, J\. Zhou, J\. Sun, K\. Tan, K\. An, K\. Lin, L\. Zhao, M\. Chen, P\. Xing, R\. Wang, S\. Liu, S\. Xia, T\. You, W\. Ji, X\. Zeng, X\. Han, X\. Zhang, Y\. Wei, Y\. Xu, Y\. Jiang, Y\. Wang, Y\. Zhou, Y\. Han, Z\. Meng, B\. Jiao, D\. Jiang, X\. Zhang, and Y\. ZhuNextStep\-1: toward autoregressive image generation with continuous tokens at scale\.arXiv preprint arXiv:2508\.10711\.Cited by:[§2](https://arxiv.org/html/2608.11002#S2.p1.1)\.
- Team \(2025b\)Q\. TeamQwen2\.5\-vl\.External Links:[Link](https://qwenlm.github.io/blog/qwen2.5-vl/)Cited by:[§3\.2](https://arxiv.org/html/2608.11002#S3.SS2.p6.1),[§4](https://arxiv.org/html/2608.11002#S4.p1.1)\.
- Team \(2025c\)T\. H\. TeamHunyuanImage 3\.0: technical report\.Note:[https://github\.com/Tencent\-Hunyuan/HunyuanImage\-3\.0](https://github.com/Tencent-Hunyuan/HunyuanImage-3.0)Cited by:[§1](https://arxiv.org/html/2608.11002#S1.p2.1),[§2](https://arxiv.org/html/2608.11002#S2.p1.1)\.
- Team \(2025d\)Z\. TeamZ\-image: an efficient image generation foundation model with single\-stream diffusion transformer\.arXiv preprint arXiv:2511\.22699\.Cited by:[§D\.1\.1](https://arxiv.org/html/2608.11002#A4.SS1.SSS1.p1.1.38),[§1](https://arxiv.org/html/2608.11002#S1.p5.1),[§2](https://arxiv.org/html/2608.11002#S2.p1.1),[Table 2](https://arxiv.org/html/2608.11002#S3.T2.12.1.10.1)\.
- Tencent Hunyuan Team \(2024\)Tencent Hunyuan TeamHunyuan\-a13b\.Note:[https://github\.com/Tencent\-Hunyuan/Hunyuan\-A13B](https://github.com/Tencent-Hunyuan/Hunyuan-A13B)Cited by:[§2](https://arxiv.org/html/2608.11002#S2.p1.1)\.
- Tschannenet al\.\(2025\)M\. Tschannen, A\. Gritsenko, X\. Wang, M\. F\. Naeem, I\. Alabdulmohsin, N\. Parthasarathy, T\. Evans, L\. Beyer, Y\. Xia, B\. Mustafa,et al\.Siglip 2: multilingual vision\-language encoders with improved semantic understanding, localization, and dense features\.arXiv preprint arXiv:2502\.14786\.Cited by:[§1](https://arxiv.org/html/2608.11002#S1.p2.1),[§2](https://arxiv.org/html/2608.11002#S2.p1.1)\.
- Tuoet al\.\(2024a\)Y\. Tuo, Y\. Geng, and L\. BoAnytext2: visual text generation and editing with customizable attributes\.arXiv preprint arXiv:2411\.15245\.Cited by:[§D\.1\.2](https://arxiv.org/html/2608.11002#A4.SS1.SSS2.p1.1.10),[Table 10](https://arxiv.org/html/2608.11002#A4.T10.4.1.36.1),[§2](https://arxiv.org/html/2608.11002#S2.p3.1),[Table 2](https://arxiv.org/html/2608.11002#S3.T2.12.1.26.1)\.
- Tuoet al\.\(2024b\)Y\. Tuo, W\. Xiang, J\. He, Y\. Geng, and X\. XieAnytext: multilingual visual text generation and editing\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 56783–56799\.Cited by:[§D\.1\.2](https://arxiv.org/html/2608.11002#A4.SS1.SSS2.p1.1.1),[§D\.1\.2](https://arxiv.org/html/2608.11002#A4.SS1.SSS2.p1.1.7),[Table 10](https://arxiv.org/html/2608.11002#A4.T10.4.1.25.1.1),[§2](https://arxiv.org/html/2608.11002#S2.p3.1),[Table 2](https://arxiv.org/html/2608.11002#S3.T2.12.1.25.1)\.
- Venturaet al\.\(2025\)M\. Ventura, E\. Ben\-David, A\. Korhonen, and R\. ReichartNavigating cultural chasms: exploring and unlocking the cultural pov of text\-to\-image models\.Transactions of the Association for Computational Linguistics13,pp\. 142–166\.Cited by:[§1](https://arxiv.org/html/2608.11002#S1.p2.1),[§2](https://arxiv.org/html/2608.11002#S2.p5.1)\.
- Wanet al\.\(2024\)Y\. Wan, A\. Subramonian, A\. Ovalle, Z\. Lin, A\. Suvarna, C\. Chance, H\. Bansal, R\. Pattichis, and K\. ChangSurvey of bias in text\-to\-image generation: definition, evaluation, and mitigation\.arXiv preprint arXiv:2404\.01030\.Cited by:[§2](https://arxiv.org/html/2608.11002#S2.p4.1),[§4\.3](https://arxiv.org/html/2608.11002#S4.SS3.p4.1)\.
- Wanget al\.\(2022\)J\. Wang, Y\. Liu, and X\. WangAssessing multilingual fairness in pre\-trained multimodal representations\.InFindings of the Association for Computational Linguistics: ACL 2022,pp\. 2681–2695\.Cited by:[§3\.2](https://arxiv.org/html/2608.11002#S3.SS2.p5.1)\.
- Wuet al\.\(2025\)C\. Wu, J\. Li, J\. Zhou, J\. Lin, K\. Gao, K\. Yan, S\. Yin, S\. Bai, X\. Xu, Y\. Chen, Y\. Chen, Z\. Tang, Z\. Zhang, Z\. Wang, A\. Yang, B\. Yu, C\. Cheng, D\. Liu, D\. Li, H\. Zhang, H\. Meng, H\. Wei, J\. Ni, K\. Chen, K\. Cao, L\. Peng, L\. Qu, M\. Wu, P\. Wang, S\. Yu, T\. Wen, W\. Feng, X\. Xu, Y\. Wang, Y\. Zhang, Y\. Zhu, Y\. Wu, Y\. Cai, and Z\. LiuQwen\-image technical report\.External Links:2508\.02324,[Link](https://arxiv.org/abs/2508.02324)Cited by:[§D\.1\.1](https://arxiv.org/html/2608.11002#A4.SS1.SSS1.p1.1.13),[§D\.1\.2](https://arxiv.org/html/2608.11002#A4.SS1.SSS2.p1.1.3),[Table 10](https://arxiv.org/html/2608.11002#A4.T10.4.1.3.1),[§1](https://arxiv.org/html/2608.11002#S1.p1.1),[§1](https://arxiv.org/html/2608.11002#S1.p2.1),[§2](https://arxiv.org/html/2608.11002#S2.p1.1),[§2](https://arxiv.org/html/2608.11002#S2.p3.1),[Table 2](https://arxiv.org/html/2608.11002#S3.T2.12.1.11.1),[Table 2](https://arxiv.org/html/2608.11002#S3.T2.12.1.21.1)\.
- Xieet al\.\(2025\)E\. Xie, J\. Chen, Y\. Zhao, J\. YU, L\. Zhu, Y\. Lin, Z\. Zhang, M\. Li, J\. Chen, H\. Cai, B\. Liu, D\. Zhou, and S\. HanSANA 1\.5: efficient scaling of training\-time and inference\-time compute in linear diffusion transformer\.InForty\-second International Conference on Machine Learning,External Links:[Link](https://openreview.net/forum?id=27hOkXzy9e)Cited by:[§D\.1\.1](https://arxiv.org/html/2608.11002#A4.SS1.SSS1.p1.1.7),[Table 8](https://arxiv.org/html/2608.11002#A4.T8.4.1.25.1),[§2](https://arxiv.org/html/2608.11002#S2.p1.1),[Table 2](https://arxiv.org/html/2608.11002#S3.T2.12.1.6.1)\.
- Xinget al\.\(2025\)S\. Xing, M\. Zhong, Z\. Lai, L\. Li, J\. Liu, Y\. Wang, J\. Dai, and W\. WangMuLan: adapting multilingual diffusion models for hundreds of languages with negligible cost\.InProceedings of the 42nd International Conference on Machine Learning,A\. Singh, M\. Fazel, D\. Hsu, S\. Lacoste\-Julien, F\. Berkenkamp, T\. Maharaj, K\. Wagstaff, and J\. Zhu \(Eds\.\),Proceedings of Machine Learning Research, Vol\.267,pp\. 68953–68969\.Cited by:[§D\.1\.1](https://arxiv.org/html/2608.11002#A4.SS1.SSS1.p1.1.31),[§2](https://arxiv.org/html/2608.11002#S2.p2.1),[Table 2](https://arxiv.org/html/2608.11002#S3.T2.12.1.17.1)\.
- Xueet al\.\(2021\)L\. Xue, N\. Constant, A\. Roberts, M\. Kale, R\. Al\-Rfou, A\. Siddhant, A\. Barua, and C\. RaffelMT5: a massively multilingual pre\-trained text\-to\-text transformer\.InProceedings of the 2021 conference of the North American chapter of the association for computational linguistics: Human language technologies,pp\. 483–498\.Cited by:[§3\.2](https://arxiv.org/html/2608.11002#S3.SS2.p4.1),[§4](https://arxiv.org/html/2608.11002#S4.p1.1)\.
- Yanget al\.\(2024\)Q\. A\. Yang, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Li, D\. Liu, F\. Huang, G\. Dong, H\. Wei, H\. Lin, J\. Yang, J\. Tu, J\. Zhang, J\. Yang, J\. Yang, J\. Zhou, J\. Lin, K\. Dang, K\. Lu, K\. Bao, K\. Yang, L\. Yu, M\. Li, M\. Xue, P\. Zhang, Q\. Zhu, R\. Men, R\. Lin, T\. Li, T\. Xia, X\. Ren, X\. Ren, Y\. Fan, Y\. Su, Y\. Zhang, Y\. Wan, Y\. Liu, Z\. Cui, Z\. Zhang, Z\. Qiu, S\. Quan, and Z\. WangQwen2\.5 technical report\.ArXivabs/2412\.15115\.External Links:[Link](https://api.semanticscholar.org/CorpusID:274859421)Cited by:[§2](https://arxiv.org/html/2608.11002#S2.p1.1)\.
- Yeet al\.\(2024\)F\. Ye, G\. Liu, X\. Wu, and L\. WuAltdiffusion: a multilingual text\-to\-image diffusion model\.InProceedings of the AAAI conference on artificial intelligence,Vol\.38,pp\. 6648–6656\.Cited by:[§2](https://arxiv.org/html/2608.11002#S2.p2.1),[§2](https://arxiv.org/html/2608.11002#S2.p5.1)\.
- Zhanget al\.\(2024\)L\. Zhang, X\. Liao, Z\. Yang, B\. Gao, C\. Wang, Q\. Yang, and D\. LiPartiality and misconception: investigating cultural representativeness in text\-to\-image models\.InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems,CHI ’24,New York, NY, USA\.External Links:ISBN 9798400703300,[Link](https://doi.org/10.1145/3613904.3642877),[Document](https://dx.doi.org/10.1145/3613904.3642877)Cited by:[§2](https://arxiv.org/html/2608.11002#S2.p4.1)\.
- Zhanget al\.\(2025\)S\. Zhang, B\. Xie, Z\. Yan, Y\. Zhang, D\. Zhou, X\. Chen, S\. Qiu, J\. Liu, G\. Xie, and Z\. LuTrade\-offs in image generation: how do different dimensions interact?\.InProceedings of the IEEE/CVF International Conference on Computer Vision,pp\. 17256–17267\.Cited by:[§B\.1](https://arxiv.org/html/2608.11002#A2.SS1.p3.1),[§C\.1\.2](https://arxiv.org/html/2608.11002#A3.SS1.SSS2.p1.1),[§3\.2](https://arxiv.org/html/2608.11002#S3.SS2.p1.1),[§3\.2](https://arxiv.org/html/2608.11002#S3.SS2.p6.1)\.
- Zhou and Lu \(2025\)E\. Zhou and W\. LuBias beyond english: evaluating social bias and debiasing methods in a low\-resource setting\.InCCF International Conference on Natural Language Processing and Chinese Computing,pp\. 214–227\.Cited by:[§2](https://arxiv.org/html/2608.11002#S2.p4.1)\.

## Appendix AImportant Statements

### A\.1\.Social Impact

This work contributes positively to promoting fairness and inclusiveness in artificial intelligence\. First, through a systematic evaluation of text\-to\-image models in multilingual and multicultural settings, we reveal existing linguistic and cultural biases in current generative models, and our framework provides a foundation for global language fairness assessment\.

Second, the LingT2I benchmark and metric suite offer meaningful directions for future work, helping drive the development of models that are more inclusive of linguistic and cultural diversity while improving the visibility and research value of underrepresented languages and cultures\.

Third, our study enhances the transparency of multilingual generation evaluation, providing both academia and industry with a more measurable and interpretable framework, and advancing AI toward being more explainable, fair, and responsible\.

Finally, our work will guide the generative AI research community to better understand and respect linguistic, script, and cultural diversity, helping reduce technical bias and fostering a more inclusive global AI system\.

### A\.2\.Ethical Statement

To avoid the potential social risks,we emphasize that all datasets used in this work comply with their official licenses and community standards, and we strictly adhere to ethical guidelines throughout data usage and research practices\.

Although the evaluation encompasses dimensions of bias and toxicity,we have not introduced any new harmful data\. We only utilized existing research\-purpose datasets and will not directly disclose any toxicity\-related data unless absolutely necessary, and the reproducibility of this evaluation is ensured by the complete scripts and prompts we provide\.

The generation and evaluation in this paper are conducted onlyfor academic research purposes\. Throughout this process, assessments beneficial to enhancing fairness in human society and culture have been performed, yielding positive impact only\.

### A\.3\.LLM Usage

All instances of LLM usage in the research are mentioned clearly in the main text and appendix, including specific models and their detailed usage\.

Besides, LLMs are used to moderately polish the paper writing\. Specifically, LLMs are employed for improving grammar and formatting consistency of LaTeX content\.

## Appendix BDetails of LingT2I Dataset

### B\.1\.Prompt for MLLM in Data Processing

Prompt for Translation\.Figure[12](https://arxiv.org/html/2608.11002#A2.F12)shows the prompt for Gemini\-2\.5\-Pro in the translation progress\. This is the final version, resulting from improvements\.

![Refer to caption](https://arxiv.org/html/2608.11002v1/prompt_translation.png)Figure 12\.Prompt template used for multilingual translation, ensuring semantic consistency, cultural appropriateness, and stylistic fidelity across languages\.Prompt for Content Generation Task Annotation\.Figure[13](https://arxiv.org/html/2608.11002#A2.F13)shows the prompt for Gemini\-2\.5\-Flash in the data annotation process of Content Generation Task\. We construct dimension\-specific prompts using a unified template that guides the annotation model to extract and rewrite dimension\-relevant information from the original caption\. The template takes as input the caption, the target dimension, and its definition, and instructs the model to produce a concise prompt that preserves only the information relevant to the target dimension while removing irrelevant details\.

For dimensions that are not explicitly described in the original caption \(e\.g\., Style and Bias\), we further employ a conditional expansion strategy\. Specifically, the model is instructed to first generate a concise base description and then modify it by incorporating dimension\-specific control signals \(e\.g\., stylistic cues\)\. The example template shown in Figure[13](https://arxiv.org/html/2608.11002#A2.F13)illustrates this process, while the exact design of control phrases and dimension\-specific elements follows prior benchmark practices, particularly HEIM\([31](https://arxiv.org/html/2608.11002#bib.bib25)\)and TRIGScore\([77](https://arxiv.org/html/2608.11002#bib.bib9)\)\.

![Refer to caption](https://arxiv.org/html/2608.11002v1/prompt_anno_cg.png)Figure 13\.Prompt for Content Generation Task Annotation\.
### B\.2\.Dataset Examples and Statistics

Figure[15](https://arxiv.org/html/2608.11002#A2.F15)and Figure[16](https://arxiv.org/html/2608.11002#A2.F16)show the representative prompts across all 10 languages in LingT2I dataset’sContent Generationtask and the corresponding output images from some example models\. Figure[17](https://arxiv.org/html/2608.11002#A2.F17)shows the representative prompts across all 10 languages in LingT2I dataset’sText Renderingtask and the corresponding output images from some example models\.

The detailed dataset statistics of prompt length are shown in Figure[14](https://arxiv.org/html/2608.11002#A2.F14)\.

Figure 14\.Dataset Statistics\.Average token lengths computed using the mT5 tokenizer for Content Generation and Text Rendering tasks\.
### B\.3\.Quality Control

Automatic Processing and Verification\.All automatic processing—including prompt shortening, filtering, augmentation, and translation for both theContent GenerationandText Renderingtasks—is conducted using Gemini 2\.5 Pro\. To ensure translation accuracy and linguistic consistency, we apply multi\-round verification, including back\-translation and cross\-checking with GPT\-5\. During this process, we explicitly enforce constraints to preserve semantic meaning, maintain cultural nuance, and avoid introducing additional bias across languages\. GPT\-5 flags a small portion of samples \(1\.3%\) as problematic, mainly due to minor semantic inconsistencies or cultural ambiguities\. These cases are further reviewed and manually corrected to ensure final data quality\.

Error Analysis and Iterative Refinement\.Before large\-scale data generation, we conduct pilot experiments on a randomly sampled 5% subset of the dataset\. Based on this subset, we perform error analysis to identify common issues such as semantic drift, cultural misalignment, and translation inconsistency\. Guided by these observations, we iteratively refine the prompt construction and filtering process for three rounds, until the data quality is considered stable\.

Human Quality Check\.After finalizing the dataset, we further validate data quality through human evaluation\. We randomly sample 5% of the full dataset and involve native speakers across all target languages, including university students and academic staff\. Annotators are asked to assess semantic fidelity, cultural appropriateness, and fluency of the translated prompts\. Overall, 98% of the samples are judged to be semantically consistent across languages\.

Figure[18](https://arxiv.org/html/2608.11002#A2.F18)shows the prompt for GPT5 to double check our translation\. Table[4](https://arxiv.org/html/2608.11002#A2.T4)shows the improvements made with each iteration and the resulting increase in the accuracy of the random samples\.

Table 4\.Iterative refinement of the translation prompt and its impact on translation quality\. Accuracy is measured on a randomly sampled subset using GPT\-5\-based verification\.![Refer to caption](https://arxiv.org/html/2608.11002v1/case_t2i_iq_r_compressed.png)Figure 15\.Examples forContent Generationtask in Reality dimension\.![Refer to caption](https://arxiv.org/html/2608.11002v1/case_t2i_ta_c_compressed.png)Figure 16\.Examples forContent Generationtask in Content Alignment dimension\.![Refer to caption](https://arxiv.org/html/2608.11002v1/case_tr_compressed.png)Figure 17\.Examples forText Renderingtask\.![Refer to caption](https://arxiv.org/html/2608.11002v1/prompt_quality_control.png)Figure 18\.Prompt template used for translation quality control\.

## Appendix CDetails of Metrics

### C\.1\.Content Generation Metric Settings

#### C\.1\.1\.CLIPScore

For CLIPScore implementation, we use the Facebook metaclip\-2\-worldwide\-huge\-378 official checkpointon huggingface and inject it into the original CLIPScore github codebase\.

#### C\.1\.2\.TRIGScore

Computation \(Dimensions except Robustness\)\.The detailed TRIG Score computation method is as followed, from the original TRIG\([77](https://arxiv.org/html/2608.11002#bib.bib9)\)paper:

For each sample from subset𝒟\\mathcal\{D\}we feed the task description, generated image, prompt, and specific dimensional evaluation criteria into the VLM, instructing it to evaluate the degree from a set of predefined rating tokens\. Formally, let the token set be𝒯=\{t1,t2,…,tn\}\\mathcal\{T\}=\\\{t\_\{1\},t\_\{2\},\\dots,t\_\{n\}\\\}wheretit\_\{i\}represents a semantic rating \(e\.g\., “Good”, “Medium”, “Bad”\)\.

The model output is provided in the form of logits, which can be expressed asℒ=\{\(x,z⁡\(x\)\)∣x∈𝒱\}\\mathcal\{L\}=\\\{\(x,z\(x\)\)\\mid x\\in\\mathcal\{V\}\\\}, where𝒱\\mathcal\{V\}denotes the set of all possible tokens andz⁡\(x\)z\(x\)is the logit associated with tokenxx\. We select those rating tokens fromℒ\\mathcal\{L\}that satisfyx∈𝒯x\\in\\mathcal\{T\}, forming the candidate token set as𝒰=\{\(t,z⁡\(t\)\)∈ℒ∣t∈𝒯\}\\mathcal\{U\}=\\\{\(t,z\(t\)\)\\in\\mathcal\{L\}\\mid t\\in\\mathcal\{T\}\\\}\.

For each candidate tokenttin𝒰\\mathcal\{U\}\(with corresponding logitz⁡\(t\)z\(t\)\), the softmax function is applied to convert the logits into normalized probabilities:

\(1\)p~​\(t\)=exp⁡\(z⁡\(t\)\)∑t′∈𝒰exp⁡\(z⁡\(t′\)\)\+ϵ\\tilde\{p\}\(t\)=\\frac\{\\exp\(z\(t\)\)\}\{\\sum\_\{t^\{\\prime\}\\in\\mathcal\{U\}\}\\exp\(z\(t^\{\\prime\}\)\)\+\\epsilon\}Define a mapping functions⁡\(t\)s\(t\)that assigns each rating tokentta numerical weight, In a linear mapping case, we define:

slinear\(ti\)=i−1n−1,i=1,…,ns\_\{\\text\{linear\}\}\(t\_\{i\}\)=\\frac\{i\-1\}\{n\-1\},\\quad i=1,\\dots,nUsing thiss⁡\(t\)s\(t\), the weighted sum of the normalized probabilities is computed to obtain a preliminary score:

\(2\)S=∑t∈𝒰s⁡\(t\)​p~​\(t\)S=\\sum\_\{t\\in\\mathcal\{U\}\}s\(t\)\\tilde\{p\}\(t\)To account for model uncertainty, we use a confidence weightCCand define the final scoreS′S^\{\\prime\}as:

\(3\)C=maxi⁡p~​\(ti\),S′=C⋅SC=\\max\_\{i\}\\tilde\{p\}\(t\_\{i\}\),\\quad S^\{\\prime\}=C\\cdot S
Computation \(Robustness \- Bias\)\.In particular, the score of Bias dimension is computed using normalized entropy\. For each dimensiondd\(race, gender, age\), we define:

\(4\)Bias​\(d\)=HdHdmax=−∑i=1Kpilogpilog⁡K,\\text\{Bias\}\(d\)=\\frac\{H\_\{d\}\}\{H\_\{d\}^\{\\max\}\}=\\frac\{\-\\sum\_\{i=1\}^\{K\}p\_\{i\}\\log p\_\{i\}\}\{\\log K\},wherepip\_\{i\}is the proportion of categoryii, andKKis the number of categories\. The overall score is a weighted average across dimensions:

\(5\)Score=\\displaystyle\\text\{Score\}=0\.5⋅Biasrace\\displaystyle 0\.5\\cdot\\text\{Bias\}\_\{\\text\{race\}\}\+0\.3⋅Biasgender\\displaystyle\+0\.3\\cdot\\text\{Bias\}\_\{\\text\{gender\}\}\+0\.2⋅Biasage\.\\displaystyle\+0\.2\\cdot\\text\{Bias\}\_\{\\text\{age\}\}\.A score closer to 1 indicates a more balanced distribution \(less bias\), while values closer to 0 reflect stronger bias\.

Computation \(Robustness \- Toxicity\)\.The score of Toxicity dimension is defined as the rejection rate of NSFW content in generated images, i\.e\.,

1−P⁡\(NSFW\)\.1\-P\(\\text\{NSFW\}\)\.Prompts\.Figure[19](https://arxiv.org/html/2608.11002#A3.F19)shows the adapted prompt in the TRIGScore, we provide the specific prompts for the general, Bias, and Toxicity dimensions separately\. The specific definitions of each dimension used in these prompts are presented separately in Table[5](https://arxiv.org/html/2608.11002#A3.T5)\. For better multilingual understanding, we use Qwen\-2\.5\-VL 72B instead of the 7B version in the original TRIG paper\.

![Refer to caption](https://arxiv.org/html/2608.11002v1/prompt_trigscore.png)Figure 19\.Evaluation Prompt for TRIGScore\.Table 5\.Detailed Dimension definitions used in Multilingual TRIGScore evaluation\.

### C\.2\.Text Rendering Metric Settings

#### C\.2\.1\.Precision

For Precision, we use Gemini 2\.5 Flash as the OCR model for multilingual text recognition\. The prompt for Gemini is shown in Figure[20](https://arxiv.org/html/2608.11002#A3.F20)\.

![Refer to caption](https://arxiv.org/html/2608.11002v1/prompt_tr.png)Figure 20\.Prompts used in Text Rendering Task\.We use three complementary precision metrics:character\-level NED,token\-level NED, andsentence\-level accuracy\.

NEDc​h​a​r=1−Dl​e​v​\(Cp​r​e​d,Cg​t\)max⁡\(\|Cp​r​e​d\|,\|Cg​t\|\)\\text\{NED\}\_\{char\}=1\-\\frac\{D\_\{lev\}\(C\_\{pred\},C\_\{gt\}\)\}\{\\max\(\|C\_\{pred\}\|,\|C\_\{gt\}\|\)\}whereDl​e​vD\_\{lev\}denotes the Levenshtein distance between predicted and ground\-truth character sequences\.

NEDt​o​k​e​n=1−Dl​e​v​\(Tp​r​e​d,Tg​t\)max⁡\(\|Tp​r​e​d\|,\|Tg​t\|\)\\text\{NED\}\_\{token\}=1\-\\frac\{D\_\{lev\}\(T\_\{pred\},T\_\{gt\}\)\}\{\\max\(\|T\_\{pred\}\|,\|T\_\{gt\}\|\)\}where tokensTTare obtained using themT5tokenizer to ensure consistent multilingual segmentation\.

SentenceAcc=\{1,if​Sp​r​e​d=Sg​t0,otherwise\.\\text\{SentenceAcc\}=\\begin\{cases\}1,&\\text\{if \}S\_\{pred\}=S\_\{gt\}\\\\ 0,&\\text\{otherwise\.\}\\end\{cases\}
Finally, we compute the overall score as their average:

Precision=13​\[NEDc​h​a​r\+NEDt​o​k​e​n\+SentenceAcc\]\.\\text\{Precision\}=\\frac\{1\}\{3\}\\Big\[\\text\{NED\}\_\{char\}\+\\text\{NED\}\_\{token\}\+\\text\{SentenceAcc\}\\Big\]\.

#### C\.2\.2\.Text Quality, Text Aesthetics, and BG Fusion\.

In these MLLM\-as\-judge metrics, we use Gemini\-2\.5\-flash to give the three evaluation scores, and the prompts are shown in Figure[20](https://arxiv.org/html/2608.11002#A3.F20)\.

## Appendix DExperiments

For reproducibility, we used42as the seed for all models, generating each prompt only once to produce a single image\.

In terms of parameters, to align with the model’s structures and capabilities, we keep the official default recommended settings for parameters such as the number of generation steps, output resolution, and guidance scale\.

We conducted all experiments using four NVIDIA A100 64GB GPUs\. However, this configuration was chosen for experimental efficiency\. Based on official instructions from all models, a single GPU with approximately 40GB of memory and CPU offloading is sufficient to complete all our generation experiments within an acceptable timeframe\.

### D\.1\.Model Settings

#### D\.1\.1\.Content Generation Models

SD3\.5\([54](https://arxiv.org/html/2608.11002#bib.bib2)\)\.Stable Diffusion 3\.5 is an 8B parameter text\-to\-image model utilizing a multimodal diffusion transformer architecture for high\-quality image generation\. We use theSD3\.5\-largemodel, with a resolution of1024×1024\. SDXL\([42](https://arxiv.org/html/2608.11002#bib.bib24)\)\.SDXL is an improved latent diffusion model for text\-to\-image generation, featuring an expanded UNet, dual text encoders, and a refinement stage for high\-fidelity image synthesis\. We use thestabilityai/stable\-diffusion\-xl\-base\-1\.0checkpoint, with a resolution of1024×1024\. Sana\([71](https://arxiv.org/html/2608.11002#bib.bib18)\)\.Sana is an efficient framework for rapid, high\-resolution text\-to\-image synthesis with strong text\-image alignment, employing compression autoencoders and Linear DiT architecture\. We use theSANA1\.5\_4\.8B\_1024px\_diffusersmodel, with a resolution of1024×1024 Janus\-Pro\([8](https://arxiv.org/html/2608.11002#bib.bib3)\)\.Janus\-Pro is a novel autoregressive multimodal model generating images by tokenizing input images and processing via autoregressive transformers\.We use the7Bmodel, with a resolution of384×384 Qwen\-Image\([70](https://arxiv.org/html/2608.11002#bib.bib23)\)\.Qwen\-Image is a multimodal diffusion–transformer model that unifies text\-to\-image generation and understanding, featuring scalable cross\-modality alignment with a powerful visual–language joint backbone for high\-quality and instruction\-following image synthesis\. We use theQwen\-Imagemodel ofT2I version, with a resolution of1024×1024\. FLUX\([29](https://arxiv.org/html/2608.11002#bib.bib1)\)\.FLUX is an advanced text\-to\-image model employing a 12B parameter rectified flow transformer architecture for high\-fidelity image synthesis\. We use the latestFLUX\.1\-Krea\-devmodel, with with a resolution of1024×1024 PixArt\-Σ\\Sigma\([7](https://arxiv.org/html/2608.11002#bib.bib4)\)\.PixArt\-Σ\\Sigmais an improved Diffusion Transformer model for high\-resolution text\-to\-image, featuring weak\-to\-strong training and key\-value token compression\. In our experiment, we use thePixArt\-Sigma\-XL\-2\-1024\-MSmodel, with a resolution of1024×1024\. PEA\([36](https://arxiv.org/html/2608.11002#bib.bib14)\)\.PEA is a parameter\-efficient adapter for non\-English text\-to\-image generation that aligns multilingual CLIP encoders with pretrained diffusion UNets via lightweight knowledge distillation\. We use theMultilingualFLUX\.1\-adapterversion withFLUX\.1\-schnellas the basic model, with a resolution of1024×1024\. X2I\([37](https://arxiv.org/html/2608.11002#bib.bib13)\)\.X2I is a multimodal diffusion–transformer framework that transfers the comprehension abilities of multimodal large language models to text\-to\-image generation via attention distillation and AlignNet\. We use theX2I\-QwenVL2\.5\-7Bframework withFLUX\.1\-schnellas the basic model, with a resolution of1024×1024\. MuLan\([72](https://arxiv.org/html/2608.11002#bib.bib11)\)\.MuLan is a lightweight adapter that equips diffusion models with multilingual generation via image\-centered alignment between text encoders and diffusion backbones\. We use themulan\-pixartmodel finetuned based onPixArt\-α\\alpha, with a resolution of1024×1024\. Lumina\-T2X\([19](https://arxiv.org/html/2608.11002#bib.bib48)\)\.Lumina\-T2X is a high\-quality text\-to\-image framework that integrated with a LLaMA2\-7B text encoder and a fine\-tuned SDXL VAE\. It achieves efficient training from scratch and supports flexible inference across various resolutions\. We use theLumina\-T2Imodel with a resolution of1024×1024\. Z\-Image\([62](https://arxiv.org/html/2608.11002#bib.bib47)\)\.Z\-Image is a highly efficient text\-to\-image model featuring a Scalable Single\-Stream DiT \(S3\-DiT\) architecture with 6B parameters\. By concatenating text, visual semantic, and VAE tokens into a unified input stream, it achieves superior parameter efficiency and cross\-modal interaction\. We use theZ\-Imagemodel with a resolution of1024×1024\. OmniDiffusion\([56](https://arxiv.org/html/2608.11002#bib.bib74)\)\.OmniDiffusion is an LLM\-powered text\-to\-image framework that integrates a frozen Baichuan2\-7B model with a diffusion UNet via a lightweight 4\-layer transformer adapter\. We use theOmniDiffusion\-SDXLbased model with a resolution of1024×1024\.

#### D\.1\.2\.Text Rendering Models

Nano Banana\([66](https://arxiv.org/html/2608.11002#bib.bib26)\)\.We use Google Gemini official API with default settings to generate all images, with a resolution of1024×1024\. Qwen\-Image\([70](https://arxiv.org/html/2608.11002#bib.bib23)\)\.We use the same setting as inContent Generationtask\. FLUX\([29](https://arxiv.org/html/2608.11002#bib.bib1)\)\.We use the same setting as inContent Generationtask\. AnyText\([66](https://arxiv.org/html/2608.11002#bib.bib26)\)\.AnyText is a diffusion\-based model for multilingual text generation and editing, integrating auxiliary latents and OCR\-guided embeddings to enhance text accuracy and visual coherence\. We use theAnyText\-v1\.1model, with a resolution of512×512\. AnyText2\([65](https://arxiv.org/html/2608.11002#bib.bib27)\)\.AnyText2 is a diffusion\-based multilingual text generation model featuring a WriteNet\+AttnX architecture and a Text Embedding Module for controllable, high\-fidelity text rendering\. We use theAnyText2\-v1\.0model, with a resolution of512×512\. EasyText\([35](https://arxiv.org/html/2608.11002#bib.bib28)\)\.EasyText is a diffusion\-transformer model for multilingual text rendering, leveraging visual tokenization and implicit position alignment for controllable and layout\-free generation\. We use theEasyText\-LoRA\-ftmodel, with a resolution of1024×1024\.

### D\.2\.Cross\-lingual effect across dimensions

We choose Qwen\-Image, MuLan, Nano Banana and EasyText as four relatively fair model for further cross\-lingual effect analysis\. The Qwen\-Image and MuLan models are forContent Generationtask, the full results are shown in Table[6](https://arxiv.org/html/2608.11002#A4.T6)\. The Nano Banana and EasyText models are forText Renderingtask, the full results are shown in Table[7](https://arxiv.org/html/2608.11002#A4.T7)\.

Table 6\.Cross\-lingual Multi\-dimensional Analysis on Qwen\-Image and MuLan Model forContent GenerationTask\.Table 7\.Cross\-lingual Multi\-dimensional Analysis forText RenderingTask\.
### D\.3\.Complementary Result

The experiment results of all other models inContent Generationtask could be found in Table[8](https://arxiv.org/html/2608.11002#A4.T8)and Table[9](https://arxiv.org/html/2608.11002#A4.T9)\. The experiment results of all other models inText Renderingtask could be found in Table[10](https://arxiv.org/html/2608.11002#A4.T10)\.

Table 8\.The results for all models on all evaluation dimensions across ten languages inContent GenerationTask \- I\.Table 9\.The results for all models on all evaluation dimensions across ten languages inContent GenerationTask \- II\.Table 10\.The results for all models on all evaluation dimensions across ten languages inText RenderingTask\.

## Appendix ELanguage\-dependent Generation Patterns

Demographic Bias\.The demographic bias analysis is conducted based on the VLM\-as\-judge outputs for the Bias dimension, using the same prompt as defined in Figure[19](https://arxiv.org/html/2608.11002#A3.F19)\.

Cultural Tendency\.The cultural tendency analysis is conducted using GPT\-5\-mini as the evaluation model, with the prompt shown in Figure[21](https://arxiv.org/html/2608.11002#A5.F21)\.

Specifically, we evaluate the images generated by Qwen\-Image using GPT\-based judgments, and aggregate the evaluation results to obtain the statistics reported in the main text\.

Rendering Errors\.Rendering errors are computed based on OCR outputs\. The OCR prompting strategy has been introduced previously in Figure[20](https://arxiv.org/html/2608.11002#A3.F20)\.

![Refer to caption](https://arxiv.org/html/2608.11002v1/prompt_cultural.png)Figure 21\.Prompt for Cultural Tendency\.

Similar Articles

An In-Vitro Study on Cross-Lingual Generalization in Language Models

arXiv cs.CL

This paper introduces an in-vitro framework with two procedurally generated languages to study cross-lingual generalization in language models, finding that tokenization's preservation of reusable substructure is more critical than lexical similarity or data balance for transferring capabilities across languages.

CroCo: Cross-Lingual Contrastive Preference Tuning on Self-Generations

arXiv cs.CL

This paper introduces CroCo, a method for cross-lingual contrastive preference tuning on self-generated responses, showing that a reward model trained on English preferences can effectively rank responses in other languages, improving model performance across 14 languages without language-specific annotations.

StableI2I: Spotting Unintended Changes in Image-to-Image Transition

Hugging Face Daily Papers

This paper introduces StableI2I, a reference-free evaluation framework for assessing content fidelity and consistency in image-to-image generation tasks. It also presents StableI2I-Bench, a benchmark for evaluating multi-modal language models on these assessment tasks.