The GaoYao Benchmark: A Comprehensive Framework for Evaluating Multilingual and Multicultural Abilities of Large Language Models

arXiv cs.CL Papers

Summary

GaoYao introduces a 182k-sample benchmark across 26 languages and 51 regions to systematically evaluate LLMs’ multilingual and multicultural capabilities, revealing large geographical performance gaps.

arXiv:2604.20225v1 Announce Type: new Abstract: Evaluating the multilingual and multicultural capabilities of Large Language Models (LLMs) is essential for their global utility. However, current benchmarks face three critical limitations: (1) fragmented evaluation dimensions that often neglect deep cultural nuances; (2) insufficient language coverage in subjective tasks relying on low-quality machine translation; and (3) shallow analysis that lacks diagnostic depth beyond simple rankings. To address these, we introduce GaoYao, a comprehensive benchmark with 182.3k samples, 26 languages and 51 nations/areas. First, GaoYao proposes a unified framework categorizing evaluation tasks into three cultural layers (General Multilingual, Cross-cultural, Monocultural) and nine cognitive sub-layers. Second, we achieve native-quality expansion by leveraging experts to rigorously localize subjective benchmarks into 19 languages and synthesizing cross-cultural test sets for 34 cultures, surpassing prior coverage by up to 111%. Third, we conduct an in-depth diagnostic analysis on 20+ flagship and compact LLMs. Our findings reveal significant geographical performance disparities and distinct gaps between tasks, offering a reliable map for future work. We release the benchmark (https://github.com/lunyiliu/GaoYao).
Original Article
View Cached Full Text

Cached at: 04/23/26, 10:03 AM

# The GaoYao Benchmark: A Comprehensive Framework for Evaluating Multilingual and Multicultural Abilities of Large Language Models
Source: [https://arxiv.org/html/2604.20225](https://arxiv.org/html/2604.20225)
Yilun Liu1, Chunguang Zhao1∗, Mengyao Piao1, Lingqi Miao1, Shimin Tao1, Minggui He1, Chenxin Liu1, Li Zhang1, Hongxia Ma1, Jiaxin Guo1, Chen Liu1, Liqun Deng1, Jiansheng Wei1, Xiaojun Meng1, Fanyi Du1, Daimeng Wei1, Yanghua Xiao2 1Huawei, China 2Fudan University, China liuyilun3@huawei\.com, zhaochunguang6@huawei\.com

###### Abstract

Evaluating the multilingual and multicultural capabilities of Large Language Models \(LLMs\) is essential for their global utility\. However, current benchmarks face three critical limitations: \(1\) fragmented evaluation dimensions that often neglect deep cultural nuances; \(2\) insufficient language coverage in subjective tasks relying on low\-quality machine translation; and \(3\) shallow analysis that lacks diagnostic depth beyond simple rankings\. To address these, we introduceGaoYao111GaoYao is derived from Chinese mythology, where he served as the first judicial officer, symbolizing fairness and comprehensiveness\., a comprehensive benchmark with 182\.3k samples, 26 languages and 51 nations/areas\. First, GaoYao proposes a unified framework categorizing evaluation tasks into three cultural layers \(General Multilingual, Cross\-cultural, Monocultural\) and nine cognitive sub\-layers\. Second, we achieve native\-quality expansion by leveraging experts to rigorously localize subjective benchmarks into 19 languages and synthesizing cross\-cultural test sets for 34 cultures, surpassing prior coverage by up to 111%\. Third, we conduct an in\-depth diagnostic analysis on 20\+ flagship and compact LLMs\. Our findings reveal significant geographical performance disparities and distinct gaps between tasks, offering a reliable map for future work\. We release the benchmark222https://github\.com/lunyiliu/GaoYao\.

The GaoYao Benchmark: A Comprehensive Framework for Evaluating Multilingual and Multicultural Abilities of Large Language Models

Yilun Liu1††thanks:Equal contribution\., Chunguang Zhao1∗, Mengyao Piao1, Lingqi Miao1, Shimin Tao1,Minggui He1, Chenxin Liu1, Li Zhang1, Hongxia Ma1, Jiaxin Guo1, Chen Liu1,Liqun Deng1, Jiansheng Wei1, Xiaojun Meng1, Fanyi Du1,Daimeng Wei1, Yanghua Xiao21Huawei, China2Fudan University, Chinaliuyilun3@huawei\.com, zhaochunguang6@huawei\.com

## 1Introduction

As Large Language Models \(LLMs\) increasingly serve a global user base, the ability to process diverse languages and navigate complex cultural contexts has become a critical measure of their inclusivity\. However, the current landscape of multilingual evaluation is fraught with challenges that hinder a holistic understanding of model performance:

\(1\) Lack of Systematicity and Cultural Neglect\.Many prominent benchmarks focus narrowly on single specific facets of language ability, such as factual knowledge\(Romanouet al\.,[2025](https://arxiv.org/html/2604.20225#bib.bib81)\)or reading comprehension\(Bandarkaret al\.,[2024](https://arxiv.org/html/2604.20225#bib.bib76)\)\. Consequently, they often overlook the deeper capabilities a model should possess \(*e\.g\.*, cultural sensitivity\), treating multilingualism merely as isolated evaluation points rather than interconnected dimensions rooted from cultural and cognitive sources\.

\(2\) Limited Language Coverage and Quality in Subjective Tasks\.Subjective tasks \(*i\.e\.*, answers are open\-ended\) such as instruction following and multi\-turn dialogue are predominantly assessed in English\(Liet al\.,[2023](https://arxiv.org/html/2604.20225#bib.bib68); Zhenget al\.,[2023](https://arxiv.org/html/2604.20225#bib.bib69)\)\. Existing multilingual extensions often rely on automated machine translation \(MT\) or cover only a handful of languages\(Zhanget al\.,[2024](https://arxiv.org/html/2604.20225#bib.bib70); Liuet al\.,[2024](https://arxiv.org/html/2604.20225#bib.bib71)\)\. This reliance on MT introduces “translationese” and fails to reflect native tongues, which can be trivial in objective tasks \(*e\.g\.*, true/false\) but is especially harmful for subjective evaluation\.

\(3\) Lack of In\-Depth Diagnostic Analysis\.Existing studies often stop at superficial leaderboard rankingsPomerenkeet al\.\([2025](https://arxiv.org/html/2604.20225#bib.bib67)\); Liuet al\.\([2024](https://arxiv.org/html/2604.20225#bib.bib71)\), failing to reveal implications under performance variance\. There is a scarcity of insights regarding how performance correlates with geographical regions, task types, or model architectures, leaving potential challenges and gaps untouched\.

To address these challenges, we introduceGaoYao, a multilingual and multicultural benchmark emphasizing systematicity, authenticity, and analytical depth, which features three aspects:

![Refer to caption](https://arxiv.org/html/2604.20225v1/images/main_fig_final.png)Figure 1:Illustration on design and construction of GaoYao\. The benchmark is grounded in theoretical models of culture and cognition, and constructed through a hybrid strategy of integration, expansion and generalization\.\(1\) A Systematic Evaluation Framework\.Grounded in cultural theory\(Hall,[1976](https://arxiv.org/html/2604.20225#bib.bib73)\)and cognitive taxonomy\(Anderson and Krathwohl,[2001](https://arxiv.org/html/2604.20225#bib.bib74)\), we propose a unified evaluation matrix\. This framework categorizes capabilities into three layers:General Multilingual Abilities\(universal concepts\),Cross\-cultural Abilities\(culturally shared concepts with variance\), andMonocultural Abilities\(unique concepts in cultures\)\. These are further expanded into nine cognitive sub\-layers, ranging from knowledge retention to creative writing\.

\(2\) Native\-Quality Data Expansion\.We address the data scarcity by leveraging a team of multilingual experts to meticulously localize English evaluation sets in two critical sub\-layers into 19 languages, ensuring native\-level quality compared to machine\-translated alternatives\. Additionally, we generalize a cultural evaluation set to cover 34 cultures through a novel expert\-verified synthesis pipeline\. In Fig\.[7](https://arxiv.org/html/2604.20225#S3.F7), our three curated test sets are able to better reflect capability stratification of LLMs due to the native\-level data quality\.

\(3\) In\-Depth Diagnostic Findings\.We conduct a tiered evaluation of representative SOTA models and go beyond simple rankings\. Our analysis reveals the severe "digital divide" across regions and the performance gap between mature and frontier tasks\. Synthesizing these findings, we propose ameta\-findingto guide the community: recommending strategic deployment methods for efficient model usage and advocating for equitable data construction to bridge the capability gaps\.

Our contributions can be summarized as:

- •We propose a systematic evaluation benchmark with 23\.3M tokens, three cultural layers and nine cognitive sub\-layers, addressing the fragmented nature of existing benchmarks\.
- •We expand critical instruction\-following and dialogue evaluations to 19 languages and generalize cultural evaluation sets to 34 cultures via a rigorous human\-in\-the\-loop method, filling the blank with high data quality and enabling better identification of LLM abilities\.
- •We provide a comprehensive capability landscape of existing multilingual LLMs through deep analysis, revealing several key insights which guide LLM usage and development\.

In addition, we release all test sets and evaluation codes, providing the community with a reliable compass for future work on multilingual LLMs\.

## 2Methodology

As shown in Fig\.[1](https://arxiv.org/html/2604.20225#S1.F1), our framework begins by defining a theoretical landscape of multilingual and multicultural capabilities critical in LLM evaluation\. Guided by this theoretical structure, we employ a three\-pronged approach—Integration, Expansion, and Generalization—to curate a comprehensive benchmark that addresses existing gaps in systematicity, coverage and quality\. Section[2\.1](https://arxiv.org/html/2604.20225#S2.SS1)details the theoretical underpinnings of our evaluation dimensions and Section[2\.2](https://arxiv.org/html/2604.20225#S2.SS2)discusses the specific processes involved in constructing the GaoYao benchmark through these three strategies\.

### 2\.1Layered Evaluation Dimensions

##### Theoretical Foundations of Three Major Layers

Drawing inspiration from the Cultural Iceberg Model\(Hall,[1976](https://arxiv.org/html/2604.20225#bib.bib73)\)and the Three\-Layer Model of organizational culture\(Schein,[2010](https://arxiv.org/html/2604.20225#bib.bib72)\), we posit that existing multilingual benchmarks often predominantly assess “surface\-level” linguistic proficiencies while overlooking the deeper, implicit cultural contexts that shape communication\. To address this, GaoYao categorizes tasks into three major layers representing cultural deepening levels:

- •General Multilingual Abilities:This layer corresponds to the “tip of the iceberg,” focusing on universal concepts that remain consistent across languages \(*e\.g\.*, applying target language to handle problems involving reasoning, knowledge or comprehension\)\.
- •Cross\-cultural Abilities:Moving beneath the surface, this layer assesses the model’s capacity to navigate shared concepts that manifest differently across cultures\. For instance, while the lexical term “dragon” translates directly, its symbolic meaning varies drastically: Western dragons are typically depicted as malevolent monsters while the eastern dragon \(or loong\) is revered as an auspicious symbol\(Zhao,[1988](https://arxiv.org/html/2604.20225#bib.bib77)\)\. An LLM must discern these subtle cultural divergences\.
- •Monocultural Abilities:The deepest layer evaluates the understanding of unique concepts exclusive to specific cultures, which often lack direct equivalents elsewhere\. An example is the Chinese phenomenon of “Chunyun” \(the massive Spring Festival travel rush\(Zhuet al\.,[2021](https://arxiv.org/html/2604.20225#bib.bib79)\)\), a culturally specific event laden with unique social implications\. Another example is “Namaste”, the special greeting etiquette in India\(Zhanget al\.,[2025](https://arxiv.org/html/2604.20225#bib.bib80)\)\.

##### Deriving Nine Sub\-layers via Cognitive Taxonomy

Within these major cultural layers, we further ensure a comprehensive evaluation matrix by structuring tasks according to Bloom’s Taxonomy of cognitive domains\(Anderson and Krathwohl,[2001](https://arxiv.org/html/2604.20225#bib.bib74)\)\. This taxonomy categorizes human thought processes into six categories along a gradient of complexity, ranging from basic remembering to complex creation\. Inspired by the six cognitive levels, the task design of GaoYao encompasses nine distinct sub\-layers to ensure a rigorous assessment:

- •Remembering & Understanding:Reflected by tasks of multilingualKnowledgeQ&A,ReadingComprehension, andTranslation\.
- •Applying & Analyzing:Assessed throughReasoningtasks andMathproblem solving\.
- •Evaluating & Creating:This highest cognitive level encompasses: \(1\) Creative Tasks:Instruction FollowingandMulti\-turn Dialogue, which demand creative writing to satisfy complex, open\-ended user intents; \(2\) Evaluative Tasks: The advancedCross\-culturalandMonoculturalassessments\. Unlike simple factual retrieval, these tasks require the model to evaluate social nuances, discern cultural appropriateness among highly plausible distractors, and make value judgments aligned with cultural normsRystrømet al\.\([2025](https://arxiv.org/html/2604.20225#bib.bib84)\)\.

### 2\.2Construction of GaoYao Benchmark

Guided by the theoretical framework above, we construct the GaoYao benchmark through a hybrid strategy combining the integration of established resources \(for seven of the nine sub\-layers\), the linguistic expansion of high\-value under\-served benchmarks \(two most critical sub\-layers: instruction following and multi\-turn dialogues\), and the generalization of cultural data through human\-in\-loop synthesis pipelines \(the cross\-cultural layer\)\.

#### 2\.2\.1Integration of Existing Test Sets

For several of the defined cognitive sub\-layers, particularly those related to objective knowledge and reasoning, the research community has already established high\-quality open\-source benchmarks\. Rather than reinventing these, we conducted literature review and quality checks to select and integrate some of the most widely\-verified and robust datasets into GaoYao, ensuring a complete coverage of our defined evaluation sub\-layers\. These datasets are mapped to our sub\-layers as follows:

\(1\)Knowledge & Reasoning: Given their coverage on factual knowledge spanning various subjects from elementary\-level knowledge up to advanced professional subjects, bothInclude\(Romanouet al\.,[2025](https://arxiv.org/html/2604.20225#bib.bib81)\)andMMMLU\(OpenAI,[2024](https://arxiv.org/html/2604.20225#bib.bib75)\)are integrated\. Compared withInclude,MMMLUfocuses more on reasoning abilities,*i\.e\.*, how LLMs apply these knowledge to solve practical problems\.

\(2\)Reading: We incorporateBelebele\(Bandarkaret al\.,[2024](https://arxiv.org/html/2604.20225#bib.bib76)\)for evaluating multilingual reading comprehension capabilities given its native passage coverage and rigorous quality assurance\.

\(3\)Translation:Flores\-101\(Goyalet al\.,[2022](https://arxiv.org/html/2604.20225#bib.bib18)\)provides a widely\-recognized standard for assessing MT across numerous language pairs\.

\(4\)Math:MGSM\(Shiet al\.,[2023](https://arxiv.org/html/2604.20225#bib.bib83)\)is also a widely\-used dataset to evaluate multilingual mathematical reasoning capabilities\.

\(5\)Cross\-culture & Monoculture: Since the research community has only recently begun to rigorously define and evaluate the multicultural capabilities of LLMs\(Rystrømet al\.,[2025](https://arxiv.org/html/2604.20225#bib.bib84)\), open\-source resources remain scarce\. We leverage two recent datasets:SAGE\(Guoet al\.,[2025](https://arxiv.org/html/2604.20225#bib.bib85)\)for identifying cultural differences in shared concepts \(cross\-culture\) andCultureScope\(Zhanget al\.,[2025](https://arxiv.org/html/2604.20225#bib.bib80)\)for understanding unique cultural concepts \(monoculture\)\. While both datasets delve deeply into culture\-specific concepts and employ rigorous design procedures, their coverage is restricted to Chinese and Spanish\. To supplement this, we constructed a cross\-cultural evaluation set spanning 34 cultures \(see Section[2\.2\.3](https://arxiv.org/html/2604.20225#S2.SS2.SSS3)\)\.

#### 2\.2\.2Expansion of Language Coverage forAlpacaEvalandMT\-Bench

Instruction following and multi\-turn dialogue represent critical capabilities reflecting an LLM’s practical utility and “human\-likeness\.” However, existing multilingual benchmarks heavily prioritize objective tasks, leaving these subjective, open\-ended abilities predominantly evaluated only in English\. To close this significant gap, we selected two widely recognized English benchmarks:AlpacaEvalLiet al\.\([2023](https://arxiv.org/html/2604.20225#bib.bib68)\), validated by over 20k human judgments for general instruction following, andMT\-BenchZhenget al\.\([2023](https://arxiv.org/html/2604.20225#bib.bib69)\), designed with challenging multi\-turn questions across intent categories such as role playing and creative writing\. We then expanded their coverage to over 19 languages, denoting asS\-AlpacaEvalandS\-MT\-Bench, respectively\.

This expansion was not a simple translation task but a rigorous localization effort\. From the language service center of a top\-tier corporation, we recruited a team of 20 native\-speaker professionals with expertise in translation, localization, and linguistic testing\. The team dedicated a total of 175 person\-days to this development\. To ensure the highest quality, a strict review\-rebuttal feedback loop was implemented for each language\. Third\-party reviewers continuously inspected samples during annotation\. Disagreements triggered a discussion phase where annotators either revised their work based on the concerns or provided justifications to persuade the reviewer to unflag the sample\.

Crucially, there is a localization process to make sure every user question is linguistically feasible, which can hardly be guaranteed using MT\. For instance, constrained English instruction like “list items starting with the letter A” will be invalid if being translated literally to a language without letter A\. Thus, such instructions were manually adapted or reconstructed to suit the phonetic and script characteristics of the target language while ensuring the cognitive task remained equivalent\.

#### 2\.2\.3Generalization of Cross\-cultural Evaluation \(SuperBLEnD\)

As discussed in Section[2\.2\.1](https://arxiv.org/html/2604.20225#S2.SS2.SSS1), existing cultural evaluation sets are limited in its culture coverage\. However, expanding such coverage presents a unique challenge: direct translation retains source\-culture concepts, while manual creation is costly\. To address this, we generalizedBLEnD\(Myunget al\.,[2024](https://arxiv.org/html/2604.20225#bib.bib82)\)intoSuperBLEnD, expanding coverage from 16 to 34 cultures via a three\-stage semi\-automated pipeline incorporating rigorous human verification\.SuperBLEnDfocuses on evaluating understanding of cultural differences regarding everyday concepts such as festivals, food and sports\. The pipelines are as follows \(full annotations and technical details are in Appendix[E](https://arxiv.org/html/2604.20225#A5)\):

![Refer to caption](https://arxiv.org/html/2604.20225v1/images/fig_example_final.png)Figure 2:An example of the linguistic enrichment process \(stage 3\), which increases complexity of MCQs without altering the underlying cultural fact\.##### Stage 1: Cultural Generalization\.

We curated a subset of high\-quality templates fromBLEnD\(inheriting answers for the original 16 cultures\) and recruited native experts to provide authentic answers for 18 additional cultures based on lived experience\. Answers underwent strict manual verification to eliminate invalid or toxic content \(discarding∼\\sim41\.1% of raw data\), ensuring high\-quality ground truth even for questions with multiple valid answers \(*e\.g\.*, accepting both "beer" and "carbonated drinks" for Malaysian nightclubs\)\.

##### Stage 2: Option Synthesis\.

To enhance diversity, verified Q&A pairs from stage 1 were converted into multiple\-choice questions \(MCQs\) by combining target answers with distractors from other cultures or plausible LLM\-generated "dummy options"\. Options underwent strict verification to exclude low\-quality cases such as hierarchical conflicts \(*e\.g\.*, rejecting "Pepsi" as a distractor if the answer is "beer", as it falls under the valid category of "carbonated drinks"\)\.

##### Stage 3: Linguistic Enrichment\.

To enhance difficulty and prevent simple pattern matching, we utilized an LLM to rephrase question stems and options via techniques like syntactic restructuring and voice alternation\. As shown in Fig\.[2](https://arxiv.org/html/2604.20225#S2.F2), this process ensures the benchmark tests deep cultural reasoning rather than superficial keyword recognition\.

## 3Experiment

### 3\.1Experimental Setups

To empirically validate GaoYao’s efficacy in mapping the global LLM landscape, we conducted a tiered evaluation across a spectrum of models representing the current SOTA\. Our selection encompasses both open\-weights models \(*e\.g\.*, DeepSeek\-V3\.1DeepSeek \([2025](https://arxiv.org/html/2604.20225#bib.bib93)\)\) inferred on standardized NPU computation nodes, and proprietary commercial models \(*e\.g\.*, GPT\-5OpenAI \([2025a](https://arxiv.org/html/2604.20225#bib.bib88)\)\) accessed via official APIs\. The specific model versions and resource addresses are in Table[7](https://arxiv.org/html/2604.20225#A7.T7)\. We make sure all evaluated LLMs are post\-trained versions \(*e\.g\.*, “instruct” or “chat” versions\) with “thinking” disabled \(except Fig\.[8](https://arxiv.org/html/2604.20225#S3.F8)\)\. FollowingYanget al\.\([2025](https://arxiv.org/html/2604.20225#bib.bib94)\), we adopted only a random 10% subset ofMMMLUdue to its unproportionate volume\. Section[3\.1\.1](https://arxiv.org/html/2604.20225#S3.SS1.SSS1)and Section[3\.1\.2](https://arxiv.org/html/2604.20225#S3.SS1.SSS2)further illustrates the setups\. See a reliability analysis of GaoYao in Appendix[C](https://arxiv.org/html/2604.20225#A3)\.

#### 3\.1\.1Statistics of Test Sets in GaoYao

![Refer to caption](https://arxiv.org/html/2604.20225v1/images/gaoyao_coverage_map.png)Figure 3:The language and culture coverage on the world map\. Colors indicate resource popularity levels\.![Refer to caption](https://arxiv.org/html/2604.20225v1/images/languages_distribution.png)

\(a\) Sample distribution by language

![Refer to caption](https://arxiv.org/html/2604.20225v1/images/categories_distribution.png)

\(b\) Distribution by evaluation dimensions

Figure 4:Distribution statistics of test sets in GaoYao \(a\) by languages and \(b\) by evaluation sub\-layers\.![Refer to caption](https://arxiv.org/html/2604.20225v1/images/Scores_distribution_per_category.png)\(a\)Open\-Source Models
![Refer to caption](https://arxiv.org/html/2604.20225v1/images/Scores_distribution_per_category_close_source.png)\(b\)Closed\-Source \(API\-based\) Models
![Refer to caption](https://arxiv.org/html/2604.20225v1/images/Scores_distribution_per_category_compact.png)\(c\)Compact Models \(<20<20B\)

Figure 5:Performance heatmaps across nine evaluation sub\-layers\. Scores are averaged across all languages\. Numbers in parentheses indicate rank within the group\. Backgrounds: Pink \(General Multilingual\), Blue \(Cultural Abilities\)\. SA, SB and CS represents specific datasets:SAGE,SuperBLEnDandCultureScope\.As shown in Fig\.[3](https://arxiv.org/html/2604.20225#S3.F3), the dataset spans 26 languages distributed across 51 nations/areas \(34 of them are culturally represented as discussed in Section[2\.2\.3](https://arxiv.org/html/2604.20225#S2.SS2.SSS3)\)\. The distribution encompasses five geopolitical clusters: Western Europe, Eastern Europe, East Asia & Southeast Asia, Middle East & Africa, and South Asia \(See Table[4](https://arxiv.org/html/2604.20225#A3.T4)for detailed statistics\)\. A crucial design principle of GaoYao is the mitigation of resource bias; as illustrated in Fig\.[4](https://arxiv.org/html/2604.20225#S3.F4)\(a\), excluding the three dominant lingua francas, the distribution is relatively balanced, with each remaining language constituting roughly 1%\-3% of the total volume in sample level\. According toJoshiet al\.\([2020](https://arxiv.org/html/2604.20225#bib.bib86)\), the languages include nine low\-resource and ten mid\-resource varieties \(See Appendix[B](https://arxiv.org/html/2604.20225#A2)\)\.

Fig\.[4](https://arxiv.org/html/2604.20225#S3.F4)\(b\) shows the distribution of test set sizes across the nine evaluation sub\-layers of the GaoYao benchmark, as introduced in Section[2\.1](https://arxiv.org/html/2604.20225#S2.SS1)\. The distribution of samples and tokens varies distinctively across sub\-layers due to their inherent task characteristics\. WhileTranslationcomprises the highest volume of samples, it accounts for a relatively modest share of total tokens, reflecting the sentence\-level brevity typical of theFlores\-101dataset\. In contrast, sub\-layers such asReasoningandReadingexhibit a significantly higher token\-to\-sample ratio, as these domains necessitate extensive context to define complex problem spaces\. Additionally, the cultural layers \(MonocultureandCross\-culture\) represent a relatively large token count to ensure sufficient depth to capture cultural nuances, reflecting GaoYao’s emphasis on cultural evaluation\.

#### 3\.1\.2Evaluation Approaches

The evaluation protocol for each sub\-dataset can be divided into two categories \(details on metrics, judges and calculation methods are in Table[6](https://arxiv.org/html/2604.20225#A7.T6)\):

Objective Evaluation:For question type with deterministic outputs \(*e\.g\.*, MCQ, calculation problems\), we utilize standardized prompt templates\(OpenAI,[2024](https://arxiv.org/html/2604.20225#bib.bib75); Romanouet al\.,[2025](https://arxiv.org/html/2604.20225#bib.bib81); Bandarkaret al\.,[2024](https://arxiv.org/html/2604.20225#bib.bib76); Goyalet al\.,[2022](https://arxiv.org/html/2604.20225#bib.bib18)\)and rule\-based extraction with regular expressions to parse answers from LLMs’ responses\. To ensure reproducibility, all pre\-processing and post\-processing scripts have been released\.

Subjective Evaluation:For open\-ended tasks \(*e\.g\.*, Q&A\), we adopt the widely\-used “LLM\-as\-Judge" paradigm\(Liet al\.,[2023](https://arxiv.org/html/2604.20225#bib.bib68); Zhenget al\.,[2023](https://arxiv.org/html/2604.20225#bib.bib69)\), where the judge model compares response from a candidate model with a reference response based on specific dimensions and concludes with “win”, “lose” or “tie”\. We standardized on DeepSeek\-v3\.1 as the judge due to its superior reasoning abilities\. The primary metric isWin Rateagainst the reference responses\. For datasets lacking inherent references \(S\-AlpacaEval,S\-MT\-Bench\), we introduced Qwen3\-235B\-A22BYanget al\.\([2025](https://arxiv.org/html/2604.20225#bib.bib94)\)as the reference anchor\.

All scores \(*e\.g\.*, accuracy, win rate\) are displayed at the scale of 0\-100 for clearer viewing\. The results are aggregated along specific axes,*e\.g\.*,Task Dimension\(averaging across all languages for a specific evaluation sub\-layer\) for Finding 1&3 andLanguage Dimension\(averaging across all sub\-layers for a specific language\) for Finding 2\.

### 3\.2Results and Findings

##### Finding 1: Performance Differentiation Among Flagship and Compact Models

Fig\.[5](https://arxiv.org/html/2604.20225#S3.F5)\(a\) and \(b\) present the landscape of flagship models, including open\-source leaders and closed\-source commercial APIs\. The results reveal distinct multilingual capability profiles rather than a uniform dominance:

- •Logic & Culture Specialist:openPangu\-Ultra\-MoE\-718B\-V1\.1Ascend Tribe \([2025](https://arxiv.org/html/2604.20225#bib.bib92)\)exhibits exceptional strength inMath\(\#1\) andReasoning\(\#2\) among open\-source LLMs, while simultaneously securing among top ranks in theCross\-culturesub\-layer\. This correlation suggests its rigorous logical training may facilitate understanding of complex cultural frameworks\.
- •Knowledge Heavyweights:Gemini\-2\.5\-ProComaniciet al\.\([2025](https://arxiv.org/html/2604.20225#bib.bib89)\)demonstrates dominance inKnowledge,ReadingandTranslation, suggesting a pre\-training corpus with extensive informational breadth\.
- •Interaction Specialists:DeepSeek\-R1DeepSeek\-AI \([2025](https://arxiv.org/html/2604.20225#bib.bib96)\)leads both open\-source and closed\-source models inInstruction FollowingandDialogue, reflecting a post\-training strategy optimized for conversational utility and complex user constraints\.

![Refer to caption](https://arxiv.org/html/2604.20225v1/images/Languages_distribution_per_geographical_groups.png)Figure 6:Impact of geography onbestperformance achieved by LLMs across languages\. Vertical dashed lines represent group averages\.Fig\.[5](https://arxiv.org/html/2604.20225#S3.F5)\(c\) illustrates the performance of compact models \(<20<20B parameters\)\. We observed a possibility of saturation for open\-source benchmark: A crucial finding is the discrepancy in performance gaps\. On established benchmarks likeBelebeleandInclude, the compact Qwen3\-14BYanget al\.\([2025](https://arxiv.org/html/2604.20225#bib.bib94)\)achieves near\-parity with the massive Qwen3\-235B\. However, on GaoYao’s newly constructed subjective sets \(S\-AlpacaEvalandS\-MT\-Bench\), a significant gap persists\. This suggests that some popular open\-source benchmarks may have become insensitive in identifying multilingual capabilities \(a phenomenon called benchmark saturation\)Akhtaret al\.\([2026](https://arxiv.org/html/2604.20225#bib.bib117)\), whereas GaoYao’s fresh, expert\-localized data exposes the true gap between compact and flagship models\. Finding 3 further investigates it\.

##### Finding 2: Digital Divide Among Language Geography and Resource\-Levels

As shown in Fig\.[6](https://arxiv.org/html/2604.20225#S3.F6)and Fig\.[9](https://arxiv.org/html/2604.20225#A3.F9), by analyzing LLMs’ best performances through a geopolitical and resource\-level lens \(*i\.e\.*, aggregating highest scores among all LLMs by languages\), a persistent "digital divide" is revealed\. Performance on a certain language is strongly correlated with geographic attributes and resource availability: Western European languages consistently score highest, while low\-resource languages in South Asia and Africa lag significantly\. This hierarchy is consistent with resource levels in Fig\.[9](https://arxiv.org/html/2604.20225#A3.F9): High \> Medium \> Low popularity across maximum, mean, and minimum scores, underscoring that current multilingual progress is uneven and largely driven by data volume rather than universal linguistic transfer\.

![Refer to caption](https://arxiv.org/html/2604.20225v1/x1.png)Figure 7:Distribution of model performance across tasks \(*i\.e\.*, sub\-layers\) using standard box plot\. Height of boxes and whiskers indicate the performance divergence among models, while the horizontal line is the median score\. Datasets constructed in this work are inRed\.
##### Finding 3: Capability Stratification Between Solved and Frontier Tasks

Fig\.[7](https://arxiv.org/html/2604.20225#S3.F7)presents a boxplot analysis of model performance distributions, revealing a severe stratification in current LLM capabilities for different tasks \(*i\.e\.*, sub\-layers\)\. Objective tasks such asReadingandMathexhibit high median scores \(\>85\>85\) with compressed boxes, indicating a closed gap between flagship and compact models due to their standardized patterns\. In contrast, subjective tasks \(especially forInstruction FollowingandDialoguewith expert\-localized datasets in GaoYao\) display significantly lower medians and elongated interquartile ranges \(box and whisker heights\)\. This high variance confirms that these tasks serve as high\-sensitivity discriminators due to their advanced requirements on creating and human\-likeness, effectively exposing the capability frontier where flagship models significantly outperform average open\-source models\.

In addition, a comparison within cultural tasks demonstrates the value of theSuperBLEnDdataset\. While the two existing cultural test sets \(*i\.e\.*,SAGEandCultureScope\) show signs of saturation \(medians≈90\\approx 90\), our synthesized dataset for the cross\-cultural layer reveals a significant drop in median score and relatively wide dispersion\. This proves that the construction strategy in Section[2\.2\.3](https://arxiv.org/html/2604.20225#S2.SS2.SSS3)successfully elevates the challenge from simple knowledge retrieval to rigorous cultural reasoning on authentic experiences, leading to a more precise test set for cultural capabilities\.

![Refer to caption](https://arxiv.org/html/2604.20225v1/images/F4_reasoning_improvment.png)Figure 8:Multilingual performance gain by “thinking” mode \(Think−BaseBase\\frac\{\\text\{Think\}\-\\text\{Base\}\}\{\\text\{Base\}\}\*100%\) across Bloom cognitive layers\.
##### Finding 4: Uneven Gain by “Thinking” in Multilingual Tasks

The paradigm of inference\-time reasoning \( or “thinking”\) has been proven effective in massive fieldsDeepSeek\-AI \([2025](https://arxiv.org/html/2604.20225#bib.bib96)\); Liuet al\.\([2025a](https://arxiv.org/html/2604.20225#bib.bib90)\); Heet al\.\([2025](https://arxiv.org/html/2604.20225#bib.bib91)\)\. In Fig\.[8](https://arxiv.org/html/2604.20225#S3.F8), we investigated the impact of “thinking” in multilingual context\. By aggregating task scores across Bloom’s cognitive taxonomy \(as described in Section[2\.1](https://arxiv.org/html/2604.20225#S2.SS1)\), two divergent behaviors are revealed:

\(1\)Selective Gain for Flagships\.For flagship models \(DeepSeek\-V3\.1, Qwen3\-235B\), the "Thinking" mode acts as a specialized tool\. In theRemembering & Understandinglayer \(*e\.g\.*, Translation, Knowledge\), gains are marginal, suggesting the bottleneck is mainly multilingual knowledge for retrieval\-heavy tasks\. However, in theEvaluating & Creatinglayer \(*e\.g\.*, Instruction Following\), we observe significant gains, probably due to a more comprehensive considerations of the constraints in users’ instructions\.

\(2\)Universal Gain for Compact Models\.In contrast, the compact Qwen3\-14B benefits universally, achieving significant gains even in basic understanding tasks\. This suggests that “Thinking” effectively compensates for the limited parameter capacity of smaller models, allowing them to punch above their weight\.

##### Meta\-Finding: From Benchmarking to Guidance

Transcending individual metrics, our findings \(F1\-F4\) combine into a meta\-finding that guides multilingual LLM utility and development:

Strategic Deployment \(Usage\)\.There is no "one\-fits\-all" model\. Users should choose flagship models wisely based on their features \(the interaction specialist or the knowledge heavyweights, per F1\) and adopt a dynamic strategy for “thinking” mode \(F4\): deploy compact models with “thinking” enabled for better performance in resource\-constrained inference, while reserving flagship models for complex creative tasks to balance cost and performance\.

Equitable Construction \(Development\)\.The geographic performance cliffs \(F2\) and task stratification \(F3\) reveal the training data gap both for low\-resource areas and highly subjective tasks\. Future data development must pivot from English\-centric translation to authentic regional curation\. We recommend leveraging thehuman\-in\-the\-loop generalizationpipeline \(in Section[2\.2\.3](https://arxiv.org/html/2604.20225#S2.SS2.SSS3)\) to efficiently fill the training data voids, ensuring both language equity and authenticity\.

## 4Discussion

### 4\.1Necessity of the Hierarchical Framework

To further examine whether the three\-layer framework reveals meaningful capability distinctions beyond an aggregation of existing benchmarks, we conduct a rank\-based transfer analysis across the three major layers in GaoYao: General multilingual, Cross\-cultural, and Monocultural abilities\. Following the model set in Fig\.[5](https://arxiv.org/html/2604.20225#S3.F5), we rank the 23 evaluated models by their average scores within each layer and compute Spearman’s rank correlation between the General multilingual ranking and the rankings of the two deeper cultural layers\.

ModelGeneralCross\-culturalMonoculturalGemini\-2\.5\-Pro\#1\#1\#8Doubao\-Seed\-1\.6\#2\#14\#6Qwen3\-235B\-A22B\#9\#11\#1DeepSeek\-V3\.1\#15\#16\#4

Table 1:Representative model rankings across the three major layers in GaoYao\. Higher general multilingual ranking does not necessarily transfer to deeper cultural layers\.The results show only modest transfer from general multilingual capability to deeper cultural capabilities: the correlation is higher for Cross\-cultural tasks \(Spearman’sρ=0\.74\\rho=0\.74\) but drops for Monocultural tasks \(ρ=0\.61\\rho=0\.61\)\. This decline is important because monocultural evaluation requires models to recognize culturally unique concepts and norms rather than solve language\-general problems\. As shown in Table[1](https://arxiv.org/html/2604.20225#S4.T1), Gemini\-2\.5\-Pro ranks first in the General Multilingual and Cross\-cultural layers but drops to eighth in the Monocultural layer, while DeepSeek\-V3\.1 rises from fifteenth in General Multilingual to fourth in Monocultural\. Similarly, Qwen3\-235B\-A22B ranks ninth in General Multilingual but first in Monocultural\. These rank shifts indicate that strong surface\-level fluency or general reasoning does not guarantee deep cultural nuance\.

This analysis validates the necessity of GaoYao’s hierarchical framework\. If all tasks were collapsed into a single multilingual score, these capability decouplings would be obscured\. By separating General Multilingual, Cross\-cultural, and Monocultural layers, GaoYao can diagnose where a model’s multilingual competence genuinely transfers and where culturally grounded evaluation remains a distinct frontier\.

### 4\.2Necessity of Linguistic Enrichment inSuperBLEnD

We also conduct an ablation study to verify whether the option synthesis and linguistic enrichment strategies inSuperBLEnDmake the benchmark more discriminative\. We compare model accuracy on the originalBLEnDsetting without our enrichment againstSuperBLEnDon the same 16 overlapping cultures\. This controls for culture coverage, so the main difference is whether our synthesized distractors and enriched phrasings are applied\.

Table[2](https://arxiv.org/html/2604.20225#S4.T2)shows that all three representative models experience accuracy drops after applying our synthesis and enrichment strategies\. The decrease is relatively modest for larger flagship models, such as Qwen3\-235B \(\-4\.51\) and GPT\-5\-chat \(\-8\.07\), but significantly more substantial for the compact Qwen3\-8B \(\-20\.81\)\. This pattern suggests that the enriched benchmark removes shortcuts based on surface forms and keyword associations, forcing models to resolve the underlying cultural context among plausible distractors\.

The ablation also restores a more expected capability hierarchy\. Without enrichment, Qwen3\-8B unexpectedly outperforms Qwen3\-235B\-A22B on the originalBLEnDsubset \(78\.06 vs\. 72\.57\), suggesting that the original format may permit shortcut exploitation\. After enrichment, Qwen3\-235B\-A22B becomes clearly more robust than Qwen3\-8B \(68\.06 vs\. 57\.25\)\. Therefore, linguistic enrichment is not merely a stylistic transformation; it improves diagnostic validity by makingSuperBLEnDbetter distinguish culturally grounded reasoning from shallow pattern matching\.

ModelBLEnDSuperBLEnDΔ\\DeltaQwen3\-235B\-A22B72\.5768\.06\-4\.51Qwen3\-8B78\.0657\.25\-20\.81GPT\-5\-chat78\.4570\.38\-8\.07

Table 2:Ablation ofSuperBLEnDon the 16 cultures overlapping withBLEnD\. Scores are average accuracy\.

## 5Conclusion

In this work, we introduced GaoYao, a holistic benchmark designed to map the full spectrum of multilingual intelligence\. Unlike fragmented prior efforts, GaoYao establishes a unified framework covering three cultural layers and nine cognitive dimensions\. Through a rigoroushuman\-in\-the\-loopconstruction pipeline, we addressed the critical scarcity of high\-quality resources for subjective and cultural tasks, proving that authentic evaluation requires native expertise rather than automated translation\. GaoYao stands as a robust alternative to English\-centric evaluations, guiding the field to move beyond surface\-level linguistic fluency towards deep cultural alignment and equitable global access\.

Future work include expanding coverage of domains \(*e\.g\.*, agent abilities\) and languages, and developing a dynamic leaderboard to keep up with the latest iterations of models\.

## 6Limitations

Despite our comprehensive efforts, GaoYao has several limitations that outline future directions:

\(1\) Domain and Task Coverage:Currently, GaoYao focuses primarily on general multilingual and multicultural capabilities\. We do not currently cover specialized vertical domains \(*e\.g\.*, legal, medical, financial\) or agentic capabilities \(*e\.g\.*, tool use, API calling\) in multilingual contexts\. However, we argue that a robust general\-purpose multilingual foundation is a prerequisite for these specialized abilities\. Constructing high\-quality benchmarks for such specific domains requires advanced expertise that falls outside the scope of this foundational work\.

\(2\) Static Nature of Benchmarking:The landscape of LLMs evolves at an unprecedented pace, and a static publication inevitably lags behind the release of the very latest models\. To address this, we have open\-sourced all test sets and evaluation code, ensuring full transparency and enabling third\-party model developers to easily verify their own systems against GaoYao\. Furthermore, we plan to launch and maintain a dynamic online leaderboard to continuously track the community’s progress\.

\(3\) Scalability of Human\-in\-the\-loop Pipeline:Our insistence on native expert curation and verification ensures high data quality but inherently limits scalability compared to fully automated pipelines\. Expanding to hundreds of low\-resource languages using this rigorous standard is resource\-intensive\. Yet, we believe that in the current era of ubiquitous machine\-generated content, establishing a high\-quality, human\-verified gold standard for a representative set of languages is more critical than broad but low\-quality coverage\.

\(4\) Task and Language Imbalance:GaoYao prioritizes high\-quality and expert\-verified coverage over perfectly uniform coverage across all sub\-layers\. As a result, some integrated resources naturally differ in language scope; for example,MGSMcovers 10 languages, while cultural resources such asSAGEandCultureScopeare limited to 2 languages/cultures\. This imbalance reflects the current scarcity of reliable multilingual and multicultural evaluation resources rather than an assumption that all languages are equally represented in every task\. Future versions will expand under\-covered task\-language pairs while preserving the same native\-expert verification standard\.

## 7Ethical Considerations

We reveal the following ethical considerations of GaoYao:

\(1\) Benchmark Usage and Contamination:We release GaoYao to facilitate the assessment of LLMs\. We explicitly discourage the inclusion of our test sets into model training corpora \(contamination\), which would render the evaluation validity null\. We urge the community to treat this benchmark as a diagnostic tool rather than a leaderboard to be gamed\.

\(2\) Annotator Fair Compensation and Well\-being:All data annotators and linguistic experts involved in the localization andSuperBLEnDsynthesis processes were full\-time employees of professional language service providers and participated as part of their regular duties\. They received standard professional salaries above local minimum wage, were informed about the intended use of the data, and were not recruited through unpaid labor or low\-paid crowdsourcing\. No personally identifiable information was collected\.

\(3\) Mitigation of Cultural Stereotypes:Constructing cultural benchmarks carries a risk of reinforcing stereotypes\. To mitigate this, we implemented a careful review process where native experts explicitly screened for offensive content, harmful generalizations, or political sensitivity as specified in Appendix[E](https://arxiv.org/html/2604.20225#A5)\. While we strive for neutrality, we acknowledge that cultural data may still reflect the subjective perspectives of the annotators\.

## References

- M\. Akhtar, A\. Reuel, P\. Soni, S\. Ahuja, P\. S\. Ammanamanchi, R\. Rawal, V\. Zouhar, S\. Yadav, C\. Whitehouse, D\. Ki,et al\.\(2026\)When ai benchmarks plateau: a systematic study of benchmark saturation\.arXiv preprint arXiv:2602\.16763\.Cited by:[§3\.2](https://arxiv.org/html/2604.20225#S3.SS2.SSS0.Px1.p3.1)\.
- L\. W\. Anderson and D\. R\. Krathwohl \(2001\)A taxonomy for learning, teaching, and assessing: a revision of bloom’s taxonomy of educational objectives: complete edition\.Addison Wesley Longman, Inc\.\.Cited by:[§1](https://arxiv.org/html/2604.20225#S1.p6.1),[§2\.1](https://arxiv.org/html/2604.20225#S2.SS1.SSS0.Px2.p1.1)\.
- Anthropic \(2026\)Note:Accessed: 2026\-01\-06External Links:[Link](https://www.anthropic.com/news/claude-sonnet-4-5)Cited by:[Table 7](https://arxiv.org/html/2604.20225#A7.T7.1.15.1)\.
- Ascend Tribe \(2025\)OpenPangu\-ultra\-moe\-718b\-v1\.1\.GitCode\.Note:[https://ai\.gitcode\.com/ascend\-tribe/openPangu\-Ultra\-MoE\-718B\-V1\.1](https://ai.gitcode.com/ascend-tribe/openPangu-Ultra-MoE-718B-V1.1)External Links:[Link](https://ai.gitcode.com/ascend-tribe/openPangu-Ultra-MoE-718B-V1.1)Cited by:[Table 7](https://arxiv.org/html/2604.20225#A7.T7.1.3.1),[1st item](https://arxiv.org/html/2604.20225#S3.I1.i1.p1.1)\.
- S\. Bai, Y\. Cai, R\. Chen, K\. Chen, X\. Chen, Z\. Cheng, L\. Deng, W\. Ding, C\. Gao, C\. Ge, W\. Ge, Z\. Guo, and et al\. \(2025\)Qwen3\-vl technical report\.External Links:2511\.21631,[Link](https://arxiv.org/abs/2511.21631)Cited by:[Table 7](https://arxiv.org/html/2604.20225#A7.T7.1.6.1)\.
- L\. Bandarkar, D\. Liang, B\. Muller, M\. Artetxe, S\. N\. Shukla, D\. Husa, N\. Goyal, A\. Krishnan, L\. Zettlemoyer, and M\. Khabsa \(2024\)The belebele benchmark: a parallel reading comprehension dataset in 122 language variants\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 749–775\.External Links:[Link](https://aclanthology.org/2024.acl-long.44/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.44)Cited by:[§A\.1](https://arxiv.org/html/2604.20225#A1.SS1.p1.1),[§A\.2](https://arxiv.org/html/2604.20225#A1.SS2.p1.1),[Table 3](https://arxiv.org/html/2604.20225#A1.T3.1.1.3.1),[§1](https://arxiv.org/html/2604.20225#S1.p2.1),[§2\.2\.1](https://arxiv.org/html/2604.20225#S2.SS2.SSS1.p3.1),[§3\.1\.2](https://arxiv.org/html/2604.20225#S3.SS1.SSS2.p2.1)\.
- L\. Barrault, O\. Bojar, M\. R\. Costa\-Jussa, C\. Federmann, M\. Fishel, Y\. Graham, B\. Haddow, M\. Huck, P\. Koehn, S\. Malmasi,et al\.\(2019\)Findings of the 2019 conference on machine translation \(wmt19\)\.Cited by:[§A\.1](https://arxiv.org/html/2604.20225#A1.SS1.p1.1)\.
- ByteDance \(2026\)SEED 1\.6\.Note:Accessed: 2026\-01\-05External Links:[Link](https://seed.bytedance.com/en/seed1_6)Cited by:[Table 7](https://arxiv.org/html/2604.20225#A7.T7.1.12.1)\.
- P\. Chen, S\. Ji, N\. Bogoychev, A\. Kutuzov, B\. Haddow, and K\. Heafield \(2024\)Monolingual or multilingual instruction tuning: which makes a better alpaca\.InFindings of the Association for Computational Linguistics: EACL 2024,pp\. 1347–1356\.Cited by:[§A\.3](https://arxiv.org/html/2604.20225#A1.SS3.p1.1)\.
- W\. Chiang, L\. Zheng, Y\. Sheng, A\. N\. Angelopoulos, T\. Li, D\. Li, B\. Zhu, H\. Zhang, M\. Jordan, J\. E\. Gonzalez,et al\.\(2024\)Chatbot arena: an open platform for evaluating llms by human preference\.InForty\-first International Conference on Machine Learning,Cited by:[Appendix C](https://arxiv.org/html/2604.20225#A3.p1.1)\.
- G\. Comanici, E\. Bieber, M\. Schaekermann, I\. Pasupat, N\. Sachdeva, I\. Dhillon, M\. Blistein, O\. Ram, D\. Zhang, E\. Rosen,et al\.\(2025\)Gemini 2\.5: pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities\.arXiv preprint arXiv:2507\.06261\.Cited by:[Table 7](https://arxiv.org/html/2604.20225#A7.T7.1.14.1),[2nd item](https://arxiv.org/html/2604.20225#S3.I1.i2.p1.1)\.
- DeepSeek\-AI \(2025\)DeepSeek\-r1: incentivizing reasoning capability in llms via reinforcement learning\.InarXiv preprint arXiv:2501\.12948,Cited by:[Table 7](https://arxiv.org/html/2604.20225#A7.T7.1.7.1),[3rd item](https://arxiv.org/html/2604.20225#S3.I1.i3.p1.1),[§3\.2](https://arxiv.org/html/2604.20225#S3.SS2.SSS0.Px4.p1.1)\.
- DeepSeek \(2025\)DeepSeek\-v3\.1 release\.Note:[https://api\-docs\.deepseek\.com/news/news250821](https://api-docs.deepseek.com/news/news250821)Cited by:[§A\.3](https://arxiv.org/html/2604.20225#A1.SS3.p2.1),[Table 7](https://arxiv.org/html/2604.20225#A7.T7.1.4.1),[§3\.1](https://arxiv.org/html/2604.20225#S3.SS1.p1.1)\.
- K\. Fujii, T\. Nakamura, M\. Loem, H\. Iida, M\. Ohi, K\. Hattori, H\. Shota, S\. Mizuki, R\. Yokota, and N\. Okazaki \(2024\)Continual pre\-training for cross\-lingual llm adaptation: enhancing japanese language capabilities\.InFirst Conference on Language Modeling,Cited by:[§A\.3](https://arxiv.org/html/2604.20225#A1.SS3.p1.1)\.
- N\. Goyal, C\. Gao, V\. Chaudhary, P\. Chen, G\. Wenzek, D\. Ju, S\. Krishnan, M\. Ranzato, F\. Guzmán, and A\. Fan \(2022\)The flores\-101 evaluation benchmark for low\-resource and multilingual machine translation\.Transactions of the Association for Computational Linguistics10,pp\. 522–538\.Cited by:[§A\.1](https://arxiv.org/html/2604.20225#A1.SS1.p1.1),[§A\.2](https://arxiv.org/html/2604.20225#A1.SS2.p1.1),[Table 3](https://arxiv.org/html/2604.20225#A1.T3.1.1.2.1),[§2\.2\.1](https://arxiv.org/html/2604.20225#S2.SS2.SSS1.p4.1),[§3\.1\.2](https://arxiv.org/html/2604.20225#S3.SS1.SSS2.p2.1)\.
- S\. Guo, S\. Jiang, Q\. He, Y\. Xiao, J\. Liang, B\. Yude, M\. He, S\. Tao, and L\. Zhang \(2025\)Do large language models truly understand cross\-cultural differences?\.External Links:2512\.07075,[Link](https://arxiv.org/abs/2512.07075)Cited by:[§A\.1](https://arxiv.org/html/2604.20225#A1.SS1.p2.1),[§A\.2](https://arxiv.org/html/2604.20225#A1.SS2.p1.1),[Table 3](https://arxiv.org/html/2604.20225#A1.T3.1.1.7.1),[§2\.2\.1](https://arxiv.org/html/2604.20225#S2.SS2.SSS1.p6.1)\.
- E\. T\. Hall \(1976\)Beyond culture\.Anchor\.Cited by:[§1](https://arxiv.org/html/2604.20225#S1.p6.1),[§2\.1](https://arxiv.org/html/2604.20225#S2.SS1.SSS0.Px1.p1.1)\.
- M\. He, Y\. Liu, S\. Tao, Y\. Luo, H\. Zeng, C\. Su, L\. Zhang, H\. Ma, D\. Wei, W\. Meng,et al\.\(2025\)R1\-t1: fully incentivizing translation capability in llms via reasoning learning\.arXiv preprint arXiv:2502\.19735\.Cited by:[§3\.2](https://arxiv.org/html/2604.20225#S3.SS2.SSS0.Px4.p1.1)\.
- A\. Q\. Jiang, A\. Sablayrolles, A\. Roux, A\. Mensch, B\. Savary, C\. Bamford, D\. S\. Chaplot, D\. de las Casas, E\. B\. Hanna, and et al\. \(2024\)Mixtral of experts\.External Links:2401\.04088,[Link](https://arxiv.org/abs/2401.04088)Cited by:[Table 7](https://arxiv.org/html/2604.20225#A7.T7.1.27.1)\.
- P\. Joshi, S\. Santy, A\. Budhiraja, K\. Bali, and M\. Choudhury \(2020\)The state and fate of linguistic diversity and inclusion in the nlp world\.InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics,pp\. 6282–6293\.Cited by:[Appendix B](https://arxiv.org/html/2604.20225#A2.p1.1),[Figure 9](https://arxiv.org/html/2604.20225#A3.F9),[§3\.1\.1](https://arxiv.org/html/2604.20225#S3.SS1.SSS1.p1.1)\.
- T\. Kocmi, E\. Avramidis, R\. Bawden, O\. Bojar, A\. Dvorkovich, C\. Federmann, M\. Fishel, M\. Freitag, T\. Gowda, R\. Grundkiewicz,et al\.\(2024\)Findings of the wmt24 general machine translation shared task: the llm era is here but mt is not solved yet\.InProceedings of the Ninth Conference on Machine Translation,pp\. 1–46\.Cited by:[§A\.1](https://arxiv.org/html/2604.20225#A1.SS1.p1.1)\.
- W\. Lai, M\. Mesgar, and A\. Fraser \(2024\)LLMs beyond english: scaling the multilingual capability of llms with cross\-lingual feedback\.arXiv preprint arXiv:2406\.01771\.Cited by:[§A\.3](https://arxiv.org/html/2604.20225#A1.SS3.p1.1)\.
- H\. Li, L\. Ding, M\. Fang, and D\. Tao \(2024\)Revisiting catastrophic forgetting in large language model tuning\.InFindings of the Association for Computational Linguistics: EMNLP 2024,pp\. 4297–4308\.Cited by:[§A\.3](https://arxiv.org/html/2604.20225#A1.SS3.p1.1)\.
- X\. Li, T\. Zhang, Y\. Dubois, R\. Taori, I\. Gulrajani, C\. Guestrin, P\. Liang, and T\. B\. Hashimoto \(2023\)AlpacaEval: an automatic evaluator of instruction\-following models\.GitHub\.Note:[https://github\.com/tatsu\-lab/alpaca\_eval](https://github.com/tatsu-lab/alpaca_eval)Cited by:[§A\.1](https://arxiv.org/html/2604.20225#A1.SS1.p2.1),[§1](https://arxiv.org/html/2604.20225#S1.p3.1),[§2\.2\.2](https://arxiv.org/html/2604.20225#S2.SS2.SSS2.p1.1),[§3\.1\.2](https://arxiv.org/html/2604.20225#S3.SS1.SSS2.p3.1)\.
- Y\. Liu, M\. Xu, S\. Wang, L\. Yang, H\. Wang, Z\. Liu, C\. Kong, Y\. Chen, M\. Sun, and E\. Yang \(2024\)OMGEval: an open multilingual generative evaluation benchmark for large language models\.arXiv preprint arXiv:2402\.13524\.Cited by:[§A\.1](https://arxiv.org/html/2604.20225#A1.SS1.p2.1),[Table 3](https://arxiv.org/html/2604.20225#A1.T3.1.1.6.1),[§1](https://arxiv.org/html/2604.20225#S1.p3.1),[§1](https://arxiv.org/html/2604.20225#S1.p4.1)\.
- Y\. Liu, Z\. Chen, S\. Xu, M\. He, S\. Tao, W\. Meng, Y\. Xie, T\. Han, C\. Zhao, J\. Du,et al\.\(2025a\)R\-log: incentivizing log analysis capability in llms via reasoning\-based reinforcement learning\.arXiv preprint arXiv:2509\.25987\.Cited by:[§3\.2](https://arxiv.org/html/2604.20225#S3.SS2.SSS0.Px4.p1.1)\.
- Y\. Liu, C\. Zhao, X\. Yang, H\. Zeng, S\. Tao, W\. Meng, M\. He, C\. Su, Y\. Yu, H\. Ma,et al\.\(2025b\)MIDB: multilingual instruction data booster for enhancing multilingual instruction synthesis\.arXiv preprint arXiv:2505\.17671\.Cited by:[§A\.3](https://arxiv.org/html/2604.20225#A1.SS3.p2.1)\.
- Meta AI \(2024\)Introducing Llama 3\.1: our most capable models to date\.Note:Accessed: 2026\-01\-05External Links:[Link](https://ai.meta.com/blog/meta-llama-3-1/)Cited by:[Table 7](https://arxiv.org/html/2604.20225#A7.T7.1.25.1),[Table 7](https://arxiv.org/html/2604.20225#A7.T7.1.26.1),[Table 7](https://arxiv.org/html/2604.20225#A7.T7.1.8.1)\.
- J\. Myung, N\. Lee, Y\. Zhou, J\. Jin, R\. Putri, D\. Antypas, H\. Borkakoty, E\. Kim, C\. Perez\-Almendros, A\. A\. Ayele,et al\.\(2024\)Blend: a benchmark for llms on everyday knowledge in diverse cultures and languages\.Advances in Neural Information Processing Systems37,pp\. 78104–78146\.Cited by:[§A\.1](https://arxiv.org/html/2604.20225#A1.SS1.p2.1),[Appendix E](https://arxiv.org/html/2604.20225#A5.SS0.SSS0.Px2.p1.1),[§2\.2\.3](https://arxiv.org/html/2604.20225#S2.SS2.SSS3.p1.1)\.
- T\. Nakamura, M\. Mishra, S\. Tedeschi, Y\. Chai, J\. T\. Stillerman, F\. Friedrich, P\. Yadav, T\. Laud, V\. M\. Chien, T\. Y\. Zhuo,et al\.\(2025\)Aurora\-m: open source continual pre\-training for multilingual language and code\.InProceedings of the 31st International Conference on Computational Linguistics: Industry Track,pp\. 656–678\.Cited by:[§A\.3](https://arxiv.org/html/2604.20225#A1.SS3.p1.1)\.
- A\. Omnilingual, G\. Keren, A\. Kozhevnikov, Y\. Meng, C\. Ropers, M\. Setzler, S\. Wang, I\. Adebara, M\. Auli, C\. Balioglu,et al\.\(2025\)Omnilingual asr: open\-source multilingual speech recognition for 1600\+ languages\.arXiv preprint arXiv:2511\.09690\.Cited by:[§A\.3](https://arxiv.org/html/2604.20225#A1.SS3.p2.1)\.
- OpenAI, :, A\. Hurst, A\. Lerer, A\. P\. Goucher, A\. Perelman, A\. Ramesh, A\. Clark, A\. Ostrow, A\. Welihinda, A\. Hayes, A\. Radford, and et al\. \(2024\)GPT\-4o system card\.External Links:2410\.21276,[Link](https://arxiv.org/abs/2410.21276)Cited by:[Table 7](https://arxiv.org/html/2604.20225#A7.T7.1.20.1)\.
- OpenAI \(2024\)Multilingual massive multitask language understanding \(mmmlu\)\.Note:[https://huggingface\.co/datasets/openai/MMMLU](https://huggingface.co/datasets/openai/MMMLU)Cited by:[§2\.2\.1](https://arxiv.org/html/2604.20225#S2.SS2.SSS1.p2.1),[§3\.1\.2](https://arxiv.org/html/2604.20225#S3.SS1.SSS2.p2.1)\.
- OpenAI \(2025a\)GPT\-5 system card\.External Links:[Link](https://cdn.openai.com/gpt-5-system-card.pdf)Cited by:[§3\.1](https://arxiv.org/html/2604.20225#S3.SS1.p1.1)\.
- OpenAI \(2025b\)Note:Accessed: 2026\-01\-06External Links:[Link](https://openai.com/en/index/introducing-o3-and-o4-mini/)Cited by:[Table 7](https://arxiv.org/html/2604.20225#A7.T7.1.17.1),[Table 7](https://arxiv.org/html/2604.20225#A7.T7.1.18.1)\.
- OpenAI \(2026\)Note:Accessed: 2026\-01\-06External Links:[Link](https://openai.com/zh-Hans-CN/index/introducing-gpt-5/)Cited by:[Table 7](https://arxiv.org/html/2604.20225#A7.T7.1.19.1)\.
- D\. Pomerenke, J\. Nothnagel, and S\. Ostermann \(2025\)The ai language proficiency monitor–tracking the progress of llms on multilingual benchmarks\.arXiv preprint arXiv:2507\.08538\.Cited by:[§A\.1](https://arxiv.org/html/2604.20225#A1.SS1.p1.1),[§1](https://arxiv.org/html/2604.20225#S1.p4.1)\.
- R\. Rei, C\. Stewart, A\. C\. Farinha, and A\. Lavie \(2020\)COMET: a neural framework for MT evaluation\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),B\. Webber, T\. Cohn, Y\. He, and Y\. Liu \(Eds\.\),Online,pp\. 2685–2702\.External Links:[Link](https://aclanthology.org/2020.emnlp-main.213/),[Document](https://dx.doi.org/10.18653/v1/2020.emnlp-main.213)Cited by:[§A\.1](https://arxiv.org/html/2604.20225#A1.SS1.p1.1)\.
- A\. Romanou, N\. Foroutan, A\. Sotnikova, S\. H\. Nelaturu, S\. Singh, R\. Maheshwary, M\. Altomare, Z\. Chen, M\. A\. Haggag, A\. Amayuelas,et al\.\(2025\)INCLUDE: evaluating multilingual language understanding with regional knowledge\.InThe Thirteenth International Conference on Learning Representations,Cited by:[§A\.1](https://arxiv.org/html/2604.20225#A1.SS1.p1.1),[Table 3](https://arxiv.org/html/2604.20225#A1.T3.1.1.4.1),[§1](https://arxiv.org/html/2604.20225#S1.p2.1),[§2\.2\.1](https://arxiv.org/html/2604.20225#S2.SS2.SSS1.p2.1),[§3\.1\.2](https://arxiv.org/html/2604.20225#S3.SS1.SSS2.p2.1)\.
- J\. Rystrøm, H\. R\. Kirk, and S\. Hale \(2025\)Multilingual\!= multicultural: evaluating gaps between multilingual capabilities and cultural alignment in llms\.arXiv preprint arXiv:2502\.16534\.Cited by:[§A\.1](https://arxiv.org/html/2604.20225#A1.SS1.p2.1),[§A\.3](https://arxiv.org/html/2604.20225#A1.SS3.p2.1),[3rd item](https://arxiv.org/html/2604.20225#S2.I2.i3.p1.1),[§2\.2\.1](https://arxiv.org/html/2604.20225#S2.SS2.SSS1.p6.1)\.
- E\. H\. Schein \(2010\)Organizational culture and leadership\.Vol\.2,John Wiley & Sons\.Cited by:[§2\.1](https://arxiv.org/html/2604.20225#S2.SS1.SSS0.Px1.p1.1)\.
- F\. Shi, M\. Suzgun, M\. Freitag, X\. Wang, S\. Srivats, S\. Vosoughi, H\. W\. Chung, Y\. Tay, S\. Ruder, D\. Zhou,et al\.\(2023\)Language models are multilingual chain\-of\-thought reasoners\.InThe Eleventh International Conference on Learning Representations,Cited by:[§2\.2\.1](https://arxiv.org/html/2604.20225#S2.SS2.SSS1.p5.1)\.
- G\. Team, A\. Kamath, J\. Ferret, S\. Pathak, N\. Vieillard, R\. Merhej, S\. Perrin, T\. Matejovicova, A\. Ramé, M\. Rivière, and et al\. \(2025a\)Gemma 3 technical report\.External Links:2503\.19786,[Link](https://arxiv.org/abs/2503.19786)Cited by:[Table 7](https://arxiv.org/html/2604.20225#A7.T7.1.24.1)\.
- K\. Team, Y\. Bai, Y\. Bao, G\. Chen, J\. Chen, N\. Chen, R\. Chen, Y\. Chen, Y\. Chen, Y\. Chen, Z\. Chen, J\. Cui, H\. Ding, and et al\. \(2025b\)Kimi k2: open agentic intelligence\.External Links:2507\.20534,[Link](https://arxiv.org/abs/2507.20534)Cited by:[Table 7](https://arxiv.org/html/2604.20225#A7.T7.1.10.1)\.
- Q\. Team \(2025\)Qwen3\-max: just scale it\.Cited by:[Table 7](https://arxiv.org/html/2604.20225#A7.T7.1.13.1)\.
- H\. Touvron, T\. Lavril, G\. Izacard, X\. Martinet, M\. Lachaux, T\. Lacroix, B\. Rozière, N\. Goyal, E\. Hambro, F\. Azhar,et al\.\(2023a\)Llama: open and efficient foundation language models\.arXiv preprint arXiv:2302\.13971\.Cited by:[§A\.3](https://arxiv.org/html/2604.20225#A1.SS3.p1.1)\.
- H\. Touvron, L\. Martin, K\. Stone, P\. Albert, A\. Almahairi, Y\. Babaei, N\. Bashlykov, S\. Batra, P\. Bhargava, S\. Bhosale,et al\.\(2023b\)Llama 2: open foundation and fine\-tuned chat models\.arXiv preprint arXiv:2307\.09288\.Cited by:[§A\.3](https://arxiv.org/html/2604.20225#A1.SS3.p1.1)\.
- xAI \(2025\)Note:Accessed: 2026\-01\-06External Links:[Link](https://x.ai/news/grok-3)Cited by:[Table 7](https://arxiv.org/html/2604.20225#A7.T7.1.16.1)\.
- A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv,et al\.\(2025\)Qwen3 technical report\.arXiv preprint arXiv:2505\.09388\.Cited by:[§A\.3](https://arxiv.org/html/2604.20225#A1.SS3.p2.1),[Table 7](https://arxiv.org/html/2604.20225#A7.T7.1.22.1),[Table 7](https://arxiv.org/html/2604.20225#A7.T7.1.23.1),[Table 7](https://arxiv.org/html/2604.20225#A7.T7.1.5.1),[§3\.1\.2](https://arxiv.org/html/2604.20225#S3.SS1.SSS2.p3.1),[§3\.1](https://arxiv.org/html/2604.20225#S3.SS1.p1.1),[§3\.2](https://arxiv.org/html/2604.20225#S3.SS2.SSS0.Px1.p3.1)\.
- B\. Yao, M\. Jiang, T\. Bobinac, D\. Yang, and J\. Hu \(2023\)Benchmarking machine translation with cultural awareness\.arXiv preprint arXiv:2305\.14328\.Cited by:[§A\.1](https://arxiv.org/html/2604.20225#A1.SS1.p2.1)\.
- A\. Zeng, X\. Lv, Q\. Zheng, Z\. Hou, B\. Chen, C\. Xie, C\. Wang, D\. Yin, H\. Zeng, J\. Zhang, K\. Wang, and et al\. \(2025\)GLM\-4\.5: agentic, reasoning, and coding \(arc\) foundation models\.External Links:2508\.06471,[Link](https://arxiv.org/abs/2508.06471)Cited by:[Table 7](https://arxiv.org/html/2604.20225#A7.T7.1.9.1)\.
- J\. Zhang, S\. Jiang, S\. Guo, S\. Chen, Y\. Xiao, H\. Feng, J\. Liang, M\. HE, S\. Tao, and H\. Ma \(2025\)CultureScope: a dimensional lens for probing cultural understanding in llms\.arXiv preprint arXiv:2509\.16188\.Cited by:[§A\.1](https://arxiv.org/html/2604.20225#A1.SS1.p2.1),[§A\.2](https://arxiv.org/html/2604.20225#A1.SS2.p1.1),[Table 3](https://arxiv.org/html/2604.20225#A1.T3.1.1.8.1),[3rd item](https://arxiv.org/html/2604.20225#S2.I1.i3.p1.1),[§2\.2\.1](https://arxiv.org/html/2604.20225#S2.SS2.SSS1.p6.1)\.
- Z\. Zhang, D\. Lee, Y\. Fang, W\. Yu, M\. Jia, M\. Jiang, and F\. Barbieri \(2024\)PLUG: leveraging pivot language in cross\-lingual instruction tuning\.InProceedings of the 62th Annual Meeting of the Association for Computational Linguistics,Cited by:[§A\.1](https://arxiv.org/html/2604.20225#A1.SS1.p2.1),[§A\.2](https://arxiv.org/html/2604.20225#A1.SS2.p1.1),[§A\.3](https://arxiv.org/html/2604.20225#A1.SS3.p1.1),[Table 3](https://arxiv.org/html/2604.20225#A1.T3.1.1.5.1),[§1](https://arxiv.org/html/2604.20225#S1.p3.1)\.
- Q\. Zhao \(1988\)A study of dragonology, east and west\.University of Massachusetts Amherst\.Cited by:[2nd item](https://arxiv.org/html/2604.20225#S2.I1.i2.p1.1)\.
- L\. Zheng, W\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Lin, Z\. Li, D\. Li, E\. Xing,et al\.\(2023\)Judging llm\-as\-a\-judge with mt\-bench and chatbot arena\.Advances in Neural Information Processing Systems36,pp\. 46595–46623\.Cited by:[§A\.1](https://arxiv.org/html/2604.20225#A1.SS1.p2.1),[§1](https://arxiv.org/html/2604.20225#S1.p3.1),[§2\.2\.2](https://arxiv.org/html/2604.20225#S2.SS2.SSS2.p1.1),[§3\.1\.2](https://arxiv.org/html/2604.20225#S3.SS1.SSS2.p3.1)\.
- R\. Zhu, Y\. Wang, D\. Lin, M\. Jendryke, M\. Xie, J\. Guo, and L\. Meng \(2021\)Exploring the rich\-club characteristic in internal migration: evidence from chinese chunyun migration\.Cities114,pp\. 103198\.Cited by:[3rd item](https://arxiv.org/html/2604.20225#S2.I1.i3.p1.1)\.
- S\. Zhu, L\. Pan, D\. Jian, and D\. Xiong \(2025\)Overcoming language barriers via machine translation with sparse mixture\-of\-experts fusion of large language models\.Information Processing & Management62\(3\),pp\. 104078\.Cited by:[§A\.3](https://arxiv.org/html/2604.20225#A1.SS3.p1.1)\.
- W\. Zhu, Y\. Lv, Q\. Dong, F\. Yuan, J\. Xu, S\. Huang, L\. Kong, J\. Chen, and L\. Li \(2023\)Extrapolating large language models to non\-english by aligning languages\.arXiv preprint arXiv:2308\.04948\.Cited by:[§A\.3](https://arxiv.org/html/2604.20225#A1.SS3.p1.1)\.
- V\. Zouhar, P\. Chen, T\. K\. Lam, N\. Moghe, and B\. Haddow \(2024\)Pitfalls and outlooks in using COMET\.InProceedings of the Ninth Conference on Machine Translation,Miami, Florida, USA,pp\. 1272–1288\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.wmt-1.121),[Link](https://aclanthology.org/2024.wmt-1.121/)Cited by:[§A\.1](https://arxiv.org/html/2604.20225#A1.SS1.p1.1)\.

## Appendix ARelated Work

BenchmarkFocus DomainTask Diversity\# Total Langs\# Subj\. Langs†\# CulturesFlores\-101\(Goyalet al\.,[2022](https://arxiv.org/html/2604.20225#bib.bib18)\)TranslationSingle101\-\-Belebele\(Bandarkaret al\.,[2024](https://arxiv.org/html/2604.20225#bib.bib76)\)ReadingSingle122\-\-Include\(Romanouet al\.,[2025](https://arxiv.org/html/2604.20225#bib.bib81)\)KnowledgeSingle44\-\-X\-AlpacaEval\(Zhanget al\.,[2024](https://arxiv.org/html/2604.20225#bib.bib70)\)InstructionSingle44\-OMGEval\(Liuet al\.,[2024](https://arxiv.org/html/2604.20225#bib.bib71)\)InstructionSingle99\-SAGE\(Guoet al\.,[2025](https://arxiv.org/html/2604.20225#bib.bib85)\)CultureSingle222CultureScope\(Zhanget al\.,[2025](https://arxiv.org/html/2604.20225#bib.bib80)\)CultureSingle222GaoYao \(Ours\)ComprehensiveAll \(9 Sub\-layers\)261934†Subj\. Langsrefers to languages supported for open\-ended subjective tasks \(*e\.g\.*, Instruction Following & Multi\-turn Dialogue\)\.

Table 3:Comparison between GaoYao and other representative benchmarks\. GaoYao distinguishes itself by its comprehensive scope, specifically surpassing others in subjective task coverage \(19 languages\) and cultural breadth \(34 cultures\) with rigorous human verification\.### A\.1Evaluation of Multilingual and Multicultural Capabilities

Evaluating the capabilities of LLMs across diverse languages and cultures is critical for their global deployment\. Traditional evaluations primarily focus on MT and objective understanding tasks\.Goyalet al\.\([2022](https://arxiv.org/html/2604.20225#bib.bib18)\)introducedFlores\-101, a benchmark covering massive languages for MT, while the WMT shared tasks\(Kocmiet al\.,[2024](https://arxiv.org/html/2604.20225#bib.bib8); Barraultet al\.,[2019](https://arxiv.org/html/2604.20225#bib.bib61)\)remain the standard for evaluating translation quality using metrics like COMET\(Reiet al\.,[2020](https://arxiv.org/html/2604.20225#bib.bib44); Zouharet al\.,[2024](https://arxiv.org/html/2604.20225#bib.bib60)\)\. Beyond translation, recent benchmarks have expanded to broader understanding capabilities\.Belebele\(Bandarkaret al\.,[2024](https://arxiv.org/html/2604.20225#bib.bib76)\)evaluates parallel reading comprehension across 122 language variants, andInclude\(Romanouet al\.,[2025](https://arxiv.org/html/2604.20225#bib.bib81)\)assesses multilingual understanding with regional knowledge\. Similarly,Pomerenkeet al\.\([2025](https://arxiv.org/html/2604.20225#bib.bib67)\)track LLM progress on various multilingual benchmarks\.

However, existing evaluations often exhibit limitations in scope and depth\. Many benchmarks rely heavily on objective formats such as multiple\-choice questions \(MCQs\), neglecting subjective generation tasks that better reflect real\-world usage\. WhileAlpacaEval\(Liet al\.,[2023](https://arxiv.org/html/2604.20225#bib.bib68)\)andMT\-Bench\(Zhenget al\.,[2023](https://arxiv.org/html/2604.20225#bib.bib69)\)provide robust evaluations for instruction following and dialogue, they are predominantly English\-centric\. Multilingual extensions likeOMGEval\(Liuet al\.,[2024](https://arxiv.org/html/2604.20225#bib.bib71)\)andX\-AlpacaEval\(Zhanget al\.,[2024](https://arxiv.org/html/2604.20225#bib.bib70)\)exist but often suffer from limited language coverage\. Furthermore, cultural evaluation remains under\-explored\.Yaoet al\.\([2023](https://arxiv.org/html/2604.20225#bib.bib57)\)andZhanget al\.\([2025](https://arxiv.org/html/2604.20225#bib.bib80)\)have initiated efforts to probe cultural awareness, andMyunget al\.\([2024](https://arxiv.org/html/2604.20225#bib.bib82)\)introducedBLEnDfor everyday cultural knowledge\. Yet, these datasets often suffer from limited culture coverages, and the problem of relying on direct translation which retains source\-culture bias still exists\(Rystrømet al\.,[2025](https://arxiv.org/html/2604.20225#bib.bib84); Guoet al\.,[2025](https://arxiv.org/html/2604.20225#bib.bib85)\)\.

### A\.2Comparison with Existing Benchmarks\.

As summarized in Table[3](https://arxiv.org/html/2604.20225#A1.T3), existing benchmarks typically exhibit a trade\-off between breadth and depth\. Massive\-scale benchmarks \(*e\.g\.*,Flores\-101Goyalet al\.\([2022](https://arxiv.org/html/2604.20225#bib.bib18)\),BelebeleBandarkaret al\.\([2024](https://arxiv.org/html/2604.20225#bib.bib76)\)\) cover over 100 languages but are restricted to objective tasks \(Translation and Reading\)\. Conversely, benchmarks focusing on subjective tasks \(*e\.g\.*,X\-AlpacaEvalZhanget al\.\([2024](https://arxiv.org/html/2604.20225#bib.bib70)\)\) are severely limited in language coverage \(<10<10\)\. Crucially, even dedicated cultural benchmarks likeSAGEGuoet al\.\([2025](https://arxiv.org/html/2604.20225#bib.bib85)\)andCultureScopeZhanget al\.\([2025](https://arxiv.org/html/2604.20225#bib.bib80)\)are constrained to specific bilingual dyads \(2 languages, 2 cultures\)\. Compared with these existing works, GaoYao establishes a more comprehensive framework\. It integrates objective tasks with rigorously localized subjective tasks across 19 languages\. Moreover, it significantly expands cultural evaluation throughSuperBLEnD, a human\-verified dataset covering 34 cultures, addressing the critical gap in deep, authentic multicultural assessment\.

### A\.3LLMs for Multilingual and Multicultural Tasks

The multilingual abilities of early LLMs are relatively underdevelopedLaiet al\.\([2024](https://arxiv.org/html/2604.20225#bib.bib104)\)due to the predominance of English in their pretraining data, such as the early LLaMA seriesTouvronet al\.\([2023a](https://arxiv.org/html/2604.20225#bib.bib111)\),*e\.g\.*, the ratio of non\-English languages in pretraining corpus of LLaMA\-2Touvronet al\.\([2023b](https://arxiv.org/html/2604.20225#bib.bib112)\)is merely around 2%\. To improve the multilingual abilities of existing foundation LLMs, researchers have explored specialized tuning strategies\. Many methods fouces on improving the quantity and quality of post\-training using curated multilingual dataset and refined training strategiesChenet al\.\([2024](https://arxiv.org/html/2604.20225#bib.bib113)\); Zhuet al\.\([2023](https://arxiv.org/html/2604.20225#bib.bib114)\); Zhanget al\.\([2024](https://arxiv.org/html/2604.20225#bib.bib70)\), while others leveraged continual pre\-training\(Fujiiet al\.,[2024](https://arxiv.org/html/2604.20225#bib.bib29); Nakamuraet al\.,[2025](https://arxiv.org/html/2604.20225#bib.bib105)\)and sparse Mixture\-of\-Experts architectures\(Zhuet al\.,[2025](https://arxiv.org/html/2604.20225#bib.bib31)\)to mitigate catastrophic forgetting\(Liet al\.,[2024](https://arxiv.org/html/2604.20225#bib.bib34)\)\.

More recently, the focus has shifted towards monolingual capabilities into multilingual tasks\. The ratio of multilingual data in pre\-training corpus of recent LLMs is significantly growing\(Yanget al\.,[2025](https://arxiv.org/html/2604.20225#bib.bib94); DeepSeek,[2025](https://arxiv.org/html/2604.20225#bib.bib93)\)\. In specific tasks \(*e\.g\.*, automatic speech recognition\), the supported number of languages reaches 1600\+Omnilingualet al\.\([2025](https://arxiv.org/html/2604.20225#bib.bib106)\)\. Despite the dramatic expansion of language coverage, the community has raised questions on the native level of responses in subjective tasksLiuet al\.\([2025b](https://arxiv.org/html/2604.20225#bib.bib108)\)and authentic cultural experience for global usersRystrømet al\.\([2025](https://arxiv.org/html/2604.20225#bib.bib84)\), where existing LLMs still lag behind\. The three curated subjective and cultural test sets in GaoYao can serve as a pioneering step for bridging the gap for existing multilingual LLM to achieve truly equitable AI access\.

## Appendix BDetailed Information of Languages Supported by GaoYao

Table[4](https://arxiv.org/html/2604.20225#A3.T4)presents the languages in our dataset, mapping full names to short codes, and detailing their geographical and resource\-level information\. Resource levels are determined using theJoshiet al\.\([2020](https://arxiv.org/html/2604.20225#bib.bib86)\)taxonomy, which rates languages from a set of 2,485 on a scale reflecting resource availability\. According to this scale, we classify languages scoring 5 as high\-resource \(*e\.g\., Chinese*\), a score of 4 as mid\-resource \(*e\.g\.*, Turkish\), and a score of≤\\leq3 as low\-resource\. Fig\.[9](https://arxiv.org/html/2604.20225#A3.F9)displays the impact on LLM’s best scores caused by resource popularity of languages, suggesting that LLM capabilities for languages with lower resources are relatively underdeveloped compared with high\-resource ones\.

## Appendix CReliability Analysis of GaoYao

A robust benchmark must strike a balance: it should reflect current user needs \(ecological validity\) while probing capabilities that users may not yet explicitly request but are essential for advanced intelligence \(comprehensive coverage\)\. To assess this, we compared the topic distribution of GaoYao against 140k authentic user queries sampled fromLMSYS Chatbot Arena\(Chianget al\.,[2024](https://arxiv.org/html/2604.20225#bib.bib87)\)\. We employed a multilingual tagging model to categorize both datasets and calculated the Tag Semantic Alignment \(TSA\) score,*i\.e\.*, the average semantic similarity between tags extracted from two groups\.

![Refer to caption](https://arxiv.org/html/2604.20225v1/images/Languages_distribution_per_popularity.png)Figure 9:Impact of language resource popularity on best performance achieved by LLMs\. Vertical dashed lines represent group averages\. The taxonomy on language resource\-levels is based onJoshiet al\.\([2020](https://arxiv.org/html/2604.20225#bib.bib86)\)\.Table[5](https://arxiv.org/html/2604.20225#A3.T5)presents the results\. We observe aHigh Alignment in Utility Tasks: TheInstruction Following\(TSA 0\.89\) andDialogue\(TSA 0\.79\) sub\-layers show strong correlation with real\-world queries, where tags like "IT Technology" and "Creative Writing" dominate\. This confirms that our localization of subjective benchmarks accurately captures the core interaction patterns of global users\.

Conversely, we seeLow Alignment in Specialized Domains: Layers likeMath\(0\.03\) andCross\-Cultural SA\(0\.13\) show lower alignment\. This is likely because real\-world users currently employ LLMs less frequently for complex tasks like advanced logic proofs or nuanced cultural philosophy\. However, these "long\-tail" capabilities are precisely where SOTA models differentiate themselves\. By including these low\-TSA but high\-value domains, GaoYao prevents overfitting to "average" user behavior and ensures models are also evaluated on the frontier of capability\.

LanguageCodeGeographical GroupResource LevelEnglishENWestern EuropeHigh ResourceFrenchFRWestern EuropeHigh ResourceGermanDEWestern EuropeHigh ResourceSpanishESWestern EuropeHigh ResourceDutchNLWestern EuropeMedium ResourceItalianITWestern EuropeMedium ResourcePortuguesePTWestern EuropeMedium ResourceGreekELWestern EuropeLow ResourceCzechCSEastern EuropeMedium ResourcePolishPLEastern EuropeMedium ResourceRussianRUEastern EuropeMedium ResourceChineseZHEast Asia & Southeast AsiaHigh ResourceJapaneseJAEast Asia & Southeast AsiaHigh ResourceKoreanKOEast Asia & Southeast AsiaMedium ResourceVietnameseVIEast Asia & Southeast AsiaMedium ResourceIndonesianIDEast Asia & Southeast AsiaLow ResourceMalayMSEast Asia & Southeast AsiaLow ResourceTagalogTLEast Asia & Southeast AsiaLow ResourceThaiTHEast Asia & Southeast AsiaLow ResourceArabicARMiddle East & AfricaHigh ResourceTurkishTRMiddle East & AfricaMedium ResourceSwahiliSWMiddle East & AfricaLow ResourceYorubaYOMiddle East & AfricaLow ResourceHindiHISouth AsiaMedium ResourceBengaliBNSouth AsiaLow ResourceTeluguTESouth AsiaLow Resource

Table 4:Language mapping: codes, geographical groups, and resource levels\.Sub\-layerTSATop\-10 Frequent Tags1Cross\-Cultural \(SA\)0\.1323Educational Philosophy, Philosophy/Religion, Values, Higher Education, Language/Writing, Humanities, Workplace Life, Politics2Cross\-Cultural \(SB\)0\.4548Workplace Life, Social Customs, Food & Cooking, IT Technology, Higher Education, Sports, Artifacts, AI3Mono\-Cultural \(CS\)0\.2116Workplace Life, Social Customs, Educational Philosophy, Humanities, Language/Writing, Business Management, Values, Banking4Math0\.0319Applied Math, Logic, Algebra, Business Management, Probability, Operations Research, Education Info5Reasoning0\.3382Philosophy/Religion, Civil Law, Politics, World History, Business Management, Economics, Constitution, Clinical Medicine6Knowledge0\.3297World History, Geography, Business Management, Physiology, Ecology, Economics, Zoology, Politics7Instruct Follow0\.8955IT Technology, Food & Cooking, Language/Writing, Literature, Movies/TV, AI, Music, Games, Business Management8Dialogue0\.7980IT Technology, Literature, Algebra, Language/Writing, Business Management, Probability, Movies/TV, AI, Physics9Reading0\.5046World History, Geography, IT Technology, Politics, Travel, Zoology, Transportation, Social Customs10Translation0\.2117Artifacts, Values, Language/Writing, Regional Characteristics, Zoology, Geography, World History, PoliticsTable 5:Tag Semantic Alignment \(TSA\) scores comparing GaoYao layers with real\-world user queries \(LMArena\)\. High TSA indicates alignment with common daily usage; low TSA indicates specialized or long\-tail capabilities\.
## Appendix DData Card for Human Annotation

Annotators were selected based on native proficiency and professional experience in translation, localization, proofreading, editing, copy\-writing, technical writing, or linguistic testing\. The annotation guidelines emphasized meaning\-preserving localization rather than literal translation, especially for instructions involving scripts, phonetics, culturally bound terms, or pragmatic conventions\. ForSuperBLEnD, annotators were instructed to answer from lived cultural experience without using AI systems or search engines, and to mark a question as "not applicable" when no clear cultural consensus existed\. All annotations followed a review\-rebuttal protocol: third\-party reviewers flagged questionable cases, after which annotators either revised the answer or provided a justification for keeping it\. Annotators were full\-time employees of the partner language service center and participated as part of their regular professional duties, receiving standard salaries above local minimum wage rather than unpaid or crowdsourced compensation\.

## Appendix ETechnical Details in the Curation ofSuperBLEnD

As discussed in Section[2\.2\.1](https://arxiv.org/html/2604.20225#S2.SS2.SSS1), existing cultural evaluation sets are limited in their culture coverage\. However, unlike the two multilingual abilities in Section[2\.2\.2](https://arxiv.org/html/2604.20225#S2.SS2.SSS2), expanding coverage of existing cultural evaluation sets presents a unique challenge: direct translation of cultural QA pairs retains the cultural perspective of the source language \(especially for the answer parts\) which fails to test true cross\-cultural generalizability, while a purely manual reconstruction is prohibitive in costs\.

To address this, we developed a three\-stage semi\-automated data expansion procedure incorporating human verification, where native speakers provide initial seeds in the first stage and inspect quality in subsequent diversifying stages to ensure both high accuracy and linguistic diversity\. The resulting evaluation set, denoted asSuperBLEnD, is a generalization of theBLEnDdataset and evaluates LLMs’ understanding of cultural differences regarding everyday concepts across 34 nations \(up from 16 inBLEnD\)\.

##### Stage 1: Generalization Q&A Seeds to More Cultures

Starting from the question templates released byBLEnD, which cover culturally shared topics ranging from festivals and food to sports, we first selected a high\-quality subset\. This filtration process excluded templates which were not universally applicable or were deemed sensitive in specific cultural contexts\. To ensure cultural authenticity and factual accuracy, answers to the selected questions were built by native members of corresponding cultures\. These annotators come from the cooperated language service center as described in Section[2\.2\.2](https://arxiv.org/html/2604.20225#S2.SS2.SSS2)\. Of the 34 cultures SuperBLEnD covers, data for 16 was inherited fromBLEnD\. For each of the 18 newly expanded cultures, three native annotators are assigned\. The instruction asks annotators to provide 1\-3 concise answers to each question based strictly on their personal life experiences within that cultural context, unassisted by AI or search engines\. To prevent forced fabrication, annotators could mark questions as “not applicable” or “no clear answer\.”

Q&A pairs across all 34 cultures underwent rigorous manual verification, discarding approximately 41\.1% of the raw data\. The removed cases mainly fall into four categories\.Semantic redundancycovers near\-duplicate answers that express the same concept, such as merging "Mum" and "Mother" into a single normalized answer\.Context errorsrefer to answers that are linguistically plausible but culturally or semantically incorrect in the target context; for example, "Wochenbett" was rejected for a German question about postpartum recovery locations because it refers to the puerperium period rather than a physical place\.Hierarchical conflictsoccur when a distractor or answer overlaps with another valid answer at a different granularity; for instance, "Pepsi" is unsuitable when "carbonated drinks" is also a valid answer\.Safety issuesinclude sensitive or potentially toxic responses that are inappropriate for a public benchmark\. The final collection contains an average of 2\.17 high\-quality answers per question template\. Note that many cultural questions accept multiple correct answers\. For example, both "beer" and "carbonated drinks" are valid, verified responses to the question: "What do young people in Malaysia usually drink at nightclubs?"

##### Stage 2: Option Synthesis

For ease of evaluation and to enhance diversity, followingMyunget al\.\([2024](https://arxiv.org/html/2604.20225#bib.bib82)\), we generalize each verified seed question into a series of multiple\-choice questions \(MCQs\)\. The volume of generated MCQs is dynamically adjusted based on the answer set size, ensuring that all verified answers are comprehensively covered as correct options while upsampling \(with different distractors\) questions with fewer answers as a balance\. For each MCQ, we synthesized options by combining the correct answer for the target country with three wrong answers as distractors derived from other countries\. In cases where insufficient real\-world distractors were available, an LLM was employed to generate plausible but incorrect "dummy options" that exist in reality but do not answer the specific question with the following prompt:

Provide \{3 \- n\} dummy option\(s\) that makes sense to be the answer\(s\) of the given question, and has to exist in real\-life \(non\-fiction\), but is totally different from the given answers without any explanation\. Make sure that the options are different from each other, and cannot be an answer from any country\. Provide as JSON format: \{"dummy\_options":\[\]\}

Synthesized MCQs underwent automated and human verification to ensure safety, answer uniqueness, and the exclusion of synonyms or hierarchical conflicts\. For instance, in the “Malaysia nightclub” scenario, “Pepsi” is an unsuitable distractor when the target answer is “beer\.” Because “Pepsi” falls under the category of “carbonated drinks” \(another acceptable cultural answer\), its inclusion introduces a hierarchical relationship that compromises the MCQ’s validity\.

##### Stage 3: Linguistic Enrichment

To further enhance linguistic diversity and reasoning difficulty, the finalized MCQs underwent a rephrasing stage\. We utilized an LLM prompted to rewrite both the question stem and options, employing techniques like paraphrasing, voice alternation, and syntactic restructuring without altering the core semantic meaning or named entities\. The prompt is as follows:

I will give you a multiple\-choice cultural question\. Your task: refine the wording of both the stem and the option as I required\. Goal: raise the overall difficulty and enrich the phrasing while keeping the underlying concepts intact\. Recommended techniques: paraphrasing, expansion, morphological variation, idioms and figurative language, voice alternation \(active↔\\leftrightarrowpassive\), syntactic restructuring, etc\. Requirements: 1\. You must not change the semantic meaning of the original stem and options\. 2\. Do not alter any entities or proper nouns \(e\.g\., personal names, company names, sport names, countries/region names, festival names\)\. 3\. Do not add a country or regional name \(or its adjective\) to the options unless the answer itself is a country or region\. 4\. Ensure the index of the original correct answer stays the same\. 5\. Use English\. 6\. Keep capitalization consistent across the options; capitalizing the first letter of each option is recommended\. 7\. Follow the specified output format exactly, including JSON punctuation\. Output format: \[…\]

This transforms straightforward MCQs into more complex versions while retaining the same cultural core, ensuring the benchmark tests cultural knowledge rather than simple pattern matching

## Appendix FDisclosure of Generative AI Usage

The technology of generative AI is partially involved in this paper for the following three scenarios: \(1\) polishing texts for language fluency, \(2\) aiding the curation of SuperBLEnD dataset as specified in Appendix[E](https://arxiv.org/html/2604.20225#A5)and \(3\) aiding in plotting art elements in figures and diagrams \(*e\.g\.*, the iceberg icon in Fig\.[1](https://arxiv.org/html/2604.20225#S1.F1)\)\.

## Appendix GDetailed Experimental Setups

Data sourceTask TypeEval\. TypeMetricJudge ModelRef\. SourceCalculation MethodS\-AlpacaEvalQASubj\.Win RateDeep Seek V3\.1Qwen3\-235B1\. Judge compares candidate vs\. reference \(correctness, richness, comprehensiveness, etc\.\); 2\.Win Rate=\#​win\+\#​tie/2\#​all\\text\{Win Rate\}=\\frac\{\\\#\\text\{win\}\+\\\#\\text\{tie\}/2\}\{\\\#\\text\{all\}\}BelebeleMCQObj\.AccuracyRule\-basedHuman \(Open Source\)1\. Reading comprehension, 4\-option regex match \(A\-D\); 2\.Accuracy=\#​correct\#​all\\text\{Accuracy\}=\\frac\{\\\#\\text\{correct\}\}\{\\\#\\text\{all\}\}INCLUDEMCQObj\.AccuracyRule\-basedHuman \(Open Source\)1\. Encyclopedic knowledge, 4\-option regex match \(A\-D\); 2\.Accuracy=\#​correct\#​all\\text\{Accuracy\}=\\frac\{\\\#\\text\{correct\}\}\{\\\#\\text\{all\}\}SuperBLEnDMCQObj\.AccuracyRule\-basedHuman and LLM Hybrid1\. Regional culture knowledge, 4\-option regex match \(A\-D\); 2\.Accuracy=\#​correct\#​all\\text\{Accuracy\}=\\frac\{\\\#\\text\{correct\}\}\{\\\#\\text\{all\}\}MGSMMathObj\.AccuracyRule\-basedHuman \(Open Source\)1\. Math reasoning, regex match for integer answers; 2\.Accuracy=\#​correct\#​all\\text\{Accuracy\}=\\frac\{\\\#\\text\{correct\}\}\{\\\#\\text\{all\}\}MMMLUMCQObj\.AccuracyRule\-basedHuman and LLM Hybrid \(Open Source\)1\. Knowledge QA, 4\-option regex match \(A\-D\)\. Uses LLM if regex fails; 2\.Accuracy=\#​correct\#​all\\text\{Accuracy\}=\\frac\{\\\#\\text\{correct\}\}\{\\\#\\text\{all\}\}Flores\-101TranslationObj\.Cometwmt22 comet \-daHuman Translated Wiki1\. Translation task; 2\. Comet ScoreSAGEMCQ\+ T/F\+ QASubj\.\+ Obj\.MixedDeep Seek V3\.1Qwen3\-max1\. MCQ and T/F uses accuracy as score; 2\. QA use LLM to recognize culture points mentioned in the answer; 3\. weighted sum score is used as final score\.CultureScopeMCQ\+ T/F\+ QASubj\.\+ Obj\.MixedDeep Seek V3\.1Human Expert1\. MCQ and T/F uses accuracy as score; 2\. QA use LLM to recognize culture points mentioned in the answer; 3\. weighted sum score is used as final score\.S\-MT\-BenchQASubj\.Win RateDeep Seek V3\.1Qwen3\-235B\-A22B1\. Judge comparison \(multi\-turn averaged\); 2\.Win Rate=\#​win\+\#​tie/2\#​all\\text\{Win Rate\}=\\frac\{\\\#\\text\{win\}\+\\\#\\text\{tie\}/2\}\{\\\#\\text\{all\}\}

Table 6:Summary of evaluation methodologies\. Task types include Multiple Choice Questions \(MCQ\), True/False \(T/F\), and Open\-ended Q&A \(QA\)\. Evaluation types distinguish between subjective \(Subj\.\) LLM\-judged approaches and objective \(Obj\.\) rule\-based approaches\. MGSM only generate Integer as final answer\.ModelModel VersionResource AddressFlagship Open\-Source ModelsopenPangu\-Ultra\-MoE\-718B\-V1\.1Ascend Tribe \([2025](https://arxiv.org/html/2604.20225#bib.bib92)\)openPangu\-Ultra\-MoE\-718B\-V1\.1[https://ai\.gitcode\.com/ascend\-tribe/openPangu\-Ultra\-MoE\-718B\-V1\.1](https://ai.gitcode.com/ascend-tribe/openPangu-Ultra-MoE-718B-V1.1)DeepSeek\-V3\.1DeepSeek \([2025](https://arxiv.org/html/2604.20225#bib.bib93)\)DeepSeek\-v3\.1\-250821[https://huggingface\.co/deepseek\-ai/DeepSeek\-V3\.1](https://huggingface.co/deepseek-ai/DeepSeek-V3.1)Qwen3\-235B\-A22BYanget al\.\([2025](https://arxiv.org/html/2604.20225#bib.bib94)\)Qwen3\-235B\-A22B\-Instruct\-2507[https://huggingface\.co/Qwen/Qwen3\-235B\-A22B\-Instruct\-2507](https://huggingface.co/Qwen/Qwen3-235B-A22B-Instruct-2507)Qwen3\-VL\-235B\-A22BBaiet al\.\([2025](https://arxiv.org/html/2604.20225#bib.bib95)\)Qwen3\-VL\-235B\-A22B\-Instruct[https://huggingface\.co/Qwen/Qwen3\-VL\-235B\-A22B\-Instruct](https://huggingface.co/Qwen/Qwen3-VL-235B-A22B-Instruct)DeepSeek\-R1DeepSeek\-AI \([2025](https://arxiv.org/html/2604.20225#bib.bib96)\)DeepSeek\-R1[https://huggingface\.co/deepseek\-ai/DeepSeek\-R1](https://huggingface.co/deepseek-ai/DeepSeek-R1)Llama\-3\.1\-405BMeta AI \([2024](https://arxiv.org/html/2604.20225#bib.bib103)\)Llama\-3\.1\-405B\-Instruct[https://huggingface\.co/meta\-llama/Llama\-3\.1\-405B\-Instruct](https://huggingface.co/meta-llama/Llama-3.1-405B-Instruct)GLM\-4\.6Zenget al\.\([2025](https://arxiv.org/html/2604.20225#bib.bib99)\)GLM\-4\.6[https://chatglm\.cn](https://chatglm.cn/)Kimi\-k2Teamet al\.\([2025b](https://arxiv.org/html/2604.20225#bib.bib100)\)kimi\-k2\-250711[https://www\.kimi\.com](https://www.kimi.com/)Closed\-Source Commercial ModelsDoubao\-seed\-1\.6ByteDance \([2026](https://arxiv.org/html/2604.20225#bib.bib97)\)doubao\-seed\-1\-6\-250615[https://www\.doubao\.com/chat](https://www.doubao.com/chat)Qwen\-maxTeam \([2025](https://arxiv.org/html/2604.20225#bib.bib98)\)Qwen\-max[https://chat\.qwen\.ai/](https://chat.qwen.ai/)Gemini\-2\.5\-ProComaniciet al\.\([2025](https://arxiv.org/html/2604.20225#bib.bib89)\)Gemini\-2\.5\-Pro[https://aistudio\.google\.com](https://aistudio.google.com/)Claude\-Sonnet\-4\.5Anthropic \([2026](https://arxiv.org/html/2604.20225#bib.bib107)\)claude\-sonnet\-4\-5\-20250929[https://claude\.ai](https://claude.ai/)Grok\-3xAI \([2025](https://arxiv.org/html/2604.20225#bib.bib109)\)Grok\-3[https://grok\.com](https://grok.com/)o3OpenAI \([2025b](https://arxiv.org/html/2604.20225#bib.bib110)\)o3[https://chatgpt\.com](https://chatgpt.com/)o4\-miniOpenAI \([2025b](https://arxiv.org/html/2604.20225#bib.bib110)\)o4\-mini[https://chatgpt\.com](https://chatgpt.com/)GPT\-5\-chatOpenAI \([2026](https://arxiv.org/html/2604.20225#bib.bib115)\)GPT\-5\-chat[https://chatgpt\.com/](https://chatgpt.com/)GPT\-4oOpenAIet al\.\([2024](https://arxiv.org/html/2604.20225#bib.bib116)\)GPT\-4o[https://chatgpt\.com](https://chatgpt.com/)Compact Models \(<20B\)Qwen3\-14BYanget al\.\([2025](https://arxiv.org/html/2604.20225#bib.bib94)\)Qwen3\-14B[https://huggingface\.co/Qwen/Qwen3\-14B](https://huggingface.co/Qwen/Qwen3-14B)Qwen3\-8BYanget al\.\([2025](https://arxiv.org/html/2604.20225#bib.bib94)\)Qwen3\-8B[https://huggingface\.co/Qwen/Qwen3\-8B](https://huggingface.co/Qwen/Qwen3-8B)Gemma\-3\-12B\-ITTeamet al\.\([2025a](https://arxiv.org/html/2604.20225#bib.bib101)\)gemma\-3\-12b\-it[https://huggingface\.co/google/gemma\-3\-12b\-it](https://huggingface.co/google/gemma-3-12b-it)Llama\-3\.1\-8BMeta AI \([2024](https://arxiv.org/html/2604.20225#bib.bib103)\)Llama\-3\.1\-8B\-Instruct[https://huggingface\.co/meta\-llama/Llama\-3\.1\-8B\-Instruct](https://huggingface.co/meta-llama/Llama-3.1-8B-Instruct)Llama\-3\-8BMeta AI \([2024](https://arxiv.org/html/2604.20225#bib.bib103)\)Llama\-3\-8B\-Instruct[https://huggingface\.co/meta\-llama/Meta\-Llama\-3\-8B\-Instruct](https://huggingface.co/meta-llama/Meta-Llama-3-8B-Instruct)Ministral\-8B\-InstructJianget al\.\([2024](https://arxiv.org/html/2604.20225#bib.bib102)\)Ministral\-8B\-Instruct\-2410[https://huggingface\.co/mistralai/Ministral\-8B\-Instruct\-2410](https://huggingface.co/mistralai/Ministral-8B-Instruct-2410)Table 7:Detailed specifications of models evaluated in GaoYao\.

Similar Articles

OmnilingualGAIA2: Evaluating the Multilingual Gap in Frontier AI Agents

arXiv cs.CL

This paper introduces OmnilingualGAIA2, a multilingual expansion of the GAIA2 agentic benchmark across ten languages, revealing a universal cross-lingual performance gap of 8.8–18.4 pass@3 points that is model-driven and persists with scale. The authors argue that multilingual agentic evaluation should become standard for globally deployed agents.