OmnilingualGAIA2: Evaluating the Multilingual Gap in Frontier AI Agents

arXiv cs.CL Papers

Summary

This paper introduces OmnilingualGAIA2, a multilingual expansion of the GAIA2 agentic benchmark across ten languages, revealing a universal cross-lingual performance gap of 8.8–18.4 pass@3 points that is model-driven and persists with scale. The authors argue that multilingual agentic evaluation should become standard for globally deployed agents.

arXiv:2608.08775v1 Announce Type: new Abstract: Agentic benchmarks aim to measure how well AI agents plan, search, execute, and recover within realistic multi-tool environments, but they are almost exclusively in English. As AI agents are globally deployed to a linguistically diverse user base, whether agentic competence measured in English transfers to other languages remains an open question. We introduce OmnilingualGAIA2, a machine-translated expansion (with partial human- expert validation) of the GAIA2 agentic benchmark, covering ten target languages spanning five writing systems, paired with a localised and human-calibrated multilingual verifier. Evaluating seven frontier and open-weight agents, we find a universal cross-lingual gap of 8.8-18.4 pass@3 points that is agent-asymmetric in magnitude, concentrates on tool-orchestration rather than quantitative reasoning, and does not close with model scale. A stratified error attribution decomposes the gap as predominantly model-driven (55%), with a bounded translation-contamination floor of only 6.4% of scenario-language pairs. Human-expert linguistic analysis further identifies morphological cue loss and amplified ambiguity as the primary failure mechanisms in non-Latin-script languages. Our results argue that multilingual agentic evaluation must become a standard part of the reporting protocol for globally deployed agents.
Original Article
View Cached Full Text

Cached at: 08/11/26, 08:09 AM

# OmnilingualGAIA2: Evaluating the Multilingual Gap in Frontier AI Agents
Source: [https://arxiv.org/html/2608.08775](https://arxiv.org/html/2608.08775)
\]Meta Superintelligence Labs

Pere\-Lluís Huguet CabotChierh ChengAlbert Ventayol\-BoadaGabriel Mejia GonzalezChristophe RopersLucas BandarkarSebastian RuderDarlene SakakiharaElliot YunPierre AndrewsGrégoire MialonRomain FrogerMarta R\. Costa\-jussà\[[costajussa](https://arxiv.org/html/2608.08775v1/mailto:costajussa)

\(August 9, 2026\)

###### Abstract

Agentic benchmarks aim to measure how well AI agents plan, search, execute, and recover within realistic multi\-tool environments, but they are almost exclusively in English\. As AI agents are globally deployed to a linguistically diverse user base, whether agentic competence measured in English transfers to other languages remains an open question\. We introduceOmnilingualGAIA2\\xspace, a machine\-translated expansion \(with partial human\-expert validation\) of theGAIA2\\xspaceagentic benchmark, covering ten target languages spanning five writing systems, paired with a localised and human\-calibrated multilingual verifier\. Evaluating seven frontier and open\-weight agents, we find a universal cross\-lingual gap of 8\.8–18\.4pass​@​3\\text\{pass\}@3points that is agent\-asymmetric in magnitude, concentrates on tool\-orchestration rather than quantitative reasoning, and does not close with model scale\. A stratified error attribution decomposes the gap as predominantly model\-driven \(55%\), with a bounded translation\-contamination floor of only 6\.4% of scenario–language pairs\. Human\-expert linguistic analysis further identifies morphological cue loss and amplified ambiguity as the primary failure mechanisms in non\-Latin\-script languages\. Our results argue that multilingual agentic evaluation must become a standard part of the reporting protocol for globally deployed agents\.

## 1Introduction

AI agents—systems that plan, invoke tools, and act autonomously within an environment to pursue a user’s goals\(Wooldridge and Jennings,[1995](https://arxiv.org/html/2608.08775#bib.bib56); Franklin and Graesser,[1996](https://arxiv.org/html/2608.08775#bib.bib10); Russell and Norvig,[2009](https://arxiv.org/html/2608.08775#bib.bib46)\)—have moved from objects of research curiosity to real products deployed and used by hundreds of millions of people worldwide, thanks to the rapid advancements of Large Language Model \(LLM\) capabilities\(Wang et al\.,[2024](https://arxiv.org/html/2608.08775#bib.bib53); Xi et al\.,[2025](https://arxiv.org/html/2608.08775#bib.bib57); Yang et al\.,[2025](https://arxiv.org/html/2608.08775#bib.bib62); Staufer et al\.,[2026](https://arxiv.org/html/2608.08775#bib.bib51)\)\. Those users do not all speak English, yet our ability to measure whether these agents actually work has barely kept pace with their deployment, and least of all beyond English\. Even in English the ground is unsteady: developers who anticipated a24%24\\%speedup from AI tools were instead slowed by19%19\\%\(Lobentanzer,[2026](https://arxiv.org/html/2608.08775#bib.bib29)\), and observed agent task coverage remains a fraction of what has been hypothesised\(Massenkoff and McCrory,[2026](https://arxiv.org/html/2608.08775#bib.bib31); Johnston et al\.,[2026](https://arxiv.org/html/2608.08775#bib.bib17)\)\. Whether the competence we certify in English survives once an agent must plan, search, and act in another language is, at the time of writing, largely unmeasured\.

In English, this gap is being rapidly closed, with considerable work invested in agentic benchmarks that probe how well these systems plan, search, and use tools\(Mialon et al\.,[2024](https://arxiv.org/html/2608.08775#bib.bib35); Qin et al\.,[2024](https://arxiv.org/html/2608.08775#bib.bib43); Xie et al\.,[2024](https://arxiv.org/html/2608.08775#bib.bib59); Patil et al\.,[2025](https://arxiv.org/html/2608.08775#bib.bib42); Deng et al\.,[2026](https://arxiv.org/html/2608.08775#bib.bib8); Barres et al\.,[2025](https://arxiv.org/html/2608.08775#bib.bib6); Merrill et al\.,[2026](https://arxiv.org/html/2608.08775#bib.bib33)\)\. Beyond English, evaluation has barely followed: for frontier agentic systems111See for instance[System Card: Claude Opus 4\.7](https://www.anthropic.com/claude-opus-4-7-system-card), multilingual assessment remains largely confined to single\-turn, knowledge\-based question answering datasets such asGlobal\-MMLU\(Singh et al\.,[2025](https://arxiv.org/html/2608.08775#bib.bib49)\), when it is performed at all222No multilingual performance reported in[GPT\-5\.4 Thinking System Card](https://openai.com/index/gpt-5-4-thinking-system-card/)\.

To help close this gap, we presentOmnilingualGAIA2, a high\-quality machine\-translated expansion ofGAIA2\(Froger et al\.,[2026](https://arxiv.org/html/2608.08775#bib.bib11)\), paired with a multilingual verifier and a frontier multilingual agentic leaderboard\. Our contributions are:

1. 1\.A multilingual agentic benchmark\.We releaseOmnilingualGAIA2, a machine\-translated \(MT\) expansion ofGAIA2that covers four capabilities, namely adaptability, ambiguity, execution, and search, across ten target languages spanning five writing systems while preserving the executable structure required by the verifier \(§[3](https://arxiv.org/html/2608.08775#S3)\)\. We select the translation system through a per\-language quality evaluation and design a MT pipeline that enforces cross\-surface terminology consistency across scenarios \(§[3\.2](https://arxiv.org/html/2608.08775#S3.SS2)\)\. See our language\-coverage contribution relative to existing datasets in Appendix[A](https://arxiv.org/html/2608.08775#A1)\.
2. 2\.A human\-calibrated multilingual verifier\.We localise the LLM\-as\-a\-Judge \(LLMaaJ\) component of theGAIA2verifier, and ensure its multilingual quality via calibration on a golden set of human\-annotated and MT agent traces \(§[3\.3](https://arxiv.org/html/2608.08775#S3.SS3)\)\.
3. 3\.A cross\-lingual multi\-agent leaderboard and empirical characterisation of the gap\. We evaluate a heterogeneous cohort of seven AI agents spanning frontier closed\-source systems and open\-weights models of various sizes and architectures, and find that: \(i\) the cross\-lingual gap is universal in direction but agent\-asymmetric in magnitude \(8\.8–18\.4 percentage points, pp\); \(ii\) the capability ordering \(search \> execution \> adaptability \> ambiguity\) is stable across languages \- translation reshapes the level without reordering difficulty; \(iii\) agents shift their behavioural strategy off English, increasing exploratory reads while reducing state\-changing writes; \(iv\) the cross\-lingual gap is concentrated in tool\-orchestration and final\-response quality, not in quantitative or categorical reasoning; and \(v\) the gap does not close with model scale within a single family \(§[4](https://arxiv.org/html/2608.08775#S4), §[5](https://arxiv.org/html/2608.08775#S5)\)\.
4. 4\.An error taxonomy and gap attribution showing the gap is predominantly model\-driven\. We define a five\-category verdict taxonomy and apply a stratified automatic triage protocol that corrects for the selection bias of deterministic\-only analysis\. The resulting decomposition attributes 55% of the cross\-lingual gap to genuine agent failures, 35% to translation defects, and 10% to verifier artefacts \(§[6\.1](https://arxiv.org/html/2608.08775#S6.SS1),§[6\.2](https://arxiv.org/html/2608.08775#S6.SS2)\)\. Extending the analysis to the whole benchmark, we bound MT\-induced contamination at only 6\.4% of all scenario–language pairs, establishing that the benchmark predominantly measures model capability rather than translation noise\.
5. 5\.A human\-expert linguistic analysis identifying failure mechanisms per error category\. Three native\-speaker linguists independently inspect agent traces across the four capabilities and surface qualitative patterns: \(i\) translation defects — MT amplifies latent English ambiguities and silently drops directional/morphological constraints that disambiguate the correct answer; \(ii\) verifier artefacts — verifier strictness is language\-dependent, stochastic, and can hallucinate rejections of correct content; \(iii\) agent failures — morphological cue loss \(articles, plural marking\) in cmn/jpn/ind drives the most consequential over\-actions; and \(iv\) measurement artefacts — action\-count grading penalises harmless self\-corrections that reach the correct end state, inflating the apparent failure rate \(§[6\.3](https://arxiv.org/html/2608.08775#S6.SS3)\)\.

OmniGAIA2 is part of Meta’s broader Omnilingual effort to extend AI capability beyond the handful of high\-resource languages that dominate current systems, alongside Omnilingual ASR\(Omnilingual\-ASR\-team et al\.,[2025](https://arxiv.org/html/2608.08775#bib.bib37)\)for speech recognition across 1,600\+ languages and Omnilingual MT\(Omnilingual\-MT\-Team et al\.,[2026](https://arxiv.org/html/2608.08775#bib.bib38)\)for translation at comparable scale\. Where those efforts expand the capability frontier, and BOUQuET\(Andrews et al\.,[2025](https://arxiv.org/html/2608.08775#bib.bib1)\)the translation\-evaluation frontier, OmniGAIA2 extends evaluation to interactive, tool\-using agentic tasks beyond English\.

## 2Related Work

#### Multilingual agentic evaluation\.

Efforts to move agentic evaluation beyond English have so far been narrow along at least one axis: the set of languages, the class of agents evaluated, or the fidelity of the verifier under translation\. In the web\-agent setting, X\-WebAgentBench\(Wang et al\.,[2025](https://arxiv.org/html/2608.08775#bib.bib54)\)translates a subset of navigation tasks into several languages, while Ticket\-Bench\(Sales Almeida et al\.,[2025](https://arxiv.org/html/2608.08775#bib.bib47)\)regionalises a customer\-service function\-calling benchmark across six European locales\. On the tool\-calling sub\-capability alone, MLCL\(Luo et al\.,[2026](https://arxiv.org/html/2608.08775#bib.bib30)\)isolates cross\-lingual robustness of API selection and argument copying, decomposing the multilingual gap into query\-comprehension and parameter\-value\-mismatch components\. MASSIVE\-Agents\(Kulkarni et al\.,[2025](https://arxiv.org/html/2608.08775#bib.bib25)\), which covers 52 languages, was created by cleaning the original MASSIVE dataset and then reformatting it for evaluation within the Berkeley Function\-Calling Leaderboard \(BFCL\) framework\. MAPS\(Hofman et al\.,[2026](https://arxiv.org/html/2608.08775#bib.bib16)\)pairs a ten\-language agent suite with a security\-focused analysis, and TelcoAgent\-Bench\(Bariah et al\.,[2026](https://arxiv.org/html/2608.08775#bib.bib5)\)evaluates multilingual telecom agents in English and Arabic, finding that performance gaps widen in unconstrained bilingual settings—a pattern consistent with our own findings across ten languages\. Closest to our setting, three recent efforts translate structurally identical benchmarks across languages while preserving their executable verifiers: SEATauBench\(Nguyen et al\.,[2026](https://arxiv.org/html/2608.08775#bib.bib36)\)adaptsτ2\\tau^\{2\}\-Bench into five Southeast Asian languages, PolyWorkBench\(Li et al\.,[2026](https://arxiv.org/html/2608.08775#bib.bib28)\)constructs a native multilingual long\-horizon agent benchmark from scratch, and our most direct sibling, GAIA\-v2\-LILT\(Kim et al\.,[2026](https://arxiv.org/html/2608.08775#bib.bib20)\), localisesGAIAacross five target languages through a translate\-then\-adapt pipeline that modifies task semantics to fit the locale\.OmnilingualGAIA2instead expands by machine translation: unlike synthetic\-generation approaches \(IntellAgent\(Levi and Kadar,[2025](https://arxiv.org/html/2608.08775#bib.bib27)\), TaskCraft\(Shi et al\.,[2026](https://arxiv.org/html/2608.08775#bib.bib48)\), AgentSynth\(Xie et al\.,[2026](https://arxiv.org/html/2608.08775#bib.bib58)\)\), MT\-based expansion preserves an existing verifier’s executable structure while scaling linguistically, following the hybrid MT\-then\-expert\-review pipeline of MMLU\-ProX\(Xuan et al\.,[2025](https://arxiv.org/html/2608.08775#bib.bib60)\)at 29 languages\. Benchmark validity grounds our analysis: the Agentic Benchmark Checklist \(ABC\)\(Zhu et al\.,[2026](https://arxiv.org/html/2608.08775#bib.bib65)\)documents systematic outcome\- and task\-validity failures in widely used agentic benchmarks; we address these by calibrating the verifier against human labels \(§[3\.3](https://arxiv.org/html/2608.08775#S3.SS3)\) and by decomposing the observed gap into translation\-, model\-, and verifier\-side factors with a bounded contamination floor \(§[6\.3](https://arxiv.org/html/2608.08775#S6.SS3)\)\. Against this landscape,OmnilingualGAIA2is, to our knowledge, the first extension ofGAIA2that\(i\)covers ten target languages spanning five writing systems \(Appendix[A](https://arxiv.org/html/2608.08775#A1)\) while preserving the executable verifier end\-to\-end,\(ii\)reports a paired comparison across a heterogeneous cohort of frontier and open\-weights agents under a shared, calibrated cross\-lingual judge, and\(iii\)decomposes the observed gap into translation\-side and model\-side factors, extending the analytic programme of MLCL from isolated tool calls to full agent trajectories—complemented by a comprehensive human\-expert linguistic validation\.

#### Multilingual LLM\-as\-a\-Judge\.

BecauseGAIA2, and by inheritanceOmnilingualGAIA2, relies on a model\-based verifier to grade the natural\-language write actions of agents, the cross\-lingual reliability of the judge itself is a first\-class concern\. LLM judging was popularised byZheng et al\. \([2023](https://arxiv.org/html/2608.08775#bib.bib64)\)and inherits from prompt\-based translation\-quality scoring in the MT community\(Kocmi and Federmann,[2023](https://arxiv.org/html/2608.08775#bib.bib22); Rei et al\.,[2022](https://arxiv.org/html/2608.08775#bib.bib45); Juraska et al\.,[2025](https://arxiv.org/html/2608.08775#bib.bib18)\); subsequent audits catalogue position, verbosity, and self\-preference biases\(Panickssery et al\.,[2024](https://arxiv.org/html/2608.08775#bib.bib41); Koo et al\.,[2024](https://arxiv.org/html/2608.08775#bib.bib23)\)that are already problematic in the monolingual setting\. In the multilingual setting these effects compound: METAL\(Hada et al\.,[2024](https://arxiv.org/html/2608.08775#bib.bib15)\)and MM\-Eval\(Son et al\.,[2024](https://arxiv.org/html/2608.08775#bib.bib50)\)document sharp drops in judge\-human agreement on lower\-resource languages, M\-RewardBench\(Gureja et al\.,[2025](https://arxiv.org/html/2608.08775#bib.bib14)\)reports the analogous collapse for multilingual reward models, BabelJudge\(KC,[2026](https://arxiv.org/html/2608.08775#bib.bib19)\)extends the analysis to full agent trajectories and finds order\-consistency near chance on some languages, the Coin\-Flip\-Judge study\(Yagubyan,[2026](https://arxiv.org/html/2608.08775#bib.bib61)\)quantifies reliability floors under adversarial pairings, andZhang et al\. \([2026](https://arxiv.org/html/2608.08775#bib.bib63)\)isolate a distinct*translationese*bias whereby judges systematically prefer back\-translated over human\-authored responses\.Doğruöz et al\. \([2026](https://arxiv.org/html/2608.08775#bib.bib9)\)synthesise these findings into concrete recommendations for multilingual judging practice\. Our verifier design \(§[3\.3](https://arxiv.org/html/2608.08775#S3.SS3)\) inherits theGAIA2LLMaaJ scaffold but explicitly recalibrates it against human\-annotated multilingual agent traces, quantifying inter\-rater agreement with standard nominal\-scale statistics\(Cohen,[1960](https://arxiv.org/html/2608.08775#bib.bib7); Artstein and Poesio,[2008](https://arxiv.org/html/2608.08775#bib.bib3); Landis and Koch,[1977](https://arxiv.org/html/2608.08775#bib.bib26); Krippendorff,[2019](https://arxiv.org/html/2608.08775#bib.bib24)\)and reporting language\-conditional confidence intervals on the resulting leaderboard\.

## 3OmnilingualGAIA2

### 3\.1Source benchmark

We build onGAIA2\(Froger et al\.,[2026](https://arxiv.org/html/2608.08775#bib.bib11)\), an agentic benchmark in which an agent operates within a simulated environment equipped with tools\. In particular, we build on the native*Mobile*environment, comprised of*apps*such as*Email*and*Calendar*, as well as of a set of corresponding collections of*app states*, aka*universe*\. Each universe identifies a synthetic*persona*and a snapshot of their digital life, with the app states containing synthetic user data such as emails or calendar events\. Each scenario then specifies a user task in this environment and a minimal \(logically and/or temporally ordered\) set of actions that the agent must have undertaken, including the final answer\. The actual actions \(*events*\) undertaken by the agent are then compared to these*oracle events*, and in particular the ones containing free\-form text \(e\.g\. the final answer\) are subject to LLMaaJ verification, to declare the scenario a success or a failure\.

The fullGAIA2capability taxonomy comprises seven splits: Execution, Search, Ambiguity, Adaptability, Time, Noise and Agent2Agent\(Froger et al\.,[2026](https://arxiv.org/html/2608.08775#bib.bib11)\)\. We initially designOmnilingualGAIA2to cover four capabilities, namelyAdaptability,Ambiguity,ExecutionandSearch, while reserving the extension to the other capabilities for future work, as those are the ones, at the time of writing, farthest away from saturation by frontier agentic systems even in English, and we posit multilingual performance to lag behind even more\.

### 3\.2Translating the benchmark

#### Translation pipeline

Translating a*verifiable*agentic task is materially harder than translating a static QA pair: the translated environment must remain internally consistent so that the verifier still admits exactly the intended solution\. To constructOmnilingualGAIA2we translate into each target language the language\-dependent surface of every scenario: the user task description, the natural\-language portion of the universes, as well as of the oracle events \(*text spans*\)\. The executable structure\+\+ of the scenario, such as identifiers and timestamps, is left untouched \(*frozen spans*\)\.

The pipeline proceeds in three stages, illustrated in Figure[1](https://arxiv.org/html/2608.08775#S3.F1)\. SinceGAIA2defines multiple scenarios over the same universe, the pipeline first aggregates the universes and translates their text spans in calls batched per app\. From these translated surfaces a reference*term table*is built—primarily derived from the already\-translated universes and augmented by a per\-scenario extraction pass—by identifying and consolidating the entities that recur across surfaces; this table enforces cross\-surface translation consistency, which is crucial to avoid constructing unsolvable tasks that mention non\-existing entities\. Task descriptions and oracle\-event text spans are then translated into the target language with the term table constraining their rendering, and a final validation pass sweeps the assembled scenario and substitutes any residual source\-language span against the table, so that the translated environment stays consistent with what the verifier expects\.

![Refer to caption](https://arxiv.org/html/2608.08775v1/x1.png)Figure 1:TheOmnilingualGAIA2translation pipeline\.EachGAIA2scenario is translated in three stages\.\(a\) Universe TranslationThe language\-dependent*text spans*of every surface—the user prompt, the text span of the app\-states of the universes, and the \(oracle\) events—are rendered into the target language by the LLM translator, in calls batched per app\.\(b\) Building the reference term tableA second LLM pass identifies the entities that recur across surfaces and consolidates them into a reference*term table*: a reference rendering of each shared entity\.\(c\) Translating prompts and eventsThe term table is applied back across the assembled scenario, substituting any residual or divergent source\-language span so that every surface refers to each entity identically—the cross\-surface consistency the verifier requires to admit exactly the intended solution\.
#### Translation system\.

We rely on BOUQuET333[facebook/bouquet](https://huggingface.co/spaces/facebook/bouquet)\(Andrews et al\.,[2025](https://arxiv.org/html/2608.08775#bib.bib1)\)to selectGemma\-4\-31B\-Instruct\(Team et al\.,[2025](https://arxiv.org/html/2608.08775#bib.bib52)\)as our translator system, as the top\-ranking open\-license translator system across all the languages we target\. We measure translation quality by pairwise LLM\-as\-judge comparison of each candidate rendering against a reference translation, adjudicated by two independent judges under position randomisation and majority voting; a per\-language head\-to\-head against the other shortlisted systems on this measure confirms this choice\. Furthermore, we investigate the effectiveness of a review and post\-editing pass with a different model, which proves a near\-no\-op: it rewrites fewer than9%9\\%of fields, with edits that are almost entirely stylistic, with no measurable impact on translation quality while roughly doubling per\-scenario latency\. See Appendix[B](https://arxiv.org/html/2608.08775#A2)for additional details\.

### 3\.3Multilingual Verifier

GAIA2evaluates every state\-changing write action that the agent performs in the environment against oracle annotations via a multi\-step*verifier*employing an LLMaaJ component for text spans\. Therefore, to properly extend the benchmark into a multilingual setting, besides the user tasks, universe app states, and oracle events, the LLMaaJ component too needs to be localised, to ensure a fair judgement of non\-English scenarios\. Here, we use the term*judge*to refer to the LLMaaJ component of the verifier, and describe the work done to localise the judge, and run calibrations to ensure parity to English and correctness on target languages\. We summarize the protocol and findings below, and refer the reader to Appendix[C](https://arxiv.org/html/2608.08775#A3)for additional details\.

#### Localising the judge\.

The \(multilingual\) judge quality can be improved along two axes: thejudge promptand thejudge model\. To evaluate both independently of agent performance, we use a set of human\-annotated scenario traces fromGAIA2, each carrying a ground\-truth pass/fail label\. The same set is translated with our pipeline \(§[3\.2](https://arxiv.org/html/2608.08775#S3.SS2)\) to produce a cross\-lingual counterpart, enabling evaluation under both source and target languages on identical trace scaffolding\.

Starting the investigation with the judge prompt, we find English\-specific assumptions that make these prompts unsound on translated content; e\.g\. the content checker is primed exclusively with English few\-shot examples, so a faithful non\-English write action can be judged non\-compliant for reasons unrelated to its correctness\.OmnilingualGAIA2therefore ships a set of localised judge prompts that remove this priming and generalise the natural\-language sub\-checks across target languages; the calibration below confirms that this localisation leaves English agreement intact while carrying over to the target languages we evaluate\. On the model axis, we require a judge that is open\-weight and light\-weight enough to score the entire multilingual leaderboard\.GPT\-OSS\-120B444the configuration the currentGAIA2reference implementation itself adopts, see[facebookresearch/meta\-agents\-research\-environment](https://github.com/facebookresearch/meta-agents-research-environments), meets both: as a mixture\-of\-experts model it activates only5\.15\.1B parameters per token\(OpenAI,[2025](https://arxiv.org/html/2608.08775#bib.bib40)\), roughly an order of magnitude fewer than the denseLlama\-3\.3\-70B\-Instructoriginal paper’s reference judge\(Meta AI,[2024](https://arxiv.org/html/2608.08775#bib.bib34)\), hence far fewer FLOPs per call\.

#### Calibrating the judge\.

We compare, on the same human\-annotated traces, the four combinations of the two judge models with the two prompt sets, scoring each against the human\-majority label with Cohen’sκ\\kappa, for chance\-corrected agreement\.

On English, no configuration reproduces the human verdicts measurably better than another: the four agree with the human labels comparably \(full\-corpusκ\\kappawithin0\.710\.71–0\.730\.73, rising toκ≥0\.93\\kappa\\geq 0\.93on the*LLM\-touched*slice the judge actually adjudicates\)\. Crucially, the proposed configuration \(GPT\-OSS\-120B, localised\) agrees with the original paper’s reference judge on almost all traces, withκ=0\.981\\kappa=0\.981\. This agreement is preserved under translation across all ten target languages, on the LLM\-touched slice Cohen’sκ\\kapparanges from0\.8860\.886\(tur\) to0\.9710\.971\(eng,cmn\)\.

## 4Experimental Setup

#### Agents

We evaluate seven frontier and open\-weight models onOmnilingualGAIA2, selected to span both the closed\-source vs open\-weights divide and, within open\-weights, across dense and mixture\-of\-experts architectures at varied scale\. Among proprietary systems, we includeClaude\-4\.7\-Opus\(Anthropic,[2025](https://arxiv.org/html/2608.08775#bib.bib2)\),GPT\-5\.4\(OpenAI,[2025](https://arxiv.org/html/2608.08775#bib.bib39)\),Gemini\-3\.1\-Pro\(Google DeepMind,[2025a](https://arxiv.org/html/2608.08775#bib.bib12)\)\. For open\-weight models, we evaluateGemma\-4\-31B\-Instruct\(Google DeepMind,[2025b](https://arxiv.org/html/2608.08775#bib.bib13)\),Kimi\-2\.6\(Kimi\-Team,[2025](https://arxiv.org/html/2608.08775#bib.bib21)\)and the Qwen 3\.6 family in both its mixture\-of\-experts \(Qwen\-3\.6\-35B\-A3B\) and dense \(Qwen\-3\.6\-27B\) configurations\(Qwen Team,[2025](https://arxiv.org/html/2608.08775#bib.bib44)\)\. We also conduct a scale ablation on the latter family\.

All agents are driven by theOpenClaw555[https://github\.com/openclaw/openclaw](https://github.com/openclaw/openclaw)harness, with the exception ofKimi\-2\.6, which is driven withOpenCode666[https://opencode\.ai](https://opencode.ai/)instead777Kimi\-2\.6scores near zero withOpenClaw\. We suspect this is because it drops its native streamingreasoning\_contentandtool\_calls\.\. For open\-weight models, we use a131131K context window, default generation parameters and enable thinking mode, setting its effort to*low*when possible\. For proprietary models, we use similar settings\.

#### Context handling\.

We serve the open\-weight agents with a context window large enough that no rollout is scored as a failure merely because its trajectory was truncated: the small fraction of rollouts that exceed the default budget—almost entirely longhin,jpnandcmntrajectories—are re\-run with a larger window and scored on that run, keeping the intended thinking\-on configuration intact\.

#### Metrics\.

We reportpass​@​3\\text\{pass\}@3as our main metric: a scenario counts as solved if any of three independent runs passes the verifier, for direct comparability withGAIA2\(Froger et al\.,[2026](https://arxiv.org/html/2608.08775#bib.bib11)\)\.

## 5Results

### 5\.1Leaderboard

Table[1](https://arxiv.org/html/2608.08775#S5.T1)reports pooledpass​@​3\\text\{pass\}@3numbers over the four capabilities \(adaptability, ambiguity, execution, search\)\.

Table 1:Per\-capabilitypass​@​3\\text\{pass\}@3\(%\) across agents; best per column inbold\. The target\-language average pools the ten non\-English languages\.The results demonstrate a clear and evident cross\-lingual gap between the English baseline and any target\-language across all agents\. This gap varies in magnitude across models, withGemini\-3\.1\-Proat8\.88\.8pass​@​3\\text\{pass\}@3points andClaude\-4\.7\-Opusat9\.19\.1\. On the other end of the range,Qwen\-3\.6\-35B\-A3Bshows a pooled gap of18\.418\.4\. The ordering amongst the seven agents is consistent withpass​@​1\\text\{pass\}@1, withGemini\-3\.1\-Prostill showing the smallest pooled gap at7\.77\.7\. However,Gemini\-3\.1\-Prohas the lowest English and multilingual performance of the three frontier agents \(62\.262\.2pass​@​3\\text\{pass\}@3\)\. Within the open\-weights classKimi\-2\.6outperformsGemini\-3\.1\-Proin English but lags behind on target languages\.

Looking at thepass​@​1\\text\{pass\}@1,pass​@​3\\text\{pass\}@3andpass3all\\text\{pass\}^\{\\text\{all\}\}\_\{3\}numbers \(reported in Appendix[D](https://arxiv.org/html/2608.08775#A4)\) we observe how for most agents extra attempts do not recover the loss and the same target\-language scenarios fail every time\.Claude\-4\.7\-Opusis the exception: its gap narrows from14\.014\.0atpass​@​1\\text\{pass\}@1to9\.19\.1atpass​@​3\\text\{pass\}@3yet is widest atpass3all\\text\{pass\}^\{\\text\{all\}\}\_\{3\}\(17\.517\.5\), so off English it still reaches the solution within three attempts, only far less consistently\. That target\-languagepass​@​3\\text\{pass\}@3stays high \(e\.g\.∼74%\\sim 74\\%forClaude\-4\.7\-Opus,58%58\\%forGPT\-5\.4\) shows these scenarios remain largely solvable off English rather than the benchmark saturating; how much of the residual difficulty reflects genuine model limitations versus translation artefacts is what we disentangle in the attribution analysis later\.

SinceGemma\-4\-31B\-Instructserves as both the model in the translation pipeline \(§[3\.2](https://arxiv.org/html/2608.08775#S3.SS2)\) and one of the evaluated agents, we verify that this leaderboard is not distorted by a*family\-match*advantage, i\.e\. an agent scoring higher on data translated by its own model family\. We run a Gemma\-vs\-Qwen agent\-translator ablation over all ten target languages, and find the agent gap invariant to which family translated the evaluation data, with both agents losing comparably on the weaker Qwen translations; full design and per\-language results are in Appendix[B\.3](https://arxiv.org/html/2608.08775#A2.SS3)\.

### 5\.2Per\-capability breakdown

We now break down the numbers of Table[1](https://arxiv.org/html/2608.08775#S5.T1)into the four capabilities\. For each \(capability, language, agent\) combination we reportpass​@​3\\text\{pass\}@3with95%95\\%CIs and we test the English\-vs\-target\-language difference with a paired McNemar test\(McNemar,[1947](https://arxiv.org/html/2608.08775#bib.bib32)\)on the matched scenario pairs, per capability\. Figure[2](https://arxiv.org/html/2608.08775#S5.F2)gives the full capability\-by\-language view at a glance, while the exact numbers, Wilson CIs and per\-cell McNemarpp\-values are in Appendix[D](https://arxiv.org/html/2608.08775#A4)\. Two patterns are visible: performance is strongest in English for nearly every \(agent, capability\) combination, and the drop concentrates on the CJK and Indic languages \(cmn,jpn,hin\)\. These three languages notably use non\-Latin scripts, and we find that part of the reason is the agents not carefully considering the various formsentitiesmay take when calling tools \(see Appendix[G](https://arxiv.org/html/2608.08775#A7)for details\), mirroring cross\-script mistakes in non\-agentic tasks\(Bandarkar et al\.,[2026](https://arxiv.org/html/2608.08775#bib.bib4)\)\. The frontier agents \(Claude\-4\.7\-Opus,GPT\-5\.4\) show the most cross\-lingually robust behaviour, while the open\-weights agents perform significantly below English on almost every target language\.

Furthermore, the capability ordering appears stable across languages, withpass​@​3\\text\{pass\}@3numbers decreasing from search through execution and adaptability to ambiguity, so translation reshapes the*level*of each capability without re\-ordering which capabilities are easy\. Search is both the strongest capability and among the more robust, withKimi\-2\.6achieving the highest target\-language search of any agent \(84\.784\.7, edgingClaude\-4\.7\-Opus’s83\.283\.2andGPT\-5\.4’s83\.183\.1\) The frontier agents are most cross\-lingually robust on*adaptability*:Claude\-4\.7\-OpusandGemini\-3\.1\-Proretain it almost intact off English \(target\-language means of73\.073\.0and48\.948\.9against English75\.675\.6and51\.251\.2, leading to gaps of2\.62\.6and2\.32\.3points\), so their pooled loss is carried instead by execution and, especially, ambiguity, which is the hardest capability everywhere and the one on which every agent scores lowest off English\.

![Refer to caption](https://arxiv.org/html/2608.08775v1/x2.png)Figure 2:Per\-capability×\\timesper\-languagepass​@​3\\text\{pass\}@3\(%\) across agents\. Cross\-lingually robust cells, i\.e\. target language cells that are*not*significantly different from the corresponding English cell from the same \(agent, capability\) combination under a paired McNemar test \(p≥0\.05p\\geq 0\.05\), are marked with anorange corner wedge\.
### 5\.3Behavioural signatures

Comparing the English and target\-language distribution of lower\-level per\-scenario signals, such as agent step count and tool\-call composition \(which tools are invoked and in what proportion\) allows us to zoom in on the agents’ behavioural modes that shift under translation: inflated step counts and tool\-call retry loops may indicate the agent struggling with a partially\-understood instruction, whereas truncated planning or a collapsed tool\-call vocabulary might suggest the agent giving up early rather than exploring\. We focus our analysis on two proprietary agents \(Claude\-4\.7\-Opus,GPT\-5\.4\) and the strongest open\-weights agent \(Kimi\-2\.6\), and find two main behavioural signatures\.

#### Tool\-call volume\.

For each scenario we pair its target\-language runs with the matching English runs and compare the average number of tool calls\. OnlyGPT\-5\.4does systematically more work off English: it issues21\.6%21\.6\\%more tool calls than in English, an increase present in*every*target language and largest forhin\(\+41%\+41\\%\),spa\(\+37%\+37\\%\),tur\(\+23%\+23\\%\),deu\(\+22%\+22\\%\) andfra\(\+21%\+21\\%\)\.Claude\-4\.7\-OpusandKimi\-2\.6do the opposite, issuing slightly*fewer*calls off English on average \(−7\.6%\-7\.6\\%\), with per\-language changes mixed in sign \(Claude\-4\.7\-Opusfrom−23%\-23\\%onporto\+7%\+7\\%oncmn;Kimi\-2\.6from−17%\-17\\%onfra/ita/spato\+6%\+6\\%onhin\)\. This extra activity does not, however, explain the regression\. ForGPT\-5\.4the scenarios that lose accuracy off English inflate their tool count no more than those that hold it \(\+4\.8\+4\.8vs\+6\.5\+6\.5calls\), so tool volume and success move independently\. The one agent whose tool count does track its accuracy isClaude\-4\.7\-Opus, and it runs the*opposite*way: its regressing scenarios issue*fewer*calls, not more \(−11\.0\-11\.0vs\+3\.6\+3\.6calls\)—a “give up early” signature rather than flailing with extra actions\.

#### Tool\-strategy composition\.

All three agents also change*which*tools they reach for off English, and this shift is far larger than a cohort\-level average suggests\. We quantify it with the Jensen–Shannon divergence \(JSD\) between an agent’s English and target\-language tool\-choice distributions, where higher means a bigger strategy shift\. Pooled over a whole cohort the shift looks negligible \(0\.0100\.010/0\.0250\.025/0\.0200\.020bits forClaude\-4\.7\-Opus/GPT\-5\.4/Kimi\-2\.6\), because opposing shifts in different scenarios cancel out, but measured per matched scenario it is several times larger \(0\.0870\.087,0\.1090\.109,0\.1160\.116bits\), and larger still within matched thirds of each trajectory \(0\.1700\.170,0\.1970\.197,0\.2160\.216\), an approximate per\-turn view\. Two things follow: the pooled figure badly understates how much each agent re\-plans its tool use off English; and, measured per scenario, the three agents look much more alike,GPT\-5\.4’s apparent2\.5×2\.5\\timeslead overClaude\-4\.7\-Opusshrinks to about1\.3×1\.3\\times, andKimi\-2\.6re\-plans its tool use as much as either frontier agent\.

#### What changes\.

The direction of the shift is consistent across all three agents and across languages: off English they spend a larger fraction of their tool budget on*exploratory reads*and a smaller fraction on*state\-changing writes*\. The categories whose share grows most are retrieval and browsing operations: product search and catalogue listing rise by\+3\+3–66pp for every agent, with further gains on product\-detail lookups and conversation listing\. The per\-tool*falling*side is noisier \(dominated by scenario\-mix effects\), so we summarise the commit side by aggregating over the mutation flag each tool carries: the write\-action share of all tool calls falls from32\.1%32\.1\\%to29\.9%29\.9\\%forClaude\-4\.7\-Opus\(−2\.1\-2\.1pp\),30\.5%30\.5\\%to28\.1%28\.1\\%forGPT\-5\.4\(−2\.5\-2\.5pp\) and29\.3%29\.3\\%to28\.0%28\.0\\%forKimi\-2\.6\(−1\.3\-1\.3pp\)\. In other words, in a non\-English scenario all three agents forage longer over the environment and commit later and less decisively – the behavioural counterpart to the tool\-count\-mismatch failures of §[5\.4](https://arxiv.org/html/2608.08775#S5.SS4), where the wrong*number*of side\-effectful calls is the dominant non\-English failure signature\.

### 5\.4Verifier decomposition

To identify*which*verifier checks drive the regression, we label each failed run by the first verifier check it fails: quantitative \(e\.g\. a mismatched count of retrieved emails\), categorical \(e\.g\. mismatched attendee\-set for a calendar event\), textual \(e\.g\. rejected message content from the LLMaaJ\), overall tool\-count mismatch, or timeout\. Then, we report each class’s*absolute*failure rate, with all runs in the denominator, rather than its share of failures\. The share\-of\-failures view is flat across languages and hides the effect; the absolute view localises it\. The cross\-lingual regression is*not*carried by quantitative or categorical reasoning: from English to the non\-English average, quantitative\-checker failures move by only\+0\.04\+0\.04percentage points \(pp\) forClaude\-4\.7\-Opus,−0\.33\-0\.33pp forGPT\-5\.4\(i\.e\. slightly*improving*\) and\+0\.06\+0\.06pp forKimi\-2\.6, and categorical\-checker failures by\+0\.08\+0\.08,\+0\.05\+0\.05and−0\.12\-0\.12pp\. None of these deltas are significant under a matched\-pair McNemar test\. The regression is instead concentrated in the tool\-count\-mismatch checker \(\+10\.6\+10\.6ppClaude\-4\.7\-Opus,\+5\.3\+5\.3ppGPT\-5\.4,\+5\.7\+5\.7ppKimi\-2\.6\) and the textual\-reply checker \(\+6\.2\+6\.2,\+3\.4\+3\.4,\+3\.5\+3\.5pp\), which together account for essentially all of the\+17\.0\+17\.0/\+9\.3\+9\.3/\+11\.1\+11\.1pp rise in the overall non\-English failure rate\. The cross\-lingual gap is thus a tool\-orchestration and final\-response\-quality effect, not a numeric\- or categorical\-reasoning effect – consistent with the ambiguity over\-commitment of §[5\.2](https://arxiv.org/html/2608.08775#S5.SS2)\(tool\-count\) and the reply\-quality degradation discussed in §[6](https://arxiv.org/html/2608.08775#S6)\(textual\)\.

### 5\.5Scale ablation: theQwen\-3\.5size ladder

We vary*scale*within a single family to ask how agentic competence and the cross\-lingual gap of §[5\.2](https://arxiv.org/html/2608.08775#S5.SS2)move with model size\. We evaluate fourQwen\-3\.5sizes: one dense \(27B\) and three mixture\-of\-experts \(35B\-A3B,122B\-A10B,397B\-A17B, with33/1010/1717B active parameters\)888We omit checkpoints smaller than9B, for which the agent harness does not reliably elicit tool calls, yielding near\-zero valid trajectories\.\. We run the full benchmark suite of1111\-language×\\times44\-capabilities, under the same experimental setup\. Note that this scaling study is run on a different model generation fromQwen\-3\.6\-35B\-A3B, so it is meant to complement rather than extends Table[1](https://arxiv.org/html/2608.08775#S5.T1)\.

Table 2:Qwen\-3\.5scale ladder:pass​@​1\\text\{pass\}@1\(%\) pooled over the eleven languages, per capability and averaged \(timeexcluded\), with the active\-parameter count per size\. Rows are ordered by average score; best per column inbold\. Same protocol as the headline \(pass​@​1\\text\{pass\}@1,thinking\_effort==low, operational judge J3\); each cell uses a denominator of160160scenarios per language\-capability pair\. Per\-language cells are in Appendix[E](https://arxiv.org/html/2608.08775#A5), Table[17](https://arxiv.org/html/2608.08775#A5.T17)\.#### Scaling is monotonic, but a dense mid\-size rivals the largest MoE\.

Averagepass​@​1\\text\{pass\}@1rises monotonically with the ladder:10\.5→12\.9→18\.1→20\.010\.5\\to 12\.9\\to 18\.1\\to 20\.0\.397B\-A17Bis the strongest size on average and on search and adaptability\. But the dense27Bactually edges the6×6\\times\-larger397B\-A17Bon execution \(21\.421\.4vs20\.720\.7\) and ambiguity \(7\.67\.6vs7\.57\.5\), and clears122B\-A10Bcomfortably\. Aggregate capability thus tracks the*active*parameter count, and the dense mid\-size model is fully competitive with the largest sparse one per active parameter\.

#### The capability ordering is stable\.

The canonical ordering of search\>\>execution\>\>adaptability\>\>ambiguity holds across sizes, with ambiguity hardest throughout; the one exception is122B\-A10B, whose execution collapses to the level of its adaptability\. Scale lifts the whole profile rather than re\-ordering which capabilities are easy\.

#### The cross\-lingual gap does not close with scale\.

English is the ceiling at every size, and the gap to the target\-language mean*widens*as models grow: English minus the mean of the other ten languages is\+8\.1\+8\.1,\+7\.5\+7\.5,\+10\.4\+10\.4, and\+12\.7\+12\.7pp for35B\-A3B→\\to397B\-A17B\.hinandjpnare the weakest languages at every size \(e\.g\.397B\-A17B:eng31\.631\.6vshin9\.29\.2,jpn10\.010\.0\), mirroring the script\-family fragility of §[5\.2](https://arxiv.org/html/2608.08775#S5.SS2)\. Scaling the model therefore raises the multilingual floor but leaves the*shape*of the cross\-lingual gap intact, so the gap identified in this work is not one that a larger same\-family checkpoint closes on its own\.

## 6Error Attribution Gap: automatic and linguistic analysis

The leaderboard numbers \(§[5](https://arxiv.org/html/2608.08775#S5)\) surface a cross\-lingual gap; this section asks*why*it exists\. First, we define an error taxonomy; second, we do an automatic evaluation, and finally, we perform human linguistic analysis\.

### 6\.1Error Taxonomy

To attribute each multilingual regression to a root cause, we apply a five\-category verdict taxonomy, summarised below\. Categories are checked top\-down; the first that fits is assigned\.

1. 1\.Infrastructure \(I1–I2\)— The run crashed or no judgement was persisted; no gradable attempt exists\.
2. 2\.Translation defect \(T1–T4\)— The MT pipeline corrupted the task inputs \(task description, universe app state field, or oracle event\) such that the target run was set up to fail\. Sub\-codes cover oracle\-event mistranslation \(T1\), universe app\-state corruption \(T2\), task description ambiguity or meaning change \(T3\), and cross\-reference/transliteration breakage \(T4\)\.999In practice T4 is near\-empty as a distinct*construction\-level*defect: a dedicated non\-Latin script audit finds no transliteration errors introduced by benchmark construction \(App\.[G](https://arxiv.org/html/2608.08775#A7)\), and the few cross\-reference cases we observe share the T1 mechanism—an oracle\-compared literal left translated rather than passed through unchanged\.
3. 3\.Verifier artefact \(K1–K3\)— The agent’s action was effectively correct but the verifier rejected it—typically a translation\-sensitive*judge*\(LLM\-as\-Judge\) rejection \(K1\), a locale/format mismatch on a hard checker \(K2\), or list/ordering strictness \(K3\)\.
4. 4\.Agent failure \(A1–A7\)— The model genuinely erred on a fair, correctly\-translated task\. It includes: wrong content/args \(A1\), missing required action \(A2\), extra/spurious action \(A3\), wrong tool/target \(A4\), wrong final answer \(A5\), refusal/no answer \(A6\), non\-convergence \(A7\)\.
5. 5\.Inconclusive \(E1\)— Multiple plausible causes coexist or the deciding evidence \(e\.g\. the judge rationale\) is unavailable\.

### 6\.2Automatic estimation

To scale error attribution beyond hand\-picked cases, we automatically triage every cross\-lingual failure into the taxonomy of §[6\.1](https://arxiv.org/html/2608.08775#S6.SS1)and estimate how the cross\-lingual gap divides across its three fault sides\. After correcting for infrastructure aborts, stratifying regressions by within\-scenario determinism, and extrapolating to scenarios that fail in*both*languages, we find the gap to be majority genuine model failure \(55\.4%55\.4\\%\), with translation defects a smaller and spatially localised contamination floor \(34\.5%34\.5\\%of the gap; only6\.4%6\.4\\%of all benchmark pairs rendered unsolvable in the target\) and verifier artefacts the remainder \(10\.1%10\.1\\%\)\. Five\-annotator human re\-adjudication confirms the automatic fault\-side label on91\.4%91\.4\\%of a stratified sample\. The remainder of this subsection develops each step\.

#### Triage protocol\.

We attribute each target\-language failure to the taxonomy of Section[6\.1](https://arxiv.org/html/2608.08775#S6.SS1)with an LLM\-agent triage protocol that operates on the per\-scenario artefacts persisted by the evaluation harness: the graded verdict, the environment action log \(the agent’s write actions\), the agent trajectory, and, where available, the LLM\-as\-Judge rationale together with the oracle reference\. One diagnostic agent handles one unit of work: it reads these artefacts and, in*regression mode*, diffs a failing target run against a passing English run for the same scenario—the English run is the ground truth the dataset does not otherwise provide—before assigning a category, sub\-code, and confidence\. Judge rejections, which are otherwise ambiguous between a verifier artefact \(K1\) and a wrong final answer from the agent \(A5\), are adjudicated from the persisted judge rationale and the oracle reference rather than inferred from which check fired\. The protocol is orchestrated as an extract–diagnose–compile fan\-out and is described in full in Appendix[F](https://arxiv.org/html/2608.08775#A6)\.

#### Gap\-representative stratification\.

The severity of a regression is itself informative\. A deterministic translation defect corrupts the task and breaks*every*target run, whereas a stochastic model slip or a flaky judge breaks only*some*\. Sampling only complete collapses \(English 3/3, target 0/3\) therefore over\-represents translation defects\. We instead stratify infra\-clean regressions by within\-scenario determinism—*deterministic*\(the target never succeeds\) versus*stochastic*\(the target succeeds on a strict subset of runs\)—triage a language\- and capability\-balanced sample of each at the granularity of individual failing runs, and reweight each stratum’s cause composition by its true failing\-run volume \(known exactly from the run index\) to obtain a gap representative estimate\. Confidence intervals are obtained by a scenario\-clustered bootstrap\.

#### Composition of the gap\.

The two strata differ sharply \(Table[3](https://arxiv.org/html/2608.08775#S6.T3)\): deterministic failures are translation\-driven \(59%59\\%\), but stochastic failures—which carry the larger share of the gap—are model\-driven \(75%75\\%\)\. After reweighting \(Table[4](https://arxiv.org/html/2608.08775#S6.T4)\),55\.4%55\.4\\%\(95% CI50\.350\.3–60\.760\.7\) of the cross\-lingual gap is a genuine model failure,34\.5%34\.5\\%\(29\.529\.5–39\.139\.1\) a translation defect, and10\.1%10\.1\\%\(7\.17\.1–13\.513\.5\) a verifier artefact\. Restricting attention to complete collapses inverts this to a translation\-dominated≈55%/33%\\approx\\\!55\\%/33\\%split, quantifying the selection bias that the stratification removes\. The model\-driven share is language\-dependent, largest for Turkish and Japanese and smallest for Portuguese and Spanish \(Appendix[F](https://arxiv.org/html/2608.08775#A6)\)\.

Table 3:Cause composition of genuine target failures by regression stratum\. Deterministic collapses are translation\-driven; stochastic partial regressions are model\-driven\. Sampling only the former biases the estimate towards translation\.Table 4:Reweighted, infra\-clean decomposition of the cross\-lingual gap \(Claude 4\.7, ten target languages, all capabilities\)\. Infrastructure failures \(8\.5%8\.5\\%of runs\) are excluded; CIs from a scenario\-clustered bootstrap\.
#### From regressions to the whole benchmark\.

The triage above conditions on English*passing*; by construction it cannot see scenarios that fail in*both*languages, which is precisely where a translation defect could deflate a target score without leaving a visible regression\. To bound contamination over the entire benchmark we therefore partition all6,0566\{,\}056English×\\timestarget scenario–language pairs by outcome \(Appendix[F](https://arxiv.org/html/2608.08775#A6)Table[18](https://arxiv.org/html/2608.08775#A6.T18)\) and use a simple observation: a translation defect can only corrupt a score by rendering the target*unsolvable*, i\.e\., the target never passes\. The74\.5%74\.5\\%of pairs solved in the target on at least one run are thus translation\-clean by construction—the translated task is demonstrably solvable—leaving only the25\.5%25\.5\\%the target never solves to account for\.

#### The both\-fail blind spot\.

We triage all three “target\-never\-solves” cells, including the both\-fail cell that regression analysis misses\. The translation\-defect rate falls monotonically as English competence drops:59%59\\%where English solves the scenario \(a clean regression\),16%16\\%where English is already flaky, and only10%10\\%\(95% CI5\.65\.6–16\.916\.9\) where English also fails\. In other words, translation defects concentrate exactly where they are already*visible*as clean regressions, and the blind spot is if anything*cleaner*than the part we can see—ruling out a hidden reservoir of contamination—and in any case both\-fail defects cannot bias the cross\-lingual ranking, since a scenario English also fails contributes nothing to the gap\. Weighting the three cells by size, translation defects render only6\.4%6\.4\\%of all pairs unsolvable in the target, clustered on a handful of oracle\- and keyword\-mistranslation bugs\. The cross\-lingual gap therefore predominantly reflects genuine model capability, over a small, localised, and fixable translation\-contamination floor\.

#### Human\-based validation\.

To test the triage beyond hand\-picked cases, five annotators re\-adjudicated a stratified sample of triaged failures through a purpose\-built review interface \(Appendix[H](https://arxiv.org/html/2608.08775#A8)\), each confirming or rejecting the assigned fault side from decision\-relevant evidence alone \(source/translated task, oracle, judge rationale, agent answer\)\. Across140items in seven languages, they confirmed the automatic label in128—91\.4%agreement \(95% CI85\.685\.6–95\.095\.0\)—with comparable rates across all three sides \(Translation19/2219/22, Verifier25/2725/27, Agent84/9184/91\) and rejections spread across category boundaries with no systematic direction\. This corroborates the decomposition of Table[4](https://arxiv.org/html/2608.08775#S6.T4): independent human judgement agrees with the automatic fault\-side assignment on the large majority of a representative sample; full breakdown in Appendix[H](https://arxiv.org/html/2608.08775#A8)\.

### 6\.3Linguistic Analysis

Three linguists, native in Indonesian, Mandarin Chinese, and Spanish respectively—all proficient in English, with the Mandarin speaker also proficient in Japanese—independently inspected agent traces, translated prompts, and universe content for 16 scenarios spanning all four capabilities; six representative cases are presented in Table[5](https://arxiv.org/html/2608.08775#S6.T5)\.

Table 5:Cross\-lingual regression cases with per\-language failure analysis \(Claude\-4\.7\-Opus,pass​@​3\\text\{pass\}@3\)\. Bold marks regressions relative to English\.We organise the observations by the three non\-infrastructure fault sides of the taxonomy \(§[6\.1](https://arxiv.org/html/2608.08775#S6.SS1)\), mirroring the automatic decomposition of §[6\.2](https://arxiv.org/html/2608.08775#S6.SS2)\.

#### Translation defects \(T\)\.

Translation neutralises disambiguating surface cues\.Several prompts are ambiguous in English yet resolvable through word order or morphology; translation collapses these cues and forces a single reading that may not be the intended one \(“whose titles” rendered as*job title*rather than*event title*in every target language; Case 1, T3\)\.Faithful\-looking translations silently drop constraints\.A directional preposition \(“booked a cab*to*”\) was dropped or weakened in every target language, removing the very constraint that fixes the correct answer \(Case 2, T3\); the Spanish “a las que…reservar”*looks*directional but is incongruent with the verb’s non\-motion reading\.Morphological cue loss drives the most consequential failures\.In article\-less, number\-neutral languages \(cmn/jpn/ind\) the singular “my contact” is parsed as a generic set, and the agent deletes*both*matching contacts instead of asking for clarification \(Case 6, T3\)\.Controlled\-vocabulary drift compounds this: an inconsistently rendered property\-type enum leaves the target filter unreachable \(Case 5, T2\)\.

#### Verifier artefacts \(K\)\.

Judge strictness is language\-dependent, stochastic, and can hallucinate\.The soft judge rejects correct target\-language content while accepting identical English content, and in one case cited agent text that was never produced \(Case 4, K1\)\.Action\-count gating penalises harmless self\-corrections\.The hard count\-gate fails any mismatch between the agent’s and the oracle’s write\-action counts even when the final state is correct—e\.g\. an agent that omits an attendee, deletes the event, and recreates it correctly \(Cases 3–5\)\. Although the taxonomy records these as extra\-action agent errors \(A3\), the correct end state means they are largely*measurement*artefacts of end\-state\-blind grading rather than genuine capability regressions\.

#### Agent failures \(A\)\.

Genuine model errors are present and, in aggregate, dominant\.Consistent with the automatic decomposition \(§[6\.2](https://arxiv.org/html/2608.08775#S6.SS2)\), where model failures account for55\.4%55\.4\\%of the cross\-lingual gap, the sample contains clean slips traceable to neither MT nor the verifier: a Spanish run skips a required save step despite naming it in its own reasoning \(Case 3, A2\), and a retrieval miss returns only one of two matching contacts \(Case 6, spa\)\.Genuine over\-actions must be separated from grading artefacts\.Not every extra write is benign: an over\-action that changes the final state is a true A3 regression, unlike the idempotent self\-corrections penalised by the count\-gate above, and the two must be disentangled before an action\-count failure is read as a capability gap\.

Overall, the linguistic analysis illustrates each non\-infrastructure fault side of §[6\.1](https://arxiv.org/html/2608.08775#S6.SS1)with concrete cross\-lingual cases and is consistent with the automatic decomposition of §[6\.2](https://arxiv.org/html/2608.08775#S6.SS2): translation defects and verifier artefacts are real but bounded, while genuine, translation\-independent agent errors surface even in this small qualitative sample\.

## 7Conclusion

We deliverOmnilingualGAIA2, a machine\-translated multilingual expansion of theGAIA2agentic benchmark spanning ten priority target languages, together with a localised and human\-calibrated verifier and a cross\-lingual leaderboard for a heterogeneous seven\-agent cohort spanning frontier closed\-source systems and open\-weights models of both dense and mixture\-of\-experts architectures\.

Our main findings are as follows\. First, the cross\-lingual gap is universal in direction but agent\-specific in magnitude \(8\.8–18\.4 pp inpass​@​3\\text\{pass\}@3\), and no agent is uniformly robust: the frontier systems narrow the gap without closing it\. The metric inspected differentiates them:Gemini\-3\.1\-Prois the most self\-consistent \(smallestpass3all\\text\{pass\}^\{\\text\{all\}\}\_\{3\}gap\) whileClaude\-4\.7\-Opuskeeps high any\-of\-three accuracy off English yet loses per\-attempt reliability\. Second, the gap is predominantly model\-driven: a stratified error attribution assigns 55% of it to genuine agent failures, 35% to translation defects and 10% to verifier artefacts, while a benchmark\-wide bound leaves only 6\.4% of scenario–language pairs MT\-unsolvable; the gap concentrates on tool\-orchestration and response quality rather than quantitative or categorical reasoning\. Third, this gap does not close with model scale: a same\-family size ladder shows the English–target delta widening from \+8 to \+13 pp as active parameters grow\. Fourth, agents adopt a more hesitant execution strategy off English, foraging longer over the environment and committing later and less decisively \(the write\-action share of tool calls falls by 1–2\.5 pp\); this shift is the behavioural signature of degraded comprehension: it accompanies the mis\-counted write actions behind the execution\-side failures rather than reflecting a successful adaptation\. Finally, linguistic analysis identifies morphological cue loss and amplified ambiguity as primary failure mechanisms, particularly in non\-Latin\-script languages\.

Taken together, our results argue that multilingual agentic evaluation belongs alongside multilingual understanding evaluation as a standard part of the reporting protocol for globally deployed agents\. Several directions remain open\. The most immediate is to broaden coverage along two axes: capabilities \(e\.g\. the one we deferred in this work\) and target languages, ideally with a focus on the lower\-resource end where we posit the cross\-lingual gap is widest\. A complementary direction is to harden the translation pipeline \(§[3\.2](https://arxiv.org/html/2608.08775#S3.SS2)\) against the translation\-defect mechanisms our error analysis surfaces \(§[6\.3](https://arxiv.org/html/2608.08775#S6.SS3)\), driving down the6\.4%6\.4\\%translation\-induced unsolvability floor of the benchmark\.

## Acknowledgements

We thank David Dale for his valuable feedback in early versions of the paper\.

## Limitations

- •Excluded capabilities\.The threeGAIA2capabilities we do not evaluate \(Time, Noise and Agent2Agent\) are initially out of scope forOmnilingualGAIA2by design \(§[4](https://arxiv.org/html/2608.08775#S4)\): their scenarios either lack a scripted oracle trajectory or require a distinct verifier path that our translation pipeline does not currently cover\. Extending the pipeline and verifier localisation to those splits is future work\.
- •Translation quality\.OmnilingualGAIA2is constructed by an automatic machine\-translation pipeline in a translator\-only configuration, without human post\-editing of the full corpus\. The on\-pipeline signals of §[3\.2](https://arxiv.org/html/2608.08775#S3.SS2)attest to translated\-content quality but do not substitute for a thorough per\-language MT\-quality audit against professional human references\. As a result, a bounded but non\-zero residual translation\-defect floor persists in the evaluated release \(§[6\.2](https://arxiv.org/html/2608.08775#S6.SS2)\), concentrated in non\-Latin, morphologically rich target languages; although we show this floor to be relatively small, scores in the most affected languages should be read as carrying it\.
- •Judge reliability ceiling\.Our judge is calibrated against human annotations, but human annotators themselves agree only moderately on this task\. The judge’s high agreement scores should therefore be read as parity with an imperfect human reference rather than as absolute correctness; a higher\-agreement re\-annotation of the calibration set would be needed to tighten this ceiling\.
- •Closed\-system extended thinking\.We evaluate the three proprietary systems we benchmark \(Claude\-4\.7\-Opus,GPT\-5\.4,Gemini\-3\.1\-Pro\) with the provider default reasoning\-effort setting, without studying the impact that extended thinking might have on the multilingual gap\. We reserve this for future work\.
- •Harness choice\.All numbers are reported under theOpenClawharness, with the exception of Kimi\. These numbers should be read as observational comparisons under a fixed harness rather than as definitive model rankings; a different harness \(e\.g\. co\-developed with the model\) or a larger step/token budget could reorder the leaderboard\.
- •Human audit\.A per\-language human audit of cross\-lingual regressions is available only on a few languages and samples \(§[6](https://arxiv.org/html/2608.08775#S6)\)\. A full per\-language regression audit across the ten\-language matrix is future work\.

## Ethics Statement

OmnilingualGAIA2is an evaluation benchmark built on top ofGAIA2and contains no personal data beyond the synthetic content of the original scenarios\. Our goal is to improve the equity of AI agents by making cross\-lingual performance gaps measurable and visible; documenting that agents underperform for non\-English users is a prerequisite for closing that gap\. We caution that machine\-translated benchmarks can embed translationese and cultural mismatches, and we therefore do not treatOmnilingualGAIA2scores as a measure of culturally appropriate behaviour; this remains a concern for future research\. We will release the dataset and human annotations to support reproducibility and further research\.

## References

- Andrews et al\. \(2025\)Pierre Andrews, Mikel Artetxe, Mariano Coria Meglioli, Marta R\. Costa\-jussà, Joe Chuang, David Dale, Mark Duppenthaler, Nathanial Paul Ekberg, Cynthia Gao, Daniel Edward Licht, Jean Maillard, Alexandre Mourachko, Christophe Ropers, Safiyyah Saleem, Eduardo Sánchez, Ioannis Tsiamas, Arina Turkatenko, Albert Ventayol\-Boada, and Shireen Yates\.BOUQuET : dataset, benchmark and open initiative for universal quality evaluation in translation\.In Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Violet Peng, editors,*Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing*, pages 27515–27535, Suzhou, China, November 2025\. Association for Computational Linguistics\.ISBN 979\-8\-89176\-332\-6\.[10\.18653/v1/2025\.emnlp\-main\.1400](https://arxiv.org/doi.org/10.18653/v1/2025.emnlp-main.1400)\.[https://aclanthology\.org/2025\.emnlp\-main\.1400/](https://aclanthology.org/2025.emnlp-main.1400/)\.
- Anthropic \(2025\)Anthropic\.Claude 4 model card\.Technical report, Anthropic, 2025\.[https://www\.anthropic\.com/research/claude\-4\-system\-card](https://www.anthropic.com/research/claude-4-system-card)\.
- Artstein and Poesio \(2008\)Ron Artstein and Massimo Poesio\.Inter\-coder agreement for computational linguistics\.*Computational Linguistics*, 34\(4\):555–596, 2008\.[https://aclanthology\.org/J08\-4004](https://aclanthology.org/J08-4004)\.
- Bandarkar et al\. \(2026\)Lucas Bandarkar, Alan Ansell, and Trevor Cohn\.Large reasoning models struggle to transfer parametric knowledge across scripts, 2026\.[https://arxiv\.org/abs/2603\.17070](https://arxiv.org/abs/2603.17070)\.
- Bariah et al\. \(2026\)Lina Bariah, Brahim Mefgouda, Farbod Tavakkoli, Enrique Molero, Louis Powell, and Merouane Debbah\.Telcoagent\-bench: A multilingual benchmark for telecom ai agents, 2026\.[https://arxiv\.org/abs/2604\.06209](https://arxiv.org/abs/2604.06209)\.
- Barres et al\. \(2025\)Victor Barres, Honghua Dong, Soham Ray, Xujie Si, and Karthik Narasimhan\.τ2\\tau^\{2\}\-bench: Evaluating conversational agents in a dual\-control environment, 2025\.[https://arxiv\.org/abs/2506\.07982](https://arxiv.org/abs/2506.07982)\.
- Cohen \(1960\)Jacob Cohen\.A coefficient of agreement for nominal scales\.*Educational and Psychological Measurement*, 20\(1\):37–46, 1960\.
- Deng et al\. \(2026\)Xiang Deng, Jeff Da, Edwin Pan, Yannis Yiming He, Charles Ide, Kanak Garg, Niklas Lauffer, Andrew Park, Chetan Rane, Karmini Sampath, Maya Krishnan, Srivatsa R Kundurthy, Sean M\. Hendryx, Zifan Wang, Chen Bo Calvin Zhang, Noah Jacobson, Bing Liu, and Brad Kenstler\.SWE\-bench pro: Can AI agents solve long\-horizon software engineering tasks?In*Forty\-third International Conference on Machine Learning*, 2026\.[https://openreview\.net/forum?id=uEVTdoAbnK](https://openreview.net/forum?id=uEVTdoAbnK)\.
- Doğruöz et al\. \(2026\)A\. Seza Doğruöz, Xixian Liao, Verena Blaschke, Jakob Prange, Senyu Li, and David Ifeoluwa Adelani\.Challenges and recommendations for LLMs\-as\-a\-judge in multilingual settings and low\-resource languages, 2026\.[https://arxiv\.org/abs/2607\.02235](https://arxiv.org/abs/2607.02235)\.
- Franklin and Graesser \(1996\)Stan Franklin and Art Graesser\.Is it an agent, or just a program?: A taxonomy for autonomous agents\.In*International workshop on agent theories, architectures, and languages*, pages 21–35\. Springer, 1996\.
- Froger et al\. \(2026\)Romain Froger, Pierre Andrews, Matteo Bettini, Amar Budhiraja, Ricardo Silveira Cabral, Virginie Do, Emilien Garreau, Jean\-Baptiste Gaya, Hugo Laurençon, Maxime Lecanu, Kunal Malkan, Dheeraj Mekala, Pierre Menard, Gerard Moreno\-Torres Bertran, Ulyana Piterbarg, Mikhail Plekhanov, Mathieu Rita, Andrey Rusakov, Vladislav Vorotilov, Mengjue Wang, Ian Yu, Amine Benhalloum, Grégoire Mialon, and Thomas Scialom\.Gaia2: Benchmarking LLM agents on dynamic and asynchronous environments\.In*The Fourteenth International Conference on Learning Representations*, 2026\.[https://openreview\.net/forum?id=9gw03JpKK4](https://openreview.net/forum?id=9gw03JpKK4)\.
- Google DeepMind \(2025a\)Google DeepMind\.Gemini: A family of highly capable multimodal models\.*arXiv preprint arXiv:2312\.11805*, 2025a\.
- Google DeepMind \(2025b\)Google DeepMind\.Gemma: Open models based on Gemini research and technology\.*arXiv preprint arXiv:2403\.08295*, 2025b\.
- Gureja et al\. \(2025\)Srishti Gureja, Lester James V\. Miranda, Shayekh Bin Islam, Rishabh Maheshwary, Drishti Sharma, Gusti Winata, Nathan Lambert, Sebastian Ruder, Sara Hooker, and Marzieh Fadaee\.M\-RewardBench: Evaluating reward models in multilingual settings\.In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar, editors,*Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 43–58, Vienna, Austria, July 2025\. Association for Computational Linguistics\.ISBN 979\-8\-89176\-251\-0\.[10\.18653/v1/2025\.acl\-long\.3](https://arxiv.org/doi.org/10.18653/v1/2025.acl-long.3)\.[https://aclanthology\.org/2025\.acl\-long\.3/](https://aclanthology.org/2025.acl-long.3/)\.
- Hada et al\. \(2024\)Rishav Hada, Varun Gumma, Mohamed Ahmed, Kalika Bali, and Sunayana Sitaram\.METAL: Towards multilingual meta\-evaluation\.In Kevin Duh, Helena Gomez, and Steven Bethard, editors,*Findings of the Association for Computational Linguistics: NAACL 2024*, pages 2280–2298, Mexico City, Mexico, June 2024\. Association for Computational Linguistics\.[10\.18653/v1/2024\.findings\-naacl\.148](https://arxiv.org/doi.org/10.18653/v1/2024.findings-naacl.148)\.[https://aclanthology\.org/2024\.findings\-naacl\.148/](https://aclanthology.org/2024.findings-naacl.148/)\.
- Hofman et al\. \(2026\)Omer Hofman, Jonathan Brokman, Oren Rachmil, Shamik Bose, Vikas Pahuja, Toshiya Shimizu, Trisha Starostina, Kelly Marchisio, Seraphina Goldfarb\-Tarrant, and Roman Vainshtein\.MAPS: A multilingual benchmark for agent performance and security\.In Vera Demberg, Kentaro Inui, and Lluís Marquez, editors,*Findings of the Association for Computational Linguistics: EACL 2026*, pages 821–845, Rabat, Morocco, March 2026\. Association for Computational Linguistics\.ISBN 979\-8\-89176\-386\-9\.[10\.18653/v1/2026\.findings\-eacl\.42](https://arxiv.org/doi.org/10.18653/v1/2026.findings-eacl.42)\.[https://aclanthology\.org/2026\.findings\-eacl\.42/](https://aclanthology.org/2026.findings-eacl.42/)\.
- Johnston et al\. \(2026\)Drew Johnston, David Holtz, Alex Martin Richmond, Christopher Ong, Prasanna Tambe, and Aaron Chatterji\.The shift to agentic ai: Evidence from codex, 2026\.[https://openai\.com/index/how\-agents\-are\-transforming\-work/](https://openai.com/index/how-agents-are-transforming-work/)\.
- Juraska et al\. \(2025\)Juraj Juraska, Tobias Domhan, Mara Finkelstein, Tetsuji Nakagawa, Geza Kovacs, Daniel Deutsch, Pidong Wang, and Markus Freitag\.MetricX\-25 and GemSpanEval: Google Translate submissions to the WMT25 evaluation shared task\.In Barry Haddow, Tom Kocmi, Philipp Koehn, and Christof Monz, editors,*Proceedings of the Tenth Conference on Machine Translation*, pages 957–968, Suzhou, China, November 2025\. Association for Computational Linguistics\.ISBN 979\-8\-89176\-341\-8\.[10\.18653/v1/2025\.wmt\-1\.70](https://arxiv.org/doi.org/10.18653/v1/2025.wmt-1.70)\.[https://aclanthology\.org/2025\.wmt\-1\.70/](https://aclanthology.org/2025.wmt-1.70/)\.
- KC \(2026\)Shreyas KC\.BabelJudge: Measuring LLM\-as\-a\-judge reliability across languages and agent trajectories, 2026\.[https://arxiv\.org/abs/2606\.22329](https://arxiv.org/abs/2606.22329)\.
- Kim et al\. \(2026\)Yunsu Kim, Kaden Uhlig, and Joern Wuebker\.Gaia\-v2\-lilt: Multilingual adaptation of agent benchmark beyond translation, 2026\.[https://arxiv\.org/abs/2604\.24929](https://arxiv.org/abs/2604.24929)\.
- Kimi\-Team \(2025\)Kimi\-Team\.Kimi 2\.6 technical report\.Technical report, 2025\.[https://www\.kimi\.com/blog/kimi\-k2\-6](https://www.kimi.com/blog/kimi-k2-6)\.
- Kocmi and Federmann \(2023\)Tom Kocmi and Christian Federmann\.Large language models are state\-of\-the\-art evaluators of translation quality\.In*Proceedings of the 24th Annual Conference of the European Association for Machine Translation*, pages 193–203, Tampere, Finland, June 2023\. European Association for Machine Translation\.[https://aclanthology\.org/2023\.eamt\-1\.19/](https://aclanthology.org/2023.eamt-1.19/)\.
- Koo et al\. \(2024\)Ryan Koo, Minhwa Lee, Vipul Raheja, Jong Inn Park, Zae Myung Kim, and Dongyeop Kang\.Benchmarking cognitive biases in large language models as evaluators\.In*Findings of the Association for Computational Linguistics \(ACL Findings\)*, 2024\.[https://aclanthology\.org/2024\.findings\-acl\.29/](https://aclanthology.org/2024.findings-acl.29/)\.
- Krippendorff \(2019\)Klaus Krippendorff\.*Content Analysis: An Introduction to Its Methodology*\.Sage Publications, 4 edition, 2019\.
- Kulkarni et al\. \(2025\)Mayank Kulkarni, Vittorio Mazzia, Judith Gaspers, Chris Hench, and Jack FitzGerald\.MASSIVE\-agents: A benchmark for multilingual function\-calling in 52 languages\.In Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Violet Peng, editors,*Findings of the Association for Computational Linguistics: EMNLP 2025*, pages 20193–20215, Suzhou, China, November 2025\. Association for Computational Linguistics\.ISBN 979\-8\-89176\-335\-7\.[10\.18653/v1/2025\.findings\-emnlp\.1099](https://arxiv.org/doi.org/10.18653/v1/2025.findings-emnlp.1099)\.[https://aclanthology\.org/2025\.findings\-emnlp\.1099/](https://aclanthology.org/2025.findings-emnlp.1099/)\.
- Landis and Koch \(1977\)J\. Richard Landis and Gary G\. Koch\.The measurement of observer agreement for categorical data\.*Biometrics*, 33\(1\):159–174, 1977\.
- Levi and Kadar \(2025\)Elad Levi and Ilan Kadar\.Intellagent: A multi\-agent framework for evaluating conversational ai systems, 2025\.[https://arxiv\.org/abs/2501\.11067](https://arxiv.org/abs/2501.11067)\.
- Li et al\. \(2026\)Hongliang Li, Yijin Liu, Zhiwei Zhang, Zihe Liu, Xinyue Lou, Jinan Xu, Fandong Meng, and Kaiyu Huang\.PolyWorkBench: Benchmarking multilingual long\-horizon LLM agents, 2026\.[https://arxiv\.org/abs/2607\.06008](https://arxiv.org/abs/2607.06008)\.
- Lobentanzer \(2026\)Sebastian Lobentanzer\.Quantifying the expectation–realisation gap for agentic ai systems, 2026\.[https://arxiv\.org/abs/2602\.20292](https://arxiv.org/abs/2602.20292)\.
- Luo et al\. \(2026\)Zheng Luo, T Pranav Kutralingam, Ogochukwu N Okoani, Wanpeng Xu, Hua Wei, and Xiyang Hu\.Lost in execution: On the multilingual robustness of tool calling in large language models, 2026\.[https://arxiv\.org/abs/2601\.05366](https://arxiv.org/abs/2601.05366)\.
- Massenkoff and McCrory \(2026\)Maxim Massenkoff and Peter McCrory\.Labor market impacts of AI: A new measure and early evidence, 2026\.[https://www\.anthropic\.com/research/labor\-market\-impacts](https://www.anthropic.com/research/labor-market-impacts)\.
- McNemar \(1947\)Quinn McNemar\.Note on the sampling error of the difference between correlated proportions or percentages\.*Psychometrika*, 12\(2\):153–157, 1947\.
- Merrill et al\. \(2026\)Mike A Merrill, Alexander G Shaw, Nicholas Carlini, Boxuan Li, Harsh Raj, Ivan Bercovich, Lin Shi, Jeong Yeon Shin, Thomas Walshe, E Kelly Buchanan, et al\.Terminal\-bench: Benchmarking agents on hard, realistic tasks in command line interfaces, 2026\.[https://arxiv\.org/abs/2601\.11868](https://arxiv.org/abs/2601.11868)\.
- Meta AI \(2024\)Meta AI\.Llama 3\.3 70B Model Card\.[https://github\.com/meta\-llama/llama\-models/blob/main/models/llama3\_3/MODEL\_CARD\.md](https://github.com/meta-llama/llama-models/blob/main/models/llama3_3/MODEL_CARD.md), 2024\.Released 2024\-12\-06\.
- Mialon et al\. \(2024\)Grégoire Mialon, Clémentine Fourrier, Craig Swift, Thomas Wolf, Yann LeCun, and Thomas Scialom\.GAIA: a benchmark for general AI assistants\.In*The Twelfth International Conference on Learning Representations \(ICLR\)*, 2024\.[https://arxiv\.org/abs/2311\.12983](https://arxiv.org/abs/2311.12983)\.
- Nguyen et al\. \(2026\)My Chiffon Nguyen, Aulia Adila, Saksorn Ruangtanusak, Kittiphat Leesombatwathana, Vissuta Gunawan Lim, Patomporn Payoungkhamdee, and Samuel Cahyawijaya\.SEATauBench: Adapting tool\-agent\-user evaluation into low\-resource southeast asian languages, 2026\.[https://arxiv\.org/abs/2606\.28715](https://arxiv.org/abs/2606.28715)\.
- Omnilingual\-ASR\-team et al\. \(2025\)Omnilingual\-ASR\-team, Gil Keren, Artyom Kozhevnikov, Yen Meng, Christophe Ropers, Matthew Setzler, Skyler Wang, Ife Adebara, Michael Auli, Can Balioglu, Kevin Chan, Chierh Cheng, Joe Chuang, Caley Droof, Mark Duppenthaler, Paul\-Ambroise Duquenne, Alexander Erben, Cynthia Gao, Gabriel Mejia Gonzalez, Kehan Lyu, Sagar Miglani, Vineel Pratap, Kaushik Ram Sadagopan, Safiyyah Saleem, Arina Turkatenko, Albert Ventayol\-Boada, Zheng\-Xin Yong, Yu\-An Chung, Jean Maillard, Rashel Moritz, Alexandre Mourachko, Mary Williamson, and Shireen Yates\.Omnilingual asr: Open\-source multilingual speech recognition for 1600\+ languages, 2025\.[https://arxiv\.org/abs/2511\.09690](https://arxiv.org/abs/2511.09690)\.
- Omnilingual\-MT\-Team et al\. \(2026\)Omnilingual\-MT\-Team, Belen Alastruey, Niyati Bafna, Andrea Caciolai, Kevin Heffernan, Artyom Kozhevnikov, Christophe Ropers, Eduardo Sánchez, Charles\-Eric Saint\-James, Ioannis Tsiamas, Xiang "Tony" Cao, Chierh Cheng, Joe Chuang, Paul\-Ambroise Duquenne, Mark Duppenthaler, Nate Ekberg, Cynthia Gao, Pere Lluís Huguet Cabot, João Maria Janeiro, Jean Maillard, Gabriel Mejia Gonzalez, Holger Schwenk, Edan Toledo, Arina Turkatenko, Albert Ventayol\-Boada, Rashel Moritz, Alexandre Mourachko, Surya Parimi, Mary Williamson, Shireen Yates, David Dale, and Marta R\. Costa\-jussà\.Omnilingual mt: Machine translation for 1,600 languages, 2026\.[https://arxiv\.org/abs/2603\.16309](https://arxiv.org/abs/2603.16309)\.
- OpenAI \(2025\)OpenAI\.GPT\-5 system card\.Technical report, OpenAI, 2025\.[https://openai\.com/index/gpt\-5\-system\-card](https://openai.com/index/gpt-5-system-card)\.
- OpenAI \(2025\)OpenAI\.gpt\-oss\-120b & gpt\-oss\-20b Model Card, 2025\.[https://arxiv\.org/abs/2508\.10925](https://arxiv.org/abs/2508.10925)\.
- Panickssery et al\. \(2024\)Arjun Panickssery, Samuel R\. Bowman, and Shi Feng\.LLM evaluators recognize and favor their own generations, 2024\.[https://arxiv\.org/abs/2404\.13076](https://arxiv.org/abs/2404.13076)\.
- Patil et al\. \(2025\)Shishir G Patil, Huanzhi Mao, Fanjia Yan, Charlie Cheng\-Jie Ji, Vishnu Suresh, Ion Stoica, and Joseph E Gonzalez\.The berkeley function calling leaderboard \(bfcl\): From tool use to agentic evaluation of large language models\.In*Forty\-second International Conference on Machine Learning*, 2025\.
- Qin et al\. \(2024\)Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, et al\.Toolllm: Facilitating large language models to master 16000\+ real\-world apis\.In*International Conference on Learning Representations*, volume 2024, pages 9695–9717, 2024\.
- Qwen Team \(2025\)Qwen Team\.Qwen3 technical report\.*arXiv preprint arXiv:2505\.09388*, 2025\.
- Rei et al\. \(2022\)Ricardo Rei, Marcos Treviso, Nuno M\. Guerreiro, Chrysoula Zerva, Ana C Farinha, Christine Maroti, José G\. C\. de Souza, Taisiya Glushkova, Duarte Alves, Luisa Coheur, Alon Lavie, and André F\. T\. Martins\.CometKiwi: IST\-unbabel 2022 submission for the quality estimation shared task\.In*Proceedings of the Seventh Conference on Machine Translation \(WMT\)*, pages 634–645, Abu Dhabi, United Arab Emirates \(Hybrid\), December 2022\. Association for Computational Linguistics\.[10\.18653/v1/2022\.wmt\-1\.60](https://arxiv.org/doi.org/10.18653/v1/2022.wmt-1.60)\.[https://aclanthology\.org/2022\.wmt\-1\.60/](https://aclanthology.org/2022.wmt-1.60/)\.
- Russell and Norvig \(2009\)Stuart J Russell and Peter Norvig\.*Artificial intelligence a modern approach*\.Pearson Education, Inc\., 2009\.
- Sales Almeida et al\. \(2025\)Thales Sales Almeida, João Guilherme Alves Santos, Thiago Laitz, and Giovana Kerche Bonás\.Ticket\-Bench: A kickoff for multilingual and regionalized agent evaluation, 2025\.[https://arxiv\.org/abs/2509\.14477](https://arxiv.org/abs/2509.14477)\.
- Shi et al\. \(2026\)Dingfeng Shi, Jingyi Cao, Qianben Chen, Weichen Sun, Weizhen Li, Hongxuan Lu, Fangchen Dong, Tianrui Qin, King Zhu, Minghao Liu, Yuchen Eleanor Jiang, Jian Yang, Ge Zhang, Jiaheng Liu, Changwang Zhang, Jun Wang, and Wangchunshu Zhou\.Taskcraft: Automated generation of agentic tasks\.In*The Fourteenth International Conference on Learning Representations*, 2026\.[https://openreview\.net/forum?id=UJFCyrYM1V](https://openreview.net/forum?id=UJFCyrYM1V)\.
- Singh et al\. \(2025\)Shivalika Singh, Angelika Romanou, Clémentine Fourrier, David Ifeoluwa Adelani, Jian Gang Ngui, Daniel Vila\-Suero, Peerat Limkonchotiwat, Kelly Marchisio, Wei Qi Leong, Yosephine Susanto, et al\.Global mmlu: Understanding and addressing cultural and linguistic biases in multilingual evaluation\.In*Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 18761–18799, 2025\.
- Son et al\. \(2024\)Guijin Son, Dongkeun Yoon, Juyoung Suk, Javier Aula\-Blasco, Mano Aslan, Vu Trong Kim, Shayekh Bin Islam, Jaume Prats\-Cristià, Lucía Tormo\-Bañuelos, and Seungone Kim\.MM\-Eval: A multilingual meta\-evaluation benchmark for LLM\-as\-a\-judge and reward models, 2024\.[https://arxiv\.org/abs/2410\.17578](https://arxiv.org/abs/2410.17578)\.
- Staufer et al\. \(2026\)Leon Staufer, Kevin Feng, Kevin Wei, Luke Bailey, Yawen Duan, Mick Yang, A Pinar Ozisik, Stephen Casper, and Noam Kolt\.The 2025 ai agent index: Documenting technical and safety features of deployed agentic ai systems, 2026\.[https://arxiv\.org/abs/2602\.17753](https://arxiv.org/abs/2602.17753)\.
- Team et al\. \(2025\)Gemma Team, Aishwarya Kamath, Johan Ferret, Shreya Pathak, Nino Vieillard, Ramona Merhej, Sarah Perrin, Tatiana Matejovicova, Alexandre Ramé, Morgane Rivière, Louis Rouillard, Thomas Mesnard, Geoffrey Cideron, Jean bastien Grill, Sabela Ramos, Edouard Yvinec, Michelle Casbon, Etienne Pot, Ivo Penchev, Gaël Liu, Francesco Visin, Kathleen Kenealy, Lucas Beyer, Xiaohai Zhai, Anton Tsitsulin, Robert Busa\-Fekete, Alex Feng, Noveen Sachdeva, Benjamin Coleman, Yi Gao, Basil Mustafa, Iain Barr, Emilio Parisotto, David Tian, Matan Eyal, Colin Cherry, Jan\-Thorsten Peter, Danila Sinopalnikov, Surya Bhupatiraju, Rishabh Agarwal, Mehran Kazemi, Dan Malkin, Ravin Kumar, David Vilar, Idan Brusilovsky, Jiaming Luo, Andreas Steiner, Abe Friesen, Abhanshu Sharma, Abheesht Sharma, Adi Mayrav Gilady, Adrian Goedeckemeyer, Alaa Saade, Alex Feng, Alexander Kolesnikov, Alexei Bendebury, Alvin Abdagic, Amit Vadi, András György, André Susano Pinto, Anil Das, Ankur Bapna, Antoine Miech, Antoine Yang, Antonia Paterson, Ashish Shenoy, Ayan Chakrabarti, Bilal Piot, Bo Wu, Bobak Shahriari, Bryce Petrini, Charlie Chen, Charline Le Lan, Christopher A\. Choquette\-Choo, CJ Carey, Cormac Brick, Daniel Deutsch, Danielle Eisenbud, Dee Cattle, Derek Cheng, Dimitris Paparas, Divyashree Shivakumar Sreepathihalli, Doug Reid, Dustin Tran, Dustin Zelle, Eric Noland, Erwin Huizenga, Eugene Kharitonov, Frederick Liu, Gagik Amirkhanyan, Glenn Cameron, Hadi Hashemi, Hanna Klimczak\-Plucińska, Harman Singh, Harsh Mehta, Harshal Tushar Lehri, Hussein Hazimeh, Ian Ballantyne, Idan Szpektor, Ivan Nardini, Jean Pouget\-Abadie, Jetha Chan, Joe Stanton, John Wieting, Jonathan Lai, Jordi Orbay, Joseph Fernandez, Josh Newlan, Ju yeong Ji, Jyotinder Singh, Kat Black, Kathy Yu, Kevin Hui, Kiran Vodrahalli, Klaus Greff, Linhai Qiu, Marcella Valentine, Marina Coelho, Marvin Ritter, Matt Hoffman, Matthew Watson, Mayank Chaturvedi, Michael Moynihan, Min Ma, Nabila Babar, Natasha Noy, Nathan Byrd, Nick Roy, Nikola Momchev, Nilay Chauhan, Noveen Sachdeva, Oskar Bunyan, Pankil Botarda, Paul Caron, Paul Kishan Rubenstein, Phil Culliton, Philipp Schmid, Pier Giuseppe Sessa, Pingmei Xu, Piotr Stanczyk, Pouya Tafti, Rakesh Shivanna, Renjie Wu, Renke Pan, Reza Rokni, Rob Willoughby, Rohith Vallu, Ryan Mullins, Sammy Jerome, Sara Smoot, Sertan Girgin, Shariq Iqbal, Shashir Reddy, Shruti Sheth, Siim Põder, Sijal Bhatnagar, Sindhu Raghuram Panyam, Sivan Eiger, Susan Zhang, Tianqi Liu, Trevor Yacovone, Tyler Liechty, Uday Kalra, Utku Evci, Vedant Misra, Vincent Roseberry, Vlad Feinberg, Vlad Kolesnikov, Woohyun Han, Woosuk Kwon, Xi Chen, Yinlam Chow, Yuvein Zhu, Zichuan Wei, Zoltan Egyed, Victor Cotruta, Minh Giang, Phoebe Kirk, Anand Rao, Kat Black, Nabila Babar, Jessica Lo, Erica Moreira, Luiz Gustavo Martins, Omar Sanseviero, Lucas Gonzalez, Zach Gleicher, Tris Warkentin, Vahab Mirrokni, Evan Senter, Eli Collins, Joelle Barral, Zoubin Ghahramani, Raia Hadsell, Yossi Matias, D\. Sculley, Slav Petrov, Noah Fiedel, Noam Shazeer, Oriol Vinyals, Jeff Dean, Demis Hassabis, Koray Kavukcuoglu, Clement Farabet, Elena Buchatskaya, Jean\-Baptiste Alayrac, Rohan Anil, Dmitry, Lepikhin, Sebastian Borgeaud, Olivier Bachem, Armand Joulin, Alek Andreev, Cassidy Hardin, Robert Dadashi, and Léonard Hussenot\.Gemma 3 technical report, 2025\.[https://arxiv\.org/abs/2503\.19786](https://arxiv.org/abs/2503.19786)\.
- Wang et al\. \(2024\)Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, et al\.A survey on large language model based autonomous agents\.*Frontiers of Computer Science*, 18\(6\):186345, 2024\.
- Wang et al\. \(2025\)Peng Wang, Ruihan Tao, Qiguang Chen, Mengkang Hu, and Libo Qin\.X\-webagentbench: A multilingual interactive web benchmark for evaluating global agentic system\.In*Findings of the Association for Computational Linguistics: ACL 2025*, 2025\.[https://aclanthology\.org/2025\.findings\-acl\.988/](https://aclanthology.org/2025.findings-acl.988/)\.
- Wilson \(1927\)Edwin B\. Wilson\.Probable inference, the law of succession, and statistical inference\.*Journal of the American Statistical Association*, 22\(158\):209–212, 1927\.
- Wooldridge and Jennings \(1995\)Michael Wooldridge and Nicholas R Jennings\.Intelligent agents: Theory and practice\.*The knowledge engineering review*, 10\(2\):115–152, 1995\.
- Xi et al\. \(2025\)Zhiheng Xi, Wenxiang Chen, Xin Guo, Wei He, Yiwen Ding, Boyang Hong, Ming Zhang, Junzhe Wang, Senjie Jin, Enyu Zhou, et al\.The rise and potential of large language model based agents: A survey\.*Science China Information Sciences*, 68\(2\):121101, 2025\.
- Xie et al\. \(2026\)Jingxu Xie, Dylan Xu, Xuandong Zhao, and Dawn Song\.Agentsynth: Scalable task generation for generalist computer\-use agents, 2026\.[https://arxiv\.org/abs/2506\.14205](https://arxiv.org/abs/2506.14205)\.
- Xie et al\. \(2024\)Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh J Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, et al\.Osworld: Benchmarking multimodal agents for open\-ended tasks in real computer environments\.*Advances in Neural Information Processing Systems*, 37:52040–52094, 2024\.
- Xuan et al\. \(2025\)Weihao Xuan, Rui Yang, Heli Qi, Qingcheng Zeng, Yunze Xiao, Aosong Feng, Dairui Liu, Yun Xing, Junjue Wang, Fan Gao, Jinghui Lu, Yuang Jiang, Huitao Li, Xin Li, Kunyu Yu, Ruihai Dong, Shangding Gu, Yuekang Li, Xiaofei Xie, Felix Juefei\-Xu, Foutse Khomh, Osamu Yoshie, Qingyu Chen, Douglas Teodoro, Nan Liu, Randy Goebel, Lei Ma, Edison Marrese\-Taylor, Shijian Lu, Yusuke Iwasawa, Yutaka Matsuo, and Irene Li\.MMLU\-ProX: A multilingual benchmark for advanced large language model evaluation\.In Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Violet Peng, editors,*Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing*, pages 1513–1532, Suzhou, China, November 2025\. Association for Computational Linguistics\.ISBN 979\-8\-89176\-332\-6\.[10\.18653/v1/2025\.emnlp\-main\.79](https://arxiv.org/doi.org/10.18653/v1/2025.emnlp-main.79)\.[https://aclanthology\.org/2025\.emnlp\-main\.79/](https://aclanthology.org/2025.emnlp-main.79/)\.
- Yagubyan \(2026\)Abel Yagubyan\.The coin flip judge? reliability and bias in LLM\-as\-a\-judge evaluation\.*arXiv preprint arXiv:2606\.13685*, 2026\.[https://arxiv\.org/abs/2606\.13685](https://arxiv.org/abs/2606.13685)\.
- Yang et al\. \(2025\)Jeremy Yang, Noah Yonack, Kate Zyskowski, Denis Yarats, Johnny Ho, and Jerry Ma\.The adoption and usage of ai agents: Early evidence from perplexity, 2025\.[https://arxiv\.org/abs/2512\.07828](https://arxiv.org/abs/2512.07828)\.
- Zhang et al\. \(2026\)Hongbin Zhang, Kehai Chen, Xuefen Bai, Youcheng Pan, Yang Xiang, Jinpeng Wang, and Min Zhang\.Mitigating translationese bias in multilingual LLM\-as\-a\-judge via disentangled information bottleneck\.*arXiv preprint arXiv:2603\.10351*, 2026\.[https://arxiv\.org/abs/2603\.10351](https://arxiv.org/abs/2603.10351)\.
- Zheng et al\. \(2023\)Lianmin Zheng, Wei\-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E\. Gonzalez, and Ion Stoica\.Judging LLM\-as\-a\-judge with MT\-bench and chatbot arena\.In*Thirty\-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track*, 2023\.[https://openreview\.net/forum?id=uccHPGDlao](https://openreview.net/forum?id=uccHPGDlao)\.
- Zhu et al\. \(2026\)Yuxuan Zhu, Tengjun Jin, Yada Pruksachatkun, Andy K Zhang, Shu Liu, Sasha Cui, Sayash Kapoor, Shayne Longpre, Kevin Meng, Rebecca Weiss, Fazl Barez, Rahul Gupta, Jwala Dhamala, Jacob Merizian, Mario Giulianelli, Harry Coppock, Cozmin Ududec, Antony Kellermann, Jasjeet S Sekhon, Jacob Steinhardt, Sarah Schwettmann, Arvind Narayanan, Matei Zaharia, Ion Stoica, Percy Liang, and Daniel Kang\.Establishing best practices in building rigorous agentic benchmarks\.In*The Thirty\-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track*, 2026\.[https://openreview\.net/forum?id=E58HNCqoaA](https://openreview.net/forum?id=E58HNCqoaA)\.

## Appendix ALanguages

Table[6](https://arxiv.org/html/2608.08775#A1.T6)specifies the languages coverage of several agentic datasets, including ours\.

Table 6:Language coverage comparison between MAPSHofman et al\. \([2026](https://arxiv.org/html/2608.08775#bib.bib16)\), GAIA\-v2\-LILTKim et al\. \([2026](https://arxiv.org/html/2608.08775#bib.bib20)\), and our benchmark\.
## Appendix BThe translation pipeline

In this section we report the results of the empirical studies we carry out that motivate the design of the final translation pipeline we adopt to produce the benchmark, as described in §[3\.2](https://arxiv.org/html/2608.08775#S3.SS2)\. In particular, these motivate the choice ofGemma\-4\-31B\-Instructas the translator, and the choice to run it in a single translator\-only pass with no downstream reviewer\. Both decisions rest on the same pairwise protocol\. Every translatable span produced by a candidate configuration is compared against the corresponding span of a*reference translation*, a fixed, high\-quality rendering of the same content\. The comparison is evaluated by an LLM judge that is tasked to return one of \{*equivalent*,*candidate\-wins*,*reference\-wins*,*both\-wrong*\}\. In particular, we consider a fixed sample of200200spans, stratified by app, and score each \(translator, language\) combination against two LLM judges,GPT\-5\.4andGemini\-3\.1\-Pro; each judge’s verdict is the majority over three position\-randomised trials \(thus controlling for the known first\-position bias of pairwise LLM judges\), so the two judges together contribute400400pooled verdicts per combination\. We report*candidate\-win*rate \(judge prefers the candidate over the reference, higher is better\) and*reference\-win*rate \(lower is better\)\.

Both candidate translators run under a1616K\-token context budget, within which each per\-scenario prompt must accommodate the source spans, the generated translation, and the term table of §[3\.2](https://arxiv.org/html/2608.08775#S3.SS2)\. Since the raw term table—up to roughly a thousand entries per universe—would overflow this budget on its own, we cap it at its200200shortest\-source entries; a reproducibility check onspafinds that the capped and uncapped tables yieldGemma\-4\-31B\-Instructtranslations differing by under2\.52\.5pp on every judged cell, confirming the cap is quality\-neutral\. Within the prompt this table is rendered inline as theTerminologyblock shown in[Figure˜3](https://arxiv.org/html/2608.08775#A2.F3), and is omitted for scenarios that share no terms across surfaces\. Both translators are executed with the reasoning \(“thinking”\) mode disabled: left enabled, it consistently exhausts the1616K\-token budget on internal reasoning traces before emitting any parse\-able translation\.

```
Translate the following text from {src_lang} to {tgt_lang}.

### Terminology (use these exact translations for the following terms):
- "{source span 1}" -> "{target rendering 1}"
- "{source span 2}" -> "{target rendering 2}"
  ...

### Text to translate:
{src_text}

Keep any proper nouns, person names, IDs, email addresses, URLs, file names,
and technical identifiers unchanged.
When the text contains quoted references to app content
(conversation titles, email subjects, event titles),
use the exact translations from the terminology section above if provided.

Respond with ONLY the translated text, no explanations or comments.
```

Figure 3:The prompt supplied to the translator for every text span\.\{src\_lang\}and\{tgt\_lang\}are the source and target language and\{src\_text\}is the span to translate; theTerminologyblock is the per\-scenario term table \(§[3\.2](https://arxiv.org/html/2608.08775#S3.SS2)\) rendered inline \(placeholder entries shown; the pipeline joins each pair with a “→\\rightarrow” arrow\), present only when the scenario shares terms across surfaces\. Oracle tool\-call arguments are translated with a structurally identical variant that additionally conditions on the user task and returns a JSON object keyed by argument index\.### B\.1Translator choice

We evaluate the two top\-ranking open\-weight systems on BOUQuET for our target languages:Gemma\-4\-31B\-InstructandQwen\-3\.6\-27B, scoring translation across all ten languages from both systems against the reference translations under the protocol above\.[Table˜7](https://arxiv.org/html/2608.08775#A2.T7)reports the per\-language verdict shares and[Figure˜4](https://arxiv.org/html/2608.08775#A2.F4)visualises them\.

We find thatGemma\-4\-31B\-InstructdominatesQwen\-3\.6\-27Bon every one of the ten target languages: its reference\-win rate is lower and its candidate\-win rate higher on all ten\. The margin is narrowest on the closest\-to\-saturation Latin\-script targets \(a−6\-6pp reference\-win gap onspa\) and widens sharply on the non\-Latin and morphologically heavy targets \(−32\-32pp ontur,−29\-29pp onhin\), where surface\-fidelity errors are harder to recover from\.Gemma\-4\-31B\-Instructconsistently wins on all ten languages and therefore we adopt it as our main and only translation system\.

Table 7:Translator head\-to\-head: pooled verdict shares \(%,n=400n=400verdicts per cell\) forGemma\-4\-31B\-InstructandQwen\-3\.6\-27Bagainst the reference translations, per target language\.Gemma\-4\-31B\-Instructwins on every language;boldmarks the winner’s candidate\-win \(higher better\) and reference\-win \(lower better\) columns\.![Refer to caption](https://arxiv.org/html/2608.08775v1/figures/multilang_02_phaseA_stacked.png)Figure 4:Translator head\-to\-head\. Pooled400400\-verdict breakdown forGemma\-4\-31B\-Instruct\(solid bars\) andQwen\-3\.6\-27B\(hatched bars\) on each of the ten target languages, each translation judged pairwise against the reference translation\. The reference\-wins share \(top segment\) is smaller forGemma\-4\-31B\-Instructon every language, and the gap widens on the non\-Latin and morphologically heavy targets\.
### B\.2The reviewer step

We further investigate whether a two\-stage*translate\-then\-review*pipeline, in which a second model flags and post\-edits the translator’s output improvesOmnilingualGAIA2and find that it does not\.

Fixing the translator toGemma\-4\-31B\-Instruct, we consider a reviewer stage with three different configurations:GPT\-OSS\-120Bas a cross\-family reviewer on all ten languages,Gemma\-4\-31B\-Instructitself as a self\-review baseline on a three\-language subset \(spa,cmn,ind\), andQwen\-3\.6\-27Bas an alternate cross\-family reviewer on the same subset\. For each configuration we measure the change in candidate\-win rate against the reviewer\-freeGemma\-4\-31B\-Instruct\-only baseline,Δ=p^reviewed−p^baseline\\Delta=\\hat\{p\}\_\{\\text\{reviewed\}\}\-\\hat\{p\}\_\{\\text\{baseline\}\}\.

Every one of the sixteen reviewer configurations is statistically indistinguishable from the reviewer\-free baseline, since all the95%95\\%confidence intervals onΔ\\Deltacross zero \(see[Table˜8](https://arxiv.org/html/2608.08775#A2.T8)\)\. On the ten\-languageGPT\-OSS\-120Bsweep the smallest uncorrectedpp\-value is0\.470\.47; pooling all ten cells gives a meanΔ=\+1\.0\\Delta=\+1\.0pp with a95%95\\%CI of\[−1\.0,\+3\.0\]\[\-1\.0,\+3\.0\]pp\. We evaluate other two choices of reviewer model on a reduced sample of languages, and find that a different choice of the reviewer model does not significantly alter the outcome\.

Manual inspection of the edits confirms the statistical picture: the reviewer rewrites under9%9\\%of fields, with changes that are almost entirely stylistic \(gender inclusion, casing, idiom normalisation\) rather than error repairs\. We therefore commit to a translation pipeline forOmnilingualGAIA2with a single translator\-only pass, eliminating a second model pass that roughly doubles per\-scenario latency for no measurable quality gain\.

LangReviewerΔ\\Deltapp95%95\\%CIpp\-valuespaGPT\-OSS\-120B\+2\.25\+2\.25\[−4\.1,\+8\.6\]\[\-4\.1,\+8\.6\]0\.490\.49porGPT\-OSS\-120B\+1\.25\+1\.25\[−4\.3,\+6\.8\]\[\-4\.3,\+6\.8\]0\.660\.66fraGPT\-OSS\-120B\+1\.50\+1\.50\[−4\.3,\+7\.3\]\[\-4\.3,\+7\.3\]0\.610\.61itaGPT\-OSS\-120B\+2\.00\+2\.00\[−4\.2,\+8\.2\]\[\-4\.2,\+8\.2\]0\.530\.53deuGPT\-OSS\-120B\+1\.00\+1\.00\[−5\.0,\+7\.0\]\[\-5\.0,\+7\.0\]0\.740\.74hinGPT\-OSS\-120B\+0\.75\+0\.75\[−5\.8,\+7\.3\]\[\-5\.8,\+7\.3\]0\.820\.82cmnGPT\-OSS\-120B−0\.75\-0\.75\[−6\.6,\+5\.1\]\[\-6\.6,\+5\.1\]0\.800\.80jpnGPT\-OSS\-120B\+2\.25\+2\.25\[−4\.0,\+8\.5\]\[\-4\.0,\+8\.5\]0\.480\.48turGPT\-OSS\-120B\+2\.25\+2\.25\[−4\.3,\+8\.8\]\[\-4\.3,\+8\.8\]0\.500\.50indGPT\-OSS\-120B−2\.25\-2\.25\[−8\.4,\+3\.9\]\[\-8\.4,\+3\.9\]0\.470\.47spaGemma\-4\-31B\-Instruct\-self\+1\.25\+1\.25\[−5\.0,\+7\.5\]\[\-5\.0,\+7\.5\]0\.700\.70cmnGemma\-4\-31B\-Instruct\-self\+1\.75\+1\.75\[−4\.2,\+7\.7\]\[\-4\.2,\+7\.7\]0\.570\.57indGemma\-4\-31B\-Instruct\-self−1\.75\-1\.75\[−7\.9,\+4\.4\]\[\-7\.9,\+4\.4\]0\.580\.58spaQwen\-3\.6\-27B−0\.75\-0\.75\[−7\.0,\+5\.5\]\[\-7\.0,\+5\.5\]0\.810\.81cmnQwen\-3\.6\-27B\+0\.50\+0\.50\[−5\.4,\+6\.4\]\[\-5\.4,\+6\.4\]0\.870\.87indQwen\-3\.6\-27B−1\.50\-1\.50\[−7\.7,\+4\.7\]\[\-7\.7,\+4\.7\]0\.630\.63Table 8:Reviewer sweep: per\-\(language, reviewer\) change in candidate\-win rateΔ\\Deltaagainst the reviewer\-freeGemma\-4\-31B\-Instruct\-only baseline\. Every configuration’s confidence interval crosses zero: no reviewer configuration produces a statistically detectable change on any language\.
### B\.3Translator\-family confound

BecauseGemma\-4\-31B\-Instructserves both as the model in the translation pipeline \(§[3\.2](https://arxiv.org/html/2608.08775#S3.SS2)\) and as one of the evaluated agents—andQwen\-3\.6\-27Bis likewise evaluated—a natural concern is a*family\-match*confound: an agent might enjoy a home\-field advantage on data translated by its own model family, inflating its apparent capability\. We test this directly with a2×22\\times 2ablation crossing agent∈\{Gemma\-4\-31B\-Instruct,Qwen\-3\.6\-27B\}\\in\\\{\\textsc\{Gemma\-4\-31B\-Instruct\},\\ \\textsc\{Qwen\-3\.6\-27B\}\\\}with translator∈\{Gemma\-4\-31B\-Instruct\-MT,Qwen\-3\.6\-27B\-MT\}\\in\\\{\\text\{\{Gemma\-4\-31B\-Instruct\}\{\}\-MT\},\\ \\text\{\{Qwen\-3\.6\-27B\}\{\}\-MT\}\\\}, whereQwen\-3\.6\-27B\-MT re\-translates the identical scenarios through the same pipeline\. The design spans all ten target languages at1,9201\{,\}920rollouts per cell \(4040cells,262262K context\), and we form the difference\-in\-differences

DiD=\[GGemma\-4\-31B\-Instruct\-MT−GQwen\-3\.6\-27B\-MT\]−\[QGemma\-4\-31B\-Instruct\-MT−QQwen\-3\.6\-27B\-MT\],\\mathrm\{DiD\}=\\bigl\[G\_\{\\text\{\{Gemma\-4\-31B\-Instruct\}\{\}\-MT\}\}\-G\_\{\\text\{\{Qwen\-3\.6\-27B\}\{\}\-MT\}\}\\bigr\]\-\\bigl\[Q\_\{\\text\{\{Gemma\-4\-31B\-Instruct\}\{\}\-MT\}\}\-Q\_\{\\text\{\{Qwen\-3\.6\-27B\}\{\}\-MT\}\}\\bigr\],\(1\)whereGGandQQare theGemma\-4\-31B\-InstructandQwen\-3\.6\-27Bagent scores;DiD\>0\\mathrm\{DiD\}\>0would indicate aGemma\-4\-31B\-Instructhome\-field advantage\.

The agent gap is essentially invariant to which family translated the evaluation data \(Table[9](https://arxiv.org/html/2608.08775#A2.T9)\): the pooled DiD is\+0\.29\+0\.29pp\. Both agents lose comparably—4\.14\.1pp \(Gemma\-4\-31B\-Instruct\) and3\.83\.8pp \(Qwen\-3\.6\-27B\)—when the data switches fromGemma\-4\-31B\-Instruct\-MT toQwen\-3\.6\-27B\-MT, a symmetric*translator\-quality*effect \(Gemma\-4\-31B\-Instructis the stronger translator, App\.[B\.1](https://arxiv.org/html/2608.08775#A2.SS1)\) rather than a family\-of\-origin advantage\. To summarise across languages and to separate a genuine family\-match effect from the overall difference in translator quality, we fit an ordinary least\-squares \(OLS\) regression over the4040cell\-level observations \(44cells×\\times1010languages\)\. Writingyyfor a cell’s completion\-robustpass​@​1\\text\{pass\}@1\(the pass rate among graded rollouts,pass/\(pass\+fail\)\\mathrm\{pass\}/\(\\mathrm\{pass\}\+\\mathrm\{fail\}\)\), we regress

y=β0\+β1​A\+β2​T\+β3​\(A⋅T\)\+ε,y=\\beta\_\{0\}\+\\beta\_\{1\}\\,A\+\\beta\_\{2\}\\,T\+\\beta\_\{3\}\\,\(A\\cdot T\)\+\\varepsilon,\(2\)whereA=1A=1for theGemma\-4\-31B\-Instructagent \(and0forQwen\-3\.6\-27B\) andT=1T=1when the evaluation data isQwen\-3\.6\-27B\-translated \(and0forGemma\-4\-31B\-Instruct\-translated\)\. The coefficientsβ1\\beta\_\{1\}andβ2\\beta\_\{2\}are the two*main effects*—the average change inpass​@​1\\text\{pass\}@1from switching the agent, respectively the translator, with the other held fixed—andβ3\\beta\_\{3\}, the coefficient on the product term, is the*family\-match interaction*: the extra change when agent and translator switch together, i\.e\. how much more theGemma\-4\-31B\-Instructagent gains fromGemma\-4\-31B\-Instruct\-translated data than theQwen\-3\.6\-27Bagent does\. Under this codingβ3=−DiD\\beta\_\{3\}=\-\\mathrm\{DiD\}, so aGemma\-4\-31B\-Instructhome\-field advantage would show up asβ3<0\\beta\_\{3\}<0\.

We estimateβ3=−0\.37\\beta\_\{3\}=\-0\.37pp \(impliedDiD=\+0\.37\\mathrm\{DiD\}=\+0\.37pp\) with a standard error \(SE, the estimated sampling uncertainty of the coefficient\) of2\.92\.9pp; the correspondingtt\-statistict=β3/SE≈0\.13t=\\beta\_\{3\}/\\mathrm\{SE\}\\approx 0\.13is far below the\|t\|≈2\|t\|\\\!\\approx\\\!2required for significance at the5%5\\%level, so the interaction is statistically indistinguishable from zero\. Re\-fitting with an added covariate—each cell’s measured translator quality, taken as the per\-language candidate\-win rate of Table[7](https://arxiv.org/html/2608.08775#A2.T7)—leavesβ3\\beta\_\{3\}essentially unchanged, so the \(already negligible\) family\-match signal is not an artefact of translator quality varying across languages\. Complementing these pooled and regression estimates, every one of the ten*per\-language*DiD values has a95%95\\%confidence interval—from a scenario\-clustered bootstrap \(resampling whole scenarios with replacement,B=1000B=1000times, so that runs of the same scenario stay together\)—that crosses zero under both the completion\-robust and the raw metric \(Figure[5](https://arxiv.org/html/2608.08775#A2.F5)\); the largest per\-language effect \(tur,−3\.3\-3\.3pp on the robust metric\) is*Qwen*\-favoured, i\.e\. directionally opposite to aGemma\-4\-31B\-Instructhome\-field advantage\. We conclude the cross\-lingual gap decomposition \(§[6](https://arxiv.org/html/2608.08775#S6)\) is not confounded by translator family\.

#### Metric note\.

TheGemma\-4\-31B\-Instruct×\\timesQwen\-3\.6\-27B\-MT cells carry∼13%\{\\sim\}13\\%non\-graded rollouts \(agent timeouts and no\-verdict cases on long search trajectories at262262K context\), asymmetrically concentrated in that single arm\. We therefore take completion\-robustpass​@​1\\text\{pass\}@1=pass/\(pass\+fail\)\{\}=\\text\{pass\}/\(\\text\{pass\}\+\\text\{fail\}\)as the primary metric and report rawpass​@​1\\text\{pass\}@1as a robustness bound; the verdict is identical under both \(pooled DiD\+0\.29\+0\.29vs\+1\.18\+1\.18pp\), and the difference\-in\-differences subtracts out most of the completion\-rate asymmetry\.

Table 9:Family\-match2×22\\times 2: pooled completion\-robustpass​@​1\\text\{pass\}@1\(%\) over all ten target languages \(1,9201\{,\}920rollouts per cell\)\. Both agents drop∼4\{\\sim\}4pp onQwen\-3\.6\-27B\-translated data; the difference\-in\-differences \(Gemma\-4\-31B\-Instructhome\-field\) is\+0\.29\+0\.29pp\. Rows are the evaluated agent; columns are the translator that produced the evaluation data\.![Refer to caption](https://arxiv.org/html/2608.08775v1/figures/family_match_did.png)Figure 5:Per\-language difference\-in\-differences \(Gemma\-4\-31B\-Instructhome\-field advantage\), raw and completion\-robust, with95%95\\%scenario\-clustered bootstrap intervals \(B=1000B=1000\)\. All ten intervals cross zero under both metrics; positive values would indicate aGemma\-4\-31B\-Instructfamily\-match advantage\. The one language with the largest magnitude \(tur\) is Qwen\-favoured\.

## Appendix CJudge calibration

In this section we report additional details on the judge model and prompt ablation results described in §[3\.3](https://arxiv.org/html/2608.08775#S3.SS3.SSS0.Px2), on the originalGAIA2human\-annotated calibration set\.

### C\.1English calibration

The four configurations cross the two judge models, the paper’s referenceLlama\-3\.3\-70B\-Instructand the newGPT\-OSS\-120B, with the two prompt sets, the upstream default and our localised overrides: J1 \(Llama\-3\.3\-70B\-Instruct, default\), J2 \(GPT\-OSS\-120B, default\), J3 \(GPT\-OSS\-120B, localised\), and J4 \(Llama\-3\.3\-70B\-Instruct, localised\)\. Table[10](https://arxiv.org/html/2608.08775#A3.T10)reports overall Cohen’sκ\\kappa\(Cohen,[1960](https://arxiv.org/html/2608.08775#bib.bib7)\)against the human\-majority label on then=414n=414non\-inconclusive traces, with bootstrap95%95\\%confidence intervals\. J3 is the judge configuration we use when constructing the agent leaderboard in §[5](https://arxiv.org/html/2608.08775#S5)\.

Table 10:Overall judge agreement against the human\-majority label on then=414n=414non\-inconclusive traces of the recovered calibration set\. All fourκ\\kappavalues fall in\[0\.714,0\.732\]\[0\.714,0\.732\]with heavily overlapping CIs: on English, neither the model swap nor prompt localization materially changes agreement with humans\. J3 matches the paper\-standard reference J1\.Because all four configurations are scored on the same frozen traces, the two design choices can be isolated directly by judge\-vs\-judge agreement\. Holding the model fixed atLlama\-3\.3\-70B\-Instruct, localizing the prompts \(J1→\\rightarrowJ4\) flips the verdict on*one single*trace out of462462\(κ=0\.994\\kappa=0\.994,95%95\\%CI\[0\.978,1\.000\]\[0\.978,1\.000\]\); holding the prompts fixed at the localised set, swapping the model \(J4→\\rightarrowJ3\) flips two \(κ=0\.987\\kappa=0\.987,95%95\\%CI\[0\.967,1\.000\]\[0\.967,1\.000\]\); and the full swap from the paper’s reference \(J1→\\rightarrowJ3\) flips three \(κ=0\.981\\kappa=0\.981,95%95\\%CI\[0\.957,1\.000\]\[0\.957,1\.000\]\), every one in the direction of J3 accepting a trajectory the reference rejected\. On English, then, both the model choice and the prompt version are immaterial to judge–human agreement; the localisation payoff is realised off\-English \(§[3\.3](https://arxiv.org/html/2608.08775#S3.SS3.SSS0.Px2)\)\.

Table[11](https://arxiv.org/html/2608.08775#A3.T11)reports per\-capabilityκ\\kappaagainst the human\-majority label alongside the human inter\-annotator ceiling where multi\-annotator coverage exists on this set\.

Table 11:Per\-capability Cohen’sκ\\kappaagainst the human\-majority label under each judge configuration\. The Human\-IAA column is Cohen’sκ\\kappaon the multi\-annotator subset of the calibration set \(adaptabilityn=83n=83, timen=45n=45\); the other three capabilities are single\-annotator on this set, so no human ceiling can be computed\.Theκ=0\\kappa=0entries on*adaptability*and*time*appear in*all four*configurations, not only under J4\. On these capabilities the judge—like the human annotators– passes essentially no trace, so the label distribution has near zero variance and Cohen’sκ\\kappacollapses to0regardless of raw agreement \(which exceeds0\.870\.87throughout\)\. This is a property of the calibration corpus, not a judge weakness; see the failure\-taxonomy discussion below\.

Because the full\-corpusκ\\kappaof Table[10](https://arxiv.org/html/2608.08775#A3.T10)is depressed both by these degenerate capabilities and, more broadly, by the deterministic\-checker\-dominated traces on which every configuration returns the same structural verdict, we recompute the2×22\\times 2on the*LLM\-touched*slice—then=153n=153traces whose verdict actually invokes an LLM checker\.101010A trace is LLM\-touched if it passes, or if its first failing turn is decided by one of the natural\-language checkers \(message\_checker,user\_message\_checker,signature\_checker,tone\_checker,content\_checker\) or returns inconclusive;timeis excluded, as in the multilingual slice\.This is the*same*frozen slice used for the multilingual comparison \(Table[14](https://arxiv.org/html/2608.08775#A3.T14)\), so the English and cross\-lingual numbers are directly comparable\. On it \(Table[12](https://arxiv.org/html/2608.08775#A3.T12)\) every configuration rises toκ≥0\.928\\kappa\\geq 0\.928, against0\.7140\.714–0\.7320\.732on the full corpus: the operational judge is strongest precisely where the LLM is exercised, and the full\-corpus figures understate rather than overstate its reliability\. The2×22\\times 2ordering is preserved—localised≥\\geqdefault on each model, andGPT\-OSS\-120B≥\\geqLlama\-3\.3\-70B\-Instructunder each prompt set—though all four CIs overlap, so the differences are directional, not significant\.

Table 12:The2×22\\times 2recomputed on the LLM\-touched slice \(n=153n=153; the same frozen traces as the multilingual Table[14](https://arxiv.org/html/2608.08775#A3.T14)\)\. Removing the deterministic\-checker\-dominated and degenerate\-κ\\kappacapabilities lifts every configuration toκ≥0\.928\\kappa\\geq 0\.928, versus0\.7140\.714–0\.7320\.732on the full corpus \(Table[10](https://arxiv.org/html/2608.08775#A3.T10)\)\. Localising the prompts on the fixedLlama\-3\.3\-70B\-Instructmodel \(J1→\\rightarrowJ4\) still moves the verdict on essentially no trace \(κ=0\.986\\kappa=0\.986judge\-vs\-judge on this slice\)\.
### C\.2Failure taxonomy on the calibration set

The calibration set is heavily deterministic\-checker\-dominated on three of the four capabilities: on adaptability, ambiguity and execution the verdict is decided by structured checks that never exercise the LLMaaJ component\. Table[13](https://arxiv.org/html/2608.08775#A3.T13)reports the failure distribution per capability\.

Table 13:Verdict provenance on the English calibration set, computed on the operational judge J3 \(not the human labels\)\.*Success*counts traces J3 passes;*Det\. fail*counts J3 failures whose first failing check is deterministic \(tool\- or SMU\-count, stuck\-loop or timeout, or a deterministic content checker\);*LLM fail*counts J3 failures whose first failing check is an LLM checker \(message,user\_message,signature,tone, orcontent\)\. The three counts sum tonnper capability\. Only*search*exercises the LLMaaJ component meaningfully\.
### C\.3Multilingual robustness

Table[14](https://arxiv.org/html/2608.08775#A3.T14)restricts the calibration set to then=153n=153traces where the LLM judge actually fires and reports per\-languageκ\\kappaunder J3 forengand all ten target languages\. This is the headline number of §[3\.3](https://arxiv.org/html/2608.08775#S3.SS3.SSS0.Px2): it is the only slice on which translation of the natural\-language content can plausibly move the verifier verdict\. Because the English human\-majority label is held fixed across languages \(a faithful translation should not change the correct verdict\),κ\\kappahere isolates whether translation alone perturbs the judge\.111111In roughly2%2\\%of translated traces \(7676of3,6903\{,\}690\) theGemma\-4\-31B\-Instructtranslator regurgitated part of its own system prompt into the oracle user message, concentrated on the*search*capability whose user messages are longest\. Manual review of the7272affected traces that fall in the LLM\-touched slice found no verdict divergence from English on any of them, so the calibration figures reported here are unaffected\.The eight languages beyondeng,jpnandspawere scored on a freshly provisionedGPT\-OSS\-120Bendpoint with identicalgpt\-oss\-120bweights andomnigaia\_v2prompts, differing only in serving host\.

Table 14:J3 agreement with human\-majority labels on the LLM\-touched slice of the calibration set, for all eleven in\-scope languages\. The human pass rate on this slice is0\.6540\.654in every language: the English human\-majority label is translation\-invariant, soκ\\kappaisolates translation\-induced verdict drift\. EveryΔ​κ\\Delta\\kappasits inside the pre\-registeredΔ​κ≤0\.10\\Delta\\kappa\\leq 0\.10hazard threshold; the largest drop isturat−0\.085\-0\.085, and the non\-Latin scripts \(cmn,jpn,hin\) rank among the strongest, so script family does not predict translation drift\.Per\-capability multilingualκ\\kappais reported only for*search*\(Table[15](https://arxiv.org/html/2608.08775#A3.T15)\)\. We deliberately omit per\-capabilityκ\\kappafor adaptability, ambiguity and execution under translation: on those subsets the verdict is deterministic \(see Table[13](https://arxiv.org/html/2608.08775#A3.T13)\) andκ\\kappais language\-invariant by construction, which would document a corpus property rather than judge quality\.

Table 15:J3 per\-language agreement on the*search*capability \(n=97n=97\), the one capability with a substantial LLM\-touched share \(43/9743/97in English\)\. Agreement holds atκ≥0\.958\\kappa\\geq 0\.958across all eleven languages;eng,jpnandcmnare invariant, and the largest loss is two verdicts \(spa,deu\)\.As a same\-family bias check on the operational judge, we additionally re\-scored a held\-out set ofGPT\-OSS\-120Brollouts—used solely for this judge self\-preference check, asGPT\-OSS\-120Bis not part of the evaluated agent leaderboard—with both J3 \(self\) and the cross\-family reference J1: the self−\-cross overallpass​@​1\\text\{pass\}@1difference is negative in all eight complete languages \(Δ∈\[−0\.025,−0\.014\]\\Delta\\in\[\-0\.025,\-0\.014\]\), soGPT\-OSS\-120Bshows no self\-preference when adjudicating its own family’s rollouts—if anything it is marginally stricter than the reference\.

## Appendix DPer\-language results

This appendix expands on the results reported in the main table \(Table[1](https://arxiv.org/html/2608.08775#S5.T1)\) and the heatmap \(Figure[2](https://arxiv.org/html/2608.08775#S5.F2)\) into the full per\-agent, per\-language, per\-capability numbers\. For each \(agent, language, capability\) cell we reportpass​@​1\\text\{pass\}@1andpass​@​3\\text\{pass\}@3with Wilson95%95\\%confidence intervals\(Wilson,[1927](https://arxiv.org/html/2608.08775#bib.bib55)\), the all\-three\-consistency metricpass3all\\text\{pass\}^\{\\text\{all\}\}\_\{3\}\. TableLABEL:tab:breakdownreports all four capabilities together, with the four capabilities laid out across the columns\.

Table 16:Full per\-agent, per\-language breakdown across the four reference\-oracle capabilities\. Each capability reportspass​@​1\\text\{pass\}@1/pass​@​3\\text\{pass\}@3/pass3all\\text\{pass\}^\{\\text\{all\}\}\_\{3\}\(%\);n=160n\{=\}160scenarios per cell\.pass​@​1\\text\{pass\}@1is the mean over*valid*runs \(agent engaged,num\_agent\_events\>\>3; infra non\-attempts excluded\),pass​@​3\\text\{pass\}@3is any\-of\-three over valid runs, and ‘all’ ispass3all\\text\{pass\}^\{\\text\{all\}\}\_\{3\}\(all three valid runs pass\)\. Sub/superscripts onpass​@​1\\text\{pass\}@1/pass​@​3\\text\{pass\}@3are Wilson95%95\\%CI bounds\.Exec\.SearchAdapt\.Ambig\.AgentLang@1@3all@1@3all@1@3all@1@3allClaude\-4\.7\-Opuseng79\.48376\{\}\_\{76\}^\{83\}88\.19282\{\}\_\{82\}^\{92\}69\.484\.08780\{\}\_\{80\}^\{87\}95\.69891\{\}\_\{91\}^\{98\}67\.763\.96859\{\}\_\{59\}^\{68\}75\.68268\{\}\_\{68\}^\{82\}49\.458\.06253\{\}\_\{53\}^\{62\}71\.67864\{\}\_\{64\}^\{78\}43\.2spa70\.47466\{\}\_\{66\}^\{74\}84\.48978\{\}\_\{78\}^\{89\}54\.474\.67870\{\}\_\{70\}^\{78\}89\.39384\{\}\_\{84\}^\{93\}57\.961\.86657\{\}\_\{57\}^\{66\}75\.68268\{\}\_\{68\}^\{82\}43\.849\.35445\{\}\_\{45\}^\{54\}68\.87661\{\}\_\{61\}^\{76\}29\.9deu65\.87061\{\}\_\{61\}^\{70\}79\.48572\{\}\_\{72\}^\{85\}48\.158\.46354\{\}\_\{54\}^\{63\}78\.08471\{\}\_\{71\}^\{84\}34\.658\.06353\{\}\_\{53\}^\{63\}70\.07762\{\}\_\{62\}^\{77\}47\.542\.24738\{\}\_\{38\}^\{47\}58\.36650\{\}\_\{50\}^\{66\}28\.2fra66\.77162\{\}\_\{62\}^\{71\}83\.18877\{\}\_\{77\}^\{88\}47\.565\.17060\{\}\_\{60\}^\{70\}81\.18674\{\}\_\{74\}^\{86\}50\.360\.56556\{\}\_\{56\}^\{65\}72\.57965\{\}\_\{65\}^\{79\}46\.242\.04738\{\}\_\{38\}^\{47\}61\.16853\{\}\_\{53\}^\{68\}24\.8ita59\.86455\{\}\_\{55\}^\{64\}76\.28269\{\}\_\{69\}^\{82\}41\.269\.07364\{\}\_\{64\}^\{73\}81\.08674\{\}\_\{74\}^\{86\}55\.762\.26658\{\}\_\{58\}^\{66\}74\.48167\{\}\_\{67\}^\{81\}48\.835\.34031\{\}\_\{31\}^\{40\}54\.16246\{\}\_\{46\}^\{62\}18\.5por64\.97060\{\}\_\{60\}^\{70\}81\.98775\{\}\_\{75\}^\{87\}46\.270\.67566\{\}\_\{66\}^\{75\}84\.99079\{\}\_\{79\}^\{90\}52\.254\.95950\{\}\_\{50\}^\{59\}73\.88066\{\}\_\{66\}^\{80\}35\.041\.34637\{\}\_\{37\}^\{46\}60\.36852\{\}\_\{52\}^\{68\}23\.7ind69\.57365\{\}\_\{65\}^\{73\}83\.18877\{\}\_\{77\}^\{88\}55\.669\.47365\{\}\_\{65\}^\{73\}86\.29180\{\}\_\{80\}^\{91\}50\.361\.66657\{\}\_\{57\}^\{66\}73\.88066\{\}\_\{66\}^\{80\}45\.644\.04940\{\}\_\{40\}^\{49\}60\.96853\{\}\_\{53\}^\{68\}26\.9tur61\.96757\{\}\_\{57\}^\{67\}77\.58370\{\}\_\{70\}^\{83\}45\.668\.37264\{\}\_\{64\}^\{72\}83\.08876\{\}\_\{76\}^\{88\}52\.252\.05648\{\}\_\{48\}^\{56\}72\.57965\{\}\_\{65\}^\{79\}31\.942\.44738\{\}\_\{38\}^\{47\}66\.77459\{\}\_\{59\}^\{74\}19\.2cmn62\.46758\{\}\_\{58\}^\{67\}75\.08168\{\}\_\{68\}^\{81\}49\.466\.07062\{\}\_\{62\}^\{70\}81\.88775\{\}\_\{75\}^\{87\}47\.256\.76152\{\}\_\{52\}^\{61\}73\.17966\{\}\_\{66\}^\{79\}38\.140\.54536\{\}\_\{36\}^\{45\}55\.86348\{\}\_\{48\}^\{63\}28\.8jpn56\.16152\{\}\_\{52\}^\{61\}70\.07762\{\}\_\{62\}^\{77\}41\.265\.36961\{\}\_\{61\}^\{69\}83\.08876\{\}\_\{76\}^\{88\}47\.855\.36051\{\}\_\{51\}^\{60\}73\.07966\{\}\_\{66\}^\{79\}33\.335\.94032\{\}\_\{32\}^\{40\}51\.96044\{\}\_\{44\}^\{60\}19\.0hin53\.75849\{\}\_\{49\}^\{58\}77\.58370\{\}\_\{70\}^\{83\}30\.063\.86859\{\}\_\{59\}^\{68\}81\.88775\{\}\_\{75\}^\{87\}44\.054\.86050\{\}\_\{50\}^\{60\}71\.97864\{\}\_\{64\}^\{78\}36\.240\.54536\{\}\_\{36\}^\{45\}59\.26751\{\}\_\{51\}^\{67\}21\.0GPT\-5\.4eng63\.06759\{\}\_\{59\}^\{67\}81\.28774\{\}\_\{74\}^\{87\}44\.482\.18578\{\}\_\{78\}^\{85\}92\.49687\{\}\_\{87\}^\{96\}70\.139\.24435\{\}\_\{35\}^\{44\}58\.86651\{\}\_\{51\}^\{66\}21\.925\.63022\{\}\_\{22\}^\{30\}41\.44934\{\}\_\{34\}^\{49\}12\.7spa51\.95647\{\}\_\{47\}^\{56\}69\.47662\{\}\_\{62\}^\{76\}33\.174\.87971\{\}\_\{71\}^\{79\}87\.39281\{\}\_\{81\}^\{92\}61\.834\.83931\{\}\_\{31\}^\{39\}51\.25944\{\}\_\{44\}^\{59\}18\.119\.02316\{\}\_\{16\}^\{23\}33\.34126\{\}\_\{26\}^\{41\}8\.2deu52\.55748\{\}\_\{48\}^\{57\}68\.87561\{\}\_\{61\}^\{75\}32\.572\.17668\{\}\_\{68\}^\{76\}85\.49079\{\}\_\{79\}^\{90\}56\.136\.04032\{\}\_\{32\}^\{40\}53\.86146\{\}\_\{46\}^\{61\}19\.421\.42518\{\}\_\{18\}^\{25\}35\.44328\{\}\_\{28\}^\{43\}8\.9fra49\.95445\{\}\_\{45\}^\{54\}65\.07257\{\}\_\{57\}^\{72\}31\.975\.17971\{\}\_\{71\}^\{79\}88\.69383\{\}\_\{83\}^\{93\}61\.435\.64031\{\}\_\{31\}^\{40\}52\.56045\{\}\_\{45\}^\{60\}17\.521\.12518\{\}\_\{18\}^\{25\}34\.44227\{\}\_\{27\}^\{42\}9\.6ita49\.85445\{\}\_\{45\}^\{54\}70\.07762\{\}\_\{62\}^\{77\}28\.773\.67769\{\}\_\{69\}^\{77\}86\.79181\{\}\_\{81\}^\{91\}57\.638\.54334\{\}\_\{34\}^\{43\}57\.56550\{\}\_\{50\}^\{65\}18\.817\.32114\{\}\_\{14\}^\{21\}29\.13723\{\}\_\{23\}^\{37\}7\.6por53\.75849\{\}\_\{49\}^\{58\}73\.88066\{\}\_\{66\}^\{80\}34\.471\.47567\{\}\_\{67\}^\{75\}85\.39079\{\}\_\{79\}^\{90\}56\.434\.53930\{\}\_\{30\}^\{39\}51\.25944\{\}\_\{44\}^\{59\}17\.519\.62316\{\}\_\{16\}^\{23\}33\.14126\{\}\_\{26\}^\{41\}8\.9ind52\.85748\{\}\_\{48\}^\{57\}70\.67763\{\}\_\{63\}^\{77\}30\.671\.97668\{\}\_\{68\}^\{76\}84\.68978\{\}\_\{78\}^\{89\}55\.835\.03931\{\}\_\{31\}^\{39\}54\.46247\{\}\_\{47\}^\{62\}16\.919\.72416\{\}\_\{16\}^\{24\}32\.54026\{\}\_\{26\}^\{40\}9\.6tur53\.15849\{\}\_\{49\}^\{58\}73\.88066\{\}\_\{66\}^\{80\}35\.066\.87162\{\}\_\{62\}^\{71\}78\.58471\{\}\_\{71\}^\{84\}52\.535\.64031\{\}\_\{31\}^\{40\}52\.56045\{\}\_\{45\}^\{60\}19\.420\.92517\{\}\_\{17\}^\{25\}34\.64228\{\}\_\{28\}^\{42\}10\.7cmn41\.84637\{\}\_\{37\}^\{46\}57\.26549\{\}\_\{49\}^\{65\}24\.565\.57061\{\}\_\{61\}^\{70\}80\.38673\{\}\_\{73\}^\{86\}48\.433\.33829\{\}\_\{29\}^\{38\}48\.85641\{\}\_\{41\}^\{56\}17\.518\.22215\{\}\_\{15\}^\{22\}29\.93823\{\}\_\{23\}^\{38\}10\.2jpn40\.54536\{\}\_\{36\}^\{45\}58\.16550\{\}\_\{50\}^\{65\}24\.461\.16557\{\}\_\{57\}^\{65\}77\.68370\{\}\_\{70\}^\{83\}42\.324\.92921\{\}\_\{21\}^\{29\}43\.15136\{\}\_\{36\}^\{51\}10\.015\.61913\{\}\_\{13\}^\{19\}26\.83420\{\}\_\{20\}^\{34\}8\.3hin45\.45041\{\}\_\{41\}^\{50\}64\.27156\{\}\_\{56\}^\{71\}24\.566\.07062\{\}\_\{62\}^\{70\}80\.88674\{\}\_\{74\}^\{86\}53\.231\.23627\{\}\_\{27\}^\{36\}49\.45742\{\}\_\{42\}^\{57\}14\.416\.42013\{\}\_\{13\}^\{20\}29\.13723\{\}\_\{23\}^\{37\}5\.7Gemini\-3\.1\-Proeng48\.45344\{\}\_\{44\}^\{53\}66\.27359\{\}\_\{59\}^\{73\}28\.867\.67263\{\}\_\{63\}^\{72\}86\.69180\{\}\_\{80\}^\{91\}47\.131\.23627\{\}\_\{27\}^\{36\}51\.25944\{\}\_\{44\}^\{59\}12\.526\.43123\{\}\_\{23\}^\{31\}43\.75136\{\}\_\{36\}^\{51\}12\.0spa37\.94234\{\}\_\{34\}^\{42\}56\.26449\{\}\_\{49\}^\{64\}20\.062\.46758\{\}\_\{58\}^\{67\}80\.58674\{\}\_\{74\}^\{86\}43\.429\.63426\{\}\_\{26\}^\{34\}46\.25439\{\}\_\{39\}^\{54\}13\.121\.22518\{\}\_\{18\}^\{25\}34\.84328\{\}\_\{28\}^\{43\}10\.1deu38\.44334\{\}\_\{34\}^\{43\}59\.46752\{\}\_\{52\}^\{67\}19\.458\.86354\{\}\_\{54\}^\{63\}79\.78573\{\}\_\{73\}^\{85\}38\.631\.23627\{\}\_\{27\}^\{36\}50\.65843\{\}\_\{43\}^\{58\}12\.520\.82517\{\}\_\{17\}^\{25\}37\.34530\{\}\_\{30\}^\{45\}6\.3fra40\.14536\{\}\_\{36\}^\{45\}55\.66348\{\}\_\{48\}^\{63\}25\.660\.06456\{\}\_\{56\}^\{64\}77\.28370\{\}\_\{70\}^\{83\}43\.031\.03527\{\}\_\{27\}^\{35\}51\.25944\{\}\_\{44\}^\{59\}11\.918\.82315\{\}\_\{15\}^\{23\}30\.23824\{\}\_\{24\}^\{38\}8\.2ita36\.04032\{\}\_\{32\}^\{40\}53\.86146\{\}\_\{46\}^\{61\}18\.160\.16456\{\}\_\{56\}^\{64\}79\.18572\{\}\_\{72\}^\{85\}37\.329\.83426\{\}\_\{26\}^\{34\}51\.25944\{\}\_\{44\}^\{59\}10\.018\.92316\{\}\_\{16\}^\{23\}32\.14025\{\}\_\{25\}^\{40\}8\.2por39\.24435\{\}\_\{35\}^\{44\}60\.06752\{\}\_\{52\}^\{67\}20\.659\.76455\{\}\_\{55\}^\{64\}77\.48370\{\}\_\{70\}^\{83\}40\.930\.83527\{\}\_\{27\}^\{35\}53\.16145\{\}\_\{45\}^\{61\}11\.218\.72215\{\}\_\{15\}^\{22\}31\.63925\{\}\_\{25\}^\{39\}8\.9ind38\.24334\{\}\_\{34\}^\{43\}58\.86651\{\}\_\{51\}^\{66\}18\.862\.26658\{\}\_\{58\}^\{66\}80\.58674\{\}\_\{74\}^\{86\}42\.131\.53627\{\}\_\{27\}^\{36\}48\.85641\{\}\_\{41\}^\{56\}16\.916\.52013\{\}\_\{13\}^\{20\}28\.93622\{\}\_\{22\}^\{36\}7\.5tur37\.74233\{\}\_\{33\}^\{42\}54\.46247\{\}\_\{47\}^\{62\}21\.953\.25849\{\}\_\{49\}^\{58\}75\.88269\{\}\_\{69\}^\{82\}31\.225\.53022\{\}\_\{22\}^\{30\}46\.25439\{\}\_\{39\}^\{54\}8\.817\.02114\{\}\_\{14\}^\{21\}28\.53622\{\}\_\{22\}^\{36\}9\.5cmn41\.84637\{\}\_\{37\}^\{46\}60\.66853\{\}\_\{53\}^\{68\}23\.156\.96152\{\}\_\{52\}^\{61\}77\.28370\{\}\_\{70\}^\{83\}34\.828\.63325\{\}\_\{25\}^\{33\}50\.65843\{\}\_\{43\}^\{58\}7\.514\.91812\{\}\_\{12\}^\{18\}24\.13118\{\}\_\{18\}^\{31\}8\.2jpn36\.34132\{\}\_\{32\}^\{41\}53\.16145\{\}\_\{45\}^\{61\}19\.454\.95950\{\}\_\{50\}^\{59\}74\.88168\{\}\_\{68\}^\{81\}34\.624\.82921\{\}\_\{21\}^\{29\}43\.15136\{\}\_\{36\}^\{51\}8\.116\.52013\{\}\_\{13\}^\{20\}26\.63420\{\}\_\{20\}^\{34\}7\.0hin31\.93628\{\}\_\{28\}^\{36\}49\.45742\{\}\_\{42\}^\{57\}16\.958\.46354\{\}\_\{54\}^\{63\}79\.08572\{\}\_\{72\}^\{85\}36\.925\.02921\{\}\_\{21\}^\{29\}47\.55540\{\}\_\{40\}^\{55\}8\.114\.31811\{\}\_\{11\}^\{18\}24\.53218\{\}\_\{18\}^\{32\}6\.3Kimi\-2\.6eng51\.75647\{\}\_\{47\}^\{56\}73\.88066\{\}\_\{66\}^\{80\}28\.781\.58578\{\}\_\{78\}^\{85\}94\.39790\{\}\_\{90\}^\{97\}63\.935\.74032\{\}\_\{32\}^\{40\}55\.06347\{\}\_\{47\}^\{63\}18\.825\.53022\{\}\_\{22\}^\{30\}42\.85135\{\}\_\{35\}^\{51\}9\.4spa36\.94133\{\}\_\{33\}^\{41\}57\.56550\{\}\_\{50\}^\{65\}18\.173\.47769\{\}\_\{69\}^\{77\}91\.29586\{\}\_\{86\}^\{95\}53\.531\.53628\{\}\_\{28\}^\{36\}48\.15641\{\}\_\{41\}^\{56\}17\.520\.52417\{\}\_\{17\}^\{24\}36\.74430\{\}\_\{30\}^\{44\}8\.2deu37\.14233\{\}\_\{33\}^\{42\}55\.66348\{\}\_\{48\}^\{63\}19\.464\.96960\{\}\_\{60\}^\{69\}82\.38776\{\}\_\{76\}^\{87\}43\.031\.43627\{\}\_\{27\}^\{36\}48\.15641\{\}\_\{41\}^\{56\}15\.617\.32114\{\}\_\{14\}^\{21\}28\.53622\{\}\_\{22\}^\{36\}7\.6fra36\.94133\{\}\_\{33\}^\{41\}56\.26449\{\}\_\{49\}^\{64\}18\.870\.97567\{\}\_\{67\}^\{75\}85\.59079\{\}\_\{79\}^\{90\}51\.631\.53628\{\}\_\{28\}^\{36\}46\.95539\{\}\_\{39\}^\{55\}17\.521\.32518\{\}\_\{18\}^\{25\}35\.04328\{\}\_\{28\}^\{43\}7\.6ita38\.34334\{\}\_\{34\}^\{43\}58\.16550\{\}\_\{50\}^\{65\}18\.169\.97466\{\}\_\{66\}^\{74\}87\.39281\{\}\_\{81\}^\{92\}48\.731\.93628\{\}\_\{28\}^\{36\}48\.85641\{\}\_\{41\}^\{56\}15\.019\.52316\{\}\_\{16\}^\{23\}30\.63824\{\}\_\{24\}^\{38\}10\.0por38\.54334\{\}\_\{34\}^\{43\}59\.46752\{\}\_\{52\}^\{67\}20\.672\.37668\{\}\_\{68\}^\{76\}88\.19282\{\}\_\{82\}^\{92\}53\.526\.13022\{\}\_\{22\}^\{30\}42\.55035\{\}\_\{35\}^\{50\}13\.816\.82014\{\}\_\{14\}^\{20\}28\.73622\{\}\_\{22\}^\{36\}6\.9ind33\.83830\{\}\_\{30\}^\{38\}50\.65843\{\}\_\{43\}^\{58\}16\.270\.77566\{\}\_\{66\}^\{75\}89\.99484\{\}\_\{84\}^\{94\}46\.225\.53022\{\}\_\{22\}^\{30\}41\.24934\{\}\_\{34\}^\{49\}10\.616\.12013\{\}\_\{13\}^\{20\}29\.43723\{\}\_\{23\}^\{37\}5\.0tur36\.74132\{\}\_\{32\}^\{41\}62\.57055\{\}\_\{55\}^\{70\}13\.863\.16759\{\}\_\{59\}^\{67\}79\.08572\{\}\_\{72\}^\{85\}44\.626\.73123\{\}\_\{23\}^\{31\}45\.65338\{\}\_\{38\}^\{53\}9\.417\.22114\{\}\_\{14\}^\{21\}30\.83824\{\}\_\{24\}^\{38\}5\.0cmn35\.03931\{\}\_\{31\}^\{39\}53\.16145\{\}\_\{45\}^\{61\}15\.664\.66960\{\}\_\{60\}^\{69\}83\.08876\{\}\_\{76\}^\{88\}45\.326\.33023\{\}\_\{23\}^\{30\}40\.04833\{\}\_\{33\}^\{48\}13\.821\.02518\{\}\_\{18\}^\{25\}34\.04227\{\}\_\{27\}^\{42\}8\.8jpn35\.44031\{\}\_\{31\}^\{40\}53\.86146\{\}\_\{46\}^\{61\}16\.263\.46859\{\}\_\{59\}^\{68\}79\.68573\{\}\_\{73\}^\{85\}43\.925\.93022\{\}\_\{22\}^\{30\}43\.15136\{\}\_\{36\}^\{51\}10\.616\.72014\{\}\_\{14\}^\{20\}30\.43824\{\}\_\{24\}^\{38\}5\.7hin33\.83830\{\}\_\{30\}^\{38\}58\.16550\{\}\_\{50\}^\{65\}15\.061\.96657\{\}\_\{57\}^\{66\}82\.48876\{\}\_\{76\}^\{88\}39\.024\.02820\{\}\_\{20\}^\{28\}40\.34833\{\}\_\{33\}^\{48\}10\.114\.11811\{\}\_\{11\}^\{18\}26\.93421\{\}\_\{21\}^\{34\}5\.6Gemma\-4\-31B\-Instructeng50\.35546\{\}\_\{46\}^\{55\}67\.57460\{\}\_\{60\}^\{74\}34\.470\.97567\{\}\_\{67\}^\{75\}85\.69079\{\}\_\{79\}^\{90\}50\.045\.25041\{\}\_\{41\}^\{50\}63\.17055\{\}\_\{55\}^\{70\}26\.223\.82820\{\}\_\{20\}^\{28\}36\.94530\{\}\_\{30\}^\{45\}10\.6spa41\.94638\{\}\_\{38\}^\{46\}56\.96449\{\}\_\{49\}^\{64\}25\.663\.66859\{\}\_\{59\}^\{68\}78\.18471\{\}\_\{71\}^\{84\}41\.239\.04335\{\}\_\{35\}^\{43\}57\.56550\{\}\_\{50\}^\{65\}23\.118\.32215\{\}\_\{15\}^\{22\}28\.13622\{\}\_\{22\}^\{36\}8\.8deu45\.65041\{\}\_\{41\}^\{50\}58\.86651\{\}\_\{51\}^\{66\}32\.559\.16455\{\}\_\{55\}^\{64\}76\.28269\{\}\_\{69\}^\{82\}40\.037\.34233\{\}\_\{33\}^\{42\}57\.56550\{\}\_\{50\}^\{65\}19\.418\.52215\{\}\_\{15\}^\{22\}31\.93925\{\}\_\{25\}^\{39\}6\.9fra42\.14738\{\}\_\{38\}^\{47\}59\.46752\{\}\_\{52\}^\{67\}23\.862\.26758\{\}\_\{58\}^\{67\}75\.68268\{\}\_\{68\}^\{82\}48\.141\.24637\{\}\_\{37\}^\{46\}56\.26449\{\}\_\{49\}^\{64\}27\.518\.92316\{\}\_\{16\}^\{23\}31\.23925\{\}\_\{25\}^\{39\}8\.8ita43\.44839\{\}\_\{39\}^\{48\}60\.06752\{\}\_\{52\}^\{67\}28\.863\.66859\{\}\_\{59\}^\{68\}76\.28269\{\}\_\{69\}^\{82\}42\.541\.84638\{\}\_\{38\}^\{46\}58\.16550\{\}\_\{50\}^\{65\}26\.916\.82014\{\}\_\{14\}^\{20\}28\.13622\{\}\_\{22\}^\{36\}8\.1por47\.35243\{\}\_\{43\}^\{52\}62\.57055\{\}\_\{55\}^\{70\}31\.260\.86556\{\}\_\{56\}^\{65\}71\.27864\{\}\_\{64\}^\{78\}40\.635\.54031\{\}\_\{31\}^\{40\}50\.65843\{\}\_\{43\}^\{58\}20\.018\.62215\{\}\_\{15\}^\{22\}28\.13622\{\}\_\{22\}^\{36\}8\.8ind45\.25041\{\}\_\{41\}^\{50\}61\.26854\{\}\_\{54\}^\{68\}27\.560\.26556\{\}\_\{56\}^\{65\}75\.08168\{\}\_\{68\}^\{81\}37\.540\.44536\{\}\_\{36\}^\{45\}59\.46752\{\}\_\{52\}^\{67\}20\.618\.52215\{\}\_\{15\}^\{22\}31\.23925\{\}\_\{25\}^\{39\}8\.1tur40\.54536\{\}\_\{36\}^\{45\}56\.96449\{\}\_\{49\}^\{64\}25\.052\.55748\{\}\_\{48\}^\{57\}65\.07257\{\}\_\{57\}^\{72\}36\.935\.84032\{\}\_\{32\}^\{40\}50\.65843\{\}\_\{43\}^\{58\}20\.618\.32215\{\}\_\{15\}^\{22\}30\.63824\{\}\_\{24\}^\{38\}6\.9cmn35\.24031\{\}\_\{31\}^\{40\}49\.45742\{\}\_\{42\}^\{57\}20\.046\.35142\{\}\_\{42\}^\{51\}63\.17055\{\}\_\{55\}^\{70\}26\.236\.24132\{\}\_\{32\}^\{41\}48\.15641\{\}\_\{41\}^\{56\}23\.118\.42215\{\}\_\{15\}^\{22\}25\.63319\{\}\_\{19\}^\{33\}10\.0jpn31\.63628\{\}\_\{28\}^\{36\}41\.95035\{\}\_\{35\}^\{50\}21\.242\.34738\{\}\_\{38\}^\{47\}59\.46752\{\}\_\{52\}^\{67\}21\.230\.53527\{\}\_\{27\}^\{35\}47\.55540\{\}\_\{40\}^\{55\}11\.913\.11610\{\}\_\{10\}^\{16\}20\.02715\{\}\_\{15\}^\{27\}7\.5hin28\.53325\{\}\_\{25\}^\{33\}45\.05337\{\}\_\{37\}^\{53\}14\.439\.94435\{\}\_\{35\}^\{44\}55\.06347\{\}\_\{47\}^\{63\}20\.633\.73830\{\}\_\{30\}^\{38\}48\.85641\{\}\_\{41\}^\{56\}18\.113\.01610\{\}\_\{10\}^\{16\}20\.02715\{\}\_\{15\}^\{27\}5\.6Qwen\-3\.6\-35B\-A3Beng49\.05345\{\}\_\{45\}^\{53\}68\.87561\{\}\_\{61\}^\{75\}28\.163\.26759\{\}\_\{59\}^\{67\}86\.29180\{\}\_\{80\}^\{91\}34\.432\.13628\{\}\_\{28\}^\{36\}48\.15641\{\}\_\{41\}^\{56\}16\.211\.7159\{\}\_\{9\}^\{15\}23\.13017\{\}\_\{17\}^\{30\}3\.1spa26\.53123\{\}\_\{23\}^\{31\}46\.25439\{\}\_\{39\}^\{54\}10\.641\.44637\{\}\_\{37\}^\{46\}69\.47662\{\}\_\{62\}^\{76\}16\.922\.32619\{\}\_\{19\}^\{26\}37\.54530\{\}\_\{30\}^\{45\}5\.66\.495\{\}\_\{5\}^\{9\}14\.42110\{\}\_\{10\}^\{21\}1\.2deu28\.03224\{\}\_\{24\}^\{32\}48\.15641\{\}\_\{41\}^\{56\}11\.236\.84133\{\}\_\{33\}^\{41\}65\.67358\{\}\_\{58\}^\{73\}10\.017\.62114\{\}\_\{14\}^\{21\}36\.94530\{\}\_\{30\}^\{45\}2\.56\.895\{\}\_\{5\}^\{9\}11\.9188\{\}\_\{8\}^\{18\}3\.1fra27\.83224\{\}\_\{24\}^\{32\}44\.45237\{\}\_\{37\}^\{52\}10\.644\.74940\{\}\_\{40\}^\{49\}71\.27864\{\}\_\{64\}^\{78\}16\.920\.32417\{\}\_\{17\}^\{24\}35\.04328\{\}\_\{28\}^\{43\}6\.28\.4116\{\}\_\{6\}^\{11\}15\.02110\{\}\_\{10\}^\{21\}2\.5ita28\.53325\{\}\_\{25\}^\{33\}48\.15641\{\}\_\{41\}^\{56\}11\.940\.94536\{\}\_\{36\}^\{45\}66\.97459\{\}\_\{59\}^\{74\}13\.119\.42316\{\}\_\{16\}^\{23\}33\.14126\{\}\_\{26\}^\{41\}6\.98\.2116\{\}\_\{6\}^\{11\}13\.8209\{\}\_\{9\}^\{20\}4\.4por28\.53325\{\}\_\{25\}^\{33\}47\.55540\{\}\_\{40\}^\{55\}10\.040\.44536\{\}\_\{36\}^\{45\}68\.87561\{\}\_\{61\}^\{75\}11\.919\.92417\{\}\_\{17\}^\{24\}32\.54026\{\}\_\{26\}^\{40\}8\.88\.5116\{\}\_\{6\}^\{11\}15\.62211\{\}\_\{11\}^\{22\}1\.9ind25\.73022\{\}\_\{22\}^\{30\}42\.55035\{\}\_\{35\}^\{50\}9\.440\.24536\{\}\_\{36\}^\{45\}66\.97459\{\}\_\{59\}^\{74\}13\.821\.32518\{\}\_\{18\}^\{25\}40\.64833\{\}\_\{33\}^\{48\}4\.47\.2105\{\}\_\{5\}^\{10\}13\.1199\{\}\_\{9\}^\{19\}1\.9tur22\.72719\{\}\_\{19\}^\{27\}40\.64833\{\}\_\{33\}^\{48\}8\.135\.74031\{\}\_\{31\}^\{40\}63\.17055\{\}\_\{55\}^\{70\}11\.915\.81913\{\}\_\{13\}^\{19\}29\.43723\{\}\_\{23\}^\{37\}3\.87\.2105\{\}\_\{5\}^\{10\}14\.42110\{\}\_\{10\}^\{21\}1\.2cmn27\.43224\{\}\_\{24\}^\{32\}41\.95035\{\}\_\{35\}^\{50\}11\.938\.74334\{\}\_\{34\}^\{43\}63\.17055\{\}\_\{55\}^\{70\}13\.117\.52114\{\}\_\{14\}^\{21\}34\.44227\{\}\_\{27\}^\{42\}3\.18\.8127\{\}\_\{7\}^\{12\}13\.1199\{\}\_\{9\}^\{19\}4\.4jpn17\.82215\{\}\_\{15\}^\{22\}30\.03823\{\}\_\{23\}^\{38\}7\.535\.74031\{\}\_\{31\}^\{40\}61\.96954\{\}\_\{54\}^\{69\}6\.912\.31610\{\}\_\{10\}^\{16\}23\.13017\{\}\_\{17\}^\{30\}5\.06\.795\{\}\_\{5\}^\{9\}13\.8209\{\}\_\{9\}^\{20\}2\.5hin15\.91913\{\}\_\{13\}^\{19\}28\.13622\{\}\_\{22\}^\{36\}5\.031\.63627\{\}\_\{27\}^\{36\}58\.86651\{\}\_\{51\}^\{66\}8\.18\.1116\{\}\_\{6\}^\{11\}15\.62211\{\}\_\{11\}^\{22\}1\.95\.384\{\}\_\{4\}^\{8\}10\.6167\{\}\_\{7\}^\{16\}0\.0Qwen\-3\.6\-27Beng54\.85950\{\}\_\{50\}^\{59\}77\.58370\{\}\_\{70\}^\{83\}26\.958\.46354\{\}\_\{54\}^\{63\}79\.48572\{\}\_\{72\}^\{85\}31\.239\.64435\{\}\_\{35\}^\{44\}61\.26854\{\}\_\{54\}^\{68\}15\.020\.02417\{\}\_\{17\}^\{24\}30\.63824\{\}\_\{24\}^\{38\}6\.2spa38\.34334\{\}\_\{34\}^\{43\}59\.46752\{\}\_\{52\}^\{67\}15\.049\.95545\{\}\_\{45\}^\{55\}76\.28269\{\}\_\{69\}^\{82\}20\.638\.84335\{\}\_\{35\}^\{43\}60\.06752\{\}\_\{52\}^\{67\}15\.615\.51912\{\}\_\{12\}^\{19\}25\.63319\{\}\_\{19\}^\{33\}4\.4deu39\.04435\{\}\_\{35\}^\{44\}59\.46752\{\}\_\{52\}^\{67\}18\.146\.35142\{\}\_\{42\}^\{51\}66\.97459\{\}\_\{59\}^\{74\}20\.036\.54132\{\}\_\{32\}^\{41\}54\.46247\{\}\_\{47\}^\{62\}14\.414\.31811\{\}\_\{11\}^\{18\}24\.43218\{\}\_\{18\}^\{32\}5\.0fra40\.34536\{\}\_\{36\}^\{45\}58\.86651\{\}\_\{51\}^\{66\}16\.247\.55243\{\}\_\{43\}^\{52\}73\.88066\{\}\_\{66\}^\{80\}18\.839\.34435\{\}\_\{35\}^\{44\}59\.46752\{\}\_\{52\}^\{67\}16\.213\.11710\{\}\_\{10\}^\{17\}22\.53017\{\}\_\{17\}^\{30\}5\.6ita40\.94536\{\}\_\{36\}^\{45\}60\.06752\{\}\_\{52\}^\{67\}18\.150\.95646\{\}\_\{46\}^\{56\}74\.48167\{\}\_\{67\}^\{81\}22\.536\.94133\{\}\_\{33\}^\{41\}56\.26449\{\}\_\{49\}^\{64\}17\.516\.02013\{\}\_\{13\}^\{20\}29\.43723\{\}\_\{23\}^\{37\}5\.6por39\.54435\{\}\_\{35\}^\{44\}58\.86651\{\}\_\{51\}^\{66\}18\.150\.55546\{\}\_\{46\}^\{55\}69\.47662\{\}\_\{62\}^\{76\}21\.932\.43728\{\}\_\{28\}^\{37\}50\.65843\{\}\_\{43\}^\{58\}10\.615\.51912\{\}\_\{12\}^\{19\}27\.53521\{\}\_\{21\}^\{35\}5\.0ind39\.44435\{\}\_\{35\}^\{44\}61\.96954\{\}\_\{54\}^\{69\}14\.448\.35344\{\}\_\{44\}^\{53\}70\.67763\{\}\_\{63\}^\{77\}18\.132\.33728\{\}\_\{28\}^\{37\}51\.95944\{\}\_\{44\}^\{59\}11\.912\.41610\{\}\_\{10\}^\{16\}21\.92916\{\}\_\{16\}^\{29\}3\.1tur33\.33829\{\}\_\{29\}^\{38\}54\.46247\{\}\_\{47\}^\{62\}13\.144\.64940\{\}\_\{40\}^\{49\}64\.47157\{\}\_\{57\}^\{71\}18\.832\.63729\{\}\_\{29\}^\{37\}53\.16145\{\}\_\{45\}^\{61\}13\.115\.92013\{\}\_\{13\}^\{20\}28\.83622\{\}\_\{22\}^\{36\}5\.0cmn33\.33829\{\}\_\{29\}^\{38\}53\.86146\{\}\_\{46\}^\{61\}15\.043\.54839\{\}\_\{39\}^\{48\}68\.87561\{\}\_\{61\}^\{75\}13\.133\.53829\{\}\_\{29\}^\{38\}53\.16145\{\}\_\{45\}^\{61\}11\.215\.81913\{\}\_\{13\}^\{19\}27\.53521\{\}\_\{21\}^\{35\}3\.1jpn27\.53224\{\}\_\{24\}^\{32\}43\.15136\{\}\_\{36\}^\{51\}11\.240\.64536\{\}\_\{36\}^\{45\}66\.27359\{\}\_\{59\}^\{73\}11\.223\.42820\{\}\_\{20\}^\{28\}38\.84632\{\}\_\{32\}^\{46\}8\.111\.6159\{\}\_\{9\}^\{15\}19\.42614\{\}\_\{14\}^\{26\}4\.4hin22\.82719\{\}\_\{19\}^\{27\}42\.55035\{\}\_\{35\}^\{50\}3\.137\.24233\{\}\_\{33\}^\{42\}63\.87156\{\}\_\{56\}^\{71\}10\.024\.32821\{\}\_\{21\}^\{28\}41\.95035\{\}\_\{35\}^\{50\}7\.59\.5137\{\}\_\{7\}^\{13\}16\.92312\{\}\_\{12\}^\{23\}3\.8
## Appendix EQwen\-3\.5scale ladder: per\-language cells

Table[17](https://arxiv.org/html/2608.08775#A5.T17)expands the pooled scale ladder of §[5\.5](https://arxiv.org/html/2608.08775#S5.SS5)\(Table[2](https://arxiv.org/html/2608.08775#S5.T2)\) into per\-languagepass​@​1\\text\{pass\}@1for the fourQwen\-3\.5sizes, averaged over the four reference\-oracle\-backed capabilities \(timeexcluded\)\. Each cell pools4×160=6404\{\\times\}160\{=\}640scenarios per \(size, language\)\. English is the strongest column at every size andhinthe weakest, withjpnclose behind; the English\-minus\-target gap \(§[5\.5](https://arxiv.org/html/2608.08775#S5.SS5)\) is visible as the spread between the first column and the rest and*widens*with scale\.

Table 17:Qwen\-3\.5scale ladder, per\-languagepass​@​1\\text\{pass\}@1\(%\) averaged over the four capabilities \(execution, search, adaptability, ambiguity;timeexcluded\), pooling4×160=6404\{\\times\}160\{=\}640scenarios per \(size, language\)\. Best per column inbold:397B\-A17Bleads every language exceptjpn\(won by the dense27B\)\. English is the ceiling andhin/jpnthe floor at every size\. Protocol and caveats are as in Table[2](https://arxiv.org/html/2608.08775#S5.T2)\.#### Reasoning budget \(thinking effort\)\.

Increasing the agent’s thinking effort from low to high does not improve agenticpass​@​1\\text\{pass\}@1on this benchmark\. On the largest397B\-A17Bsize, a high\-effort rerun matches the low\-effort ladder to within about two points per capability and is, if anything, marginally*lower*\(adaptability17\.017\.0vs17\.417\.4; ambiguity5\.25\.2vs7\.57\.5; execution tracking similarly\)\. We therefore report the ladder atthinking\_effort==low, matching the headline open\-weights configuration\.

## Appendix FTrajectory triage protocol

The automatic estimation of Section[6\.2](https://arxiv.org/html/2608.08775#S6.SS2)is produced by a reusable diagnostic procedure that reads the evaluation harness’s per\-run artefacts directly and classifies each failure under the taxonomy of Section[6\.1](https://arxiv.org/html/2608.08775#S6.SS1)\. We describe it here in full for reproducibility, together with the infrastructure correction, the stratified reweighting, the whole\-dataset extrapolation, and the human validation that back the numbers in the main text\.

#### Inputs\.

For every \(scenario, language, run\) the harness persists: \(i\) a*verdict*record with the pass/fail judgment and the failing check\(s\); \(ii\) an*environment action log*of the agent’s tool calls, from which the graded*write*actions, the user task, and the final answer are recovered; \(iii\) the raw*agent trajectory*\(LLM calls and reasoning\); and, in newer dumps, \(iv\) a*judge log*carrying, per oracle event, the LLM\-as\-Judge rationale and the*oracle reference*\(the exact expected content\)\. Runs are indexed once; the diagnostic agents then address individual run directories rather than scanning the corpus\.

#### Phase 1 — worklist construction\.

From the index we pair each target run with the English run\(s\) of the same scenario\. In*regression mode*we retain scenarios where English passes and the target does not; in*failures mode*\(no baseline\) we retain failing runs directly\. Each work item records the run directories to compare and a pre\-computed failure family \(check type, tool\-count mismatch, loop/timeout, termination\), which supplies a coarse prior before any trajectory is read\.

#### Phase 2 — per\-unit diagnosis\.

One agent processes one scenario \(regression mode\) or one failing run \(the finer, stratified mode used for the gap decomposition\)\. It applies the taxonomy*top\-down*: infrastructure first \(a run terminated before a gradable attempt is never blamed on the model\), then translation defect, verifier artefact, and agent failure, with inconclusive reserved for genuinely under\-determined cases\. The English pass run is the oracle the dataset does not otherwise provide: if the target agent took the same correct action yet the verifier rejected it, the verdict is a verifier artefact; if the translated inputs \(prompt/universe/oracle\) changed the correct action, it is a translation defect; if the inputs are faithful and the model still erred, it is an agent failure\. Following the standard priority rule, a translation defect that also induces a downstream model error is attributed to the translation\. Judge rejections—otherwise ambiguous between a verifier artefact \(K1\) and a genuine wrong answer \(A5\)—are decided by comparing the agent’s produced value against the persisted oracle reference and judge rationale, never from which check fired alone\.

#### Phase 3 — compilation and estimation\.

Per\-unit verdicts are aggregated into the fault\-side split, the category×\\timescapability and per\-language breakdowns, and a sub\-code distribution\. Two corrections make the aggregate gap\-representative\. First,*infrastructure removal*\(below\)\. Second,*stratified reweighting*: writingRℓ,sR\_\{\\ell,s\}for the true number of infra\-clean failing runs of languageℓ\\ellin determinism stratums∈\{det,stoch\}s\\in\\\{\\mathrm\{det\},\\mathrm\{stoch\}\\\}\(obtained exactly from the run index\) andp^ℓ,s​\(c\)\\hat\{p\}\_\{\\ell,s\}\(c\)for the sampled share of causeccin that cell, the per\-language composition is

ϕℓ​\(c\)=∑sRℓ,s​p^ℓ,s​\(c\)∑sRℓ,s,ϕ​\(c\)=∑ℓ,sRℓ,s​p^ℓ,s​\(c\)∑ℓ,sRℓ,s,\\phi\_\{\\ell\}\(c\)=\\frac\{\\sum\_\{s\}R\_\{\\ell,s\}\\,\\hat\{p\}\_\{\\ell,s\}\(c\)\}\{\\sum\_\{s\}R\_\{\\ell,s\}\},\\qquad\\phi\(c\)=\\frac\{\\sum\_\{\\ell,s\}R\_\{\\ell,s\}\\,\\hat\{p\}\_\{\\ell,s\}\(c\)\}\{\\sum\_\{\\ell,s\}R\_\{\\ell,s\}\},withϕ​\(c\)\\phi\(c\)the pooled estimate reported in Table[4](https://arxiv.org/html/2608.08775#S6.T4)\. Uncertainty is estimated by resampling scenarios \(not runs\) with replacement within each\(ℓ,s\)\(\\ell,s\)cell—a cluster bootstrap,B=2000B=2000—which propagates both sampling variance and the correlation among runs of the same scenario\.

#### Infrastructure removal\.

A run is labelled*infrastructure*if it carries the harness termination sentinel, a null verdict, or appears in the list of batch\-failed \(language, capability, run\) cells arising from per\-shard cold starts\. Such runs persist a non\-null but non\-gradable outcome, so a naive reading would grade them as agent errors; they account for8\.5%8\.5\\%of all runs and are excluded from every attribution rate\. The batch failures are strongly cell\-localised \(e\.g\. a single cold\-started shard can void a majority of one language–capability–run cell\), so they distort per\-language scores unevenly if not removed; correcting them was the largest single revision to earlier estimates\.

#### Sampling\.

For the gap decomposition we triage a language\- and capability\-balanced sample of infra\-clean regressions in each determinism stratum \(≈4\\approx\\\!4deterministic and88stochastic scenarios per language×\\timescapability across the ten target languages\), diagnosed at the granularity of individual failing runs\. For the whole\-dataset audit \(below\) we additionally sample≈11\\approx\\\!11–1212scenarios per language from each of the both\-fail and partial→\\tofail cells\. All samples use a fixed seed; per\-language stratum sizesRℓ,sR\_\{\\ell,s\}are taken from the full index, not the sample\.

#### Whole\-dataset extrapolation\.

The triage above conditions on English passing and is by construction blind to scenarios that fail in*both*languages—precisely where a translation defect could deflate a target score without leaving a visible regression\. To bound contamination over the entire benchmark we cross\-tabulate all6,0566\{,\}056English×\\timestarget scenario–language pairs by outcome \(Table[18](https://arxiv.org/html/2608.08775#A6.T18)\) and use the observation that a translation defect can only corrupt a score by rendering the target*unsolvable*\(the target never passes\)\. Any pair whose target succeeds on≥1\\geq\\\!1run is therefore translation\-clean by construction \(74\.5%74\.5\\%of pairs\)\. The remaining “target\-never\-solves” pairs partition into three cells—English\-solves \(pass→\\tofail\), English\-partial \(partial→\\tofail\), and English\-fails \(fail→\\tofail\)—each of which we triage; the whole\-dataset unsolvable\-due\-to\-translation rate is their size\-weighted average\. The translation\-defect rate falls monotonically as English competence drops \(59%59\\%,16%16\\%,10%10\\%respectively\), yielding the6\.4%6\.4\\%bound of Section[6\.2](https://arxiv.org/html/2608.08775#S6.SS2)and showing the both\-fail blind spot to be cleaner than the visible gap\.

Table 18:Outcome contingency over all6,0566\{,\}056English×\\timestarget scenario–language pairs \(Claude\-4\.7\-Opus, infra\-clean,≥2\\geq\\\!2valid runs per side\)\. Pairs whose target passes on≥1\\geq\\\!1run \(74\.5%74\.5\\%\) are translation\-clean by construction\. The three “target never passes” cells—pass→\\tofail \(clean regression,59%59\\%translation\), partial→\\tofail \(16%16\\%\), and fail→\\tofail \(the both\-fail blind spot,10%10\\%\)—are each triaged; their size\-weighted translation rate is the6\.4%6\.4\\%whole\-dataset bound\.
#### Auditing the both\-fail cell\.

Because both\-fail scenarios have no passing baseline in either language, we do not judge the agent there\. Instead we run an*input\-faithfulness*audit: a translation defect is a property of the inputs, so we compare the target prompt, universe, and oracle reference against their English source and flag only discrepancies that would change the correct action\. This detects translation contamination independently of any pass/fail signal\. The audit finds a translation\-defect rate of9\.9%9\.9\\%\(95% CI5\.65\.6–16\.916\.9\) in the both\-fail cell, clustered on a small number of oracle\- and keyword\-mistranslation bugs \(a single scenario accounted for defects in five languages\)\.

#### Per\-language decomposition\.

Table[19](https://arxiv.org/html/2608.08775#A6.T19)reports the reweighted composition and the resulting model\-driven share of the conditional gap by language\. The genuine \(agent\) share is largest for Turkish and Japanese and smallest for Portuguese\- and Spanish\-heavy comparisons; the translation share is largest for Hindi and Chinese\. Per\-language intervals are wide \(each rests on≈15\\approx\\\!15sampled scenarios\), so they are indicative; the pooled estimate \(Table[4](https://arxiv.org/html/2608.08775#S6.T4)\) is the reliable quantity\.

Table 19:Per\-language reweighted gap composition and model\-driven gap \(Claude\-4\.7\-Opus\)\. “Model gap” is the conditional target failure rate on English\-solved scenarios times the agent share\. Per\-language CIs are wide \(≈15\\approx\\\!15scenarios each\); the pooled estimate is authoritative\.
#### Reliability and validation\.

Each verdict carries a self\-reported confidence \(76%76\\%high on the primary regression pass\)\. As an external check, expert linguists independently re\-analysed a subset of scenarios \(Section[6\.3](https://arxiv.org/html/2608.08775#S6.SS3)\); their attributions on deterministic regressions agree with the automatic labels and with the translation\-dominated composition of that stratum, and the deep\-dive cases in Table[5](https://arxiv.org/html/2608.08775#S6.T5)were drawn from this validation\. The main residual uncertainty is that the automatic pass tends to*under*\-attribute to the verifier \(borderline judge cases are conservatively labelled translation defects\), so the translation share in Table[4](https://arxiv.org/html/2608.08775#S6.T4)is a mild upper bound and the verifier\-artefact share a lower bound; the model\-side share, and the whole\-dataset contamination bound, are unaffected\.

## Appendix GNon\-Latin Script Errors

We investigate why performance on languages with non\-Latin script were the lowest, despite these being some of the highest\-resource languages \(notably cmn\)\. This gap to non\-Latin languages is consistent across all models\. Our initial hypothesis was that translation issues were the cause, but*we find no translation/transliteration issues from benchmark construction and find all scenarios solvable*\.

We useClaude\-5\-Sonnetto identify and categorise the agentic failures ofClaude\-4\.7\-Opusin cmn, hin, jpn\. Across 939 scenarios across the three languages, 88 failures were script\-related \(9%\)\. In such failures, we categorise 35 \(4%\) to be fully the agent’s fault and 53 \(5%\) to be the agent’s mishandling of script inconsistencies in the environment data\. The most common type of error is when the agent searches/filters on only one form of an entity \(e\.g\. written in Chinese characters\) and fails because the tool was expecting it in a different form \(e\.g\. Latin characters\)\. In many cases, the agent explicitly noticed both forms in its own reasoning, but still chose to match only one\. This class of errors is similar to those identified byBandarkar et al\. \([2026](https://arxiv.org/html/2608.08775#bib.bib4)\)\. While cross\-script settings add a layer of difficulty, a highly capable reasoning model is expected to navigate such situations more effectively \(e\.g\. retrying the search with multiple forms\)\. We conservatively estimate that this agent’s lack of care in multi\-script scenarios accounts for about a22pp drop inpass​@​3\\text\{pass\}@3accuracy, about half of the gap to Latin languages like spa and ind\. We anticipate that in real\-world multilingual agentic settings, there exists as many, if not more, script inconsistencies in tools and data\.

## Appendix HHuman Validation of Triage Labels

To verify the automatic attribution \(§[6\.2](https://arxiv.org/html/2608.08775#S6.SS2)\) at scale, we had native or fluent speakers re\-adjudicate a sample of triaged failures through a purpose\-built review interface, independently of the automatic protocol\.

#### Protocol\.

Each item pairs one failing target\-language run with its automatic verdict\. The reviewer sees only the decision\-relevant evidence—the English source task, its translation, the oracle reference, the LLM\-as\-Judge rationale \(where applicable\), and the agent’s final answer—rather than the full trajectory\. For every item the reviewer either*confirms*the assigned fault side,*rejects*it and names the side they believe is correct, or marks it*undecidable*when the evidence is insufficient\. The three sides mirror the taxonomy of §[6\.1](https://arxiv.org/html/2608.08775#S6.SS1): translation defect \(T\), verifier artefact \(K\), and agent failure \(A\); infrastructure and inconclusive items are not sampled\. Items are drawn from both regression strata \(deterministic*gap*and stochastic*partfail*\) and balanced across languages and capabilities\.

#### Coverage\.

Five linguists contributed149items in seven languages \(spa, ita, hin, jpn, ind, por, cmn\) and all four capabilities\. Excluding the 9 undecidable items,128 of 140adjudications confirmed the automatic label—91\.4%agreement \(95% Wilson CI85\.685\.6–95\.095\.0\)\. Agreement is comparable across fault sides \(Table[20](https://arxiv.org/html/2608.08775#A8.T20)\) and strata \(gap92\.9%92\.9\\%, partfail88\.1%88\.1\\%\)\. Per\-language agreement \(Table[21](https://arxiv.org/html/2608.08775#A8.T21)\) is uniformly high except for Japanese \(68%68\\%\), where reviewers reattributed several agent\-failure labels to translation defects—a small, script\-specific pocket we flag for follow\-up rather than a systematic bias\.

Table 20:Human re\-adjudication of automatic triage labels by fault side \(undecidable items excluded\)\. Agreement==confirmed//adjudicated\.Table 21:Human re\-adjudication agreement by language \(undecidable items excluded\)\.

Similar Articles

Telco-GAIA: Bilingual Benchmark for Agents in Telecom Domain

arXiv cs.AI

Telco-GAIA is a bilingual, multi-modal benchmark for evaluating tool-using agents in the telecom domain, comprising 100 human-verified tasks requiring multi-hop reasoning over heterogeneous sources, with objective scoring via exact string matching.

Beyond English: Uncovering the Multilingual Gap in Vision-Language-Action Models

arXiv cs.CL

This paper presents the first systematic study of multilingual instruction following in Vision-Language-Action (VLA) models, revealing significant performance degradation when models trained on English are evaluated on other languages. The authors propose Multilingual Principal Component Alignment (MPCA) to reduce the multilingual performance gap.

Running the Gauntlet: Re-evaluating the Capabilities of Agents Beyond Familiar Environments

Hugging Face Daily Papers

GauntletBench is a new web-based benchmark that evaluates AI agents on challenging scenarios focusing on temporal perception, graphical understanding, and 3D reasoning. Results show state-of-the-art agents achieve only 19.1% success rate compared to over 80% for non-expert humans, highlighting significant limitations in current agentic systems.