Do Evaluation Metrics Detect Errors in Classical Chinese to English Translations?
Summary
This paper investigates whether automatic evaluation metrics for machine translation are reliable for Classical Chinese to English translation, using a diagnostic framework based on minimal pairs. It finds all metrics have blind spots, with MetricX24 performing best overall.
View Cached Full Text
Cached at: 08/11/26, 08:07 AM
# Do Evaluation Metrics Detect Errors in Classical Chinese to English Translations?
Source: [https://arxiv.org/html/2608.08283](https://arxiv.org/html/2608.08283)
Osvaldo Quinjica1Eric Bennett1,2Xinchen Yang1Andrew Schonebaum2Marine Carpuat11Department of Computer Science2Department of East Asian Languages and CulturesUniversity of Maryland, College Park\{quinjica,ebenne92,xcyang,schone,marine\}@umd\.edu
###### Abstract
Although large language models can translate some historical languages surprisingly well, their usefulness in digital humanities workflows is limited by the lack of reliable evaluation\. We investigate whether existing automatic evaluation metrics developed for modern languages are reliable in this setting, using translation from Classical Chinese to English as a test case\. We introduce a diagnostic framework based on minimal pairs capturing error types salient in scholarly use, probing both reference\-based and reference\-free metrics for error sensitivity and tolerance to valid variation\. We find that all metrics exhibit blind spots, however MetricX24 performs best overall\. Our findings highlight the need for more robust and interpretable metrics for historically and culturally distinct translation settings111We release the code and dataset here:[https://github\.com/cx\-olquinjica/cc\-mt\-eval](https://github.com/cx-olquinjica/cc-mt-eval)\.
## 1Introduction
Large language models \(LLMs\) have demonstrated unexpectedly strong performance on tasks involving historical languages such as Latin\(Volket al\.,[2024](https://arxiv.org/html/2608.08283#bib.bib91)\), Hanbun\(Sonet al\.,[2022](https://arxiv.org/html/2608.08283#bib.bib92)\)or Classical Chinese\(Jinet al\.,[2023](https://arxiv.org/html/2608.08283#bib.bib93)\), likely due to incidental exposure to bilingual texts during pre\-training\(Briakouet al\.,[2023](https://arxiv.org/html/2608.08283#bib.bib100)\)\. This capability opens new opportunities for digital humanities, enabling large\-scale search, translation, and computational analysis of vast historical corpora, and has the potential to democratize access to classical texts\(Nehrdich and Keutzer,[2026](https://arxiv.org/html/2608.08283#bib.bib78); Sturgeon,[2019](https://arxiv.org/html/2608.08283#bib.bib79); Songet al\.,[2025](https://arxiv.org/html/2608.08283#bib.bib81)\)\.
Yet, evaluating the quality of LLM\-based machine translation \(MT\) in these settings remains a challenge\. Current evaluation practices rely on metrics validated primarily on modern languages, and their robustness to the linguistic, cultural and historical divergences remains unclear\. For instance, Classical Chinese poses unique difficulties: its extreme conciseness and reliance on context lead to wide variability in acceptable translations, complicating both translation and evaluation\. For scholars who lack proficiency in the source language or need to process corpora at scale, reliable evaluation of LLM translations is essential\.
Despite progress in MT evaluation, from string\-based metrics\(Papineniet al\.,[2002b](https://arxiv.org/html/2608.08283#bib.bib64); Popovic,[2015](https://arxiv.org/html/2608.08283#bib.bib65)\)to learned metrics trained on human ratings\(Reiet al\.,[2020](https://arxiv.org/html/2608.08283#bib.bib66); Juraskaet al\.,[2024](https://arxiv.org/html/2608.08283#bib.bib154)\)and LLM\-based judges\(Kocmi and Federmann,[2023c](https://arxiv.org/html/2608.08283#bib.bib57); Fernandeset al\.,[2023](https://arxiv.org/html/2608.08283#bib.bib51)\), their robustness in historical and culturally distinct settings remains largely untested\. Existing studies primarily assess aggregate correlations with human ratings\(Papineniet al\.,[2002a](https://arxiv.org/html/2608.08283#bib.bib5); Reiet al\.,[2020](https://arxiv.org/html/2608.08283#bib.bib66); Nehrdichet al\.,[2025](https://arxiv.org/html/2608.08283#bib.bib63)\), but real\-world applications demand finer\-grained insights: Which errors do these metrics reliably detect, and how tolerant are they of legitimate variation? Can they help determine whether a translation is sufficiently accurate to present to scholars unfamiliar with the source language, or to construct corpora for semantic search? Or do they risk allowing errors that could distort scholarly interpretation or learning?
To address this problem, we introduce an error diagnosis framework tailored to Classical Chinese–English MT evaluation\. Building on challenge sets based on minimal pairs\(Warstadtet al\.,[2020](https://arxiv.org/html/2608.08283#bib.bib8); Isabelleet al\.,[2017](https://arxiv.org/html/2608.08283#bib.bib9)\)and automatic perturbations\(Karpinskaet al\.,[2022](https://arxiv.org/html/2608.08283#bib.bib24); Bennettet al\.,[2025](https://arxiv.org/html/2608.08283#bib.bib23)\), our framework probes both reference\-based and reference\-free metrics, assessing their sensitivity to error\-inducing perturbations and their tolerance for acceptable paraphrastic variation\. Our contributions are threefold: \(1\) a diagnostic framework for evaluating MT metrics in this domain; \(2\) benchmarks of widely used metrics on Classical Chinese\-English translation; and \(3\) insights into improving evaluators for linguistically and culturally distinct settings\.
By situating MT evaluation in a practical and historically significant scenario, this work aims to support reliable AI\-assisted scholarly translation, while advancing our understanding of the generalizability of commonly used evaluation metrics\.
## 2Background
Classical Chinese\(wenyanwen\) provides an unusually demanding and revealing test case for LLM\-based language understanding, translation and the metrics used to assess them\. As the written medium of governance, philosophy, religion, science, and literature in East Asia for over two millennia, Classical Chinese underlies an enormous and culturally foundational textual record, comparable in scope and historical importance to Latin in Europe\(Sturgeon,[2021](https://arxiv.org/html/2608.08283#bib.bib69)\)\. Yet despite its centrality, it remains largely inaccessible to non\-specialists, making high\-quality translation a key enabling technology for digital humanities research\.
Linguistically, Classical Chinese diverges sharply from the modern languages on which most MT systems and evaluation metrics are developed\. It lacks inflectional morphology, overt tense or agreement marking, and relies on highly flexible, predominantly monosyllabic lexemes whose syntactic and semantic roles are determined almost entirely by context\(Schonebaumet al\.,[2024](https://arxiv.org/html/2608.08283#bib.bib68)\)\. Meaning is frequently implicit: arguments are omitted via pervasive zero anaphora, temporal and causal relations must be inferred, and short clauses often admit multiple plausible readings\(Schonebaumet al\.,[2024](https://arxiv.org/html/2608.08283#bib.bib68)\)\. These properties create systematic failure modes for translation models, which may default to modern character meanings, fail to disambiguate polysemous items, or inadequately integrate broader discourse and cultural context when selecting an interpretation\. Such errors can produce fluent, well\-formed English translations that are nevertheless semantically misleading\. For evaluation, the challenge is therefore two\-fold: metrics should be both tolerant of legitimate paraphrase and sensitive to subtle, context\-dependent interpretation errors\. Metrics optimized for lexical overlap or generic semantic similarity may overlook these failures, allowing mistranslations that distort historical meaning while appearing acceptable under conventional automatic evaluation\.
Despite these challenges, LLMs are already being used to translate Classical Chinese\(Sturgeon,[2021](https://arxiv.org/html/2608.08283#bib.bib69)\), yet there has been comparatively little research on methods for assessing the quality of these translations\. Their translation capabilities are often attributed to incidental exposure to historical and bilingual materials during large\-scale pretraining, as well as to models’ capacity to generalize from modern Chinese and related high\-register texts\(Briakouet al\.,[2023](https://arxiv.org/html/2608.08283#bib.bib100); Changet al\.,[2023](https://arxiv.org/html/2608.08283#bib.bib48)\)\. Growing research documents LLM performance on Classical Chinese understanding tasks, including sentence interpretation, question answering, and historical knowledge probing\(Zhouet al\.,[2023](https://arxiv.org/html/2608.08283#bib.bib153); Caoet al\.,[2024b](https://arxiv.org/html/2608.08283#bib.bib1); Liet al\.,[2024](https://arxiv.org/html/2608.08283#bib.bib125)\)\. Knowledge\-grounded and retrieval\-augmented approaches have proven useful to inject historical context, canonical references, or curated commentaries, as demonstrated in systems such as TongGu and MITRA\-zh\(Caoet al\.,[2024a](https://arxiv.org/html/2608.08283#bib.bib2); Nehrdichet al\.,[2023](https://arxiv.org/html/2608.08283#bib.bib3)\)\. Translation\-specific studies report that LLMs can generate fluent and often plausible English renderings of Classical Chinese prose and poetry, sometimes rivaling or surpassing earlier neural MT systems\(Wanget al\.,[2023](https://arxiv.org/html/2608.08283#bib.bib144); Chenet al\.,[2025](https://arxiv.org/html/2608.08283#bib.bib4); Youet al\.,[2025](https://arxiv.org/html/2608.08283#bib.bib10)\)\. At the same time, these studies consistently note important translation failures: models may over\-modernize meanings, collapse distinct classical senses into a single contemporary gloss, or fail to resolve ambiguity using broader discourse or genre\-specific conventions\. These errors are difficult to detect automatically, motivating the need for reliable translation quality evaluation\.
Evaluation of Classical Chinese translationremains comparatively underexplored\. Most existing work adopts standard MT evaluation metrics like BLEU, or relies on expert human judgment along dimensions like adequacy, fluency, and literary quality, particularly for poetry\(Chenet al\.,[2025](https://arxiv.org/html/2608.08283#bib.bib4)\)\. A small number of studies have begun to assess modern automatic metrics in this domain\.Nehrdich and Keutzer \([2026](https://arxiv.org/html/2608.08283#bib.bib78)\)curated evaluation testbeds for Buddhist Chinese\.Bennettet al\.\([2025](https://arxiv.org/html/2608.08283#bib.bib23)\)evaluated both string\-based and neural translation quality metrics on Pre\-Qin Classical Chinese texts \(pre\-221 BCE\), and found that neural metrics are more reliable than BLEU or chrF when faced with artificial perturbations\. Howeveer, it remains unclear to what degree those metrics detect naturally occurring errors and acceptable variations that matter for scholarly interpretation\. This gap motivates the need for fine\-grained, diagnostic evaluation frameworks that probe metric behavior under controlled, linguistically meaningful perturbations, rather than assuming that performance in modern, high\-resource settings will generalize to historically and culturally distant languages\.
## 3Diagnostic Framework
We adopt a controlled meta\-evaluation framework inspired by DEMETR\(Karpinskaet al\.,[2022](https://arxiv.org/html/2608.08283#bib.bib24)\), using*minimal pairs*to probe MT evaluation metrics\. Each pair contains an original translation and a minimally edited variant that differs by exactly one targeted perturbation\. A well\-behaved metric should assign a higher score to the original than to the perturbed output, allowing us to attribute score differences directly to a single error type\.
For our case study on Classical Chinese\-English translation, we extend prior diagnostic setups in two ways: \(i\) we derive perturbation categories from expert philological inspection of real LLM outputs and instantiate them via semi\-automatic rewrites vetted by trained annotators; and \(ii\) we quantify metric sensitivity through mixed\-effects modeling of standardized score deltas, rather than normalizing against catastrophic baselines \(e\.g\., empty outputs\), which we find unreliable in this setting\.
### 3\.1Data
Source texts are drawn from the Chinese Text Project \(CTP\), a curated digital library of premodern Chinese writings spanning diverse genres and historical periods\(Sturgeon,[2021](https://arxiv.org/html/2608.08283#bib.bib69)\)\. CTP also provides English translations by James Legge, which we treat as high\-quality human references to anchor evaluation and support validation\. While some of Legge’s translations have been criticized for interpretive bias and 19th\-century scholarly conventions\(MacKenzie,[2024](https://arxiv.org/html/2608.08283#bib.bib155)\), he remains a foundational scholar of Classical Chinese whose translations are widely used in academic research today\.
To construct baseline machine translations, we translate selected CTP segments into English using GPT\-4o\-mini\. These outputs serve as the starting point for all minimal pairs\. Because CTP texts and their translations are widely available online, we conduct a memorization analysis followingChenet al\.\([2024](https://arxiv.org/html/2608.08283#bib.bib12)\)to assess whether results are driven by training data overlap rather than genuine translation evaluation\. As reported in Table[6](https://arxiv.org/html/2608.08283#A3.T6)in Appendix, we find little evidence of memorization\.
CategoryExampleDescriptionERRORSModern Chinese translationSu Laoquan, twenty\-seven\. He began to feel passionate and read books\.
苏老泉,二十七岁。他开始感到热情并阅读书籍。The Classical Chinese source is first translated to English, and then to modern Chinese\.Omitted objectZhang Gai said, “Please allow me to persuadethe state of Luto remain neutral\.”
Zhang Gai said, “Please allow me to persuadeto remain neutral\.”The objectthe state of Luis omitted, making it unclear what is being persuaded to remain neutral\.Omitted subjectZi Luthen said to the family, “Not to take office is not righteous\.”
Then said to the family, “Not to take office is not righteous\.”The subjectZi Luis omitted, making it unclear who is speaking\.Incorrect lexical translationThey all say, ‘We are wise; But who can distinguish the male and femalecrow?’
They all say, ‘We are wise; But who can distinguish the male and femalepigeon?’Crowis mistranslated aspigeon, substituting the referenced entity\.Pronoun substitutionThe Master said, “There is Yung\!Hemight occupy the place of a prince\.”
The Master said, “There is Yung\!Imight occupy the place of a prince\.”Heis replaced byI, introducing a coreference error\.Incorrect tenseI am not concerned thatI amnot known\.
I am not concerned thatI wasnot known\.I amis shifted toI was, creating a tense inconsistency\.Title substitutionThedukeof the state of Zhao conferred Wucheng on Lord Mengchang\.
Thekingof the state of Zhao conferred Wucheng on Lord Mengchang\.Dukeis replaced withking, misrepresenting the rank\.Modern sense substitutionFu Xie, looking veryserious, turned him down: “If I did well and no one noticed, that is simply a matter of luck\.”
Fu Xie, looking verycolorful, turned him down: “If I did well and no one noticed, that is simply a matter of luck\.”The Classical sense \(serious/complexion\) is replaced with the modern sense \(colorful\)\.ACCEPTABLE VARIATIONSName formattingZhong Yong: The Master said, “Perfect is the virtue which is according to the law\!”
ZhongYong: The Master said, “Perfect is the virtue which is according to the law\!”Whitespace removed from a romanized name\.Name annotationZi Luthen said to the family, “Not to take office is not righteous\.”
Zi Lu \(子路\)then said to the family, “Not to take office is not righteous\.”Added name gloss within parentheses\.Sentence segmentationHe, having grown old, still regrets his tardiness\. You, young one, should think early\.
Having grown old, he still regrets his tardiness; you, young one, ought to think early\.Meaning preserving segmentation and punctuation changes with minor paraphrase\.Table 1:Errors and acceptable variations for Classical Chinese–to–English metric benchmarking\.Originalandperturbedspans are highlighted\.
### 3\.2Expert\-Grounded Error Categories
A central contribution of this work is the selection of error categories informed by domain expertise\. Several co\-authors are scholars trained in Classical Chinese philology who have experimented with LLM\-based machine translation in their own research\. Their experience revealed translation failures that influence scholarly interpretation\. To formalize these observations, we conducted a systematic manual analysis of an internal benchmark\. Three experts reviewed translations of 150 source passages translated by diverse state\-of\-the\-art LLMs\(Yanget al\.,[2024](https://arxiv.org/html/2608.08283#bib.bib84); OpenAIet al\.,[2024](https://arxiv.org/html/2608.08283#bib.bib85); Comaniciet al\.,[2025](https://arxiv.org/html/2608.08283#bib.bib86); GLMet al\.,[2024](https://arxiv.org/html/2608.08283#bib.bib87); Yanget al\.,[2025](https://arxiv.org/html/2608.08283#bib.bib88);[Team,](https://arxiv.org/html/2608.08283#bib.bib89)\), targeting difficult examples, including passages containing numbers, titles, and Classical/Modern Chinese ambiguities, as well as cases with substantial disagreement across systems or evaluation metrics \(e\.g\., high GEMBA and low COMET scores\)\.
This process yields a set of 8 error categories that reflect risks in downstream use, whether for distant reading \(e\.g\., semantic search\) or close reading \(e\.g\., philosophical or historical analysis\):*\(i\) unintended translation into modern Chinese rather than English*,*\(ii\) omitted object*,*\(iii\) omitted subject*,*\(iv\) incorrect lexical translation*,*\(v\) pronoun substitution and misattribution*,*\(vi\) incorrect tense realization*,*\(vii\) title or rank substitution*, and*\(viii\) modern\-sense substitution*\. By explicitly targeting these error categories through controlled perturbations, we introduce human expertise at scale without requiring large amounts of manual annotation\.
We also examine three categories of acceptable variations of translations, including*\(i\) name formatting*,*\(ii\) name annotation*, and*\(iii\) sentence segmentation variation and minor paraphrasing*, selected to assess whether metrics appropriately tolerate variations which do not affect the underlying meaning of a source\. Examples of all variations \(whether acceptable or errorful\) are provided in Table[1](https://arxiv.org/html/2608.08283#S3.T1)\.
### 3\.3LLM\-based perturbations
For each baseline translation, we generate minimally edited variants that instantiate exactly one error type\. Perturbations are constructed using a semi\-automatic pipeline that combines rule\-based edits, LLM\-based minimal rewrites, and hybrid procedures that leverage both source\- and target\-side information\.
#### Automatic Generation
LLM\-based rewrites are used for phenomena that require semantic or discourse\-level manipulation, such as omitting an implicit subject or substituting a modern sense for a classical one while preserving fluency\. Because English is the target language, we exploit LLM strengths in controlled English rewriting while holding content constant\. Rule\-based edits are used where deterministic modifications suffice \(e\.g\., pronoun or title substitution\)\. Hybrid methods first identify candidate locations using source\-side cues and then prompt an LLM to realize the minimal change\. See Appendix[A](https://arxiv.org/html/2608.08283#A1)for prompts\.
#### Manual Vetting
To ensure perturbations faithfully instantiate the intended error and introduce no additional changes, we select a subset of at least 100 examples per category using length\-based heuristics, and have it manually vetted by two co\-authors trained in both modern and Classical Chinese\. The results on the validated subset show a moderate agreement between annotators, with a raw agreement of 67% and a Cohen’sκ\\kappaof 0\.58\. Variation inκ\\kappais concentrated in perturbations with extreme class imbalance, where most instances are labeled as accepted \(see Appendix[D](https://arxiv.org/html/2608.08283#A4)\)\. This validation setup proved more efficient for annotators than traditional translation annotation tasks, as it required only targeted judgments compared to full translation assessments\.
#### Final Benchmark
Only pairs that clearly and only realize the intended perturbation are retained in our benchmark, resulting in 1031 minimal pairs for errors, and 334 for acceptable perturbations \(see Appendix[D](https://arxiv.org/html/2608.08283#A4)for details\)\.
### 3\.4Statistical Analysis of Metric Sensitivity
Our analysis examines whether MT evaluation metrics reliably prefer an unmodified translation over a minimally perturbed one, and how their sensitivity varies across perturbation types\. For each segment, metric, and perturbation condition, we compute a*score delta*defined as the difference between the score assigned to the baseline translation and the score assigned to its perturbed counterpart\.
Because evaluation metrics operate on different scales, we standardize score deltas \(z\-scoring\) within each metric prior to analysis\. This places all metrics on a common scale while preserving relative differences in sensitivity across perturbation types\. Next, we model standardized score deltas with a single linear mixed\-effects \(LME\) model implemented with the Pythonstatsmodelspackage, fitting error types and acceptable\-variation types jointly with fixed effects for condition, metric, and their interaction, and segment ids as random effects:
δ∗∼condition×metric⏟Fixed effects\+\(1∣segment\)⏟Random effect,\\delta^\{\*\}\\sim\\underbrace\{\\texttt\{condition\}\\times\\texttt\{metric\}\}\_\{\\text\{Fixed effects\}\}\+\\underbrace\{\(1\\mid\\texttt\{segment\}\)\}\_\{\\text\{Random effect\}\},\(1\)
The fixed effectsconditionandmetric, together with their interaction, model the effects of perturbation type, evaluation metric, and whether metric sensitivity varies across perturbation types, while the random effect\(1∣segment\)\(1\\mid\\texttt\{segment\}\)accounts for repeated observations from the same source segment\. From the fitted model we report, for each \(metric, condition\) combination, an estimated mean standardized deltaμ^\\hat\{\\mu\}and its significance test obtained by re\-levelling the model so that the target condition is the reference category, indicating whether the metric consistently distinguishes baseline translations from their modified variants\. For the time\-period analysis \(Section[5\.3](https://arxiv.org/html/2608.08283#S5.SS3)\), we additionally include the source text’s time period and its interaction with metric as fixed effects\.
## 4Translation Quality Metrics
We evaluate how well off\-the\-shelf translation quality metrics assess Classical Chinese→\\rightarrowEnglish translations\. Metrics are drawn from three families: surface overlap, learned neural metrics, and LLM\-as\-a\-judge approaches—and are evaluated in both reference\-based \(src\+ref\+hyp,ref\+hyp\) and reference\-free quality estimation \(QE;src\+hyp\) settings\. Access to the source and/or a reference directly constrains the evidence available for scoring and shapes the assumptions each metric makes about adequacy\. We focus on widely adopted metrics from recent WMT evaluations, where metrics are often applied across source languages and domains beyond their original training conditions\(Lavieet al\.,[2025](https://arxiv.org/html/2608.08283#bib.bib120); Zouharet al\.,[2024](https://arxiv.org/html/2608.08283#bib.bib82)\)\. This setting allows us to assess whether current translation evaluation paradigms transfer to Classical Chinese\.
#### Surface overlap metrics\.
BLEU\(Papineniet al\.,[2002b](https://arxiv.org/html/2608.08283#bib.bib64)\)and ChrF\(Popovic,[2015](https://arxiv.org/html/2608.08283#bib.bib65)\)measure n\-gram or character overlap between a hypothesis and a reference\. While widely used, their core assumption that lexical overlap tracks adequacy might be misaligned with Classical Chinese translation, where faithful renderings may diverge substantially in phrasing\.
#### Learned neural metrics\.
We evaluate COMET\-22, XCOMET/XCOMET\-QE, and MetricX24/MetricX24\-QE\(Reiet al\.,[2020](https://arxiv.org/html/2608.08283#bib.bib66); Guerreiroet al\.,[2024](https://arxiv.org/html/2608.08283#bib.bib29); Juraskaet al\.,[2024](https://arxiv.org/html/2608.08283#bib.bib154)\)\. These metrics rely on multilingual encoders \(XLM\-R or mT5\) fine\-tuned on human ratings of translation quality as Direct Assessment or MQM error annotation primarily on MT between modern languages in the news domain from WMT evaluations\. While prior work shows that such metrics can generalize across languages\(Reiet al\.,[2020](https://arxiv.org/html/2608.08283#bib.bib66)\), their training data does not include Classical\-Chinese sources\. We therefore test each metric in its standard configuration, as well as, where applicable, in reference\-only and QE variants\. For COMET and MetricX24, we additionally evaluate reference\-only variants \(ref\+hyp\) to isolate the contribution of source conditioning
#### LLM\-as\-a\-judge metrics\.
Finally, we evaluate GEMBA, which uses zero\- or few\-shot prompting of GPT\-family models to produce translation quality judgments following DA\- or MQM\-style instructions\(Kocmi and Federmann,[2023b](https://arxiv.org/html/2608.08283#bib.bib6);[a](https://arxiv.org/html/2608.08283#bib.bib76)\)\. Unlike supervised neural metrics, GEMBA does not learn from WMT human scores via fine\-tuning, instead inheriting the LLM’s general pretraining and instruction\-following priors\.
See Appendix[B](https://arxiv.org/html/2608.08283#A2)for implementation details\.
Error PerturbationsAcceptable VariationsLex\. errorAnc\./Mod\. senseMod\. Ch\. out\.Omit\. obj\.Omit\. subj\.PronounTenseTitleCh\. annot\.Name fmt\.Sent\. seg\.Metricμ^\\hat\{\\mu\}\(SE\)Sig\.μ^\\hat\{\\mu\}\(SE\)Sig\.μ^\\hat\{\\mu\}\(SE\)Sig\.μ^\\hat\{\\mu\}\(SE\)Sig\.μ^\\hat\{\\mu\}\(SE\)Sig\.μ^\\hat\{\\mu\}\(SE\)Sig\.μ^\\hat\{\\mu\}\(SE\)Sig\.μ^\\hat\{\\mu\}\(SE\)Sig\.μ^\\hat\{\\mu\}\(SE\)Sig\.μ^\\hat\{\\mu\}\(SE\)Sig\.μ^\\hat\{\\mu\}\(SE\)Sig\.Src \+ Ref \+ MTCOMET−\-0\.171\(0\.062\)∗∗−\-0\.286\(0\.077\)∗∗∗\+\+2\.075\(0\.056\)∗∗∗\+\+0\.402\(0\.067\)∗∗∗\+\+0\.145\(0\.065\)∗\+\+0\.043\(0\.054\)−\-0\.448\(0\.056\)∗∗∗−\-0\.496\(0\.054\)∗∗∗−\-0\.488\(0\.066\)∗∗∗−\-0\.530\(0\.065\)∗∗∗−\-0\.673\(0\.064\)∗∗∗XCOMET\+\+0\.279\(0\.090\)∗∗∗−\-0\.235\(0\.113\)∗\+\+0\.973\(0\.081\)∗∗∗\+\+0\.219\(0\.096\)∗\+\+0\.456\(0\.094\)∗∗∗−\-0\.036\(0\.080\)−\-0\.319\(0\.080\)∗∗∗−\-0\.209\(0\.080\)∗∗−\-0\.626\(0\.096\)∗∗∗−\-0\.299\(0\.095\)∗∗∗−\-0\.395\(0\.092\)∗∗∗MetricX\-24\+\+0\.277\(0\.077\)∗∗∗\+\+0\.143\(0\.095\)−\-1\.266\(0\.069\)∗∗∗\+\+0\.721\(0\.083\)∗∗∗\+\+0\.462\(0\.081\)∗∗∗\+\+1\.133\(0\.067\)∗∗∗−\-0\.149\(0\.069\)∗−\-0\.051\(0\.067\)−\-0\.314\(0\.082\)∗∗∗−\-0\.249\(0\.081\)∗∗∗−\-0\.557\(0\.079\)∗∗∗GEMBA\-DA\-ref\+\+0\.874\(0\.081\)∗∗∗−\-0\.111\(0\.103\)\+\+1\.057\(0\.073\)∗∗∗\+\+0\.149\(0\.088\)\+\+0\.015\(0\.086\)\+\+0\.342\(0\.072\)∗∗∗−\-0\.447\(0\.073\)∗∗∗−\-0\.189\(0\.072\)∗∗−\-0\.728\(0\.086\)∗∗∗−\-0\.558\(0\.086\)∗∗∗−\-0\.742\(0\.083\)∗∗∗GEMBA\-MQM\-ref\+\+0\.598\(0\.085\)∗∗∗−\-0\.159\(0\.110\)\+\+0\.979\(0\.078\)∗∗∗\+\+0\.098\(0\.091\)\+\+0\.019\(0\.089\)\+\+0\.167\(0\.077\)∗−\-0\.278\(0\.076\)∗∗∗−\-0\.110\(0\.077\)−\-0\.510\(0\.091\)∗∗∗−\-0\.415\(0\.091\)∗∗∗−\-0\.726\(0\.087\)∗∗∗Src \+ MT \(QE\)XCOMET\-QE\+\+0\.561\(0\.096\)∗∗∗−\-0\.162\(0\.122\)−\-0\.248\(0\.087\)∗∗\+\+0\.217\(0\.103\)∗\+\+0\.315\(0\.101\)∗∗∗\+\+0\.055\(0\.086\)\+\+0\.040\(0\.086\)−\-0\.085\(0\.085\)−\-0\.470\(0\.103\)∗∗∗−\-0\.153\(0\.102\)−\-0\.108\(0\.099\)MetricX\-24\-QE\+\+0\.210\(0\.084\)∗∗\+\+0\.125\(0\.103\)−\-0\.962\(0\.075\)∗∗∗\+\+0\.603\(0\.090\)∗∗∗\+\+0\.495\(0\.088\)∗∗∗\+\+1\.057\(0\.073\)∗∗∗−\-0\.121\(0\.075\)−\-0\.120\(0\.073\)−\-0\.472\(0\.088\)∗∗∗−\-0\.172\(0\.087\)∗−\-0\.564\(0\.086\)∗∗∗GEMBA\-DA\+\+0\.795\(0\.079\)∗∗∗−\-0\.097\(0\.096\)\+\+1\.230\(0\.070\)∗∗∗−\-0\.003\(0\.085\)−\-0\.154\(0\.083\)\+\+0\.526\(0\.069\)∗∗∗−\-0\.489\(0\.070\)∗∗∗−\-0\.196\(0\.068\)∗∗−\-0\.688\(0\.084\)∗∗∗−\-0\.653\(0\.082\)∗∗∗−\-0\.784\(0\.082\)∗∗∗GEMBA\-MQM\+\+0\.513\(0\.094\)∗∗∗−\-0\.190\(0\.116\)\+\+0\.783\(0\.084\)∗∗∗\+\+0\.168\(0\.101\)\+\+0\.050\(0\.098\)\+\+0\.151\(0\.082\)−\-0\.238\(0\.084\)∗∗−\-0\.157\(0\.082\)−\-0\.512\(0\.099\)∗∗∗−\-0\.334\(0\.098\)∗∗∗−\-0\.461\(0\.096\)∗∗∗Ref \+ MT onlyCOMET\-noSrc−\-0\.254\(0\.054\)∗∗∗−\-0\.326\(0\.067\)∗∗∗\+\+2\.276\(0\.048\)∗∗∗\+\+0\.300\(0\.058\)∗∗∗\+\+0\.057\(0\.057\)\+\+0\.014\(0\.047\)−\-0\.438\(0\.048\)∗∗∗−\-0\.492\(0\.047\)∗∗∗−\-0\.447\(0\.057\)∗∗∗−\-0\.541\(0\.057\)∗∗∗−\-0\.663\(0\.055\)∗∗∗chrF−\-0\.356\(0\.043\)∗∗∗−\-0\.365\(0\.053\)∗∗∗\+\+2\.553\(0\.038\)∗∗∗−\-0\.089\(0\.047\)−\-0\.180\(0\.045\)∗∗∗−\-0\.323\(0\.038\)∗∗∗−\-0\.389\(0\.039\)∗∗∗−\-0\.330\(0\.037\)∗∗∗−\-0\.349\(0\.045\)∗∗∗−\-0\.355\(0\.045\)∗∗∗−\-0\.415\(0\.044\)∗∗∗BLEU−\-0\.273\(0\.091\)∗∗∗−\-0\.264\(0\.119\)∗\+\+0\.881\(0\.084\)∗∗∗−\-0\.189\(0\.097\)−\-0\.207\(0\.096\)∗\+\+0\.062\(0\.083\)−\-0\.289\(0\.082\)∗∗∗−\-0\.120\(0\.083\)−\-0\.109\(0\.098\)\+\+0\.502\(0\.098\)−\-0\.230\(0\.093\)∗MetricX\-24\-Ref\+\+0\.205\(0\.073\)∗∗\+\+0\.242\(0\.090\)∗∗−\-1\.523\(0\.065\)∗∗∗\+\+0\.701\(0\.079\)∗∗∗\+\+0\.386\(0\.077\)∗∗∗\+\+1\.149\(0\.064\)∗∗∗−\-0\.079\(0\.065\)\+\+0\.009\(0\.064\)−\-0\.241\(0\.077\)∗∗∗−\-0\.176\(0\.077\)∗−\-0\.439\(0\.075\)∗∗∗
Table 2:Metric Sensitivity to Error Perturbations and Acceptable Variations\.LME estimatesμ^m,e\\hat\{\\mu\}\_\{m,e\}, fit jointly onΔstd\\Delta\_\{\\text\{std\}\}, the z\-scored signed change in metric score between the original and modified translations, for each metricmmand perturbation typeee\. Significance levels: \* \(p<0\.05p<0\.05\), \*\* \(p<0\.01p<0\.01\), and \*\*\* \(p<0\.001p<0\.001\)\.Blue= correct penalization \(error columns\) / tolerates variation \(variation columns\);Orange= failure \(error columns\) / over\-penalizes variation \(variation columns\); no color = not significant\. Metrics show variable sensitivity across error types, with blind spots for tense errors, title substitution, and modern\-sense substitution\. No single metric is reliable across all error categories, whereas acceptable variations are tolerated by nearly all metrics\.
## 5Results
As a preliminary sanity check, we evaluate whether the metrics can distinguish reasonable translations from clearly incorrect ones\. For each source segment, we compare the original translation against both an unrelated translation of the source and a word\-shuffled version of the original translation\. All metrics consistently prefer the original translation \(Table[8](https://arxiv.org/html/2608.08283#A5.T8)\), demonstrating that they can identify catastrophic translation errors in Classical Chinese\-to\-English translation\.
Having established this basic validity, we turn to the benchmark’s central questions\. We first analyze the sensitivity of translation quality metrics to translation errors \(Section[5\.1](https://arxiv.org/html/2608.08283#S5.SS1)\) and acceptable perturbations \(Section[5\.2](https://arxiv.org/html/2608.08283#S5.SS2)\), before turning to temporal variation across texts from different historical periods \(Section[5\.3](https://arxiv.org/html/2608.08283#S5.SS3)\)\.
### 5\.1Metric Sensitivity to Errors
To study the sensitivity of metrics across error and acceptable\-variation types \(Table[1](https://arxiv.org/html/2608.08283#S3.T1)\), we fit a single linear mixed‑effects model predicting the difference in scores between an original translation and a perturbed version jointly over both error and acceptable\-variations as described in Section[3\.4](https://arxiv.org/html/2608.08283#S3.SS4)\. There was a significant interaction between error/variation type and metricχ2\(120\)=5,677\.90\\chi^\{2\}\(120\)=5\{,\}677\.90,p<\.001p<\.001, indicating that the effect of metric onδ∗\\delta^\{\*\}differed across error and acceptable\-variation types, as expected given the diverse nature of metrics considered \(Section[4](https://arxiv.org/html/2608.08283#S4)\)\.
To understand the error sensitivity patterns per metric, and how they might impact scholarly use cases, we report per\-metric LME estimates ofμ^m,e\\hat\{\\mu\}\_\{m,e\}across the eight error types in Table[2](https://arxiv.org/html/2608.08283#S4.T2)\. These values are the model intercept obtained by re\-levelling each error, acceptable variation type and metric as the reference category in the main LME model \(see Equation[1](https://arxiv.org/html/2608.08283#S3.E1)\), such thatμ^\\hat\{\\mu\}directly estimates the mean standardized score differenceδ∗\\delta^\{\*\}for metricmmon error typeee\.
First, metric sensitivity to outputs in the wrong language varies widely\. The reference\-based metrics correctly penalize*unintended translation into modern Chinese*, except for MetricX24, which along with XCOMET\-QE also fails in the reference\-free version\. Lack of sensitivity to language mismatch is a known weakness of metrics based on multilingual encoders\(Zouharet al\.,[2024](https://arxiv.org/html/2608.08283#bib.bib82); Lavieet al\.,[2025](https://arxiv.org/html/2608.08283#bib.bib120); Knowleset al\.,[2024](https://arxiv.org/html/2608.08283#bib.bib83)\), although interestingly COMET does not suffer from this issue when it has access to a reference in our settings\.
Second, COMET, XCOMET, and MetricX24 variants adequately detect omitted content \(subject and object\)\. GEMBA variants and surface metrics such as chrF and BLEU are the main failure cases, incorrectly scoring hypotheses that drop an subject or object higher than the original translations\. This behavior is particularly problematic when translations are used to support distant reading practices, such as semantic search in English over historical texts where omitted content will lead to systematic recall errors\.
Finally, sensitivity to disambiguation and substitution errors paints a more complex picture\. No metric reliably detects*incorrect tense realization, and title or rank substitution*errors\. For*modern\-sense substitution*, MetricX24\-Ref is the only metric that reliably detects the error\. For*incorrect lexical translation*, XCOMET and MetricX24 variants join the GEMBA variants in reliably detecting the error, while COMET variants and surface metrics such as chrF and BLEU fail to penalize these mistranslations\. For*pronoun substitution and misattribution*, MetricX24 variants detect it most strongly, while GEMBA only shows moderate sensitivity\. Together, these results show that all metrics have blindspots when it comes to detecting errors that matter when interpreting Classical Chinese texts, including errors that distort meaning and might mislead readers who do not have the training to verify whether translations are supported by the source text\.
Across error types, MetricX24 emerges as the metric that is most sensitive to errors overall, despite blind spots with respect to*unintended translations into modern Chinese*,*incorrect tense*,*title or rank substitution*\. Notably, only the reference\-based variant, MetricX24\-Ref, reliably detects*modern\-sense substitution*, while the base and quality\-estimation variants show only weak signal for this error type\. This suggests that the extensive use of synthetic data in training to improve its robustness to under\- and over\-translation between modern languages\(Juraskaet al\.,[2024](https://arxiv.org/html/2608.08283#bib.bib154)\)also improves generalization to an unseen source language compared to other metrics\. Interestingly, this behavior is not limited to reference\-based configurations where MetricX24 can directly compare outputs and references in English, a target language seen at training time: MetricX24 sensitivity patterns remain similar in the quality estimation mode where assessments are based on the source and translation hypothesis alone\.
### 5\.2Metric Sensitivity to Acceptable Variations
The same joint model, signedδ∗\\delta^\{\*\}, and re\-levelling procedure described in section[5\.1](https://arxiv.org/html/2608.08283#S5.SS1)were also used to generate estimatesμ^m,v\\hat\{\\mu\}\_\{m,v\}for the three acceptable\-variation types \(Table[2](https://arxiv.org/html/2608.08283#S4.T2)\)\.
Under this formulation, the interpretation of the sign is reversed relative to the error analysis\. Because a good metric should not penalize a legitimate variation, a positiveμ^\\hat\{\\mu\}indicates undesirable over\-penalization, whereas a negativeμ^\\hat\{\\mu\}indicates tolerance, the metric assigns the variation a quality score at least as high as the original translation\.
Across all three variation types,μ^\\hat\{\\mu\}is negative for essentially every metric, reflecting broad agreement that these variations should not be penalized\. All metrics tolerate*Chinese annotation*and*sentence segmentation*, while all except BLEU tolerate*name formatting*, making BLEU the only metric to penalize this variation\.*Sentence segmentation*is the most tolerated variation type overall \(meanμ^=−0\.52\\hat\{\\mu\}=\-0\.52\), suggesting that metrics generally rate re\-segmented translations at least as favourably as the originals\.
It is encouraging that most metrics are tolerant of acceptable variations, despite not having been explicitly trained to recognize them\. The strong tolerance for sentence segmentation is particularly important for the evaluation of Classical Chinese translation, where sentence boundaries are often ambiguous and segmentation decisions frequently reflect scholarly interpretation rather than objective correctness\. However, it remains to be seen whether this holds across a broader range of perturbations\.
Error PerturbationsAcceptable VariationsLex\. errorAnc\./Mod\. senseMod\. Ch\. out\.Omit\. obj\.Omit\. subj\.PronounTenseTitleCh\. annot\.Name fmt\.Sent\. seg\.Metricμ^\\hat\{\\mu\}\(SE\)Sig\.μ^\\hat\{\\mu\}\(SE\)Sig\.μ^\\hat\{\\mu\}\(SE\)Sig\.μ^\\hat\{\\mu\}\(SE\)Sig\.μ^\\hat\{\\mu\}\(SE\)Sig\.μ^\\hat\{\\mu\}\(SE\)Sig\.μ^\\hat\{\\mu\}\(SE\)Sig\.μ^\\hat\{\\mu\}\(SE\)Sig\.μ^\\hat\{\\mu\}\(SE\)Sig\.μ^\\hat\{\\mu\}\(SE\)Sig\.μ^\\hat\{\\mu\}\(SE\)Sig\.Src \+ Ref \+ MTCOMET−\-0\.158\(0\.064\)∗−\-0\.288\(0\.075\)∗∗∗\+\+2\.005\(0\.056\)∗∗∗\+\+0\.398\(0\.066\)∗∗∗\+\+0\.092\(0\.067\)\+\+0\.037\(0\.053\)−\-0\.477\(0\.056\)∗∗∗−\-0\.499\(0\.053\)∗∗∗−\-0\.530\(0\.066\)∗∗∗−\-0\.527\(0\.065\)∗∗∗−\-0\.715\(0\.067\)∗∗∗XCOMET\+\+0\.291\(0\.095\)∗∗−\-0\.222\(0\.114\)\+\+0\.974\(0\.084\)∗∗∗\+\+0\.208\(0\.098\)∗\+\+0\.410\(0\.099\)∗∗∗−\-0\.032\(0\.080\)−\-0\.352\(0\.084\)∗∗∗−\-0\.203\(0\.080\)∗∗−\-0\.606\(0\.098\)∗∗∗−\-0\.310\(0\.097\)∗∗−\-0\.444\(0\.099\)∗∗∗MetricX\-24\+\+0\.335\(0\.082\)∗∗∗\+\+0\.153\(0\.096\)−\-1\.286\(0\.072\)∗∗∗\+\+0\.723\(0\.086\)∗∗∗\+\+0\.463\(0\.087\)∗∗∗\+\+1\.128\(0\.068\)∗∗∗−\-0\.153\(0\.073\)∗−\-0\.058\(0\.068\)−\-0\.318\(0\.085\)∗∗∗−\-0\.243\(0\.084\)∗∗−\-0\.585\(0\.086\)∗∗∗GEMBA\-DA\-ref\+\+0\.871\(0\.086\)∗∗∗−\-0\.114\(0\.103\)\+\+1\.014\(0\.076\)∗∗∗\+\+0\.159\(0\.089\)\+\+0\.002\(0\.091\)\+\+0\.319\(0\.073\)∗∗∗−\-0\.465\(0\.076\)∗∗∗−\-0\.184\(0\.072\)∗∗−\-0\.730\(0\.089\)∗∗∗−\-0\.563\(0\.089\)∗∗∗−\-0\.719\(0\.090\)∗∗∗GEMBA\-MQM\-ref\+\+0\.623\(0\.090\)∗∗∗−\-0\.154\(0\.111\)\+\+0\.942\(0\.081\)∗∗∗\+\+0\.058\(0\.092\)\+\+0\.011\(0\.095\)\+\+0\.157\(0\.078\)∗−\-0\.285\(0\.080\)∗∗∗−\-0\.107\(0\.077\)−\-0\.546\(0\.094\)∗∗∗−\-0\.415\(0\.094\)∗∗∗−\-0\.700\(0\.094\)∗∗∗Src \+ MT \(QE\)XCOMET\-QE\+\+0\.562\(0\.101\)∗∗∗−\-0\.158\(0\.122\)−\-0\.197\(0\.090\)∗\+\+0\.218\(0\.105\)∗\+\+0\.277\(0\.106\)∗∗\+\+0\.077\(0\.086\)\+\+0\.073\(0\.090\)−\-0\.078\(0\.085\)−\-0\.420\(0\.106\)∗∗∗−\-0\.157\(0\.104\)−\-0\.108\(0\.105\)MetricX\-24\-QE\+\+0\.265\(0\.090\)∗∗\+\+0\.127\(0\.105\)−\-0\.976\(0\.078\)∗∗∗\+\+0\.632\(0\.093\)∗∗∗\+\+0\.494\(0\.094\)∗∗∗\+\+1\.063\(0\.074\)∗∗∗−\-0\.129\(0\.079\)−\-0\.128\(0\.074\)−\-0\.487\(0\.092\)∗∗∗−\-0\.167\(0\.091\)−\-0\.593\(0\.093\)∗∗∗GEMBA\-DA\+\+0\.791\(0\.083\)∗∗∗−\-0\.094\(0\.097\)\+\+1\.249\(0\.073\)∗∗∗−\-0\.010\(0\.087\)−\-0\.189\(0\.087\)∗\+\+0\.527\(0\.069\)∗∗∗−\-0\.495\(0\.074\)∗∗∗−\-0\.205\(0\.069\)∗∗−\-0\.683\(0\.086\)∗∗∗−\-0\.655\(0\.084\)∗∗∗−\-0\.804\(0\.087\)∗∗∗GEMBA\-MQM\+\+0\.556\(0\.099\)∗∗∗−\-0\.051\(0\.104\)\+\+0\.727\(0\.087\)∗∗∗\+\+0\.185\(0\.103\)−\-0\.254\(0\.088\)∗∗\+\+0\.118\(0\.083\)−\-0\.150\(0\.082\)−\-0\.150\(0\.082\)−\-0\.445\(0\.102\)∗∗∗−\-0\.354\(0\.100\)∗∗∗−\-0\.485\(0\.103\)∗∗∗Ref \+ MT onlyCOMET\-noSrc−\-0\.240\(0\.056\)∗∗∗−\-0\.326\(0\.066\)∗∗∗\+\+2\.218\(0\.049\)∗∗∗\+\+0\.307\(0\.059\)∗∗∗\+\+0\.023\(0\.059\)\+\+0\.013\(0\.047\)−\-0\.459\(0\.050\)∗∗∗−\-0\.496\(0\.047\)∗∗∗−\-0\.491\(0\.058\)∗∗∗−\-0\.535\(0\.057\)∗∗∗−\-0\.695\(0\.059\)∗∗∗chrF−\-0\.372\(0\.037\)∗∗∗−\-0\.365\(0\.043\)∗∗∗\+\+2\.372\(0\.032\)∗∗∗−\-0\.104\(0\.039\)∗∗−\-0\.185\(0\.038\)∗∗∗−\-0\.322\(0\.031\)∗∗∗−\-0\.400\(0\.033\)∗∗∗−\-0\.331\(0\.031\)∗∗∗−\-0\.358\(0\.037\)∗∗∗−\-0\.355\(0\.038\)∗∗∗−\-0\.446\(0\.038\)∗∗∗BLEU−\-0\.287\(0\.088\)∗∗−\-0\.256\(0\.113\)∗\+\+0\.685\(0\.082\)∗∗∗−\-0\.221\(0\.089\)∗−\-0\.214\(0\.092\)∗\+\+0\.073\(0\.079\)−\-0\.278\(0\.079\)∗∗∗−\-0\.118\(0\.079\)−\-0\.179\(0\.094\)\+\+0\.532\(0\.094\)∗∗∗−\-0\.296\(0\.092\)∗∗MetricX\-24\-Ref\+\+0\.247\(0\.078\)∗∗\+\+0\.252\(0\.092\)∗∗−\-1\.544\(0\.068\)∗∗∗\+\+0\.690\(0\.081\)∗∗∗\+\+0\.373\(0\.082\)∗∗∗\+\+1\.147\(0\.065\)∗∗∗−\-0\.083\(0\.069\)\+\+0\.007\(0\.065\)−\-0\.220\(0\.081\)∗∗−\-0\.172\(0\.079\)∗−\-0\.436\(0\.081\)∗∗∗
Table 3:Metric Sensitivity in Pre\-Qin and Han Texts\.LME estimatesμ^m,e\\hat\{\\mu\}\_\{m,e\}, fit onΔstd\\Delta\_\{\\text\{std\}\}within Pre\-Qin and Han texts only \(1600BC–220AD\), for each metricmmand error/variation typeee\. Significance:p∗<\.05\{\}^\{\*\}p<\.05,p∗∗<\.01\{\}^\{\*\*\}p<\.01,p∗∗∗<\.001\{\}^\{\*\*\*\}p<\.001\.Blue= correct penalization \(error columns\) / tolerates variation \(variation columns\);Orange= failure \(error columns\) / over\-penalizes variation \(variation columns\); no color = not significant\.Error PerturbationsAcceptable VariationsLex\. errorAnc\./Mod\. senseMod\. Ch\. out\.Omit\. obj\.Omit\. subj\.PronounTenseTitleCh\. annot\.Name fmt\.Sent\. seg\.Metricμ^\\hat\{\\mu\}\(SE\)Sig\.μ^\\hat\{\\mu\}\(SE\)Sig\.μ^\\hat\{\\mu\}\(SE\)Sig\.μ^\\hat\{\\mu\}\(SE\)Sig\.μ^\\hat\{\\mu\}\(SE\)Sig\.μ^\\hat\{\\mu\}\(SE\)Sig\.μ^\\hat\{\\mu\}\(SE\)Sig\.μ^\\hat\{\\mu\}\(SE\)Sig\.μ^\\hat\{\\mu\}\(SE\)Sig\.μ^\\hat\{\\mu\}\(SE\)Sig\.μ^\\hat\{\\mu\}\(SE\)Sig\.Src \+ Ref \+ MTCOMET−\-0\.237\(0\.218\)−\-0\.167\(0\.819\)\+\+2\.924\(0\.273\)∗∗∗\+\+0\.580\(0\.383\)\+\+0\.628\(0\.237\)∗∗\+\+0\.266\(0\.473\)−\-0\.234\(0\.223\)−\-0\.250\(0\.579\)−\-0\.035\(0\.300\)−\-0\.689\(0\.315\)∗−\-0\.355\(0\.202\)XCOMET\+\+0\.137\(0\.282\)−\-0\.970\(0\.986\)\+\+0\.939\(0\.329\)∗∗\+\+0\.263\(0\.484\)\+\+0\.785\(0\.310\)∗−\-0\.145\(0\.569\)\+\+0\.112\(0\.281\)−\-0\.574\(0\.697\)−\-0\.624\(0\.396\)−\-0\.056\(0\.439\)−\-0\.110\(0\.273\)MetricX\-24−\-0\.171\(0\.187\)−\-0\.391\(0\.809\)−\-1\.019\(0\.270\)∗∗∗\+\+1\.214\(0\.457\)∗∗\+\+0\.454\(0\.200\)∗\+\+1\.354\(0\.467\)∗∗−\-0\.309\(0\.208\)\+\+0\.363\(0\.572\)−\-0\.015\(0\.247\)−\-0\.273\(0\.251\)−\-0\.479\(0\.181\)∗∗GEMBA\-DA\-ref\+\+0\.899\(0\.260\)∗∗∗−\-0\.014\(0\.872\)\+\+1\.591\(0\.291\)∗∗∗−\-0\.071\(0\.444\)\+\+0\.116\(0\.271\)\+\+1\.308\(0\.504\)∗∗−\-0\.251\(0\.261\)−\-0\.548\(0\.617\)−\-0\.631\(0\.352\)−\-0\.484\(0\.401\)−\-0\.866\(0\.239\)∗∗∗GEMBA\-MQM\-ref\+\+0\.401\(0\.258\)−\-0\.244\(0\.892\)\+\+1\.435\(0\.297\)∗∗∗\+\+0\.891\(0\.440\)∗\+\+0\.118\(0\.270\)\+\+0\.488\(0\.515\)−\-0\.188\(0\.261\)−\-0\.244\(0\.631\)−\-0\.182\(0\.363\)−\-0\.380\(0\.374\)−\-0\.908\(0\.235\)∗∗∗Src \+ MT \(QE\)XCOMET\-QE\+\+0\.447\(0\.278\)−\-0\.401\(1\.045\)−\-0\.888\(0\.349\)∗−\-0\.059\(0\.460\)\+\+0\.600\(0\.330\)−\-0\.868\(0\.603\)−\-0\.184\(0\.298\)−\-0\.496\(0\.739\)−\-1\.082\(0\.426\)∗−\-0\.181\(0\.465\)−\-0\.109\(0\.287\)MetricX\-24\-QE−\-0\.159\(0\.209\)−\-0\.026\(0\.749\)−\-0\.794\(0\.250\)∗∗\+\+0\.117\(0\.361\)\+\+0\.543\(0\.227\)∗\+\+0\.814\(0\.433\)−\-0\.145\(0\.226\)\+\+0\.394\(0\.530\)−\-0\.216\(0\.279\)−\-0\.239\(0\.298\)−\-0\.438\(0\.195\)∗GEMBA\-DA\+\+0\.823\(0\.255\)∗∗−\-0\.319\(0\.845\)\+\+0\.990\(0\.282\)∗∗∗\+\+0\.119\(0\.422\)\+\+0\.119\(0\.263\)\+\+0\.510\(0\.488\)−\-0\.429\(0\.246\)\+\+0\.349\(0\.597\)−\-0\.758\(0\.325\)∗−\-0\.632\(0\.366\)−\-0\.667\(0\.221\)∗∗GEMBA\-MQM\+\+0\.202\(0\.256\)−\-0\.510\(0\.907\)\+\+1\.467\(0\.302\)∗∗∗−\-0\.084\(0\.459\)\+\+0\.832\(0\.271\)∗∗\+\+1\.525\(0\.524\)∗∗−\-0\.126\(0\.271\)−\-0\.615\(0\.641\)−\-1\.264\(0\.354\)∗∗∗\+\+0\.015\(0\.391\)−\-0\.398\(0\.247\)Ref \+ MT onlyCOMET\-noSrc−\-0\.358\(0\.177\)∗−\-0\.316\(0\.635\)\+\+2\.987\(0\.212\)∗∗∗\+\+0\.211\(0\.308\)\+\+0\.362\(0\.195\)\+\+0\.062\(0\.366\)−\-0\.273\(0\.183\)−\-0\.214\(0\.449\)\+\+0\.123\(0\.242\)−\-0\.686\(0\.266\)∗∗−\-0\.433\(0\.166\)∗∗chrF−\-0\.237\(0\.184\)−\-0\.394\(0\.611\)\+\+4\.798\(0\.204\)∗∗∗\+\+0\.201\(0\.303\)−\-0\.146\(0\.182\)−\-0\.362\(0\.353\)−\-0\.279\(0\.181\)−\-0\.270\(0\.432\)−\-0\.232\(0\.236\)−\-0\.361\(0\.268\)−\-0\.229\(0\.167\)BLEU−\-0\.168\(0\.338\)−\-0\.368\(1\.122\)\+\+3\.352\(0\.374\)∗∗∗\+\+0\.467\(0\.561\)−\-0\.145\(0\.353\)−\-0\.106\(0\.646\)−\-0\.189\(0\.316\)−\-0\.316\(0\.791\)\+\+0\.374\(0\.444\)−\-0\.198\(0\.500\)\+\+0\.115\(0\.310\)MetricX\-24\-Ref−\-0\.099\(0\.195\)−\-0\.276\(0\.697\)−\-1\.272\(0\.232\)∗∗∗\+\+1\.092\(0\.521\)∗\+\+0\.446\(0\.216\)∗\+\+1\.230\(0\.402\)∗∗−\-0\.138\(0\.245\)\+\+0\.146\(0\.493\)−\-0\.348\(0\.278\)−\-0\.185\(0\.282\)−\-0\.502\(0\.188\)∗∗
Table 4:Metric Sensitivity in Post\-Han Texts\.LME estimatesμ^m,e\\hat\{\\mu\}\_\{m,e\}, fit onΔstd\\Delta\_\{\\text\{std\}\}within Post\-Han texts only \(220AD–1912\), for each metricmmand error/variation typeee\. Significance:p∗<\.05\{\}^\{\*\}p<\.05,p∗∗<\.01\{\}^\{\*\*\}p<\.01,p∗∗∗<\.001\{\}^\{\*\*\*\}p<\.001\.Blue= correct penalization \(error columns\) / tolerates variation \(variation columns\);Orange= failure \(error columns\) / over\-penalizes variation \(variation columns\); no color = not significant\.
### 5\.3Metric Sensitivity Across Time Periods
To test whether metric sensitivity to errors differs across time periods, we fit a separate LME model where the time period in which the source text was written was added as a fixed effect, separating two levels: Pre\-Qin and Han \(1600BC \- 220AD\), and Post\-Han \(220AD\-1912\), since the language of Pre\-Qin and Han texts is thought to have been not very different from cultured speech of the time, while the gap between the written and spoken language began to develop in the Han dynasty and increased with time thereafter\(Peyraube,[2004](https://arxiv.org/html/2608.08283#bib.bib90)\)\.
We found a significant metric×\\timesperiod interaction,χ2\(12\)=57\.94\\chi^\{2\}\(12\)=57\.94,p=\.001p=\.001, indicating that when the source text was written modulates how metrics respond to modifications of the translations\. However, this interaction should be interpreted cautiously because Post\-Han texts constitute only 6\.4% of the dataset\. Consequently, many errors are substantially larger in the Post\-Han \(Table[4](https://arxiv.org/html/2608.08283#S5.T4)\) subset\(omitted object, omitted subject, pronoun substitution, etc\), and many effects that are significant in the larger Pre\-Qin/Han \(Table[3](https://arxiv.org/html/2608.08283#S5.T3)\) sample lose significance despite similar point estimates\. Nevertheless, sensitivity to*unintended translation into modern Chinese*remains significant for every metric in both periods\.
## 6Conclusion
This work examined whether current machine translation evaluation metrics can detect the kinds of errors that matter for Classical Chinese–English translation\. We find that neural metrics substantially outperform surface\-overlap baselines, but their performance remains uneven across error categories\. While some perturbations are detected reliably, sensitivity to omissions and meaning\-altering substitutions is inconsistent\. MetricX24 performs best overall, but it too exhibits important blind spots\. Overall, no single metric reliably captures the full range of distinctions relevant to scholarly evaluation\.
These findings suggest that progress in Classical Chinese translation evaluation depends not only on better metrics, but also on better evaluation resources\. Metrics are least reliable on the subtle semantic distinctions that matter most for scholarly interpretation, yet such distinctions are often underrepresented in conventional benchmarks\. Developing scholar\-informed evaluation data that systematically targets these phenomena is therefore a necessary step toward more reliable automatic assessment\.
Beyond Classical Chinese, we view the framework itself as a central contribution\. Synthetic minimal pairs grounded in a scholar\-informed error taxonomy provide a scalable way to convert expert judgment into evaluation signals while minimizing large\-scale annotation requirements\. Such perturbations can support both evaluation and the training of metrics that better capture domain\-specific notions of translation quality, offering a practical interface between scholarly expertise and NLP methodology in low\-resource and historically complex language settings\.
## Acknowledgments
This project was funded in part by the Artificial Intelligence Interdisciplinary Institute at Maryland \(AIM\)\. We thank the members of the CLIP Lab at the University of Maryland for their valuable feedback and support throughout this project\. We are especially grateful to Calvin Bao, Kartik Ravisankar, HyoJung Han, and Dayeon Ki for their early feedback on the project and an earlier draft of this paper\. We also thank the Chinese Text Project, whose openly available corpus of pre\-modern Chinese texts we drew on for this work\.
## References
- Evaluating evaluation metrics for Ancient Chinese to English machine translation\.InProceedings of the Second Workshop on Ancient Language Processing,A\. Anderson, S\. Gordin, B\. Li, Y\. Liu, M\. C\. Passarotti, and R\. Sprugnoli \(Eds\.\),The Albuquerque Convention Center, Laguna,pp\. 71–76\.External Links:[Link](https://aclanthology.org/2025.alp-1.9/),[Document](https://dx.doi.org/10.18653/v1/2025.alp-1.9),ISBN 979\-8\-89176\-235\-0Cited by:[§1](https://arxiv.org/html/2608.08283#S1.p4.1),[§2](https://arxiv.org/html/2608.08283#S2.p4.1)\.
- E\. Briakou, C\. Cherry, and G\. Foster \(2023\)Searching for Needles in a Haystack: On the Role of Incidental Bilingualism in PaLM’s Translation Capability\.arXiv\.External Links:2305\.10266,[Document](https://dx.doi.org/10.48550/arXiv.2305.10266)Cited by:[§1](https://arxiv.org/html/2608.08283#S1.p1.1),[§2](https://arxiv.org/html/2608.08283#S2.p3.1)\.
- J\. Cao, D\. Peng, P\. Zhang, Y\. Shi, Y\. Liu, K\. Ding, and L\. Jin \(2024a\)TongGu: mastering classical Chinese understanding with knowledge\-grounded large language models\.InFindings of the Association for Computational Linguistics: EMNLP 2024,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),Miami, Florida, USA,pp\. 4196–4210\.External Links:[Link](https://aclanthology.org/2024.findings-emnlp.243/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-emnlp.243)Cited by:[§2](https://arxiv.org/html/2608.08283#S2.p3.1)\.
- J\. Cao, Y\. Shi, D\. Peng, Y\. Liu, and L\. Jin \(2024b\)C3bench: a comprehensive classical chinese understanding benchmark for large language models\.External Links:2405\.17732,[Link](https://arxiv.org/abs/2405.17732)Cited by:[§2](https://arxiv.org/html/2608.08283#S2.p3.1)\.
- K\. Chang, M\. Cramer, S\. Soni, and D\. Bamman \(2023\)Speak, Memory: An Archaeology of Books Known to ChatGPT/GPT\-4\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,H\. Bouamor, J\. Pino, and K\. Bali \(Eds\.\),Singapore,pp\. 7312–7327\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.453)Cited by:[§2](https://arxiv.org/html/2608.08283#S2.p3.1)\.
- A\. Chen, L\. Lou, K\. Chen, X\. Bai, Y\. Xiang, M\. Yang, T\. Zhao, and M\. Zhang \(2025\)Benchmarking LLMs for translating classical Chinese poetry: evaluating adequacy, fluency, and elegance\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 33019–33036\.External Links:[Link](https://aclanthology.org/2025.emnlp-main.1678/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.1678),ISBN 979\-8\-89176\-332\-6Cited by:[§2](https://arxiv.org/html/2608.08283#S2.p3.1),[§2](https://arxiv.org/html/2608.08283#S2.p4.1)\.
- T\. Chen, A\. Asai, N\. Mireshghallah, S\. Min, J\. Grimmelmann, Y\. Choi, H\. Hajishirzi, L\. Zettlemoyer, and P\. W\. Koh \(2024\)CopyBench: measuring literal and non\-literal reproduction of copyright\-protected text in language model generation\.arXiv preprint arXiv:2407\.07087\.Cited by:[§3\.1](https://arxiv.org/html/2608.08283#S3.SS1.p2.1)\.
- G\. Comanici, E\. Bieber, M\. Schaekermann, I\. Pasupat, N\. Sachdeva, I\. Dhillon, M\. Blistein, O\. Ram, D\. Zhang, E\. Rosen, L\. Marris, S\. Petulla, C\. Gaffney, A\. Aharoni, N\. Lintz, T\. C\. Pais, H\. Jacobsson, I\. Szpektor, N\. Jiang, K\. Haridasan, A\. Omran, N\. Saunshi, D\. Bahri, G\. Mishra, E\. Chu, T\. Boyd, B\. Hekman, A\. Parisi, C\. Zhang, K\. Kawintiranon, T\. Bedrax\-Weiss, O\. Wang, Y\. Xu, O\. Purkiss, U\. Mendlovic, I\. Deutel, N\. Nguyen, A\. Langley, F\. Korn, L\. Rossazza, A\. Ramé, S\. Waghmare, H\. Miller, N\. Byrd, A\. Sheshan, R\. Hadsell, S\. Bhardwaj, P\. Janus, T\. Rissa, D\. Horgan, A\. Abdagic, L\. Belenki, J\. Allingham, A\. Singh, T\. Guidroz, S\. Srinivasan, H\. Schmit, K\. Chiafullo, A\. Elisseeff, N\. Jha, P\. Kolhar, L\. Berrada, F\. Ding, X\. Si, S\. B\. Mallick, F\. Och, S\. Erell, E\. Ni, T\. Latkar, S\. Yang, P\. Sirkovic, Z\. Feng, R\. Leland, R\. Hornung, G\. Wu, C\. Blundell, H\. Alvari, P\. Huang, C\. Yip, S\. Deur, L\. Liu, G\. Surita, P\. Duque, D\. Damen, J\. Jia, A\. Guez, M\. Mircea, A\. Sinha, A\. Magni, P\. Stradomski, T\. Marian, V\. Galić, W\. Chen, H\. Husain, A\. Singhal, D\. Grewe, F\. Aubet, S\. Song, L\. Blanco, L\. Rechis, L\. Ho, R\. Munoz, K\. Zheng, J\. Hamrick, K\. Mather, H\. Taitelbaum, E\. Rutherford, Y\. Lei, K\. Chen, A\. Shukla, E\. Moreira, E\. Doi, B\. Isik, N\. Shabat, D\. Rogozińska, K\. Kolipaka, J\. Chang, E\. Vušak, S\. Venkatachary, S\. Noghabi, T\. Bharti, Y\. Jun, A\. Zaks, S\. Green, J\. Challagundla, W\. Wong, M\. Mohammad, D\. Hirsch, Y\. Cheng, I\. Naim, L\. Proleev, D\. Vincent, A\. Singh, M\. Krikun, D\. Krishnan, Z\. Ghahramani, A\. Atias, R\. Aggarwal, C\. Kirov, D\. Vytiniotis, C\. Koh, A\. Chronopoulou, P\. Dogra, V\. Ion, G\. Tyen, J\. Lee, F\. Weissenberger, T\. Strohman, A\. Balakrishna, J\. Rae, M\. Velic, R\. de Liedekerke, O\. Elyada, W\. Yuan, C\. Liu, L\. Shani, S\. Kishchenko, B\. Alessio, Y\. Li, R\. Song, S\. Kwei, O\. Jankowski, A\. Pappu, Y\. Namiki, Y\. Ma, N\. Tripuraneni, C\. Cherry, M\. Ikonomidis, Y\. Ling, C\. Ji, B\. Westberg, A\. Wright, D\. Yu, D\. Parkinson, S\. Ramaswamy, J\. Connor, S\. H\. Yeganeh, S\. Grover, G\. Kenwright, L\. Litchev, C\. Apps, A\. Tomala, F\. Halim, A\. Castro\-Ros, Z\. Li, A\. Boral, P\. Sho, M\. Yarom, E\. Malmi, D\. Klinghoffer, R\. Lin, A\. Ansell, P\. K\. S, S\. Zhao, S\. Zuo, A\. Santoro, H\. Cheng, S\. Demmessie, Y\. Liu, N\. Brichtova, A\. Culp, N\. Braun, D\. Graur, W\. Ng, N\. Mehta, A\. Phillips, P\. Sundberg, V\. Godbole, F\. Liu, Y\. Katariya, D\. Rim, M\. Seyedhosseini, S\. Ammirati, J\. Valfridsson, M\. Malihi, T\. Knight, A\. Toor, T\. Lampe, A\. Ittycheriah, L\. Chiang, C\. Yeung, A\. Fréchette, J\. Rao, H\. Wang, H\. Srivastava, R\. Zhang, R\. Rhodes, A\. Brand, D\. Weesner, I\. Figotin, F\. Gimeno, R\. Fellinger, P\. Marcenac, J\. Leal, E\. Marcus, V\. Cotruta, R\. Cabrera, S\. Luo, D\. Garrette, V\. Axelrod, S\. Baltateanu, D\. Barker, D\. Chen, H\. Toma, B\. Ingram, J\. Riesa, C\. Kulkarni, Y\. Zhang, H\. Liu, C\. Wang, M\. Polacek, W\. Wu, K\. Hui, A\. N\. Reyes, Y\. Su, M\. Barnes, I\. Malhi, A\. Siddiqui, Q\. Feng, M\. Damaschin, D\. Pighin, A\. Steiner, S\. Yang, R\. S\. Boppana, S\. Ivanov, A\. Kandoor, A\. Shah, A\. Mujika, D\. Huang, C\. A\. Choquette\-Choo, M\. Patel, T\. Yu, T\. Creswell, Jerry, Liu, C\. Barros, Y\. Razeghi, A\. Roy, P\. Culliton, B\. Xiong, J\. Pan, T\. Strohmann, T\. Powell, B\. Seal, D\. DeCarlo, P\. Shyam, K\. Katircioglu, X\. Wang, C\. Hardin, I\. Odisho, J\. Broder, O\. Chang, A\. Nair, A\. Shtefan, M\. O’Brien, M\. Agarwal, S\. Potluri, S\. Goyal, A\. Jhindal, S\. Thakur, Y\. Stuken, J\. Lyon, K\. Toutanova, F\. Feng, A\. Wu, B\. Horn, A\. Wang, A\. Cullum, G\. Taubman, D\. Shrivastava, C\. Shi, H\. Tomlinson, R\. Patel, T\. Tu, A\. M\. Oflazer, F\. Pongetti, M\. Yang, A\. A\. Taïga, V\. Perot, N\. W\. Pierse, F\. Han, Y\. Drori, I\. Iturrate, A\. Chakrabarti, L\. Yeung, D\. Dopson, Y\. Chen, A\. Kulshreshtha, T\. Guo, P\. Pham, T\. Schuster, J\. Chen, A\. Polozov, J\. Xing, H\. Zhou, P\. Kacham, D\. Kukliansky, A\. Miech, S\. Yaroshenko, E\. Chi, S\. Douglas, H\. Fei, M\. Blondel, P\. Myla, L\. Madmoni, X\. Wu, D\. Keysers, K\. Kjems, I\. Albuquerque, L\. Yu, J\. D’sa, M\. Plantan, V\. Ionescu, J\. S\. Elias, A\. Gupta, M\. R\. Vuyyuru, F\. Alcober, T\. Zhou, K\. Ji, F\. Hartmann, S\. Puttagunta, H\. Song, E\. Amid, A\. Stefanoiu, A\. Lee, P\. Pucciarelli, E\. Wang, A\. Raul, S\. Petrov, I\. Tian, V\. Anklin, N\. Nti, V\. Gomes, M\. Schumacher, G\. Vesom, A\. Panagopoulos, K\. Bousmalis, D\. Andor, J\. Jacob, Y\. Zhang, B\. Rosgen, M\. Kecman, M\. Tung, A\. Belias, N\. Goodman, P\. Covington, B\. Wieder, N\. Saxena, E\. Davoodi, M\. Huang, S\. Maddineni, V\. Roulet, F\. Campbell\-Ajala, P\. G\. Sessa, Xintian, Wu, G\. Lai, P\. Collins, A\. Haig, V\. Sakenas, X\. Xu, M\. Giustina, L\. E\. Shafey, P\. Charoenpanit, S\. Garg, J\. Ainslie, B\. Severson, M\. G\. Arenas, S\. Pathak, S\. Rajayogam, J\. Feng, M\. Bakker, S\. Li, N\. Wichers, J\. Rogers, X\. Geng, Y\. Li, R\. Jagerman, C\. Jia, N\. Olmert, D\. Sharon, M\. Mauger, S\. Mariserla, H\. Ma, M\. Mohabey, K\. Kim, A\. Andreev, S\. Pollom, J\. Love, V\. Jain, P\. Agrawal, Y\. Schroecker, A\. Fortin, M\. Warmuth, J\. Liu, A\. Leach, I\. Blok, G\. P\. Girirajan, R\. Aharoni, B\. Uria, A\. Sozanschi, D\. Goldberg, L\. Ionita, M\. T\. Ribeiro, M\. Zlocha, V\. Birodkar, S\. Lachgar, L\. Yuan, H\. Choudhury, M\. Ginsberg, F\. Zheng, G\. Dibb, E\. Graves, S\. Lokhande, G\. Rasskin, G\. Muraru, C\. Quick, S\. Tata, P\. Sermanet, A\. Chawla, I\. Karo, Y\. Wang, S\. Zhang, O\. Keller, A\. Dragan, G\. Su, I\. Chou, X\. Liu, Y\. Tao, S\. Prabhakara, M\. Wilson, R\. Liu, S\. Wang, G\. Evans, D\. Du, A\. Castaño, G\. Prasad, M\. E\. Mahdy, S\. Gerlach, M\. Reid, J\. Kahn, A\. Zait, T\. S\. Pillai, T\. Ulrich, G\. Wang, J\. Wassenberg, E\. Farkash, K\. Yalasangi, C\. Wang, M\. Bauza, S\. Bucher, T\. Liu, J\. Yan, G\. Leung, V\. Sindhwani, P\. Barnes, A\. Singh, I\. Jurin, J\. Chang, N\. K\. Bhumihar, S\. Eiger, G\. Citovsky, B\. Withbroe, Z\. Li, S\. Xue, N\. D\. Santo, G\. Stoyanov, Y\. Raimond, S\. Zheng, Y\. Gao, V\. Listík, S\. Kwasiborski, R\. Saputro, A\. Ozturel, G\. Mallya, K\. Majmundar, R\. West, P\. Caron, J\. Wei, L\. Castrejon, S\. Vikram, D\. Ramachandran, N\. Dhawan, J\. Park, S\. Smoot, G\. van den Driessche, Y\. Blau, C\. Malik, W\. Liang, R\. Hirsch, C\. N\. dos Santos, E\. Weinstein, A\. van den Oord, S\. Lall, N\. FitzGerald, Z\. Jiang, X\. Yang, D\. Webster, A\. Elqursh, A\. Pope, G\. Rotival, D\. Raposo, W\. Zhu, J\. Dean, S\. Alabed, D\. Tran, A\. Gupta, Z\. Gleicher, J\. Austin, E\. Rosseel, M\. Umekar, D\. Das, Y\. Sun, K\. Chen, K\. Misiunas, X\. Zhou, Y\. Di, A\. Loo, J\. Newlan, B\. Li, V\. Ramasesh, Y\. Xu, A\. Chen, S\. Gandhe, R\. Soricut, N\. Gupta, S\. Hu, S\. El\-Sayed, X\. Garcia, I\. Brusilovsky, P\. Chen, A\. Bolt, L\. Huang, A\. Gurney, Z\. Zhang, A\. Pritzel, J\. Wilkiewicz, B\. Seybold, B\. K\. Shamanna, F\. Fischer, J\. Dean, K\. Gill, R\. Mcilroy, A\. Bhowmick, J\. Selier, A\. Yang, D\. Cheng, V\. Magay, J\. Tan, D\. Varma, C\. Walder, T\. Kocisky, R\. Nakashima, P\. Natsev, M\. Kwong, I\. Gog, C\. Zhang, S\. Dieleman, T\. Jimma, A\. Ryabtsev, S\. Brahma, D\. Steiner, D\. Du, A\. Žužul, M\. Žanić, M\. Raghavachari, W\. Gierke, Z\. Zheng, D\. Petrova, Y\. Dauphin, Y\. Liu, I\. Kessler, S\. Hand, C\. Duvarney, S\. Kim, H\. Lee, L\. Hussenot, J\. Hui, J\. Smith, D\. Jain, J\. Xia, G\. S\. Tomar, K\. Amiri, D\. Phan, F\. Fuchs, T\. Weyand, N\. Tomasev, A\. Cordell, X\. Liu, J\. Mallinson, P\. Joshi, A\. Crawford, A\. Suggala, S\. Chien, N\. Fernando, M\. Sanchez\-Vargas, D\. Williams, P\. Crone, X\. Luo, I\. Karpov, J\. Shan, T\. Thurk, R\. Strudel, P\. Voigtlaender, P\. Patil, T\. Dozat, A\. Khodaei, S\. Singla, P\. Ambroszczyk, Q\. Wu, Y\. Chang, B\. Roark, C\. Hegde, T\. Ding, A\. Filos, Z\. Wu, A\. S\. Pinto, S\. Liu, S\. Khanna, A\. Pandey, S\. Mcloughlin, Q\. Li, S\. Haves, A\. Zhou, E\. Buchatskaya, I\. Leal, P\. de Boursac, N\. Akazawa, N\. Anderson, T\. Chen, K\. Somandepalli, C\. Liang, S\. Goenka, S\. Winkler, A\. Grushetsky, Y\. Ding, J\. Smith, F\. Ye, J\. Pont\-Tuset, E\. Li, R\. Li, T\. Golany, D\. Wegner, T\. Jiang, O\. Barak, Y\. Shangguan, E\. Vértes, R\. Wong, J\. Bornschein, A\. Tudor, M\. Bevilacqua, T\. Schaul, A\. S\. Rawat, Y\. Zhao, K\. Axiotis, L\. Meng, C\. McLean, J\. Lai, J\. Beattie, N\. Kushman, Y\. Liu, B\. Kutzman, F\. Lang, J\. Ye, P\. Netrapalli, P\. Mishra, M\. Khan, M\. Goel, R\. Willoughby, D\. Tian, H\. Zhuang, J\. Chen, Z\. Tsai, T\. Kementsietsidis, A\. Khare, J\. Keeling, K\. Xu, N\. Waters, F\. Altché, A\. Popat, B\. Mittal, D\. Saxton, D\. E\. Badawy, M\. Mathieu, Z\. Zheng, H\. Zhou, N\. Ranka, R\. Shin, Q\. Duan, T\. Salimans, I\. Mihailescu, U\. Shaham, M\. Chang, Y\. Assael, N\. Dikkala, M\. Izzard, V\. Cohen\-Addad, C\. Graves, V\. Feinberg, G\. Chung, D\. Strouse, D\. Karmon, S\. Sharifzadeh, Z\. Ashwood, K\. Pham, J\. Blanton, A\. Vasiloff, J\. Barber, M\. Geller, A\. Zhou, F\. Zubach, T\. Huang, L\. Zhang, H\. Gupta, M\. Young, J\. Proskurnia, R\. Votel, V\. Gabeur, G\. Barcik, A\. Tripathi, H\. Yu, G\. Yan, B\. Changpinyo, F\. Pavetić, A\. Coyle, Y\. Fujii, J\. G\. Mendez, T\. Zhou, H\. Rajamani, B\. Hechtman, E\. Cao, D\. Juan, Y\. Tan, V\. Dalibard, Y\. Du, N\. Clay, K\. Yao, W\. Jia, D\. Vijaykumar, Y\. Zhou, X\. Bai, W\. Hung, S\. Pecht, G\. Todorov, N\. Khadke, P\. Gupta, P\. Lahoti, A\. Autef, K\. Duddu, J\. Lee\-Thorp, A\. Bykovsky, T\. Misiunas, S\. Flennerhag, S\. Thangaraj, J\. McGiffin, Z\. Nado, M\. Kunesch, A\. Noever, A\. Hertz, M\. Liang, V\. Stone, E\. Palmer, S\. Daruki, A\. Pramanik, S\. Põder, A\. Kyker, M\. Khan, E\. Sluzhaev, M\. Ritter, A\. Ruderman, W\. Zhou, C\. Nagpal, K\. Vodrahalli, G\. Necula, P\. Barham, E\. Pavlick, J\. Hartford, I\. Shafran, L\. Zhao, M\. Mikuła, T\. Eccles, H\. Shimokawa, K\. Garg, L\. Vilnis, H\. Chen, I\. Shumailov, K\. Lee, A\. Abdelhamed, M\. Xie, V\. Cohen, E\. Hlavnova, D\. Malkin, C\. Sitawarin, J\. Lottes, P\. Coquinot, T\. Yu, S\. Kumar, J\. Zhang, A\. Mahendru, Z\. Ahmed, J\. Martens, T\. Chen, A\. Boag, D\. Peng, C\. Devin, A\. Klimovskiy, M\. Phuong, D\. Vainstein, J\. Xie, B\. Ramabhadran, N\. Howard, X\. Yu, G\. Goswami, J\. Cui, S\. Shleifer, M\. Pinto, C\. Yeh, M\. Yang, S\. Javanmardi, D\. Ethier, C\. Lee, J\. Orbay, S\. Kotecha, C\. Bromberg, P\. Shaw, J\. Thornton, A\. G\. Rosenthal, S\. Gu, M\. Thomas, I\. Gemp, A\. Ayyar, A\. Ushio, A\. Selvan, J\. Wee, C\. Liu, M\. Majzoubi, W\. Yu, J\. Abernethy, T\. Liechty, R\. Pan, H\. Nguyen, Qiong, Hu, S\. Perrin, A\. Arora, E\. Pitler, W\. Wang, K\. Shivakumar, F\. Prost, B\. Limonchik, J\. Wang, Y\. Gao, T\. Cour, S\. Buch, H\. Gui, M\. Ivanova, P\. Neubeck, K\. Chan, L\. Kim, H\. Chen, N\. Goyal, D\. Chung, L\. Liu, Y\. Su, A\. Petrushkina, J\. Shen, A\. Joulin, Y\. Xu, S\. X\. Lin, Y\. Kulizhskaya, C\. Chelba, S\. Vasudevan, E\. Collins, V\. Bashlovkina, T\. Lu, D\. Fritz, J\. Park, Y\. Zhou, C\. Su, R\. Tanburn, M\. Sushkov, M\. Rasquinha, J\. Li, J\. Prendki, Y\. Li, P\. LV, S\. Sharma, H\. Fitoussi, H\. Huang, A\. Dai, P\. Dao, M\. Burrows, H\. Prior, D\. Qin, G\. Pundak, L\. L\. Sjoesund, A\. Khurshudov, Z\. Zhu, A\. Webson, E\. Kemp, T\. Tan, S\. Agrawal, S\. Sargsyan, L\. Cheng, J\. Stephan, T\. Kwiatkowski, D\. Reid, A\. Byravan, A\. H\. Michaely, N\. Heess, L\. Zhou, S\. Goenka, V\. Carpenter, A\. Levskaya, B\. Wang, R\. Roberts, R\. Leblond, S\. Chikkerur, S\. Ginzburg, M\. Chang, R\. Riachi, Chuqiao, Xu, Z\. Borsos, M\. Pliskin, J\. Pawar, M\. Lustman, H\. Kirkwood, A\. Anand, A\. Chaudhary, N\. Kalb, K\. Milan, S\. Augenstein, A\. Goldie, L\. Prince, K\. Raman, Y\. Sun, V\. Xia, A\. Cohen, Z\. Huo, J\. Camp, S\. Ellis, L\. Zilka, D\. V\. Torres, L\. Patel, S\. Arora, B\. Chan, J\. Adler, K\. Ayoub, J\. Liang, F\. Jamil, J\. Jiang, S\. Baumgartner, H\. Sun, Y\. Karov, Y\. Akulov, H\. Zheng, I\. Cai, C\. Fantacci, J\. Rubin, A\. R\. Acha, M\. Wang, N\. D’Souza, R\. Sathyanarayana, S\. Dai, S\. Rowe, A\. Simanovsky, O\. Goldman, Y\. Kuang, X\. Pan, A\. Rosenberg, T\. Rojas\-Esponda, P\. Dutta, A\. Zeng, I\. Jurenka, G\. Farquhar, Y\. Bansal, S\. Iqbal, B\. Roelofs, G\. Joung, P\. Beak, C\. Ryu, R\. Poplin, Y\. Wu, J\. Alayrac, S\. Buthpitiya, O\. Ronneberger, C\. Habtegebriel, W\. Li, P\. Cavallaro, A\. Wei, G\. Bensky, T\. Denk, H\. Ganapathy, J\. Stanway, P\. Joshi, F\. Bertolini, J\. Lo, O\. Ma, Z\. Charles, G\. Sampemane, H\. Sahni, X\. Chen, H\. Askham, D\. Gaddy, P\. Young, J\. Tan, M\. Eyal, A\. Bražinskas, L\. Zhong, Z\. Wu, M\. Epstein, K\. Bailey, A\. Hard, K\. Lee, S\. Goldshtein, A\. Ruiz, M\. Badawi, M\. Lochbrunner, J\. Kearns, A\. Brown, F\. Pardo, T\. Weber, H\. Yang, P\. Jiang, B\. Akin, Z\. Fu, M\. Wainwright, C\. Zou, M\. Gaba, P\. Manzagol, W\. Kan, Y\. Song, K\. Zainullina, R\. Lin, J\. Ko, S\. Deshmukh, A\. Jindal, J\. Svensson, D\. Tyam, H\. Zhao, C\. Kaeser\-Chen, S\. Baird, P\. Moradi, J\. Hall, Q\. Guo, V\. Tsang, B\. Liang, F\. Pereira, S\. Ganesh, I\. Korotkov, J\. Adamek, S\. Thiagarajan, V\. Tran, C\. Chen, C\. Tar, S\. Jain, I\. Dasgupta, T\. Bilal, D\. Reitter, K\. Zhao, G\. Vezzani, Y\. Gehman, P\. Mehta, L\. Beltrone, X\. Dotiwalla, S\. Guadarrama, Z\. Abbas, S\. Karp, P\. Georgiev, C\. Ferng, M\. Brockschmidt, L\. Peng, C\. Hirnschall, V\. Verma, Y\. Bi, Y\. Xiao, A\. Dabush, K\. Xu, P\. Wallis, R\. Parker, Q\. Wang, Y\. Xu, I\. Safarli, D\. Tewari, Y\. Zhang, S\. Kim, A\. Gesmundo, M\. Thomas, S\. Levi, A\. Chowdhury, K\. Rao, P\. Garst, S\. Conway\-Rahman, H\. Ran, K\. McKinney, Z\. Xiao, W\. Yu, R\. Agrawal, A\. Stjerngren, C\. Ionescu, J\. Chen, V\. Sharma, J\. Chiu, F\. Liu, K\. Franko, C\. Sanford, X\. Cai, P\. Michel, S\. Ganapathy, J\. Labanowski, Z\. Garrett, B\. Vargas, S\. Sun, B\. Gale, T\. Buschmann, G\. Desjardins, N\. Ghelani, P\. Jain, M\. Verma, C\. Asawaroengchai, J\. Eisenschlos, J\. Harlalka, H\. Kazawa, D\. Metzler, J\. Howland, Y\. Jian, J\. Ades, V\. Shah, T\. Gangwani, S\. Lee, R\. Ring, S\. M\. Hernandez, D\. Reich, A\. Sinha, A\. Sathe, J\. Kovac, A\. Gill, A\. Kannan, A\. D’olimpio, M\. Sevenich, J\. Whang, B\. Kim, K\. C\. Sim, J\. Chen, J\. Zhang, S\. Lall, Y\. Matias, B\. Jia, A\. Friesen, S\. Nasso, A\. Thapliyal, B\. Perozzi, T\. Yu, A\. Shekhawat, S\. Huda, P\. Grabowski, E\. Wang, A\. Sreevatsa, H\. Dib, M\. Hassen, P\. Schuh, V\. Milutinovic, C\. Welty, M\. Quinn, A\. Shah, B\. Wang, G\. Barth\-Maron, J\. Frye, N\. Axelsson, T\. Zhu, Y\. Ma, I\. Giannoumis, H\. Sedghi, C\. Ye, Y\. Luan, K\. Aydin, B\. Chandra, V\. Sampathkumar, R\. Huang, V\. Lavrenko, A\. Eleryan, Z\. Hong, S\. Hansen, S\. M\. Carthy, B\. Samanta, D\. Ćevid, X\. Wang, F\. Li, M\. Voznesensky, M\. Hoffman, A\. Terzis, V\. Sehwag, G\. Fidel, L\. He, M\. Cai, Y\. He, A\. Feng, M\. Nikoltchev, S\. Phatale, J\. Chase, R\. Lawton, M\. Zhang, T\. Ouyang, M\. Tragut, M\. H\. Manshadi, A\. Narayanan, J\. Shen, X\. Gao, T\. Bolukbasi, N\. Roy, X\. Li, D\. Golovin, L\. Panait, Z\. Qin, G\. Han, T\. Anthony, S\. Kudugunta, V\. Patraucean, A\. Ray, X\. Chen, X\. Yang, T\. Bhatia, P\. Talluri, A\. Morris, A\. Ražnatović, B\. Brownfield, J\. An, S\. Peng, P\. Kane, C\. Zheng, N\. Duduta, J\. Kessinger, J\. Noraky, S\. Liu, K\. Rong, P\. Veličković, K\. Rush, A\. Goldin, F\. Wei, S\. M\. R\. Garlapati, C\. Pantofaru, O\. Kwon, J\. Ni, E\. Noland, J\. D\. Trapani, F\. Beaufays, A\. G\. Roy, Y\. Chow, A\. Turker, G\. Cideron, L\. Mei, J\. Clark, Q\. Dou, M\. Bošnjak, R\. Leith, Y\. Du, A\. Yazdanbakhsh, M\. Nasr, C\. Kwak, S\. S\. Sheth, A\. Kaskasoli, A\. Anand, B\. Lakshminarayanan, S\. Jerome, D\. Bieber, C\. Chu, A\. Senges, T\. Shen, M\. Sridhar, N\. Ndebele, B\. Beyret, S\. Mohamed, M\. Chen, M\. Freitag, J\. Guo, L\. Liu, P\. Roit, H\. Chen, S\. Yan, T\. Stone, J\. Co\-Reyes, J\. Cole, S\. Scellato, S\. Azizi, H\. Hashemi, A\. Jin, A\. Iyer, M\. Valentine, A\. György, A\. Ahuja, D\. H\. Diaz, C\. Lee, N\. Clement, W\. Kong, D\. Garmon, I\. Watts, K\. Bhatia, K\. Gupta, M\. Miecnikowski, H\. Vallet, A\. Taly, E\. Loper, S\. Joshi, J\. Atwood, J\. Chick, M\. Collier, F\. Iliopoulos, R\. Trostle, B\. Gunel, R\. Leal\-Cavazos, A\. M\. Hrafnkelsson, M\. Guzman, X\. Ju, A\. Forbes, J\. Emond, K\. Chauhan, B\. Caine, L\. Xiao, W\. Zeng, A\. Moufarek, D\. Murphy, M\. Meng, N\. Gupta, F\. Riedel, A\. Das, E\. Lawal, S\. Narayan, T\. Sosea, J\. Swirhun, L\. Friso, B\. Neyshabur, J\. Lu, S\. Girgin, M\. Wunder, E\. Yvinec, A\. Pyne, V\. Carbune, S\. Rijhwani, Y\. Guo, T\. Doshi, A\. Briukhov, M\. Bain, A\. Hitron, X\. Wang, A\. Gupta, K\. Chen, C\. Du, W\. Zhang, D\. Shah, A\. Akula, M\. Dylla, A\. Kachra, W\. Kuo, T\. Zou, L\. Wang, L\. Xu, J\. Zhu, J\. Snyder, S\. Menon, O\. Firat, I\. Mordatch, Y\. Yuan, N\. Ponomareva, R\. Blevins, L\. Moore, W\. Wang, P\. Chen, M\. Scholz, A\. Dwornik, J\. Lin, S\. Li, D\. Antognini, T\. I, X\. Song, M\. Miller, U\. Kalra, A\. Raveret, O\. Akerlund, F\. Wu, A\. Nystrom, N\. Godbole, T\. Liu, H\. DeBalsi, J\. Zhao, B\. Liu, A\. Caciularu, L\. Lax, U\. Khandelwal, V\. Langston, E\. Bailey, S\. Lattanzi, Y\. Wang, N\. Kovelamudi, S\. Mondal, G\. Guruganesh, N\. Hua, O\. Roval, P\. Wesołowski, R\. Ingale, J\. Halcrow, T\. Sohn, C\. Angermueller, B\. Raad, E\. Stickgold, E\. Lu, A\. Kosik, J\. Xie, T\. Lillicrap, A\. Huang, L\. L\. Zhang, D\. Paulus, C\. Farabet, A\. Wertheim, B\. Wang, R\. Joshi, C\. Ko, Y\. Wu, S\. Agrawal, L\. Lin, X\. Sheng, P\. Sung, T\. Breland\-King, C\. Butterfield, S\. Gawde, S\. Singh, Q\. Zhang, R\. Apte, S\. Shetty, A\. Hutter, T\. Li, E\. Salesky, F\. Lebron, J\. Kanerva, M\. Paganini, A\. Nguyen, R\. Vallu, J\. Peter, S\. Velury, D\. Kao, J\. Hoover, A\. Bortsova, C\. Bishop, S\. Jakobovits, A\. Agostini, A\. Agarwal, C\. Liu, C\. Kwong, S\. Tavakkol, I\. Bica, A\. Greve, A\. GP, J\. Marcus, L\. Hou, T\. Duerig, R\. Moroshko, D\. Lacey, A\. Davis, J\. Amelot, G\. Wang, F\. Kim, T\. Strinopoulos, H\. Wan, C\. L\. Lan, S\. Krishnan, H\. Tang, P\. Humphreys, J\. Bai, I\. H\. Shtacher, D\. Machado, C\. Pang, K\. Burke, D\. Liu, R\. Aravamudhan, Y\. Song, E\. Hirst, A\. Singh, B\. Jou, L\. Bai, F\. Piccinno, C\. K\. Fu, R\. Alazard, B\. Meiri, D\. Winter, C\. Chen, M\. Zhang, J\. Heitkaemper, J\. Lambert, J\. Lee, A\. Frömmgen, S\. Rogulenko, P\. Nair, P\. Niemczyk, A\. Bulyenov, B\. Xu, H\. Shemtov, M\. Zadimoghaddam, S\. Toropov, M\. Wirth, H\. Dai, S\. Gollapudi, D\. Zheng, A\. Kurakin, C\. Lee, K\. Bullard, N\. Serrano, I\. Balazevic, Y\. Li, J\. Schalkwyk, M\. Murphy, M\. Zhang, K\. Sequeira, R\. Datta, N\. Agrawal, C\. Sutton, N\. Attaluri, M\. Chiang, W\. Farhan, G\. Thornton, K\. Lin, T\. Choma, H\. Nguyen, K\. Dasgupta, D\. Robinson, I\. Comşa, M\. Riley, A\. Pillai, B\. Mustafa, B\. Golan, A\. Zandieh, J\. Lespiau, B\. Porter, D\. Ross, S\. Rajayogam, M\. Agarwal, S\. Venugopalan, B\. Shahriari, Q\. Yan, H\. Xu, T\. Tobin, P\. Dubov, H\. Shi, A\. Recasens, A\. Kovsharov, S\. Borgeaud, L\. Dery, S\. Vasanth, E\. Gribovskaya, L\. Qiu, M\. Mahdieh, W\. Skut, E\. Nielsen, C\. Zheng, A\. Yu, C\. G\. Bostock, S\. Gupta, A\. Archer, C\. Rawles, E\. Davies, A\. Svyatkovskiy, T\. Tsai, Y\. Halpern, C\. Reisswig, B\. Wydrowski, B\. Chang, J\. Puigcerver, M\. H\. Taege, J\. Li, E\. Schnider, X\. Li, D\. Dena, Y\. Xu, U\. Telang, T\. Shi, H\. Zen, K\. Kastner, Y\. Ko, N\. Subramaniam, A\. Kumar, P\. Blois, Z\. Dai, J\. Wieting, Y\. Lu, Y\. Zeldes, T\. Xie, A\. Hauth, A\. Ţifrea, Y\. Li, S\. El\-Husseini, D\. Abolafia, H\. Zhou, W\. Ding, S\. Ghalebikesabi, C\. Guía, A\. Maksai, Á\. Weisz, S\. Arik, N\. Sukhanov, A\. Świetlik, X\. Jia, L\. Yu, W\. Wang, M\. Brand, D\. Bloxwich, S\. Kirmani, Z\. Chen, A\. Go, P\. Sprechmann, N\. Kannen, A\. Carin, P\. Sandhu, I\. Edkins, L\. Nooteboom, J\. Gupta, L\. Maggiore, J\. Azizi, Y\. Pritch, P\. Yin, M\. Gupta, D\. Tarlow, D\. Smith, D\. Ivanov, M\. Babaeizadeh, A\. Goel, S\. Kambala, G\. Chu, M\. Kastelic, M\. Liu, H\. Soltau, A\. Stone, S\. Agrawal, M\. Kim, K\. Soparkar, S\. Tadepalli, O\. Bunyan, R\. Soh, A\. Kannan, D\. Kim, B\. J\. Chen, A\. Halumi, S\. Roy, Y\. Wang, O\. Sercinoglu, G\. Gibson, S\. Bhatnagar, M\. Sano, D\. von Dincklage, Q\. Ren, B\. Mitrevski, M\. Olšák, J\. She, C\. Doersch, Jilei, Wang, B\. Liu, Q\. Tan, T\. Yakar, T\. Warkentin, A\. Ramirez, C\. Lebsack, J\. Dillon, R\. Mathews, T\. Cobley, Z\. Wu, Z\. Chen, J\. Simon, S\. Nath, T\. Sainath, A\. Bendebury, R\. Julian, B\. Mankalale, D\. Ćurko, P\. Zacchello, A\. R\. Brown, K\. Sodhia, H\. Howard, S\. Caelles, A\. Gupta, G\. Evans, A\. Bulanova, L\. Katzen, R\. Goldenberg, A\. Tsitsulin, J\. Stanton, B\. Schillings, V\. Kovalev, C\. Fry, R\. Shah, K\. Lin, S\. Upadhyay, C\. Li, S\. Radpour, M\. Maggioni, J\. Xiong, L\. Haas, J\. Brennan, A\. Kamath, N\. Savinov, A\. Nagrani, T\. Yacovone, R\. Kappedal, K\. Andriopoulos, L\. Lao, Y\. Li, G\. Rozhdestvenskiy, K\. Hashimoto, A\. Audibert, S\. Austin, D\. Rodriguez, A\. Ruoss, G\. Honke, D\. Karkhanis, X\. Xiong, Q\. Wei, J\. Huang, Z\. Leng, V\. Premachandran, S\. Bileschi, G\. Evangelopoulos, T\. Mensink, J\. Pavagadhi, D\. Teplyashin, P\. Chang, L\. Xue, G\. Tanzer, S\. Goldman, K\. Patel, S\. Li, J\. Wiesner, I\. Zheng, I\. Stewart\-Binks, J\. Han, Z\. Li, L\. Luo, K\. Lenc, M\. Lučić, F\. Xue, R\. Mullins, A\. Guseynov, C\. Chang, I\. Galatzer\-Levy, A\. Zhang, G\. Bingham, G\. Hu, A\. Hartman, Y\. Ma, J\. Griffith, A\. Irpan, C\. Radebaugh, S\. Yue, L\. Fan, V\. Ungureanu, C\. Sorokin, H\. Teufel, P\. Li, R\. Anil, D\. Paparas, T\. Wang, C\. Lin, H\. Peng, M\. Shum, G\. Petrovic, D\. Brady, R\. Nguyen, K\. Macherey, Z\. Li, H\. Singh, M\. Yenugula, M\. Iinuma, X\. Chen, K\. Kopparapu, A\. Stern, S\. Dave, C\. Thekkath, F\. Perot, A\. Kumar, F\. Li, Y\. Xiao, M\. Bilotti, M\. H\. Bateni, I\. Noble, L\. Lee, A\. Vázquez\-Reina, J\. Salazar, X\. Yang, B\. Wang, E\. Gruzewska, A\. Rao, S\. Raghuram, Z\. Xu, E\. Ben\-David, J\. Mei, S\. Dalmia, Z\. Zhang, Y\. Liu, G\. Bansal, H\. Pankov, S\. Schwarcz, A\. Burns, C\. Chan, S\. Sanghai, R\. Liang, E\. Liang, A\. He, A\. Stuart, A\. Narayanan, Y\. Zhu, C\. Frank, B\. Fatemi, A\. Sabne, O\. Lang, I\. Bhattacharya, S\. Settle, M\. Wang, B\. McMahan, A\. Tacchetti, L\. B\. Soares, M\. Hadian, S\. Cabi, T\. Chung, N\. Putikhin, G\. Li, J\. Chen, A\. Tarango, H\. Michalewski, M\. Kazemi, H\. Masoom, H\. Sheftel, R\. Shivanna, A\. Vadali, R\. Comanescu, D\. Reid, J\. Moore, A\. Neelakantan, M\. Sander, J\. Herzig, A\. Rosenberg, M\. Dehghani, J\. Choi, M\. Fink, R\. Hayes, E\. Ge, S\. Weng, C\. Ho, J\. Karro, K\. Krishna, L\. N\. Thiet, A\. Skerry\-Ryan, D\. Eppens, M\. Andreetto, N\. Sarma, S\. Bonacina, B\. K\. Ayan, M\. Nawhal, Z\. Shan, M\. Dusenberry, S\. Thakoor, S\. Gubbi, D\. D\. Nguyen, R\. Tsarfaty, S\. Albanie, J\. Mitrović, M\. Gandhi, B\. Chen, A\. Epasto, G\. Stephanov, Y\. Jin, S\. Gehman, A\. Amini, J\. Weber, F\. Behbahani, S\. Xu, M\. Allamanis, X\. Chen, M\. Ott, C\. Sha, M\. Jastrzebski, H\. Qi, D\. Greene, X\. Wu, A\. Toki, D\. Vlasic, J\. Shapiro, R\. Kotikalapudi, Z\. Shen, T\. Saeki, S\. Xie, A\. Cassirer, S\. Bharadwaj, T\. Kiyono, S\. Bhojanapalli, E\. Rosenfeld, S\. Ritter, J\. Mao, J\. G\. Oliveira, Z\. Egyed, B\. Bandemer, E\. Parisotto, K\. Kinoshita, J\. Pluto, P\. Maniatis, S\. Li, Y\. Guo, G\. Ghiasi, J\. Tarbouriech, S\. Chatterjee, J\. Jin, Katrina, Xu, J\. Palomaki, S\. Arnold, M\. Sewak, F\. Piccinini, M\. Sharma, B\. Albrecht, S\. Purser\-haskell, A\. Vaswani, C\. Chen, M\. Wisniewski, Q\. Cao, J\. Aslanides, N\. M\. Phu, M\. Sieb, L\. Agubuzu, A\. Zheng, D\. Sohn, M\. Selvi, A\. Andreassen, K\. Subudhi, P\. Eruvbetine, O\. Woodman, T\. Mery, S\. Krause, X\. Ren, X\. Ma, J\. Luo, D\. Chen, W\. Fan, H\. Griffiths, C\. Schuler, A\. Li, S\. Zhang, J\. Sarr, S\. Luo, R\. Patana, M\. Watson, D\. Naboulsi, M\. Collins, S\. Sidhwani, E\. Hoogeboom, S\. Silver, E\. Caveness, X\. Zhao, M\. Rodriguez, M\. Deines, L\. Bai, P\. Griffin, M\. Tagliasacchi, E\. Xue, S\. R\. Babbula, B\. Pang, N\. Ding, G\. Shen, E\. Peake, R\. Crocker, S\. S\. Raghvendra, D\. Swisher, W\. Han, R\. Singh, L\. Wu, V\. Pchelin, T\. Munkhdalai, D\. Alon, G\. Bacon, E\. Robles, J\. Bulian, M\. Johnson, G\. Powell, F\. T\. Ferreira, Y\. Li, F\. Benzing, M\. Velimirović, H\. Soyer, W\. Kong, Tony, Nguyên, Z\. Yang, J\. Liu, J\. van Amersfoort, D\. Gillick, B\. Sun, N\. Rauschmayr, K\. Zhang, S\. Zhan, T\. Zhou, A\. Frolov, C\. Yang, D\. Vnukov, L\. Rouillard, H\. Li, A\. Mandhane, N\. Fallen, R\. Venkataraman, C\. H\. Hu, J\. Brennan, J\. Lee, J\. Chang, M\. Sundermeyer, Z\. Pan, R\. Ke, S\. Tong, A\. Fabrikant, W\. Bono, J\. Gu, R\. Foley, Y\. Mao, M\. Delakis, D\. Bhaswar, R\. Frostig, N\. Li, A\. Zipori, C\. Hope, O\. Kozlova, S\. Mishra, J\. Djolonga, C\. Schiff, M\. A\. Merey, E\. Briakou, P\. Morgan, A\. Wan, A\. Hassidim, R\. Skerry\-Ryan, K\. Sengupta, M\. Jasarevic, P\. Kallakuri, P\. Kunkle, H\. Brennan, T\. Lieber, H\. Mansoor, J\. Walker, B\. Zhang, A\. Xie, G\. Žužić, A\. Chukwuka, A\. Druinsky, D\. Cho, R\. Yao, F\. Naeem, S\. Butt, E\. Kim, Z\. Jia, M\. Jordan, A\. Lelkes, M\. Kurzeja, S\. Wang, J\. Zhao, A\. Over, A\. Chakladar, M\. Prasetya, N\. Jha, S\. Ganapathy, Y\. Cong, P\. Shroff, C\. Saroufim, S\. Miryoosefi, M\. Hammad, T\. Nasir, W\. Xi, Y\. Gao, Y\. Maeng, B\. Hora, C\. Cheng, P\. Haghani, Y\. Lewenberg, C\. Lu, M\. Matysiak, N\. Raisinghani, H\. Wang, L\. Baugher, R\. Sukthankar, M\. Giang, J\. Schultz, N\. Fiedel, M\. Chen, C\. Lee, T\. Dey, H\. Zheng, S\. Paul, C\. Smith, A\. Ly, Y\. Wang, R\. Bansal, B\. Perz, S\. Ricco, S\. Blank, V\. Keshava, D\. Sharma, M\. Chow, K\. Lad, K\. Jalan, S\. Osindero, C\. Swanson, J\. Scott, A\. Ilić, X\. Li, S\. R\. Jonnalagadda, A\. S\. Soudagar, Y\. Xiong, B\. Batsaikhan, D\. Jarrett, N\. Kumar, M\. Shah, M\. Lawlor, A\. Waters, M\. Graham, R\. May, S\. Ramos, S\. Lefdal, Z\. Cankara, N\. Cano, B\. O’Donoghue, J\. Borovik, F\. Liu, J\. Grimstad, M\. Alnahlawi, K\. Tsihlas, T\. Hudson, N\. Grigorev, Y\. Jia, T\. Huang, T\. P\. Igwe, S\. Lebedev, X\. Tang, I\. Krivokon, F\. Garcia, M\. Tan, E\. Jia, P\. Stys, S\. Vashishth, Y\. Liang, B\. Venkatraman, C\. Gu, A\. Kementsietsidis, C\. Zhu, J\. Jung, Y\. Bai, M\. J\. Hosseini, F\. Ahmed, A\. Gupta, X\. Yuan, S\. Ashraf, S\. Nigam, G\. Vasudevan, P\. Awasthi, A\. M\. Gilady, Z\. Mariet, R\. Eskander, H\. Li, H\. Hu, G\. Garrido, P\. Schlattner, G\. Zhang, R\. Saxena, P\. Dević, K\. Muralidharan, A\. Murthy, Y\. Zhou, M\. Choi, A\. Wongpanich, Z\. Wang, P\. Shah, Y\. Xu, Y\. Huang, S\. Spencer, A\. Chen, J\. Cohan, J\. Wang, J\. Tompson, J\. Wu, R\. Haroun, H\. Li, B\. Huergo, F\. Yang, T\. Yin, J\. Wendt, M\. Bendersky, R\. Chaabouni, J\. Snaider, J\. Ferret, A\. Jindal, T\. Thompson, A\. Xue, W\. Bishop, S\. M\. Phal, A\. Sharma, Y\. Sung, P\. Radhakrishnan, M\. Shomrat, R\. Ingle, R\. Vij, J\. Gilmer, M\. D\. Istin, S\. Sobell, Y\. Lu, E\. Nottage, D\. Sadigh, J\. Willcock, T\. Zhang, S\. Xu, S\. Brown, K\. Lee, G\. Wang, Y\. Zhu, Y\. Tay, C\. Kim, A\. Gutierrez, A\. Sharma, Y\. Xian, S\. Seo, C\. Cui, E\. Pochernina, C\. Baetu, K\. Jastrzębski, M\. Ly, M\. Elhawaty, D\. Suh, E\. Sezener, P\. Wang, N\. Yuen, G\. Tucker, J\. Cai, Z\. Yang, C\. Wang, A\. Muzio, H\. Qian, J\. Yoo, D\. Lockhart, K\. R\. McKee, M\. Guo, M\. Mehrotra, A\. Mendonça, S\. V\. Mehta, S\. Ben, C\. Tekur, J\. Mu, M\. Zhu, V\. Krakovna, H\. Lee, A\. Maschinot, S\. Cevey, H\. Choe, A\. Bai, H\. Srinivasan, D\. Gasaway, N\. Young, P\. Siegler, D\. Holtmann\-Rice, V\. Piratla, K\. Baumli, R\. Yogev, A\. Hofer, H\. van Hasselt, S\. Grant, Y\. Chervonyi, D\. Silver, A\. Hogue, A\. Agarwal, K\. Wang, P\. Singh, F\. Flynn, J\. Lipschultz, R\. David, L\. Bellot, Y\. Yang, L\. Le, F\. Graziano, K\. Olszewska, K\. Hui, A\. Maurya, N\. Parotsidis, W\. Chen, T\. Oguntebi, J\. Kelley, A\. Baddepudi, J\. Mauerer, G\. Shaw, A\. Siegman, L\. Yang, S\. Shetty, S\. Roy, Y\. Song, W\. Stokowiec, R\. Burnell, O\. Savant, R\. Busa\-Fekete, J\. Miao, S\. Ghosh, L\. MacDermed, P\. Lippe, M\. Dektiarev, Z\. Behrman, F\. Mentzer, K\. Nguyen, M\. Wei, S\. Verma, C\. Knutsen, S\. Dasari, Z\. Yan, P\. Mitrichev, X\. Wang, V\. Shejwalkar, J\. Austin, S\. Sunkara, N\. Potti, Y\. Virin, C\. Wright, G\. Liu, O\. Riva, E\. Pot, G\. Kochanski, Q\. Le, G\. Balasubramaniam, A\. Dhar, Y\. Liao, A\. Bloniarz, D\. Shukla, E\. Cole, J\. Lee, S\. Zhang, S\. Kafle, S\. Vashishtha, P\. Mahmoudieh, G\. Chen, R\. Hoffmann, P\. Srinivasan, A\. D\. Lago, Y\. B\. Shalom, Z\. Wang, M\. Elabd, A\. Sharma, J\. Oh, S\. Kothawade, M\. Le, M\. Monteiro, S\. Yang, K\. Alarakyia, R\. Geirhos, D\. Mincu, H\. Garnes, H\. Kobayashi, S\. Mariooryad, K\. Krasowiak, Zhixin, Lai, S\. Mourad, M\. Wang, F\. Bu, O\. Aharoni, G\. Chen, A\. Goyal, V\. Zubov, A\. Bapna, E\. Dabir, N\. Kothari, K\. Lamerigts, N\. D\. Cao, J\. Shar, C\. Yew, N\. Kulkarni, D\. Mahaarachchi, M\. Joshi, Z\. Zhu, J\. Lichtarge, Y\. Zhou, H\. Muckenhirn, V\. Selo, O\. Vinyals, P\. Chen, A\. Brohan, V\. Mehta, S\. Cogan, R\. Wang, T\. Geri, W\. Ko, W\. Chen, F\. Viola, K\. Shivam, L\. Wang, M\. C\. Elish, R\. A\. Popa, S\. Pereira, J\. Liu, R\. Koster, D\. Kim, G\. Zhang, S\. Ebrahimi, P\. Talukdar, Y\. Zheng, P\. Poklukar, A\. Mikhalap, D\. Johnson, A\. Vijayakumar, M\. Omernick, M\. Dibb, A\. Dubey, Q\. Hu, A\. Suman, V\. Aggarwal, I\. Kornakov, F\. Xia, W\. Lowe, A\. Kolganov, T\. Xiao, V\. Nikolaev, S\. Hemingray, B\. Li, J\. Iljazi, M\. Rybiński, B\. Sandhu, P\. Lu, T\. Luong, R\. Jenatton, V\. Govindaraj, Hui, Li, G\. Dulac\-Arnold, W\. Park, H\. Wang, A\. Modi, J\. Pouget\-Abadie, K\. Greller, R\. Gupta, R\. Berry, P\. Ramachandran, J\. Xie, L\. McCafferty, J\. Wang, K\. Gupta, H\. Lim, B\. Bratanič, A\. Brock, I\. Akolzin, J\. Sproch, D\. Karliner, D\. Kim, A\. Goedeckemeyer, N\. Shazeer, C\. Schmid, D\. Calandriello, P\. Bhatia, K\. Choromanski, C\. Montgomery, D\. Dua, A\. Ramalho, H\. King, Y\. Gao, L\. Nguyen, D\. Lindner, D\. Pitta, O\. Johnson, K\. Salama, D\. Ardila, M\. Han, E\. Farnese, S\. Odoom, Z\. Wang, X\. Ding, N\. Rink, R\. Smith, H\. T\. Lehri, E\. Cohen, N\. Vats, T\. He, P\. Gopavarapu, A\. Paszke, M\. Patel, W\. V\. Gansbeke, L\. Loher, L\. Castro, M\. Voitovich, T\. von Glehn, N\. George, S\. Niklaus, Z\. Eaton\-Rosen, N\. Rakićević, E\. Jue, S\. Perel, C\. Zhang, Y\. Bahat, A\. Pouget, Z\. Xing, F\. Huot, A\. Shenoy, T\. Bos, V\. Coriou, B\. Richter, N\. Noy, Y\. Wang, S\. Ontanon, S\. Qin, G\. Makarchuk, D\. Hassabis, Z\. Li, M\. Sharma, K\. Venkatesan, I\. Kemaev, R\. Daniel, S\. Huang, S\. Shah, O\. Ponce, Warren, Chen, M\. Faruqui, J\. Wu, S\. Andačić, S\. Payrits, D\. McDuff, T\. Hume, Y\. Cao, M\. Tessler, Q\. Wang, Y\. Wang, I\. Rendulic, E\. Agustsson, M\. Johnson, T\. Lando, A\. Howard, S\. G\. S\. Padmanabhan, M\. Daswani, A\. Banino, M\. Kilgore, J\. Heek, Z\. Ji, A\. Caceres, C\. Li, N\. Kassner, A\. Vlaskin, Z\. Liu, A\. Grills, Y\. Hou, R\. Sukkerd, G\. Cheon, N\. Shetty, L\. Markeeva, P\. Stanczyk, T\. Iyer, Y\. Gong, S\. Gao, K\. Gopalakrishnan, T\. Blyth, M\. Reynolds, A\. Bhoopchand, M\. Bilenko, D\. Gharibian, V\. Zayats, A\. Faust, A\. Singh, M\. Ma, H\. Jiao, S\. Vijayanarasimhan, L\. Aroyo, V\. Yadav, S\. Chakera, A\. Kakarla, V\. Meshram, K\. Gregor, G\. Botea, E\. Senter, D\. Jia, G\. Kovacs, N\. Sharma, S\. Baur, K\. Kang, Y\. He, L\. Zhuo, M\. Kostelac, I\. Laish, S\. Peng, L\. O’Bryan, D\. Kasenberg, G\. R\. Rao, E\. Leurent, B\. Zhang, S\. Stevens, A\. Salazar, Y\. Zhang, I\. Lobov, J\. Walker, A\. Porter, M\. Redshaw, H\. Ke, A\. Rao, A\. Lee, H\. Lam, M\. Moffitt, J\. Kim, S\. Qiao, T\. Koo, R\. Dadashi, X\. Song, M\. Sundararajan, P\. Xu, C\. Kawamoto, Y\. Zhong, C\. Barbu, A\. Reddy, M\. Verzetti, L\. Li, G\. Papamakarios, H\. Klimczak\-Plucińska, M\. Cassin, K\. Kavukcuoglu, R\. Swavely, A\. Vaucher, J\. Zhao, R\. Hemsley, M\. Tschannen, H\. Ge, G\. Menghani, Y\. Yu, N\. Ha, W\. He, X\. Wu, M\. Song, R\. Sterneck, S\. Zinke, D\. A\. Calian, A\. Marsden, A\. C\. Ruiz, M\. Hessel, A\. Gueta, B\. Lee, B\. Farris, M\. Gupta, Y\. Li, M\. Saleh, V\. Misra, K\. Xiao, P\. Mendolicchio, G\. Buttimore, V\. Krayvanova, N\. Nayakanti, M\. Wiethoff, Y\. Pande, A\. Mirhoseini, N\. Lao, J\. Liu, Y\. Hua, A\. Chen, Y\. Malkov, D\. Kalashnikov, S\. Gupta, K\. Audhkhasi, Y\. Zhai, S\. Kopalle, P\. Jain, E\. Ofek, C\. Meyer, K\. Baatarsukh, H\. Strejček, J\. Qian, J\. Freedman, R\. Figueira, M\. Sokolik, O\. Bachem, R\. Lin, D\. Kharrat, C\. Hidey, P\. Xu, D\. Duan, Y\. Li, M\. Ersoy, R\. Everett, K\. Cen, R\. Santamaria\-Fernandez, A\. Taubenfeld, I\. Mackinnon, L\. Deng, P\. Zablotskaia, S\. Viswanadha, S\. Goel, D\. Yates, Y\. Deng, P\. Choy, M\. Chen, A\. Sinha, A\. Mossin, Y\. Wang, A\. Szlam, S\. Hao, P\. K\. Rubenstein, M\. Toksoz\-Exley, M\. Aperghis, Y\. Zhong, J\. Ahn, M\. Isard, O\. Lacombe, F\. Luisier, C\. Anastasiou, Y\. Kalley, U\. Prabhu, E\. Dunleavy, S\. Bijwadia, J\. Mao\-Jones, K\. Chen, R\. Pasumarthi, E\. Wood, A\. Dostmohamed, N\. Hurley, J\. Simsa, A\. Parrish, M\. Pajarskas, M\. Harvey, O\. Skopek, Y\. Kochinski, J\. Rey, V\. Rieser, D\. Zhou, S\. J\. Lee, T\. Acharya, G\. Li, J\. Jiang, X\. Zhang, B\. Gipson, E\. Mahintorabi, M\. Gelmi, N\. Khajehnouri, A\. Yeh, K\. Lee, L\. Matthey, L\. Baker, T\. Pham, H\. Fu, A\. Pak, P\. Gupta, C\. Vasconcelos, A\. Sadovsky, B\. Walker, S\. Hsiao, P\. Zochbauer, A\. Marzoca, N\. Velan, J\. Zeng, G\. Baechler, D\. Driess, D\. Jain, Y\. Huang, L\. Tao, J\. Maggs, N\. Levine, J\. Schneider, E\. Gemzer, S\. Petit, S\. Han, Z\. Fisher, D\. Zelle, C\. Biles, E\. Ie, A\. Fadeeva, C\. Liu, J\. V\. Franco, A\. Collister, H\. Zhang, R\. Wang, R\. Zhao, L\. Kieliger, K\. Shuster, R\. Zhu, B\. Gong, L\. Chan, R\. Sun, S\. Basu, R\. Zimmermann, J\. Hayes, A\. Bapna, J\. Snoek, W\. Yang, P\. Datta, J\. A\. Abdallah, K\. Kilgour, L\. Li, S\. Mah, Y\. Jun, M\. Rivière, A\. Karmarkar, T\. Spalink, T\. Huang, L\. Gonzalez, D\. Tran, A\. Nowak, J\. Palowitch, M\. Chadwick, E\. Talius, H\. Mehta, T\. Sellam, P\. Fränken, M\. Nicosia, K\. He, A\. Kini, D\. Amos, S\. Basu, H\. Jobe, E\. Shaw, Q\. Xu, C\. Evans, D\. Ikeda, C\. Yan, L\. Jin, L\. Wang, S\. Yadav, I\. Labzovsky, R\. Sampath, A\. Ma, C\. Schumann, A\. Siddhant, R\. Shah, J\. Youssef, R\. Agarwal, N\. Dabney, A\. Tonioni, M\. Ambar, J\. Li, I\. Guyon, B\. Li, D\. Soergel, B\. Fang, G\. Karadzhov, C\. Udrescu, T\. Trinh, V\. Raunak, S\. Noury, D\. Guo, S\. Gupta, M\. Finkelstein, D\. Petek, L\. Liang, G\. Billock, P\. Sun, D\. Wood, Y\. Song, X\. Yu, T\. Matejovicova, R\. Cohen, K\. Andra, D\. D’Ambrosio, Z\. Deng, V\. Nallatamby, E\. Songhori, R\. Dangovski, A\. Lampinen, P\. Botadra, A\. Hillier, J\. Cao, N\. Baddi, A\. Kuncoro, T\. Yoshino, A\. Bhagatwala, M\. Ranzato, R\. Schaeffer, T\. Liu, S\. Ye, O\. Sarvana, J\. Nham, C\. Kuang, I\. Gao, J\. Baek, S\. Mittal, A\. Wahid, A\. Gergely, B\. Ni, J\. Feldman, C\. Muir, P\. Lamblin, W\. Macherey, E\. Dyer, L\. Kilpatrick, V\. Campos, M\. Bhutani, S\. Fort, Y\. Ahmad, A\. Severyn, K\. Chatziprimou, O\. Ferludin, M\. Dimarco, A\. Kusupati, J\. Heyward, D\. Bahir, K\. Villela, K\. Millican, D\. Marcus, S\. Bahargam, C\. Unlu, N\. Roth, Z\. Wei, S\. Gopal, D\. Ghoshal, E\. Lee, S\. Lin, J\. Lees, D\. Lee, A\. Hosseini, C\. Fan, S\. Neel, M\. Wu, Y\. Altun, H\. Cai, E\. Piqueras, J\. Woodward, A\. Bissacco, S\. Haykal, M\. Bordbar, P\. Sundaram, S\. Hodkinson, D\. Toyama, G\. Polovets, A\. Myers, A\. Sinha, T\. Levinboim, K\. Krishnakumar, R\. Chhaparia, T\. Sholokhova, N\. B\. Gundavarapu, G\. Jawahar, H\. Qureshi, J\. Hu, N\. Momchev, M\. Rahtz, R\. Wu, A\. P\. S, K\. Dhamdhere, M\. Guo, U\. Gupta, A\. Eslami, M\. Schain, M\. Blokzijl, D\. Welling, D\. Orr, L\. Bolelli, N\. Perez\-Nieves, M\. Sirotenko, A\. Prasad, A\. Kar, B\. D\. B\. Pigem, T\. Terzi, G\. Weisz, D\. Ghosh, A\. Mavalankar, D\. Madeka, K\. Daugaard, H\. Adam, V\. Shah, D\. Berman, M\. Tran, S\. Baker, E\. Andrejczuk, G\. Chole, G\. Raboshchuk, M\. Mirzazadeh, T\. Kagohara, S\. Wu, C\. Schallhart, B\. Orlando, C\. Wang, A\. Rrustemi, H\. Xiong, H\. Liu, A\. Vezer, N\. Ramsden, S\. Chang, S\. Mudgal, Y\. Li, N\. Vieillard, Y\. Hoshen, F\. Ahmad, A\. Slone, A\. Hua, N\. Potikha, M\. Rossini, J\. Stritar, S\. Prakash, Z\. Wang, X\. Dong, A\. Nazari, E\. Nehoran, K\. Tekelioglu, Y\. Li, K\. Badola, T\. Funkhouser, Y\. Li, V\. Yerram, R\. Ganeshan, D\. Formoso, K\. Langner, T\. Shi, H\. Li, Y\. Yamamori, A\. Panda, A\. Saade, A\. S\. Scarpati, C\. Breaux, C\. Carey, Z\. Zhou, C\. Hsieh, S\. Bridgers, A\. Butryna, N\. Gupta, V\. Tulsyan, S\. Woo, E\. Eltyshev, W\. Grathwohl, C\. Parks, S\. Benjamin, R\. Panigrahy, S\. Dodhia, D\. D\. Freitas, C\. Sauer, W\. Song, F\. Alet, J\. Tolins, C\. Paduraru, X\. Zhou, B\. Albert, Z\. Zhang, L\. Shu, M\. Bansal, S\. Nguyen, A\. Globerson, O\. Xiao, J\. Manyika, T\. Hennigan, R\. Rong, J\. Matak, A\. Bakalov, A\. Sharma, D\. Sinopalnikov, A\. Pierson, S\. Roller, G\. Brown, M\. Gao, T\. Fukuzawa, A\. Ghafouri, K\. Vassigh, I\. Barr, Z\. Wang, A\. Korsun, R\. Jayaram, L\. Ren, T\. Zaman, S\. Khan, Y\. Lunts, D\. Deutsch, D\. Uthus, N\. Katz, M\. Samsikova, A\. Khalifa, N\. Sethi, J\. Sun, L\. Tang, U\. Alon, X\. Luo, D\. Yu, A\. Nayyar, B\. Petrini, W\. Truong, V\. Hellendoorn, N\. Chinaev, C\. Alberti, W\. Wang, J\. Hu, V\. Mirrokni, A\. Balashankar, A\. Aharon, A\. Mehta, A\. Iscen, J\. Kready, L\. Manning, A\. Mohananey, Y\. Chen, A\. Tripathi, A\. Wu, I\. Petrovski, D\. Hwang, M\. Baeuml, S\. Chandrakaladharan, Y\. Liu, R\. Coaguila, M\. Chen, S\. Ma, P\. Tafti, S\. Tatineni, T\. Spitz, J\. Ye, P\. Vicol, M\. Rosca, A\. Puigdomènech, Z\. Yahav, S\. Ghemawat, H\. Lin, P\. Kirk, Z\. Nabulsi, S\. Brin, B\. Bohnet, K\. Caluwaerts, A\. S\. Veerubhotla, D\. Zheng, Z\. Dai, P\. Petrov, Y\. Xu, R\. Mehran, Z\. Xu, L\. Zintgraf, J\. Choi, S\. A\. Hombaiah, R\. Thoppilan, S\. Reddi, L\. Lew, L\. Li, K\. Webster, K\. Sawhney, L\. Lamprou, S\. Shakeri, M\. Lunayach, J\. Chen, S\. Bagri, A\. Salcianu, Y\. Chen, Y\. Donchev, C\. Magister, S\. Nørly, V\. Rodrigues, T\. Izo, H\. Noga, J\. Zou, T\. Köppe, W\. Zhou, K\. Lee, X\. Long, D\. Eisenbud, A\. Chen, C\. Schenck, C\. M\. To, P\. Zhong, E\. Taropa, M\. Truong, O\. Levy, D\. Martins, Z\. Zhang, C\. Semturs, K\. Zhang, A\. Yakubovich, P\. Moreno, L\. McConnaughey, D\. Lu, S\. Redmond, L\. Weerts, Y\. Bitton, T\. Refice, N\. Lacasse, A\. Conmy, C\. Tallec, J\. Odell, H\. Forbes\-Pollard, A\. Socala, J\. Hoech, P\. Kohli, A\. Walton, R\. Wang, M\. Sazanovich, K\. Zhu, A\. Kapishnikov, R\. Galt, M\. Denton, B\. Murdoch, C\. Sikora, K\. Mohamed, W\. Wei, U\. First, T\. McConnell, L\. C\. Cobo, J\. Qin, T\. Avrahami, D\. Balle, Y\. Watanabe, A\. Louis, A\. Kraft, S\. Ariafar, Y\. Gu, E\. Rives, C\. Yoon, A\. Rusu, J\. Cobon\-Kerr, C\. Hahn, J\. Luo, Yuvein, Zhu, N\. Ahuja, R\. Benenson, R\. L\. Kaufman, H\. Yu, L\. Hightower, J\. Zhang, D\. Ni, L\. A\. Hendricks, G\. Wang, G\. Yona, L\. Jain, P\. Barrio, S\. Bhupatiraju, S\. Velusamy, A\. Dafoe, S\. Riedel, T\. Thomas, Z\. Yuan, M\. Bellaiche, S\. Panthaplackel, K\. Kloboves, S\. Jauhari, C\. Akbulut, T\. Davchev, E\. Gladchenko, D\. Madras, A\. Chuklin, T\. Hill, Q\. Yuan, M\. Madhavan, L\. Leonhard, D\. Scandinaro, Q\. Chen, N\. Niu, A\. Douillard, B\. Damoc, Y\. Onoe, F\. Pedregosa, F\. Bertsch, C\. Leichner, J\. Pagadora, J\. Malmaud, S\. Ponda, A\. Twigg, O\. Duzhyi, J\. Shen, M\. Wang, R\. Garg, J\. Chen, U\. Evci, J\. Lee, L\. Liu, K\. Kojima, M\. Yamaguchi, A\. Rajendran, A\. Piergiovanni, V\. K\. Rajendran, M\. Fornoni, G\. Ibagon, H\. Ragan, S\. M\. Khan, J\. Blitzer, A\. Bunner, G\. Sun, T\. Kosakai, S\. Lundberg, N\. Elue, K\. Guu, S\. Park, J\. Park, A\. Narayanaswamy, C\. Wu, J\. Mudigonda, T\. Cohn, H\. Mu, R\. Kumar, L\. Graesser, Y\. Zhang, R\. Killam, V\. Zhuang, M\. Giménez, W\. A\. Jishi, R\. Ley\-Wild, A\. Zhai, K\. Osawa, D\. Cedillo, J\. Liu, M\. Upadhyay, M\. Sieniek, R\. Sharma, T\. Paine, A\. Angelova, S\. Addepalli, C\. Parada, K\. Majumder, A\. Lamp, S\. Kumar, X\. Deng, A\. Myaskovsky, T\. Sabolić, J\. Dudek, S\. York, F\. de Chaumont Quitry, J\. Nie, D\. Cattle, A\. Gunjan, B\. Piot, W\. Khawaja, S\. Bang, S\. Wang, S\. Khodadadeh, R\. R, P\. Rawlani, R\. Powell, K\. Lee, J\. Griesser, G\. Oh, C\. Magalhaes, Y\. Li, S\. Tokumine, H\. N\. Vogel, D\. Hsu, A\. BC, D\. Jindal, M\. Cohen, Z\. Yang, J\. Yuan, D\. de Cesare, T\. Bruguier, J\. Xu, M\. Roy, A\. Jacovi, D\. Belov, R\. Arya, P\. Meadowlark, S\. Cohen\-Ganor, W\. Ye, P\. Morris\-Suzuki, P\. Banzal, G\. Song, P\. Ponnuramu, F\. Zhang, G\. Scrivener, S\. Zaiem, A\. R\. Rochman, K\. Han, B\. Ghazi, K\. Lee, S\. Drath, D\. Suo, A\. Girgis, P\. Shenoy, D\. Nguyen, D\. Eck, S\. Gupta, L\. Yan, J\. Carreira, A\. Gulati, R\. Sang, D\. Mirylenka, E\. Cooney, E\. Chou, M\. Ling, C\. Fan, B\. Coleman, G\. Tubone, R\. Kumar, J\. Baldridge, F\. Hernandez\-Campos, A\. Lazaridou, J\. Besley, I\. Yona, N\. Bulut, Q\. Wellens, A\. Pierigiovanni, J\. George, R\. Green, P\. Han, C\. Tao, G\. Clark, C\. You, A\. Abdolmaleki, J\. Fu, T\. Chen, A\. Chaugule, A\. Chandorkar, A\. Rahman, W\. Thompson, P\. Koanantakool, M\. Bernico, J\. Ren, A\. Vlasov, S\. Vassilvitskii, M\. Kula, Y\. Liang, D\. Kim, Y\. Huang, C\. Ye, D\. Lepikhin, and W\. Helmholz \(2025\)Gemini 2\.5: pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities\.External Links:2507\.06261,[Link](https://arxiv.org/abs/2507.06261)Cited by:[§3\.2](https://arxiv.org/html/2608.08283#S3.SS2.p1.1)\.
- P\. Fernandes, D\. Deutsch, M\. Finkelstein, P\. Riley, A\. Martins, G\. Neubig, A\. Garg, J\. Clark, M\. Freitag, and O\. Firat \(2023\)The Devil Is in the Errors: Leveraging Large Language Models for Fine\-grained Machine Translation Evaluation\.InProceedings of the Eighth Conference on Machine Translation,P\. Koehn, B\. Haddow, T\. Kocmi, and C\. Monz \(Eds\.\),Singapore,pp\. 1066–1083\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.wmt-1.100)Cited by:[§1](https://arxiv.org/html/2608.08283#S1.p3.1)\.
- T\. GLM, A\. Zeng, B\. Xu, B\. Wang, C\. Zhang, D\. Yin, D\. Rojas, G\. Feng, H\. Zhao, H\. Lai, H\. Yu, H\. Wang, J\. Sun, J\. Zhang, J\. Cheng, J\. Gui, J\. Tang, J\. Zhang, J\. Li, L\. Zhao, L\. Wu, L\. Zhong, M\. Liu, M\. Huang, P\. Zhang, Q\. Zheng, R\. Lu, S\. Duan, S\. Zhang, S\. Cao, S\. Yang, W\. L\. Tam, W\. Zhao, X\. Liu, X\. Xia, X\. Zhang, X\. Gu, X\. Lv, X\. Liu, X\. Liu, X\. Yang, X\. Song, X\. Zhang, Y\. An, Y\. Xu, Y\. Niu, Y\. Yang, Y\. Li, Y\. Bai, Y\. Dong, Z\. Qi, Z\. Wang, Z\. Yang, Z\. Du, Z\. Hou, and Z\. Wang \(2024\)ChatGLM: a family of large language models from glm\-130b to glm\-4 all tools\.External Links:2406\.12793Cited by:[§3\.2](https://arxiv.org/html/2608.08283#S3.SS2.p1.1)\.
- N\. M\. Guerreiro, R\. Rei, D\. v\. Stigt, L\. Coheur, P\. Colombo, and A\. F\. T\. Martins \(2024\)XCOMET: transparent machine translation evaluation through fine\-grained error detection\.Transactions of the Association for Computational Linguistics12,pp\. 979–995\.External Links:[Link](https://aclanthology.org/2024.tacl-1.54/),[Document](https://dx.doi.org/10.1162/tacl%5Fa%5F00683)Cited by:[§4](https://arxiv.org/html/2608.08283#S4.SS0.SSS0.Px2.p1.1)\.
- P\. Isabelle, C\. Cherry, and G\. Foster \(2017\)A challenge set approach to evaluating machine translation\.InProceedings of the 2017 Conference on Empirical Methods in Natural Language Processing,M\. Palmer, R\. Hwa, and S\. Riedel \(Eds\.\),Copenhagen, Denmark,pp\. 2486–2496\.External Links:[Link](https://aclanthology.org/D17-1263/),[Document](https://dx.doi.org/10.18653/v1/D17-1263)Cited by:[§1](https://arxiv.org/html/2608.08283#S1.p4.1)\.
- K\. Jin, D\. Zhao, and W\. Liu \(2023\)Morphological and semantic evaluation of Ancient Chinese machine translation\.InProceedings of the Ancient Language Processing Workshop,A\. Anderson, S\. Gordin, B\. Li, Y\. Liu, and M\. C\. Passarotti \(Eds\.\),Varna, Bulgaria,pp\. 96–102\.External Links:[Link](https://aclanthology.org/2023.alp-1.11/)Cited by:[§1](https://arxiv.org/html/2608.08283#S1.p1.1)\.
- J\. Juraska, D\. Deutsch, M\. Finkelstein, and M\. Freitag \(2024\)MetricX\-24: the Google submission to the WMT 2024 metrics shared task\.InProceedings of the Ninth Conference on Machine Translation,B\. Haddow, T\. Kocmi, P\. Koehn, and C\. Monz \(Eds\.\),Miami, Florida, USA,pp\. 492–504\.External Links:[Link](https://aclanthology.org/2024.wmt-1.35/),[Document](https://dx.doi.org/10.18653/v1/2024.wmt-1.35)Cited by:[§1](https://arxiv.org/html/2608.08283#S1.p3.1),[§4](https://arxiv.org/html/2608.08283#S4.SS0.SSS0.Px2.p1.1),[§5\.1](https://arxiv.org/html/2608.08283#S5.SS1.p6.1)\.
- M\. Karpinska, N\. Raj, K\. Thai, Y\. Song, A\. Gupta, and M\. Iyyer \(2022\)DEMETR: diagnosing evaluation metrics for translation\.InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing,Y\. Goldberg, Z\. Kozareva, and Y\. Zhang \(Eds\.\),Abu Dhabi, United Arab Emirates,pp\. 9540–9561\.External Links:[Link](https://aclanthology.org/2022.emnlp-main.649/),[Document](https://dx.doi.org/10.18653/v1/2022.emnlp-main.649)Cited by:[§1](https://arxiv.org/html/2608.08283#S1.p4.1),[§3](https://arxiv.org/html/2608.08283#S3.p1.1)\.
- R\. Knowles, S\. Larkin, and C\. Lo \(2024\)MSLC24: further challenges for metrics on a wide landscape of translation quality\.InProceedings of the Ninth Conference on Machine Translation,B\. Haddow, T\. Kocmi, P\. Koehn, and C\. Monz \(Eds\.\),Miami, Florida, USA,pp\. 475–491\.External Links:[Link](https://aclanthology.org/2024.wmt-1.34/),[Document](https://dx.doi.org/10.18653/v1/2024.wmt-1.34)Cited by:[§5\.1](https://arxiv.org/html/2608.08283#S5.SS1.p3.1)\.
- T\. Kocmi and C\. Federmann \(2023a\)GEMBA\-MQM: detecting translation quality error spans with GPT\-4\.InProceedings of the Eighth Conference on Machine Translation,P\. Koehn, B\. Haddow, T\. Kocmi, and C\. Monz \(Eds\.\),Singapore,pp\. 768–775\.External Links:[Link](https://aclanthology.org/2023.wmt-1.64/),[Document](https://dx.doi.org/10.18653/v1/2023.wmt-1.64)Cited by:[§4](https://arxiv.org/html/2608.08283#S4.SS0.SSS0.Px3.p1.1)\.
- T\. Kocmi and C\. Federmann \(2023b\)Large language models are state\-of\-the\-art evaluators of translation quality\.InProceedings of the 24th Annual Conference of the European Association for Machine Translation,M\. Nurminen, J\. Brenner, M\. Koponen, S\. Latomaa, M\. Mikhailov, F\. Schierl, T\. Ranasinghe, E\. Vanmassenhove, S\. A\. Vidal, N\. Aranberri, M\. Nunziatini, C\. P\. Escartín, M\. Forcada, M\. Popovic, C\. Scarton, and H\. Moniz \(Eds\.\),Tampere, Finland,pp\. 193–203\.External Links:[Link](https://aclanthology.org/2023.eamt-1.19/)Cited by:[§4](https://arxiv.org/html/2608.08283#S4.SS0.SSS0.Px3.p1.1)\.
- T\. Kocmi and C\. Federmann \(2023c\)Large Language Models Are State\-of\-the\-Art Evaluators of Translation Quality\.InProceedings of the 24th Annual Conference of the European Association for Machine Translation,M\. Nurminen, J\. Brenner, M\. Koponen, S\. Latomaa, M\. Mikhailov, F\. Schierl, T\. Ranasinghe, E\. Vanmassenhove, S\. A\. Vidal, N\. Aranberri, M\. Nunziatini, C\. P\. Escartín, M\. Forcada, M\. Popovic, C\. Scarton, and H\. Moniz \(Eds\.\),Tampere, Finland,pp\. 193–203\.Cited by:[§1](https://arxiv.org/html/2608.08283#S1.p3.1)\.
- A\. Lavie, G\. Hanneman, S\. Agrawal, D\. Kanojia, C\. Lo, V\. Zouhar, F\. Blain, C\. Zerva, E\. Avramidis, S\. Deoghare, A\. Sindhujan, J\. Wang, D\. I\. Adelani, B\. Thompson, T\. Kocmi, M\. Freitag, and D\. Deutsch \(2025\)Findings of the WMT25 Shared Task on Automated Translation Evaluation Systems: Linguistic Diversity is Challenging and References Still Help\.InProceedings of the Tenth Conference on Machine Translation,B\. Haddow, T\. Kocmi, P\. Koehn, and C\. Monz \(Eds\.\),Suzhou, China,pp\. 436–483\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.wmt-1.24),ISBN 979\-8\-89176\-341\-8Cited by:[§4](https://arxiv.org/html/2608.08283#S4.p1.1),[§5\.1](https://arxiv.org/html/2608.08283#S5.SS1.p3.1)\.
- H\. Li, Y\. Zhang, F\. Koto, Y\. Yang, H\. Zhao, Y\. Gong, N\. Duan, and T\. Baldwin \(2024\)CMMLU: Measuring massive multitask language understanding in Chinese\.InFindings of the Association for Computational Linguistics: ACL 2024,L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 11260–11285\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.findings-acl.671)Cited by:[§2](https://arxiv.org/html/2608.08283#S2.p3.1)\.
- J\. M\. MacKenzie \(2024\)Alexander chow \(ed\.\)\. 2021\. scottish missions to china: commemorating the legacy of james legge \(1815–1897\)\.Studies in World Christianity30\(3\),pp\. 399–400\.External Links:[Document](https://dx.doi.org/10.3366/swc.2024.0486),[Link](https://doi.org/10.3366/swc.2024.0486),https://doi\.org/10\.3366/swc\.2024\.0486Cited by:[§3\.1](https://arxiv.org/html/2608.08283#S3.SS1.p1.1)\.
- S\. Nehrdich, M\. Bingenheimer, J\. Brody, and K\. Keutzer \(2023\)MITRA\-zh: an efficient, open machine translation solution for buddhist Chinese\.InProceedings of the Joint 3rd International Conference on Natural Language Processing for Digital Humanities and 8th International Workshop on Computational Linguistics for Uralic Languages,M\. Hämäläinen, E\. Öhman, F\. Pirinen, K\. Alnajjar, S\. Miyagawa, Y\. Bizzoni, N\. Partanen, and J\. Rueter \(Eds\.\),Tokyo, Japan,pp\. 266–277\.External Links:[Link](https://aclanthology.org/2023.nlp4dh-1.29/)Cited by:[§2](https://arxiv.org/html/2608.08283#S2.p3.1)\.
- S\. Nehrdich, A\. Chen, M\. Bingenheimer, L\. Huang, R\. Tang, X\. Wei, L\. Zhu, and K\. Keutzer \(2025\)MITRA\-zh\-eval: Using a Buddhist Chinese Language Evaluation Dataset to Assess Machine Translation and Evaluation Metrics\.InProceedings of the 5th International Conference on Natural Language Processing for Digital Humanities,M\. Hämäläinen, E\. Öhman, Y\. Bizzoni, S\. Miyagawa, and K\. Alnajjar \(Eds\.\),Albuquerque, USA,pp\. 129–137\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.nlp4dh-1.12),ISBN 979\-8\-89176\-234\-3Cited by:[§1](https://arxiv.org/html/2608.08283#S1.p3.1)\.
- S\. Nehrdich and K\. Keutzer \(2026\)MITRA: a large\-scale parallel corpus and multilingual pretrained language model for machine translation and semantic retrieval for pāli, sanskrit, buddhist chinese, and tibetan\.External Links:2601\.06400,[Link](https://arxiv.org/abs/2601.06400)Cited by:[§1](https://arxiv.org/html/2608.08283#S1.p1.1),[§2](https://arxiv.org/html/2608.08283#S2.p4.1)\.
- OpenAI, J\. Achiam, S\. Adler, S\. Agarwal, L\. Ahmad, I\. Akkaya, F\. L\. Aleman, D\. Almeida, J\. Altenschmidt, S\. Altman, S\. Anadkat, R\. Avila, I\. Babuschkin, S\. Balaji, V\. Balcom, P\. Baltescu, H\. Bao, M\. Bavarian, J\. Belgum, I\. Bello, J\. Berdine, G\. Bernadett\-Shapiro, C\. Berner, L\. Bogdonoff, O\. Boiko, M\. Boyd, A\. Brakman, G\. Brockman, T\. Brooks, M\. Brundage, K\. Button, T\. Cai, R\. Campbell, A\. Cann, B\. Carey, C\. Carlson, R\. Carmichael, B\. Chan, C\. Chang, F\. Chantzis, D\. Chen, S\. Chen, R\. Chen, J\. Chen, M\. Chen, B\. Chess, C\. Cho, C\. Chu, H\. W\. Chung, D\. Cummings, J\. Currier, Y\. Dai, C\. Decareaux, T\. Degry, N\. Deutsch, D\. Deville, A\. Dhar, D\. Dohan, S\. Dowling, S\. Dunning, A\. Ecoffet, A\. Eleti, T\. Eloundou, D\. Farhi, L\. Fedus, N\. Felix, S\. P\. Fishman, J\. Forte, I\. Fulford, L\. Gao, E\. Georges, C\. Gibson, V\. Goel, T\. Gogineni, G\. Goh, R\. Gontijo\-Lopes, J\. Gordon, M\. Grafstein, S\. Gray, R\. Greene, J\. Gross, S\. S\. Gu, Y\. Guo, C\. Hallacy, J\. Han, J\. Harris, Y\. He, M\. Heaton, J\. Heidecke, C\. Hesse, A\. Hickey, W\. Hickey, P\. Hoeschele, B\. Houghton, K\. Hsu, S\. Hu, X\. Hu, J\. Huizinga, S\. Jain, S\. Jain, J\. Jang, A\. Jiang, R\. Jiang, H\. Jin, D\. Jin, S\. Jomoto, B\. Jonn, H\. Jun, T\. Kaftan, Ł\. Kaiser, A\. Kamali, I\. Kanitscheider, N\. S\. Keskar, T\. Khan, L\. Kilpatrick, J\. W\. Kim, C\. Kim, Y\. Kim, J\. H\. Kirchner, J\. Kiros, M\. Knight, D\. Kokotajlo, Ł\. Kondraciuk, A\. Kondrich, A\. Konstantinidis, K\. Kosic, G\. Krueger, V\. Kuo, M\. Lampe, I\. Lan, T\. Lee, J\. Leike, J\. Leung, D\. Levy, C\. M\. Li, R\. Lim, M\. Lin, S\. Lin, M\. Litwin, T\. Lopez, R\. Lowe, P\. Lue, A\. Makanju, K\. Malfacini, S\. Manning, T\. Markov, Y\. Markovski, B\. Martin, K\. Mayer, A\. Mayne, B\. McGrew, S\. M\. McKinney, C\. McLeavey, P\. McMillan, J\. McNeil, D\. Medina, A\. Mehta, J\. Menick, L\. Metz, A\. Mishchenko, P\. Mishkin, V\. Monaco, E\. Morikawa, D\. Mossing, T\. Mu, M\. Murati, O\. Murk, D\. Mély, A\. Nair, R\. Nakano, R\. Nayak, A\. Neelakantan, R\. Ngo, H\. Noh, L\. Ouyang, C\. O’Keefe, J\. Pachocki, A\. Paino, J\. Palermo, A\. Pantuliano, G\. Parascandolo, J\. Parish, E\. Parparita, A\. Passos, M\. Pavlov, A\. Peng, A\. Perelman, F\. de Avila Belbute Peres, M\. Petrov, H\. P\. de Oliveira Pinto, Michael, Pokorny, M\. Pokrass, V\. H\. Pong, T\. Powell, A\. Power, B\. Power, E\. Proehl, R\. Puri, A\. Radford, J\. Rae, A\. Ramesh, C\. Raymond, F\. Real, K\. Rimbach, C\. Ross, B\. Rotsted, H\. Roussez, N\. Ryder, M\. Saltarelli, T\. Sanders, S\. Santurkar, G\. Sastry, H\. Schmidt, D\. Schnurr, J\. Schulman, D\. Selsam, K\. Sheppard, T\. Sherbakov, J\. Shieh, S\. Shoker, P\. Shyam, S\. Sidor, E\. Sigler, M\. Simens, J\. Sitkin, K\. Slama, I\. Sohl, B\. Sokolowsky, Y\. Song, N\. Staudacher, F\. P\. Such, N\. Summers, I\. Sutskever, J\. Tang, N\. Tezak, M\. B\. Thompson, P\. Tillet, A\. Tootoonchian, E\. Tseng, P\. Tuggle, N\. Turley, J\. Tworek, J\. F\. C\. Uribe, A\. Vallone, A\. Vijayvergiya, C\. Voss, C\. Wainwright, J\. J\. Wang, A\. Wang, B\. Wang, J\. Ward, J\. Wei, C\. Weinmann, A\. Welihinda, P\. Welinder, J\. Weng, L\. Weng, M\. Wiethoff, D\. Willner, C\. Winter, S\. Wolrich, H\. Wong, L\. Workman, S\. Wu, J\. Wu, M\. Wu, K\. Xiao, T\. Xu, S\. Yoo, K\. Yu, Q\. Yuan, W\. Zaremba, R\. Zellers, C\. Zhang, M\. Zhang, S\. Zhao, T\. Zheng, J\. Zhuang, W\. Zhuk, and B\. Zoph \(2024\)GPT\-4 technical report\.External Links:2303\.08774,[Link](https://arxiv.org/abs/2303.08774)Cited by:[§3\.2](https://arxiv.org/html/2608.08283#S3.SS2.p1.1)\.
- K\. Papineni, S\. Roukos, T\. Ward, and W\. Zhu \(2002a\)Bleu: a method for automatic evaluation of machine translation\.InProceedings of the 40th Annual Meeting of the Association for Computational Linguistics,P\. Isabelle, E\. Charniak, and D\. Lin \(Eds\.\),Philadelphia, Pennsylvania, USA,pp\. 311–318\.External Links:[Link](https://aclanthology.org/P02-1040/),[Document](https://dx.doi.org/10.3115/1073083.1073135)Cited by:[§1](https://arxiv.org/html/2608.08283#S1.p3.1)\.
- K\. Papineni, S\. Roukos, T\. Ward, and W\. Zhu \(2002b\)BLEU: a method for automatic evaluation of machine translation\.InProceedings of the 40th Annual Meeting of the Association for Computational Linguistics,Philadelphia, PA\.Cited by:[§1](https://arxiv.org/html/2608.08283#S1.p3.1),[§4](https://arxiv.org/html/2608.08283#S4.SS0.SSS0.Px1.p1.1)\.
- A\. Peyraube \(2004\)Ancient Chinese\.InThe Cambridge Encyclopedia of the World’s Ancient Languages,R\. D\. Woodard \(Ed\.\),pp\. 988–1014\.Cited by:[§5\.3](https://arxiv.org/html/2608.08283#S5.SS3.p1.1)\.
- M\. Popovic \(2015\)chrF: character n\-gram F\-score for automatic MT evaluation\.InProceedings of the 10th Workshop on Statistical Machine Translation \(WMT\-15\),pp\. 392–395\.Cited by:[§1](https://arxiv.org/html/2608.08283#S1.p3.1),[§4](https://arxiv.org/html/2608.08283#S4.SS0.SSS0.Px1.p1.1)\.
- R\. Rei, C\. Stewart, A\. C\. Farinha, and A\. Lavie \(2020\)COMET: A Neural Framework for MT Evaluation\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),Online,pp\. 2685–2702\.External Links:[Document](https://dx.doi.org/10.18653/v1/2020.emnlp-main.213)Cited by:[§1](https://arxiv.org/html/2608.08283#S1.p3.1),[§4](https://arxiv.org/html/2608.08283#S4.SS0.SSS0.Px2.p1.1)\.
- A\. Schonebaum, A\. George, D\. Lattimore, H\. Hsiao\-chen, J\. Zeitlin, K\. Mei, L\. Lening, M\. B\. Wan, P\. Hanan, P\. Rouzer, R\. Llamas, S\. Wei, and X\. Tian \(2024\)Introduction to Classical Chinese\.Cited by:[§2](https://arxiv.org/html/2608.08283#S2.p2.1)\.
- J\. Son, J\. Jin, H\. Yoo, J\. Bak, K\. Cho, and A\. Oh \(2022\)Translating hanja historical documents to contemporary Korean and English\.InFindings of the Association for Computational Linguistics: EMNLP 2022,Y\. Goldberg, Z\. Kozareva, and Y\. Zhang \(Eds\.\),Abu Dhabi, United Arab Emirates,pp\. 1260–1272\.External Links:[Link](https://aclanthology.org/2022.findings-emnlp.91/),[Document](https://dx.doi.org/10.18653/v1/2022.findings-emnlp.91)Cited by:[§1](https://arxiv.org/html/2608.08283#S1.p1.1)\.
- S\. Song, H\. Yoo, J\. Jin, K\. Cho, and A\. Oh \(2025\)HERITAGE: an end\-to\-end web platform for processing korean historical documents in hanja\.External Links:2501\.11951,[Link](https://arxiv.org/abs/2501.11951)Cited by:[§1](https://arxiv.org/html/2608.08283#S1.p1.1)\.
- D\. Sturgeon \(2019\)Chinese text project: a dynamic digital library of premodern chinese\.Digital Scholarship in the Humanities\.External Links:[Document](https://dx.doi.org/10.1093/llc/fqy032),[Link](https://ctext.org/)Cited by:[§1](https://arxiv.org/html/2608.08283#S1.p1.1)\.
- D\. Sturgeon \(2021\)Chinese Text Project: A dynamic digital library of premodern Chinese\.Digital Scholarship in the Humanities36\(Supplement\_1\),pp\. i101–i112\.External Links:ISSN 2055\-7671,[Document](https://dx.doi.org/10.1093/llc/fqz046)Cited by:[§2](https://arxiv.org/html/2608.08283#S2.p1.1),[§2](https://arxiv.org/html/2608.08283#S2.p3.1),[§3\.1](https://arxiv.org/html/2608.08283#S3.SS1.p1.1)\.
- \[37\]A\. TeamThe claude 3 model family: opus, sonnet, haiku\.External Links:[Link](https://api.semanticscholar.org/CorpusID:268232499)Cited by:[§3\.2](https://arxiv.org/html/2608.08283#S3.SS2.p1.1)\.
- M\. Volk, D\. P\. Fischer, L\. Fischer, P\. Scheurer, and P\. B\. Ströbel \(2024\)LLM\-based machine translation and summarization for Latin\.InProceedings of the Third Workshop on Language Technologies for Historical and Ancient Languages \(LT4HALA\) @ LREC\-COLING\-2024,R\. Sprugnoli and M\. Passarotti \(Eds\.\),Torino, Italia,pp\. 122–128\.External Links:[Link](https://aclanthology.org/2024.lt4hala-1.15/)Cited by:[§1](https://arxiv.org/html/2608.08283#S1.p1.1)\.
- H\. Wang, H\. Shimizu, and D\. Kawahara \(2023\)Kanbun\-LM: Reading and Translating Classical Chinese in Japanese Methods by Language Models\.InFindings of the Association for Computational Linguistics: ACL 2023,A\. Rogers, J\. Boyd\-Graber, and N\. Okazaki \(Eds\.\),Toronto, Canada,pp\. 8589–8601\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.findings-acl.545)Cited by:[§2](https://arxiv.org/html/2608.08283#S2.p3.1)\.
- A\. Warstadt, A\. Parrish, H\. Liu, A\. Mohananey, W\. Peng, S\. Wang, and S\. R\. Bowman \(2020\)BLiMP: the benchmark of linguistic minimal pairs for English\.Transactions of the Association for Computational Linguistics8,pp\. 377–392\.External Links:[Link](https://aclanthology.org/2020.tacl-1.25/),[Document](https://dx.doi.org/10.1162/tacl%5Fa%5F00321)Cited by:[§1](https://arxiv.org/html/2608.08283#S1.p4.1)\.
- A\. Yang, B\. Xiao, B\. Wang, B\. Zhang, C\. Bian, C\. Yin, C\. Lv, D\. Pan, D\. Wang, D\. Yan, F\. Yang, F\. Deng, F\. Wang, F\. Liu, G\. Ai, G\. Dong, H\. Zhao, H\. Xu, H\. Sun, H\. Zhang, H\. Liu, J\. Ji, J\. Xie, J\. Dai, K\. Fang, L\. Su, L\. Song, L\. Liu, L\. Ru, L\. Ma, M\. Wang, M\. Liu, M\. Lin, N\. Nie, P\. Guo, R\. Sun, T\. Zhang, T\. Li, T\. Li, W\. Cheng, W\. Chen, X\. Zeng, X\. Wang, X\. Chen, X\. Men, X\. Yu, X\. Pan, Y\. Shen, Y\. Wang, Y\. Li, Y\. Jiang, Y\. Gao, Y\. Zhang, Z\. Zhou, and Z\. Wu \(2025\)Baichuan 2: open large\-scale language models\.External Links:2309\.10305,[Link](https://arxiv.org/abs/2309.10305)Cited by:[§3\.2](https://arxiv.org/html/2608.08283#S3.SS2.p1.1)\.
- A\. Yang, B\. Yang, B\. Hui, B\. Zheng, B\. Yu, C\. Zhou, C\. Li, C\. Li, D\. Liu, F\. Huang, G\. Dong, H\. Wei, H\. Lin, J\. Tang, J\. Wang, J\. Yang, J\. Tu, J\. Zhang, J\. Ma, J\. Yang, J\. Xu, J\. Zhou, J\. Bai, J\. He, J\. Lin, K\. Dang, K\. Lu, K\. Chen, K\. Yang, M\. Li, M\. Xue, N\. Ni, P\. Zhang, P\. Wang, R\. Peng, R\. Men, R\. Gao, R\. Lin, S\. Wang, S\. Bai, S\. Tan, T\. Zhu, T\. Li, T\. Liu, W\. Ge, X\. Deng, X\. Zhou, X\. Ren, X\. Zhang, X\. Wei, X\. Ren, X\. Liu, Y\. Fan, Y\. Yao, Y\. Zhang, Y\. Wan, Y\. Chu, Y\. Liu, Z\. Cui, Z\. Zhang, Z\. Guo, and Z\. Fan \(2024\)Qwen2 technical report\.External Links:2407\.10671,[Link](https://arxiv.org/abs/2407.10671)Cited by:[§3\.2](https://arxiv.org/html/2608.08283#S3.SS2.p1.1)\.
- M\. You, D\. Wong, J\. Zhang, and K\. Lan \(2025\)How well can state\-of\-the\-art machine translation systems render a 16th\-century chinese novel?\.Cadernos de Tradução45,pp\. 1–24\.External Links:[Document](https://dx.doi.org/10.5007/2175-7968.2025.e108394)Cited by:[§2](https://arxiv.org/html/2608.08283#S2.p3.1)\.
- B\. Zhou, Q\. Chen, T\. Wang, X\. Zhong, and Y\. Zhang \(2023\)WYWEB: A NLP Evaluation Benchmark For Classical Chinese\.InFindings of the Association for Computational Linguistics: ACL 2023,A\. Rogers, J\. Boyd\-Graber, and N\. Okazaki \(Eds\.\),Toronto, Canada,pp\. 3294–3319\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.findings-acl.204)Cited by:[§2](https://arxiv.org/html/2608.08283#S2.p3.1)\.
- V\. Zouhar, P\. Chen, T\. K\. Lam, N\. Moghe, and B\. Haddow \(2024\)Pitfalls and outlooks in using COMET\.InProceedings of the Ninth Conference on Machine Translation,B\. Haddow, T\. Kocmi, P\. Koehn, and C\. Monz \(Eds\.\),Miami, Florida, USA,pp\. 1272–1288\.External Links:[Link](https://aclanthology.org/2024.wmt-1.121/),[Document](https://dx.doi.org/10.18653/v1/2024.wmt-1.121)Cited by:[§4](https://arxiv.org/html/2608.08283#S4.p1.1),[§5\.1](https://arxiv.org/html/2608.08283#S5.SS1.p3.1)\.
## Appendix APrompt Used for Example Generation
All LLM\-generated perturbations were produced using GPT\-4o\-mini\. We selected this model on the grounds of cost efficiency and reproducibility: GPT\-4o\-mini offers sufficient instruction\-following capability for the constrained perturbation tasks defined in Section 3\.2, injecting a single targeted error into a translation\.
Prompt used for OMITTED SUBJECTPrompt:Given the following Classical Chinese text: ”\{source\_segment\}”And its English translation: ”\{machine\_translation\_text\}”Create a perturbed version of the English translation that contains an OMITTED SUBJECT error\.In Classical Chinese, subjects are often omitted when context makes them clear\. However, when translating to English, these subjects must be explicitly stated\. An omitted subject error occurs when the translator fails to properly identify or include the subject, making the English translation ambiguous or unclear\.The perturbed version should:1\.Remove or make ambiguous the subject of at least one clause or sentence2\.Create confusion about who or what is performing the action3\.Maintain the overall structure and most of the vocabulary of the original reference including capitalization, spacing, punctuation, and chinese characters if present in the translation4\.Be a realistic error that could occur in translationExample:•Original: ”The Master said, ’The gentleman studies virtue\.’”•Perturbed: ”Said, ’The gentleman studies\.’” \(subject omitted\)IMPORTANT: You must respond with ONLY valid JSON in this exact format:”\{perturbed\_translation\}”: ”the perturbed English translation with omitted subject”,”\{error\_description\}”: ”brief description of how the subject was omitted or made unclear”Do not include any text before or after the JSON\. Only return the JSON object\.
Prompt used for OMITTED OBJECTPrompt:Given the following Classical Chinese text: ”\{source\_segment\}”And its English translation: ”\{machine\_translation\_text\}”Create a perturbed version of the English translation that contains an OMITTED OBJECT error\.In Classical Chinese, objects can be omitted when context makes them clear\. When translating to English, these objects must often be explicitly stated for clarity\. An omitted object error occurs when the translator fails to properly identify or include the direct or indirect object, making the English translation incomplete or ambiguous\.The perturbed version should:1\.Remove or make ambiguous the object of at least one verb2\.Create confusion about what is being acted upon3\.Maintain the overall structure and most of the vocabulary of the original translation including capitalization, spacing, punctuation, and chinese characters if present in the translation etc\.4\.Be a realistic error that could occur in translationExample:•Original: ”The Master taught the students virtue\.”•Perturbed: ”The Master taught virtue\.” \(object ’students’ omitted\)IMPORTANT: You must respond with ONLY valid JSON in this exact format:”\{perturbed\_translation\}”: ”the perturbed English translation with omitted object”,”\{error\_description\}”: ”brief description of how the object was omitted or made unclear”Do not include any text before or after the JSON\. Only return the JSON object\.
Prompt used for INCORRECT LEXICAL TRANSLATIONPrompt:Given the following Classical Chinese source: ”\{source\_segment\}”And its English translation: ”\{machine\_translation\_text\}”Create a perturbed version of the English translation that contains an ACCURACY ERROR\.Accuracy errors involve factual or semantic inaccuracies in translation\. These can include:•Incorrect translation of specific terms of concepts•Misrepresentation of quantities, measurements, or relationships•Wrong interpretation of causal or logical connections•Factual errors about people, places, or events•Semantic shifts that change the core meaningThe perturbed version should:1\.Introduce a factual or semantic inaccuracy while maintaining plausible structure2\.Change the meaning subtly enough to seem like an honest mistake3\.Preserve most of the original wording and structure \(including capitalization, spacing, punctuation, and chinese characters if present in the translation etc\.\)4\.Create a realistic translation error that could mislead readersExample:•Original: ”He ruled for thirty years and established peace\.”•Perturbed: ”He ruled for three years and established peace\.” \(quantity error\)IMPORTANT: You must respond with ONLY valid JSON in this exact format:”\{perturbed\_translation\}”: ”the perturbed English translation with accuracy error”,”\{error\_description\}”: ”brief description of the specific accuracy error introduced”Do not include any text before or after the JSON\. Only return the JSON object\.
Prompt used for INCORRECT TENSE REALIZATIONPrompt:Given the following Classical Chinese text: ”\{source\_segment\}”And its English translation: ”\{machine\_translation\_text\}”Create a perturbed version of the English translation that contains a TENSE ERROR\.Classical Chinese has different temporal markers than English, and tense must often be inferred from context\. Tense errors occur when the translator uses incorrect verb tense \(past, present, future\)\.The perturbed version should:1\.Change the tense of at least two verbs to an incorrect tense2\.Maintain the overall structure and vocabularyExample:•Original: ”The king ruled wisely and his people prospered\.”•Perturbed: ”The king rules wisely and his people prospered\.” \(tense inconsistency\)IMPORTANT: You must respond with ONLY valid JSON in this exact format:”\{perturbed\_translation\}”: ”the perturbed English translation with tense error”,”\{error\_description\}”: ”brief description of the specific tense error introduced”Do not include any text before or after the JSON\. Only return the JSON object\.
Prompt used for UNINTENDED TRANSLATION INTO MODERN CHINESEPrompt:Translate the following ”\{english\_machine\_translation\}”222English machine translation of the Classical Chinese originalto Chinese\. Put your translation between $$ signs \(e\.g\., $$your translation here$$\)
Prompt used for MODERN CHINESE SENSE SUBSTITUTIONPrompt:Given the following Classical Chinese source: ”\{source\_segment\}”And its English translation: ”\{machine\_translation\_text\}”The Classical Chinese text contains the character ”\{character\}” which has different meanings:•Ancient meaning: ”\{ancient\_meaning\}”•Modern meaning: ”\{modern\_meaning\}”Analyze the translation and identify where the ancient meaning of ”\{character\}” appears\. Then create a perturbed version of the translation by replacing the ancient meaning with the modern meaning\.The perturbed version should:1\.Replace the ancient meaning with the modern meaning naturally2\.Maintain the overall structure and flow of the original translation3\.Preserve the meaning of the rest of the sentenceExample:•If character ”君” has ancient meaning ”gentleman” and modern meaning ”monarch”•And translation contains ”The gentleman has wine”•Then perturbed version should be ”The monarch has wine”IMPORTANT: You must respond with ONLY valid JSON in this exact format:”\{perturbed\_translation\}”: ”the translation with ancient meaning replaced by modern meaning”,”\{character\_found\}”: true/false,”\{ancient\_usage\}”: ”the specific ancient meaning usage found in translation”,”\{modern\_replacement\}”: ”the modern meaning replacement made”,”\{explanation\}”: ”brief explanation of the change made”Do not include any text before or after the JSON\. Only return the JSON object\.
## Appendix BMetrics Implementation Details
Metric\# ParamsLanguagesSourcestring\-based metricsBLEU–any[SacreBLEU](https://github.com/mjpost/sacrebleu)ChrF–any[SacreBLEU](https://github.com/mjpost/sacrebleu)learned neural metricsCOMET580580M94[Unbabel/wmt22\-comet\-da](https://huggingface.co/Unbabel/wmt22-comet-da)COMET\-REF\-ONLY580580M94[Unbabel/wmt22\-comet\-da](https://huggingface.co/Unbabel/wmt22-comet-da)XCOMET3\.53\.5B94[Unbabel/XCOMET\-XL](https://huggingface.co/Unbabel/XCOMET-XL)XCOMET\-QE3\.53\.5B94[Unbabel/XCOMET\-XL](https://huggingface.co/Unbabel/XCOMET-XL)MetricX\-243\.73\.7B49[google/metricx\-24\-hybrid\-xl\-v2p6](https://huggingface.co/google/metricx-24-hybrid-xl-v2p6)MetricX\-24\-REF\-ONLY3\.73\.7B49[google/metricx\-24\-hybrid\-xl\-v2p6](https://huggingface.co/google/metricx-24-hybrid-xl-v2p6)MetricX\-24\-QE3\.73\.7B49[google/metricx\-24\-hybrid\-xl\-v2p6](https://huggingface.co/google/metricx-24-hybrid-xl-v2p6)LLM\-as\-a\-judgeGEMBA\-DA––[gpt\-5\-mini](https://github.com/MicrosoftTranslator/GEMBA)GEMBA\-DA\-ref––[gpt\-5\-mini](https://github.com/MicrosoftTranslator/GEMBA)GEMBA\-MQM––[gpt\-5\-mini](https://github.com/MicrosoftTranslator/GEMBA)GEMBA\-MQM\-ref––[gpt\-5\-mini](https://github.com/MicrosoftTranslator/GEMBA)Table 5:Metrics used in the evaluation\. TheSourcecolumn links to the metric implementation/model card\.
## Appendix CMemorization Test
TextR\-1 \(R\)R\-2 \(R\)R\-L \(R\)Classical Chinese \(source\)33\.949\.4027\.62English \(reference\)19\.613\.0413\.31Table 6:ROUGE Recall scores for memorization testing using the CTP dataset\. Low overlap of text completion with reference suggests that there is no evidence of memorization\.
## Appendix DInter\-Annotator Agreement
SectionCategoryNA1A2Agr\.Inv\.Tot\. Agr\.Agr\. RateFail Rateκ\\kappaPerturbationsIncorrect lexical translation11210010296610286%5%0\.496PerturbationsUnintended translation into modern Chinese137128127121312488%2%0\.265PerturbationsOmitted objects168100106824412649%26%0\.474PerturbationsOmitted subjects180100102876515248%36%0\.684PerturbationsIncorrect tense realization137127127120312388%2%0\.245PerturbationsPronoun substitution and misattribution136130131126112793%1%0\.148PerturbationsSubstitution of classical senses for modern meaning1929776648314733%43%0\.532PerturbationsTitle substitution136133127127313093%2%0\.483Acceptable VariationsSentence segmentation1329911591910069%7%0\.229Acceptable VariationsChinese annotation1341129786119764%8%0\.210Acceptable VariationsName formatting1569996884913756%31%0\.740PerturbationsPerturbations total1198915898823208103169%17%0\.622Acceptable VariationsAcceptable Variations total4223103082656933463%16%0\.468OverallOverall1620122512061088277136567%17%0\.580
Table 7:Inter\-annotator agreement statisticsby category, including raw annotator counts, agreement rates, failed perturbation rates, and Cohen’sκ\\kappa\. A1/A2 = Annotator 1/2, Agr\. = Agreed, Inv\. = Both invalid\.
## Appendix EBaseline Accuracy
MetricUnrelatedShuffledSrc \+ Ref \+ MTCOMET92\.299\.7XCOMET96\.797\.5MetricX\-2496\.1100\.0GEMBA\-DA\-ref99\.999\.9GEMBA\-MQM\-ref88\.588\.5Src \+ MT \(QE\)XCOMET\-QE95\.097\.1MetricX\-24\-QE94\.299\.8GEMBA\-DA99\.999\.8GEMBA\-MQM87\.686\.2Ref \+ MT onlyCOMET\-noSrc90\.898\.8chrF77\.680\.7BLEU79\.857\.8MetricX\-24\-Ref95\.0100\.0Table 8:Baseline accuracy \(%\)\.Proportion of pairs where the metric correctly ranks the MT output above the Unrelated and Shuffled baselines\.Similar Articles
Evaluating and Preserving Lexical Stress in English-to-Chinese Speech-to-Speech Translation
A research paper proposing a new metric and stress-aware system for evaluating and preserving lexical stress in English-to-Chinese speech-to-speech translation, demonstrating significant improvements over existing approaches while maintaining translation quality.
Follow-up to my TranslateGemma-12b benchmark post: human reviewers flagged 71% of the segments automated metrics rated clean
A human review of TranslateGemma-12b's translations revealed that 71% of segments rated clean by automated metrics actually contained errors, highlighting significant gaps in metric-only evaluation for multilingual translation quality.
MetaHOPE: A Metaphor-Oriented Evaluation Framework for Analysing MT and LLM Translation Errors
MetaHOPE is a metaphor-oriented evaluation framework for analyzing translation errors in machine translation and large language models. The paper proposes an error severity-aware annotation framework and evaluates models like GoogleMT, GPT5.4, and Hunyuan-7b on English-Chinese metaphor translation.
Evaluating Multilingual Sentence Embeddings for Translation Error Detection:An English--Greek Contrastive Study
This study evaluates multilingual sentence embeddings for distinguishing correct English–Greek translations from erroneous ones, finding that embeddings provide useful semantic signals but are better integrated into broader translation evaluation frameworks.
Apples to Apples? Towards Comparable Crosslingual Language Model Evaluation
This paper investigates the fairness of crosslingual evaluation methods for language models, showing that common normalized metrics can be biased due to tokenization and orthographic differences, and proposes using sentence-level negative log likelihood on semantically equivalent sequences for more consistent crosslingual comparisons.