TransClean: A Benchmark for Detecting and Extracting Clean Translations from Large Language Model Outputs
Summary
TransClean is a benchmark for detecting and extracting clean translations from large language model outputs, analyzing over 790,000 translation outputs to identify noise patterns and evaluate extraction methods.
View Cached Full Text
Cached at: 09/11/26, 08:29 AM
# TransClean: A Benchmark for Detecting and Extracting Clean Translations from Large Language Model Outputs
Source: [https://arxiv.org/html/2609.11399](https://arxiv.org/html/2609.11399)
rmTeXGyreTermesX tt\[Extension=\.otf, UprightFont=\*\-Regular, BoldFont=\*\-Bold\]Inconsolatazi4 \[\*devanagari\]rmLohit Devanagari \[\*arabic\]rm\[Extension=\.ttf, UprightFont=\*\-Regular, BoldFont=\*\-Bold, ItalicFont=\*\-Italic, BoldItalicFont=\*\-BoldItalic\]Amiri \[\*hebrew\]rm\[Extension=\.ttf, UprightFont=\*\-Medium, BoldFont=\*\-Bold, ItalicFont=\*\-MediumOblique, BoldItalicFont=\*\-BoldOblique\]FrankRuehlCLM \[\*cjk\]rm\[Path=fonts/, Extension=\.otf, UprightFont=NotoSerifCJK\-Subset\]NotoSerifCJKCombined \[\*hangul\]rm\[Path=fonts/, Extension=\.otf, UprightFont=NotoSerifCJK\-Subset\]NotoSerifCJKCombined\\newfontfamily\\cjkfont\[Path=fonts/, Extension=\.otf, UprightFont=NotoSerifCJK\-Subset\]NotoSerifCJKCombined\\newfontfamily\\hangulfont\[Path=fonts/, Extension=\.otf, UprightFont=NotoSerifCJK\-Subset\]NotoSerifCJKCombined
Yves ScherrerAffiliation:Language Technology Group, Department of InformaticsAffiliation:University of Oslo, NorwayAffiliation:\{shenbinq, yves\.scherrer\}@ifi\.uio\.no
###### Abstract
Large language models \(LLMs\) are increasingly used for machine translation, yet their outputs often contain additional text beyond the translation itself, such as language labels, explanations or bilingual repetitions, which we termtranslation noise\. Despite its prevalence, this problem lacks dedicated benchmarks and systematic study\. We analyze over 790,000 translation outputs from 12 LLMs across 22 language pairs \(LPs\) and identify 12 recurring noise patterns, which we group intoformattingandcontentnoise\. Building on the observed patterns, we constructTransClean, a controlled benchmark of 9,900 pairs of noisy and clean translation outputs, comprising 8,800synthetically generatedinstances and 1,100 manually curatedauthenticinstances\. We evaluate two extraction approaches on the TransClean benchmark: 1\) a span\-based extraction method leveraging translation quality estimation models for span detection, and 2\) an LLM\-based extraction method that prompts an LLM to isolate the translation\. Our benchmark and analysis provide the first systematic framework to evaluate and improve the cleanliness of LLM translation outputs\.
## 1Introduction
Large language models \(LLMs\) are rapidly reshaping the landscape of machine translation \(MT\)\. LLMs can perform high\-quality translation through prompting alone and increasingly match or even surpass task\-specific MT systems in many scenarios\([Zhang et al\., 2023](https://arxiv.org/html/2609.11399#bib.bib3);[Vilar et al\., 2023](https://arxiv.org/html/2609.11399#bib.bib35);[Kocmi et al\., 2024](https://arxiv.org/html/2609.11399#bib.bib30);[Xu et al\., 2024](https://arxiv.org/html/2609.11399#bib.bib18)\)\. The flexibility and multilingual capacity of LLMs have led to widespread adoption in both research and deployment settings, from systems developed for the Conference on Machine Translation \(WMT111[https://www2\.statmt\.org/](https://www2.statmt.org/)\) shared tasks\([Kocmi et al\., 2025](https://arxiv.org/html/2609.11399#bib.bib27)\)to large\-scale commercial platforms such as Google Translate\([Caswell, 2024](https://arxiv.org/html/2609.11399#bib.bib19)\)and social media services like Instagram\([Meta AI, 2026](https://arxiv.org/html/2609.11399#bib.bib20)\)\. As LLMs increasingly serve as translation engines, understanding and standardizing their outputs becomes critical\.
Figure 1:The process of creatingTransCleanto benchmark noise detection and clean translation extraction\.However, LLM translations often contain additional text beyond the translation itself\. Instead of producing a single target\-language translation, models may prepend language labels, append explanations, repeat the source sentence, or provide cultural commentary\. While such behavior can be helpful in interactive settings, it introduces a systematic challenge for automatic evaluation and downstream integration\. Standard MT evaluation pipelines assume that model outputs consist solely of the translation\. Extra content can distort metric scores and introduce inconsistencies in large\-scale benchmarking\. We refer to this phenomenon astranslation noise: any content in an LLM output that is not part of the intended target translation\.
In preliminary experiments across multiple models and prompts, we observe that translation noise is not rare\. For most LLMs we evaluate, 3% to 99%222We formally define and quantify the noise rate in§\\lx@sectionsign[2\.2](https://arxiv.org/html/2609.11399#S2.SS2)\.of the translation outputs contain additional explanatory or formatting text\. Although carefully engineered prompts \(e\.g\., “Output the translation only\.”\) reduce this behavior, they do not fully eliminate it\. The prevalence and form of noise vary substantially across models, reflecting differences in instruction\-following abilities\. As a result, clean translation cannot be reliably guaranteed through prompting alone\.
Despite its practical importance, translation noise has not been systematically studied\. Prior work has examined related issues such as instruction following in multilingual settings\([Li et al\., 2024](https://arxiv.org/html/2609.11399#bib.bib21)\), instruction forgetting\([Chen et al\., 2023](https://arxiv.org/html/2609.11399#bib.bib22)\), and undesirable behaviors such as repetitive texts or wrong target\-language outputs\([Bawden and Yvon, 2023](https://arxiv.org/html/2609.11399#bib.bib34);[Wang et al\., 2024](https://arxiv.org/html/2609.11399#bib.bib33)\)\. Existing studies have largely centered either on mitigating these issues through model modification or fine\-tuning, or on evaluating instruction\-following capabilities by proposing new benchmarks such as IFEval\([Zhou et al\., 2023](https://arxiv.org/html/2609.11399#bib.bib23)\), InFoBench\([Qin et al\., 2024](https://arxiv.org/html/2609.11399#bib.bib32)\), and M\-IFEval\([Dussolle et al\., 2025](https://arxiv.org/html/2609.11399#bib.bib28)\)\. However, little attention has been paid to analyzing noise patterns directly in LLM translation outputs or to developing post\-processing methods that recover clean translations without altering the underlying models\. This distinction is crucial in real\-world deployment, where models may be proprietary, closed\-source, or too costly to retrain\.
In this work, we formalize the task ofclean translation extraction: given an LLM output that may contain translation noise, extract the span corresponding to the correct target\-language translation\. We approach this task in three steps:
1. 1\.We conduct a large\-scale empirical study \(§\\lx@sectionsign[2](https://arxiv.org/html/2609.11399#S2)\) of translation noise in LLM outputs, analyzing over 790,000 translations from 12 LLMs across 22 language pairs and identifying 12 recurring noise patterns and 2 main categories\.
2. 2\.We constructTransClean333[https://github\.com/shenbinqian/TransClean](https://github.com/shenbinqian/TransClean), the first benchmark \(§\\lx@sectionsign[3](https://arxiv.org/html/2609.11399#S3)\) for clean translation extraction, comprising 8,800 synthetically noised instances and 1,100 manually curated authentic noisy examples with silver clean translations\. The process of creatingTransCleanis illustrated in Figure[1](https://arxiv.org/html/2609.11399#S1.F1)\.
3. 3\.Using TransClean, we benchmark two approaches toextract clean translations\(§\\lx@sectionsign[4](https://arxiv.org/html/2609.11399#S4)\): 1\) a span\-based extraction method leveraging quality estimation models for span detection, 2\) an LLM\-based extractor that prompts an LLM to isolate the translation\. In this context, we design two evaluation metrics that enable standardized comparison\.
Table 1:The noise rate \(Noise%\), the rate of generating explanatory texts \(Expl%\) and the rate of outputting the wrong target language \(WrongL%\) for different prompts and LLMs\. We did not run all prompts on DeepSeek\-V3\.2\-Exp as we see Prompt 0 generally leads to a higher noise rate across models\.
## 2Noise in LLM Translation Outputs
In order to assess the prevalence and types of noise present in LLM\-produced translations, we generate a large sample of translations for 22 language pairs \(LPs\) using 12 representative LLMs and 3 prompt templates \(§\\lx@sectionsign[2\.1](https://arxiv.org/html/2609.11399#S2.SS1)\)\. We then systematically examine these translation outputs and analyze their noise rate and patterns \(§\\lx@sectionsign[2\.2](https://arxiv.org/html/2609.11399#S2.SS2)and§\\lx@sectionsign[2\.3](https://arxiv.org/html/2609.11399#S2.SS3)\)\.
### 2\.1Generating Noisy Translations
##### Data Sources
We identify 22 LPs with varying resource levels and translation directions and randomly sample 3,000 sentence pairs per LP from four parallel corpora collections: the TED Multilingual Parallel Corpus\([Kulkarni, 2015](https://arxiv.org/html/2609.11399#bib.bib1)\), the WMT20 Quality Estimation Dataset\([Barrault et al\., 2020](https://arxiv.org/html/2609.11399#bib.bib37)\), the SwissAdmin corpus\([Scherrer et al\., 2014](https://arxiv.org/html/2609.11399#bib.bib42)\)and the Chinese–Korean parallel corpus\([Park and Zhao, 2019](https://arxiv.org/html/2609.11399#bib.bib2)\), resulting in a test set of 66,000 instances in total\. Detailed information on LPs, dataset sizes, and their corresponding sources is provided in Table[A\.1](https://arxiv.org/html/2609.11399#A1.T1)in the appendix\.
##### Prompt Templates
We design three prompt templates \(see Figure[2](https://arxiv.org/html/2609.11399#S2.F2)\) to investigate how prompting strategies influence the generation of noisy translations and to identify which prompt yields the highest number of noisy instances\. Prompt 0 is adopted from[Zhang et al\. \(2023\)](https://arxiv.org/html/2609.11399#bib.bib3), while Prompt 1 and Prompt 2 are newly designed templates\.
Prompt 0\{src\_lang\}: \{src\_txt\}\{tgt\_lang\}:Prompt 1Translate the following \{src\_lang\} into \{tgt\_lang\}: \{src\_text\}Prompt 2Translate the following \{src\_lang\} into \{tgt\_lang\} and only output the target text: \{src\_text\}Figure 2:Prompt templates for translation\.
##### Model Selection
Different LLMs may exhibit varying tendencies in producing noisy outputs\. To capture this variability, we select 12 open\-weights LLMs that span a diverse range of model sizes, architectures, post\-training methods, and multilingual training coverage\. The selected models include decoder\-only instruction\-tuned models and their reasoning variants, likeQwen3\-4B\-Instruct\-2507andQwen3\-4B\-Thinking\-2507\([Qwen Team, 2025](https://arxiv.org/html/2609.11399#bib.bib5)\); large frontier mixture\-of\-experts models such as DeepSeek\-V3\.2\-Exp\([DeepSeek\-AI, 2025](https://arxiv.org/html/2609.11399#bib.bib17)\); smaller dense models includingLlama\-3\.2\-3B\-Instruct\([Meta AI, 2024](https://arxiv.org/html/2609.11399#bib.bib6)\)andgemma\-3\-27b\-it\([Gemma Team et al\., 2025](https://arxiv.org/html/2609.11399#bib.bib12)\); recently released instruction\-tuned encoder\-decoder models such ast5gemma\-xl\-xl\-prefixlm\-it\([Zhang et al\., 2025](https://arxiv.org/html/2609.11399#bib.bib7)\); multilingual models includingaya\-expanse\-32b\([Dang et al\., 2024](https://arxiv.org/html/2609.11399#bib.bib8)\); andTower\-Plus\-72B\([Rei et al\., 2025](https://arxiv.org/html/2609.11399#bib.bib9)\), a translation\-oriented LLM fine\-tuned on Qwen\-2\.5\-72B\([Qwen Team, 2024](https://arxiv.org/html/2609.11399#bib.bib4)\)\. Details of all our models can be found in Table[A\.2](https://arxiv.org/html/2609.11399#A1.T2)in the appendix\.
##### Inference
We generate the translation outputs exclusively in a zero\-shot setting\. The details of LLM inference used to generate the translation outputs are provided in Appendix[B](https://arxiv.org/html/2609.11399#A2)\.
Table 2:Noise patterns with examples and frequencies \(%\) in the dataset detected by Claude Opus 4\.6\. We use “…” to denote the omitted long outputs\. Text in red denotes noise\.
### 2\.2Noise Rate
To estimate how frequently translation noise appears in LLM outputs, we employ a lightweight rule\-based detector that captures two common types of noise: explanatory text and text generated in the wrong target language\. The detector is intended to provide a coarse estimate of noise prevalence rather than a complete characterization of all noise types\. The noise rate \(Noise%\) for a given prompt is formally defined as:
Noise%=\|E∪W\|N\\mathrm\{Noise\}\\%=\\frac\{\|E\\cup W\|\}\{N\}\(1\)
whereNNdenotes the total number of translation instances,EErepresents the set of outputs containing explanatory text, andWWdenotes the set of outputs containing text in the wrong target language\. We further defineExp%=\|E\|N\\mathrm\{Exp\}\\%=\\frac\{\|E\|\}\{N\}andWrongL%=\|W\|N\\mathrm\{WrongL\}\\%=\\frac\{\|W\|\}\{N\}to separately measure the proportions of explanatory noise and wrong\-language outputs\.
Explanatory text is detected using regular expressions matching common explanatory markers \(e\.g\., “explanation” and similar meta\-linguistic phrases in Appendix[C](https://arxiv.org/html/2609.11399#A3)\)\. Wrong\-language outputs are identified using the fastText language identification model\([Bojanowski et al\., 2017](https://arxiv.org/html/2609.11399#bib.bib40)\), with a confidence threshold of 60%\. An output is considered noisy if either type of signal is detected\. While this rule\-based detector does not capture all possible noise types, it is sufficient, as evaluated in§\\lx@sectionsign[4\.4](https://arxiv.org/html/2609.11399#S4.SS4), for estimating noise frequencies across prompts and models, and identifying noisy candidates for human curation in§\\lx@sectionsign[3\.2\.1](https://arxiv.org/html/2609.11399#S3.SS2.SSS1)\.
Table[1](https://arxiv.org/html/2609.11399#S1.T1)reports Noise% across prompts and models\. Prompt 0 produces the highest proportion of noisy outputs among the three prompts, particularly for wrong target\-language generation, and we therefore use its outputs for the subsequent noise analysis in§\\lx@sectionsign[2\.3](https://arxiv.org/html/2609.11399#S2.SS3)\. Among models,gemma\-3\-27b\-itexhibits the highest noise rates \(up to 99\.73%\); manual inspection confirms that it frequently generates explanatory text alongside translations\. These high noise rates across models and prompts underscore the need for a systematic study of this problem and for dedicated methods to extract clean translations for fair MT evaluation\.
### 2\.3Noise Analysis
We utilize all 792,000 translation outputs generated by 12 models with Prompt 0 across 22 language pairs for noise analysis\. To identify recurring noise patterns in such a large dataset, we conduct a two\-stage analysis\. First, we use an LLM, Claude Opus 4\.6\([Anthropic PBC, 2026](https://arxiv.org/html/2609.11399#bib.bib11)\), to assist in summarizing and grouping similar noise behaviors across the outputs\. Given a generated translation, the model is prompted to propose representative noise patterns, estimate their relative frequencies, and extract up to ten representative examples for each pattern444All instances are saved when fewer than ten are available\.\. The identified patterns and examples \(see Table[2](https://arxiv.org/html/2609.11399#S2.T2)\) are subsequently verified and refined through manual inspection\.
Based on this inspection, we categorize the twelve observed noise patterns into two groups:contentandformattingnoise, corresponding to semantic and presentation\-level artifacts, respectively\. We retain a relatively fine\-grained set of patterns to facilitate synthetic noise generation, while noting that alternative taxonomies are also possible\. Since manually validating the exact frequency of each pattern at this scale is infeasible, the reported frequencies should be interpreted as approximate estimates rather than precise measurements\.555To assess their reliability, we independently reproduce the frequency estimation using gemma\-4\-31B\-it\([Farabet and Lacombe, 2026](https://arxiv.org/html/2609.11399#bib.bib15)\), which yields a strong and statistically significant positive rank correlation with Claude Opus 4\.6 \(Spearman’sρ=0\.84\\rho=0\.84\)\.These estimated frequencies, together with the taxonomy, are used primarily to capture the overall distribution of noise patterns and to support the generation of synthetic noise that reflects realistic noise distributions, while the representative examples are used as few\-shot demonstrations for synthetic noise generation in§\\lx@sectionsign[3\.1](https://arxiv.org/html/2609.11399#S3.SS1)\.
As shown in Table[2](https://arxiv.org/html/2609.11399#S2.T2),explanationaccounts for the largest share of noisy outputs \(~33%\), primarily produced bygemma\-3\-27b\-itand DeepSeek\-V3\.2\-Exp \(see Table[1](https://arxiv.org/html/2609.11399#S1.T1)\)\. Other frequent patterns includealternativetranslations andoff\-topicresponses\. Generation in the wrong target language666Manual inspection of saved samples suggests they are mostly in English or in the source language\.also contributes a notable proportion \(~7%\)\. Overall, content\-level noise constitutes the majority of noisy translations and is generally more challenging to handle thanformattingnoise\.
## 3Benchmark Construction
To facilitate systematic research on clean translation detection and extraction, we introduceTransClean, the first benchmark designed to evaluate methods for removing noise from LLM\-generated translations\. It comprises 9,900 paired instances of LLM generated translations and their clean translation counterparts\. Although the translations generated in §[2](https://arxiv.org/html/2609.11399#S2)contain various levels of noise, their direct use would require manual annotation of the clean translation, which would be prohibitively expensive and require annotators with expertise in many languages\. Therefore, we adopt a hybrid strategy instead: we generate large\-scalesyntheticnoisy translations using LLMs, guided by the empirically observed noise patterns described in§\\lx@sectionsign[2\.3](https://arxiv.org/html/2609.11399#S2.SS3)\. To further validate the realism and usefulness of the synthetic data, we additionally curate a smaller subset ofauthenticnoisy translations paired with silver clean translations\. This subset enables comparison between synthetic and real noise scenarios\. The construction of the synthetic dataset and the curated subset are described in§\\lx@sectionsign[3\.1](https://arxiv.org/html/2609.11399#S3.SS1)and§\\lx@sectionsign[3\.2](https://arxiv.org/html/2609.11399#S3.SS2)respectively, while representative examples of the final constructed subsets are provided in Appendix[D](https://arxiv.org/html/2609.11399#A4)\.
### 3\.1Synthetic Noise Generation
Because the noise observed in LLM translation outputs originates from LLM generation behaviors, we use LLMs to simulate these noise patterns\. Synthetic noisy translations are generated by injecting noise patterns \(Table[2](https://arxiv.org/html/2609.11399#S2.T2)\) into reference translations, which serve as clean ground\-truth translations\.
##### Noise Generation
We first sample 100 instances for each LP whose source texts contain at least 10 words777Words are defined as space\-separated units for languages using whitespace\. For languages without whitespace segmentation, characters are counted instead\.\. Their reference translations are treated as clean translations, yielding 2,200 clean instances across the 22 LPs\. Based on these clean translations, we generate noisy outputs under 3 noise categories:content,formatting, and their combination \(combo\)\. Each category contains 2,200 instances\. The specific noise pattern applied to each instance is sampled according to the empirical distribution observed in Table[2](https://arxiv.org/html/2609.11399#S2.T2)\. As a result, the distribution of noise patterns in the synthetic data approximately matches the distribution observed in LLM outputs\.
In practice, noise patterns are sampled using their empirical frequencies as weights\. Forformattingnoise, includinglanguage prefix,translation prefix,extra punctuation,code block, andspecial formatting, we apply a rule\-based generator that inserts formatting artifacts into the reference translation\. For the remaining patterns, including allcontentpatterns and theformattingpatternverbose preamble, we use GPT\-5\-mini\([Singh et al\., 2025](https://arxiv.org/html/2609.11399#bib.bib13)\)to generate noisy translations in a few\-shot prompting setup\. The demonstrations consist of representative examples extracted during the noise analysis stage \(§\\lx@sectionsign[2\.3](https://arxiv.org/html/2609.11399#S2.SS3)\)\. The prompt template is provided in Appendix[E](https://arxiv.org/html/2609.11399#A5)\. For all noise patterns, the reference translation is used as the gold clean translation, exceptoff\-topicandwrong language, whose gold clean is an empty string\.
Overall, the synthetic dataset contains 8,800 instances distributed across four splits: three noisy categories and one clean category\. The clean split serves as a control set to evaluate whether extraction methods preserve already clean translations\. Detailed statistics of the synthetic dataset are shown in Table[F\.1](https://arxiv.org/html/2609.11399#A6.T1)\.
##### Manual Validation
To assess the realism and correctness of the generated noise, we conduct manual validation on a subset of the synthetic data\. Specifically, we randomly sample 10 instances for each of the seven LLM\-generated noise patterns \(i\.e\.,explanation,alternatives,off\-topic,verbose preamble,bilingual output,wrong language, andcultural note\)\. For thecombocategory, we additionally sample 30 instances\. This results in 100 manually inspected instances covering 9 language pairs\. Manual inspection confirms that the generated outputs correctly reflect the intended noise patterns and closely resemble the noise behaviors in real LLM translations\.888In the case ofwrong language, we found that the LLM typically generates text in English or in the source language, which reflects the type of language confusion found in the noise analysis in§\\lx@sectionsign[2\.3](https://arxiv.org/html/2609.11399#S2.SS3)\.
### 3\.2Curated Noisy Subset
To complement the synthetic dataset, we construct a curated subset of authentic noisy translations drawn from real LLM outputs\. The curated subset contains 1,100 instances, each annotated with a noise pattern and a silver clean translation\.
#### 3\.2\.1Authentic Noise Curation
Taking the 792,000 LLM translation outputs from Prompt 0 in §[2\.1](https://arxiv.org/html/2609.11399#S2.SS1)as a starting point, we first apply the rule\-based detector described in§\\lx@sectionsign[2\.2](https://arxiv.org/html/2609.11399#S2.SS2)to filter potentially noisy outputs, yielding 232,403 candidate instances\. We then use GPT\-5\-mini to identify authentic noisy translations among the candidate instances\. For each instance classified as noisy, the model assigns a noise pattern label\. We curate 50 instances per language pair, each annotated with its corresponding noise pattern\. We then verify the detected instances to confirm they represent authentic noise\. The prompt used for this task is provided in Appendix[G](https://arxiv.org/html/2609.11399#A7)\.
#### 3\.2\.2Clean Translation Annotation
We employ three LLMs, GPT\-5\-mini, Qwen3\.5\-122B\-A10B\([Qwen Team, 2026](https://arxiv.org/html/2609.11399#bib.bib14)\), and gemma\-4\-31B\-it to annotate the clean translation for each noisy output\. The models are provided with the noisy translation and its corresponding noise label as context\. The prompt used for this task is shown in Appendix[H](https://arxiv.org/html/2609.11399#A8)\.
Table 3:Agreement of the three LLMs on annotating clean translations for the 1100 curated examples\.We adopt a majority voting strategy to determine the final silver clean translation\. If at least two models produce identical outputs, the shared translation is used as the label\. In 48 cases where at least one model outputs an empty string, manual inspection confirms that the corresponding outputs areoff\-topicresponses, and the empty string is therefore retained as the correct label\. Agreement statistics of the three models are presented in Table[3](https://arxiv.org/html/2609.11399#S3.T3)\. Among the remaining 315 instances without majority agreement, we manually examine 143 instances for which the authors are fluent speakers of the target language\. In these cases, gemma\-4\-31B\-it produces the correct clean translation for all instances except 17 where multiple valid translations exist and all model outputs are acceptable\. Based on this observation, we adopt the output of gemma\-4\-31B\-it as the silver label for the remaining disagreement cases\.
## 4Translation Extraction
To support fair MT evaluation beyond merely detecting noise, we propose two methods that can extract clean translations from noisy LLM outputs: a span\-based method using quality estimation models in§\\lx@sectionsign[4\.1](https://arxiv.org/html/2609.11399#S4.SS1), and an LLM\-based extraction method in§\\lx@sectionsign[4\.2](https://arxiv.org/html/2609.11399#S4.SS2)\. Evaluation metrics and results are presented in§\\lx@sectionsign[4\.3](https://arxiv.org/html/2609.11399#S4.SS3)and§\\lx@sectionsign[4\.4](https://arxiv.org/html/2609.11399#S4.SS4)\. The rule\-based detector is evaluated in§\\lx@sectionsign[4\.4](https://arxiv.org/html/2609.11399#S4.SS4)for noise detection only to compare with the proposed methods\. Details for running these extraction methods are in Appendix[I](https://arxiv.org/html/2609.11399#A9)\.
### 4\.1Span\-based Extraction
The observation that lengthy explanatory text constitutes the largest source of noise in LLM outputs, and that explanations are typically separated from the translation by line breaks, motivates a span\-based approach: splitting the output into shorter spans, identifying the span most likely to contain the clean translation, and removing any residual noise from that span\.
Following common patterns observed in LLM\-generated text, we segment the output using the line feed character \(\\n\) as a delimiter, yielding a set of candidate spansS=\{s1,s2,…,sm\}S=\\\{s\_\{1\},s\_\{2\},\\dots,s\_\{m\}\\\}\. We then score each span using COMET\-KIWI\([Rei et al\., 2022](https://arxiv.org/html/2609.11399#bib.bib36)\), a reference\-free quality estimation \(QE\) model, which produces a scoreq\(sj,xsrc\)q\(s\_\{j\},x\_\{\\text\{src\}\}\)by comparing each spansjs\_\{j\}against the source textxsrcx\_\{\\text\{src\}\}\. The candidate clean translation is selected as:
s∗=\{s1,if\|S\|=1argmaxsj∈Sq\(sj,xsrc\),if\|S\|\>1s^\{\*\}=\\begin\{cases\}s\_\{1\},&\\text\{if \}\|S\|=1\\\\ \\displaystyle\\arg\\max\_\{s\_\{j\}\\in S\}\\,q\(s\_\{j\},x\_\{\\text\{src\}\}\),&\\text\{if \}\|S\|\>1\\end\{cases\}\(2\)
That is, if the output contains a single span, it is directly taken as the candidate translation without QE scoring\. Otherwise, the span with the highest QE score is selected\.
Finally, a rule\-based post\-processing step is applied tos∗s^\{\*\}to remove any residual formatting noise, such as language prefixes, yielding the extracted translationt^=g\(s∗\)\\hat\{t\}=g\(s^\{\*\}\), whereg\(⋅\)g\(\\cdot\)denotes the rule\-based cleaning function\.
### 4\.2LLM\-based Extraction
As an alternative to the span\-based method, we propose using LLMs directly as extractors to produce clean translations\. Given a noisy LLM translation outputxix\_\{i\}, the extraction is formulated as:
t^i=ℳext\(\[p;xi\]\)\\hat\{t\}\_\{i\}=\\mathcal\{M\}\_\{\\text\{ext\}\}\(\[p;x\_\{i\}\]\)\(3\)
whereℳext\\mathcal\{M\}\_\{\\text\{ext\}\}is the extractor LLM andppis a fixed prompt template instructing the model to extract the clean translation fromxix\_\{i\}\(see Appendix[J](https://arxiv.org/html/2609.11399#A10)for the full template\)\. Notably, the extractor receives only the LLM translation outputxix\_\{i\}\. No reference translation, or description of noise patterns is provided\. This constraint ensures a fair comparison with the span\-based approach, which likewise operates without access to reference information\.
This setup also distinguishes this LLM extraction approach from the silver clean translation annotation procedure described in§\\lx@sectionsign[3\.2\.2](https://arxiv.org/html/2609.11399#S3.SS2.SSS2), where the annotator LLM is given both the source text and reference translations, along with explicit noise pattern descriptions\.
In practice, we employ two backbone extractors under a zero\-shot setting: a multilingual dense LLM,aya\-expanse\-32b, and Qwen3\.5\-122B\-A10B, an English\- and Chinese\-dominant mixture\-of\-experts model\.
### 4\.3Evaluation Metrics
To evaluate how effectively our methods detect noise and extract clean translations from LLM outputs with our benchmark, we introduce two metrics:detection accuracyandextraction accuracy\.
##### Detection Accuracy
Detection accuracy measures the rate at which a method correctly identifies whether an LLM translation output is noisy or clean \(i\.e\., translation\-only\)\. Given theii\-th LLM translation outputxix\_\{i\}and an extraction methodf\(⋅\)f\(\\cdot\), the predicted label is determined by:
y^i=\{0\(clean\),iff\(xi\)=xi1\(noisy\),iff\(xi\)≠xi\\hat\{y\}\_\{i\}=\\begin\{cases\}0\\ \(\\text\{clean\}\),&\\text\{if \}f\(x\_\{i\}\)=x\_\{i\}\\\\ 1\\ \(\\text\{noisy\}\),&\\text\{if \}f\(x\_\{i\}\)\\neq x\_\{i\}\\end\{cases\}\(4\)That is, if the extraction method returns the input unchanged, the sample is classified as clean; any modification to the input implies the presence of noise\. Detection accuracy \(Accdet\) is computed as the proportion of samples for which the predicted noise labely^i\\hat\{y\}\_\{i\}matches the ground\-truth labelyiy\_\{i\}\.
##### Extraction Accuracy
Extraction accuracy is a stricter metric that measures whether the extracted translation exactly matches the annotated clean reference\. Normalization is applied to both strings prior to comparison\. It is formally defined as:
Accext=1N∑i=1N\[norm\(t^i\)=norm\(ti\)\]\\text\{Acc\}\_\{\\text\{ext\}\}=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\mathbbm\{1\}\\\!\\left\[\\texttt\{norm\}\(\\hat\{t\}\_\{i\}\)=\\texttt\{norm\}\(t\_\{i\}\)\\right\]\(5\)wheret^i\\hat\{t\}\_\{i\}is the extracted translation for theii\-th sample,tit\_\{i\}is the corresponding annotated clean reference,𝟙\[⋅\]\\mathbbm\{1\}\[\\cdot\]is the indicator function,NNis the total number of samples, andnorm\(⋅\)\\texttt\{norm\}\(\\cdot\)denotes the normalization function applied before comparison\. Specifically,norm\(⋅\)\\texttt\{norm\}\(\\cdot\)strips leading and trailing whitespace and applies Unicode NFC normalization tot^i\\hat\{t\}\_\{i\}andtit\_\{i\}\.
### 4\.4Evaluation Results
Table 4:Detection and extraction accuracy \(%\) on the synthetic and curated subsets of our benchmark\.Table 5:Extraction accuracy \(%\) for each noise split \(category\) of the synthetic subset\. The “noisy” column is the combination of the first three categories\.##### Overall Results
Table[4](https://arxiv.org/html/2609.11399#S4.T4)reports the detection and extraction accuracy on both the synthetic and curated subsets\. Detection accuracy is nearly saturated \(close to 100%\) for most methods and both subsets, indicating that identifying whether a translation contains noise is relatively easy\. Our rule\-based detector performs well, especially on the curated noisy subset, demonstrating its effectiveness in detecting noise and calculating Noise%\.
Extraction, however, remains challenging\. The best\-performing method,Qwen extractor, achieves 54\.07% accuracy on the synthetic subset and 52\.18% on the curated subset\. While promising under strict exact\-match evaluation, these results suggest substantial room for improvement in extracting clean translations from noisy outputs\.
The span\-based approach performs well on the synthetic dataset but drops sharply on the curated subset for extraction accuracy\. This behavior is expected: the synthetic data follows predefined noise patterns aligned with the rule\-based cleaning function, whereas the curated subset contains authentic noise that may not match these patterns\. In contrast, LLM\-based extraction approaches exhibit more stability across datasets, suggesting better generalization to diverse noise patterns\.
An exception isAya extractor, whose accuracy increases on the curated subset while most other methods decline\. To understand this behavior, we analyze its performance on each split of the synthetic subset\. Aya achieves only 59\.91% accuracy on the clean split, substantially lower than other methods \(above 90%\)\. Inspection shows that Aya frequently paraphrases already clean translations—about 40% of the time \(892/2200\)—altering wording, punctuation, or sentence structure\. These minor reformulations lead to mismatches under exact\-match evaluation and largely explain its lower synthetic\-set performance\.
Overall, the relative performance trends are consistent across synthetic and curated subsets, suggesting that the synthetic data reasonably approximates real noisy translations and serves as a reliable benchmark to evaluate extraction methods\.
##### Results per Noise Split
Table[5](https://arxiv.org/html/2609.11399#S4.T5)presents the extraction accuracy for each noise split of the synthetic subset\.Qwen extractorachieves the best performance across all noise categories\. The only exception is the clean split, where the span\-based method attains 100% accuracy by preserving the original translation when no noise is detected\.
Across noise categories,formattingnoise yields the highest accuracy among the noisy splits\. This is expected, as formatting noise can often be removed with simple transformations\. The most difficult category iscombo, which combinescontentandformattingnoise\. The interaction of multiple noise types substantially increases extraction difficulty, leading to lower accuracy for all methods\.
These results further support the design of the synthetic benchmark: the splits exhibit distinct difficulty levels and capture meaningful differences among noise categories, enabling more fine\-grained evaluation of extraction approaches\.
## 5Related Work
LLMs are increasingly used to generate synthetic data for Natural Language Processing tasks due to their strong language modeling and controllable generation capabilities\([Long et al\., 2024](https://arxiv.org/html/2609.11399#bib.bib31);[Nadǎş et al\., 2025](https://arxiv.org/html/2609.11399#bib.bib24)\)\. In MT, synthetic data has long been used through techniques such as back\-translation to improve performance, particularly for low\-resource languages\([Hassan et al\., 2017](https://arxiv.org/html/2609.11399#bib.bib41);[Poncelas et al\., 2018](https://arxiv.org/html/2609.11399#bib.bib39)\)\. More recent work employs LLMs directly to generate synthetic multilingual data\. For example,[de Gibert et al\. \(2025\)](https://arxiv.org/html/2609.11399#bib.bib29)generate translations for several low\-resource languages using GPT\-4o\([OpenAI et al\., 2024](https://arxiv.org/html/2609.11399#bib.bib25)\)and show that synthetic data can improve downstream MT systems despite its noise\. However, prior work primarily uses synthetic data to improve translation models rather than to study the behavior of LLM\-generated translations themselves\. In this work, we instead leverage LLMs to generate synthetic translation noise based on empirically observed patterns, enabling scalable construction of a benchmark for clean translation extraction\.
## 6Conclusion
In this work, we present a systematic study oftranslation noise\. Through large\-scale analysis of more than 790,000 LLM translation outputs across 22 language pairs, we identify 12 recurring noise patterns and categorize them intoformattingandcontentnoise\. Based on these observations, we introduceTransClean, the first benchmark designed to evaluate methods that extract clean translations from noisy LLM outputs\. The benchmark combines a large synthetic dataset with gold clean translations and a curated subset of authentic noisy translations with silver clean translation, enabling both controlled evaluation and validation on realistic data\. Using this benchmark, we evaluate span\-based and LLM\-based extraction approaches and show that, while noise detection is relatively straightforward, clean translation extraction remains a challenging task with substantial room for improvement\.
In future work, we plan to develop more robust methods for clean translation extraction and explore approaches that better generalize to diverse noise patterns across languages and models\. We hope that TransClean will facilitate further research toward more reliable use of LLMs for translation and other structured generation tasks\.
## Limitations
This work has several limitations\. First, our estimation of the noise rate relies on a coarse rule\-based detector that identifies explanatory text through English keyword matching and detects wrong\-language outputs using automatic language identification\. While this approach enables scalable analysis across hundreds of thousands of translation outputs, it may miss some noise instances that do not match the predefined patterns or may occasionally produce false positives\. Developing more reliable detection methods is therefore an important direction for future work\. One motivation of TransClean is precisely to provide a benchmark that enables systematic evaluation of improved detection and extraction approaches\.
Second, the identification of noise patterns was assisted by an LLM due to the scale of the collected outputs \(over 790,000 translations\), which makes full manual inspection impractical\. Although we subsequently verified the discovered patterns and examples through manual review, the taxonomy of noise patterns may not be exhaustive and could evolve as new models or prompting strategies produce different types of noise\.
Third, the synthetic noise of our benchmark is generated based on observed patterns and their empirical distribution\. While this design enables controlled evaluation and sufficient scale, synthetic noise may not fully capture the diversity and complexity of noise produced by LLMs in real\-world settings\. In addition, although we include a curated subset of authentic noise, the benchmark remains largely English\-centric because both the translation prompts and the noise\-generation prompts are written in English\. As a result, the generated noise may under\-represent truly multilingual or language\-specific noise phenomena\. Extending the benchmark with more diverse multilingual noise patterns remains an important direction for future work\.
Finally, exact\-match extraction accuracy may be overly stringent as the sole primary metric\. We observe that it can penalize semantically correct outputs when Aya paraphrases translations that are already clean\. This highlights a potential mismatch between exact\-match evaluation and the semantic correctness of the extracted translations\. Softer edit\-based measures or semantic similarity metrics could therefore provide a more informative complement to exact\-match accuracy in future work\.
## Ethical Considerations
This research relies exclusively on publicly accessible datasets, with all data utilization adhering to the licensing agreements specified by[Kulkarni \(2015\)](https://arxiv.org/html/2609.11399#bib.bib1),[Scherrer et al\. \(2014\)](https://arxiv.org/html/2609.11399#bib.bib42),[Park and Zhao \(2019\)](https://arxiv.org/html/2609.11399#bib.bib2), and[Barrault et al\. \(2020\)](https://arxiv.org/html/2609.11399#bib.bib37)\. It is presumed that these repositories contain no sensitive or personally identifiable information\. Consequently, their application in this study is deemed to present no significant ethical risks\. Furthermore, the systematic generation of synthetic noise and the curation of authentic noise samples are not expected to yield additional personal data or introduce further ethical complications\. In the interest of transparency and reproducibility, the resulting dataset is released to the public domain\.
All original ideas, analyses, and content in this paper were created by the authors\. AI tools were used only as supportive aids for improving writing quality and assisting with coding tasks\. The authors retain full responsibility for the intellectual content, analyses, and conclusions presented in this work\.
## Acknowledgments
This work has received funding from the European Union’s Horizon Europe research and innovation programme under the Marie Skłodowska\-Curie grant agreement No\. 101126636\.
The computations were performed on resources provided through Sigma2—the national research infrastructure provider for high\-performance computing and large\-scale data storage in Norway\. We acknowledge Norway and Sigma2 for awarding this project access to the Olivia supercomputer, through Project nn9851k\.
## References
- Anthropic PBC \(2026\)Anthropic PBCIntroducing Claude Opus 4\.6\.Note:AnthropicAccessed on 04, May 2026External Links:[Link](https://www.anthropic.com/news/claude-opus-4-6)Cited by:[§2\.3](https://arxiv.org/html/2609.11399#S2.SS3.p1.1)\.
- Barraultet al\.\(2020\)L\. Barrault, M\. Biesialska, O\. Bojar, M\. R\. Costa\-jussà, C\. Federmann, Y\. Graham, R\. Grundkiewicz, B\. Haddow, M\. Huck, E\. Joanis, T\. Kocmi, P\. Koehn, C\. Lo, N\. Ljubešić, C\. Monz, M\. Morishita, M\. Nagata, T\. Nakazawa, S\. Pal, M\. Post, and M\. ZampieriFindings of the 2020 conference on machine translation \(WMT20\)\.InProceedings of the Fifth Conference on Machine Translation,L\. Barrault, O\. Bojar, F\. Bougares, R\. Chatterjee, M\. R\. Costa\-jussà, C\. Federmann, M\. Fishel, A\. Fraser, Y\. Graham, P\. Guzman, B\. Haddow, M\. Huck, A\. J\. Yepes, P\. Koehn, A\. Martins, M\. Morishita, C\. Monz, M\. Nagata, T\. Nakazawa, and M\. Negri \(Eds\.\),Online,pp\. 1–55\.External Links:[Link](https://aclanthology.org/2020.wmt-1.1/),[Document](https://dx.doi.org/10.18653/v1/2020.wmt-1.1)Cited by:[§2\.1](https://arxiv.org/html/2609.11399#S2.SS1.SSS0.Px1.p1.1),[Ethical Considerations](https://arxiv.org/html/2609.11399#Sx2.p1.1)\.
- Bawden and Yvon \(2023\)R\. Bawden and F\. YvonInvestigating the translation performance of a large multilingual language model: the case of BLOOM\.InProceedings of the 24th Annual Conference of the European Association for Machine Translation,M\. Nurminen, J\. Brenner, M\. Koponen, S\. Latomaa, M\. Mikhailov, F\. Schierl, T\. Ranasinghe, E\. Vanmassenhove, S\. A\. Vidal, N\. Aranberri, M\. Nunziatini, C\. P\. Escartín, M\. Forcada, M\. Popovic, C\. Scarton, and H\. Moniz \(Eds\.\),Tampere, Finland,pp\. 157–170\.External Links:[Link](https://aclanthology.org/2023.eamt-1.16/)Cited by:[§1](https://arxiv.org/html/2609.11399#S1.p4.1)\.
- Bojanowskiet al\.\(2017\)P\. Bojanowski, E\. Grave, A\. Joulin, and T\. MikolovEnriching word vectors with subword information\.Transactions of the Association for Computational Linguistics5,pp\. 135–146\.External Links:[Link](https://aclanthology.org/Q17-1010/),[Document](https://dx.doi.org/10.1162/tacl%5Fa%5F00051)Cited by:[§2\.2](https://arxiv.org/html/2609.11399#S2.SS2.p4.1)\.
- Caswell \(2024\)I\. Caswell110 new languages are coming to Google Translate\.Note:Accessed on 10, Dec 2025External Links:[Link](https://blog.google/products/translate/google-translate-new-languages-2024/)Cited by:[§1](https://arxiv.org/html/2609.11399#S1.p1.1)\.
- Chenet al\.\(2023\)Y\. Chen, Y\. Liu, F\. Meng, Y\. Chen, J\. Xu, and J\. ZhouImproving translation faithfulness of large language models via augmenting instructions\.arXiv preprint\.External Links:2308\.12674,[Link](https://arxiv.org/abs/2308.12674)Cited by:[§1](https://arxiv.org/html/2609.11399#S1.p4.1)\.
- Danget al\.\(2024\)J\. Dang, S\. Singh, D\. D’souza, A\. Ahmadian, A\. Salamanca, M\. Smith, A\. Peppin, S\. Hong, M\. Govindassamy, T\. Zhao, S\. Kublik, M\. Amer, V\. Aryabumi, J\. A\. Campos, Y\. Tan, T\. Kocmi, F\. Strub, N\. Grinsztajn, Y\. Flet\-Berliac, A\. Locatelli, H\. Lin, D\. Talupuru, B\. Venkitesh, D\. Cairuz, B\. Yang, T\. Chung, W\. Ko, S\. S\. Shi, A\. Shukayev, S\. Bae, A\. Piktus, R\. Castagné, F\. Cruz\-Salinas, E\. Kim, L\. Crawhall\-Stein, A\. Morisot, S\. Roy, P\. Blunsom, I\. Zhang, A\. Gomez, N\. Frosst, M\. Fadaee, B\. Ermis, A\. Üstün, and S\. HookerAya expanse: combining research breakthroughs for a new multilingual frontier\.arXiv preprint\.External Links:2412\.04261,[Link](https://arxiv.org/abs/2412.04261)Cited by:[§2\.1](https://arxiv.org/html/2609.11399#S2.SS1.SSS0.Px3.p1.1)\.
- de Gibertet al\.\(2025\)O\. de Gibert, J\. Attieh, T\. Vahtola, M\. Aulamo, Z\. Li, R\. Vázquez, T\. Hu, and J\. TiedemannScaling low\-resource MT via synthetic data generation with LLMs\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 27674–27692\.External Links:[Link](https://aclanthology.org/2025.emnlp-main.1408/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.1408),ISBN 979\-8\-89176\-332\-6Cited by:[§5](https://arxiv.org/html/2609.11399#S5.p1.1)\.
- DeepSeek\-AIet al\.\(2025\)DeepSeek\-AI, A\. Liu, A\. Mei, B\. Lin, B\. Xue, B\. Wang, B\. Xu, B\. Wu, B\. Zhang, C\. Lin, C\. Dong, C\. Lu, C\. Zhao, C\. Deng, C\. Xu, C\. Ruan, D\. Dai, D\. Guo, D\. Yang, D\. Chen, E\. Li, F\. Zhou, F\. Lin, F\. Dai, G\. Hao, G\. Chen, G\. Li, H\. Zhang, H\. Xu, H\. Li, H\. Liang, H\. Wei, H\. Zhang, H\. Luo, H\. Ji, H\. Ding, H\. Tang, H\. Cao, H\. Gao, H\. Qu, H\. Zeng, J\. Huang, J\. Li, J\. Xu, J\. Hu, J\. Chen, J\. Xiang, J\. Yuan, J\. Cheng, J\. Zhu, J\. Ran, J\. Jiang, J\. Qiu, J\. Li, J\. Song, K\. Dong, K\. Gao, K\. Guan, K\. Huang, K\. Zhou, K\. Huang, K\. Yu, L\. Wang, L\. Zhang, L\. Wang, L\. Zhao, L\. Yin, L\. Guo, L\. Luo, L\. Ma, L\. Wang, L\. Zhang, M\. S\. Di, M\. Y\. Xu, M\. Zhang, M\. Zhang, M\. Tang, M\. Zhou, P\. Huang, P\. Cong, P\. Wang, Q\. Wang, Q\. Zhu, Q\. Li, Q\. Chen, Q\. Du, R\. Xu, R\. Ge, R\. Zhang, R\. Pan, R\. Wang, R\. Yin, R\. Xu, R\. Shen, R\. Zhang, S\. H\. Liu, S\. Lu, S\. Zhou, S\. Chen, S\. Cai, S\. Chen, S\. Hu, S\. Liu, S\. Hu, S\. Ma, S\. Wang, S\. Yu, S\. Zhou, S\. Pan, S\. Zhou, T\. Ni, T\. Yun, T\. Pei, T\. Ye, T\. Yue, W\. Zeng, W\. Liu, W\. Liang, W\. Pang, W\. Luo, W\. Gao, W\. Zhang, X\. Gao, X\. Wang, X\. Bi, X\. Liu, X\. Wang, X\. Chen, X\. Zhang, X\. Nie, X\. Cheng, X\. Liu, X\. Xie, X\. Liu, X\. Yu, X\. Li, X\. Yang, X\. Li, X\. Chen, X\. Su, X\. Pan, X\. Lin, X\. Fu, Y\. Q\. Wang, Y\. Zhang, Y\. Xu, Y\. Ma, Y\. Li, Y\. Li, Y\. Zhao, Y\. Sun, Y\. Wang, Y\. Qian, Y\. Yu, Y\. Zhang, Y\. Ding, Y\. Shi, Y\. Xiong, Y\. He, Y\. Zhou, Y\. Zhong, Y\. Piao, Y\. Wang, Y\. Chen, Y\. Tan, Y\. Wei, Y\. Ma, Y\. Liu, Y\. Yang, Y\. Guo, Y\. Wu, Y\. Wu, Y\. Cheng, Y\. Ou, Y\. Xu, Y\. Wang, Y\. Gong, Y\. Wu, Y\. Zou, Y\. Li, Y\. Xiong, Y\. Luo, Y\. You, Y\. Liu, Y\. Zhou, Z\. F\. Wu, Z\. Z\. Ren, Z\. Zhao, Z\. Ren, Z\. Sha, Z\. Fu, Z\. Xu, Z\. Xie, Z\. Zhang, Z\. Hao, Z\. Gou, Z\. Ma, Z\. Yan, Z\. Shao, Z\. Huang, Z\. Wu, Z\. Li, Z\. Zhang, Z\. Xu, Z\. Wang, Z\. Gu, Z\. Zhu, Z\. Li, Z\. Zhang, Z\. Xie, Z\. Gao, Z\. Pan, Z\. Yao, B\. Feng, H\. Li, J\. L\. Cai, J\. Ni, L\. Xu, M\. Li, N\. Tian, R\. J\. Chen, R\. L\. Jin, S\. S\. Li, S\. Zhou, T\. Sun, X\. Q\. Li, X\. Jin, X\. Shen, X\. Chen, X\. Song, X\. Zhou, Y\. X\. Zhu, Y\. Huang, Y\. Li, Y\. Zheng, Y\. Zhu, Y\. Ma, Z\. Huang, Z\. Xu, Z\. Zhang, D\. Ji, J\. Liang, J\. Guo, J\. Chen, L\. Xia, M\. Wang, M\. Li, P\. Zhang, R\. Chen, S\. Sun, S\. Wu, S\. Ye, T\. Wang, W\. L\. Xiao, W\. An, X\. Wang, X\. Sun, X\. Wang, Y\. Tang, Y\. Zha, Z\. Zhang, Z\. Ju, Z\. Zhang, and Z\. QuDeepSeek\-V3\.2: pushing the frontier of open large language models\.arXiv preprint\.External Links:2512\.02556,[Link](https://arxiv.org/abs/2512.02556)Cited by:[Appendix B](https://arxiv.org/html/2609.11399#A2.p1.1)\.
- DeepSeek\-AI \(2025\)DeepSeek\-AIDeepSeek\-v3\.2\-exp: boosting long\-context efficiency with deepseek sparse attention\.Note:Accessed on 08, Dec 2025External Links:[Link](https://github.com/deepseek-ai/DeepSeek-V3.2-Exp/blob/main/DeepSeek_V3_2.pdf)Cited by:[§2\.1](https://arxiv.org/html/2609.11399#S2.SS1.SSS0.Px3.p1.1)\.
- Dussolleet al\.\(2025\)A\. Dussolle, A\. Cardeña Díaz, S\. Sato, and P\. DevineM\-IFEval: multilingual instruction\-following evaluation\.InFindings of the Association for Computational Linguistics: NAACL 2025,L\. Chiruzzo, A\. Ritter, and L\. Wang \(Eds\.\),Albuquerque, New Mexico,pp\. 6176–6191\.External Links:[Link](https://aclanthology.org/2025.findings-naacl.344/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-naacl.344),ISBN 979\-8\-89176\-195\-7Cited by:[§1](https://arxiv.org/html/2609.11399#S1.p4.1)\.
- Farabet and Lacombe \(2026\)C\. Farabet and O\. LacombeGemma 4: Byte for byte, the most capable open models\.External Links:[Link](https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/)Cited by:[footnote 5](https://arxiv.org/html/2609.11399#footnote5)\.
- Gemma Teamet al\.\(2025\)Gemma Team, A\. Kamath, J\. Ferret, S\. Pathak, N\. Vieillard, R\. Merhej, S\. Perrin, T\. Matejovicova, A\. Ramé, M\. Rivière, L\. Rouillard, T\. Mesnard, G\. Cideron, J\. Grill, S\. Ramos, E\. Yvinec, M\. Casbon, E\. Pot, I\. Penchev, G\. Liu, F\. Visin, K\. Kenealy, L\. Beyer, X\. Zhai, A\. Tsitsulin, R\. Busa\-Fekete, A\. Feng, N\. Sachdeva, B\. Coleman, Y\. Gao, B\. Mustafa, I\. Barr, E\. Parisotto, D\. Tian, M\. Eyal, C\. Cherry, J\. Peter, D\. Sinopalnikov, S\. Bhupatiraju, R\. Agarwal, M\. Kazemi, D\. Malkin, R\. Kumar, D\. Vilar, I\. Brusilovsky, J\. Luo, A\. Steiner, A\. Friesen, A\. Sharma, A\. Sharma, A\. M\. Gilady, A\. Goedeckemeyer, A\. Saade, A\. Feng, A\. Kolesnikov, A\. Bendebury, A\. Abdagic, A\. Vadi, A\. György, A\. S\. Pinto, A\. Das, A\. Bapna, A\. Miech, A\. Yang, A\. Paterson, A\. Shenoy, A\. Chakrabarti, B\. Piot, B\. Wu, B\. Shahriari, B\. Petrini, C\. Chen, C\. L\. Lan, C\. A\. Choquette\-Choo, C\. J\. Carey, C\. Brick, D\. Deutsch, D\. Eisenbud, D\. Cattle, D\. Cheng, D\. Paparas, D\. S\. Sreepathihalli, D\. Reid, D\. Tran, D\. Zelle, E\. Noland, E\. Huizenga, E\. Kharitonov, F\. Liu, G\. Amirkhanyan, G\. Cameron, H\. Hashemi, H\. Klimczak\-Plucińska, H\. Singh, H\. Mehta, H\. T\. Lehri, H\. Hazimeh, I\. Ballantyne, I\. Szpektor, I\. Nardini, J\. Pouget\-Abadie, J\. Chan, J\. Stanton, J\. Wieting, J\. Lai, J\. Orbay, J\. Fernandez, J\. Newlan, J\. Ji, J\. Singh, K\. Black, K\. Yu, K\. Hui, K\. Vodrahalli, K\. Greff, L\. Qiu, M\. Valentine, M\. Coelho, M\. Ritter, M\. Hoffman, M\. Watson, M\. Chaturvedi, M\. Moynihan, M\. Ma, N\. Babar, N\. Noy, N\. Byrd, N\. Roy, N\. Momchev, N\. Chauhan, N\. Sachdeva, O\. Bunyan, P\. Botarda, P\. Caron, P\. K\. Rubenstein, P\. Culliton, P\. Schmid, P\. G\. Sessa, P\. Xu, P\. Stanczyk, P\. Tafti, R\. Shivanna, R\. Wu, R\. Pan, R\. Rokni, R\. Willoughby, R\. Vallu, R\. Mullins, S\. Jerome, S\. Smoot, S\. Girgin, S\. Iqbal, S\. Reddy, S\. Sheth, S\. Põder, S\. Bhatnagar, S\. R\. Panyam, S\. Eiger, S\. Zhang, T\. Liu, T\. Yacovone, T\. Liechty, U\. Kalra, U\. Evci, V\. Misra, V\. Roseberry, V\. Feinberg, V\. Kolesnikov, W\. Han, W\. Kwon, X\. Chen, Y\. Chow, Y\. Zhu, Z\. Wei, Z\. Egyed, V\. Cotruta, M\. Giang, P\. Kirk, A\. Rao, K\. Black, N\. Babar, J\. Lo, E\. Moreira, L\. G\. Martins, O\. Sanseviero, L\. Gonzalez, Z\. Gleicher, T\. Warkentin, V\. Mirrokni, E\. Senter, E\. Collins, J\. Barral, Z\. Ghahramani, R\. Hadsell, Y\. Matias, D\. Sculley, S\. Petrov, N\. Fiedel, N\. Shazeer, O\. Vinyals, J\. Dean, D\. Hassabis, K\. Kavukcuoglu, C\. Farabet, E\. Buchatskaya, J\. Alayrac, R\. Anil, Dmitry, Lepikhin, S\. Borgeaud, O\. Bachem, A\. Joulin, A\. Andreev, C\. Hardin, R\. Dadashi, and L\. HussenotGemma 3 Technical Report\.arXiv preprint\.External Links:2503\.19786,[Link](https://arxiv.org/abs/2503.19786)Cited by:[§2\.1](https://arxiv.org/html/2609.11399#S2.SS1.SSS0.Px3.p1.1)\.
- Hassanet al\.\(2017\)H\. Hassan, M\. Elaraby, and A\. Y\. TawfikSynthetic data for neural machine translation of spoken\-dialects\.InProceedings of the 14th International Conference on Spoken Language Translation,S\. Sakti and M\. Utiyama \(Eds\.\),Tokyo, Japan,pp\. 82–89\.External Links:[Link](https://aclanthology.org/2017.iwslt-1.12/)Cited by:[§5](https://arxiv.org/html/2609.11399#S5.p1.1)\.
- Kocmiet al\.\(2025\)T\. Kocmi, E\. Artemova, E\. Avramidis, R\. Bawden, O\. Bojar, K\. Dranch, A\. Dvorkovich, S\. Dukanov, M\. Fishel, M\. Freitag, T\. Gowda, R\. Grundkiewicz, B\. Haddow, M\. Karpinska, P\. Koehn, H\. Lakougna, J\. Lundin, C\. Monz, K\. Murray, M\. Nagata, S\. Perrella, L\. Proietti, M\. Popel, M\. Popović, P\. Riley, M\. Shmatova, S\. Steingrímsson, L\. Yankovskaya, and V\. ZouharFindings of the WMT25 general machine translation shared task: time to stop evaluating on easy test sets\.InProceedings of the Tenth Conference on Machine Translation,B\. Haddow, T\. Kocmi, P\. Koehn, and C\. Monz \(Eds\.\),Suzhou, China,pp\. 355–413\.External Links:[Link](https://aclanthology.org/2025.wmt-1.22/),[Document](https://dx.doi.org/10.18653/v1/2025.wmt-1.22),ISBN 979\-8\-89176\-341\-8Cited by:[§1](https://arxiv.org/html/2609.11399#S1.p1.1)\.
- Kocmiet al\.\(2024\)T\. Kocmi, E\. Avramidis, R\. Bawden, O\. Bojar, A\. Dvorkovich, C\. Federmann, M\. Fishel, M\. Freitag, T\. Gowda, R\. Grundkiewicz, B\. Haddow, M\. Karpinska, P\. Koehn, B\. Marie, C\. Monz, K\. Murray, M\. Nagata, M\. Popel, M\. Popović, M\. Shmatova, S\. Steingrímsson, and V\. ZouharFindings of the WMT24 general machine translation shared task: the LLM era is here but MT is not solved yet\.InProceedings of the Ninth Conference on Machine Translation,B\. Haddow, T\. Kocmi, P\. Koehn, and C\. Monz \(Eds\.\),Miami, Florida, USA,pp\. 1–46\.External Links:[Link](https://aclanthology.org/2024.wmt-1.1/),[Document](https://dx.doi.org/10.18653/v1/2024.wmt-1.1)Cited by:[§1](https://arxiv.org/html/2609.11399#S1.p1.1)\.
- Kulkarni \(2015\)A\. KulkarniTED Multilingual Parallel Corpus\.Note:GitHubAccessed on 08, Dec 2025External Links:[Link](https://github.com/ajinkyakulkarni14/TED-Multilingual-Parallel-Corpus)Cited by:[§2\.1](https://arxiv.org/html/2609.11399#S2.SS1.SSS0.Px1.p1.1),[Ethical Considerations](https://arxiv.org/html/2609.11399#Sx2.p1.1)\.
- Kwonet al\.\(2023\)W\. Kwon, Z\. Li, S\. Zhuang, Y\. Sheng, L\. Zheng, C\. H\. Yu, J\. Gonzalez, H\. Zhang, and I\. StoicaEfficient memory management for large language model serving with pagedattention\.InProceedings of the 29th Symposium on Operating Systems Principles,SOSP ’23,New York, NY, USA,pp\. 611–626\.External Links:ISBN 9798400702297,[Link](https://doi.org/10.1145/3600006.3613165),[Document](https://dx.doi.org/10.1145/3600006.3613165)Cited by:[Appendix B](https://arxiv.org/html/2609.11399#A2.p1.1)\.
- Liet al\.\(2024\)J\. Li, H\. Zhou, S\. Huang, S\. Cheng, and J\. ChenEliciting the translation ability of large language models via multilingual finetuning with translation instructions\.Transactions of the Association for Computational Linguistics12,pp\. 576–592\.External Links:ISSN 2307\-387X,[Document](https://dx.doi.org/10.1162/tacl%5Fa%5F00655),[Link](https://doi.org/10.1162/tacl_a_00655),https://direct\.mit\.edu/tacl/article\-pdf/doi/10\.1162/tacl\_a\_00655/2367429/tacl\_a\_00655\.pdfCited by:[§1](https://arxiv.org/html/2609.11399#S1.p4.1)\.
- Longet al\.\(2024\)L\. Long, R\. Wang, R\. Xiao, J\. Zhao, X\. Ding, G\. Chen, and H\. WangOn LLMs\-driven synthetic data generation, curation, and evaluation: a survey\.InFindings of the Association for Computational Linguistics: ACL 2024,L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 11065–11082\.External Links:[Link](https://aclanthology.org/2024.findings-acl.658/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-acl.658)Cited by:[§5](https://arxiv.org/html/2609.11399#S5.p1.1)\.
- Meta AI \(2024\)Meta AILlama 3\.2: Revolutionizing edge AI and vision with open, customizable models\.Note:Accessed on 08, Dec 2025External Links:[Link](https://ai.meta.com/blog/llama-3-2-connect-2024-vision-edge-mobile-devices/)Cited by:[§2\.1](https://arxiv.org/html/2609.11399#S2.SS1.SSS0.Px3.p1.1)\.
- Meta AI \(2026\)Meta AIExpanding Translations to More Languages to Help You Reach Bigger Audiences on Reels\.Note:Accessed on 08, May 2026External Links:[Link](https://creators.instagram.com/blog/meta-ai-translations?locale=en)Cited by:[§1](https://arxiv.org/html/2609.11399#S1.p1.1)\.
- Nadǎşet al\.\(2025\)M\. Nadǎş, L\. Dioşan, and A\. TomescuSynthetic data generation using large language models: advances in text and code\.IEEE Access13\(\),pp\. 134615–134633\.External Links:[Document](https://dx.doi.org/10.1109/ACCESS.2025.3589503)Cited by:[§5](https://arxiv.org/html/2609.11399#S5.p1.1)\.
- OpenAIet al\.\(2024\)OpenAI, A\. Hurst, A\. Lerer, A\. P\. Goucher, A\. Perelman, A\. Ramesh, A\. Clark, A\. J\. Ostrow, A\. Welihinda, A\. Hayes, A\. Radford, A\. Mądry, A\. Baker\-Whitcomb, A\. Beutel, A\. Borzunov, A\. Carney, A\. Chow, A\. Kirillov, A\. Nichol, A\. Paino, A\. Renzin, A\. T\. Passos, A\. Kirillov, A\. Christakis, A\. Conneau, A\. Kamali, A\. Jabri, A\. Moyer, A\. Tam, A\. Crookes, A\. Tootoochian, A\. Tootoonchian, A\. Kumar, A\. Vallone, A\. Karpathy, A\. Braunstein, A\. Cann, A\. Codispoti, A\. Galu, A\. Kondrich, A\. Tulloch, A\. Mishchenko, A\. Baek, A\. Jiang, A\. Pelisse, A\. Woodford, A\. Gosalia, A\. Dhar, A\. Pantuliano, A\. Nayak, A\. Oliver, B\. Zoph, B\. Ghorbani, B\. Leimberger, B\. Rossen, B\. Sokolowsky, B\. Wang, B\. Zweig, B\. Hoover, B\. Samic, B\. McGrew, B\. Spero, B\. Giertler, B\. Cheng, B\. Lightcap, B\. Walkin, B\. Quinn, B\. Guarraci, B\. Hsu, B\. Kellogg, B\. Eastman, C\. Lugaresi, C\. Wainwright, C\. Bassin, C\. Hudson, C\. Chu, C\. Nelson, C\. Li, C\. J\. Shern, C\. Conger, C\. Barette, C\. Voss, C\. Ding, C\. Lu, C\. Zhang, C\. Beaumont, C\. Hallacy, C\. Koch, C\. Gibson, C\. Kim, C\. Choi, C\. McLeavey, C\. Hesse, C\. Fischer, C\. Winter, C\. Czarnecki, C\. Jarvis, C\. Wei, C\. Koumouzelis, D\. Sherburn, D\. Kappler, D\. Levin, D\. Levy, D\. Carr, D\. Farhi, D\. Mely, D\. Robinson, D\. Sasaki, D\. Jin, D\. Valladares, D\. Tsipras, D\. Li, D\. P\. Nguyen, D\. Findlay, E\. Oiwoh, E\. Wong, E\. Asdar, E\. Proehl, E\. Yang, E\. Antonow, E\. Kramer, E\. Peterson, E\. Sigler, E\. Wallace, E\. Brevdo, E\. Mays, F\. Khorasani, F\. P\. Such, F\. Raso, F\. Zhang, F\. von Lohmann, F\. Sulit, G\. Goh, G\. Oden, G\. Salmon, G\. Starace, G\. Brockman, H\. Salman, H\. Bao, H\. Hu, H\. Wong, H\. Wang, H\. Schmidt, H\. Whitney, H\. Jun, H\. Kirchner, H\. P\. d\. O\. Pinto, H\. Ren, H\. Chang, H\. W\. Chung, I\. Kivlichan, I\. O’Connell, I\. O’Connell, I\. Osband, I\. Silber, I\. Sohl, I\. Okuyucu, I\. Lan, I\. Kostrikov, I\. Sutskever, I\. Kanitscheider, I\. Gulrajani, J\. Coxon, J\. Menick, J\. Pachocki, J\. Aung, J\. Betker, J\. Crooks, J\. Lennon, J\. Kiros, J\. Leike, J\. Park, J\. Kwon, J\. Phang, J\. Teplitz, J\. Wei, J\. Wolfe, J\. Chen, J\. Harris, J\. Varavva, J\. G\. Lee, J\. Shieh, J\. Lin, J\. Yu, J\. Weng, J\. Tang, J\. Yu, J\. Jang, J\. Q\. Candela, J\. Beutler, J\. Landers, J\. Parish, J\. Heidecke, J\. Schulman, J\. Lachman, J\. McKay, J\. Uesato, J\. Ward, J\. W\. Kim, J\. Huizinga, J\. Sitkin, J\. Kraaijeveld, J\. Gross, J\. Kaplan, J\. Snyder, J\. Achiam, J\. Jiao, J\. Lee, J\. Zhuang, J\. Harriman, K\. Fricke, K\. Hayashi, K\. Singhal, K\. Shi, K\. Karthik, K\. Wood, K\. Rimbach, K\. Hsu, K\. Nguyen, K\. Gu\-Lemberg, K\. Button, K\. Liu, K\. Howe, K\. Muthukumar, K\. Luther, L\. Ahmad, L\. Kai, L\. Itow, L\. Workman, L\. Pathak, L\. Chen, L\. Jing, L\. Guy, L\. Fedus, L\. Zhou, L\. Mamitsuka, L\. Weng, L\. McCallum, L\. Held, L\. Ouyang, L\. Feuvrier, L\. Zhang, L\. Kondraciuk, L\. Kaiser, L\. Hewitt, L\. Metz, L\. Doshi, M\. Aflak, M\. Simens, M\. Boyd, M\. Thompson, M\. Dukhan, M\. Chen, M\. Gray, M\. Hudnall, M\. Zhang, M\. Aljubeh, M\. Litwin, M\. Zeng, M\. Johnson, M\. Shetty, M\. Gupta, M\. Shah, M\. Yatbaz, M\. J\. Yang, M\. Zhong, M\. Glaese, M\. Chen, M\. Janner, M\. Lampe, M\. Petrov, M\. Wu, M\. Wang, M\. Fradin, M\. Pokrass, M\. Castro, M\. O\. T\. de Castro, M\. Pavlov, M\. Brundage, M\. Wang, M\. Khan, M\. Murati, M\. Bavarian, M\. Lin, M\. Yesildal, N\. Soto, N\. Gimelshein, N\. Cone, N\. Staudacher, N\. Summers, N\. LaFontaine, N\. Chowdhury, N\. Ryder, N\. Stathas, N\. Turley, N\. Tezak, N\. Felix, N\. Kudige, N\. Keskar, N\. Deutsch, N\. Bundick, N\. Puckett, O\. Nachum, O\. Okelola, O\. Boiko, O\. Murk, O\. Jaffe, O\. Watkins, O\. Godement, O\. Campbell\-Moore, P\. Chao, P\. McMillan, P\. Belov, P\. Su, P\. Bak, P\. Bakkum, P\. Deng, P\. Dolan, P\. Hoeschele, P\. Welinder, P\. Tillet, P\. Pronin, P\. Tillet, P\. Dhariwal, Q\. Yuan, R\. Dias, R\. Lim, R\. Arora, R\. Troll, R\. Lin, R\. G\. Lopes, R\. Puri, R\. Miyara, R\. Leike, R\. Gaubert, R\. Zamani, R\. Wang, R\. Donnelly, R\. Honsby, R\. Smith, R\. Sahai, R\. Ramchandani, R\. Huet, R\. Carmichael, R\. Zellers, R\. Chen, R\. Chen, R\. Nigmatullin, R\. Cheu, S\. Jain, S\. Altman, S\. Schoenholz, S\. Toizer, S\. Miserendino, S\. Agarwal, S\. Culver, S\. Ethersmith, S\. Gray, S\. Grove, S\. Metzger, S\. Hermani, S\. Jain, S\. Zhao, S\. Wu, S\. Jomoto, S\. Wu, Shuaiqi, Xia, S\. Phene, S\. Papay, S\. Narayanan, S\. Coffey, S\. Lee, S\. Hall, S\. Balaji, T\. Broda, T\. Stramer, T\. Xu, T\. Gogineni, T\. Christianson, T\. Sanders, T\. Patwardhan, T\. Cunninghman, T\. Degry, T\. Dimson, T\. Raoux, T\. Shadwell, T\. Zheng, T\. Underwood, T\. Markov, T\. Sherbakov, T\. Rubin, T\. Stasi, T\. Kaftan, T\. Heywood, T\. Peterson, T\. Walters, T\. Eloundou, V\. Qi, V\. Moeller, V\. Monaco, V\. Kuo, V\. Fomenko, W\. Chang, W\. Zheng, W\. Zhou, W\. Manassra, W\. Sheu, W\. Zaremba, Y\. Patil, Y\. Qian, Y\. Kim, Y\. Cheng, Y\. Zhang, Y\. He, Y\. Zhang, Y\. Jin, Y\. Dai, and Y\. MalkovGPT\-4o system card\.arXiv preprint\.External Links:2410\.21276,[Link](https://arxiv.org/abs/2410.21276)Cited by:[§5](https://arxiv.org/html/2609.11399#S5.p1.1)\.
- Park and Zhao \(2019\)J\. Park and H\. ZhaoKorean\-to\-Chinese Machine Translation using Chinese Character as Pivot Clue\.arXiv preprint\.External Links:1911\.11008,[Link](https://arxiv.org/abs/1911.11008)Cited by:[§2\.1](https://arxiv.org/html/2609.11399#S2.SS1.SSS0.Px1.p1.1),[Ethical Considerations](https://arxiv.org/html/2609.11399#Sx2.p1.1)\.
- Poncelaset al\.\(2018\)A\. Poncelas, D\. Shterionov, A\. Way, G\. Maillette de Buy Wenniger, and P\. PassbanInvestigating backtranslation in neural machine translation\.InProceedings of the 21st Annual Conference of the European Association for Machine Translation,J\. A\. Pérez\-Ortiz, F\. Sánchez\-Martínez, M\. Esplà\-Gomis, M\. Popović, C\. Rico, A\. Martins, J\. Van den Bogaert, and M\. L\. Forcada \(Eds\.\),Alicante, Spain,pp\. 269–278\.External Links:[Link](https://aclanthology.org/2018.eamt-main.25/)Cited by:[§5](https://arxiv.org/html/2609.11399#S5.p1.1)\.
- Qinet al\.\(2024\)Y\. Qin, K\. Song, Y\. Hu, W\. Yao, S\. Cho, X\. Wang, X\. Wu, F\. Liu, P\. Liu, and D\. YuInFoBench: evaluating instruction following ability in large language models\.InFindings of the Association for Computational Linguistics: ACL 2024,L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 13025–13048\.External Links:[Link](https://aclanthology.org/2024.findings-acl.772/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-acl.772)Cited by:[§1](https://arxiv.org/html/2609.11399#S1.p4.1)\.
- Qwen Team \(2024\)Qwen TeamQwen2\.5: A Party of Foundation Models\!\.Note:Accessed on 08, Dec 2025External Links:[Link](https://qwenlm.github.io/blog/qwen2.5/)Cited by:[§2\.1](https://arxiv.org/html/2609.11399#S2.SS1.SSS0.Px3.p1.1)\.
- Qwen Team \(2025\)Qwen TeamQwen3: Think Deeper, Act Faster\.Note:Accessed on 08, Dec 2025External Links:[Link](https://qwenlm.github.io/blog/qwen3/)Cited by:[§2\.1](https://arxiv.org/html/2609.11399#S2.SS1.SSS0.Px3.p1.1)\.
- Qwen Team \(2026\)Qwen TeamQwen3\.5: Accelerating Productivity with Native Multimodal Agents\.External Links:[Link](https://qwen.ai/blog?id=qwen3.5)Cited by:[§3\.2\.2](https://arxiv.org/html/2609.11399#S3.SS2.SSS2.p1.1)\.
- Reiet al\.\(2025\)R\. Rei, N\. M\. Guerreiro, J\. Pombal, J\. Alves, P\. Teixeirinha, A\. Farajian, and A\. F\. T\. MartinsTower\+: bridging generality and translation specialization in multilingual LLMs\.arXiv preprint\.External Links:2506\.17080,[Link](https://arxiv.org/abs/2412.04261)Cited by:[§2\.1](https://arxiv.org/html/2609.11399#S2.SS1.SSS0.Px3.p1.1)\.
- Reiet al\.\(2022\)R\. Rei, M\. Treviso, N\. M\. Guerreiro, C\. Zerva, A\. C\. Farinha, C\. Maroti, J\. G\. C\. de Souza, T\. Glushkova, D\. Alves, L\. Coheur, A\. Lavie, and A\. F\. T\. MartinsCometKiwi: IST\-unbabel 2022 submission for the quality estimation shared task\.InProceedings of the Seventh Conference on Machine Translation \(WMT\),P\. Koehn, L\. Barrault, O\. Bojar, F\. Bougares, R\. Chatterjee, M\. R\. Costa\-jussà, C\. Federmann, M\. Fishel, A\. Fraser, M\. Freitag, Y\. Graham, R\. Grundkiewicz, P\. Guzman, B\. Haddow, M\. Huck, A\. Jimeno Yepes, T\. Kocmi, A\. Martins, M\. Morishita, C\. Monz, M\. Nagata, T\. Nakazawa, M\. Negri, A\. Névéol, M\. Neves, M\. Popel, M\. Turchi, and M\. Zampieri \(Eds\.\),Abu Dhabi, United Arab Emirates \(Hybrid\),pp\. 634–645\.External Links:[Link](https://aclanthology.org/2022.wmt-1.60/),[Document](https://dx.doi.org/10.18653/v1/2022.wmt-1.60)Cited by:[§4\.1](https://arxiv.org/html/2609.11399#S4.SS1.p2.1)\.
- Scherreret al\.\(2014\)Y\. Scherrer, L\. Nerima, L\. Russo, M\. Ivanova, and E\. WehrliSwissAdmin: a multilingual tagged parallel corpus of press releases\.InProceedings of the Ninth International Conference on Language Resources and Evaluation \(LREC’14\),N\. Calzolari, K\. Choukri, T\. Declerck, H\. Loftsson, B\. Maegaard, J\. Mariani, A\. Moreno, J\. Odijk, and S\. Piperidis \(Eds\.\),Reykjavik, Iceland,pp\. 1832–1836\.External Links:[Link](https://aclanthology.org/L14-1602/)Cited by:[§2\.1](https://arxiv.org/html/2609.11399#S2.SS1.SSS0.Px1.p1.1),[Ethical Considerations](https://arxiv.org/html/2609.11399#Sx2.p1.1)\.
- Singhet al\.\(2025\)A\. Singh, A\. Fry, A\. Perelman, A\. Tart, A\. Ganesh, A\. El\-Kishky, A\. McLaughlin, A\. Low, A\. J\. Ostrow, A\. Ananthram, A\. Nathan, A\. Luo, A\. Helyar, A\. Madry, A\. Efremov, A\. Spyra, A\. Baker\-Whitcomb, A\. Beutel, A\. Karpenko, A\. Makelov, A\. Neitz, A\. Wei, A\. Barr, A\. Kirchmeyer, A\. Ivanov, A\. Christakis, A\. Gillespie, A\. Tam, A\. Bennett, A\. Wan, A\. Huang, A\. M\. Sandjideh, A\. Yang, A\. Kumar, A\. Saraiva, A\. Vallone, A\. Gheorghe, A\. G\. Garcia, A\. Braunstein, A\. Liu, A\. Schmidt, A\. Mereskin, A\. Mishchenko, A\. Applebaum, A\. Rogerson, A\. Rajan, A\. Wei, A\. Kotha, A\. Srivastava, A\. Agrawal, A\. Vijayvergiya, A\. Tyra, A\. Nair, A\. Nayak, B\. Eggers, B\. Ji, B\. Hoover, B\. Chen, B\. Chen, B\. Barak, B\. Minaiev, B\. Hao, B\. Baker, B\. Lightcap, B\. McKinzie, B\. Wang, B\. Quinn, B\. Fioca, B\. Hsu, B\. Yang, B\. Yu, B\. Zhang, B\. Brenner, C\. R\. Zetino, C\. Raymond, C\. Lugaresi, C\. Paz, C\. Hudson, C\. Whitney, C\. Li, C\. Chen, C\. Cole, C\. Voss, C\. Ding, C\. Shen, C\. Huang, C\. Colby, C\. Hallacy, C\. Koch, C\. Lu, C\. Kaplan, C\. Kim, C\. J\. Minott\-Henriques, C\. Frey, C\. Yu, C\. Czarnecki, C\. Reid, C\. Wei, C\. Decareaux, C\. Scheau, C\. Zhang, C\. Forbes, D\. Tang, D\. Goldberg, D\. Roberts, D\. Palmie, D\. Kappler, D\. Levine, D\. Wright, D\. Leo, D\. Lin, D\. Robinson, D\. Grabb, D\. Chen, D\. Lim, D\. Salama, D\. Bhattacharjee, D\. Tsipras, D\. Li, D\. Yu, D\. J\. Strouse, D\. Williams, D\. Hunn, E\. Bayes, E\. Arbus, E\. Akyurek, E\. Y\. Le, E\. Widmann, E\. Yani, E\. Proehl, E\. Sert, E\. Cheung, E\. Schwartz, E\. Han, E\. Jiang, E\. Mitchell, E\. Sigler, E\. Wallace, E\. Ritter, E\. Kavanaugh, E\. Mays, E\. Nikishin, F\. Li, F\. P\. Such, F\. d\. A\. B\. Peres, F\. Raso, F\. Bekerman, F\. Tsimpourlas, F\. Chantzis, F\. Song, F\. Zhang, G\. Raila, G\. McGrath, G\. Briggs, G\. Yang, G\. Parascandolo, G\. Chabot, G\. Kim, G\. Zhao, G\. Valiant, G\. Leclerc, H\. Salman, H\. Wang, H\. Sheng, H\. Jiang, H\. Wang, H\. Jin, H\. Sikchi, H\. Schmidt, H\. Aspegren, H\. Chen, H\. Qiu, H\. Lightman, I\. Covert, I\. Kivlichan, I\. Silber, I\. Sohl, I\. Hammoud, I\. Clavera, I\. Lan, I\. Akkaya, I\. Kostrikov, I\. Kofman, I\. Etinger, I\. Singal, J\. Hehir, J\. Huh, J\. Pan, J\. Wilczynski, J\. Pachocki, J\. Lee, J\. Quinn, J\. Kiros, J\. Kalra, J\. Samaroo, J\. Wang, J\. Wolfe, J\. Chen, J\. Wang, J\. Harb, J\. Han, J\. Wang, J\. Zhao, J\. Chen, J\. Yang, J\. Tworek, J\. Chand, J\. Landon, J\. Liang, J\. Lin, J\. Liu, J\. Wang, J\. Tang, J\. Yin, J\. Jang, J\. Morris, J\. Flynn, J\. Ferstad, J\. Heidecke, J\. Fishbein, J\. Hallman, J\. Grant, J\. Chien, J\. Gordon, J\. Park, J\. Liss, J\. Kraaijeveld, J\. Guay, J\. Mo, J\. Lawson, J\. McGrath, J\. Vendrow, J\. Jiao, J\. Lee, J\. Steele, J\. Wang, J\. Mao, K\. Chen, K\. Hayashi, K\. Xiao, K\. Salahi, K\. Wu, K\. Sekhri, K\. Sharma, K\. Singhal, K\. Li, K\. Nguyen, K\. Gu\-Lemberg, K\. King, K\. Liu, K\. Stone, K\. Yu, K\. Ying, K\. Georgiev, K\. Lim, K\. Tirumala, K\. Miller, L\. Ahmad, L\. Lv, L\. Clare, L\. Fauconnet, L\. Itow, L\. Yang, L\. Romaniuk, L\. Anise, L\. Byron, L\. Pathak, L\. Maksin, L\. Lo, L\. Ho, L\. Jing, L\. Wu, L\. Xiong, L\. Mamitsuka, L\. Yang, L\. McCallum, L\. Held, L\. Bourgeois, L\. Engstrom, L\. Kuhn, L\. Feuvrier, L\. Zhang, L\. Switzer, L\. Kondraciuk, L\. Kaiser, M\. Joglekar, M\. Singh, M\. Shah, M\. Stratta, M\. Williams, M\. Chen, M\. Sun, M\. Cayton, M\. Li, M\. Zhang, M\. Aljubeh, M\. Nichols, M\. Haines, M\. Schwarzer, M\. Gupta, M\. Shah, M\. Y\. Guan, M\. Huang, M\. Dong, M\. Wang, M\. Glaese, M\. Carroll, M\. Lampe, M\. Malek, M\. Sharman, M\. Zhang, M\. Wang, M\. Pokrass, M\. Florian, M\. Pavlov, M\. Wang, M\. Chen, M\. Wang, M\. Feng, M\. Bavarian, M\. Lin, M\. Abdool, M\. Rohaninejad, N\. Soto, N\. Staudacher, N\. LaFontaine, N\. Marwell, N\. Liu, N\. Preston, N\. Turley, N\. Ansman, N\. Blades, N\. Pancha, N\. Mikhaylin, N\. Felix, N\. Handa, N\. Rai, N\. Keskar, N\. Brown, O\. Nachum, O\. Boiko, O\. Murk, O\. Watkins, O\. Gleeson, P\. Mishkin, P\. Lesiewicz, P\. Baltescu, P\. Belov, P\. Zhokhov, P\. Pronin, P\. Guo, P\. Thacker, Q\. Liu, Q\. Yuan, Q\. Liu, R\. Dias, R\. Puckett, R\. Arora, R\. T\. Mullapudi, R\. Gaon, R\. Miyara, R\. Song, R\. Aggarwal, R\. J\. Marsan, R\. Yemiru, R\. Xiong, R\. Kshirsagar, R\. Nuttall, R\. Tsiupa, R\. Eldan, R\. Wang, R\. James, R\. Ziv, R\. Shu, R\. Nigmatullin, S\. Jain, S\. Talaie, S\. Altman, S\. Arnesen, S\. Toizer, S\. Toyer, S\. Miserendino, S\. Agarwal, S\. Yoo, S\. Heon, S\. Ethersmith, S\. Grove, S\. Taylor, S\. Bubeck, S\. Banesiu, S\. Amdo, S\. Zhao, S\. Wu, S\. Santurkar, S\. Zhao, S\. R\. Chaudhuri, S\. Krishnaswamy, Shuaiqi, Xia, S\. Cheng, S\. Anadkat, S\. P\. Fishman, S\. Tobin, S\. Fu, S\. Jain, S\. Mei, S\. Egoian, S\. Kim, S\. Golden, S\. Q\. Mah, S\. Lin, S\. Imm, S\. Sharpe, S\. Yadlowsky, S\. Choudhry, S\. Eum, S\. Sanjeev, T\. Khan, T\. Stramer, T\. Wang, T\. Xin, T\. Gogineni, T\. Christianson, T\. Sanders, T\. Patwardhan, T\. Degry, T\. Shadwell, T\. Fu, T\. Gao, T\. Garipov, T\. Sriskandarajah, T\. Sherbakov, T\. Korbak, T\. Kaftan, T\. Hiratsuka, T\. Wang, T\. Song, T\. Zhao, T\. Peterson, V\. Kharitonov, V\. Chernova, V\. Kosaraju, V\. Kuo, V\. Pong, V\. Verma, V\. Petrov, W\. Jiang, W\. Zhang, W\. Zhou, W\. Xie, W\. Zhan, W\. McCabe, W\. DePue, W\. Ellsworth, W\. Bain, W\. Thompson, X\. Chen, X\. Qi, X\. Xiang, X\. Shi, Y\. Dubois, Y\. Yu, Y\. Khakbaz, Y\. Wu, Y\. Qian, Y\. T\. Lee, Y\. Chen, Y\. Zhang, Y\. Xiong, Y\. Tian, Y\. Cha, Y\. Bai, Y\. Yang, Y\. Yuan, Y\. Li, Y\. Zhang, Y\. Yang, Y\. Jin, Y\. Jiang, Y\. Wang, Y\. Wang, Y\. Liu, Z\. Stubenvoll, Z\. Dou, Z\. Wu, and Z\. WangOpenAI GPT\-5 system card\.arXiv preprint\.External Links:2601\.03267,[Link](https://arxiv.org/abs/2601.03267)Cited by:[§3\.1](https://arxiv.org/html/2609.11399#S3.SS1.SSS0.Px1.p2.1)\.
- Unbabel \(2025\)UnbabelCOMET\.Note:GitHubAccessed on 09, May 2026External Links:[Link](https://github.com/Unbabel/COMET)Cited by:[Appendix I](https://arxiv.org/html/2609.11399#A9.p1.1)\.
- Vilaret al\.\(2023\)D\. Vilar, M\. Freitag, C\. Cherry, J\. Luo, V\. Ratnakar, and G\. FosterPrompting PaLM for translation: assessing strategies and performance\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),A\. Rogers, J\. Boyd\-Graber, and N\. Okazaki \(Eds\.\),Toronto, Canada,pp\. 15406–15427\.External Links:[Link](https://aclanthology.org/2023.acl-long.859/),[Document](https://dx.doi.org/10.18653/v1/2023.acl-long.859)Cited by:[§1](https://arxiv.org/html/2609.11399#S1.p1.1)\.
- Wanget al\.\(2024\)W\. Wang, Z\. Li, D\. Lian, C\. Ma, L\. Song, and Y\. WeiMitigating the language mismatch and repetition issues in LLM\-based machine translation via model editing\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),Miami, Florida, USA,pp\. 15681–15700\.External Links:[Link](https://aclanthology.org/2024.emnlp-main.879/),[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.879)Cited by:[§1](https://arxiv.org/html/2609.11399#S1.p4.1)\.
- Wolfet al\.\(2020\)T\. Wolf, L\. Debut, V\. Sanh, J\. Chaumond, C\. Delangue, A\. Moi, P\. Cistac, T\. Rault, R\. Louf, M\. Funtowicz, J\. Davison, S\. Shleifer, P\. von Platen, C\. Ma, Y\. Jernite, J\. Plu, C\. Xu, T\. Le Scao, S\. Gugger, M\. Drame, Q\. Lhoest, and A\. RushTransformers: state\-of\-the\-art natural language processing\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations,Q\. Liu and D\. Schlangen \(Eds\.\),Online,pp\. 38–45\.External Links:[Link](https://aclanthology.org/2020.emnlp-demos.6/),[Document](https://dx.doi.org/10.18653/v1/2020.emnlp-demos.6)Cited by:[Appendix B](https://arxiv.org/html/2609.11399#A2.p1.1)\.
- Xuet al\.\(2024\)H\. Xu, A\. Sharaf, Y\. Chen, W\. Tan, L\. Shen, B\. Van Durme, K\. Murray, and Y\. J\. KimContrastive preference optimization: pushing the boundaries of llm performance in machine translation\.InProceedings of the 41st International Conference on Machine Learning,ICML’24\.Cited by:[§1](https://arxiv.org/html/2609.11399#S1.p1.1)\.
- Zhanget al\.\(2023\)B\. Zhang, B\. Haddow, and A\. BirchPrompting large language model for machine translation: a case study\.InProceedings of the 40th International Conference on Machine Learning,ICML’23\.Cited by:[§1](https://arxiv.org/html/2609.11399#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.11399#S2.SS1.SSS0.Px2.p1.1)\.
- Zhanget al\.\(2025\)B\. Zhang, F\. Moiseev, J\. Ainslie, P\. Suganthan, M\. Ma, S\. Bhupatiraju, F\. Lebron, O\. Firat, A\. Joulin, and Z\. DongEncoder\-decoder Gemma: Improving the quality\-efficiency trade\-off via adaptation\.arXiv preprint\.External Links:2504\.06225,[Link](https://arxiv.org/abs/2504.06225)Cited by:[§2\.1](https://arxiv.org/html/2609.11399#S2.SS1.SSS0.Px3.p1.1)\.
- Zhouet al\.\(2023\)J\. Zhou, T\. Lu, S\. Mishra, S\. Brahma, S\. Basu, Y\. Luan, D\. Zhou, and L\. HouInstruction\-following evaluation for large language models\.External Links:2311\.07911,[Link](https://arxiv.org/abs/2311.07911)Cited by:[§1](https://arxiv.org/html/2609.11399#S1.p4.1)\.
## Appendix AAppendix: Additional Tables for Data and Models
Table A\.1:The size of our test set for each language pair and their corresponding sources\.Table A\.2:Model details including names, architectures, size and either instruction\-tuned or reasoning variants\.
## Appendix BAppendix: LLM Inference Details
We used vLLM[Kwon et al\. \(2023\)](https://arxiv.org/html/2609.11399#bib.bib10)for inference with most models with the exception of DeepSeek\-V3\.2\-Exp\([DeepSeek\-AI et al\., 2025](https://arxiv.org/html/2609.11399#bib.bib16)\)andt5gemma\-xl\-xl\-prefixlm\-it\. For these models, we obtained inference results using the respective API or the HuggingFace Transformers library[Wolf et al\. \(2020\)](https://arxiv.org/html/2609.11399#bib.bib38)\. We kept the default values of the hyperparameters with temperature and top\_p both set to 1\. With the exception of DeepSeek\-V3\.2\-Exp, all models were run without quantization on 4 NVIDIA GH200 GPUs\. On average, an instruction\-tuned model requires approximately 10 minutes to process one language pair \(3,000 instances\), whereas a reasoning model requires about 18 minutes\. For reasoning LLMs, only content after the reasoning tags \(i\.e\.,¡think¿¡/think¿\) is treated as “LLM outputs” for noise analysis\.
## Appendix CAppendix: Regular Expressions for Rule\-based Detector
We use the following regular expression patterns to detect explanatory or meta\-linguistic content\.
### C\.1Common Explanation Phrases
r’\\b\(the␣translation␣is\|here␣is\|here\\’s\|thistranslatesto\|translation:\|translatedtext:\|output:\|target:\)\\b’,
␣␣␣␣r’\\b\(in\\w\+\(this\|it\)\(means\|says\|translates\)\)\\b’,
␣␣␣␣r’\\b\(notethat\|pleasenote\|itshouldbenoted\)\\b’,
␣␣␣␣r’\\b\(explanation\|reasoning\|analysis\|breakdown\)\\b’,
␣␣␣␣r’\\n\\s\*\(translation\|explanation\|note\|original\|source\|target\)\\s\*:’,’
### C\.2Meta\-linguistic Markers
r’\\b\(literally\|figuratively\|idiomatically\|contextually\)\\b’,
r’\\b\(this␣\(word\|phrase\|sentence\|text\)\)\\b’,
r’\\b\(means\|refers␣to\|indicates\|suggests\)\\b\.\*\\b\(that\|which\)\\b’,
### C\.3Comments to user
r’\\b\(hope␣this␣helps\|let␣me␣know\|feel␣free\|if␣you\|you␣can\)\\b’,
r’\\b\(please\|kindly\|note:\|important:\)\\b’,
### C\.4Thinking Markers
r’<think\>\|</think\>\|<thought\>\|</thought\>’,
r’\\\*\\\*reasoning\\\*\\\*\|\\\*\\\*analysis\\\*\\\*\|\\\*\\\*explanation\\\*\\\*’,
### C\.5Markdown or XML Tags
r’^\#\+\\s\+’,
r’<\[a\-zA\-Z\]\+\>\.\*</\[a\-zA\-Z\]\+\>’,
### C\.6Lists or Parenthetical Explanations
r’^\\s\*\[\\d\\\-\\\*\]\+\[\\\.\\\)\]\\s\+’,
r’\\\(\[^\)\]\{50,\}\\\)’
## Appendix DAppendix: Examples of TransClean
### D\.1An Example of the Synthetic Subset
"source": "Die Wasserqualität hat sich in den letzten Jahrzehnten deutlich verbessert\.","translation": "The water quality has greatly improved over the past decades\.","src\_lang": "de", "tgt\_lang": "fr","noise\_pattern": "wrong\_language","gold\_reference": "La qualité de l’eau s’est sensiblement améliorée au cours des dernières décennies\."
### D\.2An Example of the Curated Subset
"source": "Dog control laws to be reviewed in government consultation","translation": "English: Dog control laws to be reviewed in government consultation \\nChinese:\\cjkfont政府咨询将审查狗只控制法例","src\_lang": "en", "tgt\_lang": "zh","noise\_patterns": \["language\_prefix", "bilingual\_output"\], "primary\_pattern": "bilingual\_output","silver\_reference": "\\cjkfont政府咨询将审查狗只控制法例", "silver\_agreement": 3, "silver\_votes": \["\\cjkfont政府咨询将审查狗只控制法例", "\\cjkfont政府咨询将审查狗只控制法例", "\\cjkfont政府咨询将审查狗只控制法例"\]
## Appendix EAppendix: Prompt for Synthetic Noise Generation
SYSTEM PROMPTYou are simulating a large language model that generates translations with extra noise, explanations, or formatting artifacts\. Your task is to take a clean reference translation and add realistic noise to it — exactly as a helpful\-but\-verbose LLM would\.Rules:\- Output ONLY the noisy translation \(no meta\-commentary, no JSON\)\.\- The core translation meaning must remain correct\.\- The noise must look authentic, as if a real LLM produced it\.\- Use the target language for the translation itself; explanatory text may be in English or the source/target language depending on the pattern\.USER PROMPT\{NOISE PATTERN DESCRIPTION\}\{FEW\-SHOT EXAMPLES\}Source text: \{source\}Source language: \{src\_lang\_name\}Target language: \{tgt\_lang\_name\}Clean translation: \{clean\_translation\}Generate the noisy output \{BASED ON NOISE DESCRIPTION\}:
Figure E\.1:Prompt for generating synthetic noise\.
## Appendix FAppendix: Statistics of the Synthetic Subset
\(a\)Counts per noise pattern
\(b\)Combo pattern counts
Table F\.1:Detailed statistics of the synthetic noise data: \(a\) counts per noise pattern and \(b\) combo pattern counts\.
## Appendix GAppendix: Prompt for Noise Data Curation
SYSTEM PROMPTYou are a translation quality analyst\. Your job is to determine whether a machine translation output contains ONLY the translation, or whether it also contains extra content that should NOT be part of a clean translation\.You will be given:\- source: the original text\- src\_lang / tgt\_lang: language codes\- reference: a clean reference translation\- translation: the LLM\-generated translation to judgeA “noisy” translation contains one or more of these artifacts:1\. language\_prefix: A language name label before the translation, e\.g\. “Chinese: 你好”2\. verbose\_preamble: An introductory sentence like “Here is the translation…” or “Sure\! Here’s…”3\. translation\_prefix: A “Translation:” or “Translated:” label4\. explanation: Word\-by\-word breakdown, pinyin/romanization, grammar notes, or extended commentary after the translation \(usually separated by newlines\)5\. cultural\_note: Usually a parenthetical “\(Note: …\)” explaining cultural context, idioms, or translation choices6\. bilingual\_output: Both source and target languages appear with labels, or the source text is substantially repeated7\. alternatives: Multiple numbered translation options8\. code\_block: Translation wrapped in markdown code fences \(\`\`\`\)9\. special\_formatting: Double brackets \[\[…\]\], double braces \{\{…\}\}, or XML\-like tags10\. extra\_punctuation: Excessive repeated punctuation like “\!\!\!” or “???” that isn’t in the source11\. wrong\_language: The translation is in the wrong language entirely \(not the target language\)12\. off\_topic: The output is completely unrelated to translation \(e\.g\., code, random text, instructions\)A “clean” translation contains ONLY the translated text in the target language, possibly with minor differences from the reference \(which is fine — different valid translations exist\)\.IMPORTANT: Minor differences in word choice, sentence structure, or style between the translation and the reference do NOT make it noisy\. Only extra non\-translation content counts\.Respond with ONLY valid JSON in this exact format:\{“is\_noisy”: true or false,“confidence”: “high” or “medium” or “low”,“noise\_patterns”: \[“pattern1”, “pattern2”\] or \[\],“primary\_pattern”: “the most prominent pattern” or null,“reasoning”: “brief explanation in one sentence” \}USER PROMPTSource \(\{src\_lang\} → \{tgt\_lang\}\):\{source\}Reference translation:\{reference\}LLM translation to judge:\{translation\}Is this translation noisy? Respond with JSON only\.
Figure G\.1:Prompt for curating noise data\.
## Appendix HAppendix: Prompt for Clean Translation Annotation
SYSTEM PROMPTYou are a translation quality expert\. Your task is to extract the clean translation from a noisy machine translation output\.The noisy output may contain artifacts such as:\- Language prefixes or labels \(e\.g\. “Chinese: …”\)\- Verbose preambles \(e\.g\. “Here is the translation…”\)\- Explanations, grammar notes, or word\-by\-word breakdowns\- Cultural notes or parenthetical comments\- Multiple alternative translations\- Bilingual output with both source and target text\- Code blocks, special formatting, or extra punctuation\- Wrong language or off\-topic contentYou will be given the source text, language pair, a reference translation, the noisy LLM output, and the identified noise patterns\. Use all of this context to extract ONLY the clean translation in the target language\.If the output contains multiple translation alternatives, extract the best one\. If the output is entirely off\-topic or in the wrong language, return an empty string\.Respond with ONLY valid JSON in this exact format:\{“extracted\_translation”: “the clean translation text only”\}USER PROMPTSource \(\{src\_lang\} → \{tgt\_lang\}\):\{source\}Reference translation:\{reference\}Identified noise patterns: \{noise\_patterns\}Noisy LLM output to clean:\{translation\}Extract the clean translation\. Respond with JSON only\.
Figure H\.1:Prompt for annotating clean translation\.
## Appendix IAppendix: Details for Running Noise Extractors
We used vLLM for running LLM\-based extraction methods with temperature set as00and top\_p1\.01\.0on 4 NVIDIA GH200 GPUs\. For the span\-based extraction method, we ran COMET\-KIWI via the COMET repository\([Unbabel, 2025](https://arxiv.org/html/2609.11399#bib.bib26)\)on one NVIDIA A100 40BG GPU\. On average, an LLM takes approximately 25 minutes to process all \(9900\) instances, whereas the span\-based method costs about 70 minutes\.
## Appendix JAppendix: Prompt for LLM\-based Extraction
USER PROMPTAnalyze this machine translation output\. The source language is \{src\_lang\} and the target language is \{tgt\_lang\}\.1\. Extract ONLY the clean translation \(no explanations, labels, or formatting\)\.2\. Identify the noise pattern if the output contains noise\.Return a JSON object with these fields:\- “extracted\_translation”: the clean translation text only\- “noise\_pattern”: one of “none”, “language\_prefix”, “verbose\_preamble”, ”translation\_prefix”, “explanation”, “cultural\_note”, “bilingual\_output”, “alternatives”, “code\_block”, “special\_formatting”, “extra\_punctuation”, “wrong\_language”, “off\_topic”Machine translation output: \{translation\}JSON:
Figure J\.1:Prompt for LLM\-based Extraction\.Similar Articles
Last Translation Benchmark
The Last Translation Benchmark introduces a live dataset of peer-reviewed, multimodal examples designed to evaluate and break leading machine translation models, with handcrafted verification rules for reliable assessment. It addresses the saturation of current benchmarks and the unreliability of automatic metrics.
Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness
Introduces Cultivar, a contrastive localized translation benchmark for detecting data contamination and evaluating localization robustness in multilingual translation models. It benchmarks 32 open-weight models and finds that MT-specialised models are less robust, with potential overfitting to FLORES.
An Investigation of Translationese in the Generations of Multilingual Large Language Models
This paper investigates whether text generated by multilingual large language models exhibits signs of translationese, and compares it to human-written translations to assess linguistic naturalness.
Follow-up to my TranslateGemma-12b benchmark post: human reviewers flagged 71% of the segments automated metrics rated clean
A human review of TranslateGemma-12b's translations revealed that 71% of segments rated clean by automated metrics actually contained errors, highlighting significant gaps in metric-only evaluation for multilingual translation quality.
Hy-MT2: A Family of Fast, Efficient and Powerful Multilingual Translation Models in the Wild
Hy-MT2 is a family of fast, efficient multilingual translation models from Tencent, available in 1.8B, 7B, and 30B-A3B sizes, supporting 33 languages and outperforming previous open-source and commercial models.