跨职业机器翻译评估指标的性别偏见基准测试
摘要
本研究对跨职业机器翻译评估指标的性别偏见进行基准测试,揭示男性化翻译倾向于获得更高分数,且偏见因语言和评估者而异。
arXiv:2609.21490v1 Announce Type: new
Abstract: Gender bias remains a persistent concern in machine translation (MT), affecting both generated translations and their automatic evaluation. When a source text leaves a person's gender unspecified, translations may realize that person using masculine or feminine forms, and both MT systems and evaluation metrics may exhibit systematic preferences between these alternatives despite the source providing no basis for such a distinction. We study this behavior in the WMT 2026 Automated Translation Quality Evaluation Systems Shared Task using an occupation-balanced subset of GAMBIT+. We consider seven English-source language pairs, six from the original dataset, targeting Arabic, Czech, Greek, Icelandic, Russian, and Ukrainian, and extend the original resource with German. The subset contains 1,308 masculine/feminine translation pairs per target language, with three examples for each of the 436 ISCO-08 occupational groups. We evaluate shared-task submissions and baselines for score prediction and error annotation, examining the direction, magnitude, and frequency of gender-related differences. We find an overall tendency for masculine translations to receive higher scores, as well as differences per occupation following stereotypical gender representations, although the strength and consistency of this preference vary considerably across evaluators and languages. Our results show that gender bias remains present in MT evaluation, but that capturing its extent requires looking beyond a single aggregate measure to complementary dimensions of evaluator behavior.
查看缓存全文
缓存时间: 2026/09/21 09:08
# Benchmarking Gender Bias in Machine Translation Evaluation Metricsacross Occupations Source: [https://arxiv.org/html/2609.21490](https://arxiv.org/html/2609.21490) Orfeas Menis Mastromichalakis††thanks:Corresponding author:[menisorfeas@gmail\.com](mailto:[email protected])\.Giorgos FilandrianosAffiliation:National Technical University of Athens, GreeceWafaa MohammedAffiliation:University of Amsterdam, NetherlandsGiuseppe AttanasioAffiliation:Instituto de Telecomunicações, Lisbon, PortugalChrysoula ZervaAffiliation:National Technical University of Athens, Greece ###### Abstract Gender bias remains a persistent concern in machine translation \(MT\), affecting both generated translations and their automatic evaluation\. When a source text leaves a person’s gender unspecified, translations may realize that person using masculine or feminine forms, and both MT systems and evaluation metrics may exhibit systematic preferences between these alternatives despite the source providing no basis for such a distinction\. We study this behavior in the WMT 2026 Automated Translation Quality Evaluation Systems Shared Task using an occupation\-balanced subset of GAMBIT\+111The dataset is available at:[https://huggingface\.co/datasets/ailsntua/gambit\-plus](https://huggingface.co/datasets/ailsntua/gambit-plus)\.\. We consider seven English\-source language pairs, six from the original dataset, targeting Arabic, Czech, Greek, Icelandic, Russian, and Ukrainian, and extend the original resource with German\. The subset contains 1,308 masculine/feminine translation pairs per target language, with three examples for each of the 436 ISCO\-08 occupational groups\. We evaluate shared\-task submissions and baselines for score prediction and error annotation, examining the direction, magnitude, and frequency of gender\-related differences\. We find an overall tendency for masculine translations to receive higher scores, as well as differences per occupation following stereotypical gender representations, although the strength and consistency of this preference vary considerably across evaluators and languages\. Our results show that gender bias remains present in MT evaluation, but that capturing its extent requires looking beyond a single aggregate measure to complementary dimensions of evaluator behavior\. ## 1Introduction Gender\-related preferences can arise both in machine translation outputs and in the metrics used to assess their quality\. When a source text leaves a person’s gender unspecified, masculine and feminine translations may both be valid, yet MT systems may systematically favor one realization over the other\([Menis Mastromichalakis et al\., 2025b](https://arxiv.org/html/2609.21490#bib.bib9)\)\. This is common for gender\-ambiguous occupational terms in English, such asdoctor,teacher, orlegislator, where the source itself provides no evidence for assigning either gender\. Gender\-neutral or gender\-fair formulations may also be possible in many cases; however, the present work focuses specifically on the contrast between masculine and feminine realizations\. The same concern extends to automatic evaluation\. If two translations preserve the information expressed in the source and differ primarily in the gender used to realize an otherwise ambiguous case, an evaluation metric should not systematically reward one variant over the other\. Previous work has nevertheless shown that MT evaluation metrics can exhibit such preferences\([Filandrianos et al\., 2025](https://arxiv.org/html/2609.21490#bib.bib8);[Zaranis et al\., 2025](https://arxiv.org/html/2609.21490#bib.bib4)\)\. This is particularly relevant because automatic metrics are widely used to compare MT systems, rank candidate translations, and guide model development, meaning that systematic evaluator preferences can affect how MT quality is ultimately progressing\. GAMBIT\+\([Filandrianos et al\., 2025](https://arxiv.org/html/2609.21490#bib.bib8)\)was introduced to study this form of evaluator bias through controlled masculine/feminine translation pairs\. Building on GAMBIT\([Menis Mastromichalakis et al\., 2025b](https://arxiv.org/html/2609.21490#bib.bib9)\), it organizes gender\-ambiguous occupational references according to the 436 four\-digit groups of the ISCO\-08 classification222https://ilostat\.ilo\.org/methods/concepts\-and\-definitions/classification\-occupation/and provides paired translations in which the relevant occupation is realized once in masculine and once in feminine form\. The original study evaluated 33 source–target language pairs, combining three source languages with eleven target languages, and used the resulting challenge set to analyze quality estimation metrics at WMT 2025\. In this work, we revisit this evaluation for the WMT 2026 Automated Translation Quality Evaluation Systems Shared Task[Lavie et al\. \(2026\)](https://arxiv.org/html/2609.21490#bib.bib25)using a smaller, occupation\-balanced subset of GAMBIT\+\. The full benchmark contains nearly 290,000 paired instances, making large\-scale evaluation increasingly costly as the number and complexity of evaluation systems grows\. We therefore retain three examples for each of the 436 ISCO\-08 occupational groups, preserving complete occupational coverage while substantially reducing the computational cost of evaluation\. We use English exclusively as the source language and consider seven target languages: Arabic, Czech, German, Greek, Icelandic, Russian, and Ukrainian\. Six of these target languages are inherited from the original resource, while German is newly added in this work, yielding 1,308 masculine/feminine translation pairs per target language\. We use this benchmark to evaluate both shared\-task submissions and baselines for score prediction and error annotation\. However, identifying whether an evaluator exhibits gender bias is not fully captured by a single average score difference\. A metric may react strongly to the gendered realization of many examples while showing little net preference because differences in opposite directions cancel out\. Conversely, relatively small score differences may consistently favor the same gender across many examples\. We therefore examine several complementary aspects of evaluator behavior: signed score differences to capture the direction of preference, absolute paired differences to capture sensitivity regardless of direction, and preference frequencies to measure how consistently one variant is favored\. For error\-annotation systems, we additionally compare how often masculine and feminine variants are predicted to be error\-free\. Our contributions are threefold\. First, we construct an occupation\-balanced English\-source subset of GAMBIT\+ for the seven language pairs considered in this work, including the German extension\. Second, we evaluate gender\-related differences across score\-prediction submissions and baselines, jointly analyzing their direction, magnitude, and frequency\. Third, we extend the analysis to error\-annotation systems, examining whether masculine and feminine translations differ in their likelihood of being labeled error\-free\. Overall, our results show that gender bias remains present in MT evaluation, but that its patterns are not fully captured by a single measure and require a more nuanced analysis of evaluator behavior\. ## 2Related Work Gender bias in machine translation has been extensively studied across languages, model architectures, and evaluation settings\([Savoldi et al\., 2021](https://arxiv.org/html/2609.21490#bib.bib1);[Vanmassenhove, 2024](https://arxiv.org/html/2609.21490#bib.bib15);[Savoldi et al\., 2024](https://arxiv.org/html/2609.21490#bib.bib16);[Rescigno et al\., 2020](https://arxiv.org/html/2609.21490#bib.bib17);[Paolucci et al\., 2023](https://arxiv.org/html/2609.21490#bib.bib18);[Ghosh and Caliskan, 2023](https://arxiv.org/html/2609.21490#bib.bib19);[Kostikova et al\., 2023](https://arxiv.org/html/2609.21490#bib.bib20);[Piazzolla et al\., 2023](https://arxiv.org/html/2609.21490#bib.bib21)\)\. A particularly relevant strand of this work concerns occupational gender bias, where translation systems may associate professions with stereotypical gender distributions\([Menis Mastromichalakis et al\., 2025a](https://arxiv.org/html/2609.21490#bib.bib5);[Menis Mastromichalakis et al\., 2025b](https://arxiv.org/html/2609.21490#bib.bib9);[Menis Mastromichalakis et al\., 2024](https://arxiv.org/html/2609.21490#bib.bib7);[Tal et al\., 2022](https://arxiv.org/html/2609.21490#bib.bib22);[Gorti et al\., 2024](https://arxiv.org/html/2609.21490#bib.bib23)\)\. Recent work continues to find such effects in modern MT architectures, suggesting that default gender preferences remain an open problem even as model capabilities improve\([Manna et al\., 2026](https://arxiv.org/html/2609.21490#bib.bib12)\)\. Evaluating these biases is particularly challenging when the source does not provide enough information to determine a person’s gender\. Benchmarks such as WinoMT and MT\-GenEval focus primarily on cases where contextual evidence identifies the appropriate gender\([Stanovsky et al\., 2019](https://arxiv.org/html/2609.21490#bib.bib3);[Currey et al\., 2022](https://arxiv.org/html/2609.21490#bib.bib2)\), whereas GAMBIT targets occupational references for which masculine and feminine realizations may both be valid\([Menis Mastromichalakis et al\., 2025b](https://arxiv.org/html/2609.21490#bib.bib9)\)\. This distinction has received increasing attention: recent work has explored uncertainty as a way of evaluating model behavior under gender ambiguity\([Staliūnaitė et al\., 2026](https://arxiv.org/html/2609.21490#bib.bib11)\), while other studies emphasize preserving ambiguity through gender\-neutral or gender\-inclusive translation strategies\([Dawkins et al\., 2025](https://arxiv.org/html/2609.21490#bib.bib14);[Savoldi et al\., 2025](https://arxiv.org/html/2609.21490#bib.bib24)\)\. Our setting is complementary to these approaches, focusing specifically on systematic preferences between masculine and feminine realizations rather than evaluating neutral alternatives\. While most research has focused on biases in MT outputs, less attention has been paid to biases introduced during automatic evaluation\.[Zaranis et al\. \(2025\)](https://arxiv.org/html/2609.21490#bib.bib4)showed that quality\-estimation metrics can exhibit systematic gender disparities, and GAMBIT\+ extended this analysis to a multilingual, occupation\-indexed challenge set\([Filandrianos et al\., 2025](https://arxiv.org/html/2609.21490#bib.bib8)\)\. More recently, FairQE has explored explicitly mitigating gender bias in quality estimation\([Jang et al\., 2026](https://arxiv.org/html/2609.21490#bib.bib13)\)\. Building on this line of work, we use GAMBIT\+ to evaluate WMT 2026 metrics and extend the analysis beyond average score differences, considering complementary dimensions of evaluator preference and sensitivity\. ## 3The Challenge Set ### 3\.1GAMBIT\+ extension GAMBIT\+\([Filandrianos et al\., 2025](https://arxiv.org/html/2609.21490#bib.bib8)\)provides English texts containing gender\-ambiguous occupational references, together with paired target\-language translations in which the relevant occupation is realized once in masculine and once in feminine form\. The dataset covers the 436 four\-digit occupational groups of ISCO\-08 and was originally released for multiple source and target languages\. For this year’s evaluation, we focus exclusively on English as the source language and consider seven target languages: Arabic, Czech, German, Greek, Icelandic, Russian, and Ukrainian\. Six of these language pairs are inherited from the original GAMBIT\+ release, while German is newly added in this work\. We follow the same general generation procedure described in GAMBIT\+ for constructing paired masculine and feminine translations, using LongCat\-2\.0333https://longcat\.ai/for the German data\. As in the original benchmark, the two target variants are intended to preserve the same source information while differing in the gender realization of the ambiguous occupational reference and any grammatical elements that depend on it\. Table[1](https://arxiv.org/html/2609.21490#S3.T1)shows an example\. Table 1:An English source text with the paired German translations introduced in the GAMBIT\+ extension\. The underlined text marks the gender\-dependent difference between the masculine and feminine variants; the remaining text is identical\. ### 3\.2Occupation\-balanced subset To make the benchmark more practical for evaluating a larger and increasingly computationally demanding set of metrics, we construct a smaller occupation\-balanced subset of GAMBIT\+\. We sample three English source examples for each of the 436 ISCO\-08 occupational groups, yielding 1,308 source instances per language pair\. This preserves complete occupational coverage while substantially reducing the size of the benchmark\. Each source instance is still associated with two target translations: one masculine and one feminine\. The resulting benchmark therefore contains seven English–target language pairs, with 1,308 paired translations per target language\. ### 3\.3Evaluated systems We evaluate the WMT 2026 Automated Translation Quality Evaluation Shared Task systems available for these seven language pairs, including both participant submissions and shared\-task baselines\. The analysis covers two types of evaluator output: numerical quality scores and error\-span annotations\. For score prediction, each evaluator assigns a numerical quality score to the masculine and feminine translation of every source example\. For error annotation, systems instead identify translation errors through spans and omission flags\. These two settings provide complementary views of evaluator behavior: the former allows us to measure differences in assigned quality scores, while the latter allows us to examine whether masculine and feminine variants are treated differently in terms of predicted errors\. All analyses are performed on the complete set of 1,308 masculine/feminine pairs available for each evaluated language–system combination\. We additionally verify the consistency of segment identifiers and pair alignment before computing the results\. Appendix[A](https://arxiv.org/html/2609.21490#A1)reports the detailed system coverage and score ranges\. Since the benchmark is stratified by occupation, the analysis below treats occupation as the primary grouping unit when assessing the stability and consistency of gender\-related score differences\. ## 4Analysis Protocol ### 4\.1Direction, magnitude and frequency LetqkℓiMq\_\{k\\ell i\}^\{M\}andqkℓiFq\_\{k\\ell i\}^\{F\}be the scores for metrickk, targetℓ\\elland sampleii, for the masculine and feminine translations respectively\. We define: dkℓi\\displaystyle d\_\{k\\ell i\}=qkℓiM−qkℓiF,\\displaystyle=q\_\{k\\ell i\}^\{M\}\-q\_\{k\\ell i\}^\{F\},\(1\)Skℓ\\displaystyle S\_\{k\\ell\}=1nℓ∑idkℓi,\\displaystyle=\\frac\{1\}\{n\_\{\\ell\}\}\\sum\_\{i\}d\_\{k\\ell i\},\(2\)Akℓ\\displaystyle A\_\{k\\ell\}=1nℓ∑i\|dkℓi\|\.\\displaystyle=\\frac\{1\}\{n\_\{\\ell\}\}\\sum\_\{i\}\|d\_\{k\\ell i\}\|\.\(3\)PositiveSSvalues indicate higher masculine scores, whileAAmeasures sensitivity regardless of direction\. In particular,A≠\|S\|A\\neq\|S\|in general: opposing differences cancel inSSbut not inAA\. For comparability with the range\-based analysis of GAMBIT\+, we also report S%kℓ=100Skℓ/Rk,A%kℓ=100Akℓ/Rk,S^\{\\%\}\_\{k\\ell\}=100S\_\{k\\ell\}/R\_\{k\},\\quad A^\{\\%\}\_\{k\\ell\}=100A\_\{k\\ell\}/R\_\{k\},\(4\)whereRkR\_\{k\}is the maximum minus minimum*individual*score over both variants and all available target languages for that submission\. These percentages express the differences as a fraction of the observed range\. They remain sensitive to outliers and coverage\. We additionally calculate masculine\-win, feminine\-win and tie rates, with all pairs considered in the denominator\. The preference balance then corresponds to100\[Pr\(d\>0\)−Pr\(d<0\)\]100\[\\Pr\(d\>0\)\-\\Pr\(d<0\)\]percentage points\. ForA\>0A\>0, the cancellation ratio1−\|S\|/A1\-\|S\|/Aquantifies the discrepancy between net and absolute effects\. ### 4\.2Aggregation, confidence intervals, and significance testing Each target language contains three examples for each of the 436 ISCO\-08 occupational groups\. Since every occupation is represented equally, averaging over all 1,308 translation pairs is equivalent to first averaging within each occupation and then averaging over occupations\. For multilingual summaries, we compute macro\-averages across languages, giving each language equal weight\. We report averages over all available languages for each system, and use a fixed six\-language panel \(AR, CS, DE, IS, RU, and UK\) when comparing the 22 systems with common coverage\. To assess the stability of the aggregate effects across occupations, we use an occupation\-block bootstrap\. We resample the 436 occupations with replacement, retaining the three paired examples belonging to each selected occupation, and recompute the signed and absolute differences over 5,000 bootstrap replicates\. For multilingual summaries, all observations associated with the same occupation across the included languages remain in the same block\. We report percentile 95% intervals while keeping the observed normalization ranges fixed\. Because the benchmark already covers all 436 ISCO\-08 groups, these intervals should be interpreted as sensitivity to changes in the relative weighting of occupations, rather than as sampling uncertainty over unseen occupations or languages\. For each system–language run, we test whether the mean signed difference across the 436 occupations differs from zero using a two\-sided one\-samplett\-test\. We correct the resulting 148pp\-values for multiple comparisons using the Benjamini–Hochberg procedure\([Benjamini and Hochberg, 1995](https://arxiv.org/html/2609.21490#bib.bib10)\)\. A significant result indicates a consistent directional preference across occupations; the sign of the mean determines whether masculine or feminine variants receive higher scores on average\. ## 5Score\-Prediction Results Table 2:Signed / mean absolute paired differences \(S%/A%S^\{\\%\}/A^\{\\%\}\), in percent of each evaluator's observed score range\. Positive signed values indicate higher masculine scores\. Each language has 1,308 pairs\. Macro\-6 equally averages AR, CS, DE, IS, RU and UK; – denotes missing coverage, never zero\. RB/RF retain returned reference\-based/reference\-free labels, and CA abbreviates Confidence Aware; inference configurations were not independently verified\. Native scales appear in Appendix[A](https://arxiv.org/html/2609.21490#A1)\.†denotes a WMT 2026 organizer baseline; unmarked systems are participant submissions\.### 5\.1Systematic preference and cancellation Table[2](https://arxiv.org/html/2609.21490#S5.T2)reports both normalized summaries for every Task 2 \(Quality Score Prediction\) system\. After Benjamini–Hochberg correction atq=0\.05q=0\.05, 131 of the 148 system–language runs show a statistically detectable directional difference\. Of these, 126 favor masculine variants on average and five favor feminine variants\. The five significant negative runs areVertical\_870257on Arabic,Vertical\_870554on Arabic and German, andfluency2\-gemini35\-esaon German and Ukrainian\. Significance should also be read together with magnitude: a consistent but small preference and a large, heterogeneous effect have different implications\. Figure[1](https://arxiv.org/html/2609.21490#S5.F1)compares the signed and mean absolute differences on the common six\-language set\. The contrast between the two quantities is particularly clear forfluency2\-gemini35\-esa: although its mean absolute difference is large \(A=5\.7234A=5\.7234,A%=12\.69A^\{\\%\}=12\.69\), its signed mean is close to zero \(S=−0\.0996S=\-0\.0996,S%=−0\.22S^\{\\%\}=\-0\.22\)\. This indicates that substantial gender\-related score changes occur in both directions and largely cancel in the signed average\. Qwen 3\.6, by comparison, hasS%=4\.06S^\{\\%\}=4\.06andA%=7\.34A^\{\\%\}=7\.34, whilexCOMET XXLreference\-free hasS%=5\.75S^\{\\%\}=5\.75andA%=7\.01A^\{\\%\}=7\.01, indicating a more consistently positive preference relative to the overall magnitude of the differences\. COMETKiwi22gemba\-polyAEGISVertical\_870257Vertical\_870554bytepop\_pro9bytepop\_pro10xCOMET XL RBxCOMET XL RFGemini\-3\.6\-FlashFACET\_869543FACET\_869546Cohere CAT\+ ensembleMQM\-LLMMQM\-LLM CAGemma 4 RBGemma 4xCOMET XXL RBxCOMET XXL RFQwen 3\.6 RBQwen 3\.6fluency2\-gemini35\-esa024681012Difference \(% of observed metric range\)∙\\bulletSigned mean \(95% interval\)■\\blacksquareMean absolute differenceFigure 1:Signed and mean absolute gender\-related score differences on the common six\-language set across 22 qualifying evaluators\. The signed difference captures the direction of preference, while the mean absolute difference captures its magnitude regardless of direction\. Whiskers show percentile 95% intervals from 5,000 resamples of the 436 ISCO blocks, keeping paired observations and included languages together\. Values closer to zero indicate smaller gender\-related differences\.COMETKiwi22\([Rei et al\., 2022](https://arxiv.org/html/2609.21490#bib.bib6)\)illustrates a different pattern\. Its six\-language mean absolute difference is small \(A=0\.00674A=0\.00674, corresponding to 1\.19% of its observed range\), yet it assigns higher scores to masculine variants in 78\.1% of pairs, compared with 20\.8% for feminine variants and 1\.1% ties\. Thus, a small average magnitude does not necessarily imply balanced preference frequencies\. Conversely,gemba\-polyproduces ties for 61\.8% of pairs\. ### 5\.2Language and occupation differences To compare target languages using the same set of evaluators, we average the normalized differences over the 13 systems available for all seven languages\. The mean absolute difference,A%A^\{\\%\}, ranges from 3\.14 in German to 4\.64 in Icelandic\. Russian shows the largest average signed difference \(S%=2\.72S^\{\\%\}=2\.72\), while German is much closer to zero \(S%=0\.26S^\{\\%\}=0\.26\)\. The direction of the preference can also vary across evaluators\. For German, for example, Qwen 3\.6 has a positive signed difference \(S%=2\.09S^\{\\%\}=2\.09\), whereasfluency2\-gemini35\-esahas a negative one \(S%=−2\.59S^\{\\%\}=\-2\.59\)\. Thus, even within the same target language, evaluators do not always favor the same gender\. This suggests that the observed differences are more strongly associated with the evaluator than with the target language\. We also examine differences across occupations\. Table[3](https://arxiv.org/html/2609.21490#S5.T3)reports the five occupations with the largest feminine\- and masculine\-preferring signed differences, based on the mean normalized signed differenceS%S^\{\\%\}across the seven target languages\. To make these comparisons consistent, all language\-level values are averaged over the same 13 systems with coverage for all seven languages\. Several of the strongest differences align with familiar occupational gender stereotypes\. Midwifery professionals and child care workers receive higher feminine scores on average, whereas plumbers, hunters, and trappers receive higher masculine scores\. This pattern is broadly consistent with observations in our previous GAMBIT\+ evaluation\. However, the association is not uniform across occupations or languages: for example, Fashion and Other Models shows a feminine preference overall despite positive differences in Arabic and Czech\. We therefore treat these results as descriptive evidence of occupation\-level variation rather than as a direct measure of gender stereotypes\. Table 3:Occupation\-level extremes across the seven target languages, using the fixed panel of 13 Task 2 systems available for all languages\. Each language column reports the mean normalized signed differenceS%S^\{\\%\}across these systems, and the final column gives the macro\-average across languages\. Negative values indicate higher scores for feminine variants and positive values higher scores for masculine variants\.Table 4:Task 1 differences in predicted error\-free rate, in percentage points \(masculine minus feminine\)\. Positive values indicate that masculine variants are more often predicted as error\-free, while negative values indicate a higher error\-free rate for feminine variants\.†denotes a baseline; unmarked systems are participant submissions\. ## 6Error\-Annotation Results Task 1 provides a complementary view of gender\-related evaluator behavior by allowing us to examine whether masculine and feminine variants differ in how often they are predicted to be error\-free\. We define a translation as error\-free when the system predicts neither an error span nor an omission flag\. This closely parallels Task 3 of the shared task, which focuses on the detection of error\-free segments\. For each system and language, we compute the difference between the masculine and feminine error\-free rates, in percentage points; positive values indicate that masculine variants are more often predicted as error\-free\. Table[4](https://arxiv.org/html/2609.21490#S5.T4)reports the results across all available systems and languages\. The overall pattern is strongly skewed toward masculine variants: most system–language combinations have positive values, indicating that masculine translations are more often judged error\-free than their feminine counterparts\. In several cases the differences are substantial, reaching 27\.8 percentage points forGemma 4in Russian and around 30 points for theLexicalavariants in Arabic\. At the same time, the magnitude of the effect varies considerably across both systems and languages, and the dominant direction is not universal\. For example,fluency2\-gemini35\-spansfavors feminine variants in Czech, German, and Ukrainian, while showing positive differences in the remaining languages\. More generally, systems often show similar tendencies across several languages, but different evaluators can behave very differently on the same language\. This mirrors the Task 2 results and suggests that the observed gender\-related differences are strongly shaped by the evaluator itself rather than a uniform language\-level effect\. ## 7Conclusion In this work, we revisited gender bias in MT evaluation using an occupation\-balanced subset of GAMBIT\+, covering seven English\-source language pairs and extending the benchmark with German\. Across WMT 2026 score\-prediction systems, masculine translations are generally favored: 126 of the 131 statistically detectable system–language differences are positive\. However, the strength and form of this preference vary substantially across evaluators\. Signed differences, absolute differences, and preference frequencies often capture different aspects of the same behavior, while the error\-annotation analysis shows a similar overall tendency for masculine variants to be predicted as error\-free more often, with important exceptions\. Overall, our results show that gender\-related bias remains present in modern MT evaluation systems, but cannot be adequately characterized by a single aggregate measure\. The observed effects vary across evaluators, languages, and occupations, with some of the strongest occupation\-level differences aligning with familiar gender stereotypes\. We therefore recommend assessing evaluator bias through complementary measures of direction, magnitude, and consistency\. The occupation\-balanced GAMBIT\+ subset introduced here provides a more computationally practical resource for such analyses while preserving complete ISCO\-08 occupational coverage\. ## Limitations German was generated separately from the inherited GAMBIT\+ targets, and the language\-specific subsets do not necessarily contain the same English contexts\. In addition, each occupation–language combination contains only three examples, so occupation\-level differences should be interpreted descriptively rather than as stable estimates\. The analysis also has certain methodological limitations\. Range\-based normalization is sensitive to extreme scores, and small gender\-related differences should not be interpreted as evidence of overall evaluator quality\. We also observe only one returned run per system and therefore cannot separate systematic effects from stochastic variation\. Finally, our evaluation considers only masculine and feminine variants; gender\-neutral or inclusive alternatives are outside the scope of the submitted benchmark\. Accordingly, we do not draw causal conclusions about language effects, occupation\-specific stereotypes, or evaluator fairness beyond the pairs and systems studied here\. ## Acknowledgments O\. Menis Mastromichalakis and G\. Attanasio were supported by the project DECOLLAGE \(ERC\-2022\-CoG 101088763\), and FCT/MECI through national funds and co\-funded EU funds under UID/50008: Instituto de Telecomunicações\. G\. Filandrianos and C\. Zerva were partly supported within the framework of the Pharos AI Factory project, funded by the European High\-Performance Computing Joint Undertaking \(EuroHPC JU\) under Grant Agreement No\. 101234269 as part of Horizon Europe and by the Greek Public Investments Program programme\. G\. Attanasio was additionally supported by the AMALIA project under Measure RE\-C05\-i08 of the Portuguese national Programa de Recuperação e Resiliência, and by the Portuguese Recovery and Resilience Plan through project C645008882\-00000055 \(Center for Responsible AI\)\. C\. Zerva was additionally supported by a Google Research Scholar award\. We would also like to thank Manuel Lardelli for the help in reviewing selected German masculine and feminine translations for quality and validity\. ## References - Benjamini and Hochberg \(1995\)Y\. Benjamini and Y\. HochbergControlling the false discovery rate: a practical and powerful approach to multiple testing\.Journal of the Royal statistical society: series B \(Methodological\)57\(1\),pp\. 289–300\.Cited by:[§4\.2](https://arxiv.org/html/2609.21490#S4.SS2.p3.1)\. - Curreyet al\.\(2022\)A\. Currey, M\. Nadejde, R\. R\. Pappagari, M\. Mayer, S\. Lauly, X\. Niu, B\. Hsu, and G\. DinuMT\-GenEval: a counterfactual and contextual dataset for evaluating gender accuracy in machine translation\.InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing,Y\. Goldberg, Z\. Kozareva, and Y\. Zhang \(Eds\.\),Abu Dhabi, United Arab Emirates,pp\. 4287–4299\.External Links:[Link](https://aclanthology.org/2022.emnlp-main.288/),[Document](https://dx.doi.org/10.18653/v1/2022.emnlp-main.288)Cited by:[§2](https://arxiv.org/html/2609.21490#S2.p2.1)\. - Dawkinset al\.\(2025\)H\. Dawkins, I\. Nejadgholi, and C\. LoGender\-neutral machine translation strategies in practice\.InProceedings of the 3rd workshop on gender\-inclusive translation technologies \(GITT 2025\),pp\. 74–88\.Cited by:[§2](https://arxiv.org/html/2609.21490#S2.p2.1)\. - Filandrianoset al\.\(2025\)G\. Filandrianos, O\. Menis Mastromichalakis, W\. Mohammed, G\. Attanasio, and C\. ZervaGambit\+: a challenge set for evaluating gender bias in machine translation quality estimation metrics\.InProceedings of the Tenth Conference on Machine Translation,pp\. 314–326\.Cited by:[§1](https://arxiv.org/html/2609.21490#S1.p2.1),[§1](https://arxiv.org/html/2609.21490#S1.p3.1),[§2](https://arxiv.org/html/2609.21490#S2.p3.1),[§3\.1](https://arxiv.org/html/2609.21490#S3.SS1.p1.1)\. - Ghosh and Caliskan \(2023\)S\. Ghosh and A\. CaliskanChatgpt perpetuates gender bias in machine translation and ignores non\-gendered pronouns: findings across bengali and five other low\-resource languages\.InProceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society,pp\. 901–912\.Cited by:[§2](https://arxiv.org/html/2609.21490#S2.p1.1)\. - Gortiet al\.\(2024\)A\. Gorti, A\. Chadha, and M\. GaurUnboxing occupational bias: debiasing llms with us labor data\.InProceedings of the AAAI Symposium Series,Vol\.4,pp\. 48–55\.Cited by:[§2](https://arxiv.org/html/2609.21490#S2.p1.1)\. - Janget al\.\(2026\)J\. Jang, J\. Choi, D\. Lee, S\. Yu, and Y\. KimFairQE: multi\-agent framework for mitigating gender bias in translation quality estimation\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 37891–37911\.Cited by:[§2](https://arxiv.org/html/2609.21490#S2.p3.1)\. - Kostikovaet al\.\(2023\)A\. Kostikova, J\. Daems, and T\. LazarovHow adaptive is adaptive machine translation, really? a gender\-neutral language use case\.InProceedings of the First Workshop on Gender\-Inclusive Translation Technologies,pp\. 95–97\.Cited by:[§2](https://arxiv.org/html/2609.21490#S2.p1.1)\. - Lavieet al\.\(2026\)A\. Lavie, G\. Hanneman, S\. Perrella, S\. Ding, E\. Avramidis, L\. Proietti, C\. Lo, A\. Shurtz, C\. Zerva, A\. Sindhujan, V\. Zouhar, D\. Kanojia, F\. Blain, B\. Thompson, G\. Filandrianos, O\. Menis Mastromichalakis, T\. Kocmi, and P\. GuptaFindings of the wmt26 shared task on automated translation quality evaluation: compact open models are competitive quality evaluators\.InProceedings of the Eleventh Conference on Machine Translation,P\. Koehn, B\. Haddow, T\. Kocmi, and C\. Monz \(Eds\.\),Budapest, Hungary\.Cited by:[§1](https://arxiv.org/html/2609.21490#S1.p4.1)\. - Mannaet al\.\(2026\)C\. Manna, H\. Mohebbi, A\. Alishahi, F\. Blain, and E\. VanmassenhoveGender disambiguation in machine translation: diagnostic evaluation in decoder\-only architectures\.InProceedings of the Fifteenth Language Resources and Evaluation Conference,S\. Piperidis, N\. Bel, H\. van den Heuvel, N\. Ide, S\. Krek, and A\. Toral \(Eds\.\),Palma de Mallorca, Spain,pp\. 8535–8550\.External Links:[Link](https://aclanthology.org/2026.lrec-1.673/),[Document](https://dx.doi.org/10.63317/4wphxianzxf6)Cited by:[§2](https://arxiv.org/html/2609.21490#S2.p1.1)\. - Menis Mastromichalakiset al\.\(2025a\)O\. Menis Mastromichalakis, G\. Filandrianos, M\. Symeonaki, G\. Stamatopoulou, D\. Parsanoglou, and G\. StamouGender bias in machine learning: insights from official labour statistics and textual analysis\.Quality & Quantity,pp\. 1–35\.Cited by:[§2](https://arxiv.org/html/2609.21490#S2.p1.1)\. - Menis Mastromichalakiset al\.\(2025b\)O\. Menis Mastromichalakis, G\. Filandrianos, M\. Symeonaki, and G\. StamouAssumed identities: quantifying gender bias in machine translation of gender\-ambiguous occupational terms\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 32221–32237\.Cited by:[§1](https://arxiv.org/html/2609.21490#S1.p1.1),[§1](https://arxiv.org/html/2609.21490#S1.p3.1),[§2](https://arxiv.org/html/2609.21490#S2.p1.1),[§2](https://arxiv.org/html/2609.21490#S2.p2.1)\. - Menis Mastromichalakiset al\.\(2024\)O\. Menis Mastromichalakis, G\. Filandrianos, E\. Tsouparopoulou, D\. Parsanoglou, M\. Symeonaki, and G\. StamouGost\-mt: a knowledge graph for occupation\-related gender biases in machine translation\.arXiv preprint arXiv:2409\.10989\.Cited by:[§2](https://arxiv.org/html/2609.21490#S2.p1.1)\. - Paolucciet al\.\(2023\)A\. B\. Paolucci, M\. Lardelli, and D\. GromannGender\-fair language in translation: a case study\.InProceedings of the First Workshop on Gender\-Inclusive Translation Technologies,pp\. 13–23\.Cited by:[§2](https://arxiv.org/html/2609.21490#S2.p1.1)\. - Piazzollaet al\.\(2023\)S\. A\. Piazzolla, B\. Savoldi, and L\. BentivogliGood, but not always fair: an evaluation of gender bias for three commercial machine translation systems\.arXiv preprint arXiv:2306\.05882\.Cited by:[§2](https://arxiv.org/html/2609.21490#S2.p1.1)\. - Reiet al\.\(2022\)R\. Rei, M\. Treviso, N\. M\. Guerreiro, C\. Zerva, A\. C\. Farinha, C\. Maroti, J\. G\. C\. de Souza, T\. Glushkova, D\. Alves, L\. Coheur, A\. Lavie, and A\. F\. T\. MartinsCometKiwi: IST\-unbabel 2022 submission for the quality estimation shared task\.InProceedings of the Seventh Conference on Machine Translation \(WMT\),P\. Koehn, L\. Barrault, O\. Bojar, F\. Bougares, R\. Chatterjee, M\. R\. Costa\-jussà, C\. Federmann, M\. Fishel, A\. Fraser, M\. Freitag, Y\. Graham, R\. Grundkiewicz, P\. Guzman, B\. Haddow, M\. Huck, A\. Jimeno Yepes, T\. Kocmi, A\. Martins, M\. Morishita, C\. Monz, M\. Nagata, T\. Nakazawa, M\. Negri, A\. Névéol, M\. Neves, M\. Popel, M\. Turchi, and M\. Zampieri \(Eds\.\),Abu Dhabi, United Arab Emirates \(Hybrid\),pp\. 634–645\.External Links:[Link](https://aclanthology.org/2022.wmt-1.60/)Cited by:[§5\.1](https://arxiv.org/html/2609.21490#S5.SS1.p3.1)\. - Rescignoet al\.\(2020\)A\. A\. Rescigno, E\. Vanmassenhove, J\. Monti, and A\. WayA case study of natural gender phenomena in translation\. a comparison of google translate, bing microsoft translator and deepl for english to italian, french and spanish\.InProceedings of the Seventh Italian Conference on Computational Linguistics \(CLiC\-it 2020\),pp\. 257–262\.Cited by:[§2](https://arxiv.org/html/2609.21490#S2.p1.1)\. - Savoldiet al\.\(2025\)B\. Savoldi, G\. Attanasio, E\. Cupin, E\. Gkovedarou, J\. Hackenbuchner, A\. Lauscher, M\. Negri, A\. Piergentili, M\. Thind, and L\. BentivogliMind the inclusivity gap: multilingual gender\-neutral translation evaluation with mGeNTE\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 13698–13720\.External Links:[Link](https://aclanthology.org/2025.emnlp-main.692/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.692),ISBN 979\-8\-89176\-332\-6Cited by:[§2](https://arxiv.org/html/2609.21490#S2.p2.1)\. - Savoldiet al\.\(2021\)B\. Savoldi, M\. Gaido, L\. Bentivogli, M\. Negri, and M\. TurchiGender bias in machine translation\.Transactions of the Association for Computational Linguistics9,pp\. 845–874\.External Links:[Link](https://aclanthology.org/2021.tacl-1.51/),[Document](https://dx.doi.org/10.1162/tacl%5Fa%5F00401)Cited by:[§2](https://arxiv.org/html/2609.21490#S2.p1.1)\. - Savoldiet al\.\(2024\)B\. Savoldi, S\. Papi, M\. Negri, A\. Guerberof\-Arenas, and L\. BentivogliWhat the harm? quantifying the tangible impact of gender bias in machine translation with a human\-centered study\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,pp\. 18048–18076\.Cited by:[§2](https://arxiv.org/html/2609.21490#S2.p1.1)\. - Staliūnaitėet al\.\(2026\)I\. Staliūnaitė, J\. Cheng, and A\. VlachosUncertainty quantification for evaluating gender bias in machine translation\.InFindings of the Association for Computational Linguistics: EACL 2026,pp\. 2204–2225\.Cited by:[§2](https://arxiv.org/html/2609.21490#S2.p2.1)\. - Stanovskyet al\.\(2019\)G\. Stanovsky, N\. A\. Smith, and L\. ZettlemoyerEvaluating gender bias in machine translation\.InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics,A\. Korhonen, D\. Traum, and L\. Màrquez \(Eds\.\),Florence, Italy,pp\. 1679–1684\.External Links:[Link](https://aclanthology.org/P19-1164/),[Document](https://dx.doi.org/10.18653/v1/P19-1164)Cited by:[§2](https://arxiv.org/html/2609.21490#S2.p2.1)\. - Talet al\.\(2022\)Y\. Tal, I\. Magar, and R\. SchwartzFewer errors, but more stereotypes? the effect of model size on gender bias\.InProceedings of the 4th Workshop on Gender Bias in Natural Language Processing \(GeBNLP\),pp\. 112–120\.Cited by:[§2](https://arxiv.org/html/2609.21490#S2.p1.1)\. - Vanmassenhove \(2024\)E\. VanmassenhoveGender bias in machine translation and the era of large language models\.Gendered Technology in Translation and Interpreting,pp\. 225–252\.Cited by:[§2](https://arxiv.org/html/2609.21490#S2.p1.1)\. - Zaraniset al\.\(2025\)E\. Zaranis, G\. Attanasio, S\. Agrawal, and A\. MartinsWatching the watchers: exposing gender disparities in machine translation quality estimation\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 25261–25284\.External Links:[Link](https://aclanthology.org/2025.acl-long.1228/),[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.1228),ISBN 979\-8\-89176\-251\-0Cited by:[§1](https://arxiv.org/html/2609.21490#S1.p2.1),[§2](https://arxiv.org/html/2609.21490#S2.p3.1)\. ## Appendix ACoverage, Native Scales and Macro Results Table[5](https://arxiv.org/html/2609.21490#A1.T5)gives the interval of individual scores observed across each submission’s available languages and the corresponding rangeRkR\_\{k\}\. Available\-language means use all returned targets for that submission\. Six\-language means use only AR, CS, DE, IS, RU and UK\. Tables and figures abbreviate reference\-based/free as RB/RF, Confidence Aware as CA, and Lexicala\-QE\-Ensemble as Lexicala\. The unsuffixed Gemma 4 and Qwen 3\.6 names refer to their returned “Reasoning Baseline” configurations\. Table 5:Observed score scales and macro means in native units\.LLis the number of returned target languages\. Absolute means are means of pairwise absolute differences, not absolute values of signed means\.†denotes a baseline; unmarked systems are participant submissions\.
相似文章
将LLM性别偏见锚定于人类基线:一项跨语言审计
本文对六种大型语言模型在英语、韩语、中文和日语中的性别刻板印象进行审计,并以人类基线作为锚定。研究发现,LLM的刻板印象程度往往超过人类跨国差异,且可能跨语言叠加,为此引入了一个四模式框架来表征此类行为。
回答格式变化对大型语言模型中性别偏见的影响
本文使用BBQ和OpinionQA基准测试,评估回答格式的变化如何影响大型语言模型中性别偏见的测量,发现结果有显著变化,这突显了模型评估中多格式设计的必要性。
缓解英语到罗马尼亚语机器翻译中的性别偏见
本文提出了一种混合流水线,将基于微调LLaMA的性别分类与标签感知的神经机器翻译相结合,以缓解英语到罗马尼亚语机器翻译中的性别偏见,引入了新数据集,并在基准测试上将性别准确率提高了40多个百分点。
两个Emojis之差:多语言情感生成基准测试到底在测什么
本文审计了多语言情感生成基准测试,揭示系统排名由测量伪影而非真实性能差异驱动,标注者变异性在其中扮演关键角色。
语言模型认为谁是称职的?职业偏见的机制分析
论文提出了一个因果框架,用于分析语言模型的职业偏见,发现即使行为指标未显示差异,表征偏见也可能持续存在,并且在干预时这些偏见可能影响下游行为。