G-IdiomAlign: A Gloss-Pivoted Benchmark for Cross-Lingual Idiom Alignment

arXiv cs.CL Papers

Summary

G-IdiomAlign is a gloss-pivoted benchmark for evaluating cross-lingual idiom alignment in LLMs, featuring controlled multiple-choice and gloss-contrastive protocols to diagnose literal translation bias and the effect of semantic pivots.

arXiv:2606.18989v1 Announce Type: new Abstract: Idioms are difficult to transfer across languages due to their non-compositionality and weak surface-form grounding, making literal mappings unreliable. We present G-IdiomAlign, a gloss-pivoted benchmark where each idiom is anchored by an English gloss from Wiktionary. We further construct a high-confidence reference alignment set for reproducible evaluation. G-IdiomAlign supports two protocols: (1) a controlled Multiple-Choice Idiom Equivalence with typed distractors for error attribution; and (2) a Gloss-Contrastive Generation contrasting No-gloss and With-gloss inputs to isolate the effect of an explicit semantic pivot. Across diverse LLMs, a bias to literal translation is a dominant failure mode, especially when the target is a low-resource language. Glosses consistently improve Gloss-Contrastive Generation under an embedding-based semantic proxy, but performance remains modest, indicating substantial headroom in the open output space. Subsequent analysis on Qwen3-8B further suggests that cross-condition differences are concentrated more in attention heads than in layers, while better With-gloss generations coincide with stronger gloss anchoring.
Original Article
View Cached Full Text

Cached at: 06/18/26, 05:47 AM

# A Gloss-Pivoted Benchmark for Cross-Lingual Idiom Alignment
Source: [https://arxiv.org/html/2606.18989](https://arxiv.org/html/2606.18989)
Fengying Ye1,11footnotemark:1Yanming Sun1,Runzhe Zhan1 Zheqi Zhang2Lidia S\. Chao1Derek F\. Wong1, 1NLP2CT Lab, Department of Computer and Information Science, University of Macau 2Faculty of Arts and Humanities, University of Macau nlp2ct\.\{fengying, yanming, runzhe\}@gmail\.com, \{lidiasc, derekfw\}@um\.edu\.mo

###### Abstract

Idioms are difficult to transfer across languages due to their non\-compositionality and weak surface\-form grounding, making literal mappings unreliable\. We presentG\-IdiomAlign, a gloss\-pivoted benchmark where each idiom is anchored by an English gloss from Wiktionary\. We further construct a high\-confidence reference alignment set for reproducible evaluation\. G\-IdiomAlign supports two protocols: \(1\) a controlled Multiple\-Choice Idiom Equivalence with typed distractors for error attribution; and \(2\) a Gloss\-Contrastive Generation contrastingNo\-glossandWith\-glossinputs to isolate the effect of an explicit semantic pivot\. Across diverse LLMs, a bias to literal translation is a dominant failure mode, especially when the target is a low\-resource language\. Glosses consistently improve Gloss\-Contrastive Generation under an embedding\-based semantic proxy, but performance remains modest, indicating substantial headroom in the open output space\. Subsequent analysis on Qwen3\-8B further suggests that cross\-condition differences are concentrated more in attention heads than in layers, while betterWith\-glossgenerations coincide with stronger gloss anchoring111Dataset:[https://github\.com/NLP2CT/G\-IdiomAlign](https://github.com/NLP2CT/G-IdiomAlign)\.

G\-IdiomAlign: A Gloss\-Pivoted Benchmark for Cross\-Lingual Idiom Alignment

Fengying Ye1,11footnotemark:1Yanming Sun1,††thanks:Equal Contribution\.Runzhe Zhan1Zheqi Zhang2Lidia S\. Chao1Derek F\. Wong1,††thanks:Corresponding Author\.1NLP2CT Lab, Department of Computer and Information Science, University of Macau2Faculty of Arts and Humanities, University of Macaunlp2ct\.\{fengying, yanming, runzhe\}@gmail\.com,\{lidiasc, derekfw\}@um\.edu\.mo

## 1Introduction

Idioms pose a persistent challenge for cross\-lingual meaning transfer because their figurative meanings are non\-compositional and culturally grounded, making word\-by\-word composition unreliableHeet al\.\([2025](https://arxiv.org/html/2606.18989#bib.bib20)\)\. Recent evidence suggests that large language models \(LLMs\) often over\-index on surface statistical cues \(such as collocational frequency or sentence probability\) rather than recovering figurative intent from contextMiet al\.\([2025](https://arxiv.org/html/2606.18989#bib.bib2)\); Yanget al\.\([2025c](https://arxiv.org/html/2606.18989#bib.bib13)\); Yeet al\.\([2026](https://arxiv.org/html/2606.18989#bib.bib38)\)\. Accurate cross\-lingual idiom alignment is crucial not only for cross\-cultural communication and machine\-assisted localization, but also as a litmus test for whether LLMs genuinely grasp culturally embedded semantics beyond surface patterns, motivating evaluation protocols that target idiom\-to\-idiom semantic equivalence over mere lexical overlap\.

However, existing resources offer limited support for controlled and diagnostic evaluation of cross\-lingual idiom\-to\-idiom equivalence\. Strong systems frequently produce literal, partial, or missing idiom renderingsYanget al\.\([2025c](https://arxiv.org/html/2606.18989#bib.bib13)\), yet current datasets lack unified benchmarks with standardized protocols for systematic error attribution\. To enable explicit and comparable semantic grounding across languages, we adopt English glosses as a shared*semantic pivot*, modeling a resource\-augmented setting where models can leverage external semantic support \(e\.g\., lexicons or knowledge bases\)\. We introduce a contrastive setup betweenNo\-glossandWith\-glossinputs to isolate the effect of this explicit semantic signal\.

Table 1:Example idiom pairs fromG\-IdiomAlignacross different language pairs\.IIandGGdenote idioms and their English glosses; subscripts indicate languages\.In this work, we introduceG\-IdiomAlign, a gloss\-pivoted idiom alignment benchmark across nine core languages, with coverage of four languages from underrepresented families, where each idiom is linked to a meaning\-equivalent English gloss from Wiktionary \(examples are shown in Table[1](https://arxiv.org/html/2606.18989#S1.T1)\)\. We construct a high\-confidence reference set via a precision\-first pipeline that combines distribution\-aware filtering with bidirectional one\-to\-one constraints\. On top of this benchmark, we provide two complementary evaluation settings: a Multiple\-Choice Idiom Equivalence task with typed distractors, and a Gloss\-Contrastive Generation underNo\-glossandWith\-glossinputs\.

Across diverse LLMs, models show a pervasive bias to literal translations\. Although adding glosses yields consistent but limited improvements under an embedding\-based semantic proxy, this underscores the difficulty of producing canonical, meaning\-equivalent idioms in an unconstrained space\. Subsequent attention\-based correlational analyses on Qwen3\-8B further suggest that improvedWith\-glossgenerations align with stronger gloss anchoring, with cross\-condition differences concentrating mainly at the level of attention heads rather than broad layer\-level shifts\. Our analysis positions G\-IdiomAlign as a foundation for future work on robust cross\-lingual idiom modeling\.

Our contributions are as follows: \(1\) We releaseG\-IdiomAlign, a gloss\-pivoted dataset covering 36 language pairs across nine core languages \(18,785 idiom pairs\), filtered via a bidirectional pipeline to ensure high semantic equivalence\. \(2\) We establish two diagnostic protocols for cross\-lingual idiom alignment: Multiple\-Choice Idiom Equivalence with typed distractors, and Gloss\-Contrastive Generation that contrastsNo\-glossandWith\-glossinputs to test semantic grounding\. \(3\) We reveal a widespread bias towards literal translation in LLMs and provide attention\-based evidence linking gloss usage to improved semantic anchoring\.

## 2Related Work

### 2\.1Idiom Benchmarks

Existing idiom benchmarks support two core tasks: detection, which identifies whether a phrase is used idiomatically, and disambiguation, which resolves whether an expression should be interpreted literally or figuratively\. Detection datasets test whether a model distinguishes idiomatic from literal usages, ranging from linguist\-curated contrastive setsMiet al\.\([2025](https://arxiv.org/html/2606.18989#bib.bib2)\)to English test suites such as IdioTSDe Luca Fornaciariet al\.\([2024](https://arxiv.org/html/2606.18989#bib.bib5)\), multilingual benchmarks ID10MTedeschiet al\.\([2022](https://arxiv.org/html/2606.18989#bib.bib6)\)and CLCL frameworkZhouet al\.\([2023](https://arxiv.org/html/2606.18989#bib.bib25)\)\. Disambiguation datasets including EPIESaxena and Paul \([2020](https://arxiv.org/html/2606.18989#bib.bib4)\), MAGPIEHaagsmaet al\.\([2020](https://arxiv.org/html/2606.18989#bib.bib1)\), and MultiCoPIESentsovaet al\.\([2025](https://arxiv.org/html/2606.18989#bib.bib3)\)label potential idiomatic expressions with literal versus idiomatic readings, supporting contextual sense selectionFakharian and Cook \([2021](https://arxiv.org/html/2606.18989#bib.bib24)\); Zhouet al\.\([2021](https://arxiv.org/html/2606.18989#bib.bib26)\)\. Complementary resources broaden coverage further: LIdiomsMoussallemet al\.\([2018](https://arxiv.org/html/2606.18989#bib.bib9)\)links idioms across languages as linked data, andFuet al\.\([2025](https://arxiv.org/html/2606.18989#bib.bib10)\)evaluate Chinese idioms across multiple competencies\. While these efforts are valuable for idiom recognition and interpretation, they do not directly target*idiom\-to\-idiom*meaning\-equivalence alignment within a unified cross\-lingual evaluation\.

### 2\.2Idiom Alignment

Cross\-lingual idiom alignment remains challenging because figurative meanings often diverge from literal forms and are shaped by language\- and culture\-specific conventionsMoussallemet al\.\([2018](https://arxiv.org/html/2606.18989#bib.bib9)\); Donthiet al\.\([2025](https://arxiv.org/html/2606.18989#bib.bib11)\)\. Recent evaluations confirm persistent failures in both NMT systems and LLMsYanget al\.\([2025c](https://arxiv.org/html/2606.18989#bib.bib13)\); Sunet al\.\([2026](https://arxiv.org/html/2606.18989#bib.bib37)\), prompting approaches that decompose translation into semantic analysis and candidate selectionQian \([2024](https://arxiv.org/html/2606.18989#bib.bib14)\)or inject external signals, such as retrieval\-augmented MT with loss weightingLiuet al\.\([2023](https://arxiv.org/html/2606.18989#bib.bib15)\)or multilingual idiom knowledge basesLiet al\.\([2024](https://arxiv.org/html/2606.18989#bib.bib16)\)\. However, these methods often rely on surface\-level cues:Sentsovaet al\.\([2025](https://arxiv.org/html/2606.18989#bib.bib3)\)report substantially higher performance on idioms with direct English lexical counterparts, and cross\-lingual evaluations note strong prompt sensitivity and performance gaps in low\-overlap language pairsKhoshtabet al\.\([2025](https://arxiv.org/html/2606.18989#bib.bib17)\)\. This reliance is exacerbated in retrieval\-based alignment frameworks like bilingual lexicon induction \(BLI\), which formulate cross\-lingual matching as nearest\-neighbor search over candidate setsLiet al\.\([2023](https://arxiv.org/html/2606.18989#bib.bib18)\)\. Such approaches are prone to false positivesDinget al\.\([2024](https://arxiv.org/html/2606.18989#bib.bib19)\), unless constrained by precision\-oriented criteria like bidirectional agreement\. To address these limitations, we move beyond surface\-driven retrieval by anchoring alignment in meaning\-equivalent English glosses and adopt bidirectional constraints to ensure high\-precision idiom\-to\-idiom pairing, thus enabling controlled evaluation that isolates semantic equivalence from lexical shortcuts\.

![Refer to caption](https://arxiv.org/html/2606.18989v1/x1.png)Figure 1:Overview of the G\-IdiomAlign construction pipeline\. Using English glosses as a shared semantic pivot, we extract idiom entries and core glosses from Wiktionary, retrieve top\-kkcandidates in a gloss\-embedding space, keep MNN pairs, and apply a pair\-specific distribution\-aware filter, yielding the final G\-IdiomAlign benchmark\.

## 3G\-IdiomAlign

We introduceG\-IdiomAlign, a gloss\-pivoted benchmark for cross\-lingual idiom alignment across nine core languages, with coverage of four languages from low\-resource language families\. English glosses from Wiktionary222https://www\.wiktionary\.org/serve as a shared semantic pivot, supporting cross\-lingual meaning comparison while mitigating shortcuts based on surface lexical overlap\. We construct G\-IdiomAlign with a precision\-first, staged construction pipeline: \(i\) extract idiom entries and core glosses \(excluding usage notes/examples\), \(ii\) retrieve top\-kkcandidates in a shared gloss\-embedding space, \(iii\) retain mutual nearest neighbor \(MNN\) pairs via bidirectional agreement, and \(iv\) apply distribution\-aware filtering to produce a high\-confidence alignment set\. The resulting benchmark is intended for evaluation and diagnostic analysis\. Figure[1](https://arxiv.org/html/2606.18989#S2.F1)summarizes the construction pipeline\.

### 3\.1Language Coverage & Data Collection

##### Language Scope\.

G\-IdiomAlign covers nine core languages: De, En, Es, Fi, Fr, Ja, Pl, Zh, and Pt, for which our extraction pipeline yields sufficient high\-quality aligned pairs\. To broaden language coverage, we further include four languages \(Arabic, Korean, Thai, and Vietnamese\) and report their results separately in Appendix[A](https://arxiv.org/html/2606.18989#A1)\.

##### Collection Pipeline\.

We collect idiom entries from language\-specific Wiktionary category pages and extract a cleaned core gloss from each entry’s sense definition, excluding auxiliary material such as usage notes and examples; see Appendix[B](https://arxiv.org/html/2606.18989#A2)for implementation details\. These glosses provide a consistent meaning description and function as the semantic pivot throughout construction\.

##### Single\-Sense Filtering\.

To preserve interpretability of the reference, we keep idioms with a single Wiktionary sense \(one gloss\) and remove polysemous entries\. This avoids one\-to\-many sense correspondences that would make idiom\-to\-idiom equivalence ambiguous at construction time\. The resulting reference set is smaller but cleaner, supporting more controlled evaluation and diagnosis\.

Table 2:G\-IdiomAlign language\-pair composition\.Ndenotes the count of aligned idiom pairs and%denotes the proportion of the dataset \(out of all aligned pairs\)\. We report each pair once using a canonical ordering\.
##### Gloss\-based Candidate Retrieval\.

For each directed language pairA→BA\\rightarrow B, we embed the glosses associated with idioms in both languages using the OpenAI text\-embedding\-3\-largeOpenAI \([2024](https://arxiv.org/html/2606.18989#bib.bib35)\)\. For a source idiomxxwith glossgxg\_\{x\}and a candidate idiomyywith glossgyg\_\{y\}, we define gloss similarity as

s​\(x,y\)=cos⁡\(E​\(gx\),E​\(gy\)\),s\(x,y\)=\\cos\\\!\\big\(E\(g\_\{x\}\),E\(g\_\{y\}\)\\big\),whereE​\(⋅\)E\(\\cdot\)denotes the embedding function\. For eachxx, we retrieve the top\-kkcandidates inBBbys​\(x,y\)s\(x,y\)withk=10k=10, producing a candidate set for subsequent bidirectional filtering\. Although final alignments are determined by rank\-1 agreement \(see below\), usingk\>1k\>1improves candidate coverage and robustness to embedding noise before enforcing one\-to\-one constraints\.

##### MNN Alignment\.

To obtain unambiguous evaluation pairs, we enforce a one\-to\-one matching constraint via mutual nearest neighbors \(MNN\)\. We retain a pair\(x,y\)\(x,y\)if and only ifxxandyyare rank\-1 nearest neighbors of each other under both directions \(A→BA\\rightarrow BandB→AB\\rightarrow A\)\. This bidirectional criterion removes asymmetric or many\-to\-one associations that may arise from retrieval artifacts, ensuring that retained alignments reflect strong mutual semantic correspondence\.

##### Distribution\-Aware Filtering\.

Since similarity score scales differ substantially across language pairs, fixed global thresholds can be poorly calibrated\. Moreover, even under MNN, nearest\-neighbor retrieval always returns a best match within the dataset, which can force alignments even when no true equivalent exists, leading to spurious pairs\. For example, a Chinese idiom “洞房花燭夜” \(gloss: the wedding night\) may be aligned with the English idiom “white marriage” \(gloss: an unconsummated marriage\): although their glosses share salient words \(wedding and marriage\), the underlying meanings are not equivalent\. Accordingly, we apply a language\-pair\-specific, parameter\-light cutoff to remove weak matches while preserving high\-confidence alignments\.

For each language pair, we collect the rank\-1 similarity scores of MNN\-confirmed pairs and discretize similarity scores within\-pair range into 10 equal\-width bins\. Letbbdenote the modal bin\. We retain pairs whose scores fall in binbbor higher, using the lower edge of the modal bin as a cutoff\. Similarity scores are used here as a diagnostic signal for relative strength within each language pair, rather than as an absolute criterion of semantic correctness \(details are shown in Appendix[C](https://arxiv.org/html/2606.18989#A3)\)\.

To ensure deterministic reporting, we compute similarity scores using a fixed canonical direction for each unordered language pair, while the MNN criterion itself is always enforced bidirectionally\.

### 3\.2Benchmark Statistics

G\-IdiomAlign comprises 18,785 aligned idiom pairs across 36*unordered*language pairs drawn from nine languages\. Although alignments are reported without direction, each pair supports evaluation in either direction \(e\.g\., Zh→\\rightarrowEn or En→\\rightarrowZh\)\. For reporting and aggregation, each unordered language pair is listed once using a canonical ordering\. Pair sizes range from 186 to 1,782, with a median of 422 \(interquartile range: 323–574\); the full breakdown is provided in Table[2](https://arxiv.org/html/2606.18989#S3.T2)\.

### 3\.3Alignment Quality Evaluation

We assess the semantic alignment quality of G\-IdiomAlign using both human evaluation and LLM\-based majority voting\. Sampled pairs are rated on a 3\-point scale: 2 denotes equivalent meaning and interchangeability across contexts; 1 denotes partial equivalence, where meanings are close but differ in tone, intensity, or pragmatics; and 0 denotes non\-equivalence, where lexical or topical relatedness does not imply semantic equivalence\.

For Zh–En idiom pairs, we randomly sample 200 pairs and evaluate them with native speakers and senior Ph\.D\. students with expertise in relevant languages\. For the remaining language pairs, we sample 50 pairs per language pair and score them independently with GPT\-5\.1OpenAI \([2026](https://arxiv.org/html/2606.18989#bib.bib36)\), Gemini\-2\.5\-ProGemini Team \([2025](https://arxiv.org/html/2606.18989#bib.bib29)\), and Claude\-4\.5\-HaikuAnthropic \([2025](https://arxiv.org/html/2606.18989#bib.bib30)\)\. LLM judges are prompted to follow the same annotation instructions as human annotators\. We use majority voting as the final label; when all judges disagree, we assign score 1 to reflect partial equivalence\.

We report strict accuracy \(only score\-2 pairs\), and lenient accuracy \(score\-1 and score\-2 pairs\)\. Across non\-Zh–En language pairs, the LLM\-based evaluation yields a mean strict accuracy of 0\.685 and a lenient accuracy of 0\.923, where each language pair is treated as one observation\. The corresponding 95% confidence intervals are computed using a t\-interval over language pairs \(strict: \[0\.645, 0\.724\]; lenient: \[0\.907, 0\.940\]\)\. Details are provided in Appendix[D](https://arxiv.org/html/2606.18989#A4)\.

In addition, performance varies substantially across language pairs\. High\-resource or closely related pairs such as En–De, En–Es, En–Pt, and Pt–Es achieve very high strict accuracy \(up to 0\.96\), with most annotations assigned score 2\. In contrast, more distant pairs such as De–Ja, Fr–Ja, and Zh–Es show lower accuracy and a larger proportion of score 1 and score 0 cases, reflecting the greater difficulty of establishing idiomatic equivalence across typologically distant languages\.

### 3\.4Similarity Characterization

Complementing the alignment quality evaluation, we analyze gloss\-based similarity scores to understand how embedding\-space signals support and characterize the constructed alignments\.

##### Overall distribution\.

Across 18,785 aligned pairs, construction\-time similarity scores span a wide range \(min=0\.368=0\.368, max≈1\.0\\approx 1\.0\) with moderately high central tendency \(mean=0\.670=0\.670, median=0\.651=0\.651\)\. Near\-saturation scores are rare and typically correspond to highly formulaic or nearly identical glosses \(see Appendix[E\.1](https://arxiv.org/html/2606.18989#A5.SS1)\)\.

##### Threshold sensitivity\.

To assess the effect of encoder calibration on coverage, we sweep nine thresholdsttover the shared overlap interval of the two encoders’ score distributions and count pairs withs​i​m≥tsim\\geq t\(see Appendix[E\.2](https://arxiv.org/html/2606.18989#A5.SS2)\)\. In Figure[2](https://arxiv.org/html/2606.18989#S3.F2), both encoders produce monotonic retention curves but diverge substantially at the same absolute threshold, reflecting calibration and scaling differences rather than semantic disagreement\. This supports treating similarity as a relative diagnostic signal and motivates distribution\-aware filtering\.

![Refer to caption](https://arxiv.org/html/2606.18989v1/x2.png)Figure 2:Retention curves for OpenAI text\-embedding\-3\-large and Qwen3\-Embedding\-8B as the similarity thresholdttvaries, illustrating encoder\-dependent calibration effects under fixed absolute thresholds\.
##### Embedding consistency and calibration\.

Embedding consistency refers to the agreement in relative similarity structure across different embedding models\. When we recompute similarities with independent multilingual encoder, Qwen3\-Embedding\-8BYanget al\.\([2025b](https://arxiv.org/html/2606.18989#bib.bib34)\), we observe a systematic increase in absolute cosine similarity while maintaining strong agreement in relative structure \(Pearsonr=0\.807r=0\.807, Spearmanρ=0\.784\\rho=0\.784; see Appendix[E\.3](https://arxiv.org/html/2606.18989#A5.SS3)\)\. This suggests that absolute similarity values are encoder\-dependent, while relative similarity structure is largely preserved\.

## 4Experiments

We evaluate cross\-lingual idiom alignment under two complementary settings\. Task 1 formulates alignment as a controlled multiple\-choice problem, enabling fine\-grained diagnosis of error patterns through typed distractors\. Task 2 evaluates open\-ended target\-idiom generation in a large output space and contrastsNo\-glossandWith\-glossinputs to assess the effect of an explicit semantic pivot under semantic ambiguity and surface\-form mismatch\. These settings support reproducible quantitative comparison and reveal recurring failure modes in cross\-lingual idiom alignment\.

Table 3:Multiple\-choice accuracy aggregated by target\-language groups\. Micro is instance\-weighted within each target group \(Zh\-/En\-/Other\-target\), while Macro averages over directions\. All numbers are percentages\.### 4\.1Task 1:Multiple\-Choice Idiom Equivalence

Task formulation\.Given a source\-language idiom, the model selects the meaning\-equivalent target\-language option from a 4\-way candidate set\. Each instance contains one canonical target idiom from G\-IdiomAlign and three typed distractors, enabling error analysis by distractor type in addition to accuracy\. To reduce positional bias, we shuffle the option order per instance and store option\-type labels \(reference, LT, LC, CA; defined below\)\. The model outputs a single choice \(A\-D\) under greedy decoding \(temperature=0=0, top\-p=1p=1\)\. We evaluate 30 direction settings in total, including 16 directions with*high\-resource*target languages \(Chinese or English; 8 each\) and 14 additional directions with other target languages\. Here a*direction*is an ordered mapping from a source language to a target language, and we group directions by the target language \(Zh\-target, En\-target, and other targets\)\.

Candidate construction\.The correct \(reference\) option is the canonical target idiom from G\-IdiomAlign\. We construct three types of target\-language distractors:Literal Translation Trap\(LT\), a word\-for\-word translation of the source idiom;Lexical Cue Trap\(LC\), a target\-language idiom that shares a*partial*lexical cue with the literal translation \(typically one salient content word\) but conveys an unrelated meaning; andContextual Association Trap\(CA\), a target\-language idiom that is contextually plausible yet semantically opposite\. Distractors are generated using Qwen\-Max under a unified prompt with explicit type constraints; the details are provided in the Appendix[F\.1](https://arxiv.org/html/2606.18989#A6.SS1)\.

##### Validity checks for LLM\-generated distractors\.

To assess whether LLM\-generated distractors introduce superficial shortcuts, we conduct an option\-only control and a manual validity check\. The results show no significant preference for the gold option over chance and confirm 81\.5% distractor\-type validity\. Details are provided in Appendix[F\.2](https://arxiv.org/html/2606.18989#A6.SS2)\.

### 4\.2Task 2: Gloss\-Contrastive Generation

##### Task formulation\.

In the open\-ended setting, the model is given a source\-language idiom and must generate a meaning\-equivalent idiom in the target language\. We compare two input conditions:No\-gloss\(source idiom only\) andWith\-gloss\(source idiom plus its English gloss\), where the gloss provides an explicit semantic pivot\. We evaluate Task 2 on all 72 directions available in G\-IdiomAlign, using greedy decoding \(temperature=0=0, top\-p=1p=1\)\. We require models to output exactly one target\-language idiom with no additional explanation, enabling deterministic parsing and automatic scoring \(see Appendix[G](https://arxiv.org/html/2606.18989#A7)\)\.

##### Automatic evaluation\.

Because multiple outputs can be valid in open\-ended generation, we use an embedding\-based semantic matching proxy for coarse\-grained scoring\. We embed the model output and the canonical target idiom in G\-IdiomAlign using Qwen3\-Embedding\-8B and compute cosine similarity\. We report Acc@tt, counting a prediction as correct if the similarity exceeds a thresholdttin the same embedding space\. We choose two operating points,t=0\.70t=0\.70andt=0\.80t=0\.80:0\.700\.70is a more permissive threshold, while0\.800\.80is a stricter threshold close to the median similarity of canonical aligned pairs in G\-IdiomAlign under Qwen3\-Embedding\-8B \(median≈0\.78\\approx 0\.78\)\. This proxy supports aggregate comparison but is not a definitive correctness criterion; for example, it can under\-count valid synonymous idioms that diverge from the canonical reference\. We therefore treat embedding similarity as a high\-confidence semantic indicator rather than a complete estimate of idiom\-form correctness, and report a small human evaluation together with auxiliary surface\-form metrics \(EM/BLEU/ChrF\) in Appendix[J](https://arxiv.org/html/2606.18989#A10)\. We interpret Task 2 results primarily as comparative trends\.

### 4\.3Models

We evaluate several open\-source LLMs and proprietary, including DeepSeek\-V3\.2 \(T/NT\)DeepSeek\-AIet al\.\([2025](https://arxiv.org/html/2606.18989#bib.bib28)\), Gemini\-2\.5\-Pro, Claude\-4\.5\-Haiku, Ministral\-8B\-InstructMistral AI team \([2024](https://arxiv.org/html/2606.18989#bib.bib31)\), and Qwen3\-8B \(T/NT\)Yanget al\.\([2025a](https://arxiv.org/html/2606.18989#bib.bib32)\)\. T denotes*Thinking mode*and NT is*No thinking mode*\. During dataset preprocessing, we embed glosses with text\-embedding\-3\-large to select high\-confidence reference alignments\. For Task 1, we generate typed distractors using Qwen\-MaxTeam \([2025](https://arxiv.org/html/2606.18989#bib.bib33)\)\. For Task 2 automatic scoring, we use Qwen3\-Embedding\-8B\. We run open\-source models on a NVIDIA A40 GPU, while proprietary models are accessed via official APIs\.

Table 4:Gloss\-Contrastive Generation results\. MeanSim \(100×s​i​m100\\times sim\) and Acc@ttare direction\-averaged \(macro\) over 72 language directions; Acc@ttcounts a prediction as correct if cosines​i​m≥tsim\\geq t\. We reportt=0\.70t=0\.70andt=0\.80t=0\.80\.Δ\\DeltaAcc@0\.80 is theWith\-glossminusNo\-glossimprovement\. Bold indicates the best performance\.![Refer to caption](https://arxiv.org/html/2606.18989v1/x3.png)Figure 3:Task 1 outcomes by target\-language regime\. Stacked bars \(normalized within each regime\) decompose specific instances into Correct predictions and LT/LC/CA, highlighting differences between high\-resource targets \(Zh/En\) and other target languages\.

## 5Results and Analysis

### 5\.1Multiple\-Choice Idiom Equivalence

#### 5\.1\.1Overall Accuracy Across Target Groups

Table[3](https://arxiv.org/html/2606.18989#S4.T3)reports Task 1 accuracy by target\-language group\. Across models, performance is consistently lower on Other targets directions than on Zh\-target or En\-target, suggesting that idiom alignment is more challenging when the target language is not a high\-resource language such as Chinese or English\. This gap is most pronounced for lower\-performing models\. Gemini\-2\.5\-Pro achieves the strongest overall performance, particularly on Other targets\. Enabling*thinking mode*yields consistent gains across models, with clear improvements for DeepSeek\-V3\.2 and Qwen3\-8B\.

#### 5\.1\.2Outcome Decomposition

Figure[3](https://arxiv.org/html/2606.18989#S4.F3)decomposes Task 1 outcomes by target\-language regime intoCorrectpredictions and three distractor types \(LT/LC/CA\); proportions are in Appendix[H\.1](https://arxiv.org/html/2606.18989#A8.SS1)\. These patterns are further illustrated with representative examples and error cases in Appendix[H\.2](https://arxiv.org/html/2606.18989#A8.SS2), which provide concrete instances of LT\-, LC\-, and CA\-driven errors\. The Other targets regime has fewer correct predictions, consistent with lower accuracy\.

Across models,Literal Translation Trap\(LT\) dominates errors, especially for lower\-resource targets, indicating stronger literal\-transfer attraction\. In contrast,Zh\-targetdirections show moreLexical Cue\(LC\) errors, suggesting partial lexical overlap misleads models when Chinese is the target\. Lower\-performing models also exhibit higherContextual Association\(CA\) rates, particularly in Other targets\.

Enabling*thinking mode*improves accuracy for DeepSeek\-V3\.2 \(reducing LT and LC errors\) and Qwen3\-8B \(mainly decreasing LC and CA, with LT still dominant\)\. Overall, literal\-translation attraction remains the primary challenge in cross\-lingual idiom alignment\.

### 5\.2Gloss\-Contrastive Generation

Overall performance\.Table[4](https://arxiv.org/html/2606.18989#S4.T4)summarizes Task 2 results under the semantic\-similarity proxy\. Providing an English gloss \(With\-gloss\) improves MeanSim and Acc@ttfor every model, with consistent gains at the stricter threshold Acc@0\.80, suggesting that glosses help constrain generation toward the intended meaning\. Despite this, Acc@0\.80 remains modest even with gloss, highlighting the difficulty of producing meaning\-equivalent idioms in an unconstrained output space\.

UnderWith\-gloss, DeepSeek\-V3\.2 \(T\) achieves the strongest overall performance\. For models with both variants, enabling thinking yields consistent improvements, with the advantage more apparent at higher similarity thresholds\. Results at additional thresholds are reported in Appendix[I\.1](https://arxiv.org/html/2606.18989#A9.SS1)\. Representative examples and error cases for open\-ended generation are provided in Appendix[I\.2](https://arxiv.org/html/2606.18989#A9.SS2), illustrating both correct outputs and acceptable paraphrases that may be under\- or over\-estimated by the similarity\-based metric\.

### 5\.3Attention\-based Diagnostics

We introduce attention\-based diagnostics for Task 2 to characterize how the model allocates attention over*input spans*during decoding and how these allocations relate to output quality\. Throughout, we treat attention strictly as acorrelational diagnostic signal\. All analyses are conducted onQwen3\-8Bwith greedy decoding \(temperature=0=0, top\-p=1p=1\)\. Post\-softmax self\-attention weights are extracted usingTransformerLenshooks333https://github\.com/TransformerLensOrg/TransformerLens\.

For each input condition, we annotate 200 generations \(400 total\) and exclude Type 0 invalid outputs \(empty, garbled, or not interpretable as meaningful target\-language text\), yieldingn=189n=189valid instances inWith\-glossandn=189n=189inNo\-gloss\. Table[5](https://arxiv.org/html/2606.18989#S5.T5)reports the outcome breakdown by error type\. Two collaborators independently label the outputs \(Cohen’sκ=0\.81\\kappa=0\.81\) and resolve disagreements by discussion\. We use the following error taxonomy: Type 2 \(T2, literal word\-by\-word translation missing the idiom’s figurative meaning\), Type 3 \(T3, meaning\-correct but non\-idiomatic\), and Type 4 \(T4, meaning\-incorrect and not a word\-by\-word literal translation\)\.

Table 5:Outcome composition for the annotated subset \(Type 0 excluded\)\. T2/T3/T4 denote the breakdown of error types within Wrong outputs\.##### Head/layer\-level structural overlap\.

To assess whether cross\-condition differences are driven more by head changes or layer shifts, we compare the cross\-condition overlap of the salient heads and salient layers \(Table[6](https://arxiv.org/html/2606.18989#S5.T6); see Appendix[L](https://arxiv.org/html/2606.18989#A12)for details\)\.

As shown in Table[6](https://arxiv.org/html/2606.18989#S5.T6), head overlap is substantially lower than layer overlap overall \(Jheads=0\.32J\_\{\\text\{heads\}\}=0\.32vs\.Jlayers=0\.90J\_\{\\text\{layers\}\}=0\.90\)\. This pattern also holds across subsets, although layer overlap is lower for Correct\-only than for Wrong\-only outputs\. These results suggest that cross\-condition differences are expressed more strongly through head\-level reconfiguration than broad layer\-level shifts: the two conditions largely recruit similar layers but differ in which heads within those layers are salient\. Lower layer overlap for Correct\-only outputs further suggests less shared layer\-level structure across conditions for correct than for wrong cases\.

Table 6:Cross\-condition overlap betweenWith\-glossandNo\-glossfor salient head sets and salient layer sets \(Jaccard; higher indicates more overlap\)\.Table 7:Token\-level attention diagnostics \(Correct vs\. Wrong\)\. Each metric measures where the model attends during generation: IAR, GAR, DR, and OtherTop1\. We report median \(IQR\) for each group, Cliff’sδ\\delta, andqq\-values from two\-sided Mann\-Whitney U tests with BH\-FDR correction within each condition\.Table 8:Token\-level attention diagnostics across error types \(Type 2/3/4\)\. We report median \(IQR\) per error type andqq\-values from Kruskal\-Wallis tests with BH\-FDR correction within each condition\.
##### Token\-level diagnostics\.

LetA¯​\(k\)\\bar\{A\}\(k\)denote the aggregated post\-softmax attention mass assigned to input key tokenkk, computed by averaging attention over layers and heads and over template\-defined generation positions corresponding to the content\-bearing segment\. Herekkranges over key\-token positions in the full prompt sequence\. We interpretA¯​\(k\)\\bar\{A\}\(k\)as the model’s average attention mass assigned to tokenkkduring the generation of the content\-bearing output \(see Appendix[M](https://arxiv.org/html/2606.18989#A13)\)\.

Using tokenizer offset mapping, we map the idiom span and \(when available\) the gloss span to token\-index setsKidiomK\_\{\\text\{idiom\}\}andKglossK\_\{\\text\{gloss\}\}in the full prompt sequence\. We then define:

IAR=∑k∈KidiomA¯​\(k\)\\mathrm\{IAR\}=\\sum\_\{k\\in K\_\{\\text\{idiom\}\}\}\\bar\{A\}\(k\)IAR\\mathrm\{IAR\}\(Idiom Attention Ratio\) is the total attention mass on the idiom span; larger values indicate stronger concentration on idiom tokens\.

GAR=∑k∈KglossA¯​\(k\)\(With\-gloss\)\\mathrm\{GAR\}=\\sum\_\{k\\in K\_\{\\text\{gloss\}\}\}\\bar\{A\}\(k\)\\ \\ \(\\textit\{With\-gloss\}\)GAR\\mathrm\{GAR\}\(Gloss Attention Ratio\) is the total attention mass on the gloss span \(defined only underWith\-gloss\); larger values indicate stronger anchoring to the provided gloss\.

DR=1−IAR−GAR\\mathrm\{DR\}=1\-\\mathrm\{IAR\}\-\\mathrm\{GAR\}DR\\mathrm\{DR\}\(Diffuse Ratio\) captures the residual attention mass outside the tracked spans; larger values indicate greater allocation to other context tokens\. UnderNo\-gloss,GAR\\mathrm\{GAR\}is defined as0, soDR=1−IAR\\mathrm\{DR\}=1\-\\mathrm\{IAR\}\.

OtherTop1=maxk∈Kother⁡A¯​\(k\)∑j∈KotherA¯​\(j\)\\mathrm\{OtherTop1\}=\\max\_\{k\\in K\_\{\\mathrm\{other\}\}\}\\frac\{\\bar\{A\}\(k\)\}\{\\sum\_\{j\\in K\_\{\\mathrm\{other\}\}\}\\bar\{A\}\(j\)\}
OtherTop1\\mathrm\{OtherTop1\}is the maximum share among off\-span tokens after renormalizing within the off\-span set; higher values indicate a stronger off\-span peak\.

For Correct\-vs\.\-Wrong comparisons, we use two\-sided Mann\-Whitney U tests, a non\-parametric two\-group comparison, and report Cliff’sδ\\delta, whose sign indicates the direction of the difference and whose magnitude reflects its strength\. For comparisons across Types 2/3/4 within wrong outputs, we use Kruskal\-Wallis tests, which assess whether the error types differ overall without assuming normality\. Within each condition,pp\-values are adjusted across metric\-wise tests using BH\-FDR\. Table[7](https://arxiv.org/html/2606.18989#S5.T7)and[8](https://arxiv.org/html/2606.18989#S5.T8)summarize the result of token\-level diagnostics\.

*With\-gloss*\(Type 0 excluded;n=189n=189\)\. Correct outputs exhibit stronger gloss anchoring and reduced off\-span allocation: GAR increases, while both DR and OtherTop1 decrease \(allq<0\.05q<0\.05\)\. By contrast, idiom\-span mass does not distinguish Correct from Wrong \(IAR;q=0\.250q=0\.250\)\. Within wrong outputs, cross\-type differences are not robust after correction\.

*No\-gloss*\(Type 0 excluded;n=189n=189; GAR not applicable\)\. The Correct\-vs\.\-Wrong contrast is weaker and does not survive correction for IAR or DR, and OtherTop1 shows no meaningful difference\. Nevertheless, WrongType comparisons are strongly structured by error type across the available diagnostics \(IAR/DR/OtherTop1; allq<0\.001q<0\.001\), suggesting more heterogeneous failure modes in the absence of an explicit gloss anchor\.

Overall, glosses consistently improve open\-ended idiom generation, and attention diagnostics suggest that correctness underWith\-glossaligns with stronger gloss anchoring, whereasNo\-glosserrors exhibit more heterogeneous attention patterns across error types\.

## 6Conclusion

We presentG\-IdiomAlign, a gloss\-pivoted benchmark supporting Multiple\-Choice Idiom Equivalence and Gloss\-Contrastive Generation to diagnose literal\-translation biases across LLMs\. Our results, spanning diverse proprietary and open\-source models, highlight literal\-translation attraction as a persistent obstacle in cross\-lingual idiom alignment\. Attention\-based diagnostics further suggest that successfulWith\-glossgenerations are associated with stronger anchoring to gloss information\. While providing explicit glosses consistently improves open\-ended generation under an embedding\-based semantic proxy, performance remains far from saturated\. This motivates further developments in robust idiom translation\.

## Limitations

G\-IdiomAlign is designed as a precision\-first benchmark for diagnosing cross\-lingual idiom alignment rather than exhaustive idiomatic equivalence\. This improves interpretability and reproducibility, but limits coverage and external validity\.

English\-pivot bias\.We use English Wiktionary glosses as a single semantic pivot across nine languages\. This improves consistency, but may introduce English\-centric bias because glosses can compress pragmatic or culture\-specific meaning and vary in style and granularity across editions\.

Trade\-offs in reference alignments\.To ensure unambiguous supervision, we apply single\-sense filtering and exclude polysemous idioms\. This improves interpretability but removes sense selection and reduces coverage\. We further impose a one\-to\-one constraint via MNN, which favors high\-precision pairs but under\-represents many\-to\-many relations such as synonym clusters\.

Dependence on embeddings, proxies, and tools\.The pipeline relies on embedding\-based retrieval and filtering, as well as an LLM for distractor generation\. As a result, benchmark construction is sensitive to embedding calibration, dataset size, and gloss noise, and may inherit model\-specific biases\. Task 2 further uses fixed\-threshold embedding similarity as a scalable proxy, which may miss valid non\-canonical generations and may not be fully comparable across languages, such dependence remains a limitation\.

Limited generality\.Our attention analyses are correlational rather than causal, and are based on a single model with a modest annotated sample under greedy decoding\. The observed patterns may therefore not generalize across models, decoding settings, or language directions\.

## Ethical Considerations

This work involves several value\-sensitive design choices\. First, we use English glosses as a shared semantic pivot to enable controlled cross\-lingual idiom alignment\. While this results in high\-confidence alignment, it may introduce English\-centric bias and compress culture\-specific pragmatic or stylistic distinctions encoded in non\-English idioms\. We treat this as a deliberate trade\-off for diagnostic clarity, rather than as a claim of cultural neutrality\.

Second, idioms are culturally grounded expressions, and operationalizing idiomatic equivalence through glosses and embedding\-based similarity necessarily abstracts away contextual and sociocultural nuance\. Our benchmark is therefore intended to support analysis of model behavior under controlled conditions, not to define authoritative judgments of idiomatic correctness across cultures\.

Finally, our evaluation metrics, especially the embedding\-based proxy in Gloss\-Contrastive Generation, are designed for consistent comparison rather than deployment\. We caution against using benchmark scores as standalone indicators of translation quality or fairness in real\-world applications\.

## Acknowledgments

This work was supported in part by the Science and Technology Development Fund of Macau SAR \(Grant Nos\. FDCT/0007/2024/AKP, EF2024\-00185\-FST\), the UM and UMDF \(Grant Nos\. MYRG\-GRG2024\-00165\-FST\-UMDF, MYRG\-GRG2025\-00236\-FST\), the Tencent AI Lab Rhino\-Bird Research Program \(Grant No\. EF2023\-00151\-FST\), the Stanley Ho Medical Development Foundation \(Grant No\. SHMDF\-AI/2026/001\), and the National Natural Science Foundation of China \(Grant No\. 62266013\)\.

## References

- Anthropic \(2025\)Claude Haiku 4\.5\.External Links:[Link](https://www.anthropic.com/claude/haiku)Cited by:[§3\.3](https://arxiv.org/html/2606.18989#S3.SS3.p2.1)\.
- F\. De Luca Fornaciari, B\. Altuna, I\. Gonzalez\-Dios, and M\. Melero \(2024\)A Hard Nut to Crack: Idiom Detection with Conversational Large Language Models\.InProceedings of the 4th Workshop on Figurative Language Processing \(FigLang 2024\),D\. Ghosh, S\. Muresan, A\. Feldman, T\. Chakrabarty, and E\. Liu \(Eds\.\),Mexico City, Mexico \(Hybrid\),pp\. 35–44\.External Links:[Link](https://aclanthology.org/2024.figlang-1.5/),[Document](https://dx.doi.org/10.18653/v1/2024.figlang-1.5)Cited by:[§2\.1](https://arxiv.org/html/2606.18989#S2.SS1.p1.1)\.
- DeepSeek\-AI, A\. Liu, A\. Mei, B\. Lin, B\. Xue, B\. Wang, B\. Xu, B\. Wu, B\. Zhang, C\. Lin, C\. Dong, C\. Lu, C\. Zhao, C\. Deng, C\. Xu, C\. Ruan, D\. Dai, D\. Guo, D\. Yang, D\. Chen, E\. Li, F\. Zhou, F\. Lin, F\. Dai, G\. Hao, G\. Chen, G\. Li, H\. Zhang, H\. Xu, H\. Li, H\. Liang, H\. Wei, H\. Zhang, H\. Luo, H\. Ji, H\. Ding, H\. Tang, H\. Cao, H\. Gao, H\. Qu, H\. Zeng, J\. Huang, J\. Li, J\. Xu, J\. Hu, J\. Chen, J\. Xiang, J\. Yuan, J\. Cheng, J\. Zhu, J\. Ran, J\. Jiang, J\. Qiu, J\. Li, J\. Song, K\. Dong, K\. Gao, K\. Guan, K\. Huang, K\. Zhou, K\. Huang, K\. Yu, L\. Wang, L\. Zhang, L\. Wang, L\. Zhao, L\. Yin, L\. Guo, L\. Luo, L\. Ma, L\. Wang, L\. Zhang, M\. S\. Di, M\. Y\. Xu, M\. Zhang, M\. Zhang, M\. Tang, M\. Zhou, P\. Huang, P\. Cong, P\. Wang, Q\. Wang, Q\. Zhu, Q\. Li, Q\. Chen, Q\. Du, R\. Xu, R\. Ge, R\. Zhang, R\. Pan, R\. Wang, R\. Yin, R\. Xu, R\. Shen, R\. Zhang, S\. H\. Liu, S\. Lu, S\. Zhou, S\. Chen, S\. Cai, S\. Chen, S\. Hu, S\. Liu, S\. Hu, S\. Ma, S\. Wang, S\. Yu, S\. Zhou, S\. Pan, S\. Zhou, T\. Ni, T\. Yun, T\. Pei, T\. Ye, T\. Yue, W\. Zeng, W\. Liu, W\. Liang, W\. Pang, W\. Luo, W\. Gao, W\. Zhang, X\. Gao, X\. Wang, X\. Bi, X\. Liu, X\. Wang, X\. Chen, X\. Zhang, X\. Nie, X\. Cheng, X\. Liu, X\. Xie, X\. Liu, X\. Yu, X\. Li, X\. Yang, X\. Li, X\. Chen, X\. Su, X\. Pan, X\. Lin, X\. Fu, Y\. Q\. Wang, Y\. Zhang, Y\. Xu, Y\. Ma, Y\. Li, Y\. Li, Y\. Zhao, Y\. Sun, Y\. Wang, Y\. Qian, Y\. Yu, Y\. Zhang, Y\. Ding, Y\. Shi, Y\. Xiong, Y\. He, Y\. Zhou, Y\. Zhong, Y\. Piao, Y\. Wang, Y\. Chen, Y\. Tan, Y\. Wei, Y\. Ma, Y\. Liu, Y\. Yang, Y\. Guo, Y\. Wu, Y\. Wu, Y\. Cheng, Y\. Ou, Y\. Xu, Y\. Wang, Y\. Gong, Y\. Wu, Y\. Zou, Y\. Li, Y\. Xiong, Y\. Luo, Y\. You, Y\. Liu, Y\. Zhou, Z\. F\. Wu, Z\. Z\. Ren, Z\. Zhao, Z\. Ren, Z\. Sha, Z\. Fu, Z\. Xu, Z\. Xie, Z\. Zhang, Z\. Hao, Z\. Gou, Z\. Ma, Z\. Yan, Z\. Shao, Z\. Huang, Z\. Wu, Z\. Li, Z\. Zhang, Z\. Xu, Z\. Wang, Z\. Gu, Z\. Zhu, Z\. Li, Z\. Zhang, Z\. Xie, Z\. Gao, Z\. Pan, Z\. Yao, B\. Feng, H\. Li, J\. L\. Cai, J\. Ni, L\. Xu, M\. Li, N\. Tian, R\. J\. Chen, R\. L\. Jin, S\. S\. Li, S\. Zhou, T\. Sun, X\. Q\. Li, X\. Jin, X\. Shen, X\. Chen, X\. Song, X\. Zhou, Y\. X\. Zhu, Y\. Huang, Y\. Li, Y\. Zheng, Y\. Zhu, Y\. Ma, Z\. Huang, Z\. Xu, Z\. Zhang, D\. Ji, J\. Liang, J\. Guo, J\. Chen, L\. Xia, M\. Wang, M\. Li, P\. Zhang, R\. Chen, S\. Sun, S\. Wu, S\. Ye, T\. Wang, W\. L\. Xiao, W\. An, X\. Wang, X\. Sun, X\. Wang, Y\. Tang, Y\. Zha, Z\. Zhang, Z\. Ju, Z\. Zhang, and Z\. Qu \(2025\)DeepSeek\-V3\.2: Pushing the Frontier of Open Large Language Models\.External Links:2512\.02556,[Link](https://arxiv.org/abs/2512.02556)Cited by:[§4\.3](https://arxiv.org/html/2606.18989#S4.SS3.p1.1)\.
- Q\. Ding, H\. Cao, and T\. Zhao \(2024\)Enhancing bilingual lexicon induction via bi\-directional translation pair retrieving\.Proceedings of the AAAI Conference on Artificial Intelligence38\(16\),pp\. 17898–17906\.External Links:[Link](https://ojs.aaai.org/index.php/AAAI/article/view/29744),[Document](https://dx.doi.org/10.1609/aaai.v38i16.29744)Cited by:[§2\.2](https://arxiv.org/html/2606.18989#S2.SS2.p1.1)\.
- S\. Donthi, M\. Spencer, O\. B\. Patel, J\. Y\. Doh, E\. Rodan, K\. Zhu, and S\. O’Brien \(2025\)Improving LLM abilities in idiomatic translation\.InProceedings of the First Workshop on Language Models for Low\-Resource Languages,H\. Hettiarachchi, T\. Ranasinghe, P\. Rayson, R\. Mitkov, M\. Gaber, D\. Premasiri, F\. A\. Tan, and L\. Uyangodage \(Eds\.\),Abu Dhabi, United Arab Emirates,pp\. 175–181\.External Links:[Link](https://aclanthology.org/2025.loreslm-1.13/)Cited by:[§2\.2](https://arxiv.org/html/2606.18989#S2.SS2.p1.1)\.
- S\. Fakharian and P\. Cook \(2021\)Contextualized embeddings encode monolingual and cross\-lingual knowledge of idiomaticity\.InProceedings of the 17th Workshop on Multiword Expressions \(MWE 2021\),P\. Cook, J\. Mitrović, C\. P\. Escartín, A\. Vaidya, P\. Osenova, S\. Taslimipoor, and C\. Ramisch \(Eds\.\),Online,pp\. 23–32\.External Links:[Link](https://aclanthology.org/2021.mwe-1.4/),[Document](https://dx.doi.org/10.18653/v1/2021.mwe-1.4)Cited by:[§2\.1](https://arxiv.org/html/2606.18989#S2.SS1.p1.1)\.
- Y\. Fu, Z\. Huang, L\. Yang, Y\. Lu, and Z\. Dai \(2025\)CHENGYU\-BENCH: benchmarking large language models for Chinese idiom understanding and use\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 2355–2366\.External Links:[Link](https://aclanthology.org/2025.emnlp-main.119/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.119),ISBN 979\-8\-89176\-332\-6Cited by:[§2\.1](https://arxiv.org/html/2606.18989#S2.SS1.p1.1)\.
- G\. Gemini Team \(2025\)Gemini 2\.5: pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities\.arXiv preprint arXiv:2507\.06261\.External Links:[Link](https://arxiv.org/abs/2507.06261)Cited by:[§3\.3](https://arxiv.org/html/2606.18989#S3.SS3.p2.1)\.
- H\. Haagsma, J\. Bos, and M\. Nissim \(2020\)MAGPIE: a large corpus of potentially idiomatic expressions\.InProceedings of the Twelfth Language Resources and Evaluation Conference,N\. Calzolari, F\. Béchet, P\. Blache, K\. Choukri, C\. Cieri, T\. Declerck, S\. Goggi, H\. Isahara, B\. Maegaard, J\. Mariani, H\. Mazo, A\. Moreno, J\. Odijk, and S\. Piperidis \(Eds\.\),Marseille, France,pp\. 279–287\(eng\)\.External Links:[Link](https://aclanthology.org/2020.lrec-1.35/),ISBN 979\-10\-95546\-34\-4Cited by:[§2\.1](https://arxiv.org/html/2606.18989#S2.SS1.p1.1)\.
- W\. He, T\. K\. Vieira, M\. Garcia, C\. Scarton, M\. Idiart, and A\. Villavicencio \(2025\)Investigating idiomaticity in word representations\.Computational Linguistics51,pp\. 505–555\.External Links:[Link](https://aclanthology.org/2025.cl-2.4/),[Document](https://dx.doi.org/10.1162/coli%5Fa%5F00546)Cited by:[§1](https://arxiv.org/html/2606.18989#S1.p1.1)\.
- P\. Khoshtab, D\. Namazifard, M\. Masoudi, A\. Akhgary, S\. Mahdizadeh Sani, and Y\. Yaghoobzadeh \(2025\)Comparative study of multilingual idioms and similes in large language models\.InProceedings of the 31st International Conference on Computational Linguistics,O\. Rambow, L\. Wanner, M\. Apidianaki, H\. Al\-Khalifa, B\. D\. Eugenio, and S\. Schockaert \(Eds\.\),Abu Dhabi, UAE,pp\. 8680–8698\.External Links:[Link](https://aclanthology.org/2025.coling-main.580/)Cited by:[§2\.2](https://arxiv.org/html/2606.18989#S2.SS2.p1.1)\.
- S\. Li, J\. Chen, S\. Yuan, X\. Wu, H\. Yang, S\. Tao, and Y\. Xiao \(2024\)Translate meanings, not just words: IdiomKB’s role in optimizing idiomatic translation with language models\.Proceedings of the AAAI Conference on Artificial Intelligence38\(17\),pp\. 18554–18563\.External Links:[Link](https://ojs.aaai.org/index.php/AAAI/article/view/29817),[Document](https://dx.doi.org/10.1609/aaai.v38i17.29817)Cited by:[§2\.2](https://arxiv.org/html/2606.18989#S2.SS2.p1.1)\.
- Y\. Li, A\. Korhonen, and I\. Vulić \(2023\)On bilingual lexicon induction with large language models\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,H\. Bouamor, J\. Pino, and K\. Bali \(Eds\.\),Singapore,pp\. 9577–9599\.External Links:[Link](https://aclanthology.org/2023.emnlp-main.595/),[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.595)Cited by:[§2\.2](https://arxiv.org/html/2606.18989#S2.SS2.p1.1)\.
- E\. Liu, A\. Chaudhary, and G\. Neubig \(2023\)Crossing the threshold: idiomatic machine translation through retrieval augmentation and loss weighting\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,H\. Bouamor, J\. Pino, and K\. Bali \(Eds\.\),Singapore,pp\. 15095–15111\.External Links:[Link](https://aclanthology.org/2023.emnlp-main.933/),[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.933)Cited by:[§2\.2](https://arxiv.org/html/2606.18989#S2.SS2.p1.1)\.
- M\. Mi, A\. Villavicencio, and N\. S\. Moosavi \(2025\)Rolling the DICE on idiomaticity: how LLMs fail to grasp context\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 7314–7332\.External Links:[Link](https://aclanthology.org/2025.acl-long.362/),[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.362),ISBN 979\-8\-89176\-251\-0Cited by:[§1](https://arxiv.org/html/2606.18989#S1.p1.1),[§2\.1](https://arxiv.org/html/2606.18989#S2.SS1.p1.1)\.
- Mistral AI team \(2024\)Un ministral, des ministraux\.External Links:[Link](https://mistral.ai/news/ministraux/)Cited by:[§4\.3](https://arxiv.org/html/2606.18989#S4.SS3.p1.1)\.
- D\. Moussallem, M\. A\. Sherif, D\. Esteves, M\. Zampieri, and A\. Ngonga Ngomo \(2018\)LIdioms: a multilingual linked idioms data set\.InProceedings of the Eleventh International Conference on Language Resources and Evaluation \(LREC 2018\),N\. Calzolari, K\. Choukri, C\. Cieri, T\. Declerck, S\. Goggi, K\. Hasida, H\. Isahara, B\. Maegaard, J\. Mariani, H\. Mazo, A\. Moreno, J\. Odijk, S\. Piperidis, and T\. Tokunaga \(Eds\.\),Miyazaki, Japan\.External Links:[Link](https://aclanthology.org/L18-1392/)Cited by:[§2\.1](https://arxiv.org/html/2606.18989#S2.SS1.p1.1),[§2\.2](https://arxiv.org/html/2606.18989#S2.SS2.p1.1)\.
- OpenAI \(2024\)New embedding models and API updates\.External Links:[Link](https://openai.com/index/new-embedding-models-and-api-updates/)Cited by:[§3\.1](https://arxiv.org/html/2606.18989#S3.SS1.SSS0.Px4.p1.5)\.
- OpenAI \(2026\)GPT‑5\.1: A smarter, more conversational ChatGPT\.External Links:[Link](https://openai.com/zh-Hans-CN/index/gpt-5-1/)Cited by:[§3\.3](https://arxiv.org/html/2606.18989#S3.SS3.p2.1)\.
- M\. Qian \(2024\)Automating idiom translation with cross\-lingual natural language generation grounded in semantic analyses using large language models\.InProceedings of the 16th Conference of the Association for Machine Translation in the Americas \(Volume 2: Presentations\),M\. Martindale, J\. Campbell, K\. Savenkov, and S\. Goel \(Eds\.\),Chicago, USA,pp\. 95–115\.External Links:[Link](https://aclanthology.org/2024.amta-presentations.7/)Cited by:[§2\.2](https://arxiv.org/html/2606.18989#S2.SS2.p1.1)\.
- P\. Saxena and S\. Paul \(2020\)EPIE Dataset: A Corpus For Possible Idiomatic Expressions\.External Links:2006\.09479Cited by:[§2\.1](https://arxiv.org/html/2606.18989#S2.SS1.p1.1)\.
- U\. Sentsova, D\. Ciminari, J\. V\. Genabith, and C\. España\-Bonet \(2025\)MultiCoPIE: a multilingual corpus of potentially idiomatic expressions for cross\-lingual PIE disambiguation\.InProceedings of the 21st Workshop on Multiword Expressions \(MWE 2025\),A\. Kr\. Ojha, V\. Giouli, V\. B\. Mititelu, M\. Constant, G\. Korvel, A\. S\. Doğruöz, and A\. Rademaker \(Eds\.\),Albuquerque, New Mexico, U\.S\.A\.,pp\. 67–81\.External Links:[Link](https://aclanthology.org/2025.mwe-1.8/),[Document](https://dx.doi.org/10.18653/v1/2025.mwe-1.8),ISBN 979\-8\-89176\-243\-5Cited by:[§2\.1](https://arxiv.org/html/2606.18989#S2.SS1.p1.1),[§2\.2](https://arxiv.org/html/2606.18989#S2.SS2.p1.1)\.
- Y\. Sun, R\. Zhan, C\. S\. Cheang, H\. Wu, X\. Liu, Y\. Niu, F\. Ye, K\. Lan, L\. S\. Chao, and D\. F\. Wong \(2026\)Exposing the cracks: vulnerabilities of retrieval\-augmented llm\-based machine translation\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.40,pp\. 33135–33143\.Cited by:[§2\.2](https://arxiv.org/html/2606.18989#S2.SS2.p1.1)\.
- Q\. Team \(2025\)Qwen3\-Max: Just Scale it\.External Links:[Link](https://qwen.ai/blog?id=qwen3-max)Cited by:[§4\.3](https://arxiv.org/html/2606.18989#S4.SS3.p1.1)\.
- S\. Tedeschi, F\. Martelli, and R\. Navigli \(2022\)ID10M: idiom identification in 10 languages\.InFindings of the Association for Computational Linguistics: NAACL 2022,M\. Carpuat, M\. de Marneffe, and I\. V\. Meza Ruiz \(Eds\.\),Seattle, United States,pp\. 2715–2726\.External Links:[Link](https://aclanthology.org/2022.findings-naacl.208/),[Document](https://dx.doi.org/10.18653/v1/2022.findings-naacl.208)Cited by:[§2\.1](https://arxiv.org/html/2606.18989#S2.SS1.p1.1)\.
- A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv, C\. Zheng, D\. Liu, F\. Zhou, F\. Huang, F\. Hu, H\. Ge, H\. Wei, H\. Lin, J\. Tang, J\. Yang, J\. Tu, J\. Zhang, J\. Yang, J\. Yang, J\. Zhou, J\. Zhou, J\. Lin, K\. Dang, K\. Bao, K\. Yang, L\. Yu, L\. Deng, M\. Li, M\. Xue, M\. Li, P\. Zhang, P\. Wang, Q\. Zhu, R\. Men, R\. Gao, S\. Liu, S\. Luo, T\. Li, T\. Tang, W\. Yin, X\. Ren, X\. Wang, X\. Zhang, X\. Ren, Y\. Fan, Y\. Su, Y\. Zhang, Y\. Zhang, Y\. Wan, Y\. Liu, Z\. Wang, Z\. Cui, Z\. Zhang, Z\. Zhou, and Z\. Qiu \(2025a\)Qwen3 technical report\.External Links:2505\.09388,[Link](https://arxiv.org/abs/2505.09388)Cited by:[§4\.3](https://arxiv.org/html/2606.18989#S4.SS3.p1.1)\.
- A\. Yang, J\. Lin, J\. Zhou, and et al\. \(2025b\)Qwen3\-embedding\-8B\.External Links:[Link](https://huggingface.co/Qwen/Qwen3-Embedding-8B)Cited by:[§3\.4](https://arxiv.org/html/2606.18989#S3.SS4.SSS0.Px3.p1.2)\.
- C\. Yang, Y\. Dou, D\. Heineman, X\. Wu, and W\. Xu \(2025c\)Evaluating LLMs on chinese idiom translation\.ArXivabs/2508\.10421\.External Links:[Link](https://api.semanticscholar.org/CorpusID:280649964)Cited by:[§1](https://arxiv.org/html/2606.18989#S1.p1.1),[§1](https://arxiv.org/html/2606.18989#S1.p2.1),[§2\.2](https://arxiv.org/html/2606.18989#S2.SS2.p1.1)\.
- F\. Ye, S\. Wang, L\. S\. Chao, and D\. F\. Wong \(2026\)Probing semantic alignment, lexical invariance, and syntactic influence in llm metaphor processing\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),San Diego, California\.Cited by:[§1](https://arxiv.org/html/2606.18989#S1.p1.1)\.
- J\. Zhou, H\. Gong, and S\. Bhat \(2021\)PIE: a parallel idiomatic expression corpus for idiomatic sentence generation and paraphrasing\.InProceedings of the 17th Workshop on Multiword Expressions \(MWE 2021\),P\. Cook, J\. Mitrović, C\. P\. Escartín, A\. Vaidya, P\. Osenova, S\. Taslimipoor, and C\. Ramisch \(Eds\.\),Online,pp\. 33–48\.External Links:[Link](https://aclanthology.org/2021.mwe-1.5/),[Document](https://dx.doi.org/10.18653/v1/2021.mwe-1.5)Cited by:[§2\.1](https://arxiv.org/html/2606.18989#S2.SS1.p1.1)\.
- J\. Zhou, Z\. Zeng, and S\. Bhat \(2023\)CLCL: non\-compositional expression detection with contrastive learning and curriculum learning\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),A\. Rogers, J\. Boyd\-Graber, and N\. Okazaki \(Eds\.\),Toronto, Canada,pp\. 730–743\.External Links:[Link](https://aclanthology.org/2023.acl-long.43/),[Document](https://dx.doi.org/10.18653/v1/2023.acl-long.43)Cited by:[§2\.1](https://arxiv.org/html/2606.18989#S2.SS1.p1.1)\.

## Appendix AAdditional Language Pairs

Table 9:Additional language\-pair composition in G\-IdiomAlign\.Ndenotes the count of aligned idiom pairs\. Each pair is reported once using a canonical ordering\.In addition to the core G\-IdiomAlign benchmark reported in Table[2](https://arxiv.org/html/2606.18989#S3.T2), we construct a supplementary set of additional language pairs to broaden cross\-lingual coverage in Table[9](https://arxiv.org/html/2606.18989#A1.T9)\. This extension introduces four new languages: Arabic \(Ar\), Korean \(Ko\), Thai \(Th\), and Vietnamese \(Vi\), and matches them with both the original benchmark languages and one another using the same gloss\-pivoted pipeline\.

Overall, this supplementary set contains 42 language pairs and 1,014 bidirectional rank\-1 aligned idiom pairs\. As in the core benchmark, these pairs are obtained through gloss\-based candidate retrieval, bidirectional agreement, and distribution\-aware filtering\. Because coverage remains limited after precision\-oriented filtering, we do not include these pairs in the main evaluation; instead, we release them to support future research on broader cross\-lingual idiom alignment\.

## Appendix BWiktionary Idiom Harvesting and Gloss Cleaning

This appendix describes our cross\-lingual harvesting framework for extracting idiom entries and their definition\-based glosses from Wiktionary\.

##### License\.

The data used in this section are derived from Wiktionary, a collaboratively constructed resource available under the Creative Commons Attribution\-ShareAlike 3\.0 License \(CC BY\-SA 3\.0\)\.

##### Harvesting\.

For each language, we enumerate idiom entry pages by traversing the corresponding idiom category on English Wiktionary and collecting linked entry pages across all “next page” partitions\. We optionally apply conservative filters to remove obvious auxiliary pages introduced by category\-page organization\.

##### Gloss extraction\.

From each entry page, we extract an example\-free definition string as the gloss representation\. The extraction prioritizes sense\-definition text and excludes usage examples and other non\-definitional material\.

##### Precision\-oriented screening\.

We enforce a strict single\-sense criterion: an entry is retained only if extraction yields exactly one candidate gloss string\. We then apply lightweight normalization to remove leading labels and standardize whitespace\.

## Appendix CDistribution\-Aware Filtering: Implementation Details

This appendix provides an implementation\-level specification of the distribution\-aware filtering step described in Section[3\.1](https://arxiv.org/html/2606.18989#S3.SS1)\. The filtering procedure operates on the set of MNN\-confirmed idiom pairs for a given language pair, each associated with its rank\-1 gloss similarity score\.

##### Input\.

For each*unordered*language pair, after enforcing mutual nearest neighbors \(MNN\), we obtain a set of candidate alignments𝒫=\{\(xi,yi\)\}i=1n,\\mathcal\{P\}=\\\{\(x\_\{i\},y\_\{i\}\)\\\}\_\{i=1\}^\{n\},where each retained pair is associated with a rank\-1 similarity scoresi∈𝐑s\_\{i\}\\in\\mathbf\{R\}, computed as cosine similarity between the corresponding English\-gloss embeddings \(as defined in Section[3\.1](https://arxiv.org/html/2606.18989#S3.SS1)\)\. The MNN criterion itself is always enforced bidirectionally\.

##### Equal\-width binning\.

Letsmin=mini⁡sis\_\{\\min\}=\\min\_\{i\}s\_\{i\}andsmax=maxi⁡sis\_\{\\max\}=\\max\_\{i\}s\_\{i\}\. We partition the interval\[smin,smax\]\[s\_\{\\min\},s\_\{\\max\}\]into 10 equal\-width bins using 11 bin edges:

ej=smin\+j⋅smax−smin10,j=0,1,…,10\.e\_\{j\}=s\_\{\\min\}\+j\\cdot\\frac\{s\_\{\\max\}\-s\_\{\\min\}\}\{10\},\\quad j=0,1,\\dots,10\.

##### Modal\-bin cutoff\.

Letcjc\_\{j\}denote the number of MNN\-confirmed pairs whose rank\-1 scores fall into binjj\. We define the modal bin as

b=arg⁡maxj∈\{1,…,10\}⁡cj\.b=\\arg\\max\_\{j\\in\\\{1,\\dots,10\\\}\}c\_\{j\}\.When multiple bins tie for the maximum count, ties are resolved by selecting the first maximizer returned by the implementation\.

##### Filtering rule\.

An MNN\-confirmed pair\(xi,yi\)\(x\_\{i\},y\_\{i\}\)is retained if and only if its rank\-1 similarity score falls in the modal bin or any higher bin:

\(xi,yi\)​is kept⇔bin​\(si\)≥b\.\(x\_\{i\},y\_\{i\}\)\\ \\text\{is kept\}\\iff\\text\{bin\}\(s\_\{i\}\)\\geq b\.Equivalently, the lower edge of the modal bin acts as a language\-pair\-specific cutoff\. As in the main text, similarity scores are treated as a relative diagnostic signal within each language pair, rather than as an absolute criterion of semantic correctness\.

## Appendix DAlignment Quality Evaluation Details

This appendix provides additional details on the annotation protocol, LLM prompting strategy, and statistical estimation used in Section[3\.3](https://arxiv.org/html/2606.18989#S3.SS3)\.

##### Annotation protocol\.

All idiom pairs are evaluated using a 3\-point semantic equivalence scale: 2 \(fully equivalent\), 1 \(partially equivalent\), and 0 \(non\-equivalent\)\. A score of 2 indicates that two idioms express the same core meaning and can be reasonably substituted in similar contexts; a score of 1 indicates partial equivalence with differences in tone, intensity, or pragmatic usage; and a score of 0 indicates non\-equivalence\. Annotators are instructed to focus on semantic meaning rather than literal form\. Both human annotators and LLM judges follow the same annotation instructions and scoring criteria\.

##### LLM prompting strategy\.

We implement LLM\-based evaluation by embedding the annotation rubric directly into structured prompts\. Each prompt takes as input a pair of idioms and their English glosses and requires the model to output an equivalence label \(0/1/2\)\.

To improve robustness and reduce prompt sensitivity, we adopt a multi\-view prompting strategy with complementary perspectives, including \(i\) direct semantic comparison, \(ii\) a substitutability\-based test that evaluates whether the two idioms can be used interchangeably in similar contexts, and \(iii\) comparison of pragmatic function and strength\. All prompts share the same core rules: prioritize English glosses over surface forms and avoid relying on literal similarity\.

Prompts are shown below:

> Task: Given two expressions in different languages and their English glosses, assign one equivalence label: \- 2 = Correct Equivalent \(same core meaning; substitutable in similar contexts\) \- 1 = Partially Correct \(overlapping meaning but differences in tone or usage\) \- 0 = Incorrect \(different core meaning or function\)\. \- Do not rely on literal similarity\. Input: \- Idiom A: <IDIOM\_A\> \- GLOSS A: <GLOSS\_A\_EN\> \- Idiom B: <IDIOM\_B\> \- GLOSS B: <GLOSS\_B\_EN\> Output: \- Equivalence: <0/1/2\>

> Task: Judge equivalence using a substitutable test\. Step 1: Based on the English glosses, imagine 2 short English contexts where this meaning is used\. Step 2: Decide if A and B could reasonably substitute each other in those contexts\. \- 2 = Correct Equivalent \(same core meaning; substitutable in similar contexts\) \- 1 = Partially Correct \(overlapping meaning but differences in tone or usage\) \- 0 = Incorrect \(different core meaning or function\)\. \- Do not rely on literal similarity\. Input: \- Idiom A: <IDIOM\_A\> \- GLOSS A: <GLOSS\_A\_EN\> \- Idiom B: <IDIOM\_B\> \- GLOSS B: <GLOSS\_B\_EN\> Output: \- Equivalence: <0/1/2\>

> Task: Judge equivalence by explicitly comparing pragmatic function and strength\. \- 2 = Correct Equivalent \(same core meaning; substitutable in similar contexts\) \- 1 = Partially Correct \(overlapping meaning but differences in tone or usage\) \- 0 = Incorrect \(different core meaning or function\)\. \- Do not rely on literal similarity\. Input: \- Idiom A: <IDIOM\_A\> \- GLOSS A: <GLOSS\_A\_EN\> \- Idiom B: <IDIOM\_B\> \- GLOSS B: <GLOSS\_B\_EN\> Output: \- Equivalence: <0/1/2\>

For each idiom pair, we obtain independent judgments from three LLMs \(GPT\-5\.1, Gemini\-2\.5\-pro, and Claude\-4\.5\-Haiku\)\. The final label is determined via majority voting across models\. When all three models disagree, we assign a score of 1 to avoid over\-claiming full equivalence while retaining borderline cases instead of discarding them\. This combination of multi\-prompt design and multi\-model aggregation improves the robustness and stability of LLM\-based judgments\.

Table 10:Global similarity statistics for the OpenAI text\-embedding\-3\-large and Qwen3\-Embedding\-8B, and the per\-instance differenceΔi=siQwen3−siorig\\Delta\_\{i\}=s\_\{i\}^\{\\text\{Qwen3\}\}\-s\_\{i\}^\{\\text\{orig\}\}overN=18,785N=18\{,\}785aligned pairs\.Table 11:Coverage under fixed similarity thresholds: number of aligned pairs withs​i​m≥tsim\\geq tunder each encoder\.
##### Sampling and statistical estimation\.

For Zh–En, we randomly sample 200 idiom pairs and evaluate them using human annotators\. For all other language pairs, we sample 50 idiom pairs per language pair and evaluate them using the LLM\-based protocol described above\.

For non\-Zh–En language pairs, each language pair is treated as one observation\. We compute the mean strict and lenient accuracy across language pairs and report 95% confidence intervals computed as t\-intervals over language pairs\.

For Zh–En, statistics are computed over individual samples \(n=200n=200\)\. The reported accuracies correspond to the proportion of samples satisfying each criterion, where strict accuracy counts only score\-2 pairs and lenient accuracy counts score\-1 and score\-2 pairs\.

##### Results summary\.

For non\-Zh–En language pairs, the mean strict accuracy is 0\.685 \(95% CI \[0\.645, 0\.724\]\) and the mean lenient accuracy is 0\.923 \(95% CI \[0\.907, 0\.940\]\)\. These correspond to 68\.5% fully equivalent pairs \(score = 2\) and 92\.3% partially or fully equivalent pairs \(score≥1\\geq 1\)\.

For Zh–En human evaluation, the strict accuracy is 0\.655 \(95% CI \[0\.589, 0\.721\]\) and the lenient accuracy is 0\.895 \(95% CI \[0\.852, 0\.938\]\)\. Despite being slightly more conservative, human evaluation yields results consistent with LLM\-based estimates\.

Overall, both LLM\-based and human evaluations provide converging evidence that G\-IdiomAlign achieves high semantic alignment quality, supporting the effectiveness of the proposed mining and filtering pipeline\.

## Appendix EModel Dependence of Similarity Scores

We assess the sensitivity of gloss\-based similarity scores to encoders by recomputing scores for the same aligned idiom pairs using Qwen3\-Embedding\-8B, an alternative multilingual embedding model\. We examine three aspects: global score calibration, threshold\-based coverage, and cross\-encoder consistency in relative score structure\.

### E\.1Global Statistics and Calibration Shift

For the sameN=18,785N=18\{,\}785aligned pairs, we recompute cosine similarities between source and target glosses using Qwen3\-Embedding\-8B and compare them with the construction\-time scores obtained using OpenAI text\-embedding\-3\-large\. Both encoders use cosine similarity overL2L\_\{2\}\-normalized embeddings\.

Table[10](https://arxiv.org/html/2606.18989#A4.T10)summarizes the original scoressorigs^\{\\text\{orig\}\}, the recomputed scoressQwen3s^\{\\text\{Qwen3\}\}, and the per\-instance differenceΔi=siQwen3−siorig\\Delta\_\{i\}=s\_\{i\}^\{\\text\{Qwen3\}\}\-s\_\{i\}^\{\\text\{orig\}\}\. Qwen3 produces systematically higher absolute similarity values than the original encoder \(mean0\.67→0\.780\.67\\rightarrow 0\.78; median0\.65→0\.780\.65\\rightarrow 0\.78\), with an average shift ofΔ​μ=0\.11\\Delta\\mu=0\.11\. However, the shift is heterogeneous across instances \(p10=0\.01p\_\{10\}=0\.01,p90=0\.21p\_\{90\}=0\.21\) and includes negative values, indicating that the difference is not reducible to a simple global offset or rescaling\.

### E\.2Threshold Sensitivity Under Calibration Shift

To illustrate the practical effect of calibration differences, we count the number of aligned pairs satisfyings≥ts\\geq tunder fixed absolute thresholdstt\. We sweep nine thresholds uniformly over the intersection of the two encoders’ 10th–90th percentile score intervals\.

Table[11](https://arxiv.org/html/2606.18989#A4.T11)shows that the number of retained pairs differs substantially across encoders at the same threshold\. This shows that absolute similarity thresholds are not directly comparable across embedding models in this setting, supporting our treatment of similarity as a relative diagnostic signal in the main text\.

### E\.3Embedding Consistency

Despite the shift in absolute similarity values, the relative score structure is largely preserved across encoders\. Across allN=18,785N=18\{,\}785aligned pairs, the two score sets are strongly correlated \(Pearsonr=0\.807r=0\.807\) and show substantial rank agreement \(Spearmanρ=0\.784\\rho=0\.784\)\. Thus, encoder choice has a larger effect on absolute calibration and threshold\-based coverage than on the comparative ordering of aligned pairs\.

## Appendix FTask 1 Distractor Generation

This appendix reports the unified prompt template used to generate the three typed distractors for Task 1 \(Multiple\-Choice Idiom Equivalence\)\. For each instance, the canonical target idiom from G\-IdiomAlign is used as the reference option, while Qwen\-Max is used*only*to generate the remaining three distractors under a single prompt: Literal Translation Trap \(LT\), Lexical Cue Trap \(LC\), and Contextual Association Trap \(CA\)\.

### F\.1Prompt

##### Model and decoding\.

We generate distractors usingQwen\-Maxwith greedy decoding \(temperature=0=0, top\-p=1p=1\)\.

##### Prompt template\.

The following prompt is used verbatim in our implementation, with placeholders instantiated per instance\.

> You are an expert linguist specializing in cross\-cultural idiom translation and test design\. Your task is to create a multiple\-choice question dataset to test whether an AI model truly understands idioms or just relies on literal translation\. Input Data: \- Source Idiom: <SOURCE\_IDIOM\> \- Source Meaning: <SOURCE\_MEANING\> \- Source Language: <SOURCE\_LANGUAGE\> \- Target Language: <TARGET\_LANGUAGE\> Task: Generate 3 options for a multiple\-choice question\. 1\. Option \(Literal Translation Trap\): A direct, word\-for\-word translation of the source in <TARGET\_LANGUAGE\>\. 2\. Option \(Lexical Cue Trap\): A real idiom in <TARGET\_LANGUAGE\> that shares only part of a salient keyword from the literal translation but has a completely DIFFERENT meaning\. 3\. Option \(Contextual Association Trap\): A real idiom in <TARGET\_LANGUAGE\> that has a related context but opposite meaning\. Constraints: \- The ’Lexical Cue Trap’ should clearly reflect a salient lexical cue from the literal translation\. Output strictly in JSON format\.

##### Distractor intent\.

The prompt defines three distractor types for diagnosis by error category: \(i\)LT \(Literal Translation Trap\)is a word\-by\-word rendering of the source idiom into the target language, designed to be surface\-faithful but not meaning\-equivalent; \(ii\)LC \(Lexical Cue Trap\)is a target\-language idiom that overlaps with a salient lexical cue from the literal translation while conveying a different meaning; \(iii\)CA \(Contextual Association Trap\)is a target\-language idiom that is contextually related yet semantically opposite to the intended meaning\.

##### Instance assembly and shuffling\.

For each benchmark alignment pair\(x,y\)\(x,y\)in G\-IdiomAlign, we take the canonical reference target idiomyyas the correct option and populate the remaining three options with the generated LT/LC/CA distractors\. We then shuffle the option order per instance to reduce positional bias\.

### F\.2Validity of LLM\-Generated Distractors

We conduct a control in which the model is shown only the answer options, without the question stem\. To avoid a trivial signal, we exclude the literal\-translation hard negatives in this control, since they are intentionally designed to be non\-idiomatic\. The goal is to test whether the model can systematically prefer the gold idiomatic option based on surface properties alone\.

Across three runs, the average selection rate for the gold option type is 0\.3567, compared with a random baseline of approximately 0\.3333\. A chi\-square test does not show a significant deviation from a uniform distribution \(χ2=2\.94\\chi^\{2\}=2\.94,p=0\.23p=0\.23\)\. This suggests that there is no strong evidence that the model can reliably identify the gold option from stylistic cues alone\. To further reduce superficial shortcuts, we also shuffle option order and normalize option formatting\.

Table 12:Task 1 outcome proportions by target\-language regime\. Each row reports themicrofraction \(%\) of instances that areCorrector correspond to choosing one of the three typed distractors: Literal Translation Trap \(LT\), Lexical Cue Trap \(LC\), and Contextual Association Trap \(CA\)\. Values sum to 100% within each model–regime block \(up to rounding\)\.Table 13:Representative examples for Task 1 \(multiple\-choice\)\. Each instance contains one correct target idiom and three typed distractors\. ✓ indicates a correct prediction and ✗ an incorrect one\.##### Manual verification of distractor\-type validity\.

We additionally manually inspect 200 questions to verify whether the generated distractors match their intended categories \(e\.g\., LT, LC, and CA\)\. We adopt a strict per\-question criterion: a question is counted as valid only if all four options match their intended types\. Under this criterion, 163 out of 200 questions are valid, corresponding to an accuracy of 81\.5%\.

These results indicate that the generated distractors largely satisfy the intended hard\-negative constraints and are not trivially distinguishable by superficial signals alone\.

## Appendix GTask 2 Generation Prompts and Output Constraints

This appendix reports the prompts used for Task 2 \(Gloss\-Contrastive Generation\) under the two input conditions:With\-gloss\(source idiom plus its English gloss\) andNo\-gloss\(source idiom only\)\.

##### Prompt template \(With\-gloss\)\.

In theWith\-glosscondition, we provide the source idiom together with its English gloss as an explicit semantic pivot, and ask the model to generate a meaning\-equivalent idiom in the target language:

> Output the <TARGET\_LANGUAGE\> idiom corresponding to "<SOURCE\_IDIOM\>" with meaning "<GLOSS\>"\. Return only one idiom and do not include any explanation or additional text\.

##### Prompt template \(No\-gloss\)\.

In theNo\-glosscondition, we provide only the source idiom and ask the model to generate the corresponding idiom in the target language:

> Output the <TARGET\_LANGUAGE\> idiom corresponding to "<SOURCE\_IDIOM\>"\. Return only one idiom and do not include any explanation or additional text\.

Table 14:Task 2 threshold sensitivity in a single table with two panels\. MeanSim is100×s​i​m100\\times simmacro\-averaged across directions\. Acc@ttis macro\-averaged accuracy \(%\) at thresholdtt\. In theWith\-glosspanel, each Acc@ttcell reports the absolute accuracy followed by the per\-threshold gain overNo\-glossin parentheses \(percentage points\)\.

## Appendix HTask 1 Outcomes

### H\.1Breakdown by Distractor Type

This appendix reports the numeric outcome proportions corresponding to Figure[3](https://arxiv.org/html/2606.18989#S4.F3)\. For each model and target\-language regime \(Zh\-target, En\-target, and Other targets\), we decompose outcomes into theCorrectselection and three typed distractor selections:Literal Translation Trap\(LT\),Lexical Cue Trap\(LC\), andContextual Association Trap\(CA\)\. All values in Table[12](https://arxiv.org/html/2606.18989#A6.T12)are regime\-levelmicroproportions \(instance\-weighted within each regime\) computed over the Task 1 evaluation instances for that regime; within each model–regime block, the four percentages sum to 100% \(up to rounding\)\. These numeric breakdowns support the error\-type comparisons discussed in Section[5\.1\.2](https://arxiv.org/html/2606.18989#S5.SS1.SSS2)\.

### H\.2Case Study: Task 1 \(Multiple\-choice\)

Table[13](https://arxiv.org/html/2606.18989#A6.T13)presents representative examples from Task 1 to illustrate how the multiple\-choice design probes different types of distractors\. Each instance includes one correct target idiom and three typed distractors: a Literal Translation Trap, a Lexical Cue Trap, and a Contextual Association Trap\.

The examples show that models can succeed when they recover the intended figurative meaning \(e\.g\., 一丈差九尺→\\rightarrowwide of the mark\), but models may select \(i\) a literal translation that closely mirrors the source form \(e\.g\., 殺人不眨眼\), \(ii\) an option triggered by salient lexical cues \(e\.g\., 仆心仆肺\), or \(iii\) an expression that is semantically related but not equivalent \(e\.g\., 善罷甘休\)\.

These cases highlight that Task 1 not only evaluates overall accuracy, but also reveals distinct and interpretable failure modes through the use of typed distractors\.

Table 15:Representative examples for Task 2 \(open\-ended generation\) with embedding\-based similarity scores\. ✓ indicates correct or acceptable paraphrase; ✗ indicates incorrect predictions\.

## Appendix ITask 2 Outcomes

### I\.1Threshold Sensitivity

Table 16:Auxiliary surface\-form metrics for Task 2\.We report Task 2 results under additional similarity thresholdst∈\{0\.65,0\.70,0\.75,0\.80\}t\\in\\\{0\.65,0\.70,0\.75,0\.80\\\}in Table[14](https://arxiv.org/html/2606.18989#A7.T14)\. MeanSim is reported as100×s​i​m100\\times simand macro\-averaged over directions\. Acc@ttis macro\-averaged accuracy \(%\) at thresholdtt\. For theWith\-glosspanel, we additionally report the per\-threshold improvement overNo\-glossin parentheses \(percentage points\)\.

### I\.2Case Study: Task 2 \(Open\-ended Generation\)

Table[15](https://arxiv.org/html/2606.18989#A8.T15)presents representative examples from Task 2 to illustrate model behavior in open\-ended idiom generation\. Unlike Task 1, this setting does not constrain models to a fixed set of options, and predictions are evaluated based on semantic equivalence or acceptable paraphrases of the target idiom\.

The examples show that models can produce correct idiomatic expressions \(e\.g\., ir al grano→\\rightarrowcut to the chase\) or acceptable paraphrases \(e\.g\., 倒錢落海→\\rightarrowpour money down the drain\), but errors often reflect difficulties in capturing the precise figurative meaning\. In particular, models may generate expressions that are overly general \(e\.g\., 修身養性→\\rightarrowpractice restraint\), semantically shifted \(e\.g\., turn over a new leaf\), or unrelated to the intended meaning \(e\.g\., de plantilla\)\.

## Appendix JTask 2 Supplementary Evaluation

##### Human calibration\.

To provide a point of calibration for the embedding\-based metric, we conduct a small human evaluation on Task 2 outputs\. We manually annotateN=86N=86examples for idiom\-to\-idiom correctness, obtaining an overall correct ratio of0\.4070\.407\. Using these annotations as ground truth, embedding similarity yields an AUC of0\.770\.77, with a 95% confidence interval of\[0\.67,0\.86\]\[0\.67,0\.86\], computed by nonparametric bootstrap resampling over the 86 examples\.

At a representative thresholdt=0\.75t=0\.75, the proxy achieves precision0\.8750\.875\(T​P=7TP=7,F​P=1FP=1\) with coverage9\.3%9\.3\\%\. These values indicate that the metric is most reliable for identifying a relatively high\-confidence subset of semantically correct outputs, rather than for exhaustively capturing all acceptable generations\.

##### Surface\-form metrics\.

As a complement to semantic matching, we also report auxiliary surface\-form metrics against the canonical gold reference in Table[16](https://arxiv.org/html/2606.18989#A9.T16)\. Specifically, we compute Exact Match \(EM\), and additionally report BLEU and ChrF as reference\-based overlap measures\. These metrics provide a view of whether a model output matches an attested or canonical idiom form\.

As expected in open\-ended generation, surface\-form metrics should be interpreted with caution: semantically valid idiomatic paraphrases may still receive low scores if they differ from the canonical reference string\. For this reason, we treat EM, BLEU, and ChrF as supplementary indicators\. Although absolute EM values are low, the directional trends remain informative\.

## Appendix KHuman Annotation Scheme for Gloss\-Contrastive Generation

We categorize each idiom chosen from Gloss\-Contrastive Generation into one of five mutually exclusive labels\{0,1,2,3,4\}\\\{0,1,2,3,4\\\}\. Labels are assigned following the decision procedure below\.

##### Label definitions\.

- •Label 0 \(Invalid / generation failure\)\.The output is empty, gibberish, or otherwise not interpretable as meaningful text in the target language\.
- •Label 1 \(Correct idiomatic translation\)\.The output is a fluent and semantically correct translation that realizes an appropriate idiomatic expression in the target language, conveying the intended meaning of the source idiom\.
- •Label 2 \(Literal word\-by\-word translation\)\.The output is semantically linked to the idiom’s surface form and appears to translate the idiom compositionally \(word\-by\-word or phrase\-by\-phrase\), rather than conveying the intended idiomatic meaning\.
- •Label 3 \(Meaning paraphrase, non\-idiomatic form\)\.The output correctly conveys the intended meaning of the source idiom, but does so using a non\-idiomatic paraphrase \(i\.e\., not an idiom or conventional idiomatic expression in the target language\)\.
- •Label 4 \(Incorrect meaning\)\.The output is meaningful text in the target language but fails to convey the intended meaning of the source idiom, including cases of mistranslation, wrong sense, or unrelated content\.

##### Decision procedure\.

Labels are assigned in the following order to ensure mutual exclusivity:

1. 1\.If the output is not interpretable as meaningful text in the target language, assignLabel 0\.
2. 2\.Otherwise, if the output is a fluent and correct idiomatic translation that appropriately realizes the source idiom in the target language, assignLabel 1\.
3. 3\.Otherwise, if the output translates the idiom literally based on its surface form without conveying the intended idiomatic meaning, assignLabel 2\.
4. 4\.Otherwise, if the output correctly conveys the intended meaning but does not use an idiomatic expression in the target language, assignLabel 3\.
5. 5\.All remaining meaningful but semantically incorrect outputs are assignedLabel 4\.

## Appendix LAttention Diagnostics: Salient Heads and Layer Aggregation

This appendix defines the salient head and layer sets used in Table[6](https://arxiv.org/html/2606.18989#S5.T6)\. All computations use post\-softmax attention weights from Qwen3\-8B usingTransformerLensunder greedy decoding\. We consider three subsetsuu:All,Correct\-only, andWrong\-only, under two conditionscc:With\-glossandNo\-gloss\.

##### Per\-sample head scores\.

LetAi,t→k\(ℓ,h\)A^\{\(\\ell,h\)\}\_\{i,t\\rightarrow k\}denote the attention weight from generation positionttto key positionkkat layerℓ\\elland headhhfor sampleii\. Let𝒯i\\mathcal\{T\}\_\{i\}be the set of template\-defined generation positions andKiK\_\{i\}the tracked token span \(e\.g\., idiom span or gloss span\)\. The head score is

Si\(ℓ,h\)=1\|𝒯i\|​∑t∈𝒯i∑k∈KiAi,t→k\(ℓ,h\),S\_\{i\}^\{\(\\ell,h\)\}=\\frac\{1\}\{\|\\mathcal\{T\}\_\{i\}\|\}\\sum\_\{t\\in\\mathcal\{T\}\_\{i\}\}\\sum\_\{k\\in K\_\{i\}\}A^\{\(\\ell,h\)\}\_\{i,t\\rightarrow k\},or the corresponding ratio\-based variant for contrastive analyses\.

##### Salient heads\.

For each sampleii, we retain the top\-kkheads ranked bySi\(ℓ,h\)S\_\{i\}^\{\(\\ell,h\)\}\(k=10k\{=\}10\)\. We then count how often each head appears across samples and define the salient head setℋsal\(c,u\)\\mathcal\{H\}\_\{\\mathrm\{sal\}\}^\{\(c,u\)\}as them=50m\{=\}50most frequent heads\.

##### Salient layers\.

We aggregate layers from the per\-sample top\-kkheads\. For each layerℓ\\ell, we sum the number of per\-sample top\-kkentries belonging toℓ\\ellacross all samples, and retain the top\-nnlayers by this count as the salient layer setℒsal\(c,u\)\\mathcal\{L\}\_\{\\mathrm\{sal\}\}^\{\(c,u\)\}\(n=10n\{=\}10\)\.

##### Cross\-condition overlap\.

For each subsetuu, we report the Jaccard similarity betweenWith\-glossandNo\-glosssalient sets, for both heads and layers\. Higher values indicate greater structural overlap across conditions\. These are the values reported in Table[6](https://arxiv.org/html/2606.18989#S5.T6)\.

## Appendix MToken\-level Attention Aggregation Details

This section details the implementation of the token\-level attention diagnostics used in Section[5\.3](https://arxiv.org/html/2606.18989#S5.SS3)\. All attention weights are extracted from Qwen3\-8B hooks and correspond to post\-softmax decoder self\-attention under greedy decoding\.

##### Attention extraction\.

For each generation, we collect the self\-attention tensor at each layer and head, yielding attention weightsAt→k\(ℓ,h\)A^\{\(\\ell,h\)\}\_\{t\\rightarrow k\}from generation query positionttto key token positionkk\. Attention is defined over the full prompt token sequence, including template tokens\.

##### Generation positions\.

Let𝒯\\mathcal\{T\}denote the set of template\-defined generation positions used for analysis\. In practice,𝒯\\mathcal\{T\}corresponds to the decoding positions associated with the content\-bearing segment of the output\. No additional filtering based on token type \(e\.g\., punctuation\) is applied beyond this template\-based selection\.

##### Token span identification\.

Token spans corresponding to the idiom and \(when present\) the gloss are identified by mapping character offsets in the prompt text to token index ranges using the model tokenizer\. These spans are defined with respect to the full prompt tokenization and are not re\-indexed to exclude template tokens\.

##### Token\-level attention mass\.

For each key token positionkk, we compute the aggregated attention mass

A¯​\(k\)=1\|𝒯\|​L​H​∑ℓ=1L∑h=1H∑t∈𝒯At→k\(ℓ,h\)\.\\bar\{A\}\(k\)=\\frac\{1\}\{\|\\mathcal\{T\}\|LH\}\\sum\_\{\\ell=1\}^\{L\}\\sum\_\{h=1\}^\{H\}\\sum\_\{t\\in\\mathcal\{T\}\}A^\{\(\\ell,h\)\}\_\{t\\rightarrow k\}\.This quantity represents the average post\-softmax attention mass assigned to tokenkkacross layers, heads, and selected generation positions\.

Using the aggregated attention massA¯​\(k\)\\bar\{A\}\(k\), we define the token\-level diagnostics reported in Section[5\.3](https://arxiv.org/html/2606.18989#S5.SS3)\.

Similar Articles

IdioLink: Retrieving Meaning Beyond Words Across Idiomatic and Literal Expressions

arXiv cs.CL

Introduces IdioLink, a retrieval benchmark of 10,700 documents and 2,140 queries across 107 idioms that tests whether models can link idiomatic expressions to conceptually equivalent literal or paraphrased meanings. Evaluations show current embedding models struggle with this task, highlighting gaps in idiom-aware semantic retrieval.

PolyAlign: Conditional Human-Distribution Alignment

arXiv cs.CL

PolyAlign is a distribution-aware alignment framework that aligns language models to context-specific human response distributions rather than a single global style, improving naturalness and faithfulness across bilingual settings.