ALEE: Any-Language Evaluation of Embeddings via English-Centric Minimal Pairs

arXiv cs.CL Papers

Summary

Introduces ALEE, a framework that uses Abstract Meaning Representations to generate English minimal pairs with controlled semantic shifts and translates them for evaluating text embeddings across 275+ languages, revealing persistent gaps in cross-lingual semantic representation.

arXiv:2607.00171v1 Announce Type: new Abstract: Text embeddings are standard for semantic similarity tasks, yet their evaluation remains an open challenge. Current benchmarks are static, cover only a limited set of languages, are often domain-specific, susceptible to overfitting, and poorly representative of low-resource languages. To address these limitations, we introduce ALEE, a framework that extends Sentence Smith (Li et al., 2025) to the cross-lingual and paragraph level. ALEE uses Abstract Meaning Representations (AMR) to generate English minimal pairs with controlled, fine-grained semantic shifts, which are paired with translations in target languages. This approach enables targeted diagnostics for models in any language with English parallel data. We conduct a large-scale empirical study across a diverse set of embedding models and 275+ languages spanning three parallel datasets. On ALEE, performance varies substantially across languages, text lengths, and linguistic phenomena, exposing persistent gaps in cross-lingual semantic representation that track language prevalence in training resources and subword tokenization. We release ALEE at https://github.com/Andrian0s/any-lang-embed-eval
Original Article
View Cached Full Text

Cached at: 07/02/26, 05:36 AM

# ALEE: Any-Language Evaluation of Embeddings via English-Centric Minimal Pairs
Source: [https://arxiv.org/abs/2607.00171](https://arxiv.org/abs/2607.00171)
[View PDF](https://arxiv.org/pdf/2607.00171)

> Abstract:Text embeddings are standard for semantic similarity tasks, yet their evaluation remains an open challenge\. Current benchmarks are static, cover only a limited set of languages, are often domain\-specific, susceptible to overfitting, and poorly representative of low\-resource languages\. To address these limitations, we introduce ALEE, a framework that extends Sentence Smith \(Li et al\., 2025\) to the cross\-lingual and paragraph level\. ALEE uses Abstract Meaning Representations \(AMR\) to generate English minimal pairs with controlled, fine\-grained semantic shifts, which are paired with translations in target languages\. This approach enables targeted diagnostics for models in any language with English parallel data\. We conduct a large\-scale empirical study across a diverse set of embedding models and 275\+ languages spanning three parallel datasets\. On ALEE, performance varies substantially across languages, text lengths, and linguistic phenomena, exposing persistent gaps in cross\-lingual semantic representation that track language prevalence in training resources and subword tokenization\. We release ALEE at[this https URL](https://github.com/Andrian0s/any-lang-embed-eval)

## Submission history

From: Andrianos Michail \[[view email](https://arxiv.org/show-email/c8a3d487/2607.00171)\] **\[v1\]**Tue, 30 Jun 2026 20:45:17 UTC \(2,903 KB\)

Similar Articles