LAMAR: An Open Language-Aware Multilingual Alignment Reranker

Hugging Face Daily Papers Papers

Summary

LAMAR is a language-aware multilingual cross-encoder reranker that uses English-anchored relevance distillation and preference alignment to prioritize documents in the same language as the query while preserving semantic relevance, achieving strong performance on multilingual benchmarks.

In multilingual retrieval augmented generation, a retriever can retrieve relevant documents written in multiple languages, which are subsequently reranked before answer generation. However, it remains unclear whether existing multilingual rerankers consider document language when ordering semantically relevant candidates. Our analysis shows that these rerankers do not consistently prioritize documents written in the same language as the query when semantically equivalent documents are available across languages, even though document language can affect answer generation. We release LAMAR, a language aware multilingual cross encoder trained to account for both semantic relevance and language coherence. LAMAR first uses English anchored relevance distillation to establish consistent relevance scoring across multilingual inputs and then applies preference alignment for language coherence to encourage documents written in the same language as the query to receive higher rankings while retaining semantic relevance. In a controlled experiment designed to assess language coherence, LAMAR achieves the best performance overall and across all languages examined individually. LAMAR also remains competitive on established multilingual reranking benchmarks. In practical retrieval settings, LAMAR achieves the best results across all reported metrics when reranking candidates retrieved in the first stage. These results demonstrate that LAMAR accounts for language coherence while achieving strong performance on general multilingual reranking benchmarks.
Original Article
View Cached Full Text

Cached at: 07/27/26, 05:40 AM

Paper page - LAMAR: An Open Language-Aware Multilingual Alignment Reranker

Source: https://huggingface.co/papers/2607.22042

Abstract

Inmultilingualretrievalaugmentedgeneration,aretrievercanretrieverelevantdocumentswritteninmultiplelanguages,whicharesubsequentlyrerankedbeforeanswergeneration.However,itremainsunclearwhetherexistingmultilingualrerankersconsiderdocumentlanguagewhenorderingsemanticallyrelevantcandidates.Ouranalysisshowsthatthesererankersdonotconsistentlyprioritizedocumentswritteninthesamelanguageasthequerywhensemanticallyequivalentdocumentsareavailableacrosslanguages,eventhoughdocumentlanguagecanaffectanswergeneration.WereleaseLAMAR,alanguageawaremultilingualcrossencodertrainedtoaccountforbothsemanticrelevanceandlanguagecoherence.LAMARfirstusesEnglishanchoredrelevancedistillationtoestablishconsistentrelevancescoringacrossmultilingualinputsandthenappliespreferencealignmentforlanguagecoherencetoencouragedocumentswritteninthesamelanguageasthequerytoreceivehigherrankingswhileretainingsemanticrelevance.Inacontrolledexperimentdesignedtoassesslanguagecoherence,LAMARachievesthebestperformanceoverallandacrossalllanguagesexaminedindividually.LAMARalsoremainscompetitiveonestablishedmultilingualrerankingbenchmarks.Inpracticalretrievalsettings,LAMARachievesthebestresultsacrossallreportedmetricswhenrerankingcandidatesretrievedinthefirststage.TheseresultsdemonstratethatLAMARaccountsforlanguagecoherencewhileachievingstrongperformanceongeneralmultilingualrerankingbenchmarks.

View arXiv pageView PDFProject pageAdd to collection

Get this paper in your agent:

hf papers read 2607\.22042

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper1

#### nlpai-lab/LAMAR-600m Text Ranking• 0.6B• Updatedabout 1 hour ago • 127 • 8

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2607.22042 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2607.22042 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Language Chain in Alignment: Cross-lingual Ranking Preference Optimization

Hugging Face Daily Papers

Cross-lingual Ranking Preference Optimization (CRPO) is a novel framework that enhances multilingual LLM alignment by transferring English preference knowledge to target languages through hierarchical ranking optimization, demonstrating improved performance in instruction-following and knowledge utilization across multiple languages.