LAMAR: An Open Language-Aware Multilingual Alignment Reranker
Summary
LAMAR is a language-aware multilingual cross-encoder reranker that uses English-anchored relevance distillation and preference alignment to prioritize documents in the same language as the query while preserving semantic relevance, achieving strong performance on multilingual benchmarks.
View Cached Full Text
Cached at: 07/27/26, 05:40 AM
Paper page - LAMAR: An Open Language-Aware Multilingual Alignment Reranker
Source: https://huggingface.co/papers/2607.22042
Abstract
Inmultilingualretrievalaugmentedgeneration,aretrievercanretrieverelevantdocumentswritteninmultiplelanguages,whicharesubsequentlyrerankedbeforeanswergeneration.However,itremainsunclearwhetherexistingmultilingualrerankersconsiderdocumentlanguagewhenorderingsemanticallyrelevantcandidates.Ouranalysisshowsthatthesererankersdonotconsistentlyprioritizedocumentswritteninthesamelanguageasthequerywhensemanticallyequivalentdocumentsareavailableacrosslanguages,eventhoughdocumentlanguagecanaffectanswergeneration.WereleaseLAMAR,alanguageawaremultilingualcrossencodertrainedtoaccountforbothsemanticrelevanceandlanguagecoherence.LAMARfirstusesEnglishanchoredrelevancedistillationtoestablishconsistentrelevancescoringacrossmultilingualinputsandthenappliespreferencealignmentforlanguagecoherencetoencouragedocumentswritteninthesamelanguageasthequerytoreceivehigherrankingswhileretainingsemanticrelevance.Inacontrolledexperimentdesignedtoassesslanguagecoherence,LAMARachievesthebestperformanceoverallandacrossalllanguagesexaminedindividually.LAMARalsoremainscompetitiveonestablishedmultilingualrerankingbenchmarks.Inpracticalretrievalsettings,LAMARachievesthebestresultsacrossallreportedmetricswhenrerankingcandidatesretrievedinthefirststage.TheseresultsdemonstratethatLAMARaccountsforlanguagecoherencewhileachievingstrongperformanceongeneralmultilingualrerankingbenchmarks.
View arXiv pageView PDFProject pageAdd to collection
Get this paper in your agent:
hf papers read 2607\.22042
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper1
#### nlpai-lab/LAMAR-600m Text Ranking• 0.6B• Updatedabout 1 hour ago • 127 • 8
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2607.22042 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2607.22042 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
All Languages Matter: Understanding and Mitigating Language Bias in Multilingual RAG
Researchers identify systematic English and query-language bias in multilingual RAG rerankers and introduce LAURA, a utility-driven alignment method that boosts performance by retrieving answer-critical documents across languages.
KaLM-Reranker-V1: Fast but Not Late Interaction for Compressed Document Reranking
KaLM-Reranker-V1 is a fast reranker that decouples query and passage computation using an encoder-decoder architecture with Matryoshka embedding pooling and cross-attention, achieving state-of-the-art reranking performance on BEIR and competitive results on multilingual benchmarks.
Language Chain in Alignment: Cross-lingual Ranking Preference Optimization
Cross-lingual Ranking Preference Optimization (CRPO) is a novel framework that enhances multilingual LLM alignment by transferring English preference knowledge to target languages through hierarchical ranking optimization, demonstrating improved performance in instruction-following and knowledge utilization across multiple languages.
One Domain, Many Tongues: Composing Domain and Language LoRAs for Cross-Lingual Remote-Sensing MLLMs without Paired Data
This paper introduces Modl, a technique for creating cross-lingual remote-sensing multimodal large language models by composing domain and language LoRAs with mutual orthogonality, achieving superior performance without paired multilingual data.
Transforming LLMs into Efficient Cross-Encoders via Knowledge Distillation for RAG Reranking
This paper presents a method to fine-tune LLaMA 3 8B as an efficient reranker for Retrieval-Augmented Generation using knowledge distillation and 4-bit quantization, achieving 14-21% gains in retrieval metrics over cross-encoder baselines with reduced inference cost.