L\"etzCross: A Cross-Lingual Page-Level Benchmark for Multimodal Retrieval over Luxembourgish Documents

arXiv cs.CL Papers

Summary

LëtzCross is a benchmark for cross-lingual page-level retrieval over Luxembourgish PDF documents, comparing text-only and multimodal retrievers in low-resource settings.

arXiv:2608.21714v1 Announce Type: new Abstract: Recent page-image retrievers such as ColPali have improved retrieval over visually rich documents, yet little is known about how they behave in cross-lingual, low-resource settings. We introduce L\"etzCross, a benchmark for cross-lingual page-level retrieval over Luxembourgish PDF documents, with document pages indexed as images and queries provided in English, French, German, and Luxembourgish. The benchmark combines text-focused QA pairs with visually grounded QA pairs, covering both textual and visual retrieval needs in PDF-based RAG. We use L\"etzCross to compare OCR-based text-only retrievers with ColPali-style page-image retrievers and find that the latter perform better across query languages in this system-level comparison. We also examine single-language and multilingual fine-tuning. Fine-tuning transfers across query languages, with French yielding the highest mean performance on Luxembourgish queries among the single-language settings. In the multilingual setting, including Luxembourgish gives the strongest results and substantially improves retrieval for Luxembourgish queries.
Original Article
View Cached Full Text

Cached at: 08/25/26, 04:15 AM

# LëtzCross: A Cross-Lingual Page-Level Benchmark for Multimodal Retrieval over Luxembourgish Documents
Source: [https://arxiv.org/html/2608.21714](https://arxiv.org/html/2608.21714)
Omar El BachyrFred PhilippyAffiliation:University of Luxembourg, LuxembourgLaura Maria BernardyAffiliation:University of Luxembourg, Luxembourg\[0\.5em\]Saad EzziniAffiliation:King Fahd University of Petroleum and Minerals, Saudi ArabiaJacques KleinAffiliation:University of Luxembourg, LuxembourgTegawendé F\. BissyandéAffiliation:University of Luxembourg, Luxembourg

###### Abstract

Recent page\-image retrievers such as ColPali\([6](https://arxiv.org/html/2608.21714#bib.bib13)\)have improved retrieval over visually rich documents, yet little is known about how they behave in cross\-lingual, low\-resource settings\. We introduceLëtzCross, a benchmark for cross\-lingual page\-level retrieval over Luxembourgish PDF documents, with document pages indexed as images and queries provided in English, French, German, and Luxembourgish\. The benchmark combines text\-focused QA pairs with visually grounded QA pairs, covering both textual and visual retrieval needs in PDF\-based RAG\. We useLëtzCrossto compare OCR\-based text\-only retrievers with ColPali\-style page\-image retrievers and find that the latter perform better across query languages in this system\-level comparison\. We also examine single\-language and multilingual fine\-tuning\. Fine\-tuning transfers across query languages, with French yielding the highest mean performance on Luxembourgish queries among the single\-language settings\. In the multilingual setting, including Luxembourgish gives the strongest results and substantially improves retrieval for Luxembourgish queries\.

## 1Introduction

Retrieval\-augmented generation \(RAG\) and document retrieval systems increasingly operate in multilingual environments, where user queries and source documents may be written in different languages\. In this setting, cross\-lingual information retrieval \(CLIR\) has advanced substantially with dense retrieval architectures and multilingual encoders\([12](https://arxiv.org/html/2608.21714#bib.bib9);[4](https://arxiv.org/html/2608.21714#bib.bib10)\)\. However, most current evaluations focus on text\-centric, high\-resource scenarios and do not reflect the challenges of visually rich PDF documents\.

Recent vision–language retrieval models, such as DSE\([17](https://arxiv.org/html/2608.21714#bib.bib14)\)and ColPali\([6](https://arxiv.org/html/2608.21714#bib.bib13)\), improve retrieval over document pages by jointly modeling textual and visual evidence\. Yet their behavior in low\-resource language settings remains underexplored\. We study this challenge through the case of Luxembourgish, a low\-resource language for which NLP resources are growing but still limited compared to high\-resource languages\([16](https://arxiv.org/html/2608.21714#bib.bib18);[2](https://arxiv.org/html/2608.21714#bib.bib17);[21](https://arxiv.org/html/2608.21714#bib.bib22)\)\.

To address this gap, we introduceLëtzCross, a benchmark for cross\-lingual page\-level retrieval over Luxembourgish PDF pages, where pages are represented as images\. The benchmark combines text\-focused and visually grounded QA pairs, and is designed around realistic page\-level retrieval needs in RAG pipelines, where answers may depend on textual content, visual structure, or both, and where queries may be issued in multiple languages\.

Our main contributions are:

- •We introduce LëtzCross, a benchmark for cross\-lingual page\-level retrieval over Luxembourgish PDF pages, where pages are represented as images\.
- •We benchmark multilingual text\-only retrievers and late\-interaction page\-image retrievers on this task\.
- •We study query\-language\-specific fine\-tuning for cross\-lingual retrieval over Luxembourgish document pages\.
- •We analyze multilingual fine\-tuning settings for retrieval over Luxembourgish document pages as a low\-resource language use case\.

Overall, our findings show that late\-interaction page\-image retrievers are strong candidates for cross\-lingual retrieval over Luxembourgish document pages, outperforming the evaluated OCR\-based text\-only baselines in a system\-level comparison\. We further show that query\-language\-specific fine\-tuning transfers across languages, while multilingual fine\-tuning is most effective when it includes the target low\-resource language\. Together, these results positionLëtzCrossas a useful benchmark for studying cross\-lingual page\-level retrieval over Luxembourgish PDF documents as a low\-resource language in RAG settings\.

We release LëtzCross and all experimental code and resources through our[GitHub repository](https://github.com/OmarElbachyr/letzcross-benchmark)\.

## 2Related Work

#### Luxembourgish NLP\-Resources

The Luxembourgish NLP landscape is still under development\. Natural Language Processing techniques for Luxembourgish are primarily developed within a low\-resource setting and rely heavily on manual efforts, such as data collection and annotation\. In recent years, this approach has enabled the development of several foundational resources and applications, including sentiment analysis systems\([24](https://arxiv.org/html/2608.21714#bib.bib19);[7](https://arxiv.org/html/2608.21714#bib.bib20)\), a Luxembourgish dependency parsing treebank\([20](https://arxiv.org/html/2608.21714#bib.bib21)\), and zero\-shot topic classification\([19](https://arxiv.org/html/2608.21714#bib.bib26)\)\. These initiatives have contributed to a better understanding of the linguistic properties of Luxembourgish through computational methods\. Within this context, Lothritz et al\. introduced LuxemBERT\([16](https://arxiv.org/html/2608.21714#bib.bib18)\), a language model trained using artificial data augmentation strategies\. In addition to the model itself, they expanded Luxembourgish resources for several downstream tasks, including part\-of\-speech tagging, named entity recognition, and news classification\. Two further language models have since been established for Luxembourgish: LuXGPT\([2](https://arxiv.org/html/2608.21714#bib.bib17)\), developed using transfer learning techniques, and, more recently, LUXT5\([21](https://arxiv.org/html/2608.21714#bib.bib22)\), which benefits from multilingual pretraining\. Beyond language modeling, practical NLP applications have also emerged\. These include LUX\-ASR\([8](https://arxiv.org/html/2608.21714#bib.bib25)\)for automatic speech recognition, as well as systems for automatic comment moderation\([23](https://arxiv.org/html/2608.21714#bib.bib23)\)and orthographic correction\([22](https://arxiv.org/html/2608.21714#bib.bib24)\)\.

#### Cross\-lingual retrieval

Transformer\-based Cross\-Lingual Information Retrieval \(CLIR\) builds on multilingual representation learning and dense retrieval architectures that enable semantic matching across languages\. Early foundations include BERT\([5](https://arxiv.org/html/2608.21714#bib.bib11)\), which introduced contextual transformer representations for retrieval models, and XLM\-R\([4](https://arxiv.org/html/2608.21714#bib.bib10)\), a multilingual transformer that provides strong cross\-lingual embeddings\. Dense Passage Retrieval \(DPR\)\([12](https://arxiv.org/html/2608.21714#bib.bib9)\)later established the dual\-encoder architecture widely used in dense retrieval systems\. Building on these foundations, more recent work explores unsupervised and generative approaches, such as Unsupervised Multilingual Dense Retrieval \(UMR\) proposed by\([11](https://arxiv.org/html/2608.21714#bib.bib8)\), which uses generative pseudo\-labeling to train dense retrievers without parallel supervision, and research on adapting LLMs for retrieval tasks\([10](https://arxiv.org/html/2608.21714#bib.bib7)\)\. Recent multilingual embedding models such as BGE\-M3\([3](https://arxiv.org/html/2608.21714#bib.bib5)\), Qwen3 Embedding\([26](https://arxiv.org/html/2608.21714#bib.bib4)\), and voyage\-4\-nano\([1](https://arxiv.org/html/2608.21714#bib.bib16)\)further improve cross\-lingual semantic representations and are widely used as baselines in modern CLIR systems\.

#### Vision Language Retrieval

Text\-only encoders have achieved strong performance in information retrieval \(IR\), with models such as\([3](https://arxiv.org/html/2608.21714#bib.bib5);[26](https://arxiv.org/html/2608.21714#bib.bib4);[1](https://arxiv.org/html/2608.21714#bib.bib16)\)demonstrating high effectiveness across multilingual and dense retrieval tasks\. However, recent studies in IR have advanced beyond conventional text\-only systems by incorporating visual features\. DSE\([17](https://arxiv.org/html/2608.21714#bib.bib14)\)propose the use of Vision–Language Models \(VLMs\) to encode full document screenshots, successfully preserving both text and visual information without the need for optical character recognition \(OCR\)\. Similarly, ColPali\([6](https://arxiv.org/html/2608.21714#bib.bib13)\)addresses the shortcomings of text\-focused retrieval in the context of visually rich documents by adapting ColBERT\([13](https://arxiv.org/html/2608.21714#bib.bib12)\)late\-interaction similarity scoring mechanism to work with page images\. This produces multi\-vector visual embeddings that enable fine\-grained matching between query tokens and specific visual areas on document pages\.

![Refer to caption](https://arxiv.org/html/2608.21714v1/m-luxRAG_datasets.png)

Figure 1:Overview of the LëtzCross benchmark construction pipeline\.

## 3LëtzCross Benchmark Construction

LëtzCrossis a cross\-lingual page\-level retrieval benchmark for Question\-Answering \(QA\), built from Luxembourgish documents, with queries in English \(EN\), French \(FR\), German \(DE\), and Luxembourgish \(LB\)\. Documents are represented as page images, and answers are grounded in the corresponding pages\. Figure[1](https://arxiv.org/html/2608.21714#S2.F1)summarizes the process, with steps described below\.

### 3\.1PDF Documents Acquisition

We obtain the original PDF documents by iterating over the Luxembourgish \(LB\) split of FinePDFs\([14](https://arxiv.org/html/2608.21714#bib.bib15)\)dataset and retrieving each document from its source\. When documents are archived in Common Crawl, we extract the corresponding PDF directly from the web archive; otherwise, we download the PDF from the original URL provided in the dataset\. This process allows us to recover the original PDF files associated with the dataset entries\.

### 3\.2Filtering and Curation of PDF Documents

After downloading the PDFs, we perform an automatic validation step to remove corrupted or unreadable documents by verifying that each PDF can be successfully parsed and processed\. We then manually classify the validated PDFs into visually rich and text\-dominant subsets\. A document is considered visually rich when it contains informative visual elements, such as charts, diagrams, or tables, that convey substantive information beyond the surrounding text\. Non\-informative visuals, such as portraits, landscapes, and logos, are disregarded\. Documents whose information is conveyed primarily through continuous text, with few or no informative visual elements, are classified as text\-dominant\. The resulting corpus serves as the foundation for the subsequent Question–Answer generation step\.

### 3\.3Question–Answer Generation

We then use the validated PDFs from the previous step to generate Question–Answer \(QA\) pairs\. We generate queries in English to ensure high\-quality and consistent question formulation, as this yields more reliable QA pairs on Luxembourgish pages\. The benchmark combines two complementary QA sources: manually authored visually grounded QA pairs for visually rich pages, and automatically generated text\-focused QA pairs for text\-dominant pages\.

#### Manual visually grounded QA authoring\.

For the subset of PDFs containing rich visual components, a native Luxembourgish speaker manually crafts page\-level QA pairs that intentionally target visually grounded evidence, such as charts, plots, diagrams, figures, and other layout\-dependent elements, rather than relying solely on textual content\. Questions are phrased to be self\-contained and answerable from a single page image, with concise answers grounded in the visual content\.

#### Automatic text\-focused QA generation\.

Using the text\-dominant PDF subset, we automatically generate page\-level QA pairs with GPT\-5\-mini\. Each page is processed independently as an image and prompted to produce self\-contained questions with concise answers grounded in its textual content\. To validate the generated QA pairs, each query and its corresponding PDF are provided to Gemini\-2\.5\-Flash, which extracts an answer from the document\. We then compare the extracted answer with thegold answerusing both exact and semantic matching; the latter is performed by prompting GPT\-5\-mini to assess whether the two answers convey the same meaning\. Overall, 90\.28% of the QA pairs were judged semantically equivalent by GPT\-5\-mini, and 71\.50% contained the exact literal gold answer\. QA pairs that do not yield consistent answers are discarded, resulting in a consistency\-filtered set for the subsequent manual validation step\.

### 3\.4Quality Control and RAG Alignment

As a follow\-up step, the manually authored QA pairs are included directly in the final benchmark\. In contrast, the automatically generated QA pairs undergo a manual validation phase focused exclusively on question formulation: annotators are shown only the questions and revise them when needed to produce clear, natural queries that reflect RAG\-relevant information needs \(e\.g\., realistic, retrieval\-oriented queries\)\.

### 3\.5Cross\-Lingual Query Translation

To enable cross\-lingual retrieval evaluation, we translate the finalized English queries into German, French, and Luxembourgish using a controlled LLM prompt that enforces semantic preservation and prevents paraphrasing or additional explanations\. This produces aligned query sets across languages, where each translated query is intended to preserve the same retrieval intent as the original English query\. We use Gemini\-3\-Flash for German and French, and Gemini\-3\-Pro for Luxembourgish\. This choice was made after manually inspecting a small set of trial translations: Gemini\-3\-Pro produced better Luxembourgish translations, while Gemini\-3\-Flash was sufficient for German and French\.

The final benchmark contains 579 QA pairs over 908 document pages\. Table[1](https://arxiv.org/html/2608.21714#S3.T1)summarizes the benchmark statistics, including the distribution between automatically generated text\-focused QA pairs and manually authored visually grounded QA pairs\. All prompts used during benchmark construction are available in the accompanying repository\.

Benchmark statisticValueNumber of pages908Number of queries579Avg\. query length19\.5Avg\. corpus item length722\.5QA type distributionAuto\-generated text\-focused530Manual visually grounded49Table 1:Overview statistics of theLëtzCrossevaluation benchmark\. Query and corpus item lengths are reported in tokens\.

## 4Methods

This section presents our method for investigating cross\-lingual retrieval over Luxembourgish document pages as a low\-resource case study using late\-interaction page\-image retrievers\.

### 4\.1Training Data Construction

We construct the training data by scraping Luxembourgish Wikipedia articles and converting them into PDF documents\. We first apply size\-based filtering with minimum and maximum page thresholds to exclude trivial or overly long documents \(e\.g\., pages containing only footnotes\)\. QA pairs are then generated from 4,977 PDFs with 11,833 pages using the same automatic text\-focused QA generation procedure as in the benchmark \(Section[3\.3](https://arxiv.org/html/2608.21714#S3.SS3.SSS0.Px2)\)\. This resulted in 22,028 QA pairs\. In addition, we construct 450 QA pairs from the same Wikipedia\-derived pipeline and reserve them as a held\-out development split for fine\-tuning\. To extend the training set for cross\-lingual retrieval, we translate queries into German and French, and a 5,000\-query subset into Luxembourgish, following the same query translation procedure as in the benchmark \(Section[3\.5](https://arxiv.org/html/2608.21714#S3.SS5)\)\.

### 4\.2Retriever and Training Objective

We fine\-tune a late\-interaction page\-image retriever based on ColPali\([6](https://arxiv.org/html/2608.21714#bib.bib13)\)using the training setup provided by the underlying framework\. Each document page is treated as an image and split into visual patches, which are encoded into contextualized image patch embeddings, while the query text is tokenized and encoded into query token embeddings\. Both are represented in a shared latent space, and relevance is computed through late interaction between query tokens and page patches\.

Letqi=\{qi​1,…,qi​\|qi\|\}q\_\{i\}=\\\{q\_\{i1\},\\dots,q\_\{i\|q\_\{i\}\|\}\\\}denote the query token embeddings of queryii, and letdj=\{dj​1,…,dj​\|dj\|\}d\_\{j\}=\\\{d\_\{j1\},\\dots,d\_\{j\|d\_\{j\}\|\}\\\}denote the document image patch embeddings of pagejj\. The score between queryiiand pagejjis:

score​\(qi,dj\)=∑t=1\|qi\|maxu=1,…,\|dj\|⁡\(qi​t⊤​dj​u\)\\text\{score\}\(q\_\{i\},d\_\{j\}\)=\\sum\_\{t=1\}^\{\|q\_\{i\}\|\}\\max\_\{u=1,\\dots,\|d\_\{j\}\|\}\\left\(q\_\{it\}^\{\\top\}d\_\{ju\}\\right\)\(1\)
We use the original ColBERT\([13](https://arxiv.org/html/2608.21714#bib.bib12)\)contrastive objective with in\-batch negatives:

ℒ=−1B∑i=1Blogexp⁡\(score​\(qi,di\)/τ\)∑j=1Bexp⁡\(score​\(qi,dj\)/τ\)\\mathcal\{L\}=\-\\frac\{1\}\{B\}\\sum\_\{i=1\}^\{B\}\\log\\frac\{\\exp\\left\(\\text\{score\}\(q\_\{i\},d\_\{i\}\)/\\tau\\right\)\}\{\\sum\_\{j=1\}^\{B\}\\exp\\left\(\\text\{score\}\(q\_\{i\},d\_\{j\}\)/\\tau\\right\)\}\(2\)
whereτ\\tauis a temperature hyperparameter\. Each training instance consists of a query paired with its corresponding document page image as the positive example, while the remainingB−1B\-1pages in the batch serve as in\-batch negatives\. We do not use external hard negatives or additional negative sampling\.

### 4\.3Fine\-Tuning Configurations

We consider two fine\-tuning settings\. In theper\-languagesetting, the retriever is fine\-tuned separately on English, French, and German queries\. In themultilingualsetting, a single model is trained on mixed\-language data, using either EN\+FR\+DE or EN\+FR\+DE\+LB\. In all cases, document pages remain in Luxembourgish, and only the query language varies\. To analyze visual adaptation, we also compare the default setting, which updates both textual and visual components, with a text\-component\-only variant\. The EN, FR, and DE settings each use 22,028/450 train/dev QA pairs, while EN\+FR\+DE and EN\+FR\+DE\+LB use 21,000/1,350 and 20,000/1,800 train/dev samples, respectively\.

RetrieverQuery LanguageAvg\.Index \(ms/page\)↓\\downarrowLatency \(ms/query\)↓\\downarrowENFRDELBtext\-only retrieversbge\-m372\.4871\.6372\.4776\.1273\.1821\.714±0\.01521\.714\\pm 0\.0153\.859±0\.0273\.859\\pm 0\.027Qwen3\-Embedding\-0\.6B73\.0271\.6375\.2167\.9771\.9642\.863±0\.11542\.863\\pm 0\.1156\.935±0\.0086\.935\\pm 0\.008multilingual\-e5\-large75\.9675\.0277\.2468\.3274\.1411\.068±0\.07011\.068\\pm 0\.0703\.858±0\.0123\.858\\pm 0\.012jina\-embeddings\-v476\.5175\.3977\.3073\.8675\.76113\.007±0\.287113\.007\\pm 0\.28726\.271±0\.33626\.271\\pm 0\.336multi\-vector page\-image retrieverscolSmol\-256M71\.6159\.4552\.7655\.1159\.73367\.108±1\.415367\.108\\pm 1\.4153\.914±0\.0353\.914\\pm 0\.035colSmol\-500M73\.9067\.5665\.9263\.8767\.81361\.103±11\.169361\.103\\pm 11\.1694\.125±0\.0174\.125\\pm 0\.017colpali\-v1\.377\.0377\.4378\.4675\.7877\.1861\.801±0\.11061\.801\\pm 0\.1102\.997±0\.0072\.997\\pm 0\.007colqwen2\.5\-v0\.278\.0778\.7179\.9278\.6078\.83133\.876±0\.813133\.876\\pm 0\.8134\.892±0\.1134\.892\\pm 0\.113colnomic\-embed\-multimodal\-3b78\.7378\.4479\.5977\.8978\.66145\.193±0\.277145\.193\\pm 0\.2774\.981±0\.0674\.981\\pm 0\.067

Table 2:Cross\-lingual retrieval over Luxembourgish document pages\. Query\-language columns report nDCG@10 where best results are shown inbold, and second\-best results areunderlined\. Efficiency values are mean±\\pmstandard deviation over three runs on the same hardware\.

## 5Experimental Setup

#### Models\.

Based on the input type, we consider two families of retrievers in our experiments: text\-only retrievers and late\-interaction page\-image retrievers\. The text\-only baselines are bge\-m3\([3](https://arxiv.org/html/2608.21714#bib.bib5)\), Qwen3\-Embedding\-0\.6B\([26](https://arxiv.org/html/2608.21714#bib.bib4)\), multilingual\-e5\-large\([25](https://arxiv.org/html/2608.21714#bib.bib2)\), and jina\-embeddings\-v4\([9](https://arxiv.org/html/2608.21714#bib.bib3)\)111Althoughjina\-embeddings\-v4is a multimodal embedding model, we use it here in a text\-only setting\.\. The late\-interaction page\-image retrievers are colSmol\-256M, colSmol\-500M, colpali\-v1\.3, and colqwen2\.5\-v0\.2\([6](https://arxiv.org/html/2608.21714#bib.bib13)\), and colnomic\-embed\-multimodal\-3b\([18](https://arxiv.org/html/2608.21714#bib.bib1)\)\.

#### Text Extraction Pipeline\.

The text\-only retrievers described above use page\-level text extracted withunstructured\[pdf\]v0\.18\.15 under thehi\_resstrategy, combining Poppler page rendering, YOLOX layout detection, Luxembourgish Tesseract OCR \(ltz\), and table\-structure extraction\. Existing PDF text was incorporated when available, and all PDFs were processed using the same pipeline\. We consider this a practical system\-level comparison between text\-based and page\-image retrieval, rather than an evaluation of PDF extraction methods\.

#### Fine\-Tuning Details\.

We fine\-tune the retriever using the ColBERT late\-interaction contrastive loss defined in Equation \([2](https://arxiv.org/html/2608.21714#S4.E2)\) with temperatureτ=0\.02\\tau=0\.02and LoRA adapters \(r=32,alpha=32, dropout=0\.1=0\.1\)\. In the text\-component\-only setting, LoRA targets the query encoder’s attention projections \(q\_proj,k\_proj,v\_proj, ando\_proj\), feed\-forward projections \(down\_proj,gate\_proj, andup\_proj\), andcustom\_text\_proj; the vision encoder remains frozen\. In the text\+visual setting, the same attention and feed\-forward projections are adapted in both the textual and visual transformer blocks, where the visual component denotes the blocks encoding page\-image patches\. All untargeted parameters remain frozen\. Training is performed for 3 epochs with a learning rate of5×10−55\\times 10^\{\-5\}and 100 warmup steps, using a per\-device batch size of 16 on 4 NVIDIA L40S GPUs and no gradient accumulation, resulting in an effective batch size of 64\. All runs use bfloat16 precision\. Each fine\-tuning configuration is repeated with three seeds \(42, 43, and 44\), and results are reported as mean±\\pmstandard deviation\.

#### Evaluation Setup and Metrics\.

We evaluate retrieval at the page level: for each query, the retriever ranks all document pages in the benchmark against the gold relevant page\. Results are reported separately for English, French, German, and Luxembourgish queries\. Our main metric is nDCG@10; for analyses across retrieval depths, we additionally report nDCG@kkfork∈\{1,3,5,10\}k\\in\\\{1,3,5,10\\\}\. Indexing time per page and query latency per query are aggregated across the four query languages and reported as mean±\\pmstandard deviation over three runs\. All evaluation runs, including effectiveness and efficiency measurements, use one NVIDIA L40S GPU and eight CPU cores\.

## 6Results and Analysis

Our experiments evaluate cross\-lingual retrieval over Luxembourgish document pages\. We first compare text\-only and page\-image retrievers, then analyze query\-language\-specific fine\-tuning, including an ablation on the visual component, and finally evaluate multilingual fine\-tuning\.

### 6\.1Page\-Image vs\. OCR\-Based Text Retrieval over Luxembourgish Documents

We compare multilingual OCR\-based text\-only embedding baselines, which index page\-level text produced by the OCR, layout\-analysis, and table\-extraction pipeline described in Section[5](https://arxiv.org/html/2608.21714#S5.SS0.SSS0.Px2), with late\-interaction page\-image retrievers ranging from lightweight models below 500M parameters to models of up to 3B parameters\. Table[2](https://arxiv.org/html/2608.21714#S4.T2)shows that the strongest page\-image retrievers outperform the evaluated text\-only baselines across all query languages\. Since the two families differ in input representation, extraction pipeline, and model architecture, we interpret this as a practical system\-level comparison rather than a controlled modality ablation\.

Performance remains strongly model\-dependent\. The smaller colSmol models underperform strong multilingual text\-only retrievers such as jina\-embeddings\-v4\. In contrast, colqwen2\.5\-v0\.2 and colnomic\-embed\-multimodal\-3b outperform all evaluated text\-only models across every query language, while colpali\-v1\.3 also performs strongly on English, French, and German\. This shows that sufficiently capable multi\-vector page\-image retrievers can outperform multilingual text embedding models in cross\-lingual retrieval\. The best page\-image retriever achieves an average nDCG@10 of 78\.83, improving by 3\.07 points over the strongest text\-only baseline\. A more customized OCR and layout\-detection pipeline could nevertheless affect this gap\.

These results are notable because the text\-only baselines are explicitly designed for multilingual retrieval, whereas the page\-image retrievers primarily target visual\-document retrieval\. Their strong zero\-shot cross\-lingual performance may partly reflect the multilingual VLM backbones used by several models\.

Training setupQuery languageENFRDELBAvg\.Baseline78\.7378\.4479\.5977\.8978\.66FT \(EN\)81\.66±\\pm0\.33\(\+2\.93\)81\.34±\\pm0\.40\(\+2\.89\)81\.84±\\pm0\.11\(\+2\.24\)78\.48±\\pm0\.65\(\+0\.60\)80\.83±\\pm0\.23\(\+2\.17\)FT \(FR\)81\.11±\\pm0\.57\(\+2\.38\)81\.50±\\pm0\.49\(\+3\.06\)81\.76±\\pm0\.40\(\+2\.17\)79\.45±\\pm0\.75\(\+1\.57\)80\.96±\\pm0\.55\(\+2\.29\)FT \(DE\)79\.48±\\pm0\.27\(\+0\.75\)80\.51±\\pm0\.88\(\+2\.07\)81\.56±\\pm0\.34\(\+1\.97\)79\.34±\\pm0\.33\(\+1\.45\)80\.22±\\pm0\.14\(\+1\.56\)Avg\. FT80\.75±\\pm0\.1581\.12±\\pm0\.2281\.72±\\pm0\.0879\.09±\\pm0\.3480\.67±\\pm0\.11

Table 3:nDCG@10 by query language \(columns\) for the baseline model and models fine\-tuned on different training languages \(rows\)\. Values are reported as mean±\\pmstandard deviation over three runs; values in parentheses indicate gains over the baseline on the same query language\. Best values are shown inbold\.Table[4](https://arxiv.org/html/2608.21714#S6.T4)further examines performance by QA source\. The page\-image retrievers outperform the strongest text\-only baseline on both the text\-focused and visually grounded subsets, with larger gains on the visually grounded questions\. This provides additional evidence that page\-image retrieval is beneficial beyond the predominantly text\-focused portion of the benchmark\. However, the visually grounded subset contains only 49 QA pairs and yields higher scores for all retrievers, so these results remain descriptive and should be interpreted cautiously\.

RetrieverText\-focusedVisually groundedOveralljina\-embeddings\-v475\.2781\.1075\.76colqwen2\.5\-v0\.278\.4283\.2378\.83colnomic\-embed\-multimodal\-3b77\.9186\.8278\.66

Table 4:Descriptive nDCG@10 split by QA source, averaged across query languages\.#### Efficiency\.

Table[2](https://arxiv.org/html/2608.21714#S4.T2)shows clear efficiency differences across retrievers\. Among text\-only models, multilingual\-e5\-large indexes fastest, while jina\-embeddings\-v4 is slower, likely due to its larger multimodal backbone\. Among page\-image retrievers, colpali\-v1\.3 is the fastest, possibly because of its fixed\-resolution processing and fixed visual\-token budget\. The slower ColSmol indexing further shows that runtime depends on visual\-token count and implementation choices, not only parameter count\. Query latency remains similar across most page\-image models because page representations are precomputed\.

For the text\-only retrievers, page parsing introduces an additional shared preprocessing cost\. Using the extraction pipeline described in Section[5](https://arxiv.org/html/2608.21714#S5.SS0.SSS0.Px2), parsing takes5\.382±0\.0545\.382\\pm 0\.054s/page across three runs\. Including this cost, the end\-to\-end preprocessing and indexing time ranges from5\.393±0\.0545\.393\\pm 0\.054s/page for multilingual\-e5\-large to5\.495±0\.0545\.495\\pm 0\.054s/page for jina\-embeddings\-v4\. Thus, for text\-only retrieval, document parsing dominates the overall offline preprocessing cost\.

### 6\.2Cross\-Lingual Transfer of Query\-Language\-Specific Fine\-Tuning

This subsection studies how fine\-tuning transfers across evaluation languages for the late\-interaction page\-image retrievers considered in this work\. We fine\-tune the retriever separately on English, French, and German queries, and evaluate each variant on all four query languages\. We use colnomic\-embed\-multimodal\-3b for this analysis, as both the baseline model and the backbone for fine\-tuning, because its performance remains very close to the top model in our experiments, while ViDoRe leaderboard\([15](https://arxiv.org/html/2608.21714#bib.bib6)\)results suggest it is a strong page\-image retriever overall\.

Table[3](https://arxiv.org/html/2608.21714#S6.T3)shows that all language\-specific fine\-tuning settings improve over the baseline on every query language\. FT \(FR\) achieves the highest overall average at 80\.96, closely followed by FT \(EN\) at 80\.83, while FT \(DE\) reaches 80\.22\. The largest language\-specific gains are obtained by FT \(EN\) on English \(\+2\.93\) and FT \(FR\) on French \(\+3\.06\)\. Cross\-lingual transfer is also substantial: FT \(EN\) achieves the highest German score at 81\.84 \(\+2\.24\), while FT \(FR\) performs best on Luxembourgish at 79\.45 \(\+1\.57\)\.

These results show that the strongest configuration does not always correspond to matching the fine\-tuning and evaluation languages\. In particular, English fine\-tuning transfers strongly to German, while French fine\-tuning provides the strongest Luxembourgish performance\. German fine\-tuning improves all four languages but produces the lowest overall average among the three variants\. Overall, query\-language\-specific fine\-tuning consistently improves multilingual retrieval, while the magnitude of transfer depends on the training and evaluation language pair\.

Figure[2](https://arxiv.org/html/2608.21714#S6.F2)shows that these trends are broadly maintained across retrieval cutoffs\. FT \(EN\) remains strongest for English queries, whereas the fine\-tuned variants are much closer for French and German\. For Luxembourgish, FT \(FR\) and FT \(DE\) generally perform better than FT \(EN\)\. The uncertainty bands also overlap in several cases, indicating that small differences between fine\-tuning languages should not be over\-interpreted\. Overall, the improvements persist across different values ofkk, showing that the cross\-lingual benefits of fine\-tuning are not limited to nDCG@10\.

Figure 2:Multilingual retrieval performance \(nDCG@kk\) across query languages\. Shaded bands indicate±1\\pm 1standard deviation across three fine\-tuning runs\.#### Effect of Updating the Visual Component\.

In all previous fine\-tuning experiments, we use a setting in which both the textual and visual components of the page\-image retriever are updated\. To isolate the contribution of visual adaptation, we compare this setting with the corresponding text\-component\-only fine\-tuning variant\. Table[5](https://arxiv.org/html/2608.21714#S6.T5)reports the paired difference in nDCG@10 between the two settings across three runs\.

The mean differences are positive for every training and evaluation language combination, indicating that updating the visual component provides an additional benefit beyond text\-component\-only adaptation\. The magnitude of this gain varies substantially across settings, ranging from 0\.21 points for FT \(EN\) evaluated on Luxembourgish to 2\.53 points for FT \(DE\) evaluated on French\. The larger gains observed for some FT \(DE\) configurations are accompanied by higher run\-to\-run variability, while several EN and FR configurations show smaller but more stable improvements\. Overall, these results suggest that visual adaptation provides complementary gains, although its contribution depends on both the fine\-tuning language and the evaluation language\.

Training setupQuery languageENFRDELBFT \(EN\)1\.69±\\pm0\.571\.49±\\pm0\.531\.41±\\pm0\.120\.21±\\pm0\.38FT \(FR\)2\.27±\\pm0\.161\.32±\\pm0\.341\.78±\\pm0\.441\.27±\\pm0\.98FT \(DE\)2\.49±\\pm1\.252\.53±\\pm1\.461\.34±\\pm1\.070\.82±\\pm0\.48

Table 5:Difference in nDCG@10 between text\+visual and text\-component\-only fine\-tuning across query languages\. Values are reported as mean±\\pmstandard deviation over three runs; positive values indicate gains from updating the visual component\. Best values are shown inbold\.Training setupQuery languageENFRDELBAvg\.Baseline78\.7378\.4479\.5977\.8978\.66FT \(FR\)81\.11±\\pm0\.5781\.50±\\pm0\.4981\.76±\\pm0\.4079\.45±\\pm0\.7580\.96±\\pm0\.55Multilingual FTEN\+FR\+DE80\.47±\\pm0\.39\(\-0\.63\)80\.46±\\pm0\.16\(\-1\.04\)81\.30±\\pm0\.26\(\-0\.46\)77\.48±\\pm0\.48\(\-1\.97\)79\.93±\\pm0\.12\(\-1\.03\)EN\+FR\+DE\+LB81\.03±\\pm0\.51\(\-0\.08\)81\.04±\\pm0\.44\(\-0\.46\)81\.66±\\pm0\.44\(\-0\.10\)80\.02±\\pm0\.56\(\+0\.57\)80\.94±\\pm0\.41\(\-0\.02\)

Table 6:nDCG@10 across query languages for the baseline, the best single\-language fine\-tuning setting, and multilingual fine\-tuning variants\. Results are reported as mean±\\pmstandard deviation over three runs\. Values in parentheses indicate changes relative to FT \(FR\); the best value is shown inboldand the second\-best isunderlined\.

### 6\.3Multilingual Fine\-Tuning for Improved Cross\-Lingual Retrieval

We further investigate whether multilingual fine\-tuning improves cross\-lingual retrieval beyond the single\-language fine\-tuning studied earlier\. We compare the best single\-language setting, FT \(FR\), with two multilingual variants: one fine\-tuned on English, French, and German, and another fine\-tuned on English, French, German, and Luxembourgish\. We again use colnomic\-embed\-multimodal\-3b as both the baseline model and the backbone used for fine\-tuning\.

As shown in Table[6](https://arxiv.org/html/2608.21714#S6.T6), multilingual fine\-tuning on English, French, and German underperforms FT \(FR\) across all query languages, with an average decrease of 1\.03 nDCG points\. The largest decrease occurs on Luxembourgish queries \(−1\.97\-1\.97\)\. This degradation may partly reflect the smaller number of training examples available per language in the multilingual setting compared with single\-language FT \(FR\), which can reduce language\-specific adaptation\. The absence of Luxembourgish supervision may further contribute to the larger drop observed on Luxembourgish queries\.

Adding Luxembourgish substantially changes this pattern\. The EN\+FR\+DE\+LB setting reaches an average nDCG@10 of 80\.94, essentially matching FT \(FR\) at 80\.96\. Relative to EN\+FR\+DE, adding Luxembourgish improves performance across all query languages, with the largest increase on Luxembourgish queries, from 77\.48 to 80\.02 \(\+2\.54\+2\.54points\)\. Compared with FT \(FR\), the multilingual setting remains close on English, French, and German, with differences of−0\.08\-0\.08,−0\.46\-0\.46, and−0\.10\-0\.10, respectively, while improving Luxembourgish retrieval by\+0\.57\+0\.57\.

Since the corpus pages are in Luxembourgish, direct exposure to Luxembourgish queries may help align query representations with the corresponding document pages\. However, the multilingual settings also differ from single\-language fine\-tuning in the amount and distribution of training data per language, so the effect cannot be attributed to Luxembourgish supervision alone\. Overall, the results show that multilingual fine\-tuning including Luxembourgish can recover the performance lost in the EN\+FR\+DE setting, match the strongest single\-language configuration overall, and provide the strongest Luxembourgish retrieval performance\. Given the variability across three runs, the small differences between FT \(FR\) and EN\+FR\+DE\+LB should be interpreted as comparable overall performance rather than a clear advantage for either setting\.

## 7Conclusion

Luxembourgish provides a representative low\-resource use case for studying cross\-lingual retrieval, given the scarcity of dedicated resources and retrieval benchmarks for the language\. In this work, we introducedLëtzCross, a cross\-lingual benchmark for page\-level retrieval over Luxembourgish PDF documents in a page\-image RAG setting\. The benchmark combines automatically generated text\-focused QA pairs with manually authored visually grounded QA pairs\. Our results show that late\-interaction page\-image retrievers outperform the evaluated OCR\-based text\-only baselines in a system\-level comparison, and that fine\-tuning improves retrieval with gains that transfer across query languages\. Among the evaluated single\-language settings, French fine\-tuning yields the highest mean performance on Luxembourgish queries, while among the multilingual configurations, including Luxembourgish produces the strongest results and substantially improves Luxembourgish\-query retrieval\. Together, these findings provide a first benchmark and empirical study for cross\-lingual page\-level retrieval over Luxembourgish document pages\.

## Limitations

This work has several limitations\. First,LëtzCrossis relatively small, with a limited manually authored visually grounded subset compared with the automatically generated text\-focused QA pairs\. Consequently, the evaluation is dominated by text\-focused queries, and results on the visually grounded subset should be considered complementary rather than representative of the benchmark\. Second, the comparison between text\-only and page\-image retrievers is system\-level: the text\-only baselines rely on document parsing with OCR, layout analysis, and table extraction, while page\-image retrievers operate on rendered pages\. Alternative parsing or hybrid pipelines could therefore affect the observed gap\. Third, all non\-English queries are translated from English rather than independently authored, which may introduce translation artifacts or English\-oriented formulations\. Moreover, query translation and automatic QA generation rely on LLMs, although we apply validation and manual review to improve consistency and quality\. Finally, our findings are specific to Luxembourgish and should be validated on other low\-resource languages before broader generalization\.

## Acknowledgments

This research was funded in whole or in part by the Luxembourg National Research Fund \(FNR\), grant reference NCER22/IS/16570468/NCER\-FT\. We acknowledge Google\.org for providing Gemini credits used in this work\. We also thank the anonymous reviewers for their careful reading and constructive feedback\.

## References

- AI \(2026\)V\. AIVoyage\-4\-nano\.Note:[https://huggingface\.co/voyageai/voyage\-4\-nano](https://huggingface.co/voyageai/voyage-4-nano)State\-of\-the\-art text embedding model with 32k token context lengthCited by:[§2](https://arxiv.org/html/2608.21714#S2.SS0.SSS0.Px2.p1.1),[§2](https://arxiv.org/html/2608.21714#S2.SS0.SSS0.Px3.p1.1)\.
- Bernardy \(2022\)L\. BernardyA Luxembourgish GPT\-2 Approach Based on Transfer Learning\.Master’s Thesis,University of Trier\.Cited by:[§1](https://arxiv.org/html/2608.21714#S1.p2.1),[§2](https://arxiv.org/html/2608.21714#S2.SS0.SSS0.Px1.p1.1)\.
- Chenet al\.\(2024\)J\. Chen, S\. Xiao, P\. Zhang, K\. Luo, D\. Lian, and Z\. LiuBGE m3\-embedding: multi\-lingual, multi\-functionality, multi\-granularity text embeddings through self\-knowledge distillation\.External Links:2402\.03216Cited by:[§2](https://arxiv.org/html/2608.21714#S2.SS0.SSS0.Px2.p1.1),[§2](https://arxiv.org/html/2608.21714#S2.SS0.SSS0.Px3.p1.1),[§5](https://arxiv.org/html/2608.21714#S5.SS0.SSS0.Px1.p1.1)\.
- Conneauet al\.\(2020\)A\. Conneau, K\. Khandelwal, N\. Goyal, V\. Chaudhary, G\. Wenzek, F\. Guzmán, E\. Grave, M\. Ott, L\. Zettlemoyer, and V\. StoyanovUnsupervised cross\-lingual representation learning at scale\.InProceedings of the 58th annual meeting of the association for computational linguistics,pp\. 8440–8451\.Cited by:[§1](https://arxiv.org/html/2608.21714#S1.p1.1),[§2](https://arxiv.org/html/2608.21714#S2.SS0.SSS0.Px2.p1.1)\.
- Devlinet al\.\(2019\)J\. Devlin, M\. Chang, K\. Lee, and K\. ToutanovaBert: pre\-training of deep bidirectional transformers for language understanding\.InProceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 \(long and short papers\),pp\. 4171–4186\.Cited by:[§2](https://arxiv.org/html/2608.21714#S2.SS0.SSS0.Px2.p1.1)\.
- Faysseet al\.\(2025\)M\. Faysse, H\. Sibille, T\. Wu, B\. Omrani, G\. Viaud, C\. Hudelot, and P\. ColomboColpali: efficient document retrieval with vision language models\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 61424–61449\.Cited by:[§1](https://arxiv.org/html/2608.21714#S1.p2.1),[§2](https://arxiv.org/html/2608.21714#S2.SS0.SSS0.Px3.p1.1),[§4\.2](https://arxiv.org/html/2608.21714#S4.SS2.p1.1),[§5](https://arxiv.org/html/2608.21714#S5.SS0.SSS0.Px1.p1.1),[Abstract](https://arxiv.org/html/2608.21714#abstract1.1)\.
- Gierschek \(2022\)D\. GierschekDetection of Sentiment in Luxembourgish User Comments\.Ph\.D\. Thesis,University of Luxembourg\.External Links:[Link](http://hdl.handle.net/10993/50533)Cited by:[§2](https://arxiv.org/html/2608.21714#S2.SS0.SSS0.Px1.p1.1)\.
- Gilleset al\.\(2023\)P\. Gilles, N\. H\. Kivanani, and L\. E\. A\. HillahLUX\-asr: building an asr system for the luxembourgish language\.InProceedings of the 2022 IEEE Spoken Language Technology Workshop \(SLT\),Cited by:[§2](https://arxiv.org/html/2608.21714#S2.SS0.SSS0.Px1.p1.1)\.
- Güntheret al\.\(2025\)M\. Günther, S\. Sturua, M\. K\. Akram, I\. Mohr, A\. Ungureanu, S\. Eslami, S\. Martens, B\. Wang, N\. Wang, and H\. XiaoJina\-embeddings\-v4: universal embeddings for multimodal multilingual retrieval\.External Links:2506\.18902,[Link](https://arxiv.org/abs/2506.18902)Cited by:[§5](https://arxiv.org/html/2608.21714#S5.SS0.SSS0.Px1.p1.1)\.
- Guoet al\.\(2024\)P\. Guo, Y\. Ren, Y\. Hu, Y\. Cao, Y\. Li, and H\. HuangSteering large language models for cross\-lingual information retrieval\.InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval,pp\. 585–596\.Cited by:[§2](https://arxiv.org/html/2608.21714#S2.SS0.SSS0.Px2.p1.1)\.
- Huanget al\.\(2024\)C\. Huang, C\. Li, T\. Hsu, C\. Hsu, and Y\. ChenUnsupervised multilingual dense retrieval via generative pseudo labeling\.InFindings of the Association for Computational Linguistics: EACL 2024,pp\. 736–746\.Cited by:[§2](https://arxiv.org/html/2608.21714#S2.SS0.SSS0.Px2.p1.1)\.
- Karpukhinet al\.\(2020\)V\. Karpukhin, B\. Oguz, S\. Min, P\. Lewis, L\. Wu, S\. Edunov, D\. Chen, and W\. YihDense passage retrieval for open\-domain question answering\.InProceedings of the 2020 conference on empirical methods in natural language processing \(EMNLP\),pp\. 6769–6781\.Cited by:[§1](https://arxiv.org/html/2608.21714#S1.p1.1),[§2](https://arxiv.org/html/2608.21714#S2.SS0.SSS0.Px2.p1.1)\.
- Khattab and Zaharia \(2020\)O\. Khattab and M\. ZahariaColbert: efficient and effective passage search via contextualized late interaction over bert\.InProceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval,pp\. 39–48\.Cited by:[§2](https://arxiv.org/html/2608.21714#S2.SS0.SSS0.Px3.p1.1),[§4\.2](https://arxiv.org/html/2608.21714#S4.SS2.p3.1)\.
- Kydlíčeket al\.\(2025\)H\. Kydlíček, G\. Penedo, and L\. von WerraFinePDFs\.Hugging Face\.Note:[https://huggingface\.co/datasets/HuggingFaceFW/finepdfs](https://huggingface.co/datasets/HuggingFaceFW/finepdfs)Cited by:[§3\.1](https://arxiv.org/html/2608.21714#S3.SS1.p1.1)\.
- Loisonet al\.\(2026\)A\. Loison, Q\. Macé, A\. Edy, V\. Xing, T\. Balough, G\. Moreira, B\. Liu, M\. Faysse, C\. Hudelot, and G\. ViaudViDoRe v3: a comprehensive evaluation of retrieval augmented generation in complex real\-world scenarios\.External Links:2601\.08620,[Link](https://arxiv.org/abs/2601.08620)Cited by:[§6\.2](https://arxiv.org/html/2608.21714#S6.SS2.p1.1)\.
- Lothritzet al\.\(2022\)C\. Lothritz, B\. Lebichot, K\. Allix, L\. Veiber, T\. Bissyande, J\. Klein, A\. Boytsov, C\. Lefebvre, and A\. GoujonLuxemBERT: Simple and Practical Data Augmentation in Language Model Pre\-Training for Luxembourgish\.InProceedings of LREC,External Links:[Link](https://aclanthology.org/2022.lrec-1.543)Cited by:[§1](https://arxiv.org/html/2608.21714#S1.p2.1),[§2](https://arxiv.org/html/2608.21714#S2.SS0.SSS0.Px1.p1.1)\.
- Maet al\.\(2024\)X\. Ma, S\. Lin, M\. Li, W\. Chen, and J\. LinUnifying multimodal retrieval via document screenshot embedding\.arXiv preprint arXiv:2406\.11251\.Cited by:[§1](https://arxiv.org/html/2608.21714#S1.p2.1),[§2](https://arxiv.org/html/2608.21714#S2.SS0.SSS0.Px3.p1.1)\.
- Nomic AI \(2025\)Nomic AIColNomic embed multimodal 3b\.Note:[https://huggingface\.co/nomic\-ai/colnomic\-embed\-multimodal\-3b](https://huggingface.co/nomic-ai/colnomic-embed-multimodal-3b)Hugging Face model cardCited by:[§5](https://arxiv.org/html/2608.21714#S5.SS0.SSS0.Px1.p1.1)\.
- Philippyet al\.\(2024\)F\. Philippy, S\. Haddadan, and S\. GuoForget NLI, Use a Dictionary: Zero\-Shot Topic Classification for Low\-Resource Languages with Application to Luxembourgish\.InProceedings of SIGUL \(LREC\-COLING\),External Links:[Link](https://aclanthology.org/2024.sigul-1.13)Cited by:[§2](https://arxiv.org/html/2608.21714#S2.SS0.SSS0.Px1.p1.1)\.
- Plumet al\.\(2024\)A\. Plum, C\. Döhmer, E\. Milano, A\. Lutgen, and C\. PurschkeLuxBank: The First Universal Dependency Treebank for Luxembourgish\.InProceedings of TLT,D\. Dakota, S\. Jablotschkin, S\. Kübler, and H\. Zinsmeister \(Eds\.\),External Links:[Link](https://aclanthology.org/2024.tlt-1.4/)Cited by:[§2](https://arxiv.org/html/2608.21714#S2.SS0.SSS0.Px1.p1.1)\.
- Plumet al\.\(2025\)A\. Plum, T\. Ranasinghe, and C\. PurschkeText Generation Models for Luxembourgish with Limited Data: A Balanced Multilingual Strategy\.InProceedings of VarDial \(COLING\),Y\. Scherrer, T\. Jauhiainen, N\. Ljubešić, P\. Nakov, J\. Tiedemann, and M\. Zampieri \(Eds\.\),External Links:[Link](https://aclanthology.org/2025.vardial-1.7/)Cited by:[§1](https://arxiv.org/html/2608.21714#S1.p2.1),[§2](https://arxiv.org/html/2608.21714#S2.SS0.SSS0.Px1.p1.1)\.
- Purschke \(2020\)C\. PurschkeAttitudes toward multilingualism in luxembourg\. a comparative analysis of online news comments and crowdsourced questionnaire data\.Frontiers in Artificial Intelligence3\.External Links:[Link](https://api.semanticscholar.org/CorpusID:224818791)Cited by:[§2](https://arxiv.org/html/2608.21714#S2.SS0.SSS0.Px1.p1.1)\.
- Ranasingheet al\.\(2023\)T\. Ranasinghe, A\. Plum, C\. Purschke, and M\. ZampieriPublish or hold? automatic comment moderation in Luxembourgish news articles\.InProceedings of the 14th International Conference on Recent Advances in Natural Language Processing,R\. Mitkov and G\. Angelova \(Eds\.\),Varna, Bulgaria,pp\. 968–978\.External Links:[Link](https://aclanthology.org/2023.ranlp-1.104/)Cited by:[§2](https://arxiv.org/html/2608.21714#S2.SS0.SSS0.Px1.p1.1)\.
- Sirajzadeet al\.\(2020\)J\. Sirajzade, D\. Gierschek, and C\. SchommerAn Annotation Framework for Luxembourgish Sentiment Analysis\.InProceedings of SLTU\-CCURL \(LREC\),Cited by:[§2](https://arxiv.org/html/2608.21714#S2.SS0.SSS0.Px1.p1.1)\.
- Wanget al\.\(2024\)L\. Wang, N\. Yang, X\. Huang, L\. Yang, R\. Majumder, and F\. WeiMultilingual e5 text embeddings: a technical report\.arXiv preprint arXiv:2402\.05672\.Cited by:[§5](https://arxiv.org/html/2608.21714#S5.SS0.SSS0.Px1.p1.1)\.
- Zhanget al\.\(2025\)Y\. Zhang, M\. Li, D\. Long, X\. Zhang, H\. Lin, B\. Yang, P\. Xie, A\. Yang, D\. Liu, J\. Lin, F\. Huang, and J\. ZhouQwen3 embedding: advancing text embedding and reranking through foundation models\.arXiv preprint arXiv:2506\.05176\.Cited by:[§2](https://arxiv.org/html/2608.21714#S2.SS0.SSS0.Px2.p1.1),[§2](https://arxiv.org/html/2608.21714#S2.SS0.SSS0.Px3.p1.1),[§5](https://arxiv.org/html/2608.21714#S5.SS0.SSS0.Px1.p1.1)\.

Similar Articles