Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist Domains

Hugging Face Daily Papers Papers

Summary

This paper presents an end-to-end adaptation of NVIDIA's Nemotron retrieval stack for Modern Greek, including a new benchmark HERA and models fine-tuned for retrieval, reranking, and grounded generation across specialist domains.

Modern Greek is absent from NVIDIA's Nemotron retrieval models and from major multilingual retrieval benchmarks, despite being important for retrieval-augmented generation (RAG) in legal, energy, financial, and medical applications. We present an end-to-end adaptation of the Nemotron retrieval stack for Modern Greek, including corpus mining, synthetic supervision, retrieval model training, reranker adaptation, reader fine-tuning, and a new benchmark called HERA. Our study shows that a parameter-free BM25 baseline outperforms several off-the-shelf multilingual dense retrieval models on specialist Greek corpora. After fine-tuning on 65,773 Greek retrieval pairs, a Nemotron 1B embedder improves nDCG@10 from 0.362 to 0.835 and substantially outperforms its unadapted counterpart. The learned language competence transfers to general-domain Greek, although the advantage over BM25 remains domain-dependent. We further adapt a cross-encoder reranker and demonstrate consistent improvements across specialist domains. Finally, we LoRA-tune a Nemotron 30B-A3B mixture-of-experts reader for grounded generation, increasing judged answer correctness from 29.4% to 66.9% while significantly improving faithfulness and citation quality. We also introduce HERA, the first large-scale Greek benchmark for retrieval-augmented generation, and release our adapted models and benchmark to support future research on Greek-language RAG systems.
Original Article
View Cached Full Text

Cached at: 08/07/26, 09:56 AM

Paper page - Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist Domains

Source: https://huggingface.co/papers/2608.05138

Abstract

ModernGreekisabsentfromNVIDIA’sNemotronretrievalmodelsandfrommajormultilingualretrievalbenchmarks,despitebeingimportantforretrieval-augmentedgeneration(RAG)inlegal,energy,financial,andmedicalapplications.Wepresentanend-to-endadaptationoftheNemotronretrievalstackforModernGreek,includingcorpusmining,syntheticsupervision,retrievalmodeltraining,rerankeradaptation,readerfine-tuning,andanewbenchmarkcalledHERA.Ourstudyshowsthataparameter-freeBM25baselineoutperformsseveraloff-the-shelfmultilingualdenseretrievalmodelsonspecialistGreekcorpora.Afterfine-tuningon65,773Greekretrievalpairs,aNemotron1BembedderimprovesnDCG@10from0.362to0.835andsubstantiallyoutperformsitsunadaptedcounterpart.Thelearnedlanguagecompetencetransferstogeneral-domainGreek,althoughtheadvantageoverBM25remainsdomain-dependent.Wefurtheradaptacross-encoderrerankeranddemonstrateconsistentimprovementsacrossspecialistdomains.Finally,weLoRA-tuneaNemotron30B-A3Bmixture-of-expertsreaderforgroundedgeneration,increasingjudgedanswercorrectnessfrom29.4%to66.9%whilesignificantlyimprovingfaithfulnessandcitationquality.WealsointroduceHERA,thefirstlarge-scaleGreekbenchmarkforretrieval-augmentedgeneration,andreleaseouradaptedmodelsandbenchmarktosupportfutureresearchonGreek-languageRAGsystems.

View arXiv pageView PDFAdd to collection

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2608.05138 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2608.05138 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2608.05138 in a Space README.md to link it from this page.

Collections including this paper1

Similar Articles

Building a Fast Multilingual OCR Model with Synthetic Data

Hugging Face Blog

NVIDIA introduces Nemotron OCR v2, a fast multilingual OCR model built using synthetic data generation. The model achieves 34.7 pages/second on a single A100 GPU by using a unified FOTS-based architecture with feature reuse across detection, recognition, and relational components.

HEBATRON: A Hebrew-Specialized Open-Weight Mixture-of-Experts Language Model

arXiv cs.CL

Hebatron is a new open-weight Hebrew-specialized Large Language Model built on NVIDIA's Nemotron-3 Mixture-of-Experts architecture, achieving strong reasoning performance with efficient inference. It is the first language-specific adaptation of this architecture and supports native long-context processing.