Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist Domains
Summary
This paper presents an end-to-end adaptation of NVIDIA's Nemotron retrieval stack for Modern Greek, including a new benchmark HERA and models fine-tuned for retrieval, reranking, and grounded generation across specialist domains.
View Cached Full Text
Cached at: 08/07/26, 09:56 AM
Paper page - Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist Domains
Source: https://huggingface.co/papers/2608.05138
Abstract
ModernGreekisabsentfromNVIDIA’sNemotronretrievalmodelsandfrommajormultilingualretrievalbenchmarks,despitebeingimportantforretrieval-augmentedgeneration(RAG)inlegal,energy,financial,andmedicalapplications.Wepresentanend-to-endadaptationoftheNemotronretrievalstackforModernGreek,includingcorpusmining,syntheticsupervision,retrievalmodeltraining,rerankeradaptation,readerfine-tuning,andanewbenchmarkcalledHERA.Ourstudyshowsthataparameter-freeBM25baselineoutperformsseveraloff-the-shelfmultilingualdenseretrievalmodelsonspecialistGreekcorpora.Afterfine-tuningon65,773Greekretrievalpairs,aNemotron1BembedderimprovesnDCG@10from0.362to0.835andsubstantiallyoutperformsitsunadaptedcounterpart.Thelearnedlanguagecompetencetransferstogeneral-domainGreek,althoughtheadvantageoverBM25remainsdomain-dependent.Wefurtheradaptacross-encoderrerankeranddemonstrateconsistentimprovementsacrossspecialistdomains.Finally,weLoRA-tuneaNemotron30B-A3Bmixture-of-expertsreaderforgroundedgeneration,increasingjudgedanswercorrectnessfrom29.4%to66.9%whilesignificantlyimprovingfaithfulnessandcitationquality.WealsointroduceHERA,thefirstlarge-scaleGreekbenchmarkforretrieval-augmentedgeneration,andreleaseouradaptedmodelsandbenchmarktosupportfutureresearchonGreek-languageRAGsystems.
View arXiv pageView PDFAdd to collection
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2608.05138 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2608.05138 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2608.05138 in a Space README.md to link it from this page.
Collections including this paper1
Similar Articles
NVIDIA Nemotron 3 Embed Ranks #1 Overall on RTEB, Advancing Agentic Retrieval
NVIDIA releases Nemotron 3 Embed, a collection of open embedding models that top the RTEB leaderboard, featuring an 8B flagship model and efficient 1B variants for production-scale retrieval.
From a Multilingual Streaming ASR Backbone to Kenyan-Language Systems: Data-Centric Adaptation of Nemotron 3.5 for Kikuyu, Dholuo, and Kalenjin
This paper presents an engineering study adapting NVIDIA Nemotron 3.5 ASR Streaming 0.6B to Kikuyu, Dholuo, and Kalenjin, achieving 42.97% and 33.98% WER on internal sets for Kikuyu and Dholuo, respectively, through data-centric techniques including corpus auditing, normalization, and streaming evaluation.
Building a Fast Multilingual OCR Model with Synthetic Data
NVIDIA introduces Nemotron OCR v2, a fast multilingual OCR model built using synthetic data generation. The model achieves 34.7 pages/second on a single A100 GPU by using a unified FOTS-based architecture with feature reuse across detection, recognition, and relational components.
NVIDIA has released Nemotron-TwoTower-30B-A3B-Base-BF16, an unusual diffusion-based language model built from the Nemotron 3 Nano 30B-A3B backbone.
NVIDIA released Nemotron-TwoTower-30B-A3B-Base-BF16, a diffusion-based language model that uses block-wise autoregressive diffusion to generate text by iterative denoising of token blocks, achieving 2.42× the generation throughput of the autoregressive baseline while retaining 98.7% of benchmark quality.
HEBATRON: A Hebrew-Specialized Open-Weight Mixture-of-Experts Language Model
Hebatron is a new open-weight Hebrew-specialized Large Language Model built on NVIDIA's Nemotron-3 Mixture-of-Experts architecture, achieving strong reasoning performance with efficient inference. It is the first language-specific adaptation of this architecture and supports native long-context processing.