Padamitra: Grounded Glossary Generation for Classical Sanskrit

arXiv cs.CL Papers

Summary

This paper introduces grounded glossary generation for Classical Sanskrit, a task involving recovering Sanskrit phrases and producing translation-grounded meanings from sloka-translation pairs. It constructs a benchmark from Hindu texts and evaluates various AI models, finding that instruction fine-tuning improves performance, with morphological modeling identified as a key challenge.

arXiv:2608.25038v1 Announce Type: new Abstract: We introduce grounded glossary generation, a structured task requiring models to recover semantically meaningful Sanskrit phrases and produce translation-grounded meanings from a sloka-translation pair, formalizing the traditional patha commentary practice as an evaluable NLP objective. We construct a benchmark of 31,316 sloka-translation-glossary triples from the Valmiki Ramayana and Srimad Bhagavatam, paired with two metrics: Jaccard for phrase recovery and Meaning Faithfulness for semantic consistency. Across zero-shot, few-shot, and instruction fine-tuned variants of Gemma-3n-E4B, Gemma-3-12B, Phi-4, and Qwen3.5-9B, instruction fine-tuning substantially outperforms prompting, while explicit segmentation yields gains. Error analysis identifies over-segmentation of sandhi and samasa compounds as the dominant failure mode, pointing to morphological modeling as the key bottleneck for faithful Sanskrit lexical decomposition.
Original Article
View Cached Full Text

Cached at: 08/27/26, 09:14 AM

# Padamitra: Grounded Glossary Generation for Classical Sanskrit
Source: [https://arxiv.org/abs/2608.25038](https://arxiv.org/abs/2608.25038)
[View PDF](https://arxiv.org/pdf/2608.25038)

> Abstract:We introduce grounded glossary generation, a structured task requiring models to recover semantically meaningful Sanskrit phrases and produce translation\-grounded meanings from a sloka\-translation pair, formalizing the traditional patha commentary practice as an evaluable NLP objective\. We construct a benchmark of 31,316 sloka\-translation\-glossary triples from the Valmiki Ramayana and Srimad Bhagavatam, paired with two metrics: Jaccard for phrase recovery and Meaning Faithfulness for semantic consistency\. Across zero\-shot, few\-shot, and instruction fine\-tuned variants of Gemma\-3n\-E4B, Gemma\-3\-12B, Phi\-4, and Qwen3\.5\-9B, instruction fine\-tuning substantially outperforms prompting, while explicit segmentation yields gains\. Error analysis identifies over\-segmentation of sandhi and samasa compounds as the dominant failure mode, pointing to morphological modeling as the key bottleneck for faithful Sanskrit lexical decomposition\.

## Submission history

From: Manoj Balaji Jagadeeshan \[[view email](https://arxiv.org/show-email/a27a4631/2608.25038)\] **\[v1\]**Tue, 25 Aug 2026 18:27:28 UTC \(378 KB\)

Similar Articles