Tag
This paper introduces BHARATI, a set of SentencePiece BPE tokenizers trained on a balanced multilingual corpus for classical Indian languages. Subword fertility analysis shows significant improvements in tokenization efficiency, reducing sequence length by up to 90% compared to baseline tokenizers, thereby enhancing effective context length for downstream language models.
Introduces PaliBench, a multi-reference benchmark for Pali-to-English translation using independent translations from multiple scholars, and a reusable methodology for creating similar benchmarks for classical languages.