From Vajrayana Tara to Bengali Baul: A Computational Study of Lexical Transmission Across Buddhist, Shakta, and Vaishnava Traditions in Bengal
Summary
This paper presents a computational analysis of lexical transmission across Buddhist, Shakta, and Vaishnava traditions in Bengal, examining how words and concepts moved between these religious communities.
View Cached Full Text
Cached at: 06/26/26, 05:19 AM
# From Vajrayana Tara to Bengali Baul: A Computational Study of Lexical Transmission Across Buddhist, Shakta, and Vaishnava Traditions in Bengal Source: [https://arxiv.org/abs/2606.26803](https://arxiv.org/abs/2606.26803) Bibliographic Tools ## Bibliographic and Citation Tools Bibliographic Explorer Toggle Code, Data, Media ## Code, Data and Media Associated with this Article Demos ## Demos Related Papers ## Recommenders and Search Tools About arXivLabs ## arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website\. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy\. arXiv is committed to these values and only works with partners that adhere to them\. Have an idea for a project that will add value for arXiv's community?[**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html)\.
Similar Articles
Three Buddhist Vocabularies: Computational Stylometry of the English Pali Canon across Sutta, Vinaya, and Abhidhamma
This paper applies computational stylometry to English translations of the Pali Canon, examining vocabulary differences across the Sutta, Vinaya, and Abhidhamma divisions.
A Word-Level Digital Reader of the Prasthanatrayi with Sankara's Bhasya: Corpus, Method, and an Open, Offline Reading Aid for the Advaita Vedanta Canon
Presents an open, offline word-level digital reader of the Prasthānatrayī with Śaṅkara's Bhāṣya, featuring clickable word analysis, concordance, and a hybrid pipeline using rule-based and LLM-assisted methods.
5-Dialects-BN: Unmasking the Impact of Transliteration on Bangla Dialectal LLMs
The paper presents 5-Dialects-BN, a manually annotated benchmark dataset for five Bangla dialects with aligned transliterations and translations, designed to improve evaluation of dialect-aware LLMs.
Tokenizer Transplantation: Mitigating Autoregressive Collapse in Edge-Efficient Bengali ASR
This paper proposes a tokenizer transplantation pipeline for lightweight ASR models like Moonshine to address autoregressive collapse in Bengali. By replacing the English-centric tokenizer with a BanglaBERT WordPiece vocabulary, token fertility drops from 9.16 to 1.30 and sequence length by 85.8%, achieving 21.54% WER on the Lipi-Ghor dataset.
Which India Survives Translation? Narrative Homogenisation Across Indian Oral Traditions in LLMs
A computational study reveals that large language models partially homogenize distinct Indian oral traditions, with regional language prompts surprisingly reducing fidelity to authentic narratives.