IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languages

arXiv cs.CL Papers

Summary

IndicTalk is a large-scale multilingual conversational corpus covering 9 Indic languages with code-mixed dialogues, generated via an automated pipeline with news grounding and persona conditioning, aimed at advancing conversational AI for underrepresented languages.

arXiv:2607.23242v1 Announce Type: new Abstract: Large Language Models (LLMs) have transformed conversational AI, yet high-quality multilingual code-mixed dialogue resources remain scarce, particularly for Indic languages where speakers naturally alternate between English and their native language in both native-script and Romanized forms. We present IndicTalk, one of the largest multilingual Indic code-mixed conversational corpora, comprising over 13,28,604 event-grounded multi-turn conversations across 18 language varieties covering 9 Indic languages. The corpus is generated through a fully automated pipeline that combines real-world news grounding, persona-conditioned dialogue generation using multilingual LLMs, and automatic quality validation. Extensive linguistic, automatic, and human evaluations demonstrate that IndicTalk produces fluent, coherent, and naturally code-mixed conversations across both script variants. We will release IndicTalk to support the development and evaluation of multilingual conversational AI for underrepresented Indic languages. The dataset is available at: https://huggingface.co/datasets/LingoIITGN/IndicTalk .
Original Article
View Cached Full Text

Cached at: 07/28/26, 06:28 AM

# IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languages
Source: [https://arxiv.org/abs/2607.23242](https://arxiv.org/abs/2607.23242)
[View PDF](https://arxiv.org/pdf/2607.23242)

> Abstract:Large Language Models \(LLMs\) have transformed conversational AI, yet high\-quality multilingual code\-mixed dialogue resources remain scarce, particularly for Indic languages where speakers naturally alternate between English and their native language in both native\-script and Romanized forms\. We present IndicTalk, one of the largest multilingual Indic code\-mixed conversational corpora, comprising over 13,28,604 event\-grounded multi\-turn conversations across 18 language varieties covering 9 Indic languages\. The corpus is generated through a fully automated pipeline that combines real\-world news grounding, persona\-conditioned dialogue generation using multilingual LLMs, and automatic quality validation\. Extensive linguistic, automatic, and human evaluations demonstrate that IndicTalk produces fluent, coherent, and naturally code\-mixed conversations across both script variants\. We will release IndicTalk to support the development and evaluation of multilingual conversational AI for underrepresented Indic languages\. The dataset is available at:[this https URL](https://huggingface.co/datasets/LingoIITGN/IndicTalk)\.

## Submission history

From: Sahil Gawande \[[view email](https://arxiv.org/show-email/da6b7807/2607.23242)\] **\[v1\]**Sat, 25 Jul 2026 14:56:15 UTC \(405 KB\)

Similar Articles