IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languages
Summary
IndicTalk is a large-scale multilingual conversational corpus covering 9 Indic languages with code-mixed dialogues, generated via an automated pipeline with news grounding and persona conditioning, aimed at advancing conversational AI for underrepresented languages.
View Cached Full Text
Cached at: 07/28/26, 06:28 AM
# IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languages Source: [https://arxiv.org/abs/2607.23242](https://arxiv.org/abs/2607.23242) [View PDF](https://arxiv.org/pdf/2607.23242) > Abstract:Large Language Models \(LLMs\) have transformed conversational AI, yet high\-quality multilingual code\-mixed dialogue resources remain scarce, particularly for Indic languages where speakers naturally alternate between English and their native language in both native\-script and Romanized forms\. We present IndicTalk, one of the largest multilingual Indic code\-mixed conversational corpora, comprising over 13,28,604 event\-grounded multi\-turn conversations across 18 language varieties covering 9 Indic languages\. The corpus is generated through a fully automated pipeline that combines real\-world news grounding, persona\-conditioned dialogue generation using multilingual LLMs, and automatic quality validation\. Extensive linguistic, automatic, and human evaluations demonstrate that IndicTalk produces fluent, coherent, and naturally code\-mixed conversations across both script variants\. We will release IndicTalk to support the development and evaluation of multilingual conversational AI for underrepresented Indic languages\. The dataset is available at:[this https URL](https://huggingface.co/datasets/LingoIITGN/IndicTalk)\. ## Submission history From: Sahil Gawande \[[view email](https://arxiv.org/show-email/da6b7807/2607.23242)\] **\[v1\]**Sat, 25 Jul 2026 14:56:15 UTC \(405 KB\)
Similar Articles
Conversational Domain Adaptation of IndicTrans2 across 21 Indic Languages via Experience Replay and Model Soups
This paper adapts IndicTrans2-1B to conversational register across 21 Indic languages using experience replay and model souping, achieving conversational gains without sacrificing general-domain performance, though human evaluation shows the metric-based gains may not reflect perceived quality improvements.
IndicMedDialog: A Parallel Multi-Turn Medical Dialogue Dataset for Accessible Healthcare in Indic Languages
IndicMedDialog is a parallel multi-turn medical dialogue dataset spanning English and nine Indic languages, with a fine-tuned model for personalized symptom elicitation. The dataset is derived from MDDial, enhanced with LLM-generated synthetic consultations and expert verification, supporting multilingual healthcare AI.
Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages
Indic DiarBench is a multilingual joint diarization and ASR benchmark covering all 22 scheduled languages of India with 108 hours of human-corrected multi-speaker audio, capturing conversational nuances like code-mixing and speaker overlap.
Voice of India: A Large-Scale Benchmark for Real-World Speech Recognition in India
Researchers introduce Voice of India, a 536-hour closed benchmark of unscripted telephonic conversations across 15 Indian languages and 139 regional clusters, exposing geographic and demographic ASR performance disparities.
From Lexicon to AI: A Structured-Data Pipeline for Specialized Conversational Systems in Low-Resource Languages
Presents a systematic methodology for converting Hindi WordNet into 1.25 million instruction-response pairs to fine-tune a 12B-parameter language model using LoRA, demonstrating improved pedagogical effectiveness for specialized conversational systems in low-resource languages.