Tag
The paper introduces VakyArth, the first pragmatic benchmark for Indic languages, evaluating LLMs on cultural and contextual reasoning across Hindi, Punjabi, Tamil, and Malayalam. It reveals consistent failures in models on pragmatic meanings, with systematic differences across languages and tasks.
SuTRA is a morphology-aware tokenization algorithm that preserves akshara indivisibility for Indic languages, reducing morphological shattering and achieving improvements in machine translation metrics over standard BPE methods.
The article presents L3Cube-IndicQuest v2, a large-scale multilingual benchmark for evaluating factual knowledge of Large Language Models across Indic languages, with evaluation results for six models.
This paper presents a retrieval-augmented translation system using BM25 and Gemini 2.5 Flash for low-resource North-Eastern Indian languages, submitted to the WMT26 shared task without model fine-tuning.
Introduces SurakshaEval, a safety benchmark for LLMs covering ten Indian languages and English, with human-written prompts spanning seven harm types. Benchmarks multilingual LLMs and finds issues like over-refusal and missed implicit bias in Indic contexts.
An open-source gateway, sarvam-bridge, lets existing ElevenLabs/OpenAI/Deepgram apps switch to Sarvam AI by changing only the base URL, handling Indic language quirks like chunking, audio reassembly, and language codes. The author details technical decisions, stress testing, and a crash bug fix.
IndicTalk is a large-scale multilingual conversational corpus covering 9 Indic languages with code-mixed dialogues, generated via an automated pipeline with news grounding and persona conditioning, aimed at advancing conversational AI for underrepresented languages.
Introduces Indi-RomCoM, a benchmark for evaluating LLMs on Romanized Code-Mixed (RCM) instructions in four Indic languages, finding that LLMs underperform on RCM tasks and performance degrades with higher code-mixing density.
This paper adapts IndicTrans2-1B to conversational register across 21 Indic languages using experience replay and model souping, achieving conversational gains without sacrificing general-domain performance, though human evaluation shows the metric-based gains may not reflect perceived quality improvements.
This paper proposes a benchmark suite grounded in Pāṇinian grammar to unify Indic language processing across languages, aiming to improve accuracy, data efficiency, and transferability.
This paper geometrically analyzes why LLMs acting as judges agree strongly with each other but weakly with humans, finding that inter-LLM consensus reflects a collapsed subspace rather than true human alignment on subjective rubrics. Post-hoc calibration on human data improves alignment, but even calibrated LLMs fall short of human reliability.
This paper proposes a multi-stage training pipeline using language-based preprocessing and an ensemble of models to detect abusive comments in Indic languages, aiming to minimize false positives while preserving freedom of expression.
SCRIBE is a diagnostic evaluation framework for automatic speech recognition that provides categorical error decomposition for Indic languages, releasing benchmarks and open-weight rich transcription models for Hindi, Malayalam, and Kannada.
Released a free 9.8 million document multilingual Indic corpus (11 languages, CC0 license) on HuggingFace, containing approximately 8.4 billion tokens, built for multilingual research.
This paper presents a case study on visually-guided movie subtitle translation for low-resource Indic languages, demonstrating that selective visual grounding improves translation quality while addressing temporal misalignment challenges.
IndicMedDialog is a parallel multi-turn medical dialogue dataset spanning English and nine Indic languages, with a fine-tuned model for personalized symptom elicitation. The dataset is derived from MDDial, enhanced with LLM-generated synthetic consultations and expert verification, supporting multilingual healthcare AI.
Researchers introduce Voice of India, a 536-hour closed benchmark of unscripted telephonic conversations across 15 Indian languages and 139 regional clusters, exposing geographic and demographic ASR performance disparities.