indic-languages

Tag

Cards List
#indic-languages

VakyArth: Evaluating Pragmatic Competence in LLMs across Indic Languages

arXiv cs.CL ↗ · 2026-09-03 Cached

The paper introduces VakyArth, the first pragmatic benchmark for Indic languages, evaluating LLMs on cultural and contextual reasoning across Hindi, Punjabi, Tamil, and Malayalam. It reveals consistent failures in models on pragmatic meanings, with systematic differences across languages and tasks.

0 favorites 0 likes
#indic-languages

SuTRA : Structurally-Unified Tokenization with Root Awareness

arXiv cs.CL ↗ · 2026-08-20 Cached

SuTRA is a morphology-aware tokenization algorithm that preserves akshara indivisibility for Indic languages, reducing morphological shattering and achieving improvements in machine translation metrics over standard BPE methods.

0 favorites 0 likes
#indic-languages

L3Cube-IndicQuest v2: A Large-Scale Multilingual Benchmark for Evaluating Factual Knowledge of Large Language Models Across Indic Languages

arXiv cs.CL ↗ · 2026-08-18 Cached

The article presents L3Cube-IndicQuest v2, a large-scale multilingual benchmark for evaluating factual knowledge of Large Language Models across Indic languages, with evaluation results for six models.

0 favorites 0 likes
#indic-languages

BM25-Augmented Many-Shot Translation for Low-Resource North-Eastern Indian Languages

arXiv cs.CL ↗ · 2026-08-17 Cached

This paper presents a retrieval-augmented translation system using BM25 and Gemini 2.5 Flash for low-resource North-Eastern Indian languages, submitted to the WMT26 shared task without model fine-tuning.

0 favorites 0 likes
#indic-languages

SurakshaEval: An Indic Safety Benchmark for Multilingual LLMs

arXiv cs.CL ↗ · 2026-08-11 Cached

Introduces SurakshaEval, a safety benchmark for LLMs covering ten Indian languages and English, with human-written prompts spanning seven harm types. Benchmarks multilingual LLMs and finds issues like over-refusal and missed implicit bias in Indic contexts.

0 favorites 0 likes
#indic-languages

Built an open-source gateway that lets existing ElevenLabs / OpenAI / Deepgram apps run on Sarvam AI by changing one line.

Reddit r/ArtificialInteligence ↗ · 2026-08-08

An open-source gateway, sarvam-bridge, lets existing ElevenLabs/OpenAI/Deepgram apps switch to Sarvam AI by changing only the base URL, handling Indic language quirks like chunking, audio reassembly, and language codes. The author details technical decisions, stress testing, and a crash bug fix.

0 favorites 0 likes
#indic-languages

IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languages

arXiv cs.CL ↗ · 2026-07-28 Cached

IndicTalk is a large-scale multilingual conversational corpus covering 9 Indic languages with code-mixed dialogues, generated via an automated pipeline with news grounding and persona conditioning, aimed at advancing conversational AI for underrepresented languages.

0 favorites 0 likes
#indic-languages

Indi-RomCoM: Code-Mixed Benchmark for Evaluating LLMs on Romanized Indic-English Instructions

arXiv cs.CL ↗ · 2026-07-01 Cached

Introduces Indi-RomCoM, a benchmark for evaluating LLMs on Romanized Code-Mixed (RCM) instructions in four Indic languages, finding that LLMs underperform on RCM tasks and performance degrades with higher code-mixing density.

0 favorites 0 likes
#indic-languages

Conversational Domain Adaptation of IndicTrans2 across 21 Indic Languages via Experience Replay and Model Soups

arXiv cs.CL ↗ · 2026-06-30 Cached

This paper adapts IndicTrans2-1B to conversational register across 21 Indic languages using experience replay and model souping, achieving conversational gains without sacrificing general-domain performance, though human evaluation shows the metric-based gains may not reflect perceived quality improvements.

0 favorites 0 likes
#indic-languages

A P\={a}ninian Foundation for Indic Language Processing

arXiv cs.CL ↗ · 2026-06-24 Cached

This paper proposes a benchmark suite grounded in Pāṇinian grammar to unify Indic language processing across languages, aiming to improve accuracy, data efficiency, and transferability.

0 favorites 0 likes
#indic-languages

The Geometry of LLM-as-Judge: Why Inter-LLM Consensus Is Not Human Alignment

arXiv cs.CL ↗ · 2026-06-03 Cached

This paper geometrically analyzes why LLMs acting as judges agree strongly with each other but weakly with humans, finding that inter-LLM consensus reflects a collapsed subspace rather than true human alignment on subjective rubrics. Post-hoc calibration on human data improves alignment, but even calibrated LLMs fall short of human reliability.

0 favorites 0 likes
#indic-languages

Multi-Stage Training for Abusive Comment Detection in Indic Languages

arXiv cs.CL ↗ · 2026-05-22 Cached

This paper proposes a multi-stage training pipeline using language-based preprocessing and an ensemble of models to detect abusive comments in Indic languages, aiming to minimize false positives while preserving freedom of expression.

0 favorites 0 likes
#indic-languages

SCRIBE: Diagnostic Evaluation and Rich Transcription Models for Indic ASR

arXiv cs.CL ↗ · 2026-05-21 Cached

SCRIBE is a diagnostic evaluation framework for automatic speech recognition that provides categorical error decomposition for Indic languages, releasing benchmarks and open-weight rich transcription models for Hindi, Malayalam, and Kannada.

0 favorites 0 likes
#indic-languages

Released a free 9.8M doc Indic multilingual corpus — Hindi, Bengali, Tamil, Telugu + 7 more (CC0, HuggingFace) [P]

Reddit r/MachineLearning ↗ · 2026-05-18

Released a free 9.8 million document multilingual Indic corpus (11 languages, CC0 license) on HuggingFace, containing approximately 8.4 billion tokens, built for multilingual research.

0 favorites 0 likes
#indic-languages

Towards Visually-Guided Movie Subtitle Translation for Indic Languages

arXiv cs.CL ↗ · 2026-05-13 Cached

This paper presents a case study on visually-guided movie subtitle translation for low-resource Indic languages, demonstrating that selective visual grounding improves translation quality while addressing temporal misalignment challenges.

0 favorites 0 likes
#indic-languages

IndicMedDialog: A Parallel Multi-Turn Medical Dialogue Dataset for Accessible Healthcare in Indic Languages

Hugging Face Daily Papers ↗ · 2026-05-13 Cached

IndicMedDialog is a parallel multi-turn medical dialogue dataset spanning English and nine Indic languages, with a fine-tuned model for personalized symptom elicitation. The dataset is derived from MDDial, enhanced with LLM-generated synthetic consultations and expert verification, supporting multilingual healthcare AI.

0 favorites 0 likes
#indic-languages

Voice of India: A Large-Scale Benchmark for Real-World Speech Recognition in India

arXiv cs.CL ↗ · 2026-04-22 Cached

Researchers introduce Voice of India, a 536-hour closed benchmark of unscripted telephonic conversations across 15 Indian languages and 139 regional clusters, exposing geographic and demographic ASR performance disparities.

0 favorites 0 likes
← Back to home

Submit Feedback