Inspect India Evals: An Open Benchmarking Framework for Evaluating Large Language Models in the Indian Linguistic and Cultural Context

arXiv cs.CL Papers

Summary

Introduces Inspect India Evals, an open-source framework for evaluating LLMs in Indian linguistic and cultural contexts, with six benchmarks testing multilingual ability, bias, safety, and cultural knowledge. Tests on five models show Sarvam-M 24B and Gemma 2 27B lead.

arXiv:2607.25375v1 Announce Type: new Abstract: India is a vast nation of over 1.4 billion people, varied by hundreds of diverse and locally specific traditions and cultures and 22 officially recognized languages. Large language models (LLMs) are now being deployed on a massive scale throughout the mainland as well as in remote villages. However, the common benchmarks - MMLU, BIG-Bench, and TruthfulQA are almost exclusively English- and Western-centric. They do not identify those safety, fairness, and accuracy failures unique to the Indian context. That is the gap Inspect India Evals seeks to fill. It is an open-source framework built on top of UK AISI's Inspect AI platform. It has six benchmarks: Multilingual MMLU across sixteen Indian languages, BharatBBQ (our adaptation of BBQ for Indian social bias), a safety evaluation for Digital Public Infrastructure, a multilingual safety test using harmful prompts in Indian languages, a multi-turn jailbreak resistance test, and an Indian cultural knowledge benchmark scored using LLM-as-judge rubrics. In this study, we tested five open-weight models ranging from 8B to 32B parameters. Sarvam-M 24B and Gemma 2 27B came out on top, both scoring 80% on the composite India Fairness Index, with Sarvam-M even beating larger 32B models on Indian cultural knowledge and DPI safety compliance. All models scored 100% refusal on Multilingual Safety, whereas DPI safety varied from 20% to 100%. The framework is public. It's built to work with the UK AISI registry. Anyone can reproduce or extend this work.
Original Article
View Cached Full Text

Cached at: 07/29/26, 09:55 AM

# Inspect India Evals: An Open Benchmarking Framework for Evaluating Large Language Models in the Indian Linguistic and Cultural Context
Source: [https://arxiv.org/abs/2607.25375](https://arxiv.org/abs/2607.25375)
[View PDF](https://arxiv.org/pdf/2607.25375)

> Abstract:India is a vast nation of over 1\.4 billion people, varied by hundreds of diverse and locally specific traditions and cultures and 22 officially recognized languages\. Large language models \(LLMs\) are now being deployed on a massive scale throughout the mainland as well as in remote villages\. However, the common benchmarks \- MMLU, BIG\-Bench, and TruthfulQA are almost exclusively English\- and Western\-centric\. They do not identify those safety, fairness, and accuracy failures unique to the Indian context\. That is the gap Inspect India Evals seeks to fill\. It is an open\-source framework built on top of UK AISI's Inspect AI platform\. It has six benchmarks: Multilingual MMLU across sixteen Indian languages, BharatBBQ \(our adaptation of BBQ for Indian social bias\), a safety evaluation for Digital Public Infrastructure, a multilingual safety test using harmful prompts in Indian languages, a multi\-turn jailbreak resistance test, and an Indian cultural knowledge benchmark scored using LLM\-as\-judge rubrics\. In this study, we tested five open\-weight models ranging from 8B to 32B parameters\. Sarvam\-M 24B and Gemma 2 27B came out on top, both scoring 80% on the composite India Fairness Index, with Sarvam\-M even beating larger 32B models on Indian cultural knowledge and DPI safety compliance\. All models scored 100% refusal on Multilingual Safety, whereas DPI safety varied from 20% to 100%\. The framework is public\. It's built to work with the UK AISI registry\. Anyone can reproduce or extend this work\.

## Submission history

From: Abhishek Singh \[[view email](https://arxiv.org/show-email/a7477d01/2607.25375)\] **\[v1\]**Tue, 28 Jul 2026 07:30:12 UTC \(1,132 KB\)

Similar Articles

Benchmarking Frontier LLMs on Arabic Cultural and Sociolinguistic Knowledge: A Cross-Evaluation Framework with Human SME Ground Truth

arXiv cs.CL

This paper introduces a cross-evaluation framework for benchmarking LLMs on Arabic cultural and sociolinguistic knowledge, using human SME ground truth and automated judges. The authors contribute a dataset of prompt-rubric pairs for Egyptian and Iraqi Arabic, evaluating frontier LLMs and finding that cultural reasoning remains a primary failure mode for automated grading.

Introducing IndQA

OpenAI Blog

OpenAI introduced IndQA, a new benchmark with 2,278 questions across 12 Indian languages and 10 cultural domains, designed to evaluate AI models' understanding of culturally nuanced and reasoning-heavy tasks that existing benchmarks fail to capture. Created with 261 domain experts, IndQA addresses the saturation of existing multilingual benchmarks like MMMLU and focuses on real-world cultural comprehension rather than translation or multiple-choice tasks.

When English Rewrites Local Knowledge: Global Narrative Dominance in Large Language Models

arXiv cs.CL

This paper introduces CulturalNB, a dataset of Bengali cultural question-answer pairs, and evaluates nine LLMs for cross-lingual cultural bias. Findings show that English prompting increases global narrative substitution and reduces local perspectives, revealing that cultural failures in LLMs are grounding and prioritization issues, not just missing knowledge.