benchmark-integrity

Tag

Cards List
#benchmark-integrity

Mizan: A National Benchmark for Evaluating Large Language Models on Iraqi Arabic and the Iraqi Civic Context

arXiv cs.CL · 5d ago Cached

Mizan introduces a national benchmark for evaluating large language models on Iraqi Arabic and civic context, highlighting gaps in MSA-focused evaluations and revealing issues like over-refusal in safety-hardened models.

0 favorites 0 likes
#benchmark-integrity

Beyond Shapley: An Influence-Based Data Auditing Pipeline for LLM Alignment and Evaluation

arXiv cs.LG · 2026-07-28 Cached

Introduces a scalable, inference-only data valuation pipeline that approximates Shapley values to audit LLM alignment datasets, reducing manual audit search space by 99.1% and uncovering hidden label failures in HelpSteer2 and HH-RLHF.

0 favorites 0 likes
← Back to home

Submit Feedback