Tag
Mizan introduces a national benchmark for evaluating large language models on Iraqi Arabic and civic context, highlighting gaps in MSA-focused evaluations and revealing issues like over-refusal in safety-hardened models.
Introduces a scalable, inference-only data valuation pipeline that approximates Shapley values to audit LLM alignment datasets, reducing manual audit search space by 99.1% and uncovering hidden label failures in HelpSteer2 and HH-RLHF.