benchmark-validity

Tag

Cards List
#benchmark-validity

When benchmark inferences do not compose: Projectibility in AI evaluation

arXiv cs.AI · 22h ago Cached

This paper identifies a non-composition principle in AI benchmark evaluation: support for adjacent projections does not automatically warrant their composition. It proposes a projectibility audit to diagnose unsupported joins in benchmark-to-use arguments, with a legal-research case study and simulations.

0 favorites 0 likes
#benchmark-validity

Auditing the Audit: Five Failure Modes in Benchmark-Validity Audits

arXiv cs.LG · 2026-07-07 Cached

This paper identifies five failure modes in perturbation-based benchmark-validity audits used for AI governance, demonstrating that implementation details can silently manufacture conclusions. It proposes a due-diligence gate to improve the reliability of evaluation evidence.

0 favorites 0 likes
#benchmark-validity

The Translation Tax Is Not a Scalar: A Counterfactual Audit of English-Source Cue Inheritance in Chinese Multilingual Benchmarks

arXiv cs.CL · 2026-05-11 Cached

This paper audits the 'Translation Tax' in Chinese multilingual benchmarks, arguing it is not a scalar but a set of estimator- and item-dependent validity risks. It introduces a naturalization stress test to quantify how English-source cues inflate model scores.

0 favorites 0 likes
← Back to home

Submit Feedback