Tag
This paper identifies a non-composition principle in AI benchmark evaluation: support for adjacent projections does not automatically warrant their composition. It proposes a projectibility audit to diagnose unsupported joins in benchmark-to-use arguments, with a legal-research case study and simulations.
This paper identifies five failure modes in perturbation-based benchmark-validity audits used for AI governance, demonstrating that implementation details can silently manufacture conclusions. It proposes a due-diligence gate to improve the reliability of evaluation evidence.
This paper audits the 'Translation Tax' in Chinese multilingual benchmarks, arguing it is not a scalar but a set of estimator- and item-dependent validity risks. It introduces a naturalization stress test to quantify how English-source cues inflate model scores.