Fairness Auditing: Lower Bounds on Company Manipulation
Summary
This paper derives lower bounds on how much a company can manipulate fairness metrics after an audit, showing that finite-budget fairness certification cannot fully eliminate post-audit manipulation.
View Cached Full Text
Cached at: 08/04/26, 07:42 AM
# Fairness Auditing: Lower Bounds on Company Manipulation
Source: [https://arxiv.org/abs/2608.00568](https://arxiv.org/abs/2608.00568)
[View PDF](https://arxiv.org/pdf/2608.00568)
> Abstract:Fairness audits are increasingly mandated in high\-stakes applications such as hiring, lending, and automated decision\-making\. Recent work has established fundamental impossibility results for black\-box fairness auditing, showing that sufficiently expressive models can evade any auditing strategy\. We complement these results by quantifying the extent of unavoidable post\-audit manipulation under finite audit resources\. We formulate fairness auditing as a min\-max optimization between a computationally unbounded company and a budget\-constrained auditor\. We study two auditing regimes: \(i\) a budgeted auditor that certifies fairness using a fixed\-size audit set, and \(ii\) a budgeted \{\\alpha\}\-tolerant auditor that additionally requires the audit set to estimate the fairness of the certified model within an \{\\alpha\} approximation\. For both settings, we derive explicit lower bounds on the worst\-case post\-audit demographic parity deviation as functions of the audit budget, group imbalance, and fairness tolerance\. Finally, we empirically illustrate these theoretical limits using simple audit\-set construction heuristics with linear and neural network classifiers\. Our results demonstrate that increasing audit resources reduces, but does not eliminate, the scope for post\-audit manipulation, highlighting fundamental limitations of finite\-budget fairness certification\.
## Submission history
From: Rachit Verma \[[view email](https://arxiv.org/show-email/840c701e/2608.00568)\] **\[v1\]**Sat, 1 Aug 2026 10:11:20 UTC \(47 KB\)Similar Articles
Manipulation-Proof Oblivious Audits against Deceptive Model Providers
This paper introduces a novel audit protocol using Private Information Retrieval to make audits manipulation-proof, forcing deceptive model providers to falsify more responses and increasing detection likelihood, with theoretical guarantees and experimental validation.
Auditing the Audit: Five Failure Modes in Benchmark-Validity Audits
This paper identifies five failure modes in perturbation-based benchmark-validity audits used for AI governance, demonstrating that implementation details can silently manufacture conclusions. It proposes a due-diligence gate to improve the reliability of evaluation evidence.
The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection
This paper identifies distribution shift and scale constraints as critical failure modes for statistical contamination detection methods in LLM benchmark auditing. Evaluating three paradigms across 27 models reveals only 199 correct outcomes out of 335 evaluations, indicating a systematic reliability gap that prevents these methods from replacing transparent data provenance.
FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation
Introduces FairFund-Bench, a benchmark for evaluating distributive bias in LLM resource allocation, showing that audit format changes the direction and magnitude of bias, and that causal framing effects dominate demographic effects.
When Is Benchmark Contamination Detectable? Information Limits and Power-Calibrated Audits
This paper formalizes when benchmark contamination is detectable, deriving information-theoretic limits and proposing power-calibrated audits that distinguish a clean benchmark from a powerless detector. It reports two-sided empirical findings on calibration efficacy and validity gates.