Fairness Auditing: Lower Bounds on Company Manipulation

arXiv cs.LG Papers

Summary

This paper derives lower bounds on how much a company can manipulate fairness metrics after an audit, showing that finite-budget fairness certification cannot fully eliminate post-audit manipulation.

arXiv:2608.00568v1 Announce Type: new Abstract: Fairness audits are increasingly mandated in high-stakes applications such as hiring, lending, and automated decision-making. Recent work has established fundamental impossibility results for black-box fairness auditing, showing that sufficiently expressive models can evade any auditing strategy. We complement these results by quantifying the extent of unavoidable post-audit manipulation under finite audit resources. We formulate fairness auditing as a min-max optimization between a computationally unbounded company and a budget-constrained auditor. We study two auditing regimes: (i) a budgeted auditor that certifies fairness using a fixed-size audit set, and (ii) a budgeted {\alpha}-tolerant auditor that additionally requires the audit set to estimate the fairness of the certified model within an {\alpha} approximation. For both settings, we derive explicit lower bounds on the worst-case post-audit demographic parity deviation as functions of the audit budget, group imbalance, and fairness tolerance. Finally, we empirically illustrate these theoretical limits using simple audit-set construction heuristics with linear and neural network classifiers. Our results demonstrate that increasing audit resources reduces, but does not eliminate, the scope for post-audit manipulation, highlighting fundamental limitations of finite-budget fairness certification.
Original Article
View Cached Full Text

Cached at: 08/04/26, 07:42 AM

# Fairness Auditing: Lower Bounds on Company Manipulation
Source: [https://arxiv.org/abs/2608.00568](https://arxiv.org/abs/2608.00568)
[View PDF](https://arxiv.org/pdf/2608.00568)

> Abstract:Fairness audits are increasingly mandated in high\-stakes applications such as hiring, lending, and automated decision\-making\. Recent work has established fundamental impossibility results for black\-box fairness auditing, showing that sufficiently expressive models can evade any auditing strategy\. We complement these results by quantifying the extent of unavoidable post\-audit manipulation under finite audit resources\. We formulate fairness auditing as a min\-max optimization between a computationally unbounded company and a budget\-constrained auditor\. We study two auditing regimes: \(i\) a budgeted auditor that certifies fairness using a fixed\-size audit set, and \(ii\) a budgeted \{\\alpha\}\-tolerant auditor that additionally requires the audit set to estimate the fairness of the certified model within an \{\\alpha\} approximation\. For both settings, we derive explicit lower bounds on the worst\-case post\-audit demographic parity deviation as functions of the audit budget, group imbalance, and fairness tolerance\. Finally, we empirically illustrate these theoretical limits using simple audit\-set construction heuristics with linear and neural network classifiers\. Our results demonstrate that increasing audit resources reduces, but does not eliminate, the scope for post\-audit manipulation, highlighting fundamental limitations of finite\-budget fairness certification\.

## Submission history

From: Rachit Verma \[[view email](https://arxiv.org/show-email/840c701e/2608.00568)\] **\[v1\]**Sat, 1 Aug 2026 10:11:20 UTC \(47 KB\)

Similar Articles

Manipulation-Proof Oblivious Audits against Deceptive Model Providers

arXiv cs.LG

This paper introduces a novel audit protocol using Private Information Retrieval to make audits manipulation-proof, forcing deceptive model providers to falsify more responses and increasing detection likelihood, with theoretical guarantees and experimental validation.

Auditing the Audit: Five Failure Modes in Benchmark-Validity Audits

arXiv cs.LG

This paper identifies five failure modes in perturbation-based benchmark-validity audits used for AI governance, demonstrating that implementation details can silently manufacture conclusions. It proposes a due-diligence gate to improve the reliability of evaluation evidence.

The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection

arXiv cs.AI

This paper identifies distribution shift and scale constraints as critical failure modes for statistical contamination detection methods in LLM benchmark auditing. Evaluating three paradigms across 27 models reveals only 199 correct outcomes out of 335 evaluations, indicating a systematic reliability gap that prevents these methods from replacing transparent data provenance.