Competence-Preserving Resume Perturbations Expose Presentation Sensitivity in LLM Screening

arXiv cs.CL Papers

Summary

This paper introduces a controlled audit of resume screening by LLMs, revealing that decisions can be unstable under presentation changes even when competence evidence is identical.

arXiv:2609.16517v1 Announce Type: new Abstract: Resume screeners must infer job-relevant competence from resumes whose presentation can vary substantially in wording, structure, stylistic polish, and document extraction quality. Ideally, such surface variation should not change decisions when the underlying qualification evidence is unchanged. We introduce a controlled audit of this property, constructing occupation-grounded candidate profiles at controlled competence levels and rendering each profile into multiple resume presentations. A deterministic validation gate excludes variants that alter the underlying evidence before scoring. Across six open instruction-tuned LLM conditions, we find a clear disconnect between screening validity and presentation stability. Llama-3.1-8B with its native chat template achieves the strongest validity ($0.781$) yet reverses $29.6\%$ of matched pairwise decisions under competence-preserving presentation changes; Mistral-7B-v0.3 reaches validity $0.644$ with a $41.4\%$ flip rate. Native chat formatting improves validity for several chat-tuned models but does not remove this instability. These results show that resume-screening evaluations should assess not only whether a system identifies stronger candidates, but also whether those decisions remain stable when the same competence evidence is presented differently.
Original Article
View Cached Full Text

Cached at: 09/16/26, 08:48 AM

# Competence-Preserving Resume Perturbations Expose Presentation Sensitivity in LLM Screening
Source: [https://arxiv.org/html/2609.16517](https://arxiv.org/html/2609.16517)
Yang XiaoAffiliation:The University of Melbourne

###### Abstract

Resume screeners must infer job\-relevant competence from resumes whose presentation can vary substantially in wording, structure, stylistic polish, and document extraction quality\. Ideally, such surface variation should not change decisions when the underlying qualification evidence is unchanged\. We introduce a controlled audit of this property, constructing occupation\-grounded candidate profiles at controlled competence levels and rendering each profile into multiple resume presentations\. A deterministic validation gate excludes variants that alter the underlying evidence before scoring\. Across six open instruction\-tuned LLM conditions, we find a clear disconnect between screening validity and presentation stability\. Llama\-3\.1\-8B with its native chat template achieves the strongest validity \(0\.7810\.781\) yet reverses29\.6%29\.6\\%of matched pairwise decisions under competence\-preserving presentation changes; Mistral\-7B\-v0\.3 reaches validity0\.6440\.644with a41\.4%41\.4\\%flip rate\. Native chat formatting improves validity for several chat\-tuned models but does not remove this instability\. These results show that resume\-screening evaluations should assess not only whether a system identifies stronger candidates, but also whether those decisions remain stable when the same competence evidence is presented differently\.

## 1Introduction

Automated resume screening requires systems to infer job\-relevant competence from highly variable natural\-language documents\. The same qualification evidence can appear in concise bullet points, paragraph\-style descriptions, AI\-polished prose, or text affected by document extraction and layout artifacts\. These differences are often orthogonal to whether a candidate is actually qualified, yet they change the surface form presented to the screening model\. A central challenge is therefore to distinguish variation in candidate competence from variation in how that competence is expressed\.

Algorithmic hiring has long raised questions about the validity, fairness, and reliability of automated decision systems\([Raghavan et al\., 2020](https://arxiv.org/html/2609.16517#bib.bib12);[Köchling and Wehner, 2020](https://arxiv.org/html/2609.16517#bib.bib13);[De\-Arteaga et al\., 2019](https://arxiv.org/html/2609.16517#bib.bib14)\)\. More recent work on LLM\-based hiring has examined screening validity\([Castleman et al\., 2026](https://arxiv.org/html/2609.16517#bib.bib1)\), demographic bias\([Gao et al\., 2026](https://arxiv.org/html/2609.16517#bib.bib2);[Iso et al\., 2025](https://arxiv.org/html/2609.16517#bib.bib15)\), self\-preference for AI\-written resumes\([Xu et al\., 2025](https://arxiv.org/html/2609.16517#bib.bib3)\), and the influence of AI recommendations on human screeners\([Wilson et al\., 2026](https://arxiv.org/html/2609.16517#bib.bib4)\)\. These studies address important questions about whether screening systems make useful or fair decisions, but leave open a complementary robustness question:*when the underlying competence evidence is unchanged, does the screening decision remain stable under changes in presentation?*This question is practically relevant because applicants differ in writing assistance, resume templates, formatting conventions, and document\-conversion pipelines\. Presentation variation is therefore not merely an artificial perturbation, but a natural source of heterogeneity in the inputs that screening systems receive\.

We study this problem through a controlled O\*NET\-backed audit that separates candidate competence from resume presentation\. Candidate profiles are constructed from occupation\-grounded rubrics at controlled competence levels and then rendered into multiple resume forms\. A deterministic validation gate removes variants that alter the underlying evidence before scoring\. This design allows us to evaluate two properties on the same candidate set: whether a screener recovers known\-superiority ordering \(*validity*\) and whether those decisions survive competence\-preserving presentation changes \(*presentation invariance*\)\. We find that the two properties can diverge sharply: Llama\-3\.1\-chat achieves validity0\.7810\.781yet reverses29\.6%29\.6\\%of matched decisions, while Mistral reaches validity0\.6440\.644with a41\.4%41\.4\\%flip rate\. Native chat formatting improves validity for several chat\-tuned models but does not eliminate this sensitivity, consistent with broader evidence that LLM behavior can depend strongly on prompt format and other semantically incidental prompt choices\([Sclar et al\., 2024](https://arxiv.org/html/2609.16517#bib.bib16);[Chatterjee et al\., 2024](https://arxiv.org/html/2609.16517#bib.bib17)\)\.

## 2Benchmark Framework

Our benchmark is designed to isolate a single question:does a resume screener preserve its decision when competence is fixed but presentation changes?To make this measurable, we separate benchmark construction into three stages as Figure[1](https://arxiv.org/html/2609.16517#S2.F1)\. We first instantiate occupation\-grounded candidate profiles with controlled competence differences, then render each profile into multiple presentation forms, and finally admit only variants that pass deterministic fact\-preservation checks\. Scorers operate on the validated resume text rather than hidden competence labels\. This design allows known\-superiority validity and presentation invariance to be evaluated on the same underlying candidate set\.

### 2\.1O\*NET\-Grounded Candidate Profiles

We construct the benchmark from the O\*NET 30\.3\([National Center for O\*NET Development, 2026](https://arxiv.org/html/2609.16517#bib.bib18)\)database, selecting 17 occupations spanning technical, administrative, customer\-facing, and health\-related roles\. For each occupation, we derive a screening rubric from its title, required skills, knowledge elements, and representative core tasks\. We then instantiate six synthetic candidate profiles for a total of 102 profiles, per occupation, with two qualified, two borderline, and two under\-qualified\. Profiles differ in rubric\-aligned skill evidence, years of experience, and achievement strength, providing controlled competence differences from which known\-superiority comparisons can be constructed\.

![Refer to caption](https://arxiv.org/html/2609.16517v1/figures/alta2.png)Figure 1:Benchmark construction and scoring flow\.This construction separates candidate competence from its eventual textual realization\. Competence is specified at the profile level before any resume is rendered, so multiple presentation variants can later be generated from the same underlying candidate record\. Known\-superiority pairs support the validity evaluation, while same\-tier pairs provide controlled equal\-competence comparisons\. This design allows presentation changes to be studied without redefining the underlying candidate qualifications\.

### 2\.2Competence\-Preserving Resume Perturbations

Each candidate profile is rendered into five resume forms: an original bullet\-style resume and four controlled presentation perturbations\.Verbosityadds redundant summary language around the same facts;structurereorganizes bullet\-style evidence into paragraph\-like prose;AI polishrewrites tone and fluency without introducing new skills; andlayoutsimulates extraction artifacts and local ordering noise\. Together, these axes vary length, discourse organization, stylistic polish, and document\-extraction quality while keeping the intended job\-relevant evidence fixed\. Across 102 candidate profiles, this process produces 510 resume variants\.

Because an intended rewrite is not necessarily competence\-preserving, every generated variant passes through a deterministic validation gate before scoring\. A variant is accepted only if required fact fields remain present, no unsupported rubric skill is introduced, no hidden competence\-tier label is exposed, and perturbation generation is independent of the downstream scorer\. Of the 510 variants, 505 pass validation\. For same\-tier base pairs, we additionally require exact agreement in the benchmark evidence vector \(required\-skill hits, task hits, and years of experience\)\. The 255 base candidate comparisons then expand across presentation axes into 1,250 validation\-approved pair rows\. For the invariance audit, each non\-original axis is matched to its corresponding original\-axis decision, yielding 1,000 flip comparisons\.

### 2\.3Independent Screening Protocol

Each resume is scored independently against the rubric of its corresponding occupation\. Given a resumerir\_\{i\}and occupation rubricqq, a screening function produces a scalar score\.

si=f⁡\(ri,q\)\.s\_\{i\}=f\(r\_\{i\},q\)\.\(1\)
For two candidatesiiandjj, the pairwise decision is computed only after both resumes have been scored:

D⁡\(i,j\)=sign⁡\(si−sj\)\.D\(i,j\)=\\operatorname\{sign\}\(s\_\{i\}\-s\_\{j\}\)\.\(2\)
The scorer therefore never receives two resumes in the same comparison prompt\. This avoids pair\-order and comparative prompt\-position effects in the known\-superiority evaluation\. It is also important for the invariance audit: an original\-to\-perturbed decision change reflects movement in independently assigned resume scores rather than a change in the surrounding pairwise prompt context\.

### 2\.4Evaluation Metrics

We evaluate screening systems along two primary dimensions:validityandpresentation invariance\.

#### Known\-superiority validity\.

For a known\-superiority pair\(i,j\)\(i,j\), where candidateiiis constructed to be stronger than candidatejj, validity measures whether the scorer assigns the stronger candidate the higher score:

Validity=1N∑\(i,j\)𝕀\[si\>sj\]\.\\mathrm\{Validity\}=\\frac\{1\}\{N\}\\sum\_\{\(i,j\)\}\\mathbb\{I\}\[s\_\{i\}\>s\_\{j\}\]\.Validity therefore measures recovery of the benchmark’s controlled competence ordering\.

#### Presentation flip rate\.

For each non\-original presentation axisaa, we compare its pairwise decisionDa​\(i,j\)D\_\{a\}\(i,j\)with the decision obtained from the corresponding original resumes,D0​\(i,j\)D\_\{0\}\(i,j\):

Flip=1M∑\(i,j,a\)𝕀\[Da\(i,j\)≠D0\(i,j\)\]\.\\mathrm\{Flip\}=\\frac\{1\}\{M\}\\sum\_\{\(i,j,a\)\}\\mathbb\{I\}\[D\_\{a\}\(i,j\)\\neq D\_\{0\}\(i,j\)\]\.\(3\)
A flip indicates that the pairwise decision changes when presentation changes while the underlying competence record is held fixed\. Flip rate measures*decision instability*, not error: a flip may either introduce an incorrect decision or correct an originally incorrect one\. We therefore treat validity and flip rate as complementary properties\.

#### Complementary stability metrics\.

For same\-tier pairs, we report the fraction assigned equal scores as the equal\-tier tie rate\. We also compute Kendall’sτ\\taubetween original\-axis and perturbed\-axis rankings within each occupation and average it across occupations\. Flip rate captures matched pairwise decision changes, whereas Kendall’sτ\\tauprovides an occupation\-level view of ranking stability\.

## 3Experimental Setup

#### Screening systems\.

We evaluate lexical baselines and direct LLM scorers under the validation\-gated benchmark\. The lexical baselines are BM25\([Robertson and Zaragoza, 2009](https://arxiv.org/html/2609.16517#bib.bib5)\)and TF–IDF\([Salton and Buckley, 1988](https://arxiv.org/html/2609.16517#bib.bib6)\)\. The direct LLM scorers span five model families: Qwen2\.5\-7B\([Qwen Team, 2025](https://arxiv.org/html/2609.16517#bib.bib7)\), Llama\-3\.1\-8B and Llama\-3\.2\-3B\([Meta AI, 2024](https://arxiv.org/html/2609.16517#bib.bib9)\), Phi\-3\.5\-mini\([Abdin et al\., 2024](https://arxiv.org/html/2609.16517#bib.bib10)\), Mistral\-7B\-v0\.3\([Jiang et al\., 2023](https://arxiv.org/html/2609.16517#bib.bib11)\), and Gemma\-2\-2B\([Gemma Team, 2024](https://arxiv.org/html/2609.16517#bib.bib8)\)\.

We additionally include a deterministic phrase\-preserving control that matches rubric skill and task phrases together with years of experience and aggregates the resulting evidence deterministically\.

#### Prompt conditions\.

We use raw prompts where available and native chat templates for Llama\-3\.1, Llama\-3\.2, and Gemma as a prompt\-format ablation\. The main table reports the stronger native\-template condition for Llama\-3\.2 and Gemma\.

## 4Results

Table 1:Validity and presentation stability across screening systems\. Higher validity indicates better recovery of known\-superiority ordering, while lower flip rate indicates greater stability under competence\-preserving presentation changes\.### 4\.1Validity and Presentation Invariance

Table[1](https://arxiv.org/html/2609.16517#S4.T1)shows that validity and presentation invariance capture distinct properties of a screening system\. The strongest direct LLM condition is Llama\-3\.1\-8B with its native chat template, which achieves validity 0\.781\. Under paired occupation\-cluster bootstrap, it exceeds BM25 by 0\.293 and TF–IDF by 0\.305\. Mistral\-7B\-v0\.3 also clears both lexical baselines, reaching validity 0\.644\. Phi\-3\.5\-mini clears TF–IDF but not BM25, while the remaining direct LLM conditions provide directional rather than confirmatory validity evidence\.

Stronger validity, however, does not imply presentation invariance\. Llama\-3\.1\-chat reverses29\.6%29\.6\\%of matched pairwise decisions under competence\-preserving presentation changes, while Mistral flips41\.4%41\.4\\%\. The same pattern is visible across the broader direct LLM set, with flip rates ranging from 0\.285 to 0\.456\. Thus, even systems that recover known\-superiority ordering relatively well can remain substantially sensitive to how the same competence evidence is presented\.

This separation motivates treating validity and presentation invariance as complementary evaluation dimensions\. Validity measures whether a screener recovers the intended competence ordering, whereas invariance measures whether that decision survives irrelevant surface variation\. High invariance alone is therefore not sufficient: a system may preserve consistently poor decisions\. The desirable regime is one in which screening decisions are both valid and stable\.

BM25 and TF–IDF show much lower flip rates \(0\.040 and 0\.052\), but these values should not be interpreted as evidence that lexical matching is generally a better resume screener\. The benchmark deliberately preserves many rubric phrases across presentation variants, allowing lexical systems to retain the same pairwise sign under this controlled phrase\-preserving condition\. They therefore serve as stability references rather than deployment\-quality screening methods\. The deterministic structured control is narrower still: its validity 1\.000 and flip rate 0\.000 arise by construction and serve only as a benchmark sanity check\.

### 4\.2Prompt\-Format Ablation

A potential confound is that chat\-tuned models may be miscalibrated under raw completion prompts\. Table[2](https://arxiv.org/html/2609.16517#S4.T2)tests this explanation using native chat templates for Llama\-3\.1, Llama\-3\.2, and Gemma\. Chat formatting substantially improves validity: Llama\-3\.1 rises from 0\.622 to 0\.781, Llama\-3\.2 from 0\.321 to 0\.551, and Gemma from 0\.287 to 0\.576\. Weak raw\-prompt validity should therefore not be interpreted as an intrinsic inability of these chat\-tuned models to perform the task\.

The presentation\-sensitivity result nevertheless survives this correction\. Native\-template flip rates remain 0\.296 for Llama\-3\.1, 0\.285 for Llama\-3\.2, and 0\.456 for Gemma\. Prompt\-format mismatch can therefore explain part of the validity degradation under raw prompting, but it is not a sufficient explanation for presentation instability\.

Table 2:Native chat\-template ablation\. Raw completion prompts can understate validity for chat\-tuned models, but presentation sensitivity remains high after the correction\.

## 5Conclusion

We introduced a controlled audit for studying how resume\-screening systems respond to presentation variation when job\-relevant competence is held fixed\. By separating candidate competence from resume realization, the benchmark makes it possible to evaluate screening validity and presentation stability on the same underlying candidates\. Our results show that these properties can diverge substantially: the strongest\-validity LLM conditions still reverse a large fraction of decisions under competence\-preserving rewrites, and native chat formatting does not eliminate the effect\. These findings highlight presentation stability as an important dimension of resume\-screening evaluation\. Future audits should therefore examine not only whether a system ranks stronger candidates correctly, but also whether those decisions remain consistent across equivalent forms of the same evidence\.

## 6Limitations

Our controlled construction trades breadth for attribution: the benchmark covers 17 occupations and 102 synthetic O\*NET\-grounded candidates, allowing presentation effects to be isolated while competence evidence is held fixed\. Extending the audit to more occupations, human\-authored resumes, and production resume\-processing pipelines would test how well these findings generalize beyond the controlled setting\. The occupation\-cluster analysis is also based on 17 clusters, so its bootstrap intervals are best interpreted as pilot\-scale evidence rather than precise population estimates\. Finally, the evaluated systems are open\-weight instruction models; broader audits could include proprietary screening models and end\-to\-end applicant\-tracking systems\.

## References

- Abdinet al\.\(2024\)M\. Abdin, S\. A\. Jacobs, A\. A\. Awan, J\. Aneja, A\. Awadallah, H\. Awadalla, N\. Bach, A\. Bahree, A\. Bakhtiari, H\. Behl,et al\.Phi\-3 technical report: a highly capable language model locally on your phone\.External Links:2404\.14219Cited by:[§3](https://arxiv.org/html/2609.16517#S3.SS0.SSS0.Px1.p1.1)\.
- Castlemanet al\.\(2026\)J\. Castleman, Z\. Shen, B\. Metevier, M\. Springer, and A\. KorolovaMeasuring validity in LLM\-based resume screening\.External Links:2602\.18550Cited by:[§1](https://arxiv.org/html/2609.16517#S1.p2.1)\.
- Chatterjeeet al\.\(2024\)A\. Chatterjee, H\. S\. V\. N\. S\. K\. Renduchintala, S\. Bhatia, and T\. ChakrabortyPOSIX: a prompt sensitivity index for large language models\.InFindings of the Association for Computational Linguistics: EMNLP 2024,Miami, Florida, USA,pp\. 14550–14565\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.findings-emnlp.852),[Link](https://aclanthology.org/2024.findings-emnlp.852/)Cited by:[§1](https://arxiv.org/html/2609.16517#S1.p3.1)\.
- De\-Arteagaet al\.\(2019\)M\. De\-Arteaga, A\. Romanov, H\. Wallach, J\. Chayes, C\. Borgs, A\. Chouldechova, S\. Geyik, K\. Kenthapadi, and A\. T\. KalaiBias in bios: a case study of semantic representation bias in a high\-stakes setting\.InProceedings of the Conference on Fairness, Accountability, and Transparency,pp\. 120–128\.External Links:[Document](https://dx.doi.org/10.1145/3287560.3287572),[Link](https://doi.org/10.1145/3287560.3287572)Cited by:[§1](https://arxiv.org/html/2609.16517#S1.p2.1)\.
- Gaoet al\.\(2026\)Z\. Gao, W\. Jiang, and Y\. YanCan LLMs hire fairly? racial bias in resume screening\.External Links:2606\.28978Cited by:[§1](https://arxiv.org/html/2609.16517#S1.p2.1)\.
- Gemma Team \(2024\)Gemma TeamGemma: open models based on Gemini research and technology\.External Links:2403\.08295Cited by:[§3](https://arxiv.org/html/2609.16517#S3.SS0.SSS0.Px1.p1.1)\.
- Isoet al\.\(2025\)H\. Iso, P\. Pezeshkpour, N\. Bhutani, and E\. HruschkaEvaluating bias in LLMs for job\-resume matching: gender, race, and education\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 3: Industry Track\),Albuquerque, New Mexico,pp\. 672–683\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.naacl-industry.55),[Link](https://aclanthology.org/2025.naacl-industry.55/)Cited by:[§1](https://arxiv.org/html/2609.16517#S1.p2.1)\.
- Jianget al\.\(2023\)A\. Q\. Jiang, A\. Sablayrolles, A\. Mensch, C\. Bamford, D\. S\. Chaplot, D\. d\. L\. Casas, F\. Bressand, G\. Lengyel, G\. Lample, L\. Saulnier,et al\.Mistral 7b\.External Links:2310\.06825Cited by:[§3](https://arxiv.org/html/2609.16517#S3.SS0.SSS0.Px1.p1.1)\.
- Köchling and Wehner \(2020\)A\. Köchling and M\. C\. WehnerDiscriminated by an algorithm: a systematic review of discrimination and fairness by algorithmic decision\-making in the context of hr recruitment and hr development\.Business Research13,pp\. 795–848\.External Links:[Document](https://dx.doi.org/10.1007/s40685-020-00134-w),[Link](https://doi.org/10.1007/s40685-020-00134-w)Cited by:[§1](https://arxiv.org/html/2609.16517#S1.p2.1)\.
- Meta AI \(2024\)Meta AIThe Llama 3 herd of models\.External Links:2407\.21783Cited by:[§3](https://arxiv.org/html/2609.16517#S3.SS0.SSS0.Px1.p1.1)\.
- National Center for O\*NET Development \(2026\)National Center for O\*NET DevelopmentO\*NET 30\.3 Database\.Note:O\*NET Resource CenterVersion 30\.3External Links:[Link](https://www.onetcenter.org/database.html)Cited by:[§2\.1](https://arxiv.org/html/2609.16517#S2.SS1.p1.1)\.
- Qwen Team \(2025\)Qwen TeamQwen2\.5 technical report\.External Links:2412\.15115Cited by:[§3](https://arxiv.org/html/2609.16517#S3.SS0.SSS0.Px1.p1.1)\.
- Raghavanet al\.\(2020\)M\. Raghavan, S\. Barocas, J\. Kleinberg, and K\. LevyMitigating bias in algorithmic hiring: evaluating claims and practices\.InProceedings of the 2020 Conference on Fairness, Accountability, and Transparency,pp\. 469–481\.External Links:[Document](https://dx.doi.org/10.1145/3351095.3372828),[Link](https://doi.org/10.1145/3351095.3372828)Cited by:[§1](https://arxiv.org/html/2609.16517#S1.p2.1)\.
- Robertson and Zaragoza \(2009\)S\. Robertson and H\. ZaragozaThe probabilistic relevance framework: BM25 and beyond\.InFoundations and Trends in Information Retrieval,Vol\.3,pp\. 333–389\.Cited by:[§3](https://arxiv.org/html/2609.16517#S3.SS0.SSS0.Px1.p1.1)\.
- Salton and Buckley \(1988\)G\. Salton and C\. BuckleyTerm\-weighting approaches in automatic text retrieval\.Information Processing & Management24\(5\),pp\. 513–523\.Cited by:[§3](https://arxiv.org/html/2609.16517#S3.SS0.SSS0.Px1.p1.1)\.
- Sclaret al\.\(2024\)M\. Sclar, Y\. Choi, Y\. Tsvetkov, and A\. SuhrQuantifying language models’ sensitivity to spurious features in prompt design or: how i learned to start worrying about prompt formatting\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/)Cited by:[§1](https://arxiv.org/html/2609.16517#S1.p3.1)\.
- Wilsonet al\.\(2026\)K\. Wilson, M\. Sim, A\. Gueorguieva, S\. Chatterjee, and A\. CaliskanResume screening, fast and slow: \(biased\) AI recommendations’ influence on human decision making\.External Links:2606\.22213Cited by:[§1](https://arxiv.org/html/2609.16517#S1.p2.1)\.
- Xuet al\.\(2025\)J\. Xu, G\. Li, and J\. Y\. JiangAI self\-preferencing in algorithmic hiring: empirical evidence and insights\.External Links:2509\.00462Cited by:[§1](https://arxiv.org/html/2609.16517#S1.p2.1)\.

## Appendix AImplementation Details

Phi\-3\.5\-mini and Mistral\-7B\-v0\.3 checkpoints are obtained from Hugging Face\. Llama, Phi, Mistral, and the evaluated chat\-template conditions are executed through vLLM on a single NVIDIA L40S GPU with 48 GB of memory\. Each resume is scored independently against its corresponding occupation rubric; no model receives both resumes from a comparison pair in the same prompt\. Pairwise decisions are formed only after the individual scalar scores have been produced, avoiding pair\-order and comparative prompt\-position effects\.

For uncertainty estimation, we report both row\-level bootstrap intervals and occupation\-cluster percentile bootstrap intervals using 10,000 resamples\. The cluster bootstrap resamples the 17 occupations and serves as a robustness check against row\-level pseudo\-replication\. Because the benchmark contains only 17 occupation clusters, we interpret these intervals as exploratory small\-cluster evidence rather than precise nominal population coverage\.

## Appendix BAxis\-Level Presentation Sensitivity

As a directional failure\-analysis slice, Figure[2](https://arxiv.org/html/2609.16517#A2.F2)decomposes Qwen2\.5\-7B raw flips by presentation axis\. Layout artifacts produce the highest flip rate \(0\.472\), followed by AI\-polished prose \(0\.404\), structural reorganization \(0\.380\), and verbosity \(0\.252\)\. Because Qwen’s validity advantage over the lexical baselines does not clear zero under paired occupation\-cluster bootstrap, we treat this decomposition as an illustrative diagnostic rather than evidence that these axes have a universal ordering of difficulty\.

Figure 2:Axis\-level flip\-rate diagnostic for a directional raw\-model slice\. AI\-polished prose and layout artifacts are among the most unstable axes in this slice\.

Similar Articles

I analyzed 25,500 LLM resume screenings to measure hiring bias. The results are a wake-up call.

Reddit r/artificial

A study analyzing 25,500 LLM resume evaluations across 10 models found a 45% bias rate driven by 'silent bias', with models inventing professional-sounding excuses to penalize candidates. It highlights significant variability in fairness and stability, with Claude, Mistral-Large, and Llama 4 being most stable, while Qwen and older Gemini models were volatile.

Can LLMs Hire Fairly? Racial Bias in Resume Screening

arXiv cs.CL

This paper audits 14 large language models for hiring discrimination using a paired-resume methodology, finding that older models exhibit pro-White bias while newer models show null or pro-Black bias, indicating a reversal in algorithmic hiring bias across model generations.