Disentangling Algorithmic Bias from Archival Artifacts: A Controlled Audit of Vision-Language Model Valuation in Metropolitan Museum Archives
Summary
This study audits CLIP models for gender bias in Metropolitan Museum artwork metadata, finding no statistically significant bias but emphasizing the need for multivariate confound control in AI fairness assessments.
View Cached Full Text
Cached at: 09/17/26, 08:42 AM
# Disentangling Algorithmic Bias from Archival Artifacts: A Controlled Audit of Vision-Language Model Valuation in Metropolitan Museum Archives
Source: [https://arxiv.org/html/2609.17572](https://arxiv.org/html/2609.17572)
\[1\]\\fnmManpreet\\surSingh
\[1\]\\orgnameBoston University,\\orgaddress\\cityBoston,\\stateMA,\\countryUSA 2\]\\orgnameUniversity of Eastern Finland,\\countryFinland 3\]\\orgnameSymbiosis Institute of Technology, Pune, Symbiosis International \(Deemed University\),\\orgaddress\\cityPune,\\countryIndia
###### Abstract
Auditing vision\-language models \(VLMs\) for societal bias requires distinguishing direct algorithmic valuation disparities from confounders embedded within archival metadata\. In this study, we audit Contrastive Language\-Image Pre\-training \(CLIP\) models using historical artwork metadata harvested from the Metropolitan Museum of Art Open Access collection \(N=1,500N=1,500total objects;N=743N=743attributed works: Malen=534n=534, Femalen=209n=209;n=618n=618anonymous\)\. We establish a quantitative audit framework evaluating zero\-shot CLIP logit differential scores across three semantic prompt pairs \(masterpiece,quality, andinfluence\)\. Unadjusted evaluations demonstrate high score convergence without a statistically significant main gender effect under OpenAI CLIP \(μF=−0\.0067\\mu\_\{F\}=\-0\.0067vsμM=−0\.0035,p=0\.1829\\mu\_\{M\}=\-0\.0035,p=0\.1829\) or OpenCLIP \(μF=0\.0171\\mu\_\{F\}=0\.0171vsμM=0\.0237,p=0\.1224\\mu\_\{M\}=0\.0237,p=0\.1224\)\. Two One\-Sided Tests \(TOST\) confirm statistical equivalence across Cohen’sd≥0\.25d\\geq 0\.25bounds \(pTOST<0\.005p\_\{\\text\{TOST\}\}<0\.005\)\. Multivariate OLS regression controlling for artwork medium, creation era, and aspect ratio \(R2<0\.02R^\{2\}<0\.02\) confirms that artist gender has no statistically significant conditional effect \(p\>0\.20p\>0\.20\)\. High residual embedding variance \(R2<2%R^\{2\}<2\\%\) indicates that global zero\-shot valuation metrics operate near an embedding noise floor, demonstrating that broad zero\-shot prompt logit differentials function as a coarse, insensitive measurement instrument for visual art evaluation rather than proving absolute model fairness\. We highlight two key caveats: \(i\) macro\-level score equivalence reflects metric insensitivity to fine\-grained visual\-semantic features and does not preclude localized micro\-level visual biases, and \(ii\) excluding 41\.2% unattributed holdings reflects institutional survival bias\. These results demonstrate the necessity of multivariate confound control, equivalence testing, and archival provenance auditing when assessing AI fairness in cultural heritage collections\.
###### keywords:
Vision\-Language Models, Algorithmic Bias Audits, Museum Archives, Digital Humanities, Confound Control, Responsible AI
## 1Introduction
### 1\.1Background and Motivation
Large\-scale vision\-language models \(VLMs\), such as Contrastive Language\-Image Pre\-training \(CLIP\)\[radford2021learning\], are increasingly deployed across digital humanities, automated collection cataloging, and computational art history\[fiorucci2020machine,garcia2020bias\]\. Because these models align visual features with text representations learned from uncurated web scrapes, concern has grown that they inherit and amplify historical socio\-cultural biases\[birhane2021multimodal,wolfe2022evidence,bender2021stochastic\]\. In cultural heritage institutions, historical collections already reflect systemic representational imbalances, including severe gender and regional disparities in acquisition and archival documentation\[topaz2019diversity,meier2021gender,carlson2022museum,bailey2020gender,nochlin1971why,pollock1988vision\]\.
When auditing VLMs for representational bias, a critical methodological challenge arises: distinguishing direct algorithmic valuation bias from confounding variables embedded within archival metadata\[hall2023auditing,buolamwini2018gender,noorthuis2020bias\]\. Grounded in established Responsible AI governance frameworks\[dignum2019responsible,stahl2021responsible,jobin2019global\]and critical studies of classification infrastructures\[bowker2000sorting,noble2018algorithms,pasquale2015black\], evaluating AI fairness in cultural archives requires controlled quantitative audit pipelines capable of isolating demographic main effects from structural collection artifacts\. In critical archival studies, classification taxonomies are recognized not as neutral administrative tools, but as sociotechnical infrastructures that embed historical power relations, institutional boundaries, and gendered labor exclusions\[bowker2000sorting,noble2018algorithms\]\. If female artists in a historical collection are disproportionately represented in specific media \(such as textiles or works on paper\) relative to male artists \(who dominate large\-scale oil paintings or monumental sculptures\), an unadjusted evaluation of visual\-text similarity scores may misattribute medium\-specific model behaviors to gender bias\. Embedding multivariate confound control directly into Responsible AI audit architectures\[mitchell2019model,stahl2021responsible\]ensures that institutional AI deployments do not misinterpret historical curation artifacts as active model evaluation bias\.
Furthermore, the rapid integration of zero\-shot vision\-language classifiers into public search systems, digital library indexing, and collection exploration interfaces makes this distinction practical rather than purely theoretical\[srinivasan2021arts,luccioni2023stable\]\. If a museum search index uses uncalibrated CLIP embeddings to surface ”masterpieces” or ”influential works,” model biases interacting with archival metadata gaps could systematically deprioritize works by historically underrepresented artists\[zhao2017men,bianchi2023easily\]\. Addressing these risks requires systematic empirical audits that evaluate both raw model outputs and multivariate confound interactions across diverse pretraining architectures\.
### 1\.2Research Questions and Core Contributions
This study investigates how vision\-language models evaluate historical artworks and whether observed score disparities reflect direct artist gender bias or underlying archival confounders\. We address three central research questions:
1. 1\.RQ1:Do CLIP visual\-text similarity scores assign significantly lower aesthetic valuation to artworks created by female artists compared to male artists in public museum collections?
2. 2\.RQ2:How robust are algorithmic valuation scores across distinct semantic prompt formulations operationalizing artistic importance?
3. 3\.RQ3:When controlling for structural archival confounders such as artwork medium, historical creation era, and aspect ratio, does artist gender remain a statistically significant predictor of model evaluation?
To answer these questions, we audit a complete corpus of historical artworks \(N=1,500N=1,500total records:N=743N=743attributed works, Malen=534n=534, Femalen=209n=209;n=618n=618anonymous/unattributed\) harvested from the Metropolitan Museum of Art Open Access API\[met\_api\_2023,garcia2020bias\]\. We construct a multi\-tiered quantitative framework that measures CLIP softmax probability differential scores across three distinct prompt pairs \(masterpiece,quality, andinfluence\)\. We apply non\-parametric hypothesis testing \(Mann\-WhitneyUU\), rank\-biserial effect sizes \(rr\), bootstrapped 95% confidence intervals, and multivariate Ordinary Least Squares \(OLS\) regression with HC3 heteroskedasticity\-robust standard errors across two model pretraining regimes: OpenAI CLIP \(ViT\-B/32, trained on curated WIT\) and OpenCLIP \(ViT\-B/32, trained on uncurated LAION\-2B\)\[cherti2023reproducible,schuhmann2022laion\]\.
Our primary empirical contributions include:
- •Unadjusted evaluations across the completeN=743N=743attributed corpus demonstrate high score convergence without a statistically significant composite gender valuation gap between female and male artists under OpenAI CLIP \(μF=−0\.0067\\mu\_\{F\}=\-0\.0067vsμM=−0\.0035,U=59,307\.00,p=0\.1829,r=−0\.0628\\mu\_\{M\}=\-0\.0035,U=59,307\.00,p=0\.1829,r=\-0\.0628\) or OpenCLIP \(μF=0\.0171\\mu\_\{F\}=0\.0171vsμM=0\.0237,U=59,867\.00,p=0\.1224,r=−0\.0728\\mu\_\{M\}=0\.0237,U=59,867\.00,p=0\.1224,r=\-0\.0728\)\.
- •Two One\-Sided Tests \(TOST\) establish formal statistical equivalence between male and female artwork score distributions across Cohen’sd≥0\.25d\\geq 0\.25equivalence bounds \(pTOST=0\.0042p\_\{\\text\{TOST\}\}=0\.0042atd=0\.30d=0\.30for OpenAI CLIP;pTOST=0\.0024p\_\{\\text\{TOST\}\}=0\.0024for OpenCLIP\), supported by sensitivity checks across boundsd∈\[0\.15,0\.40\]d\\in\[0\.15,0\.40\]\.
- •Prompt sensitivity checks demonstrate robust consistency across all semantic descriptors \(masterpiece,quality, andinfluence\), with no statistically significant individual prompt shifts surviving baseline controls or Bonferroni adjustment\.
- •Multivariate OLS regression controlling for medium, creation century, and framing aspect ratio \(R2=0\.018,F\(11,731\)=0\.9286,p=0\.512R^\{2\}=0\.018,F\(11,731\)=0\.9286,p=0\.512for CLIP;R2=0\.017,F\(11,731\)=0\.9888,p=0\.455R^\{2\}=0\.017,F\(11,731\)=0\.9888,p=0\.455for OpenCLIP\) confirms that artist gender \(B=0\.0037,p=0\.202B=0\.0037,p=0\.202for OpenAI CLIP;B=0\.0065,p=0\.356B=0\.0065,p=0\.356for OpenCLIP\) and aspect ratio \(B=−0\.0064,p=0\.141B=\-0\.0064,p=0\.141for CLIP;B=−0\.0023,p=0\.516B=\-0\.0023,p=0\.516for OpenCLIP\) exhibit non\-significant conditional effects\. Low totalR2R^\{2\}\(<2%<2\\%\) reflects high residual embedding variance, demonstrating that zero\-shot global prompt metrics operate near a noise floor where macro\-level evaluations are dominated by embedding variance rather than systematic demographic bias\.
- •We contextualize these findings within institutional archival boundaries, emphasizing theArchival Survival Bias Paradox: filtering out41\.2%41\.2\\%\(n=618n=618\) unattributed objects introduces dataset selection bias by auditing only named creators who already survived institutional gatekeeping while omitting the archival strata where female domestic craft labor was historically erased\.
### 1\.3Scope and Target Framework
This work aligns directly with Responsible AI governance frameworks in archival and digital humanities practice\[dignum2019responsible,stahl2021responsible,mitchell2019model,jobin2019global\]\. By demonstrating that apparent algorithmic disparities can stem from archival curation artifacts rather than standalone model evaluation bias, we highlight multivariate confound control as an essential requirement for auditing AI across cultural heritage lifecycles\.
The remainder of this paper is organized as follows: Section[2](https://arxiv.org/html/2609.17572#S2)reviews literature on archival bias, vision\-language model evaluation, and confound control\. Section[3](https://arxiv.org/html/2609.17572#S3)describes the Metropolitan Museum metadata collection and archival audit findings\. Section[4](https://arxiv.org/html/2609.17572#S4)details the CLIP scoring metrics, statistical hypothesis testing, and OLS regression framework\. Section[5](https://arxiv.org/html/2609.17572#S5)presents empirical results\. Section[6](https://arxiv.org/html/2609.17572#S6)discusses implications for archival AI deployment, and Section[7](https://arxiv.org/html/2609.17572#S7)concludes\.
## 2Related Work
### 2\.1Representational Disparities in Cultural Heritage Archives
Cultural heritage archives and museum collections reflect historical patterns of institutional acquisition, patronage, and societal exclusion\[carlson2022museum,meier2021gender,bailey2020gender\]\. Quantitative surveys of major Western art archives demonstrate substantial representational imbalances across artist gender and geographic origin\[topaz2019diversity\]\. Systemic analysis of holdings across prominent U\.S\. art museums indicates that over 85% of cataloged artists are male and predominantly Euro\-American\[topaz2019diversity\]\. These structural imbalances stem from historical barriers to institutional access, restricted entry to royal academies, patron preferences, and archival documentation practices where metadata fields for marginalized groups remain sparse or unindexed\[garcia2020bias,noorthuis2020bias\]\.
Digitization of museum collections via public Application Programming Interfaces \(APIs\)—such as those provided by the Metropolitan Museum of Art and Europeana—has enabled large\-scale computational art history and digital humanities research\[fiorucci2020machine,met\_api\_2023\]\. However, digitized metadata carries forward the structural biases of physical archives\. When computational pipelines process museum APIs without accounting for historical metadata gaps, representational asymmetries risk being interpreted as inherent features of cultural production rather than artifacts of institutional curation\[dignum2019responsible\]\.
In critical archival studies, classification systems are recognized not as neutral administrative tools, but as sociotechnical infrastructures that embed historical power relations and institutional boundaries\[bowker2000sorting,noble2018algorithms,caswell2017archival\]\. When vision\-language models process digitized museum metadata, they do not merely extract visual features; they re\-encode historical taxonomic classifications that historically prioritized canonical Western fine art \(such as oil paintings and monumental sculpture\) over decorative, domestic, and textile media\[bailey2020gender,bowker2000sorting\]\. Evaluating AI fairness in digital archives therefore requires inspecting how classification infrastructures interact with pretraining visual\-text representations\.
### 2\.2Algorithmic Valuation in Vision\-Language Models
Contrastive Language\-Image Pre\-training \(CLIP\)\[radford2021learning\]and related multimodal vision\-language models \(VLMs\) align visual and text representations by joint optimization over web\-scraped image\-text pairs\[he2016deep,dosovitskiy2020image\]\. While CLIP achieves strong zero\-shot classification performance, pretraining datasets ingest pervasive social stereotypes, demographic biases, and valuation skews\[birhane2021multimodal,wolfe2022evidence,bender2021stochastic,steed2021image\]\.
In visual domain audits, VLMs exhibit bias when evaluating human traits, professional roles, and aesthetic concepts\[hall2023auditing,wang2022revisiting,luccioni2023stable\]\. Specifically, when prompted with value\-laden terms \(such asmasterpiece,high quality, orinfluential\), VLMs compute cosine similarities that reflect learned cultural associations between visual features and subjective value judgments\[garcia2020bias,bianchi2023easily\]\. When applied to cultural artifacts, these models risk operationalizing normative value judgments that privilege traditional Western fine art over decorative arts, textiles, or non\-canonical media, indirectly penalizing artist demographics historically associated with those media\[wolfe2022evidence,zhao2017men\]\.
### 2\.3Confound Control in Multimodal AI Audits
A growing body of algorithmic fairness literature highlights the risk of confounding in observational AI audits\[buolamwini2018gender,mitchell2019model,bolukbasi2016man\]\. Observational evaluations that compare raw model score outputs across demographic groups without controlling for correlated structural covariates often report spurious bias metrics\[hall2023auditing\]\. Early audits in facial analysis demonstrated that classification disparities were driven by lighting and skin tone interactions rather than standalone gender predictors\[buolamwini2018gender\]\.
In digital heritage, observational evaluations of model behavior face severe confounding from physical artwork attributes, including artistic medium, physical dimensions, historical era, and digitized image resolution\[fiorucci2020machine,meier2021gender\]\. Medium distributions in historical collections are non\-randomly correlated with artist gender due to historical restrictions on female artists’ access to specific materials and academies\[topaz2019diversity,bailey2020gender\]\. Consequently, an unadjusted comparison of model valuation scores across artist gender risks confusing medium\-specific visual feature scoring \(e\.g\., model response to 3D sculptures versus 2D canvas paintings\) with direct demographic bias\[hall2023auditing\]\. Establishing rigorous audit methodologies requires multivariate confound control frameworks—such as Ordinary Least Squares \(OLS\) regression—to isolate demographic main effects from structural collection artifacts\.
## 3Data Collection and Archival Audit
This section describes the data acquisition pipeline, metadata enrichment methodology, data validation controls, and an empirical representation audit of the Metropolitan Museum of Art’s public collection archives\.
### 3\.1Corpus Acquisition via RESTful Museum APIs
Primary data was harvested from the Metropolitan Museum of Art Open Access RESTful API\[met\_api\_2023\]\. The API provides machine\-readable access to over470,000470,000cultural heritage artifacts in the public domain\. To maximize representational diversity across visual art forms with established canon histories, we harvested objects from nine curatorial divisions: American Decorative Arts \(Dept\. 1\), Asian Art \(Dept\. 6\), The Costume Institute \(Dept\. 8\), Drawings and Prints \(Dept\. 9\), European Paintings \(Dept\. 11\), European Sculpture and Decorative Arts \(Dept\. 12\), Musical Instruments \(Dept\. 15\), Photographs \(Dept\. 19\), and Modern and Contemporary Art \(Dept\. 21\)\.
The harvesting pipeline queried object endpoints sequentially under HTTP rate\-limiting controls \(1010requests per second\) with local JSON response caching \(‘\.met\_cache‘\) to ensure experiment reproducibility\. Inclusion criteria mandated: \(i\) ‘isPublicDomain == True‘, \(ii\) a non\-null high\-resolution visual asset URL \(‘primaryImageSmall‘\), and \(iii\) basic cataloging metadata \(‘objectID‘, ‘title‘, ‘medium‘, ‘objectBeginDate‘\)\. A complete raw corpus ofN=1,500N=1,500primary artwork records was harvested and processed\.
### 3\.2Metadata Enrichment and Demographic Categorization
Museum archive catalogs frequently suffer from incomplete or unstandardized demographic records due to legacy curatorial documentation practices\[carlson2022museum,garcia2020bias,bailey2020gender\]\. To resolve these gaps, we engineered a multi\-stage deterministic enrichment pipeline:
#### 3\.2\.1Artist Gender Resolution
We first query the Met’s explicit ‘artistGender‘ field\. In the Met Open Access dataset, this field is populated primarily for female artists \(’Female’\) and left unpopulated otherwise\. Blanks are treated as non\-informative fall\-through to name\-based inference rather than assumed male attributions\. When ‘artistGender‘ is unpopulated, we parse ‘artistDisplayName‘, stripping delimiters, honorifics, and qualifying attribution prefixes\. In historical art museum cataloging, attributions frequently include qualifying studio descriptors—such as”Workshop of…”,”Studio of…”,”Circle of…”,”Follower of…”, or”Manner of…”\. Our deterministic pre\-processing pipeline strips these qualifying prefixes to isolate individual artist name tokens while routing ambiguous collective attributions \(e\.g\.,”Unidentified 17th Century Master”,”Flemish Painter”\) directly to theUnknowncategory \(41\.20%41\.20\\%,n=618n=618\)\. For named master workshops \(e\.g\.,”Circle of Artemisia Gentileschi”\), the extracted master name token is evaluated under dictionary lookup\. The extracted first name token is evaluated using ‘gender\-guesser‘\[gender\_guesser\], a dictionary detector based on Jörg Michael’s ‘gender\.c‘ database\. Names classified as ‘male‘/‘mostly\_male‘ or ‘female‘/‘mostly\_female‘ are assigned accordingly; unisex, unknown, or non\-Western names are assigned to anUnknowncategory \(41\.20%41\.20\\%of raw objects,n=618n=618\)\.
The dictionary lookup evaluates explicit first\-name gender associations, mapping unambiguously recognized male or female tokens while routing ambiguous strings, unisex names, pre\-18th\-century Latinized variants \(e\.g\.,Johannes,Nicolaes\), and patronymic mononyms \(e\.g\.,di Bondone\) directly to theUnknowncategory \(41\.20%41\.20\\%,n=618n=618\)\. This conservative rule prevents false\-positive demographic attributions at the expense of catalog coverage\. Categorical breakdown of this excluded anonymous cohort \(n=618n=618\) reveals heavy concentration in textiles \(38\.2%38\.2\\%,n=236n=236\), decorative ceramics and woodwork \(31\.6%31\.6\\%,n=195n=195\), graphics/prints \(18\.4%18\.4\\%,n=114n=114\), and unclassified domestic items \(11\.8%11\.8\\%,n=73n=73\)\. To empirically benchmark the accuracy ofgender\-guesseron this historical corpus, we manually audited a stratified random subsample ofn=100n=100attributed creators \(50 male\-inferred, 50 female\-inferred\) against verified Union List of Artist Names \(ULAN\) and Getty biographical records\. The manual audit confirmed96\.0%96\.0\\%overall classification accuracy \(96/10096/100\), with2\.0%2\.0\\%false\-positive female attributions \(primarily from Latinized diminutive suffixes\) and2\.0%2\.0\\%false\-positive male attributions \(driven by non\-Western transliterated mononyms\)\. Stratifying accuracy across creation eras demonstrates era\-dependent performance variance:90\.0%90\.0\\%accuracy \(18/2018/20\) for15th15^\{\\text\{th\}\}–16th16^\{\\text\{th\}\}century creators \(driven by Latinized mononyms and guild patronyms\),96\.0%96\.0\\%\(48/5048/50\) for17th17^\{\\text\{th\}\}–18th18^\{\\text\{th\}\}century creators, and98\.0%98\.0\\%\(29/3029/30\) for19th19^\{\\text\{th\}\}–20th20^\{\\text\{th\}\}century catalog entries\.
Applying binary automated gender recognition to historical creators carries inherent ethical and validity constraints\[keyes2018misgendering\]\. Name\-based classification imposes a contemporary binary schema on historical individuals whose self\-identifications or archival records may not align with modern taxonomy, and exhibits higher error rates on non\-Western or Latinized naming conventions\. Crucially, the presence of≈4\.0%\\approx 4\.0\\%measurement error in automated binary regressor attributions introduces classical regressor measurement error, which theoretically induces slight attenuation bias \(B^1→0\\hat\{B\}\_\{1\}\\to 0\) in OLS slope estimation\. Furthermore, accounting for41\.20%41\.20\\%\(n=618n=618\) unattributed holdings reflects structural power dynamics in physical archive curation\[bowker2000sorting,carlson2022museum,bailey2020gender,topaz2019diversity,parker1984subversive\]\. In historical European collections, female creators were systematically denied guild membership, forcing production into domestic workshops cataloged anonymously or under male family heads\[bailey2020gender,parker1984subversive\]\. We term this theArchival Survival Bias Paradox: filtering out unattributed objects to construct named audit cohorts inherently introduces dataset selection bias by evaluating a sanitized survival subset of named creators while removing the very archival strata where female domestic labor was historically erased\. Name\-based AI auditing pipelines must acknowledge this uneliminable boundary condition in institutional data governance\. Retaining an explicit41\.20%41\.20\\%\(n=618n=618\)Unknownbuffer ensures that ambiguous, collective, or non\-binary historical attributions are not force\-fitted into binary categories, maintaining conservative data ethics standards in archival research\. Additionally, thegender\-guessertool collapses androgynous\-classified names with truly unknown entries into a singleUnknowncategory; future work should disaggregate these subcategories to assess whether androgynous\-named artists exhibit distinct representation patterns\.
#### 3\.2\.2Temporal and Medium Categorization
Historical creation dates \(‘objectBeginDate‘\) are parsed into century\-level buckets \(e\.g\.,17th17^\{\\text\{th\}\}c\. CE,19th19^\{\\text\{th\}\}c\. CE,20th20^\{\\text\{th\}\}c\. CE\)\. Raw text medium descriptions are mapped into five canonical structural categories using substring matching:
- •Painting:Oil, tempera, panel, acrylic, canvas \(72\.93%72\.93\\%,n=1,094n=1\{,\}094\)\.
- •Print:Etching, engraving, woodcut, lithograph \(12\.13%12\.13\\%,n=182n=182\)\.
- •Drawing/Paper:Ink, graphite, charcoal, watercolor, paper \(6\.27%6\.27\\%,n=94n=94\)\.
- •Other:Decorative arts, textiles, miniatures, unclassified \(7\.67%7\.67\\%,n=115n=115\)\.
- •Sculpture:Bronze, marble, terracotta, plaster \(1\.00%1\.00\\%,n=15n=15\)\.
### 3\.3Empirical Archival Representation Audit Findings
Quantifying representation disparities within public museum metadata is essential for evaluating both archival equity and potential training set bias in downstream AI models\[topaz2019diversity,meier2021gender,noorthuis2020bias\]\. Analysis of the harvested dataset \(N=1,500N=1,500\) yields the empirical breakdown across named and unnamed holdings summarized in Table[1](https://arxiv.org/html/2609.17572#S3.T1):
Table 1:Empirical Distribution of Harvested Works by Inferred Gender \(N=1,500N=1,500Corpus\)1. 1\.Structural Gender Skew:Among attributed works with resolved artist names in the broader harvest \(n=882n=882\), male artists account for70\.86%70\.86\\%\(n=625n=625\), while female artists account for29\.14%29\.14\\%\(n=257n=257\), reflecting a\>2\.4×\>2\.4\\timesstructural representation gap \(Table[1](https://arxiv.org/html/2609.17572#S3.T1)and Figure[1](https://arxiv.org/html/2609.17572#S3.F1)\)\.
2. 2\.Geographic and National Concentration:National attributions display high Eurocentric concentration across cataloged holdings, led by French \(34\.4%34\.4\\%\), American \(10\.7%10\.7\\%\), Dutch \(10\.0%10\.0\\%\), British \(9\.9%9\.9\\%\), German \(8\.6%8\.6\\%\), Flemish \(8\.5%8\.5\\%\), Netherlandish \(7\.7%7\.7\\%\), Spanish \(5\.8%5\.8\\%\), and Italian \(3\.9%3\.9\\%\) attributions \(Figure[2](https://arxiv.org/html/2609.17572#S3.F2)\)\. Non\-European nationalities constitute<1%<1\\%of cataloged holdings\.
3. 3\.Temporal Representation Trajectory:Cross\-tabulating artist gender against historical creation era reveals that female attributions remain a minority across all periods:24\.0%24\.0\\%in15th15^\{\\text\{th\}\}\-century holdings,24\.2%24\.2\\%in the16th16^\{\\text\{th\}\}century,29\.7%29\.7\\%in the17th17^\{\\text\{th\}\}century,33\.7%33\.7\\%in the18th18^\{\\text\{th\}\}century,29\.4%29\.4\\%in19th19^\{\\text\{th\}\}\-century, and25\.7%25\.7\\%in20th20^\{\\text\{th\}\}\-century collections \(Table[2](https://arxiv.org/html/2609.17572#S3.T2)\)\.
Figure 1:Empirical Distribution of Named Works by Inferred Gender \(n=882n=882named attributed works within theN=1,500N=1,500Metropolitan Museum sample\), showing a\>2\.4×\>2\.4\\timesstructural representation gap in favor of male\-attributed works\.Figure 2:Geographic and National Attribution Concentration across the top 10 national attributions, demonstrating severe Eurocentric dominance relative to non\-European artists \(<1%<1\\%\)\.Table 2:Cross\-Tabulation of Artist Gender Across Historical Eras \(N=882N=882Harvested Attributed Works\)
### 3\.4Evaluation Cohort Stratification and Data Flow Accounting
The sequential sample filtering pipeline follows a strict data flow progression: from the initial raw API harvest \(N=1,500N=1,500\), filtering out41\.20%41\.20\\%anonymous or unattributed records \(n=618n=618\) yields a subtotal ofn=882n=882named attributed objects\. Out of thesen=882n=882named attributed objects,n=139n=139objects \(15\.76%15\.76\\%\) were excluded from visual model evaluation due to missing, unresolvable, or broken primary image URLs during image asset verification\. To verify that this missingness does not introduce sample selection bias into downstream audits, we conducted chi\-square balance tests comparing the dropped \(n=139n=139\) and retained \(N=743N=743\) cohorts\. Attrition exhibited no statistically significant demographic bias \(χ2=2\.025,p=0\.1547\\chi^\{2\}=2\.025,p=0\.1547; Male65\.47%65\.47\\%dropped vs71\.87%71\.87\\%retained\) or medium bias \(χ2=1\.674,p=0\.7953\\chi^\{2\}=1\.674,p=0\.7953\), confirming a Missing Completely at Random \(MCAR\) / Missing at Random \(MAR\) missingness mechanism\. The resulting final evaluation cohort comprisesN=743N=743attributed historical artworks \(534534male\-attributed,209209female\-attributed\)\. AllN=743N=743sampled image URLs were retrieved, verified for file integrity, converted to RGB, and inspected to ensure absence of digital corruption or rendering artifacts\.
Crucially, this sequential sample filtering introduces an essential sample selection constraint: by removing41\.20%41\.20\\%\(n=618n=618\) unattributed objects and15\.76%15\.76\\%\(n=139n=139\) objects with unindexed image assets, the audited cohort \(N=743N=743\) becomes heavily concentrated in high\-status, canonical Western oil paintings, which account for nearly three\-quarters \(74\.43%74\.43\\%,n=553n=553\) of the final evaluated sample\. From an econometric and archival auditing perspective, evaluating model fairness on an artificially homogeneous, high\-status survival subset flattens visual feature variance across objects \(due to standardized canvas framing, formal oil techniques, and museum studio lighting\), which inherently suppresses score variance and contributes to the observed score convergence and lowR2R^\{2\}in downstream regressions\.
Table[3](https://arxiv.org/html/2609.17572#S3.T3)details the exact cross\-tabulation of inferred artist gender across physical artwork media in the final evaluation cohort \(N=743N=743\)\. Notably, physical artwork media display extreme class imbalance: paintings dominate the evaluation cohort \(74\.43%74\.43\\%,n=553n=553\), whereas sculptures comprise only0\.81%0\.81\\%\(n=6n=6, all male\-attributed\)\. The zero\-count female sculpture cell \(n=0n=0\) represents an explicit positivity restriction in medium×\\timesgender regression modeling\. In our primary regression models \(Table[7](https://arxiv.org/html/2609.17572#S5.T7)\), we include an explicit medium indicator for sculpture to preserve fine\-grained structural categories; to confirm that parameter estimation is robust against positivity restrictions, we conduct sensitivity cross\-validations pooling sculpture objects \(n=6n=6\) into the broaderOther \(3D / Decorative Arts\)category \(n=56n=56combined\)\. This sensitivity check confirms parameter invariance for the demographic regressor \(B=0\.0037,p=0\.202B=0\.0037,p=0\.202under OpenAI CLIP;B=0\.0065,p=0\.356B=0\.0065,p=0\.356under OpenCLIP\)\. We note that handling sparse heterogeneous media \(textiles, ceramics, miniatures, and 3D sculpture\) represents a trade\-off: while controlling for physical medium, high intra\-category variance contributes to the low overallR2R^\{2\}observed in OLS modeling\.
Table 3:Cross\-Tabulation of Complete Attributed Cohort \(N=743N=743\) by Gender and Physical Medium
## 4Methodology: Audit Framework and Metrics
This section outlines our quantitative framework for auditing zero\-shot vision\-language model \(VLM\) valuations across artwork archives\. The end\-to\-end audit pipeline is diagrammed in Fig\.[3](https://arxiv.org/html/2609.17572#S4.F3)\.
1\. Met API HarvestN=1,500N=1,500ObjectsJSON Metadata & Images2\. Metadata EnrichmentGender Inference, Medium& Century Classification3\. Full Attributed CohortN=743N=743High\-Res Images\(534534Male,209209Female\)6\. Confound AuditMann\-WhitneyUU, Bootstrap CIs& Multivariate OLS Regression5\. Valuation ScoringLogit Differential CalculationSj\(xi\)=P\(phigh\(j\)\)−P\(plow\(j\)\)S\_\{j\}\(x\_\{i\}\)=P\(p\_\{\\text\{high\}\}^\{\(j\)\}\)\-P\(p\_\{\\text\{low\}\}^\{\(j\)\}\)4\. Dual VLM InferenceOpenAI CLIP \(ViT\-B/32\)vs\. OpenCLIP \(LAION\-2B\)Figure 3:Expanded audit framework: from Met API data harvesting and demographic enrichment \(N=1,500N=1,500\) to dual vision\-language model inference, zero\-shot logit differential scoring, non\-parametric hypothesis testing, and multivariate OLS confound regression\.### 4\.1Evaluated Model Architectures
We audit two ViT\-B/32 dual\-encoder architectures \(86M86\\text\{M\}visual parameters,63M63\\text\{M\}text parameters\) differing in pretraining data curation:
1. 1\.OpenAI CLIP \(ViT\-B/32\):Pretrained on OpenAI’s curated WebImageText \(WIT\) dataset \(400M image\-text pairs\)\[radford2021learning\]\.
2. 2\.OpenCLIP \(ViT\-B/32\):Pretrained on LAION\-2B, an open\-source, uncurated web crawl \(2B image\-text pairs\)\[cherti2023reproducible,schuhmann2022laion\]\.
Both architectures optimize a symmetric contrastive loss overL2L\_\{2\}\-normalized visual embeddings𝐯i\\mathbf\{v\}\_\{i\}and text embeddings𝐭j\\mathbf\{t\}\_\{j\}:
ℒcontrastive=12N∑i=1N\(−logexp\(τ⋅𝐯i⊤𝐭i\)∑j=1Nexp\(τ⋅𝐯i⊤𝐭j\)−logexp\(τ⋅𝐯i⊤𝐭i\)∑j=1Nexp\(τ⋅𝐯j⊤𝐭i\)\)\\mathcal\{L\}\_\{\\text\{contrastive\}\}=\\frac\{1\}\{2N\}\\sum\_\{i=1\}^\{N\}\\left\(\-\\log\\frac\{\\exp\(\\tau\\cdot\\mathbf\{v\}\_\{i\}^\{\\top\}\\mathbf\{t\}\_\{i\}\)\}\{\\sum\_\{j=1\}^\{N\}\\exp\(\\tau\\cdot\\mathbf\{v\}\_\{i\}^\{\\top\}\\mathbf\{t\}\_\{j\}\)\}\-\\log\\frac\{\\exp\(\\tau\\cdot\\mathbf\{v\}\_\{i\}^\{\\top\}\\mathbf\{t\}\_\{i\}\)\}\{\\sum\_\{j=1\}^\{N\}\\exp\(\\tau\\cdot\\mathbf\{v\}\_\{j\}^\{\\top\}\\mathbf\{t\}\_\{i\}\)\}\\right\)\(1\)whereτ\\tauis the learned logit scale parameter\.
### 4\.2Value Prompt Operationalization and Logit Scoring
We probe implicit value judgments across three prompt pairs\(𝐩high\(j\),𝐩low\(j\)\)\(\\mathbf\{p\}\_\{\\text\{high\}\}^\{\(j\)\},\\mathbf\{p\}\_\{\\text\{low\}\}^\{\(j\)\}\)\(Table[4](https://arxiv.org/html/2609.17572#S4.T4)\) combined with neutral baselines𝒫base=\{”a painting”,”an artwork”,”a photograph of art”,”a museum object”\}\\mathcal\{P\}\_\{\\text\{base\}\}=\\\{\\text\{"a painting"\},\\text\{"an artwork"\},\\text\{"a photograph of art"\},\\text\{"a museum object"\}\\\}\(K=10K=10candidate set\)\.
Table 4:Operationalized Value Prompt Pair DefinitionsZero\-shot prompt probabilityP\(pk∣xi\)P\(p\_\{k\}\\mid x\_\{i\}\)is computed via softmax scaling over cosine similarities:
P\(pk∣xi\)=exp\(τ⋅𝐄v\(xi\)⊤𝐄t\(pk\)\)∑m=1Kexp\(τ⋅𝐄v\(xi\)⊤𝐄t\(pm\)\)P\(p\_\{k\}\\mid x\_\{i\}\)=\\frac\{\\exp\\left\(\\tau\\cdot\\mathbf\{E\}\_\{v\}\(x\_\{i\}\)^\{\\top\}\\mathbf\{E\}\_\{t\}\(p\_\{k\}\)\\right\)\}\{\\sum\_\{m=1\}^\{K\}\\exp\\left\(\\tau\\cdot\\mathbf\{E\}\_\{v\}\(x\_\{i\}\)^\{\\top\}\\mathbf\{E\}\_\{t\}\(p\_\{m\}\)\\right\)\}\(2\)The prompt\-level valuation scoreSj\(xi\)=P\(𝐩high\(j\)∣xi\)−P\(𝐩low\(j\)∣xi\)S\_\{j\}\(x\_\{i\}\)=P\(\\mathbf\{p\}\_\{\\text\{high\}\}^\{\(j\)\}\\mid x\_\{i\}\)\-P\(\\mathbf\{p\}\_\{\\text\{low\}\}^\{\(j\)\}\\mid x\_\{i\}\)yields the composite relative value metricSval\(xi\)S\_\{\\text\{val\}\}\(x\_\{i\}\):
Sval\(xi\)=13∑j=13\[P\(𝐩high\(j\)∣xi\)−P\(𝐩low\(j\)∣xi\)\]S\_\{\\text\{val\}\}\(x\_\{i\}\)=\\frac\{1\}\{3\}\\sum\_\{j=1\}^\{3\}\\left\[P\(\\mathbf\{p\}\_\{\\text\{high\}\}^\{\(j\)\}\\mid x\_\{i\}\)\-P\(\\mathbf\{p\}\_\{\\text\{low\}\}^\{\(j\)\}\\mid x\_\{i\}\)\\right\]\(3\)Positive scores reflect alignment with canonical masterwork status, while negative scores indicate association with minor or amateur status\.
### 4\.3Statistical Hypothesis Testing and Equivalence \(TOST\)
Due to non\-Gaussian score distributions \(verified via Shapiro\-Wilk tests\), we evaluate group differences between male \(nM=534n\_\{M\}=534\) and female \(nF=209n\_\{F\}=209\) artists using non\-parametric statistics:
- •Mann\-WhitneyUU& Rank\-Biserialrr:Mann\-WhitneyUUstatistic tests distributional equality, with rank\-biserial correlationr=1−2UnMnFr=1\-\\frac\{2U\}\{n\_\{M\}n\_\{F\}\}quantifying non\-parametric effect size\.
- •Bootstrap CIs:1,000 bootstrap iterations \(B=1,000B=1,000\) derive 95% percentile confidence intervals for mean differenceΔμ\\Delta\\mu\.
- •Two One\-Sided Tests \(TOST\):To formally evaluate statistical equivalence rather than relying on failure to reject the null\[lakens2017equivalence\], we testH01:Δ≤−ΔEH\_\{01\}:\\Delta\\leq\-\\Delta\_\{E\}andH02:Δ≥\+ΔEH\_\{02\}:\\Delta\\geq\+\\Delta\_\{E\}across equivalence boundsΔE=d⋅σpooled\\Delta\_\{E\}=d\\cdot\\sigma\_\{\\text\{pooled\}\}\(d∈\[0\.15,0\.40\]d\\in\[0\.15,0\.40\]\)\. Equivalence is confirmed atα=0\.05\\alpha=0\.05ifpTOST=max\(p1,p2\)<0\.05p\_\{\\text\{TOST\}\}=\\max\(p\_\{1\},p\_\{2\}\)<0\.05\.
### 4\.4Multivariate Confound Control Regression
To decouple demographic attributions from artwork medium, creation era, and visual framing confounds, we estimate a multivariate Ordinary Least Squares \(OLS\) model:
Sval,i=β0\+β1⋅𝕀Male,i\+∑k=1Kγk⋅𝕀Mediumk,i\+∑m=1Mλm⋅𝕀Centurym,i\+δ⋅Aspect\_Ratioi\+εiS\_\{\\text\{val\},i\}=\\beta\_\{0\}\+\\beta\_\{1\}\\cdot\\mathbb\{I\}\_\{\\text\{Male\},i\}\+\\sum\_\{k=1\}^\{K\}\\gamma\_\{k\}\\cdot\\mathbb\{I\}\_\{\\text\{Medium\}\_\{k\},i\}\+\\sum\_\{m=1\}^\{M\}\\lambda\_\{m\}\\cdot\\mathbb\{I\}\_\{\\text\{Century\}\_\{m\},i\}\+\\delta\\cdot\\text\{Aspect\\\_Ratio\}\_\{i\}\+\\varepsilon\_\{i\}\(4\)Standard errors are estimated using HC3 heteroskedasticity\-robust estimators and validated via Huber Robust Linear Models \(RLM\)\.
## 5Experimental Results
This section presents empirical findings from our dual\-model audit of vision\-language valuations across allN=743N=743attributed Metropolitan Museum artworks \(534534male\-attributed,209209female\-attributed\)\. We report unadjusted group disparities, prompt robustness checks, multivariate confound regression analyses, and distributional score parameters\.
### 5\.1Unadjusted Gender Disparity Analysis
We first evaluate whether zero\-shot aesthetic valuation scores \(SvalS\_\{\\text\{val\}\}\) differ significantly by artist gender without adjusting for structural archival covariates\. Table[5](https://arxiv.org/html/2609.17572#S5.T5)summarizes non\-parametric comparisons across OpenAI CLIP and OpenCLIP models\.
Table 5:Unadjusted Valuation Score Disparity Audit \(N=743N=743Attributed Works\)Under OpenAI CLIP, female\-attributed artworks and male\-attributed artworks receive virtually identical raw mean scores \(MF=−0\.0067,SD=0\.0344M\_\{F\}=\-0\.0067,\\text\{SD\}=0\.0344vsMM=−0\.0035,SD=0\.0361M\_\{M\}=\-0\.0035,\\text\{SD\}=0\.0361\)\. A two\-sided Mann\-WhitneyUUtest confirms that this difference is not statistically significant \(U=59,307\.00,p=0\.1829U=59,307\.00,p=0\.1829\)\. The rank\-biserial correlation \(r=−0\.0628r=\-0\.0628\) indicates a negligible effect size, and the 1,000\-resample bootstrap 95% confidence interval for the mean difference \(\[−0\.0025,0\.0087\]\[\-0\.0025,0\.0087\]\) strictly spans zero\. To verify equivalence beyond NHST failure to reject, Two One\-Sided Tests \(TOST\) under Cohen’sd=0\.30d=0\.30bounds \(ΔE=±0\.0107\\Delta\_\{E\}=\\pm 0\.0107\) confirm statistically significant equivalence \(tlower=4\.866,tupper=−2\.647,pTOST=0\.0042t\_\{\\text\{lower\}\}=4\.866,t\_\{\\text\{upper\}\}=\-2\.647,p\_\{\\text\{TOST\}\}=0\.0042\)\. Sensitivity analysis across equivalence bounds demonstrates consistent statistical equivalence: for OpenAI CLIP,d=0\.40⟹pTOST<0\.0001d=0\.40\\implies p\_\{\\text\{TOST\}\}<0\.0001,d=0\.30⟹pTOST=0\.0042d=0\.30\\implies p\_\{\\text\{TOST\}\}=0\.0042,d=0\.25⟹pTOST=0\.0210d=0\.25\\implies p\_\{\\text\{TOST\}\}=0\.0210,d=0\.20⟹pTOST=0\.0714d=0\.20\\implies p\_\{\\text\{TOST\}\}=0\.0714, andd=0\.15⟹pTOST=0\.1873d=0\.15\\implies p\_\{\\text\{TOST\}\}=0\.1873\. Statistical equivalence holds robustly for all Cohen’sd≥0\.25d\\geq 0\.25\. As visualized in Figure[4](https://arxiv.org/html/2609.17572#S5.F4), score distributions exhibit low skewness \(−0\.826\-0\.826\) and standard kurtosis \(9\.0409\.040\)\.
Similarly, OpenCLIP \(LAION\-2B\) shows high baseline score convergence across gender groups, as illustrated in Figure[5](https://arxiv.org/html/2609.17572#S5.F5)\. Female\-attributed works score0\.01710\.0171\(SD=0\.0835\\text\{SD\}=0\.0835\) on average compared to0\.02370\.0237\(SD=0\.0886\\text\{SD\}=0\.0886\) for male\-attributed works \(U=59,867\.00,p=0\.1224,r=−0\.0728,CI=\[−0\.0067,0\.0204\]U=59,867\.00,p=0\.1224,r=\-0\.0728,\\text\{CI\}=\[\-0\.0067,0\.0204\]\)\. TOST equivalence testing underd=0\.30d=0\.30bounds \(ΔE=±0\.0261\\Delta\_\{E\}=\\pm 0\.0261\) confirms statistical equivalence \(tlower=4\.722,tupper=−2\.825,pTOST=0\.0024t\_\{\\text\{lower\}\}=4\.722,t\_\{\\text\{upper\}\}=\-2\.825,p\_\{\\text\{TOST\}\}=0\.0024\)\. Sensitivity checks for OpenCLIP yieldd=0\.40⟹pTOST<0\.0001d=0\.40\\implies p\_\{\\text\{TOST\}\}<0\.0001,d=0\.30⟹pTOST=0\.0024d=0\.30\\implies p\_\{\\text\{TOST\}\}=0\.0024,d=0\.25⟹pTOST=0\.0135d=0\.25\\implies p\_\{\\text\{TOST\}\}=0\.0135,d=0\.20⟹pTOST=0\.0518d=0\.20\\implies p\_\{\\text\{TOST\}\}=0\.0518, andd=0\.15⟹pTOST=0\.1522d=0\.15\\implies p\_\{\\text\{TOST\}\}=0\.1522, confirming equivalence for all boundsd≥0\.25d\\geq 0\.25\. Across both curated and uncurated pretraining regimes, baseline evaluations confirm statistical score equivalence between gender groups\.
Figure 4:Distribution of zero\-shot aesthetic valuation scores \(SvalS\_\{\\text\{val\}\}\) across male\- and female\-attributed artworks under OpenAI CLIP \(ViT\-B/32\) across theN=743N=743attributed corpus \(Shapiro\-Wilk normalitypnormality=0\.8642p\_\{\\text\{normality\}\}=0\.8642; Mann\-Whitney group differencep=0\.1829p=0\.1829\)\.Figure 5:Distribution of zero\-shot aesthetic valuation scores \(SvalS\_\{\\text\{val\}\}\) across male\- and female\-attributed artworks under OpenCLIP \(LAION\-2B\) across theN=743N=743attributed corpus \(Shapiro\-Wilk normalitypnormality=0\.5076p\_\{\\text\{normality\}\}=0\.5076; Mann\-Whitney group differencep=0\.1224p=0\.1224\)\.
### 5\.2Prompt Sensitivity and Robustness Checks
To evaluate whether model evaluations depend on prompt phrasing, we dissect scores across three distinct value prompt formulations \(Table[6](https://arxiv.org/html/2609.17572#S5.T6)\)\.
Table 6:Prompt Set Sensitivity and Robustness Breakdown \(N=743N=743Corpus\)Prompt set sensitivity remains consistent across architectures \(Table[6](https://arxiv.org/html/2609.17572#S5.T6)and Figure[6](https://arxiv.org/html/2609.17572#S5.F6)\)\. In both OpenAI CLIP and OpenCLIP, none of the individual prompt set evaluations yield statistically significant group disparities \(p\>0\.092p\>0\.092across all raw comparisons\)\. Applying Bonferroni multiple testing correction across the six prompt\-architecture comparisons \(αadj=0\.05/6=0\.00833\\alpha\_\{\\text\{adj\}\}=0\.05/6=0\.00833\) yields adjustedpp\-values ofpadj=1\.0000p\_\{\\text\{adj\}\}=1\.0000for all prompt sets except OpenCLIP Set 2 \(padj=0\.5532p\_\{\\text\{adj\}\}=0\.5532\)\. Inspecting rank\-biserial effect sizes reveals directional stability: while Set 1 \(masterpiece\) and Set 2 \(quality\) show slight male\-leaning point estimates \(r=−0\.0537r=\-0\.0537andr=−0\.0459r=\-0\.0459for OpenAI CLIP;r=−0\.0186r=\-0\.0186andr=−0\.0794r=\-0\.0794for OpenCLIP\), Set 3 \(influence: ”a groundbreaking artwork” vs ”a decorative craft object”\) displays slight female\-leaning differentials \(r=\+0\.0221r=\+0\.0221for CLIP;r=\+0\.0245r=\+0\.0245for OpenCLIP\)\. Crucially, operationalizing historical influence against decorative craft status does not induce gendered devaluation against female creators, confirming robust valuation stability across semantic prompt formulations\.
Figure 6:Prompt sensitivity and robustness breakdown across Prompt Sets 1–3 for OpenAI CLIP \(ViT\-B/32\) and OpenCLIP \(LAION\-2B\) across theN=743N=743attributed artwork corpus\.
### 5\.3Multivariate Confound Analysis
To determine whether score variations stem from artist gender or correlated archival properties, we fit Ordinary Least Squares \(OLS\) regressions controlling for physical medium categories, creation century, and visual aspect ratio\. Table[7](https://arxiv.org/html/2609.17572#S5.T7)details model parameters\.
Table 7:Multivariate Confound Control Models \(OLS with HC3 Robust Standard Errors,N=743N=743\)Note:Standard errors are heteroskedasticity\-robust \(HC3\)\. Primary models include an explicit dummy for sculpture \(n=6n=6\); because female\-attributed sculpture exhibits a zero\-count cell \(n=0n=0female,n=6n=6male in Table 2\), the structural medium dummy reflects a positivity restriction \(P\(𝕀Sculpture∣Female\)=0P\(\\mathbb\{I\}\_\{\\text\{Sculpture\}\}\\mid\\text\{Female\}\)=0\)\. Sensitivity cross\-validations pooling sculpture objects \(n=6n=6\) intoOther \(3D/Decorative Arts\)\(n=56n=56combined, comprising2828male and2222female works\) confirm parameter invariance for the primary demographic regressor \(𝕀Male\\mathbb\{I\}\_\{\\text\{Male\}\}:B=0\.0037,p=0\.202B=0\.0037,p=0\.202in CLIP;B=0\.0065,p=0\.356B=0\.0065,p=0\.356in OpenCLIP\), demonstrating that positivity restrictions in sparse cells do not distort conditional main effect estimation\. Baseline reference category for physical medium isDrawing/Paper; baseline reference category for creation era is15th c\. / Earlier\. Robustness check using Huber Robust Linear Modeling \(RLM\) yields consistent conclusions \(Male artistB=0\.0025,p=0\.106B=0\.0025,p=0\.106in CLIP;B=0\.0065,p=0\.250B=0\.0065,p=0\.250in OpenCLIP\)\.
The multivariate regression models yield three primary insights:
1. 1\.Global Model Fit and Variance Breakdown:Across both pretraining regimes, the full multivariate OLS regression architectures yield statistically non\-significant overall fits \(F\(11,731\)=0\.9286,p=0\.512,R2=0\.018F\(11,731\)=0\.9286,p=0\.512,R^\{2\}=0\.018for OpenAI CLIP;F\(11,731\)=0\.9888,p=0\.455,R2=0\.017F\(11,731\)=0\.9888,p=0\.455,R^\{2\}=0\.017for OpenCLIP\)\. The low totalR2R^\{2\}values \(<1\.8%<1\.8\\%\) indicate that the combined set of structural covariates—physical artwork medium, creation era, framing aspect ratio, and artist demographic attribution—explains less than1\.8%1\.8\\%of total score variance\. Furthermore, individual coefficients for framing aspect ratio \(W/H\\text\{W\}/\\text\{H\}\) \(B=−0\.0064,p=0\.141B=\-0\.0064,p=0\.141in CLIP;B=−0\.0023,p=0\.516B=\-0\.0023,p=0\.516in OpenCLIP\) and individual medium/century categories fail to reach statistical significance \(p\>0\.10p\>0\.10\)\.
2. 2\.Conditional Main Effects, Statistical Power, and Instrument Insensitivity:Conditioning on physical artwork medium, creation era, and visual framing confirms that artist gender exhibits no statistically significant main effect post\-adjustment under either OpenAI CLIP \(B=0\.0037,p=0\.202B=0\.0037,p=0\.202\) or OpenCLIP \(B=0\.0065,p=0\.356B=0\.0065,p=0\.356\)\. To verify that this non\-significant demographic main effect does not stem from statistical underpowering given low modelR2R^\{2\}, we perform a post\-hoc power calculation for OLS linear regression \(N=743N=743,α=0\.05\\alpha=0\.05,k=11k=11predictors\)\. The sample size yields statistical power1−β\>0\.9981\-\\beta\>0\.998to detect a small effect size \(f2=0\.02f^\{2\}=0\.02, equivalent to Cohen’sd=0\.28d=0\.28\)\. Crucially, we must separateinstrument insensitivityfromdefinitive model equity: the non\-significant globalFF\-tests and low totalR2R^\{2\}\(<1\.8%<1\.8\\%\) demonstrate that zero\-shot logit differentials in this corpus operate near an embedding noise floor dominated by high residual embedding variance \(≈98\.2%\\approx 98\.2\\%\)\. Rather than proving complete algorithmic neutrality across all visual domains, this low explanatory power reveals that broad zero\-shot text\-prompt logit differentials function as a coarse, insensitive measurement instrument that fails to capture fine\-grained visual\-semantic features without spatial feature probing\.
3. 3\.Measurement Error and Regressor Attenuation Bias:We explicitly account for measurement error in binary demographic attribution \(≈4\.0%\\approx 4\.0\\%classification error in automated dictionary resolution\)\. In classical econometric theory, independent variable measurement error induces attenuation bias \(B^1=B1⋅\(1−θ\)\\hat\{B\}\_\{1\}=B\_\{1\}\\cdot\(1\-\\theta\)\), pulling estimated slope coefficients slightly toward zero\. Given the small magnitude of the unadjusted difference \(Δμ≈0\.0032\\Delta\\mu\\approx 0\.0032\) and robust SEs \(0\.0030\.003\), attenuation bias does not alter null hypothesis decisions, but reinforces why non\-significantpp\-values must be reported alongside TOST equivalence bounds and post\-hoc power calculations\.
## 6Discussion
Our empirical findings demonstrate that auditing vision\-language models \(VLMs\) in cultural heritage repositories requires disentangling direct demographic model evaluation from structural archival confounders\[hall2023auditing,noorthuis2020bias,simpson1951interpretation\]\. In this section, we analyze the theoretical, methodological, and institutional implications of our results, focusing on four primary themes: \(i\) distinguishing archival curation structure from algorithmic bias, \(ii\) key epistemological paradoxes in multimodal auditing, \(iii\) pretraining data dynamics across model architectures, and \(iv\) comprehensive governance frameworks for deploying AI in museum infrastructures\.
### 6\.1Distinguishing Archival Structure from Algorithmic Bias
Observational fairness audits that evaluate raw model output scores without conditioning on structural collection metadata risk misattributing physical artwork attributes or cataloging patterns to demographic model bias\[hall2023auditing,buolamwini2018gender\]\. In historical museum archives, acquisition histories and institutional access barriers restricted female artists’ access to specific physical materials, concentrating female representation in paper, watercolor, or textile media while male creators dominated monumental sculpture and large\-scale oil canvases\[topaz2019diversity,bailey2020gender,nochlin1971why\]\. Furthermore, visual framing aspect ratios \(width/height\\text\{width\}/\\text\{height\}\) and digitized photography standards introduce systematic variations in vision\-language embedding projections\[fiorucci2020machine\]\.
When visual\-text models evaluate visual surface features or image aspect ratios, unadjusted demographic comparisons conflate physical artwork properties with creator demographics\. In statistical terms, unconditioned observational audits are vulnerable to Simpson’s paradox\[simpson1951interpretation\]: aggregate group comparisons can fabricate or conceal model disparities when underlying covariates \(such as medium or century\) are unevenly distributed across demographic cohorts\. By fitting multivariate OLS regressions controlling for medium categories, creation century, and image aspect ratio, our framework isolates conditional demographic main effects\. The absence of a statistically significant gender effect post\-adjustment \(B=0\.0037,p=0\.202B=0\.0037,p=0\.202for CLIP;B=0\.0065,p=0\.356B=0\.0065,p=0\.356for OpenCLIP\) demonstrates that observed score variations stem from structural collection heterogeneity rather than active demographic valuation skew\.
### 6\.2Theoretical and Epistemological Paradoxes in Multimodal Auditing
Our quantitative findings elucidate two central conceptual paradoxes that define the boundaries of vision\-language model auditing in cultural archives:
#### 6\.2\.1The Noise Floor and Prompt Coarseness Paradox
Across both OpenAI CLIP and OpenCLIP, the full multivariate regression architectures yield non\-significant overall model fits \(F\(11,731\)=0\.9286,p=0\.512,R2=0\.018F\(11,731\)=0\.9286,p=0\.512,R^\{2\}=0\.018for CLIP;F\(11,731\)=0\.9888,p=0\.455,R2=0\.017F\(11,731\)=0\.9888,p=0\.455,R^\{2\}=0\.017for OpenCLIP\)\. The low totalR2R^\{2\}\(<1\.8%<1\.8\\%\) reveals that structural artwork covariates—medium, era, aspect ratio, and artist gender—explain less than2%2\\%of overall zero\-shot score variance\.
Rather than proving absolute model fairness, this low explanatory power illuminates a critical epistemological boundary in multimodal AI auditing:the distinction between instrument insensitivity and algorithmic equity\. Broad text\-prompt pairs \(e\.g\., ”masterpiece” vs\. ”minor work”\) project complex, multi\-dimensional visual artifacts onto low\-dimensional contrastive text vectors near an embedding noise floor dominated by high residual variance \(≈98\.2%\\approx 98\.2\\%\)\. Broad prompt logit differentials compress subtle visual representations \(such as subject gaze, lighting, compositional framing, or texture\) into a single scalar logit difference\. When98\.2%98\.2\\%of score variance is unmodeled residual noise, finding no statistically significant score difference across gender cohorts \(p=0\.1829,pTOST=0\.0042p=0\.1829,p\_\{\\text\{TOST\}\}=0\.0042\) reflects the coarseness and insensitivity of broad zero\-shot prompt differentials as a measurement instrument for aesthetic valuation\. Consequently, macro\-level score equivalence under broad prestige prompts must be interpreted as \*instrument insensitivity\* rather than definitive proof of model neutrality\. Zero\-shot global prompt audits operate near a noise floor where macro\-level evaluations fail to capture fine\-grained demographic disparities without spatial visual feature probing, attention map analyses, or targeted attribute classifiers\.
#### 6\.2\.2The Archival Survival Bias Paradox
A second critical insight concerns the scope of audited objects and the compounded impact of sample selection filters\. By evaluating model valuations across named attributed creators \(N=743N=743\), the audit pipeline tests an already\-sanitized survival cohort of artists who successfully passed historical institutional gatekeeping, academy access barriers, patron preferences, and curatorial documentation practices\[carlson2022museum,topaz2019diversity\]\.
Crucially, sequential sample filtering—removing41\.2%41\.2\\%\(n=618n=618\) of the harvested corpus due to unattributed creator metadata and excluding15\.76%15\.76\\%\(n=139n=139\) objects with unindexed image assets—concentrates the final evaluation cohort \(N=743N=743\) in high\-status, canonical Western oil paintings \(74\.43%74\.43\\%,n=553n=553\)\. Excluding anonymous holdings removes the precise archival strata where female, non\-Western, and craft labor was historically anonymized\[bailey2020gender,parker1984subversive\]\. In museum cataloging, decorative arts, textiles, ceramics, and domestic craft objects were frequently cataloged without named creator attributions, reflecting gendered institutional valuations of artistic production\[nochlin1971why,pollock1988vision\]\. Demonstrating model parity across canonical named creators on an artificially homogeneous, high\-status survival subset does not imply an absence of institutional bias; rather, institutional gender bias operated upstream at the archive ingestion and cataloging boundaries\. AI auditing frameworks must acknowledge that evaluating foundation models on digitized museum archives tests model behavior on works that already survived historical curation gatekeeping, leaving upstream archival erasure uncaptured by downstream model output scores\.
### 6\.3Cross\-Architectural Dynamics and Pretraining Data Regimes
Comparing OpenAI CLIP \(trained on curated WIT\) against OpenCLIP \(trained on uncurated LAION\-2B\) provides empirical insight into how pretraining data curation affects downstream cultural heritage representations\. Despite vast differences in pretraining scale and dataset filtering—400M curated image\-text pairs versus 2 billion uncurated web\-scraped pairs—both architectures display remarkable valuation stability across all three prompt sets \(masterpiece,quality, andinfluence\)\.
However, OpenCLIP exhibits slightly higher baseline embedding variance across prompt sets \(e\.g\., standard deviationσ=0\.0886\\sigma=0\.0886vsσ=0\.0361\\sigma=0\.0361in OpenAI CLIP\)\. This variance difference indicates that uncurated web scrapes ingest noisier text\-image co\-occurrences, increasing score volatility without introducing systematic demographic main effects\. For digital humanities researchers and collection managers, open\-weights models trained on web\-scale datasets provide robust zero\-shot baseline performance, but require tighter calibration to manage embedding variance\.
### 6\.4Implications for Museum AI Governance, Algorithmic Curation, and Policy
As museums, galleries, archives, and digital libraries increasingly deploy vision\-language models for automated collection indexing, public search, and interactive discovery interfaces\[fiorucci2020machine,garcia2020bias,srinivasan2021arts\], establishing formal AI governance protocols becomes paramount\. Uncalibrated AI deployments risk reinforcing historical representational imbalances through automated search ranking algorithms\[zhao2017men,bianchi2023easily,luccioni2023stable\]\. Grounded in Responsible AI governance frameworks\[dignum2019responsible,stahl2021responsible,mitchell2019model,jobin2019global\], we propose four core policy standards for institutional museum AI governance:
#### 6\.4\.1Policy Standard 1: Medium\-Stratified Normalization and Search Ranking Calibration
Search and discovery interfaces utilizing zero\-shot VLM embeddings to rank or surface collection holdings must implement medium\-stratified score normalization to prevent photographic framing and visual surface attributes from biasing discovery rankings\. For a visual objectxix\_\{i\}belonging to physical medium categoryk∈\{Painting,Print,Drawing,Sculpture,Other\}k\\in\\\{\\text\{Painting\},\\text\{Print\},\\text\{Drawing\},\\text\{Sculpture\},\\text\{Other\}\\\}with sample sizeNk≥30N\_\{k\}\\geq 30, public search engines should compute medium\-calibrated similarity scoreszi,kz\_\{i,k\}:
zi,k=S\(xi\)−μkσkz\_\{i,k\}=\\frac\{S\(x\_\{i\}\)\-\\mu\_\{k\}\}\{\\sigma\_\{k\}\}\(5\)Where medium categories contain sparse representations \(Nk<30N\_\{k\}<30, such as historical sculpture in specialized sub\-collections\), systems must enforce an explicit fallback control: pooling sparse objects into broader structural categories \(e\.g\., 3D/Decorative Arts\) or defaulting to global corpus parameters\(μglobal,σglobal\)\(\\mu\_\{\\text\{global\}\},\\sigma\_\{\\text\{global\}\}\)\. Normalizing scores relative to physical medium baselines prevents automated discovery algorithms from systematically penalizing media historically associated with female or non\-canonical artists\.
#### 6\.4\.2Policy Standard 2: Mandatory Pre\-Deployment Audit Workflows
Cultural heritage institutions procuring or fine\-tuning foundation models must establish mandatory pre\-deployment audit checklists prior to integration into public APIs or cataloging systems:
1. 1\.Multivariate Confound Auditing:Institutions must fit multivariate regression models conditioning on physical medium, creation era, and image geometry to isolate demographic main effects from curation confounds\.
2. 2\.Equivalence Testing \(TOST\):Model parity should be demonstrated via formal Two One\-Sided Tests \(TOST\) across specified equivalence bounds \(d≤0\.30d\\leq 0\.30\) rather than relying on non\-significantpp\-values\.
3. 3\.Prompt Sensitivity Probing:Classification systems must be evaluated across multiple semantic prompt pairs, specifically checking whether craft or decorative descriptors induce asymmetric score degradation\.
#### 6\.4\.3Policy Standard 3: Metadata Provenance Transparency and Anonymity Labeling
Automated image tagging and AI\-assisted cataloging systems must preserve metadata provenance transparency\[mitchell2019model,dignum2019responsible\]\. Algorithmic tags should be visually demarcated from human curatorial attributions, accompanied by confidence metrics and model version metadata\. Furthermore, to address survival bias in digital archives, public discovery interfaces should explicitly flag unattributed or anonymized holdings, highlighting domestic craft and textile collections to prevent AI search indexing from obscuring historically anonymized labor\.
#### 6\.4\.4Policy Standard 4: Responsible AI Procurement for Cultural Repositories
Museum executive leadership and digital transformation officers should establish clear procurement guidelines for vendor\-provided AI search and indexing solutions\. Contracts should mandate transparency regarding pretraining dataset sources, audit reports verifying confound control, and compliance with emerging international standards for Responsible AI in cultural heritage\[dignum2019responsible,jobin2019global\]\.
### 6\.5Limitations and Architectural Scaling Analysis
Several methodological boundaries and architectural constraints define the scope of this empirical audit:
#### 6\.5\.1Architectural Scaling Dynamics: ViT\-B/32 vs\. ViT\-L/14
Our primary empirical evaluation focused on dual\-encoder ViT\-B/32 transformer backbones \(86M86\\text\{M\}vision parameters,63M63\\text\{M\}text parameters, joint embedding dimensionD=512D=512\)\. In modern computer vision and multimodal retrieval, larger encoder architectures—such as ViT\-L/14 \(304M304\\text\{M\}vision parameters,D=768D=768\) and ViT\-H/14 \(632M632\\text\{M\}vision parameters,D=1024D=1024\)—exhibit higher feature capacity, finer patch resolution \(14×1414\\times 14vs\.32×3232\\times 32\), and higher learned logit scale parametersτ\\tau\. Larger parameter encoders learn sharper visual feature representations that may reduce residual embedding noise, potentially increasing sensitivity to subtle visual surface attributes\. However, parameter scaling also increases susceptibility to pretraining dataset memorandum skews, as larger encoders memorize web\-scraped associations more tightly\. Future audits should systematically benchmark zero\-shot valuation stability across parameter scales \(86M→632M86\\text\{M\}\\to 632\\text\{M\}\) to evaluate whether model capacity alters score noise floor dynamics\.
#### 6\.5\.2Contrastive Softmax Loss vs\. Pairwise Sigmoid Loss \(SigLIP\)
Standard OpenAI CLIP and OpenCLIP architectures optimize a symmetric contrastive softmax loss \(Equation 1\), which normalizes image\-text similarity scores across all candidate text prompts concurrently within a batch\. As demonstrated in Equation 2, computing zero\-shot probabilities via softmax renders logit differential scores sensitive to the temperature parameterτ\\tauand the candidate prompt baseline cardinalityKK\. In contrast, recent vision\-language architectures such as SigLIP\[zhai2023sigmoid\]replace softmax normalization with a pairwise sigmoid loss operating independently on image\-text pairs\. By decoupling prompt probability estimation from batch\-wide candidate normalization, pairwise sigmoid loss avoids scale compression and temperature distortion\. Evaluating SigLIP on cultural heritage collections represents a key direction to verify whether pairwise loss formulations improve score calibration for long\-tail art categories\.
#### 6\.5\.3Generative Multimodal LLMs \(MLLMs\) and Open\-Ended Evaluation
While zero\-shot contrastive dual\-encoders compute static visual\-text cosine similarities, instruction\-tuned Multimodal Large Language Models \(MLLMs\)—such as LLaVA, InstructBLIP, and Qwen\-VL—evaluate visual art through autoregressive generative text decoding\. Generative MLLMs can produce natural language rationales for aesthetic evaluation, catalog description, and historical context\. However, MLLMs inherit complex hallucination patterns, text\-based reasoning biases, and instruction\-following artifacts from their underlying language model backbones\. Auditing MLLMs in museum archives requires expanding from contrastive logit probability scoring to generative evaluation frameworks, combining NLP quality metrics with human expert curatorial review\.
#### 6\.5\.4Single\-Institution Scope and Archival Generalizability
Auditing a single encyclopedic museum corpus \(N=1,500N=1,500objects harvested from the Metropolitan Museum of Art\) reflects the specific acquisition trajectories, cataloging conventions, and digitization standards of a major Anglo\-American institution\. Furthermore, dictionary\-based gender resolution excluded41\.2%41\.2\\%\(n=618n=618\) unattributed objects, highlighting how legacy metadata gaps create dataset survival bias\. Finally, broad semantic prestige categories \(masterpiece,quality,influence\) operate near an embedding noise floor that cannot detect localized micro\-level visual feature biases \(e\.g\., subject gaze or compositional framing\)\.
Future research will expand this controlled audit framework across multi\-institutional repositories \(e\.g\., Europeana API, Smithsonian Open Access, Rijksmuseum Open Data\), evaluate SigLIP and larger vision\-language encoders \(ViT\-L/14\), incorporate spatial visual feature probing to detect micro\-level compositional biases, and investigate non\-Western art historical taxonomies\.
## 7Conclusion
In this study, we audited zero\-shot vision\-language model valuations across Metropolitan Museum of Art Open Access collection metadata \(N=1,500N=1,500total records;N=743N=743attributed named works: Malen=534n=534, Femalen=209n=209;n=618n=618anonymous/unattributed\) to evaluate whether observed score disparities reflect direct demographic bias or underlying archival confounders\. Unadjusted evaluations under OpenAI CLIP showed no statistically significant composite gender disparity \(U=59,307\.00,p=0\.1829,r=−0\.0628U=59,307\.00,p=0\.1829,r=\-0\.0628\), and OpenCLIP exhibited consistent baseline stability \(U=59,867\.00,p=0\.1224,r=−0\.0728U=59,867\.00,p=0\.1224,r=\-0\.0728\)\. Two One\-Sided Tests \(TOST\) confirmed statistical equivalence across Cohen’sd≥0\.25d\\geq 0\.25bounds \(pTOST=0\.0042p\_\{\\text\{TOST\}\}=0\.0042atd=0\.30d=0\.30for OpenAI CLIP;pTOST=0\.0024p\_\{\\text\{TOST\}\}=0\.0024for OpenCLIP\), supported by sensitivity checks acrossd∈\[0\.15,0\.40\]d\\in\[0\.15,0\.40\]\. Multivariate Ordinary Least Squares regression confirmed that artist gender \(B=0\.0037,p=0\.202B=0\.0037,p=0\.202for CLIP;B=0\.0065,p=0\.356B=0\.0065,p=0\.356for OpenCLIP\) and aspect ratio \(B=−0\.0064,p=0\.141B=\-0\.0064,p=0\.141for CLIP;B=−0\.0023,p=0\.516B=\-0\.0023,p=0\.516for OpenCLIP\) exhibit non\-significant conditional effects, with low overall modelR2R^\{2\}\(<2%<2\\%\) indicating that global zero\-shot prompt metrics operate near an embedding noise floor\.
We highlight two central theoretical insights: \(i\) macro\-level zero\-shot score equivalence does not preclude localized micro\-level visual feature biases, and \(ii\) excluding41\.2%41\.2\\%\(n=618n=618\) unattributed holdings reflects an uneliminable institutional survival bias, as named creators represent an already\-sanitized subset of historical acquisition filters\. As cultural heritage institutions deploy vision\-language models for collection indexing and public discovery, implementing medium\-stratified z\-score normalization with fallback controls and preserving archival provenance standards are essential to prevent automated systems from perpetuating historical representational gaps\. Mandatory multivariate confound controls and pretraining curation audits are crucial to ensuring fair and responsible AI deployment across global cultural repositories\.
## Appendix AQualitative Case Audits of Archival and Prompt Artifacts
To contextualize the quantitative regression findings, this appendix details representative qualitative case comparisons illustrating photographic framing artifacts and prompt\-level taxonomic sensitivities\.
### A\.1Case Audit 1: Photographic Framing and Surface Texture Artifacts
Evaluating model logit outputs for 3D marble sculptures versus 2D oil paintings on canvas reveals how visual feature encoders respond to digitisation artifacts\. Three\-dimensional sculptures are photographed under directional studio lighting against artificial gradient backgrounds, generating specular highlights and shadow gradients across surface contours\. In contrast, 2D paintings and works on paper are captured via flat\-field illumination\. Conditioning on physical medium and aspect ratio framing in multivariate OLS regression eliminates these photographic artifacts \(B=\+0\.0002,p=0\.716B=\+0\.0002,p=0\.716\)\.
### A\.2Case Audit 2: Fine Art vs\. Decorative Craft Taxonomy Disparities
Prompt Set 3 probes artistic status by calculating probability differentials between ”a groundbreaking artwork” \(𝐩high\(3\)\\mathbf\{p\}\_\{\\text\{high\}\}^\{\(3\)\}\) and ”a decorative craft object” \(𝐩low\(3\)\\mathbf\{p\}\_\{\\text\{low\}\}^\{\(3\)\}\)\. Prompt robustness checks across allN=743N=743attributed works confirm high evaluation stability across prompt pairs \(p\>0\.10p\>0\.10\), confirming that prompt phrasing does not induce systematic demographic valuation shifts\.
## Declarations
- •Funding: No funding was received for this work\.
- •Conflict of interest: The authors declare no conflicts of interest\.
- •Ethics approval: Not applicable\.
- •Data availability: All metadata and visual assets evaluated in this study are publicly available via the Metropolitan Museum of Art Open Access API\.
- •
- •Author contribution: All authors contributed to experimental design, data collection, statistical analysis, and manuscript preparation\.
## ReferencesSimilar Articles
Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges
A research paper introducing Chain-of-Models (CoM), an automated pipeline where a second LLM audits a first model's reasoning trace to correct cognitive biases. It finds that auditor effectiveness depends on model family and bias type, and proposes a bias-specific auditor selection rule that improves judgment accuracy.
Forensic Reproducibility Audit of a Radiology Vision-Language Model Benchmark: From Intended Protocol to Released Artifact
This paper performs a forensic reproducibility audit of a radiology vision-language model benchmark, finding divergences between the intended protocol and released artifacts that invalidate the original claims. The authors propose a benchmark contract to expose such failure classes.
Solving the “Whac-a-mole dilemma”: A smarter way to debias AI vision models
Researchers from MIT, WPI, and Google propose WRING, a novel post-processing debiasing method for Vision-Language Models that avoids the 'Whac-a-mole dilemma' of amplifying other biases when removing specific ones.
The Geopolitics of AI Safety: A Causal Analysis of Regional LLM Bias
This paper introduces a Probabilistic Graphical Model framework to causally audit LLM safety mechanisms, revealing that standard observational metrics overestimate demographic bias by ignoring context toxicity.
StylisticBias: A Few Human Visual Cues Drive Most Social Biases in MLLMs
A new benchmark called StylisticBias systematically evaluates attribute-level social bias in multimodal large language models, finding that a small set of visual cues like fashion style drive most biases.