Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science
Summary
This paper tests how different LLM families evaluate ethnonationalist pseudo-science across time and interfaces, finding that epistemic stance is contingent on deployment configuration rather than stable model properties, raising concerns about epistemic accountability.
View Cached Full Text
Cached at: 07/27/26, 07:42 AM
# Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science Source: [https://arxiv.org/abs/2607.22513](https://arxiv.org/abs/2607.22513) [View PDF](https://arxiv.org/pdf/2607.22513) > Abstract:Commercial large language models are increasingly used as knowledge references, yet their stance on contested scientific claims is neither stable nor transparent\. We tested how four major LLM families \(Claude, Grok, GPT, Gemini\) evaluate ethnonationalist pseudo\-science derived from Frank Salter's biosocial framework across four temporal snapshots \(October 2025\-February 2026\), via both API and web interfaces\. Grok's Fast versions \(which power the default user experience on X\) consistently assigned credibility scores of 70\-75, two to five times higher than all other models \(which scored 15\-40\)\. This pattern was absent from control prompts testing basic evolutionary consensus and refuted Lamarckian claims, where all models performed comparably\. Three additional findings emerged: \(1\) a silent patch reversed Grok's behaviour from chaotic to stably high validation overnight, without any public documentation; \(2\) the same Grok model identifier produced radically divergent outputs via API \(75\) and web \(5\.5\) three months later; \(3\) refusal to rate the pseudo\-scientific claim, the most defensible response observed, appeared in two model families through different interfaces \(Claude Opus 4\.1 categorically via web, GPT\-5\.1 Chat intermittently via API\) and eroded in the successor version of each\. These results indicate that the epistemic stance of a commercial LLM is not a stable property of the model but a contingent effect of deployment configuration: system prompts, safety layers, interface routing, and silent updates\. This remains opaque to users and researchers alike\. We argue this constitutes a matter of public concern requiring new forms of epistemic accountability\. ## Submission history From: Davide Scarso \[[view email](https://arxiv.org/show-email/40e0ee65/2607.22513)\] **\[v1\]**Fri, 24 Jul 2026 17:32:43 UTC \(341 KB\)
Similar Articles
Validating LLMs in social science: Epistemic threats and emerging norms
This paper analyzes validation practices for using LLMs as measurement instruments in social science, identifying epistemic threats and proposing emerging norms for robust validation.
Trust, but Don't Verify: Epistemic Blind Spots in LLM Source Evaluation
This paper identifies a failure mode in LLMs where they do not verify the validity of numerical statistics when synthesizing multiple sources, instead relying on the stylistic markers of analytical rigor. The authors term this 'epistemic alignment' and show that it persists across models and domains, resisting prompting-based mitigations.
Measuring Epistemic Resilience of LLMs Under Misleading Medical Context
Introduces MedMisBench to measure LLMs' ability to maintain correct medical reasoning under misleading context. Shows that accuracy drops sharply from 71.1% to 38.0% under adversarial conditions, with potential harm flagged by clinical panel.
Polistemics: Evaluating LLMs as Information Mediators in Politics & Elections
This paper introduces Polistemics, a theory-grounded benchmark for evaluating how LLMs mediate political information in elections, and finds that while aggregate scores look good, models systematically fail under ambiguous or contradictory information.
Evaluating Second-Order Bias of LLMs Through Epistemic Entitlement
This paper introduces 'second-order bias', the bias LLMs exhibit when judging biased content, and proposes a reasoning task grounded in epistemic entitlement to evaluate it. Experiments show that the task evades safety guardrails and reveals systematic demographic biases in LLM judges.