Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science

arXiv cs.CL Papers

Summary

This paper tests how different LLM families evaluate ethnonationalist pseudo-science across time and interfaces, finding that epistemic stance is contingent on deployment configuration rather than stable model properties, raising concerns about epistemic accountability.

arXiv:2607.22513v1 Announce Type: cross Abstract: Commercial large language models are increasingly used as knowledge references, yet their stance on contested scientific claims is neither stable nor transparent. We tested how four major LLM families (Claude, Grok, GPT, Gemini) evaluate ethnonationalist pseudo-science derived from Frank Salter's biosocial framework across four temporal snapshots (October 2025-February 2026), via both API and web interfaces. Grok's Fast versions (which power the default user experience on X) consistently assigned credibility scores of 70-75, two to five times higher than all other models (which scored 15-40). This pattern was absent from control prompts testing basic evolutionary consensus and refuted Lamarckian claims, where all models performed comparably. Three additional findings emerged: (1) a silent patch reversed Grok's behaviour from chaotic to stably high validation overnight, without any public documentation; (2) the same Grok model identifier produced radically divergent outputs via API (75) and web (5.5) three months later; (3) refusal to rate the pseudo-scientific claim, the most defensible response observed, appeared in two model families through different interfaces (Claude Opus 4.1 categorically via web, GPT-5.1 Chat intermittently via API) and eroded in the successor version of each. These results indicate that the epistemic stance of a commercial LLM is not a stable property of the model but a contingent effect of deployment configuration: system prompts, safety layers, interface routing, and silent updates. This remains opaque to users and researchers alike. We argue this constitutes a matter of public concern requiring new forms of epistemic accountability.
Original Article
View Cached Full Text

Cached at: 07/27/26, 07:42 AM

# Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science
Source: [https://arxiv.org/abs/2607.22513](https://arxiv.org/abs/2607.22513)
[View PDF](https://arxiv.org/pdf/2607.22513)

> Abstract:Commercial large language models are increasingly used as knowledge references, yet their stance on contested scientific claims is neither stable nor transparent\. We tested how four major LLM families \(Claude, Grok, GPT, Gemini\) evaluate ethnonationalist pseudo\-science derived from Frank Salter's biosocial framework across four temporal snapshots \(October 2025\-February 2026\), via both API and web interfaces\. Grok's Fast versions \(which power the default user experience on X\) consistently assigned credibility scores of 70\-75, two to five times higher than all other models \(which scored 15\-40\)\. This pattern was absent from control prompts testing basic evolutionary consensus and refuted Lamarckian claims, where all models performed comparably\. Three additional findings emerged: \(1\) a silent patch reversed Grok's behaviour from chaotic to stably high validation overnight, without any public documentation; \(2\) the same Grok model identifier produced radically divergent outputs via API \(75\) and web \(5\.5\) three months later; \(3\) refusal to rate the pseudo\-scientific claim, the most defensible response observed, appeared in two model families through different interfaces \(Claude Opus 4\.1 categorically via web, GPT\-5\.1 Chat intermittently via API\) and eroded in the successor version of each\. These results indicate that the epistemic stance of a commercial LLM is not a stable property of the model but a contingent effect of deployment configuration: system prompts, safety layers, interface routing, and silent updates\. This remains opaque to users and researchers alike\. We argue this constitutes a matter of public concern requiring new forms of epistemic accountability\.

## Submission history

From: Davide Scarso \[[view email](https://arxiv.org/show-email/40e0ee65/2607.22513)\] **\[v1\]**Fri, 24 Jul 2026 17:32:43 UTC \(341 KB\)

Similar Articles

Trust, but Don't Verify: Epistemic Blind Spots in LLM Source Evaluation

arXiv cs.LG

This paper identifies a failure mode in LLMs where they do not verify the validity of numerical statistics when synthesizing multiple sources, instead relying on the stylistic markers of analytical rigor. The authors term this 'epistemic alignment' and show that it persists across models and domains, resisting prompting-based mitigations.

Evaluating Second-Order Bias of LLMs Through Epistemic Entitlement

arXiv cs.CL

This paper introduces 'second-order bias', the bias LLMs exhibit when judging biased content, and proposes a reasoning task grounded in epistemic entitlement to evaluate it. Experiments show that the task evades safety guardrails and reveals systematic demographic biases in LLM judges.