Same Facts, Different Diagnosis: Measuring and Mitigating Narrative Anchoring in Clinical Language Models
Summary
This paper introduces Narrative Anchoring, a failure mode where clinical language models produce divergent diagnoses when identical clinical facts are expressed in different sociolinguistic registers. The authors release a USMLE-derived dataset and propose NarrativeShield, a three-agent pipeline that reduces the anchoring gap to near-zero.
View Cached Full Text
Cached at: 07/31/26, 10:01 AM
# Measuring and Mitigating Narrative Anchoring in Clinical Language ModelsCode and data available at https://github.com/prabhjotschugh/narrativeshield-sdoh-medqa Source: [https://arxiv.org/html/2607.27384](https://arxiv.org/html/2607.27384) Prabhjot Singh University of Texas at Austin Austin, Texas, USA [prabhjot\.singh@utexas\.edu](https://arxiv.org/html/2607.27384v1/mailto:[email protected])&Pritam Deka Queen’s University Belfast Belfast, United Kingdom [pdeka01@qub\.ac\.uk](https://arxiv.org/html/2607.27384v1/mailto:[email protected])&Vijay Chennareddy Middlesex University London, United Kingdom [Vc381@live\.mdx\.ac\.uk](https://arxiv.org/html/2607.27384v1/mailto:[email protected]) ###### Abstract Large language models used for clinical diagnostic reasoning are sensitive to sociolinguistic register, not just clinical content\. We term this failure mode Narrative Anchoring: identical clinical facts expressed in different registers cause diagnostic outputs to diverge\. Unlike prior demographic\-bias work, which manipulates explicit identity tokens such as race or income, our benchmark isolates register as the sole channel of variation, with no demographic marker present in any form\. We construct a dataset of 1,000 USMLE clinical vignettes, each rewritten into three sociolinguistically distinct personas under an independently audited fact\-preservation guarantee, verified by a separate model that never sees the generation prompt\. Across seven language models spanning three architecture families and scales, Narrative Anchoring is statistically significant under direct prompting in every model tested, with a Narrative Anchoring Gap of 0\.064 to 0\.151\. Chain\-of\-thought reasoning and explicit debiasing instructions reduce the bias only partially, and their apparent gains are frequently confounded by accuracy collapse\. We introduce NarrativeShield, a three\-agent pipeline that structurally extracts and verifies clinical facts before diagnostic reasoning begins, reducing the Narrative Anchoring Gap to near\-zero \(−0\.004\-0\.004to0\.0370\.037\) and achieving the lowest rate of severely unstable decisions \(DSS<<0\.8\) of any method across all models, at a modest and mechanistically expected accuracy cost for most models\. A stress test using a non\-instruction\-tuned base model shows that executing a debiasing intervention at all is gated by zero\-shot instruction\-following ability, not prompt content alone\. We release our dataset, human\-validated for fact preservation, as a standalone resource for studying register\-based clinical bias\. ## 1Introduction Large language models are increasingly proposed as decision support tools for clinical diagnostic reasoning, from triage assistance to differential diagnosis generation, carrying a corresponding risk of introducing harm or exacerbating health disparities in deployment\. Much of the resulting research has focused on demographic bias: whether a model’s diagnostic or treatment recommendation changes when a patient’s race, gender, or socioeconomic status is stated explicitly\. This concern is well founded\.Zacket al\.\([2024](https://arxiv.org/html/2607.27384#bib.bib1)\)show that GPT\-4 consistently produces clinical vignettes that stereotype demographic presentations, with differential diagnoses more likely to include conditions associated with particular races, ethnicities, and genders\. Benchmarks such as EquityMedQA\(Pfohlet al\.,[2024](https://arxiv.org/html/2607.27384#bib.bib2)\)and DiversityMedQA\(Rawatet al\.,[2024](https://arxiv.org/html/2607.27384#bib.bib3)\)have made this failure mode measurable by perturbing patient demographics directly\.Poulainet al\.\([2026](https://arxiv.org/html/2607.27384#bib.bib4)\)extend this further, finding that reflection\-style prompting can reduce biased outcomes, while larger and medically fine\-tuned models are not necessarily less biased\. This line of work shares a structural assumption that leaves a clinically realistic failure mode unexamined\. Real patients do not announce their income bracket or cultural background to a clinician, or to a model\. They describe their own symptoms, in their own words, shaped by how they were taught to talk about pain, what they can afford to do about it, and what their family expects of someone who is unwell\. Two patients with identical symptoms, duration, and vital signs can produce vignettes a model reads very differently, not because either mentioned a demographic category, but because the register in which their story is told differs\. We term this failure modeNarrative Anchoring: a shift in a model’s diagnostic output that tracks sociolinguistic register rather than clinical content, occurring even when every clinical fact is held constant and no demographic marker is ever stated\. This mechanism already carries diagnostic weight independent of any stated label\. Across 1\.7 million outputs from nine models,Omaret al\.\([2025](https://arxiv.org/html/2607.27384#bib.bib5)\)find that cases described with certain sociodemographic framings were directed toward more aggressive care than clinically indicated, relative to a matched control description of the same facts, while holding clinical content constant\. Narrative Anchoring isolates this mechanism precisely: our three personas contain no demographic marker of any kind, no race, income figure, or named cultural identifier, anywhere in the text\. The bias channel is register and framing alone, not identity disclosure, which distinguishes our benchmark from EquityMedQA, DiversityMedQA, and the demographic\-label paradigm generally, and shows that a model need not be told who a patient is to treat that patient differently\. Narrative Anchoring is largely invisible to benchmarks that present each case exactly once: a model anchored to register never gets the chance to reveal it if it is never shown the same patient twice, told two different ways\. Measuring it requires a dataset that varies register while holding clinical content fixed, with that fact preservation independently verified rather than assumed, since any drift would confound register effects with genuine case difficulty\. We make three contributions: - •A dataset of 1,000 patient vignettes, filtered from the MedQA\-USMLE corpus\(Jinet al\.,[2020](https://arxiv.org/html/2607.27384#bib.bib6)\), each rewritten into three sociolinguistically distinct personas: control, socioeconomic, and cultural\. Fact preservation is certified not by the generating model itself but by a second, independently invoked model with no access to the generation prompt, whose only task is to verify that every clinical fact survives the rewrite, following the finding that separately invoked evaluator models approximate human judgment more reliably than models grading their own output\(Zhenget al\.,[2023](https://arxiv.org/html/2607.27384#bib.bib7)\)\. We further validate this guarantee through human annotation\. - •Evidence that Narrative Anchoring is not marginal\. Across seven models spanning three architecture families and a range of scales, direct prompting produces a statistically significant divergence in recommendations between the control persona and each marked persona, in every model tested, with a Narrative Anchoring Gap of 0\.064 to 0\.151\. Chain\-of\-thought reasoning\(Weiet al\.,[2023](https://arxiv.org/html/2607.27384#bib.bib8)\)and an explicit debiasing instruction reduce this only partially, and chain\-of\-thought’s apparent gains are frequently confounded by a drop in accuracy, consistent with evidence that a chain\-of\-thought explanation does not reliably reflect the true reason for a model’s prediction and can be steered by contextual, non\-clinical cues\(Turpinet al\.,[2023](https://arxiv.org/html/2607.27384#bib.bib9)\)\. - •NarrativeShield, a three\-agent pipeline that addresses Narrative Anchoring architecturally: it structurally extracts and verifies clinical facts before diagnostic reasoning begins, decoupling what the model reasons over from how the patient’s story was told\. Across all seven models, NarrativeShield reduces the Narrative Anchoring Gap to near\-zero \(−0\.004\-0\.004to0\.0370\.037\) and achieves the lowest rate of severely unstable, persona\-dependent decisions \(Decision Stability Score below 0\.8\) of any method we test\. This comes with a modest, mechanistically explicable accuracy cost for most models, and a striking failure in a domain\-adapted but non\-instruction\-tuned base model\(Labraket al\.,[2024](https://arxiv.org/html/2607.27384#bib.bib10)\), which we show is evidence that any debiasing intervention presupposes a baseline zero\-shot instruction\-following capacity that not all models possess\. The risk, then, is not only that a model treats two patients differently because it was told they belong to different demographic groups\. It is that a model treats two patients differently because of how each one, or their family, describes what is wrong, a channel of bias that requires no demographic disclosure and survives interventions aimed only at prompt content rather than the architecture that processes it\. ## 2Dataset Construction We constructNarrativeShield\-SDoH, 1,000 USMLE\-style clinical vignettes, each expressed in three sociolinguistic personas that hold clinical content fixed while varying register\. Construction has two stages: deterministic filtering and sampling, and persona generation with independent fact\-preservation auditing\. ### 2\.1Source Filtering and Sampling We draw candidates from MedQA\-USMLE\(Jinet al\.,[2020](https://arxiv.org/html/2607.27384#bib.bib6)\), pooling train and test splits, and apply six filters as a single boolean mask with fixed seed 42, making filtering fully deterministic\. Each filter targets a property required for persona rewriting: \(1\) a patient\-vignette stemˆ\(A\|An\)\\s \+\\d \+\[\\\-\\s \]?\(year\|month\|week\|day\)\[\\\-\\s \]?old\|\(2\) minimum length of 300 characters, ensuring narrative depth beyond pure recall; \(3\) at least one clinical signal term from a 49\-term list spanning symptoms, presentation context, vitals/labs, and physical exam; \(4\) exclusion of 21 phrase patterns indicating pure mechanism or pathophysiology questions, which carry no patient\-reported framing to vary; \(5\) a target type of diagnosis, treatment, or investigation; and \(6\) at least one narrative\-richness marker from 27 phrases indicating first\-person or caregiver\-reported history \(e\.g\., “she states”\)\. Full regex and phrase lists are in Appendix[A](https://arxiv.org/html/2607.27384#A1)\. From the 2,921\-question filtered pool, we draw 1,000 via proportional stratified sampling over the cross of USMLE exam step and correct\-answer letter, preserving this joint distribution in the sample \(Appendix[B](https://arxiv.org/html/2607.27384#A2)\)\. ### 2\.2Persona Generation Each vignette is rewritten into three personas using Gemini\-3\.1\-Flash\-Lite: acontrol\(PαP\_\{\\alpha\}\) preserving the original register, asocioeconomic\(PβP\_\{\\beta\}\) reframed through financial or occupational strain, and acultural\(PγP\_\{\\gamma\}\) reframed through culturally situated metaphor\. No persona contains an explicit demographic marker; the bias channel is register alone\. Four design decisions distinguish this pipeline from naive rewriting, each closing a failure mode observed in an earlier iteration\.Independent generation calls:each persona is generated via a separate call with no shared context, since joint generation risks anchoring phrasing across personas into cosmetic synonym substitution\.Differentiated temperature:0\.25 forPαP\_\{\\alpha\}, favoring fidelity, and 0\.90 forPβP\_\{\\beta\}/PγP\_\{\\gamma\}, favoring genuine register divergence, with high temperature permitting linguistic divergence while the fact\-preservation constraint below prevents clinical content loss\.Structural, not cosmetic, register:PβP\_\{\\beta\}/PγP\_\{\\gamma\}prompts explicitly forbid patterns identified in earlier iterations as producing detectable templating, generic phrasing \(e\.g\., “traditional herbal teas”\), formulaic framing \(e\.g\., “my family insisted”\), and end\-loaded cues, and each generation is assigned one of five opening styles sampled independently per attempt, preventing a detectable structural template across the dataset \(full definitions in Appendix[C](https://arxiv.org/html/2607.27384#A3)\)\.Independent fact\-preservation audit:each persona is passed to a separate auditor model at temperature 0 with no access to the generation prompt, whose sole task is to extract every clinical fact from the original and verify its presence in the persona, clinical, lay, or metaphorical, returning a pass/fail verdict; the generating model cannot pass its own test\. A persona enters the dataset only onpass; failures retry up to 10 times with a freshly sampled opening style, and every prompt states the fact\-preservation requirement directly, with a fallback forPγP\_\{\\gamma\}that an unmetaphorizable fact must be stated plainly\. ### 2\.3Human Validation All 3,000 persona generations reachedSUCCESSunder audit, with zeroPARTIALorFAILEDrows released; we report this completeness figure rather than intermediate pass rates, which are pipeline diagnostics rather than properties of the released data\. To corroborate this automated guarantee independently, three annotators with clinical backgrounds, ranging from an attending physician with 20 years of experience to residents and interns, none involved in construction and none with access to auditor verdicts, rated a stratified sample of 100 questions \(300 encounters\) on fact preservation, persona register match, and 5\-point Likert narrative realism\. Full protocol and raw ratings are released with the dataset\. #### Fact preservation and register match\. All three annotators rated 100% of encounters as preserving facts and matching register, a raw agreement ceiling for which chance\-corrected statistics are undefined by construction\. Since the sample was drawn only from audit\-passed personas, this reflects the fact\-preservation gate working on this sample rather than an unfiltered failure rate, consistent with comparably high post\-filter fidelity reported for physician\-reviewed AI\-generated clinical vignettes elsewhere\(Yanagitaet al\.,[2024](https://arxiv.org/html/2607.27384#bib.bib15)\)\. #### Narrative realism\. This subjective dimension diverged sharply\. One annotator rated all 300 items at ceiling \(confirmed as genuine judgment, not artifact\); restricting to the two informative annotators, exact agreement was 42\.0% but adjacent \(±1\\pm 1\) agreement was 85\.3%, and Krippendorff’sα\\alpha\(Krippendorff,[2004](https://arxiv.org/html/2607.27384#bib.bib16)\)was 0\.226, below the 0\.667 threshold for tentative conclusions, though comparable to reported agreement on other subjective NLP tasks \(27%κ\\kappaon GoEmotions\(Demszkyet al\.,[2020](https://arxiv.org/html/2607.27384#bib.bib17)\), 46% on HateXplain\(Mathewet al\.,[2021](https://arxiv.org/html/2607.27384#bib.bib18)\)\)\. Critically, disagreement was not random: exact agreement fell from 74\.0% \(PαP\_\{\\alpha\}\) to 34\.0% \(PβP\_\{\\beta\}\) to 18\.0% \(PγP\_\{\\gamma\}\), tracking persona difficulty\. Despite this calibration gap, ordering was stable: pooled means werePα=4\.89\>Pβ=4\.64\>Pγ=4\.19P\_\{\\alpha\}=4\.89\>P\_\{\\beta\}=4\.64\>P\_\{\\gamma\}=4\.19, holding independently for every annotator \(Kruskal\-WallisH=138\.69H=138\.69,p<\.0001p<\.0001, all pairwise comparisons significant after Bonferroni correction\)\. Annotators disagree on where exactly aPγP\_\{\\gamma\}narrative sits on the scale but agree without exception that it is harder to render naturalistically thanPβP\_\{\\beta\}, which is harder thanPαP\_\{\\alpha\}, the ordering persona construction was designed to produce, consistent with evidence that annotator disagreement on graded tasks can track conceptual difficulty rather than inconsistency\(Kellertet al\.,[2026](https://arxiv.org/html/2607.27384#bib.bib19)\)\. We report this as a limitation: judging realism of non\-clinical register is a harder, more subjective task than judging fact preservation, echoing this paper’s broader argument that narrative style, not clinical content, is where models and annotators alike diverge most\. Full statistics and per\-persona breakdowns are in Appendix[I](https://arxiv.org/html/2607.27384#A9)\. ## 3Experimental Setup ### 3\.1Models We evaluate seven open\-weight models: Llama\-3\.1\-8B\-Instruct and Llama\-3\.2\-3B\-Instruct\(Grattafioriet al\.,[2024](https://arxiv.org/html/2607.27384#bib.bib11)\), Mistral\-7B\-Instruct\-v0\.3\(Jianget al\.,[2023](https://arxiv.org/html/2607.27384#bib.bib12)\), Qwen2\.5\-7B\-Instruct\(Qwenet al\.,[2025](https://arxiv.org/html/2607.27384#bib.bib13)\), gemma\-3\-12b\-it and gemma\-4\-E4B\-it\(Teamet al\.,[2025](https://arxiv.org/html/2607.27384#bib.bib14)\), and BioMistral\-7B\(Labraket al\.,[2024](https://arxiv.org/html/2607.27384#bib.bib10)\), a PubMed Central\-adapted Mistral\-7B\-Instruct\-v0\.3 base model without instruction tuning\. BioMistral is included as an architectural stress test: it isolates zero\-shot instruction\-following capacity from medical domain knowledge\. All models use greedy decoding\. ### 3\.2Conditions Four conditions, applied identically across all three personas of every vignette: B1 \(Direct\)\.Persona and options presented directly, no reasoning scaffold\. Zero\-shot floor\. B2 \(Chain\-of\-Thought\)\.Step\-by\-step reasoning before answering\(Weiet al\.,[2023](https://arxiv.org/html/2607.27384#bib.bib8)\)\. For BioMistral, a two\-shot demonstration of the reasoning\-then\-answer format is included, since the base model cannot infer novel output structure from instruction alone\. B3 \(Explicit Debiasing\)\.A single capitalized instruction to disregard sociolinguistic framing, administered zero\-shot with no demonstration, including for BioMistral\. The asymmetry with B2 is intentional: comparing them isolates the effect of in\-context demonstration on format commitment independent of reasoning quality\. Figure 1:NarrativeShield architecture\. Agent 1 strips sociolinguistic register while preserving all clinical facts; a deterministic, non\-model tool router injects clinical decision support; Agent 2 reasons only over Agent 1’s structured output; Agent 3 activates only when Agent 2 fails to commit to a parseable answer\.NarrativeShield \(NS\)\.Our proposed architectural intervention \(§[3\.3](https://arxiv.org/html/2607.27384#S3.SS3)\)\. We report parse rate for every condition: the fraction of responses yielding a parseable final answer letter\. Under B1–B3 this is computed on raw output; under NS it accounts for Agent 3 fallback\. Parse rate is architecturally informative, not merely a data\-quality footnote: a model that reasons adequately but cannot commit to format produces a qualitatively different failure than one that reasons poorly, a distinction we rely on directly for BioMistral\. ### 3\.3NarrativeShield Architecture NarrativeShield is a three\-agent pipeline with a deterministic tool layer between Agents 1 and 2\. Its principle: register is removed structurally by an extraction step blind to downstream reasoning, not suppressed by instructing a reasoning model to ignore it\. Figure[1](https://arxiv.org/html/2607.27384#S3.F1)summarizes the pipeline end to end\. #### Agent 1: Adversarial Intake Extractor\. Receives only the raw persona narrative; returns structured JSON \(age, sex, chief complaint, onset, present and absent symptoms, vitals, physical exam, labs, imaging, current and administered medications, history, question type\)\. The prompt enumerates what must be stripped \(lay phrasing, economic framing, cultural references, emotional language, somatic metaphors\) and what must be preserved exactly \(every number, every symptom present or absent, all medications by pharmacological name, all history and exam findings\)\. This is the load\-bearing debiasing step: persona voice is gone before any downstream component sees the case\. Agent 2 has no access to the original narrative\. #### Deterministic Tool Router\. A fixed Python layer, not a model, inspects Agent 1’s JSON and the question, invoking up to four tools: a lab interpreter using static reference ranges; at most one of five clinical risk calculators \(Wells DVT/PE, CURB\-65, Glasgow Coma Scale, Apgar\), each gated on mutually exclusive keyword patterns; and a static 47\-drug knowledge base, queried against current medications or answer options, capped at two lookups for treatment or management questions\. Routing is rule\-based because smaller models do not reliably emit well\-formed tool\-call syntax\. Tool outputs are injected into Agent 2’s prompt as decision support before reasoning begins\. #### Agent 2: Clinical Reasoning Engine\. Receives Agent 1’s JSON, tool outputs, and the question with options\. Follows a fixed protocol: summarize the clinical picture, enumerate a differential grounded in specific extracted values, state a conclusion, commit to a single answer letter\. Never sees the original narrative\. #### Agent 3: Emergency Fallback Extractor\. Invoked only when Agent 2’s output contains no parseable answer\. Extracts a single letter from Agent 2’s already\-generated text, performing no reasoning\. This separates two failure modes: poor reasoning versus adequate reasoning with format\-commitment failure\. ### 3\.4Metrics Five metrics, computed identically across all conditions\. Option Match Rate \(OMR\)\.Binary accuracy, per persona and overall, with 95% Wilson score intervals\. Narrative Anchoring Gap \(NAG\)\.OMRα−min\(OMRβ,OMRγ\)\\mathrm\{OMR\}\_\{\\alpha\}\-\\min\(\\mathrm\{OMR\}\_\{\\beta\},\\mathrm\{OMR\}\_\{\\gamma\}\)\. Primary equity metric: the accuracy advantage a neutral register confers over the more disadvantaged marked persona\. Near\-zero NAG indicates register\-insensitive diagnostic accuracy\. Diagnostic Stability Score \(DSS\)\.Mean pairwise cosine similarity of full reasoning outputs across the three persona presentations of the same question, usingall\-MiniLM\-L6\-v2\. Primary stability metric: captures semantic drift in reasoning even when final answers agree\. We report the fraction of questions with DSS below 0\.80 \(severely unstable\) alongside the mean\. Cohen’sκ\\kappaand McNemar’s test\.Secondary: chance\-corrected inter\-persona agreement and directional bias\. We noteκ\\kappa’s high\-accuracy paradox and treat DSS as interpretively primary\. For NS specifically:Agent 1 parse rate\(valid JSON fraction\),mean tools invoked per question, andmean end\-to\-end latencyas a deployment\-relevant descriptor, not a primary criterion\. ## 4Results ### 4\.1Narrative Anchoring Is Pervasive Under Direct Prompting Under B1, every one of the seven models shows a positive Narrative Anchoring Gap, ranging from 0\.064 \(BioMistral\) to 0\.151 \(Llama\-3\.1\-8B\-Instruct\), with McNemar’s test significant \(p<0\.05p<0\.05, uncorrected\) on the control\-versus\-socioeconomic and control\-versus\-cultural pairs for all seven models\.111Across all 21 pairwise tests in B1 \(7 models×\\times3 persona pairs\), Bonferroni correction \(α=0\.05/21\\alpha=0\.05/21\) flips two comparisons from significant to non\-significant, both secondaryβ\\beta\-versus\-γ\\gammaor borderline pairs; every control\-versus\-socioeconomic and control\-versus\-cultural comparison, the pairs our central claim rests on, remains significant under correction\. Full values are in Appendix[E](https://arxiv.org/html/2607.27384#A5)\.This holds across architecture families, parameter scales from 3B to 12B, and both general\-purpose and medically domain\-adapted models\. This establishes Narrative Anchoring as a pervasive property of direct clinical prompting rather than an artifact of any single model family, and directly supports our first claim: sociolinguistic register alone, with no demographic marker present, is sufficient to shift diagnostic output\. BioMistral’s B1 NAG is the smallest of the seven models, but this is not evidence of relatively unbiased behavior: BioMistral simultaneously has the lowest overall OMR \(0\.426\) and the highest rate of severely unstable decisions \(DSS<0\.8<0\.8: 80\.8%\) of any model under B1\. NAG is a bounded difference of accuracies and mechanically compresses as overall accuracy falls toward a floor; DSS, an unbounded similarity measure, shows the opposite and more informative story\. We return to BioMistral’s role directly in §[4\.4](https://arxiv.org/html/2607.27384#S4.SS4)\. ### 4\.2Chain\-of\-Thought and Explicit Debiasing Only Partially Mitigate Anchoring Table[1](https://arxiv.org/html/2607.27384#S4.T1)reports NAG across all four conditions; the corresponding DSS<<0\.8% synthesis is in Appendix[H](https://arxiv.org/html/2607.27384#A8)\. Under B2, five of seven models show reduced NAG relative to B1, but this improvement is confounded by substantial accuracy loss for three of them: gemma\-3\-12b\-it loses 21 points of overall OMR, gemma\-4\-E4B\-it loses 34 points, and BioMistral’s OMR collapses alongside a parse rate of only 54\.07%, the lowest of any non\-degenerate condition \(Figure[3](https://arxiv.org/html/2607.27384#S4.F3)\)\. A lower NAG computed over a smaller or differently\-distributed set of correct answers is not evidence of improved fairness\. We therefore do not read B2 as a clean mitigation result\. B3 is the more informative comparison\. For the six models capable of executing the instruction, NAG falls modestly from B1 \(e\.g\., Llama\-3\.1\-8B\-Instruct: 0\.151 to 0\.135; Qwen2\.5: 0\.091 to 0\.067\), overall OMR is essentially unchanged, and mean DSS improves in every case\. This is our cleanest baseline finding: a single explicit debiasing instruction helps a little, at no accuracy cost, for models that can follow it\. BioMistral cannot follow it: under B3, its parse rate is 0\.00%, every OMR value is exactly zero, and Cohen’sκ\\kappais undefined for all three persona pairs\. We discuss why in §[4\.4](https://arxiv.org/html/2607.27384#S4.SS4)\. Table 1:Narrative Anchoring Gap across all four conditions\. \*BioMistral’s B3 value is a degenerate artifact of a 0\.00% parse rate \(§[4\.4](https://arxiv.org/html/2607.27384#S4.SS4)\), not evidence of eliminated bias\.Table[1](https://arxiv.org/html/2607.27384#S4.T1)reports NAG across all four conditions, and Figure[2](https://arxiv.org/html/2607.27384#S4.F2)visualizes the same trajectory across models; the corresponding DSS<<0\.8% synthesis is in Appendix[H](https://arxiv.org/html/2607.27384#A8)\. Figure 2:Narrative Anchoring Gap across all four conditions, per model\. Six of seven models converge toward near\-zero NAG under NarrativeShield \(NS\); BioMistral \(dashed\) follows a qualitatively different trajectory, discussed in §[4\.4](https://arxiv.org/html/2607.27384#S4.SS4)\.Figure 3:Overall Option Match Rate \(OMR\) across B1 \(Direct\), B2 \(Chain\-of\-Thought\), and B3 \(Explicit Debiasing\), per model\. ### 4\.3NarrativeShield Reduces Anchoring to Near\-Zero NarrativeShield reduces NAG to within\[−0\.004,0\.037\]\[\-0\.004,0\.037\]for every model, a reduction of at least four\-fold relative to B1 for every model, and considerably more where NAG falls to near\-zero or below \(Table[1](https://arxiv.org/html/2607.27384#S4.T1)\)\. Table[2](https://arxiv.org/html/2607.27384#S4.T2)reports full NarrativeShield results\. NarrativeShield also achieves the lowest DSS<0\.8<0\.8rate of any condition for every one of the seven models, ranging from 16\.3% \(gemma\-3\-12b\-it\) to 31\.2% \(Llama\-3\.2\-3B\-Instruct\), compared to a B1 range of 35\.6% to 80\.8%\. Because Agent 1’s fact\-extraction step is verified against the same independently audited fact\-preservation guarantee that underlies the released dataset \(§[2\.3](https://arxiv.org/html/2607.27384#S2.SS3)\), we attribute this collapse in both NAG and DSS to NarrativeShield’s reasoning architecture rather than to any possibility that persona difficulty happened to equalize across variants: the dataset’s own construction rules out that explanation directly\. Table 2:NarrativeShield primary results\. Full pipeline metrics are in Appendix[G](https://arxiv.org/html/2607.27384#A7)\.This reduction carries a modest, model\-dependent accuracy cost\. Relative to B1, overall OMR under NarrativeShield changes by less than 1\.5 points for three models, drops by 4\.2 to 6\.9 points for two others, and improves by 6\.7 points for gemma\-4\-E4B\-it, the only model to gain accuracy under the pipeline\. We attribute the cost, where present, to Agent 1 acting as a lossy compression step: any detail Agent 1 fails to extract is permanently unavailable to Agent 2, regardless of whether the original narrative contained it\. gemma\-4\-E4B\-it’s gain is consistent with the same mechanism running in the opposite direction: for a model whose direct\-prompting accuracy is otherwise held back by narrative noise irrelevant to the clinical question, Agent 1’s structured extraction can remove more confounding signal than clinical signal, improving rather than degrading downstream reasoning\. This is a direct, mechanistically expected consequence of the architecture, not an unexplained side effect, and one we consider a favorable trade against a near\-total elimination of register\-driven diagnostic instability\. ### 4\.4BioMistral as an Architectural Stress Test BioMistral’s behavior across all four conditions traces a single, monotonically worsening pattern that we argue is evidence for our thesis rather than noise requiring explanation away\. Each condition places an increasing demand on zero\-shot instruction\-following capacity, independent of medical reasoning ability, and BioMistral’s degree of failure tracks that demand precisely: under B1, which requires no format compliance beyond a direct answer, BioMistral answers with partial success \(OMR 0\.426\); under B2, given a two\-shot demonstration of the required format, its parse rate recovers to 54\.07%; under B3, given only a bare instruction with no demonstration, its parse rate falls to 0\.00%; and under NarrativeShield’s Agent 1, which requires zero\-shot emission of a structured JSON schema with no worked example, its parse rate is 1\.10%\. Inspection of its raw output under B3 and Agent 1 shows BioMistral consistently produces reasoning text addressing the clinical question; it fails specifically to resolve that reasoning into the required output structure\. This is precisely what a domain\-adapted base model without equivalent instruction\-tuning predicts: BioMistral can imitate a structure demonstrated in\-context, as its B2 recovery shows, but cannot infer a novel output format from instruction alone\. This pattern has a direct implication for how NarrativeShield’s own results should be read\. BioMistral’s near\-zero NAG under NarrativeShield is not evidence that the architecture successfully debiased BioMistral; with a 1\.10% Agent 1 parse rate, its NarrativeShield results are computed over a residual set of cases too small to support any claim about generalized debiasing, and we report them for completeness rather than as a further success case\. The broader claim NarrativeShield supports is that any debiasing intervention, whether a single instruction or a structural pipeline, first requires a baseline capacity to execute the intervention at all; that capacity is what BioMistral consistently lacks, and its absence, not the content of any specific prompt, is what explains its trajectory across all four conditions\. ## 5Conclusion We introduced Narrative Anchoring, a bias in clinical language models that operates through sociolinguistic register rather than demographic disclosure, and showed it is present, significant, and largely unaddressed by chain\-of\-thought or explicit debiasing across seven models spanning three architecture families\. We releasedNarrativeShield\-SDoH, a dataset built to isolate this channel precisely, with fact preservation certified by an independent auditor model and corroborated by human clinical annotation\. We introduced NarrativeShield, a three\-agent pipeline that reduces this bias to near\-zero by structurally separating clinical fact extraction from diagnostic reasoning, and we used a domain\-adapted, non\-instruction\-tuned base model as a stress test\. We show that any debiasing intervention presupposes a baseline instruction\-following capacity not all models have\. Taken together, this argues that clinical bias in language models is not solely a question of what a patient is labeled, but of how a patient’s own words are allowed to reach a model’s reasoning, and that closing this gap is a matter of pipeline design as much as prompt design\. ## Limitations We report five limitations of the present study\. First, all seven models we evaluate are open\-weight and span 3B to 12B parameters\. We do not test proprietary frontier models, and whether Narrative Anchoring persists, weakens, or strengthens at larger scale or under different training regimes is a question our design cannot answer\. Second, all personas inNarrativeShield\-SDoHare generated by a single model, Gemini\-3\.1\-Flash\-Lite\. Systematic tendencies in how that specific model renders socioeconomic or cultural register, rather than properties of the underlying phenomenon we study, could in principle shape our results\. The independent fact\-preservation audit and human validation constrain this risk to a question of framing rather than clinical fidelity, but do not eliminate it, since both checks were themselves designed around the same generation pipeline\. Third, our benchmark isolates two persona categories beyond the control condition, socioeconomic and cultural framing\. Other axes of sociolinguistic register, including disability framing, age\-related speech patterns, and non\-native speaker phrasing, may anchor language models differently and are not covered by this study\. Fourth, our vignettes are single\-turn USMLE examination questions rather than multi\-turn clinical dialogue or real patient encounters\. We make no claim about how Narrative Anchoring manifests, or how NarrativeShield performs, in extended, interactive clinical use, where a patient’s register may itself evolve over the course of a conversation\. Fifth, NarrativeShield’s accuracy cost, while modest for most models tested, is a genuine trade\-off rather than a free improvement, and this paper does not evaluate whether that cost is acceptable relative to any specific clinical deployment context\. Such a determination depends on the stakes and setting of a particular application, which we are not positioned to specify in the abstract, and which any real deployment of an architecture like NarrativeShield would need to establish independently\. These limitations bound the scope of our claims rather than undermine them\. We do not claim Narrative Anchoring is the only channel through which sociolinguistic bias enters clinical language model reasoning, nor that NarrativeShield is a complete solution to it\. We claim that register\-driven bias exists independent of demographic disclosure, that it survives the mitigations most readily available at the prompt level, and that a structural intervention addresses it more reliably than an instructional one, within the scope of models, personas, and task format we test\. #### Future directions\. Each limitation points to a natural extension\. Applying the same audited persona\-generation pipeline to proprietary frontier models would test how far Narrative Anchoring generalizes beyond open weights\. Adding further persona axes, such as disability framing or non\-native speaker phrasing, would test whether the effect and its mitigation hold across register types or are specific to the two studied here\. Extending the benchmark to multi\-turn dialogue, with register shifting across turns rather than fixed at intake, would test whether NarrativeShield’s architectural separation of register from content survives interactive use\. Finally, evaluating NarrativeShield’s accuracy trade\-off against clinician judgment on live cases, rather than against exam ground truth alone, is a necessary step before any claim of deployment readiness\. ## Ethical Considerations #### Dataset and human subjects\. NarrativeShield\-SDoHis derived entirely from MedQA\-USMLE, a publicly available, de\-identified corpus of medical examination questions; no real patient data, clinical records, or personally identifiable information is used at any stage of construction\. All three personas are synthetic rewrites of exam vignettes and do not describe or reference any real individual\. Human validation of the dataset was conducted by physician annotators who volunteered their clinical expertise, were not exposed to any patient data, and rated only the fidelity of already\-synthetic text; no institutional patient\-facing risk was involved in this process\. #### Dual\-use and deployment risk\. This work demonstrates that clinical language models are sensitive to a patient’s sociolinguistic register, a finding with a clear protective motivation, but one that could in principle be misused to deliberately construct adversarial patient narratives designed to elicit a specific diagnosis or treatment recommendation from a deployed system\. We believe the balance of this risk favors publication: the underlying vulnerability already exists in deployed models regardless of whether it is documented, and identifying it, together with an architectural mitigation, gives practitioners a concrete tool to close the gap rather than leaving it undocumented and unaddressed\. #### Clinical deployment\. Neither the base models we evaluate nor NarrativeShield itself are validated, certified, or intended for real clinical decision\-making\. All experiments in this paper are conducted on retrospective examination questions with known ground\-truth answers, not on live patients or clinical workflows\. This paper does not endorse any of the evaluated models, or NarrativeShield, as suitable for unsupervised clinical use\. We view our contribution as diagnostic of a failure mode and a research\-stage mitigation, not as a deployment\-ready clinical tool, and we recommend that any translation of NarrativeShield\-style architectures toward real clinical settings undergo the validation, regulatory review, and prospective evaluation appropriate to a clinical decision support system\. #### Representation of socioeconomic and cultural register\. Varying socioeconomic and cultural framing risks encoding the very stereotypes we study if personas rely on caricatured markers\. We mitigate this by forbidding generic or formulaic markers during generation \(§[2\.2](https://arxiv.org/html/2607.27384#S2.SS2)\) and by having physician annotators, not the generating model, judge personas as naturalistic rather than caricatured\. We nonetheless acknowledge that any synthetic rendering of illness\-talk risks flattening real diversity within a group, and do not claim our three personas are exhaustive or representative\. #### AI writing assistance\. Portions of this paper’s prose were drafted and revised with the assistance of a generative language model, for language polishing and editing of the authors’ original content, consistent with the ACL Policy on AI Writing Assistance\. All technical claims, citations, and results were verified by the authors against the underlying code, data, and literature\. ## Acknowledgements We express our sincere gratitude to the medical professionals who generously volunteered their clinical expertise to annotate and validate the dataset for this work:Dr\. Roopam Deka, DM \(Assistant Professor, Department of Pathology, All India Institute of Medical Sciences, Guwahati\);Dr\. Benjina Ahmed, MBBS \(Junior Resident, Department of Pathology and Lab Medicine, All India Institute of Medical Sciences, Guwahati\); andDr\. Meghna Kashyap, MBBS \(Intern, Laxmi Chandravansi Medical College and Hospital\)\. Their rigorous evaluation of clinical fact preservation and narrative fidelity was fundamental to ensuring the benchmark’s quality\. ## References - GoEmotions: a dataset of fine\-grained emotions\.InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics,D\. Jurafsky, J\. Chai, N\. Schluter, and J\. Tetreault \(Eds\.\),Online,pp\. 4040–4054\.External Links:[Link](https://aclanthology.org/2020.acl-main.372/),[Document](https://dx.doi.org/10.18653/v1/2020.acl-main.372)Cited by:[§2\.3](https://arxiv.org/html/2607.27384#S2.SS3.SSS0.Px2.p1.12)\. - A\. Grattafiori, A\. Dubey, A\. Jauhri,et al\.\(2024\)The llama 3 herd of models\.External Links:2407\.21783,[Link](https://arxiv.org/abs/2407.21783)Cited by:[§3\.1](https://arxiv.org/html/2607.27384#S3.SS1.p1.1)\. - A\. Q\. Jiang, A\. Sablayrolles, A\. Mensch, C\. Bamford, D\. S\. Chaplot, D\. de las Casas, F\. Bressand, G\. Lengyel, G\. Lample,et al\.\(2023\)Mistral 7b\.External Links:2310\.06825,[Link](https://arxiv.org/abs/2310.06825)Cited by:[§3\.1](https://arxiv.org/html/2607.27384#S3.SS1.p1.1)\. - D\. Jin, E\. Pan, N\. Oufattole, W\. Weng, H\. Fang, and P\. Szolovits \(2020\)What disease does this patient have? a large\-scale open domain question answering dataset from medical exams\.External Links:2009\.13081,[Link](https://arxiv.org/abs/2009.13081)Cited by:[1st item](https://arxiv.org/html/2607.27384#S1.I1.i1.p1.1),[§2\.1](https://arxiv.org/html/2607.27384#S2.SS1.p1.1)\. - O\. Kellert, S\. Kondury, C\. Koo, N\. Tyagi, and S\. Eikenberry \(2026\)Structured disagreement in health\-literacy annotation: epistemic stability, conceptual difficulty, and agreement\-stratified inference\.External Links:2604\.19943,[Link](https://arxiv.org/abs/2604.19943)Cited by:[§2\.3](https://arxiv.org/html/2607.27384#S2.SS3.SSS0.Px2.p1.12)\. - K\. Krippendorff \(2004\)Content analysis: an introduction to its methodology\.2nd edition,Sage Publications,Thousand Oaks, CA\.Cited by:[§2\.3](https://arxiv.org/html/2607.27384#S2.SS3.SSS0.Px2.p1.12)\. - Y\. Labrak, A\. Bazoge, E\. Morin, P\. Gourraud, M\. Rouvier, and R\. Dufour \(2024\)BioMistral: a collection of open\-source pretrained large language models for medical domains\.InFindings of the Association for Computational Linguistics: ACL 2024,L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 5848–5864\.External Links:[Link](https://aclanthology.org/2024.findings-acl.348/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-acl.348)Cited by:[3rd item](https://arxiv.org/html/2607.27384#S1.I1.i3.p1.2),[§3\.1](https://arxiv.org/html/2607.27384#S3.SS1.p1.1)\. - B\. Mathew, P\. Saha, S\. M\. Yimam, C\. Biemann, P\. Goyal, and A\. Mukherjee \(2021\)HateXplain: a benchmark dataset for explainable hate speech detection\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.35,pp\. 14867–14875\.Cited by:[§2\.3](https://arxiv.org/html/2607.27384#S2.SS3.SSS0.Px2.p1.12)\. - M\. Omar, S\. Soffer, R\. Agbareia, N\. L\. Bragazzi, D\. U\. Apakama, C\. R\. Horowitz, A\. W\. Charney, R\. Freeman, B\. Kummer, B\. S\. Glicksberg, G\. N\. Nadkarni, and E\. Klang \(2025\)Sociodemographic biases in medical decision making by large language models\.Nature Medicine31\(6\),pp\. 1873–1881\.External Links:[Document](https://dx.doi.org/10.1038/s41591-025-03626-6),[Link](https://doi.org/10.1038/s41591-025-03626-6),ISSN 1546\-170XCited by:[§1](https://arxiv.org/html/2607.27384#S1.p3.1)\. - S\. R\. Pfohl, H\. Cole\-Lewis, R\. Sayres, D\. Neal, M\. Asiedu, A\. Dieng, N\. Tomasev, Q\. M\. Rashid, S\. Azizi, N\. Rostamzadeh, L\. G\. McCoy, L\. A\. Celi, Y\. Liu, M\. Schaekermann, A\. Walton, A\. Parrish, C\. Nagpal, P\. Singh, A\. Dewitt, P\. Mansfield,et al\.\(2024\)A toolbox for surfacing health equity harms and biases in large language models\.Nature Medicine30\(12\),pp\. 3590–3600\.External Links:[Document](https://dx.doi.org/10.1038/s41591-024-03258-2),[Link](https://doi.org/10.1038/s41591-024-03258-2),ISSN 1546\-170XCited by:[§1](https://arxiv.org/html/2607.27384#S1.p1.1)\. - R\. Poulain, F\. I\. Adiba, H\. Fayyaz, and R\. Beheshti \(2026\)Bias patterns in the application of LLMs for clinical decision support: a comprehensive study\.Delaware Journal of Public Health12\(1\),pp\. 54–67\.External Links:[Document](https://dx.doi.org/10.32481/djph.2026.03.10)Cited by:[§1](https://arxiv.org/html/2607.27384#S1.p1.1)\. - Qwen, :, A\. Yang, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Li, D\. Liu, F\. Huang, H\. Wei, H\. Lin, J\. Yang, J\. Tu, J\. Zhang, J\. Yang, J\. Yang, J\. Zhou, J\. Lin, K\. Dang, K\. Lu, K\. Bao, K\. Yang, L\. Yu, M\. Li, M\. Xue, P\. Zhang, Q\. Zhu, R\. Men, R\. Lin, T\. Li, T\. Tang, T\. Xia, X\. Ren, X\. Ren, Y\. Fan, Y\. Su, Y\. Zhang, Y\. Wan, Y\. Liu, Z\. Cui, Z\. Zhang, and Z\. Qiu \(2025\)Qwen2\.5 technical report\.External Links:2412\.15115,[Link](https://arxiv.org/abs/2412.15115)Cited by:[§3\.1](https://arxiv.org/html/2607.27384#S3.SS1.p1.1)\. - R\. Rawat, H\. McBride, R\. Ghosh, D\. Nirmal, J\. Moon, D\. Alamuri, S\. O’Brien, and K\. Zhu \(2024\)DiversityMedQA: a benchmark for assessing demographic biases in medical diagnosis using large language models\.InProceedings of the Third Workshop on NLP for Positive Impact,D\. Dementieva, O\. Ignat, Z\. Jin, R\. Mihalcea, G\. Piatti, J\. Tetreault, S\. Wilson, and J\. Zhao \(Eds\.\),Miami, Florida, USA,pp\. 334–348\.External Links:[Link](https://aclanthology.org/2024.nlp4pi-1.29/),[Document](https://dx.doi.org/10.18653/v1/2024.nlp4pi-1.29)Cited by:[§1](https://arxiv.org/html/2607.27384#S1.p1.1)\. - G\. Team, A\. Kamath, J\. Ferret, S\. Pathak, N\. Vieillard,et al\.\(2025\)Gemma 3 technical report\.External Links:2503\.19786,[Link](https://arxiv.org/abs/2503.19786)Cited by:[§3\.1](https://arxiv.org/html/2607.27384#S3.SS1.p1.1)\. - M\. Turpin, J\. Michael, E\. Perez, and S\. R\. Bowman \(2023\)Language models don’t always say what they think: unfaithful explanations in chain\-of\-thought prompting\.InThirty\-seventh Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=bzs4uPLXvi)Cited by:[2nd item](https://arxiv.org/html/2607.27384#S1.I1.i2.p1.1)\. - J\. Wei, X\. Wang, D\. Schuurmans, M\. Bosma, B\. Ichter, F\. Xia, E\. Chi, Q\. Le, and D\. Zhou \(2023\)Chain\-of\-thought prompting elicits reasoning in large language models\.External Links:2201\.11903,[Link](https://arxiv.org/abs/2201.11903)Cited by:[2nd item](https://arxiv.org/html/2607.27384#S1.I1.i2.p1.1),[§3\.2](https://arxiv.org/html/2607.27384#S3.SS2.p3.1)\. - Y\. Yanagita, D\. Yokokawa, S\. Uchida, Y\. Li, T\. Uehara, and M\. Ikusaka \(2024\)Can AI\-generated clinical vignettes in Japanese be used medically and linguistically?\.Journal of General Internal Medicine39\(16\),pp\. 3282–3289\.External Links:[Document](https://dx.doi.org/10.1007/s11606-024-09031-y)Cited by:[§2\.3](https://arxiv.org/html/2607.27384#S2.SS3.SSS0.Px1.p1.1)\. - T\. Zack, E\. Lehman, M\. Suzgun, J\. A\. Rodriguez, L\. A\. Celi, J\. Gichoya, D\. Jurafsky, P\. Szolovits, D\. W\. Bates, R\. E\. Abdulnour, A\. J\. Butte, and E\. Alsentzer \(2024\)Assessing the potential of gpt\-4 to perpetuate racial and gender biases in health care: a model evaluation study\.The Lancet Digital Health6\(1\),pp\. e12–e22\.External Links:[Document](https://dx.doi.org/10.1016/S2589-7500%2823%2900225-X),[Link](https://doi.org/10.1016/S2589-7500(23)00225-X),ISSN 2589\-7500Cited by:[§1](https://arxiv.org/html/2607.27384#S1.p1.1)\. - L\. Zheng, W\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Lin, Z\. Li, D\. Li, E\. P\. Xing, H\. Zhang, J\. E\. Gonzalez, and I\. Stoica \(2023\)Judging llm\-as\-a\-judge with mt\-bench and chatbot arena\.External Links:2306\.05685,[Link](https://arxiv.org/abs/2306.05685)Cited by:[1st item](https://arxiv.org/html/2607.27384#S1.I1.i1.p1.1)\. ## Appendix AFilter Definitions and Regular Expressions We apply six filters as a single combined boolean mask over the pooled MedQA\-USMLE train and test splits, with a fixed random seed \(SEED = 42\) for all sampling operations, making the entire filtering and sampling process deterministic and reproducible\. #### Filter 1: Patient\-vignette stem\. The question must open with a patient\-vignette stem matching: ``` ^(A|An)\s+\d+[\-\s]? (year|month|week|day)[\-\s]?old ``` e\.g\., “A 23\-year\-old woman…”, “A 6\-month\-old boy…”\. #### Filter 2: Minimum length\. Question text must be at least 300 characters, ensuring at least two to three sentences of clinical context beyond pure factual recall\. #### Filter 3: Clinical signal terms\. The question must contain at least one term from a curated 49\-term list spanning four categories: symptoms \(25 terms, e\.g\., pain, fever, cough, dyspnea, nausea, syncope\), presentation context \(6 terms, e\.g\., presents to, emergency department, admitted\), vitals and labs \(12 terms, e\.g\., blood pressure, hemoglobin, creatinine, mmHg\), and physical exam findings \(6 terms, e\.g\., palpation, auscultation, tenderness, murmur, edema, jaundice\)\. #### Filter 4: Exclusion of mechanism and pathophysiology questions\. The question must not match any of 21 excluded phrase patterns across four categories: mechanism and pathophysiology \(e\.g\., “most likely mechanism”, “pathogenesis”, “which enzyme”, “which receptor”\), anatomy, genetics and biochemistry, and histology \(e\.g\., “which gene”, “histological”, “biopsy shows”\)\. These question types carry no patient\-reported symptom framing to vary across personas\. #### Filter 5: Target question type\. The question must match at least one of 24 phrase patterns across three categories: diagnosis \(6 patterns, e\.g\., “most likely diagnosis”, “most likely cause”, “most likely etiology”\), treatment and management \(12 patterns, e\.g\., “best treatment”, “most appropriate next step”, “which medication”\), and investigation \(6 patterns, e\.g\., “best initial test”, “most appropriate test”, “confirm the diagnosis”\)\. #### Filter 6: Narrative richness markers\. The question must contain at least one of 27 phrases indicating the patient or a caregiver is reporting symptoms in narrative form, e\.g\., “she states”, “he reports”, “the patient denies”, “history of”, “for the past”, “worsening”, “mother reports”\. ## Appendix BSampling Statistics The combined filter mask yields 2,921 eligible questions from the pooled MedQA\-USMLE corpus\. We draw 1,000 questions via proportional stratified sampling over the cross of USMLE exam step \(Step 1 vs\. Step 2&3\) and correct\-answer letter \(A–D\)\. Each stratum contributes a number of rows proportional to its share of the filtered pool, computed asn=min\(\|stratum\|,max\(1,⌊1000×\|stratum\|/2921⌋\)\)n=\\min\(\|\\text\{stratum\}\|,\\max\(1,\\lfloor 1000\\times\|\\text\{stratum\}\|/2921\\rfloor\)\), with a top\-up step drawing any shortfall from the remaining filtered pool, followed by a final shuffle \(random\_state=42\) before truncation to exactly 1,000 rows\. The released dataset preserves the answer\-letter and exam\-step distribution of the filtered pool to within rounding\. We note one detail for exact reproducibility: the top\-up step draws any shortfall from the entire remaining filtered pool rather than proportionally from under\-filled strata specifically\. Because the filtered pool \(2,921\) comfortably exceeds the target sample size \(1,000\), this branch is not expected to trigger under typical random seeds, but we flag it as a minor implementation detail relevant to exact stratification guarantees under different sampling parameters\. ## Appendix CPersona Generation: Forbidden Patterns and Opening Styles ### C\.1Forbidden Structural Patterns To prevent detectable structural templating, generation prompts forPβP\_\{\\beta\}andPγP\_\{\\gamma\}explicitly forbid patterns identified in an earlier pipeline iteration as producing cosmetic rather than genuine register variation\.Forbidden inPβP\_\{\\beta\}\(socioeconomic\):sentences beginning with a formulaic third\-person hedge \(e\.g\., “He mentioned he was hesitant to come back so soon because…”\); economic framing appended only to the final sentence of an otherwise standard narrative; the verbatim phrase “I can’t afford to miss work”; apologetic openers \(e\.g\., “I’m sorry to bother you”, “I know you’re busy”\); and ending the persona with the economic detail rather than integrating it throughout\.Forbidden inPγP\_\{\\gamma\}\(cultural\):the verbatim phrases “my family insisted I come”, “traditional herbal teas” \(flagged as too generic\), “my mother insisted”, and “fulfill my duties”; cultural framing appearing only in the final one to two sentences; and somatic metaphors that reduce to simple synonym substitution \(the prompt explicitly notes that “burning sensation” for dysuria is not a somatic metaphor, since it remains standard medical language\)\. ### C\.2Opening Styles EachPβP\_\{\\beta\}andPγP\_\{\\gamma\}generation is independently assigned one of five opening styles per attempt, sampled fresh on each retry, to prevent a detectable structural template across the dataset\.PβP\_\{\\beta\}opening styles:Impact\-First\(open with the functional or daily\-life consequence before naming the symptom\);Timeline\-First\(open with how long the patient delayed seeking care and what was tried\);Remedy\-First\(open with a specific named over\-the\-counter product or home remedy that failed\);Worry\-First\(open with the patient’s fear of job, family, or financial consequences, with symptoms following as the reason for that worry\);Symptom\-First\-Lay\(open directly with the chief complaint in maximally informal language, with no preamble\)\.PγP\_\{\\gamma\}opening styles:Somatic\-Metaphor\-First\(open with the body metaphor before any chronological or contextual framing\);Family\-Context\-First\(open with a specific family or community member’s specific action, e\.g\., “My husband placed his hand on my forehead and said I must go today”, rather than a generic statement that family insisted\);Remedy\-Failure\-First\(open with a culturally specific failed home remedy, e\.g\., ginger water, cumin seed tea, turmeric compress, prayer, or elder\-prescribed rest, rather than a generic reference\);Duty\-Frame\-First\(open with obligations fulfilled or unfulfilled due to illness, before symptoms\);Quiet\-Onset\-First\(open with an understated, formally phrased onset description, carrying cultural register only through language and metaphor, without family framing\)\. We note that the opening style used per successful generation is sampled at generation time but not persisted in the released dataset schema; reproducing an opening\-style distribution statistic would require either instrumenting the pipeline to log the winning style per row, or a post\-hoc classification of released text\. ## Appendix DAgent 1 Output Schema Agent 1 returns a single JSON object per persona narrative, used unmodified as input to the deterministic tool router and Agent 2: ``` { "age": <int or null>, "sex": "<male|female|unknown>", "gestation_weeks": <int or null>, "chief_complaint": "<neutral clinical sentence>", "symptom_onset_days": <float or null>, "symptoms_present": ["<s1>"], "symptoms_absent": ["<a1>"], "vitals": { "temp_f": <float or null>, "bp_systolic": <int or null>, "bp_diastolic": <int or null>, "hr": <int or null>, "rr": <int or null>, "spo2": <float or null> }, "physical_exam": ["<f1>"], "labs": [{"name": "<lab>", "value": <float>, "unit": "<unit>"}], "imaging": ["<finding1>"], "medications_current": ["<drug dose>"], "medications_given_ed": ["<drug dose>"], "past_medical_history": ["<condition>"], "past_surgical_history": ["<procedure>"], "allergies": ["<allergen>"], "relevant_history": "<sentence or null>", "question_type": "<diagnosis|treatment| mechanism|next_step| prognosis>" } ``` Scalar fields usenullwhen absent; list fields use\[\]\. Agent 1’s system prompt instructs that fields describing narrative register, economic framing, cultural references, or emotional language have no corresponding schema field and must not appear anywhere in the output\. ## Appendix EFull B1 Results Table[3](https://arxiv.org/html/2607.27384#A5.T3)reports full per\-model OMR, NAG, and DSS statistics under B1\. Table[4](https://arxiv.org/html/2607.27384#A5.T4)reports the corresponding Cohen’sκ\\kappaand McNemar’s test values by persona pair, underlying the significance claims in §[4\.1](https://arxiv.org/html/2607.27384#S4.SS1)\. Table 3:B1 \(Direct Prompting\) full results\.Table 4:B1 Cohen’sκ\\kappaand McNemar’s testpp\-values \(uncorrected\) by persona pair\. Under Bonferroni correction \(α=0\.05/21\\alpha=0\.05/21\), Llama\-3\.1\-8B\-Instruct’s\(β,γ\)\(\\beta,\\gamma\)pair and BioMistral’s\(α,β\)\(\\alpha,\\beta\)pair no longer reach significance; all other entries are unaffected\. ## Appendix FFull B2 and B3 Results Table 5:B2 \(Chain\-of\-Thought\) full results\.Table 6:B3 \(Explicit Debiasing\) full results\. \*BioMistral’s DSS of 1\.0000 is a degenerate artifact of comparing empty\-response embeddings across all three personas, not genuine reasoning stability; its parse rate is 0\.00% \(§[4\.4](https://arxiv.org/html/2607.27384#S4.SS4)\)\.Table[5](https://arxiv.org/html/2607.27384#A6.T5)and Table[6](https://arxiv.org/html/2607.27384#A6.T6)report full per\-model OMR, NAG, DSS, and parse rate statistics under B2 and B3 respectively, underlying the discussion in §[4\.2](https://arxiv.org/html/2607.27384#S4.SS2)\. Cohen’sκ\\kappaand McNemar’s test for B2 and B3, computed identically to the B1 procedure in Appendix[E](https://arxiv.org/html/2607.27384#A5), are reported in the released supplementary evaluation files rather than reproduced here, since the paper’s central significance claim rests on B1 \(§[4\.1](https://arxiv.org/html/2607.27384#S4.SS1)\) and B2/B3 are discussed primarily through OMR, NAG, and DSS in the main text\. ## Appendix GFull NarrativeShield Results Table[7](https://arxiv.org/html/2607.27384#A7.T7)reports full per\-model OMR, NAG, and DSS statistics under NarrativeShield\. Table[8](https://arxiv.org/html/2607.27384#A7.T8)reports the corresponding Cohen’sκ\\kappa, McNemar’s test, and DSS<<0\.8 values by persona pair, and Table[9](https://arxiv.org/html/2607.27384#A7.T9)reports pipeline operational metrics \(Agent 1 parse rate, Agent 3 fallback rate, tool usage, and latency\), underlying the discussion in §[4\.3](https://arxiv.org/html/2607.27384#S4.SS3)and §[4\.4](https://arxiv.org/html/2607.27384#S4.SS4)\. Table 7:NarrativeShield full results with 95% Wilson CIs omitted here for space; overall OMR 95% CIs are within±0\.02\\pm 0\.02of the point estimate for all models\.Table 8:NarrativeShield Cohen’sκ\\kappa, McNemar’s test \(pp\-values, uncorrected\), and DSS<<0\.8 rate by persona pair\.κ\\kappavalues are substantially higher than the corresponding B1 values \(Table[4](https://arxiv.org/html/2607.27384#A5.T4)\) for every model, consistent with the NAG collapse in Table[7](https://arxiv.org/html/2607.27384#A7.T7)\. Five of seven models show no significant directional bias on any McNemar pair; Llama\-3\.1\-8B\-Instruct retains significance on\(α,β\)\(\\alpha,\\beta\)and\(α,γ\)\(\\alpha,\\gamma\), consistent with its comparatively higher residual NAG \(0\.037\) among the six non\-degenerate models\. BioMistral’s near\-perfectκ\\kappa\(\>0\.96\>0\.96\) reflects near\-total lack of behavioral variance from a 1\.10% Agent 1 parse rate \(Table[9](https://arxiv.org/html/2607.27384#A7.T9)\) rather than well\-calibrated agreement, and should not be read as evidence of successful debiasing\.Table 9:NarrativeShield pipeline operational metrics\. Final parse rate is 100\.00% for every model, since Agent 3 fallback resolves any case Agent 2 leaves unparsed, so it is omitted as a column\. Agent 3 fallback rate is near\-zero for six models, indicating Agent 2 reliably commits to a parseable answer once given Agent 1’s structured output\. BioMistral’s combination of a 1\.10% Agent 1 parse rate and only a 3\.03% Agent 3 fallback rate indicates most of its unresolved cases originate in Agent 1’s extraction failure rather than Agent 2’s format compliance \(§[4\.4](https://arxiv.org/html/2607.27384#S4.SS4)\)\. ## Appendix HDSS Synthesis Across All Conditions Table[10](https://arxiv.org/html/2607.27384#A8.T10)reports the fraction of questions with DSS<0\.8<0\.8across all four conditions for every model, synthesizing the stability results discussed throughout §[4](https://arxiv.org/html/2607.27384#S4)\. Table 10:Fraction of questions with DSS<0\.8<0\.8across all four conditions\. Lower is more stable\. \*BioMistral’s B3 value is a degenerate artifact of its 0\.00% parse rate\. NarrativeShield achieves the lowest rate of any condition for every model\. ## Appendix IHuman Validation: Full Protocol and Statistics ### I\.1Protocol Three annotators with clinical backgrounds, none involved in dataset construction and none with access to auditor verdicts during rating, independently rated a stratified sample of 100 questions \(300 persona\-conditioned encounters: 100 each forPαP\_\{\\alpha\},PβP\_\{\\beta\},PγP\_\{\\gamma\}\) on three dimensions: clinical fact preservation \(binary\), persona register match \(binary\), and narrative realism \(5\-point Likert\)\. Annotators were shown the original vignette alongside a single persona rewrite and asked to flag any clinical fact, symptom, duration, vital sign, lab value, medication, or negated finding, present in the original but absent or altered in the rewrite\. Full annotation instructions and raw per\-annotator ratings are released with the dataset\. \\cellcolorblue\!10Criterion\\cellcolorblue\!10Raw / Adjacent Agree\.\\cellcolorblue\!10Chance\-Corr\. Agree\.\\cellcolorblue\!10NoteFacts Preserved100% \(3/3\)UndefinedPost\-filter ceilingRegister Match100% \(3/3\)UndefinedPost\-filter ceilingRealism \(all 3 annotators\)—α=−0\.014\\alpha=\-0\.0141 zero\-variance raterRealism \(2 informative\)42\.0% / 85\.3%α=0\.226\\alpha=0\.226Below tentative threshold; ordinal trendp<\.0001p<\.0001Realism agreement between the two informative annotators, by persona:\\cellcolorblue\!10Persona\\cellcolorblue\!10Exact\\cellcolorblue\!10Adjacent \(±\\pm1\)\\cellcolorblue\!10Mean Signed DiffPαP\_\{\\alpha\}\(control\)74\.0%95\.0%−0\.32\-0\.32PβP\_\{\\beta\}\(socioeconomic\)34\.0%91\.0%−0\.49\-0\.49PγP\_\{\\gamma\}\(cultural\)18\.0%70\.0%−1\.04\-1\.04Table 11:Human validation inter\-annotator agreement, 300 persona\-conditioned encounters, and per\-persona realism agreement between the two informative annotators \(mean signed difference is Annotator 1 minus Annotator 3\)\.One annotator rated all 300 items at ceiling \(5/5,SD=0\.000SD=0\.000\); we confirmed directly with this annotator that the flat rating reflected genuine judgment rather than a data\-entry artifact\. Restricting to the two annotators with non\-degenerate ratings, pooled realism means werePα=4\.89P\_\{\\alpha\}=4\.89,Pβ=4\.64P\_\{\\beta\}=4\.64,Pγ=4\.19P\_\{\\gamma\}=4\.19, an ordering that held independently for every one of the three annotators \(Kruskal\-WallisH=138\.69H=138\.69,p<\.0001p<\.0001; all pairwise Mann\-WhitneyUUcomparisons significant after Bonferroni correction\)\. The direction of disagreement between the two informative annotators was systematic rather than random: Annotator 3 rated realism higher than Annotator 1 on 157 of 300 items, versus the reverse on only 17 \(Wilcoxon signed\-rank test,p<\.0001p<\.0001\), and this gap widened monotonically with persona difficulty\.
Similar Articles
Same Patient, Different Words, Different Diagnosis? Evaluating Semantic Stability in Clinical LLMs
This paper proposes a semantic verification framework using Natural Language Inference (NLI) to evaluate the sensitivity of clinical LLMs to meaning-preserving prompt variations, introducing metrics such as MVS, ΔC, and WCI. Results show that domain specialization does not consistently improve robustness, with both domain-specific and general-purpose models showing mixed performance.
Localizing Anchoring Pathways in Language Models
This paper investigates how irrelevant numbers in prompts cause anchoring effects in language models and localizes the internal pathways carrying this signal using attribution-based circuit methods on Qwen and Llama models.
Artificial Intolerance: Stigmatizing Language in Clinical Documentation Skews Large Language Model Decision-Making
This study demonstrates that large language models inherit and amplify biases from stigmatizing language in clinical notes, leading to less aggressive patient management, and that current mitigation strategies are insufficient.
LLM Doesn't Know What It Doesn't Know: Detecting Epistemic Blind Spots via Cross-Model Attribution Divergence on Clinical Tabular Data
This paper explores Large Language Models' inability to recognize their knowledge limits on structured clinical data, proposing a cross-model attribution divergence method to detect epistemic blind spots. The approach improves calibration and accuracy without training by combining few-shot examples and SHAP-derived feature evidence.
Do Large Language Models Always Tell The Same Stories?
This paper investigates whether large language models generate diverse stories. Using narrative similarity analysis, the authors find that LLM-generated narratives are consistently more similar to each other than human-written stories, and that common mitigation strategies like negative prompting and temperature scaling fail to address this homogeneity.