Measuring Epistemic Resilience of LLMs Under Misleading Medical Context
Summary
Introduces MedMisBench to measure LLMs' ability to maintain correct medical reasoning under misleading context. Shows that accuracy drops sharply from 71.1% to 38.0% under adversarial conditions, with potential harm flagged by clinical panel.
View Cached Full Text
Cached at: 06/15/26, 09:03 AM
Paper page - Measuring Epistemic Resilience of LLMs Under Misleading Medical Context
Source: https://huggingface.co/papers/2606.12291 Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
Large language models demonstrate reduced medical reasoning accuracy when exposed to misleading context, highlighting a critical gap in current evaluation methods that fails to assess epistemic resilience under adversarial conditions.
Large language models (LLMs) now reach expert-level scores onmedical licensing exams, encouraging the assumption that high scores imply safe medical judgment while patients increasingly use them for health advice. We show this assumption is fragile: whenmisleading contextis injected into questions that LLMs originally answer correctly, they abandon the correct answer. We call the ability to maintain correct judgment under adversarial contextepistemic resilience, and introduceMedMisBenchto measure it.MedMisBenchcontains 10,932 medical question items and 48,889misleading context-option pairs spanning medical reasoning, agentic capability, and patient-journey evaluation. Across 11 model configurations, mean accuracy falls from 71.1% on original questions to 38.0% under focusedmisleading context, with 51.5%attack success. The most damaging injections are formal, rule-like fabrications:authority-framed falsehoodsreach 69.5%attack successandexception-poisoning claimsreach 64.1%. A 14-member clinical panel from 7 countries identified serious potential harm in 38.2% of reviewed cases.MedMisBenchexposes a structural blind spot in LLM evaluation in medical settings: existing benchmarks measure what models know, but not whether they preserve correct medical judgment undermisleading context.
View arXiv pageView PDFProject pageAdd to collection
Get this paper in your agent:
hf papers read 2606\.12291
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2606.12291 in a model README.md to link it from this page.
Datasets citing this paper1
#### HongjianZhou/MedMisBench Viewer• Updatedabout 4 hours ago • 10.9k • 1
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2606.12291 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
When Correct Beliefs Collapse: Epistemic Resilience of LLMs under Clinical Pressure
This paper investigates how large language models maintain correct beliefs under adversarial pressure in clinical settings, proposing R-FT fine-tuning to improve epistemic resilience while balancing corrigibility, and demonstrating significant robustness gains on medical benchmarks.
Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning
Introduces MedPIC-Bench, a benchmark with counterfactual questions to evaluate whether LLMs correctly apply medication-safety rules when patient-specific conditions change; across 28 LLMs, accuracy drops significantly on counterfactual questions, revealing a common failure to revise judgments.
Evaluating medical AI under missing information: same-provider judges and human raters change apparent safety
This paper extends missing-information stress-testing to open-ended medical conversation, finding that LLM judge choice materially changes apparent safety and that LLM judges are more permissive than clinicians.
Can LLMs Follow Medical Expert Logic? A Benchmark for Hierarchical Logical Consistency in Risk-of-Bias Assessment
This paper introduces LogiMed-RoB, a benchmark for evaluating large language models' hierarchical logical consistency in medical risk-of-bias assessment, revealing that high atomic consistency can conceal critical reasoning flaws in clinical deployment.
Diagnostic Foundation for Evaluating LLMs' Research Integrity as Co-Scientists
This paper introduces IntegrityBench, a benchmark for evaluating whether LLMs uphold research integrity when acting as co-scientists under institutional pressure. Findings show frontier models fail roughly 1 in 3 integrity-critical decisions under peak pressure, and that ethical action does not require accurate misconduct classification.