Measuring Epistemic Resilience of LLMs Under Misleading Medical Context

Hugging Face Daily Papers 06/10/26, 12:00 AM Papers

Summary

Introduces MedMisBench to measure LLMs' ability to maintain correct medical reasoning under misleading context. Shows that accuracy drops sharply from 71.1% to 38.0% under adversarial conditions, with potential harm flagged by clinical panel.

Large language models (LLMs) now reach expert-level scores on medical licensing exams, encouraging the assumption that high scores imply safe medical judgment while patients increasingly use them for health advice. We show this assumption is fragile: when misleading context is injected into questions that LLMs originally answer correctly, they abandon the correct answer. We call the ability to maintain correct judgment under adversarial context epistemic resilience, and introduce MedMisBench to measure it. MedMisBench contains 10,932 medical question items and 48,889 misleading context-option pairs spanning medical reasoning, agentic capability, and patient-journey evaluation. Across 11 model configurations, mean accuracy falls from 71.1% on original questions to 38.0% under focused misleading context, with 51.5% attack success. The most damaging injections are formal, rule-like fabrications: authority-framed falsehoods reach 69.5% attack success and exception-poisoning claims reach 64.1%. A 14-member clinical panel from 7 countries identified serious potential harm in 38.2% of reviewed cases. MedMisBench exposes a structural blind spot in LLM evaluation in medical settings: existing benchmarks measure what models know, but not whether they preserve correct medical judgment under misleading context.

Original Article

View Cached Full Text

Cached at: 06/15/26, 09:03 AM

Paper page - Measuring Epistemic Resilience of LLMs Under Misleading Medical Context

Source: https://huggingface.co/papers/2606.12291 Authors:

Abstract

Large language models demonstrate reduced medical reasoning accuracy when exposed to misleading context, highlighting a critical gap in current evaluation methods that fails to assess epistemic resilience under adversarial conditions.

Large language models (LLMs) now reach expert-level scores onmedical licensing exams, encouraging the assumption that high scores imply safe medical judgment while patients increasingly use them for health advice. We show this assumption is fragile: whenmisleading contextis injected into questions that LLMs originally answer correctly, they abandon the correct answer. We call the ability to maintain correct judgment under adversarial contextepistemic resilience, and introduceMedMisBenchto measure it.MedMisBenchcontains 10,932 medical question items and 48,889misleading context-option pairs spanning medical reasoning, agentic capability, and patient-journey evaluation. Across 11 model configurations, mean accuracy falls from 71.1% on original questions to 38.0% under focusedmisleading context, with 51.5%attack success. The most damaging injections are formal, rule-like fabrications:authority-framed falsehoodsreach 69.5%attack successandexception-poisoning claimsreach 64.1%. A 14-member clinical panel from 7 countries identified serious potential harm in 38.2% of reviewed cases.MedMisBenchexposes a structural blind spot in LLM evaluation in medical settings: existing benchmarks measure what models know, but not whether they preserve correct medical judgment undermisleading context.

View arXiv page View PDF Project page Add to collection

Get this paper in your agent:

hf papers read 2606\.12291

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2606.12291 in a model README.md to link it from this page.

Datasets citing this paper1

#### HongjianZhou/MedMisBench Viewer• Updatedabout 4 hours ago • 10.9k • 1

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2606.12291 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context

Paper page - Measuring Epistemic Resilience of LLMs Under Misleading Medical Context

Abstract

Models citing this paper0

Datasets citing this paper1

Spaces citing this paper0

Collections including this paper0

Similar Articles

When Correct Beliefs Collapse: Epistemic Resilience of LLMs under Clinical Pressure

Do Benchmarks Underestimate LLM Performance? Evaluating Hallucination Detection With LLM-First Human-Adjudicated Assessment

Auditing LLM Benchmarks with Item Response Theory

Evaluating Second-Order Bias of LLMs Through Epistemic Entitlement

Human-LLM Dialogue Improves Diagnostic Accuracy in Emergency Care

Submit Feedback

Similar Articles

When Correct Beliefs Collapse: Epistemic Resilience of LLMs under Clinical Pressure

Do Benchmarks Underestimate LLM Performance? Evaluating Hallucination Detection With LLM-First Human-Adjudicated Assessment

Auditing LLM Benchmarks with Item Response Theory

Evaluating Second-Order Bias of LLMs Through Epistemic Entitlement

Human-LLM Dialogue Improves Diagnostic Accuracy in Emergency Care