medical-benchmark

Tag

Cards List
#medical-benchmark

Same Facts, Different Diagnosis: Measuring and Mitigating Narrative Anchoring in Clinical Language Models

arXiv cs.CL ↗ · 2026-07-31 Cached

This paper introduces Narrative Anchoring, a failure mode where clinical language models produce divergent diagnoses when identical clinical facts are expressed in different sociolinguistic registers. The authors release a USMLE-derived dataset and propose NarrativeShield, a three-agent pipeline that reduces the anchoring gap to near-zero.

0 favorites 0 likes
#medical-benchmark

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context

Hugging Face Daily Papers ↗ · 2026-06-10 Cached

Introduces MedMisBench to measure LLMs' ability to maintain correct medical reasoning under misleading context. Shows that accuracy drops sharply from 71.1% to 38.0% under adversarial conditions, with potential harm flagged by clinical panel.

0 favorites 0 likes
← Back to home

Submit Feedback