Evaluating Large Language Models in Dynamic Clinical Decision-Making with Standardized Patient Cases
Summary
Researchers introduce MedSP1000, a 1,638-case interactive benchmark derived from standardized patient scenarios to evaluate LLMs as dynamic clinical agents across multi-turn encounters. Results show even the best model (GPT-5.5) completes only 60.4% of expert rubric items, suggesting current LLMs are not yet reliable enough for clinical practice.
Similar Articles
Large Language Models Versus Physicians in Traditional Chinese Medicine: A Real-World Clinical Case Evaluation
The study evaluates 16 large language models against 60 practicing TCM physicians using real-world clinical cases, finding that LLMs achieve higher expert scores in some areas but show discrepancies and safety concerns. It highlights the potential and limitations of LLMs in TCM decision support.
ClinicalMC: A Benchmark for Multi-Course Clinical Decision-Making with Large Language Models
ClinicalMC is a benchmark designed to evaluate large language models in multi-course clinical decision-making, featuring datasets in Chinese and English and a multi-agent evaluation framework.
Ask Before You Diagnose: Safe-Psych, a Sequential Evaluation Benchmark for LLMs in Psychiatry
Introduces Safe-Psych, a sequential benchmark for evaluating how large language models handle diagnostic uncertainty in psychiatry, revealing that even strong models often fail to abstain or seek clarification when clinical evidence is incomplete.
MedGuideX: Internalizing Decision Logic from Executable Guidelines into Large Language Models for Clinical Reasoning
MedGuideX transforms clinical practice guidelines into executable decision logic to generate factual and counterfactual QA data for training medical LLMs, achieving a 10.28% relative improvement in average accuracy across clinical reasoning benchmarks.
MTDiag: A Multi-Turn Diagnostic Dataset Towards Clinically Meaningful LLM Evaluation
This paper introduces MTDiag, a multi-turn diagnostic dialogue dataset for evaluating Large Language Models in realistic clinical diagnostic scenarios, addressing limitations of static QA benchmarks.