Large Language Models Versus Physicians in Traditional Chinese Medicine: A Real-World Clinical Case Evaluation
Summary
The study evaluates 16 large language models against 60 practicing TCM physicians using real-world clinical cases, finding that LLMs achieve higher expert scores in some areas but show discrepancies and safety concerns. It highlights the potential and limitations of LLMs in TCM decision support.
View Cached Full Text
Cached at: 09/17/26, 08:47 AM
# Large Language Models Versus Physicians in Traditional Chinese Medicine: A Real-World Clinical Case Evaluation Source: [https://arxiv.org/abs/2609.17544](https://arxiv.org/abs/2609.17544) Authors:[Jiacheng Xie](https://arxiv.org/search/cs?searchtype=author&query=Xie,+J),[Xiaoting Tang](https://arxiv.org/search/cs?searchtype=author&query=Tang,+X),[Yang Yu](https://arxiv.org/search/cs?searchtype=author&query=Yu,+Y),[Jinpu Li](https://arxiv.org/search/cs?searchtype=author&query=Li,+J),[Shouli Li](https://arxiv.org/search/cs?searchtype=author&query=Li,+S),[Congcong Jing](https://arxiv.org/search/cs?searchtype=author&query=Jing,+C),[Yantao Yang](https://arxiv.org/search/cs?searchtype=author&query=Yang,+Y),[Zhiyong Zhao](https://arxiv.org/search/cs?searchtype=author&query=Zhao,+Z),[Ziyang Zhang](https://arxiv.org/search/cs?searchtype=author&query=Zhang,+Z),[Qilin Song](https://arxiv.org/search/cs?searchtype=author&query=Song,+Q),[Guanghui An](https://arxiv.org/search/cs?searchtype=author&query=An,+G),[Dong Xu](https://arxiv.org/search/cs?searchtype=author&query=Xu,+D) [View PDF](https://arxiv.org/pdf/2609.17544) > Abstract:Large language models \(LLMs\) are increasingly being explored for clinical applications, yet their assessment for real\-world traditional Chinese medicine \(TCM\) practice remains limited We constructed a clinical case library comprising 349 de\-identified outpatient cases from 62 hospitals and evaluated 16 LLMs and a comparator cohort of 60 practicing TCM physicians using 60 representative cases selected from this library\. Model outputs and physician reports were anonymized and scored by five senior TCM experts across nine diagnostic and therapeutic dimensions\. Cutting\-edge general\-purpose LLMs achieved higher expert scores than the physician comparators, particularly for medical advice, treatment principles and selected diagnostic tasks\. However, prescription\-level analyses revealed discrepancies in herb selection, dosage, and treatment strategy, and qualitative safety review identified hallucinations and undesirable template\-driven outputs\. These findings highlight the potential of LLMs for TCM decision support while underscoring the need for physician oversight, safety constraints and prospective clinical evaluation\. ## Submission history From: Jiacheng Xie \[[view email](https://arxiv.org/show-email/98da81a0/2609.17544)\] **\[v1\]**Wed, 15 Jul 2026 03:18:48 UTC \(6,006 KB\)
Similar Articles
Evaluating Large Language Models in Dynamic Clinical Decision-Making with Standardized Patient Cases
Researchers introduce MedSP1000, a 1,638-case interactive benchmark derived from standardized patient scenarios to evaluate LLMs as dynamic clinical agents across multi-turn encounters. Results show even the best model (GPT-5.5) completes only 60.4% of expert rubric items, suggesting current LLMs are not yet reliable enough for clinical practice.
ClinicalMC: A Benchmark for Multi-Course Clinical Decision-Making with Large Language Models
ClinicalMC is a benchmark designed to evaluate large language models in multi-course clinical decision-making, featuring datasets in Chinese and English and a multi-agent evaluation framework.
Large Language Models as Unified Multimodal Learners for Clinical Prediction
The paper proposes converting multimodal patient data (text, labs, vitals) into a single natural language sequence and fine-tuning LLMs for clinical prediction, achieving comparable or better performance than specialized fusion architectures across three tasks.
Accuracy and Reliability of Large Language Models in Cosmetic Chemistry and Skin Health: A Benchmarking Study
This study benchmarks 14 large language models on cosmetic chemistry and skin health topics, finding poor accuracy and reliability, especially in technical and quantitative tasks. It concludes that current general-purpose LLMs are not reliable for informed consumer decision-making without further fine-tuning and algorithmic improvements.
LLMs as Post-hoc Auditors of Physiological Plausibility in Symbolic Regression: A Clinician-Evaluated Case Study
The paper investigates using Large Language Models as post-hoc auditors to evaluate symbolic regression models for interpretability and medical plausibility, with clinician assessments showing comparative model rankings are more favorably perceived than term-level interpretations.