Large Language Models Versus Physicians in Traditional Chinese Medicine: A Real-World Clinical Case Evaluation

arXiv cs.CL Papers

Summary

The study evaluates 16 large language models against 60 practicing TCM physicians using real-world clinical cases, finding that LLMs achieve higher expert scores in some areas but show discrepancies and safety concerns. It highlights the potential and limitations of LLMs in TCM decision support.

arXiv:2609.17544v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly being explored for clinical applications, yet their assessment for real-world traditional Chinese medicine (TCM) practice remains limited We constructed a clinical case library comprising 349 de-identified outpatient cases from 62 hospitals and evaluated 16 LLMs and a comparator cohort of 60 practicing TCM physicians using 60 representative cases selected from this library. Model outputs and physician reports were anonymized and scored by five senior TCM experts across nine diagnostic and therapeutic dimensions. Cutting-edge general-purpose LLMs achieved higher expert scores than the physician comparators, particularly for medical advice, treatment principles and selected diagnostic tasks. However, prescription-level analyses revealed discrepancies in herb selection, dosage, and treatment strategy, and qualitative safety review identified hallucinations and undesirable template-driven outputs. These findings highlight the potential of LLMs for TCM decision support while underscoring the need for physician oversight, safety constraints and prospective clinical evaluation.
Original Article
View Cached Full Text

Cached at: 09/17/26, 08:47 AM

# Large Language Models Versus Physicians in Traditional Chinese Medicine: A Real-World Clinical Case Evaluation
Source: [https://arxiv.org/abs/2609.17544](https://arxiv.org/abs/2609.17544)
Authors:[Jiacheng Xie](https://arxiv.org/search/cs?searchtype=author&query=Xie,+J),[Xiaoting Tang](https://arxiv.org/search/cs?searchtype=author&query=Tang,+X),[Yang Yu](https://arxiv.org/search/cs?searchtype=author&query=Yu,+Y),[Jinpu Li](https://arxiv.org/search/cs?searchtype=author&query=Li,+J),[Shouli Li](https://arxiv.org/search/cs?searchtype=author&query=Li,+S),[Congcong Jing](https://arxiv.org/search/cs?searchtype=author&query=Jing,+C),[Yantao Yang](https://arxiv.org/search/cs?searchtype=author&query=Yang,+Y),[Zhiyong Zhao](https://arxiv.org/search/cs?searchtype=author&query=Zhao,+Z),[Ziyang Zhang](https://arxiv.org/search/cs?searchtype=author&query=Zhang,+Z),[Qilin Song](https://arxiv.org/search/cs?searchtype=author&query=Song,+Q),[Guanghui An](https://arxiv.org/search/cs?searchtype=author&query=An,+G),[Dong Xu](https://arxiv.org/search/cs?searchtype=author&query=Xu,+D)

[View PDF](https://arxiv.org/pdf/2609.17544)

> Abstract:Large language models \(LLMs\) are increasingly being explored for clinical applications, yet their assessment for real\-world traditional Chinese medicine \(TCM\) practice remains limited We constructed a clinical case library comprising 349 de\-identified outpatient cases from 62 hospitals and evaluated 16 LLMs and a comparator cohort of 60 practicing TCM physicians using 60 representative cases selected from this library\. Model outputs and physician reports were anonymized and scored by five senior TCM experts across nine diagnostic and therapeutic dimensions\. Cutting\-edge general\-purpose LLMs achieved higher expert scores than the physician comparators, particularly for medical advice, treatment principles and selected diagnostic tasks\. However, prescription\-level analyses revealed discrepancies in herb selection, dosage, and treatment strategy, and qualitative safety review identified hallucinations and undesirable template\-driven outputs\. These findings highlight the potential of LLMs for TCM decision support while underscoring the need for physician oversight, safety constraints and prospective clinical evaluation\.

## Submission history

From: Jiacheng Xie \[[view email](https://arxiv.org/show-email/98da81a0/2609.17544)\] **\[v1\]**Wed, 15 Jul 2026 03:18:48 UTC \(6,006 KB\)

Similar Articles