Towards Precision Therapy in Hepatocellular Carcinoma: A Clinical-Reasoning LLM for Risk Stratification and Treatment Guidance
Summary
This paper presents HCC-STAR, a clinically aligned large language model for risk stratification and treatment guidance in hepatocellular carcinoma, aiming to improve precision therapy by leveraging electronic medical records.
View Cached Full Text
Cached at: 07/10/26, 06:09 AM
# Towards Precision Therapy in Hepatocellular Carcinoma: A Clinical-Reasoning LLM for Risk Stratification and Treatment Guidance Source: [https://arxiv.org/abs/2607.08602](https://arxiv.org/abs/2607.08602) ## Computer Science \> Artificial Intelligence **arXiv:2607\.08602**\(cs\) Authors:[Peng Cui](https://arxiv.org/search/cs?searchtype=author&query=Cui,+P),[Jitao Wang](https://arxiv.org/search/cs?searchtype=author&query=Wang,+J),[Siyan Xue](https://arxiv.org/search/cs?searchtype=author&query=Xue,+S),[Yao Huang](https://arxiv.org/search/cs?searchtype=author&query=Huang,+Y),[Haoming Xia](https://arxiv.org/search/cs?searchtype=author&query=Xia,+H),[Dong Li](https://arxiv.org/search/cs?searchtype=author&query=Li,+D),[Dengxiang Liu](https://arxiv.org/search/cs?searchtype=author&query=Liu,+D),[Weilin Wang](https://arxiv.org/search/cs?searchtype=author&query=Wang,+W),[Liping Liu](https://arxiv.org/search/cs?searchtype=author&query=Liu,+L),[Leida Zhang](https://arxiv.org/search/cs?searchtype=author&query=Zhang,+L),[Yunfu Cui](https://arxiv.org/search/cs?searchtype=author&query=Cui,+Y),[Tao Peng](https://arxiv.org/search/cs?searchtype=author&query=Peng,+T),[Daolin Ji](https://arxiv.org/search/cs?searchtype=author&query=Ji,+D),[Haitao Zhao](https://arxiv.org/search/cs?searchtype=author&query=Zhao,+H),[Wei Zhang](https://arxiv.org/search/cs?searchtype=author&query=Zhang,+W),[Xiaojuan Wang](https://arxiv.org/search/cs?searchtype=author&query=Wang,+X),[Weijie Ma](https://arxiv.org/search/cs?searchtype=author&query=Ma,+W),[Zongren Ding](https://arxiv.org/search/cs?searchtype=author&query=Ding,+Z),[Jinlong Li](https://arxiv.org/search/cs?searchtype=author&query=Li,+J),[Yuan Ding](https://arxiv.org/search/cs?searchtype=author&query=Ding,+Y),[Jiajing Zhao](https://arxiv.org/search/cs?searchtype=author&query=Zhao,+J),[Zhiyu Chen](https://arxiv.org/search/cs?searchtype=author&query=Chen,+Z),[Chengkun Yang](https://arxiv.org/search/cs?searchtype=author&query=Yang,+C),[Ziyue Huang](https://arxiv.org/search/cs?searchtype=author&query=Huang,+Z),[Jiaqi Liu](https://arxiv.org/search/cs?searchtype=author&query=Liu,+J),[Fusheng Liu](https://arxiv.org/search/cs?searchtype=author&query=Liu,+F),[Yang Zhou](https://arxiv.org/search/cs?searchtype=author&query=Zhou,+Y),[Xiaojuan Wang](https://arxiv.org/search/cs?searchtype=author&query=Wang,+X),[Zhongquan Sun](https://arxiv.org/search/cs?searchtype=author&query=Sun,+Z),[Shiyun Bao](https://arxiv.org/search/cs?searchtype=author&query=Bao,+S),[Xiaojun Wang](https://arxiv.org/search/cs?searchtype=author&query=Wang,+X),[Ming Yang](https://arxiv.org/search/cs?searchtype=author&query=Yang,+M),[Guangxin Li](https://arxiv.org/search/cs?searchtype=author&query=Li,+G),[Bin Shu](https://arxiv.org/search/cs?searchtype=author&query=Shu,+B),[Yong Liao](https://arxiv.org/search/cs?searchtype=author&query=Liao,+Y),[Hongxuan Li](https://arxiv.org/search/cs?searchtype=author&query=Li,+H),[Yao Tang](https://arxiv.org/search/cs?searchtype=author&query=Tang,+Y),[Shizhong Yang](https://arxiv.org/search/cs?searchtype=author&query=Yang,+S),[Yongyi Zeng](https://arxiv.org/search/cs?searchtype=author&query=Zeng,+Y),[Yufeng Yuan](https://arxiv.org/search/cs?searchtype=author&query=Yuan,+Y),[Yinpeng Dong](https://arxiv.org/search/cs?searchtype=author&query=Dong,+Y),[Jihui Hao](https://arxiv.org/search/cs?searchtype=author&query=Hao,+J),[Jun Zhu](https://arxiv.org/search/cs?searchtype=author&query=Zhu,+J),[Jiahong Dong](https://arxiv.org/search/cs?searchtype=author&query=Dong,+J) [View PDF](https://arxiv.org/pdf/2607.08602) > Abstract:Hepatocellular carcinoma \(HCC\) is a common malignancy and a leading cause of cancer\-related mortality\. Current guidelines and staging systems provide coarse categories, but often miss within\-stage heterogeneity and the clinical context in electronic medical records \(EMRs\)\. We present HCC\-STAR \(Hepatocellular Carcinoma Staging, Treatment And pRognosis\), a clinically aligned large language model that reads routine EMR narratives and jointly outputs risk score\-based staging, ranked guideline\-consistent treatments with evidence\-based rationales, and individualized survival estimates\. We curated about 30,000 HCC cases from SEER and expanded them into EMR\-style narrative training data using a clinician\-validated, prompt\-based augmentation workflow\. On this corpus, we developed a knowledge\-aligned reasoning framework optimized with a step\-verifiable composite reward, moving beyond text\-level memorization of clinical guidelines\. In a multi\-center cohort of 6,668 patients from 12 hospitals in China, HCC\-STAR achieved state\-of\-the\-art performance in treatment recommendation and risk stratification compared with clinical guidelines and competitive models, including GPT\-5 and Gemini\-2\.5 Pro\. Hypothetical overall\-survival analysis showed a median survival of 51 months under adherence to HCC\-STAR recommendations, compared with 29 and 32 months under BCLC and CNLC\. In clinician\-centric evaluations, blinded hepatobiliary specialists rated HCC\-STAR's reasoning and evidence\-based justifications as trustworthy\. The model surpassed resident and attending physicians in treatment accuracy and helped physicians make more accurate decisions faster when used as an assistant\. These findings support HCC\-STAR as a reliable and verifiable decision\-support system for risk stratification and precision therapy in HCC\. ## Submission history From: Peng Cui \[[view email](https://arxiv.org/show-email/dccd4cb5/2607.08602)\] **\[v1\]**Thu, 9 Jul 2026 15:33:08 UTC \(17,722 KB\) Bibliographic Tools ## Bibliographic and Citation Tools Bibliographic Explorer Toggle Code, Data, Media ## Code, Data and Media Associated with this Article Demos ## Demos Related Papers ## Recommenders and Search Tools About arXivLabs ## arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website\. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy\. arXiv is committed to these values and only works with partners that adhere to them\. Have an idea for a project that will add value for arXiv's community?[**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html)\.
Similar Articles
The Large Cancer Assistant (LCA): A Model-Agnostic Orchestration Framework for Scalable Clinical Decision Support in Oncology
This paper introduces the Large Cancer Assistant (LCA), a model-agnostic orchestration framework for scalable clinical decision support in oncology that decouples multimodal data ingestion from AI inference using a 7-tuple architecture and Algorithmic Impermeability.
LLMs for Cardiovascular Risk Prediction from Structured Clinical Data
This paper presents a hybrid framework that combines structured clinical data with LLM-generated narratives for coronary artery disease prediction, achieving high fidelity in variable extraction and comparing ML models with LLM-based zero-shot and few-shot classification.
Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support
This Perspective paper argues that large language models are not yet safe for autonomous clinical decision support, particularly in triage of undifferentiated patients, due to lack of robust evaluation under incomplete information and asymmetric costs of missed diagnoses.
ChatHealthAI: Aligning Electronic Health Record Representations with Large Language Models for Grounded Clinical Reasoning
ChatHealthAI is a multimodal reasoning framework that aligns structured EHR representations with a frozen LLM to enable grounded clinical reasoning while maintaining predictive performance.
Evaluating Large Language Models in Dynamic Clinical Decision-Making with Standardized Patient Cases
Researchers introduce MedSP1000, a 1,638-case interactive benchmark derived from standardized patient scenarios to evaluate LLMs as dynamic clinical agents across multi-turn encounters. Results show even the best model (GPT-5.5) completes only 60.4% of expert rubric items, suggesting current LLMs are not yet reliable enough for clinical practice.