Auditable agentic AI for evidence-grounded thyroid ultrasound diagnosis and reporting

arXiv cs.AI Papers

Summary

This paper presents ThyroidXAgent, a clinician-interactive agentic AI system that coordinates specialized diagnostic tools for thyroid ultrasound, storing outputs as auditable evidence records. Developed on a large multicentre dataset, it achieves strong results in nodule segmentation, benign-malignant classification, and report generation.

arXiv:2608.12590v1 Announce Type: new Abstract: Thyroid ultrasound diagnosis requires coordinated lesion localization, measurement, risk stratification and reporting, yet most AI systems address these tasks in isolation and provide limited support for clinical review. We present ThyroidXAgent, a clinician-interactive agentic AI system that coordinates specialized diagnostic tools and stores their outputs as an auditable case-level evidence record. The system was developed using OpenThyroidDB, a multicentre, multitask resource integrating approximately 0.3 million ultrasound images and 24,000 paired reports, and was evaluated on 28,458 non-overlapping test cases, including 8,721 cases from 35 centres in the private NHC-MISD-TUS cohort. Across heterogeneous datasets, ThyroidXAgent achieved a mean Dice score of 87.21 percent for nodule segmentation and a mean AUROC of 0.9466 for benign-malignant classification. The same workflow supported lymph-node metastasis prediction and follicular versus papillary thyroid carcinoma classification, with AUROCs of 0.864 and 0.805, respectively. For report generation, evidence-grounded assembly outperformed multimodal language-model baselines across three cohorts. ThyClinScore, a lesion-level clinical semantic metric introduced here, showed the strongest correlation with a location-aware language-model judge. ThyroidXAgent improved physician classification accuracy, increased report diagnostic consistency from 70.3 percent to 86.2 percent, and reduced segmentation and reporting time by 35.9 percent and 27.4 percent, respectively. These findings support auditable, clinician-correctable agentic AI for thyroid ultrasound diagnosis and reporting.
Original Article
View Cached Full Text

Cached at: 08/14/26, 09:26 AM

# Auditable agentic AI for evidence-grounded thyroid ultrasound diagnosis and reporting
Source: [https://arxiv.org/html/2608.12590](https://arxiv.org/html/2608.12590)
Haifan GongShiyu ChenAffiliation:School of Computer Science and Engineering, Sun Yat\-sen University, Guangzhou, ChinaBodong WangAffiliation:School of Software Engineering, Sun Yat\-sen University, Zhuhai, ChinaYuqi WangAffiliation:Independent Researcher, Jersey City, NJ, USAShijie WangAffiliation:School of Mathematical Sciences, Zhejiang University, Hangzhou, ChinaGuoliang YouAffiliation:Department of Radiology, Perelman School of Medicine, University of Pennsylvania, Philadelphia, PA, USAXinyu XiongAffiliation:School of Computer Science and Engineering, Sun Yat\-sen University, Guangzhou, ChinaHaowei WangAffiliation:Department of Pathology, Zhujiang Hospital, Southern Medical University, Guangzhou, ChinaMingzhi MaoAffiliation:School of Software Engineering, Sun Yat\-sen University, Zhuhai, ChinaDexing KongAffiliation:School of Mathematical Sciences, Zhejiang University, Hangzhou, ChinaQinghua LiuAffiliation:Department of Health Management, Zhujiang Hospital, Southern Medical University, Guangzhou, ChinaAffiliation:Corresponding authors:Qinghua Liu \([13760633321@163\.com](mailto:[email protected])\), Wei Lou \([louwei@zjnu\.edu\.cn](mailto:[email protected])\), Fei Chen \([gzchenfei@126\.com](mailto:[email protected])\), and Guanbin Li \([liguanbin@mail\.sysu\.edu\.cn](mailto:[email protected])\)Wei LouAffiliation:College of Mathematical Medicine, Zhejiang Normal University, Jinhua, ChinaAffiliation:Corresponding authors:Qinghua Liu \([13760633321@163\.com](mailto:[email protected])\), Wei Lou \([louwei@zjnu\.edu\.cn](mailto:[email protected])\), Fei Chen \([gzchenfei@126\.com](mailto:[email protected])\), and Guanbin Li \([liguanbin@mail\.sysu\.edu\.cn](mailto:[email protected])\)Fei ChenAffiliation:Department of Thyroid Surgery, Zhujiang Hospital, Southern Medical University, Guangzhou, ChinaAffiliation:Corresponding authors:Qinghua Liu \([13760633321@163\.com](mailto:[email protected])\), Wei Lou \([louwei@zjnu\.edu\.cn](mailto:[email protected])\), Fei Chen \([gzchenfei@126\.com](mailto:[email protected])\), and Guanbin Li \([liguanbin@mail\.sysu\.edu\.cn](mailto:[email protected])\)Guanbin LiAffiliation:School of Computer Science and Engineering, Sun Yat\-sen University, Guangzhou, ChinaAffiliation:Corresponding authors:Qinghua Liu \([13760633321@163\.com](mailto:[email protected])\), Wei Lou \([louwei@zjnu\.edu\.cn](mailto:[email protected])\), Fei Chen \([gzchenfei@126\.com](mailto:[email protected])\), and Guanbin Li \([liguanbin@mail\.sysu\.edu\.cn](mailto:[email protected])\)

###### 摘要

Thyroid ultrasound diagnosis requires coordinated lesion localization, measurement, risk stratification and reporting, yet most AI systems address these tasks in isolation and provide limited support for clinical review\. We present ThyroidXAgent, a clinician\-interactive agentic AI system that coordinates specialized diagnostic tools and stores their outputs as an auditable case\-level evidence record\. The system was developed using OpenThyroidDB, a multicentre, multitask resource integrating approximately 0\.3 million ultrasound images and 24,000 paired reports, and was evaluated on 28,458 non\-overlapping test cases, including 8,721 cases from 35 centres in the private NHC\-MISD\-TUS cohort\. Across heterogeneous datasets, ThyroidXAgent achieved a mean Dice score of 87\.21% for nodule segmentation and a mean AUROC of 0\.9466 for benign–malignant classification\. The same workflow supported lymph\-node metastasis prediction and follicular versus papillary thyroid carcinoma classification, with AUROCs of 0\.864 and 0\.805, respectively\. For report generation, evidence\-grounded assembly outperformed multimodal language\-model baselines across three cohorts\. ThyClinScore, a lesion\-level clinical semantic metric introduced here, showed the strongest correlation with a location\-aware language\-model judge\. ThyroidXAgent improved physician classification accuracy, increased report diagnostic consistency from 70\.3% to 86\.2%, and reduced segmentation and reporting time by 35\.9% and 27\.4%, respectively\. These findings support auditable, clinician\-correctable agentic AI for thyroid ultrasound diagnosis and reporting\.

## 1Introduction

Thyroid ultrasound diagnosis depends on a sequence of lesion\-level observations and clinical decisions\[[1](https://arxiv.org/html/2608.12590#bib.bib1),[2](https://arxiv.org/html/2608.12590#bib.bib2)\]\. A clinically useful examination localizes and measures thyroid nodules\[[1](https://arxiv.org/html/2608.12590#bib.bib1),[2](https://arxiv.org/html/2608.12590#bib.bib2)\], characterizes sonographic features\[[2](https://arxiv.org/html/2608.12590#bib.bib2),[3](https://arxiv.org/html/2608.12590#bib.bib3),[4](https://arxiv.org/html/2608.12590#bib.bib4)\], assigns risk using systems such as the Thyroid Imaging Reporting and Data System \(TI\-RADS\)\[[3](https://arxiv.org/html/2608.12590#bib.bib3)\], determines whether fine\-needle aspiration is indicated\[[1](https://arxiv.org/html/2608.12590#bib.bib1),[3](https://arxiv.org/html/2608.12590#bib.bib3)\], integrates cytology when available\[[5](https://arxiv.org/html/2608.12590#bib.bib5)\]and communicates the findings in a structured report\[[2](https://arxiv.org/html/2608.12590#bib.bib2),[3](https://arxiv.org/html/2608.12590#bib.bib3)\]\. Each step introduces variability\. Descriptors such as margins and echogenic foci\[[3](https://arxiv.org/html/2608.12590#bib.bib3),[4](https://arxiv.org/html/2608.12590#bib.bib4)\]show substantial reader dependence, and small changes in these features can alter biopsy or follow\-up recommendations\[[3](https://arxiv.org/html/2608.12590#bib.bib3),[4](https://arxiv.org/html/2608.12590#bib.bib4)\]\. These requirements make thyroid ultrasound a workflow\-level problem for clinical AI\[[6](https://arxiv.org/html/2608.12590#bib.bib6),[7](https://arxiv.org/html/2608.12590#bib.bib7)\]: useful systems must support prediction, preserve the evidence behind each step\[[8](https://arxiv.org/html/2608.12590#bib.bib8)\]and allow clinicians to revise that evidence when needed\[[9](https://arxiv.org/html/2608.12590#bib.bib9),[10](https://arxiv.org/html/2608.12590#bib.bib10)\]\.

Most thyroid ultrasound AI systems have focused on individual components of this workflow, including segmentation\[[11](https://arxiv.org/html/2608.12590#bib.bib11),[12](https://arxiv.org/html/2608.12590#bib.bib12),[13](https://arxiv.org/html/2608.12590#bib.bib13)\], classification\[[14](https://arxiv.org/html/2608.12590#bib.bib14),[15](https://arxiv.org/html/2608.12590#bib.bib15),[16](https://arxiv.org/html/2608.12590#bib.bib16),[17](https://arxiv.org/html/2608.12590#bib.bib17)\], and report generation\[[18](https://arxiv.org/html/2608.12590#bib.bib18),[19](https://arxiv.org/html/2608.12590#bib.bib19),[20](https://arxiv.org/html/2608.12590#bib.bib20)\]\. Multicenter systems\[[15](https://arxiv.org/html/2608.12590#bib.bib15),[21](https://arxiv.org/html/2608.12590#bib.bib21)\]and feature\-aligned multimodal models\[[16](https://arxiv.org/html/2608.12590#bib.bib16),[17](https://arxiv.org/html/2608.12590#bib.bib17),[19](https://arxiv.org/html/2608.12590#bib.bib19)\]have connected image predictions to risk descriptors\[[16](https://arxiv.org/html/2608.12590#bib.bib16),[17](https://arxiv.org/html/2608.12590#bib.bib17)\]or management recommendations\[[15](https://arxiv.org/html/2608.12590#bib.bib15),[17](https://arxiv.org/html/2608.12590#bib.bib17)\]\. Recent studies have extended thyroid AI to fine\-needle aspiration cytology\[[21](https://arxiv.org/html/2608.12590#bib.bib21)\], lateral lymph\-node metastasis prediction\[[22](https://arxiv.org/html/2608.12590#bib.bib22)\]and rare thyroid cancer subtype classification\[[23](https://arxiv.org/html/2608.12590#bib.bib23)\]\. Many systems still present their outputs as endpoints: a mask, probability, label or report\-like text\. The intermediate evidence that supports these outputs is often unavailable for clinical review\[[8](https://arxiv.org/html/2608.12590#bib.bib8),[9](https://arxiv.org/html/2608.12590#bib.bib9)\], correction\[[9](https://arxiv.org/html/2608.12590#bib.bib9),[10](https://arxiv.org/html/2608.12590#bib.bib10)\]or reuse across downstream tasks\[[8](https://arxiv.org/html/2608.12590#bib.bib8),[24](https://arxiv.org/html/2608.12590#bib.bib24),[25](https://arxiv.org/html/2608.12590#bib.bib25)\]\. This endpoint\-oriented design makes it difficult to determine whether an AI result is supported by appropriate lesion localization, measurement, sonographic features and report statements\.

This limitation reflects a broader challenge in medical AI\. High\-impact clinical AI studies increasingly emphasize workflow integration\[[26](https://arxiv.org/html/2608.12590#bib.bib26),[9](https://arxiv.org/html/2608.12590#bib.bib9),[27](https://arxiv.org/html/2608.12590#bib.bib27)\], human\-AI collaboration\[[6](https://arxiv.org/html/2608.12590#bib.bib6),[28](https://arxiv.org/html/2608.12590#bib.bib28),[29](https://arxiv.org/html/2608.12590#bib.bib29),[30](https://arxiv.org/html/2608.12590#bib.bib30),[31](https://arxiv.org/html/2608.12590#bib.bib31),[32](https://arxiv.org/html/2608.12590#bib.bib32),[33](https://arxiv.org/html/2608.12590#bib.bib33)\], and evidence beyond retrospective performance\[[34](https://arxiv.org/html/2608.12590#bib.bib34),[6](https://arxiv.org/html/2608.12590#bib.bib6),[7](https://arxiv.org/html/2608.12590#bib.bib7),[26](https://arxiv.org/html/2608.12590#bib.bib26)\]\. The cognitive consequences of AI\-supported clinical work\[[24](https://arxiv.org/html/2608.12590#bib.bib24)\], clinician interaction with algorithmic recommendations\[[10](https://arxiv.org/html/2608.12590#bib.bib10)\]and the transition of AI from a tool to a clinical teammate\[[35](https://arxiv.org/html/2608.12590#bib.bib35),[32](https://arxiv.org/html/2608.12590#bib.bib32)\]have also become central considerations\. Generalist medical AI\[[36](https://arxiv.org/html/2608.12590#bib.bib36),[37](https://arxiv.org/html/2608.12590#bib.bib37),[38](https://arxiv.org/html/2608.12590#bib.bib38),[39](https://arxiv.org/html/2608.12590#bib.bib39)\]and multimodal foundation models\[[36](https://arxiv.org/html/2608.12590#bib.bib36),[39](https://arxiv.org/html/2608.12590#bib.bib39),[40](https://arxiv.org/html/2608.12590#bib.bib40),[41](https://arxiv.org/html/2608.12590#bib.bib41),[42](https://arxiv.org/html/2608.12590#bib.bib42),[43](https://arxiv.org/html/2608.12590#bib.bib43)\]extend this ambition to flexible inputs and outputs across tasks\. For thyroid ultrasound, however, generality must be connected to specialty\-specific requirements: lesion\-level measurement\[[1](https://arxiv.org/html/2608.12590#bib.bib1),[3](https://arxiv.org/html/2608.12590#bib.bib3)\], sonographic feature attribution\[[3](https://arxiv.org/html/2608.12590#bib.bib3),[16](https://arxiv.org/html/2608.12590#bib.bib16)\], anatomical context\[[2](https://arxiv.org/html/2608.12590#bib.bib2)\], guideline\-aligned management\[[1](https://arxiv.org/html/2608.12590#bib.bib1),[3](https://arxiv.org/html/2608.12590#bib.bib3),[15](https://arxiv.org/html/2608.12590#bib.bib15)\]and auditable reporting\[[8](https://arxiv.org/html/2608.12590#bib.bib8),[19](https://arxiv.org/html/2608.12590#bib.bib19),[9](https://arxiv.org/html/2608.12590#bib.bib9)\]\. We therefore treat the case\-level evidence record, rather than a single prediction endpoint, as the central object of AI assistance\.

Agent\-based workflows\[[44](https://arxiv.org/html/2608.12590#bib.bib44),[35](https://arxiv.org/html/2608.12590#bib.bib35),[45](https://arxiv.org/html/2608.12590#bib.bib45),[46](https://arxiv.org/html/2608.12590#bib.bib46),[47](https://arxiv.org/html/2608.12590#bib.bib47),[48](https://arxiv.org/html/2608.12590#bib.bib48)\]provide one way to implement this evidence\-centerd formulation\. In this setting, the agent’s role is coordination rather than direct image interpretation\[[44](https://arxiv.org/html/2608.12590#bib.bib44),[35](https://arxiv.org/html/2608.12590#bib.bib35),[45](https://arxiv.org/html/2608.12590#bib.bib45)\]: it acts as a workflow controller that plans case\-specific analysis\[[49](https://arxiv.org/html/2608.12590#bib.bib49),[50](https://arxiv.org/html/2608.12590#bib.bib50),[51](https://arxiv.org/html/2608.12590#bib.bib51)\], routes inputs to specialized tools\[[52](https://arxiv.org/html/2608.12590#bib.bib52),[45](https://arxiv.org/html/2608.12590#bib.bib45),[46](https://arxiv.org/html/2608.12590#bib.bib46)\], maintains intermediate state\[[53](https://arxiv.org/html/2608.12590#bib.bib53)\]and exposes structured evidence for human review\[[8](https://arxiv.org/html/2608.12590#bib.bib8),[9](https://arxiv.org/html/2608.12590#bib.bib9),[35](https://arxiv.org/html/2608.12590#bib.bib35)\]\. Medical\-agent benchmarks increasingly emphasize these capabilities in interactive settings\[[54](https://arxiv.org/html/2608.12590#bib.bib54),[50](https://arxiv.org/html/2608.12590#bib.bib50),[51](https://arxiv.org/html/2608.12590#bib.bib51)\]that require retrieval, action execution and workflow\-level reasoning\[[52](https://arxiv.org/html/2608.12590#bib.bib52),[54](https://arxiv.org/html/2608.12590#bib.bib54),[53](https://arxiv.org/html/2608.12590#bib.bib53)\]\. For thyroid ultrasound, the relevant evidence objects include images, lesion masks and measurements\[[11](https://arxiv.org/html/2608.12590#bib.bib11),[12](https://arxiv.org/html/2608.12590#bib.bib12),[55](https://arxiv.org/html/2608.12590#bib.bib55)\], radiomic descriptors and risk estimates\[[56](https://arxiv.org/html/2608.12590#bib.bib56),[16](https://arxiv.org/html/2608.12590#bib.bib16),[17](https://arxiv.org/html/2608.12590#bib.bib17)\], report clauses\[[57](https://arxiv.org/html/2608.12590#bib.bib57),[18](https://arxiv.org/html/2608.12590#bib.bib18),[19](https://arxiv.org/html/2608.12590#bib.bib19)\], uncertainty signals\[[58](https://arxiv.org/html/2608.12590#bib.bib58)\]and clinician corrections\[[30](https://arxiv.org/html/2608.12590#bib.bib30),[31](https://arxiv.org/html/2608.12590#bib.bib31),[9](https://arxiv.org/html/2608.12590#bib.bib9)\]\. The agent is therefore useful insofar as it can coordinate these objects into an auditable clinical workflow\.

Here we present ThyroidXAgent, a clinician\-interactive agentic system that reframes thyroid ultrasound AI around an auditable case\-level evidence record rather than a collection of isolated prediction endpoints\. Instead of directly interpreting ultrasound images with a general\-purpose multimodal model, ThyroidXAgent acts as a workflow controller that plans case\-specific analyses, routes inputs to specialized tools and maintains structured intermediate evidence, including lesion masks, measurements, class probabilities, radiomic descriptors, uncertainty signals and report clauses\. These evidence objects remain inspectable and editable by clinicians, and corrections can be propagated to subsequent analysis and reporting steps\. We evaluate this formulation across nodule segmentation, benign\-malignant classification, malignant\-lesion stratification and case\-level report generation using 40 heterogeneous multicentre datasets and clinician reader studies\. We further introduce ThyClinScore, a lesion\-level semantic metric that evaluates whether generated reports preserve clinically relevant evidence rather than surface\-level wording alone\. ThyroidXAgent improved produced more clinically consistent reports and reduced clinician workload while retaining an editable evidence trace\. Together, these results establish evidence\-centred orchestration as an alternative to both standalone predictive models and unconstrained end\-to\-end medical agents\.

![Refer to caption](https://arxiv.org/html/2608.12590v1/Introduction3.png)Figure 1:Multicenter thyroid ultrasound data and the ThyroidXAgent evidence workflow\.a, OpenThyroidDB integrates curated public thyroid ultrasound resources and institutionally governed clinical cohorts into a full\-spectrum resource for segmentation, benign\-malignant classification, report generation and malignant\-lesion stratification across diverse scanners and acquisition settings\.b, The clinician\-feedback workflow stores AI\-generated masks, predictions, report clauses and clinician corrections as case\-level evidence, allowing intermediate outputs to be reviewed, revised and reused in later steps\.c, ThyroidXAgent acts as a workflow controller for four evidence\-producing workflows: expert\-refined nodule segmentation, clinician\-verified malignancy classification, evidence\-grounded structured reporting, and advanced diagnosis for lymph\-node metastasis assessment and PTC/FTC subtype analysis\.d, Multicenter validation design, with model development/internal validation followed by independent external validation\.![Refer to caption](https://arxiv.org/html/2608.12590v1/ThyroidXAgent_SegCls_performance.png)Figure 2:Agent\-routed evidence for thyroid nodule segmentation, benign\-malignant classification and clinician review\.a, Interactive review workflow: clinicians assess the predicted nodule mask, accept the classification if the mask is reliable, or refine it through box annotation and interactive segmentation; SHAP\-based feature attributions are then recomputed from the corrected mask\.b\-e, Segmentation and classification performance across heterogeneous thyroid ultrasound benchmarks, evaluated using Dice, HD95, AUROC and AUPRC\.f, Cohort\-level SHAP beeswarm analysis for benign\-malignant classification\.g, ROC curves on the 500\-image physician comparison set, showing ThyroidXAgent, representative AI baselines and clinician operating points before and after ThyroidXAgent support; both clinicians improved with evidence support\.h, Pooled performance on the NHC\-MISD\-TUS private external test set across nine models and four metrics: nodule Dice, HD95, binary AUROC and AUPRC\. Bars show point estimates with 95% CIs, and the dashed line indicates the classification chance level of 0\.5\.i, Segmentation time for manual and AI\-assisted workflows\.j, Ranked within\-case time savings\.k, Paired Dice distributions showing preserved segmentation quality with improved efficiency\.
## 2Results

### 2\.1ThyroidXAgent organizes thyroid ultrasound diagnosis as an auditable case\-level evidence workflow

We developed ThyroidXAgent to organize thyroid ultrasound diagnosis as a case\-level evidence workflow \(Fig\.[1](https://arxiv.org/html/2608.12590#S1.F1)\)\. For each examination, the agent routes the case through tool\-callable steps, including nodule segmentation, measurement, benign\-malignant classification, radiomics extraction, malignant\-lesion stratification and report generation\. The workflow stores tool outputs in a shared evidence record that can be inspected, corrected and reused across downstream tasks\.

OpenThyroidDB provides the multicenter data resource underlying this workflow\-level formulation\. The segmentation and classification analyses included seven cohorts from six clinical centres across China, Vietnam and Colombia, acquired on at least eight ultrasound platforms \(Supplementary Table[S1](https://arxiv.org/html/2608.12590#Ax1.T1)\)\. Four cohorts served as internal training data\. TN3K\[[11](https://arxiv.org/html/2608.12590#bib.bib11)\], collected at Zhujiang Hospital, Southern Medical University, Guangzhou, comprised 4,633 training, 100 validation and 614 test images acquired on GE Logiq E9, ARIETTA 850 and RESONA 70B scanners\. TN5K\[[55](https://arxiv.org/html/2608.12590#bib.bib55)\], from the Cancer Hospital, Chinese Academy of Medical Sciences, Beijing, comprised 3,500 training, 500 validation and 1,000 test images acquired on GE Logiq E9 and GE S7 scanners with 5\-12 or 8\-15 MHz probes\. ThyroidXL\[[59](https://arxiv.org/html/2608.12590#bib.bib59)\], from the Vietnam National Hospital of Endocrinology, Hanoi, comprised 9,441 training, 100 validation and 2,090 test images acquired on a Hitachi Aloka Arietta V70 scanner\. PKTN\[[13](https://arxiv.org/html/2608.12590#bib.bib13)\], from Peking University First Hospital, Beijing, comprised 703 training, 150 validation and 150 test images; the acquisition device was not disclosed\. Three cohorts provided independent external test sets\. DDTI\[[60](https://arxiv.org/html/2608.12590#bib.bib60)\], from IDIME, Bogotá, Colombia, contributed 637 images \(637 for segmentation, 349 of which also carry benign\-malignant classification labels\) acquired on TOSHIBA Nemio 30 and Nemio MX scanners with a 12 MHz probe\. RJH\-7K\[[61](https://arxiv.org/html/2608.12590#bib.bib61)\], from Ruijin Hospital, Shanghai Jiao Tong University School of Medicine, Shanghai, contributed 7,288 segmentation\-only images acquired on multiple scanners whose models were not specified\. ZJH\-8K, collected at Zhujiang Hospital, Southern Medical University, Guangzhou, contributed 7,958 images \(426 benign cases with 3,202 images and 723 malignant cases with 4,756 images\) used for both segmentation and classification; the acquisition device was not disclosed\. Overall, the segmentation and classification benchmark comprised 38,864 images \(18,277 training, 850 validation, 19,737 test\) from seven cohorts spanning six centres\.

The report\-generation analyses included four cohorts from three centres\. SMU\-HMC comprised 23,955 reports from 21,954 patients, with 248,194 associated ultrasound images\. A patient\-level held\-out set of 400 patients, comprising 5,984 images, was reserved for internal testing\. After excluding all examinations from these internally held\-out patients, the remaining SMU\-HMC cohort was used to develop the relevant component models, and 7,147 quality\-controlled reports were selected to construct the retrieval template library\. The KMVE analysis cohort\[[18](https://arxiv.org/html/2608.12590#bib.bib18)\]comprised 2,457 cases and 4,914 image assignments, partitioned according to the original dataset split into 1,719 training, 246 validation and 492 test cases; its training partition was additionally used for template\-library construction\. External evaluation included two complementary cohorts\. ZJH\-TS provided an independent external\-centre test set of 150 reports selected from the original collection of 353 reports, after excluding reports dominated by postoperative findings\. TNVideo provided an independently assembled case\-level cohort for the external reader study, comprising 148 ultrasound examinations, of which 145 had annotation\-derived diagnostic labels\. Overall, the report\-generation analyses comprised four cohorts from three centres, including an internal SMU\-HMC test set, an independent external\-centre test set and an external reader\-study cohort\.

For malignant\-lesion stratification, the 4,756 malignant images from ZJH\-8K \(also an external test set for segmentation and classification\) served as the primary training set, of which 20 cases \(183 images\) were held out for validation\. For lateral lymph\-node metastasis \(LNM\) prediction, 180 images from LymphUs Center 1 were additionally included as training data, and the 158 images from LymphUs Center 2 served as the independent external test set\. For follicular \(FTC\) versus papillary \(PTC\) thyroid carcinoma subtype classification, the 200 public images released by Daiet al\.\[[62](https://arxiv.org/html/2608.12590#bib.bib62)\]served as the external test set\. Within the ZJH\-8K malignant cohort, 533 cases \(3,394 images\) were classified as CN0 and 190 cases \(1,362 images\) as CN1 for lymph\-node status, and 656 cases \(4,312 images\) were classic PTC and 67 cases \(444 images\) were follicular variant for subtype\.

Collectively, these resources cover heterogeneous acquisition settings, file formats, annotation types and clinical tasks, including nodule segmentation, benign\-malignant classification, report generation and advanced malignant\-lesion analysis \(Fig\.[1](https://arxiv.org/html/2608.12590#S1.F1)a and Supplementary Table[S2](https://arxiv.org/html/2608.12590#Ax1.T2)\)\. Using this resource, ThyroidXAgent coordinates four evidence\-producing workflow branches: expert\-refined nodule segmentation, clinician\-verified malignancy classification, expert\-edited report generation and advanced diagnosis for lateral lymph\-node metastasis and PTC/FTC subtype analysis \(Fig\.[1](https://arxiv.org/html/2608.12590#S1.F1)b,c\)\. Each branch contributes structured evidence to the same case\-level record, including masks, measurements, class probabilities, radiomic descriptors, feature attributions, uncertainty signals and report clauses\. This design allows intermediate outputs to be reused across tasks\. A corrected mask can support radiomics extraction, a malignancy estimate can inform the report impression, and structured report evidence can be reviewed and edited rather than accepted as opaque text\.

The resulting workflow makes auditability a property of the diagnostic process rather than a post hoc explanation attached to a final answer\. Clinicians can inspect generated masks, correct segmentation errors, review SHAP\- or Grad\-CAM\-based explanations, edit report statements and return corrected outputs to the evidence store\. The subsequent results evaluate this formulation across automatic image analysis, clinician correction, malignant\-lesion stratification, clinical semantic report scoring and evidence\-grounded report assembly\.

### 2\.2Agent\-routed evidence improves cross\-dataset segmentation and classification

We first evaluated the two image\-analysis tasks that anchor the downstream workflow: nodule segmentation and benign\-malignant classification\. For each case, ThyroidXAgent collected candidate masks and class probabilities from DINOv3\-based experts, selected or fused outputs using case\-level quality signals, extracted radiomic features from the selected lesion mask and stored confidence, disagreement and tabular predictions as structured evidence \(Supplementary Fig\.[S1](https://arxiv.org/html/2608.12590#Ax1.F1)\)\. This design tests whether tool routing and evidence consolidation improve robustness across heterogeneous ultrasound datasets, where dataset bias remains a major source of performance degradation\[[63](https://arxiv.org/html/2608.12590#bib.bib63)\]\.

On seven segmentation test sets, including the independent DDTI, RJH\-7K and ZJH\-8K cohorts, ThyroidXAgent achieved a mean Dice coefficient of 87\.21% and a mean 95th\-percentile Hausdorff distance \(HD95\) of 6\.90 mm \(Fig\.[2](https://arxiv.org/html/2608.12590#S1.F2)b,c and Supplementary Table[S3](https://arxiv.org/html/2608.12590#Ax1.T3)\)\. It obtained the highest Dice score on six of seven test sets and the lowest HD95 on all test sets, indicating improved boundary robustness across heterogeneous acquisition conditions\. The strongest baseline, MedSAM2\[[64](https://arxiv.org/html/2608.12590#bib.bib64)\], achieved a mean Dice of 85\.66%, with the largest gap on the external ZJH\-8K cohort \(86\.29% versus 94\.30% for ThyroidXAgent\)\. UltraFedFM\[[65](https://arxiv.org/html/2608.12590#bib.bib65)\], a medical imaging foundation model, reached 79\.66%\.

For benign\-malignant classification, ThyroidXAgent achieved a mean AUROC of 0\.9466 and a mean AUPRC of 0\.8361 across five test sets, including the independent DDTI and ZJH\-8K cohorts \(Fig\.[2](https://arxiv.org/html/2608.12590#S1.F2)d,e and Supplementary Table[S4](https://arxiv.org/html/2608.12590#Ax1.T4)\)\. Specialized image models showed weaker cross\-dataset consistency; for example, RepViT\[[66](https://arxiv.org/html/2608.12590#bib.bib66)\]reached AUROC of 0\.777 on ThyroidXL but 0\.556 on TN3K\. General\-purpose vision\-language models\[[67](https://arxiv.org/html/2608.12590#bib.bib67)\]also underperformed; GPT\-5\[[68](https://arxiv.org/html/2608.12590#bib.bib68)\]reached AUROC of 0\.611\-0\.774 across test sets, and Gemini\-2\.5\-Pro\[[69](https://arxiv.org/html/2608.12590#bib.bib69)\]reached 0\.616\-0\.687 \(Supplementary Table[S4](https://arxiv.org/html/2608.12590#Ax1.T4)\)\. These results support a division of labour in which domain\-specific image and radiomics tools generate the evidence, while the agent routes and consolidates tool outputs\. We further validated on NHC\-MISD\-TUS, an independent private multicentre test set \(Fig\.[2](https://arxiv.org/html/2608.12590#S1.F2)h and Supplementary Tables[S9](https://arxiv.org/html/2608.12590#Ax1.T9)\-[S12](https://arxiv.org/html/2608.12590#Ax1.T12)\)\. ThyroidXAgent led on both nodule segmentation \(Dice 82\.31%, HD95 9\.41 mm\) and binary classification \(AUROC 0\.819, AUPRC 0\.823\)\. By contrast, MedSAM2 and MedSegX failed on segmentation \(Dice ¡0\.3\), and zero\-shot vision\-language baselines \(BiomedCLIP\[[70](https://arxiv.org/html/2608.12590#bib.bib70)\], MedSigLIP\[[71](https://arxiv.org/html/2608.12590#bib.bib71)\]\) performed at or below chance on classification\. GPT\-5 and Gemini\-2\.5\-Pro could not be evaluated on this private intranet dataset\.

Beyond prediction accuracy, we examined whether the structured evidence provided interpretable signals for clinician review\. Cohort\-level SHAP profiles\[[72](https://arxiv.org/html/2608.12590#bib.bib72)\]showed that morphology\-related radiomic descriptors, especially Sphericity and Elongation, dominated benign\-malignant classification, whereas texture and intensity features contributed complementary information \(Fig\.[2](https://arxiv.org/html/2608.12590#S1.F2)f\)\. Representative cases confirmed that accurate segmentation produced SHAP attributions and Grad\-CAM maps aligned with visible nodule characteristics, whereas poor segmentation degraded these explanations \(Supplementary Fig\.[S3](https://arxiv.org/html/2608.12590#Ax1.F3)\)\.

To assess whether this evidence supports clinician decision\-making, we conducted a blinded 500\-image physician comparison \(Fig\.[2](https://arxiv.org/html/2608.12590#S1.F2)g\)\. ThyroidXAgent achieved AUROC of 0\.9256 and AUPRC of 0\.9250\. In the same comparison, UltraFedFM\[[65](https://arxiv.org/html/2608.12590#bib.bib65)\]reached AUROC of 0\.880, GPT\-5\[[68](https://arxiv.org/html/2608.12590#bib.bib68)\]0\.660 and Gemini\-2\.5\-Pro\[[69](https://arxiv.org/html/2608.12590#bib.bib69)\]0\.619\. When clinicians were provided with ThyroidXAgent’s structured evidence, including SHAP\-based feature attributions and nodule segmentation boundaries, both clinicians improved across all metrics\. For clinician 1, accuracy rose from 79\.2% to 84\.6% \(\+5\.4%\), precision from 83\.5% to 87\.5%, recall from 72\.8% to 80\.8%, specificity from 85\.6% to 88\.4%, and F1 from 0\.778 to 0\.840, while the false\-positive rate fell from 14\.4% to 11\.6%\. Clinician 2 improved more markedly: accuracy from 74\.0% to 84\.8% \(\+10\.8%\), precision from 75\.2% to 84\.3%, recall from 71\.6% to 85\.6%, specificity from 76\.4% to 84\.0%, and F1 from 0\.734 to 0\.849, with the false\-positive rate decreasing from 23\.6% to 16\.0%; clinician 2 thereby narrowly surpassed clinician 1 as the higher\-performing reader\.

### 2\.3Clinician correction reduces segmentation time while preserving quality

The same evidence representation supported clinician\-interactive segmentation review\. Clinicians inspected predicted masks, corrected segmentation errors when needed and returned the refined masks to the case\-level evidence store\. SHAP\-based feature attributions were then recomputed using the corrected mask, allowing downstream classification evidence to reflect clinician refinement \(Fig\.[2](https://arxiv.org/html/2608.12590#S1.F2)a\)\.

AI assistance reduced mean segmentation time from 14\.21 s to 9\.11 s per image, a 1\.6\-fold speedup, while preserving segmentation quality \(Fig\.[2](https://arxiv.org/html/2608.12590#S1.F2)i\-k\)\. AI\-assisted Dice \(0\.903\) matched or exceeded manual Dice \(0\.879\) in approximately two\-thirds of paired cases\. These results indicate that the intermediate evidence layer can reduce repetitive annotation work while keeping the segmentation boundary available for clinician correction\.

![Refer to caption](https://arxiv.org/html/2608.12590v1/Malignant_Image_tasks.png)Figure 3:Shared ThyroidXAgent tools support malignant\-lesion stratification with task\-specific radiomic attributions\.a, Workflow for SHAP\-based interpretation of malignant\-lesion tasks\.b, Performance comparison between ThyroidXAgent and the corresponding specialist baselines for FTC/PTC subtype classification and lymph node metastasis prediction, reported as AUROC and AUPRC; percentages denote the relative improvement of ThyroidXAgent over each baseline\.c,d, Global and representative local SHAP analyses for lymph node metastasis prediction\.e,f, Global and representative local SHAP analyses for FTC/PTC subtype classification, showing stronger contributions from texture heterogeneity and shape descriptors\.
### 2\.4Shared agent tools generate task\-specific evidence for malignant\-lesion stratification

We next tested whether the shared ThyroidXAgent tools could be redirected to clinically distinct malignant\-lesion stratification tasks after the primary thyroid nodule assessment\. Lateral lymph\-node metastasis \(LNM\) prediction informs surgical planning, whereas follicular \(FTC\) versus papillary \(PTC\) thyroid carcinoma subtype discrimination informs treatment strategy and follow\-up\. In a conventional development pipeline, each task would require a separate workflow for preprocessing, feature extraction, prediction, interpretation and reporting\. In ThyroidXAgent, the segmentation, radiomics extraction, tabular classification and routing logic were reused, with task\-specific classifier fine\-tuning and task instructions changed\.

This reuse enabled rapid adaptation to new clinical questions while preserving the same evidence structure\. ThyroidXAgent achieved AUROC of 0\.864 for LNM prediction on the 158\-image LymphUs Center 2 test set and 0\.805 for FTC/PTC subtype classification on the 200\-image Daiet al\.test set\[[62](https://arxiv.org/html/2608.12590#bib.bib62)\]\(Fig\.[3](https://arxiv.org/html/2608.12590#S2.F3)b and Supplementary Table[S5](https://arxiv.org/html/2608.12590#Ax1.T5)\), outperforming the specialist baselines LLNM\-Net\[[22](https://arxiv.org/html/2608.12590#bib.bib22)\]\(0\.767\) for LNM and Tiger\-Model\[[23](https://arxiv.org/html/2608.12590#bib.bib23)\]\(0\.714\) for FTC/PTC\.

The attribution profiles changed with the clinical task \(Fig\.[3](https://arxiv.org/html/2608.12590#S2.F3)a,c\-f\)\. LNM prediction relied more on lymph\-node position and size features, including distance to the thyroid capsule and lesion area\. FTC/PTC subtype classification relied more on texture heterogeneity and shape descriptors\. These task\-dependent attribution patterns indicate that ThyroidXAgent generated radiomic evidence according to the requested clinical question, rather than simply reusing the benign\-malignant decision rule\.

![Refer to caption](https://arxiv.org/html/2608.12590v1/ThyClinScore.png)Figure 4:ThyClinScore for clinical semantic evaluation of thyroid ultrasound reports\.a, ThyClinScore first structures the ground\-truth and predicted reports into lesion\-level entries, matches corresponding lesions, and scores clinically relevant attributes, including size, vascularity, morphology, lesion\-level F1 and completeness\.b, Pearson correlation matrix comparing conventional natural\-language generation metrics with clinical semantic metrics\. Overlap\-based metrics were highly correlated with one another, whereas the clinical semantic metrics captured complementary report\-quality dimensions\.c, Pearson correlations between each metric and a location\-aware LLM judge\. ThyClinScore showed the strongest correlation among the evaluated metrics; asterisks denote statistical significance \(\*p<0\.05p<0\.05, \*\*p<0\.01p<0\.01, \*\*\*p<0\.001p<0\.001\)\.d, Qualitative comparison showing that reports with similar wording overlap can differ in clinically important lesion attributes, which is reflected by ThyClinScore\.![Refer to caption](https://arxiv.org/html/2608.12590v1/ReportGen.png)Figure 5:Evidence\-grounded report assembly in ThyroidXAgent\.a, Case\-level thyroid ultrasound input, including video or image sequences, multiple views and anatomical regions, and greyscale and CDFI modalities\.b, Input\-preparation skills crop regions of interest, parse image context, detect CDFI information, classify nodules and anatomical regions, and convert the resulting preprocessing outputs into agent\-ready image priors\.c, Diagnostic planning combines the skill instructions, preprocessing information and tool contract with an LLM planner to generate a staged diagnostic task graph\.d, A ReAct\-style execution loop observes intermediate evidence, reasons over the next step and calls tools through an MCP server that exposes the public workflow API, including case initialization, plan approval, status checking, evidence retrieval and report generation\.e, Structured evidence is transformed into report text through query construction, template retrieval, slot filling and clause combination, so that report clauses remain linked to measurements, lesion descriptors, risk estimates and other case\-level evidence\.f, Report\-generation performance comparison on SMU\-HMC, KMVE and ZJH\-TS\.g, Radar plots comparing ThyroidXAgent with the static pipeline across lexical and clinical semantic metrics\.
### 2\.5ThyClinScore captures lesion\-level report errors missed by overlap metrics

We then addressed the evaluation of thyroid ultrasound reports\. Conventional natural\-language generation metrics, such as BLEU\[[73](https://arxiv.org/html/2608.12590#bib.bib73)\], ROUGE\[[74](https://arxiv.org/html/2608.12590#bib.bib74)\], and METEOR\[[75](https://arxiv.org/html/2608.12590#bib.bib75)\], primarily reward surface overlap and can miss clinically important disagreements in lesion location, size, vascularity, morphology or impression\. We therefore developed ThyClinScore, a clinical semantic metric that structures ground\-truth and generated reports, matches lesion entries and scores clinically relevant attributes \(Fig\.[4](https://arxiv.org/html/2608.12590#S2.F4)a\)\.

ThyClinScore captured report\-quality dimensions that were complementary to wording overlap\. Overlap\-based metrics were strongly correlated with one another, whereas the clinical semantic metrics captured distinct lesion\-level and feature\-level information \(Fig\.[4](https://arxiv.org/html/2608.12590#S2.F4)b\)\. Using a Pearson correlation\-based evaluation approach similar to that of Liet al\.\[[20](https://arxiv.org/html/2608.12590#bib.bib20)\], ThyClinScore showed the strongest correlation with a location\-aware LLM judge among the evaluated metrics \(Pearson’sr=0\.696r=0\.696,p<0\.001p<0\.001; Fig\.[4](https://arxiv.org/html/2608.12590#S2.F4)c\)\. Qualitative examples further show that reports with similar wording overlap can differ in clinically important attributes, which is reflected by the ThyClinScore components \(Fig\.[4](https://arxiv.org/html/2608.12590#S2.F4)d\)\.

### 2\.6Evidence\-grounded report assembly improves reporting consistency and efficiency

For report generation, ThyroidXAgent converted the case\-level evidence record into structured report text\. The agent first used multi\-view and multimodal thyroid ultrasound inputs to construct image priors, invoked diagnostic tools through planning and execution, and then assembled report clauses from structured facts using BM25 template retrieval, slot filling and clause combination \(Fig\.[5](https://arxiv.org/html/2608.12590#S2.F5)a\-g\)\. This design links report statements to intermediate evidence, including gland measurements, nodule location, lesion size, sonographic descriptors, vascularity, lymph\-node findings and diagnostic impressions, rather than generating unconstrained free text\. The report\-generation workflow was packaged as a reusable skill and exposed to external agents through a Workflow MCP server\. This interface provided high\-level case operations, including input preparation, case initiation, plan review, approval, report retrieval and evidence retrieval, while low\-level model calls remained internal to the workflow\. Auxiliary tool performance is summarized in Supplementary Table[S13](https://arxiv.org/html/2608.12590#Ax1.T13)\. Clinicians can review the structured evidence, edit generated statements and return corrected report content to the case\-level evidence store\.

On conventional natural\-language generation metrics, ThyroidXAgent achieved the strongest overall performance across SMU\-HMC, KMVE\[[18](https://arxiv.org/html/2608.12590#bib.bib18)\]and ZJH\-TS \(Fig\.[5](https://arxiv.org/html/2608.12590#S2.F5)f and Supplementary Table[S14](https://arxiv.org/html/2608.12590#Ax1.T14)\)\. On SMU\-HMC, BLEU\-1, BLEU\-4 and ROUGELreached 0\.5961, 0\.3405 and 0\.5450, respectively\. On KMVE, ThyroidXAgent ranked first on all reported overlap metrics, with BLEU\-1 of 0\.6209, BLEU\-4 of 0\.4465, METEOR of 0\.3596 and ROUGELof 0\.5826\. On ZJH\-TS, ThyroidXAgent achieved the highest BLEU\-1, BLEU\-4 and ROUGELvalues, reaching 0\.5134, 0\.2942 and 0\.5529, respectively\.

We further compared the adaptive tool\-routing workflow with a fixed rule\-based tool\-calling pipeline \(Fig\.[5](https://arxiv.org/html/2608.12590#S2.F5)g and Supplementary Table[S16](https://arxiv.org/html/2608.12590#Ax1.T16)\)\. Across all three report\-generation test sets, ThyroidXAgent produced larger multi\-metric profiles than the static pipeline\. ThyClinScore increased from 0\.4293 to 0\.5238 on SMU\-HMC, from 0\.3346 to 0\.4465 on KMVE and from 0\.3648 to 0\.4775 on ZJH\-TS\.

Clinical semantic evaluation showed that the reports retained clinically relevant information beyond wording similarity \(Fig\.[4](https://arxiv.org/html/2608.12590#S2.F4)and Supplementary Table[S15](https://arxiv.org/html/2608.12590#Ax1.T15)\)\. ThyroidXAgent achieved the highest ThyClinScore on SMU\-HMC \(0\.5238\), KMVE \(0\.4465\) and ZJH\-TS \(0\.4775\)\. It also achieved the highest lesion\-level F1 and report completeness on SMU\-HMC and ZJH\-TS, whereas KMVE showed a different submetric profile in which several individual components were led by other baselines\. Together with the metric\-correlation analysis \(Fig\.[4](https://arxiv.org/html/2608.12590#S2.F4)b,c\), these results indicate that clinical semantic scoring captures report\-quality information not represented by conventional overlap metrics\.

Finally, we assessed human\-AI cooperation in a cross\-over reader study\. Two physicians wrote reports for 145 thyroid ultrasound videos under manual and AI\-assisted workflows, with each case evaluated in both workflows by different physicians to reduce memory bias \(Fig\.[6](https://arxiv.org/html/2608.12590#S2.F6)a\)\. In the 145 cases with annotation\-derived diagnostic\-direction labels, AI\-assisted reports showed higher consistency than manual reports overall and within benign and malignant subsets \(Fig\.[6](https://arxiv.org/html/2608.12590#S2.F6)c\)\. AI assistance reduced mean reporting time from 2\.5 to 1\.8 min per case, a 27\.4% reduction, with similar time savings for both readers \(Fig\.[6](https://arxiv.org/html/2608.12590#S2.F6)d,e\)\. A representative malignant case shows that the structured evidence supported statements on gland morphology, nodule location, measurements, sonographic features and diagnostic impression, while leaving partially correct or incorrect statements available for clinician review \(Fig\.[6](https://arxiv.org/html/2608.12590#S2.F6)b\)\.

![Refer to caption](https://arxiv.org/html/2608.12590v1/ReaderStudy.png)Figure 6:Reader study of AI\-assisted thyroid ultrasound reporting\.a, Cross\-over reader\-study design\. Each ultrasound video was interpreted under both manual and AI\-assisted conditions by different physicians, reducing recall bias while enabling paired case\-level comparisons\.b, Representative malignant thyroid nodule case comparing reports generated by Qwen 3\.5, GPT\-5 and ThyroidXAgent with the reference report\. Text spans are annotated as correct, partially correct or incorrect according to medical\-semantic concordance; ThyroidXAgent shows closer agreement with the reference in this example\.c, Annotation\-based diagnostic\-direction consistency of manual and AI\-assisted reports, shown overall and stratified by benign and malignant cases\.d, Case\-level reporting\-time reduction, defined as manual minus AI\-assisted reporting time and ranked across paired cases\.e, Physician\-level reporting time under the manual and AI\-assisted conditions\.

## 3Discussion

This study reframes thyroid ultrasound AI as an auditable case\-level evidence workflow rather than a collection of isolated prediction tasks\. The principal contribution of ThyroidXAgent is not a new standalone segmentation, classification or report\-generation model, but a shared evidence architecture in which lesion masks, measurements, probabilities, radiomic descriptors, uncertainty signals and report clauses remain visible, editable and reusable across the diagnostic process\. Across heterogeneous datasets, this formulation improved nodule segmentation and benign\-malignant classification, supported adaptation to additional malignant\-lesion tasks, enhanced evidence\-grounded reporting and reduced clinician workload\. Auditability is therefore embedded in the workflow itself rather than appended retrospectively to a final prediction\.

ThyroidXAgent also defines a constrained and clinically practical role for agentic AI\. Rather than asking a general\-purpose vision\-language model to interpret ultrasound images directly, the agent plans case\-specific analyses, routes inputs to specialized tools, maintains case state and exposes intermediate results for review\. Image interpretation and quantitative measurement remain assigned to task\-specific models, whereas the agent provides coordination and evidence management across them\. This division of labour extends emerging medical\-agent frameworks centred on planning, tool use and coordinated clinical workflows\[[44](https://arxiv.org/html/2608.12590#bib.bib44),[35](https://arxiv.org/html/2608.12590#bib.bib35),[45](https://arxiv.org/html/2608.12590#bib.bib45)\]\. Its value lies less in unrestricted autonomy than in making heterogeneous model outputs coherent, traceable and correctable\.

The shared evidence record further distinguishes ThyroidXAgent from a conventional fixed pipeline\. A corrected lesion mask can be propagated to measurement, radiomics, classification and reporting, while anatomical context, malignancy estimates and lesion descriptors can be reused when constructing the diagnostic impression\. The same evidence\-producing tools could consequently be redirected towards lymph\-node metastasis prediction and thyroid carcinoma subtype classification without rebuilding the complete workflow\. Report generation provides a particularly stringent demonstration of this design: rather than producing unconstrained free text, ThyroidXAgent assembles report statements from structured clinical evidence\. ThyClinScore complements this approach by evaluating lesion matching, clinically relevant attributes and report completeness, thereby capturing errors in location, size or morphology that may be missed by conventional lexical\-overlap metrics\.

The reader studies indicate that this evidence\-centred formulation can support human\-AI cooperation without requiring clinicians to accept an opaque recommendation\. AI assistance reduced segmentation and reporting time while preserving meaningful points of clinical control, including lesion\-boundary correction, evidence inspection and report editing\. These findings should not be interpreted as evidence for replacing clinical judgement\. Instead, they suggest that agentic systems can reduce repetitive work while keeping consequential intermediate outputs available for verification and correction\. The effectiveness of such systems will ultimately depend not only on predictive accuracy, but also on whether clinicians can identify errors, understand their downstream consequences and efficiently intervene\.

Several limitations remain\. The evaluations were retrospective, and prospective multicentre studies are needed under real acquisition conditions, changing scanner settings and institution\-specific reporting practices\. The reader studies included a limited number of physicians and primarily assessed workflow feasibility and efficiency rather than patient\-level clinical benefit\. ThyroidXAgent also remains dependent on the validity and calibration of its component tools: routing cannot compensate for systematically biased segmentation, incomplete metadata or poorly generalized classifiers, although visible intermediate evidence may make these failures easier to detect\[[34](https://arxiv.org/html/2608.12590#bib.bib34),[76](https://arxiv.org/html/2608.12590#bib.bib76)\]\. More generalizable tool design may further improve cross\-domain robustness\[[77](https://arxiv.org/html/2608.12590#bib.bib77),[78](https://arxiv.org/html/2608.12590#bib.bib78)\]\. Finally, the template\-based reporting strategy may require local adaptation, and learning from clinician corrections will require governance for data quality, privacy, provenance, model versioning and distribution shift\. Future work should therefore determine whether evidence\-level corrections improve downstream clinical decisions prospectively and whether the shared evidence architecture can support additional imaging tasks without sacrificing reliability or interpretability\.

## 4Methods

### 4\.1Datasets and task definitions

For image segmentation and benign\-malignant classification, OpenThyroidDB integrates seven public and institutional ultrasound sources comprising 38,864 images \(18,277 training, 850 validation, 19,737 test; Supplementary Table[S1](https://arxiv.org/html/2608.12590#Ax1.T1)\)\. TN3K\[[11](https://arxiv.org/html/2608.12590#bib.bib11)\], TN5K\[[55](https://arxiv.org/html/2608.12590#bib.bib55)\], ThyroidXL\[[59](https://arxiv.org/html/2608.12590#bib.bib59)\]and PKTN\[[13](https://arxiv.org/html/2608.12590#bib.bib13)\]served as internal cohorts for nodule segmentation training; TN3K, TN5K and ThyroidXL additionally provided benign\-malignant classification labels\. DDTI\[[60](https://arxiv.org/html/2608.12590#bib.bib60)\], RJH\-7K\[[61](https://arxiv.org/html/2608.12590#bib.bib61)\]and ZJH\-8K \(collected at Zhujiang Hospital, Southern Medical University\) served as independent external test sets: DDTI and ZJH\-8K for both tasks, RJH\-7K for segmentation only\. To construct the expert pool, we merged training portions across datasets into stacked training sets, with the largest containing 18,277 images\. The included cohorts differed markedly in sample size, class balance and lesion characteristics \(Supplementary Fig\.[S2](https://arxiv.org/html/2608.12590#Ax1.F2)\)\.

For malignant\-lesion stratification, two additional tasks were evaluated: lateral lymph\-node metastasis \(LNM\) prediction and follicular \(FTC\) versus papillary \(PTC\) thyroid carcinoma subtype classification\. The training data for both tasks comprised the 4,756 malignant images from ZJH\-8K, of which 20 cases \(183 images\) were held out for validation\. Within this cohort, 533 cases \(3,394 images\) were CN0 and 190 cases \(1,362 images\) CN1 for lymph\-node status, and 656 cases \(4,312 images\) were classic PTC and 67 cases \(444 images\) follicular variant; all labels were confirmed by post\-surgical histopathology\. For LNM prediction, 180 images from LymphUs Center 1\[[79](https://arxiv.org/html/2608.12590#bib.bib79),[80](https://arxiv.org/html/2608.12590#bib.bib80)\]were added to the training set \(total 4,753 images\), and the 158 images from LymphUs Center 2 served as the independent external test set\. LymphUs is a multicenter open\-access database of patients with histologically confirmed PTC and binary LNM labels confirmed by fine\-needle aspiration biopsy, acquired using a Samsung Medison RS80 scanner \(5\-12 MHz L5\-12/60 transducer\) at Center 1 and a SuperSonic Imagine AixPlorer Ultimate scanner \(4\-15 MHz SL15\-4 transducer\) at Center 2\. For FTC\-versus\-PTC classification, the training and validation splits were unchanged, and the 200 public images from Dai et al\.\[[62](https://arxiv.org/html/2608.12590#bib.bib62)\]served as the external test set\.

### 4\.2Agent workflow controller and evidence store

ThyroidXAgent decomposes each case into tool calls and structured intermediate outputs\. The workflow contains an expert pool for image segmentation and classification, a radiomics branch, a tabular prediction branch, post hoc explanation modules, anatomical\-context parsers, measurement tools and report\-generation modules\. The LLM router operates on structured summaries rather than raw ultrasound images\. During inference, it receives candidate masks, class probabilities, confidence estimates, radiomic descriptors and metadata such as image resolution, device and data source\. It emits a strict JSON decision that records the selected output, supporting evidence and uncertainty signals\. These outputs are normalized into a case\-level evidence store containing masks, measurements, class probabilities, radiomic descriptors, explanation objects, warnings and report clauses\. The evidence store is used for final prediction, report assembly and clinician review, and clinician corrections can be written back to the same record for subsequent use\.

### 4\.3Segmentation, classification and radiomics

ThyroidXAgent coordinates a heterogeneous expert pool, an LLM router and a radiomics branch for image segmentation and classification: the expert pool generates candidate masks and class probabilities, the router selects the most reliable candidate based on quality metrics and case\-level metadata, and the radiomics branch provides an independent classification signal\. For image segmentation and classification, the expert pool is designed as a heterogeneous ensemble rather than a single model to reduce sensitivity to dataset bias\[[81](https://arxiv.org/html/2608.12590#bib.bib81),[63](https://arxiv.org/html/2608.12590#bib.bib63)\]\(Supplementary Fig\.[S1](https://arxiv.org/html/2608.12590#Ax1.F1)\)\. Each expert typically uses a DINOv3\-based backbone\[[82](https://arxiv.org/html/2608.12590#bib.bib82)\]with a task\-specific lightweight head, though task\-specific external models can also be assembled for specialized classification targets\. To encourage complementary generalization profiles, these experts were trained under varying configurations along three dimensions: stacked\-training composition, input resolution \(128, 224 and 448 pixels\), and whether DINOv3 pretrained weights were loaded or the backbone was trained from scratch\. For the segmentation branch, backbone dilation rates were additionally varied to produce experts with different receptive\-field profiles\. This heterogeneity exposes individual experts to progressively broader data distributions, so that the router can select the most reliable candidate for each case rather than relying on a single model’s bias\. The effect of these stacked\-training configurations on cross\-dataset performance is reported in Supplementary Table[S6](https://arxiv.org/html/2608.12590#Ax1.T6)\.

Within this pool, the segmentation branch uses a U\-Net\-style decoder with skip fusion to preserve fine boundary detail, and is optimized with a combined loss that balances pixel\-level supervision with region\-level overlap:

ℒseg=ℒwBCE​\(m,m^\)\+ℒIoU​\(m,m^\),\\mathcal\{L\}\_\{\\text\{seg\}\}=\\mathcal\{L\}\_\{\\text\{wBCE\}\}\(m,\\hat\{m\}\)\+\\mathcal\{L\}\_\{\\text\{IoU\}\}\(m,\\hat\{m\}\),wheremmdenotes the ground\-truth mask,m^\\hat\{m\}the prediction,ℒwBCE\\mathcal\{L\}\_\{\\text\{wBCE\}\}the class\-weighted binary cross\-entropy andℒIoU\\mathcal\{L\}\_\{\\text\{IoU\}\}the intersection\-over\-union loss\. The classification branch pools backbone features using global average and max pooling, followed by a compact attention head\. To mitigate class imbalance, generalized logit adjustment \(GLA\)\[[83](https://arxiv.org/html/2608.12590#bib.bib83)\]is applied to the classification logits:

z~c=zc\+τ​log⁡\(πc\),\\tilde\{z\}\_\{c\}=z\_\{c\}\+\\tau\\log\(\\pi\_\{c\}\),wherezcz\_\{c\}is the logit for classcc,πc\\pi\_\{c\}is the empirical class prior estimated from the training set, andτ\\tauis adaptively set based on the class imbalance ratio\. Binary classification tasks use BCE on the adjusted logits, whereas multi\-class tasks use cross\-entropy on the adjusted logits\. Each expert produces a class probabilitypip\_\{i\}; the per\-expert confidencecic\_\{i\}is taken as the maximum softmax output of the classification head for expertii\.

Once these per\-expert outputs are available, task\-specific quality metrics are computed across them to inform the selection: morphological plausibility such as area, circularity and compactness, and inter\-model agreement measured by pairwise IoU and HD95, for segmentation; and prediction uncertainty such as entropy and margin, and class consensus, for classification\. The LLM router then selects the best mask and classification result by reasoning over these quality metrics, per\-expert confidence estimates and case\-level metadata, rather than by simple confidence maximization or majority voting\. Depending on the configured ensemble size, the router selects either the single best expert or a subset of experts for weighted ensemble fusion\. This routing design allows the system to adapt its selection to the acquisition conditions of each case, rather than relying on a fixed model ranking\.

To provide a classification signal independent of the image\-based experts, the radiomics branch processes the selected mask\. The mask is first refined via connected\-component analysis to remove isolated noisy regions when multiple disconnected components are present\. Two\-dimensional PyRadiomics descriptors\[[56](https://arxiv.org/html/2608.12590#bib.bib56)\], including shape, intensity and texture features, are then extracted from the refined mask\-image pair, yielding the radiomic descriptor vector\. These descriptors are passed to AutoGluon\-tabular classifiers\[[84](https://arxiv.org/html/2608.12590#bib.bib84)\], which ensemble multiple tabular models under automated hyperparameter optimization\. The resulting tabular class prediction is stored alongside the router’s selection in the case\-level evidence store, providing a complementary interpretive signal for clinician review\. After selection, SHAP analysis\[[72](https://arxiv.org/html/2608.12590#bib.bib72)\]is applied to the tabular classifier to estimate global and local feature contributions, while Grad\-CAM\[[85](https://arxiv.org/html/2608.12590#bib.bib85)\]is applied to the selected segmentation model to visualize the image regions driving its mask prediction\. These post hoc explanations are stored in the evidence store for clinician inspection\. Importantly, the same segmentation, radiomics, tabular classification and explanation workflow is reused across benign\-malignant classification and additional malignant\-lesion stratification tasks, including LNM prediction and FTC/PTC subtype classification; only task\-specific classifier fine\-tuning, the task description supplied to the router, and, where applicable, external model assembly are changed\. This reuse allows the agent workflow to be redirected to new clinical questions without rebuilding the diagnostic pipeline\.

### 4\.4ThyClinScore

ThyClinScore evaluates thyroid ultrasound reports as a structured clinical semantic agreement task\. It measures both report completeness and semantic consistency with the reference report\. Semantic consistency is assessed at two levels: gland\-level agreement for thyroid measurements, parenchymal morphology and gland\-level vascularity; and lesion\-level agreement for lesion detection, lesion size, lesion descriptors and lesion\-level vascularity\. Size and vascularity can therefore be scored at either level when the corresponding fields are available, whereas morphology mainly captures concept\-level agreement in gland and parenchymal descriptions\.

For each case, the reference reportRgtR^\{\\mathrm\{gt\}\}and generated reportRpredR^\{\\mathrm\{pred\}\}were converted into a shared schema containing thyroid measurements, parenchymal findings, lesion attributes, lymph\-node findings and diagnostic impressions\. Measurements were standardized in millimetres, categorical fields were restricted to predefined thyroid ultrasound descriptors, and absent information was represented as null\. The schema retained clinically relevant information, including lesion location, size, composition, echogenicity, margin, shape, echogenic foci, vascularity and TI\-RADS category\. These fields supported subsequent lesion\-level matching and semantic scoring\.

At the lesion level, reference and predicted lesions were first matched before lesion detection and attribute agreement were evaluated\. LetG=\{gi\}i=1mG=\\\{g\_\{i\}\\\}\_\{i=1\}^\{m\}denote reference lesions andP=\{pj\}j=1nP=\\\{p\_\{j\}\\\}\_\{j=1\}^\{n\}denote predicted lesions\. For each candidate pair, we computed

Si​j=Mi​jloc​\(αs​Si​jsize\+βf​Si​jfeat\),Si​jsize=min⁡\(di,dj\)max⁡\(di,dj\),S\_\{ij\}=M^\{\\mathrm\{loc\}\}\_\{ij\}\\left\(\\alpha\_\{s\}S^\{\\mathrm\{size\}\}\_\{ij\}\+\\beta\_\{f\}S^\{\\mathrm\{feat\}\}\_\{ij\}\\right\),\\qquad S^\{\\mathrm\{size\}\}\_\{ij\}=\\frac\{\\min\(d\_\{i\},d\_\{j\}\)\}\{\\max\(d\_\{i\},d\_\{j\}\)\},wheredid\_\{i\}anddjd\_\{j\}are maximum lesion diameters,Mi​jloc∈\{0,1\}M^\{\\mathrm\{loc\}\}\_\{ij\}\\in\\\{0,1\\\}is a hard anatomical\-location gate, andSi​jfeat∈\{0,0\.5,1\}S^\{\\mathrm\{feat\}\}\_\{ij\}\\in\\\{0,0\.5,1\\\}scores lesion\-composition agreement\. Explicit left\-right mismatches were assignedMi​jloc=0M^\{\\mathrm\{loc\}\}\_\{ij\}=0\. We solved the bipartite assignment usingci​j=1−Si​jc\_\{ij\}=1\-S\_\{ij\}and retained pairs above a predefined threshold\. Matched pairs were treated as true positives, unmatched reference lesions as false negatives and unmatched predicted lesions as false positives, from which lesion\-level precision, recall, F1 and false discovery rate were calculated\.

Numerical measurements were scored using mean relative error \(MRE\), applied to thyroid lobe and isthmus measurements at the gland level and lesion dimensions at the lesion level\. For paired dimensionsKK,

MREK=1\|K\|​∑k∈K\|pk−gk\|max⁡\(\|gk\|,ϵ\),ssize​\(MREK,τ\)=2−\(MREK/τ\)2,\\mathrm\{MRE\}\_\{K\}=\\frac\{1\}\{\|K\|\}\\sum\_\{k\\in K\}\\frac\{\|p\_\{k\}\-g\_\{k\}\|\}\{\\max\(\|g\_\{k\}\|,\\epsilon\)\},\\qquad s\_\{\\mathrm\{size\}\}\(\\mathrm\{MRE\}\_\{K\};\\tau\)=2^\{\-\(\\mathrm\{MRE\}\_\{K\}/\\tau\)^\{2\}\},whereϵ\\epsilonprovides numerical stability andτ\\taucontrols tolerance to size error\. Vascularity was evaluated at either level when available\. It was discretized into four grades, from absent flow to markedly increased flow, and scored as

svasc=max⁡\(0,1−\|vgt−vpred\|3\)\.s\_\{\\mathrm\{vasc\}\}=\\max\\left\(0,1\-\\frac\{\|v^\{\\mathrm\{gt\}\}\-v^\{\\mathrm\{pred\}\}\|\}\{3\}\\right\)\.Morphological consistency was computed at the concept level for gland and parenchymal descriptions\. A thyroid\-specific lexicon mapped report phrases to concepts covering echogenicity, texture, margin, shape, calcification, posterior acoustic features and composition\. For concept setsCgtC^\{\\mathrm\{gt\}\}andCpredC^\{\\mathrm\{pred\}\}, morphology agreement was scored by concept F1\. For matched lesions, categorical descriptors were scored by mean accuracy over reference\-present fields, including composition, echogenicity, margin, shape and echogenic foci\.

Gland\-level and lesion\-level assessments were combined into a clinical consistency scoreCC, which aggregates thyroid\-size agreement, lesion\-size agreement, vascularity agreement, lesion\-detection F1, lesion\-feature accuracy and morphology agreement\. Components not applicable to either report were excluded from the denominator, whereas reference\-present fields missing from the generated report contributed zero:

C=∑q∈𝒬wq​sq∑q∈𝒬wq,C=\\frac\{\\sum\_\{q\\in\\mathcal\{Q\}\}w\_\{q\}s\_\{q\}\}\{\\sum\_\{q\\in\\mathcal\{Q\}\}w\_\{q\}\},wheresqs\_\{q\}denotes a valid component score andwqw\_\{q\}its predefined weight\. Report completenessBBwas defined as the weighted fraction of required information present in the generated structured report:

B=∑r∈ℛur​br∑r∈ℛur\.B=\\frac\{\\sum\_\{r\\in\\mathcal\{R\}\}u\_\{r\}b\_\{r\}\}\{\\sum\_\{r\\in\\mathcal\{R\}\}u\_\{r\}\}\.The completeness groups were thyroid measurements, parenchyma, lesions, impression and lymph\-node description;brb\_\{r\}indicates presence of grouprr, anduru\_\{r\}denotes its predefined weight\. The final ThyClinScore was

ThyClinScore=\[λ​B\+\(1−λ\)​C\]​\[η\+\(1−η\)​FL\],\\mathrm\{ThyClinScore\}=\\left\[\\lambda B\+\(1\-\\lambda\)C\\right\]\\left\[\\eta\+\(1\-\\eta\)F\_\{L\}\\right\],for cases with reference lesions, whereλ\\lambdabalances completeness and clinical consistency, andη\\etacontrols the minimum lesion\-detection gate\. For reports without reference lesions, the gate was omitted\. This design rewards complete and semantically consistent reports while penalizing missed or hallucinated lesions\.

### 4\.5Report generation

For the reporting branch, thyroid ultrasound reporting was formulated as a case\-level evidence\-to\-report task rather than as single\-image captioning\. Given an image setℐ=\{Ij\}j=1N\\mathcal\{I\}=\\\{I\_\{j\}\\\}\_\{j=1\}^\{N\}, which could include video frames, static images, multiple anatomical views, greyscale ultrasound and colour Doppler images, the workflow generated a structured reportRRand an accompanying evidence trace\. The report covered thyroid gland morphology, nodule\-level findings, lymph\-node findings when present and the diagnostic impression\. The trace recorded the intermediate observations supporting each report component\.

Input preparation converted heterogeneous case files into image priors for subsequent planning\. Images were standardized and processed by auxiliary modules for region\-of\-interest cropping, anatomical\-context parsing, colour Doppler identification, nodule\-presence triage and pixel\-spacing estimation when required\. The resulting priors encoded region, view, modality, Doppler status and nodule likelihood\. They were used to guide downstream analysis, rather than inserted directly into report text\.

The reporting workflow was packaged as a single reusable skill\. The skill defined the reporting objective, evidence schema, staged diagnostic process and constraints for evidence\-grounded writing\. External agents invoked this skill through a Workflow Model Context Protocol \(MCP\) server, which exposed high\-level case operations for input preparation, case initiation, plan review and approval, and report and evidence retrieval\. Low\-level operations, including segmentation, classification, measurement and captioning, remained internal to the workflow\. This separated a stable public interface from the image\-processing and model\-execution details needed to construct reliable evidence\. It also made the resulting report traceable for clinician review\.

After input preparation, ThyroidXAgent used a planner\-executor design\[[49](https://arxiv.org/html/2608.12590#bib.bib49)\]\. The planner received the skill instructions, tool contract and image priors, and generated a case\-specific staged task graph covering gland assessment, nodule analysis, lymph\-node assessment, evidence fusion and report generation\. This graph provided global diagnostic constraints, while allowing case\-specific execution\. In the deployable workflow, the plan could be inspected and approved before model execution\. The executor then followed the approved graph and used ReAct\-style local decision\-making\[[52](https://arxiv.org/html/2608.12590#bib.bib52)\]to select images and internal tools according to intermediate observations\. For example, cases without nodule priors could bypass nodule\-feature classification, whereas lateral\-neck images could trigger lymph\-node screening before report assembly\.

Selected model operations were executed by the internal diagnostic runtime\. In deployment, the Workflow MCP process remained lightweight and delegated GPU inference to a separate tool service\. The runtime could call preprocessing modules, gland captioning, spacing prediction, thyroid and nodule measurement, nodule segmentation, nodule\-feature classification, malignancy classification, cervical lymph\-node screening and nodule\-level fusion\. Their outputs were normalized into a shared evidence object before language generation\. Gland\-level evidence included thyroid lobe and isthmus measurements, parenchymal morphology and vascularity\. Nodule\-level evidence included anatomical location, size, composition, echogenicity, margin, shape, echogenic foci, vascularity, segmentation\-derived measurements and risk\-related predictions\. Lymph\-node evidence summarized cervical\-region screening results\. When multiple views described the same lesion, the fusion stage consolidated compatible findings into case\-level nodule entries and retained warnings for incomplete or conflicting outputs\.

Report text was generated by controlled data\-to\-text assembly\[[57](https://arxiv.org/html/2608.12590#bib.bib57)\]\. To reduce factual hallucination risk\[[58](https://arxiv.org/html/2608.12590#bib.bib58)\], ThyroidXAgent used training\-free BM25 template retrieval, slot filling and clause assembly rather than unconstrained free\-text decoding\. Corresponding template libraries were constructed from the SMU\-HMC and KMVE datasets\. The SMU\-HMC was derived from 7,147 quality\-controlled reports after excluding all reports from the 400 patients reserved for internal testing\. The KMVE was constructed from all 1,719 reports in its training partition, with the validation and test partitions excluded\. During template construction, whole reports were decomposed into five clinically defined clause categories: thyroid measurement, gland morphology, nodule findings, lymph\-node findings and ultrasound impression\. During inference, the evidence object assembled through tool execution was partitioned into corresponding evidence blocks and converted into category\-specific BM25 queries\. Retrieved templates served as linguistic frames; slot filling inserted case\-specific measurements, descriptors, TI\-RADS\-relevant information, risk categories and other tool\-derived evidence, and clause assembly combined the completed findings and impression clauses into the final report\.

Notably, KMVE contains only the ultrasound findings section and does not provide original measurement values\. The template library constructed from KMVE therefore retained the original structure and linguistic style of the KMVE findings and was used in accordance with its original findings\-only evaluation protocol\. Because tool execution was decoupled from report realization, the retrieval libraries could be exchanged in a plug\-and\-play manner to align generated reports with centre\- or corpus\-specific writing conventions, without retraining the upstream image\-analysis tools or modifying the underlying evidence schema\. The same procedure can be used to construct updated retrieval libraries from new report sources \(Supplementary Fig\.[S4](https://arxiv.org/html/2608.12590#Ax1.F4)c\)\.

The final artefact consisted of the report text and its supporting evidence trace\. Clinicians or external agents could inspect the diagnostic plan, verify the evidence used for each statement, review warnings and edit the generated report\. Corrected reports could then be returned to the case\-level evidence store for subsequent review and model updating\. This design kept report generation constrained by structured clinical evidence while preserving a reviewable path from image inputs to final report statements\.

### 4\.6Reader studies and statistical analysis

For segmentation review, clinicians corrected AI\-generated masks, and correction time and paired Dice scores were compared with manual segmentation\. For report writing, two physicians evaluated 145 thyroid ultrasound videos with annotation\-derived diagnostic labels in a cross\-over design\. Each case was interpreted under both manual and AI\-assisted conditions, but by different physicians, to reduce memory and recall bias\. In the manual condition, physicians reviewed the ultrasound video and completed a blank structured\-report template\. In the AI\-assisted condition, they reviewed an AI\-generated draft together with the supporting tool\-derived evidence and revised the report as required\. We recorded reporting time, video\-review time and the final submitted report for each assessment\. Reporting efficiency was assessed at both the case and physician levels, with case\-level time saving defined as the manual minus AI\-assisted reporting time\.

For annotation\-based diagnostic\-direction analysis, frame\-level benign and malignant bounding\-box annotations were aggregated into case\-level labels\. Cases containing any malignant annotation were classified as malignant or suspicious, whereas cases containing only benign annotations were classified as benign\. This yielded 145 labelled cases, comprising 72 benign and 73 malignant or suspicious cases\. A prespecified rule\-based procedure assigned each submitted report to a benign or low\-risk, malignant or suspicious, or ambiguous direction using TI\-RADS categories and diagnostic terms\. Directional consistency was defined by agreement with the annotation\-derived case label and was summarized overall and within the benign and malignant or suspicious strata\. For statistical analysis, performance was summarized using the primary metric for each task: Dice and HD95 for segmentation, AUROC and AUPRC for classification, conventional natural\-language generation metrics for report wording, and ThyClinScore and its submetrics for report semantics\. Confidence intervals in the supplementary tables were estimated by nonparametric bootstrap resampling of the test set\. Correlations between report metrics and the location\-aware LLM judge were evaluated using two\-sided Pearson correlation tests\.

## 5Data availability

The public datasets used in this study are available from their original sources, as cited in the Methods and Supplementary Table[S1](https://arxiv.org/html/2608.12590#Ax1.T1)\. The publicly released data associated with this study are available through ThyroidOpenDB at[https://huggingface\.co/datasets/MedXAgent/ThyroidOpenDB](https://huggingface.co/datasets/MedXAgent/ThyroidOpenDB)\. The study was approved by the Medical Ethics Committee of Zhujiang Hospital, Southern Medical University \(approval no\. 2026\-KY\-081\-01\)\.

The NHCMISD dataset was collected from real\-world clinical ultrasound cases provided by the National Health Commission Medical Imaging Standard Database, for which the co\-author, Prof\. Dexing Kong, has authorized access\. All data usage complied with relevant institutional and regulatory requirements\.

## 6Code availability

The source code for ThyroidXAgent is publicly available on GitHub at[https://github\.com/MedXAgent/ThyroidXAgent](https://github.com/MedXAgent/ThyroidXAgent)\. The website will be made available after acceptance\. The trained model weights are available on Hugging Face at[https://huggingface\.co/MedXAgent/ThyroidXAgent](https://huggingface.co/MedXAgent/ThyroidXAgent)\. The repositories contain the resources required to reproduce the reported computational analyses\.

## 7Acknowledgements

This work was supported in part by the National Natural Science Foundation of China \(Grant No\. 62322608\), the Zhejiang Provincial Natural Science Foundation of China \(Grant No\. LQN26F020029\), and the Natural Science Foundation of Guangdong Province \(Grant No\. 2024A1515010255\)\.

## 8Author contributions

- •Conceptualization:H\.G\. and G\.L\.
- •System development:S\.C\., B\.W\. and X\.X\.
- •Computational evaluation:S\.C\., B\.W\. and S\.W\.
- •Visualization:S\.C\., B\.W\. and H\.G\.
- •Reader study:H\.W\., Q\.L\. and F\.C\.
- •Data curation and organization:H\.G\., S\.C\., B\.W\., Y\.W\., G\.Y\., H\.W\., M\.M\., D\.K\., Q\.L\., W\.L\. and F\.C\.
- •Writing–original draft:H\.G\., Y\.W\., S\.C\. and B\.W\.
- •Writing–review and editing:All authors\.
- •Supervision:Q\.L\., W\.L\., F\.C\. and G\.L\.

## 9Confilts

The authors has no confilt of interests\.

## Supplementary information

#### Supplementary figures

#### Supplementary tables

![Refer to caption](https://arxiv.org/html/2608.12590v1/ThyroidXAgent_for_seg_and_cls.png)Figure S1:ThyroidXAgent architecture for thyroid nodule segmentation and classification\. DINOv3\-based experts generate candidate segmentation masks and malignancy probabilities\. The radiomics branch extracts PyRadiomics features from the selected lesion mask, applies an AutoGluon classifier and computes SHAP attributions\. The LLM router integrates the candidate outputs with image metadata, including resolution, device and data source, to select the final prediction and its supporting evidence\.![Refer to caption](https://arxiv.org/html/2608.12590v1/SegCls_statistics.png)Figure S2:Dataset composition and lesion distributions across thyroid ultrasound cohorts\. Top, grouped bar chart comparing the numbers of benign and malignant images in TN3K, TN5K, ThyroidXL, DDTI and ZJH\-8K on a logarithmic y axis, highlighting marked variation in cohort size and class balance across cohorts\. Bottom, for each of the five datasets, two\-dimensional kernel density estimates of normalized lesion\-mask centroid positions \(left\) and relative lesion size distributions, defined as mask area divided by image area \(right\)\. All size distributions share a common x\-axis range, and all spatial maps use a common density scale, enabling direct cross\-dataset comparison of where lesions appear within the ultrasound frame and how large they tend to be\.![Refer to caption](https://arxiv.org/html/2608.12590v1/BM_cases.png)Figure S3:Case\-level explanations for thyroid nodule classification\. Left, SHAP values for the most influential radiomic features; red and blue indicate contributions towards malignant and benign predictions, respectively\. Right, ultrasound images with segmentation contours and Grad\-CAM maps from the selected segmentation model\. Representative benign and malignant cases with accurate and inaccurate segmentation are shown\.![Refer to caption](https://arxiv.org/html/2608.12590v1/ReportGenAgent1.png)Figure S4:Multimodal input processing and case interaction in ThyroidXAgent\.a, Agent Chat interface for initiating and monitoring case\-level thyroid ultrasound analysis, with the case context and workflow activity displayed alongside the conversation\.b, Visual perception of case\-level, multiview thyroid ultrasound inputs, including image selection, anatomical\-region recognition and organization of image\-context priors\.c, Video\-input processing, in which an ultrasound video is sampled into representative frames and organized with anatomical\-region predictions and image\-context priors for downstream analysis\.![Refer to caption](https://arxiv.org/html/2608.12590v1/ReportGenAgent2.png)Figure S5:Planning, tool execution and clinician review in the report\-generation workflow\.a, Case\-specific planning, in which preprocessing outputs, workflow instructions and diagnostic objectives are converted into a staged plan for review and approval\.b, Execution after plan approval, showing internal tool calls, intermediate outputs, evidence collection and progression from the approved plan to a reviewable report\.c, Report Review Dashboard displaying the generated report alongside source images, structured evidence and tool traces for clinician inspection and editing\.![Refer to caption](https://arxiv.org/html/2608.12590v1/ReportGenAgent3.png)Figure S6:Template\-bank construction for centre\-specific report generation\. The interface enables users to select an existing centre\-specific report\-generation profile or import a new report corpus\. Generic template\-construction scripts derive dataset\-specific template banks to support adaptation to additional reporting data\.![Refer to caption](https://arxiv.org/html/2608.12590v1/RG_CaseReview.png)Figure S7:Qualitative comparison of thyroid ultrasound report generation\. Representative benign and malignant cases compare reports generated by Qwen 3\.5, GPT\-5 and ThyroidXAgent with the reference reports\. Text spans are annotated as clinically correct, partially correct or incorrect\. The malignant example corresponds to Fig\.[6](https://arxiv.org/html/2608.12590#S2.F6)b; the benign example provides an additional complementary case\.Table S1:Composition and data splits of the multicentre thyroid ultrasound benchmark\. Numbers and percentages show the images contributed by each dataset to the full benchmark \(n=38,864\) and to the training \(n=18,277\), validation \(n=850\) and test \(n=19,737\) cohorts\. DDTI, RJH\-7K and ZJH\-8K are independent external test cohorts\. DDTI comprises 637 images, of which 349 carry benign\-malignant classification labels\. ZJH\-8K serves as an external test set for segmentation and classification; its 4,756 malignant images additionally serve as the primary training set for malignant\-lesion stratification \(Supplementary Table[S5](https://arxiv.org/html/2608.12590#Ax1.T5)\)\.DatasetTaskCenterUltrasound deviceTotalTrainValidTestTN3K\[[11](https://arxiv.org/html/2608.12590#bib.bib11)\]Segmentation,ClassificationZhujiang Hospital,Southern MedicalUniversity, Guangzhou,ChinaGE Logiq E9,ARIETTA 850,RESONA 70B5,347 \(13\.76%\)4,633 \(25\.35%\)100 \(11\.76%\)614 \(3\.11%\)TN5K\[[55](https://arxiv.org/html/2608.12590#bib.bib55)\]Segmentation,ClassificationCancer Hospital,Chinese Academy ofMedical Sciences,Beijing, ChinaGE Logiq E9,GE S7 \(5\-12 MHzor 8\-15 MHz\)5,000 \(12\.87%\)3,500 \(19\.15%\)500 \(58\.82%\)1,000 \(5\.07%\)ThyroidXL\[[59](https://arxiv.org/html/2608.12590#bib.bib59)\]Segmentation,ClassificationVietnam NationalHospital ofEndocrinology,Hanoi, VietnamHitachi Aloka Arietta V7011,631 \(29\.93%\)9,441 \(51\.66%\)100 \(11\.76%\)2,090 \(10\.59%\)PKTN\[[13](https://arxiv.org/html/2608.12590#bib.bib13)\]SegmentationPeking UniversityFirst Hospital,Beijing, China–1,003 \(2\.58%\)703 \(3\.85%\)150 \(17\.65%\)150 \(0\.76%\)DDTI\[[60](https://arxiv.org/html/2608.12590#bib.bib60)\]Segmentation,ClassificationIDIME,Bogotá, ColombiaTOSHIBA Nemio 30,TOSHIBA Nemio MX\(12 MHz probe\)637 \(1\.64%\)––637 \(3\.23%\) seg349 clsRJH\-7K\[[61](https://arxiv.org/html/2608.12590#bib.bib61)\]SegmentationRuijin Hospital,Shanghai Jiao TongUniversity School ofMedicine, Shanghai,ChinaDifferent machines,not specified7,288 \(18\.75%\)––7,288 \(36\.93%\)ZJH\-8KSegmentation,ClassificationZhujiang Hospital,Southern MedicalUniversity, Guangzhou,China–7,958 \(20\.48%\)––7,958 \(40\.32%\)

Table S2:Characteristics of public and institutional thyroid ultrasound datasets\. Dataset size, data split, file format, task, geographical source and ultrasound scanner information are summarized for the included resources\.DatasetDataset SizeDevelopment/ValidationEvaluationFile FormatTaskLocationUltrasonic Imaging DeviceTGVideo\[[86](https://arxiv.org/html/2608.12590#bib.bib86)\]15,186 \(16 cases\)15,186N/ASegmentationGermanyGE Logiq E9DDTI\[[60](https://arxiv.org/html/2608.12590#bib.bib60)\]637N/A637Image: PNGMask: PNGLabel: CSVSegmentationClassificationColombiaTOSHIBA Nemio 30TOSHIBA Nemio MXTN3K\[[11](https://arxiv.org/html/2608.12590#bib.bib11),[14](https://arxiv.org/html/2608.12590#bib.bib14),[12](https://arxiv.org/html/2608.12590#bib.bib12)\]5,3474,733614Image: JPGMask: JPGLabel: CSVSegmentationClassificationGuangzhou, ChinaGE Logiq E9ARIETTA 850RESONA 70BTN5K\[[55](https://arxiv.org/html/2608.12590#bib.bib55)\]5,0004,0001,000Beijing, ChinaThyUS2Path\[[87](https://arxiv.org/html/2608.12590#bib.bib87)\]8,5085,4573,051ClassificationZhejiang, ChinaEsaote MyLab \(Portable\)Cine\-clip\[[88](https://arxiv.org/html/2608.12590#bib.bib88)\]17,412 frames192 casesavg\. 90 frames/caseN/AN/ASegmentationClassificationCalifornia, USAN/AAHU\[[89](https://arxiv.org/html/2608.12590#bib.bib89)\]1,833 cases125,896 imagesN/AN/AImage: JPGLabel: Folder\-levelClassificationChina \(web scraping\)HeterogeneousThyroidXL\[[59](https://arxiv.org/html/2608.12590#bib.bib59)\]11,6319,5412,090Image: PNGMask: PNGLabel: TXTSegmentationClassificationDetectionVietnamHitachi Aloka Arietta V70PKTN\[[13](https://arxiv.org/html/2608.12590#bib.bib13)\]1,003N/AN/ASegmentationBeijing, ChinaN/AKMVE\[[18](https://arxiv.org/html/2608.12590#bib.bib18)\]2,457 cases4,914 image assignments1,719 training cases246 validation cases492 test casesReport GenerationBeijing, ChinaN/ASMU\-HMC23,955 reports248,194 images23,555 reports242,210 imagesReport GenerationComponent DevelopmentGuangzhou, ChinaHeterogeneousZJH\-TSN/AReport GenerationExternal ValidationGuangzhou, ChinaHeterogeneousTNVideo148 casesN/A145 labelled casesVideo: AVISegmentationReader StudyGuangzhou, ChinaN/A

Table S3:Cross\-dataset generalization for thyroid nodule segmentation\. Dice coefficient and 95th\-percentile Hausdorff distance \(HD95\) are reported with 95% confidence intervals\.ModelTN3KThyroidXLPKTNTN5KDDTIZJH\-8KRJH\-7KDice \(%\)↑\\uparrowTransUnet\[[90](https://arxiv.org/html/2608.12590#bib.bib90)\]81\.84±1\.6281\.84\\pm 1\.6285\.75±0\.5785\.75\\pm 0\.5776\.89±3\.5676\.89\\pm 3\.5678\.54±1\.5178\.54\\pm 1\.5176\.58±1\.6276\.58\\pm 1\.6280\.72±0\.9780\.72\\pm 0\.9784\.83±0\.3784\.83\\pm 0\.37MedSegX\[[91](https://arxiv.org/html/2608.12590#bib.bib91)\]83\.93±0\.7983\.93\\pm 0\.7979\.98±0\.3679\.98\\pm 0\.3680\.63±0\.4280\.63\\pm 0\.4283\.10±0\.4883\.10\\pm 0\.4875\.12±1\.6875\.12\\pm 1\.6884\.06±0\.3984\.06\\pm 0\.3985\.40±0\.1885\.40\\pm 0\.18MedSAM2\[[64](https://arxiv.org/html/2608.12590#bib.bib64)\]84\.47±1\.0284\.47\\pm 1\.0286\.94±0\.3686\.94\\pm 0\.3683\.46±\\pm2\.6083\.03±1\.2983\.03\\pm 1\.2984\.72±1\.2684\.72\\pm 1\.2686\.29±0\.7386\.29\\pm 0\.7390\.72±0\.2190\.72\\pm 0\.21UltraFedFM\[[65](https://arxiv.org/html/2608.12590#bib.bib65)\]81\.18±1\.4681\.18\\pm 1\.4684\.70±0\.5384\.70\\pm 0\.5375\.31±1\.1275\.31\\pm 1\.1277\.13±1\.3877\.13\\pm 1\.3875\.57±1\.6775\.57\\pm 1\.6780\.64±0\.8480\.64\\pm 0\.8483\.10±0\.3383\.10\\pm 0\.33ThyroidXAgent85\.28±\\pm1\.2887\.58±\\pm0\.4482\.99±2\.1082\.99\\pm 2\.1083\.26±\\pm1\.3485\.62±\\pm1\.0794\.30±\\pm0\.3891\.46±\\pm0\.14HD95 \(mm\)↓\\downarrowTransUnet\[[90](https://arxiv.org/html/2608.12590#bib.bib90)\]27\.27±5\.5227\.27\\pm 5\.5222\.42±1\.3422\.42\\pm 1\.3426\.88±9\.6626\.88\\pm 9\.6622\.32±3\.4322\.32\\pm 3\.4317\.12±1\.5517\.12\\pm 1\.5518\.37±0\.7518\.37\\pm 0\.7518\.81±0\.7418\.81\\pm 0\.74MedSegX\[[91](https://arxiv.org/html/2608.12590#bib.bib91)\]10\.95±0\.6410\.95\\pm 0\.6411\.07±0\.3211\.07\\pm 0\.3210\.83±0\.7010\.83\\pm 0\.7011\.76±0\.7611\.76\\pm 0\.7618\.39±1\.6518\.39\\pm 1\.6510\.96±0\.3510\.96\\pm 0\.359\.37±0\.189\.37\\pm 0\.18MedSAM2\[[64](https://arxiv.org/html/2608.12590#bib.bib64)\]11\.51±1\.5311\.51\\pm 1\.535\.46±0\.445\.46\\pm 0\.4410\.56±3\.6410\.56\\pm 3\.6410\.94±1\.1210\.94\\pm 1\.1210\.06±1\.2110\.06\\pm 1\.216\.79±0\.576\.79\\pm 0\.572\.92±0\.172\.92\\pm 0\.17UltraFedFM\[[65](https://arxiv.org/html/2608.12590#bib.bib65)\]14\.98±2\.1014\.98\\pm 2\.108\.10±0\.588\.10\\pm 0\.5816\.08±1\.6716\.08\\pm 1\.6714\.96±1\.6514\.96\\pm 1\.6518\.12±1\.4718\.12\\pm 1\.478\.69±0\.808\.69\\pm 0\.809\.06±0\.389\.06\\pm 0\.38ThyroidXAgent10\.31±\\pm1\.705\.43±\\pm0\.539\.01±\\pm3\.5810\.12±\\pm1\.239\.24±\\pm1\.072\.25±\\pm0\.391\.92±\\pm0\.08Table S4:Cross\-dataset generalization for benign\-malignant thyroid nodule classification\. Area under the receiver operating characteristic curve \(AUROC\) and area under the precision\-recall curve \(AUPRC\) are reported with 95% confidence intervals\.MethodTN3KThyroidXLTN5KDDTIZJH\-8KAUROC↑\\uparrowResNet\-50\[[92](https://arxiv.org/html/2608.12590#bib.bib92)\]0\.7674±0\.03940\.7674\\pm 0\.03940\.9044±0\.01180\.9044\\pm 0\.01180\.9322±0\.01680\.9322\\pm 0\.01680\.6704±0\.08420\.6704\\pm 0\.08420\.6704±0\.08420\.6704\\pm 0\.0842RepViT\[[66](https://arxiv.org/html/2608.12590#bib.bib66)\]0\.5556±0\.04630\.5556\\pm 0\.04630\.7774±0\.01880\.7774\\pm 0\.01880\.6603±0\.03750\.6603\\pm 0\.03750\.6162±0\.08040\.6162\\pm 0\.08040\.8538±0\.01850\.8538\\pm 0\.0185LSNet\[[93](https://arxiv.org/html/2608.12590#bib.bib93)\]0\.8095±0\.03330\.8095\\pm 0\.03330\.9178±0\.01140\.9178\\pm 0\.01140\.9091±0\.02010\.9091\\pm 0\.02010\.7581±0\.06580\.7581\\pm 0\.06580\.8631±0\.02010\.8631\\pm 0\.0201UltraFedFM\[[65](https://arxiv.org/html/2608.12590#bib.bib65)\]0\.8461±0\.06970\.8461\\pm 0\.06970\.9239±0\.01040\.9239\\pm 0\.01040\.9298±0\.01750\.9298\\pm 0\.01750\.7518±0\.17120\.7518\\pm 0\.17120\.9115±0\.01400\.9115\\pm 0\.0140MedGemma\[[71](https://arxiv.org/html/2608.12590#bib.bib71)\]0\.8492±0\.03050\.8492\\pm 0\.03050\.9371±0\.00950\.9371\\pm 0\.00950\.9442±0\.01560\.9442\\pm 0\.01560\.8255±\\pm0\.06500\.8976±0\.01660\.8976\\pm 0\.0166Qwen3\-VL\-8B\-Instruct\[[94](https://arxiv.org/html/2608.12590#bib.bib94)\]0\.8237±0\.03280\.8237\\pm 0\.03280\.9050±0\.01150\.9050\\pm 0\.01150\.9214±0\.01870\.9214\\pm 0\.01870\.7361±0\.06920\.7361\\pm 0\.06920\.8659±0\.01890\.8659\\pm 0\.0189GPT\-5\[[68](https://arxiv.org/html/2608.12590#bib.bib68)\]0\.6924±0\.04210\.6924\\pm 0\.04210\.7059±0\.04690\.7059\\pm 0\.04690\.7737±0\.09960\.7737\\pm 0\.09960\.6346±0\.09140\.6346\\pm 0\.09140\.6109±0\.05150\.6109\\pm 0\.0515Gemini\-2\.5\-Pro\[[69](https://arxiv.org/html/2608.12590#bib.bib69)\]0\.6587±0\.04550\.6587\\pm 0\.04550\.6246±0\.06400\.6246\\pm 0\.06400\.6873±0\.06910\.6873\\pm 0\.06910\.6156±0\.13080\.6156\\pm 0\.13080\.6493±0\.05160\.6493\\pm 0\.0516ThyroidXAgent0\.8692±\\pm0\.03490\.9676±\\pm0\.00660\.9472±\\pm0\.01520\.7991±0\.07410\.7991\\pm 0\.07410\.9175±\\pm0\.0167AUPRC↑\\uparrowResNet\-50\[[92](https://arxiv.org/html/2608.12590#bib.bib92)\]0\.6882±0\.06320\.6882\\pm 0\.06320\.8882±0\.01740\.8882\\pm 0\.01740\.9674±0\.02680\.9674\\pm 0\.02680\.3755±0\.11760\.3755\\pm 0\.11760\.2755±0\.11670\.2755\\pm 0\.1167RepViT\[[66](https://arxiv.org/html/2608.12590#bib.bib66)\]0\.4275±0\.05280\.4275\\pm 0\.05280\.7161±0\.02760\.7161\\pm 0\.02760\.8403±0\.02160\.8403\\pm 0\.02160\.3924±0\.09330\.3924\\pm 0\.09330\.9486±0\.00780\.9486\\pm 0\.0078LSNet\[[93](https://arxiv.org/html/2608.12590#bib.bib93)\]0\.7581±0\.04520\.7581\\pm 0\.04520\.9040±0\.01420\.9040\\pm 0\.01420\.9551±0\.01340\.9551\\pm 0\.01340\.4180±0\.14100\.4180\\pm 0\.14100\.9449±0\.01130\.9449\\pm 0\.0113UltraFedFM\[[65](https://arxiv.org/html/2608.12590#bib.bib65)\]0\.8531±0\.02840\.8531\\pm 0\.02840\.9354±0\.01140\.9354\\pm 0\.01140\.8422±0\.04210\.8422\\pm 0\.04210\.4487±0\.14520\.4487\\pm 0\.14520\.9669±0\.00840\.9669\\pm 0\.0084MedGemma\[[71](https://arxiv.org/html/2608.12590#bib.bib71)\]0\.8047±0\.04300\.8047\\pm 0\.04300\.9201±0\.01390\.9201\\pm 0\.01390\.9747±0\.00840\.9747\\pm 0\.00840\.5537±0\.16630\.5537\\pm 0\.16630\.9589±0\.00960\.9589\\pm 0\.0096Qwen3\-VL\-8B\-Instruct\[[94](https://arxiv.org/html/2608.12590#bib.bib94)\]0\.7617±0\.05110\.7617\\pm 0\.05110\.8787±0\.03790\.8787\\pm 0\.03790\.9636±0\.01060\.9636\\pm 0\.01060\.4112±0\.14150\.4112\\pm 0\.14150\.9498±0\.00960\.9498\\pm 0\.0096GPT\-5\[[68](https://arxiv.org/html/2608.12590#bib.bib68)\]0\.6627±0\.06330\.6627\\pm 0\.06330\.6237±0\.06660\.6237\\pm 0\.06660\.8920±0\.03160\.8920\\pm 0\.03160\.3578±0\.10890\.3578\\pm 0\.10890\.8311±0\.03770\.8311\\pm 0\.0377Gemini\-2\.5\-Pro\[[69](https://arxiv.org/html/2608.12590#bib.bib69)\]0\.6205±0\.05870\.6205\\pm 0\.05870\.4914±0\.08410\.4914\\pm 0\.08410\.8462±0\.04460\.8462\\pm 0\.04460\.3924±0\.15270\.3924\\pm 0\.15270\.8403±0\.03620\.8403\\pm 0\.0362ThyroidXAgent0\.8545±\\pm0\.06000\.9653±\\pm0\.00780\.9752±\\pm0\.00890\.5863±\\pm0\.13800\.9711±\\pm0\.0006Table S5:Performance on malignant thyroid lesion stratification\. AUROC and AUPRC with 95% confidence intervals are reported for lateral lymph\-node metastasis prediction and follicular versus papillary thyroid carcinoma subtype classification\. Training data comprised 4,756 malignant images from ZJH\-8K, of which 20 cases \(183 images\) were held out for validation\. Em dashes indicate tasks that were not evaluated\.MethodLymph Node MetastasisFTC/PTC subtypeAUROC↑\\uparrowAUPRC↑\\uparrowAUROC↑\\uparrowAUPRC↑\\uparrowRepViT\[[66](https://arxiv.org/html/2608.12590#bib.bib66)\]0\.7905±0\.06760\.7905\\pm 0\.06760\.8152±0\.06380\.8152\\pm 0\.06380\.6419±0\.08390\.6419\\pm 0\.08390\.6297±0\.09420\.6297\\pm 0\.0942LSNet\[[93](https://arxiv.org/html/2608.12590#bib.bib93)\]0\.5878±0\.08650\.5878\\pm 0\.08650\.6301±0\.08750\.6301\\pm 0\.08750\.4858±0\.09250\.4858\\pm 0\.09250\.4845±0\.09080\.4845\\pm 0\.0908UltraFedFM\[[65](https://arxiv.org/html/2608.12590#bib.bib65)\]0\.7757±0\.07310\.7757\\pm 0\.07310\.7902±0\.08450\.7902\\pm 0\.08450\.7365±0\.07440\.7365\\pm 0\.07440\.7582±0\.08240\.7582\\pm 0\.0824MedGemma\[[71](https://arxiv.org/html/2608.12590#bib.bib71)\]0\.8403±0\.04610\.8403\\pm 0\.04610\.8585±0\.05660\.8585\\pm 0\.05660\.6598±0\.08240\.6598\\pm 0\.08240\.6142±0\.10560\.6142\\pm 0\.1056Qwen3\-VL\-8B\-Instruct\[[94](https://arxiv.org/html/2608.12590#bib.bib94)\]0\.8070±0\.06320\.8070\\pm 0\.06320\.8055±0\.08000\.8055\\pm 0\.08000\.6056±0\.08660\.6056\\pm 0\.08660\.5539±0\.11180\.5539\\pm 0\.1118GPT\-5\[[68](https://arxiv.org/html/2608.12590#bib.bib68)\]0\.8410±0\.05750\.8410\\pm 0\.05750\.8629±0\.05330\.8629\\pm 0\.05330\.1604±0\.07060\.1604\\pm 0\.07060\.3638±0\.08470\.3638\\pm 0\.0847Gemini\-2\.5\-Pro\[[69](https://arxiv.org/html/2608.12590#bib.bib69)\]0\.5414±0\.07360\.5414\\pm 0\.07360\.5492±0\.09150\.5492\\pm 0\.09150\.3324±0\.08720\.3324\\pm 0\.08720\.4187±0\.08370\.4187\\pm 0\.0837LLNM\-Net\[[22](https://arxiv.org/html/2608.12590#bib.bib22)\]0\.7665±0\.06920\.7665\\pm 0\.06920\.7363±0\.08490\.7363\\pm 0\.0849––Tiger\-Model\[[23](https://arxiv.org/html/2608.12590#bib.bib23)\]––0\.7136±0\.08140\.7136\\pm 0\.08140\.7117±0\.11010\.7117\\pm 0\.1101ThyroidXAgent0\.8642±\\pm0\.05500\.8808±\\pm0\.05370\.8053±\\pm0\.05990\.7863±\\pm0\.0793Table S6:Effects of cumulative training\-data integration on segmentation and classification\. Models were trained on progressively expanded configurations and evaluated on independent test sets\. Segmentation performance is reported as Dice coefficient \(%,↑\\uparrowhigher is better\) and HD95 \(mm,↓\\downarrowlower is better\); classification performance as AUROC and AUPRC \(↑\\uparrowhigher is better\)\. Values are means±\\pm95% confidence intervals across five independent runs\. Dashes indicate that the test set was not applicable for the given task\. Segmentation training configurations: dataset1 \(TN3K\), dataset2 \(TN3K \+ ThyroidXL\), dataset3 \(TN3K \+ ThyroidXL \+ PKTN\), dataset4 \(TN3K \+ ThyroidXL \+ PKTN \+ TN5K\)\. Classification training configurations: dataset1 \(TN3K\), dataset2 \(TN3K \+ ThyroidXL\), dataset3 \(TN3K \+ ThyroidXL \+ TN5K\)\.TrainTestTN3KThyroidXLPKTNTN5KDDTIZJH\-8KRJH\-7KSegmentation – Dice \(%\)↑\\uparrowdataset182\.76±\\pm3\.5481\.97±\\pm2\.6379\.21±\\pm2\.7172\.18±\\pm5\.0878\.08±\\pm3\.0994\.57±\\pm0\.4380\.77±\\pm0\.44dataset281\.63±\\pm3\.8186\.84±\\pm1\.9081\.73±\\pm2\.2671\.07±\\pm5\.3576\.72±\\pm3\.3994\.85±\\pm0\.4082\.38±\\pm0\.42dataset380\.81±\\pm3\.8286\.00±\\pm2\.3181\.91±\\pm2\.2672\.82±\\pm4\.9184\.81±\\pm2\.4394\.82±\\pm0\.3991\.44±\\pm0\.15dataset481\.86±\\pm3\.7086\.97±\\pm2\.1983\.28±\\pm2\.1982\.57±\\pm3\.4684\.86±\\pm2\.3894\.77±\\pm0\.3991\.46±\\pm0\.15Segmentation – HD95 \(mm\)↓\\downarrowdataset113\.49±\\pm3\.838\.34±\\pm2\.3411\.73±\\pm3\.0711\.37±\\pm3\.4916\.93±\\pm3\.122\.30±\\pm0\.4711\.46±\\pm0\.50dataset215\.92±\\pm4\.584\.99±\\pm1\.589\.72±\\pm2\.7113\.64±\\pm4\.2818\.54±\\pm3\.171\.98±\\pm0\.419\.65±\\pm0\.44dataset315\.94±\\pm4\.345\.46±\\pm1\.6210\.92±\\pm3\.5211\.07±\\pm3\.2811\.97±\\pm3\.981\.93±\\pm0\.381\.87±\\pm0\.07dataset417\.00±\\pm5\.344\.74±\\pm1\.428\.89±\\pm2\.934\.64±\\pm1\.439\.91±\\pm2\.182\.07±\\pm0\.401\.88±\\pm0\.07Classification – AUROC↑\\uparrowdataset10\.7666±\\pm0\.030\.8713±\\pm0\.01–0\.8272±\\pm0\.020\.7244±\\pm0\.070\.9924±\\pm0\.01–dataset20\.7724±\\pm0\.040\.9254±\\pm0\.01–0\.8151±\\pm0\.020\.5762±\\pm0\.100\.9932±\\pm0\.01–dataset30\.7906±\\pm0\.030\.9288±\\pm0\.01–0\.9515±\\pm0\.010\.7623±\\pm0\.070\.9937±\\pm0\.01–Classification – AUPRC↑\\uparrowdataset10\.6806±\\pm0\.060\.8147±\\pm0\.03–0\.9151±\\pm0\.020\.3578±\\pm0\.130\.9951±\\pm0\.01–dataset20\.7237±\\pm0\.050\.9140±\\pm0\.01–0\.9074±\\pm0\.010\.3190±\\pm0\.140\.9968±\\pm0\.01–dataset30\.7188±\\pm0\.050\.9144±\\pm0\.01–0\.9803±\\pm0\.010\.4029±\\pm0\.140\.9967±\\pm0\.01–

Table S7:Per\-centre Dice similarity coefficient \(%\) for thyroid gland segmentation on the NHC\-MISD\-TUS external test set\. Values are reported as point estimates with 95% confidence intervals\. Bold indicates the best result in each row\.CenterNNThyroidXAgentMedSAM2MedSegXTransUNetUltraFedFMDice \(%\)↑\\uparrowDice \(%\)↑\\uparrowDice \(%\)↑\\uparrowDice \(%\)↑\\uparrowDice \(%\)↑\\uparrowOverall8,72159\.28\(58\.70, 59\.86\)56\.80\(56\.35, 57\.24\)54\.98\(54\.57, 55\.39\)53\.56\(53\.06, 54\.12\)33\.30\(32\.73, 33\.88\)THYB\_S\_ZJ241,17480\.64\(79\.84, 81\.40\)52\.96\(51\.96, 54\.03\)50\.32\(49\.25, 51\.42\)72\.91\(72\.03, 73\.79\)8\.49\(7\.59, 9\.42\)THYB\_S\_EN0494965\.30\(63\.82, 66\.78\)51\.05\(49\.46, 52\.51\)59\.83\(58\.66, 60\.96\)59\.65\(58\.11, 61\.15\)35\.33\(33\.81, 36\.91\)THYB\_S\_BJ0187661\.05\(59\.14, 62\.88\)61\.22\(59\.93, 62\.48\)51\.36\(49\.83, 52\.83\)56\.68\(54\.95, 58\.35\)27\.42\(25\.73, 29\.23\)THYB\_S\_SH0183243\.33\(41\.36, 45\.33\)64\.88\(63\.37, 66\.28\)55\.60\(54\.14, 56\.95\)42\.72\(40\.94, 44\.52\)49\.82\(48\.11, 51\.48\)THYB\_S\_ZJ0569455\.25\(53\.12, 57\.40\)54\.83\(53\.19, 56\.51\)53\.77\(52\.33, 55\.13\)50\.71\(48\.70, 52\.55\)30\.49\(28\.66, 32\.35\)THYB\_S\_SH0566961\.58\(59\.74, 63\.48\)59\.92\(58\.52, 61\.38\)56\.63\(55\.27, 57\.98\)50\.26\(48\.46, 52\.13\)34\.59\(32\.66, 36\.40\)THYB\_S\_NX0158252\.69\(50\.28, 54\.89\)45\.59\(44\.00, 47\.18\)52\.80\(51\.32, 54\.40\)54\.11\(51\.83, 56\.22\)32\.64\(30\.64, 34\.63\)THYB\_S\_ZJ0646048\.58\(46\.12, 50\.98\)62\.74\(60\.83, 64\.42\)55\.20\(53\.35, 56\.97\)42\.03\(39\.70, 44\.34\)35\.19\(32\.95, 37\.35\)THYB\_S\_QX0728966\.05\(63\.24, 68\.91\)50\.61\(48\.09, 53\.10\)59\.57\(57\.62, 61\.64\)62\.24\(59\.24, 65\.06\)46\.38\(43\.64, 49\.40\)THYB\_S\_AN0124336\.90\(33\.32, 40\.25\)74\.15\(71\.96, 76\.34\)56\.14\(53\.14, 59\.07\)36\.30\(33\.35, 39\.72\)60\.76\(57\.91, 63\.56\)THYB\_S\_CQ0323863\.69\(60\.50, 66\.62\)55\.01\(52\.63, 57\.49\)58\.52\(56\.01, 60\.89\)46\.06\(42\.68, 49\.56\)33\.93\(30\.82, 37\.30\)THYB\_S\_GZ0223359\.92\(56\.20, 63\.38\)62\.23\(60\.11, 64\.35\)55\.84\(53\.35, 58\.40\)54\.40\(51\.00, 58\.00\)37\.71\(34\.20, 41\.27\)THYB\_S\_JX0621554\.16\(50\.51, 58\.01\)63\.46\(60\.77, 65\.99\)55\.71\(52\.48, 58\.52\)45\.96\(42\.07, 49\.54\)46\.71\(43\.30, 50\.23\)THYB\_S\_JS0219036\.44\(32\.39, 40\.43\)64\.41\(61\.96, 66\.78\)56\.60\(54\.36, 58\.91\)34\.57\(31\.09, 38\.09\)50\.48\(46\.08, 54\.64\)THYB\_S\_YN0513755\.44\(50\.24, 60\.32\)52\.03\(48\.59, 55\.63\)63\.75\(60\.52, 66\.75\)63\.10\(58\.34, 67\.44\)14\.28\(10\.73, 18\.14\)THYB\_S\_EN0212055\.75\(50\.96, 60\.59\)39\.06\(36\.22, 42\.01\)51\.87\(48\.88, 54\.83\)38\.76\(33\.77, 43\.55\)40\.34\(35\.64, 44\.90\)THYB\_S\_AH0410165\.76\(59\.94, 70\.88\)52\.92\(49\.38, 56\.65\)61\.94\(58\.10, 65\.51\)56\.38\(50\.58, 61\.47\)47\.68\(41\.39, 54\.39\)THYB\_S\_GS039450\.96\(45\.22, 56\.84\)59\.17\(54\.56, 63\.60\)51\.74\(47\.42, 55\.76\)48\.24\(42\.53, 53\.76\)40\.09\(34\.46, 45\.45\)THYB\_S\_FJ038566\.63\(60\.76, 72\.05\)46\.90\(43\.13, 50\.64\)54\.13\(50\.57, 57\.52\)50\.83\(44\.94, 56\.45\)36\.40\(30\.87, 41\.92\)THYB\_S\_SD127569\.59\(64\.15, 74\.62\)64\.10\(60\.34, 67\.84\)59\.41\(55\.70, 63\.10\)58\.28\(52\.76, 63\.56\)33\.57\(27\.12, 39\.83\)THYB\_S\_JS017256\.33\(49\.82, 62\.97\)49\.38\(45\.18, 53\.56\)54\.34\(50\.59, 58\.03\)49\.09\(42\.61, 55\.51\)28\.27\(23\.18, 33\.88\)THYB\_S\_NM026657\.71\(52\.20, 63\.08\)52\.72\(47\.87, 57\.77\)52\.81\(47\.81, 57\.29\)46\.24\(40\.41, 52\.18\)36\.74\(30\.84, 42\.37\)THYB\_S\_JL045443\.92\(34\.91, 52\.81\)56\.22\(50\.23, 62\.17\)55\.83\(50\.74, 61\.29\)40\.85\(32\.13, 49\.16\)42\.49\(34\.52, 50\.97\)THYB\_S\_SH064349\.75\(41\.79, 57\.64\)51\.33\(45\.78, 56\.54\)45\.09\(37\.42, 51\.92\)46\.30\(37\.76, 54\.02\)22\.32\(15\.88, 29\.32\)THYB\_S\_BJ094154\.80\(46\.03, 63\.29\)55\.54\(49\.59, 61\.66\)55\.00\(50\.40, 59\.85\)39\.65\(31\.89, 47\.97\)52\.67\(44\.89, 60\.73\)THYB\_S\_SX044161\.37\(55\.03, 67\.56\)56\.50\(50\.56, 62\.42\)48\.79\(42\.56, 54\.12\)44\.04\(34\.60, 52\.15\)42\.36\(33\.90, 51\.28\)THYB\_S\_SC062938\.63\(29\.91, 48\.19\)69\.92\(62\.86, 76\.12\)58\.20\(49\.69, 65\.60\)32\.00\(25\.56, 38\.77\)59\.45\(50\.88, 67\.48\)THYB\_S\_XJ012568\.57\(58\.62, 77\.48\)55\.07\(48\.48, 61\.59\)57\.03\(50\.98, 62\.37\)60\.51\(51\.17, 68\.75\)14\.00\(7\.83, 21\.61\)THYB\_S\_YN012555\.83\(43\.67, 66\.82\)58\.55\(51\.86, 65\.06\)56\.48\(48\.32, 64\.21\)46\.59\(35\.73, 57\.93\)29\.32\(19\.14, 40\.82\)THYB\_S\_SD132269\.21\(57\.84, 80\.57\)72\.83\(63\.53, 80\.86\)63\.22\(55\.37, 71\.01\)41\.87\(31\.62, 52\.02\)46\.04\(34\.90, 56\.40\)THYB\_S\_GX011964\.80\(52\.35, 75\.87\)53\.73\(46\.83, 61\.09\)60\.04\(51\.68, 67\.03\)61\.17\(48\.99, 71\.80\)24\.77\(12\.11, 38\.80\)THYB\_S\_ZJ291685\.21\(82\.65, 87\.63\)52\.40\(43\.77, 61\.90\)67\.08\(59\.58, 74\.31\)47\.06\(32\.00, 61\.91\)3\.30\(0\.48, 7\.60\)THYB\_S\_HB07831\.68\(9\.84, 58\.05\)69\.62\(56\.11, 82\.55\)52\.10\(37\.92, 62\.81\)19\.55\(0\.08, 43\.94\)29\.09\(16\.83, 41\.56\)THYB\_S\_FJ01267\.14\(64\.04, 70\.23\)44\.00\(38\.25, 49\.75\)49\.95\(43\.08, 56\.82\)45\.18\(32\.48, 57\.89\)39\.20\(29\.38, 49\.01\)THYB\_S\_SD1423\.13\(2\.55, 3\.71\)46\.73\(36\.29, 57\.18\)58\.38\(54\.22, 62\.54\)44\.78\(31\.59, 57\.97\)37\.57\(31\.23, 43\.91\)Table S8:Per\-centre 95% Hausdorff distance \(HD95, mm\) for thyroid gland segmentation on the NHC\-MISD\-TUS external test set\. Values are reported as point estimates with 95% confidence intervals\. Bold indicates the best result in each row\.CenterNNThyroidXAgentMedSAM2MedSegXTransUNetUltraFedFMHD95 \(mm\)↓\\downarrowHD95 \(mm\)↓\\downarrowHD95 \(mm\)↓\\downarrowHD95 \(mm\)↓\\downarrowHD95 \(mm\)↓\\downarrowOverall8,72135\.83\(35\.14, 36\.49\)65\.71\(65\.07, 66\.38\)45\.87\(45\.36, 46\.37\)40\.65\(39\.95, 41\.35\)58\.17\(57\.35, 58\.99\)THYB\_S\_ZJ241,17414\.74\(13\.90, 15\.63\)75\.31\(73\.88, 76\.74\)64\.28\(62\.83, 65\.85\)22\.39\(21\.43, 23\.43\)53\.71\(50\.40, 56\.90\)THYB\_S\_EN0494930\.01\(28\.39, 31\.72\)73\.59\(71\.54, 75\.87\)42\.31\(40\.95, 43\.81\)34\.27\(32\.43, 36\.13\)63\.80\(61\.88, 65\.55\)THYB\_S\_BJ0187636\.67\(34\.70, 38\.76\)55\.12\(53\.38, 56\.96\)41\.33\(40\.11, 42\.83\)40\.85\(38\.94, 43\.05\)60\.27\(57\.54, 63\.41\)THYB\_S\_SH0183250\.89\(48\.25, 53\.54\)55\.08\(52\.59, 57\.51\)39\.63\(38\.18, 41\.11\)51\.90\(49\.43, 54\.61\)51\.25\(49\.37, 53\.41\)THYB\_S\_ZJ0569437\.22\(34\.87, 39\.57\)70\.81\(68\.16, 73\.22\)49\.93\(48\.19, 51\.64\)40\.86\(38\.46, 43\.25\)64\.30\(61\.86, 67\.00\)THYB\_S\_SH0566936\.52\(34\.23, 39\.01\)58\.39\(56\.18, 60\.42\)43\.68\(42\.20, 45\.23\)44\.98\(42\.65, 47\.48\)67\.24\(64\.77, 69\.87\)THYB\_S\_NX0158239\.66\(37\.22, 42\.54\)82\.07\(79\.47, 84\.56\)51\.93\(49\.68, 54\.07\)39\.06\(36\.74, 41\.62\)57\.51\(54\.52, 60\.33\)THYB\_S\_ZJ0646043\.86\(40\.98, 47\.06\)53\.93\(51\.41, 56\.72\)44\.67\(42\.75, 46\.81\)48\.71\(45\.68, 51\.91\)62\.98\(60\.17, 65\.91\)THYB\_S\_QX0728930\.43\(27\.06, 33\.67\)84\.32\(81\.13, 87\.54\)37\.11\(35\.21, 39\.31\)34\.16\(30\.75, 37\.53\)58\.07\(53\.90, 61\.53\)THYB\_S\_AN0124359\.89\(54\.84, 65\.27\)39\.11\(35\.71, 42\.69\)36\.76\(34\.09, 39\.70\)57\.02\(52\.09, 61\.87\)42\.73\(39\.24, 46\.33\)THYB\_S\_CQ0323836\.80\(33\.17, 40\.64\)67\.69\(64\.19, 71\.21\)39\.40\(36\.37, 42\.33\)49\.95\(45\.43, 54\.40\)63\.03\(58\.68, 67\.50\)THYB\_S\_GZ0223337\.00\(33\.22, 40\.65\)50\.49\(47\.63, 53\.64\)39\.41\(37\.30, 41\.67\)40\.19\(35\.91, 44\.49\)57\.61\(53\.08, 62\.19\)THYB\_S\_JX0621542\.43\(38\.02, 47\.27\)54\.75\(50\.60, 59\.00\)37\.59\(35\.11, 40\.19\)48\.81\(44\.42, 53\.58\)52\.03\(47\.75, 56\.31\)THYB\_S\_JS0219049\.35\(44\.19, 55\.17\)52\.35\(48\.11, 56\.54\)53\.87\(50\.49, 57\.35\)56\.88\(51\.46, 62\.82\)52\.07\(47\.01, 57\.13\)THYB\_S\_YN0513737\.57\(32\.03, 43\.12\)73\.49\(68\.84, 77\.86\)35\.07\(31\.77, 38\.50\)39\.92\(34\.22, 46\.49\)50\.39\(41\.86, 58\.71\)THYB\_S\_EN0212037\.62\(32\.53, 42\.62\)100\.8\(95\.9, 105\.8\)45\.67\(42\.65, 48\.58\)44\.64\(39\.27, 50\.38\)54\.04\(48\.46, 59\.63\)THYB\_S\_AH0410125\.82\(20\.55, 31\.83\)70\.36\(64\.77, 75\.45\)33\.52\(29\.78, 37\.56\)35\.53\(29\.34, 42\.79\)52\.02\(44\.23, 60\.59\)THYB\_S\_GS039444\.87\(37\.84, 51\.56\)61\.22\(54\.74, 67\.47\)44\.51\(40\.21, 49\.25\)48\.16\(40\.83, 55\.61\)52\.29\(45\.50, 59\.13\)THYB\_S\_FJ038527\.04\(21\.49, 33\.08\)83\.26\(77\.91, 88\.32\)48\.29\(43\.55, 53\.23\)40\.21\(34\.71, 46\.20\)57\.81\(51\.88, 63\.89\)THYB\_S\_SD127523\.80\(19\.40, 28\.64\)57\.62\(52\.08, 62\.79\)36\.03\(32\.04, 39\.97\)35\.69\(30\.37, 41\.96\)51\.62\(43\.29, 59\.76\)THYB\_S\_JS017240\.00\(32\.29, 47\.62\)79\.39\(72\.78, 85\.69\)45\.34\(40\.63, 50\.53\)40\.73\(33\.99, 47\.88\)64\.43\(55\.62, 73\.62\)THYB\_S\_NM026645\.94\(38\.55, 53\.94\)63\.04\(56\.60, 69\.43\)44\.14\(39\.41, 49\.57\)51\.02\(42\.31, 59\.30\)62\.80\(54\.21, 70\.52\)THYB\_S\_JL045453\.99\(42\.73, 65\.60\)67\.66\(59\.00, 76\.78\)39\.38\(34\.57, 44\.49\)50\.55\(39\.81, 61\.07\)57\.53\(46\.74, 67\.69\)THYB\_S\_SH064345\.77\(34\.95, 57\.09\)63\.35\(55\.55, 71\.99\)45\.14\(38\.78, 51\.03\)41\.91\(30\.31, 53\.31\)63\.26\(50\.44, 76\.60\)THYB\_S\_BJ094136\.63\(27\.94, 47\.32\)63\.20\(55\.22, 70\.78\)42\.93\(36\.63, 49\.75\)41\.72\(31\.74, 52\.87\)44\.38\(34\.57, 53\.98\)THYB\_S\_SX044138\.80\(31\.32, 47\.14\)66\.23\(57\.04, 75\.13\)48\.62\(44\.02, 53\.56\)53\.31\(42\.58, 66\.15\)68\.42\(55\.93, 82\.59\)THYB\_S\_SC062951\.99\(36\.85, 66\.32\)45\.67\(35\.65, 56\.18\)38\.52\(30\.69, 46\.92\)56\.62\(43\.06, 70\.35\)43\.15\(33\.11, 53\.56\)THYB\_S\_XJ012529\.06\(21\.16, 38\.25\)54\.51\(46\.58, 63\.05\)44\.52\(37\.28, 52\.45\)44\.96\(34\.41, 56\.69\)52\.44\(36\.21, 68\.45\)THYB\_S\_YN012535\.59\(22\.18, 54\.31\)62\.96\(52\.87, 73\.71\)38\.48\(32\.35, 44\.59\)54\.30\(37\.89, 71\.41\)66\.12\(49\.33, 83\.02\)THYB\_S\_SD132223\.75\(12\.47, 36\.80\)43\.51\(30\.68, 58\.21\)36\.56\(28\.65, 44\.41\)52\.16\(38\.57, 65\.81\)66\.39\(52\.20, 84\.55\)THYB\_S\_GX011933\.19\(22\.74, 43\.33\)78\.95\(67\.13, 88\.94\)29\.49\(22\.85, 36\.30\)38\.98\(26\.03, 54\.00\)38\.20\(20\.84, 60\.52\)THYB\_S\_ZJ291612\.12\(6\.47, 20\.65\)65\.31\(54\.30, 76\.11\)22\.34\(17\.11, 28\.07\)38\.17\(22\.88, 54\.66\)45\.26\(21\.23, 69\.87\)THYB\_S\_HB07830\.47\(7\.26, 58\.20\)44\.74\(23\.64, 68\.97\)32\.40\(23\.17, 45\.39\)39\.33\(7\.94, 74\.12\)80\.20\(66\.41, 92\.63\)THYB\_S\_FJ01258\.89\(23\.60, 94\.19\)75\.44\(66\.29, 84\.60\)33\.92\(27\.53, 40\.31\)47\.04\(34\.00, 60\.08\)49\.86\(33\.24, 66\.48\)THYB\_S\_SD14278\.47\(78\.16, 78\.77\)98\.68\(81\.39, 115\.97\)55\.96\(43\.32, 68\.59\)34\.84\(19\.69, 50\.00\)40\.44\(36\.06, 44\.82\)Table S9:Per\-centre Dice similarity coefficient \(%\) for thyroid nodule segmentation on the NHC\-MISD\-TUS external test set\. Values are reported as point estimates with 95% confidence intervals\. Bold indicates the best result in each row\.CenterNNThyroidXAgentMedSAM2MedSegXTransUNetUltraFedFMDice \(%\)↑\\uparrowDice \(%\)↑\\uparrowDice \(%\)↑\\uparrowDice \(%\)↑\\uparrowDice \(%\)↑\\uparrowOverall6,38482\.31\(81\.78, 82\.83\)23\.20\(22\.58, 23\.86\)28\.82\(28\.11, 29\.53\)72\.99\(72\.29, 73\.68\)58\.87\(58\.04, 59\.68\)THYB\_S\_EN0494682\.92\(81\.77, 84\.00\)15\.54\(14\.29, 16\.78\)22\.25\(20\.68, 23\.66\)74\.43\(72\.85, 75\.95\)56\.67\(54\.69, 58\.62\)THYB\_S\_SH0183488\.78\(88\.01, 89\.61\)39\.13\(37\.10, 41\.01\)46\.16\(44\.13, 48\.29\)84\.48\(83\.49, 85\.58\)71\.64\(69\.99, 73\.33\)THYB\_S\_ZJ0569082\.45\(81\.21, 83\.64\)15\.47\(13\.92, 17\.01\)16\.32\(14\.69, 17\.96\)67\.34\(65\.14, 69\.41\)55\.66\(52\.98, 58\.07\)THYB\_S\_SH0566585\.53\(84\.17, 86\.86\)20\.11\(18\.43, 21\.73\)29\.48\(27\.34, 31\.47\)75\.56\(73\.63, 77\.41\)60\.40\(57\.93, 62\.74\)THYB\_S\_NX0155777\.41\(75\.39, 79\.40\)16\.92\(15\.32, 18\.56\)23\.51\(21\.59, 25\.42\)68\.50\(66\.26, 70\.91\)47\.56\(44\.76, 50\.35\)THYB\_S\_ZJ0646084\.52\(82\.82, 86\.14\)25\.61\(23\.25, 27\.95\)28\.94\(26\.49, 31\.44\)75\.20\(73\.03, 77\.38\)61\.29\(58\.63, 63\.89\)THYB\_S\_ZJ2432145\.53\(41\.66, 48\.73\)5\.47\(4\.61, 6\.45\)4\.17\(3\.42, 5\.08\)20\.51\(17\.52, 23\.78\)14\.92\(12\.15, 17\.89\)THYB\_S\_QX0728085\.51\(83\.52, 87\.35\)22\.27\(19\.68, 24\.88\)28\.59\(25\.55, 31\.77\)77\.89\(74\.95, 80\.54\)68\.19\(65\.16, 71\.49\)THYB\_S\_AN0124391\.60\(89\.81, 92\.94\)53\.29\(49\.81, 56\.91\)61\.68\(58\.00, 65\.00\)88\.34\(86\.28, 90\.14\)78\.19\(75\.20, 81\.15\)THYB\_S\_CQ0318682\.72\(79\.86, 85\.39\)18\.16\(15\.15, 21\.46\)27\.06\(23\.17, 31\.10\)72\.57\(68\.81, 76\.37\)55\.25\(50\.23, 60\.01\)THYB\_S\_JS0218487\.67\(85\.58, 89\.59\)34\.39\(30\.61, 38\.22\)35\.40\(30\.90, 39\.94\)82\.96\(80\.24, 85\.39\)71\.70\(67\.73, 75\.66\)THYB\_S\_JX0616787\.50\(84\.81, 89\.82\)35\.52\(31\.67, 39\.88\)45\.33\(40\.68, 50\.25\)83\.78\(80\.89, 86\.39\)73\.76\(70\.40, 77\.15\)THYB\_S\_BJ0114269\.73\(64\.15, 74\.65\)16\.56\(13\.36, 19\.94\)21\.83\(17\.92, 25\.72\)58\.88\(53\.18, 64\.45\)46\.92\(41\.12, 52\.74\)THYB\_S\_EN0210784\.38\(81\.60, 86\.98\)11\.99\(9\.13, 15\.33\)19\.36\(15\.55, 23\.76\)73\.29\(69\.08, 77\.46\)50\.71\(44\.52, 57\.00\)THYB\_S\_GZ029289\.21\(87\.74, 90\.46\)27\.23\(22\.31, 32\.19\)30\.27\(24\.62, 36\.27\)84\.00\(81\.49, 86\.01\)73\.15\(68\.35, 77\.38\)THYB\_S\_GS039185\.66\(81\.56, 88\.93\)28\.49\(22\.48, 34\.72\)35\.22\(28\.72, 41\.50\)78\.93\(74\.06, 83\.12\)64\.75\(58\.31, 70\.89\)THYB\_S\_JS016685\.58\(81\.87, 88\.54\)13\.65\(10\.66, 16\.77\)18\.99\(15\.28, 23\.17\)76\.18\(70\.10, 81\.21\)58\.55\(50\.67, 65\.99\)THYB\_S\_FJ036479\.73\(74\.98, 83\.85\)10\.49\(7\.58, 13\.92\)14\.62\(10\.01, 20\.30\)64\.17\(55\.49, 71\.69\)50\.24\(41\.88, 59\.00\)THYB\_S\_NM025083\.34\(75\.81, 89\.18\)21\.78\(15\.16, 29\.35\)27\.11\(20\.14, 34\.56\)76\.31\(67\.18, 83\.75\)61\.40\(52\.34, 69\.28\)THYB\_S\_SH063170\.85\(59\.24, 79\.88\)4\.60\(3\.47, 5\.86\)8\.28\(5\.59, 11\.59\)55\.16\(39\.98, 68\.39\)46\.54\(33\.59, 59\.77\)THYB\_S\_AH043091\.71\(89\.82, 93\.31\)39\.78\(30\.88, 48\.49\)48\.60\(37\.90, 58\.61\)84\.52\(77\.22, 89\.78\)79\.09\(70\.70, 84\.79\)THYB\_S\_SC062990\.16\(88\.07, 91\.95\)44\.20\(32\.95, 53\.60\)49\.38\(38\.45, 58\.61\)88\.29\(85\.51, 90\.65\)76\.57\(67\.27, 84\.35\)THYB\_S\_SX042789\.20\(86\.29, 91\.70\)17\.84\(12\.89, 22\.84\)16\.51\(11\.20, 23\.50\)87\.47\(84\.52, 89\.90\)68\.19\(57\.72, 77\.31\)THYB\_S\_BJ092286\.20\(81\.55, 90\.25\)30\.06\(19\.75, 40\.78\)34\.98\(22\.71, 47\.67\)81\.62\(74\.86, 87\.53\)70\.13\(63\.36, 77\.01\)THYB\_S\_YN012287\.24\(79\.41, 92\.13\)23\.20\(13\.43, 34\.02\)28\.13\(18\.31, 39\.88\)80\.58\(70\.28, 87\.54\)44\.07\(29\.99, 57\.81\)THYB\_S\_SD131892\.74\(91\.15, 94\.06\)52\.54\(39\.34, 64\.38\)61\.58\(48\.47, 73\.53\)90\.83\(88\.24, 92\.78\)77\.31\(70\.93, 82\.80\)THYB\_S\_JL041681\.27\(67\.59, 93\.13\)36\.60\(22\.73, 50\.94\)48\.61\(31\.53, 65\.36\)76\.87\(62\.21, 89\.36\)70\.49\(53\.00, 86\.26\)THYB\_S\_SD121683\.71\(69\.71, 93\.91\)53\.71\(39\.32, 67\.16\)62\.86\(47\.03, 78\.03\)79\.72\(62\.56, 93\.40\)74\.86\(58\.02, 88\.45\)THYB\_S\_YN051291\.81\(88\.17, 94\.81\)59\.02\(42\.58, 71\.76\)70\.91\(51\.90, 85\.39\)91\.01\(87\.57, 93\.88\)68\.71\(47\.83, 85\.32\)THYB\_S\_XJ01882\.71\(73\.69, 90\.34\)13\.18\(4\.03, 24\.95\)14\.95\(4\.71, 29\.10\)56\.32\(33\.11, 79\.31\)45\.53\(20\.28, 71\.40\)THYB\_S\_GX01485\.84\(83\.15, 88\.54\)20\.06\(9\.70, 30\.41\)38\.09\(22\.52, 49\.68\)84\.85\(78\.42, 89\.87\)70\.97\(60\.58, 83\.12\)THYB\_S\_FJ01279\.74\(78\.18, 81\.31\)6\.47\(4\.46, 8\.48\)12\.32\(8\.26, 16\.37\)81\.94\(77\.73, 86\.14\)46\.04\(36\.60, 55\.49\)THYB\_S\_HB07221\.28\(18\.58, 23\.99\)1\.90\(1\.48, 2\.31\)1\.86\(1\.61, 2\.11\)23\.49\(0\.00, 46\.97\)17\.79\(11\.76, 23\.82\)Table S10:Per\-centre 95% Hausdorff distance \(HD95, mm\) for thyroid nodule segmentation on the NHC\-MISD\-TUS external test set\. Values are reported as point estimates with 95% confidence intervals\. Bold indicates the best result in each row\.CenterNNThyroidXAgentMedSAM2MedSegXTransUNetUltraFedFMHD95 \(mm\)↓\\downarrowHD95 \(mm\)↓\\downarrowHD95 \(mm\)↓\\downarrowHD95 \(mm\)↓\\downarrowHD95 \(mm\)↓\\downarrowOverall6,3849\.41\(8\.89, 9\.93\)101\.2\(100\.2, 102\.0\)86\.96\(86\.08, 87\.84\)10\.62\(10\.14, 11\.12\)27\.04\(26\.28, 27\.75\)THYB\_S\_EN049466\.95\(6\.05, 7\.91\)113\.0\(111\.0, 114\.9\)91\.23\(89\.41, 92\.99\)9\.27\(8\.31, 10\.37\)33\.50\(31\.47, 35\.76\)THYB\_S\_SH018346\.41\(5\.50, 7\.36\)79\.94\(77\.11, 82\.81\)65\.42\(63\.21, 67\.68\)8\.32\(7\.35, 9\.27\)21\.31\(19\.64, 22\.92\)THYB\_S\_ZJ056905\.00\(4\.31, 5\.76\)108\.5\(106\.1, 110\.6\)106\.8\(104\.6, 109\.0\)7\.73\(6\.88, 8\.61\)21\.22\(19\.59, 23\.13\)THYB\_S\_SH056657\.60\(6\.34, 8\.89\)101\.8\(99\.4, 104\.1\)81\.02\(78\.74, 83\.33\)10\.17\(8\.84, 11\.51\)28\.23\(25\.93, 30\.37\)THYB\_S\_NX0155712\.95\(10\.96, 14\.86\)112\.7\(110\.1, 115\.1\)92\.51\(89\.99, 94\.97\)13\.02\(11\.18, 14\.80\)30\.17\(27\.56, 32\.82\)THYB\_S\_ZJ064607\.62\(6\.22, 9\.24\)92\.99\(90\.07, 95\.87\)88\.02\(85\.09, 91\.28\)10\.88\(9\.34, 12\.66\)26\.09\(23\.54, 28\.75\)THYB\_S\_ZJ2432145\.03\(39\.74, 50\.67\)142\.9\(139\.7, 145\.9\)140\.9\(137\.7, 143\.8\)30\.45\(26\.09, 34\.77\)45\.71\(40\.11, 51\.74\)THYB\_S\_QX072806\.69\(5\.08, 8\.55\)110\.2\(106\.4, 113\.9\)85\.79\(82\.50, 89\.13\)8\.11\(6\.54, 10\.16\)22\.53\(19\.71, 25\.68\)THYB\_S\_AN012434\.70\(3\.51, 6\.11\)60\.30\(55\.67, 64\.85\)49\.32\(45\.45, 53\.68\)7\.19\(5\.65, 9\.00\)16\.62\(14\.15, 19\.21\)THYB\_S\_CQ031868\.81\(5\.97, 12\.38\)105\.0\(100\.0, 109\.4\)81\.35\(77\.26, 85\.23\)12\.51\(9\.28, 16\.26\)33\.57\(28\.12, 38\.94\)THYB\_S\_JS021846\.68\(4\.75, 8\.90\)80\.55\(75\.49, 85\.50\)84\.81\(78\.76, 90\.71\)8\.78\(6\.74, 10\.90\)24\.40\(20\.23, 28\.77\)THYB\_S\_JX061677\.25\(5\.01, 9\.88\)82\.29\(76\.67, 87\.66\)63\.70\(58\.48, 68\.68\)9\.18\(6\.44, 12\.50\)20\.12\(16\.76, 23\.56\)THYB\_S\_BJ0114220\.72\(15\.43, 26\.10\)101\.8\(96\.3, 107\.7\)92\.30\(87\.28, 97\.40\)20\.10\(15\.47, 25\.01\)34\.87\(29\.55, 40\.13\)THYB\_S\_EN021074\.49\(3\.07, 6\.59\)128\.9\(123\.9, 133\.4\)92\.21\(87\.64, 96\.59\)8\.45\(6\.20, 11\.49\)29\.38\(24\.47, 34\.60\)THYB\_S\_GZ02924\.51\(3\.43, 5\.77\)84\.00\(78\.53, 89\.58\)81\.50\(74\.63, 88\.26\)7\.53\(5\.65, 9\.73\)17\.57\(13\.41, 22\.16\)THYB\_S\_GS03917\.80\(4\.49, 11\.81\)89\.69\(81\.70, 97\.29\)75\.21\(68\.30, 82\.38\)8\.70\(5\.54, 12\.57\)24\.98\(19\.70, 31\.13\)THYB\_S\_JS01665\.69\(3\.49, 8\.71\)111\.8\(106\.2, 117\.2\)94\.67\(89\.47, 99\.84\)9\.84\(6\.27, 14\.06\)27\.35\(19\.68, 35\.58\)THYB\_S\_FJ03646\.99\(3\.86, 10\.72\)115\.9\(109\.6, 122\.4\)107\.7\(100\.0, 115\.2\)8\.87\(5\.71, 12\.92\)22\.36\(16\.53, 28\.70\)THYB\_S\_NM02506\.75\(3\.95, 10\.41\)95\.61\(85\.83, 104\.48\)82\.78\(75\.37, 90\.43\)6\.49\(4\.04, 9\.79\)21\.73\(16\.43, 27\.52\)THYB\_S\_SH063116\.01\(6\.79, 25\.90\)115\.0\(110\.3, 120\.7\)94\.85\(88\.68, 100\.91\)7\.08\(3\.20, 12\.90\)28\.90\(17\.04, 42\.21\)THYB\_S\_AH04303\.09\(2\.14, 4\.32\)81\.37\(70\.67, 92\.67\)66\.34\(54\.18, 78\.80\)9\.13\(4\.11, 16\.74\)18\.64\(11\.53, 27\.65\)THYB\_S\_SC06294\.72\(2\.94, 6\.80\)71\.23\(58\.43, 85\.23\)61\.37\(51\.00, 72\.75\)5\.60\(3\.63, 7\.90\)20\.14\(10\.14, 31\.55\)THYB\_S\_SX04273\.54\(2\.39, 4\.80\)106\.5\(98\.0, 114\.6\)108\.8\(98\.9, 118\.0\)4\.05\(3\.12, 5\.09\)25\.37\(14\.31, 39\.60\)THYB\_S\_BJ09225\.56\(3\.29, 8\.33\)85\.51\(71\.97, 98\.12\)77\.91\(63\.39, 92\.08\)7\.06\(3\.88, 11\.19\)18\.84\(11\.45, 26\.82\)THYB\_S\_YN01224\.61\(2\.12, 8\.49\)96\.72\(83\.10, 109\.97\)73\.60\(64\.25, 81\.65\)4\.98\(3\.42, 6\.90\)40\.06\(27\.32, 54\.75\)THYB\_S\_SD13183\.32\(2\.22, 4\.57\)61\.73\(47\.66, 76\.68\)45\.00\(32\.04, 58\.04\)4\.51\(2\.74, 6\.60\)21\.77\(15\.46, 28\.42\)THYB\_S\_JL041614\.24\(2\.17, 31\.43\)88\.05\(70\.97, 104\.04\)58\.06\(41\.17, 75\.76\)12\.51\(3\.43, 23\.13\)16\.36\(6\.05, 28\.72\)THYB\_S\_SD12169\.76\(3\.28, 17\.71\)57\.71\(38\.35, 78\.90\)50\.35\(31\.21, 71\.00\)8\.08\(2\.22, 15\.94\)18\.16\(7\.70, 30\.26\)THYB\_S\_YN05123\.22\(2\.05, 4\.34\)61\.17\(44\.75, 80\.17\)38\.97\(24\.66, 57\.04\)4\.86\(2\.25, 9\.08\)13\.71\(6\.45, 21\.71\)THYB\_S\_XJ0186\.99\(1\.71, 16\.39\)111\.3\(92\.8, 131\.6\)99\.22\(85\.22, 115\.90\)20\.73\(4\.35, 42\.34\)24\.55\(5\.64, 47\.07\)THYB\_S\_GX0145\.57\(3\.00, 8\.47\)101\.0\(81\.2, 123\.7\)78\.07\(56\.31, 92\.33\)5\.12\(2\.62, 8\.46\)15\.24\(6\.58, 25\.12\)THYB\_S\_FJ0128\.32\(4\.00, 12\.65\)112\.3\(104\.0, 120\.7\)86\.49\(73\.93, 99\.04\)4\.41\(2\.83, 6\.00\)56\.71\(26\.29, 87\.13\)THYB\_S\_HB07230\.80\(18\.03, 43\.57\)109\.0\(98\.6, 119\.4\)104\.4\(95\.2, 113\.5\)4\.70\(0\.00, 9\.39\)19\.30\(8\.60, 30\.00\)Table S11:Per\-centre AUROC for benign versus malignant thyroid nodule classification on the NHC\-MISD\-TUS external test set\. Values are reported as point estimates with 95% confidence intervals\. Bold indicates the best result in each row\.CenterNNThyroidXAgentBiomedCLIPMedSigLIPUltraFedFMAUROC↑\\uparrowAUROC↑\\uparrowAUROC↑\\uparrowAUROC↑\\uparrowOverall4,9990\.819\(0\.808, 0\.831\)0\.434\(0\.418, 0\.450\)0\.520\(0\.503, 0\.535\)0\.436\(0\.420, 0\.452\)THYB\_S\_SH017920\.896\(0\.873, 0\.916\)0\.327\(0\.290, 0\.368\)0\.552\(0\.517, 0\.594\)0\.331\(0\.295, 0\.368\)THYB\_S\_ZJ056470\.811\(0\.773, 0\.844\)0\.526\(0\.480, 0\.576\)0\.586\(0\.546, 0\.627\)0\.494\(0\.455, 0\.533\)THYB\_S\_EN046260\.846\(0\.808, 0\.876\)0\.345\(0\.289, 0\.404\)0\.538\(0\.472, 0\.597\)0\.528\(0\.476, 0\.584\)THYB\_S\_SH055390\.719\(0\.672, 0\.758\)0\.331\(0\.280, 0\.381\)0\.543\(0\.492, 0\.595\)0\.370\(0\.324, 0\.423\)THYB\_S\_ZJ064520\.811\(0\.768, 0\.849\)0\.468\(0\.412, 0\.520\)0\.594\(0\.541, 0\.651\)0\.539\(0\.486, 0\.592\)THYB\_S\_QX072790\.870\(0\.814, 0\.917\)0\.457\(0\.381, 0\.537\)0\.630\(0\.534, 0\.711\)0\.428\(0\.338, 0\.520\)THYB\_S\_ZJ242690\.590\(0\.346, 0\.832\)0\.412\(0\.270, 0\.544\)0\.673\(0\.493, 0\.823\)0\.463\(0\.255, 0\.678\)THYB\_S\_AN012360\.862\(0\.784, 0\.923\)0\.241\(0\.165, 0\.319\)0\.601\(0\.519, 0\.687\)0\.292\(0\.202, 0\.390\)THYB\_S\_CQ031720\.784\(0\.716, 0\.849\)0\.452\(0\.366, 0\.550\)0\.521\(0\.431, 0\.608\)0\.440\(0\.350, 0\.532\)THYB\_S\_JS021710\.746\(0\.667, 0\.815\)0\.504\(0\.417, 0\.596\)0\.675\(0\.599, 0\.750\)0\.415\(0\.328, 0\.501\)THYB\_S\_JX061530\.903\(0\.853, 0\.951\)0\.399\(0\.315, 0\.487\)0\.616\(0\.524, 0\.702\)0\.480\(0\.390, 0\.575\)THYB\_S\_EN02990\.851\(0\.741, 0\.941\)0\.383\(0\.248, 0\.543\)0\.749\(0\.602, 0\.878\)0\.329\(0\.166, 0\.488\)THYB\_S\_BJ01930\.749\(0\.635, 0\.851\)0\.245\(0\.154, 0\.345\)0\.400\(0\.235, 0\.562\)0\.460\(0\.319, 0\.609\)THYB\_S\_GZ02900\.693\(0\.578, 0\.795\)0\.325\(0\.198, 0\.471\)0\.499\(0\.374, 0\.614\)0\.366\(0\.223, 0\.530\)THYB\_S\_GS03840\.838\(0\.735, 0\.924\)0\.244\(0\.125, 0\.373\)0\.443\(0\.311, 0\.566\)0\.470\(0\.328, 0\.607\)THYB\_S\_JS01590\.781\(0\.491, 1\.000\)0\.509\(0\.069, 0\.947\)0\.281\(0\.035, 0\.552\)0\.070\(0\.017, 0\.500\)THYB\_S\_NM02490\.653\(0\.497, 0\.800\)0\.408\(0\.253, 0\.567\)0\.719\(0\.549, 0\.875\)0\.393\(0\.202, 0\.595\)THYB\_S\_SX04270\.860\(0\.500, 1\.000\)0\.640\(0\.231, 0\.962\)0\.500\(0\.077, 0\.924\)0\.620\(0\.400, 0\.846\)THYB\_S\_AH04240\.844\(0\.663, 0\.979\)0\.617\(0\.368, 0\.838\)0\.586\(0\.305, 0\.838\)0\.266\(0\.076, 0\.523\)THYB\_S\_SC06220\.876\(0\.702, 1\.000\)0\.256\(0\.050, 0\.484\)0\.612\(0\.350, 0\.839\)0\.141\(0\.009, 0\.325\)THYB\_S\_YN01210\.853\(0\.618, 1\.000\)0\.353\(0\.000, 0\.778\)0\.500\(0\.105, 0\.895\)0\.176\(0\.000, 0\.400\)THYB\_S\_BJ09200\.586\(0\.319, 0\.849\)0\.505\(0\.213, 0\.798\)0\.636\(0\.341, 0\.885\)0\.404\(0\.150, 0\.687\)THYB\_S\_SD13180\.875\(0\.636, 1\.000\)0\.536\(0\.125, 1\.000\)0\.714\(0\.415, 0\.956\)0\.304\(0\.000, 0\.623\)THYB\_S\_SD12150\.929\(0\.500, 1\.000\)0\.143\(0\.000, 0\.500\)0\.571\(0\.356, 0\.786\)0\.071\(0\.000, 0\.500\)THYB\_S\_JL04120\.407\(0\.000, 1\.000\)0\.074\(0\.000, 0\.446\)0\.148\(0\.000, 0\.500\)0\.296\(0\.000, 0\.727\)THYB\_S\_YN05120\.500\(0\.500, 0\.500\)0\.500\(0\.500, 0\.500\)0\.500\(0\.500, 0\.500\)0\.500\(0\.500, 0\.500\)THYB\_S\_XJ0151\.000\(0\.500, 1\.000\)0\.667\(0\.000, 1\.000\)0\.333\(0\.000, 1\.000\)0\.833\(0\.250, 1\.000\)THYB\_S\_GX0140\.500\(0\.500, 0\.500\)0\.500\(0\.500, 0\.500\)0\.500\(0\.500, 0\.500\)0\.500\(0\.500, 0\.500\)THYB\_S\_SH0641\.000\(0\.500, 1\.000\)0\.000\(0\.000, 0\.500\)0\.000\(0\.000, 0\.500\)0\.500\(0\.000, 1\.000\)THYB\_S\_FJ0120\.500\(0\.500, 0\.500\)0\.500\(0\.500, 0\.500\)0\.500\(0\.500, 0\.500\)0\.500\(0\.500, 0\.500\)THYB\_S\_FJ0320\.500\(0\.500, 0\.500\)0\.500\(0\.500, 0\.500\)0\.500\(0\.500, 0\.500\)0\.500\(0\.500, 0\.500\)THYB\_S\_NX0110\.500\(0\.500, 0\.500\)0\.500\(0\.500, 0\.500\)0\.500\(0\.500, 0\.500\)0\.500\(0\.500, 0\.500\)THYB\_S\_HB070––––THYB\_S\_SD140––––THYB\_S\_ZJ290––––Table S12:Per\-centre AUPRC for benign versus malignant thyroid nodule classification on the NHC\-MISD\-TUS external test set\. Values are reported as point estimates with 95% confidence intervals\. Bold indicates the best result in each row\.CenterNNThyroidXAgentBiomedCLIPMedSigLIPUltraFedFMAUPRC↑\\uparrowAUPRC↑\\uparrowAUPRC↑\\uparrowAUPRC↑\\uparrowOverall4,9990\.823\(0\.807, 0\.838\)0\.450\(0\.433, 0\.467\)0\.519\(0\.499, 0\.539\)0\.465\(0\.447, 0\.483\)THYB\_S\_SH017920\.859\(0\.817, 0\.893\)0\.323\(0\.290, 0\.356\)0\.450\(0\.404, 0\.504\)0\.340\(0\.304, 0\.380\)THYB\_S\_ZJ056470\.746\(0\.690, 0\.798\)0\.460\(0\.412, 0\.518\)0\.519\(0\.469, 0\.580\)0\.445\(0\.393, 0\.504\)THYB\_S\_EN046260\.969\(0\.957, 0\.978\)0\.796\(0\.758, 0\.840\)0\.851\(0\.812, 0\.888\)0\.871\(0\.837, 0\.904\)THYB\_S\_SH055390\.832\(0\.787, 0\.865\)0\.539\(0\.496, 0\.589\)0\.680\(0\.630, 0\.734\)0\.555\(0\.509, 0\.610\)THYB\_S\_ZJ064520\.811\(0\.752, 0\.864\)0\.502\(0\.442, 0\.573\)0\.597\(0\.526, 0\.670\)0\.554\(0\.490, 0\.626\)THYB\_S\_QX072790\.575\(0\.432, 0\.706\)0\.163\(0\.107, 0\.235\)0\.270\(0\.173, 0\.385\)0\.158\(0\.107, 0\.245\)THYB\_S\_ZJ242690\.164\(0\.026, 0\.423\)0\.032\(0\.014, 0\.054\)0\.068\(0\.027, 0\.137\)0\.051\(0\.017, 0\.137\)THYB\_S\_AN012360\.588\(0\.430, 0\.756\)0\.110\(0\.082, 0\.147\)0\.195\(0\.141, 0\.261\)0\.121\(0\.087, 0\.172\)THYB\_S\_CQ031720\.876\(0\.819, 0\.922\)0\.603\(0\.510, 0\.706\)0\.637\(0\.544, 0\.735\)0\.559\(0\.474, 0\.651\)THYB\_S\_JS021710\.734\(0\.629, 0\.822\)0\.502\(0\.414, 0\.621\)0\.649\(0\.544, 0\.750\)0\.459\(0\.369, 0\.561\)THYB\_S\_JX061530\.897\(0\.839, 0\.946\)0\.351\(0\.276, 0\.438\)0\.545\(0\.422, 0\.671\)0\.409\(0\.314, 0\.522\)THYB\_S\_EN02990\.969\(0\.937, 0\.992\)0\.812\(0\.710, 0\.908\)0\.939\(0\.888, 0\.983\)0\.775\(0\.669, 0\.899\)THYB\_S\_BJ01930\.930\(0\.880, 0\.970\)0\.692\(0\.583, 0\.825\)0\.743\(0\.634, 0\.858\)0\.764\(0\.661, 0\.884\)THYB\_S\_GZ02900\.850\(0\.760, 0\.921\)0\.558\(0\.460, 0\.695\)0\.702\(0\.581, 0\.825\)0\.569\(0\.457, 0\.697\)THYB\_S\_GS03840\.921\(0\.847, 0\.970\)0\.578\(0\.459, 0\.706\)0\.692\(0\.559, 0\.816\)0\.674\(0\.553, 0\.808\)THYB\_S\_JS01590\.990\(0\.966, 1\.000\)0\.964\(0\.885, 1\.000\)0\.950\(0\.862, 1\.000\)0\.921\(0\.808, 1\.000\)THYB\_S\_NM02490\.810\(0\.662, 0\.920\)0\.594\(0\.441, 0\.781\)0\.777\(0\.611, 0\.941\)0\.562\(0\.420, 0\.729\)THYB\_S\_SX04270\.988\(0\.959, 1\.000\)0\.960\(0\.873, 1\.000\)0\.931\(0\.792, 1\.000\)0\.965\(0\.895, 1\.000\)THYB\_S\_AH04240\.746\(0\.430, 0\.967\)0\.463\(0\.222, 0\.800\)0\.547\(0\.203, 0\.819\)0\.258\(0\.130, 0\.450\)THYB\_S\_SC06220\.873\(0\.653, 1\.000\)0\.390\(0\.229, 0\.609\)0\.704\(0\.409, 0\.898\)0\.356\(0\.209, 0\.573\)THYB\_S\_YN01210\.968\(0\.901, 1\.000\)0\.748\(0\.537, 0\.990\)0\.832\(0\.615, 0\.993\)0\.737\(0\.498, 0\.947\)THYB\_S\_BJ09200\.683\(0\.415, 0\.910\)0\.687\(0\.406, 0\.895\)0\.699\(0\.431, 0\.925\)0\.559\(0\.292, 0\.806\)THYB\_S\_SD13180\.567\(0\.200, 1\.000\)0\.446\(0\.059, 1\.000\)0\.415\(0\.111, 0\.900\)0\.196\(0\.056, 0\.415\)THYB\_S\_SD12150\.500\(0\.000, 1\.000\)0\.077\(0\.000, 0\.177\)0\.143\(0\.000, 0\.365\)0\.071\(0\.000, 0\.150\)THYB\_S\_JL04120\.491\(0\.083, 1\.000\)0\.187\(0\.083, 0\.379\)0\.199\(0\.083, 0\.404\)0\.241\(0\.083, 0\.610\)THYB\_S\_YN05120\.000\(0\.000, 0\.000\)0\.000\(0\.000, 0\.000\)0\.000\(0\.000, 0\.000\)0\.000\(0\.000, 0\.000\)THYB\_S\_XJ0151\.000\(1\.000, 1\.000\)0\.806\(0\.333, 1\.000\)0\.589\(0\.200, 1\.000\)0\.917\(0\.417, 1\.000\)THYB\_S\_GX0141\.000\(1\.000, 1\.000\)1\.000\(1\.000, 1\.000\)1\.000\(1\.000, 1\.000\)1\.000\(1\.000, 1\.000\)THYB\_S\_SH0641\.000\(0\.000, 1\.000\)0\.417\(0\.000, 1\.000\)0\.417\(0\.000, 1\.000\)0\.583\(0\.000, 1\.000\)THYB\_S\_FJ0121\.000\(1\.000, 1\.000\)1\.000\(1\.000, 1\.000\)1\.000\(1\.000, 1\.000\)1\.000\(1\.000, 1\.000\)THYB\_S\_FJ0321\.000\(1\.000, 1\.000\)1\.000\(1\.000, 1\.000\)1\.000\(1\.000, 1\.000\)1\.000\(1\.000, 1\.000\)THYB\_S\_NX0111\.000\(1\.000, 1\.000\)1\.000\(1\.000, 1\.000\)1\.000\(1\.000, 1\.000\)1\.000\(1\.000, 1\.000\)THYB\_S\_HB070––––THYB\_S\_SD140––––THYB\_S\_ZJ290––––Table S13:Performance of preprocessing and executor tools in ThyroidXAgent\. Held\-out test results are reported for preprocessing tools, including image normalization, nodule\-presence triage and anatomical\-context parsing, and for executor\-stage tools, including measurement support, gland localization, lymph\-node screening, gland captioning and nodule\-feature extraction\. Anatomical\-context parsing and nodule\-feature extraction are additionally reported at the class level\. For the binary margin and shape classifiers, AUROC and AUPRC are reported once across the paired class rows\. AP, average precision; MAE, mean absolute error; MSE, mean squared error; MAPE, mean absolute percentage error\.Agent stagePreprocessing toolsTool groupToolTrainValTestPrimary resultSecondary resultPreprocessingImage normalizationUltrasound ROI cropping1521779Dice, 0\.9822; IoU, 0\.9658Precision, 0\.9904; recall, 0\.9749; pixel accuracy, 0\.9829Case triageNodule\-presence detection82,31210,98216,467Accuracy, 0\.9830; F1, 0\.9749AUROC, 0\.9981; AP, 0\.9961; sensitivity, 0\.9848; specificity, 0\.9821Anatomical context parsingToolClassTrainValTestTotalPrecisionRecallF1AUROCAUPRCThyroid\-regionclassificationLeft\-lobe lateral view7801101591,0490\.66430\.59750\.62910\.85210\.7185Right\-lobe lateral view9461331991,2780\.64600\.73370\.68710\.83270\.7368Bilateral thyroid view12019301690\.89290\.83330\.86210\.98960\.8911Left\-lobe transverse view26752413600\.70450\.75610\.72940\.95670\.8198Right\-lobe transverse view27261794120\.80880\.69620\.74830\.95530\.8585Neck region26721123001\.00000\.91670\.95650\.99740\.9524Agent stageExecutor\-stage toolsTool groupToolTrainValTestPrimary resultSecondary resultExecutorMeasurement supportSpacing prediction5,288661662MAE, 0\.0131;R2R^\{2\}, 0\.8520MSE,5\.66×10−45\.66\\times 10^\{\-4\}; MAPE, 21\.39%Gland localizationGland segmentation335\-90Dice, 0\.8006; IoU, 0\.6866Precision, 0\.8025; recall, 0\.8339Neck\-region screeningCervical lymph\-node detection251\-49Accuracy, 0\.7959; F1, 0\.7368AUROC, 0\.8163Gland descriptionGland captioning22,782200612BLEU\-4, 0\.5898; METEOR, 0\.4582ROUGEL, 0\.7450; CIDEr, 2\.7736Nodule feature extractionTool familyFeature classifierClassTrainValTestTotalSpecificitySensitivityAUROCAUPRCNodule\-featureclassificationCompositionCystic1,8242272282,2790\.83970\.85530\.91660\.9073Mixed cystic and solid8831081101,1010\.93580\.36360\.82570\.5943Solid1,3911731771,7410\.80180\.79660\.90110\.8033EchogenicityAnechoic1,6912122132,1160\.86030\.89670\.93970\.9160Hyperechoic21827272720\.98470\.25930\.84740\.3591Hypoechoic1,4161731731,7620\.81730\.67050\.83910\.7514Isoechoic58073727250\.92740\.54170\.88040\.5440Echogenic fociMacrocalcifications1,1651451461,4560\.81070\.47260\.72790\.5601None2,1502702702,6900\.51300\.80370\.74310\.7483Punctate echogenic foci66883848350\.94230\.13100\.64110\.2615MarginIll\-defined1,1771471481,4720\.82000\.74320\.87020\.8847Smooth1,6022002002,0020\.74320\.8200ShapeTaller\-than\-wide811091000\.98720\.55560\.91740\.9899Wider\-than\-tall60279787590\.55560\.9872

Table S14:Lexical performance of thyroid ultrasound report generation across datasets\. BLEU\-1 to BLEU\-4, METEOR and ROUGELare reported on the SMU\-HMC, KMVE and ZJH\-TS test sets\. Values are means±\\pmthe half\-width of the bootstrap 95% percentile confidence interval\.ModelBLEU\-1BLEU\-2BLEU\-3BLEU\-4METEORROUGELSMU\-HMC Testset\(n=400n=400\)GPT\-4o\[[95](https://arxiv.org/html/2608.12590#bib.bib95)\]0\.3500±\\pm0\.01000\.2535±\\pm0\.00800\.1842±\\pm0\.00660\.1330±\\pm0\.00570\.3247±\\pm0\.00470\.3577±\\pm0\.0077GPT\-5\[[68](https://arxiv.org/html/2608.12590#bib.bib68)\]0\.3836±\\pm0\.00850\.2749±\\pm0\.00690\.1965±\\pm0\.00590\.1374±\\pm0\.00540\.3254±\\pm0\.00370\.3732±\\pm0\.0068Gemini2\.5 Pro\[[69](https://arxiv.org/html/2608.12590#bib.bib69)\]0\.3702±\\pm0\.00940\.2660±\\pm0\.00810\.1907±\\pm0\.00700\.1373±\\pm0\.00630\.3308±\\pm0\.00380\.3584±\\pm0\.0071Qwen3\.5 Plus\[[96](https://arxiv.org/html/2608.12590#bib.bib96)\]0\.4483±\\pm0\.01300\.3623±\\pm0\.01110\.2959±\\pm0\.00950\.2427±\\pm0\.00810\.3628±\\pm0\.00400\.5147±\\pm0\.0088Claude\-Sonnet\-4\.60\.3326±\\pm0\.00990\.2543±\\pm0\.00800\.1968±\\pm0\.00670\.1528±\\pm0\.00560\.3417±\\pm0\.00350\.3782±\\pm0\.0074MedGemma\[[71](https://arxiv.org/html/2608.12590#bib.bib71)\]0\.0457±\\pm0\.00720\.0343±\\pm0\.00560\.0265±\\pm0\.00440\.0207±\\pm0\.00350\.1736±\\pm0\.00510\.0829±\\pm0\.0075LLaVA\-Med\[[97](https://arxiv.org/html/2608.12590#bib.bib97)\]0\.1670±\\pm0\.00770\.0581±\\pm0\.00660\.0290±\\pm0\.00400\.0159±\\pm0\.00240\.1439±\\pm0\.00530\.1482±\\pm0\.0073KMVE\[[18](https://arxiv.org/html/2608.12590#bib.bib18)\]0\.1743±\\pm0\.01310\.1110±\\pm0\.00820\.0684±\\pm0\.00520\.0398±\\pm0\.00350\.1719±\\pm0\.00630\.2212±\\pm0\.0039ThyroidXAgent0\.5924±\\pm0\.01430\.4806±\\pm0\.01390\.4006±\\pm0\.01370\.3381±\\pm0\.01350\.3627±\\pm0\.00880\.5422±\\pm0\.0122KMVE Testset\(n=492n=492\)GPT\-4o\[[95](https://arxiv.org/html/2608.12590#bib.bib95)\]0\.4467±\\pm0\.01540\.3423±\\pm0\.01430\.2683±\\pm0\.01180\.2136±\\pm0\.01020\.2799±\\pm0\.01160\.3850±\\pm0\.0148GPT\-5\[[68](https://arxiv.org/html/2608.12590#bib.bib68)\]0\.4149±\\pm0\.00730\.2992±\\pm0\.00660\.2190±\\pm0\.00590\.1668±\\pm0\.00620\.3042±\\pm0\.00700\.4128±\\pm0\.0089Gemini2\.5 Pro\[[69](https://arxiv.org/html/2608.12590#bib.bib69)\]0\.5065±\\pm0\.01170\.3851±\\pm0\.01030\.2981±\\pm0\.00950\.2394±\\pm0\.00920\.2863±\\pm0\.00830\.4742±\\pm0\.0111Qwen3\.5 Plus\[[96](https://arxiv.org/html/2608.12590#bib.bib96)\]0\.5425±\\pm0\.01350\.4243±\\pm0\.01260\.3415±\\pm0\.01210\.2783±\\pm0\.01200\.2952±\\pm0\.00840\.5066±\\pm0\.0118Claude\-Sonnet\-4\.60\.4331±\\pm0\.01190\.3380±\\pm0\.01120\.2638±\\pm0\.01070\.2089±\\pm0\.01060\.3284±\\pm0\.00790\.4639±\\pm0\.0115MedGemma\[[71](https://arxiv.org/html/2608.12590#bib.bib71)\]0\.0374±\\pm0\.01950\.0299±\\pm0\.01750\.0252±\\pm0\.01610\.0219±\\pm0\.01510\.1075±\\pm0\.01190\.1862±\\pm0\.0205LLaVA\-Med\[[97](https://arxiv.org/html/2608.12590#bib.bib97)\]0\.2842±\\pm0\.01490\.2135±\\pm0\.01240\.1595±\\pm0\.00990\.1244±\\pm0\.00880\.2138±\\pm0\.01010\.3939±\\pm0\.0127ThyroidXAgent0\.6357±\\pm0\.01710\.5606±\\pm0\.01570\.5008±\\pm0\.01510\.4535±\\pm0\.01510\.3672±\\pm0\.00990\.5880±\\pm0\.0120ZJH\-TS Testset\(n=150n=150\)GPT\-4o\[[95](https://arxiv.org/html/2608.12590#bib.bib95)\]0\.3402±\\pm0\.02730\.2447±\\pm0\.02060\.1788±\\pm0\.01550\.1332±\\pm0\.01190\.2495±\\pm0\.01630\.3679±\\pm0\.0219GPT\-5\[[68](https://arxiv.org/html/2608.12590#bib.bib68)\]0\.4736±\\pm0\.01580\.3499±\\pm0\.01300\.2539±\\pm0\.01080\.1789±\\pm0\.00950\.3332±\\pm0\.00600\.4790±\\pm0\.0098Gemini2\.5 Pro\[[69](https://arxiv.org/html/2608.12590#bib.bib69)\]0\.1932±\\pm0\.01120\.1307±\\pm0\.00800\.0887±\\pm0\.00580\.0601±\\pm0\.00450\.2765±\\pm0\.00520\.2536±\\pm0\.0091Qwen3\.5 Plus\[[96](https://arxiv.org/html/2608.12590#bib.bib96)\]0\.4964±\\pm0\.01820\.3965±\\pm0\.01620\.3196±\\pm0\.01450\.2581±\\pm0\.01310\.3508±\\pm0\.00760\.5314±\\pm0\.0125Claude\-Sonnet\-4\.60\.4480±\\pm0\.01810\.3518±\\pm0\.01570\.2792±\\pm0\.01350\.2229±\\pm0\.01180\.3455±\\pm0\.00730\.4864±\\pm0\.0127MedGemma\[[71](https://arxiv.org/html/2608.12590#bib.bib71)\]0\.1236±\\pm0\.02200\.0958±\\pm0\.01740\.0750±\\pm0\.01380\.0591±\\pm0\.01100\.2268±\\pm0\.00940\.1592±\\pm0\.0211LLaVA\-Med\[[97](https://arxiv.org/html/2608.12590#bib.bib97)\]0\.2318±\\pm0\.01940\.1316±\\pm0\.01410\.0891±\\pm0\.01010\.0606±\\pm0\.00740\.1658±\\pm0\.01060\.2479±\\pm0\.0201KMVE\[[18](https://arxiv.org/html/2608.12590#bib.bib18)\]0\.1682±\\pm0\.01460\.1041±\\pm0\.00900\.0598±\\pm0\.00540\.0289±\\pm0\.00380\.1648±\\pm0\.00690\.2244±\\pm0\.0038ThyroidXAgent0\.5051±\\pm0\.02280\.4137±\\pm0\.02060\.3447±\\pm0\.01880\.2909±\\pm0\.01740\.3293±\\pm0\.01190\.5480±\\pm0\.0160Table S15:Clinical semantic performance of thyroid ultrasound report generation across datasets\. False discovery rate \(FDR\), feature accuracy, lesion\-level F1 score, completeness, consistency and ThyClinScore are reported on the SMU\-HMC, KMVE and ZJH\-TS test sets\. Values are means±\\pmthe half\-width of the bootstrap 95% percentile confidence interval\. Lower FDR indicates better performance\.ModelFDR↓\\downarrowFeat AccF1 ScoreComplete\.Consist\.ThyClinSMU\-HMC Testset\(n=400n=400\)GPT\-4o\[[95](https://arxiv.org/html/2608.12590#bib.bib95)\]0\.7189±\\pm0\.03620\.5555±\\pm0\.03510\.2526±\\pm0\.03320\.8208±\\pm0\.01320\.4023±\\pm0\.01820\.4105±\\pm0\.0172GPT\-5\[[68](https://arxiv.org/html/2608.12590#bib.bib68)\]0\.4170±\\pm0\.04170\.5644±\\pm0\.03510\.4390±\\pm0\.04020\.8585±\\pm0\.00770\.4980±\\pm0\.01660\.4882±\\pm0\.0182Gemini2\.5 Pro\[[69](https://arxiv.org/html/2608.12590#bib.bib69)\]0\.6797±\\pm0\.03880\.5586±\\pm0\.03850\.2965±\\pm0\.03660\.8789±\\pm0\.01070\.4040±\\pm0\.01870\.4280±\\pm0\.0176Qwen3\.5 Plus\[[96](https://arxiv.org/html/2608.12590#bib.bib96)\]0\.7426±\\pm0\.02950\.5804±\\pm0\.03730\.2656±\\pm0\.02900\.9620±\\pm0\.00630\.4809±\\pm0\.01420\.4883±\\pm0\.0137Claude\-Sonnet\-4\.60\.8040±\\pm0\.03050\.5580±\\pm0\.04640\.2049±\\pm0\.02980\.8786±\\pm0\.01080\.3886±\\pm0\.01470\.4015±\\pm0\.0144MedGemma\[[71](https://arxiv.org/html/2608.12590#bib.bib71)\]0\.8943±\\pm0\.02290\.5784±\\pm0\.04340\.1027±\\pm0\.02020\.9657±\\pm0\.00550\.3353±\\pm0\.01190\.3697±\\pm0\.0131LLaVA\-Med\[[97](https://arxiv.org/html/2608.12590#bib.bib97)\]0\.1625±\\pm0\.03630\.4100±\\pm0\.21250\.2582±\\pm0\.04260\.4576±\\pm0\.02270\.1387±\\pm0\.01660\.1752±\\pm0\.0150KMVE\[[18](https://arxiv.org/html/2608.12590#bib.bib18)\]0\.4813±\\pm0\.04880\.5944±\\pm0\.12330\.2103±\\pm0\.03900\.5995±\\pm0\.00850\.2013±\\pm0\.01440\.2265±\\pm0\.0145ThyroidXAgent0\.2238±\\pm0\.03770\.6238±\\pm0\.03580\.5467±\\pm0\.04360\.9691±\\pm0\.00600\.5016±\\pm0\.01660\.5189±\\pm0\.0212KMVE Testset\(n=492n=492\)GPT\-4o\[[95](https://arxiv.org/html/2608.12590#bib.bib95)\]0\.5843±\\pm0\.04270\.6608±\\pm0\.06910\.2137±\\pm0\.03500\.6180±\\pm0\.01130\.4275±\\pm0\.02310\.3344±\\pm0\.0188GPT\-5\[[68](https://arxiv.org/html/2608.12590#bib.bib68)\]0\.4133±\\pm0\.04230\.6967±\\pm0\.04920\.3749±\\pm0\.04020\.6506±\\pm0\.00380\.4504±\\pm0\.01890\.3853±\\pm0\.0189Gemini2\.5 Pro\[[69](https://arxiv.org/html/2608.12590#bib.bib69)\]0\.3211±\\pm0\.04070\.5093±\\pm0\.04180\.4648±\\pm0\.04010\.6270±\\pm0\.00450\.4135±\\pm0\.02130\.3810±\\pm0\.0187Qwen3\.5 Plus\[[96](https://arxiv.org/html/2608.12590#bib.bib96)\]0\.4553±\\pm0\.04220\.6154±\\pm0\.04480\.3823±\\pm0\.03970\.6525±\\pm0\.00500\.4035±\\pm0\.02300\.3674±\\pm0\.0193Claude\-Sonnet\-4\.60\.5803±\\pm0\.04320\.6605±\\pm0\.05780\.2721±\\pm0\.03750\.6707±\\pm0\.00430\.3779±\\pm0\.02110\.3362±\\pm0\.0182MedGemma\[[71](https://arxiv.org/html/2608.12590#bib.bib71)\]0\.1877±\\pm0\.03270\.6820±\\pm0\.04630\.2868±\\pm0\.03810\.6622±\\pm0\.00330\.5172±\\pm0\.02400\.3932±\\pm0\.0203LLaVA\-Med\[[97](https://arxiv.org/html/2608.12590#bib.bib97)\]0\.0000±\\pm0\.0000\*N/A\*0\.2541±\\pm0\.03860\.6508±\\pm0\.00070\.5669±\\pm0\.02250\.3928±\\pm0\.0208ThyroidXAgent0\.4858±\\pm0\.03860\.7366±\\pm0\.03550\.3623±\\pm0\.03580\.6416±\\pm0\.00290\.5654±\\pm0\.02170\.4407±\\pm0\.0194ZJH\-TS Testset\(n=150n=150\)GPT\-4o\[[95](https://arxiv.org/html/2608.12590#bib.bib95)\]0\.4583±\\pm0\.07060\.5510±\\pm0\.04680\.2706±\\pm0\.05230\.7290±\\pm0\.03050\.3415±\\pm0\.02940\.3139±\\pm0\.0260GPT\-5\[[68](https://arxiv.org/html/2608.12590#bib.bib68)\]0\.5201±\\pm0\.06580\.5263±\\pm0\.04870\.3771±\\pm0\.05310\.9534±\\pm0\.01120\.4735±\\pm0\.02430\.4472±\\pm0\.0248Gemini2\.5 Pro\[[69](https://arxiv.org/html/2608.12590#bib.bib69)\]0\.5450±\\pm0\.06890\.5507±\\pm0\.05470\.3157±\\pm0\.05230\.7859±\\pm0\.01830\.3910±\\pm0\.02920\.3557±\\pm0\.0244Qwen3\.5 Plus\[[96](https://arxiv.org/html/2608.12590#bib.bib96)\]0\.5781±\\pm0\.05580\.5464±\\pm0\.04650\.3902±\\pm0\.04850\.9575±\\pm0\.00860\.4767±\\pm0\.02190\.4520±\\pm0\.0220Claude\-Sonnet\-4\.60\.6186±\\pm0\.05710\.5748±\\pm0\.04550\.3405±\\pm0\.04960\.8226±\\pm0\.01760\.4298±\\pm0\.02330\.3861±\\pm0\.0219MedGemma\[[71](https://arxiv.org/html/2608.12590#bib.bib71)\]0\.6537±\\pm0\.06440\.5463±\\pm0\.04980\.2959±\\pm0\.05200\.9437±\\pm0\.01060\.3868±\\pm0\.02250\.3807±\\pm0\.0238LLaVA\-Med\[[97](https://arxiv.org/html/2608.12590#bib.bib97)\]0\.2733±\\pm0\.07330\.6595±\\pm0\.10770\.1241±\\pm0\.04540\.6136±\\pm0\.04940\.1505±\\pm0\.03080\.1812±\\pm0\.0275KMVE\[[18](https://arxiv.org/html/2608.12590#bib.bib18)\]0\.7911±\\pm0\.06170\.6288±\\pm0\.14900\.0445±\\pm0\.02530\.6252±\\pm0\.00850\.1581±\\pm0\.01620\.1650±\\pm0\.0124ThyroidXAgent0\.2833±\\pm0\.06000\.5820±\\pm0\.04080\.4889±\\pm0\.05700\.9641±\\pm0\.00900\.4564±\\pm0\.02280\.4676±\\pm0\.0268
- \*On the KMVE dataset, LLaVA\-Med collapsed and predicted “no abnormality” for all test samples\. Therefore, Feat Acc is N/A and FDR is 0 because no positive predictions were made\.

Table S16:Static\-pipeline metrics used for report\-generation radar plots\. Conventional language\-generation metrics and clinical semantic metrics are reported for the SMU\-HMC, KMVE and ZJH\-TS test sets used in Fig\.[5](https://arxiv.org/html/2608.12590#S2.F5)g\. Values are means±\\pmthe half\-width of the 95% confidence interval\. Lower FDR indicates better performance\.Conventional natural\-language generation metricsDatasetnnBLEU\-1BLEU\-2BLEU\-3BLEU\-4METEORROUGELSMU\-HMC4000\.4586±\\pm0\.01490\.3817±\\pm0\.01390\.3271±\\pm0\.01360\.2849±\\pm0\.01330\.3209±\\pm0\.00730\.4725±\\pm0\.0115KMVE4920\.3266±\\pm0\.00690\.2105±\\pm0\.00480\.1306±\\pm0\.00340\.0624±\\pm0\.00340\.2724±\\pm0\.00510\.2750±\\pm0\.0045ZJH\-TS1500\.4242±\\pm0\.02340\.3527±\\pm0\.02060\.2986±\\pm0\.01860\.2534±\\pm0\.01720\.2916±\\pm0\.01010\.4994±\\pm0\.0140
Clinical semantic metricsDatasetnnFDR↓\\downarrowFeat AccF1 ScoreComplete\.Consist\.ThyClinSMU\-HMC4000\.3162±\\pm0\.04310\.6227±\\pm0\.03820\.5060±\\pm0\.04530\.7821±\\pm0\.00640\.4325±\\pm0\.01610\.4293±\\pm0\.0192KMVE4920\.5894±\\pm0\.04320\.5616±\\pm0\.06310\.2644±\\pm0\.03750\.7665±\\pm0\.00590\.3391±\\pm0\.01810\.3346±\\pm0\.0174ZJH\-TS1500\.4433±\\pm0\.07170\.5517±\\pm0\.05160\.3894±\\pm0\.06030\.8127±\\pm0\.01200\.3734±\\pm0\.02010\.3648±\\pm0\.0239

Table S17:Ablation of tool integration for thyroid ultrasound report generation\. Segmentation, classification, captioning and measurement tools were added cumulatively, and performance was evaluated using conventional language\-generation metrics\. Values are means±\\pmthe half\-width of the 95% confidence interval\.ConfigurationBLEU\-1BLEU\-2BLEU\-3BLEU\-4METEORROUGELSMU\-HMC Testset\(n=400n=400\)Segmentation only0\.0897±\\pm0\.00950\.0692±\\pm0\.00750\.0562±\\pm0\.00630\.0459±\\pm0\.00530\.1521±\\pm0\.00470\.2603±\\pm0\.0092\- Classification0\.1877±\\pm0\.01690\.1431±\\pm0\.01300\.1148±\\pm0\.01040\.0929±\\pm0\.00850\.1849±\\pm0\.00720\.2870±\\pm0\.0106\- Captioning0\.4503±\\pm0\.01630\.3724±\\pm0\.01500\.3179±\\pm0\.01420\.2764±\\pm0\.01390\.3052±\\pm0\.00810\.4601±\\pm0\.0126\- Measurement \(full\)0\.5924±\\pm0\.01430\.4806±\\pm0\.01390\.4006±\\pm0\.01370\.3381±\\pm0\.01350\.3627±\\pm0\.00880\.5422±\\pm0\.0122KMVE Testset\(n=492n=492\)Segmentation only0\.6130±\\pm0\.01910\.5456±\\pm0\.01860\.4907±\\pm0\.01820\.4462±\\pm0\.01860\.3658±\\pm0\.01030\.5747±\\pm0\.0126\- Classification0\.6315±\\pm0\.01890\.5637±\\pm0\.01840\.5081±\\pm0\.01800\.4633±\\pm0\.01830\.3751±\\pm0\.01020\.5799±\\pm0\.0127\- Captioning0\.6357±\\pm0\.01720\.5606±\\pm0\.01600\.5008±\\pm0\.01520\.4535±\\pm0\.01530\.3672±\\pm0\.00990\.5880±\\pm0\.0122\- Measurement \(full\)0\.6357±\\pm0\.01720\.5606±\\pm0\.01600\.5008±\\pm0\.01520\.4535±\\pm0\.01530\.3672±\\pm0\.00990\.5880±\\pm0\.0122ZJH\-TS Testset\(n=150n=150\)Segmentation only0\.0964±\\pm0\.01500\.0799±\\pm0\.01230\.0674±\\pm0\.01050\.0567±\\pm0\.00910\.1500±\\pm0\.00730\.3348±\\pm0\.0129\- Classification0\.2456±\\pm0\.02460\.1990±\\pm0\.01970\.1648±\\pm0\.01630\.1356±\\pm0\.01370\.2053±\\pm0\.01050\.4003±\\pm0\.0142\- Captioning0\.4092±\\pm0\.02520\.3409±\\pm0\.02200\.2891±\\pm0\.01970\.2461±\\pm0\.01800\.2848±\\pm0\.01150\.4903±\\pm0\.0162\- Measurement \(full\)0\.5051±\\pm0\.02270\.4137±\\pm0\.02050\.3447±\\pm0\.01880\.2909±\\pm0\.01740\.3293±\\pm0\.01180\.5480±\\pm0\.0158
- •The KMVE dataset retains only the findings section and does not provide original measurement values\. To match the original evaluation protocol, only the generated findings section was evaluated and measurement values were masked; therefore, the captioning and full configurations have identical KMVE scores\.

## References

- \[1\]Alexander, E\. K\. & Cibas, E\. S\.Diagnosis of thyroid nodules\.*The Lancet Diabetes & Endocrinology*10, 533–539, DOI:[10\.1016/S2213\-8587\(22\)00101\-2](https://doi.org/10.1016/S2213-8587(22)00101-2)\(2022\)\.
- \[2\]Grani, G\., Sponziello, M\., Filetti, S\. & Durante, C\.Thyroid nodules: diagnosis and management\.*Nature Reviews Endocrinology*DOI:[10\.1038/s41574\-024\-01025\-4](https://doi.org/10.1038/s41574-024-01025-4)\(2024\)\.
- \[3\]Tessler, F\. N\.*et al\.*Acr thyroid imaging, reporting and data system \(ti\-rads\): White paper of the acr ti\-rads committee\.*Journal of the American College of Radiology*14, 587–595, DOI:[10\.1016/j\.jacr\.2017\.01\.046](https://doi.org/10.1016/j.jacr.2017.01.046)\(2017\)\.
- \[4\]Hoang, J\. K\.*et al\.*Interobserver variability of sonographic features used in the american college of radiology thyroid imaging reporting and data system\.*American Journal of Roentgenology*211, 162–167, DOI:[10\.2214/AJR\.17\.19192](https://doi.org/10.2214/AJR.17.19192)\(2018\)\.
- \[5\]Cibas, E\. S\. & Ali, S\. Z\.The 2017 bethesda system for reporting thyroid cytopathology\.*Thyroid*27, 1341–1346, DOI:[10\.1089/thy\.2017\.0500](https://doi.org/10.1089/thy.2017.0500)\(2017\)\.
- \[6\]Topol, E\. J\.High\-performance medicine: the convergence of human and artificial intelligence\.*Nature Medicine*25, 44–56, DOI:[10\.1038/s41591\-018\-0300\-7](https://doi.org/10.1038/s41591-018-0300-7)\(2019\)\.
- \[7\]Rajpurkar, P\., Chen, E\., Banerjee, O\. & Topol, E\. J\.Ai in health and medicine\.*Nature Medicine*28, 31–38, DOI:[10\.1038/s41591\-021\-01614\-0](https://doi.org/10.1038/s41591-021-01614-0)\(2022\)\.
- \[8\]Chen, H\., Gomez, C\., Huang, C\.\-M\.*et al\.*Explainable medical imaging AI needs human\-centered design: Guidelines and evidence from a systematic review\.*npj Digital Medicine*5, 156, DOI:[10\.1038/s41746\-022\-00699\-2](https://doi.org/10.1038/s41746-022-00699-2)\(2022\)\.
- \[9\]Wekenborg, M\. K\., Gilbert, S\. & Kather, J\. N\.Examining human–AI interaction in real\-world healthcare beyond the laboratory\.*npj Digital Medicine*8, 169, DOI:[10\.1038/s41746\-025\-01559\-5](https://doi.org/10.1038/s41746-025-01559-5)\(2025\)\.
- \[10\]Giddings, R\.*et al\.*Factors influencing clinician and patient interaction with machine learning\-based risk prediction models: A systematic review\.*The Lancet Digital Health*6, e131–e144, DOI:[10\.1016/S2589\-7500\(23\)00241\-8](https://doi.org/10.1016/S2589-7500(23)00241-8)\(2024\)\.
- \[11\]Gong, H\.*et al\.*Multi\-task learning for thyroid nodule segmentation with thyroid region prior\.In*2021 IEEE 18th international symposium on biomedical imaging \(ISBI\)*, 257–261 \(2021\)\.
- \[12\]Gong, H\.*et al\.*Thyroid region prior guided attention for ultrasound segmentation of thyroid nodules\.*Computers in biology and medicine*155, 106389 \(2023\)\.
- \[13\]Sun, X\., Wei, B\., Jiang, Y\., Mao, L\. & Zhao, Q\.Clip\-tnseg: A multi\-modal hybrid framework for thyroid nodule segmentation in ultrasound images\.*arXiv preprint arXiv:2412\.05530*\(2024\)\.
- \[14\]Gong, H\.*et al\.*Less is more: adaptive curriculum learning for thyroid nodule diagnosis\.In*International Conference on Medical Image Computing and Computer\-Assisted Intervention*, 248–257 \(2022\)\.
- \[15\]Peng, S\.*et al\.*Deep learning\-based artificial intelligence model to assist thyroid nodule diagnosis and management: a multicentre diagnostic study\.*The Lancet Digital Health*3, e250–e259, DOI:[10\.1016/S2589\-7500\(21\)00041\-8](https://doi.org/10.1016/S2589-7500(21)00041-8)\(2021\)\.
- \[16\]Chen, Y\.*et al\.*An artificial intelligence model based on acr ti\-rads characteristics for us diagnosis of thyroid nodules\.*Radiology*303, 613–619, DOI:[10\.1148/radiol\.211455](https://doi.org/10.1148/radiol.211455)\(2022\)\.
- \[17\]Yao, J\.*et al\.*Multimodal gpt model for assisting thyroid nodule diagnosis and management\.*npj Digital Medicine*8, 245, DOI:[10\.1038/s41746\-025\-01652\-9](https://doi.org/10.1038/s41746-025-01652-9)\(2025\)\.
- \[18\]Li, J\., Su, T\.*et al\.*Ultrasound report generation with cross\-modality feature alignment via unsupervised guidance\.*IEEE Transactions on Medical Imaging*44, 19–30 \(2024\)\.
- \[19\]Tanno, R\., Barrett, D\. G\. T\., Sellergren, A\.*et al\.*Collaboration between clinicians and vision–language models in radiology report generation\.*Nature Medicine*31, 599–608, DOI:[10\.1038/s41591\-024\-03302\-1](https://doi.org/10.1038/s41591-024-03302-1)\(2025\)\.
- \[20\]Li, C\.\-Y\., Chang, K\.\-J\.*et al\.*Towards a holistic framework for multimodal llm in 3d brain ct radiology report generation\.*Nature Communications*16, 2258 \(2025\)\.
- \[21\]Wang, J\.*et al\.*Deep learning models for thyroid nodules diagnosis of fine\-needle aspiration biopsy: a retrospective, prospective, multicentre study in china\.*The Lancet Digital Health*6, e458–e469, DOI:[10\.1016/S2589\-7500\(24\)00085\-2](https://doi.org/10.1016/S2589-7500(24)00085-2)\(2024\)\.
- \[22\]Shen, P\.*et al\.*Explainable multimodal deep learning for predicting thyroid cancer lateral lymph node metastasis using ultrasound imaging\.*Nature Communications*16, 7052, DOI:[10\.1038/s41467\-025\-62042\-z](https://doi.org/10.1038/s41467-025-62042-z)\(2025\)\.
- \[23\]Dai, F\.*et al\.*Improving ai models for rare thyroid cancer subtype by text guided diffusion models\.*Nature Communications*16, 4449, DOI:[10\.1038/s41467\-025\-59478\-8](https://doi.org/10.1038/s41467-025-59478-8)\(2025\)\.
- \[24\]Tikhomirov, L\.*et al\.*Medical artificial intelligence for clinicians: The lost cognitive perspective\.*The Lancet Digital Health*6, e589–e594, DOI:[10\.1016/S2589\-7500\(24\)00095\-5](https://doi.org/10.1016/S2589-7500(24)00095-5)\(2024\)\.
- \[25\]You, G\., Li, H\., Zhang, Y\. & Fan, Y\.Learning anatomy\-grounded CT vision\-language representations with organ\-hierarchical report knowledge\.*arXiv preprint arXiv:2607\.10953*DOI:[10\.48550/arXiv\.2607\.10953](https://doi.org/10.48550/arXiv.2607.10953)\(2026\)\.
- \[26\]Vasey, B\.*et al\.*Reporting guideline for the early\-stage clinical evaluation of decision support systems driven by artificial intelligence: Decide\-ai\.*Nature Medicine*28, 924–933, DOI:[10\.1038/s41591\-022\-01772\-9](https://doi.org/10.1038/s41591-022-01772-9)\(2022\)\.
- \[27\]Dreyer, M\.*et al\.*Mechanistic understanding and validation of large ai models with semanticlens\.*Nature Machine Intelligence*7, 1572–1585 \(2025\)\.
- \[28\]Patel, B\. N\., Rosenberg, L\., Willcox, G\.*et al\.*Human–machine partnership with artificial intelligence for chest radiograph diagnosis\.*npj Digital Medicine*2, 111, DOI:[10\.1038/s41746\-019\-0189\-7](https://doi.org/10.1038/s41746-019-0189-7)\(2019\)\.
- \[29\]Leibig, C\.*et al\.*Combining the strengths of radiologists and AI for breast cancer screening: A retrospective analysis\.*The Lancet Digital Health*4, e507–e519, DOI:[10\.1016/S2589\-7500\(22\)00070\-X](https://doi.org/10.1016/S2589-7500(22)00070-X)\(2022\)\.
- \[30\]Yu, F\.*et al\.*Heterogeneity and predictors of the effects of AI assistance on radiologists\.*Nature Medicine*30, 837–849, DOI:[10\.1038/s41591\-024\-02850\-w](https://doi.org/10.1038/s41591-024-02850-w)\(2024\)\.
- \[31\]Chen, M\., Wang, Y\., Wang, Q\.*et al\.*Impact of human and artificial intelligence collaboration on workload reduction in medical image interpretation\.*npj Digital Medicine*7, 349, DOI:[10\.1038/s41746\-024\-01328\-w](https://doi.org/10.1038/s41746-024-01328-w)\(2024\)\.
- \[32\]Everett, S\. S\., Bunning, B\. J\., Jain, P\.*et al\.*From tool to teammate in a randomized controlled trial of clinician–AI collaborative workflows for diagnosis\.*npj Digital Medicine*9, 409, DOI:[10\.1038/s41746\-026\-02545\-1](https://doi.org/10.1038/s41746-026-02545-1)\(2026\)\.
- \[33\]Strong, J\., Rogers, H\., Sun, E\.*et al\.*Human–AI collaboration in healthcare: A scoping review\.*npj Digital Medicine*DOI:[10\.1038/s41746\-026\-02918\-6](https://doi.org/10.1038/s41746-026-02918-6)\(2026\)\.
- \[34\]Wiens, J\.*et al\.*Do no harm: a roadmap for responsible machine learning for health care\.*Nature Medicine*25, 1337–1340, DOI:[10\.1038/s41591\-019\-0548\-6](https://doi.org/10.1038/s41591-019-0548-6)\(2019\)\.
- \[35\]Zou, J\. & Topol, E\. J\.The rise of agentic ai teammates in medicine\.*The Lancet*405, 457, DOI:[10\.1016/S0140\-6736\(25\)00202\-8](https://doi.org/10.1016/S0140-6736(25)00202-8)\(2025\)\.
- \[36\]Moor, M\.*et al\.*Foundation models for generalist medical artificial intelligence\.*Nature*616, 259–265, DOI:[10\.1038/s41586\-023\-05881\-4](https://doi.org/10.1038/s41586-023-05881-4)\(2023\)\.
- \[37\]Kohane, I\. S\.Injecting artificial intelligence into medicine\.*NEJM AI*1, 1–3, DOI:[10\.1056/AIe2300197](https://doi.org/10.1056/AIe2300197)\(2024\)\.
- \[38\]Katz, U\., Cohen, E\., Shachar, E\.*et al\.*GPT versus resident physicians—a benchmark based on official board scores\.*NEJM AI*1, DOI:[10\.1056/AIdbp2300192](https://doi.org/10.1056/AIdbp2300192)\(2024\)\.
- \[39\]Zhou, H\.\-Y\.*et al\.*Medversa: A generalist foundation model for diverse medical imaging tasks\.*NEJM AI*3, DOI:[10\.1056/AIoa2500595](https://doi.org/10.1056/AIoa2500595)\(2026\)\.
- \[40\]Tu, T\., Schaekermann, M\., Palepu, A\.*et al\.*Towards conversational diagnostic artificial intelligence\.*Nature*642, 442–450, DOI:[10\.1038/s41586\-025\-08866\-7](https://doi.org/10.1038/s41586-025-08866-7)\(2025\)\.
- \[41\]McDuff, D\., Schaekermann, M\., Tu, T\.*et al\.*Towards accurate differential diagnosis with large language models\.*Nature*642, 451–457, DOI:[10\.1038/s41586\-025\-08869\-4](https://doi.org/10.1038/s41586-025-08869-4)\(2025\)\.
- \[42\]Deltadahl, S\.*et al\.*Deep generative classification of blood cell morphology\.*Nature Machine Intelligence*7, 1791–1803 \(2025\)\.
- \[43\]Pontikos, N\.*et al\.*Next\-generation phenotyping of inherited retinal diseases from multimodal imaging with eye2gene\.*Nature Machine Intelligence*7, 967–978 \(2025\)\.
- \[44\]Qiu, J\.*et al\.*Llm\-based agentic systems in medicine and healthcare\.*Nature Machine Intelligence*6, 1418–1420, DOI:[10\.1038/s42256\-024\-00944\-1](https://doi.org/10.1038/s42256-024-00944-1)\(2024\)\.
- \[45\]Moritz, M\., Topol, E\. & Rajpurkar, P\.Coordinated ai agents for advancing healthcare\.*Nature Biomedical Engineering*9, 432–438, DOI:[10\.1038/s41551\-025\-01363\-2](https://doi.org/10.1038/s41551-025-01363-2)\(2025\)\.
- \[46\]Ferber, D\., Hilgers, L\., H”oper, C\.*et al\.*Towards autonomous medical artificial intelligence agents\.*Nature*DOI:[10\.1038/s41586\-026\-10675\-5](https://doi.org/10.1038/s41586-026-10675-5)\(2026\)\.
- \[47\]Collaco, B\. G\., Haider, S\. A\., Prabha, S\.*et al\.*The role of agentic artificial intelligence in healthcare: A scoping review\.*npj Digital Medicine*9, 345, DOI:[10\.1038/s41746\-026\-02517\-5](https://doi.org/10.1038/s41746-026-02517-5)\(2026\)\.
- \[48\]Kong, Q\.*et al\.*Ai agent\-based discovery of d\-enantiomeric antimicrobial peptides against multidrug\-resistant bacterial infection\.*Biomaterials*123927 \(2025\)\.
- \[49\]Wang, L\., Xu, W\.*et al\.*Plan\-and\-solve prompting: Improving zero\-shot chain\-of\-thought reasoning by large language models\.In*Proceedings of the 61st annual meeting of the association for computational linguistics \(volume 1: long papers\)*, 2609–2634 \(2023\)\.
- \[50\]Schmidgall, S\., Ziaei, R\., Harris, C\.*et al\.*AgentClinic: A multimodal benchmark for tool\-using clinical AI agents\.*npj Digital Medicine*9, 499, DOI:[10\.1038/s41746\-026\-02674\-7](https://doi.org/10.1038/s41746-026-02674-7)\(2026\)\.
- \[51\]Liu, Y\., Carrero, Z\. I\., Jiang, X\.*et al\.*Benchmarking large language model\-based agent systems for clinical decision tasks\.*npj Digital Medicine*9, 259, DOI:[10\.1038/s41746\-026\-02443\-6](https://doi.org/10.1038/s41746-026-02443-6)\(2026\)\.
- \[52\]Yao, S\., Zhao, J\.*et al\.*React: Synergizing reasoning and acting in language models\.In*The eleventh international conference on learning representations*\(2022\)\.
- \[53\]Tian, J\., Fard, P\., Cagan, C\.*et al\.*An autonomous agentic workflow for clinical detection of cognitive concerns using large language models\.*npj Digital Medicine*9, 51, DOI:[10\.1038/s41746\-025\-02324\-4](https://doi.org/10.1038/s41746-025-02324-4)\(2026\)\.
- \[54\]Jiang, Y\.*et al\.*Medagentbench: A virtual ehr environment to benchmark medical llm agents\.*NEJM AI*AIdbp2500144, DOI:[10\.1056/AIdbp2500144](https://doi.org/10.1056/AIdbp2500144)\(2025\)\.
- \[55\]Zhang, H\., Liu, Q\., Han, X\.*et al\.*Tn5000: An ultrasound image dataset for thyroid nodule detection and classification\.*Scientific Data*12, 1437, DOI:[10\.1038/s41597\-025\-05757\-4](https://doi.org/10.1038/s41597-025-05757-4)\(2025\)\.
- \[56\]Van Griethuysen, J\. J\.*et al\.*Computational radiomics system to decode the radiographic phenotype\.*Cancer research*77, e104–e107 \(2017\)\.
- \[57\]Rebuffel, C\., Soulier, L\.*et al\.*A hierarchical model for data\-to\-text generation\.In*European Conference on Information Retrieval*, 65–80 \(Springer, 2020\)\.
- \[58\]Farquhar, S\., Kossen, J\., Kuhn, L\. & Gal, Y\.Detecting hallucinations in large language models using semantic entropy\.*Nature*630, 625–630 \(2024\)\.
- \[59\]Duong, V\. H\.*et al\.*Thyroidxl: Advancing thyroid nodule diagnosis with an expert\-labeled, pathology\-validated dataset\.In*Medical Image Computing and Computer Assisted Intervention – MICCAI 2025*, vol\. 15974, 616–626, DOI:[10\.1007/978\-3\-032\-05182\-0˙60](https://doi.org/10.1007/978-3-032-05182-0_60)\(Springer Nature Switzerland, 2025\)\.
- \[60\]Pedraza, L\.*et al\.*An open access thyroid ultrasound image database\.In*10th International Symposium on Medical Information Processing and Analysis*, vol\. 9287, 188–193 \(2015\)\.
- \[61\]Shusharina, N\., Heinrich, M\. P\. & Huang, R\.*Segmentation, Classification, and Registration of Multi\-modality Medical Imaging Data: MICCAI 2020 Challenges, ABCs 2020, L2R 2020, TN\-SCUI 2020, Held in Conjunction with MICCAI 2020, Lima, Peru, October 4–8, 2020, Proceedings*\(Springer Nature, 2021\)\.
- \[62\]Dai, F\.*et al\.*Improving ai models for rare thyroid cancer subtype by text guided diffusion models\.*Nature Communications*16, 4449 \(2025\)\.
- \[63\]Liu, Z\. & He, K\.A decade’s battle on dataset bias: Are we there yet?In*International Conference on Learning Representations*\(2025\)\.
- \[64\]Ma, J\.*et al\.*Medsam2: Segment anything in 3d medical images and videos\.*arXiv preprint arXiv:2504\.03600*\(2025\)\.
- \[65\]Jiang, Y\.*et al\.*From pretraining to privacy: Federated ultrasound foundation model with self\-supervised learning\.*npj Digital Medicine*8, 714, DOI:[10\.1038/s41746\-025\-02085\-0](https://doi.org/10.1038/s41746-025-02085-0)\(2025\)\.
- \[66\]Wang, A\., Chen, H\., Lin, Z\., Han, J\. & Ding, G\.Repvit: Revisiting mobile cnn from vit perspective\.In*Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition*, 15909–15920 \(2024\)\.
- \[67\]Dong, C\.*et al\.*A survey of natural language generation\.*ACM Computing Surveys*55, 1–38 \(2022\)\.
- \[68\]OpenAI\.Gpt\-5 system card\.[https://openai\.com/index/gpt\-5\-system\-card/](https://openai.com/index/gpt-5-system-card/)\(2025\)\.Published August 7, 2025\.
- \[69\]Comanici, G\., Bieber, E\.*et al\.*Gemini 2\.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities\.*arXiv preprint arXiv:2507\.06261*\(2025\)\.
- \[70\]Zhang, S\.*et al\.*A multimodal biomedical foundation model trained from fifteen million image–text pairs\.*NEJM AI*2, DOI:[10\.1056/AIoa2400640](https://doi.org/10.1056/AIoa2400640)\(2024\)\.
- \[71\]Sellergren, A\.*et al\.*Medgemma technical report\.*arXiv preprint arXiv:2507\.05201*\(2025\)\.
- \[72\]Lundberg, S\. M\. & Lee, S\.\-I\.A unified approach to interpreting model predictions\.In*Advances in Neural Information Processing Systems*, vol\. 30 \(2017\)\.
- \[73\]Papineni, K\., Roukos, S\.*et al\.*Bleu: a method for automatic evaluation of machine translation\.In*Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics*, 311–318 \(2002\)\.
- \[74\]Lin, C\.\-Y\.Rouge: A package for automatic evaluation of summaries\.In*Text Summarization Branches Out*, 74–81 \(2004\)\.
- \[75\]Banerjee, S\. & Lavie, A\.Meteor: An automatic metric for mt evaluation with improved correlation with human judgments\.In*Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization*, 65–72 \(2005\)\.
- \[76\]Obermeyer, Z\., Powers, B\., Vogeli, C\. & Mullainathan, S\.Dissecting racial bias in an algorithm used to manage the health of populations\.*Science*366, 447–453, DOI:[10\.1126/science\.aax2342](https://doi.org/10.1126/science.aax2342)\(2019\)\.
- \[77\]Gong, H\., Lu, Y\., Wan, X\. & Li, H\.Domain generalized medical landmark detection via robust boundary\-aware pre\-training\.In*Proceedings of the AAAI Conference on Artificial Intelligence*, vol\. 39, 3140–3148 \(2025\)\.
- \[78\]Gong, H\.*et al\.*Intermediate domain alignment and morphology analogy for patent\-product image retrieval\.*Advances in Neural Information Processing Systems*38, 14501–14523 \(2026\)\.
- \[79\]Mohammadi, A\.*et al\.*Lymphus: A multicenter open\-access database of lymph node ultrasound images in patients with papillary thyroid carcinoma for clinical and artificial intelligence research\.*Data in Brief*66, 112694, DOI:[10\.1016/j\.dib\.2026\.112694](https://doi.org/10.1016/j.dib.2026.112694)\(2026\)\.
- \[80\]Abbasian Ardakani, A\.*et al\.*Diagnosis of metastatic lymph nodes in patients with papillary thyroid cancer: A comparative multi\-center study of semantic features and deep learning\-based models\.*Journal of Ultrasound in Medicine*42, 1211–1221, DOI:[10\.1002/jum\.16131](https://doi.org/10.1002/jum.16131)\(2023\)\.
- \[81\]Torralba, A\. & Efros, A\. A\.Unbiased look at dataset bias\.In*Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition*, 1521–1528 \(2011\)\.
- \[82\]Siméoni, O\.*et al\.*Dinov3\.*arXiv preprint arXiv:2508\.10104*\(2025\)\.
- \[83\]Menon, A\. K\.*et al\.*Long\-tail learning via logit adjustment\.*arXiv preprint arXiv:2007\.07314*\(2020\)\.
- \[84\]Erickson, N\.*et al\.*Autogluon\-tabular: Robust and accurate automl for structured data\.*arXiv preprint arXiv:2003\.06505*\(2020\)\.
- \[85\]Selvaraju, R\. R\.*et al\.*Grad\-cam: Visual explanations from deep networks via gradient\-based localization\.In*Proceedings of the IEEE International Conference on Computer Vision*, 618–626 \(2017\)\.
- \[86\]Wunderling, T\.*et al\.*Comparison of thyroid segmentation techniques for 3d ultrasound\.In*Medical Imaging 2017: Image Processing*, vol\. 10133, 346–352 \(2017\)\.
- \[87\]Hou, X\.*et al\.*An ultrasonography of thyroid nodules dataset with pathological diagnosis annotation for deep learning\.*Scientific Data*11, 1272, DOI:[10\.1038/s41597\-024\-04156\-5](https://doi.org/10.1038/s41597-024-04156-5)\(2024\)\.
- \[88\]Stanford AIMI\.Thyroid ultrasound cine\-clip dataset, DOI:[10\.71718/7m5n\-rh16](https://doi.org/10.71718/7m5n-rh16)\(2024\)\.
- \[89\]Yang, Y\.*et al\.*An annotated heterogeneous ultrasound database\.*Scientific Data*12, 148 \(2025\)\.
- \[90\]Chen, J\.*et al\.*Transunet: Rethinking the u\-net architecture design for medical image segmentation through the lens of transformers\.*Medical Image Analysis*97, 103280, DOI:[10\.1016/j\.media\.2024\.103280](https://doi.org/10.1016/j.media.2024.103280)\(2024\)\.
- \[91\]Zhang, S\.*et al\.*A generalist foundation model and database for open\-world medical image segmentation\.*Nature Biomedical Engineering*10, 1026–1041, DOI:[10\.1038/s41551\-025\-01497\-3](https://doi.org/10.1038/s41551-025-01497-3)\(2026\)\.Published online 5 September 2025\.
- \[92\]He, K\., Zhang, X\., Ren, S\. & Sun, J\.Deep residual learning for image recognition\.In*Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition \(CVPR\)*, 770–778 \(2016\)\.
- \[93\]Wang, A\., Chen, H\., Lin, Z\., Han, J\. & Ding, G\.Lsnet: See large, focus small\.In*Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition*, 9718–9729 \(2025\)\.
- \[94\]Bai, S\., Cai, Y\., Chen, R\.*et al\.*Qwen3\-vl technical report\.*arXiv preprint arXiv:2511\.21631*\(2025\)\.
- \[95\]OpenAI\.Gpt\-4o system card\.*arXiv preprint arXiv:2410\.21276*\(2024\)\.
- \[96\]Yang, A\., Li, A\.*et al\.*Qwen3 technical report\.*arXiv preprint arXiv:2505\.09388*\(2025\)\.
- \[97\]Li, C\., Wong, C\.*et al\.*Llava\-med: Training a large language\-and\-vision assistant for biomedicine in one day\.*Advances in Neural Information Processing Systems*36, 28541–28564 \(2023\)\.

Similar Articles

RadAgent: A tool-using AI agent for stepwise interpretation of chest computed tomography

Hugging Face Daily Papers

RadAgent is a tool-using AI agent that generates chest CT reports through interpretable step-by-step reasoning, improving clinical accuracy by 36.4% relative and achieving 37% faithfulness—a capability absent in existing 3D vision-language models. The system provides fully inspectable reasoning traces allowing clinicians to validate and refine diagnostic outputs.

Skill-Augmented AI Agents for Medical Research Analysis: An Exploratory Multi-Model Human Evaluation in an NSCLC Transcriptomic Biomarker Task

arXiv cs.AI

This exploratory study evaluates whether augmenting AI agents with a medical research skill package improves the quality of transcriptomic research analysis outputs compared to native AI, using a multi-model human evaluation in an NSCLC biomarker task. Results show a directional but statistically non-significant improvement, highlighting the need for larger, more robust evaluations.

Expert Evaluation of Clinical AI Tools on Real Point-of-Care Clinical Queries

arXiv cs.AI

This paper presents a blinded evaluation of clinical AI tools using real point-of-care queries from physicians, comparing specialized and general-purpose models across five dimensions. The specialized tool (OpenEvidence) outperformed general-purpose models on all axes, and the authors release the Real-POCQi benchmark.