From Continuous Predictors to Clinical Thresholds: Early Evidence on Performance Trade-offs of Guideline-Based Categorisation for Ischaemic Stroke Outcome Prediction
摘要
The paper investigates whether guideline-based categorical encodings of continuous predictors can replace continuous inputs in stroke outcome prediction models without sacrificing accuracy, finding comparable performance in most treatment cohorts.
查看缓存全文
缓存时间: 2026/08/07 07:45
# From Continuous Predictors to Clinical Thresholds: Early Evidence on Performance Trade-offs of Guideline-Based Categorisation for Ischaemic Stroke Outcome Prediction Source: [https://arxiv.org/html/2608.05203](https://arxiv.org/html/2608.05203) \\copyrightclause Copyright for this paper by its authors\. Use permitted under Creative Commons License Attribution 4\.0 International \(CC BY 4\.0\)\. \\conference EXPLIMED 2026 \- Third Workshop on Explainable Artificial Intelligence for the medical domain \- 15\-17 August 2026, Bremen, Germany \[orcid=0000\-0003\-2288\-2406, email=esra\.zihni@tudublin\.ie, url=https://github\.com/esrazihni/, \]\\cormark\[1\] \[orcid=0000\-0002\-2888\-9943, \] \[orcid=0000\-0001\-9288\-6520 \] \[orcid=0000\-0003\-3950\-8453 \] \[orcid=0000\-0002\-7458\-5166 \] \[orcid=0000\-0001\-6462\-3248 \] Esra ZihniKatryna CisekHamzah ZiadehAalborg University, Aalborg, DenmarkHendrik KnocheRobert MikulikInternational Clinical Research Center, St\. Anne’s University Hospital, Brno, CzechiaHealth Management Institute, Brno, CzechiaJohn D\. KelleherTrinity College Dublin, Ireland \(2026\) ###### Abstract Machine learning models achieve strong predictive accuracy for 90\-day outcome prediction in acute ischaemic stroke, yet clinical adoption is limited by the misalignment of model explanations with clinicians’ reasoning\. Motivated by a clinician user study calling for clinical guideline\-aligned cut\-offs, we ask whether continuous predictors can be replaced by clinically informed categorical encodings without sacrificing performance\. On a multi\-centre European registry stratified into three treatment cohorts, we compare standard and fully categorised gradient\-boosted models, the latter using stroke guideline\-aligned, treatment\-specific thresholds\. The fully categorised models are statistically indistinguishable from their continuous counterparts in two of the treatment cohorts, with a significant drop in predictive accuracy in one cohort\. Global feature importance rankings remain consistent, suggesting that discretising continuous predictors into guideline‑based categories preserves the core hierarchy of prognostic factors across all treatment groups\. Guideline\-based categorisation is thus a viable design choice for stroke\-outcome models\. ###### keywords: explainable AI\\sepclinical decision support\\sepstroke outcome prediction\\sepSHAP\\sepgradient boosted decision trees\\sepuser\-centred evaluation ## 1Introduction Machine\-learning models for outcome prediction in acute ischaemic stroke have matured rapidly\. Recent gradient\-boosting and ensemble models, paired with SHapley Additive exPlanations \(SHAP\)\[lundberg\_2017\], now routinely achieve AUROCs in the 0\.81\-0\.91 range for 90\-day functional outcomes on retrospective tabular data\[diprose\_2024,abujaber\_2025,chen\_2025\]\. Predictive accuracy, however, is no longer the limiting step for clinical adoption\. Recent studies of human\-centred XAI evaluation in clinical decision support identify a different bottleneck: explanations that are technically faithful to the model but not aligned with clinicians’ reasoning impose cognitive load, induce either mistrust or over\-reliance, and fail to translate into actionable decisions\[sivaraman\_2023,karagoz\_2024\]\. This is the gap our work targets, in the specific setting of dichotomised 90\-day mRS prediction for AIS\. Surveys of human\-centred xAI evaluation suggest that meaningful explanations cannot be defined from inside the model and must be elicited from the intended end\-users\[kim\_2024\]\. In the clinical setting this elicitation typically combines qualitative methods, such as semi\-structured interviews and think\-aloud studies, with standardised quantitative instruments\. To this end, we recently conducted a structured user study with 12 stroke clinicians, evaluating an earlier iteration of the standard models considered here\. The study combined a quantitative arm \(validated questionnaires covering causability, workload, and technology acceptance\) with a qualitative arm \(semi\-structured interviews\) in which participants interacted with individual SHAP explanations and a what\-if interface\. Eight of the twelve clinicians reported that continuous predictors \(blood pressure, cholesterol, and blood glucose\) added unnecessary detail and reduced their understanding of the model: small perturbations to raw values during ’what\-if’ analysis produced visible feature attribution changes that the clinicians considered clinically meaningless, since their own reasoning operates over guideline\-aligned ranges rather than exact values\. Six participants suggested training two models, one as presented to them in the study and one with clinically informed cut\-off points to represent a prediction more aligned with their mental process\. The implication of this feedback is that the limitation lies not in SHAP\-based explanations themselves but in the units and reference frame in which feature attributions are presented to clinicians, an implication inferred in other recent studies\[panigutti\_2022,hur\_2025\]\. Acting on the findings of this user study requires re\-encoding continuous predictors as categories aligned with clinical guidelines\. The biostatistics literature is notably cautious about such re\-encoding since categorisation can result in loss of statistical power\[royston\_2006\]\. Furthermore, most relevant work in clinical ML typically relies on data\-driven cut\-off points\[maslove\_2013,chen\_2019,tustumi\_2022\], which makes the resulting models prone to systematic bias\[naggara\_2011\]\. Clinically driven feature engineering, more broadly, has nonetheless been shown to reduce model complexity at little or no cost to accuracy\[roe\_2020\], suggesting that categorisation driven by guideline\-aligned thresholds may maintain predictive power while aligning explanations with clinical reasoning\. Whether this is the case has not, to our knowledge, been examined\. In this work, we address this gap using 90\-day outcome prediction in acute ischaemic stroke patients as a case study\. We train gradient\-boosted decision trees on a multi\-centre stroke registry, stratified into three treatment cohorts\. For each cohort we compare a standard model, in which continuous predictors are retained on their original scale, against a fully categorised model, in which every eligible continuous predictor is replaced by a categorical encoding whose thresholds are derived from stroke\-specific clinical guidelines\. Where guideline targets differ across treatment pathways, the same underlying predictor is encoded with treatment pathway\-specific thresholds\. To our knowledge, this is the first paired comparison of guideline\-based categorisation of continuous predictors for 90\-day mRS prediction in acute ischaemic stroke, motivated by a documented clinician user study rather than by statistical or computational considerations\. ## 2Methods ### 2\.1Data and cohort We used a subset of the multinational RES\-Q stroke registry data which included 81,735 records from Czechia, Bulgaria, Poland, Greece, Romania\[mikulik\_registry\_2017\]\.‘ Modelling was restricted to acute ischaemic stroke patients, and the prediction target was the 90\-day modified Rankin Scale \(mRS\), which measures the degree of disability or dependence in daily activities\. Cases were excluded if any of the following held: 90\-day mRS missing; no imaging performed; discharge mRS = 6 \(in\-hospital death\); or transfer to another hospital for treatment\. Time\-related predictors outside clinically relevant ranges \(e\.g\., onset\-to\-door∉\\notin\[0, 120\] hours\) and negative values for laboratory or numeric scoring predictors were removed, and IQR\-based outlier filtering was applied to age \(retained range 40\-104 years\)\. The final modelling cohort comprised 3,017 patients\. Since different treatment pathways represent distinct patient profiles, where different clinical and procedural factors influence outcomes, patients were stratified into three treatment cohorts, and a separate model was trained for each: - •Cohort 1— no recanalisation \(n = 512\) - •Cohort 2— thrombolysis \(n = 1,981\); - •Cohort 3— thrombectomy ± thrombolysis \(n = 524\); ### 2\.2Predictors and preprocessing Predictors were chosen with clinical partners, prioritising predictors available before discharge so the model can support real\-time decisions, and with an emphasis on modifiable factors\. Predictors recorded at the time of discharge \(e\.g\., discharge NIHSS, discharge medication\) were intentionally excluded from this iteration\. The candidate set covers baseline demographics and risk factors, prior medication, baseline labs and vitals, imaging, treatment workflow times, and early post\-treatment findings \(full list in Appendix A\)\. For each treatment cohort, predictors with \>10% missing values or with a dominant category \(most\-frequent category \> 100× the combined frequency of all others\) were removed\. This yielded 28 / 30 / 32 predictors for cohorts 1 / 2 / 3, respectively\. Missing values were imputed within each train/test split using median/mode imputation computed from the training set only, to avoid leakage\. ### 2\.3Guideline\-based categorisation of continuous predictors The contribution of this work is replacing the standard continuous predictors with clinically grounded categorical encodings chosen to align with actionable decision thresholds clinicians use at the bedside\. Categorised predictors span baseline severity and demographics \(age, NIHSS\), metabolic markers \(blood pressure, blood glucose, cholesterol\), and workflow times \(onset\-to\-door, door\-to\-imaging, door\-to\-needle, door\-to\-groin, door\-to\-reperfusion\)\. Cut\-offs for metabolic markers were taken from AHA/ASA guidelines\[prabhakaran\_2026\], whereas relevant stroke studies were used to decide on the age and NIHSS cut\-off points\[adams\_1999,mortensen\_2020\]\. Workflow\-time encodings reflect graded performance targets; for door\-to\-needle, for instance, we used≤\\leq15 / 15\-30 / 30\-45 / \>45 min, capturing quality\-improvement benchmarks motivated by evidence that each additional 15\-min door\-to\-needle delay is associated with worse one\-year outcomes\. Where guidelines differ by treatment pathway \(notably blood pressure and onset\-to\-door\), the same predictor is encoded with treatment pathway\-specific thresholds across the three cohorts\. The full set of thresholds is summarised in Table[1](https://arxiv.org/html/2608.05203#S2.T1)\. Table 1:Guideline\-aligned thresholds used to categorise continuous predictors\. Treatment pathway\-specific encodings are indicated in the*Cohort\(s\)*column\. BP: blood pressure\.PredictorCohort\(s\)CategoriesAge \(years\)1, 2, 3<<65 / 65–<<80 /≥\\geq80NIHSS score1, 2, 3<<6 / 6–<<16 /≥\\geq16Systolic BP \(mmHg\)1≤\\leq110 / 110–<<220 /≥\\geq2202, 3≤\\leq110 / 110–<<180 /≥\\geq180Diastolic BP \(mmHg\)1<<120 /≥\\geq1202, 3<<105 /≥\\geq105Blood glucose \(mg/dL\)1, 2, 3<<60 / 60–140 /\>\>140Cholesterol \(mg/dL\)1, 2, 3<<70 / 70–100 /\>\>100Onset\-to\-door \(h\)1≤\\leq24 /\>\>242≤\\leq4\.5 / 4\.5–24 /\>\>243≤\\leq6 / 6–24 /\>\>24Door\-to\-imaging \(min\)1, 2, 3≤\\leq20 /\>\>20Door\-to\-needle \(min\)2≤\\leq15 / 15–30 / 30–45 /\>\>45Door\-to\-groin \(min\)3≤\\leq60 / 60–90 /\>\>90Door\-to\-reperfusion \(min\)3≤\\leq90 / 90–120 /\>\>120We refer to the models with continuous predictors retained on their original scale asstandard modelsand to the new variants \(all eligible continuous predictors re\-coded as in Table[1](https://arxiv.org/html/2608.05203#S2.T1)\) asfully categorised models\. All other predictors and modelling choices are held fixed across the two variants, isolating the effect of guideline\-based categorisation\. ### 2\.4Multicollinearity Gradient\-boosted trees are robust to multicollinearity in terms of predictive performance, since each split selects the most informative feature from the candidate predictors, but high inter\-feature correlation can still distort SHAP\-based importance attribution\. Because feature importance is central to this work, we ran a multicollinearity check on the full predictor set of every cohort × variant combination \(six predictor sets in total\)\. We computed the adjusted squared GVIF, which extends the standard VIF to categorical predictors with multiple levels and uses the same interpretive threshold \(\>10 indicates problematic collinearity\)\[fox\_monette\_1992\]\. Across all three cohorts and both the standard and fully categorised variants, every adjusted squared GVIF was < 5, so no problematic multicollinearity is present in any of the predictor sets, and SHAP attributions can be interpreted without correction\. ### 2\.5Prediction target The prediction target is the dichotomised 90\-day mRS \(favourable, mRS≤\\leq2, vs\. unfavourable, mRS \> 2\)\. The class imbalance varies across the three treatment cohorts: - •Cohort 1— no recanalisation: favourable / unfavourable≈\\approx332 / 180 \(≈\\approx2 : 1\) - •Cohort 2— thrombolysis: favourable / unfavourable≈\\approx1453 / 528 \(≈\\approx3 : 1\) - •Cohort 3— thrombectomy ± thrombolysis: favourable / unfavourable≈\\approx312 / 212 \(≈\\approx1\.5 : 1\) No resampling, class weighting, or threshold tuning was applied; both standard and fully categorised models are trained on the same class distributions\. ### 2\.6Models and hyperparameter tuning All six models \(3 cohorts × 2 variants\) were developed using CatBoost gradient\-boosted decision trees, chosen for its native handling of categorical predictors and built\-in overfitting control\[dorogush\_catboost\_2018\]\. Hyperparameters were tuned by 5\-fold cross\-validated grid search on each training set: tree depth∈\\in4, 6, 8, learning rate∈\\in0\.01, 0\.03, 0\.1, L2 regularisation rate∈\\in3, 10, 100\. The number of trees was set automatically by CatBoost’s overfitting detector\. ### 2\.7Performance estimation and final model training Generalisation performance was estimated via Monte\-Carlo cross\-validation with 50 random 80/20 train/test splits\. The hyperparameter tuning processed described in the previous section was repeated for the training partition of each split\. The workflow is summarised in Figure[1](https://arxiv.org/html/2608.05203#S2.F1)\. Figure 1:Performance\-estimation workflow\. The data is randomly partitioned into training and test sets 50 times, where 5\-fold CV is used to optimise hyperparameters on each training set\. The resulting optimal models are then evaluated on each test set\.Final models were trained on the complete available dataset for each cohort, using hyperparameters set to the mean or mode of the values obtained across the 50 tuned training splits \(Table[2](https://arxiv.org/html/2608.05203#S2.T2)\)\. SHAP values were subsequently computed for each of these final models\. Table 2:Hyperparameters of the final standard and fully categorised models across the three treatment cohorts\.HyperparameterValue \(Standard/Fully Categorised\)Cohort 1Cohort 2Cohort 3Number of trees186/171194/127214/158Tree depth6/44/44/4Learning rate0\.03/0\.030\.03/0\.10\.03/0\.03L2 regularisation rate3/310/1003/3 ## 3Results ### 3\.1Generalisation performance We report the mean test\-set metric across the 50 splits, with 95% confidence intervals, using AUROC, balanced accuracy, F1 score, sensitivity, and specificity \(Table[3](https://arxiv.org/html/2608.05203#S3.T3)\)\. Differences between the standard and fully categorised models are tested with the Wilcoxon signed\-rank test atα\\alpha= 0\.05\. Across the three cohorts, the fully categorised models retain most of the discrimination of the standard models, with a significant difference observed only in the thrombolysis cohort\. Table 3:Mean test\-set performance \(95% CI\) over 50 Monte\-Carlo splits\. Asterisks mark metrics where the paired difference between standard and fully categorised variants is significant \(Wilcoxon signed\-rank,p<0\.05p<0\.05\)\.CohortVariantAUROCBal\. Acc\.F1SensitivitySpecificityNo recanalisationStandard0\.8200\.7140\.7250\.5140\.913\(0\.809–0\.831\)\(0\.703–0\.724\)\(0\.714–0\.736\)\(0\.491–0\.537\)\(0\.903–0\.923\)Fully0\.8020\.6930\.7020\.4910\.895Categorised\(0\.788–0\.815\)\(0\.679–0\.707\)\(0\.687–0\.718\)\(0\.466–0\.516\)\(0\.882–0\.908\)ThrombolysisStandard0\.807\*0\.662\*0\.682\*0\.386\*0\.938\(0\.802–0\.813\)\(0\.656–0\.668\)\(0\.675–0\.688\)\(0\.375–0\.397\)\(0\.934–0\.942\)Fully0\.785\*0\.634\*0\.650\*0\.324\*0\.945Categorised\(0\.779–0\.790\)\(0\.629–0\.640\)\(0\.643–0\.657\)\(0\.311–0\.336\)\(0\.941–0\.949\)Thrombectomy±\\pmStandard0\.8120\.7270\.7310\.6230\.831thrombolysis\(0\.802–0\.822\)\(0\.717–0\.738\)\(0\.720–0\.741\)\(0\.602–0\.644\)\(0\.818–0\.845\)Fully0\.7970\.7240\.7270\.6330\.816Categorised\(0\.786–0\.809\)\(0\.715–0\.734\)\(0\.717–0\.736\)\(0\.613–0\.652\)\(0\.802–0\.831\)In the no recanalisation and the thrombectomy±\\pmthrombolysis cohorts, all five metrics of the fully categorised variant overlap the 95% CIs of the corresponding standard models\. None of the paired differences reach significance, indicating that guideline\-based categorisation preserves predictive performance in these two cohorts\. Categorisation produces the clearest cost in the thrombolysis cohort: AUROC, balanced accuracy, F1, and sensitivity all decrease significantly, with sensitivity showing the largest drop from 0\.393 to 0\.320, while specificity is unchanged \(0\.938 → 0\.945\)\. Overall, the cost in predictive accuracy of guideline\-based categorisation is isolated to a single cohort, leaving cohorts 1 and 3 statistically equivalent to their continuous counterparts\. ### 3\.2Global feature importance For global importance ranking, we use unit\-norm\-normalised mean absolute SHAP values to allow direct comparison across treatment cohorts and across the standard / fully categorised variants\. Figure[2](https://arxiv.org/html/2608.05203#S3.F2)shows the top\-10 predictors for the standard and fully categorised models in each cohort\. Figure 2:Top\-10 global feature importances in the standard and fully categorised models for each treatment cohort\.The two model variants produce largely similar feature\-importance ratings\. NIHSS and age are among the top predictors in all six models; glucose, diabetes risk, swallowing screening, perfusion imaging, and post\-treatment fever recur across cohorts\. The top\-10 overlap between the two model variants is 8/10 in cohort 1, 10/10 in cohort 2, and 7/10 in cohort 3, with top 2, 3, and 4 predictors overlapping in each cohort, respectively\. One of the main differences between the variants is the redistribution of importance from a continuous predictor to correlated comorbidity indicators: under categorisation, glucose drops out of the top ranks while the presence of hyperlipidaemia and diabetes becomes more prominent\. ### 3\.3Patient\-level explanations Patient\-level SHAP visualisations were generated under the fully categorised encoding for representative cases in each cohort\. Preliminary inspection indicates good agreement between the standard and fully categorised explanations on the same patients, with attributions on categorical predictors rendered against guideline\-aligned reference levels rather than continuous baselines\. We do not draw substantive conclusions from these examples here\. The planned next step is getting clinical feedback, in which clinicians will evaluate the fully categorised model explanations against the standard model ones on alignment with their decision criteria, perceived trust, and actionability\. ## 4Discussion Our work provides early evidence for guideline\-based categorisation of continuous predictors as a viable encoding choice for gradient\-boosted 90\-day mRS prediction in acute ischaemic stroke\. In the no\-recanalisation and the thrombectomy±\\pmthrombolysis cohorts, the fully categorised models are statistically indistinguishable from their continuous counterparts on every metric, indicating that the prognostic information carried by continuous laboratory and workflow predictors is largely retained when those predictors are re\-expressed with clinically informed thresholds\. Furthermore, the global SHAP rankings between the two model variants are largely concordant, with some attribution redistributed within clinical domains \(top predictors shift from continuous blood glucose to diabetes and hyperglycaemia indicators and from continuous blood pressure to heart disease indicators\)\. These suggest that categorisation re\-frames explanations without changing the underlying prognostic structure\. On the other hand, we directly observe the cost of categorisation in the thrombolysis cohort, where AUROC, balanced accuracy, F1, and sensitivity all decrease significantly\. The thrombolysis cohort represents the largest but also the most class\-imbalanced \(≈\\approx1:3 favourable to unfavourable\) cohort\. Note that specificity is preserved while the largest decrease occurs in sensitivity, suggesting that the fully categorised encoding is more susceptible to information loss when discriminating a minority of unfavourable outcomes\. Whether a sensitivity drop of this magnitude is acceptable cannot be answered from predictive metrics alone\. It depends on whether the fully categorised explanations produce measurable improvements in clinician trust, comprehension, and decision quality\. We plan to address this trade\-off with a clinician evaluation of the fully categorised model; for now, it is the central question we leave open here\. This study has several limitations\. First, the analysis is restricted to a single registry; replication on independent stroke registries is needed before claims about generalisability can be made\. Second, our cohort sizes, in particular for the no recanalisation and thrombectomy±\\pmthrombolysis cohorts \(n=512 and 524\), are modest, and we plan to repeat the analysis on a substantially larger sample drawn from the same registry as more data become available\. A larger sample may sharpen estimates and shift the thrombolysis cohort result in either direction: the differences we currently flag as significant could attenuate as the minority class is better represented, or they could persist and become more precisely characterised\. Finally, we have not yet evaluated the patient\-level explanations under the new encoding with clinicians; the qualitative claim that fully categorised SHAP plots are easier to reason about than continuous ones rests, for the moment, on prior user\-study evidence rather than on direct evaluation of the present models\. Closing that loop with a structured comparison of standard and fully categorised individual explanations on the same patients with stroke neurologists is the planned follow\-up to the present work\. ###### Acknowledgements\. This work was supported by the RESQ\+ project funded by EU’s Horizon Europe research and innovation programme under grant agreement No\. 101057603\. ## Declaration on Generative AI During the preparation of this work, the author\(s\) used Claude Sonnet 4\.6 and Grammarly in order to: Paraphrase and reword, grammar and spelling check, and generate images\. After using these tool\(s\)/service\(s\), the author\(s\) reviewed and edited the content as needed and take\(s\) full responsibility for the publication’s content\. ## References ## Appendix ASummary of Predictors Table 4:Summary of predictors used to model dichotomised 90\-day mRS outcomes, by treatment cohort \(no recanalisation, thrombolysis, thrombectomy±\\pmthrombolysis\)\. Continuous predictors are reported as median \(IQR\) and "yes/no" predictors as the percentage of*yes*\. C: cohort; BP: blood pressure\.PredictorC1C2C3PredictorC1C2C3Baseline — admissionBaseline — medicationage \(years\)73 \(15\)73 \(15\)73 \(16\)any antiplatelets25\.435\.523\.7sex \(% male\)55\.452\.850\.4any anticoagulants19\.95\.719\.9NIHSS score3 \(4\)5 \(4\)14 \(9\)Imagingwakeup stroke5\.94\.43\.2door\-to\-imaging \(min\)15 \(23\)9 \(10\)10 \(10\)arrival mode \(%\)perfusion imaging done14\.017\.029\.7EMS from home/scene80\.491\.891\.1old infarcts96\.895\.393\.6private transportation15\.26\.20\.6occlusion found13\.019\.199\.8from stroke centre2\.41\.67\.0Treatmentin\-hospital stroke2\.00\.41\.3thrombolysis——71\.0hospitalised in \(%\)IVT department \(%\)ICU / stroke unit68\.699\.599\.4radiology—50\.5—standard bed25\.40\.20\.2emergency—27\.9—monitored bed6\.00\.30\.4stroke unit / ICU—21\.4—onset\-to\-door \(hours\)11\.4 \(19\.5\)1\.7 \(1\.6\)1\.7 \(1\.9\)door\-to\-needle \(min\)∗—22 \(17\)—door\-to\-groin \(min\)——67 \(35\)Baseline — comorbiditiesdoor\-to\-reperfusion \(min\)——102 \(50\)atrial fibrillation20\.911\.027\.1mTICI \(%\>\>2A\)——84\.5hypertension78\.974\.571\.5Post\-treatmentdiabetes57\.447\.043\.8stroke etiology known69\.672\.858\.5hyperlipidaemia42\.842\.038\.1fever diagnosed7\.36\.311\.2congestive heart failure7\.05\.08\.8hyperglycaemia diagnosed14\.813\.413\.1smoker25\.827\.125\.6swallowing screening \(%\)previous stroke16\.820\.413\.1yes86\.292\.788\.4coronary heart disease14\.811\.914\.5no9\.53\.20\.4not applicable4\.34\.111\.2Baseline — labstherapy \(%\)systolic BP \(mmHg\)153 \(34\)160 \(33\)152\.0 \(30\)yes81\.583\.784\.1diastolic BP \(mmHg\)87 \(15\)87 \(16\)85\.0 \(16\)no6\.53\.78\.9blood glucose \(mg/dL\)113 \(45\)121 \(47\)121 \(40\)not required12\.012\.67\.0cholesterol \(mg/dL\)47 \(29\)47 \(27\)45 \(25\) ∗Door\-to\-needle time is included in door\-to\-groin time when thrombolysis precedes thrombectomy\.
相似文章
基于CPRD的可扩展临床数据基础设施及比较性机器学习评估在多重慢性病老年患者住院风险预测中的应用
本文提出了一种可扩展的临床数据基础设施,并比较了深度学习(TG-CNN)与传统机器学习模型(LASSO和Random Forests)在多重慢性病老年患者住院风险预测中的应用,结论是由于LASSO具有更优的校准性能,更适合临床部署。
治疗诱发的标签不确定性下利用专家对反事实结果的标注进行学习:神经系统预后预测案例研究
本文针对临床预测中治疗导致的标签不确定性问题,以心脏骤停后神经预后判断为案例研究。作者提出了一个纳入专家对反事实结果标注的框架,并强调了确定病例与不确定病例之间准确性的权衡。
一种可迁移的基于阈值的可解释医学数据分类框架
本文介绍了一种基于统计的框架,使用伯努利朴素贝叶斯和χ²引导的二值化进行可解释的临床分类,在基准数据集上取得了有竞争力的AUC分数,同时提供了明确的决策规则和校准的风险估计。
机器学习评估炎症生物标志物对老年西班牙裔成人队列认知障碍的预测价值
本文介绍了一种可解释的机器学习方法,使用朴素贝叶斯分类器在小型临床数据集上从炎症生物标志物预测认知障碍,并识别出I-309 (CCL1)作为关键预测特征。
基于LLM探针的主要ICD类别预测
本文提出了一种方法,利用冻结的医学大型语言模型(LLM)表示作为共享嵌入空间,从结构化和非结构化电子健康记录数据中预测主要ICD诊断类别,在MIMIC-IV上取得了优于基线方法的准确率,并展示了向MIMIC-III的迁移能力。