A Machine-Learned Comorbidity Index
Summary
This paper proposes a Machine-Learned Comorbidity Index (MLCI) that uses diagnosis codes and nonlinear learning to improve risk adjustment across multiple clinical outcomes, outperforming traditional mortality-centric indices.
View Cached Full Text
Cached at: 06/17/26, 05:36 AM
# A Machine-Learned Comorbidity Index
Source: [https://arxiv.org/html/2606.17450](https://arxiv.org/html/2606.17450)
###### Abstract
Traditional comorbidity scores \(e\.g\., Charlson and Elixhauser\) are widely used for risk adjustment and patient stratification, but they have two key limitations: \(i\) they are largely mortality\-centric and do not align well with other clinical outcomes, and \(ii\) their linear, rule\-based structure cannot capture nonlinear, outcome\-specific risk relationships\. We propose a Machine\-Learned Comorbidity Index \(MLCI\) that maps diagnosis codes to a single scalar by maximizing the normalized Hilbert–Schmidt Independence Criterion \(nHSIC\) between the learned score and multiple clinical outcomes\. MLCI captures nonlinear risk–outcome dependence and is supported by a theory that characterizes when a unified, informative admission\-level ordering can be achieved across outcomes\. Empirical results on multiple benchmark electronic health record \(EHR\) datasets show that MLCI outperforms strong baselines across multiple evaluation metrics\.
Machine Learning, ICML
## 1Introduction
We definecomorbidityas the overall severity and complexity of the diagnoses recorded for a hospital admission, reflected in diagnosis codes\. Comorbidity scores are widely used for patient stratification\(Eninget al\.,[2015](https://arxiv.org/html/2606.17450#bib.bib8)\), risk stratification\(O’Haraet al\.,[2024](https://arxiv.org/html/2606.17450#bib.bib9)\), and risk adjustment\(Ouet al\.,[2012](https://arxiv.org/html/2606.17450#bib.bib16); Quanet al\.,[2011](https://arxiv.org/html/2606.17450#bib.bib19)\)\. These applications rely largely on hand\-engineered comorbidity indices that compress diagnosis information into a single scalar score\. However, traditional indices such as Charlson\(Charlsonet al\.,[1987](https://arxiv.org/html/2606.17450#bib.bib3)\)and Elixhauser\(Elixhauseret al\.,[1998](https://arxiv.org/html/2606.17450#bib.bib4); Van Walravenet al\.,[2009](https://arxiv.org/html/2606.17450#bib.bib25)\)have two key limitations\.
First, as risk indicators, these indices were designed for mortality and often generalize poorly to other clinical outcomes because they rely on predefined comorbidity categories and mortality\-calibrated fixed weights\. This undermines the implicit assumption behind comorbidity\-based stratification: that a patient’s diagnosis burden during a hospital admission yields a shared ordering of admissions that is meaningful across multiple clinical outcomes\(Byleset al\.,[2005](https://arxiv.org/html/2606.17450#bib.bib56)\)\. For example, admissions involving “sicker” patients typically face higher risk across outcomes such as mortality and ICU\-level care than admissions involving “less sick” patients\(Kuswardhaniet al\.,[2020](https://arxiv.org/html/2606.17450#bib.bib21)\)\. However, existing indices lack a principled way to learn a data\-driven severity ordering that is consistent across these outcomes\. Clinicians also need more than a ranking: they need a cutoff that identifies high\-severity admissions for intervention\(Billingset al\.,[2006](https://arxiv.org/html/2606.17450#bib.bib2); Patelet al\.,[2021](https://arxiv.org/html/2606.17450#bib.bib55)\)\. Traditional indices identify such high\-risk groups using mortality\-calibrated thresholds, rather than thresholds learned to reflect consistent risk across outcomes\.
Second, their linear, rule\-based nature limits their ability to capture nonlinear relationships between risk and outcomes\. Risk can change nonlinearly with comorbidity burden\(Weiet al\.,[2023](https://arxiv.org/html/2606.17450#bib.bib58); Lyet al\.,[2025](https://arxiv.org/html/2606.17450#bib.bib57)\), as certain diagnosis combinations may amplify risk\(Willadsenet al\.,[2018](https://arxiv.org/html/2606.17450#bib.bib59)\)while additional diagnoses can have diminishing impact at high baseline severity\. Clinical risk tools often reflect this nonlinear structure: SAPS II maps severity to mortality through a logistic\-style link, while NEWS uses score cutoffs to flag patients at higher risk of rapid deterioration and severe outcomes\(Le Gallet al\.,[2005](https://arxiv.org/html/2606.17450#bib.bib5); Kimet al\.,[2021](https://arxiv.org/html/2606.17450#bib.bib6)\)\.
Together, these limitations motivate three central questions that guide our approach:
1. 1\.To what extent do common hospital outcomes share an underlying admission\-level severity ordering, so that a single score can rank admissions consistently across outcomes?
2. 2\.If such an ordering exists, can we learn it in a principled, data\-driven way while modeling outcome\-specific nonlinear severity–risk links?
3. 3\.Beyond ranking admissions, can we learn a cutoff that isolates a high\-severity group consistently across outcomes?
Answering these questions requires a model that produces a single severity score while capturing nonlinear diagnosis\-code interactions\. A multi\-endpoint kernel dependence objective lets the score jointly capture nonlinear associations with clinical outcomes, thereby learning shared outcome\-relevant signal beyond linear correlation\.
To address these challenges, we proposeMLCI,a Machine\-Learned Comorbidity Indexthat maps the diagnosis\-code representationXiX\_\{i\}for admissioniito a scalar scoresi=sθ\(Xi\)∈ℝs\_\{i\}=s\_\{\\theta\}\(X\_\{i\}\)\\in\\mathbb\{R\}\. MLCI is trained to recover a shared, clinically useful admission\-level ordering when such an ordering exists, while allowing each outcome to have its own nonlinear severity–risk relationship\. It uses normalized Hilbert–Schmidt Independence Criterion \(nHSIC\), a kernel\-based, scale\-comparable dependence measure across outcomes\. We learnsi=sθ\(Xi\)s\_\{i\}=s\_\{\\theta\}\(X\_\{i\}\)by maximizing normalized dependence with multiple clinical endpoints, so the score captures shared outcome\-relevant signal without being dominated by any single endpoint\. After training, we estimate separate risk curves for each outcome, mapping the shared score to outcome\-specific risk\.
We motivate our approach by atheoretical analysisthat characterizes when a single learned comorbidity score can serve as an approximate shared ordering across multiple clinical outcomes\. In the theory, the finite\-sample units are hospital admissions\. We posit that each admissioniihas an unobserved latent scalar severityziz\_\{i\}, representing admission\-level comorbidity burden\. The ordering of theziz\_\{i\}’s represents a ranking of admissions by latent severity\. Because outcomes may depend on severity through unknown, potentially nonlinear relationships, the goal is to recover the ordering of theziz\_\{i\}’s rather than their absolute scale\. The learned scoresθ\(Xi\)s\_\{\\theta\}\(X\_\{i\}\)is therefore intended to preserve this latent admission\-level ordering from diagnosis features, so that admissions with larger latent severity receive larger learned scores\. Under this view, maximizing nHSIC compares admissions pairwise\. On the score side, it asks which admissions are close under the learned score\. On the label side, it asks which admissions have the same binary outcome\. A high nHSIC value means these two pairwise patterns agree: admissions close in score tend to have similar labels\. We then study what happens when this comparison is made across several outcomes\. For each clinical outcome, we form a binary label vector over admissions, whose entries indicate whether that endpoint occurred for a given admission; we then center and stack these vectors across outcomes\. If the stacked matrix is approximately rank one, then the outcomes share a dominant admission\-level direction\. This direction gives a simple threshold rule: choose the admission split whose centered step vector is most aligned with the shared direction\. We use the strength of this shared component and the implied threshold split as diagnostics for shared ordering, rather than outcome\-specific deviations\.
Overall, the main contributions of this paper are summarized as follows:
1. 1\.A novel single\-score machine\-learned comorbidity index\.We introduce MLCI, a data\-driven comorbidity score that learns a shared admission\-level latent risk by maximizing nHSIC with multiple clinical outcomes, capturing nonlinear effects in a one\-dimensional summary\.
2. 2\.Theory for shared severity ordering and threshold stratification\.To our knowledge, this is the first finite\-sample analysis linking multi\-outcome nHSIC to a shared monotone admission\-level ordering\. When outcomes share a dominant severity signal, the objective identifies a common admission\-level direction and motivates a principled high\-severity cutoff, which we evaluate on MIMIC\-III and MIMIC\-IV\.
3. 3\.Consistent gains in dependence metrics\.On MIMIC\-III/IV, MLCI shows the strongest score–outcome dependence, outperforming strong single\-index clinical baselines in statistical dependence measures\.
## 2Related Work
Comorbidity indices are a longstanding foundation for clinical risk modeling\. The Charlson Comorbidity Index \(CCI\)\(Charlsonet al\.,[1987](https://arxiv.org/html/2606.17450#bib.bib3)\)assigns fixed weights to predefined chronic conditions to predict mortality\. Elixhauser measures expand the CCI with more diagnosis\-derived conditions and often improve inpatient outcome prediction\(Elixhauseret al\.,[1998](https://arxiv.org/html/2606.17450#bib.bib4)\)\. Van Walraven et al\. distilled them into a single scalar score for in\-hospital mortality that improves usability while retaining strong discrimination\(Van Walravenet al\.,[2009](https://arxiv.org/html/2606.17450#bib.bib25)\)\. A common alternative to rule\-based indices treats diagnosis codes as high\-dimensional sparse features and learns outcome models directly from data; classical machine\-learning methods remain competitive for clinical outcome prediction: logistic regression is interpretable yet can match complex diagnosis\-code mortality models\(Cowlinget al\.,[2021](https://arxiv.org/html/2606.17450#bib.bib26)\); gradient\-boosted trees capture nonlinearities\(Liet al\.,[2025](https://arxiv.org/html/2606.17450#bib.bib27)\); and factorization machines model low\-rank sparse\-code interactions for multi\-task prediction\(Yinet al\.,[2025](https://arxiv.org/html/2606.17450#bib.bib28)\)\. More recently, neural models also embed diagnosis codes and learn nonlinear aggregations, including attention over ICD codes for early length of stay and in\-hospital mortality prediction\(Liuet al\.,[2020](https://arxiv.org/html/2606.17450#bib.bib32); Harerimanaet al\.,[2021](https://arxiv.org/html/2606.17450#bib.bib33)\)\. Most clinical prediction models optimize a single\-outcome likelihood, producing task\-specific representations\. A complementary approach is to learn representations by maximizing dependence with different outcomes\. HSIC is a kernel\-based dependence measure with an empirical centered\-Gram estimator enabling sample\-based independence testing\(Grettonet al\.,[2005](https://arxiv.org/html/2606.17450#bib.bib35),[2007](https://arxiv.org/html/2606.17450#bib.bib36)\)\. Centered kernel alignment connects dependence measurement to normalized, scale\-invariant HSIC\-style criteria\(Corteset al\.,[2012](https://arxiv.org/html/2606.17450#bib.bib37)\), and HSIC has also been used directly as a dependence\-maximization training signal, including self\-supervised objectives that optimize kernel dependence between views\(Liet al\.,[2021](https://arxiv.org/html/2606.17450#bib.bib38)\)\. HSIC is also used for nonlinear feature selection and biomarker discovery in prediction, reducing redundancy and dimensionality in clinical domains\(Takahashiet al\.,[2020](https://arxiv.org/html/2606.17450#bib.bib39); Yuet al\.,[2023](https://arxiv.org/html/2606.17450#bib.bib41); Daiet al\.,[2025](https://arxiv.org/html/2606.17450#bib.bib40)\)\.
## 3Preliminaries
#### Comorbidity indices\.
The two widely used comorbidity indices in the current literature are the Charlson Comorbidity Index \(CCI\) and the van Walraven\-weighted Elixhauser Comorbidity Index \(ECI\)\. We implemented both indices using Quan et al\.’s published coding algorithms\(Quanet al\.,[2005](https://arxiv.org/html/2606.17450#bib.bib53)\)\. Detailed notation and scoring formulas are provided in Appendix[D](https://arxiv.org/html/2606.17450#A4)\. To make the connection between classic indices and our approach explicit, we describe CCI and ECI through a common input–output lens\. Both CCI and the van Walraven ECI map an admission’s diagnosis\-code representationXiX\_\{i\}, restricted to ICD codes, to a scalar index scoreci∈ℝc\_\{i\}\\in\\mathbb\{R\}by forming comorbidity\-category indicators fromXiX\_\{i\}and summing them with fixed weights to summarize baseline illness burden\.
## 4Problem Formulation
Our goal is to develop a*machine\-learned comorbidity score*that preserves the CCI/ECI input/output contract: given admission featuresXiX\_\{i\}for admissionii, output a single scalar scoresi:=sθ\(Xi\)∈ℝ,s\_\{i\}:=s\_\{\\theta\}\(X\_\{i\}\)\\in\\mathbb\{R\},wheresθs\_\{\\theta\}is a neural network mapping admission features to a real\-valued comorbidity score\. Following the single\-index philosophy of CCI/ECI, we aim forsθ\(Xi\)s\_\{\\theta\}\(X\_\{i\}\)to provide a*common admission\-level ordering*that captures shared comorbidity burden across multiple clinical outcomes\. We observe hospital admissions indexed byi∈\{1,…,n\}i\\in\\\{1,\\dots,n\\\}\. Each admission has diagnosis\-code featuresXiX\_\{i\}andTTbinary clinical outcomes\. For taskt∈\[T\]:=\{1,…,T\}t\\in\[T\]:=\\\{1,\\dots,T\\\}, writeyi\(t\)∈\{0,1\}y\_\{i\}^\{\(t\)\}\\in\\\{0,1\\\}for the outcome label of admissioniion tasktt\. Outcomes may be missing; letMi\(t\)∈\{0,1\}M\_\{i\}^\{\(t\)\}\\in\\\{0,1\\\}indicate whetheryi\(t\)y\_\{i\}^\{\(t\)\}is observed\.
Shared latent severity model\.Comorbidity scores, while designed specifically for mortality, are commonly used to quantify the risks of other clinical outcomes such as ICU transfer\(Katzet al\.,[2023](https://arxiv.org/html/2606.17450#bib.bib47)\)\. Based on this practice, we posit that each admissioniihas an unobserved real\-valued latent severityzi∈ℝz\_\{i\}\\in\\mathbb\{R\}, representing overall sickness or disease burden\. Tasktthas an outcome\-specific response curvePr\{yi\(t\)=1∣zi\}=ft\(zi\),\\Pr\\\{y\_\{i\}^\{\(t\)\}=1\\mid z\_\{i\}\\\}=f\_\{t\}\(z\_\{i\}\),whereft:ℝ→\(0,1\)f\_\{t\}:\\mathbb\{R\}\\to\(0,1\)is not assumed to have a parametric form\. Thus, different outcomes may have different prevalence, onset, saturation, and noise patterns, while still sharing a common one\-dimensional severity signal\. The learned scoresθ\(Xi\)s\_\{\\theta\}\(X\_\{i\}\)is intended to recover this shared signal at the level of ordering\. That is, after an optional global sign flip, it should be monotone in the realized latent severity:zi<zj⟹sθ\(Xi\)≤sθ\(Xj\),z\_\{i\}<z\_\{j\}\\implies s\_\{\\theta\}\(X\_\{i\}\)\\leq s\_\{\\theta\}\(X\_\{j\}\),allowing ties\. This matches the role of comorbidity indices in practice: they provide a scalar ordering of admissions rather than a separate risk model for each endpoint\.
We interpret larger learned scores as indicating greater latent severity, but we do not require the observed binary labels to be perfectly monotone in the score\. Clinical outcomes are noisy, and different outcomes may activate at different parts of the severity range\. The theory therefore studies an ideal monotone oracle scorerrthat preserves the latent severity ordering, while the learned scoresθ\(Xi\)s\_\{\\theta\}\(X\_\{i\}\)is treated as a noisy, data\-driven approximation to that ordering\.
Learning from multiple outcomes under heterogeneity\.Training on a single outcome, such as mortality, can yield an outcome\-specific score that may not generalize to other clinical outcomes\. At the same time, naive multi\-task training can suffer from interference or domination by one task because outcomes differ in prevalence, noise level, and severity\-response pattern\. We therefore treat outcomes as heterogeneous signals of a shared latent severity axis, with task\-specific residual structure\. This motivates learning one scalar score that captures the shared axis while preventing any single outcome from dominating\.
## 5Methodology
Figure 1:MLCI: learning a shared severity score via multi\-outcome nHSIC \(centered kernel alignment\)\.We learn a*single*scalar comorbidity/severity scoresi=sθ\(Xi\)∈ℝs\_\{i\}=s\_\{\\theta\}\(X\_\{i\}\)\\in\\mathbb\{R\}from each admission’s diagnosis\-code representationXiX\_\{i\}\. Diagnoses are represented as variable\-length collections of ICD\-prefix tokens \(ICD\-9 in MIMIC\-III, ICD\-10 in MIMIC\-IV\)\. This single\-index representation compresses diagnosis\-code information into one score that is evaluated for dependence with each outcomeyi\(1\),…,yi\(T\)y\_\{i\}^\{\(1\)\},\\dots,y\_\{i\}^\{\(T\)\}\.
Diagnosis representation\.Each raw ICD code is normalized by converting to uppercase and removing punctuation and whitespace, then mapped to a prefix token using the firstk=4k=4characters\. We build a diagnosis vocabulary from the training split only and augment it with<PAD\>\(index 0\) and<UNK\>\(index 1\)\. Using this fixed training vocabulary, we map diagnoses in the train/val/test splits to a variable\-length token collectionXi=\{Ci,1,…,Ci,mi\}X\_\{i\}=\\\{C\_\{i,1\},\\ldots,C\_\{i,m\_\{i\}\}\\\}, wheremim\_\{i\}is the number of retained diagnosis tokens after truncation andDmax=256D\_\{\\max\}=256is the maximum retained length\.
Permutation\-invariant encoder\.We implementsθs\_\{\\theta\}using a DeepSets\-style encoder\(Zaheer and others,[2017](https://arxiv.org/html/2606.17450#bib.bib46)\)\. Letej∈ℝde\_\{j\}\\in\\mathbb\{R\}^\{d\}be an embedding of tokenjj, andϕ:ℝd→ℝd\\phi:\\mathbb\{R\}^\{d\}\\rightarrow\\mathbb\{R\}^\{d\}be an elementwise MLP applied to each token, producinghj=ϕ\(ej\)h\_\{j\}=\\phi\(e\_\{j\}\)\. We aggregate token features using masked mean pooling and masked max pooling and concatenate the results to form an admission representation, then map it to a scalar through a second MLPρ\\rho:si=sθ\(Xi\)=ρ\(Agg\(\{ϕ\(ej\)\}\)\)∈ℝ\.s\_\{i\}=s\_\{\\theta\}\(X\_\{i\}\)=\\rho\\\!\\Big\(\\mathrm\{Agg\}\\big\(\\\{\\phi\(e\_\{j\}\)\\\}\\big\)\\Big\)\\in\\mathbb\{R\}\.Architecture and training details are in Appendix[A](https://arxiv.org/html/2606.17450#A1)\.
Motivation for outcome\-aligned dependence objective\.We assume multiple binary outcomes share a latent severity signal but may follow distinct, possibly nonlinear response curves\. Thus linear correlation can miss score–outcome dependence\. We trainsθ\(X\)s\_\{\\theta\}\(X\)with a kernel dependence objective based on normalized HSIC \(nHSIC\), which measures general \(linear or nonlinear\) dependence via kernels\.
Normalized HSIC \(nHSIC\)\.To compute nHSIC on a mini\-batch of sizenbn\_\{b\}, with scores\{si\}i=1nb\\\{s\_\{i\}\\\}\_\{i=1\}^\{n\_\{b\}\}and task labels\{yi\(t\)\}i=1nb\\\{y\_\{i\}^\{\(t\)\}\\\}\_\{i=1\}^\{n\_\{b\}\}, defineKij\(b\):=k\(si,sj\),Lij\(b,t\):=𝕀\{yi\(t\)=yj\(t\)\},K\_\{ij\}^\{\(b\)\}:=k\(s\_\{i\},s\_\{j\}\),\\qquad L\_\{ij\}^\{\(b,t\)\}:=\\mathbb\{I\}\\\{y\_\{i\}^\{\(t\)\}=y\_\{j\}^\{\(t\)\}\\\},and the mini\-batch centering matrixHb:=I−1nb𝟏𝟏⊤\.H\_\{b\}:=I\-\\frac\{1\}\{n\_\{b\}\}\\mathbf\{1\}\\mathbf\{1\}^\{\\top\}\.The centered mini\-batch Gram matrices areKc\(b\):=HbK\(b\)Hb,Lt,c\(b\):=HbL\(b,t\)Hb\.K\_\{c\}^\{\(b\)\}:=H\_\{b\}K^\{\(b\)\}H\_\{b\},\\qquad L\_\{t,c\}^\{\(b\)\}:=H\_\{b\}L^\{\(b,t\)\}H\_\{b\}\.We optimize the normalized criterion
nHSIC^\(s,y\(t\)\)=⟨Kc\(b\),Lt,c\(b\)⟩Fmax\{‖Kc\(b\)‖F,ε0\}‖Lt,c\(b\)‖F,\\widehat\{\\mathrm\{nHSIC\}\}\\\!\\left\(s,y^\{\(t\)\}\\right\)=\\frac\{\\langle K\_\{c\}^\{\(b\)\},L\_\{t,c\}^\{\(b\)\}\\rangle\_\{F\}\}\{\\max\\\{\\\|K\_\{c\}^\{\(b\)\}\\\|\_\{F\},\\varepsilon\_\{0\}\\\}\\;\\\|L\_\{t,c\}^\{\(b\)\}\\\|\_\{F\}\},\(1\)whereε0\>0\\varepsilon\_\{0\}\>0floors‖Kc\(b\)‖F\\\|K\_\{c\}^\{\(b\)\}\\\|\_\{F\}for numerical stability\.
Instantiation\.We use the Gaussian RBF kernel on the scalar score,kσ\(s,s′\)=exp\(−\|s−s′\|22σ2\),k\_\{\\sigma\}\(s,s^\{\\prime\}\)=\\exp\\\!\\left\(\-\\frac\{\|s\-s^\{\\prime\}\|^\{2\}\}\{2\\sigma^\{2\}\}\\right\),and the delta label Gram matrixL\(b,t\)L^\{\(b,t\)\}defined above\. We set the RBF bandwidthσ\\sigmausing a stabilized median heuristic on training scores \(Appendix[A\.5](https://arxiv.org/html/2606.17450#A1.SS5)\)\.
Missing labels and validity\.LetMi\(t\)∈\{0,1\}M\_\{i\}^\{\(t\)\}\\in\\\{0,1\\\}indicate whether labelyi\(t\)y\_\{i\}^\{\(t\)\}is observed\. MLCI uses the intersection\-valid cohort,Mi∩=∏t=1TMi\(t\)=1,M\_\{i\}^\{\\cap\}=\\prod\_\{t=1\}^\{T\}M\_\{i\}^\{\(t\)\}=1,for multi\-task nHSIC training, validation, and model selection; see Appendix[A\.2](https://arxiv.org/html/2606.17450#A1.SS2)\. BCE baselines use task\-specific masks because their losses decompose by outcome\. For fair per\-outcome comparison, Tables 1–2 evaluate every model on the same task\-specific valid test admissions for that outcome; see Appendix[A\.10](https://arxiv.org/html/2606.17450#A1.SS10)\.
#### Task\-weighted multi\-task dependence training\.
We train a single model to produce one shared scoresi=sθ\(Xi\)s\_\{i\}=s\_\{\\theta\}\(X\_\{i\}\)that is simultaneously dependent on multiple binary outcome vectorsy\(1\),…,y\(T\)y^\{\(1\)\},\\dots,y^\{\(T\)\}\. Our training objective maximizes a weighted sum of per\-task nHSIC values:
maxθ∑t=1TαtnHSIC^\(sθ\(X\),y\(t\)\),\\max\_\{\\theta\}\\ \\sum\_\{t=1\}^\{T\}\\alpha\_\{t\}\\,\\widehat\{\\mathrm\{nHSIC\}\}\\\!\\left\(s\_\{\\theta\}\(X\),y^\{\(t\)\}\\right\),\(2\)wheresθ\(X\)s\_\{\\theta\}\(X\)denotes the vector of scores in the current mini\-batch,y\(t\)y^\{\(t\)\}denotes the corresponding mini\-batch labels for tasktt, andαt≥0\\alpha\_\{t\}\\geq 0controls each task’s contribution\.
Two\-stage weighting\.Outcomes vary in prevalence and alignment with the shared signal, so unweighted training can be dominated by a few tasks\. We set task weights from single\-task nHSIC baselines: \(Stage 1\) train one model per outcome and record best validation nHSICh^t\\widehat\{h\}\_\{t\}; \(Stage 2\) train a new multi\-task model with stabilized inverse\-strength weights
αt∝\(h^maxmax\(h^t,εwt\)\)γwt,h^max=maxth^t,\\alpha\_\{t\}\\propto\\left\(\\frac\{\\widehat\{h\}\_\{\\max\}\}\{\\max\(\\widehat\{h\}\_\{t\},\\varepsilon\_\{\\mathrm\{wt\}\}\)\}\\right\)^\{\\gamma\_\{\\mathrm\{wt\}\}\},\\qquad\\widehat\{h\}\_\{\\max\}=\\max\_\{t\}\\widehat\{h\}\_\{t\},\(3\)withεwt\>0\\varepsilon\_\{\\mathrm\{wt\}\}\>0andγwt∈\(0,1\)\\gamma\_\{\\mathrm\{wt\}\}\\in\(0,1\)\. We clip and renormalize for robustness \(Appendix[A](https://arxiv.org/html/2606.17450#A1)\)\.
Score orientation\.Because the RBF kernel depends only on pairwise score distances, the objective is sign\-invariant inss\. For reporting, we orientssusing validation mortality: if the Pearson correlation betweensis\_\{i\}andyi\(mort\)y\_\{i\}^\{\(\\mathrm\{mort\}\)\}on validation admissions is negative, we setsi←−sis\_\{i\}\\leftarrow\-s\_\{i\}for all admissions, so larger scores correspond to higher mortality risk\.
## 6Theoretical Analysis
Let us assumennadmissions in a given system, such as a hospital system, have latent severities ordered asz1<⋯<zn\.z\_\{1\}<\\cdots<z\_\{n\}\.An oracle score is a bounded nondecreasing mapr:\{z1,…,zn\}→\[−B,B\],zi<zj⟹r\(zi\)≤r\(zj\)\.r:\\\{z\_\{1\},\\dots,z\_\{n\}\\\}\\to\[\-B,B\],\\qquad z\_\{i\}<z\_\{j\}\\implies r\(z\_\{i\}\)\\leq r\(z\_\{j\}\)\.For each ordered admissionii, defineri:=r\(zi\)\.r\_\{i\}:=r\(z\_\{i\}\)\.Thus the oracle score induces the finite score vectorr=\(r1,…,rn\)⊤∈\[−B,B\]n,r=\(r\_\{1\},\\dots,r\_\{n\}\)^\{\\top\}\\in\[\-B,B\]^\{n\},withr1≤r2≤⋯≤rnr\_\{1\}\\leq r\_\{2\}\\leq\\cdots\\leq r\_\{n\}\. The theory studies nHSIC objectives between this finite monotone score vector and the observed binary label vectors\. The goal is to characterize when the multi\-outcome nHSIC objective supports a shared severity ordering across outcomes\. The argument has three steps\. First, each binary outcome is converted into a centered label profile over admissions\. Second, these profiles are stacked across tasks, producing an admission\-to\-admission label\-alignment matrix\. Third, when the stacked label profiles are approximately rank one, the leading admission\-level directionvvyields an explicit monotone threshold rule: choose the admission split whose centered step vector is most aligned withvv\.
#### Oracle setup\.
For a fixedB\>0B\>0, define the finite\-sample monotone score classℛB↑:=\{r∈\[−B,B\]n:r1≤⋯≤rn\}\.\\mathcal\{R\}\_\{B\}^\{\\uparrow\}:=\\\{r\\in\[\-B,B\]^\{n\}:r\_\{1\}\\leq\\cdots\\leq r\_\{n\}\\\}\.ForTToutcomes, write\[T\]:=\{1,…,T\}\[T\]:=\\\{1,\\dots,T\\\}\. For each taskt∈\[T\]t\\in\[T\], lety\(t\)=\(y1\(t\),…,yn\(t\)\)⊤∈\{0,1\}ny^\{\(t\)\}=\(y^\{\(t\)\}\_\{1\},\\dots,y^\{\(t\)\}\_\{n\}\)^\{\\top\}\\in\\\{0,1\\\}^\{n\}be the observed binary label vector across thennadmissions\. Define the delta label Gram matrix\(Lt\)ij:=𝕀\{yi\(t\)=yj\(t\)\},\(L\_\{t\}\)\_\{ij\}:=\\mathbb\{I\}\\\{y\_\{i\}^\{\(t\)\}=y\_\{j\}^\{\(t\)\}\\\},where𝕀\{⋅\}\\mathbb\{I\}\\\{\\cdot\\\}denotes the indicator function\. For scores, use a bounded continuous positive semidefinite radial kernelk\(x,x′\)=κ\(\|x−x′\|\),k\(x,x^\{\\prime\}\)=\\kappa\(\|x\-x^\{\\prime\}\|\),whereκ\\kappais strictly decreasing on\[0,2B\]\[0,2B\]\. DefineKij\(r\):=k\(ri,rj\)=κ\(\|ri−rj\|\)\.K\_\{ij\}\(r\):=k\(r\_\{i\},r\_\{j\}\)=\\kappa\(\|r\_\{i\}\-r\_\{j\}\|\)\.LetH:=I−1n𝟏𝟏⊤H:=I\-\\frac\{1\}\{n\}\\mathbf\{1\}\\mathbf\{1\}^\{\\top\}, where𝟏∈ℝn\\mathbf\{1\}\\in\\mathbb\{R\}^\{n\}is the all\-ones vector\. This is the centering matrix\. The centered score Gram matrix isKc\(r\):=HK\(r\)H\.K\_\{c\}\(r\):=HK\(r\)H\.SinceK\(r\)⪰0K\(r\)\\succeq 0for everyr∈ℛB↑r\\in\\mathcal\{R\}\_\{B\}^\{\\uparrow\}, it follows thatKc\(r\)⪰0\.K\_\{c\}\(r\)\\succeq 0\.The centered label Gram matrix for taskttisLt,c:=HLtHL\_\{t,c\}:=HL\_\{t\}H\. The per\-task normalized HSIC objective between the oracle score vector and tasktt’s labels is
nHSIC\(r,y\(t\)\):=⟨Kc\(r\),Lt,c⟩F‖Kc\(r\)‖F‖Lt,c‖F\.\\mathrm\{nHSIC\}\\\!\\left\(r,y^\{\(t\)\}\\right\):=\\frac\{\\langle K\_\{c\}\(r\),L\_\{t,c\}\\rangle\_\{F\}\}\{\\\|K\_\{c\}\(r\)\\\|\_\{F\}\\\|L\_\{t,c\}\\\|\_\{F\}\}\.We use the conventionnHSIC\(r,y\(t\)\)=0\\mathrm\{nHSIC\}\(r,y^\{\(t\)\}\)=0whenever either denominator factor is zero\. The multi\-task objective is
J\(r\):=∑t=1TαtnHSIC\(r,y\(t\)\),αt\>0,J\(r\):=\\sum\_\{t=1\}^\{T\}\\alpha\_\{t\}\\,\\mathrm\{nHSIC\}\\\!\\left\(r,y^\{\(t\)\}\\right\),\\qquad\\alpha\_\{t\}\>0,with fixed task weightsαt\>0\\alpha\_\{t\}\>0\.
#### Link to the learned score\.
The analysis is stated for deterministic oracle scoresr∈ℛB↑r\\in\\mathcal\{R\}\_\{B\}^\{\\uparrow\}\. The learned score vectors=\(s1,…,sn\)⊤s=\(s\_\{1\},\\dots,s\_\{n\}\)^\{\\top\}, withsi=sθ\(Xi\)s\_\{i\}=s\_\{\\theta\}\(X\_\{i\}\), is a noisy, data\-driven proxy forrr\. Appendix[B\.6](https://arxiv.org/html/2606.17450#A2.Thmtheorem6)shows that whenKc\(s\)K\_\{c\}\(s\)concentrates around a nondegenerate mean centered kernel, the realized learned\-scorenHSIC\(s,y\(t\)\)\\mathrm\{nHSIC\}\(s,y^\{\(t\)\}\)is close to the normalized HSIC computed from that averaged centered kernel and the task labels\.
#### From binary labels to an admission\-level label\-alignment matrix\.
Assume each included task is nonconstant\. Define the centered label profileℓ\(t\):=Hy\(t\)\.\\ell^\{\(t\)\}:=Hy^\{\(t\)\}\.Then‖ℓ\(t\)‖2\>0\\\|\\ell^\{\(t\)\}\\\|\_\{2\}\>0\. For binary labels under the delta kernel, Appendix[B\.1](https://arxiv.org/html/2606.17450#A2.Thmtheorem1)shows that each centered label Gram matrix has the rank\-one formLt,c=2ℓ\(t\)ℓ\(t\)⊤\.L\_\{t,c\}=2\\ell^\{\(t\)\}\\ell^\{\(t\)\\top\}\.Therefore,‖Lt,c‖F=2‖ℓ\(t\)‖22\.\\\|L\_\{t,c\}\\\|\_\{F\}=2\\\|\\ell^\{\(t\)\}\\\|\_\{2\}^\{2\}\.Whenever‖Kc\(r\)‖F\>0\\\|K\_\{c\}\(r\)\\\|\_\{F\}\>0,
nHSIC\(r,y\(t\)\)=ℓ\(t\)⊤Kc\(r\)ℓ\(t\)‖Kc\(r\)‖F‖ℓ\(t\)‖22\.\\mathrm\{nHSIC\}\\\!\\left\(r,y^\{\(t\)\}\\right\)=\\frac\{\\ell^\{\(t\)\\top\}K\_\{c\}\(r\)\\ell^\{\(t\)\}\}\{\\\|K\_\{c\}\(r\)\\\|\_\{F\}\\\|\\ell^\{\(t\)\}\\\|\_\{2\}^\{2\}\}\.When‖Kc\(r\)‖F=0\\\|K\_\{c\}\(r\)\\\|\_\{F\}=0, we keep the conventionnHSIC\(r,y\(t\)\)=0\\mathrm\{nHSIC\}\(r,y^\{\(t\)\}\)=0\. So each task contributes a quadratic form ofKc\(r\)K\_\{c\}\(r\)along its centered label vectorℓ\(t\)\\ell^\{\(t\)\}\. Including the fixed task coefficientαt\\alpha\_\{t\}, whenever‖Kc\(r\)‖F\>0\\\|K\_\{c\}\(r\)\\\|\_\{F\}\>0, the multi\-task objective becomes
J\(r\)=1‖Kc\(r\)‖F∑t=1Tαt‖ℓ\(t\)‖22ℓ\(t\)⊤Kc\(r\)ℓ\(t\)\.J\(r\)=\\frac\{1\}\{\\\|K\_\{c\}\(r\)\\\|\_\{F\}\}\\sum\_\{t=1\}^\{T\}\\frac\{\\alpha\_\{t\}\}\{\\\|\\ell^\{\(t\)\}\\\|\_\{2\}^\{2\}\}\\ell^\{\(t\)\\top\}K\_\{c\}\(r\)\\ell^\{\(t\)\}\.
The purpose of the following stacking is to establish the equality between this weighted sum of task\-wise quadratic forms and a single Frobenius inner product\. To simplify notation, combine the fixed task coefficient with the nHSIC normalization by definingβt:=αt/‖ℓ\(t\)‖22\\beta\_\{t\}:=\\alpha\_\{t\}/\\\|\\ell^\{\(t\)\}\\\|\_\{2\}^\{2\}\. Now form the stacked label\-profile matrixW~∈ℝT×n\\widetilde\{W\}\\in\\mathbb\{R\}^\{T\\times n\}, whosett\-th row is the weighted centered label vectorβtℓ\(t\)⊤\.\\sqrt\{\\beta\_\{t\}\}\\,\\ell^\{\(t\)\\top\}\.Equivalently,
W~=\[β1ℓ\(1\)⊤β2ℓ\(2\)⊤⋮βTℓ\(T\)⊤\]\.\\widetilde\{W\}=\\begin\{bmatrix\}\\sqrt\{\\beta\_\{1\}\}\\ell^\{\(1\)\\top\}\\\\ \\sqrt\{\\beta\_\{2\}\}\\ell^\{\(2\)\\top\}\\\\ \\vdots\\\\ \\sqrt\{\\beta\_\{T\}\}\\ell^\{\(T\)\\top\}\\end\{bmatrix\}\.
By construction,W~⊤W~=∑t=1Tβtℓ\(t\)ℓ\(t\)⊤\.\\widetilde\{W\}^\{\\top\}\\widetilde\{W\}=\\sum\_\{t=1\}^\{T\}\\beta\_\{t\}\\,\\ell^\{\(t\)\}\\ell^\{\(t\)\\top\}\.Therefore,⟨Kc\(r\),W~⊤W~⟩F=∑t=1Tβt⟨Kc\(r\),ℓ\(t\)ℓ\(t\)⊤⟩F\.\\left\\langle K\_\{c\}\(r\),\\widetilde\{W\}^\{\\top\}\\widetilde\{W\}\\right\\rangle\_\{F\}=\\sum\_\{t=1\}^\{T\}\\beta\_\{t\}\\left\\langle K\_\{c\}\(r\),\\ell^\{\(t\)\}\\ell^\{\(t\)\\top\}\\right\\rangle\_\{F\}\.Since⟨Kc\(r\),ℓ\(t\)ℓ\(t\)⊤⟩F=ℓ\(t\)⊤Kc\(r\)ℓ\(t\),\\left\\langle K\_\{c\}\(r\),\\ell^\{\(t\)\}\\ell^\{\(t\)\\top\}\\right\\rangle\_\{F\}=\\ell^\{\(t\)\\top\}K\_\{c\}\(r\)\\ell^\{\(t\)\},we obtain
⟨Kc\(r\),W~⊤W~⟩F=∑t=1Tαt‖ℓ\(t\)‖22ℓ\(t\)⊤Kc\(r\)ℓ\(t\)\.\\left\\langle K\_\{c\}\(r\),\\widetilde\{W\}^\{\\top\}\\widetilde\{W\}\\right\\rangle\_\{F\}=\\sum\_\{t=1\}^\{T\}\\frac\{\\alpha\_\{t\}\}\{\\\|\\ell^\{\(t\)\}\\\|\_\{2\}^\{2\}\}\\ell^\{\(t\)\\top\}K\_\{c\}\(r\)\\ell^\{\(t\)\}\.Hence, for‖Kc\(r\)‖F\>0\\\|K\_\{c\}\(r\)\\\|\_\{F\}\>0, the multi\-task objective can be written as
J\(r\)=⟨Kc\(r\),W~⊤W~⟩F‖Kc\(r\)‖F\.J\(r\)=\\frac\{\\langle K\_\{c\}\(r\),\\widetilde\{W\}^\{\\top\}\\widetilde\{W\}\\rangle\_\{F\}\}\{\\\|K\_\{c\}\(r\)\\\|\_\{F\}\}\.\(4\)When‖Kc\(r\)‖F=0\\\|K\_\{c\}\(r\)\\\|\_\{F\}=0, we keep the conventionJ\(r\)=0J\(r\)=0\.
The matrixW~⊤W~\\widetilde\{W\}^\{\\top\}\\widetilde\{W\}is a cross\-task label\-alignment matrix: its\(i,j\)\(i,j\)\-entry compares admissionsiiandjjusing all centered label profiles across tasks,\(W~⊤W~\)ij=∑t=1Tβtℓi\(t\)ℓj\(t\)\.\(\\widetilde\{W\}^\{\\top\}\\widetilde\{W\}\)\_\{ij\}=\\sum\_\{t=1\}^\{T\}\\beta\_\{t\}\\,\\ell\_\{i\}^\{\(t\)\}\\ell\_\{j\}^\{\(t\)\}\.Thus it summarizes how similarly two admissions behave across all weighted outcomes\. The scorerrenters only throughKc\(r\)K\_\{c\}\(r\), and the objective asks this score\-similarity matrix to align with the combined cross\-task label structure\. We then ask whether this combined label structure is mainly driven by one unified admission direction\.
#### Rank\-one projection to a shared admission direction\.
LetW~1=σ1uv⊤\\widetilde\{W\}\_\{1\}=\\sigma\_\{1\}uv^\{\\top\}be the best rank\-one approximation ofW~\\widetilde\{W\}, with‖u‖2=‖v‖2=1\\\|u\\\|\_\{2\}=\\\|v\\\|\_\{2\}=1\. Hereu∈ℝTu\\in\\mathbb\{R\}^\{T\}is the task\-side direction andv∈ℝnv\\in\\mathbb\{R\}^\{n\}is the admission\-level direction\. Whenσ1\>0\\sigma\_\{1\}\>0,vvis a leading eigenvector ofW~⊤W~\\widetilde\{W\}^\{\\top\}\\widetilde\{W\}, with eigenvalueσ12\\sigma\_\{1\}^\{2\}, and captures the dominant admission\-level pattern shared across tasks\. Ifσ1=0\\sigma\_\{1\}=0, thenW~=0\\widetilde\{W\}=0, and the projected objective below is identically zero\. Define the projected objective
J1\(r\):=⟨Kc\(r\),W~1⊤W~1⟩F‖Kc\(r\)‖F\.J\_\{1\}\(r\):=\\frac\{\\langle K\_\{c\}\(r\),\\widetilde\{W\}\_\{1\}^\{\\top\}\\widetilde\{W\}\_\{1\}\\rangle\_\{F\}\}\{\\\|K\_\{c\}\(r\)\\\|\_\{F\}\}\.We setJ1\(r\)=0J\_\{1\}\(r\)=0whenever‖Kc\(r\)‖F=0\\\|K\_\{c\}\(r\)\\\|\_\{F\}=0\.
###### Lemma 6\.1\(Reduction under rank\-one projection\)\.
For everyr∈ℛB↑r\\in\\mathcal\{R\}\_\{B\}^\{\\uparrow\}with‖Kc\(r\)‖F\>0\\\|K\_\{c\}\(r\)\\\|\_\{F\}\>0,J1\(r\)=σ12Δ¯v\(r\),J\_\{1\}\(r\)=\\sigma\_\{1\}^\{2\}\\bar\{\\Delta\}\_\{v\}\(r\),where
Δ¯v\(r\):=v⊤Kc\(r\)v‖Kc\(r\)‖F\.\\bar\{\\Delta\}\_\{v\}\(r\):=\\frac\{v^\{\\top\}K\_\{c\}\(r\)v\}\{\\\|K\_\{c\}\(r\)\\\|\_\{F\}\}\.Therefore, ifσ1\>0\\sigma\_\{1\}\>0,argmaxr∈ℛB↑:‖Kc\(r\)‖F\>0J1\(r\)=argmaxr∈ℛB↑:‖Kc\(r\)‖F\>0Δ¯v\(r\)\.\\arg\\max\_\{r\\in\\mathcal\{R\}\_\{B\}^\{\\uparrow\}:\\,\\\|K\_\{c\}\(r\)\\\|\_\{F\}\>0\}J\_\{1\}\(r\)=\\arg\\max\_\{r\\in\\mathcal\{R\}\_\{B\}^\{\\uparrow\}:\\,\\\|K\_\{c\}\(r\)\\\|\_\{F\}\>0\}\\bar\{\\Delta\}\_\{v\}\(r\)\.Ifσ1=0\\sigma\_\{1\}=0, thenJ1≡0J\_\{1\}\\equiv 0\.
###### Proof sketch\.
SinceW~1=σ1uv⊤\\widetilde\{W\}\_\{1\}=\\sigma\_\{1\}uv^\{\\top\}with‖u‖2=1\\\|u\\\|\_\{2\}=1,W~1⊤W~1=σ12vv⊤\.\\widetilde\{W\}\_\{1\}^\{\\top\}\\widetilde\{W\}\_\{1\}=\\sigma\_\{1\}^\{2\}vv^\{\\top\}\.Substituting this intoJ1\(r\)J\_\{1\}\(r\)gives
J1\(r\)=σ12⟨Kc\(r\),vv⊤⟩F‖Kc\(r\)‖F=σ12v⊤Kc\(r\)v‖Kc\(r\)‖F\.J\_\{1\}\(r\)=\\sigma\_\{1\}^\{2\}\\frac\{\\langle K\_\{c\}\(r\),vv^\{\\top\}\\rangle\_\{F\}\}\{\\\|K\_\{c\}\(r\)\\\|\_\{F\}\}=\\sigma\_\{1\}^\{2\}\\frac\{v^\{\\top\}K\_\{c\}\(r\)v\}\{\\\|K\_\{c\}\(r\)\\\|\_\{F\}\}\.The maximizer statement follows because multiplication by the positive constantσ12\\sigma\_\{1\}^\{2\}does not change the argmax\. Ifσ1=0\\sigma\_\{1\}=0, thenW~1⊤W~1=0\\widetilde\{W\}\_\{1\}^\{\\top\}\\widetilde\{W\}\_\{1\}=0, soJ1≡0J\_\{1\}\\equiv 0\. ∎
Thus, under a rank\-one shared label structure, the projected multi\-task objective reduces to aligning the centered score Gram matrix withvv\. This same direction determines the monotone admission\-level cutoff whose centered step vector is most aligned with the shared signal\.
#### Threshold certificate from the shared direction\.
For a splitj∈\{1,…,n−1\}j\\in\\\{1,\\dots,n\-1\\\}, defineLj:=\{1,…,j\},Rj:=\{j\+1,…,n\}\.L\_\{j\}:=\\\{1,\\dots,j\\\},\\qquad R\_\{j\}:=\\\{j\+1,\\dots,n\\\}\.Letgj:=H𝟏Rj\.g\_\{j\}:=H\\mathbf\{1\}\_\{R\_\{j\}\}\.Heregjg\_\{j\}is the centered indicator of the candidate upper\-tail group\.
###### Lemma 6\.2\(Global bound and threshold certificate\)\.
Fix any centered target directionw∈ℝnw\\in\\mathbb\{R\}^\{n\}, so𝟏⊤w=0\\mathbf\{1\}^\{\\top\}w=0\. This lemma is stated for a general directionww; after the lemma we apply it to the centered shared admission directionvv\. Define
Δ¯w\(r\):=w⊤Kc\(r\)w‖Kc\(r\)‖F,\\bar\{\\Delta\}\_\{w\}\(r\):=\\frac\{w^\{\\top\}K\_\{c\}\(r\)w\}\{\\\|K\_\{c\}\(r\)\\\|\_\{F\}\},with value0whenever‖Kc\(r\)‖F=0\\\|K\_\{c\}\(r\)\\\|\_\{F\}=0\. Ifw=0w=0, thenΔ¯w≡0\\bar\{\\Delta\}\_\{w\}\\equiv 0, and the statements are trivial\. Assumew≠0w\\neq 0below\. ThenΔ¯w\(r\)≤‖w‖22\\bar\{\\Delta\}\_\{w\}\(r\)\\leq\\\|w\\\|\_\{2\}^\{2\}for allr∈ℛB↑r\\in\\mathcal\{R\}\_\{B\}^\{\\uparrow\}\. Moreover, consider a nontrivial two\-level threshold score on the splitLj/RjL\_\{j\}/R\_\{j\}, meaning a scorer\(j\)∈ℛB↑r^\{\(j\)\}\\in\\mathcal\{R\}\_\{B\}^\{\\uparrow\}of the form
ri\(j\)=\{a,i∈Lj,b,i∈Rj,witha<b\.r\_\{i\}^\{\(j\)\}=\\begin\{cases\}a,&i\\in L\_\{j\},\\\\ b,&i\\in R\_\{j\},\\end\{cases\}\\qquad\\text\{with \}a<b\.That is, all admissions below the cutoff receive one score level, and all admissions above the cutoff receive another score level\. For any such threshold score, we have the exact value
Δ¯w\(r\(j\)\)=‖w‖22ρj2,ρj:=w⊤gj‖w‖2‖gj‖2\.\\bar\{\\Delta\}\_\{w\}\(r^\{\(j\)\}\)=\\\|w\\\|\_\{2\}^\{2\}\\rho\_\{j\}^\{2\},\\qquad\\rho\_\{j\}:=\\frac\{w^\{\\top\}g\_\{j\}\}\{\\\|w\\\|\_\{2\}\\\|g\_\{j\}\\\|\_\{2\}\}\.Thus the best two\-level threshold is obtained byj⋆∈argmax1≤j≤n−1ρj2\.j^\{\\star\}\\in\\arg\\max\_\{1\\leq j\\leq n\-1\}\\rho\_\{j\}^\{2\}\.The achieved value certifies a fractionρ⋆2:=max1≤j≤n−1ρj2\\rho\_\{\\star\}^\{2\}:=\\max\_\{1\\leq j\\leq n\-1\}\\rho\_\{j\}^\{2\}of the global upper bound‖w‖22\\\|w\\\|\_\{2\}^\{2\}\.
###### Proof sketch\.
Becausewwis centered,w⊤Kc\(r\)w=w⊤K\(r\)ww^\{\\top\}K\_\{c\}\(r\)w=w^\{\\top\}K\(r\)w, andw⊤Kc\(r\)w=⟨Kc\(r\),ww⊤⟩F\.w^\{\\top\}K\_\{c\}\(r\)w=\\langle K\_\{c\}\(r\),ww^\{\\top\}\\rangle\_\{F\}\.By Cauchy–Schwarz inequality,⟨Kc\(r\),ww⊤⟩F≤‖Kc\(r\)‖F‖ww⊤‖F=‖Kc\(r\)‖F‖w‖22\.\\langle K\_\{c\}\(r\),ww^\{\\top\}\\rangle\_\{F\}\\leq\\\|K\_\{c\}\(r\)\\\|\_\{F\}\\\|ww^\{\\top\}\\\|\_\{F\}=\\\|K\_\{c\}\(r\)\\\|\_\{F\}\\\|w\\\|\_\{2\}^\{2\}\.Dividing by‖Kc\(r\)‖F\\\|K\_\{c\}\(r\)\\\|\_\{F\}gives the global bound whenever‖Kc\(r\)‖F\>0\\\|K\_\{c\}\(r\)\\\|\_\{F\}\>0; the zero\-norm case follows from the convention\. For a nontrivial two\-level threshold scorer\(j\)r^\{\(j\)\},K\(r\(j\)\)K\(r^\{\(j\)\}\)is block\-constant\. Letc0:=κ\(0\),c1:=κ\(\|a−b\|\)\.c\_\{0\}:=\\kappa\(0\),\\qquad c\_\{1\}:=\\kappa\(\|a\-b\|\)\.The valuec0c\_\{0\}is the within\-block similarity, andc1c\_\{1\}is the across\-block similarity\. Sincea<ba<bandκ\\kappais strictly decreasing, we havec0−c1\>0\.c\_\{0\}\-c\_\{1\}\>0\.After centering, the constant all\-ones component disappears and only the block contrast remains\. HenceKc\(r\(j\)\)=2\(c0−c1\)gjgj⊤\.K\_\{c\}\(r^\{\(j\)\}\)=2\(c\_\{0\}\-c\_\{1\}\)g\_\{j\}g\_\{j\}^\{\\top\}\.Equivalently, writeKc\(r\(j\)\)=cjgjgj⊤,cj:=2\(c0−c1\)\>0\.K\_\{c\}\(r^\{\(j\)\}\)=c\_\{j\}\\,g\_\{j\}g\_\{j\}^\{\\top\},\\qquad c\_\{j\}:=2\(c\_\{0\}\-c\_\{1\}\)\>0\.Therefore,
Δ¯w\(r\(j\)\)=cj\(w⊤gj\)2cj‖gjgj⊤‖F=\(w⊤gj\)2‖gj‖22=‖w‖22\(w⊤gj‖w‖2‖gj‖2\)2=‖w‖22ρj2\.∎\\begin\{aligned\} \\bar\{\\Delta\}\_\{w\}\(r^\{\(j\)\}\)&=\\frac\{c\_\{j\}\(w^\{\\top\}g\_\{j\}\)^\{2\}\}\{c\_\{j\}\\\|g\_\{j\}g\_\{j\}^\{\\top\}\\\|\_\{F\}\}=\\frac\{\(w^\{\\top\}g\_\{j\}\)^\{2\}\}\{\\\|g\_\{j\}\\\|\_\{2\}^\{2\}\}\\\\ &=\\\|w\\\|\_\{2\}^\{2\}\\left\(\\frac\{w^\{\\top\}g\_\{j\}\}\{\\\|w\\\|\_\{2\}\\\|g\_\{j\}\\\|\_\{2\}\}\\right\)^\{2\}=\\\|w\\\|\_\{2\}^\{2\}\\rho\_\{j\}^\{2\}\.\\end\{aligned\}\\qed
Assumeσ1\>0\\sigma\_\{1\}\>0\. Applying Lemma[6\.2](https://arxiv.org/html/2606.17450#S6.Thmtheorem2)to the shared admission direction turnsvvinto an explicit monotone admission\-level cutoff\. The lemma requires a centered direction\. Since the rows ofW~\\widetilde\{W\}are centered,W~𝟏=0\\widetilde\{W\}\\mathbf\{1\}=0\. Thus, any right singular vector associated with a positive singular value is orthogonal to𝟏\\mathbf\{1\}, and therefore satisfiesHv=vHv=v\. Hence we may takew=vw=v\. The rank\-one shared direction gives the explicit cutoff
j⋆∈argmax1≤j≤n−1\(v⊤gj‖v‖2‖gj‖2\)2\.j^\{\\star\}\\in\\arg\\max\_\{1\\leq j\\leq n\-1\}\\left\(\\frac\{v^\{\\top\}g\_\{j\}\}\{\\\|v\\\|\_\{2\}\\\|g\_\{j\}\\\|\_\{2\}\}\\right\)^\{2\}\.
#### Approximate rank\-one structure\.
The previous reduction is exact for the projected matrixW~1\\widetilde\{W\}\_\{1\}\. In real data, outcomes also contain task\-specific variation, soW~\\widetilde\{W\}is generally only approximately rank one\. We therefore compare the original objective to the projected rank\-one objective using anε0\\varepsilon\_\{0\}\-floored denominator
Jε0\(r\)\\displaystyle J\_\{\\varepsilon\_\{0\}\}\(r\):=⟨Kc\(r\),W~⊤W~⟩Fmax\{‖Kc\(r\)‖F,ε0\},\\displaystyle=\\frac\{\\langle K\_\{c\}\(r\),\\widetilde\{W\}^\{\\top\}\\widetilde\{W\}\\rangle\_\{F\}\}\{\\max\\\{\\\|K\_\{c\}\(r\)\\\|\_\{F\},\\varepsilon\_\{0\}\\\}\},J1,ε0\(r\)\\displaystyle J\_\{1,\\varepsilon\_\{0\}\}\(r\):=⟨Kc\(r\),W~1⊤W~1⟩Fmax\{‖Kc\(r\)‖F,ε0\}\.\\displaystyle=\\frac\{\\langle K\_\{c\}\(r\),\\widetilde\{W\}\_\{1\}^\{\\top\}\\widetilde\{W\}\_\{1\}\\rangle\_\{F\}\}\{\\max\\\{\\\|K\_\{c\}\(r\)\\\|\_\{F\},\\varepsilon\_\{0\}\\\}\}\.
###### Theorem 6\.3\(Uniform approximation and near\-optimality transfer\)\.
Define the exact uniform gapΔgap:=supr∈ℛB↑\|Jε0\(r\)−J1,ε0\(r\)\|\.\\Delta\_\{\\mathrm\{gap\}\}:=\\sup\_\{r\\in\\mathcal\{R\}\_\{B\}^\{\\uparrow\}\}\|J\_\{\\varepsilon\_\{0\}\}\(r\)\-J\_\{1,\\varepsilon\_\{0\}\}\(r\)\|\.Letεgap\\varepsilon\_\{\\mathrm\{gap\}\}denote the explicit upper bound from Appendix[B\.4](https://arxiv.org/html/2606.17450#A2.Thmtheorem4), so thatΔgap≤εgap\.\\Delta\_\{\\mathrm\{gap\}\}\\leq\\varepsilon\_\{\\mathrm\{gap\}\}\.Ifr1⋆∈argmaxr∈ℛB↑J1,ε0\(r\)r\_\{1\}^\{\\star\}\\in\\arg\\max\_\{r\\in\\mathcal\{R\}\_\{B\}^\{\\uparrow\}\}J\_\{1,\\varepsilon\_\{0\}\}\(r\), thenJε0\(r1⋆\)≥supr∈ℛB↑Jε0\(r\)−2εgap\.J\_\{\\varepsilon\_\{0\}\}\(r\_\{1\}^\{\\star\}\)\\geq\\sup\_\{r\\in\\mathcal\{R\}\_\{B\}^\{\\uparrow\}\}J\_\{\\varepsilon\_\{0\}\}\(r\)\-2\\varepsilon\_\{\\mathrm\{gap\}\}\.
###### Proof sketch\.
By definition ofΔgap\\Delta\_\{\\mathrm\{gap\}\}, for allr∈ℛB↑r\\in\\mathcal\{R\}\_\{B\}^\{\\uparrow\},Jε0\(r\)≥J1,ε0\(r\)−Δgap,J1,ε0\(r\)≥Jε0\(r\)−Δgap\.J\_\{\\varepsilon\_\{0\}\}\(r\)\\geq J\_\{1,\\varepsilon\_\{0\}\}\(r\)\-\\Delta\_\{\\mathrm\{gap\}\},\\qquad J\_\{1,\\varepsilon\_\{0\}\}\(r\)\\geq J\_\{\\varepsilon\_\{0\}\}\(r\)\-\\Delta\_\{\\mathrm\{gap\}\}\.By optimality ofr1⋆r\_\{1\}^\{\\star\}forJ1,ε0J\_\{1,\\varepsilon\_\{0\}\}, for everyr∈ℛB↑r\\in\\mathcal\{R\}\_\{B\}^\{\\uparrow\},J1,ε0\(r1⋆\)≥J1,ε0\(r\)\.J\_\{1,\\varepsilon\_\{0\}\}\(r\_\{1\}^\{\\star\}\)\\geq J\_\{1,\\varepsilon\_\{0\}\}\(r\)\.Therefore, for everyr∈ℛB↑r\\in\\mathcal\{R\}\_\{B\}^\{\\uparrow\},
Jε0\(r1⋆\)\\displaystyle J\_\{\\varepsilon\_\{0\}\}\(r\_\{1\}^\{\\star\}\)≥J1,ε0\(r1⋆\)−Δgap≥J1,ε0\(r\)−Δgap\\displaystyle\\geq J\_\{1,\\varepsilon\_\{0\}\}\(r\_\{1\}^\{\\star\}\)\-\\Delta\_\{\\mathrm\{gap\}\}\\geq J\_\{1,\\varepsilon\_\{0\}\}\(r\)\-\\Delta\_\{\\mathrm\{gap\}\}≥Jε0\(r\)−2Δgap\.\\displaystyle\\geq J\_\{\\varepsilon\_\{0\}\}\(r\)\-2\\Delta\_\{\\mathrm\{gap\}\}\.Taking the supremum overr∈ℛB↑r\\in\\mathcal\{R\}\_\{B\}^\{\\uparrow\}and usingΔgap≤εgap\\Delta\_\{\\mathrm\{gap\}\}\\leq\\varepsilon\_\{\\mathrm\{gap\}\}, we obtainJε0\(r1⋆\)≥supr∈ℛB↑Jε0\(r\)−2εgap\.J\_\{\\varepsilon\_\{0\}\}\(r\_\{1\}^\{\\star\}\)\\geq\\sup\_\{r\\in\\mathcal\{R\}\_\{B\}^\{\\uparrow\}\}J\_\{\\varepsilon\_\{0\}\}\(r\)\-2\\varepsilon\_\{\\mathrm\{gap\}\}\.∎
Thus, when the explicit residual boundεgap\\varepsilon\_\{\\mathrm\{gap\}\}is small, a maximizer of the floored rank\-one projected objectiveJ1,ε0J\_\{1,\\varepsilon\_\{0\}\}is near\-optimal for the original floored multi\-task objectiveJε0J\_\{\\varepsilon\_\{0\}\}\.
## 7Experiments
Table 1:Distance Correlation between each model’s unified risk score and clinical outcomes\.Higher is better\. Bolded values highlight the best performer per outcome\.For readability, each column is scaled by the power of 10 shown in the header; divide by that factor to recover the original dCorr\.Table 2:Mutual Information between each model’s unified risk score and clinical outcomes\.Higher is better\. Bolded values highlight the best performer per outcome\.For readability, each column is scaled by the power of 10 shown in the header; divide by that factor to recover the original MI\.#### Datasets\.
We used two large, de\-identified real\-world EHR benchmark datasets:MIMIC\-IVandMIMIC\-III\. MIMIC\-IV\(Johnsonet al\.,[2023](https://arxiv.org/html/2606.17450#bib.bib43)\)contains546,028546\{,\}028hospital admissions from223,452223\{,\}452patients; we restrict to an ICD\-10 cohort \(admissions with at least one ICD\-10 diagnosis\), yielding254,377254\{,\}377admissions from122,905122\{,\}905patients in our final dataset\. MIMIC\-III\(Johnsonet al\.,[2016](https://arxiv.org/html/2606.17450#bib.bib44)\)contains58,97658\{,\}976admissions from46,52046\{,\}520patients and is ICU\-heavy:57,78657\{,\}786admissions include an ICU stay \(ratio0\.9800\.980\) compared to0\.1560\.156in MIMIC\-IV, which motivates a stricter late\-transfer ICU label for MIMIC\-III \(Appendix[A\.9](https://arxiv.org/html/2606.17450#A1.SS9)\)\.
Splits\.We use patient\-disjoint train/validation/test splits \(70%/10%/20%\) by hashingsubject\_id\. This yields, for MIMIC\-IV ICD\-10:178,178/25,228/50,971178\{,\}178/25\{,\}228/50\{,\}971admissions \(train/val/test\) and85,980/12,235/24,69085\{,\}980/12\{,\}235/24\{,\}690patients\. For MIMIC\-III:41,200/5,867/11,90941\{,\}200/5\{,\}867/11\{,\}909admissions and32,488/4,630/9,40232\{,\}488/4\{,\}630/9\{,\}402patients\.
Evaluation Metrics\.We evaluategeneral dependencebetween a*single*severity scoresi=sθ\(Xi\)s\_\{i\}=s\_\{\\theta\}\(X\_\{i\}\)and each binary outcome, to capture nonlinear score–outcome relationships\. We reportdistance correlation \(dCorr\)\(Székelyet al\.,[2007](https://arxiv.org/html/2606.17450#bib.bib45)\), which detects systematic linear or nonlinear dependence betweensis\_\{i\}andyi\(t\)y\_\{i\}^\{\(t\)\}, andmutual information \(MI\), which quantifies how much the score reduces uncertainty about the outcome\. Details are provided in Appendix[A\.10](https://arxiv.org/html/2606.17450#A1.SS10)\.
Clinical Outcomes\.We consider four clinically relevant binary outcomes: \(i\)in\-hospital mortality \(MORT\), defined by a recorded in\-hospital death time for the admission; \(ii\)30\-day mortality \(30M\), defined as all\-cause death within 30 days of admission using patient\-level date of death; \(iii\)length of stay \(LOS\), defined as a binary outcome representing length of stay\>7\>7days whenADMITTIMEandDISCHTIMEare available; and \(iv\)ICU transfer \(ICU\), defined from the first ICUINTIMErelative to admission\. For ICU transfer, MIMIC\-IV uses any ICU entry after admission, whereas MIMIC\-III uses first ICU entry\>24\>24hours after admission\. Details are provided in Appendix[A\.9](https://arxiv.org/html/2606.17450#A1.SS9)\.
Baselines\.We evaluate our model against several competitive single\-index baselines: \(i\)traditional clinical indices, including CCI\(Charlsonet al\.,[1987](https://arxiv.org/html/2606.17450#bib.bib3)\), van Walraven–weighted ECI\(Van Walravenet al\.,[2009](https://arxiv.org/html/2606.17450#bib.bib25)\), and a CCI\+ECI aggregate model; \(ii\)classical ML baselines, includingkkNN\(Choiet al\.,[2022](https://arxiv.org/html/2606.17450#bib.bib31)\), Naive Bayes\(Wolfsonet al\.,[2015](https://arxiv.org/html/2606.17450#bib.bib29)\), logistic regression\(Cowlinget al\.,[2021](https://arxiv.org/html/2606.17450#bib.bib26)\), factorization machines\(Yinet al\.,[2025](https://arxiv.org/html/2606.17450#bib.bib28)\), and gradient\-boosted trees\(Liet al\.,[2025](https://arxiv.org/html/2606.17450#bib.bib27)\); and \(iii\)deep learning baselines, including neural set/sequence models such as bag\-of\-codes MLPs\(Yinet al\.,[2025](https://arxiv.org/html/2606.17450#bib.bib28)\), attention\-based MIL pooling\(Harerimanaet al\.,[2021](https://arxiv.org/html/2606.17450#bib.bib33)\), and DeepSets\(Liuet al\.,[2020](https://arxiv.org/html/2606.17450#bib.bib32)\)\. Full details are provided in Appendices[A\.7](https://arxiv.org/html/2606.17450#A1.SS7)and[A\.8](https://arxiv.org/html/2606.17450#A1.SS8)\.
Experimental protocol and uncertainty\.Deep models use the same tokenization, patient\-disjoint splits, batch size, epochs, and validation\-based checkpoint selection\. For neural models, we report mean±\\pmSD over seeds\{11,101,1001\}\\\{11,101,1001\\\}; seeds were fixed, but CUDA/cuDNN was not forced to be bitwise deterministic, so exact reruns may show small variation\. Classical baselines are deterministic or use fixed randomness\. Metrics use valid\-label admissions only; details are provided in Appendix[A\.10](https://arxiv.org/html/2606.17450#A1.SS10)\.
Main results\.Across both datasets and dependence measures, MLCI is substantially more informative for mortality, while remaining competitive on other outcomes\. On both MIMIC\-IV/III, our method achieves thestrongest dependencefor both in\-hospital mortality and 30\-day mortality\. For LOS, our model isstill best or essentially tiedin both datasets \(MIMIC\-IV dCorr=51\.15±0\.88=51\.15\\pm 0\.88; MIMIC\-III dCorr=49\.84±1\.38=49\.84\\pm 1\.38; MI shows the same pattern\)\. For ICU transfer, our method isbeston MIMIC\-IV \(dCorr=61\.97±0\.64=61\.97\\pm 0\.64\), while on MIMIC\-III the strongest ICU\-transfer dependence is achieved by a DeepSets BCE baseline\. This is consistent with MIMIC\-III being ICU\-heavy and our ICU label capturing*late*transfers influenced by workflow/triage beyond a shared severity axis\.
## 8Theory Diagnostics
The theory motivates two empirical diagnostics for when a*single monotone severity score*captures shared structure across binary endpoints\. The first is the rank\-one energy ratioσ12/‖W~‖F2\\sigma\_\{1\}^\{2\}/\\\|\\widetilde\{W\}\\\|\_\{F\}^\{2\}, which measures how much of the stacked centered label structure is explained by one shared admission direction\. The second is the threshold scanj⋆∈argmaxj\(v⊤gj\)2/\(∥v∥22∥gj∥22\)j^\{\\star\}\\in\\arg\\max\_\{j\}\(v^\{\\top\}g\_\{j\}\)^\{2\}/\(\\\|v\\\|\_\{2\}^\{2\}\\\|g\_\{j\}\\\|\_\{2\}^\{2\}\), which tests whether this direction identifies a compact high\-severity upper\-tail group\. We compute these diagnostics on MIMIC\-IV and MIMIC\-III using the four clinical outcomes, restricted to admissions with all labels present\. We usen=5000n=5000uniformly sampled admissions for MIMIC\-IV and the full intersection\-valid MIMIC\-III cohort\(n=11909\)\(n=11909\), ordered by the learned scoressi=sθ\(Xi\)s\_\{i\}=s\_\{\\theta\}\(X\_\{i\}\)\. Reported diagnostics are from the original experimental runs\.
Diagnostics computed\.Letℓ\(t\):=Hy\(t\)\\ell^\{\(t\)\}:=Hy^\{\(t\)\}denote the centered label vector for tasktt\. For task weightsαt\>0\\alpha\_\{t\}\>0, form the stacked normalized label\-profile matrixW~∈ℝT×n\\widetilde\{W\}\\in\\mathbb\{R\}^\{T\\times n\}with rowsW~t,:=w~\(t\)⊤=αt‖ℓ\(t\)‖2ℓ\(t\)⊤\.\\widetilde\{W\}\_\{t,:\}=\\widetilde\{w\}^\{\(t\)\\top\}=\\frac\{\\sqrt\{\\alpha\_\{t\}\}\}\{\\\|\\ell^\{\(t\)\}\\\|\_\{2\}\}\\,\\ell^\{\(t\)\\top\}\.Equivalently,W~⊤W~=∑t=1Tαt‖ℓ\(t\)‖22ℓ\(t\)ℓ\(t\)⊤\.\\widetilde\{W\}^\{\\top\}\\widetilde\{W\}=\\sum\_\{t=1\}^\{T\}\\frac\{\\alpha\_\{t\}\}\{\\\|\\ell^\{\(t\)\}\\\|\_\{2\}^\{2\}\}\\ell^\{\(t\)\}\\ell^\{\(t\)\\top\}\.For the diagnostic results reported below, we setαt≡1\\alpha\_\{t\}\\equiv 1for all tasks\. Compute the SVDW~=UΣV⊤\\widetilde\{W\}=U\\Sigma V^\{\\top\}and define the shared admission directionv:=V:,1v:=V\_\{:,1\}\. We report: \(i\) the rank\-one energy ratioσ12/‖W~‖F2\\sigma\_\{1\}^\{2\}/\\\|\\widetilde\{W\}\\\|\_\{F\}^\{2\}, \(ii\) the full multi\-task objectiveJ\(s\)J\(s\)evaluated at the learned score vector, and \(iii\) the rank\-one projected objectiveJ1\(s\)=σ12Δ¯v\(s\)J\_\{1\}\(s\)=\\sigma\_\{1\}^\{2\}\\,\\bar\{\\Delta\}\_\{v\}\(s\)\.
Diagnostic 1: Rank\-one alignment across tasks\.Table[3](https://arxiv.org/html/2606.17450#S8.T3)shows thatW~\\widetilde\{W\}is*moderately*rank\-one aligned in both cohorts:σ12/‖W~‖F2=0\.49\\sigma\_\{1\}^\{2\}/\\\|\\widetilde\{W\}\\\|\_\{F\}^\{2\}=0\.49\(MIMIC\-IV\) and0\.450\.45\(MIMIC\-III\)\. The corresponding rank\-one surrogate captures a substantial portion of the learned\-score objective \(MIMIC\-IV:J1/J≈0\.64J\_\{1\}/J\\approx 0\.64; MIMIC\-III:J1/J≈0\.34J\_\{1\}/J\\approx 0\.34\), indicating a meaningful shared cross\-task component while allowing for additional task\-specific structure\.
Table 3:Rank\-one alignment and objective values on learned scores\.J\(s\)J\(s\)is the full multi\-task nHSIC objective evaluated on the learned score vector;J1\(s\)J\_\{1\}\(s\)is the rank\-one projected objective\.Diagnostic 2: Shared direction predicts the optimal monotone threshold split\.Lemma[6\.2](https://arxiv.org/html/2606.17450#S6.Thmtheorem2), applied to the shared directionvv, implies that the optimal split for the rank\-one projected objective within two\-level monotone threshold scores can be found by scanningρj2\(v\)=\(v⊤gj\)2‖v‖22‖gj‖22\\rho\_\{j\}^\{2\}\(v\)=\\frac\{\(v^\{\\top\}g\_\{j\}\)^\{2\}\}\{\\\|v\\\|\_\{2\}^\{2\}\\\|g\_\{j\}\\\|\_\{2\}^\{2\}\}, wheregj:=H𝟏\{j\+1,…,n\}g\_\{j\}:=H\\mathbf\{1\}\_\{\\\{j\+1,\\dots,n\\\}\}\. Hereρj\(v\)\\rho\_\{j\}\(v\)is the cosine similarity betweenvvand the centered step vectorgjg\_\{j\}, corresponding to a two\-level threshold split after indexjj\. Since both vectors are centered, this is equivalently their Pearson correlation\.
Table 4:Two\-level monotone threshold diagnostics\.jJj\_\{J\}is the best split for the full multi\-task threshold objective within the threshold family;jvj\_\{v\}maximizesρj2\(v\)\\rho\_\{j\}^\{2\}\(v\), equivalently selecting the best threshold for the rank\-one projected objective\.Table[4](https://arxiv.org/html/2606.17450#S8.T4)compares thevv\-selected split with the best split for the full multi\-task threshold objective\. In the original diagnostic analysis reported here, the two coincide in both datasets\. This supports the rank\-one diagnostic interpretation:vvidentifies a compact high\-severity upper\-tail subgroup\. In general, deviations betweenjvj\_\{v\}andjJj\_\{J\}would quantify residual task\-specific heterogeneity beyond the leading rank\-one component\.
Task\-wise heterogeneity\.Table[5](https://arxiv.org/html/2606.17450#S8.T5)shows that different endpoints relate to the learned score in different ways\. The per\-task nHSIC values measure continuous dependence with the learned score: in MIMIC\-III, this dependence is strongest for length of stay, whereas in MIMIC\-IV it is strongest for ICU transfer and length of stay\. The threshold certificates show a complementary pattern: mortality endpoints have strong upper\-tail threshold structure, especially in MIMIC\-IV\. Thus, some outcomes vary smoothly along the severity score, while others are mainly captured by a high\-risk cutoff, helping explain why a single monotone threshold can still support multiple endpoints\.
Table 5:Task\-wise diagnostics: \(top\) nHSIC of learned scores and \(bottom\) each task’s best threshold certificate from Lemma[6\.2](https://arxiv.org/html/2606.17450#S6.Thmtheorem2)\.
## 9Limitations
This study has several limitations\. First, MLCI assumes that different clinical outcomes share a common severity signal; this may be weaker when clinical outcomes are dominated by task\-specific variation\. Second, MLCI uses only diagnosis codes, which may be incomplete or noisy and may miss severity information contained in notes, labs, or vital signs\. Thus, it should not be interpreted as a complete physiological severity measure\. Third, some outcomes reflect care processes as well as patient severity; for example, ICU transfer can depend on bed availability, triage practice, and hospital workflow in addition to disease burden\.
## 10Conclusion
Across two EHR benchmarks, dependence results show that MLCI learns a severity score with strong general dependence on multiple clinical outcomes, while theory diagnostics indicate that centered admission\-level outcome profiles across these outcomes exhibit moderate rank\-one structure and that the induced shared admission directionvvyields a monotone threshold rule identifying compact high\-severity subgroups\. More broadly, MLCI shows how data\-driven comorbidity scoring can move beyond fixed linear indices by learning nonlinear severity structure while preserving the practical simplicity of a one\-dimensional clinical score\.
## Acknowledgements
This work was supported by NSF CAREER Award IIS\-2442159 \(B\.A\. and S\.B\.\) and NSF SCH\-2500344 \(K\.J\.\)\.
## Impact Statement
In this study, we introduce MLCI, a machine\-learned comorbidity index for more flexible and outcome\-aware severity scoring in electronic health records\. Unlike traditional indices such as Charlson and Elixhauser, which rely on fixed mortality\-centric weights, MLCI learns a single scalar severity score from diagnosis codes by maximizing normalized Hilbert–Schmidt Independence Criterion across multiple clinical endpoints\. This framework supports data\-driven severity summarization, multi\-outcome risk stratification, and retrospective risk adjustment while preserving the usability of a one\-dimensional clinical score\. Future work should extend MLCI to richer EHR modalities and prospectively validate it before clinical deployment\.
## References
- J\. Billings, J\. Dixon, T\. Mijanovich, and D\. Wennberg \(2006\)Case finding for patients at risk of readmission to hospital: development of algorithm to identify high risk patients\.Bmj333\(7563\),pp\. 327\.Cited by:[§1](https://arxiv.org/html/2606.17450#S1.p2.1)\.
- J\. E\. Byles, C\. D’Este, L\. Parkinson, R\. O’Connell, and C\. Treloar \(2005\)Single index of multimorbidity did not predict multiple outcomes\.Journal of clinical epidemiology58\(10\),pp\. 997–1005\.Cited by:[§1](https://arxiv.org/html/2606.17450#S1.p2.1)\.
- M\. E\. Charlson, P\. Pompei, K\. L\. Ales, and C\. R\. MacKenzie \(1987\)A new method of classifying prognostic comorbidity in longitudinal studies: development and validation\.Journal of chronic diseases40\(5\),pp\. 373–383\.Cited by:[§1](https://arxiv.org/html/2606.17450#S1.p1.1),[§2](https://arxiv.org/html/2606.17450#S2.p1.1),[§7](https://arxiv.org/html/2606.17450#S7.SS0.SSS0.Px1.p5.1)\.
- M\. H\. Choi, D\. Kim, E\. J\. Choi, Y\. J\. Jung, Y\. J\. Choi, J\. H\. Cho, and S\. H\. Jeong \(2022\)Mortality prediction of patients in intensive care units using machine learning algorithms based on electronic health records\.Scientific reports12\(1\),pp\. 7180\.Cited by:[§7](https://arxiv.org/html/2606.17450#S7.SS0.SSS0.Px1.p5.1)\.
- C\. Cortes, M\. Mohri, and A\. Rostamizadeh \(2012\)Algorithms for learning kernels based on centered alignment\.The Journal of Machine Learning Research13\(1\),pp\. 795–828\.Cited by:[§2](https://arxiv.org/html/2606.17450#S2.p1.1)\.
- T\. E\. Cowling, D\. A\. Cromwell, A\. Bellot, L\. D\. Sharples, and J\. van der Meulen \(2021\)Logistic regression and machine learning predicted patient mortality from large sets of diagnosis codes comparably\.Journal of Clinical Epidemiology133,pp\. 43–52\.Cited by:[§2](https://arxiv.org/html/2606.17450#S2.p1.1),[§7](https://arxiv.org/html/2606.17450#S7.SS0.SSS0.Px1.p5.1)\.
- Y\. Dai, D\. Wu, I\. Carroll, F\. Zou, and B\. Zou \(2025\)High\-dimensional biomarker identification for interpretable disease prediction via machine learning models\.Bioinformatics41\(5\),pp\. btaf266\.Cited by:[§2](https://arxiv.org/html/2606.17450#S2.p1.1)\.
- A\. Elixhauser, C\. Steiner, D\. R\. Harris, and R\. M\. Coffey \(1998\)Comorbidity measures for use with administrative data\.Medical care36\(1\),pp\. 8–27\.Cited by:[§1](https://arxiv.org/html/2606.17450#S1.p1.1),[§2](https://arxiv.org/html/2606.17450#S2.p1.1)\.
- G\. Ening, F\. Osterheld, D\. Capper, K\. Schmieder, and C\. Brenke \(2015\)Charlson comorbidity index: an additional prognostic parameter for preoperative glioblastoma patient stratification\.Journal of cancer research and clinical oncology141\(6\),pp\. 1131–1137\.Cited by:[§1](https://arxiv.org/html/2606.17450#S1.p1.1)\.
- A\. Gretton, O\. Bousquet, A\. Smola, and B\. Schölkopf \(2005\)Measuring statistical dependence with hilbert\-schmidt norms\.InInternational conference on algorithmic learning theory,pp\. 63–77\.Cited by:[§2](https://arxiv.org/html/2606.17450#S2.p1.1)\.
- A\. Gretton, K\. Fukumizu, C\. Teo, L\. Song, B\. Schölkopf, and A\. Smola \(2007\)A kernel statistical test of independence\.Advances in neural information processing systems20\.Cited by:[§2](https://arxiv.org/html/2606.17450#S2.p1.1)\.
- G\. Harerimana, J\. W\. Kim, and B\. Jang \(2021\)A deep attention model to forecast the length of stay and the in\-hospital mortality right on admission from icd codes and demographic data\.Journal of biomedical informatics118,pp\. 103778\.Cited by:[§2](https://arxiv.org/html/2606.17450#S2.p1.1),[§7](https://arxiv.org/html/2606.17450#S7.SS0.SSS0.Px1.p5.1)\.
- A\. E\. Johnson, L\. Bulgarelli, L\. Shen, A\. Gayles, A\. Shammout, S\. Horng, T\. J\. Pollard, S\. Hao, B\. Moody, B\. Gow,et al\.\(2023\)MIMIC\-iv, a freely accessible electronic health record dataset\.Scientific data10\(1\),pp\. 1\.Cited by:[§7](https://arxiv.org/html/2606.17450#S7.SS0.SSS0.Px1.p1.9)\.
- A\. E\. Johnson, T\. J\. Pollard, L\. Shen, L\. H\. Lehman, M\. Feng, M\. Ghassemi, B\. Moody, P\. Szolovits, L\. Anthony Celi, and R\. G\. Mark \(2016\)MIMIC\-iii, a freely accessible critical care database\.Scientific data3\(1\),pp\. 1–9\.Cited by:[§7](https://arxiv.org/html/2606.17450#S7.SS0.SSS0.Px1.p1.9)\.
- D\. E\. Katz, G\. Leibner, Y\. Esayag, N\. Kaufman, S\. Brammli\-Greenberg, and A\. J\. Rose \(2023\)Using the elixhauser risk adjustment model to predict outcomes among patients hospitalized in internal medicine at a large, tertiary\-care hospital in israel\.Israel Journal of Health Policy Research12\(1\),pp\. 32\.Cited by:[§4](https://arxiv.org/html/2606.17450#S4.p2.7)\.
- S\. H\. Kim, H\. S\. Choi, E\. S\. Jin, H\. Choi, H\. Lee, S\. Lee, C\. Y\. Lee, M\. G\. Lee, and Y\. Kim \(2021\)Predicting severe outcomes using national early warning score \(news\) in patients identified by a rapid response system: a retrospective cohort study\.Scientific reports11\(1\),pp\. 18021\.Cited by:[§1](https://arxiv.org/html/2606.17450#S1.p3.1)\.
- R\. T\. Kuswardhani, J\. Henrina, R\. Pranata, M\. A\. Lim, S\. Lawrensia, and K\. Suastika \(2020\)Charlson comorbidity index and a composite of poor outcomes in covid\-19 patients: a systematic review and meta\-analysis\.Diabetes & Metabolic Syndrome: Clinical Research & Reviews14\(6\),pp\. 2103–2109\.Cited by:[§1](https://arxiv.org/html/2606.17450#S1.p2.1)\.
- J\. R\. Le Gall, A\. Neumann, F\. Hemery, J\. P\. Bleriot, J\. P\. Fulgencio, B\. Garrigues, C\. Gouzes, E\. Lepage, P\. Moine, and D\. Villers \(2005\)Mortality prediction using saps ii: an update for french intensive care units\.Critical Care9\(6\),pp\. R645\.Cited by:[§1](https://arxiv.org/html/2606.17450#S1.p3.1)\.
- J\. Li, Y\. Sun, J\. Ren, Y\. Wu, and Z\. He \(2025\)Machine learning for in\-hospital mortality prediction in critically ill patients with acute heart failure: a retrospective analysis based on the mimic\-iv database\.Journal of Cardiothoracic and Vascular Anesthesia39\(3\),pp\. 666–674\.Cited by:[§2](https://arxiv.org/html/2606.17450#S2.p1.1),[§7](https://arxiv.org/html/2606.17450#S7.SS0.SSS0.Px1.p5.1)\.
- Y\. Li, R\. Pogodin, D\. J\. Sutherland, and A\. Gretton \(2021\)Self\-supervised learning with kernel dependence maximization\.Advances in Neural Information Processing Systems34,pp\. 15543–15556\.Cited by:[§2](https://arxiv.org/html/2606.17450#S2.p1.1)\.
- W\. Liu, C\. Stansbury, K\. Singh, A\. M\. Ryan, D\. Sukul, E\. Mahmoudi, A\. Waljee, J\. Zhu, and B\. K\. Nallamothu \(2020\)Predicting 30\-day hospital readmissions using artificial neural networks with medical code embedding\.PloS one15\(4\),pp\. e0221606\.Cited by:[§2](https://arxiv.org/html/2606.17450#S2.p1.1),[§7](https://arxiv.org/html/2606.17450#S7.SS0.SSS0.Px1.p5.1)\.
- K\. Ly, D\. Wakefield, and R\. ZuWallack \(2025\)The usefulness of charlson comorbidity index \(cci\) scoring in predicting all\-cause mortality in outpatients with clinical diagnoses of copd\.Journal of Multimorbidity and Comorbidity15,pp\. 26335565251315876\.Cited by:[§1](https://arxiv.org/html/2606.17450#S1.p3.1)\.
- A\. O’Hara, J\. Pozin, M\. Abourahma, R\. Gigstad, D\. Torres, B\. Knapp, B\. Kantarcioglu, J\. Fareed, and A\. Darki \(2024\)Charlson and elixhauser comorbidity indices for prediction of mortality and hospital readmission in patients with acute pulmonary embolism\.Clinical and Applied Thrombosis/Hemostasis30,pp\. 10760296241253844\.Cited by:[§1](https://arxiv.org/html/2606.17450#S1.p1.1)\.
- H\. Ou, B\. Mukherjee, S\. R\. Erickson, J\. D\. Piette, R\. P\. Bagozzi, and R\. Balkrishnan \(2012\)Comparative performance of comorbidity indices in predicting health care\-related behaviors and outcomes among medicaid enrollees with type 2 diabetes\.Population health management15\(4\),pp\. 220–229\.Cited by:[§1](https://arxiv.org/html/2606.17450#S1.p1.1)\.
- B\. S\. Patel, E\. Steinberg, S\. R\. Pfohl, and N\. H\. Shah \(2021\)Learning decision thresholds for risk stratification models from aggregate clinician behavior\.Journal of the American Medical Informatics Association28\(10\),pp\. 2258–2264\.Cited by:[§1](https://arxiv.org/html/2606.17450#S1.p2.1)\.
- H\. Quan, B\. Li, C\. M\. Couris, K\. Fushimi, P\. Graham, P\. Hider, J\. Januel, and V\. Sundararajan \(2011\)Updating and validating the charlson comorbidity index and score for risk adjustment in hospital discharge abstracts using data from 6 countries\.American journal of epidemiology173\(6\),pp\. 676–682\.Cited by:[§1](https://arxiv.org/html/2606.17450#S1.p1.1)\.
- H\. Quan, V\. Sundararajan, P\. Halfon, A\. Fong, B\. Burnand, J\. Luthi, L\. D\. Saunders, C\. A\. Beck, T\. E\. Feasby, and W\. A\. Ghali \(2005\)Coding algorithms for defining comorbidities in icd\-9\-cm and icd\-10 administrative data\.Medical care43\(11\),pp\. 1130–1139\.Cited by:[§D\.1](https://arxiv.org/html/2606.17450#A4.SS1.p1.4),[§3](https://arxiv.org/html/2606.17450#S3.SS0.SSS0.Px1.p1.3)\.
- G\. J\. Székely, M\. L\. Rizzo, and N\. K\. Bakirov \(2007\)Measuring and testing dependence by correlation of distances\.Cited by:[§7](https://arxiv.org/html/2606.17450#S7.SS0.SSS0.Px1.p3.3)\.
- Y\. Takahashi, M\. Ueki, M\. Yamada, G\. Tamiya, I\. N\. Motoike, D\. Saigusa, M\. Sakurai, F\. Nagami, S\. Ogishima, S\. Koshiba,et al\.\(2020\)Improved metabolomic data\-based prediction of depressive symptoms using nonlinear machine learning with feature selection\.Translational psychiatry10\(1\),pp\. 157\.Cited by:[§2](https://arxiv.org/html/2606.17450#S2.p1.1)\.
- C\. Van Walraven, P\. C\. Austin, A\. Jennings, H\. Quan, and A\. J\. Forster \(2009\)A modification of the elixhauser comorbidity measures into a point system for hospital death using administrative data\.Medical care47\(6\),pp\. 626–633\.Cited by:[§1](https://arxiv.org/html/2606.17450#S1.p1.1),[§2](https://arxiv.org/html/2606.17450#S2.p1.1),[§7](https://arxiv.org/html/2606.17450#S7.SS0.SSS0.Px1.p5.1)\.
- D\. Wei, Y\. Sun, R\. Chen, Y\. Meng, and W\. Wu \(2023\)Age\-adjusted charlson comorbidity index and in\-hospital mortality in critically ill patients with cardiogenic shock: a retrospective cohort study\.Experimental and Therapeutic Medicine25\(6\),pp\. 299\.Cited by:[§1](https://arxiv.org/html/2606.17450#S1.p3.1)\.
- T\. Willadsen, V\. Siersma, D\. Nicolaisdóttir, R\. Køster\-Rasmussen, D\. Jarbøl, S\. Reventlow, S\. Mercer, and N\. d\. F\. Olivarius \(2018\)Multimorbidity and mortality: a 15\-year longitudinal registry\-based nationwide danish population study\.Journal of comorbidity8\(1\),pp\. 2235042X18804063\.Cited by:[§1](https://arxiv.org/html/2606.17450#S1.p3.1)\.
- J\. Wolfson, S\. Bandyopadhyay, M\. Elidrisi, G\. Vazquez\-Benitez, D\. M\. Vock, D\. Musgrove, G\. Adomavicius, P\. E\. Johnson, and P\. J\. O’Connor \(2015\)A naive bayes machine learning approach to risk prediction using censored, time\-to\-event data\.Statistics in medicine34\(21\),pp\. 2941–2957\.Cited by:[§7](https://arxiv.org/html/2606.17450#S7.SS0.SSS0.Px1.p5.1)\.
- R\. Yin, J\. Li, Q\. Yang, X\. Chen, X\. Zhang, M\. Lin, J\. Bian, and A\. Subramaniam \(2025\)MTLNFM: a multi\-task framework using neural factorization machines to predict patient clinical outcomes\.Applied Sciences15\(15\),pp\. 8733\.Cited by:[§2](https://arxiv.org/html/2606.17450#S2.p1.1),[§7](https://arxiv.org/html/2606.17450#S7.SS0.SSS0.Px1.p5.1)\.
- Y\. Yu, W\. Zhang, D\. O’Gara, J\. Li, and S\. Chang \(2023\)A moment kernel machine for clinical data mining to inform medical decision making\.Scientific reports13\(1\),pp\. 10459\.Cited by:[§2](https://arxiv.org/html/2606.17450#S2.p1.1)\.
- M\. Zaheeret al\.\(2017\)Deep sets, in proceedings of the 31st international conference on neural information processing systems\.Long Beach, CA, USA\.Cited by:[§5](https://arxiv.org/html/2606.17450#S5.p3.7)\.
## Appendix
This appendix accompanies “a Machine\-Learned Comorbidity Index” and is organized as follows:
- •Appendix A: Implementation Details\.Data preprocessing, model architecture, training procedure, and hyperparameters\.
- •Appendix B: Multi\-task nHSIC Theory\.Theoretical results underpinning the multi\-task dependence objective\.
- •Appendix C: Additional Experiments\.Supplementary dependence evaluations and ablations\.
- •Appendix D: Comorbidity Indices\.Notation and scoring formulas for the Charlson and van Walraven Elixhauser comorbidity indices\.
- •Appendix E: Post\-hoc Risk Curves\.Estimation and visualization of outcome\-specific risk curves as a function of the learned score\.
## Appendix AImplementation Details
### A\.1Data processing and diagnosis tokenization
#### ICD normalization and prefixing\.
Each diagnosis code is converted to uppercase, punctuation and whitespace are removed, and the firstk=4k=4characters are used as the diagnosis token\. A vocabulary is built*from the training split only*and augmented with<PAD\>\(index 0\) and<UNK\>\(index 1\)\. Each admissioniiis represented as a variable\-length token collectionXi=\{Ci,1,…,Ci,mi\}X\_\{i\}=\\\{C\_\{i,1\},\\ldots,C\_\{i,m\_\{i\}\}\\\}, wheremim\_\{i\}is the number of retained diagnosis tokens after truncation andDmax=256D\_\{\\max\}=256is the maximum retained length\.
#### Train/val/test mapping\.
We build the vocabulary from the training diagnoses file, then map diagnoses into token collections separately for train/val/test using the shared training vocabulary\.
### A\.2Outcomes, missingness masks, and validity
#### Task set\.
We considerTTbinary outcomes, such as in\-hospital mortality, 30\-day mortality, length of stay, and ICU transfer, with possible missing labels\. For taskt∈\[T\]t\\in\[T\], the label for admissioniiis denotedyi\(t\)∈\{0,1\}y\_\{i\}^\{\(t\)\}\\in\\\{0,1\\\}\.
#### Mask definition\.
LetMi\(t\)∈\{0,1\}M\_\{i\}^\{\(t\)\}\\in\\\{0,1\\\}indicate whetheryi\(t\)y\_\{i\}^\{\(t\)\}is defined\. For most outcomes,Mi\(t\)=1M\_\{i\}^\{\(t\)\}=1if and only if the label value is non\-missing\. Forlong\_stay, validity is defined by a dedicated indicator columnlong\_stay\_defined; we set
Mi\(los\)=long\_stay\_definedM\_\{i\}^\{\(\\mathrm\{los\}\)\}=\\texttt\{long\\\_stay\\\_defined\}and treatlong\_staylabels outside this mask as invalid\. In storage, missing labels are filled with0, but are excluded from objectives and metrics using the mask\.
#### Intersection validity for MLCI\.
For MLCI’s multi\-task nHSIC training, validation, and model selection, we enforce intersection validity
Mi∩=∏t=1TMi\(t\),M\_\{i\}^\{\\cap\}=\\prod\_\{t=1\}^\{T\}M\_\{i\}^\{\(t\)\},\(5\)and compute every per\-task nHSIC contribution only on rows withMi∩=1M\_\{i\}^\{\\cap\}=1\. This ensures that all task contributions in the multi\-task nHSIC objective are evaluated on the same admission set\. Decomposable BCE baselines instead use task\-specific validity masks, as described below\.
### A\.3DeepSets architecture for a single scalar score
#### Embedding and token MLP\.
We use an embedding dimensiond=128d=128\. Each token embedding is transformed by an elementwise MLP
ϕ:Linear\(d,d\)→ReLU→Linear\(d,d\)→ReLU\.\\phi:\\ \\mathrm\{Linear\}\(d,d\)\\rightarrow\\mathrm\{ReLU\}\\rightarrow\\mathrm\{Linear\}\(d,d\)\\rightarrow\\mathrm\{ReLU\}\.
#### Pooling\.
Given masked token features\{hj\}\\\{h\_\{j\}\\\}, we compute masked mean pooling and masked max pooling, then concatenate and apply LayerNorm:
p=LN\(\[h¯;h^\]\)∈ℝ2d\.p=\\mathrm\{LN\}\(\[\\bar\{h\};\\hat\{h\}\]\)\\in\\mathbb\{R\}^\{2d\}\.
#### Score head\.
We mapppto a scalar via
ρ:Linear\(2d,d\)→ReLU→Linear\(d,d\)→ReLU→Linear\(d,1\),\\rho:\\ \\mathrm\{Linear\}\(2d,d\)\\rightarrow\\mathrm\{ReLU\}\\rightarrow\\mathrm\{Linear\}\(d,d\)\\rightarrow\\mathrm\{ReLU\}\\rightarrow\\mathrm\{Linear\}\(d,1\),producing the learned scoresi=sθ\(Xi\)∈ℝs\_\{i\}=s\_\{\\theta\}\(X\_\{i\}\)\\in\\mathbb\{R\}\.
### A\.4nHSIC computation
#### Kernels\.
On a mini\-batch of scores\{si\}i=1nb\\\{s\_\{i\}\\\}\_\{i=1\}^\{n\_\{b\}\}, we use an RBF score kernel
Kij\(b\)=exp\(−\(si−sj\)22σ2\)\.K\_\{ij\}^\{\(b\)\}=\\exp\\\!\\left\(\-\\frac\{\(s\_\{i\}\-s\_\{j\}\)^\{2\}\}\{2\\sigma^\{2\}\}\\right\)\.For tasktt, the binary label Gram matrix is
Lij\(b,t\)=𝕀\{yi\(t\)=yj\(t\)\}\.L\_\{ij\}^\{\(b,t\)\}=\\mathbb\{I\}\\\{y\_\{i\}^\{\(t\)\}=y\_\{j\}^\{\(t\)\}\\\}\.
#### Centering\.
Let
Hb:=I−1nb𝟏𝟏⊤H\_\{b\}:=I\-\\frac\{1\}\{n\_\{b\}\}\\mathbf\{1\}\\mathbf\{1\}^\{\\top\}be the mini\-batch centering matrix\. The centered mini\-batch Gram matrices are
Kc\(b\):=HbK\(b\)Hb,Lt,c\(b\):=HbL\(b,t\)Hb\.K\_\{c\}^\{\(b\)\}:=H\_\{b\}K^\{\(b\)\}H\_\{b\},\\qquad L\_\{t,c\}^\{\(b\)\}:=H\_\{b\}L^\{\(b,t\)\}H\_\{b\}\.Equivalently, this subtracts row and column means and adds the grand mean\.
#### Normalized criterion with floor\.
For tasktt, we compute
nHSIC^\(s,y\(t\)\)=⟨Kc\(b\),Lt,c\(b\)⟩Fmax\{‖Kc\(b\)‖F,ε0\}‖Lt,c\(b\)‖F\.\\widehat\{\\mathrm\{nHSIC\}\}\\\!\\left\(s,y^\{\(t\)\}\\right\)=\\frac\{\\langle K\_\{c\}^\{\(b\)\},L\_\{t,c\}^\{\(b\)\}\\rangle\_\{F\}\}\{\\max\\\{\\\|K\_\{c\}^\{\(b\)\}\\\|\_\{F\},\\varepsilon\_\{0\}\\\}\\,\\\|L\_\{t,c\}^\{\(b\)\}\\\|\_\{F\}\}\.We set the nHSIC floor to
ε0=NHSIC\_FLOOR=10−4\\varepsilon\_\{0\}=\\texttt\{NHSIC\\\_FLOOR\}=10^\{\-4\}during training\.
#### Per\-task masking and guards\.
For MLCI, each per\-task nHSIC term is computed on the intersection\-valid mini\-batch indices\{i:Mi∩=1\}\.\\\{i:M\_\{i\}^\{\\cap\}=1\\\}\.If fewer than 3 valid points are available in a batch, or if the valid labels for taskttare single\-class, we setnHSIC^\(s,y\(t\)\)=0\\widehat\{\\mathrm\{nHSIC\}\}\(s,y^\{\(t\)\}\)=0for that batch\.
### A\.5Bandwidth selection
We set the RBF bandwidthσ\\sigmaonce per run using a stabilized median heuristic on training scores\. Concretely, we run the current model on a fixed\-size training subsample, standardize the sampled scores to zero mean and unit variance, compute pairwise absolute differences on a deterministic subsample, and setσ\\sigmato the median of the nonzero differences\. We apply a lower bound of0\.050\.05to avoid degenerate micro\-kernel behavior when scores are nearly tied\.
This standardization is used only for selecting the bandwidth\. The learned score itself is not standardized in the nHSIC objective: training uses the scalar model outputs directly in
Kij\(b\)=exp\(−\(si−sj\)22σ2\)\.K\_\{ij\}^\{\(b\)\}=\\exp\\\!\\left\(\-\\frac\{\(s\_\{i\}\-s\_\{j\}\)^\{2\}\}\{2\\sigma^\{2\}\}\\right\)\.
### A\.6Optimization and training protocol
#### Optimizer and schedule\.
We use AdamW with learning rate10−310^\{\-3\}, batch size256256, and no weight decay\. Training uses two stages: Stage 1 trains single\-task models for 10 epochs each to estimate task weights; Stage 2 trains a multi\-task model for 10 epochs using the weighted objective\.
#### Two\-stage weights\.
Leth^t\\widehat\{h\}\_\{t\}be the best validation nHSIC for taskttin Stage 1, and let
h^max:=maxth^t\.\\widehat\{h\}\_\{\\max\}:=\\max\_\{t\}\\widehat\{h\}\_\{t\}\.We compute stabilized inverse\-strength weights
αt∝\(h^maxmax\(h^t,εwt\)\)γwt,\\alpha\_\{t\}\\propto\\left\(\\frac\{\\widehat\{h\}\_\{\\max\}\}\{\\max\(\\widehat\{h\}\_\{t\},\\varepsilon\_\{\\mathrm\{wt\}\}\)\}\\right\)^\{\\gamma\_\{\\mathrm\{wt\}\}\},with
εwt=0\.02,γwt=0\.25\.\\varepsilon\_\{\\mathrm\{wt\}\}=0\.02,\\qquad\\gamma\_\{\\mathrm\{wt\}\}=0\.25\.We clip weights to a bounded range with factor3\.03\.0and renormalize so the mean weight over active tasks is11\. The assumptionαt\>0\\alpha\_\{t\}\>0is intended for non\-degenerate tasks included in the objective\. If a task has no valid labels or only one observed class in the training split, it is not included in the empirical objective; setting its code weight to zero is only the implementation equivalent of omitting that task from the sum\.
#### Model selection\.
We select checkpoints by validation objective: Stage 1 selects the checkpoint maximizing single\-task validation nHSIC; Stage 2 selects the checkpoint maximizing the weighted validation sum\.
#### Random seeds and deterministic execution\.
For neural models, we fixed Python, NumPy, and PyTorch random seeds for each run and report mean±\\pmSD over seeds\{11,101,1001\}\\\{11,101,1001\\\}\. CUDA/cuDNN execution was not forced to be bitwise deterministic, so exact reruns may exhibit small numerical variation\. Reported neural results correspond to the runs with the specified seeds\.
#### Score orientation\.
After Stage 2 training, we orient the score using validation mortality\. If the Pearson correlation betweensis\_\{i\}andyi\(mort\)y\_\{i\}^\{\(\\mathrm\{mort\}\)\}on valid validation rows is negative, we multiply all scores by−1\-1\.
### A\.7Deep Learning Baselines
All deep learning baseline models use the*same*data processing, vocabulary construction, padding rules, train/validation/test splits, optimizer settings, and evaluation code described above\. They use task\-specific validity masks for BCE training and validation, as detailed below\. The only component that changes across baselines is the architecture, which maps an unordered multiset of diagnosis tokens for admissioniito a single scalar baseline scoresibase=fθ\(Xi\)s\_\{i\}^\{\\mathrm\{base\}\}=f\_\{\\theta\}\(X\_\{i\}\)\.
#### Single\-index multi\-outcome setting\.
Each admission produces one shared scalar baseline scoresibase∈ℝs\_\{i\}^\{\\mathrm\{base\}\}\\in\\mathbb\{R\}\. This scalar is reused across all outcomes and is not an outcome\-specific calibrated probability\. Missing labels are handled by the task maskMi\(t\)M\_\{i\}^\{\(t\)\}\(Section[A\.2](https://arxiv.org/html/2606.17450#A1.SS2)\); padded tokens never contribute to the pooled representation\.
#### Two\-stage training with weighted masked BCE\.
All BCE baselines use a two\-stage protocol\. In Stage 1, we train separate single\-task models, one per outcome, and select the best checkpoint by*minimum*masked validation BCE\. Stage 1 validation losses are converted into task weights that emphasize tasks that are harder under the shared score constraint\. In Stage 2, a new shared model is trained from scratch to minimize a normalized weighted masked BCE across outcomes:
ℒBCE=∑t=1TαtBCE∑iMi\(t\)BCEWithLogits\(sibase,yi\(t\)\)∑t=1TαtBCE∑iMi\(t\)\.\\mathcal\{L\}\_\{\\mathrm\{BCE\}\}=\\frac\{\\sum\_\{t=1\}^\{T\}\\alpha\_\{t\}^\{\\mathrm\{BCE\}\}\\sum\_\{i\}M\_\{i\}^\{\(t\)\}\\,\\mathrm\{BCEWithLogits\}\(s\_\{i\}^\{\\mathrm\{base\}\},y\_\{i\}^\{\(t\)\}\)\}\{\\sum\_\{t=1\}^\{T\}\\alpha\_\{t\}^\{\\mathrm\{BCE\}\}\\sum\_\{i\}M\_\{i\}^\{\(t\)\}\}\.
#### Weight computation and skipped tasks\.
Letℒ^t\\widehat\{\\mathcal\{L\}\}\_\{t\}denote the best Stage 1 validation BCE for outcomett\. We compute a nonnegative score
scoret=max\(log2−ℒ^t,0\)\\mathrm\{score\}\_\{t\}=\\max\(\\log 2\-\\widehat\{\\mathcal\{L\}\}\_\{t\},0\)and set baseline task weights using a capped inverse power law:
αtBCE=clip\(\(maxuscoreumax\(scoret,εwt\)\)γwt,1c,c\),\\alpha\_\{t\}^\{\\mathrm\{BCE\}\}=\\mathrm\{clip\}\\\!\\left\(\\left\(\\frac\{\\max\_\{u\}\\mathrm\{score\}\_\{u\}\}\{\\max\(\\mathrm\{score\}\_\{t\},\\varepsilon\_\{\\mathrm\{wt\}\}\)\}\\right\)^\{\\gamma\_\{\\mathrm\{wt\}\}\},\\frac\{1\}\{c\},c\\right\),with
εwt=0\.02,γwt=0\.25,c=3\.0\.\\varepsilon\_\{\\mathrm\{wt\}\}=0\.02,\\qquad\\gamma\_\{\\mathrm\{wt\}\}=0\.25,\\qquad c=3\.0\.Weights are then mean\-normalized over active tasks\. Outcomes that are single\-class, or have no valid labels in the training split, are skipped and assignedαtBCE=0\\alpha\_\{t\}^\{\\mathrm\{BCE\}\}=0, so they do not contribute to Stage 2 training\.
### A\.8Classical ML baselines
We implement several classical baselines on top of the same admission\-level ICD\-prefix representation and the same masking conventions as in Section[A\.2](https://arxiv.org/html/2606.17450#A1.SS2)\. For all classical models, diagnosis codes are normalized by converting to uppercase and removing punctuation and whitespace, then truncated to the firstk=4k=4characters to form ICD\-prefix tokens\. A prefix vocabulary is constructed*from the training split only*and augmented with a dedicated<UNK\>token for out\-of\-vocabulary prefixes\. Each admission is represented by at mostDmax=256D\_\{\\max\}=256retained diagnosis tokens in encounter order\. Unless otherwise noted, classical baselines use the resulting sparse bag\-of\-words features; when TF–IDF is applied, its statistics are fit on TRAIN and then applied unchanged to VAL/TEST\.
#### Label validity and masking\.
For each tasktt, we extract binary labelsyi\(t\)y\_\{i\}^\{\(t\)\}and a validity maskMi\(t\)M\_\{i\}^\{\(t\)\}, whereMi\(t\)=1M\_\{i\}^\{\(t\)\}=1denotes that the label is defined for admissionii\. For most tasks,Mi\(t\)=1M\_\{i\}^\{\(t\)\}=1if and only if the label entry is non\-missing; forlong\_stay, validity is given by the dedicated indicatorlong\_stay\_defined\. All training, validation objectives, and evaluation metrics for taskttuse only rows withMi\(t\)=1M\_\{i\}^\{\(t\)\}=1\.
#### DL\-style validation loss\.
To match the deep\-learning evaluation convention used elsewhere, we compute validation binary cross\-entropy as a mean over fixed\-size batches\. For each batch, we compute BCE on the subset of rows withMi\(t\)=1M\_\{i\}^\{\(t\)\}=1; if a batch contains zero valid labels, its contribution is defined as0, and we then average over batches\.
#### Two\-stage task weighting\.
For multi\-task fusion baselines, we derive per\-task weights from Stage 1 validation BCE\. Letℒ^t\\widehat\{\\mathcal\{L\}\}\_\{t\}denote the Stage 1 validation BCE for tasktt, computed with masking and batch\-averaging as above\. We convertℒ^t\\widehat\{\\mathcal\{L\}\}\_\{t\}to a nonnegative strength score
scoret=max\(log2−ℒ^t,0\),\\mathrm\{score\}\_\{t\}=\\max\(\\log 2\-\\widehat\{\\mathcal\{L\}\}\_\{t\},0\),and then set
αtBCE=clip\(\(maxuscoreumax\(scoret,εwt\)\)γwt,1c,c\)\.\\alpha\_\{t\}^\{\\mathrm\{BCE\}\}=\\mathrm\{clip\}\\\!\\left\(\\left\(\\frac\{\\max\_\{u\}\\mathrm\{score\}\_\{u\}\}\{\\max\(\\mathrm\{score\}\_\{t\},\\varepsilon\_\{\\mathrm\{wt\}\}\)\}\\right\)^\{\\gamma\_\{\\mathrm\{wt\}\}\},\\frac\{1\}\{c\},c\\right\)\.Tasks that are unusable on TRAIN, because they have too few valid labels or are single\-class on valid TRAIN rows, or have non\-finite validation loss, receiveαtBCE=0\\alpha\_\{t\}^\{\\mathrm\{BCE\}\}=0\. Active weights are normalized to have mean11over\{t:αtBCE\>0\}\\\{t:\\alpha\_\{t\}^\{\\mathrm\{BCE\}\}\>0\\\}\.
#### Stage\-1 per\-task models\.
We train per\-task models for logistic regression \(LR\), factorization machines \(FM\), gradient\-boosted trees \(GBT\),kk\-nearest neighbors \(kkNN\), and Naive Bayes \(NB\)\. Each Stage 1 model is trained on TRAIN rows withMi\(t\)=1M\_\{i\}^\{\(t\)\}=1for that task and produces an admission\-level score, logit, or marginai\(t\),basea\_\{i\}^\{\(t\),\\mathrm\{base\}\}for all admissions in TRAIN/VAL/TEST\. When a model outputs probabilities, such askkNN, we convert to logits by
ai\(t\),base=log\(p^i\(t\)1−p^i\(t\)\),a\_\{i\}^\{\(t\),\\mathrm\{base\}\}=\\log\\\!\\left\(\\frac\{\\hat\{p\}\_\{i\}^\{\(t\)\}\}\{1\-\\hat\{p\}\_\{i\}^\{\(t\)\}\}\\right\),with clipping for numerical stability\.
#### Stage\-2 fusion into a single scalar score\.
To obtain a single scalar score per admission for downstream dependence evaluation, we use one of two fusion strategies depending on the baseline\. First,*pooled learner fusion*pools examples across tasks by repeating admissions for each valid task label and fits a single Stage 2 model with per\-example sample weightαtBCE\\alpha\_\{t\}^\{\\mathrm\{BCE\}\}\. This is used for FM/GBT and for LR fusion over Stage 1 logits\. Second,*weighted logit averaging*defines the final score as a weighted mean of Stage 1 logits across tasks,
sibase=∑tα~tBCEai\(t\),base,α~tBCE∝αtBCEs\_\{i\}^\{\\mathrm\{base\}\}=\\sum\_\{t\}\\widetilde\{\\alpha\}\_\{t\}^\{\\mathrm\{BCE\}\}\\,a\_\{i\}^\{\(t\),\\mathrm\{base\}\},\\qquad\\widetilde\{\\alpha\}\_\{t\}^\{\\mathrm\{BCE\}\}\\propto\\alpha\_\{t\}^\{\\mathrm\{BCE\}\}over active tasks\. This is used for lightweight fusion variants such askkNN and NB when no Stage 2 learner is trained\.
### A\.9Outcomes and dataset construction
#### Label construction overview\.
We construct four admission\-level binary outcomes from MIMIC\-III and MIMIC\-IV: in\-hospital mortality, 30\-day all\-cause mortality, length of stay, and ICU transfer\. Labels are derived from structured hospital tables, such asADMISSIONS,PATIENTS, andICUSTAYS, using timestamp\-based rules relative to each admission’sADMITTIME\. Unless otherwise noted, labels are fully observed for all included admissions; when required timestamps are missing, we treat the label as missing and exclude the admission for that task\.
#### In\-hospital mortality\.
We define an admission\-level in\-hospital mortality label using the hospitalADMISSIONStable\. For each admissionii, we set
yi\(mort\)=1y\_\{i\}^\{\(\\mathrm\{mort\}\)\}=1if and only ifDEATHTIMEis non\-null for that admission, and setyi\(mort\)=0y\_\{i\}^\{\(\\mathrm\{mort\}\)\}=0otherwise\. For MIMIC\-IV, we restrict to an ICD\-10 cohort by keeping admissions with at least one ICD\-10 diagnosis record \(DIAGNOSES\_ICD\.icd\_version=10\)\. The label is fully observed for all included admissions, soMi\(mort\)=1M\_\{i\}^\{\(\\mathrm\{mort\}\)\}=1\. We additionally compute age at admission, capped at 90 for ages\>89\>89, using dataset\-specific conventions\.
#### 30\-day all\-cause mortality\.
For each admissionii, we define a 30\-day all\-cause mortality label relative to the admission start time\. We use patient\-level date\-of\-death,PATIENTS\.DOD, as the all\-cause death timestamp and compute the elapsed time in days from the admission timeADMITTIME\. For admissions with validADMITTIME, we setyi\(mort30\)=1y\_\{i\}^\{\(\\mathrm\{mort30\}\)\}=1if and only ifDODis observed and
0≤\(DOD−ADMITTIME\)≤300\\leq\(\\texttt\{DOD\}\-\\texttt\{ADMITTIME\}\)\\leq 30days, and setyi\(mort30\)=0y\_\{i\}^\{\(\\mathrm\{mort30\}\)\}=0otherwise\. Thus the label includes both in\-hospital deaths and post\-discharge deaths within 30 days of admission\. Admissions with missingADMITTIMEare excluded from this label and treated as missing\. For MIMIC\-IV, we restrict to an ICD\-10 cohort by keeping admissions with at least one ICD\-10 diagnosis record \(DIAGNOSES\_ICD\.icd\_version=10\)\.
#### Length of stay\.
We define an admission\-level length\-of\-stay label using stay duration computed fromADMISSIONS\.ADMITTIMEandADMISSIONS\.DISCHTIME\. For each admissioniiwith valid timestamps, we compute
los\_days=\(DISCHTIME−ADMITTIME\)\\texttt\{los\\\_days\}=\(\\texttt\{DISCHTIME\}\-\\texttt\{ADMITTIME\}\)in days and setyi\(los\)=1y\_\{i\}^\{\(\\mathrm\{los\}\)\}=1if and only iflos\_days\>7\\texttt\{los\\\_days\}\>7, and setyi\(los\)=0y\_\{i\}^\{\(\\mathrm\{los\}\)\}=0otherwise\. Admissions with missing admission or discharge times are excluded from this label\. For MIMIC\-IV, we restrict to an ICD\-10 cohort by keeping admissions with at least one ICD\-10 diagnosis record \(DIAGNOSES\_ICD\.icd\_version=10\)\.
#### ICU transfer\.
We derive ICU\-transfer labels by linking hospital admissions,hadm\_id, to ICU stays and extracting the first ICU entry time\. For each admissionii, we setfirst\_icu\_intimeto the earliest validICUSTAYS\.INTIMEassociated with thathadm\_idand computetime\_to\_icu\_hoursas the elapsed hours fromADMISSIONS\.ADMITTIMEtofirst\_icu\_intime\. We defineicu\_any=1 if and only if an ICU stay exists for the admission\. Admissions with missingADMITTIMEare excluded from this label\. Because MIMIC\-III is ICU\-heavy, meaning nearly all admissions have an ICU stay, we defineicu\_transferas a late ICU transfer indicator:
yi\(icu\)=1if and only iftime\_to\_icu\_hours\>24\.y\_\{i\}^\{\(\\mathrm\{icu\}\)\}=1\\quad\\text\{if and only if\}\\quad\\texttt\{time\\\_to\\\_icu\\\_hours\}\>24\.In MIMIC\-IV, where ICU admission is less ubiquitous, we use an any\-ICU\-after\-admission definition:
yi\(icu\)=1if and only iftime\_to\_icu\_hours\>0\.y\_\{i\}^\{\(\\mathrm\{icu\}\)\}=1\\quad\\text\{if and only if\}\\quad\\texttt\{time\\\_to\\\_icu\\\_hours\}\>0\.We restrict the MIMIC\-IV cohort to ICD\-10 admissions for consistency with the rest of our pipeline\.
#### Dataset construction and splits\.
For each dataset, MIMIC\-III and MIMIC\-IV, we construct an admission\-level cohort anchored on the set of unique\(subject\_id, hadm\_id\)pairs appearing in our in\-hospital mortality label file, and merge all other outcome labels by\(subject\_id, hadm\_id\)\. For MIMIC\-IV, this cohort is restricted to ICD\-10 admissions\.
We then create train/validation/test splits at the*patient*level by assigning eachsubject\_idto a split via a deterministic hash with proportions 70%/10%/20%, and placing all of that patient’s admissions in the same split\. This avoids leakage across repeated admissions from the same individual and yields an evaluation that reflects generalization to previously unseen patients\.
### A\.10Metric computation
The held\-out dependence metrics reported in Tables 1–2 are computed*per outcome*using the task\-specific validity mask for that outcome\. Thus, for each outcome column, all models are evaluated on the same valid test admissions\. For outcomett, let
ℐt=\{i:Mi\(t\)=1\}\\mathcal\{I\}\_\{t\}=\\\{i:M\_\{i\}^\{\(t\)\}=1\\\}denote admissions with a defined label\. For any scalar scoreqiq\_\{i\}being evaluated, we compute each metric using pairs
\{\(qi,yi\(t\)\)\}i∈ℐt\.\\\{\(q\_\{i\},y\_\{i\}^\{\(t\)\}\)\\\}\_\{i\\in\\mathcal\{I\}\_\{t\}\}\.If\|ℐt\|<3\|\\mathcal\{I\}\_\{t\}\|<3oryi\(t\)y\_\{i\}^\{\(t\)\}is single\-class onℐt\\mathcal\{I\}\_\{t\}, we return0\.
#### Distance correlation \(dCorr\)\.
We compute distance correlation between the scalar score and the binary outcome onℐt\\mathcal\{I\}\_\{t\}using the standard distance correlation implementation \(dcor\.distance\_correlation\)\. Because dCorr is signless, it measures the strength of dependence rather than direction\.
#### Mutual information \(MI, nats\)\.
We estimate mutual information between a scalar scoreqiq\_\{i\}and binary labelyi\(t\)y\_\{i\}^\{\(t\)\}onℐt\\mathcal\{I\}\_\{t\}using thekk\-nearest\-neighbor estimator implemented insklearn\.feature\_selection\.mutual\_info\_classif\. We first standardize scores on the valid subset by z\-scoring, and usekMI=min\(k0,\|ℐt\|−1\)k\_\{\\mathrm\{MI\}\}=\\min\(k\_\{0\},\|\\mathcal\{I\}\_\{t\}\|\-1\)neighbors with fixed random seed to avoid invalid settings when few valid labels are available\. The resulting estimate is reported in nats, the natural\-logarithm units of mutual information\.
## Appendix BRank\-One Multi\-Task nHSIC Theory: Roadmap, Assumptions, and Notation
#### Overview\.
We study the multi\-task dependence structure induced by normalized HSIC, where each task compares a centered score Gram matrix with a centered label Gram matrix\. For binary labels, the centered label Gram matrix reduces to a rank\-one object determined by the centered label profile\. This lets us rewrite the weighted multi\-task objective as an alignment between the centered score Gram matrix and a single cross\-task label\-alignment matrix\.
The main idea is then to isolate the dominant shared component of this cross\-task label structure\. We stack the weighted centered label profiles across tasks and project the resulting matrix onto its best rank\-one approximation\. The projected objective reduces exactly to a single\-admission\-direction problem governed by the leading right singular vectorvv\. Applying the global nHSIC ratio lemma to this direction yields an explicit two\-level monotone threshold rule: the best split is the one whose centered step vector is most aligned withvv\.
Finally, when the rank\-one residual makes the explicit boundεgap\\varepsilon\_\{\\mathrm\{gap\}\}small, the appendix transfers near\-optimality from the projected objective back to the original multi\-task objective using anε0\\varepsilon\_\{0\}\-floored denominator\. We also include a noise\-averaging bridge showing that, when the learned\-score Gram matrix concentrates around its noise average, the realized learned\-score nHSIC is close to the corresponding averaged\-kernel target\.
#### Finite\-sample nHSIC objective\.
For each tasktt, the appendix studies
nHSIC\(r,y\(t\)\):=⟨Kc\(r\),Lt,c⟩F‖Kc\(r\)‖F‖Lt,c‖F,Kc\(r\):=HK\(r\)H,Lt,c:=HLtH,\\mathrm\{nHSIC\}\\\!\\left\(r,y^\{\(t\)\}\\right\):=\\frac\{\\langle K\_\{c\}\(r\),L\_\{t,c\}\\rangle\_\{F\}\}\{\\\|K\_\{c\}\(r\)\\\|\_\{F\}\\\|L\_\{t,c\}\\\|\_\{F\}\},\\qquad K\_\{c\}\(r\):=HK\(r\)H,\\quad L\_\{t,c\}:=HL\_\{t\}H,with the convention that the ratio is0whenever a required denominator factor is zero\. The multi\-task objective is
J\(r\):=∑t=1TαtnHSIC\(r,y\(t\)\),αt\>0\.J\(r\):=\\sum\_\{t=1\}^\{T\}\\alpha\_\{t\}\\,\\mathrm\{nHSIC\}\\\!\\left\(r,y^\{\(t\)\}\\right\),\\qquad\\alpha\_\{t\}\>0\.
#### Proof dependencies\.
The appendix proceeds as follows\.
1. 1\.Lemma[B\.1](https://arxiv.org/html/2606.17450#A2.Thmtheorem1)shows that binary delta\-label Grams satisfyLt,c=2ℓ\(t\)ℓ\(t\)⊤,ℓ\(t\):=Hy\(t\)\.L\_\{t,c\}=2\\ell^\{\(t\)\}\\ell^\{\(t\)\\top\},\\qquad\\ell^\{\(t\)\}:=Hy^\{\(t\)\}\.
2. 2\.Lemma[B\.2](https://arxiv.org/html/2606.17450#A2.Thmtheorem2)studies Δ¯w\(r\):=w⊤Kc\(r\)w‖Kc\(r\)‖F,𝟏⊤w=0,\\bar\{\\Delta\}\_\{w\}\(r\):=\\frac\{w^\{\\top\}K\_\{c\}\(r\)w\}\{\\\|K\_\{c\}\(r\)\\\|\_\{F\}\},\\qquad\\mathbf\{1\}^\{\\top\}w=0,provesΔ¯w\(r\)≤‖w‖22\\bar\{\\Delta\}\_\{w\}\(r\)\\leq\\\|w\\\|\_\{2\}^\{2\}, and gives the exact threshold value Δ¯w\(r\(j\)\)=‖w‖22ρj2,ρj:=w⊤gj‖w‖2‖gj‖2,gj:=H𝟏Rj\.\\bar\{\\Delta\}\_\{w\}\(r^\{\(j\)\}\)=\\\|w\\\|\_\{2\}^\{2\}\\rho\_\{j\}^\{2\},\\qquad\\rho\_\{j\}:=\\frac\{w^\{\\top\}g\_\{j\}\}\{\\\|w\\\|\_\{2\}\\\|g\_\{j\}\\\|\_\{2\}\},\\qquad g\_\{j\}:=H\\mathbf\{1\}\_\{R\_\{j\}\}\.
3. 3\.Lemma[B\.3](https://arxiv.org/html/2606.17450#A2.Thmtheorem3)proves the rank\-one projection reductionJ1\(r\)=σ12Δ¯v\(r\)J\_\{1\}\(r\)=\\sigma\_\{1\}^\{2\}\\bar\{\\Delta\}\_\{v\}\(r\)forW~1=σ1uv⊤\\widetilde\{W\}\_\{1\}=\\sigma\_\{1\}uv^\{\\top\}\.
4. 4\.Theorem[B\.4](https://arxiv.org/html/2606.17450#A2.Thmtheorem4)defines the exact uniform gapΔgap\\Delta\_\{\\mathrm\{gap\}\}and proves the explicit boundΔgap≤εgap\.\\Delta\_\{\\mathrm\{gap\}\}\\leq\\varepsilon\_\{\\mathrm\{gap\}\}\.Corollary[B\.5](https://arxiv.org/html/2606.17450#A2.Thmtheorem5)then transfers near\-optimality from the floored projected objective to the original floored objective using this bound\.
5. 5\.Lemma[B\.6](https://arxiv.org/html/2606.17450#A2.Thmtheorem6)shows that ifKc\(s\)K\_\{c\}\(s\)concentrates aroundK¯c:=𝔼η\[Kc\(s\)\],\\bar\{K\}\_\{c\}:=\\mathbb\{E\}\_\{\\eta\}\[K\_\{c\}\(s\)\],thennHSIC\(s,y\(t\)\)\\mathrm\{nHSIC\}\(s,y^\{\(t\)\}\)is close to the averaged\-kernel normalized target\.
### B\.1Core assumptions and conventions
Unless stated otherwise, we assume:
1. \(A1\)Admissions are indexed by latent severity:z1<⋯<zn\.z\_\{1\}<\\cdots<z\_\{n\}\.
2. \(A2\)Each included task is nonconstant:‖Hy\(t\)‖2\>0\.\\\|Hy^\{\(t\)\}\\\|\_\{2\}\>0\.
3. \(A3\)Scores lie in the bounded monotone class:ℛB↑:=\{r∈\[−B,B\]n:r1≤⋯≤rn\}\.\\mathcal\{R\}\_\{B\}^\{\\uparrow\}:=\\\{r\\in\[\-B,B\]^\{n\}:r\_\{1\}\\leq\\cdots\\leq r\_\{n\}\\\}\.
4. \(A4\)The score kernel is radial, bounded, continuous, and positive semidefinite: Kij\(r\)=k\(ri,rj\)=κ\(\|ri−rj\|\),K\_\{ij\}\(r\)=k\(r\_\{i\},r\_\{j\}\)=\\kappa\(\|r\_\{i\}\-r\_\{j\}\|\),whereκ\\kappais strictly decreasing on the relevant range\[0,2B\]\[0,2B\]\. We assume thatkkis positive semidefinite on\[−B,B\]\[\-B,B\], so that for everyr∈ℛB↑r\\in\\mathcal\{R\}\_\{B\}^\{\\uparrow\}, the Gram matrixK\(r\)K\(r\)is symmetric positive semidefinite\. Consequently, Kc\(r\)=HK\(r\)HK\_\{c\}\(r\)=HK\(r\)His also symmetric positive semidefinite\.
5. \(A5\)All objectives use centered Gram matrices: Kc\(r\)=HK\(r\)H,Lt,c=HLtH,H=I−1n𝟏𝟏⊤\.K\_\{c\}\(r\)=HK\(r\)H,\\qquad L\_\{t,c\}=HL\_\{t\}H,\\qquad H=I\-\\frac\{1\}\{n\}\\mathbf\{1\}\\mathbf\{1\}^\{\\top\}\.
6. \(A6\)When invoked, the best rank\-one approximation isW~1=σ1uv⊤\\widetilde\{W\}\_\{1\}=\\sigma\_\{1\}uv^\{\\top\}, withσ1≥0\\sigma\_\{1\}\\geq 0,‖u‖2=‖v‖2=1\\\|u\\\|\_\{2\}=\\\|v\\\|\_\{2\}=1\. Since each row ofW~\\widetilde\{W\}is centered,W~𝟏n=0\\widetilde\{W\}\\mathbf\{1\}\_\{n\}=0, where𝟏n∈ℝn\\mathbf\{1\}\_\{n\}\\in\\mathbb\{R\}^\{n\}denotes the all\-ones vector\. Hence, wheneverσ1\>0\\sigma\_\{1\}\>0, the associated right singular vector satisfies𝟏⊤v=0\\mathbf\{1\}^\{\\top\}v=0, equivalentlyHv=vHv=v\.
7. \(A7\)nHSIC\-type ratios are set to0whenever a required denominator factor is zero\. Floored objectives usemax\{‖Kc\(r\)‖F,ε0\}\.\\max\\\{\\\|K\_\{c\}\(r\)\\\|\_\{F\},\\varepsilon\_\{0\}\\\}\.
### B\.2Notation at a glance
###### Lemma B\.1\(Centered delta\-kernel is rank one for binary labels\)\.
Lety∈\{0,1\}ny\\in\\\{0,1\\\}^\{n\}and define the delta label Gram matrixLij:=𝕀\{yi=yj\},L\_\{ij\}:=\\mathbb\{I\}\\\{y\_\{i\}=y\_\{j\}\\\},where𝕀\{⋅\}\\mathbb\{I\}\\\{\\cdot\\\}denotes the indicator function\. LetLc:=HLHL\_\{c\}:=HLH\. ThenLc=2\(Hy\)\(Hy\)⊤,L\_\{c\}=2\(Hy\)\(Hy\)^\{\\top\},sorank\(Lc\)≤1\\mathrm\{rank\}\(L\_\{c\}\)\\leq 1\. Moreover,‖Lc‖F=2‖Hy‖22\.\\\|L\_\{c\}\\\|\_\{F\}=2\\\|Hy\\\|\_\{2\}^\{2\}\.
###### Proof\.
UsingLij=yiyj\+\(1−yi\)\(1−yj\)=1−yi−yj\+2yiyj,L\_\{ij\}=y\_\{i\}y\_\{j\}\+\(1\-y\_\{i\}\)\(1\-y\_\{j\}\)=1\-y\_\{i\}\-y\_\{j\}\+2y\_\{i\}y\_\{j\},we getL=𝟏𝟏⊤−y𝟏⊤−𝟏y⊤\+2yy⊤\.L=\\mathbf\{1\}\\mathbf\{1\}^\{\\top\}\-y\\mathbf\{1\}^\{\\top\}\-\\mathbf\{1\}y^\{\\top\}\+2yy^\{\\top\}\.SinceH𝟏=0H\\mathbf\{1\}=0, centering kills the first three terms and
Lc=H\(2yy⊤\)H=2\(Hy\)\(Hy\)⊤\.L\_\{c\}=H\(2yy^\{\\top\}\)H=2\(Hy\)\(Hy\)^\{\\top\}\.The Frobenius norm follows from‖uu⊤‖F=‖u‖22\\\|uu^\{\\top\}\\\|\_\{F\}=\\\|u\\\|\_\{2\}^\{2\}\. ∎
###### Lemma B\.2\(Global upper bound and threshold certificate for the nHSIC ratio\)\.
Considernnadmissions with ordered latent severitiesz1<⋯<znz\_\{1\}<\\cdots<z\_\{n\}, and fix any centered target directionw∈ℝnw\\in\\mathbb\{R\}^\{n\}, so𝟏⊤w=0\\mathbf\{1\}^\{\\top\}w=0\. Ifw=0w=0, thenΔ¯w≡0\\bar\{\\Delta\}\_\{w\}\\equiv 0and all statements are trivial\. Hence assumew≠0w\\neq 0below\.
Letk\(x,x′\)=κ\(\|x−x′\|\)k\(x,x^\{\\prime\}\)=\\kappa\(\|x\-x^\{\\prime\}\|\)be a bounded continuous positive semidefinite radial kernel, whereκ\\kappais strictly decreasing on\[0,2B\]\[0,2B\]\. For any score vectorr∈ℛB↑r\\in\\mathcal\{R\}\_\{B\}^\{\\uparrow\}, define
Kij\(r\):=κ\(\|ri−rj\|\),Kc\(r\):=HK\(r\)H,H:=I−1n𝟏𝟏⊤\.K\_\{ij\}\(r\):=\\kappa\(\|r\_\{i\}\-r\_\{j\}\|\),\\qquad K\_\{c\}\(r\):=HK\(r\)H,\\qquad H:=I\-\\frac\{1\}\{n\}\\mathbf\{1\}\\mathbf\{1\}^\{\\top\}\.Define
Δ¯w\(r\):=w⊤Kc\(r\)w‖Kc\(r\)‖F,\\bar\{\\Delta\}\_\{w\}\(r\):=\\frac\{w^\{\\top\}K\_\{c\}\(r\)w\}\{\\\|K\_\{c\}\(r\)\\\|\_\{F\}\},with the conventionΔ¯w\(r\)=0\\bar\{\\Delta\}\_\{w\}\(r\)=0whenever‖Kc\(r\)‖F=0\\\|K\_\{c\}\(r\)\\\|\_\{F\}=0\.
For each split indexj∈\{1,…,n−1\}j\\in\\\{1,\\dots,n\-1\\\}, define
Lj:=\{1,…,j\},Rj:=\{j\+1,…,n\},L\_\{j\}:=\\\{1,\\dots,j\\\},\\qquad R\_\{j\}:=\\\{j\+1,\\dots,n\\\},and
gj:=H𝟏Rj=𝟏Rj−\|Rj\|n𝟏,g\_\{j\}:=H\\mathbf\{1\}\_\{R\_\{j\}\}=\\mathbf\{1\}\_\{R\_\{j\}\}\-\\frac\{\|R\_\{j\}\|\}\{n\}\\mathbf\{1\},where𝟏Rj∈ℝn\\mathbf\{1\}\_\{R\_\{j\}\}\\in\\mathbb\{R\}^\{n\}is the indicator vector ofRjR\_\{j\}\. Sincew≠0w\\neq 0andgj≠0g\_\{j\}\\neq 0, define
ρj:=w⊤gj‖w‖2‖gj‖2\.\\rho\_\{j\}:=\\frac\{w^\{\\top\}g\_\{j\}\}\{\\\|w\\\|\_\{2\}\\\|g\_\{j\}\\\|\_\{2\}\}\.
Then:
1. 1\.For everyr∈ℛB↑r\\in\\mathcal\{R\}\_\{B\}^\{\\uparrow\}, Δ¯w\(r\)≤‖w‖22\.\\bar\{\\Delta\}\_\{w\}\(r\)\\leq\\\|w\\\|\_\{2\}^\{2\}\.
2. 2\.For any nontrivial two\-level threshold scorer\(j\)∈ℛB↑r^\{\(j\)\}\\in\\mathcal\{R\}\_\{B\}^\{\\uparrow\}of the form ri\(j\)=\{a,i∈Lj,b,i∈Rj,a<b,r\_\{i\}^\{\(j\)\}=\\begin\{cases\}a,&i\\in L\_\{j\},\\\\ b,&i\\in R\_\{j\},\\end\{cases\}\\qquad a<b,we have Δ¯w\(r\(j\)\)=‖w‖22ρj2\.\\bar\{\\Delta\}\_\{w\}\(r^\{\(j\)\}\)=\\\|w\\\|\_\{2\}^\{2\}\\rho\_\{j\}^\{2\}\.In particular, this value is independent of the separation\|a−b\|\|a\-b\|\.
3. 3\.Let ρ⋆2:=max1≤j≤n−1ρj2,\\rho\_\{\\star\}^\{2\}:=\\max\_\{1\\leq j\\leq n\-1\}\\rho\_\{j\}^\{2\},and letj⋆j^\{\\star\}attain the maximum\. Then any nontrivial two\-level threshold onLj⋆/Rj⋆L\_\{j^\{\\star\}\}/R\_\{j^\{\\star\}\}satisfies Δ¯w\(r\(j⋆\)\)=‖w‖22ρ⋆2\\bar\{\\Delta\}\_\{w\}\(r^\{\(j^\{\\star\}\)\}\)=\\\|w\\\|\_\{2\}^\{2\}\\rho\_\{\\star\}^\{2\}and hence Δ¯w\(r\(j⋆\)\)≥ρ⋆2supu∈ℛB↑Δ¯w\(u\)\.\\bar\{\\Delta\}\_\{w\}\(r^\{\(j^\{\\star\}\)\}\)\\geq\\rho\_\{\\star\}^\{2\}\\sup\_\{u\\in\\mathcal\{R\}\_\{B\}^\{\\uparrow\}\}\\bar\{\\Delta\}\_\{w\}\(u\)\.Thusρ⋆2\\rho\_\{\\star\}^\{2\}is an explicit certificate for the fraction of the global upper bound achieved by the best threshold\.
4. 4\.There exists a nontrivial two\-level thresholdr\(j\)r^\{\(j\)\}that achieves the global upper bound Δ¯w\(r\(j\)\)=‖w‖22\\bar\{\\Delta\}\_\{w\}\(r^\{\(j\)\}\)=\\\|w\\\|\_\{2\}^\{2\}if and only ifρ⋆2=1\\rho\_\{\\star\}^\{2\}=1, equivalentlywwis proportional to some centered step vectorgjg\_\{j\}\.
###### Proof\.
Since𝟏⊤w=0\\mathbf\{1\}^\{\\top\}w=0, we haveHw=wHw=w\. Therefore
w⊤Kc\(r\)w=w⊤HK\(r\)Hw=w⊤K\(r\)w\.w^\{\\top\}K\_\{c\}\(r\)w=w^\{\\top\}HK\(r\)Hw=w^\{\\top\}K\(r\)w\.Also,
w⊤Kc\(r\)w=⟨Kc\(r\),ww⊤⟩F\.w^\{\\top\}K\_\{c\}\(r\)w=\\langle K\_\{c\}\(r\),ww^\{\\top\}\\rangle\_\{F\}\.Hence
Δ¯w\(r\)=⟨Kc\(r\),ww⊤⟩F‖Kc\(r\)‖F\.\\bar\{\\Delta\}\_\{w\}\(r\)=\\frac\{\\langle K\_\{c\}\(r\),ww^\{\\top\}\\rangle\_\{F\}\}\{\\\|K\_\{c\}\(r\)\\\|\_\{F\}\}\.
If‖Kc\(r\)‖F=0\\\|K\_\{c\}\(r\)\\\|\_\{F\}=0, thenΔ¯w\(r\)=0\\bar\{\\Delta\}\_\{w\}\(r\)=0by convention, so the bound is immediate\. Otherwise, by Cauchy–Schwarz in the Frobenius inner product,
Δ¯w\(r\)≤‖Kc\(r\)‖F‖ww⊤‖F‖Kc\(r\)‖F=‖ww⊤‖F=‖w‖22,\\bar\{\\Delta\}\_\{w\}\(r\)\\leq\\frac\{\\\|K\_\{c\}\(r\)\\\|\_\{F\}\\\|ww^\{\\top\}\\\|\_\{F\}\}\{\\\|K\_\{c\}\(r\)\\\|\_\{F\}\}=\\\|ww^\{\\top\}\\\|\_\{F\}=\\\|w\\\|\_\{2\}^\{2\},which proves the global upper bound\.
Now fix a nontrivial two\-level threshold scorer\(j\)r^\{\(j\)\}, with valueaaonLjL\_\{j\}and valuebbonRjR\_\{j\}, wherea<ba<b\. Let
c0:=κ\(0\),c1:=κ\(\|a−b\|\)\.c\_\{0\}:=\\kappa\(0\),\\qquad c\_\{1\}:=\\kappa\(\|a\-b\|\)\.Sinceκ\\kappais strictly decreasing anda<ba<b, we havec0−c1\>0c\_\{0\}\-c\_\{1\}\>0\. The score Gram matrix is block\-constant and can be written as
K\(r\(j\)\)=c1𝟏𝟏⊤\+\(c0−c1\)\(𝟏Lj𝟏Lj⊤\+𝟏Rj𝟏Rj⊤\)\.K\(r^\{\(j\)\}\)=c\_\{1\}\\mathbf\{1\}\\mathbf\{1\}^\{\\top\}\+\(c\_\{0\}\-c\_\{1\}\)\\left\(\\mathbf\{1\}\_\{L\_\{j\}\}\\mathbf\{1\}\_\{L\_\{j\}\}^\{\\top\}\+\\mathbf\{1\}\_\{R\_\{j\}\}\\mathbf\{1\}\_\{R\_\{j\}\}^\{\\top\}\\right\)\.Centering removes the first term becauseH𝟏=0H\\mathbf\{1\}=0, so
Kc\(r\(j\)\)=\(c0−c1\)H\(𝟏Lj𝟏Lj⊤\+𝟏Rj𝟏Rj⊤\)H\.K\_\{c\}\(r^\{\(j\)\}\)=\(c\_\{0\}\-c\_\{1\}\)H\\left\(\\mathbf\{1\}\_\{L\_\{j\}\}\\mathbf\{1\}\_\{L\_\{j\}\}^\{\\top\}\+\\mathbf\{1\}\_\{R\_\{j\}\}\\mathbf\{1\}\_\{R\_\{j\}\}^\{\\top\}\\right\)H\.Since
𝟏Lj=𝟏−𝟏Rj,\\mathbf\{1\}\_\{L\_\{j\}\}=\\mathbf\{1\}\-\\mathbf\{1\}\_\{R\_\{j\}\},we have
H𝟏Lj=H𝟏−H𝟏Rj=−gj,H𝟏Rj=gj\.H\\mathbf\{1\}\_\{L\_\{j\}\}=H\\mathbf\{1\}\-H\\mathbf\{1\}\_\{R\_\{j\}\}=\-g\_\{j\},\\qquad H\\mathbf\{1\}\_\{R\_\{j\}\}=g\_\{j\}\.Therefore
H\(𝟏Lj𝟏Lj⊤\+𝟏Rj𝟏Rj⊤\)H=\(H𝟏Lj\)\(H𝟏Lj\)⊤\+\(H𝟏Rj\)\(H𝟏Rj\)⊤H\\left\(\\mathbf\{1\}\_\{L\_\{j\}\}\\mathbf\{1\}\_\{L\_\{j\}\}^\{\\top\}\+\\mathbf\{1\}\_\{R\_\{j\}\}\\mathbf\{1\}\_\{R\_\{j\}\}^\{\\top\}\\right\)H=\(H\\mathbf\{1\}\_\{L\_\{j\}\}\)\(H\\mathbf\{1\}\_\{L\_\{j\}\}\)^\{\\top\}\+\(H\\mathbf\{1\}\_\{R\_\{j\}\}\)\(H\\mathbf\{1\}\_\{R\_\{j\}\}\)^\{\\top\}=\(−gj\)\(−gj\)⊤\+gjgj⊤=2gjgj⊤\.=\(\-g\_\{j\}\)\(\-g\_\{j\}\)^\{\\top\}\+g\_\{j\}g\_\{j\}^\{\\top\}=2g\_\{j\}g\_\{j\}^\{\\top\}\.Thus
Kc\(r\(j\)\)=2\(c0−c1\)gjgj⊤\.K\_\{c\}\(r^\{\(j\)\}\)=2\(c\_\{0\}\-c\_\{1\}\)g\_\{j\}g\_\{j\}^\{\\top\}\.Using‖gjgj⊤‖F=‖gj‖22\\\|g\_\{j\}g\_\{j\}^\{\\top\}\\\|\_\{F\}=\\\|g\_\{j\}\\\|\_\{2\}^\{2\}, we get
‖Kc\(r\(j\)\)‖F=2\(c0−c1\)‖gj‖22\.\\\|K\_\{c\}\(r^\{\(j\)\}\)\\\|\_\{F\}=2\(c\_\{0\}\-c\_\{1\}\)\\\|g\_\{j\}\\\|\_\{2\}^\{2\}\.Therefore
Δ¯w\(r\(j\)\)=⟨2\(c0−c1\)gjgj⊤,ww⊤⟩F2\(c0−c1\)‖gj‖22\.\\bar\{\\Delta\}\_\{w\}\(r^\{\(j\)\}\)=\\frac\{\\left\\langle 2\(c\_\{0\}\-c\_\{1\}\)g\_\{j\}g\_\{j\}^\{\\top\},ww^\{\\top\}\\right\\rangle\_\{F\}\}\{2\(c\_\{0\}\-c\_\{1\}\)\\\|g\_\{j\}\\\|\_\{2\}^\{2\}\}\.Since
⟨gjgj⊤,ww⊤⟩F=\(w⊤gj\)2,\\langle g\_\{j\}g\_\{j\}^\{\\top\},ww^\{\\top\}\\rangle\_\{F\}=\(w^\{\\top\}g\_\{j\}\)^\{2\},we obtain
Δ¯w\(r\(j\)\)=\(w⊤gj\)2‖gj‖22=‖w‖22\(w⊤gj‖w‖2‖gj‖2\)2=‖w‖22ρj2\.\\bar\{\\Delta\}\_\{w\}\(r^\{\(j\)\}\)=\\frac\{\(w^\{\\top\}g\_\{j\}\)^\{2\}\}\{\\\|g\_\{j\}\\\|\_\{2\}^\{2\}\}=\\\|w\\\|\_\{2\}^\{2\}\\left\(\\frac\{w^\{\\top\}g\_\{j\}\}\{\\\|w\\\|\_\{2\}\\\|g\_\{j\}\\\|\_\{2\}\}\\right\)^\{2\}=\\\|w\\\|\_\{2\}^\{2\}\\rho\_\{j\}^\{2\}\.This proves the exact threshold value and shows that\|a−b\|\|a\-b\|cancels\.
The best threshold statement follows by choosingj⋆∈argmaxjρj2j^\{\\star\}\\in\\arg\\max\_\{j\}\\rho\_\{j\}^\{2\}\. Since the global upper bound is‖w‖22\\\|w\\\|\_\{2\}^\{2\}, the value
‖w‖22ρ⋆2\\\|w\\\|\_\{2\}^\{2\}\\rho\_\{\\star\}^\{2\}is aρ⋆2\\rho\_\{\\star\}^\{2\}\-fraction of the global upper bound\.
Finally, for a nontrivial two\-level threshold,
Kc\(r\(j\)\)=2\(c0−c1\)gjgj⊤,K\_\{c\}\(r^\{\(j\)\}\)=2\(c\_\{0\}\-c\_\{1\}\)g\_\{j\}g\_\{j\}^\{\\top\},with2\(c0−c1\)\>02\(c\_\{0\}\-c\_\{1\}\)\>0\. Hence this threshold attains the global upper bound‖w‖22\\\|w\\\|\_\{2\}^\{2\}if and only if
‖w‖22ρj2=‖w‖22,\\\|w\\\|\_\{2\}^\{2\}\\rho\_\{j\}^\{2\}=\\\|w\\\|\_\{2\}^\{2\},equivalentlyρj2=1\\rho\_\{j\}^\{2\}=1\. This is equivalent to equality in the Cauchy–Schwarz inequality betweenwwandgjg\_\{j\}, hence tow∝gjw\\propto g\_\{j\}\. Maximizing overjjgives the conditionρ⋆2=1\\rho\_\{\\star\}^\{2\}=1\. ∎
Unfloored vs\.ε0\\varepsilon\_\{0\}\-floored normalization\.Lemma[B\.2](https://arxiv.org/html/2606.17450#A2.Thmtheorem2)characterizes the exact geometry of the unfloored ratio
Δ¯w\(r\)=w⊤Kc\(r\)w‖Kc\(r\)‖F\.\\bar\{\\Delta\}\_\{w\}\(r\)=\\frac\{w^\{\\top\}K\_\{c\}\(r\)w\}\{\\\|K\_\{c\}\(r\)\\\|\_\{F\}\}\.In later approximation and stability results, we introduce anε0\\varepsilon\_\{0\}\-floor in the denominator to avoid degeneracy for nearly constant scores\. When‖Kc\(r\)‖F≥ε0\\\|K\_\{c\}\(r\)\\\|\_\{F\}\\geq\\varepsilon\_\{0\}, the unfloored and floored objectives coincide\.
#### Set up the multi\-task finite\-sample problem \(nHSIC form\)\.
Fixnnadmissions with strictly ordered latent severities
z1<z2<⋯<zn\.z\_\{1\}<z\_\{2\}<\\cdots<z\_\{n\}\.For each taskt∈\[T\]:=\{1,…,T\}t\\in\[T\]:=\\\{1,\\dots,T\\\}, let
y\(t\)∈\{0,1\}ny^\{\(t\)\}\\in\\\{0,1\\\}^\{n\}denote the observed binary labels on thesennadmissions\. Assume each included task is nonconstant, equivalently
‖Hy\(t\)‖2\>0\.\\\|Hy^\{\(t\)\}\\\|\_\{2\}\>0\.LetB\>0B\>0be fixed, and define
ℛB↑:=\{r∈\[−B,B\]n:r1≤⋯≤rn\}\.\\mathcal\{R\}\_\{B\}^\{\\uparrow\}:=\\\{r\\in\[\-B,B\]^\{n\}:r\_\{1\}\\leq\\cdots\\leq r\_\{n\}\\\}\.Writer=\(r1,…,rn\)⊤r=\(r\_\{1\},\\dots,r\_\{n\}\)^\{\\top\}\. Letk\(x,x′\)=κ\(\|x−x′\|\)k\(x,x^\{\\prime\}\)=\\kappa\(\|x\-x^\{\\prime\}\|\)be a bounded continuous positive semidefinite radial kernel, whereκ\\kappais strictly decreasing on\[0,2B\]\[0,2B\]\. Define
Kij\(r\):=κ\(\|ri−rj\|\),Kc\(r\):=HK\(r\)H,H:=I−1n𝟏𝟏⊤\.K\_\{ij\}\(r\):=\\kappa\(\|r\_\{i\}\-r\_\{j\}\|\),\\qquad K\_\{c\}\(r\):=HK\(r\)H,\\qquad H:=I\-\\frac\{1\}\{n\}\\mathbf\{1\}\\mathbf\{1\}^\{\\top\}\.
#### Single\-task nHSIC objective\.
For each tasktt, letLtL\_\{t\}be the delta label Gram matrix
\(Lt\)ij:=𝕀\{yi\(t\)=yj\(t\)\},\(L\_\{t\}\)\_\{ij\}:=\\mathbb\{I\}\\\{y\_\{i\}^\{\(t\)\}=y\_\{j\}^\{\(t\)\}\\\},and define its centered version
Lt,c:=HLtH\.L\_\{t,c\}:=HL\_\{t\}H\.The finite\-sample normalized HSIC objective is
nHSIC\(r,y\(t\)\):=⟨Kc\(r\),Lt,c⟩F‖Kc\(r\)‖F‖Lt,c‖F\.\\mathrm\{nHSIC\}\\\!\\left\(r,y^\{\(t\)\}\\right\):=\\frac\{\\langle K\_\{c\}\(r\),L\_\{t,c\}\\rangle\_\{F\}\}\{\\\|K\_\{c\}\(r\)\\\|\_\{F\}\\\|L\_\{t,c\}\\\|\_\{F\}\}\.We setnHSIC\(r,y\(t\)\)=0\\mathrm\{nHSIC\}\(r,y^\{\(t\)\}\)=0whenever either denominator factor is zero\. By Lemma[B\.1](https://arxiv.org/html/2606.17450#A2.Thmtheorem1), for binary labels,
Lt,c=2ℓ\(t\)ℓ\(t\)⊤,ℓ\(t\):=Hy\(t\),L\_\{t,c\}=2\\ell^\{\(t\)\}\\ell^\{\(t\)\\top\},\\qquad\\ell^\{\(t\)\}:=Hy^\{\(t\)\},and
‖Lt,c‖F=2‖ℓ\(t\)‖22,⟨Kc\(r\),Lt,c⟩F=2ℓ\(t\)⊤Kc\(r\)ℓ\(t\)\.\\\|L\_\{t,c\}\\\|\_\{F\}=2\\\|\\ell^\{\(t\)\}\\\|\_\{2\}^\{2\},\\qquad\\langle K\_\{c\}\(r\),L\_\{t,c\}\\rangle\_\{F\}=2\\ell^\{\(t\)\\top\}K\_\{c\}\(r\)\\ell^\{\(t\)\}\.Therefore, when‖Kc\(r\)‖F\>0\\\|K\_\{c\}\(r\)\\\|\_\{F\}\>0,
nHSIC\(r,y\(t\)\)=ℓ\(t\)⊤Kc\(r\)ℓ\(t\)‖Kc\(r\)‖F‖ℓ\(t\)‖22\.\\mathrm\{nHSIC\}\\\!\\left\(r,y^\{\(t\)\}\\right\)=\\frac\{\\ell^\{\(t\)\\top\}K\_\{c\}\(r\)\\ell^\{\(t\)\}\}\{\\\|K\_\{c\}\(r\)\\\|\_\{F\}\\\|\\ell^\{\(t\)\}\\\|\_\{2\}^\{2\}\}\.When‖Kc\(r\)‖F=0\\\|K\_\{c\}\(r\)\\\|\_\{F\}=0, we keep the conventionnHSIC\(r,y\(t\)\)=0\\mathrm\{nHSIC\}\(r,y^\{\(t\)\}\)=0\.
#### Multi\-task weighted\-sum nHSIC objective\.
Letαt\>0\\alpha\_\{t\}\>0be fixed task weights and define
J\(r\):=∑t=1TαtnHSIC\(r,y\(t\)\)\.J\(r\):=\\sum\_\{t=1\}^\{T\}\\alpha\_\{t\}\\,\\mathrm\{nHSIC\}\\\!\\left\(r,y^\{\(t\)\}\\right\)\.Using the previous display, on the event‖Kc\(r\)‖F\>0\\\|K\_\{c\}\(r\)\\\|\_\{F\}\>0,
J\(r\)=1‖Kc\(r\)‖F∑t=1Tαt‖ℓ\(t\)‖22ℓ\(t\)⊤Kc\(r\)ℓ\(t\)\.J\(r\)=\\frac\{1\}\{\\\|K\_\{c\}\(r\)\\\|\_\{F\}\}\\sum\_\{t=1\}^\{T\}\\frac\{\\alpha\_\{t\}\}\{\\\|\\ell^\{\(t\)\}\\\|\_\{2\}^\{2\}\}\\ell^\{\(t\)\\top\}K\_\{c\}\(r\)\\ell^\{\(t\)\}\.If‖Kc\(r\)‖F=0\\\|K\_\{c\}\(r\)\\\|\_\{F\}=0, thenJ\(r\)=0J\(r\)=0by convention\.
#### Absorbing label norms into weights\.
Define
βt:=αt‖ℓ\(t\)‖22,w~\(t\):=βtℓ\(t\)\.\\beta\_\{t\}:=\\frac\{\\alpha\_\{t\}\}\{\\\|\\ell^\{\(t\)\}\\\|\_\{2\}^\{2\}\},\\qquad\\widetilde\{w\}^\{\(t\)\}:=\\sqrt\{\\beta\_\{t\}\}\\,\\ell^\{\(t\)\}\.LetW~∈ℝT×n\\widetilde\{W\}\\in\\mathbb\{R\}^\{T\\times n\}have rows
W~t,:=w~\(t\)⊤=βtℓ\(t\)⊤\.\\widetilde\{W\}\_\{t,:\}=\\widetilde\{w\}^\{\(t\)\\top\}=\\sqrt\{\\beta\_\{t\}\}\\,\\ell^\{\(t\)\\top\}\.Then
W~⊤W~=∑t=1Tβtℓ\(t\)ℓ\(t\)⊤\.\\widetilde\{W\}^\{\\top\}\\widetilde\{W\}=\\sum\_\{t=1\}^\{T\}\\beta\_\{t\}\\,\\ell^\{\(t\)\}\\ell^\{\(t\)\\top\}\.Hence, when‖Kc\(r\)‖F\>0\\\|K\_\{c\}\(r\)\\\|\_\{F\}\>0,
J\(r\)=⟨Kc\(r\),W~⊤W~⟩F‖Kc\(r\)‖F\.J\(r\)=\\frac\{\\langle K\_\{c\}\(r\),\\widetilde\{W\}^\{\\top\}\\widetilde\{W\}\\rangle\_\{F\}\}\{\\\|K\_\{c\}\(r\)\\\|\_\{F\}\}\.When‖Kc\(r\)‖F=0\\\|K\_\{c\}\(r\)\\\|\_\{F\}=0, we keep the conventionJ\(r\)=0J\(r\)=0\.
#### Centering of task weight vectors\.
Sinceℓ\(t\)=Hy\(t\)\\ell^\{\(t\)\}=Hy^\{\(t\)\}, we have𝟏⊤ℓ\(t\)=0\\mathbf\{1\}^\{\\top\}\\ell^\{\(t\)\}=0\. Therefore
𝟏⊤w~\(t\)=βt1⊤ℓ\(t\)=0\\mathbf\{1\}^\{\\top\}\\widetilde\{w\}^\{\(t\)\}=\\sqrt\{\\beta\_\{t\}\}\\,\\mathbf\{1\}^\{\\top\}\\ell^\{\(t\)\}=0for alltt\. Equivalently,
W~𝟏=0\.\\widetilde\{W\}\\mathbf\{1\}=0\.
#### Rank\-one projected objective\.
LetW~1\\widetilde\{W\}\_\{1\}be the best rank\-one approximation toW~\\widetilde\{W\}in Frobenius norm:
W~1:=σ1uv⊤,σ1≥0,‖u‖2=‖v‖2=1\.\\widetilde\{W\}\_\{1\}:=\\sigma\_\{1\}uv^\{\\top\},\\qquad\\sigma\_\{1\}\\geq 0,\\qquad\\\|u\\\|\_\{2\}=\\\|v\\\|\_\{2\}=1\.SinceW~𝟏=0\\widetilde\{W\}\\mathbf\{1\}=0, we have
\(W~⊤W~\)𝟏=0\.\(\\widetilde\{W\}^\{\\top\}\\widetilde\{W\}\)\\mathbf\{1\}=0\.Ifσ1\>0\\sigma\_\{1\}\>0, thenvvis an eigenvector ofW~⊤W~\\widetilde\{W\}^\{\\top\}\\widetilde\{W\}with eigenvalueσ12\\sigma\_\{1\}^\{2\}, so
v⟂𝟏,Hv=v\.v\\perp\\mathbf\{1\},\\qquad Hv=v\.Ifσ1=0\\sigma\_\{1\}=0, thenW~1=0\\widetilde\{W\}\_\{1\}=0and the projected objective below is identically zero\.
Define
J1\(r\):=⟨Kc\(r\),W~1⊤W~1⟩F‖Kc\(r\)‖F,J\_\{1\}\(r\):=\\frac\{\\langle K\_\{c\}\(r\),\\widetilde\{W\}\_\{1\}^\{\\top\}\\widetilde\{W\}\_\{1\}\\rangle\_\{F\}\}\{\\\|K\_\{c\}\(r\)\\\|\_\{F\}\},withJ1\(r\)=0J\_\{1\}\(r\)=0whenever‖Kc\(r\)‖F=0\\\|K\_\{c\}\(r\)\\\|\_\{F\}=0\.
###### Lemma B\.3\(Reduction under rank\-one projection\)\.
For everyr∈ℛB↑r\\in\\mathcal\{R\}\_\{B\}^\{\\uparrow\}with‖Kc\(r\)‖F\>0\\\|K\_\{c\}\(r\)\\\|\_\{F\}\>0,
J1\(r\)=σ12Δ¯v\(r\),Δ¯v\(r\):=v⊤Kc\(r\)v‖Kc\(r\)‖F\.J\_\{1\}\(r\)=\\sigma\_\{1\}^\{2\}\\bar\{\\Delta\}\_\{v\}\(r\),\\qquad\\bar\{\\Delta\}\_\{v\}\(r\):=\\frac\{v^\{\\top\}K\_\{c\}\(r\)v\}\{\\\|K\_\{c\}\(r\)\\\|\_\{F\}\}\.Ifσ1\>0\\sigma\_\{1\}\>0, then
argmaxr∈ℛB↑:‖Kc\(r\)‖F\>0J1\(r\)=argmaxr∈ℛB↑:‖Kc\(r\)‖F\>0Δ¯v\(r\)\.\\arg\\max\_\{r\\in\\mathcal\{R\}\_\{B\}^\{\\uparrow\}:\\,\\\|K\_\{c\}\(r\)\\\|\_\{F\}\>0\}J\_\{1\}\(r\)=\\arg\\max\_\{r\\in\\mathcal\{R\}\_\{B\}^\{\\uparrow\}:\\,\\\|K\_\{c\}\(r\)\\\|\_\{F\}\>0\}\\bar\{\\Delta\}\_\{v\}\(r\)\.Ifσ1=0\\sigma\_\{1\}=0, thenJ1≡0J\_\{1\}\\equiv 0\.
###### Proof\.
SinceW~1=σ1uv⊤\\widetilde\{W\}\_\{1\}=\\sigma\_\{1\}uv^\{\\top\}with‖u‖2=1\\\|u\\\|\_\{2\}=1,
W~1⊤W~1=σ12vv⊤\.\\widetilde\{W\}\_\{1\}^\{\\top\}\\widetilde\{W\}\_\{1\}=\\sigma\_\{1\}^\{2\}vv^\{\\top\}\.Substituting intoJ1\(r\)J\_\{1\}\(r\)gives
J1\(r\)=σ12⟨Kc\(r\),vv⊤⟩F‖Kc\(r\)‖F=σ12v⊤Kc\(r\)v‖Kc\(r\)‖F=σ12Δ¯v\(r\)\.J\_\{1\}\(r\)=\\sigma\_\{1\}^\{2\}\\frac\{\\langle K\_\{c\}\(r\),vv^\{\\top\}\\rangle\_\{F\}\}\{\\\|K\_\{c\}\(r\)\\\|\_\{F\}\}=\\sigma\_\{1\}^\{2\}\\frac\{v^\{\\top\}K\_\{c\}\(r\)v\}\{\\\|K\_\{c\}\(r\)\\\|\_\{F\}\}=\\sigma\_\{1\}^\{2\}\\bar\{\\Delta\}\_\{v\}\(r\)\.The maximizer statement follows because multiplication by the positive constantσ12\\sigma\_\{1\}^\{2\}does not change the argmax\. Ifσ1=0\\sigma\_\{1\}=0, thenW~1⊤W~1=0\\widetilde\{W\}\_\{1\}^\{\\top\}\\widetilde\{W\}\_\{1\}=0, soJ1≡0J\_\{1\}\\equiv 0\. ∎
#### Threshold implication for the rank\-one direction\.
Assumeσ1\>0\\sigma\_\{1\}\>0\. SinceHv=vHv=v, Lemma[B\.2](https://arxiv.org/html/2606.17450#A2.Thmtheorem2)applies withw=vw=v\. For each split
Lj:=\{1,…,j\},Rj:=\{j\+1,…,n\},L\_\{j\}:=\\\{1,\\dots,j\\\},\\qquad R\_\{j\}:=\\\{j\+1,\\dots,n\\\},define
gj:=H𝟏Rj,ρj:=v⊤gj‖v‖2‖gj‖2\.g\_\{j\}:=H\\mathbf\{1\}\_\{R\_\{j\}\},\\qquad\\rho\_\{j\}:=\\frac\{v^\{\\top\}g\_\{j\}\}\{\\\|v\\\|\_\{2\}\\\|g\_\{j\}\\\|\_\{2\}\}\.Then every nontrivial monotone two\-level threshold on splitjjsatisfies
Δ¯v\(r\(j\)\)=‖v‖22ρj2\.\\bar\{\\Delta\}\_\{v\}\(r^\{\(j\)\}\)=\\\|v\\\|\_\{2\}^\{2\}\\rho\_\{j\}^\{2\}\.Thus the best threshold split is
j⋆∈argmax1≤j≤n−1ρj2,j^\{\\star\}\\in\\arg\\max\_\{1\\leq j\\leq n\-1\}\\rho\_\{j\}^\{2\},and it achieves a fraction
ρ⋆2:=max1≤j≤n−1ρj2\\rho\_\{\\star\}^\{2\}:=\\max\_\{1\\leq j\\leq n\-1\}\\rho\_\{j\}^\{2\}of the global upper bound for the rank\-one direction objective\. SinceJ1=σ12Δ¯vJ\_\{1\}=\\sigma\_\{1\}^\{2\}\\bar\{\\Delta\}\_\{v\}, the same split also maximizesJ1J\_\{1\}over two\-level thresholds\.
#### Uniform approximation bound for the rank\-one projection\.
Define
κmaxabs:=supd∈\[0,2B\]\|κ\(d\)\|<∞\.\\kappa\_\{\\max\}^\{\\mathrm\{abs\}\}:=\\sup\_\{d\\in\[0,2B\]\}\|\\kappa\(d\)\|<\\infty\.Let
E:=W~−W~1,e\(t\):=w~\(t\)−w~^\(t\),E:=\\widetilde\{W\}\-\\widetilde\{W\}\_\{1\},\\qquad e^\{\(t\)\}:=\\widetilde\{w\}^\{\(t\)\}\-\\widehat\{\\widetilde\{w\}\}^\{\(t\)\},where
w~^\(t\):=σ1utv\.\\widehat\{\\widetilde\{w\}\}^\{\(t\)\}:=\\sigma\_\{1\}u\_\{t\}v\.Whenσ1\>0\\sigma\_\{1\}\>0, we haveHv=vHv=v, sow~^\(t\)\\widehat\{\\widetilde\{w\}\}^\{\(t\)\}is centered\. Whenσ1=0\\sigma\_\{1\}=0,w~^\(t\)=0\\widehat\{\\widetilde\{w\}\}^\{\(t\)\}=0, so it is centered as well\. Fixε0\>0\\varepsilon\_\{0\}\>0and define
Jε0\(r\):=⟨Kc\(r\),W~⊤W~⟩Fmax\{‖Kc\(r\)‖F,ε0\},J1,ε0\(r\):=⟨Kc\(r\),W~1⊤W~1⟩Fmax\{‖Kc\(r\)‖F,ε0\}\.J\_\{\\varepsilon\_\{0\}\}\(r\):=\\frac\{\\langle K\_\{c\}\(r\),\\widetilde\{W\}^\{\\top\}\\widetilde\{W\}\\rangle\_\{F\}\}\{\\max\\\{\\\|K\_\{c\}\(r\)\\\|\_\{F\},\\varepsilon\_\{0\}\\\}\},\\qquad J\_\{1,\\varepsilon\_\{0\}\}\(r\):=\\frac\{\\langle K\_\{c\}\(r\),\\widetilde\{W\}\_\{1\}^\{\\top\}\\widetilde\{W\}\_\{1\}\\rangle\_\{F\}\}\{\\max\\\{\\\|K\_\{c\}\(r\)\\\|\_\{F\},\\varepsilon\_\{0\}\\\}\}\.
###### Theorem B\.4\(Uniform error bound forJε0J\_\{\\varepsilon\_\{0\}\}versusJ1,ε0J\_\{1,\\varepsilon\_\{0\}\}\)\.
For everyr∈ℛB↑r\\in\\mathcal\{R\}\_\{B\}^\{\\uparrow\},
\|Jε0\(r\)−J1,ε0\(r\)\|≤κmaxabsε0∑t=1T∑i,j=1n\|w~i\(t\)w~j\(t\)−w~^i\(t\)w~^j\(t\)\|\.\|J\_\{\\varepsilon\_\{0\}\}\(r\)\-J\_\{1,\\varepsilon\_\{0\}\}\(r\)\|\\leq\\frac\{\\kappa\_\{\\max\}^\{\\mathrm\{abs\}\}\}\{\\varepsilon\_\{0\}\}\\sum\_\{t=1\}^\{T\}\\sum\_\{i,j=1\}^\{n\}\\left\|\\widetilde\{w\}\_\{i\}^\{\(t\)\}\\widetilde\{w\}\_\{j\}^\{\(t\)\}\-\\widehat\{\\widetilde\{w\}\}\_\{i\}^\{\(t\)\}\\widehat\{\\widetilde\{w\}\}\_\{j\}^\{\(t\)\}\\right\|\.Moreover, for each tasktt,
∑i,j=1n\|w~i\(t\)w~j\(t\)−w~^i\(t\)w~^j\(t\)\|≤\(‖w~\(t\)‖1\+‖w~^\(t\)‖1\)‖e\(t\)‖1\.\\sum\_\{i,j=1\}^\{n\}\\left\|\\widetilde\{w\}\_\{i\}^\{\(t\)\}\\widetilde\{w\}\_\{j\}^\{\(t\)\}\-\\widehat\{\\widetilde\{w\}\}\_\{i\}^\{\(t\)\}\\widehat\{\\widetilde\{w\}\}\_\{j\}^\{\(t\)\}\\right\|\\leq\\left\(\\\|\\widetilde\{w\}^\{\(t\)\}\\\|\_\{1\}\+\\\|\\widehat\{\\widetilde\{w\}\}^\{\(t\)\}\\\|\_\{1\}\\right\)\\\|e^\{\(t\)\}\\\|\_\{1\}\.Consequently, the exact uniform gap
Δgap:=supr∈ℛB↑\|Jε0\(r\)−J1,ε0\(r\)\|\\Delta\_\{\\mathrm\{gap\}\}:=\\sup\_\{r\\in\\mathcal\{R\}\_\{B\}^\{\\uparrow\}\}\|J\_\{\\varepsilon\_\{0\}\}\(r\)\-J\_\{1,\\varepsilon\_\{0\}\}\(r\)\|satisfies
Δgap≤εgap,\\Delta\_\{\\mathrm\{gap\}\}\\leq\\varepsilon\_\{\\mathrm\{gap\}\},where the explicit upper bound is
εgap:=κmaxabsε0∑t=1T\(‖w~\(t\)‖1\+‖w~^\(t\)‖1\)‖e\(t\)‖1\.\\varepsilon\_\{\\mathrm\{gap\}\}:=\\frac\{\\kappa\_\{\\max\}^\{\\mathrm\{abs\}\}\}\{\\varepsilon\_\{0\}\}\\sum\_\{t=1\}^\{T\}\\left\(\\\|\\widetilde\{w\}^\{\(t\)\}\\\|\_\{1\}\+\\\|\\widehat\{\\widetilde\{w\}\}^\{\(t\)\}\\\|\_\{1\}\\right\)\\\|e^\{\(t\)\}\\\|\_\{1\}\.
###### Proof\.
If‖Kc\(r\)‖F=0\\\|K\_\{c\}\(r\)\\\|\_\{F\}=0, thenJε0\(r\)=J1,ε0\(r\)=0J\_\{\\varepsilon\_\{0\}\}\(r\)=J\_\{1,\\varepsilon\_\{0\}\}\(r\)=0\. Otherwise, sinceri∈\[−B,B\]r\_\{i\}\\in\[\-B,B\],
\|ri−rj\|∈\[0,2B\],\|κ\(\|ri−rj\|\)\|≤κmaxabs\.\|r\_\{i\}\-r\_\{j\}\|\\in\[0,2B\],\\qquad\|\\kappa\(\|r\_\{i\}\-r\_\{j\}\|\)\|\\leq\\kappa\_\{\\max\}^\{\\mathrm\{abs\}\}\.Becausew~\(t\)\\widetilde\{w\}^\{\(t\)\}andw~^\(t\)\\widehat\{\\widetilde\{w\}\}^\{\(t\)\}are centered,
\(w~\(t\)\)⊤Kc\(r\)w~\(t\)=\(w~\(t\)\)⊤K\(r\)w~\(t\),\(\\widetilde\{w\}^\{\(t\)\}\)^\{\\top\}K\_\{c\}\(r\)\\widetilde\{w\}^\{\(t\)\}=\(\\widetilde\{w\}^\{\(t\)\}\)^\{\\top\}K\(r\)\\widetilde\{w\}^\{\(t\)\},and similarly forw~^\(t\)\\widehat\{\\widetilde\{w\}\}^\{\(t\)\}\. Hence
Jε0\(r\)−J1,ε0\(r\)=∑t=1T∑i,j=1n\(w~i\(t\)w~j\(t\)−w~^i\(t\)w~^j\(t\)\)Kij\(r\)max\{‖Kc\(r\)‖F,ε0\}\.J\_\{\\varepsilon\_\{0\}\}\(r\)\-J\_\{1,\\varepsilon\_\{0\}\}\(r\)=\\frac\{\\sum\_\{t=1\}^\{T\}\\sum\_\{i,j=1\}^\{n\}\\left\(\\widetilde\{w\}\_\{i\}^\{\(t\)\}\\widetilde\{w\}\_\{j\}^\{\(t\)\}\-\\widehat\{\\widetilde\{w\}\}\_\{i\}^\{\(t\)\}\\widehat\{\\widetilde\{w\}\}\_\{j\}^\{\(t\)\}\\right\)K\_\{ij\}\(r\)\}\{\\max\\\{\\\|K\_\{c\}\(r\)\\\|\_\{F\},\\varepsilon\_\{0\}\\\}\}\.Taking absolute values and usingKij\(r\)=κ\(\|ri−rj\|\)K\_\{ij\}\(r\)=\\kappa\(\|r\_\{i\}\-r\_\{j\}\|\)proves the first bound\.
For the second bound, expand
w~i\(t\)w~j\(t\)−w~^i\(t\)w~^j\(t\)=w~i\(t\)ej\(t\)\+w~^j\(t\)ei\(t\)\.\\widetilde\{w\}\_\{i\}^\{\(t\)\}\\widetilde\{w\}\_\{j\}^\{\(t\)\}\-\\widehat\{\\widetilde\{w\}\}\_\{i\}^\{\(t\)\}\\widehat\{\\widetilde\{w\}\}\_\{j\}^\{\(t\)\}=\\widetilde\{w\}\_\{i\}^\{\(t\)\}e\_\{j\}^\{\(t\)\}\+\\widehat\{\\widetilde\{w\}\}\_\{j\}^\{\(t\)\}e\_\{i\}^\{\(t\)\}\.Therefore
\|w~i\(t\)w~j\(t\)−w~^i\(t\)w~^j\(t\)\|≤\|w~i\(t\)\|\|ej\(t\)\|\+\|w~^j\(t\)\|\|ei\(t\)\|\.\\left\|\\widetilde\{w\}\_\{i\}^\{\(t\)\}\\widetilde\{w\}\_\{j\}^\{\(t\)\}\-\\widehat\{\\widetilde\{w\}\}\_\{i\}^\{\(t\)\}\\widehat\{\\widetilde\{w\}\}\_\{j\}^\{\(t\)\}\\right\|\\leq\|\\widetilde\{w\}\_\{i\}^\{\(t\)\}\|\\,\|e\_\{j\}^\{\(t\)\}\|\+\|\\widehat\{\\widetilde\{w\}\}\_\{j\}^\{\(t\)\}\|\\,\|e\_\{i\}^\{\(t\)\}\|\.Summing overi,ji,jgives
∑i,j=1n\|w~i\(t\)w~j\(t\)−w~^i\(t\)w~^j\(t\)\|≤\(‖w~\(t\)‖1\+‖w~^\(t\)‖1\)‖e\(t\)‖1\.\\sum\_\{i,j=1\}^\{n\}\\left\|\\widetilde\{w\}\_\{i\}^\{\(t\)\}\\widetilde\{w\}\_\{j\}^\{\(t\)\}\-\\widehat\{\\widetilde\{w\}\}\_\{i\}^\{\(t\)\}\\widehat\{\\widetilde\{w\}\}\_\{j\}^\{\(t\)\}\\right\|\\leq\\left\(\\\|\\widetilde\{w\}^\{\(t\)\}\\\|\_\{1\}\+\\\|\\widehat\{\\widetilde\{w\}\}^\{\(t\)\}\\\|\_\{1\}\\right\)\\\|e^\{\(t\)\}\\\|\_\{1\}\.Substituting this into the first bound yields the stated uniform bound\. ∎
#### Consequence: near\-optimality of the projected maximizer for the floored objective\.
###### Corollary B\.5\.
Let
r1⋆∈argmaxr∈ℛB↑J1,ε0\(r\)\.r\_\{1\}^\{\\star\}\\in\\arg\\max\_\{r\\in\\mathcal\{R\}\_\{B\}^\{\\uparrow\}\}J\_\{1,\\varepsilon\_\{0\}\}\(r\)\.LetΔgap\\Delta\_\{\\mathrm\{gap\}\}andεgap\\varepsilon\_\{\\mathrm\{gap\}\}be as in Theorem[B\.4](https://arxiv.org/html/2606.17450#A2.Thmtheorem4)\. Then
Jε0\(r1⋆\)≥supr∈ℛB↑Jε0\(r\)−2εgap\.J\_\{\\varepsilon\_\{0\}\}\(r\_\{1\}^\{\\star\}\)\\geq\\sup\_\{r\\in\\mathcal\{R\}\_\{B\}^\{\\uparrow\}\}J\_\{\\varepsilon\_\{0\}\}\(r\)\-2\\varepsilon\_\{\\mathrm\{gap\}\}\.
###### Proof\.
By definition ofΔgap\\Delta\_\{\\mathrm\{gap\}\}, for everyr∈ℛB↑r\\in\\mathcal\{R\}\_\{B\}^\{\\uparrow\},
Jε0\(r\)≥J1,ε0\(r\)−Δgap,J1,ε0\(r\)≥Jε0\(r\)−Δgap\.J\_\{\\varepsilon\_\{0\}\}\(r\)\\geq J\_\{1,\\varepsilon\_\{0\}\}\(r\)\-\\Delta\_\{\\mathrm\{gap\}\},\\qquad J\_\{1,\\varepsilon\_\{0\}\}\(r\)\\geq J\_\{\\varepsilon\_\{0\}\}\(r\)\-\\Delta\_\{\\mathrm\{gap\}\}\.By optimality ofr1⋆r\_\{1\}^\{\\star\}forJ1,ε0J\_\{1,\\varepsilon\_\{0\}\}, for everyr∈ℛB↑r\\in\\mathcal\{R\}\_\{B\}^\{\\uparrow\},
J1,ε0\(r1⋆\)≥J1,ε0\(r\)\.J\_\{1,\\varepsilon\_\{0\}\}\(r\_\{1\}^\{\\star\}\)\\geq J\_\{1,\\varepsilon\_\{0\}\}\(r\)\.Therefore, for everyr∈ℛB↑r\\in\\mathcal\{R\}\_\{B\}^\{\\uparrow\},
Jε0\(r1⋆\)≥J1,ε0\(r1⋆\)−Δgap≥J1,ε0\(r\)−Δgap≥Jε0\(r\)−2Δgap\.J\_\{\\varepsilon\_\{0\}\}\(r\_\{1\}^\{\\star\}\)\\geq J\_\{1,\\varepsilon\_\{0\}\}\(r\_\{1\}^\{\\star\}\)\-\\Delta\_\{\\mathrm\{gap\}\}\\geq J\_\{1,\\varepsilon\_\{0\}\}\(r\)\-\\Delta\_\{\\mathrm\{gap\}\}\\geq J\_\{\\varepsilon\_\{0\}\}\(r\)\-2\\Delta\_\{\\mathrm\{gap\}\}\.Since this holds for everyr∈ℛB↑r\\in\\mathcal\{R\}\_\{B\}^\{\\uparrow\}, it also holds after taking the supremum overrr\. UsingΔgap≤εgap\\Delta\_\{\\mathrm\{gap\}\}\\leq\\varepsilon\_\{\\mathrm\{gap\}\}, we obtain
Jε0\(r1⋆\)≥supr∈ℛB↑Jε0\(r\)−2εgap\.J\_\{\\varepsilon\_\{0\}\}\(r\_\{1\}^\{\\star\}\)\\geq\\sup\_\{r\\in\\mathcal\{R\}\_\{B\}^\{\\uparrow\}\}J\_\{\\varepsilon\_\{0\}\}\(r\)\-2\\varepsilon\_\{\\mathrm\{gap\}\}\.∎
###### Lemma B\.6\(Approximate noise\-averaging for normalized nHSIC\)\.
Condition on a fixed taskttand the fixed data\(r,y\(t\)\)\(r,y^\{\(t\)\}\)\. Let
si=ri\+ηis\_\{i\}=r\_\{i\}\+\\eta\_\{i\}with i\.i\.d\. noise\. Let
K¯c:=𝔼η\[Kc\(s\)\]\\bar\{K\}\_\{c\}:=\\mathbb\{E\}\_\{\\eta\}\[K\_\{c\}\(s\)\]and
nHSIC¯\(r,y\(t\)\):=⟨K¯c,Lt,c⟩F‖K¯c‖F‖Lt,c‖F\.\\overline\{\\mathrm\{nHSIC\}\}\\\!\\left\(r,y^\{\(t\)\}\\right\):=\\frac\{\\langle\\bar\{K\}\_\{c\},L\_\{t,c\}\\rangle\_\{F\}\}\{\\\|\\bar\{K\}\_\{c\}\\\|\_\{F\}\\\|L\_\{t,c\}\\\|\_\{F\}\}\.Assume‖Lt,c‖F\>0\\\|L\_\{t,c\}\\\|\_\{F\}\>0and‖K¯c‖F≥γ\>0\\\|\\bar\{K\}\_\{c\}\\\|\_\{F\}\\geq\\gamma\>0\. On any event where
‖Kc\(s\)−K¯c‖F≤γ/2,\\\|K\_\{c\}\(s\)\-\\bar\{K\}\_\{c\}\\\|\_\{F\}\\leq\\gamma/2,we have both‖Kc\(s\)‖F≥γ/2\\\|K\_\{c\}\(s\)\\\|\_\{F\}\\geq\\gamma/2and
\|nHSIC\(s,y\(t\)\)−nHSIC¯\(r,y\(t\)\)\|≤4γ‖Kc\(s\)−K¯c‖F\.\\left\|\\mathrm\{nHSIC\}\\\!\\left\(s,y^\{\(t\)\}\\right\)\-\\overline\{\\mathrm\{nHSIC\}\}\\\!\\left\(r,y^\{\(t\)\}\\right\)\\right\|\\leq\\frac\{4\}\{\\gamma\}\\\|K\_\{c\}\(s\)\-\\bar\{K\}\_\{c\}\\\|\_\{F\}\.
###### Proof\.
Let
F\(K\):=⟨K,Lt,c⟩F‖K‖F\.F\(K\):=\\frac\{\\langle K,L\_\{t,c\}\\rangle\_\{F\}\}\{\\\|K\\\|\_\{F\}\}\.On the event‖Kc\(s\)−K¯c‖F≤γ/2\\\|K\_\{c\}\(s\)\-\\bar\{K\}\_\{c\}\\\|\_\{F\}\\leq\\gamma/2, the reverse triangle inequality gives
‖Kc\(s\)‖F≥‖K¯c‖F−‖Kc\(s\)−K¯c‖F≥γ/2\.\\\|K\_\{c\}\(s\)\\\|\_\{F\}\\geq\\\|\\bar\{K\}\_\{c\}\\\|\_\{F\}\-\\\|K\_\{c\}\(s\)\-\\bar\{K\}\_\{c\}\\\|\_\{F\}\\geq\\gamma/2\.Now take anyK,K′K,K^\{\\prime\}with‖K‖F,‖K′‖F≥γ/2\\\|K\\\|\_\{F\},\\\|K^\{\\prime\}\\\|\_\{F\}\\geq\\gamma/2\. Then
\|F\(K\)−F\(K′\)\|\\displaystyle\|F\(K\)\-F\(K^\{\\prime\}\)\|=\|⟨K,Lt,c⟩F‖K‖F−⟨K′,Lt,c⟩F‖K′‖F\|\\displaystyle=\\left\|\\frac\{\\langle K,L\_\{t,c\}\\rangle\_\{F\}\}\{\\\|K\\\|\_\{F\}\}\-\\frac\{\\langle K^\{\\prime\},L\_\{t,c\}\\rangle\_\{F\}\}\{\\\|K^\{\\prime\}\\\|\_\{F\}\}\\right\|≤\|⟨K−K′,Lt,c⟩F‖K‖F\|\+\|⟨K′,Lt,c⟩F\|\|1‖K‖F−1‖K′‖F\|\\displaystyle\\leq\\left\|\\frac\{\\langle K\-K^\{\\prime\},L\_\{t,c\}\\rangle\_\{F\}\}\{\\\|K\\\|\_\{F\}\}\\right\|\+\|\\langle K^\{\\prime\},L\_\{t,c\}\\rangle\_\{F\}\|\\left\|\\frac\{1\}\{\\\|K\\\|\_\{F\}\}\-\\frac\{1\}\{\\\|K^\{\\prime\}\\\|\_\{F\}\}\\right\|≤2γ‖K−K′‖F‖Lt,c‖F\+‖Lt,c‖F2γ‖K−K′‖F\\displaystyle\\leq\\frac\{2\}\{\\gamma\}\\\|K\-K^\{\\prime\}\\\|\_\{F\}\\\|L\_\{t,c\}\\\|\_\{F\}\+\\\|L\_\{t,c\}\\\|\_\{F\}\\frac\{2\}\{\\gamma\}\\\|K\-K^\{\\prime\}\\\|\_\{F\}=4‖Lt,c‖Fγ‖K−K′‖F\.\\displaystyle=\\frac\{4\\\|L\_\{t,c\}\\\|\_\{F\}\}\{\\gamma\}\\\|K\-K^\{\\prime\}\\\|\_\{F\}\.In the second term, we used
\|⟨K′,Lt,c⟩F\|≤‖K′‖F‖Lt,c‖F\|\\langle K^\{\\prime\},L\_\{t,c\}\\rangle\_\{F\}\|\\leq\\\|K^\{\\prime\}\\\|\_\{F\}\\\|L\_\{t,c\}\\\|\_\{F\}and
\|1‖K‖F−1‖K′‖F\|=\|‖K′‖F−‖K‖F\|‖K‖F‖K′‖F≤‖K−K′‖F‖K‖F‖K′‖F\.\\left\|\\frac\{1\}\{\\\|K\\\|\_\{F\}\}\-\\frac\{1\}\{\\\|K^\{\\prime\}\\\|\_\{F\}\}\\right\|=\\frac\{\\big\|\\\|K^\{\\prime\}\\\|\_\{F\}\-\\\|K\\\|\_\{F\}\\big\|\}\{\\\|K\\\|\_\{F\}\\\|K^\{\\prime\}\\\|\_\{F\}\}\\leq\\frac\{\\\|K\-K^\{\\prime\}\\\|\_\{F\}\}\{\\\|K\\\|\_\{F\}\\\|K^\{\\prime\}\\\|\_\{F\}\}\.TakingK=Kc\(s\)K=K\_\{c\}\(s\)andK′=K¯cK^\{\\prime\}=\\bar\{K\}\_\{c\}, then dividing by‖Lt,c‖F\>0\\\|L\_\{t,c\}\\\|\_\{F\}\>0, convertsF\(Kc\(s\)\)F\(K\_\{c\}\(s\)\)intonHSIC\(s,y\(t\)\)\\mathrm\{nHSIC\}\(s,y^\{\(t\)\}\)andF\(K¯c\)F\(\\bar\{K\}\_\{c\}\)intonHSIC¯\(r,y\(t\)\)\\overline\{\\mathrm\{nHSIC\}\}\(r,y^\{\(t\)\}\)\. This proves the result\. ∎
## Appendix CAdditional Experiments: Single\-Index Baselines and nHSIC Ablations
We report additional dependence\-based evaluations to further compare MLCI against BCE\-trained single\-index models and to assess the sensitivity of the proposed nHSIC framework to encoder architecture and score\-kernel choice\. The BCE\-trained models serve as additional predictive baselines: they use the same diagnosis\-code inputs and produce scalar scores, but they are optimized with supervised binary cross\-entropy objectives rather than the proposed multi\-outcome nHSIC objective\. In contrast, the nHSIC models in this appendix should be interpreted as ablations of the MLCI framework rather than as external competing baselines, since they share the same goal of learning a single scalar diagnosis\-code index by maximizing multi\-outcome nHSIC\.
The default MLCI configuration uses the DeepSets mean⊕\\oplusmax encoder with a single RBF score kernel\. This choice is motivated by the structure and intended use of comorbidity scores\. Diagnosis codes form unordered sets, so a permutation\-invariant DeepSets encoder is appropriate\. Mean pooling summarizes distributed comorbidity burden across the diagnosis set, while max pooling can retain strong diagnosis\-level feature activations, including signals from rare or high\-severity codes\. Combining the two gives a compact representation that captures both broad disease burden and salient severe\-code signals, while preserving the simplicity and single\-score form expected of a comorbidity index\.
Multi\-RBF kernels and higher\-capacity encoder variants can improve selected dependence metrics, showing that the nHSIC framework is extensible\. However, these variants introduce additional architectural and kernel\-design choices\. We therefore use the DeepSets mean⊕\\oplusmax encoder with a single RBF kernel as the default MLCI configuration, and report the remaining nHSIC variants here as architecture and kernel ablations\.
All models reported in this section are single\-index models: each maps diagnosis\-code inputs to one scalar score, and we evaluate the dependence between that scalar score and each clinical outcome\. Specifically, we report mutual information \(MI\) between each model’s scalar score and each outcome, as well as distance correlation \(dCorr\), which captures both linear and nonlinear dependence\. All results are computed on the same evaluation splits\. Per\-outcome MI and dCorr values are computed using the corresponding outcome’s valid\-label mask\.
Table 6:Distance Correlation between each model’s unified risk score and clinical outcomes\.Higher is better\. Bolded values indicate the best performer per outcome; underlined values indicate the second\-best performer\.For readability, each column is scaled by the power of 10 shown in the header; divide by that factor to recover the original dCorr\.Table 7:Mutual Information between each model’s unified risk score and clinical outcomes\.Higher is better\. Bolded values highlight the best performer per outcome; underlined values highlight the second\-best performer\.For readability, each column is scaled by the power of 10 shown in the header; divide by that factor to recover the original MI\.
## Appendix DComorbidity indices and scoring details
### D\.1Input data and coding algorithms
Letiiindex admissions\. Each admission has a set or sequence of diagnosis codes
Xi=\(di1,…,dimi\),X\_\{i\}=\(d\_\{i1\},\\ldots,d\_\{im\_\{i\}\}\),wheredijd\_\{ij\}is an ICD code andmim\_\{i\}is the number of recorded diagnosis codes for admissionii\. Unless otherwise noted, we use diagnosis codes recorded for the hospital admission to summarize admission\-level comorbidity burden; these administrative codes may include diagnoses finalized after discharge and should not be interpreted as strictly present\-on\-admission measurements\. CCI and Elixhauser comorbidities were assigned using Quan et al\.’s coding algorithms\(Quanet al\.,[2005](https://arxiv.org/html/2606.17450#bib.bib53)\): ICD\-10 diagnoses used Quan’s ICD\-10 code lists, and ICD\-9\-CM diagnoses used the corresponding enhanced ICD\-9\-CM mappings\.
### D\.2Charlson Comorbidity Index \(CCI\)
The Charlson Comorbidity Index \(CCI\) is a weighted summary of chronic disease burden, with higher scores indicating greater comorbidity and mortality risk\. Following standard implementations, including Quan’s mapping, ICD codes are first mapped into Charlson comorbidity categories, and each category contributes at most once\.
Letk∈\{1,…,K\}k\\in\\\{1,\\ldots,K\\\}index Charlson comorbidity categories, and define the category indicator
cikCCI=𝕀\{any ICD code inXimaps to Charlson categoryk\}\.c\_\{ik\}^\{\\mathrm\{CCI\}\}=\\mathbb\{I\}\\\{\\text\{any ICD code in \}X\_\{i\}\\text\{ maps to Charlson category \}k\\\}\.Letak∈\{1,2,3,6\}a\_\{k\}\\in\\\{1,2,3,6\\\}denote the Charlson weight for categorykk\. Categories not present contribute0\. The CCI for admissioniiis
CCIi=∑k=1KakcikCCI\.\\mathrm\{CCI\}\_\{i\}=\\sum\_\{k=1\}^\{K\}a\_\{k\}\\,c\_\{ik\}^\{\\mathrm\{CCI\}\}\.\(6\)
### D\.3Elixhauser comorbidities and the van Walraven weighted summary \(ECI\)
Elixhauser et al\. \(1998\) represent comorbidity burden asQQbinary indicators included separately as covariates, commonlyQ=30Q=30\. Letq∈\{1,…,Q\}q\\in\\\{1,\\ldots,Q\\\}index Elixhauser comorbidity groups, and define
eiqECI=𝕀\{any ICD code inXimaps to Elixhauser groupq\}\.e\_\{iq\}^\{\\mathrm\{ECI\}\}=\\mathbb\{I\}\\\{\\text\{any ICD code in \}X\_\{i\}\\text\{ maps to Elixhauser group \}q\\\}\.The original Elixhauser approach defines no single summary score; rather, theQQindicators are entered separately\.
Van Walraven et al\. summarize theQQElixhauser indicators into a single weighted score by converting mortality\-model coefficients into integer weights\. Letwqw\_\{q\}denote the van Walraven weight for groupqq\. The van Walraven weighted Elixhauser score is
ECIVW,i=∑q=1QwqeiqECI\.\\mathrm\{ECI\}\_\{\\mathrm\{VW\},i\}=\\sum\_\{q=1\}^\{Q\}w\_\{q\}\\,e\_\{iq\}^\{\\mathrm\{ECI\}\}\.\(7\)Across this paper, the termElixhauser Comorbidity Index \(ECI\)refers to the van Walraven weighted summary in \([7](https://arxiv.org/html/2606.17450#A4.E7)\)\.
### D\.4Common input–output formulation
To make the connection between classic indices and our approach explicit, we describe CCI and ECI through a common input–output lens\. Both CCI and the van Walraven ECI map the admission\-level diagnosis historyXiX\_\{i\}, restricted to ICD codes, to a scalar scoreci∈ℝc\_\{i\}\\in\\mathbb\{R\}by \(i\) forming comorbidity\-category indicators fromXiX\_\{i\}and \(ii\) summing them with fixed weights to summarize baseline illness burden\.
## Appendix ERisk Curves for the Learned Severity Score
This appendix visualizes how the learned scalar scoresi=sθ\(Xi\)s\_\{i\}=s\_\{\\theta\}\(X\_\{i\}\)relates to empirical outcome risk across the four evaluation tasks: in\-hospital mortality, 30\-day mortality, length of stay\>7\>7days, and ICU transfer\. Because our theory is ordering\-based, we focus on the mapping from the learned score ordering to empirical outcome rates\. In particular, for each tasktt, we plot estimates of
pt\(s\)≈Pr\{yi\(t\)=1∣si=s\}p\_\{t\}\(s\)\\approx\\Pr\\\{y\_\{i\}^\{\(t\)\}=1\\mid s\_\{i\}=s\\\}as a function of the learned score\.
#### Estimators\.
We report two complementary curve estimators: \(i\) anunconstrained binnedestimate, which partitions test scores into quantile bins and plots the empirical event rate per bin; and \(ii\) amonotone isotonicestimate, which fits a nondecreasing function of the score using isotonic regression on valid labels for the given task\. The binned curves reveal local non\-monotonicities due to finite\-sample noise and label sparsity, while the isotonic curves summarize the best monotone risk calibration implied by the learned ordering\.
#### Validity masking\.
For each tasktt, curves are computed over the subset of test admissions with valid labels,\{i:Mi\(t\)=1\}\\\{i:M\_\{i\}^\{\(t\)\}=1\\\}\. For length of stay, validity is defined by the dedicated indicator columnlong\_stay\_defined\. This matches the evaluation protocol used in the main paper\.
#### Summary statistics\.
Table[8](https://arxiv.org/html/2606.17450#A5.T8)summarizes whether risk generally increases with the learned score for each outcome and reports the end\-to\-end change in the monotone fit\.
Table 8:Risk\-curve summary from isotonic fits ofPr\{yi\(t\)=1∣si=s\}\\Pr\\\{y\_\{i\}^\{\(t\)\}=1\\mid s\_\{i\}=s\\\}as a function of the learned score\. We report a qualitative monotonicity label \(inc\./weakly inc\.\) and the end\-to\-end changeΔ:=p95−p5\\Delta:=p\_\{95\}\-p\_\{5\}, wherepqp\_\{q\}denotes theqq\-th percentile of the fitted risk values along the score axis\. For both datasets,Δ\\Deltais reported as mean±\\pmstandard deviation over seeds\{11,101,1001\}\\\{11,101,1001\\\}\.
#### Figures\.
Figures[2](https://arxiv.org/html/2606.17450#A5.F2)–[5](https://arxiv.org/html/2606.17450#A5.F5)report per\-task risk curves for MIMIC\-IV and MIMIC\-III\. For each dataset, we show \(i\)unconstrained binnedempirical curves using quantile bins, to visualize sampling noise and possible local non\-monotonicities, and \(ii\)isotonicfits, to estimate the best nondecreasing risk mapping implied by the learned ordering\. The plotted curves correspond to seed 11; Table[8](https://arxiv.org/html/2606.17450#A5.T8)reports the correspondingΔ\\Deltasummary across seeds\{11,101,1001\}\\\{11,101,1001\\\}\.
\(a\)Mortality
\(b\)30\-day mortality
\(c\)Length of stay
\(d\)ICU transfer
Figure 2:MIMIC\-IV: binned \(unconstrained\) risk curvesPr\{yi\(t\)=1∣si=s\}\\Pr\\\{y\_\{i\}^\{\(t\)\}=1\\mid s\_\{i\}=s\\\}as a function of the learned score \(60 quantile bins\)\.\(a\)Mortality
\(b\)30\-day mortality
\(c\)Length of stay
\(d\)ICU transfer
Figure 3:MIMIC\-IV: isotonic \(monotone\) risk curvesPr\{yi\(t\)=1∣si=s\}\\Pr\\\{y\_\{i\}^\{\(t\)\}=1\\mid s\_\{i\}=s\\\}fit via isotonic regression on valid labels for each task\.\(a\)Mortality
\(b\)30\-day mortality
\(c\)Length of stay
\(d\)ICU transfer
Figure 4:MIMIC\-III: binned \(unconstrained\) risk curvesPr\{yi\(t\)=1∣si=s\}\\Pr\\\{y\_\{i\}^\{\(t\)\}=1\\mid s\_\{i\}=s\\\}as a function of the learned score \(60 quantile bins\)\.\(a\)Mortality
\(b\)30\-day mortality
\(c\)Length of stay
\(d\)ICU transfer
Figure 5:MIMIC\-III: isotonic \(monotone\) risk curvesPr\{yi\(t\)=1∣si=s\}\\Pr\\\{y\_\{i\}^\{\(t\)\}=1\\mid s\_\{i\}=s\\\}fit via isotonic regression on valid labels for each task\.Similar Articles
Primary ICD Category Prediction using LLM-based Probing
This paper presents a method that uses frozen medical large language model (LLM) representations as a shared embedding space to predict primary ICD diagnosis categories from both structured and unstructured electronic health record data, achieving improved accuracy over baseline methods on MIMIC-IV and showing transferability to MIMIC-III.
Multi-Modal Machine Learning for Breast Cancer Recurrence Prediction
This paper examines the integration of multi-modal clinical data, including treatment records, pathology reports, and clinician notes, using rule-based extraction and machine learning to improve breast cancer recurrence prediction compared to single-modal approaches.
Language Models as Interfaces, Not Oracles: A Hybrid LLM-ML System for Pediatric Appendicitis
This paper presents ClaMPAPP, a hybrid architecture that uses an LLM as an interface to extract features from clinical narratives, which are then passed to an XGBoost classifier for pediatric appendicitis diagnosis, demonstrating improved robustness and safety over end-to-end LLM baselines.
ClinicalMC: A Benchmark for Multi-Course Clinical Decision-Making with Large Language Models
ClinicalMC is a benchmark designed to evaluate large language models in multi-course clinical decision-making, featuring datasets in Chinese and English and a multi-agent evaluation framework.
LLMs for Cardiovascular Risk Prediction from Structured Clinical Data
This paper presents a hybrid framework that combines structured clinical data with LLM-generated narratives for coronary artery disease prediction, achieving high fidelity in variable extraction and comparing ML models with LLM-based zero-shot and few-shot classification.