HCOE:基于生物医学语言模型的双曲临床本体嵌入

arXiv cs.AI 论文

摘要

本文介绍了双曲临床本体嵌入(HCOE),这是一种将生物医学语言模型嵌入映射到双曲空间的方法,能更好地捕捉医学代码的层级结构,并在临床任务上优于现有方法。

arXiv:2609.30763v1 Announce Type: new Abstract: Biomedical language models (LMs) encode textual semantics but do not explicitly preserve medical code hierarchies. We present Hyperbolic Clinical Ontology Embeddings (HCOE) for hierarchy-aware clinical concept representation. HCOE maps frozen BioBERT embeddings into a Poincare ball, combining parent-side and child-side ontology-guided contrastive learning with coarse-to-fine ontology-path aggregation. It uses International Classification of Diseases (ICD) codes organized by Clinical Classifications Software (CCS) and Anatomical Therapeutic Chemical (ATC) medication hierarchies. Evaluations show that HCOE performs best on ICD/ATC clinical relation prediction and CCS-to-PheCode hierarchy transfer. On the MIMIC-IV dataset, HCOE also achieves the best performance on mortality prediction, readmission prediction, medication recommendation, and rare drug prediction.
查看原文
查看缓存全文

缓存时间: 2026/09/28 09:46

# Hyperbolic Clinical Ontology Embeddings from Biomedical Language Models
Source: [https://arxiv.org/html/2609.30763](https://arxiv.org/html/2609.30763)
## HCOE: Hyperbolic Clinical Ontology Embeddings from Biomedical Language Models Thanks:This work was supported in part by the NVIDIA Academic Grant Program and the Google Cloud Research Credits program\.

Weihao LiAffiliation:Northwestern UniversityZiyang Song\*Affiliation:Ohio University

###### Abstract

Biomedical language models \(LMs\) encode textual semantics but do not explicitly preserve medical code hierarchies\. We presentHyperbolic Clinical Ontology Embeddings \(HCOE\)for hierarchy\-aware clinical concept representation\. HCOE maps frozen BioBERT embeddings into a Poincar´e ball, combining parent\-side and child\-side ontology\-guided contrastive learning with coarse\-to\-fine ontology\-path aggregation\. It uses International Classification of Diseases \(ICD\) codes organized by Clinical Classifications Software \(CCS\) and Anatomical Therapeutic Chemical \(ATC\) medication hierarchies\. Evaluations show that HCOE performs best on ICD/ATC clinical relation prediction and CCS\-to\-PheCode hierarchy transfer\. On the MIMIC\-IV dataset, HCOE also achieves the best performance on mortality prediction, readmission prediction, medication recommendation, and rare drug prediction\.

###### Index Terms:

Clinical ontology embeddings, biomedical language models, hyperbolic representation learning, ontology\-guided contrastive learning, electronic health records\.

††footnotetext:\*Corresponding author:ziyangs@ohio\.edu\. Accepted at the 2026 IEEE International Conference on Bioinformatics and Biomedicine \(BIBM\)\.## IIntroduction

In electronic health record \(EHR\) modeling, diagnosis, procedure, and medication codes are used to represent patient visits and longitudinal histories for tasks such as diagnosis prediction, drug recommendation, and patient risk stratification\[[1](https://arxiv.org/html/2609.30763#bib.bib1),[2](https://arxiv.org/html/2609.30763#bib.bib2),[3](https://arxiv.org/html/2609.30763#bib.bib3)\]\. Biomedical language models \(LMs\) provide transferable textual representation of medical concepts\[[4](https://arxiv.org/html/2609.30763#bib.bib4)\], but their pretraining does not explicitly capture structural relationships among clinical concepts\. The International Classification of Diseases \(ICD\) and Anatomical Therapeutic Chemical \(ATC\) classification organize diagnoses, procedures, and medications into multi\-level clinical coding hierarchies, ranging from broad categories to fine\-grained codes\[[5](https://arxiv.org/html/2609.30763#bib.bib5),[6](https://arxiv.org/html/2609.30763#bib.bib6)\]\. These clinical ontologies contain clinically meaningful coarse\-to\-fine relationships that are important for clinical concept representation\[[7](https://arxiv.org/html/2609.30763#bib.bib22)\]\.

![Refer to caption](https://arxiv.org/html/2609.30763v1/figures/motivation.png)Fig\. 1:Comparison of 1\-hop parent–child relation prediction in the ICD hierarchy\.We assess whether BioBERT and HCOE embed each child concept closer to its true parent than to unrelated concepts across five levels\. HCOE consistently outperforms BioBERT, with larger gains at finer\-grained levels\.To motivate the need for hierarchy\-aware clinical concept embeddings, we evaluate BioBERT, a biomedical LM pre\-trained on PubMed abstracts\[[4](https://arxiv.org/html/2609.30763#bib.bib4)\], on 1\-hop parent–child relation prediction in the ICD hierarchy\. The task tests whether each child concept is embedded closer to its true parent than to unrelated concepts\. As shown in Fig\.[1](https://arxiv.org/html/2609.30763#S1.F1), BioBERT performs reasonably well at coarse hierarchy levels, but its accuracy decreases as the hierarchy becomes more fine\-grained, reaching 58\.6% at the finest level\. In contrast, our proposed method maintains higher accuracy across levels\. These results suggest that biomedical LM semantics alone are insufficient to preserve fine\-grained ICD hierarchy relations, motivating hierarchy\-aware clinical concept embeddings that incorporate clinical ontology structure\.

In EHR modeling, deep learning methods learn clinical code embeddings from co\-occurrence patterns and longitudinal patient histories to support clinical prediction tasks\[[1](https://arxiv.org/html/2609.30763#bib.bib1),[8](https://arxiv.org/html/2609.30763#bib.bib16)\]\. To incorporate domain knowledge, several methods further extend this paradigm with clinical ontologies or graph structures for more informative patient and concept representations\[[9](https://arxiv.org/html/2609.30763#bib.bib14),[10](https://arxiv.org/html/2609.30763#bib.bib15),[3](https://arxiv.org/html/2609.30763#bib.bib3)\]\. However, these methods are often trained from scratch for specific prediction tasks and do not exploit the transferable knowledge learned from large\-scale biomedical corpora\. Biomedical LMs address this limitation by pretraining on large\-scale biomedical data, providing transferable textual semantics for medical concepts, but they are not explicitly optimized to distinguish hierarchical relations among clinical codes\[[4](https://arxiv.org/html/2609.30763#bib.bib4),[11](https://arxiv.org/html/2609.30763#bib.bib19)\]\. Recent hierarchy\-aware LMs incorporate structured knowledge to learn more hierarchy\-aware representations, but their Euclidean embeddings are less suitable for representing clinical ontologies\[[12](https://arxiv.org/html/2609.30763#bib.bib20)\]\. In addition, hyperbolic embedding methods provide a natural geometry for modeling hierarchical relations, but they generally do not leverage pre\-trained LM representations\[[13](https://arxiv.org/html/2609.30763#bib.bib7),[14](https://arxiv.org/html/2609.30763#bib.bib9),[15](https://arxiv.org/html/2609.30763#bib.bib12)\]\. This motivates integrating pre\-trained biomedical LM representations with clinical ontologies in hyperbolic space for hierarchy\-aware clinical concept representation\.

We presentHyperbolic Clinical Ontology Embeddings \(HCOE\), a hyperbolic representation learning framework for hierarchy\-aware clinical concept embeddings\. HCOE learns parent–child relations from CCS\-organized ICD codes and ATC hierarchies using ontology\-guided contrastive learning, then aggregates embeddings along each coarse\-to\-fine ontology path\. We evaluate clinical relation prediction, CCS\-to\-PheCode hierarchy transfer, and downstream MIMIC\-IV tasks\. The main contributions are:

1. 1\.We develop HCOE for hierarchy\-aware medical code embeddings from frozen biomedical LM representations and clinical ontologies\.
2. 2\.We introduce ontology\-guided hyperbolic contrastive learning and coarse\-to\-fine Möbius path aggregation\.
3. 3\.We demonstrate the advantage of HCOE on ICD/ATC relation prediction, CCS\-to\-PheCode hierarchy transfer, and MIMIC\-IV mortality prediction, readmission prediction, medication recommendation, and rare drug prediction\.

## IIRelated Work

Clinical concept and medical code representation has been widely studied for EHR modeling\. Early methods learn medical code embeddings from co\-occurrence patterns and longitudinal patient records, such as cui2vec\[[8](https://arxiv.org/html/2609.30763#bib.bib16)\]\. To incorporate domain knowledge, ontology\- and graph\-enhanced representation methods further use clinical hierarchies or graph structures to improve clinical concept embeddings\. For example, GRAM uses ontology ancestors with attention to learn clinically meaningful medical code representations\[[9](https://arxiv.org/html/2609.30763#bib.bib14)\], while RotatE represents relations in structured knowledge graphs as rotations in a complex embedding space and captures relational patterns among concepts\[[16](https://arxiv.org/html/2609.30763#bib.bib17)\]\. However, these methods are trained from scratch on task\-specific data and do not fully exploit transferable knowledge learned from large\-scale corpora\.

Biomedical LMs, such as BioBERT\[[4](https://arxiv.org/html/2609.30763#bib.bib4)\]and ClinicalBERT\[[11](https://arxiv.org/html/2609.30763#bib.bib19)\], provide transferable contextual semantics for medical concepts by pretraining on large\-scale biomedical or clinical corpora and have shown strong performance on downstream clinical prediction tasks\. However, these models primarily capture textual semantics and are not explicitly optimized to preserve hierarchical relations among clinical codes, limiting their ability to model clinical ontology structure\. Recent hierarchy\- and ontology\-aware LM methods further use structured biomedical knowledge to improve clinical representations\. For example, SapBERT uses UMLS\-based synonym alignment to learn better clinical concept embeddings\[[12](https://arxiv.org/html/2609.30763#bib.bib20)\], while G\-BERT combines EHR sequences with medical ontology graphs for medication recommendation\[[10](https://arxiv.org/html/2609.30763#bib.bib15)\]\. However, these methods learn embeddings in Euclidean space, which is less suitable for modeling the tree\-structured hierarchies of clinical ontologies\. In contrast, hyperbolic representation learning provides a more suitable geometry for hierarchical structures because the volume of hyperbolic space grows exponentially with radius\. Methods such as Poincaré Embeddings\[[13](https://arxiv.org/html/2609.30763#bib.bib7)\], Hyperbolic Entailment Cones\[[14](https://arxiv.org/html/2609.30763#bib.bib9)\], and Poincaré GloVe\[[15](https://arxiv.org/html/2609.30763#bib.bib12)\]have shown the effectiveness of hyperbolic spaces for modeling hierarchical or relational structures\. However, these methods are not designed to integrate pre\-trained biomedical LM semantics with expert\-curated clinical ontologies\. To bridge this gap, HCOE learns hierarchy\-aware medical code embeddings by integrating biomedical LM semantics with clinical ontologies in hyperbolic space\.

## IIIBackground of Hyperbolic Representation

Hyperbolic representation learning embeds hierarchical structures in a negatively curved space, whose exponential volume growth naturally fits tree\-like structures\[[13](https://arxiv.org/html/2609.30763#bib.bib7)\]\. It motivates hierarchy\-preserving hyperbolic embedding methods such as Poincaré embeddings\[[13](https://arxiv.org/html/2609.30763#bib.bib7)\], Poincaré GloVe\[[15](https://arxiv.org/html/2609.30763#bib.bib12)\], and entailment cones\[[14](https://arxiv.org/html/2609.30763#bib.bib9)\]\. Here, we use thedd\-dimensional Poincaré ball with a curvature−κ\-\\kappa\(κ\>0\\kappa\>0\):

𝔹κd=\{𝐱∈ℝd:κ​‖𝐱‖2<1\}\\mathbb\{B\}^\{d\}\_\{\\kappa\}=\\left\\\{\\mathbf\{x\}\\in\\mathbb\{R\}^\{d\}:\\kappa\\,\\\|\\mathbf\{x\}\\\|^\{2\}<1\\right\\\}\(1\)The hyperbolic distance between two points𝐮,𝐯∈𝔹κd\\mathbf\{u\},\\mathbf\{v\}\\in\\mathbb\{B\}^\{d\}\_\{\\kappa\}is:

dκ\(𝐮,𝐯\)=2κtanh−1\(κ∥\(−𝐮\)⊕κ𝐯∥\)d\_\{\\kappa\}\(\\mathbf\{u\},\\mathbf\{v\}\)=\\frac\{2\}\{\\sqrt\{\\kappa\}\}\\tanh^\{\-1\}\\\!\\left\(\\sqrt\{\\kappa\}\\,\\\|\(\-\\mathbf\{u\}\)\\oplus\_\{\\kappa\}\\mathbf\{v\}\\\|\\right\)\(2\)where∥⋅∥\\\|\\cdot\\\|is Euclidean norm and⊕κ\\oplus\_\{\\kappa\}is Möbius addition:

𝐮⊕κ𝐯=\(1\+2​κ​⟨𝐮,𝐯⟩\+κ​‖𝐯‖2\)​𝐮\+\(1−κ​‖𝐮‖2\)​𝐯1\+2​κ​⟨𝐮,𝐯⟩\+κ2​‖𝐮‖2​‖𝐯‖2\\mathbf\{u\}\\oplus\_\{\\kappa\}\\mathbf\{v\}=\\frac\{\(1\+2\\kappa\\langle\\mathbf\{u\},\\mathbf\{v\}\\rangle\+\\kappa\\\|\\mathbf\{v\}\\\|^\{2\}\)\\mathbf\{u\}\+\(1\-\\kappa\\\|\\mathbf\{u\}\\\|^\{2\}\)\\mathbf\{v\}\}\{1\+2\\kappa\\langle\\mathbf\{u\},\\mathbf\{v\}\\rangle\+\\kappa^\{2\}\\\|\\mathbf\{u\}\\\|^\{2\}\\\|\\mathbf\{v\}\\\|^\{2\}\}\(3\)where⟨⋅,⋅⟩\\langle\\cdot,\\cdot\\rangledenotes the Euclidean inner product\.

## IVMethodology

### IV\-AICD and ATC Clinical Ontologies

Clinical coding systems, such as ICD and ATC, organize diagnoses, procedures, and medications into multi\-level taxonomies from broad categories to fine\-grained codes\. For this study, we map each clinical concept to a unique five\-level ontology path, which defines the parent\-child relations used to train HCOE\. For diagnoses, we use the Clinical Classifications Software \(CCS\) hierarchy developed by the Agency for Healthcare Research and Quality \(AHRQ\) to organize ICD\-CM codes into five levels: major category, sub\-category, CCS code, integer\-level ICD\-CM code, and decimal\-level ICD\-CM code\. For procedures, we similarly organize ICD\-PCS codes using the CCS procedure hierarchy into major category, sub\-category, CCS code, three\-character ICD\-PCS prefix, and full ICD\-PCS code\. For medications, we use the five\-level ATC hierarchy: Anatomical Main Group→\\rightarrowTherapeutic Subgroup→\\rightarrowPharmacological Subgroup→\\rightarrowChemical Subgroup→\\rightarrowChemical Substance\.

Fig\. 2:Overview of HCOE\.a\. An expert\-curated medical ontology organizes clinical concepts into coarse\-to\-fine hierarchy levels\.b\. HCOE uses hyperbolic contrastive learning to organize biomedical LM\-encoded clinical codes, pulling parent–child concepts closer and pushing unrelated concepts apart\.c\. For a target clinical concept, HCOE aggregates embeddings along its ontology path to form a coarse\-to\-fine hierarchy\-aware representation\.
### IV\-BClinical Ontology Paths and LM Initialization

Clinical concepts are organized in an expert\-curated medical ontology withLLlevels of granularity\. LetC\(ℓ\)=\{c1\(ℓ\),…,c\|C\(ℓ\)\|\(ℓ\)\}C^\{\(\\ell\)\}=\\\{c^\{\(\\ell\)\}\_\{1\},\\dots,c^\{\(\\ell\)\}\_\{\|C^\{\(\\ell\)\}\|\}\\\}denote the set of clinical concepts at levelℓ∈\{1,…,L\}\\ell\\in\\\{1,\\dots,L\\\}\. For each clinical conceptii, we define its ontology path as the ordered sequence of ancestor concepts from coarse to fine levels:

𝒫⁡\(i\)≡\(ai\(1\),…,ai\(L\)\)\\mathcal\{P\}\(i\)\\equiv\\big\(a\_\{i\}^\{\(1\)\},\\dots,a\_\{i\}^\{\(L\)\}\\big\)\(4\)whereai\(ℓ\)∈C\(ℓ\)a\_\{i\}^\{\(\\ell\)\}\\in C^\{\(\\ell\)\}denotes the ancestor of conceptiiat levelℓ\\ell\. As shown in Fig\.[2](https://arxiv.org/html/2609.30763#S4.F2)c, the ICD\-10\-CM codeE11\.2\(type 2 diabetes mellitus with kidney complications\) follows the path\[3, 3\.3, 50, E11, E11\.2\]corresponding to endocrine disease \(Major category\), diabetes mellitus with complication \(Sub\-category\), diabetes mellitus with complication \(CCS code\), type 2 diabetes mellitus \(integer\-level ICD\), and itself \(decimal\-level ICD\)\.

We construct ahierarchy\-aware semantic embeddingfor conceptiiby aggregating level\-specific embeddings of its ancestor codes based on its ontology path𝒫⁡\(i\)\\mathcal\{P\}\(i\)\. For each clinical codecj\(ℓ\)c^\{\(\\ell\)\}\_\{j\}at levelℓ\\ell, we use a frozen biomedical LMEnc​\(⋅\)\\text\{Enc\}\(\\cdot\)to encode its textual description and obtain its embedding from the pooled LM output:

𝐱j\(ℓ\)=Pooler⁡\(Enc⁡\(cj\(ℓ\)\)\)\.\\mathbf\{x\}^\{\(\\ell\)\}\_\{j\}=\\mathrm\{Pooler\}\\big\(\\mathrm\{Enc\}\(c^\{\(\\ell\)\}\_\{j\}\)\\big\)\.\(5\)Let𝐗\(ℓ\)∈ℝ\|C\(ℓ\)\|×d0\\mathbf\{X\}^\{\(\\ell\)\}\\in\\mathbb\{R\}^\{\|C^\{\(\\ell\)\}\|\\times d\_\{0\}\}denote the LM\-initialized embedding table at levelℓ\\ell, whered0d\_\{0\}is the output dimension of the biomedical LM\. These embeddings provide semantic initialization for HCOE, while the ontology paths provide parent\-child relations for hyperbolic representation learning\.

### IV\-CHyperbolic Clinical Ontology Embeddings

HCOE learns hyperbolic clinical ontology embeddings from biomedical LM representations using ontology\-guided hyperbolic contrastive learning and coarse\-to\-fine path aggregation\. Hyperbolic space is well suited to hierarchical clinical ontologies because it provides greater capacity for representing increasingly fine\-grained concepts\[[13](https://arxiv.org/html/2609.30763#bib.bib7)\]\. The CCS/ICD and ATC hierarchies provide the parent\-child relations and ontology paths used to train HCOE\. We map frozen biomedical LM embeddings into a shared Poincaré ball using level\-specific linear projections followed by the exponential map, with the projection parameters optimized through ontology\-guided hyperbolic contrastive learning\. For each ontology nodecj\(ℓ\)c\_\{j\}^\{\(\\ell\)\}, we compute:

𝐳j\(ℓ\)=Wℓ​𝐱j\(ℓ\)\+𝐛ℓ,𝐡j\(ℓ\)=exp𝟎κ⁡\(𝐳j\(ℓ\)\),\\mathbf\{z\}^\{\(\\ell\)\}\_\{j\}=W\_\{\\ell\}\\mathbf\{x\}^\{\(\\ell\)\}\_\{j\}\+\\mathbf\{b\}\_\{\\ell\},\\qquad\\mathbf\{h\}^\{\(\\ell\)\}\_\{j\}=\\exp^\{\\kappa\}\_\{\\mathbf\{0\}\}\\big\(\\mathbf\{z\}^\{\(\\ell\)\}\_\{j\}\\big\),\(6\)whereWℓW\_\{\\ell\}and𝐛ℓ\\mathbf\{b\}\_\{\\ell\}are level\-specific projection parameters,exp𝟎κ⁡\(⋅\)\\exp^\{\\kappa\}\_\{\\mathbf\{0\}\}\(\\cdot\)is the exponential map at the origin, and𝐡j\(ℓ\)∈𝔹κdh\\mathbf\{h\}^\{\(\\ell\)\}\_\{j\}\\in\\mathbb\{B\}^\{d\_\{h\}\}\_\{\\kappa\}is the hyperbolic embedding of ontology nodecj\(ℓ\)c^\{\(\\ell\)\}\_\{j\}\. As a result, all nodes are mapped to the samedhd\_\{h\}\-dimensional Poincaré ball, allowing parent\-child distances to be computed directly\.

To encode clinical hierarchy, we optimize HCOE using parent\-child relations in the ontology\. As shown in Fig\.[2](https://arxiv.org/html/2609.30763#S4.F2)b, we use a hyperbolic contrastive learning strategy that uses both parent\-side and child\-side triplet losses\. The triplet losses encourage semantically related concepts to be close, while pushing unrelated ones farther apart\. For a given anchor conceptuu, we construct a triplet\(u,p,n\)\(u,p,n\)by selecting a positive sampleppand a negative samplenn\. The resulting hyperbolic triplet loss is:

ℒ=∑\(u,p,n\)max⁡\(0,dκ​\(𝐡u,𝐡p\)−dκ​\(𝐡u,𝐡n\)\+α\)\\mathcal\{L\}=\\sum\_\{\(u,p,n\)\}\\max\\left\(0,\\;d\_\{\\kappa\}\(\\mathbf\{h\}\_\{u\},\\mathbf\{h\}\_\{p\}\)\-d\_\{\\kappa\}\(\\mathbf\{h\}\_\{u\},\\mathbf\{h\}\_\{n\}\)\+\\alpha\\right\)\(7\)wheredκ​\(⋅,⋅\)d\_\{\\kappa\}\(\\cdot,\\cdot\)denotes the hyperbolic distance in the Poincaré ball andα\\alphais a margin hyperparameter\.

TABLE I:Clinical relation prediction on ICD→\\rightarrowICD, ATC→\\rightarrowATC, ICD→\\rightarrowATC, and CCS→\\rightarrowPheCode\. Thresholds are selected on validation and fixed for testing without variance\. The best results in each column are highlighted in bold\.We compute triplet losses from both the parent\-side and the child\-side\. The parent\-side objective uses a child concept as the anchor and its parent as the positive concept, encouraging each child to remain close to its ancestor\. The child\-side objective uses a parent concept as the anchor and one of its children as the positive concept, encouraging parent concepts to preserve their local descendant structure\. Negatives are sampled from sibling concepts or unrelated concepts in the ontology\. The final HCOE training objective is

ℒHCOE=ℒparent\+ℒchild\\mathcal\{L\}\_\{\\text\{HCOE\}\}=\\mathcal\{L\}\_\{\\text\{parent\}\}\+\\mathcal\{L\}\_\{\\text\{child\}\}\(8\)
After training, each ontology node has a hyperbolic embedding𝐡j\(ℓ\)\\mathbf\{h\}^\{\(\\ell\)\}\_\{j\}\. For a fine\-grained clinical codeii, we obtain its final representation by aggregating the hyperbolic embeddings of nodes along its ontology path using Möbius addition:

𝐡iHCOE=⨁ℓ=1L𝐡ai\(ℓ\)\(ℓ\)\\mathbf\{h\}\_\{i\}^\{\\text\{HCOE\}\}=\\bigoplus\_\{\\ell=1\}^\{L\}\\mathbf\{h\}\_\{a\_\{i\}^\{\(\\ell\)\}\}^\{\(\\ell\)\}\(9\)where⨁\\bigoplusdenotes sequential Möbius addition from the coarsest to the finest ontology level \(ℓ=1\\ell=1toLL\), keeping the aggregated representation inside the Poincaré ball\. The resulting representation𝐡iHCOE\\mathbf\{h\}\_\{i\}^\{\\text\{HCOE\}\}integrates biomedical LM semantics with the coarse\-to\-fine structure of the clinical ontology\.

## VExperiments

### V\-AMIMIC\-IV Dataset and Preprocessing

We evaluated HCOE on MIMIC‑IV, a publicly available, de\-identified EHR database from the Beth Israel Deaconess Medical Center\[[17](https://arxiv.org/html/2609.30763#bib.bib13)\]\. The database contains comprehensive longitudinal hospital and intensive care data collected from emergency department and inpatient admissions between 2008 and 2022\. We use the inpatient hospitalization data, including diagnoses, procedures, medication prescriptions, and admission and discharge timestamps\. We extracted clinical codes from all patients and represented them using their corresponding five\-level codes\. For each patient, all clinical codes within a visit are treated as an unordered set, while visits are chronologically ordered\. After preprocessing, the cohort contains 223,452 patients and 546,028 hospital admissions, with 9,143 ICD\-9 diagnosis codes, 19,440 ICD\-10 diagnosis codes, 2,557 ICD\-9 procedure codes, 12,354 ICD\-10 procedure codes, and 1,286 ATC level\-5 medication codes\.

### V\-BICD/ATC Relation Prediction and PheCode Transfer

To evaluate whether HCOE captures clinical relationships in biomedical taxonomies, we assess its performance across four clinical relation prediction settings: ICD→\\rightarrowICD, ATC→\\rightarrowATC, ICD→\\rightarrowATC, and CCS→\\rightarrowPheCode\. We first perform multi\-hop relation prediction within the ICD and ATC hierarchies to predict ancestor–descendant relations beyond directly observed parent\-child links\. For each positive relation pair, we sample ten negative candidates, using sibling concepts as hard negatives when available and otherwise sampling unrelated concepts uniformly at random\. We also evaluate a cross\-ontology one\-hop relation prediction task between medications and diagnoses using ICD\-10→\\rightarrowATC links\[[18](https://arxiv.org/html/2609.30763#bib.bib10)\]\. Given an ATC medication code, the task is to predict whether a candidate ICD\-10 diagnosis code is associated with it\. Since each ATC code is paired with one ICD\-10 label in our dataset, we sample one negative ICD\-10 code for each positive pair\. Finally, we evaluate HCOE on an unseen PheCode\-based ICD hierarchy after training on the CCS\-based ICD ontology\[[3](https://arxiv.org/html/2609.30763#bib.bib3)\]\. For each query ICD code, candidates sharing its integer\-level PheCode are positives, whereas candidates from other PheCodes are hard negatives\. We compare candidates by hyperbolic distance and test whether positives rank above negatives\. For each method, the decision threshold is selected on the validation set and fixed for the test set\. All reported Precision, Recall, and F1\-score values are computed on the held\-out test set\. We compare HCOE with four groups of baselines: \(1\) a hierarchy\-agnostic LM baseline BioBERT\[[4](https://arxiv.org/html/2609.30763#bib.bib4)\]; \(2\) structure\-aware LM baselines, including SapBERT\[[12](https://arxiv.org/html/2609.30763#bib.bib20)\], HiTs\[[19](https://arxiv.org/html/2609.30763#bib.bib11)\], and OnT\[[20](https://arxiv.org/html/2609.30763#bib.bib21)\]; \(3\) structure\-aware embedding baselines, including RotatE\[[16](https://arxiv.org/html/2609.30763#bib.bib17)\]and cui2vec\[[8](https://arxiv.org/html/2609.30763#bib.bib16)\]; and \(4\) hyperbolic embedding baselines, including Poincaré Embedding\[[13](https://arxiv.org/html/2609.30763#bib.bib7)\], Hyperbolic Entailment Cone\[[14](https://arxiv.org/html/2609.30763#bib.bib9)\], and Poincaré GloVe embedding\[[15](https://arxiv.org/html/2609.30763#bib.bib12)\]\.

### V\-CMIMIC\-IV Downstream Prediction Setup

For clinical prediction, we extracted diagnosis, procedure, and medication codes from MIMIC\-IV and evaluated three common tasks, including mortality and readmission prediction as well as drug recommendation\. Mortality prediction assesses whether a patient passes away within 90 days after discharge\. Readmission prediction assesses whether a patient is readmitted within the next 15 days after discharge\. Drug recommendation predicts the set of medications for each visit given the patient’s prior clinical history\[[2](https://arxiv.org/html/2609.30763#bib.bib2)\]\. Mortality and readmission are binary classification tasks evaluated with Area Under Precision\-Recall Curve \(AUPRC\) and Area Under Receiver Operating Characteristic Curve \(AUROC\)\. Drug recommendation is a multi\-label prediction task evaluated with AUPRC, F1\-score, and Jaccard\. We also predicted rare drugs on 98 ATC codes with frequencies between 10 and 20\. For each of the 98 rare drugs, we compute Recall@15 over visits containing that drug\. For all downstream tasks, we split patients into training, validation, and test sets using a 70%/10%/20% split\.

To facilitate a controlled comparison of embedding quality, all methods are evaluated using the same mean\-pooling strategy and linear prediction head\. We map clinical codes to their HCOE embeddings and use mean pooling over code embeddings to obtain a patient\-level representation\. For all downstream tasks, we use a single linear prediction head on top of the patient\-level representation\. Mortality and readmission are modeled as binary classification tasks with a sigmoid activation and binary cross\-entropy loss\. Drug recommendation is modeled as a multi\-label classification task, where a linear layer outputs logits for all medication codes and an element\-wise sigmoid activation is applied\. The model is trained using binary cross\-entropy loss over all medication labels\. The best model is selected according to the validation loss on the target task\.

We compare HCOE with three groups of baselines: \(1\) structure\-aware embedding baselines, including GRAM\[[9](https://arxiv.org/html/2609.30763#bib.bib14)\], cui2vec\[[8](https://arxiv.org/html/2609.30763#bib.bib16)\], and RotatE\[[16](https://arxiv.org/html/2609.30763#bib.bib17)\]; \(2\) biomedical LM baselines, including BioBERT\[[4](https://arxiv.org/html/2609.30763#bib.bib4)\], G\-BERT\[[10](https://arxiv.org/html/2609.30763#bib.bib15)\], ClinicalBERT\[[11](https://arxiv.org/html/2609.30763#bib.bib19)\], and BEHRT\[[21](https://arxiv.org/html/2609.30763#bib.bib18)\]; \(3\) structure\-aware LM baselines, including SapBERT\[[12](https://arxiv.org/html/2609.30763#bib.bib20)\], HiTs\[[19](https://arxiv.org/html/2609.30763#bib.bib11)\], and OnT\[[20](https://arxiv.org/html/2609.30763#bib.bib21)\]\.

TABLE II:Clinical prediction results on MIMIC\-IV for mortality, readmission, drug recommendation, and rare drug prediction tasks\. Metrics are reported as mean \(standard error\) from bootstrap\. The best results in each column are highlighted in bold\.
### V\-DImplementation Details

We use BioBERT as the frozen biomedical LM backbone and obtain concept embeddings from its pooled output\[[4](https://arxiv.org/html/2609.30763#bib.bib4)\]\. BioBERT producesd0=768d\_\{0\}=768\-dimensional embeddings, which are projected todh=256d\_\{h\}=256\-dimensional hyperbolic embeddings\. We set the curvature parameter of the Poincaré ball toκ=1/dh\\kappa=1/d\_\{h\}and the margin parameter in the hyperbolic triplet loss toα=5\\alpha=5\. All models were implemented in PyTorch and optimized with Adam optimizer with a learning rate of1×10−41\\times 10^\{\-4\}and weight decay of1×10−51\\times 10^\{\-5\}\. The training process uses a batch size of 64 for up to 20 epochs with early stopping\.

Fig\. 3:Visualization of HCOE embeddings for ICD diagnosis codes under the endocrine disease category \(major category 3\) using a Poincaré map\. The embeddings of the endocrine disease category exhibit a radial hierarchy, with coarse concepts near the center and fine\-grained codes toward the boundary\.

## VIResults

### VI\-AHCOE Preserves Clinical Ontology Relations

As shown in Table[I](https://arxiv.org/html/2609.30763#S4.T1), HCOE achieves the best F\-score across all four clinical relation prediction tasks\. For within\-ontology prediction, HCOE obtains F\-scores of 75\.6% on ICD→\\rightarrowICD and 76\.3% on ATC→\\rightarrowATC, corresponding to absolute gains of 8\.5% and 1\.8% over the strongest baselines, respectively\. These results indicate that combining biomedical LM semantics with hyperbolic ontology contrastive learning better preserves ancestor\-descendant relations in expert\-defined clinical taxonomies\. HCOE also outperforms BioBERT, indicating that hyperbolic contrastive learning better distinguishes true ontology relations from unrelated candidate pairs than textual similarity alone\. HCOE further outperforms structure\-aware embedding methods, showing that combining biomedical LM semantics with clinical ontology yields more discriminative concept representations\. It also surpasses hyperbolic embedding baselines, indicating that hyperbolic geometry alone is insufficient without LM\-based semantics\. Cross\-ontology prediction is more challenging because diagnosis and medication concepts come from separate taxonomies, and within\-ontology training does not directly provide alignment between ICD and ATC concepts\. Nevertheless, HCOE still ranks first with F\-score 56\.7%, outperforming the strongest baseline HiTs with F\-score 53\.2%\. These results suggest that HCOE can support relation prediction across heterogeneous clinical coding systems, where diagnosis and medication concepts come from different taxonomies\. On the CCS→\\rightarrowPheCode transfer task, HCOE also achieves the highest F\-score 80\.6%, outperforming the strongest baseline HiTs \(76\.5%\) by 4\.1%\. Because PheCode labels are not used during training, this result suggests that HCOE learns clinical relationships among ICD diagnostic codes that generalize beyond the CCS\-based training hierarchy to an unseen PheCode\-based hierarchy\.

TABLE III:Ablation of HCOE on four clinical relation prediction tasks, including ICD→\\rightarrowICD, ATC→\\rightarrowATC, ICD→\\rightarrowATC, and CCS→\\rightarrowPheCode\. We report F\-score for each task\. The best results in each column are highlighted in bold\.
### VI\-BQualitative Visualization of HCOE Representations

To examine whether HCOE learns hierarchy\-consistent geometry, we visualize the embeddings of all descendant concepts under the endocrine disease category \(major category 3\)\. We use a Poincaré map, a two\-dimensional visualization of hyperbolic embeddings, to illustrate all ontology nodes under the endocrine disease category\[[22](https://arxiv.org/html/2609.30763#bib.bib8)\]\. As shown in Fig\.[3](https://arxiv.org/html/2609.30763#S5.F3), the visualization shows a clear radial structure: coarse\-grained concepts, such as 3 endocrine disease and 3\.2 diabetes mellitus without complication, lie near the center of the Poincaré ball, whereas fine\-grained ICD codes are positioned closer to the boundary\. This pattern is consistent with hyperbolic geometry, where the available volume grows exponentially with radius and provides greater capacity near the boundary for fine\-grained concepts\. This result suggests that HCOE captures the coarse\-to\-fine structure of the clinical ontology while producing interpretable hyperbolic concept embeddings\.

### VI\-CHCOE Improves Clinical Prediction

As shown in Table[II](https://arxiv.org/html/2609.30763#S5.T2), HCOE consistently improves clinical prediction performance across all downstream clinical prediction metrics on MIMIC\-IV\. For mortality and readmission, HCOE outperforms the strongest baseline in mortality \(AUPRC 44\.3% vs\. 43\.5% and AUROC 87\.2% vs\. 86\.2%\) and in readmission \(AUPRC 48\.7% vs\. 47\.3% and AUROC 72\.3% vs\. 71\.5%\)\. These results indicate that hierarchy\-aware clinical concept embeddings produce more discriminative patient\-level features than pre\-trained biomedical LM representations alone\. The gains are most evident on medication recommendation tasks\. For drug recommendation, HCOE achieves the best scores across all three metrics, indicating that incorporating clinical ontology structure helps capture relationships among medication codes, leading to more effective multi\-label drug prediction\. For rare drug prediction, HCOE achieves the highest Recall@15 \(44\.2%\), outperforming the best baseline \(42\.7%\)\. Because rare drugs have limited training examples, this improvement suggests that HCOE can leverage shared ontology structure among related medications to improve retrieval of long\-tail drug codes\.

### VI\-DAblation Studies

We conduct ablation studies to isolate the effects of path aggregation and hyperbolic contrastive learning in HCOE\. As shown in Table[III](https://arxiv.org/html/2609.30763#S6.T3), we first evaluate the role of path aggregation by removing it while keeping the full hyperbolic contrastive learning objective\. We then evaluate the role of the contrastive objective by keeping path aggregation and using only the parent\-side objective, only the child\-side objective, or no contrastive learning\. We report F\-score on four clinical relation prediction tasks: ICD→\\rightarrowICD, ATC→\\rightarrowATC, ICD→\\rightarrowATC, and CCS→\\rightarrowPheCode\.

As shown in Table[III](https://arxiv.org/html/2609.30763#S6.T3), the full HCOE model achieves the best overall performance on all four tasks, demonstrating the effectiveness of jointly using path aggregation and hyperbolic contrastive learning for clinical concept representation\. Removing the hyperbolic contrastive learning objective causes the largest performance drop, reducing the F\-score from 75\.6 to 47\.6 on ICD→\\rightarrowICD and from 76\.3 to 51\.6 on ATC→\\rightarrowATC\. This indicates that removing the hyperbolic contrastive objective substantially weakens the hierarchy\-aware representations, whereas using both parent\-side and child\-side contrastive objectives produces more hierarchy\-consistent hyperbolic embeddings\. Comparing the two one\-sided contrastive objectives, the parent\-side objective consistently outperforms the child\-side objective across all tasks, suggesting that parent concepts serve as more stable anchors for organizing fine\-grained clinical codes\. Moreover, the full contrastive objective consistently outperforms both one\-sided variants, indicating that parent\-side and child\-side supervision provide complementary hierarchical information\. Overall, these results demonstrate that HCOE improves clinical relation prediction by integrating biomedical LM semantics, path\-level ontology aggregation, and the proposed hyperbolic contrastive learning\.

### VI\-ESensitivity Analysis of Dimension

TABLE IV:Sensitivity of HCOE to the hyperbolic latent dimension\. We report F\-score on clinical relation prediction tasks\.We assess the sensitivity of HCOE to the hyperbolic embedding dimension using three settings,dh∈\{128,256,768\}d\_\{h\}\\in\\\{128,256,768\\\}, on three representative clinical relation prediction tasks: ICD→\\rightarrowICD, ICD→\\rightarrowATC, and CCS→\\rightarrowPheCode\. As shown in Table[IV](https://arxiv.org/html/2609.30763#S6.T4), HCOE achieves the best performance withdh=256d\_\{h\}=256on all three tasks\. The original LM embedding size,dh=768d\_\{h\}=768, does not provide additional gains and may introduce unnecessary complexity, whereasdh=128d\_\{h\}=128appears to limit representation capacity\. The ICD→\\rightarrowATC results vary only slightly across dimensions, suggesting that cross\-ontology prediction is less sensitive to embedding dimension\. We thus usedh=256d\_\{h\}=256for all experiments, as it achieves the best overall performance\.

## VIIConclusion

HCOE combines frozen BioBERT representations, ontology\-guided hyperbolic contrastive learning, and coarse\-to\-fine path aggregation for hierarchy\-aware medical code embeddings\. Evaluations cover ICD/ATC clinical relation prediction, CCS\-to\-PheCode hierarchy transfer, and MIMIC\-IV mortality prediction, readmission prediction, medication recommendation, and rare drug prediction\. The reported results support its use for clinical ontology representation and downstream EHR modeling, with ablations assessing contrastive learning and path aggregation\. This study focuses on common clinical ontologies and standard EHR prediction tasks, leaving broader biomedical knowledge sources for future exploration\. Future work will explore richer biomedical knowledge graphs, including UMLS, and multimodal clinical inputs such as clinical notes\.

## References

- \[1\]Z\. Song, Q\. Lu, H\. Zhu, D\. Buckeridge, and Y\. Li\(2026\)TrajGPT: irregular time\-series representation learning of health trajectory\.IEEE Journal of Biomedical and Health Informatics30\(5\),pp\. 3888–3899\.External Links:[Document](https://dx.doi.org/10.1109/JBHI.2025.3620205)Cited by:[§I](https://arxiv.org/html/2609.30763#S1.p1.1),[§I](https://arxiv.org/html/2609.30763#S1.p3.1)\.
- \[2\]X\. Wang, H\. Wu, Z\. Tai, S\. Lyu, Q\. Lu, Z\. Zhao, J\. Chi, J\. Tian, X\. Chang, and Z\. Song\(2026\)SafeRx\-agent: a knowledge\-grounded multi\-agent framework for safe and explainable medication recommendation\.External Links:2605\.29146,[Link](https://arxiv.org/abs/2605.29146)Cited by:[§I](https://arxiv.org/html/2609.30763#S1.p1.1),[§V\-C](https://arxiv.org/html/2609.30763#S5.SS3.p1.1)\.
- \[3\]Z\. Yang, Z\. Song, S\. Zabad, M\. Legault, and Y\. Li\(2026\)PheCode\-guided multi\-modal topic modeling of electronic health records improves disease incidence prediction and GWAS discovery from UK biobank\.Brief\. Bioinform\.27\(1\) \(en\)\.Cited by:[§I](https://arxiv.org/html/2609.30763#S1.p1.1),[§I](https://arxiv.org/html/2609.30763#S1.p3.1),[§V\-B](https://arxiv.org/html/2609.30763#S5.SS2.p1.1)\.
- \[4\]J\. Lee, W\. Yoon, S\. Kim, D\. Kim, S\. Kim, C\. H\. So, and J\. Kang\(2020\)BioBERT: a pre\-trained biomedical language representation model for biomedical text mining\.Bioinformatics36\(4\),pp\. 1234–1240\.Cited by:[§I](https://arxiv.org/html/2609.30763#S1.p1.1),[§I](https://arxiv.org/html/2609.30763#S1.p2.1),[§I](https://arxiv.org/html/2609.30763#S1.p3.1),[§II](https://arxiv.org/html/2609.30763#S2.p2.1),[§V\-B](https://arxiv.org/html/2609.30763#S5.SS2.p1.1),[§V\-C](https://arxiv.org/html/2609.30763#S5.SS3.p3.1),[§V\-D](https://arxiv.org/html/2609.30763#S5.SS4.p1.1)\.
- \[5\]S\. J\. Steindel\(2010\)International classification of diseases, 10th edition, clinical modification and procedure coding system: descriptive overview of the next generation HIPAA code sets\.J\. Am\. Med\. Inform\. Assoc\.17\(3\),pp\. 274–282\(en\)\.Cited by:[§I](https://arxiv.org/html/2609.30763#S1.p1.1)\.
- \[6\]L\. Chen, W\. Zeng, Y\. Cai, K\. Feng, and K\. Chou\(2012\)Predicting anatomical therapeutic chemical \(atc\) classification of drugs by integrating chemical\-chemical interactions and similarities\.PLOS ONE7\(4\),pp\. 1–7\.External Links:[Document](https://dx.doi.org/10.1371/journal.pone.0035254),[Link](https://doi.org/10.1371/journal.pone.0035254)Cited by:[§I](https://arxiv.org/html/2609.30763#S1.p1.1)\.
- \[7\]L\. Shen, Q\. Lu, H\. Zhu, and Z\. Song\(2025\)SMI: semantic medical ID using medical ontology\.InSocially Responsible and Trustworthy Foundation Models at NeurIPS 2025,External Links:[Link](https://openreview.net/forum?id=Z8RT75Prj1)Cited by:[§I](https://arxiv.org/html/2609.30763#S1.p1.1)\.
- \[8\]A\. L\. Beam, B\. Kompa, A\. Schmaltz, I\. Fried, G\. Weber, N\. Palmer, X\. Shi, T\. Cai, and I\. S\. Kohane\(2020\)Clinical concept embeddings learned from massive sources of multimodal medical data\.InPacific Symposium on Biocomputing\. Pacific Symposium on Biocomputing,Vol\.25,pp\. 295\.Cited by:[§I](https://arxiv.org/html/2609.30763#S1.p3.1),[§II](https://arxiv.org/html/2609.30763#S2.p1.1),[§V\-B](https://arxiv.org/html/2609.30763#S5.SS2.p1.1),[§V\-C](https://arxiv.org/html/2609.30763#S5.SS3.p3.1)\.
- \[9\]E\. Choi, M\. T\. Bahadori, L\. Song, W\. F\. Stewart, and J\. Sun\(2017\)GRAM: graph\-based attention model for healthcare representation learning\.InProceedings of the 23rd ACM SIGKDD international conference on knowledge discovery and data mining,pp\. 787–795\.Cited by:[§I](https://arxiv.org/html/2609.30763#S1.p3.1),[§II](https://arxiv.org/html/2609.30763#S2.p1.1),[§V\-C](https://arxiv.org/html/2609.30763#S5.SS3.p3.1)\.
- \[10\]J\. Shang, T\. Ma, C\. Xiao, and J\. Sun\(2019\)Pre\-training of graph augmented transformers for medication recommendation\.arXiv preprint arXiv:1906\.00346\.Cited by:[§I](https://arxiv.org/html/2609.30763#S1.p3.1),[§II](https://arxiv.org/html/2609.30763#S2.p2.1),[§V\-C](https://arxiv.org/html/2609.30763#S5.SS3.p3.1)\.
- \[11\]E\. Alsentzer, J\. Murphy, W\. Boag, W\. Weng, D\. Jindi, T\. Naumann, and M\. McDermott\(2019\)Publicly available clinical bert embeddings\.InProceedings of the 2nd clinical natural language processing workshop,pp\. 72–78\.Cited by:[§I](https://arxiv.org/html/2609.30763#S1.p3.1),[§II](https://arxiv.org/html/2609.30763#S2.p2.1),[§V\-C](https://arxiv.org/html/2609.30763#S5.SS3.p3.1)\.
- \[12\]F\. Liu, I\. Vulić, A\. Korhonen, and N\. Collier\(2021\)Learning domain\-specialised representations for cross\-lingual biomedical entity linking\.InProceedings of ACL\-IJCNLP 2021,pp\. 565–574\.Cited by:[§I](https://arxiv.org/html/2609.30763#S1.p3.1),[§II](https://arxiv.org/html/2609.30763#S2.p2.1),[§V\-B](https://arxiv.org/html/2609.30763#S5.SS2.p1.1),[§V\-C](https://arxiv.org/html/2609.30763#S5.SS3.p3.1)\.
- \[13\]M\. Nickel and D\. Kiela\(2017\)Poincaré embeddings for learning hierarchical representations\.InAdvances in Neural Information Processing Systems 30,I\. Guyon, U\. V\. Luxburg, S\. Bengio, H\. Wallach, R\. Fergus, S\. Vishwanathan, and R\. Garnett \(Eds\.\),pp\. 6341–6350\.Cited by:[§I](https://arxiv.org/html/2609.30763#S1.p3.1),[§II](https://arxiv.org/html/2609.30763#S2.p2.1),[§III](https://arxiv.org/html/2609.30763#S3.p1.1),[§IV\-C](https://arxiv.org/html/2609.30763#S4.SS3.p1.1),[§V\-B](https://arxiv.org/html/2609.30763#S5.SS2.p1.1)\.
- \[14\]O\. Ganea, G\. Becigneul, and T\. Hofmann\(2018\)Hyperbolic entailment cones for learning hierarchical embeddings\.InProceedings of the 35th International Conference on Machine Learning,J\. Dy and A\. Krause \(Eds\.\),Proceedings of Machine Learning Research, Vol\.80,pp\. 1646–1655\.External Links:[Link](https://proceedings.mlr.press/v80/ganea18a.html)Cited by:[§I](https://arxiv.org/html/2609.30763#S1.p3.1),[§II](https://arxiv.org/html/2609.30763#S2.p2.1),[§III](https://arxiv.org/html/2609.30763#S3.p1.1),[§V\-B](https://arxiv.org/html/2609.30763#S5.SS2.p1.1)\.
- \[15\]A\. Tifrea, G\. Becigneul, and O\. Ganea\(2019\)Poincare glove: hyperbolic word embeddings\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=Ske5r3AqK7)Cited by:[§I](https://arxiv.org/html/2609.30763#S1.p3.1),[§II](https://arxiv.org/html/2609.30763#S2.p2.1),[§III](https://arxiv.org/html/2609.30763#S3.p1.1),[§V\-B](https://arxiv.org/html/2609.30763#S5.SS2.p1.1)\.
- \[16\]Z\. Sun, Z\. Deng, J\. Nie, and J\. Tang\(2019\)Rotate: knowledge graph embedding by relational rotation in complex space\.arXiv preprint arXiv:1902\.10197\.Cited by:[§II](https://arxiv.org/html/2609.30763#S2.p1.1),[§V\-B](https://arxiv.org/html/2609.30763#S5.SS2.p1.1),[§V\-C](https://arxiv.org/html/2609.30763#S5.SS3.p3.1)\.
- \[17\]A\. E\. Johnson, L\. Bulgarelli, L\. Shen, A\. Gayles, A\. Shammout, S\. Horng, T\. J\. Pollard, S\. Hao, B\. Moody, B\. Gow,et al\.\(2023\)MIMIC\-iv, a freely accessible electronic health record dataset\.Scientific data10\(1\),pp\. 1\.Cited by:[§V\-A](https://arxiv.org/html/2609.30763#S5.SS1.p1.1)\.
- \[18\]Y\. Zou, A\. Pesaranghader, Z\. Song, A\. Verma, D\. L\. Buckeridge, and Y\. Li\(2022\)Modeling electronic health record data using an end\-to\-end knowledge\-graph\-informed topic model\.Sci\. Rep\.12\(1\),pp\. 17868\(en\)\.Cited by:[§V\-B](https://arxiv.org/html/2609.30763#S5.SS2.p1.1)\.
- \[19\]Y\. He, M\. Yuan, J\. Chen, and I\. Horrocks\(2024\)Language models as hierarchy encoders\.Advances in Neural Information Processing Systems37,pp\. 14690–14711\.Cited by:[§V\-B](https://arxiv.org/html/2609.30763#S5.SS2.p1.1),[§V\-C](https://arxiv.org/html/2609.30763#S5.SS3.p3.1)\.
- \[20\]H\. Yang, J\. Chen, Y\. He, Y\. Gao, and I\. Horrocks\(2025\)Language models as ontology encoders\.InThe Semantic Web – ISWC 2025: 24th International Semantic Web Conference, Nara, Japan, November 2–6, 2025, Proceedings, Part I,Berlin, Heidelberg,pp\. 443–461\.External Links:ISBN 978\-3\-032\-09526\-8,[Link](https://doi.org/10.1007/978-3-032-09527-5_24),[Document](https://dx.doi.org/10.1007/978-3-032-09527-5%5F24)Cited by:[§V\-B](https://arxiv.org/html/2609.30763#S5.SS2.p1.1),[§V\-C](https://arxiv.org/html/2609.30763#S5.SS3.p3.1)\.
- \[21\]Y\. Li, S\. Rao, J\. R\. A\. Solares, A\. Hassaine, R\. Ramakrishnan, D\. Canoy, Y\. Zhu, K\. Rahimi, and G\. Salimi\-Khorshidi\(2020\)BEHRT: transformer for electronic health records\.Scientific reports10\(1\),pp\. 7155\.Cited by:[§V\-C](https://arxiv.org/html/2609.30763#S5.SS3.p3.1)\.
- \[22\]A\. Klimovskaia, D\. Lopez\-Paz, L\. Bottou, and M\. Nickel\(2020\)Poincaré maps for analyzing complex hierarchies in single\-cell data\.Nat\. Commun\.11\(1\),pp\. 2966\(en\)\.Cited by:[§VI\-B](https://arxiv.org/html/2609.30763#S6.SS2.p1.1)\.

相似文章

EHR基础模型中ICD代码的分层建模

arXiv cs.AI

本文研究了在EHR基础模型中显式编码ICD-10-CM层级结构的方法,采用层级令牌增强和基于图结构的代码表示。在MIMIC-IV和eICU上的实验表明,与扁平代码表示相比,该方法在域内和跨数据集预测任务中均有改进。