Scalable partial information decomposition for symptom networks via supervised embeddings

arXiv cs.LG Papers

Summary

The paper introduces ePID, a scalable pipeline using supervised embeddings to compute partial information decomposition for symptom networks, enabling the separation of redundant and synergistic information in mental health data.

arXiv:2609.13203v1 Announce Type: new Abstract: Pairwise relationships among mental-health symptoms are routinely summarised asscalar edge weights, which cannot express whether two symptoms carry overlapping information about a third or information that appears only in combination. Partial information decomposition (PID) addresses this gap but is computationally intractable beyond a few sources. We introduce embedding-based PID (ePID), a scalable pipeline that compresses all nonfocal symptoms into a low-cardinality discrete embedding and computes a tractable two-source PID, yielding source-unique, remainder-unique, redundant, and synergistic components for each ordered source-target pair. We benchmarked 13 candidate embeddings on synthetic Bayesian networks calibrated to PHQ-9 and on 83 real-world datasets across five PID measures. A supervised Agglomerative Conditional Information Bottleneck (ACIB) embedding recovered the reference decomposition most accurately of the 13 embeddings tested, and did so for every PID measure yielding non-negative atoms once four or more symptoms were compressed (synergy recovery r = 0.92). The two instruments then diverged sharply. In PHQ-9 networks (UK Biobank,N = 154,291; Xinxiang student sample, N = 24,292) the surrounding symptom context carried most pairwise dependence through redundant and remainder-unique channels; synergy contributed 6 to 9%, and no directed edge was synergy-dominated in either cohort. In the 28-item Interpersonal Reactivity Index, 45% of source pairs were. The identical pipeline, applied without parameter changes, therefore returned opposite profiles for the two instruments, each consistent with how that instrument was constructed. By separating overlapping from interaction-dependent information, ePID provides a scalable, model-agnostic complement to standard symptom-network methodology to distinguish redundant and synergistic contributions to observed correlations.
Original Article
View Cached Full Text

Cached at: 09/15/26, 08:36 AM

# Scalable partial information decomposition for symptom networks via supervised embeddings
Source: [https://arxiv.org/html/2609.13203](https://arxiv.org/html/2609.13203)
###### Abstract

Background\.Pairwise relationships among mental\-health symptoms are routinely summarised as scalar edge weights, which cannot express whether two symptoms carry overlapping information about a third or information that appears only in combination\. Partial information decomposition \(PID\) addresses this gap but is computationally intractable beyond a few sources\.

Method\.We introduce embedding\-based PID \(ePID\), a scalable pipeline that compresses all non\-focal symptoms into a low\-cardinality discrete embedding and computes a tractable two\-source PID, yielding source\-unique, remainder\-unique, redundant, and synergistic components for each ordered source\-target pair\. We benchmarked 13 candidate embeddings on synthetic Bayesian networks calibrated to PHQ\-9 and on 83 real\-world datasets across five PID measures\.

Results\.A supervised Agglomerative Conditional Information Bottleneck \(ACIB\) embedding recovered the reference decomposition most accurately of the 13 embeddings tested, and did so for every PID measure yielding non\-negative atoms once four or more symptoms were compressed \(synergy recoveryr=0\.92r=0\.92\)\. The two instruments then diverged sharply\. In PHQ\-9 networks \(UK Biobank,N=154,291N=154\{,\}291; Xinxiang student sample,N=24,292N=24\{,\}292\) the surrounding symptom context carried most pairwise dependence through redundant and remainder\-unique channels; synergy contributed 6 to 9%, and no directed edge was synergy\-dominated in either cohort\. In the 28\-item Interpersonal Reactivity Index, 45% of source pairs were\. The identical pipeline, applied without parameter changes, therefore returned opposite profiles for the two instruments, each consistent with how that instrument was constructed\.

Conclusions\.By separating overlapping from interaction\-dependent information, ePID provides a scalable, model\-agnostic complement to standard symptom\-network methodology to distinguish redundant and synergistic contributions to observed correlations\.

1\. Computational Science Lab, Institute for Informatics, University of Amsterdam

Public Significance Statement

Symptom networks map how mental\-health symptoms relate to one another\. These relationships are complex, yet standard methods report only whether two symptoms are associated—not whether they carry the same information about a third symptom or only matter in combination\. We developed ePID, a pipeline that splits each pairwise relationship into overlapping \(redundant\) information, shared by both symptoms, and joint\-only \(synergistic\) information, available only when the two occur together\. Applied to the PHQ\-9 \(depressive symptoms\) and the Interpersonal Reactivity Index \(empathy\), ePID returned profiles matching how each questionnaire was built\. The PHQ\-9 is designed so its items measure a single severity dimension, and ePID found its information almost entirely overlapping\. The Interpersonal Reactivity Index is designed around four distinct facets of empathy, and ePID found much of its information available only from combinations of items\. The method recovered each design without being told anything about it, which is what makes the profiles credible\. Such profiles could, with further validation, help indicate which symptoms to assess or address together rather than in isolation\.

Keywords:symptom networks, partial information decomposition, higher\-order interactions, information theory, psychopathology

## 1Introduction

Network models represent multivariate systems as nodes and their statistical associations as edges, offering a framework in which the structure of interest emerges from the observed relationships themselves rather than from an assumed latent cause\. This approach has been widely adopted in psychopathology, where symptom networks represent mental\-health symptoms as nodes and conditional dependencies as edges\([Borsboom and Cramer, 2013](https://arxiv.org/html/2609.13203#bib.bib8);[Borsboom, 2017](https://arxiv.org/html/2609.13203#bib.bib7)\)\. Rather than treating symptoms as passive indicators of a single disease entity, the network perspective treats them as mutually reinforcing elements whose interaction patterns constitute the disorder itself\([Guloksuz et al\., 2017](https://arxiv.org/html/2609.13203#bib.bib43);[Malgaroli et al\., 2021](https://arxiv.org/html/2609.13203#bib.bib74)\)\. In the most common formulation, edges quantify pairwise conditional associations between symptoms after adjusting for the remaining symptoms, yielding a sparse network often interpreted as a “direct” dependency structure\.

For example, after a stressful event, an individual may experience insomnia that induces fatigue, which might affect concentration, in turn contributing to depressed mood and feelings of worthlessness\([Cramer et al\., 2016](https://arxiv.org/html/2609.13203#bib.bib19)\)\. This perspective has shifted the conceptualisation of mental disorders from a centralised latent\-disease model to a decentralised network model\. However, commonly employed estimators of symptom networks share a fundamental limitation that no choice of pairwise measure can overcome\.

Each edge summarises the association between two symptoms as a single number\. Even when that association is estimated without assuming a particular functional form—a useful safeguard—a scalar edge still cannot express how several symptoms act*together*to inform a third\. That is the gap we address\.

Consider how symptoms behave in context\. Two symptoms may act together, so that their combined presence is informative even though neither tells us much alone; or they may act as overlapping sources, each carrying largely the same information about a third symptom\. For instance, the sleep and fatigue items of the PHQ\-9 remain correlated after conditioning on overall depression severity, so that observing one adds little once the other is known\([Horton and Perry, 2016](https://arxiv.org/html/2609.13203#bib.bib53);[Fried et al\., 2017](https://arxiv.org/html/2609.13203#bib.bib33)\); by contrast, other symptom pairs may carry information about suicidal ideation jointly that neither carries alone, a possibility that pairwise screening has rarely tested\([Franklin et al\., 2017](https://arxiv.org/html/2609.13203#bib.bib31)\)\. A single number per pair cannot tell these cases apart\. They can, however, be operationalised as distinct kinds of information: overlapping information that two symptoms share about a target \(*redundancy*\) and information that becomes available only when they are considered jointly \(*synergy*\), alongside the information each contributes uniquely\.

This distinction has practical implications under an interventionist reading—one that, we stress, presupposes a causal structure and is vulnerable to hidden confounding, so the readings below are strong causal assumptions rather than established conclusions\. If two source symptoms are highly redundant, changing the state of only one would change little for the target, as the other continues to convey the same information\. Conversely, if their joint information is primarily synergistic, then whether disrupting either source suffices to break the joint channel depends on the underlying causal system: that the joint configuration carries the information does not by itself license the intervention\. These interpretations are hypothesis\-generating and would require longitudinal or experimental validation; we revisit them in light of the empirical findings in the Discussion\.

Information theory supplies the tools to make this operationalisation precise\. A source variable provides*information*about a target when observing the source reduces uncertainty about the target’s state\([Shannon, 1948](https://arxiv.org/html/2609.13203#bib.bib88)\)\. Conditional mutual information \(CMI\) uses this idea to generalise partial correlation \(with which it coincides as a conditional\-independence criterion under joint Gaussianity\) into a model\-free \(with respect to functional form\) measure of conditional dependence\. Partial information decomposition \(PID\) goes further, decomposing the information that two or more*sources*provide about a designated*target*into the non\-overlapping components, or*atoms*\([Williams and Beer, 2010](https://arxiv.org/html/2609.13203#bib.bib114)\), sketched above—redundant, unique, and synergistic—so that a conditional association can be read as overlapping or interaction\-only rather than as a single undifferentiated number\.

The obstacle to applying PID is scale\. The number of atoms grows super\-exponentially with the number of sources, from 4 atoms for 2 sources to over 7 million for 6, following the Dedekind numbersM⁡\(N\)M\(N\)\([Kleitman, 1969](https://arxiv.org/html/2609.13203#bib.bib64)\)\. Estimating so many atoms requires high\-dimensional joint distributions that are severely undersampled at typical clinical sample sizes, which in practice restricts most analyses to the triplet level \(two sources and one target\)\.

Although triplet\-level decompositions are tractable, they miss interactions that emerge only when three or more sources act together\. In symptom networks the relevant comparator for a given symptom is its full context—all remaining symptoms—so triplet decompositions cannot answer the system\-level attribution question of primary interest: whether a symptom contributes information about a target*beyond what is already present in the rest of the symptom profile*\.

To address this, we introduce an embedding\-based PID pipeline, which we refer to as*ePID*\. For a given pair of symptoms, ePID first compresses the rest of the network into a compact summary variable—a small number of categories learned to preserve as much information as possible about the*target*symptom specifically—and then asks how the focal symptom \(i\.e\., the source in the pair\) relates to the target compared with that summary\. Because such a decomposition is defined entirely by information about the target, a summary optimised to preserve the remainder’s target\-relevant information retains what it depends on, while what the summary discards is, by construction, largely irrelevant to the target; we make this approximation, and the information it loses, precise in the Methods\. This target\-supervised compression is specific to ePID; partial correlation and CMI, by contrast, are not directed at any particular target\. The procedure yields four components for each ordered symptom pair: information that only the focal symptom provides \(source\-unique\), information that only the rest of the network provides \(remainder\-unique\), overlapping information \(redundancy\), and information that emerges only when the focal symptom and the summary are considered together \(synergy\)\. The resulting directed edges reflect the source–target orientation that the decomposition requires\.

A complementary strand of psychometric network research tackles higher\-order structure with Moderated Network Models \(MNM\), which add multiplicative interaction terms so that the association between two symptoms can vary with a third\([Haslbeck et al\., 2021](https://arxiv.org/html/2609.13203#bib.bib49)\)\. The contrast with ePID is essentially parametric versus non\-parametric: MNM provides a parametric test of whether a chosen moderator significantly modifies a given edge, whereas ePID asks, without specifying interactions in advance, how the information underlying each edge is distributed across redundant, unique, and synergistic channels relative to the rest of the symptom profile\. We develop the fuller methodological contrast in the Discussion\.

This study estimates and compares three approaches to symptom network estimation that progressively relax modelling assumptions and move from pairwise to higher\-order characterisations \(Figure[1](https://arxiv.org/html/2609.13203#S1.F1)\)\. Partial correlations \(PC\), in their Pearson \(PPC\) and Spearman \(SPC\) variants, provide a baseline for conditional linear and monotonic associations; the two variants yield highly similar networks on the ordinal symptom items analysed here \(Appendix[H](https://arxiv.org/html/2609.13203#A8), Figure[S14](https://arxiv.org/html/2609.13203#A8.F14)\), and we therefore present them jointly as PC\. CMI generalises to arbitrary conditional dependence while remaining at the pairwise level\. To characterise how that dependence is composed, we apply PID via the ePID pipeline, yielding for each source–target pair a decomposition into the four atoms above\. The degree to which these methods agree or diverge is itself informative about the statistical structure of the symptom system under study\.

We evaluate ePID in two stages\. First, we benchmark multiple embedding strategies in simulation settings calibrated to PHQ\-9 response distributions, then extend the evaluation across 83 real\-world datasets spanning five PID measures to assess generalisability, quantifying how much information the compression loses \(the approximation error\)\. Second, we apply the approach to item\-level depressive\-symptom networks in two datasets \(UK Biobank\([Sudlow et al\., 2015](https://arxiv.org/html/2609.13203#bib.bib93);[Davis et al\., 2020](https://arxiv.org/html/2609.13203#bib.bib20)\)and the Xinxiang student sample\([Su et al\., 2024a](https://arxiv.org/html/2609.13203#bib.bib91);[Su et al\., 2024b](https://arxiv.org/html/2609.13203#bib.bib92)\)\), using harmonised PHQ\-9 items \(excluding suicidal ideation\) to assess cross\-cohort stability, and to the Interpersonal Reactivity Index\([Davis, 1983](https://arxiv.org/html/2609.13203#bib.bib21)\), a multi\-facet empathy instrument that serves as a contrasting psychometric profile\. In each application we first treat the agreement between partial correlations and CMI as an assumption check on whether conditional structure is largely monotonic at the item level, and then use ePID to add an interaction\-aware interpretation layer, separating redundant from synergistic channels relative to the full symptom context\.

By combining model\-free measures of dependence with a scalable decomposition method, we characterise what kinds of statistical structure underlie observed symptom networks and assess whether, and to what extent, interaction\-aware models are warranted for a given instrument and population\. To our knowledge this is the first application and systematic validation of supervised\-embedding PID to symptom networks, and although we develop it here for psychometric instruments, the pipeline is generic and should apply to other multivariate systems in which network\-context higher\-order structure is of interest—a generalisation we do not yet validate empirically\. We emphasise throughout that these patterns describe statistical associations in cross\-sectional data rather than causal mechanisms, and should be read as constraints and hypotheses for future longitudinal and experimental work; their empirical magnitudes, and their behaviour across cohorts and instruments, are reported in the Results\.

Figure 1:Methods comparison on the toy structural causal model
Note\.The toy system is a five\-variable mixture structural causal model \(X1,X2,W∼Uniform⁡\{0,…,4\}X\_\{1\},X\_\{2\},W\\sim\\mathrm\{Uniform\}\\\{0,\\ldots,4\\\}independent;X4=X5=WX\_\{4\}=X\_\{5\}=W;X3=h⁡\(X1,X2\)X\_\{3\}=h\(X\_\{1\},X\_\{2\}\)with probability0\.50\.5elseX3=WX\_\{3\}=W, wherehhis the non\-monotonic stylised mapping defined in Appendix[J](https://arxiv.org/html/2609.13203#A10)\)\. Top row:\(a\)True graph:X1,X2,X4→X3X\_\{1\},X\_\{2\},X\_\{4\}\\to X\_\{3\}withX4=X5X\_\{4\}=X\_\{5\}marked\.\(b\)PC network: only theX1→X3X\_\{1\}\\to X\_\{3\}edge survives \(ρi​j\|rest=\+0\.38\\rho\_\{ij\\mid\\text\{rest\}\}=\+0\.38\);X4X\_\{4\},X5X\_\{5\}edges are undefined under naive PC because the deterministic equalityX4=X5X\_\{4\}=X\_\{5\}makes the conditioning covariance singular\.\(c\)CMI network:X1→X3X\_\{1\}\\to X\_\{3\}and the non\-monotonicX2→X3X\_\{2\}\\to X\_\{3\}edge \(0\.84 and 0\.19 bits respectively\);X4X\_\{4\},X5X\_\{5\}contributions vanish because conditioning on the remaining system already includes the duplicateX5=X4X\_\{5\}=X\_\{4\}, giving each zero conditional entropy given the rest and henceI⁡\(X4;X3∣rest\)=0I\(X\_\{4\};X\_\{3\}\\mid\\text\{rest\}\)=0; the edges disappear from vanishing conditional mutual information, not from the conditioning rule itself, and the same attenuation arises under near\-redundancy when the sources are highly but not perfectly correlated\. Bottom row:\(d\)the ePID pipeline\. Six triplet\-level two\-source PIDs \(one per source pair targetingX3X\_\{3\}, shown at lower opacity to signal pair\-dependent attribution\) collapse via ACIB\-supervised compression of the remainder\{X2,X4,X5\}\\\{X\_\{2\},X\_\{4\},X\_\{5\}\\\}into a four\-atom ePID for focal sourceX1X\_\{1\}\. All atom values are computed exactly underIminI\_\{\\min\}\(Williams & Beer 2010\); the corresponding BROJA decomposition is shown in Appendix Figure[S1](https://arxiv.org/html/2609.13203#A3.F1)\. PC==partial correlation; CMI==conditional mutual information; PID==partial information decomposition; ePID==embedding\-based PID; ACIB==Agglomerative Conditional Information Bottleneck\.
## 2Methodology

The methodology proceeds as follows\. We first introduce notation and then describe two correlation\-based methods: Pearson and Spearman partial correlations \(PPC and SPC\)\. Next, we explain conditional mutual information \(CMI\), which relaxes the linearity and monotonicity assumptions of PPC and SPC respectively\. These three methods capture only pairwise associations\.

Partial information decomposition \(PID\) goes further by decomposing the information that multiple source variables provide about a target into redundant, synergistic, and unique components, illustrated with a stylised example \(Appendix[J](https://arxiv.org/html/2609.13203#A10)\)\. To make this decomposition tractable beyond a handful of sources, we introduce an embedding\-based PID method \(ePID\) that compresses the remaining variables into a low\-dimensional proxy while retaining information about the target \(Section[2\.5\.2](https://arxiv.org/html/2609.13203#S2.SS5.SSS2)\)\. We then describe the simulation and multi\-dataset benchmarks used to validate this approximation, define the normalisation and significance\-testing framework, and outline the empirical\-analysis procedures applied to the two PHQ\-9 cohorts and the IRI empathy instrument: permutation testing with false discovery rate control for edge significance, entropy\-normalised edge ranking for network visualisation, and cross\-cohort stability assessment between the two PHQ\-9 cohorts\.

### 2\.1Notation

Let𝐗=\{X1,X2,…,Xp\}\\mathbf\{X\}=\\\{X\_\{1\},X\_\{2\},\\ldots,X\_\{p\}\\\}denoteppdiscrete random variables\. EachXiX\_\{i\}can represent the presence or severity of symptomiiin an individual, for example, or other covariates such as education or neighbourhood characteristics\. The cardinality of a discrete variableXiX\_\{i\}is denoted\|Xi\|\|X\_\{i\}\|and equals the number of categories\. For any ordered pair\(Xi,Xj\)\(X\_\{i\},X\_\{j\}\), we writeZi​j=𝐗∖\{Xi,Xj\}Z\_\{ij\}=\\mathbf\{X\}\\setminus\\\{X\_\{i\},X\_\{j\}\\\}for the remaining variables used as a conditioning set\. For PID, we consider the triplets\{Xi,Xj,Xk\}\\\{X\_\{i\},X\_\{j\},X\_\{k\}\\\}to be directed, where we refer to\{Xi,Xj\}\\\{X\_\{i\},X\_\{j\}\\\}\(inputs\) as sources and\{Xk\}\\\{X\_\{k\}\\\}\(output\) as the target\.

For embedding\-based decomposition, we consider an ordered source–target pair\(Xi,Xk\)\(X\_\{i\},X\_\{k\}\)and denote the remaining variables by

Zi→k:=\{Xℓ:ℓ∈\{1,…,p\}∖\{i,k\}\}\.Z\_\{i\\to k\}:=\\\{X\_\{\\ell\}:\\ell\\in\\\{1,\\ldots,p\\\}\\setminus\\\{i,k\\\}\\\}\.In words,Zi→kZ\_\{i\\to k\}collects every symptom other than the focal sourceXiX\_\{i\}and the targetXkX\_\{k\}: the rest of the network, whose joint state we will compress into a low\-dimensional embedding\. We construct a supervised discrete embedding of this remainder,

Ei→k:=fi→k​\(Zi→k\),E\_\{i\\to k\}:=f\_\{i\\to k\}\(Z\_\{i\\to k\}\),wherefi→kf\_\{i\\to k\}is fit to predict the targetXkX\_\{k\}fromZi→kZ\_\{i\\to k\}andEi→kE\_\{i\\to k\}has finite cardinality\|Ei→k\|\|E\_\{i\\to k\}\|\. We then compute a two\-source PID of\(Xi,Ei→k,Xk\)\(X\_\{i\},E\_\{i\\to k\};X\_\{k\}\), interpretingEi→kE\_\{i\\to k\}as a compressed representation of the network context \(the “remainder embedding”\)\.

### 2\.2Data

We analyse Patient Health Questionnaire \(PHQ\-9\) item responses from two datasets: the UK Biobank online mental health questionnaire\([Davis et al\., 2020](https://arxiv.org/html/2609.13203#bib.bib20)\)and a large online sample of Chinese university students \(Xinxiang Medical University; Su et al\. 2024\)\([Su et al\., 2024a](https://arxiv.org/html/2609.13203#bib.bib91);[Su et al\., 2024b](https://arxiv.org/html/2609.13203#bib.bib92)\)\. The PHQ\-9 is a widely used and validated instrument for screening and severity assessment of depressive symptoms in clinical and population settings\([Kroenke et al\., 2001](https://arxiv.org/html/2609.13203#bib.bib68);[Spitzer et al\., 1999](https://arxiv.org/html/2609.13203#bib.bib90)\)\. Each item asks how often, in the last two weeks, respondents were bothered by a symptom, with ordinal response categories coded as 0 \(“Not at all”\), 1 \(“Several days”\), 2 \(“More than half the days”\), and 3 \(“Nearly every day”\)\.

For the primary analyses, we focus on eight PHQ\-9 items \(anhedonia, depressed mood, sleep problems, fatigue/low energy, appetite change, feelings of worthlessness or excessive guilt, concentration problems, and psychomotor changes\)\. We exclude the suicidality item \(thoughts of self\-harm\) because it is extremely sparse in both datasets, leading to unstable estimation in conditional and higher\-order analyses\. To reduce sparsity in the remaining items and harmonise measurement across datasets, we merge response categories 2 and 3 and analyse a three\-level ordinal scale\{0,1,2\}\\\{0,1,2\\\}for each retained item\. Because categories 2 and 3 are the rarest, this adjacent merge preserves ordinality but yields an asymmetric scale on which the lowest\-severity level is well sampled and the highest is compressed into a single cell, which reduces the attainable entropy, and hence the absolute mutual information in bits, for high\-severity symptom combinations\. Because we report atoms as shares of the target entropy \(Section[2\.8](https://arxiv.org/html/2609.13203#S2.SS8)\), however, this reduction acts largely as a common rescaling that cancels in the shares: recomputing a direct two\-source PID at two, three, and four levels yields a near\-constant synergy share of0\.1930\.193,0\.1880\.188, and0\.1900\.190respectively, with the redundancy share stable at0\.5340\.534to0\.5420\.542and hence dominant throughout\. We therefore treat the merge as likely conservative for synergy rather than as a fixed downward bias, and return to this point in Section[4\.4](https://arxiv.org/html/2609.13203#S4.SS4)\. After preprocessing, the UK Biobank dataset comprisedN=154,291N=154\{,\}291participants and the Xinxiang student sample comprisedN=24,292N=24\{,\}292respondents\.

To assess whether the ePID pipeline generalises beyond depressive\-symptom networks to instruments with different higher\-order structure, we additionally analyse responses to the Interpersonal Reactivity Index \(IRI\), a widely used multidimensional measure of dispositional empathy\([Davis, 1983](https://arxiv.org/html/2609.13203#bib.bib21)\)\. The IRI comprises 28 items distributed across four 7\-item subscales: Perspective\-Taking \(PT\) and Fantasy \(FS\), indexing cognitive\-empathy facets, and Empathic Concern \(EC\) and Personal Distress \(PD\), indexing affective\-empathy facets\([Davis, 1983](https://arxiv.org/html/2609.13203#bib.bib21)\)\. Responses are scored on a five\-point Likert scale; for the present analyses we discretise to a three\-level ordinal scale to harmonise cardinality with the PHQ\-9 protocol\. The dataset analysed here comprisesN=1,973N=1\{,\}973respondents from the Open\-Source Psychometrics Project \([https://openpsychometrics\.org/\_rawdata/](https://openpsychometrics.org/_rawdata/)\)\. We apply ePID using the same protocol as for PHQ\-9 \(nsources=5n\_\{\\mathrm\{sources\}\}=5ACIB\-embedded remainder context,ImmiI\_\{\\mathrm\{mmi\}\}redundancy measure\), supplemented by within\-subscale and cross\-subscale analyses reported in Section[3\.4](https://arxiv.org/html/2609.13203#S3.SS4)\.

### 2\.3Correlation Networks

We represent the symptom system as an undirected graphical model in which nodes are symptoms and edges represent conditional dependence after adjusting for the remaining symptoms\. Under a Gaussian \(or Gaussian\-copula\) graphical model, the absence of an edge corresponds to conditional independence given all other symptoms, while a nonzero partial correlation indicates a conditional association within that model class\.

The \(population\) partial correlation betweenXiX\_\{i\}andXjX\_\{j\}givenZi​jZ\_\{ij\}, denotedρi​j\|Zi​j\\rho\_\{ij\\mid Z\_\{ij\}\}, is defined as the correlation between the residuals obtained by regressingXiX\_\{i\}andXjX\_\{j\}onZi​jZ\_\{ij\}, respectively\.

To relax marginal normality and linearity assumptions, we compute rank\-based partial correlations \(SPC\)\. Specifically, each variableXkX\_\{k\}is transformed by its empirical marginal distribution to obtain ranksR⁡\(Xk\)R\(X\_\{k\}\), and partial correlations are computed on the transformed variables\. This procedure estimates conditional associations under a Gaussian copula graphical model: dependencies are encoded by a latent multivariate normal structure while marginal distributions are left unrestricted\. Rank\-based partial correlations therefore capture conditional monotonic dependence, irrespective of marginal scale or skewness\.

Partial correlations take values in\[−1,1\]\[\-1,1\]and define an undirected network in which edges represent conditional associations between symptom pairs\.

### 2\.4Information\-Theoretic Networks

To quantify conditional dependence without restricting the functional form of associations, we construct information\-theoretic networks based on mutual information\.

For a discrete random variableXX, Shannon entropyH⁡\(X\)H\(X\)quantifies uncertainty about its outcome on a logarithmic scale and provides a natural uncertainty baseline for categorical data\. For a variable with\|X\|\|X\|categories,H⁡\(X\)≤log2⁡\|X\|H\(X\)\\leq\\log\_\{2\}\|X\|\. Mutual information between two variablesXiX\_\{i\}andXjX\_\{j\},

I⁡\(Xi,Xj\)=H⁡\(Xi\)−H⁡\(Xi∣Xj\),I\(X\_\{i\};X\_\{j\}\)=H\(X\_\{i\}\)\-H\(X\_\{i\}\\mid X\_\{j\}\),measures the reduction in uncertainty about one variable given knowledge of the other\. Mutual information is symmetric, non\-negative, and makes no assumptions about linearity or monotonicity\.

To isolate direct associations in a network setting, we compute conditional mutual information \(CMI\),

I⁡\(Xi;Xj∣Zi​j\)=H⁡\(Xi∣Zi​j\)−H⁡\(Xi∣Xj,Zi​j\),I\(X\_\{i\};X\_\{j\}\\mid Z\_\{ij\}\)=H\(X\_\{i\}\\mid Z\_\{ij\}\)\-H\(X\_\{i\}\\mid X\_\{j\},Z\_\{ij\}\),which measures the remaining dependence betweenXiX\_\{i\}andXjX\_\{j\}after conditioning on all other variables\. Unlike correlation\-based measures, CMI detects arbitrary conditional dependence\. At the population level,I⁡\(Xi;Xj∣Zi​j\)=0I\(X\_\{i\};X\_\{j\}\\mid Z\_\{ij\}\)=0if and only ifXiX\_\{i\}andXjX\_\{j\}are conditionally independent givenZi​jZ\_\{ij\}for discrete variables, so CMI provides a model\-free criterion for conditional independence\. In finite samples, plug\-in estimates of CMI are positively biased; we address this via permutation testing \(Section[2\.9](https://arxiv.org/html/2609.13203#S2.SS9)\)\.

Correlation\-based measures and CMI estimate the same conditional dependence graph but differ in the class of dependencies they can detect\. Partial correlations detect only linear or monotonic conditional associations, whereas CMI provides a necessary and sufficient test of conditional independence for discrete data\. Because mutual information is in turn bounded above by the relevant entropies, we report a normalised CMI that can be interpreted as the fraction of remaining uncertainty explained \(Section[2\.8](https://arxiv.org/html/2609.13203#S2.SS8)\)\.

### 2\.5Higher\-order interactions

Having established the pairwise measures, we now move from pairwise dependence to interactions among more than two variables\. We proceed in three steps: \(i\) define the two\-source PID and its redundancy function, \(ii\) introduce embedding\-based PID \(ePID\), which compresses the multivariate remainder of the symptom system into a supervised discrete embeddingEi→kE\_\{i\\to k\}for each focal source–target pair, and \(iii\) map the resulting atoms onto directed network visualisations\. A stylised three\-variable example \(Appendix[J](https://arxiv.org/html/2609.13203#A10)\) builds intuition before the formal definitions\.

Multivariate extensions of mutual information quantify whether dependence among groups of variables is primarily overlapping or interaction\-specific\. A core diagnostic is that conditioning can reveal dependencies invisible to marginal analysis: in an exclusive\-or \(XOR\) relationship between two independent, uniformly distributed binary inputs, where the output indicates whether the inputs differ, each input alone carries no information about the output, whereas the joint state is fully informative\. We apply partial information decomposition \(PID\) to obtain non\-negative unique, redundant, and synergistic components, using an embedding of the remaining symptoms to make the decomposition tractable at PHQ\-9 scale\.

Several complementary frameworks provide scalable summaries of higher\-order dependence without estimating a full multivariate PID\. System\-level measures such as the O\-information \(and related local or dynamical variants\) quantify whether multivariate dependence is dominated by redundancy or synergy, but do not attribute information to specific source–target relations\([Rosas et al\., 2019](https://arxiv.org/html/2609.13203#bib.bib82);[Scagliarini et al\., 2023](https://arxiv.org/html/2609.13203#bib.bib86)\)\. Decision\-theoretic and entropy\-centric decompositions offer principled alternatives to the standard PID lattice: the redundancy bottleneck recasts redundancy as a constrained optimisation over compressed representations of the sources\([Kolchinsky, 2024](https://arxiv.org/html/2609.13203#bib.bib66)\), while partial entropy decomposition operates on joint entropy rather than mutual information\([Ince, 2017](https://arxiv.org/html/2609.13203#bib.bib55)\)\. Both can scale via convex optimisation or algebraic structure\. Finally, synergy\-backbone approaches aim to identify irreducible synergistic subsets in large systems while avoiding the full lattice\([Varley, 2024](https://arxiv.org/html/2609.13203#bib.bib110)\)\. These methods highlight a common trade\-off between scalability and attributional resolution; here we retain PID’s directional, source–target interpretability while scaling to the full symptom context via ePID’s remainder embedding \(Section[2\.5\.2](https://arxiv.org/html/2609.13203#S2.SS5.SSS2)\)\.

As a concrete anchor, consider a stylised three\-variable system \(stress, sleep problems, and concentration\) in which concentration tracks sleep under normal conditions but, when sleep is very poor, depends on stress in a non\-monotonic way\. A partial correlation recovers only the sleep\-to\-concentration link; conditional mutual information additionally flags the stress\-to\-concentration dependence; and PID makes the structure explicit, attributing about71%71\\%of the information about concentration uniquely to sleep,7%7\\%to redundancy, and21%21\\%to synergy between stress and sleep \(the interaction\-only signal that partial correlation misses entirely and conditional mutual information detects only indirectly\)\. The full worked construction is given in Appendix[J](https://arxiv.org/html/2609.13203#A10)\. This pattern is not an artefact of the construction: in the real IRI empathy data, the analogous configuration \(sources Fantasy and Personal Distress, target Perspective\-Taking\) is synergy\-dominant in69\.4%69\.4\\%of informative cross\-subscale triplets, with a mean synergy fraction of about21%21\\%, albeit only about0\.0060\.006bits in total\. We now formalise these atoms\.

#### 2\.5\.1Two\-source partial information decomposition

Partial information decomposition \(PID\) refines mutual information by decomposing the information that two sources\(Xi,Xj\)\(X\_\{i\},X\_\{j\}\)provide about a targetXkX\_\{k\}into four non\-negative components \(information “atoms”\): redundant, unique\-to\-XiX\_\{i\}, unique\-to\-XjX\_\{j\}, and synergistic information\. For discrete variables, we write

I⁡\(Xi,Xj,Xk\)=Red\+UnqXi\+UnqXj\+Syn,I\(X\_\{i\},X\_\{j\};X\_\{k\}\)=\\mathrm\{Red\}\+\\mathrm\{Unq\}\_\{X\_\{i\}\}\+\\mathrm\{Unq\}\_\{X\_\{j\}\}\+\\mathrm\{Syn\},\(1\)subject to the consistency relations

I⁡\(Xi,Xk\)\\displaystyle I\(X\_\{i\};X\_\{k\}\)=Red\+UnqXi,\\displaystyle=\\mathrm\{Red\}\+\\mathrm\{Unq\}\_\{X\_\{i\}\},\(2\)I⁡\(Xj,Xk\)\\displaystyle I\(X\_\{j\};X\_\{k\}\)=Red\+UnqXj\.\\displaystyle=\\mathrm\{Red\}\+\\mathrm\{Unq\}\_\{X\_\{j\}\}\.\(3\)That is, the four atoms partition the joint mutual informationI⁡\(Xi,Xj,Xk\)I\(X\_\{i\},X\_\{j\};X\_\{k\}\)into non\-overlapping non\-negative pieces, and individually they must sum to the marginal informationsI⁡\(Xi,Xk\)I\(X\_\{i\};X\_\{k\}\)andI⁡\(Xj,Xk\)I\(X\_\{j\};X\_\{k\}\)via the consistency constraints\. HereRed\\mathrm\{Red\}is information that both sources provide about the target,UnqXi\\mathrm\{Unq\}\_\{X\_\{i\}\}andUnqXj\\mathrm\{Unq\}\_\{X\_\{j\}\}are information that one source provides and the other does not, andSyn\\mathrm\{Syn\}is information that only becomes available when the two sources are considered jointly \(interaction\-only effects\)\.

Different PID measures pin down these four atoms in different ways\. Most, including the minimum mutual information \(MMI\) redundancy measure we use in all embedding\-based analyses, define the redundant information first and obtain the others from it; the BROJA measure \(below\) instead defines the unique information first, as the smallest unique information consistent with the source and target marginals\. MMI satisfies self\-redundancy, symmetry, and monotonicity and has been widely applied in discrete systems\. MMI defines redundancy as the minimum information that any single source provides about the target,RedMMI=min⁡\{I⁡\(Xi,Xk\),I⁡\(Xj,Xk\)\}\\mathrm\{Red\}\_\{\\mathrm\{MMI\}\}=\\min\\\{I\(X\_\{i\};X\_\{k\}\),\\,I\(X\_\{j\};X\_\{k\}\)\\\}, and derives the remaining atoms via the consistency relations above\. This definition is intuitive \(redundancy cannot exceed what any one source individually reveals\) and is computationally efficient for low\-cardinality systems\. A known limitation of MMI is that it can overestimate redundancy when two sources carry different information about the target of equal magnitude, because the minimum over marginal mutual informations equates equal MI values with shared content regardless of whether the sources encode different aspects of the target\([Bertschinger et al\., 2014](https://arxiv.org/html/2609.13203#bib.bib6)\)\. Consequently, MMI may underestimate synergy\. Alternative redundancy definitions address this limitation from different angles\. For example, the BROJA measure\([Bertschinger et al\., 2014](https://arxiv.org/html/2609.13203#bib.bib6)\)defines redundancy as the smallest amount of unique information consistent with the observed source–target marginals, expressed as a decision\-theoretic optimisation; but each introduces additional computational cost or conceptual trade\-offs\. To assess sensitivity to the redundancy definition, we additionally compute decompositions using four alternative measures as robustness checks: theIminI\_\{\\min\}\(Williams–Beer\) measure\([Williams and Beer, 2010](https://arxiv.org/html/2609.13203#bib.bib114)\), the original PID redundancy that takes the pointwise minimum specificity across sources; a rescaled redundancy measureIrrI\_\{\\mathrm\{rr\}\}\([Goodwell and Kumar, 2017](https://arxiv.org/html/2609.13203#bib.bib40)\)that normalises by source entropy; the pointwise partial information measureI±I\_\{\\pm\}\([Finn and Lizier, 2018](https://arxiv.org/html/2609.13203#bib.bib28)\), which decomposes information at the level of individual joint outcomes rather than averaged distributions; and the Gács–Körner common\-information measureI∧I\_\{\\wedge\}\([Gács and Körner, 1973](https://arxiv.org/html/2609.13203#bib.bib34)\), which retains only the information that two sources jointly encode in a perfectly aligned \(deterministic\) way\. Qualitative patterns in our empirical analyses \(in particular the ordering of edges by synergy fraction\) were consistent across all five measures\.

ComputingI±I\_\{\\pm\}andI∧I\_\{\\wedge\}for five or more sources via conventional lattice enumeration is computationally prohibitive\. The PID redundancy lattice, the hierarchy that organises all non\-trivial ways source subsets can overlap in their information about a target, contains7,5797\{,\}579antichains \(the lattice nodes that index distinct PID atoms\) atN=5N=5sources\. We bypass this bottleneck using the fast Möbius transform \(FMT\)\([Jansma et al\., 2025](https://arxiv.org/html/2609.13203#bib.bib59)\), which precomputes the lattice structure and obtains each target’s PID atoms via a single matrix–vector multiplication, reducing per\-target computation from hours to minutes\.

#### 2\.5\.2Embedding the remaining sources for symptom networks

In principle, PID can be extended to more than two sources, but the number of information atoms grows super\-exponentially with the number of sources \(from 4 atoms at two sources to over 7,500 at five; Section[2\.5\.1](https://arxiv.org/html/2609.13203#S2.SS5.SSS1)\), quickly becoming computationally and statistically intractable\. Two factors compound this\. First, the decomposition is not computed once but re\-estimated for every directed source–target pair, in each dataset, across the five redundancy measures, and across the calibrated synthetic ensemble used for validation, so the nominal per\-target cost is incurred many times over\. Second, stable estimation of a high\-order joint distribution requires adequate counts in each joint cell: the eight non\-target symptoms together with the target \(nine three\-level variables\) span39≈2×1043^\{9\}\\approx 2\\times 10^\{4\}joint states, leaving on the order of one observation per state in the Xinxiang student sample \(N≈24,000N\\approx 24\{,\}000\) and about eight in UK Biobank \(N≈154,000N\\approx 154\{,\}000\), far too sparse for reliable higher\-order estimates\. To approximate the high\-order structure while retaining interpretability, we adopt an embedding\-based strategy\. Intuitively, instead of estimating the full joint distribution of all remaining symptoms, we replace them with a single compact summary variable—a learned grouping of the possible remainder configurations that keeps what is informative about the target and discards the rest—so the decomposition need only handle the focal symptom together with this summary\. Throughout this section and in all empirical analyses, the two\-source decomposition uses the MMI redundancy measureImmiI\_\{\\mathrm\{mmi\}\}defined in Section[2\.5\.1](https://arxiv.org/html/2609.13203#S2.SS5.SSS1)\. The four alternative measures introduced there are reported as robustness checks and as a benchmark axis in the multi\-dataset evaluation, whose results are in Appendix[E\.1](https://arxiv.org/html/2609.13203#A5.SS1)\.

For each ordered symptom pair\(Xi,Xk\)\(X\_\{i\},X\_\{k\}\)withi≠ki\\neq k, we treatXkX\_\{k\}as the target andXiX\_\{i\}as a focal source\. The remaining variables𝐗∖\{Xi,Xk\}\\mathbf\{X\}\\setminus\\\{X\_\{i\},X\_\{k\}\\\}are compressed into a single discrete embedding

Ei→k=fi→k​\(𝐗∖\{Xi,Xk\}\)\.E\_\{i\\to k\}\\;=\\;f\_\{i\\to k\}\\\!\\left\(\\mathbf\{X\}\\setminus\\\{X\_\{i\},X\_\{k\}\\\}\\right\)\.This embedding provides a compact proxy for the joint configuration of the remaining symptom network when paired with the focal source in a two\-source PID\.

We consider multiple constructions offi→kf\_\{i\\to k\}, including \(i\)*target\-directed*embeddings that learnfi→kf\_\{i\\to k\}usingXkX\_\{k\}and \(ii\)*unsupervised*embeddings that compress𝐗∖\{Xi,Xk\}\\mathbf\{X\}\\setminus\\\{X\_\{i\},X\_\{k\}\\\}without access toXkX\_\{k\}\. All candidate embeddings are evaluated in a simulation benchmark against tractable multivariate ground truth \(Section[2\.5\.3](https://arxiv.org/html/2609.13203#S2.SS5.SSS3)\)\.

Our primary embedding, which we term the Agglomerative Conditional Information Bottleneck \(ACIB\), groups the joint states of the remainder symptoms into a small number of clusters and then iteratively merges the most similar clusters, choosing each merge so that the conditional information that the remainder carries about the target \(given the focal source\) is preserved as much as possible\. The result is a low\-cardinality discrete embeddingEi→kE\_\{i\\to k\}that summarises the rest of the symptom network without discarding the part of its dependence with the target that the focal source does not already explain\.

Formally, ACIB adapts conditional information\-bottleneck clustering\([Gondek and Hofmann, 2003](https://arxiv.org/html/2609.13203#bib.bib39)\)with agglomerative merging\([Slonim and Tishby, 1999](https://arxiv.org/html/2609.13203#bib.bib89)\)to the PID setting\. It operates in two phases: first, joint states of the remainder are assigned to clusters by minimising a conditional Kullback–Leibler\-divergence distortion; second, the closest cluster pair is greedily merged while constraining the relative loss in conditional mutual information,1−I⁡\(Xi;Xk∣Ei→k\)/I⁡\(Xi;Xk∣Zi→k\)1\-I\(X\_\{i\};X\_\{k\}\\mid E\_\{i\\to k\}\)/I\(X\_\{i\};X\_\{k\}\\mid Z\_\{i\\to k\}\), to remain below a specified tolerance\. Because this constraint operates on conditional rather than unconditional mutual information, the compression preserves the source–target relationship that PID subsequently decomposes\.

ACIB is fitted separately for each\(Xi,Xk\)\(X\_\{i\},X\_\{k\}\)pair, using𝐗∖\{Xi,Xk\}\\mathbf\{X\}\\setminus\\\{X\_\{i\},X\_\{k\}\\\}as inputs\. It has three hyperparameters: the conditional\-information loss tolerance, the cardinality capKmaxK\_\{\\max\}, and a Dirichlet smoothing constant, which we set to5%5\\%relative information loss,Kmax=12K\_\{\\max\}=12, and0\.50\.5throughout\. A dedicated sensitivity analysis \(Appendix[K](https://arxiv.org/html/2609.13203#A11)\) shows that fidelity is essentially insensitive to the loss tolerance and the smoothing constant and thatKmax=12K\_\{\\max\}=12is statistically indistinguishable from neighbouring values; ACIB was selected over alternative embeddings for its superior fidelity atN≥4N\\geq 4sources \(Section[2\.5\.3](https://arxiv.org/html/2609.13203#S2.SS5.SSS3)\)\.

Given the embedded system\(Xi,Ei→k,Xk\)\(X\_\{i\},E\_\{i\\to k\};X\_\{k\}\), we compute a two\-source PID using the MMI measure described above\. The resulting atoms,

Unqi→k,UnqEi→k,Redi→k,Syni→k,\\mathrm\{Unq\}\_\{i\\to k\},\\quad\\mathrm\{Unq\}\_\{E\_\{i\\to k\}\},\\quad\\mathrm\{Red\}\_\{i\\to k\},\\quad\\mathrm\{Syn\}\_\{i\\to k\},are interpreted as follows for each targetXkX\_\{k\}:

- •Unqi→k\\mathrm\{Unq\}\_\{i\\to k\}: information aboutXkX\_\{k\}that is carried uniquely by symptomXiX\_\{i\}beyond what is available from the embedded remainderEi→kE\_\{i\\to k\};
- •UnqEi→k\\mathrm\{Unq\}\_\{E\_\{i\\to k\}\}: information aboutXkX\_\{k\}that is carried uniquely by the rest of the network \(viaEi→kE\_\{i\\to k\}\) and not byXiX\_\{i\};
- •Redi→k\\mathrm\{Red\}\_\{i\\to k\}: information aboutXkX\_\{k\}that is shared betweenXiX\_\{i\}and the embedded remainder;
- •Syni→k\\mathrm\{Syn\}\_\{i\\to k\}: information aboutXkX\_\{k\}that becomes available only whenXiX\_\{i\}and the embedded remainder are considered together\.

These atoms provide an approximate decomposition of the information that all other symptoms collectively carry aboutXkX\_\{k\}into components attributable uniquely toXiX\_\{i\}, uniquely to the remainder, shared, and synergistic channels\. Because the embedding is many\-to\-one, the decomposition is approximate rather than exact for the original multivariate system; its accuracy is characterised in the simulation study below\.

In ePID, the remainder\-unique atomUnqEi→k\\mathrm\{Unq\}\_\{E\_\{i\\to k\}\}quantifies predictability of the target from the remainder of the symptom system alone\. In contrast, the redundancy and synergy atoms necessarily involve the focal sourceXiX\_\{i\}and therefore quantify*source involvement*in the source→\\rightarrowtarget dependence\. That is,UnqEi→k\\mathrm\{Unq\}\_\{E\_\{i\\to k\}\}quantifies*remainder\-only predictability*ofXkX\_\{k\},Redi→k\\mathrm\{Red\}\_\{i\\to k\}quantifies*overlap*between the focal symptom and the remainder, andSyni→k\\mathrm\{Syn\}\_\{i\\to k\}quantifies*interaction\-only information*that is available only from their joint configuration\.

Before computing the two\-source PID, each of the three variables \(focal source, embedding, and target\) is capped to at most five categories: the four most frequent values are retained and all remaining, rarer values are pooled into a single “other” category\. Unlike the adjacency\-preserving PHQ\-9 merge of Section[2\.2](https://arxiv.org/html/2609.13203#S2.SS2), this is a purely frequency\-based cap that does not preserve ordinal adjacency; for the three\-level PHQ\-9 items it is therefore a no\-op on the source and target and binds only on the embedding \(a nominal cluster label, for which order is irrelevant\), whereas for higher\-cardinality instruments such as the IRI it can pool non\-adjacent ordinal levels\. This cardinality cap ensures that the joint distribution\(Xi,Ei→k,Xk\)\(X\_\{i\},E\_\{i\\to k\},X\_\{k\}\)remains well\-sampled at typical clinical sample sizes\.

#### 2\.5\.3Simulation benchmark for embedding validation

Because the embedding\-based PID is approximate and the ground\-truth multivariate PID cannot be estimated from observational data at scale, we validate against calibrated synthetic systems where exact computation is tractable\. This benchmark serves two purposes: \(i\) quantifying how well ePID recovers network\-context unique, redundant, and synergistic structure, and \(ii\) selecting an embedding method that is accurate under empirically realistic marginal distributions\. We generate an ensemble of Bayesian networks calibrated to PHQ\-9 data and compare the embedding\-derived two\-source PIDs to the ground\-truth multivariate PID\.

##### Bayesian network ensemble\.

Using the UK Biobank PHQ\-9 item distributions as a reference, we learn a directed acyclic graph over nine discrete variables for benchmarking purposes\. We denote each synthetic Bayesian network byGGand write\(G,Xt,N\)\(G,X\_\{t\},N\)for the configuration in which targetXtX\_\{t\}and number of sourcesNNare evaluated on networkGG\. The empirical analyses exclude suicidal ideation due to sparsity \(only4\.3%4\.3\\%and7\.6%7\.6\\%of respondents endorsed any thoughts of self\-harm at severity≥1\\geq 1in UK Biobank and the Xinxiang student sample respectively\), but including it in the synthetic calibration provides a stress test in the presence of low\-prevalence states\. Conditional probability tables are estimated from the empirical data and then perturbed with small Dirichlet noise to create 50 distinct Bayesian networks\. For each network we draw600 000600\\,000samples by forward sampling\. This yields an ensemble of synthetic datasets that preserve the marginal distributions and broad dependency patterns of the PHQ\-9 items while allowing exact computation of information\-theoretic quantities\. Calibration diagnostics confirm that the synthetic ensembles closely match the empirical mutual information structure \(mean MI error6\.10%6\.10\\%; Kolmogorov–Smirnov statistic0\.1660\.166; see Figure[S2](https://arxiv.org/html/2609.13203#A4.F2)\)\.

##### Ground\-truth multivariate PID\.

For each Bayesian networkGG, each targetXtX\_\{t\}\(one of the nine variables\), and each number of sourcesN∈\{3,4,5\}N\\in\\\{3,4,5\\\}, we selectNNdistinct source variables from the remaining items and compute the fullNN\-source PID of\(Xi,…,XN,Xt\)\(X\_\{i\},\\ldots,X\_\{N\};X\_\{t\}\)using the MMI PID measure\. This yields a redundancy lattice whose atoms specify unique, redundant, and synergistic contributions for all combinations of sources\.

##### Embedding\-based PID on the same systems\.

For the same\(G,Xt,N\)\(G,X\_\{t\},N\)configurations, we approximate the multivariate PID using two\-source decompositions as follows\. For each choice of focal sourceXiX\_\{i\}among theNNsources, we embed the remainingN−1N\-1sources\{Xj:j≠i\}\\\{X\_\{j\}:j\\neq i\\\}into a one\-dimensional discrete variableEEusing each of 13 embedding methods spanning three families:

- •*Manifold and dimensionality reduction\.*ACIB, the agglomerative conditional information bottleneck described above\([Gondek and Hofmann, 2003](https://arxiv.org/html/2609.13203#bib.bib39);[Slonim and Tishby, 1999](https://arxiv.org/html/2609.13203#bib.bib89)\); multiple correspondence analysis followed bykk\-means clustering of factor scores\([Greenacre, 2017](https://arxiv.org/html/2609.13203#bib.bib41);[Halford, 2023](https://arxiv.org/html/2609.13203#bib.bib46)\); non\-negative matrix factorisation followed bykk\-means\([Lee and Seung, 1999](https://arxiv.org/html/2609.13203#bib.bib70)\); sliced inverse regression with quantile binning, which estimates a target\-supervised low\-dimensional projection\([Li, 1991](https://arxiv.org/html/2609.13203#bib.bib72);[Koepke, 2018](https://arxiv.org/html/2609.13203#bib.bib65)\); partial least squares with quantile binning, which extracts components maximising covariance with the target\([Wold et al\., 1984](https://arxiv.org/html/2609.13203#bib.bib115)\); and truncated singular value decomposition followed bykk\-means\([Golub and Van Loan, 2013](https://arxiv.org/html/2609.13203#bib.bib38)\)\.
- •*Feature selection\.*Conditional mutual information maximisation \(CMIM\), which selects features that maximise CMI with the target conditional on previously chosen features\([Fleuret, 2004](https://arxiv.org/html/2609.13203#bib.bib29)\); joint mutual information \(JMI\), which selects features by their average pairwise joint information with the target\([Yang and Moody, 1999](https://arxiv.org/html/2609.13203#bib.bib119)\); a greedy CMI\-based selector that adds variables one at a time to maximise the running conditional mutual information\([Brown et al\., 2012](https://arxiv.org/html/2609.13203#bib.bib14)\); and ReliefF, which scores features by their ability to distinguish nearest hits from nearest misses across classes\([Kononenko, 1994](https://arxiv.org/html/2609.13203#bib.bib67);[Urbanowicz et al\., 2018](https://arxiv.org/html/2609.13203#bib.bib107)\)\.
- •*Clustering\.*Agglomerative clustering with Hamming distance over the joint state vector\([Murtagh and Contreras, 2012](https://arxiv.org/html/2609.13203#bib.bib78);[Finch, 2005](https://arxiv.org/html/2609.13203#bib.bib27)\);kk\-modes, the categorical analogue ofkk\-means using mode\-based centroid updates\([Huang, 1998](https://arxiv.org/html/2609.13203#bib.bib54)\); and spectral clustering with a Hamming\-distance\-based affinity matrix\([Ng et al\., 2001](https://arxiv.org/html/2609.13203#bib.bib79)\)\.

We then compute a two\-source PID for\(Xi,E,Xt\)\(X\_\{i\},E;X\_\{t\}\)using the MMI measure\.

##### Amalgamation of multivariate PID\.

To compare the two\-source PID on\(Xi,E,Xt\)\(X\_\{i\},E;X\_\{t\}\)with theNN\-source PID on\(X1,…,XN,Xt\)\(X\_\{1\},\\ldots,X\_\{N\};X\_\{t\}\), we collapse the many atoms of theNN\-source decomposition into four two\-source atoms via a structural amalgamation that groups theN−1N\-1non\-focal sources into a single composite source, so that the projected atoms answer the same question as ePID: how much information aboutXtX\_\{t\}is attributable to the focal sourceXiX\_\{i\}beyond what the remainder provides\. The mapping is conservative with respect to synergy, classifying an atom as synergistic only when it necessarily requires contributions spanning both the focal source and the remainder group, so the amalgamated synergy is a lower bound on the true interaction\-only information betweenXiX\_\{i\}and the remainder\. The grouping preserves the two\-source consistency relations\. The full antichain classification rule and its derivation are given in Appendix[B\.6](https://arxiv.org/html/2609.13203#A2.SS6)\. Figure[3](https://arxiv.org/html/2609.13203#S2.F3)visualises this amalgamation step on the toy structural causal model used in the methods\-comparison figure of the main text\.

##### Error metrics\.

For each PID atomc∈\{UnqXi,UnqE,Red,Syn\}c\\in\\\{\\mathrm\{Unq\}\_\{X\_\{i\}\},\\mathrm\{Unq\}\_\{E\},\\mathrm\{Red\},\\mathrm\{Syn\}\\\}and each embedding method, we compute the absolute error

AEc=\|PIDc\(N​\-source\)−PIDc\(embed\)\|,\\mathrm\{AE\}\_\{c\}=\\left\|\\mathrm\{PID\}^\{\(N\\text\{\-source\}\)\}\_\{c\}\-\\mathrm\{PID\}^\{\(\\text\{embed\}\)\}\_\{c\}\\right\|,\(4\)where the first term is the amalgamated ground\-truth value and the second term is the corresponding atom from the two\-source PID on\(Xi,E,Xt\)\(X\_\{i\},E;X\_\{t\}\)\. In words, the absolute error captures the per\-atom gap between embedding\-based and ground\-truth values; smaller values indicate that the embedding more faithfully reproduces the multivariate decomposition for that atom\. We summarise errors across all\(G,Xt,N\)\(G,X\_\{t\},N\)configurations and focal sources\. Figure[4](https://arxiv.org/html/2609.13203#S3.F4)displays the mean absolute error by embedding method and PID atom; Appendix Figure[S8](https://arxiv.org/html/2609.13203#A6.F8)compares approximated versus ground\-truth synergy for one representative method from each family tier \(ACIB, JMI,kk\-modes\); and Appendix Figure[S9](https://arxiv.org/html/2609.13203#A6.F9)shows how synergy error scales with the number of sources\.

For the multi\-dataset benchmark \(Appendix[E\.1](https://arxiv.org/html/2609.13203#A5.SS1)\), we additionally report the total absolute errorTAE=∑cAEc\\mathrm\{TAE\}=\\sum\_\{c\}\\mathrm\{AE\}\_\{c\}summed across all four PID atoms, and the relative errorRE=TAE/I⁡\(Xi,E,Xt\)\\mathrm\{RE\}=\\mathrm\{TAE\}\\,/\\,I\(X\_\{i\},E;X\_\{t\}\), where the denominator is the ground\-truth total mutual information that the focal source and its embedded remainder jointly carry about the target \(equivalently, the sum of the four ground\-truth atoms\)\. RE expresses the approximation error as a fraction of the total information being decomposed, enabling comparison across datasets with different baseline information levels\.

The simulation study serves two purposes\. First, it establishes that embedding\-based PID can approximate the multivariate decomposition with low error when suitable embeddings are used\. Second, it provides empirical guidance for choosing embeddings in the empirical PHQ\-9 analyses, where the ground\-truth PID is unknown\.

##### Applicability beyond PHQ\-9\.

The ePID pipeline is designed to accommodate datasets with varying numbers of variables, response cardinalities, and sample sizes\. For variables with many categories, the general pipeline supports cardinality capping that merges infrequent categories to reduce the dimensionality of joint distributions before embedding and PID computation\. This step is exercised in the multi\-dataset benchmark \(Section[2\.6](https://arxiv.org/html/2609.13203#S2.SS6)\), where datasets vary in cardinality\. For the PHQ\-9 items analysed here, the three\-level ordinal scale after category merging is already low\-cardinality, making this additional step unnecessary\. Computational details for the embedding benchmark, including the subsampling strategy for distance\-based methods, are provided in Appendix[B\.6](https://arxiv.org/html/2609.13203#A2.SS6)\. Figure[2](https://arxiv.org/html/2609.13203#S2.F2)provides an overview of the full ePID evaluation pipeline, including both the embedding\-based path and the ground\-truth benchmark path\.

Input dataImputation &discretisationmerge cats\. 2 & 3→\\tok=3k\{=\}3Path A: ePIDFor each pair\(Xi,Xk\)\(X\_\{i\},X\_\{k\}\)CompressremainderEi→kE\_\{i\\to k\}\|E\|=k\|E\|=kstates2\-source PID\(Xi,E,Xk\)\(X\_\{i\},E;X\_\{k\}\)ePIDnetwork∀\(i,k\)\\forall\\;\(i,k\)13 embedding methods: Dim\. reduction:ACIB, SIR, PLS, MCA, NMF, SVD \(\+kk\-means\) Feature selection:CMIM, JMI, greedy CMI, ReliefF Clustering:agglomerative,kk\-modes, spectral \(Hamming\)Path B: ground truth83 datasets \+ synthetic BNsFor each pair\(Xi,Xk\)\(X\_\{i\},X\_\{k\}\)FullNN\-sourcePIDFMT forN≥5N\{\\geq\}5Amalgamate to2\-source atomsCompare\|Δ​AE\|\|\\Delta\\mathrm\{AE\}\|select bestembedding

Figure 2:Overview of the ePID evaluation pipeline
Note\.An end\-to\-end view of how raw symptom data become a PID\-decomposed network: a scalable embedding\-based path used for the main analyses runs in parallel with a small\-scale ground\-truth path used to validate it\. Raw data are preprocessed \(imputation, discretisation\) and then processed along these two paths\.*Path A \(top, blue\):*for each source–target pair, the remainder symptoms are compressed into a discrete embedding using one of 13 methods, and a two\-source PID is computed on \(source, embedding, target\)\.*Path B \(bottom, orange\):*for benchmark validation, the fullNN\-source PID is computed via the fast Möbius transform \(FMT\) and amalgamated to two\-source atoms for comparison with Path A\.Figure 3:ePID on the toy SCM: redundancy lattice, amalgamated atoms, truth\-versus\-ePID recovery, and the embedding shortcut
Note\.Same SCM as Figure[1](https://arxiv.org/html/2609.13203#S1.F1); visualises the four\-source PID lattice that the main pipeline amalgamates over\.\(a\)FullN=4N=4antichain lattice for\{X1,X2,X4,X5\}→X3\\\{X\_\{1\},X\_\{2\},X\_\{4\},X\_\{5\}\\\}\\to X\_\{3\}underIminI\_\{\\min\}: 166 atoms in total, of which 9 carry non\-zero information \(sum==1\.489 bits\)\. Node area scales with atom value; colour encodes the type\-1 amalgamation class\.\(b\)The nine non\-zero atoms grouped by their type\-1 amalgamation class \(synergyΣ=0\.839\\Sigma=0\.839, remainder\-uniqueΣ=0\.242\\Sigma=0\.242, redundancyΣ=0\.408\\Sigma=0\.408bits\), each labelled with its individual value\.\(c\)Amalgamated four\-atom view for focal sourceX1X\_\{1\}: ground truth versus ePID \(ACIB embedding,\|E\|=5\\lvert E\\rvert=5\); the arrow marks the amalgamation step that collapses the lattice into four atoms\. ACIB recoversRed\\mathrm\{Red\}exactly,Syn\\mathrm\{Syn\}at85%85\\%, andUnqE\\mathrm\{Unq\}\_\{E\}at73%73\\%, for87%87\\%of total joint information\. The lower row makes the ePID shortcut explicit: rather than enumerating the full lattice, the remaining sources\{X2,X4,X5\}\\\{X\_\{2\},X\_\{4\},X\_\{5\}\\\}are compressed into a single embeddingEEand a two\-source PID is computed on\(X1,E,X3\)\(X\_\{1\},E;X\_\{3\}\), recovering the amalgamated atoms directly \(synergy0\.7110\.711, remainder\-unique0\.1760\.176, redundancy0\.4080\.408bits\)\. Grey nodes carry zero information\. ACIB==Agglomerative Conditional Information Bottleneck\.

### 2\.6Multi\-dataset embedding benchmark

The synthetic Bayesian network benchmark \(Section[2\.5\.3](https://arxiv.org/html/2609.13203#S2.SS5.SSS3)\) provides controlled validation with known ground truth but is limited to a single data\-generating mechanism calibrated to PHQ\-9\. To assess generalisability, we additionally evaluated embedding fidelity across 83 real\-world datasets spanning clinical, epidemiological, and psychometric domains\.

##### Dataset assembly and preprocessing\.

We assembled 83 datasets from publicly available repositories spanning clinical, epidemiological, and psychometric domains, with 5–56 variables per dataset and sample sizes ranging from 32 to 445,000 observations\. Full details of dataset assembly, preprocessing, imputation, and discretisation are provided in Appendix[A\.1](https://arxiv.org/html/2609.13203#A1.SS1)\. In total, the benchmark comprises 2,477,784 individual comparisons across 5 PID measures, 83 datasets, 13 embedding methods, andN∈\{3,4,5\}N\\in\\\{3,4,5\\\}sources\.

##### Embedding hyperparameters\.

The 13 embedding methods were configured as follows: ACIB used a loss tolerance of 5% relative information loss with maximum embedding cardinalityKmax=12K\_\{\\max\}=12; clustering methods \(kk\-modes, spectral Hamming, agglomerative Hamming\) usedK=6K=6clusters; SIR and PLS used 2 latent dimensions withB=4B=4bins; MCA, SVD, and NMF used 2 components withK=6K=6kk\-means clusters; ReliefF used 100 neighbours withk=3k=3selected features; JMI and CMIM selectedk=3k=3features\. For distance\-based methods \(spectral and agglomerative Hamming\), datasets with more than 5,000 observations were subsampled to 5,000 rows for distance matrix computation\. All methods used a fixed random seed of 42 for reproducibility\.

##### Error metrics\.

For each embedding method and dataset, we computed the ground\-truthNN\-source PID via the fast Möbius transform\([Jansma et al\., 2025](https://arxiv.org/html/2609.13203#bib.bib59)\)and amalgamated it to a bivariate decomposition using the conservative amalgamation scheme described in Section[2\.5\.3](https://arxiv.org/html/2609.13203#S2.SS5.SSS3)\. Embedding error was quantified using the total absolute error \(TAE\) and relative error \(RE\) defined in Section[2\.5\.3](https://arxiv.org/html/2609.13203#S2.SS5.SSS3)\.

### 2\.7Multi\-measure robustness framework

PID decompositions are measure\-dependent: the choice of redundancy function determines how atoms are partitioned\. To assess the robustness of embedding fidelity across different PID axiomatisations, we computed ground\-truthNN\-source PIDs using the five redundancy measures introduced in Section[2\.5\.1](https://arxiv.org/html/2609.13203#S2.SS5.SSS1)\(ImmiI\_\{\\mathrm\{mmi\}\},IminI\_\{\\min\},IrrI\_\{\\mathrm\{rr\}\},I∧I\_\{\\wedge\},I±I\_\{\\pm\}\); we adoptImmiI\_\{\\mathrm\{mmi\}\}as the primary measure throughout and treat the remaining four as benchmark axes here\.

ForN≥5N\\geq 5sources,I±I\_\{\\pm\}andI∧I\_\{\\wedge\}were computed via the fast Möbius transform\([Jansma et al\., 2025](https://arxiv.org/html/2609.13203#bib.bib59)\)using precomputed antichains and Möbius matrices for the redundancy lattice, reducing per\-target computation from hours to minutes\.

For each dataset, target, number of sourcesN∈\{3,4,5\}N\\in\\\{3,4,5\\\}, and PID measure, we computed the fullNN\-source ground\-truth PID, amalgamated the resulting atoms onto two\-source equivalents using the conservative scheme described in Section[2\.5\.3](https://arxiv.org/html/2609.13203#S2.SS5.SSS3), and compared these against the corresponding embedding\-based two\-source PID\. This yields error profiles stratified by embedding method, PID measure, and number of sources, enabling a systematic assessment of which embedding methods are robust across different axiomatisations of redundancy and synergy\.

##### Dataset coverage\.

Coverage was identical across all five measures: complete ground\-truth PIDs were available for 83 datasets atN=3N=3, 79 atN=4N=4, and 77 atN=5N=5, the shortfall atN=4N=4arising from four NHANES datasets that lack sufficient variables\.

### 2\.8Normalisation and effect\-size scaling

Because information\-theoretic measures are expressed in bits and are bounded above by the uncertainty of the target, we report normalised effect sizes—a target\-normalised CMI \(nCMI\\mathrm\{nCMI\}\) and entropy\-share PID atoms—to facilitate comparisons across targets and cohorts; the definitions, equations, and the network\-visualisation pruning rule are given in Appendix[B\.7](https://arxiv.org/html/2609.13203#A2.SS7)\.

### 2\.9Significance testing for pairwise measures

Statistical significance of pairwise conditional\-dependence edges \(PC and CMI\) was assessed using permutation tests\. For each source–target pair\(Xi,Xj\)\(X\_\{i\},X\_\{j\}\), the target columnXjX\_\{j\}was permuted across participants while all other columns \(includingXiX\_\{i\}and the conditioning setZi​jZ\_\{ij\}\) were held fixed\. This preserves the joint distribution of the non\-target variables while destroying the focal source–target association\. The pairwise statistic was recomputed on each ofB=1,000B=1\{,\}000permuted datasets to form an empirical null distribution\. The one\-sidedpp\-value was computed asp=\(m\+1\)/\(B\+1\)p=\(m\+1\)/\(B\+1\), wheremmis the number of replicates with a test statistic at least as large as the observed value\. Multiple testing across all edges was controlled using the Benjamini–Hochberg false discovery rate \(FDR\) procedure atq=0\.05q=0\.05\([Benjamini and Hochberg, 1995](https://arxiv.org/html/2609.13203#bib.bib5)\)\.

##### Cross\-cohort stability\.

To assess replicability of the PID decomposition, we match directed source–target edges that are significant in both UK Biobank and the Xinxiang student sample and compute Spearman rank correlations of PID atom fractions \(redundancy, synergy, remainder\-unique\) across matching edges\. Mean absolute differences between dataset\-specific fractions are also reported\.

## 3Results

We present results in four stages\. First, we benchmark candidate embeddings against ground\-truth PID on synthetic Bayesian networks calibrated to PHQ\-9\. Second, we extend that benchmark across 83 real\-world datasets and five PID measures to assess generalisability\. Third, we apply the validated pipeline to the empirical PHQ\-9 networks in UK Biobank and the Xinxiang student sample, comparing pairwise conditional\-dependence estimates from partial correlations and CMI before decomposing each source–target dependence via ePID into redundancy, synergy, and remainder\-unique channels\. Fourth, we apply the same pipeline to the Interpersonal Reactivity Index as a contrasting psychometric instrument\.

### 3\.1Embedding benchmark for ePID

The central question for the benchmark is whether this compression preserves the resulting PID atoms\. This subsection serves to evaluate the quality of preserving PID atoms for different embedding algorithms and determine the optimal embedding algorithm for all subsequent analyses\.

To determine which embedding methods provide reliable approximations to high\-order PID, we evaluated 13 candidates on an ensemble of 50 Bayesian networks\{G1,…,G50\}\\\{G\_\{1\},\\ldots,G\_\{50\}\\\}calibrated to the joint distribution of the PHQ\-9 items\. BN calibration statistics are shown in Figures[S2](https://arxiv.org/html/2609.13203#A4.F2),[S3](https://arxiv.org/html/2609.13203#A4.F3)\. For each networkGG, each targetXtX\_\{t\}, and each number of sourcesN∈\{3,4,5\}N\\in\\\{3,4,5\\\}, we computed \(i\) the exactNN\-source PID using the MMI\-based redundancy measure indit, and \(ii\) a family of two\-source PIDs where one source was a focal variable and the other was a discrete embedding of the remainingN−1N\-1sources\. We then amalgamated the ground\-truthNN\-source atoms onto the four components of a two\-source PID and quantified the absolute error for each atom \(source\-unique, remainder\-unique, redundancy, synergy\) and each embedding method \(Methods[2\.5\.3](https://arxiv.org/html/2609.13203#S2.SS5.SSS3)\)\. This synthetic benchmark uses MMI \(minimum mutual information\) throughout, both for the ground\-truthNN\-source decomposition and for the embedding\-based approximation, so comparisons are internally consistent\. Because these results characterise approximation fidelity conditional on a given PID measure, they do not speak to the separate question of which measure best captures the information structure of interest; we return to measure dependence in the Discussion \(Section[4\.4](https://arxiv.org/html/2609.13203#S4.SS4)\)\. Cross\-measure generalisation, comparing each of five PID measures against its own ground truth, is evaluated separately in Appendix[E\.1](https://arxiv.org/html/2609.13203#A5.SS1)\.

Figure 4:Mean absolute error by embedding method and PID atom
Note\.Each panel shows mean absolute error \(in bits\) between the amalgamated ground\-truthNN\-source PID and the embedding\-based two\-source PID approximation, for one of the four PID atoms:\(a\)source\-uniqueUnqXi\\mathrm\{Unq\}\_\{X\_\{i\}\},\(b\)remainder\-uniqueUnqE\\mathrm\{Unq\}\_\{E\},\(c\)redundancyRed\\mathrm\{Red\},\(d\)synergySyn\\mathrm\{Syn\}\. Methods are coloured by family:*manifold / dim\-reduction*\(blue: ACIB, SIR, PLS, MCA\+kk\-means, SVD\+kk\-means, NMF\+kk\-means\),*information\-theoretic feature selection*\(pink: CMIM, JMI, ReliefF, Greedy CMI\), and*clustering*\(green: Agglomerative Hamming, Spectral Hamming,kk\-modes\)\. Errors are pooled across 50 synthetic Bayesian networks calibrated to PHQ\-9, nine targets, and three source counts \(N∈\{3,4,5\}N\\in\\\{3,4,5\\\}\); within each panel, methods are sorted by mean error\. Within the manifold family, ACIB \(the recommended default\) achieves the lowest error across all four atoms\. The clustering family is heterogeneous:kk\-modes performs comparably to manifold methods onRed\\mathrm\{Red\}andUnqXi\\mathrm\{Unq\}\_\{X\_\{i\}\}, while Hamming\-based clustering collapsesSyn\\mathrm\{Syn\}toward zero \(see also Figure[S4](https://arxiv.org/html/2609.13203#A5.F4)\)\. ACIB==Agglomerative Conditional Information Bottleneck; JMI==Joint Mutual Information; CMIM==Conditional Mutual Information Maximisation\.##### Method\-family stratification

Errors stratified strongly by embedding family \(Figure[4](https://arxiv.org/html/2609.13203#S3.F4)\)\. Manifold and dimensionality\-reduction methods \(ACIB, SIR, PLS, MCA\+kk\-means, SVD\+kk\-means, NMF\+kk\-means\) achieved the lowest overall mean absolute error across components \(≈0\.014\\approx 0\.014bits\), ahead of information\-theoretic feature selection approaches \(CMIM, JMI, greedy CMI, ReliefF; mean error≈0\.018\\approx 0\.018bits\) and clustering methods \(agglomerative Hamming, spectral Hamming,kk\-modes; mean error≈0\.022\\approx 0\.022bits\)\. Among all methods, the Agglomerative Conditional Information Bottleneck \(ACIB\) embedding showed the highest overall fidelity\. AtN=5N=5, ACIB achieved a mean total absolute error \(summed across all four PID atoms\) of0\.0480\.048bits, or about19%19\\%of the total mutual information \(TMI\); synergy\-specific errors remained small in both mean \(0\.00950\.0095bits\) and median \(0\.0080\.008bits\)\.

##### Component\-resolved performance\.

Most methods reproduced the source\-unique and redundancy atoms with near\-numerical precision \(median errors<10−10<10^\{\-10\}bits for 11 of 13 methods\), whereas discrepancies concentrated in the remainder\-unique and synergy atoms \(bottom panels of Figure[4](https://arxiv.org/html/2609.13203#S3.F4)\)\. A relative\-error view of the same data, in which per\-atom errors are normalised by the magnitude of the corresponding ground\-truth atom, is reported in Appendix Figure[S6](https://arxiv.org/html/2609.13203#A5.F6)\. For these challenging atoms, ACIB again performed best: the median error for the remainder\-unique atom was0\.01630\.0163bits and the median synergy error was0\.00580\.0058bits, whereas typical feature\-selection methods exhibited larger median synergy errors\. We report these errors as fractions of TMI, the exact ground\-truth total mutual information for each edge \(the sum of its four ground\-truth atoms\), rather than as fractions of each atom’s own true magnitude, because the synergy and remainder\-unique atoms are themselves near\-zero in many of the calibrated networks: a per\-atom denominator is then numerically unstable and can inflate relative errors without bound, whereas TMI provides a stable common scale across atoms\. The complementary per\-atom relative\-error view is the convention used in Appendix Figure[S6](https://arxiv.org/html/2609.13203#A5.F6), where exactly this instability is visible for the near\-zero atoms\. This pattern indicates that the main difficulty in approximating PID lies in preserving higher\-order and source\-specific information; embeddings that explicitly optimise information preservation fare better than those that optimise only geometric cohesion or pairwise criteria\.

##### Synergy fidelity across method families\.

Per\-pair comparison of approximated versus ground\-truth synergy illustrates the family\-level spread \(Appendix Figure[S8](https://arxiv.org/html/2609.13203#A6.F8)\)\. For ACIB \(manifold/dim\-reduction; Pearsonr=0\.92r=0\.92\), synergy estimates aligned closely with the identity line, indicating that the one\-dimensional embedding captured the range of synergistic information present in the original multi\-source systems\. For JMI \(feature\-selection;r=0\.76r=0\.76\), estimates were more scattered around the identity line but no longer strongly attenuated\. Forkk\-modes \(clustering;r=0\.75r=0\.75\), estimates aligned reasonably well, illustrating the upper end of clustering performance\. Under this calibration the Hamming\-based agglomerative and spectral clustering methods no longer collapse synergy toward zero \(agglomerativer=0\.68r=0\.68, spectralr=0\.62r=0\.62; see Figure[S4](https://arxiv.org/html/2609.13203#A5.F4)\); they recover a moderate share of the synergistic structure, though still less than the manifold methods\.

Feature selection methods \(CMIM, JMI, greedy CMI, ReliefF\) showed a distinct failure mode\. For systems with three sources, these methods approximated synergy almost perfectly \(errors near machine precision\)\. This is expected: withN=3N=3sources the remainder consists ofN−1=2N\-1=2variables, and feature\-selection embeddings that retain up to two features can store both losslessly, so the joint state of the remainder is preserved exactly and the resulting two\-source PID coincides with the three\-source ground truth\. The embedding step is therefore active \(two sources are still compressed into a single discrete embedding for the PID\), but no information is lost in that compression\. When the number of sources increased to five, synergy errors rose to≈0\.017\\approx 0\.017bits, a more modest increase, but still indicating that methods which select or weight features via greedy, pairwise, or univariate criteria capture the emergent joint information less completely as more sources interact simultaneously\.

##### Scaling with the number of sources\.

We next examined how approximation error changed with the number of sourcesNN\. For ACIB, mean synergy error increased smoothly from0\.00460\.0046bits atN=3N=3to0\.00950\.0095bits atN=5N=5, remaining below0\.010\.01bits throughout\. Because the UK Biobank\-calibrated networks carry little synergy per edge, these small absolute errors correspond to a larger fraction of the \(small\) per\-edge synergy: the synergy atom’s median per\-edge error, normalised by each edge’s synergy, rose from about21%21\\%atN=3N=3to76%76\\%atN=5N=5\. We therefore emphasise the absolute error, which stays below0\.010\.01bits; the relative figure is inflated by the near\-zero synergy denominator in these low\-synergy networks\. A sliced inverse regression \(SIR\) embedding with binning showed synergy error that increased withNN\(from≈0\.008\\approx 0\.008bits atN=3N=3to≈0\.016\\approx 0\.016bits atN=5N=5\), comparable to the other manifold methods, whereas clustering methods had both larger baseline errors and steeper increases withNN\. Thus, ACIB provides the best absolute fidelity, while SIR offers the most stable performance as systems become more complex\. The extended multi\-dataset benchmark \(Appendix[E\.1](https://arxiv.org/html/2609.13203#A5.SS1)\) confirms that ACIB’s advantage persists at higherNN, with nominally the shallowest error growth slope among all methods\.

##### Embedding choice for symptom networks\.

Based on these benchmarks, we use ACIB as the primary embedding in all empirical symptom\-network analyses\. The manifold/dimensionality family consistently preserved the structure of unique, redundant, and synergistic information with low error, whereas clustering and feature\-selection methods either underestimated synergy or collapsed it toward zero\. The main text therefore focuses on ACIB\-based two\-source PIDs\.

To confirm that embedding\-based ePID generalises beyond the PHQ\-9\-calibrated synthetic networks, we ran an extended benchmark of all 13 embeddings across 83 real\-world datasets \(clinical, epidemiological, and psychometric; 5–56 variables\) and all five PID measures atN∈\{3,4,5\}N\\in\\\{3,4,5\\\}sources—over 2\.4 million atom\-level comparisons\. The best\-performing embedding varied from dataset to dataset, but a single manifold\-family method, ACIB, was the most consistent choice, ranking first on 75–84% of datasets atN=5N=5forIminI\_\{\\min\},ImmiI\_\{\\mathrm\{mmi\}\}, andI∧I\_\{\\wedge\}; we adopt it as the default and report the full method\-by\-dataset comparison in Appendix[E\.1](https://arxiv.org/html/2609.13203#A5.SS1)\. The one measure no embedding approximates well is the signedI±I\_\{\\pm\}, whose pointwise structure does not survive compression\.

### 3\.2PHQ\-9 symptom networks

Having validated the embedding pipeline on synthetic and real\-world data, we apply it to the empirical PHQ\-9 networks in UK Biobank and the Xinxiang student sample\. We report two complementary views of the same data\. Section[3\.2\.1](https://arxiv.org/html/2609.13203#S3.SS2.SSS1)characterises pairwise conditional dependence using partial correlations and CMI; Section[3\.2\.2](https://arxiv.org/html/2609.13203#S3.SS2.SSS2)then decomposes those dependencies via ePID\. Reporting both side by side allows the ePID atoms to be read against the conventional partial\-correlation network used elsewhere in the symptom\-network literature\.

#### 3\.2\.1Pairwise conditional dependence: partial correlations vs conditional mutual information

We found that PHQ\-9 items in both cohorts have largely monotonic conditional dependence at the level detectable by Spearman partial correlation \(SPC\), while CMI provides a model\-free scale in bits for the same conditional\-dependence structure\. Figure[S12](https://arxiv.org/html/2609.13203#A8.F12)\(Appendix[H](https://arxiv.org/html/2609.13203#A8)\) compares pairwise conditional\-dependence estimates derived from partial correlations and conditional mutual information \(CMI\) in the Xinxiang student sample\. Pearson \(PPC\) and Spearman \(SPC\) partial correlations were highly similar at the edge level \(Fig\.[S14](https://arxiv.org/html/2609.13203#A8.F14)\), and both aligned closely with CMI\. Permutation tests with false discovery rate control \(q=\.05q=\.05\) yielded statistically significant CMI for 56/56 directed source–target pairs in UK Biobank and 54/56 in the Xinxiang student sample \(though many of the UK Biobank edges have nCMI well below 0\.005, reflecting the high power afforded byN≈154,000N\\approx 154\{,\}000; see Supplementary Table[S7](https://arxiv.org/html/2609.13203#A8.T7)for effect\-size–pruned counts\)\.

In the Xinxiang student sample the CMI and Spearman partial\-correlation networks were comparably dense \(54 significant directed CMI edges versus 54 for SPC atq=\.05q=\.05\), indicating that monotonic partial correlations and CMI flag a similar set of conditional dependencies at the item level\. Normalised CMI values \(nCMI=I⁡\(X;Y∣𝐙\)/H⁡\(Y\)\\mathrm\{nCMI\}=I\(X;Y\\mid\\mathbf\{Z\}\)/H\(Y\)\) ranged from 0\.02 to 0\.07 in the Xinxiang student sample and from 0\.0003 to 0\.06 in UK Biobank\. For context, the largest nCMI in the Xinxiang student sample indicates that the source symptom carries conditional information about the target amounting to about 7% of the target’s marginal entropy, after accounting for the rest of the network; most values are far smaller\. Conditional dependencies at the item level are therefore modest in absolute magnitude\.

#### 3\.2\.2Network\-context decomposition via ePID

We now decompose each conditional dependence into redundancy, synergy, and remainder\-unique channels, asking whether this finer\-grained view reveals structure that pairwise methods cannot express\.

##### Aggregate atom distributions and cross\-cohort stability\.

The composition of source–target information was similar in UK Biobank and the Xinxiang student sample: redundancy accounted for 51% of total mutual information \(TMI\) in UK Biobank and 45% in the Xinxiang student sample, synergy for 6% and 8%, and remainder\-unique information for 42% and 47%, respectively\. Under MMI with a target\-optimised embedding the source\-unique atom was zero or near\-zero for all edges, so the decomposition effectively partitions dependence into three non\-trivial channels: redundancy, synergy, and remainder\-unique\. We report atom fractions as shares of TMI,I⁡\(Xi,Ei→k,Xk\)I\(X\_\{i\},E\_\{i\\to k\};X\_\{k\}\), so that shares sum to one for each edge; target\-entropy\-normalised shares \(Eq\.[6](https://arxiv.org/html/2609.13203#A2.E6)\) are reported alongside in bits when needed\. Edge\-level ranges were also comparable across datasets: redundancy 11–79% \(UK Biobank\) and 19–68% \(Xinxiang student sample\), synergy 0\.4–28% and 1–20%, and remainder\-unique 20–79% and 13–80%\.

The cross\-dataset similarity of atom shares was matched by edge\-level concordance\. Across the 54 directed edges that were significant in both datasets \(FDR\-correctedq<\.05q<\.05\), PID atom fractions were positively correlated \(Spearmanρ\\rho: redundancy=0\.59=0\.59; synergy=0\.78=0\.78; remainder\-unique=0\.68=0\.68\), with mean absolute differences in atom share of0\.050\.05,0\.020\.02, and0\.060\.06\(i\.e\.55,22, and66percentage points of TMI\) respectively\. The redundancy concordance \(ρ=0\.59\\rho=0\.59\) is more modest than the synergy concordance\. Together, these aggregate and edge\-level convergences indicate that the relative importance of redundant, synergistic, and remainder\-unique channels is a replicable feature of PHQ\-9 symptom networks rather than a property of either dataset\.

![Refer to caption](https://arxiv.org/html/2609.13203v1/phq9_atom_heatmaps_combined.png)Figure 5:Entropy\-normalised PID atom shares across source–target pairs
Note\.Each cell is the atom fractionIatom​\(X→Y\)/H⁡\(Y\)I\_\{\\mathrm\{atom\}\}\(X\\to Y\)/H\(Y\)for a directed pair among the eight PHQ\-9 symptoms \(suicidality excluded\); rows are sources, columns are targets, and diagonals are blanked\. Top row \(a\): Xinxiang student sample; bottom row \(b\): UK Biobank\. Columns show the three non\-trivial PID channels under ACIB \(redundancy, synergy, and remainder\-unique\) with shared colour scales within each atom to make cross\-cohort comparisons visible\. The source\-unique atom is zero or near\-zero across all edges under MMI with a target\-optimised embedding \(see text above\) and is omitted; the conditional\-mutual\-information share at the edge level is visualised in Fig\.[6](https://arxiv.org/html/2609.13203#S3.F6)\.
##### Network visualisation\.

To make the PID networks interpretable at a glance, we report two complementary visualisations\. In the*full*view \(Fig\.[13\(a\)](https://arxiv.org/html/2609.13203#A8.F13.sf1), Appendix[H](https://arxiv.org/html/2609.13203#A8)\), each directed edgeX→YX\\to Yis segmented into remainder\-unique, redundancy, and synergy; edge thickness encodes the target\-normalised total informationI⁡\(X,EX→Y,Y\)/H⁡\(Y\)I\(X,E\_\{X\\to Y\};Y\)/H\(Y\)\. The*no\-remainder\-unique*view \(Fig\.[13\(b\)](https://arxiv.org/html/2609.13203#A8.F13.sf2)\) removes the remainder\-unique atom and rescales the remaining segments, so that colour differences directly compare redundant versus synergistic source involvement\. Read together, these views distinguish targets that are predictable primarily from the broader symptom context from targets where the focal symptom participates through overlap \(redundancy\) or interaction\-only structure \(synergy\)\.

Figure[6](https://arxiv.org/html/2609.13203#S3.F6)shows the resulting networks for UK Biobank and the Xinxiang student sample\. PID decomposes each retained source–target relation into overlapping versus interaction\-only channels\. Across retained edges, most dependence is carried by redundant and remainder\-unique components, indicating that the apparent influence of a focal symptom on a target typically reflects information already present in the surrounding symptom context\. Nonetheless, several edges display substantial synergistic components, consistent with the presence of interaction\-dependent configurations in which a symptom becomes informative about a target primarily in conjunction with the remainder of the network\.

\(a\)Xinxiang student sample\(b\)UK Biobank
Figure 6:Replication of ePID networks \(top\-3 incoming edges per target\) in UK Biobank and the Xinxiang student sample
Note\.Nodes are eight PHQ\-9 symptoms \(suicidality excluded\)\. Edges were first screened using permutation\-tested conditional mutual information \(CMI\) with Benjamini–Hochberg FDR control \(q=\.05q=\.05\) and then pruned for visualisation by retaining, for each target symptomYY, the three incoming edgesX→YX\\rightarrow Ywith the largest entropy\-normalised conditional mutual informationnCMIX→Y=I⁡\(X;Y∣𝐙\)/H⁡\(Y\)\\mathrm\{nCMI\}\_\{X\\rightarrow Y\}=I\(X;Y\\mid\\mathbf\{Z\}\)/H\(Y\)\. For each retained directed edge, we compute an embedding\-based two\-source partial information decomposition \(PID\) on\(X,EX→Y,Y\)\(X,E\_\{X\\rightarrow Y\};Y\)using an ACIB embedding of the remaining symptoms\. Edge thickness is proportional to total mutual informationI⁡\(X,EX→Y,Y\)I\(X,E\_\{X\\rightarrow Y\};Y\)\. Coloured segments indicate the fraction of this total carried uniquely by the embedded remainder \(blue\), redundantly byXXand the remainder \(orange\), or synergistically \(green\); a terminal black segment indicates direction\. Edges were selected for display using CMI; decompositions for all 56 directed pairs are reported in Appendix[H](https://arxiv.org/html/2609.13203#A8)\. Because nCMI conditions on the full remaining symptom set, it captures the source’s unique\-plus\-synergistic information and excludes redundancy; since the source\-unique atom is near zero throughout these networks, the top\-three\-per\-target display, if anything, emphasises synergistic edges rather than hiding them\. The reported atom fractions \(Fig\.[5](https://arxiv.org/html/2609.13203#S3.F5)\) are in any case computed from all 56 directed decompositions and all informative edges, not only the displayed subset\.

### 3\.3Depression\-status contrast in PHQ\-9 network composition

To test whether the compositional balance of the PHQ\-9 network depends on symptom severity, we split the UK Biobank sample by the standard PHQ\-9 cutoff \(total score≥10\\geq 10versus<10<10\) and estimated the ePID network separately in each group\. To remove sample size as a confound, the two arms were equal\-NNmatched atn=8,879n=8\{,\}879per group and each decomposition was averaged over five random subsamples of the larger \(non\-depressed\) group\. Relative to non\-depressed respondents, the depressed network shifts away from redundancy toward unique and synergistic information \(mean shares across the displayed top\-three\-per\-target edges: redundancy0\.34→0\.150\.34\\rightarrow 0\.15, unique0\.47→0\.590\.47\\rightarrow 0\.59, synergy0\.19→0\.260\.19\\rightarrow 0\.26; Fig\.[7](https://arxiv.org/html/2609.13203#S3.F7)\), the clearest signal being the contraction of redundancy\. Because the absolute synergy share is sample\-size dependent \(the full\-cohort PHQ\-9 synergy share is≈0\.08\\approx 0\.08\), these values should be read as a matched\-NN*contrast*between depression groups rather than as absolute synergy levels; the same qualitative shift holds over all 56 significant edges \(Fig\.[S16](https://arxiv.org/html/2609.13203#A8.F16)\)\.

\(a\)\(b\)Figure 7:ePID symptom networks in non\-depressed versus depressed UK Biobank respondents
Note\.Directed embedding\-PID networks for the eight retained PHQ\-9 symptoms \(suicidality excluded\), estimated separately by depression status\.\(a\)non\-depressed \(PHQ\-9 total<10<10\);\(b\)depressed \(PHQ\-9 total≥10\\geq 10\)\. The groups were equal\-NNmatched atn=8,879n=8\{,\}879per arm and each decomposition was averaged over five random subsamples of the larger group\. Both panels display the same 24\-edge skeleton \(the top three incoming edges per target by entropy\-normalised CMI, pooled across groups\)\. For each directed edge, coloured segments partition the pair–target total mutual information into unique information carried by the source plus embedded remainder context \(blue\), redundancy \(orange\), and synergy \(green\); a terminal black segment indicates direction, and edge thickness is proportional to total mutual information on a scale shared across both panels\. Relative to non\-depressed respondents, the depressed network shifts from redundancy toward unique and synergistic information \(mean shares: redundancy0\.34→0\.150\.34\\rightarrow 0\.15, unique0\.47→0\.590\.47\\rightarrow 0\.59, synergy0\.19→0\.260\.19\\rightarrow 0\.26\); the clearest visual signal is the contraction of the orange \(redundancy\) segments\. Because the absolute synergy share is sample\-size dependent \(full\-cohort PHQ\-9 synergy≈0\.08\\approx 0\.08\), the figure communicates the matched\-NN*contrast*between groups rather than an absolute synergy level\. All 56 significant directed edges are shown in Fig\.[S16](https://arxiv.org/html/2609.13203#A8.F16)\.
### 3\.4PID and ePID applied to the Interpersonal Reactivity Index

The PHQ\-9 atom decomposition \(Fig\.[5](https://arxiv.org/html/2609.13203#S3.F5)\) returned a modest synergy share: 6–9% of pair–target information across the cohorts analysed, with no individual edge synergy\-dominant\. This raises an interpretive question\. Does ePID systematically return low synergy, or do PHQ\-9 items share substantial overlapping information about a common depressive\-symptom construct by design? To distinguish these possibilities we applied the identical pipeline to the Interpersonal Reactivity Index \(IRI; Section[2\.2](https://arxiv.org/html/2609.13203#S2.SS2)\)\([Davis, 1983](https://arxiv.org/html/2609.13203#bib.bib21)\)\. Using protocol parameters matched to the multi\-dataset benchmark \(nsources=5n\_\{\\mathrm\{sources\}\}=5ACIB embedding,ImmiI\_\{\\mathrm\{mmi\}\}measure\), the global 28\-item IRI yielded a mean synergy fraction of 23\.1%, with 71% \(95/133\) of informative directed pairs synergy\-dominant\.

The IRI analyses combine three protocols matched to system size\. For within\-subscale analyses \(7 items per subscale, applied to each of the four subscales separately\), we apply ACIB\-based ePID atnsources=5n\_\{\\mathrm\{sources\}\}=5and pool atom fractions across the four subscales for reporting\. For cross\-subscale triplet analyses \(one item from each of three different subscales asX1X\_\{1\},X2X\_\{2\}, target\), we report exact two\-source PID without embedding, since the triplet system has only two sources and the embedding step is unnecessary\. For the cross\-subscale triplet analysis, synergy\-dominance percentages are computed on the subset of triplets with total mutual informationI⁡\(Xi,Xj,Xk\)≥0\.01I\(X\_\{i\},X\_\{j\};X\_\{k\}\)\\geq 0\.01bits, excluding those whose four atoms are individually too small to distinguish synergy and redundancy reliably\. For the full 28\-item analysis, we report two complementary runs: a primary protocol matched to PHQ\-9 \(nsources=5n\_\{\\mathrm\{sources\}\}=5ACIB\-embedded remainder,ImmiI\_\{\\mathrm\{mmi\}\}measure\) supplies the headline atom fractions reported below, and a full\-context run \(conditioning on all 26 remaining items\) supplies the per\-pair atom heatmaps and the source–target retention diagnostic in this section\. The two full\-instrument runs differ in conditioning\-set saturation: at\|Z\|=26\|Z\|=26many directed pairs haveI⁡\(Xi;Xk∣𝐙\)≈0I\(X\_\{i\};X\_\{k\}\\mid\\mathbf\{Z\}\)\\approx 0before any embedding, so the full\-context run is reported chiefly as a structural diagnostic rather than as the source of the synergy\-dominance claim\. Atom fractions across these protocols are comparable in interpretation only when read together as a gradient \(Table[1](https://arxiv.org/html/2609.13203#S3.T1)\)\.

The atoms reported here use the same two\-stage pipeline as the PHQ\-9 application: ACIB compresses each remainder context into an embedding withnsources=5n\_\{\\mathrm\{sources\}\}=5, after which a two\-source PID is computed using theImmiI\_\{\\mathrm\{mmi\}\}redundancy measure\. For the full\-context run \(conditioning on all 26 remaining items\), we restrict the atom\-fraction analysis to pairs whose ACIB embedding respected its 5% conditional\-mutual\-information loss tolerance, retaining 133 of 756 directed source–target pairs\.

Table 1:Scale\-dependent synergy\-dominance gradient across IRI configurations
Note\.Synergy fraction increases from within\-subscale IRI through cross\-subscale triplet IRI to the global 28\-item IRI; redundancy decreases in mirror image\. Atom fractions are means over informative directed pairs whose ACIB embedding respected its 5% conditional\-mutual\-information loss tolerance \(where applicable\); synergy\-dominant pairs are those for which the synergy share exceeds the redundancy share\.Atom fractions in Table[1](https://arxiv.org/html/2609.13203#S3.T1)change systematically with analysis scale\. Within a single 7\-item subscale \(e\.g\. Perspective\-Taking analysed in isolation\), items are designed to converge on a single empathy facet: redundancy is high \(29%\) and synergy is modest \(16%\)\. At the full 28\-item scale, synergy reaches 23% and dominates in 71% of pairs, while redundancy contracts to 18%\. When the analysis widens to triplets that cross subscales, redundancy drops to 23% and synergy rises to 22%\. This pattern is consistent with the multi\-facet construct theory of empathy\([Davis, 1983](https://arxiv.org/html/2609.13203#bib.bib21)\): within\-facet convergent validity manifests as information\-theoretic redundancy, and cross\-facet integration manifests as synergy\.

Three patterns within the IRI atoms warrant a closer reading\. First, among cross\-subscale source pairs, the combination of Fantasy and Personal Distress targeting Perspective\-Taking items is the most strongly synergy\-dominant configuration: 69% of such triplets cross the synergy–redundancy threshold\. Mature cognitive empathy is informed synergistically by the affective extremes: neither Fantasy nor Personal Distress alone is strongly informative, but their joint state is\. Second, the Fantasy subscale, rather than Davis’s cognitive and affective division, organises redundancy\. Every source pair containing Fantasy ranked above every pair without it \(24\.9% versus 20\.3% of TMI; difference 4\.6 percentage points, 95% CI \[2\.1, 6\.9\], cluster bootstrap over target items\), although individual pairs within each group were not separable\. Perspective\-Taking paired with Empathic Concern was the least redundant combination \(17\.2%; lowest in 94% of replicates\)\. Fantasy items therefore duplicate information the rest of the scale already carries, consistent with long\-standing doubts about the subscale’s validity\([Lawrence et al\., 2004](https://arxiv.org/html/2609.13203#bib.bib69)\); the cognitive and affective core of the instrument contributes the most distinct information\. Third, when pooled across all source pairs targeting any Perspective\-Taking item, 49% of pairs are synergy\-dominant, the highest subscale\-level synergy\-dominance rate\. Mature cognitive empathy emerges as the IRI facet most synergistically informed by the others, consistent with developmental accounts in which cognitive empathy emerges later than its affective counterpart and comes to integrate it\([Decety and Holvoet, 2021](https://arxiv.org/html/2609.13203#bib.bib22)\)\.

The dominance regime is not an artefact of pipeline choices\. The global 28\-item mean synergy is stable acrossnsources∈\{3,4,5\}n\_\{\\mathrm\{sources\}\}\\in\\\{3,4,5\\\}\(range 22\.2–24\.2%\) and across the twelve embedding methods compared in the multi\-dataset benchmark \(range 22–28%\), indicating that neither the ACIB cardinality nor the supervised\-embedding family drives the result\.

## 4Discussion

Symptom networks have become a widely used framework for representing psychopathology, with edges encoding the \(conditional\) dependence between symptom pairs\. Yet a scalar edge cannot express how several symptoms act together to inform a third, and partial information decomposition \(PID\), the natural tool for recovering that higher\-order structure, becomes intractable and hard to interpret beyond a handful of sources\. Embedding\-based partial information decomposition \(ePID\) makes network\-context partial information decomposition tractable for psychometric instruments: by compressing the multivariate remainder of a symptom network into a target\-directed discrete proxy, it decomposes each directed pair’s conditional dependence into redundant, synergistic, and unique channels at realistic item counts\. In both PHQ\-9 datasets, redundancy and remainder\-unique channels jointly dominated pair–target information, with a smaller but consistently non\-zero synergistic component, and this compositional profile replicated across two large, demographically distinct datasets \(synergy fraction cross\-datasetρ=0\.78\\rho=0\.78; redundancyρ=0\.59\\rho=0\.59\), one a population cohort \(UK Biobank\) and the other an open online sample \(the Xinxiang student sample\); the redundancy concordance is more modest than the synergy concordance\. Applied without parameter changes to the Interpersonal Reactivity Index, the same pipeline instead returned a synergy\-dominated decomposition, indicating that ePID resolves structural differences between instruments rather than imposing them\. That this added structure sits atop a monotonic pairwise backbone is an assumption check rather than a result in its own right: the close agreement of the Pearson and Spearman partial\-correlation networks \(Fig\.[S14](https://arxiv.org/html/2609.13203#A8.F14)\), together with the rarity of correlation\-invisible non\-linear dependencies across datasets \(Appendix[G](https://arxiv.org/html/2609.13203#A7)\), shows that ePID’s contribution is the compositional layer, not a relaxation of pairwise linearity\. After situating ePID against existing approaches, we organise the discussion around three questions: what PHQ\-9 redundancy dominance implies for depression\-instrument design, what the IRI contrast implies about construct heterogeneity, and what the principal methodological limitations and next steps are\.

### 4\.1What ePID adds, and how it relates to existing approaches

PPC, SPC, and CMI quantify whether a symptom pair is conditionally associated, but not whether that association reflects information shared with the broader symptom profile or information expressed only when symptoms are considered jointly\. PID converts the scalar edge weight into a compositional profile and so represents this distinction directly: two pairs with identical CMI may differ in whether their information is largely redundant with the surrounding context \(suggesting substitutability\) or partly synergistic \(suggesting interaction\-dependent structure\)\. The cross\-cohort stability of these fractions supports that the decomposition captures reproducible structure rather than cohort\-specific noise\.

Triplet\-level PIDs provide a useful local lens but cannot attribute information to the remainder of the symptom system \(all symptoms except the focal source and target\), so information redundant with unmodelled symptoms can be misclassified as unique or synergistic within a restricted triplet\. The embedding\-based approach fills this gap by including a proxy for the remainder as a source, yielding decompositions aligned with the network question: how a focal symptom behaves*in context*of the full symptom set\.

Embedding\-based PID provides an information\-theoretic analogue of “incremental contribution” that is not restricted to linear models and does not require enumerating all higher\-order interactions\. The simulation results indicate that manifold and dimensionality\-reduction methods, in particular target\-directed ones such as ACIB, preserve synergy structure, whereas Hamming\-based clustering substantially distorts it \(Fig\.[4](https://arxiv.org/html/2609.13203#S3.F4)\)\.

The stylised symptom triplet \(Appendix[J](https://arxiv.org/html/2609.13203#A10)\) illustrates this principle in miniature: when synergy is high, nonlinear models dramatically outperform linear models, and partial correlations fail to detect the underlying interaction\. Although the deterministic structure of the example exaggerates the effect relative to empirical data, the implication carries over: edges with high synergy fractions in the PHQ\-9 networks may represent dependence patterns that additive clinical models cannot capture\. Testing that possibility would require predictive validation against outcomes, which the present cross\-sectional design cannot provide, and would be subject to the estimation caveats set out in Section[4\.4](https://arxiv.org/html/2609.13203#S4.SS4)\.

A central methodological contribution is the validation of ACIB as a practical embedding for network\-context PID\. The simulation benchmark demonstrates that this supervised design choice is consequential: target\-directed embeddings preserve synergy with high fidelity, whereas similarity\-based clustering collapses synergy toward zero\. This supervised\-embedding step therefore provides a general recipe for extending interaction\-aware information decompositions to symptom systems of realistic size\.

Recent work has argued that common dimensionality\-reduction techniques, particularly variance\-preserving methods such as PCA, can preferentially retain redundant structure while failing to preserve synergistic higher\-order dependencies\([Varley et al\., 2025](https://arxiv.org/html/2609.13203#bib.bib111)\)\. Our embedding\-based PID does not treat the symptom system as lying on a low\-dimensional manifold and does not optimise a geometric or variance criterion\. Instead, we use a target\-directed discrete coarse\-graining of the remaining symptoms to form a proxy conditioning variable and then quantify unique, redundant, and synergistic contributions in the embedded two\-source system\. This distinction matters for interpretation: our goal is not to recover a continuous latent geometry of symptoms, but to obtain an interpretable approximation to the distribution of information about a target across a focal symptom and the remainder of the network\.

Several complementary frameworks quantify higher\-order statistical structure in multivariate systems, and situating ePID relative to these clarifies its specific contribution\. Unlike system\-level measures such as the O\-information \(Section[2\.5](https://arxiv.org/html/2609.13203#S2.SS5)\), ePID provides per\-edge resolution, decomposing each directed pair’s dependence into redundant, synergistic, and unique channels\. In neuroscience, PID has been applied to characterise information processing in neural circuits, revealing synergistic coding and redundant representations across neural populations\([Timme and Lapish, 2018](https://arxiv.org/html/2609.13203#bib.bib100)\)\. Symptom networks differ from neural data in important respects: fewer variables, ordinal rather than continuous measurement, and clinical constraints on data collection\. ePID addresses the scalability challenge differently from neural applications, by embedding the multivariate remainder rather than restricting analysis to small subsets of variables, making it applicable to the moderate\-dimensional, discrete systems typical of psychopathological research\.

Among psychometric methods specifically, the closest existing approach to higher\-order structure is the Moderated Network Model \(MNM\) framework\([Haslbeck et al\., 2021](https://arxiv.org/html/2609.13203#bib.bib49)\), which warrants direct contrast on three axes\. First, MNM is parametric: it extends a mixed graphical model with multiplicative interaction terms and reads edge moderation off estimated coefficients, whereas ePID is model\-free and information\-theoretic, with synergy and redundancy obtained from joint distributions via PID atoms and no functional\-form assumption on the interaction\. The trade\-off is between interpretable coefficients with standard parametric inference on the one hand, and freedom from any assumed functional form for the interaction on the other\. Second, MNM requires the analyst to nominate one or a few candidate moderators and enter them explicitly, which suits confirmatory questions about specific moderators; ePID instead treats the entire remainder of the symptom profile simultaneously as conditioning context via the target\-directed embedding, which suits exploratory characterisation of system\-wide higher\-order structure\. Third, the two methods identify different objects: MNM tests*which*edges are moderated by a chosen variable and by how much, yielding a moderated\-edge map, whereas ePID quantifies*how much*of each edge’s information is interaction\-dependent \(synergy\), overlapping with the remainder \(redundancy\), or carried uniquely: a per\-edge compositional profile\. The two outputs are not interchangeable: a moderated edge under MNM and a synergistic edge under ePID need not coincide\. MNM and ePID therefore occupy complementary positions in the modelling\-flexibility and pre\-specification trade\-off, and the choice between them depends on whether the analyst seeks confirmatory tests of specific moderators or exploratory decomposition of system\-level higher\-order information\.

### 4\.2What does PHQ\-9 redundancy dominance imply for instrument design?

In both cohorts the decomposition was dominated by redundancy and remainder\-unique channels, and this composition is itself informative about how the instrument is built\. It is the information\-theoretic signature of a well\-targeted, largely unidimensional severity scale whose items function as partial substitutes for one another, which aligns with longstanding critiques of depression sum scores assembled from overlapping, content\-similar items\([Fried et al\., 2016](https://arxiv.org/html/2609.13203#bib.bib32);[Fried et al\., 2017](https://arxiv.org/html/2609.13203#bib.bib33)\)\. What that redundancy looks like in detail, and where synergy nonetheless appears, is the question we now address, beginning with the dominant channels and working towards the synergistic edges that pairwise measures cannot represent\.

##### Redundancy and remainder\-unique channels dominate\.

In both cohorts, redundancy and remainder\-unique channels jointly accounted for the vast majority of pair–target information \(the Xinxiang student sample: 45% redundancy, 47% remainder\-unique; UK Biobank: 51% redundancy, 42% remainder\-unique\), with synergy contributing 6–9%\. This finding should be interpreted as a structural characterisation of the symptom network rather than a limitation of the method\. High redundancy between a focal symptom and the remainder means that the information the focal symptom carries about the target is also available through other routes in the symptom network, regardless of the underlying generative model\. This is a statement about the information topology of the system: many symptoms share overlapping information channels about a given target\. The practical value of redundancy maps is that they identify which edges are “substitutable” \(where the focal symptom’s contribution to predicting the target is largely duplicated by other symptoms\) and which carry unique or interaction\-dependent information that would be lost if the focal symptom were removed from the model\. This distinction is not available from PC or CMI alone\.

##### Specific synergistic symptom pairs\.

Synergistic contributions were not uniformly distributed across targets, but the distribution proved reproducible without being attributable to particular symptoms\. The edge\-level ranking of synergy agreed closely between the two cohorts \(Spearmanρ=0\.78\\rho=0\.78\), and the agreement was insensitive to normalisation \(ρ=0\.76\\rho=0\.76with synergy expressed in bits\), so the relative distribution of synergy is a reproducible property of the PHQ\-9 network rather than of the chosen scale\. Per\-symptom rankings, by contrast, were not stable\. The target ranked highest on synergy fraction differed between cohorts \(psychomotor disturbance in the Xinxiang student sample, energy in UK Biobank\), and both rankings reordered substantially when synergy was expressed in bits rather than as a fraction of pair–target information \(Fig\.[5](https://arxiv.org/html/2609.13203#S3.F5)\)\. We therefore report the reproducibility of the synergy profile and refrain from a symptom\-specific interpretation, which the present sample sizes and the sensitivity to normalisation do not support\. We record this as a negative result, since naming individual synergistic symptom pairs is the use to which a decomposition of this kind would most naturally be put: ePID recovers the composition of dependence reliably at the level of the whole network, but these data do not identify which PHQ\-9 pairs carry the synergy\. Localisation of that kind will require either uncertainty intervals on individual atoms, which demand participant\-level resampling, or instruments whose synergy is concentrated rather than diffuse, as the Interpersonal Reactivity Index proved to be\. In contrast, for many pairs the embedded remainder alone captured most of the pairwise information, indicating that these targets are largely determined by the broader symptom context rather than by any single additional source\.

##### Clinical reading\.

The introduction states that “if two source symptoms are highly redundant, changing the state of only one source would change little to the target’s state\.” The empirical results allow us to revisit this observation concretely\. Redundancy\-dominated edges suggest that the focal symptom’s predictive contribution to the target is largely duplicated by other symptoms; removing it from a predictive model would change little, because the same information is available through other routes\. Synergistic edges suggest a different pattern: the information about the target would be lost if either the focal symptom or the remainder configuration were unobserved, because it arises only from their joint state\. These are hypothesis\-generating insights rather than treatment recommendations, given the cross\-sectional design and the observational nature of the data\. Nonetheless, they illustrate how the redundancy–synergy distinction can generate qualitatively different clinical hypotheses from the same conditional\-dependence backbone\.

Whether such hypotheses generalise depends on whether the compositional profile itself is reproducible across samples, an empirical question that connects the within\-cohort findings above to the broader symptom\-network replicability literature\.

The cross\-cohort stability of PID atom fractions speaks to the broader replicability debate in symptom network research\. Previous work has raised concerns about the replicability of network edge weights and centrality indices across samples\([Forbes et al\., 2017](https://arxiv.org/html/2609.13203#bib.bib30);[Fried et al\., 2017](https://arxiv.org/html/2609.13203#bib.bib33)\)\. The*composition*of dependence, and not only its magnitude, replicates across two large and demographically distinct datasets\. It does not, however, replicate better than the pairwise quantities it decomposes\. Over the same 28 symptom pairs, cross\-cohort rank agreement wasρ=0\.82\\rho=0\.82for CMI andρ=0\.74\\rho=0\.74for both partial\-correlation variants, againstρ=0\.76\\rho=0\.76,0\.590\.59and0\.590\.59for the synergy, redundancy and remainder\-unique fractions respectively\. The compositional layer should therefore be read as reproducible to a degree comparable with standard edge weights, rather than as a more stable representation of network structure\.

### 4\.3What does the IRI contrast imply about construct heterogeneity?

##### Synergy dominance and the construct\-structure hypothesis\.

The cross\-instrument application to the Interpersonal Reactivity Index \(Section[3\.4](https://arxiv.org/html/2609.13203#S3.SS4)\) demonstrates that the redundancy dominance observed in PHQ\-9 is not an artefact of the embedding pipeline\. The same ACIB\-based ePID, applied without parameter changes, returned a synergy\-dominated decomposition on the IRI: 71% of informative directed pairs had higher synergy than redundancy, against 0% for PHQ\-9\. The monotonic gradient across analysis scales \(within\-subscale 16% synergy; cross\-subscale triplets 22%; full instrument 23%\) indicates that the dominance regime is shaped by the construct structure of the instrument rather than by the embedding choice\. The same ePID pipeline that returned a redundancy\-dominated PHQ\-9 recovered a synergy\-dominated IRI without parameter changes, indicating that the dominance regime is a property of the instrument rather than of the method\. Together with the PHQ\-9 anchor, these findings indicate that ePID discriminates between redundancy\-dominated and synergy\-dominated psychometric instruments using identical machinery: the decomposition captures structural differences in how information is distributed across items, rather than artefacts of the instrument or the embedding choice\.

##### Specific synergy patterns\.

Three patterns within the IRI atoms warrant a closer reading\. First, in the cross\-subscale raw 2\-source PID \(no remainder embedding\), the combination of Fantasy and Personal Distress targeting Perspective\-Taking items is the most strongly synergy\-dominant configuration: 69% of such triplets cross the synergy–redundancy threshold\. This suggests that mature cognitive empathy is informed synergistically by affective extremes: neither Fantasy nor Personal Distress alone is strongly informative, but their joint state is\. Second, also within the cross\-subscale 2\-source PID, redundancy is organised by the Fantasy subscale rather than by Davis’s cognitive and affective division: pairs containing Fantasy carried reliably more redundant information than pairs without it, while pairs within each group were not separable, and Perspective\-Taking with Empathic Concern was the least redundant combination\. Fantasy items therefore duplicate information the rest of the scale already carries, consistent with long\-standing doubts about the subscale’s validity\([Lawrence et al\., 2004](https://arxiv.org/html/2609.13203#bib.bib69)\)\. Third, in the full 28\-item ePID \(nsources=5n\_\{\\mathrm\{sources\}\}=5ACIB embedding,ImmiI\_\{\\mathrm\{mmi\}\}\), 49% of source pairs targeting any Perspective\-Taking item are synergy\-dominant, the highest subscale\-level synergy\-dominance rate, consistent with developmental accounts in which cognitive empathy emerges later than its affective counterpart and comes to integrate it\([Decety and Holvoet, 2021](https://arxiv.org/html/2609.13203#bib.bib22)\)\.

##### Theoretical implications and limitations\.

The redundancy–synergy distinction therefore captures structural differences between psychometric instruments, not just between symptoms within an instrument\. Instruments designed around a single underlying construct \(PHQ\-9 depressive symptoms\) yield high redundancy because items are partial substitutes for one another, whereas multi\-facet instruments \(IRI’s four empathy facets\) yield high synergy because diagnostic information about a target item often resides in the joint configuration of items from different facets\. This suggests an information\-theoretic axis for psychometric construct validation that complements conventional reliability and factor\-analytic criteria\. Because self\-report items are imperfectly reliable, and measurement noise attenuates higher\-order structure more than linear structure, by analogy with the reliability attenuation of degree\-kkinteractions\([Schulz and Ritter, 2026](https://arxiv.org/html/2609.13203#bib.bib87)\)the reported synergy fractions are best read as a plausible lower bound\. Two caveats qualify the IRI\-specific findings\. The IRI sample is drawn from the Open\-Source Psychometrics Project \(a self\-selected online sample whose composition differs from clinical or community samples\), and, in the full\-context analysis \(conditioning on all 26 remaining items\), the atom\-fraction analysis is restricted, by design, to the 133 of 756 directed pairs \(17\.6%\) for which the ACIB embedding preserved at least 95% of the source–target conditional mutual information\. The 5% loss tolerance is applied because the atoms of a PID are unstable summaries when the embedding has thrown away an appreciable share of the quantity being decomposed; the remaining pairs therefore represent the IRI source–target combinations on which the decomposition can be read with confidence\. The IRI patterns should therefore be read as a within\-instrument illustration of discriminative validity rather than a population\-level inference about empathy structure\.

### 4\.4Methodological limitations and next steps

The choice of PID measure is itself a non\-trivial modelling choice, since PID has no single canonical measure; we therefore reportImmiI\_\{\\mathrm\{mmi\}\}as primary withIminI\_\{\\min\}as a robustness check, and the qualitative patterns \(edge rankings by synergy fraction, cross\-cohort stability\) held across both\. The multi\-dataset benchmark across 83 real\-world datasets and five PID measures provides a comprehensive assessment of embedding fidelity for this class of methods\. Three findings deserve emphasis\.

First, ACIB’s dominance is robust across four of five measures \(ImmiI\_\{\\mathrm\{mmi\}\},IminI\_\{\\min\},IrrI\_\{\\mathrm\{rr\}\},I∧I\_\{\\wedge\}\) and across diverse data structures, from small clinical samples to large population surveys\. AtN=5N=5, ACIB won 84\.4% of per\-dataset contests forI∧I\_\{\\wedge\}\[95% CI: 78\.4, 91\.1\], 75–78% forIminI\_\{\\min\}/ImmiI\_\{\\mathrm\{mmi\}\}, and 64\.9% forIrrI\_\{\\mathrm\{rr\}\}\[56\.2, 73\.3\] \(its most modest margin\), with the lowest error growth slope among all methods \(0\.041 \[0\.035, 0\.046\] bits per additional source\)\. This consistency suggests that target\-directed coarse\-graining is a generally effective strategy for preserving the information structure that PID quantifies, not an artefact of the synthetic benchmark’s calibration to PHQ\-9\.

Second, the striking failure of ACIB on theI±I\_\{\\pm\}\(pointwise partial information\) measure, where ACIB won only 2\.6% of contests, has a specific mechanistic origin\. Unlike non\-negative measures,I±I\_\{\\pm\}produces signed atoms: 77\.7% of ground\-truth unique\-source values were negative atN=5N=5\(compared to 0\.7% forIminI\_\{\\min\}\)\. The embedding compression disrupts the structured cancellation between positive and negative contributions, inflating absolute error beyond the total signal \(RE\>100\>100%\)\. This finding has practical implications: researchers who wish to use pointwise PID measures should consider alternative embeddings such as ReliefF or SVD\+kk\-means, though errors remain substantial\.

Third, atN=3N=3all methods were statistically indistinguishable \(Appendix[E\.1](https://arxiv.org/html/2609.13203#A5.SS1)\), consistent with the minimal compression required when only two sources are embedded into one variable\. Method choice becomes consequential only atN≥4N\\geq 4, where ACIB’s information\-bottleneck objective provides a clear advantage\.

More broadly, the finding that no single embedding method dominates all PID measures cautions against treating any embedding as a universal approximation\. The choice of embedding should be aligned with the choice of PID measure, and sensitivity analyses across both dimensions are advisable\.

##### Limitations\.

Three limitations qualify the interpretation of these findings\. First, embedding fidelity has been validated forN=3N=3toN=5N=5sources, whereas the empirical setting embedsN=6N=6remainder symptoms\. Computing a ground\-truth PID atN=6N=6is infeasible \(the redundancy lattice contains approximately7\.8×1067\.8\\times 10^\{6\}antichains\)\. The sub\-linear error growth fromN=3N=3toN=5N=5and the bounded atom magnitudes \(≤H⁡\(Xk\)\\leq H\(X\_\{k\}\)\) support but do not guarantee extrapolation toN=6N=6; accordingly, the empirical atom fractions should be interpreted as approximate compositional profiles rather than exact quantities\.

Second, sampling uncertainty in the atom estimates themselves is not formally quantified\. We report point estimates from plug\-in \(frequency\-based\) estimators, for which the entropy term is negatively biased in finite samples while the mutual\-information functionals built from it are typically biased upward\([Paninski, 2003](https://arxiv.org/html/2609.13203#bib.bib80)\); participant\-level bootstrap confidence intervals for PID atoms would be informative but are computationally demanding because each resample requires re\-fitting the ACIB embedding for all source–target pairs\. The confidence intervals we do report, for contrasts between groups of source–target pairs, are obtained by resampling analysis units \(target items or directed edges\) rather than participants, and therefore describe variability across pairs within a cohort, not estimation error in any individual atom\. They should not be read as sampling intervals on the atom values\. Bias is partially mitigated by the large sample sizes \(N\>24,000N\>24\{,\}000for the Xinxiang student sample;N\>154,000N\>154\{,\}000for UK Biobank\) and the low cardinality \(three levels per item\), but no formal bias correction was applied\. Small differences in atom fractions across edges or cohorts should therefore be interpreted cautiously\.

Third, three design choices in the empirical pipeline systematically affect atom estimates\. We merged the two highest PHQ\-9 response categories to reduce sparsity, which may underestimate synergy by collapsing fine\-grained joint configurations that distinguish interaction\-dependent patterns from additive ones\. Separately, the ACIB embedding is optimised to preserve information about the target, which by design ensures the remainder embedding is maximally informative aboutXkX\_\{k\}and could inflate the remainder\-unique component relative to a target\-agnostic embedding\. The decomposition should therefore be interpreted as conditional on these design choices: estimated synergy fractions are likely conservative, and the atoms describe how information aboutXkX\_\{k\}is distributed between the focal source and a*target\-optimised*summary of the remainder\. A third design choice concerns the 5% relative\-CMI tolerance used by ACIB, which serves both as an internal stopping rule and as a downstream retention filter in the full\-context IRI analysis\. The 17\.6% retention rate \(133 of 756 pairs\) reflects principally that conditioning on 26 of the 28 IRI items saturates the conditioning set: most directed pairs haveI⁡\(Xi;Xk∣𝐙\)I\(X\_\{i\};X\_\{k\}\\mid\\mathbf\{Z\}\)at or below numerical zero before any embedding step, and the filter therefore drops mathematically vacuous pairs rather than badly compressed ones\. The PHQ\-9 application restricts the conditioning set to the five highest\-mutual\-information sources and so does not encounter this saturation regime \(45/56 directed pairs retained\); the divergent retention rates between PHQ\-9 and IRI reflect the protocol’s interaction with instrument size, not a property of the embedding family\. The 5% threshold is supported by the sensitivity analysis in Appendix[K](https://arxiv.org/html/2609.13203#A11), where the atom estimates are insensitive to the tolerance and to the smoothing prior across the tested grid\.

A further interpretive limitation concerns sign\. Because the measures used in the main analyses \(ImmiI\_\{\\mathrm\{mmi\}\},IminI\_\{\\min\},IrrI\_\{\\mathrm\{rr\}\},I∧I\_\{\\wedge\}\) are non\-negative, they carry no sign: unlike partial correlation, ePID has no analogue of a negative or inhibitory association and cannot, on its own, distinguish facilitative from suppressive relations\. The signedI±I\_\{\\pm\}measure is the only exception, and for the reasons above it is not used in the main analyses\.

##### Directedness does not imply causality\.

Directed edges in the ePID framework are a computational convention required by PID’s source–target formulation, not causal claims\. The decomposition is not symmetric:PID⁡\(Xi→Xk\)\\mathrm\{PID\}\(X\_\{i\}\\to X\_\{k\}\)andPID⁡\(Xk→Xi\)\\mathrm\{PID\}\(X\_\{k\}\\to X\_\{i\}\)generally yield different atom profiles because the remainder embedding is constructed separately for each target\. This asymmetry reflects the information structure of the system \(how predictive each symptom is of a given target in the context of the remaining symptoms\) rather than causal direction\. Cross\-sectional observational data cannot distinguish causal influence from confounding, reverse causation, or shared latent causes\. The clinical hypotheses described in the introduction \(e\.g\., that insomnia can induce fatigue\) motivate the choice of symptom instrument and provide domain context, but they are not tested or validated by the present analyses\.

##### Next steps\.

Two extensions would most directly strengthen the ePID framework\. First, bootstrap or permutation\-based confidence intervals for PID atoms would enable formal statistical inference on atom differences across edges, targets, and cohorts, addressing the estimation uncertainty noted in Section[4\.4](https://arxiv.org/html/2609.13203#S4.SS4)\. Second, a systematic sensitivity analysis of category merging \(varying the number of ordinal levels from 2 to 4\) would quantify how discretisation affects atom estimates, particularly synergy\.

### Conclusion

ePID makes network\-context partial information decomposition tractable for psychometric instruments by compressing the multivariate remainder of a symptom network into a target\-directed discrete proxy, validated up toN=5N=5remainder sources\. A comprehensive benchmark, spanning calibrated synthetic data and 83 real\-world datasets and five PID measures, establishes that ACIB preserves the information structure PID quantifies while identifying measure\-specific failure modes, notably for the signedI±I\_\{\\pm\}\. Applied to two PHQ\-9 datasets and to the Interpersonal Reactivity Index, the method shows that decomposing conditional dependence into redundant and synergistic channels adds interpretive value beyond scalar association measures: the composition of dependence replicates across datasets within an instrument, the same pipeline discriminates redundancy\-dominated from synergy\-dominated instruments, and the redundancy/synergy distinction generates qualitatively different clinical hypotheses about symptom substitutability and interaction dependence\. Together, these results support embedding\-based PID as a practical complement to standard symptom\-network methodology\.

## Ethics

This research has been conducted using the UK Biobank Resource under Application Number 68746; UK Biobank holds generic ethical approval from the North West Multi\-centre Research Ethics Committee, and its participants provided written informed consent\. All other datasets, namely the Xinxiang student sample \(openly available; Su et al\. 2024\), the Interpersonal Reactivity Index from the Open\-Source Psychometrics Project, and the 83 datasets in the multi\-dataset benchmark \(Supplemental Table[S1](https://arxiv.org/html/2609.13203#A1.T1)\), are publicly available and de\-identified\.

## Data and Code Availability

UK Biobank data are available through the UK Biobank Access Management System \([https://www\.ukbiobank\.ac\.uk](https://www.ukbiobank.ac.uk/)\)\. The Xinxiang student sample is openly available on Zenodo \([https://zenodo\.org/records/10423537](https://zenodo.org/records/10423537); DOI 10\.5281/zenodo\.10423537\)\. The Interpersonal Reactivity Index is openly available from the Open\-Source Psychometrics Project \([https://openpsychometrics\.org](https://openpsychometrics.org/)\)\. Sources for the 83 multi\-dataset benchmark datasets are listed in Supplemental Table[S1](https://arxiv.org/html/2609.13203#A1.T1)\. Code implementing the ePID pipeline will be made available upon publication at[https://github\.com/CillianHourican/ePID](https://github.com/CillianHourican/ePID)\.

This appendix comprises the following sections: Appendix[A](https://arxiv.org/html/2609.13203#A1)describes the multi\-dataset benchmark design and participating datasets\. Appendix[B](https://arxiv.org/html/2609.13203#A2)provides detailed descriptions and hyperparameter settings for all 13 embedding methods\. Appendix[C](https://arxiv.org/html/2609.13203#A3)discusses the information\-bottleneck interpretation of the ACIB embedding and computational scaling properties\. Appendix[D](https://arxiv.org/html/2609.13203#A4)presents calibration statistics for the synthetic Bayesian Network ensemble\. Appendix[E](https://arxiv.org/html/2609.13203#A5)reports omnibus and post\-hoc statistical tests for the embedding benchmark\. Appendix[F](https://arxiv.org/html/2609.13203#A6)reports additional embedding\-fidelity diagnostics and symptom\-level PID examples\. Appendix[G](https://arxiv.org/html/2609.13203#A7)examines the prevalence of non\-linear dependencies in clinical datasets\. Appendix[H](https://arxiv.org/html/2609.13203#A8)provides cross\-cohort comparisons of information\-theoretic PHQ\-9 networks\. Appendix[J](https://arxiv.org/html/2609.13203#A10)presents an illustrative three\-variable example contrasting how partial correlation, conditional mutual information, and PID reveal progressively richer dependence structure\. Appendix[K](https://arxiv.org/html/2609.13203#A11)demonstrates that the reported decompositions are robust to the ACIB embedding hyperparameters and justifies embedding only a bounded remainder context\.

## Appendix AMulti\-Dataset Benchmark

This section describes the assembly and preprocessing of the 83 real\-world datasets used to evaluate embedding fidelity across diverse data structures and PID measures \(Section[2\.6](https://arxiv.org/html/2609.13203#S2.SS6)\)\.

### A\.1Benchmark dataset assembly and preprocessing

We assembled 83 datasets from publicly available repositories spanning clinical, epidemiological, and psychometric domains, including NHANES\([Centers for Disease Control and Prevention , National Center for Health Statistics \(NCHS\), 2020 CDC](https://arxiv.org/html/2609.13203#bib.bib18)\), SHARE\([Börsch\-Supan, 2022b](https://arxiv.org/html/2609.13203#bib.bib10)\), ELSA\([Banks et al\., 2023](https://arxiv.org/html/2609.13203#bib.bib4)\), HRS\([Health and Retirement Study, 2024](https://arxiv.org/html/2609.13203#bib.bib51)\), PROMIS\([Cella et al\., 2010](https://arxiv.org/html/2609.13203#bib.bib15)\), OpenPsychometrics, the UCI Machine Learning Repository, OSF, BRFSS\([Centers for Disease Control and Prevention , 2023 CDC](https://arxiv.org/html/2609.13203#bib.bib17)\), MEPS\([Agency for Healthcare Research and Quality , 2024 AHRQ](https://arxiv.org/html/2609.13203#bib.bib1)\), the Household Pulse Survey\([U\.S\. Census Bureau, 2025](https://arxiv.org/html/2609.13203#bib.bib108)\), ICPSR, Zenodo, Figshare, and Mendeley Data\. Datasets span depression, anxiety, PTSD, stress, personality, eating disorders, physical functioning, chronic conditions, sleep, and other clinical domains, with 5–56 variables per dataset and sample sizes ranging from 32 to 445,000 observations\. For each dataset, we selected only substantive item\-level or symptom\-level variables, excluding identifiers, demographic covariates, composite scores, and diagnostic labels\. Source\-specific missing\-value codes \(e\.g\., 7/77/777 and 9/99/999 in NHANES; negative codes in SHARE\) were recoded to missing\. For NHANES SAS transport files, the XPT format encodes zero as a tiny floating\-point number \(≈5\.4×10−79\{\\approx\}5\.4\\times 10^\{\-79\}\); these values were rounded to zero before further processing\.

Columns with more than 30% missing values were dropped, followed by rows with more than 50% missing values across the remaining columns\. Remaining missing values were imputed usingkk\-nearest neighbours \(k=5k=5\) via scikit\-learn’sKNNImputer\. For four datasets where this filtering left fewer than three columns or thirty rows \(nhanes\_phys\_func,covidistress\_28,ncsr\_depression\_18,nhanes\_alcohol\), listwise deletion \(complete\-case analysis\) was used instead, retaining all variables and only rows with no missing values\. After imputation, any genuinely constant columns \(zero variance\) were removed\.

For each dataset, two association matrices were estimated\. Partial correlations were computed from anℓ1\\ell\_\{1\}\-regularised precision matrix obtained viaGraphicalLassoCVwith 5\-fold cross\-validation, which estimates partial correlations as scaled off\-diagonal entries of the regularised precision matrixΣ−1\\Sigma^\{\-1\}, yielding a sparse estimate of conditional dependencies\. For mutual information estimation, all imputed values were first rounded to the nearest integer to recover ordinal response levels\. Variables with five or fewer unique integer values were retained at their original levels; variables with more than five unique values were discretised into eight equal\-frequency bins\. For partial information decomposition, all variables were further reduced to at mostkkdiscrete states by iteratively merging adjacent categories with the smallest combined frequency, wherekkwas the largest integer satisfyingk\(Nsources\+1\)≤N/5k^\{\(N\_\{\\mathrm\{sources\}\}\+1\)\}\\leq N/5, ensuring adequate coverage of the joint state space\. Pairwise mutual information was estimated using the discrete estimator from thenpeetlibrary, with diagonal entries set to the Shannon entropy of each variable\.

Table[S1](https://arxiv.org/html/2609.13203#A1.T1)lists all 83 datasets, their domain, number of selected variables, sample size, source repository, and citation\.

Table S1:Benchmark datasetsNote\.The 83 datasets used in the multi\-dataset embedding benchmark\.*Items*indicates the number of substantive variables selected before preprocessing;*N*is the raw sample size before row\-level missingness filtering\. easySHARE waves \(6a–6g\) are drawn from the same longitudinal study but are treated as separate networks because each wave has an independent cross\-section of respondents\.DatasetDomainItemsNNSourceCitationCSWS ItemsSelf\-worth35680OSF\([Briganti et al\., 2019](https://arxiv.org/html/2609.13203#bib.bib13)\)CSWS SubscalesSelf\-worth7680OSF\([Briganti et al\., 2019](https://arxiv.org/html/2609.13203#bib.bib13)\)SHARE W7Depression1777,000SHARE\([Börsch\-Supan, 2022b](https://arxiv.org/html/2609.13203#bib.bib10)\)ELSA W6Depression1710,000UK Data\([Banks et al\., 2023](https://arxiv.org/html/2609.13203#bib.bib4)\)HRS W13Sleep/pain/health1821,000HRS\([Health and Retirement Study, 2024](https://arxiv.org/html/2609.13203#bib.bib51)\)easySHARE 2004 \(W1\)Summary health830,000SHARE\([Börsch\-Supan, 2022a](https://arxiv.org/html/2609.13203#bib.bib9);[Gruber et al\., 2014](https://arxiv.org/html/2609.13203#bib.bib42)\)easySHARE 2007 \(W2\)Summary health937,000SHARE\([Börsch\-Supan, 2022a](https://arxiv.org/html/2609.13203#bib.bib9);[Gruber et al\., 2014](https://arxiv.org/html/2609.13203#bib.bib42)\)easySHARE 2011 \(W4\)Summary health858,000SHARE\([Börsch\-Supan, 2022a](https://arxiv.org/html/2609.13203#bib.bib9);[Gruber et al\., 2014](https://arxiv.org/html/2609.13203#bib.bib42)\)easySHARE 2013 \(W5\)Summary health866,000SHARE\([Börsch\-Supan, 2022a](https://arxiv.org/html/2609.13203#bib.bib9);[Gruber et al\., 2014](https://arxiv.org/html/2609.13203#bib.bib42)\)easySHARE 2015 \(W6\)Summary health968,000SHARE\([Börsch\-Supan, 2022a](https://arxiv.org/html/2609.13203#bib.bib9);[Gruber et al\., 2014](https://arxiv.org/html/2609.13203#bib.bib42)\)easySHARE 2017 \(W7\)Summary health677,000SHARE\([Börsch\-Supan, 2022a](https://arxiv.org/html/2609.13203#bib.bib9);[Gruber et al\., 2014](https://arxiv.org/html/2609.13203#bib.bib42)\)easySHARE 2020 \(W8\)Summary health947,000SHARE\([Börsch\-Supan, 2022a](https://arxiv.org/html/2609.13203#bib.bib9);[Gruber et al\., 2014](https://arxiv.org/html/2609.13203#bib.bib42)\)NHANES PHQ\-9Depression95,500CDC\([Centers for Disease Control and Prevention , National Center for Health Statistics \(NCHS\), 2020 CDC](https://arxiv.org/html/2609.13203#bib.bib18);[Kroenke et al\., 2001](https://arxiv.org/html/2609.13203#bib.bib68)\)NHANES SleepSleep106,200CDC\([Centers for Disease Control and Prevention , National Center for Health Statistics \(NCHS\), 2020 CDC](https://arxiv.org/html/2609.13203#bib.bib18)\)NHANES Phys\. Func\.Disability248,400CDC\([Centers for Disease Control and Prevention , National Center for Health Statistics \(NCHS\), 2020 CDC](https://arxiv.org/html/2609.13203#bib.bib18)\)NHANES Health StatusGeneral health88,400CDC\([Centers for Disease Control and Prevention , National Center for Health Statistics \(NCHS\), 2020 CDC](https://arxiv.org/html/2609.13203#bib.bib18)\)NHANES AlcoholAlcohol95,500CDC\([Centers for Disease Control and Prevention , National Center for Health Statistics \(NCHS\), 2020 CDC](https://arxiv.org/html/2609.13203#bib.bib18)\)NHANES Blood PressureCardiovascular106,200CDC\([Centers for Disease Control and Prevention , National Center for Health Statistics \(NCHS\), 2020 CDC](https://arxiv.org/html/2609.13203#bib.bib18)\)NHANES Med\. Cond\.Chronic conditions178,900CDC\([Centers for Disease Control and Prevention , National Center for Health Statistics \(NCHS\), 2020 CDC](https://arxiv.org/html/2609.13203#bib.bib18)\)NHANES SmokingSmoking106,700CDC\([Centers for Disease Control and Prevention , National Center for Health Statistics \(NCHS\), 2020 CDC](https://arxiv.org/html/2609.13203#bib.bib18)\)NHANES Phys\. ActivityActivity165,900CDC\([Centers for Disease Control and Prevention , National Center for Health Statistics \(NCHS\), 2020 CDC](https://arxiv.org/html/2609.13203#bib.bib18)\)NHANES Hospital Util\.Healthcare99,300CDC\([Centers for Disease Control and Prevention , National Center for Health Statistics \(NCHS\), 2020 CDC](https://arxiv.org/html/2609.13203#bib.bib18)\)DASS\-42Dep/Anx/Stress4239,800OpenPsych\([Lovibond and Lovibond, 1995](https://arxiv.org/html/2609.13203#bib.bib73)\)TMASTrait anxiety505,400OpenPsych\([Taylor, 1953](https://arxiv.org/html/2609.13203#bib.bib97)\)Big Five \(IPIP\)Personality5019,700OpenPsych\([Goldberg, 1999](https://arxiv.org/html/2609.13203#bib.bib35)\)RSESelf\-esteem1048,000OpenPsych\([Rosenberg, 1965](https://arxiv.org/html/2609.13203#bib.bib83)\)ECRAttachment3651,500OpenPsych\([Brennan et al\., 1998](https://arxiv.org/html/2609.13203#bib.bib12)\)HSNS\+DDNarcissism/dep1054,000OpenPsych\([Hendin and Cheek, 1997](https://arxiv.org/html/2609.13203#bib.bib52);[Jonason and Webster, 2010](https://arxiv.org/html/2609.13203#bib.bib60)\)KIMSMindfulness39601OpenPsych\([Baer et al\., 2004](https://arxiv.org/html/2609.13203#bib.bib3)\)SCSSexual compulsiv\.123,400OpenPsych\([Kalichman and Rompa, 1995](https://arxiv.org/html/2609.13203#bib.bib61)\)DermatologySkin disease34366UCI\([Guvenir and Demiroz, 1998](https://arxiv.org/html/2609.13203#bib.bib44);[Guvenir et al\., 1998](https://arxiv.org/html/2609.13203#bib.bib45)\)Primary TumorCancer features17339UCI\([UCI Machine Learning Repository, 1988](https://arxiv.org/html/2609.13203#bib.bib102)\)Early\-Stage DiabetesDiabetes symptoms16520UCI\([Islam et al\., 2020](https://arxiv.org/html/2609.13203#bib.bib57);[Islam et al\., 2019](https://arxiv.org/html/2609.13203#bib.bib56)\)Thyroid Cancer Recur\.Cancer clinical16383UCI\([Borzooei and Tarokhian, 2023](https://arxiv.org/html/2609.13203#bib.bib11)\)Autism Screen\. AdultASD screening21704UCI\([Thabtah, 2017](https://arxiv.org/html/2609.13203#bib.bib98);[Thabtah, 2018](https://arxiv.org/html/2609.13203#bib.bib99)\)Thoracic SurgerySurgical symptoms16470UCI\([Zieba et al\., 2014](https://arxiv.org/html/2609.13203#bib.bib120);[Zięba et al\., 2014](https://arxiv.org/html/2609.13203#bib.bib121)\)Mammographic MassBreast imaging5961UCI\([Elter, 2007](https://arxiv.org/html/2609.13203#bib.bib24);[Elter et al\., 2007](https://arxiv.org/html/2609.13203#bib.bib25)\)Chronic Kidney DiseaseKidney signs/labs24400UCI\([Rubini et al\., 2015](https://arxiv.org/html/2609.13203#bib.bib84)\)MI ComplicationsCardiac symptoms1111,700UCI\([Golovenkin et al\., 2020a](https://arxiv.org/html/2609.13203#bib.bib36);[Golovenkin et al\., 2020b](https://arxiv.org/html/2609.13203#bib.bib37)\)Glioma GradingCancer mutations25839UCI\([Tasci and Camphausen, 2022](https://arxiv.org/html/2609.13203#bib.bib96)\)Cervical Cancer RiskSTD/risk factors36858UCI\([Fernandes et al\., 2017](https://arxiv.org/html/2609.13203#bib.bib26)\)CDC Diabetes Indic\.Health indicators21253,700UCI/CDC\([Centers for Disease Control and Prevention, 2017](https://arxiv.org/html/2609.13203#bib.bib16)\)AIDS Clinical TrialsHIV clinical252,139UCI\([Hammer and Katzenstein, 1996](https://arxiv.org/html/2609.13203#bib.bib47);[Hammer et al\., 1996](https://arxiv.org/html/2609.13203#bib.bib48)\)MesotheliomaCancer symptoms34324UCI\([Tanrikulu and Er, 2012](https://arxiv.org/html/2609.13203#bib.bib95)\)NPHA Healthy AgingHealth self\-report15714UCI\([University of Michigan Institute for Healthcare Policy and Innovation, 2023](https://arxiv.org/html/2609.13203#bib.bib103)\)HCV Egyptian PatientsLiver symptoms281,385UCI\([Kamal and ElEleimy, 2017](https://arxiv.org/html/2609.13203#bib.bib62)\)Xinxiang PHQ\-9Depression924,292Zenodo\([Su et al\., 2024a](https://arxiv.org/html/2609.13203#bib.bib91);[Su et al\., 2024b](https://arxiv.org/html/2609.13203#bib.bib92)\)Xinxiang GAD\-7Anxiety724,292Zenodo\([Su et al\., 2024a](https://arxiv.org/html/2609.13203#bib.bib91);[Su et al\., 2024b](https://arxiv.org/html/2609.13203#bib.bib92)\)Xinxiang ISIInsomnia724,292Zenodo\([Su et al\., 2024a](https://arxiv.org/html/2609.13203#bib.bib91);[Su et al\., 2024b](https://arxiv.org/html/2609.13203#bib.bib92)\)Xinxiang PSSStress1024,292Zenodo\([Su et al\., 2024a](https://arxiv.org/html/2609.13203#bib.bib91);[Su et al\., 2024b](https://arxiv.org/html/2609.13203#bib.bib92)\)PHQ\+GAD\+ESS MexicoDep/Anx/Sleep33783Figshare\([Unknown, 2015](https://arxiv.org/html/2609.13203#bib.bib104)\)MHP StudentsAnx/Stress/Dep262,000Figshare\([Syeed, 2024](https://arxiv.org/html/2609.13203#bib.bib94)\)Bangladesh MHMulti\-scale dep45502Mendeley\([Unknown, 2023](https://arxiv.org/html/2609.13203#bib.bib105)\)PHQ\-9 StudentsDepression9682Mendeley\([Unknown, 2024](https://arxiv.org/html/2609.13203#bib.bib106)\)Colombia Well\-BeingStress/Anx/Dep263,000Mendeley\([Martínez et al\., 2024b](https://arxiv.org/html/2609.13203#bib.bib76);[Martínez et al\., 2024a](https://arxiv.org/html/2609.13203#bib.bib75)\)McNally 2014PTSD17344OSF\([McNally et al\., 2015](https://arxiv.org/html/2609.13203#bib.bib77)\)ED NetworkEating disorders22245OSF\([Vervaet et al\., 2021](https://arxiv.org/html/2609.13203#bib.bib113)\)IRI EmpathyEmpathy281,973OSF\([Davis, 1983](https://arxiv.org/html/2609.13203#bib.bib21)\)CPS ChinesePurpose in life12598OSF\([Wu et al\., 2024a](https://arxiv.org/html/2609.13203#bib.bib116);[Wu et al\., 2024b](https://arxiv.org/html/2609.13203#bib.bib117)\)COVIDiSTRESSStress/Loneliness28173,000OSF\([Yamada et al\., 2021](https://arxiv.org/html/2609.13203#bib.bib118)\)PROMIS AnxietyAnxiety56817Harvard DV\([Cella et al\., 2010](https://arxiv.org/html/2609.13203#bib.bib15)\)PROMIS DepressionDepression56811Harvard DV\([Cella et al\., 2010](https://arxiv.org/html/2609.13203#bib.bib15)\)PROMIS AngerAnger56918Harvard DV\([Cella et al\., 2010](https://arxiv.org/html/2609.13203#bib.bib15)\)PROMIS Fatigue \(Exp\)Fatigue56821Harvard DV\([Cella et al\., 2010](https://arxiv.org/html/2609.13203#bib.bib15)\)PROMIS Fatigue \(Imp\)Fatigue56820Harvard DV\([Cella et al\., 2010](https://arxiv.org/html/2609.13203#bib.bib15)\)PROMIS Pain Interf\.Pain56866Harvard DV\([Cella et al\., 2010](https://arxiv.org/html/2609.13203#bib.bib15)\)PROMIS Pain QualityPain56862Harvard DV\([Cella et al\., 2010](https://arxiv.org/html/2609.13203#bib.bib15)\)PROMIS Pain BehaviorPain56859Harvard DV\([Cella et al\., 2010](https://arxiv.org/html/2609.13203#bib.bib15)\)PROMIS PhysFun APhysical function56812Harvard DV\([Cella et al\., 2010](https://arxiv.org/html/2609.13203#bib.bib15)\)PROMIS PhysFun BPhysical function56814Harvard DV\([Cella et al\., 2010](https://arxiv.org/html/2609.13203#bib.bib15)\)PROMIS Social PerfSocial function56864Harvard DV\([Cella et al\., 2010](https://arxiv.org/html/2609.13203#bib.bib15)\)PROMIS Social SatSocial function56851Harvard DV\([Cella et al\., 2010](https://arxiv.org/html/2609.13203#bib.bib15)\)PROMIS AlcoholAlcohol56903Harvard DV\([Cella et al\., 2010](https://arxiv.org/html/2609.13203#bib.bib15)\)PROMIS Profile 29Multi\-domain294,500Harvard DV\([Hays et al\., 2018](https://arxiv.org/html/2609.13203#bib.bib50)\)BRFSS 2022Health indicators25445,000CDC\([Centers for Disease Control and Prevention , 2023 CDC](https://arxiv.org/html/2609.13203#bib.bib17)\)PHQ\+GAD\+ISI\+PSS comb\.All 4 scales3324,292Zenodo\([Su et al\., 2024a](https://arxiv.org/html/2609.13203#bib.bib91);[Su et al\., 2024b](https://arxiv.org/html/2609.13203#bib.bib92)\)NCS\-R DepressionMDE symptoms182,300ICPSR\([Kessler and Merikangas, 2004](https://arxiv.org/html/2609.13203#bib.bib63);[Alegria et al\., 2007](https://arxiv.org/html/2609.13203#bib.bib2)\)MIDUS 2 Phys\. Symp\.Somatic symptoms104,000ICPSR\([Ryff et al\., 2019](https://arxiv.org/html/2609.13203#bib.bib85)\)MIDUS 2 Chronic Cond\.Multimorbidity204,000ICPSR\([Ryff et al\., 2019](https://arxiv.org/html/2609.13203#bib.bib85)\)MIDUS 2 Health LimitsFunctional limits104,000ICPSR\([Ryff et al\., 2019](https://arxiv.org/html/2609.13203#bib.bib85)\)MEPS 2022 SF\-12Health status1711,300AHRQ\([Agency for Healthcare Research and Quality , 2024 AHRQ](https://arxiv.org/html/2609.13203#bib.bib1)\)MEPS 2022 ConditionsChronic conditions1722,300AHRQ\([Agency for Healthcare Research and Quality , 2024 AHRQ](https://arxiv.org/html/2609.13203#bib.bib1)\)HPS Dec 2024PHQ\-2\+GAD\-2\+Disab\.149,400Census\([U\.S\. Census Bureau, 2025](https://arxiv.org/html/2609.13203#bib.bib108)\)Per\-dataset variable selections, exclusion criteria, and recoding rules are available in the project’s data documentation repository\.

## Appendix BEmbedding Methods: Descriptions and Hyperparameters

To compress the multivariate remainder𝐙i→k=𝐗∖\{Xi,Xk\}\\mathbf\{Z\}\_\{i\\to k\}=\\mathbf\{X\}\\setminus\\\{X\_\{i\},X\_\{k\}\\\}into a single discrete embedding variableEi→kE\_\{i\\to k\}for two\-source PID computation, we benchmarked 13 embedding methods spanning five families\. Table[S2](https://arxiv.org/html/2609.13203#A2.T2)summarises the method families, software implementations, and key hyperparameters used throughout the benchmark\. All methods were applied with a global random seed of 42 for reproducibility\. Below, we describe each method and its hyperparameter settings\.

Table S2:Summary of the 13 embedding methods, grouped by family\. All methods share a global random seed of 42\. “KK” denotes the number of clusters or output cardinality; “kk” denotes the number of selected features; “BB” denotes the number of quantile bins per dimension\.MethodFamilyPackage/LibraryKey HyperparametersACIBInfo\-theoreticCustom \(npeet\)Kmax=12K\_\{\\max\}=12,Kmin=2K\_\{\\min\}=2,α=0\.5\\alpha=0\.5,loss tol\.=5%=5\\%, max\_iter=40=40Greedy CMIInfo\-theoreticCustom \(npeet\)max\_vars=4=4, loss tol\.=5%=5\\%JMIMI\-based selectionscikit\-featurek=3k=3CMIMMI\-based selectionscikit\-featurek=3k=3ReliefFMI\-based selectionskrebatek=3k=3,n​\_​neighbours=100n\\\_\\text\{neighbours\}=100kk\-modesClusteringkmodesK=6K=6, init==Huang,n​\_​init=5n\\\_\\text\{init\}=5Spectral HammingClusteringscikit\-learnK=6K=6,γ=5\.0\\gamma=5\.0, labels==kk\-meansAgglom\. HammingClusteringscikit\-learnK=6K=6, linkage==averageMCA\+\+kk\-meansDim\. red\.\+\+clust\.prince, scikit\-learnn​\_​comp=2n\\\_\\text\{comp\}=2,K=6K=6,n​\_​init=10n\\\_\\text\{init\}=10SVD\+\+kk\-meansDim\. red\.\+\+clust\.scikit\-learnn​\_​comp=2n\\\_\\text\{comp\}=2,K=6K=6,n​\_​init=10n\\\_\\text\{init\}=10NMF\+\+kk\-meansDim\. red\.\+\+clust\.scikit\-learnn​\_​comp=2n\\\_\\text\{comp\}=2,K=6K=6, init==nndsvda,n​\_​init=10n\\\_\\text\{init\}=10SIR\+\+binningSupervised SDRslicedn​\_​dir=2n\\\_\\text\{dir\}=2,B=4B=4PLS\+\+binningSupervised SDRscikit\-learnn​\_​comp=2n\\\_\\text\{comp\}=2,B=4B=4### B\.1Information\-theoretic methods

##### Agglomerative Conditional Information Bottleneck \(ACIB\)\.

ACIB is a two\-phase, deterministic, target\-directed compression algorithm that maps the joint states of the multivariate remainder𝐙i→k\\mathbf\{Z\}\_\{i\\to k\}into a low\-cardinality discrete variableEi→kE\_\{i\\to k\}while preserving as much information as possible about the targetXkX\_\{k\}\. In the first phase, akk\-CIB \(Conditional Information Bottleneck\) initialisation is performed\. The unique joint states of𝐙i→k\\mathbf\{Z\}\_\{i\\to k\}are assigned toKmax=12K\_\{\\max\}=12initial clusters using a farthest\-first seeding procedure based on the Jensen–Shannon \(JS\) divergence of the conditional target distributionsP⁡\(Xk∣𝐙i→k=z,Xi=s\)P\(X\_\{k\}\\mid\\mathbf\{Z\}\_\{i\\to k\}=z,X\_\{i\}=s\)across source statesss\. Cluster assignments are then refined by iteratively reassigning each joint state to the cluster whose prototype minimises the Kullback–Leibler divergence, with conditional distributions smoothed via Laplace smoothing \(α=0\.5\\alpha=0\.5\) to regularise sparse cells\. This alternating assignment–update loop runs for up to 40 iterations or until assignments stabilise\.

In the second phase, a greedy agglomerative merging procedure reduces the number of clusters belowKmaxK\_\{\\max\}\. At each step, the pair of clusters whose merger incurs the smallest JS divergence loss is identified and merged\. After each merge, the relative information loss is evaluated as\|I⁡\(Xi;Xk∣𝐙\)−I⁡\(Xi;Xk∣E\)\|/I⁡\(Xi;Xk∣𝐙\)\|I\(X\_\{i\};X\_\{k\}\\mid\\mathbf\{Z\}\)\-I\(X\_\{i\};X\_\{k\}\\mid E\)\|\\,/\\,I\(X\_\{i\};X\_\{k\}\\mid\\mathbf\{Z\}\), whereI\(⋅;⋅∣⋅\)I\(\\cdot;\\cdot\\mid\\cdot\)denotes conditional mutual information estimated via plug\-in discrete estimators\([Ver Steeg and Galstyan, 2014](https://arxiv.org/html/2609.13203#bib.bib112), npeet;\)\. Merging continues as long as the relative loss remains below a tolerance of5%5\\%, and stops at a minimum ofKmin=2K\_\{\\min\}=2clusters\. The result is a deterministic mappingfi→kf\_\{i\\to k\}from joint remainder states to a compact discrete code whose cardinality adapts to the complexity of each edge\.

##### Greedy CMI subset selection\.

Greedy CMI is a forward feature\-selection wrapper that selects a small subset of the remainder variables whose joint state preserves the conditional mutual informationI⁡\(Xi;Xk∣𝐙\)I\(X\_\{i\};X\_\{k\}\\mid\\mathbf\{Z\}\)\. Starting from an empty set, it greedily adds the variable that minimises the relative information loss at each step, stopping when the loss falls below5%5\\%or when a maximum of 4 variables have been selected\. The embedding is the joint categorical code of the selected variables\. This method directly targets information preservation without intermediate dimensionality reduction\.

### B\.2MI\-based feature selection methods

##### Joint Mutual Information \(JMI\)\.

JMI\([Yang and Moody, 1999](https://arxiv.org/html/2609.13203#bib.bib119)\)is a filter\-based feature selection criterion that scores each candidate feature by its joint mutual information with the target, accounting for redundancy among already\-selected features\. We use the implementation from the scikit\-feature library\([Li et al\., 2018](https://arxiv.org/html/2609.13203#bib.bib71)\)withk=3k=3features selected\. The embedding is the joint categorical code of the three selected remainder variables\.

##### Conditional Mutual Information Maximisation \(CMIM\)\.

CMIM\([Fleuret, 2004](https://arxiv.org/html/2609.13203#bib.bib29)\)selects features by maximising the minimum conditional mutual information between each candidate feature and the target, given every previously selected feature\. This criterion is more conservative than JMI, as it penalises redundancy more aggressively\. We use the scikit\-feature implementation withk=3k=3selected features, and the embedding is again the joint categorical code\.

##### ReliefF\.

ReliefF\([Kononenko, 1994](https://arxiv.org/html/2609.13203#bib.bib67)\)is an instance\-based feature weighting algorithm that evaluates features by their ability to distinguish between instances from different classes and instances from the same class, using nearest\-neighbour distances\. We use the skrebate implementation\([Urbanowicz et al\., 2018](https://arxiv.org/html/2609.13203#bib.bib107)\)with 100 neighbours andk=3k=3top\-ranked features\. The embedding is the joint categorical code of the selected variables\.

### B\.3Clustering methods

##### kk\-modes\.

Thekk\-modes algorithm\([Huang, 1998](https://arxiv.org/html/2609.13203#bib.bib54)\)is a categorical analogue ofkk\-means that uses the Hamming distance and a frequency\-based mode update rule\. We use the kmodes Python package with Huang initialisation,n​\_​init=5n\\\_\\text\{init\}=5random restarts, andK=6K=6clusters\. Cluster labels serve directly as the discrete embedding\.

##### Spectral clustering on Hamming affinity\.

Spectral clustering\([Ng et al\., 2001](https://arxiv.org/html/2609.13203#bib.bib79)\)is applied to a Hamming\-based affinity matrix\. The pairwise Hamming distance matrix is converted to an affinity matrix via a Gaussian kernelAi​j=exp\(−γ⋅dH\(zi,zj\)\)A\_\{ij\}=\\exp\(\-\\gamma\\cdot d\_\{H\}\(z\_\{i\},z\_\{j\}\)\)with bandwidthγ=5\.0\\gamma=5\.0\. Spectral clustering withK=6K=6clusters andkk\-means label assignment is then performed on the Laplacian eigenvectors using scikit\-learn\([Pedregosa et al\., 2011](https://arxiv.org/html/2609.13203#bib.bib81)\)\. To manage memory, rows are subsampled to 5,000 when the dataset is larger; remaining rows are assigned to the cluster of their nearest subsampled neighbour by Hamming distance\.

##### Agglomerative clustering on Hamming distance\.

Agglomerative \(hierarchical\) clustering with average linkage is applied to the pairwise Hamming distance matrix using scikit\-learn, withK=6K=6clusters\. As with spectral clustering, datasets exceeding 5,000 rows are subsampled and out\-of\-sample rows are assigned via nearest\-neighbour lookup\.

### B\.4Dimensionality reduction followed by clustering

##### MCA\+\+kk\-means\.

Multiple Correspondence Analysis \(MCA\) is a dimensionality reduction technique for categorical data that generalises PCA to indicator matrices\([Greenacre, 2017](https://arxiv.org/html/2609.13203#bib.bib41)\)\. We extract the first two MCA components using the prince Python library\([Halford, 2023](https://arxiv.org/html/2609.13203#bib.bib46)\), then applykk\-means clustering \(K=6K=6,n​\_​init=10n\\\_\\text\{init\}=10\) to the two\-dimensional scores\. Cluster labels form the discrete embedding\.

##### Truncated SVD\+\+kk\-means\.

The remainder variables are one\-hot encoded and projected onto two components via truncated singular value decomposition \(SVD\) using scikit\-learn\. The two\-dimensional projections are then clustered withkk\-means \(K=6K=6,n​\_​init=10n\\\_\\text\{init\}=10\)\. This approach treats categorical data as sparse binary features before linear dimensionality reduction\.

##### NMF\+\+kk\-means\.

Non\-negative Matrix Factorisation \(NMF\) decomposes the one\-hot encoded remainder matrix into two non\-negative low\-rank factors\([Lee and Seung, 1999](https://arxiv.org/html/2609.13203#bib.bib70)\)\. We extract two components using thenndsvdainitialisation in scikit\-learn, then applykk\-means \(K=6K=6,n​\_​init=10n\\\_\\text\{init\}=10\) to the coefficient matrix\. The non\-negativity constraint can produce parts\-based representations that may capture interpretable response patterns\.

### B\.5Supervised sufficient dimension reduction followed by binning

##### Sliced Inverse Regression \(SIR\)\+\+binning\.

SIR\([Li, 1991](https://arxiv.org/html/2609.13203#bib.bib72)\)estimates the central dimension\-reduction subspace by exploiting the inverse regression of predictors on the response\. Using the sliced Python package\([Koepke, 2018](https://arxiv.org/html/2609.13203#bib.bib65)\), we project the one\-hot encoded remainder onto two SIR directions with respect to the targetXkX\_\{k\}\. The two\-dimensional projections are discretised via aB×B=4×4B\\times B=4\\times 4quantile grid, yielding up to 16 discrete categories\. Because SIR uses the target for projection, this method produces a target\-directed embedding\.

##### Partial Least Squares \(PLS\)\+\+binning\.

PLS regression\([Wold et al\., 1984](https://arxiv.org/html/2609.13203#bib.bib115)\)finds latent components that maximise the covariance between the \(one\-hot encoded\) remainder and the target variable\. We extract two PLS components using scikit\-learn’sPLSRegressionand discretise the scores with a4×44\\times 4quantile grid, producing up to 16 categories\. Like SIR, PLS is target\-directed, making the embedding tailored to information aboutXkX\_\{k\}\.

### B\.6Computational considerations

Two practical decisions affect the computational footprint of the embedding benchmark\. First, for the two distance\-based methods—spectral clustering on Hamming affinity and agglomerative clustering on Hamming distance—the pairwise distance matrix requiresO⁡\(n2\)O\(n^\{2\}\)memory\. For datasets exceeding 5,000 rows, we subsample to 5,000 observations \(drawn without replacement using the global random seed\) and assign out\-of\-sample observations to the cluster of their nearest subsampled neighbour\. This cap ensures that then×nn\\times ndistance matrix fits in RAM while retaining the full dataset size for all other methods\.

Second, all methods that involve random initialisation \(e\.g\.,kk\-means,kk\-modes, spectral label assignment\) use a fixed random seed of 42 to ensure reproducibility across runs and computing environments\. The ACIB and Greedy CMI methods are deterministic given the data, so the seed affects only the subsampling step for distance\-based methods and the initialisation of downstream clustering routines\.

##### Amalgamation of the multivariate PID lattice\.

This appendix gives the full antichain classification rule summarised in the main text \(Section[2\.5\.3](https://arxiv.org/html/2609.13203#S2.SS5.SSS3)\)\. Each multivariate atom corresponds to an antichain: a collection of source subsets, none of which contains another\. We first label each subset according to whether it contains onlyXiX\_\{i\}\(label “A”\), only non\-focal sources \(label “B”\), or bothXiX\_\{i\}and at least one other source \(label “AB”\)\. We then apply a conservative classification rule: if any A or B label is present in the antichain, all AB labels are discarded, and the atom is assigned based on the remaining labels:

- •source\-unique if only A labels remain;
- •remainder\-unique if only B labels remain;
- •redundancy if both A and B labels remain;
- •synergy only if the antichain contained exclusively AB labels \(i\.e\., no pure A or B subset existed\)\.

This rule is conservative with respect to synergy: an atom is classified as synergistic only when every subset in its antichain necessarily spans both the focal source and the remainder group, with no subset attributable to either alone\. When a pure\-A or pure\-B subset coexists with an AB subset, the atom is attributed to the group that can account for it without invoking cross\-group interaction\. The grouping preserves the two\-source consistency relations: the amalgamated source\-unique plus redundancy equalsI⁡\(Xi,Xt\)I\(X\_\{i\};X\_\{t\}\), and the amalgamated remainder\-unique plus redundancy equalsI⁡\(\{Xj:j≠i\},Xt\)I\(\\\{X\_\{j\}:j\\neq i\\\};X\_\{t\}\)\.

##### Software\.

PID was computed using theditPython library\([James et al\., 2018](https://arxiv.org/html/2609.13203#bib.bib58)\), with theImmiI\_\{\\text\{mmi\}\}\(minimum mutual information\) redundancy measure as the primary PID specification\. As robustness checks we additionally computed decompositions with theIminI\_\{\\text\{min\}\}\(Williams–Beer\)\([Williams and Beer, 2010](https://arxiv.org/html/2609.13203#bib.bib114)\),I±I\_\{\\pm\}\(Finn–Lizier\)\([Finn and Lizier, 2018](https://arxiv.org/html/2609.13203#bib.bib28)\), andI∧I\_\{\\wedge\}\(Gács–Körner\)\([Gács and Körner, 1973](https://arxiv.org/html/2609.13203#bib.bib34)\)measures\. ForN≥5N\\geq 5sources,I±I\_\{\\pm\}andI∧I\_\{\\wedge\}were computed via the fast Möbius transform using precomputed lattice data from thealgebraicPIDpackage\([Jansma et al\., 2025](https://arxiv.org/html/2609.13203#bib.bib59)\)\. Mutual information was estimated using plug\-in frequency\-based estimators vianpeet\([Ver Steeg and Galstyan, 2014](https://arxiv.org/html/2609.13203#bib.bib112)\); conditional mutual information for the PHQ\-9 analyses was computed usingpyitlib\. Rank\-based partial correlations for the PHQ\-9 analyses were computed using thepingouinpackage\([Vallat, 2018](https://arxiv.org/html/2609.13203#bib.bib109)\)\.

### B\.7Normalisation and effect\-size scaling

Because information\-theoretic measures are expressed in bits and are bounded above by the uncertainty of the target, we report normalised effect sizes to facilitate comparisons across targets and across cohorts\.

For conditional mutual information \(CMI\), we report the target\-normalised quantity

nCMIX→Y=I⁡\(X;Y∣𝐙\)H⁡\(Y\)∈\[0,1\],\\mathrm\{nCMI\}\_\{X\\rightarrow Y\}\\;=\\;\\frac\{I\(X;Y\\mid\\mathbf\{Z\}\)\}\{H\(Y\)\}\\in\[0,1\],\(5\)where𝐙\\mathbf\{Z\}denotes the remaining symptoms\.

For embedding\-based PID atoms about a given targetYY, we report shares of the target entropy,

πX→Y\(⋅\)=Atom\(⋅\)​\(X,EX→Y,Y\)H⁡\(Y\)∈\[0,1\],\\pi^\{\(\\cdot\)\}\_\{X\\rightarrow Y\}\\;=\\;\\frac\{\\mathrm\{Atom\}^\{\(\\cdot\)\}\(X,E\_\{X\\rightarrow Y\};Y\)\}\{H\(Y\)\}\\in\[0,1\],\(6\)whereAtom\(⋅\)\\mathrm\{Atom\}^\{\(\\cdot\)\}denotes source\-unique, remainder\-unique, redundant, or synergistic information from the two\-source decomposition on\(X,EX→Y,Y\)\(X,E\_\{X\\rightarrow Y\};Y\), andEX→YE\_\{X\\rightarrow Y\}is the embedding of the remaining symptoms\. Both quantities lie in\[0,1\]\[0,1\]and read as the fraction of the target’s marginal uncertainty explained via a given information channel\.

PHQ\-9 items are ordinal and empirically skewed toward low\-severity response categories, which reduces their effective entropy and constrains the attainable magnitude of mutual information and related quantities; this motivates reporting normalised effect sizes rather than raw bits\. ThenCMI\\mathrm\{nCMI\}normalisation \(Eq\.[5](https://arxiv.org/html/2609.13203#A2.E5)\) is conservative becauseH⁡\(Y\)≥H⁡\(Y∣𝐙\)H\(Y\)\\geq H\(Y\\mid\\mathbf\{Z\}\), and it allowsnCMI\\mathrm\{nCMI\}to be interpreted as the fraction of the target’s marginal uncertainty associated with the source after conditioning on the rest of the symptom set\. The PID shares \(Eq\.[6](https://arxiv.org/html/2609.13203#A2.E6)\) read as “the fraction ofH⁡\(Y\)H\(Y\)explained via a given information channel” and are reported alongside unnormalised totals in bits when needed\.

##### Visualisation pruning\.

Statistical significance alone can yield dense networks in large samples \(especially in UK Biobank\)\. To obtain readable network figures and to enable direct cross\-cohort comparison, we apply a two\-stage rule: \(i\) edge screening by CMI significance, followed by \(ii\) rank\-based pruning for visualisation by retaining, for each target, the incoming edges with the largest normalised effect sizenCMIX→Y\\mathrm\{nCMI\}\_\{X\\rightarrow Y\}\(Eq\.[5](https://arxiv.org/html/2609.13203#A2.E5)\)\. The number of retained edges per target is specified in the figure captions\.

## Appendix CEmbedding Background and Computational Scaling

##### Relation to information bottleneck and redundancy bottleneck\.

In the main analyses with real PHQ\-9 data, we use the Agglomerative Conditional Information Bottleneck \(ACIB\) embedding, which is target\-directed: it iteratively merges states of the multivariate remainder to minimise information loss aboutXkX\_\{k\}under the fixed\-cardinality constraint\. ACIB is estimated separately for each\(Xi,Xk\)\(X\_\{i\},X\_\{k\}\)pair using𝐗∖\{Xi,Xk\}\\mathbf\{X\}\\setminus\\\{X\_\{i\},X\_\{k\}\\\}as inputs andXkX\_\{k\}as the supervised target\. ACIB can be viewed as an instance of the information bottleneck principle\([Tishby et al\., 1999](https://arxiv.org/html/2609.13203#bib.bib101)\): it constructs a compressed discrete representation of the multivariate remainderEi→k=fi→k​\(𝐗∖\{Xi,Xk\}\)E\_\{i\\to k\}=f\_\{i\\to k\}\(\\mathbf\{X\}\\setminus\\\{X\_\{i\},X\_\{k\}\\\}\)that preserves information about the targetXkX\_\{k\}\. More precisely, ACIB solves a discrete variant of the information bottleneck: it seeks to minimiseI⁡\(Zi→k,Ei→k\)I\(Z\_\{i\\to k\};E\_\{i\\to k\}\)\(compression\) subject to the constraintI⁡\(Ei→k,Xk\)≥\(1−δ\)⋅I⁡\(Zi→k,Xk\)I\(E\_\{i\\to k\};X\_\{k\}\)\\geq\(1\-\\delta\)\\cdot I\(Z\_\{i\\to k\};X\_\{k\}\), whereδ\\deltais a loss tolerance set to 5% in our analyses\. The greedy merging phase traverses the IB rate–distortion curve from high cardinality \(many clusters\) to low cardinality, stopping when further compression would violate the information\-preservation constraint\. The 5% loss tolerance corresponds to operating near the “knee” of the IB curve, the regime where additional compression incurs disproportionately large information loss about the target\. This connection provides principled guidance for the tolerance hyperparameter: it should be set small enough to preserve target\-relevant structure while allowing sufficient compression for tractable PID computation\. Recent theoretical work has further shown that a principled notion of PID redundancy can itself be formulated as a bottleneck trade\-off via the*redundancy bottleneck*, yielding redundancy–compression curves and efficient optimisation procedures\([Kolchinsky, 2024](https://arxiv.org/html/2609.13203#bib.bib66)\)\. Whereas redundancy\-bottleneck formulations provide a decision\-theoretic grounding for redundancy, we use bottleneck\-style compression as a practical approximation step that enables scalable two\-source PID \(including synergy\) in high\-dimensional symptom systems\. BecauseEi→kE\_\{i\\to k\}is learned without access to the focal sourceXiX\_\{i\}, a theoretical failure mode arises when the multivariate remainder is informative aboutXkX\_\{k\}primarily through its interaction withXiX\_\{i\}\(i\.e\., interaction\-only motifs\); in such settings, any lossy compression of the remainder can attenuate synergy estimates\. In the present application, symptoms exhibit broad marginal dependence and do not resemble near\-deterministic XOR\-like constructions in which marginal associations vanish, reducing the plausibility of this edge case for PHQ\-9 networks\.

##### Computational scaling\.

A fullnn\-source PID requires bookkeeping over a lattice whose number of atoms grows according to the Dedekind numbers, making exact multivariate decompositions infeasible beyond smallnn\. By contrast, our approximation replaces the multivariate remainder by a low\-cardinality discrete embedding and then computes a two\-source PID on\(Xi,Ei→k,Xk\)\(X\_\{i\},E\_\{i\\to k\};X\_\{k\}\)\. For fixed item cardinalityKKand embedding size\|Ei→k\|=K\|E\_\{i\\to k\}\|=K, the PID step depends only on aK3K^\{3\}contingency table, while the dominant cost lies in constructingEi→kE\_\{i\\to k\}for each ordered pair\(i,k\)\(i,k\)\. This yields a scalableO⁡\(p2\)O\(p^\{2\}\)pipeline in the number of symptom pairs, with the per\-pair embedding cost determined by the number of observed joint states of the remainder\.

### C\.1Toy SCM: methods comparison under BROJA

Figure[S1](https://arxiv.org/html/2609.13203#A3.F1)recomputes the bottom\-row ePID pipeline of main\-text Figure[1](https://arxiv.org/html/2609.13203#S1.F1)under the BROJA redundancy measure\([Bertschinger et al\., 2014](https://arxiv.org/html/2609.13203#bib.bib6)\)on the identical SCM\. The top\-row panels \(true graph, PC, CMI\) are measure\-independent and therefore reproduced unchanged from the main\-text figure\. The redundancy lattice underlying the type\-1 amalgamation, alongside the truth\-versus\-ePID recovery for the focal sourceX1X\_\{1\}, is shown in main\-text Figure[3](https://arxiv.org/html/2609.13203#S2.F3)\.

Figure S1:Methods comparison under BROJA
Note\.Companion to main\-text Figure[1](https://arxiv.org/html/2609.13203#S1.F1), identical structural model and pipeline, with the bottom\-row two\-source PIDs computed under the BROJA redundancy measure\([Bertschinger et al\., 2014](https://arxiv.org/html/2609.13203#bib.bib6)\)instead ofIminI\_\{\\min\}\([Williams and Beer, 2010](https://arxiv.org/html/2609.13203#bib.bib114)\)\. The top row \(panelsa–c: true graph, PC, CMI\) is measure\-independent\. On the homogeneous source pair\(X1,X2\)\(X\_\{1\},X\_\{2\}\), both source variables live in thehh\-channel ofX3X\_\{3\}and the two measures agree to several decimals\. On heterogeneous source pairs\(Xa,Xb\)\(X\_\{a\},X\_\{b\}\)where one source lives in thehh\-channel and the other in theWW\-channel ofX3X\_\{3\},IminI\_\{\\min\}allocates information across redundancy and synergy while BROJA assigns the same information to pure uniques: a known divergence between the two measures that motivates BROJA’s operationalist definition\. BROJA==Bertschinger–Rauh–Olbrich–Jost–Ay; PC==partial correlation; CMI==conditional mutual information\.

## Appendix DBayesian Network Calibration Statistics

The following figures report calibration diagnostics for the ensemble of 50 synthetic Bayesian networks used in the embedding benchmark \(Section[2\.5\.3](https://arxiv.org/html/2609.13203#S2.SS5.SSS3)\)\. We compare the mutual information distributions of the synthetic networks against the empirical UK Biobank data to verify that the synthetic data preserve the statistical properties relevant to PID computation\.

![Refer to caption](https://arxiv.org/html/2609.13203v1/embedding_figures/BN_Calibration/mi_comparison_1_histogram.png)\(a\)MI Distribution Comparison
![Refer to caption](https://arxiv.org/html/2609.13203v1/embedding_figures/BN_Calibration/mi_comparison_2_cdf.png)\(b\)CDF Comparison \(KS stat: 0\.166\)

Figure S2:Mutual information distribution comparison between real data and synthetic Bayesian network ensemble
Note\.\(a\) Histogram overlay showing density distributions of MI values\. \(b\) Cumulative distribution functions \(CDFs\) with Kolmogorov–Smirnov \(KS\) statistic quantifying the distributional match\. The ensemble of 50 synthetic BNs closely matches the real\-data distribution, validating the calibration procedure\.![Refer to caption](https://arxiv.org/html/2609.13203v1/embedding_figures/BN_Calibration/mi_comparison_3_qqplot.png)\(a\)Q\-Q Plot
![Refer to caption](https://arxiv.org/html/2609.13203v1/embedding_figures/BN_Calibration/ensemble_variability_1_mi_mean_per_bn.png)\(b\)Ensemble Variability

Figure S3:Detailed validation of Bayesian network ensemble calibration
Note\.\(a\) Quantile–quantile \(Q–Q\) plot comparing MI quantiles between real data and synthetic BNs\. Points close to the diagonal indicate good distributional match\. \(b\) Mean MI per network across the ensemble, showing consistent calibration with low variability \(std==0\.0382\)\. The red dashed line indicates the ensemble average, with shaded region showing±1\\pm 1standard deviation\.
## Appendix EMulti\-Dataset Benchmark: Full Results and Statistical Tests

This appendix reports the full multi\-dataset embedding benchmark whose headline is summarised in the main text \(Section[3\.1](https://arxiv.org/html/2609.13203#S3.SS1)\): per\-measure fidelity and method rankings, theI±I\_\{\\pm\}exception, and the omnibus and post\-hoc statistical tests\.

### E\.1Multi\-dataset embedding benchmark

To assess embedding fidelity beyond the synthetic benchmark, we evaluated 13 embedding methods across 83 real\-world datasets, 5 PID measures, andN∈\{3,4,5\}N\\in\\\{3,4,5\\\}sources \(Section[2\.6](https://arxiv.org/html/2609.13203#S2.SS6)\)\. Figure[S5](https://arxiv.org/html/2609.13203#A5.F5)summarises the results across all methods; Table[S3](https://arxiv.org/html/2609.13203#A5.T3)reports key statistics for the recommended method \(ACIB\) atN=5N=5\.

##### Embedding fidelity is measure\-dependent\.

AtN=5N=5, mean relative error \(RE\) for ACIB ranged from 19\.4% \(I∧I\_\{\\wedge\}\) to 115\.1% \(I±I\_\{\\pm\}\) \(Table[S3](https://arxiv.org/html/2609.13203#A5.T3)\), with corresponding mean absolute errors \(MAE\) of 0\.057–0\.247 bits\. To put these in scale, typical ground\-truth atom magnitudes are 0\.05–0\.20 bits, so an RE of 45% reflects MAE on the same order as a single atom rather than near\-total signal loss; for comparison, a trivial “predict\-zero” baseline yields RE \> 130% forIminI\_\{\\min\}\.I∧I\_\{\\wedge\}\(Gács–Körner\) was best preserved because its decomposition is sparsest, with most atoms exactly zero and the few non\-zero atoms recovered with <13% per\-atom RE\.IminI\_\{\\min\}andImmiI\_\{\\mathrm\{mmi\}\}showed moderate distortion \(RE≈\\approx45%; AE≈\\approx0\.10 bits\), driven almost entirely by the remainder\-unique and synergy atoms \(Section[E\.1\.1](https://arxiv.org/html/2609.13203#A5.SS1.SSS1)\); the redundancy and source\-unique atoms were recovered with <3% RE in absolute terms\.IrrI\_\{\\mathrm\{rr\}\}showed substantial distortion \(RE==70\.9%\), andI±I\_\{\\pm\}errors exceeded the total signal \(RE\>\>100%\) due to a signed\-atom cancellation failure analysed in Section[E\.1\.1](https://arxiv.org/html/2609.13203#A5.SS1.SSS1), which persists even when the dominant remainder\-unique error component is removed\.

##### ACIB is the recommended default for non\-I±I\_\{\\pm\}measures\.

ACIB delivered the lowest approximation error on the largest share of datasets across non\-I±I\_\{\\pm\}measures, and its advantage grew with the number of sources\. To make this claim concrete we ranked the 13 embedding methods on each dataset by their per\-atom absolute error and counted, for every measure×\\timesNNcombination, how often each method ranked first\. Across all 15 measure×\\timesNNcombinations ACIB ranked first in 8 \(allN≥4N\\geq 4exceptI±I\_\{\\pm\}\); pooling per\-dataset comparisons, ACIB ranked first in 44\.3% of 1,141 cases overall, and in 55\.7% onceI±I\_\{\\pm\}is excluded\. AtN=5N=5, ACIB ranked first on 84\.4% of datasets forI∧I\_\{\\wedge\}\(95% CI: \[78\.4, 91\.1\]%\), 77\.9% forIminI\_\{\\min\}\[71\.1, 85\.0\], and 75\.3% forImmiI\_\{\\mathrm\{mmi\}\}\[68\.6, 82\.9\]; the next\-best methods \(CMIM: 10\.5%, ReliefF: 7\.4%\) trailed substantially\. ACIB’s advantage is not that it is dramatically better at any singleNN, but that it degrades most slowly asNNgrows: its error growth slope was 0\.041 \[95% CI: 0\.035, 0\.046\], versus 0\.044 \[0\.039, 0\.049\] for SIR and PLS, and 0\.046 \[0\.040, 0\.053\] for ReliefF\. AtN=3N=3, where the embedding compresses only two sources, methods were statistically indistinguishable \(Friedmanχ2=12\.0\\chi^\{2\}=12\.0,p=0\.45p=0\.45\)\. Figure[S4](https://arxiv.org/html/2609.13203#A5.F4)summarises method ordering on the synthetic benchmark by per\-pair Pearson correlation between approximated and ground\-truth synergy; the family\-level separation is visible at a glance, and the win\-rate ordering reported above is preserved\.

Figure S4:Ranking of 13 embedding methods by ground\-truth synergy correlation
Note\.Pearson correlationrrbetween embedding\-approximated and ground\-truth synergy across the synthetic Bayesian network benchmark \(50 networks×\\times9 targets×\\times3 source countsN∈\{3,4,5\}N\\in\\\{3,4,5\\\},n=5,377n=5\{,\}377source–target configurations; type\-1 amalgamation throughout\)\. Bars are coloured by method family:*manifold / dim\-reduction*\(blue\),*feature selection*\(green\),*clustering*\(orange\)\. ACIB ranks first by this criterion; manifold\-family methods occupy the top six positions\. ACIB==Agglomerative Conditional Information Bottleneck; JMI==Joint Mutual Information; CMIM==Conditional Mutual Information Maximisation; SIR==Sliced Inverse Regression; MCA==Multiple Correspondence Analysis; PLS==Partial Least Squares; SVD==Singular Value Decomposition; NMF==Non\-negative Matrix Factorisation\.
##### Number of sources matters more than method choice\.

Approximation error roughly tripled fromN=3N=3toN=5N=5regardless of method\. AtN=3N=3, methods were indistinguishable \(as noted above; Nemenyi post\-hoc: 0 significant pairs for all measures\)\. AtN=4N=4andN=5N=5, method differences became highly significant \(Friedmanp<10−22p<10^\{\-22\}\)\. The Nemenyi post\-hoc test, which compares mean ranks across all\(132\)=78\\binom\{13\}\{2\}=78method pairs while controlling the family\-wise error rate, identified 14–37 significant pairs atN=4N=4\([Demšar, 2006](https://arxiv.org/html/2609.13203#bib.bib23)\)\. This pattern—visible as convergent lines atN=3N=3separating sharply at higherNNin Figure[S5](https://arxiv.org/html/2609.13203#A5.F5)—indicates that the fundamental compression challenge of higher\-dimensional remainder sets dominates the choice of compression algorithm\. Full statistical test results are provided in the following subsections\.

Table S3:ACIB embedding fidelity by PID measure atN=5N=5sources
Note\.Mean absolute error \(MAE, bits\) is the per\-atom absolute error averaged across the four PID atoms \(source\-unique, remainder\-unique, redundancy, synergy\) and across datasets\. Mean relative error \(RE\) is the same quantity divided by the magnitude of the corresponding ground\-truth atom\. Win rate is the proportion of per\-dataset contests \(across 77–83 datasets\) in which ACIB achieved the lowest error among the 13 candidate embeddings, with bootstrap 95% confidence intervals\. ForIrrI\_\{\\mathrm\{rr\}\}, 12 of the 13 embeddings were available \(greedy CMI lacked anIrrI\_\{\\mathrm\{rr\}\}ground truth\)\.
#### E\.1\.1TheI±I\_\{\\pm\}exception

The high relative\-error figures forI±I\_\{\\pm\}reflect a structural mismatch betweenI±I\_\{\\pm\}’s information\-flow semantics and the compression that any ePID pipeline performs\.I±I\_\{\\pm\}decomposes mutual information*pointwise*, atom by atom over individual outcomes, and the resulting signed atoms encode local information flow that cancels in structured ways across the joint distribution\([Finn and Lizier, 2018](https://arxiv.org/html/2609.13203#bib.bib28)\)\. AtN=5N=5, 77\.7% of ground\-truthUnqS1\\mathrm\{Unq\}\_\{S\_\{1\}\}values were negative \(compared to 0\.7% forIminI\_\{\\min\}\)\. Embedding methods are designed to compress the joint distribution of the remainder into a single discrete summary; this compression is well\-suited to atoms that average over the joint distribution, but it aggregates over precisely the local outcome\-level structure thatI±I\_\{\\pm\}relies on, so cancellations no longer line up after compression\. The mismatch is therefore conceptual, not a tunable embedding choice\.

##### Practical recommendation\.

Based on the benchmark, we recommend ACIB for analyses usingIminI\_\{\\min\},ImmiI\_\{\\mathrm\{mmi\}\},IrrI\_\{\\mathrm\{rr\}\}, orI∧I\_\{\\wedge\}\(best atN≥4N\\geq 4\); forI±I\_\{\\pm\}, ReliefF \(N≤4N\\leq 4\) or SVD\+k\+k\-means \(N=5N=5\) achieved the lowest absolute errors in our runs, but the structural mismatch means embedding\-based estimates ofI±I\_\{\\pm\}should be interpreted as approximate compositional profiles rather than exact quantities\. AtN=3N=3, method choice is immaterial\. The empirical PHQ\-9 analyses in this paper useImmiI\_\{\\mathrm\{mmi\}\}, for which ACIB is the best\-validated embedding\.

### E\.2Friedman omnibus test

Table[S4](https://arxiv.org/html/2609.13203#A5.T4)reports the Friedman test statistic andpp\-value for each combination of PID measure and number of sources\. The Friedman test assesses whether the ranking of embedding methods differs significantly across datasets\.

Table S4:Friedman test results by PID measure and number of sources
Note\.χ2\\chi^\{2\}statistic andpp\-value testing whether method rankings differ across datasets\. AtN=3N=3, no significant differences were detected for any measure\.
### E\.3Nemenyi post\-hoc tests

Table[S5](https://arxiv.org/html/2609.13203#A5.T5)reports the number of significantly different method pairs \(out of 78 possible\) identified by the Nemenyi post\-hoc test following significant Friedman results\.

Table S5:Number of significantly different method pairs \(Nemenyi post\-hoc\)
Note\.Out of 78 possible pairwise comparisons between 13 methods\. Zero pairs atN=3N=3confirms that all methods perform equivalently when compressing only two sources\.
### E\.4Supplementary benchmark figures

Figure[S5](https://arxiv.org/html/2609.13203#A5.F5)reports embedding approximation error across all 13 methods, 5 PID measures, andN∈\{3,4,5\}N\\in\\\{3,4,5\\\}sources\.

Figure S5:Embedding approximation error across PID measures and number of sources \(all 13 methods\)
Note\.Top row: mean absolute error \(bits\); bottom row: mean relative error \(%\)\. Each column corresponds to one PID measure, ordered by increasing difficulty\. ACIB \(bold orange\) is the lowest\-error method atN≥4N\\geq 4for all measures exceptI±I\_\{\\pm\}\. The dashed line at RE=100=100% marks the threshold where approximation error equals the total signal\.IrrI\_\{\\mathrm\{rr\}\}atN=5N=5is based on 23 datasets; all others on 77–83 datasets\.Figure[S6](https://arxiv.org/html/2609.13203#A5.F6)reports the relative\-error analogue of the main\-text Figure[4](https://arxiv.org/html/2609.13203#S3.F4)\. Normalising each absolute error by the magnitude of the corresponding ground\-truth atom is informative because absolute errors are bounded above by atom magnitude: methods that appear comparable in bits can diverge sharply once normalised\. The relative\-error ranking is consistent with the absolute\-error ranking, with ACIB dominant across all four atoms; the gap between manifold and clustering families widens slightly under relative\-error scaling\.

Figure S6:Mean relative error by embedding method and PID atom
Note\.Per\-atom error normalised by the magnitude of the corresponding ground\-truth atom, complementing the absolute\-error view in Figure[4](https://arxiv.org/html/2609.13203#S3.F4)of the main text\. Panels, methods, and family colour coding are identical to the main\-text figure\. Relative\-error normalisation makes small absolute errors visible when ground\-truth atoms are small in magnitude; the ranking of methods within each panel is preserved against the absolute\-error view\. ACIB==Agglomerative Conditional Information Bottleneck\.
### E\.5Atom\-level error profile across PID measures

For non\-I±I\_\{\\pm\}measures, embedding error in the multi\-dataset benchmark is concentrated in the atoms that involve the compressed remainder:UnqE\\mathrm\{Unq\}\_\{E\}\(remainder\-unique\) andSyn\\mathrm\{Syn\}\(synergy\) carry most of the approximation error, whileUnqS1\\mathrm\{Unq\}\_\{S\_\{1\}\}\(focal\-source unique\) andRed\\mathrm\{Red\}\(redundancy\) are well\-preserved \(Fig\.[S7](https://arxiv.org/html/2609.13203#A5.F7)\)\. ForI±I\_\{\\pm\}, all four atoms exhibit high error, consistent with the disruption of signed cancellation across the full decomposition discussed in the main text\.

Figure S7:Atom\-level error profiles by PID measure \(ACIB,N=5N=5\)
Note\.For non\-I±I\_\{\\pm\}measures, error is concentrated in the remainder\-dependent atoms \(UnqE\\mathrm\{Unq\}\_\{E\},Syn\\mathrm\{Syn\}\)\. ForI±I\_\{\\pm\}, all atoms show high error due to disruption of signed cancellation structure\.

## Appendix FAdditional Fidelity Diagnostics and Symptom\-Level Examples

This appendix collects three diagnostic items referenced from the main text: per\-pair synergy fidelity for representative methods \(Section[F\.1](https://arxiv.org/html/2609.13203#A6.SS1)\), the scaling of synergy approximation error with the number of embedded sources \(Section[F\.2](https://arxiv.org/html/2609.13203#A6.SS2)\), and target\-centred PID examples for two representative PHQ\-9 symptoms in the Xinxiang student sample \(Section[F\.3](https://arxiv.org/html/2609.13203#A6.SS3)\)\.

### F\.1Per\-pair synergy fidelity for representative methods

Figure[S8](https://arxiv.org/html/2609.13203#A6.F8)reports per\-pair scatter for one representative method from each family tier, illustrating the spread visible in the aggregate atom errors of main\-text Figure[4](https://arxiv.org/html/2609.13203#S3.F4)and in the 13\-method ranking of main\-text Figure[S4](https://arxiv.org/html/2609.13203#A5.F4)\. ACIB \(manifold/dim\-reduction\) aligns closely with the identity line \(r=0\.98r=0\.98\); JMI \(feature selection\) shows moderate scatter \(r=0\.41r=0\.41\);kk\-modes \(clustering\) aligns reasonably well \(r=0\.90r=0\.90\), illustrating the upper end of clustering performance\. Hamming\-based clustering methods \(collapse tor≈0r\\approx 0; see main\-text Figure[S4](https://arxiv.org/html/2609.13203#A5.F4)\) are not shown in this trio\.

\(a\)ACIB \(r=0\.98r=0\.98\)\(b\)JMI \(r=0\.41r=0\.41\)\(c\)kk\-modes \(r=0\.90r=0\.90\)
Figure S8:Approximated versus ground\-truth synergy for one representative method per family tier
Note\.Per\-pair scatter of embedding\-approximated synergy \(y\-axis\) against ground\-truth synergy \(x\-axis\), in bits, across 5,377 \(G,Xt,NG,X\_\{t\},N\) configurations from the synthetic Bayesian network benchmark\. Each panel shows one representative method per family\.\(a\)ACIB, manifold/dim\-reduction family, Pearsonr=0\.98r=0\.98\.\(b\)JMI, feature\-selection family,r=0\.41r=0\.41\.\(c\)kk\-modes, clustering family,r=0\.90r=0\.90\. Dashed diagonal is the identity line\. The visible density clustering near ground\-truth synergy≈0\.10\\approx 0\.10reflects multiple synthetic networks producing similar synergy values\. Methods within each family vary in fidelity; main\-text Figure[S4](https://arxiv.org/html/2609.13203#A5.F4)ranks all 13 methods\. The corresponding scaling of synergy approximation error with the number of embedded sources is reported in Section[F\.2](https://arxiv.org/html/2609.13203#A6.SS2)\. ACIB==Agglomerative Conditional Information Bottleneck; JMI==Joint Mutual Information\.
### F\.2Synergy approximation error scaling with the number of sources

For ACIB on the synthetic Bayesian network ensemble, mean absolute error in the synergy atom increased smoothly from0\.00080\.0008bits atN=3N=3to0\.00560\.0056bits atN=5N=5, remaining below0\.0060\.006bits throughout \(Fig\.[S9](https://arxiv.org/html/2609.13203#A6.F9)\)\. A sliced inverse regression \(SIR\) embedding with binning showed slightly higher but remarkably stable synergy error acrossNN\(≈0\.0045\\approx 0\.0045–0\.00480\.0048bits\), whereas linear SVD\+k\+k\-means and clustering methods had both larger baseline errors and steeper increases withNN\.

Figure S9:Synergy approximation error as a function of the number of embedded sources \(N=3,4,5N=3,4,5\)
Note\.ACIB error increases sub\-linearly from0\.00080\.0008bits atN=3N=3to0\.00560\.0056bits atN=5N=5\.
### F\.3Symptom\-level PID examples in the Xinxiang student sample

Figure[S10](https://arxiv.org/html/2609.13203#A6.F10)visualises the embedding\-based PID for two representative targets in the Xinxiang student sample \(anhedonia, sad mood\) as target\-centred radial graphs\. The central dark grey node represents the target; surrounding coloured nodes represent source variables\. Each edge is segmented into four PID components \(teal: redundancy; blue: unique source; purple: unique embedding; orange: synergy\), with edge thickness scaled to total mutual information\. These examples illustrate how strongly connected affective symptoms \(e\.g\., anhedonia and sad mood\) combine substantial unique contributions with non\-negligible synergy, whereas somatic symptoms such as sleep and motor disturbance more often contribute via synergy or redundancy rather than purely unique channels\.

\(a\)Anhedonia as target\(b\)Sad mood as target
Figure S10:Target\-centred PID radial graphs for two representative targets in the Xinxiang student sample
Note\.Each edge is segmented into four PID atoms \(teal: redundancy; blue: unique source; purple: unique embedding; orange: synergy\); edge thickness scales with total mutual information\.

## Appendix GDependency Structure in Clinical Datasets

To assess whether non\-linear conditional dependencies are prevalent in the types of datasets used for symptom network analysis, we compared three unconditional bivariate dependency measures across 80 clinical and psychometric datasets \(32,889 variable pairs in total\)\.

For each dataset, we computed \(i\) absolute Pearson correlation on the imputed continuous data, \(ii\) absolute Spearman correlation on rank\-transformed data, and \(iii\) mutual information on discretised data \(8 equal\-frequency bins\)\. All three measures are unconditional and bivariate, ensuring direct comparability without confounding by conditioning structure\. Statistical significance was assessed using permutation tests \(1,000 permutations per dataset, columns shuffled independently\) with Benjamini–Hochberg FDR correction atα=0\.05\\alpha=0\.05, applied uniformly across all three measures\.

Figure[S11](https://arxiv.org/html/2609.13203#A7.F11)summarises the results\. At the level of statistical significance, Pearson and Spearman correlation agreed almost perfectly \(panel a\): fewer than 1% of significant pairs were detected by one but not the other, confirming that monotonic\-but\-non\-linear dependencies are negligible in these data\. Mutual information detected fewer significant pairs than either correlation measure \(panel b\), consistent with discretisation\-induced information loss rather than with MI capturing additional non\-linear structure\. Across all 80 datasets, only 193 of 30,467 MI\-significant pairs \(0\.6%\) were detected by MI but not by either correlation measure \(panel c\), indicating that non\-linear dependencies invisible to correlation are rare in clinical and psychometric datasets with ordinal response scales\.

These results suggest that the linearity assumption underlying partial correlation networks is largely met in practice for ordinal symptom items\. The primary motivation for information\-theoretic methods in this context is therefore not the detection of non\-linear pairwise associations, but the decomposition of multivariate dependence into redundant, unique, and synergistic components—a distinction that correlation\-based methods cannot provide regardless of functional form\.

Figure S11:Comparison of bivariate dependency measures across 80 clinical and psychometric datasets
Note\.\(a\) Number of significant pairs detected by Pearson vs\. Spearman correlation per dataset; near\-perfect agreement indicates negligible non\-linear monotonic structure\. \(b\) Number of significant pairs detected by Pearson correlation vs\. mutual information; MI generally detects fewer pairs, consistent with discretisation\-induced information loss\. \(c\) Distribution of the percentage of MI\-significant pairs not detected by either correlation measure; the vast majority of datasets show fewer than 2\.5% non\-linear\-only dependencies\.
## Appendix HCross\-cohort Comparison of Information\-Theoretic PHQ\-9 Networks

\(a\)Pearson Partial Correlation Network\(b\)Spearman Partial Correlation Network\(c\)Conditional Mutual Information Network
Figure S12:Xinxiang student sample PHQ\-9 networks under three association measures
Note\.Three measures that progressively relax modelling assumptions of linearity and monotonicity, each leading to a similar network structure: \(a\) Pearson partial correlation, \(b\) Spearman partial correlation, \(c\) conditional mutual information\.\(a\)Full ePID view \(all four atoms\)\(b\)No\-remainder\-unique view \(redundancy and synergy only\)
Figure S13:Two ePID network views in the Xinxiang student sample
Note\.Directed edges represent ordered source→\\rightarrowtarget relations\.\(a\) Full ePID view:each edge is partitioned into remainder\-unique \(blue\), redundancy \(orange\), and synergy \(green\) atoms from the two\-source PID of\(Xi,Ei→k,Xk\)\(X\_\{i\},E\_\{i\\to k\};X\_\{k\}\); the terminal black segment indicates direction\.\(b\) No\-remainder\-unique view:the remainder\-unique atom is removed and the remaining redundancy and synergy atoms are rescaled, highlighting dependence that involves the focal source beyond remainder\-only predictability\. Edge thickness encodes target\-normalised total informationI⁡\(Xi,Ei→k,Xk\)/H⁡\(Xk\)I\(X\_\{i\},E\_\{i\\to k\};X\_\{k\}\)/H\(X\_\{k\}\)\.\(a\)Edge\-wise scatter comparison![Refer to caption](https://arxiv.org/html/2609.13203v1/zenodo_pc_difference_heatmap.png)\(b\)Edge\-wise difference heatmap

Figure S14:Agreement between partial Pearson and partial Spearman correlation networks in the Xinxiang student sample
Note\.Partial Pearson correlations \(PPC\) and Spearman partial correlations \(SPC\) were computed for each symptom pair while conditioning on all remaining PHQ\-9 symptoms \(suicidality excluded\)\.\(a\)Edge\-wise comparison of absolute partial correlations \(\|r\|\|r\|\) shows near\-identical weighting across pairs \(Pearsonr=0\.997r=0\.997; Spearmanρ=0\.995\\rho=0\.995\)\.\(b\)Edge\-wise differences \(SPC−\-PPC\) are small across the matrix, indicating that rank\-based partial correlation and linear partial correlation yield highly similar conditional\-association structures in this setting\.Figure[S15](https://arxiv.org/html/2609.13203#A8.F15)quantifies cross\-dataset agreement in dependence strength and PID composition\. Edge\-wise differences in TMI are systematically positive \(reflecting larger absolute dependence in UK Biobank\), whereas synergy\-proportion differences are concentrated near zero, and the synergy proportions show substantial cross\-dataset agreement \(Pearsonr=0\.663r=0\.663; Spearmanρ=0\.777\\rho=0\.777\)\.

Figure S15:Cross\-dataset comparison of information\-theoretic dependence strength and decomposition stability \(UK Biobank vs the Xinxiang student sample\)
Note\.For each directed source→\\totarget symptom pair \(eight PHQ\-9 symptoms; suicidality excluded\), we compare total mutual information \(TMI\) and the*synergy proportion*\(synergy/TMI\) derived from the embedding\-based PID framework\. Edge\-wise differences \(UK Biobank−\-Xinxiang\) show systematically larger dependence strengths in UK Biobank, whereas synergy\-proportion differences are concentrated near zero\. Synergy proportions show substantial cross\-dataset agreement \(Pearsonr=0\.663r=0\.663; Spearmanρ=0\.777\\rho=0\.777\), indicating that the*composition*of dependence is more stable across datasets than the overall dependence magnitude\.### H\.1Edge definition and comparability across cohorts

We compared information\-theoretic symptom networks estimated separately in two datasets using the same eight PHQ\-9 symptom variables\. For each symptom pair\(X,Y\)\(X,Y\), we estimated the conditional mutual information \(CMI\)I⁡\(X;Y∣𝐙\)I\(X;Y\\mid\\mathbf\{Z\}\), where𝐙\\mathbf\{Z\}denotes the remaining six symptoms\. The “significant edge” sets reported below correspond to the cohort\-specific CMI screening procedure exported in the significant\-edge files\. Although CMI is symmetric inXXandYYgiven the same conditioning set, we represent edges as directed \(X→YX\\to Y\) for two reasons: \(i\) to align with the directional partial information decomposition \(PID\) that is defined for a designated target variable, and \(ii\) to permit visualisation of asymmetries in the PID atoms between the two orientations\. Accordingly, edge direction in these plots should not be interpreted as causal direction\.

Because the cohorts differed in marginal response distributions, their marginal entropiesH⁡\(X\)H\(X\)also differed across items\. Since CMI in bits is upper\-bounded by target uncertainty \(e\.g\.,I⁡\(X;Y∣𝐙\)≤H⁡\(Y\)I\(X;Y\\mid\\mathbf\{Z\}\)\\leq H\(Y\)\), absolute thresholds expressed in bits are not directly comparable across cohorts whenH⁡\(Y\)H\(Y\)differs\. To improve cross\-cohort comparability, we therefore implemented an entropy\-normalised effect size,

nCMIX→Y=I⁡\(X;Y∣𝐙\)H⁡\(Y\),\\mathrm\{nCMI\}\_\{X\\rightarrow Y\}\\;=\\;\\frac\{I\(X;Y\\mid\\mathbf\{Z\}\)\}\{H\(Y\)\},\(7\)which can be interpreted as the fraction of the target’s marginal uncertainty associated with the source after conditioning on the other symptoms\. This normalisation is conservative becauseH⁡\(Y\)H\(Y\)upper\-boundsH⁡\(Y∣𝐙\)H\(Y\\mid\\mathbf\{Z\}\)\.

### H\.2Do cohorts yield the same significant CMI edges?

At the level of statistical significance alone \(i\.e\., without additional pruning\), the UK Biobank cohort exhibited statistically significant CMI for all5656directed symptom pairs among eight nodes, whereas the Xinxiang student sample exhibited5454significant directed edges\. Thus, the Xinxiang student sample’s significant edge set was a strict subset of UK Biobank’s\. Table[S6](https://arxiv.org/html/2609.13203#A8.T6)lists the22directed edges that were significant in UK Biobank but not in the Xinxiang student sample; equivalently, the Xinxiang student sample contained all remaining directed edges\.

Table S6:Directed CMI edges that were statistically significant in UK Biobank but not in the Xinxiang student sample \(2 edges\)\. UK Biobank contained all 56 directed symptom pairs; the Xinxiang student sample contained the complementary set of 54 edges\.
### H\.3Pruning rules for visualisation and “clinical relevance”

Statistical significance alone can yield dense graphs \(notably in UK Biobank\)\. To obtain readable figures while maintaining transparency, we evaluated several simple pruning rules applied*after*significance filtering\. Table[S7](https://arxiv.org/html/2609.13203#A8.T7)summarizes the number of retained edges in each cohort and the overlap between cohorts under: \(i\) an absolute cutoff on raw CMI \(CMI≥0\.05\\geq 0\.05bits\), and \(ii\) rank\-based rules computed onnCMI\\mathrm\{nCMI\}\(top\-10 edges overall, top quartile \[top 25%\], and top\-3 incoming edges per target\)\. The absolute CMI cutoff retained more edges in the Xinxiang student sample than in UK Biobank, consistent with the datasets’ differing entropy scales\. In contrast, rank\-based rules onnCMI\\mathrm\{nCMI\}produced more comparable densities and larger overlaps, thereby supporting interpretable cross\-dataset comparisons\.

Table S7:Overlap of statistically significant CMI edges across cohorts under alternative pruning rules\. Rank\-based rules use entropy\-normalized CMI,nCMI=I⁡\(X;Y∣𝐙\)/H⁡\(Y\)\\mathrm\{nCMI\}=I\(X;Y\\mid\\mathbf\{Z\}\)/H\(Y\)\. The Jaccard index is\|E∩\|/\|E∪\|\|E\_\{\\cap\}\|/\|E\_\{\\cup\}\|for the directed edge sets\.
### H\.4Depression\-status contrast: all significant edges

Figure[S16](https://arxiv.org/html/2609.13203#A8.F16)reproduces the depression\-status contrast of main\-text Figure[7](https://arxiv.org/html/2609.13203#S3.F7)over all 56 significant directed edges rather than the top\-three\-per\-target subset\. The redundancy\-to\-unique/synergy shift is larger on the full edge set \(mean shares: redundancy0\.26→0\.110\.26\\rightarrow 0\.11, unique0\.60→0\.660\.60\\rightarrow 0\.66, synergy0\.14→0\.230\.14\\rightarrow 0\.23\), consistent with the pruned main\-text display but computed from the complete significant\-edge network\.

\(a\)\(b\)Figure S16:ePID networks in non\-depressed versus depressed UK Biobank respondents \(all significant edges\)
Note\.Companion to main\-text Fig\.[7](https://arxiv.org/html/2609.13203#S3.F7), showing all 56 significant directed edges rather than the top\-three\-per\-target subset\.\(a\)non\-depressed \(PHQ\-9 total<10<10\);\(b\)depressed \(PHQ\-9 total≥10\\geq 10\); equal\-NNmatched atn=8,879n=8\{,\}879per arm, five\-seed average, eight PHQ\-9 symptoms \(suicidality excluded\)\. Segment colours \(unique = blue, redundancy = orange, synergy = green\) and the shared total\-mutual\-information thickness scale are as in Fig\.[7](https://arxiv.org/html/2609.13203#S3.F7)\. Over all 56 edges the redundancy\-to\-unique/synergy shift is larger than for the pruned display \(mean shares: redundancy0\.26→0\.110\.26\\rightarrow 0\.11, unique0\.60→0\.660\.60\\rightarrow 0\.66, synergy0\.14→0\.230\.14\\rightarrow 0\.23\)\. As in the main figure, the absolute synergy share isNN\-dependent and the panels should be read as a matched\-NNcontrast, not an absolute level\.

## Appendix IIRI cross\-subscale synergy and redundancy patterns

Figure[S17](https://arxiv.org/html/2609.13203#A9.F17)provides a visual summary of the cross\-subscale triplet PID analysis whose headline percentages appear in main\-text Section[3\.4](https://arxiv.org/html/2609.13203#S3.SS4)\. The three panels report \(a\) synergy\-dominance by source pair for Perspective\-Taking targets, \(b\) redundancy ranking across source pairs pooled across targets, and \(c\) synergy\-dominance rate per target subscale\. Percentages match those quoted in the main text under the same informative\-pair filter \(I⁡\(Xi,Xj,Xk\)≥0\.01I\(X\_\{i\},X\_\{j\};X\_\{k\}\)\\geq 0\.01bits\) and the same cross\-subscale criterion \(one item from each of three distinct subscales asX1X\_\{1\},X2X\_\{2\}, target\)\.

Figure S17:IRI cross\-subscale synergy and redundancy patterns
Note\.Each panel is computed over informative triplets \(total mutual informationI⁡\(Xi,Xj,Xk\)≥0\.01I\(X\_\{i\},X\_\{j\};X\_\{k\}\)\\geq 0\.01bits\) in which the source pair and target item belong to three distinct IRI subscales\.\(a\)Synergy\-dominance rate \(Synergy\>\>Redundancy\) for triplets targeting a Perspective\-Taking item, by cross\-subscale source pair; FS\+\+PD tops the ranking at 69\.4%, consistent with the main\-text claim that mature cognitive empathy is informed synergistically by the affective extremes\.\(b\)Mean redundancy fraction per source pair, pooled across all target subscales; EC\+\+FS tops the ranking at 25\.7%, and all three pairs containing Fantasy rank above all three pairs without it\.\(c\)Synergy\-dominance rate per target subscale, pooled across all cross\-subscale source pairs; PT ranks highest at 48\.9%, ahead of PD \(46\.8%\), EC \(45\.6%\) and FS \(41\.0%\), consistent with developmental accounts in which cognitive empathy emerges later than its affective counterpart and comes to integrate it\([Decety and Holvoet, 2021](https://arxiv.org/html/2609.13203#bib.bib22)\)\. PT==Perspective Taking; FS==Fantasy Scale; EC==Empathic Concern; PD==Personal Distress\.
## Appendix JIllustrative example: stress, sleep, and concentration

To illustrate how PC, CMI, and PID provide complementary information, we construct a hypothetical three\-variable system: stress \(X2X\_\{2\}\), sleep problems \(X3X\_\{3\}\), and concentration \(X1X\_\{1\}\), each on a five\-level ordinal scale \(\{0,1,2,3,4\}\\\{0,1,2,3,4\\\}\)\. Concentration closely tracks sleep quality under typical conditions \(X1=X3X\_\{1\}=X\_\{3\}whenX3∈\{0,1,2,3\}X\_\{3\}\\in\\\{0,1,2,3\\\}\), but when sleep is extremely poor \(X3=4X\_\{3\}=4\), concentration depends non\-monotonically on stress111Note that this equation can be rewritten with use of the modulo operator, but we omit this for ease of explanation\.:

X1=\{X3,ifX3∈\{0,1,2,3\}\(normal sleep\),0,ifX3=4,X2=0\(very poor sleep, no stress\),4,ifX3=4,X2=1\(very poor sleep, mild stress\),3,ifX3=4,X2=2\(very poor sleep, moderate stress\),2,ifX3=4,X2=3\(very poor sleep, high stress\),1,ifX3=4,X2=4\(very poor sleep, extreme stress\)\.X\_\{1\}=\\begin\{cases\}X\_\{3\},&\\text\{if $X\_\{3\}\\in\\\{0,1,2,3\\\}$ \(normal sleep\)\},\\\\\[6\.0pt\] 0,&\\text\{if $X\_\{3\}=4,\\,X\_\{2\}=0$ \(very poor sleep, no stress\)\},\\\\\[2\.0pt\] 4,&\\text\{if $X\_\{3\}=4,\\,X\_\{2\}=1$ \(very poor sleep, mild stress\)\},\\\\\[2\.0pt\] 3,&\\text\{if $X\_\{3\}=4,\\,X\_\{2\}=2$ \(very poor sleep, moderate stress\)\},\\\\\[2\.0pt\] 2,&\\text\{if $X\_\{3\}=4,\\,X\_\{2\}=3$ \(very poor sleep, high stress\)\},\\\\\[2\.0pt\] 1,&\\text\{if $X\_\{3\}=4,\\,X\_\{2\}=4$ \(very poor sleep, extreme stress\)\}\.\\end\{cases\}
We simulate10,00010\{,\}000samples and estimate the network using partial correlations, CMI, and PID \(Figure[S18](https://arxiv.org/html/2609.13203#A10.F18)\)\. The three methods reveal progressively richer structure:

- •Partial correlation:Detects only theX1X\_\{1\}–X3X\_\{3\}edge\. The stress–concentration link is missed because it is non\-monotonic and emerges through a synergistic interaction with sleep\.
- •CMI:Detects both theX1X\_\{1\}–X2X\_\{2\}andX1X\_\{1\}–X3X\_\{3\}edges, revealing that stress influences concentration conditionally—its effect appears only when sleep is very poor\.
- •PID\(usingIminI\_\{\\min\}, a redundancy measure formalised in Section[2\.5\.1](https://arxiv.org/html/2609.13203#S2.SS5.SSS1)\):Decomposes the joint influence into71\.5%71\.5\\%unique to sleep,21\.1%21\.1\\%synergistic between stress and sleep,7\.2%7\.2\\%redundant, and0\.1%0\.1\\%unique to stress\. The synergy fraction quantifies precisely the interaction that PC misses and CMI detects only indirectly\.

Concentration \(X1X\_\{1\}\)Stress \(X2X\_\{2\}\)Sleep \(X3X\_\{3\}\)PC \(0\.709\)CMI \(0\.470 bits\)CMI \(2 bits\)Unique: 0\.1%\(0\.00 bits\)Unique: 71\.5%\(1\.54 bits\)Synergy: 21\.1%\(0\.46 bits\)Redundancy: 7\.2%\(0\.155 bits\)Figure S18:The same three\-variable system as seen by three methods
Note\.PC \(blue, dotted\) only recovers the linearX1X\_\{1\}–X3X\_\{3\}edge; CMI \(red, dashed\) additionally detects the conditionalX1X\_\{1\}–X2X\_\{2\}dependence; PID \(green, solid\) further splits the joint information \(2\.16 bits total\) into unique, redundant, and synergistic atoms, making explicit the synergy that PC misses and CMI detects only indirectly\. This example is constructed for methodological demonstration and does not represent an empirically identified mechanism\.The deterministic structure of this example exaggerates the effect relative to empirical data\. Nonetheless, the decomposition illustrates a general principle: when synergy is high, predictive models that accommodate interactions \(e\.g\., random forests, which achieved perfect accuracy here\) substantially outperform additive models \(adjustedR2=0\.45R^\{2\}=0\.45for linear regression\)\. The information\-theoretic decomposition thus provides principled guidance on when interaction terms are needed\.

## Appendix KRobustness of the ACIB embedding to its hyperparameters

### K\.1Question

The embedding step at the heart of ePID \(the agglomerative conditional information bottleneck, ACIB\) has three free hyperparameters: the conditional\-mutual\-information*loss tolerance*\(default5%5\\%\), the embedding cardinality capKmaxK\_\{\\max\}\(set to1212in the main analyses\), and a Dirichlet*smoothing*constantα\\alpha\(default0\.50\.5\)\. A reviewer will reasonably ask whether the decompositions we report are robust to these settings, or artefacts of a particular tuning\. This appendix answers that with two complementary sensitivity analyses, and in doing so also justifies the choice to embed onlynsources=5n\_\{\\mathrm\{sources\}\}=5remainder symptoms rather than the full network\.

### K\.2What we did

##### Fidelity against ground truth \(synthetic\)\.

On the synthetic networks used to validate the embedding \(which have known, exact ground\-truth PID atoms\), we re\-ran ACIB across a3×33\\times 3grid of loss tolerance∈\{2%,5%,10%\}\\in\\\{2\\%,5\\%,10\\%\\\}andKmax∈\{6,10,14\}K\_\{\\max\}\\in\\\{6,10,14\\\}, plus a one\-dimensional slice over the smoothing constantα∈\{0\.1,0\.5,1\.0\}\\alpha\\in\\\{0\.1,0\.5,1\.0\\\}\. For each setting we recorded the total absolute error of the recovered two\-source atoms against ground truth \(measureImmiI\_\{\\mathrm\{mmi\}\},N=5N=5sources\), the realised embedding cardinality, and the retained conditional mutual information \(the information\-bottleneck “knee”\)\.

##### Stability of the empirical signal \(cohorts\)\.

On the two empirical instruments \(the Xinxiang student sample and the IRI empathy network\) we recomputed the per\-edge synergy fraction across the same grid\. We additionally contrasted thensources=5n\_\{\\mathrm\{sources\}\}=5embedding protocol against a*full\-context*embedding that folds all remaining items into a single variable\.

### K\.3What we found

##### \(i\) Fidelity depends only onKmaxK\_\{\\max\}, and the default is in the stable region\.

The approximation error is*insensitive*to the loss tolerance: it is identical to four decimal places across2%2\\%,5%5\\%, and10%10\\%at everyKmaxK\_\{\\max\}, and to the smoothing constantα\\alpha\. The only consequential knob isKmaxK\_\{\\max\}: error falls monotonically as more clusters are allowed, with clearly diminishing returns \(Table[S8](https://arxiv.org/html/2609.13203#A11.T8), Fig\.[S20](https://arxiv.org/html/2609.13203#A11.F20)\)\. The information\-bottleneck retention curve plateaus at a realised cardinality of roughly ten \(Fig\.[S20](https://arxiv.org/html/2609.13203#A11.F20)\)\. The cap used in the main analyses,Kmax=12K\_\{\\max\}=12, therefore sits above the realised cardinality and within this stable region: it is statistically indistinguishable from the neighbouring grid valuesKmax=10K\_\{\\max\}=10andKmax=14K\_\{\\max\}=14, which differ from one another by only0\.0110\.011bits\.

Table S8:Mean total absolute error of the recovered atoms against ground truth \(N=5N=5,ImmiI\_\{\\mathrm\{mmi\}\},α=0\.5\\alpha=0\.5\)\. Rows \(loss tolerance\) are identical; columns \(KmaxK\_\{\\max\}\) carry all the variation\.
##### \(ii\) The empirical synergy contrast is not a hyperparameter artefact\.

Across the whole grid the per\-edge synergy fraction is stable:6\.16\.1to6\.8%6\.8\\%in the Xinxiang student sample and25\.525\.5to26\.4%26\.4\\%in IRI\. The headline contrast, namely depression networks redundancy\-dominated and empathy networks synergy\-dominated, therefore holds regardless of the embedding settings\.

##### \(iii\) Full\-context embedding saturates, validatingnsources=5n\_\{\\mathrm\{sources\}\}=5\.

Folding*all*remaining items into one embedding collapses the decomposition for the 28\-item IRI network: synergy falls from0\.2560\.256undernsources=5n\_\{\\mathrm\{sources\}\}=5to0\.0440\.044under full context \(Fig\.[S21](https://arxiv.org/html/2609.13203#A11.F21)\)\. For the 9\-item PHQ\-9 the change is modest \(0\.0610\.061to0\.0920\.092, no collapse\)\. This is the expected saturation: when too many items are conditioned on at once, the residual conditional dependence approaches zero and the atoms degenerate into numerical noise\. Limiting the embedding to a small remainder context avoids this regime\.

![Refer to caption](https://arxiv.org/html/2609.13203v1/fig_fidelity_surface.png)Figure S19:Fidelity surface \(mean total absolute error\)\.KmaxK\_\{\\max\}is the only active knob\.
Figure S20:Information\-bottleneck retention plateaus near a realised cardinality of ten\.
Figure S21:Full\-context versusnsources=5n\_\{\\mathrm\{sources\}\}=5synergy fraction\. The 28\-item IRI network saturates under full context \(synergy collapses\); the 9\-item PHQ\-9 does not\.

### K\.4What we conclude

The reported decompositions are robust to the ACIB hyperparameters\. The loss tolerance and the smoothing constant have no measurable effect on fidelity; the only meaningful knob,KmaxK\_\{\\max\}, sits in the stable region at its main\-analysis value of1212, above the realised embedding cardinality\. The empirical synergy estimates, and the redundancy\-versus\-synergy contrast between instruments, are stable across the grid\. Finally, the saturation of the full\-context embedding for large networks independently justifies the decision to embed only a bounded remainder context\.

### K\.5Caveats

- •The synthetic sweep atN=5N=5was run forImmiI\_\{\\mathrm\{mmi\}\}; the other measures \(IminI\_\{\\min\},I∧I\_\{\\wedge\}\) are covered atN=3,4N=3,4and in the full multi\-dataset benchmark\.
- •The cohort sensitivity used a modest sample of edges per configuration: it is a robustness and saturation*demonstration*, not a precise re\-estimate of the headline synergy levels \(which are reported in the main text on the full samples\)\.

## References

- Agency for Healthcare Research and Quality , 2024 \(AHRQ\)Agency for Healthcare Research and Quality \(AHRQ\) \(2024\)\.Medical Expenditure Panel Survey \(MEPS\), 2022 Full Year Consolidated Data File \(HC\-243\)\.U\.S\. Department of Health and Human Services, Agency for Healthcare Research and Quality\.Accessed: 2026\-03\-19\.
- Alegria et al\., \(2007\)Alegria, M\., Jackson, J\. S\., Kessler, R\. C\., and Takeuchi, D\. \(2007\)\.Collaborative Psychiatric Epidemiology Surveys \(CPES\), 2001–2003 \[United States\]\.ICPSR – Interuniversity Consortium for Political and Social Research\.Dataset\. NCS\-R depression module used from DS0002 \(20240\-0002\)\.
- Baer et al\., \(2004\)Baer, R\. A\., Smith, G\. T\., and Allen, K\. B\. \(2004\)\.Assessment of mindfulness by self\-report: The Kentucky Inventory of Mindfulness Skills\.Assessment, 11\(3\):191–206\.
- Banks et al\., \(2023\)Banks, J\., Batty, G\. D\., Breedvelt, J\., Coughlin, K\., Crawford, R\., Marmot, M\., Nazroo, J\., Oldfield, Z\., Steel, N\., Steptoe, A\., Wood, M\., and Zaninotto, P\. \(2023\)\.English Longitudinal Study of Ageing: Waves 0–9, 1998–2019\.UK Data Service\.Wave 6 used in this study\.
- Benjamini and Hochberg, \(1995\)Benjamini, Y\. and Hochberg, Y\. \(1995\)\.Controlling the false discovery rate: A practical and powerful approach to multiple testing\.Journal of the Royal Statistical Society: Series B \(Methodological\), 57\(1\):289–300\.
- Bertschinger et al\., \(2014\)Bertschinger, N\., Rauh, J\., Olbrich, E\., Jost, J\., and Ay, N\. \(2014\)\.Quantifying unique information\.Entropy, 16\(4\):2161–2183\.
- Borsboom, \(2017\)Borsboom, D\. \(2017\)\.A network theory of mental disorders\.World psychiatry, 16\(1\):5–13\.
- Borsboom and Cramer, \(2013\)Borsboom, D\. and Cramer, A\. O\. \(2013\)\.Network analysis: an integrative approach to the structure of psychopathology\.Annual review of clinical psychology, 9\(1\):91–121\.
- \(9\)Börsch\-Supan, A\. \(2022a\)\.easySHARE\.SHARE\-ERIC\.Release version: 8\.0\.0\. Waves 1, 2, 4, 5, 6, 7, 8 used in this study\.
- \(10\)Börsch\-Supan, A\. \(2022b\)\.Survey of Health, Ageing and Retirement in Europe \(SHARE\) Wave 7\.SHARE\-ERIC\.Release version: 8\.0\.0\.
- Borzooei and Tarokhian, \(2023\)Borzooei, S\. and Tarokhian, A\. \(2023\)\.Differentiated thyroid cancer recurrence\.UCI Machine Learning Repository\.
- Brennan et al\., \(1998\)Brennan, K\. A\., Clark, C\. L\., and Shaver, P\. R\. \(1998\)\.Self\-report measurement of adult attachment: An integrative overview\.In Simpson, J\. A\. and Rholes, W\. S\., editors,Attachment Theory and Close Relationships, pages 46–76\. Guilford Press, New York\.Data obtained from the Open\-Source Psychometrics Project,[https://openpsychometrics\.org/\_rawdata/](https://openpsychometrics.org/_rawdata/)\.
- Briganti et al\., \(2019\)Briganti, G\., Fried, E\. I\., and Linkowski, P\. \(2019\)\.Network analysis of Contingencies of Self\-Worth Scale in 680 university students\.Psychiatry Research, 272:252–257\.
- Brown et al\., \(2012\)Brown, G\., Pocock, A\., Zhao, M\.\-J\., and Luján, M\. \(2012\)\.Conditional likelihood maximisation: A unifying framework for information theoretic feature selection\.Journal of Machine Learning Research, 13:27–66\.
- Cella et al\., \(2010\)Cella, D\., Riley, W\., Stone, A\., Rothrock, N\., Reeve, B\., Yount, S\., Amtmann, D\., Bode, R\., Buysse, D\., Choi, S\., Cook, K\., DeVellis, R\., DeWalt, D\., Fries, J\. F\., Gershon, R\., Hahn, E\. A\., Lai, J\.\-S\., Pilkonis, P\., Revicki, D\., Rose, M\., Weinfurt, K\., and Hays, R\. \(2010\)\.The Patient\-Reported Outcomes Measurement Information System \(PROMIS\) developed and tested its first wave of adult self\-reported health outcome item banks: 2005–2008\.Journal of Clinical Epidemiology, 63\(11\):1179–1194\.
- Centers for Disease Control and Prevention, \(2017\)Centers for Disease Control and Prevention \(2017\)\.CDC Diabetes Health Indicators\.UCI Machine Learning Repository\.Dataset, derived from CDC BRFSS 2015 survey data\.
- Centers for Disease Control and Prevention , 2023 \(CDC\)Centers for Disease Control and Prevention \(CDC\) \(2023\)\.Behavioral Risk Factor Surveillance System Survey Data, 2022\.U\.S\. Department of Health and Human Services, Centers for Disease Control and Prevention\.Accessed: 2026\-03\-19\.
- Centers for Disease Control and Prevention , National Center for Health Statistics \(NCHS\), 2020 \(CDC\)Centers for Disease Control and Prevention \(CDC\), National Center for Health Statistics \(NCHS\) \(2020\)\.National Health and Nutrition Examination Survey \(NHANES\), 2017–2018\.U\.S\. Department of Health and Human Services, Centers for Disease Control and Prevention\.Accessed: 2026\-03\-19\.
- Cramer et al\., \(2016\)Cramer, A\. O\. J\., van Borkulo, C\. D\., Giltay, E\. J\., van der Maas, H\. L\. J\., Kendler, K\. S\., Scheffer, M\., and Borsboom, D\. \(2016\)\.Major depression as a complex dynamic system\.PLOS ONE, 11\(12\):e0167490\.
- Davis et al\., \(2020\)Davis, K\. A\. S\., Coleman, J\. R\. I\., Adams, M\., et al\. \(2020\)\.Mental health in uk biobank: development, implementation and results from an online questionnaire completed by 157,366 participants: a reanalysis\.BJPsych Open\.
- Davis, \(1983\)Davis, M\. H\. \(1983\)\.Measuring individual differences in empathy: Evidence for a multidimensional approach\.Journal of Personality and Social Psychology, 44\(1\):113–126\.
- Decety and Holvoet, \(2021\)Decety, J\. and Holvoet, C\. \(2021\)\.The emergence of empathy: A developmental neuroscience perspective\.Developmental Review, 62:100999\.
- Demšar, \(2006\)Demšar, J\. \(2006\)\.Statistical comparisons of classifiers over multiple data sets\.Journal of Machine Learning Research, 7:1–30\.
- Elter, \(2007\)Elter, M\. \(2007\)\.Mammographic mass\.UCI Machine Learning Repository\.
- Elter et al\., \(2007\)Elter, M\., Schulz\-Wendtland, R\., and Wittenberg, T\. \(2007\)\.The prediction of breast cancer biopsy outcomes using two CAD approaches that both emphasize an intelligible decision process\.Medical Physics, 34\(11\):4164–4172\.
- Fernandes et al\., \(2017\)Fernandes, K\., Cardoso, J\. S\., and Fernandes, J\. \(2017\)\.Cervical Cancer \(Risk Factors\)\.UCI Machine Learning Repository\.Dataset\.
- Finch, \(2005\)Finch, H\. \(2005\)\.Comparison of distance measures in cluster analysis with dichotomous data\.Journal of Data Science, 3\(1\):85–100\.
- Finn and Lizier, \(2018\)Finn, C\. and Lizier, J\. T\. \(2018\)\.Pointwise partial information decomposition using the specificity and ambiguity lattices\.Entropy, 20\(4\):297\.
- Fleuret, \(2004\)Fleuret, F\. \(2004\)\.Fast binary feature selection with conditional mutual information\.Journal of Machine Learning Research, 5:1531–1555\.
- Forbes et al\., \(2017\)Forbes, M\. K\., Wright, A\. G\. C\., Markon, K\. E\., and Krueger, R\. F\. \(2017\)\.Evidence that psychopathology symptom networks have limited replicability\.Journal of Abnormal Psychology, 126\(7\):969–988\.
- Franklin et al\., \(2017\)Franklin, J\. C\., Ribeiro, J\. D\., Fox, K\. R\., Bentley, K\. H\., Kleiman, E\. M\., Huang, X\., Musacchio, K\. M\., Jaroszewski, A\. C\., Chang, B\. P\., and Nock, M\. K\. \(2017\)\.Risk factors for suicidal thoughts and behaviors: A meta\-analysis of 50 years of research\.Psychological Bulletin, 143\(2\):187–232\.
- Fried et al\., \(2016\)Fried, E\. I\., Epskamp, S\., Nesse, R\. M\., Tuerlinckx, F\., and Borsboom, D\. \(2016\)\.What are ‘good’ depression symptoms? Comparing the centrality of DSM and non\-DSM symptoms of depression in a network analysis\.Journal of Affective Disorders, 189:314–320\.
- Fried et al\., \(2017\)Fried, E\. I\., van Borkulo, C\. D\., Cramer, A\. O\. J\., Boschloo, L\., Schoevers, R\. A\., and Borsboom, D\. \(2017\)\.Moving forward: Challenges and directions for psychopathological network theory and methodology\.Perspectives on Psychological Science, 12\(6\):999–1020\.
- Gács and Körner, \(1973\)Gács, P\. and Körner, J\. \(1973\)\.Common information is far less than mutual information\.Problems of Control and Information Theory, 2\(2\):149–162\.
- Goldberg, \(1999\)Goldberg, L\. R\. \(1999\)\.A broad\-bandwidth, public domain, personality inventory measuring the lower\-level facets of several five\-factor models\.In Mervielde, I\., Deary, I\. J\., De Fruyt, F\., and Ostendorf, F\., editors,Personality Psychology in Europe, volume 7, pages 7–28\. Tilburg University Press, Tilburg, The Netherlands\.Data obtained from the Open\-Source Psychometrics Project,[https://openpsychometrics\.org/\_rawdata/](https://openpsychometrics.org/_rawdata/)\.
- \(36\)Golovenkin, S\. E\., Bac, J\., Chervov, A\., Mirkes, E\. M\., Orlova, Y\. V\., Barillot, E\., Gorban, A\. N\., and Zinovyev, A\. \(2020a\)\.Myocardial infarction complications\.UCI Machine Learning Repository\.Dataset\.
- \(37\)Golovenkin, S\. E\., Bac, J\., Chervov, A\., Mirkes, E\. M\., Orlova, Y\. V\., Barillot, E\., Gorban, A\. N\., and Zinovyev, A\. \(2020b\)\.Trajectories, bifurcations, and pseudo\-time in large clinical datasets: Applications to myocardial infarction and diabetes data\.GigaScience, 9\(11\):giaa128\.
- Golub and Van Loan, \(2013\)Golub, G\. H\. and Van Loan, C\. F\. \(2013\)\.Matrix Computations\.Johns Hopkins University Press, 4th edition\.
- Gondek and Hofmann, \(2003\)Gondek, D\. and Hofmann, T\. \(2003\)\.Conditional information bottleneck clustering\.InProceedings of the 3rd IEEE International Conference on Data Mining, Workshop on Clustering Large Data Sets, pages 36–42\.
- Goodwell and Kumar, \(2017\)Goodwell, A\. E\. and Kumar, P\. \(2017\)\.Temporal information partitioning: Characterizing synergy, uniqueness, and redundancy in interacting environmental variables\.Water Resources Research, 53\(7\):5920–5942\.
- Greenacre, \(2017\)Greenacre, M\. \(2017\)\.Correspondence Analysis in Practice\.Chapman and Hall/CRC, 3rd edition\.
- Gruber et al\., \(2014\)Gruber, S\., Hunkler, C\., and Stuck, S\. \(2014\)\.Generating easySHARE: Guidelines, structure, content and programming\.SHARE Working Paper Series 17\-2014, Munich Center for the Economics of Aging \(MEA\), Max Planck Institute for Social Law and Social Policy\.
- Guloksuz et al\., \(2017\)Guloksuz, S\., Pries, L\., and Van Os, J\. \(2017\)\.Application of network methods for understanding mental disorders: pitfalls and promise\.Psychological medicine, 47\(16\):2743–2752\.
- Guvenir and Demiroz, \(1998\)Guvenir, H\. A\. and Demiroz, G\. \(1998\)\.Dermatology\.UCI Machine Learning Repository\.
- Guvenir et al\., \(1998\)Guvenir, H\. A\., Demiroz, G\., and Ilter, N\. \(1998\)\.Learning differential diagnosis of erythemato\-squamous diseases using voting feature intervals\.Artificial Intelligence in Medicine, 13\(3\):147–165\.
- Halford, \(2023\)Halford, M\. \(2023\)\.prince: Multivariate exploratory data analysis in Python\.Python package\.
- Hammer and Katzenstein, \(1996\)Hammer, S\. M\. and Katzenstein, D\. A\. \(1996\)\.AIDS Clinical Trials Group Study 175\.UCI Machine Learning Repository\.Dataset\.
- Hammer et al\., \(1996\)Hammer, S\. M\., Katzenstein, D\. A\., Hughes, M\. D\., Gundacker, H\., Schooley, R\. T\., Haubrich, R\. H\., Henry, W\. K\., Lederman, M\. M\., Phair, J\. P\., Niu, M\., Hirsch, M\. S\., and Merigan, T\. C\. \(1996\)\.A trial comparing nucleoside monotherapy with combination therapy in HIV\-infected adults with CD4 cell counts from 200 to 500 per cubic millimeter\.New England Journal of Medicine, 335\(15\):1081–1090\.
- Haslbeck et al\., \(2021\)Haslbeck, J\. M\. B\., Borsboom, D\., and Waldorp, L\. J\. \(2021\)\.Moderated network models\.Multivariate Behavioral Research, 56\(2\):256–287\.
- Hays et al\., \(2018\)Hays, R\. D\., Spritzer, K\. L\., Schalet, B\. D\., and Cella, D\. \(2018\)\.PROMIS®\-29 v2\.0 profile physical and mental health summary scores\.Quality of Life Research, 27\(7\):1885–1891\.
- Health and Retirement Study, \(2024\)Health and Retirement Study \(2024\)\.Harmonized HRS version C\.University of Michigan, Institute for Social Research\.Wave 13 \(2016\) used in this study\. Harmonized HRS Version C\. DOI 10\.7910/DVN/8KXWPK could not be verified – using institutional URL instead\.
- Hendin and Cheek, \(1997\)Hendin, H\. M\. and Cheek, J\. M\. \(1997\)\.Assessing hypersensitive narcissism: A reexamination of Murray’s Narcism scale\.Journal of Research in Personality, 31\(4\):588–599\.
- Horton and Perry, \(2016\)Horton, M\. and Perry, A\. E\. \(2016\)\.Screening for depression in primary care: A Rasch analysis of the PHQ\-9\.BJPsych Bulletin, 40\(5\):237–243\.
- Huang, \(1998\)Huang, Z\. \(1998\)\.Extensions to thekk\-means algorithm for clustering large data sets with categorical values\.Data Mining and Knowledge Discovery, 2\(3\):283–304\.
- Ince, \(2017\)Ince, R\. A\. A\. \(2017\)\.The partial entropy decomposition: Decomposing multivariate entropy and mutual information via pointwise common surprisal\.
- Islam et al\., \(2019\)Islam, M\. M\. F\., Ferdousi, R\., Rahman, S\., and Bushra, H\. Y\. \(2019\)\.Likelihood prediction of diabetes at early stage using data mining techniques\.InComputer Vision and Machine Intelligence in Medical Image Analysis, volume 992 ofAdvances in Intelligent Systems and Computing, pages 113–125\. Springer\.CORRECTION: year 2019 per verified metadata\.
- Islam et al\., \(2020\)Islam, M\. M\. F\., Ferdousi, R\., Rahman, S\., and Bushra, H\. Y\. \(2020\)\.Early stage diabetes risk prediction\.UCI Machine Learning Repository\.
- James et al\., \(2018\)James, R\. G\., Ellison, C\. J\., and Crutchfield, J\. P\. \(2018\)\.dit: a Python package for discrete information theory\.Journal of Open Source Software, 3\(25\):738\.
- Jansma et al\., \(2025\)Jansma, A\., Mediano, P\. A\. M\., and Rosas, F\. E\. \(2025\)\.Fast Möbius transform: An algebraic approach to information decomposition\.Physical Review Research, 7\(3\):033049\.
- Jonason and Webster, \(2010\)Jonason, P\. K\. and Webster, G\. D\. \(2010\)\.The dirty dozen: A concise measure of the dark triad\.Psychological Assessment, 22\(2\):420–432\.
- Kalichman and Rompa, \(1995\)Kalichman, S\. C\. and Rompa, D\. \(1995\)\.Sexual sensation seeking and sexual compulsivity scales: Validity, and predicting HIV risk behavior\.Journal of Personality Assessment, 65\(3\):586–601\.
- Kamal and ElEleimy, \(2017\)Kamal, S\. and ElEleimy, M\. \(2017\)\.Hepatitis C Virus \(HCV\) for Egyptian patients\.UCI Machine Learning Repository\.Dataset\.
- Kessler and Merikangas, \(2004\)Kessler, R\. C\. and Merikangas, K\. R\. \(2004\)\.The National Comorbidity Survey Replication \(NCS\-R\): Background and aims\.International Journal of Methods in Psychiatric Research, 13\(2\):60–68\.
- Kleitman, \(1969\)Kleitman, D\. J\. \(1969\)\.On dedekind’s problem: The number of monotone boolean functions\.Proceedings of the American Mathematical Society, 21\(3\):677–682\.
- Koepke, \(2018\)Koepke, J\. \(2018\)\.sliced: Sliced inverse regression in Python\.Python package\.
- Kolchinsky, \(2024\)Kolchinsky, A\. \(2024\)\.Partial information decomposition: Redundancy as information bottleneck\.Entropy, 26:546\.
- Kononenko, \(1994\)Kononenko, I\. \(1994\)\.Estimating attributes: Analysis and extensions of RELIEF\.Machine Learning: ECML\-94, pages 171–182\.
- Kroenke et al\., \(2001\)Kroenke, K\., Spitzer, R\. L\., and Williams, J\. B\. W\. \(2001\)\.The phq\-9: Validity of a brief depression severity measure\.Journal of General Internal Medicine, 16\(9\):606–613\.
- Lawrence et al\., \(2004\)Lawrence, E\. J\., Shaw, P\., Baker, D\., Baron\-Cohen, S\., and David, A\. S\. \(2004\)\.Measuring empathy: Reliability and validity of the empathy quotient\.Psychological Medicine, 34\(5\):911–920\.
- Lee and Seung, \(1999\)Lee, D\. D\. and Seung, H\. S\. \(1999\)\.Learning the parts of objects by non\-negative matrix factorization\.Nature, 401\(6755\):788–791\.
- Li et al\., \(2018\)Li, J\., Cheng, K\., Wang, S\., Morstatter, F\., Trevino, R\. P\., Tang, J\., and Liu, H\. \(2018\)\.scikit\-feature: Feature selection repository in Python\.Python package\.
- Li, \(1991\)Li, K\.\-C\. \(1991\)\.Sliced inverse regression for dimension reduction\.Journal of the American Statistical Association, 86\(414\):316–327\.
- Lovibond and Lovibond, \(1995\)Lovibond, P\. F\. and Lovibond, S\. H\. \(1995\)\.The structure of negative emotional states: Comparison of the Depression Anxiety Stress Scales \(DASS\) with the Beck Depression and Anxiety Inventories\.Behaviour Research and Therapy, 33\(3\):335–343\.
- Malgaroli et al\., \(2021\)Malgaroli, M\., Calderon, A\., and Bonanno, G\. A\. \(2021\)\.Networks of major depressive disorder: A systematic review\.Clinical Psychology Review, 85:102000\.
- \(75\)Martínez, L\., Robles, E\., Trofimoff, V\., Vidal, N\., Espada, A\. D\., Mosquera, N\., Franco, B\., Sarmiento, V\., and Zafra, M\. I\. \(2024a\)\.Subjective well\-being and mental health among college students: Dataset\.Mendeley Data\.Dataset accompanying Martínez et al\. \(2024\),Data, 9\(3\), 44\.
- \(76\)Martínez, L\., Robles, E\., Trofimoff, V\., Vidal, N\., Espada, A\. D\., Mosquera, N\., Franco, B\., Sarmiento, V\., and Zafra, M\. I\. \(2024b\)\.Subjective well\-being and mental health among college students: Two datasets for diagnosis and program evaluation\.Data, 9\(3\):44\.
- McNally et al\., \(2015\)McNally, R\. J\., Robinaugh, D\. J\., Wu, G\. W\. Y\., Wang, L\., Deserno, M\. K\., and Borsboom, D\. \(2015\)\.Mental disorders as causal systems: A network approach to posttraumatic stress disorder\.Clinical Psychological Science, 3\(6\):836–849\.
- Murtagh and Contreras, \(2012\)Murtagh, F\. and Contreras, P\. \(2012\)\.Algorithms for hierarchical clustering: An overview\.WIREs Data Mining and Knowledge Discovery, 2\(1\):86–97\.
- Ng et al\., \(2001\)Ng, A\. Y\., Jordan, M\. I\., and Weiss, Y\. \(2001\)\.On spectral clustering: Analysis and an algorithm\.InAdvances in Neural Information Processing Systems, volume 14\.
- Paninski, \(2003\)Paninski, L\. \(2003\)\.Estimation of entropy and mutual information\.Neural Computation, 15\(6\):1191–1253\.
- Pedregosa et al\., \(2011\)Pedregosa, F\., Varoquaux, G\., Gramfort, A\., Michel, V\., Thirion, B\., Grisel, O\., Blondel, M\., Prettenhofer, P\., Weiss, R\., Dubourg, V\., et al\. \(2011\)\.Scikit\-learn: Machine learning in Python\.Journal of Machine Learning Research, 12:2825–2830\.
- Rosas et al\., \(2019\)Rosas, F\. E\., Mediano, P\. A\. M\., Gastpar, M\., and Jensen, H\. J\. \(2019\)\.Quantifying high\-order interdependencies via multivariate extensions of the mutual information\.Physical Review E, 100:032305\.
- Rosenberg, \(1965\)Rosenberg, M\. \(1965\)\.Society and the Adolescent Self\-Image\.Princeton University Press, Princeton, NJ\.Data obtained from the Open\-Source Psychometrics Project,[https://openpsychometrics\.org/\_rawdata/](https://openpsychometrics.org/_rawdata/)\.
- Rubini et al\., \(2015\)Rubini, L\., Soundarapandian, P\., and Eswaran, P\. \(2015\)\.Chronic Kidney Disease\.UCI Machine Learning Repository\.Dataset\.
- Ryff et al\., \(2019\)Ryff, C\. D\., Almeida, D\. M\., Ayanian, J\. Z\., Carr, D\. S\., Cleary, P\. D\., Coe, C\., Davidson, R\. J\., Krueger, R\. F\., Lachman, M\. E\., Marks, N\. F\., Mroczek, D\. K\., Seeman, T\. E\., Seltzer, M\. M\., Singer, B\. H\., Sloan, R\. P\., Tun, P\. A\., Weinstein, M\., and Williams, D\. R\. \(2019\)\.Midlife in the United States \(MIDUS 2\), 2004–2006\.ICPSR – Interuniversity Consortium for Political and Social Research\.Dataset, version 8\. Datasets \#89–91 use variables from DS0001 \(04652\-0001\)\.
- Scagliarini et al\., \(2023\)Scagliarini, T\. et al\. \(2023\)\.Gradients of o\-information: Low\-order descriptors of high\-order dependencies\.Physical Review Research, 5:013025\.
- Schulz and Ritter, \(2026\)Schulz, M\.\-A\. and Ritter, K\. \(2026\)\.Measurement noise limits the advantage of nonlinear models over linear models in biomedical prediction\.
- Shannon, \(1948\)Shannon, C\. E\. \(1948\)\.A mathematical theory of communication\.The Bell system technical journal, 27\(3\):379–423\.
- Slonim and Tishby, \(1999\)Slonim, N\. and Tishby, N\. \(1999\)\.Agglomerative information bottleneck\.InAdvances in Neural Information Processing Systems, volume 12, pages 617–623\.
- Spitzer et al\., \(1999\)Spitzer, R\. L\., Kroenke, K\., Williams, J\. B\. W\., and the Patient Health Questionnaire Primary Care Study Group \(1999\)\.Validation and utility of a self\-report version of PRIME\-MD: The PHQ primary care study\.JAMA, 282\(18\):1737–1744\.
- \(91\)Su, Z\., Liu, R\., Wei, Y\., Zhang, R\., Xu, X\., Wang, Y\., Zhu, Y\., Wang, L\., Liang, L\., Wang, F\., and Zhang, X\. \(2024a\)\.Temporal dynamics in psychological assessments: A novel dataset with scales and response times\.Scientific Data, 11\(1\):1046\.
- \(92\)Su, Z\., Liu, R\., Wei, Y\., Zhang, R\., Xu, X\., Wang, Y\., Zhu, Y\., Wang, L\., Liang, L\., Wang, F\., and Zhang, X\. \(2024b\)\.Temporal dynamics in psychological assessments: Scales and response times dataset\.Zenodo\.Dataset accompanying Su et al\. \(2024\),Scientific Data, 11, 1046\.
- Sudlow et al\., \(2015\)Sudlow, C\., Gallacher, J\., Allen, N\., Beral, V\., Burton, P\., Danesh, J\., Downey, P\., Elliott, P\., Green, J\., Landray, M\., et al\. \(2015\)\.Uk biobank: an open access resource for identifying the causes of a wide range of complex diseases of middle and old age\.PLoS medicine, 12\(3\):e1001779\.
- Syeed, \(2024\)Syeed, M\. M\. \(2024\)\.Mental health problems among university students: GAD\-7, PSS\-10, and PHQ\-9 dataset\.Figshare\.Author list may be incomplete; verify via DataCite\.
- Tanrikulu and Er, \(2012\)Tanrikulu, A\. C\. and Er, O\. \(2012\)\.Mesothelioma’s disease data set\.UCI Machine Learning Repository\.Dataset\.
- Tasci and Camphausen, \(2022\)Tasci, E\. and Camphausen, K\. \(2022\)\.Glioma Grading Clinical and Mutation Features\.UCI Machine Learning Repository\.Dataset\.
- Taylor, \(1953\)Taylor, J\. A\. \(1953\)\.A personality scale of manifest anxiety\.The Journal of Abnormal and Social Psychology, 48\(2\):285–290\.
- Thabtah, \(2017\)Thabtah, F\. \(2017\)\.Autism screening adult\.UCI Machine Learning Repository\.
- Thabtah, \(2018\)Thabtah, F\. \(2018\)\.Autism spectrum disorder screening: Machine learning adaptation framework\.Informatics for Health and Social Care, 44:278–297\.CORRECTION: year 2018, volume 44, pages 278–297 per verified metadata\.
- Timme and Lapish, \(2018\)Timme, N\. M\. and Lapish, C\. \(2018\)\.A tutorial for information theory in neuroscience\.eNeuro, 5\(3\):ENEURO\.0052–18\.2018\.
- Tishby et al\., \(1999\)Tishby, N\., Pereira, F\. C\., and Bialek, W\. \(1999\)\.The information bottleneck method\.Proceedings of the 37th Annual Allerton Conference on Communication, Control, and Computing, pages 368–377\.
- UCI Machine Learning Repository, \(1988\)UCI Machine Learning Repository \(1988\)\.Primary tumor\.UCI Machine Learning Repository\.
- University of Michigan Institute for Healthcare Policy and Innovation, \(2023\)University of Michigan Institute for Healthcare Policy and Innovation \(2023\)\.National Poll on Healthy Aging \(NPHA\)\.UCI Machine Learning Repository\.No working DOI available\. Accessed: 2026\-03\-19\.
- Unknown, \(2015\)Unknown \(2015\)\.PHQ\-9, GAD\-7, and Epworth Sleepiness Scale data from Mexico\.Figshare\.Author and title to be confirmed via DataCite or landing page\.
- Unknown, \(2023\)Unknown \(2023\)\.Mental health dataset from Bangladesh: BDI\-II and UCLA loneliness items\.Mendeley Data\.Author and title to be confirmed via DataCite or landing page\.
- Unknown, \(2024\)Unknown \(2024\)\.PHQ\-9 student depression data\.Mendeley Data\.Author and title to be confirmed via DataCite or landing page\.
- Urbanowicz et al\., \(2018\)Urbanowicz, R\. J\., Meeker, M\., La Cava, W\., Olson, R\. S\., and Moore, J\. H\. \(2018\)\.Benchmarking relief\-based feature selection methods for bioinformatics data mining\.Journal of Biomedical Informatics, 85:168–188\.
- U\.S\. Census Bureau, \(2025\)U\.S\. Census Bureau \(2025\)\.Household Pulse Survey: Measuring Social and Economic Impacts During the Coronavirus Pandemic, Phase 4\.2, Cycle 4 \(December 2024\)\.U\.S\. Census Bureau\.Accessed: 2026\-03\-19\.
- Vallat, \(2018\)Vallat, R\. \(2018\)\.Pingouin: statistics in Python\.Journal of Open Source Software, 3\(31\):1026\.
- Varley, \(2024\)Varley, T\. F\. \(2024\)\.A scalable synergy\-first backbone decomposition of higher\-order structures in complex systems\.npj Complexity\.
- Varley et al\., \(2025\)Varley, T\. F\., Mediano, P\. A\. M\., Patania, A\., and Bongard, J\. \(2025\)\.The topology of synergy: Linking topological and information\-theoretic approaches to higher\-order interactions in complex systems\.PLOS Computational Biology, 21\(11\):e1013649\.
- Ver Steeg and Galstyan, \(2014\)Ver Steeg, G\. and Galstyan, A\. \(2014\)\.NPEET: Non\-parametric entropy estimation toolbox\.Python package\.
- Vervaet et al\., \(2021\)Vervaet, M\., Puttevils, L\., Hoekstra, R\. H\. A\., Fried, E\., and Vanderhasselt, M\.\-A\. \(2021\)\.Transdiagnostic vulnerability factors in eating disorders: A network analysis\.European Eating Disorders Review, 29\(1\):86–100\.
- Williams and Beer, \(2010\)Williams, P\. L\. and Beer, R\. D\. \(2010\)\.Nonnegative decomposition of multivariate information\.arXiv preprint arXiv:1004\.2515\.
- Wold et al\., \(1984\)Wold, S\., Ruhe, A\., Wold, H\., and Dunn, III, W\. J\. \(1984\)\.The collinearity problem in linear regression\. The partial least squares \(PLS\) approach to generalized inverses\.SIAM Journal on Scientific and Statistical Computing, 5\(3\):735–743\.
- \(116\)Wu, Y\., Yan, W\., Wu, Y\., and Peng, K\. \(2024a\)\.Adaptation and validation of the Claremont Purpose Scale in Chinese adolescents\.Journal of Research on Adolescence, 34:776–790\.
- \(117\)Wu, Y\., Yan, W\., Wu, Y\., and Peng, K\. \(2024b\)\.Claremont purpose scale chinese adaptation: Dataset\.Open Science Framework\.Dataset accompanying Wu et al\. \(2024\),Journal of Research on Adolescence, 34, 776–790\.
- Yamada et al\., \(2021\)Yamada, Y\., Ćepulić, D\.\-B\., Coll\-Martín, T\., Debove, S\., Gautreau, G\., Han, H\., Rasmussen, J\., Tran, T\. P\., Travaglino, G\. A\., et al\. \(2021\)\.COVIDiSTRESS global survey dataset on psychological and behavioural consequences of the COVID\-19 outbreak\.Scientific Data, 8\(1\):3\.
- Yang and Moody, \(1999\)Yang, H\. H\. and Moody, J\. \(1999\)\.Data visualization and feature selection: New algorithms for nongaussian data\.Advances in Neural Information Processing Systems, 12\.
- Zieba et al\., \(2014\)Zieba, M\., Tomczak, J\. M\., Lubicz, M\., and Swiatek, J\. \(2014\)\.Thoracic surgery data\.UCI Machine Learning Repository\.CORRECTION: year 2014 per Zięba et al\. \(2014\),Applied Soft Computing, 14, 99–108\.
- Zięba et al\., \(2014\)Zięba, M\., Tomczak, J\. M\., Lubicz, M\., and Świątek, J\. \(2014\)\.Boosted SVM for extracting rules from imbalanced data in application to prediction of the post\-operative life expectancy in the lung cancer patients\.Applied Soft Computing, 14:99–108\.

Similar Articles

Interpretable Depression Detection from Social Media Text Using LLM-Derived Embeddings

arXiv cs.CL

This paper investigates the use of large language models (LLMs) and supervised classifiers for depression detection from social media text, proposing a prompt-based embedding method that enhances interpretability. Experiments on multiple datasets show that zero-shot LLMs perform well for binary classification but struggle with fine-grained severity, while supervised models on LLM summary embeddings achieve more consistent performance across multi-class and ordinal tasks.

A Granularity-Aware EEG Feature Framework for Psychopathology Dimension Prediction

arXiv cs.LG

This paper presents a granularity-aware EEG feature framework that organizes multi-scale descriptors into global, regional, and channel levels to predict dimensional psychopathology. Using the HBN cohort, it shows that tree-based models and granularity-balanced feature selection yield modest improvements, suggesting multi-scale EEG features contain weak but detectable signals for pediatric mental health.