Tree-Structured Orthonormal Decomposition of the Aitchison Simplex
Summary
This paper introduces PolyILR, a canonical orthonormal decomposition of the Aitchison tangent space that aligns with any tree topology, providing stable and interpretable features for compositional data such as microbiome and single-cell profiles.
View Cached Full Text
Cached at: 06/11/26, 01:50 PM
# Tree-Structured Orthonormal Decomposition of the Aitchison Simplex
Source: [https://arxiv.org/html/2606.11646](https://arxiv.org/html/2606.11646)
Qijun ZhangTravis PenceBarbara B\. BendlinFederico ReyVikas Singh
###### Abstract
Compositional data—vectors encoding relative proportions—arise across scientific domains, including ecology, geochemistry, and genomics\. The features in these data often come with known hierarchical structure \(e\.g\., taxonomies, phylogenies, ontologies\), yet existing methods either ignore this structure, discard the intrinsic Aitchison geometry, are designed for binary trees, or yield incomplete coordinate systems\. We describe*PolyILR*, a canonical orthonormal decomposition of the Aitchison tangent space aligned with any tree topology\. Our construction defines a weighted local geometry at each internal node capturing full branching structure, then lifts these to a global orthonormal basis where every coordinate corresponds to a specific tree location\. On microbiome and single\-cell benchmarks, PolyILR yields stable, interpretable features and enables inference at multiscale tree resolution\. We also establish a novel theoretical connection to softmax classifiers, suggesting possible applications to probabilistic modeling\.
Machine Learning, ICML
## 1Introduction
*Compositional data*—nonnegative vectors \(or components\) whose summed total carries no intrinsic meaning—arise across scientific and statistical settings, including microbiome profiles, cell type proportions, ecological counts, and probabilistic model outputs\(Glooret al\.,[2017](https://arxiv.org/html/2606.11646#bib.bib13); Billheimeret al\.,[2001](https://arxiv.org/html/2606.11646#bib.bib11); Buettneret al\.,[2021](https://arxiv.org/html/2606.11646#bib.bib24)\)\. Such data live on the simplex, where inference is based on*relative*not absolute values\. Often, the components are not exchangeable and are organized by domain hierarchies reflecting evolutionary, functional, or semantic relationships\(Silvermanet al\.,[2017](https://arxiv.org/html/2606.11646#bib.bib4); Harmon,[2019](https://arxiv.org/html/2606.11646#bib.bib16)\)\. These trees describe which comparisons among components are meaningful and at what resolution\.
The standard tool in compositional data analysis \(CoDA\) is*Aitchison geometry*\(Aitchison,[1982](https://arxiv.org/html/2606.11646#bib.bib9)\), which formalizes that only ratios are informative\. This is achieved via an isometric log‑ratio \(ILR\) embedding of compositional data into Euclidean space\(Egozcueet al\.,[2003](https://arxiv.org/html/2606.11646#bib.bib1)\)\. Here, one treats components*symmetrically*, partly reflecting origins in domains where no external structure was assumed\(Egozcue and Pawlowsky\-Glahn,[2005](https://arxiv.org/html/2606.11646#bib.bib6); Mandalet al\.,[2015](https://arxiv.org/html/2606.11646#bib.bib20)\)\. But in many applications, domain\-specific hierarchies are known: trees specify which comparisons are meaningful and how they are related\. To address this gap, the literature provides strategies to incorporate tree structure, such as tree‑based balances and phylogeny‑aware coordinates\. Doing so improves interpretability and downstream analysis\. Yet, existing approaches typically address specific regimes \(e\.g\., binary trees as in phylogenetic reconstruction, selected contrasts\)\(Silvermanet al\.,[2017](https://arxiv.org/html/2606.11646#bib.bib4); Washburneet al\.,[2017](https://arxiv.org/html/2606.11646#bib.bib5); Mortonet al\.,[2017](https://arxiv.org/html/2606.11646#bib.bib22)\)or rely on construction choicesexternalto the geometry\(Lozupone and Knight,[2005](https://arxiv.org/html/2606.11646#bib.bib39); Mao and Ma,[2022](https://arxiv.org/html/2606.11646#bib.bib40)\)\. For instance, a polytomous node admits no canonical binary resolution, yet binary\-tree methods force an arbitrary choice \(Figure[1](https://arxiv.org/html/2606.11646#S2.F1)\)\. Hence, the key question we study is whether hierarchies can be incorporated in a way that is general, geometrically principled, and canonical for arbitrary trees\.
Main difficulty\.The core issue is structural incompatibility\. Aitchison geometry identifies compositions up to a*global*scaling: a\(d−1\)\(d\-1\)\-dimensional Euclidean tangent space where ILR coordinates live\. But trees impose*local*and*nested*constraints: distinctions are meaningful within clades \(i\.e\., subtrees\) and comparisons happen at multiple resolutions\. Aligning these requiresdecomposingthe Aitchison tangent space to respect branching structure at every internal node\. Binary trees sidestep this issue by reducing each node to a single contrast \(Figure[1](https://arxiv.org/html/2606.11646#S2.F1)\)\. For multi‑branching \(i\.e\.,*polytomous*\) hierarchies, no canonical construction exists\.
This paper\.We ask if an orthonormal decomposition of the Aitchison tangent space can be compatible with arbitrary tree structure and the simplex invariances—without sacrificing isometry or introducing arbitrary choices\. Such a decomposition would give each internal node a geometrically motivated signature and provide a multiscale coordinate system for simplex\-valued data on the same footing as standard ILR methods\(Pawlowsky\-Glahn and Egozcue,[2001](https://arxiv.org/html/2606.11646#bib.bib10)\), while grounding the analysis firmly in tree structure\.
The absence of such a coordinate system extends beyond traditional CoDA to probabilistic model representations\. Softmax outputs are simplex\-valued but often represented in flat probability coordinates, while structure over outcomes is increasingly explicit: class taxonomies, semantic hierarchies, grammars, tree\-structured search spaces\(Silla Jr and Freitas,[2011](https://arxiv.org/html/2606.11646#bib.bib59)\)\. The absence of a canonical tree\-aligned coordinate system beyond heuristics limits analysis of where probability mass, errors, or learning signals concentrate\.
Contributions\.We \(1\) constructPolyILR\(Polytomous ILR\), a canonical orthonormal decomposition of the Aitchison tangent space aligned with arbitrary trees, answering affirmatively the question raised above, \(2\) demonstrate its utility in CoDA; stable feature selection and tree\-level inference in standard microbiome and single\-cell datasets, \(3\) establish a novel theoretical connection to softmax classifiers via shared invariance structure\.Our goal is interpretability of the representation itself, not downstream performance: each coordinate corresponds to a specific tree location \(i\.e\., a node\-contrast pair\) that yields consistent, tree\-grounded analysis and inference \(see Table[1](https://arxiv.org/html/2606.11646#S5.T1)\)\.Code is available at[https://github\.com/vsingh\-group/polyilr](https://github.com/vsingh-group/polyilr)\.
Conflict of Interest Disclosure\.The authors declare no financial conflicts of interest related to this work\.
## 2Background
Original𝒯\\mathcal\{T\}1234k=3k\\\!=\\\!3V∈ℝ4×3V\\in\\mathbb\{R\}^\{4\\times 3\}Binarized𝒯b\\mathcal\{T\}\_\{b\}1234artificialVb∈ℝ4×3V\_\{b\}\\in\\mathbb\{R\}^\{4\\times 3\}
Figure 1:*Polytomous vs\. Binarized Tree\.*\(Left\)𝒯\\mathcal\{T\}with a polytomous root \(k=3k=3\)\. PolyILR assignsk−1=2k\-1=2basis vectors \(column ofVV\) to the root \(red\)\. \(Right\) Binarized𝒯b\\mathcal\{T\}\_\{b\}required by PhILR \(*one coordinate per node*\) introduces an artificial node \(yellow\) encoding an arbitrary grouping of leaves 1 and 2—a choice not justified by original𝒯\\mathcal\{T\}\.We introduce the geometry of compositional data and formalize the tree alignment problem \(seeAitchison \([1982](https://arxiv.org/html/2606.11646#bib.bib9)\)\)\.
### 2\.1Aitchison Geometry
Compositional data\.Compositional data are nonnegative vectors whose totals are uninformative and only relative proportions matter, e\.g\., microbial abundances \(counts with varying sequencing depth\) or chemical concentrations \(parts of a mixture\)\(Glooret al\.,[2017](https://arxiv.org/html/2606.11646#bib.bib13); Jackson,[1997](https://arxiv.org/html/2606.11646#bib.bib14)\)\. We normalize to unit sum, placing data in the open simplex:
Δd−1=\{x∈ℝ\>0d:∑i=1dxi=1\}\.\\Delta^\{d\-1\}=\\left\\\{x\\in\\mathbb\{R\}^\{d\}\_\{\>0\}:\\sum\_\{i=1\}^\{d\}x\_\{i\}=1\\right\\\}\.\(1\)The key constraint is that only ratiosxi/xjx\_\{i\}/x\_\{j\}carry information, not absolute values\.
Geometry is fixed\.Aitchison geometry\(Aitchison,[1982](https://arxiv.org/html/2606.11646#bib.bib9)\)formalizes this by equipping the setΔd−1\\Delta^\{d\-1\}with perturbationx⊕y=𝒞\(x1y1,…,xdyd\)x\\oplus y=\\mathcal\{C\}\(x\_\{1\}y\_\{1\},\\ldots,x\_\{d\}y\_\{d\}\)and poweringα⊙x=𝒞\(x1α,…,xdα\)\\alpha\\odot x=\\mathcal\{C\}\(x\_\{1\}^\{\\alpha\},\\ldots,x\_\{d\}^\{\\alpha\}\), where𝒞\(⋅\)\\mathcal\{C\}\(\\cdot\)is the closure\. Under these operations,\(Δd−1,⊕,⊙\)\(\\Delta^\{d\-1\},\\oplus,\\odot\)forms a\(d−1\)\(d\-1\)\-dimensional Hilbert space with inner product:
⟨x,y⟩A=1d∑i=1d∑j=1dlogxixjlogyiyj\.\\langle x,y\\rangle\_\{A\}=\\frac\{1\}\{d\}\\sum\_\{i=1\}^\{d\}\\sum\_\{j=1\}^\{d\}\\log\\frac\{x\_\{i\}\}\{x\_\{j\}\}\\log\\frac\{y\_\{i\}\}\{y\_\{j\}\}\.The induced Aitchison distancedA\(x,y\)=‖x⊖y‖Ad\_\{A\}\(x,y\)=\\\|x\\ominus y\\\|\_\{A\}is perturbation\-invariant:dA\(x⊕z,y⊕z\)=dA\(x,y\)d\_\{A\}\(x\\oplus z,y\\oplus z\)=d\_\{A\}\(x,y\)\. Note that this geometry is*not a modeling choice*—it is the unique structure respecting compositional invariance\.
### 2\.2ILR Basis
Basis is a choice\.The centered log\-ratio \(CLR\) transform mapsx∈Δd−1x\\in\\Delta^\{d\-1\}\(under Aitchison geometry\) isometrically to the CLR hyperplane \(i\.e\.,*Aitchison tangent space*\)
ℋ=\{z∈ℝd:𝟏⊤z=0\}\\mathcal\{H\}=\\\{z\\in\\mathbb\{R\}^\{d\}:\\mathbf\{1\}^\{\\top\}z=0\\\}\(2\)viaclr\(x\)=\(logx1/g\(x\),…,logxd/g\(x\)\)\\text\{clr\}\(x\)=\(\\log x\_\{1\}/g\(x\),\\ldots,\\log x\_\{d\}/g\(x\)\), whereg\(x\)=\(∏ixi\)1/dg\(x\)=\(\\prod\_\{i\}x\_\{i\}\)^\{1/d\}is the geometric mean\. Any matrixV∈ℝd×d−1V\\in\\mathbb\{R\}^\{d\\times d\-1\}whose*columns form an orthonormal basis*ofℋ\\mathcal\{H\}yields isometric log\-ratio \(ILR\) coordinates and admits an isometric bijection
φ\(x\)=V⊤logx,\\varphi\(x\)=V^\{\\top\}\\log x,\(3\)from\(Δd−1,⟨⋅,⋅⟩A\)\(\\Delta^\{d\-1\},\\langle\\cdot,\\cdot\\rangle\_\{A\}\)to\(ℝd−1,⟨⋅,⋅⟩2\)\(\\mathbb\{R\}^\{d\-1\},\\langle\\cdot,\\cdot\\rangle\_\{2\}\)\(Egozcueet al\.,[2003](https://arxiv.org/html/2606.11646#bib.bib1)\)\. All such basesVVare related by orthogonal transformations\. That is, they induce the same geometry but different decompositions of the simplex\. The question is thus not whether to use an ILR basis, but*which basis to choose*, and this choice largely determines interpretability\.
Basis choice controls interpretability\.This parallels classical signal processing: Fourier bases yield frequency components from translation symmetry\(Brigham,[1988](https://arxiv.org/html/2606.11646#bib.bib25)\); wavelet bases yield scale\-localized components from dyadic partitions\(Mallat,[2002](https://arxiv.org/html/2606.11646#bib.bib26)\)\. Aligning the basis with domain structure produces interpretable coordinates\.
## 3Problem Setup
Compositions come with tree\.Theddcomponents of a composition are often organized by a known rooted tree𝒯\\mathcal\{T\}, e\.g\., phylogenetic or taxonomic trees in ecology, gene ontologies in genomics\(Ashburneret al\.,[2000](https://arxiv.org/html/2606.11646#bib.bib41); Lozupone and Knight,[2005](https://arxiv.org/html/2606.11646#bib.bib39)\)\. The tree encodes domain structure: which comparisons are meaningful and at what resolution\. Hence, we seek a basisVValigned with𝒯\\mathcal\{T\}\.
Binary trees\.When𝒯\\mathcal\{T\}is binary, each internal nodeuuhas exactly two children, yielding one*contrast*: the log\-ratio of geometric means of the two descendant clades\. This is a special case of sequential binary partitioning \(SBP\)\(Egozcue and Pawlowsky\-Glahn,[2005](https://arxiv.org/html/2606.11646#bib.bib6)\), whichSilvermanet al\.\([2017](https://arxiv.org/html/2606.11646#bib.bib4)\)applied to phylogenies as*PhILR*\. PhILR is well\-matched to that setting, as phylogenies are typically inferred as bifurcating, and here, orthonormality is straightforward: contrasts at disjoint nodes have disjoint support, and contrasts at nested nodes are orthogonal because the inner contrast sums to zero on each child clade\. This construction works because binary branching imposes minimal local structure: each node requires exactly one contrast, yielding a one\-to\-one correspondence between internal nodes and basis vectors\. The global consistency problem asking that local contrasts compose into an orthonormal basis on leaves reduces to verifyingpairwiseorthogonality, which holds by support structure and zero\-sum constraints\.
What happens with polytomies?For general trees𝒯\\mathcal\{T\}, the simplicity above breaks down \(Figure[1](https://arxiv.org/html/2606.11646#S2.F1)\)\. Consider a nodeuuwithku\>2k\_\{u\}\>2children\. Comparingkuk\_\{u\}clades requiresku−1k\_\{u\}\-1orthogonal contrasts—a subspace, not a single vector\. Several challenges arise: \(i\) defining canonical local contrasts atuu, \(ii\) extending them to global vectors on leaves, and \(iii\) ensuring orthogonality across all nodes\. Standard approaches fail because subtrees of different sizes contribute unequally to inner products, as detailed shortly\.
Polytomies are common in practice\.Polytomies arise within phylogenies as both hard polytomies \(e\.g\., rapid radiations\) and soft polytomies \(e\.g\., collapsed low\-support nodes\)\(Maddison,[1989](https://arxiv.org/html/2606.11646#bib.bib62)\)\. About 64% of taxonomic branch points in the NCBI Taxonomy Database have three or more children\(Linet al\.,[2011](https://arxiv.org/html/2606.11646#bib.bib29)\)and biomedical ontologies routinely encode multi\-way groupings, e\.g\., the cell ontology\(Diehlet al\.,[2016](https://arxiv.org/html/2606.11646#bib.bib63)\)\. Such curated trees often carry*meaningful internal structure*\(e\.g\., independently named internal nodes\) that can inform downstream representations\.However, in practice, one typically*arbitrarily refines*them into binary trees\(Linet al\.,[2011](https://arxiv.org/html/2606.11646#bib.bib29)\)\. This introduces additional internal nodes and splitsnot present in the original hierarchy, making resulting coordinates and interpretations binarization\-dependent\. Alternative approaches aim to identify predictive log\-ratio features via log\-contrast regression, greedy balance selection\(Rivera\-Pintoet al\.,[2018](https://arxiv.org/html/2606.11646#bib.bib21)\), pairwise log\-ratio testing\(Mandalet al\.,[2015](https://arxiv.org/html/2606.11646#bib.bib20)\), or phylogeny\-guided ILR factors via edge selection\(Washburneet al\.,[2017](https://arxiv.org/html/2606.11646#bib.bib5)\)\. But these methods yield isolated contrasts rather than a full, node\-grouped orthonormal coordinate system canonically tied to a given multifurcating tree\.
No existing method provides a canonical, complete orthonormal decomposition for general trees that simultaneously defines local contrasts, extends them globally, and ensures consistency\. Obtaining such a decomposition without sacrificing isometry or introducing arbitrary choices is our goal\.
## 4PolyILR
u1u\_\{1\}11u2u\_\{2\}2233u3u\_\{3\}445566k1=3k\_\{1\}\\\!=\\\!3k2=2k\_\{2\}\\\!=\\\!2k3=3k\_\{3\}\\\!=\\\!3*PolyILR*u1u\_\{1\}u2u\_\{2\}u3u\_\{3\}v1v\_\{1\}v2v\_\{2\}v3v\_\{3\}v4v\_\{4\}v5v\_\{5\}112233445566
Figure 2:*PolyILR*from𝒯\\mathcal\{T\}ond=6d=6leaves toV∈ℝd×\(d−1\)V\\in\\mathbb\{R\}^\{d\\times\(d\-1\)\}\. Colors indicate generating nodes; white entries are zeros\.We describe the high\-level idea in §[4\.1](https://arxiv.org/html/2606.11646#S4.SS1), the formal construction in §[4\.2](https://arxiv.org/html/2606.11646#S4.SS2), and properties in §[4\.3](https://arxiv.org/html/2606.11646#S4.SS3)\.
### 4\.1From Hierarchy to Geometry
Goal\.We seek an orthonormal decomposition ofℋ\\mathcal\{H\}reflecting the full branching structure of𝒯\\mathcal\{T\}\. We do not just want predictive log\-ratios or useful balances, but a complete coordinate system where each coordinate corresponds to a location in𝒯\\mathcal\{T\}\. Such a basis must capture the full local structure at each node, maintain orthonormality within and across nodes, spanℋ\\mathcal\{H\}, and be canonical\. The first ensures complete encoding, the next two define a valid ILR basis in \([3](https://arxiv.org/html/2606.11646#S2.E3)\)\. And the last one ensures reproducibility\. It is not obvious such a construction exists\.
Key insight\.To address this, we \(i\) associate local geometric structure to each internal node and \(ii\) assemble them into global structure\. We attach to each node a structured object encoding all relative comparisons among its children\. The specific structure and canonicity follow from requiring the global basis to be a valid ILR basis and a deterministic choice of local basis\. We make this precise in §[4\.2](https://arxiv.org/html/2606.11646#S4.SS2)\.
\(i\) Local structure\.Consider an internal nodeuuwithkuk\_\{u\}children\. The global basis will act on compositions, comparing leaves \(via geometric mean\) within each child clade\. Since compositions carry only relative information \(as in §[2\.1](https://arxiv.org/html/2606.11646#S2.SS1)\), we encode relative differences among thesekuk\_\{u\}clades and not absolute levels\. This requiresku−1k\_\{u\}\-1degrees of freedom: one contrast distinguishes two children, two contrasts distinguish three, and so on\. Each internal node thus contributes a\(ku−1\)\(k\_\{u\}\-1\)\-dimensional structure\.
\(ii\) Global assembly\.Local contrasts live at nodes, but ILR coordinates must be global vectors on leaves\. We write a weighted inner product at each node to account for the number of descendant leaves in each child clade\. Orthonormality under this weighted inner product \(locally\) guarantees global orthonormality inℝd\\mathbb\{R\}^\{d\}after spreading to leaves\.
We now formalize our construction,PolyILR\(Figure[2](https://arxiv.org/html/2606.11646#S4.F2)\)\.
### 4\.2PolyILR Construction
Setup\.Let𝒯\\mathcal\{T\}be any rooted tree withddleaves\. Our data live in the Aitchison simplexΔd−1\\Delta^\{d\-1\}, where each component corresponds to a leaf of𝒯\\mathcal\{T\}\. Our goal is to construct a valid ILR basisV∈ℝd×\(d−1\)V\\in\\mathbb\{R\}^\{d\\times\(d\-1\)\}such that each column ofVVcorresponds to a specific internal node of𝒯\\mathcal\{T\}\.
Local contrast subspace\.Consider a nodeuuwithkuk\_\{u\}children\. We define the local contrast subspace atuuas
Su=\{h∈ℝku:∑r=1kuhr=0\}\.\\displaystyle S\_\{u\}=\\left\\\{h\\in\\mathbb\{R\}^\{k\_\{u\}\}:\\sum\_\{r=1\}^\{k\_\{u\}\}h\_\{r\}=0\\right\\\}\.
\(4\)This is the\(ku−1\)\(k\_\{u\}\-1\)\-dimensional subspace orthogonal to𝟏\\mathbf\{1\}, capturing all relative comparisons among thekuk\_\{u\}children \(see §[4\.1](https://arxiv.org/html/2606.11646#S4.SS1)\)\. Notice that this zero\-sum constraint is not arbitrary: it is, in fact, forced by the ILR requirement thatV⊤𝟏=0V^\{\\top\}\\mathbf\{1\}=0\. Since each column ofVVmust be orthogonal to𝟏\\mathbf\{1\}, and each column is formed by spreading a local vector𝐡\\mathbf\{h\}from nodeuu, we require𝐡⊤𝟏=0\\mathbf\{h\}^\{\\top\}\\mathbf\{1\}=0locally\.
Weighted inner product\.To ensure that local orthonormality extends globally, we equipSuS\_\{u\}with a*weighted inner product*\. Letnrn\_\{r\}denote the number of leaves descending from childr∈\{1,…,ku\}r\\in\\\{1,\\dots,k\_\{u\}\\\}\. We define:
⟨h,h′⟩w=∑r=1kuhrhr′nr\.\\langle h,h^\{\\prime\}\\rangle\_\{w\}=\\sum\_\{r=1\}^\{k\_\{u\}\}\\frac\{h\_\{r\}h^\{\\prime\}\_\{r\}\}\{n\_\{r\}\}\.\(5\)This accounts for unequal subtree sizes: children with more descendants contribute less per leaf to the global inner product when spread\. Note that\(Su,⟨⋅,⋅⟩w\)\(S\_\{u\},\\langle\\cdot,\\cdot\\rangle\_\{w\}\)is a Hilbert space\.
Local basis\.We choose an orthonormal basis ofSuS\_\{u\}\. Any orthonormal basis works mathematically, but we use Helmert contrasts\(Lancaster,[1965](https://arxiv.org/html/2606.11646#bib.bib18)\)for canonicity\. The standard Helmert matrixH∈ℝk×\(k−1\)H\\in\\mathbb\{R\}^\{k\\times\(k\-1\)\}has columns:
Hr,m=\{1m\(m\+1\)ifr≤m,−mm\+1ifr=m\+1,0ifr\>m\+1\.H\_\{r,m\}=\\begin\{cases\}\\sqrt\{\\frac\{1\}\{m\(m\+1\)\}\}&\\text\{if \}r\\leq m,\\\\ \-\\sqrt\{\\frac\{m\}\{m\+1\}\}&\\text\{if \}r=m\+1,\\\\ 0&\\text\{if \}r\>m\+1\.\\end\{cases\}\(6\)Themm\-th column compares childm\+1m\+1against the average of children1,…,m1,\\ldots,m\. For example, withku=3k\_\{u\}=3children:
H\(u\)=\(1216−12160−26\)\.H^\{\(u\)\}=\\begin\{pmatrix\}\\frac\{1\}\{\\sqrt\{2\}\}&\\frac\{1\}\{\\sqrt\{6\}\}\\\\\[4\.0pt\] \-\\frac\{1\}\{\\sqrt\{2\}\}&\\frac\{1\}\{\\sqrt\{6\}\}\\\\\[4\.0pt\] 0&\-\\frac\{2\}\{\\sqrt\{6\}\}\\end\{pmatrix\}\.The first column contrasts child 2 versus child 1, the second contrasts child 3 versus the average of children 1 and 2\. Helmert contrasts provide a canonical choice: given an ordering of children, the basis is deterministic\. Alternative orthonormal bases ofSuS\_\{u\}\(e\.g\., QR decomposition\) can also work but lack this sequential interpretability\.
Here, the columns ofH\(u\)H^\{\(u\)\}are orthonormal under the standard inner product and lie inSuS\_\{u\}\. To obtain orthonormality under⟨⋅,⋅⟩w\\langle\\cdot,\\cdot\\rangle\_\{w\}, we apply Gram\-Schmidt toH\(u\)H^\{\(u\)\}under this weighted inner product, yieldingH~\(u\)∈ℝku×\(ku−1\)\\widetilde\{H\}^\{\(u\)\}\\in\\mathbb\{R\}^\{k\_\{u\}\\times\(k\_\{u\}\-1\)\}\. This choice ensures that given𝒯\\mathcal\{T\}and a fixed ordering of children at each node, the local basis isuniquelyobtained\.
𝐡⊤=\(12,−12\)\\mathbf\{h\}^\{\\top\}=\\Big\(\\frac\{1\}\{\\sqrt\{2\}\}\\;,\\;\-\\frac\{1\}\{\\sqrt\{2\}\}\\Big\)uu11223344n1=1n\_\{1\}\\\!=\\\!1n2=3n\_\{2\}\\\!=\\\!3uniform𝐯u⊤=\(12,−12,−12,−12\)\\mathbf\{v\}\_\{u\}^\{\\top\}=\\Big\(\\frac\{1\}\{\\sqrt\{2\}\}\\;,\\;\-\\frac\{1\}\{\\sqrt\{2\}\}\\;,\\;\-\\frac\{1\}\{\\sqrt\{2\}\}\\;,\\;\-\\frac\{1\}\{\\sqrt\{2\}\}\\Big\)𝐡~⊤=\(32,−32\)\\tilde\{\\mathbf\{h\}\}^\{\\top\}=\\Big\(\\frac\{\\sqrt\{3\}\}\{2\}\\;,\\;\-\\frac\{\\sqrt\{3\}\}\{2\}\\Big\)uu11223344n1=1n\_\{1\}\\\!=\\\!1n2=3n\_\{2\}\\\!=\\\!3÷nr\\div n\_\{r\}𝐯u⊤=\(32,−36,−36,−36\)\\mathbf\{v\}\_\{u\}^\{\\top\}=\\Big\(\\frac\{\\sqrt\{3\}\}\{2\}\\;,\\;\-\\frac\{\\sqrt\{3\}\}\{6\}\\;,\\;\-\\frac\{\\sqrt\{3\}\}\{6\}\\;,\\;\-\\frac\{\\sqrt\{3\}\}\{6\}\\Big\)
Figure 3:*Naive vs\. weighted spreading \(Example[4\.1](https://arxiv.org/html/2606.11646#S4.Thmexample1)\)\.*Uniform spreading \(Left\) vs\. Weighted spreading \(Right\)#### Spreading to leaves\.
Each column of the local basisH~\(u\)\\widetilde\{H\}^\{\(u\)\}is a vector inℝku\\mathbb\{R\}^\{k\_\{u\}\}, defined on the children ofuu\. We spread it to a global vectorv∈ℝdv\\in\\mathbb\{R\}^\{d\}on all leaves:
vi=\{H~r,m\(u\)/nrif leafidescends from childr,0otherwise\.v\_\{i\}=\\begin\{cases\}\\widetilde\{H\}^\{\(u\)\}\_\{r,m\}/n\_\{r\}&\\text\{if leaf \}i\\text\{ descends from child \}r,\\\\ 0&\\text\{otherwise\.\}\\end\{cases\}Division bynrn\_\{r\}is key, as it ensures that orthonormality under⟨⋅,⋅⟩w\\langle\\cdot,\\cdot\\rangle\_\{w\}at nodeuuimplies orthonormality inℝd\\mathbb\{R\}^\{d\}after spreading\. Consider two local vectors𝐡,𝐡′\\mathbf\{h\},\\mathbf\{h\}^\{\\prime\}orthonormal under⟨⋅,⋅⟩w\\langle\\cdot,\\cdot\\rangle\_\{w\}:∑rhrhr′/nr=δ𝐡,𝐡′\\sum\_\{r\}h\_\{r\}h^\{\\prime\}\_\{r\}/n\_\{r\}=\\delta\_\{\\mathbf\{h\},\\mathbf\{h\}^\{\\prime\}\}\. After spreading with the1/nr1/n\_\{r\}weighting, their global inner product becomes:
⟨𝐯,𝐯′⟩\\displaystyle\\langle\\mathbf\{v\},\\mathbf\{v\}^\{\\prime\}\\rangle=∑i=1dvivi′=∑r=1ku∑i∈Cr\(u\)\(hrnr\)\(hr′nr\)\\displaystyle=\\sum\_\{i=1\}^\{d\}v\_\{i\}v^\{\\prime\}\_\{i\}=\\sum\_\{r=1\}^\{k\_\{u\}\}\\sum\_\{i\\in C\_\{r\}^\{\(u\)\}\}\\left\(\\frac\{h\_\{r\}\}\{n\_\{r\}\}\\right\)\\left\(\\frac\{h^\{\\prime\}\_\{r\}\}\{n\_\{r\}\}\\right\)=∑r=1kuhrhr′nr2⋅nr=∑r=1kuhrhr′nr=δ𝐡,𝐡′,\\displaystyle=\\sum\_\{r=1\}^\{k\_\{u\}\}\\frac\{h\_\{r\}h^\{\\prime\}\_\{r\}\}\{n\_\{r\}^\{2\}\}\\cdot n\_\{r\}=\\sum\_\{r=1\}^\{k\_\{u\}\}\\frac\{h\_\{r\}h^\{\\prime\}\_\{r\}\}\{n\_\{r\}\}=\\delta\_\{\\mathbf\{h\},\\mathbf\{h\}^\{\\prime\}\},where the second equality holds because childrrcontributes exactlynrn\_\{r\}leaves, each with coefficienthr/nrh\_\{r\}/n\_\{r\}\(similarly for𝐡′\\mathbf\{h\}^\{\\prime\}\)\. So, local orthonormality means global orthonormality\.
###### Example 4\.1\.
Consideruuwith two childrenn1=1n\_\{1\}=1andn2=3n\_\{2\}=3\. The Helmert contrast is𝐡=\(1/2,−1/2\)⊤\\mathbf\{h\}=\(1/\\sqrt\{2\},\-1/\\sqrt\{2\}\)^\{\\top\}, which satisfies‖𝐡‖22=1\\\|\\mathbf\{h\}\\\|\_\{2\}^\{2\}=1, but spreading uniformly gives𝐯=\(1/2,−1/2,−1/2,−1/2\)⊤\\mathbf\{v\}=\(1/\\sqrt\{2\},\-1/\\sqrt\{2\},\-1/\\sqrt\{2\},\-1/\\sqrt\{2\}\)^\{\\top\}with‖𝐯‖2=2≠1\\\|\\mathbf\{v\}\\\|^\{2\}=2\\neq 1\. The weighted norm in \([5](https://arxiv.org/html/2606.11646#S4.E5)\) gives‖𝐡‖w2=\(1/2\)/1\+\(1/2\)/3=2/3\\\|\\mathbf\{h\}\\\|\_\{w\}^\{2\}=\(1/2\)/1\+\(1/2\)/3=2/3; normalizing gives𝐡~=\[3/2,−3/2\]⊤\\tilde\{\\mathbf\{h\}\}=\[\\sqrt\{3\}/2,\-\\sqrt\{3\}/2\]^\{\\top\}\. Spreading with÷nr\\div n\_\{r\}yields‖𝐯‖2=1\\\|\\mathbf\{v\}\\\|^\{2\}=1\. See Figure[3](https://arxiv.org/html/2606.11646#S4.F3)\.
Algorithm 1PolyILR Basis Construction0:Rooted
TTwith
ddleaves, internal node ordering
π\\pi\(DFS\)
0:ILR basis
V∈ℝd×\(d−1\)V\\in\\mathbb\{R\}^\{d\\times\(d\-1\)\}
1:
j←1j\\leftarrow 1
2:foreach internal node
uufrom
π\\pido
3:
ku←k\_\{u\}\\leftarrownumber of children of
uu
4:for
r=1,…,kur=1,\\ldots,k\_\{u\}do
5:
Cu\(r\)←C\_\{u\}^\{\(r\)\}\\leftarrowleaves descending from
rr\-th child
6:
nr←\|Cu\(r\)\|n\_\{r\}\\leftarrow\|C\_\{u\}^\{\(r\)\}\|
7:endfor
8:
𝒮u←\{𝐡∈ℝku:∑rhr=0\}\\mathcal\{S\}\_\{u\}\\leftarrow\\\{\\mathbf\{h\}\\in\\mathbb\{R\}^\{k\_\{u\}\}:\\sum\_\{r\}h\_\{r\}=0\\\}
9:
⟨𝐡,𝐡′⟩w←∑rhrhr′/nr\\langle\\mathbf\{h\},\\mathbf\{h\}^\{\\prime\}\\rangle\_\{w\}\\leftarrow\\sum\_\{r\}h\_\{r\}h^\{\\prime\}\_\{r\}/n\_\{r\}
10:
H\(u\)←H^\{\(u\)\}\\leftarrowHelmert matrix in
ℝku×\(ku−1\)\\mathbb\{R\}^\{k\_\{u\}\\times\(k\_\{u\}\-1\)\}
11:
H~\(u\)←\\widetilde\{H\}^\{\(u\)\}\\leftarrowGram\-Schmidt on
H\(u\)H^\{\(u\)\}under
⟨⋅,⋅⟩w\\langle\\cdot,\\cdot\\rangle\_\{w\}
12:for
m=1,…,ku−1m=1,\\ldots,k\_\{u\}\-1do
13:for
i=1,…,di=1,\\ldots,ddo
14:if
i∈Cu\(r\)i\\in C\_\{u\}^\{\(r\)\}for some
rrthen
15:
Vi,j←H~r,m\(u\)/nrV\_\{i,j\}\\leftarrow\\widetilde\{H\}^\{\(u\)\}\_\{r,m\}/n\_\{r\}
16:else
17:
Vi,j←0V\_\{i,j\}\\leftarrow 0
18:endif
19:endfor
20:
j←j\+1j\\leftarrow j\+1
21:endfor
22:endfor
23:return
VV
Assembling it all\.Applying this procedure at every internal node of𝒯\\mathcal\{T\}and collecting all spread vectors yields the PolyILR basisVV\. The columns ofVVare indexed by pairs\(u,m\)\(u,m\)for some internal nodeuuand a contrast indexm∈\{1,…,ku−1\}m\\in\\\{1,\\ldots,k\_\{u\}\-1\\\}\. For example, if𝒯\\mathcal\{T\}hasℓ\\ellinternal nodes ordered asu1,…,uℓu\_\{1\},\\ldots,u\_\{\\ell\}, withu1u\_\{1\}having 4 children,u2u\_\{2\}having 3,…\\ldots, anduℓu\_\{\\ell\}having 3 children, then the basis is
V=\(v1,v2,v3⏟nodeu1,v4,v5⏟nodeu2,…,vd−2,vd−1⏟nodeuℓ\)\.V=\\bigl\(\\underbrace\{v\_\{1\},v\_\{2\},v\_\{3\}\}\_\{\\text\{node \}u\_\{1\}\},\\underbrace\{v\_\{4\},v\_\{5\}\}\_\{\\text\{node \}u\_\{2\}\},\\ldots,\\underbrace\{v\_\{d\-2\},v\_\{d\-1\}\}\_\{\\text\{node \}u\_\{\\ell\}\}\\bigr\)\.See Algorithm[1](https://arxiv.org/html/2606.11646#alg1)and Figure[2](https://arxiv.org/html/2606.11646#S4.F2)\.
###### Theorem 4\.1\(PolyILR\)\.
Let𝒯\\mathcal\{T\}be any rooted tree withddleaves\. The matrixV∈ℝd×\(d−1\)V\\in\\mathbb\{R\}^\{d\\times\(d\-1\)\}from Alg\.[1](https://arxiv.org/html/2606.11646#alg1)satisfies:
1. 1\.V⊤𝟏=0V^\{\\top\}\\mathbf\{1\}=0 \(contrast property\),
2. 2\.V⊤V=Id−1V^\{\\top\}V=I\_\{d\-1\} \(orthonormality\)\.
Consequently,φ\(x\)=V⊤logx\\varphi\(x\)=V^\{\\top\}\\log xis an isometry from\(Δd−1,⟨⋅,⋅⟩A\)\(\\Delta^\{d\-1\},\\langle\\cdot,\\cdot\\rangle\_\{A\}\)to\(ℝd−1,⟨⋅,⋅⟩2\)\(\\mathbb\{R\}^\{d\-1\},\\langle\\cdot,\\cdot\\rangle\_\{2\}\)\.
*Proof idea\.*The contrast property follows from each local contrast summing to zero\. For orthonormality: vectors from the same node are orthonormal by construction; disjoint nodes have disjoint support\. Nested nodes \(one ancestor of the other\) are orthogonal because the descendant’s spread vector sums to zero on each child clade\. Full proof and algorithm details in Appendix[A](https://arxiv.org/html/2606.11646#A1)\.
Interpretation\.Given compositionxx, each coordinatezj=vj⊤logxz\_\{j\}=v\_\{j\}^\{\\top\}\\log xis a*balance*: a log\-ratio comparing geometric means of child clades at an internal node\. The transformationz=V⊤logx∈ℝd−1z=V^\{\\top\}\\log x\\in\\mathbb\{R\}^\{d\-1\}decomposesxxinto interpretable contrasts at every level of the hierarchy, see §[5](https://arxiv.org/html/2606.11646#S5)\.
### 4\.3Properties of PolyILR
Relation to existing methods\.When𝒯\\mathcal\{T\}is binary \(ku=2k\_\{u\}=2for alluu\), each node contributesonecontrast, and PolyILR reduces to PhILR \(see Appendix[A](https://arxiv.org/html/2606.11646#A1)\)\. Unlike greedy balance selection\(Rivera\-Pintoet al\.,[2018](https://arxiv.org/html/2606.11646#bib.bib21)\)or edge\-based factorization, which yield isolated contrasts, PolyILR provides a complete orthonormal basis aligned with the full tree\.*Unlike arbitrary binarization, PolyILR respects the original topology without introducing artificial splits*\(Figure[1](https://arxiv.org/html/2606.11646#S2.F1)\)\.
Uniqueness and recoverability\.PolyILR provides a*canonical*basisVValigned with𝒯\\mathcal\{T\}as follows\.
###### Proposition 4\.2\.
Given a rooted tree𝒯\\mathcal\{T\}with fixed leaf labels and child orderings, PolyILR produces a unique basisV\(𝒯\)V\(\\mathcal\{T\}\)\. Moreover,𝒯↦V\(𝒯\)\\mathcal\{T\}\\mapsto V\(\\mathcal\{T\}\)is injective:𝒯\\mathcal\{T\}can be recovered fromVVvia the clade supports of its columns\.
We point out that PolyILR’s*canonicity*rests on two structural conventions: \(i\) a child ordering at each internal node and \(ii\) a sign convention for the Helmert columns \(first nonzero entry positive\)\. These fix the representation but not the underlying geometry: different orderings yield bases related by an orthogonal transformation within each node’s block \(and permutations across blocks\) where sign flips change the orientation of contrasts\. In practice, the ordering is inherited from the input tree and held fixed\. Full proofs are in Appendix[A](https://arxiv.org/html/2606.11646#A1)\.
*Summary\.*PolyILR provides a canonical, tree\-aligned orthonormal coordinate system for the Aitchison simplex\. Each coordinate corresponds to a contrast at a specific internal node, enabling interpretable analysis at any resolution\.
## 5Structured Analysis with PolyILR
We describe how PolyILR coordinates enable structured analysis beyond what standard log\-ratio transforms provide\. Letx∈Δd−1x\\in\\Delta^\{d\-1\}be a composition with components as leaves of𝒯\\mathcal\{T\}\. The PolyILR transform yieldsz=φ\(x\)∈ℝd−1z=\\varphi\(x\)\\in\\mathbb\{R\}^\{d\-1\}\.
### 5\.1Tree\-Aligned Coordinates
Coordinate indexing\.By construction, each coordinate indexjjcorresponds bijectively to a pair\(u,m\)\(u,m\): an internal nodeuuand a contrast indexm∈\{1,…,ku−1\}m\\in\\\{1,\\ldots,k\_\{u\}\-1\\\}\. The coordinatezj=z\(u,m\)z\_\{j\}=z\_\{\(u,m\)\}is a*balance*—a log\-ratio comparing the geometric mean of leaves under childm\+1m\+1against that under children1,…,m1,\\ldots,mat nodeuu\(see \([6](https://arxiv.org/html/2606.11646#S4.E6)\)\)\. This association is intrinsic\.
Multiscale structure\.By construction, each internal nodeuucontributesku−1k\_\{u\}\-1coordinates to the basisVV\(see §[4\.2](https://arxiv.org/html/2606.11646#S4.SS2)\)\. Because these coordinates are orthonormal, the coordinates at distinct nodes span orthogonal subspaces, yielding a disjoint partition ofℝd−1\\mathbb\{R\}^\{d\-1\}indexed by tree nodes\. We can thus reason about nodeuuas a unit: do the coordinates atuujointly explain an outcome? Does variation concentrate atuu?
This node\-level partition extends to coarser groupings\. Aggregating nodes by tree depth or by subtree membership yields alternative orthogonal partitions of the same space \(Table[1](https://arxiv.org/html/2606.11646#S5.T1)\)\. PolyILR inherits them directly from the tree\. We illustrate these partitions in Figure[5](https://arxiv.org/html/2606.11646#A2.F5)\(Appendix[B](https://arxiv.org/html/2606.11646#A2)\)\.
AggregationQuestion AnsweredNodeWhich*splits*drive signal?DepthWhat*resolution*matters?SubtreeIs an entire*clade*informative?Table 1:Tree substructure aggregations enabled by PolyILR\.
### 5\.2Implications for Inference
Given data\{\(xi,yi\)\}i=1N\\\{\(x\_\{i\},y\_\{i\}\)\\\}\_\{i=1\}^\{N\}with outcomeyiy\_\{i\}, we transformzi=φ\(xi\)z\_\{i\}=\\varphi\(x\_\{i\}\)and fit any model on\{\(zi,yi\)\}\\\{\(z\_\{i\},y\_\{i\}\)\\\}\. Sinceφ\\varphiis an isometry, the full geometry is preserved\.
TaskCLRPhILRPolyILRRF / SVM / LRRF / SVM / LRRF / SVM / LR*Acc \(%\)*HMPbody sites \(5\)\.956/\.971/\.962\.961/\.971/\.962\.963/\.971/\.962body subsites \(18\)\.597/\.672/\.646\.608/\.672/\.646\.622/\.672/\.646cMD3westernized \(2\)\.972/\.979/\.968\.966/\.979/\.967\.967/\.979/\.967age category \(5\)\.785/\.814/\.739\.797/\.814/\.738\.797/\.814/\.738DISCOleukemia \(2\)\.925/\.932/\.927\.940/\.932/\.927\.935/\.932/\.927HCC \(2\)\.921/\.927/\.944\.910/\.927/\.944\.910/\.927/\.944*AUROC*HMPbody sites \(5\)\.987/\.995/\.994\.992/\.995/\.994\.992/\.995/\.994body subsites \(18\)\.957/\.974/\.967\.965/\.974/\.967\.966/\.974/\.967cMD3westernized \(2\)\.975/\.978/\.967\.966/\.978/\.966\.966/\.978/\.967age category \(5\)\.836/\.867/\.833\.837/\.867/\.833\.843/\.867/\.833DISCOleukemia \(2\)\.968/\.984/\.982\.984/\.984/\.982\.984/\.984/\.982HCC \(2\)\.921/\.996/\.998\.984/\.996/\.998\.974/\.996/\.998
TaskPolyILRPhILR \(index / semantic\)KK=5KK=10KK=50KK=5KK=10KK=50HMPbody sites \(5\)\.66\.65\.84\.01 / \.08\.01 / \.06\.07 / \.05body subsites \(18\)\.71\.72\.88\.01 / \.22\.02 / \.13\.07 / \.04cMD3westernized \(2\)\.43\.69\.81\.00 / \.01\.00 / \.01\.01 / \.02age category \(5\)\.73\.58\.76\.12 / \.04\.09 / \.03\.03 / \.03healthy vs disease \(2\)\.56\.86\.80\.00 / \.02\.00 / \.03\.02 / \.02DISCOleukemia \(2\)\.75\.81\.92\.05 / \.09\.12 / \.16\.73 / \.12HCC \(2\)\.77\.78\.84\.03 / \.07\.06 / \.09\.35 / \.11
Table 2:\(Top\)*Classification accuracy and AUROC*\(5 runs\) across CLR, PhILR, and PolyILR\. Each cell reports RF/SVM/LR\. SVM and LR match across the three ILR representations within each task, as expected from isometry; differences are confined to RF but modest\. Full statistical variability \(95% CIs\) is in Appendix[B\.7](https://arxiv.org/html/2606.11646#A2.SS7)but observed to be small\.\(Bottom\)*Feature stability*\(Jaccard of top\-KKfeatures across CV folds\)\. PolyILR stable; PhILR \(index/semantic\) unstable from arbitrary binarization\.RankContrastRF Importance \(%\)Rank rangeHMP
body sites1Streptococcaceae vs Lactobacillus \+ Leuconostocaceae3\.9412Lactococcus vs Streptococcus3\.422–33Pseudomonadales vs Cardiobacteriaceae \+ Vibrionaceae \+ Legionellales \+ …3\.402–34Bacillales vs Gemella \+ Exiguobacterium \+ Turicibacter \+ …2\.954–5cMD3
westernized1Prevotella vs Alistipes \+ Anaerotruncus \+ Bacteroides \+ …1\.7712Murimonas vs Alistipes \+ Anaerotruncus \+ Bacteroides \+ …1\.492–43Prevotella vs Bacteroides \+ Alistipes1\.452–54Lactobacillus vs Alistipes \+ Anaerotruncus \+ Bacteroides \+ …1\.384–6cMD3
health/dis\.1Lachnoclostridium vs Bacteroides \+ Turicibacter \+ Ruminococcus \+ …0\.5512Bifidobacterium vs Blautia \+ Enterococcus0\.382–53Flavonifractor vs Bifidobacterium \+ Actinomyces0\.352–64Fusicatenibacter vs Gemmiger0\.342–6DISCO
leukemia1Myeloid cell vs Erythrocyte/Megakaryocyte \+ Hematopoietic precursor cell14\.5012Cycling T/NK cell vs ILC \+ T cell \+ NK cell11\.7923Lymphoid cell vs Erythrocyte/Megakaryocyte \+ Hematopoietic precursor \+ Myeloid cell6\.303–44MAIT cell vs Naive T cell \+ Memory T cell \+ CD8 T cell \+ …5\.453–5DISCO
HCC1Erythrocyte/Megakaryocyte vs Myeloid cell \+ Lymphoid cell \+ Hematopoietic precursor cell10\.0812B cell precursor vs Plasma cell \+ INF\-activated naive B cell \+ Memory B cell \+ …6\.522–43Hematopoietic precursor cell vs Myeloid cell \+ Lymphoid cell6\.342–34Venous EC vs LSEC5\.452–8Table 3:Top\-4 PolyILR contrasts by RF importance \(%\), with rank range across 5\-fold CV\. Each contrast is a coordinate representing the log\-ratio of geometric means between two groups\.“A vs B \+ C” means the log\-ratio of geometric means of A’s descendants against the pooled descendants of B and C; “\+ …” marks additional siblings omitted\.Full importance variability \(mean±\\pmstd\) is in Appendix[B\.7](https://arxiv.org/html/2606.11646#A2.SS7)\.Feature selection\.Identifying*which features drive outcomes*is a key scientific goal\. In genomics, neuroscience, and microbiome alike, the goal is often not just prediction but understanding which variables matter and why\(Rudin,[2019](https://arxiv.org/html/2606.11646#bib.bib42); Marcos\-Zambranoet al\.,[2021](https://arxiv.org/html/2606.11646#bib.bib43)\)\. With standard ILR, important coordinates are anonymous indices with no semantics\. With PolyILR, when coordinatej=\(u,m\)j=\(u,m\)is identified as important, we know which node and contrast drive the signal: a log\-ratio comparing specific groups of leaves\. This interpretability is intrinsic to the representation, no post\-hoc processing needed\.
Tree\-level aggregation\.The multiscale structure of PolyILR enables inference at any tree substructure\. Letωj\\omega\_\{j\}denote importance of coordinatejjfrom any method \(e\.g\., random forest\)\. We aggregate per\-coordinate importances over any disjoint setSSof coordinates \(node, depth, or subtree\) viaω\(S\)=∑j∈Sωj\\omega\(S\)=\\sum\_\{j\\in S\}\\omega\_\{j\}\. This aggregation is well\-defined because coordinates at distinct nodes span orthogonal subspaces, any partition of the tree into disjoint substructures yields a partition ofℝd−1\\mathbb\{R\}^\{d\-1\}, and the corresponding importances sum to the total without double\-counting\.
Leaf\-level importance\.To quantify importance of an individual leafℓ\\ell\(e\.g\., a taxon\), we cannot directly aggregate coordinates because coordinates are not exclusive to any single leaf\. Instead, we can distribute importance weighted by participation\. By construction,VℓjV\_\{\\ell j\}quantifies how much leafℓ\\ellparticipates in coordinatejj\. We defineω\(ℓ\)=∑j=1d−1Vℓj2⋅ωj\\omega\(\\ell\)=\\sum\_\{j=1\}^\{d\-1\}V\_\{\\ell j\}^\{2\}\\cdot\\omega\_\{j\}\. Since columns ofVVare unit vectors,∑ℓVℓj2=1\\sum\_\{\\ell\}V\_\{\\ell j\}^\{2\}=1, so leaf importances sum to total importance\.
## 6Experiments
We evaluate PolyILR on standard microbiome and single\-cell benchmarks\. Our goals are to demonstrate PolyILR provides:\(G1\)valid ILR representations \(§[6\.2](https://arxiv.org/html/2606.11646#S6.SS2)\),\(G2\)stable feature selection, unlike PhILR with arbitrary binarization \(§[6\.3](https://arxiv.org/html/2606.11646#S6.SS3)\),\(G3\)interpretable features grounded in the tree \(§[6\.4](https://arxiv.org/html/2606.11646#S6.SS4)\), and\(G4\)structured inference by tree subparts \(§[6\.5](https://arxiv.org/html/2606.11646#S6.SS5)\)\.
### 6\.1Setup
Datasets\.We use three large datasets from two domains\. For microbiome:HMP\(Human Microbiome Project; 4,743 samples, 402 taxa\)\(Human Microbiome Project Consortium,[2012](https://arxiv.org/html/2606.11646#bib.bib28)\)andcMD3\(curatedMetagenomicData v3; 20,238 samples from 86 studies, 2,047 taxa\)\(Pasolliet al\.,[2017](https://arxiv.org/html/2606.11646#bib.bib27)\), with taxonomies from NCBI\. For single\-cell biology:DISCO\(Database of Immune Single\-Cell Omics; 751 sample\-level composition profiles derived from∼\{\\sim\}5\.3M cells, 62–99 cell types\)\(Liet al\.,[2022](https://arxiv.org/html/2606.11646#bib.bib49)\), with cell types organized by the Cell Ontology\. Tasks include predictions on body site \(5–18 sites\), westernization, age category, healthy vs\. disease \(microbiome\), and healthy vs\. leukemia/HCC \(single\-cell\)\.As with any log\-ratio method, PolyILR requires zero handling before the coordinate transform; we use a small additive pseudocount per dataset \(Appendix[B\.4](https://arxiv.org/html/2606.11646#A2.SS4)\)\.
Methods\.We compare PolyILR against CLR \(same geometry, no tree alignment\) andunweightedPhILR with random binarization \(tree\-aligned, butdefined on binary trees, so polytomies must be resolved arbitrarily\)\. We use random forest \(RF\), SVM, and logistic regression \(LR\) with 5\-fold cross\-validation\.Additional experiments, hyperparameter, and dataset construction details are in Appendix[B](https://arxiv.org/html/2606.11646#A2)\.
### 6\.2Representation Validity
We verify PolyILR is a geometrically valid representation\(G1\)\. CLR projects compositions into the tangent spaceℋ\\mathcal\{H\}, spanned by PolyILR coordinates \(Thm\.[4\.1](https://arxiv.org/html/2606.11646#S4.Thmtheorem1)\)\. Table[2](https://arxiv.org/html/2606.11646#S5.T2)\(top\) supports this equivalence:CLR, PhILR, and PolyILR yield identical SVM and LR accuracy, as expected from isometry\. RF accuracy varies modestly across representations \(within±\\pm1\.5% on most tasks\), since RF is sensitive to the choice of axes; differences in either direction are consistent with the geometric equivalence\.
### 6\.3Feature selection is stable
A key advantage of PolyILR over PhILR is*canonical*decomposition on any tree\(G2\)\. PhILR requires binarizing polytomies, making feature selection unstable across binarizations with*no correct choice*\. We measure stability via: \(i\)*index stability*\(Jaccard similarity of top\-KKfeature indices across runs\) measuring if the same coordinate positions are selected; and \(ii\)*semantic stability*\(similarity of the corresponding taxonomic/ontological contrasts\) measuring whether selected features represent the same biological comparisons regardless of index\. The latter is fairer to PhILR as it ignores arbitrary index assignment\. For PhILR, we vary the random binarization of the same polytomous tree across runs; for PolyILR, \(i\) and \(ii\) coincide since coordinates are canonical given the tree\. Table[2](https://arxiv.org/html/2606.11646#S5.T2)\(bottom\) shows PolyILR achieves high stability \(0\.43–0\.92\) while PhILR collapses \(near 0\) under both metrics across all three datasets\. Even when comparing semantically, PhILR’s artificial binary splits yield different partitions across binarizations, confirming that the instability is structural \(Figure[1](https://arxiv.org/html/2606.11646#S2.F1)\)\.
### 6\.4Features are interpretable
PolyILR coordinates are directly interpretable as taxonomic/ontological contrasts\(G3\)\. Table[3](https://arxiv.org/html/2606.11646#S5.T3)shows the top\-4 features by RF importance, with rank range across 5 runs indicating stability of the ranking\. Each feature is a log\-ratio contrast between groups at a specific node \(§[5\.2](https://arxiv.org/html/2606.11646#S5.SS2)\)\.
*Scientific interpretation \(Table[3](https://arxiv.org/html/2606.11646#S5.T3)\)\.*The recovered contrasts are consistent with previously reported observations in the literature\. For HMP body sites, Streptococcus vs\. Lactococcus \(3\.4%\) reflects known niche specialization within Streptococcaceae across oral subsites\(Human Microbiome Project Consortium,[2012](https://arxiv.org/html/2606.11646#bib.bib28); Dewhirstet al\.,[2010](https://arxiv.org/html/2606.11646#bib.bib30)\), with Lactococcus lactis reported as a prevalent lactic\-acid bacterium in the gut\(Pasolliet al\.,[2020](https://arxiv.org/html/2606.11646#bib.bib64)\)\. For westernization, Prevotella vs\. Bacteroides/Alistipes \(1\.5–1\.8%\) captures the lifestyle axis, with Prevotella enriched in non\-Western populations consuming plant\-rich diets\(De Filippoet al\.,[2010](https://arxiv.org/html/2606.11646#bib.bib31); Yatsunenkoet al\.,[2012](https://arxiv.org/html/2606.11646#bib.bib32)\)\. For healthy vs\. disease, Lachnoclostridium \(0\.6%\) aligns with documented links to colorectal cancer and atherosclerosis\(Caiet al\.,[2022](https://arxiv.org/html/2606.11646#bib.bib33); Lianget al\.,[2020](https://arxiv.org/html/2606.11646#bib.bib34)\)\. For leukemia, myeloid vs\. erythroid/precursor imbalance \(14\.5%\) reflects lineage disruption in hematological malignancies\(Löwenberget al\.,[1999](https://arxiv.org/html/2606.11646#bib.bib50)\)\. For HCC, venous EC vs\. LSEC \(5\.5%\) captures the well\-documented dedifferentiation of liver sinusoidal endothelial cells in hepatocellular carcinoma\(Sørensenet al\.,[2015](https://arxiv.org/html/2606.11646#bib.bib51)\)\.
*Geometric structure\.*Figure[4](https://arxiv.org/html/2606.11646#S6.F4)projectsHMPsamples onto the top\-2 coordinates\. Unlike PCA,*each axis here is a single interpretable contrast*rather than a linear combination of all features\. We do not claim maximal variance explained; rather, biologically meaningful features alone suffice to separate body sites\. The linear substructures within classes may reflect shared sparsity: samples with identical zero\-count taxa map to parallel manifolds in ILR space\.
Figure 4:Projection onto top\-2 PolyILR features forHMPbody site classification \(left\) andDISCOleukemia classification \(right\)\. Axes are the two most predictive contrasts \(Table[3](https://arxiv.org/html/2606.11646#S5.T3)\)\.body sites \(5\)body subsites \(18\)StructureComponentAccImp\.ComponentAccImp\.Depth≤\\leq0 \(coarsest\)\.876\.8%≤\\leq0 \(coarsest\)\.424\.8%≤\\leq2\.9647%≤\\leq2\.5944%≤\\leq4 \(all\)\.96100%≤\\leq4 \(all\)\.62100%SubtreeFirmicutes\.9547%Firmicutes\.5536%Proteobacteria\.8720%Proteobacteria\.4426%Actinobacteria\.9119%Actinobacteria\.4522%NodeActinomycetales–11%Actinomycetales–12%Lactobacillales–9\.1%Lachnospiraceae–8\.4%Gammaproteobacteria–8\.0%Clostridiales–3\.8%TaxonStreptococcus–3\.1%Oribacterium–1\.0%Lactococcus–3\.1%Corynebacterium–1\.0%Pasteurella–2\.6%Lactococcus–1\.0%Table 4:Tree\-level inference onHMP\. Importance aggregated by depth \(cumulative\), subtree \(phylum\), node, and taxon\.westernized \(2\)healthy/disease \(2\)StructureComponentAcc\.Imp\.ComponentAcc\.Imp\.Depth≤\\leq0 \(coarsest\)\.9340\.1%≤\\leq0 \(coarsest\)\.6130\.3%≤\\leq3\.9565\.3%≤\\leq3\.6676\.4%≤\\leq6 \(all\)\.967100%≤\\leq6 \(all\)\.682100%TaxonPrevotella–1\.7%Lachnoclostridium–0\.6%Murimonas–1\.4%Bifidobacterium–0\.3%Bacteroides–1\.3%Enterococcus–0\.3%Table 5:Tree\-level inference oncMD3\. Importance aggregated by depth \(cumulative\) and taxon\.leukemia \(2\)HCC \(2\)StructureComponentAcc\.Imp\.ComponentAcc\.Imp\.Depth≤\\leq0 \(coarsest\)\.624\.6%≤\\leq0 \(coarsest\)\.888\.1%≤\\leq1\.8826%≤\\leq1\.9138%≤\\leq6 \(all\)\.94100%≤\\leq6 \(all\)\.91100%SubtreeImmune cell\.9395%Immune cell\.9279%–––Endothelial cell\.807\.8%–––Epithelial cell\.784\.1%NodeImmune cell–22%Immune cell–19%T cell–16%Hema\. precursor–9\.6%T/NK cell–14%Endothelial cell–7\.8%Cell typeCycling T/NK cell–11\.3%GMP–4\.4%MAIT cell–5\.0%Erythroblast \(int\.\)–4\.2%Naive CD8 T cell–4\.1%Erythroblast \(late\)–4\.2%Table 6:Tree\-level inference onDISCO\. Importance aggregated by depth \(cumulative\), subtree, node, and cell type\.
### 6\.5Tree\-Level Inference
PolyILR enables structured hypothesis testing at multiple resolutions\(G4\)\. RF importance can be aggregated by depth, subtree, node, or leaf \(see §[5\.2](https://arxiv.org/html/2606.11646#S5.SS2)\)\. We report all four levels forHMP\(Table[4](https://arxiv.org/html/2606.11646#S6.T4)\) andDISCO\(Table[6](https://arxiv.org/html/2606.11646#S6.T6)\), but only depth and taxon forcMD3\(Table[5](https://arxiv.org/html/2606.11646#S6.T5)\) whose meta\-analytic tree lacks consistent intermediate labels\. Subtree\-level partitions by root’s children \(root omitted\)\.
*Scientific interpretation \(Tables[4](https://arxiv.org/html/2606.11646#S6.T4)–[6](https://arxiv.org/html/2606.11646#S6.T6)\)\.*Aggregations agree with known structure\. For HMP, Firmicutes \(47%\) and Proteobacteria \(19%\) dominate body site signals\(Human Microbiome Project Consortium,[2012](https://arxiv.org/html/2606.11646#bib.bib28); Costelloet al\.,[2009](https://arxiv.org/html/2606.11646#bib.bib35); Maet al\.,[2024](https://arxiv.org/html/2606.11646#bib.bib61)\)\. At node level, Actinomycetales \(11%\) and Lactobacillales \(9%\) capture skin vs\. oral distinctions\. For cMD3 westernization, coarse contrasts \(≤\\leq3\) achieve 95\.6% acc\., consistent with diet\-associated shifts at coarse resolution\(De Filippoet al\.,[2010](https://arxiv.org/html/2606.11646#bib.bib31); Arumugamet al\.,[2011](https://arxiv.org/html/2606.11646#bib.bib37)\)\. For DISCO leukemia \(single\-cell\), 95% of importance concentrates in Immune cells, with T cell \(16%\) and T/NK cell \(14%\) nodes dominating\(Löwenberget al\.,[1999](https://arxiv.org/html/2606.11646#bib.bib50)\)\. For HCC, importance distributes across Immune \(79%\), Endothelial \(8%\), and Epithelial \(4%\) subtrees, reflecting multi\-compartment remodeling\(Sørensenet al\.,[2015](https://arxiv.org/html/2606.11646#bib.bib51)\)\.
We should note that the biological interpretations above are*plausibility checks consistent with prior literature*\. Any causal or clinical conclusions will require much deeper analyses beyond the scope of this methodological work\.
In summary, PolyILR addresses all goals\(G1–G4\)while recovering features consistent with known biomarkers\.
## 7Beyond Compositional Data
We establish a connection between compositional data and probabilistic modeling via shared underlying geometry\. Further analysis and validation may be of independent interest\.
Aitchison geometry as quotient\.Compositional data identifies vectors up to equivalence classes\[𝐜\]\[\\mathbf\{c\}\]induced by𝐜∼cλ𝐜\\mathbf\{c\}\\sim\_\{c\}\\lambda\\mathbf\{c\}forλ\>0\\lambda\>0\(i\.e\., scaling\), since only ratios carry information\. We observe that the quotientℝ\>0d/∼c\\mathbb\{R\}^\{d\}\_\{\>0\}/\{\\sim\_\{c\}\}is the Aitchison simplex, whose tangent space isℋ\\mathcal\{H\}via CLR\(Aitchison,[1982](https://arxiv.org/html/2606.11646#bib.bib9); Egozcueet al\.,[2003](https://arxiv.org/html/2606.11646#bib.bib1)\)\.
Probabilistic modeling\.Consider a modelfθf\_\{\\theta\}outputting logits𝐳=fθ\(𝐱\)∈ℝd\\mathbf\{z\}=f\_\{\\theta\}\(\\mathbf\{x\}\)\\in\\mathbb\{R\}^\{d\}, with predicted distribution𝐩=softmax\(𝐳\)\\mathbf\{p\}=\\mathrm\{softmax\}\(\\mathbf\{z\}\)trained via cross\-entropy\. Since softmax is*shift\-invariant*, i\.e\.,softmax\(𝐳\+c𝟏\)=softmax\(𝐳\)\\mathrm\{softmax\}\(\\mathbf\{z\}\+c\\mathbf\{1\}\)=\\mathrm\{softmax\}\(\\mathbf\{z\}\), this induces an equivalence relation𝐳∼ℓ𝐳\+c𝟏\\mathbf\{z\}\\sim\_\{\\ell\}\\mathbf\{z\}\+c\\mathbf\{1\}\. The loss and predictions are invariant to shifts along𝟏\\mathbf\{1\}\.
###### Proposition 7\.1\.
Letℒ=ℝd/∼ℓ\\mathcal\{L\}=\\mathbb\{R\}^\{d\}/\{\\sim\_\{\\ell\}\}be the quotient of logits under shift equivalence\. Thenℒ≅ℋ\\mathcal\{L\}\\cong\\mathcal\{H\}\(isomorphism\)\. For𝐩=softmax\(𝐳\)\\mathbf\{p\}=\\mathrm\{softmax\}\(\\mathbf\{z\}\), we have𝐳−z¯𝟏=clr\(𝐩\),wherez¯=1d∑izi\\mathbf\{z\}\-\\bar\{z\}\\mathbf\{1\}=\\mathrm\{clr\}\(\\mathbf\{p\}\),\\quad\\text\{where \}\\bar\{z\}=\\tfrac\{1\}\{d\}\\textstyle\\sum\_\{i\}z\_\{i\}, and\[𝐳\]→𝐳−z¯𝟏\[\\mathbf\{z\}\]\\to\\mathbf\{z\}\-\\bar\{z\}\\mathbf\{1\}is well\-defined\.
The individual components \(i\.e\., the CLR hyperplane, centering map\(Egozcueet al\.,[2003](https://arxiv.org/html/2606.11646#bib.bib1)\), and softmax shift\-invariance\) are well\-known\. Proposition[7\.1](https://arxiv.org/html/2606.11646#S7.Thmtheorem1)newly establishes thatafter quotienting by the shift symmetry, logit space aligns with the CLR/Aitchison tangent space\.
Implications\.Many datasets have tree structure over classes, e\.g\., a superclass hierarchy forCIFAR\-100and WordNet forImageNet\(Denget al\.,[2009](https://arxiv.org/html/2606.11646#bib.bib44); Miller,[1995](https://arxiv.org/html/2606.11646#bib.bib45)\)\. Given such a tree𝒯\\mathcal\{T\}, PolyILR can transform model logits as𝐚=𝐕⊤𝐳\\mathbf\{a\}=\\mathbf\{V\}^\{\\top\}\\mathbf\{z\}in tree\-aligned coordinates where each component corresponds to a node and contrast\. Since𝐩=softmax\(𝐕𝐚\)\\mathbf\{p\}=\\mathrm\{softmax\}\(\\mathbf\{V\}\\mathbf\{a\}\), predictions can be analyzed in this interpretable space\. For instance, the gradient w\.r\.t\. logits is∇𝐳ℓ=𝐩−𝐞y\\nabla\_\{\\mathbf\{z\}\}\\ell=\\mathbf\{p\}\-\\mathbf\{e\}\_\{y\}, with one\-hot target𝐞y\\mathbf\{e\}\_\{y\}\. The gradient w\.r\.t\. PolyILR coordinates thus is∇𝐚ℓ=𝐕⊤∇𝐳ℓ=𝐕⊤\(𝐩−𝐞y\)\\nabla\_\{\\mathbf\{a\}\}\\ell=\\mathbf\{V\}^\{\\top\}\\nabla\_\{\\mathbf\{z\}\}\\ell=\\mathbf\{V\}^\{\\top\}\(\\mathbf\{p\}\-\\mathbf\{e\}\_\{y\}\)\. Each\(∇𝐚ℓ\)\(u,m\)\(\\nabla\_\{\\mathbf\{a\}\}\\ell\)\_\{\(u,m\)\}measures how strongly the loss pushes probability mass along contrastmmat nodeuu, localizing model errors in𝒯\\mathcal\{T\}\. See Appendix[A](https://arxiv.org/html/2606.11646#A1)and[B\.6](https://arxiv.org/html/2606.11646#A2.SS6)for proof and preliminary results\. Practical applications to model training or analysis remain an open direction\.
## 8Related Work
Compositional hierarchy methods\.ILR transform provides orthonormal coordinates for compositional data\(Egozcueet al\.,[2003](https://arxiv.org/html/2606.11646#bib.bib1)\)\.PhILR\(Silvermanet al\.,[2017](https://arxiv.org/html/2606.11646#bib.bib4)\)aligns this transform with phylogenetic trees, the setting it was designed for, where binary topology is the standard convention; applying it to polytomous trees requires arbitrary binary resolution\.UniFrac\(Lozupone and Knight,[2005](https://arxiv.org/html/2606.11646#bib.bib39)\)incorporates phylogenetic information but produces a dissimilarity measure, not a coordinate system\. Phylofactorization\(Washburneet al\.,[2017](https://arxiv.org/html/2606.11646#bib.bib5)\)and selbal\(Rivera\-Pintoet al\.,[2018](https://arxiv.org/html/2606.11646#bib.bib21)\)target biomarker discovery andidentify predictive balances via greedy selection, but yield task\-specific contrasts rather than a complete basis\. Dirichlet\-tree models\(Mao and Ma,[2022](https://arxiv.org/html/2606.11646#bib.bib40)\)use tree structure for clustering but operate probabilistically, not geometrically\. PolyILR complements this line of work by providing a complete orthonormal decomposition for*any*tree\.Other work compares proportion\-based and compositional normalizations\(Yerkeet al\.,[2024](https://arxiv.org/html/2606.11646#bib.bib65)\)\.
Trees and geometric representations\.A separate line of work studies geometric representations of trees themselves\. Hyperbolic embeddings learn representations of hierarchical data in spaces of constant negative curvature\(Nickel and Kiela,[2017](https://arxiv.org/html/2606.11646#bib.bib54); Chamiet al\.,[2019](https://arxiv.org/html/2606.11646#bib.bib55); Salaet al\.,[2018](https://arxiv.org/html/2606.11646#bib.bib60)\), while tropical geometry and BHV tree space study geodesics and statistics over spaces*of*trees\(Billeraet al\.,[2001](https://arxiv.org/html/2606.11646#bib.bib56); Owen and Provan,[2010](https://arxiv.org/html/2606.11646#bib.bib57); Monodet al\.,[2018](https://arxiv.org/html/2606.11646#bib.bib58)\)\. These embed trees or treat them as data; PolyILR differs in that the tree is a fixed input that structures a decomposition of the data space\.
## 9Conclusion
We introduced PolyILR, a canonical orthonormal decomposition of the Aitchison simplex for arbitrary tree topologies, including polytomous ones common in real taxonomies and ontologies\. The construction equips each internal node with a weighted local geometry that assembles into a global ILR basis, addressing the incompatibility between Aitchison geometry and polytomous hierarchies\. Experiments on microbiome and single\-cell data show stable, interpretable features with inference at multiple tree resolutions\. A connection to softmax classifiers suggests applications to hierarchical probabilistic modeling\.
Limitations\.PolyILR requires a*known, fixed tree*as input and does not accommodate topological uncertainty \(e\.g\., bootstrap support, posterior distributions over trees\)\. We also know that external trees such as taxonomies and phylogenies may also contain noise\. We provide a robustness analysis under nearest\-neighbor interchange perturbations in Appendix[B\.5](https://arxiv.org/html/2606.11646#A2.SS5), and view extensions that propagate tree uncertainty into the coordinate system \(e\.g\., support\-weighted local geometries\) as future work\. PolyILR’s coordinates are also only locally interpretable by construction, as those under different internal nodes live in distinct weighted subspaces and are not directly comparable, though cross\-subtree summaries are recovered via tree\-substructure aggregation in Section[5\.2](https://arxiv.org/html/2606.11646#S5.SS2)\. Finally, like all log\-ratio methods, PolyILR requires zero replacement as preprocessing \(Section[6\.1](https://arxiv.org/html/2606.11646#S6.SS1); Appendix[B\.4](https://arxiv.org/html/2606.11646#A2.SS4)\)\.
## Acknowledgments
We thank the anonymous reviewers for their constructive feedback\. We are grateful to Prof\. Colin Dewey and Prof\. Christina Kendziorski for discussions and feedback\. Authors were all partly supported by NIH R01AG092220\.
## Impact Statement
This paper presents work whose goal is to advance the field of Machine Learning\. There are many potential societal consequences of our work, none which we feel must be specifically highlighted here\.
## References
- J\. Aitchison \(1982\)The statistical analysis of compositional data\.Journal of the Royal Statistical Society: Series B \(Methodological\)44\(2\),pp\. 139–160\.Cited by:[§1](https://arxiv.org/html/2606.11646#S1.p2.1),[§2\.1](https://arxiv.org/html/2606.11646#S2.SS1.p2.6),[§2](https://arxiv.org/html/2606.11646#S2.p1.1),[§7](https://arxiv.org/html/2606.11646#S7.p2.5)\.
- M\. Arumugam, J\. Raes, E\. Pelletier, D\. Le Paslier, T\. Yamada, D\. R\. Mende, G\. R\. Fernandes, J\. Tap, T\. Bruls, J\. Batto,et al\.\(2011\)Enterotypes of the human gut microbiome\.Nature473\(7346\),pp\. 174–180\.Cited by:[§6\.5](https://arxiv.org/html/2606.11646#S6.SS5.p2.1)\.
- M\. Ashburner, C\. A\. Ball, J\. A\. Blake, D\. Botstein, H\. Butler, J\. M\. Cherry, A\. P\. Davis, K\. Dolinski, S\. S\. Dwight, J\. T\. Eppig,et al\.\(2000\)Gene ontology: tool for the unification of biology\.Nature Genetics25\(1\),pp\. 25–29\.Cited by:[§3](https://arxiv.org/html/2606.11646#S3.p1.4)\.
- L\. J\. Billera, S\. P\. Holmes, and K\. Vogtmann \(2001\)Geometry of the space of phylogenetic trees\.Advances in Applied Mathematics27\(4\),pp\. 733–767\.Cited by:[§8](https://arxiv.org/html/2606.11646#S8.p2.1)\.
- D\. Billheimer, P\. Guttorp, and W\. F\. Fagan \(2001\)Statistical interpretation of species composition\.Journal of the American Statistical Association96\(456\),pp\. 1205–1214\.Cited by:[§1](https://arxiv.org/html/2606.11646#S1.p1.1)\.
- E\. O\. Brigham \(1988\)The Fast Fourier Transform and Its Applications\.Prentice\-Hall, Inc\.\.Cited by:[§2\.2](https://arxiv.org/html/2606.11646#S2.SS2.p2.1)\.
- M\. Buettner, J\. Ostner, C\. L\. Mueller, F\. J\. Theis, and B\. Schubert \(2021\)scCODA is a Bayesian model for compositional single\-cell data analysis\.Nature Communications12\(1\),pp\. 6876\.Cited by:[§B\.2](https://arxiv.org/html/2606.11646#A2.SS2.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2606.11646#S1.p1.1)\.
- Y\. Cai, F\. Huang, X\. Lao, Y\. Lu, X\. Gao, R\. N\. Alolga, K\. Yin, X\. Zhou, Y\. Wang, B\. Liu,et al\.\(2022\)Integrated metagenomics identifies a crucial role for trimethylamine\-producing lachnoclostridium in promoting atherosclerosis\.NPJ Biofilms and Microbiomes8\(1\),pp\. 11\.Cited by:[§6\.4](https://arxiv.org/html/2606.11646#S6.SS4.p2.1)\.
- I\. Chami, Z\. Ying, C\. Ré, and J\. Leskovec \(2019\)Hyperbolic graph convolutional neural networks\.Advances in neural information processing systems32\.Cited by:[§8](https://arxiv.org/html/2606.11646#S8.p2.1)\.
- E\. K\. Costello, C\. L\. Lauber, M\. Hamady, N\. Fierer, J\. I\. Gordon, and R\. Knight \(2009\)Bacterial community variation in human body habitats across space and time\.Science326\(5960\),pp\. 1694–1697\.Cited by:[§6\.5](https://arxiv.org/html/2606.11646#S6.SS5.p2.1)\.
- C\. De Filippo, D\. Cavalieri, M\. Di Paola, M\. Ramazzotti, J\. B\. Poullet, S\. Massart, S\. Collini, G\. Pieraccini, and P\. Lionetti \(2010\)Impact of diet in shaping gut microbiota revealed by a comparative study in children from europe and rural africa\.Proceedings of the National Academy of Sciences107\(33\),pp\. 14691–14696\.Cited by:[§6\.4](https://arxiv.org/html/2606.11646#S6.SS4.p2.1),[§6\.5](https://arxiv.org/html/2606.11646#S6.SS5.p2.1)\.
- J\. Deng, W\. Dong, R\. Socher, L\. Li, K\. Li, and L\. Fei\-Fei \(2009\)ImageNet: a large\-scale hierarchical image database\.In2009 IEEE conference on computer vision and pattern recognition,pp\. 248–255\.Cited by:[§7](https://arxiv.org/html/2606.11646#S7.p5.10)\.
- F\. E\. Dewhirst, T\. Chen, J\. Izard, B\. J\. Paster, A\. C\. Tanner, W\. Yu, A\. Lakshmanan, and W\. G\. Wade \(2010\)The human oral microbiome\.Journal of Bacteriology192\(19\),pp\. 5002–5017\.Cited by:[§6\.4](https://arxiv.org/html/2606.11646#S6.SS4.p2.1)\.
- A\. D\. Diehl, T\. F\. Meehan, Y\. M\. Bradford, M\. H\. Brush, W\. M\. Dahdul, D\. S\. Dougall, Y\. He, D\. Osumi\-Sutherland, A\. Ruttenberg, S\. Sarntivijai,et al\.\(2016\)The Cell Ontology 2016: enhanced content, modularization, and ontology interoperability\.Journal of Biomedical Semantics7\(1\),pp\. 44\.Cited by:[§3](https://arxiv.org/html/2606.11646#S3.p4.1.1.1)\.
- J\. J\. Egozcue, V\. Pawlowsky\-Glahn, G\. Mateu\-Figueras, and C\. Barcelo\-Vidal \(2003\)Isometric logratio transformations for compositional data analysis\.Mathematical Geology35\(3\),pp\. 279–300\.Cited by:[§1](https://arxiv.org/html/2606.11646#S1.p2.1),[§2\.2](https://arxiv.org/html/2606.11646#S2.SS2.p1.8),[§7](https://arxiv.org/html/2606.11646#S7.p2.5),[§7](https://arxiv.org/html/2606.11646#S7.p4.1.1),[§8](https://arxiv.org/html/2606.11646#S8.p1.1)\.
- J\. J\. Egozcue and V\. Pawlowsky\-Glahn \(2005\)Groups of parts and their balances in compositional data analysis\.Mathematical Geology37\(7\),pp\. 795–828\.Cited by:[§1](https://arxiv.org/html/2606.11646#S1.p2.1),[§3](https://arxiv.org/html/2606.11646#S3.p2.2)\.
- J\. J\. Egozcue and V\. Pawlowsky\-Glahn \(2016\)Changing the reference measure in the simplex and its weighting effects\.Austrian Journal of Statistics45\(4\),pp\. 25–44\.Cited by:[§A\.2](https://arxiv.org/html/2606.11646#A1.SS2.SSS0.Px4.p1.13.1)\.
- G\. B\. Gloor, J\. M\. Macklaim, V\. Pawlowsky\-Glahn, and J\. J\. Egozcue \(2017\)Microbiome datasets are compositional: and this is not optional\.Frontiers in Microbiology8,pp\. 2224\.Cited by:[§1](https://arxiv.org/html/2606.11646#S1.p1.1),[§2\.1](https://arxiv.org/html/2606.11646#S2.SS1.p1.2)\.
- L\. Harmon \(2019\)Phylogenetic comparative methods: learning from trees\.Cited by:[§1](https://arxiv.org/html/2606.11646#S1.p1.1)\.
- Y\. Hu, G\. A\. Satten, and Y\. Hu \(2022\)LOCOM: a logistic regression model for testing differential abundance in compositional microbiome data with false discovery rate control\.Proceedings of the National Academy of Sciences119\(30\),pp\. e2122788119\.Cited by:[§B\.4](https://arxiv.org/html/2606.11646#A2.SS4.p1.2)\.
- Human Microbiome Project Consortium \(2012\)Structure, function and diversity of the healthy human microbiome\.Nature486\(7402\),pp\. 207–214\.Cited by:[§6\.1](https://arxiv.org/html/2606.11646#S6.SS1.p1.1),[§6\.4](https://arxiv.org/html/2606.11646#S6.SS4.p2.1),[§6\.5](https://arxiv.org/html/2606.11646#S6.SS5.p2.1)\.
- D\. A\. Jackson \(1997\)Compositional data in community ecology: the paradigm or peril of proportions?\.Ecology78\(3\),pp\. 929–940\.Cited by:[§2\.1](https://arxiv.org/html/2606.11646#S2.SS1.p1.2)\.
- A\. Kaul, S\. Mandal, O\. Davidov, and S\. D\. Peddada \(2017\)Analysis of microbiome data in the presence of excess zeros\.Frontiers in microbiology8,pp\. 2114\.Cited by:[§B\.4](https://arxiv.org/html/2606.11646#A2.SS4.p1.2)\.
- H\. Lancaster \(1965\)The Helmert matrices\.The American Mathematical Monthly72\(1\),pp\. 4–12\.Cited by:[§4\.2](https://arxiv.org/html/2606.11646#S4.SS2.p4.2)\.
- M\. Li, X\. Zhang, K\. S\. Ang, J\. Ling, R\. Sethi, N\. Y\. S\. Lee, F\. Ginhoux, and J\. Chen \(2022\)DISCO: a database of deeply integrated human single\-cell omics data\.Nucleic acids research50\(D1\),pp\. D596–D602\.Cited by:[§B\.2](https://arxiv.org/html/2606.11646#A2.SS2.SSS0.Px2.p1.1),[§6\.1](https://arxiv.org/html/2606.11646#S6.SS1.p1.1)\.
- J\. Q\. Liang, T\. Li, G\. Nakatsu, Y\. Chen, T\. O\. Yau, E\. Chu, S\. Wong, C\. H\. Szeto, S\. C\. Ng, F\. K\. Chan,et al\.\(2020\)A novel faecal lachnoclostridium marker for the non\-invasive diagnosis of colorectal adenoma and cancer\.Gut69\(7\),pp\. 1248–1257\.Cited by:[§6\.4](https://arxiv.org/html/2606.11646#S6.SS4.p2.1)\.
- G\. N\. Lin, C\. Zhang, and D\. Xu \(2011\)Polytomy identification in microbial phylogenetic reconstruction\.BMC systems Biology5\(Suppl 3\),pp\. S2\.Cited by:[§3](https://arxiv.org/html/2606.11646#S3.p4.1),[§3](https://arxiv.org/html/2606.11646#S3.p4.1.1.1)\.
- B\. Löwenberg, J\. R\. Downing, and A\. Burnett \(1999\)Acute myeloid leukemia\.New England Journal of Medicine341\(14\),pp\. 1051–1062\.Cited by:[§6\.4](https://arxiv.org/html/2606.11646#S6.SS4.p2.1),[§6\.5](https://arxiv.org/html/2606.11646#S6.SS5.p2.1)\.
- C\. Lozupone and R\. Knight \(2005\)UniFrac: a new phylogenetic method for comparing microbial communities\.Applied and environmental microbiology71\(12\),pp\. 8228–8235\.Cited by:[§1](https://arxiv.org/html/2606.11646#S1.p2.1),[§3](https://arxiv.org/html/2606.11646#S3.p1.4),[§8](https://arxiv.org/html/2606.11646#S8.p1.1)\.
- Z\. Ma, T\. Zuo, N\. Frey, and A\. Y\. Rangrez \(2024\)A systematic framework for understanding the microbiome in human health and disease: from basic principles to clinical translation\.Signal Transduction and Targeted Therapy9\(1\),pp\. 237\.Cited by:[§6\.5](https://arxiv.org/html/2606.11646#S6.SS5.p2.1)\.
- W\. Maddison \(1989\)Reconstructing character evolution on polytomous cladograms\.Cladistics5\(4\),pp\. 365–377\.Cited by:[§3](https://arxiv.org/html/2606.11646#S3.p4.1.1.1)\.
- R\. Malashin, Y\. Valeria, and A\. V\. Mullin \(2025\)Hypernym bias: unraveling deep classifier training dynamics through the lens of class hierarchy\.InThe 28th International Conference on Artificial Intelligence and Statistics,External Links:[Link](https://openreview.net/forum?id=DobXzInjnV)Cited by:[§B\.6](https://arxiv.org/html/2606.11646#A2.SS6.SSS0.Px1.p3.1)\.
- S\. G\. Mallat \(2002\)A theory for multiresolution signal decomposition: the wavelet representation\.IEEE transactions on pattern analysis and machine intelligence11\(7\),pp\. 674–693\.Cited by:[§2\.2](https://arxiv.org/html/2606.11646#S2.SS2.p2.1)\.
- S\. Mandal, W\. Van Treuren, R\. A\. White, M\. Eggesbø, R\. Knight, and S\. D\. Peddada \(2015\)Analysis of composition of microbiomes: a novel method for studying microbial composition\.Microbial ecology in health and disease26\(1\),pp\. 27663\.Cited by:[§1](https://arxiv.org/html/2606.11646#S1.p2.1),[§3](https://arxiv.org/html/2606.11646#S3.p4.1)\.
- J\. Mao and L\. Ma \(2022\)Dirichlet\-tree multinomial mixtures for clustering microbiome compositions\.The Annals of Applied Statistics16\(3\),pp\. 1476\.Cited by:[§1](https://arxiv.org/html/2606.11646#S1.p2.1),[§8](https://arxiv.org/html/2606.11646#S8.p1.1)\.
- L\. J\. Marcos\-Zambrano, K\. Karaduzovic\-Hadziabdic, T\. Loncar Turukalo, P\. Przymus, V\. Trajkovik, O\. Aasmets, M\. Berland, A\. Gruca, J\. Hasic, K\. Hron,et al\.\(2021\)Applications of machine learning in human microbiome studies: a review on feature selection, biomarker identification, disease prediction and treatment\.Frontiers in Microbiology12,pp\. 634511\.Cited by:[§5\.2](https://arxiv.org/html/2606.11646#S5.SS2.p2.1)\.
- G\. A\. Miller \(1995\)WordNet: a lexical database for English\.Communications of the ACM38\(11\),pp\. 39–41\.Cited by:[§7](https://arxiv.org/html/2606.11646#S7.p5.10)\.
- A\. Monod, B\. Lin, R\. Yoshida, and Q\. Kang \(2018\)Tropical geometry of phylogenetic tree space: a statistical perspective\.arXiv preprint arXiv:1805\.12400\.Cited by:[§8](https://arxiv.org/html/2606.11646#S8.p2.1)\.
- J\. T\. Morton, J\. Sanders, R\. A\. Quinn, D\. McDonald, A\. Gonzalez, Y\. Vázquez\-Baeza, J\. A\. Navas\-Molina, S\. J\. Song, J\. L\. Metcalf, E\. R\. Hyde,et al\.\(2017\)Balance trees reveal microbial niche differentiation\.MSystems2\(1\),pp\. e00162–16\.Cited by:[§1](https://arxiv.org/html/2606.11646#S1.p2.1)\.
- M\. Nickel and D\. Kiela \(2017\)Poincaré embeddings for learning hierarchical representations\.Advances in neural information processing systems30\.Cited by:[§8](https://arxiv.org/html/2606.11646#S8.p2.1)\.
- M\. Owen and J\. S\. Provan \(2010\)A fast algorithm for computing geodesic distances in tree space\.IEEE/ACM Transactions on Computational Biology and Bioinformatics8\(1\),pp\. 2–13\.Cited by:[§8](https://arxiv.org/html/2606.11646#S8.p2.1)\.
- E\. Palumbo, M\. Vandenhirtz, A\. Ryser, I\. Daunhawer, and J\. E\. Vogt \(2025\)From logits to hierarchies: hierarchical clustering made simple\.InForty\-second International Conference on Machine Learning,External Links:[Link](https://openreview.net/forum?id=t0x2VnBskT)Cited by:[§B\.6](https://arxiv.org/html/2606.11646#A2.SS6.SSS0.Px1.p2.3)\.
- E\. Pasolli, F\. De Filippis, I\. E\. Mauriello, F\. Cumbo, A\. M\. Walsh, J\. Leech, P\. D\. Cotter, N\. Segata, and D\. Ercolini \(2020\)Large\-scale genome\-wide analysis links lactic acid bacteria from food with the gut microbiome\.Nature Communications11\(1\),pp\. 2610\.Cited by:[§6\.4](https://arxiv.org/html/2606.11646#S6.SS4.p2.1.2)\.
- E\. Pasolli, L\. Schiffer, P\. Manghi, A\. Renson, V\. Obenchain, D\. T\. Truong, F\. Beghini, F\. Malik, M\. Ramos, J\. B\. Dowd,et al\.\(2017\)Accessible, curated metagenomic data through experimenthub\.Nature Methods14\(11\),pp\. 1023–1024\.Cited by:[§6\.1](https://arxiv.org/html/2606.11646#S6.SS1.p1.1)\.
- V\. Pawlowsky\-Glahn and J\. J\. Egozcue \(2001\)Geometric approach to statistical analysis on the simplex\.Stochastic Environmental Research and Risk Assessment15\(5\),pp\. 384–398\.Cited by:[§1](https://arxiv.org/html/2606.11646#S1.p4.1)\.
- B\. Phipson, C\. B\. Sim, E\. R\. Porrello, A\. W\. Hewitt, J\. Powell, and A\. Oshlack \(2022\)propeller: testing for differences in cell type proportions in single cell data\.Bioinformatics38\(20\),pp\. 4720–4726\.Cited by:[§B\.2](https://arxiv.org/html/2606.11646#A2.SS2.SSS0.Px1.p1.1)\.
- J\. Rivera\-Pinto, J\. J\. Egozcue, V\. Pawlowsky\-Glahn, R\. Paredes, M\. Noguera\-Julian, and M\. L\. Calle \(2018\)Balances: a new perspective for microbiome analysis\.MSystems3\(4\),pp\. 10–1128\.Cited by:[§3](https://arxiv.org/html/2606.11646#S3.p4.1),[§4\.3](https://arxiv.org/html/2606.11646#S4.SS3.p1.3),[§8](https://arxiv.org/html/2606.11646#S8.p1.1)\.
- C\. Rudin \(2019\)Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead\.Nature Machine Intelligence1\(5\),pp\. 206–215\.Cited by:[§5\.2](https://arxiv.org/html/2606.11646#S5.SS2.p2.1)\.
- F\. Sala, C\. De Sa, A\. Gu, and C\. Ré \(2018\)Representation tradeoffs for hyperbolic embeddings\.InInternational Conference on Machine Learning,pp\. 4460–4469\.Cited by:[§8](https://arxiv.org/html/2606.11646#S8.p2.1)\.
- C\. N\. Silla Jr and A\. A\. Freitas \(2011\)A survey of hierarchical classification across different application domains\.Data mining and knowledge discovery22\(1\),pp\. 31–72\.Cited by:[§1](https://arxiv.org/html/2606.11646#S1.p5.1)\.
- J\. D\. Silverman, A\. D\. Washburne, S\. Mukherjee, and L\. A\. David \(2017\)A phylogenetic transform enhances analysis of compositional microbiota data\.elife6,pp\. e21887\.Cited by:[§A\.2](https://arxiv.org/html/2606.11646#A1.SS2.SSS0.Px4.p1.13),[§A\.2](https://arxiv.org/html/2606.11646#A1.SS2.SSS0.Px4.p1.9),[§B\.3](https://arxiv.org/html/2606.11646#A2.SS3.p2.1),[§1](https://arxiv.org/html/2606.11646#S1.p1.1),[§1](https://arxiv.org/html/2606.11646#S1.p2.1),[§3](https://arxiv.org/html/2606.11646#S3.p2.2),[§8](https://arxiv.org/html/2606.11646#S8.p1.1.2)\.
- K\. K\. Sørensen, J\. Simon\-Santamaria, R\. S\. McCuskey, and B\. Smedsrød \(2015\)Liver sinusoidal endothelial cells\.Comprehensive Physiology5\(4\),pp\. 1751–1774\.Cited by:[§6\.4](https://arxiv.org/html/2606.11646#S6.SS4.p2.1),[§6\.5](https://arxiv.org/html/2606.11646#S6.SS5.p2.1)\.
- A\. D\. Washburne, J\. D\. Silverman, J\. W\. Leff, D\. J\. Bennett, J\. L\. Darcy, S\. Mukherjee, N\. Fierer, and L\. A\. David \(2017\)Phylogenetic factorization of compositional data yields lineage\-level associations in microbiome datasets\.PeerJ5,pp\. e2969\.Cited by:[§1](https://arxiv.org/html/2606.11646#S1.p2.1),[§3](https://arxiv.org/html/2606.11646#S3.p4.1),[§8](https://arxiv.org/html/2606.11646#S8.p1.1)\.
- T\. Yatsunenko, F\. E\. Rey, M\. J\. Manary, I\. Trehan, M\. G\. Dominguez\-Bello, M\. Contreras, M\. Magris, G\. Hidalgo, R\. N\. Baldassano, A\. P\. Anokhin,et al\.\(2012\)Human gut microbiome viewed across age and geography\.Nature486\(7402\),pp\. 222–227\.Cited by:[§6\.4](https://arxiv.org/html/2606.11646#S6.SS4.p2.1)\.
- A\. Yerke, D\. Fry Brumit, and A\. A\. Fodor \(2024\)Proportion\-based normalizations outperform compositional data transformations in machine learning applications\.Microbiome12\(1\),pp\. 45\.Cited by:[§8](https://arxiv.org/html/2606.11646#S8.p1.1.5)\.
## Appendix AProofs and Algorithm Details
This section contains full proofs of the main theoretical results and detailed algorithm specifications\.
### A\.1Proofs
###### Theorem A\.1\(Restatement of Theorem[4\.1](https://arxiv.org/html/2606.11646#S4.Thmtheorem1)\)\.
Let𝒯\\mathcal\{T\}be any rooted tree withddleaves\. The matrixV∈ℝd×\(d−1\)V\\in\\mathbb\{R\}^\{d\\times\(d\-1\)\}from Algorithm[1](https://arxiv.org/html/2606.11646#alg1)satisfies:
1. 1\.V⊤𝟏=0V^\{\\top\}\\mathbf\{1\}=0\(contrast property\),
2. 2\.V⊤V=Id−1V^\{\\top\}V=I\_\{d\-1\}\(orthonormality\)\.
Consequently,φ\(x\)=V⊤logx\\varphi\(x\)=V^\{\\top\}\\log xis an isometry from\(Δd−1,⟨⋅,⋅⟩A\)\(\\Delta^\{d\-1\},\\langle\\cdot,\\cdot\\rangle\_\{A\}\)to\(ℝd−1,⟨⋅,⋅⟩2\)\(\\mathbb\{R\}^\{d\-1\},\\langle\\cdot,\\cdot\\rangle\_\{2\}\)\.
###### Proof\.
LetTTbe a rooted tree withddleaves\. For each internal nodeuuwithkuk\_\{u\}children, letCu\(r\)C\_\{u\}^\{\(r\)\}denote the set of leaves descending from childrr, and letnr=\|Cu\(r\)\|n\_\{r\}=\|C\_\{u\}^\{\(r\)\}\|\. The construction produces a local basisH~\(u\)∈ℝku×\(ku−1\)\\tilde\{H\}^\{\(u\)\}\\in\\mathbb\{R\}^\{k\_\{u\}\\times\(k\_\{u\}\-1\)\}that is orthonormal under the weighted inner product⟨h,h′⟩w=∑r=1kuhrhr′/nr\\langle h,h^\{\\prime\}\\rangle\_\{w\}=\\sum\_\{r=1\}^\{k\_\{u\}\}h\_\{r\}h^\{\\prime\}\_\{r\}/n\_\{r\}and satisfies the contrast condition∑r=1kuH~r,m\(u\)=0\\sum\_\{r=1\}^\{k\_\{u\}\}\\tilde\{H\}^\{\(u\)\}\_\{r,m\}=0for allmm\. Each local vector is spread to a global vectorv\(u,m\)∈ℝdv^\{\(u,m\)\}\\in\\mathbb\{R\}^\{d\}via
vi\(u,m\)=\{H~r,m\(u\)/nrifi∈Cu\(r\)for somer∈\{1,…,ku\},0otherwise\.v^\{\(u,m\)\}\_\{i\}=\\begin\{cases\}\\tilde\{H\}^\{\(u\)\}\_\{r,m\}/n\_\{r\}&\\text\{if \}i\\in C\_\{u\}^\{\(r\)\}\\text\{ for some \}r\\in\\\{1,\\ldots,k\_\{u\}\\\},\\\\ 0&\\text\{otherwise\}\.\\end\{cases\}The precise procedure is given in Algorithm[1](https://arxiv.org/html/2606.11646#alg1)\. It suffices for us to show thatVV\(i\) satisfies the contrast property \(hence its columns span the tangent spaceℋ\\mathcal\{H\}\), \(ii\) is orthonormal within each node, and \(iii\) is orthonormal across distinct nodes\.
*\(i\) Contrast property\.*For any internal nodeuuand contrast indexmm,
∑i=1dvi\(u,m\)=∑r=1ku∑i∈Cu\(r\)H~r,m\(u\)nr=∑r=1kunr⋅H~r,m\(u\)nr=∑r=1kuH~r,m\(u\)=0,\\sum\_\{i=1\}^\{d\}v^\{\(u,m\)\}\_\{i\}=\\sum\_\{r=1\}^\{k\_\{u\}\}\\sum\_\{i\\in C\_\{u\}^\{\(r\)\}\}\\frac\{\\tilde\{H\}^\{\(u\)\}\_\{r,m\}\}\{n\_\{r\}\}=\\sum\_\{r=1\}^\{k\_\{u\}\}n\_\{r\}\\cdot\\frac\{\\tilde\{H\}^\{\(u\)\}\_\{r,m\}\}\{n\_\{r\}\}=\\sum\_\{r=1\}^\{k\_\{u\}\}\\tilde\{H\}^\{\(u\)\}\_\{r,m\}=0,where the last equality holds becauseH~\(u\)\\tilde\{H\}^\{\(u\)\}is obtained by applying Gram\-Schmidt to Helmert contrasts, which lie in the subspaceSu=\{h∈ℝku:∑rhr=0\}S\_\{u\}=\\\{h\\in\\mathbb\{R\}^\{k\_\{u\}\}:\\sum\_\{r\}h\_\{r\}=0\\\}, and Gram\-Schmidt preserves this subspace\.
*\(ii\) Within\-node orthonormality\.*For any internal nodeuuand contrast indicesm1,m2∈\{1,…,ku−1\}m\_\{1\},m\_\{2\}\\in\\\{1,\\ldots,k\_\{u\}\-1\\\},
⟨v\(u,m1\),v\(u,m2\)⟩\\displaystyle\\langle v^\{\(u,m\_\{1\}\)\},v^\{\(u,m\_\{2\}\)\}\\rangle=∑i=1dvi\(u,m1\)vi\(u,m2\)=∑r=1ku∑i∈Cu\(r\)H~r,m1\(u\)nr⋅H~r,m2\(u\)nr\\displaystyle=\\sum\_\{i=1\}^\{d\}v^\{\(u,m\_\{1\}\)\}\_\{i\}\\,v^\{\(u,m\_\{2\}\)\}\_\{i\}=\\sum\_\{r=1\}^\{k\_\{u\}\}\\sum\_\{i\\in C\_\{u\}^\{\(r\)\}\}\\frac\{\\tilde\{H\}^\{\(u\)\}\_\{r,m\_\{1\}\}\}\{n\_\{r\}\}\\cdot\\frac\{\\tilde\{H\}^\{\(u\)\}\_\{r,m\_\{2\}\}\}\{n\_\{r\}\}=∑r=1kunr⋅H~r,m1\(u\)H~r,m2\(u\)nr2=∑r=1kuH~r,m1\(u\)H~r,m2\(u\)nr\\displaystyle=\\sum\_\{r=1\}^\{k\_\{u\}\}n\_\{r\}\\cdot\\frac\{\\tilde\{H\}^\{\(u\)\}\_\{r,m\_\{1\}\}\\,\\tilde\{H\}^\{\(u\)\}\_\{r,m\_\{2\}\}\}\{n\_\{r\}^\{2\}\}=\\sum\_\{r=1\}^\{k\_\{u\}\}\\frac\{\\tilde\{H\}^\{\(u\)\}\_\{r,m\_\{1\}\}\\,\\tilde\{H\}^\{\(u\)\}\_\{r,m\_\{2\}\}\}\{n\_\{r\}\}=⟨H~⋅,m1\(u\),H~⋅,m2\(u\)⟩w=δm1,m2,\\displaystyle=\\langle\\tilde\{H\}^\{\(u\)\}\_\{\\cdot,m\_\{1\}\},\\tilde\{H\}^\{\(u\)\}\_\{\\cdot,m\_\{2\}\}\\rangle\_\{w\}=\\delta\_\{m\_\{1\},m\_\{2\}\},where the last equality holds by construction ofH~\(u\)\\tilde\{H\}^\{\(u\)\}as an orthonormal basis under⟨⋅,⋅⟩w\\langle\\cdot,\\cdot\\rangle\_\{w\}\.
*\(iii\) Across\-node orthogonality\.*Consider two distinct internal nodesu≠u′u\\neq u^\{\\prime\}with contrast indicesmmandm′m^\{\\prime\}respectively\. We consider two cases\. \(a\) If the subtrees rooted atuuandu′u^\{\\prime\}share no leaves, thenv\(u,m\)v^\{\(u,m\)\}andv\(u′,m′\)v^\{\(u^\{\\prime\},m^\{\\prime\}\)\}have disjoint supports, so⟨v\(u,m\),v\(u′,m′\)⟩=0\\langle v^\{\(u,m\)\},v^\{\(u^\{\\prime\},m^\{\\prime\}\)\}\\rangle=0\. \(b\) Suppose without loss of generality thatu′u^\{\\prime\}is a descendant ofuu\. Thenu′u^\{\\prime\}lies entirely within the subtree of exactly one child ofuu, say childss, so the support ofv\(u′,m′\)v^\{\(u^\{\\prime\},m^\{\\prime\}\)\}is contained inCu\(s\)C\_\{u\}^\{\(s\)\}\. OnCu\(s\)C\_\{u\}^\{\(s\)\}, the vectorv\(u,m\)v^\{\(u,m\)\}takes the constant valueH~s,m\(u\)/ns\\tilde\{H\}^\{\(u\)\}\_\{s,m\}/n\_\{s\}\. Sincevi\(u′,m′\)=0v^\{\(u^\{\\prime\},m^\{\\prime\}\)\}\_\{i\}=0fori∉Cu\(s\)i\\notin C\_\{u\}^\{\(s\)\},
⟨v\(u,m\),v\(u′,m′\)⟩=∑i∈Cu\(s\)vi\(u,m\)vi\(u′,m′\)=H~s,m\(u\)ns∑i∈Cu\(s\)vi\(u′,m′\)=H~s,m\(u\)ns⋅0=0,\\langle v^\{\(u,m\)\},v^\{\(u^\{\\prime\},m^\{\\prime\}\)\}\\rangle=\\sum\_\{i\\in C\_\{u\}^\{\(s\)\}\}v^\{\(u,m\)\}\_\{i\}\\,v^\{\(u^\{\\prime\},m^\{\\prime\}\)\}\_\{i\}=\\frac\{\\tilde\{H\}^\{\(u\)\}\_\{s,m\}\}\{n\_\{s\}\}\\sum\_\{i\\in C\_\{u\}^\{\(s\)\}\}v^\{\(u^\{\\prime\},m^\{\\prime\}\)\}\_\{i\}=\\frac\{\\tilde\{H\}^\{\(u\)\}\_\{s,m\}\}\{n\_\{s\}\}\\cdot 0=0,where the last equality uses \(i\): the entries ofv\(u′,m′\)v^\{\(u^\{\\prime\},m^\{\\prime\}\)\}sum to zero, and all nonzero entries lie withinCu\(s\)C\_\{u\}^\{\(s\)\}\.
*Conclusion\.*\(i\), \(ii\), and \(iii\) establish thatV⊤𝟏=0V^\{\\top\}\\mathbf\{1\}=0andV⊤V=Id−1V^\{\\top\}V=I\_\{d\-1\}\. For the dimension count, letNintN\_\{\\mathrm\{int\}\}denote the number of internal nodes\. Every non\-root node has exactly one parent, so the number of edges satisfies\|E\|=d\+Nint−1\|E\|=d\+N\_\{\\mathrm\{int\}\}\-1\. Since\|E\|=∑uku\|E\|=\\sum\_\{u\}k\_\{u\}, we have
∑u\(ku−1\)=∑uku−Nint=\(d\+Nint−1\)−Nint=d−1=dim\(ℋ\)\.\\sum\_\{u\}\(k\_\{u\}\-1\)=\\sum\_\{u\}k\_\{u\}\-N\_\{\\mathrm\{int\}\}=\(d\+N\_\{\\mathrm\{int\}\}\-1\)\-N\_\{\\mathrm\{int\}\}=d\-1=\\dim\(\\mathcal\{H\}\)\.By standard ILR theory, anyV∈ℝd×\(d−1\)V\\in\\mathbb\{R\}^\{d\\times\(d\-1\)\}satisfyingV⊤𝟏=0V^\{\\top\}\\mathbf\{1\}=0andV⊤V=Id−1V^\{\\top\}V=I\_\{d\-1\}has columns forming an orthonormal basis ofℋ\\mathcal\{H\}, the image ofclr\\mathrm\{clr\}\. Thusϕ\(x\)=V⊤logx\\phi\(x\)=V^\{\\top\}\\log xis an isometry from\(Δd−1,⟨⋅,⋅⟩A\)\(\\Delta^\{d\-1\},\\langle\\cdot,\\cdot\\rangle\_\{A\}\)to\(ℝd−1,⟨⋅,⋅⟩2\)\(\\mathbb\{R\}^\{d\-1\},\\langle\\cdot,\\cdot\\rangle\_\{2\}\)\. ∎
###### Proposition A\.2\(Restatement of Proposition[4\.2](https://arxiv.org/html/2606.11646#S4.Thmtheorem2)\)\.
Given a rooted tree𝒯\\mathcal\{T\}with fixed leaf labels, fixed child orderings,and a fixed orderingπ\\piof internal nodes \(e\.g\., DFS\), PolyILR produces a unique basisV\(𝒯\)V\(\\mathcal\{T\}\)\. Moreover,𝒯↦V\(𝒯\)\\mathcal\{T\}\\mapsto V\(\\mathcal\{T\}\)is injective:𝒯\\mathcal\{T\}can be recovered constructively fromVV\.
###### Proof\.
We prove uniqueness and recoverability separately\.
*\(i\) Uniqueness\.*The PolyILR construction \(Algorithm[1](https://arxiv.org/html/2606.11646#alg1)\) is deterministic given the treeTT, leaf labels, child orderings at each internal node, and an orderingπ\\piof internal nodes\. At each nodeuuwithkuk\_\{u\}children, the construction follows:
- •The descendant setsCu\(r\)C\_\{u\}^\{\(r\)\}and countsnr=\|Cu\(r\)\|n\_\{r\}=\|C\_\{u\}^\{\(r\)\}\|are determined byTTand the leaf labels, and the corresponding weighted inner product⟨h,h′⟩w=∑rhrhr′/nr\\langle h,h^\{\\prime\}\\rangle\_\{w\}=\\sum\_\{r\}h\_\{r\}h^\{\\prime\}\_\{r\}/n\_\{r\}is determined by\{nr\}\\\{n\_\{r\}\\\}\.
- •The Helmert matrixH\(u\)∈ℝku×\(ku−1\)H^\{\(u\)\}\\in\\mathbb\{R\}^\{k\_\{u\}\\times\(k\_\{u\}\-1\)\}is determined bykuk\_\{u\}and the child ordering\. The weighted\-orthonormal basisH~\(u\)\\tilde\{H\}^\{\(u\)\}is obtained by applying Gram\-Schmidt toH\(u\)H^\{\(u\)\}under⟨⋅,⋅⟩w\\langle\\cdot,\\cdot\\rangle\_\{w\}\. Gram\-Schmidt is made unique by normalizing each output vector to unitww\-norm and adopting a fixed sign convention \(e\.g\., the first nonzero entry in child order is positive\)\. Gram\-Schmidt works as the weighted inner product forms a Hilbert space over the subspace𝒮u\\mathcal\{S\}\_\{u\}\.
- •The spreading operation is deterministic givenCu\(r\)C\_\{u\}^\{\(r\)\}andnrn\_\{r\}\.
ThusV\(T\)V\(T\)is uniquely determined\. This makes PolyILR*canonical*, i\.e\., the construction does not incur intractable randomness\.
*\(ii\) Recoverability\.*We show that the tree topology can be recovered fromVV\.
*Support structure:*By the structure of Helmert contrasts, columnmmofH\(u\)H^\{\(u\)\}has nonzero entries only in rows1,…,m\+11,\\ldots,m\+1\. Since Gram\-Schmidt orthonormalizes sequentially,H~⋅,m\(u\)\\tilde\{H\}^\{\(u\)\}\_\{\\cdot,m\}is a linear combination ofH⋅,1\(u\),…,H⋅,m\(u\)H^\{\(u\)\}\_\{\\cdot,1\},\\ldots,H^\{\(u\)\}\_\{\\cdot,m\}, soH~r,m\(u\)=0\\tilde\{H\}^\{\(u\)\}\_\{r,m\}=0forr\>m\+1r\>m\+1\. Moreover, columns1,…,m−11,\\ldots,m\-1ofH\(u\)H^\{\(u\)\}have zeros in rowm\+1m\+1, so orthogonalizing against them does not affect entry\(m\+1,m\)\(m\+1,m\); thusH~m\+1,m\(u\)=Hm\+1,m\(u\)≠0\\tilde\{H\}^\{\(u\)\}\_\{m\+1,m\}=H^\{\(u\)\}\_\{m\+1,m\}\\neq 0\. After spreading, the support ofv\(u,m\)v^\{\(u,m\)\}is⋃r=1m\+1Cu\(r\)\\bigcup\_\{r=1\}^\{m\+1\}C\_\{u\}^\{\(r\)\}, and the supports at nodeuuform a*strictly nested chain*:
supp\(v\(u,1\)\)⊊supp\(v\(u,2\)\)⊊⋯⊊supp\(v\(u,ku−1\)\)=⋃r=1kuCu\(r\)\.\\mathrm\{supp\}\(v^\{\(u,1\)\}\)\\subsetneq\\mathrm\{supp\}\(v^\{\(u,2\)\}\)\\subsetneq\\cdots\\subsetneq\\mathrm\{supp\}\(v^\{\(u,k\_\{u\}\-1\)\}\)=\\bigcup\_\{r=1\}^\{k\_\{u\}\}C\_\{u\}^\{\(r\)\}\.
*Recovering the child clades of a fixed node\.*Fix an internal nodeuuwith childrenCu\(1\),…,Cu\(ku\)C\_\{u\}^\{\(1\)\},\\dots,C\_\{u\}^\{\(k\_\{u\}\)\}\. From the support\-structure argument above,
supp\(v\(u,m\)\)=⋃r=1m\+1Cu\(r\)\.\\mathrm\{supp\}\\\!\\bigl\(v^\{\(u,m\)\}\\bigr\)=\\bigcup\_\{r=1\}^\{m\+1\}C\_\{u\}^\{\(r\)\}\.Form=2,…,ku−1m=2,\\dots,k\_\{u\}\-1, the child cladesCu\(m\+1\)=supp\(v\(u,m\)\)∖supp\(v\(u,m−1\)\)C\_\{u\}^\{\(m\+1\)\}=\\mathrm\{supp\}\\\!\\bigl\(v^\{\(u,m\)\}\\bigr\)\\setminus\\mathrm\{supp\}\\\!\\bigl\(v^\{\(u,m\-1\)\}\\bigr\)are determined by the support chain alone\. It remains to recoverCu\(1\)C\_\{u\}^\{\(1\)\}andCu\(2\)C\_\{u\}^\{\(2\)\}fromsupp\(v\(u,1\)\)=Cu\(1\)∪Cu\(2\)\\mathrm\{supp\}\(v^\{\(u,1\)\}\)=C\_\{u\}^\{\(1\)\}\\cup C\_\{u\}^\{\(2\)\}\. Since there are no previous columns to orthogonalize against, the first weighted Gram\-Schmidt output satisfiesH~⋅,1\(u\)=αH⋅,1\(u\)\\tilde\{H\}^\{\(u\)\}\_\{\\cdot,1\}=\\alpha\\,H^\{\(u\)\}\_\{\\cdot,1\}for someα\>0\\alpha\>0\. The first Helmert column has opposite signs in rows11and22, soH~1,1\(u\)\>0\\tilde\{H\}^\{\(u\)\}\_\{1,1\}\>0andH~2,1\(u\)<0\\tilde\{H\}^\{\(u\)\}\_\{2,1\}<0\. After spreading, the valuesH~1,1\(u\)/n1\>0\>H~2,1\(u\)/n2\\tilde\{H\}^\{\(u\)\}\_\{1,1\}/n\_\{1\}\>0\>\\tilde\{H\}^\{\(u\)\}\_\{2,1\}/n\_\{2\}are distinct, so the two level sets ofv\(u,1\)v^\{\(u,1\)\}on its support are preciselyCu\(1\)C\_\{u\}^\{\(1\)\}andCu\(2\)C\_\{u\}^\{\(2\)\}\.
*Recovering the full tree\.*Applying the preceding child\-clade recovery argument at the root and then recursively to each non\-singleton child clade recovers all clades of𝒯\\mathcal\{T\}\. A rooted tree with labeled leaves is uniquely determined by its clades via inclusion:u′u^\{\\prime\}is a child ofuuiff the clade ofu′u^\{\\prime\}is a maximal proper subset of the clade ofuu\(i\.e\., a strict subsetC′⊊CC^\{\\prime\}\\subsetneq Cwith no cladeC′′C^\{\\prime\\prime\}satisfyingC′⊊C′′⊊CC^\{\\prime\}\\subsetneq C^\{\\prime\\prime\}\\subsetneq C\), and leaves are singleton clades\. Thus𝒯\\mathcal\{T\}is recoverable fromVV\.∎
Here, we comment on child orderings\. Different child orderings at nodeuuyield local bases that differ by an orthogonal transformation on\(Su,⟨⋅,⋅⟩w\)\(S\_\{u\},\\langle\\cdot,\\cdot\\rangle\_\{w\}\)\. This induces an orthogonal transformation on the\(ku−1\)\(k\_\{u\}\-1\)\-dimensional coordinate block ofVVcorresponding touu\. The geometry is preserved; only coordinate labels change\. Now, we present the proof for Proposition[7\.1](https://arxiv.org/html/2606.11646#S7.Thmtheorem1)\.
###### Proposition A\.3\(Restatement of Proposition[7\.1](https://arxiv.org/html/2606.11646#S7.Thmtheorem1)\)\.
Letℒ=ℝd/∼ℓ\\mathcal\{L\}=\\mathbb\{R\}^\{d\}/\{\\sim\_\{\\ell\}\}be the quotient space of logits under shift equivalence\. Thenℒ≅ℋ\\mathcal\{L\}\\cong\\mathcal\{H\}\(linear isomorphism\)\. Concretely, for𝐩=softmax\(𝐳\)\\mathbf\{p\}=\\mathrm\{softmax\}\(\\mathbf\{z\}\), we have
𝐳−z¯𝟏=clr\(𝐩\),wherez¯=1d∑izi,\\mathbf\{z\}\-\\bar\{z\}\\mathbf\{1\}=\\mathrm\{clr\}\(\\mathbf\{p\}\),\\quad\\text\{where \}\\bar\{z\}=\\tfrac\{1\}\{d\}\\textstyle\\sum\_\{i\}z\_\{i\},and\[𝐳\]→𝐳−z¯𝟏\[\\mathbf\{z\}\]\\to\\mathbf\{z\}\-\\bar\{z\}\\mathbf\{1\}is well\-defined\.
###### Proof\.
We establish both the concrete identity and the isomorphism\.
*\(i\) Centered logits equal CLR coordinates\.*Letp=softmax\(z\)p=\\mathrm\{softmax\}\(z\), sopi=ezi/∑jezjp\_\{i\}=e^\{z\_\{i\}\}/\\sum\_\{j\}e^\{z\_\{j\}\}\. Taking logarithms,
logpi=zi−log∑jezj\.\\log p\_\{i\}=z\_\{i\}\-\\log\\sum\_\{j\}e^\{z\_\{j\}\}\.The CLR transform isclr\(p\)i=logpi−1d∑klogpk\\mathrm\{clr\}\(p\)\_\{i\}=\\log p\_\{i\}\-\\frac\{1\}\{d\}\\sum\_\{k\}\\log p\_\{k\}\. Substituting,
clr\(p\)i\\displaystyle\\mathrm\{clr\}\(p\)\_\{i\}=\(zi−log∑jezj\)−1d∑k\(zk−log∑jezj\)\\displaystyle=\\left\(z\_\{i\}\-\\log\\sum\_\{j\}e^\{z\_\{j\}\}\\right\)\-\\frac\{1\}\{d\}\\sum\_\{k\}\\left\(z\_\{k\}\-\\log\\sum\_\{j\}e^\{z\_\{j\}\}\\right\)=zi−log∑jezj−z¯\+log∑jezj\\displaystyle=z\_\{i\}\-\\log\\sum\_\{j\}e^\{z\_\{j\}\}\-\\bar\{z\}\+\\log\\sum\_\{j\}e^\{z\_\{j\}\}=zi−z¯,\\displaystyle=z\_\{i\}\-\\bar\{z\},wherez¯=1d∑kzk\\bar\{z\}=\\frac\{1\}\{d\}\\sum\_\{k\}z\_\{k\}\. Thusclr\(p\)=z−z¯𝟏\\mathrm\{clr\}\(p\)=z\-\\bar\{z\}\\mathbf\{1\}\.
*\(ii\) Isomorphism\.*Define the centering mapπ:ℝd→ℋ\\pi\\colon\\mathbb\{R\}^\{d\}\\to\\mathcal\{H\}byπ\(z\)=z−z¯𝟏\\pi\(z\)=z\-\\bar\{z\}\\mathbf\{1\}\. We showπ\\piinduces a linear isomorphism fromℒ=ℝd/∼ℓ\\mathcal\{L\}=\\mathbb\{R\}^\{d\}/\{\\sim\_\{\\ell\}\}toℋ\\mathcal\{H\}\. Ifz′=z\+c𝟏z^\{\\prime\}=z\+c\\mathbf\{1\}for somec∈ℝc\\in\\mathbb\{R\}, thenz′¯=z¯\+c\\bar\{z^\{\\prime\}\}=\\bar\{z\}\+c, so
π\(z′\)=z\+c𝟏−\(z¯\+c\)𝟏=z−z¯𝟏=π\(z\)\.\\pi\(z^\{\\prime\}\)=z\+c\\mathbf\{1\}\-\(\\bar\{z\}\+c\)\\mathbf\{1\}=z\-\\bar\{z\}\\mathbf\{1\}=\\pi\(z\)\.Thusπ\\piis constant on equivalence classes and descends to a mapπ~:ℒ→ℋ\\tilde\{\\pi\}\\colon\\mathcal\{L\}\\to\\mathcal\{H\}\. For anyz∈ℝdz\\in\\mathbb\{R\}^\{d\},𝟏⊤π\(z\)=∑i\(zi−z¯\)=dz¯−dz¯=0\\mathbf\{1\}^\{\\top\}\\pi\(z\)=\\sum\_\{i\}\(z\_\{i\}\-\\bar\{z\}\)=d\\bar\{z\}\-d\\bar\{z\}=0, soπ\(z\)∈ℋ\\pi\(z\)\\in\\mathcal\{H\}\. Now, for anyh∈ℋh\\in\\mathcal\{H\}, we haveh¯=0\\bar\{h\}=0, soπ\(h\)=h−0⋅𝟏=h\\pi\(h\)=h\-0\\cdot\\mathbf\{1\}=h\. Thusπ\\piis surjective ontoℋ\\mathcal\{H\}\. Supposeπ\(z\)=π\(z′\)\\pi\(z\)=\\pi\(z^\{\\prime\}\)\. Thenz−z¯𝟏=z′−z′¯𝟏z\-\\bar\{z\}\\mathbf\{1\}=z^\{\\prime\}\-\\bar\{z^\{\\prime\}\}\\mathbf\{1\}, which givesz−z′=\(z¯−z′¯\)𝟏z\-z^\{\\prime\}=\(\\bar\{z\}\-\\bar\{z^\{\\prime\}\}\)\\mathbf\{1\}\. Thusz∼ℓz′z\\sim\_\{\\ell\}z^\{\\prime\}, soπ~\\tilde\{\\pi\}is injective\. Sinceπ\\piis linear and constant on equivalence classes, the induced mapπ~\\tilde\{\\pi\}is linear\. Thus,π~:ℒ→ℋ\\tilde\{\\pi\}\\colon\\mathcal\{L\}\\to\\mathcal\{H\}is a linear isomorphism\. Moreover, equippingℒ\\mathcal\{L\}with the quotient metricdℒ\(\[z\],\[z′\]\)=infc∈ℝ‖z−z′\+c𝟏‖d\_\{\\mathcal\{L\}\}\(\[z\],\[z^\{\\prime\}\]\)=\\inf\_\{c\\in\\mathbb\{R\}\}\\\|z\-z^\{\\prime\}\+c\\mathbf\{1\}\\\|, the infimum is attained atc=z′¯−z¯c=\\bar\{z^\{\\prime\}\}\-\\bar\{z\}, which givesdℒ\(\[z\],\[z′\]\)=‖π\(z\)−π\(z′\)‖2d\_\{\\mathcal\{L\}\}\(\[z\],\[z^\{\\prime\}\]\)=\\\|\\pi\(z\)\-\\pi\(z^\{\\prime\}\)\\\|\_\{2\}\. Thusπ~\\tilde\{\\pi\}is an isometry\. ∎
### A\.2Algorithm Details
#### Time complexity and runtime of PolyILR\.
Algorithm[1](https://arxiv.org/html/2606.11646#alg1)constructsVVinO\(d2\)O\(d^\{2\}\)time\. For each internal nodeuuwithkuk\_\{u\}children, building the Helmert matrix takesO\(ku2\)O\(k\_\{u\}^\{2\}\)time, Gram\-Schmidt orthonormalization takesO\(ku3\)O\(k\_\{u\}^\{3\}\)time, and spreading to thedd\-dimensional vectors takesO\(nu⋅ku\)O\(n\_\{u\}\\cdot k\_\{u\}\)time wherenu=∑rnrn\_\{u\}=\\sum\_\{r\}n\_\{r\}is the number of leaves in the subtree rooted atuu\. Summing over all internal nodes, the total cost is dominated by the spreading step:∑unu⋅ku≤d⋅∑uku=O\(d2\)\\sum\_\{u\}n\_\{u\}\\cdot k\_\{u\}\\leq d\\cdot\\sum\_\{u\}k\_\{u\}=O\(d^\{2\}\)in the worst case \(e\.g\., a star tree\)\. For balanced trees, the complexity reduces toO\(dlogd\)O\(d\\log d\)\. In practice,VVis computed once as a preprocessing step, so this cost is negligible compared to downstream tasks such as model training or statistical inference, which typically scale with sample sizeNNand involve iterative optimization\. In all of our experiments, constructingVVusing PolyILR took<5<5seconds\.
#### On Gram\-Schmidt\.
The weighted Gram\-Schmidt procedure orthonormalizes the columns of the Helmert matrixH\(u\)H^\{\(u\)\}under the inner product⟨h,h′⟩w=∑r=1kuhrhr′/nr\\langle h,h^\{\\prime\}\\rangle\_\{w\}=\\sum\_\{r=1\}^\{k\_\{u\}\}h\_\{r\}h^\{\\prime\}\_\{r\}/n\_\{r\}\. This is well\-defined since\(Su,⟨⋅,⋅⟩w\)\(S\_\{u\},\\langle\\cdot,\\cdot\\rangle\_\{w\}\)is a finite\-dimensional Hilbert space\. Explicitly, form=1,…,ku−1m=1,\\ldots,k\_\{u\}\-1:
H~⋅,m\(u\)=H⋅,m\(u\)−∑j=1m−1⟨H⋅,m\(u\),H~⋅,j\(u\)⟩wH~⋅,j\(u\),then normalize:H~⋅,m\(u\)←H~⋅,m\(u\)‖H~⋅,m\(u\)‖w\.\\tilde\{H\}^\{\(u\)\}\_\{\\cdot,m\}=H^\{\(u\)\}\_\{\\cdot,m\}\-\\sum\_\{j=1\}^\{m\-1\}\\langle H^\{\(u\)\}\_\{\\cdot,m\},\\tilde\{H\}^\{\(u\)\}\_\{\\cdot,j\}\\rangle\_\{w\}\\,\\tilde\{H\}^\{\(u\)\}\_\{\\cdot,j\},\\quad\\text\{then normalize: \}\\tilde\{H\}^\{\(u\)\}\_\{\\cdot,m\}\\leftarrow\\frac\{\\tilde\{H\}^\{\(u\)\}\_\{\\cdot,m\}\}\{\\\|\\tilde\{H\}^\{\(u\)\}\_\{\\cdot,m\}\\\|\_\{w\}\}\.The weighting by1/nr1/n\_\{r\}accounts for clade sizes: larger clades contribute less per\-component to the inner product, ensuring that the resulting ILR coordinates treat leaves uniformly regardless of tree imbalance\.
#### On Child ordering conventions\.
The child ordering at each internal node affects which contrasts appear in which columns ofVV, but does not affect the subspace spanned or the geometric properties as mentioned before\. Two natural conventions are: \(i\)*Lexicographic:*Order children by the smallest leaf index in each subtree\. This is deterministic given leaf labels\. \(ii\)*By clade size:*Order children by descendingnrn\_\{r\}\. This places contrasts involving larger clades in earlier columns\. As noted in Proposition[4\.2](https://arxiv.org/html/2606.11646#S4.Thmtheorem2), different child orderings yield bases related by orthogonal transformations within each node’s coordinate block\. For reproducibility, we used the lexicographic ordering in all of our experiments\.
#### Reduction to PhILR\.
When the tree𝒯\\mathcal\{T\}is strictly binary \(every internal node has exactlyku=2k\_\{u\}=2children\), PolyILR reduces to PhILR\(Silvermanet al\.,[2017](https://arxiv.org/html/2606.11646#bib.bib4)\)\. We verify this algebraically\. For a binary nodeuuwith child cladesCu\(1\)C\_\{u\}^\{\(1\)\}andCu\(2\)C\_\{u\}^\{\(2\)\}of sizesn1n\_\{1\}andn2n\_\{2\}, the Helmert matrix isH\(u\)=\[1/2,−1/2\]⊤H^\{\(u\)\}=\[1/\\sqrt\{2\},\-1/\\sqrt\{2\}\]^\{\\top\}\. Under the weighted inner product⟨h,h′⟩w=h1h1′/n1\+h2h2′/n2\\langle h,h^\{\\prime\}\\rangle\_\{w\}=h\_\{1\}h^\{\\prime\}\_\{1\}/n\_\{1\}\+h\_\{2\}h^\{\\prime\}\_\{2\}/n\_\{2\}, the squared norm is
‖H\(u\)‖w2=1/2n1\+1/2n2=n1\+n22n1n2\.\\\|H^\{\(u\)\}\\\|\_\{w\}^\{2\}=\\frac\{1/2\}\{n\_\{1\}\}\+\\frac\{1/2\}\{n\_\{2\}\}=\\frac\{n\_\{1\}\+n\_\{2\}\}\{2n\_\{1\}n\_\{2\}\}\.Applying Gram\-Schmidt under⟨⋅,⋅⟩w\\langle\\cdot,\\cdot\\rangle\_\{w\}givesH~\(u\)=\[1/2,−1/2\]⊤⋅2n1n2/\(n1\+n2\)\\tilde\{H\}^\{\(u\)\}=\[1/\\sqrt\{2\},\-1/\\sqrt\{2\}\]^\{\\top\}\\cdot\\sqrt\{2n\_\{1\}n\_\{2\}/\(n\_\{1\}\+n\_\{2\}\)\}\. After spreading viavi=H~r/nrv\_\{i\}=\\tilde\{H\}\_\{r\}/n\_\{r\}:
vi\(u\)=\{\+n2n1\(n1\+n2\)ifi∈Cu\(1\),−n1n2\(n1\+n2\)ifi∈Cu\(2\),0otherwise\.v^\{\(u\)\}\_\{i\}=\\begin\{cases\}\+\\sqrt\{\\dfrac\{n\_\{2\}\}\{n\_\{1\}\(n\_\{1\}\+n\_\{2\}\)\}\}&\\text\{if \}i\\in C\_\{u\}^\{\(1\)\},\\\\\[6\.0pt\] \-\\sqrt\{\\dfrac\{n\_\{1\}\}\{n\_\{2\}\(n\_\{1\}\+n\_\{2\}\)\}\}&\\text\{if \}i\\in C\_\{u\}^\{\(2\)\},\\\\\[6\.0pt\] 0&\\text\{otherwise\}\.\\end\{cases\}This matches the PhILR balance formula exactly\(Silvermanet al\.,[2017](https://arxiv.org/html/2606.11646#bib.bib4)\)\. Thus PolyILR is a strict generalization of PhILR to arbitrary tree structures, with the binary case recovering the original method\.We note that PhILR also supports optional weighting schemes: taxon weights\(Egozcue and Pawlowsky\-Glahn,[2016](https://arxiv.org/html/2606.11646#bib.bib3)\)that modify the simplex metric itself, and branch length weights that scale coordinates\. These define a different but well\-posed weighted geometry on the simplex\(Egozcue and Pawlowsky\-Glahn,[2016](https://arxiv.org/html/2606.11646#bib.bib3)\), rather than the standard Aitchison geometry\. PolyILR is built on the standard \(unweighted\) Aitchison geometry to make the construction directly comparable to canonical ILR; integrating weighted\-simplex variants is compatible in principle and left to future work\.
## Appendix BAdditional Experiments and Details
By Nodeu1u\_\{1\}11u2u\_\{2\}2233u3u\_\{3\}u4u\_\{4\}44556677u1u\_\{1\}u2u\_\{2\}u3u\_\{3\}u4u\_\{4\}V∈ℝ7×6V\\in\\mathbb\{R\}^\{7\\times 6\}By Depthu1u\_\{1\}11u2u\_\{2\}2233u3u\_\{3\}u4u\_\{4\}44556677depth 0depth 1depth 2V∈ℝ7×6V\\in\\mathbb\{R\}^\{7\\times 6\}By Subtreeu1u\_\{1\}11u2u\_\{2\}2233u3u\_\{3\}u4u\_\{4\}44556677u1u\_\{1\}u2u\_\{2\}subtreeu3u\_\{3\}subtreeV∈ℝ7×6V\\in\\mathbb\{R\}^\{7\\times 6\}
Figure 5:*Orthogonal subspace partitions\.*The coordinates ofVVcan be grouped into orthogonal subspaces indexed by tree structure: by individual node, by depth level, or by subtree \(one example\)\. Each grouping enables inference at a different resolution\.### B\.1Microbiome Dataset and Taxonomy Construction Details
#### Datasets\.
We use two microbiome datasets\.HMP\(Human Microbiome Project; 4,743 samples, 402 taxa\) provides samples from 18 body sites with a curated NCBI\-derived taxonomy\.cMD3\(curatedMetagenomicData v3; 20,238 samples from 86 studies, 2,047 taxa\) aggregates shotgun metagenomic data across diverse cohorts\. Disease labels denote any non\-healthy diagnosis, pooled across conditions and body sites\. Note that 748HMPsamples also appear incMD3; we do not deduplicate, as our focus is methodological comparison\. TheHMPdata were obtained from theHMP16SDatapackage in R\. Operational taxonomic unit \(OTU\) abundances and sample metadata were extracted, and samples were filtered to those with complete body site annotations\. The taxonomic tree was constructed from the NCBI\-derived lineage strings provided with each OTU, parsed from Kingdom through Genus, and converted to Newick format using theapepackage\. ThecMD3metagenomic data were obtained from thecuratedMetagenomicDatapackage \(v3\.0\) in R\. We retrieved relative abundance data for all available studies, excludingIaniroG\_2022\. Taxonomic abundance matrices and sample metadata were extracted fromTreeSummarizedExperimentobjects and merged usingmergeData\. Only samples present in both the abundance matrix and metadata were retained\. Taxonomic lineages were parsed from strings of the formk\_\_Kingdom\|p\_\_Phylum\|c\_\_Class\|o\_\_Order\|f\_\_Family\|g\_\_Genus\|s\_\_Species, with rank prefixes removed\. Taxon identifiers were generated via MD5 hashing \(first 8 characters, prefixedtaxon\_\) for reproducibility\.
#### Taxonomic Trees\.
ForHMP, we use the NCBI\-derived taxonomy provided with the dataset\. ForcMD3, the taxonomic tree was constructed from lineage strings using thedata\.treepackage with a common root, converted to Newick format viaape, and assigned ultrametric branch lengths using Grafen’s method\. Tip labels were replaced with hash\-based identifiers to match the abundance matrix\. BecausecMD3aggregates studies with heterogeneous taxonomic resolution, some internal nodes have children spanning multiple taxonomic ranks; we label such nodes by their lowest common rank \(e\.g\.,*Bacteria \(Kingdom\)*or*Mixed*\)\.
### B\.2Single\-Cell Dataset and Ontology Construction Details
#### Single\-Cell Dataset\.
We useDISCO\(Database of Immune Single\-Cell Omics\), a curated atlas of human immune single\-cell transcriptomes\. Data were extracted via theDISCOtoolkitR package\. We selected samples with≥\\geq500 cells and randomly sampled up to 200 samples per condition\. We extracted blood tissue \(healthy, COVID\-19, leukemia\) and liver tissue \(healthy, hepatocellular carcinoma\)\. Cell type proportions were computed by counting cells per annotated type and normalizing to sum to one, following standard practice in single\-cell compositional analysis\(Phipsonet al\.,[2022](https://arxiv.org/html/2606.11646#bib.bib52); Buettneret al\.,[2021](https://arxiv.org/html/2606.11646#bib.bib24)\)\. After processing: blood \(199 healthy, 200 COVID\-19, 199 leukemia; 62 cell types\) and liver \(200 healthy, 153 HCC; 99 cell types\)\. For the main experiments, we use the leukemia and HCC tasks\.
#### Cell Ontology Tree\.
The cell type hierarchy was obtained from DISCO’s cell ontology API\(Liet al\.,[2022](https://arxiv.org/html/2606.11646#bib.bib49)\)\. For each tissue, we constructed a subtree by: \(1\) identifying cell types present in the data, \(2\) filtering to true leaves \(types with no children in the data\), and \(3\) tracing ancestor paths to the root\. Unit branch lengths were assigned\. The cell ontology contains extensive polytomies: blood has 14 nodes with\>\>2 children \(“T cell” has 9 children\); liver has similar structure with 99 leaves and 72 internal nodes\.
### B\.3Software and Hyperparameters
Classification models were tuned via 5\-fold cross\-validation on training data\. Grid search ranges and selected values are summarized in Table[7](https://arxiv.org/html/2606.11646#A2.T7)\. Feature importance was computed via mean decrease in impurity \(random forest\), coefficient magnitude \(logistic regression\), or permutation importance \(SVM\)\.
Table 7:Hyperparameter search grid and selected values\.ModelParameterSearch RangeUsedRandom Forestn\_estimators\{100,300,500\}\\\{100,300,500\\\}500max\_depth\{10,20,None\}\\\{10,20,\\texttt\{None\}\\\}20max\_features\{sqrt,log2\}\\\{\\texttt\{sqrt\},\\texttt\{log2\}\\\}sqrtLogistic Reg\.CC\(inverse reg\.\)\{0\.01,0\.1,1,10\}\\\{0\.01,0\.1,1,10\\\}1penalty\{ℓ1,ℓ2\}\\\{\\ell\_\{1\},\\ell\_\{2\}\\\}ℓ2\\ell\_\{2\}SVM \(RBF\)CC\{0\.1,1,10\}\\\{0\.1,1,10\\\}1γ\\gamma\{scale,0\.01,0\.1\}\\\{\\texttt\{scale\},0\.01,0\.1\\\}scaleWe compare PolyILR against: \(i\)CLR\(centered log\-ratio\), which ignores tree structure; and \(ii\)PhILR\(Silvermanet al\.,[2017](https://arxiv.org/html/2606.11646#bib.bib4)\), which requires arbitrary binary refinement of polytomies\. Both use identical zero handling\.
### B\.4Handling Zeros in PolyILR
The raw compositional data \(e\.g\., microbiome counts, single\-cell type counts\) typically exhibit high sparsity, with zeros comprising 70% or more of entries\. Since log\-ratio transformations are undefined at zero, zero replacement is required as preprocessing\. There is no consensus on the optimal choice; common choices include 0\.5, 1, or smaller values, depending on whether the raw data are counts or proportions\(Kaulet al\.,[2017](https://arxiv.org/html/2606.11646#bib.bib47); Huet al\.,[2022](https://arxiv.org/html/2606.11646#bib.bib48)\)\. Because our focus is on comparing representations rather than zero\-handling strategies, we adopt simple per\-dataset choices appropriate to each data scale as follows\.HMPandcMD3are count data:HMPprovides raw 16S counts, andcMD3\(curatedMetagenomicData v3\) is retrieved withcounts=TRUE, which multiplies relative abundances by read depth\. For both, we add a pseudocount of 1 prior to normalization, i\.e\.,x~i=\(xi\+1\)/∑j\(xj\+1\)\\tilde\{x\}\_\{i\}=\(x\_\{i\}\+1\)/\\sum\_\{j\}\(x\_\{j\}\+1\), which is commonly done at count scale\. ForcMD3, one study \(IaniroG\_2022\) was excluded because read depth metadata was unavailable\.*DISCO*provides cell\-type proportions rather than counts, so we instead useϵ=10−10\\epsilon=10^\{\-10\}applied additively before renormalization\. We note that PolyILR partially mitigates the effect of zero replacement, because each coordinate is a contrast between*geometric means over clades*rather than between individual taxa \(Section[4\.2](https://arxiv.org/html/2606.11646#S4.SS2)\), but a rigorous study is for future work\.
### B\.5PolyILR’s Robustness under Topological Noise
As discussed in Section[9](https://arxiv.org/html/2606.11646#S9), PolyILR assumes a fixed input tree𝒯\\mathcal\{T\}, and its coordinates are defined relative to𝒯\\mathcal\{T\}\. This is reasonable because PolyILR is designed for*curated scientific hierarchies*provided by domain experts \(e\.g\., phylogenies, taxonomies, or ontologies\) which can largely be trusted\. In practice, however, some tree noise or error is unavoidable\. One common form is*topological noise*, e\.g\., small misestimations of branching order or leaf placement\. Therefore, we provide a preliminary analysis of how the PolyILR basis and resulting representations respond to such local perturbations in the input tree, both empirically \(as semantic stability of selected features\) and structurally \(as how much of the basisVVis preserved\)\.
We introduce topological noise via random*nearest\-neighbor interchange*\(NNI\) operations on the NCBI taxonomy of theHMPdataset\. Each NNI swaps two subtrees across an internal edge, locally rearranging the tree without changing leaf labels\. We report perturbation strength as the absolute number of NNI applied randomly\. We use theHMPbody\-subsite task \(18 classes,d=402d=402taxa\), matching the setting of Table[2](https://arxiv.org/html/2606.11646#S5.T2)\(bottom\)\.
\#NNIsPolyILR semantic top\-10 stability00\.73±\\pm0\.1010\.71±\\pm0\.1120\.68±\\pm0\.0930\.65±\\pm0\.08Table 8:Semantic Jaccard similarity of top\-10 PolyILR features under NNI perturbations on theHMPtaxonomy \(18 body subsites, 402 taxa\)\. Mean±\\pmstd over 10 perturbed trees×\\times10 seeds\.Semantic stability under NNI\.For each NNI count, we perturb the original tree, recompute the PolyILR basisVV, retrain the downstream RF model, and re\-extract the top\-10 important features\. We then measure semantic Jaccard similarity of the top\-10 features against the unperturbed baseline, averaged over 10 perturbed trees and 10 random seeds \(Table[8](https://arxiv.org/html/2606.11646#A2.T8)\)\. Other setups mirror those used for Table[2](https://arxiv.org/html/2606.11646#S5.T2)\. At 0 NNIs \(no noise\), the result matches Table[2](https://arxiv.org/html/2606.11646#S5.T2)\(bottom\)\. As we introduce more NNIs, PolyILR degrades smoothly but*remains substantially more stable than PhILR*even at 3 NNIs: PolyILR at 3 NNIs \(0\.65\) is still well above PhILR at 0 NNIs \(0\.13, from Table[2](https://arxiv.org/html/2606.11646#S5.T2)bottom\), which suffers from binarization\-induced instability on top of any tree noise\. This shows that PolyILR remains effective even at moderate noise levels\.
Locality of NNI effects on the basisVV\.The robustness above has a structural explanation: the PolyILR basis factors node by node\. For each internal nodeuuwith childrenc1,…,ckuc\_\{1\},\\ldots,c\_\{k\_\{u\}\},VVcontains exactlyku−1k\_\{u\}\-1local basis vectors associated withuu\. Once child ordering is fixed, this local block depends only on the partition of descendant leaves induced by the children ofuuand their subtree sizes\{nr\}\\\{n\_\{r\}\\\}\(Section[4\.2](https://arxiv.org/html/2606.11646#S4.SS2)\)\. After spreading, each basis is supported only on the leaves descending fromuu, and no node’s construction references any other node’s local structure\. Therefore, a local topological change has only a local effect onVV\.
Consider an NNI at an internal edge\(p,c\)\(p,c\), whereppis the parent ofcc\. Atppandcc, the child partitions directly change, so the local coordinate blocks ofppandccinVValso change; abovepp, the swap happens entirely withinpp’s subtree, so the descendant leaf sets under each ancestor’s children, and their subtree sizes, are unchanged, so their local blocks inVVare*exactly preserved*; belowcc, the internal structure of each subtree is untouched, so again their local blocks inVVare*exactly preserved*\. This way, an NNI perturbation modifies only the coordinate blocks attached to a small set of affected nodes; all others remain preserved\.
\#NNIsCols at perturbed nodesCols changedCols preserved \(%\)118\.88\.4392\.6 / 401 \(97\.9%\)226\.613\.2387\.8 / 401 \(96\.7%\)336\.618\.6382\.4 / 401 \(95\.4%\)Table 9:NNI effects onVV\(HMP, 402 taxa, 401 coords; mean over 5 seeds\)\. We report the number of cumulative NNIs, the numbers of columns at perturbed nodes, of columns that changed, and of columns that are preserved\.We verify this empirically on theHMPtaxonomy \(402 taxa, 401 PolyILR coordinates\) above\. For each NNI count, we apply\{1,2,3\}\\\{1,2,3\\\}cumulative random NNI moves and compareVVbefore and after, averaged over 5 seeds\. In Table[9](https://arxiv.org/html/2606.11646#A2.T9), we observe that the number of changed columns never exceeds the number of columns at perturbed nodes, and most ofVVremains exactly preserved\. This empirically supports the structural argument and helps explain the mild degradation seen in Table[8](https://arxiv.org/html/2606.11646#A2.T8)\.
### B\.6Extended Classification Results
We provide a preliminary empirical illustration of the geometric connection described in Section[7](https://arxiv.org/html/2606.11646#S7)\. UsingCIFAR\-100with its known 20\-superclass hierarchy and a ResNet\-34 trained to 80\.4% test accuracy, we ask two questions: \(i\) Does the true semantic hierarchy capture structure in the model’s error distribution that arbitrary hierarchies do not? \(ii\) If so, how does this structure develop during training?
#### Setup and discussion\.
TheCIFAR\-100hierarchy has 100 fine classes grouped into 20 superclasses, yielding a two\-level tree with 19 depth\-0 coordinates \(superclass contrasts\) and 80 depth\-1 coordinates \(within\-superclass contrasts\)\. For each test sample, we compute the gradient∇𝐚ℓ=𝐕⊤\(𝐩−𝐞y\)\\nabla\_\{\\mathbf\{a\}\}\\ell=\\mathbf\{V\}^\{\\top\}\(\\mathbf\{p\}\-\\mathbf\{e\}\_\{y\}\)in PolyILR coordinates and measure how gradient norm distributes across depths\. We summarize this via*gradient entropy*:H=−∑dpdlogpdH=\-\\sum\_\{d\}p\_\{d\}\\log p\_\{d\}, wherepdp\_\{d\}is the fraction of total gradient norm at depthdd\. Lower entropy indicates concentration at specific depths, which we view as*some structure*emerging\.
*\(i\) Known tree vs\. shuffled trees\.*We compare the true hierarchy against 1000 random trees constructed by permuting leaf assignments while preserving structure\. For the trained model, the true tree yields significantly lower entropy \(0\.576 vs\.0\.632±0\.0010\.632\\pm 0\.001;p<0\.001p<0\.001\), indicating that errors concentrate at specific depths under the true hierarchy but not under arbitrary ones\. For a randomly initialized model, no difference exists \(p=0\.73p=0\.73\)\. This confirms that this structure emerges from learning rather than architectural bias\. This aligns with recent work showing that flat classifiers implicitly encode semantic hierarchies recoverable from logits alone\(Palumboet al\.,[2025](https://arxiv.org/html/2606.11646#bib.bib38)\)\.
*\(ii\): Learning dynamics\.*We track gradient distribution across 200 training epochs\. The fraction at depth 1 \(within\-superclass\) increases from 0\.68 to 0\.74 over training, consistent with the model progressively resolving coarse superclass distinctions before fine\-grained class boundaries\. This coarse\-to\-fine pattern—recently termed*hypernym bias*\(Malashinet al\.,[2025](https://arxiv.org/html/2606.11646#bib.bib53)\)—reflects curriculum\-like learning where easier \(coarser\) distinctions are learned first\.
#### Implication\.
These results suggest that PolyILR coordinates may provide a meaningful lens for analyzing softmax classifiers and steering model training when class hierarchies are available\. The gradient decomposition partly reveals where in the hierarchy a model’s errors concentrate and how this evolves during training\. We view this as opening a research direction rather than a complete empirical study, which we leave to future work\.
Figure 6:Learning dynamics onCIFAR\-100\.Left:Gradient norm distribution by tree depth across training\. Depth 0 \(superclass contrasts\) decreases while depth 1 \(within\-superclass\) increases, indicating coarse\-to\-fine learning\.Right:Gradient entropy and test accuracy over training epochs\.
### B\.7Full Experimental Results
This appendix provides complete experimental results; biological interpretation is in Section[6](https://arxiv.org/html/2606.11646#S6)\. The patterns are consistent with those discussed in the main text: PolyILR recovers interpretable taxonomic and ontological contrasts across all tasks\. Deeper biological validation—including wet\-lab experiments and large\-scale cohort studies—is beyond the scope of this methodological work\.Throughout these tables, we report point estimates alongside variability \(e\.g\., 95% confidence intervals for accuracy and AUROC, std for importance, and rank range across CV folds\), which are small in our observations and demonstrate robustness of the reported numbers\.
Table[10](https://arxiv.org/html/2606.11646#A2.T10)reports classification accuracyand AUROC with 95% CIsacross all tasks: body sites/subsites \(HMP\), westernization/age/health/body site \(cMD3\), and COVID\-19\(blood\)/leukemia\(blood\)/HCC\(liver\) \(DISCO\)\.SVM and LR are identical across CLR, PhILR, and PolyILR within each task, as expected from isometry; RF varies modestly\.Table[11](https://arxiv.org/html/2606.11646#A2.T11)reports feature stability \(Jaccard similarity of top\-KKfeatures across CV folds\); PolyILR remains stable while PhILR suffers from arbitrary binarization for both index and semantic stabilities\. Tables[12](https://arxiv.org/html/2606.11646#A2.T12)–[14](https://arxiv.org/html/2606.11646#A2.T14)list the top\-10 PolyILR contrasts by RF importance forHMP\(microbiome, body sites\),cMD3\(microbiome, health/lifestyle\), andDISCO\(single\-cell, disease\), with mean importance, std, and rank range across 5\-fold CV\. Each contrast is a log\-ratio of geometric means between two groups at a tree node\. Tables[15](https://arxiv.org/html/2606.11646#A2.T15)–[17](https://arxiv.org/html/2606.11646#A2.T17)report tree\-level aggregation results, with importance aggregated by depth \(cumulative\), subtree, node, and taxon/cell\-type levels\. Accuracy columns show predictive performance using only features at that level\.
Table 10:*Classification accuracy and AUROC with 95% confidence intervals*\(all tasks\)\. Each cell reports RF / SVM / LR\. SVM and LR are identical across CLR, PhILR, and PolyILR within each task, as expected from isometry; RF varies modestly\. Upper CI bounds clipped at 1\.00\.TaskCLRPhILRPolyILRRF / SVM / LRRF / SVM / LRRF / SVM / LR*Acc \(%\) \(95% CI\)*HMPbody sites \(5\)\.956 \(\.952–\.959\) / \.971 \(\.968–\.974\) / \.962 \(\.956–\.968\)\.961 \(\.958–\.965\) / \.971 \(\.968–\.974\) / \.962 \(\.955–\.968\)\.963 \(\.961–\.966\) / \.971 \(\.968–\.974\) / \.962 \(\.955–\.968\)body subsites \(18\)\.597 \(\.584–\.610\) / \.672 \(\.660–\.685\) / \.646 \(\.637–\.655\)\.608 \(\.588–\.627\) / \.672 \(\.660–\.685\) / \.646 \(\.638–\.654\)\.622 \(\.608–\.635\) / \.672 \(\.660–\.685\) / \.646 \(\.637–\.656\)cMD3westernized \(2\)\.972 \(\.967–\.977\) / \.979 \(\.972–\.986\) / \.968 \(\.956–\.979\)\.966 \(\.959–\.972\) / \.979 \(\.972–\.986\) / \.967 \(\.955–\.979\)\.967 \(\.962–\.972\) / \.979 \(\.972–\.986\) / \.967 \(\.956–\.979\)age category \(5\)\.785 \(\.755–\.815\) / \.814 \(\.796–\.833\) / \.739 \(\.712–\.766\)\.797 \(\.782–\.813\) / \.814 \(\.796–\.833\) / \.738 \(\.712–\.765\)\.797 \(\.784–\.811\) / \.814 \(\.796–\.833\) / \.738 \(\.711–\.765\)healthy/disease \(2\)\.692 \(\.665–\.719\) / \.701 \(\.681–\.721\) / \.664 \(\.638–\.690\)\.686 \(\.651–\.721\) / \.702 \(\.682–\.721\) / \.664 \(\.638–\.690\)\.681 \(\.648–\.714\) / \.702 \(\.682–\.721\) / \.664 \(\.638–\.690\)disease stool \(2\)\.688 \(\.668–\.707\) / \.691 \(\.662–\.721\) / \.662 \(\.623–\.701\)\.677 \(\.643–\.710\) / \.691 \(\.662–\.721\) / \.663 \(\.624–\.701\)\.674 \(\.640–\.707\) / \.691 \(\.662–\.721\) / \.663 \(\.624–\.701\)body site \(6\)\.982 \(\.976–\.988\) / \.989 \(\.984–\.994\) / \.987 \(\.983–\.991\)\.987 \(\.980–\.993\) / \.989 \(\.984–\.994\) / \.987 \(\.983–\.991\)\.987 \(\.982–\.992\) / \.989 \(\.984–\.994\) / \.987 \(\.983–\.991\)DISCOCOVID\-19 \(2\)\.747 \(\.661–\.832\) / \.777 \(\.688–\.866\) / \.737 \(\.679–\.794\)\.762 \(\.668–\.856\) / \.777 \(\.688–\.866\) / \.737 \(\.679–\.794\)\.777 \(\.688–\.865\) / \.777 \(\.688–\.866\) / \.739 \(\.683–\.796\)leukemia \(2\)\.925 \(\.883–\.967\) / \.932 \(\.893–\.971\) / \.927 \(\.880–\.974\)\.940 \(\.900–\.980\) / \.932 \(\.893–\.971\) / \.927 \(\.880–\.974\)\.935 \(\.891–\.978\) / \.932 \(\.893–\.971\) / \.927 \(\.880–\.974\)HCC \(2\)\.921 \(\.807–1\.00\) / \.927 \(\.835–1\.00\) / \.944 \(\.863–1\.00\)\.910 \(\.794–1\.00\) / \.927 \(\.835–1\.00\) / \.944 \(\.863–1\.00\)\.910 \(\.800–1\.00\) / \.927 \(\.835–1\.00\) / \.944 \(\.863–1\.00\)*AUROC \(95% CI\)*HMPbody sites \(5\)\.987 \(\.981–\.992\) / \.995 \(\.994–\.996\) / \.994 \(\.993–\.996\)\.992 \(\.991–\.994\) / \.995 \(\.994–\.996\) / \.994 \(\.993–\.996\)\.992 \(\.990–\.994\) / \.995 \(\.994–\.996\) / \.994 \(\.993–\.996\)body subsites \(18\)\.957 \(\.954–\.960\) / \.974 \(\.973–\.975\) / \.967 \(\.966–\.969\)\.965 \(\.962–\.967\) / \.974 \(\.973–\.975\) / \.967 \(\.966–\.969\)\.966 \(\.964–\.969\) / \.974 \(\.973–\.975\) / \.967 \(\.966–\.969\)cMD3westernized \(2\)\.975 \(\.965–\.986\) / \.978 \(\.958–\.999\) / \.967 \(\.940–\.995\)\.966 \(\.951–\.980\) / \.978 \(\.958–\.999\) / \.966 \(\.938–\.994\)\.966 \(\.951–\.982\) / \.978 \(\.958–\.999\) / \.967 \(\.939–\.995\)age category \(5\)\.836 \(\.802–\.870\) / \.867 \(\.838–\.895\) / \.833 \(\.811–\.855\)\.837 \(\.804–\.871\) / \.867 \(\.838–\.895\) / \.833 \(\.811–\.855\)\.843 \(\.813–\.873\) / \.867 \(\.838–\.895\) / \.833 \(\.811–\.855\)healthy/disease \(2\)\.662 \(\.634–\.691\) / \.690 \(\.665–\.715\) / \.629 \(\.598–\.660\)\.637 \(\.607–\.668\) / \.690 \(\.665–\.715\) / \.629 \(\.598–\.660\)\.631 \(\.596–\.667\) / \.690 \(\.665–\.715\) / \.629 \(\.598–\.660\)disease stool \(2\)\.663 \(\.630–\.696\) / \.683 \(\.645–\.722\) / \.633 \(\.587–\.680\)\.634 \(\.590–\.678\) / \.683 \(\.645–\.722\) / \.633 \(\.587–\.680\)\.632 \(\.586–\.677\) / \.683 \(\.645–\.722\) / \.633 \(\.587–\.680\)body site \(6\)\.945 \(\.907–\.982\) / \.994 \(\.991–\.997\) / \.996 \(\.993–\.999\)\.998 \(\.996–1\.00\) / \.994 \(\.992–\.997\) / \.996 \(\.993–\.999\)\.981 \(\.949–1\.00\) / \.994 \(\.992–\.997\) / \.996 \(\.993–\.999\)DISCOCOVID\-19 \(2\)\.808 \(\.699–\.918\) / \.897 \(\.818–\.977\) / \.847 \(\.784–\.910\)\.875 \(\.802–\.948\) / \.897 \(\.818–\.977\) / \.847 \(\.784–\.911\)\.882 \(\.806–\.957\) / \.897 \(\.817–\.977\) / \.848 \(\.785–\.911\)leukemia \(2\)\.968 \(\.946–\.990\) / \.984 \(\.969–\.999\) / \.982 \(\.963–1\.00\)\.984 \(\.964–1\.00\) / \.984 \(\.969–\.999\) / \.982 \(\.963–1\.00\)\.984 \(\.966–1\.00\) / \.984 \(\.969–\.999\) / \.982 \(\.963–1\.00\)HCC \(2\)\.921 \(\.794–1\.00\) / \.996 \(\.990–1\.00\) / \.998 \(\.996–1\.00\)\.984 \(\.960–1\.00\) / \.996 \(\.990–1\.00\) / \.998 \(\.996–1\.00\)\.974 \(\.935–1\.00\) / \.996 \(\.990–1\.00\) / \.998 \(\.996–1\.00\)
Table 11:*Feature stability*: Jaccard similarity of top\-KKfeatures across CV folds \(all tasks\)\. PolyILR stable; PhILR \(index/semantic\) unstable from arbitrary binarization\.TaskPolyILRPhILR \(index / semantic\)KK=5KK=10KK=50KK=5KK=10KK=50HMPbody sites \(5\)\.66\.65\.84\.01 / \.08\.01 / \.06\.07 / \.05body subsites \(18\)\.71\.72\.88\.01 / \.22\.02 / \.13\.07 / \.04cMD3westernized \(2\)\.43\.69\.81\.00 / \.01\.00 / \.01\.01 / \.02age category \(5\)\.73\.58\.76\.12 / \.04\.09 / \.03\.03 / \.03healthy/disease \(2\)\.56\.86\.80\.00 / \.02\.00 / \.03\.02 / \.02disease stool \(2\)\.53\.67\.80\.00 / \.02\.00 / \.03\.02 / \.03body site \(6\)\.33\.32\.62\.00 / \.00\.00 / \.00\.01 / \.01DISCOCOVID\-19 \(2\)1\.0\.87\.99\.06 / \.07\.12 / \.11\.73 / \.12leukemia \(2\)\.75\.81\.92\.05 / \.09\.12 / \.16\.73 / \.12HCC \(2\)\.77\.78\.84\.03 / \.07\.06 / \.09\.35 / \.11Table 12:*Top\-10 contrasts*by RF importance \(%\) with mean±\\pmstd and rank range across 5\-fold CV:HMP\.RkContrastImp\. \(mean±\\pmstd\)Rank rangebody sites\(5 classes\)1Streptococcaceae vs Lactobacillus \+ Leuconostocaceae3\.94±\\pm0\.0612Lactococcus vs Streptococcus3\.42±\\pm0\.162–33Pseudomonadales vs Cardiobacteriaceae \+ Vibrionaceae \+ Legionellales \+ …3\.40±\\pm0\.172–34Bacillales vs Gemella \+ Exiguobacterium \+ Turicibacter \+ …2\.95±\\pm0\.074–55Pasteurella vs Haemophilus \+ Actinobacillus \+ Aggregatibacter2\.85±\\pm0\.084–56Veillonellaceae vs Dehalobacterium \+ Fusibacter \+ Peptococcus \+ …2\.61±\\pm0\.0667Pasteurellaceae vs Cardiobacteriaceae \+ Vibrionaceae \+ Legionellales \+ …2\.37±\\pm0\.067–98Leuconostocaceae vs Lactobacillus2\.32±\\pm0\.147–99Actinomycetaceae vs Streptomyces \+ Microbispora \+ Brevibacterium \+ …2\.29±\\pm0\.087–910Propionibacteriaceae vs Streptomyces \+ Microbispora \+ Brevibacterium \+ …1\.81±\\pm0\.1710–17body subsites\(18 classes\)1Oribacterium vs Blautia \+ Coprococcus \+ Eubacterium \+ …1\.22±\\pm0\.0312Clostridiales vs Bacilli1\.10±\\pm0\.072–43Streptococcaceae vs Lactobacillus \+ Leuconostocaceae1\.08±\\pm0\.042–54Lactococcus vs Streptococcus1\.03±\\pm0\.033–55Corynebacterium vs Streptomyces \+ Microbispora \+ Brevibacterium \+ …1\.02±\\pm0\.033–56Bacillales vs Gemella \+ Exiguobacterium \+ Turicibacter \+ …0\.92±\\pm0\.026–87Propionibacteriaceae vs Streptomyces \+ Microbispora \+ Brevibacterium \+ …0\.91±\\pm0\.016–88Pseudonocardiaceae vs Streptomyces \+ Microbispora \+ Brevibacterium \+ …0\.89±\\pm0\.056–129Ruminococcaceae vs Dehalobacterium \+ Fusibacter \+ Peptococcus \+ …0\.88±\\pm0\.018–1110Propionibacterium vs Tessaracoccus0\.85±\\pm0\.047–12Table 13:*Top\-10 contrasts*by RF importance \(%\) with mean±\\pmstd and rank range across 5\-fold CV:cMD3\.RkContrastImp\. \(mean±\\pmstd\)Rank rangewesternized\(2 classes\)1Prevotella vs Alistipes \+ Anaerotruncus \+ Bacteroides \+ …1\.77±\\pm0\.1212Murimonas vs Alistipes \+ Anaerotruncus \+ Bacteroides \+ …1\.49±\\pm0\.072–43Prevotella vs Bacteroides \+ Alistipes1\.45±\\pm0\.032–54Lactobacillus vs Alistipes \+ Anaerotruncus \+ Bacteroides \+ …1\.38±\\pm0\.044–65Bacteroides vs Alistipes \+ Anaerotruncus \+ Bacteroides \+ …1\.36±\\pm0\.113–86Dialister vs Alistipes \+ Anaerotruncus \+ Bacteroides \+ …1\.32±\\pm0\.152–137Tissierellia\_unclassified vs Alistipes \+ Anaerotruncus \+ Bacteroides \+ …1\.31±\\pm0\.095–108Peptoniphilus vs Alistipes \+ Anaerotruncus \+ Bacteroides \+ …1\.28±\\pm0\.055–109Megasphaera vs Alistipes \+ Anaerotruncus \+ Bacteroides \+ …1\.23±\\pm0\.046–1010Prevotella vs Citrobacter \+ Bacteroides \+ Leuconostoc \+ …1\.18±\\pm0\.067–12age category\(5 classes\)1Mixed vs Phascolarctobacterium \+ Tychonema \+ Sediminibacterium \+ …1\.32±\\pm0\.041–22Bacteria vs Bacteria \+ Bacteria \+ Bacteria \+ …1\.24±\\pm0\.071–33Mixed vs Bacteria \+ Mixed1\.21±\\pm0\.072–34Mixed vs Phascolarctobacterium \+ Tychonema \+ Sediminibacterium \+ …0\.90±\\pm0\.0345Holdemanella vs Faecalibacterium0\.75±\\pm0\.045–106Enterococcus vs Blautia0\.73±\\pm0\.065–97Mixed vs Dialister \+ Bacteria \+ Mixed \+ …0\.69±\\pm0\.036–88Mixed vs Mixed0\.69±\\pm0\.035–99Bacteria vs Dialister \+ Bacteria \+ Mixed \+ …0\.67±\\pm0\.057–1410Mixed vs Bacteria \+ Bacteria \+ Bacteria \+ …0\.66±\\pm0\.046–11healthy vsdisease \(2\)1Lachnoclostridium vs Bacteroides \+ Turicibacter \+ Ruminococcus \+ …0\.55±\\pm0\.0612Bifidobacterium vs Blautia \+ Enterococcus0\.38±\\pm0\.042–53Flavonifractor vs Bifidobacterium \+ Actinomyces0\.35±\\pm0\.022–64Fusicatenibacter vs Gemmiger0\.34±\\pm0\.022–65Enterococcus vs Blautia0\.34±\\pm0\.023–66Lachnoclostridium vs Parabacteroides \+ Bacteroides \+ Barnesiella \+ …0\.33±\\pm0\.024–87Coprococcus vs Ruthenibacterium \+ Coprobacter \+ Oscillibacter \+ …0\.30±\\pm0\.023–98Actinomyces vs Bifidobacterium \+ Actinomyces \+ Flavonifractor0\.29±\\pm0\.026–109Streptococcus vs Enterococcus \+ Stenotrophomonas \+ Clostridioides0\.28±\\pm0\.028–1010Gemella vs Aeriscardovia \+ Enterococcus \+ Dickeya \+ …0\.27±\\pm0\.027–16stool disease\(2 classes\)1Lachnoclostridium vs Bacteroides \+ Turicibacter \+ Ruminococcus \+ …0\.58±\\pm0\.0412Gemella vs Aeriscardovia \+ Enterococcus \+ Dickeya \+ …0\.38±\\pm0\.032–53Bifidobacterium vs Blautia \+ Enterococcus0\.37±\\pm0\.022–44Fusicatenibacter vs Gemmiger0\.33±\\pm0\.024–105Lachnoclostridium vs Parabacteroides \+ Bacteroides \+ Barnesiella \+ …0\.33±\\pm0\.033–116Flavonifractor vs Bifidobacterium \+ Actinomyces0\.32±\\pm0\.025–87Enterococcus vs Blautia0\.31±\\pm0\.025–88Actinomyces vs Bifidobacterium \+ Actinomyces \+ Flavonifractor0\.31±\\pm0\.044–149Streptococcus vs Enterococcus \+ Stenotrophomonas \+ Clostridioides0\.30±\\pm0\.035–1110Coprococcus vs Ruthenibacterium \+ Coprobacter \+ Oscillibacter \+ …0\.29±\\pm0\.033–14body site\(6 classes\)1Bacillus vs Aggregatibacter \+ Prevotella \+ Leptotrichia \+ …1\.43±\\pm0\.051–22Bacteria vs Dialister \+ Bacteria \+ Mixed \+ …1\.30±\\pm0\.121–43Cutibacterium vs Parascardovia \+ Streptococcus \+ Erysipelatoclostridium \+ …1\.30±\\pm0\.052–34Arcobacter vs Aggregatibacter \+ Prevotella \+ Leptotrichia \+ …1\.21±\\pm0\.142–75Lactobacillus vs Aggregatibacter \+ Prevotella \+ Leptotrichia \+ …1\.05±\\pm0\.056–116Mixed vs Phascolarctobacterium \+ Tychonema \+ Sediminibacterium \+ …1\.04±\\pm0\.105–157Bacteria vs Bacteria1\.04±\\pm0\.034–118Cand\. Gastranaerophilales vs Aggregatibacter \+ Prevotella \+ Leptotrichia \+ …1\.02±\\pm0\.066–139Providencia vs Aggregatibacter \+ Prevotella \+ Leptotrichia \+ …1\.00±\\pm0\.088–1910Helicobacter vs Aggregatibacter \+ Prevotella \+ Leptotrichia \+ …0\.99±\\pm0\.045–14Table 14:*Top\-10 contrasts*by RF importance \(%\) with mean±\\pmstd and rank range across 5\-fold CV:DISCO\.RkContrastImp\. \(mean±\\pmstd\)Rank rangeCOVID\-19\(blood\)1Treg cell vs Naive CD4 T cell \+ Memory CD4 T cell \+ Tfh cell9\.03±\\pm0\.8812CD4 T cell vs Naive T cell \+ Memory T cell \+ CD8 T cell \+ …7\.34±\\pm0\.642–43Cycling T/NK cell vs ILC \+ T cell \+ NK cell6\.68±\\pm0\.432–44Tfh cell vs Naive CD4 T cell \+ Memory CD4 T cell6\.46±\\pm0\.283–45Plasma cell vs Pre\-GC B cell \+ B cell precursor \+ Memory B cell \+ …5\.03±\\pm0\.2656INF\-activated T cell vs Naive T cell \+ Memory T cell \+ CD8 T cell \+ …3\.96±\\pm0\.356–87Naive CD8 T cell vs Memory CD8 T cell3\.73±\\pm0\.486–88Double negative T cell vs Naive T cell \+ Memory T cell \+ CD8 T cell \+ …2\.93±\\pm0\.298–99Gamma delta T cell vs Naive T cell \+ Memory T cell \+ CD8 T cell \+ …2\.90±\\pm0\.736–1810Epithelial cell vs Immune cell \+ Neuron2\.37±\\pm0\.189–12leukemia\(blood\)1Myeloid cell vs Erythrocyte/Megakaryocyte \+ Hematopoietic precursor cell14\.50±\\pm0\.4612Cycling T/NK cell vs ILC \+ T cell \+ NK cell11\.79±\\pm0\.5423Lymphoid cell vs Erythrocyte/Megakaryocyte \+ Hematopoietic precursor \+ Myeloid6\.30±\\pm0\.453–44MAIT cell vs Naive T cell \+ Memory T cell \+ CD8 T cell \+ …5\.45±\\pm0\.513–55Naive CD8 T cell vs Memory CD8 T cell4\.73±\\pm0\.873–86Monocyte vs Granulocyte4\.22±\\pm0\.376–77T/NK cell vs B cell4\.09±\\pm0\.675–88Dendritic cell vs Granulocyte \+ Monocyte3\.13±\\pm0\.447–109Gamma delta T cell vs Naive T cell \+ Memory T cell \+ CD8 T cell \+ …2\.83±\\pm0\.337–1310Cycling T cell vs Naive T cell \+ Memory T cell \+ CD8 T cell2\.66±\\pm0\.428–14HCC\(liver\)1Erythrocyte/Megakaryocyte vs Myeloid \+ Lymphoid \+ Hematopoietic precursor10\.08±\\pm0\.3212B cell precursor vs Plasma cell \+ INF\-activated naive B cell \+ Memory B cell \+ …6\.52±\\pm0\.322–43Hematopoietic precursor cell vs Myeloid cell \+ Lymphoid cell6\.34±\\pm0\.552–34Venous EC vs LSEC5\.45±\\pm0\.852–85Granulocyte\-monocyte progenitor vs Multipotent progenitor \+ Megakaryocyte \+ …4\.37±\\pm0\.515–76Immune cell vs Cycling villous cytotrophoblast \+ CCL19/21 pericyte \+ Muscle \+ …4\.37±\\pm0\.394–87MHCII high CD14 monocyte vs MHCII low CD14 monocyte3\.97±\\pm0\.385–108Intermediate EPCAM\+ erythroblast vs Late hemoglobin\+ erythroblast3\.78±\\pm0\.406–109Alveolar macrophage vs Kupffer cell3\.52±\\pm0\.277–1010CD14 monocyte vs Placenta defensin\+ monocyte \+ Placenta CAMP\+ monocyte \+ …3\.22±\\pm0\.348–12Table 15:*Tree\-level inference*:HMP\. Importance \(mean±\\pmstd\) aggregated by depth \(cumulative\), subtree \(phylum\), node, and taxon\. Accuracy uses only features at that level\.body sites \(5\)body subsites \(18\)StructureComponentAccImpComponentAccImpDepth≤\\leq0 \(coarsest\)\.87±\\pm\.016\.8±\\pm0\.2%≤\\leq0 \(coarsest\)\.42±\\pm\.014\.8±\\pm0\.1%≤\\leq1\.92±\\pm\.0112±\\pm0\.2%≤\\leq1\.52±\\pm\.0112±\\pm0\.2%≤\\leq2\.96±\\pm\.0147±\\pm0\.3%≤\\leq2\.59±\\pm\.0344±\\pm0\.3%≤\\leq3\.96±\\pm\.0186±\\pm0\.5%≤\\leq3\.62±\\pm\.0287±\\pm0\.0%≤\\leq4 \(all\)\.96±\\pm\.00100%≤\\leq4 \(all\)\.62±\\pm\.01100%SubtreeFirmicutes\.95±\\pm\.0147±\\pm0\.7%Firmicutes\.55±\\pm\.0236±\\pm0\.3%Proteobacteria\.87±\\pm\.0120±\\pm0\.3%Proteobacteria\.44±\\pm\.0227±\\pm0\.3%Actinobacteria\.91±\\pm\.0119±\\pm0\.6%Actinobacteria\.45±\\pm\.0222±\\pm0\.2%Bacteroidetes\.82±\\pm\.016\.1±\\pm0\.2%Bacteroidetes\.35±\\pm\.017\.4±\\pm0\.1%NodeActinomycetales–11±\\pm0\.2%Actinomycetales–12±\\pm0\.2%Lactobacillales–9\.1±\\pm0\.2%Lachnospiraceae–8\.4±\\pm0\.2%Gammaproteobacteria–8\.0±\\pm0\.2%Clostridiales–3\.8±\\pm0\.1%TaxonStreptococcus–3\.1%Oribacterium–1\.0%Lactococcus–3\.1%Corynebacterium–1\.0%Pasteurella–2\.6%Lactococcus–1\.0%Lactobacillus–2\.3%Streptococcus–1\.0%Table 16:*Tree\-level inference*:cMD3\. Importance \(mean±\\pmstd\) aggregated by depth \(cumulative\) and taxon across westernization, age, healthy/disease, body site, and stool disease tasks\.westernized \(2\)age category \(5\)healthy/disease \(2\)body site \(6\)stool disease \(2\)Struct\.Comp\.AccImpComp\.AccImpComp\.AccImpComp\.AccImpComp\.AccImpDepth≤\\leq0\.930\.1%≤\\leq0\.711\.3%≤\\leq0\.610\.3%≤\\leq0\.910\.4%≤\\leq0\.580\.3%≤\\leq1\.940\.9%≤\\leq1\.764\.4%≤\\leq1\.641\.8%≤\\leq1\.983\.2%≤\\leq1\.621\.7%≤\\leq2\.941\.9%≤\\leq2\.776\.4%≤\\leq2\.653\.2%≤\\leq2\.985\.2%≤\\leq2\.623\.2%≤\\leq3\.965\.3%≤\\leq3\.7911%≤\\leq3\.676\.4%≤\\leq3\.9810%≤\\leq3\.656\.4%≤\\leq4\.969\.1%≤\\leq4\.7918%≤\\leq4\.6812%≤\\leq4\.9914%≤\\leq4\.6612%≤\\leq5\.9624%≤\\leq5\.8035%≤\\leq5\.6831%≤\\leq5\.9923%≤\\leq5\.6631%≤\\leq6\.97100%≤\\leq6\.80100%≤\\leq6\.68100%≤\\leq6\.99100%≤\\leq6\.67100%TaxonPrevotella–1\.7%Blastocystis–0\.6%Lachnoclost\.–0\.6%Bacillus–1\.4%Lachnoclost\.–0\.6%Murimonas–1\.4%Agathobaculum–0\.5%Bifidobact\.–0\.3%Cutibacterium–1\.2%Gemella–0\.4%Tissierellia–1\.4%Staphylococcus–0\.5%Enterococcus–0\.3%Malassezia–1\.1%Bifidobact\.–0\.3%Bacteroides–1\.4%Ruminococcus–0\.5%Coprococcus–0\.3%Arcobacter–1\.1%Coprococcus–0\.3%Table 17:*Tree\-level inference*:DISCO\. Importance \(mean±\\pmstd\) aggregated by depth \(cumulative\), subtree, node, and cell type across COVID\-19, leukemia, and HCC tasks\.COVID\-19 \(blood\)leukemia \(blood\)HCC \(liver\)StructureComponentAccImpComponentAccImpComponentAccImpDepth≤\\leq0\.506\.9±\\pm0\.3%≤\\leq0\.624\.6±\\pm0\.4%≤\\leq0\.888\.1±\\pm0\.7%≤\\leq1\.5611±\\pm0\.4%≤\\leq1\.8826±\\pm0\.5%≤\\leq1\.9138±\\pm1\.5%≤\\leq2\.6420±\\pm0\.6%≤\\leq2\.9143±\\pm1\.4%≤\\leq2\.9155±\\pm0\.8%≤\\leq3\.7350±\\pm0\.5%≤\\leq3\.9468±\\pm0\.9%≤\\leq3\.9177±\\pm0\.5%≤\\leq4\.7776±\\pm0\.8%≤\\leq4\.9486±\\pm1\.1%≤\\leq4\.9196±\\pm0\.3%≤\\leq5\.7797±\\pm0\.2%≤\\leq5\.9498±\\pm0\.2%≤\\leq5\.9199±\\pm0\.1%≤\\leq6 \(all\)\.77100%≤\\leq6 \(all\)\.94100%≤\\leq6 \(all\)\.91100%SubtreeImmune cell\.8093±\\pm0\.3%Immune cell\.9395±\\pm0\.4%Immune cell\.9279±\\pm1\.4%Epithelial cell\.510\.0%Fibroblast\.500\.0%Endothelial cell\.807\.8±\\pm1\.0%Fibroblast\.500\.0%Epithelial cell\.510\.0%Epithelial cell\.784\.1±\\pm0\.4%NodeT cell–23±\\pm0\.6%Immune cell–22±\\pm0\.5%Immune cell–19±\\pm1\.0%CD4 T cell–18±\\pm0\.8%T cell–16±\\pm0\.4%Hema\. precursor–9\.6±\\pm0\.8%T/NK cell–10±\\pm0\.6%T/NK cell–14±\\pm0\.3%Endothelial cell–7\.8±\\pm1\.0%B cell–7\.2±\\pm0\.6%Myeloid cell–7\.9±\\pm0\.5%B cell–7\.0±\\pm0\.4%Cell typeTreg cell–8\.5%Cycling T/NK cell–11\.3%GMP–4\.4%Cycling T/NK cell–6\.5%MAIT cell–5\.0%Intermediate erythroblast–4\.2%Tfh cell–6\.2%Naive CD8 T cell–4\.1%Late erythroblast–4\.2%Plasma cell–4\.8%Gamma delta T cell–3\.2%MHCII high monocyte–3\.6%Similar Articles
Orthogonal Dendritic Intrinsic Networks: An Architecture for Significance-Ordered, Orthogonal Latent Spaces
This paper introduces ODIN, a novel autoencoder architecture that enforces orthogonality and importance ordering of latent dimensions, recovering PCA-like interpretability in a fully non-linear regime. The method integrates geometric constraints into the training objective, theoretically grounded and empirically validated on synthetic and real-world datasets.
Riemannian Archetypal Analysis: Interpretable non-linear data analysis on deformed star distributions
This paper introduces a Riemannian version of archetypal analysis using data-driven pullback geometry to combine interpretability with non-linear expressiveness, proposing the Riemannian Archetypal Mapping (RAM) and demonstrating its effectiveness on synthetic data and MNIST.
Sparse Cholesky Elimination Tree
The article derives the column elimination tree for the right-looking sparse Cholesky algorithm, explaining how it predicts fill-in and task dependencies without performing dense factorization.
RRB-Trees: Efficient Immutable Vectors
This paper presents RRB-Trees, a data structure for efficient immutable vectors, enabling logarithmic time concatenation and slicing.
GRALIS: A Unified Canonical Framework for Linear Attribution Methods via Riesz Representation
This arXiv preprint introduces GRALIS, a unified mathematical framework using Riesz Representation Theory to formalize and compare linear attribution methods like SHAP, LIME, and Integrated Gradients.