Hierarchical Clustering Can Jointly Satisfy Richness, Consistency, and Scale Invariance

arXiv cs.LG Papers

Summary

This paper shows that hierarchical clustering can simultaneously satisfy scale invariance, richness, and consistency axioms, resolving Kleinberg's Impossibility Theorem for flat clustering. It constructs admissible hierarchical methods and analyzes their diversity and common backbone.

arXiv:2609.11173v1 Announce Type: new Abstract: Despite its ubiquity, clustering lacks a universally accepted definition of what is a cluster. Kleinberg's Impossibility Theorem formalizes this difficulty by showing that no flat clustering method can simultaneously satisfy three natural axioms: scale invariance, richness, and consistency. In this paper, we ask whether this impossibility persists when the output is a hierarchy rather than a single partition. We show that, in contrast to the flat clustering setting, the hierarchical analog of these axioms are jointly satisfiable. In fact, there exist uncountably many hierarchical clustering methods satisfying these axioms, which we call admissible. We explicitly construct several admissible methods, including methods based on well-separated clusters and a non-binary version of single linkage. For certain pairs of admissible methods, the hierarchy produced by one always refines that produced by the other. This refinement relation defines a partial order on the class of admissible methods. This partially ordered set has no greatest element and contains uncountably many pairwise incompatible maximal elements, revealing substantial diversity among admissible methods. Nevertheless, this diversity is constrained: every admissible method contains a hierarchy of sufficiently well-separated clusters, and every finite collection of admissible methods shares such a nontrivial common backbone.
Original Article
View Cached Full Text

Cached at: 09/11/26, 08:31 AM

# Hierarchical Clustering Can Jointly Satisfy Richness, Consistency, and Scale Invariance
Source: [https://arxiv.org/html/2609.11173](https://arxiv.org/html/2609.11173)
Daichi Kuroda, Maximilien Dreveton, Matthias Grossglauser, and Patrick Thiran

Daichi Kuroda daichi\.kuroda@epfl\.chAffiliation:School of Computer and Communication SciencesAffiliation:École Polytechnique Fédérale de Lausanne \(EPFL\)Affiliation:Lausanne, 1015, SwitzerlandMaximilien Dreveton maximilien\.dreveton@univ\-eiffel\.frAffiliation:LAMA, UMR\-CNRS 8050,Affiliation:Université Gustave EiffelAffiliation:5 Bd Descartes, 77454 Marne\-la\-Vallée, FranceMatthias Grossglauser matthias\.grossglauser@epfl\.chAffiliation:School of Computer and Communication SciencesAffiliation:École Polytechnique Fédérale de Lausanne \(EPFL\)Affiliation:Lausanne, 1015, SwitzerlandPatrick Thiran patrick\.thiran@epfl\.chAffiliation:School of Computer and Communication SciencesAffiliation:École Polytechnique Fédérale de Lausanne \(EPFL\)Affiliation:Lausanne, 1015, Switzerland

###### Abstract

Despite its ubiquity, clustering lacks a universally accepted definition of what is a cluster\. Kleinberg’s Impossibility Theorem formalizes this difficulty by showing that no flat clustering method can simultaneously satisfy three natural axioms: scale invariance, richness, and consistency\. In this paper, we ask whether this impossibility persists when the output is a hierarchy rather than a single partition\. We show that, in contrast to the flat clustering setting, the hierarchical analog of these axioms are jointly satisfiable\. In fact, there exist uncountably many hierarchical clustering methods satisfying these axioms, which we call admissible\. We explicitly construct several admissible methods, including methods based on well\-separated clusters and a non\-binary version of single linkage\. For certain pairs of admissible methods, the hierarchy produced by one always refines that produced by the other\. This refinement relation defines a partial order on the class of admissible methods\. This partially ordered set has no greatest element and contains uncountably many pairwise incompatible maximal elements, revealing substantial diversity among admissible methods\. Nevertheless, this diversity is constrained: every admissible method contains a hierarchy of sufficiently well\-separated clusters, and every finite collection of admissible methods shares such a nontrivial common backbone\.

††heading:23 2026 1\-1/21; Revised 5/22 9/22 21\-0000††shortheadings:Axiomatic Hierarchical Clustering / Kuroda, Dreveton, Grossglauser, and Thiran††firstpage:1††editor:My editor###### keywords

clustering; hierarchical clustering; axiomatic clustering; unsupervised learning; ultrametrics\.

## 1Introduction

Clustering is one of the most fundamental tasks in unsupervised learning\. Given pairwise dissimilarities between data points, clustering seeks to uncover meaningful group structure without access to labels or ground truth\. Yet this task is intrinsically underdetermined: there is no universally accepted definition of a cluster\. Axiomatic frameworks therefore provide a principled way to state and compare desirable properties of clustering methods\.

A central result in this direction is Kleinberg’s Impossibility Theorem\([Kleinberg, 2002](https://arxiv.org/html/2609.11173#bib.bib8)\), that shows that no flat clustering method mapping dissimilarities to partitions can simultaneously satisfy three natural axioms: scale invariance, richness, and consistency\. This theorem has had a profound influence on the theory of clustering: It implies that some trade\-off among the axioms, the problem formulation, and/or the clustering output, is unavoidable\. A large body of subsequent work explored ways to circumvent this impossibility by weakening or modifying the axioms, or by modifying the input or output space\([Ben\-David and Ackerman, 2008](https://arxiv.org/html/2609.11173#bib.bib6);[Zadeh and Ben\-David, 2009](https://arxiv.org/html/2609.11173#bib.bib15);[Strazzeri and Sánchez\-García, 2022](https://arxiv.org/html/2609.11173#bib.bib14);[Willson and Warnow, 2024](https://arxiv.org/html/2609.11173#bib.bib7)\)\. In contrast, in this paper, we ask a different question:

> *Does Kleinberg’s impossibility persist when the output is a hierarchy rather than a flat partition?*

Out of the three Kleinberg axioms,*consistency*is arguably the most contentious for the flat clustering setting\. It requires that if all within\-cluster dissimilarities \(with respect to the output partition\) are decreased and all cross\-cluster dissimilarities are increased, then the output clustering must remain unchanged\. This formalizes the intuition that strengthening the evidence for a clustering should not overturn it\. However, such transformations can also create an arbitrarily strong*substructure*within a cluster, thus revealing finer distinctions that a flat partition is unable to express\. This tension lies at the heart of Kleinberg’s impossibility result\. Figure[1](https://arxiv.org/html/2609.11173#S1.F1)illustrates this phenomenon: Starting from a single clusterC1∪C2C\_\{1\}\\cup C\_\{2\}, a permissible strengthening can make each subcluster,C1C\_\{1\}andC2C\_\{2\}, much tighter than their union, thereby revealing two new clusters that a flat clustering method that respects the consistency axiom would be forced to ignore\.

Figure 1:Strengthening the clusterC1∪C2C\_\{1\}\\cup C\_\{2\}can create two well\-separated subclusters,C1C\_\{1\}andC2C\_\{2\}, which a Kleinberg\-consistent flat clustering method is nevertheless forced to ignore\. Whereas hierarchical clustering methods can have bothC1∪C2C\_\{1\}\\cup C\_\{2\}, andC1C\_\{1\}andC2C\_\{2\}as they are properly nested\.This observation suggests moving beyond flat clustering by considering hierarchical clustering methods\. Indeed, a hierarchical clustering can preserve the desirable clusterC1∪C2C\_\{1\}\\cup C\_\{2\}while simultaneously incorporating the newly emerging subclustersC1C\_\{1\}andC2C\_\{2\}: strengthening the evidence for an existing cluster need not prevent the method from expressing finer structure within that cluster\. Formally, a hierarchical clustering method is a map from dissimilarities to hierarchies, represented as laminar families of clusters\. We formulate hierarchical analogs of scale invariance, richness, and consistency, and additionally impose permutation invariance as a basic symmetry requirement\.

Our first main result is positive: Contrary to the flat setting, these axioms are*jointly satisfiable*in the hierarchical setting\. In fact, there exist uncountably many admissible methods\. We explicitly construct several admissible methods, including hierarchies based on well\-separated clusters and a non\-binary version of single linkage\. In contrast, non\-binary variants of other classical linkage methods \(such as complete, average, Ward, centroid, and median linkage\) fail to satisfy the axioms\.

Beyond existence, we study the global structure of the family of admissible methods under the refinement order\. This family forms a remarkably diverse partially ordered set: it has uncountable height, width, and cellularity, and contains uncountably many pairwise incompatible maximal elements\. In particular, there is no greatest admissible method\. Nevertheless, the axioms impose a nontrivial common structure\. We prove a*backbone property*: every admissible method refines a hierarchy of sufficiently well\-separated clusters\. Moreover, every finite collection of admissible methods shares such a common backbone\. Thus, the axioms allow substantial diversity while still enforcing agreement on sufficiently well\-separated cluster structure\.

We then consider a stronger requirement that is specific to hierarchical clustering\. When the input dissimilarity is an ultrametric,111An ultrametric is a dissimilarity satisfying the strong triangle inequalityu⁡\(x,y\)≤max⁡\{u⁡\(x,z\),u⁡\(y,z\)\}u\(x,y\)\\leq\\max\\\{u\(x,z\),u\(y,z\)\\\}for allx,y,zx,y,z; such a dissimilarity canonically encodes a hierarchy through its nested distance balls\.it already encodes a canonical hierarchical structure\. It is therefore natural to require a hierarchical clustering method to recover this hierarchy exactly\. We call this property*exactness on ultrametrics*\. The resulting class of strongly admissible methods remains uncountable and retains the backbone and maximality phenomena described above\. However, the additional requirement sharpens the order\-theoretic structure: unlike the canonical admissible class, the strongly admissible class has a least element under refinement\.

Finally, motivated by practical clustering pipelines, we study preprocessing transformations of the input dissimilarity\. We derive general conditions under which composition with such a transformation preserves each axiom, and illustrate these principles with several common preprocessing operations\.

Our work connects several strands of the clustering literature\. It contributes to the efforts to understand and bypass Kleinberg’s impossibility theorem by modifying the axioms and/or the clustering problem formulation\([Ben\-David and Ackerman, 2008](https://arxiv.org/html/2609.11173#bib.bib6);[Cohen\-Addad et al\., 2018](https://arxiv.org/html/2609.11173#bib.bib5);[Willson and Warnow, 2024](https://arxiv.org/html/2609.11173#bib.bib7)\)\. It is also related to axiomatic and structural characterizations of hierarchical clustering methods\([Carlsson and Mémoli, 2010](https://arxiv.org/html/2609.11173#bib.bib4);[Ackerman et al\., 2010](https://arxiv.org/html/2609.11173#bib.bib10);[Ackerman and Ben\-David, 2016](https://arxiv.org/html/2609.11173#bib.bib13)\), as well as to population\-level axiomatizations of hierarchical clustering\([Thomann et al\., 2015](https://arxiv.org/html/2609.11173#bib.bib26);[Arias\-Castro and Coda, 2025](https://arxiv.org/html/2609.11173#bib.bib9)\)\.

In particular, hierarchical outputs have previously been shown to support positive axiomatic results:[Carlsson and Mémoli \(2010\)](https://arxiv.org/html/2609.11173#bib.bib4)characterize single linkage using an axiom system that differ substantially from Kleinberg’s framework, whereas[Ackerman and Ben\-David \(2016\)](https://arxiv.org/html/2609.11173#bib.bib13)characterize linkage\-based hierarchical methods using locality and a weaker consistency that only considers moving farther apart clusters that are already well\-separated\. Both works retain the numerical scales at which clusters merge, so their outputs are height\-labeled hierarchies \(dendrograms\), equivalently represented by ultrametrics\. They therefore study maps from input dissimilarities to output ultrametrics\. Instead, we consider methods returning unweighted hierarchies, which retain only the nested cluster structure\. Within this less structured output framework, our axioms more closely mirror Kleinberg’s original requirements while imposing fewer structural constraints\. Accordingly, rather than characterizing a unique method or a prescribed algorithmic family, we study the much more diverse class of hierarchical methods satisfying these axioms\. Section[6](https://arxiv.org/html/2609.11173#S6)provides a detailed comparison\.

### 1\.1Definitions and Notation

Throughout the paper,𝒳\\mathcal\{X\}is a finite set of items with cardinalitynn\. Moreover, as the labeling of the elements of𝒳\\mathcal\{X\}is irrelevant to an unsupervised task such as clustering, we implicitly assume𝒳=\[n\]\\mathcal\{X\}=\[n\], wheren=\|𝒳\|n=\|\\mathcal\{X\}\|is finite and\[n\]=\{1,…,n\}\[n\]=\\\{1,\\dots,n\\\}\. To avoid trivial cases, we always assumen≥4n\\geq 4\.222Forn≤2n\\leq 2, a hierarchy can contain only the root and singleton leaves and is therefore necessarily the star hierarchy\. Forn=3n=3, there is only one non\-star hierarchy up to relabeling\.A*dissimilarity function*on𝒳\\mathcal\{X\}is a mapd:𝒳×𝒳→R≥0d\\colon\\mathcal\{X\}\\times\\mathcal\{X\}\\to\\mathbb\{R\}\_\{\\geq 0\}such thatd⁡\(x,x\)=0d\(x,x\)=0andd⁡\(x,y\)=d⁡\(y,x\)\>0d\(x,y\)=d\(y,x\)\>0for all distinctx,y∈𝒳x,y\\in\\mathcal\{X\}\. We do not assume thatddobeys the triangle inequality, henceddis not necessarily a distance\. We denote by𝒟⁡\(𝒳\)\\mathcal\{D\}\(\\mathcal\{X\}\)the set of all dissimilarity functions on𝒳\\mathcal\{X\}\. We use the conventionmin⁡∅=∞\\min\\emptyset=\\infty\.

A*cluster*CCis a nonempty subset of𝒳\\mathcal\{X\}, and a*partition*of𝒳\\mathcal\{X\}is a set𝒞=\{C1,…,Ck\}\\mathcal\{C\}=\\\{C\_\{1\},\\dots,C\_\{k\}\\\}of pairwise disjoint clusters whose union is𝒳\\mathcal\{X\}\. LetC¯:=𝒳∖C\\bar\{C\}:=\\mathcal\{X\}\\setminus Cdenote the complement of the clusterCCwith respect to𝒳\\mathcal\{X\}\. We write𝒫⁡\(𝒳\)\\mathcal\{P\}\(\\mathcal\{X\}\)for the set of all partitions of𝒳\\mathcal\{X\}\.

### 1\.2Structure of the Paper

The remainder of the paper is organized as follows\. In Section[2](https://arxiv.org/html/2609.11173#S2), we state the main axioms and establish the corresponding achievability results\. In Section[3](https://arxiv.org/html/2609.11173#S3), we present several admissible hierarchical clustering methods\. In Section[4](https://arxiv.org/html/2609.11173#S4), we study the structural properties of the set of admissible methods\. In Section[5](https://arxiv.org/html/2609.11173#S5), we add exactness on ultrametrics to the axiom system and characterize preprocessing transformations that preserve the axioms\. In Section[6](https://arxiv.org/html/2609.11173#S6), we discuss related work and, in Section[7](https://arxiv.org/html/2609.11173#S7), conclude the paper\. Omitted proofs and technical lemmas are provided in the Appendix\.

### 1\.3Use of Large Language Models \(LLM\)

Whereas the conceptualization of the project underlying this paper was carried out entirely by the authors, we also used GPT\-5\.6 Sol to improve clarity, wording, presentation, and to assist in identifying and correcting minor errors and typos in earlier drafts\. However, the vast majority of the mathematical content and proofs were generated by us\. The only mathematical results for which LLMs were used to develop proof techniques are Proposition[13](https://arxiv.org/html/2609.11173#Thmtheorem13), Lemmas[16](https://arxiv.org/html/2609.11173#Thmtheorem16)and[55](https://arxiv.org/html/2609.11173#Thmtheorem55), and Proposition[40](https://arxiv.org/html/2609.11173#Thmtheorem40)\(iv\)\. We also used the model to assist with the numerical simulations in Appendix[E](https://arxiv.org/html/2609.11173#A5)\. All text proposed by LLM was rigorously checked and edited by the authors, and we take full responsibility for this article\.

## 2From Impossibility to Achievability

In this section, we first recall Kleinberg’s impossibility theorem for flat clustering and then show that, in contrast, the hierarchical analogs of his axioms are jointly satisfiable\.

### 2\.1Kleinberg’s Impossibility Theorem for Flat Clustering

We begin by recalling the axiomatic framework for flat clustering, introduced by[Kleinberg \(2002\)](https://arxiv.org/html/2609.11173#bib.bib8)\. A*flat clustering method*is a map

f:𝒟⁡\(𝒳\)→𝒫⁡\(𝒳\)f\\colon\\mathcal\{D\}\(\\mathcal\{X\}\)\\to\\mathcal\{P\}\(\\mathcal\{X\}\)such thatf⁡\(d\)f\(d\)is a partition of𝒳\\mathcal\{X\}for every dissimilarityd∈𝒟⁡\(𝒳\)d\\in\\mathcal\{D\}\(\\mathcal\{X\}\)\.

We first introduce a transformation of dissimilarities, called strengthening333This transformation is called aΓ\\Gamma\-transformation in[Kleinberg \(2002\)](https://arxiv.org/html/2609.11173#bib.bib8), whereΓ\\Gammais a partition\.that reinforces a given clustering\. Intuitively, a strengthening moves points within the same cluster closer to each other, and points in different clusters further apart\.

###### Definition 1\(𝒞\\mathcal\{C\}\-strengthening\)\.

Let𝒞∈𝒫⁡\(𝒳\)\\mathcal\{C\}\\in\\mathcal\{P\}\(\\mathcal\{X\}\)andd,d′∈𝒟⁡\(𝒳\)d,d^\{\\prime\}\\in\\mathcal\{D\}\(\\mathcal\{X\}\)be two dissimilarity functions\. We say thatd′∈𝒟⁡\(𝒳\)d^\{\\prime\}\\in\\mathcal\{D\}\(\\mathcal\{X\}\)is a𝒞\\mathcal\{C\}\-strengthening ofddif for all clustersC∈𝒞C\\in\\mathcal\{C\},

- •\(Intra\-cluster contraction\)d′​\(x,y\)≤d⁡\(x,y\)d^\{\\prime\}\(x,y\)\\leq d\(x,y\)for allx,y∈Cx,y\\in C;
- •\(Inter\-cluster expansion\)d′​\(x,y\)≥d⁡\(x,y\)d^\{\\prime\}\(x,y\)\\geq d\(x,y\)for allx∈Cx\\in C,y∉Cy\\notin C\.

We now formalize the properties that a clustering method should satisfy\. These axioms capture natural requirements such as invariance to scaling, expressiveness, and stability under strengthening\.

###### Definition 2\.

A flat clustering methodffis:

- •scale invariantiff⁡\(β​d\)=f⁡\(d\)f\(\\beta\\,d\)=f\(d\)for everyd∈𝒟⁡\(𝒳\)d\\in\\mathcal\{D\}\(\\mathcal\{X\}\)and everyβ\>0\\beta\>0;
- •richifffis surjective, that is, for every𝒞∈𝒫⁡\(𝒳\)\\mathcal\{C\}\\in\\mathcal\{P\}\(\\mathcal\{X\}\), there existsd∈𝒟⁡\(𝒳\)d\\in\\mathcal\{D\}\(\\mathcal\{X\}\)such thatf⁡\(d\)=𝒞f\(d\)=\\mathcal\{C\};
- •consistentiff⁡\(d\)=f⁡\(d′\)f\(d\)=f\(d^\{\\prime\}\)for everyd∈𝒟⁡\(𝒳\)d\\in\\mathcal\{D\}\(\\mathcal\{X\}\)and everyf⁡\(d\)f\(d\)\-strengtheningd′d^\{\\prime\}ofdd\.

These axioms appear natural when considered individually, although the consistency axiom is sometimes viewed as contentious because it permits the creation of new clusters \(see Section[1](https://arxiv.org/html/2609.11173#S1)\)\. Kleinberg established that these three requirements are not jointly satisfiable\.

###### Theorem 3\([Kleinberg \(2002\)](https://arxiv.org/html/2609.11173#bib.bib8)\)\.

There is no flat clustering methodffthat is simultaneously scale invariant, rich, and consistent\.

### 2\.2Achievability of Hierarchical Clustering

We now turn to hierarchical clustering\. Unlike flat clustering, which outputs a single partition of the dataset, hierarchical clustering outputs*nested clusters*\. Formally, a hierarchy on𝒳\\mathcal\{X\}is a rooted tree whose leaves are the singleton sets\{x\}\\\{x\\\}forx∈𝒳x\\in\\mathcal\{X\}and whose root is the full set𝒳\\mathcal\{X\}\. We represent such a tree by its associated nested set of clusters\.

###### Definition 4\(Hierarchy\)\.

Let2𝒳2^\{\\mathcal\{X\}\}denote the power set of𝒳\\mathcal\{X\}\. A setΨ⊆2𝒳∖\{∅\}\\Psi\\subseteq 2^\{\\mathcal\{X\}\}\\setminus\\\{\\emptyset\\\}is a*hierarchy*on𝒳\\mathcal\{X\}if it satisfies:

1. 1\.*Laminarity*: For allC1,C2∈ΨC\_\{1\},C\_\{2\}\\in\\Psi, we have eitherC1∩C2=∅C\_\{1\}\\cap C\_\{2\}=\\emptyset, orC1⊆C2C\_\{1\}\\subseteq C\_\{2\}, orC2⊆C1C\_\{2\}\\subseteq C\_\{1\}\.
2. 2\.*Root and leaves:*𝒳∈Ψ\\mathcal\{X\}\\in\\Psiand\{x\}∈Ψ\\\{x\\\}\\in\\Psifor everyx∈𝒳x\\in\\mathcal\{X\}\.

We denote by𝒯⁡\(𝒳\)\\mathcal\{T\}\(\\mathcal\{X\}\)the set of all hierarchies on𝒳\\mathcal\{X\}\.

###### Example 5\.

Let𝒳=\{1,2,3,4\}\\mathcal\{X\}=\\\{1,2,3,4\\\}\. The setΨ=\{\{1,2,3,4\},\{1,2\},\{3,4\},\{1\},\{2\},\{3\},\{4\}\}\\Psi=\\big\\\{\\\{1,2,3,4\\\},\\\{1,2\\\},\\\{3,4\\\},\\\{1\\\},\\\{2\\\},\\\{3\\\},\\\{4\\\}\\big\\\}is a hierarchy representing a binary tree with root\{1,2,3,4\}\\\{1,2,3,4\\\}and two children\{1,2\}\\\{1,2\\\}and\{3,4\}\\\{3,4\\\}\.

A*hierarchical clustering method*is a map

T:𝒟⁡\(𝒳\)→𝒯⁡\(𝒳\)T\\colon\\mathcal\{D\}\(\\mathcal\{X\}\)\\to\\mathcal\{T\}\(\\mathcal\{X\}\)such that for everyd∈𝒟⁡\(𝒳\)d\\in\\mathcal\{D\}\(\\mathcal\{X\}\),T⁡\(d\)T\(d\)is a hierarchy on𝒳\\mathcal\{X\}\. Kleinberg’s axioms of scale invariance, consistency, and richness admit natural analog in the hierarchical setting\.

###### Definition 6\.

A hierarchical clustering methodTTon𝒳\\mathcal\{X\}is

1. 1\.scale invariantifT⁡\(β​d\)=T⁡\(d\)T\(\\beta\\,d\)=T\(d\)for everyβ\>0\\beta\>0and everyd∈𝒟⁡\(𝒳\)d\\in\\mathcal\{D\}\(\\mathcal\{X\}\);
2. 2\.partition richif for every partition𝒞∈𝒫⁡\(𝒳\)\\mathcal\{C\}\\in\\mathcal\{P\}\(\\mathcal\{X\}\), there existsd∈𝒟⁡\(𝒳\)d\\in\\mathcal\{D\}\(\\mathcal\{X\}\)such that𝒞⊆T⁡\(d\)\\mathcal\{C\}\\subseteq T\(d\);
3. 3\.partition consistentif for everyd∈𝒟⁡\(𝒳\)d\\in\\mathcal\{D\}\(\\mathcal\{X\}\)and every partition𝒞∈𝒫⁡\(𝒳\)\\mathcal\{C\}\\in\\mathcal\{P\}\(\\mathcal\{X\}\)such that𝒞⊆T⁡\(d\)\\mathcal\{C\}\\subseteq T\(d\), we have𝒞⊆T⁡\(d′\)\\mathcal\{C\}\\subseteq T\(d^\{\\prime\}\)for every𝒞\\mathcal\{C\}\-strengtheningd′d^\{\\prime\}ofdd*444For a hierarchy represented as a rooted tree, partitions𝒞⊆T⁡\(d\)\\mathcal\{C\}\\subseteq T\(d\)are in bijection with vertex cuts separating the root from the leaves: the vertices in the cut are precisely the nodes corresponding to the clusters in𝒞\\mathcal\{C\}\.*;
4. 4\.permutation invariantifT⁡\(dϕ\)=ϕ⋅T⁡\(d\)T\(d\_\{\\phi\}\)\\ =\\ \\phi\\cdot T\(d\)for every permutationϕ\\phiof𝒳\\mathcal\{X\}, wheredϕ​\(x,y\):=d⁡\(ϕ−1​\(x\),ϕ−1​\(y\)\)d\_\{\\phi\}\(x,y\):=d\(\\phi^\{\-1\}\(x\),\\phi^\{\-1\}\(y\)\)andϕ⋅T⁡\(d\)=\{\{ϕ⁡\(x\):x∈C\}:C∈T⁡\(d\)\}\\phi\\cdot T\(d\)=\\\{\\\{\\phi\(x\)\\colon x\\in C\\\}\\colon C\\in T\(d\)\\\}\.

A hierarchical clustering methodTTis*admissible*if it satisfies all four axioms\.

Definition[6](https://arxiv.org/html/2609.11173#Thmtheorem6)is a direct extension of Kleinberg’s original axioms to the hierarchical setting\. The flat clustering methodffis replaced by a hierarchical clustering methodTT, and the three axioms of Definition[2](https://arxiv.org/html/2609.11173#Thmtheorem2)are reformulated accordingly\. We also explicitly include permutation invariance, requiring that the output depends only on the underlying dissimilarity structure and not on the labeling of the points\. In Kleinberg’s original framework, this requirement is implicit, as clusterings are defined directly on sets rather than on a labeled representation\.

We refer to the set of four axioms in Definition[6](https://arxiv.org/html/2609.11173#Thmtheorem6)as the canonical axiom systemαadm\\alpha\_\{\\mathrm\{adm\}\}, and we call*admissible*the hierarchical clustering methods that satisfyαadm\\alpha\_\{\\mathrm\{adm\}\}\. There are, however, other reasonable axiomatizations of hierarchical clustering\. In particular, there is a well\-known correspondence between ultrametrics and hierarchies, which is not addressed by the axioms ofαadm\\alpha\_\{\\mathrm\{adm\}\}because it has no direct counterpart in flat clustering\. In Section[5\.1](https://arxiv.org/html/2609.11173#S5.SS1), we study the consequences of such an additional requirement\. Moreover, further alternative axioms are discussed in Appendix[A](https://arxiv.org/html/2609.11173#A1), where we show that most of the results established for this canonical axiom system in fact remain valid under these alternative formulations\. All the axioms and their alternatives are cataloged in Table[1](https://arxiv.org/html/2609.11173#A1.T1)\.

The first main result of this paper is that the hierarchical clustering problem does admit \(many\) methods that satisfyαadm\\alpha\_\{\\mathrm\{adm\}\}, in contrast with flat clustering\.

###### Theorem 7\.

There exist uncountably many admissible methods\.

The proof of Theorem[7](https://arxiv.org/html/2609.11173#Thmtheorem7)is constructive\. In the next section, we give explicit examples of admissible hierarchical clustering methods\.

## 3Explicit Construction of Admissible Methods

In this section, we construct admissible methods from three complementary perspectives\. We first examine the admissibility of some standardlinkage methods, as they are the most widely used class of hierarchical clustering methods\. We then introduce two families of methods defined through explicit separation conditions and finally study Bryant\-Berry stable clusters\.

These two parameterized families play a central role in the analysis of the class of admissible hierarchical clustering methods made in Section[4](https://arxiv.org/html/2609.11173#S4): the separation methods will be useful to order the admissible methods based on the cluster structures they produce, whereas the Bryant\-Berry cluster construction will be useful for showing the existence of many admissible methods that are fundamentally different \(more precisely, incompatible\)\.

### 3\.1Linkage Methods

Linkage\-based methods form a classical and widely used class of hierarchical clustering algorithms\. They are typically defined as greedy agglomerative procedures: Starting from the singleton partition, clusters are iteratively merged according to a prescribed inter\-cluster dissimilarity rule, such as single, complete, average, or Ward linkage\. We refer to Appendix[C\.2](https://arxiv.org/html/2609.11173#A3.SS2)for formal definitions of these classic linkage rules\. Standard formulations of linkage algorithms merge exactly two clusters at each step, and hence always produce binary hierarchies\.555A hierarchyΨ⊆2𝒳\\Psi\\subseteq 2^\{\\mathcal\{X\}\}on𝒳\\mathcal\{X\}is a*binary hierarchy*if, for every non\-singleton clusterC∈ΨC\\in\\Psi, there exist exactly two proper subsetsC1,C2∈ΨC\_\{1\},C\_\{2\}\\in\\Psisuch that: \(i\)C1,C2⊊CC\_\{1\},C\_\{2\}\\subsetneq C; \(ii\)C1∩C2=∅C\_\{1\}\\cap C\_\{2\}=\\emptyset; \(iii\)C1∪C2=CC\_\{1\}\\cup C\_\{2\}=C\.

We first observe that this restriction to pairwise merges and hence to binary hierarchies is incompatible with permutation invariance, regardless of the specific linkage rule employed\.

###### Lemma 8\.

Any hierarchical clustering method that is constrained to always produce a binary hierarchy is not permutation invariant\.

###### Proof\.

Given a dissimilaritydd, we say that two elementsx≠yx\\neq yhave pairwise identical dissimilarity profiles ifxxandyyare equidistant from every other element of𝒳\\mathcal\{X\}\(i\.e\.,∀z∈𝒳∖\{x,y\}:d⁡\(x,z\)=d⁡\(y,z\)\\forall z\\in\\mathcal\{X\}\\setminus\\\{x,y\\\}\\colon\\quad d\(x,z\)=d\(y,z\)\)\. Letddbe a dissimilarity on𝒳\\mathcal\{X\}, and suppose that𝒳\\mathcal\{X\}has three distinct elementsx1x\_\{1\},x2x\_\{2\}andx3x\_\{3\}, such that any two have pairwise identical profiles\. As the elements ofS:=\{x1,x2,x3\}S:=\\\{x\_\{1\},x\_\{2\},x\_\{3\}\\\}have pairwise identical dissimilarity profiles, every permutationσ\\sigmaofSSthat fixes𝒳∖S\\mathcal\{X\}\\setminus Spreservesdd\. Permutation invariance therefore requires thatC∈T⁡\(d\)⟹σ⁡\(C\)∈T⁡\(d\)\.C\\in T\(d\)\\Longrightarrow\\sigma\(C\)\\in T\(d\)\.

LetC∈T⁡\(d\)C\\in T\(d\)be non\-singleton and suppose thatC∩S≠∅C\\cap S\\neq\\varnothing\. IfC∩S=\{xi\}C\\cap S=\\\{x\_\{i\}\\\}, thenC=U∪\{xi\}C=U\\cup\\\{x\_\{i\}\\\}for some nonemptyUU, and exchangingxix\_\{i\}with another element ofSSproduces a clusterU∪\{xj\}∈T⁡\(d\)U\\cup\\\{x\_\{j\}\\\}\\in T\(d\)incompatible withCC\. If\|C∩S\|=2\|C\\cap S\|=2, exchanging one of these two elements with the third likewise produces a cluster incompatible withCC\. Both cases contradict laminarity\. Hence every non\-singleton cluster having a non\-empty intersection withSScontains all elements ofSS, and thusTTmust be non\-binary\. ∎

Motivated by Lemma[8](https://arxiv.org/html/2609.11173#Thmtheorem8), we consider variants of linkage methods whose output is not constrained to binary hierarchies\. In these variants, whenever multiple pairs of clusters attain the same minimal inter\-cluster dissimilarity, all clusters involved in this tie are merged together simultaneously\. This produces hierarchies that are not necessarily binary, but that respect the symmetries of the input dissimilarity function\. Among the classic linkage rules mentioned earlier, single linkage is the only one that produces admissible hierarchies, as established by the following proposition\.

###### Proposition 9\(Admissibility of Linkage Methods\)\.

The non\-binary variant of the single linkage methodTSLT\_\{\\mathrm\{SL\}\}is admissible\. In contrast, non\-binary variants of the complete, average, Ward, centroid, and median linkage violate at least one of the axioms in Definition[6](https://arxiv.org/html/2609.11173#Thmtheorem6)\.

###### Proof\.

We prove the admissibility ofTSLT\_\{\\mathrm\{SL\}\}in Appendix[C\.1](https://arxiv.org/html/2609.11173#A3.SS1), and the non\-admissibility of other linkage methods in Appendix[C\.2](https://arxiv.org/html/2609.11173#A3.SS2)\. ∎

### 3\.2Separation Methods

We now introduce two classes of admissible hierarchical clustering methods, based on explicit separation conditions between within\-cluster and cross\-cluster dissimilarities\. These methods declare a nonempty subsetC⊆𝒳C\\subseteq\\mathcal\{X\}to be a cluster if the internal dissimilarities withinCCare sufficiently small compared to the dissimilarities between elements ofCCand of its complementC¯\\bar\{C\}\. The resulting hierarchies are defined directly fromdd, without resorting to an agglomerative procedure\. We quantify the separability of a clusterCCby the ratios

d⁡\(x1,y\)d⁡\(x2,z\),x1,x2,y∈C,z∉C,\\frac\{d\(x\_\{1\},y\)\}\{d\(x\_\{2\},z\)\},\\qquad x\_\{1\},x\_\{2\},y\\in C,\\ z\\notin C,and we consider below two families of methods resulting from thresholding these ratios\. The two families differ in the choice of the reference pointsx1,x2∈Cx\_\{1\},x\_\{2\}\\in C: the first family enforces a worst\-case comparison globally for any pairx1,x2∈Cx\_\{1\},x\_\{2\}\\in C, whereas the second family enforces this comparison locally by settingx1=x2x\_\{1\}=x\_\{2\}\.

###### Definition 10\.

For any finite set𝒳\\mathcal\{X\}, define the global and local separabilities with respect tod∈𝒟⁡\(𝒳\)d\\in\\mathcal\{D\}\(\\mathcal\{X\}\)and a proper subsetC⊊𝒳C\\subsetneq\\mathcal\{X\}with\|C\|≥2\|C\|\\geq 2, respectively by

ϱ⁡\(d,C\):=maxx1,x2,y∈C,z∈C¯⁡d⁡\(x1,y\)d⁡\(x2,z\)andτ⁡\(d,C\):=maxx,y∈C,z∈C¯⁡d⁡\(x,y\)d⁡\(x,z\)\.\\displaystyle\\varrho\(d,C\)\\,:=\\,\\max\_\{\\begin\{subarray\}\{c\}x\_\{1\},x\_\{2\},y\\in C,z\\in\\bar\{C\}\\end\{subarray\}\}\\frac\{d\(x\_\{1\},y\)\}\{d\(x\_\{2\},z\)\}\\quad\\text\{ and \}\\quad\\tau\(d,C\)\\,:=\\,\\max\_\{\\begin\{subarray\}\{c\}x,y\\in C,z\\in\\bar\{C\}\\end\{subarray\}\}\\frac\{d\(x,y\)\}\{d\(x,z\)\}\.We call*separation margin sequence*any sequence𝛈=\(ηm\)1≤m≤n−2\\boldsymbol\{\\eta\}=\(\\eta\_\{m\}\)\_\{1\\leq m\\leq n\-2\}such that0<ηs≤10<\\eta\_\{s\}\\leq 1for everyss\. For such sequence𝛈\\boldsymbol\{\\eta\}, we define the set of𝛈\\boldsymbol\{\\eta\}\-globally and𝛈\\boldsymbol\{\\eta\}\-locally separated clusters on𝒳\\mathcal\{X\}, respectively as

Tglob𝜼​\(d\)\\displaystyle T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}\(d\):=\{C⊊𝒳:\|C\|≥2,ϱ\(d,C\)<η\|C\|−1\}∪\{\{x\}:x∈𝒳\}∪\{𝒳\};\\displaystyle:=\\left\\\{C\\subsetneq\\mathcal\{X\}:\|C\|\\geq 2,\\,\\varrho\(d,C\)<\\eta\_\{\|C\|\-1\}\\right\\\}\\cup\\\{\\\{x\\\}:x\\in\\mathcal\{X\}\\\}\\cup\\\{\\mathcal\{X\}\\\};Tloc𝜼​\(d\)\\displaystyle T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}\(d\):=\{C⊊𝒳:\|C\|≥2,τ\(d,C\)<η\|C\|−1\}∪\{\{x\}:x∈𝒳\}∪\{𝒳\}\.\\displaystyle:=\\left\\\{C\\subsetneq\\mathcal\{X\}:\|C\|\\geq 2,\\,\\tau\(d,C\)<\\eta\_\{\|C\|\-1\}\\right\\\}\\cup\\\{\\\{x\\\}:x\\in\\mathcal\{X\}\\\}\\cup\\\{\\mathcal\{X\}\\\}\.

The global \(resp\., local\) separabilityϱ⁡\(d,C\)\\varrho\(d,C\)\(resp\.,τ⁡\(d,C\)\\tau\(d,C\)\) measures the sensitivity of the global \(resp\., local\) separability of the clusterCCto the dissimilarity functiondd\. The following proposition establishes that the mapsTglob𝜼:d↦Tglob𝜼​\(d\)T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}\\colon d\\mapsto T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}\(d\)andTloc𝜼:d↦Tloc𝜼​\(d\)T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}\\colon d\\mapsto T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}\(d\)are admissible hierarchical clustering methods\.

###### Proposition 11\.

For any separation margin sequence𝛈\\boldsymbol\{\\eta\},Tglob𝛈T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}andTloc𝛈T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}are admissible\. Furthermore, if𝛈≠𝛈′\\boldsymbol\{\\eta\}\\neq\\boldsymbol\{\\eta\}^\{\\prime\}, thenTglob𝛈≠Tglob𝛈′T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}\\neq T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}^\{\\prime\}\}andTloc𝛈≠Tloc𝛈′T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}\\neq T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}^\{\\prime\}\}\.

When𝜼=𝟏=\(1,⋯,1\)\\boldsymbol\{\\eta\}=\\boldsymbol\{1\}=\(1,\\cdots,1\), we simply writeTglob,TlocT\_\{\\mathrm\{glob\}\},T\_\{\\mathrm\{loc\}\}instead ofTglob𝟏,Tloc𝟏T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{1\}\},T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{1\}\}\. The clusters belonging toTglob​\(d\)T\_\{\\mathrm\{glob\}\}\(d\)andTloc​\(d\)T\_\{\\mathrm\{loc\}\}\(d\)are the sets of points that are strictly closer to any point in the set than to any point outside the set\. Among these methods, the most studied one isTloc​\(d\)T\_\{\\mathrm\{loc\}\}\(d\), which is called the*Apresjan hierarchy*\([Apresjan, 1966](https://arxiv.org/html/2609.11173#bib.bib22)\)\. The clusters belonging toTloc​\(d\)T\_\{\\mathrm\{loc\}\}\(d\)have been studied independently by different authors, and are referred to as Apresjan clusters,KK\-clumps, strong clusters, nice clusters, or valid clusters in the literature\([Ackerman and Dasgupta, 2014](https://arxiv.org/html/2609.11173#bib.bib18);[Diatta and Fichet, 1994](https://arxiv.org/html/2609.11173#bib.bib19);[Balcan et al\., 2008](https://arxiv.org/html/2609.11173#bib.bib11);[Bryant and Berry, 2001](https://arxiv.org/html/2609.11173#bib.bib20)\)\.

The separation margins𝜼\\boldsymbol\{\\eta\}control the sensitivity to separabilityϱ⁡\(d,C\)\\varrho\(d,C\)andτ⁡\(d,C\)\\tau\(d,C\)\. The higher the values of𝜼\\boldsymbol\{\\eta\}are, the more sensitive these methods become to close\-to\-one separabilities, but at the price of being too prone to noise\. For example, consider𝒳=\{1,2,3,4\}\\mathcal\{X\}=\\\{1,2,3,4\\\}andd∈𝒟⁡\(𝒳\)d\\in\\mathcal\{D\}\(\\mathcal\{X\}\)such that every pairwise dissimilarity is exactly equal to 1\. For thisdd, because no nontrivial subset satisfies the global separation condition,TglobT\_\{\\mathrm\{glob\}\}identifies only the root and leaves as clusters\. However, ifd⁡\(1,2\)d\(1,2\)is slightly decreased to be equal to1−ε1\-\\varepsilon, even for an extremely smallε\>0\\varepsilon\>0such asε=10−10\\varepsilon=10^\{\-10\}, then\{1,2\}\\\{1,2\\\}belongs to the hierarchy returned byTglobT\_\{\\mathrm\{glob\}\}\. But, setting𝜼=\(1−ε′\)⋅𝟏\\boldsymbol\{\\eta\}=\(1\-\\varepsilon^\{\\prime\}\)\\cdot\\mathbf\{1\}withε′\>ε\\varepsilon^\{\\prime\}\>\\varepsilon, the cluster\{1,2\}\\\{1,2\\\}does not belong to the hierarchyTglob𝜼T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}, asTglob𝜼T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}outputs only the root and leaves\.

###### Example 12\.

Let𝒳=\{1,2,3,4,5\}\\mathcal\{X\}=\\\{1,2,3,4,5\\\}and

d=\(0124310456240564550736670\)\.\\displaystyle d=\\begin\{pmatrix\}0&1&2&4&3\\\\ 1&0&4&5&6\\\\ 2&4&0&5&6\\\\ 4&5&5&0&7\\\\ 3&6&6&7&0\\end\{pmatrix\}\.\(1\)Then, as shown in Figure[2](https://arxiv.org/html/2609.11173#S3.F2), the outputs ofTSL,Tloc,TglobT\_\{\\mathrm\{SL\}\},T\_\{\\mathrm\{loc\}\},T\_\{\\mathrm\{glob\}\}, andTglob𝛈T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}with𝛈=12​𝟏\\boldsymbol\{\\eta\}=\\frac\{1\}\{2\}\\boldsymbol\{1\}areTSL​\(d\)=\{\{1,2,3,5\},\{1,2,3\},\{1,2\}\}∪\{𝒳\}∪\{\{x\}:x∈𝒳\};T\_\{\\mathrm\{SL\}\}\(d\)=\\\{\\\{1,2,3,5\\\},\\\{1,2,3\\\},\\\{1,2\\\}\\\}\\cup\\\{\\mathcal\{X\}\\\}\\cup\\\{\\\{x\\\}:x\\in\\mathcal\{X\}\\\};Tloc​\(d\)=\{\{1,2,3\},\{1,2\}\}∪\{𝒳\}∪\{\{x\}:x∈𝒳\};T\_\{\\mathrm\{loc\}\}\(d\)=\\\{\\\{1,2,3\\\},\\\{1,2\\\}\\\}\\cup\\\{\\mathcal\{X\}\\\}\\cup\\\{\\\{x\\\}:x\\in\\mathcal\{X\}\\\};Tglob​\(d\)=\{\{1,2\}\}∪\{𝒳\}∪\{\{x\}:x∈𝒳\};T\_\{\\mathrm\{glob\}\}\(d\)=\\\{\\\{1,2\\\}\\\}\\cup\\\{\\mathcal\{X\}\\\}\\cup\\\{\\\{x\\\}:x\\in\\mathcal\{X\}\\\};Tglob𝛈\(d\)=\{\{𝒳\}∪\{\{x\}:x∈𝒳\}\.T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}\(d\)=\\\{\\\{\\mathcal\{X\}\\\}\\cup\\\{\\\{x\\\}:x\\in\\mathcal\{X\}\\\}\.

\{1,2,3,4,5\}\\\{1,2,3,4,5\\\}\{4\}\\\{4\\\}\{1,2,3,5\}\\\{1,2,3,5\\\}\{1,2,3\}\\\{1,2,3\\\}\{1,2\}\\\{1,2\\\}\{1\}\\\{1\\\}\{2\}\\\{2\\\}\{3\}\\\{3\\\}\{5\}\\\{5\\\}

\(a\)TSL​\(d\)T\_\{\\mathrm\{SL\}\}\(d\)

\{1,2,3,4,5\}\\\{1,2,3,4,5\\\}\{1,2,3\}\\\{1,2,3\\\}\{1,2\}\\\{1,2\\\}\{1\}\\\{1\\\}\{2\}\\\{2\\\}\{3\}\\\{3\\\}\{4\}\\\{4\\\}\{5\}\\\{5\\\}

\(b\)Tloc​\(d\)T\_\{\\mathrm\{loc\}\}\(d\)

\{1,2,3,4,5\}\\\{1,2,3,4,5\\\}\{1,2\}\\\{1,2\\\}\{1\}\\\{1\\\}\{2\}\\\{2\\\}\{3\}\\\{3\\\}\{4\}\\\{4\\\}\{5\}\\\{5\\\}

\(c\)Tglob​\(d\)T\_\{\\mathrm\{glob\}\}\(d\)

\{1,2,3,4,5\}\\\{1,2,3,4,5\\\}\{1\}\\\{1\\\}\{2\}\\\{2\\\}\{3\}\\\{3\\\}\{4\}\\\{4\\\}\{5\}\\\{5\\\}

\(d\)Tglob𝛈​\(d\)T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}\(d\)

Figure 2:Output ofTSL,Tloc,Tglob,T\_\{\\mathrm\{SL\}\},T\_\{\\mathrm\{loc\}\},T\_\{\\mathrm\{glob\}\},andTglob𝛈T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}with𝛈=12​𝟏\\boldsymbol\{\\eta\}=\\frac\{1\}\{2\}\\boldsymbol\{1\}on the dissimilarityddgiven in Equation \([1](https://arxiv.org/html/2609.11173#S3.E1)\)\.

### 3\.3Bryant\-Berry Stable Clusters

[Bryant and Berry \(2001\)](https://arxiv.org/html/2609.11173#bib.bib20)introduced the notion of*stable clusters*, a combinatorial criterion that is strictly weaker than the separation conditions in Definition[10](https://arxiv.org/html/2609.11173#Thmtheorem10), yet still yields sets of clusters that form hierarchies\. In the following, we recall their construction, by translating the similarity framework in[Bryant and Berry \(2001\)](https://arxiv.org/html/2609.11173#bib.bib20)to the dissimilarity\-based framework used in this paper\. We then demonstrate that the induced hierarchical clustering method is admissible\.

Forx,y,z∈𝒳x,y,z\\in\\mathcal\{X\}, the*Bryant–Berry isolation weight*of the pair\(x,y\)\(x,y\)with respect tozzis

ρd​\(x​y∣z\):=min⁡\{d⁡\(x,z\),d⁡\(y,z\)\}−d⁡\(x,y\)\.\\rho\_\{d\}\(xy\\mid z\):=\\min\\\{d\(x,z\),d\(y,z\)\\\}\-d\(x,y\)\.Intuitively,ρd​\(x​y∣z\)\>0\\rho\_\{d\}\(xy\\mid z\)\>0means that the pointzzis farther from bothxxandyythan they are from each other\. For disjoint nonempty subsetsU,V,Z⊆𝒳U,V,Z\\subseteq\\mathcal\{X\}, we define the average Bryant–Berry isolation weight by

ρ¯d​\(U​V∣Z\):=1\|U​‖V‖​Z\|​∑u∈U,v∈V,z∈Zρd​\(u​v∣z\),\\overline\{\\rho\}\_\{d\}\(UV\\mid Z\):=\\frac\{1\}\{\|U\|\|V\|\|Z\|\}\\sum\_\{u\\in U,v\\in V,z\\in Z\}\\rho\_\{d\}\(uv\\mid z\),and the*stable\-cluster index*of a proper subsetC⊆𝒳C\\subseteq\\mathcal\{X\}with\|C\|≥2\|C\|\\geq 2by

ιd​\(C\):=minU,V≠∅,U∩V=∅U∪V=C∅≠Z⊆𝒳∖C⁡ρ¯d​\(U​V∣Z\)\.\\iota^\{d\}\(C\):=\\min\_\{\\begin\{subarray\}\{c\}U,V\\neq\\emptyset,\\ U\\cap V=\\emptyset\\\\ U\\cup V=C\\\\ \\emptyset\\neq Z\\subseteq\\mathcal\{X\}\\setminus C\\end\{subarray\}\}\\overline\{\\rho\}\_\{d\}\(UV\\mid Z\)\.A subsetCCis*stable*ifιd​\(C\)\>0\\iota^\{d\}\(C\)\>0, i\.e\., if for every bipartition\(U,V\)\(U,V\)ofCCand every nonempty setZZexternal toCC, the average isolation weight is strictly positive\. The Bryant\-Berry method is then666[Bryant and Berry \(2001\)](https://arxiv.org/html/2609.11173#bib.bib20)work with similaritiesss; by lettingd:=M−sd:=M\-swithM≥max⁡sM\\geq\\max s, the isolation weightρ\\rhois invariant, so the stable\-cluster family transfers verbatim to the dissimilarity setting\.

Tstable\(d\):=\{𝒳\}∪\{\{x\}:x∈𝒳\}∪\{C⊊𝒳:\|C\|≥2,ιd\(C\)\>0\}\.T\_\{\\mathrm\{stable\}\}\(d\):=\\\{\\mathcal\{X\}\\\}\\cup\\\{\\\{x\\\}:x\\in\\mathcal\{X\}\\\}\\cup\\\{C\\subsetneq\\mathcal\{X\}:\|C\|\\geq 2,\\,\\iota^\{d\}\(C\)\>0\\\}\.[Bryant and Berry \(2001\)](https://arxiv.org/html/2609.11173#bib.bib20)show that the set of stable clusters is laminar, soTstable​\(d\)T\_\{\\mathrm\{stable\}\}\(d\)is indeed a hierarchy\. Computingιd​\(C\)\\iota^\{d\}\(C\)requires minimizing over all bipartitions ofCCand all external witness sets, and[Bryant and Berry \(2001\)](https://arxiv.org/html/2609.11173#bib.bib20)show that deciding stability is NP\-hard in general\. They further establish that the stable\-cluster family is contained in the average\-linkage hierarchy, thus providing a tractable outer approximation of their method\.

We also consider variants of this method resulting from preprocessing the dissimilarity function with power transformations\. This yields the following parameterized family of admissible methods, which turns out to be useful for the analysis in the next section\. As shown in the next proposition, the Bryant–Berry method and its power\-transformation variants satisfy the admissibility axioms\.

###### Proposition 13\.

TstableT\_\{\\mathrm\{stable\}\}is admissible\. Furthermore, for everyp\>0p\>0, the composed methodTstable\(p\):=Tstable∘μpowerpT\_\{\\mathrm\{stable\}\}^\{\(p\)\}:=T\_\{\\mathrm\{stable\}\}\\circ\\mu\_\{\\mathrm\{power\}\}^\{p\}is admissible, whereμpowerp​\(d\)​\(x,y\)=\(d⁡\(x,y\)\)p\\mu\_\{\\mathrm\{power\}\}^\{p\}\(d\)\(x,y\)=\(d\(x,y\)\)^\{p\}for everyx,yx,y\.

## 4The Structure of the Set of Admissible Methods

This section studies the admissible class as a whole, by focusing on structural properties shared across all of its members rather than on any individual method\. We establish three main insights, each developed in its own subsection\. First, in Section[4\.1](https://arxiv.org/html/2609.11173#S4.SS1), we highlight the*diversity*of the admissible class: there exist uncountably many pairwise incompatible admissible methods, and both the width and the height of the class \(suitably defined\) are uncountable\. Second, we show in Section[4\.2](https://arxiv.org/html/2609.11173#S4.SS2)that despite their diversity, all admissible methods agree on a nontrivial*common core*of conservative cluster structures\. Finally, in Section[4\.3](https://arxiv.org/html/2609.11173#S4.SS3), we study*maximal*admissible methods and show that there are uncountably many of them\.

### 4\.1Comparing Admissible Methods: Refinement and Incompatibility

Section[3](https://arxiv.org/html/2609.11173#S3)introduced several admissible methods, including a number of parameterized variants\. This naturally raises the question of how restrictive the admissibility axioms are, and to what extent they constrain the class of admissible methods\.

#### 4\.1\.1Refinement: A Partial Order on Hierarchical Clustering Methods

Addressing the questions above requires comparing hierarchical clustering methods\. We begin by formalizing the intuition that some methods always return more fine\-grained cluster structures than others\.

###### Definition 14\.

LetT1T\_\{1\}andT2T\_\{2\}be two hierarchical clustering methods\. We sayT2T\_\{2\}*refines*T1T\_\{1\}\(or equivalently, thatT2T\_\{2\}is*finer*thanT1T\_\{1\}\), and denote it byT1⊑T2T\_\{1\}\\sqsubseteq T\_\{2\}, if

T1​\(d\)⊆T2​\(d\)for all​d∈𝒟⁡\(𝒳\)\.T\_\{1\}\(d\)\\subseteq T\_\{2\}\(d\)\\quad\\text\{for all \}d\\in\\mathcal\{D\}\(\\mathcal\{X\}\)\.

The binary relation⊑\\sqsubseteqis a partial order between hierarchical clustering methods\. We denote byℋ⁡\(𝒳\)\\mathcal\{H\}\(\\mathcal\{X\}\)andℋαadm​\(𝒳\)\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\)the sets of hierarchical clustering methods on𝒳\\mathcal\{X\}and hierarchical clustering methods on𝒳\\mathcal\{X\}that satisfyαadm\\alpha\_\{\\mathrm\{adm\}\}, respectively \(thusℋαadm​\(𝒳\)⊊ℋ⁡\(𝒳\)\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\)\\subsetneq\\mathcal\{H\}\(\\mathcal\{X\}\)\)\. Then,\(ℋ⁡\(𝒳\),⊑\)\(\\mathcal\{H\}\(\\mathcal\{X\}\),\\,\\sqsubseteq\)is a partially ordered set, and so is\(ℋαadm​\(𝒳\),⊑\)\(\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\),\\,\\sqsubseteq\)\.

###### Example 15\(Known and immediate refinement relations\)\.

Among the admissible methods introduced in Section[3](https://arxiv.org/html/2609.11173#S3), we have the following refinement relationships:

1. \(i\)Tloc⊑TSLT\_\{\\mathrm\{loc\}\}\\sqsubseteq T\_\{\\mathrm\{SL\}\}andTloc⊑TstableT\_\{\\mathrm\{loc\}\}\\sqsubseteq T\_\{\\mathrm\{stable\}\}
2. \(ii\)Tglob𝜼⊑Tloc𝜼T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}\\sqsubseteq T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}for any𝜼\\boldsymbol\{\\eta\};
3. \(iii\)Tglob𝜼′⊑Tglob𝜼T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}^\{\\prime\}\}\\sqsubseteq T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}andTloc𝜼′⊑Tloc𝜼T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}^\{\\prime\}\}\\sqsubseteq T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}for any𝜼′≤𝜼\\boldsymbol\{\\eta\}^\{\\prime\}\\leq\\boldsymbol\{\\eta\}\(where inequalities are element\-wise, i\.e\.,ηs′≤ηs\\eta^\{\\prime\}\_\{s\}\\leq\\eta\_\{s\}for allss\)\.

Point \(i\) is established by[Bryant and Berry \(2001\)](https://arxiv.org/html/2609.11173#bib.bib20)\. \(Tloc⊑TSLT\_\{\\mathrm\{loc\}\}\\sqsubseteq T\_\{\\mathrm\{SL\}\}has also been re\-proven by[Balcan et al\. \(2008\)](https://arxiv.org/html/2609.11173#bib.bib11)and[Dreveton et al\. \(2025\)](https://arxiv.org/html/2609.11173#bib.bib12)\.\) Point \(ii\) holds becauseτ⁡\(d,C\)≤ϱ⁡\(d,C\)\\tau\(d,C\)\\leq\\varrho\(d,C\)for everyC⊆𝒳C\\subseteq\\mathcal\{X\}andd∈𝒟⁡\(𝒳\)d\\in\\mathcal\{D\}\(\\mathcal\{X\}\)\. For point \(iii\), if𝛈′≤𝛈\\boldsymbol\{\\eta\}^\{\\prime\}\\leq\\boldsymbol\{\\eta\}, thenϱ⁡\(d,C\)<η\|C\|−1′\\varrho\(d,C\)<\\eta^\{\\prime\}\_\{\|C\|\-1\}impliesϱ⁡\(d,C\)<η\|C\|−1\\varrho\(d,C\)<\\eta\_\{\|C\|\-1\}, and likewise forτ⁡\(d,C\)\\tau\(d,C\)\. HenceTglob𝛈′⊑Tglob𝛈T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}^\{\\prime\}\}\\sqsubseteq T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}andTloc𝛈′⊑Tloc𝛈T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}^\{\\prime\}\}\\sqsubseteq T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}\.

#### 4\.1\.2Order\-theoretic Terminology

Before analyzingℋαadm​\(𝒳\)\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\)as a subset of the partially ordered set \(poset\)\(ℋ⁡\(𝒳\),⊑\)\(\\mathcal\{H\}\(\\mathcal\{X\}\),\\,\\sqsubseteq\), we first briefly recall the standard order\-theoretic terminology\.

In a poset\(P,⊑\)\(P,\\sqsubseteq\), and a subsetS⊆PS\\subseteq P, an elementg∈Sg\\in Sis a*greatest element*ofSSifs⊑gs\\sqsubseteq gfor alls∈Ss\\in S\. An elementm∈Sm\\in Sis a*maximal element*ofSSif there is nos∈Ss\\in Ssuch thatm⋤sm\\sqsubsetneq s\. A poset may contain multiple maximal elements but at most one greatest element; if a greatest element exists, it is necessarily the unique maximal element\. An elementa∈Pa\\in Pis a*lower bound*ofSSif for every elementb∈Sb\\in S,a⊑ba\\sqsubseteq b\. Similarly,a∈Pa\\in Pis an*upper bound*ofSSif for every elementb∈Sb\\in S,b⊑ab\\sqsubseteq a\. The*infimum*\(i\.e\., the greatest lower bound\) of the subsetSSis the lower bounda∈Pa\\in PofSSsuch thatb⊑ab\\sqsubseteq afor any lower boundbbofSSinPP\. The*supremum*\(i\.e\., the least upper bound or join\) of the subsetSSis the upper bounda∈Pa\\in PofSSsatisfyinga⊑ba\\sqsubseteq bfor any upper boundbbofSSinPP\.

Furthermore, two elementsx,y∈Px,y\\in Pare*comparable*if eitherx⊑yx\\sqsubseteq yory⊑xy\\sqsubseteq x; otherwise, they are incomparable\. A*chain*is a subsetC⊆PC\\subseteq Pin which every pair of distinct elements is comparable\. The*height*of the posetPPis the supremum of the cardinalities of all chains\.

A nonempty subsetS⊆PS\\subseteq Pis*upward directed*if, for everyx,y∈Sx,y\\in S, there existsz∈Sz\\in Ssuch thatx⊑zx\\sqsubseteq zandy⊑zy\\sqsubseteq z\. Equivalently, every finite nonempty subset ofSShas an upper bound that belongs toSS\. Every chain is upward directed, but an upward\-directed family need not be totally ordered\.

We say that two elementsx,y∈Px,y\\in Pare*upward compatible*if they share a common upper bound inPP; otherwise, they are said to be \(upward\) incompatible\. When analyzing the structure of subsets withinPP, these concepts of comparability and compatibility give rise to different types of antichains\. A standard \(or weak\)*antichain*is a subsetA⊆PA\\subseteq Pin which no two distinct elements are comparable\. The quantity used to capture the maximal size of these mutually incomparable sets is the*width*of the poset, defined as the supremum of the cardinalities of all standard antichains inPP\.

Besides incomparability, a stricter structural condition requires subsets to be mutually incompatible\. A*strong upward antichain*is a subsetA⊆PA\\subseteq Pwhere no two distinct elements are upward compatible \(for anyx≠y∈Ax\\neq y\\in A, there is noz∈Pz\\in Psuch thatx⊑zx\\sqsubseteq zandy⊑zy\\sqsubseteq z\)\. Analogous to how width measures the maximum size of standard antichains, the*cellularity*of a poset is defined as the supremum of the cardinalities of all its strong upward antichains\.

#### 4\.1\.3Diversity of Admissible Methods

###### Lemma 16\.

The family of admissible methods\{Tstable\(p\):p\>0\}⊊ℋαadm​\(𝒳\)\\left\\\{T\_\{\\mathrm\{stable\}\}^\{\(p\)\}:p\>0\\right\\\}\\subsetneq\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\)is pairwise incompatible and thus forms a strong upward antichain\.

Lemma[16](https://arxiv.org/html/2609.11173#Thmtheorem16)reveals thatℋαadm​\(𝒳\)\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\)is genuinely diverse, as it shows that no single admissible method refines all the others: not only do admissible methods differ, but uncountably many pairs admit no common refinement \(or upper bound in the order\-theoretic terminology\) even in the larger classℋ⁡\(𝒳\)\\mathcal\{H\}\(\\mathcal\{X\}\)\. In particular,\(ℋαadm​\(𝒳\),⊑\)\(\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\),\\sqsubseteq\)has no greatest element: there is no universal admissible method that simultaneously refines all others\. We obtain the following theorem\.

###### Theorem 17\.

The height, width, and cellularity ofℋαadm​\(𝒳\)\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\)are uncountable\.

###### Proof\.

Uncountable height follows immediately from the family of admissible methods\{Tloc𝜼:η∈\(0,1\],𝜼=η⋅𝟏\}\\left\\\{T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}:\\eta\\in\(0,1\],\\boldsymbol\{\\eta\}=\\eta\\cdot\\boldsymbol\{1\}\\right\\\}, which is itself uncountable\. Uncountable width and cellularity hold because the set\{Tstable\(p\):p\>0\}\\left\\\{T\_\{\\mathrm\{stable\}\}^\{\(p\)\}:p\>0\\right\\\}is an uncountable set of admissible methods whose elements are pairwise incompatible by Lemma[16](https://arxiv.org/html/2609.11173#Thmtheorem16), and therefore mutually incomparable\. Finally, we show thatTSLT\_\{\\mathrm\{SL\}\}andTstable\(p\)T\_\{\\mathrm\{stable\}\}^\{\(p\)\}with anyp\>0p\>0are incompatible in Appendix[D\.1](https://arxiv.org/html/2609.11173#A4.SS1)\. ∎

### 4\.2Uniformity: A Well\-separated Backbone

Section[4\.1\.3](https://arxiv.org/html/2609.11173#S4.SS1.SSS3)established the diversity of the set of admissible methodsℋαadm​\(𝒳\)\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\)\. This raises another natural question: do admissible methods nevertheless share some common cluster structure? As a motivating example, considerTSLT\_\{\\mathrm\{SL\}\}andTstableT\_\{\\mathrm\{stable\}\}\. Both refineTlocT\_\{\\mathrm\{loc\}\}: as noted in Section[4\.1](https://arxiv.org/html/2609.11173#S4.SS1),Tloc⊑TSLT\_\{\\mathrm\{loc\}\}\\sqsubseteq T\_\{\\mathrm\{SL\}\}andTloc⊑TstableT\_\{\\mathrm\{loc\}\}\\sqsubseteq T\_\{\\mathrm\{stable\}\}\. Hence, despite being incompatible,TSLT\_\{\\mathrm\{SL\}\}andTstableT\_\{\\mathrm\{stable\}\}shareTlocT\_\{\\mathrm\{loc\}\}as common cluster structure\. The following theorem shows that this is not a coincidence specific toTSLT\_\{\\mathrm\{SL\}\}andTstableT\_\{\\mathrm\{stable\}\}: every method inℋαadm​\(𝒳\)\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\)refines a globally separated hierarchyTglob𝜼T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}for some margin sequence𝜼\\boldsymbol\{\\eta\}\.

###### Theorem 18\(Backbone hierarchy\)\.

For anyT∈ℋαadm​\(𝒳\)T\\in\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\), there exists a separation margin sequence𝛈\\boldsymbol\{\\eta\}\(that may depend onTTand on\|𝒳\|\|\\mathcal\{X\}\|\) such thatTglob𝛈⊑TT\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}\\sqsubseteq T\.

Although admissible methods may disagree on individual clusters, Theorem[18](https://arxiv.org/html/2609.11173#Thmtheorem18)reveals a structural common ground: every admissible method refines a globally𝜼\\boldsymbol\{\\eta\}\-separated hierarchyTglob𝜼T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}\.

Whereas the separation margin sequence𝜼\\boldsymbol\{\\eta\}is method\-dependent, the following corollary shows that any finite collection of admissible methods admits a common backbone hierarchy, obtained by taking the pointwise minimum of their individual margin sequences\.

###### Corollary 19\.

LetT1,⋯,TmT\_\{1\},\\cdots,T\_\{m\}be methods inℋαadm​\(𝒳\)\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\), withmmfinite\. There exists a separation margin sequence𝛈\\boldsymbol\{\\eta\}such thatTglob𝛈⊑TℓT\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}\\sqsubseteq T\_\{\\ell\}for allℓ∈\[m\]\\ell\\in\[m\]\.

###### Proof\.

For eachℓ∈\[m\]\\ell\\in\[m\], Theorem[18](https://arxiv.org/html/2609.11173#Thmtheorem18)provides a separation margin sequence𝜼\(ℓ\)=\(ηs\(ℓ\)\)1≤s≤n−2\\boldsymbol\{\\eta\}^\{\(\\ell\)\}=\(\\eta\_\{s\}^\{\(\\ell\)\}\)\_\{1\\leq s\\leq n\-2\}such thatTglob𝜼\(ℓ\)⊑TℓT\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}^\{\(\\ell\)\}\}\\sqsubseteq T\_\{\\ell\}\. Define𝜼=\(ηs\)1≤s≤n−2\\boldsymbol\{\\eta\}=\(\\eta\_\{s\}\)\_\{1\\leq s\\leq n\-2\}byηs:=minℓ∈\[m\]⁡ηs\(ℓ\),\\eta\_\{s\}:=\\min\_\{\\ell\\in\[m\]\}\\eta\_\{s\}^\{\(\\ell\)\},which satisfiesηs∈\(0,1\]\\eta\_\{s\}\\in\(0,1\]for everysssince the minimum is taken over a finite set of positive values\. By construction,ηs≤ηs\(ℓ\)\\eta\_\{s\}\\leq\\eta\_\{s\}^\{\(\\ell\)\}for allℓ∈\[m\]\\ell\\in\[m\]and alls≥1s\\geq 1, soTglob𝜼⊑Tglob𝜼\(ℓ\)⊑TℓT\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}\\sqsubseteq T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}^\{\(\\ell\)\}\}\\sqsubseteq T\_\{\\ell\}\. ∎

The preceding corollary shows that every finite family of admissible methods has an admissible common lower bound\. This conclusion does not, however, extend to the entire admissible class\. To make this precise, define the*trivial method*Ttriv∈ℋ⁡\(𝒳\)T\_\{\\mathrm\{triv\}\}\\in\\mathcal\{H\}\(\\mathcal\{X\}\)by

Ttriv​\(d\):=\{𝒳\}∪\{\{x\}:x∈𝒳\}∀d∈𝒟⁡\(𝒳\)\.\\displaystyle T\_\{\\mathrm\{triv\}\}\(d\):=\\\{\\mathcal\{X\}\\\}\\cup\\bigl\\\{\\\{x\\\}:x\\in\\mathcal\{X\}\\bigr\\\}\\qquad\\forall d\\in\\mathcal\{D\}\(\\mathcal\{X\}\)\.\(2\)In other words,TtrivT\_\{\\mathrm\{triv\}\}is the method that always returns only the root and the singleton leaves\. The following proposition shows thatTtrivT\_\{\\mathrm\{triv\}\}is the infimum ofℋαadm​\(𝒳\)\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\)\. Moreover,TtrivT\_\{\\mathrm\{triv\}\}does not satisfy partition richness and hence is not admissible\. Hence Proposition[20](https://arxiv.org/html/2609.11173#Thmtheorem20)also implies that\(ℋαadm​\(𝒳\),⊑\)\(\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\),\\sqsubseteq\)has no least element\.

###### Proposition 20\.

The infimum ofℋαadm​\(𝒳\)\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\), computed in the poset\(ℋ⁡\(𝒳\),⊑\)\(\\mathcal\{H\}\(\\mathcal\{X\}\),\\sqsubseteq\), isTtrivT\_\{\\mathrm\{triv\}\}\.

###### Proof\.

As every hierarchy on𝒳\\mathcal\{X\}contains the root and all singleton clusters,Ttriv⊑TT\_\{\\mathrm\{triv\}\}\\sqsubseteq Tfor everyT∈ℋ⁡\(𝒳\)T\\in\\mathcal\{H\}\(\\mathcal\{X\}\)\. In particular,TtrivT\_\{\\mathrm\{triv\}\}is a lower bound ofℋαadm​\(𝒳\)\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\)\.

We now show that no nontrivial proper cluster belongs to every admissible method\. Fixd∈𝒟⁡\(𝒳\)d\\in\\mathcal\{D\}\(\\mathcal\{X\}\)and a nontrivial proper subsetC⊊𝒳C\\subsetneq\\mathcal\{X\}\. Sinceϱ⁡\(d,C\)\>0\\varrho\(d,C\)\>0, we may choose a separation margin sequence𝜼\\boldsymbol\{\\eta\}such that0<η\|C\|−1≤ϱ⁡\(d,C\)\.0<\\eta\_\{\|C\|\-1\}\\leq\\varrho\(d,C\)\.Then, by definition,C∉Tglob𝜼​\(d\)C\\notin T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}\(d\), whereasTglob𝜼T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}is admissible\. HenceCCdoes not belong to the intersection of all admissible methods\. Because this holds for everyd∈𝒟⁡\(𝒳\)d\\in\\mathcal\{D\}\(\\mathcal\{X\}\)and every nontrivial proper subsetC⊊𝒳C\\subsetneq\\mathcal\{X\}, and because the root and singleton clusters belong to every hierarchy, for everyd∈𝒟⁡\(𝒳\)d\\in\\mathcal\{D\}\(\\mathcal\{X\}\)⋂T∈ℋαadm​\(𝒳\)T⁡\(d\)=Ttriv​\(d\)\.\\bigcap\_\{T\\in\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\)\}T\(d\)=T\_\{\\mathrm\{triv\}\}\(d\)\.ThereforeTtrivT\_\{\\mathrm\{triv\}\}is the infimum ofℋαadm​\(𝒳\)\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\)inℋ⁡\(𝒳\)\\mathcal\{H\}\(\\mathcal\{X\}\)\. ∎

Theorem[18](https://arxiv.org/html/2609.11173#Thmtheorem18)carries an additional, perhaps unexpected, implication: the partition richness axiom can be upgraded to a seemingly stronger one,*hierarchical richness*\. Recall that partition richness only requires that for every partition𝒞\\mathcal\{C\}, there exists a dissimilarityddsuch that𝒞⊆T⁡\(d\)\\mathcal\{C\}\\subseteq T\(d\)\. Combined with the other axioms ofαadm\\alpha\_\{\\mathrm\{adm\}\}, it in fact implies thatTTsatisfies hierarchical richness, that is, for anyΨ∈𝒯⁡\(𝒳\)\\Psi\\in\\mathcal\{T\}\(\\mathcal\{X\}\), there existsd∈𝒟⁡\(𝒳\)d\\in\\mathcal\{D\}\(\\mathcal\{X\}\)such thatΨ⊆T⁡\(d\)\\Psi\\subseteq T\(d\)\.777This is a direct consequence ofTTrefiningTglob𝜼T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}for some𝜼\\boldsymbol\{\\eta\}, and ofTglob𝜼:𝒟⁡\(𝒳\)→𝒯⁡\(𝒳\)T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}\\colon\\mathcal\{D\}\(\\mathcal\{X\}\)\\to\\mathcal\{T\}\(\\mathcal\{X\}\)being surjective\.

### 4\.3Existence of Uncountably Many Maximal Admissible Methods

#### 4\.3\.1Intersection and Union of Methods

Corollary[19](https://arxiv.org/html/2609.11173#Thmtheorem19)shows that every finite collection of admissible methods has a lower bound inℋαadm​\(𝒳\)\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\), whereas by Proposition[20](https://arxiv.org/html/2609.11173#Thmtheorem20)an infinite collection of admissible methods may not admit a lower bound inℋαadm​\(𝒳\)\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\)\. There are, however, particular cases for which the infimum and the supremum, of an arbitrary, possibly uncountable, collection of admissible methods remain inℋαadm​\(𝒳\)\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\)\. We introduce intersections and unions of hierarchical clustering methods and determine when these operations remain hierarchical clustering methods and preserve admissibility\.

###### Definition 21\.

For two hierarchical clustering methodsT1,T2∈ℋ⁡\(𝒳\)T\_\{1\},T\_\{2\}\\in\\mathcal\{H\}\(\\mathcal\{X\}\), define their intersectionT1⊓T2T\_\{1\}\\sqcap T\_\{2\}and unionT1⊔T2T\_\{1\}\\sqcup T\_\{2\}by

\(T1⊓T2\)​\(d\):=T1​\(d\)∩T2​\(d\)and\(T1⊔T2\)​\(d\):=T1​\(d\)∪T2​\(d\)∀d∈𝒟⁡\(𝒳\)\.\(T\_\{1\}\\sqcap T\_\{2\}\)\(d\):=T\_\{1\}\(d\)\\cap T\_\{2\}\(d\)\\quad\\text\{ and \}\\quad\(T\_\{1\}\\sqcup T\_\{2\}\)\(d\):=T\_\{1\}\(d\)\\cup T\_\{2\}\(d\)\\qquad\\forall d\\in\\mathcal\{D\}\(\\mathcal\{X\}\)\.More generally, for any nonempty subsetL⊆ℋ⁡\(𝒳\)L\\subseteq\\mathcal\{H\}\(\\mathcal\{X\}\)of hierarchical clustering methods, defineT⊓LT\_\{\\sqcap L\}andT⊔LT\_\{\\sqcup L\}by

T⊓L​\(d\):=⋂Ti∈LTi​\(d\)andT⊔L​\(d\):=⋃Ti∈LTi​\(d\)∀d∈𝒟⁡\(𝒳\)\.T\_\{\\sqcap L\}\(d\)\\;:=\\;\\bigcap\_\{T\_\{i\}\\in L\}T\_\{i\}\(d\)\\quad\\text\{and\}\\quad T\_\{\\sqcup L\}\(d\)\\;:=\\;\\bigcup\_\{T\_\{i\}\\in L\}T\_\{i\}\(d\)\\qquad\\forall d\\in\\mathcal\{D\}\(\\mathcal\{X\}\)\.

For any dissimilarity functionddand any nonemptyL⊆ℋ⁡\(𝒳\)L\\subseteq\\mathcal\{H\}\(\\mathcal\{X\}\),T⊓L​\(d\)T\_\{\\sqcap L\}\(d\)is always a hierarchy, as the intersection of laminar families remains laminar and still contains𝒳\\mathcal\{X\}and all singletons\. In contrast,T⊔L​\(d\)T\_\{\\sqcup L\}\(d\)need not be laminar, and therefore need not define a hierarchy in general\. However, closure under union holds under an additional compatibility assumption, as captured by the following lemma\.

###### Lemma 22\.

LetL⊆ℋ⁡\(𝒳\)L\\subseteq\\mathcal\{H\}\(\\mathcal\{X\}\)be a nonempty set of hierarchical clustering methods\.

1. \(i\)T⊓L∈ℋ⁡\(𝒳\)T\_\{\\sqcap L\}\\in\\mathcal\{H\}\(\\mathcal\{X\}\)\.
2. \(ii\)T⊔L∈ℋ⁡\(𝒳\)T\_\{\\sqcup L\}\\in\\mathcal\{H\}\(\\mathcal\{X\}\)if and only if the methods inLLare pairwise compatible, i\.e\., for allTi,Tj∈LT\_\{i\},T\_\{j\}\\in L,Ti⊔Tj∈ℋ⁡\(𝒳\)T\_\{i\}\\sqcup T\_\{j\}\\in\\mathcal\{H\}\(\\mathcal\{X\}\)\.

###### Proof\.

\(i\) For anyd∈𝒟⁡\(𝒳\)d\\in\\mathcal\{D\}\(\\mathcal\{X\}\), eachT⁡\(d\)T\(d\)is a laminar family containing𝒳\\mathcal\{X\}and all singletons\. The intersectionT⊓L​\(d\)=⋂T∈LT⁡\(d\)T\_\{\\sqcap L\}\(d\)=\\bigcap\_\{T\\in L\}T\(d\)still contains𝒳\\mathcal\{X\}and all singletons, and remains laminar since laminarity is a pairwise condition preserved under intersection\.

\(ii\) If the methods inLLare pairwise compatible, then for anyd∈𝒟⁡\(𝒳\)d\\in\\mathcal\{D\}\(\\mathcal\{X\}\)and anyA,B∈T⊔L​\(d\)A,B\\in T\_\{\\sqcup L\}\(d\), there existTi,Tj∈LT\_\{i\},T\_\{j\}\\in LwithA∈Ti​\(d\)A\\in T\_\{i\}\(d\)andB∈Tj​\(d\)B\\in T\_\{j\}\(d\)\. Pairwise compatibility implies thatA,BA,Bare laminar\. Together with𝒳\\mathcal\{X\}and all singletons being inT⊔L​\(d\)T\_\{\\sqcup L\}\(d\), this provesT⊔L∈ℋT\_\{\\sqcup L\}\\in\\mathcal\{H\}\. Conversely, ifT⊔L∈ℋT\_\{\\sqcup L\}\\in\\mathcal\{H\}, then for anyTi,Tj∈LT\_\{i\},T\_\{j\}\\in L,Ti⊔Tj⊆T⊔LT\_\{i\}\\sqcup T\_\{j\}\\subseteq T\_\{\\sqcup L\}inherits laminarity, soTi⊔Tj∈ℋ⁡\(𝒳\)T\_\{i\}\\sqcup T\_\{j\}\\in\\mathcal\{H\}\(\\mathcal\{X\}\)\. ∎

Lemma[22](https://arxiv.org/html/2609.11173#Thmtheorem22)settles the structural question of whenT⊓LT\_\{\\sqcap L\}andT⊔LT\_\{\\sqcup L\}define hierarchies\. We now turn to the axiomatic question: under what conditions are these operations closed within the admissible classℋαadm​\(𝒳\)\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\)?

###### Lemma 23\.

LetL⊆ℋαadm​\(𝒳\)L\\subseteq\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\)be nonempty\.

1. \(i\)IfLLis finite, thenT⊓L∈ℋαadm​\(𝒳\)T\_\{\\sqcap L\}\\in\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\)\.
2. \(ii\)IfLLis upward directed under⊑\\sqsubseteq, thenT⊔L∈ℋαadm​\(𝒳\)T\_\{\\sqcup L\}\\in\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\)\.

Observe that finiteness ofLLis required to ensureT⊓L∈ℋαadm​\(𝒳\)T\_\{\\sqcap L\}\\in\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\)\. As a simple counterexample, considerL=\{Tglobα​𝟏:α∈\(0,1\]\}L=\\\{T\_\{\\mathrm\{glob\}\}^\{\\alpha\\boldsymbol\{1\}\}\\colon\\alpha\\in\(0,1\]\\\}\. ThenLLis an infinite subset ofℋαadm​\(𝒳\)\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\), andT⊓L=TtrivT\_\{\\sqcap L\}=T\_\{\\mathrm\{triv\}\}is the trivial method defined in Equation \([2](https://arxiv.org/html/2609.11173#S4.E2)\), which is not admissible\.

A partially ordered set admitting an infimum \(resp\. supremum\) for every finite nonempty subset is called a*meet\-semilattice*\(resp\.*join\-semilattice*\)\. Lemma[23](https://arxiv.org/html/2609.11173#Thmtheorem23)thus implies that\(ℋαadm​\(𝒳\),⊑\)\(\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\),\\sqsubseteq\)is a meet\-semilattice but not a join\-semilattice\.

#### 4\.3\.2Existence of Maximal Admissible Methods

A consequence of Lemma[16](https://arxiv.org/html/2609.11173#Thmtheorem16)is that\(ℋαadm​\(𝒳\),⊑\)\(\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\),\\sqsubseteq\)has no greatest element and does not even admit a supremum inℋ⁡\(𝒳\)\\mathcal\{H\}\(\\mathcal\{X\}\)\. Nevertheless, closure under unions of chains yields, via Zorn’s lemma, the existence of*maximal*elements\.

###### Theorem 24\.

The setℋαadm​\(𝒳\)\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\)has maximal elements, and everyT∈ℋαadm​\(𝒳\)T\\in\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\)is refined by at least one maximal elementTM∈ℋαadm​\(𝒳\)T^\{M\}\\in\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\)\. Moreover,ℋαadm​\(𝒳\)\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\)contains an uncountable family of pairwise\-incompatible maximal elements\.

###### Proof\.

We apply Zorn’s lemma \(see, e\.g\.,[Shen and Vereshchagin \(2002, Theorem 30\)](https://arxiv.org/html/2609.11173#bib.bib21)\)\. LetL⊆ℋαadm​\(𝒳\)L\\subseteq\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\)be any chain \(i\.e\., a subset totally ordered by⊑\\sqsubseteq\)\. We show thatT⊔L∈ℋαadm​\(𝒳\)T\_\{\\sqcup L\}\\in\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\), providing an upper bound forLLinℋαadm​\(𝒳\)\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\)\.

Because every chain is upward directed, Lemma[23](https://arxiv.org/html/2609.11173#Thmtheorem23)\(ii\) yields thatT⊔L∈ℋαadm​\(𝒳\)T\_\{\\sqcup L\}\\in\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\)\. By Zorn’s lemma,ℋαadm​\(𝒳\)\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\)has maximal elements, and everyT∈ℋαadm​\(𝒳\)T\\in\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\)is refined by some maximal element\.

For cardinality and structural incompatibility, consider the uncountable family\{Tstable\(p\):p\>0\}\\\{T\_\{\\mathrm\{stable\}\}^\{\(p\)\}:p\>0\\\}from Lemma[16](https://arxiv.org/html/2609.11173#Thmtheorem16)\. Extend eachTstable\(p\)T\_\{\\mathrm\{stable\}\}^\{\(p\)\}to a maximal methodMpM\_\{p\}\. Ifp≠qp\\neq q, then there is no hierarchy\-valued method that can refine bothMpM\_\{p\}andMqM\_\{q\}, because such a method would also refine bothTstable\(p\)T\_\{\\mathrm\{stable\}\}^\{\(p\)\}andTstable\(q\)T\_\{\\mathrm\{stable\}\}^\{\(q\)\}\. Thus\{Mp:p\>0\}\\\{M\_\{p\}:p\>0\\\}is an uncountable pairwise incompatible family of maximal elements\. ∎

Incompatibility arises when two methods select clusters that cannot coexist in a single hierarchy for some input\. Conversely, any pairwise\-compatible family of methods can be combined through their union, which remains hierarchy\-valued\.

Althoughℋαadm​\(𝒳\)\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\)has no greatest element, Theorem[24](https://arxiv.org/html/2609.11173#Thmtheorem24)guarantees that every admissible method is refined by a maximal admissible method that cannot be further refined withinℋαadm​\(𝒳\)\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\)\.

Figure[3](https://arxiv.org/html/2609.11173#S4.F3)summarizes the results of this section on the structure ofℋαadm​\(𝒳\)\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\)\.

ℋ⁡\(𝒳\)\\mathcal\{H\}\(\\mathcal\{X\}\)ℋαadm​\(𝒳\)\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\)TSLMT\_\{\\mathrm\{SL\}\}^\{M\}TstableMT\_\{\\mathrm\{stable\}\}^\{M\}Tstable\(2\),MT\_\{\\mathrm\{stable\}\}^\{\(2\),M\}Tstable\(3\),MT\_\{\\mathrm\{stable\}\}^\{\(3\),M\}…\\dotsTwlocMT\_\{\\mathrm\{wloc\}\}^\{M\}TSLT\_\{\\mathrm\{SL\}\}TstableT\_\{\\mathrm\{stable\}\}Tstable\(2\)T\_\{\\mathrm\{stable\}\}^\{\(2\)\}Tstable\(3\)T\_\{\\mathrm\{stable\}\}^\{\(3\)\}…\\dotsTwlocT\_\{\\mathrm\{wloc\}\}TlocT\_\{\\mathrm\{loc\}\}Tloc𝜼\(1\)T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}^\{\(1\)\}\}Tloc𝜼\(2\)T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}^\{\(2\)\}\}⋮\\vdotsTglobT\_\{\\mathrm\{glob\}\}Tglob𝜼\(1\)T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}^\{\(1\)\}\}Tglob𝜼\(2\)T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}^\{\(2\)\}\}⋮\\vdots

Figure 3:Hasse diagrams of a subset ofℋαadm​\(𝒳\)\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\)under the refinement order⊑\\sqsubseteq, where transitive edges are omitted and an arrow fromT′T^\{\\prime\}toTTindicates thatT′⊑TT^\{\\prime\}\\sqsubseteq T\. The blue solid and red dashed boundaries delimitℋαadm​\(𝒳\)\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\)andℋ⁡\(𝒳\)\\mathcal\{H\}\(\\mathcal\{X\}\), respectively\. The superscriptMMdenotes the maximal admissible element of each chain\. In particular, the diagram includes the strong upward antichain\{Tstable\(p\)\}p∈\[n−2\]\\\{T\_\{\\mathrm\{stable\}\}^\{\(p\)\}\\\}\_\{p\\in\[n\-2\]\}, and chains\{Tglob𝜼\(i\)\}i∈\[n−2\]\\\{T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}^\{\(i\)\}\}\\\}\_\{i\\in\[n\-2\]\}, and\{Tloc𝜼\(i\)\}i∈\[n−2\]\\\{T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}^\{\(i\)\}\}\\\}\_\{i\\in\[n\-2\]\}with𝜼\(i\)=ηi​𝟏\\boldsymbol\{\\eta\}^\{\(i\)\}=\\eta^\{i\}\\mathbf\{1\}, for someη∈\(0,1\)\\eta\\in\(0,1\)\.

## 5Extensions of the Canonical Framework

The previous sections study hierarchical clustering under the canonical axiom system\. We now examine the robustness of this framework in two directions that are specific to hierarchical clustering and its practical use\. First, we add exactness on ultrametric inputs, a fidelity requirement with no direct analog in flat clustering\. Second, we allow the input dissimilarity to be preprocessed before the hierarchy is constructed and ask which axioms are preserved under composition\.

### 5\.1Exactness on Ultrametric Inputs

The admissibility axioms closely follow Kleinberg’s original setting\. However, in the particular case where the input distance is an ultrametric, hierarchical clustering induces a known correspondence with dendrograms, which motivates an additional requirement tailored to hierarchical outputs\. Recall that a distanceddis an*ultrametric*ifddsatisfies the strong triangle inequality

d⁡\(x,y\)≤max⁡\{d⁡\(x,z\),d⁡\(y,z\)\}∀x,y,z∈𝒳,d\(x,y\)\\ \\leq\\ \\max\\\{d\(x,z\),d\(y,z\)\\\}\\quad\\forall x,y,z\\in\\mathcal\{X\},and that a*dendrogram*\(Ψ,θ\)\(\\Psi,\\theta\)is a hierarchyΨ∈𝒯\\Psi\\in\\mathcal\{T\}in which each node is assigned height through the dendrogram’s height functionθ:Ψ→R\+\\theta\\colon\\Psi\\to\\mathbb\{R\}\_\{\+\}, with leaves \(singletons\) having height00\. The functionθ\\thetasatisfiesθ⁡\(\{x\}\)=0\\theta\(\\\{x\\\}\)=0for everyx∈𝒳x\\in\\mathcal\{X\}and is strictly increasing from the leaves towards the root, that is,θ⁡\(C\)<θ⁡\(C′\)\\theta\(C\)<\\theta\(C^\{\\prime\}\)for everyC⊊C′C\\subsetneq C^\{\\prime\}\.

Ultrametric distances and dendrograms are in one\-to\-one correspondence \(seee\.g\.,[Carlsson and Mémoli \(2010\)](https://arxiv.org/html/2609.11173#bib.bib4)\): for every ultrametricuu, there exists a unique dendrogram\(Ψu,θu\)\(\\Psi\_\{u\},\\theta\_\{u\}\)such thatu⁡\(x,y\)=θu​\(lcaΨu​\(x,y\)\)u\(x,y\)=\\theta\_\{u\}\(\\mathrm\{lca\}\_\{\\Psi\_\{u\}\}\(x,y\)\)for allx,y∈𝒳x,y\\in\\mathcal\{X\}, wherelcaΨu​\(x,y\)\\mathrm\{lca\}\_\{\\Psi\_\{u\}\}\(x,y\)denotes the smallest cluster inΨu\\Psi\_\{u\}containing bothxxandyy\. Moreover, the hierarchyΨu\\Psi\_\{u\}can be constructed explicitly from the ultrametricuuby

Ψu=\{Bu\(x,r\):x∈𝒳,r≥0\}whereBu\(x,r\):=\{y∈𝒳:u\(x,y\)≤r\},\\displaystyle\\Psi\_\{u\}=\\\{B\_\{u\}\(x,r\):x\\in\\mathcal\{X\},\\ r\\geq 0\\\}\\quad\\text\{ where \}B\_\{u\}\(x,r\):=\\\{y\\in\\mathcal\{X\}:u\(x,y\)\\leq r\\\},\(3\)and the dendrogram’s height functionθu\\theta\_\{u\}is given byθu​\(C\)=maxx,y∈C⁡u⁡\(x,y\)\.\\theta\_\{u\}\(C\)=\\max\_\{x,y\\in C\}u\(x,y\)\.

This correspondence motivates an additional fidelity requirement: when the input is already an ultrametric, the method should recover its canonical hierarchy \(albeit we do not requires to output the associated height\)\.

###### Definition 25\.

A hierarchical clustering methodT∈ℋ⁡\(𝒳\)T\\in\\mathcal\{H\}\(\\mathcal\{X\}\)is*exact on ultrametrics*if, for every ultrametricuu, we haveT⁡\(u\)=ΨuT\(u\)=\\Psi\_\{u\}\.

In particular, when restricted to ultrametric inputs, all methods that are exact on ultrametrics coincide\.

We say that a hierarchical clustering method is*strongly admissible*if it is admissible and exact on ultrametrics\. We denote this strengthened axiom system byαadm\+\\alpha\_\{\\mathrm\{adm\+\}\}, and writeℋαadm\+​\(𝒳\)\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\+\}\}\}\(\\mathcal\{X\}\)for the class of strongly admissible hierarchical clustering methods on𝒳\\mathcal\{X\}\.

We prove in the appendix that the non\-binary variant of single linkage is exact on ultrametrics\. Moreover,Tglob𝜼T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}andTloc𝜼T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}are exact on ultrametrics if and only if𝜼=𝟏\\boldsymbol\{\\eta\}=\\boldsymbol\{1\}\. Finally, we also prove that for everyp\>0p\>0, the composed methodTstable\(p\):=Tstable∘μpowerpT\_\{\\mathrm\{stable\}\}^\{\(p\)\}:=T\_\{\\mathrm\{stable\}\}\\circ\\mu\_\{\\mathrm\{power\}\}^\{p\}is exact on ultrametrics\. Together, these results show that there exist uncountably many strongly admissible methods\. Moreover, the results of Section[4](https://arxiv.org/html/2609.11173#S4)regarding the existence of uncountably many maximal elements and of a backbone hierarchy hold when we replaceαadm\\alpha\_\{\\mathrm\{adm\}\}byαadm\+\\alpha\_\{\\mathrm\{adm\+\}\}\. There are, however, some important differences\. Indeed, for every strongly admissible methodTT, we haveTglob⊑TT\_\{\\mathrm\{glob\}\}\\sqsubseteq T\. Thus, whereas\(ℋαadm​\(𝒳\),⊑\)\(\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\}\(\\mathcal\{X\}\),\\,\\sqsubseteq\)has no least element \(Proposition[20](https://arxiv.org/html/2609.11173#Thmtheorem20)\),\(ℋαadm\+​\(𝒳\),⊑\)\(\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\+\}\}\}\(\\mathcal\{X\}\),\\,\\sqsubseteq\)has one least element,TglobT\_\{\\mathrm\{glob\}\}, which is thus the coarsest strongly admissible hierarchical method\. In practice,TglobT\_\{\\mathrm\{glob\}\}returns a hierarchy composed of many non\-singleton clusters\. To illustrate it, we provide in Appendix[E](https://arxiv.org/html/2609.11173#A5)numerical experiments to evaluate the size of the hierarchy returned byTglobT\_\{\\mathrm\{glob\}\}on real data sets\.

###### Theorem 26\.

The height, width, and cellularity ofℋαadm\+​\(𝒳\)\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\+\}\}\}\(\\mathcal\{X\}\)are uncountable\. In particular,TSLT\_\{\\mathrm\{SL\}\},TglobT\_\{\\mathrm\{glob\}\},TlocT\_\{\\mathrm\{loc\}\}, andTstable\(p\)T\_\{\\mathrm\{stable\}\}^\{\(p\)\}for everyp\>0p\>0are strongly admissible\. Moreover,

1. 1\.The backbone and maximality results of Section 4 remain valid;
2. 2\.ℋαadm\+​\(𝒳\)\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\+\}\}\}\(\\mathcal\{X\}\)admits uncountably many pairwise incompatible maximal elements;
3. 3\.TglobT\_\{\\mathrm\{glob\}\}is the least strongly admissible method under refinement\.

### 5\.2Axiom\-Preserving Preprocessing

Modern clustering pipelines rarely apply a hierarchical clustering method directly to the raw observations or their original pairwise dissimilarities\. Instead, the data is typically preprocessed before the hierarchy is constructed\. Common examples include dimensionality reduction, feature re\-weighting, and the choice or transformation of a metric \(for example using a kernel\)\. Density\-aware procedures also fit this framework: they first construct a nearest\-neighbor\-based dissimilarity and then apply a single\-linkage\-type method to the transformed data\. These examples explain treating preprocessing as a part of the clustering pipeline and asking which axioms are preserved by it\.

#### 5\.2\.1General Preservation Principles

We model preprocessing as a mapμ:𝒟⁡\(𝒳\)→𝒟⁡\(𝒳\)\\mu\\colon\\mathcal\{D\}\(\\mathcal\{X\}\)\\to\\mathcal\{D\}\(\\mathcal\{X\}\), called a*transformation*\. Given a hierarchical clustering methodTT, the composed methodT∘μ:d↦T⁡\(μ⁡\(d\)\)T\\circ\\mu\\colon d\\mapsto T\(\\mu\(d\)\)first transforms the input dissimilarity and then appliesTT\. Becauseμ⁡\(d\)\\mu\(d\)is a dissimilarity on the same domain asdd, the compositionT∘μT\\circ\\muis again a hierarchical clustering method\.

Our goal is to study when the axiomatic properties ofTTare inherited by the composed methodT∘μT\\circ\\mu\. Rather than studying this question separately for each method, we lift each axiom to transformations: a transformation preserves an axiom if composition with it preserves that axiom for every method satisfying it\.

###### Definition 27\.

A transformationμ:𝒟⁡\(𝒳\)→𝒟⁡\(𝒳\)\\mu\\colon\\mathcal\{D\}\(\\mathcal\{X\}\)\\to\\mathcal\{D\}\(\\mathcal\{X\}\)*preserves*an axiomα\\alphaif, for every methodTTthat satisfiesα\\alpha, the composed methodT∘μT\\circ\\mualso satisfiesα\\alpha\.

###### Lemma 28\.

Letμ:𝒟⁡\(𝒳\)→𝒟⁡\(𝒳\)\\mu\\colon\\mathcal\{D\}\(\\mathcal\{X\}\)\\to\\mathcal\{D\}\(\\mathcal\{X\}\)\. Each of the following conditions is sufficient forμ\\muto preserve the corresponding axiom:

- •*Scale invariance*: for everyd∈𝒟⁡\(𝒳\)d\\in\\mathcal\{D\}\(\\mathcal\{X\}\)and everyβ\>0\\beta\>0, there existsβ′\>0\\beta^\{\\prime\}\>0such thatμ⁡\(β​d\)=β′​μ​\(d\)\\mu\(\\beta d\)=\\beta^\{\\prime\}\\mu\(d\);
- •*Partition richness:*μ\\muis surjective;
- •*Partition consistency*: for every partition𝒞∈𝒫⁡\(𝒳\)\\mathcal\{C\}\\in\\mathcal\{P\}\(\\mathcal\{X\}\)and everyd,d′∈𝒟⁡\(𝒳\)d,d^\{\\prime\}\\in\\mathcal\{D\}\(\\mathcal\{X\}\), ifd′d^\{\\prime\}is a𝒞\\mathcal\{C\}\-strengthening ofdd, thenμ⁡\(d′\)\\mu\(d^\{\\prime\}\)is a𝒞\\mathcal\{C\}\-strengthening ofμ⁡\(d\)\\mu\(d\);
- •*Permutation invariance:*for everyd∈𝒟⁡\(𝒳\)d\\in\\mathcal\{D\}\(\\mathcal\{X\}\)and every permutationϕ\\phi,μ⁡\(dϕ\)=μ​\(d\)ϕ\\mu\(d\_\{\\phi\}\)=\\mu\(d\)\_\{\\phi\};
- •*Exactness on ultrametrics:*for every ultrametricuu,μ⁡\(u\)\\mu\(u\)is an ultrametric andΨμ⁡\(u\)=Ψu\\Psi\_\{\\mu\(u\)\}=\\Psi\_\{u\}\.

Except for partition richness, each condition in Lemma[28](https://arxiv.org/html/2609.11173#Thmtheorem28)ensures that the input structure relevant to the corresponding axiom is preserved by the transformation: scaling remains scaling, strengthenings remain strengthenings, and permutations commute with the transformation\. Similarly, ultrametrics are mapped to ultrametrics that induce the same hierarchy\.

These observations are instances of a more general principle: the axioms above can be expressed by imposing a condition on the outputs associated with one or two dissimilarities satisfying a prescribed relation\. In Appendix[B](https://arxiv.org/html/2609.11173#A2), we formalize this class of relational properties and recover the corresponding parts of Lemma[28](https://arxiv.org/html/2609.11173#Thmtheorem28)as corollaries\.

###### Corollary 29\.

Letg:R≥0→R≥0g\\colon\\mathbb\{R\}\_\{\\geq 0\}\\to\\mathbb\{R\}\_\{\\geq 0\}satisfyg⁡\(0\)=0g\(0\)=0andg⁡\(t\)\>0g\(t\)\>0for everyt\>0t\>0\. Define the pointwise transformationμg:𝒟⁡\(𝒳\)→𝒟⁡\(𝒳\)\\mu\_\{g\}\\colon\\mathcal\{D\}\(\\mathcal\{X\}\)\\to\\mathcal\{D\}\(\\mathcal\{X\}\)by

μg​\(d\)​\(x,y\):=g⁡\(d⁡\(x,y\)\)for all​d∈𝒟⁡\(𝒳\)​and​x,y∈𝒳\.\\mu\_\{g\}\(d\)\(x,y\):=g\\bigl\(d\(x,y\)\\bigr\)\\qquad\\text\{for all \}d\\in\\mathcal\{D\}\(\\mathcal\{X\}\)\\text\{ and \}x,y\\in\\mathcal\{X\}\.Thenμg\\mu\_\{g\}preserves permutation invariance\. Moreover,

- •ifggis strictly increasing, thenμg\\mu\_\{g\}preserves partition consistency and exactness on ultrametrics;
- •ifggis surjective ontoR≥0\\mathbb\{R\}\_\{\\geq 0\}, thenμg\\mu\_\{g\}preserves partition richness;
- •if, for everyβ\>0\\beta\>0, there existscβ\>0c\_\{\\beta\}\>0such thatg⁡\(β​t\)=cβ​g​\(t\)g\(\\beta t\)=c\_\{\\beta\}g\(t\)for everyt≥0t\\geq 0, thenμg\\mu\_\{g\}preserves scale invariance\.

In particular, for everyp\>0p\>0, the power transformationμpowerp​\(d\)​\(x,y\):=d​\(x,y\)p\\mu\_\{\\mathrm\{power\}\}^\{p\}\(d\)\(x,y\):=d\(x,y\)^\{p\}preserves scale invariance, partition richness, partition consistency, permutation invariance, and exactness on ultrametrics\. Thus,μpowerp\\mu\_\{\\mathrm\{power\}\}^\{p\}preserves both admissibility and strong admissibility\.

#### 5\.2\.2Examples and Applications

*Minimax\-path preprocessing and single linkage\.*Our first example gives an exact preprocessing representation of single linkage\. Given a dissimilarityd∈𝒟⁡\(𝒳\)d\\in\\mathcal\{D\}\(\\mathcal\{X\}\), consider the weighted complete graph whose vertices are the points in𝒳\\mathcal\{X\}with edge\-weightdd\. The bottleneck of a pathγ\\gammais the largest weight along the path, that is,maxe∈γ⁡d⁡\(e\)\\max\_\{e\\in\\gamma\}d\(e\)\. The*minimum bottleneck*between two pointsx,y∈𝒳x,y\\in\\mathcal\{X\}is

B∗​\(d\)​\(x,y\):=\{minγ∈Γx,y⁡max\{u,v\}∈γ⁡d⁡\(u,v\),x≠y,0,x=y,\\displaystyle B^\{\*\}\(d\)\(x,y\)\\ :=\\ \\begin\{cases\}\\displaystyle\\min\_\{\\gamma\\in\\Gamma\_\{x,y\}\}\\max\_\{\\\{u,v\\\}\\in\\gamma\}d\(u,v\),&x\\neq y,\\\\\[4\.30554pt\] 0,&x=y,\\end\{cases\}\(4\)whereΓx,y\\Gamma\_\{x,y\}denotes the set of paths fromxxtoyy\.

Although the definition ofB∗B^\{\*\}given in \([4](https://arxiv.org/html/2609.11173#S5.E4)\) involves a minimum over all paths, this quantity can efficiently be computed from any minimum spanning treeMM\. Indeed, ifγx,yM\\gamma^\{M\}\_\{x,y\}denotes the unique path fromxxtoyyinMM, then

B∗​\(d\)​\(x,y\)=max\{u,v\}∈γx,yM⁡d⁡\(u,v\)\.B^\{\*\}\(d\)\(x,y\)\\ =\\ \\max\_\{\\\{u,v\\\}\\in\\gamma^\{M\}\_\{x,y\}\}d\(u,v\)\.The equivalence between this minimum\-bottleneck construction and single linkage is standard; see, for example,[Carlsson and Mémoli \(2010\)](https://arxiv.org/html/2609.11173#bib.bib4)\.

The dissimilarityB∗​\(d\)B^\{\*\}\(d\)is an ultrametric, commonly called the*subdominant ultrametric*associated withdd; its values are also known as minimax\-path distances\. Moreover, for everyr≥0r\\geq 0, two pointsxxandyysatisfyB∗​\(d\)​\(x,y\)≤rB^\{\*\}\(d\)\(x,y\)\\leq rif and only if they belong to the same connected component of the graph whose edges are the pairs\{u,v\}\\\{u,v\\\}satisfyingd⁡\(u,v\)≤rd\(u,v\)\\leq r\. These connected components, asrrvaries, are exactly the clusters produced by single linkage\. Hence, by the ultrametric\-hierarchy correspondence detailed in Equation \([3](https://arxiv.org/html/2609.11173#S5.E3)\), we haveTSL​\(d\)=ΨB∗​\(d\)\.T\_\{\\mathrm\{SL\}\}\(d\)\\ =\\ \\Psi\_\{B^\{\*\}\(d\)\}\.Moreover, becauseTglobT\_\{\\mathrm\{glob\}\}is exact on ultrametric inputs, we obtainTSL​\(d\)=Tglob​\(B∗​\(d\)\)T\_\{\\mathrm\{SL\}\}\(d\)=T\_\{\\mathrm\{glob\}\}\\bigl\(B^\{\*\}\(d\)\\bigr\), and hence

TSL=Tglob∘B∗\.\\displaystyle T\_\{\\mathrm\{SL\}\}\\ =\\ T\_\{\\mathrm\{glob\}\}\\circ B^\{\*\}\.\(5\)Thus, single linkage can be viewed as first replacing the original dissimilarity by its minimax\-path ultrametric and then extracting the globally separated clusters\. Observe also that the same reasoning applies to any method that is exact on ultrametric inputs; in particular, we may replaceTglobT\_\{\\mathrm\{glob\}\}in \([5](https://arxiv.org/html/2609.11173#S5.E5)\) withTlocT\_\{\\mathrm\{loc\}\}orTstableT\_\{\\mathrm\{stable\}\}\.

Finally, the transformationB∗:d∈𝒟⁡\(𝒳\)↦B∗​\(d\)B^\{\*\}\\colon d\\in\\mathcal\{D\}\(\\mathcal\{X\}\)\\mapsto B^\{\*\}\(d\)satisfies

B∗\(βd\)=βB∗\(d\),B∗\(dϕ\)=B∗\(d\)ϕ,andB∗\(u\)=u,B^\{\*\}\(\\beta d\)=\\beta B^\{\*\}\(d\),\\qquad B^\{\*\}\(d\_\{\\phi\}\)=B^\{\*\}\(d\)\_\{\\phi\},\\quad\\text\{ and \}\\quad B^\{\*\}\(u\)=u,for everyβ\>0\\beta\>0, every permutationϕ\\phi, and every ultrametricuu\. Hence,B∗B^\{\*\}preserves scale invariance, permutation invariance, and exactness on ultrametrics\. Together with the corresponding properties ofTglobT\_\{\\mathrm\{glob\}\}, the factorizationTSL=Tglob∘B∗T\_\{\\mathrm\{SL\}\}=T\_\{\\mathrm\{glob\}\}\\circ B^\{\*\}yields these properties for single linkage\.

*PCA preprocessing\.*In this paragraph only, we restrict to Euclidean distancesdd, meaning that there exist points\(zx\)x∈𝒳⊂Rp\(z\_\{x\}\)\_\{x\\in\\mathcal\{X\}\}\\subset\\mathbb\{R\}^\{p\}such thatd⁡\(x,y\)=∥zx−zy∥2d\(x,y\)\\ =\\ \\lVert z\_\{x\}\-z\_\{y\}\\rVert\_\{2\}for allx,y∈𝒳x,y\\in\\mathcal\{X\}\. Center these points, that is definez~x:=zx−z¯\\tilde\{z\}\_\{x\}:=z\_\{x\}\-\\bar\{z\}wherez¯=1\|𝒳\|​∑x∈𝒳zx\\bar\{z\}=\\frac\{1\}\{\|\\mathcal\{X\}\|\}\\sum\_\{x\\in\\mathcal\{X\}\}z\_\{x\}, and letPrP\_\{r\}be the orthogonal projection onto the subspace spanned by the firstrrprincipal components of the centered configuration\(z~x\)x∈𝒳\(\\tilde\{z\}\_\{x\}\)\_\{x\\in\\mathcal\{X\}\}\. Assuming that this principal subspace is uniquely defined, we define

μPCA,r\(d\)\(x,y\):=∥Prz~x−Prz~y∥2\.\\mu\_\{\\mathrm\{PCA\},r\}\(d\)\(x,y\):=\\bigl\\lVert P\_\{r\}\\widetilde\{z\}\_\{x\}\-P\_\{r\}\\widetilde\{z\}\_\{y\}\\bigr\\rVert\_\{2\}\.This definition does not depend on the chosen Euclidean realization, since centered realizations of the same dissimilarity differ only by an orthogonal transformation\. Under the assumption that the projected pointsPr​z~xP\_\{r\}\\widetilde\{z\}\_\{x\},x∈𝒳x\\in\\mathcal\{X\}are pairwise distinct,888Otherwise, two distinct pointsx≠yx\\neq ymay have the same projection, in which caseμPCA,r​\(d\)​\(x,y\)=0\\mu\_\{\\mathrm\{PCA\},r\}\(d\)\(x,y\)=0and the construction yields a pseudodissimilarity rather than an element of𝒟⁡\(𝒳\)\\mathcal\{D\}\(\\mathcal\{X\}\)\. Such collisions may occur in particular for highly symmetric configurations, including some Euclidean realizations of ultrametrics\. In that case, the chosen projection dimensionrris unsuitable if the preprocessing is required to remain dissimilarity\-valued\.μPCA,r​\(d\)\\mu\_\{\\mathrm\{PCA\},r\}\(d\)is a dissimilarity\.

Multiplyingddbyβ\>0\\beta\>0amounts to multiplying the realizing configuration byβ\\beta, which leaves its principal subspace unchanged and multiplies all projected distances byβ\\beta\. HenceμPCA,r​\(β​d\)=β​μPCA,r​\(d\)\\mu\_\{\\mathrm\{PCA\},r\}\(\\beta d\)=\\beta\\,\\mu\_\{\\mathrm\{PCA\},r\}\(d\), and thus PCA preprocessing preserves scale invariance\.

PCA preprocessing is also equivariant under relabelling\. Indeed, permuting the points leaves the covariance operator, and hence the principalrr\-dimensional subspace, unchanged, while merely relabelling the projected points\. Therefore,μPCA,r​\(dϕ\)=μPCA,r​\(d\)ϕ\\mu\_\{\\mathrm\{PCA\},r\}\(d\_\{\\phi\}\)=\\mu\_\{\\mathrm\{PCA\},r\}\(d\)\_\{\\phi\}, and thus PCA preprocessing also preserves permutation invariance\.

In general, however, PCA preprocessing need not preserve the remaining axioms\. A partition strengthening may alter the principal subspace, so its image need not remain a strengthening of the projected dissimilarity\. Moreover, for fixedrr, every output ofμPCA,r\\mu\_\{\\mathrm\{PCA\},r\}is a Euclidean dissimilarity realizable in dimension at mostrr\. Because a general dissimilarity need not have this form,μPCA,r\\mu\_\{\\mathrm\{PCA\},r\}is not surjective onto𝒟⁡\(𝒳\)\\mathcal\{D\}\(\\mathcal\{X\}\)\. Finally, projecting an ultrametric realization need not produce an ultrametric inducing the same hierarchy\.

*HDBSCAN and mutual\-reachability preprocessing\.*Fixk≥1k\\geq 1, and letdcore​\(x\)d\_\{\\mathrm\{core\}\}\(x\)be the dissimilarity fromxxto itskk\-th nearest neighbor\. The mutual\-reachability transformation used in HDBSCAN is defined, forx≠yx\\neq y, by

μHDBSCAN​\(d\)​\(x,y\):=max⁡\{dcore​\(x\),dcore​\(y\),d⁡\(x,y\)\},\\mu\_\{\\rm HDBSCAN\}\(d\)\(x,y\):=\\max\\bigl\\\{d\_\{\\mathrm\{core\}\}\(x\),d\_\{\\mathrm\{core\}\}\(y\),d\(x,y\)\\bigr\\\},with zero diagonal\. Since all nearest\-neighbor dissimilarities scale together withdd,

μHDBSCAN​\(β​d\)=β​μHDBSCAN​\(d\)for every​β\>0\.\\mu\_\{\\rm HDBSCAN\}\(\\beta d\)=\\beta\\,\\mu\_\{\\rm HDBSCAN\}\(d\)\\qquad\\text\{for every \}\\beta\>0\.Moreover, the construction is equivariant under relabelling:μHDBSCAN​\(dϕ\)=μHDBSCAN​\(d\)ϕ\.\\mu\_\{\\rm HDBSCAN\}\(d\_\{\\phi\}\)=\\mu\_\{\\rm HDBSCAN\}\(d\)\_\{\\phi\}\.Hence mutual\-reachability preprocessing preserves both scale invariance and permutation invariance of the downstream hierarchical method\.

*Similarity\-valued preprocessing\.*The same preservation principle extends to pipelines expressed in terms of similarities\. For example, the Gaussian kernelkσ​\(x,y\):=exp⁡\{−\(d⁡\(x,y\)\)22​σ2\}k\_\{\\sigma\}\(x,y\):=\\exp\\left\\\{\-\\frac\{\\left\(d\(x,y\)\\right\)^\{2\}\}\{2\\sigma^\{2\}\}\\right\\\}is a strictly decreasing function ofd⁡\(x,y\)d\(x,y\)\. It therefore converts a dissimilarity strengthening into the corresponding similarity strengthening: within\-cluster similarities increase, whereas cross\-cluster similarities decrease\. The transformation also commutes with permutations\. \(Note that this example lies outside the dissimilarity\-valued framework adopted above, so it should be understood using the similarity analog of the relevant axioms\.\)

## 6Related Work

Prior to Kleinberg,[Puzicha et al\. \(2000\)](https://arxiv.org/html/2609.11173#bib.bib17)developed an axiomatic framework for clustering objective functions, based on properties such as monotonicity, invariance, and robustness, and identified objectives satisfying these requirements\. Kleinberg’s impossibility theorem\([Kleinberg, 2002](https://arxiv.org/html/2609.11173#bib.bib8)\)subsequently motivated a broad line of work that modifies the axioms, the input, or the problem formulation in order to circumvent the impossibility result\([Ben\-David and Ackerman, 2008](https://arxiv.org/html/2609.11173#bib.bib6);[Zadeh and Ben\-David, 2009](https://arxiv.org/html/2609.11173#bib.bib15);[Strazzeri and Sánchez\-García, 2022](https://arxiv.org/html/2609.11173#bib.bib14);[Willson and Warnow, 2024](https://arxiv.org/html/2609.11173#bib.bib7)\)\. Our work contributes to this line by asking whether the impossibility can be resolved by changing the form of the output, from a flat partition to a hierarchy, while retaining a dissimilarity\-based input and direct hierarchical analogs of Kleinberg’s axioms\.

More precisely, these works relax Kleinberg’s framework in different ways\.[Ben\-David and Ackerman \(2008\)](https://arxiv.org/html/2609.11173#bib.bib6)consider an analogous axiomatic setting for clustering\-quality functions and establish the existence of functions satisfying their axioms\.[Zadeh and Ben\-David \(2009\)](https://arxiv.org/html/2609.11173#bib.bib15)take the desired number of clusters as an additional input and correspondingly restrict richness to partitions having that number of clusters\.[Cohen\-Addad et al\. \(2018\)](https://arxiv.org/html/2609.11173#bib.bib5)introduce a cost function to estimate the number of clusters and weaken consistency by requiring it only when this estimate remains unchanged\.[Strazzeri and Sánchez\-García \(2022\)](https://arxiv.org/html/2609.11173#bib.bib14)restrict the transformations allowed by the consistency axiom and extend the setting to graph clustering\. In the graph setting,[Van Laarhoven and Marchiori \(2014\)](https://arxiv.org/html/2609.11173#bib.bib16)adapt axioms for clustering\-quality functions and introduce adaptive scale modularity, while[Willson and Warnow \(2024\)](https://arxiv.org/html/2609.11173#bib.bib7)translate Kleinberg’s axioms to unweighted graphs and identify clustering methods satisfying their axioms\. A different structural perspective is developed by[Carlsson and Mémoli \(2013\)](https://arxiv.org/html/2609.11173#bib.bib23), who study clustering schemes through functoriality, requiring compatibility of clustering outputs with suitable maps between input metric spaces\.

A distinct line of work develops axiomatic characterizations specific to hierarchical clustering\.[Carlsson and Mémoli \(2010\)](https://arxiv.org/html/2609.11173#bib.bib4)view hierarchical clustering as a map from finite metric spaces to proximity dendrograms, equivalently ultrametrics, and characterize single linkage using normalization, separation, and functoriality\. Functoriality requires every distance\-nonincreasing map between input metric spaces to remain distance\-nonincreasing between the corresponding output ultrametric spaces, including maps between datasets of different sizes\. As distance\-preserving relabelings are examples of such maps, functoriality is substantially stronger than permutation invariance\. In fact, it has no direct counterpart in our framework, because it compares outputs across different ground sets and constrains their merge scales, whereas our outputs are unweighted hierarchies on a fixed set\.

In a complementary direction,[Ackerman et al\. \(2010\)](https://arxiv.org/html/2609.11173#bib.bib10)and[Ackerman and Ben\-David \(2016\)](https://arxiv.org/html/2609.11173#bib.bib13)study characterizations of linkage\-based methods\. In particular, the latter show that locality together with outer consistency characterizes hierarchical linkage methods\. Outer consistency increases dissimilarities between the clusters of the partition induced by cutting a dendrogram at a selected height, while keeping within\-block dissimilarities fixed\. In contrast, partition consistency is thus more faithful to Kleinberg’s consistency as it allows within\-block contractions and applies to every partition represented in the hierarchy\. As a result, several linkage methods other than single linkage satisfy outer consistency whereas failing partition consistency\.

Related axiomatic frameworks have also been developed for asymmetric dissimilarities\. In particular,[Carlsson et al\. \(2014\)](https://arxiv.org/html/2609.11173#bib.bib24)introduce hierarchical quasi\-clustering for asymmetric networks and obtain a uniqueness result under their axioms, while[Carlsson et al\. \(2017\)](https://arxiv.org/html/2609.11173#bib.bib25)characterize broader families of admissible hierarchical methods for asymmetric networks\.

These earlier frameworks use additional structural requirements to characterize particular methods or algorithmic families\. Our axioms remain closer to Kleinberg’s original requirements and admit a much more diverse class of methods, whose structure under refinement is a key focus of our work\.

A complementary perspective axiomatizes objective functions for evaluating hierarchical clusterings rather than the behavior of the clustering method itself\.[Dasgupta \(2016\)](https://arxiv.org/html/2609.11173#bib.bib1)introduced a cost function for similarity\-based hierarchical clustering, thereby formulating hierarchical clustering as an optimization problem\. Building on this perspective,[Cohen\-Addad et al\. \(2019\)](https://arxiv.org/html/2609.11173#bib.bib2)give an axiomatic characterization of a broad class of admissible objective functions for both similarity\- and dissimilarity\-based hierarchical clustering, including Dasgupta’s objective\. This differs from our approach: their axioms determine which objective functions appropriately score a hierarchy, whereas ours directly constrain the behavior of the map from dissimilarities to hierarchies\.

Population\-level approaches provide another, substantially different, axiomatic perspective on hierarchical clustering\.[Thomann et al\. \(2015\)](https://arxiv.org/html/2609.11173#bib.bib26)axiomatize hierarchical clustering of probability measures without assuming an underlying metric or dissimilarity\. Their starting point is a user\-specified clustering on a class of elementary measures, and they study how this clustering can be extended to more general distributions\. Under suitable conditions, additivity and continuity requirements determine a unique such extension\. More recently,[Arias\-Castro and Coda \(2025\)](https://arxiv.org/html/2609.11173#bib.bib9)propose an axiomatic definition of hierarchical clustering based on the topology of density level sets\. Rather than axiomatizing a clustering method and its behavior under transformations of the input, they seek to characterize which subsets of the support of a density should constitute population clusters\. They first consider piecewise\-constant densities and require clusters to have connected interior, not to split connected regions of constant density, and to be surrounded by regions of lower density\. Among the cluster trees satisfying these requirements, they select the finest one, and subsequently extend the construction to more general densities, recovering Hartigan’s cluster tree under suitable conditions\.

Our framework is different in both its input and its object of study\. In the spirit of Kleinberg, we axiomatize maps from pairwise dissimilarities to hierarchies: The axioms constrain how the output hierarchy behaves when the input dissimilarity is rescaled, strengthened, or relabeled, as well as which hierarchical structures the method can realize\. Population\-level approaches such as[Thomann et al\. \(2015\)](https://arxiv.org/html/2609.11173#bib.bib26)and[Arias\-Castro and Coda \(2025\)](https://arxiv.org/html/2609.11173#bib.bib9)specify what a population hierarchy should be based on distributional or topological information, whereas ours impose structural requirements on hierarchical clustering methods that act directly on dissimilarities\.

## 7Conclusion

Kleinberg’s impossibility theorem shows that scale invariance, richness, and consistency cannot be jointly satisfied by a flat clustering method\. In this work, we have shown that this incompatibility disappears when the output is allowed to be hierarchical\. Natural hierarchical analogs of Kleinberg’s axioms are jointly satisfiable; in fact, there exist uncountably many hierarchical clustering methods satisfying them\.

Moreover, our work reveals that the set of admissible hierarchical methods is extremely complex\. Under the refinement order, the class of admissible methods is a poset that has uncountable height, width, and cellularity, has neither a greatest nor a least element, and contains uncountably many pairwise incompatible maximal elements\. Nevertheless, admissible methods cannot differ arbitrarily\. Every admissible method contains a hierarchy of sufficiently well\-separated clusters, and every finite collection of admissible methods therefore shares a nontrivial common backbone\. Thus, the axioms simultaneously allow for substantial freedom in the additional clusters reported by a method while enforcing a common conservative hierarchical structure\.

We also considered extensions that are specific to the hierarchical setting\. Adding exactness on ultrametric inputs preserves the main structural picture while imposing a stronger connection between dissimilarities and their canonical hierarchies\. In addition, we studied preprocessing transformations and identified conditions under which they preserve the axioms, allowing the framework to apply to clustering pipelines in which the input dissimilarity is transformed before the hierarchy is constructed\.

This study raises several open questions\. A first one is to determine which other hierarchical clustering methods are admissible, and, in particular, to identify explicit maximal admissible methods\. Another one is to understand which additional axioms meaningfully reduce the large admissible class\. Our results on ultrametric exactness provide one step in this direction, yielding, in particular, a least element among strongly admissible methods\.

Finally, hierarchical clustering has an expressive advantage over flat clustering, as it does not require the number of clusters to be specified in advance and returns a richer, nested cluster structure that cannot be represented by a single partition\. Nevertheless, some downstream tasks ultimately require a flat partition, thereby requiring a cut of the hierarchy, that is, a selection of a set of clusters from the hierarchy that forms a partition\. Kleinberg’s theorem implies that no such cut\-selection rule can induce a flat clustering method that simultaneously satisfies scale invariance, richness, and consistency\. Characterizing the trade\-offs inherent in cut selection, as well as the benefits of hierarchy\-aware downstream procedures that avoid this projection altogether, remains an interesting direction for future work\.

## Appendix AAlternative Axioms and Robustness

Whereas the main text focuses on the canonical axiom systemαadm\\alpha\_\{\\mathrm\{adm\}\}and its strengthening by exactness on ultrametrics, several closely related formulations are also natural in the hierarchical setting\. This appendix collects the variants that are useful for comparison, describes their logical relations, and shows that the principal structural conclusions of Section[4](https://arxiv.org/html/2609.11173#S4)are robust to many of these choices\.

We denote byα\\alphaa requirement, or a property, that may or may not be satisfied by a given hierarchical clustering method\. In particular, all axioms considered in this paper are requirements\. For a requirementα\\alpha, let

ℋα​\(𝒳\):=\{T∈ℋ⁡\(𝒳\):T​satisfies​α\}\.\\mathcal\{H\}\_\{\\alpha\}\(\\mathcal\{X\}\):=\\\{T\\in\\mathcal\{H\}\(\\mathcal\{X\}\):T\\text\{ satisfies \}\\alpha\\\}\.We say that a requirementα2\\alpha\_\{2\}is*stronger*than another requirementα1\\alpha\_\{1\}, denotedα2⊳α1\\alpha\_\{2\}\\triangleright\\alpha\_\{1\}, ifℋα2​\(𝒳\)⊆ℋα1​\(𝒳\)\.\\mathcal\{H\}\_\{\\alpha\_\{2\}\}\(\\mathcal\{X\}\)\\subseteq\\mathcal\{H\}\_\{\\alpha\_\{1\}\}\(\\mathcal\{X\}\)\.If bothα2⊳α1\\alpha\_\{2\}\\triangleright\\alpha\_\{1\}andα1⊳α2\\alpha\_\{1\}\\triangleright\\alpha\_\{2\}hold, we say the two requirements are equivalent, and writeα1≡α2\\alpha\_\{1\}\\equiv\\alpha\_\{2\}\. Finally, the conjunction ofα1\\alpha\_\{1\}andα2\\alpha\_\{2\}is denoted byα1∧α2\\alpha\_\{1\}\\wedge\\alpha\_\{2\}\.

We retain the notationαsca,αpr,αpc,αperm\\alpha\_\{\\mathrm\{sca\}\},\\alpha\_\{\\mathrm\{pr\}\},\\alpha\_\{\\mathrm\{pc\}\},\\alpha\_\{\\mathrm\{perm\}\}for scale invariance, partition richness, partition consistency, and permutation invariance, respectively, andαsuhr\\alpha\_\{\\mathrm\{suhr\}\}for exactness on ultrametrics\. Thus

αadm:=αsca∧αpr∧αpc∧αperm,αadm\+:=αadm∧αsuhr≡αsca∧αsuhr∧αpc∧αperm,\\alpha\_\{\\mathrm\{adm\}\}:=\\alpha\_\{\\mathrm\{sca\}\}\\wedge\\alpha\_\{\\mathrm\{pr\}\}\\wedge\\alpha\_\{\\mathrm\{pc\}\}\\wedge\\alpha\_\{\\mathrm\{perm\}\},\\qquad\\alpha\_\{\\mathrm\{adm\+\}\}:=\\alpha\_\{\\mathrm\{adm\}\}\\wedge\\alpha\_\{\\mathrm\{suhr\}\}\\ \\equiv\\ \\alpha\_\{\\mathrm\{sca\}\}\\wedge\\alpha\_\{\\mathrm\{suhr\}\}\\wedge\\alpha\_\{\\mathrm\{pc\}\}\\wedge\\alpha\_\{\\mathrm\{perm\}\},where the last equivalence holds because exactness on ultrametrics implies partition richness\. More generally, an*axiom system*is a conjunction of requirements\.

Table[1](https://arxiv.org/html/2609.11173#A1.T1)summarizes all the requirements/properties considered in this paper, grouped into four categories: admissibility, invariance, richness, and consistency\.

*Category**Symbol / name*γ⊓\\gamma\_\{\\sqcap\}γ⊔up\\gamma^\{\\mathrm\{up\}\}\_\{\\sqcup\}γ⊔\\gamma\_\{\\sqcup\}γrp\\gamma\_\{\\mathrm\{rp\}\}Admissibilityαadm\\alpha\_\{\\mathrm\{adm\}\}\(admissibility\)✓✓×\\times×\\timesαadm\+\\alpha\_\{\\mathrm\{adm\+\}\}\(strong admissibility\)✓✓×\\times✓Invarianceαsca\\alpha\_\{\\mathrm\{sca\}\}\(scale invariance\)✓✓✓✓αord\\alpha\_\{\\mathrm\{ord\}\}\(order invariance\)✓✓✓✓αperm\\alpha\_\{\\mathrm\{perm\}\}\(permutation invariance\)✓✓✓✓Richness variantsαpr\\alpha\_\{\\mathrm\{pr\}\}\(partition richness\)×\\times✓\\checkmark✓×\\timesαcwr\\alpha\_\{\\mathrm\{cwr\}\}\(cluster\-wise richness\)×\\times✓✓×\\timesαhr\\alpha\_\{\\mathrm\{hr\}\}\(hierarchical richness\)×\\times✓✓×\\timesαshr\\alpha\_\{\\mathrm\{shr\}\}\(strict hierarchical richness\)×\\times×\\times×\\times×\\timesαuhr\\alpha\_\{\\mathrm\{uhr\}\}\(refinement on ultrametrics\)✓✓✓✓αsuhr\\alpha\_\{\\mathrm\{suhr\}\}\(exact on ultrametrics\)✓✓✓✓αbh\\alpha\_\{\\mathrm\{bh\}\}\(backbone\-hierarchy\)✓✓✓×\\timesαbh1\\alpha\_\{\\mathrm\{bh\}\}^\{1\}\(unit\-backbone\-hierarchy\)✓✓✓×\\timesConsistency variantsαpc\\alpha\_\{\\mathrm\{pc\}\}\(partition consistency\)✓✓×\\times✓αcwc\\alpha\_\{\\mathrm\{cwc\}\}\(cluster\-wise consistency\)✓✓✓✓αhc\\alpha\_\{\\mathrm\{hc\}\}\(hierarchical consistency\)✓✓✓✓αabc\\alpha\_\{\\mathrm\{abc\}\}\(absence consistency\)✓✓×\\times✓Table 1:Summary of the \(alternative\) axioms and their structural properties\. Gray shading indicates the canonical axioms; magenta shading indicates exactness on ultrametrics and strong admissibility\. Unshaded rows correspond to requirements introduced in this appendix\. A checkmark in theγ⊓\\gamma\_\{\\sqcap\},γ⊔up\\gamma^\{\\mathrm\{up\}\}\_\{\\sqcup\}, orγ⊔\\gamma\_\{\\sqcup\}column means that the corresponding method class is closed, respectively, under finite nonempty intersections, unions of nonempty upward\-directed families, or unions of pairwise compatible nonempty families; a cross indicates that no such closure property is asserted\. The three closure propertiesγ⊓\\gamma\_\{\\sqcap\},γ⊔up\\gamma^\{\\mathrm\{up\}\}\_\{\\sqcup\}, orγ⊔\\gamma\_\{\\sqcup\}are preserved under finite conjunctions of axioms; see Proposition[38](https://arxiv.org/html/2609.11173#Thmtheorem38)\. Similarly, a checkmark in the columnγrp\\gamma\_\{\\mathrm\{rp\}\}means that the axiom is representable as a relational property or as a conjunction of relational properties \(we refer to Section[B](https://arxiv.org/html/2609.11173#A2)for a precise definition and to Section[B\.2](https://arxiv.org/html/2609.11173#A2.SS2)for a proof of the checkmarks in the last column\)\.### A\.1Alternative Axiom Formulations

#### A\.1\.1Order invariance

Scale invariance can be strengthened by requiring the output to depend only on the ordering of the pairwise dissimilarities\.

###### Definition 30\(Order invarianceαord\\alpha\_\{\\mathrm\{ord\}\}\)\.

A methodT∈ℋ⁡\(𝒳\)T\\in\\mathcal\{H\}\(\\mathcal\{X\}\)is*order invariant*if, for everyd∈𝒟⁡\(𝒳\)d\\in\\mathcal\{D\}\(\\mathcal\{X\}\)and every strictly increasingg:R≥0→R≥0g:\\mathbb\{R\}\_\{\\geq 0\}\\to\\mathbb\{R\}\_\{\\geq 0\}satisfyingg⁡\(0\)=0g\(0\)=0, we haveT⁡\(g∘d\)=T⁡\(d\),T\(g\\circ d\)=T\(d\),where\(g∘d\)​\(x,y\):=g⁡\(d⁡\(x,y\)\)\(g\\circ d\)\(x,y\):=g\(d\(x,y\)\)for all elementsx,y∈𝒳x,y\\in\\mathcal\{X\}\.

#### A\.1\.2Richness variants

Partition richness requires that for any partition, there is a dissimilarity functionddthat makes the partition part of the output hierarchy\. In the hierarchical setting, one may instead require the realization of individual clusters or of entire hierarchies\.

###### Definition 31\(Richness variants\)\.

A methodT∈ℋ⁡\(𝒳\)T\\in\\mathcal\{H\}\(\\mathcal\{X\}\)satisfies

- •*cluster\-wise richness*αcwr\\alpha\_\{\\mathrm\{cwr\}\}if, for every nonemptyC⊆𝒳C\\subseteq\\mathcal\{X\}, there existsd∈𝒟⁡\(𝒳\)d\\in\\mathcal\{D\}\(\\mathcal\{X\}\)such thatC∈T⁡\(d\)C\\in T\(d\);
- •*hierarchical richness*αhr\\alpha\_\{\\mathrm\{hr\}\}if, for everyΨ∈𝒯⁡\(𝒳\)\\Psi\\in\\mathcal\{T\}\(\\mathcal\{X\}\), there existsd∈𝒟⁡\(𝒳\)d\\in\\mathcal\{D\}\(\\mathcal\{X\}\)such thatΨ⊆T⁡\(d\)\\Psi\\subseteq T\(d\);
- •*strict hierarchical richness*αshr\\alpha\_\{\\mathrm\{shr\}\}if, for everyΨ∈𝒯⁡\(𝒳\)\\Psi\\in\\mathcal\{T\}\(\\mathcal\{X\}\), there existsd∈𝒟⁡\(𝒳\)d\\in\\mathcal\{D\}\(\\mathcal\{X\}\)such thatT⁡\(d\)=ΨT\(d\)=\\Psi;
- •*refinement on ultrametrics*αuhr\\alpha\_\{\\mathrm\{uhr\}\}if,Ψu⊆T⁡\(u\)\\Psi\_\{u\}\\subseteq T\(u\)for every ultrametricuu\.

Refinement on ultrametrics is an alternative to exactness on ultrametrics \(Definition[25](https://arxiv.org/html/2609.11173#Thmtheorem25); denotedαsuhr\\alpha\_\{\\mathrm\{suhr\}\}\), and parallel toαsuhr\\alpha\_\{\\mathrm\{suhr\}\}, this can be regarded as a richness variant\.

#### A\.1\.3Consistency variants

Partition consistency preserves all blocks of a partition under a strengthening of that partition\. A natural alternative is to require the same property cluster by cluster\. To compare the two formulations, we first extend the notion of strengthening\.

###### Definition 32\(ψ\\psi\-strengthening\)\.

Letψ⊆2𝒳∖\{∅\}\\psi\\subseteq 2^\{\\mathcal\{X\}\}\\setminus\\\{\\emptyset\\\}be nonempty\. A dissimilarityd′∈𝒟⁡\(𝒳\)d^\{\\prime\}\\in\\mathcal\{D\}\(\\mathcal\{X\}\)is aψ\\psi\-strengthening ofd∈𝒟⁡\(𝒳\)d\\in\\mathcal\{D\}\(\\mathcal\{X\}\)if, for everyC∈ψC\\in\\psi,

- •\(Intra\-cluster contraction\)d′​\(x,y\)≤d⁡\(x,y\)d^\{\\prime\}\(x,y\)\\leq d\(x,y\)for allx,y∈Cx,y\\in C;
- •\(Inter\-cluster expansion\)d′​\(x,y\)≥d⁡\(x,y\)d^\{\\prime\}\(x,y\)\\geq d\(x,y\)for allx∈Cx\\in Candy∉Cy\\notin C;
- •\(External pairs unconstrained\) Pairs with both endpoints outside⋃C∈ψC\\bigcup\_\{C\\in\\psi\}Care unconstrained\.

Observe that in the case of a𝒞\\mathcal\{C\}\-strengthening as defined in Definition[1](https://arxiv.org/html/2609.11173#Thmtheorem1),𝒞\\mathcal\{C\}is a partition, so there exists no pairx,y∉⋃C∈𝒞Cx,y\\notin\\bigcup\_\{C\\in\\mathcal\{C\}\}Cas𝒞\\mathcal\{C\}covers𝒳\\mathcal\{X\}\. However, because the setψ\\psiin Definition[32](https://arxiv.org/html/2609.11173#Thmtheorem32)does not necessarily cover𝒳\\mathcal\{X\}, we make the treatment of such external pairs explicit\.

###### Definition 33\(Consistency variants\)\.

A methodT∈ℋ⁡\(𝒳\)T\\in\\mathcal\{H\}\(\\mathcal\{X\}\)satisfies

- •*cluster\-wise consistency*αcwc\\alpha\_\{\\mathrm\{cwc\}\}if, wheneverC∈T⁡\(d\)C\\in T\(d\), every\{C\}\\\{C\\\}\-strengtheningd′d^\{\\prime\}ofddsatisfiesC∈T⁡\(d′\)C\\in T\(d^\{\\prime\}\);
- •*hierarchical consistency*αhc\\alpha\_\{\\mathrm\{hc\}\}if, wheneverψ⊆T⁡\(d\)\\psi\\subseteq T\(d\), everyψ\\psi\-strengtheningd′d^\{\\prime\}ofddsatisfiesψ⊆T⁡\(d′\)\\psi\\subseteq T\(d^\{\\prime\}\)\.

For completeness, one can also state partition consistency in the reverse direction\. Say thatd′d^\{\\prime\}is a𝒞\\mathcal\{C\}\-*weakening*ofddwhenddis a𝒞\\mathcal\{C\}\-strengthening ofd′d^\{\\prime\}, and call a method*absence consistent*\(αabc\\alpha\_\{\\mathrm\{abc\}\}\) if

𝒞⊈T⁡\(d\)⟹𝒞⊈T⁡\(d′\)\\mathcal\{C\}\\not\\subseteq T\(d\)\\quad\\Longrightarrow\\quad\\mathcal\{C\}\\not\\subseteq T\(d^\{\\prime\}\)for every𝒞\\mathcal\{C\}\-weakeningd′d^\{\\prime\}ofdd\. As shown below, this is simply the contrapositive formulation of partition consistency\.

### A\.2Relations Among the Axioms

The elementary implications among the richness and invariance requirements areαord⊳αsca\\alpha\_\{\\mathrm\{ord\}\}\\,\\triangleright\\,\\alpha\_\{\\mathrm\{sca\}\},αsuhr⊳αshr⊳αhr⊳αpr⊳αcwr,\\alpha\_\{\\mathrm\{suhr\}\}\\,\\triangleright\\,\\alpha\_\{\\mathrm\{shr\}\}\\,\\triangleright\\,\\alpha\_\{\\mathrm\{hr\}\}\\,\\triangleright\\,\\alpha\_\{\\mathrm\{pr\}\}\\,\\triangleright\\,\\alpha\_\{\\mathrm\{cwr\}\},as well asαsuhr⊳αuhr⊳αhr\\alpha\_\{\\mathrm\{suhr\}\}\\,\\triangleright\\,\\alpha\_\{\\mathrm\{uhr\}\}\\,\\triangleright\\,\\alpha\_\{\\mathrm\{hr\}\}\. The following proposition shows that the consistency variants collapse more strongly than their definitions suggest\.

###### Proposition 34\.

We haveαcwc≡αhc⊳αpc≡αabc\\alpha\_\{\\mathrm\{cwc\}\}\\equiv\\alpha\_\{\\mathrm\{hc\}\}\\triangleright\\alpha\_\{\\mathrm\{pc\}\}\\equiv\\alpha\_\{\\mathrm\{abc\}\}\.

###### Proof\.

Hierarchical consistency implies cluster\-wise consistency by takingψ=\{C\}\\psi=\\\{C\\\}\. Conversely, ifd′d^\{\\prime\}is aψ\\psi\-strengthening ofdd, then it is a\{C\}\\\{C\\\}\-strengthening for everyC∈ψC\\in\\psi\. Cluster\-wise consistency therefore preserves everyC∈ψC\\in\\psi, provingαcwc≡αhc\\alpha\_\{\\mathrm\{cwc\}\}\\equiv\\alpha\_\{\\mathrm\{hc\}\}\. Takingψ=𝒞\\psi=\\mathcal\{C\}for a partition𝒞⊆T⁡\(d\)\\mathcal\{C\}\\subseteq T\(d\)givesαhc⊳αpc\\alpha\_\{\\mathrm\{hc\}\}\\triangleright\\alpha\_\{\\mathrm\{pc\}\}\.

Finally, ifd′d^\{\\prime\}is a𝒞\\mathcal\{C\}\-weakening ofdd, thenddis a𝒞\\mathcal\{C\}\-strengthening ofd′d^\{\\prime\}\. Hence

𝒞⊆T⁡\(d′\)⟹𝒞⊆T⁡\(d\)\\mathcal\{C\}\\subseteq T\(d^\{\\prime\}\)\\Longrightarrow\\mathcal\{C\}\\subseteq T\(d\)is precisely partition consistency applied to the pair\(d′,d\)\(d^\{\\prime\},d\), and its contrapositive is absence consistency\. ∎

We next record the relation between richness and the backbone theorem \(Theorem[18](https://arxiv.org/html/2609.11173#Thmtheorem18)\)\. Introduce the auxiliary*backbone property*αbh\\alpha\_\{\\mathrm\{bh\}\}:

Tsatisfiesαbh⟺Tglob𝜼⊑Tfor some separation margin sequence𝜼\.T\\text\{ satisfies \}\\alpha\_\{\\mathrm\{bh\}\}\\quad\\Longleftrightarrow\\quad T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}\\sqsubseteq T\\text\{ for some separation margin sequence \}\\boldsymbol\{\\eta\}\.SinceTglob𝜼T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}is hierarchically rich for all separation margin𝜼\\boldsymbol\{\\eta\},αbh⊳αhr\.\\alpha\_\{\\mathrm\{bh\}\}\\,\\triangleright\\,\\alpha\_\{\\mathrm\{hr\}\}\.

###### Proposition 35\.

Letα1⊳\(αsca∧αpc\)\\alpha\_\{1\}\\triangleright\(\\alpha\_\{\\mathrm\{sca\}\}\\wedge\\alpha\_\{\\mathrm\{pc\}\}\)andα2⊳\(αsca∧αcwc\)\\alpha\_\{2\}\\triangleright\(\\alpha\_\{\\mathrm\{sca\}\}\\wedge\\alpha\_\{\\mathrm\{cwc\}\}\)\. Then,

\(α1∧αpr\)≡\(α1∧αhr\)≡\(α1∧αbh\),and\(α2∧αcwr\)≡\(α2∧αpr\)≡\(α2∧αhr\)≡\(α2∧αbh\)\.\(\\alpha\_\{1\}\\wedge\\alpha\_\{\\mathrm\{pr\}\}\)\\equiv\(\\alpha\_\{1\}\\wedge\\alpha\_\{\\mathrm\{hr\}\}\)\\equiv\(\\alpha\_\{1\}\\wedge\\alpha\_\{\\mathrm\{bh\}\}\),\\quad\\text\{and\}\\quad\(\\alpha\_\{2\}\\wedge\\alpha\_\{\\mathrm\{cwr\}\}\)\\equiv\(\\alpha\_\{2\}\\wedge\\alpha\_\{\\mathrm\{pr\}\}\)\\equiv\(\\alpha\_\{2\}\\wedge\\alpha\_\{\\mathrm\{hr\}\}\)\\equiv\(\\alpha\_\{2\}\\wedge\\alpha\_\{\\mathrm\{bh\}\}\)\.

###### Proof\.

We prove the equivalences by the sandwiching argument i\.e, by showing

\(α1∧αpr\)⊳\(α1∧αhr\)⊳\(α1∧αbh\)and\(α1∧αpr\)⊲\(α1∧αhr\)⊲\(α1∧αbh\)\.\(\\alpha\_\{1\}\\wedge\\alpha\_\{\\mathrm\{pr\}\}\)\\triangleright\(\\alpha\_\{1\}\\wedge\\alpha\_\{\\mathrm\{hr\}\}\)\\triangleright\(\\alpha\_\{1\}\\wedge\\alpha\_\{\\mathrm\{bh\}\}\)\\quad\\text\{and\}\\quad\(\\alpha\_\{1\}\\wedge\\alpha\_\{\\mathrm\{pr\}\}\)\\triangleleft\(\\alpha\_\{1\}\\wedge\\alpha\_\{\\mathrm\{hr\}\}\)\\triangleleft\(\\alpha\_\{1\}\\wedge\\alpha\_\{\\mathrm\{bh\}\}\)\.At the beginning of the subsection, we already observedαbh⊳αhr⊳αpr⊳αcwr\\alpha\_\{\\mathrm\{bh\}\}\\,\\triangleright\\,\\alpha\_\{\\mathrm\{hr\}\}\\,\\triangleright\\,\\alpha\_\{\\mathrm\{pr\}\}\\,\\triangleright\\,\\alpha\_\{\\mathrm\{cwr\}\}while Lemma[57](https://arxiv.org/html/2609.11173#Thmtheorem57)gives the other direction:

αsca∧αpr∧αpc⊳αbh\.\\alpha\_\{\\mathrm\{sca\}\}\\wedge\\alpha\_\{\\mathrm\{pr\}\}\\wedge\\alpha\_\{\\mathrm\{pc\}\}\\,\\triangleright\\,\\alpha\_\{\\mathrm\{bh\}\}\.The same argument with cluster\-wise richness and cluster\-wise consistency gives

αsca∧αcwr∧αcwc⊳αbh\.\\alpha\_\{\\mathrm\{sca\}\}\\wedge\\alpha\_\{\\mathrm\{cwr\}\}\\wedge\\alpha\_\{\\mathrm\{cwc\}\}\\triangleright\\alpha\_\{\\mathrm\{bh\}\}\.The two chains of equivalences stated in the proposition follow by sandwiching\. ∎

The following corollary follows by applying Proposition[35](https://arxiv.org/html/2609.11173#Thmtheorem35)toα1=αsca∧αpc∧αperm\\alpha\_\{1\}=\\alpha\_\{\\mathrm\{sca\}\}\\wedge\\alpha\_\{\\mathrm\{pc\}\}\\wedge\\alpha\_\{\\mathrm\{perm\}\}\.

###### Corollary 36\.

The canonical axiom system admits the equivalent formulation

αadm≡αsca∧αbh∧αpc∧αperm\.\\alpha\_\{\\mathrm\{adm\}\}\\equiv\\alpha\_\{\\mathrm\{sca\}\}\\wedge\\alpha\_\{\\mathrm\{bh\}\}\\wedge\\alpha\_\{\\mathrm\{pc\}\}\\wedge\\alpha\_\{\\mathrm\{perm\}\}\.

Corollary[36](https://arxiv.org/html/2609.11173#Thmtheorem36)is useful because partition richness itself is not naturally preserved by intersections, whereas a common positive\-margin backbone is\. Hence this corollary, combined with Proposition[38](https://arxiv.org/html/2609.11173#Thmtheorem38), proves Lemma[23](https://arxiv.org/html/2609.11173#Thmtheorem23)\(i\)\.

### A\.3Robustness of the Structural Results

We now ask to what extent the structural results of Section[4](https://arxiv.org/html/2609.11173#S4)persist under the alternative requirements introduced above\. To state these results compactly, we introduce notation for several order\-theoretic properties of the class of methods induced by an axiom system\.

###### Definition 37\.

For an axiom systemα\\alpha, we write

α∈γ⊓\\displaystyle\\alpha\\in\\gamma\_\{\\sqcap\}⇔T⊓L∈ℋα​\(𝒳\)for every finite nonempty​L⊆ℋα​\(𝒳\),\\displaystyle\\iff T\_\{\\sqcap L\}\\in\\mathcal\{H\}\_\{\\alpha\}\(\\mathcal\{X\}\)\\quad\\text\{for every finite nonempty \}L\\subseteq\\mathcal\{H\}\_\{\\alpha\}\(\\mathcal\{X\}\),α∈γ⊔up\\displaystyle\\alpha\\in\\gamma^\{\\mathrm\{up\}\}\_\{\\sqcup\}⇔T⊔L∈ℋα​\(𝒳\)for every nonempty upward\-directed​L⊆ℋα​\(𝒳\),\\displaystyle\\iff T\_\{\\sqcup L\}\\in\\mathcal\{H\}\_\{\\alpha\}\(\\mathcal\{X\}\)\\quad\\text\{for every nonempty upward\-directed \}L\\subseteq\\mathcal\{H\}\_\{\\alpha\}\(\\mathcal\{X\}\),α∈γ⊔\\displaystyle\\alpha\\in\\gamma\_\{\\sqcup\}⇔T⊔L∈ℋα​\(𝒳\)for every nonempty pairwise\-compatible​L⊆ℋα​\(𝒳\),\\displaystyle\\iff T\_\{\\sqcup L\}\\in\\mathcal\{H\}\_\{\\alpha\}\(\\mathcal\{X\}\)\\quad\\text\{for every nonempty pairwise\-compatible \}L\\subseteq\\mathcal\{H\}\_\{\\alpha\}\(\\mathcal\{X\}\),α∈γem\\displaystyle\\alpha\\in\\gamma\_\{\\mathrm\{em\}\}⇔every​T∈ℋα​\(𝒳\)​is refined by some maximal element of​ℋα​\(𝒳\)\.\\displaystyle\\iff\\text\{every \}T\\in\\mathcal\{H\}\_\{\\alpha\}\(\\mathcal\{X\}\)\\text\{ is refined by some maximal element of \}\\mathcal\{H\}\_\{\\alpha\}\(\\mathcal\{X\}\)\.

Thus,γ⊓\\gamma\_\{\\sqcap\},γ⊔up\\gamma^\{\\mathrm\{up\}\}\_\{\\sqcup\},γ⊔\\gamma\_\{\\sqcup\}are the sets of requirements/properties that satisfy closure under finite intersections, upward\-directed unions, and compatible unions, respectively, whileγem\\gamma\_\{\\mathrm\{em\}\}is the set of requirements/properties that admit the maximal elements of the induced poset\.

Since every upward\-directed family is pairwise compatible, we haveγ⊔⊆γ⊔up\\gamma\_\{\\sqcup\}\\subseteq\\gamma^\{\\mathrm\{up\}\}\_\{\\sqcup\}\. Moreover, closure under upward\-directed unions implies that every refinement chain has an upper bound in the same class; hence Zorn’s lemma givesγ⊔up⊆γem\.\\gamma^\{\\mathrm\{up\}\}\_\{\\sqcup\}\\subseteq\\gamma\_\{\\mathrm\{em\}\}\.Therefore,

γ⊔⊆γ⊔up⊆γem\.\\displaystyle\\gamma\_\{\\sqcup\}\\subseteq\\gamma^\{\\mathrm\{up\}\}\_\{\\sqcup\}\\subseteq\\gamma\_\{\\mathrm\{em\}\}\.\(6\)
In addition to*refinement on ultrametrics*αuhr\\alpha\_\{\\mathrm\{uhr\}\}and to the*backbone\-hierarchy property*αbh\\alpha\_\{\\mathrm\{bh\}\}introduced earlier, we will use the following stronger backbone property:TTsatisfies the*unit\-backbone property*αbh1\\alpha\_\{\\mathrm\{bh\}\}^\{1\}ifTglob⊑TT\_\{\\mathrm\{glob\}\}\\sqsubseteq T\.

###### Proposition 38\(Closure principles\)\.

The closure properties indicated by checkmarks in theγ⊓\\gamma\_\{\\sqcap\},γ⊔up\\gamma^\{\\mathrm\{up\}\}\_\{\\sqcup\}, andγ⊔\\gamma\_\{\\sqcup\}columns of Table[1](https://arxiv.org/html/2609.11173#A1.T1)hold\. Moreover, these closure properties are preserved under finite conjunctions: ifα=⋀i=1kαi\\alpha=\\bigwedge\_\{i=1\}^\{k\}\\alpha\_\{i\}, then any of the three closure properties \(γ⊓\\gamma\_\{\\sqcap\},γ⊔up\\gamma^\{\\mathrm\{up\}\}\_\{\\sqcup\}, andγ⊔\\gamma\_\{\\sqcup\}\) shared by allαi\\alpha\_\{i\}is also satisfied byα\\alpha\.

###### Proof\.

We first establish the closure properties of the individual axioms, following the categories in Table[1](https://arxiv.org/html/2609.11173#A1.T1)\.

Throughout this proof,LLdenotes a nonempty family of hierarchical clustering methods\. Whenever a unionT⊔LT\_\{\\sqcup L\}is considered,LLis assumed to be either upward directed \(when studyingγ⊔up\\gamma^\{\\mathrm\{up\}\}\_\{\\sqcup\}\) or pairwise compatible \(when studyingγ⊔\\gamma\_\{\\sqcup\}\)\. In either case,T⊔LT\_\{\\sqcup L\}is hierarchy\-valued by Lemma[22](https://arxiv.org/html/2609.11173#Thmtheorem22), because every upward\-directed family is pairwise compatible\.

*Invariance properties\.*Scale invariance, order invariance, and permutation invariance are preserved under both intersections and unions\. For example, if everyT∈LT\\in Lis scale invariant, then

T⊓L​\(β​d\)=⋂T∈LT⁡\(β​d\)=⋂T∈LT⁡\(d\)=T⊓L​\(d\),T\_\{\\sqcap L\}\(\\beta d\)=\\bigcap\_\{T\\in L\}T\(\\beta d\)=\\bigcap\_\{T\\in L\}T\(d\)=T\_\{\\sqcap L\}\(d\),and the same calculation with unions givesT⊔L​\(β​d\)=T⊔L​\(d\)T\_\{\\sqcup L\}\(\\beta d\)=T\_\{\\sqcup L\}\(d\)\. The arguments for order and permutation invariance are identical\. Therefore,αsca,αord,αperm∈γ⊓∩γ⊔up∩γ⊔\.\\alpha\_\{\\mathrm\{sca\}\},\\alpha\_\{\\mathrm\{ord\}\},\\alpha\_\{\\mathrm\{perm\}\}\\in\\gamma\_\{\\sqcap\}\\cap\\gamma^\{\\mathrm\{up\}\}\_\{\\sqcup\}\\cap\\gamma\_\{\\sqcup\}\.

*Ultrametric properties\.*Refinement on ultrametrics is also preserved under intersections and compatible unions\. Indeed, if everyT∈LT\\in LsatisfiesΨu⊆T⁡\(u\)\\Psi\_\{u\}\\subseteq T\(u\), then

Ψu⊆⋂T∈LT⁡\(u\)⊆⋃T∈LT⁡\(u\)\.\\Psi\_\{u\}\\subseteq\\bigcap\_\{T\\in L\}T\(u\)\\subseteq\\bigcup\_\{T\\in L\}T\(u\)\.Similarly, if everyT∈LT\\in Lis exact on ultrametrics, thenT⊓L​\(u\)=T⊔L​\(u\)=Ψu\.T\_\{\\sqcap L\}\(u\)=T\_\{\\sqcup L\}\(u\)=\\Psi\_\{u\}\.Henceαuhr,αsuhr∈γ⊓∩γ⊔up∩γ⊔\.\\alpha\_\{\\mathrm\{uhr\}\},\\alpha\_\{\\mathrm\{suhr\}\}\\in\\gamma\_\{\\sqcap\}\\cap\\gamma^\{\\mathrm\{up\}\}\_\{\\sqcup\}\\cap\\gamma\_\{\\sqcup\}\.

*Richness properties\.*Cluster\-wise, partition, and hierarchical richness are preserved by every nonempty compatible union\. To see this, fix anyT0∈LT\_\{0\}\\in L\. Any dissimilarity realizing the relevant richness requirement forT0T\_\{0\}also realizes it forT⊔LT\_\{\\sqcup L\}, becauseT0​\(d\)⊆T⊔L​\(d\)T\_\{0\}\(d\)\\subseteq T\_\{\\sqcup L\}\(d\)\. Thereforeαcwr,αpr,αhr∈γ⊔⊆γ⊔up\.\\alpha\_\{\\mathrm\{cwr\}\},\\alpha\_\{\\mathrm\{pr\}\},\\alpha\_\{\\mathrm\{hr\}\}\\in\\gamma\_\{\\sqcup\}\\subseteq\\gamma^\{\\mathrm\{up\}\}\_\{\\sqcup\}\.The same argument does not apply to strict hierarchical richness, because additional clusters contributed by other methods inLLmay prevent the union from realizing a prescribed hierarchy exactly\.

*Consistency properties\.*We begin with cluster\-wise consistency, so we letLLbe a finite nonempty family of cluster\-wise consistent methods\.

Suppose thatC∈T⊓L​\(d\)C\\in T\_\{\\sqcap L\}\(d\)\. HenceC∈T⁡\(d\)C\\in T\(d\)for everyT∈LT\\in L, and thus, for every\{C\}\\\{C\\\}\-strengtheningd′d^\{\\prime\}ofdd, we haveC∈T⁡\(d′\)C\\in T\(d^\{\\prime\}\)for everyT∈LT\\in L\. This means thatC∈T⊓L​\(d′\)C\\in T\_\{\\sqcap L\}\(d^\{\\prime\}\), which provesαcwc∈γ⊓\\alpha\_\{\\mathrm\{cwc\}\}\\in\\gamma\_\{\\sqcap\}\.

Now suppose thatLLis pairwise compatible and letC∈T⊔L​\(d\)C\\in T\_\{\\sqcup L\}\(d\)\. ThenC∈T0​\(d\)C\\in T\_\{0\}\(d\)for someT0∈LT\_\{0\}\\in L\. Cluster\-wise consistency ofT0T\_\{0\}ensuresC∈T0​\(d′\)C\\in T\_\{0\}\(d^\{\\prime\}\)for any\{C\}\\\{C\\\}\-strengtheningd′d^\{\\prime\}ofdd, and henceC∈T0​\(d′\)⊆T⊔L​\(d′\)\.C\\in T\_\{0\}\(d^\{\\prime\}\)\\subseteq T\_\{\\sqcup L\}\(d^\{\\prime\}\)\.Thusαcwc∈γ⊓∩γ⊔up∩γ⊔\.\\alpha\_\{\\mathrm\{cwc\}\}\\in\\gamma\_\{\\sqcap\}\\cap\\gamma^\{\\mathrm\{up\}\}\_\{\\sqcup\}\\cap\\gamma\_\{\\sqcup\}\.Asαhc≡αcwc\\alpha\_\{\\mathrm\{hc\}\}\\equiv\\alpha\_\{\\mathrm\{cwc\}\}by Proposition[34](https://arxiv.org/html/2609.11173#Thmtheorem34), the same conclusions hold for hierarchical consistency\.

Partition consistency is likewise preserved under intersections\. Indeed, let𝒞\\mathcal\{C\}be a partition such that𝒞⊆T⊓L​\(d\)\\mathcal\{C\}\\subseteq T\_\{\\sqcap L\}\(d\)\. Then𝒞⊆T⁡\(d\)\\mathcal\{C\}\\subseteq T\(d\)for everyT∈LT\\in L, so every𝒞\\mathcal\{C\}\-strengtheningd′d^\{\\prime\}ofddsatisfies𝒞⊆T⁡\(d′\)\\mathcal\{C\}\\subseteq T\(d^\{\\prime\}\)for everyT∈LT\\in L, and therefore𝒞⊆T⊓L​\(d′\)\\mathcal\{C\}\\subseteq T\_\{\\sqcap L\}\(d^\{\\prime\}\)\. This provesαpc∈γ⊓\\alpha\_\{\\mathrm\{pc\}\}\\in\\gamma\_\{\\sqcap\}\.

To proveαpc∈γ⊔up\\alpha\_\{\\mathrm\{pc\}\}\\in\\gamma^\{\\mathrm\{up\}\}\_\{\\sqcup\}, an additional argument is needed because the different clusters of a partition may initially come from different methods\. LetLLbe upward directed and suppose𝒞=\{C1,…,Cm\}⊆T⊔L​\(d\)\.\\mathcal\{C\}=\\\{C\_\{1\},\\ldots,C\_\{m\}\\\}\\subseteq T\_\{\\sqcup L\}\(d\)\.For eachj∈\[m\]j\\in\[m\], chooseTj∈LT\_\{j\}\\in Lsuch thatCj∈Tj​\(d\)C\_\{j\}\\in T\_\{j\}\(d\)\. Since𝒞\\mathcal\{C\}is finite andLLis upward directed, there existsT⋆∈LT^\{\\star\}\\in Lrefining allT1,…,TmT\_\{1\},\\ldots,T\_\{m\}\. Hence𝒞⊆T⋆​\(d\)\.\\mathcal\{C\}\\subseteq T^\{\\star\}\(d\)\.Ifd′d^\{\\prime\}is a𝒞\\mathcal\{C\}\-strengthening ofdd, partition consistency ofT⋆T^\{\\star\}yields𝒞⊆T⋆​\(d′\)⊆T⊔L​\(d′\)\.\\mathcal\{C\}\\subseteq T^\{\\star\}\(d^\{\\prime\}\)\\subseteq T\_\{\\sqcup L\}\(d^\{\\prime\}\)\.Thusαpc∈γ⊔up\.\\alpha\_\{\\mathrm\{pc\}\}\\in\\gamma^\{\\mathrm\{up\}\}\_\{\\sqcup\}\.Finally,αabc≡αpc\\alpha\_\{\\mathrm\{abc\}\}\\equiv\\alpha\_\{\\mathrm\{pc\}\}, so absence consistency has the same closure properties\.

*Backbone property\.*LetL=\{T1,…,Tm\}L=\\\{T\_\{1\},\\ldots,T\_\{m\}\\\}be finite and suppose that eachTiT\_\{i\}satisfies the backbone property with margin sequence𝜼\(i\)\\boldsymbol\{\\eta\}^\{\(i\)\},i\.e\.,Tglob𝜼\(i\)⊑Ti\.T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}^\{\(i\)\}\}\\sqsubseteq T\_\{i\}\.Define the coordinate\-wise minimumηs:=mini∈\[m\]⁡ηs\(i\)\.\\eta\_\{s\}:=\\min\_\{i\\in\[m\]\}\\eta\_\{s\}^\{\(i\)\}\.BecauseLLis finite,𝜼\\boldsymbol\{\\eta\}is a separation margin sequence, andTglob𝜼⊑TiT\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}\\sqsubseteq T\_\{i\}for everyi∈\[m\]i\\in\[m\]\. ThereforeTglob𝜼⊑T⊓L,T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}\\sqsubseteq T\_\{\\sqcap L\},which provesαbh∈γ⊓\\alpha\_\{\\mathrm\{bh\}\}\\in\\gamma\_\{\\sqcap\}\. For any nonempty union, the backbone of an arbitrary fixed memberT0∈LT\_\{0\}\\in Lis contained inT⊔LT\_\{\\sqcup L\}; henceαbh∈γ⊓∩γ⊔up∩γ⊔\.\\alpha\_\{\\mathrm\{bh\}\}\\in\\gamma\_\{\\sqcap\}\\cap\\gamma^\{\\mathrm\{up\}\}\_\{\\sqcup\}\\cap\\gamma\_\{\\sqcup\}\.

This proves all positive closure claims for the individual axioms in the first three structural columns of Table[1](https://arxiv.org/html/2609.11173#A1.T1)\.

Finally, we prove closure under conjunction\. Letα=⋀i=1kαi\.\\alpha=\\bigwedge\_\{i=1\}^\{k\}\\alpha\_\{i\}\.Thenℋα​\(𝒳\)=⋂i=1kℋαi​\(𝒳\)\.\\mathcal\{H\}\_\{\\alpha\}\(\\mathcal\{X\}\)=\\bigcap\_\{i=1\}^\{k\}\\mathcal\{H\}\_\{\\alpha\_\{i\}\}\(\\mathcal\{X\}\)\.Suppose thatαi∈γ⊓\\alpha\_\{i\}\\in\\gamma\_\{\\sqcap\}for everyi∈\[k\]i\\in\[k\], and letL⊆ℋα​\(𝒳\)L\\subseteq\\mathcal\{H\}\_\{\\alpha\}\(\\mathcal\{X\}\)be finite and nonempty\. ThenL⊆ℋαi​\(𝒳\)L\\subseteq\\mathcal\{H\}\_\{\\alpha\_\{i\}\}\(\\mathcal\{X\}\)for everyii\. Moreover, because eachαi\\alpha\_\{i\}satisfiesγ⊓\\gamma\_\{\\sqcap\}, we haveT⊓L∈ℋαi​\(𝒳\)T\_\{\\sqcap L\}\\in\\mathcal\{H\}\_\{\\alpha\_\{i\}\}\(\\mathcal\{X\}\)for everyi∈\[k\]i\\in\[k\]\. Therefore,T⊓L∈⋂i=1kℋαi​\(𝒳\)=ℋα​\(𝒳\),T\_\{\\sqcap L\}\\in\\bigcap\_\{i=1\}^\{k\}\\mathcal\{H\}\_\{\\alpha\_\{i\}\}\(\\mathcal\{X\}\)=\\mathcal\{H\}\_\{\\alpha\}\(\\mathcal\{X\}\),proving thatα∈γ⊓\\alpha\\in\\gamma\_\{\\sqcap\}\. The arguments forγ⊔up\\gamma^\{\\mathrm\{up\}\}\_\{\\sqcup\}andγ⊔\\gamma\_\{\\sqcup\}are identical, replacingT⊓LT\_\{\\sqcap L\}byT⊔LT\_\{\\sqcup L\}\. ∎

###### Corollary 39\(Closure of the admissible classes\)\.

The canonical admissible class satisfiesαadm∈γ⊓∩γ⊔up\.\\alpha\_\{\\mathrm\{adm\}\}\\in\\gamma\_\{\\sqcap\}\\cap\\gamma^\{\\mathrm\{up\}\}\_\{\\sqcup\}\.Moreover,ℋαadm\+​\(𝒳\)\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\+\}\}\}\(\\mathcal\{X\}\)is closed under arbitrary nonempty intersections and under nonempty upward\-directed unions\.

###### Proof\.

By Corollary[36](https://arxiv.org/html/2609.11173#Thmtheorem36),αadm≡αsca∧αbh∧αpc∧αperm\.\\alpha\_\{\\mathrm\{adm\}\}\\equiv\\alpha\_\{\\mathrm\{sca\}\}\\wedge\\alpha\_\{\\mathrm\{bh\}\}\\wedge\\alpha\_\{\\mathrm\{pc\}\}\\wedge\\alpha\_\{\\mathrm\{perm\}\}\.Each of these four requirements belongs toγ⊓\\gamma\_\{\\sqcap\}, so Proposition[38](https://arxiv.org/html/2609.11173#Thmtheorem38)givesαadm∈γ⊓\\alpha\_\{\\mathrm\{adm\}\}\\in\\gamma\_\{\\sqcap\}\. Directed\-union closure follows directly fromαadm=αsca∧αpr∧αpc∧αperm,\\alpha\_\{\\mathrm\{adm\}\}=\\alpha\_\{\\mathrm\{sca\}\}\\wedge\\alpha\_\{\\mathrm\{pr\}\}\\wedge\\alpha\_\{\\mathrm\{pc\}\}\\wedge\\alpha\_\{\\mathrm\{perm\}\},as each of these requirements belongs toγ⊔up\\gamma^\{\\mathrm\{up\}\}\_\{\\sqcup\}\.

For strong admissibility, recall thatαadm\+≡αsca∧αsuhr∧αpc∧αperm\.\\alpha\_\{\\mathrm\{adm\+\}\}\\equiv\\alpha\_\{\\mathrm\{sca\}\}\\wedge\\alpha\_\{\\mathrm\{suhr\}\}\\wedge\\alpha\_\{\\mathrm\{pc\}\}\\wedge\\alpha\_\{\\mathrm\{perm\}\}\.Scale invariance, exactness on ultrametrics, partition consistency, and permutation invariance are all preserved under arbitrary nonempty intersections\. Exactness on ultrametrics also guarantees partition richness, so the resulting method remains strongly admissible\. Directed\-union closure follows from Proposition[38](https://arxiv.org/html/2609.11173#Thmtheorem38)\. ∎

Proposition[38](https://arxiv.org/html/2609.11173#Thmtheorem38)gives a compact way to transfer the order\-theoretic arguments of Section[4](https://arxiv.org/html/2609.11173#S4)to alternative axiom systems\. The next result records the consequences that are most relevant for comparison with the canonical framework\.

###### Proposition 40\(Robustness under alternative axioms\)\.

Letα\\alphabe any conjunction of the axioms introduced in Section[A\.1](https://arxiv.org/html/2609.11173#A1.SS1)\. Then:

1. \(i\)*Achievability\.*The setℋα​\(𝒳\)\\mathcal\{H\}\_\{\\alpha\}\(\\mathcal\{X\}\)is nonempty\. In particular,Tglob,Tloc,TSL∈ℋα​\(𝒳\)T\_\{\\mathrm\{glob\}\},T\_\{\\mathrm\{loc\}\},T\_\{\\mathrm\{SL\}\}\\in\\mathcal\{H\}\_\{\\alpha\}\(\\mathcal\{X\}\)\.
2. \(ii\)*Maximal extensions\.*If strict hierarchical richness is not among the requirements definingα\\alphaor if exactness on ultrametrics is required, then everyT∈ℋα​\(𝒳\)T\\in\\mathcal\{H\}\_\{\\alpha\}\(\\mathcal\{X\}\)is refined by a maximal member ofℋα​\(𝒳\)\\mathcal\{H\}\_\{\\alpha\}\(\\mathcal\{X\}\)\.
3. \(iii\)*Finite intersections\.*Ifα⊳αbh\\alpha\\triangleright\\alpha\_\{\\mathrm\{bh\}\}and either strict hierarchical richness is not required or exactness on ultrametrics is required, thenℋα​\(𝒳\)\\mathcal\{H\}\_\{\\alpha\}\(\\mathcal\{X\}\)is closed under finite nonempty intersections\.
4. \(iv\)*Diversity versus order invariance\.*If order invariance is not among the requirements definingα\\alpha, thenℋα​\(𝒳\)\\mathcal\{H\}\_\{\\alpha\}\(\\mathcal\{X\}\)has uncountable height, width, and cellularity\. If order invariance is required, thenℋα​\(𝒳\)\\mathcal\{H\}\_\{\\alpha\}\(\\mathcal\{X\}\)is finite\.
5. \(v\)*Least elements\.*Letαbh1\\alpha\_\{\\mathrm\{bh\}\}^\{1\}denote the unit\-backbone propertyTglob⊑TT\_\{\\mathrm\{glob\}\}\\sqsubseteq T\. We have αuhr∧αpc⊳αbh1andαord∧αpc∧αpr⊳αbh1\.\\alpha\_\{\\mathrm\{uhr\}\}\\wedge\\alpha\_\{\\mathrm\{pc\}\}\\triangleright\\alpha\_\{\\mathrm\{bh\}\}^\{1\}\\quad\\text\{ and \}\\quad\\alpha\_\{\\mathrm\{ord\}\}\\wedge\\alpha\_\{\\mathrm\{pc\}\}\\wedge\\alpha\_\{\\mathrm\{pr\}\}\\triangleright\\alpha\_\{\\mathrm\{bh\}\}^\{1\}\.Moreover, ifα⊳αbh1\\alpha\\triangleright\\alpha\_\{\\mathrm\{bh\}\}^\{1\}, thenTglobT\_\{\\mathrm\{glob\}\}is the least element of the poset\(ℋα​\(𝒳\),⊑\)\(\\mathcal\{H\}\_\{\\alpha\}\(\\mathcal\{X\}\),\\sqsubseteq\)\.

###### Proof\.

For \(i\), we will show in later sections that the three methodsTglobT\_\{\\mathrm\{glob\}\},TlocT\_\{\\mathrm\{loc\}\}, andTSLT\_\{\\mathrm\{SL\}\}are order invariant, cluster\-wise consistent, exact on ultrametrics, and permutation invariant\. Exactness on ultrametrics implies all three richness requirements as well as refinement on ultrametrics, and cluster\-wise consistency implies partition consistency\. Thus these methods satisfy every requirement introduced in this paper\.

For \(ii\), with the exception of strict hierarchical richness, every other requirement is preserved by directed unions according to Proposition[38](https://arxiv.org/html/2609.11173#Thmtheorem38)\. If exactness on ultrametrics is required, strict hierarchical richness is redundant\. Hence,α\\alphasatisfiesγ⊔up\\gamma^\{\\mathrm\{up\}\}\_\{\\sqcup\}; thus every nonempty chain has an upper bound inℋα​\(𝒳\)\\mathcal\{H\}\_\{\\alpha\}\(\\mathcal\{X\}\), and Zorn’s lemma gives a maximal refinement above everyT∈ℋα​\(𝒳\)T\\in\\mathcal\{H\}\_\{\\alpha\}\(\\mathcal\{X\}\)\(see the relationship \([6](https://arxiv.org/html/2609.11173#A1.E6)\)\)\.

For \(iii\), all requirements other than the richness variants are preserved by finite intersections\. The assumptionα⊳αbh\\alpha\\triangleright\\alpha\_\{\\mathrm\{bh\}\}ensures the existence of a common backbone to every finite intersection, andαbh⊳αhr⊳αpr⊳αcwr\\alpha\_\{\\mathrm\{bh\}\}\\triangleright\\alpha\_\{\\mathrm\{hr\}\}\\triangleright\\alpha\_\{\\mathrm\{pr\}\}\\triangleright\\alpha\_\{\\mathrm\{cwr\}\}restores every non\-strict richness requirement\. If strict hierarchical richness is required together with exactness on ultrametrics, exactness is preserved by intersections and implies strict hierarchical richness\.

For \(iv\), suppose first that order invariance is not required\. By Lemma[54](https://arxiv.org/html/2609.11173#Thmtheorem54)TstableT\_\{\\mathrm\{stable\}\}satisfies all the requirements except possibly order invariance, and because power transformations preserve these requirements by Corollary[29](https://arxiv.org/html/2609.11173#Thmtheorem29), the methods\{Tstable\(p\):p\>0\}\\\{T\_\{\\mathrm\{stable\}\}^\{\(p\)\}:p\>0\\\}also satisfy them\. Lemma[16](https://arxiv.org/html/2609.11173#Thmtheorem16)shows that they are pairwise incompatible, yielding uncountable width and cellularity\. To establish the uncountable height, fora∈\(0,1\]a\\in\(0,1\], define a margin sequence by𝜼\(a\)=a⋅𝟏\\boldsymbol\{\\eta\}^\{\(a\)\}=a\\cdot\\boldsymbol\{1\}and set

Sa​\(d\):=Tglob​\(d\)∪Tloc𝜼\(a\)​\(d\)\.S\_\{a\}\(d\):=T\_\{\\mathrm\{glob\}\}\(d\)\\cup T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}^\{\(a\)\}\}\(d\)\.BothTglob​\(d\)T\_\{\\mathrm\{glob\}\}\(d\)andTloc𝜼\(a\)​\(d\)T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}^\{\(a\)\}\}\(d\)are subhierarchies ofTloc​\(d\)T\_\{\\mathrm\{loc\}\}\(d\), so the union is a hierarchy\. By Lemma[49](https://arxiv.org/html/2609.11173#Thmtheorem49)and Proposition[38](https://arxiv.org/html/2609.11173#Thmtheorem38),SaS\_\{a\}is scale invariant, cluster\-wise consistent, and permutation invariant\. Moreover, for every ultrametricuu,

Ψu=Tglob​\(u\)⊆Sa​\(u\)⊆Tloc​\(u\)=Ψu,\\Psi\_\{u\}=T\_\{\\mathrm\{glob\}\}\(u\)\\subseteq S\_\{a\}\(u\)\\subseteq T\_\{\\mathrm\{loc\}\}\(u\)=\\Psi\_\{u\},soSaS\_\{a\}is exact on ultrametrics and therefore satisfies every requirement except possibly order invariance\.

The family\{Sa:a∈\(0,1\]\}\\\{S\_\{a\}:a\\in\(0,1\]\\\}is an increasing chain\. It is strict: for0<a<b≤10<a<b\\leq 1, chooser∈\[a,b\)r\\in\[a,b\),L\>1/rL\>1/r, and distinctx1,x2,x3,z∈𝒳x\_\{1\},x\_\{2\},x\_\{3\},z\\in\\mathcal\{X\}\. WithC=\{x1,x2,x3\}C=\\\{x\_\{1\},x\_\{2\},x\_\{3\}\\\}, set

d⁡\(x1,x2\)=d⁡\(x1,x3\)=r,d⁡\(x2,x3\)=r​L,d⁡\(x1,z\)=1,d⁡\(x2,z\)=d⁡\(x3,z\)=L,d\(x\_\{1\},x\_\{2\}\)=d\(x\_\{1\},x\_\{3\}\)=r,\\quad d\(x\_\{2\},x\_\{3\}\)=rL,\\quad d\(x\_\{1\},z\)=1,\\quad d\(x\_\{2\},z\)=d\(x\_\{3\},z\)=L,and, for every remainingw∉C∪\{z\}w\\notin C\\cup\\\{z\\\}, setd⁡\(x1,w\)=1d\(x\_\{1\},w\)=1andd⁡\(x2,w\)=d⁡\(x3,w\)=Ld\(x\_\{2\},w\)=d\(x\_\{3\},w\)=L\. Complete the remaining dissimilarities arbitrarily\. Thenτ⁡\(d,C\)=r\\tau\(d,C\)=randϱ⁡\(d,C\)=r​L\>1\\varrho\(d,C\)=rL\>1, soC∉Sa​\(d\)C\\notin S\_\{a\}\(d\)butC∈Sb​\(d\)C\\in S\_\{b\}\(d\)\.

Conversely, suppose order invariance is required\. We say that two dissimilaritiesd,d′∈𝒟⁡\(𝒳\)d,d^\{\\prime\}\\in\\mathcal\{D\}\(\\mathcal\{X\}\)are order\-equivalent, writtend∼d′d\\sim d^\{\\prime\}, if they induce the same weak ordering of the\(n2\)\\binom\{n\}\{2\}unordered pairs, that is

d⁡\(x1,y1\)≤d⁡\(x2,y2\)⇔d′​\(x1,y1\)≤d′​\(x2,y2\)for all​\(x1,y1\),\(x2,y2\)∈𝒳\.d\(x\_\{1\},y\_\{1\}\)\\ \\leq\\ d\(x\_\{2\},y\_\{2\}\)\\iff d^\{\\prime\}\(x\_\{1\},y\_\{1\}\)\\ \\leq\\ d^\{\\prime\}\(x\_\{2\},y\_\{2\}\)\\quad\\text\{ for all \}\(x\_\{1\},y\_\{1\}\),\(x\_\{2\},y\_\{2\}\)\\in\\mathcal\{X\}\.Each equivalence class corresponds to a weak ordering of the\(n2\)\\binom\{n\}\{2\}pairs, and there are only finitely many such weak orderings, so the quotient space𝒟\(𝒳\)/∼\\mathcal\{D\}\(\\mathcal\{X\}\)/\{\\sim\}is finite\. Ifd∼d′d\\sim d^\{\\prime\}, a strictly increasing interpolation of the finitely many distinct values taken byddandd′d^\{\\prime\}givesd′=g∘dd^\{\\prime\}=g\\circ d; hence an order invariance hierarchical method is constant on each equivalence class\. Because𝒯⁡\(𝒳\)\\mathcal\{T\}\(\\mathcal\{X\}\)is finite,\|ℋαord\(𝒳\)\|≤\|𝒯\(𝒳\)\|\|𝒟\(𝒳\)/∼\|<∞,\|\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{ord\}\}\}\(\\mathcal\{X\}\)\|\\ \\leq\\ \|\\mathcal\{T\}\(\\mathcal\{X\}\)\|^\{\\,\|\\mathcal\{D\}\(\\mathcal\{X\}\)/\{\\sim\}\|\}\\ <\\ \\infty,and therefore every subclassℋα​\(𝒳\)\\mathcal\{H\}\_\{\\alpha\}\(\\mathcal\{X\}\)satisfying order invariance is finite\.

For \(v\), the first claimαuhr∧αpc⊳αbh1\\alpha\_\{\\mathrm\{uhr\}\}\\wedge\\alpha\_\{\\mathrm\{pc\}\}\\triangleright\\alpha\_\{\\mathrm\{bh\}\}^\{1\}is established in Lemma[58](https://arxiv.org/html/2609.11173#Thmtheorem58), whereas the second claim is established in Lemma[59](https://arxiv.org/html/2609.11173#Thmtheorem59)\. Finally, ifα⊳αbh1\\alpha\\triangleright\\alpha\_\{\\mathrm\{bh\}\}^\{1\}, then everyT∈ℋα​\(𝒳\)T\\in\\mathcal\{H\}\_\{\\alpha\}\(\\mathcal\{X\}\)refinesTglobT\_\{\\mathrm\{glob\}\}, whileTglob∈ℋα​\(𝒳\)T\_\{\\mathrm\{glob\}\}\\in\\mathcal\{H\}\_\{\\alpha\}\(\\mathcal\{X\}\)by \(i\)\. ThereforeTglobT\_\{\\mathrm\{glob\}\}is the least element ofℋα​\(𝒳\)\\mathcal\{H\}\_\{\\alpha\}\(\\mathcal\{X\}\)\. ∎

## Appendix BRelational Properties and Axiom\-Preserving Transformations

Section[5\.2](https://arxiv.org/html/2609.11173#S5.SS2)introduced transformations of dissimilarities and the notion of preservation of an axiom under preprocessing\. In this appendix, we formalize a common mechanism behind several of the preservation results stated there\. Indeed, many of the axioms considered in this paper compare the outputs of a method on two inputs that are related in a prescribed way\. Scale invariance, for instance, comparesT⁡\(d\)T\(d\)andT⁡\(β​d\)T\(\\beta d\), while partition consistency comparesT⁡\(d\)T\(d\)andT⁡\(d′\)T\(d^\{\\prime\}\)whend′d^\{\\prime\}is a strengthening ofdd\. Relational properties provide a common language for such requirements\.

### B\.1Relational Properties

###### Definition 41\(Relational property\)\.

A*relational property*α\\alphais specified by

- •an*input relation*Rα⊆𝒟⁡\(𝒳\)×𝒟⁡\(𝒳\)R^\{\\alpha\}\\subseteq\\mathcal\{D\}\(\\mathcal\{X\}\)\\times\\mathcal\{D\}\(\\mathcal\{X\}\);
- •an*output relation*Sα⊆𝒯⁡\(𝒳\)×𝒯⁡\(𝒳\)S^\{\\alpha\}\\subseteq\\mathcal\{T\}\(\\mathcal\{X\}\)\\times\\mathcal\{T\}\(\\mathcal\{X\}\)\.

A hierarchical clustering methodT∈ℋ⁡\(𝒳\)T\\in\\mathcal\{H\}\(\\mathcal\{X\}\)satisfies the relational propertyα\\alphaif

\(d,d′\)∈Rα⟹\(T⁡\(d\),T⁡\(d′\)\)∈Sα\.\(d,d^\{\\prime\}\)\\in R^\{\\alpha\}\\quad\\Longrightarrow\\quad\\left\(T\(d\),T\(d^\{\\prime\}\)\\right\)\\in S^\{\\alpha\}\.

More generally, we writeα∈γrp\\alpha\\in\\gamma\_\{\\mathrm\{rp\}\}if the requirementα\\alphacan be expressed as a relational property or as a conjunction of relational properties\. This is the meaning of theγrp\\gamma\_\{\\mathrm\{rp\}\}column in Table[1](https://arxiv.org/html/2609.11173#A1.T1)\.

The usefulness of this formulation for preprocessing comes from the following simple principle: to preserve a relational property, it is sufficient for the transformation to preserve its input relation\.

###### Proposition 42\(Relational preservation principle\)\.

Letα\\alphabe a relational property with input relationRαR^\{\\alpha\}, and letμ:𝒟⁡\(𝒳\)→𝒟⁡\(𝒳\)\\mu:\\mathcal\{D\}\(\\mathcal\{X\}\)\\to\\mathcal\{D\}\(\\mathcal\{X\}\)be a transformation\. If

\(d,d′\)∈Rα⟹\(μ⁡\(d\),μ⁡\(d′\)\)∈Rα,\(d,d^\{\\prime\}\)\\in R^\{\\alpha\}\\quad\\Longrightarrow\\quad\\bigl\(\\mu\(d\),\\mu\(d^\{\\prime\}\)\\bigr\)\\in R^\{\\alpha\},thenμ\\mupreservesα\\alpha\.

###### Proof\.

LetT∈ℋα​\(𝒳\)T\\in\\mathcal\{H\}\_\{\\alpha\}\(\\mathcal\{X\}\)\. If\(d,d′\)∈Rα\(d,d^\{\\prime\}\)\\in R^\{\\alpha\}, the assumption gives\(μ⁡\(d\),μ⁡\(d′\)\)∈Rα\\bigl\(\\mu\(d\),\\mu\(d^\{\\prime\}\)\\bigr\)\\in R^\{\\alpha\}\. BecauseTTsatisfiesα\\alpha, we have\(T⁡\(μ⁡\(d\)\),T⁡\(μ⁡\(d′\)\)\)∈Sα\\bigl\(T\(\\mu\(d\)\),T\(\\mu\(d^\{\\prime\}\)\)\\bigr\)\\in S^\{\\alpha\}\. ThusT∘μT\\circ\\musatisfiesα\\alpha, and thereforeμ\\mupreservesα\\alpha\. ∎

We will also use the following elementary observation when an axiom is represented as a conjunction of relational properties\.

###### Lemma 43\(Preservation under conjunction\)\.

Let\(αj\)j∈J\(\\alpha\_\{j\}\)\_\{j\\in J\}be a family of requirements\. If a transformationμ\\mupreservesαj\\alpha\_\{j\}for everyj∈Jj\\in J, then it preserves the conjunction⋀j∈Jαj\.\\bigwedge\_\{j\\in J\}\\alpha\_\{j\}\.

###### Proof\.

LetTTsatisfy⋀j∈Jαj\\bigwedge\_\{j\\in J\}\\alpha\_\{j\}\. ThenTTsatisfies everyαj\\alpha\_\{j\}\. Becauseμ\\mupreserves eachαj\\alpha\_\{j\}, the compositionT∘μT\\circ\\mualso satisfies everyαj\\alpha\_\{j\}, and hence satisfies their conjunction\. ∎

Consequently, whenα∈γrp\\alpha\\in\\gamma\_\{\\mathrm\{rp\}\}, preservation ofα\\alphacan be verified by checking thatμ\\mupreserves the input relation of each relational property appearing in its representation\. Notice that partition richness is handled differently in Lemma[28](https://arxiv.org/html/2609.11173#Thmtheorem28): its preservation follows from surjectivity of the transformation rather than from the relational principle above\.

### B\.2Relational Formulations of the Axioms

We now give relational representations of the axioms marked by a checkmark in theγrp\\gamma\_\{\\mathrm\{rp\}\}column of Table[1](https://arxiv.org/html/2609.11173#A1.T1)\. These representations also make explicit which structure of the input must be preserved by a preprocessing transformation\.

We denote the diagonal relation on hierarchies by

Δ𝒯:=\{\(Ψ,Ψ\):Ψ∈𝒯⁡\(𝒳\)\}\.\\Delta\_\{\\mathcal\{T\}\}:=\\\{\(\\Psi,\\Psi\):\\Psi\\in\\mathcal\{T\}\(\\mathcal\{X\}\)\\\}\.
*Scale and order invariance\.*Scale invariance is the relational property defined by

Rαsca=\{\(d,βd\):d∈𝒟\(𝒳\),β\>0\},Sαsca=Δ𝒯\.R^\{\\alpha\_\{\\mathrm\{sca\}\}\}=\\\{\(d,\\beta d\):d\\in\\mathcal\{D\}\(\\mathcal\{X\}\),\\ \\beta\>0\\\},\\qquad S^\{\\alpha\_\{\\mathrm\{sca\}\}\}=\\Delta\_\{\\mathcal\{T\}\}\.\(7\)Indeed, the corresponding implication is exactlyT⁡\(β​d\)=T⁡\(d\)T\(\\beta d\)=T\(d\)\.

Likewise, order invariance is relational\. Let

𝒢↑:=\{g:R≥0→R≥0:g\(0\)=0,gstrictly increasing\}\.\\mathcal\{G\}\_\{\\uparrow\}:=\\left\\\{g:\\mathbb\{R\}\_\{\\geq 0\}\\to\\mathbb\{R\}\_\{\\geq 0\}:g\(0\)=0,\\ g\\text\{ strictly increasing\}\\right\\\}\.Then order invariance is obtained from

Rαord=\{\(d,g∘d\):d∈𝒟\(𝒳\),g∈𝒢↑\},Sαord=Δ𝒯\.R^\{\\alpha\_\{\\mathrm\{ord\}\}\}=\\\{\(d,g\\circ d\):d\\in\\mathcal\{D\}\(\\mathcal\{X\}\),\\ g\\in\\mathcal\{G\}\_\{\\uparrow\}\\\},\\qquad S^\{\\alpha\_\{\\mathrm\{ord\}\}\}=\\Delta\_\{\\mathcal\{T\}\}\.\(8\)
*Permutation invariance\.*Here the required output relation depends on the permutation\. For each permutationϕ∈Π⁡\(𝒳\)\\phi\\in\\Pi\(\\mathcal\{X\}\), define the relational propertyαpermϕ\\alpha\_\{\\mathrm\{perm\}\}^\{\\phi\}by

Rαpermϕ=\{\(d,dϕ\):d∈𝒟\(𝒳\)\},Sαpermϕ=\{\(Ψ,ϕ⋅Ψ\):Ψ∈𝒯\(𝒳\)\}\.\\displaystyle R^\{\\alpha\_\{\\mathrm\{perm\}\}^\{\\phi\}\}=\\\{\(d,d\_\{\\phi\}\):d\\in\\mathcal\{D\}\(\\mathcal\{X\}\)\\\},\\qquad S^\{\\alpha\_\{\\mathrm\{perm\}\}^\{\\phi\}\}=\\\{\(\\Psi,\\phi\\cdot\\Psi\):\\Psi\\in\\mathcal\{T\}\(\\mathcal\{X\}\)\\\}\.\(9\)Thenαperm≡⋀ϕ∈Π⁡\(𝒳\)αpermϕ\.\\alpha\_\{\\mathrm\{perm\}\}\\equiv\\bigwedge\_\{\\phi\\in\\Pi\(\\mathcal\{X\}\)\}\\alpha\_\{\\mathrm\{perm\}\}^\{\\phi\}\.Indeed, satisfying everyαpermϕ\\alpha\_\{\\mathrm\{perm\}\}^\{\\phi\}is precisely the requirement

T⁡\(dϕ\)=ϕ⋅T⁡\(d\)for every​d∈𝒟⁡\(𝒳\),ϕ∈Π⁡\(𝒳\)\.T\(d\_\{\\phi\}\)=\\phi\\cdot T\(d\)\\qquad\\text\{for every \}d\\in\\mathcal\{D\}\(\\mathcal\{X\}\),\\ \\phi\\in\\Pi\(\\mathcal\{X\}\)\.
*Consistency\.*Fix a partition𝒞∈𝒫⁡\(𝒳\)\\mathcal\{C\}\\in\\mathcal\{P\}\(\\mathcal\{X\}\)\. Defineαpc𝒞\\alpha\_\{\\mathrm\{pc\}\}^\{\\mathcal\{C\}\}by

Rαpc𝒞\\displaystyle R^\{\\alpha\_\{\\mathrm\{pc\}\}^\{\\mathcal\{C\}\}\}=\{\(d,d′\)∈𝒟​\(𝒳\)2:d′​is a​𝒞​\-strengthening of​d\},\\displaystyle=\\left\\\{\(d,d^\{\\prime\}\)\\in\\mathcal\{D\}\(\\mathcal\{X\}\)^\{2\}:d^\{\\prime\}\\text\{ is a \}\\mathcal\{C\}\\text\{\-strengthening of \}d\\right\\\},\(10\)Sαpc𝒞\\displaystyle S^\{\\alpha\_\{\\mathrm\{pc\}\}^\{\\mathcal\{C\}\}\}=\{\(Ψ,Ψ′\)∈𝒯​\(𝒳\)2:𝒞⊆Ψ⟹𝒞⊆Ψ′\}\.\\displaystyle=\\left\\\{\(\\Psi,\\Psi^\{\\prime\}\)\\in\\mathcal\{T\}\(\\mathcal\{X\}\)^\{2\}:\\mathcal\{C\}\\subseteq\\Psi\\Longrightarrow\\mathcal\{C\}\\subseteq\\Psi^\{\\prime\}\\right\\\}\.Partition consistency is thereforeαpc≡⋀𝒞∈𝒫⁡\(𝒳\)αpc𝒞\.\\alpha\_\{\\mathrm\{pc\}\}\\equiv\\bigwedge\_\{\\mathcal\{C\}\\in\\mathcal\{P\}\(\\mathcal\{X\}\)\}\\alpha\_\{\\mathrm\{pc\}\}^\{\\mathcal\{C\}\}\.

Cluster\-wise consistency admits the analog representation\. For every nonemptyC⊆𝒳C\\subseteq\\mathcal\{X\}, letαcwcC\\alpha\_\{\\mathrm\{cwc\}\}^\{C\}be defined by

RαcwcC\\displaystyle R^\{\\alpha\_\{\\mathrm\{cwc\}\}^\{C\}\}=\{\(d,d′\)∈𝒟​\(𝒳\)2:d′​is a​\{C\}​\-strengthening of​d\},\\displaystyle=\\left\\\{\(d,d^\{\\prime\}\)\\in\\mathcal\{D\}\(\\mathcal\{X\}\)^\{2\}:d^\{\\prime\}\\text\{ is a \}\\\{C\\\}\\text\{\-strengthening of \}d\\right\\\},\(11\)SαcwcC\\displaystyle S^\{\\alpha\_\{\\mathrm\{cwc\}\}^\{C\}\}=\{\(Ψ,Ψ′\)∈𝒯​\(𝒳\)2:C∈Ψ⟹C∈Ψ′\}\.\\displaystyle=\\left\\\{\(\\Psi,\\Psi^\{\\prime\}\)\\in\\mathcal\{T\}\(\\mathcal\{X\}\)^\{2\}:C\\in\\Psi\\Longrightarrow C\\in\\Psi^\{\\prime\}\\right\\\}\.Henceαcwc≡⋀∅≠C⊆𝒳αcwcC\.\\alpha\_\{\\mathrm\{cwc\}\}\\equiv\\bigwedge\_\{\\emptyset\\neq C\\subseteq\\mathcal\{X\}\}\\alpha\_\{\\mathrm\{cwc\}\}^\{C\}\.By Proposition[34](https://arxiv.org/html/2609.11173#Thmtheorem34),αhc≡αcwc\\alpha\_\{\\mathrm\{hc\}\}\\equiv\\alpha\_\{\\mathrm\{cwc\}\}andαabc≡αpc\\alpha\_\{\\mathrm\{abc\}\}\\equiv\\alpha\_\{\\mathrm\{pc\}\}, so hierarchical consistency and absence consistency are also representable as conjunctions of relational properties\.

*Ultrametric properties\.*For a hierarchyΨ∈𝒯⁡\(𝒳\)\\Psi\\in\\mathcal\{T\}\(\\mathcal\{X\}\), let𝒰Ψ:=\{u∈𝒰⁡\(𝒳\):Ψu=Ψ\}\\mathcal\{U\}\_\{\\Psi\}:=\\\{u\\in\\mathcal\{U\}\(\\mathcal\{X\}\):\\Psi\_\{u\}=\\Psi\\\}be the set of ultrametrics inducingΨ\\Psi\. To represent refinement on ultrametrics, defineαuhrΨ\\alpha\_\{\\mathrm\{uhr\}\}^\{\\Psi\}by

RαuhrΨ=\{\(u,u\):u∈𝒰Ψ\},SαuhrΨ=\{\(Ψ′,Ψ′\):Ψ′∈𝒯\(𝒳\),Ψ⊆Ψ′\}\.R^\{\\alpha\_\{\\mathrm\{uhr\}\}^\{\\Psi\}\}=\\\{\(u,u\):u\\in\\mathcal\{U\}\_\{\\Psi\}\\\},\\qquad S^\{\\alpha\_\{\\mathrm\{uhr\}\}^\{\\Psi\}\}=\\\{\(\\Psi^\{\\prime\},\\Psi^\{\\prime\}\):\\Psi^\{\\prime\}\\in\\mathcal\{T\}\(\\mathcal\{X\}\),\\ \\Psi\\subseteq\\Psi^\{\\prime\}\\\}\.\(12\)Thenαuhr≡⋀Ψ∈𝒯⁡\(𝒳\)αuhrΨ\.\\alpha\_\{\\mathrm\{uhr\}\}\\equiv\\bigwedge\_\{\\Psi\\in\\mathcal\{T\}\(\\mathcal\{X\}\)\}\\alpha\_\{\\mathrm\{uhr\}\}^\{\\Psi\}\.

Exactness on ultrametrics is obtained by strengthening the output relation\. For eachΨ∈𝒯⁡\(𝒳\)\\Psi\\in\\mathcal\{T\}\(\\mathcal\{X\}\), defineαsuhrΨ\\alpha\_\{\\mathrm\{suhr\}\}^\{\\Psi\}by

RαsuhrΨ=\{\(u,u\):u∈𝒰Ψ\},SαsuhrΨ=\{\(Ψ,Ψ\)\}\.R^\{\\alpha\_\{\\mathrm\{suhr\}\}^\{\\Psi\}\}=\\\{\(u,u\):u\\in\\mathcal\{U\}\_\{\\Psi\}\\\},\\qquad S^\{\\alpha\_\{\\mathrm\{suhr\}\}^\{\\Psi\}\}=\\\{\(\\Psi,\\Psi\)\\\}\.\(13\)Thusαsuhr≡⋀Ψ∈𝒯⁡\(𝒳\)αsuhrΨ\.\\alpha\_\{\\mathrm\{suhr\}\}\\equiv\\bigwedge\_\{\\Psi\\in\\mathcal\{T\}\(\\mathcal\{X\}\)\}\\alpha\_\{\\mathrm\{suhr\}\}^\{\\Psi\}\.

These representations establish the corresponding entries in theγrp\\gamma\_\{\\mathrm\{rp\}\}column of Table[1](https://arxiv.org/html/2609.11173#A1.T1)\. In particular, asαadm\+≡αsca∧αsuhr∧αpc∧αperm,\\alpha\_\{\\mathrm\{adm\+\}\}\\equiv\\alpha\_\{\\mathrm\{sca\}\}\\wedge\\alpha\_\{\\mathrm\{suhr\}\}\\wedge\\alpha\_\{\\mathrm\{pc\}\}\\wedge\\alpha\_\{\\mathrm\{perm\}\},αadm\+\\alpha\_\{\\mathrm\{adm\+\}\}also belongs toγrp\\gamma\_\{\\mathrm\{rp\}\}\.

### B\.3Applications to Axiom\-Preserving Transformations

###### Proof of Lemma[28](https://arxiv.org/html/2609.11173#Thmtheorem28)\.

For scale invariance, let\(d,β​d\)∈Rαsca\(d,\\beta d\)\\in R^\{\\alpha\_\{\\mathrm\{sca\}\}\}\. By assumption, there existsβ′\>0\\beta^\{\\prime\}\>0such thatμ⁡\(β​d\)=β′​μ​\(d\)\.\\mu\(\\beta d\)=\\beta^\{\\prime\}\\mu\(d\)\.Hence\(μ⁡\(d\),μ⁡\(β​d\)\)=\(μ⁡\(d\),β′​μ​\(d\)\)∈Rαsca\\bigl\(\\mu\(d\),\\mu\(\\beta d\)\\bigr\)=\\bigl\(\\mu\(d\),\\beta^\{\\prime\}\\mu\(d\)\\bigr\)\\in R^\{\\alpha\_\{\\mathrm\{sca\}\}\}, and Proposition[42](https://arxiv.org/html/2609.11173#Thmtheorem42)ensures thatμ\\mupreserves scale invariance\.

Partition richness is the one property in the lemma that we verify directly, as it is not a relational property\. LetTTbe partition rich and let𝒞∈𝒫⁡\(𝒳\)\\mathcal\{C\}\\in\\mathcal\{P\}\(\\mathcal\{X\}\)\. There existsd0∈𝒟⁡\(𝒳\)d\_\{0\}\\in\\mathcal\{D\}\(\\mathcal\{X\}\)such that𝒞⊆T⁡\(d0\)\.\\mathcal\{C\}\\subseteq T\(d\_\{0\}\)\.Becauseμ\\muis surjective, choosed∈𝒟⁡\(𝒳\)d\\in\\mathcal\{D\}\(\\mathcal\{X\}\)such thatμ⁡\(d\)=d0\\mu\(d\)=d\_\{0\}\. Then𝒞⊆T⁡\(μ⁡\(d\)\),\\mathcal\{C\}\\subseteq T\(\\mu\(d\)\),soT∘μT\\circ\\muis partition rich\.

For partition consistency, fix𝒞∈𝒫⁡\(𝒳\)\\mathcal\{C\}\\in\\mathcal\{P\}\(\\mathcal\{X\}\)\. The assumption states precisely that

\(d,d′\)∈Rαpc𝒞⟹\(μ⁡\(d\),μ⁡\(d′\)\)∈Rαpc𝒞,\(d,d^\{\\prime\}\)\\in R^\{\\alpha\_\{\\mathrm\{pc\}\}^\{\\mathcal\{C\}\}\}\\quad\\Longrightarrow\\quad\\bigl\(\\mu\(d\),\\mu\(d^\{\\prime\}\)\\bigr\)\\in R^\{\\alpha\_\{\\mathrm\{pc\}\}^\{\\mathcal\{C\}\}\},and thus Proposition[42](https://arxiv.org/html/2609.11173#Thmtheorem42)shows thatμ\\mupreservesαpc𝒞\\alpha\_\{\\mathrm\{pc\}\}^\{\\mathcal\{C\}\}\. Because𝒞\\mathcal\{C\}was arbitrary, Lemma[43](https://arxiv.org/html/2609.11173#Thmtheorem43)gives preservation ofαpc\\alpha\_\{\\mathrm\{pc\}\}\.

For permutation invariance, fixϕ∈Π⁡\(𝒳\)\\phi\\in\\Pi\(\\mathcal\{X\}\)\. If\(d,dϕ\)∈Rαpermϕ\(d,d\_\{\\phi\}\)\\in R^\{\\alpha\_\{\\mathrm\{perm\}\}^\{\\phi\}\}, thenμ⁡\(dϕ\)=μ​\(d\)ϕ,\\mu\(d\_\{\\phi\}\)=\\mu\(d\)\_\{\\phi\},and therefore\(μ⁡\(d\),μ⁡\(dϕ\)\)=\(μ⁡\(d\),μ​\(d\)ϕ\)∈Rαpermϕ\.\\bigl\(\\mu\(d\),\\mu\(d\_\{\\phi\}\)\\bigr\)=\\bigl\(\\mu\(d\),\\mu\(d\)\_\{\\phi\}\\bigr\)\\in R^\{\\alpha\_\{\\mathrm\{perm\}\}^\{\\phi\}\}\.Proposition[42](https://arxiv.org/html/2609.11173#Thmtheorem42)and the conjunction Lemma[43](https://arxiv.org/html/2609.11173#Thmtheorem43)give preservation ofαperm\\alpha\_\{\\mathrm\{perm\}\}\.

Finally, fixΨ∈𝒯⁡\(𝒳\)\\Psi\\in\\mathcal\{T\}\(\\mathcal\{X\}\)andu∈𝒰Ψu\\in\\mathcal\{U\}\_\{\\Psi\}\. By assumption,μ⁡\(u\)\\mu\(u\)is an ultrametric satisfyingΨμ⁡\(u\)=Ψu=Ψ\.\\Psi\_\{\\mu\(u\)\}=\\Psi\_\{u\}=\\Psi\.Thusμ⁡\(u\)∈𝒰Ψ\\mu\(u\)\\in\\mathcal\{U\}\_\{\\Psi\}, and hence\(u,u\)∈RαsuhrΨ⟹\(μ⁡\(u\),μ⁡\(u\)\)∈RαsuhrΨ\.\(u,u\)\\in R^\{\\alpha\_\{\\mathrm\{suhr\}\}^\{\\Psi\}\}\\Longrightarrow\\bigl\(\\mu\(u\),\\mu\(u\)\\bigr\)\\in R^\{\\alpha\_\{\\mathrm\{suhr\}\}^\{\\Psi\}\}\.Again, Proposition[42](https://arxiv.org/html/2609.11173#Thmtheorem42)and Lemma[43](https://arxiv.org/html/2609.11173#Thmtheorem43)prove preservation of exactness on ultrametrics\. ∎

###### Proof of Corollary[29](https://arxiv.org/html/2609.11173#Thmtheorem29)\.

The conditiong⁡\(0\)=0g\(0\)=0, together withg⁡\(t\)\>0g\(t\)\>0fort\>0t\>0, ensures thatμg​\(d\)∈𝒟​\(𝒳\)\\mu\_\{g\}\(d\)\\in\\mathcal\{D\}\(\\mathcal\{X\}\)wheneverd∈𝒟⁡\(𝒳\)d\\in\\mathcal\{D\}\(\\mathcal\{X\}\)\.

Suppose first thatggis strictly increasing\. Thenggpreserves all inequalities defining a𝒞\\mathcal\{C\}\-strengthening\. Hence, wheneverd′d^\{\\prime\}is a𝒞\\mathcal\{C\}\-strengthening ofdd,μg​\(d′\)\\mu\_\{g\}\(d^\{\\prime\}\)is a𝒞\\mathcal\{C\}\-strengthening ofμg​\(d\)\\mu\_\{g\}\(d\)\. By Lemma[28](https://arxiv.org/html/2609.11173#Thmtheorem28),μg\\mu\_\{g\}therefore preserves partition consistency\.

Moreover, ifuuis an ultrametric, then

g⁡\(u⁡\(x,y\)\)≤g⁡\(max⁡\{u⁡\(x,z\),u⁡\(y,z\)\}\)=max⁡\{g⁡\(u⁡\(x,z\)\),g⁡\(u⁡\(y,z\)\)\},g\(u\(x,y\)\)\\ \\leq\\ g\\\!\\left\(\\max\\\{u\(x,z\),u\(y,z\)\\\}\\right\)\\ =\\ \\max\\\{g\(u\(x,z\)\),g\(u\(y,z\)\)\\\},soμg​\(u\)\\mu\_\{g\}\(u\)is an ultrametric\. Sinceggis strictly increasing, it preserves the ordering of all pairwise dissimilarities, and thereforeuuandμg​\(u\)\\mu\_\{g\}\(u\)induce the same hierarchy\. Thus Lemma[28](https://arxiv.org/html/2609.11173#Thmtheorem28)also gives preservation of exactness on ultrametrics\.

For every permutationϕ∈Π⁡\(𝒳\)\\phi\\in\\Pi\(\\mathcal\{X\}\),

μg​\(dϕ\)​\(x,y\)=g⁡\(d⁡\(ϕ−1​\(x\),ϕ−1​\(y\)\)\)=\(μg​\(d\)\)ϕ​\(x,y\),\\mu\_\{g\}\(d\_\{\\phi\}\)\(x,y\)\\ =\\ g\\\!\\left\(d\(\\phi^\{\-1\}\(x\),\\phi^\{\-1\}\(y\)\)\\right\)\\ =\\ \\bigl\(\\mu\_\{g\}\(d\)\\bigr\)\_\{\\phi\}\(x,y\),and henceμg​\(dϕ\)=\(μg​\(d\)\)ϕ\.\\mu\_\{g\}\(d\_\{\\phi\}\)=\\bigl\(\\mu\_\{g\}\(d\)\\bigr\)\_\{\\phi\}\.Thereforeμg\\mu\_\{g\}preserves permutation invariance\.

Ifggis surjective, letd0∈𝒟⁡\(𝒳\)d\_\{0\}\\in\\mathcal\{D\}\(\\mathcal\{X\}\)\. For every unordered pair\{x,y\}⊆𝒳\\\{x,y\\\}\\subseteq\\mathcal\{X\}withx≠yx\\neq y, choosetx​y\>0t\_\{xy\}\>0such thatg⁡\(tx​y\)=d0​\(x,y\)\.g\(t\_\{xy\}\)=d\_\{0\}\(x,y\)\.Such a positive preimage exists becaused0​\(x,y\)\>0d\_\{0\}\(x,y\)\>0,ggis surjective, andg⁡\(0\)=0g\(0\)=0\. Defined⁡\(x,x\):=0d\(x,x\):=0, andd⁡\(x,y\)=d⁡\(y,x\):=tx​yd\(x,y\)=d\(y,x\):=t\_\{xy\}forx≠yx\\neq y\. Thend∈𝒟⁡\(𝒳\)d\\in\\mathcal\{D\}\(\\mathcal\{X\}\)andμg​\(d\)=d0\\mu\_\{g\}\(d\)=d\_\{0\}\. Henceμg\\mu\_\{g\}is surjective and, by Lemma[28](https://arxiv.org/html/2609.11173#Thmtheorem28), preserves partition richness\.

Finally, if for everyβ\>0\\beta\>0there existscβ\>0c\_\{\\beta\}\>0such thatg⁡\(β​t\)=cβ​g​\(t\)g\(\\beta t\)=c\_\{\\beta\}g\(t\)for everyt≥0t\\geq 0, thenμg​\(β​d\)=cβ​μg​\(d\)\.\\mu\_\{g\}\(\\beta d\)=c\_\{\\beta\}\\,\\mu\_\{g\}\(d\)\.Lemma[28](https://arxiv.org/html/2609.11173#Thmtheorem28)therefore gives preservation of scale invariance\.

Forg⁡\(t\)=tpg\(t\)=t^\{p\}, withp\>0p\>0, all the preceding conditions are satisfied, withcβ=βpc\_\{\\beta\}=\\beta^\{p\}\. Hence the power transformation preserves all five properties, and thus preserves both admissibility and strong admissibility\. ∎

## Appendix CProofs for Section[3](https://arxiv.org/html/2609.11173#S3)\(Admissible Methods\)

We now prove the requirements satisfied by the methods introduced in Section[3](https://arxiv.org/html/2609.11173#S3)\. Whenever convenient, we establish stronger requirements than those required for admissibility: e\.g\., we establish thatTSL,Tglob,Tloc∈ℋαord∧αcwc∧αadm\+​\(𝒳\)T\_\{\\mathrm\{SL\}\},T\_\{\\mathrm\{glob\}\},T\_\{\\mathrm\{loc\}\}\\in\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{ord\}\}\\wedge\\alpha\_\{\\mathrm\{cwc\}\}\\wedge\\alpha\_\{\\mathrm\{adm\+\}\}\}\(\\mathcal\{X\}\), andTstable∈ℋαcwc∧αadm\+​\(𝒳\)T\_\{\\mathrm\{stable\}\}\\in\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{cwc\}\}\\wedge\\alpha\_\{\\mathrm\{adm\+\}\}\}\(\\mathcal\{X\}\)\.

### C\.1Single Linkage

#### C\.1\.1Single linkage and minimum spanning tree

Recall that the*minimum bottleneck*between two pointsx,y∈𝒳x,y\\in\\mathcal\{X\}is defined by

B∗​\(d\)​\(x,y\):=\{minγ∈Γx,y⁡max\{u,v\}∈γ⁡d⁡\(u,v\),x≠y,0,x=y,\\displaystyle B^\{\*\}\(d\)\(x,y\)\\ :=\\ \\begin\{cases\}\\displaystyle\\min\_\{\\gamma\\in\\Gamma\_\{x,y\}\}\\max\_\{\\\{u,v\\\}\\in\\gamma\}d\(u,v\),&x\\neq y,\\\\\[4\.30554pt\] 0,&x=y,\\end\{cases\}\(14\)whereΓx,y\\Gamma\_\{x,y\}denotes the set of paths fromxxtoyy\. The valueB∗​\(d\)​\(x,y\)B^\{\*\}\(d\)\(x,y\)can be computed by finding a minimum spanning tree \(MST\) of the complete graph on𝒳\\mathcal\{X\}with edge weightsdd\.

###### Lemma 44\.

Letd∈𝒟⁡\(𝒳\)d\\in\\mathcal\{D\}\(\\mathcal\{X\}\), and letMMbe a minimum spanning tree of the complete weighted graph on𝒳\\mathcal\{X\}with edge weightsdd\. For everyx,y∈𝒳x,y\\in\\mathcal\{X\}, letγx​yd\\gamma^\{d\}\_\{xy\}denote the unique path fromxxtoyyinMM\. ThenB∗​\(d\)​\(x,y\)=max\(x′,y′\)∈γx​yd⁡d⁡\(x′,y′\)\.B^\{\*\}\(d\)\(x,y\)=\\max\_\{\(x^\{\\prime\},y^\{\\prime\}\)\\in\\gamma^\{d\}\_\{xy\}\}d\(x^\{\\prime\},y^\{\\prime\}\)\.

###### Proof\.

Lete∗e^\{\\ast\}be an edge of maximal weight onγx​yd\\gamma^\{d\}\_\{xy\}, and denotet:=d⁡\(e∗\)=maxe∈γx​yd⁡d⁡\(e\)\.t:=d\(e^\{\\ast\}\)=\\max\_\{e\\in\\gamma^\{d\}\_\{xy\}\}d\(e\)\.Becauseγx​yd\\gamma^\{d\}\_\{xy\}is a path fromxxtoyy, we obtain from \([14](https://arxiv.org/html/2609.11173#A3.E14)\)B∗​\(d\)​\(x,y\)≤t\.B^\{\*\}\(d\)\(x,y\)\\leq t\.To prove the reverse inequality, observe that removinge∗e^\{\\ast\}disconnectsMMinto two components, separatingxxandyy\. Any path fromxxtoyymust cross the resulting cut\. If some crossing edge had weight strictly smaller thantt, replacinge∗e^\{\\ast\}by that edge would produce a spanning tree of smaller total weight, contradicting the minimality ofMM\. Hence everyxx–yypath contains an edge of weight at leasttt, soB∗​\(d\)​\(x,y\)≥t\.B^\{\*\}\(d\)\(x,y\)\\geq t\.∎

###### Lemma 45\.

For everyd∈𝒟⁡\(𝒳\)d\\in\\mathcal\{D\}\(\\mathcal\{X\}\),B∗​\(d\)B^\{\*\}\(d\)is an ultrametric andTSL​\(d\)=ΨB∗​\(d\)T\_\{\\mathrm\{SL\}\}\(d\)=\\Psi\_\{B^\{\*\}\(d\)\}\. Equivalently, for everyr≥0r\\geq 0, the clusters ofTSL​\(d\)T\_\{\\mathrm\{SL\}\}\(d\)at levelrrare the connected components of the graphGr​\(d\):=\(𝒳,\{\{x,y\}:d⁡\(x,y\)≤r\}\)\.G\_\{r\}\(d\):=\\bigl\(\\mathcal\{X\},\\\{\\\{x,y\\\}:d\(x,y\)\\leq r\\\}\\bigr\)\.

Lemma[45](https://arxiv.org/html/2609.11173#Thmtheorem45)is the standard minimax\-path characterization of single linkage\(see, for example,[Carlsson and Mémoli, 2010](https://arxiv.org/html/2609.11173#bib.bib4), Proposition 8\), and thus we omit the proof\.

#### C\.1\.2Properties of Single Linkage

We first record the transformation properties of the minimax mapB∗B^\{\*\}\.

###### Lemma 46\.

Transformationd↦B∗​\(d\)d\\mapsto B^\{\*\}\(d\)preservesαord,αsuhr\\alpha\_\{\\mathrm\{ord\}\},\\alpha\_\{\\mathrm\{suhr\}\}, andαperm\\alpha\_\{\\mathrm\{perm\}\}\.

###### Proof\.

*1\. Order invariance\.*Letg:R≥0→R≥0g:\\mathbb\{R\}\_\{\\geq 0\}\\to\\mathbb\{R\}\_\{\\geq 0\}be strictly increasing withg⁡\(0\)=0g\(0\)=0\. For everyd∈𝒟⁡\(𝒳\)d\\in\\mathcal\{D\}\(\\mathcal\{X\}\)and everyx,y∈𝒳x,y\\in\\mathcal\{X\},

B∗​\(g∘d\)​\(x,y\)\\displaystyle B^\{\*\}\(g\\circ d\)\(x,y\):=minγ∈Γx​y⁡max\{u,v\}∈γ⁡g⁡\(d⁡\(u,v\)\)=minγ∈Γx​y⁡g⁡\(max\{u,v\}∈γ⁡d⁡\(u,v\)\)\\displaystyle:=\\min\_\{\\gamma\\in\\Gamma\_\{xy\}\}\\max\_\{\\\{u,v\\\}\\in\\gamma\}g\(d\(u,v\)\)=\\min\_\{\\gamma\\in\\Gamma\_\{xy\}\}g\\\!\\left\(\\max\_\{\\\{u,v\\\}\\in\\gamma\}d\(u,v\)\\right\)=g⁡\(minγ∈Γx​y⁡max\{u,v\}∈γ⁡d⁡\(u,v\)\)=g⁡\(B∗​\(d\)​\(x,y\)\)\.\\displaystyle=g\\\!\\left\(\\min\_\{\\gamma\\in\\Gamma\_\{xy\}\}\\max\_\{\\\{u,v\\\}\\in\\gamma\}d\(u,v\)\\right\)=g\\bigl\(B^\{\*\}\(d\)\(x,y\)\\bigr\)\.The second and third equalities use the strict monotonicity ofgg\. HenceB∗​\(g∘d\)=g∘B∗​\(d\)\.B^\{\*\}\(g\\circ d\)=g\\circ B^\{\*\}\(d\)\.Therefore the transformationB∗B^\{\*\}preserves order equivalence of dissimilarities, and consequently preserves order invariance\.

*2\. Exactness on ultrametric:*Letuube an ultrametric\. We proveB∗​\(u\)=uB^\{\*\}\(u\)=uby sandwiching\. The inequalityB∗​\(u\)≤uB^\{\*\}\(u\)\\leq uis immediate from the definition ofB∗B^\{\*\}, as the one\-edge path fromxxtoyygivesB∗​\(u\)​\(x,y\)≤u⁡\(x,y\)\.B^\{\*\}\(u\)\(x,y\)\\leq u\(x,y\)\.For the reverse inequality, fixx,y∈𝒳x,y\\in\\mathcal\{X\}and introduce an arbitraryxxtoyypathγ=\(\(x0,x1\),…,\(xm−1,xm\)\)\\gamma=\(\(x\_\{0\},x\_\{1\}\),\\dots,\(x\_\{m\-1\},x\_\{m\}\)\)\. Repeatedly applying the strong triangle inequality givesu⁡\(x,y\)≤maxi=1,…,m⁡u⁡\(xi−1,xi\)\.u\(x,y\)\\leq\\max\_\{i=1,\\dots,m\}u\(x\_\{i\-1\},x\_\{i\}\)\.Because this holds for anyxx\-to\-yypathγ\\gamma,u⁡\(x,y\)≤minγ∈Γx,y⁡max\(u,v\)∈γ⁡u⁡\(u,v\)=B∗​\(u\)​\(x,y\)\.u\(x,y\)\\leq\\min\_\{\\gamma\\in\\Gamma\_\{x,y\}\}\\max\_\{\(u,v\)\\in\\gamma\}u\(u,v\)=B^\{\*\}\(u\)\(x,y\)\.HenceB∗​\(u\)=uB^\{\*\}\(u\)=u\.

*3\. Permutation invariance:*For a pathγ=\(\(u1,u2\),\(u2,u3\),⋯,\(um−1,um\)\)\\gamma=\(\(u\_\{1\},u\_\{2\}\),\(u\_\{2\},u\_\{3\}\),\\cdots,\(u\_\{m\-1\},u\_\{m\}\)\), define the pathϕ−1​γ=\(\(ϕ−1​\(u1\),ϕ−1​\(u2\)\),\(ϕ−1​\(u2\),ϕ−1​\(u3\)\),⋯,\(ϕ−1​\(um−1\),ϕ−1​\(um\)\)\)\\phi^\{\-1\}\\gamma=\(\(\\phi^\{\-1\}\(u\_\{1\}\),\\phi^\{\-1\}\(u\_\{2\}\)\),\(\\phi^\{\-1\}\(u\_\{2\}\),\\phi^\{\-1\}\(u\_\{3\}\)\),\\cdots,\(\\phi^\{\-1\}\(u\_\{m\-1\}\),\\phi^\{\-1\}\(u\_\{m\}\)\)\)\. The mapγ↦ϕ−1​γ\\gamma\\mapsto\\phi^\{\-1\}\\gammais a bijection fromΓx,y\\Gamma\_\{x,y\}ontoΓϕ−1​\(x\),ϕ−1​\(y\)\\Gamma\_\{\\phi^\{\-1\}\(x\),\\phi^\{\-1\}\(y\)\}\. Hence

B∗​\(dϕ\)​\(x,y\)=B∗​\(d\)​\(ϕ−1​\(x\),ϕ−1​\(y\)\)=B∗​\(d\)ϕ​\(x,y\)\.B^\{\*\}\(d\_\{\\phi\}\)\(x,y\)\\ =\\ B^\{\*\}\(d\)\(\\phi^\{\-1\}\(x\),\\phi^\{\-1\}\(y\)\)\\ =\\ B^\{\*\}\(d\)\_\{\\phi\}\(x,y\)\.∎

###### Lemma 47\.

Single linkageTSLT\_\{\\mathrm\{SL\}\}satisfiesαord,αsuhr,αperm\\alpha\_\{\\mathrm\{ord\}\},\\alpha\_\{\\mathrm\{suhr\}\},\\alpha\_\{\\mathrm\{perm\}\}, andαcwc\\alpha\_\{\\mathrm\{cwc\}\}\.

###### Proof\.

BecauseTglobT\_\{\\mathrm\{glob\}\}satisfiesαord,αsuhr\\alpha\_\{\\mathrm\{ord\}\},\\alpha\_\{\\mathrm\{suhr\}\}, andαperm\\alpha\_\{\\mathrm\{perm\}\}, order invariance, exactness on ultrametrics, and permutation invariance follow immediately from Proposition[42](https://arxiv.org/html/2609.11173#Thmtheorem42), Lemma[45](https://arxiv.org/html/2609.11173#Thmtheorem45), and Lemma[46](https://arxiv.org/html/2609.11173#Thmtheorem46)\.

*Cluster\-wise consistency:*LetC∈TSL​\(d\)C\\in T\_\{\\mathrm\{SL\}\}\(d\), and letd′d^\{\\prime\}be a\{C\}\\\{C\\\}\-strengthening ofdd\. By Lemma[45](https://arxiv.org/html/2609.11173#Thmtheorem45), there existsr≥0r\\geq 0such thatCCis a connected component ofGr​\(d\)G\_\{r\}\(d\)\.

Becaused′d^\{\\prime\}is a\{C\}\\\{C\\\}\-strengthening, every edge withinCCwhosedd\-weight is at mostrrstill hasd′d^\{\\prime\}\-weight at mostrr\. ThusCCremains connected inGr​\(d′\)G\_\{r\}\(d^\{\\prime\}\)\. Conversely, every edge betweenCCandC¯\\bar\{C\}hasdd\-weight greater thanrr, and its weight can only increase under the strengthening\. Hence no edge ofGr​\(d′\)G\_\{r\}\(d^\{\\prime\}\)joinsCCtoC¯\\bar\{C\}\.

ThereforeCCis also a connected component ofGr​\(d′\)G\_\{r\}\(d^\{\\prime\}\), and henceC∈TSL​\(d′\)\.C\\in T\_\{\\mathrm\{SL\}\}\(d^\{\\prime\}\)\.ThusTSLT\_\{\\mathrm\{SL\}\}is cluster\-wise consistent\. ∎

### C\.2Other Linkage Methods

#### C\.2\.1Definitions of the Linkage Methods Considered

We briefly recall the non\-binary variants of the standard linkage rules considered in Section[3\.1](https://arxiv.org/html/2609.11173#S3.SS1)\. Recall that linkage methods proceed by iteratively merging clusters \(starting with singletons\)\. At a given step of the algorithm, let𝒞active\\mathcal\{C\}\_\{\\mathrm\{active\}\}denote the current partition into active clusters, and letD:𝒞active×𝒞active→R≥0D:\\mathcal\{C\}\_\{\\mathrm\{active\}\}\\times\\mathcal\{C\}\_\{\\mathrm\{active\}\}\\to\\mathbb\{R\}\_\{\\geq 0\}denote the linkage dissimilarity between active clusters, as determined by the chosen linkage rule\. The next merge occurs at the minimum inter\-cluster dissimilarity

r:=minC1,C2∈𝒞activeC1≠C2⁡D⁡\(C1,C2\)\.r:=\\min\_\{\\begin\{subarray\}\{c\}C\_\{1\},C\_\{2\}\\in\\mathcal\{C\}\_\{\\mathrm\{active\}\}\\\\ C\_\{1\}\\neq C\_\{2\}\\end\{subarray\}\}D\(C\_\{1\},C\_\{2\}\)\.However, several pairs of active clusters may attain this minimum simultaneously\. To merge all clusters involved in such ties simultaneously, form the graph whose vertices are the active clustersC∈𝒞activeC\\in\\mathcal\{C\}\_\{\\mathrm\{active\}\}, with an edge betweenC1C\_\{1\}andC2C\_\{2\}wheneverD⁡\(C1,C2\)=r\.D\(C\_\{1\},C\_\{2\}\)=r\.Each nontrivial connected component of this graph is then merged into a single new cluster, obtained as the union of its vertices\. In particular, ifC1C\_\{1\}is tied withC2C\_\{2\}andC2C\_\{2\}is tied withC3C\_\{3\}, then all three are merged together even ifD⁡\(C1,C3\)\>rD\(C\_\{1\},C\_\{3\}\)\>r\. Each cluster created in this way is added to the output hierarchy, and the procedure is repeated with the resulting collection of active clusters\.

The linkage dissimilarities considered here admit the Lance\-Williams update\([Murtagh and Contreras, 2017](https://arxiv.org/html/2609.11173#bib.bib3)\)

D⁡\(C11∪C12,C2\)\\displaystyle D\(C\_\{11\}\\cup C\_\{12\},C\_\{2\}\)=α1​D​\(C11,C2\)\+α2​D​\(C12,C2\)\+β​D​\(C11,C12\)\+γ​\|D⁡\(C11,C2\)−D⁡\(C12,C2\)\|,\\displaystyle=\\alpha\_\{1\}D\(C\_\{11\},C\_\{2\}\)\+\\alpha\_\{2\}D\(C\_\{12\},C\_\{2\}\)\+\\beta D\(C\_\{11\},C\_\{12\}\)\+\\gamma\\left\|D\(C\_\{11\},C\_\{2\}\)\-D\(C\_\{12\},C\_\{2\}\)\\right\|,where the coefficients are given in Table[2](https://arxiv.org/html/2609.11173#A3.T2)\. We takeD⁡\(\{x\},\{y\}\)=d⁡\(x,y\)D\(\\\{x\\\},\\\{y\\\}\)=d\(x,y\)\.

Table 2:Classical hierarchical clustering linkages and their Lance–Williams coefficients\.mCm\_\{C\}denotes the centroid of clusterCC\.Ward and Centroid linkages are typically defined for squared Euclidean dissimilarities, as the updates can become negative on arbitrary dissimilarities\. We implicitly restrict to squared Euclidean dissimilarities when working with these linkages\. Moreover, for linkage rules such as WPGMA and median linkage, the updated dissimilarity after a multiway merge may depend on the order in which clusters in a tied component are merged\. To obtain a well\-defined hierarchical clustering method, we therefore regard a fixed deterministic tie\-breaking convention as part of the linkage rule\. This choice will not affect the counterexamples below, since all comparisons determining the relevant merges are strict\.

#### C\.2\.2Non\-admissibility of the Other Linkage Methods

We demonstrate the non\-admissibility of complete, average, Ward, UPGMC, and WPGMC linkage methods by providing counterexamples that show they do not satisfy consistency\.

###### Lemma 48\.

Complete, unweighted and weighted average, Ward, centroid, and median linkages do not satisfy partition consistency\. Thus, none of them is admissible, regardless of the tie\-breaking rule\.

###### Proof\.

Fix distinct points1,2,3,4∈𝒳1,2,3,4\\in\\mathcal\{X\}, letY:=𝒳∖\{1,2,3,4\}Y:=\\mathcal\{X\}\\setminus\\\{1,2,3,4\\\},A:=\{1,2,3\}A:=\\\{1,2,3\\\}, and𝒞:=\{A,\{4\}\}∪\{\{y\}:y∈Y\}\.\\mathcal\{C\}:=\\\{A,\\\{4\\\}\\\}\\cup\\bigl\\\{\\\{y\\\}:y\\in Y\\bigr\\\}\.Defined,d′∈𝒟⁡\(𝒳\)d,d^\{\\prime\}\\in\\mathcal\{D\}\(\\mathcal\{X\}\)by

d\(1,2\)=d\(1,3\)=6,d\(1,4\)=d\(2,4\)=8,d\(2,3\)=4,d\(3,4\)=q,andd\(1,2\)=d\(1,3\)=6,\\qquad d\(1,4\)=d\(2,4\)=8,\\qquad d\(2,3\)=4,\\qquad d\(3,4\)=q,\\text\{ and\}d′\(1,2\)=3,d′\(i,j\)=d\(i,j\)for all other pairs\.d^\{\\prime\}\(1,2\)=3,\\qquad d^\{\\prime\}\(i,j\)=d\(i,j\)\\quad\\text\{for all other pairs\}\.Whenever at least one ofx,yx,ybelongs toYY, setd⁡\(x,y\)=d′​\(x,y\)=100d\(x,y\)=d^\{\\prime\}\(x,y\)=100\. Thusd′d^\{\\prime\}is a𝒞\\mathcal\{C\}\-strengthening ofdd\.

Chooseqqas in table below\. For every value ofqqin the table, bothddandd′d^\{\\prime\}are squared Euclidean dissimilarities\. The relevant linkage values are also given in the table\.

Underdd, every rule first merges\{2,3\}\\\{2,3\\\}, sinced⁡\(2,3\)=4d\(2,3\)=4is the unique smallest initial dissimilarity\. The table then givesD⁡\(\{2,3\},\{1\}\)<min⁡\{D⁡\(\{2,3\},\{4\}\),d⁡\(1,4\)\},D\(\\\{2,3\\\},\\\{1\\\}\)<\\min\\bigl\\\{D\(\\\{2,3\\\},\\\{4\\\}\),d\(1,4\)\\bigr\\\},while all dissimilarities involvingYYare larger\. Hence the second merge formsAA, soA∈T⁡\(d\)A\\in T\(d\)and therefore𝒞⊆T⁡\(d\)\\mathcal\{C\}\\subseteq T\(d\)\.

Underd′d^\{\\prime\}, the unique first merge is\{1,2\}\\\{1,2\\\}\. The last two columns show thatq<min⁡\{D′​\(\{1,2\},\{3\}\),D′​\(\{1,2\},\{4\}\)\},q<\\min\\bigl\\\{D^\{\\prime\}\(\\\{1,2\\\},\\\{3\\\}\),D^\{\\prime\}\(\\\{1,2\\\},\\\{4\\\}\)\\bigr\\\},so the second merge is\{3,4\}\\\{3,4\\\}\. Consequently, every subsequent cluster containing33also contains44, and henceA∉T⁡\(d′\)A\\notin T\(d^\{\\prime\}\)\.

ThusA∈T⁡\(d\)∖T⁡\(d′\)A\\in T\(d\)\\setminus T\(d^\{\\prime\}\), althoughd′d^\{\\prime\}is a𝒞\\mathcal\{C\}\-strengthening ofddand𝒞⊆T⁡\(d\)\\mathcal\{C\}\\subseteq T\(d\)\. Therefore each of the six linkage rules violates partition consistency\. All comparisons are strict, so the conclusion is independent of the tie\-breaking convention\. ∎

### C\.3Separation Methods: Proof of Proposition[11](https://arxiv.org/html/2609.11173#Thmtheorem11)

Lemma[49](https://arxiv.org/html/2609.11173#Thmtheorem49)proves thatTglob𝜼T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}andTloc𝜼T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}are hierarchical clustering methods satisfying scale invariance, cluster\-wise consistency, and permutation invariance, and Lemma[52](https://arxiv.org/html/2609.11173#Thmtheorem52)establishes strict hierarchical richness\. Because cluster\-wise consistency implies partition consistency and strict hierarchical richness implies partition richness, we conclude thatTglob𝜼T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}andTloc𝜼T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}are admissible for every separation margin sequence𝜼\\boldsymbol\{\\eta\}\. Lemma[50](https://arxiv.org/html/2609.11173#Thmtheorem50)shows that distinct margin sequences define distinct methods, completing the proof of Proposition[11](https://arxiv.org/html/2609.11173#Thmtheorem11)\. We also record the stronger properties of exactness on ultrametrics and order invariance for the unit\-margin methods in Lemmas[51](https://arxiv.org/html/2609.11173#Thmtheorem51)and[53](https://arxiv.org/html/2609.11173#Thmtheorem53)\.

#### C\.3\.1Basic properties

###### Lemma 49\.

For every separation margin sequence𝛈\\boldsymbol\{\\eta\},Tglob𝛈T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}andTloc𝛈T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}are hierarchical clustering methods satisfyingαsca\\alpha\_\{\\mathrm\{sca\}\},αcwc\\alpha\_\{\\mathrm\{cwc\}\}, andαperm\\alpha\_\{\\mathrm\{perm\}\}\.

###### Proof\.

TlocT\_\{\\mathrm\{loc\}\}is a known hierarchical clustering method, and for any dissimilaritydd,Tloc​\(d\)T\_\{\\mathrm\{loc\}\}\(d\)satisfies all the conditions in Definition[4](https://arxiv.org/html/2609.11173#Thmtheorem4)\([Dreveton et al\., 2025](https://arxiv.org/html/2609.11173#bib.bib12);[Balcan et al\., 2008](https://arxiv.org/html/2609.11173#bib.bib11)\)\. Let𝜼\\boldsymbol\{\\eta\}be a separation margin sequence\. By definition of the corresponding separation conditions, for every dissimilaritydd, we haveTglob𝜼​\(d\)⊆Tloc𝜼​\(d\)⊆Tloc\.T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}\(d\)\\subseteq T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}\(d\)\\subseteq T\_\{\\mathrm\{loc\}\}\.BecauseTloc​\(d\)T\_\{\\mathrm\{loc\}\}\(d\)is laminar, each of its subfamilies is also laminar\. Hence the outputs ofTglob𝜼T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}andTloc𝜼T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}are laminar as well\. Moreover, by construction, each of these outputs contains the root𝒳\\mathcal\{X\}and all leaves\{x\}\\\{x\\\},x∈𝒳x\\in\\mathcal\{X\}\. Therefore,Tglob𝜼T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}andTloc𝜼T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}are hierarchical clustering methods\.

Next, we show thatTloc𝜼T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}satisfies each property\. The proofs forTglob𝜼T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}are analogous and hence omitted\.

*1\.Tloc𝛈T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}satisfiesαsca\\alpha\_\{\\mathrm\{sca\}\}\.*Letβ\>0\\beta\>0\. For everyx,y,z∈𝒳x,y,z\\in\\mathcal\{X\}we haveβ⋅d⁡\(x,y\)β⋅d⁡\(x,z\)=d⁡\(x,y\)d⁡\(x,z\)\\frac\{\\beta\\cdot d\(x,y\)\}\{\\beta\\cdot d\(x,z\)\}=\\frac\{d\(x,y\)\}\{d\(x,z\)\}\. HenceC∈Tloc𝜼​\(β​d\)C\\in T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}\(\\beta d\)iffC∈Tloc𝜼​\(d\),C\\in T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}\(d\),and thereforeTloc𝜼​\(β​d\)=Tloc𝜼​\(d\)T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}\(\\beta d\)=T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}\(d\)\.

*2\.Tloc𝛈T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}satisfiesαcwc\\alpha\_\{\\mathrm\{cwc\}\}\.*Letddbe a dissimilarity function, and takeC∈Tloc𝜼​\(d\)C\\in T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}\(d\)\. Because the root and singleton clusters belong toTloc𝜼​\(d\)T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}\(d\)for every dissimilarity, it suffices to consider a nontrivial proper clusterCC\. Consider a dissimilarityd′d^\{\\prime\}such thatd′d^\{\\prime\}is a\{C\}\\\{C\\\}\-strengthening ofdd\. Then, for everyx,y∈Cx,y\\in C,z∈𝒳\\Cz\\in\\mathcal\{X\}\\backslash C,d′​\(x,y\)≤d⁡\(x,y\)d^\{\\prime\}\(x,y\)\\leq d\(x,y\)andd′​\(x,z\)≥d⁡\(x,z\)d^\{\\prime\}\(x,z\)\\geq d\(x,z\)hold\. Thus

τ⁡\(d′,C\)=maxx,y∈C,z∈C¯⁡d′​\(x,y\)d′​\(x,z\)≤maxx,y∈C,z∈C¯⁡d⁡\(x,y\)d⁡\(x,z\)=τ⁡\(d,C\)<η\|C\|−1,\\tau\(d^\{\\prime\},C\)\\ =\\ \\max\_\{x,y\\in C,z\\in\\bar\{C\}\}\\frac\{d^\{\\prime\}\(x,y\)\}\{d^\{\\prime\}\(x,z\)\}\\ \\leq\\ \\max\_\{x,y\\in C,z\\in\\bar\{C\}\}\\frac\{d\(x,y\)\}\{d\(x,z\)\}\\ =\\ \\tau\(d,C\)\\ <\\ \\eta\_\{\|C\|\-1\},establishing thatC∈Tloc𝜼​\(d′\)C\\in T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}\(d^\{\\prime\}\)\.

*3\.Tloc𝛈T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}satisfiesαperm\\alpha\_\{\\mathrm\{perm\}\}\.*Letϕ∈Π⁡\(𝒳\)\\phi\\in\\Pi\(\\mathcal\{X\}\)be a permutation, and recall thatdϕ​\(x,y\):=d⁡\(ϕ−1​\(x\),ϕ−1​\(y\)\)d\_\{\\phi\}\(x,y\):=d\(\\phi^\{\-1\}\(x\),\\phi^\{\-1\}\(y\)\)\(Definition[6](https://arxiv.org/html/2609.11173#Thmtheorem6)\)\. Hence,dϕ​\(ϕ⁡\(x\),ϕ⁡\(y\)\)=d⁡\(x,y\)d\_\{\\phi\}\(\\phi\(x\),\\phi\(y\)\)=d\(x,y\)for allx,y∈𝒳x,y\\in\\mathcal\{X\}, and

τ⁡\(d,C\):=maxx,y∈C,z∈C¯⁡d⁡\(x,y\)d⁡\(x,z\)\\displaystyle\\tau\(d,C\):=\\max\_\{x,y\\in C,z\\in\\bar\{C\}\}\\frac\{d\(x,y\)\}\{d\(x,z\)\}=maxx,y∈C,z∈C¯⁡dϕ​\(ϕ⁡\(x\),ϕ⁡\(y\)\)dϕ​\(ϕ⁡\(x\),ϕ⁡\(z\)\)=maxx,y∈ϕ⁡\(C\),z∈ϕ⁡\(C\)¯⁡dϕ​\(x,y\)dϕ​\(x,z\)=τ⁡\(dϕ,ϕ⁡\(C\)\),\\displaystyle\\ =\\ \\max\_\{x,y\\in C,z\\in\\bar\{C\}\}\\frac\{d\_\{\\phi\}\(\\phi\(x\),\\phi\(y\)\)\}\{d\_\{\\phi\}\(\\phi\(x\),\\phi\(z\)\)\}\\ =\\ \\max\_\{\\begin\{subarray\}\{c\}x,y\\in\\phi\(C\),\\\\ z\\in\\overline\{\\phi\(C\)\}\\end\{subarray\}\}\\frac\{d\_\{\\phi\}\(x,y\)\}\{d\_\{\\phi\}\(x,z\)\}\\ =\\ \\tau\(d\_\{\\phi\},\\phi\(C\)\),where we used the fact thatϕ\\phiis a bijection andϕ⁡\(C¯\)=ϕ⁡\(C\)¯\\phi\(\\bar\{C\}\)=\\overline\{\\phi\(C\)\}\. Moreover,\|ϕ⁡\(C\)\|=\|C\|\|\\phi\(C\)\|=\|C\|\. Henceτ⁡\(d,C\)<η\|C\|−1\\tau\(d,C\)<\\eta\_\{\|C\|\-1\}iffτ⁡\(dϕ,ϕ⁡\(C\)\)<η\|ϕ⁡\(C\)\|−1\.\\tau\(d\_\{\\phi\},\\phi\(C\)\)<\\eta\_\{\|\\phi\(C\)\|\-1\}\.Therefore:C∈Tloc𝜼​\(d\)C\\in T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}\(d\)iffϕ⁡\(C\)∈Tloc𝜼​\(dϕ\)\.\\phi\(C\)\\in T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}\(d\_\{\\phi\}\)\.ThusTloc𝜼​\(dϕ\)=ϕ⋅Tloc𝜼​\(d\)\.T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}\(d\_\{\\phi\}\)=\\phi\\cdot T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}\(d\)\.∎

###### Lemma 50\.

Let𝛈≠𝛈′\\boldsymbol\{\\eta\}\\neq\\boldsymbol\{\\eta\}^\{\\prime\}be two distinct separation margin sequences\. ThenTglob𝛈≠Tglob𝛈′T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}\\neq T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}^\{\\prime\}\}andTloc𝛈≠Tloc𝛈′T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}\\neq T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}^\{\\prime\}\}\.

###### Proof\.

Because𝜼≠𝜼′\\boldsymbol\{\\eta\}\\neq\\boldsymbol\{\\eta\}^\{\\prime\}, there existss∈\[n−2\]s\\in\[n\-2\]such thatηs≠ηs′\\eta\_\{s\}\\neq\\eta\_\{s\}^\{\\prime\}\. Without loss of generality, assumeηs<ηs′\.\\eta\_\{s\}<\\eta\_\{s\}^\{\\prime\}\.ChooseC⊊𝒳C\\subsetneq\\mathcal\{X\}with\|C\|=s\+1\|C\|=s\+1, setr:=ηs\+ηs′2,r:=\\frac\{\\eta\_\{s\}\+\\eta\_\{s\}^\{\\prime\}\}\{2\},and define

d\(x,y\)=rfor distinctx,y∈Corx,y∈C¯,andd\(x,y\)=1for other distinctx,yd\(x,y\)=r\\text\{ for distinct \}x,y\\in C\\text\{ or \}x,y\\in\\bar\{C\},\\qquad\\text\{and\}\\qquad d\(x,y\)=1\\text\{ for other distinct \}x,yObserve thatϱ⁡\(d,C\)=τ⁡\(d,C\)=r\\varrho\(d,C\)=\\tau\(d,C\)=r\. Therefore, the setCCbelongs toTloc𝜼′​\(d\)T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}^\{\\prime\}\}\(d\)and toTglob𝜼′​\(d\)T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}^\{\\prime\}\}\(d\)but belongs neither toTloc𝜼​\(d\)T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}\(d\)nor toTglob𝜼​\(d\)T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}\(d\), provingTloc𝜼​\(d\)≠Tloc𝜼′​\(d\)T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}\(d\)\\neq T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}^\{\\prime\}\}\(d\)andTglob𝜼​\(d\)≠Tglob𝜼′​\(d\)T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}\(d\)\\neq T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}^\{\\prime\}\}\(d\)\. ∎

#### C\.3\.2Specific properties for𝜼=𝟏\\boldsymbol\{\\eta\}=\\mathbf\{1\}

###### Lemma 51\.

Let𝛈\\boldsymbol\{\\eta\}be a separation margin sequence\. These statements are equivalent: \(i\)𝛈=𝟏\\boldsymbol\{\\eta\}=\\mathbf\{1\}; \(ii\)Tloc𝛈T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}satisfiesαsuhr\\alpha\_\{\\mathrm\{suhr\}\}; \(iii\)Tglob𝛈T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}satisfiesαsuhr\\alpha\_\{\\mathrm\{suhr\}\}\.

###### Proof\.

We first show that \(i\) implies both \(ii\) and \(iii\)\. Assume that𝜼=𝟏\\boldsymbol\{\\eta\}=\\mathbf\{1\}, so thatTloc𝜼=TlocT\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}=T\_\{\\mathrm\{loc\}\}andTglob𝜼=TglobT\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}=T\_\{\\mathrm\{glob\}\}\. Fix an ultrametricu∈𝒰⁡\(𝒳\)u\\in\\mathcal\{U\}\(\\mathcal\{X\}\), and letΨu\\Psi\_\{u\}denote its associated hierarchy\. We prove thatΨu⊆Tglob​\(u\)\\Psi\_\{u\}\\subseteq T\_\{\\mathrm\{glob\}\}\(u\)andTloc​\(u\)⊆ΨuT\_\{\\mathrm\{loc\}\}\(u\)\\subseteq\\Psi\_\{u\}\. Combined withTglob​\(u\)⊆Tloc​\(u\)T\_\{\\mathrm\{glob\}\}\(u\)\\subseteq T\_\{\\mathrm\{loc\}\}\(u\), this establishesΨu=Tglob​\(u\)=Tloc​\(u\)\\Psi\_\{u\}=T\_\{\\mathrm\{glob\}\}\(u\)=T\_\{\\mathrm\{loc\}\}\(u\)

We first proveΨu⊆Tglob​\(u\)\\Psi\_\{u\}\\subseteq T\_\{\\mathrm\{glob\}\}\(u\)\. The root and singleton clusters belong toTglob​\(u\)T\_\{\\mathrm\{glob\}\}\(u\)by definition, so letC∈ΨuC\\in\\Psi\_\{u\}be a non\-root and non\-singleton cluster\. By the ultrametric\-ball representation ofΨu\\Psi\_\{u\}\(see Equation \([3](https://arxiv.org/html/2609.11173#S5.E3)\)\), there existsr≥0r\\geq 0such thatu⁡\(x,y\)≤r<u⁡\(x,z\)u\(x,y\)\\leq r<u\(x,z\)for allx,y∈C,z∉C\.x,y\\in C,\\ z\\notin C\.Hence, for allx1,x2,y∈Cx\_\{1\},x\_\{2\},y\\in Candz∉Cz\\notin C, we haveu⁡\(x1,y\)u⁡\(x2,z\)<1\.\\frac\{u\(x\_\{1\},y\)\}\{u\(x\_\{2\},z\)\}<1\.Thereforeϱ⁡\(u,C\)<1,\\varrho\(u,C\)<1,and thusC∈Tglob​\(u\)C\\in T\_\{\\mathrm\{glob\}\}\(u\)\. This establishesΨu⊆Tglob​\(u\)\\Psi\_\{u\}\\subseteq T\_\{\\mathrm\{glob\}\}\(u\)

It remains to proveTloc​\(u\)⊆ΨuT\_\{\\mathrm\{loc\}\}\(u\)\\subseteq\\Psi\_\{u\}\. Again, the root and singleton clusters are immediate, so letC∈Tloc​\(u\)C\\in T\_\{\\mathrm\{loc\}\}\(u\)be a non\-root and non\-singleton cluster\. Becauseτ⁡\(u,C\)<1\\tau\(u,C\)<1, we haveu⁡\(x,y\)<u⁡\(x,z\)u\(x,y\)<u\(x,z\)for allx,y∈C,z∉C\.x,y\\in C,\\ z\\notin C\.Fixx0∈Cx\_\{0\}\\in C, and setRin:=maxy∈C⁡u⁡\(x0,y\)R\_\{\\mathrm\{in\}\}:=\\max\_\{y\\in C\}u\(x\_\{0\},y\)andRout:=minz∉C⁡u⁡\(x0,z\)\.R\_\{\\mathrm\{out\}\}:=\\min\_\{z\\notin C\}u\(x\_\{0\},z\)\.Because𝒳\\mathcal\{X\}is finite, these extrema are attained, and the preceding strict inequalities implyRin<Rout\.R\_\{\\mathrm\{in\}\}<R\_\{\\mathrm\{out\}\}\.Therefore,Bu​\(x0,Rin\)=C\.B\_\{u\}\(x\_\{0\},R\_\{\\mathrm\{in\}\}\)=C\.By the ultrametric\-ball representation ofΨu\\Psi\_\{u\}, this givesC∈ΨuC\\in\\Psi\_\{u\}\. This establishesTloc​\(u\)⊆ΨuT\_\{\\mathrm\{loc\}\}\(u\)\\subseteq\\Psi\_\{u\}\.

Hence,Ψu=Tglob​\(u\)=Tloc​\(u\)\\Psi\_\{u\}=T\_\{\\mathrm\{glob\}\}\(u\)=T\_\{\\mathrm\{loc\}\}\(u\)for every ultrametricuu, proving \(i\) implies \(ii\) and \(iii\)\.

We now prove the converses by contraposition\. Suppose that𝜼≠𝟏\\boldsymbol\{\\eta\}\\neq\\mathbf\{1\}\. Then there existss∈\[n−2\]s\\in\[n\-2\]such thatηs<1\\eta\_\{s\}<1\. Choose a subsetC⊊𝒳C\\subsetneq\\mathcal\{X\}with\|C\|=s\+1\|C\|=s\+1and chooserrsuch thatηs<r<1\.\\eta\_\{s\}<r<1\.Define a dissimilarityuuby

u\(x,y\)=rfor distinctx,y∈C,andu\(x,y\)=1for other distinctx,y,u\(x,y\)=r\\text\{ for distinct \}x,y\\in C,\\qquad\\text\{and\}\\qquad u\(x,y\)=1\\text\{ for other distinct\}x,y,and observe thatuuis an ultrametric\. Moreover, for everyx∈Cx\\in C,Bu​\(x,r\)=C,B\_\{u\}\(x,r\)=C,soC∈ΨuC\\in\\Psi\_\{u\}\. However,τ⁡\(u,C\)=ϱ⁡\(u,C\)=r\>ηs=η\|C\|−1\.\\tau\(u,C\)=\\varrho\(u,C\)=r\>\\eta\_\{s\}=\\eta\_\{\|C\|\-1\}\.HenceC∉Tloc𝜼​\(u\)C\\notin T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}\(u\)andC∉Tglob𝜼​\(u\)\.C\\notin T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}\(u\)\.Thus neitherTloc𝜼​\(u\)T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}\(u\)norTglob𝜼​\(u\)T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}\(u\)equalsΨu\\Psi\_\{u\}\. Therefore neither method is exact on ultrametrics\. This proves \(ii\) implies \(i\) and \(iii\) implies \(i\)\. ∎

Similarly, we obtain the strict hierarchical richness ofTglob𝜼T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}andTloc𝜼T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}\.

###### Lemma 52\.

Tglob𝜼T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}andTloc𝛈T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}satisfyαshr\\alpha\_\{\\mathrm\{shr\}\}for any separation margin sequence𝛈\\boldsymbol\{\\eta\}\.

###### Proof\.

Fix an arbitrary hierarchyΨ∈𝒯⁡\(𝒳\)\\Psi\\in\\mathcal\{T\}\(\\mathcal\{X\}\), and letq∈\(0,min1≤s≤n−2⁡ηs\)\.q\\in\\left\(0,\\min\_\{1\\leq s\\leq n\-2\}\\eta\_\{s\}\\right\)\.We will construct a dissimilarityddsuch thatTglob𝜼​\(d\)=Tloc𝜼​\(d\)=ΨT\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}\(d\)=T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}\(d\)=\\Psi\. ViewΨ\\Psias a rooted tree, with root𝒳\\mathcal\{X\}and leaves the singletons\{x\}\\\{x\\\}forx∈𝒳x\\in\\mathcal\{X\}; letdepth⁡\(C\)\\operatorname\{depth\}\(C\)denote the depth of clusterC∈ΨC\\in\\Psi, withdepth⁡\(𝒳\)=0\\operatorname\{depth\}\(\\mathcal\{X\}\)=0\. Assign to every non\-singletonC∈ΨC\\in\\Psithe heighth⁡\(C\):=qdepth⁡\(C\),h\(C\):=q^\{\\operatorname\{depth\}\(C\)\},and assign height00to the singleton leaves\. Define for every distinctx,yx,y,

uΨ​\(x,y\):=h⁡\(lcaΨ⁡\(x,y\)\)\.u\_\{\\Psi\}\(x,y\):=h\\bigl\(\\operatorname\{lca\}\_\{\\Psi\}\(x,y\)\\bigr\)\.ThenuΨu\_\{\\Psi\}is an ultrametric whose associated hierarchy is exactlyΨ\\Psi\.

LetC∈ΨC\\in\\Psibe non\-singleton and proper cluster, and denote bypar⁡\(C\)\\operatorname\{par\}\(C\)its parent inΨ\\Psi\. For allx1,x2,y∈Cx\_\{1\},x\_\{2\},y\\in Candz∉Cz\\notin C, we haveuΨ​\(x1,y\)≤h⁡\(C\)u\_\{\\Psi\}\(x\_\{1\},y\)\\leq h\(C\)anduΨ​\(x2,z\)≥h⁡\(par⁡\(C\)\)u\_\{\\Psi\}\(x\_\{2\},z\)\\geq h\(\\operatorname\{par\}\(C\)\)\. Usingh⁡\(C\)h⁡\(par⁡\(C\)\)=q\\frac\{h\(C\)\}\{h\(\\operatorname\{par\}\(C\)\)\}=qand the definition ofϱ⁡\(uΨ,C\)\\varrho\(u\_\{\\Psi\},C\), we obtainϱ≤q<η\|C\|−1\.\\varrho\\leq q<\\eta\_\{\|C\|\-1\}\.Asτ⁡\(uΨ,C\)≤ρ⁡\(uΨ,C\)\\tau\(u\_\{\\Psi\},C\)\\leq\\rho\(u\_\{\\Psi\},C\), it follows that everyC∈ΨC\\in\\Psibelongs toTglobη​\(uΨ\)T\_\{\\mathrm\{glob\}\}^\{\\eta\}\(u\_\{\\Psi\}\)and toTlocη​\(uΨ\)T\_\{\\mathrm\{loc\}\}^\{\\eta\}\(u\_\{\\Psi\}\), and thus

Ψ⊆Tglobη​\(uΨ\)andΨ⊆Tlocη​\(uΨ\)\.\\Psi\\subseteq T\_\{\\mathrm\{glob\}\}^\{\\eta\}\(u\_\{\\Psi\}\)\\quad\\text\{ and \}\\quad\\Psi\\subseteq T\_\{\\mathrm\{loc\}\}^\{\\eta\}\(u\_\{\\Psi\}\)\.Moreover, becauseTglobT\_\{\\mathrm\{glob\}\}andTlocT\_\{\\mathrm\{loc\}\}are exact on ultrametrics \(Lemma[51](https://arxiv.org/html/2609.11173#Thmtheorem51)\), we also have

Tglobη​\(uΨ\)⊆Tglob​\(uΨ\)=Ψ,andTlocη​\(uΨ\)⊆Tloc​\(uΨ\)=Ψ,T\_\{\\mathrm\{glob\}\}^\{\\eta\}\(u\_\{\\Psi\}\)\\subseteq T\_\{\\mathrm\{glob\}\}\(u\_\{\\Psi\}\)=\\Psi,\\quad\\text\{ and \}\\quad T\_\{\\mathrm\{loc\}\}^\{\\eta\}\(u\_\{\\Psi\}\)\\subseteq T\_\{\\mathrm\{loc\}\}\(u\_\{\\Psi\}\)=\\Psi,HenceTglobη​\(uΨ\)=Tlocη​\(uΨ\)=ΨT\_\{\\mathrm\{glob\}\}^\{\\eta\}\(u\_\{\\Psi\}\)=T\_\{\\mathrm\{loc\}\}^\{\\eta\}\(u\_\{\\Psi\}\)=\\Psi, and both methods satisfy strict hierarchical richness\. ∎

###### Lemma 53\.

Let𝛈\\boldsymbol\{\\eta\}be a separation margin sequence\. These statements are equivalent: \(i\)𝛈=𝟏\\boldsymbol\{\\eta\}=\\mathbf\{1\}; \(ii\)Tloc𝛈T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}satisfiesαord\\alpha\_\{\\mathrm\{ord\}\}; \(iii\)Tglob𝛈T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}satisfiesαord\\alpha\_\{\\mathrm\{ord\}\}\.

###### Proof\.

We only prove the equivalence between \(i\) and \(ii\)\. The proof of the equivalence between \(i\) and \(iii\) is almost identical and thus is omitted\.

To show that \(i\) implies \(ii\), let𝜼=𝟏\\boldsymbol\{\\eta\}=\\mathbf\{1\}, so thatTloc𝜼=TlocT\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}=T\_\{\\mathrm\{loc\}\}\. Letg:R≥0→R≥0g:\\mathbb\{R\}\_\{\\geq 0\}\\to\\mathbb\{R\}\_\{\\geq 0\}be strictly increasing withg⁡\(0\)=0g\(0\)=0\. For every nonsingleton clusterC⊊𝒳C\\subsetneq\\mathcal\{X\},

τ⁡\(d,C\)<1⇔d⁡\(x,y\)<d⁡\(x,z\)for all​x,y∈C,z∉C\.\\tau\(d,C\)<1\\iff d\(x,y\)<d\(x,z\)\\quad\\text\{for all \}x,y\\in C,\\ z\\notin C\.Becauseggis strictly increasing, this is equivalent to

g⁡\(d⁡\(x,y\)\)<g⁡\(d⁡\(x,z\)\)for all​x,y∈C,z∉C,g\(d\(x,y\)\)<g\(d\(x,z\)\)\\quad\\text\{for all \}x,y\\in C,\\ z\\notin C,and hence toτ⁡\(g∘d,C\)<1\.\\tau\(g\\circ d,C\)<1\.ThereforeTloc​\(g∘d\)=Tloc​\(d\)\.T\_\{\\mathrm\{loc\}\}\(g\\circ d\)=T\_\{\\mathrm\{loc\}\}\(d\)\.

We now establish that \(ii\) implies \(i\)\. We prove it by contraposition\. Consider a separation margin sequence𝜼≠𝟏\\boldsymbol\{\\eta\}\\neq\\mathbf\{1\}; we will establish thatTloc𝜼T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}is not order\-invariant by constructing two dissimilaritiesd,d′d,d^\{\\prime\}such thatd′=g∘dd^\{\\prime\}=g\\circ dfor some strictly increasing functionggsatisfyingg⁡\(0\)=0g\(0\)=0, but for whichTloc𝜼​\(d\)≠Tloc𝜼​\(d′\)T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}\(d\)\\neq T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}\(d^\{\\prime\}\)\.

Because𝜼≠𝟏\\boldsymbol\{\\eta\}\\neq\\mathbf\{1\}, there existss∈\[n−2\]s\\in\[n\-2\]such thatηs<1\\eta\_\{s\}<1\. Choose a subsetC⊆𝒳C\\subseteq\\mathcal\{X\}withs=\|C\|−1s=\|C\|\-1\. Defined,d′∈𝒟⁡\(𝒳\)d,d^\{\\prime\}\\in\\mathcal\{D\}\(\\mathcal\{X\}\)such that for distinctx,y∈𝒳x,y\\in\\mathcal\{X\},

d⁡\(x,y\):=\{ηs/2,if​x,y∈C​or​x,y∈C¯,1,ifx∈C,y∈C¯,d\(x,y\):=\\begin\{cases\}\\eta\_\{s\}/2,&\\text\{if \}x,y\\in C\\text\{ or \}x,y\\in\\bar\{C\},\\\\ 1,&\\text\{if \}x\\in C,\\ y\\in\\bar\{C\},\\end\{cases\}and

d′​\(x,y\):=\{ηs,if​x,y∈C​or​x,y∈C¯,1,ifx∈C,y∈C¯\.d^\{\\prime\}\(x,y\):=\\begin\{cases\}\\eta\_\{s\},&\\text\{if \}x,y\\in C\\text\{ or \}x,y\\in\\bar\{C\},\\\\ 1,&\\text\{if \}x\\in C,\\ y\\in\\bar\{C\}\.\\end\{cases\}Because0<ηs/2<ηs<10<\\eta\_\{s\}/2<\\eta\_\{s\}<1there exists a strictly increasingg:R≥0→R≥0g:\\mathbb\{R\}\_\{\\geq 0\}\\to\\mathbb\{R\}\_\{\\geq 0\}satisfyingg⁡\(0\)=0g\(0\)=0,g⁡\(ηs/2\)=ηsg\(\\eta\_\{s\}/2\)=\\eta\_\{s\}, andg⁡\(1\)=1g\(1\)=1\. Henced′=g∘d\.d^\{\\prime\}=g\\circ d\.

Observe that for everyx,y∈Cx,y\\in Candz∉Cz\\notin C,d⁡\(x,y\)d⁡\(x,z\)=ηs/21<ηs,\\frac\{d\(x,y\)\}\{d\(x,z\)\}=\\frac\{\\eta\_\{s\}/2\}\{1\}<\\eta\_\{s\},henceC∈Tloc𝜼​\(d\)C\\in T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}\(d\)holds\. However, fix anyz0∈C¯z\_\{0\}\\in\\bar\{C\}\. For everyx,y∈Cx,y\\in Cwithx≠yx\\neq y,d′​\(x,y\)d′​\(x,z0\)=ηs1=ηs≮ηs\.\\frac\{d^\{\\prime\}\(x,y\)\}\{d^\{\\prime\}\(x,z\_\{0\}\)\}=\\frac\{\\eta\_\{s\}\}\{1\}=\\eta\_\{s\}\\not<\\eta\_\{s\}\.ThereforeC∉Tloc𝜼​\(d′\)\.C\\notin T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}\(d^\{\\prime\}\)\.HenceTloc𝜼T\_\{\\mathrm\{loc\}\}^\{\\boldsymbol\{\\eta\}\}is not order invariant\. ∎

### C\.4Bryant\-Berry Stable Clusters: Proof of Proposition[13](https://arxiv.org/html/2609.11173#Thmtheorem13)

We first establish the strong admissibility of the Bryant–Berry method\. Moreover, Corollary[29](https://arxiv.org/html/2609.11173#Thmtheorem29)shows thatμpowerp\\mu\_\{\\mathrm\{power\}\}^\{p\}preserves strong admissibility for everyp\>0p\>0\. HenceTstable\(p\)=Tstable∘μpowerpT\_\{\\mathrm\{stable\}\}^\{\(p\)\}=T\_\{\\mathrm\{stable\}\}\\circ\\mu\_\{\\mathrm\{power\}\}^\{p\}is strongly admissible for everyp\>0p\>0\. In particular,TstableT\_\{\\mathrm\{stable\}\}and all its power\-transformed variants are admissible, proving Proposition[13](https://arxiv.org/html/2609.11173#Thmtheorem13)\.

###### Lemma 54\.

TstableT\_\{\\mathrm\{stable\}\}is a hierarchical clustering method and satisfiesαsca,αsuhr,αcwc\\alpha\_\{\\mathrm\{sca\}\},\\alpha\_\{\\mathrm\{suhr\}\},\\alpha\_\{\\mathrm\{cwc\}\}andαperm\\alpha\_\{\\mathrm\{perm\}\}\. In particular,TstableT\_\{\\mathrm\{stable\}\}is strongly admissible\.

###### Proof\.

For everyd∈𝒟⁡\(𝒳\)d\\in\\mathcal\{D\}\(\\mathcal\{X\}\), Bryant and Berry\([Bryant and Berry, 2001](https://arxiv.org/html/2609.11173#bib.bib20)\)show that the stable clusters\{C⊆𝒳:ιd​\(C\)\>0\}\\\{C\\subseteq\\mathcal\{X\}:\\iota^\{d\}\(C\)\>0\\\}form a laminar family\. BecauseTstable​\(d\)T\_\{\\mathrm\{stable\}\}\(d\)additionally contains the root𝒳\\mathcal\{X\}and all singleton clusters, it is a hierarchy\. ThusTstableT\_\{\\mathrm\{stable\}\}is a hierarchical clustering method\. We now verify the four axioms\.

*1\. Scale invariance\.*For allβ\>0\\beta\>0andu,v,z∈𝒳u,v,z\\in\\mathcal\{X\}, we haveρβ​d​\(u​v∣z\)=β​ρd​\(u​v∣z\)\.\\rho\_\{\\beta d\}\(uv\\mid z\)=\\beta\\,\\rho\_\{d\}\(uv\\mid z\)\.Hence

ρ¯β​d​\(U​V∣Z\)=β​ρ¯d​\(U​V∣Z\)\\overline\{\\rho\}\_\{\\beta d\}\(UV\\mid Z\)=\\beta\\,\\overline\{\\rho\}\_\{d\}\(UV\\mid Z\)for every admissible triplet\(U,V,Z\)\(U,V,Z\), and thereforeιβ​d​\(C\)=β​ιd​\(C\)\.\\iota^\{\\beta d\}\(C\)=\\beta\\,\\iota^\{d\}\(C\)\.Becauseβ\>0\\beta\>0, we also haveιβ​d​\(C\)\>0\\iota^\{\\beta d\}\(C\)\>0iffιd​\(C\)\>0\.\\iota^\{d\}\(C\)\>0\.ThusTstable​\(β​d\)=Tstable​\(d\)T\_\{\\mathrm\{stable\}\}\(\\beta d\)=T\_\{\\mathrm\{stable\}\}\(d\)\.

*2\. Exactness on ultrametrics\.*Letuube an ultrametric, and letΨu\\Psi\_\{u\}be its associated hierarchy\. BecauseTloc⊑TstableT\_\{\\mathrm\{loc\}\}\\sqsubseteq T\_\{\\mathrm\{stable\}\}andTloc​\(u\)=ΨuT\_\{\\mathrm\{loc\}\}\(u\)=\\Psi\_\{u\}, we immediately obtainΨu⊆Tstable​\(u\)\.\\Psi\_\{u\}\\subseteq T\_\{\\mathrm\{stable\}\}\(u\)\.It remains to prove the reverse inclusion\.

LetC∈Tstable​\(u\)C\\in T\_\{\\mathrm\{stable\}\}\(u\)\. The root and singleton cases are immediate, so assume thatCCis a non\-singleton proper subset of𝒳\\mathcal\{X\}\. Thenιu​\(C\)\>0\\iota^\{u\}\(C\)\>0\.

We first show that every external point is equidistant from all points ofCC\. To do so, fixz∈𝒳∖Cz\\in\\mathcal\{X\}\\setminus C, and we will prove that the functionx↦u⁡\(x,z\)x\\mapsto u\(x,z\)is constant onCC\. By contradiction, suppose this is not the case\. Let

r:=minx∈C⁡u⁡\(x,z\),U:=\{x∈C:u⁡\(x,z\)=r\},V:=C∖U,Z:=\{z\}\.r:=\\min\_\{x\\in C\}u\(x,z\),\\qquad U:=\\\{x\\in C:u\(x,z\)=r\\\},\\qquad V:=C\\setminus U,\\qquad Z:=\\\{z\\\}\.ThenU,V,ZU,V,Zare nonempty, andU∩V=∅U\\cap V=\\emptysetandC=U∪VC=U\\cup V\. For everyx∈Ux\\in Uandy∈Vy\\in V, we haver=u⁡\(x,z\)<u⁡\(y,z\)\.r=u\(x,z\)<u\(y,z\)\.Becauseuuis an ultrametric, the two larger distances amongu⁡\(x,y\),u⁡\(x,z\),u⁡\(y,z\)u\(x,y\),u\(x,z\),u\(y,z\)are equal\. Henceu⁡\(x,y\)=u⁡\(y,z\)\.u\(x,y\)=u\(y,z\)\.Therefore

ρu​\(x​y∣z\)=min⁡\{u⁡\(x,z\),u⁡\(y,z\)\}−u⁡\(x,y\)=u⁡\(x,z\)−u⁡\(y,z\)<0\.\\rho\_\{u\}\(xy\\mid z\)=\\min\\\{u\(x,z\),u\(y,z\)\\\}\-u\(x,y\)=u\(x,z\)\-u\(y,z\)<0\.Averaging overx∈Ux\\in U,y∈Vy\\in V, andZ=\{z\}Z=\\\{z\\\}, we obtainρ¯u​\(U​V∣Z\)<0,\\overline\{\\rho\}\_\{u\}\(UV\\mid Z\)<0,contradictingιu​\(C\)\>0\\iota^\{u\}\(C\)\>0\. Thusu⁡\(x,z\)u\(x,z\)is constant onCC\. Denote its common value byλz:=u⁡\(x,z\),x∈C\.\\lambda\_\{z\}:=u\(x,z\),\\,x\\in C\.

We next show that every internal dissimilarity is strictly smaller than this common external dissimilarity\. Suppose, toward a contradiction, that there exist distinctx,y∈Cx,y\\in Csuch thatu⁡\(x,y\)≥λz\.u\(x,y\)\\geq\\lambda\_\{z\}\.Becauseu⁡\(x,z\)=u⁡\(y,z\)=λzu\(x,z\)=u\(y,z\)=\\lambda\_\{z\}, the ultrametric inequality gives

u⁡\(x,y\)≤max⁡\{u⁡\(x,z\),u⁡\(y,z\)\}=λz\.u\(x,y\)\\leq\\max\\\{u\(x,z\),u\(y,z\)\\\}=\\lambda\_\{z\}\.Henceu⁡\(x,y\)=λz\.u\(x,y\)=\\lambda\_\{z\}\.Define a relation∼\\simonCCbya∼b⟺u⁡\(a,b\)<λz\.a\\sim b\\,\\Longleftrightarrow\\,u\(a,b\)<\\lambda\_\{z\}\.By the ultrametric inequality,∼\\simis an equivalence relation onCC\. Becauseu⁡\(x,y\)=λzu\(x,y\)=\\lambda\_\{z\}, it has at least two equivalence classes\. LetUUbe one equivalence class and setV:=C∖U,Z:=\{z\}\.V:=C\\setminus U,\\,Z:=\\\{z\\\}\.For everya∈Ua\\in Uandb∈Vb\\in V, we haveu⁡\(a,b\)≥λzu\(a,b\)\\geq\\lambda\_\{z\}\. On the other hand,u⁡\(a,b\)≤max⁡\{u⁡\(a,z\),u⁡\(b,z\)\}=λz\.u\(a,b\)\\leq\\max\\\{u\(a,z\),u\(b,z\)\\\}=\\lambda\_\{z\}\.Thusu⁡\(a,b\)=λz\.u\(a,b\)=\\lambda\_\{z\}\.Thereforeρu​\(a​b∣z\)=min⁡\{u⁡\(a,z\),u⁡\(b,z\)\}−u⁡\(a,b\)=λz−λz=0\.\\rho\_\{u\}\(ab\\mid z\)=\\min\\\{u\(a,z\),u\(b,z\)\\\}\-u\(a,b\)=\\lambda\_\{z\}\-\\lambda\_\{z\}=0\.Henceρ¯u​\(U​V∣Z\)=0,\\overline\{\\rho\}\_\{u\}\(UV\\mid Z\)=0,again contradictingιu​\(C\)\>0\\iota^\{u\}\(C\)\>0\. Therefore, for everyz∈𝒳∖Cz\\in\\mathcal\{X\}\\setminus Cand all distinctx,y∈Cx,y\\in C,u⁡\(x,y\)<u⁡\(x,z\)=u⁡\(y,z\)\.u\(x,y\)<u\(x,z\)=u\(y,z\)\.

Now setR:=maxx,y∈C⁡u⁡\(x,y\)\.R:=\\max\_\{x,y\\in C\}u\(x,y\)\.The preceding inequality implies that, for everyx∈Cx\\in C,Bu​\(x,R\)=C,B\_\{u\}\(x,R\)=C,whereBuB\_\{u\}is defined in Equation \([3](https://arxiv.org/html/2609.11173#S5.E3)\)\. HenceC∈ΨuC\\in\\Psi\_\{u\}by the ultrametric\-ball representation ofΨu\\Psi\_\{u\}\. This proves thatTstable​\(u\)⊆Ψu\.T\_\{\\mathrm\{stable\}\}\(u\)\\subseteq\\Psi\_\{u\}\.

*3\. Cluster\-wise consistency:*LetC∈Tstable​\(d\)C\\in T\_\{\\mathrm\{stable\}\}\(d\), and letd′d^\{\\prime\}be a\{C\}\\\{C\\\}\-strengthening ofdd\. The root and singleton cases are immediate, so assume thatCCis nontrivial and proper\.

Consider two nonempty disjoint setsU,V⊆𝒳U,V\\subseteq\\mathcal\{X\}withU∪V=CU\\cup V=C, and a nonemptyZ⊆𝒳∖CZ\\subseteq\\mathcal\{X\}\\setminus C\. For everya∈Ua\\in U,b∈Vb\\in V,z∈Zz\\in Z, we haved′​\(a,b\)≤d⁡\(a,b\)d^\{\\prime\}\(a,b\)\\leq d\(a,b\),d′​\(a,z\)≥d⁡\(a,z\)d^\{\\prime\}\(a,z\)\\geq d\(a,z\), andd′​\(b,z\)≥d⁡\(b,z\)d^\{\\prime\}\(b,z\)\\geq d\(b,z\)\. Thus,

ρd′​\(a​b∣z\)=min⁡\{d′​\(a,z\),d′​\(b,z\)\}−d′​\(a,b\)≥min⁡\{d⁡\(a,z\),d⁡\(b,z\)\}−d⁡\(a,b\)=ρd​\(a​b∣z\)\.\\displaystyle\\rho\_\{d^\{\\prime\}\}\(ab\\mid z\)=\\min\\\{d^\{\\prime\}\(a,z\),d^\{\\prime\}\(b,z\)\\\}\-d^\{\\prime\}\(a,b\)\\ \\geq\\ \\min\\\{d\(a,z\),d\(b,z\)\\\}\-d\(a,b\)=\\rho\_\{d\}\(ab\\mid z\)\.Averaging and then minimizing over all admissible triples givesιd′​\(C\)≥ιd​\(C\)\>0\.\\iota^\{d^\{\\prime\}\}\(C\)\\geq\\iota^\{d\}\(C\)\>0\.HenceC∈Tstable​\(d′\)C\\in T\_\{\\mathrm\{stable\}\}\(d^\{\\prime\}\), proving cluster\-wise consistency\.

*4\. Permutation invariance\.*Letϕ\\phibe a permutation of𝒳\\mathcal\{X\}, and recalldϕ​\(x,y\):=d⁡\(ϕ−1​\(x\),ϕ−1​\(y\)\)d\_\{\\phi\}\(x,y\):=d\(\\phi^\{\-1\}\(x\),\\phi^\{\-1\}\(y\)\)\. For everyu,v,z∈𝒳u,v,z\\in\\mathcal\{X\}, we haveρdϕ​\(ϕ⁡\(u\)​ϕ​\(v\)∣ϕ⁡\(z\)\)=ρd​\(u​v∣z\)\.\\rho\_\{d\_\{\\phi\}\}\\bigl\(\\phi\(u\)\\phi\(v\)\\mid\\phi\(z\)\\bigr\)=\\rho\_\{d\}\(uv\\mid z\)\.Moreover, the map\(U,V,Z\)⟼\(ϕ⁡\(U\),ϕ⁡\(V\),ϕ⁡\(Z\)\)\(U,V,Z\)\\longmapsto\\bigl\(\\phi\(U\),\\phi\(V\),\\phi\(Z\)\\bigr\)is a bijection between the triples admissible in the definition ofιd​\(C\)\\iota^\{d\}\(C\)and those admissible in the definition ofιdϕ​\(ϕ​\(C\)\)\\iota^\{d\_\{\\phi\}\}\(\\phi\(C\)\)\. Thereforeιdϕ​\(ϕ⁡\(C\)\)=ιd​\(C\)\\iota^\{d\_\{\\phi\}\}\(\\phi\(C\)\)=\\iota^\{d\}\(C\)for everyC⊆𝒳C\\subseteq\\mathcal\{X\}\. Hence

C∈Tstable​\(d\)⇔ϕ⁡\(C\)∈Tstable​\(dϕ\),C\\in T\_\{\\mathrm\{stable\}\}\(d\)\\iff\\phi\(C\)\\in T\_\{\\mathrm\{stable\}\}\(d\_\{\\phi\}\),provingTstable​\(dϕ\)=ϕ⋅Tstable​\(d\)\.T\_\{\\mathrm\{stable\}\}\(d\_\{\\phi\}\)=\\phi\\cdot T\_\{\\mathrm\{stable\}\}\(d\)\.∎

## Appendix DProofs for Section[4](https://arxiv.org/html/2609.11173#S4)

### D\.1Incompatibility among Methods

###### Proof of Lemma[16](https://arxiv.org/html/2609.11173#Thmtheorem16)\.

First, we consider𝒳=\[4\]\\mathcal\{X\}=\[4\]\. We show that for any0<p<q0<p<q, the methodsTstable\(p\)T\_\{\\mathrm\{stable\}\}^\{\(p\)\}andTstable\(q\)T\_\{\\mathrm\{stable\}\}^\{\(q\)\}admit no common refinement that outputs hierarchies\. Let

𝒳=\{1,2,3,4\},C1:=\{1,2,3\},C2:=\{2,3,4\}\.\\mathcal\{X\}=\\\{1,2,3,4\\\},\\qquad C\_\{1\}:=\\\{1,2,3\\\},\\qquad C\_\{2\}:=\\\{2,3,4\\\}\.ThenC1∩C2=\{2,3\}C\_\{1\}\\cap C\_\{2\}=\\\{2,3\\\}is nonempty while neitherC1⊆C2C\_\{1\}\\subseteq C\_\{2\}norC2⊆C1C\_\{2\}\\subseteq C\_\{1\}, soC1C\_\{1\}andC2C\_\{2\}are not laminar\-compatible; in particular, no hierarchy on𝒳\\mathcal\{X\}contains both\. We constructd∈𝒟⁡\(𝒳\)d\\in\\mathcal\{D\}\(\\mathcal\{X\}\)such thatC1∈Tstable\(p\)​\(d\)​and​C2∈Tstable\(q\)​\(d\),C\_\{1\}\\in T\_\{\\mathrm\{stable\}\}^\{\(p\)\}\(d\)\\,\\text\{and\}\\,C\_\{2\}\\in T\_\{\\mathrm\{stable\}\}^\{\(q\)\}\(d\),hence prove thatTstable\(p\)T\_\{\\mathrm\{stable\}\}^\{\(p\)\}andTstable\(q\)T\_\{\\mathrm\{stable\}\}^\{\(q\)\}are mutually incompatible\.

*Construction ofdd\.*Letr:=q/p\>1r:=q/p\>1\. Fixα\>0\\alpha\>0small enough that\(1\+α/2\)r<2\\bigl\(1\+\\alpha/2\\bigr\)^\{r\}<2; such anα\\alphaexists because\(1\+α/2\)r→1\(1\+\\alpha/2\)^\{r\}\\to 1asα→0\+\\alpha\\to 0^\{\+\}\. Seth:=1\+α,t0:=1\+α2=1\+h2\.h:=1\+\\alpha,\\,t\_\{0\}:=1\+\\frac\{\\alpha\}\{2\}=\\frac\{1\+h\}\{2\}\.Noticet0r=\(1\+h2\)r<1\+hr2,t\_\{0\}^\{r\}=\\left\(\\frac\{1\+h\}\{2\}\\right\)^\{r\}<\\frac\{1\+h^\{r\}\}\{2\},equivalently1\+hr\>2​t0r1\+h^\{r\}\>2t\_\{0\}^\{r\}holds\. By continuity, both strict inequalitiest0r<2t\_\{0\}^\{r\}<2and1\+hr\>2​t0r1\+h^\{r\}\>2t\_\{0\}^\{r\}persist after a small perturbation oft0t\_\{0\}: there existsδ∈\(0,α/2\)\\delta\\in\(0,\\alpha/2\)such thatt:=t0\+δt:=t\_\{0\}\+\\deltasatisfiestr<2t^\{r\}<2and1\+hr\>2​tr\.1\+h^\{r\}\>2t^\{r\}\.The constraintδ<α/2\\delta<\\alpha/2ensures1<t<h1<t<h\. Finally, chooseε\\varepsilonwith0<ε<1​and​εr<2−tr;0<\\varepsilon<1\\,\\text\{and\}\\,\\varepsilon^\{r\}<2\-t^\{r\};both constraints are compatible becausetr<2t^\{r\}<2\.

Pick anyM\>hM\>h, and defined∈𝒟⁡\(𝒳\)d\\in\\mathcal\{D\}\(\\mathcal\{X\}\)by prescribing thepp\-th powers of its values:

d​\(1,2\)p=1,\\displaystyle d\(1,2\)^\{p\}=1,d​\(1,3\)p=h,\\displaystyle d\(1,3\)^\{p\}=h,d​\(1,4\)p=M,\\displaystyle d\(1,4\)^\{p\}=M,d​\(2,3\)p=ε,\\displaystyle d\(2,3\)^\{p\}=\\varepsilon,d​\(2,4\)p=t,\\displaystyle d\(2,4\)^\{p\}=t,d​\(3,4\)p=t\.\\displaystyle d\(3,4\)^\{p\}=t\.All six prescribed values are strictly positive, soddis a valid dissimilarity\. By using the orderingε<1<t<h<M,\\varepsilon<1<t<h<M,we can check by direct computation thatC1∈Tstable\(p\)​\(d\)C\_\{1\}\\in T\_\{\\mathrm\{stable\}\}^\{\(p\)\}\(d\)whereasC2∈Tstable\(q\)​\(d\)C\_\{2\}\\in T\_\{\\mathrm\{stable\}\}^\{\(q\)\}\(d\)\.

Suppose, for contradiction, thatTTrefines bothTstable\(p\)T\_\{\\mathrm\{stable\}\}^\{\(p\)\}andTstable\(q\)T\_\{\\mathrm\{stable\}\}^\{\(q\)\}\. ThenT⁡\(d\)T\(d\)contains bothC1C\_\{1\}andC2C\_\{2\}, contradicting laminarity\.

For a domain𝒳′\\mathcal\{X\}^\{\\prime\}with\|𝒳′\|\>4\|\\mathcal\{X\}^\{\\prime\}\|\>4, retain the above dissimilarities on\{1,2,3,4\}\\\{1,2,3,4\\\}and setd⁡\(w,i\)=Ld\(w,i\)=Lfor every new pointwwand everyi∈\{1,2,3,4\}i\\in\\\{1,2,3,4\\\}, whereLLis larger than every dissimilarity in the four\-point construction; assign arbitrary positive symmetric dissimilarities among the new points\. For any new external witnessww, every isolation weight contributing to the stability ofC1C\_\{1\}under powerppor ofC2C\_\{2\}under powerqqis strictly positive\. An average over an external set containing both old and new witnesses is therefore a positive average of the already verified terms and these new positive terms\. HenceC1∈Tstable\(p\)​\(d\)C\_\{1\}\\in T\_\{\\mathrm\{stable\}\}^\{\(p\)\}\(d\)andC2∈Tstable\(q\)​\(d\)C\_\{2\}\\in T\_\{\\mathrm\{stable\}\}^\{\(q\)\}\(d\)still hold on𝒳′\\mathcal\{X\}^\{\\prime\}, proving incompatibility for every\|𝒳′\|≥4\|\\mathcal\{X\}^\{\\prime\}\|\\geq 4\. ∎

###### Lemma 55\.

Letp\>0p\>0\. The methodsTSLT\_\{\\mathrm\{SL\}\}andTstable\(p\)T\_\{\\mathrm\{stable\}\}^\{\(p\)\}are incompatible\.

###### Proof\.

We first construct a dissimilarity functiond0d\_\{0\}on𝒳=\{1,2,3,4\}\\mathcal\{X\}=\\\{1,2,3,4\\\}such thatTSL​\(d0\)T\_\{\\mathrm\{SL\}\}\(d\_\{0\}\)andTstable\(p\)​\(d0\)T\_\{\\mathrm\{stable\}\}^\{\(p\)\}\(d\_\{0\}\)are incompatible\. Defined0∈𝒟⁡\(𝒳\)d\_\{0\}\\in\\mathcal\{D\}\(\\mathcal\{X\}\)such that

d0​\(1,2\)=3,d0​\(1,3\)=d0​\(1,4\)=10,d0​\(2,3\)=1,d0​\(2,4\)=d0​\(3,4\)=4\.d\_\{0\}\(1,2\)=3,\\quad d\_\{0\}\(1,3\)=d\_\{0\}\(1,4\)=10,\\quad d\_\{0\}\(2,3\)=1,\\quad d\_\{0\}\(2,4\)=d\_\{0\}\(3,4\)=4\.We can verify that\{1,2,3\}∈TSL​\(d0\)\\\{1,2,3\\\}\\in T\_\{\\mathrm\{SL\}\}\(d\_\{0\}\)while\{2,3,4\}∈Tstable​\(d0\)\.\\\{2,3,4\\\}\\in T\_\{\\mathrm\{stable\}\}\(d\_\{0\}\)\.The clusters\{1,2,3\}\\\{1,2,3\\\}and\{2,3,4\}\\\{2,3,4\\\}overlap, but neither contains the other\. ThusTSL​\(d0\)∪Tstable​\(d0\)T\_\{\\mathrm\{SL\}\}\(\{d\_\{0\}\}\)\\cup T\_\{\\mathrm\{stable\}\}\(\{d\_\{0\}\}\)is not laminar\.

Now fixp\>0p\>0and defineddentry\-wise byd⁡\(x,y\):=d0​\(x,y\)1/p\.d\(x,y\):=\{d\_\{0\}\}\(x,y\)^\{1/p\}\.Thenμp​\(d\)=d0\\mu^\{p\}\(d\)=\{d\_\{0\}\}\. Because single linkage is order invariant,TSL​\(d\)=TSL​\(d0\),T\_\{\\mathrm\{SL\}\}\(d\)=T\_\{\\mathrm\{SL\}\}\(\{d\_\{0\}\}\),whereas, by definition,Tstable\(p\)​\(d\)=Tstable​\(μp​\(d\)\)=Tstable​\(d0\)\.T\_\{\\mathrm\{stable\}\}^\{\(p\)\}\(d\)=T\_\{\\mathrm\{stable\}\}\(\\mu^\{p\}\(d\)\)=T\_\{\\mathrm\{stable\}\}\(\{d\_\{0\}\}\)\.Hence the same two non\-laminar clusters belong toTSL​\(d\)∪Tstable\(p\)​\(d\)\.T\_\{\\mathrm\{SL\}\}\(d\)\\cup T\_\{\\mathrm\{stable\}\}^\{\(p\)\}\(d\)\.

Forn\>4n\>4, extendd0\{d\_\{0\}\}by settingd0​\(i,j\)=20\{d\_\{0\}\}\(i,j\)=20whenever at least one ofi,ji,jbelongs to\{5,…,n\}\\\{5,\\ldots,n\\\}\. This does not affect the single\-linkage cluster\{1,2,3\}\\\{1,2,3\\\}\. Moreover, for every new pointzzand everyu,v∈\{2,3,4\}u,v\\in\\\{2,3,4\\\},ρd0​\(u​v∣z\)=20−d0​\(u,v\)\>0\.\\rho\_\{d\_\{0\}\}\(uv\\mid z\)=20\-\{d\_\{0\}\}\(u,v\)\>0\.Thusιd0​\(\{2,3,4\}\)\>0\\iota^\{d\_\{0\}\}\(\\\{2,3,4\\\}\)\>0remains valid\. Takingd⁡\(i,j\)=d0​\(i,j\)1/pd\(i,j\)=\{d\_\{0\}\}\(i,j\)^\{1/p\}completes the proof for everyn≥4n\\geq 4\. ∎

### D\.2Well\-separated Backbone Hierarchy

#### D\.2\.1Proof of Theorem[18](https://arxiv.org/html/2609.11173#Thmtheorem18)

The canonical partition consistency axiom is formulated at the level of partitions: if a partition already appears in the hierarchy, then strengthening that partition preserves all of its blocks\. To establish the existence of a well\-separated backbone hierarchy, however, we need a cluster\-wise consequence of this axiom: a cluster that is sufficiently well separated from its complement must be selected by any admissible method\.

We prove it in two steps\. First, we show in Lemma[56](https://arxiv.org/html/2609.11173#Thmtheorem56)that partition consistency yields a cluster\-wise consequence: if there exists a dissimilarityd0d\_\{0\}such that\{C,C¯\}⊆T⁡\(d0\)\\\{C,\\bar\{C\}\\\}\\subseteq T\(d\_\{0\}\), the clusterCCremains selected under any dissimilarity that does not increase dissimilarities withinCCand does not decrease dissimilarities fromCCtoC¯\\bar\{C\}\. Second, partition richness provides a reference dissimilarityd0d\_\{0\}realizing\{C,C¯\}\\\{C,\\bar\{C\}\\\}\. Moreover, by scale invariance, we can rescale any dissimilarityddsatisfyingϱ⁡\(d,C\)<ηC\\varrho\(d,C\)<\\eta\_\{C\}, for a sufficiently small constantηC\>0\\eta\_\{C\}\>0, by a factorβ\>0\\beta\>0so thatβ​d​\(x,y\)≤d0​\(x,y\)\\beta d\(x,y\)\\leq d\_\{0\}\(x,y\)for allx,y∈Cx,y\\in C, andβ​d​\(x,z\)≥d0​\(x,z\)\\beta d\(x,z\)\\geq d\_\{0\}\(x,z\)for allx∈C,z∈C¯\.x\\in C,\\ z\\in\\bar\{C\}\.Lemma[56](https://arxiv.org/html/2609.11173#Thmtheorem56)then givesC∈T⁡\(β​d\)C\\in T\(\\beta d\), and scale invariance yieldsC∈T⁡\(d\)C\\in T\(d\)\. This is the content of Lemma[57](https://arxiv.org/html/2609.11173#Thmtheorem57)\.

###### Lemma 56\.

LetTTbe a hierarchical method satisfying partition consistency, and letC⊊𝒳C\\subsetneq\\mathcal\{X\}be a nonempty proper cluster\. Letd0,d∈𝒟⁡\(𝒳\)d\_\{0\},d\\in\\mathcal\{D\}\(\\mathcal\{X\}\)such thatddis a\{C\}\\\{C\\\}\-strengthening ofd0d\_\{0\}\. If\{C,C¯\}⊆T⁡\(d0\)\\\{C,\\bar\{C\}\\\}\\subseteq T\(d\_\{0\}\)thenC∈T⁡\(d\)C\\in T\(d\)\.

###### Proof\.

Defined1∈𝒟⁡\(𝒳\)d\_\{1\}\\in\\mathcal\{D\}\(\\mathcal\{X\}\)by setting

d1\(x,y\):=min\{d\(x,y\),d0\(x,y\)\}whenx,y∈C¯,andd1\(x,y\):=d0\(x,y\)otherwise\.d\_\{1\}\(x,y\):=\\min\\\{d\(x,y\),d\_\{0\}\(x,y\)\\\}\\ \\text\{when \}x,y\\in\\bar\{C\},\\quad\\text\{and\}\\quad d\_\{1\}\(x,y\):=d\_\{0\}\(x,y\)\\ \\text\{otherwise\}\.Thend1d\_\{1\}is a\{C,C¯\}\\\{C,\\bar\{C\}\\\}\-strengthening ofd0d\_\{0\}\. Because\{C,C¯\}⊆T⁡\(d0\)\\\{C,\\bar\{C\}\\\}\\subseteq T\(d\_\{0\}\), partition consistency ensures that\{C,C¯\}⊆T⁡\(d1\)\\\{C,\\bar\{C\}\\\}\\subseteq T\(d\_\{1\}\)\. In particular, the partition𝒫C:=\{C\}∪\{\{x\}:x∈C¯\}\\mathcal\{P\}\_\{C\}:=\\\{C\\\}\\cup\\\{\\\{x\\\}:x\\in\\bar\{C\}\\\}satisfies𝒫C⊆T⁡\(d1\)\\mathcal\{P\}\_\{C\}\\subseteq T\(d\_\{1\}\), because every hierarchy contains all singleton clusters\.

We now prove thatddis a𝒫C\\mathcal\{P\}\_\{C\}\-strengthening ofd1d\_\{1\}\. Indeed, we have

- •d⁡\(x,y\)≤d0​\(x,y\)=d1​\(x,y\)d\(x,y\)\\leq d\_\{0\}\(x,y\)=d\_\{1\}\(x,y\)for everyx,y∈Cx,y\\in C,
- •d⁡\(x,y\)≥d0​\(x,y\)=d1​\(x,y\)d\(x,y\)\\geq d\_\{0\}\(x,y\)=d\_\{1\}\(x,y\)for everyx∈Cx\\in C,y∈C¯y\\in\\bar\{C\},
- •d\(x,y\)≥min\{\(d\(x,y\),d0\(x,y\)\}=d1\(x,y\)d\(x,y\)\\geq\\min\\left\\\{\(d\(x,y\),d\_\{0\}\(x,y\)\\right\\\}=d\_\{1\}\(x,y\)for everyx,y∈C¯x,y\\in\\bar\{C\}\.

Thus, partition consistency ensures𝒫C⊆T⁡\(d\)\\mathcal\{P\}\_\{C\}\\subseteq T\(d\), and henceC∈T⁡\(d\)C\\in T\(d\)\. ∎

The previous lemma reduces the problem to finding a rescaling of a given dissimilarityddwhich is dominated byd0d\_\{0\}insideCC, but which dominatesd0d\_\{0\}across the cut\(C,C¯\)\(C,\\bar\{C\}\)\. Such a rescaling exists whenever the internal diameter ofCCunderddis sufficiently small compared to its distance from the complement\.

###### Lemma 57\.

LetTTsatisfy scale invariance, partition richness, and partition consistency\. For any clusterC⊊𝒳C\\subsetneq\\mathcal\{X\}, with\|C\|≥2\|C\|\\geq 2, there existsηC\>0\\eta\_\{C\}\>0such that, for everyd∈𝒟⁡\(𝒳\)d\\in\\mathcal\{D\}\(\\mathcal\{X\}\),

ϱ⁡\(d,C\)<ηC⟹C∈T⁡\(d\)\.\\varrho\(d,C\)<\\eta\_\{C\}\\quad\\Longrightarrow\\quad C\\in T\(d\)\.IfTTfurther satisfies permutation invariance,ηC\\eta\_\{C\}only depends on the size\|C\|\|C\|andηC∈\(0,1\]\\eta\_\{C\}\\in\(0,1\]\.

Applying Lemma[57](https://arxiv.org/html/2609.11173#Thmtheorem57)to an admissible methodTT, define the separation margin sequence𝜼=\(ηs\)1≤s≤n−2\\boldsymbol\{\\eta\}=\(\\eta\_\{s\}\)\_\{1\\leq s\\leq n\-2\}as provided by Lemma[57](https://arxiv.org/html/2609.11173#Thmtheorem57)\. Then every non\-singleton proper clusterC∈Tglob𝜼​\(d\)C\\in T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}\(d\)satisfiesϱ⁡\(d,C\)<η\|C\|−1,\\varrho\(d,C\)<\\eta\_\{\|C\|\-1\},and therefore belongs toT⁡\(d\)T\(d\)\. Because both hierarchies contain the root and all singleton clusters,Tglob𝜼​\(d\)⊆T⁡\(d\)T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}\(d\)\\subseteq T\(d\)for everyd∈𝒟⁡\(𝒳\)d\\in\\mathcal\{D\}\(\\mathcal\{X\}\), and henceTglob𝜼⊑T\.T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}\\sqsubseteq T\.This proves Theorem[18](https://arxiv.org/html/2609.11173#Thmtheorem18)\.

###### Proof of Lemma[57](https://arxiv.org/html/2609.11173#Thmtheorem57)\.

By partition richness, choosed0∈𝒟⁡\(𝒳\)d\_\{0\}\\in\\mathcal\{D\}\(\\mathcal\{X\}\)such that\{C,C¯\}⊆T⁡\(d0\)\\\{C,\\bar\{C\}\\\}\\subseteq T\(d\_\{0\}\)\. Definea:=minx,y∈Cx≠y⁡d0​\(x,y\)a:=\\min\_\{\\begin\{subarray\}\{c\}x,y\\in C\\\\ x\\neq y\\end\{subarray\}\}d\_\{0\}\(x,y\)andb:=maxx∈C,z∈C¯⁡d0​\(x,z\)\.b:=\\max\_\{x\\in C,\\ z\\in\\bar\{C\}\}d\_\{0\}\(x,z\)\.Bothaaandbbare positive\. SetηC:=ab\.\\eta\_\{C\}:=\\frac\{a\}\{b\}\.

Now letddbe a dissimilarity satisfyingϱ⁡\(d,C\)<ηC\\varrho\(d,C\)<\\eta\_\{C\}\. WriteM:=maxx,y∈C⁡d⁡\(x,y\)M:=\\max\_\{x,y\\in C\}d\(x,y\)andm:=minx∈C,z∈C¯⁡d⁡\(x,z\)m:=\\min\_\{x\\in C,\\ z\\in\\bar\{C\}\}d\(x,z\)\. ThenMm=ϱ⁡\(d,C\)<ab\\frac\{M\}\{m\}=\\varrho\(d,C\)<\\frac\{a\}\{b\}, and hencebm<aM\\frac\{b\}\{m\}<\\frac\{a\}\{M\}\. Next, chooseβ\>0\\beta\>0such thatbm<β<aM\\frac\{b\}\{m\}<\\beta<\\frac\{a\}\{M\}\. This choice ensures that,

β​d​\(x,y\)≤β​M<a≤d0​\(x,y\),for all​x,y∈C,and\\beta d\(x,y\)\\leq\\beta M<a\\leq d\_\{0\}\(x,y\),\\quad\\text\{for all \}x,y\\in C,\\text\{ and\}β​d​\(x,z\)≥β​m\>b≥d0​\(x,z\)for all​x∈C,z∈C¯\.\\beta d\(x,z\)\\geq\\beta m\>b\\geq d\_\{0\}\(x,z\)\\quad\\text\{for all \}x\\in C,z\\in\\bar\{C\}\.Thus, by Lemma[56](https://arxiv.org/html/2609.11173#Thmtheorem56), applied toβ​d\\beta d, we getC∈T⁡\(β​d\)\.C\\in T\(\\beta d\)\.Finally, scale invariance givesT⁡\(β​d\)=T⁡\(d\)T\(\\beta d\)=T\(d\), henceC∈T⁡\(d\)C\\in T\(d\)\.

Assume now thatTTis permutation invariant\. For every non\-singleton proper clusterC⊊𝒳C\\subsetneq\\mathcal\{X\}, define

ΓT​\(C\):=supd0∈𝒟⁡\(𝒳\)\{C,C¯\}⊆T⁡\(d0\)minx,y∈Cx≠y⁡d0​\(x,y\)maxx∈C,z∈C¯⁡d0​\(x,z\)\.\\Gamma\_\{T\}\(C\):=\\sup\_\{\\begin\{subarray\}\{c\}d\_\{0\}\\in\\mathcal\{D\}\(\\mathcal\{X\}\)\\\\ \\\{C,\\bar\{C\}\\\}\\subseteq T\(d\_\{0\}\)\\end\{subarray\}\}\\frac\{\\min\_\{\\begin\{subarray\}\{c\}x,y\\in C\\\\ x\\neq y\\end\{subarray\}\}d\_\{0\}\(x,y\)\}\{\\max\_\{x\\in C,\\ z\\in\\bar\{C\}\}d\_\{0\}\(x,z\)\}\.By the previous paragraph, for everyd∈𝒟⁡\(𝒳\)d\\in\\mathcal\{D\}\(\\mathcal\{X\}\), we have:ϱ⁡\(d,C\)<ΓT​\(C\)⟹C∈T⁡\(d\)\.\\varrho\(d,C\)<\\Gamma\_\{T\}\(C\)\\Longrightarrow C\\in T\(d\)\.

We now prove thatΓT​\(C\)\\Gamma\_\{T\}\(C\)depends only on the cardinality ofCC\. LetC,C′⊊𝒳C,C^\{\\prime\}\\subsetneq\\mathcal\{X\}be two nontrivial subsets such that\|C\|=\|C′\|\|C\|=\|C^\{\\prime\}\|\. Because𝒳\\mathcal\{X\}is finite, there exists a permutationϕ∈Π⁡\(𝒳\)\\phi\\in\\Pi\(\\mathcal\{X\}\)such thatϕ⁡\(C\)=C′\.\\phi\(C\)=C^\{\\prime\}\.Because permutation invariance ofTTimpliesΓT​\(ϕ⁡\(C\)\)=ΓT​\(C\)\\Gamma\_\{T\}\(\\phi\(C\)\)=\\Gamma\_\{T\}\(C\), we haveΓT​\(C′\)=ΓT​\(C\)\.\\Gamma\_\{T\}\(C^\{\\prime\}\)=\\Gamma\_\{T\}\(C\)\.HenceΓT​\(C\)\\Gamma\_\{T\}\(C\)is constant over all subsets of𝒳\\mathcal\{X\}having the same cardinality, and thus depends only on\|C\|\|C\|\. Thus, for1≤s≤\|𝒳\|−21\\leq s\\leq\|\\mathcal\{X\}\|\-2, we may define

ηs:=ΓT​\(C\)\\eta\_\{s\}:=\\Gamma\_\{T\}\(C\)for any subsetC⊆𝒳C\\subseteq\\mathcal\{X\}withs=\|C\|−1s=\|C\|\-1\. This is well\-defined by the preceding argument, and, for everyC⊆𝒳C\\subseteq\\mathcal\{X\},

ϱ⁡\(d,C\)<η\|C\|−1⟹C∈T⁡\(d\)\.\\varrho\(d,C\)<\\eta\_\{\|C\|\-1\}\\quad\\Longrightarrow\\quad C\\in T\(d\)\.
It remains to prove that, for1≤s≤\|𝒳\|−21\\leq s\\leq\|\\mathcal\{X\}\|\-2, one has0<ηs≤1\.0<\\eta\_\{s\}\\leq 1\.Observe thatηs\>0\\eta\_\{s\}\>0follows from partition richness, as for eachCCthere existsd0d\_\{0\}such that\{C,C¯\}⊆T⁡\(d0\)\\\{C,\\bar\{C\}\\\}\\subseteq T\(d\_\{0\}\), and the corresponding ratio is strictly positive\.

To establishηs≤1\\eta\_\{s\}\\leq 1, suppose by contradiction thatηs\>1\\eta\_\{s\}\>1for some1≤s≤\|𝒳\|−21\\leq s\\leq\|\\mathcal\{X\}\|\-2\. Letddbe the uniform dissimilarity on𝒳\\mathcal\{X\}, namelyd⁡\(x,y\)=1d\(x,y\)=1for allx≠yx\\neq y\. Thenϱ⁡\(d,C\)=1<ηs\\varrho\(d,C\)=1<\\eta\_\{s\}for every subsetC⊊𝒳C\\subsetneq\\mathcal\{X\}of sizes\+1s\+1\. Hence every suchCCbelongs toT⁡\(d\)T\(d\)\. But there exist two subsets of𝒳\\mathcal\{X\}of sizes\+1s\+1that overlap without either containing the other\. For instance, since1≤s≤n−21\\leq s\\leq n\-2, the setsC1:=\{1,…,s\+1\}C\_\{1\}:=\\\{1,\\ldots,s\+1\\\}andC2:=\{2,…,s\+2\}C\_\{2\}:=\\\{2,\\ldots,s\+2\\\}both have cardinalitys\+1s\+1, overlap nontrivially, and neither contains the other\. This contradicts the laminarity ofT⁡\(d\)T\(d\)\. Thereforeηs≤1\\eta\_\{s\}\\leq 1\. ∎

### D\.3Additional Lemmas

###### Lemma 58\.

We haveαuhr∧αpc⊳αbh1\\alpha\_\{\\mathrm\{uhr\}\}\\wedge\\alpha\_\{\\mathrm\{pc\}\}\\triangleright\\alpha\_\{\\mathrm\{bh\}\}^\{1\}\.

###### Proof\.

LetT∈ℋαuhr∧αpc​\(𝒳\)\.T\\in\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{uhr\}\}\\wedge\\alpha\_\{\\mathrm\{pc\}\}\}\(\\mathcal\{X\}\)\.We prove thatTglob⊑TT\_\{\\mathrm\{glob\}\}\\sqsubseteq T\. Fixd∈𝒟⁡\(𝒳\)d\\in\\mathcal\{D\}\(\\mathcal\{X\}\)andC∈Tglob​\(d\)C\\in T\_\{\\mathrm\{glob\}\}\(d\)\. The casesC=𝒳C=\\mathcal\{X\}and\|C\|=1\{\\lvert C\\rvert\}=1are immediate because every hierarchy contains the root and all singleton clusters\. We therefore assume thatC⊊𝒳C\\subsetneq\\mathcal\{X\}and\|C\|≥2\{\\lvert C\\rvert\}\\geq 2\. Defineα:=maxx,y∈Cx≠y⁡d⁡\(x,y\)\\alpha:=\\max\_\{\\begin\{subarray\}\{c\}x,y\\in C\\\\ x\\neq y\\end\{subarray\}\}d\(x,y\)andβ:=minx∈Cz∈C¯⁡d⁡\(x,z\)\.\\beta:=\\min\_\{\\begin\{subarray\}\{c\}x\\in C\\\\ z\\in\\bar\{C\}\\end\{subarray\}\}d\(x,z\)\.BecauseC∈Tglob​\(d\)C\\in T\_\{\\mathrm\{glob\}\}\(d\), we haveα<β\\alpha<\\beta\. Chooseα<β~<β\\alpha<\\widetilde\{\\beta\}<\\betaand setγ:=min\(\{β~\}∪\{d\(z,w\):z,w∈C¯,z≠w\}\)\.\\gamma:=\\min\\left\(\\\{\\widetilde\{\\beta\}\\\}\\cup\\\{d\(z,w\):z,w\\in\\bar\{C\},\\ z\\neq w\\\}\\right\)\.In particular,0<γ≤β~0<\\gamma\\leq\\widetilde\{\\beta\}\.

Defineu∈𝒟⁡\(𝒳\)u\\in\\mathcal\{D\}\(\\mathcal\{X\}\)byu⁡\(x,x\):=0u\(x,x\):=0and, for distinctx,y∈𝒳x,y\\in\\mathcal\{X\},

u⁡\(x,y\):=\{α,x,y∈C,β~,x∈C,y∈C¯ory∈C,x∈C¯,γ,x,y∈C¯\.u\(x,y\):=\\begin\{cases\}\\alpha,&x,y\\in C,\\\\\[2\.84526pt\] \\widetilde\{\\beta\},&x\\in C,\\ y\\in\\bar\{C\}\\text\{ or \}y\\in C,\\ x\\in\\bar\{C\},\\\\\[2\.84526pt\] \\gamma,&x,y\\in\\bar\{C\}\.\\end\{cases\}Becauseα<β~\\alpha<\\widetilde\{\\beta\}andγ≤β~\\gamma\\leq\\widetilde\{\\beta\},uuis an ultrametric: every triangle meeting bothCCandC¯\\bar\{C\}has two dissimilarities equal toβ~\\widetilde\{\\beta\}, while the third is eitherα\\alphaorγ\\gamma, and triangles contained entirely inCCor inC¯\\bar\{C\}are equilateral\.

Recall the ultrametric\-ball representation in Equation \([3](https://arxiv.org/html/2609.11173#S5.E3)\)\. For everyx∈Cx\\in C,Bu​\(x,α\)=CB\_\{u\}\(x,\\alpha\)=Cbecauseu⁡\(x,y\)≤αu\(x,y\)\\leq\\alphafory∈Cy\\in C, whereasu⁡\(x,z\)=β~\>αu\(x,z\)=\\widetilde\{\\beta\}\>\\alphaforz∈C¯z\\in\\bar\{C\}\. HenceC∈ΨuC\\in\\Psi\_\{u\}\. AsTTsatisfies refinement on ultrametrics,C∈Ψu⊆T⁡\(u\)\.C\\in\\Psi\_\{u\}\\subseteq T\(u\)\.

Now consider the partition𝒫C:=\{C\}∪\{\{z\}:z∈C¯\}\.\\mathcal\{P\}\_\{C\}:=\\\{C\\\}\\cup\\bigl\\\{\\\{z\\\}:z\\in\\bar\{C\}\\bigr\\\}\.Because every hierarchy contains all singleton clusters, we have𝒫C⊆T⁡\(u\)\.\\mathcal\{P\}\_\{C\}\\subseteq T\(u\)\.We claim thatddis a𝒫C\\mathcal\{P\}\_\{C\}\-strengthening ofuu\. Indeed, for distinctx,y∈Cx,y\\in C,d⁡\(x,y\)≤α=u⁡\(x,y\)\.d\(x,y\)\\leq\\alpha=u\(x,y\)\.Forx∈Cx\\in Candz∈C¯z\\in\\bar\{C\},d⁡\(x,z\)≥β\>β~=u⁡\(x,z\)\.d\(x,z\)\\geq\\beta\>\\widetilde\{\\beta\}=u\(x,z\)\.Finally, for distinctz,w∈C¯z,w\\in\\bar\{C\}, the pointszzandwwbelong to different singleton blocks of𝒫C\\mathcal\{P\}\_\{C\}, and the definition ofγ\\gammagivesd⁡\(z,w\)≥γ=u⁡\(z,w\)\.d\(z,w\)\\geq\\gamma=u\(z,w\)\.

Partition consistency therefore yields𝒫C⊆T⁡\(d\),\\mathcal\{P\}\_\{C\}\\subseteq T\(d\),and henceC∈T⁡\(d\)C\\in T\(d\)\. Becaused∈𝒟⁡\(𝒳\)d\\in\\mathcal\{D\}\(\\mathcal\{X\}\)andC∈Tglob​\(d\)C\\in T\_\{\\mathrm\{glob\}\}\(d\)were arbitrary, we conclude thatTglob⊆T\.T\_\{\\mathrm\{glob\}\}\\subseteq T\.∎

###### Lemma 59\.

We haveαadm∧αord⊳αbh1,\\alpha\_\{\\mathrm\{adm\}\}\\wedge\\alpha\_\{\\mathrm\{ord\}\}\\triangleright\\alpha\_\{\\mathrm\{bh\}\}^\{1\},i\.e\., for anyT∈ℋαadm∧αordT\\in\\mathcal\{H\}\_\{\\alpha\_\{\\mathrm\{adm\}\}\\wedge\\alpha\_\{\\mathrm\{ord\}\}\},Tglob⊑TT\_\{\\mathrm\{glob\}\}\\sqsubseteq Tholds\.

###### Proof\.

LetTTbe admissible and order invariant\. By Theorem[18](https://arxiv.org/html/2609.11173#Thmtheorem18), there exists a separation margin sequence𝜼=\(ηs\)1≤s≤n−2\\boldsymbol\{\\eta\}=\(\\eta\_\{s\}\)\_\{1\\leq s\\leq n\-2\}such thatTglob𝜼⊑T\.T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}\\sqsubseteq T\.Setηmin:=min1≤s≤n−2⁡ηs\>0\.\\eta\_\{\\min\}:=\\min\_\{1\\leq s\\leq n\-2\}\\eta\_\{s\}\>0\.

Fixd∈𝒟⁡\(𝒳\)d\\in\\mathcal\{D\}\(\\mathcal\{X\}\)\. We construct a strictly increasing functiong:R≥0→R≥0g:\\mathbb\{R\}\_\{\\geq 0\}\\to\\mathbb\{R\}\_\{\\geq 0\}, withg⁡\(0\)=0g\(0\)=0, such thatTglob​\(d\)⊆Tglob𝜼​\(g∘d\)\.T\_\{\\mathrm\{glob\}\}\(d\)\\subseteq T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}\(g\\circ d\)\.

Let0<δ1<δ2<⋯<δm0<\\delta\_\{1\}<\\delta\_\{2\}<\\cdots<\\delta\_\{m\}be the distinct positive values taken bydd\. Chooseq\>1ηmin,q\>\\frac\{1\}\{\\eta\_\{\\min\}\},and prescribeg⁡\(0\):=0g\(0\):=0andg⁡\(δi\):=qig\(\\delta\_\{i\}\):=q^\{i\}for alli∈\[m\]\.i\\in\[m\]\.Because the sequences\(δi\)\(\\delta\_\{i\}\)and\(qi\)\(q^\{i\}\)are strictly increasing, these prescribed values can be extended to a strictly increasing functiong:R≥0→R≥0g:\\mathbb\{R\}\_\{\\geq 0\}\\to\\mathbb\{R\}\_\{\\geq 0\}withg⁡\(0\)=0g\(0\)=0; for instance, using piecewise\-linear interpolation on\[0,δm\]\[0,\\delta\_\{m\}\]and extending linearly with positive slope beyondδm\\delta\_\{m\}\.

Now letC∈Tglob​\(d\)C\\in T\_\{\\mathrm\{glob\}\}\(d\)be a nontrivial proper cluster\. By definition,d⁡\(x1,y\)<d⁡\(x2,z\)d\(x\_\{1\},y\)<d\(x\_\{2\},z\)for everyx1,x2,y∈Cx\_\{1\},x\_\{2\},y\\in Candz∉Cz\\notin Cwheneverd⁡\(x1,y\)\>0d\(x\_\{1\},y\)\>0\. Hence, ifd⁡\(x1,y\)=δid\(x\_\{1\},y\)=\\delta\_\{i\}andd⁡\(x2,z\)=δj,d\(x\_\{2\},z\)=\\delta\_\{j\},theni<ji<j, and therefore

g⁡\(d⁡\(x1,y\)\)g⁡\(d⁡\(x2,z\)\)=qi−j≤1q<ηmin≤η\|C\|−1\.\\frac\{g\(d\(x\_\{1\},y\)\)\}\{g\(d\(x\_\{2\},z\)\)\}\\ =\\ q^\{i\-j\}\\ \\leq\\ \\frac\{1\}\{q\}\\ <\\ \\eta\_\{\\min\}\\ \\leq\\ \\eta\_\{\|C\|\-1\}\.Ifd⁡\(x1,y\)=0d\(x\_\{1\},y\)=0, the same inequality holds trivially becauseg⁡\(0\)=0g\(0\)=0\.

Therefore,ϱ⁡\(g∘d,C\)<η\|C\|−1\\varrho\(g\\circ d,C\)<\\eta\_\{\|C\|\-1\}, and henceC∈Tglob𝜼​\(g∘d\)\.C\\in T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}\(g\\circ d\)\.The root and singleton clusters belong to both hierarchies by definition, soTglob​\(d\)⊆Tglob𝜼​\(g∘d\)\.T\_\{\\mathrm\{glob\}\}\(d\)\\subseteq T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}\(g\\circ d\)\.UsingTglob𝜼⊑TT\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}\\sqsubseteq T, we further obtain

Tglob​\(d\)⊆Tglob𝜼​\(g∘d\)⊆T⁡\(g∘d\)\.T\_\{\\mathrm\{glob\}\}\(d\)\\subseteq T\_\{\\mathrm\{glob\}\}^\{\\boldsymbol\{\\eta\}\}\(g\\circ d\)\\subseteq T\(g\\circ d\)\.Finally, order invariance ofTTgivesT⁡\(g∘d\)=T⁡\(d\),T\(g\\circ d\)=T\(d\),and thereforeTglob​\(d\)⊆T⁡\(d\)\.T\_\{\\mathrm\{glob\}\}\(d\)\\subseteq T\(d\)\.Becaused∈𝒟⁡\(𝒳\)d\\in\\mathcal\{D\}\(\\mathcal\{X\}\)was arbitrary,Tglob⊑T\.T\_\{\\mathrm\{glob\}\}\\sqsubseteq T\.∎

## Appendix EEmpirical Size of Backbone Hierarchy

We quantify the size of the backbone hierarchy across four standard scikit\-learn datasets\([Pedregosa et al\., 2011](https://arxiv.org/html/2609.11173#bib.bib27)\): Iris, Wine, Breast Cancer, and Digits\. We remove exact duplicate feature vectors, standardize each feature, and use Euclidean dissimilarities\. We count only nontrivial clustersCCsatisfying1<\|C\|<n1<\|C\|<n, and report the ratio of the number returned byTglobT\_\{\\mathrm\{glob\}\}to the number returned byTSLT\_\{\\mathrm\{SL\}\}\.

Table 3:Size ofTglobT\_\{\\mathrm\{glob\}\}relative to the non\-binary single\-linkage hierarchyTSLT\_\{\\mathrm\{SL\}\}\.Across these datasets,TglobT\_\{\\mathrm\{glob\}\}contains approximately19%19\\%–29%29\\%of the nontrivial clusters inTSLT\_\{\\mathrm\{SL\}\}\(23\.0%23\.0\\%in aggregate\)\. Thus, the backbone is nontrivial on these examples\.

## References

- M\. Ackerman, S\. Ben\-David, and D\. LokerCharacterization of linkage\-based clustering\.\.InCOLT,Vol\.2010,pp\. 270–281\.Cited by:[§1](https://arxiv.org/html/2609.11173#S1.p9.1),[§6](https://arxiv.org/html/2609.11173#S6.p4.1)\.
- Ackerman and Ben\-David \(2016\)M\. Ackerman and S\. Ben\-DavidA characterization of linkage\-based hierarchical clustering\.Journal of Machine Learning Research17\(231\),pp\. 1–17\.Cited by:[§1](https://arxiv.org/html/2609.11173#S1.p10.1),[§1](https://arxiv.org/html/2609.11173#S1.p9.1),[§6](https://arxiv.org/html/2609.11173#S6.p4.1)\.
- Ackerman and Dasgupta \(2014\)M\. Ackerman and S\. DasguptaIncremental clustering: the case for extra clusters\.Advances in Neural Information Processing Systems27\.Cited by:[§3\.2](https://arxiv.org/html/2609.11173#S3.SS2.p3.1)\.
- Apresjan \(1966\)J\. D\. ApresjanAn algorithm for constructing clusters from a distance matrix\.Mashinnyi perevod: prikladnaja lingvistika9,pp\. 3–18\.Cited by:[§3\.2](https://arxiv.org/html/2609.11173#S3.SS2.p3.1)\.
- Arias\-Castro and Coda \(2025\)E\. Arias\-Castro and E\. CodaAn axiomatic definition of hierarchical clustering\.Journal of Machine Learning Research26\(10\),pp\. 1–26\.Cited by:[§1](https://arxiv.org/html/2609.11173#S1.p9.1),[§6](https://arxiv.org/html/2609.11173#S6.p8.1),[§6](https://arxiv.org/html/2609.11173#S6.p9.1)\.
- Balcanet al\.\(2008\)M\. Balcan, A\. Blum, and S\. VempalaA discriminative framework for clustering via similarity functions\.InProceedings of the Fortieth annual ACM Symposium on Theory of Computing,pp\. 671–680\.Cited by:[§C\.3\.1](https://arxiv.org/html/2609.11173#A3.SS3.SSS1.p1.1.1),[§3\.2](https://arxiv.org/html/2609.11173#S3.SS2.p3.1),[Example 15](https://arxiv.org/html/2609.11173#Thmtheorem15.p1.2.1)\.
- Ben\-David and Ackerman \(2008\)S\. Ben\-David and M\. AckermanMeasures of clustering quality: a working set of axioms for clustering\.Advances in Neural Information Processing Systems21\.Cited by:[§1](https://arxiv.org/html/2609.11173#S1.p2.1),[§1](https://arxiv.org/html/2609.11173#S1.p9.1),[§6](https://arxiv.org/html/2609.11173#S6.p1.1),[§6](https://arxiv.org/html/2609.11173#S6.p2.1)\.
- Bryant and Berry \(2001\)D\. Bryant and V\. BerryA structured family of clustering and tree construction methods\.Advances in Applied Mathematics27\(4\),pp\. 705–732\.Cited by:[§C\.4](https://arxiv.org/html/2609.11173#A3.SS4.p2.1.1),[§3\.2](https://arxiv.org/html/2609.11173#S3.SS2.p3.1),[§3\.3](https://arxiv.org/html/2609.11173#S3.SS3.p1.1),[§3\.3](https://arxiv.org/html/2609.11173#S3.SS3.p2.5),[Example 15](https://arxiv.org/html/2609.11173#Thmtheorem15.p1.2.1),[footnote 6](https://arxiv.org/html/2609.11173#footnote6)\.
- Carlsson and Mémoli \(2010\)G\. E\. Carlsson and F\. MémoliCharacterization, stability and convergence of hierarchical clustering methods\.\.Journal of Machine Learning Research11\(Apr\),pp\. 1425–1470\.Cited by:[§C\.1\.1](https://arxiv.org/html/2609.11173#A3.SS1.SSS1.p3.1),[§1](https://arxiv.org/html/2609.11173#S1.p10.1),[§1](https://arxiv.org/html/2609.11173#S1.p9.1),[§5\.1](https://arxiv.org/html/2609.11173#S5.SS1.p2.1),[§5\.2\.2](https://arxiv.org/html/2609.11173#S5.SS2.SSS2.p2.2),[§6](https://arxiv.org/html/2609.11173#S6.p3.1)\.
- Carlssonet al\.\(2014\)G\. Carlsson, F\. Mémoli, A\. Ribeiro, and S\. SegarraHierarchical quasi\-clustering methods for asymmetric networks\.InInternational Conference on Machine Learning,pp\. 352–360\.Cited by:[§6](https://arxiv.org/html/2609.11173#S6.p5.1)\.
- Carlssonet al\.\(2017\)G\. Carlsson, F\. Mémoli, A\. Ribeiro, and S\. SegarraAdmissible hierarchical clustering methods and algorithms for asymmetric networks\.IEEE Transactions on Signal and Information Processing over Networks3\(4\),pp\. 711–727\.Cited by:[§6](https://arxiv.org/html/2609.11173#S6.p5.1)\.
- Carlsson and Mémoli \(2013\)G\. Carlsson and F\. MémoliClassifying clustering schemes\.Foundations of Computational Mathematics13\(2\),pp\. 221–252\.Cited by:[§6](https://arxiv.org/html/2609.11173#S6.p2.1)\.
- Cohen\-Addadet al\.\(2019\)V\. Cohen\-Addad, V\. Kanade, F\. Mallmann\-Trenn, and C\. MathieuHierarchical clustering: objective functions and algorithms\.Journal of the ACM \(JACM\)66\(4\),pp\. 1–42\.Cited by:[§6](https://arxiv.org/html/2609.11173#S6.p7.1)\.
- Cohen\-Addadet al\.\(2018\)V\. Cohen\-Addad, V\. Kanade, and F\. Mallmann\-TrennClustering redemption–beyond the impossibility of Kleinberg’s axioms\.Advances in Neural Information Processing Systems31\.Cited by:[§1](https://arxiv.org/html/2609.11173#S1.p9.1),[§6](https://arxiv.org/html/2609.11173#S6.p2.1)\.
- Dasgupta \(2016\)S\. DasguptaA cost function for similarity\-based hierarchical clustering\.InProceedings of the Forty\-Eighth Annual ACM Symposium on Theory of Computing,STOC ’16,New York, NY, USA,pp\. 118–127\.External Links:ISBN 9781450341325Cited by:[§6](https://arxiv.org/html/2609.11173#S6.p7.1)\.
- Diatta and Fichet \(1994\)J\. Diatta and B\. FichetFrom Apresjan hierarchies and Bandelt\-Dress weak hierarchies to quasi\-hierarchies\.InNew approaches in classification and data analysis,pp\. 111–118\.Cited by:[§3\.2](https://arxiv.org/html/2609.11173#S3.SS2.p3.1)\.
- Drevetonet al\.\(2025\)M\. Dreveton, M\. Grossglauser, D\. Kuroda, and P\. ThiranHierarchical linkage clustering beyond binary trees and ultrametrics\.External Links:2511\.18056,[Link](https://arxiv.org/abs/2511.18056)Cited by:[§C\.3\.1](https://arxiv.org/html/2609.11173#A3.SS3.SSS1.p1.1.1),[Example 15](https://arxiv.org/html/2609.11173#Thmtheorem15.p1.2.1)\.
- Kleinberg \(2002\)J\. KleinbergAn impossibility theorem for clustering\.Advances in Neural Information Processing Systems15\.Cited by:[§1](https://arxiv.org/html/2609.11173#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.11173#S2.SS1.p1.1),[§6](https://arxiv.org/html/2609.11173#S6.p1.1),[Theorem 3](https://arxiv.org/html/2609.11173#Thmtheorem3),[footnote 3](https://arxiv.org/html/2609.11173#footnote3)\.
- Murtagh and Contreras \(2017\)F\. Murtagh and P\. ContrerasAlgorithms for hierarchical clustering: an overview, II\.Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery7\(6\),pp\. e1219\.Cited by:[§C\.2\.1](https://arxiv.org/html/2609.11173#A3.SS2.SSS1.p2.1)\.
- Pedregosaet al\.\(2011\)F\. Pedregosa, G\. Varoquaux, A\. Gramfort, V\. Michel, B\. Thirion, O\. Grisel, M\. Blondel, P\. Prettenhofer, R\. Weiss, V\. Dubourg, J\. Vanderplas, A\. Passos, D\. Cournapeau, M\. Brucher, M\. Perrot, and E\. DuchesnayScikit\-learn: machine learning in Python\.Journal of Machine Learning Research12,pp\. 2825–2830\.Cited by:[Appendix E](https://arxiv.org/html/2609.11173#A5.p1.1)\.
- Puzichaet al\.\(2000\)J\. Puzicha, T\. Hofmann, and J\. M\. BuhmannA theory of proximity based clustering: structure detection by optimization\.Pattern Recognition33\(4\),pp\. 617–634\.Cited by:[§6](https://arxiv.org/html/2609.11173#S6.p1.1)\.
- Shen and Vereshchagin \(2002\)A\. Shen and N\. K\. VereshchaginBasic set theory\.American Mathematical Society\.Cited by:[§4\.3\.2](https://arxiv.org/html/2609.11173#S4.SS3.SSS2.p2.1.1)\.
- Strazzeri and Sánchez\-García \(2022\)F\. Strazzeri and R\. J\. Sánchez\-GarcíaPossibility results for graph clustering: a novel consistency axiom\.Pattern Recognition128,pp\. 108687\.Cited by:[§1](https://arxiv.org/html/2609.11173#S1.p2.1),[§6](https://arxiv.org/html/2609.11173#S6.p1.1),[§6](https://arxiv.org/html/2609.11173#S6.p2.1)\.
- Thomannet al\.\(2015\)P\. Thomann, I\. Steinwart, and N\. SchmidTowards an axiomatic approach to hierarchical clustering of measures\.Journal of Machine Learning Research16\(1\),pp\. 1949–2002\.Cited by:[§1](https://arxiv.org/html/2609.11173#S1.p9.1),[§6](https://arxiv.org/html/2609.11173#S6.p8.1),[§6](https://arxiv.org/html/2609.11173#S6.p9.1)\.
- Van Laarhoven and Marchiori \(2014\)T\. Van Laarhoven and E\. MarchioriAxioms for graph clustering quality functions\.Journal of Machine Learning Research15\(1\),pp\. 193–215\.Cited by:[§6](https://arxiv.org/html/2609.11173#S6.p2.1)\.
- Willson and Warnow \(2024\)J\. Willson and T\. WarnowAxioms for clustering simple unweighted graphs: no impossibility result\.PLOS Complex Systems1\(2\),pp\. e0000011\.Cited by:[§1](https://arxiv.org/html/2609.11173#S1.p2.1),[§1](https://arxiv.org/html/2609.11173#S1.p9.1),[§6](https://arxiv.org/html/2609.11173#S6.p1.1),[§6](https://arxiv.org/html/2609.11173#S6.p2.1)\.
- Zadeh and Ben\-David \(2009\)R\. B\. Zadeh and S\. Ben\-DavidA uniqueness theorem for clustering\.InProceedings of the Twenty\-Fifth Conference on Uncertainty in Artificial Intelligence,pp\. 639–646\.Cited by:[§1](https://arxiv.org/html/2609.11173#S1.p2.1),[§6](https://arxiv.org/html/2609.11173#S6.p1.1),[§6](https://arxiv.org/html/2609.11173#S6.p2.1)\.

Similar Articles

RHEA: Reliability-Harmonized Reconstruction and Assignment for Robust Multimodal-Attributed Graph Clustering

arXiv cs.LG

This paper proposes RHEA, a reliability-aware framework for multimodal-attributed graph clustering that estimates node-specific modality reliability from neighborhood consensus, reconstructs unreliable modalities, and uses reliability-aware fusion and optimal transport clustering. Experiments on four benchmarks show consistent gains, especially under noisy or missing attributes.

Hierarchical Domain Generalization

arXiv cs.LG

This paper introduces hierarchical domain generalization, formalizing extrapolation from finite observed regions to an entire instance space. It shows that no matter how simple the hypothesis class, certain domain partitions make generalization impossible, arguing that modern generalization theory must incorporate domain structure.

Heterogeneous Graph Condensation via Role-Aware Clustering

arXiv cs.LG

This paper proposes HGC-RC, a role-aware heterogeneous graph condensation framework that uses lightweight propagation and a hybrid clustering strategy to produce compact heterogeneous graphs, enabling efficient HGNN training on large-scale graphs without sacrificing performance.