Any-Dimensional Invariant Universality

arXiv cs.LG Papers

Summary

This paper develops a systematic framework for establishing universality of machine learning models that handle inputs of varying dimensions (e.g., graphs with different node counts). It shows that many existing architectures fail to be universal and proposes simple modifications to restore universality.

arXiv:2605.23156v1 Announce Type: new Abstract: Several machine learning models are defined for inputs of any size, such as graphs with different numbers of nodes and point clouds containing varying numbers of points. The universality properties of such any-dimensional models remain poorly understood, as universality is traditionally studied for models accepting inputs of a fixed size, defined on a compact subset of their domain. In sharp contrast, any-dimensional models can be viewed as sequences of functions defined on growing-sized inputs, and it is not clear in which sense they can be universal. We develop a systematic approach to establish any-dimensional universality, by identifying any-dimensional functions with a unique function taking inputs in a suitable infinite-dimensional limit space containing inputs of all finite sizes as well as their limits. Using the symmetries of these inputs and relations between inputs of different sizes, we show that this limit space admits a natural topology with rich families of compact sets on which any-dimensional universality can be established. We illustrate our approach by showing that several existing architectures fail to be universal, and we propose simple modifications that restore universality.
Original Article
View Cached Full Text

Cached at: 05/25/26, 09:01 AM

# Any-Dimensional Invariant Universality
Source: [https://arxiv.org/html/2605.23156](https://arxiv.org/html/2605.23156)
Shengtai YaoDepartment of Applied Mathematics and Statistics, Johns Hopkins University, Baltimore, MD 21218, USA\. MD is partially supported by NSF awards CCF 2442615 and DMS 2502377, and a Sloan Research Fellowship\.Eitan LevinDepartment of Computing and Mathematical Sciences, Caltech, Pasadena, CA 91125, USA\. EL is partially supported by AFOSR FA9550\-23\-1\-0070 and FA9550\-23\-1\-0204

###### Abstract

Several machine learning models are defined for inputs of any size, such as graphs with different numbers of nodes and point clouds containing varying numbers of points\. The universality properties of such any\-dimensional models remain poorly understood, as universality is traditionally studied for models accepting inputs of a fixed size, defined on a compact subset of their domain\. In sharp contrast, any\-dimensional models can be viewed as sequences of functions defined on growing\-sized inputs, and it is not clear in which sense they can be universal\. We develop a systematic approach to establish any\-dimensional universality, by identifying any\-dimensional functions with a unique function taking inputs in a suitable infinite\-dimensional limit space containing inputs of all finite sizes as well as their limits\. Using the symmetries of these inputs and relations between inputs of different sizes, we show that this limit space admits a natural topology with rich families of compact sets on which any\-dimensional universality can be established\. We illustrate our approach by showing that several existing architectures fail to be universal, and we propose simple modifications that restore universality\.

## 1Introduction

Traditional supervised learning aims to learn mappings defined on fixed, finite\-dimensional spaces by using training examples lying in those same spaces\. In contrast, many modern learning tasks involve maps defined on inputs of varying dimensions\. For instance, graph parameters, particle system dynamics, games, and regularizers remain meaningful regardless of the number of nodes, particles, players, or variables\. There are a number of architectures in the literature that are defined on inputs of any dimension, e\.g\., DeepSets\[[1](https://arxiv.org/html/2605.23156#bib.bib1)\]for sets, Graph Neural Networks \(GNN\)\[[2](https://arxiv.org/html/2605.23156#bib.bib2)\]for graphs, and PointNet\[[3](https://arxiv.org/html/2605.23156#bib.bib3)\]for point clouds\. Effectively, these any\-dimensional architectures finitely parametrize an infinite sequence of invariant functions\(f^n:Vn→ℝ\)n\(\\widehat\{f\}\_\{n\}\\colon V\_\{n\}\\to\\mathbb\{R\}\)\_\{n\}with increasingly larger input dimensions\.111By ‘invariant’ we mean that there is a sequence of groups\(Gn\)\(G\_\{n\}\), such as permutations or rotations, such that for eachnnthe groupGnG\_\{n\}acts onVnV\_\{n\}and for allg∈Gng\\in G\_\{n\}andv∈Vnv\\in V\_\{n\}, we havefn​\(g⋅v\)=fn​\(v\)f\_\{n\}\(g\\cdot v\)=f\_\{n\}\(v\)\.Despite the proliferation of any\-dimensional architectures, their expressivity remains relatively unexplored\. Specifically, the universality of most of the any\-dimensional architectures above has mostly been studied in a fixed dimension or in finitely many dimensions\[[4](https://arxiv.org/html/2605.23156#bib.bib4),[5](https://arxiv.org/html/2605.23156#bib.bib5),[6](https://arxiv.org/html/2605.23156#bib.bib6)\]\. This motivates our main question

Can any\-dimensional models universally approximate functions across dimensions?

As stated, this question is ill\-posed\. To understand why, recall that classical universality concerns approximating a*single*continuous functionf:K→ℝf\\colon K\\to\\mathbb\{R\}on a fixed compact setKKby a modelf^\\widehat\{f\}uniformly onKK\. Any\-dimensional models, however, produce a*sequence*of maps with varying domains\. To resolve this mismatch, we adopt the framework of\[[7](https://arxiv.org/html/2605.23156#bib.bib7)\]\. Specifically, we view the input spaces as a nested sequenceV1⊆V2⊆⋯V\_\{1\}\\subseteq V\_\{2\}\\subseteq\\cdotsof increasing\-dimensional vector spaces, and embed them all into an infinite\-dimensional limit spaceV∞V\_\{\\infty\}we construct containing eachVnV\_\{n\}as a subspace\. Under a mild compatibility condition, an any\-dimensional model\(f^n\)\(\\widehat\{f\}\_\{n\}\)admits a unique extensionf^∞:V∞→ℝ\\widehat\{f\}\_\{\\infty\}\\colon V\_\{\\infty\}\\to\\mathbb\{R\}satisfyingf^∞​\(x\)=f^n​\(x\)\\widehat\{f\}\_\{\\infty\}\(x\)=\\widehat\{f\}\_\{n\}\(x\)for anyx∈Vnx\\in V\_\{n\}andnn\. This identification lets us pose “universality across dimensions” as standard universality on an application\-specific infinite\-dimensional space\.

##### Main contributions\.

Let us summarize our contributions\.

1. \(A recipe\) We develop a general strategy for proving any\-dimensional invariant universality\. The strategy highlights two principles:\(i\)\(i\)exploiting symmetries by passing to orbit spaces can yield rich families of compact sets, even in infinite dimensions, on which universality can be proved; and\(i​i\)\(ii\)unlike in finite dimensions—where all norms induce the same topology—the choice of norm onV∞V\_\{\\infty\}is essential, as continuity \(and hence admissible activations and architectural primitives\) depends on it\.
2. \(Instantiations\) We apply this recipe to three representative domains: sets, graphs, and point clouds\. We show that several widely used architectures yield extensionsf^∞\\widehat\{f\}\_\{\\infty\}that are either discontinuous in the natural topology induced by the learning task or fail to be universal\. We then propose modifications—and, when necessary, new architectures—that restore continuity and achieve universality\.

##### Outline\.

The remainder of this section reviews related work\. Section[2](https://arxiv.org/html/2605.23156#S2)reviews the mathematical preliminaries of any dimensional learning necessary for the rest of the paper\. Section[3](https://arxiv.org/html/2605.23156#S3)formalizes the type of any\-dimensional universality that we wish to establish and outlines a general recipe for establishing universality\. Section[4](https://arxiv.org/html/2605.23156#S4)presents concrete instantiations of this recipe for models that apply to sets, graphs, and point clouds\. Finally, Section[5](https://arxiv.org/html/2605.23156#S5)concludes the paper with limitations and opportunities for future work\.

### 1\.1Related work

##### Dimension\-free learning\.

Many modern architectures are any\-dimensional\. Convolutional neural networks apply the same translation\-invariant filters to images of arbitrary resolution\[[8](https://arxiv.org/html/2605.23156#bib.bib8),[9](https://arxiv.org/html/2605.23156#bib.bib9)\]\. Recurrent neural networks and transformers handle variable\-length sequences via recurrence or attention\[[10](https://arxiv.org/html/2605.23156#bib.bib10),[11](https://arxiv.org/html/2605.23156#bib.bib11),[12](https://arxiv.org/html/2605.23156#bib.bib12)\]\. Neural operators such as DeepONet and FNO learn maps between \(infinite\-dimensional\) function spaces and can be evaluated on grids of different resolutions\[[13](https://arxiv.org/html/2605.23156#bib.bib13),[14](https://arxiv.org/html/2605.23156#bib.bib14)\]\. In geometric deep learning, several models have been proposed to handle sets of arbitrary cardinality\[[1](https://arxiv.org/html/2605.23156#bib.bib1),[3](https://arxiv.org/html/2605.23156#bib.bib3)\], graphs with varying numbers of nodes and edges\[[2](https://arxiv.org/html/2605.23156#bib.bib2),[15](https://arxiv.org/html/2605.23156#bib.bib15)\], and point clouds with any size\[[16](https://arxiv.org/html/2605.23156#bib.bib16),[17](https://arxiv.org/html/2605.23156#bib.bib17)\]\. Recently,\[[18](https://arxiv.org/html/2605.23156#bib.bib18),[19](https://arxiv.org/html/2605.23156#bib.bib19),[20](https://arxiv.org/html/2605.23156#bib.bib20)\]leveraged*representation stability*\[[21](https://arxiv.org/html/2605.23156#bib.bib21),[22](https://arxiv.org/html/2605.23156#bib.bib22)\]to derive general\-purpose dimension\-free equivariant and invariant model classes, encoding neural networks, convex sets, and kernel machines\. However, the ability to accept variable sizes does not guarantee consistent \(or transferable\) behavior across sizes\. The study of transferability originated in the GNN literature\[[23](https://arxiv.org/html/2605.23156#bib.bib23)\]and has since been explored extensively in that context\[[24](https://arxiv.org/html/2605.23156#bib.bib24),[25](https://arxiv.org/html/2605.23156#bib.bib25),[26](https://arxiv.org/html/2605.23156#bib.bib26),[27](https://arxiv.org/html/2605.23156#bib.bib27),[28](https://arxiv.org/html/2605.23156#bib.bib28),[29](https://arxiv.org/html/2605.23156#bib.bib29)\]\. More recently,\[[7](https://arxiv.org/html/2605.23156#bib.bib7)\]proposed a unified framework—beyond graphs—for formulating and certifying transferability across dimensions\.

##### Universality\.

Universality is one of the pillars of deep learning theory\. Let us start by commenting on finite\-dimensional results\. Classical universal approximation theorems show that fully connected networks approximate continuous functions on compact sets\[[30](https://arxiv.org/html/2605.23156#bib.bib30),[31](https://arxiv.org/html/2605.23156#bib.bib31),[32](https://arxiv.org/html/2605.23156#bib.bib32),[33](https://arxiv.org/html/2605.23156#bib.bib33)\], with recent extensions beyond compact domains under suitable growth and activation assumptions\[[34](https://arxiv.org/html/2605.23156#bib.bib34),[35](https://arxiv.org/html/2605.23156#bib.bib35),[36](https://arxiv.org/html/2605.23156#bib.bib36)\]\. The expressivity of symmetric models has only recently been studied\. For graphs, expressive power is subtle: message\-passing GNNs are known to have intrinsic limitations\[[37](https://arxiv.org/html/2605.23156#bib.bib37)\], commonly formalized via the Weisfeiler–Lehman \(WL\) test\[[38](https://arxiv.org/html/2605.23156#bib.bib38)\], and many variants have been analyzed through this lens\[[37](https://arxiv.org/html/2605.23156#bib.bib37),[39](https://arxiv.org/html/2605.23156#bib.bib39),[40](https://arxiv.org/html/2605.23156#bib.bib40)\]\. Alternative GNN based on tensor layers\[[4](https://arxiv.org/html/2605.23156#bib.bib4),[5](https://arxiv.org/html/2605.23156#bib.bib5)\], and those based on homomorphism density functions\[[41](https://arxiv.org/html/2605.23156#bib.bib41),[42](https://arxiv.org/html/2605.23156#bib.bib42)\]do achieve universality\. For permutation\-invariant problems, DeepSets is universal\[[1](https://arxiv.org/html/2605.23156#bib.bib1),[43](https://arxiv.org/html/2605.23156#bib.bib43)\]\. Beyond these specific examples,\[[6](https://arxiv.org/html/2605.23156#bib.bib6)\]established universality of polynomial\-invariant\-based models that are symmetric under the action of any compact group\. Nonetheless, these results are largely formulated for a fixed input dimension\. By contrast, universality for any\-dimensional models remains comparatively underdeveloped\. The closest related results are\[[44](https://arxiv.org/html/2605.23156#bib.bib44),[45](https://arxiv.org/html/2605.23156#bib.bib45)\]for sets and\[[5](https://arxiv.org/html/2605.23156#bib.bib5),[46](https://arxiv.org/html/2605.23156#bib.bib46)\]for graphs\. Our results and strategy highlight a more general perspective based on the transferability framework developed in\[[7](https://arxiv.org/html/2605.23156#bib.bib7)\], applicable beyond these concrete scenarios\. Let us briefly comment on more specific differences\. First,\[[45](https://arxiv.org/html/2605.23156#bib.bib45)\]studies lower bounds on approximation for certain RKHS classes; these classes are not our focus\. Our DeepSets results extend\[[44](https://arxiv.org/html/2605.23156#bib.bib44)\]by covering a broader class of compact sets\. On the graph side,\[[25](https://arxiv.org/html/2605.23156#bib.bib25)\]proves universality of graphon neural networks in simplified regimes, while\[[46](https://arxiv.org/html/2605.23156#bib.bib46)\]establishes universality with respect to theδp\\delta\_\{p\}metric\. In contrast, we work with the arguably more natural and widely studied cut metricδ□\\delta\_\{\\square\}, which induces a coarser topology\.

## 2Preliminaries

In this section, we introduce the notation and technical background necessary for the paper\.

##### Notation\.

We useℝ\\mathbb\{R\}andℕ\{\\mathbb\{N\}\}to denote the reals and positive integers, respectively\. Defineℕ0=ℕ∪\{0\}\{\{\\mathbb\{N\}\}\}\_\{0\}=\{\\mathbb\{N\}\}\\cup\\\{0\\\}and\[n\]≔\{1,…,n\}\[n\]\\coloneqq\\\{1,\\dots,n\\\}forn∈ℕn\\in\{\\mathbb\{N\}\}\. We writeO​\(k\)\\mathrm\{O\}\(k\)for the orthogonal group in dimensionkkandSn\\mathrm\{S\}\_\{n\}for the symmetric group onnnletters\. For a given metric space\(𝒳,d\)\(\\mathcal\{X\},d\), we denote the open ball centered atxxof radiusη\\etabyB​\(x,η\)≔\{u∈𝒳∣d​\(u,x\)<η\}B\(x,\\eta\)\\coloneqq\\\{u\\in\\mathcal\{X\}\\mid d\(u,x\)<\\eta\\\}\. LetC​\(𝒳,𝒴\)C\(\\mathcal\{X\},\\mathcal\{Y\}\)denote the space of continuous maps from a topological space𝒳\\mathcal\{X\}to𝒴\\mathcal\{Y\}, and writeC​\(𝒳\)C\(\\mathcal\{X\}\)when𝒴=ℝ\\mathcal\{Y\}=\\mathbb\{R\}\. Letℬ​\(𝒳,𝒴\)\\mathcal\{B\}\(\\mathcal\{X\},\\mathcal\{Y\}\)be the space of bounded linear operators between normed spaces𝒳\\mathcal\{X\}and𝒴\\mathcal\{Y\}\. Given an arbitrary norm∥⋅∥ℝk\\\|\\cdot\\\|\_\{\\mathbb\{R\}^\{k\}\}onℝk\\mathbb\{R\}^\{k\}, the symbolℓp​\(ℝk\)\\ell\_\{p\}\(\\mathbb\{R\}^\{k\}\)denotes the space of sequencesX=\(Xj:∈ℝk\)j∈ℕX=\(X\_\{j:\}\\in\\mathbb\{R\}^\{k\}\)\_\{j\\in\{\\mathbb\{N\}\}\}such that∑j=1∞‖Xj:‖ℝkp<∞\\sum\_\{j=1\}^\{\\infty\}\\\|X\_\{j:\}\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{p\}<\\infty\. Similarly, for a given measure space\(ℝm,μ\)\(\\mathbb\{R\}^\{m\},\\mu\), letLp​\(ℝm;ℝk\)L^\{p\}\(\\mathbb\{R\}^\{m\};\\mathbb\{R\}^\{k\}\)be the space of measurable functionsX:ℝm→ℝkX\\colon\\mathbb\{R\}^\{m\}\\to\\mathbb\{R\}^\{k\}with‖X‖p≔\(∫‖X​\(t\)‖ℝkp​𝑑μ​\(t\)\)1/p<∞\\\|X\\\|\_\{p\}\\coloneqq\\left\(\\int\\\|X\(t\)\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{p\}d\\mu\(t\)\\right\)^\{1/p\}<\\infty\. The symbol∥⋅∥p\\\|\\cdot\\\|\_\{p\}refers to either theℓp\\ell^\{p\}\- orLpL^\{p\}\-norms, as appropriate\. Finally, let𝒫​\(ℝk\)\\mathcal\{P\}\(\\mathbb\{R\}^\{k\}\)be the set of probability measures onℝk\\mathbb\{R\}^\{k\}, and let𝒫p​\(ℝk\)⊆𝒫​\(ℝk\)\\mathcal\{P\}\_\{p\}\(\\mathbb\{R\}^\{k\}\)\\subseteq\\mathcal\{P\}\\left\(\\mathbb\{R\}^\{k\}\\right\)denote those with finiteppth moments\. We use the symbolWpW\_\{p\}to denote the Wasserstein\-ppdistance\.

##### Any\-dimensional learning\.

Next, we recall a few basic notions pertaining to any\-dimensional models and their generalization across dimensions\. Our exposition here follows\[[7](https://arxiv.org/html/2605.23156#bib.bib7)\]; we refer the interested reader to that reference for additional details\. We focus exclusively on real\-valued functions, and accordingly streamline several definitions, though all notions in this section extend to more general codomains\. The key idea of this section is that a sequence of functions\(fn:Vn→ℝ\)\(f\_\{n\}\\colon V\_\{n\}\\to\\mathbb\{R\}\)with growing\-sized inputs can be identified with a single extensionf∞:V∞¯→ℝf\_\{\\infty\}\\colon\\overline\{V\_\{\\infty\}\}\\to\\mathbb\{R\}, defined on an infinite\-dimensional spaceV∞¯\\overline\{V\_\{\\infty\}\}that ‘contains’ eachVnV\_\{n\}as a subspace\. This viewpoint lets us pose questions such as universality on one common domain, rather than across a sequence of domains\. To state these ideas precisely, we start from a family of vector spaces encoding inputs of growing size, together with maps and relations linking them\.

###### Definition 2\.1\.

Aconsistent sequenceis a triple𝕍=\{\(Vn\)n∈ℕ,\(φN,n\)n⪯N,\(Gn\)n∈ℕ\}\\mathbb\{V\}=\\\{\\left\(V\_\{n\}\\right\)\_\{n\\in\{\\mathbb\{N\}\}\},\\left\(\\varphi\_\{N,n\}\\right\)\_\{n\\preceq N\},\\left\(\\mathrm\{G\}\_\{n\}\\right\)\_\{n\\in\{\\mathbb\{N\}\}\}\\\}indexed by a directed poset\(ℕ,⪯\)\(\{\\mathbb\{N\}\},\\preceq\),222A directed poset is a partial order⪯\\preceqonℕ\{\\mathbb\{N\}\}such that every two elements have a common upper bound\.of finite dimensional vector spacesVnV\_\{n\}, mapsφN,n\\varphi\_\{N,n\}and groupsGn\\mathrm\{G\}\_\{n\}acting linearly onVnV\_\{n\}such that, for alln⪯Nn\\preceq N,\(i\)\(i\)the groupGn\\mathrm\{G\}\_\{n\}is embedded intoGN\\mathrm\{G\}\_\{N\}and\(i​i\)\(ii\)φN,n:Vn↪VN\\varphi\_\{N,n\}\\colon V\_\{n\}\\hookrightarrow V\_\{N\}is a linear,Gn\\mathrm\{G\}\_\{n\}\-equivariant embedding\.

The elements of different vector spaces in a consistent sequence represent objects of different sizes, while the groups and embeddings between them represent their symmetries and relations between objects of different sizes, respectively\. To crystallize this notion, we provide two examples of consistent sequences that will play a crucial role in this work\.

###### Example 2\.2\.

LetVn=ℝnV\_\{n\}=\\mathbb\{R\}^\{n\}andGn=Sn\\mathrm\{G\}\_\{n\}=\\mathrm\{S\}\_\{n\}be the symmetric group acting by coordinate permutation\. One might endow this sequence with two different types of orderings and embeddings\.

1. 1\.\(Zero\-padding embedding\) Index byℕ\{\\mathbb\{N\}\}with the standard ordering≤\\leq\. DefineφN,n:Vn↪VN\\varphi\_\{N,n\}:V\_\{n\}\\hookrightarrow V\_\{N\}by φN,n​\(x1,…,xn\)≔\(x1,…,xn,0,…,0⏟\(N−n\)​zeros\)\.\\varphi\_\{N,n\}\\left\(x\_\{1\},\\ldots,x\_\{n\}\\right\)\\coloneqq\(x\_\{1\},\\ldots,x\_\{n\},\\underbrace\{0,\\ldots,0\}\_\{\(N\-n\)\\text\{ zeros\}\}\)\.We use the symbol𝕍zero\\mathbb\{V\}\_\{\\mathrm\{zero\}\}to denote this sequence\.
2. 2\.\(Duplication embedding\) Index byℕ\{\\mathbb\{N\}\}withn⪯Nn\\preceq NifnndividesNN\. DefineφN,n:Vn↪VN\\varphi\_\{N,n\}\\colon V\_\{n\}\\hookrightarrow V\_\{N\}by repeating each coordinateN/nN/ntimes, φN,n​\(x1,…,xn\)=\(x1,…,x1⏟N/n​copies,…,xn,…,xn⏟N/n​copies\)\.\\varphi\_\{N,n\}\\left\(x\_\{1\},\\ldots,x\_\{n\}\\right\)=\(\\underbrace\{x\_\{1\},\\ldots,x\_\{1\}\}\_\{N/n\\text\{ copies\}\},\\ldots,\\underbrace\{x\_\{n\},\\ldots,x\_\{n\}\}\_\{N/n\\text\{ copies\}\}\)\.Similarly, we use𝕍dup\\mathbb\{V\}\_\{\\mathrm\{dup\}\}to denote this sequence\.

To define and study continuous any\-dimensional models, we further define a single, infinite\-dimensional space containing inputs of all finite sizes\. This is done in the following definition\.

###### Definition 2\.3\.

DefineV∞V\_\{\\infty\}as the disjoint union⨆Vn\\bigsqcup V\_\{n\}modulo an equivalence relation

V∞≔⨆nVn/∼,V\_\{\\infty\}\\coloneqq\\bigsqcup\_\{n\}V\_\{n\}/\\sim,wherev∼φN,n​\(v\)v\\sim\\varphi\_\{N,n\}\(v\)whenevern⪯Nn\\preceq N, and denote by\[v\]∈V∞\[v\]\\in V\_\{\\infty\}the equivalence class ofv∈Vnv\\in V\_\{n\}\. The limiting groupG∞G\_\{\\infty\}and its equivalence classes\[g\]\[g\]are defined analogously\.

In turn, this allows us to identify any finite\-dimensional elementv∈Vnv\\in V\_\{n\}with its infinite\-dimensional counterpart\[v\]∈V∞\.\[v\]\\in V\_\{\\infty\}\.The next definition will allow us to extend a sequence of functions to this infinite dimensional space\.

###### Definition 2\.4\.

Let𝕍=\{\(Vn\),\(φN,n\),\(Gn\)\}\\mathbb\{V\}=\\left\\\{\\left\(V\_\{n\}\\right\),\\left\(\\varphi\_\{N,n\}\\right\),\\left\(\\mathrm\{G\}\_\{n\}\\right\)\\right\\\}be a consistent sequence indexed by\(ℕ,⪯\)\(\{\\mathbb\{N\}\},\\preceq\)\. A sequence\(fn:Vn→ℝ\)\\left\(f\_\{n\}:V\_\{n\}\\rightarrow\\mathbb\{R\}\\right\)iscompatiblewith respect to𝕍\\mathbb\{V\}iffN∘φN,n=fnf\_\{N\}\\circ\\varphi\_\{N,n\}=f\_\{n\}for alln⪯Nn\\preceq N, and eachfnf\_\{n\}isGn\\mathrm\{G\}\_\{n\}\-equivariant\.

In other words, a compatible sequence of functions is one that takes the same value on “equivalent” inputs, i\.e\., those related by symmetries and embeddings\. A sequence of functions\(fn\)\(f\_\{n\}\)is compatible if and only if there exists a uniqueG∞G\_\{\\infty\}\-equivariant mapf∞:V∞→ℝf\_\{\\infty\}\\colon V\_\{\\infty\}\\to\\mathbb\{R\}such thatf∞\|Vn=fnf\_\{\\infty\}\|\_\{V\_\{n\}\}=f\_\{n\}for alln∈ℕn\\in\{\\mathbb\{N\}\}\[[7](https://arxiv.org/html/2605.23156#bib.bib7)\]\. Thus, we can identify any compatible sequence of functions with a unique extension that matches the functions when restricted to finite dimensions\.

In order to reason about continuity off∞f\_\{\\infty\}, we will endowV∞V\_\{\\infty\}with a metric\. To do so, we consider a sequence of*compatible invariant norms*\(∥⋅∥Vn:Vn→ℝ\)\(\\\|\\cdot\\\|\_\{V\_\{n\}\}\\colon V\_\{n\}\\to\\mathbb\{R\}\)\. Compatibility yields the existence of aG∞G\_\{\\infty\}\-invariant norm∥⋅∥V∞\\\|\\cdot\\\|\_\{V\_\{\\infty\}\}onV∞V\_\{\\infty\}\. This construction allows us to define a limit space that contains not only \(equivalent classes of\) finite\-dimensional objects, but also their limits\.

###### Definition 2\.5\.

Thelimit spaceis the pair\(V∞¯,G∞\)\\left\(\\overline\{V\_\{\\infty\}\},\\mathrm\{G\}\_\{\\infty\}\\right\)whereV∞¯\\overline\{V\_\{\\infty\}\}denotes the completion ofV∞V\_\{\\infty\}with respect to∥⋅∥V∞\\\|\\cdot\\\|\_\{V\_\{\\infty\}\}, endowed with the symmetrized metric

d¯​\(x,y\)≔infg∈G∞‖g⋅x−y‖V∞for​x,y∈V∞¯\.\\overline\{\\mathrm\{d\}\}\(x,y\)\\coloneqq\\inf\_\{g\\in\\mathrm\{G\}\_\{\\infty\}\}\\\|g\\cdot x\-y\\\|\_\{V\_\{\\infty\}\}\\quad\\text\{ for \}x,y\\in\\overline\{V\_\{\\infty\}\}\.\(1\)Theorbit spaceis the metric space\(V∞¯/G∞,d¯\)\(\\overline\{V\_\{\\infty\}\}/G\_\{\\infty\},\\overline\{\\mathrm\{d\}\}\)with

V∞¯/G∞≔\{G∞⋅x∣x∈V∞¯\}¯,\\overline\{V\_\{\\infty\}\}/G\_\{\\infty\}\\coloneqq\\overline\{\\left\\\{G\_\{\\infty\}\\cdot x\\mid x\\in\\overline\{V\_\{\\infty\}\}\\right\\\}\},where the outer completion is taken with respect tod¯\\overline\{\\mathrm\{d\}\}\.

This symmetrized metric is a pseudometric onV∞¯\\overline\{V\_\{\\infty\}\}and a metric on the space of orbit closures ofV∞¯\\overline\{V\_\{\\infty\}\}under the action ofG∞G\_\{\\infty\}\[[7](https://arxiv.org/html/2605.23156#bib.bib7)\]\. We slightly abuse notation, by denoting this setV∞¯/G∞\\overline\{V\_\{\\infty\}\}/G\_\{\\infty\}and calling it “orbit space\.” Two elementsx,y∈V∞¯x,y\\in\\overline\{V\_\{\\infty\}\}are in the same orbit \(closure\) if, and only if,d¯​\(x,y\)=0\.\\overline\{\\mathrm\{d\}\}\(x,y\)=0\.Let us provide a quick example that will be used recurrently\.

###### Example 2\.6\(Duplication sequence withℓp\\ell\_\{p\}norms\)\.

Consider the duplication embedding from Example[2\.2](https://arxiv.org/html/2605.23156#S2.Thmtheorem2)\. Endow eachVn=ℝnV\_\{n\}=\\mathbb\{R\}^\{n\}with the normalizedℓp\\ell\_\{p\}norm forp∈\[1,∞\)p\\in\[1,\\infty\), i\.e\.,‖x‖Vn=\(1n​∑inxip\)1/p\\\|x\\\|\_\{V\_\{n\}\}=\(\\frac\{1\}\{n\}\\sum\_\{i\}^\{n\}x\_\{i\}^\{p\}\)^\{1/p\}\. Then, the limit spaceV∞V\_\{\\infty\}can be identified withLp​\(\[0,1\]\)L^\{p\}\(\[0,1\]\), with permutationsSn\\mathrm\{S\}\_\{n\}acting by permuting thennconsecutive intervals of length1/n1/nin the domain\[0,1\]\[0,1\]of functions inLp​\(\[0,1\]\)L^\{p\}\(\[0,1\]\)\. The orbit spaceV∞¯/G∞\\overline\{V\_\{\\infty\}\}/G\_\{\\infty\}corresponds to the space of probability distributions𝒫p​\(ℝ\)\\mathcal\{P\}\_\{p\}\(\\mathbb\{R\}\)\. Furthermore, the symmetrized distanced¯\\overline\{\\mathrm\{d\}\}coincides with the Wasserstein\-ppmetric\. Details appear in Appendix[C\.1\.2](https://arxiv.org/html/2605.23156#A3.SS1.SSS2)\.

Armed with these notions, we are now ready to define continuity across dimensions\.

###### Definition 2\.7\.

Let𝕍\\mathbb\{V\}be a consistent sequence endowed with a compatible invariant norm\(∥⋅∥Vn\)\(\\\|\\cdot\\\|\_\{V\_\{n\}\}\)\. A sequence of invariant functions\(fn\(f\_\{n\}:Vn→ℝ\)V\_\{n\}\\rightarrow\\mathbb\{R\}\)iscontinuously transferableif\(i\)\(i\)there exists an extensionf:V∞¯→ℝf\\colon\\overline\{V\_\{\\infty\}\}\\to\\mathbb\{R\}, i\.e\.,fn=f∞\|Vnf\_\{n\}=\\left\.f\_\{\\infty\}\\right\|\_\{V\_\{n\}\}for allnn, and\(i​i\)\(ii\)the extension is continuous with respect to∥⋅∥V∞\.\\\|\\cdot\\\|\_\{V\_\{\\infty\}\}\.

We remark that if\(fn\)\\left\(f\_\{n\}\\right\)is continuously transferable, then it must be compatible\. Further, an extensionf∞f\_\{\\infty\}is continuous in the limit space with respect to∥⋅∥V∞\\\|\\cdot\\\|\_\{V\_\{\\infty\}\}if, and only if, it is continuous in the orbit spaceV∞¯/G∞\\overline\{V\_\{\\infty\}\}/G\_\{\\infty\}with respect tod¯\\overline\{\\mathrm\{d\}\}; see Lemma[A\.4](https://arxiv.org/html/2605.23156#A1.Thmtheorem4)in Appendix[A](https://arxiv.org/html/2605.23156#A1)\. The term ‘transferable’ is motivated by the results in\[[7](https://arxiv.org/html/2605.23156#bib.bib7)\], which show that if a high\-dimensional model is transferable, then it can be trained on low\-dimensional inputs and generalize to higher\-dimensional ones\.

## 3How to establish any\-dimensional invariant universality?

In this section, we outline a general recipe for establishing any\-dimensional invariant universality\. We begin by formally stating our goal\. Consider a continuously transferable, consistent sequence of invariant functions\(fn:Vn→ℝ\)\(f\_\{n\}:V\_\{n\}\\rightarrow\\mathbb\{R\}\)\. Building upon the foundational concepts from Section[2](https://arxiv.org/html/2605.23156#S2), our primary focus is on approximating itsG∞G\_\{\\infty\}\-invariant extensionf∞f\_\{\\infty\}\. To approximate this function we consider a class of continuously transferable any\-dimensional modelsℱ\\mathcal\{F\}\. We identify each of these models with its limiting extensionf^∞\\widehat\{f\}\_\{\\infty\}and our goal is to show that: for any fixed accuracyε\>0\\varepsilon\>0and compact setKKin the domain off∞f\_\{\\infty\}, there existsf^∞∈ℱ\\widehat\{f\}\_\{\\infty\}\\in\\mathcal\{F\}such that

supx∈K\|f∞​\(x\)−f^∞​\(x\)\|≤ε\.\\sup\_\{x\\in K\}\|f\_\{\\infty\}\(x\)\-\\widehat\{f\}\_\{\\infty\}\(x\)\|\\leq\\varepsilon\.\(2\)We make two remarks concerning domains and topology\.

##### Domain of the extension\.

Since invariant functions are constant on orbits, they admit two equivalent viewpoints: as functions on the limit spacef∞:V∞¯→ℝf\_\{\\infty\}\\colon\\overline\{V\_\{\\infty\}\}\\to\\mathbb\{R\}, or as functions on the orbit spacef∞:V∞¯/G∞→ℝf\_\{\\infty\}\\colon\\overline\{V\_\{\\infty\}\}/G\_\{\\infty\}\\to\\mathbb\{R\}\. Thus, the question of approximation can be posed over any of these domains\. Any consistent norm∥⋅∥V∞\\\|\\cdot\\\|\_\{V\_\{\\infty\}\}induces a topology over the limit space, the corresponding topology on the orbit space is given by the symmetrized metric \([1](https://arxiv.org/html/2605.23156#S2.E1)\)\. As it turns out, compact sets in the orbit space are ‘richer’ than those in the limit space\. The underlying reason for this phenomenon is that by taking orbits, we collapse ‘big’ subsets \(potentially noncompact\) into single orbits\. To illustrate this point, consider the setting in Example[2\.6](https://arxiv.org/html/2605.23156#S2.Thmtheorem6)whereV∞¯=Lp​\(\[0,1\]\)\\overline\{V\_\{\\infty\}\}=L^\{p\}\(\[0,1\]\)are integrable functions andV∞¯/G∞=𝒫p​\(ℝ\)\\overline\{V\_\{\\infty\}\}/G\_\{\\infty\}=\\mathcal\{P\}\_\{p\}\(\\mathbb\{R\}\)are probability distributions withppth moments\. While the setK~=\{X∈Lp​\(\[0,1\]\)∣range⁡\(X\)⊆\[0,1\]\}\\widetilde\{K\}=\\\{X\\in L^\{p\}\(\[0,1\]\)\\mid\\operatorname\{range\}\\left\(X\\right\)\\subseteq\[0,1\]\\\}is not compact inLp​\(\[0,1\]\),L^\{p\}\(\[0,1\]\),the collection of its equivalence classesK=\{μ∈𝒫p​\(ℝ\)∣supp⁡\(μ\)⊆\[0,1\]\}K=\\\{\\mu\\in\\mathcal\{P\}\_\{p\}\(\\mathbb\{R\}\)\\mid\\operatorname\{supp\}\(\\mu\)\\subseteq\[0,1\]\\\}is compact in𝒫p​\(ℝ\);\\mathcal\{P\}\_\{p\}\(\\mathbb\{R\}\);we defer details to Appendix[B](https://arxiv.org/html/2605.23156#A2)\. In Section[4](https://arxiv.org/html/2605.23156#S4), we will see more examples of this phenomena\. In what follows, we focus on universality on the orbit space equipped with the symmetrized metric, i\.e\., we prove \([2](https://arxiv.org/html/2605.23156#S3.E2)\) for compact subsetsKKof the orbit space\.

##### Induced topology and activation functions\.

In finite dimensional settings, all norms induce equivalent topologies, which means that continuity with respect to one norm guarantees continuity under any other\. In sharp contrast, the topology of the limit spaceV∞¯\\overline\{V\_\{\\infty\}\}\(and correspondingly that of the orbit space\) is heavily dependent on the choice of the consistent norm∥⋅∥V∞\.\\\|\\cdot\\\|\_\{V\_\{\\infty\}\}\.This choice is often dictated by the learning task, i\.e\., the mapf∞f\_\{\\infty\}we aim to approximate \(or in practice learn\) is naturally continuous with respect to certain norms but not others\. To illustrate this point, consider again the setting in Example[2\.6](https://arxiv.org/html/2605.23156#S2.Thmtheorem6)withp=2p=2, and the sequence of ‘second moment’ functions

fn​\(x\)=1n​∑j=1nxj2​with extension​f∞​\(μX\)=𝔼Y∼μX⁡\[Y2\]f\_\{n\}\(x\)=\\frac\{1\}\{n\}\\sum\_\{j=1\}^\{n\}x\_\{j\}^\{2\}\\,\\,\\,\\text\{with extension\}\\,\\,\\,f\_\{\\infty\}\(\\mu\_\{X\}\)=\\operatorname\{\\mathbb\{E\}\}\_\{Y\\sim\\mu\_\{X\}\}\[Y^\{2\}\]whereμX∈𝒫2​\(ℝ\)\\mu\_\{X\}\\in\\mathcal\{P\}\_\{2\}\(\\mathbb\{R\}\)denotes the probability law associated withX∈L2​\(\[0,1\]\)X\\in L^\{2\}\(\[0,1\]\)\. In this case,f∞f\_\{\\infty\}is continuous with respect toW2W\_\{2\}, but discontinuous with respect toW1W\_\{1\}\. As we will see in Section[4](https://arxiv.org/html/2605.23156#S4), several existing any\-dimensional architectures are not continuous in the natural topologies induced by norms of interest\. A recurring theme is that the chosen topology constrains which activation functions and architectural primitives can be used\.

##### Desiderata and a recipe\.

We now outline our strategy for proving universality\. Our proof strategy is based on the Stone–Weierstrass theorem, which we recall for convenience; see\[[47](https://arxiv.org/html/2605.23156#bib.bib47)\]\[Theorem 5\.7\] for a proof\.

###### Theorem 3\.1\(Stone\-Weierstrass\)\.

LetC​\(K\)C\(K\)be the set of continuous real\-valued functions on a compact metric spaceKKendowed with the norm‖f‖∞=supx∈K\|f​\(x\)\|\\\|f\\\|\_\{\\infty\}=\\sup\_\{x\\in K\}\|f\(x\)\|\. Letℱ\\mathcal\{F\}be a subalgebra ofC​\(K\)C\(K\)which contains a non\-zero constant function\.333A subalgebra of functions is a set closed under addition, scalar multiplication, and products\.Thenℱ\\mathcal\{F\}is dense inC​\(K\)C\(K\)if, and only if, it separates points\.444The function classℱ\\mathcal\{F\}separates points onKKif for allx,y∈Kx,y\\in Kandx≠yx\\neq y, there is anf∈ℱf\\in\\mathcal\{F\}such thatf​\(x\)≠f​\(y\)\.f\(x\)\\neq f\(y\)\.

Density in the sup\-norm is equivalent to \([2](https://arxiv.org/html/2605.23156#S3.E2)\)\. This discussion suggests two desiderata for practical model classesℱ\\mathcal\{F\}:\(i\)\(i\)functions inℱ\\mathcal\{F\}should be continuous with respect to the natural topology induced by the learning task; and\(i​i\)\(ii\)the classℱ\\mathcal\{F\}should form a subalgebra that separates points\. Somewhat surprisingly, many existing architectures fail to meet one \(or both\) of these requirements, likely because most prior work studies approximation on a fixed dimensionVnV\_\{n\}rather than on the orbit space\. Guided by these desiderata, we propose the following three\-step recipe\.

1. Step 1 \(Compact sets\)\.Given a consistent norm∥⋅∥∞\\\|\\cdot\\\|\_\{\\infty\}dictated by the learning task, characterize compact sets in the orbit spaceV∞¯/G∞\\overline\{V\_\{\\infty\}\}/G\_\{\\infty\}\.
2. Step 2 \(Continuity\)\.Construct a collectionℱ\\mathcal\{F\}of invariant any\-dimensional models that are continuously transferable with respect to the symmetrized metric \([1](https://arxiv.org/html/2605.23156#S2.E1)\) on the orbit spaceV∞¯/G∞\\overline\{V\_\{\\infty\}\}/G\_\{\\infty\}\.
3. Step 3 \(Universality\)\.Establish universality ofℱ\\mathcal\{F\}via the Stone–Weierstrass theorem by showing thatℱ\\mathcal\{F\}forms a subalgebra that separates points\.

## 4Instantiations of the recipe

In this section, we instantiate our general recipe for any\-dimensional models on three representative domains: sets \(Section[4\.1](https://arxiv.org/html/2605.23156#S4.SS1)\), graphs \(Section[4\.2](https://arxiv.org/html/2605.23156#S4.SS2)\), and point clouds \(Section[4\.3](https://arxiv.org/html/2605.23156#S4.SS3)\)\. For brevity, we work directly with the corresponding infinite\-dimensional extensions of our models, rather than the underlying sequences\. Across all domains, we find that standard architectures are either discontinuous, or fail to be universal\. We therefore either introduce minor, principled modifications to these existing architectures or propose new architectures, and we prove that the resulting models are both continuous and universal\.

### 4\.1Set functions

We start with models for sets of any size\. Since fully connected feedforward neural networks appear repeatedly throughout, we define the family

NNk,mϕ≔⋃r=1∞\{L2∘ϕ∘L1∣L1∈ℒr,m,L2∈ℒk,r\},\\texttt\{NN\}^\{\\phi\}\_\{k,m\}\\coloneqq\\bigcup\_\{r=1\}^\{\\infty\}\\big\\\{L\_\{2\}\\circ\\phi\\circ L\_\{1\}\\mid L\_\{1\}\\in\\mathcal\{L\}\_\{r,m\},L\_\{2\}\\in\\mathcal\{L\}\_\{k,r\}\\big\\\},\(3\)whereℒm,n≔\{x↦W​x\+θ∣W∈ℝm×n,θ∈ℝm\}\\mathcal\{L\}\_\{m,n\}\\coloneqq\\left\\\{x\\mapsto Wx\+\\theta\\mid W\\in\\mathbb\{R\}^\{m\\times n\},\\theta\\in\\mathbb\{R\}^\{m\}\\right\\\}is the collection of affine maps fromℝn\\mathbb\{R\}^\{n\}toℝm\\mathbb\{R\}^\{m\}andϕ:ℝ→ℝ\\phi\\colon\\mathbb\{R\}\\to\\mathbb\{R\}is an activation function applied component\-wise\. Observe that these networks can be arbitrarily wide\.

We consider two models whose input spaces parallel the consistent sequences in Example[2\.2](https://arxiv.org/html/2605.23156#S2.Thmtheorem2)\. Although both sequences share the same finite\-dimensional spacesVn=ℝn×kV\_\{n\}=\\mathbb\{R\}^\{n\\times k\}, they use different embeddings and consistent norms, and therefore induce substantially different orbit spaces\. In both settings we start from the DeepSets architecture\[[1](https://arxiv.org/html/2605.23156#bib.bib1)\], which maps

DeepSetsρ,σ⁡\(X\)=σ​\(∑i=1∞ρ​\(Xi:\)\),\\operatorname\{DeepSets\}^\{\\rho,\\sigma\}\(X\)=\\sigma\\left\(\\sum\_\{i=1\}^\{\\infty\}\\rho\(X\_\{i:\}\)\\right\),\(4\)whereρ∈NNr,kϕ\\rho\\in\\texttt\{NN\}^\{\\phi\}\_\{r,k\},σ∈NN1,rϕ\\sigma\\in\\texttt\{NN\}^\{\\phi\}\_\{1,r\},r∈ℕr\\in\{\\mathbb\{N\}\}\. For each of the two orbit spaces we consider, we introduce a mild modification of DeepSets that restores continuity and yields universality\.

#### 4\.1\.1Universality over sequences

Additional details about the following constructions can be found in Appendix[C\.1\.1](https://arxiv.org/html/2605.23156#A3.SS1.SSS1)\. For the first example, we consider thekk\-fold direct sum of the zero\-padding consistent sequence𝕍zero⊕k=\{\(ℝn×k\),\(φN,n⊕k\),\(Sn\)\}\\mathbb\{V\}\_\{\\mathrm\{zero\}\}^\{\\oplus k\}=\\\{\(\\mathbb\{R\}^\{n\\times k\}\),\(\\varphi\_\{N,n\}^\{\\oplus k\}\),\(\\mathrm\{S\}\_\{n\}\)\\\}\. Given an arbitrary norm∥⋅∥ℝk\\\|\\cdot\\\|\_\{\\mathbb\{R\}^\{k\}\}inℝk\\mathbb\{R\}^\{k\}, and a real numberp∈\[1,∞\)p\\in\[1,\\infty\), we endow𝕍zero⊕k\\mathbb\{V\}\_\{\\mathrm\{zero\}\}^\{\\oplus k\}with the compatible sequence ofℓp\\ell\_\{p\}norm‖X‖Vn=\(∑i=1n‖Xi:‖ℝkp\)1/p\\\|X\\\|\_\{V\_\{n\}\}=\(\\sum\_\{i=1\}^\{n\}\\\|X\_\{i:\}\\\|^\{p\}\_\{\\mathbb\{R\}^\{k\}\}\)^\{1/p\}, which are unchanged under zero\-padding\. The resulting limit space is the space of sequencesV∞¯=ℓp​\(ℝk\)\\overline\{V\_\{\\infty\}\}=\\ell\_\{p\}\(\\mathbb\{R\}^\{k\}\), with the groupG∞G\_\{\\infty\}acting by permuting finitely many entries of these sequences\.

##### Step 1 \(Compact sets\)\.

For this example, we consider orbit closures of compact setsK⊆ℓp​\(ℝk\)K\\subseteq\\ell\_\{p\}\(\\mathbb\{R\}^\{k\}\), which we denote byK/G∞K/G\_\{\\infty\}\. These are compact subsets ofV∞¯/G∞\\overline\{V\_\{\\infty\}\}/G\_\{\\infty\}by the definition of the quotient topology onV∞¯/G∞\\overline\{V\_\{\\infty\}\}/G\_\{\\infty\}\. Compact sets inℓp​\(ℝk\)\\ell\_\{p\}\(\\mathbb\{R\}^\{k\}\)are well\-understood\.

###### Proposition 4\.1\(Compactness inℓp​\(ℝk\)\\ell\_\{p\}\(\\mathbb\{R\}^\{k\}\),\[[48](https://arxiv.org/html/2605.23156#bib.bib48)\]\)\.

Forp∈\[1,∞\)p\\in\[1,\\infty\), a setK⊆ℓp​\(ℝk\)K\\subseteq\\ell\_\{p\}\(\\mathbb\{R\}^\{k\}\)is compact if, and only if, it satisfies that\(i\)\(i\)it is closed;\(i​i\)\(ii\)it is bounded, i\.e\.,supX∈K‖X‖p<∞;\\sup\_\{X\\in K\}\\\|X\\\|\_\{p\}<\\infty;and\(i​i​i\)\(iii\)it exhibits uniform tail decay, i\.e\.,limN→∞supX∈K\(∑i≥N‖Xi:‖ℝkp\)1/p=0\.\\lim\_\{N\\rightarrow\\infty\}\\sup\_\{X\\in K\}\(\\sum\_\{i\\geq N\}\\\|X\_\{i:\}\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{p\}\)^\{1/p\}=0\.

For instance, the set\{\(xi\)i=1∞:\|xi\|≤2−i​for all​i∈ℕ\}\\\{\(x\_\{i\}\)\_\{i=1\}^\{\\infty\}:\|x\_\{i\}\|\\leq 2^\{\-i\}\\textrm\{ for all \}i\\in\{\\mathbb\{N\}\}\\\}is compact inℓp​\(ℝ\)\\ell\_\{p\}\(\\mathbb\{R\}\)for allp∈\[1,∞\)p\\in\[1,\\infty\)\. We now turn to constructing a family of continuous models onV∞¯/G∞\\overline\{V\_\{\\infty\}\}/G\_\{\\infty\}\.

##### Step 2 \(Continuity\)\.

Our architecture is based on DeepSets \([4](https://arxiv.org/html/2605.23156#S4.E4)\)\. To ensure that our architecture is compatible with respect to our consistent sequence structure, i\.e\., that it is unchanged under zero\-padding, we assume thatρ​\(𝟎k\)=𝟎r\\rho\\left\(\\mathbf\{0\}\_\{k\}\\right\)=\\mathbf\{0\}\_\{r\}\. In general, \([4](https://arxiv.org/html/2605.23156#S4.E4)\) is not a continuous model as the infinite sum may diverge since the functionρ\\rhocan decay arbitrarily slowly\. Motivated by this observation, we propose

DeepSets∞ρ,σ⁡\(X\)=σ​\(∑i=1∞‖Xi:‖ℝkp⋅ρ​\(Xi:\)\),\\operatorname\{DeepSets\}\_\{\\infty\}^\{\\rho,\\sigma\}\(X\)=\\sigma\\left\(\\sum\_\{i=1\}^\{\\infty\}\\\|X\_\{i:\}\\\|^\{p\}\_\{\\mathbb\{R\}^\{k\}\}\\cdot\\rho\(X\_\{i:\}\)\\right\),\(5\)whereρ∈NNr,kϕ\\rho\\in\\texttt\{NN\}^\{\\phi\}\_\{r,k\}andσ∈NN1,rϕ\\sigma\\in\\texttt\{NN\}^\{\\phi\}\_\{1,r\}withr∈ℕr\\in\{\\mathbb\{N\}\}\. The following proposition shows thatDeepSets∞ρ,σ\{\\mathrm\{DeepSets\}\}\_\{\\infty\}^\{\\rho,\\sigma\}is indeed continuous on the spaceV∞¯/G∞\\overline\{V\_\{\\infty\}\}/G\_\{\\infty\}of orbit closures when bothρ\\rhoandσ\\sigmaare continuous\. The proof is deferred to Appendix[C\.2](https://arxiv.org/html/2605.23156#A3.SS2)\.

###### Theorem 4\.2\(Continuity on sequences\)\.

Supposeρ∈C​\(ℝk,ℝr\)\\rho\\in C\(\\mathbb\{R\}^\{k\},\\mathbb\{R\}^\{r\}\)andσ∈C​\(ℝr,ℝ\)\\sigma\\in C\(\\mathbb\{R\}^\{r\},\\mathbb\{R\}\)are continuous functions\. Then, the mapDeepSets∞ρ,σ:V∞¯/G∞→ℝ\\operatorname\{DeepSets\}\_\{\\infty\}^\{\\rho,\\sigma\}\\colon\\overline\{V\_\{\\infty\}\}/G\_\{\\infty\}\\to\\mathbb\{R\}defined in \([5](https://arxiv.org/html/2605.23156#S4.E5)\) is continuous with respect to the symmetrized metric \([1](https://arxiv.org/html/2605.23156#S2.E1)\)\.

As an immediate consequence, we get that all functions in

ℱDS≔\{DeepSets∞ρ,σ∣r∈ℕ,ρ∈NNr,kϕ,σ∈NN1,rϕ\}​are continuous\.\\mathcal\{F\}\_\{\\mathrm\{DS\}\}\\coloneqq\\left\\\{\{\\mathrm\{DeepSets\}\}\_\{\\infty\}^\{\\rho,\\sigma\}\\mid r\\in\{\\mathbb\{N\}\},\\rho\\in\\texttt\{NN\}^\{\\phi\}\_\{r,k\},\\sigma\\in\\texttt\{NN\}^\{\\phi\}\_\{1,r\}\\right\\\}\\text\{ are continuous\.\}

##### Step 3 \(Universality\)\.

Finally, we prove that the above variation of DeepSets is universal for sequences\. The proof of the following result is deferred to Appendix[C\.3](https://arxiv.org/html/2605.23156#A3.SS3)\.

###### Theorem 4\.3\(Universality ofDeepSets∞\{\\mathrm\{DeepSets\}\}\_\{\\infty\}\)\.

Suppose that the activationϕ\\phiis continuous and non\-polynomial\. LetK⊆ℓp​\(ℝk\)K\\subseteq\\ell\_\{p\}\(\\mathbb\{R\}^\{k\}\)be any compact set\. Then,ℱDS\\mathcal\{F\}\_\{\\mathrm\{DS\}\}is dense inC​\(K/G∞\)C\(K/G\_\{\\infty\}\)\.

#### 4\.1\.2Universality over measures

Additional details about the constructions in this section are deferred to Appendix[C\.1\.2](https://arxiv.org/html/2605.23156#A3.SS1.SSS2)\. We consider thekk\-fold direct sum of the duplication embedding consistent sequence𝕍dup⊕k≔\{\(ℝn×k\),\(φN,n⊕k\),\(Sn\)\}\\mathbb\{V\}\_\{\\mathrm\{dup\}\}^\{\\oplus k\}\\coloneqq\\\{\(\\mathbb\{R\}^\{n\\times k\}\),\(\\varphi\_\{N,n\}^\{\\oplus k\}\),\(\\mathrm\{S\}\_\{n\}\)\\\}defined as in Example[2\.2](https://arxiv.org/html/2605.23156#S2.Thmtheorem2)\. Given an arbitrary norm∥⋅∥ℝk\\\|\\cdot\\\|\_\{\\mathbb\{R\}^\{k\}\}onℝk\\mathbb\{R\}^\{k\}, and a real numberp∈\[1,∞\)p\\in\[1,\\infty\), we endow eachℝn×k\\mathbb\{R\}^\{n\\times k\}with the normalizedℓp\\ell\_\{p\}norm, i\.e\.,‖X‖p¯=\(1n​∑i=1n‖Xi:‖ℝkp\)1/p\\\|X\\\|\_\{\\bar\{p\}\}=\(\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\\|X\_\{i:\}\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{p\}\)^\{1/p\}\. As explained in Example[2\.6](https://arxiv.org/html/2605.23156#S2.Thmtheorem6), the resulting limit space is the space of functionsLp​\(\[0,1\],ℝk\)L^\{p\}\(\[0,1\],\\mathbb\{R\}^\{k\}\), and the space of orbit closuresV∞¯/G∞\\overline\{V\_\{\\infty\}\}/G\_\{\\infty\}can be identified with the Wasserstein space𝒫p​\(ℝk\)\\mathcal\{P\}\_\{p\}\(\\mathbb\{R\}^\{k\}\)equipped with the Wasserstein\-ppdistanceWpW\_\{p\}\.

##### Step 1 \(Compact sets\)\.

We begin with a characterization of compact sets in the orbit space\.

###### Proposition 4\.4\(Compacta in𝒫p​\(ℝk\)\\mathcal\{P\}\_\{p\}\(\\mathbb\{R\}^\{k\}\)\[[49](https://arxiv.org/html/2605.23156#bib.bib49)\]\)\.

Forp∈\[1,∞\)p\\in\[1,\\infty\), a subsetQ⊆𝒫p​\(ℝk\)Q\\subseteq\\mathcal\{P\}\_\{p\}\(\\mathbb\{R\}^\{k\}\)is compact if, and only if, it is\(i\)\(i\)closed;\(i​i\)\(ii\)tight, i\.e\., for anyε\>0\\varepsilon\>0, there is a compactKε⊆ℝkK\_\{\\varepsilon\}\\subseteq\\mathbb\{R\}^\{k\}such thatsupμ∈Qμ​\(ℝk\\Kε\)≤ε;\\sup\_\{\\mu\\in Q\}\\mu\\left\(\\mathbb\{R\}^\{k\}\\backslash K\_\{\\varepsilon\}\\right\)\\leq\\varepsilon;and\(i​i​i\)\(iii\)pp\-uniformly integrable, i\.e\.,

limR→∞supμ∈Q∫‖x‖ℝk\>R‖x‖ℝkp​dμ​\(x\)=0\.\\lim\_\{R\\rightarrow\\infty\}\\sup\_\{\\mu\\in Q\}\\int\_\{\\\|x\\\|\_\{\\mathbb\{R\}^\{k\}\}\>R\}\\\|x\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{p\}\\mathrm\{~d\}\\mu\(x\)=0\.

For example, the collection of all measures supported on a closed ball of radiusR\>0R\>0,Q=\{μ∈𝒫p​\(ℝk\)∣supp⁡\(μ\)⊆B​\(0,R\)¯\},Q=\\bigl\\\{\\mu\\in\\mathcal\{P\}\_\{p\}\(\\mathbb\{R\}^\{k\}\)\\mid\\operatorname\{supp\}\(\\mu\)\\subseteq\\overline\{B\(0,R\)\}\\bigr\\\},is compact in this space\. We now turn to constructing a family of continuous models on𝒫p​\(ℝk\)\\mathcal\{P\}\_\{p\}\(\\mathbb\{R\}^\{k\}\)that can approximate any continuous function on a compact subset to arbitrary precision\.

##### Step 2 \(Continuity\)\.

We consider the normalized DeepSets architecture proposed in\[[44](https://arxiv.org/html/2605.23156#bib.bib44)\],

DeepSets¯∞ρ,σ​\(μ\)=σ​\(∫ρ​𝑑μ\),\\overline\{\\mathrm\{DeepSets\}\}\_\{\\infty\}^\{\\rho,\\sigma\}\(\\mu\)=\\sigma\{\\left\(\\int\\rho d\\mu\\right\)\},\(6\)whereρ∈NNr,kϕ\\rho\\in\\texttt\{NN\}^\{\\phi\}\_\{r,k\}andσ∈NN1,rϕ\\sigma\\in\\texttt\{NN\}^\{\\phi\}\_\{1,r\}withr∈ℕr\\in\{\\mathbb\{N\}\}\. To establish the continuity of this architecture we require that the entries ofρ\\rhodo not grow too fast\. In particular, we denote byℱp​\(ℝk,ℝr\)\\mathcal\{F\}\_\{p\}\(\\mathbb\{R\}^\{k\},\\mathbb\{R\}^\{r\}\)the class of continuous functionsρ=\(ρ1,…,ρr\)⊤\\rho=\(\\rho\_\{1\},\\ldots,\\rho\_\{r\}\)^\{\\top\}such that each componentρj\\rho\_\{j\}satisfies app\-th order growth condition; specifically, for eachj∈\[r\]j\\in\[r\], there exists a constantMj\>0M\_\{j\}\>0such that

\|ρj​\(x\)\|≤Mj​\(1\+‖x‖ℝkp\),∀x∈ℝk\.\|\\rho\_\{j\}\(x\)\|\\leq M\_\{j\}\(1\+\\\|x\\\|^\{p\}\_\{\\mathbb\{R\}^\{k\}\}\),\\quad\\forall x\\in\\mathbb\{R\}^\{k\}\.The following theorem establishes the continuity ofDeepSets¯∞ρ,σ\\overline\{\\mathrm\{DeepSets\}\}\_\{\\infty\}^\{\\rho,\\sigma\}under this growth condition\. We defer its proof to Appendix[C\.4](https://arxiv.org/html/2605.23156#A3.SS4)\.

###### Theorem 4\.5\(Continuity on probability measures\)\.

Supposeρ∈ℱp​\(ℝk,ℝr\)\\rho\\in\\mathcal\{F\}\_\{p\}\(\\mathbb\{R\}^\{k\},\\mathbb\{R\}^\{r\}\)andσ∈C​\(ℝr,ℝ\)\\sigma\\in C\(\\mathbb\{R\}^\{r\},\\mathbb\{R\}\)are continuous functions\. Then, the mapDeepSets¯∞ρ,σ:𝒫p​\(ℝk\)→ℝ\\overline\{\\mathrm\{DeepSets\}\}\_\{\\infty\}^\{\\rho,\\sigma\}\\colon\\mathcal\{P\}\_\{p\}\\left\(\\mathbb\{R\}^\{k\}\\right\)\\rightarrow\\mathbb\{R\}, defined in \([6](https://arxiv.org/html/2605.23156#S4.E6)\), is continuous with respect toWpW\_\{p\}\.

As an immediate consequence of this theorem, we obtain that all functions in

ℱDS¯≔\{DeepSets¯∞ρ,σ∣r∈ℕ,ρ∈NNr,kϕ,σ∈NN1,rϕ\}\\mathcal\{F\}\_\{\\overline\{\\mathrm\{DS\}\}\}\\coloneqq\\left\\\{\\overline\{\\mathrm\{DeepSets\}\}\_\{\\infty\}^\{\\rho,\\sigma\}\\mid r\\in\{\\mathbb\{N\}\},\\rho\\in\\texttt\{NN\}^\{\\phi\}\_\{r,k\},\\sigma\\in\\texttt\{NN\}^\{\\phi\}\_\{1,r\}\\right\\\}are continuous provided thatϕ\\phiis Lipschitz, since it ensures thatNNr,kϕ⊆ℱp​\(ℝk,ℝr\)\\texttt\{NN\}^\{\\phi\}\_\{r,k\}\\subseteq\\mathcal\{F\}\_\{p\}\(\\mathbb\{R\}^\{k\},\\mathbb\{R\}^\{r\}\), forp∈\[1,∞\)p\\in\[1,\\infty\)\.

##### Step 3 \(Universality\)

Finally, we establish the universality of the normalized DeepSets architecture in \([6](https://arxiv.org/html/2605.23156#S4.E6)\)\. We defer the proof of this result to Appendix[C\.5](https://arxiv.org/html/2605.23156#A3.SS5)\.

###### Theorem 4\.6\(Universality ofDeepSets¯∞\\overline\{\\mathrm\{DeepSets\}\}\_\{\\infty\}\)\.

Suppose thatϕ\\phiis a Lipschitz continuous, non\-polynomial activation that is asymptotically polynomial at±∞\\pm\\infty\.555We say thatϕ​\(t\)\\phi\(t\)is asymptotically polynomial at±∞\\pm\\inftyif there exist polynomialsP\+P\_\{\+\}andP−P\_\{\-\}\(possibly constant\) such thatϕ​\(t\)/P±​\(t\)→1\\phi\(t\)/P\_\{\\pm\}\(t\)\\rightarrow 1ast→±∞t\\rightarrow\\pm\\infty, respectively\.LetQ⊆𝒫p​\(ℝk\)Q\\subseteq\\mathcal\{P\}\_\{p\}\(\\mathbb\{R\}^\{k\}\)be a compact set\. Then,ℱDS¯\\mathcal\{F\}\_\{\\overline\{\\mathrm\{DS\}\}\}is dense inC​\(Q\)C\(Q\)\.

The special case of Theorem[4\.6](https://arxiv.org/html/2605.23156#S4.Thmtheorem6)withQQconsisting of all measures supported on a fixed compact set was proved in\[[44](https://arxiv.org/html/2605.23156#bib.bib44)\]\. The requirement thatϕ\\phiis asymptotically polynomial at±∞\\pm\\inftyguarantees the universality of neural networks on non\-compact domains\[[34](https://arxiv.org/html/2605.23156#bib.bib34)\]\. Representative activation functions meeting these conditions include ReLU, Softplus, and Sigmoid, among others\.

### 4\.2Graph functions

Additional details about the constructions in this section can be found in Appendix[D\.1](https://arxiv.org/html/2605.23156#A4.SS1)\. Recall the divisibility ordering from the duplication embedding in Example[2\.2](https://arxiv.org/html/2605.23156#S2.Thmtheorem2)\. We consider a consistent sequence indexed with this ordering and given by𝕍dupG=\{\(ℝsymn×n\),\(φN,n\),\(Sn\)\},\\mathbb\{V\}\_\{\\mathrm\{dup\}\}^\{G\}=\\\{\(\\mathbb\{R\}\_\{\\text\{sym\}\}^\{n\\times n\}\),\(\\varphi\_\{N,n\}\),\(\\mathrm\{S\}\_\{n\}\)\\\},whereℝsymn×n\\mathbb\{R\}^\{n\\times n\}\_\{\\text\{sym\}\}denotes the space ofn×nn\\times nsymmetric matrices with real entries, andφN,n\\varphi\_\{N,n\}converts each entry intoN/n×N/nN/n\\times N/nblocks with the same entry\. We endowℝsymn×n\\mathbb\{R\}\_\{\\text\{sym\}\}^\{n\\times n\}with the cut norm,‖A‖□=1n2​maxI,J⊆\[n\]⁡\|∑i∈I,j∈JAi​j\|\\\|A\\\|\_\{\\square\}=\\tfrac\{1\}\{n^\{2\}\}\\max\_\{I,J\\subseteq\[n\]\}\|\\sum\_\{i\\in I,j\\in J\}A\_\{ij\}\|\. The limit spaceV∞¯\\overline\{V\_\{\\infty\}\}can be identified with the space of kernels𝒲\\mathcal\{W\}, i\.e\., measurable symmetric functionsW:\[0,1\]2→ℝW\\colon\[0,1\]^\{2\}\\rightarrow\\mathbb\{R\}equipped with the cut norm∥⋅∥□\\\|\\cdot\\\|\_\{\\square\}\. The symmetrized metric \([1](https://arxiv.org/html/2605.23156#S2.E1)\) corresponds to the so\-called cut metricδ□\.\\delta\_\{\\square\}\.This limit space has been widely studied in the literature on large graphs or graphons\[[50](https://arxiv.org/html/2605.23156#bib.bib50)\]\. Intuitively, the functionsWWmodel adjacency matrices of infinite\-node graphs\.

##### Step 1 \(Compact sets\)\.

As is standard in the graphon literature, we consider graphs with bounded edge weights; in particular, we restrict them to the unit interval\[0,1\]\[0,1\]\. With this restriction, we focus on the subset of orbits corresponding to graphons, that is,

𝒲0≔\{W∈𝒲∣range\(W\)⊆\[0,1\]\}/∼,\{\\mathcal\{W\}\_\{0\}\}\\coloneqq\\\{W\\in\\mathcal\{W\}\\mid\\operatorname\{range\}\\left\(W\\right\)\\subseteq\[0,1\]\\\}\\ /\\sim,\(7\)whereW1∼W2W\_\{1\}\\sim W\_\{2\}wheneverδ□​\(W1,W2\)=0\\delta\_\{\\square\}\(W\_\{1\},W\_\{2\}\)=0\[[50](https://arxiv.org/html/2605.23156#bib.bib50)\]\. This is the collection of all limits of simple graphs, i\.e\., graphs with edge weights in\{0,1\}\\\{0,1\\\}\. This entire graphon space is compact\.

###### Proposition 4\.7\(Compactness of𝒲0\{\\mathcal\{W\}\_\{0\}\},\[[50](https://arxiv.org/html/2605.23156#bib.bib50)\]\)\.

The space𝒲0\{\\mathcal\{W\}\_\{0\}\}endowed with the cut metricδ□\\delta\_\{\\square\}is compact\.

##### Step 2 \(Continuity\)\.

Achieving continuity in the cut metric is subtle\. In particular, pointwise nonlinearities are not continuous in cut metric: for any pointwise mapϕ\\phi, the operatorW↦ϕ​\(W\)W\\mapsto\\phi\(W\)is continuous with respect to∥⋅∥□\\\|\\cdot\\\|\_\{\\square\}if and only ifϕ\\phiis linear\[[46](https://arxiv.org/html/2605.23156#bib.bib46)\]\. Since continuity with respect to∥⋅∥□\\\|\\cdot\\\|\_\{\\square\}on the limit space is equivalent to continuity with respect toδ□\\delta\_\{\\square\}on the orbit space \(Lemma[A\.4](https://arxiv.org/html/2605.23156#A1.Thmtheorem4)\), this precludes using pointwise activations\. As it turns out, many existing graph architectures are discontinuous in\(𝒲0,δ□\)\(\{\\mathcal\{W\}\_\{0\}\},\\delta\_\{\\square\}\)\[[4](https://arxiv.org/html/2605.23156#bib.bib4),[5](https://arxiv.org/html/2605.23156#bib.bib5)\]\. One way to circumvent this issue is to coarsen the topology to make pointwise nonlinearities continuous, as in\[[46](https://arxiv.org/html/2605.23156#bib.bib46)\]\. In contrast, we work directly with the cut distanceδ□\\delta\_\{\\square\}, which metrizes a standard notion of graph limits\[[50](https://arxiv.org/html/2605.23156#bib.bib50)\]\.

The continuity with respect to this distance is closely tied to*homomorphism densities*\. Intuitively, these are functionals that quantify how frequently a fixed motif \(a finite graph\) appears in a graphon\. Formally, given a motifF=\(V,E\)F=\(V,E\)and a graphonW∈𝒲0W\\in\{\\mathcal\{W\}\_\{0\}\}, the homomorphism density is

t​\(F,W\)=∫\[0,1\]\|V\|∏i​j∈EW​\(xi,xj\)​∏i∈Vd​xi\.t\(F,W\)=\\int\_\{\[0,1\]^\{\|V\|\}\}\\prod\_\{ij\\in E\}W\\left\(x\_\{i\},x\_\{j\}\\right\)\\prod\_\{i\\in V\}dx\_\{i\}~\.\(8\)In turn, the linear span of homomorphism densities of simple graphs,ℋ​𝒟≔span⁡\{t​\(F,⋅\)∣F​is simple\},\\mathcal\{HD\}\\coloneqq\\operatorname\{span\}\\\{\\,t\(F,\\cdot\)\\mid F\\;\\text\{is simple\}\\,\\\},is dense in the space of continuous functionsC​\(𝒲0,δ□\)C\\left\(\{\\mathcal\{W\}\_\{0\}\},\\delta\_\{\\square\}\\right\)\[[50](https://arxiv.org/html/2605.23156#bib.bib50)\]\. Thus, it suffices to find a model class that can be expressed as linear combinations of homomorphism densities of simple graphs\. Importantly, the desired model class must also not contain any homomorphism densities of non\-simple graphs, as these are discontinuous in the cut metric\[[50](https://arxiv.org/html/2605.23156#bib.bib50)\]\. With these considerations in mind, we propose the following family of models

hma,b​\(W\)=∫\[0,1\]m∏1≤i<j≤m\(ai​j​W​\(xi,xj\)\+bi​j\)​∏ℓ=1md​xℓ,h\_\{m\}^\{a,b\}\(W\)=\\int\_\{\[0,1\]^\{m\}\}\\prod\_\{1\\leq i<j\\leq m\}\\left\(a\_\{ij\}W\(x\_\{i\},x\_\{j\}\)\+b\_\{ij\}\\right\)\\prod\_\{\\ell=1\}^\{m\}dx\_\{\\ell\},\(9\)whereai​j,bi​j∈ℝa\_\{ij\},b\_\{ij\}\\in\\mathbb\{R\}are parameters\. In particular, the restrictioni<ji<jensures that only simple motifs can be parameterized\. Based on this, we define the function class

ℱ𝒲≔span​\{hma,b​\(⋅\)∣m∈ℕ,m≥2,a,b∈ℝsymm×m\}\.\\mathcal\{F\}\_\{\\mathcal\{W\}\}\\coloneqq\\mathrm\{span\}\\\{h\_\{m\}^\{a,b\}\(\\cdot\)\\mid m\\in\\mathbb\{N\},\\,m\\geq 2,\\,a,b\\in\\mathbb\{R\}^\{m\\times m\}\_\{\\text\{sym\}\}\\\}\.\(10\)
As we detail in Appendix[D\.4](https://arxiv.org/html/2605.23156#A4.SS4), this family can be represented by a neural\-network\-like deep architecture, making it directly amenable to gradient\-based optimization\. This property is informally stated as follows\.

###### Theorem 4\.8\(Informal\)\.

The family of functionsℱ𝒲\\mathcal\{F\}\_\{\\mathcal\{W\}\}can be parameterized using a neural network architecture\.

The neural network architecture in this result is somewhat involved; to streamline the presentation, we defer to Appendix[D\.4](https://arxiv.org/html/2605.23156#A4.SS4)\. We mention in passing that the nonlinearities correspond to tensor contractions, which have been proposed in the past to enhance the expressivity of invariant architectures\[[51](https://arxiv.org/html/2605.23156#bib.bib51)\]\. Our discussion thus far yields the following result; its formal proof appears in Appendix[D\.2](https://arxiv.org/html/2605.23156#A4.SS2)\.

###### Theorem 4\.9\(Continuity ofℱ𝒲\\mathcal\{F\}\_\{\\mathcal\{W\}\}\)\.

All functions inℱ𝒲\\mathcal\{F\}\_\{\\mathcal\{W\}\}are continuous with respect toδ□\\delta\_\{\\square\}\.

##### Step 3 \(Universality\)\.

Separating points in the graphon space\(𝒲0,δ□\)\(\{\\mathcal\{W\}\_\{0\}\},\\delta\_\{\\square\}\)corresponds to distinguishing graphs up to weak isomorphism—a fundamental task in the study of graph neural networks\. Prior work on the expressive power of standard GNN architectures\[[37](https://arxiv.org/html/2605.23156#bib.bib37),[39](https://arxiv.org/html/2605.23156#bib.bib39),[40](https://arxiv.org/html/2605.23156#bib.bib40)\]has linked their expressivity to the Weisfeiler–Lehman \(WL\) graph isomorphism test\[[38](https://arxiv.org/html/2605.23156#bib.bib38)\], which fails to distinguish certain non\-isomorphic graphs\. In contrast, the following theorem establishes that our model class \([10](https://arxiv.org/html/2605.23156#S4.E10)\) is universal; its proof is deferred to Appendix[D\.3](https://arxiv.org/html/2605.23156#A4.SS3)\.

###### Theorem 4\.10\(Universality ofℱ𝒲\\mathcal\{F\}\_\{\\mathcal\{W\}\}\)\.

The class of functionsℱ𝒲\\mathcal\{F\}\_\{\\mathcal\{W\}\}in \([10](https://arxiv.org/html/2605.23156#S4.E10)\) is dense inC​\(𝒲0,δ□\)C\\left\(\{\\mathcal\{W\}\_\{0\}\},\\delta\_\{\\square\}\\right\)\.

The proof of Theorems[4\.9](https://arxiv.org/html/2605.23156#S4.Thmtheorem9)and[4\.10](https://arxiv.org/html/2605.23156#S4.Thmtheorem10)essentially proceed by expanding out the product in the definition of our model \([9](https://arxiv.org/html/2605.23156#S4.E9)\) and observing that it is a linear combination of homomorphism densities\. By appropriately choosing the parametersa,ba,bin \([9](https://arxiv.org/html/2605.23156#S4.E9)\), we further show that all homomorphism densities lie within our parametric familyℱ𝒲\\mathcal\{F\}\_\{\\mathcal\{W\}\}\.

We emphasize at this point that there are two key advantages to working with the parametrization \([9](https://arxiv.org/html/2605.23156#S4.E9)\) over directly learning coefficients in a linear combination of homomorphism densities, as done in\[[41](https://arxiv.org/html/2605.23156#bib.bib41),[42](https://arxiv.org/html/2605.23156#bib.bib42)\]for example\. The first advantage pertains to the number of parameters needed to represent all homomorphism densities of simple graphs onmmnodes\. While the number of such densities scales as𝒪​\(2\(m2\)/m\!\)\\mathcal\{O\}\\left\(2^\{\\binom\{m\}\{2\}\}/m\!\\right\)and these densities are all linearly independent, our model is able to represent all of them using onlym​\(m−1\)m\(m\-1\)parameters\. The second advantage pertains to the implementation and training of our model\. Learning coefficients in a linear combination of homomorphism densities requires computing each homomorphism density separately for every motif and every input graph, which is costly\. Our parameterization avoids this by encoding graph patterns implicitly through the learnable coefficientsa,ba,bin \([9](https://arxiv.org/html/2605.23156#S4.E9)\)\. Moreover, our model is directly amenable to gradient\-based optimization thanks to its neural network\-like decomposition in Appendix[D\.4](https://arxiv.org/html/2605.23156#A4.SS4)\.

The key enabler of universality here is the use of high\-order tensors as hidden layers; the order of these tensors—as well as the width and depth of our networks—grows withmm\. Whether universal approximation is achievable with fixed width or depth remains an interesting open question\.

### 4\.3Point cloud functions

Details of the constructions in this section appear in Appendix[E\.1](https://arxiv.org/html/2605.23156#A5.SS1)\. We consider a consistent sequence analogous to the one used for sets \(measures\), namely𝕍dupP=\{\(ℝn×k\),\(φN,n⊕k\),\(Sn×O​\(k\)\)\}\\mathbb\{V\}\_\{\\mathrm\{dup\}\}^\{P\}=\\\{\(\\mathbb\{R\}^\{n\\times k\}\),\(\\varphi\_\{N,n\}^\{\\oplus k\}\),\(\\mathrm\{S\}\_\{n\}\\times\\mathrm\{O\}\(k\)\)\\\}, whereℝn×k\\mathbb\{R\}^\{n\\times k\}represents sets ofnnpoints inℝk\\mathbb\{R\}^\{k\}for fixedk∈ℕk\\in\{\\mathbb\{N\}\}, and the mapsφN,n\\varphi\_\{N,n\}are duplication embeddings\. The key difference is that the group is nowSn×O​\(k\)\\mathrm\{S\}\_\{n\}\\times\\mathrm\{O\}\(k\), whereSn\\mathrm\{S\}\_\{n\}is the permutation group andO​\(k\)\\mathrm\{O\}\(k\)is the orthogonal group, acting via\(g,h\)⋅X=g​X​h⊤\.\(g,h\)\\cdot X=gXh^\{\\top\}\.We endowℝk\\mathbb\{R\}^\{k\}with the Euclidean norm∥⋅∥ℝk\\\|\\cdot\\\|\_\{\\mathbb\{R\}^\{k\}\}, and equip eachVnV\_\{n\}with the normalizedℓ2\\ell\_\{2\}norm; the limit space can then be identified withV∞¯=L2​\(\[0,1\];ℝk\)\.\\overline\{V\_\{\\infty\}\}=L^\{2\}\(\[0,1\];\\mathbb\{R\}^\{k\}\)\.

Since equivariant operators appear repeatedly throughout, we define the family

LEk→l≔\{L∈𝔹​\(k,l\)∣L​is​S\[0,1\]​\-equivariant\},\\texttt\{LE\}\_\{k\\rightarrow l\}\\coloneqq\\left\\\{L\\in\\mathbb\{B\}\(k,l\)\\mid L\\text\{ is \}S\_\{\[0,1\]\}\\text\{\-equivariant\}\\right\\\},\(11\)where𝔹​\(k,l\)=ℬ​\(L2​\(\[0,1\]k\),L2​\(\[0,1\]l\)\)\\mathbb\{B\}\(k,l\)=\\mathcal\{B\}\\left\(L^\{2\}\(\[0,1\]^\{k\}\),L^\{2\}\(\[0,1\]^\{l\}\)\\right\)denotes the space of bounded linear operators andS\[0,1\]S\_\{\[0,1\]\}is the group of measure\-preserving bijectionsφ:\[0,1\]→\[0,1\]\\varphi\\colon\[0,1\]\\to\[0,1\]\. We say thatLLisS\[0,1\]S\_\{\[0,1\]\}\-equivariant if, for everyφ∈S\[0,1\]\\varphi\\in S\_\{\[0,1\]\}and everyW∈L2​\(\[0,1\]k\)W\\in L^\{2\}\(\[0,1\]^\{k\}\), one hasL​\(Wφ\)=L​\(W\)φL\(W^\{\\varphi\}\)=L\(W\)^\{\\varphi\}almost everywhere, whereWφ​\(x1,…,xk\)≔W​\(φ​\(x1\),…,φ​\(xk\)\)W^\{\\varphi\}\(x\_\{1\},\\dots,x\_\{k\}\)\\coloneqq W\(\\varphi\(x\_\{1\}\),\\dots,\\varphi\(x\_\{k\}\)\)\.

##### Step 1 \(Compact sets\)\.

For this setting, we consider the set of orbit closures arising from a set of bounded functions in the limit spaceV∞¯\\overline\{V\_\{\\infty\}\}, namely

KR≔\{X∈L2​\(\[0,1\];ℝk\)∣‖X‖L∞≤R\},K\_\{R\}\\coloneqq\\\{X\\in L^\{2\}\\left\(\[0,1\];\\mathbb\{R\}^\{k\}\\right\)\\mid\\\|X\\\|\_\{L^\{\\infty\}\}\\leq R\\\},whereR\>0R\>0is a positive constant\. The next proposition establishes that the space of orbit closures ofKRK\_\{R\}, which we denote byKR/G∞K\_\{R\}/G\_\{\\infty\}, is compact inV∞¯/G∞\\overline\{V\_\{\\infty\}\}/G\_\{\\infty\}\. The proof is deferred to Appendix[E\.2](https://arxiv.org/html/2605.23156#A5.SS2)\.

###### Proposition 4\.11\(Compactness ofKR/G∞K\_\{R\}/G\_\{\\infty\}\)\.

The set of orbit closuresKR/G∞K\_\{R\}/G\_\{\\infty\}is a compact subset ofV∞¯/G∞\\overline\{V\_\{\\infty\}\}/G\_\{\\infty\}\.

##### Step 2 \(Continuity\)\.

Next, we introduce a class of continuous models on the space of orbit closures\. The architecture decomposes into two stages: given a point cloudXX, we first form its Gram \(inner\-product\) matrixX⊤​XX^\{\\top\}X, and then apply a graph neural network \(GNN\) to this matrix\. Intuitively, the Gram map enforces invariance toO​\(k\)\\mathrm\{O\}\(k\), while the GNN provides invariance to permutationsSnS\_\{n\}\. Closely related architectures were first proposed in\[[16](https://arxiv.org/html/2605.23156#bib.bib16)\]\.

Let us formally define the architecture in the orbit space\. For the first stage, we map eachX∈KR/G∞X\\in K\_\{R\}/G\_\{\\infty\}to a graphonWX:\[0,1\]2→\[0,1\]W\_\{X\}\\colon\[0,1\]^\{2\}\\to\[0,1\]defined by

WX​\(x,y\)≔12​R2​⟨X​\(x\),X​\(y\)⟩\+12,W\_\{X\}\(x,y\)\\coloneqq\\frac\{1\}\{2R^\{2\}\}\\langle X\(x\),X\(y\)\\rangle\+\\frac\{1\}\{2\}~,\(12\)where we interpretXXas any representative in its equivalence class\. Notice thatWXW\_\{X\}is a symmetric measurable function, and so can be seen as a graphon\. We endow the space of graphons with theδ2\\delta\_\{2\}metric given by

δ2​\(W,W~\)≔infφ∈S\[0,1\]‖W−W~φ‖2\.\\delta\_\{2\}\(W,\\widetilde\{W\}\)\\coloneqq\\inf\_\{\\varphi\\in S\_\{\[0,1\]\}\}\\\|W\-\\widetilde\{W\}^\{\\varphi\}\\\|\_\{2\}\.For the second stage, we use Invariant Graphon Networks \(IGNs\)\[[46](https://arxiv.org/html/2605.23156#bib.bib46)\], i\.e\., functionsIϱ,M,L,b:𝒲0→ℝI\_\{\\varrho,M,L,b\}\\colon\\mathcal\{W\}\_\{0\}\\to\\mathbb\{R\}of the form

Iϱ,M,L,b​\(W\)≔∑m=1MLm\(2\)​\(ϱ​\(Lm\(1\)​\(W\)\+bm\(1\)\)\)\+b\(2\),I\_\{\\varrho,M,L,b\}\(W\)\\coloneqq\\sum\_\{m=1\}^\{M\}L\_\{m\}^\{\(2\)\}\\\!\\left\(\\varrho\\\!\\left\(L\_\{m\}^\{\(1\)\}\(W\)\+b\_\{m\}^\{\(1\)\}\\right\)\\right\)\+b^\{\(2\)\},\(13\)whereM∈ℕ0M\\in\{\{\\mathbb\{N\}\}\}\_\{0\},Lm\(1\)∈LE2→kmL\_\{m\}^\{\(1\)\}\\in\\texttt\{LE\}\_\{2\\rightarrow k\_\{m\}\},Lm\(2\)∈LEkm→0L\_\{m\}^\{\(2\)\}\\in\\texttt\{LE\}\_\{k\_\{m\}\\rightarrow 0\},bm\(1\),b\(2\)∈ℝb\_\{m\}^\{\(1\)\},b^\{\(2\)\}\\in\\mathbb\{R\}, andkm∈ℕk\_\{m\}\\in\{\\mathbb\{N\}\}for eachm∈\{1,…,M\}m\\in\\\{1,\\ldots,M\\\}\. Hereϱ:ℝ→ℝ\\varrho\\colon\\mathbb\{R\}\\to\\mathbb\{R\}is a nonlinearity acting pointwise\.

Altogether, we consider the family of models given by

ℱP​Cϱ≔\{X↦Iϱ,M,L,b​\(WX\)\}\.\\mathcal\{F\}\_\{PC\}^\{\\varrho\}\\coloneqq\\\{X\\mapsto I\_\{\\varrho,M,L,b\}\(W\_\{X\}\)\\\}\.\(14\)Here the architectural hyperparameters and parameters \(i\.e\.,MM, the operatorsLL, and the biasesbb\) range over the admissible sets described after \([13](https://arxiv.org/html/2605.23156#S4.E13)\); we omit the full specification for brevity\. The next theorem shows that this architecture is indeed continuous with respect to the symmetrized distance\. The proof is in Appendix[E\.3](https://arxiv.org/html/2605.23156#A5.SS3)\.

###### Theorem 4\.12\(Continuity ofℱP​Cϱ\\mathcal\{F\}\_\{PC\}^\{\\varrho\}\)\.

Letϱ\\varrhobe continuous\. Then, all models inℱP​Cϱ\\mathcal\{F\}\_\{PC\}^\{\\varrho\}are continuous with respect to the symmetrized metric\.

##### Step 3 \(Universality\)\.

To close, we show that our model class for point clouds \([14](https://arxiv.org/html/2605.23156#S4.E14)\) is universal; the proof is deferred to Appendix[E\.4](https://arxiv.org/html/2605.23156#A5.SS4)\.

###### Theorem 4\.13\(Universality ofℱP​Cϱ\\mathcal\{F\}\_\{PC\}^\{\\varrho\}\)\.

Letϱ\\varrhobe continuous and non\-polynomial\. Then,ℱP​Cϱ\\mathcal\{F\}\_\{PC\}^\{\\varrho\}is dense inC​\(KR/G∞\)C\\left\(K\_\{R\}/G\_\{\\infty\}\\right\)\.

The above theorem gives universality in the space of continuous functions of point clouds that are invariant under rotations of the cloud\. If we center the input point cloud before applying our proposed architecture, we also obtain universality in the smaller space of continuous functions invariant under all rigid motions of their inputs\.

## 5Conclusions and future work

To summarize, we have \(i\) formalized universality for any\-dimensional invariant architectures, \(ii\) proposed a general recipe for constructing and certifying any\-dimensional universal model classes, and \(iii\) used this recipe to prove universality for several concrete families\. A key limitation of our framework is that it applies only to scalar\-valued outputs\. Many practically relevant any\-dimensional architectures have both input and output sizes growing and are naturally equivariant as opposed to invariant\. In this setting, the problem no longer reduces to studying functionals on orbit spaces\. Developing tools that exploit such equivariance to obtain universality results for more general architectures is an important direction for future work\. Another limitation is that we only focus on dense graphs in our current instantiations\. There are several notions of limits for sparse graphs including\[[52](https://arxiv.org/html/2605.23156#bib.bib52),[53](https://arxiv.org/html/2605.23156#bib.bib53),[54](https://arxiv.org/html/2605.23156#bib.bib54)\], and applying our framework to these limits is another interesting direction\.

## References

- \[1\]Manzil Zaheer, Satwik Kottur, Siamak Ravanbakhsh, Barnabas Poczos, Russ R Salakhutdinov, and Alexander J Smola\.Deep sets\.Advances in neural information processing systems, 30, 2017\.
- \[2\]Marco Gori, Gabriele Monfardini, and Franco Scarselli\.A new model for learning in graph domains\.InProceedings\. 2005 IEEE international joint conference on neural networks, 2005\., volume 2, pages 729–734\. IEEE, 2005\.
- \[3\]Charles R Qi, Hao Su, Kaichun Mo, and Leonidas J Guibas\.Pointnet: Deep learning on point sets for 3d classification and segmentation\.InProceedings of the IEEE conference on computer vision and pattern recognition, pages 652–660, 2017\.
- \[4\]Haggai Maron, Ethan Fetaya, Nimrod Segol, and Yaron Lipman\.On the universality of invariant networks\.InInternational conference on machine learning, pages 4363–4371\. PMLR, 2019\.
- \[5\]Nicolas Keriven and Gabriel Peyré\.Universal invariant and equivariant graph neural networks\.Advances in neural information processing systems, 32, 2019\.
- \[6\]Dmitry Yarotsky\.Universal approximations of invariant maps by neural networks\.Constructive Approximation, 55\(1\):407–474, 2022\.
- \[7\]Eitan Levin, Yuxin Ma, Mateo Díaz, and Soledad Villar\.On transferring transferability: Towards a theory for size generalization\.arXiv preprint arXiv:2505\.23599, 2025\.
- \[8\]Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner\.Gradient\-based learning applied to document recognition\.Proceedings of the IEEE, 86\(11\):2278–2324, 2002\.
- \[9\]Mahdi Hashemi\.Enlarging smaller images before inputting into convolutional neural network: zero\-padding vs\. interpolation\.Journal of Big Data, 6\(1\):1–13, 2019\.
- \[10\]Sepp Hochreiter and Jürgen Schmidhuber\.Long short\-term memory\.Neural computation, 9\(8\):1735–1780, 1997\.
- \[11\]Richard Socher, Cliff C Lin, Chris Manning, and Andrew Y Ng\.Parsing natural scenes and natural language with recursive neural networks\.InProceedings of the 28th international conference on machine learning \(ICML\-11\), pages 129–136, 2011\.
- \[12\]Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin\.Attention is all you need\.Advances in neural information processing systems, 30, 2017\.
- \[13\]Lu Lu, Pengzhan Jin, Guofei Pang, Zhongqiang Zhang, and George Em Karniadakis\.Learning nonlinear operators via deeponet based on the universal approximation theorem of operators\.Nature machine intelligence, 3\(3\):218–229, 2021\.
- \[14\]Zongyi Li, Nikola Kovachki, Kamyar Azizzadenesheli, Burigede Liu, Kaushik Bhattacharya, Andrew Stuart, and Anima Anandkumar\.Fourier neural operator for parametric partial differential equations\.arXiv preprint arXiv:2010\.08895, 2020\.
- \[15\]Franco Scarselli, Marco Gori, Ah Chung Tsoi, Markus Hagenbuchner, and Gabriele Monfardini\.The graph neural network model\.IEEE transactions on neural networks, 20\(1\):61–80, 2008\.
- \[16\]Soledad Villar, David W Hogg, Kate Storey\-Fisher, Weichi Yao, and Ben Blum\-Smith\.Scalars are universal: Equivariant machine learning, structured like classical physics\.Advances in neural information processing systems, 34:28848–28863, 2021\.
- \[17\]Ben Blum\-Smith, Ningyuan Huang, Marco Cuturi, and Soledad Villar\.Functions on symmetric matrices and point clouds via lightweight invariant features from galois theory\.SIAM Journal on Applied Algebra and Geometry, 9\(4\):902–938, 2025\.
- \[18\]Eitan Levin and Venkat Chandrasekaran\.Free descriptions of convex sets\.arXiv preprint arXiv:2307\.04230, 2023\.
- \[19\]Eitan Levin and Mateo Díaz\.Any\-dimensional equivariant neural networks\.InInternational Conference on Artificial Intelligence and Statistics, pages 2773–2781\. PMLR, 2024\.
- \[20\]Mateo Díaz, Dmitriy Drusvyatskiy, Jack Kendrick, and Rekha R Thomas\.Invariant kernels: Rank stabilization and generalization across dimensions\.arXiv preprint arXiv:2502\.01886, 2025\.
- \[21\]Thomas Church and Benson Farb\.Representation theory and homological stability\.Advances in Mathematics, 245:250–314, 2013\.
- \[22\]Thomas Church, Jordan S Ellenberg, and Benson Farb\.Fi\-modules and stability for representations of symmetric groups\.Duke Mathematical Journal, 164\(9\), 2015\.
- \[23\]Luana Ruiz, Luiz Chamon, and Alejandro Ribeiro\.Graphon neural networks and the transferability of graph neural networks\.Advances in Neural Information Processing Systems, 33:1702–1712, 2020\.
- \[24\]Ron Levie, Wei Huang, Lorenzo Bucci, Michael Bronstein, and Gitta Kutyniok\.Transferability of spectral graph convolutional neural networks\.Journal of Machine Learning Research, 22\(272\):1–59, 2021\.
- \[25\]Nicolas Keriven, Alberto Bietti, and Samuel Vaiter\.Convergence and stability of graph convolutional networks on large random graphs\.Advances in Neural Information Processing Systems, 33:21512–21523, 2020\.
- \[26\]Luana Ruiz, Luiz FO Chamon, and Alejandro Ribeiro\.Transferability properties of graph neural networks\.IEEE Transactions on Signal Processing, 71:3474–3489, 2023\.
- \[27\]Sohir Maskey, Ron Levie, and Gitta Kutyniok\.Transferability of graph neural networks: an extended graphon approach\.Applied and Computational Harmonic Analysis, 63:48–83, 2023\.
- \[28\]Matthieu Cordonnier, Nicolas Keriven, Nicolas Tremblay, and Samuel Vaiter\.Convergence of message passing graph neural networks with generic aggregation on random graphs\.InGSP 2023\-6th Graph Signal Processing workshop, pages 1–3, 2023\.
- \[29\]Chen Cai and Yusu Wang\.Convergence of invariant graph networks\.InInternational Conference on Machine Learning, pages 2457–2484\. PMLR, 2022\.
- \[30\]George Cybenko\.Approximation by superpositions of a sigmoidal function\.Mathematics of control, signals and systems, 2\(4\):303–314, 1989\.
- \[31\]Kurt Hornik, Maxwell Stinchcombe, and Halbert White\.Multilayer feedforward networks are universal approximators\.Neural networks, 2\(5\):359–366, 1989\.
- \[32\]Moshe Leshno, Vladimir Ya Lin, Allan Pinkus, and Shimon Schocken\.Multilayer feedforward networks with a nonpolynomial activation function can approximate any function\.Neural networks, 6\(6\):861–867, 1993\.
- \[33\]Kurt Hornik\.Approximation capabilities of multilayer feedforward networks\.Neural networks, 4\(2\):251–257, 1991\.
- \[34\]Teun DH van Nuland\.Noncompact uniform universal approximation\.Neural Networks, 173:106181, 2024\.
- \[35\]Ariel Neufeld and Philipp Schmocker\.Universal approximation results for neural networks with non\-polynomial activation function over non\-compact domains\.arXiv preprint arXiv:2410\.14759, 2024\.
- \[36\]Ahmed Abdeljawad and Thomas Dittrich\.Weighted sobolev approximation rates for neural networks on unbounded domains\.arXiv preprint arXiv:2411\.04108, 2024\.
- \[37\]Keyulu Xu, Weihua Hu, Jure Leskovec, and Stefanie Jegelka\.How powerful are graph neural networks?arXiv preprint arXiv:1810\.00826, 2018\.
- \[38\]Andrei Leman and Boris Weisfeiler\.A reduction of a graph to a canonical form and an algebra arising during this reduction\.Nauchno\-Technicheskaya Informatsiya, 2\(9\):12–16, 1968\.
- \[39\]Haggai Maron, Heli Ben\-Hamu, Hadar Serviansky, and Yaron Lipman\.Provably powerful graph networks\.Advances in neural information processing systems, 32, 2019\.
- \[40\]Christopher Morris, Martin Ritzert, Matthias Fey, William L Hamilton, Jan Eric Lenssen, Gaurav Rattan, and Martin Grohe\.Weisfeiler and leman go neural: Higher\-order graph neural networks\.InProceedings of the AAAI conference on artificial intelligence, volume 33, pages 4602–4609, 2019\.
- \[41\]Takanori Maehara and Hoang NT\.A simple proof of the universality of invariant/equivariant graph neural networks\.arXiv preprint arXiv:1910\.03802, 2019\.
- \[42\]Hoang Nguyen and Takanori Maehara\.Graph homomorphism convolution\.InInternational Conference on Machine Learning, pages 7306–7316\. PMLR, 2020\.
- \[43\]Akiyoshi Sannai, Yuuki Takai, and Matthieu Cordonnier\.Universal approximations of permutation invariant/equivariant functions by deep neural networks\.arXiv preprint arXiv:1903\.01939, 2019\.
- \[44\]Christian Bueno and Alan Hylton\.On the representation power of set pooling networks\.Advances in Neural Information Processing Systems, 34:17170–17182, 2021\.
- \[45\]Aaron Zweig and Joan Bruna\.A functional perspective on learning symmetric functions with neural networks\.InInternational Conference on Machine Learning, pages 13023–13032\. PMLR, 2021\.
- \[46\]Daniel Herbst and Stefanie Jegelka\.Higher\-order graphon neural networks: Approximation and cut distance\.arXiv preprint arXiv:2503\.14338, 2025\.
- \[47\]Walter Rudin\.Functional Analysis\.International Series in Pure and Applied Mathematics\. McGraw\-Hill, New York, 2 edition, 1991\.
- \[48\]Joseph Diestel\.Sequences and series in Banach spaces, volume 92\.Springer Science & Business Media, 2012\.
- \[49\]Luigi Ambrosio, Nicola Gigli, and Giuseppe Savaré\.Gradient flows: in metric spaces and in the space of probability measures\.Springer, 2005\.
- \[50\]László Lovász\.Large networks and graph limits\.American Mathematical Soc\., 2012\.
- \[51\]Marc Finzi, Max Welling, and Andrew Gordon Wilson\.A practical method for constructing equivariant multilayer perceptrons for arbitrary matrix groups\.InInternational conference on machine learning, pages 3318–3328\. PMLR, 2021\.
- \[52\]Christian Borgs, Jennifer Chayes, Henry Cohn, and Yufei Zhao\.AnLpL^\{p\}theory of sparse graph convergence I: Limits, sparse random graph models, and power law distributions\.Transactions of the American Mathematical Society, 372\(5\):3019–3062, 2019\.
- \[53\]Christian Borgs, Jennifer T\. Chayes, Henry Cohn, and Nina Holden\.Sparse exchangeable graphs and their limits via graphon processes\.Journal of Machine Learning Research, 18\(210\):1–71, 2018\.
- \[54\]Ágnes Backhausz and Balázs Szegedy\.Action convergence of operators and graphs\.Canadian Journal of Mathematics, 74\(1\):72–121, 2022\.
- \[55\]Stephen Willard\.General topology\.Courier Corporation, 2012\.
- \[56\]Walter Rudin\.Principles of mathematical analysis\.3rd ed\., 1976\.
- \[57\]Cédric Villani et al\.Optimal transport: old and new, volume 338\.Springer, 2008\.
- \[58\]Vladimir I Bogachev\.Measure theory\.Springer, 2007\.
- \[59\]C\. Borgs, J\.T\. Chayes, L\. Lovász, V\.T\. Sós, and K\. Vesztergombi\.Convergent sequences of dense graphs I: Subgraph frequencies, metric properties and testing\.Advances in Mathematics, 219\(6\):1801–1851, 2008\.
- \[60\]Peter Diao, Dominique Guillot, Apoorva Khare, and Bala Rajaratnam\.Differential calculus on graphon space\.Journal of Combinatorial Theory, Series A, 133:183–227, 2015\.
- \[61\]Sashi Mohan Srivastava\.A course on Borel sets\.Springer, 1998\.
- \[62\]Stephen H Friedberg, Arnold J Insel, and Lawrence E Spence\.Linear algebra\.Prentice Hall, 1997\.

## Appendix AMissing details from Section[2](https://arxiv.org/html/2605.23156#S2)

In this section, we present missing details from Section[2](https://arxiv.org/html/2605.23156#S2)\. We start with a more detailed definition of a consistent sequence\.

###### Definition A\.1\(Detailed version of Def\.[2\.1](https://arxiv.org/html/2605.23156#S2.Thmtheorem1)\)\.

Aconsistent sequenceof group representations over directed poset\(ℕ,⪯\)\(\{\\mathbb\{N\}\},\\preceq\)is a triple𝕍=\{\(Vn\)n∈ℕ,\(φN,n\)n⪯N,\(Gn\)n∈ℕ\}\\mathbb\{V\}=\\\{\\left\(V\_\{n\}\\right\)\_\{n\\in\{\\mathbb\{N\}\}\},\\left\(\\varphi\_\{N,n\}\\right\)\_\{n\\preceq N\},\\left\(\\mathrm\{G\}\_\{n\}\\right\)\_\{n\\in\{\\mathbb\{N\}\}\}\\\}consisting of the following elements\.

1. 1\.\(Groups\) A sequence of groups\(Gn\)\\left\(\\mathrm\{G\}\_\{n\}\\right\)indexed byℕ\{\\mathbb\{N\}\}that embed into each other\. Specifically, whenevern⪯Nn\\preceq Nthere is an injective group homomorphismθN,n:Gn→GN\\theta\_\{N,n\}\\colon\\mathrm\{G\}\_\{n\}\\rightarrow\\mathrm\{G\}\_\{N\}with θi,i\\displaystyle\\theta\_\{i,i\}=idGifor all​i∈ℕ,\\displaystyle=\\operatorname\{id\}\_\{\\mathrm\{G\}\_\{i\}\}\\quad\\text\{ for all \}i\\in\{\\mathbb\{N\}\},θk,j∘θj,i\\displaystyle\\theta\_\{k,j\}\\circ\\theta\_\{j,i\}=θk,iwhenever​i⪯j⪯k​in​ℕ\.\\displaystyle=\\theta\_\{k,i\}\\quad\\ \\text\{ whenever \}i\\preceq j\\preceq k\\text\{ in \}\{\\mathbb\{N\}\}\.\\
2. 2\.\(Vector spaces\) A sequence of finite\-dimensional, real vector spaces\(Vn\)\\left\(V\_\{n\}\\right\)indexed byℕ\{\\mathbb\{N\}\}such that eachVnV\_\{n\}is aGn\\mathrm\{G\}\_\{n\}\-representation\.
3. 3\.\(Embeddings\) A collection of embeddings\(φN,n:Vn↪VN\)n⪯N\(\\varphi\_\{N,n\}\\colon V\_\{n\}\\hookrightarrow V\_\{N\}\)\_\{n\\preceq N\}such thatφN,n\\varphi\_\{N,n\}isGn\\mathrm\{G\}\_\{n\}\-equivariant, i\.e\., φN,n​\(g⋅v\)=θN,n​\(g\)⋅φN,n​\(v\)​for all​g∈Gn,v∈Vn\.\\varphi\_\{N,n\}\(g\\cdot v\)=\\theta\_\{N,n\}\(g\)\\cdot\\varphi\_\{N,n\}\(v\)\\text\{ for all \}g\\in\\mathrm\{G\}\_\{n\},v\\in V\_\{n\}\.and such that φi,i\\displaystyle\\varphi\_\{i,i\}=idVifor all​i∈ℕ,\\displaystyle=\\operatorname\{id\}\_\{V\_\{i\}\}\\quad\\text\{ for all \}i\\in\{\\mathbb\{N\}\},φk,j∘φj,i\\displaystyle\\varphi\_\{k,j\}\\circ\\varphi\_\{j,i\}=φk,iwhenever​i⪯j⪯k​in​ℕ\.\\displaystyle=\\varphi\_\{k,i\}\\quad\\text\{ whenever \}i\\preceq j\\preceq k\\text\{ in \}\{\\mathbb\{N\}\}\.\\

We can take direct sums of consistent sequences to obtain richer consistent sequences, as done in several of the examples of Section[4](https://arxiv.org/html/2605.23156#S4)\.

###### Definition A\.2\.

Thekk\-fold direct sum of𝕍\\mathbb\{V\}is defined as

𝕍⊕k≔\{\(Vn⊕k\),\(φN,n⊕k\),\(Gn\)\}\\mathbb\{V\}^\{\\oplus k\}\\coloneqq\\\{\(V\_\{n\}^\{\\oplus k\}\),\(\\varphi\_\{N,n\}^\{\\oplus k\}\),\\left\(\\mathrm\{G\}\_\{n\}\\right\)\\\}whereVn⊕dV\_\{n\}^\{\\oplus d\}denotes the direct sum ofddcopies ofVnV\_\{n\}andφN,n⊕d:Vn⊕d→VN⊕d\\varphi\_\{N,n\}^\{\\oplus d\}:V\_\{n\}^\{\\oplus d\}\\rightarrow V\_\{N\}^\{\\oplus d\}is defined by applyingφN,n\\varphi\_\{N,n\}to each component\. The groupGn\\mathrm\{G\}\_\{n\}acts onVn⊕dV\_\{n\}^\{\\oplus d\}by simultaneously acting on every copy ofVnV\_\{n\}, i\.e\.g⋅\(v1,…,vd\)≔\(g⋅v1,…,g⋅vd\)g\\cdot\\left\(v\_\{1\},\\ldots,v\_\{d\}\\right\)\\coloneqq\\left\(g\\cdot v\_\{1\},\\ldots,g\\cdot v\_\{d\}\\right\)\.

The next two propositions show that compatibility \(Definition[2\.4](https://arxiv.org/html/2605.23156#S2.Thmtheorem4)\) is equivalent to the existence of a limiting extension\.

###### Proposition A\.3\(\[[7](https://arxiv.org/html/2605.23156#bib.bib7), Prop\. B\.6\]\)\.

Let𝕍=\{\(Vn\),\(φN,n\),\(Gn\)\}\\mathbb\{V\}=\\left\\\{\\left\(V\_\{n\}\\right\),\\left\(\\varphi\_\{N,n\}\\right\),\\left\(\\mathrm\{G\}\_\{n\}\\right\)\\right\\\}and𝕌=\{\(Un\),\(ψN,n\),\(Gn\)\}\{\\mathbb\{U\}\}=\\left\\\{\\left\(U\_\{n\}\\right\),\\left\(\\psi\_\{N,n\}\\right\),\\left\(\\mathrm\{G\}\_\{n\}\\right\)\\right\\\}be consistent sequences\. A sequence of maps\(fn:Vn→Un\)\\left\(f\_\{n\}:V\_\{n\}\\rightarrow U\_\{n\}\\right\)is compatible if, and only if, it extends to the limit, i\.e\., there exists aG∞G\_\{\\infty\}\-equivariant mapf∞:V∞→U∞f\_\{\\infty\}\\colon V\_\{\\infty\}\\rightarrow U\_\{\\infty\}such thatfn=f∞\|Vnf\_\{n\}=\\left\.f\_\{\\infty\}\\right\|\_\{V\_\{n\}\}for allnn\.

In particular, a sequence of norms∥⋅∥Vn\\\|\\cdot\\\|\_\{V\_\{n\}\}is compatible if and only if they extend to a norm∥⋅∥V∞\\\|\\cdot\\\|\_\{V\_\{\\infty\}\}preserved by the action ofG∞G\_\{\\infty\}\.

Finally, we show that continuity of invariant functions in the limit space is equivalent to continuity of the corresponding induced function on the orbit space\.

###### Lemma A\.4\.

Let∥⋅∥V∞\\\|\\cdot\\\|\_\{V\_\{\\infty\}\}be the extension of a compatible sequence of norms∥⋅∥Vn\\\|\\cdot\\\|\_\{V\_\{n\}\}\. Letf∞:V∞¯→ℝf\_\{\\infty\}\\colon\\overline\{V\_\{\\infty\}\}\\to\\mathbb\{R\}beG∞G\_\{\\infty\}\-invariant function, and letf¯∞\\bar\{f\}\_\{\\infty\}be the induced function on the space of orbit closuresV∞¯/G∞\\overline\{V\_\{\\infty\}\}/G\_\{\\infty\}\. Then,f∞f\_\{\\infty\}is continuous on\(V∞¯,∥⋅∥V∞\)\\left\(\\overline\{V\_\{\\infty\}\},\\\|\\cdot\\\|\_\{V\_\{\\infty\}\}\\right\)if, and only if,f¯∞\\bar\{f\}\_\{\\infty\}is continuous on\(V∞¯/G∞,d¯\)\\left\(\\overline\{V\_\{\\infty\}\}/G\_\{\\infty\},\\overline\{\\mathrm\{d\}\}\\right\)\.

###### Proof of Lemma[A\.4](https://arxiv.org/html/2605.23156#A1.Thmtheorem4)\.

Letπ:V∞¯→V∞¯/G∞\\pi\\colon\\overline\{V\_\{\\infty\}\}\\rightarrow\\overline\{V\_\{\\infty\}\}/G\_\{\\infty\}be the quotient mapπ​\(x\)=\[x\]\\pi\(x\)=\[x\]\. Becausef∞f\_\{\\infty\}isG∞G\_\{\\infty\}\-invariant, we can writef¯∞∘π\\bar\{f\}\_\{\\infty\}\\circ\\pifor somef¯∞:V∞¯/G∞→ℝ\\bar\{f\}\_\{\\infty\}\\colon\\overline\{V\_\{\\infty\}\}/G\_\{\\infty\}\\to\\mathbb\{R\}\. We prove the two implications\.

##### \(⇐\\Leftarrow\)

Supposef¯∞:V∞¯/G∞→ℝ\\bar\{f\}\_\{\\infty\}\\colon\\overline\{V\_\{\\infty\}\}/G\_\{\\infty\}\\rightarrow\\mathbb\{R\}is continuous with respect tod¯\\overline\{\\mathrm\{d\}\}\. Note thatπ\\piis continuous with respect to the norm∥⋅∥V∞\\\|\\cdot\\\|\_\{V\_\{\\infty\}\}onV∞¯\\overline\{V\_\{\\infty\}\}and symmetrized metricd¯\\overline\{\\mathrm\{d\}\}onV∞¯/G∞\\overline\{V\_\{\\infty\}\}/G\_\{\\infty\}, sinced¯​\(π​\(x\),π​\(y\)\)=infg∈G∞‖g⋅x−y‖V∞≤‖x−y‖V∞\\overline\{\\mathrm\{d\}\}\\left\(\\pi\(x\),\\pi\(y\)\\right\)=\\inf\_\{g\\in G\_\{\\infty\}\}\\\|g\\cdot x\-y\\\|\_\{V\_\{\\infty\}\}\\leq\\\|x\-y\\\|\_\{V\_\{\\infty\}\}\. Therefore, the compositionf∞=f¯∞∘πf\_\{\\infty\}=\\bar\{f\}\_\{\\infty\}\\circ\\piis continuous with respect to∥⋅∥V∞\\\|\\cdot\\\|\_\{V\_\{\\infty\}\}\.

##### \(⇒\\Rightarrow\)

Supposef∞:V∞¯→ℝf\_\{\\infty\}\\colon\\overline\{V\_\{\\infty\}\}\\rightarrow\\mathbb\{R\}is continuous with respect to∥⋅∥V∞\\\|\\cdot\\\|\_\{V\_\{\\infty\}\}\. Then for anyε\>0\\varepsilon\>0and any\[x\]∈V∞¯/G∞\[x\]\\in\\overline\{V\_\{\\infty\}\}/G\_\{\\infty\}there existsδ\>0\\delta\>0such that anyy∈V∞¯y\\in\\overline\{V\_\{\\infty\}\}with‖x−y‖V∞<δ\\\|x\-y\\\|\_\{V\_\{\\infty\}\}<\\deltasatisfies\|f∞​\(x\)−f∞​\(y\)\|<ε\.\|f\_\{\\infty\}\(x\)\-f\_\{\\infty\}\(y\)\|<\\varepsilon\.Now, consider any\[y′\]∈V∞¯/G∞\[y^\{\\prime\}\]\\in\\overline\{V\_\{\\infty\}\}/G\_\{\\infty\}satisfyingd¯​\(\[x\],\[y′\]\)<δ\\overline\{\\mathrm\{d\}\}\\left\(\[x\],\[y^\{\\prime\}\]\\right\)<\\delta, then by the definition ofd¯\\overline\{\\mathrm\{d\}\}, there exists an elementg0∈G∞g\_\{0\}\\in G\_\{\\infty\}such that‖g0⋅x−y′‖V∞<δ\\\|g\_\{0\}\\cdot x\-y^\{\\prime\}\\\|\_\{V\_\{\\infty\}\}<\\delta\. Since for anyg∈G∞g\\in G\_\{\\infty\}, the mapx↦g⋅xx\\mapsto g\\cdot xis an isometry onV∞¯\\overline\{V\_\{\\infty\}\}, we have‖x−g0−1⋅y′‖V∞=‖g0⋅x−y′‖V∞<δ\\\|x\-g\_\{0\}^\{\-1\}\\cdot y^\{\\prime\}\\\|\_\{V\_\{\\infty\}\}=\\\|g\_\{0\}\\cdot x\-y^\{\\prime\}\\\|\_\{V\_\{\\infty\}\}<\\delta, which implies\|f∞​\(x\)−f∞​\(g0−1⋅y′\)\|<ε\|f\_\{\\infty\}\(x\)\-f\_\{\\infty\}\(g\_\{0\}^\{\-1\}\\cdot y^\{\\prime\}\)\|<\\varepsilon\. Sincef∞f\_\{\\infty\}isG∞G\_\{\\infty\}\-invariant, we obtain\|f¯∞​\(\[x\]\)−f¯∞​\(\[y′\]\)\|=\|f∞​\(x\)−f∞​\(g0−1⋅y′\)\|<ε\|\\bar\{f\}\_\{\\infty\}\(\[x\]\)\-\\bar\{f\}\_\{\\infty\}\(\[y^\{\\prime\}\]\)\|=\|f\_\{\\infty\}\(x\)\-f\_\{\\infty\}\(g\_\{0\}^\{\-1\}\\cdot y^\{\\prime\}\)\|<\\varepsilon, and conclude thatf¯∞\\bar\{f\}\_\{\\infty\}is continuous\. ∎

## Appendix BMissing details from Section[3](https://arxiv.org/html/2605.23156#S3)

In this section, we establish the claim that the set of functions inLp​\[0,1\]L^\{p\}\[0,1\]with uniformly bounded range is not compact\. This is a well known fact, but we prove it for completeness\.

###### Lemma B\.1\.

The setK~=\{X:\[0,1\]→\[0,1\]\}\\widetilde\{K\}=\\\{X\\colon\[0,1\]\\to\[0,1\]\\\}is not compact inLp​\(\[0,1\]\)L^\{p\}\(\[0,1\]\)for anyp∈\[1,∞\]p\\in\[1,\\infty\]\.

###### Proof of Lemma[B\.1](https://arxiv.org/html/2605.23156#A2.Thmtheorem1)\.\.

It suffices to exhibit an infinite sequence\(Xn\)⊆K~\(X\_\{n\}\)\\subseteq\\widetilde\{K\}that has no convergent subsequence\. To this end, letXn​\(t\)=1X\_\{n\}\(t\)=1ift∈⋃i=12n−1\[2​i−12n,2​i2n\]t\\in\\bigcup\_\{i=1\}^\{2^\{n\-1\}\}\\left\[\\frac\{2i\-1\}\{2^\{n\}\},\\frac\{2i\}\{2^\{n\}\}\\right\]andXn​\(t\)=0X\_\{n\}\(t\)=0otherwise\. Note that‖Xn−Xm‖p=141/p\\\|X\_\{n\}\-X\_\{m\}\\\|\_\{p\}=\\frac\{1\}\{4^\{1/p\}\}for alln≠mn\\neq m, so\(Xn\)\(X\_\{n\}\)cannot have a convergent subsequence inLpL^\{p\}\.

∎

## Appendix CMissing proofs and additional details from Section[4\.1](https://arxiv.org/html/2605.23156#S4.SS1)

In this section, we present the proofs and additional details about the constructions for set problems\.

### C\.1Consistent sequences on sets

We start by elaborating on the consistent sequences from Section[4\.1](https://arxiv.org/html/2605.23156#S4.SS1)\. These were studied in detail in\[[7](https://arxiv.org/html/2605.23156#bib.bib7), Appx\. F\], and we refer the reader there for more details and proofs\.

#### C\.1\.1Zero\-padding consistent sequence𝕍zero\\mathbb\{V\}\_\{\\mathrm\{zero\}\}withℓp\\ell\_\{p\}norm

The zero\-padding consistent sequence𝕍zero=\{\(Vn\),\(φN,n\),\(Gn\)\}\\mathbb\{V\}\_\{\\mathrm\{zero\}\}=\\left\\\{\\left\(V\_\{n\}\\right\),\\left\(\\varphi\_\{N,n\}\\right\),\\left\(\\mathrm\{G\}\_\{n\}\\right\)\\right\\\}is defined as Example[2\.2](https://arxiv.org/html/2605.23156#S2.Thmtheorem2)\. Here we setGn=Sn\\mathrm\{G\}\_\{n\}=\\mathrm\{S\}\_\{n\}to be the permutation group acting onℝn\\mathbb\{R\}^\{n\}by permuting coordinates, i\.e\.,\(g⋅x\)i=xg−1​\(i\)\(g\\cdot x\)\_\{i\}=x\_\{g^\{\-1\}\(i\)\}forg∈Sng\\in\\mathrm\{~S\}\_\{n\}\. Forn≤Nn\\leq N, the embedding of groupsθN,n:Sn→SN\\theta\_\{N,n\}\\colon\\mathrm\{S\}\_\{n\}\\rightarrow\\mathrm\{~S\}\_\{N\}is given by

θN,n​\(g\)=\[g00IN−n\]for​g∈Sn\.\\theta\_\{N,n\}\(g\)=\\left\[\\begin\{array\}\[\]\{cc\}g&0\\\\ 0&I\_\{N\-n\}\\end\{array\}\\right\]\\quad\\text\{for \}g\\in S\_\{n\}\.Forp∈\[1,∞\)p\\in\[1,\\infty\), eachVnV\_\{n\}is equipped with theℓp\\ell\_\{p\}\-norms‖x‖p=\(∑i=1n\|xi\|p\)1/p\\\|x\\\|\_\{p\}=\\left\(\\sum\_\{i=1\}^\{n\}\\left\|x\_\{i\}\\right\|^\{p\}\\right\)^\{1/p\}which are permutation\-invariant\. By Proposition[A\.3](https://arxiv.org/html/2605.23156#A1.Thmtheorem3), this induces a norm onV∞V\_\{\\infty\}, also denoted as∥⋅∥p\\\|\\cdot\\\|\_\{p\}\. Consequently, the limit space is identified with the classical sequence space

V∞¯=ℓp=\{x=\(xi\)i=1∞:‖x‖p=\(∑i=1∞\|xi\|p\)1/p<∞\}\.\\overline\{V\_\{\\infty\}\}=\\ell\_\{p\}=\\left\\\{x=\\left\(x\_\{i\}\\right\)\_\{i=1\}^\{\\infty\}:\\\|x\\\|\_\{p\}=\\left\(\\sum\_\{i=1\}^\{\\infty\}\\left\|x\_\{i\}\\right\|^\{p\}\\right\)^\{1/p\}<\\infty\\right\\\}\.The symmetrized distanced¯p​\(x,y\)=infg∈G∞‖g⋅x−y‖p\\overline\{\\mathrm\{d\}\}\_\{p\}\(x,y\)=\\inf\_\{g\\in G\_\{\\infty\}\}\\\|g\\cdot x\-y\\\|\_\{p\}defines a metric on the space of orbit closuresV∞¯/G∞\\overline\{V\_\{\\infty\}\}/G\_\{\\infty\}\.

Next, consider thekk\-fold direct sum𝕍zero⊕k=\{\(ℝn×k\),\(φN,n⊕k\),\(Sn\)\}\\mathbb\{V\}\_\{\\mathrm\{zero\}\}^\{\\oplus k\}=\\\{\(\\mathbb\{R\}^\{n\\times k\}\),\(\\varphi\_\{N,n\}^\{\\oplus k\}\),\(\\mathrm\{S\}\_\{n\}\)\\\}of𝕍zero\\mathbb\{V\}\_\{\\mathrm\{zero\}\}, as defined in Definition[A\.2](https://arxiv.org/html/2605.23156#A1.Thmtheorem2)\. Fix an arbitrary norm∥⋅∥ℝk\\\|\\cdot\\\|\_\{\\mathbb\{R\}^\{k\}\}onℝk\\mathbb\{R\}^\{k\}, and define theℓp\\ell\_\{p\}\-norm onℝn×d\\mathbb\{R\}^\{n\\times d\}as

‖X‖p=\(∑i=1n‖Xi:‖ℝkp\)1/p,\\\|X\\\|\_\{p\}=\\left\(\\sum\_\{i=1\}^\{n\}\\left\\\|X\_\{i:\}\\right\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{p\}\\right\)^\{1/p\},whereXi:X\_\{i:\}denotes theii\-th row ofXX\. The corresponding limit space can be represented as:

V∞¯=ℓp​\(ℝk\)=\{X=\(Xi:\)i=1∞:‖X‖p=\(∑i=1∞‖Xi:‖ℝkp\)1/p<∞\}\.\\overline\{V\_\{\\infty\}\}=\\ell\_\{p\}\(\\mathbb\{R\}^\{k\}\)=\\left\\\{X=\\left\(X\_\{i:\}\\right\)\_\{i=1\}^\{\\infty\}:\\\|X\\\|\_\{p\}=\\left\(\\sum\_\{i=1\}^\{\\infty\}\\left\\\|X\_\{i:\}\\right\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{p\}\\right\)^\{1/p\}<\\infty\\right\\\}\.

#### C\.1\.2Duplication consistent sequence𝕍dup\\mathbb\{V\}\_\{\\mathrm\{dup\}\}with normalizedℓp\\ell\_\{p\}\-norms

The duplication embedding consistent sequence𝕍dup=\{\(Vn\),\(φN,n\),\(Gn\)\}\\mathbb\{V\}\_\{\\mathrm\{dup\}\}=\\left\\\{\\left\(V\_\{n\}\\right\),\\left\(\\varphi\_\{N,n\}\\right\),\\left\(\\mathrm\{G\}\_\{n\}\\right\)\\right\\\}is defined as Example[2\.2](https://arxiv.org/html/2605.23156#S2.Thmtheorem2)\. The groupGn=Sn\\mathrm\{G\}\_\{n\}=\\mathrm\{S\}\_\{n\}again acts onℝn\\mathbb\{R\}^\{n\}by permuting coordinates\. Forn\|Nn\|N, the embedding of groupsθN,n:Sn→SN\\theta\_\{N,n\}\\colon\\mathrm\{S\}\_\{n\}\\rightarrow\\mathrm\{~S\}\_\{N\}is given byθN,n​\(g\)=g⊗IN/n\\theta\_\{N,n\}\(g\)=g\\otimes I\_\{N/n\}, where⊗\\otimesdenotes the Kronecker product\.

The spaceV∞V\_\{\\infty\}can be identified with the space of step functions, whose discontinuity points are rational\. More precisely, eachx∈ℝnx\\in\\mathbb\{R\}^\{n\}corresponds to a step functionfx:\[0,1\]→ℝf\_\{x\}:\[0,1\]\\rightarrow\\mathbb\{R\}given by

fx​\(t\)=xifor​t∈\(i−1n,in\],i∈\[n\],f\_\{x\}\(t\)=x\_\{i\}\\quad\\text\{ for \}t\\in\\left\(\\frac\{i\-1\}\{n\},\\frac\{i\}\{n\}\\right\],\\,i\\in\[n\],andfx​\(0\)=x1f\_\{x\}\(0\)=x\_\{1\}\. Forp∈\[1,∞\)p\\in\[1,\\infty\), eachVnV\_\{n\}is equipped with the normalizedℓp\\ell\_\{p\}\-norms‖x‖p¯=\(1n​∑i=1n\|xi\|p\)1/p\\\|x\\\|\_\{\\overline\{p\}\}=\\left\(\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\left\|x\_\{i\}\\right\|^\{p\}\\right\)^\{1/p\}, which are compatible\. Under the identification with step functions, this corresponds to theLpL^\{p\}norm of functions, given by‖fx‖p=\(∫01\|fx​\(t\)\|p​𝑑t\)1/p\\\|f\_\{x\}\\\|\_\{p\}=\\left\(\\int\_\{0\}^\{1\}\|f\_\{x\}\(t\)\|^\{p\}dt\\right\)^\{1/p\}\. By Proposition[A\.3](https://arxiv.org/html/2605.23156#A1.Thmtheorem3), this induces a norm onV∞V\_\{\\infty\}\. Consequently, the limit space can be identified with theLpL^\{p\}space, specifically,

V∞¯≅Lp​\(\[0,1\]\)=\{f:\[0,1\]→ℝ​measurable :​∫01\|f​\(t\)\|p​𝑑t<∞\}\.\\overline\{V\_\{\\infty\}\}\\cong L^\{p\}\(\[0,1\]\)=\\left\\\{f:\[0,1\]\\rightarrow\\mathbb\{R\}\\text\{ measurable : \}\\int\_\{0\}^\{1\}\|f\(t\)\|^\{p\}dt<\\infty\\right\\\}\.The permutations inSn\\mathrm\{S\}\_\{n\}act on functions onLp​\(\[0,1\]\)L^\{p\}\(\[0,1\]\)by permuting consecutive intervals of length1/n1/n\. Formally, eachσ∈Sn\\sigma\\in\\mathrm\{S\}\_\{n\}defines a measure\-preserving bijectionσ~:\[0,1\]→\[0,1\]\\widetilde\{\\sigma\}\\colon\[0,1\]\\to\[0,1\]viaσ~​\(\(i−1\)/n\+t\)=\(σ​\(i\)−1\)/n\+t\\widetilde\{\\sigma\}\(\(i\-1\)/n\+t\)=\(\\sigma\(i\)\-1\)/n\+tfort∈\[0,1/n\)t\\in\[0,1/n\)andi∈\[n\]i\\in\[n\]\(andσ~​\(1\)=1\\widetilde\{\\sigma\}\(1\)=1for simplicity\)\.

A functionf∈V∞¯f\\in\\overline\{V\_\{\\infty\}\}gives rise to a probability measureμf\\mu\_\{f\}defined as the distribution off​\(T\)f\(T\)forTTuniformly sampled from\[0,1\]\[0,1\]\. Equivalently,μf=f\#​λ\\mu\_\{f\}=f\_\{\\\#\}\\lambdais the pushforward underffof the Lebesgue measureλ\\lambdaon\[0,1\]\[0,1\]\. All elements in the orbit offfcorrespond to the same measureμf\\mu\_\{f\}, and conversely, two functions are in the orbit\-closures of each other if and only if they correspond to the same measure\. Thus, the orbit spaceV∞¯/G∞\\overline\{V\_\{\\infty\}\}/G\_\{\\infty\}can be identified with𝒫p​\(ℝ\)\\mathcal\{P\}\_\{p\}\(\\mathbb\{R\}\)\. Furthermore, the symmetrized distance coincides with the Wasserstein\-ppdistance in this case\.

Likewise, for thekk\-fold direct sum of𝕍dup\\mathbb\{V\}\_\{\\mathrm\{dup\}\}we fix an arbitrary norm onℝk\\mathbb\{R\}^\{k\}\. The limit spaceV∞¯\\overline\{V\_\{\\infty\}\}can be identified withLp​\(\[0,1\];ℝk\)L^\{p\}\\left\(\[0,1\];\\mathbb\{R\}^\{k\}\\right\), and the orbit space can be identified with𝒫p​\(ℝk\)\\mathcal\{P\}\_\{p\}\(\\mathbb\{R\}^\{k\}\)endowed with the Wassersteinpp\-distance with respect to∥⋅∥ℝk\\\|\\cdot\\\|\_\{\\mathbb\{R\}^\{k\}\}in the same way\. More details and proofs can be found in\[[7](https://arxiv.org/html/2605.23156#bib.bib7), Appx\. F\]\.

### C\.2Proof of Theorem[4\.2](https://arxiv.org/html/2605.23156#S4.Thmtheorem2)

###### Proof of Theorem[4\.2](https://arxiv.org/html/2605.23156#S4.Thmtheorem2)\.

We show that the modelDeepSets∞ρ,σ\\mathrm\{DeepSets\}\_\{\\infty\}^\{\\rho,\\sigma\}can be decomposed asDeepSets∞ρ,σ=σ∘τ\\mathrm\{DeepSets\}\_\{\\infty\}^\{\\rho,\\sigma\}=\\sigma\\circ\\tauwithτ:V∞¯/G∞→ℝr\\tau\\colon\\overline\{V\_\{\\infty\}\}/G\_\{\\infty\}\\to\\mathbb\{R\}^\{r\}continuous\. SinceV∞¯/G∞\\overline\{V\_\{\\infty\}\}/G\_\{\\infty\}is a metric space equipped with the symmetrized distance, continuity onV∞¯/G∞\\overline\{V\_\{\\infty\}\}/G\_\{\\infty\}is equivalent to continuity on each of its compact subset\[[55](https://arxiv.org/html/2605.23156#bib.bib55)\]\. Finally, by definition,σ:ℝr→ℝ\\sigma\\colon\\mathbb\{R\}^\{r\}\\to\\mathbb\{R\}is a continuous function, this decomposition implies continuity of the entire model\.

For anyn∈ℕn\\in\{\\mathbb\{N\}\}, define the mapsτ:V∞→ℝr\\tau\\colon V\_\{\\infty\}\\rightarrow\\mathbb\{R\}^\{r\}andτn:V∞→ℝr\\tau\_\{n\}\\colon V\_\{\\infty\}\\rightarrow\\mathbb\{R\}^\{r\}by

τ​\(X\)=∑i=1∞‖Xi:‖ℝkp​ρ​\(Xi:\)andτn​\(X\)=∑i=1n−1‖Xi:‖ℝkp​ρ​\(Xi:\),\\tau\(X\)=\\sum\_\{i=1\}^\{\\infty\}\\\|X\_\{i:\}\\\|^\{p\}\_\{\\mathbb\{R\}^\{k\}\}\\rho\(X\_\{i:\}\)\\quad\\text\{and\}\\quad\\tau\_\{n\}\(X\)=\\sum\_\{i=1\}^\{n\-1\}\\\|X\_\{i:\}\\\|^\{p\}\_\{\\mathbb\{R\}^\{k\}\}\\rho\(X\_\{i:\}\),respectively\. We equipℝr\\mathbb\{R\}^\{r\}with a standard norm∥⋅∥ℝr\\\|\\cdot\\\|\_\{\\mathbb\{R\}^\{r\}\}\. Fix an arbitrary compact subsetK⊆ℓp​\(ℝk\)K\\subseteq\\ell\_\{p\}\(\\mathbb\{R\}^\{k\}\)\. By Proposition[4\.1](https://arxiv.org/html/2605.23156#S4.Thmtheorem1), there existsM\>0M\>0such thatsupX∈K\(∑j=1∞‖Xj:‖ℝkp\)1/p≤M\.\\sup\_\{X\\in K\}\\left\(\\sum\_\{j=1\}^\{\\infty\}\\\|X\_\{j:\}\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{p\}\\right\)^\{1/p\}\\leq M\.This impliessupX∈Ksupi∈ℕ‖Xi:‖ℝk≤M\.\\sup\_\{X\\in K\}\\sup\_\{i\\in\{\\mathbb\{N\}\}\}\\\|X\_\{i:\}\\\|\_\{\\mathbb\{R\}^\{k\}\}\\leq M\.Sinceρ:ℝk→ℝr\\rho\\colon\\mathbb\{R\}^\{k\}\\rightarrow\\mathbb\{R\}^\{r\}is continuous, it maps compact sets to compact sets; thus, there existsMρ\>0M\_\{\\rho\}\>0such thatsupX∈Ksupi∈ℕ‖ρ​\(Xi:\)‖ℝr≤Mρ\.\\sup\_\{X\\in K\}\\sup\_\{i\\in\{\\mathbb\{N\}\}\}\\\|\\rho\(X\_\{i:\}\)\\\|\_\{\\mathbb\{R\}^\{r\}\}\\leq M\_\{\\rho\}\.Eachτn\\tau\_\{n\}is continuous onKKas it is a finite sum of continuous functions \(coordinate projections andρ\\rho\)\. According to the uniform tail decay property from Proposition[4\.1](https://arxiv.org/html/2605.23156#S4.Thmtheorem1), we have

limn→∞supX∈K‖τn​\(X\)−τ​\(X\)‖ℝr≤Mρ​limn→∞supX∈K∑i≥n‖Xi:‖ℝkp=0\.\\displaystyle\\quad\\lim\_\{n\\rightarrow\\infty\}\\sup\_\{X\\in K\}\\\|\\tau\_\{n\}\(X\)\-\\tau\(X\)\\\|\_\{\\mathbb\{R\}^\{r\}\}\\leq M\_\{\\rho\}\\lim\_\{n\\rightarrow\\infty\}\\sup\_\{X\\in K\}\\sum\_\{i\\geq n\}\\\|X\_\{i:\}\\\|^\{p\}\_\{\\mathbb\{R\}^\{k\}\}=0\.Hence,τn\{\\tau\}\_\{n\}uniformly converges toτ\\tauonKK\. By invoking the uniform limit theorem\[[56](https://arxiv.org/html/2605.23156#bib.bib56)\], we conclude thatτ\\tauis continuous onKK\. SinceKKwas an arbitrary compact subset, we conclude thatτ\\tauis continuous on all ofV∞¯\\overline\{V\_\{\\infty\}\}\. Furthermore, sinceτ\\tauisG∞G\_\{\\infty\}\-invariant, it induces a function on the orbit spaceV∞¯/G∞\\overline\{V\_\{\\infty\}\}/G\_\{\\infty\}, which we still denote byτ\\tau\. By Lemma[A\.4](https://arxiv.org/html/2605.23156#A1.Thmtheorem4), this induced mapτ:V∞¯/G∞→ℝr\\tau\\colon\\overline\{V\_\{\\infty\}\}/G\_\{\\infty\}\\to\\mathbb\{R\}^\{r\}is continuous, which implies thatDeepSets∞ρ,σ=σ∘τ\\mathrm\{DeepSets\}\_\{\\infty\}^\{\\rho,\\sigma\}=\\sigma\\circ\\tauis continuous\. ∎

### C\.3Proof of Theorem[4\.3](https://arxiv.org/html/2605.23156#S4.Thmtheorem3)

To facilitate the proof, we introduce the set of invariant functionsℱcont\\mathcal\{F\}\_\{\\mathrm\{cont\}\}given by

ℱcont≔\{f:K/G∞→ℝ,f​\(X\)=σ​\(∑i=1∞‖Xi:‖ℝkp​ρ​\(Xi:\)\)\|r∈ℕ,ρ∈C​\(ℝk,ℝr\),σ∈C​\(ℝr,ℝ\)\}\.\\mathcal\{F\}\_\{\\mathrm\{cont\}\}\\coloneqq\\left\\\{f:K/G\_\{\\infty\}\\to\\mathbb\{R\},\\;f\(X\)=\\sigma\\left\(\\sum\_\{i=1\}^\{\\infty\}\\\|X\_\{i:\}\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{p\}\\rho\(X\_\{i:\}\)\\right\)\\,\\Big\|\\;r\\in\{\\mathbb\{N\}\},\\;\\rho\\in C\\left\(\\mathbb\{R\}^\{k\},\\mathbb\{R\}^\{r\}\\right\),\\;\\sigma\\in C\\left\(\\mathbb\{R\}^\{r\},\\mathbb\{R\}\\right\)\\right\\\}\.This class serves as an intermediate bridge between the target spaceC​\(K/G∞\)C\(K/G\_\{\\infty\}\)and the neural network classℱDS\\mathcal\{F\}\_\{\\mathrm\{DS\}\}\.

The proof of Theorem[4\.3](https://arxiv.org/html/2605.23156#S4.Thmtheorem3)is split into two main components\. We first invoke the Stone\-Weierstrass Theorem to prove thatℱcont\\mathcal\{F\}\_\{\\mathrm\{cont\}\}is dense inC​\(K/G∞\)C\(K/G\_\{\\infty\}\)\. Next, we show that any element inℱcont\\mathcal\{F\}\_\{\\mathrm\{cont\}\}can be uniformly approximated byℱDS\\mathcal\{F\}\_\{\\mathrm\{DS\}\}through the universal approximation property of neural networks\. Finally, the result follows by the transitivity of the density property\. The details are organized into the three lemmas below, Lemma[C\.1](https://arxiv.org/html/2605.23156#A3.Thmtheorem1),[C\.2](https://arxiv.org/html/2605.23156#A3.Thmtheorem2),[C\.4](https://arxiv.org/html/2605.23156#A3.Thmtheorem4)\.

###### Lemma C\.1\.

The function classℱcont\\mathcal\{F\}\_\{\\mathrm\{cont\}\}is a subalgebra ofC​\(K/G∞\)C\\left\(K/G\_\{\\infty\}\\right\)\.

###### Proof of Lemma[C\.1](https://arxiv.org/html/2605.23156#A3.Thmtheorem1)\.

To establish thatℱcont\\mathcal\{F\}\_\{\\mathrm\{cont\}\}is a subalgebra ofC​\(K/G∞\)C\(K/G\_\{\\infty\}\), we verify its closure under scalar multiplication, addition, and pointwise multiplication\. Letf1,f2∈ℱcontf\_\{1\},f\_\{2\}\\in\\mathcal\{F\}\_\{\\mathrm\{cont\}\}withfj​\(X\)=σj​\(∑i=1∞‖Xi:‖ℝkp​ρj​\(Xi:\)\)f\_\{j\}\(X\)=\\sigma\_\{j\}\\left\(\\sum\_\{i=1\}^\{\\infty\}\\\|X\_\{i:\}\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{p\}\\rho\_\{j\}\(X\_\{i:\}\)\\right\)forj∈\{1,2\}j\\in\\\{1,2\\\}, whereρj∈C​\(ℝk,ℝrj\)\\rho\_\{j\}\\in C\(\\mathbb\{R\}^\{k\},\\mathbb\{R\}^\{r\_\{j\}\}\)andσj∈C​\(ℝrj,ℝ\)\\sigma\_\{j\}\\in C\(\\mathbb\{R\}^\{r\_\{j\}\},\\mathbb\{R\}\)\.

##### Scalar Multiplication\.

For anyλ∈ℝ\\lambda\\in\\mathbb\{R\},\(λ​f1\)​\(X\)=σλ​\(∑i=1∞‖Xi:‖ℝkp​ρ1​\(Xi:\)\)\(\\lambda f\_\{1\}\)\(X\)=\\sigma\_\{\\lambda\}\\left\(\\sum\_\{i=1\}^\{\\infty\}\\\|X\_\{i:\}\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{p\}\\rho\_\{1\}\(X\_\{i:\}\)\\right\)whereσλ≔λ​σ1\\sigma\_\{\\lambda\}\\coloneqq\\lambda\\sigma\_\{1\}\. Sinceσλ\\sigma\_\{\\lambda\}is also continuous,λ​f1∈ℱcont\\lambda f\_\{1\}\\in\\mathcal\{F\}\_\{\\mathrm\{cont\}\}\.

##### Addition\.

Define the concatenated mapρ0:ℝk→ℝr1\+r2\\rho\_\{0\}\\colon\\mathbb\{R\}^\{k\}\\to\\mathbb\{R\}^\{r\_\{1\}\+r\_\{2\}\}asρ0​\(x\)=\(ρ1​\(x\),ρ2​\(x\)\)\\rho\_\{0\}\(x\)=\(\\rho\_\{1\}\(x\),\\rho\_\{2\}\(x\)\)\. Since its componentsρ1\\rho\_\{1\}andρ2\\rho\_\{2\}are continuous,ρ0:ℝk→ℝr1\+r2\\rho\_\{0\}\\colon\\mathbb\{R\}^\{k\}\\to\\mathbb\{R\}^\{r\_\{1\}\+r\_\{2\}\}is continuous by the property of product topologies\. Defineσadd:ℝr1\+r2→ℝ\\sigma\_\{\\rm add\}\\colon\\mathbb\{R\}^\{r\_\{1\}\+r\_\{2\}\}\\to\\mathbb\{R\}byσadd​\(u1,u2\)=σ1​\(u1\)\+σ2​\(u2\)\\sigma\_\{\\rm add\}\(u\_\{1\},u\_\{2\}\)=\\sigma\_\{1\}\(u\_\{1\}\)\+\\sigma\_\{2\}\(u\_\{2\}\)foru1∈ℝr1,u2∈ℝr2u\_\{1\}\\in\\mathbb\{R\}^\{r\_\{1\}\},u\_\{2\}\\in\\mathbb\{R\}^\{r\_\{2\}\}\.σadd\\sigma\_\{\\rm add\}is also continuous since the addition operator onℝ\\mathbb\{R\}is continuous\. Consequently,

\(f1\+f2\)​\(X\)=σ1​\(∑i=1∞‖Xi:‖ℝkp​ρ1​\(Xi:\)\)\+σ2​\(∑i=1∞‖Xi:‖ℝkp​ρ2​\(Xi:\)\)=σadd​\(∑i=1∞‖Xi:‖ℝkp​ρ0​\(Xi:\)\)\.\(f\_\{1\}\+f\_\{2\}\)\(X\)=\\sigma\_\{1\}\\left\(\\sum\_\{i=1\}^\{\\infty\}\\\|X\_\{i:\}\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{p\}\\rho\_\{1\}\(X\_\{i:\}\)\\right\)\+\\sigma\_\{2\}\\left\(\\sum\_\{i=1\}^\{\\infty\}\\\|X\_\{i:\}\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{p\}\\rho\_\{2\}\(X\_\{i:\}\)\\right\)=\\sigma\_\{\\rm add\}\\left\(\\sum\_\{i=1\}^\{\\infty\}\\\|X\_\{i:\}\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{p\}\\rho\_\{0\}\(X\_\{i:\}\)\\right\)\.Thus,ℱcont\\mathcal\{F\}\_\{\\mathrm\{cont\}\}is closed under addition\.

##### Pointwise Multiplication\.

Similarly, define the concatenated mapρ0:ℝk→ℝr1\+r2\\rho\_\{0\}\\colon\\mathbb\{R\}^\{k\}\\to\\mathbb\{R\}^\{r\_\{1\}\+r\_\{2\}\}asρ0​\(V\)=\(ρ1​\(V\),ρ2​\(V\)\)\\rho\_\{0\}\(V\)=\(\\rho\_\{1\}\(V\),\\rho\_\{2\}\(V\)\), which is continuous\. Defineσmul:ℝr1\+r2→ℝ\\sigma\_\{\\rm mul\}\\colon\\mathbb\{R\}^\{r\_\{1\}\+r\_\{2\}\}\\to\\mathbb\{R\}byσmul​\(u1,u2\)=σ1​\(u1\)⋅σ2​\(u2\)\\sigma\_\{\\rm mul\}\(u\_\{1\},u\_\{2\}\)=\\sigma\_\{1\}\(u\_\{1\}\)\\cdot\\sigma\_\{2\}\(u\_\{2\}\)foru1∈ℝr1,u2∈ℝr2u\_\{1\}\\in\\mathbb\{R\}^\{r\_\{1\}\},u\_\{2\}\\in\\mathbb\{R\}^\{r\_\{2\}\}, which is also continuous, because of the continuity of multiplication operator onℝ\\mathbb\{R\}\. We then have

\(f1⋅f2\)​\(X\)=σ1​\(∑i=1∞‖Xi:‖ℝkp​ρ1​\(Xi:\)\)​σ2​\(∑i=1∞‖Xi:‖ℝkp​ρ2​\(Xi:\)\)=σmul​\(∑i=1∞‖Xi:‖ℝkp​ρ0​\(Xi:\)\)\.\(f\_\{1\}\\cdot f\_\{2\}\)\(X\)=\\sigma\_\{1\}\\left\(\\sum\_\{i=1\}^\{\\infty\}\\\|X\_\{i:\}\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{p\}\\rho\_\{1\}\(X\_\{i:\}\)\\right\)\\sigma\_\{2\}\\left\(\\sum\_\{i=1\}^\{\\infty\}\\\|X\_\{i:\}\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{p\}\\rho\_\{2\}\(X\_\{i:\}\)\\right\)=\\sigma\_\{\\rm mul\}\\left\(\\sum\_\{i=1\}^\{\\infty\}\\\|X\_\{i:\}\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{p\}\\rho\_\{0\}\(X\_\{i:\}\)\\right\)\.Thus,ℱcont\\mathcal\{F\}\_\{\\mathrm\{cont\}\}is closed under pointwise multiplication\. ∎

###### Lemma C\.2\.

ℱcont\\mathcal\{F\}\_\{\\mathrm\{cont\}\}separates points inK/G∞K/G\_\{\\infty\}, and contains a nonzero constant function\.

To establish point separation on the quotient spaceK/G∞K/G\_\{\\infty\}, we first characterize the conditions under which the metricd¯​\(X,Y\)\\overline\{\\mathrm\{d\}\}\(X,Y\)vanishes\. Specifically, we identify the properties that distinguish points in distinct orbits\. We claim that points in different orbits must differ in their first finite coordinate indices, see Claim[C\.3](https://arxiv.org/html/2605.23156#A3.Thmtheorem3)\. Based on this, we can construct functions inℱcont\\mathcal\{F\}\_\{\\mathrm\{cont\}\}to separate them\.

###### Claim C\.3\.

For any compactK⊆ℓp​\(ℝk\)K\\subseteq\\ell\_\{p\}\(\\mathbb\{R\}^\{k\}\)andX,Y∈KX,Y\\in K, if the multisets\{Xi::‖Xi:‖ℝk\>ε\}\\\{X\_\{i:\}:\\\|X\_\{i:\}\\\|\_\{\\mathbb\{R\}^\{k\}\}\>\\varepsilon\\\}and\{Yi::‖Yi:‖ℝk\>ε\}\\\{Y\_\{i:\}:\\\|Y\_\{i:\}\\\|\_\{\\mathbb\{R\}^\{k\}\}\>\\varepsilon\\\}are equal for allε\>0\\varepsilon\>0, thend¯​\(X,Y\)=0\\overline\{\\mathrm\{d\}\}\(X,Y\)=0\.

###### Proof of Claim[C\.3](https://arxiv.org/html/2605.23156#A3.Thmtheorem3)\.

Define the index sets byIx​\(ε\)≔\{i∈ℕ:‖Xi:‖ℝk\>ε\}I\_\{x\}\(\\varepsilon\)\\coloneqq\\left\\\{i\\in\{\\mathbb\{N\}\}:\\left\\\|X\_\{i:\}\\right\\\|\_\{\\mathbb\{R\}^\{k\}\}\>\\varepsilon\\right\\\}, andIy​\(ε\)≔\{i∈ℕ:‖Yi:‖ℝk\>ε\}I\_\{y\}\(\\varepsilon\)\\coloneqq\\left\\\{i\\in\{\\mathbb\{N\}\}:\\left\\\|Y\_\{i:\}\\right\\\|\_\{\\mathbb\{R\}^\{k\}\}\>\\varepsilon\\right\\\}, respectively\. SinceX,Y∈KX,Y\\in K, by Proposition[4\.1](https://arxiv.org/html/2605.23156#S4.Thmtheorem1),supx∈K\(∑i=1∞‖Xi:‖ℝkp\)<∞\\sup\_\{x\\in K\}\\left\(\\sum\_\{i=1\}^\{\\infty\}\\left\\\|X\_\{i:\}\\right\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{p\}\\right\)<\\infty\. Therefore,Ix​\(ε\)I\_\{x\}\(\\varepsilon\)andIy​\(ε\)I\_\{y\}\(\\varepsilon\)are finite and have equal cardinality\. Thus, there existsσε∈G∞\\sigma\_\{\\varepsilon\}\\in G\_\{\\infty\}, such thatXσε−1​\(i\):=Yi:X\_\{\\sigma\_\{\\varepsilon\}^\{\-1\}\(i\):\}=Y\_\{i:\}, for alli∈Iy​\(ε\)i\\in I\_\{y\}\(\\varepsilon\)\. We then have

d¯​\(X,Y\)p≤‖σε⋅X−Y‖pp=∑i∉Iy​\(ε\)‖Xσε−1​\(i\):−Yi:‖ℝkp\.\\overline\{\\mathrm\{d\}\}\(X,Y\)^\{p\}\\leq\\\|\\sigma\_\{\\varepsilon\}\\cdot X\-Y\\\|\_\{p\}^\{p\}=\\sum\_\{i\\notin I\_\{y\}\(\\varepsilon\)\}\\\|X\_\{\\sigma\_\{\\varepsilon\}^\{\-1\}\(i\):\}\-Y\_\{i:\}\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{p\}\.Fori∉Iy​\(ε\)i\\notin I\_\{y\}\(\\varepsilon\), we have‖Yi:‖ℝk≤ε\\\|Y\_\{i:\}\\\|\_\{\\mathbb\{R\}^\{k\}\}\\leq\\varepsilon\. Furthermore, sinceσε\\sigma\_\{\\varepsilon\}is a bijection, the index set\{σε−1​\(i\):i∉Iy​\(ε\)\}\\\{\\sigma\_\{\\varepsilon\}^\{\-1\}\(i\):i\\notin I\_\{y\}\(\\varepsilon\)\\\}is exactly\{j∈ℕ:j∉Ix​\(ε\)\}\\\{j\\in\{\\mathbb\{N\}\}:j\\notin I\_\{x\}\(\\varepsilon\)\\\}, where‖Xj:‖ℝk≤ε\\\|X\_\{j:\}\\\|\_\{\\mathbb\{R\}^\{k\}\}\\leq\\varepsilon\. Sincep≥1p\\geq 1, apply the inequality‖a−b‖p≤2p−1​\(‖a‖p\+‖b‖p\)\\\|a\-b\\\|^\{p\}\\leq 2^\{p\-1\}\(\\\|a\\\|^\{p\}\+\\\|b\\\|^\{p\}\)and obtain

d¯​\(X,Y\)p≤2p−1​\(∑i∉Ix​\(ε\)‖Xi:‖ℝkp\+∑i∉Iy​\(ε\)‖Yi:‖ℝkp\)\.\\overline\{\\mathrm\{d\}\}\(X,Y\)^\{p\}\\leq 2^\{p\-1\}\\left\(\\sum\_\{i\\notin I\_\{x\}\(\\varepsilon\)\}\\\|X\_\{i:\}\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{p\}\+\\sum\_\{i\\notin I\_\{y\}\(\\varepsilon\)\}\\\|Y\_\{i:\}\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{p\}\\right\)\.Lettingε→0\\varepsilon\\to 0shows thatd¯​\(X,Y\)=0\\overline\{\\mathrm\{d\}\}\(X,Y\)=0, meaningXXandYYrepresent the same point inK/G∞K/G\_\{\\infty\}\. ∎

We are ready to prove separation of points\.

###### Proof of Lemma[C\.2](https://arxiv.org/html/2605.23156#A3.Thmtheorem2)\.

First,ℱcont\\mathcal\{F\}\_\{\\mathrm\{cont\}\}contains a nonzero constant function by takingσ≡1\\sigma\\equiv 1, which yieldsf​\(X\)=1f\(X\)=1for allXX\.

To establish point separation, consider two distinct points inK/G∞K/G\_\{\\infty\}, represented byX,Y∈KX,Y\\in Ksuch thatd¯​\(X,Y\)\>0\\overline\{\\mathrm\{d\}\}\(X,Y\)\>0\. By the contrapositive of Claim[C\.3](https://arxiv.org/html/2605.23156#A3.Thmtheorem3), there existsε\>0\\varepsilon\>0such that the finite multisetsMX≔\{Xi::‖Xi:‖ℝk\>ε\}M\_\{X\}\\coloneqq\\\{X\_\{i:\}:\\\|X\_\{i:\}\\\|\_\{\\mathbb\{R\}^\{k\}\}\>\\varepsilon\\\}andMY≔\{Yi::‖Yi:‖ℝk\>ε\}M\_\{Y\}\\coloneqq\\\{Y\_\{i:\}:\\\|Y\_\{i:\}\\\|\_\{\\mathbb\{R\}^\{k\}\}\>\\varepsilon\\\}are distinct\. Thus, there exists somev∈ℝkv\\in\\mathbb\{R\}^\{k\}with‖v‖ℝk\>ε\\\|v\\\|\_\{\\mathbb\{R\}^\{k\}\}\>\\varepsilonsuch that their multiplicities differ:mX​\(v\)≠mY​\(v\)m\_\{X\}\(v\)\\neq m\_\{Y\}\(v\), wheremX​\(v\)≔\#​\{i∈ℕ:Xi:=v,Xi:∈MX\}m\_\{X\}\(v\)\\coloneqq\\\#\\\{i\\in\{\\mathbb\{N\}\}:X\_\{i:\}=v,X\_\{i:\}\\in M\_\{X\}\\\}andmY​\(v\)≔\#​\{i∈ℕ:Yi:=v,Yi:∈MY\}m\_\{Y\}\(v\)\\coloneqq\\\#\\\{i\\in\{\\mathbb\{N\}\}:Y\_\{i:\}=v,Y\_\{i:\}\\in M\_\{Y\}\\\}\. LetQ≔\(MX∪MY\)∖\{v\}Q\\coloneqq\(M\_\{X\}\\cup M\_\{Y\}\)\\setminus\\\{v\\\}, which is a finite set as bothMXM\_\{X\}andMYM\_\{Y\}are finite\. We define the separation radius

η≔min⁡\(minu∈Q⁡‖u−v‖ℝk,‖v‖ℝk−ε\)\.\\eta\\coloneqq\\min\\left\(\\min\_\{u\\in Q\}\\\|u\-v\\\|\_\{\\mathbb\{R\}^\{k\}\},\\\|v\\\|\_\{\\mathbb\{R\}^\{k\}\}\-\\varepsilon\\right\)\.Consider the ballB​\(v,η\)B\(v,\\eta\)aroundvv\. For anyi∈ℕi\\in\{\\mathbb\{N\}\}, we distinguish the following two cases\.

##### Case 1\.

If‖Xi:‖ℝk≤ε\\\|X\_\{i:\}\\\|\_\{\\mathbb\{R\}^\{k\}\}\\leq\\varepsilon, then by the reverse triangle inequality,‖Xi:−v‖ℝk≥‖v‖ℝk−‖Xi:‖ℝk≥‖v‖ℝk−ε≥η\\\|X\_\{i:\}\-v\\\|\_\{\\mathbb\{R\}^\{k\}\}\\geq\\\|v\\\|\_\{\\mathbb\{R\}^\{k\}\}\-\\\|X\_\{i:\}\\\|\_\{\\mathbb\{R\}^\{k\}\}\\geq\\\|v\\\|\_\{\\mathbb\{R\}^\{k\}\}\-\\varepsilon\\geq\\eta\.

##### Case 2\.

IfXi:∈MXX\_\{i:\}\\in M\_\{X\}andXi:≠vX\_\{i:\}\\neq v, then by the definition ofQQ,‖Xi:−v‖ℝk≥minu∈Q⁡‖u−v‖ℝk≥η\\\|X\_\{i:\}\-v\\\|\_\{\\mathbb\{R\}^\{k\}\}\\geq\\min\_\{u\\in Q\}\\\|u\-v\\\|\_\{\\mathbb\{R\}^\{k\}\}\\geq\\eta\.

Combined, these cases imply thatXi:∈B​\(v,η\)X\_\{i:\}\\in B\(v,\\eta\)if and only ifXi:=vX\_\{i:\}=v\. The same logic applies to the sequenceYY\. By Urysohn’s Lemma\[[55](https://arxiv.org/html/2605.23156#bib.bib55)\], there exists a continuous functionφ:ℝk→\[0,1\]\\varphi\\colon\\mathbb\{R\}^\{k\}\\rightarrow\[0,1\], such thatφ​\(B​\(v,η2\)¯\)=1\\varphi\\left\(\\overline\{B\(v,\\frac\{\\eta\}\{2\}\)\}\\right\)=1andφ​\(ℝk∖B​\(v,η\)¯\)=0\\varphi\\left\(\\overline\{\\mathbb\{R\}^\{k\}\\setminus B\(v,\\eta\)\}\\right\)=0\. Definef∈ℱcontf\\in\\mathcal\{F\}\_\{\\mathrm\{cont\}\}by settingr=1r=1,σ​\(u\)=u\\sigma\(u\)=u, andρ=φ\\rho=\\varphi, then,

f​\(X\)=σ​\(∑i=1∞‖Xi:‖ℝkp​ρ​\(Xi:\)\)=‖v‖ℝkp​mX​\(v\)≠‖v‖ℝkp​mY​\(v\)=σ​\(∑i=1∞‖Yi:‖ℝkp​ρ​\(Yi:\)\)=f​\(Y\)\.f\(X\)=\\sigma\\left\(\\sum\_\{i=1\}^\{\\infty\}\\\|X\_\{i:\}\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{p\}\\rho\\left\(X\_\{i:\}\\right\)\\right\)=\\\|v\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{p\}m\_\{X\}\(v\)\\neq\\\|v\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{p\}m\_\{Y\}\(v\)=\\sigma\\left\(\\sum\_\{i=1\}^\{\\infty\}\\\|Y\_\{i:\}\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{p\}\\rho\\left\(Y\_\{i:\}\\right\)\\right\)=f\(Y\)\.Therefore, we conclude thatℱcont\\mathcal\{F\}\_\{\\mathrm\{cont\}\}separates points inK/G∞K/G\_\{\\infty\}\. ∎

###### Lemma C\.4\.

ℱDS\\mathcal\{F\}\_\{\\mathrm\{DS\}\}is dense inℱcont\\mathcal\{F\}\_\{\\mathrm\{cont\}\}onKK\.

###### Proof of Lemma[C\.4](https://arxiv.org/html/2605.23156#A3.Thmtheorem4)\.

We invoke the Universal Approximation Theorem \(UAT\)\[[32](https://arxiv.org/html/2605.23156#bib.bib32)\]: for any continuous non\-polynomial activation functionϕ\\phi, the setNNr,kϕ\\texttt\{NN\}^\{\\phi\}\_\{r,k\}is dense inC​\(ℝk,ℝr\)C\(\\mathbb\{R\}^\{k\},\\mathbb\{R\}^\{r\}\)under the supremum norm on compact sets\. More formally, fix a compact setQ⊆ℝkQ\\subseteq\\mathbb\{R\}^\{k\}and a norm∥⋅∥ℝr\\\|\\cdot\\\|\_\{\\mathbb\{R\}^\{r\}\}onℝr\\mathbb\{R\}^\{r\}\. For any continuous functionh∈C​\(ℝk,ℝr\)h\\in C\(\\mathbb\{R\}^\{k\},\\mathbb\{R\}^\{r\}\), and any approximation precisionε\>0\\varepsilon\>0, there existsh^∈NNr,kϕ\\hat\{h\}\\in\\texttt\{NN\}^\{\\phi\}\_\{r,k\}such thatsupv∈Q‖h​\(v\)−h^​\(v\)‖ℝr≤ε\\sup\_\{v\\in Q\}\\\|h\(v\)\-\\hat\{h\}\(v\)\\\|\_\{\\mathbb\{R\}^\{r\}\}\\leq\\varepsilon\.

Given any functionf∈ℱcontf\\in\\mathcal\{F\}\_\{\\mathrm\{cont\}\}, it admits the formf​\(X\)=σ​\(∑i=1∞‖Xi:‖ℝkp​ρ​\(Xi:\)\),f\(X\)=\\sigma\\left\(\\sum\_\{i=1\}^\{\\infty\}\\\|X\_\{i:\}\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{p\}\\rho\(X\_\{i:\}\)\\right\),for somer∈ℕr\\in\{\\mathbb\{N\}\},σ∈C​\(ℝr,ℝ\),ρ∈C​\(ℝk,ℝr\)\.\\sigma\\in C\(\\mathbb\{R\}^\{r\},\\mathbb\{R\}\),\\,\\rho\\in C\(\\mathbb\{R\}^\{k\},\\mathbb\{R\}^\{r\}\)\.SinceK∈ℓp​\(ℝk\)K\\in\\ell\_\{p\}\(\\mathbb\{R\}^\{k\}\)is compact, by Proposition[4\.1](https://arxiv.org/html/2605.23156#S4.Thmtheorem1), there existsM\>0M\>0such thatsupX∈K\(∑j=1∞‖Xj:‖ℝkp\)1/p≤M\.\\sup\_\{X\\in K\}\\left\(\\sum\_\{j=1\}^\{\\infty\}\\\|X\_\{j:\}\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{p\}\\right\)^\{1/p\}\\leq M\.This impliessupX∈Ksupi∈ℕ‖Xi:‖ℝk≤M\.\\sup\_\{X\\in K\}\\sup\_\{i\\in\{\\mathbb\{N\}\}\}\\\|X\_\{i:\}\\\|\_\{\\mathbb\{R\}^\{k\}\}\\leq M\.Sinceρ∈C​\(ℝk,ℝr\)\\rho\\in C\(\\mathbb\{R\}^\{k\},\\mathbb\{R\}^\{r\}\), it maps compact sets to compact sets, i\.e\., there existsMρ\>0M\_\{\\rho\}\>0such thatsupX∈Ksupi∈ℕ‖ρ​\(Xi:\)‖ℝr≤Mρ\.\\sup\_\{X\\in K\}\\sup\_\{i\\in\{\\mathbb\{N\}\}\}\\\|\\rho\(X\_\{i:\}\)\\\|\_\{\\mathbb\{R\}^\{r\}\}\\leq M\_\{\\rho\}\.Thus for anyX∈KX\\in K,‖∑i=1∞‖​Xi:∥ℝkp​ρ​\(Xi:\)∥ℝr≤Mp​Mρ\.\\left\\\|\\sum\_\{i=1\}^\{\\infty\}\\\|X\_\{i:\}\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{p\}\\rho\(X\_\{i:\}\)\\right\\\|\_\{\\mathbb\{R\}^\{r\}\}\\leq M^\{p\}M\_\{\\rho\}\.This means that the input domains of functionsσ\\sigmaandρ\\rhoare both compact, which we denote byQσQ\_\{\\sigma\}andQρQ\_\{\\rho\}, respectively\.

##### Approximatingσ\\sigma\.

For anyε\>0\\varepsilon\>0, UAT ensures there existsσ^∈NN1,rϕ\\hat\{\\sigma\}\\in\\texttt\{NN\}^\{\\phi\}\_\{1,r\}such thatsupu∈Qσ\|σ​\(u\)−σ^​\(u\)\|≤ε2\.\\sup\_\{u\\in Q\_\{\\sigma\}\}\|\\sigma\(u\)\-\\hat\{\\sigma\}\(u\)\|\\leq\\frac\{\\varepsilon\}\{2\}\.Moreover, sinceσ^\\hat\{\\sigma\}is uniformly continuous on compact sets, there existsδ\>0\\delta\>0such that‖u1−u2‖ℝr≤δ⇒\|σ^​\(u1\)−σ^​\(u2\)\|≤ε2\.\\\|u\_\{1\}\-u\_\{2\}\\\|\_\{\\mathbb\{R\}^\{r\}\}\\leq\\delta\\Rightarrow\|\\hat\{\\sigma\}\(u\_\{1\}\)\-\\hat\{\\sigma\}\(u\_\{2\}\)\|\\leq\\frac\{\\varepsilon\}\{2\}\.

##### Approximatingρ\\rho\.

Applying UAT toρ\\rho, there existsρ^∈NNr,kϕ\\hat\{\\rho\}\\in\\texttt\{NN\}^\{\\phi\}\_\{r,k\}such thatsupv∈Qρ‖ρ​\(v\)−ρ^​\(v\)‖ℝr≤δMp\.\\sup\_\{v\\in Q\_\{\\rho\}\}\\\|\\rho\(v\)\-\\hat\{\\rho\}\(v\)\\\|\_\{\\mathbb\{R\}^\{r\}\}\\leq\\frac\{\\delta\}\{M^\{p\}\}\.Consequently,

supX∈K‖∑i=1∞‖​Xi:∥ℝkp​\(ρ​\(Xi:\)−ρ^​\(Xi:\)\)∥ℝr≤supX∈K∑i=1∞‖Xi:‖ℝkp​supX∈K‖ρ​\(Xi:\)−ρ^​\(Xi:\)‖ℝr≤Mp​δMp=δ\.\\sup\_\{X\\in K\}\\left\\\|\\sum\_\{i=1\}^\{\\infty\}\\\|X\_\{i:\}\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{p\}\(\\rho\(X\_\{i:\}\)\-\\hat\{\\rho\}\(X\_\{i:\}\)\)\\right\\\|\_\{\\mathbb\{R\}^\{r\}\}\\leq\\sup\_\{X\\in K\}\\sum\_\{i=1\}^\{\\infty\}\\\|X\_\{i:\}\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{p\}\\sup\_\{X\\in K\}\\left\\\|\\rho\(X\_\{i:\}\)\-\\hat\{\\rho\}\(X\_\{i:\}\)\\right\\\|\_\{\\mathbb\{R\}^\{r\}\}\\leq M^\{p\}\\frac\{\\delta\}\{M^\{p\}\}=\\delta\.We define the overall approximating functionf^∈ℱDS\\hat\{f\}\\in\\mathcal\{F\}\_\{\\mathrm\{DS\}\}byf^​\(X\)=σ^​\(∑i=1∞‖Xi:‖ℝkp​ρ^​\(Xi:\)\)\\hat\{f\}\(X\)=\\hat\{\\sigma\}\\left\(\\sum\_\{i=1\}^\{\\infty\}\\\|X\_\{i:\}\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{p\}\\hat\{\\rho\}\(X\_\{i:\}\)\\right\), then

supX∈K\|f​\(X\)−f^​\(X\)\|\\displaystyle\\sup\_\{X\\in K\}\\left\|f\(X\)\-\\hat\{f\}\(X\)\\right\|≤supX∈K\|σ​\(∑i=1∞‖Xi:‖ℝkp​ρ​\(Xi:\)\)−σ^​\(∑i=1∞‖Xi:‖ℝkp​ρ​\(Xi:\)\)\|\\displaystyle\\quad\\leq\\sup\_\{X\\in K\}\\left\|\\sigma\\left\(\\sum\_\{i=1\}^\{\\infty\}\\\|X\_\{i:\}\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{p\}\\rho\(X\_\{i:\}\)\\right\)\-\\hat\{\\sigma\}\\left\(\\sum\_\{i=1\}^\{\\infty\}\\\|X\_\{i:\}\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{p\}\{\\rho\}\(X\_\{i:\}\)\\right\)\\right\|\+\|σ^​\(∑i=1∞‖Xi:‖ℝkp​ρ​\(Xi:\)\)−σ^​\(∑i=1∞‖Xi:‖ℝkp​ρ^​\(Xi:\)\)\|\\displaystyle\\qquad\\qquad\\,\\,\+\\left\|\\hat\{\\sigma\}\\left\(\\sum\_\{i=1\}^\{\\infty\}\\\|X\_\{i:\}\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{p\}\\rho\(X\_\{i:\}\)\\right\)\-\\hat\{\\sigma\}\\left\(\\sum\_\{i=1\}^\{\\infty\}\\\|X\_\{i:\}\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{p\}\{\\hat\{\\rho\}\}\(X\_\{i:\}\)\\right\)\\right\|≤ε2\+ε2=ε,\\displaystyle\\quad\\leq\\frac\{\\varepsilon\}\{2\}\+\\frac\{\\varepsilon\}\{2\}=\\varepsilon,which shows thatℱDS\\mathcal\{F\}\_\{\\text\{DS\}\}uniformly approximates every function inℱcont\\mathcal\{F\}\_\{\\text\{cont\}\}\. By the definition of both of the function classes,ℱDS⊆ℱcont\\mathcal\{F\}\_\{\\mathrm\{DS\}\}\\subseteq\\mathcal\{F\}\_\{\\mathrm\{cont\}\}, therefore,ℱDS\\mathcal\{F\}\_\{\\mathrm\{DS\}\}is dense inℱcont\\mathcal\{F\}\_\{\\text\{cont\}\}onKKwith respect to the supremum norm\. ∎

Finally, combining the above lemmas, we prove Theorem[4\.3](https://arxiv.org/html/2605.23156#S4.Thmtheorem3)\.

###### Proof of Theorem[4\.3](https://arxiv.org/html/2605.23156#S4.Thmtheorem3)\.

Invoking the continuity property of our model, Theorem[4\.2](https://arxiv.org/html/2605.23156#S4.Thmtheorem2),ℱcont⊆C​\(K/G∞\)\\mathcal\{F\}\_\{\\mathrm\{cont\}\}\\subseteq C\(K/G\_\{\\infty\}\)\. By Proposition[3\.1](https://arxiv.org/html/2605.23156#S3.Thmtheorem1), the Stone–Weierstrass Theorem, and Lemmas[C\.1](https://arxiv.org/html/2605.23156#A3.Thmtheorem1)and[C\.2](https://arxiv.org/html/2605.23156#A3.Thmtheorem2), we conclude thatℱcont\\mathcal\{F\}\_\{\\mathrm\{cont\}\}is dense inC​\(K/G∞\)C\\left\(K/G\_\{\\infty\}\\right\)with respect to the supremum norm\. Combining this with Lemma[C\.4](https://arxiv.org/html/2605.23156#A3.Thmtheorem4)and the transitivity of density, it follows thatℱDS\\mathcal\{F\}\_\{\\mathrm\{DS\}\}is dense inC​\(K/G∞\)C\\left\(K/G\_\{\\infty\}\\right\), completing the proof that the model is universal\. ∎

### C\.4Proof of Theorem[4\.5](https://arxiv.org/html/2605.23156#S4.Thmtheorem5)

###### Proof of Theorem[4\.5](https://arxiv.org/html/2605.23156#S4.Thmtheorem5)\.

This theorem follows directly from the fact thatWpW\_\{p\}\-convergence is equivalent to weak convergence in𝒫p\\mathcal\{P\}\_\{p\}\. Consider any sequence of measures\(μi\)i∈ℕ⊆𝒫p​\(ℝk\)\(\\mu\_\{i\}\)\_\{i\\in\{\\mathbb\{N\}\}\}\\subseteq\\mathcal\{P\}\_\{p\}\(\\mathbb\{R\}^\{k\}\)converging toμ\\muin the Wassersteinpp\-metric, i\.e\.,Wp​\(μi,μ\)→0W\_\{p\}\(\\mu\_\{i\},\\mu\)\\to 0\. According to the characterization of convergence in𝒫p​\(ℝk\)\\mathcal\{P\}\_\{p\}\(\\mathbb\{R\}^\{k\}\)\[[57](https://arxiv.org/html/2605.23156#bib.bib57)\], this is equivalent to∫φ​𝑑μi→∫φ​𝑑μ\\int\\varphi\\,d\\mu\_\{i\}\\to\\int\\varphi\\,d\\mufor all continuous functionsφ:ℝk→ℝ\\varphi\\colon\\mathbb\{R\}^\{k\}\\rightarrow\\mathbb\{R\}satisfying the growth condition\|φ​\(x\)\|≤C​\(1\+‖x‖ℝkp\)\|\\varphi\(x\)\|\\leq C\(1\+\\\|x\\\|^\{p\}\_\{\\mathbb\{R\}^\{k\}\}\)for someC\>0C\>0\.

In our framework,ρ∈ℱp​\(ℝk,ℝr\)\\rho\\in\\mathcal\{F\}\_\{p\}\(\\mathbb\{R\}^\{k\},\\mathbb\{R\}^\{r\}\)is a continuous mapping whose componentsρj\\rho\_\{j\}satisfy\|ρj​\(x\)\|≤Mj​\(1\+‖x‖ℝkp\)\|\\rho\_\{j\}\(x\)\|\\leq M\_\{j\}\(1\+\\\|x\\\|^\{p\}\_\{\\mathbb\{R\}^\{k\}\}\)\. It follows that∫ρ​𝑑μi→∫ρ​𝑑μ\\int\\rho\\,d\\mu\_\{i\}\\to\\int\\rho\\,d\\muasi→∞i\\to\\infty, establishing the continuity of the mapμ↦∫ρ​𝑑μ\\mu\\mapsto\\int\\rho\\,d\\mu\. Sinceσ:ℝr→ℝ\\sigma:\\mathbb\{R\}^\{r\}\\to\\mathbb\{R\}is continuous, the compositionDeepSets¯∞ρ,σ​\(μ\)=σ​\(∫ρ​𝑑μ\)\\overline\{\\mathrm\{DeepSets\}\}\_\{\\infty\}^\{\\rho,\\sigma\}\(\\mu\)=\\sigma\(\\int\\rho\\,d\\mu\)is continuous with respect to theWpW\_\{p\}topology\. ∎

### C\.5Proof of Theorem[4\.6](https://arxiv.org/html/2605.23156#S4.Thmtheorem6)

We first define an intermediate class of functions,ℱcont¯\\mathcal\{F\}\_\{\\overline\{\\mathrm\{cont\}\}\}, to bridgeC​\(Q\)C\(Q\)andℱDS¯\\mathcal\{F\}\_\{\\overline\{\\mathrm\{DS\}\}\}

ℱcont¯≔\{f:Q→ℝ,f​\(μ\)=σ​\(∫ρ​𝑑μ\)\|r∈ℕ,ρ∈ℱp​\(ℝk,ℝr\),σ∈C​\(ℝr,ℝ\)\}\.\\mathcal\{F\}\_\{\\overline\{\\mathrm\{cont\}\}\}\\coloneqq\\Bigl\\\{f\\colon Q\\to\\mathbb\{R\},\\;f\(\\mu\)=\\sigma\\;\\left\(\\int\\rho\\;d\\mu\\right\)\\;\\Big\|\\;r\\in\{\\mathbb\{N\}\},\\;\\rho\\in\\mathcal\{F\}\_\{p\}\\left\(\\mathbb\{R\}^\{k\},\\mathbb\{R\}^\{r\}\\right\),\\;\\sigma\\in C\\left\(\\mathbb\{R\}^\{r\},\\mathbb\{R\}\\right\)\\Bigr\\\}\.The proof strategy of Theorem[4\.6](https://arxiv.org/html/2605.23156#S4.Thmtheorem6)is similar to Theorem[4\.3](https://arxiv.org/html/2605.23156#S4.Thmtheorem3)by showing that:\(i\)\(i\)ℱcont¯\\mathcal\{F\}\_\{\\overline\{\\mathrm\{cont\}\}\}is dense inC​\(Q\)C\(Q\)via the Stone\-Weierstrass Theorem, and\(i​i\)\(ii\)ℱDS¯\\mathcal\{F\}\_\{\\overline\{\\mathrm\{DS\}\}\}is dense inℱcont¯\\mathcal\{F\}\_\{\\overline\{\\mathrm\{cont\}\}\}via the universal approximation property\. We decompose the proof into three auxiliary Lemmas:[C\.5](https://arxiv.org/html/2605.23156#A3.Thmtheorem5),[C\.6](https://arxiv.org/html/2605.23156#A3.Thmtheorem6), and[C\.7](https://arxiv.org/html/2605.23156#A3.Thmtheorem7)\.

###### Lemma C\.5\.

The function classℱcont¯\\mathcal\{F\}\_\{\\overline\{\\mathrm\{cont\}\}\}is a subalgebra ofC​\(Q\)\.C\(Q\)\.

###### Proof of Lemma[C\.5](https://arxiv.org/html/2605.23156#A3.Thmtheorem5)\.

Invoking Theorem[4\.5](https://arxiv.org/html/2605.23156#S4.Thmtheorem5),ℱcont¯⊆C​\(Q\)\\mathcal\{F\}\_\{\\overline\{\\mathrm\{cont\}\}\}\\subseteq C\(Q\)\. To showℱcont¯\\mathcal\{F\}\_\{\\overline\{\\mathrm\{cont\}\}\}is a subalgebra, it suffices to verify that it is closed under scalar multiplication, addition and pointwise multiplication\. Letf1,f2∈ℱcont¯f\_\{1\},f\_\{2\}\\in\\mathcal\{F\}\_\{\\overline\{\\mathrm\{cont\}\}\}withfj​\(μ\)=σj​\(∫ρj​𝑑μ\)f\_\{j\}\(\\mu\)=\\sigma\_\{j\}\\left\(\\int\\rho\_\{j\}d\\mu\\right\)forj∈\{1,2\}j\\in\\\{1,2\\\}, whereρj∈ℱp​\(ℝk,ℝrj\)\\rho\_\{j\}\\in\\mathcal\{F\}\_\{p\}\(\\mathbb\{R\}^\{k\},\\mathbb\{R\}^\{r\_\{j\}\}\)andσj∈C​\(ℝrj,ℝ\)\\sigma\_\{j\}\\in C\(\\mathbb\{R\}^\{r\_\{j\}\},\\mathbb\{R\}\)\.

##### Scalar Multiplication\.

For anyλ∈ℝ\\lambda\\in\\mathbb\{R\},\(λ​f1\)​\(μ\)=σλ​\(∫ρ1​𝑑μ\)\(\\lambda f\_\{1\}\)\(\\mu\)=\\sigma\_\{\\lambda\}\\left\(\\int\\rho\_\{1\}d\\mu\\right\)whereσλ≔λ​σ1\\sigma\_\{\\lambda\}\\coloneqq\\lambda\\sigma\_\{1\}\. Sinceσλ\\sigma\_\{\\lambda\}is also continuous,λ​f1∈ℱcont¯\\lambda f\_\{1\}\\in\\mathcal\{F\}\_\{\\overline\{\\mathrm\{cont\}\}\}\.

##### Addition\.

Define the concatenated mapρ0:ℝk→ℝr1\+r2\\rho\_\{0\}\\colon\\mathbb\{R\}^\{k\}\\to\\mathbb\{R\}^\{r\_\{1\}\+r\_\{2\}\}asρ0​\(x\)=\(ρ1​\(x\),ρ2​\(x\)\)\\rho\_\{0\}\(x\)=\(\\rho\_\{1\}\(x\),\\rho\_\{2\}\(x\)\)\. Since its componentsρ1\\rho\_\{1\}andρ2\\rho\_\{2\}are continuous,ρ0:ℝk→ℝr1\+r2\\rho\_\{0\}\\colon\\mathbb\{R\}^\{k\}\\to\\mathbb\{R\}^\{r\_\{1\}\+r\_\{2\}\}is continuous by the property of product topologies\. Furthermore, sinceρj∈ℱp​\(ℝk,ℝrj\)\\rho\_\{j\}\\in\\mathcal\{F\}\_\{p\}\(\\mathbb\{R\}^\{k\},\\mathbb\{R\}^\{r\_\{j\}\}\)forj∈\{1,2\}j\\in\\\{1,2\\\}, each individual component function ofρ0=\(ρ1,1,…,ρ1,r1,ρ2,1,…,ρ2,r2\)⊤\\rho\_\{0\}=\(\\rho\_\{1,1\},\\dots,\\rho\_\{1,r\_\{1\}\},\\rho\_\{2,1\},\\dots,\\rho\_\{2,r\_\{2\}\}\)^\{\\top\}satisfies thepp\-th order growth condition\. Consequently, we haveρ0∈ℱp​\(ℝk,ℝr1\+r2\)\\rho\_\{0\}\\in\\mathcal\{F\}\_\{p\}\(\\mathbb\{R\}^\{k\},\\mathbb\{R\}^\{r\_\{1\}\+r\_\{2\}\}\)\. Defineσadd:ℝr1\+r2→ℝ\\sigma\_\{\\rm add\}\\colon\\mathbb\{R\}^\{r\_\{1\}\+r\_\{2\}\}\\to\\mathbb\{R\}byσadd​\(u1,u2\)=σ1​\(u1\)\+σ2​\(u2\)\\sigma\_\{\\rm add\}\(u\_\{1\},u\_\{2\}\)=\\sigma\_\{1\}\(u\_\{1\}\)\+\\sigma\_\{2\}\(u\_\{2\}\)foru1∈ℝr1,u2∈ℝr2u\_\{1\}\\in\\mathbb\{R\}^\{r\_\{1\}\},u\_\{2\}\\in\\mathbb\{R\}^\{r\_\{2\}\}\. Observe thatσadd\\sigma\_\{\\rm add\}is also continuous since the addition operator onℝ\\mathbb\{R\}is continuous\. Consequently,

\(f1\+f2\)​\(μ\)=σ1​\(∫ρ1​𝑑μ\)\+σ2​\(∫ρ2​𝑑μ\)=σadd​\(\[∫ρ1​𝑑μ,∫ρ2​𝑑μ\]⊤\)=σadd​\(∫ρ0​𝑑μ\)\.\\left\(f\_\{1\}\+f\_\{2\}\\right\)\(\\mu\)=\\sigma\_\{1\}\\left\(\\int\\rho\_\{1\}d\\mu\\right\)\+\\sigma\_\{2\}\\left\(\\int\\rho\_\{2\}d\\mu\\right\)=\\sigma\_\{\\rm add\}\\left\(\\left\[\\int\\rho\_\{1\}d\\mu,\\;\\int\\rho\_\{2\}d\\mu\\right\]^\{\\top\}\\right\)=\\sigma\_\{\\rm add\}\\left\(\\int\\rho\_\{0\}d\\mu\\right\)\.Thus,ℱcont¯\\mathcal\{F\}\_\{\\overline\{\\mathrm\{cont\}\}\}is closed under addition\.

##### Pointwise Multiplication\.

Similarly, define the concatenated mapρ0:ℝk→ℝr1\+r2\\rho\_\{0\}\\colon\\mathbb\{R\}^\{k\}\\to\\mathbb\{R\}^\{r\_\{1\}\+r\_\{2\}\}asρ0​\(V\)=\(ρ1​\(V\),ρ2​\(V\)\)\\rho\_\{0\}\(V\)=\(\\rho\_\{1\}\(V\),\\rho\_\{2\}\(V\)\), which is continuous\. Defineσmul:ℝr1\+r2→ℝ\\sigma\_\{\\rm mul\}\\colon\\mathbb\{R\}^\{r\_\{1\}\+r\_\{2\}\}\\to\\mathbb\{R\}byσmul​\(u1,u2\)=σ1​\(u1\)×σ2​\(u2\)\\sigma\_\{\\rm mul\}\(u\_\{1\},u\_\{2\}\)=\\sigma\_\{1\}\(u\_\{1\}\)\\times\\sigma\_\{2\}\(u\_\{2\}\)foru1∈ℝr1,u2∈ℝr2u\_\{1\}\\in\\mathbb\{R\}^\{r\_\{1\}\},u\_\{2\}\\in\\mathbb\{R\}^\{r\_\{2\}\}, which is also continuous, because of the continuity of multiplication operator onℝ\\mathbb\{R\}\. We then have

\(f1⋅f2\)​\(μ\)=σ1​\(∫ρ1​𝑑μ\)​σ2​\(∫ρ2​𝑑μ\)=σmul​\(\[∫ρ1​𝑑μ,∫ρ2​𝑑μ\]⊤\)=σmul​\(∫ρ0​𝑑μ\)\.\\left\(f\_\{1\}\\cdot f\_\{2\}\\right\)\(\\mu\)=\\sigma\_\{1\}\\left\(\\int\\rho\_\{1\}d\\mu\\right\)\\sigma\_\{2\}\\left\(\\int\\rho\_\{2\}d\\mu\\right\)=\\sigma\_\{\\rm mul\}\\left\(\\left\[\\int\\rho\_\{1\}d\\mu,\\;\\int\\rho\_\{2\}d\\mu\\right\]^\{\\top\}\\right\)=\\sigma\_\{\\rm mul\}\\left\(\\int\\rho\_\{0\}d\\mu\\right\)\.Thus,ℱcont¯\\mathcal\{F\}\_\{\\overline\{\\mathrm\{cont\}\}\}is closed under pointwise multiplication\. ∎

###### Lemma C\.6\.

The function classℱcont¯\\mathcal\{F\}\_\{\\overline\{\\mathrm\{cont\}\}\}separates points inQQ, and contains a nonzero constant function\.

###### Proof of Lemma[C\.6](https://arxiv.org/html/2605.23156#A3.Thmtheorem6)\.

The fact thatℱcont¯\\mathcal\{F\}\_\{\\overline\{\\mathrm\{cont\}\}\}contains a nonzero constant function is trivial\. To show it separates points inQQ, supposeμ,ν∈Q⊆𝒫p​\(ℝk\)\\mu,\\nu\\in Q\\subseteq\\mathcal\{P\}\_\{p\}\(\\mathbb\{R\}^\{k\}\), andμ≠ν\\mu\\neq\\nu\. There must exist a Borel setB⊆ℝkB\\subseteq\\mathbb\{R\}^\{k\}such thatμ​\(B\)≠ν​\(B\)\\mu\(B\)\\neq\\nu\(B\)\. Without loss of generality, assumeμ​\(B\)\>ν​\(B\)\\mu\(B\)\>\\nu\(B\), and letδ≔μ​\(B\)−ν​\(B\)\>0\\delta\\coloneqq\\mu\(B\)\-\\nu\(B\)\>0\. Sinceℝk\\mathbb\{R\}^\{k\}is a metric space, any measure in𝒫p​\(ℝk\)\\mathcal\{P\}\_\{p\}\(\\mathbb\{R\}^\{k\}\)is regular\[[58](https://arxiv.org/html/2605.23156#bib.bib58)\]\. By the property of regularity, for anyε\>0\\varepsilon\>0, there exists a closed setE⊆BE\\subseteq Band an open setO⊃BO\\supset Bsuch that

μ​\(B\)≤μ​\(E\)\+ε,ν​\(B\)≥ν​\(O\)−ε\.\\mu\(B\)\\leq\\mu\(E\)\+\\varepsilon,\\quad\\nu\(B\)\\geq\\nu\(O\)\-\\varepsilon\.SinceEEandℝk∖O\\mathbb\{R\}^\{k\}\\setminus Oare disjoint closed sets inℝk\\mathbb\{R\}^\{k\}, by Urysohn’s lemma\[[55](https://arxiv.org/html/2605.23156#bib.bib55)\], there exists a continuous functionφ:ℝk→\[0,1\]\\varphi\\colon\\mathbb\{R\}^\{k\}\\rightarrow\[0,1\]such thatφ​\(E\)=1\\varphi\(E\)=1andφ​\(ℝk∖O\)=0\\varphi\(\\mathbb\{R\}^\{k\}\\setminus O\)=0\. It follows that

∫φ​𝑑μ≥μ​\(E\)≥μ​\(B\)−ε,∫φ​𝑑ν≤ν​\(O\)≤ν​\(B\)\+ε\.\\int\\varphi\\;d\\mu\\geq\\mu\(E\)\\geq\\mu\(B\)\-\\varepsilon,\\quad\\int\\varphi\\;d\\nu\\leq\\nu\(O\)\\leq\\nu\(B\)\+\\varepsilon\.Setε=δ3\>0\\varepsilon=\\frac\{\\delta\}\{3\}\>0, we obtain

∫φ​𝑑μ−∫φ​𝑑ν≥μ​\(B\)−ν​\(B\)−2​ε=δ3\>0\.\\int\\varphi\\;d\\mu\-\\int\\varphi\\;d\\nu\\geq\\mu\(B\)\-\\nu\(B\)\-2\\varepsilon=\\frac\{\\delta\}\{3\}\>0\.
Recall thatℱcont¯≔\{f:Q→ℝ,f​\(μ\)=σ​\(∫ρ​𝑑μ\)\|r∈ℕ,ρ∈ℱp​\(ℝk,ℝr\),σ∈C​\(ℝr,ℝ\)\}\\mathcal\{F\}\_\{\\overline\{\\mathrm\{cont\}\}\}\\coloneqq\\left\\\{f\\colon Q\\to\\mathbb\{R\},\\;f\(\\mu\)=\\sigma\\left\(\\int\\rho d\\mu\\right\)\\;\\Big\|\\;r\\in\{\\mathbb\{N\}\},\\rho\\in\\mathcal\{F\}\_\{p\}\\left\(\\mathbb\{R\}^\{k\},\\mathbb\{R\}^\{r\}\\right\),\\sigma\\in C\\left\(\\mathbb\{R\}^\{r\},\\mathbb\{R\}\\right\)\\right\\\}\. By choosingr=1r=1,σ​\(u\)=u\\sigma\(u\)=u, andρ=φ\\rho=\\varphi, we observe thatφ\\varphiis bounded and continuous, which implies\|φ​\(x\)\|≤1≤\(1\+‖x‖ℝkp\)\|\\varphi\(x\)\|\\leq 1\\leq\(1\+\\\|x\\\|^\{p\}\_\{\\mathbb\{R\}^\{k\}\}\), and thusρ∈ℱp​\(ℝk,ℝ\)\\rho\\in\\mathcal\{F\}\_\{p\}\(\\mathbb\{R\}^\{k\},\\mathbb\{R\}\)\. Therefore, it follows that for anyμ,ν∈Q⊆𝒫p​\(ℝk\)\\mu,\\nu\\in Q\\subseteq\\mathcal\{P\}\_\{p\}\(\\mathbb\{R\}^\{k\}\)withμ≠ν\\mu\\neq\\nu, there exists a functionf∈ℱcont¯f\\in\\mathcal\{F\}\_\{\\overline\{\\mathrm\{cont\}\}\}such thatf​\(μ\)≠f​\(ν\)f\(\\mu\)\\neq f\(\\nu\)\. ∎

###### Lemma C\.7\.

The function classℱDS¯\\mathcal\{F\}\_\{\\overline\{\\mathrm\{DS\}\}\}is dense inℱcont¯\\mathcal\{F\}\_\{\\overline\{\\mathrm\{cont\}\}\}onQQ\.

###### Proof of Lemma[C\.7](https://arxiv.org/html/2605.23156#A3.Thmtheorem7)\.

The proof for this lemma proceeds by analogy with Lemma[C\.4](https://arxiv.org/html/2605.23156#A3.Thmtheorem4)except for the following subtle issue: the compactness ofQ⊆𝒫p​\(ℝk\)Q\\subseteq\\mathcal\{P\}\_\{p\}\(\\mathbb\{R\}^\{k\}\), as described in Proposition[4\.4](https://arxiv.org/html/2605.23156#S4.Thmtheorem4), does not ensure that the input domain ofρ\\rhois compact\. Consequently, the standard Universal Approximation Theorem cannot be directly applied\. To address this issue, we leverage the property that compactness in𝒫p​\(ℝk\)\\mathcal\{P\}\_\{p\}\(\\mathbb\{R\}^\{k\}\)guarantees the probability mass in the tails of the measures is negligible\. This allows us to replaceρ\\rhowith a function vanishing at infinity, which can then be uniformly approximated by neural networks on the entire spaceℝk\\mathbb\{R\}^\{k\}using recent results for non\-compact domains\[[34](https://arxiv.org/html/2605.23156#bib.bib34)\]\.

For any functionf∈ℱcont¯f\\in\\mathcal\{F\}\_\{\\overline\{\\mathrm\{cont\}\}\}, we can represent it asf​\(μ\)=σ​\(∫ρ​𝑑μ\),f\(\\mu\)=\\sigma\\left\(\\int\\rho d\\mu\\right\),for somer∈ℕr\\in\{\\mathbb\{N\}\},ρ∈ℱp​\(ℝk,ℝr\)\\rho\\in\\mathcal\{F\}\_\{p\}\(\\mathbb\{R\}^\{k\},\\mathbb\{R\}^\{r\}\), andσ∈C​\(ℝr,ℝ\)\\sigma\\in C\\left\(\\mathbb\{R\}^\{r\},\\mathbb\{R\}\\right\)\. We fix a compact setQ⊆𝒫p​\(ℝk\)Q\\subseteq\\mathcal\{P\}\_\{p\}\(\\mathbb\{R\}^\{k\}\), and letℝr\\mathbb\{R\}^\{r\}be equipped with the Euclidean norm∥⋅∥ℝr\\\|\\cdot\\\|\_\{\\mathbb\{R\}^\{r\}\}\. By the growth condition onρ\\rho, there exist positive constantsMj\>0M\_\{j\}\>0such that for each component functionρj\\rho\_\{j\}, we have\|ρj​\(x\)\|≤Mj​\(1\+‖x‖ℝkp\)\|\\rho\_\{j\}\(x\)\|\\leq M\_\{j\}\(1\+\\\|x\\\|^\{p\}\_\{\\mathbb\{R\}^\{k\}\}\)for allx∈ℝkx\\in\\mathbb\{R\}^\{k\}\. Defining the vectorM≔\(M1,…,Mr\)⊤∈ℝrM\\coloneqq\(M\_\{1\},\\ldots,M\_\{r\}\)^\{\\top\}\\in\\mathbb\{R\}^\{r\}, it follows that‖ρ​\(x\)‖ℝr≤‖M‖ℝr​\(1\+‖x‖ℝkp\),∀x∈ℝk\.\\\|\\rho\(x\)\\\|\_\{\\mathbb\{R\}^\{r\}\}\\leq\\\|M\\\|\_\{\\mathbb\{R\}^\{r\}\}\(1\+\\\|x\\\|^\{p\}\_\{\\mathbb\{R\}^\{k\}\}\),\\;\\forall x\\in\\mathbb\{R\}^\{k\}\.Moreover, Theorem[4\.5](https://arxiv.org/html/2605.23156#S4.Thmtheorem5)ensures the continuity of the mapμ↦∫ρ​𝑑μ\\mu\\mapsto\\int\\rho d\\mu\. Thus, the image of the compact setQQunder this map, denoted byQσ⊆ℝrQ\_\{\\sigma\}\\subseteq\\mathbb\{R\}^\{r\}, is also compact, forming the input domain forσ\\sigma\.

##### Approximatingσ\\sigma\.

For anyε\>0\\varepsilon\>0, UAT ensures there existsσ^∈NN1,rϕ\\hat\{\\sigma\}\\in\\texttt\{NN\}^\{\\phi\}\_\{1,r\}such thatsupu∈Qσ\|σ​\(u\)−σ^​\(u\)\|≤ε2\.\\sup\_\{u\\in Q\_\{\\sigma\}\}\|\\sigma\(u\)\-\\hat\{\\sigma\}\(u\)\|\\leq\\frac\{\\varepsilon\}\{2\}\.Moreover, sinceσ^\\hat\{\\sigma\}is uniformly continuous on compact sets, there existsδ\>0\\delta\>0such that‖u1−u2‖ℝr≤δ⇒\|σ^​\(u1\)−σ^​\(u2\)\|≤ε2\.\\\|u\_\{1\}\-u\_\{2\}\\\|\_\{\\mathbb\{R\}^\{r\}\}\\leq\\delta\\Rightarrow\|\\hat\{\\sigma\}\(u\_\{1\}\)\-\\hat\{\\sigma\}\(u\_\{2\}\)\|\\leq\\frac\{\\varepsilon\}\{2\}\.

##### Tail Control and Truncation ofρ\\rho\.

SinceQQis compact, it is tight andpp\-uniformly integrable, by Proposition[4\.4](https://arxiv.org/html/2605.23156#S4.Thmtheorem4)\. Then for the constantδ4​‖M‖ℝr\\frac\{\\delta\}\{4\\\|M\\\|\_\{\\mathbb\{R\}^\{r\}\}\}, there exists a positive constantR\>0R\>0such that

supμ∈Qμ​\(ℝk∖B​\(0,R\)¯\)\\displaystyle\\sup\_\{\\mu\\in Q\}\\mu\\left\(\\mathbb\{R\}^\{k\}\\setminus\\overline\{B\(0,R\)\}\\right\)≤δ4​‖M‖ℝr\\displaystyle\\leq\\frac\{\\delta\}\{4\\\|M\\\|\_\{\\mathbb\{R\}^\{r\}\}\}\(tightness\),\\displaystyle\\text\{\(tightness\)\},supμ∈Q∫ℝk∖B​\(0,R\)¯‖x‖ℝkp​dμ\\displaystyle\\sup\_\{\\mu\\in Q\}\\int\_\{\\mathbb\{R\}^\{k\}\\setminus\\overline\{B\(0,R\)\}\}\\\|x\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{p\}\\mathrm\{~d\}\\mu≤δ4​‖M‖ℝr\\displaystyle\\leq\\frac\{\\delta\}\{4\\\|M\\\|\_\{\\mathbb\{R\}^\{r\}\}\}\(p\-uniformly integrability\)\.\\displaystyle\\text\{\($p$\-uniformly integrability\)\}\.Using Urysohn’s Lemma\[[55](https://arxiv.org/html/2605.23156#bib.bib55)\], there exists a continuous functionφ:ℝk→\[0,1\]\\varphi\\colon\\mathbb\{R\}^\{k\}\\rightarrow\[0,1\]such thatφ​\(B​\(0,R\)¯\)=1\\varphi\\left\(\\overline\{B\(0,R\)\}\\right\)=1andφ​\(ℝk∖B​\(0,2​R\)¯\)=0\\varphi\\left\(\\overline\{\\mathbb\{R\}^\{k\}\\setminus\{B\(0,2R\)\}\}\\right\)=0\. We truncate functionρ\\rhowithφ\\varphiby

ρφ​\(x\)≔φ​\(x\)​ρ​\(x\),∀x∈ℝk\.\\rho\_\{\\varphi\}\(x\)\\coloneqq\\varphi\(x\)\\rho\(x\),\\quad\\forall x\\in\\mathbb\{R\}^\{k\}\.Evidently,ρφ∈C0​\(ℝk,ℝr\)\\rho\_\{\\varphi\}\\in C\_\{0\}\(\\mathbb\{R\}^\{k\},\\mathbb\{R\}^\{r\}\), the space of continuous functions vanishing at infinity\. Furthermore, by the tail control ofρ\\rho, the discrepancy betweenρ\\rhoand its truncated counterpartρφ\\rho\_\{\\varphi\}can also be controlled onQQ

supμ∈Q∫‖ρφ−ρ‖ℝr​𝑑μ≤supμ∈Q∫ℝk∖B​\(0,R\)¯‖ρ​\(x\)‖ℝr​𝑑μ​\(x\)≤‖M‖ℝr​supμ∈Q∫ℝk∖B​\(0,R\)¯\(1\+‖x‖ℝkp\)​𝑑μ​\(x\)≤δ2\.\\sup\_\{\\mu\\in Q\}\\int\\\|\\rho\_\{\\varphi\}\-\\rho\\\|\_\{\\mathbb\{R\}^\{r\}\}d\\mu\\leq\\sup\_\{\\mu\\in Q\}\\int\_\{\\mathbb\{R\}^\{k\}\\setminus\\overline\{B\(0,R\)\}\}\\\|\\rho\(x\)\\\|\_\{\\mathbb\{R\}^\{r\}\}d\\mu\(x\)\\leq\\\|M\\\|\_\{\\mathbb\{R\}^\{r\}\}\\sup\_\{\\mu\\in Q\}\\int\_\{\\mathbb\{R\}^\{k\}\\setminus\\overline\{B\(0,R\)\}\}\\left\(1\+\\\|x\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{p\}\\right\)d\\mu\(x\)\\leq\\frac\{\\delta\}\{2\}\.

##### Approximatingρφ\\rho\_\{\\varphi\}\.

For the approximation ofC0C\_\{0\}functions, we invoke a specialized UAT\[[34](https://arxiv.org/html/2605.23156#bib.bib34)\]: if the activation functionϕ\\phiis continuous, nonpolynomial, and asymptotically polynomial at±∞\\pm\\infty, then any function inC0​\(ℝk,ℝ\)C\_\{0\}\(\\mathbb\{R\}^\{k\},\\mathbb\{R\}\)can be uniformly approximated byNN1,kϕ\\texttt\{NN\}^\{\\phi\}\_\{1,k\}onℝk\\mathbb\{R\}^\{k\}\. This result extends naturally to the vector\-valued spaceC0​\(ℝk,ℝr\)C\_\{0\}\(\\mathbb\{R\}^\{k\},\\mathbb\{R\}^\{r\}\)andNNr,kϕ\\texttt\{NN\}^\{\\phi\}\_\{r,k\}by approximating each component function independently\. Sinceρφ∈C0​\(ℝk,ℝr\),\\rho\_\{\\varphi\}\\in C\_\{0\}\\left\(\\mathbb\{R\}^\{k\},\\mathbb\{R\}^\{r\}\\right\),there existsρ^∈NNr,kϕ\\hat\{\\rho\}\\in\\texttt\{NN\}^\{\\phi\}\_\{r,k\}such thatsupx∈ℝk‖ρ^​\(x\)−ρφ​\(x\)‖ℝr≤δ2\.\\sup\_\{x\\in\\mathbb\{R\}^\{k\}\}\{\\\|\\hat\{\\rho\}\(x\)\-\\rho\_\{\\varphi\}\(x\)\\\|\_\{\\mathbb\{R\}^\{r\}\}\}\\leq\\frac\{\\delta\}\{2\}\.Thus,

supμ∈Q‖∫\(ρ−ρ^\)​𝑑μ‖ℝr≤supμ∈Q∫‖ρφ−ρ‖ℝr​𝑑μ\+supμ∈Q∫‖ρφ−ρ^‖ℝr​𝑑μ≤δ\.\\sup\_\{\\mu\\in Q\}\\left\\\|\\int\(\\rho\-\\hat\{\\rho\}\)d\\mu\\right\\\|\_\{\\mathbb\{R\}^\{r\}\}\\leq\\sup\_\{\\mu\\in Q\}\\int\\\|\\rho\_\{\\varphi\}\-\\rho\\\|\_\{\\mathbb\{R\}^\{r\}\}d\\mu\+\\sup\_\{\\mu\\in Q\}\\int\\\|\\rho\_\{\\varphi\}\-\\hat\{\\rho\}\\\|\_\{\\mathbb\{R\}^\{r\}\}d\\mu\\leq\\delta\.
We define the overall architecture asf^=σ^​\(∫ρ^​𝑑μ\)∈ℱDS¯\\hat\{f\}=\\hat\{\\sigma\}\(\\int\\hat\{\\rho\}d\\mu\)\\in\\mathcal\{F\}\_\{\\overline\{\\mathrm\{DS\}\}\}, and estimate the approximation error

supμ∈Q\|f​\(μ\)−f^​\(μ\)\|\\displaystyle\\sup\_\{\\mu\\in Q\}\\left\|f\(\\mu\)\-\\hat\{f\}\(\\mu\)\\right\|=supμ∈Q\|σ​\(∫ρ​𝑑μ\)−σ^​\(∫ρ^​𝑑μ\)\|\\displaystyle=\\sup\_\{\\mu\\in Q\}\\left\|\\sigma\\left\(\\int\\rho d\\mu\\right\)\-\\hat\{\\sigma\}\\left\(\\int\\hat\{\\rho\}d\\mu\\right\)\\right\|≤supμ∈Q\|σ​\(∫ρ​𝑑μ\)−σ^​\(∫ρ​𝑑μ\)\|\+supμ∈Q\|σ^​\(∫ρ​𝑑μ\)−σ^​\(∫ρ^​𝑑μ\)\|\\displaystyle\\leq\\sup\_\{\\mu\\in Q\}\\left\|\\sigma\\left\(\\int\\rho d\\mu\\right\)\-\\hat\{\\sigma\}\\left\(\\int\\rho d\\mu\\right\)\\right\|\+\\sup\_\{\\mu\\in Q\}\\left\|\\hat\{\\sigma\}\\left\(\\int\\rho d\\mu\\right\)\-\\hat\{\\sigma\}\\left\(\\int\\hat\{\\rho\}d\\mu\\right\)\\right\|≤ε2\+ε2=ε,\\displaystyle\\leq\\frac\{\\varepsilon\}\{2\}\+\\frac\{\\varepsilon\}\{2\}=\\varepsilon,which concludes thatℱDS¯\\mathcal\{F\}\_\{\\overline\{\\mathrm\{DS\}\}\}is dense inℱcont¯\\mathcal\{F\}\_\{\\overline\{\\mathrm\{cont\}\}\}\. ∎

###### Proof of Theorem[4\.6](https://arxiv.org/html/2605.23156#S4.Thmtheorem6)\.

By Proposition[3\.1](https://arxiv.org/html/2605.23156#S3.Thmtheorem1), the Stone–Weierstrass Theorem, and Lemmas[C\.5](https://arxiv.org/html/2605.23156#A3.Thmtheorem5)and[C\.6](https://arxiv.org/html/2605.23156#A3.Thmtheorem6), we conclude thatℱcont¯\\mathcal\{F\}\_\{\\overline\{\\mathrm\{cont\}\}\}is dense inC​\(Q\)C\(Q\)\. Combining this with Lemma[C\.7](https://arxiv.org/html/2605.23156#A3.Thmtheorem7)and the transitivity of density, it follows thatℱDS¯\\mathcal\{F\}\_\{\\overline\{\\mathrm\{DS\}\}\}is dense inC​\(Q\)C\(Q\)\. ∎

## Appendix DDetails and missing proofs from Section[4\.2](https://arxiv.org/html/2605.23156#S4.SS2)

In this section, we present additional details and proofs for our results on graphs\.

### D\.1Duplication consistent sequence for graphs

We start by elaborating on the consistent sequence for graphs used in Section[4\.2](https://arxiv.org/html/2605.23156#S4.SS2)\. More details and proofs for the following assertions can be found in\[[7](https://arxiv.org/html/2605.23156#bib.bib7), Appdx\. G\]\.

The duplication embedding consistent sequence for graphs𝕍dupG=\{\(Vn\),\(φN,n\),\(Gn\)\}\\mathbb\{V\}\_\{\\mathrm\{dup\}\}^\{G\}=\\left\\\{\\left\(V\_\{n\}\\right\),\\left\(\\varphi\_\{N,n\}\\right\),\\left\(\\mathrm\{G\}\_\{n\}\\right\)\\right\\\}is defined as follows\. The index set\(ℕ,⋅∣⋅\)\(\{\\mathbb\{N\}\},\\cdot\\mid\\cdot\)is the set of natural numbers with divisibility partial order, wheren⪯Nn\\preceq Nif and only ifn∣Nn\\mid N\. For eachn∈ℕn\\in\{\\mathbb\{N\}\},Vn=ℝsymn×nV\_\{n\}=\\mathbb\{R\}^\{n\\times n\}\_\{\\text\{sym\}\}\. Forn⪯Nn\\preceq N, the duplication embeddingφN,n:ℝsymn×n↪ℝsymN×N\\varphi\_\{N,n\}\\colon\\mathbb\{R\}^\{n\\times n\}\_\{\\text\{sym\}\}\\hookrightarrow\\mathbb\{R\}^\{N\\times N\}\_\{\\text\{sym\}\}is given byφN,n​\(A\)=A⊗\(𝟙N/n​𝟙N/n⊤\),\\varphi\_\{N,n\}\(A\)=A\\otimes\\left\(\\mathbbm\{1\}\_\{N/n\}\\mathbbm\{1\}\_\{N/n\}^\{\\top\}\\right\),forA∈ℝsymn×nA\\in\\mathbb\{R\}^\{n\\times n\}\_\{\\text\{sym\}\}, where⊗\\otimesdenotes the Kronecker product\. The group is the symmetric groupSnS\_\{n\}and the group embedding is the same as the case in Appendix[C\.1\.2](https://arxiv.org/html/2605.23156#A3.SS1.SSS2)\.SnS\_\{n\}acts onVnV\_\{n\}viag⋅A=g​A​g⊤g\\cdot A=gAg^\{\\top\}\.

The spaceV∞V\_\{\\infty\}can be identified with the space of step graphonsWA:\[0,1\]2→ℝW\_\{A\}\\colon\[0,1\]^\{2\}\\rightarrow\\mathbb\{R\}given by

WA​\(x,y\)=Ai​jfor​\(x,y\)∈\(i−1n,in\]×\(j−1n,jn\],i,j∈\[n\]\.W\_\{A\}\(x,y\)=A\_\{ij\}\\quad\\text\{ for \}\(x,y\)\\in\\left\(\\frac\{i\-1\}\{n\},\\frac\{i\}\{n\}\\right\]\\times\\left\(\\frac\{j\-1\}\{n\},\\frac\{j\}\{n\}\\right\],i,j\\in\[n\]\.The symmetric group acts on the induced step graphon by

σ⋅WA=WAσ−1≔WA​\(σ−1​\(x\),σ−1​\(y\)\)\.\\sigma\\cdot W\_\{A\}=W\_\{A\}^\{\\sigma^\{\-1\}\}\\coloneqq W\_\{A\}\(\\sigma^\{\-1\}\(x\),\\sigma^\{\-1\}\(y\)\)\.We endow eachVnV\_\{n\}with the cut norm, which is defined as

‖A‖□≔1n2​maxS⊆\[n\],T⊆\[n\]⁡\|∑i∈S,j∈TAi​j\|for​A∈ℝsymn×n\.\\\|A\\\|\_\{\\square\}\\coloneqq\\frac\{1\}\{n^\{2\}\}\\max\_\{S\\subseteq\[n\],T\\subseteq\[n\]\}\\left\|\\sum\_\{i\\in S,j\\in T\}A\_\{ij\}\\right\|\\quad\\text\{for \}A\\in\\mathbb\{R\}^\{n\\times n\}\_\{\\text\{sym\}\}\.This cut norm extends to a norm inV∞V\_\{\\infty\}, which coincides with the cut norm on graphons, defined as:

‖W‖□≔supS,T⊆\[0,1\]\|∫S×TW​\(x,y\)​𝑑x​𝑑y\|\.\\\|W\\\|\_\{\\square\}\\coloneqq\\sup\_\{S,T\\subseteq\[0,1\]\}\\left\|\\int\_\{S\\times T\}W\(x,y\)dxdy\\right\|\.The symmetrized distance coincides with theδ□\\delta\_\{\\square\}distance, which is defined by

d¯​\(W,U\)=δ□​\(W,U\)≔infφ∈S\[0,1\]‖\(Uφ−W\)‖□,\\overline\{\\mathrm\{d\}\}\(W,U\)=\\delta\_\{\\square\}\(W,U\)\\coloneqq\\inf\_\{\\varphi\\in S\_\{\[0,1\]\}\}\\\|\\left\(U^\{\\varphi\}\-W\\right\)\\\|\_\{\\square\},whereS\[0,1\]S\_\{\[0,1\]\}denotes the set of all measure\-preserving bijections\[0,1\]→\[0,1\]\[0,1\]\\rightarrow\[0,1\]; andUφ​\(x,y\)≔U​\(φ​\(x\),φ​\(y\)\)U^\{\\varphi\}\(x,y\)\\coloneqq U\(\\varphi\(x\),\\varphi\(y\)\)\. The orbit space can be identified with the space of symmetric measurable functions\[0,1\]2→ℝ\[0,1\]^\{2\}\\rightarrow\\mathbb\{R\}, modulo the equivalence relationW1∼W2W\_\{1\}\\sim W\_\{2\}wheneverδ□​\(W1,W2\)=0\.\\delta\_\{\\square\}\(W\_\{1\},W\_\{2\}\)=0\.

### D\.2Proof of Theorem[4\.9](https://arxiv.org/html/2605.23156#S4.Thmtheorem9)

###### Proof of Theorem[4\.9](https://arxiv.org/html/2605.23156#S4.Thmtheorem9)\.

We expand the formulation \([9](https://arxiv.org/html/2605.23156#S4.E9)\) and get

hm​\(W\)=∫\[0,1\]m∏1≤i<j≤m\(ai​j​W​\(xi,xj\)\+bi​j\)​∏ℓ=1md​xℓ=∑S⊆Em\(∏e∈Sae​∏e∉Sbe\)​t​\(S;W\),\\displaystyle h\_\{m\}\(W\)=\\int\_\{\[0,1\]^\{m\}\}\\prod\_\{1\\leq i<j\\leq m\}\\left\(a\_\{ij\}W\(x\_\{i\},x\_\{j\}\)\+b\_\{ij\}\\right\)\\prod\_\{\\ell=1\}^\{m\}dx\_\{\\ell\}=\\sum\_\{S\\subseteq E\_\{m\}\}\\left\(\\prod\_\{e\\in S\}a\_\{e\}\\prod\_\{e\\notin S\}b\_\{e\}\\right\)t\(S;W\),whereEmE\_\{m\}denotes the set of edges of themm\-node complete graph\. Thus, eachhm​\(⋅\)h\_\{m\}\(\\cdot\)is a linear combination of homomorphism densities of simple graphs, which are continuous with respect to the cut distanceδ□\\delta\_\{\\square\}\[[59](https://arxiv.org/html/2605.23156#bib.bib59), Thm\. 2\.7\], and henceℱ𝒲⊆C​\(𝒲0,δ□\)\\mathcal\{F\}\_\{\\mathcal\{W\}\}\\subseteq C\\left\(\{\\mathcal\{W\}\_\{0\}\},\\delta\_\{\\square\}\\right\)as claimed\. ∎

### D\.3Proof of Theorem[4\.10](https://arxiv.org/html/2605.23156#S4.Thmtheorem10)

###### Proof of Theorem[4\.10](https://arxiv.org/html/2605.23156#S4.Thmtheorem10)\.

To establish the density ofℱ𝒲\\mathcal\{F\}\_\{\\mathcal\{W\}\}, it suffices to show thatℋ​𝒟⊆ℱ𝒲\\mathcal\{HD\}\\subseteq\\mathcal\{F\}\_\{\\mathcal\{W\}\}, whereℋ​𝒟≔span⁡\{t​\(F,⋅\)∣F​is a simple graph\}\\mathcal\{HD\}\\coloneqq\\operatorname\{span\}\\\{\\,t\(F,\\cdot\)\\mid F\\text\{ is a simple graph\}\\,\\\}is known to be dense inC​\(𝒲0,δ□\)C\\left\(\{\\mathcal\{W\}\_\{0\}\},\\delta\_\{\\square\}\\right\)by\[[60](https://arxiv.org/html/2605.23156#bib.bib60), Thm\. 2\.2\]\.

For any simple graphF=\(V,E\)F=\(V,E\)with\|V\|=k≤m\|V\|=k\\leq m, consider the functionalhm​\(W\)h\_\{m\}\(W\)defined in \([9](https://arxiv.org/html/2605.23156#S4.E9)\) with parameters

ai​j=\{1if​\{i,j\}∈E0if​\{i,j\}∉E,andbi​j=\{0if​\{i,j\}∈E1if​\{i,j\}∉E,a\_\{ij\}=\\begin\{cases\}1&\\text\{if \}\\\{i,j\\\}\\in E\\\\ 0&\\text\{if \}\\\{i,j\\\}\\notin E\\end\{cases\},\\quad\\text\{and\}\\quad b\_\{ij\}=\\begin\{cases\}0&\\text\{if \}\\\{i,j\\\}\\in E\\\\ 1&\\text\{if \}\\\{i,j\\\}\\notin E\\end\{cases\},and note thathm​\(W\)=t​\(F,W\)h\_\{m\}\(W\)=t\(F,W\), proving thatℋ​𝒟⊆ℱ𝒲\\mathcal\{HD\}\\subseteq\\mathcal\{F\}\_\{\\mathcal\{W\}\}as desired\. ∎

### D\.4A universal deep model for graphs

First, we introduce a tensor contraction operator that serves as the nonlinear activation function in the network\. Fork∈ℕ0k\\in\{\{\\mathbb\{N\}\}\}\_\{0\}, define a class of functions𝒢k≔\{G:\[0,1\]k→ℝmeasurable, bounded\}/∼a\.e\.\\mathcal\{G\}\_\{k\}\\coloneqq\\\{G:\[0,1\]^\{k\}\\rightarrow\\mathbb\{R\}~\\text\{ measurable, bounded \}\\\}/\\sim\_\{\\text\{a\.e\.\}\}whereG1∼a\.e\.G2G\_\{1\}\\sim\_\{\\text\{a\.e\.\}\}G\_\{2\}ifG1=G2G\_\{1\}=G\_\{2\}almost everywhere, and set𝒢0=ℝ\\mathcal\{G\}\_\{0\}=\\mathbb\{R\}\.

###### Definition D\.1\.

We define the operatorT:𝒢k\+1×∏i=1k𝒢2→𝒢kT\\colon\\mathcal\{G\}\_\{k\+1\}\\times\\prod\_\{i=1\}^\{k\}\\mathcal\{G\}\_\{2\}\\to\\mathcal\{G\}\_\{k\}as

\(T​\(P,G1,…,Gk\)\)​\(y1,…,yk\)=∫\[0,1\]P​\(x,y1,…,yk\)​∏i=1kGi​\(x,yi\)​d​x\.\\displaystyle\\left\(T\(P,G\_\{1\},\\dots,G\_\{k\}\)\\right\)\(y\_\{1\},\\ldots,y\_\{k\}\)=\\int\_\{\[0,1\]\}P\(x,y\_\{1\},\\ldots,y\_\{k\}\)\\prod\_\{i=1\}^\{k\}G\_\{i\}\(x,y\_\{i\}\)dx\.Fork=0k=0, the operatorT:𝒢1→ℝT\\colon\\mathcal\{G\}\_\{1\}\\rightarrow\\mathbb\{R\}is given by∫\[0,1\]P​\(x\)​𝑑x\\int\_\{\[0,1\]\}P\(x\)dx\.

We define thejj\-th linear layer as

ℒj​\(Pj,W\)≔\(Pjaj1​W\+bj1⋮ajkj​W\+bjkj\),\\mathcal\{L\}\_\{j\}\(P\_\{j\},W\)\\coloneqq\\left\(\\begin\{array\}\[\]\{ll\}\\quad P\_\{j\}\\\\ a\_\{j\}^\{1\}W\+b\_\{j\}^\{1\}\\\\ \\quad\\;\\;\\vdots\\\\ a\_\{j\}^\{k\_\{j\}\}W\+b\_\{j\}^\{k\_\{j\}\}\\\\ \\end\{array\}\\right\),wherePj∈𝒢kj\+1P\_\{j\}\\in\{\\mathcal\{G\}\}\_\{k\_\{j\}\+1\}, for somekj∈ℕ0k\_\{j\}\\in\{\{\\mathbb\{N\}\}\}\_\{0\};aji,bji∈ℝa\_\{j\}^\{i\},b\_\{j\}^\{i\}\\in\\mathbb\{R\}, fori∈\[kj\]i\\in\[k\_\{j\}\];W∈𝒲0W\\in\{\\mathcal\{W\}\_\{0\}\}\. We denote thejj\-th linear layer map asℒ~j:𝒢kj\+1×𝒲0→\(𝒢kj\+1×∏i=1kj𝒢2\)×𝒲0\\widetilde\{\\mathcal\{L\}\}\_\{j\}\\colon\{\\mathcal\{G\}\}\_\{k\_\{j\}\+1\}\\times\{\\mathcal\{W\}\_\{0\}\}\\rightarrow\\left\(\{\\mathcal\{G\}\}\_\{k\_\{j\}\+1\}\\times\{\\prod\}\_\{i=1\}^\{k\_\{j\}\}\\mathcal\{G\}\_\{2\}\\right\)\\times\{\\mathcal\{W\}\_\{0\}\}given byℒ~j​\(Pj,W\)=\(ℒj​\(Pj,W\),W\)\.\\widetilde\{\\mathcal\{L\}\}\_\{j\}\(P\_\{j\},W\)=\(\\mathcal\{L\}\_\{j\}\(P\_\{j\},W\),W\)~\.Furthermore, we denote the nonlinearity in the network asT~:\(𝒢k\+1×∏i=1k𝒢2\)×𝒲0→𝒢k×𝒲0\\widetilde\{T\}\\colon\\left\(\\mathcal\{G\}\_\{k\+1\}\\times\\prod\_\{i=1\}^\{k\}\\mathcal\{G\}\_\{2\}\\right\)\\times\{\\mathcal\{W\}\_\{0\}\}\\to\\mathcal\{G\}\_\{k\}\\times\{\\mathcal\{W\}\_\{0\}\}given byT~​\(\(P,G1,…,Gk\),W\)=\(T​\(P,G1,…,Gk\),W\)\\widetilde\{T\}\\left\(\(P,G\_\{1\},\\ldots,G\_\{k\}\),W\\right\)=\\left\(T\(P,G\_\{1\},\\ldots,G\_\{k\}\),W\\right\), whereTTis the operator in Definition[D\.1](https://arxiv.org/html/2605.23156#A4.Thmtheorem1)\.

We then contract the output ofℒj\\mathcal\{L\}\_\{j\}usingTTdefined in Definition[D\.1](https://arxiv.org/html/2605.23156#A4.Thmtheorem1)to get the input of\(j\+1\)\(j\+1\)\-th linear layer

T∘ℒj​\(Pj,W\)≔Pj\+1∈𝒢kj\.T\\circ\\mathcal\{L\}\_\{j\}\(P\_\{j\},W\)\\coloneqq P\_\{j\+1\}\\in\{\\mathcal\{G\}\}\_\{k\_\{j\}\}\.
Suppose we havem≥1m\\geq 1layers, setP1∈𝒢m,P1​\(x1,…,xm\)≡1P\_\{1\}\\in\\mathcal\{G\}\_\{m\},P\_\{1\}\(x\_\{1\},\\ldots,x\_\{m\}\)\\equiv 1, and letkj=m−jk\_\{j\}=m\-j\. For thejj\-th tensor output,1≤j≤m\+11\\leq j\\leq m\+1, we havePj∈𝒢m\+1−jP\_\{j\}\\in\\mathcal\{G\}\_\{m\+1\-j\}\. The overall architecturehDm:𝒲0→ℝh^\{m\}\_\{D\}:\{\\mathcal\{W\}\_\{0\}\}\\rightarrow\\mathbb\{R\}can be written as

\(hDm​\(W\),W\)=T~∘ℒ~m∘⋯∘T~∘ℒ~1​\(P1,W\)\.\(h\_\{D\}^\{m\}\(W\),W\)=\\widetilde\{T\}\\circ\\widetilde\{\\mathcal\{L\}\}\_\{m\}\\circ\\cdots\\circ\\widetilde\{T\}\\circ\\widetilde\{\\mathcal\{L\}\}\_\{1\}\(P\_\{1\},W\)\.\(15\)
###### Theorem D\.2\.

Letℱd​e​e​p\\mathcal\{F\}\_\{deep\}be the class of functions defined asℱd​e​e​p≔span​\{hDm∣m∈ℕ\}\\mathcal\{F\}\_\{deep\}\\coloneqq\\mathrm\{span\}\\\{\{h^\{m\}\_\{D\}\\mid m\\in\{\\mathbb\{N\}\}\}\\\}, where eachhDmh\_\{D\}^\{m\}has the form of \([15](https://arxiv.org/html/2605.23156#A4.E15)\)\. Thenℱd​e​e​p\\mathcal\{F\}\_\{deep\}is dense inC​\(𝒲0,δ□\)C\\left\(\{\\mathcal\{W\}\_\{0\}\},\\delta\_\{\\square\}\\right\)\.

###### Proof of Theorem[D\.2](https://arxiv.org/html/2605.23156#A4.Thmtheorem2)\.

We establish the result by induction, demonstrating that the deep model form \([15](https://arxiv.org/html/2605.23156#A4.E15)\) recovers the structure of \([9](https://arxiv.org/html/2605.23156#S4.E9)\), and complete the proof by Theorem[4\.10](https://arxiv.org/html/2605.23156#S4.Thmtheorem10)\. We prove inductively that for anyr∈ℕr\\in\{\\mathbb\{N\}\}such that2≤r≤m\+12\\leq r\\leq m\+1, the output tensorPr∈𝒢m\+1−rP\_\{r\}\\in\\mathcal\{G\}\_\{m\+1\-r\}of therrth layer can be represented as

Pr​\(𝐱≥r\)=∫\[0,1\]r−1∏1≤i<j≤r−1Lij−i​\(W​\(xi,xj\)\)​∏i=1r−1∏j=rmLij−i​\(W​\(xi,xj\)\)​d​𝐱<r,P\_\{r\}\(\\mathbf\{x\}\_\{\\geq r\}\)=\\int\_\{\[0,1\]^\{r\-1\}\}\\prod\_\{1\\leq i<j\\leq r\-1\}L\_\{i\}^\{j\-i\}\\left\(W\(x\_\{i\},x\_\{j\}\)\\right\)\\prod\_\{i=1\}^\{r\-1\}\\prod\_\{j=r\}^\{m\}L\_\{i\}^\{j\-i\}\\left\(W\(x\_\{i\},x\_\{j\}\)\\right\)d\\mathbf\{x\}\_\{<r\},\(16\)whereLiv​\(W\)≔aiv​W\+bivL\_\{i\}^\{v\}\(W\)\\coloneqq a\_\{i\}^\{v\}W\+b\_\{i\}^\{v\}withaiv,biv∈ℝa\_\{i\}^\{v\},b\_\{i\}^\{v\}\\in\\mathbb\{R\},𝐱≥r≔\(xr,…,xm\)\\mathbf\{x\}\_\{\\geq r\}\\coloneqq\(x\_\{r\},\\dots,x\_\{m\}\), andd​𝐱<r≔d​x1​…​d​xr−1d\\mathbf\{x\}\_\{<r\}\\coloneqq dx\_\{1\}\\dots dx\_\{r\-1\}\. Having done so, we setr=m\+1r=m\+1and conclude that

Pm\+1=∫\[0,1\]m∏1≤i<j≤mLij−i​\(W​\(xi,xj\)\)​d​𝐱<m\+1=∫\[0,1\]m∏1≤i<j≤m\(aij−i​W​\(xi,xj\)\+bij−i\)​d​𝐱<m\+1\.P\_\{m\+1\}=\\int\_\{\[0,1\]^\{m\}\}\\prod\_\{1\\leq i<j\\leq m\}L\_\{i\}^\{j\-i\}\(W\(x\_\{i\},x\_\{j\}\)\)d\\mathbf\{x\}\_\{<m\+1\}=\\int\_\{\[0,1\]^\{m\}\}\\prod\_\{1\\leq i<j\\leq m\}\\left\(a\_\{i\}^\{j\-i\}W\(x\_\{i\},x\_\{j\}\)\+b\_\{i\}^\{j\-i\}\\right\)d\\mathbf\{x\}\_\{<m\+1\}\.By settingai​j≔aij−ia\_\{ij\}\\coloneqq a\_\{i\}^\{j\-i\}andbi​j≔bij−ib\_\{ij\}\\coloneqq b\_\{i\}^\{j\-i\}, this expression exactly matches the form \([9](https://arxiv.org/html/2605.23156#S4.E9)\), completing the proof\. Thus, we proceed to establish \([16](https://arxiv.org/html/2605.23156#A4.E16)\)\.

##### Base Case \(r=2r=2\)\.

Directly applying Definition[D\.1](https://arxiv.org/html/2605.23156#A4.Thmtheorem1), we have

P2​\(x2,…,xm\)=∫\[0,1\]∏u=1m−1\(a1u​W​\(x1,xu\+1\)\+b1u\)​d​x1=∫\[0,1\]∏j=2mL1j−1​\(W​\(x1,xj\)\)​d​x1P\_\{2\}\(x\_\{2\},\\ldots,x\_\{m\}\)=\\int\_\{\[0,1\]\}\\prod\_\{u=1\}^\{m\-1\}\\left\(a\_\{1\}^\{u\}W\(x\_\{1\},x\_\{u\+1\}\)\+b\_\{1\}^\{u\}\\right\)dx\_\{1\}=\\int\_\{\[0,1\]\}\\prod\_\{j=2\}^\{m\}L\_\{1\}^\{j\-1\}\(W\(x\_\{1\},x\_\{j\}\)\)dx\_\{1\}which is consistent with the inductive hypothesis \([16](https://arxiv.org/html/2605.23156#A4.E16)\)\.

##### Inductive Step\.

Suppose the hypothesis holds for somer∈ℕr\\in\{\\mathbb\{N\}\}\(2≤r≤m2\\leq r\\leq m\), such thatPr∈𝒢m\+1−rP\_\{r\}\\in\\mathcal\{G\}\_\{m\+1\-r\}takes the form of \([16](https://arxiv.org/html/2605.23156#A4.E16)\)\. Then, forr\+1r\+1, we obtain

Pr\+1​\(xr\+1,…,xm\)\\displaystyle P\_\{r\+1\}\(x\_\{r\+1\},\\ldots,x\_\{m\}\)=∫\[0,1\]Pr​\(xr,xr\+1,…,xm\)​∏u=1m−r\(aru​W​\(xr,xr\+u\)\+bru\)​d​xr\\displaystyle=\\int\_\{\[0,1\]\}P\_\{r\}\(x\_\{r\},x\_\{r\+1\},\\ldots,x\_\{m\}\)\\prod\_\{u=1\}^\{m\-r\}\\left\(a\_\{r\}^\{u\}W\(x\_\{r\},x\_\{r\+u\}\)\+b\_\{r\}^\{u\}\\right\)dx\_\{r\}=∫\[0,1\]\(∫\[0,1\]r−1∏1≤i<j≤r−1Lij−i​\(W​\(xi,xj\)\)​∏i=1r−1∏j=rmLij−i​\(W​\(xi,xj\)\)​d​𝐱<r\)​∏j=r\+1mLrj−r​\(W​\(xr,xj\)\)​d​xr\\displaystyle=\\int\_\{\[0,1\]\}\\left\(\\int\_\{\[0,1\]^\{r\-1\}\}\\prod\_\{1\\leq i<j\\leq r\-1\}L\_\{i\}^\{j\-i\}\(W\(x\_\{i\},x\_\{j\}\)\)\\prod\_\{i=1\}^\{r\-1\}\\prod\_\{j=r\}^\{m\}L\_\{i\}^\{j\-i\}\(W\(x\_\{i\},x\_\{j\}\)\)d\\mathbf\{x\}\_\{<r\}\\right\)\\prod\_\{j=r\+1\}^\{m\}L\_\{r\}^\{j\-r\}\(W\(x\_\{r\},x\_\{j\}\)\)dx\_\{r\}=∫\[0,1\]r\(∏1≤i<j≤r−1Lij−i​\(W​\(xi,xj\)\)\)​\(∏i=1r−1Lir−i​\(W​\(xi,xr\)\)\)​\(∏i=1r∏j=r\+1mLij−i​\(W​\(xi,xj\)\)\)​𝑑𝐱<r\+1\\displaystyle=\\int\_\{\[0,1\]^\{r\}\}\\left\(\\prod\_\{1\\leq i<j\\leq r\-1\}L\_\{i\}^\{j\-i\}\(W\(x\_\{i\},x\_\{j\}\)\)\\right\)\\left\(\\prod\_\{i=1\}^\{r\-1\}L\_\{i\}^\{r\-i\}\(W\(x\_\{i\},x\_\{r\}\)\)\\right\)\\left\(\\prod\_\{i=1\}^\{r\}\\prod\_\{j=r\+1\}^\{m\}L\_\{i\}^\{j\-i\}\(W\(x\_\{i\},x\_\{j\}\)\)\\right\)d\\mathbf\{x\}\_\{<r\+1\}=∫\[0,1\]r∏1≤i<j≤rLij−i​\(W​\(xi,xj\)\)​∏i=1r∏j=r\+1mLij−i​\(W​\(xi,xj\)\)​d​𝐱<r\+1\.\\displaystyle=\\int\_\{\[0,1\]^\{r\}\}\\prod\_\{1\\leq i<j\\leq r\}L\_\{i\}^\{j\-i\}\(W\(x\_\{i\},x\_\{j\}\)\)\\prod\_\{i=1\}^\{r\}\\prod\_\{j=r\+1\}^\{m\}L\_\{i\}^\{j\-i\}\(W\(x\_\{i\},x\_\{j\}\)\)d\\mathbf\{x\}\_\{<r\+1\}\.This confirms that the hypothesis holds forr\+1r\+1\. ∎

We remark that the above architecture can be made more expressive by allowing general widths for the linear layers, and by using general linear equivariant maps\[[46](https://arxiv.org/html/2605.23156#bib.bib46)\]\. The latter would require adapting the nonlinear operators in Definition[D\.1](https://arxiv.org/html/2605.23156#A4.Thmtheorem1)appropriately, and we do not further explore these extensions\.

## Appendix EDetails and missing proofs from Section[4\.3](https://arxiv.org/html/2605.23156#S4.SS3)

### E\.1Duplication consistent sequence for point clouds

We elaborate on the consistent sequence used in Section[4\.3](https://arxiv.org/html/2605.23156#S4.SS3)\. See\[[7](https://arxiv.org/html/2605.23156#bib.bib7), Appx\. H\]for more details and proofs of the following assertions\.

Similar to the case of graphs, the duplication embedding consistent sequence for point clouds𝕍dupP=\{\(Vn\),\(φN,n\),\(Gn\)\}\\mathbb\{V\}\_\{\\mathrm\{dup\}\}^\{P\}=\\left\\\{\\left\(V\_\{n\}\\right\),\\left\(\\varphi\_\{N,n\}\\right\),\\left\(\\mathrm\{G\}\_\{n\}\\right\)\\right\\\}is defined as follows\. The index set\(ℕ,⋅∣⋅\)\(\{\\mathbb\{N\}\},\\cdot\\mid\\cdot\)is the set of natural numbers with divisibility partial order, wheren⪯Nn\\preceq Nif and only ifn∣Nn\\mid N\. For eachn∈ℕn\\in\{\\mathbb\{N\}\},Vn=ℝn×kV\_\{n\}=\\mathbb\{R\}^\{n\\times k\}, which represents sets ofnnpoints inℝk\\mathbb\{R\}^\{k\}, andkkis fixed\. The group isGn=Sn×O​\(k\)\\mathrm\{G\}\_\{n\}=\\mathrm\{S\}\_\{n\}\\times\\mathrm\{O\}\(k\), whereSn\\mathrm\{S\}\_\{n\}is the permutation group, andO​\(k\)\\mathrm\{O\}\(k\)is the orthogonal group\. The group action onVnV\_\{n\}is defined as

\(g,h\)⋅X=g​X​h⊤\.\(g,h\)\\cdot X=gXh^\{\\top\}\.Forn⪯Nn\\preceq N, the duplication embeddingφN,n:ℝn×k↪ℝN×k\\varphi\_\{N,n\}\\colon\\mathbb\{R\}^\{n\\times k\}\\hookrightarrow\\mathbb\{R\}^\{N\\times k\}is given byφN,n​\(X\)=X⊗𝟙N/n,\\varphi\_\{N,n\}\(X\)=X\\otimes\\mathbbm\{1\}\_\{N/n\},and the group embeddingθN,n:Sn×O​\(k\)↪SN×O​\(k\)\\theta\_\{N,n\}\\colon\\mathrm\{S\}\_\{n\}\\times\\mathrm\{O\}\(k\)\\hookrightarrow\\mathrm\{S\}\_\{N\}\\times\\mathrm\{O\}\(k\)is given byθN,n​\(g,h\)=\(g⊗IN/n,h\)\.\\theta\_\{N,n\}\(g,h\)=\\left\(g\\otimes I\_\{N/n\},h\\right\)\.We consider the Euclidean norm onℝk\\mathbb\{R\}^\{k\}denoted by∥⋅∥ℝk\\\|\\cdot\\\|\_\{\\mathbb\{R\}^\{k\}\}, which corresponds to the inner product preserved by elements ofO​\(k\)\\mathrm\{O\}\(k\)\. We equip eachVnV\_\{n\}with the normalizedℓ2\\ell\_\{2\}norm:

‖X‖2¯=\(1n​∑i=1n‖Xi:‖ℝk2\)1/2\.\\\|X\\\|\_\{\\bar\{2\}\}=\\left\(\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\left\\\|X\_\{i:\}\\right\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{2\}\\right\)^\{1/2\}\.Similarly to the case of sets, we can identify each matrixX∈ℝn×kX\\in\\mathbb\{R\}^\{n\\times k\}with a step function\[0,1\]→ℝk\[0,1\]\\rightarrow\\mathbb\{R\}^\{k\}\. Then the limit space can be identified withV∞¯=L2​\(\[0,1\];ℝk\)\\overline\{V\_\{\\infty\}\}=L^\{2\}\\left\(\[0,1\];\\mathbb\{R\}^\{k\}\\right\), and the symmetrized metric can be written as:

d¯​\(X,Y\)=infh∈O​\(k\)infφ∈S\[0,1\]\(∫01‖h​\(X​\(t\)\)−Y​\(φ​\(t\)\)‖ℝk2​𝑑t\)1/2,\\overline\{\\mathrm\{d\}\}\(X,Y\)=\\inf\_\{h\\in\\mathrm\{O\}\(k\)\}\\inf\_\{\\varphi\\in S\_\{\[0,1\]\}\}\\left\(\\int\_\{0\}^\{1\}\\\|h\(X\(t\)\)\-Y\(\\varphi\(t\)\)\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{2\}dt\\right\)^\{1/2\},whereS\[0,1\]S\_\{\[0,1\]\}is the set of measure preserving bijections\.

We can further identify each matrixX∈ℝn×kX\\in\\mathbb\{R\}^\{n\\times k\}with an empirical probability measure inℝk\\mathbb\{R\}^\{k\}\. The space of orbit closuresV∞¯/G∞\\overline\{V\_\{\\infty\}\}/G\_\{\\infty\}can then be identified with the space of orbits of probability measures onℝk\\mathbb\{R\}^\{k\}under theO​\(k\)\\mathrm\{O\}\(k\)action by pushforwardsh⋅μ=h\#​μh\\cdot\\mu=h\_\{\\\#\}\\muforh∈O​\(k\)h\\in\\mathrm\{O\}\(k\)\. From this perspective, we can further rewrite the symmetrized distance as

d¯​\(X,Y\)=infh∈O​\(k\)W2​\(h⋅μX,μY\)for​X,Y∈V∞¯,\\overline\{\\mathrm\{d\}\}\(X,Y\)=\\inf\_\{h\\in\\mathrm\{O\}\(k\)\}W\_\{2\}\\left\(h\\cdot\\mu\_\{X\},\\mu\_\{Y\}\\right\)\\quad\\text\{ for \}X,Y\\in\\overline\{V\_\{\\infty\}\},whereμX=X\#​λ\\mu\_\{X\}=X\_\{\\\#\}\\lambda,λ\\lambdais the Lebesgue measure on\[0,1\]\[0,1\]\.

### E\.2Proof of Proposition[4\.11](https://arxiv.org/html/2605.23156#S4.Thmtheorem11)

###### Proof of Proposition[4\.11](https://arxiv.org/html/2605.23156#S4.Thmtheorem11)\.

Consider the set of probability measures corresponding to the setKRK\_\{R\}, denoted byQR≔\{X\#​λ∣X∈KR\}Q\_\{R\}\\coloneqq\\\{X\_\{\\\#\}\\lambda\\mid X\\in K\_\{R\}\\\}\. We show thatQRQ\_\{R\}is equal to the set of probability measures withRR\-bounded support𝒫2​\(B​\(0,R\)¯\)≔\{μ∈𝒫2​\(ℝk\)∣supp⁡μ⊆B​\(0,R\)¯\}\\mathcal\{P\}\_\{2\}\\left\(\\overline\{B\(0,R\)\}\\right\)\\coloneqq\\\{\\mu\\in\\mathcal\{P\}\_\{2\}\\left\(\\mathbb\{R\}^\{k\}\\right\)\\mid\\operatorname\{supp\}\\mu\\subseteq\\overline\{B\(0,R\)\}\\\}\.

First, we show thatQR⊆𝒫2​\(B​\(0,R\)¯\)Q\_\{R\}\\subseteq\\mathcal\{P\}\_\{2\}\\left\(\\overline\{B\(0,R\)\}\\right\)\. For anyμ∈QR\\mu\\in Q\_\{R\}, there exists a functionX∈KRX\\in K\_\{R\}such thatμ=X\#​λ\\mu=X\_\{\\\#\}\\lambda\. By the definition of the push\-forward measure, for any Borel setB⊆ℝkB\\subseteq\\mathbb\{R\}^\{k\}, we haveμ​\(B\)=λ​\(X−1​\(B\)\)\\mu\(B\)=\\lambda\(X^\{\-1\}\(B\)\)\. SinceX∈KRX\\in K\_\{R\}, for almost everyt∈\[0,1\]t\\in\[0,1\]we have‖X​\(t\)‖ℝk≤R\\\|X\(t\)\\\|\_\{\\mathbb\{R\}^\{k\}\}\\leq R\. Consequently,supp⁡μ⊆B​\(0,R\)¯\\operatorname\{supp\}\\mu\\subseteq\\overline\{B\(0,R\)\}, which impliesQR⊆𝒫2​\(B​\(0,R\)¯\)Q\_\{R\}\\subseteq\\mathcal\{P\}\_\{2\}\\left\(\\overline\{B\(0,R\)\}\\right\)\.

Conversely, we show that𝒫2​\(B​\(0,R\)¯\)⊆QR\\mathcal\{P\}\_\{2\}\\left\(\\overline\{B\(0,R\)\}\\right\)\\subseteq Q\_\{R\}\. For anyμ∈𝒫2​\(B​\(0,R\)¯\)\\mu\\in\\mathcal\{P\}\_\{2\}\\left\(\\overline\{B\(0,R\)\}\\right\), there exists a Borel mappingX:\[0,1\]→B​\(0,R\)¯X\\colon\[0,1\]\\rightarrow\\overline\{B\(0,R\)\}such thatX\#​λ=μX\_\{\\\#\}\\lambda=\\mubecauseB​\(0,R\)¯\\overline\{B\(0,R\)\}is a standard Borel space\[[61](https://arxiv.org/html/2605.23156#bib.bib61), Thm\. 3\.3\.13\]\. Since the image ofXXis contained within the ball, we have‖X​\(t\)‖ℝk≤R\\\|X\(t\)\\\|\_\{\\mathbb\{R\}^\{k\}\}\\leq Rfor alltt\. Moreover, since∫\[0,1\]‖X​\(t\)‖ℝk2​𝑑λ​\(t\)≤R2<∞\\int\_\{\[0,1\]\}\\\|X\(t\)\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{2\}d\\lambda\(t\)\\leq R^\{2\}<\\infty, it follows thatX∈L2​\(\[0,1\];ℝk\)X\\in L^\{2\}\(\[0,1\];\\mathbb\{R\}^\{k\}\), and thusX∈KRX\\in K\_\{R\}\. This concludes that𝒫2​\(B​\(0,R\)¯\)⊆QR\\mathcal\{P\}\_\{2\}\\left\(\\overline\{B\(0,R\)\}\\right\)\\subseteq Q\_\{R\}\.

Consequently, the setQR=𝒫2​\(B​\(0,R\)¯\)Q\_\{R\}=\\mathcal\{P\}\_\{2\}\\left\(\\overline\{B\(0,R\)\}\\right\)is a compact subset of the Wasserstein space𝒫2​\(ℝk\)\\mathcal\{P\}\_\{2\}\\left\(\\mathbb\{R\}^\{k\}\\right\)with respect to theW2W\_\{2\}metric by Proposition[4\.4](https://arxiv.org/html/2605.23156#S4.Thmtheorem4)\. Since the canonical projection mapπ:𝒫2​\(ℝk\)→𝒫2​\(ℝk\)/O​\(k\)\\pi\\colon\\mathcal\{P\}\_\{2\}\\left\(\\mathbb\{R\}^\{k\}\\right\)\\rightarrow\\mathcal\{P\}\_\{2\}\\left\(\\mathbb\{R\}^\{k\}\\right\)/\\mathrm\{O\}\(k\)is continuous \(see the proof of Lemma[A\.4](https://arxiv.org/html/2605.23156#A1.Thmtheorem4)\), we conclude thatKR/G∞K\_\{R\}/G\_\{\\infty\}, which can be identified withQR/O​\(k\)Q\_\{R\}/\\mathrm\{O\}\(k\), is compact inV∞¯/G∞\\overline\{V\_\{\\infty\}\}/G\_\{\\infty\}\. ∎

### E\.3Proof of Theorem[4\.12](https://arxiv.org/html/2605.23156#S4.Thmtheorem12)

###### Proof of Theorem[4\.12](https://arxiv.org/html/2605.23156#S4.Thmtheorem12)\.

By\[[46](https://arxiv.org/html/2605.23156#bib.bib46)\], the IGN architecture in \([13](https://arxiv.org/html/2605.23156#S4.E13)\) is continuous with respect toδ2\\delta\_\{2\}distance\. Therefore, to establish the overall model is continuous, it suffices to show thatX↦WXX\\mapsto W\_\{X\}is continuous with symmetrized distanced¯\\overline\{\\mathrm\{d\}\}in its input andδ2\\delta\_\{2\}distance in its image\. ForX,Y∈KR=\{X∈L2​\(\[0,1\];ℝk\):‖X‖∞≤R\},X,Y\\in K\_\{R\}=\\left\\\{X\\in L^\{2\}\\left\(\[0,1\];\\mathbb\{R\}^\{k\}\\right\):\\\|X\\\|\_\{\\infty\}\\leq R\\right\\\},the distance of their corresponding covariance functions can be bounded by

δ2​\(WX,WY\)2\\displaystyle\\delta\_\{2\}\(W\_\{X\},W\_\{Y\}\)^\{2\}=14​R4​infφ∈S\[0,1\]\(∫\[0,1\]2\|⟨X​\(x\),X​\(y\)⟩−⟨Y​\(φ​\(x\)\),Y​\(φ​\(y\)\)⟩\|2​𝑑x​𝑑y\)\\displaystyle=\\frac\{1\}\{4R^\{4\}\}\\inf\_\{\\varphi\\in S\_\{\[0,1\]\}\}\\left\(\\int\_\{\[0,1\]^\{2\}\}\|\\langle X\(x\),X\(y\)\\rangle\-\\langle Y\(\\varphi\(x\)\),Y\(\\varphi\(y\)\)\\rangle\|^\{2\}dxdy\\right\)=14​R4​infφ∈S\[0,1\]\(∫\[0,1\]2\|⟨X​\(x\)−Y​\(φ​\(x\)\),X​\(y\)⟩\+⟨Y​\(φ​\(x\)\),X​\(y\)−Y​\(φ​\(y\)\)⟩\|2​𝑑x​𝑑y\)\\displaystyle=\\frac\{1\}\{4R^\{4\}\}\\inf\_\{\\varphi\\in S\_\{\[0,1\]\}\}\\left\(\\int\_\{\[0,1\]^\{2\}\}\|\\langle X\(x\)\-Y\(\\varphi\(x\)\),X\(y\)\\rangle\+\\langle Y\(\\varphi\(x\)\),X\(y\)\-Y\(\\varphi\(y\)\)\\rangle\|^\{2\}dxdy\\right\)≤14​R2​infφ∈S\[0,1\]\(∫\[0,1\]2\(‖X​\(x\)−Y​\(φ​\(x\)\)‖ℝk\+‖X​\(y\)−Y​\(φ​\(y\)\)‖ℝk\)2​𝑑x​𝑑y\)\.\\displaystyle\\leq\\frac\{1\}\{4R^\{2\}\}\\inf\_\{\\varphi\\in S\_\{\[0,1\]\}\}\\left\(\\int\_\{\[0,1\]^\{2\}\}\\left\(\\\|X\(x\)\-Y\(\\varphi\(x\)\)\\\|\_\{\\mathbb\{R\}^\{k\}\}\+\\\|X\(y\)\-Y\(\\varphi\(y\)\)\\\|\_\{\\mathbb\{R\}^\{k\}\}\\right\)^\{2\}dxdy\\right\)\.Applying Young’s inequality, namely,\(a\+b\)2≤2​\(a2\+b2\)\(a\+b\)^\{2\}\\leq 2\(a^\{2\}\+b^\{2\}\), we obtain

δ2​\(WX,WY\)2\\displaystyle\\delta\_\{2\}\(W\_\{X\},W\_\{Y\}\)^\{2\}≤12​R2​infφ∈S\[0,1\]\(∫\[0,1\]2\(‖X​\(x\)−Y​\(φ​\(x\)\)‖ℝk2\+‖X​\(y\)−Y​\(φ​\(y\)\)‖ℝk2\)​𝑑x​𝑑y\)\\displaystyle\\leq\\frac\{1\}\{2R^\{2\}\}\\inf\_\{\\varphi\\in S\_\{\[0,1\]\}\}\\left\(\\int\_\{\[0,1\]^\{2\}\}\\left\(\\\|X\(x\)\-Y\(\\varphi\(x\)\)\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{2\}\+\\\|X\(y\)\-Y\(\\varphi\(y\)\)\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{2\}\\right\)dxdy\\right\)=1R2​infφ∈S\[0,1\]\(∫01‖X​\(x\)−Y​\(φ​\(x\)\)‖ℝk2​𝑑x\)\.\\displaystyle=\\frac\{1\}\{R^\{2\}\}\\inf\_\{\\varphi\\in S\_\{\[0,1\]\}\}\\left\(\\int\_\{0\}^\{1\}\\\|X\(x\)\-Y\(\\varphi\(x\)\)\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{2\}dx\\right\)\.For anyh∈O​\(k\)h\\in\\mathrm\{O\}\(k\), we haveWh⋅X=⟨h⋅X​\(x\),h⋅X​\(y\)⟩=⟨X​\(x\),X​\(y\)⟩=WX,W\_\{h\\cdot X\}=\\langle h\\cdot X\(x\),h\\cdot X\(y\)\\rangle=\\langle X\(x\),X\(y\)\\rangle=W\_\{X\},so

δ2​\(WX,WY\)2\\displaystyle\\delta\_\{2\}\(W\_\{X\},W\_\{Y\}\)^\{2\}=δ2​\(Wh⋅X,WY\)2≤1R2​infφ∈S\[0,1\]\(∫01‖h⋅X​\(x\)−Y​\(φ​\(x\)\)‖ℝk2​𝑑x\)\.\\displaystyle=\\delta\_\{2\}\(W\_\{h\\cdot X\},W\_\{Y\}\)^\{2\}\\leq\\frac\{1\}\{R^\{2\}\}\\inf\_\{\\varphi\\in S\_\{\[0,1\]\}\}\\left\(\\int\_\{0\}^\{1\}\\\|h\\cdot X\(x\)\-Y\(\\varphi\(x\)\)\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{2\}dx\\right\)\.Taking the infimum over the orthogonal groupO​\(k\)\\mathrm\{O\}\(k\), the inequality still holds:

δ2​\(WX,WY\)≤1R​infh∈O​\(k\)infφ∈S\[0,1\]\(∫01‖h⋅X​\(x\)−Y​\(φ​\(x\)\)‖ℝk2​𝑑x\)1/2=1R​d¯​\(X,Y\),\\delta\_\{2\}\(W\_\{X\},W\_\{Y\}\)\\leq\\frac\{1\}\{R\}\\inf\_\{h\\in\\mathrm\{O\}\(k\)\}\\inf\_\{\\varphi\\in S\_\{\[0,1\]\}\}\\left\(\\int\_\{0\}^\{1\}\\\|h\\cdot X\(x\)\-Y\(\\varphi\(x\)\)\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{2\}dx\\right\)^\{1/2\}=\\frac\{1\}\{R\}\\overline\{\\mathrm\{d\}\}\(X,Y\),which implies that the mapX↦WXX\\mapsto W\_\{X\}is Lipschitz continuous\. ∎

### E\.4Proof of Theorem[4\.13](https://arxiv.org/html/2605.23156#S4.Thmtheorem13)

The proof relies on the correspondence betweenKR/G∞K\_\{R\}/G\_\{\\infty\}and the orbit space of\(𝒲0,δ2\)\\left\(\{\\mathcal\{W\}\_\{0\}\},\\delta\_\{2\}\\right\)under the action of the orthogonal groupO​\(k\)\\mathrm\{O\}\(k\), together with the universality of the modelIϱ,M,L,bϱI\_\{\\varrho,M,L,b\}^\{\\varrho\}on compact subsets of\(𝒲0,δ2\)\\left\(\{\\mathcal\{W\}\_\{0\}\},\\delta\_\{2\}\\right\)\[[46](https://arxiv.org/html/2605.23156#bib.bib46)\]\. We first state and prove several auxiliary lemmas\.

###### Lemma E\.1\.

ForX,Y∈KRX,Y\\in K\_\{R\}, andWXW\_\{X\},WYW\_\{Y\}are the corresponding covariance functions defined in \([12](https://arxiv.org/html/2605.23156#S4.E12)\), thend¯​\(X,Y\)=0\\overline\{\\mathrm\{d\}\}\(X,Y\)=0if and only ifδ2​\(WX,WY\)=0\\delta\_\{2\}\(W\_\{X\},W\_\{Y\}\)=0\.

This lemma not only ensures that the architecture is well\-defined, but also implies that point separation onKR/G∞K\_\{R\}/G\_\{\\infty\}is equivalent to separate points on\(𝒲0,δ2\)/O​\(k\)\\left\(\{\\mathcal\{W\}\_\{0\}\},\\delta\_\{2\}\\right\)/\\mathrm\{O\}\(k\)\. The proof is built on the following claim that if the covariance functions are equal almost everywhere, then there must exist an orthogonal transformation between the original functions\. More formally,

###### Claim E\.2\.

LetX,Y∈L2​\(\[0,1\],ℝk\)X,Y\\in L^\{2\}\\left\(\[0,1\],\\mathbb\{R\}^\{k\}\\right\)\. Suppose that

⟨X​\(x\),X​\(y\)⟩=⟨Y​\(x\),Y​\(y\)⟩for almost every​x,y∈\[0,1\]\.\\langle X\(x\),X\(y\)\\rangle=\\langle Y\(x\),Y\(y\)\\rangle\\quad\\text\{ for almost every \}x,y\\in\[0,1\]\.Then there exists an orthogonal transformationh∈O​\(k\)h\\in\\mathrm\{O\}\(k\)such thatY=h​XY=hXalmost everywhere\.

###### Proof of Claim[E\.2](https://arxiv.org/html/2605.23156#A5.Thmtheorem2)\.

First, we can identify the spaceL2​\(\[0,1\],ℝk\)L^\{2\}\\left\(\[0,1\],\\mathbb\{R\}^\{k\}\\right\)withℬ​\(L2​\(\[0,1\]\),ℝk\)\\mathcal\{B\}\\left\(L^\{2\}\(\[0,1\]\),\\mathbb\{R\}^\{k\}\\right\), which is the set of bounded linear operators between two Hilbert spacesL2​\(\[0,1\]\)→ℝkL^\{2\}\(\[0,1\]\)\\rightarrow\\mathbb\{R\}^\{k\}, equipped with the Hilbert–Schmidt norm∥⋅∥H​S\\\|\\cdot\\\|\_\{HS\}\. In more detail, eachX∈L2​\(\[0,1\],ℝk\)X\\in L^\{2\}\\left\(\[0,1\],\\mathbb\{R\}^\{k\}\\right\)can be written as

X=\(X1,…,Xk\)⊤withXi∈L2​\(\[0,1\]\),X=\\left\(X\_\{1\},\\ldots,X\_\{k\}\\right\)^\{\\top\}\\quad\\text\{with\}\\quad X\_\{i\}\\in L^\{2\}\(\[0,1\]\),which defines a bounded linear mapΦ:L2​\(\[0,1\]\)→ℝk\\Phi\\colon L^\{2\}\(\[0,1\]\)\\rightarrow\\mathbb\{R\}^\{k\}by

ΦX​φ=\(⟨X1,φ⟩,…,⟨Xk,φ⟩\)⊤\.\\Phi\_\{X\}\\varphi=\\left\(\\left\\langle X\_\{1\},\\varphi\\right\\rangle,\\ldots,\\left\\langle X\_\{k\},\\varphi\\right\\rangle\\right\)^\{\\top\}\.Conversely, by Riesz Representation Theorem any bounded linear operatorΦ∈ℬ​\(L2​\(\[0,1\]\),ℝk\)\\Phi\\in\\mathcal\{B\}\\left\(L^\{2\}\(\[0,1\]\),\\mathbb\{R\}^\{k\}\\right\)can be written as this form for someX1,⋯,Xk∈L2​\(\[0,1\]\)X\_\{1\},\\cdots,X\_\{k\}\\in L^\{2\}\(\[0,1\]\)\.

Define𝒱m≔span⁡\{X1,…,Xk\}\\mathcal\{V\}\_\{m\}\\coloneqq\\operatorname\{span\}\\left\\\{X\_\{1\},\\ldots,X\_\{k\}\\right\\\}\. ThenΦX\\Phi\_\{X\}vanishes on the orthogonal complement𝒱k⟂\\mathcal\{V\}\_\{k\}^\{\\perp\}\. SinceΦX:𝒱k→ℝk\\Phi\_\{X\}\\colon\\mathcal\{V\}\_\{k\}\\rightarrow\\mathbb\{R\}^\{k\}is a linear map between finite\-dimensional vector spaces, it admits a singular value decomposition\[[62](https://arxiv.org/html/2605.23156#bib.bib62)\]\. Thus, there exist non\-negative singular valuesσ1,…,σk∈ℝ≥0\\sigma\_\{1\},\\ldots,\\sigma\_\{k\}\\in\\mathbb\{R\}\_\{\\geq 0\}, orthonormal bases\{v1,…,vk\}⊆ℝk\\left\\\{v\_\{1\},\\ldots,v\_\{k\}\\right\\\}\\subseteq\\mathbb\{R\}^\{k\}, and\{u1,…,uk\}⊆L2​\(\[0,1\]\)\\left\\\{u\_\{1\},\\ldots,u\_\{k\}\\right\\\}\\subseteq L^\{2\}\(\[0,1\]\)such that

ΦX=∑i=1kσi​⟨ui,⋅⟩​vi\.\\Phi\_\{X\}=\\sum\_\{i=1\}^\{k\}\\sigma\_\{i\}\\left\\langle u\_\{i\},\\cdot\\right\\rangle v\_\{i\}\.Moreover, for eachσi\>0\\sigma\_\{i\}\>0, theuiu\_\{i\}is an eigenvector of the self\-adjoint operatorΦX∗​ΦX\\Phi\_\{X\}^\{\*\}\\Phi\_\{X\}corresponding to the eigenvalueσi2\\sigma\_\{i\}^\{2\}\. For indicesiiwithσi=0\\sigma\_\{i\}=0, the vectorsuiu\_\{i\}can be chosen to complete the set\{u1,…,uk\}\\left\\\{u\_\{1\},\\ldots,u\_\{k\}\\right\\\}into an orthonormal basis of𝒱k\\mathcal\{V\}\_\{k\}\.

The self\-adjoint operatorΦX∗​ΦX\\Phi\_\{X\}^\{\*\}\\Phi\_\{X\}evaluated onφ∈L2​\(\[0,1\]\)\\varphi\\in L^\{2\}\(\[0,1\]\)yields

\(ΦX∗​ΦX​φ\)​\(x\)=∑i=1k⟨Xi,φ⟩​Xi​\(x\)=∑i=1kXi​\(x\)​∫01Xi​\(y\)​φ​\(y\)​𝑑y=∫01⟨X​\(x\),X​\(y\)⟩​φ​\(y\)​𝑑y\.\(\\Phi\_\{X\}^\{\*\}\\Phi\_\{X\}\\varphi\)\(x\)=\\sum\_\{i=1\}^\{k\}\\left\\langle X\_\{i\},\\varphi\\right\\rangle X\_\{i\}\(x\)=\\sum\_\{i=1\}^\{k\}X\_\{i\}\(x\)\\int\_\{0\}^\{1\}X\_\{i\}\(y\)\\varphi\(y\)dy=\\int\_\{0\}^\{1\}\\left\\langle X\(x\),X\(y\)\\right\\rangle\\varphi\(y\)dy\.Since⟨X​\(x\),X​\(y\)⟩=⟨Y​\(x\),Y​\(y\)⟩\\langle X\(x\),X\(y\)\\rangle=\\langle Y\(x\),Y\(y\)\\ranglealmost everywhere, thenΦX∗​ΦX=ΦY∗​ΦY\\Phi\_\{X\}^\{\*\}\\Phi\_\{X\}=\\Phi\_\{Y\}^\{\*\}\\Phi\_\{Y\}\. Therefore, we can choose the same orthonormal bases\{u1,…,uk\}⊆L2​\(\[0,1\]\)\\left\\\{u\_\{1\},\\ldots,u\_\{k\}\\right\\\}\\subseteq L^\{2\}\(\[0,1\]\)forΦX\\Phi\_\{X\}andΦY\\Phi\_\{Y\}, such that

ΦX=∑i=1kσi​⟨ui,⋅⟩​vi,ΦY=∑i=1kσi​⟨ui,⋅⟩​wi,\\Phi\_\{X\}=\\sum\_\{i=1\}^\{k\}\\sigma\_\{i\}\\left\\langle u\_\{i\},\\cdot\\right\\rangle v\_\{i\},\\quad\\Phi\_\{Y\}=\\sum\_\{i=1\}^\{k\}\\sigma\_\{i\}\\left\\langle u\_\{i\},\\cdot\\right\\rangle w\_\{i\},where\{v1,…,vk\}\\left\\\{v\_\{1\},\\ldots,v\_\{k\}\\right\\\}and\{w1,…,wk\}\\left\\\{w\_\{1\},\\ldots,w\_\{k\}\\right\\\}are both the orthonormal bases inℝk\\mathbb\{R\}^\{k\}\. Then there existsh∈O​\(k\)h\\in\\mathrm\{O\}\(k\)such thath​\(vi\)=wih\(v\_\{i\}\)=w\_\{i\}, which implies

h​ΦX=∑i=1kσi​⟨ui,⋅⟩​\(h​vi\)=∑i=1kσi​⟨ui,⋅⟩​wi=ΦY\.h\\Phi\_\{X\}=\\sum\_\{i=1\}^\{k\}\\sigma\_\{i\}\\left\\langle u\_\{i\},\\cdot\\right\\rangle\(hv\_\{i\}\)=\\sum\_\{i=1\}^\{k\}\\sigma\_\{i\}\\left\\langle u\_\{i\},\\cdot\\right\\rangle w\_\{i\}=\\Phi\_\{Y\}\.For anyω∈ℝk,\\omega\\in\\mathbb\{R\}^\{k\},⟨Y​\(⋅\),ω⟩=ΦY∗​ω=\(h​ΦX\)∗​ω=⟨h​X​\(⋅\),ω⟩,\\langle Y\(\\cdot\),\\omega\\rangle=\\Phi\_\{Y\}^\{\*\}\\omega=\\left\(h\\Phi\_\{X\}\\right\)^\{\*\}\\omega=\\langle hX\(\\cdot\),\\omega\\rangle,which yieldsY=h​XY=hXinL2​\(\[0,1\],ℝk\)L^\{2\}\(\[0,1\],\\mathbb\{R\}^\{k\}\)\. ∎

###### Proof of Lemma[E\.1](https://arxiv.org/html/2605.23156#A5.Thmtheorem1)\.

First, by Theorem[4\.12](https://arxiv.org/html/2605.23156#S4.Thmtheorem12), ifd¯​\(X,Y\)=0\\overline\{\\mathrm\{d\}\}\(X,Y\)=0, thenδ2​\(WX,WY\)=0\\delta\_\{2\}\(W\_\{X\},W\_\{Y\}\)=0\. We only need to show thatδ2​\(WX,WY\)=0\\delta\_\{2\}\(W\_\{X\},W\_\{Y\}\)=0impliesd¯​\(X,Y\)=0\\overline\{\\mathrm\{d\}\}\(X,Y\)=0\. Sinceδ2​\(WX,WY\)=0\\delta\_\{2\}\\left\(W\_\{X\},W\_\{Y\}\\right\)=0, there exists measure\-preserving mapsφ,ψ∈S¯\[0,1\]\\varphi,\\psi\\in\\bar\{S\}\_\{\[0,1\]\}, such thatWXφ=WYψW\_\{X\}^\{\\varphi\}=W^\{\\psi\}\_\{Y\}almost everywhere\[[50](https://arxiv.org/html/2605.23156#bib.bib50), Cor\. 10\.35\], hence

⟨X​\(φ​\(x\)\),X​\(φ​\(y\)\)⟩=⟨Y​\(ψ​\(x\)\),Y​\(ψ​\(y\)\)⟩for almost every​x,y∈\[0,1\]\.\\langle X\(\\varphi\(x\)\),X\(\\varphi\(y\)\)\\rangle=\\langle Y\(\\psi\(x\)\),Y\(\\psi\(y\)\)\\rangle\\quad\\text\{ for almost every \}x,y\\in\[0,1\]\.By Claim[E\.2](https://arxiv.org/html/2605.23156#A5.Thmtheorem2), there existsh∈O​\(k\)h\\in\\mathrm\{O\}\(k\)such thath​Xφ=Yψ,hX^\{\\varphi\}=Y^\{\\psi\},whereXφ​\(x\)=X​\(φ​\(x\)\)X^\{\\varphi\}\(x\)=X\(\\varphi\(x\)\)\. Therefore,

d¯​\(X,Y\)\\displaystyle\\overline\{\\mathrm\{d\}\}\(X,Y\)=infh∈O​\(k\)infφ∈S\[0,1\]\(∫01‖h​\(X​\(t\)\)−Y​\(φ​\(t\)\)‖ℝk2​𝑑t\)1/2\\displaystyle=\\inf\_\{h\\in\\mathrm\{O\}\(k\)\}\\inf\_\{\\varphi\\in S\_\{\[0,1\]\}\}\\left\(\\int\_\{0\}^\{1\}\\\|h\(X\(t\)\)\-Y\(\\varphi\(t\)\)\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{2\}dt\\right\)^\{1/2\}=infh∈O​\(k\)infφ,ψ∈S¯\[0,1\]\(∫01‖h​\(X​\(φ​\(t\)\)\)−Y​\(ψ​\(t\)\)‖ℝk2​𝑑t\)1/2\\displaystyle=\\inf\_\{h\\in\\mathrm\{O\}\(k\)\}\\inf\_\{\\varphi,\\psi\\in\\overline\{S\}\_\{\[0,1\]\}\}\\left\(\\int\_\{0\}^\{1\}\\\|h\(X\(\\varphi\(t\)\)\)\-Y\(\\psi\(t\)\)\\\|\_\{\\mathbb\{R\}^\{k\}\}^\{2\}dt\\right\)^\{1/2\}≤0,\\displaystyle\\leq 0,whereS¯\[0,1\]\\overline\{S\}\_\{\[0,1\]\}denotes the set of measure\-preserving maps\[0,1\]→\[0,1\]\[0,1\]\\rightarrow\[0,1\]\. This proves thatd¯​\(X,Y\)=0\\overline\{\\mathrm\{d\}\}\(X,Y\)=0, as desired\. ∎

###### Proof of Theorem[4\.13](https://arxiv.org/html/2605.23156#S4.Thmtheorem13)\.

We define the class of functionsℱh\\mathcal\{F\}\_\{h\}as

ℱh≔span⁡\{X↦t​\(F,WX\)∣F​is simple\},\\mathcal\{F\}\_\{h\}\\coloneqq\\operatorname\{span\}\\\{X\\mapsto t\(F,W\_\{X\}\)\\mid F\\text\{ is simple \}\\\},wherettdenotes the homomorphism density defined in \([8](https://arxiv.org/html/2605.23156#S4.E8)\)\. It is straightforward to see thatℱh\\mathcal\{F\}\_\{h\}is a subalgebra, sincet​\(F1,⋅\)​t​\(F2,⋅\)=t​\(F1⊔F2,⋅\)t\\left\(F\_\{1\},\\cdot\\right\)t\\left\(F\_\{2\},\\cdot\\right\)=t\\left\(F\_\{1\}\\sqcup F\_\{2\},\\cdot\\right\)\. By Lemmas[E\.1](https://arxiv.org/html/2605.23156#A5.Thmtheorem1)and\[[50](https://arxiv.org/html/2605.23156#bib.bib50), Cor\. 10\.34\],ℱh\\mathcal\{F\}\_\{h\}separates points\. Therefore,ℱh\\mathcal\{F\}\_\{h\}is dense inC​\(KR/G∞\)C\\left\(K\_\{R\}/G\_\{\\infty\}\\right\)\.

By Theorem[4\.12](https://arxiv.org/html/2605.23156#S4.Thmtheorem12),X↦WXX\\mapsto W\_\{X\}is a continuous mapping\. Since the inputKR/G∞K\_\{R\}/G\_\{\\infty\}is compact by Proposition[4\.11](https://arxiv.org/html/2605.23156#S4.Thmtheorem11), the image ofKR/G∞K\_\{R\}/G\_\{\\infty\}is a compact subset of\(𝒲0,δ2\)\\left\(\{\\mathcal\{W\}\_\{0\}\},\\delta\_\{2\}\\right\)\. The IGN model is universal on compact subsets of\(𝒲0,δ2\)\\left\(\{\\mathcal\{W\}\_\{0\}\},\\delta\_\{2\}\\right\), and thus can approximate arbitrarily well any homomorphism density of simple graphs\[[46](https://arxiv.org/html/2605.23156#bib.bib46)\]\. Consequently,ℱP​Cϱ\\mathcal\{F\}\_\{PC\}^\{\\varrho\}is dense inC​\(KR/G∞\)C\\left\(K\_\{R\}/G\_\{\\infty\}\\right\)\. ∎

Similar Articles

Unified Neural Scaling Laws

Hugging Face Daily Papers

Presents a unified neural scaling law that accurately models deep neural network scaling across multiple dimensions including parameters, dataset size, training steps, and compute, validated across diverse architectures and tasks.

Universality of Gradient Descent Neural Network Training

Hacker News Top

The paper explores whether any neural network can be redesigned to train effectively with gradient descent, proving a universality result that for any network, there exists an extension that reproduces given weights and outputs via gradient descent.