Equivariant Sheaf Neural Networks: Learning Geometric Transport on Graphs
Summary
This paper introduces Equivariant Sheaf Neural Networks (ESNN), which learn directed geometric transport on graphs to enhance equivariant message passing for physical systems while preserving exact Euclidean symmetry.
View Cached Full Text
Cached at: 09/01/26, 12:57 PM
# Equivariant Sheaf Neural Networks: Learning Geometric Transport on Graphs
Source: [https://arxiv.org/html/2608.28853](https://arxiv.org/html/2608.28853)
Alessio BorgiAffiliation:Department of Computer Science and Technology, University of CambridgeAffiliation:Department of Computer, Control and Management Engineering, Sapienza University of RomeMario Severino11footnotemark:1Affiliation:Department of Computer Science and Technology, University of CambridgeAffiliation:Department of Information Engineering, University of PaduaFabrizio SilvestriAffiliation:Department of Computer, Control and Management Engineering, Sapienza University of Rome
###### Abstract
Equivariant graph neural networks provide a principled way to model geometric systems, but efficient first\-order architectures remain limited in how vector information can be transformed as it moves across a graph\. We introduceESNN, an Equivariant Sheaf Neural Network that enriches this interaction by learning directed, matrix\-valued transport between neighboring vector features while preserving exact Euclidean equivariance\. Rather than increasing the order of the representation, ESNN keeps scalar and vector features first\-order and places the additional geometric flexibility in the edge transport itself\. We characterize this transport theoretically, showing that when relative displacement is the only covariant geometric input, every linearO\(n\)O\(n\)\-equivariant map decomposes into independent radial and tangential components, while learned covariant features enable richer feature\-conditioned transformations\. We also introduce controlled symmetry relaxation for systems with a preferred ambient direction, which may be prescribed or inferred from data while recovering fullE\(n\)E\(n\)\-equivariance when the directional pathway is inactive\. Across particle dynamics, mesh\-based simulation, point\-cloud classification, and molecular property prediction, ESNN improves dynamics prediction, recovers the gravity axis when symmetry is broken, yields substantial gains on selected mesh tasks and long\-horizon rollouts, and remains robust to unseen rotations\. These results show that learning how geometric information is transported across edges offers a complementary route to expressive equivariant message passing without requiring higher\-order representations\.
## 1Introduction
Many physical systems, from molecular dynamics to fluid mechanics, are naturally embedded in annn\-dimensional Euclidean space and governed by the symmetries of the Euclidean groupE\(n\)E\(n\)\([Satorras et al\., 2021](https://arxiv.org/html/2608.28853#bib.bib1);[Brandstetter et al\., 2022](https://arxiv.org/html/2608.28853#bib.bib8);[Du et al\., 2023](https://arxiv.org/html/2608.28853#bib.bib2);[Wang et al\., 2024b](https://arxiv.org/html/2608.28853#bib.bib4);[Aykent and Xia, 2025](https://arxiv.org/html/2608.28853#bib.bib6)\)\. When represented as geometric graphs, their node states may combine invariant scalar attributes with vector or tensor features that transform with the geometry\([Schütt et al\., 2021](https://arxiv.org/html/2608.28853#bib.bib7);[Simeon and de Fabritiis, 2023](https://arxiv.org/html/2608.28853#bib.bib3);[Liao et al\., 2024](https://arxiv.org/html/2608.28853#bib.bib5)\)\. A physically consistent model should therefore respond predictably to translations, rotations, and, when appropriate, reflections\. Equivariance constrains this transformation behavior, but it does not by itself determine how geometric information should be exchanged between neighboring nodes\. In many physical systems, such interactions are inherently direction\-dependent: longitudinal and transverse responses, shear, and anisotropic propagation depend on relative orientation rather than on pairwise distance alone\. Geometric message passing should therefore capture this directional structure while preserving exactE\(n\)E\(n\)\-equivariance\.
EGNNSNNESNNedge messagej→ij\\\!\\to\\\!i𝐫ij\\mathbf\{r\}\_\{ij\}iijjϕ\(∥𝐫ij∥\)\\phi\(\\lVert\\mathbf\{r\}\_\{ij\}\\rVert\)isotropic scalar transportSchematic spatial motifs:compression✓\\checkmarkshear×\\timestorsion×\\timesequivariance testinput𝐱\\mathbf\{x\}iif\(𝐱\)f\(\\mathbf\{x\}\)𝐐\\mathbf\{Q\}rotated𝐐𝐱\\mathbf\{Q\}\\mathbf\{x\}iiexp\.𝐐f\(𝐱\)\\mathbf\{Q\}f\(\\mathbf\{x\}\)✓\\checkmarkf\(𝐐𝐱\)=𝐐f\(𝐱\)f\(\\mathbf\{Q\}\\mathbf\{x\}\)=\\mathbf\{Q\}f\(\\mathbf\{x\}\)equivariant✓\\checkmarkanisotropic×\\timessheaf messagej→ij\\\!\\to\\\!i𝐫ij\\mathbf\{r\}\_\{ij\}ℱi\\mathcal\{F\}\_\{i\}ℱj\\mathcal\{F\}\_\{j\}=𝐖ij=\\mathbf\{W\}\_\{ij\}unconstrained coupling mapSchematic spatial motifs:compression✓\\checkmarkshear✓\\checkmarktorsion✓\\checkmarkequivariance testinput𝐱\\mathbf\{x\}iif\(𝐱\)f\(\\mathbf\{x\}\)𝐐\\mathbf\{Q\}rotated𝐐𝐱\\mathbf\{Q\}\\mathbf\{x\}iiexp\.𝐐f\(𝐱\)\\mathbf\{Q\}f\(\\mathbf\{x\}\)f\(𝐐𝐱\)f\(\\mathbf\{Q\}\\mathbf\{x\}\)≠\\neq×\\timesf\(𝐐𝐱\)≠𝐐f\(𝐱\)f\(\\mathbf\{Q\}\\mathbf\{x\}\)\\neq\\mathbf\{Q\}f\(\\mathbf\{x\}\)equivariant×\\timesanisotropic✓\\checkmarksheaf messagej→ij\\\!\\to\\\!i𝐫ij\\mathbf\{r\}\_\{ij\}ℱi\\mathcal\{F\}\_\{i\}ℱj\\mathcal\{F\}\_\{j\}⋅\\cdot𝐒ij\\mathbf\{S\}\_\{ij\}spatial𝐌ij\\mathbf\{M\}\_\{ij\}channelsSchematic spatial motifs:compression✓\\checkmarkshear✓\\checkmarktorsion✓\\checkmarkequivariance testinput𝐱\\mathbf\{x\}iif\(𝐱\)f\(\\mathbf\{x\}\)𝐐\\mathbf\{Q\}rotated𝐐𝐱\\mathbf\{Q\}\\mathbf\{x\}iiexp\.𝐐f\(𝐱\)\\mathbf\{Q\}f\(\\mathbf\{x\}\)✓\\checkmarkf\(𝐐𝐱\)=𝐐f\(𝐱\)f\(\\mathbf\{Q\}\\mathbf\{x\}\)=\\mathbf\{Q\}f\(\\mathbf\{x\}\)equivariant✓\\checkmarkanisotropic✓\\checkmark
Figure 1:*Edge\-wise geometric transport in EGNN\-style message passing, generic sheaf neural networks, and ESNNs\.*EGNN\-style models construct geometric interactions from invariant scalar coefficients and relative displacements, while generic sheaf networks allow matrix\-valued transformations between local feature spaces without necessarily respecting the transformation law of geometric vectors\. ESNN combines matrix\-valued transport with exact equivariance by separating anO\(n\)O\(n\)\-covariant spatial action from invariant channel mixing\. The deformation motifs illustrate edge\-wise spatial actions and are not intended as claims about the full expressive capacity of each architecture\.Existing equivariant architectures capture directional structure through different mechanisms\. Cartesian models keep the representation space simple, working with invariant scalars and first\-order vectors\([Satorras et al\., 2021](https://arxiv.org/html/2608.28853#bib.bib1);[Schütt et al\., 2021](https://arxiv.org/html/2608.28853#bib.bib7)\), while steerable architectures increase the richness of the propagated representations, using higher\-order features, spherical harmonics, and tensor products to capture more detailed angular structure\([Thomas et al\., 2018](https://arxiv.org/html/2608.28853#bib.bib9);[Batatia et al\., 2022a](https://arxiv.org/html/2608.28853#bib.bib10);[Batatia et al\., 2022b](https://arxiv.org/html/2608.28853#bib.bib11)\)\. A complementary possibility is to keep the feature representation first\-order and instead enrich the way information is transformed as it passes between neighboring nodes\. This is precisely the perspective suggested by Sheaf Neural Networks \(SNNs\), where interactions between neighboring nodes are mediated by learned linear maps between local feature spaces\([Hansen and Gebhart, 2020](https://arxiv.org/html/2608.28853#bib.bib17);[Bodnar et al\., 2022](https://arxiv.org/html/2608.28853#bib.bib16)\)\. These maps provide matrix\-valued edge interactions, offering a natural mechanism for richer vector transport, but generic sheaf parameterizations do not account for the Euclidean transformation laws of geometric vector features\.
To bridge this gap, we propose*Equivariant Sheaf Neural Networks \(ESNN\)*, a family of equivariant graph neural networks that learn matrix\-valued transport between neighboring vector features while preserving exactE\(n\)E\(n\)\-equivariance\. ESNN factorizes each transport into anO\(n\)O\(n\)\-covariant111A matrix\-valued spatial transport functionSijS\_\{ij\}isO\(n\)O\(n\)\-covariant if, for everyQ∈O\(n\)Q\\in O\(n\), transforming its geometric inputs byQQinducesSij′=QSijQ⊤S^\{\\prime\}\_\{ij\}=QS\_\{ij\}Q^\{\\top\}\.spatial action and invariant mixing across feature channels\. This makes it possible to model anisotropic interactions while remaining entirely within first\-order scalar and vector representations\. A particularly simple instantiation distinguishes the component of a vector message along the relative displacement from the component orthogonal to it, and transforms the two independently\. We show that, when the relative displacement is the only covariant geometric input, this radial–tangential decomposition characterizes the full class of linearO\(n\)O\(n\)\-equivariant transports\.
ESNN further accommodates systems in which external structure, such as gravity or background flow, selects a preferred direction and reduces the symmetry of the dynamics\. This direction may be fixed from prior knowledge or learned as a global model parameter, while a learnable relaxation coefficient controls how strongly it affects the transport\. When the coefficient is zero, the model is exactlyE\(n\)E\(n\)\-equivariant; when the directional signal is used, the symmetry is reduced to the subgroup that preserves the preferred direction\.
Contributions\.Our main contributions are:
- •We introduce*Equivariant Sheaf Neural Networks \(ESNN\)*, a first\-order Cartesian framework that learns matrix\-valued transport of vector features across edges while preserving exactE\(n\)E\(n\)\-equivariance\.
- •We develop a family of equivariant transport mechanisms with increasing geometric flexibility, from identity and isotropic transformations to feature\-dependent rotations and anisotropic radial–tangential transport\. We show that, when relative displacement is the only covariant geometric input, the radial–tangential form captures the complete class of linearO\(n\)O\(n\)\-equivariant transports\.
- •We introduce controlled symmetry relaxation for systems with a preferred ambient direction, which may be prescribed or learned from data\. A learnable coefficient controls its influence: at zero, fullE\(n\)E\(n\)\-equivariance is recovered exactly; when active, equivariance is retained to the subgroup that preserves the preferred direction\.
- •We evaluate ESNN across particle dynamics, mesh\-based physical simulation, point\-cloud classification, and molecular property prediction, testing richer transport under full symmetry, data\-driven recovery of a symmetry\-breaking direction, and transfer across distinct geometric domains\.
## 2Related Work and Background
Related Work\.ESNN lies at the intersection of equivariant message passing and sheaf\-based graph learning, combining the geometric inductive biases of the former with the matrix\-valued transport perspective of the latter\. Cartesian architectures such as EGNN\([Satorras et al\., 2021](https://arxiv.org/html/2608.28853#bib.bib1)\)and scalar–vector models such as PaiNN and GVP\([Schütt et al\., 2021](https://arxiv.org/html/2608.28853#bib.bib7);[Jing et al\., 2021](https://arxiv.org/html/2608.28853#bib.bib18)\)achieve efficient Euclidean equivariance using invariant scalars and low\-order covariant features\. A complementary line of work develops steerable models, which represent features according to how they transform under rotations and build equivariant interactions by combining these representation types in symmetry\-preserving ways\([Thomas et al\., 2018](https://arxiv.org/html/2608.28853#bib.bib9);[Fuchs et al\., 2020](https://arxiv.org/html/2608.28853#bib.bib19);[Brandstetter et al\., 2022](https://arxiv.org/html/2608.28853#bib.bib8);[Batatia et al\., 2022b](https://arxiv.org/html/2608.28853#bib.bib11);[Liao et al\., 2024](https://arxiv.org/html/2608.28853#bib.bib5)\)\. ESNN retains the first\-order Cartesian setting, but increase the expressivity of the edge interaction by learning how neighboring vector features are transformed before aggregation\. This perspective connects naturally to SNNs, which generalize scalar edge weighting to learned linear maps between local feature spaces\([Hansen and Gebhart, 2020](https://arxiv.org/html/2608.28853#bib.bib17);[Bodnar et al\., 2022](https://arxiv.org/html/2608.28853#bib.bib16)\)\. Connection\-based sheaf models have considered orthogonal geometric maps\([Barbero et al\., 2022a](https://arxiv.org/html/2608.28853#bib.bib38)\), while recent work has extended sheaf\-based propagation to directional and asymmetric interactions\([Ribeiro et al\., 2025](https://arxiv.org/html/2608.28853#bib.bib44);[Fiorini et al\., 2025](https://arxiv.org/html/2608.28853#bib.bib45)\)\. Copresheaf formulations provide a related view in which information is propagated through directed maps between local feature spaces\([Hajij et al\., 2025](https://arxiv.org/html/2608.28853#bib.bib46)\)\. ESNN builds on this local, matrix\-valued perspective while imposing the Euclidean transformation laws required by geometric vector features\. ESNN also allow this symmetry prior to be relaxed when the environment has a preferred direction\. Related subequivariant models enforce the subgroup associated with a prescribed field\([Han et al\., 2022](https://arxiv.org/html/2608.28853#bib.bib27)\), while relaxed\-equivariant methods learn departures from an underlying symmetry\([Wang et al\., 2022a](https://arxiv.org/html/2608.28853#bib.bib50);[Hofgard et al\., 2024](https://arxiv.org/html/2608.28853#bib.bib52)\)\. In ESNN, this additional degree of freedom is incorporated through a learnable preferred\-direction signal, with exactE\(n\)E\(n\)\-equivariance recovered when the signal is inactive\. A broader discussion of these connections and related geometric and sheaf\-based approaches is provided in Appendix[A](https://arxiv.org/html/2608.28853#A1)\.
### 2\.1Geometric Background
We consider a physical system represented by a graph𝒢=\(𝒱,ℰ\)\\mathcal\{G\}=\(\\mathcal\{V\},\\mathcal\{E\}\), where each nodei∈𝒱i\\in\\mathcal\{V\}has spatial coordinates𝐱i∈ℝn\\mathbf\{x\}\_\{i\}\\in\\mathbb\{R\}^\{n\}\. Our goal is to construct message\-passing operators that respect Euclidean symmetries while allowing interactions between neighboring nodes to depend on their relative geometry\. We first specify the relevant transformation laws and feature representation, and then introduce the cellular\-sheaf construction that motivates the transport formulation of Section[3](https://arxiv.org/html/2608.28853#S3)\.
Equivariance\.The Euclidean groupE\(n\)=O\(n\)⋉ℝnE\(n\)=O\(n\)\\ltimes\\mathbb\{R\}^\{n\}acts on the coordinates through rigid motions:
𝐱i↦Q𝐱i\+𝐭,Q∈O\(n\),𝐭∈ℝn\\mathbf\{x\}\_\{i\}\\mapsto Q\\mathbf\{x\}\_\{i\}\+\\mathbf\{t\},\\qquad Q\\in O\(n\),\\quad\\mathbf\{t\}\\in\\mathbb\{R\}^\{n\}\(1\)AnE\(n\)E\(n\)\-equivariant layer transforms its outputs consistently with this action\. Graph interactions are expressed through relative displacements𝐫ij=𝐱i−𝐱j\\mathbf\{r\}\_\{ij\}=\\mathbf\{x\}\_\{i\}\-\\mathbf\{x\}\_\{j\}for which translations cancel, and orthogonal transformations act as𝐫ij↦Q𝐫ij\\mathbf\{r\}\_\{ij\}\\mapsto Q\\mathbf\{r\}\_\{ij\}\. For operations constructed from relative geometry, translation equivariance is therefore built into the representation, while the remaining geometric requirement is to enforce the appropriateO\(n\)O\(n\)transformation laws\.
Feature Representation\.Each node carries invariant scalar features and covariant vector features:
𝐡i=\(𝐬i,𝐕i\)\\mathbf\{h\}\_\{i\}=\(\\mathbf\{s\}\_\{i\},\\mathbf\{V\}\_\{i\}\)\(2\)where𝐬i∈ℝcs\\mathbf\{s\}\_\{i\}\\in\\mathbb\{R\}^\{c\_\{s\}\}contains invariant scalar channels and𝐕i∈ℝn×cv\\mathbf\{V\}\_\{i\}\\in\\mathbb\{R\}^\{n\\times c\_\{v\}\}containscvc\_\{v\}vector channels\. Each column of𝐕i\\mathbf\{V\}\_\{i\}is annn\-dimensional vector that transforms as:
𝐬i↦𝐬i,𝐕i↦Q𝐕i\\mathbf\{s\}\_\{i\}\\mapsto\\mathbf\{s\}\_\{i\},\\qquad\\mathbf\{V\}\_\{i\}\\mapsto Q\\mathbf\{V\}\_\{i\}\(3\)Accordingly, the corresponding node feature space isℱ\(i\)=ℝcs⊕\(ℝn⊗ℝcv\)\\mathcal\{F\}\(i\)=\\mathbb\{R\}^\{c\_\{s\}\}\\oplus\\left\(\\mathbb\{R\}^\{n\}\\otimes\\mathbb\{R\}^\{c\_\{v\}\}\\right\)\. In the sheaf interpretation below,ℱ\(i\)\\mathcal\{F\}\(i\)plays the role of the node*stalk*, namely the local vector space attached to nodeii\. ESNN applies geometric transport to its vector component, while the scalar component is propagated through invariant message\-passing operations\.
Cellular Sheaves and Transport\.Standard graph message passing implicitly treats neighboring features as elements of a common space that can be compared and aggregated directly\. Cellular sheaves generalize this picture by assigning local vector spaces to nodes and edges and relating them through linear maps\([Hansen and Gebhart, 2020](https://arxiv.org/html/2608.28853#bib.bib17);[Bodnar et al\., 2022](https://arxiv.org/html/2608.28853#bib.bib16)\)\. For an edgee=\{i,j\}e=\\\{i,j\\\}, a cellular sheaf assigns node stalksℱ\(i\)\\mathcal\{F\}\(i\)andℱ\(j\)\\mathcal\{F\}\(j\), an edge stalkℱ\(e\)\\mathcal\{F\}\(e\), and restriction maps:
ρi→e:ℱ\(i\)→ℱ\(e\),ρj→e:ℱ\(j\)→ℱ\(e\)\\rho\_\{i\\to e\}:\\mathcal\{F\}\(i\)\\rightarrow\\mathcal\{F\}\(e\),\\qquad\\rho\_\{j\\to e\}:\\mathcal\{F\}\(j\)\\rightarrow\\mathcal\{F\}\(e\)\(4\)These maps express the node features in a common edge space\. After choosing an orientation foree, the degree\-00coboundary measures their disagreement:
\(δℱ𝐡\)e=ρi→e𝐡i−ρj→e𝐡j\(\\delta\_\{\\mathcal\{F\}\}\\mathbf\{h\}\)\_\{e\}=\\rho\_\{i\\to e\}\\mathbf\{h\}\_\{i\}\-\\rho\_\{j\\to e\}\\mathbf\{h\}\_\{j\}\(5\)With standard Euclidean inner products, these restriction maps define the sheaf Laplacian𝐋ℱ=δℱ∗δℱ\\mathbf\{L\}\_\{\\mathcal\{F\}\}=\\delta\_\{\\mathcal\{F\}\}^\{\*\}\\delta\_\{\\mathcal\{F\}\}, which reduces to the ordinary graph Laplacian when all stalks share the same feature space and the restriction maps are identities\. For an edgee=\{i,j\}e=\\\{i,j\\\}, its off\-diagonal coupling is, up to the conventional minus sign:
ρi→e∗ρj→e:ℱ\(j\)→ℱ\(i\)\\rho\_\{i\\to e\}^\{\*\}\\rho\_\{j\\to e\}:\\mathcal\{F\}\(j\)\\rightarrow\\mathcal\{F\}\(i\)\(6\)
This node\-to\-node coupling is the point of departure for ESNN: instead of learning separate incidence maps through an intermediate edge stalk, ESNN parameterizes the transport between neighboring nodes directly\. For geometric vector features, this transport must additionally respect rotations and reflections, leading to theO\(n\)O\(n\)\-equivariant construction developed in Section[3](https://arxiv.org/html/2608.28853#S3)\. In this sense, ESNN retains the sheaf perspective of local feature spaces connected by linear compatibility maps, while parameterizing the effective node\-to\-node transport directly; Appendix[C](https://arxiv.org/html/2608.28853#A3)gives the precise relationship with classical sheaf diffusion\.
## 3Equivariant Spatial Transport
The vector features introduced in Section[2\.1](https://arxiv.org/html/2608.28853#S2.SS1)have two distinct axes: the spatial dimensionℝn\\mathbb\{R\}^\{n\}, which carries theO\(n\)O\(n\)action, and the channel dimensionℝcv\\mathbb\{R\}^\{c\_\{v\}\}, on which the group acts trivially\. This distinction naturally separates the geometric transformation of a vector message from the mixing of its feature channels\. ESNN therefore represents edge transport as a finite sum of left–right actions: covariant spatial operators act on the spatial axis, while invariant channel operators act on the channel axis\. This allows the vector representation of a neighboring node to be transformed before aggregation without breaking equivariance\.
Let𝒞ij\\mathcal\{C\}\_\{ij\}denote the local edge context used to construct the transport for the interactionj→ij\\rightarrow i\. It collects the geometric quantities available at that edge, including the relative displacement𝐫ij\\mathbf\{r\}\_\{ij\}and, for feature\-dependent transports, the covariant vector features at its endpoints\. We writeQ⋅𝒞ijQ\\\!\\cdot\\\!\\mathcal\{C\}\_\{ij\}for the simultaneous action ofQQon all covariant quantities in this context, and𝐳ij\\mathbf\{z\}\_\{ij\}for the invariant features derived from it\.
###### Definition 3\.1\(*Equivariant Transport Map*\)\.
For a directed interactionj→ij\\rightarrow iand a fixed local edge context, ESNN defines a transport map𝒯i←j:ℝn×cv⟶ℝn×cv\\mathcal\{T\}\_\{i\\leftarrow j\}:\\mathbb\{R\}^\{n\\times c\_\{v\}\}\\longrightarrow\\mathbb\{R\}^\{n\\times c\_\{v\}\}of the form:
𝒯i←j\(𝐕j\)=∑k=1K𝐒ij\(k\)𝐕j𝐌ij\(k\)\\mathcal\{T\}\_\{i\\leftarrow j\}\(\\mathbf\{V\}\_\{j\}\)=\\sum\_\{k=1\}^\{K\}\\mathbf\{S\}\_\{ij\}^\{\(k\)\}\\mathbf\{V\}\_\{j\}\\mathbf\{M\}\_\{ij\}^\{\(k\)\}\(7\)whereKKdenotes the number of components in the transport expansion\. Each term combines two actions:
- •The*spatial operator*𝐒ij\(k\)∈ℝn×n\\mathbf\{S\}\_\{ij\}^\{\(k\)\}\\in\\mathbb\{R\}^\{n\\times n\}acts on the spatial axis and transforms covariantly: 𝐒ij\(k\)\(Q⋅𝒞ij\)=Q𝐒ij\(k\)\(𝒞ij\)Q⊤∀Q∈O\(n\)\\mathbf\{S\}\_\{ij\}^\{\(k\)\}\(Q\\\!\\cdot\\\!\\mathcal\{C\}\_\{ij\}\)=Q\\,\\mathbf\{S\}\_\{ij\}^\{\(k\)\}\(\\mathcal\{C\}\_\{ij\}\)Q^\{\\top\}\\qquad\\forall Q\\in O\(n\)\(8\)
- •The*channel operator*𝐌ij\(k\)∈ℝcv×cv\\mathbf\{M\}\_\{ij\}^\{\(k\)\}\\in\\mathbb\{R\}^\{c\_\{v\}\\times c\_\{v\}\}mixes vector channels and is constructed fromO\(n\)O\(n\)\-invariant edge features: 𝐌ij\(k\)\(Q⋅𝒞ij\)=𝐌ij\(k\)\(𝒞ij\)∀Q∈O\(n\)\\mathbf\{M\}\_\{ij\}^\{\(k\)\}\(Q\\\!\\cdot\\\!\\mathcal\{C\}\_\{ij\}\)=\\mathbf\{M\}\_\{ij\}^\{\(k\)\}\(\\mathcal\{C\}\_\{ij\}\)\\qquad\\forall Q\\in O\(n\)\(9\)
These two transformation laws are sufficient to make the resulting transport equivariant\.
###### Proposition 3\.2\(Equivariance of the Transport Map\)\.
Suppose that, for every componentkk, the spatial operator satisfies:
𝐒ij\(k\)\(Q⋅𝒞ij\)=Q𝐒ij\(k\)\(𝒞ij\)Q⊤\\mathbf\{S\}\_\{ij\}^\{\(k\)\}\(Q\\\!\\cdot\\\!\\mathcal\{C\}\_\{ij\}\)=Q\\mathbf\{S\}\_\{ij\}^\{\(k\)\}\(\\mathcal\{C\}\_\{ij\}\)Q^\{\\top\}\(10\)and the channel operator𝐌ij\(k\)\\mathbf\{M\}\_\{ij\}^\{\(k\)\}is unchanged under theO\(n\)O\(n\)action\. Then the transport map in Equation[7](https://arxiv.org/html/2608.28853#S3.E7)isO\(n\)O\(n\)\-equivariant:
𝒯i←j′\(Q𝐕j\)=Q𝒯i←j\(𝐕j\)\\mathcal\{T\}^\{\\prime\}\_\{i\\leftarrow j\}\(Q\\mathbf\{V\}\_\{j\}\)=Q\\,\\mathcal\{T\}\_\{i\\leftarrow j\}\(\\mathbf\{V\}\_\{j\}\)\(11\)where𝒯i←j′\\mathcal\{T\}^\{\\prime\}\_\{i\\leftarrow j\}denotes the transport evaluated from the transformed geometric inputs\. \(Proof is provided in Appendix[D\.1](https://arxiv.org/html/2608.28853#A4.SS1)\)\.
For a fixed edge context𝒞ij\\mathcal\{C\}\_\{ij\}, the resulting transport is linear in the vector feature being propagated\. The spatial and channel operators are themselves determined from𝒞ij\\mathcal\{C\}\_\{ij\}and may therefore depend non\-linearly on the local node, edge, and geometric information\. This allows ESNN to adapt the transport to each interaction while retaining the equivariance guaranteed by Proposition[3\.2](https://arxiv.org/html/2608.28853#S3.Thmtheorem2)\.
In ESNN, the invariant channel operators use the structured parameterization:
𝐌ij\(k\)=𝐖\(k\)𝐃\(𝐠ij\(k\)\)\\mathbf\{M\}\_\{ij\}^\{\(k\)\}=\\mathbf\{W\}^\{\(k\)\}\\mathbf\{D\}\\\!\\left\(\\mathbf\{g\}\_\{ij\}^\{\(k\)\}\\right\)\(12\)where𝐖\(k\)∈ℝcv×cv\\mathbf\{W\}^\{\(k\)\}\\in\\mathbb\{R\}^\{c\_\{v\}\\times c\_\{v\}\}is learned and shared across edges, while𝐃\(𝐠ij\(k\)\)\\mathbf\{D\}\(\\mathbf\{g\}\_\{ij\}^\{\(k\)\}\)is a diagonal gate predicted from the invariant edge features𝐳ij\\mathbf\{z\}\_\{ij\}\. The shared matrix mixes vector channels, while the edge\-dependent gate adapts this mixing to the local interaction\. Together with the spatial operators in Equation[7](https://arxiv.org/html/2608.28853#S3.E7), this yields a structured separation between two complementary roles:*channel operators control how features are mixed*, while*spatial operators control how vector messages are transformed geometrically*\. Section[4](https://arxiv.org/html/2608.28853#S4)develops several choices for the spatial action, ranging from trivial and isotropic transport to orthogonal and radial–tangential transformations, while Appendix[B](https://arxiv.org/html/2608.28853#A2)provides a more detailed algebraic account of this spatial–channel decomposition\.
Table 1:ESNN transport families\.Each transport is an instantiation of Equation[7](https://arxiv.org/html/2608.28853#S3.E7)\. The spatial operators act on the Euclidean dimension, while the channel operators act on thecvc\_\{v\}vector channels\.TransportSpatial operatorChannel operatorGeometric actionIdentity𝐈n\\mathbf\{I\}\_\{n\}𝐈cv\\mathbf\{I\}\_\{c\_\{v\}\}Trivial transportDiagonalλij𝐈n\\lambda\_\{ij\}\\mathbf\{I\}\_\{n\}𝐖𝐃\(𝐠ij\)\\mathbf\{W\}\\mathbf\{D\}\(\\mathbf\{g\}\_\{ij\}\)Isotropic scalingOrthogonal𝐑ij∈SO\(n\)\\mathbf\{R\}\_\{ij\}\\in SO\(n\)𝐖𝐃\(𝐠ij\)\\mathbf\{W\}\\mathbf\{D\}\(\\mathbf\{g\}\_\{ij\}\)Feature\-dependent rotationRadial–Tangential𝐏ij∥,𝐏ij⟂\\mathbf\{P\}^\{\\parallel\}\_\{ij\},\\,\\mathbf\{P\}^\{\\perp\}\_\{ij\}𝐖∥𝐃\(𝐠ij∥\),𝐖⟂𝐃\(𝐠ij⟂\)\\mathbf\{W\}\_\{\\parallel\}\\mathbf\{D\}\(\\mathbf\{g\}^\{\\parallel\}\_\{ij\}\),\\,\\mathbf\{W\}\_\{\\perp\}\\mathbf\{D\}\(\\mathbf\{g\}^\{\\perp\}\_\{ij\}\)Independent longitudinal and transverse transport
IdentityDiagonalOrthogonalRadial–Tangential𝐫^\\widehat\{\\mathbf\{r\}\}tang\.𝐯′=𝐯\\mathbf\{v\}^\{\\prime\}=\\mathbf\{v\}𝐒ij=𝐈n\\mathbf\{S\}\_\{ij\}=\\mathbf\{I\}\_\{n\}𝐌ij=𝐈cv\\mathbf\{M\}\_\{ij\}=\\mathbf\{I\}\_\{c\_\{v\}\}Trivial transport𝐫^\\widehat\{\\mathbf\{r\}\}tang\.𝐯\\mathbf\{v\}λ𝐯\\lambda\\,\\mathbf\{v\}𝐒ij=λij𝐈n\\mathbf\{S\}\_\{ij\}=\\lambda\_\{ij\}\\mathbf\{I\}\_\{n\}𝐌ij=𝐖𝐃\(𝐠ij\)\\mathbf\{M\}\_\{ij\}=\\mathbf\{W\}\\mathbf\{D\}\(\\mathbf\{g\}\_\{ij\}\)Isotropic spatial action𝐫^\\widehat\{\\mathbf\{r\}\}tang\.𝐯\\mathbf\{v\}𝐑ij𝐯\\mathbf\{R\}\_\{ij\}\\mathbf\{v\}𝐑ij\\mathbf\{R\}\_\{ij\}𝐒ij=𝐑ij∈SO\(n\)\\mathbf\{S\}\_\{ij\}=\\mathbf\{R\}\_\{ij\}\\in SO\(n\)𝐌ij=𝐖𝐃\(𝐠ij\)\\mathbf\{M\}\_\{ij\}=\\mathbf\{W\}\\mathbf\{D\}\(\\mathbf\{g\}\_\{ij\}\)Feature\-dependent spatial rotation𝐫^\\widehat\{\\mathbf\{r\}\}tang\.𝐯\\mathbf\{v\}𝐏ij∥\\mathbf\{P\}^\{\\parallel\}\_\{ij\}𝐏ij⟂\\mathbf\{P\}^\{\\perp\}\_\{ij\}𝐯′\\mathbf\{v\}^\{\\prime\}𝐒ij\(1\)=𝐏ij∥,𝐒ij\(2\)=𝐏ij⟂\\mathbf\{S\}\_\{ij\}^\{\(1\)\}=\\mathbf\{P\}^\{\\parallel\}\_\{ij\},\\quad\\mathbf\{S\}\_\{ij\}^\{\(2\)\}=\\mathbf\{P\}^\{\\perp\}\_\{ij\}𝐌ij∥,𝐌ij⟂\\mathbf\{M\}\_\{ij\}^\{\\parallel\},\\quad\\mathbf\{M\}\_\{ij\}^\{\\perp\}Displacement\-conditioned radial–tangential formPrincipal ESNN spatial transport families
Figure 2:*Principal ESNN transport families\.*The four constructions differ in their spatial action while their channel actions remain invariant underO\(n\)O\(n\)\. Identity Transport leaves the vector representation unchanged; Diagonal Transport applies an edge\-dependent but spatially isotropic scaling; Orthogonal Transport learns a feature\-conditioned spatial rotation; and Radial–Tangential Transport transforms components parallel and orthogonal to the relative displacement independently\. When the displacement is the only covariant geometric input, the Radial–Tangential form characterizes the complete class of linearO\(n\)O\(n\)\-equivariant transports \(Theorem[4\.3](https://arxiv.org/html/2608.28853#S4.Thmtheorem3)\)\.
## 4Equivariant Transport Maps
We now instantiate the spatial–channel decomposition of Section[3](https://arxiv.org/html/2608.28853#S3)through four transport families, summarized in Table[1](https://arxiv.org/html/2608.28853#S3.T1)and illustrated in Figure[2](https://arxiv.org/html/2608.28853#S3.F2)\. Each family corresponds to a different choice of spatial operator𝐒ij\(k\)\\mathbf\{S\}\_\{ij\}^\{\(k\)\}in Equation[7](https://arxiv.org/html/2608.28853#S3.E7), while the associated scalar coefficients and channel gates are computed fromO\(n\)O\(n\)\-invariant features\. The constructions below act on covariant vector features𝐕j∈ℝn×cv\\mathbf\{V\}\_\{j\}\\in\\mathbb\{R\}^\{n\\times c\_\{v\}\}; scalar features follow the invariant message\-passing pathway introduced with the full ESNN layer in Section[5](https://arxiv.org/html/2608.28853#S5)\.
Identity Transport\.The simplest choice is the*Identity Transport*:
𝒯i←j\(𝐕j\)=𝐕j\\mathcal\{T\}\_\{i\\leftarrow j\}\(\\mathbf\{V\}\_\{j\}\)=\\mathbf\{V\}\_\{j\}\(13\)corresponding to𝐒ij=𝐈n\\mathbf\{S\}\_\{ij\}=\\mathbf\{I\}\_\{n\}and𝐌ij=𝐈cv\\mathbf\{M\}\_\{ij\}=\\mathbf\{I\}\_\{c\_\{v\}\}\. Vector features are therefore aggregated in the common frame without an edge\-dependent spatial transformation, giving the trivial\-transport limit of ESNN\.
Diagonal Transport\.A more expressive but still spatially isotropic choice is to set𝐒ij=λij𝐈n,\\mathbf\{S\}\_\{ij\}=\\lambda\_\{ij\}\\mathbf\{I\}\_\{n\},and𝐌ij=𝐖𝐃\(𝐠ij\)\\mathbf\{M\}\_\{ij\}=\\mathbf\{W\}\\mathbf\{D\}\(\\mathbf\{g\}\_\{ij\}\), whereλij∈ℝ\\lambda\_\{ij\}\\in\\mathbb\{R\}is predicted from invariant edge features\. The resulting*Diagonal Transport*is:
𝒯i←j\(𝐕j\)=λij\(𝐕j𝐖\)𝐃\(𝐠ij\)\\mathcal\{T\}\_\{i\\leftarrow j\}\(\\mathbf\{V\}\_\{j\}\)=\\lambda\_\{ij\}\(\\mathbf\{V\}\_\{j\}\\mathbf\{W\}\)\\mathbf\{D\}\(\\mathbf\{g\}\_\{ij\}\)\(14\)Hereλij\\lambda\_\{ij\}acts identically along every spatial direction,𝐖\\mathbf\{W\}mixes vector channels, and𝐠ij\\mathbf\{g\}\_\{ij\}adapts this mixing to the local edge context\. The transport can therefore vary from edge to edge while remaining isotropic in physical space\.
Orthogonal Transport\.Diagonal Transport can modulate a vector message but its spatial action remains proportional to the identity\. Orthogonal Transport introduces a non\-trivial edge\-dependent spatial transformation by constructing a rotation from the current vector features\. We first form the cross\-feature matrix and its normalized counterpart:
𝐂ij=𝐕j𝐕i⊤∈ℝn×n,𝐂~ij=𝐂ij‖𝐂ij‖F\+ε\\mathbf\{C\}\_\{ij\}=\\mathbf\{V\}\_\{j\}\\mathbf\{V\}\_\{i\}^\{\\top\}\\in\\mathbb\{R\}^\{n\\times n\},\\qquad\\widetilde\{\\mathbf\{C\}\}\_\{ij\}=\\frac\{\\mathbf\{C\}\_\{ij\}\}\{\\\|\\mathbf\{C\}\_\{ij\}\\\|\_\{F\}\+\\varepsilon\}\(15\)Its skew\-symmetric component𝛀ij=𝐂~ij−𝐂~ij⊤∈𝔰𝔬\(n\)\\mathbf\{\\Omega\}\_\{ij\}=\\widetilde\{\\mathbf\{C\}\}\_\{ij\}\-\\widetilde\{\\mathbf\{C\}\}\_\{ij\}^\{\\top\}\\in\\mathfrak\{so\}\(n\)provides a generator of a rotation\. An invariant edge network predicts a scalar coefficientβij\\beta\_\{ij\}, from which we construct:
𝐑ij=exp\(βij𝛀ij\)∈SO\(n\)\\mathbf\{R\}\_\{ij\}=\\exp\\\!\\left\(\\beta\_\{ij\}\\mathbf\{\\Omega\}\_\{ij\}\\right\)\\in SO\(n\)\(16\)The resulting*Orthogonal Transport*is:
𝒯i←j\(𝐕j\)=𝐑ij\(𝐕j𝐖\)𝐃\(𝐠ij\)\\mathcal\{T\}\_\{i\\leftarrow j\}\(\\mathbf\{V\}\_\{j\}\)=\\mathbf\{R\}\_\{ij\}\(\\mathbf\{V\}\_\{j\}\\mathbf\{W\}\)\\mathbf\{D\}\(\\mathbf\{g\}\_\{ij\}\)\(17\)
Because𝐑ij\\mathbf\{R\}\_\{ij\}is inferred from the current vector representations, the spatial transformation adapts to the local feature geometry around the edge\(i,j\)\(i,j\)\. Once this local context is fixed,𝐑ij∈SO\(n\)\\mathbf\{R\}\_\{ij\}\\in SO\(n\)acts linearly on the transported vector features\. Under a global orthogonal transformation, both its generator and matrix exponential transform by conjugation, giving the required equivariant transport\.
###### Lemma 4\.1\(O\(n\)O\(n\)\-Equivariance of Orthogonal Transport\)\.
LetQ∈O\(n\)Q\\in O\(n\)act on the vector features as𝐕i↦Q𝐕i\\mathbf\{V\}\_\{i\}\\mapsto Q\\mathbf\{V\}\_\{i\}\. Then:
𝒯i←j′\(Q𝐕j\)=Q𝒯i←j\(𝐕j\)\\mathcal\{T\}^\{\\prime\}\_\{i\\leftarrow j\}\(Q\\mathbf\{V\}\_\{j\}\)=Q\\,\\mathcal\{T\}\_\{i\\leftarrow j\}\(\\mathbf\{V\}\_\{j\}\)\(18\)where the primed transport is evaluated from the transformed geometric features\. \(Proof is provided in Appendix[D\.2](https://arxiv.org/html/2608.28853#A4.SS2)\)\.
The term*Orthogonal Transport*refers to the spatial factor𝐑ij∈SO\(n\)\\mathbf\{R\}\_\{ij\}\\in SO\(n\)\. The complete edge map also includes learned channel mixing and edge\-dependent gating and is therefore not, in general, an orthogonal transformation of the full vector feature space\.
Radial–Tangential Transport\.The transports above apply a single spatial transformation to the entire vector message\. Many geometric interactions, however, distinguish components along an edge from those orthogonal to it\. Radial–Tangential Transport makes this distinction explicit using the relative displacement𝐫ij=𝐱i−𝐱j\\mathbf\{r\}\_\{ij\}=\\mathbf\{x\}\_\{i\}\-\\mathbf\{x\}\_\{j\}\. For𝐫ij≠𝟎\\mathbf\{r\}\_\{ij\}\\neq\\mathbf\{0\}, let𝐫^ij=𝐫ij‖𝐫ij‖\\widehat\{\\mathbf\{r\}\}\_\{ij\}=\\frac\{\\mathbf\{r\}\_\{ij\}\}\{\\\|\\mathbf\{r\}\_\{ij\}\\\|\}and define the orthogonal projectors:
𝐏ij∥=𝐫^ij𝐫^ij⊤,𝐏ij⟂=𝐈n−𝐏ij∥\\mathbf\{P\}^\{\\parallel\}\_\{ij\}=\\widehat\{\\mathbf\{r\}\}\_\{ij\}\\widehat\{\\mathbf\{r\}\}\_\{ij\}^\{\\top\},\\qquad\\mathbf\{P\}^\{\\perp\}\_\{ij\}=\\mathbf\{I\}\_\{n\}\-\\mathbf\{P\}^\{\\parallel\}\_\{ij\}\(19\)The first extracts the component parallel to the edge and the second its orthogonal complement\. ESNN assigns an independent channel transformation to each component:
𝐌ij∥=𝐖∥𝐃\(𝐠ij∥\),𝐌ij⟂=𝐖⟂𝐃\(𝐠ij⟂\)\\mathbf\{M\}^\{\\parallel\}\_\{ij\}=\\mathbf\{W\}\_\{\\parallel\}\\mathbf\{D\}\(\\mathbf\{g\}^\{\\parallel\}\_\{ij\}\),\\qquad\\mathbf\{M\}^\{\\perp\}\_\{ij\}=\\mathbf\{W\}\_\{\\perp\}\\mathbf\{D\}\(\\mathbf\{g\}^\{\\perp\}\_\{ij\}\)\(20\)giving the*Radial–Tangential Transport*:
𝒯i←j\(𝐕j\)=𝐏ij∥𝐕j𝐌ij∥\+𝐏ij⟂𝐕j𝐌ij⟂\\mathcal\{T\}\_\{i\\leftarrow j\}\(\\mathbf\{V\}\_\{j\}\)=\\mathbf\{P\}^\{\\parallel\}\_\{ij\}\\mathbf\{V\}\_\{j\}\\mathbf\{M\}^\{\\parallel\}\_\{ij\}\+\\mathbf\{P\}^\{\\perp\}\_\{ij\}\\mathbf\{V\}\_\{j\}\\mathbf\{M\}^\{\\perp\}\_\{ij\}\(21\)The longitudinal and transverse components can therefore be transformed independently before aggregation, providing anisotropic transport while remaining entirely within first\-order Cartesian features\. The projectors need not be materialized as densen×nn\\times nmatrices\. Their action can be evaluated directly as:
𝐏ij∥𝐕j=𝐫^ij\(𝐫^ij⊤𝐕j\),𝐏ij⟂𝐕j=𝐕j−𝐏ij∥𝐕j\\mathbf\{P\}^\{\\parallel\}\_\{ij\}\\mathbf\{V\}\_\{j\}=\\widehat\{\\mathbf\{r\}\}\_\{ij\}\\left\(\\widehat\{\\mathbf\{r\}\}\_\{ij\}^\{\\top\}\\mathbf\{V\}\_\{j\}\\right\),\\qquad\\mathbf\{P\}^\{\\perp\}\_\{ij\}\\mathbf\{V\}\_\{j\}=\\mathbf\{V\}\_\{j\}\-\\mathbf\{P\}^\{\\parallel\}\_\{ij\}\\mathbf\{V\}\_\{j\}\(22\)The spatial projection therefore costs𝒪\(ncv\)\\mathcal\{O\}\(nc\_\{v\}\)\. With densecv×cvc\_\{v\}\\times c\_\{v\}channel transformations, the overall complexity becomes𝒪\(ncv2\)\\mathcal\{O\}\(nc\_\{v\}^\{2\}\)up to constant factors, while remaining linear in the spatial dimensionnn\. Self\-interactions require separate treatment because𝐫ii=𝟎\\mathbf\{r\}\_\{ii\}=\\mathbf\{0\}does not define radial and tangential directions\. ESNN therefore handles self\-information through the identity self\-loop of the diffusion operator\.
###### Lemma 4\.2\(O\(n\)O\(n\)\-Equivariance of Radial–Tangential Transport\)\.
For𝐫ij≠𝟎\\mathbf\{r\}\_\{ij\}\\neq\\mathbf\{0\}, the transport map in Equation[21](https://arxiv.org/html/2608.28853#S4.E21)isO\(n\)O\(n\)\-equivariant\. In particular:
𝐏ij∥↦Q𝐏ij∥Q⊤,𝐏ij⟂↦Q𝐏ij⟂Q⊤\\mathbf\{P\}^\{\\parallel\}\_\{ij\}\\mapsto Q\\mathbf\{P\}^\{\\parallel\}\_\{ij\}Q^\{\\top\},\\qquad\\mathbf\{P\}^\{\\perp\}\_\{ij\}\\mapsto Q\\mathbf\{P\}^\{\\perp\}\_\{ij\}Q^\{\\top\}\(23\)while the channel operators remain invariant\. \(Proof is provided in Appendix[D\.3](https://arxiv.org/html/2608.28853#A4.SS3)\)\.
Radial–Tangential Transport is not merely one equivariant choice among many\. When the relative displacement is the only covariant geometric input, it captures the most general linear transport compatible with fullO\(n\)O\(n\)\-equivariance\.
###### Theorem 4\.3\.
Letn≥2n\\geq 2, and consider a linear transport onℝn⊗ℝcv\\mathbb\{R\}^\{n\}\\otimes\\mathbb\{R\}^\{c\_\{v\}\}whose only covariant geometric conditioning variable is a non\-zero relative displacement𝐫ij\\mathbf\{r\}\_\{ij\}, with arbitrary additionalO\(n\)O\(n\)\-invariant scalar conditioning\. Every suchO\(n\)O\(n\)\-equivariant transport can be written as:
𝒯i←j\(𝐕j\)=𝐏ij∥𝐕j𝐀ij\+𝐏ij⟂𝐕j𝐁ij\\mathcal\{T\}\_\{i\\leftarrow j\}\(\\mathbf\{V\}\_\{j\}\)=\\mathbf\{P\}^\{\\parallel\}\_\{ij\}\\mathbf\{V\}\_\{j\}\\mathbf\{A\}\_\{ij\}\+\\mathbf\{P\}^\{\\perp\}\_\{ij\}\\mathbf\{V\}\_\{j\}\\mathbf\{B\}\_\{ij\}\(24\)for invariant channel endomorphisms𝐀ij,𝐁ij∈ℝcv×cv\\mathbf\{A\}\_\{ij\},\\mathbf\{B\}\_\{ij\}\\in\\mathbb\{R\}^\{c\_\{v\}\\times c\_\{v\}\}\. \(Proof is given in Appendix[B\.2](https://arxiv.org/html/2608.28853#A2.SS2)using the stabilizer of a non\-zero displacement\. \)
The theorem shows that unrestricted radial and tangential channel transformations span the complete class of displacement\-conditioned linearO\(n\)O\(n\)\-equivariant transports\. ESNN realizes these transformations through the structured factorization𝐖𝐃\(𝐠ij\)\\mathbf\{W\}\\mathbf\{D\}\(\\mathbf\{g\}\_\{ij\}\), which lies within this theoretical class\. This completeness result is specific to the displacement\-conditionedO\(n\)O\(n\)setting: when learned covariant vector features are also available, the admissible spatial operators become richer, and underSO\(n\)SO\(n\)additional orientation\-sensitive constructions may also arise\.
Unified Transport\.The completeness result above assumes that the relative displacement is the only covariant geometric input\. Once the learned vector features are also available, they can be used to construct additional covariant spatial operators\. In the unified formulation, let𝒦⊆\{id,∥,⟂,skew\}\\mathcal\{K\}\\subseteq\\left\\\{\\mathrm\{id\},\\parallel,\\perp,\\mathrm\{skew\}\\right\\\}denote the active spatial operators, with:
𝐁\(id\)ij=𝐈n,𝐁\(∥\)ij=𝐏∥ij,𝐁\(⟂\)ij=𝐏⟂ij,𝐁\(skew\)ij=𝛀^Vij\\mathbf\{B\}^\{\(\\mathrm\{id\}\)\}\_\{ij\}=\\mathbf\{I\}\_\{n\},\\qquad\\mathbf\{B\}^\{\(\\parallel\)\}\_\{ij\}=\\mathbf\{P\}^\{\\parallel\}\_\{ij\},\\qquad\\mathbf\{B\}^\{\(\\perp\)\}\_\{ij\}=\\mathbf\{P\}^\{\\perp\}\_\{ij\},\\qquad\\mathbf\{B\}^\{\(\\mathrm\{skew\}\)\}\_\{ij\}=\\widehat\{\\mathbf\{\\Omega\}\}^\{V\}\_\{ij\}\(25\)where we have the feature\-dependent covariant skew\-symmetric operator and its normalized form to be defined as:
𝛀ijV=𝐕j𝐕i⊤−𝐕i𝐕j⊤,𝛀^ijV=𝛀ijV‖𝛀ijV‖F\+ε\\mathbf\{\\Omega\}^\{V\}\_\{ij\}=\\mathbf\{V\}\_\{j\}\\mathbf\{V\}\_\{i\}^\{\\top\}\-\\mathbf\{V\}\_\{i\}\\mathbf\{V\}\_\{j\}^\{\\top\},\\qquad\\widehat\{\\mathbf\{\\Omega\}\}^\{V\}\_\{ij\}=\\frac\{\\mathbf\{\\Omega\}^\{V\}\_\{ij\}\}\{\\\|\\mathbf\{\\Omega\}^\{V\}\_\{ij\}\\\|\_\{F\}\+\\varepsilon\}\(26\)Because the Frobenius norm is invariant under orthogonal conjugation,𝛀^ijV\\widehat\{\\mathbf\{\\Omega\}\}^\{V\}\_\{ij\}transforms as𝛀^ijV↦Q𝛀^ijVQ⊤\\widehat\{\\mathbf\{\\Omega\}\}^\{V\}\_\{ij\}\\mapsto Q\\widehat\{\\mathbf\{\\Omega\}\}^\{V\}\_\{ij\}Q^\{\\top\}\. All spatial operators in𝒦\\mathcal\{K\}therefore satisfy Equation[8](https://arxiv.org/html/2608.28853#S3.E8)\. Together with invariant channel maps𝐌ij\(b\)\\mathbf\{M\}^\{\(b\)\}\_\{ij\}, Proposition[3\.2](https://arxiv.org/html/2608.28853#S3.Thmtheorem2)guaranteesO\(n\)O\(n\)\-equivariance\. We can define the*Unified Transport*therefore as:
𝒯i←j\(𝐕j\)=∑b∈𝒦𝐁ij\(b\)𝐕j𝐌ij\(b\)\\mathcal\{T\}\_\{i\\leftarrow j\}\(\\mathbf\{V\}\_\{j\}\)=\\sum\_\{b\\in\\mathcal\{K\}\}\\mathbf\{B\}^\{\(b\)\}\_\{ij\}\\mathbf\{V\}\_\{j\}\\mathbf\{M\}^\{\(b\)\}\_\{ij\}\(27\)This formulation*extends the transport beyond displacement\-only geometry*by allowing its spatial action to depend on the learned vector features\. It therefore lies outside the setting characterized by Theorem[4\.3](https://arxiv.org/html/2608.28853#S4.Thmtheorem3), while following the same equivariance principle established in Section[3](https://arxiv.org/html/2608.28853#S3)\.
## 5The ESNN Architecture
The full ESNN layer combines the transport maps of Section[4](https://arxiv.org/html/2608.28853#S4)with invariant scalar messaging and, when required, coordinate dynamics\. At each nodeii, the layer receives coordinates𝐱i∈ℝn\\mathbf\{x\}\_\{i\}\\in\\mathbb\{R\}^\{n\}, invariant scalar features𝐬i∈ℝcs\\mathbf\{s\}\_\{i\}\\in\\mathbb\{R\}^\{c\_\{s\}\}, and covariant vector features𝐕i∈ℝn×cv\\mathbf\{V\}\_\{i\}\\in\\mathbb\{R\}^\{n\\times c\_\{v\}\}\. The update follows five stages: an invariant edge context parameterizes scalar messages and vector transport, the resulting messages are aggregated through a normalized transport operator, scalar and vector features are updated within their respective representation spaces, and an optional kinematic update evolves the coordinates\.
1\. Invariant Edge Context\.For each non\-self directed interactionj→ij\\rightarrow iwith𝐫ij≠𝟎\\mathbf\{r\}\_\{ij\}\\neq\\mathbf\{0\}, ESNN first constructs an invariant description of the local interaction\. Let𝐫ij=𝐱i−𝐱j,𝐫^ij=𝐫ij‖𝐫ij‖\\mathbf\{r\}\_\{ij\}=\\mathbf\{x\}\_\{i\}\-\\mathbf\{x\}\_\{j\},\\widehat\{\\mathbf\{r\}\}\_\{ij\}=\\frac\{\\mathbf\{r\}\_\{ij\}\}\{\\\|\\mathbf\{r\}\_\{ij\}\\\|\}, and denote by𝐧\(𝐕i\)=\(‖𝐯i,1‖,…,‖𝐯i,cv‖\)\\mathbf\{n\}\(\\mathbf\{V\}\_\{i\}\)=\\bigl\(\\\|\\mathbf\{v\}\_\{i,1\}\\\|,\\ldots,\\\|\\mathbf\{v\}\_\{i,c\_\{v\}\}\\\|\\bigr\)the channel\-wise vector norms\. The edge context is:
𝐳ij=\[𝐬i,𝐬j,𝐧\(𝐕i\),𝐧\(𝐕j\),ϕr\(‖𝐫ij‖\),diag\(𝐕i⊤𝐕j\),𝐕i⊤𝐫^ij,𝐕j⊤𝐫^ij,𝐞ijattr\]\\mathbf\{z\}\_\{ij\}=\\Big\[\\mathbf\{s\}\_\{i\},\\mathbf\{s\}\_\{j\},\\,\\mathbf\{n\}\(\\mathbf\{V\}\_\{i\}\),\\mathbf\{n\}\(\\mathbf\{V\}\_\{j\}\),\\,\\phi\_\{r\}\(\\\|\\mathbf\{r\}\_\{ij\}\\\|\),\\operatorname\{diag\}\(\\mathbf\{V\}\_\{i\}^\{\\top\}\\mathbf\{V\}\_\{j\}\),\\,\\mathbf\{V\}\_\{i\}^\{\\top\}\\widehat\{\\mathbf\{r\}\}\_\{ij\},\\,\\mathbf\{V\}\_\{j\}^\{\\top\}\\widehat\{\\mathbf\{r\}\}\_\{ij\},\\,\\mathbf\{e\}^\{\\mathrm\{attr\}\}\_\{ij\}\\Big\]\(28\)whereϕr\\phi\_\{r\}provides a radial representation of the distance and𝐞ijattr\\mathbf\{e\}^\{\\mathrm\{attr\}\}\_\{ij\}collects optional invariant edge attributes\. Every component of𝐳ij\\mathbf\{z\}\_\{ij\}isO\(n\)O\(n\)\-invariant and can therefore be used to predict edge\-dependent gates and transport coefficients without breaking equivariance\. Self\-information does not require a directional edge description and is instead propagated through the explicit identity self\-loop of the diffusion operator, so𝐳ij\\mathbf\{z\}\_\{ij\}is only constructed for non\-self interactions\.
2\. Vector Transport and Scalar Messaging\.The vector message is obtained by applying one of the transport maps from Section[4](https://arxiv.org/html/2608.28853#S4):
𝐦i←jV=𝒯i←j\(𝐕j\)\\mathbf\{m\}^\{V\}\_\{i\\leftarrow j\}=\\mathcal\{T\}\_\{i\\leftarrow j\}\(\\mathbf\{V\}\_\{j\}\)\(29\)Its coefficients are determined by the invariant edge context and, for feature\-dependent transport families, by covariant geometric quantities satisfying Equation[8](https://arxiv.org/html/2608.28853#S3.E8)\. Scalar features are propagated through a separate invariant pathway:
𝐦i←js=𝐬j\+ϕs\(𝐳ijs\)\\mathbf\{m\}^\{s\}\_\{i\\leftarrow j\}=\\mathbf\{s\}\_\{j\}\+\\phi\_\{s\}\(\\mathbf\{z\}^\{s\}\_\{ij\}\)\(30\)where𝐳ijs\\mathbf\{z\}^\{s\}\_\{ij\}contains invariant scalar and radial features and may also include the invariant vector inner products in Equation[28](https://arxiv.org/html/2608.28853#S5.E28)\. The two pathways therefore allow geometric information to influence both scalar and vector representations while preserving their respective transformation laws\.
3\. Normalized Transport Diffusion\.The scalar and vector messages are then aggregated over the neighborhood of each node\. Letωij≥0\\omega\_\{ij\}\\geq 0denote an optionalO\(n\)O\(n\)\-invariant weight for the directed interactionj→ij\\rightarrow i, withωij=1\\omega\_\{ij\}=1in the unweighted case\. For learned directional weights, we use the symmetrized weighted degree:
d¯i=12\(∑jωij\+∑jωji\)\\bar\{d\}\_\{i\}=\\frac\{1\}\{2\}\\left\(\\sum\_\{j\}\\omega\_\{ij\}\+\\sum\_\{j\}\\omega\_\{ji\}\\right\)\(31\)whereas for the unweighted operatord¯i=\|𝒩\(i\)\|\\bar\{d\}\_\{i\}=\|\\mathcal\{N\}\(i\)\|\. The normalized coefficients are:
νij=ωij\(d¯i\+1\)\(d¯j\+1\),νii=1d¯i\+1\\nu\_\{ij\}=\\frac\{\\omega\_\{ij\}\}\{\\sqrt\{\(\\bar\{d\}\_\{i\}\+1\)\(\\bar\{d\}\_\{j\}\+1\)\}\},\\qquad\\nu\_\{ii\}=\\frac\{1\}\{\\bar\{d\}\_\{i\}\+1\}\(32\)The additional unit accounts for the explicit identity self\-loop\. The diffused representations are:
𝐕idiff=νii𝐕i\+∑j∈𝒩\(i\)νij𝐦i←jV,𝐬idiff=νii𝐬i\+∑j∈𝒩\(i\)νij𝐦i←js\\mathbf\{V\}^\{\\mathrm\{diff\}\}\_\{i\}=\\nu\_\{ii\}\\mathbf\{V\}\_\{i\}\+\\sum\_\{j\\in\\mathcal\{N\}\(i\)\}\\nu\_\{ij\}\\,\\mathbf\{m\}^\{V\}\_\{i\\leftarrow j\},\\qquad\\mathbf\{s\}^\{\\mathrm\{diff\}\}\_\{i\}=\\nu\_\{ii\}\\mathbf\{s\}\_\{i\}\+\\sum\_\{j\\in\\mathcal\{N\}\(i\)\}\\nu\_\{ij\}\\,\\mathbf\{m\}^\{s\}\_\{i\\leftarrow j\}\(33\)For a fixed layer context, the vector branch defines the normalized transport operator:
\(𝒜𝒯𝐕\)i=νii𝐕i\+∑j∈𝒩\(i\)νij𝒯i←j\(𝐕j\)\(\\mathcal\{A\}\_\{\\mathcal\{T\}\}\\mathbf\{V\}\)\_\{i\}=\\nu\_\{ii\}\\mathbf\{V\}\_\{i\}\+\\sum\_\{j\\in\\mathcal\{N\}\(i\)\}\\nu\_\{ij\}\\,\\mathcal\{T\}\_\{i\\leftarrow j\}\(\\mathbf\{V\}\_\{j\}\)\(34\)The two orientations of an edge may carry different transport maps, so the general ESNN operator is directional and need not be self\-adjoint\. Recent directed sheaf models similarly distinguish edge orientations through asymmetric sheaf operators\([Ribeiro et al\., 2025](https://arxiv.org/html/2608.28853#bib.bib44);[Fiorini et al\., 2025](https://arxiv.org/html/2608.28853#bib.bib45)\)\. ESNN instead parameterizes the directed node\-to\-node transport directly and impose the ambientO\(n\)O\(n\)covariance constraint on this map\. A self\-adjoint specialization is recovered on a bidirected graph whenνij=νji\\nu\_\{ij\}=\\nu\_\{ji\}and𝒯j←i=𝒯i←j∗\\mathcal\{T\}\_\{j\\leftarrow i\}=\\mathcal\{T\}\_\{i\\leftarrow j\}^\{\*\}\. Under the corresponding orthogonal specialization, this yields the connection\-style coupling of classical connection sheaves\. Appendix[C](https://arxiv.org/html/2608.28853#A3)discusses these relationships in greater detail\.
Degree normalization is the default aggregation mechanism\. ESNN can also support invariant multi\-head attention, whereνij\\nu\_\{ij\}is replaced by receiver\-normalized weights computed from invariant edge features\. Since these weights are invariant scalars, the resulting aggregation preserves the sameO\(n\)O\(n\)\-equivariance\.
4\. Residual Feature Update\.After aggregation, scalar and vector features are updated within their respective representation spaces\. The vector branch first applies a learned channel transformation:
𝐕~i=𝐕idiff𝐖V,𝐖V∈ℝcv×cv\\widetilde\{\\mathbf\{V\}\}\_\{i\}=\\mathbf\{V\}^\{\\mathrm\{diff\}\}\_\{i\}\\mathbf\{W\}\_\{V\},\\qquad\\mathbf\{W\}\_\{V\}\\in\\mathbb\{R\}^\{c\_\{v\}\\times c\_\{v\}\}\(35\)Since𝐖V\\mathbf\{W\}\_\{V\}acts only on the channel dimension,\(Q𝐕\)𝐖V=Q\(𝐕𝐖V\)\(Q\\mathbf\{V\}\)\\mathbf\{W\}\_\{V\}=Q\(\\mathbf\{V\}\\mathbf\{W\}\_\{V\}\), so channel mixing preserves equivariance\. A radial non\-linearity is then applied independently to each vector channel,σV\(𝐯\)=a\(‖𝐯‖\)𝐯\\sigma\_\{V\}\(\\mathbf\{v\}\)=a\(\\\|\\mathbf\{v\}\\\|\)\\mathbf\{v\}, withaaa learned scalar function\. Because the scaling depends only on the vector norm, the output transforms covariantly\. The scalar branch uses an ordinary linear map and scalar non\-linearity\. Including residual connections, the update is:
𝐕i′=𝐕i\+σV\(𝐕idiff𝐖V\),𝐬i′=𝐬i\+σs\(𝐖s𝐬idiff\)\\mathbf\{V\}^\{\\prime\}\_\{i\}=\\mathbf\{V\}\_\{i\}\+\\sigma\_\{V\}\\\!\\left\(\\mathbf\{V\}^\{\\mathrm\{diff\}\}\_\{i\}\\mathbf\{W\}\_\{V\}\\right\),\\qquad\\mathbf\{s\}^\{\\prime\}\_\{i\}=\\mathbf\{s\}\_\{i\}\+\\sigma\_\{s\}\\\!\\left\(\\mathbf\{W\}\_\{s\}\\mathbf\{s\}^\{\\mathrm\{diff\}\}\_\{i\}\\right\)\(36\)Thus the layer can mix information freely across channels without changing the transformation type of either representation\.
5\. Coordinate Kinematics Update\.For tasks with evolving geometry, ESNN includes an EGNN\-style coordinate update driven by invariant features\. From the updated representation we form𝐡iinv=\[𝐧\(𝐕i′\),𝐬i′\]\\mathbf\{h\}^\{\\mathrm\{inv\}\}\_\{i\}=\\big\[\\mathbf\{n\}\(\\mathbf\{V\}^\{\\prime\}\_\{i\}\),\\mathbf\{s\}^\{\\prime\}\_\{i\}\\big\]and predict the invariant scalarγij=ϕx\(𝐡iinv,𝐡jinv,ϕr\(‖𝐫ij‖\),𝐞ijattr\)\\gamma\_\{ij\}=\\phi\_\{x\}\\\!\\left\(\\mathbf\{h\}^\{\\mathrm\{inv\}\}\_\{i\},\\mathbf\{h\}^\{\\mathrm\{inv\}\}\_\{j\},\\phi\_\{r\}\(\\\|\\mathbf\{r\}\_\{ij\}\\\|\),\\mathbf\{e\}^\{\\mathrm\{attr\}\}\_\{ij\}\\right\)\. The coordinate displacement is:
Δ𝐱i=1\|𝒩\(i\)\|∑j∈𝒩\(i\)γij𝐫ij,𝐱i′=𝐱i\+Δ𝐱i\\Delta\\mathbf\{x\}\_\{i\}=\\frac\{1\}\{\|\\mathcal\{N\}\(i\)\|\}\\sum\_\{j\\in\\mathcal\{N\}\(i\)\}\\gamma\_\{ij\}\\mathbf\{r\}\_\{ij\},\\qquad\\mathbf\{x\}^\{\\prime\}\_\{i\}=\\mathbf\{x\}\_\{i\}\+\\Delta\\mathbf\{x\}\_\{i\}\(37\)Sinceγij\\gamma\_\{ij\}is invariant and𝐫ij↦Q𝐫ij\\mathbf\{r\}\_\{ij\}\\mapsto Q\\mathbf\{r\}\_\{ij\}, the displacement transforms covariantly\. The transport mechanism can therefore enrich the latent vector representation while retaining the familiar first\-order coordinate update of equivariant GNNs\.
Each stage preserves the transformation type of its inputs, yielding anE\(n\)E\(n\)\-equivariant layer under the assumptions below\.
###### Theorem 5\.1\(E\(n\)E\(n\)\-Equivariance\)\.
Assume that the graph topology is fixed or constructed fromE\(n\)E\(n\)\-invariant geometric quantities, that the edge attributes areO\(n\)O\(n\)\-invariant, and that every non\-self interaction for which𝐫^ij\\widehat\{\\mathbf\{r\}\}\_\{ij\}is used satisfies𝐫ij≠𝟎\\mathbf\{r\}\_\{ij\}\\neq\\mathbf\{0\}\. Assume further that the transport maps satisfy the conditions of Proposition[3\.2](https://arxiv.org/html/2608.28853#S3.Thmtheorem2)\. In the absence of explicit symmetry\-relaxing inputs, the ESNN layer defined by Equations[28](https://arxiv.org/html/2608.28853#S5.E28)–[37](https://arxiv.org/html/2608.28853#S5.E37)isE\(n\)E\(n\)\-equivariant\. That is, under:
𝐱i↦Q𝐱i\+𝐭,𝐕i↦Q𝐕i,𝐬i↦𝐬i\\mathbf\{x\}\_\{i\}\\mapsto Q\\mathbf\{x\}\_\{i\}\+\\mathbf\{t\},\\qquad\\mathbf\{V\}\_\{i\}\\mapsto Q\\mathbf\{V\}\_\{i\},\\qquad\\mathbf\{s\}\_\{i\}\\mapsto\\mathbf\{s\}\_\{i\}\(38\)the updated features satisfy:
𝐱i′↦Q𝐱i′\+𝐭,𝐕i′↦Q𝐕i′,𝐬i′↦𝐬i′\\mathbf\{x\}^\{\\prime\}\_\{i\}\\mapsto Q\\mathbf\{x\}^\{\\prime\}\_\{i\}\+\\mathbf\{t\},\\qquad\\mathbf\{V\}^\{\\prime\}\_\{i\}\\mapsto Q\\mathbf\{V\}^\{\\prime\}\_\{i\},\\qquad\\mathbf\{s\}^\{\\prime\}\_\{i\}\\mapsto\\mathbf\{s\}^\{\\prime\}\_\{i\}\(39\)\(A proof is provided in Appendix[D\.5](https://arxiv.org/html/2608.28853#A4.SS5)\)\.
Extension to Dynamical Systems\.For dynamical systems, nodes may additionally carry a velocity𝐮i∈ℝn\\mathbf\{u\}\_\{i\}\\in\\mathbb\{R\}^\{n\}, treated as a covariant vector\. The same equivariant displacementΔ𝐱i\\Delta\\mathbf\{x\}\_\{i\}can then be used to update both velocity and position:
𝐮i′=ai𝐮i\+Δ𝐱i,𝐱i′=𝐱i\+𝐮i′\\mathbf\{u\}^\{\\prime\}\_\{i\}=a\_\{i\}\\mathbf\{u\}\_\{i\}\+\\Delta\\mathbf\{x\}\_\{i\},\\qquad\\mathbf\{x\}^\{\\prime\}\_\{i\}=\\mathbf\{x\}\_\{i\}\+\\mathbf\{u\}^\{\\prime\}\_\{i\}\(40\)whereaia\_\{i\}is an invariant scalar predicted from𝐡iinv\\mathbf\{h\}^\{\\mathrm\{inv\}\}\_\{i\}\. Since both𝐮i\\mathbf\{u\}\_\{i\}andΔ𝐱i\\Delta\\mathbf\{x\}\_\{i\}transform covariantly, Equation[40](https://arxiv.org/html/2608.28853#S5.E40)preservesE\(n\)E\(n\)\-equivariance\.
FullE\(n\)E\(n\)EquivarianceNo Preferred Direction𝐱i\\mathbf\{x\}\_\{i\}The fullO\(n\)O\(n\)actionis enforcedAdaptive Symmetry RelaxationControlled Symmetry ReductionDirectional conditioning viapreferred direction𝐠\\mathbf\{g\}λ=0\\lambda=0FullE\(n\)E\(n\)λ≠0\\lambda\\neq 0ReducedE𝐠\(n\)E\_\{\\mathbf\{g\}\}\(n\)Learnableλ\\lambdaDirectional influence vanishescontinuously asλ→0\\lambda\\to 0StabilizerO𝐠\(n\)O\_\{\\mathbf\{g\}\}\(n\)Fixed Ambient Direction𝐠\\mathbf\{g\}✓\\checkmark×\\timesOnly transformations satisfyingQ𝐠=𝐠Q\\mathbf\{g\}=\\mathbf\{g\}are enforced
Figure 3:*Controlled symmetry relaxation through a preferred ambient direction\.*Withλ=0\\lambda=0, the directional pathway is inactive and ESNN retains fullE\(n\)E\(n\)\-equivariance\. When the directional featureλ⟨𝐫ij,𝐠⟩\\lambda\\langle\\mathbf\{r\}\_\{ij\},\\mathbf\{g\}\\rangleis active, the guaranteed orthogonal symmetry is reduced to the stabilizerO𝐠\(n\)O\_\{\\mathbf\{g\}\}\(n\), while translation equivariance is preserved\. The resulting symmetry group isE𝐠\(n\)=O𝐠\(n\)⋉ℝnE\_\{\\mathbf\{g\}\}\(n\)=O\_\{\\mathbf\{g\}\}\(n\)\\ltimes\\mathbb\{R\}^\{n\}\. The learnable relaxation coefficient is initialized at zero\.
## 6Controlled Symmetry Relaxation
FullE\(n\)E\(n\)\-equivariance is a natural inductive bias when the system has no preferred direction\. In many physical settings, however, external structure such as gravity or background flow introduces a distinguished direction and thereby reduces the symmetry of the problem\([Smidt et al\., 2021](https://arxiv.org/html/2608.28853#bib.bib15);[Gibb et al\., 2024](https://arxiv.org/html/2608.28853#bib.bib14);[Weidinger et al\., 2017](https://arxiv.org/html/2608.28853#bib.bib13);[Baek et al\., 2017](https://arxiv.org/html/2608.28853#bib.bib12)\)\. ESNN accommodates this setting by introducing an orientation\-dependent scalar into the edge context\. The spatial transport remains unchanged, while its scalar coefficients can now depend on how an edge is oriented relative to the preferred direction\.
Symmetry\-Relaxed Edge Context\.When symmetry relaxation is enabled, let𝐠∈ℝn\\mathbf\{g\}\\in\\mathbb\{R\}^\{n\}denote a global preferred direction andλ∈ℝ\\lambda\\in\\mathbb\{R\}a learnable relaxation coefficient initialized at zero\. The direction𝐠\\mathbf\{g\}may be prescribed or learned as a global model parameter; in both cases, it is treated as fixed in the ambient frame when the input geometry is transformed\. For𝐠≠𝟎\\mathbf\{g\}\\neq\\mathbf\{0\}and a directed interactionj→ij\\rightarrow i, we define the signed projection:
zij𝐠=⟨𝐫ij,𝐠⟩z^\{\\mathbf\{g\}\}\_\{ij\}=\\left\\langle\\mathbf\{r\}\_\{ij\},\\mathbf\{g\}\\right\\rangle\(41\)Unlike the invariant quantities in Equation[28](https://arxiv.org/html/2608.28853#S5.E28), this scalar records whether an edge is aligned or opposed to the preferred direction\. We augment the edge context as:
𝐳ijrelaxed=\[𝐳ij,λ⟨𝐫ij,𝐠⟩\]\\mathbf\{z\}^\{\\mathrm\{relaxed\}\}\_\{ij\}=\\left\[\\mathbf\{z\}\_\{ij\},\\;\\lambda\\left\\langle\\mathbf\{r\}\_\{ij\},\\mathbf\{g\}\\right\\rangle\\right\]\(42\)where𝐳ij\\mathbf\{z\}\_\{ij\}is the invariant edge context from Equation[28](https://arxiv.org/html/2608.28853#S5.E28)\. The added scalar can influence transport coefficients, channel gates, and aggregation weights through the same edge networks used by the fully equivariant layer\. Atλ=0\\lambda=0, its contribution vanishes exactly and the originalE\(n\)E\(n\)\-equivariant model is recovered\.
###### Theorem 6\.1\(Stabilizer Subequivariance\)\.
Let𝐠≠𝟎\\mathbf\{g\}\\neq\\mathbf\{0\}be held fixed in the ambient coordinate frame and define its stabilizer:
O𝐠\(n\)=\{Q∈O\(n\):Q𝐠=𝐠\}O\_\{\\mathbf\{g\}\}\(n\)=\\left\\\{Q\\in O\(n\)\\;:\\;Q\\mathbf\{g\}=\\mathbf\{g\}\\right\\\}\(43\)Forλ≠0\\lambda\\neq 0, an ESNN layer conditioned on Equation[42](https://arxiv.org/html/2608.28853#S6.E42)remains equivariant under translations and under everyQ∈O𝐠\(n\)Q\\in O\_\{\\mathbf\{g\}\}\(n\)\. Hence the guaranteed equivariance group is:
E𝐠\(n\)=O𝐠\(n\)⋉ℝnE\_\{\\mathbf\{g\}\}\(n\)=O\_\{\\mathbf\{g\}\}\(n\)\\ltimes\\mathbb\{R\}^\{n\}\(44\)Equivariance under orthogonal transformations outsideO𝐠\(n\)O\_\{\\mathbf\{g\}\}\(n\)is not enforced by the architecture\. Forλ=0\\lambda=0, the directional conditioning vanishes and fullE\(n\)E\(n\)\-equivariance is recovered\.
The result follows from the fact that the signed projection is unchanged by any transformation that preserves𝐠\\mathbf\{g\}\. ForQ∈O𝐠\(n\)Q\\in O\_\{\\mathbf\{g\}\}\(n\):
⟨Q𝐫ij,𝐠⟩=⟨𝐫ij,Q⊤𝐠⟩=⟨𝐫ij,𝐠⟩\\left\\langle Q\\mathbf\{r\}\_\{ij\},\\mathbf\{g\}\\right\\rangle=\\left\\langle\\mathbf\{r\}\_\{ij\},Q^\{\\top\}\\mathbf\{g\}\\right\\rangle=\\left\\langle\\mathbf\{r\}\_\{ij\},\\mathbf\{g\}\\right\\rangle\(45\)so the relaxed edge context remains invariant under the stabilizer of𝐠\\mathbf\{g\}\. Together with the covariant spatial operators of Section[4](https://arxiv.org/html/2608.28853#S4), this gives the stated subgroup equivariance\. A proof for the complete ESNN layer is provided in Appendix[D\.7](https://arxiv.org/html/2608.28853#A4.SS7)\.
Symmetry Prior\.FullE\(n\)E\(n\)\-equivariance remains the default ESNN setting\. When symmetry relaxation is disabled, no preferred\-direction pathway is present\. When it is enabled, the relaxation coefficients are initialized atλ=0\\lambda=0, so the directional contribution still vanishes exactly at initialization\. Training can then activate this pathway by movingλ\\lambdaaway from zero, whileλ=0\\lambda=0always recovers the fully equivariant regime\.
## 7Experiments
We organize the empirical study around four questions that progressively probe the role of geometric transport and symmetry in ESNN\.Q1:Does richer equivariant transport improve dynamics prediction when fullE\(3\)E\(3\)symmetry is the correct inductive bias?Q2:Can ESNN exploit a known reduction in symmetry, or recover the corresponding preferred direction directly from data?Q3:How does geometric transport behave in mesh\-based physical systems with directional fields, irregular geometry, and long\-horizon dynamics?Q4:Does the same framework remain robust to unseen rotations when transferred beyond physical simulation to point\-cloud classification? We study these questions on charged and gravity\-augmented N\-body dynamics\([Kipf et al\., 2018](https://arxiv.org/html/2608.28853#bib.bib35)\), three physical\-simulation benchmarks from MeshGraphNets\([Pfaff et al\., 2021](https://arxiv.org/html/2608.28853#bib.bib32)\), and ModelNet40\([Wu et al\., 2015](https://arxiv.org/html/2608.28853#bib.bib34)\)\. We additionally evaluate molecular\-property prediction on QM9 as a complementary test of the same transport mechanisms on invariant graph\-level targets \(Appendix[E\.5](https://arxiv.org/html/2608.28853#A5.SS5)\)\.
### 7\.1Particle Dynamics
#### Q1: Charged N\-Body Dynamics\.
We first consider the standard charged N\-body benchmark, where five particles interact through attractive or repulsive Coulomb forces\. From the positions, velocities, and charges at an observed time, the model predicts the particle coordinates at a later time, with performance measured by mean squared error \(MSE\)\. The underlying dynamics retain full Euclidean symmetry, so both EGNN and ESNN operate under the sameE\(3\)E\(3\)symmetry prior\. This setting therefore isolates the effect of enriching the geometric transport without changing the assumed symmetry of the problem\. Table[2](https://arxiv.org/html/2608.28853#S7.T2)shows that every ESNN variant improves over EGNN\. Even ESNN\-Id reduces the MSE from0\.00710\.0071to0\.00600\.0060, while learned transport provides a further gain\. ESNN\-Ortho achieves the best result at0\.00510\.0051, followed closely by ESNN\-RadTan at0\.00520\.0052and ESNN\-Diag at0\.00540\.0054\. The best model reduces the error by approximately28%28\\%relative to EGNN\. The improvement beyond Identity Transport suggests that explicitly learning how vector information is transformed across edges can strengthen first\-order equivariant message passing even when fullE\(3\)E\(3\)symmetry is already the appropriate inductive bias\.
#### Q2: Gravity\-Augmented N\-Body Dynamics\.
We next introduce a uniform gravitational field into the charged N\-body system, adding the acceleration𝐚𝐠=\(0,0,−9\.81\)⊤\\mathbf\{a\}\_\{\\mathbf\{g\}\}=\(0,0,\-9\.81\)^\{\\top\}to the pairwise Coulomb dynamics\. This selects the preferred direction𝐠^true=\(0,0,−1\)⊤\\widehat\{\\mathbf\{g\}\}\_\{\\mathrm\{true\}\}=\(0,0,\-1\)^\{\\top\}and breaks rotational symmetry while preserving translations, providing a direct test of the controlled symmetry relaxation introduced in Section[6](https://arxiv.org/html/2608.28853#S6)\. As before, the model predicts future particle coordinates from an observed state and is evaluated using coordinate MSE\. We compare three matched settings\.*None*receives no preferred direction and remains fullyE\(3\)E\(3\)\-equivariant\.*Fixed*is given the true gravity direction, while*Learned*uses a single trainable global vector whose orientation must be inferred from the dynamics\. In the two symmetry\-relaxed settings, the directional feature enters throughλ⟨𝐫ij,𝐠⟩\\lambda\\langle\\mathbf\{r\}\_\{ij\},\\mathbf\{g\}\\rangle, with the relaxation coefficients initialized at zero\. We therefore also reportmaxℓ\|λg\(ℓ\)\|‖𝐠‖2\\max\_\{\\ell\}\|\\lambda\_\{g\}^\{\(\\ell\)\}\|\\\|\\mathbf\{g\}\\\|\_\{2\}, which measures the effective scale of the directional pathway and indicates whether this initially inactive signal is used after training\. Table[2](https://arxiv.org/html/2608.28853#S7.T2)shows a clear gap between the fully equivariant and symmetry\-relaxed models\. Across all transport families, both*Fixed*and*Learned*reduce the MSE from approximately0\.100\.10–0\.130\.13to around0\.020\.02, while the nonzero directional scales confirm that the relaxed pathway is actively used\. The learned setting closely matches the fixed\-direction setting despite receiving no prior information about the gravity axis\. To determine whether the learned vector also recovers the correct physical direction, we report the sign\-invariant alignmentAg=\|⟨𝐠^,𝐠^true⟩\|A\_\{g\}=\|\\langle\\widehat\{\\mathbf\{g\}\},\\widehat\{\\mathbf\{g\}\}\_\{\\mathrm\{true\}\}\\rangle\|\. HereAg=1A\_\{g\}=1denotes perfect alignment with the gravity axis up to sign, whereasAg=0A\_\{g\}=0corresponds to an orthogonal direction\. The learned models achieveAg=1A\_\{g\}=1in every reported case\. Together, these results show that ESNN can both benefit from the appropriate reduction in symmetry and recover the associated symmetry\-breaking axis directly from data\.
Table 2:*N\-body dynamics benchmarks\.*Mean squared error \(MSE\) for future\-position prediction on the charged\-particle and gravity\-augmented systems\. Charged N\-body baselines follow[Satorras et al\. \(2021\)](https://arxiv.org/html/2608.28853#bib.bib1)\. ESNN models are shown inbold, with the three best charged N\-body results highlighted asFirst,Second, andThird\. Gravity results are mean±\\pmstandard deviation over five runs\. The quantitymaxℓ\|λg\(ℓ\)\|‖𝐠‖2\\max\_\{\\ell\}\|\\lambda\_\{g\}^\{\(\\ell\)\}\|\\\|\\mathbf\{g\}\\\|\_\{2\}measures the effective scale of the directional pathway, whileAg=\|⟨𝐠^,𝐠^true⟩\|A\_\{g\}=\|\\langle\\widehat\{\\mathbf\{g\}\},\\widehat\{\\mathbf\{g\}\}\_\{\\mathrm\{true\}\}\\rangle\|measures sign\-invariant alignment with the gravity axis\. For*Fixed*,Ag=1A\_\{g\}=1by construction; for*Learned*, it is measured from the inferred direction\. Lower MSE and higherAgA\_\{g\}are better\.\(a\)*N\-body*MethodMSE↓\\downarrowLinear0\.0819SE\(3\) Transformer0\.0244Tensor Field Network0\.0155Graph Neural Network0\.0107Radial Field0\.0104EGNN0\.0071ESNN\-Id0\.0060ESNN\-Diag0\.0054ESNN\-Ortho0\.0051ESNN\-RadTan0\.0052
\(b\)*N\-body \+ Gravity*ModeTransportMSE↓\\downarrow𝐦𝐚𝐱ℓ\|𝝀𝒈\(ℓ\)\|‖𝐠‖𝟐\\bm\{\\max\_\{\\ell\}\|\\lambda\_\{g\}^\{\(\\ell\)\}\|\\\|\\mathbf\{g\}\\\|\_\{2\}\}𝑨𝒈\\bm\{A\_\{g\}\}↑\\uparrow*None*ESNN\-Diag0\.107160±0\.0345660\.107160\\pm 0\.034566––ESNN\-Ortho0\.125181±0\.0489700\.125181\\pm 0\.048970––ESNN\-RadTan0\.101860±0\.0324440\.101860\\pm 0\.032444––*Learned*ESNN\-Diag0\.020548±0\.0016410\.020548\\pm 0\.0016410\.267400±0\.0704060\.267400\\pm 0\.0704061\.0001\.000ESNN\-Ortho0\.019687±0\.0014520\.019687\\pm 0\.0014520\.400903±0\.0721510\.400903\\pm 0\.0721511\.0001\.000ESNN\-RadTan0\.021276±0\.0033450\.021276\\pm 0\.0033452\.055211±0\.6337432\.055211\\pm 0\.6337431\.0001\.000*Fixed*ESNN\-Diag0\.020838±0\.0026250\.020838\\pm 0\.0026250\.346428±0\.0357240\.346428\\pm 0\.0357241\.0001\.000ESNN\-Ortho0\.019823±0\.0016510\.019823\\pm 0\.0016510\.609956±0\.1770710\.609956\\pm 0\.1770711\.0001\.000ESNN\-RadTan0\.023381±0\.0054140\.023381\\pm 0\.0054141\.379156±0\.3013971\.379156\\pm 0\.3013971\.0001\.000
### 7\.2Mesh\-Based Physical Dynamics
#### Q3: Direction\-Dependent Fields and Mesh Dynamics\.
We next evaluate ESNN on three physical\-simulation benchmarks from MeshGraphNets\([Pfaff et al\., 2021](https://arxiv.org/html/2608.28853#bib.bib32)\):CylinderFlow,DeformingPlate, andAirfoil, covering incompressible flow, structural deformation, and compressible aerodynamics\. These tasks are defined on irregular simulation meshes and combine vector fields with invariant scalar quantities and node types\. Their dynamics contain strong directional structure arising from flow, deformation, spatial gradients, and boundary geometry, making them a natural test bed for the transport mechanisms introduced in Sections[3](https://arxiv.org/html/2608.28853#S3)and[4](https://arxiv.org/html/2608.28853#S4)\. Following[Pfaff et al\. \(2021\)](https://arxiv.org/html/2608.28853#bib.bib32), we report root mean squared error \(RMSE\) for one\-step prediction, 50\-step autoregressive rollout, and rollout over the full trajectory\. Table[3](https://arxiv.org/html/2608.28853#S7.T3)shows that the benefit of geometric transport depends on both the physical system and the prediction horizon\. The clearest gains occur onDeformingPlate, where all nontrivial ESNN variants improve over MeshGraphNets at every horizon\. In particular, ESNN\-RadTan reduces the RMSE from0\.250\.25to0\.080\.08for one\-step prediction, from1\.81\.8to1\.01\.0over 50 steps, and from15\.115\.1to5\.85\.8over the full trajectory\. OnCylinderFlow, ESNN\-Ortho remains close to MeshGraphNets at one step and improves the full\-trajectory error from40\.8840\.88to35\.9435\.94, although MeshGraphNets performs better at the intermediate 50\-step horizon\. The horizon dependence is even more pronounced onAirfoil: ESNN is less accurate for one\-step and 50\-step prediction, but ESNN\-RadTan reduces the full\-trajectory error from1152911529to77877787\. Overall, the mesh benchmarks do not show a uniform advantage across all systems and horizons\. Instead, the strongest gains appear in structural deformation and in selected long\-horizon rollouts, where geometric information must be propagated repeatedly through the evolving state\. This is notable because MeshGraphNets is designed specifically for learned simulation on unstructured meshes, whereas ESNN uses the same general transport framework across all geometric domains considered in this work\.
Table 3:*Mesh\-based physical dynamics\.*RMSE \(×10−3\\times 10^\{\-3\}\) for one\-step prediction, 50\-step autoregressive rollout, and rollout over the complete trajectory onCylinderFlow,DeformingPlate, andAirfoil\. MeshGraphNets results are taken from[Pfaff et al\. \(2021\)](https://arxiv.org/html/2608.28853#bib.bib32)\. The best full\-trajectory result for each benchmark is shown inbold\. A dash denotes a pending ESNN\-Id result\. Lower is better\.CylinderFlowDeformingPlateAirfoilMethod1\-Step50\-StepFull1\-Step50\-StepFull1\-Step50\-StepFullMeshGraphNets2\.34±0\.122\.34\\pm 0\.126\.3±0\.76\.3\\pm 0\.740\.88±7\.240\.88\\pm 7\.20\.25±0\.050\.25\\pm 0\.051\.8±0\.51\.8\\pm 0\.515\.1±4\.015\.1\\pm 4\.0314±36314\\pm 36582±37582\\pm 3711529±120311529\\pm 1203ESNN\-Id3\.9615\.869\.60\.151\.37\.75529991017974ESNN\-Diag2\.8210\.337\.040\.171\.49\.25130961611598ESNN\-Ortho2\.407\.735\.940\.131\.18\.13014472210328ESNN\-RadTan3\.0810\.949\.270\.081\.05\.8258433467787
Table 4:*ModelNet40 classification under rotations\.*Classification accuracy \(%\) under thez/zz/z,z/SO\(3\)z/\\mathrm\{SO\}\(3\), andSO\(3\)/SO\(3\)\\mathrm\{SO\}\(3\)/\\mathrm\{SO\}\(3\)train/test rotation protocols\. Thez/SO\(3\)z/\\mathrm\{SO\}\(3\)setting is the out\-of\-distribution rotation regime: training examples are rotated only around the vertical axis, whereas arbitrary three\-dimensional rotations are encountered at test time\.ΔOOD\\Delta\_\{\\mathrm\{OOD\}\}denotes the absolute change in accuracy betweenz/zz/zandz/SO\(3\)z/\\mathrm\{SO\}\(3\); lower values indicate greater robustness to this train–test rotation shift\. Baseline results follow[Lippmann et al\. \(2024\)](https://arxiv.org/html/2608.28853#bib.bib20)\. Higher classification accuracy is better\.Method𝒛/𝒛\\bm\{z/z\}𝒛/𝐒𝐎\(𝟑\)\\bm\{z/\\mathrm\{SO\}\(3\)\}𝐒𝐎\(𝟑\)/𝐒𝐎\(𝟑\)\\bm\{\\mathrm\{SO\}\(3\)/\\mathrm\{SO\}\(3\)\}𝚫𝐎𝐎𝐃\\bm\{\\Delta\_\{\\mathrm\{OOD\}\}\}↓\\downarrowPointNet85\.919\.674\.766\.3RS\-CNN90\.348\.782\.641\.6DGCNN90\.333\.888\.656\.5RI\-Conv86\.586\.486\.40\.1GC\-Conv89\.089\.189\.20\.1Luo et al\. DGCNN88\.488\.488\.90\.0LGR\-Net90\.990\.991\.10\.0Li et al\. \(w/ TTA\)91\.691\.691\.60\.0CRIN91\.891\.891\.80\.0TFN88\.585\.387\.63\.2VN\-PointNet77\.577\.577\.20\.0VN\-DGCNN89\.589\.590\.20\.0ESNN\-Id84\.68485\.73785\.5751\.1ESNN\-Diag84\.60384\.31984\.4000\.3ESNN\-Ortho84\.92785\.17084\.6840\.2ESNN\-RadTan85\.37384\.64386\.2640\.7
### 7\.3Rotation Generalization
#### Q4: ModelNet40 Rotation Generalization\.
We finally evaluate whether the same geometric transport framework transfers beyond physical dynamics to point\-cloud classification\. ModelNet40 contains CAD models from 40 object categories, represented as point clouds and converted into localkk\-nearest\-neighbor graphs\([Wu et al\., 2015](https://arxiv.org/html/2608.28853#bib.bib34)\)\. Rather than using an architecture specialized for point\-cloud recognition, ESNN processes these graphs with the same general geometric framework used throughout the other experiments\. We therefore use ModelNet40 primarily to assess rotation generalization and cross\-domain transfer\. We report classification accuracy under thez/zz/z,z/SO\(3\)z/\\mathrm\{SO\}\(3\), andSO\(3\)/SO\(3\)\\mathrm\{SO\}\(3\)/\\mathrm\{SO\}\(3\)train/test protocols\. Thez/SO\(3\)z/\\mathrm\{SO\}\(3\)setting provides the most informative robustness test: the model is trained only on rotations around the vertical axis and evaluated on arbitrary three\-dimensional orientations\. Table[4](https://arxiv.org/html/2608.28853#S7.T4)also reportsΔOOD\\Delta\_\{\\mathrm\{OOD\}\}, the absolute change in accuracy between thez/zz/zandz/SO\(3\)z/\\mathrm\{SO\}\(3\)settings, with lower values indicating greater robustness to this rotation shift\. Baseline results follow[Lippmann et al\. \(2024\)](https://arxiv.org/html/2608.28853#bib.bib20)\. Across all three protocols, ESNN maintains accuracies of approximately8484–86%86\\%and changes only marginally under unseen rotations\.ΔOOD\\Delta\_\{\\mathrm\{OOD\}\}is at most1\.11\.1percentage points across the ESNN variants, and is only0\.20\.2points for ESNN\-Ortho\. By contrast, orientation\-sensitive baselines such as PointNet and DGCNN lose more than5050percentage points when arbitrary three\-dimensional rotations are introduced only at test time\. This stability is consistent with the Euclidean symmetry built into ESNN rather than with exposure to the full range of test orientations during training\. Specialized rotation\-robust point\-cloud architectures achieve higher absolute classification accuracy, and several are similarly insensitive to the rotation shift\. ModelNet40 therefore plays a complementary role in our evaluation: it shows that the same matrix\-valued transport framework used for physical dynamics can be transferred to a substantially different geometric domain while retaining robustness to unseen global rotations\.
## 8Conclusions and Limitations
We introduced ESNN, a first\-order Cartesian framework that enriches equivariant message passing by learning structured matrix\-valued transport between neighboring vector features\. Rather than increasing representation order, ESNN places additional geometric flexibility in the edge transport itself\. We showed that, when relative displacement is the only covariant geometric input, linearO\(n\)O\(n\)\-equivariant transport reduces to independent radial and tangential actions, while learned vector features enable richer feature\-conditioned transformations\. We also introduced controlled symmetry relaxation for systems with a preferred ambient direction, recovering fullE\(n\)E\(n\)\-equivariance when the directional pathway is inactive\. Across the experiments, ESNN improves particle dynamics, recovers the gravity axis from data, achieves strong gains on selected mesh\-based dynamics and long\-horizon rollouts, and remains robust to unseen rotations\. The current formulation is limited to scalar and first\-order vector features, and the completeness result applies specifically to displacement\-conditioned linear transport\. More general covariant inputs admit a broader class of spatial operators, while the present symmetry\-relaxation mechanism assumes a single global preferred direction\. Extending ESNN to richer representation types, local or gauge\-aware transport, and more general symmetry\-breaking fields is a natural direction for future work\.
## References
- Aykent and Xia \(2025\)S\. Aykent and T\. XiaGotennet: rethinking efficient 3D equivariant graph neural networks\.InThe Thirteenth International Conference on Learning Representations,Cited by:[§A\.1](https://arxiv.org/html/2608.28853#A1.SS1.p1.1),[Table 5](https://arxiv.org/html/2608.28853#A5.T5),[§1](https://arxiv.org/html/2608.28853#S1.p1.1)\.
- Baeket al\.\(2017\)Y\. Baek, Y\. Kafri, and V\. LecomteDynamical symmetry breaking and phase transitions in driven diffusive systems\.Physical Review Letters118\(3\),pp\. 030604\.External Links:[Link](http://dx.doi.org/10.1103/PhysRevLett.118.030604)Cited by:[§6](https://arxiv.org/html/2608.28853#S6.p1.1)\.
- Barberoet al\.\(2022a\)F\. Barbero, C\. Bodnar, H\. Sáez de Ocáriz Borde, M\. Bronstein, P\. Veličković, and P\. LiòSheaf neural networks with connection laplacians\.arXiv preprint arXiv:2206\.08702\.External Links:[Link](https://arxiv.org/abs/2206.08702)Cited by:[§A\.3](https://arxiv.org/html/2608.28853#A1.SS3.p1.1),[§C\.3](https://arxiv.org/html/2608.28853#A3.SS3.p7.1),[§2](https://arxiv.org/html/2608.28853#S2.p1.1)\.
- Barberoet al\.\(2022b\)F\. Barbero, C\. Bodnar, H\. Sáez de Ocáriz Borde, and P\. LiòSheaf attention networks\.InNeurIPS 2022 Workshop on Symmetry and Geometry in Neural Representations,Cited by:[§A\.3](https://arxiv.org/html/2608.28853#A1.SS3.p1.1)\.
- Batatiaet al\.\(2022a\)I\. Batatia, S\. Batzner, D\. P\. Kovács, A\. Musaelian, G\. N\. C\. Simm, R\. Drautz, C\. Ortner, B\. Kozinsky, and G\. CsányiThe design space of E\(3\)\-equivariant atom\-centered interatomic potentials\.arXiv preprint arXiv:2205\.06643\.Cited by:[§1](https://arxiv.org/html/2608.28853#S1.p2.1)\.
- Batatiaet al\.\(2022b\)I\. Batatia, D\. P\. Kovács, G\. Simm, C\. Ortner, and G\. CsányiMace: higher order equivariant message passing neural networks for fast and accurate force fields\.InAdvances in Neural Information Processing Systems,Vol\.35,pp\. 11423–11436\.Cited by:[§A\.1](https://arxiv.org/html/2608.28853#A1.SS1.p1.1),[§1](https://arxiv.org/html/2608.28853#S1.p2.1),[§2](https://arxiv.org/html/2608.28853#S2.p1.1)\.
- Battiloroet al\.\(2024\)C\. Battiloro, E\. Karaismailoğlu, M\. Tec, G\. Dasoulas, M\. Audirac, and F\. DominiciE\(n\) equivariant topological neural networks\.arXiv preprint arXiv:2405\.15429\.External Links:[Link](https://arxiv.org/abs/2405.15429)Cited by:[§A\.2](https://arxiv.org/html/2608.28853#A1.SS2.p1.1)\.
- Battiloroet al\.\(2023\)C\. Battiloro, Z\. Wang, H\. Riess, P\. Di Lorenzo, and A\. RibeiroTangent bundle convolutional learning: from manifolds to cellular sheaves and back\.arXiv preprint arXiv:2303\.11323\.External Links:[Link](https://arxiv.org/abs/2303.11323)Cited by:[§A\.2](https://arxiv.org/html/2608.28853#A1.SS2.p1.1)\.
- Bodnaret al\.\(2022\)C\. Bodnar, F\. Di Giovanni, B\. P\. Chamberlain, P\. Lió, and M\. M\. BronsteinNeural sheaf diffusion: a topological perspective on heterophily and oversmoothing in GNNs\.InAdvances in Neural Information Processing Systems,Vol\.35,pp\. 18561–18577\.External Links:[Link](https://arxiv.org/abs/2202.04579)Cited by:[§A\.3](https://arxiv.org/html/2608.28853#A1.SS3.p1.1),[§1](https://arxiv.org/html/2608.28853#S1.p2.1),[§2\.1](https://arxiv.org/html/2608.28853#S2.SS1.p4.1),[§2](https://arxiv.org/html/2608.28853#S2.p1.1)\.
- Borgiet al\.\(2025\)A\. Borgi, F\. Silvestri, and P\. LiòPolynomial neural sheaf diffusion: a spectral filtering approach on cellular sheaves\.arXiv preprint arXiv:2512\.00242\.Cited by:[§A\.3](https://arxiv.org/html/2608.28853#A1.SS3.p1.1)\.
- Braithwaiteet al\.\(2024\)L\. Braithwaite, A\. Borgi, G\. Onorato, K\. Tarantelli, F\. Restuccia, F\. Silvestri, and P\. LiòHeterogeneous sheaf neural networks\.arXiv preprint arXiv:2409\.08036\.Cited by:[§A\.3](https://arxiv.org/html/2608.28853#A1.SS3.p1.1)\.
- Brandstetteret al\.\(2022\)J\. Brandstetter, R\. Hesselink, E\. van der Pol, E\. J\. Bekkers, and M\. WellingGeometric and physical quantities improve E\(3\) equivariant message passing\.InInternational Conference on Learning Representations,External Links:[Link](https://arxiv.org/abs/2110.02905)Cited by:[§A\.1](https://arxiv.org/html/2608.28853#A1.SS1.p1.1),[§1](https://arxiv.org/html/2608.28853#S1.p1.1),[§2](https://arxiv.org/html/2608.28853#S2.p1.1)\.
- Cenet al\.\(2024\)J\. Cen, A\. Li, N\. Lin, Y\. R\. Ren, Z\. Wang, and W\. HuangAre high\-degree representations really unnecessary in equivariant graph neural networks?\.arXiv preprint arXiv:2410\.11443\.External Links:[Link](https://arxiv.org/abs/2410.11443)Cited by:[§A\.1](https://arxiv.org/html/2608.28853#A1.SS1.p1.1)\.
- Cohenet al\.\(2019\)T\. S\. Cohen, M\. Weiler, B\. Kicanaoglu, and M\. WellingGauge equivariant convolutional networks and the icosahedral CNN\.InProceedings of the 36th International Conference on Machine Learning,Vol\.97,pp\. 1321–1330\.External Links:[Link](https://arxiv.org/abs/1902.04615)Cited by:[§A\.2](https://arxiv.org/html/2608.28853#A1.SS2.p1.1)\.
- Duet al\.\(2023\)Y\. Du, L\. Wang, D\. Feng, G\. Wang, S\. Ji, C\. P\. Gomes, Z\. Ma,et al\.A new perspective on building efficient and expressive 3D equivariant graph neural networks\.Advances in Neural Information Processing Systems36,pp\. 66647–66674\.Cited by:[§1](https://arxiv.org/html/2608.28853#S1.p1.1)\.
- Dutaet al\.\(2023\)I\. Duta, G\. Cassarà, F\. Silvestri, and P\. LiòSheaf hypergraph networks\.Advances in Neural Information Processing Systems36,pp\. 12087–12099\.Cited by:[§A\.3](https://arxiv.org/html/2608.28853#A1.SS3.p1.1)\.
- Duvalet al\.\(2023\)A\. Duval, V\. Schmidt, A\. Hernández\-García, S\. Miret, F\. D\. Malliaros, Y\. Bengio, and D\. RolnickFAENet: frame averaging equivariant GNN for materials modeling\.InProceedings of the 40th International Conference on Machine Learning,Vol\.202,pp\. 9013–9033\.External Links:[Link](https://proceedings.mlr.press/v202/duval23a.html)Cited by:[§A\.2](https://arxiv.org/html/2608.28853#A1.SS2.p1.1)\.
- Fioriniet al\.\(2025\)S\. Fiorini, H\. Aktas, I\. Duta, S\. Coniglio, P\. Morerio, A\. Del Bue, and P\. LiòSheaves reloaded: a directional awakening\.arXiv preprint arXiv:2506\.02842\.External Links:[Link](https://arxiv.org/abs/2506.02842)Cited by:[§A\.3](https://arxiv.org/html/2608.28853#A1.SS3.p1.1),[§C\.2](https://arxiv.org/html/2608.28853#A3.SS2.p3.2),[§2](https://arxiv.org/html/2608.28853#S2.p1.1),[§5](https://arxiv.org/html/2608.28853#S5.p4.5)\.
- Fuchset al\.\(2020\)F\. B\. Fuchs, D\. E\. Worrall, V\. Fischer, and M\. WellingSE\(3\)\-transformers: 3D roto\-translation equivariant attention networks\.InAdvances in Neural Information Processing Systems,Vol\.33,pp\. 1970–1981\.External Links:[Link](https://arxiv.org/abs/2006.10503)Cited by:[§A\.1](https://arxiv.org/html/2608.28853#A1.SS1.p1.1),[§E\.1](https://arxiv.org/html/2608.28853#A5.SS1.p1.1),[§2](https://arxiv.org/html/2608.28853#S2.p1.1)\.
- Gibbet al\.\(2024\)C\. J\. Gibb, J\. Hobbs, D\. I\. Nikolova, T\. Raistrick, S\. R\. Berrow, A\. Mertelj, N\. Osterman, N\. Sebastián, H\. F\. Gleeson, and R\. J\. MandleSpontaneous symmetry breaking in polar fluids\.Nature Communications15\(1\)\.External Links:[Link](http://dx.doi.org/10.1038/s41467-024-50230-2)Cited by:[§6](https://arxiv.org/html/2608.28853#S6.p1.1)\.
- Hajijet al\.\(2025\)M\. Hajij, L\. Bastian, S\. Osentoski, H\. Kabaria, J\. L\. Davenport, S\. Dawood, B\. Cherukuri, J\. G\. Kocheemoolayil, N\. Shahmansouri, A\. Lew, T\. Papamarkou, and T\. BirdalCopresheaf topological neural networks: a generalized deep learning framework\.InAdvances in Neural Information Processing Systems,Vol\.38\.Cited by:[§A\.3](https://arxiv.org/html/2608.28853#A1.SS3.p1.1),[§C\.2](https://arxiv.org/html/2608.28853#A3.SS2.p4.1),[§2](https://arxiv.org/html/2608.28853#S2.p1.1)\.
- Hanet al\.\(2022\)J\. Han, W\. Huang, H\. Ma, J\. Li, J\. Tenenbaum, and C\. GanLearning physical dynamics with subequivariant graph neural networks\.Advances in Neural Information Processing Systems35,pp\. 26256–26268\.Cited by:[§A\.4](https://arxiv.org/html/2608.28853#A1.SS4.p1.1),[§2](https://arxiv.org/html/2608.28853#S2.p1.1)\.
- Hansen and Gebhart \(2020\)J\. Hansen and T\. GebhartSheaf neural networks\.InNeurIPS 2020 Workshop on TDA and Beyond,External Links:[Link](https://arxiv.org/abs/2012.06333)Cited by:[§A\.3](https://arxiv.org/html/2608.28853#A1.SS3.p1.1),[§1](https://arxiv.org/html/2608.28853#S1.p2.1),[§2\.1](https://arxiv.org/html/2608.28853#S2.SS1.p4.1),[§2](https://arxiv.org/html/2608.28853#S2.p1.1)\.
- Hernandez Caraltet al\.\(2026\)F\. Hernandez Caralt, M\. Gonzàlez i Català, A\. Bazaga, and P\. LiòOn the necessity of learnable sheaf laplacians\.InICLR 2026 Workshop on Geometry\-grounded Representation Learning and Generative Modeling,Note:Tiny Paper TrackExternal Links:[Link](https://arxiv.org/abs/2603.05395)Cited by:[§A\.3](https://arxiv.org/html/2608.28853#A1.SS3.p1.1)\.
- Hofgardet al\.\(2024\)E\. Hofgard, R\. Wang, R\. Walters, and T\. SmidtRelaxed equivariant graph neural networks\.arXiv preprint arXiv:2407\.20471\.External Links:[Link](https://arxiv.org/abs/2407.20471)Cited by:[§A\.4](https://arxiv.org/html/2608.28853#A1.SS4.p1.1),[§2](https://arxiv.org/html/2608.28853#S2.p1.1)\.
- Jinget al\.\(2021\)B\. Jing, S\. Eismann, P\. N\. Soni, and R\. O\. DrorEquivariant graph neural networks for 3D macromolecular structure\.arXiv preprint arXiv:2106\.03843\.External Links:[Link](https://arxiv.org/abs/2106.03843)Cited by:[§A\.1](https://arxiv.org/html/2608.28853#A1.SS1.p1.1),[§2](https://arxiv.org/html/2608.28853#S2.p1.1)\.
- Kabaet al\.\(2023\)S\. Kaba, A\. K\. Mondal, Y\. Zhang, Y\. Bengio, and S\. RavanbakhshEquivariance with learned canonicalization functions\.InProceedings of the 40th International Conference on Machine Learning,Vol\.202,pp\. 15546–15566\.External Links:[Link](https://proceedings.mlr.press/v202/kaba23a.html)Cited by:[§A\.2](https://arxiv.org/html/2608.28853#A1.SS2.p1.1)\.
- Kipfet al\.\(2018\)T\. Kipf, E\. Fetaya, K\. Wang, M\. Welling, and R\. ZemelNeural relational inference for interacting systems\.InInternational conference on machine learning,pp\. 2688–2697\.Cited by:[§7](https://arxiv.org/html/2608.28853#S7.p1.1)\.
- Kofinaset al\.\(2023\)M\. Kofinas, E\. J\. Bekkers, N\. S\. Nagaraja, and E\. GavvesLatent field discovery in interacting dynamical systems with neural fields\.Advances in Neural Information Processing Systems \(NeurIPS\)\.External Links:[Link](https://arxiv.org/abs/2310.20679)Cited by:[§A\.4](https://arxiv.org/html/2608.28853#A1.SS4.p1.1)\.
- Kovačet al\.\(2024\)V\. Kovač, E\. J\. Bekkers, P\. Liò, and F\. EijkelboomE\(n\) equivariant message passing cellular networks\.arXiv preprint arXiv:2406\.03145\.External Links:[Link](https://arxiv.org/abs/2406.03145)Cited by:[§A\.2](https://arxiv.org/html/2608.28853#A1.SS2.p1.1)\.
- Levyet al\.\(2023\)D\. Levy, S\. Kaba, C\. Gonzales, S\. Miret, and S\. RavanbakhshUsing multiple vector channels improves E\(n\)\-equivariant graph neural networks\.arXiv preprint arXiv:2309\.03139\.External Links:[Link](https://arxiv.org/abs/2309.03139)Cited by:[§A\.1](https://arxiv.org/html/2608.28853#A1.SS1.p1.1)\.
- Liet al\.\(2025\)D\. Li, S\. Arya, and R\. GhristLearning from frustration: torsor CNNs on graphs\.arXiv preprint arXiv:2510\.23288\.External Links:[Link](https://arxiv.org/abs/2510.23288)Cited by:[§A\.2](https://arxiv.org/html/2608.28853#A1.SS2.p1.1)\.
- Liaoet al\.\(2024\)Y\. Liao, B\. M\. Wood, A\. Das, and T\. SmidtEquiformerv2: improved equivariant transformer for scaling to higher\-degree representations\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 39282–39309\.External Links:[Link](https://openreview.net/forum?id=mCOBKZmrzD)Cited by:[§A\.1](https://arxiv.org/html/2608.28853#A1.SS1.p1.1),[§1](https://arxiv.org/html/2608.28853#S1.p1.1),[§2](https://arxiv.org/html/2608.28853#S2.p1.1)\.
- Lippmannet al\.\(2024\)P\. Lippmann, G\. Gerhartz, R\. Remme, and F\. A\. HamprechtBeyond canonicalization: how tensorial messages improve equivariant message passing\.arXiv preprint arXiv:2405\.15389\.External Links:[Link](https://arxiv.org/abs/2405.15389)Cited by:[§A\.2](https://arxiv.org/html/2608.28853#A1.SS2.p1.1),[§7\.3](https://arxiv.org/html/2608.28853#S7.SS3.SSS0.Px1.p1.1),[Table 4](https://arxiv.org/html/2608.28853#S7.T4)\.
- Maruyama \(2026\)Y\. MaruyamaFoundations of equivariant deep learning: unifying graph and sheaf neural networks\.arXiv preprint arXiv:2607\.03798\.External Links:[Link](https://arxiv.org/abs/2607.03798)Cited by:[§A\.2](https://arxiv.org/html/2608.28853#A1.SS2.p1.1),[§A\.3](https://arxiv.org/html/2608.28853#A1.SS3.p1.1)\.
- Penget al\.\(2026\)Y\. Peng, J\. Dong, Y\. Zeng, H\. Li, C\. Ju, H\. Feng, D\. Taha, A\. Wienhard, and K\. XiaSheaf neural networks on SPD manifolds: second\-order geometric representation learning\.arXiv preprint arXiv:2604\.20308\.External Links:[Link](https://arxiv.org/abs/2604.20308)Cited by:[§A\.3](https://arxiv.org/html/2608.28853#A1.SS3.p1.1)\.
- Pfaffet al\.\(2021\)T\. Pfaff, M\. Fortunato, A\. Sanchez\-Gonzalez, and P\. W\. BattagliaLearning mesh\-based simulation with graph networks\.International Conference on Learning Representations \(ICLR\)\.External Links:[Link](https://arxiv.org/abs/2010.03409)Cited by:[§E\.4](https://arxiv.org/html/2608.28853#A5.SS4.p1.1),[§7\.2](https://arxiv.org/html/2608.28853#S7.SS2.SSS0.Px1.p1.1),[Table 3](https://arxiv.org/html/2608.28853#S7.T3),[§7](https://arxiv.org/html/2608.28853#S7.p1.1)\.
- Punyet al\.\(2022\)O\. Puny, M\. Atzmon, H\. Ben\-Hamu, E\. J\. Smith, H\. Maron, and Y\. LipmanFrame averaging for invariant and equivariant network design\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=zIU4k8hK3H)Cited by:[§A\.2](https://arxiv.org/html/2608.28853#A1.SS2.p1.1)\.
- Ramakrishnanet al\.\(2014\)R\. Ramakrishnan, P\. O\. Dral, M\. Rupp, and O\. A\. Von LilienfeldQuantum chemistry structures and properties of 134 kilo molecules\.Scientific data1\(1\),pp\. 1–7\.Cited by:[§E\.5](https://arxiv.org/html/2608.28853#A5.SS5.p1.1)\.
- Ribeiroet al\.\(2025\)A\. Ribeiro, A\. L\. Tenório, J\. Belieni, A\. H\. Souza, and D\. MesquitaCooperative sheaf neural networks\.arXiv preprint arXiv:2507\.00647\.External Links:[Link](https://arxiv.org/abs/2507.00647)Cited by:[§A\.3](https://arxiv.org/html/2608.28853#A1.SS3.p1.1),[§C\.2](https://arxiv.org/html/2608.28853#A3.SS2.p3.2),[§2](https://arxiv.org/html/2608.28853#S2.p1.1),[§5](https://arxiv.org/html/2608.28853#S5.p4.5)\.
- Satorraset al\.\(2021\)V\. G\. Satorras, E\. Hoogeboom, and M\. WellingE\(n\) equivariant graph neural networks\.InInternational Conference on Machine Learning,pp\. 9323–9332\.External Links:[Link](https://arxiv.org/abs/2102.09844)Cited by:[§A\.1](https://arxiv.org/html/2608.28853#A1.SS1.p1.1),[§E\.1](https://arxiv.org/html/2608.28853#A5.SS1.p1.1),[§E\.1](https://arxiv.org/html/2608.28853#A5.SS1.p2.1),[§E\.5](https://arxiv.org/html/2608.28853#A5.SS5.p2.1),[§1](https://arxiv.org/html/2608.28853#S1.p1.1),[§1](https://arxiv.org/html/2608.28853#S1.p2.1),[§2](https://arxiv.org/html/2608.28853#S2.p1.1),[Table 2](https://arxiv.org/html/2608.28853#S7.T2)\.
- Schüttet al\.\(2021\)K\. Schütt, O\. Unke, and M\. GasteggerEquivariant message passing for the prediction of tensorial properties and molecular spectra\.InInternational Conference on Machine Learning,pp\. 9377–9388\.Cited by:[§A\.1](https://arxiv.org/html/2608.28853#A1.SS1.p1.1),[§1](https://arxiv.org/html/2608.28853#S1.p1.1),[§1](https://arxiv.org/html/2608.28853#S1.p2.1),[§2](https://arxiv.org/html/2608.28853#S2.p1.1)\.
- Simeon and de Fabritiis \(2023\)G\. Simeon and G\. de FabritiisTensorNet: cartesian tensor representations for efficient learning of molecular potentials\.InAdvances in Neural Information Processing Systems,Vol\.36,pp\. 56877–56893\.Cited by:[§A\.1](https://arxiv.org/html/2608.28853#A1.SS1.p1.1),[§1](https://arxiv.org/html/2608.28853#S1.p1.1)\.
- Smidtet al\.\(2021\)T\. Smidt, M\. Geiger, and B\. MillerFinding symmetry breaking order parameters with euclidean neural networks\.Physical Review Research3\(1\),pp\. L012002\.External Links:[Link](https://link.aps.org/doi/10.1103/PhysRevResearch.3.L012002)Cited by:[§6](https://arxiv.org/html/2608.28853#S6.p1.1)\.
- Thomaset al\.\(2018\)N\. Thomas, T\. Smidt, S\. Kearnes, L\. Yang, L\. K\. Li, K\. Kohlhoff, and P\. RileyTensor field networks: rotation\-and translation\-equivariant neural networks for 3D point clouds\.arXiv preprint arXiv:1802\.08219\.Cited by:[§A\.1](https://arxiv.org/html/2608.28853#A1.SS1.p1.1),[§1](https://arxiv.org/html/2608.28853#S1.p2.1),[§2](https://arxiv.org/html/2608.28853#S2.p1.1)\.
- Wanget al\.\(2024a\)R\. Wang, E\. Hofgard, H\. Gao, R\. Walters, and T\. E\. SmidtDiscovering symmetry breaking in physical systems with relaxed group convolution\.InProceedings of the 41st International Conference on Machine Learning,External Links:[Link](https://arxiv.org/abs/2310.02299)Cited by:[§A\.4](https://arxiv.org/html/2608.28853#A1.SS4.p1.1)\.
- Wanget al\.\(2022a\)R\. Wang, R\. Walters, and R\. YuApproximately equivariant networks for imperfectly symmetric dynamics\.InProceedings of the 39th International Conference on Machine Learning,pp\. 23078–23091\.External Links:[Link](https://arxiv.org/abs/2201.11969)Cited by:[§A\.4](https://arxiv.org/html/2608.28853#A1.SS4.p1.1),[§2](https://arxiv.org/html/2608.28853#S2.p1.1)\.
- Wanget al\.\(2022b\)R\. Wang, R\. Walters, and R\. YuRelaxing equivariance constraints with non\-stationary continuous filters\.InAdvances in Neural Information Processing Systems,Vol\.35,pp\. 25776–25790\.External Links:[Link](https://papers.neurips.cc/paper_files/paper/2022/file/dafd116ac8c735f149558b79fd48e090-Paper-Conference.pdf)Cited by:[§A\.4](https://arxiv.org/html/2608.28853#A1.SS4.p1.1)\.
- Wanget al\.\(2024b\)Y\. Wang, S\. Li, X\. He, M\. Li, Z\. Wang, N\. Zheng, B\. Shao, T\. Liu, and T\. WangViSNet: an equivariant geometry\-enhanced graph neural network with vector\-scalar interactive message passing for molecules\.Nature Communications15\(1\),pp\. 228\.Cited by:[§A\.1](https://arxiv.org/html/2608.28853#A1.SS1.p1.1),[§1](https://arxiv.org/html/2608.28853#S1.p1.1)\.
- Weidingeret al\.\(2017\)S\. A\. Weidinger, M\. Heyl, A\. Silva, and M\. KnapDynamical quantum phase transitions in systems with continuous symmetry breaking\.Physical Review B96,pp\. 134313\.External Links:[Link](https://link.aps.org/doi/10.1103/PhysRevB.96.134313)Cited by:[§6](https://arxiv.org/html/2608.28853#S6.p1.1)\.
- Weileret al\.\(2021\)M\. Weiler, P\. Forré, E\. Verlinde, and M\. WellingCoordinate independent convolutional networks–isometry and gauge equivariant convolutions on riemannian manifolds\.arXiv preprint arXiv:2106\.06020\.Cited by:[§A\.2](https://arxiv.org/html/2608.28853#A1.SS2.p1.1)\.
- Wuet al\.\(2015\)Z\. Wu, S\. Song, A\. Khosla, F\. Yu, L\. Zhang, X\. Tang, and J\. Xiao3d shapenets: a deep representation for volumetric shapes\.InProceedings of the IEEE conference on computer vision and pattern recognition,pp\. 1912–1920\.Cited by:[§E\.3](https://arxiv.org/html/2608.28853#A5.SS3.p1.1),[§7\.3](https://arxiv.org/html/2608.28853#S7.SS3.SSS0.Px1.p1.1),[§7](https://arxiv.org/html/2608.28853#S7.p1.1)\.
Equivariant Sheaf Neural Networks: Learning Geometric Transport on Graphs
Supplementary Material
## Appendix AExtended Related Work
This appendix extends the discussion in Section[2](https://arxiv.org/html/2608.28853#S2)by placing ESNN within three closely related areas: equivariant geometric learning, sheaf\-based representation learning, and methods that relax or reduce symmetry in the presence of external structure\. We focus on approaches that act on geometric vector features, learn edge\-dependent transformations between local representations, or adapt the symmetry imposed on message passing\.
### A\.1Cartesian and Steerable Equivariant Models
E\(n\)E\(n\)\-equivariant graph networks differ primarily in the representations they propagate and the operations used to couple neighboring features\. A broad family of Cartesian architectures works directly with invariant scalars and low\-order covariant features\. EGNN\([Satorras et al\., 2021](https://arxiv.org/html/2608.28853#bib.bib1)\)constructs coordinate updates from invariant edge functions multiplying relative displacement vectors, while PaiNN\([Schütt et al\., 2021](https://arxiv.org/html/2608.28853#bib.bib7)\)and GVP\-based models\([Jing et al\., 2021](https://arxiv.org/html/2608.28853#bib.bib18)\)maintain scalar and vector channels and couple them through equivariant operations\. Subsequent approaches have enriched this low\-order design space through vector–scalar interactions\([Wang et al\., 2024b](https://arxiv.org/html/2608.28853#bib.bib4)\), multiple vector channels\([Levy et al\., 2023](https://arxiv.org/html/2608.28853#bib.bib39)\), and Cartesian tensor representations\([Simeon and de Fabritiis, 2023](https://arxiv.org/html/2608.28853#bib.bib3)\)\. More recent architectures such as GotenNet\([Aykent and Xia, 2025](https://arxiv.org/html/2608.28853#bib.bib6)\)similarly pursue the trade\-off between geometric expressivity and computational efficiency using geometric tensor representations without explicit Clebsch–Gordan contractions\. A complementary family of architectures uses steerable representations transforming under irreducible representations of the rotation group\. Tensor Field Networks\([Thomas et al\., 2018](https://arxiv.org/html/2608.28853#bib.bib9)\), SE\(3\)\-Transformers\([Fuchs et al\., 2020](https://arxiv.org/html/2608.28853#bib.bib19)\), SEGNN\([Brandstetter et al\., 2022](https://arxiv.org/html/2608.28853#bib.bib8)\), MACE\([Batatia et al\., 2022b](https://arxiv.org/html/2608.28853#bib.bib11)\), and EquiformerV2\([Liao et al\., 2024](https://arxiv.org/html/2608.28853#bib.bib5)\)construct angular interactions using spherical harmonics and tensor\-product couplings between representation types\. Increasing the maximum representation degree provides access to richer angular structure, but also increases the representation and contraction costs\. Recent work has therefore investigated how much high\-degree information is necessary and how it can be processed more efficiently\([Cen et al\., 2024](https://arxiv.org/html/2608.28853#bib.bib36)\)\. ESNN addresses a different point in this design space: it deliberately remains within first\-order Cartesian vector features and increases expressivity through matrix\-valued edge transport rather than higher representation degree\.
### A\.2Gauges and Topological Domains
Another related strategy obtains equivariance by choosing or averaging local reference frames\. Frame Averaging\([Puny et al\., 2022](https://arxiv.org/html/2608.28853#bib.bib25)\)constructs exactly equivariant models by averaging a backbone over an equivariant frame, with FAENet\([Duval et al\., 2023](https://arxiv.org/html/2608.28853#bib.bib26)\)adapting this idea to atomistic modeling\. Learned canonicalization\([Kaba et al\., 2023](https://arxiv.org/html/2608.28853#bib.bib22)\)instead predicts a canonical representative before applying a generic backbone, while more recent work has investigated the advantages of retaining tensorial messages beyond canonicalized scalar representations\([Lippmann et al\., 2024](https://arxiv.org/html/2608.28853#bib.bib20)\)\. These approaches resolve orientation dependence through a choice or averaging of frames\. Gauge\-equivariant learning addresses a related but mathematically distinct problem\. Gauge\-equivariant CNNs\([Cohen et al\., 2019](https://arxiv.org/html/2608.28853#bib.bib23)\)and coordinate\-independent convolutions\([Weiler et al\., 2021](https://arxiv.org/html/2608.28853#bib.bib31)\)describe signals expressed in independently chosen local frames and require the network to transform consistently under changes of those frames\. Tangent\-bundle convolutional networks connect this perspective to connection Laplacians and cellular sheaves by discretizing vector\-field diffusion on Riemannian manifolds\([Battiloro et al\., 2023](https://arxiv.org/html/2608.28853#bib.bib40)\)\. Torsor CNNs further extend local group\-valued transport to arbitrary graphs through edge potentials relating neighboring frames\([Li et al\., 2025](https://arxiv.org/html/2608.28853#bib.bib21)\)\. These constructions concern*local gauge equivariance*, whereas the core ESNN architecture enforces equivariance under a common ambientO\(n\)O\(n\)transformation\. The Orthogonal Transport of ESNN is therefore connection\-style rather than a general gauge\-equivariant construction\. Finally, recent work has combined Euclidean equivariance with topological domains directly\.E\(n\)E\(n\)\-Equivariant Topological Neural Networks\([Battiloro et al\., 2024](https://arxiv.org/html/2608.28853#bib.bib41)\)extend equivariant message passing from graphs to combinatorial complexes, whileE\(n\)E\(n\)\-equivariant message\-passing cellular networks\([Kovač et al\., 2024](https://arxiv.org/html/2608.28853#bib.bib42)\)develop related constructions on cellular structures\. These approaches enrich the*combinatorial domain*through higher\-order cells\. ESNN is complementary: it remains graph\-based in the present work and enriches the*edge transport*between geometric feature spaces through a sheaf\-inspired construction\. A complementary theoretical perspective is provided by\([Maruyama, 2026](https://arxiv.org/html/2608.28853#bib.bib49)\), who formulate order\-equivariant neural networks on equivariant vector bundles over face posets and show that graph and sheaf layers arise within a common order\-equivariant framework\. Their equivariance acts on the combinatorial indexing structure through poset automorphisms, whereas ESNN considers the continuous Euclidean action on geometric feature fibers and constrains the learned edge transport accordingly\. The two constructions therefore address distinct, compatible symmetry structures: order equivariance of the underlying topological domain and ambientO\(n\)O\(n\)\-equivariance of geometric transport, respectively\.
### A\.3Sheaf Learning and Directionality
Cellular sheaves associate local vector spaces with graph cells and relate them through incidence restriction maps\. Early SNNs\([Hansen and Gebhart, 2020](https://arxiv.org/html/2608.28853#bib.bib17)\)used the resulting sheaf Laplacian to generalize graph convolution, while Neural Sheaf Diffusion\([Bodnar et al\., 2022](https://arxiv.org/html/2608.28853#bib.bib16)\)introduced learnable restriction maps and studied how the resulting diffusion affects heterophily, class separation, and oversmoothing\. Together, these works established neural sheaf diffusion as propagation through learned compatibility maps between local feature spaces\. Several subsequent architectures have modified either the structure of the sheaf or the operator acting on it\. Connection\-Laplacian SNNs\([Barbero et al\., 2022a](https://arxiv.org/html/2608.28853#bib.bib38)\)construct orthogonal restriction maps motivated by local tangent\-space alignment, providing an especially relevant precedent for geometric transport\. Sheaf Attention Networks\([Barbero et al\., 2022b](https://arxiv.org/html/2608.28853#bib.bib43)\)introduce attention into sheaf propagation\. Sheaf Hypergraph Networks\([Duta et al\., 2023](https://arxiv.org/html/2608.28853#bib.bib30)\)extend the construction to higher\-order relations, while Heterogeneous Sheaf Neural Networks\([Braithwaite et al\., 2024](https://arxiv.org/html/2608.28853#bib.bib28)\)use heterogeneous stalk and restriction structures to represent typed graphs\. Polynomial Neural Sheaf Diffusion\([Borgi et al\., 2025](https://arxiv.org/html/2608.28853#bib.bib29)\)develops higher\-order spectral filters of the sheaf operator\. Recent work has also begun to explore richer non\-Euclidean stalk geometries, including second\-order representations on SPD manifolds\([Peng et al\., 2026](https://arxiv.org/html/2608.28853#bib.bib48)\)\. Directionality has received increasing attention in the sheaf literature\. Classical sheaf Laplacians are self\-adjoint and therefore do not directly encode independent propagation rules for the two orientations of an edge\. Cooperative Sheaf Neural Networks\([Ribeiro et al\., 2025](https://arxiv.org/html/2608.28853#bib.bib44)\)introduce cellular sheaves on directed graphs together with in\- and out\-degree sheaf Laplacians, while Directed Sheaf Neural Networks\([Fiorini et al\., 2025](https://arxiv.org/html/2608.28853#bib.bib45)\)develop a directed cellular\-sheaf construction and a directional sheaf Laplacian\. These works establish that asymmetric information flow can be incorporated into sheaf\-based learning without forcing the two edge orientations to represent the same interaction\. A related categorical perspective is provided by Copresheaf Topological Neural Networks\([Hajij et al\., 2025](https://arxiv.org/html/2608.28853#bib.bib46)\), which formulate neural architectures in terms of covariant maps between local spaces and their compositions\. More recent categorical formulations have similarly investigated broader notions of equivariance that encompass graph and sheaf neural networks\([Maruyama, 2026](https://arxiv.org/html/2608.28853#bib.bib49)\)\. These approaches are useful for formalizing directed maps and compositional structure, but their notion of equivariance is distinct from the ambient EuclideanE\(n\)E\(n\)symmetry considered here\. Appendix[C](https://arxiv.org/html/2608.28853#A3)develops the formal distinction relevant to ESNN\. Classical incidence maps induce node\-to\-node blocks such asρi→e∗ρj→e\\rho\_\{i\\to e\}^\{\*\}\\rho\_\{j\\to e\}, whereas ESNN directly parameterizes the directed node\-to\-node transport\. Adjoint consistency and full\-fiber orthogonality identify the exact connection\-sheaf specialization; the general directional operator remains a broader connection\-style construction\. This distinction is also relevant in light of recent empirical analyses of sheaf learning\. In particular, identity\-sheaf baselines can perform competitively with learned sheaf operators on several standard heterophilic benchmarks\([Hernandez Caralt et al\., 2026](https://arxiv.org/html/2608.28853#bib.bib47)\), suggesting that learning unrestricted restriction maps is not uniformly beneficial across graph\-learning tasks\. ESNN does not rely on a generic claim that non\-trivial sheaf maps are always advantageous\. Instead, it targets geometric settings in which vector features carry a known Euclidean transformation law and asks how edge\-wise transport can be made both directional and symmetry\-compatible\.
### A\.4Symmetry Relaxation and Subequivariant Inductive Biases
External fields or other environmental structure can reduce the symmetry of an observed system\. Existing approaches address this mismatch by enforcing a known subgroup, learning approximate departures from equivariance, or separating external effects from the equivariant interaction model\. Subequivariant Graph Neural Networks\([Han et al\., 2022](https://arxiv.org/html/2608.28853#bib.bib27)\)take the first approach\. They explicitly incorporate external fields such as gravity and construct message\-passing operations equivariant to the subgroup that preserves the field direction\. This provides an exact physical inductive bias when the external field and the resulting subgroup are known in advance\. ESNN adopts the same stabilizer\-group principle but incorporates the directional signal within the transport framework and controls its contribution through a learnable relaxation coefficient\. A second line of work allows equivariance to be violated gradually\. Non\-stationary continuous filters\([Wang et al\., 2022b](https://arxiv.org/html/2608.28853#bib.bib24)\)introduce learnable departures from weight sharing, while approximately equivariant networks\([Wang et al\., 2022a](https://arxiv.org/html/2608.28853#bib.bib50)\)bias dynamics models toward a symmetry without imposing it exactly\. Relaxed group convolutions\([Wang et al\., 2024a](https://arxiv.org/html/2608.28853#bib.bib51)\)use symmetry\-dependent weights to identify and quantify symmetry breaking in physical data, and Relaxed EGNNs\([Hofgard et al\., 2024](https://arxiv.org/html/2608.28853#bib.bib52)\)extend this idea to continuousE\(3\)E\(3\)\-equivariant graph architectures\. These methods are designed to learn*approximate*deviations from a prescribed group action\. External effects can also be separated from the equivariant interaction model rather than absorbed into its symmetry\. Latent Field Discovery\([Kofinas et al\., 2023](https://arxiv.org/html/2608.28853#bib.bib37)\)decomposes interacting dynamics into an equivariant local interaction and an additional global neural field, allowing spatially varying external effects to be inferred from data\. This provides greater flexibility for unknown fields, including effects that depend on absolute position, but introduces a separate field model alongside the equivariant interaction network\. ESNN follows the first regime at nonzero relaxation: a fixed preferred direction yields exact equivariance to its stabilizer\. The zero\-initialized coefficient controls whether that directional feature is used, with fullE\(n\)E\(n\)\-equivariance recovered exactly at zero\. The construction therefore has a group\-theoretic guarantee distinct from a generic approximately equivariant perturbation and is narrower than a separately learned spatial field\.
## Appendix BAlgebraic Foundations of Equivariant Spatial Transport
This appendix develops the algebraic foundations of the transport maps introduced in Section[3](https://arxiv.org/html/2608.28853#S3)\. We first derive the separable spatial–channel form, then establish completeness when the relative displacement is the only covariant input, and finally clarify how feature\-conditioned operators extend beyond this setting\. Throughout, “linear transport” refers to linearity in the transported vector feature once the local edge context is fixed\.
### B\.1Tensor\-Product Structure of Vector Transport
The vector component of an ESNN stalk is:
ℱvec\(i\)=ℝn⊗ℝcv\\mathcal\{F\}\_\{\\mathrm\{vec\}\}\(i\)=\\mathbb\{R\}^\{n\}\\otimes\\mathbb\{R\}^\{c\_\{v\}\}\(46\)whereℝn\\mathbb\{R\}^\{n\}represents the spatial dimension andℝcv\\mathbb\{R\}^\{c\_\{v\}\}indexes the vector channels\. TheO\(n\)O\(n\)action affects only the spatial component, so after identifying the tensor product withℝn×cv\\mathbb\{R\}^\{n\\times c\_\{v\}\}we have:
𝐕↦Q𝐕,Q∈O\(n\)\\mathbf\{V\}\\mapsto Q\\mathbf\{V\},\\qquad Q\\in O\(n\)\(47\)This separation between spatial and channel dimensions also carries over to linear operators\. In finite dimensions:
End\(ℝn⊗ℝcv\)≅End\(ℝn\)⊗End\(ℝcv\)\\operatorname\{End\}\\left\(\\mathbb\{R\}^\{n\}\\otimes\\mathbb\{R\}^\{c\_\{v\}\}\\right\)\\cong\\operatorname\{End\}\(\\mathbb\{R\}^\{n\}\)\\otimes\\operatorname\{End\}\(\\mathbb\{R\}^\{c\_\{v\}\}\)\(48\)so every linear operator on the vector stalk can be written as a finite sum of separable spatial and channel transformations\.
###### Proposition B\.1\(Separable Expansion of Linear Vector Transport\)\.
Letℒ:ℝn×cv⟶ℝn×cv\\mathcal\{L\}:\\mathbb\{R\}^\{n\\times c\_\{v\}\}\\longrightarrow\\mathbb\{R\}^\{n\\times c\_\{v\}\}be linear\. Then there exist spatial matrices𝐒\(k\)∈ℝn×n\\mathbf\{S\}^\{\(k\)\}\\in\\mathbb\{R\}^\{n\\times n\}and channel matrices𝐌\(k\)∈ℝcv×cv\\mathbf\{M\}^\{\(k\)\}\\in\\mathbb\{R\}^\{c\_\{v\}\\times c\_\{v\}\}such that:
ℒ\(𝐕\)=∑k=1K𝐒\(k\)𝐕𝐌\(k\)\\mathcal\{L\}\(\\mathbf\{V\}\)=\\sum\_\{k=1\}^\{K\}\\mathbf\{S\}^\{\(k\)\}\\mathbf\{V\}\\mathbf\{M\}^\{\(k\)\}\(49\)for some finiteKK\.
###### Proof\.
Equation[48](https://arxiv.org/html/2608.28853#A2.E48)implies that any endomorphism ofℝn⊗ℝcv\\mathbb\{R\}^\{n\}\\otimes\\mathbb\{R\}^\{c\_\{v\}\}can be expressed as a finite sum of elementary tensor\-product operators\. Under the matrix identificationℝn⊗ℝcv≃ℝn×cv\\mathbb\{R\}^\{n\}\\otimes\\mathbb\{R\}^\{c\_\{v\}\}\\simeq\\mathbb\{R\}^\{n\\times c\_\{v\}\}, the standard vectorization identity:
vec\(𝐒𝐕𝐌\)=\(𝐌⊤⊗𝐒\)vec\(𝐕\)\\operatorname\{vec\}\\left\(\\mathbf\{S\}\\mathbf\{V\}\\mathbf\{M\}\\right\)=\\left\(\\mathbf\{M\}^\{\\top\}\\otimes\\mathbf\{S\}\\right\)\\operatorname\{vec\}\(\\mathbf\{V\}\)\(50\)maps each elementary tensor\-product operator to a left spatial action and a right channel action, yielding Equation[49](https://arxiv.org/html/2608.28853#A2.E49)\. ∎
Proposition[B\.1](https://arxiv.org/html/2608.28853#A2.Thmtheorem1)is purely algebraic and does not by itself enforce equivariance\. In ESNN, the spatial and channel matrices depend on the local edge context, so equivariance must be imposed on this dependence\. For a transformed contextQ⋅𝒞ijQ\\\!\\cdot\\\!\\mathcal\{C\}\_\{ij\}, the sufficient component\-wise conditions introduced in Section[3](https://arxiv.org/html/2608.28853#S3)are:
𝐒ij\(k\)\(Q⋅𝒞ij\)=Q𝐒ij\(k\)\(𝒞ij\)Q⊤,𝐌ij\(k\)\(Q⋅𝒞ij\)=𝐌ij\(k\)\(𝒞ij\)\\mathbf\{S\}^\{\(k\)\}\_\{ij\}\(Q\\\!\\cdot\\\!\\mathcal\{C\}\_\{ij\}\)=Q\\mathbf\{S\}^\{\(k\)\}\_\{ij\}\(\\mathcal\{C\}\_\{ij\}\)Q^\{\\top\},\\qquad\\mathbf\{M\}^\{\(k\)\}\_\{ij\}\(Q\\\!\\cdot\\\!\\mathcal\{C\}\_\{ij\}\)=\\mathbf\{M\}^\{\(k\)\}\_\{ij\}\(\\mathcal\{C\}\_\{ij\}\)\(51\)These are precisely the conditions used in Proposition[3\.2](https://arxiv.org/html/2608.28853#S3.Thmtheorem2): the decomposition separates the spatial and channel actions, while the covariance constraints determine which context\-dependent choices preserve compatibility with the ambientO\(n\)O\(n\)action\.
Canonical ESNN Transport𝒯i←j\(𝐕j\)=∑k=1K𝐒ij\(k\)𝐕j𝐌ij\(k\)\\displaystyle\\mathcal\{T\}\_\{i\\leftarrow j\}\(\\mathbf\{V\}\_\{j\}\)=\\sum\_\{k=1\}^\{K\}\{\\color\[rgb\]\{0\.4219,0\.2461,0\.6289\}\\mathbf\{S\}^\{\(k\)\}\_\{ij\}\}\\mathbf\{V\}\_\{j\}\{\\color\[rgb\]\{0\.3555,0\.418,0\.4805\}\\mathbf\{M\}^\{\(k\)\}\_\{ij\}\}𝐕j\\mathbf\{V\}\_\{j\}vector stalk𝐕j∈ℝn×cv\\mathbf\{V\}\_\{j\}\\in\\mathbb\{R\}^\{n\\times c\_\{v\}\}ambient spaceℝn\\mathbb\{R\}^\{n\}\(nnrows\)vector channelsℝcv\\mathbb\{R\}^\{c\_\{v\}\}\(cvc\_\{v\}columns\)𝐒ij\(k\)∈ℝn×n\\mathbf\{S\}^\{\(k\)\}\_\{ij\}\\in\\mathbb\{R\}^\{n\\times n\}Acts on ambient spaceLeft multiplication\(acts on spatial rows\)𝐌ij\(k\)∈ℝcv×cv\\mathbf\{M\}^\{\(k\)\}\_\{ij\}\\in\\mathbb\{R\}^\{c\_\{v\}\\times c\_\{v\}\}Mixes vector channelsRight multiplication\(acts on channel columns\)
Figure 4:*Spatial–channel factorization of ESNN transport\.*Each term acts on𝐕j∈ℝn×cv\\mathbf\{V\}\_\{j\}\\in\\mathbb\{R\}^\{n\\times c\_\{v\}\}along two independent axes:𝐒ij\(k\)\\mathbf\{S\}\_\{ij\}^\{\(k\)\}acts on the ambient spatial dimension by left multiplication, while𝐌ij\(k\)\\mathbf\{M\}\_\{ij\}^\{\(k\)\}acts on the vector channels by right multiplication\. Summing the separable terms gives the canonical transport𝒯i←j\\mathcal\{T\}\_\{i\\leftarrow j\}\. ESNN enforces equivariance by constraining the spatial operators to beO\(n\)O\(n\)\-covariant and the channel maps to be invariant\.
### B\.2Completeness Under Displacement\-Only Conditioning
This subsection establishes the completeness result stated in Theorem[4\.3](https://arxiv.org/html/2608.28853#S4.Thmtheorem3)\. The key assumption is that the relative displacement is the only covariant geometric quantity available to the transport, while any additional conditioning isO\(n\)O\(n\)\-invariant\. Under this restriction, equivariance with respect to the stabilizer of a non\-zero displacement forces the spatial action to separate into radial and tangential components\. The proof below makes this constraint explicit and characterizes the resulting transport class\.
Theorem[4\.3](https://arxiv.org/html/2608.28853#S4.Thmtheorem3)\.Letn≥2n\\geq 2, and consider a linear transport onℝn⊗ℝcv\\mathbb\{R\}^\{n\}\\otimes\\mathbb\{R\}^\{c\_\{v\}\}whose only covariant geometric conditioning variable is a non\-zero displacement𝐫∈ℝn\\mathbf\{r\}\\in\\mathbb\{R\}^\{n\}, with arbitrary additionalO\(n\)O\(n\)\-invariant scalar conditioning\. Every suchO\(n\)O\(n\)\-equivariant transport can be written as:𝒯𝐫,𝝃\(𝐕\)=𝐏∥\(𝐫\)𝐕𝐀𝐫,𝝃\+𝐏⟂\(𝐫\)𝐕𝐁𝐫,𝝃\\mathcal\{T\}\_\{\\mathbf\{r\},\\bm\{\\xi\}\}\(\\mathbf\{V\}\)=\\mathbf\{P\}^\{\\parallel\}\(\\mathbf\{r\}\)\\mathbf\{V\}\\mathbf\{A\}\_\{\\mathbf\{r\},\\bm\{\\xi\}\}\+\\mathbf\{P\}^\{\\perp\}\(\\mathbf\{r\}\)\\mathbf\{V\}\\mathbf\{B\}\_\{\\mathbf\{r\},\\bm\{\\xi\}\}\(52\)for invariant channel endomorphisms𝐀𝐫,𝛏,𝐁𝐫,𝛏∈ℝcv×cv\\mathbf\{A\}\_\{\\mathbf\{r\},\\bm\{\\xi\}\},\\mathbf\{B\}\_\{\\mathbf\{r\},\\bm\{\\xi\}\}\\in\\mathbb\{R\}^\{c\_\{v\}\\times c\_\{v\}\}\.
###### Proof\.
Let:
𝐫^=𝐫‖𝐫‖,𝐏∥=𝐫^𝐫^⊤,𝐏⟂=𝐈n−𝐏∥\\widehat\{\\mathbf\{r\}\}=\\frac\{\\mathbf\{r\}\}\{\\\|\\mathbf\{r\}\\\|\},\\qquad\\mathbf\{P\}^\{\\parallel\}=\\widehat\{\\mathbf\{r\}\}\\widehat\{\\mathbf\{r\}\}^\{\\top\},\\qquad\\mathbf\{P\}^\{\\perp\}=\\mathbf\{I\}\_\{n\}\-\\mathbf\{P\}^\{\\parallel\}\(53\)
For a fixed non\-zero displacement𝐫\\mathbf\{r\}, consider its stabilizer:
H𝐫=\{Q∈O\(n\):Q𝐫=𝐫\}≅O\(n−1\)H\_\{\\mathbf\{r\}\}=\\left\\\{Q\\in O\(n\):Q\\mathbf\{r\}=\\mathbf\{r\}\\right\\\}\\cong O\(n\-1\)\(54\)Under this subgroup, the spatial representation decomposes as:
ℝn=span\{𝐫^\}⊕𝐫^⟂\\mathbb\{R\}^\{n\}=\\operatorname\{span\}\\\{\\widehat\{\\mathbf\{r\}\}\\\}\\oplus\\widehat\{\\mathbf\{r\}\}^\{\\perp\}\(55\)The stabilizer acts trivially on the radial componentspan\{𝐫^\}\\operatorname\{span\}\\\{\\widehat\{\\mathbf\{r\}\}\\\}and through the standardO\(n−1\)O\(n\-1\)representation on the tangential component𝐫^⟂\\widehat\{\\mathbf\{r\}\}^\{\\perp\}\.
Let𝝃\\bm\{\\xi\}denote any additionalO\(n\)O\(n\)\-invariant scalar conditioning available to the transport, and consider:
𝒯𝐫,𝝃:ℝn⊗ℝcv⟶ℝn⊗ℝcv\\mathcal\{T\}\_\{\\mathbf\{r\},\\bm\{\\xi\}\}:\\mathbb\{R\}^\{n\}\\otimes\\mathbb\{R\}^\{c\_\{v\}\}\\longrightarrow\\mathbb\{R\}^\{n\}\\otimes\\mathbb\{R\}^\{c\_\{v\}\}Since𝝃\\bm\{\\xi\}is invariant,O\(n\)O\(n\)\-equivariance requires:
𝒯Q𝐫,𝝃\(Q𝐕\)=Q𝒯𝐫,𝝃\(𝐕\)∀Q∈O\(n\)\\mathcal\{T\}\_\{Q\\mathbf\{r\},\\bm\{\\xi\}\}\(Q\\mathbf\{V\}\)=Q\\,\\mathcal\{T\}\_\{\\mathbf\{r\},\\bm\{\\xi\}\}\(\\mathbf\{V\}\)\\qquad\\forall Q\\in O\(n\)\(56\)
For everyQ∈H𝐫Q\\in H\_\{\\mathbf\{r\}\}, the displacement is unchanged, so Equation[56](https://arxiv.org/html/2608.28853#A2.E56)reduces to:
𝒯𝐫,𝝃\(Q𝐕\)=Q𝒯𝐫,𝝃\(𝐕\)\\mathcal\{T\}\_\{\\mathbf\{r\},\\bm\{\\xi\}\}\(Q\\mathbf\{V\}\)=Q\\,\\mathcal\{T\}\_\{\\mathbf\{r\},\\bm\{\\xi\}\}\(\\mathbf\{V\}\)\(57\)Thus, for fixed\(𝐫,𝝃\)\(\\mathbf\{r\},\\bm\{\\xi\}\), the transport must commute with the action of the stabilizerH𝐫H\_\{\\mathbf\{r\}\}\.
The two spatial subspaces in Equation[55](https://arxiv.org/html/2608.28853#A2.E55)carry inequivalent representations ofH𝐫H\_\{\\mathbf\{r\}\}: the radial line carries the trivial representation, whereas𝐫^⟂\\widehat\{\\mathbf\{r\}\}^\{\\perp\}carries the standard representation ofO\(n−1\)O\(n\-1\)\. An intertwining operator therefore cannot mix the radial and tangential components\. On the tangential subspace, the spatial part of the commutant is proportional to the identity\. Sinceℝcv\\mathbb\{R\}^\{c\_\{v\}\}carries no non\-trivialO\(n\)O\(n\)action, arbitrary linear maps remain available on the channel dimension\. Consequently:
EndH𝐫\(ℝn⊗ℝcv\)=\(𝐏∥⊗End\(ℝcv\)\)⊕\(𝐏⟂⊗End\(ℝcv\)\)\\operatorname\{End\}\_\{H\_\{\\mathbf\{r\}\}\}\\left\(\\mathbb\{R\}^\{n\}\\otimes\\mathbb\{R\}^\{c\_\{v\}\}\\right\)=\\left\(\\mathbf\{P\}^\{\\parallel\}\\otimes\\operatorname\{End\}\(\\mathbb\{R\}^\{c\_\{v\}\}\)\\right\)\\oplus\\left\(\\mathbf\{P\}^\{\\perp\}\\otimes\\operatorname\{End\}\(\\mathbb\{R\}^\{c\_\{v\}\}\)\\right\)\(58\)Hence the transport must take the form:
𝒯𝐫,𝝃\(𝐕\)=𝐏∥\(𝐫\)𝐕𝐀𝐫,𝝃\+𝐏⟂\(𝐫\)𝐕𝐁𝐫,𝝃\\boxed\{\\mathcal\{T\}\_\{\\mathbf\{r\},\\bm\{\\xi\}\}\(\\mathbf\{V\}\)=\\mathbf\{P\}^\{\\parallel\}\(\\mathbf\{r\}\)\\mathbf\{V\}\\mathbf\{A\}\_\{\\mathbf\{r\},\\bm\{\\xi\}\}\+\\mathbf\{P\}^\{\\perp\}\(\\mathbf\{r\}\)\\mathbf\{V\}\\mathbf\{B\}\_\{\\mathbf\{r\},\\bm\{\\xi\}\}\}\(59\)for channel endomorphisms𝐀𝐫,𝝃,𝐁𝐫,𝝃∈ℝcv×cv\\mathbf\{A\}\_\{\\mathbf\{r\},\\bm\{\\xi\}\},\\mathbf\{B\}\_\{\\mathbf\{r\},\\bm\{\\xi\}\}\\in\\mathbb\{R\}^\{c\_\{v\}\\times c\_\{v\}\}\.
It remains to determine how these channel maps may depend on the orientation of𝐫\\mathbf\{r\}\. From Equation[56](https://arxiv.org/html/2608.28853#A2.E56)and:
𝐏∥\(Q𝐫\)=Q𝐏∥\(𝐫\)Q⊤,𝐏⟂\(Q𝐫\)=Q𝐏⟂\(𝐫\)Q⊤\\mathbf\{P\}^\{\\parallel\}\(Q\\mathbf\{r\}\)=Q\\mathbf\{P\}^\{\\parallel\}\(\\mathbf\{r\}\)Q^\{\\top\},\\qquad\\mathbf\{P\}^\{\\perp\}\(Q\\mathbf\{r\}\)=Q\\mathbf\{P\}^\{\\perp\}\(\\mathbf\{r\}\)Q^\{\\top\}the uniqueness of the radial–tangential block decomposition gives:
𝐀Q𝐫,𝝃=𝐀𝐫,𝝃,𝐁Q𝐫,𝝃=𝐁𝐫,𝝃\\mathbf\{A\}\_\{Q\\mathbf\{r\},\\bm\{\\xi\}\}=\\mathbf\{A\}\_\{\\mathbf\{r\},\\bm\{\\xi\}\},\\qquad\\mathbf\{B\}\_\{Q\\mathbf\{r\},\\bm\{\\xi\}\}=\\mathbf\{B\}\_\{\\mathbf\{r\},\\bm\{\\xi\}\}\(60\)The channel endomorphisms can therefore depend on the displacement only throughO\(n\)O\(n\)\-invariant quantities such as‖𝐫‖\\\|\\mathbf\{r\}\\\|, together with the additional invariant conditioning𝝃\\bm\{\\xi\}\. This establishes Equation[59](https://arxiv.org/html/2608.28853#A2.E59)as the complete class of linearO\(n\)O\(n\)\-equivariant transports under the stated conditioning assumptions\. ∎
#### Scope of the result\.
First, arbitrary additionalO\(n\)O\(n\)\-invariant scalar information does not alter the radial–tangential form: it may change the channel maps𝐀𝐫,𝝃\\mathbf\{A\}\_\{\\mathbf\{r\},\\bm\{\\xi\}\}and𝐁𝐫,𝝃\\mathbf\{B\}\_\{\\mathbf\{r\},\\bm\{\\xi\}\}, but introduces no additional spatial operators\. ESNN parameterizes these channel transformations through the structured factorization𝐖𝐃\(𝐠ij\)\\mathbf\{W\}\\mathbf\{D\}\(\\mathbf\{g\}\_\{ij\}\)\. Second, the completeness statement relies on the full orthogonal groupO\(n\)O\(n\)\. UnderSO\(n\)SO\(n\)alone, additional orientation\-sensitive spatial operators may be admissible\. In three dimensions, for example, consider\[𝐫^\]×\[\\widehat\{\\mathbf\{r\}\}\]\_\{\\times\}defined by:
\[𝐫^\]×𝐯=𝐫^×𝐯\[\\widehat\{\\mathbf\{r\}\}\]\_\{\\times\}\\mathbf\{v\}=\\widehat\{\\mathbf\{r\}\}\\times\\mathbf\{v\}\(61\)ForQ∈O\(3\)Q\\in O\(3\), this operator transforms as:
\[Q𝐫^\]×=det\(Q\)Q\[𝐫^\]×Q⊤\[Q\\widehat\{\\mathbf\{r\}\}\]\_\{\\times\}=\\det\(Q\)\\,Q\[\\widehat\{\\mathbf\{r\}\}\]\_\{\\times\}Q^\{\\top\}\(62\)It therefore satisfies the required conjugation law forQ∈SO\(3\)Q\\in SO\(3\)but acquires an additional sign under reflections\. Theorem[4\.3](https://arxiv.org/html/2608.28853#S4.Thmtheorem3)should consequently be understood as a completeness result for displacement\-conditioned linearO\(n\)O\(n\)\-equivariant transport, not for the larger class of arbitrarySO\(n\)SO\(n\)\-equivariant kernels\.
### B\.3Feature\-Conditioned Spatial Operators
The completeness result of Appendix[B\.2](https://arxiv.org/html/2608.28853#A2.SS2)assumes that the relative displacement is the only covariant geometric quantity available to the transport\. ESNN can also use the learned vector features at the endpoints of an edge to construct spatial operators\. These additional covariant quantities enlarge the local geometric context and therefore allow transport mechanisms beyond the radial–tangential form\. The resulting transport may depend nonlinearly on the evolving feature state, while remaining linear in the transported vector feature once the local context is fixed\.
A simple example is the cross\-feature operator:
𝐂ij=𝐕j𝐕i⊤\\mathbf\{C\}\_\{ij\}=\\mathbf\{V\}\_\{j\}\\mathbf\{V\}\_\{i\}^\{\\top\}\(63\)Under a common orthogonal transformation of the vector features, it transforms by conjugation:
𝐂ij↦Q𝐂ijQ⊤\\mathbf\{C\}\_\{ij\}\\mapsto Q\\mathbf\{C\}\_\{ij\}Q^\{\\top\}\(64\)Its skew\-symmetric component:
𝛀ijV=𝐕j𝐕i⊤−𝐕i𝐕j⊤\\mathbf\{\\Omega\}^\{V\}\_\{ij\}=\\mathbf\{V\}\_\{j\}\\mathbf\{V\}\_\{i\}^\{\\top\}\-\\mathbf\{V\}\_\{i\}\\mathbf\{V\}\_\{j\}^\{\\top\}\(65\)inherits the same transformation law:
𝛀ijV↦Q𝛀ijVQ⊤\\mathbf\{\\Omega\}^\{V\}\_\{ij\}\\mapsto Q\\mathbf\{\\Omega\}^\{V\}\_\{ij\}Q^\{\\top\}\(66\)Since the Frobenius norm is invariant under orthogonal conjugation, the normalized operator:
𝛀^ijV=𝛀ijV‖𝛀ijV‖F\+ε\\widehat\{\\mathbf\{\\Omega\}\}^\{V\}\_\{ij\}=\\frac\{\\mathbf\{\\Omega\}^\{V\}\_\{ij\}\}\{\\\|\\mathbf\{\\Omega\}^\{V\}\_\{ij\}\\\|\_\{F\}\+\\varepsilon\}\(67\)is also a validO\(n\)O\(n\)\-covariant spatial operator\. This is the feature\-conditioned skew operator used in the Unified Transport of Section[4](https://arxiv.org/html/2608.28853#S4)\.
Orthogonal Transport uses the same principle with a slightly different normalization\. Starting from𝐂ij=𝐕j𝐕i⊤\\mathbf\{C\}\_\{ij\}=\\mathbf\{V\}\_\{j\}\\mathbf\{V\}\_\{i\}^\{\\top\}, define:
𝐂~ij=𝐂ij‖𝐂ij‖F\+ε,𝛀ij=𝐂~ij−𝐂~ij⊤\\widetilde\{\\mathbf\{C\}\}\_\{ij\}=\\frac\{\\mathbf\{C\}\_\{ij\}\}\{\\\|\\mathbf\{C\}\_\{ij\}\\\|\_\{F\}\+\\varepsilon\},\\qquad\\mathbf\{\\Omega\}\_\{ij\}=\\widetilde\{\\mathbf\{C\}\}\_\{ij\}\-\\widetilde\{\\mathbf\{C\}\}\_\{ij\}^\{\\top\}\(68\)The invariance of the Frobenius norm again gives:
𝛀ij↦Q𝛀ijQ⊤\\mathbf\{\\Omega\}\_\{ij\}\\mapsto Q\\mathbf\{\\Omega\}\_\{ij\}Q^\{\\top\}\(69\)Ifβij\\beta\_\{ij\}is an invariant scalar, the matrix exponential preserves this conjugation law:
𝐑ij=exp\(βij𝛀ij\)↦Q𝐑ijQ⊤\\mathbf\{R\}\_\{ij\}=\\exp\\\!\\left\(\\beta\_\{ij\}\\mathbf\{\\Omega\}\_\{ij\}\\right\)\\mapsto Q\\mathbf\{R\}\_\{ij\}Q^\{\\top\}\(70\)Because𝛀ij\\mathbf\{\\Omega\}\_\{ij\}is skew\-symmetric,𝐑ij∈SO\(n\)\\mathbf\{R\}\_\{ij\}\\in SO\(n\), recovering the spatial factor used by the Orthogonal Transport of Section[4](https://arxiv.org/html/2608.28853#S4)\.
These constructions illustrate why the displacement\-only completeness result does not extend directly to feature\-conditioned transport\. Additional covariant vector features provide more geometric structure than the single direction𝐫ij\\mathbf\{r\}\_\{ij\}, so the stabilizer argument that restricts the spatial action to𝐏∥\\mathbf\{P\}^\{\\parallel\}and𝐏⟂\\mathbf\{P\}^\{\\perp\}no longer applies in general\. The Unified Transport exploits this additional freedom by combining displacement\- and feature\-conditioned spatial operators\. We do not claim that its particular operator set spans the complete class of feature\-conditionedO\(n\)O\(n\)\-equivariant transports\.
## Appendix CCellular Sheaves, Directed Transport, and Connections
ESNN directly learns transport maps between neighboring vector features\. This appendix clarifies how these maps relate to the sheaf perspective introduced in Section[2\.1](https://arxiv.org/html/2608.28853#S2.SS1)\. We first show how classical cellular\-sheaf restriction maps induce effective node\-to\-node couplings, and then explain how directly parameterizing these couplings leads naturally to directed transport on a quiver\. Finally, we identify the additional conditions under which ESNN recovers a self\-adjoint, connection\-style specialization\. Throughout, the ambientO\(n\)O\(n\)\-equivariance of ESNN refers to a common transformation of the physical geometry and should be distinguished from covariance under independent changes of local reference frame\.
### C\.1Cellular Sheaves and Induced Node\-to\-Node Coupling
The connection between cellular sheaves and ESNN is most easily seen through the node\-to\-node interaction induced by the sheaf Laplacian\. LetG=\(𝒱,ℰ\)G=\(\\mathcal\{V\},\\mathcal\{E\}\)be a graph\. A cellular sheafℱ\\mathcal\{F\}assigns a vector spaceℱ\(i\)\\mathcal\{F\}\(i\)to each vertexi∈𝒱i\\in\\mathcal\{V\}and a vector spaceℱ\(e\)\\mathcal\{F\}\(e\)to each edgee∈ℰe\\in\\mathcal\{E\}\. For every vertex–edge incidencei⊴ei\\trianglelefteq e, these local spaces are related by a linear restriction map:
ρi→e:ℱ\(i\)⟶ℱ\(e\)\\rho\_\{i\\to e\}:\\mathcal\{F\}\(i\)\\longrightarrow\\mathcal\{F\}\(e\)\(71\)The restriction maps place the features associated with the endpoints of an edge in a common edge space, where their compatibility can be compared\. After choosing an orientation for each edgee=\{i,j\}e=\\\{i,j\\\}, the degree\-00coboundary measures this disagreement as:
\(δℱ𝐡\)e=ρi→e𝐡i−ρj→e𝐡j\(\\delta\_\{\\mathcal\{F\}\}\\mathbf\{h\}\)\_\{e\}=\\rho\_\{i\\to e\}\\mathbf\{h\}\_\{i\}\-\\rho\_\{j\\to e\}\\mathbf\{h\}\_\{j\}\(72\)up to the chosen orientation convention\. With standard Euclidean inner products, the corresponding cellular\-sheaf Laplacian is𝐋ℱ=δℱ∗δℱ\\mathbf\{L\}\_\{\\mathcal\{F\}\}=\\delta\_\{\\mathcal\{F\}\}^\{\*\}\\delta\_\{\\mathcal\{F\}\}, where∗denotes the adjoint\. For a single edgee=\{i,j\}e=\\\{i,j\\\}, its contribution to the Laplacian on the two endpoint stalks is:
𝐋e=\(ρi→e∗ρi→e−ρi→e∗ρj→e−ρj→e∗ρi→eρj→e∗ρj→e\)\\mathbf\{L\}\_\{e\}=\\begin\{pmatrix\}\\rho\_\{i\\to e\}^\{\*\}\\rho\_\{i\\to e\}&\-\\rho\_\{i\\to e\}^\{\*\}\\rho\_\{j\\to e\}\\\\\[5\.69054pt\] \-\\rho\_\{j\\to e\}^\{\*\}\\rho\_\{i\\to e\}&\\rho\_\{j\\to e\}^\{\*\}\\rho\_\{j\\to e\}\\end\{pmatrix\}\(73\)
The off\-diagonal blocks reveal the effective interaction between neighboring node stalks\. In particular, information from nodejjcontributes to nodeiithrough the composition:
ρi→e∗ρj→e:ℱ\(j\)⟶ℱ\(i\)\\rho\_\{i\\to e\}^\{\*\}\\rho\_\{j\\to e\}:\\mathcal\{F\}\(j\)\\longrightarrow\\mathcal\{F\}\(i\)\(74\)up to the conventional minus sign in the Laplacian\. The first restriction map expresses the feature atjjin the common edge space, while the adjoint of the restriction atiimaps the resulting representation back to the stalk atii\. Their composition therefore acts as an induced linear transport between the two neighboring node spaces\.
This effective node\-to\-node coupling is the point of departure for ESNN\. Rather than parameterizing separate restriction maps through an explicit edge stalk and obtaining their composition indirectly, ESNN learns the neighboring vector transport itself:
𝒯i←j:ℱvec\(j\)⟶ℱvec\(i\)\\mathcal\{T\}\_\{i\\leftarrow j\}:\\mathcal\{F\}\_\{\\mathrm\{vec\}\}\(j\)\\longrightarrow\\mathcal\{F\}\_\{\\mathrm\{vec\}\}\(i\)\(75\)The transport is then constrained to satisfy the transformation law of the ambient Euclidean representation, as developed in Sections[3](https://arxiv.org/html/2608.28853#S3)and[4](https://arxiv.org/html/2608.28853#S4)\. This preserves the sheaf perspective of local feature spaces connected by linear maps while allowing the two orientations of an interaction to be parameterized directly\. In general, however, the resulting ESNN operator need not arise from a classical cellular\-sheaf Laplacian\. The additional conditions under which such a realization is recovered are developed in the following subsections\.
### C\.2Directed Transport as a Quiver Representation
Directly parameterizing node\-to\-node transport naturally accommodates asymmetric interactions\. A convenient language for describing this structure is that of quiver representations\. A quiver is a directed graph whose vertices are assigned vector spaces and whose arrows are assigned linear maps between those spaces\. For ESNN, the vertices correspond to node vector stalks and the arrows to the directed transport maps used during message passing\.
For a fixed layer context, let:
𝒬=\(𝒱,𝒜\)\\mathcal\{Q\}=\(\\mathcal\{V\},\\mathcal\{A\}\)\(76\)denote the directed interaction graph, with an arrowj→ij\\to ifor every directed message\-passing interaction\. Assign the vector space:
ℱvec\(i\)=ℝn⊗ℝcv\\mathcal\{F\}\_\{\\mathrm\{vec\}\}\(i\)=\\mathbb\{R\}^\{n\}\\otimes\\mathbb\{R\}^\{c\_\{v\}\}\(77\)to each vertexii, and the linear transport:
𝒯i←j:ℱvec\(j\)⟶ℱvec\(i\)\\mathcal\{T\}\_\{i\\leftarrow j\}:\\mathcal\{F\}\_\{\\mathrm\{vec\}\}\(j\)\\longrightarrow\\mathcal\{F\}\_\{\\mathrm\{vec\}\}\(i\)\(78\)to each arrowj→ij\\to i\. For a fixed layer context, these assignments define a representation of the quiver𝒬\\mathcal\{Q\}\. The same construction can equivalently be viewed through the free path category of𝒬\\mathcal\{Q\}\. Each directed path:
i0⟶i1⟶⋯⟶imi\_\{0\}\\longrightarrow i\_\{1\}\\longrightarrow\\cdots\\longrightarrow i\_\{m\}\(79\)is associated with the composition of its edge transports:
𝒯im←im−1∘⋯∘𝒯i1←i0\\mathcal\{T\}\_\{i\_\{m\}\\leftarrow i\_\{m\-1\}\}\\circ\\cdots\\circ\\mathcal\{T\}\_\{i\_\{1\}\\leftarrow i\_\{0\}\}\(80\)so the vertex spaces and directed transports extend naturally to a covariant functor from paths in the interaction graph to finite\-dimensional vector spaces\. This categorical viewpoint is not required to define an ESNN layer, but makes explicit that directed transports can be composed consistently along paths\. The ambient equivariance of the individual edge maps is preserved under this composition\. If every arrow satisfies:
𝒯i←j′∘Q=Q∘𝒯i←j\\mathcal\{T\}^\{\\prime\}\_\{i\\leftarrow j\}\\circ Q=Q\\circ\\mathcal\{T\}\_\{i\\leftarrow j\}\(81\)whereQQdenotes its natural action onℱvec=ℝn⊗ℝcv\\mathcal\{F\}\_\{\\mathrm\{vec\}\}=\\mathbb\{R\}^\{n\}\\otimes\\mathbb\{R\}^\{c\_\{v\}\}, then for any pathi0→i1→⋯→imi\_\{0\}\\to i\_\{1\}\\to\\cdots\\to i\_\{m\}:
𝒯′im←im−1∘⋯∘𝒯′i1←i0∘Q\\displaystyle\\mathcal\{T\}^\{\\prime\}\_\{i\_\{m\}\\leftarrow i\_\{m\-1\}\}\\circ\\cdots\\circ\\mathcal\{T\}^\{\\prime\}\_\{i\_\{1\}\\leftarrow i\_\{0\}\}\\circ Q=Q∘𝒯im←im−1∘⋯∘𝒯i1←i0\\displaystyle=Q\\circ\\mathcal\{T\}\_\{i\_\{m\}\\leftarrow i\_\{m\-1\}\}\\circ\\cdots\\circ\\mathcal\{T\}\_\{i\_\{1\}\\leftarrow i\_\{0\}\}\(82\)Thus the transport associated with a path obeys the same globalO\(n\)O\(n\)intertwining relation as the individual edge maps\. This describes the algebraic composition of the transports; a single ESNN layer still performs the one\-hop aggregation defined in Section[5](https://arxiv.org/html/2608.28853#S5)\.
Importantly, the quiver formulation does not require any compatibility condition between opposite edge orientations\. If bothi→ji\\to jandj→ij\\to iare present, the general ESNN construction does not impose:
𝒯i←j=𝒯j←i∗,𝒯i←j=𝒯j←i−1\\mathcal\{T\}\_\{i\\leftarrow j\}=\\mathcal\{T\}\_\{j\\leftarrow i\}^\{\*\},\\qquad\\mathcal\{T\}\_\{i\\leftarrow j\}=\\mathcal\{T\}\_\{j\\leftarrow i\}^\{\-1\}\(83\)The two directions may therefore represent genuinely different local interactions\. This distinguishes the general directional ESNN operator from diffusion generated by a classical self\-adjoint sheaf Laplacian and is consistent with recent directed extensions of sheaf\-based learning\([Ribeiro et al\., 2025](https://arxiv.org/html/2608.28853#bib.bib44);[Fiorini et al\., 2025](https://arxiv.org/html/2608.28853#bib.bib45)\)\.
This perspective is also closely related to copresheaf neural constructions, which organize information through directed maps between local feature spaces and their compositions\([Hajij et al\., 2025](https://arxiv.org/html/2608.28853#bib.bib46)\)\. ESNN shares this covariant, directed view of information flow, but imposes an additional geometric requirement: each transport must intertwine the common ambientO\(n\)O\(n\)action carried by the vector features\. The resulting structure is therefore simultaneously directional at the level of the interaction graph and equivariant with respect to the Euclidean geometry\. Finally, the linearity discussed here is conditional on the local layer context\. The coefficients defining𝒯i←j\\mathcal\{T\}\_\{i\\leftarrow j\}may depend non\-linearly on the current node, edge, and geometric features, as described in Section[3](https://arxiv.org/html/2608.28853#S3)\. Once that context is fixed, however, each directed transport is linear in the vector feature being propagated\.
### C\.3Adjoint\-Consistent and Connection\-Style Specializations
The directed formulation of Appendix[C\.2](https://arxiv.org/html/2608.28853#A3.SS2)allows the two orientations of an edge to carry independent transport maps\. Classical sheaf diffusion is more structured: the off\-diagonal blocks of a cellular\-sheaf Laplacian occur in adjoint pairs, which makes the resulting operator self\-adjoint\. We now identify the corresponding specialization of ESNN and then show the stronger conditions under which its edge transport admits an exact connection\-sheaf realization\.
Recall the normalized transport operator:
\(𝒜𝒯𝐕\)i=νii𝐕i\+∑j∈𝒩\(i\)νij𝒯i←j\(𝐕j\)\(\\mathcal\{A\}\_\{\\mathcal\{T\}\}\\mathbf\{V\}\)\_\{i\}=\\nu\_\{ii\}\\mathbf\{V\}\_\{i\}\+\\sum\_\{j\\in\\mathcal\{N\}\(i\)\}\\nu\_\{ij\}\\mathcal\{T\}\_\{i\\leftarrow j\}\(\\mathbf\{V\}\_\{j\}\)\(84\)For the left\-right transport form introduced in Equation[7](https://arxiv.org/html/2608.28853#S3.E7):
𝒯\(𝐕\)=∑k𝐒\(k\)𝐕𝐌\(k\)\\mathcal\{T\}\(\\mathbf\{V\}\)=\\sum\_\{k\}\\mathbf\{S\}^\{\(k\)\}\\mathbf\{V\}\\mathbf\{M\}^\{\(k\)\}\(85\)the adjoint with respect to the Frobenius inner product is:
𝒯∗\(𝐔\)=∑k\(𝐒\(k\)\)⊤𝐔\(𝐌\(k\)\)⊤\\mathcal\{T\}^\{\*\}\(\\mathbf\{U\}\)=\\sum\_\{k\}\\left\(\\mathbf\{S\}^\{\(k\)\}\\right\)^\{\\top\}\\mathbf\{U\}\\left\(\\mathbf\{M\}^\{\(k\)\}\\right\)^\{\\top\}\(86\)Indeed:
⟨𝐔,𝐒𝐕𝐌⟩F=⟨𝐒⊤𝐔𝐌⊤,𝐕⟩F\\left\\langle\\mathbf\{U\},\\mathbf\{S\}\\mathbf\{V\}\\mathbf\{M\}\\right\\rangle\_\{F\}=\\left\\langle\\mathbf\{S\}^\{\\top\}\\mathbf\{U\}\\mathbf\{M\}^\{\\top\},\\mathbf\{V\}\\right\\rangle\_\{F\}\(87\)For Radial–Tangential Transport, the spatial projectors are symmetric, so the adjoint acts only by transposing the corresponding channel maps:
𝒯∗\(𝐔\)=𝐏∥𝐔\(𝐌∥\)⊤\+𝐏⟂𝐔\(𝐌⟂\)⊤\\mathcal\{T\}^\{\*\}\(\\mathbf\{U\}\)=\\mathbf\{P\}^\{\\parallel\}\\mathbf\{U\}\\left\(\\mathbf\{M\}^\{\\parallel\}\\right\)^\{\\top\}\+\\mathbf\{P\}^\{\\perp\}\\mathbf\{U\}\\left\(\\mathbf\{M\}^\{\\perp\}\\right\)^\{\\top\}\(88\)
This makes the condition for self\-adjoint transport explicit\. On a bidirected graph, the scalar normalization must agree across the two orientations and the transport in one direction must be the adjoint of the transport in the other\.
###### Proposition C\.1\(Self\-Adjoint Transport Operator\)\.
Fix the layer context so that all transport maps are linear\. Suppose the graph is bidirected and that, for every adjacent pairi,ji,j:
νij=νji,𝒯j←i=𝒯i←j∗\\nu\_\{ij\}=\\nu\_\{ji\},\\qquad\\mathcal\{T\}\_\{j\\leftarrow i\}=\\mathcal\{T\}\_\{i\\leftarrow j\}^\{\*\}\(89\)Then𝒜𝒯\\mathcal\{A\}\_\{\\mathcal\{T\}\}is self\-adjoint with respect to the direct\-sum Frobenius inner product on the node vector features\.
###### Proof\.
For two node signals𝐔\\mathbf\{U\}and𝐕\\mathbf\{V\}, the off\-diagonal contribution to⟨𝐔,𝒜𝒯𝐕⟩\\langle\\mathbf\{U\},\\mathcal\{A\}\_\{\\mathcal\{T\}\}\\mathbf\{V\}\\ranglecontains:
νij⟨𝐔i,𝒯i←j𝐕j⟩\\nu\_\{ij\}\\left\\langle\\mathbf\{U\}\_\{i\},\\mathcal\{T\}\_\{i\\leftarrow j\}\\mathbf\{V\}\_\{j\}\\right\\rangle\(90\)Using the adjoint relation and symmetry of the scalar coefficient:
νij⟨𝐔i,𝒯i←j𝐕j⟩=νji⟨𝒯j←i𝐔i,𝐕j⟩\\nu\_\{ij\}\\left\\langle\\mathbf\{U\}\_\{i\},\\mathcal\{T\}\_\{i\\leftarrow j\}\\mathbf\{V\}\_\{j\}\\right\\rangle=\\nu\_\{ji\}\\left\\langle\\mathcal\{T\}\_\{j\\leftarrow i\}\\mathbf\{U\}\_\{i\},\\mathbf\{V\}\_\{j\}\\right\\rangle\(91\)Summing over the bidirected edge set gives:
⟨𝐔,𝒜𝒯𝐕⟩=⟨𝒜𝒯𝐔,𝐕⟩\\langle\\mathbf\{U\},\\mathcal\{A\}\_\{\\mathcal\{T\}\}\\mathbf\{V\}\\rangle=\\langle\\mathcal\{A\}\_\{\\mathcal\{T\}\}\\mathbf\{U\},\\mathbf\{V\}\\rangle\(92\)The diagonal self\-loop terms are real scalar multiples of the identity and are therefore self\-adjoint\. ∎
Proposition[C\.1](https://arxiv.org/html/2608.28853#A3.Thmtheorem1)identifies the adjoint\-consistent specialization of the general directional operator\. This brings ESNN closer to classical sheaf diffusion, but self\-adjointness alone is not sufficient to make an arbitrary𝒜𝒯\\mathcal\{A\}\_\{\\mathcal\{T\}\}a cellular\-sheaf Laplacian\. In particular, a classical connection sheaf imposes additional structure: transport between neighboring fibers is orthogonal, and reversing the edge applies its inverse, which coincides with its adjoint\. Under these stronger conditions, the ESNN edge transport can be realized exactly through classical sheaf restriction maps\.
###### Proposition C\.2\(Connection\-Sheaf Realization\)\.
Letℋ=ℝn⊗ℝcv\\mathcal\{H\}=\\mathbb\{R\}^\{n\}\\otimes\\mathbb\{R\}^\{c\_\{v\}\}be the vector fiber, and consider an undirected edgee=\{i,j\}e=\\\{i,j\\\}\. Suppose the full transport:
𝒰i←j:ℋ⟶ℋ\\mathcal\{U\}\_\{i\\leftarrow j\}:\\mathcal\{H\}\\longrightarrow\\mathcal\{H\}\(93\)is orthogonal and that the reverse transport is its adjoint:
𝒰j←i=𝒰i←j∗=𝒰i←j−1\\mathcal\{U\}\_\{j\\leftarrow i\}=\\mathcal\{U\}\_\{i\\leftarrow j\}^\{\*\}=\\mathcal\{U\}\_\{i\\leftarrow j\}^\{\-1\}\(94\)Choose vertex and edge stalks equal toℋ\\mathcal\{H\}and restrictions:
ρi→e=𝐈ℋ,ρj→e=𝒰i←j\\rho\_\{i\\to e\}=\\mathbf\{I\}\_\{\\mathcal\{H\}\},\\qquad\\rho\_\{j\\to e\}=\\mathcal\{U\}\_\{i\\leftarrow j\}\(95\)Then the contribution ofeeto the cellular\-sheaf Laplacian is:
𝐋e=\(𝐈ℋ−𝒰i←j−𝒰i←j∗𝐈ℋ\)\\mathbf\{L\}\_\{e\}=\\begin\{pmatrix\}\\mathbf\{I\}\_\{\\mathcal\{H\}\}&\-\\mathcal\{U\}\_\{i\\leftarrow j\}\\\\\[2\.84526pt\] \-\\mathcal\{U\}\_\{i\\leftarrow j\}^\{\*\}&\\mathbf\{I\}\_\{\\mathcal\{H\}\}\\end\{pmatrix\}\(96\)Thus the off\-diagonal node couplings are exactly the orthogonal transport and its adjoint, up to the conventional Laplacian sign\.
###### Proof\.
Substituting Equation[95](https://arxiv.org/html/2608.28853#A3.E95)into Equation[73](https://arxiv.org/html/2608.28853#A3.E73)gives:
ρi→e∗ρi→e=𝐈ℋ,ρi→e∗ρj→e=𝒰i←j\\rho\_\{i\\to e\}^\{\*\}\\rho\_\{i\\to e\}=\\mathbf\{I\}\_\{\\mathcal\{H\}\},\\qquad\\rho\_\{i\\to e\}^\{\*\}\\rho\_\{j\\to e\}=\\mathcal\{U\}\_\{i\\leftarrow j\}\(97\)Since𝒰i←j\\mathcal\{U\}\_\{i\\leftarrow j\}is orthogonal:
ρj→e∗ρj→e=𝒰i←j∗𝒰i←j=𝐈ℋ\\rho\_\{j\\to e\}^\{\*\}\\rho\_\{j\\to e\}=\\mathcal\{U\}\_\{i\\leftarrow j\}^\{\*\}\\mathcal\{U\}\_\{i\\leftarrow j\}=\\mathbf\{I\}\_\{\\mathcal\{H\}\}\(98\)while:
ρj→e∗ρi→e=𝒰i←j∗\\rho\_\{j\\to e\}^\{\*\}\\rho\_\{i\\to e\}=\\mathcal\{U\}\_\{i\\leftarrow j\}^\{\*\}\(99\)Equation[96](https://arxiv.org/html/2608.28853#A3.E96)follows\. ∎
Proposition[C\.2](https://arxiv.org/html/2608.28853#A3.Thmtheorem2)is the graph analogue of a discrete orthogonal connection and is closely related to connection\-Laplacian SNNs\([Barbero et al\., 2022a](https://arxiv.org/html/2608.28853#bib.bib38)\)\. It also clarifies the precise sense in which ESNN Orthogonal Transport is connection\-style\. Its spatial factor𝐑i←j∈SO\(n\)\\mathbf\{R\}\_\{i\\leftarrow j\}\\in SO\(n\)has the required orthogonal structure and recovers the proposition whencv=1c\_\{v\}=1with trivial channel action, or when that spatial factor is considered in isolation\. The complete ESNN transport, however, also contains channel mixing and edge\-dependent gating and is therefore not generally orthogonal onℝn⊗ℝcv\\mathbb\{R\}^\{n\}\\otimes\\mathbb\{R\}^\{c\_\{v\}\}\.
There is consequently a hierarchy of increasingly restrictive cases\. General ESNN transport permits independent directed edge maps\. Imposing Equation[89](https://arxiv.org/html/2608.28853#A3.E89)yields a self\-adjoint transport operator, while additionally requiring the complete fiber maps to be orthogonal gives the classical connection\-sheaf realization of Proposition[C\.2](https://arxiv.org/html/2608.28853#A3.Thmtheorem2)\. Outside this final specialization, the term connection\-style refers to the geometric role of the transport rather than to an exact classical connection sheaf\.
### C\.4Ambient Equivariance and Local Gauge Transformations
The connection\-style interpretation above should not be confused with local gauge equivariance\. In a classical cellular sheaf, the bases of different stalks may be changed independently\. If𝐐i\\mathbf\{Q\}\_\{i\}and𝐐e\\mathbf\{Q\}\_\{e\}are orthogonal changes of basis on a vertex stalk and an incident edge stalk, the coordinate representation of the corresponding restriction map transforms as:
ρi→e′=𝐐eρi→e𝐐i⊤\\rho^\{\\prime\}\_\{i\\to e\}=\\mathbf\{Q\}\_\{e\}\\rho\_\{i\\to e\}\\mathbf\{Q\}\_\{i\}^\{\\top\}\(100\)This expresses covariance under independent changes of local reference frame\.
ESNN enforces a different symmetry\. A single orthogonal transformationQ∈O\(n\)Q\\in O\(n\)acts simultaneously on the physical coordinates and all vector features:
𝐱i↦Q𝐱i\+𝐭,𝐕i↦Q𝐕i\\mathbf\{x\}\_\{i\}\\mapsto Q\\mathbf\{x\}\_\{i\}\+\\mathbf\{t\},\\qquad\\mathbf\{V\}\_\{i\}\\mapsto Q\\mathbf\{V\}\_\{i\}\(101\)while the spatial transport transforms as:
𝐒ij↦Q𝐒ijQ⊤\\mathbf\{S\}\_\{ij\}\\mapsto Q\\mathbf\{S\}\_\{ij\}Q^\{\\top\}\(102\)This is the ambientE\(n\)E\(n\)\-equivariance established in Sections[3](https://arxiv.org/html/2608.28853#S3)–[5](https://arxiv.org/html/2608.28853#S5)\. It corresponds to transforming the entire geometric system in a common Cartesian frame rather than independently changing the basis attached to each node\. In particular, the feature\-derived Orthogonal Transport is not designed to be covariant under arbitrary node\-wise transformations\. For𝐂ij=𝐕j𝐕i⊤\\mathbf\{C\}\_\{ij\}=\\mathbf\{V\}\_\{j\}\\mathbf\{V\}\_\{i\}^\{\\top\}, independent changes:
𝐕i↦𝐐i𝐕i,𝐕j↦𝐐j𝐕j\\mathbf\{V\}\_\{i\}\\mapsto\\mathbf\{Q\}\_\{i\}\\mathbf\{V\}\_\{i\},\\qquad\\mathbf\{V\}\_\{j\}\\mapsto\\mathbf\{Q\}\_\{j\}\\mathbf\{V\}\_\{j\}\(103\)give:
𝐂ij′=𝐐j𝐂ij𝐐i⊤\\mathbf\{C\}^\{\\prime\}\_\{ij\}=\\mathbf\{Q\}\_\{j\}\\mathbf\{C\}\_\{ij\}\\mathbf\{Q\}\_\{i\}^\{\\top\}\(104\)and the skew\-symmetric construction used by Orthogonal Transport does not in general transform as a gauge\-covariant map between the two independently chosen frames\. When𝐐i=𝐐j=Q\\mathbf\{Q\}\_\{i\}=\\mathbf\{Q\}\_\{j\}=Q, however, the same construction reduces to conjugation by the common ambient transformation, which is precisely the covariance required by ESNN\.
Thus the Orthogonal Transport should be understood as a*connection\-style ambient transport*\. The orthogonal, adjoint\-consistent specialization of Proposition[C\.2](https://arxiv.org/html/2608.28853#A3.Thmtheorem2)admits an exact classical connection\-sheaf realization, whereas the general ESNN architecture does not claim covariance under arbitrary local gauge transformations\.
This distinction also suggests a possible extension of the framework\. A gauge\-equivariant ESNN would allow each node to carry its own local frame and would require a directed transport to transform according to:
𝒯i←j′=𝐐i𝒯i←j𝐐j⊤\\mathcal\{T\}^\{\\prime\}\_\{i\\leftarrow j\}=\\mathbf\{Q\}\_\{i\}\\mathcal\{T\}\_\{i\\leftarrow j\}\\mathbf\{Q\}\_\{j\}^\{\\top\}\(105\)under independent local transformations𝐐i\\mathbf\{Q\}\_\{i\}and𝐐j\\mathbf\{Q\}\_\{j\}\. Achieving this property would require transport generators constructed directly from gauge\-covariant quantities rather than the ambient feature\-derived operators used here\. Developing such a locally gauge\-equivariant extension is left for future work\.
## Appendix DProofs of Equivariance and Symmetry Results
This appendix provides the proofs of the equivariance and symmetry results stated in the main text\. The algebraic proof of the completeness of displacement\-conditioned linear transport is given separately in Appendix[B\.2](https://arxiv.org/html/2608.28853#A2.SS2), while the self\-adjoint and connection\-sheaf specializations are established in Appendix[C](https://arxiv.org/html/2608.28853#A3)\.
### D\.1Proof of Proposition[3\.2](https://arxiv.org/html/2608.28853#S3.Thmtheorem2):O\(n\)O\(n\)\-Equivariance of the Transport Map
Proposition[3\.2](https://arxiv.org/html/2608.28853#S3.Thmtheorem2)\(O\(n\)O\(n\)\-Equivariance of the Transport Map\)\.Suppose that, for every componentkk, the spatial operator satisfies:𝐒ij\(k\)\(Q⋅𝒞ij\)=Q𝐒ij\(k\)\(𝒞ij\)Q⊤\\mathbf\{S\}\_\{ij\}^\{\(k\)\}\(Q\\\!\\cdot\\\!\\mathcal\{C\}\_\{ij\}\)=Q\\mathbf\{S\}\_\{ij\}^\{\(k\)\}\(\\mathcal\{C\}\_\{ij\}\)Q^\{\\top\}\(106\)and the channel operator𝐌ij\(k\)\\mathbf\{M\}\_\{ij\}^\{\(k\)\}is unchanged under theO\(n\)O\(n\)action\. Then the transport map in Equation[7](https://arxiv.org/html/2608.28853#S3.E7)isO\(n\)O\(n\)\-equivariant:𝒯i←j′\(Q𝐕j\)=Q𝒯i←j\(𝐕j\)\\mathcal\{T\}^\{\\prime\}\_\{i\\leftarrow j\}\(Q\\mathbf\{V\}\_\{j\}\)=Q\\,\\mathcal\{T\}\_\{i\\leftarrow j\}\(\\mathbf\{V\}\_\{j\}\)\(107\)where𝒯i←j′\\mathcal\{T\}^\{\\prime\}\_\{i\\leftarrow j\}denotes the transport evaluated from the transformed geometric inputs\.
###### Proof\.
LetQ∈O\(n\)Q\\in O\(n\)and consider the transport evaluated from the transformed local edge contextQ⋅𝒞ijQ\\\!\\cdot\\\!\\mathcal\{C\}\_\{ij\}\. By assumption, each spatial operator transforms covariantly:
𝐒ij\(k\)\(Q⋅𝒞ij\)=Q𝐒ij\(k\)\(𝒞ij\)Q⊤\\mathbf\{S\}\_\{ij\}^\{\(k\)\}\(Q\\\!\\cdot\\\!\\mathcal\{C\}\_\{ij\}\)=Q\\mathbf\{S\}\_\{ij\}^\{\(k\)\}\(\\mathcal\{C\}\_\{ij\}\)Q^\{\\top\}\(108\)while each channel operator is invariant:
𝐌ij\(k\)\(Q⋅𝒞ij\)=𝐌ij\(k\)\(𝒞ij\)\\mathbf\{M\}\_\{ij\}^\{\(k\)\}\(Q\\\!\\cdot\\\!\\mathcal\{C\}\_\{ij\}\)=\\mathbf\{M\}\_\{ij\}^\{\(k\)\}\(\\mathcal\{C\}\_\{ij\}\)\(109\)Evaluating the transport on the transformed vector feature therefore gives:
𝒯i←j′\(Q𝐕j\)\\displaystyle\\mathcal\{T\}^\{\\prime\}\_\{i\\leftarrow j\}\(Q\\mathbf\{V\}\_\{j\}\)=∑k=1K\(Q𝐒ij\(k\)Q⊤\)\(Q𝐕j\)𝐌ij\(k\)\\displaystyle=\\sum\_\{k=1\}^\{K\}\\left\(Q\\mathbf\{S\}\_\{ij\}^\{\(k\)\}Q^\{\\top\}\\right\)\(Q\\mathbf\{V\}\_\{j\}\)\\mathbf\{M\}\_\{ij\}^\{\(k\)\}\(110\)=Q∑k=1K𝐒ij\(k\)𝐕j𝐌ij\(k\)\\displaystyle=Q\\sum\_\{k=1\}^\{K\}\\mathbf\{S\}\_\{ij\}^\{\(k\)\}\\mathbf\{V\}\_\{j\}\\mathbf\{M\}\_\{ij\}^\{\(k\)\}\(111\)=Q𝒯i←j\(𝐕j\)\\displaystyle=Q\\,\\mathcal\{T\}\_\{i\\leftarrow j\}\(\\mathbf\{V\}\_\{j\}\)\(112\)which is exactly the requiredO\(n\)O\(n\)\-equivariance relation\. ∎
### D\.2Proof of Lemma[4\.1](https://arxiv.org/html/2608.28853#S4.Thmtheorem1):O\(n\)O\(n\)\-Equivariance of Orthogonal Transport
Lemma[4\.1](https://arxiv.org/html/2608.28853#S4.Thmtheorem1)\(O\(n\)O\(n\)\-Equivariance of Orthogonal Transport\)\.LetQ∈O\(n\)Q\\in O\(n\)act on the vector features as𝐕i↦Q𝐕i\\mathbf\{V\}\_\{i\}\\mapsto Q\\mathbf\{V\}\_\{i\}\. Then the Orthogonal Transport of Equation[17](https://arxiv.org/html/2608.28853#S4.E17)satisfies:𝒯i←j′\(Q𝐕j\)=Q𝒯i←j\(𝐕j\)\\mathcal\{T\}^\{\\prime\}\_\{i\\leftarrow j\}\(Q\\mathbf\{V\}\_\{j\}\)=Q\\,\\mathcal\{T\}\_\{i\\leftarrow j\}\(\\mathbf\{V\}\_\{j\}\)\(113\)where the primed transport is evaluated from the transformed local context\.
###### Proof\.
Recall that the Orthogonal Transport is:
𝒯i←j\(𝐕j\)=𝐑ij𝐕j𝐌ij,𝐌ij=𝐖𝐃\(𝐠ij\)\\mathcal\{T\}\_\{i\\leftarrow j\}\(\\mathbf\{V\}\_\{j\}\)=\\mathbf\{R\}\_\{ij\}\\mathbf\{V\}\_\{j\}\\mathbf\{M\}\_\{ij\},\\qquad\\mathbf\{M\}\_\{ij\}=\\mathbf\{W\}\\mathbf\{D\}\(\\mathbf\{g\}\_\{ij\}\)\(114\)with𝐑ij=exp\(βij𝛀ij\)\\mathbf\{R\}\_\{ij\}=\\exp\\\!\\left\(\\beta\_\{ij\}\\mathbf\{\\Omega\}\_\{ij\}\\right\)and:
𝛀ij=𝐂~ij−𝐂~ij⊤,𝐂~ij=𝐂ij‖𝐂ij‖F\+ε,𝐂ij=𝐕j𝐕i⊤\\mathbf\{\\Omega\}\_\{ij\}=\\widetilde\{\\mathbf\{C\}\}\_\{ij\}\-\\widetilde\{\\mathbf\{C\}\}\_\{ij\}^\{\\top\},\\qquad\\widetilde\{\\mathbf\{C\}\}\_\{ij\}=\\frac\{\\mathbf\{C\}\_\{ij\}\}\{\\\|\\mathbf\{C\}\_\{ij\}\\\|\_\{F\}\+\\varepsilon\},\\qquad\\mathbf\{C\}\_\{ij\}=\\mathbf\{V\}\_\{j\}\\mathbf\{V\}\_\{i\}^\{\\top\}\(115\)
Since𝛀ij⊤=−𝛀ij\\mathbf\{\\Omega\}\_\{ij\}^\{\\top\}=\-\\mathbf\{\\Omega\}\_\{ij\},βij𝛀ij∈𝔰𝔬\(n\)\\beta\_\{ij\}\\mathbf\{\\Omega\}\_\{ij\}\\in\\mathfrak\{so\}\(n\)and therefore𝐑ij∈SO\(n\)\\mathbf\{R\}\_\{ij\}\\in SO\(n\)\. Under a common orthogonal transformation:
𝐕i′=Q𝐕i,𝐕j′=Q𝐕j\\mathbf\{V\}\_\{i\}^\{\\prime\}=Q\\mathbf\{V\}\_\{i\},\\qquad\\mathbf\{V\}\_\{j\}^\{\\prime\}=Q\\mathbf\{V\}\_\{j\}\(116\)The cross\-feature matrix therefore transforms as:
𝐂ij′=\(Q𝐕j\)\(Q𝐕i\)⊤=Q𝐕j𝐕i⊤Q⊤=Q𝐂ijQ⊤\\mathbf\{C\}^\{\\prime\}\_\{ij\}=\(Q\\mathbf\{V\}\_\{j\}\)\(Q\\mathbf\{V\}\_\{i\}\)^\{\\top\}=Q\\mathbf\{V\}\_\{j\}\\mathbf\{V\}\_\{i\}^\{\\top\}Q^\{\\top\}=Q\\mathbf\{C\}\_\{ij\}Q^\{\\top\}\(117\)
The Frobenius norm is invariant under orthogonal left and right multiplication, so‖𝐂ij′‖F=‖𝐂ij‖F\\\|\\mathbf\{C\}^\{\\prime\}\_\{ij\}\\\|\_\{F\}=\\\|\\mathbf\{C\}\_\{ij\}\\\|\_\{F\}\. Consequently:
𝐂~ij′=Q𝐂~ijQ⊤\\widetilde\{\\mathbf\{C\}\}^\{\\prime\}\_\{ij\}=Q\\widetilde\{\\mathbf\{C\}\}\_\{ij\}Q^\{\\top\}\(118\)and hence:
𝛀ij′=𝐂~ij′−\(𝐂~ij′\)⊤=Q\(𝐂~ij−𝐂~ij⊤\)Q⊤=Q𝛀ijQ⊤\\mathbf\{\\Omega\}^\{\\prime\}\_\{ij\}=\\widetilde\{\\mathbf\{C\}\}^\{\\prime\}\_\{ij\}\-\(\\widetilde\{\\mathbf\{C\}\}^\{\\prime\}\_\{ij\}\)^\{\\top\}=Q\\left\(\\widetilde\{\\mathbf\{C\}\}\_\{ij\}\-\\widetilde\{\\mathbf\{C\}\}\_\{ij\}^\{\\top\}\\right\)Q^\{\\top\}=Q\\mathbf\{\\Omega\}\_\{ij\}Q^\{\\top\}\(119\)
The coefficientβij\\beta\_\{ij\}is predicted from invariant edge features, soβij′=βij\\beta^\{\\prime\}\_\{ij\}=\\beta\_\{ij\}\. For any square matrix𝐀\\mathbf\{A\}and invertible matrix𝐐\\mathbf\{Q\}:
exp\(𝐐𝐀𝐐−1\)=𝐐exp\(𝐀\)𝐐−1\\exp\(\\mathbf\{Q\}\\mathbf\{A\}\\mathbf\{Q\}^\{\-1\}\)=\\mathbf\{Q\}\\exp\(\\mathbf\{A\}\)\\mathbf\{Q\}^\{\-1\}\(120\)SinceQ−1=Q⊤Q^\{\-1\}=Q^\{\\top\}, Equation[119](https://arxiv.org/html/2608.28853#A4.E119)gives:
𝐑ij′=exp\(βijQ𝛀ijQ⊤\)=Qexp\(βij𝛀ij\)Q⊤=Q𝐑ijQ⊤\\mathbf\{R\}^\{\\prime\}\_\{ij\}=\\exp\\\!\\left\(\\beta\_\{ij\}Q\\mathbf\{\\Omega\}\_\{ij\}Q^\{\\top\}\\right\)=Q\\exp\\\!\\left\(\\beta\_\{ij\}\\mathbf\{\\Omega\}\_\{ij\}\\right\)Q^\{\\top\}=Q\\mathbf\{R\}\_\{ij\}Q^\{\\top\}\(121\)
The gate𝐠ij\\mathbf\{g\}\_\{ij\}is likewise constructed from invariant quantities, and therefore𝐌ij′=𝐌ij\\mathbf\{M\}^\{\\prime\}\_\{ij\}=\\mathbf\{M\}\_\{ij\}\. Evaluating the transport from the transformed context yields:
𝒯i←j′\(Q𝐕j\)=𝐑ij′\(Q𝐕j\)𝐌ij=\(Q𝐑ijQ⊤\)\(Q𝐕j\)𝐌ij\\displaystyle\\mathcal\{T\}^\{\\prime\}\_\{i\\leftarrow j\}\(Q\\mathbf\{V\}\_\{j\}\)=\\mathbf\{R\}^\{\\prime\}\_\{ij\}\(Q\\mathbf\{V\}\_\{j\}\)\\mathbf\{M\}\_\{ij\}=\(Q\\mathbf\{R\}\_\{ij\}Q^\{\\top\}\)\(Q\\mathbf\{V\}\_\{j\}\)\\mathbf\{M\}\_\{ij\}=Q𝐑ij𝐕j𝐌ij\\displaystyle=Q\\mathbf\{R\}\_\{ij\}\\mathbf\{V\}\_\{j\}\\mathbf\{M\}\_\{ij\}\(122\)=Q𝒯i←j\(𝐕j\)\\displaystyle=Q\\,\\mathcal\{T\}\_\{i\\leftarrow j\}\(\\mathbf\{V\}\_\{j\}\)\(123\)Thus Orthogonal Transport isO\(n\)O\(n\)\-equivariant\. ∎
### D\.3Proof of Lemma[4\.2](https://arxiv.org/html/2608.28853#S4.Thmtheorem2):O\(n\)O\(n\)\-Equivariance of Radial–Tangential Transport
Lemma[4\.2](https://arxiv.org/html/2608.28853#S4.Thmtheorem2)\(O\(n\)O\(n\)\-Equivariance of Radial–Tangential Transport\)\.For𝐫ij≠𝟎\\mathbf\{r\}\_\{ij\}\\neq\\mathbf\{0\}, the Radial–Tangential Transport:𝒯i←j\(𝐕j\)=𝐏ij∥𝐕j𝐌ij∥\+𝐏ij⟂𝐕j𝐌ij⟂\\mathcal\{T\}\_\{i\\leftarrow j\}\(\\mathbf\{V\}\_\{j\}\)=\\mathbf\{P\}^\{\\parallel\}\_\{ij\}\\mathbf\{V\}\_\{j\}\\mathbf\{M\}^\{\\parallel\}\_\{ij\}\+\\mathbf\{P\}^\{\\perp\}\_\{ij\}\\mathbf\{V\}\_\{j\}\\mathbf\{M\}^\{\\perp\}\_\{ij\}\(124\)isO\(n\)O\(n\)\-equivariant\.
###### Proof\.
LetQ∈O\(n\)Q\\in O\(n\)\. Since𝐫ij′=Q𝐫ij\\mathbf\{r\}\_\{ij\}^\{\\prime\}=Q\\mathbf\{r\}\_\{ij\}, orthogonality implies‖𝐫ij′‖=‖𝐫ij‖\\\|\\mathbf\{r\}\_\{ij\}^\{\\prime\}\\\|=\\\|\\mathbf\{r\}\_\{ij\}\\\|, and hence:
𝐫^ij′=Q𝐫ij‖𝐫ij‖=Q𝐫^ij\\widehat\{\\mathbf\{r\}\}^\{\\prime\}\_\{ij\}=\\frac\{Q\\mathbf\{r\}\_\{ij\}\}\{\\\|\\mathbf\{r\}\_\{ij\}\\\|\}=Q\\widehat\{\\mathbf\{r\}\}\_\{ij\}\(125\)
The radial projector consequently transforms by conjugation:
\(𝐏ij∥\)′=𝐫^ij′\(𝐫^ij′\)⊤=Q𝐫^ij𝐫^ij⊤Q⊤=Q𝐏ij∥Q⊤\(\\mathbf\{P\}^\{\\parallel\}\_\{ij\}\)^\{\\prime\}=\\widehat\{\\mathbf\{r\}\}^\{\\prime\}\_\{ij\}\(\\widehat\{\\mathbf\{r\}\}^\{\\prime\}\_\{ij\}\)^\{\\top\}=Q\\widehat\{\\mathbf\{r\}\}\_\{ij\}\\widehat\{\\mathbf\{r\}\}\_\{ij\}^\{\\top\}Q^\{\\top\}=Q\\mathbf\{P\}^\{\\parallel\}\_\{ij\}Q^\{\\top\}\(126\)Since𝐏ij⟂=𝐈n−𝐏ij∥\\mathbf\{P\}^\{\\perp\}\_\{ij\}=\\mathbf\{I\}\_\{n\}\-\\mathbf\{P\}^\{\\parallel\}\_\{ij\}:
\(𝐏ij⟂\)′=𝐈n−Q𝐏ij∥Q⊤=Q\(𝐈n−𝐏ij∥\)Q⊤=Q𝐏ij⟂Q⊤\(\\mathbf\{P\}^\{\\perp\}\_\{ij\}\)^\{\\prime\}=\\mathbf\{I\}\_\{n\}\-Q\\mathbf\{P\}^\{\\parallel\}\_\{ij\}Q^\{\\top\}=Q\\left\(\\mathbf\{I\}\_\{n\}\-\\mathbf\{P\}^\{\\parallel\}\_\{ij\}\\right\)Q^\{\\top\}=Q\\mathbf\{P\}^\{\\perp\}\_\{ij\}Q^\{\\top\}\(127\)The channel maps𝐌ij∥\\mathbf\{M\}^\{\\parallel\}\_\{ij\}and𝐌ij⟂\\mathbf\{M\}^\{\\perp\}\_\{ij\}are constructed from invariant edge features and are therefore unchanged by theO\(n\)O\(n\)action\. Hence:
𝒯i←j′\(Q𝐕j\)\\displaystyle\\mathcal\{T\}^\{\\prime\}\_\{i\\leftarrow j\}\(Q\\mathbf\{V\}\_\{j\}\)=\(Q𝐏ij∥Q⊤\)\(Q𝐕j\)𝐌ij∥\+\(Q𝐏ij⟂Q⊤\)\(Q𝐕j\)𝐌ij⟂\\displaystyle=\(Q\\mathbf\{P\}^\{\\parallel\}\_\{ij\}Q^\{\\top\}\)\(Q\\mathbf\{V\}\_\{j\}\)\\mathbf\{M\}^\{\\parallel\}\_\{ij\}\+\(Q\\mathbf\{P\}^\{\\perp\}\_\{ij\}Q^\{\\top\}\)\(Q\\mathbf\{V\}\_\{j\}\)\\mathbf\{M\}^\{\\perp\}\_\{ij\}\(128\)=Q\(𝐏ij∥𝐕j𝐌ij∥\+𝐏ij⟂𝐕j𝐌ij⟂\)=Q𝒯i←j\(𝐕j\)\\displaystyle=Q\\left\(\\mathbf\{P\}^\{\\parallel\}\_\{ij\}\\mathbf\{V\}\_\{j\}\\mathbf\{M\}^\{\\parallel\}\_\{ij\}\+\\mathbf\{P\}^\{\\perp\}\_\{ij\}\\mathbf\{V\}\_\{j\}\\mathbf\{M\}^\{\\perp\}\_\{ij\}\\right\)=Q\\,\\mathcal\{T\}\_\{i\\leftarrow j\}\(\\mathbf\{V\}\_\{j\}\)\(129\)Thus the Radial–Tangential Transport isO\(n\)O\(n\)\-equivariant\. ∎
### D\.4Completeness of Displacement\-Conditioned Linear Transport
Theorem[4\.3](https://arxiv.org/html/2608.28853#S4.Thmtheorem3)is proved in Appendix[B\.2](https://arxiv.org/html/2608.28853#A2.SS2)\. We do not repeat the argument here\. The proof uses the stabilizer:
H𝐫=\{Q∈O\(n\):Q𝐫=𝐫\}≅O\(n−1\)H\_\{\\mathbf\{r\}\}=\\\{Q\\in O\(n\):Q\\mathbf\{r\}=\\mathbf\{r\}\\\}\\cong O\(n\-1\)of a non\-zero displacement and characterizes the commutant of its action onℝn⊗ℝcv\\mathbb\{R\}^\{n\}\\otimes\\mathbb\{R\}^\{c\_\{v\}\}\. This yields precisely the two spatial projectors𝐏∥\\mathbf\{P\}^\{\\parallel\}and𝐏⟂\\mathbf\{P\}^\{\\perp\}, each accompanied by an arbitrary invariant endomorphism of the vector\-channel space\.
### D\.5Proof of Theorem[5\.1](https://arxiv.org/html/2608.28853#S5.Thmtheorem1):E\(n\)E\(n\)\-Equivariance of the ESNN Layer
Theorem[5\.1](https://arxiv.org/html/2608.28853#S5.Thmtheorem1)\(E\(n\)E\(n\)\-Equivariance\)\.Assume that the graph topology is fixed or constructed fromE\(n\)E\(n\)\-invariant geometric quantities, that the edge attributes areO\(n\)O\(n\)\-invariant, and that every non\-self interaction for which𝐫^ij\\widehat\{\\mathbf\{r\}\}\_\{ij\}is used satisfies𝐫ij≠𝟎\\mathbf\{r\}\_\{ij\}\\neq\\mathbf\{0\}\. Assume further that the transport maps satisfy the conditions of Proposition[3\.2](https://arxiv.org/html/2608.28853#S3.Thmtheorem2)\. In the absence of explicit symmetry\-relaxing inputs, the ESNN layer defined by Equations[28](https://arxiv.org/html/2608.28853#S5.E28)–[37](https://arxiv.org/html/2608.28853#S5.E37)isE\(n\)E\(n\)\-equivariant\.
###### Proof\.
Let\(Q,𝐭\)∈E\(n\)\(Q,\\mathbf\{t\}\)\\in E\(n\), withQ∈O\(n\)Q\\in O\(n\)and𝐭∈ℝn\\mathbf\{t\}\\in\\mathbb\{R\}^\{n\}\. We denote quantities evaluated from the transformed input by an overbar:
𝐱¯i=Q𝐱i\+𝐭,𝐕¯i=Q𝐕i,𝐬¯i=𝐬i\\overline\{\\mathbf\{x\}\}\_\{i\}=Q\\mathbf\{x\}\_\{i\}\+\\mathbf\{t\},\\qquad\\overline\{\\mathbf\{V\}\}\_\{i\}=Q\\mathbf\{V\}\_\{i\},\\qquad\\overline\{\\mathbf\{s\}\}\_\{i\}=\\mathbf\{s\}\_\{i\}\(130\)
By assumption, the graph construction isE\(n\)E\(n\)\-invariant, so the neighborhood sets𝒩\(i\)\\mathcal\{N\}\(i\)are unchanged by this transformation\.
Invariant edge context\.For every non\-self interaction:
𝐫¯ij=𝐱¯i−𝐱¯j=Q\(𝐱i−𝐱j\)=Q𝐫ij\\overline\{\\mathbf\{r\}\}\_\{ij\}=\\overline\{\\mathbf\{x\}\}\_\{i\}\-\\overline\{\\mathbf\{x\}\}\_\{j\}=Q\(\\mathbf\{x\}\_\{i\}\-\\mathbf\{x\}\_\{j\}\)=Q\\mathbf\{r\}\_\{ij\}\(131\)SinceQQis orthogonal:
‖𝐫¯ij‖=‖𝐫ij‖,𝐫¯^ij=Q𝐫^ij\\\|\\overline\{\\mathbf\{r\}\}\_\{ij\}\\\|=\\\|\\mathbf\{r\}\_\{ij\}\\\|,\\qquad\\widehat\{\\overline\{\\mathbf\{r\}\}\}\_\{ij\}=Q\\widehat\{\\mathbf\{r\}\}\_\{ij\}\(132\)For every vector channel:
‖Q𝐯i,c‖=‖𝐯i,c‖\\\|Q\\mathbf\{v\}\_\{i,c\}\\\|=\\\|\\mathbf\{v\}\_\{i,c\}\\\|\(133\)and therefore:
𝐧\(𝐕¯i\)=𝐧\(𝐕i\)\\mathbf\{n\}\(\\overline\{\\mathbf\{V\}\}\_\{i\}\)=\\mathbf\{n\}\(\\mathbf\{V\}\_\{i\}\)\(134\)Likewise:
𝐕¯i⊤𝐕¯j\\displaystyle\\overline\{\\mathbf\{V\}\}\_\{i\}^\{\\top\}\\overline\{\\mathbf\{V\}\}\_\{j\}=𝐕i⊤Q⊤Q𝐕j=𝐕i⊤𝐕j,\\displaystyle=\\mathbf\{V\}\_\{i\}^\{\\top\}Q^\{\\top\}Q\\mathbf\{V\}\_\{j\}=\\mathbf\{V\}\_\{i\}^\{\\top\}\\mathbf\{V\}\_\{j\},\(135\)𝐕¯i⊤𝐫¯^ij\\displaystyle\\overline\{\\mathbf\{V\}\}\_\{i\}^\{\\top\}\\widehat\{\\overline\{\\mathbf\{r\}\}\}\_\{ij\}=𝐕i⊤Q⊤Q𝐫^ij=𝐕i⊤𝐫^ij,\\displaystyle=\\mathbf\{V\}\_\{i\}^\{\\top\}Q^\{\\top\}Q\\widehat\{\\mathbf\{r\}\}\_\{ij\}=\\mathbf\{V\}\_\{i\}^\{\\top\}\\widehat\{\\mathbf\{r\}\}\_\{ij\},\(136\)𝐕¯j⊤𝐫¯^ij\\displaystyle\\overline\{\\mathbf\{V\}\}\_\{j\}^\{\\top\}\\widehat\{\\overline\{\\mathbf\{r\}\}\}\_\{ij\}=𝐕j⊤𝐫^ij\\displaystyle=\\mathbf\{V\}\_\{j\}^\{\\top\}\\widehat\{\\mathbf\{r\}\}\_\{ij\}\(137\)The scalar node features and optional edge attributes are invariant by assumption\. Consequently every component of Equation[28](https://arxiv.org/html/2608.28853#S5.E28)is unchanged𝐳¯ij=𝐳ij\\overline\{\\mathbf\{z\}\}\_\{ij\}=\\mathbf\{z\}\_\{ij\}\.
Vector and scalar messages\.By Proposition[3\.2](https://arxiv.org/html/2608.28853#S3.Thmtheorem2):
𝐦¯i←jV=𝒯¯i←j\(𝐕¯j\)=Q𝒯i←j\(𝐕j\)=Q𝐦i←jV\\overline\{\\mathbf\{m\}\}^\{V\}\_\{i\\leftarrow j\}=\\overline\{\\mathcal\{T\}\}\_\{i\\leftarrow j\}\(\\overline\{\\mathbf\{V\}\}\_\{j\}\)=Q\\mathcal\{T\}\_\{i\\leftarrow j\}\(\\mathbf\{V\}\_\{j\}\)=Q\\mathbf\{m\}^\{V\}\_\{i\\leftarrow j\}\(138\)This argument applies to any transport family satisfying the covariance conditions of Section[3](https://arxiv.org/html/2608.28853#S3), including feature\-conditioned families\. The scalar message is:
𝐦i←js=𝐬j\+ϕs\(𝐳ijs\)\\mathbf\{m\}^\{s\}\_\{i\\leftarrow j\}=\\mathbf\{s\}\_\{j\}\+\\phi\_\{s\}\(\\mathbf\{z\}^\{s\}\_\{ij\}\)\(139\)where every argument of𝐳ijs\\mathbf\{z\}^\{s\}\_\{ij\}is invariant\. Therefore:
𝐦¯i←js=𝐦i←js\\overline\{\\mathbf\{m\}\}^\{s\}\_\{i\\leftarrow j\}=\\mathbf\{m\}^\{s\}\_\{i\\leftarrow j\}\(140\)
Normalized transport diffusion\.The optional edge weightsωij\\omega\_\{ij\}are invariant scalars\. Because the graph topology is unchanged, every degree quantity constructed from these weights is also invariant\. Hence:
ν¯ij=νij,ν¯ii=νii\\overline\{\\nu\}\_\{ij\}=\\nu\_\{ij\},\\qquad\\overline\{\\nu\}\_\{ii\}=\\nu\_\{ii\}\(141\)Using Equation[138](https://arxiv.org/html/2608.28853#A4.E138):
𝐕¯idiff=νiiQ𝐕i\+∑j∈𝒩\(i\)νijQ𝐦i←jV=Q\(νii𝐕i\+∑j∈𝒩\(i\)νij𝐦i←jV\)=Q𝐕idiff\\displaystyle\\overline\{\\mathbf\{V\}\}^\{\\mathrm\{diff\}\}\_\{i\}=\\nu\_\{ii\}Q\\mathbf\{V\}\_\{i\}\+\\sum\_\{j\\in\\mathcal\{N\}\(i\)\}\\nu\_\{ij\}Q\\mathbf\{m\}^\{V\}\_\{i\\leftarrow j\}=Q\\left\(\\nu\_\{ii\}\\mathbf\{V\}\_\{i\}\+\\sum\_\{j\\in\\mathcal\{N\}\(i\)\}\\nu\_\{ij\}\\mathbf\{m\}^\{V\}\_\{i\\leftarrow j\}\\right\)=Q\\mathbf\{V\}^\{\\mathrm\{diff\}\}\_\{i\}\(142\)Similarly, Equation[140](https://arxiv.org/html/2608.28853#A4.E140)gives:𝐬¯idiff=𝐬idiff\\overline\{\\mathbf\{s\}\}^\{\\mathrm\{diff\}\}\_\{i\}=\\mathbf\{s\}^\{\\mathrm\{diff\}\}\_\{i\}\.
Residual feature update\.The vector\-channel mixer acts only on the multiplicity dimension:
𝐕~¯i=𝐕¯idiff𝐖V=\(Q𝐕idiff\)𝐖V=Q𝐕~i\\displaystyle\\overline\{\\widetilde\{\\mathbf\{V\}\}\}\_\{i\}=\\overline\{\\mathbf\{V\}\}^\{\\mathrm\{diff\}\}\_\{i\}\\mathbf\{W\}\_\{V\}=\(Q\\mathbf\{V\}^\{\\mathrm\{diff\}\}\_\{i\}\)\\mathbf\{W\}\_\{V\}=Q\\widetilde\{\\mathbf\{V\}\}\_\{i\}\(143\)For the radial non\-linearityσV\(𝐯\)=a\(‖𝐯‖\)𝐯\\sigma\_\{V\}\(\\mathbf\{v\}\)=a\(\\\|\\mathbf\{v\}\\\|\)\\mathbf\{v\}, orthogonality ofQQgives:
σV\(Q𝐯\)=a\(‖Q𝐯‖\)Q𝐯=a\(‖𝐯‖\)Q𝐯=QσV\(𝐯\)\\displaystyle\\sigma\_\{V\}\(Q\\mathbf\{v\}\)=a\(\\\|Q\\mathbf\{v\}\\\|\)Q\\mathbf\{v\}=a\(\\\|\\mathbf\{v\}\\\|\)Q\\mathbf\{v\}=Q\\sigma\_\{V\}\(\\mathbf\{v\}\)\(144\)Applied independently to every vector channel, this implies:
σV\(Q𝐕~i\)=QσV\(𝐕~i\)\\sigma\_\{V\}\(Q\\widetilde\{\\mathbf\{V\}\}\_\{i\}\)=Q\\sigma\_\{V\}\(\\widetilde\{\\mathbf\{V\}\}\_\{i\}\)\(145\)Therefore the residual vector update satisfies:
𝐕¯i′=Q𝐕i\+σV\(Q𝐕~i\)=Q\(𝐕i\+σV\(𝐕~i\)\)=Q𝐕i′\\displaystyle\\overline\{\\mathbf\{V\}\}^\{\\prime\}\_\{i\}=Q\\mathbf\{V\}\_\{i\}\+\\sigma\_\{V\}\(Q\\widetilde\{\\mathbf\{V\}\}\_\{i\}\)=Q\\left\(\\mathbf\{V\}\_\{i\}\+\\sigma\_\{V\}\(\\widetilde\{\\mathbf\{V\}\}\_\{i\}\)\\right\)=Q\\mathbf\{V\}^\{\\prime\}\_\{i\}\(146\)The scalar update contains only linear maps and nonlinearities acting on invariant scalar quantities\. Thus:
𝐬¯i′=𝐬i′\\overline\{\\mathbf\{s\}\}^\{\\prime\}\_\{i\}=\\mathbf\{s\}^\{\\prime\}\_\{i\}\(147\)
Coordinate update\.From Equations[146](https://arxiv.org/html/2608.28853#A4.E146)and[147](https://arxiv.org/html/2608.28853#A4.E147), we have𝐡¯iinv=𝐡iinv\\overline\{\\mathbf\{h\}\}^\{\\mathrm\{inv\}\}\_\{i\}=\\mathbf\{h\}^\{\\mathrm\{inv\}\}\_\{i\}, because vector norms are invariant underQQ\. All arguments of the coordinate network are therefore invariant, and hence:
γ¯ij=γij\\overline\{\\gamma\}\_\{ij\}=\\gamma\_\{ij\}\(148\)Using Equation[131](https://arxiv.org/html/2608.28853#A4.E131):
Δ𝐱¯i=1\|𝒩\(i\)\|∑j∈𝒩\(i\)γ¯ij𝐫¯ij=1\|𝒩\(i\)\|∑j∈𝒩\(i\)γijQ𝐫ij=QΔ𝐱i\\displaystyle\\overline\{\\Delta\\mathbf\{x\}\}\_\{i\}=\\frac\{1\}\{\|\\mathcal\{N\}\(i\)\|\}\\sum\_\{j\\in\\mathcal\{N\}\(i\)\}\\overline\{\\gamma\}\_\{ij\}\\overline\{\\mathbf\{r\}\}\_\{ij\}=\\frac\{1\}\{\|\\mathcal\{N\}\(i\)\|\}\\sum\_\{j\\in\\mathcal\{N\}\(i\)\}\\gamma\_\{ij\}Q\\mathbf\{r\}\_\{ij\}=Q\\Delta\\mathbf\{x\}\_\{i\}\(149\)Finally:
𝐱¯i′=𝐱¯i\+Δ𝐱¯i=Q𝐱i\+𝐭\+QΔ𝐱i=Q\(𝐱i\+Δ𝐱i\)\+𝐭=Q𝐱i′\+𝐭\\displaystyle\\overline\{\\mathbf\{x\}\}^\{\\prime\}\_\{i\}=\\overline\{\\mathbf\{x\}\}\_\{i\}\+\\overline\{\\Delta\\mathbf\{x\}\}\_\{i\}=Q\\mathbf\{x\}\_\{i\}\+\\mathbf\{t\}\+Q\\Delta\\mathbf\{x\}\_\{i\}=Q\\left\(\\mathbf\{x\}\_\{i\}\+\\Delta\\mathbf\{x\}\_\{i\}\\right\)\+\\mathbf\{t\}=Q\\mathbf\{x\}^\{\\prime\}\_\{i\}\+\\mathbf\{t\}\(150\)Equations[146](https://arxiv.org/html/2608.28853#A4.E146),[147](https://arxiv.org/html/2608.28853#A4.E147), and[150](https://arxiv.org/html/2608.28853#A4.E150)are exactly the transformation laws required forE\(n\)E\(n\)\-equivariance\. ∎
### D\.6Equivariance of the Dynamical Extension
###### Corollary D\.1\(Equivariance of the Velocity Update\)\.
Suppose𝐮i↦Q𝐮i\\mathbf\{u\}\_\{i\}\\mapsto Q\\mathbf\{u\}\_\{i\}underQ∈O\(n\)Q\\in O\(n\)and letaia\_\{i\}be an invariant scalar\. IfΔ𝐱i↦QΔ𝐱i\\Delta\\mathbf\{x\}\_\{i\}\\mapsto Q\\Delta\\mathbf\{x\}\_\{i\}, then Equation[40](https://arxiv.org/html/2608.28853#S5.E40):
𝐮i′=ai𝐮i\+Δ𝐱i,𝐱i′=𝐱i\+𝐮i′\\mathbf\{u\}^\{\\prime\}\_\{i\}=a\_\{i\}\\mathbf\{u\}\_\{i\}\+\\Delta\\mathbf\{x\}\_\{i\},\\qquad\\mathbf\{x\}^\{\\prime\}\_\{i\}=\\mathbf\{x\}\_\{i\}\+\\mathbf\{u\}^\{\\prime\}\_\{i\}\(151\)isE\(n\)E\(n\)\-equivariant\.
###### Proof\.
Under\(Q,𝐭\)∈E\(n\)\(Q,\\mathbf\{t\}\)\\in E\(n\):
𝐮¯i′=aiQ𝐮i\+QΔ𝐱i=Q\(ai𝐮i\+Δ𝐱i\)=Q𝐮i′\\displaystyle\\overline\{\\mathbf\{u\}\}^\{\\prime\}\_\{i\}=a\_\{i\}Q\\mathbf\{u\}\_\{i\}\+Q\\Delta\\mathbf\{x\}\_\{i\}=Q\\left\(a\_\{i\}\\mathbf\{u\}\_\{i\}\+\\Delta\\mathbf\{x\}\_\{i\}\\right\)=Q\\mathbf\{u\}^\{\\prime\}\_\{i\}\(152\)Therefore:
𝐱¯i′=Q𝐱i\+𝐭\+Q𝐮i′=Q\(𝐱i\+𝐮i′\)\+𝐭=Q𝐱i′\+𝐭\\displaystyle\\overline\{\\mathbf\{x\}\}^\{\\prime\}\_\{i\}=Q\\mathbf\{x\}\_\{i\}\+\\mathbf\{t\}\+Q\\mathbf\{u\}^\{\\prime\}\_\{i\}=Q\\left\(\\mathbf\{x\}\_\{i\}\+\\mathbf\{u\}^\{\\prime\}\_\{i\}\\right\)\+\\mathbf\{t\}=Q\\mathbf\{x\}^\{\\prime\}\_\{i\}\+\\mathbf\{t\}\(153\)∎
### D\.7Proof of Theorem[6\.1](https://arxiv.org/html/2608.28853#S6.Thmtheorem1): Stabilizer Subequivariance
Theorem[6\.1](https://arxiv.org/html/2608.28853#S6.Thmtheorem1)\(Stabilizer Subequivariance\)\.Let𝐠≠𝟎\\mathbf\{g\}\\neq\\mathbf\{0\}be fixed in the ambient coordinate frame and define:O𝐠\(n\)=\{Q∈O\(n\):Q𝐠=𝐠\}O\_\{\\mathbf\{g\}\}\(n\)=\\left\\\{Q\\in O\(n\):Q\\mathbf\{g\}=\\mathbf\{g\}\\right\\\}\(154\)Forλ≠0\\lambda\\neq 0, an ESNN layer conditioned on Equation[42](https://arxiv.org/html/2608.28853#S6.E42)is equivariant under translations and everyQ∈O𝐠\(n\)Q\\in O\_\{\\mathbf\{g\}\}\(n\)\. Thus the architecture guarantees equivariance to:E𝐠\(n\)=O𝐠\(n\)⋉ℝnE\_\{\\mathbf\{g\}\}\(n\)=O\_\{\\mathbf\{g\}\}\(n\)\\ltimes\\mathbb\{R\}^\{n\}\(155\)No equivariance under transformations outsideO𝐠\(n\)O\_\{\\mathbf\{g\}\}\(n\)is enforced by the construction\. Forλ=0\\lambda=0, fullE\(n\)E\(n\)\-equivariance is recovered\.
###### Proof\.
The relaxed edge context is defined as:
𝐳ijrelaxed=\[𝐳ij,λ⟨𝐫ij,𝐠⟩\]\\mathbf\{z\}^\{\\mathrm\{relaxed\}\}\_\{ij\}=\\left\[\\mathbf\{z\}\_\{ij\},\\lambda\\langle\\mathbf\{r\}\_\{ij\},\\mathbf\{g\}\\rangle\\right\]\(156\)The original context𝐳ij\\mathbf\{z\}\_\{ij\}isO\(n\)O\(n\)\-invariant by the invariant\-edge\-context argument in Appendix[D\.5](https://arxiv.org/html/2608.28853#A4.SS5)\. It remains to characterize the transformation of the additional directional scalar\. First consider a translation𝐱i↦𝐱i\+𝐭\\mathbf\{x\}\_\{i\}\\mapsto\\mathbf\{x\}\_\{i\}\+\\mathbf\{t\}\. Relative displacements are unchanged:
\(𝐱i\+𝐭\)−\(𝐱j\+𝐭\)=𝐫ij\(\\mathbf\{x\}\_\{i\}\+\\mathbf\{t\}\)\-\(\\mathbf\{x\}\_\{j\}\+\\mathbf\{t\}\)=\\mathbf\{r\}\_\{ij\}\(157\)and therefore⟨𝐫ij,𝐠⟩\\langle\\mathbf\{r\}\_\{ij\},\\mathbf\{g\}\\rangleis translation invariant\. Now letQ∈O𝐠\(n\)Q\\in O\_\{\\mathbf\{g\}\}\(n\)\. SinceQ𝐠=𝐠Q\\mathbf\{g\}=\\mathbf\{g\}, orthogonality also impliesQ⊤𝐠=𝐠Q^\{\\top\}\\mathbf\{g\}=\\mathbf\{g\}\. Hence:
⟨Q𝐫ij,𝐠⟩=⟨𝐫ij,Q⊤𝐠⟩\\displaystyle\\left\\langle Q\\mathbf\{r\}\_\{ij\},\\mathbf\{g\}\\right\\rangle=\\left\\langle\\mathbf\{r\}\_\{ij\},Q^\{\\top\}\\mathbf\{g\}\\right\\rangle=⟨𝐫ij,𝐠⟩\\displaystyle=\\left\\langle\\mathbf\{r\}\_\{ij\},\\mathbf\{g\}\\right\\rangle\(158\)Thus the full relaxed context is invariant:
𝐳¯ijrelaxed=𝐳ijrelaxed∀Q∈O𝐠\(n\)\\overline\{\\mathbf\{z\}\}^\{\\mathrm\{relaxed\}\}\_\{ij\}=\\mathbf\{z\}^\{\\mathrm\{relaxed\}\}\_\{ij\}\\qquad\\forall Q\\in O\_\{\\mathbf\{g\}\}\(n\)\(159\)
The relaxed scalar is used only to parameterize scalar coefficients such as transport gates or aggregation weights\. Since these coefficients remain unchanged underO𝐠\(n\)O\_\{\\mathbf\{g\}\}\(n\), while the spatial transport operators retain the covariance laws of Section[4](https://arxiv.org/html/2608.28853#S4), every step in the proof of Theorem[5\.1](https://arxiv.org/html/2608.28853#S5.Thmtheorem1)remains valid after restrictingQQfromO\(n\)O\(n\)toO𝐠\(n\)O\_\{\\mathbf\{g\}\}\(n\)\. Together with translation equivariance, this proves equivariance toE𝐠\(n\)E\_\{\\mathbf\{g\}\}\(n\)\.
It remains to clarify what happens outside the stabilizer\. LetQ∉O𝐠\(n\)Q\\notin O\_\{\\mathbf\{g\}\}\(n\)\. Then:
Q⊤𝐠≠𝐠Q^\{\\top\}\\mathbf\{g\}\\neq\\mathbf\{g\}\(160\)The two linear functionals:
𝐫⟼⟨𝐫,Q⊤𝐠⟩and𝐫⟼⟨𝐫,𝐠⟩\\mathbf\{r\}\\longmapsto\\langle\\mathbf\{r\},Q^\{\\top\}\\mathbf\{g\}\\rangle\\qquad\\text\{and\}\\qquad\\mathbf\{r\}\\longmapsto\\langle\\mathbf\{r\},\\mathbf\{g\}\\rangle\(161\)are therefore distinct\. Consequently, there exists a displacement𝐫\\mathbf\{r\}such that:
⟨Q𝐫,𝐠⟩≠⟨𝐫,𝐠⟩\\langle Q\\mathbf\{r\},\\mathbf\{g\}\\rangle\\neq\\langle\\mathbf\{r\},\\mathbf\{g\}\\rangle\(162\)
Sinceλ≠0\\lambda\\neq 0, Equation[162](https://arxiv.org/html/2608.28853#A4.E162)also implies:
λ⟨Q𝐫,𝐠⟩≠λ⟨𝐫,𝐠⟩\\lambda\\langle Q\\mathbf\{r\},\\mathbf\{g\}\\rangle\\neq\\lambda\\langle\\mathbf\{r\},\\mathbf\{g\}\\rangle\(163\)Hence the relaxed edge context is not structurally invariant under such a transformation, so the architecture does not enforce equivariance outsideO𝐠\(n\)O\_\{\\mathbf\{g\}\}\(n\)\. Particular learned parameters may ignore the directional feature and exhibit a larger symmetry; the theorem concerns the symmetry guaranteed by the architecture\.
Finally, whenλ=0\\lambda=0, the second component of Equation[156](https://arxiv.org/html/2608.28853#A4.E156)vanishes identically, independently of𝐠\\mathbf\{g\}\. The edge context then reduces to the original invariant context𝐳ij\\mathbf\{z\}\_\{ij\}, and Theorem[5\.1](https://arxiv.org/html/2608.28853#S5.Thmtheorem1)recovers fullE\(n\)E\(n\)\-equivariance\. ∎
## Appendix EExperimental Details
This section provides additional details on dataset construction, prediction targets, and evaluation procedures for the experiments presented in Section[7](https://arxiv.org/html/2608.28853#S7)\. We also report the fullQM9molecular\-property benchmark, providing a property\-wise view of ESNN performance\.
### E\.1Charged N\-Body Dynamics
The first task is a three\-dimensional charged N\-body system, a standard benchmark for equivariant dynamics prediction\([Fuchs et al\., 2020](https://arxiv.org/html/2608.28853#bib.bib19);[Satorras et al\., 2021](https://arxiv.org/html/2608.28853#bib.bib1)\)\. The system contains five particles, each associated with a position, velocity, and positive or negative charge\. Their trajectories are determined by pairwise attractive or repulsive interactions, and the learning problem consists of predicting the future particle coordinates from an earlier system state\.
Dataset and Physical System\.The benchmark consists ofN=5N=5particles moving in three\-dimensional Euclidean space\. Particleiiis described by its position𝐱i∈ℝ3\\mathbf\{x\}\_\{i\}\\in\\mathbb\{R\}^\{3\}, velocity𝐯i∈ℝ3\\mathbf\{v\}\_\{i\}\\in\\mathbb\{R\}^\{3\}, and chargeqi∈\{−1,\+1\}q\_\{i\}\\in\\\{\-1,\+1\\\}\. Charges are sampled independently with equal probability, and particles interact through pairwise Coulomb forces\. The original dataset contains 50,000 training trajectories, 2,000 validation trajectories, and 2,000 test trajectories\. Following[Satorras et al\. \(2021\)](https://arxiv.org/html/2608.28853#bib.bib1), we use the reduced\-data setting with 3,000 training, 2,000 validation, and 2,000 test trajectories\.
Initial Conditions and Numerical Integration\.Initial positions are sampled independently as:
𝐱i0∼𝒩\(𝟎,I3\)\\mathbf\{x\}\_\{i\}^\{0\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},I\_\{3\}\)\(164\)while initial velocity directions are sampled from a Gaussian distribution and subsequently normalized such that‖𝐯i0‖2=0\.5\\\|\\mathbf\{v\}\_\{i\}^\{0\}\\\|\_\{2\}=0\.5\. Trajectories are generated using leapfrog integration with numerical time stepδt=10−3\\delta t=10^\{\-3\}\. Each trajectory contains 5,000 integration steps, and particle states are recorded every 100 steps, corresponding to0\.10\.1simulation\-time units between consecutive stored frames\. For numerical stability, every Cartesian component of the interaction force is clipped to the interval\[−100,100\]\[\-100,100\]before the velocity update\.
Graph Construction\.Each system state is represented as a complete directed graph without self\-loops:
ℰ=\{\(i,j\)\|i,j∈\{1,…,5\},i≠j\}\\mathcal\{E\}=\\left\\\{\(i,j\)\\;\\middle\|\\;i,j\\in\\\{1,\\ldots,5\\\},\\;i\\neq j\\right\\\}\(165\)Every graph therefore contains five nodes and twenty directed edges\. The dataset stores the charge product associated with the interactionj→ij\\rightarrow i:
eij=qiqje\_\{ij\}=q\_\{i\}q\_\{j\}\(166\)ESNN uses the particle charges as invariant scalar node features, while the stored charge products are retained for compatibility with the standard N\-body data format\. Positions and velocities are treated as covariant three\-dimensional vectors\.
Prediction Task and Evaluation\.For each trajectory, the model receives the particle state at stored frame66and predicts the coordinates at stored frame88:
\{𝐱i\(6\),𝐯i\(6\),qi\}i=15⟼\{𝐱^i\(8\)\}i=15\\left\\\{\\mathbf\{x\}\_\{i\}^\{\(6\)\},\\mathbf\{v\}\_\{i\}^\{\(6\)\},q\_\{i\}\\right\\\}\_\{i=1\}^\{5\}\\longmapsto\\left\\\{\\widehat\{\\mathbf\{x\}\}\_\{i\}^\{\(8\)\}\\right\\\}\_\{i=1\}^\{5\}\(167\)The prediction horizon therefore corresponds to200200numerical integration steps, or0\.20\.2simulation\-time units\. Models are trained by minimizing the mean squared coordinate error:
ℒNBody=1BN∑b=1B∑i=1N‖𝐱^b,i\(8\)−𝐱b,i\(8\)‖22\\mathcal\{L\}\_\{\\mathrm\{NBody\}\}=\\frac\{1\}\{BN\}\\sum\_\{b=1\}^\{B\}\\sum\_\{i=1\}^\{N\}\\left\\\|\\widehat\{\\mathbf\{x\}\}\_\{b,i\}^\{\(8\)\}\-\\mathbf\{x\}\_\{b,i\}^\{\(8\)\}\\right\\\|\_\{2\}^\{2\}\(168\)Validation MSE is used for model selection, and the checkpoint obtaining the lowest validation error is evaluated on the held\-out test partition\.
### E\.2Gravity\-Augmented N\-Body Dynamics
The second N\-body task extends the charged\-particle benchmark of Appendix[E\.1](https://arxiv.org/html/2608.28853#A5.SS1)with a uniform gravitational field\. Unless stated otherwise, we use the same particle system, initial\-condition distribution, graph construction, numerical integration scheme, and coordinate prediction objective as in the standard benchmark\. The key difference is that gravity introduces a preferred ambient direction, reducing the symmetry of the dynamics while preserving translations\. This setting therefore provides a direct test of the controlled symmetry relaxation introduced in Section[6](https://arxiv.org/html/2608.28853#S6)\.
Figure 5:*Charged N\-body dynamics with and without a preferred ambient direction\.*\(a\) In the standard benchmark, particle motion is governed only by pairwise Coulomb interactions and the dynamics retain full Euclidean symmetry\. \(b\) The gravity variant additionally applies the uniform acceleration𝐚𝐠\\mathbf\{a\}\_\{\\mathbf\{g\}\}along the preferred direction𝐠^true\\widehat\{\\mathbf\{g\}\}\_\{\\mathrm\{true\}\}\. Translations remain symmetries, while the orthogonal symmetry is reduced to the subgroupO𝐠\(3\)O\_\{\\mathbf\{g\}\}\(3\)that preserves the gravity axis\.Gravity\-Augmented Dynamics\.In addition to the pairwise Coulomb interactions of Appendix[E\.1](https://arxiv.org/html/2608.28853#A5.SS1), every particle experiences the uniform gravitational acceleration:
𝐚𝐠=\(0,0,−9\.81\)⊤\\mathbf\{a\}\_\{\\mathbf\{g\}\}=\(0,0,\-9\.81\)^\{\\top\}\(169\)The Coulomb contribution is computed and component\-wise clipped using the same procedure as in the standard benchmark, after which the gravitational acceleration is added to the particle dynamics\. The corresponding unit gravity direction is:
𝐠^true=\(0,0,−1\)⊤\\widehat\{\\mathbf\{g\}\}\_\{\\mathrm\{true\}\}=\(0,0,\-1\)^\{\\top\}\(170\)Because the gravitational field is uniform and the simulation contains no fixed spatial boundary, translations remain symmetries of the system\. The orthogonal symmetry is instead restricted to transformations that preserve the gravity direction:
E𝐠\(3\)=O𝐠\(3\)⋉ℝ3,O𝐠\(3\)=\{Q∈O\(3\)\|Q𝐠^true=𝐠^true\}E\_\{\\mathbf\{g\}\}\(3\)=O\_\{\\mathbf\{g\}\}\(3\)\\ltimes\\mathbb\{R\}^\{3\},\\qquad O\_\{\\mathbf\{g\}\}\(3\)=\\left\\\{Q\\in O\(3\)\\;\\middle\|\\;Q\\widehat\{\\mathbf\{g\}\}\_\{\\mathrm\{true\}\}=\\widehat\{\\mathbf\{g\}\}\_\{\\mathrm\{true\}\}\\right\\\}\(171\)withO𝐠\(3\)≅O\(2\)O\_\{\\mathbf\{g\}\}\(3\)\\cong O\(2\)\. Orthogonal transformations that change the gravity direction are therefore no longer symmetries of the dynamics\.
We generate a separate gravity\-augmented dataset using the same initial position, velocity, and charge distributions as Appendix[E\.1](https://arxiv.org/html/2608.28853#A5.SS1)\. The dataset contains 3,000 training, 2,000 validation, and 2,000 test trajectories generated with random seed4343, using the same integration and frame\-sampling convention as the standard benchmark\. No observation noise is added\. The graph topology is also unchanged\. ESNN receives the particle charges as invariant scalar node features together with the covariant positions and velocities; the stored charge productsqiqjq\_\{i\}q\_\{j\}are retained for compatibility with the standard N\-body baseline data format\.
Prediction Protocol\.Unlike the standard task, the gravity benchmark uses an earlier observation time and a longer prediction horizon\. For each trajectory, the model observes stored frame00and predicts the particle coordinates at stored frame1010:
\{𝐱i\(0\),𝐯i\(0\),qi\}i=15⟼\{𝐱^i\(10\)\}i=15\\left\\\{\\mathbf\{x\}^\{\(0\)\}\_\{i\},\\mathbf\{v\}^\{\(0\)\}\_\{i\},q\_\{i\}\\right\\\}\_\{i=1\}^\{5\}\\longmapsto\\left\\\{\\widehat\{\\mathbf\{x\}\}^\{\(10\)\}\_\{i\}\\right\\\}\_\{i=1\}^\{5\}\(172\)Under the dataset sampling convention, stored frame00is the first recorded state after 100 integration steps, while stored frame1010corresponds to the state after 1,100 steps\. The prediction horizon is therefore one simulation\-time unit\.
Using an early observation time is important for the intended diagnostic\. As the trajectory evolves, the accumulated free\-fall velocity itself becomes a covariant cue for the gravity axis and can reveal directional information even to a strictlyE\(3\)E\(3\)\-equivariant architecture\. Observing the system early reduces this cue, while the longer prediction horizon gives the gravitational field sufficient time to influence the target coordinates\. Training, checkpoint selection, and test evaluation otherwise follow the coordinate\-MSE protocol of Appendix[E\.1](https://arxiv.org/html/2608.28853#A5.SS1)\.
Preferred\-Direction Variants\.We compare three matched ESNN variants that differ only in how the preferred ambient direction is represented\. In the*None*setting, no preferred direction or symmetry\-relaxation coefficient is introduced, and the model remains fullyE\(3\)E\(3\)\-equivariant\. In the*Fixed*setting, the true unit gravity direction𝐠^true\\widehat\{\\mathbf\{g\}\}\_\{\\mathrm\{true\}\}is supplied as a non\-trainable global vector\. In the*Learned*setting, the preferred direction is instead represented by a single trainable global vector𝐠θ∈ℝ3\\mathbf\{g\}\_\{\\theta\}\\in\\mathbb\{R\}^\{3\}, shared across all layers and initialized as:
𝐠θ\(0\)=0\.01ϵ,ϵ∼𝒩\(𝟎,𝐈3\)\\mathbf\{g\}\_\{\\theta\}^\{\(0\)\}=0\.01\\,\\bm\{\\epsilon\},\\qquad\\bm\{\\epsilon\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\_\{3\}\)\(173\)so that its orientation must be inferred from the observed dynamics\.
For each directed interactionj→ij\\rightarrow i, an active symmetry\-relaxation pathway augments the invariant edge context with the directional scalar:
zij𝐠,\(ℓ\)=λ𝐠\(ℓ\)⟨𝐫ij,𝐠⟩,𝐫ij=𝐱i−𝐱jz\_\{ij\}^\{\\mathbf\{g\},\(\\ell\)\}=\\lambda\_\{\\mathbf\{g\}\}^\{\(\\ell\)\}\\left\\langle\\mathbf\{r\}\_\{ij\},\\mathbf\{g\}\\right\\rangle,\\qquad\\mathbf\{r\}\_\{ij\}=\\mathbf\{x\}\_\{i\}\-\\mathbf\{x\}\_\{j\}\(174\)where𝐠=𝐠^true\\mathbf\{g\}=\\widehat\{\\mathbf\{g\}\}\_\{\\mathrm\{true\}\}for*Fixed*and𝐠=𝐠θ\\mathbf\{g\}=\\mathbf\{g\}\_\{\\theta\}for*Learned*\. The preferred direction is shared globally, whereas the implementation uses a separate relaxation coefficientλ𝐠\(ℓ\)\\lambda\_\{\\mathbf\{g\}\}^\{\(\\ell\)\}for each active transport layer\. Every coefficient is initialized exactly at zero, so the directional pathway is initially inactive and fullE\(3\)E\(3\)\-equivariance is recovered at initialization\. During training, the coefficients may depart from zero when the preferred direction is useful for the prediction task\. In the*Learned*variant, this also enables gradients to update𝐠θ\\mathbf\{g\}\_\{\\theta\}and infer its orientation from the data\. No additional penalty on𝐠θ\\mathbf\{g\}\_\{\\theta\}orλ𝐠\(ℓ\)\\lambda\_\{\\mathbf\{g\}\}^\{\(\\ell\)\}is used in this benchmark; the initial symmetry prior is imposed through the zero initialization of the relaxation coefficients\.
Learned\-Direction Diagnostic\.Prediction error alone does not establish whether the*Learned*variant has recovered the physical symmetry\-breaking direction\. We therefore measure the orientation of the learned global vector relative to the true gravity axis\. Because the directional pathway depends onλ𝐠\(ℓ\)⟨𝐫ij,𝐠θ⟩\\lambda\_\{\\mathbf\{g\}\}^\{\(\\ell\)\}\\langle\\mathbf\{r\}\_\{ij\},\\mathbf\{g\}\_\{\\theta\}\\rangle, simultaneously reversing the sign of𝐠θ\\mathbf\{g\}\_\{\\theta\}and the relaxation coefficients leaves the directional contribution unchanged\. We consequently use the sign\-independent alignment:
A𝐠=\|⟨𝐠θ‖𝐠θ‖2,𝐠^true⟩\|A\_\{\\mathbf\{g\}\}=\\left\|\\left\\langle\\frac\{\\mathbf\{g\}\_\{\\theta\}\}\{\\\|\\mathbf\{g\}\_\{\\theta\}\\\|\_\{2\}\},\\widehat\{\\mathbf\{g\}\}\_\{\\mathrm\{true\}\}\\right\\rangle\\right\|\(175\)whereA𝐠=1A\_\{\\mathbf\{g\}\}=1corresponds to recovery of the gravity axis up to sign andA𝐠=0A\_\{\\mathbf\{g\}\}=0to an orthogonal direction\.
We report this alignment together with the effective directional scalemaxℓ\|λ𝐠\(ℓ\)\|‖𝐠‖2\\max\_\{\\ell\}\|\\lambda\_\{\\mathbf\{g\}\}^\{\(\\ell\)\}\|\\\|\\mathbf\{g\}\\\|\_\{2\}\. The magnitude of𝐠θ\\mathbf\{g\}\_\{\\theta\}alone is not meaningful because the directional signal depends jointly on𝐠θ\\mathbf\{g\}\_\{\\theta\}andλ𝐠\(ℓ\)\\lambda\_\{\\mathbf\{g\}\}^\{\(\\ell\)\}\. High alignment is therefore interpreted as evidence of recovered directional structure only when the directional pathway is also active\.
### E\.3ModelNet40 Point\-Cloud Classification
ModelNet40\([Wu et al\., 2015](https://arxiv.org/html/2608.28853#bib.bib34)\)tests graph\-level classification under controlled changes in object orientation\. The dataset contains CAD models from 40 object categories; each object is represented as a point cloud and then as a local geometric graph\.
Figure 6:*ModelNet40 graph construction and rotation\-generalization protocol\.*\(a\)–\(c\) Each CAD model is sampled as a point cloud and converted into a localkk\-nearest\-neighbor graph\. \(d\) Classification uses an invariant graph\-level readout, so a global rotation of the input should leave the predicted class unchanged\. \(e\)–\(g\) We evaluate thez/zz/z,z/SO\(3\)z/\\mathrm\{SO\}\(3\), andSO\(3\)/SO\(3\)\\mathrm\{SO\}\(3\)/\\mathrm\{SO\}\(3\)train/test protocols\. Thez/SO\(3\)z/\\mathrm\{SO\}\(3\)setting directly tests generalization to arbitrary three\-dimensional orientations that are not observed during training\.Point Cloud and Graph Construction\.For an input object, letX=\{𝐱1,…,𝐱N\},𝐱i∈ℝ3X=\\left\\\{\\mathbf\{x\}\_\{1\},\\ldots,\\mathbf\{x\}\_\{N\}\\right\\\},\\mathbf\{x\}\_\{i\}\\in\\mathbb\{R\}^\{3\}denote its sampled point cloud\. We construct akk\-nearest\-neighbor graph using Euclidean distances in the input geometry\. Geometric interactions are expressed through the relative displacement:
𝐫ij=𝐱i−𝐱j\\mathbf\{r\}\_\{ij\}=\\mathbf\{x\}\_\{i\}\-\\mathbf\{x\}\_\{j\}\(176\)which transforms covariantly under a global rotation while its norm remains invariant\. Node\-level outputs are aggregated by an invariant graph readout for 40\-way classification\.
Rotation Protocol\.Following the standard protocol for rotation\-robust point\-cloud classification, we consider three train/test settings:
z/z\\displaystyle z/z:Rtrain∈SO\(2\)z,Rtest∈SO\(2\)z,\\displaystyle:\\quad R\_\{\\mathrm\{train\}\}\\in SO\(2\)\_\{z\},\\qquad R\_\{\\mathrm\{test\}\}\\in SO\(2\)\_\{z\},\(177\)z/SO\(3\)\\displaystyle z/\\mathrm\{SO\}\(3\):Rtrain∈SO\(2\)z,Rtest∈SO\(3\),\\displaystyle:\\quad R\_\{\\mathrm\{train\}\}\\in SO\(2\)\_\{z\},\\qquad R\_\{\\mathrm\{test\}\}\\in SO\(3\),\(178\)SO\(3\)/SO\(3\)\\displaystyle\\mathrm\{SO\}\(3\)/\\mathrm\{SO\}\(3\):Rtrain∈SO\(3\),Rtest∈SO\(3\)\\displaystyle:\\quad R\_\{\\mathrm\{train\}\}\\in SO\(3\),\\qquad R\_\{\\mathrm\{test\}\}\\in SO\(3\)\(179\)Here,SO\(2\)zSO\(2\)\_\{z\}denotes rotations around the verticalzzaxis, whereasSO\(3\)SO\(3\)denotes arbitrary three\-dimensional rotations\. Thez/zz/zregime matches the restricted train and test rotation families\. Thez/SO\(3\)z/\\mathrm\{SO\}\(3\)regime measures out\-of\-distribution rotation generalization, whileSO\(3\)/SO\(3\)\\mathrm\{SO\}\(3\)/\\mathrm\{SO\}\(3\)evaluates performance when arbitrary rotations are observed during both training and testing\.
Prediction Objective and Evaluation\.The graph representation is mapped to logits over the 40 object categories and trained with categorical cross\-entropy:
ℒcls=−1B∑b=1Blogpθ\(yb∣Xb\)\\mathcal\{L\}\_\{\\mathrm\{cls\}\}=\-\\frac\{1\}\{B\}\\sum\_\{b=1\}^\{B\}\\log p\_\{\\theta\}\\left\(y\_\{b\}\\mid X\_\{b\}\\right\)\(180\)We report test classification accuracy for each rotation protocol\. Since the target is invariant, an exactly rotation\-invariant classifier should satisfy:
f\(RX\)=f\(X\),∀R∈SO\(3\)f\(RX\)=f\(X\),\\qquad\\forall R\\in SO\(3\)\(181\)up to numerical precision\. The corresponding classification results are reported in Table[4](https://arxiv.org/html/2608.28853#S7.T4)\.
### E\.4Mesh\-Based Physical Dynamics
We use three MeshGraphNets\([Pfaff et al\., 2021](https://arxiv.org/html/2608.28853#bib.bib32)\)physical\-simulation domains:CylinderFlow,DeformingPlate, andAirfoil\. All three use unstructured meshes but differ in physical regime, state representation, and boundary geometry\. They provide complementary settings for comparing the ESNN transport families on direction\-dependent fields\.
Figure 7:*Mesh\-based physical simulation benchmarks\.*\(a\)CylinderFlow: incompressible flow around a cylindrical obstacle, shown through the velocity field\. \(b\)DeformingPlate: deformation of a hyper\-elastic plate over time, illustrated using von Mises stress\. \(c\)Airfoil: compressible flow around an airfoil, shown through the pressure field\. \(d\)–\(f\) Corresponding local views of the unstructured meshes and boundary geometries\. Together, the three domains span fluid and structural dynamics on irregular discretizations with both scalar and vector physical states\.Common Graph and Feature Representation\.At timett, a physical state is represented by a mesh:
ℳt=\(𝒱,ℰ,𝐪t\)\\mathcal\{M\}^\{t\}=\\left\(\\mathcal\{V\},\\mathcal\{E\},\\mathbf\{q\}^\{t\}\\right\)\(182\)where𝒱\\mathcal\{V\}denotes mesh vertices,ℰ\\mathcal\{E\}the mesh connectivity, and𝐪t\\mathbf\{q\}^\{t\}the dynamical state sampled at the vertices\. Mesh connections are represented as bidirectional graph edges\. For an edge\(i,j\)\(i,j\), ESNN receives the relative geometric displacement𝐫ij\\mathbf\{r\}\_\{ij\}together with invariant geometric quantities such as‖𝐫ij‖\\\|\\mathbf\{r\}\_\{ij\}\\\|\. Vector\-valued physical quantities are stored in covariant channels, whereas node types and scalar physical quantities are represented by invariant channels\. Thus the field type determines the ESNN representation, while mesh displacements supply the local geometry\.
CylinderFlow: State and Target\.CylinderFlowmodels incompressible flow around a cylindrical obstacle on a fixed two\-dimensional Eulerian mesh\. Mesh nodes represent the fluid domain and its boundaries, while node types distinguish fluid, wall, inflow, and outflow locations\. The dynamical state contains the two\-dimensional flow momentum, from which the velocity field is obtained, together with the scalar pressure field\. The model predicts the temporal change of the momentum field and the pressure at the next state\. The fixed obstacle and inflow/outflow boundaries distinguish spatial directions in the flow domain\.
DeformingPlate: State and Target\.DeformingPlateis a Lagrangian structural\-mechanics problem in which a hyper\-elastic plate is deformed by a kinematic actuator\. Each mesh vertex has a reference position and a time\-dependent world\-space position\. Node types distinguish the deformable plate from actuator vertices\. The model predicts the Lagrangian velocity used to advance the mesh coordinates and the scalar von\-Mises stressσv\\sigma\_\{v\}at every node\. The reference and world\-space geometries distinguish material configuration from the evolving deformation\.
Airfoil: State and Target\.Airfoilmodels compressible aerodynamic flow around a two\-dimensional airfoil cross\-section\. The state contains the vector momentum field together with scalar density and pressure\. The model predicts changes in momentum and density and directly estimates the pressure field\. The airfoil boundary and incident flow establish preferred directions in the simulation domain\.
One\-Step Training Objective\.Following the MeshGraphNets evaluation protocol, the model is first evaluated on the prediction of the next physical state from the current state:
ℳ^t\+1=Fθ\(ℳt\)\\widehat\{\\mathcal\{M\}\}^\{t\+1\}=F\_\{\\theta\}\\left\(\\mathcal\{M\}^\{t\}\\right\)\(183\)Training uses per\-node supervision on the task\-specific predicted dynamical quantities\. For a generic vector or scalar target𝐲it\+1\\mathbf\{y\}\_\{i\}^\{t\+1\}, the one\-step objective takes the form:
ℒstep=1\|𝒱\|∑i∈𝒱‖𝐲^it\+1−𝐲it\+1‖22\\mathcal\{L\}\_\{\\mathrm\{step\}\}=\\frac\{1\}\{\|\\mathcal\{V\}\|\}\\sum\_\{i\\in\\mathcal\{V\}\}\\left\\\|\\widehat\{\\mathbf\{y\}\}\_\{i\}^\{t\+1\}\-\\mathbf\{y\}\_\{i\}^\{t\+1\}\\right\\\|\_\{2\}^\{2\}\(184\)
Rollout Evaluation and Metrics\.A one\-step predictor is recursively applied at inference time to produce a trajectory:
ℳ^t\+h=Fθ\(ℳ^t\+h−1\),h=1,…,H\\widehat\{\\mathcal\{M\}\}^\{t\+h\}=F\_\{\\theta\}\\left\(\\widehat\{\\mathcal\{M\}\}^\{t\+h\-1\}\\right\),\\qquad h=1,\\ldots,H\(185\)Predicted states are fed back to the model, so rollout errors include accumulation across steps\. We report RMSE for one\-step prediction, a 50\-step rollout, and the complete trajectory, as shown in Table[3](https://arxiv.org/html/2608.28853#S7.T3)\.
### E\.5QM9 Molecular\-Property Prediction
As a supplementary benchmark, we evaluate invariant quantum\-chemical property prediction onQM9\([Ramakrishnan et al\., 2014](https://arxiv.org/html/2608.28853#bib.bib33)\)\. Unlike the dynamics tasks, QM9 maps a static three\-dimensional molecular geometry to a graph\-level invariant target\.
Dataset and Preprocessing\.QM9 contains 133,885 equilibrium molecular geometries together with quantum\-chemical properties computed using density functional theory\. The molecules contain hydrogen and up to nine heavy atoms selected from carbon, nitrogen, oxygen, and fluorine\. Following the standard preprocessing protocol in[Satorras et al\. \(2021\)](https://arxiv.org/html/2608.28853#bib.bib1), molecules failing geometric\-consistency checks are removed, resulting in 130,831 examples\.
Graph and Input Features\.Each molecule is represented as a geometric graph𝒢=\(𝒱,ℰ\)\\mathcal\{G\}=\(\\mathcal\{V\},\\mathcal\{E\}\)\. Atomiihas Cartesian coordinate𝐱i∈ℝ3\\mathbf\{x\}\_\{i\}\\in\\mathbb\{R\}^\{3\}and atomic number:
Zi∈\{1,6,7,8,9\}Z\_\{i\}\\in\\\{1,6,7,8,9\\\}\(186\)corresponding respectively to H, C, N, O, and F\. Atomic numbers are embedded as invariant scalar features\. Geometric interactions use relative displacements:
𝐫ij=𝐱i−𝐱j\\mathbf\{r\}\_\{ij\}=\\mathbf\{x\}\_\{i\}\-\\mathbf\{x\}\_\{j\}\(187\)ensuring that the representation is independent of absolute molecular position\. When chemical bond information is used, its embedding is included as an invariant edge attribute\.
Prediction Targets\.We evaluate the twelve standard quantum\-chemical properties:
α,Δϵ,ϵHOMO,ϵLUMO,μ,Cv,G,H,⟨R2⟩,U,U0,ZPVE\\alpha,\\;\\Delta\\epsilon,\\;\\epsilon\_\{\\mathrm\{HOMO\}\},\\;\\epsilon\_\{\\mathrm\{LUMO\}\},\\;\\mu,\\;C\_\{v\},\\;G,\\;H,\\;\\langle R^\{2\}\\rangle,\\;U,\\;U\_\{0\},\\;\\mathrm\{ZPVE\}\(188\)These correspond to isotropic polarizability, the HOMO–LUMO energy gap, HOMO and LUMO energies, dipole moment, heat capacity, free energy, enthalpy, electronic spatial extent, internal energies at298\.15298\.15K and00K, and zero\-point vibrational energy\. A separate model is trained for each target\. Performance is measured using mean absolute error:
MAE=1\|𝒟test\|∑m∈𝒟test\|y^m−ym\|\\operatorname\{MAE\}=\\frac\{1\}\{\|\\mathcal\{D\}\_\{\\mathrm\{test\}\}\|\}\\sum\_\{m\\in\\mathcal\{D\}\_\{\\mathrm\{test\}\}\}\\left\|\\widehat\{y\}\_\{m\}\-y\_\{m\}\\right\|\(189\)
Results\.Table[5](https://arxiv.org/html/2608.28853#A5.T5)reports property\-wise MAE across the twelve QM9 targets\. Relative to EGNN, the strongest ESNN variant achieves lower error on nine of the twelve properties, including the polarizabilityα\\alpha, orbital gapΔϵ\\Delta\\epsilon, dipole momentμ\\mu, and several orbital\-energy and thermodynamic targets\. EGNN retains lower error onCvC\_\{v\},HH, and⟨R2⟩\\langle R^\{2\}\\rangle\. Among the ESNN transport families, Radial–Tangential Transport gives the best result on the majority of properties, indicating that the additional geometric transport remains useful even though the final molecular predictions are invariant scalars\. Several architectures developed specifically for molecular\-property prediction still achieve lower absolute errors on individual QM9 targets\. We therefore use this benchmark as a complementary test of the transport mechanism rather than as a molecular state\-of\-the\-art comparison: within the first\-order geometric setting represented by EGNN, matrix\-valued equivariant transport improves prediction across most of the evaluated properties\.
Table 5:*QM9 molecular\-property prediction\.*Mean absolute error \(MAE\) on twelve invariant molecular properties\. Baseline values and data\-partition annotations follow the comparison reported by[Aykent and Xia \(2025\)](https://arxiv.org/html/2608.28853#bib.bib6)\. A†\\daggerdenotes results obtained using different data partitions and should therefore not be interpreted as a strictly matched comparison\. Within the ESNN block, the best result for each property is shown inbold\. Target units follow the standard QM9 reporting convention shown in the second header row\. Lower is better\.Method𝜶\\bm\{\\alpha\}𝚫ϵ\\bm\{\\Delta\\epsilon\}ϵ𝐇𝐎𝐌𝐎\\bm\{\\epsilon\_\{\\mathrm\{HOMO\}\}\}ϵ𝐋𝐔𝐌𝐎\\bm\{\\epsilon\_\{\\mathrm\{LUMO\}\}\}𝝁\\bm\{\\mu\}𝑪𝒗\\bm\{C\_\{v\}\}𝑮\\bm\{G\}𝑯\\bm\{H\}⟨𝑹𝟐⟩\\bm\{\\langle R^\{2\}\\rangle\}𝑼\\bm\{U\}𝑼𝟎\\bm\{U\_\{0\}\}ZPVEUnitsma03\\mathrm\{m\}a\_\{0\}^\{3\}meVmeVmeVmDmcalmol−1K−1\\mathrm\{mcal\\,mol^\{\-1\}\\,K^\{\-1\}\}meVmeVma02\\mathrm\{m\}a\_\{0\}^\{2\}meVmeVmeV*Invariant models*Cormorant856134383826202196121222\.03NMP926943383040191718020201\.50DimeNet\+\+†\\dagger4432\.624\.619\.529\.7237\.566\.533316\.286\.321\.21ComENet†\\dagger4532\.423\.119\.824\.5227\.986\.862596\.826\.691\.20SphereNet†\\dagger4631\.122\.818\.924\.5227\.786\.332686\.366\.261\.12*Scalarization\-based models*ClofNet63533325402799610981\.23EGNN714829252931121210612111\.55PaiNN†\\dagger4545\.727\.620\.412\.0247\.355\.98665\.835\.851\.28LEFTNet48402418122376109761\.33EQGAT533220161124232438225252\.00ET5936\.120\.317\.511267\.626\.16336\.386\.151\.84Geoformer4033\.818\.415\.410226\.134\.39284\.414\.431\.28SaVeNet\-B†\\dagger3924\.818\.416\.39\.3236\.645\.43585\.485\.431\.18*Equivariant Sheaf Neural Networks*ESNN\-Id63\.0865\.3543\.1328\.5258\.8137\.8315\.7618\.71749\.44–10\.541\.54ESNN\-Diag62\.1643\.6927\.4421\.4317\.1531\.5711\.4812\.71134\.9911\.7713\.571\.62ESNN\-Ortho79\.4243\.3525\.1023\.1721\.7736\.5017\.7626\.07323\.5218\.5925\.251\.94ESNN\-RadTan60\.3239\.0125\.1619\.1716\.0831\.5714\.7713\.86119\.1110\.6310\.341\.45Similar Articles
Equivariant Cellular Sheaves for Molecular Electronic Structure: Bridging Sheaf Cohomology and E(3)-Equivariant Hamiltonian Learning
This paper introduces a framework linking sheaf cohomology with E(3)-equivariant neural networks for predicting molecular Hamiltonians, proposing Equivariant Cellular Sheaf Networks that generalize existing methods and provide topological insights.
Oversmoothing as Representation Degeneracy in Neural Sheaf Diffusion
This paper analyzes oversmoothing in Neural Sheaf Diffusion (NSD) as a representation degeneracy phenomenon using quiver theory and Geometric Invariant Theory. It proposes moment-map-inspired regularizers and explores non-uniform stalk dimensions to mitigate this issue in heterophilic graph benchmarks.
Group-Equivariant Poincar\'e Convolutional Networks
This paper proposes Equivariant Poincaré ResNets, combining hyperbolic geometry with discrete symmetry groups to improve efficiency in learning visual representations by treating rotated features as symmetric rather than distinct hierarchical concepts.
Symmetry in the Wild: The Role of Equivariance in Neural Fluid Surrogates
This paper investigates the role of group-equivariant architectures in neural fluid dynamics surrogates, introducing the AB-GATr model. It finds that equivariance is beneficial when data lacks strong alignment, but can degrade performance on highly aligned datasets.
Can SAEs Capture Neural Geometry? (6 minute read)
This article explores how sparse autoencoders (SAEs) can capture curved neural geometry, revealing three distinct ways SAE features represent manifolds, and presents an unsupervised pipeline to uncover geometric structure in neural representations.