TAGTorch: A PyTorch Library for Geometry, Topology, and Symmetry-Aware Machine Learning
摘要
TAGTorch is an open-source PyTorch library that unifies tools for topology, algebra, and geometry-aware machine learning, covering preprocessing, architectures, training techniques, and model analysis.
arXiv:2607.28755v1 Announce Type: new
Abstract: Over the last decade, neural networks have been applied to an increasingly diverse range of applications, including data with rich geometric, topological, or symmetry-related structure. As a result, researchers have increasingly drawn inspiration from topology, algebra, and geometry. Despite this rich algorithmic development, the supporting software ecosystem remains fragmented. Many important methods exist only as research prototypes in unmaintained repositories. We address this by introducing Topology, Algebra, and Geometry Torch (TAGTorch), an open-source, PyTorch-based library that unifies tools inspired by topology, algebra, and geometry, including data-preprocessing methods, architectures, training techniques, and model analysis tools. We describe the design philosophy of TAGTorch and then discuss its current architecture and capabilities, highlighting areas where it can fill gaps in the current software ecosystem. We conclude with a discussion of our future development priorities for the library.
查看缓存全文
缓存时间: 2026/08/03 07:32
# TAGTorch: A PyTorch Library for Geometry, Topology, and Symmetry-Aware Machine Learning
Source: [https://arxiv.org/html/2607.28755](https://arxiv.org/html/2607.28755)
\\theorembodyfont\\theoremheaderfont\\theorempostheader
:\\theoremsep \\jmlrvolume334\\jmlryear2026\\jmlrworkshopTopology, Algebra, and Geometry in Data Science
\\NameTegan Emerson\\Email \\addrPacific Northwest National LaboratoryUniversity of Texas at El Paso\\NameGregory Roek\\Email \\addrPacific Northwest National Laboratory\\NameEmilie Purvine\\Email \\addrPacific Northwest National Laboratory\\NameHenry Kvinge\\Emailhenry\.kvinge@pnnl\.gov \\addrUniversity of WashingtonPacific Northwest National Laboratory
###### Abstract
Over the last decade, neural networks have been applied to an increasingly diverse range of applications, including data with rich geometric, topological, or symmetry\-related structure\. As a result, researchers have increasingly drawn inspiration from topology, algebra, and geometry\. Despite this rich algorithmic development, the supporting software ecosystem remains fragmented\. Many important methods exist only as research prototypes in unmaintained repositories\. We address this by introducing*Topology, Algebra, and Geometry Torch \(TAGTorch\)*, an open\-source, PyTorch\-based library that unifies tools inspired by topology, algebra, and geometry, including data\-preprocessing methods, architectures, training techniques, and model analysis tools\. We describe the design philosophy ofTAGTorchand then discuss its current architecture and capabilities, highlighting areas where it can fill gaps in the current software ecosystem\. We conclude with a discussion of our future development priorities for the library\.
###### keywords:
Software, Equivariance, TAG algorithms, Topological data analysis, Geometric deep learning
## 1Introduction
![[Uncaptioned image]](https://arxiv.org/html/2607.28755v1/tagtorch_logo.png)The rapid improvement of neural network performance on tasks in natural language and computer vision has triggered an explosion in applications of deep learning to more complex scientific settings\. Many of these feature data with complex structure, symmetries, or geometry motivating domain\-aware models that benefit from specialized mathematical priors\. These can range from group equivariant layer types to input features extracted via topological data analysis\. At the same time, the science of deep learning has advanced considerably in the last decade with complex analytical tools now available to understand the behavior and learning dynamics of neural networks\. As model internals and the learning process they evolve from are intrinsically mathematical, this field of research has also borrowed heavily from mathematics\. Examples include the use of intrinsic dimension statistics to understand the process of learning\(ansuini2019intrinsic;joshi2026geometry\), weight symmetries to better understand loss landscape structure\(simsek2021geometry;ainsworth2022git;godfrey2022symmetries\), and geometric notions of feature structure such as superposition\(elhage2022toy\)and the linear representation hypothesis\(park2023linear\)\.
This transfer of methods and ideas from mathematics to machine learning has been valuable to both communities\. Machine learning has received an influx of new ways of looking at models and data and mathematics has been exposed to a host of problems that have the potential to motivate new research directions\. But mathematics differs from machine learning in some important ways, one being the role of software\. Software remains the bedrock of progress in machine learning\. Even work that is primarily theory\-based often includes experiments as a way to validate results\. Because of this, the success of a machine learning field depends on the software that is available to support it\.
What software is available for a researcher interested in using mathematically\-inspired tools in AI development? The answer varies considerably based on what area one focuses on\. There are some mature, well\-maintained libraries with large user\-bases\. For example, thee3nnlibrary captures a range of methods for buildingSO\(3\)SO\(3\)\-equivariant architectures for molecular modeling\(geiger2022e3nn\)\. Another example isPyTorch Geometric, a package aimed at building graph neural networks for learning on structured data\(fey2019fast\)\. While these packages are often effective within the scope that they support they can be hard to work with outside their intended uses\. Considerable work might be required to extract a base method \(e\.g\., an equivariance layer construction\) from a domain\-specific set up \(e\.g\., molecular modeling\)\. This can make it hard to leverage these tools for applications with novel data types or tasks\. As another example, chaining together multiple tools in a single workflow \(e\.g\., topological data analysis features with an equivariant architecture\) usually requires using several different libraries, leading to potential dependency conflicts\.
To address this, we introduce*Topology, Algebra, and Geometry Torch \(TAGTorch\)*111[https://github\.com/hkvinge/tagtorch](https://github.com/hkvinge/tagtorch), a PyTorch\-based library that provides implementations of mathematically\-inspired methods in deep learning, particularly those based on the fields of topology, algebra, and geometry\. The motivating goals ofTAGTorchare the following\. \(1\)TAGTorchis a single entry point to a diverse assortment of TAG\-related methods while remaining maximally domain agnostic to accommodate new applications\. \(2\)TAGTorchis designed to minimize the barrier to entry for non\-mathematicians\. \(3\)TAGTorchhas a mathematically\-informed foundation and structure making it extensible to future developments in the field\. \(4\) Where strong and mature libraries exist,TAGTorchuses these rather than replicating capabilities\.
In this paper we discuss these and other design goals ofTAGTorch, including our strategy of prioritizing one method over another for integration\. We also describe our plans to include a strong visualization component to help users without an extensive background build robust intuition for the methods\. Finally, we give a brief overview of some of the capabilities that currently exist inTAGTorch\. SinceTAGTorchis still in the early stages of development, we also outline priority features that we are actively planning to integrate in the future\.
In summary, our contributions include the following\.
- •We introduceTAGTorch, a PyTorch\-based library that makes mathematical tools for AI development accessible to users beyond experts in machine learning and mathematics\.
- •We describe the design principles that informTAGTorchdevelopment\.
- •Finally, we describe the features currently inTAGTorch, as well as features that will be added in the near future\.
## 2Related work
Ideas from the fields of topology, algebra, and geometry have spurred a range of fruitful research directions in deep learning\. Many of these are supported by software libraries that enable efficient experimentation by the community\. Examples includee3nnwhich is specialized toO\(3\)O\(3\),SO\(3\)SO\(3\), andE\(3\)E\(3\)symmetries of points inℝ3\\mathbb\{R\}^\{3\}\(geiger2022e3nn\),escnnwhich captures planar symmetries on grids viaE\(n\)E\(n\)\-equivariant steerable CNNs\(cesa2022a;e2cnn\)\. Also notable iscuEquivariancewhich provides low\-level tensor operations to support equivariant layers\(nvidia\-cuequivariance\)\. Beyond equivariance specifically,PyTorch\-Geometric \(PyG\)is one of the most prominent libraries for geometric deep learning, focusing on learning on graphs\(fey2019fast\)\. Other geometry specific libraries includeGeomstatswhich provides support for several core geometric constructions\(JMLR:v21:19\-027;miolane2020introduction;guigui2023introduction;le2023parametric;pereira2025learning\)andGeooptwhich provides manifold optimizers\(geoopt2020kochurov\)\.
The field of topological data analysis \(TDA\) and topological deep learning has its own ecosystem\. Of thesescikit\-tdais most similar toTAGTorch\(scikittda2019\)\. It primarily provides wrappers for other libraries, exposing lots of different packages to users in a convenient format\.GUDHIis a well\-established, open\-source library\(maria2014gudhi\)\.Topology Toolkit \(TTK\)contains TDA tools for data in dimensions 2 and 3\(tierny2017topology\)\. Finally, Teaspoon is a TDA package for studying dynamics and time\-series data\(Khasawneh2025\)\.
Lastly, there is an emerging ecosystem of interpretability libraries\. Among the most prominent of these isTransformerLens, which contains a range of tools for analysis of transformer\-based language models drawn from mechanistic interpretability\(nanda2022transformerlens\)\. Other important libraries includeNNsight\(fiotto2025nnsight\)which aims to enable analysis of model internals at a large scale andSAELens\(bloom2024saetrainingcodebase\)which is primarily aimed at providing support for sparse autoencoders on language models\.
## 3Design Goals
TAGTorchaims to eventually organize and centralize existing tools from the fields of geometric deep learning, topological deep learning and topological data analysis, representation geometry, model symmetries, loss landscape geometry, and mathematical interpretability techniques\. We particularly hope to implement missing links and engineer high\-level flows to reduce the barrier to entry into the world of TAG and deep learning\. Our strategy blends an appreciation for the wide array of exciting research in the space, as well as closing gaps and building bridges between communities and sub\-fields\. By building this package, we aim to start a community\-wide initiative for contributing new methods to a centralized project, which will amplify individual research efforts and ease direct comparisons, evaluations, and applications of methods in the TAG toolbox\.
Figure 1:A diagram of the current structure ofTAGTorch\. This will be a starting point upon which we will build out specific capabilities\.### 3\.1Organizing and centralizing existing tools
As a community, there are many scattered implementations of algorithms and specialized tools that all operate on universal mathematical constructions\. However, reconciling these implementations runs into the issue that there is significant heterogeneity in verbiage and formalisms used across software packages\. Some of this comes from the fact that the implementations originate from different disciplines \(e\.g\., mathematics vs\. physics\)\. Some arises from the applications that the implementations are geared toward\. In either case, the consequence is it can take some work to translate between tools, even when they operate within an identical mathematical framework\.
To make this concrete, consider the example of representation theory, which serves as the foundation for equivariant neural networks\. Each group has a well\-defined set of isomorphism classes of irreducible linear representations\(serre1977linear;fulton2013representation\)\. These serve as universal building blocks for all symmetries of the group\. They also provide a unifying grammar for understanding and comparing different approaches to equivariant machine learning that leverage that group\. However, it is very hard to see these connections at the level of software\.
Beyond the fact that the current ecosystem fails to elevate structure that could help unify different implementations, the field also suffers from maladies common to any growing research area\. This includes situations where the only implementations of algorithms are often one\-off, unmaintained research repositories\. These sometimes allow replication, but they usually lack the engineering necessary to provide a general capability to a larger community\.
In this space,TAGTorchaims to amplify existing works and contextualize them within a common framework\. This will involve:
- •Implementing methods within a well\-engineered software architecture, which standardizes access points to methods, uses abstractions and object\-oriented programming, and adds extensive testing for robustness and improvements to efficiency\.
- •Defining abstractions such that the myriad formalisms used across subfields and software packages can be interacted with in a standardized way \(e\.g\., a universal language for irreducible representations\)\.
- •Connecting methods to the PyTorch ecosystem\. This unifying backend brings methods into contact with sophisticated training capabilities and optimizations, evaluation utilities, and methods from other deep learning disciplines such as computer vision or natural language processing\.
### 3\.2Implementation of New Methods
In cases where methods lack accepted implementations from the community, our goal will be to add them toTAGTorch\. We will aim to do this in a way that aligns them with competing approaches\. For example, when a new tool for analyzing representation geometry appears, it will be added alongside existing representation geometry tools aligning available hyperparameters, assumptions on input data, etc\. as much as reasonably possible\. By aiming for implementation alignment, we hope to make it easy to compare competing methods without the friction of having to integrate a new code base with different assumptions and formalisms\.
### 3\.3Reducing the Barrier to Entry
Centralizing and rounding out the TAG toolbox only goes so far in terms of making a tool that can be of use to a wider community\. We aim forTAGTorchto occupy a shared space between fundamental and applied science, connecting real data and use cases to researchers in deep learning and mathematics\. This is a symbiotic relationship: fundamental science often develops in a closed community, leveraging data that is useful for innovation but untethered from the domain sciences which are, in theory, benefiting from this novel research; meanwhile, domain scientists may not be aware of relevant methods coming from the machine learning community\. This divergence is compounded by the lack of unifying software\.
To reduce the barrier to entry for domain scientists into the world of TAG\-flavored AI tools, and in turn drive AI research with real applications,TAGTorchprovides high\-level workflows that help “connect the dots” between distinct phases in the research cycle, from data processing and transformation, visualization, building mathematically\-aware architectures, and training and evaluating models informed by the geometry of data\.
### 3\.4Extensibility
Ultimately the success ofTAGTorchwill largely depend on community engagement, both in terms of feedback on the package, but also direct contributions\. We have made two major design choices to try and encourage this\. First, we have implemented object\-oriented design principles in various modules in the package\. With this, contributors can add new functionality to the package with minimal changes required to the codebase, adhering to the encapsulation principle of software architecture\. Designed class hierarchies inTAGTorchinclude: data loaders and processors, which currently are wrappers around other libraries’ data loading capabilities; torch data transforms using methods from topological and symmetry\-based analysis; equivariant torch architectures; and implementations of groups and group actions for symmetry analysis\. Each of these can be integrated into stable, package\-level flows that implement multiple steps in analysis, modeling, and evaluation\.
## 4Design and Architecture
### 4\.1A standardized API for data transforms
Transforms are central in modern mathematics\. As category theory tells us, we should focus on the morphisms, not the objects\. Transforms are also central to the world of traditional deep learning and PyTorch, deep learning essentially being the process of composing many simple transformations together\. Beyond that, familiar transforms like augmentation, mutation, and sampling are integrated into workflows in popular libraries, including ‘torchvision‘\(marcel2010torchvision\)and ‘torch\-geometric‘\(fey2019fast\)\. PyTorch\-compatible transform layers can be directly integrated into data loaders and training and evaluation processes\.
InTAGTorch, we use the notion of a transform to incorporate many central capabilities within a single framework\. Currently, the library includes topological transforms and data symmetries, but we will add weight symmetries and representation statistics as well\. Workflows in TAG require non\-trivial transformations among data types, particularly when working with graphical data\.TAGTorchcombines native implementations of common graph utilities with imports of standard libraries\.
### 4\.2Visualization and interactive utilities
It is non\-trivial to interpret data and train models without strong intuition\. This is particularly true when techniques are built on methods from graduate level mathematics\. For instance, equivariant architectures are generally harder to interpret than their non\-equivariant counterparts, especially if one does not have some basic intuition about representation theory\.
In some cases, specific types of visualizations can be leveraged to aid in intuition building\. InTAGTorch, some methods will be supported by custom visualization procedures\. One feature that is currently in development is a visualization pipeline that helps a user understand parameter choice in persistence homology \(Figure[2](https://arxiv.org/html/2607.28755#A1.F2)in the Appendix\)\. In this particular instance, the user can generate visualizations that show the impact of different hyperparameter choices \(specifically density floor andϵ\\epsilon\-net which filter out points and maximum distance and maximum number of neighbors which filter out edges\)\.
## 5Capabilities
In this section we outlineTAGTorch’s current functionality as well as functionality in development\.
### 5\.1Symmetry and Equivariance
We have taken a mathematics\-first approach to building symmetry and equivariance capabilities inTAGTorch\. This means starting with groups, then defining actions and linear representations of these groups on specific datatypes, and then moving to downstream functionality\. By building up from algebraic foundations, we make the package more extensible\. We have protocols that define a class interface to build a new action or representation for an existing group\. These are all currently in thetagtorch\.symmetrymodule\. This framework also helps us to connect different packages that use the same underlying groups and representations but use different syntax to invoke them\. For example,e3nnandescnnboth utilize the same underlying mathematics, but you would not know this from looking at the way the code is set up\. We hope to smooth out these differences\.
Beyond these foundations, we have started by including basic functionality such as incorporating symmetry actions into a dataloader for convenient augmentation\. We have also elevated some of the best resources for equivariant architectures, includinge3nn\(geiger2022e3nn\)andescnn\(e2cnn;cesa2022a\), making these easily callable from withinTAGTorch\. Other equivariant architectures that we consider essential are those associated with sets, such as DeepSets\(zaheer2017deep\)and Set Transformer\(lee2019set\)\. We plan to also include other prominent methods in equivariant learning that are not architectural, such as frame averaging methods\(puny2021frame;ma2024canonicalization\)and methods which promote a softer inductive bias toward equivariance\(finzi2021residual\)\.
### 5\.2Topology
Our ambition is to include two types of topology\-inspired tools inTAGTorch\. The first are tools coming from topological data analysis\. These involve extraction of topological statistics from data either for input features to a downstream neural network or as a tool for analyzing training or inference data\.TAGTorchalready includes the Euler characteristic transform \(ECT\)\(turner2014persistent\)and basic persistent homology capabilities\. We aim to eventually include skeletonization via Morse theory approaches\(gyulassy2008efficient\), Mapper\(singh2007topological\), standard topological vectorizations, zig\-zag persistence\(carlsson2010zigzag\), and sheaf\-theoretic methods\(curry2014sheaves\)\. In the longer term we also plan to include support for topology\-aware layers\(papamarkou2024position\)which is a different flavor of topology\-inspired tool\.
As noted in Section[2](https://arxiv.org/html/2607.28755#S2), there are several mature Python\-based software packages for TDA\. Our goal withTAGTorchis to utilize these where possible, including extra functionality where it will make integration into PyTorch\-based workflows easier\. In certain cases, where no maintained tool is available, we plan to provide our own implementations\.
Beyond the algorithms themselves, we also provide resources for hyperparameter optimization\. This includes identifying the specific parameters most relevant to a given dataset, maximizing the information extracted under a restricted compute budget, and rapidly visualising simplicial complexes for qualitative analysis/validation \(identifying data points excluded as outliers, included/excluded edges and triangles, etc\.\)\.
### 5\.3Model Properties
One fundamental area where development has been limited is ourTAGTorch\.model\_propertiesmodules which will include mathematical tools for better understanding neural networks\. This will include functions to estimate intrinsic dimension\(camastra2016intrinsic\), functions for analyzing representation geometry\(lin2024topology\), functions for representation similarity comparison\(klabunde2025similarity\), and functions to measure other mathematical properties of a model\. At present we have a function for calculating the group empirical equivariance deviation\(kvinge2022ways\)which measures the extent to which a neural network is equivariant\.
## 6Conclusion
In addition to providing new, cross\-disciplinary capabilities,TAGTorchis a vehicle for interdisciplinary cohesion and the sharing of methods, utilities, and common practices\. Much as other disciplines in machine learning have benefited from centralized software and model sharing software \(e\.g\., torch hub222[https://pytorch\.org/hub/](https://pytorch.org/hub/), HuggingFace333[https://huggingface\.co/](https://huggingface.co/), or torchvision444[https://docs\.pytorch\.org/vision/stable/index\.html](https://docs.pytorch.org/vision/stable/index.html)\), the communities focused on developing and applying TAG methods will benefit from having comprehensive method implementations integrated into the PyTorch ecosystem, continuously updated by the package maintainers and supported by contributions from the community\.
TAGTorchis in its early stages, and will continue evolving in terms of content and user flows but maintain the initial structures and architecture described here\. These structures will create simple, accessible means for community members to contribute to the package, taking advantage of abstractions to add support for new models, data structures and transforms, tools for model property analysis, and visual tools\. Completed implementations are highlighted by topological and symmetry\-based transforms, a unifying implementation and formalism for implementing symmetries in PyTorch, interactive tools for visualizing data and applying persistent homology, and wrapper functions and classes around existing python libraries for equivariant architectures\.
At the same time,TAGTorchis actively being fleshed out beyond the initial architecture and structural design\. Active areas of development include: \(1\) implementations of new equivariant architectures beyond the wrapper functionality we provide fore3nnandescnn; \(2\) interactive Graph User Interfaces \(GUI\) for managing human\-in\-the\-loop workflows, such as hyperparameter tuning; and \(3\) specific flows for data loading and processing and model development\. By open\-sourcingTAGTorch, we invite the community to contribute to these and other areas in the package, adding diversity to the present implementations and providing a platform for others to use and extend their work\.
## References
## Appendix AWorkflows and Interactive Applications
#### A\.0\.1Interactive widgets for hyperparameter optimization
In many cases, data visualization provides valuable signal that can help a model developer evaluate goodness of fit, guide the design of data augmentation or transformation, and model design\. InTAGTorch, we are developing a suite of end\-to\-end methods for tying data visualization to model development\. In particular, the initial release ofTAGTorchfeatures an interactive Graphical User Interface for selecting hyperparameters when applying persistent homology\.
Figure 2:TAGTorch enables Human\-in\-the\-Loop \(HIL\) workflows that are necessary to tune the process of topological data transforms\.
#### A\.0\.2Building equivariant neural network models
There is a diverse collection of strategies for building equivariance into neural networks, spanning data transformation and augmentation to designing structurally equivariant architectures\.TAGTorchis designed to provide a unified access point to tools and model definitions in other packages in the community and implements other architectures\. But defining networks is only half the battle — significant domain knowledge and practical expertise, gained through years of familiarity with the research topic and underlying mathematics, make configuring models complex, often requiring experimentation and matching model design and specification to data types and modeling tasks\.
As it matures,TAGTorchwill support model specification through serialized configs \(e\.g\., JSON or YAML files\), defining model sizes, layer types, and preprocessing\. Doing so will allow internal and community contributors to add configurations that define “recipes” for building networks specifically suited to certain data modalities and modeling objectives\.
The other half of TAG and deep learning is training models\. As with all model training, this requires expertise with hyperparameter tuning, tricks of the trade, and knowing which metrics and indicators of a well\-trained model\. Training workflows inTAGTorchwill integrate with established training libraries and utilities in PyTorch, as well as specialized hyperparameter tuning software, pairing equivariant architectures and training procedures with traditional deep learning infrastructure\.相似文章
TorchKM:面向GPU的核学习与模型选择库
TorchKM是一个开源的GPU加速核机器库(支持向量机、核逻辑回归等),采用scikit-learn风格的API。通过重用矩阵运算加速训练和模型选择,相比标准基线实现了显著的加速比。
Transformer Geometry Observatory TGO-IV:发展拓扑观测站
本文介绍了TGO-IV,一个使用持续同调来分析Transformer表示如何跨层演化的拓扑框架,补充了先前的谱和几何观测站。
TorchMorph:CUDA加速的形态学变换
TorchMorph是一个CUDA加速的PyTorch扩展,提供GPU优化的形态学和距离变换算子,与SciPy等CPU实现相比,实现了显著的加速,并保持了兼容的API。
PyTorch 生态圈
PyTorch 生态系统的概述,涵盖关键工具、库和社区资源。
@Alacritic_Super: 以为 PyTorch 只是一个库?再想想。以下是每个机器学习工程师都应该了解的 PyTorch 生态系统:PyTorch – C…
本文概述了 PyTorch 生态系统,重点介绍了 PyTorch、LibTorch、ExecuTorch、TorchServe 等库和工具,适用于各种 AI/ML 任务。